Benchmark Methodology and Corrections, September 2026

How We Benchmark LLM Inference Honestly,and the Four Conclusions We Retracted

Short answer

On 24 September 2026 we put a reviewer on our own GPU benchmark corpus with one instruction: find every way these numbers could mislead a buyer. The audit covered 47 result files measured between 21 and 24 September 2026, the concurrency harness that produced them, the launcher scripts, and a page we had staged for publication. It came back with a ranked list of 16 biases and a plain verdict: four of the conclusions on that staged page were wrong, and three of them were wrong in a way that reversed their direction. All four were pulled before the page shipped, and a fifth figure had already been removed earlier in the same week. The largest single cause was comparing configurations at a fixed total stream count. At 32 total streams a GB10 pair carried 16 streams per unit while an eight-unit group carried 4, which manufactured a per-unit decline of 150.0, 106.7 and 67.2 tokens per second. At a matched 4 streams per unit the same three result files give 70.4, 70.9 and 67.2 tokens per second per unit, which is essentially flat.

Every figure below names the result file it came from and the caveat that limits it. Where the corpus cannot support a claim, this page says so and names the measurement Petronella Technology Group, Inc. owes before it will make the claim.

Standing note on every decode figure

Every decode figure on this page is prose decode: prompts of 39 to 42 tokens, a 256 token output cap, thinking off, one distinct prompt per request so that prefix caching cannot flatter the result. Decode content class is stated because it changes the number. On one of our own models measured in August 2026, decode ran at 339 tokens per second on mathematics prompts, 246 on prose and 233 on code, on identical hardware.

The reviewer had not run any of the benchmarks. No model was launched and no engine was restarted for the audit. The only live reads were power telemetry and process listings.

Why this page exists

This page is the method that came out of that audit. It is written for a technical reader who wants to check the arithmetic rather than trust the conclusion. Every number below carries its source file and its caveat. Where the corpus cannot support a claim, the page says so and names the measurement we owe before we will make it.

Petronella Technology Group, Inc. sells the hardware and the hosting that these benchmarks describe. That is a conflict of interest, so the only defence is method: publish the corrections, publish the losing configurations, publish the flags, and make the numbers checkable. Our results index lives at the fleet benchmark hub, and the per-model scores sit on the open-weight model leaderboard.

The hardware in question is a small research fleet. Eight GB10 units, four of them MSI EdgeXpert MS-C931 boxes and four of them NVIDIA DGX Spark systems, all sm_121 with 121.7 GiB of unified memory each, cabled to a 200 gigabit RoCE fabric and runnable as single units, pairs, four-way groups or one eight-way tensor-parallel group. One NVIDIA DGX Station GB300. One workstation with four RTX PRO 6000 Blackwell cards. One production host with six H200 NVL cards, which is shared with live customer traffic and therefore mostly off limits.

That is an unusually direct way to answer the question buyers actually ask, which is not "how fast is this card" but "how many of my users can this box serve at a speed they will tolerate, and what does the next box add". The risk is equally direct. When one team owns the hardware, writes the harness, picks the workload and publishes the conclusion, every convenient shortcut points the same way.

The audit found six shortcuts that mattered. One of them alone was enough to invert two headline conclusions.

The four conclusions we retracted

Each row is a statement that was written, reviewed and pulled before it reached the public site. The third column is what the corpus supports today, which in one case is nothing at all.

Four retracted conclusions from a staged benchmark page, with the replacement each one does or does not have. Reviewed 24 September 2026.
Staged conclusionWhat is wrong with itWhat the corpus supports today
"Throughput per unit fell from 150 to 107 to 67 tok/s, so every added unit adds less"Measured at a fixed 32 total streams, so a pair carried 16 streams per unit while eight units carried 4 each. The decline is an artifact of the load setting.At a matched 4 streams per unit, per-unit throughput is 70.4, 70.9 and 67.2 tok/s. Essentially flat.
"Four independent pairs deliver about 1,200 tok/s, 2.2 times the 537.6 of the eight-way cluster"Compares 128 users at 10.5 tok/s each against 32 users at 20.1 tok/s each.At 32 users on both sides: 563 against 537.6, a ratio of 1.05. A tie.
"The GB300 is roughly at parity with GB10 pairs on throughput per dollar"Holds only at 32 streams, where the GB300 was nowhere near saturated. Its engine allowed 64 sequences and it was still climbing at 48.Retracted and not replaced. The replacement number is blocked behind the controlled runs described below.
"The eight-unit cluster was the worst value measured in every scenario we priced"Same fixed-stream basis. The cluster and the pair land close together once both serve users at the same speed.At a 20 tok/s per-user floor the pair and the eight-way group are within a few percent of each other on throughput per thousand dollars of hardware.

A fifth figure, an 8.6 times ratio between four H200 cards and the eight-unit GB10 cluster, was removed before the audit. It combined four separate mismatches in one division: FP8 against NVFP4, five-token speculative decoding against none, an engine sequence cap of at least 96 against a cap of 32, and four GPUs against eight complete systems. Measured against the better of the two GB10 runs, the same arithmetic produces 7.0 instead of 8.6, which is a clear sign that the ratio was measuring our choices rather than the hardware.

The same pattern appears in an earlier post on our own blog, where a pooled measurement across six units at 18 streams is compared with a four-unit endpoint at 6 streams. That comparison is a statement about aggregate capacity, not about efficiency per unit, and it is being restated on a matched basis.

Bias one: the fixed-concurrency fallacy

This is a common error in published multi-node GB10 work, including our own, and it is worth spelling out because it looks like the careful thing to do.

You are comparing a pair of units, a four-way group and an eight-way group. You want a fair comparison, so you hold the workload constant: 32 concurrent streams against every configuration. The number of units changes and the load does not. That feels controlled.

It is not. Thirty-two streams across two units is 16 streams per unit. Thirty-two streams across eight units is 4 streams per unit. A pair is being pushed four times harder per box than the cluster. Batched inference gets more aggregate throughput out of a device the more requests it has in flight, up to its scheduler limit, so the pair is measured deep into its efficient region while the cluster is measured barely off idle. Divide the result by unit count and you produce a smooth, convincing, entirely manufactured curve of diminishing returns.

Here is our own data, all three configurations measured on 23 September 2026 with the same harness, the same container image, the same NVFP4 checkpoint of Qwen3.8-Flash-Next, the same engine arguments and the same 39 to 42 token prompts with a 256 token cap. Prose decode.

The same three GB10 configurations read two ways. Aggregate and per-unit tokens per second, prose decode, measured 23 September 2026. Source files: 2026-09-23-gb10-tp2-spark34-qwen38-flash-next.json, 2026-09-23-gb10-tp4-msi1234-qwen38-flash-next.json, 2026-09-23-gb10-tp8-qwen38-flash-next.json.
Comparison basisPair, 2 unitsFour unitsEight units
Aggregate at 32 total streams299.9426.6537.6
Per unit at 32 total streams150.0106.767.2
Streams per unit at that point1684
Aggregate at 4 streams per unit140.8 at 8 streams283.6 at 16 streams537.6 at 32 streams
Per unit at 4 streams per unit70.470.967.2
Per-user median rate at those points21.722.720.1

Read the last two rows. At a matched load per unit and a matched per-user speed, a GB10 contributes about 70 tokens per second of prose decode whether it sits in a pair or in an eight-way group. The 150 to 107 to 67 decline does not exist as a property of the hardware. It is the shape of the batching curve, sampled at three different depths.

What the larger group genuinely buys is visible in the same files, and it is not throughput per unit. Single-stream decode rises from 33.1 to 44.4 to 47.5 tokens per second per user across two, four and eight units. That is a latency improvement for one user, and it flattens hard after four units: two to four units bought 34 percent, four to eight bought 7 percent. The other thing the larger group buys is capacity, which is binary. A 126 GB checkpoint does not load on one 121.7 GiB unit at all. The eight-way group holds a 433 GiB model that no smaller configuration on our fleet can hold. "Does not fit" is a benchmark result, and we publish it as one.

The generalisable rule is short. When the number of devices is a variable, total concurrency cannot be the constant. Hold streams per device constant, or better, hold the delivered per-user speed constant and let each configuration find its own stream count.

Bias two: compare at a matched per-user speed floor

The stronger basis is the one MLCommons uses for MLPerf Inference and that SemiAnalysis uses for its InferenceX Pareto charts: define the service level first, then ask how much work each configuration does while meeting it. Nobody buys throughput. They buy a number of users at a speed those users will accept, with a first-token latency they will accept.

Set the floor at roughly 20 tokens per second per user, which is a readable streaming rate for chat, and take each configuration at the highest measured concurrency that still clears it.

The same GB10 configurations at a matched per-user speed floor of roughly 20 tokens per second. Prose decode, measured 23 September 2026, same three source files as the table above.
ConfigurationAggregate at the floorStreamsPer-user median
Pair, 2 GB10 units140.8821.7
Four GB10 units283.61622.7
Eight GB10 units537.63220.1

Now the staged claim about four pairs beating the cluster can be checked. Four independent pairs, each serving 8 streams, serve 32 users in total and deliver 4 times 140.8, which is 563 tokens per second. The eight-way cluster serves the same 32 users at 537.6 tokens per second. The ratio is 1.05. That is a tie inside the run-to-run variance discussed in the next section, not a 2.2 times win.

The 2.2 figure came from running each pair at 32 streams, which is 128 users in total, each receiving 10.5 tokens per second. Users at 10.5 tokens per second are not being served the same product as users at 20.1 tokens per second. Raise the floor to about 30 tokens per second and the ordering shifts again: four pairs give 4 times 98.9, which is 396, against the cluster's 341.2, while two four-unit groups give 2 times 220.3, which is 441. The layout that wins depends on the service level, which is exactly why a single number is the wrong output and a curve is the right one.

Two consequences follow for anyone running the same comparison. First, publish the per-user rate next to every aggregate figure, because the aggregate alone cannot be interpreted. Second, ladder every configuration to saturation. Our GB10 engines were capped at 32 sequences, the GB300 at 64, and the production H200 service at 96 or more. A "peak aggregate" column that mixes those caps is comparing engine configuration, not hardware. The GB300 gained 30 percent going from 32 to 48 streams on Qwen3.8-Flash-Next, and 55 percent going from 32 to 64 streams on Qwen3.8-27B in NVFP4. Any peak we had quoted at 32 was an understatement of that platform and we had no way of knowing by how much until the ladder was extended.

Bias three: run-to-run variance, and why one ladder is not publishable

The audit's most uncomfortable finding is how unstable a single-shot ladder is on this fleet.

GLM-5.3-Flash on the eight-unit group, same model, same layout, two runs one day apart: 257.8 tokens per second at 32 streams on 22 September and 316.7 on 23 September. That is a 23 percent swing. At 16 streams the same pair of runs gives 247.4 and 148.8, a 66 percent swing, and the direction flips between the two levels. Single-stream decode, meanwhile, was stable to within a percent: 33.1 and 33.3. A model or fabric problem would not behave like that. Scheduler stalls would, and the first-token traces say so directly: those GLM levels carry median first-token latencies of 6, 12 and 21 seconds, which land inside a level that only lasts 15 to 30 seconds in total.

That is a harness problem, not just a reporting problem. Our aggregate metric was total tokens divided by wall time, where wall time runs from the first thread starting to the last thread finishing. With a 256-token output cap, a level completes in 5 to 45 seconds, so one 12-second stall removes a large fraction of the level's apparent throughput. The same defect explains the non-monotonic dips that had been published without comment, such as 100.6 tokens per second at 4 streams sitting above 57.7 at 8 streams in the same ladder.

It is not confined to the fragile GB10 lane. The GB300 was measured twice on the same afternoon with the same NVFP4 checkpoint. At 32 streams the two runs agree to within 0.03 percent: 3,553.3 and 3,552.3. At 8 streams they differ by 31 percent: 740.4 against 1,080.1, and the low run carries a 1.06 second median first-token latency where the high run carries 0.18. One outlier level per ladder, on the most stable platform we own.

So these are the rules we now apply before a figure is eligible for a public page.

  • At least three full ladders per configuration, and five for anything on a page. At least one repeat must follow a full engine restart, because the largest swing we have seen sits between launches, not within a session.
  • Randomise level order inside each repeat, so thermal state and drift do not line up with concurrency.
  • Report median, minimum and maximum, the number of runs, and the coefficient of variation. Any level above 10 percent variation is investigated before it is quoted, not after.
  • Record first-token latency at the 50th and 95th percentile per level, and flag any level whose 95th percentile exceeds five times its median. Our 48-stream rows are a worked example: the eight-unit group shows a 0.63 second median against a 14.0 second 95th percentile, and the pair shows 0.86 against 26.9. Those levels are past the knee and a single median would hide it completely.
  • Replace burst runs with steady-state windows. Keep N requests in flight for at least 120 seconds, count only tokens produced inside the window after a 20 second ramp, and keep the old burst number as a separate cold-start metric rather than the headline.

With two runs there is no interval worth quoting. The honest statement for GLM-5.3-Flash at 32 streams on eight units is "two runs, 258 and 317, not yet resolved", and that is what we publish.

Bias four: speculative decoding and engine mode, the unrecorded confound

Speculative decoding drafts several tokens at once and verifies them in a single forward pass. When acceptance is good it is close to free throughput, and the gain is largest at batch size one and shrinks as the batch grows. Eager mode, meaning CUDA graphs disabled, is the opposite: a flat tax on every step.

Neither setting was recorded in most of our result files, and the two were not consistent across the corpus. The GB10 launcher script passes a four-token speculative budget to its container image, which builds its own serve arguments from the environment. A container inspection recorded earlier in the week for the same image shows both that speculative configuration and an explicit eager flag. None of the pair, four-way or eight-way result files mention either. The GB300 file records its image but not its speculative configuration. Separately, our H200 service runs with a five-token budget and the RTX PRO 6000 runs at two.

The size of the effect is not hypothetical. On our own four-card RTX PRO 6000 workstation, eager mode cost a factor of 2.4 on decode throughput, 22.6 tokens per second against 55. A single unrecorded eager flag can therefore dominate any cross-platform ratio and leave no trace in the result file. In our GB10 against GB300 comparison, the two unrecorded settings push in opposite directions: the speculative budget helps the GB10 side and eager mode hurts it. We do not know the net, which means we do not know the ratio.

The correction is procedural rather than clever. Read the serve arguments from the running process or the container, never retype them. Record the speculative configuration, the graph mode, the key-value cache dtype, the sequence cap, the batched-token cap and the prefix-cache setting in the result file itself. Then run the comparison four ways: speculation off on both sides, on on both sides, and each platform at its own preferred setting. Publish the off-against-off ratio as the hardware comparison and the best-against-best ratio as the deployment comparison, clearly labelled as two different questions.

One further rule, borrowed from the published literature and from MLPerf's own restriction of speculative decoding to a single scenario: never measure a speculative speedup on random-token prompts. Acceptance rates depend on the text being predictable. Synthetic filler will give a number that no real workload reproduces.

Bias five: precision is a choice, not a property of the hardware

Our GB10 numbers ran an NVFP4 checkpoint. Our GB300 numbers ran NVFP4. Our H200 numbers ran FP8, because that generation has no FP4 path at all. Three platforms, three answers to a question that has nothing to do with which box is faster.

Two things go wrong at once. The obvious one is bytes: a four-bit weight format reads roughly half the memory of an eight-bit one per token, and decode is bandwidth bound, so the lower precision gets a speed advantage that belongs to the format rather than the silicon. The subtler one is that four-bit does not mean four-bit. On sm_121 the published vLLM builds route NVFP4 mixture-of-experts layers through a Marlin kernel that dequantises before compute, so the hardware never executes a native four-bit tensor operation. Our own result files carry that caveat explicitly: the numbers are a floor for that hardware, not a ceiling.

The public record says the same thing more sharply. A vLLM issue reports that published wheels and images send NVFP4 mixture-of-experts work to Marlin on sm_121 while builds compiled for that architecture select the native path, and a second issue shows a four-bit dense layer picking a W4A16 kernel and losing 31 percent of prefill throughput, 2,986 tokens per second down to 2,041, with a named workaround. An independent write-up measured plain FP8 beating Marlin-routed NVFP4 by 32 percent on a single unit, 53.8 against 40.8 tokens per second. The claim that this generation "lacks native FP4" was retracted by the person who first made it once the kernel selection was understood.

So a quantisation label on a chart is not a specification. It is a description of which kernel the engine happened to pick, and it can be wrong in either direction. Our rule is two tracks, both published. A matched-configuration track uses the same checkpoint and the same precision on every platform, which in practice means FP8 because it is the only format all four run natively, and isolates the hardware. A best-per-platform track lets each box use its strongest supported path and reflects what a buyer would actually deploy. Neither track is honest on its own: matched configuration penalises hardware whose best format is not the common denominator, and best-per-platform lets a precision choice masquerade as a hardware difference. That specific confusion is what turned an earlier industry accelerator comparison into a public dispute, and it is entirely avoidable by labelling both.

Precision also has a quality cost that none of our throughput files measured. A configuration that is fast because it is lossy is not fast, it is broken, and we had no mechanism to tell the difference beyond a four-prompt sanity check that catches total garbage and nothing subtle. That gap is named in the section on what we will not claim yet. Clients who need reproducible behaviour from a model, rather than just tokens per second, should read our notes on AI system auditing alongside these figures.

Bias six: decode-only prompts hide prefill, and TTFT needs a prompt length

Every ladder in the audited corpus used prompts of 39 to 42 tokens and an output cap of 256. That shape is almost pure decode, and decode is memory-bandwidth bound. Prefill, the pass that processes the prompt, is compute bound. Published research on phase-split inference is consistent on that point, and a cross-accelerator study from June 2026 concluded plainly that decode-only evaluation gives an incomplete picture because platform strengths differ by phase.

A decode-only suite therefore rewards whichever platform has the most memory bandwidth per byte of active model, and it conceals the behaviour that matters for retrieval-augmented work, document processing and long-context agents, where a request may carry thousands of prompt tokens and return a short answer. Our fleet has the evidence for how large that gap is. On the four-card RTX PRO 6000 workstation, the same engine measured with long padded prompts instead of 40-token prompts falls to roughly 330 tokens per second aggregate with first-token latencies between 11 and 106 seconds. The decode-shaped run on the same box reports numbers many times higher. Both are true. Only one of them describes a document pipeline.

That makes one of our staged comparisons indefensible as written: a first-token latency of 0.223 seconds on one platform against 1.506 seconds on another at 32 streams. Those figures are real, they are measured, and they are valid only for a 41-token prompt. There is good reason to expect the gap to widen with prompt length, because prefill is where the compute difference lives, but we had not measured it. Every first-token figure now carries its prompt length in the same sentence, and the suite reports two shapes side by side: the decode-heavy shape we already had, and a prefill-heavy shape of two to four thousand unique prompt tokens with a short output.

Prefill also needs a curve rather than a point. First-token latency plotted against input length from one thousand to one hundred and twenty-eight thousand tokens at a single stream isolates compute-bound behaviour cleanly, and it is the plot most likely to change a purchasing decision for anyone doing retrieval work on a private language model deployment.

The metadata every result file must carry

Most of the biases above are not measurement errors. They are recording errors. A run whose exact configuration is written down can be re-examined a week later and either defended or corrected. A run whose configuration lives in somebody's shell history cannot be defended at all, and three of our four retractions were only checkable because a second document happened to preserve the missing detail.

This is the checklist a result file has to satisfy before it is eligible for publication. It is short enough to adopt and mechanical enough to enforce.

  • Host name and hardware description, including how many devices of which model took part.
  • Engine name, version and container image digest.
  • The complete serve command line, read from the running process or the container inspection output, never retyped from memory.
  • Speculative decoding configuration, or an explicit statement that it is off.
  • Graph mode, meaning whether eager execution was forced.
  • Key-value cache dtype, maximum sequence count, maximum model length, maximum batched tokens, and whether prefix caching was enabled.
  • Precision and the exact checkpoint revision, plus which kernel the engine reported selecting.
  • Tensor, expert and pipeline parallel sizes, and the node count.
  • Fabric details for multi-node runs, including the collective library settings that were tuned.
  • A pre-run parity check with a timestamp, for any fleet where nodes can drift out of alignment.
  • A concurrent-load snapshot for any shared or production host, sampled before, during and after the run.
  • Harness git revision, sampling parameters, temperature and seed.
  • Filenames of the paired power trace and quality-gate result.

The last item is the one that changes behaviour most. If a speed number cannot be published without a paired accuracy number from the same endpoint, then no configuration can win by being quietly broken.

What changed in the harness

The audit was a design review, not a code change, and the fixes that came out of it are deliberate and boring.

1. Steady-state windows. Measurement windows rather than bursts, with a ramp period excluded, and at least 300 seconds per level when power is being sampled from a rack power distribution unit whose telemetry only refreshes about every 30 seconds.

2. A fixed output length. Throughput runs use the engine's end-of-sequence override plus a minimum token count, with temperature and seed fixed and recorded. Per-request completion token counts are stored so that percentiles can be recomputed later without a rerun.

3. Count every generated token, including reasoning tokens emitted on a separate field. Our old harness dropped a request that produced no visible content from the success set while still counting its time in the wall clock, which deflated throughput for reasoning models and forced a separate one-off script that was never committed.

4. Two latencies recorded, not one: time to first token of any kind, and time to first answer token after the reasoning block closes. For a reasoning model the user-relevant speed is answers per second, and a token rate that includes an invisible reasoning preamble is not a reading rate.

5. A recorded warmup pass at every batch size before measurement, so lazy graph capture and just-in-time compilation do not land inside the first measured level. The rule is written down before the first scored run, because an unwritten warmup rule is how a benchmark becomes a fairness dispute rather than a result.

6. A run nonce prefixed to every prompt, so that rerunning against a warm server cannot quietly score cache hits as speed.

7. Ladders extended to at least twice the engine sequence cap, plus one ladder with the cap equalised across platforms wherever memory allows, so that an engine setting is never mistaken for a hardware limit.

Decode content class is now stated on every decode figure, because it is not a detail. On one of our own models measured in August 2026, decode ran at 339 tokens per second on mathematics prompts, 246 on prose and 233 on code, on identical hardware. Prose is the conservative choice and it understates speculative decoding gains substantially, which is exactly why the class has to appear next to the number.

What we will not claim yet, and what we owe first

This is the section that makes the rest of the page worth reading.

We are not publishing a cross-platform figure for throughput per dollar or per watt between the GB10 units and the GB300. The retracted parity claim has not been replaced with a better number, and it will not be until the controlled runs below are done. Publishing a corrected ratio built on the same unrecorded settings would repeat the original error with more decimal places.

What is owed before that figure returns:

  • Speculative decoding on against off, and eager against graph mode, on the eight-unit GB10 group with the same image, and the same pair of runs on the GB300 with its speculative configuration recorded rather than inferred. The ratio then gets reported four ways.
  • Wall power, not device power. Our GB10 units draw about 46 watts at the wall with no engine loaded while the accelerator rail reports under 5 watts, so a device-power efficiency figure for these boxes would be wrong by an order of magnitude. Most of the draw sits in the processor cores, the memory, the network adapter and the supply. The GB300 and the H200 host are not on a metered outlet at all and need out-of-band management access before any watt-normalised claim is defensible. A wall number and a device number must never appear in the same comparison.
  • A quality gate attached to every throughput row. Our corpus contains a four-prompt correctness check and nothing else. The designed replacement has three tiers: a 50-prompt mechanical gate that blocks the throughput run on failure, a fixed-subset generative evaluation once per model, precision, platform and image combination, and a quantisation drift check against a reference-precision baseline scored on paired items. Accuracy differences smaller than about five points cannot be resolved by unpaired single runs at these sample sizes, so the pairing is not optional.
  • Prices with a stated basis. A per-dollar table needs the networking to be inside it: a pair needs one cable, a four-way or eight-way group needs switch fabric and breakout cables, and a workstation needs neither. Two of the platforms have no published manufacturer list price at all, which means either a sourced market figure with a market-conditions disclaimer or their exclusion from the table.

Two of our earlier findings also need revisiting, and saying so is part of the method. Our conclusion that a particular large model cannot be served on this hardware class is contradicted by at least two public four-node deployments that got it running with a different block size and an on-disk table layout. And our Marlin kernel observation is most likely a build and kernel-selection problem that others have already worked around, rather than a hardware limit. Neither correction is comfortable and both belong on the record.

We have also not benchmarked training or fine-tuning at all, and the public evidence says that is where multi-unit scaling is closest to linear. We have only tested tensor parallelism, while published results show pipeline parallelism winning by a wide margin at large batch sizes on the same hardware class. Those are gaps, not conclusions, and we list them rather than filling the space with the measurements we happen to have.

Where this sits next to the public GB10 work

Eight-unit clusters of this hardware are not new ground. Alex Ziskind published an eight-unit build in February 2026, covered in the trade press at the time. ServeTheHome published an eight-unit cluster in April 2026 with a detailed account of the networking and storage. Forum contributors have run very large mixture-of-experts models across eight units and reported their numbers openly, including the failures. A 2026 white paper measured per-user capacity at a latency service level across one, two and four units, which is the metric this page argues for, and reached it before we did. Independent write-ups have documented the kernel selection trap, the dual-rail collective settings and the driver-version effects. Where our numbers agree with theirs, that is corroboration and we say so.

Our contribution is narrower than a headline. It is a same-day one, two, four and eight unit series measured on a single harness with the launcher differences reduced to the node count, published alongside the audit that found four errors in our first attempt at interpreting it, and alongside the corrections. Our prefill scaling of roughly 13 to 20 percent from two units to four is in line with everything else published, including the vendor's own figures, which improve first-token latency by far less than they improve decode when units are added. That agreement is the useful signal: no published result we have found shows tensor-parallel prefill scaling across these units, so a buyer who needs prefill throughput should scale out with independent replicas behind a router rather than scale up one wide group.

How to check our work

Three things make a benchmark checkable, and all three are cheap.

The first is the arithmetic in the tables above. Take the aggregate figure, divide by the stream count, and see whether the per-user rate you get matches the per-user rate we published. Where our result files record a per-stream median separately from the aggregate, the two differ because the aggregate includes first-token latency and stragglers inside the wall clock. That difference is itself diagnostic: when it is large, the level contained a stall and the aggregate is not a steady-state number.

The second is the configuration. Every figure on this page names the result file it came from, and those files carry the container image, the checkpoint, the engine arguments and the fabric settings. If a setting that would change the number is missing from a file, that is a defect in the file, and it is the reason the metadata checklist exists.

The third is the losing runs. We publish configurations that did not work, levels that came out worse than the level below them, and models that would not load. A corpus with no failures in it has been curated, and a curated corpus cannot be used to check anything.

If you are evaluating this class of hardware for regulated work, the performance question is the easy half. The harder half is whether the deployment satisfies your obligations, which for defence suppliers means the CMMC requirements and the underlying NIST 800-171 controls, and for everyone else means the ordinary discipline covered in our cybersecurity services. Keeping model inference on hardware you control is a defensible answer to a data-handling question, and we describe how we build that in on-premise AI and managed inference hosting.

Frequently asked questions about this benchmark methodology

What exactly did you retract, and had it been published?

Four conclusions from a page that had been staged for publication and was pulled during review, so the four incorrect statements did not reach the public site. A fifth figure had already been removed earlier in the same week. One related comparison in an earlier blog post of ours has the same structural problem and is being restated on a matched basis. We are describing all of it here because the correction is more useful than the original claim would have been.

Why is comparing at a fixed number of concurrent streams wrong?

Because when the number of devices is the variable, fixed total concurrency changes the load per device. Thirty-two streams is 16 per unit on a pair and 4 per unit on eight units. Batched inference is more efficient per device at higher in-flight counts, so the smaller configuration is measured in its efficient region and the larger one is measured near idle. Divide by unit count and you get a decline that describes the load setting, not the hardware.

What should be held constant instead?

The delivered per-user speed. Pick a floor, for example 20 or 30 tokens per second with a stated first-token latency at the 95th percentile and a stated prompt length, then let each configuration run at the highest concurrency that still meets it. Publish the resulting curve rather than a single point, so a reader can pick the operating point that matches their own service level.

How much does your fleet vary between identical runs?

More than is comfortable. On the eight-unit group with one reasoning model, two runs a day apart gave 257.8 and 316.7 tokens per second at 32 streams, a 23 percent swing, and 247.4 against 148.8 at 16 streams, a 66 percent swing with the direction reversed. Single-stream decode was stable to within one percent across the same two runs, which points at scheduler stalls rather than the model or the fabric. Even on our most stable platform, two runs on the same afternoon agreed to 0.03 percent at 32 streams and differed by 31 percent at 8. All of those figures are prose decode at a 256 token output cap.

Why does speculative decoding have to be recorded?

Because it changes throughput substantially, the change is largest at low batch sizes, and it leaves no trace in a result file that does not record it. In our corpus one platform ran a four-token budget, another a five-token budget, another two tokens, and one was unknown, while a separate flag that disables graph capture cost a factor of 2.4 on one of our own machines. Any ratio built across those configurations is measuring our settings.

Does a four-bit quantisation label mean the hardware is doing four-bit maths?

Not necessarily. On the architecture in our GB10 units, published engine builds have been observed routing four-bit mixture-of-experts work through a kernel that dequantises before computing, so the numbers are a floor for that hardware rather than a ceiling. A second reported case shows a dense four-bit layer selecting an eight-bit-activation kernel and losing 31 percent of prefill throughput. Check which kernel the engine says it selected, and record it.

Why does a first-token latency figure need a prompt length next to it?

Because first-token latency is dominated by prefill, prefill scales with prompt length, and prefill is compute bound while decode is bandwidth bound. A 0.2 second figure measured on a 41-token prompt tells you nothing about a retrieval workload. On one of our own machines, switching from 40-token prompts to long padded prompts moved first-token latency into the 11 to 106 second range and cut aggregate throughput to a fraction of the decode-shaped figure.

Why will you not publish a per-watt or per-dollar comparison between these platforms yet?

Three reasons, all fixable. The speculative decoding and graph-mode settings are not yet recorded consistently, so the underlying throughput ratio is not trustworthy. Two of the platforms are not on metered outlets, and device-level power understates a small unified-memory box by roughly an order of magnitude, so a mixed-scope efficiency figure would be meaningless. And a per-dollar table has to include the networking each configuration requires, with a stated and sourced price basis. We would rather leave the cell empty than fill it with a number we cannot defend.

Can I reproduce these results?

The result files record the container image, the checkpoint, the engine arguments, the fabric tuning and the harness settings for each run, which is what reproduction needs. Where a file is missing a setting that would change the number, that is a defect we have listed rather than hidden. Reproduction on different hardware will not match our absolute figures and is not supposed to; what should reproduce is the shape of the curve and the direction of the comparisons.

Do you sell this hardware, and does that bias the results?

Yes to the first, which is why the second question is fair. Petronella Technology Group, Inc. sells and hosts systems in this class, so we have an interest in the conclusions. The only useful response is method: publish the audit, publish the retractions, publish the losing configurations and the models that would not load, record the settings so a reader can find a mistake, and refuse to publish a comparison whose confounds are known and unmeasured.

About the author

Craig Petronella, founding principal of Petronella Technology Group, Inc.

Craig Petronella is a CMMC Registered Practitioner (CMMC-RP), Cisco CCNA, CWNE and holder of Digital Forensic Examiner License 604180-DFE. He is the Amazon #1 best-selling author of 14+ cybersecurity books and the founding principal of Petronella Technology Group, Inc., which has served regulated and defense-adjacent organizations since 2002.

Petronella Technology Group, Inc. is a Cyber AB Registered Provider Organization (RPO-1449). Every figure on this page was produced on hardware the company owns and operates in Raleigh, North Carolina, using the harness described above. No vendor-supplied numbers are used, and the audit that produced the retractions is described rather than summarised. Read more about Craig Petronella.

Work with the team that publishes its own corrections

If you are sizing on-premise inference hardware, the question worth asking a vendor is not which benchmark they lead. It is which of their own numbers they have retracted, and why. A supplier who cannot answer that has not audited anything.

Petronella Technology Group, Inc. builds and operates private inference on hardware our clients own or lease, including the compliance work that has to sit around it. The benchmark corpus behind this page, including the corrections, is at the fleet benchmark hub, and the deployment side is described under GPU server hosting and self-hosted language models. To talk through a sizing exercise, a per-user capacity target or a regulated deployment, call 919-348-4912 or use the contact form. Bring your workload shape, your prompt lengths and the per-user speed your users will actually tolerate. Those three numbers decide the answer, and no single benchmark figure can substitute for them.

Audit conducted 24 September 2026 against result files measured 21 to 24 September 2026 by Petronella Technology Group, Inc. Result files are retained with this report. Page published 1 October 2026.