Open-Weight LLM Leaderboard: Coding, Research and Writing Scores by GPU

This open-weight LLM leaderboard scores models benchmarked on the private GPU fleet of Petronella Technology Group, Inc. (6x H200 NVL, RTX PRO 6000 Blackwell, GB10). Every published score is cross-judged: the generator never grades itself, and the campaign judge is locked to GPT-4.1 (Session-13/15 lock-in; Haiku-4.5 or validated-equivalent DeepSeek-V4-Flash on legacy rows). Last updated: 2026-08-26 17:07 EDT. Auto-generated from harness result files; the per-row Measured column is the data-collection date.

How to Read This LLM Leaderboard

Every table on this page is generated from the harness result files, so each number is a measured value, not a vendor figure. Scores run from 0 to 1. Coding and research cases blend automatic checks (0.6) with a blind language-model judge (0.4); blog writing is 0.5 structural gates plus 0.5 judge on a 10-point rubric, which is why that table also shows the raw judge score out of 10. The Runs column reports n=, the number of full repetitions behind the row, with the min to max range across them: the score shown is the mean of the model's largest rep family, and scores within 0.03 of each other are statistical ties. The Temp column is the sampling temperature stamped on the run; "default" marks legacy rows that predate the stamp. The Measured column is the date the result file was written, in Eastern time, so the age of any row is visible. Rows labelled "cloud reference" are API-served frontier models included for comparison; every other row ran locally on the fleet hardware described in the benchmark overview and best-in-slot picks.

Read the tables together with the per-hardware picks below: the picks name the model, quantization and serving flags Petronella Technology Group, Inc. actually runs on each tier, from a single GB10 appliance through a four-way H200 pool. If you are planning a deployment, the self-hosted LLM planning guide covers the decisions that come before hardware, the on-premise AI hardware page covers the tiers themselves, and our private AI deployment services team builds and operates the result. Talk to Petronella Technology Group, Inc. if you want a candidate model scored against your own workload before you buy.

Four longer write-ups explain the hardware findings behind these tables: the GB10 versus RTX PRO 6000 throughput write-up, Ollama and vLLM measured on Blackwell, Mistral 3.2 and Gemma-4 on four GPUs, and M5 Ultra against DGX Spark.

Current Picks by Hardware Role

Verdicts dated 2026-08-29; hard-set (v2) verdicts from 2026-07-06 are retained where no newer run exists.
4-way H200 NVL bridge (575 GB)
GLM-5.3-Flash-FP8 (MIT, 320B-A18B, 1M ctx, released 2026-08-26): coding 0.987 (n=3, 0.975-0.995) under the campaign judge; the three-judge panel (gpt-5.4 / opus-4.8 / deepseek, per-case median) scores it 0.9955, tying #1 of 51 panel rows, and #1 on panel research (0.912). Throughput on the NVLink quad: prefill 20.6k tok/s, decode math 339 / prose 246 / code 233 t/s, KV 2.86M tokens. Hard-set (v2) holder is still Ornith-1.0-397B-FP8 (0.971, n=3, 20/20 every rep; blog 1.000). Traps: thinking is on by default and gated by reasoning_effort (a default long-form request can spend the whole budget on hidden reasoning); free-form summarization fabricates (0.736) while citation accuracy is 0.991.
2-way H200 NVL bridge (287 GB)
Qwen3-Coder-Next-FP8 (80B-A3B, Apache-2.0): best coder in the 2-way class on the v2 hard set (0.939, n=3, at 2.6 s per case; the nearest class rival needs 20x the latency for less), 85.2% on the 88-case tool corpus, 178 t/s single-stream and 2,789 t/s aggregate at c=32. Long-form writer alternative: MiniMax-M2.7-NVFP4 (blog 0.94 mean across three topics, v2 0.901).
Single 96 GB card
Qwen3.8-27B-NVFP4 + DSpark (gittensor build, block-draft speculative decoding): coding 0.984 GPT-4.1 / 0.980 panel, 186.7 t/s math and 125.5 code on one RTX PRO 6000; the no-speculation control (0.9785 / 0.9875) shows no measurable quality cost. Displaces Qwen3.6-35B-A3B-FP8 (0.975), which stays the MoE choice for research and cited RAG. NInfer int4 retired (0.909 at 135.5 t/s).
Long context (524K)
DeepSeek-V4-Flash-0731: coding 0.964, blog 0.944; SWE-bench Verified 75.3% (113/150) at vendor-spec sampling on 2x H200 (2026-08-13), tying GLM-5.2's 76.0%. The earlier 67.3% was a temperature-0 arm; at temp 0 the 0731 build produced 2.7% usable patches. Also the fleet's validated free grader.
Voice tool-caller
gpt-oss-20b: confirmed a fifth time (2026-08-17): 125/125 correct declines, 0% tool hallucination on the 88-case corpus. Nemotron-3.5-Lightning has the best raw tool accuracy measured (86.7%, N=1) but no small general-purpose variant; Ling-3.0-tiny is the fastest ever (34 ms TTFT) and the worst on decline-safety (7.95% hallucination).
Single GB10 (128 GB)
Qwen3.6-35B-A3B-NVFP4 + DSpark-8: coding 0.977/0.977 GPT-4.1 (N=2), panel 0.9659, tying the FP8 no-speculation baseline 0.975 (2026-08-17). MoE beats dense about 5.5x on this bandwidth-bound platform, so dense models are retired from the GB10 tier. Nemotron-3.5-Lightning-NVFP4 + DSpark is faster (184.0 t/s math) but loses about 1.8 points under speculation.
GB10 pair (256 GB)
DeepSeek-V4-Flash-DSpark (2x GB10): research 0.963, coding 0.986 (GPT-4.1), 524K ctx; the pair is a model-size and context play, not a speed upgrade.
4-way GB10 fabric (512 GB, 200G RoCE)
Qwen3.8-Flash-Next-NVFP4 (125B-A6B + 51B n-gram table), TP4 + expert-parallel, MTP-4: 39.6 / 57.0 / 76.9 t/s prose / code / math single-stream, prefill 3,550, 167-184 t/s at 6 streams, KV 2.64M tokens (2026-08-29). Quality alternative GLM-5.3-Flash-NVFP4 (26.5 / 38.1 / 46.7, 91 agg; 69 t/s math single-user with DFlash2). Throughput only; fabric quality evals pending.
Air-gap edge appliance
Ministral-3-8B (2026-05-27 re-test); Granite-4.1-8B and Gemma-4-e4b for tool-driving. All Mistral models need vLLM --tool-call-parser mistral.

Want one of these running inside your compliance boundary?

Petronella Technology Group, Inc. designs, builds and operates private AI on hardware you own, from a single GB10 appliance to a multi-H200 enclave, for firms that handle CUI, PHI or privileged data. Every pick above is a configuration we run ourselves.

Coding v1: 22-Task Longitudinal Set

22 real operations, SEO and debugging tasks; the longitudinal reference set, saturated at the top.
ModelScoreRuns (range)PassLatencyTempMeasured
Claude Opus-4.8 (cloud reference) BEST0.995n=122/222.6sdefault2026-07-06
Claude Opus-4.70.991n=122/222.7sdefault2026-05-25
Laguna-S-2.1-NVFP40.991n=122/223.3sdefault2026-07-27
Qwen3.6-27B BF16 (dense)0.990n=120/2266.4sdefault2026-06-08
gpt-oss:20b0.989n=122/222.5sdefault2026-05-28
glm-5.3-flash0.987n=3 (0.975-0.995)22/224.4s0.22026-08-26
Qwen3.6-27B (dense)0.984n=122/2225.2sdefault2026-06-30
MiniMax-M3 (MXFP8)0.984n=3 (0.975-0.991)21/225.2sdefault2026-07-05
Laguna-M.1-NVFP40.984n=122/226.4sdefault2026-07-26
DeepSeek-V4-Pro (cloud reference)0.982n=122/224.1sdefault2026-06-28
GPT-5.2-Codex (cloud reference)0.982n=122/222.6sdefault2026-07-06
Ornith-1.0-397B (FP8)0.980n=3 (0.977-0.982)22/222.7sdefault2026-07-05
Qwen3.6-35B-A3B-NVFP4-Fast (unsloth)0.980n=5 (0.971-0.986)22/2220.0sdefault2026-07-12
Gemma-4-31B0.980n=122/2228.9sdefault2026-05-27
Nex-N2-Pro (NVFP4)0.980n=122/2218.1sdefault2026-06-10
North-Mini-Code-1.0 (FP8, Cohere)0.980n=5 (0.976-0.986)22/223.4sdefault2026-06-20
gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX50900.978n=2 (0.977-0.980)22/22110.9sdefault2026-08-17
Qwen3-Coder-Next-80B0.977n=3 (0.977-0.977)22/220.6sdefault2026-07-05
Gemma-4-12B0.977n=122/2215.8sdefault2026-06-03
Qwen3.8-27B (BF16)0.977n=122/2235.7sdefault2026-08-14
MiniMax-M2.7 (NVFP4)0.977n=3 (0.965-0.984)22/2217.3sdefault2026-07-05
gemma-4-26b-a4b-nvfp40.977n=3 (0.970-0.980)22/221.5sdefault2026-07-15
Qwen3.6-35B-A3B0.975n=5 (0.961-0.984)22/220.6sdefault2026-05-29
gpt-oss-20b0.975n=122/2211.8sdefault2026-05-28
ornith-35b0.975n=14/226.6sdefault2026-06-29
glm-5.20.975n=122/226.3sdefault2026-07-13
GLM-4.7-Flash0.973n=122/228.9sdefault2026-05-28
Qwen3.6-35B-A3B (BF16)0.973n=122/2222.7sdefault2026-06-08
GLM-5.2 (NVFP4)0.972n=5 (0.959-0.986)22/2246.0sdefault2026-07-03
qwen3.8-27b-fp80.972n=2 (0.968-0.975)22/2222.1sdefault2026-08-16
Ornith-1.0-397B (W4A16 int4)0.970n=5 (0.959-0.982)21/223.6sdefault2026-07-03
Gemma-4-26B-A4B (BF16)0.970n=122/221.0sdefault2026-06-09
GPT-5.4 (cloud reference)0.970n=122/221.7sdefault2026-07-06
Laguna-XS-2.1-NVFP40.970n=121/220.7sdefault2026-07-26
Mistral-Medium-3.5-128B0.968n=122/224.6sdefault2026-05-26
Qwen3.6-27B-NVFP4 (unsloth, dense)0.967n=5 (0.966-0.970)21/2214.4sdefault2026-07-12
qwen3-coder:480b0.966n=122/222.1sdefault2026-06-28
Qwen3.8-27B (FP8)0.966n=121/22124.6sdefault2026-08-14
gpt-oss-120b0.964n=121/228.4sdefault2026-05-28
glm-5.2-nvfp40.964n=122/2280.0s0.22026-08-26
DeepSeek-V4-Flash0.964n=5 (0.961-0.970)21/220.9sdefault2026-05-29
GLM-5.2 (IQ2_M 2-bit)0.959n=121/2226.9sdefault2026-06-20
gpt-4.10.957n=121/221.4sdefault2026-05-25
Mistral-Small-4-119B0.957n=121/221.4sdefault2026-05-27
gemma4-coder0.957n=122/223.1sdefault2026-07-14
GLM-4.5-Air0.956n=117/22105.3sdefault2026-05-28
claude-haiku-4-5-202510010.956n=5 (0.955-0.959)21/221.2sdefault2026-05-29
Gemma-4-31B (BF16)0.955n=121/224.9sdefault2026-06-09
z-ai/glm-5.20.955n=121/2217.3sdefault2026-06-27
Qwen3-Coder-30B0.952n=122/220.8sdefault2026-06-09
openai/gpt-oss-20b0.952n=121/221.8sdefault2026-05-26
qwen3-coder:30b0.952n=122/220.8sdefault2026-05-27
Gemma-4-12B (BF16)0.952n=121/222.5sdefault2026-06-09
Gemma-4-e4b (BF16)0.949n=122/221.8sdefault2026-06-09
Muse-Glimmer-30B (Meta)0.949n=121/2214.6sdefault2026-08-15
Ornith-1.0-35B (FP8)0.948n=5 (0.913-0.970)21/223.7sdefault2026-07-02
Qwen3.8-27B (NInfer int4+MTP3)0.941n=121/2218.8sdefault2026-08-15
command-a-plus0.939n=121/2218.1sdefault2026-05-26
nemotron-3.5-lightning-30b-a3b-nvfp40.935n=3 (0.926-0.945)22/221.1sdefault2026-08-16
Gemma-4-e2b0.935n=121/2212.0sdefault2026-05-26
Qwen3.5-Opus-distill (27B)0.934n=121/2213.1sdefault2026-05-26
Qwen3.8-27B (Q4_K_XL)0.934n=120/22171.1sdefault2026-08-15
Devstral-Small-2-24B0.932n=120/226.7sdefault2026-05-28
gemma-4-12b-nvfp40.931n=3 (0.930-0.932)21/222.1sdefault2026-07-15
Gemma-4-e4b0.931n=120/2213.6sdefault2026-05-26
DeepSeek-V4-Flash-DSpark (2x GB10)0.929n=3 (0.924-0.935)20/229.9sdefault2026-08-16
Falcon-H1R-7B0.926n=119/2284.2sdefault2026-05-27
devstral-small-2-24b0.925n=5 (0.900-0.936)20/220.7sdefault2026-07-03
qwen3.8-27b0.914n=2 (0.898-0.930)19/2218.5sdefault2026-08-16
Mixtral-8x22B0.909n=121/221.9sdefault2026-05-26
Mistral-Small-24B0.909n=120/221.4sdefault2026-06-09
Granite-4.1-8B0.903n=119/220.8sdefault2026-06-09
hf.co/tiiuae/Falcon-H1R-7B-GGUF:Q4_K_M0.885n=119/2249.9sdefault2026-05-27
q25-gptq0.885n=119/222.4sdefault2026-05-29
Apriel-1.6-15B0.880n=119/22110.6sdefault2026-05-27
Gemma-4-26B-A4B0.877n=118/2218.4sdefault2026-05-26
ornith-9b0.875n=119/2254.0sdefault2026-06-29
Ministral-3-8B0.872n=118/222.9sdefault2026-05-27
hf.co/unsloth/Ministral-3-8B-Instruct-2512-GGUF:Q4_K_M0.858n=118/220.8sdefault2026-05-27
qwen2.5:7b-instruct0.847n=118/222.1sdefault2026-05-29
l31-awq0.847n=119/223.4sdefault2026-05-29
LFM2.5-8B-A1B (Liquid)0.846n=117/222.0sdefault2026-06-03
q25-awq0.820n=117/222.5sdefault2026-05-29
llama3.1:8b0.820n=117/222.7sdefault2026-05-29
l31-gptq0.794n=115/222.9sdefault2026-05-29
Mixtral-8x7B0.773n=117/220.7sdefault2026-05-26
MiniCPM5-1B0.644n=110/226.2sdefault2026-05-27
Incumbent fleet coder Qwen3.6-35B-A3B now co-leads. North-Mini-Code-1.0 (Cohere, Apache-2.0, 30B-total/3B-active MoE) matches the very top on Petronella Technology Group coding - 0.980 +/- 0.004 (N=5), statistically tied with Nex-N2-Pro and edging Qwen3.6-35B-FP8 (0.975) - and loop-amplifies on agentic tasks (cap=1 0.833 to cap=5 0.917, fabrication 2 to 1), unlike the non-amplifiers below. Apache-2.0, FP8 fits one card, fast, clean tools; a genuine alternative fleet coder (Session-22). It is not a research/RAG model (0.821) and blogs short of 3,000 words. Gemma-4-12B (0.977) ties the 31B/Qwen3.6 on single-shot coding but is single-shot only - it does NOT loop-amplify (cap=1=cap=5=0.917 on agentic tasks) and fabricates "done" under loop pressure; use it as a fast coder, not an agent. LFM2.5-8B-A1B is an on-device model (edge-tier 0.846, 3.41% tool-call hallucination) - not a fleet upgrade. Score = mean of each model's largest rep family (newest on ties); scores within 0.03 of each other are statistical ties. Claude Opus-4.7 reference = 1.00. Scores cross-judged by GPT-4.1, Haiku-4.5, or validated-equivalent DeepSeek-V4-Flash. Temperature column: stamped value where present; "default" = pre-Session-13 runs (eval_coding.py default 0.0).
Per-case breakdown: coding_tasks.jsonl (22 cases; full prompt texts withheld to keep the benchmark uncontaminated)
Case idCategoryAutomatic checksJudge scale
01_bash_backup_recent_htmlbash_opssyntax:bash, must_include(3), must_include_any(3)1-5
02_fish_fleet_uptimebash_opssyntax:fish, must_include(3), must_include_any(2)1-5
03_bash_safe_rsync_deploybash_opssyntax:bash, must_include(3), must_include_any(2)1-5
04_py_openai_compatible_callpython_automationsyntax:python, must_include(3), must_not_include(2)1-5
05_py_factorial_constraintsinstruction_followingsyntax:python, must_include(1), must_not_include(2)1-5
06_py_retry_backoff_decoratorpython_automationsyntax:python, must_include(2), must_include_any(3)1-5
07_py_concurrent_endpoint_pingpython_automationsyntax:python, must_include(2), must_include_any(3)1-5
08_py_parse_nvidia_smi_powerpython_automationsyntax:python, must_include(3)1-5
09_py_phone_regexpython_automationsyntax:python, must_include(2), must_include_any(2)1-5
10_systemd_timer_oncalendarconfig_editmust_include(3), must_not_include(1)1-5
11_yaml_add_fleet_tierconfig_editmust_include(5)1-5
12_htaccess_301_httpsconfig_editmust_include(4), must_include_any(3)1-5
13_php_canonical_tagpython_automationmust_include(2), must_include_any(2)1-5
14_sql_top_referring_domainssqlmust_include(5), must_include_any(4)1-5
15_debug_keyerrordebug_fixsyntax:python, must_include(1), must_include_any(4)1-5
16_debug_ollama_cpu_onlydebug_fixmust_include_any(6)1-5
17_refactor_bash_looprefactorsyntax:bash, must_include(3), must_not_include(1)1-5
18_refactor_preserve_signatureinstruction_followingsyntax:python, must_include(1)1-5
19_git_branch_commit_pushinstruction_followingsyntax:bash, must_include(3), must_include_any(2)1-5
20_py_argparse_clipython_automationsyntax:python, must_include(4)1-5
21_multifile_config_not_loadedmulti_file_reasoningsyntax:python, must_include(2), must_include_any(3)1-5
22_py_idempotent_insert_guardpython_automationsyntax:python, must_include(2), must_include_any(4)1-5

Coding v2: 20-Case Hard Set

20 cases, added 2026-07-05.
ModelScoreRuns (range)PassLatencyTempMeasured
Claude Opus-4.8 (cloud reference) BEST0.974n=120/207.7sdefault2026-07-06
Ornith-1.0-397B (FP8)0.971n=3 (0.966-0.977)20/2013.0sdefault2026-07-05
GLM-5.2 FP8 (Z.ai serving, reference)0.952n=3 (0.927-0.965)19/2078.5sdefault2026-07-06
GLM-5.2 (NVFP4)0.942n=3 (0.918-0.968)20/20147.6sdefault2026-07-06
Qwen3-Coder-Next-80B0.939n=3 (0.930-0.950)20/203.0sdefault2026-07-05
GPT-5.4 (cloud reference)0.935n=120/204.0sdefault2026-07-06
MiniMax-M3 (MXFP8)0.931n=3 (0.923-0.936)18/2030.0sdefault2026-07-06
GPT-5.2-Codex (cloud reference)0.928n=119/207.1sdefault2026-07-06
MiniMax-M2.7 (NVFP4)0.901n=119/2048.8sdefault2026-07-05
Qwen3.6-35B-A3B0.892n=3 (0.879-0.903)18/2027.4sdefault2026-07-06
The original 22-case suite saturated (the top ten sit within 0.02, below the noise floor), so v2 was built to discriminate at the top: multi-file bug tracing (import cycles, config precedence, systemd environment, cache-header rollouts), hard debugging (mutable defaults, DST arithmetic, subprocess deadlock, asyncio fan-out), strict spec-compliance tasks where any missed constraint costs points, gaps-and-islands SQL, and infrastructure tasks (Quadlet units, hotfix git surgery, parallel-safe bash). Same scoring blend and GPT-4.1 judge as v1. The hard set separates models the saturated suite could not: leaders that tie at 0.97-0.99 on v1 spread across 0.89-0.97 here. v1 remains the longitudinal reference; v2 decides picks.
Per-case breakdown: coding_tasks_v2.jsonl (20 cases; full prompt texts withheld to keep the benchmark uncontaminated)
Case idCategoryAutomatic checksJudge scale
v2_01_interval_merge_edgealgorithmssyntax:python, must_include(3), must_include_any(2)1-5
v2_02_toposort_cyclealgorithmssyntax:python, must_include(4), must_include_any(3)1-5
v2_03_lru_ttlalgorithmssyntax:python, must_include(6), must_not_include(2)1-5
v2_04_debug_mutable_defaulthard_debugsyntax:python, must_include(2), must_include_any(4), must_not_include(1)1-5
v2_05_debug_tz_dsthard_debugsyntax:python, must_include(2), must_include_any(4), must_not_include(1)1-5
v2_06_debug_subprocess_deadlockhard_debugsyntax:python, must_include(4), must_not_include(1)1-5
v2_07_debug_async_gatherhard_debugsyntax:python, must_include(2), must_include_any(2)1-5
v2_08_spec_versioned_configspec_compliancesyntax:python, must_include(4), must_not_include(3)1-5
v2_09_spec_redact_loggerspec_compliancesyntax:python, must_include(6), must_include_any(2)1-5
v2_10_spec_atomic_writespec_compliancesyntax:python, must_include(5), must_not_include(2)1-5
v2_11_spec_cli_exitcodesspec_compliancesyntax:python, must_include(4), must_include_any(2), must_not_include(2)1-5
v2_12_multifile_import_cyclemulti_file_reasoningsyntax:python, must_include(3), must_include_any(3)1-5
v2_13_multifile_env_precedencemulti_file_reasoningsyntax:python, must_include(2), must_include_any(4)1-5
v2_14_multifile_systemd_envmulti_file_reasoningmust_include(3), must_include_any(4)1-5
v2_15_multifile_nginx_cachemulti_file_reasoningmust_include(3), must_include_any(4)1-5
v2_16_sql_window_dedupsql_hardmust_include(3), must_include_any(3)1-5
v2_17_sql_upsert_countersql_hardmust_include(5), must_include_any(2)1-5
v2_18_bash_parallel_safeinfra_hardsyntax:bash, must_include(4), must_include_any(4)1-5
v2_19_podman_quadletinfra_hardmust_include(7), must_include_any(2)1-5
v2_20_gitops_hotfixinfra_hardsyntax:bash, must_include(6), must_include_any(3)1-5

Research and Reasoning: 27 Cases

27-case set, no retrieval context, task-appropriate temperature 0.0 to 0.3.
ModelScoreRuns (range)PassLatencyTempMeasured
Qwen3.6-27B (dense) BEST0.963n=126/2737.6sdefault2026-06-30
Qwen3-Coder-Next-80B0.963n=126/272.0sdefault2026-07-05
Gemma-4-31B (BF16)0.963n=126/2711.7sdefault2026-06-09
Gemma-4-26B-A4B (BF16)0.963n=126/272.2sdefault2026-06-09
Nex-N2-Pro (NVFP4)0.963n=126/273.2sdefault2026-06-10
DeepSeek-V4-Flash-DSpark (2x GB10)0.963n=126/2725.9sdefault2026-06-30
Ornith-1.0-397B (FP8)0.963n=126/2722.9sdefault2026-07-05
MiniMax-M3 (MXFP8)0.963n=126/2715.6sdefault2026-07-05
gemma-4-12b-nvfp40.963n=126/275.0sdefault2026-07-15
gemma-4-26b-a4b-nvfp40.963n=126/273.5sdefault2026-07-15
DeepSeek-V4-Flash0.926n=125/2715.1sdefault2026-05-27
Claude Opus-4.70.926n=125/279.5sdefault2026-05-25
Qwen3.6-27B BF16 (dense)0.926n=125/27161.2sdefault2026-06-09
Gemma-4-12B (BF16)0.926n=125/275.6sdefault2026-06-09
Gemma-4-e4b (BF16)0.926n=125/272.7sdefault2026-06-09
GLM-5.2 (NVFP4)0.926n=125/2746.1sdefault2026-07-05
GLM-5.2 FP8 (Z.ai serving, reference)0.926n=125/2735.0sdefault2026-07-06
Qwen3.8-27B (FP8)0.926n=125/27719.9sdefault2026-08-14
Granite-4.1-8B0.889n=124/272.6sdefault2026-06-09
Qwen3.6-35B-A3B (BF16)0.889n=124/2724.2sdefault2026-06-08
devstral-small-2-24b0.889n=124/273.1sdefault2026-07-03
Gemma-4-e2b0.852n=123/2712.2sdefault2026-05-26
Gemma-4-31B0.852n=123/2743.2sdefault2026-05-27
Qwen3.5-Opus-distill (27B)0.852n=123/27270.5sdefault2026-05-27
Ministral-3-8B0.852n=123/2711.5sdefault2026-05-27
Ornith-1.0-35B (FP8)0.852n=123/278.8sdefault2026-07-02
Ornith-1.0-397B (W4A16 int4)0.852n=123/2726.5sdefault2026-07-05
MiniMax-M2.7 (NVFP4)0.852n=123/2712.2sdefault2026-07-05
glm-5.3-flash0.852n=123/2713.1s0.22026-08-26
glm-5.2-nvfp40.852n=123/27146.1s0.22026-08-26
Qwen3.6-35B-A3B0.815n=122/271.7sdefault2026-06-02
GLM-4.5-Air0.815n=122/27121.8sdefault2026-05-28
Qwen3-Coder-30B0.815n=122/271.7sdefault2026-06-09
Muse-Glimmer-30B (Meta)0.815n=122/2733.8sdefault2026-08-15
Gemma-4-e4b0.778n=121/2720.0sdefault2026-05-26
North-Mini-Code-1.0 (FP8, Cohere)0.778n=121/276.9sdefault2026-06-20
nemotron-3.5-lightning-30b-a3b-nvfp40.778n=121/274.8sdefault2026-08-11
Mistral-Small-4-119B0.741n=120/2711.0sdefault2026-05-26
Mistral-Small-24B0.741n=120/274.2sdefault2026-06-09
Gemma-4-26B-A4B0.556n=115/2746.4sdefault2026-05-26
Heretic-9B0.000n=10/27-default2026-05-27
Re-baselined on a single cloud judge (Haiku-4.5 or GPT-4.1) per Session-13 lock-in. Qwen3.6-35B-A3B, DeepSeek-V4-Flash and Claude Opus-4.7 tie at 0.926 - fleet reasoning is at Opus parity. The Claude-Opus reasoning-distill (Qwen3.5-Opus, 0.852) did NOT beat native dense.
Per-case breakdown: research_tasks.jsonl (27 cases; full prompt texts withheld to keep the benchmark uncontaminated)
Case idCategoryAutomatic checksJudge scale
summ_01long_doc_summarizationmust_include(4), must_not_include(3)1-5
summ_02long_doc_summarizationmust_include(3), must_not_include(2)1-5
summ_03long_doc_summarizationmust_include(4), must_not_include(2)1-5
summ_04long_doc_summarizationmust_include(3), must_not_include(2)1-5
summ_05long_doc_summarizationmust_include(4), must_not_include(2)1-5
summ_06long_doc_summarizationmust_include(3), must_not_include(2)1-5
summ_07long_doc_summarizationmust_include(7), must_not_include(2)1-5
mhqa_01multi_hop_qamust_include(3), must_include_any(3), must_not_include(2)1-5
mhqa_02multi_hop_qamust_include(2), must_not_include(4)1-5
mhqa_03multi_hop_qamust_include(2), must_include_any(5), must_not_include(2)1-5
mhqa_04multi_hop_qamust_include(3), must_include_any(6)1-5
mhqa_05multi_hop_qamust_include(2), must_include_any(3), must_not_include(1)1-5
mhqa_06multi_hop_qamust_include_any(7), must_not_include(2)1-5
mhqa_07multi_hop_qamust_include(2), must_include_any(4)1-5
cite_01citation_accuracymust_include(3), must_include_any(2), must_not_include(4)1-5
cite_02citation_accuracymust_include_any(6), must_not_include(4)1-5
cite_03citation_accuracymust_include(1), must_include_any(2), must_not_include(3)1-5
cite_04citation_accuracymust_include_any(7), must_not_include(4)1-5
cite_05citation_accuracymust_include(4), must_include_any(2), must_not_include(3)1-5
cite_06citation_accuracymust_include_any(9), must_not_include(4)1-5
code_01code_generationsyntax:bash, must_include(5), must_include_any(2), must_not_include(1)1-5
code_02code_generationsyntax:python, must_include(6), must_include_any(2), must_include_all(3)1-5
code_03code_generationsyntax:fish, must_include(5), must_include_any(2), must_not_include(2)1-5
code_04code_generationsyntax:python, must_include(10), must_include_any(3), must_not_include(4)1-5
code_05code_generationsyntax:bash, must_include(5), must_include_any(2), must_not_include(2)1-5
code_06code_generationsyntax:python, must_include(7), must_include_any(2), must_not_include(1)1-5
code_07code_generationsyntax:bash, must_include(10), must_include_any(2), must_not_include(1)1-5

Blog Writing: 10-Criteria Rubric

Combined score = 0.5 structural + 0.5 judge; each row at the model's best-per-model temperature.
ModelScoreRuns (range)WordsJudgeGen speedTempMeasured
DeepSeek-V4-Flash BEST1.000n=23774-64s 0.32026-05-28
Qwen3.6-27B BF16 (dense)1.000n=1382910/10477s 28 t/sdefault2026-06-08
Nex-N2-Pro (NVFP4)1.000n=1441110/1095s 91 t/sdefault2026-06-10
Qwen3.6-27B (dense)1.000n=1330510/1081s 98 t/sdefault2026-06-30
qwen3.6-27b-bf161.000n=3 (1.000-1.000)315710/10148s 60 t/sdefault2026-07-02
Ornith-1.0-397B (FP8)1.000n=3 (1.000-1.000)385710/1064s 120 t/sdefault2026-07-05
GLM-5.2 (NVFP4)1.000n=1343910/10207s 40 t/sdefault2026-07-06
MiniMax-M3 (MXFP8)1.000n=3 (1.000-1.000)331710/1046s 123 t/sdefault2026-07-05
Claude Opus-4.8 (cloud reference)1.000n=1329710/10125s 77 t/sdefault2026-07-06
Claude Fable 5 (cloud reference)1.000n=1363010/10135s 81 t/sdefault2026-07-06
Ornith-1.0-35B (FP8)0.963n=3 (0.945-1.000)221010/1066s 166 t/sdefault2026-07-02
gpt-oss-120b0.945n=21938-132s 0.02026-05-28
Gemma-4-31B (BF16)0.945n=1239010/10169s 24 t/sdefault2026-06-09
Gemma-4-12B (BF16)0.945n=1229110/1078s 53 t/sdefault2026-06-09
North-Mini-Code-1.0 (FP8, Cohere)0.945n=1226910/1043s 139 t/sdefault2026-06-20
DeepSeek-V4-Flash-DSpark (2x GB10)0.945n=1264610/10131s 46 t/sdefault2026-06-30
Laguna-S-2.1-NVFP40.945n=1325510/1055s 106 t/sdefault2026-07-26
Ornith-1.0-397B (W4A16 int4)0.945n=3 (0.944-0.945)32519/1070s 104 t/sdefault2026-07-03
GLM-4.7-Flash0.945n=22640-59s 0.02026-05-28
Qwen3.6-35B-A3B (BF16)0.944n=132939/1049s 171 t/sdefault2026-06-08
GLM-5.2 (IQ2_M 2-bit)0.944n=141279/10199s 39 t/sdefault2026-06-20
GPT-5.4 (cloud reference)0.944n=132469/1064s 102 t/sdefault2026-07-06
MiniMax-M2.7 (NVFP4)0.907n=3 (0.833-1.000)39669/1050s 129 t/sdefault2026-07-05
Qwen3.6-35B-A3B (FP8)0.889n=22597-38s 0.02026-05-28
Gemma-4-26B-A4B0.889n=122809/10128s 44 t/sdefault2026-05-26
Gemma-4-31B0.889n=120639/10553s 7 t/sdefault2026-05-27
Gemma-4-26B-A4B (BF16)0.889n=1212810/1026s 147 t/sdefault2026-06-09
Gemma-4-e4b (BF16)0.889n=1267610/1040s 122 t/sdefault2026-06-09
Laguna-XS-2.1-NVFP40.889n=117039/1030s 192 t/sdefault2026-07-26
glm-5.3-flash0.889n=132038/1025s 216 t/s0.42026-08-26
Qwen3-Coder-Next-80B0.870n=3 (0.833-0.945)32109/1036s 180 t/sdefault2026-07-05
gemma-4-26b-a4b-nvfp40.870n=3 (0.833-0.945)18279/1034s 97 t/sdefault2026-07-15
Qwen3.6-27B-NVFP4 (unsloth, dense)0.855n=5 (0.333-1.000)32119/1089s 98 t/sdefault2026-07-12
Gemma-4-e2b0.833n=128649/1080s 78 t/sdefault2026-05-26
command-a-plus0.833n=119539/10103s 49 t/sdefault2026-05-26
Qwen3.5-Opus-distill (27B)0.833n=141979/10240s 37 t/sdefault2026-05-26
Laguna-M.1-NVFP40.833n=119269/1058s 79 t/sdefault2026-07-26
nemotron-3.5-lightning-30b-a3b-nvfp40.833n=135829/10104s 73 t/sdefault2026-08-11
glm-5.2-nvfp40.833n=135337/10920s 15 t/s0.42026-08-26
Claude Opus-4.70.805n=22342-107s default (~0.0)2026-05-28
Qwen3.6-35B-A3B-NVFP4-Fast (unsloth)0.800n=5 (0.278-0.945)271910/1056s 138 t/sdefault2026-07-12
devstral-small-2-24b0.796n=3 (0.778-0.833)13079/1032s 124 t/sdefault2026-07-03
gpt-oss-20b0.778n=117939/1034s 243 t/sdefault2026-05-25
GLM-4-9B0.778n=18987/1012s 181 t/sdefault2026-05-25
Mixtral-8x22B0.778n=19317/1041s 60 t/sdefault2026-05-26
Granite-4.1-8B0.778n=18689/1013s 169 t/sdefault2026-06-09
Qwen3.8-27B (FP8)0.778n=125228/10700s 8 t/sdefault2026-08-14
Mistral-Small-24B0.722n=113427/1039s 74 t/sdefault2026-06-09
Gemma-4-e4b0.722n=124347/10105s 47 t/sdefault2026-05-26
Mistral-Small-4-119B0.722n=132628/10281s 28 t/sdefault2026-05-26
gemma-4-12b-nvfp40.704n=3 (0.333-0.889)61383/10202s 59 t/sdefault2026-07-15
Mistral-Medium-3.5-128B0.611n=114697/10450s 18 t/sdefault2026-05-26
GLM-4.5-Air0.500n=145413/101360s 6 t/sdefault2026-05-25
nemotron-3-nano:30b0.445n=156182/10156s 51 t/sdefault2026-05-25
Qwen3.6-35B-A3B0.278n=11953/10212s 38 t/sdefault2026-05-25
Mixtral-8x7B0.222n=1792/101s 137 t/sdefault2026-05-26
Qwen3-Coder-30B0.111n=101/1011s 0 t/sdefault2026-06-09
3,000+ word SEO CMMC blog from a fixed prompt. Each row reports the model's best-scoring temperature when the Session-14 Bench-2 sweep covered it (the 2026-05-28 model set only); models benched after that date show their single measured temperature. Scores within 0.03 are statistical ties. Autoblog runs overnight in batch, so quality decides.

Blog Temperature Sweep

Session-14 Bench-2; judge GPT-4.1; N=2 reruns per cell.
ModelTempMean scoreRange (min-max)Mean wordsMean gen
Claude Opus-4.7default (~0.0)0.8050.778-0.8332342107.2s
Claude Opus-4.7default (~0.3)0.8050.778-0.8332342107.2s
Claude Opus-4.7default (~0.7)0.8050.778-0.8332342107.2s
Claude Opus-4.7default (~1.0)0.8050.778-0.8332342107.2s
DeepSeek-V4-Flash0.00.9720.944-1.000419478.4s
DeepSeek-V4-Flash0.3 BEST1.0001.000-1.000377464.2s
DeepSeek-V4-Flash0.71.0001.000-1.000373265.0s
DeepSeek-V4-Flash1.00.9720.944-1.000402572.4s
GLM-4.7-Flash0.0 BEST0.9450.944-0.945264058.8s
GLM-4.7-Flash0.30.8050.722-0.889221053.1s
GLM-4.7-Flash0.70.8610.833-0.889274857.9s
GLM-4.7-Flash1.00.8890.889-0.889207249.5s
Qwen3.6-35B-A3B (FP8)0.0 BEST0.8890.889-0.889259737.5s
Qwen3.6-35B-A3B (FP8)0.30.8060.667-0.945320439.8s
Qwen3.6-35B-A3B (FP8)0.70.5840.445-0.722251744.2s
Qwen3.6-35B-A3B (FP8)1.00.6110.611-0.611254941.5s
gpt-oss-120b0.0 BEST0.9450.945-0.9451938131.8s
gpt-oss-120b0.30.8890.889-0.8892145137.4s
gpt-oss-120b0.70.9450.945-0.9452144127.8s
gpt-oss-120b1.00.9170.889-0.9451710124.2s
Session-14 Bench-2 temperature sweep measured 2026-05-28 13:48 EDT. Reasoning models (DeepSeek-V4-Flash, Qwen3.6-35B-A3B) may treat temperature as a hint during their reasoning phase; a flat row across temps is itself a finding. Claude Opus-4.7 rejects the temperature parameter; its rows are copied from a single default-temperature run.

How the Scores Are Produced

Coding - 22 real operations tasks, not textbook puzzles. Drawn from Petronella Technology Group's day-to-day infrastructure and SEO work: 8 Python automation tasks (OpenAI-compatible API clients, retry/backoff decorators, concurrent endpoint health checks, log parsing, argparse CLIs), 3 bash/fish shell-ops tasks (timestamped backups, safe rsync deploys, fleet uptime sweeps), 3 production config edits (systemd OnCalendar timers, YAML inventories, .htaccess 301 rules), 3 instruction-following traps (tasks that fail if a stated constraint is ignored, such as preserving a function signature), 2 debug-and-fix cases, 1 SQL analytics query, 1 shell refactor, and 1 multi-file reasoning case. Each case is scored 0.6 x automatic checks (required content, regex gates, and real syntax validation: bash -n, python compile, fish -n) + 0.4 x a blind LM-judge correctness score (1-5; the judge sees only task and solution, never the model's identity). Deterministic decoding (temp 0.0). The leaderboard number is the mean across all 22 cases.
Research & reasoning - 27 cases in four categories. 7 long-document summarization cases (word bounds, bullet counts, must-include / must-not-include gates against supplied source documents), 7 multi-hop QA cases (the answer requires chaining facts across documents), 6 citation-grounding cases (presence and format of citations regex-verified against the supplied sources; semantic citation correctness is measured separately in Petronella Technology Group's cited-RAG evaluations, which score a much harsher 0.65-0.80 F1), and 7 grounded code-generation cases with syntax gates. Same 0.6 auto + 0.4 blind-judge blend as coding.
Blog writing - one fixed brief, structural gates plus publish-readiness. Every model gets the identical brief: a 3,000+ word CMMC compliance post in clean HTML. Nine automatic structural gates (word count, H1/H2/H3 hierarchy with 8+ sections, no leftover markdown artifacts, FAQ with 5+ substantive Q&As, a well-formed comparison table, internal links, correct call-to-action, and brand rules) plus a 10-criteria publish-readiness judge score (1-10) covering factual accuracy (CMMC 2.0 levels, NIST 800-171), tone, and SEO structure. Combined = 0.5 x structural pass rate + 0.5 x judge. Blog rows report each model's best temperature from the Session-14 sweep where covered.
Voice tool-calling - 88-case corpus with decline-safety distractors. Real assistant tool schemas; the corpus mixes valid calls with hard distractors that look like tool calls but must be refused. Scored on correct calls, correct declines, and hallucinated-call rate. This suite decides the voice tool-caller pick; decline-safety on distractors is the deciding metric.
Agentic coding - SWE-bench Verified, no LM judge. Stratified 150-instance subset of SWE-bench Verified run through the official harness with mini-swe-agent; a task counts only if the generated patch resolves the issue's test suite. Referenced in picks and findings (for example DeepSeek-V4-Flash 67.3%, measured 2026-06-28); not a leaderboard column.
Throughput - measured on the serving hardware. Single-stream (c=1) plus concurrent c=8 and c=32 sweeps of 800-token generations against the exact serve config named in each finding. Speeds quoted in findings are decode tokens per second on the stated GPUs and quant.
Judging integrity - the cross-judge gate. A generator never grades itself: self-judging inflated scores 20-37% in our measurements. The campaign judge is locked to GPT-4.1 (code-enforced since 2026-05-30, with a guard that flags any same-family judge pairing). Legacy rows graded by Claude Haiku-4.5 are labeled; DeepSeek-V4-Flash is accepted for bulk grading after validating equivalence with GPT-4.1 (r=0.962, mean absolute delta 0.18 on the 1-5 scale; re-audited 2026-07-06 on 40 blog outputs: mean abs delta 0.45 on the 1-10 scale, 87.5% within 1 point, second judge slightly stricter). Excluded from all tables: misconfigured or errored runs, runs without a cross-family judge, and hardware-specific Mac coding runs (see the Apple Silicon matrix).
Reproducibility. Every row carries an absolute measurement date; the exact quantization is recorded in the run label (quant changes the result: the same model at FP8, NVFP4, and int4 scores differently and is listed separately). Headline claims run N=3 to N=5 repetitions. As of 2026-07-05 the board shows the MEAN of each model's largest rep family (newest on ties) with n and min-max range in the Runs column (previously best-run, which rewarded models benched more often); scores within 0.03 are statistical ties.

Findings

Measured 2026-08-28/29
What does a 4-way GB10 fabric buy, and what does it take to use it? Four GB10 units on a switched 200G RoCE fabric (two MikroTik CRS812; 196 Gb/s node-to-node RDMA measured through the switch with no PFC/ECN tuning) pool 512 GB. Qwen3.8-Flash-Next NVFP4 (125B-A6B plus a 51B n-gram table) split four ways decodes 39.6 / 57.0 / 76.9 tokens/s on prose / code / math single-stream, 3,550 tokens/s prefill, and 167-184 tokens/s aggregate at six streams with a 2.6M-token KV cache; the same model on a two-unit pair does 35 / 36 / 53 and 118 aggregate. GLM-5.3-Flash NVFP4 (18B active) on the same four units: 26.5 / 38.1 / 46.7, prefill 3,840, 91 aggregate. The platform is memory-bandwidth-bound, so the smaller-active model wins decode and the fabric is never the limiter (11-17 GB of RoCE traffic per run). Three practical findings. Plain 4-way tensor parallelism cannot load the NVFP4 MoE (expert intermediate width 640 does not divide four ways for the 4-bit kernels): expert parallelism is required, and using both 100G halves of each node's link adds 5-15% (EP all-to-all benefits; GLM's all-reduce does not). MTP-4 is the speculative sweet spot: MTP-6 (3.6 of 6 drafts accepted) and piecewise CUDA graphs (25.5 / 37.1 / 44.6, KV cache shrank 6x) both lose, and a DFlash2 block-draft head lifts GLM to 69 tokens/s on math single-stream but halves the 6-stream aggregate (acceptance collapses to about 2). One GB10 can serve the 125B model alone by memory-mapping the n-gram table from NVMe (24.8 / 31.1 / 33.8, 86 aggregate), so four independent units give about 345 tokens/s aggregate versus 184 pooled. Ops caveat that cost a night: a gpu-memory-utilization of 0.85 leaves 1-4% of host memory free on a 121 GB unified-memory unit and an earlyoom daemon will silently kill the worker; 0.65 (4-way) and 0.72-0.78 (single) with a memory guard are the working floors. Throughput only; fabric quality evaluations pending.
Measured 2026-08-26
GLM-5.3-Flash: what replaced GLM-5.2 on the H200 quad, and its two traps. zai-org/GLM-5.3-Flash (MIT, 320B-A18B, hybrid KDA + NoPE sparse MLA, native FP8, 1M ctx) leads or ties every panel-graded axis: coding 0.9955 (ties #1 of 51 panel rows), research 0.912 (#1 of 21); campaign-judge coding 0.987 (n=3). On 4x NVLink H200: prefill 10.3k -> 20.6k tok/s from 1.5k to 23k context, decode math 339 / prose 246 / code 233 t/s, KV 2.86M tokens (21.8x concurrency at 131k). A 320B model decoding like a 35B. Trap 1: thinking is on by default and gated by reasoning_effort (low | high | max), not a Qwen-style enable_thinking flag; at default effort a long-form request can return empty content with finish_reason=length. Trap 2: free-form summarization fabricates (long_doc_summarization 0.736, judges penalize invented cost bands, dates and regulatory claims) while citation_accuracy is 0.991; use it grounded, never as a free summarizer of compliance material. Launch traps: the vendor's fp8 KV cache is invalid on Hopper (BF16 KV required); --model must be a local path. GLM-5.2 was deleted from the fleet the same day.
Measured 2026-08-18
Does llama.cpp on a GB10 have a speculative path? Yes, since mainline merged MTP drafting (ggml-org PR #22673: --spec-type draft-mtp --spec-draft-n-max 3). The forum claim of ~27 t/s for a Q4_K_XL GGUF on GB10 was real and used it. Within dense-on-GB10, llama.cpp Q4_K_XL + MTP (20-25 t/s) now beats vLLM FP8 + MTP (12.2 t/s). Trap: --spec-type draft-dspark on a plain GGUF is a silent no-op (healthy server, zero log lines, baseline speed); verify the "creating MTP draft context" log line before believing any spec flag.
Measured 2026-08-15
Why did a cabled 4-way GB10 cluster get slower, and what is a GB10 actually good at? Direct-attached (two ConnectX-7 ports per node), four units in TP=4 on a 27B dense FP8 model decoded 3.5 t/s and prefilled 539 tok/s versus 8.3 t/s and 1,963 tok/s on ONE unit: all-to-all RDMA is physically impossible with two ports (three links per node needed), so a 4-way cluster needs a switch, not cables. That switch arrived on 2026-08-28 (fabric card above). Also found: one card's firmware had throttled RDMA to 12.74 Gb/s while every status tool reported healthy; only a bandwidth test finds it (111.86 Gb/s after apt full-upgrade + reboot). Re-measure after every GB10 rebuild. The durable number: prefill:decode on one GB10 is 237:1 (1,963 vs 8.3), and decode lands within 12% of the LPDDR5X prediction (273 GB/s / 29 GB = 9.4 t/s). Routing rule: long-prompt, short-output work (RAG, summarization, classification, reranking) to the GB10; long generation to the RTX PRO 6000.
Measured 2026-08-13
How much does sampling config move SWE-bench? DeepSeek-V4-Flash-0731 at vendor-spec sampling scored 75.3% (113/150) on the stratified SWE-bench Verified set on 2x H200, tying GLM-5.2's 76.0% (itself a temperature-0, off-spec run). The temperature-0 arm of the same model collapsed to 2.7% usable patches. A QA audit found the earlier temp-0 SWE series (Qwen 55.3, Opus 82.0 anchors, DeepSeek 67.3) was accidental and inconsistent across models; treat any agentic-coding number without its sampling config as non-reproducible.
Measured 2026-08-16/17
Does speculative decoding cost quality, and how much speed does block-drafting buy? Block-draft speculative decoding (DSpark) took a Nemotron-3.5 MoE on a GB10 node from 75.1 to 184.0 tokens/s (2.45x), beating standard multi-token prediction by 79% on the same hardware; a mixture-of-experts model also beat a dense model of similar quality tier by about 5.5x on this bandwidth-bound platform, so dense models are retired from the GB10 tier entirely. Quality is a separate question and the answer is per-model: against purpose-built controls that change ONLY the speculative flags, Nemotron-3.5 loses about 1.8 points (three-judge panel) under both MTP and DSpark, while Qwen3.6-35B-A3B (0.977 vs its 0.975 baseline, N=2) and Qwen3.8-27B (0.984 DSpark vs 0.979 control, deltas inside run spread) show no measurable cost. Vendor claims of "lossless by construction" should be treated as a hypothesis to test per model, with a control that changes one variable. One measurement caveat that moved every published number: decode speed depends strongly on content class - the same DSpark config measured 184.0 tokens/s on math, 120.6 on code, and 105.0 on prose - so a decode rate without its content class understates or overstates by up to 75%.
Measured 2026-08-16
Does 4-bit quantization cost quality? It depends on the format, not the bit-width. NVFP4 (W4A4 with lm_head, embeddings, and speculative heads kept in BF16) statistically TIES FP8 on the 22-case coding set under a three-judge frontier panel (0.9795 vs 0.9762 on identical prompts, N=2), while naive 4-bit formats of the same model lose 4-7 points (Q4_K_XL GGUF 0.934, a groupwise-int4 engine 0.909). The earlier rule of thumb "4-bit costs about 4 points of coding quality" is true for naive 4-bit and false for NVFP4. Practical read for 96 GB-class cards: NVFP4 buys FP8 quality at roughly half the weight memory, and the freed VRAM buys context and co-tenancy - decode speed itself is bandwidth-bound and does not improve.
Measured 2026-07-15
Is Gemma-4-26B-A4B better than Qwen3.6, and can it run in the agent gateway? It depends on task difficulty, judged by a three-model frontier panel (GPT-5.4 + Claude Opus-4.8 + GLM-5.2, the panel in force until DeepSeek-V4 took the third seat in August 2026; median score on identical generations, N=5). On easy/routine coding Qwen wins (dense 27B 0.983, 35B-FP8 0.979 vs Gemma 0.962); on the hard coding set Gemma wins decisively (0.915 vs every Qwen variant 0.78-0.81), corroborated by the deterministic auto-checks. Research and blog are ties (about 0.86 and 0.79). Agentic: served with the correct vLLM tool parser (gemma4, not pythonic - the wrong one produces zero tool calls), Gemma is the strongest tool-caller (55 correct calls vs Qwen's 51) and solves every multi-turn task (pass@max 1.00 vs Qwen-FP8 0.778). Both models were verified end-to-end through the real agent gateway - writing and executing code and returning correct results - which requires serving the model at 131k context (the agent system prompt is about 58k tokens). NVFP4 is near-lossless for Gemma-4 but not for Qwen3.6-35B (a long-form runaway appears at FP4; use FP8).
Measured 2026-07-06
Does GLM-5.2 lose anything at FP4, and is the six-GPU FP8 layout justified? No measurable loss. GLM-5.2 at FULL FP8 fidelity (Z.ai's own serving, via OpenRouter) scored 0.952 on the v2 hard set (n=3: 0.927/0.964/0.965) and 0.92 on research - versus the FP4 quant on the fleet's 4-way bridge at 0.942 and 0.94. The FP8 delta (+0.010 coding, -0.02 research) is inside the 0.03 tie threshold: FP4 is effectively lossless for GLM-5.2 on these tests, and both configurations score below Ornith-1.0-397B-FP8 (0.971). Implication for fleet layout: an FP8 deployment spanning six GPUs and both NVLink bridges buys no measurable quality over FP4 on four GPUs.
Measured 2026-07-06
Do two independent judges agree on these scores? Yes. A dual-judge audit re-graded 40 stored blog outputs with DeepSeek-V4-Flash (served on the DGX Spark pair) against their original GPT-4.1 scores using the identical rubric: mean absolute difference 0.45 points on the 1-10 scale, 87.5% of pairs within one point, and the second judge ran slightly STRICTER on average (signed mean -0.3) with its five largest disagreements spread across Anthropic, MiniMax, and Qwen outputs - no family favoritism pattern. This is the periodic cross-judge integrity check the methodology commits to.
Measured 2026-07-06
How does the fleet compare to the July 2026 cloud frontier on the same tests? Cloud reference rows added: Claude Opus-4.8 v1 0.995 / v2 0.974 / blog 1.000 (GPT-4.1 judge); GPT-5.4 v1 0.970 / v2 0.935 / blog 0.944 and GPT-5.2-Codex v1 0.982 / v2 0.928 (both Haiku-judged per the cross-family gate; OpenAI rows and Anthropic/fleet rows use different judges, so treat cross-vendor deltas under 0.05 with care). Read: fleet-owned Ornith-1.0-397B-FP8 (v2 0.971, zero marginal cost, CUI-capable) sits between GPT-5.4 (0.935) and Opus-4.8 (0.974) on the hard set. Also measured: Claude Fable 5's API safety layer CONTENT-FILTERED 9 of 22 routine infrastructure tasks (finish_reason=content_filter after ~3 tokens: rsync deploy scripts, a phone-number regex, an argparse CLI), making it unusable as a coding baseline here; its blog run scored 1.000. GPT-5.2-Codex is Responses-API-only (the harness gained --use-responses-api).
Measured 2026-07-06
Do the bridge picks survive a harder test set and multi-topic writing? Yes, and the ordering sharpens. The 20-case v2 coding set (built because the original suite saturated) spreads the leaders decisively: Ornith-1.0-397B-FP8 0.971 (n=3, 20/20 every rep, 13 s per case), GLM-5.2-NVFP4 0.942 (n=3, high variance 0.918-0.968 at 150 s per case), Qwen3-Coder-Next 0.939 (n=3, 2.6 s per case), MiniMax-M3 0.931 (n=3, drops 1-2 cases per rep), incumbent Qwen3.6-35B-A3B 0.892 (n=3). MiniMax-M3's v1 crown (0.984) inverts on v2: the saturated suite was measuring judge quibbles, not capability. Blog re-tested at 3 topics x 3 reps: Ornith-397B a perfect 1.000 on all nine runs; MiniMax-M3 0.975 mean (short only on the CMMC brief); MiniMax-M2.7 0.94 mean (its earlier single-run 1.000 was variance); Qwen3-Coder-Next 0.887 mean. Qwen3-Coder-Next also posted 85.2% on the 88-case tool corpus (53 correct calls, 24 correct declines, 1 hallucinated call): strong for agentic coding, not voice-grade (gpt-oss-20b remains 0% hallucination). Remaining before any production routing change: a local SWE-bench Verified run for Qwen3-Coder-Next.
Measured 2026-07-05
Which open model is best on each NVLink island of the 6x H200 fleet? First head-to-head since the 4-way bridge install. On the 4-way bridge (575 GB), Ornith-1.0-397B-FP8 (MIT) is the balanced winner: coding 0.980 (N=3), blog 1.000, research 0.91 at 119.7 t/s single-stream and 639 t/s at c=8, roughly 2x the speed of speed-tuned GLM-5.2-NVFP4 at equal or better quality on two of three axes. MiniMax-M3-MXFP8 posted the campaign's best coding rep (0.991; N=3 mean 0.984) with 1M context and multimodal input, held back only by its non-MIT community license. GLM-5.2-NVFP4 keeps the research crown (0.94, citation accuracy 0.992). On the 2-way bridge (287 GB), Qwen3-Coder-Next-FP8 (80B-A3B) is the coder/agentic pick: 0.977 coding with zero variance across three full reps, clean native tool calls, 178 t/s single-stream and 2,789 t/s aggregate at c=32. Ruled out: Kimi-K2.6 (594 GB INT4 exceeds any island), GLM-5.x-FP8 (754 GB needs all six cards), Nemotron-3-Ultra-NVFP4 (numerically broken on Hopper: loads but emits gibberish). Serving notes: Ornith-397B-FP8 requires VLLM_TEST_FORCE_FP8_MARLIN=1 plus an explicit chat template; MiniMax-M3 requires --block-size 128; the 2026-07-03 vLLM nightly has broken block-scaled FP8 kernels on Hopper (pin v0.24.0).
Which models do we keep resident and route to? Coding: Qwen3.6-35B-A3B (0.975, GPT-4.1). Research: Qwen3.6-35B (speed) / DeepSeek-V4-Flash (long-context), both Opus-parity. Blog to DeepSeek-V4-Flash / GLM-4.7-Flash. Voice tool-call: gpt-oss-20b. Edge appliance: Granite-4.1-8B (+ Gemma-4-e4b for tool-driving). Reserve Claude Opus for the top few percent.
Do self-improving loops help small models? A loop is a capability amplifier, not an equalizer: Qwen3.6-35B goes 0.917 to 1.0 with iterations; small models (Gemma-4-e4b, GLM-4.7-Flash) stay flat at 0.917. The done-gate makes a small model honest (no silent fabrication), not capable.
Self-improving loop vs a general agent (pi.dev) on a 4B? Gemma-4-e4b scored 0.917 in a done-gated loop vs 0.0 in pi.dev (it fabricated all 12 tasks). Air-gap appliances should pair a small model with a programmatic verifier loop, never a general agent framework.
Are there other open models worth adding? A live Feb to May 2026 scan found none that beat the incumbents; Qwen3.7/Qwen4, DeepSeek-R2, Phi-5 and Grok-3 weights are unreleased or hosted-only.
What is the best model for a single RTX PRO 6000 96GB (Blackwell) card, and is a 35B a waste of it? No displacer. Qwen3.6-35B-A3B (3B active) is the best all-rounder that fits one card: it wins research and cited-RAG outright and leads coding (0.975, GPT-4.1). The models large enough to “fill” the card (gpt-oss-120b 0.964, Mistral-Small-4-119B 0.957) are slower and weaker on the role’s core axes. A low-active MoE is the correct shape for a 96GB concurrency server: comparable NVFP4 models scale to ~2,000 t/s aggregate at c=32 on this card. Spare VRAM is best spent on KV/concurrency, or on NVFP4 (same quality at half the VRAM, freeing room to co-locate a second model), not on a bigger-but-worse model. Mistral-Small-4-119B is the lone alternative, and only if the card is redefined as a cited-RAG / vision / compliance resident.
Is a purpose-built Rust inference engine (Atlas) faster than our tuned vLLM on Blackwell? (measured 2026-06-07) No. On identical GB10 (DGX-Spark-class) hardware and the same Qwen3.6-35B-A3B-NVFP4 model, our tuned vLLM (NVIDIA MTP recipe) ran 116 to 119 tok/s steady-state vs Atlas’s 88.9; Atlas’s advertised “130 to 133 tok/s” and “3.1x faster than vLLM” did not reproduce (the 3.1x is vs an untuned vLLM). Atlas serving is quality-preserving - blog 0.944 (ties our blog leader) and 6/6 on a coding spot-check - and ships an ~8x-smaller (2.98 GB) no-Python single binary. That makes it a candidate packaging vehicle for an air-gapped compliance appliance, not a throughput upgrade. Its multi-node expert-parallel mode is not yet shipping (runtime is single-node only).
Does an agentic multi-hop retriever beat single-shot RAG for compliance Q&A? (measured 2026-06-07) On hard multi-hop CMMC / NIST 800-171 questions, an RL-trained search agent (Harness-1, 21B, gpt-oss-20b base) found every gold control (retrieval recall 1.000) where single-shot dense top-8 reached only 0.881 - it recovers the deep 2nd/3rd-hop controls single-shot drops at production cutoffs. But its curated answer (0.929) only matched single-shot top-15 (the curation step, not the search, is the bottleneck) and cost ~1,000x the latency - so the value is exhaustive batch retrieval (audit / SSP gap analysis), not interactive RAG. Control-id deduplication remains the cheap universal lever: it lifts both single-shot (0.786 to 0.881) and the agent (0.905 to 1.000).

Apple Silicon (MLX) Results

These rows come from the Apple Silicon workload matrix measured 2026-05-26 with MLX (mlx_lm.server 0.31.3), one model resident at a time, by unified-memory tier. Coding is cross-judged by Claude Haiku-4.5 (the generator never equals the judge). The first table is the dense-versus-MoE head-to-head on the 64 GB M1 Max at the same quantization; the rest are per-tier matrices for the 16 GB M4 Mac mini, the 32 GB M5 and the 64 GB M1 Max, followed by the macOS GPU wired-memory ceiling rule. Bench-only: nothing here changes production routing.

Dense vs MoE: the decision rule

AxisDense Llama-70B (4-bit)MoE Qwen3.6-35B-A3B (4-bit)
Decode6.54 t/s54.3 t/s  (8.3x faster)
Coding (Haiku-judged)0.8990.952
Memory footprint41 GB14.2 GB
Tool-call accuracyn/a (no MLX template)97%

16 GB: M4 Mac mini

ModelQuantSizeFits no-swapDecode t/sCodingTool-call
Llama-3.1-8B-Instruct BEST FITMLX 4-bit5 GByes18.60.84386.4%
Qwen2.5-14B-InstructMLX 4-bit7.7 GBtight10.30.915n/a*
gpt-oss-20bMXFP410 GBedge34-390.69**n/a*
Qwen3.6-35B-A3BMLX 3-bit14.5 GBOOMn/an/an/a

32 GB: M5 (the value sweet spot)

ModelQuantSizeDecode t/sTTFTCodingTool-call
Qwen3.6-35B-A3B DAILY DRIVERMLX 4-bit MoE21 GB50.50.2-0.5s0.93197.0%
Qwen2.5-7B-InstructMLX 8-bit8.1 GB15.40.67s0.889n/a*
Llama-3.1-8B-InstructMLX 4-bit4.8 GB26.30.69s0.82986.4%
gpt-oss-20bMXFP411 GB46.61.21s0.694**n/a*

64 GB: M1 Max

ModelQuantSizeDecode t/sCodingTool-call
Qwen3.6-35B-A3B DAILY DRIVERMLX 4-bit MoE14.2 GB54.30.95297.0%
Llama-3.3-70B-InstructMLX 4-bit dense41 GB6.540.899n/a
gpt-oss-120bMXFP459 GBOOMn/an/a

The ceiling rule (macOS GPU wired cap about 67-75% of RAM)

Unified RAMGPU budgetLargest comfortable modelHard ceiling
16 GB (M4)~10-11 GB8B-4bit (14B tight)~10 GB / no 20B+ headroom on a shared box
32 GB (M5)~20-24 GB30-35B-A3B-4bit MoEno 70B
64 GB (M1 Max)~43-48 GB35B MoE (or 70B-4bit single-stream)no 120B
* mlx_lm.server 0.31.3 silently drops native tool_calls for Qwen2.5 and gpt-oss, so those tool-call cells are n/a; measure those models via Ollama. ** gpt-oss-20b coding is depressed by a reasoning-channel parsing artifact, not capability. n/a in a speed or score cell means the model did not fit (OOM) on that machine.
Benchmarked and published by Petronella Technology Group, Inc. All scores are produced by the Petronella Technology Group llm-benchmark harness described in "How the scores are produced" above.

About the Author

Craig Petronella, founding principal of Petronella Technology Group, Inc.

Craig Petronella is a CMMC Registered Practitioner (CMMC-RP), Cisco CCNA, CWNE and holder of Digital Forensic Examiner License 604180-DFE. He is the Amazon #1 best-selling author of 14+ cybersecurity books and the founding principal of Petronella Technology Group, Inc., which has served regulated and defense-adjacent organizations since 2002.

Petronella Technology Group, Inc. is a Cyber AB Registered Provider Organization (RPO-1449). Every score on this page was produced on hardware the company owns and operates in Raleigh, North Carolina, using the evaluation harness described in the methodology section; no vendor-supplied numbers are used.

Talk to the team that measured every number on this page

A 30-minute call is enough to tell you which tier fits your users, your data and your budget, and what it will cost to run.

Petronella Technology Group, Inc. · 5540 Centerview Dr., Suite 200, Raleigh, NC 27606 · 919-348-4912 · info@petronellatech.com · Last Updated: August 26, 2026