Open-Weight LLM Leaderboard: Coding, Research and Writing Scores by GPU
This open-weight LLM leaderboard scores models benchmarked on the private GPU fleet of Petronella Technology Group, Inc. (6x H200 NVL, RTX PRO 6000 Blackwell, GB10). Every published score is cross-judged: the generator never grades itself, and the campaign judge is locked to GPT-4.1 (Session-13/15 lock-in; Haiku-4.5 or validated-equivalent DeepSeek-V4-Flash on legacy rows). Last updated: 2026-08-26 17:07 EDT. Auto-generated from harness result files; the per-row Measured column is the data-collection date.
How to Read This LLM Leaderboard
Every table on this page is generated from the harness result files, so each number is a measured value, not a vendor figure. Scores run from 0 to 1. Coding and research cases blend automatic checks (0.6) with a blind language-model judge (0.4); blog writing is 0.5 structural gates plus 0.5 judge on a 10-point rubric, which is why that table also shows the raw judge score out of 10. The Runs column reports n=, the number of full repetitions behind the row, with the min to max range across them: the score shown is the mean of the model's largest rep family, and scores within 0.03 of each other are statistical ties. The Temp column is the sampling temperature stamped on the run; "default" marks legacy rows that predate the stamp. The Measured column is the date the result file was written, in Eastern time, so the age of any row is visible. Rows labelled "cloud reference" are API-served frontier models included for comparison; every other row ran locally on the fleet hardware described in the benchmark overview and best-in-slot picks.
Read the tables together with the per-hardware picks below: the picks name the model, quantization and serving flags Petronella Technology Group, Inc. actually runs on each tier, from a single GB10 appliance through a four-way H200 pool. If you are planning a deployment, the self-hosted LLM planning guide covers the decisions that come before hardware, the on-premise AI hardware page covers the tiers themselves, and our private AI deployment services team builds and operates the result. Talk to Petronella Technology Group, Inc. if you want a candidate model scored against your own workload before you buy.
Four longer write-ups explain the hardware findings behind these tables: the GB10 versus RTX PRO 6000 throughput write-up, Ollama and vLLM measured on Blackwell, Mistral 3.2 and Gemma-4 on four GPUs, and M5 Ultra against DGX Spark.
Current Picks by Hardware Role
reasoning_effort (a default long-form request can spend the whole budget on hidden
reasoning); free-form summarization fabricates (0.736) while citation accuracy is 0.991.--tool-call-parser mistral.Want one of these running inside your compliance boundary?
Petronella Technology Group, Inc. designs, builds and operates private AI on hardware you own, from a single GB10 appliance to a multi-H200 enclave, for firms that handle CUI, PHI or privileged data. Every pick above is a configuration we run ourselves.
Coding v1: 22-Task Longitudinal Set
| Model | Score | Runs (range) | Pass | Latency | Temp | Measured |
|---|---|---|---|---|---|---|
| Claude Opus-4.8 (cloud reference) BEST | 0.995 | n=1 | 22/22 | 2.6s | default | 2026-07-06 |
| Claude Opus-4.7 | 0.991 | n=1 | 22/22 | 2.7s | default | 2026-05-25 |
| Laguna-S-2.1-NVFP4 | 0.991 | n=1 | 22/22 | 3.3s | default | 2026-07-27 |
| Qwen3.6-27B BF16 (dense) | 0.990 | n=1 | 20/22 | 66.4s | default | 2026-06-08 |
| gpt-oss:20b | 0.989 | n=1 | 22/22 | 2.5s | default | 2026-05-28 |
| glm-5.3-flash | 0.987 | n=3 (0.975-0.995) | 22/22 | 4.4s | 0.2 | 2026-08-26 |
| Qwen3.6-27B (dense) | 0.984 | n=1 | 22/22 | 25.2s | default | 2026-06-30 |
| MiniMax-M3 (MXFP8) | 0.984 | n=3 (0.975-0.991) | 21/22 | 5.2s | default | 2026-07-05 |
| Laguna-M.1-NVFP4 | 0.984 | n=1 | 22/22 | 6.4s | default | 2026-07-26 |
| DeepSeek-V4-Pro (cloud reference) | 0.982 | n=1 | 22/22 | 4.1s | default | 2026-06-28 |
| GPT-5.2-Codex (cloud reference) | 0.982 | n=1 | 22/22 | 2.6s | default | 2026-07-06 |
| Ornith-1.0-397B (FP8) | 0.980 | n=3 (0.977-0.982) | 22/22 | 2.7s | default | 2026-07-05 |
| Qwen3.6-35B-A3B-NVFP4-Fast (unsloth) | 0.980 | n=5 (0.971-0.986) | 22/22 | 20.0s | default | 2026-07-12 |
| Gemma-4-31B | 0.980 | n=1 | 22/22 | 28.9s | default | 2026-05-27 |
| Nex-N2-Pro (NVFP4) | 0.980 | n=1 | 22/22 | 18.1s | default | 2026-06-10 |
| North-Mini-Code-1.0 (FP8, Cohere) | 0.980 | n=5 (0.976-0.986) | 22/22 | 3.4s | default | 2026-06-20 |
| gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 | 0.978 | n=2 (0.977-0.980) | 22/22 | 110.9s | default | 2026-08-17 |
| Qwen3-Coder-Next-80B | 0.977 | n=3 (0.977-0.977) | 22/22 | 0.6s | default | 2026-07-05 |
| Gemma-4-12B | 0.977 | n=1 | 22/22 | 15.8s | default | 2026-06-03 |
| Qwen3.8-27B (BF16) | 0.977 | n=1 | 22/22 | 35.7s | default | 2026-08-14 |
| MiniMax-M2.7 (NVFP4) | 0.977 | n=3 (0.965-0.984) | 22/22 | 17.3s | default | 2026-07-05 |
| gemma-4-26b-a4b-nvfp4 | 0.977 | n=3 (0.970-0.980) | 22/22 | 1.5s | default | 2026-07-15 |
| Qwen3.6-35B-A3B | 0.975 | n=5 (0.961-0.984) | 22/22 | 0.6s | default | 2026-05-29 |
| gpt-oss-20b | 0.975 | n=1 | 22/22 | 11.8s | default | 2026-05-28 |
| ornith-35b | 0.975 | n=1 | 4/22 | 6.6s | default | 2026-06-29 |
| glm-5.2 | 0.975 | n=1 | 22/22 | 6.3s | default | 2026-07-13 |
| GLM-4.7-Flash | 0.973 | n=1 | 22/22 | 8.9s | default | 2026-05-28 |
| Qwen3.6-35B-A3B (BF16) | 0.973 | n=1 | 22/22 | 22.7s | default | 2026-06-08 |
| GLM-5.2 (NVFP4) | 0.972 | n=5 (0.959-0.986) | 22/22 | 46.0s | default | 2026-07-03 |
| qwen3.8-27b-fp8 | 0.972 | n=2 (0.968-0.975) | 22/22 | 22.1s | default | 2026-08-16 |
| Ornith-1.0-397B (W4A16 int4) | 0.970 | n=5 (0.959-0.982) | 21/22 | 3.6s | default | 2026-07-03 |
| Gemma-4-26B-A4B (BF16) | 0.970 | n=1 | 22/22 | 1.0s | default | 2026-06-09 |
| GPT-5.4 (cloud reference) | 0.970 | n=1 | 22/22 | 1.7s | default | 2026-07-06 |
| Laguna-XS-2.1-NVFP4 | 0.970 | n=1 | 21/22 | 0.7s | default | 2026-07-26 |
| Mistral-Medium-3.5-128B | 0.968 | n=1 | 22/22 | 4.6s | default | 2026-05-26 |
| Qwen3.6-27B-NVFP4 (unsloth, dense) | 0.967 | n=5 (0.966-0.970) | 21/22 | 14.4s | default | 2026-07-12 |
| qwen3-coder:480b | 0.966 | n=1 | 22/22 | 2.1s | default | 2026-06-28 |
| Qwen3.8-27B (FP8) | 0.966 | n=1 | 21/22 | 124.6s | default | 2026-08-14 |
| gpt-oss-120b | 0.964 | n=1 | 21/22 | 8.4s | default | 2026-05-28 |
| glm-5.2-nvfp4 | 0.964 | n=1 | 22/22 | 80.0s | 0.2 | 2026-08-26 |
| DeepSeek-V4-Flash | 0.964 | n=5 (0.961-0.970) | 21/22 | 0.9s | default | 2026-05-29 |
| GLM-5.2 (IQ2_M 2-bit) | 0.959 | n=1 | 21/22 | 26.9s | default | 2026-06-20 |
| gpt-4.1 | 0.957 | n=1 | 21/22 | 1.4s | default | 2026-05-25 |
| Mistral-Small-4-119B | 0.957 | n=1 | 21/22 | 1.4s | default | 2026-05-27 |
| gemma4-coder | 0.957 | n=1 | 22/22 | 3.1s | default | 2026-07-14 |
| GLM-4.5-Air | 0.956 | n=1 | 17/22 | 105.3s | default | 2026-05-28 |
| claude-haiku-4-5-20251001 | 0.956 | n=5 (0.955-0.959) | 21/22 | 1.2s | default | 2026-05-29 |
| Gemma-4-31B (BF16) | 0.955 | n=1 | 21/22 | 4.9s | default | 2026-06-09 |
| z-ai/glm-5.2 | 0.955 | n=1 | 21/22 | 17.3s | default | 2026-06-27 |
| Qwen3-Coder-30B | 0.952 | n=1 | 22/22 | 0.8s | default | 2026-06-09 |
| openai/gpt-oss-20b | 0.952 | n=1 | 21/22 | 1.8s | default | 2026-05-26 |
| qwen3-coder:30b | 0.952 | n=1 | 22/22 | 0.8s | default | 2026-05-27 |
| Gemma-4-12B (BF16) | 0.952 | n=1 | 21/22 | 2.5s | default | 2026-06-09 |
| Gemma-4-e4b (BF16) | 0.949 | n=1 | 22/22 | 1.8s | default | 2026-06-09 |
| Muse-Glimmer-30B (Meta) | 0.949 | n=1 | 21/22 | 14.6s | default | 2026-08-15 |
| Ornith-1.0-35B (FP8) | 0.948 | n=5 (0.913-0.970) | 21/22 | 3.7s | default | 2026-07-02 |
| Qwen3.8-27B (NInfer int4+MTP3) | 0.941 | n=1 | 21/22 | 18.8s | default | 2026-08-15 |
| command-a-plus | 0.939 | n=1 | 21/22 | 18.1s | default | 2026-05-26 |
| nemotron-3.5-lightning-30b-a3b-nvfp4 | 0.935 | n=3 (0.926-0.945) | 22/22 | 1.1s | default | 2026-08-16 |
| Gemma-4-e2b | 0.935 | n=1 | 21/22 | 12.0s | default | 2026-05-26 |
| Qwen3.5-Opus-distill (27B) | 0.934 | n=1 | 21/22 | 13.1s | default | 2026-05-26 |
| Qwen3.8-27B (Q4_K_XL) | 0.934 | n=1 | 20/22 | 171.1s | default | 2026-08-15 |
| Devstral-Small-2-24B | 0.932 | n=1 | 20/22 | 6.7s | default | 2026-05-28 |
| gemma-4-12b-nvfp4 | 0.931 | n=3 (0.930-0.932) | 21/22 | 2.1s | default | 2026-07-15 |
| Gemma-4-e4b | 0.931 | n=1 | 20/22 | 13.6s | default | 2026-05-26 |
| DeepSeek-V4-Flash-DSpark (2x GB10) | 0.929 | n=3 (0.924-0.935) | 20/22 | 9.9s | default | 2026-08-16 |
| Falcon-H1R-7B | 0.926 | n=1 | 19/22 | 84.2s | default | 2026-05-27 |
| devstral-small-2-24b | 0.925 | n=5 (0.900-0.936) | 20/22 | 0.7s | default | 2026-07-03 |
| qwen3.8-27b | 0.914 | n=2 (0.898-0.930) | 19/22 | 18.5s | default | 2026-08-16 |
| Mixtral-8x22B | 0.909 | n=1 | 21/22 | 1.9s | default | 2026-05-26 |
| Mistral-Small-24B | 0.909 | n=1 | 20/22 | 1.4s | default | 2026-06-09 |
| Granite-4.1-8B | 0.903 | n=1 | 19/22 | 0.8s | default | 2026-06-09 |
| hf.co/tiiuae/Falcon-H1R-7B-GGUF:Q4_K_M | 0.885 | n=1 | 19/22 | 49.9s | default | 2026-05-27 |
| q25-gptq | 0.885 | n=1 | 19/22 | 2.4s | default | 2026-05-29 |
| Apriel-1.6-15B | 0.880 | n=1 | 19/22 | 110.6s | default | 2026-05-27 |
| Gemma-4-26B-A4B | 0.877 | n=1 | 18/22 | 18.4s | default | 2026-05-26 |
| ornith-9b | 0.875 | n=1 | 19/22 | 54.0s | default | 2026-06-29 |
| Ministral-3-8B | 0.872 | n=1 | 18/22 | 2.9s | default | 2026-05-27 |
| hf.co/unsloth/Ministral-3-8B-Instruct-2512-GGUF:Q4_K_M | 0.858 | n=1 | 18/22 | 0.8s | default | 2026-05-27 |
| qwen2.5:7b-instruct | 0.847 | n=1 | 18/22 | 2.1s | default | 2026-05-29 |
| l31-awq | 0.847 | n=1 | 19/22 | 3.4s | default | 2026-05-29 |
| LFM2.5-8B-A1B (Liquid) | 0.846 | n=1 | 17/22 | 2.0s | default | 2026-06-03 |
| q25-awq | 0.820 | n=1 | 17/22 | 2.5s | default | 2026-05-29 |
| llama3.1:8b | 0.820 | n=1 | 17/22 | 2.7s | default | 2026-05-29 |
| l31-gptq | 0.794 | n=1 | 15/22 | 2.9s | default | 2026-05-29 |
| Mixtral-8x7B | 0.773 | n=1 | 17/22 | 0.7s | default | 2026-05-26 |
| MiniCPM5-1B | 0.644 | n=1 | 10/22 | 6.2s | default | 2026-05-27 |
Per-case breakdown: coding_tasks.jsonl (22 cases; full prompt texts withheld to keep the benchmark uncontaminated)
| Case id | Category | Automatic checks | Judge scale |
|---|---|---|---|
| 01_bash_backup_recent_html | bash_ops | syntax:bash, must_include(3), must_include_any(3) | 1-5 |
| 02_fish_fleet_uptime | bash_ops | syntax:fish, must_include(3), must_include_any(2) | 1-5 |
| 03_bash_safe_rsync_deploy | bash_ops | syntax:bash, must_include(3), must_include_any(2) | 1-5 |
| 04_py_openai_compatible_call | python_automation | syntax:python, must_include(3), must_not_include(2) | 1-5 |
| 05_py_factorial_constraints | instruction_following | syntax:python, must_include(1), must_not_include(2) | 1-5 |
| 06_py_retry_backoff_decorator | python_automation | syntax:python, must_include(2), must_include_any(3) | 1-5 |
| 07_py_concurrent_endpoint_ping | python_automation | syntax:python, must_include(2), must_include_any(3) | 1-5 |
| 08_py_parse_nvidia_smi_power | python_automation | syntax:python, must_include(3) | 1-5 |
| 09_py_phone_regex | python_automation | syntax:python, must_include(2), must_include_any(2) | 1-5 |
| 10_systemd_timer_oncalendar | config_edit | must_include(3), must_not_include(1) | 1-5 |
| 11_yaml_add_fleet_tier | config_edit | must_include(5) | 1-5 |
| 12_htaccess_301_https | config_edit | must_include(4), must_include_any(3) | 1-5 |
| 13_php_canonical_tag | python_automation | must_include(2), must_include_any(2) | 1-5 |
| 14_sql_top_referring_domains | sql | must_include(5), must_include_any(4) | 1-5 |
| 15_debug_keyerror | debug_fix | syntax:python, must_include(1), must_include_any(4) | 1-5 |
| 16_debug_ollama_cpu_only | debug_fix | must_include_any(6) | 1-5 |
| 17_refactor_bash_loop | refactor | syntax:bash, must_include(3), must_not_include(1) | 1-5 |
| 18_refactor_preserve_signature | instruction_following | syntax:python, must_include(1) | 1-5 |
| 19_git_branch_commit_push | instruction_following | syntax:bash, must_include(3), must_include_any(2) | 1-5 |
| 20_py_argparse_cli | python_automation | syntax:python, must_include(4) | 1-5 |
| 21_multifile_config_not_loaded | multi_file_reasoning | syntax:python, must_include(2), must_include_any(3) | 1-5 |
| 22_py_idempotent_insert_guard | python_automation | syntax:python, must_include(2), must_include_any(4) | 1-5 |
Coding v2: 20-Case Hard Set
| Model | Score | Runs (range) | Pass | Latency | Temp | Measured |
|---|---|---|---|---|---|---|
| Claude Opus-4.8 (cloud reference) BEST | 0.974 | n=1 | 20/20 | 7.7s | default | 2026-07-06 |
| Ornith-1.0-397B (FP8) | 0.971 | n=3 (0.966-0.977) | 20/20 | 13.0s | default | 2026-07-05 |
| GLM-5.2 FP8 (Z.ai serving, reference) | 0.952 | n=3 (0.927-0.965) | 19/20 | 78.5s | default | 2026-07-06 |
| GLM-5.2 (NVFP4) | 0.942 | n=3 (0.918-0.968) | 20/20 | 147.6s | default | 2026-07-06 |
| Qwen3-Coder-Next-80B | 0.939 | n=3 (0.930-0.950) | 20/20 | 3.0s | default | 2026-07-05 |
| GPT-5.4 (cloud reference) | 0.935 | n=1 | 20/20 | 4.0s | default | 2026-07-06 |
| MiniMax-M3 (MXFP8) | 0.931 | n=3 (0.923-0.936) | 18/20 | 30.0s | default | 2026-07-06 |
| GPT-5.2-Codex (cloud reference) | 0.928 | n=1 | 19/20 | 7.1s | default | 2026-07-06 |
| MiniMax-M2.7 (NVFP4) | 0.901 | n=1 | 19/20 | 48.8s | default | 2026-07-05 |
| Qwen3.6-35B-A3B | 0.892 | n=3 (0.879-0.903) | 18/20 | 27.4s | default | 2026-07-06 |
Per-case breakdown: coding_tasks_v2.jsonl (20 cases; full prompt texts withheld to keep the benchmark uncontaminated)
| Case id | Category | Automatic checks | Judge scale |
|---|---|---|---|
| v2_01_interval_merge_edge | algorithms | syntax:python, must_include(3), must_include_any(2) | 1-5 |
| v2_02_toposort_cycle | algorithms | syntax:python, must_include(4), must_include_any(3) | 1-5 |
| v2_03_lru_ttl | algorithms | syntax:python, must_include(6), must_not_include(2) | 1-5 |
| v2_04_debug_mutable_default | hard_debug | syntax:python, must_include(2), must_include_any(4), must_not_include(1) | 1-5 |
| v2_05_debug_tz_dst | hard_debug | syntax:python, must_include(2), must_include_any(4), must_not_include(1) | 1-5 |
| v2_06_debug_subprocess_deadlock | hard_debug | syntax:python, must_include(4), must_not_include(1) | 1-5 |
| v2_07_debug_async_gather | hard_debug | syntax:python, must_include(2), must_include_any(2) | 1-5 |
| v2_08_spec_versioned_config | spec_compliance | syntax:python, must_include(4), must_not_include(3) | 1-5 |
| v2_09_spec_redact_logger | spec_compliance | syntax:python, must_include(6), must_include_any(2) | 1-5 |
| v2_10_spec_atomic_write | spec_compliance | syntax:python, must_include(5), must_not_include(2) | 1-5 |
| v2_11_spec_cli_exitcodes | spec_compliance | syntax:python, must_include(4), must_include_any(2), must_not_include(2) | 1-5 |
| v2_12_multifile_import_cycle | multi_file_reasoning | syntax:python, must_include(3), must_include_any(3) | 1-5 |
| v2_13_multifile_env_precedence | multi_file_reasoning | syntax:python, must_include(2), must_include_any(4) | 1-5 |
| v2_14_multifile_systemd_env | multi_file_reasoning | must_include(3), must_include_any(4) | 1-5 |
| v2_15_multifile_nginx_cache | multi_file_reasoning | must_include(3), must_include_any(4) | 1-5 |
| v2_16_sql_window_dedup | sql_hard | must_include(3), must_include_any(3) | 1-5 |
| v2_17_sql_upsert_counter | sql_hard | must_include(5), must_include_any(2) | 1-5 |
| v2_18_bash_parallel_safe | infra_hard | syntax:bash, must_include(4), must_include_any(4) | 1-5 |
| v2_19_podman_quadlet | infra_hard | must_include(7), must_include_any(2) | 1-5 |
| v2_20_gitops_hotfix | infra_hard | syntax:bash, must_include(6), must_include_any(3) | 1-5 |
Research and Reasoning: 27 Cases
| Model | Score | Runs (range) | Pass | Latency | Temp | Measured |
|---|---|---|---|---|---|---|
| Qwen3.6-27B (dense) BEST | 0.963 | n=1 | 26/27 | 37.6s | default | 2026-06-30 |
| Qwen3-Coder-Next-80B | 0.963 | n=1 | 26/27 | 2.0s | default | 2026-07-05 |
| Gemma-4-31B (BF16) | 0.963 | n=1 | 26/27 | 11.7s | default | 2026-06-09 |
| Gemma-4-26B-A4B (BF16) | 0.963 | n=1 | 26/27 | 2.2s | default | 2026-06-09 |
| Nex-N2-Pro (NVFP4) | 0.963 | n=1 | 26/27 | 3.2s | default | 2026-06-10 |
| DeepSeek-V4-Flash-DSpark (2x GB10) | 0.963 | n=1 | 26/27 | 25.9s | default | 2026-06-30 |
| Ornith-1.0-397B (FP8) | 0.963 | n=1 | 26/27 | 22.9s | default | 2026-07-05 |
| MiniMax-M3 (MXFP8) | 0.963 | n=1 | 26/27 | 15.6s | default | 2026-07-05 |
| gemma-4-12b-nvfp4 | 0.963 | n=1 | 26/27 | 5.0s | default | 2026-07-15 |
| gemma-4-26b-a4b-nvfp4 | 0.963 | n=1 | 26/27 | 3.5s | default | 2026-07-15 |
| DeepSeek-V4-Flash | 0.926 | n=1 | 25/27 | 15.1s | default | 2026-05-27 |
| Claude Opus-4.7 | 0.926 | n=1 | 25/27 | 9.5s | default | 2026-05-25 |
| Qwen3.6-27B BF16 (dense) | 0.926 | n=1 | 25/27 | 161.2s | default | 2026-06-09 |
| Gemma-4-12B (BF16) | 0.926 | n=1 | 25/27 | 5.6s | default | 2026-06-09 |
| Gemma-4-e4b (BF16) | 0.926 | n=1 | 25/27 | 2.7s | default | 2026-06-09 |
| GLM-5.2 (NVFP4) | 0.926 | n=1 | 25/27 | 46.1s | default | 2026-07-05 |
| GLM-5.2 FP8 (Z.ai serving, reference) | 0.926 | n=1 | 25/27 | 35.0s | default | 2026-07-06 |
| Qwen3.8-27B (FP8) | 0.926 | n=1 | 25/27 | 719.9s | default | 2026-08-14 |
| Granite-4.1-8B | 0.889 | n=1 | 24/27 | 2.6s | default | 2026-06-09 |
| Qwen3.6-35B-A3B (BF16) | 0.889 | n=1 | 24/27 | 24.2s | default | 2026-06-08 |
| devstral-small-2-24b | 0.889 | n=1 | 24/27 | 3.1s | default | 2026-07-03 |
| Gemma-4-e2b | 0.852 | n=1 | 23/27 | 12.2s | default | 2026-05-26 |
| Gemma-4-31B | 0.852 | n=1 | 23/27 | 43.2s | default | 2026-05-27 |
| Qwen3.5-Opus-distill (27B) | 0.852 | n=1 | 23/27 | 270.5s | default | 2026-05-27 |
| Ministral-3-8B | 0.852 | n=1 | 23/27 | 11.5s | default | 2026-05-27 |
| Ornith-1.0-35B (FP8) | 0.852 | n=1 | 23/27 | 8.8s | default | 2026-07-02 |
| Ornith-1.0-397B (W4A16 int4) | 0.852 | n=1 | 23/27 | 26.5s | default | 2026-07-05 |
| MiniMax-M2.7 (NVFP4) | 0.852 | n=1 | 23/27 | 12.2s | default | 2026-07-05 |
| glm-5.3-flash | 0.852 | n=1 | 23/27 | 13.1s | 0.2 | 2026-08-26 |
| glm-5.2-nvfp4 | 0.852 | n=1 | 23/27 | 146.1s | 0.2 | 2026-08-26 |
| Qwen3.6-35B-A3B | 0.815 | n=1 | 22/27 | 1.7s | default | 2026-06-02 |
| GLM-4.5-Air | 0.815 | n=1 | 22/27 | 121.8s | default | 2026-05-28 |
| Qwen3-Coder-30B | 0.815 | n=1 | 22/27 | 1.7s | default | 2026-06-09 |
| Muse-Glimmer-30B (Meta) | 0.815 | n=1 | 22/27 | 33.8s | default | 2026-08-15 |
| Gemma-4-e4b | 0.778 | n=1 | 21/27 | 20.0s | default | 2026-05-26 |
| North-Mini-Code-1.0 (FP8, Cohere) | 0.778 | n=1 | 21/27 | 6.9s | default | 2026-06-20 |
| nemotron-3.5-lightning-30b-a3b-nvfp4 | 0.778 | n=1 | 21/27 | 4.8s | default | 2026-08-11 |
| Mistral-Small-4-119B | 0.741 | n=1 | 20/27 | 11.0s | default | 2026-05-26 |
| Mistral-Small-24B | 0.741 | n=1 | 20/27 | 4.2s | default | 2026-06-09 |
| Gemma-4-26B-A4B | 0.556 | n=1 | 15/27 | 46.4s | default | 2026-05-26 |
| Heretic-9B | 0.000 | n=1 | 0/27 | - | default | 2026-05-27 |
Per-case breakdown: research_tasks.jsonl (27 cases; full prompt texts withheld to keep the benchmark uncontaminated)
| Case id | Category | Automatic checks | Judge scale |
|---|---|---|---|
| summ_01 | long_doc_summarization | must_include(4), must_not_include(3) | 1-5 |
| summ_02 | long_doc_summarization | must_include(3), must_not_include(2) | 1-5 |
| summ_03 | long_doc_summarization | must_include(4), must_not_include(2) | 1-5 |
| summ_04 | long_doc_summarization | must_include(3), must_not_include(2) | 1-5 |
| summ_05 | long_doc_summarization | must_include(4), must_not_include(2) | 1-5 |
| summ_06 | long_doc_summarization | must_include(3), must_not_include(2) | 1-5 |
| summ_07 | long_doc_summarization | must_include(7), must_not_include(2) | 1-5 |
| mhqa_01 | multi_hop_qa | must_include(3), must_include_any(3), must_not_include(2) | 1-5 |
| mhqa_02 | multi_hop_qa | must_include(2), must_not_include(4) | 1-5 |
| mhqa_03 | multi_hop_qa | must_include(2), must_include_any(5), must_not_include(2) | 1-5 |
| mhqa_04 | multi_hop_qa | must_include(3), must_include_any(6) | 1-5 |
| mhqa_05 | multi_hop_qa | must_include(2), must_include_any(3), must_not_include(1) | 1-5 |
| mhqa_06 | multi_hop_qa | must_include_any(7), must_not_include(2) | 1-5 |
| mhqa_07 | multi_hop_qa | must_include(2), must_include_any(4) | 1-5 |
| cite_01 | citation_accuracy | must_include(3), must_include_any(2), must_not_include(4) | 1-5 |
| cite_02 | citation_accuracy | must_include_any(6), must_not_include(4) | 1-5 |
| cite_03 | citation_accuracy | must_include(1), must_include_any(2), must_not_include(3) | 1-5 |
| cite_04 | citation_accuracy | must_include_any(7), must_not_include(4) | 1-5 |
| cite_05 | citation_accuracy | must_include(4), must_include_any(2), must_not_include(3) | 1-5 |
| cite_06 | citation_accuracy | must_include_any(9), must_not_include(4) | 1-5 |
| code_01 | code_generation | syntax:bash, must_include(5), must_include_any(2), must_not_include(1) | 1-5 |
| code_02 | code_generation | syntax:python, must_include(6), must_include_any(2), must_include_all(3) | 1-5 |
| code_03 | code_generation | syntax:fish, must_include(5), must_include_any(2), must_not_include(2) | 1-5 |
| code_04 | code_generation | syntax:python, must_include(10), must_include_any(3), must_not_include(4) | 1-5 |
| code_05 | code_generation | syntax:bash, must_include(5), must_include_any(2), must_not_include(2) | 1-5 |
| code_06 | code_generation | syntax:python, must_include(7), must_include_any(2), must_not_include(1) | 1-5 |
| code_07 | code_generation | syntax:bash, must_include(10), must_include_any(2), must_not_include(1) | 1-5 |
Blog Writing: 10-Criteria Rubric
| Model | Score | Runs (range) | Words | Judge | Gen speed | Temp | Measured |
|---|---|---|---|---|---|---|---|
| DeepSeek-V4-Flash BEST | 1.000 | n=2 | 3774 | - | 64s | 0.3 | 2026-05-28 |
| Qwen3.6-27B BF16 (dense) | 1.000 | n=1 | 3829 | 10/10 | 477s 28 t/s | default | 2026-06-08 |
| Nex-N2-Pro (NVFP4) | 1.000 | n=1 | 4411 | 10/10 | 95s 91 t/s | default | 2026-06-10 |
| Qwen3.6-27B (dense) | 1.000 | n=1 | 3305 | 10/10 | 81s 98 t/s | default | 2026-06-30 |
| qwen3.6-27b-bf16 | 1.000 | n=3 (1.000-1.000) | 3157 | 10/10 | 148s 60 t/s | default | 2026-07-02 |
| Ornith-1.0-397B (FP8) | 1.000 | n=3 (1.000-1.000) | 3857 | 10/10 | 64s 120 t/s | default | 2026-07-05 |
| GLM-5.2 (NVFP4) | 1.000 | n=1 | 3439 | 10/10 | 207s 40 t/s | default | 2026-07-06 |
| MiniMax-M3 (MXFP8) | 1.000 | n=3 (1.000-1.000) | 3317 | 10/10 | 46s 123 t/s | default | 2026-07-05 |
| Claude Opus-4.8 (cloud reference) | 1.000 | n=1 | 3297 | 10/10 | 125s 77 t/s | default | 2026-07-06 |
| Claude Fable 5 (cloud reference) | 1.000 | n=1 | 3630 | 10/10 | 135s 81 t/s | default | 2026-07-06 |
| Ornith-1.0-35B (FP8) | 0.963 | n=3 (0.945-1.000) | 2210 | 10/10 | 66s 166 t/s | default | 2026-07-02 |
| gpt-oss-120b | 0.945 | n=2 | 1938 | - | 132s | 0.0 | 2026-05-28 |
| Gemma-4-31B (BF16) | 0.945 | n=1 | 2390 | 10/10 | 169s 24 t/s | default | 2026-06-09 |
| Gemma-4-12B (BF16) | 0.945 | n=1 | 2291 | 10/10 | 78s 53 t/s | default | 2026-06-09 |
| North-Mini-Code-1.0 (FP8, Cohere) | 0.945 | n=1 | 2269 | 10/10 | 43s 139 t/s | default | 2026-06-20 |
| DeepSeek-V4-Flash-DSpark (2x GB10) | 0.945 | n=1 | 2646 | 10/10 | 131s 46 t/s | default | 2026-06-30 |
| Laguna-S-2.1-NVFP4 | 0.945 | n=1 | 3255 | 10/10 | 55s 106 t/s | default | 2026-07-26 |
| Ornith-1.0-397B (W4A16 int4) | 0.945 | n=3 (0.944-0.945) | 3251 | 9/10 | 70s 104 t/s | default | 2026-07-03 |
| GLM-4.7-Flash | 0.945 | n=2 | 2640 | - | 59s | 0.0 | 2026-05-28 |
| Qwen3.6-35B-A3B (BF16) | 0.944 | n=1 | 3293 | 9/10 | 49s 171 t/s | default | 2026-06-08 |
| GLM-5.2 (IQ2_M 2-bit) | 0.944 | n=1 | 4127 | 9/10 | 199s 39 t/s | default | 2026-06-20 |
| GPT-5.4 (cloud reference) | 0.944 | n=1 | 3246 | 9/10 | 64s 102 t/s | default | 2026-07-06 |
| MiniMax-M2.7 (NVFP4) | 0.907 | n=3 (0.833-1.000) | 3966 | 9/10 | 50s 129 t/s | default | 2026-07-05 |
| Qwen3.6-35B-A3B (FP8) | 0.889 | n=2 | 2597 | - | 38s | 0.0 | 2026-05-28 |
| Gemma-4-26B-A4B | 0.889 | n=1 | 2280 | 9/10 | 128s 44 t/s | default | 2026-05-26 |
| Gemma-4-31B | 0.889 | n=1 | 2063 | 9/10 | 553s 7 t/s | default | 2026-05-27 |
| Gemma-4-26B-A4B (BF16) | 0.889 | n=1 | 2128 | 10/10 | 26s 147 t/s | default | 2026-06-09 |
| Gemma-4-e4b (BF16) | 0.889 | n=1 | 2676 | 10/10 | 40s 122 t/s | default | 2026-06-09 |
| Laguna-XS-2.1-NVFP4 | 0.889 | n=1 | 1703 | 9/10 | 30s 192 t/s | default | 2026-07-26 |
| glm-5.3-flash | 0.889 | n=1 | 3203 | 8/10 | 25s 216 t/s | 0.4 | 2026-08-26 |
| Qwen3-Coder-Next-80B | 0.870 | n=3 (0.833-0.945) | 3210 | 9/10 | 36s 180 t/s | default | 2026-07-05 |
| gemma-4-26b-a4b-nvfp4 | 0.870 | n=3 (0.833-0.945) | 1827 | 9/10 | 34s 97 t/s | default | 2026-07-15 |
| Qwen3.6-27B-NVFP4 (unsloth, dense) | 0.855 | n=5 (0.333-1.000) | 3211 | 9/10 | 89s 98 t/s | default | 2026-07-12 |
| Gemma-4-e2b | 0.833 | n=1 | 2864 | 9/10 | 80s 78 t/s | default | 2026-05-26 |
| command-a-plus | 0.833 | n=1 | 1953 | 9/10 | 103s 49 t/s | default | 2026-05-26 |
| Qwen3.5-Opus-distill (27B) | 0.833 | n=1 | 4197 | 9/10 | 240s 37 t/s | default | 2026-05-26 |
| Laguna-M.1-NVFP4 | 0.833 | n=1 | 1926 | 9/10 | 58s 79 t/s | default | 2026-07-26 |
| nemotron-3.5-lightning-30b-a3b-nvfp4 | 0.833 | n=1 | 3582 | 9/10 | 104s 73 t/s | default | 2026-08-11 |
| glm-5.2-nvfp4 | 0.833 | n=1 | 3533 | 7/10 | 920s 15 t/s | 0.4 | 2026-08-26 |
| Claude Opus-4.7 | 0.805 | n=2 | 2342 | - | 107s | default (~0.0) | 2026-05-28 |
| Qwen3.6-35B-A3B-NVFP4-Fast (unsloth) | 0.800 | n=5 (0.278-0.945) | 2719 | 10/10 | 56s 138 t/s | default | 2026-07-12 |
| devstral-small-2-24b | 0.796 | n=3 (0.778-0.833) | 1307 | 9/10 | 32s 124 t/s | default | 2026-07-03 |
| gpt-oss-20b | 0.778 | n=1 | 1793 | 9/10 | 34s 243 t/s | default | 2026-05-25 |
| GLM-4-9B | 0.778 | n=1 | 898 | 7/10 | 12s 181 t/s | default | 2026-05-25 |
| Mixtral-8x22B | 0.778 | n=1 | 931 | 7/10 | 41s 60 t/s | default | 2026-05-26 |
| Granite-4.1-8B | 0.778 | n=1 | 868 | 9/10 | 13s 169 t/s | default | 2026-06-09 |
| Qwen3.8-27B (FP8) | 0.778 | n=1 | 2522 | 8/10 | 700s 8 t/s | default | 2026-08-14 |
| Mistral-Small-24B | 0.722 | n=1 | 1342 | 7/10 | 39s 74 t/s | default | 2026-06-09 |
| Gemma-4-e4b | 0.722 | n=1 | 2434 | 7/10 | 105s 47 t/s | default | 2026-05-26 |
| Mistral-Small-4-119B | 0.722 | n=1 | 3262 | 8/10 | 281s 28 t/s | default | 2026-05-26 |
| gemma-4-12b-nvfp4 | 0.704 | n=3 (0.333-0.889) | 6138 | 3/10 | 202s 59 t/s | default | 2026-07-15 |
| Mistral-Medium-3.5-128B | 0.611 | n=1 | 1469 | 7/10 | 450s 18 t/s | default | 2026-05-26 |
| GLM-4.5-Air | 0.500 | n=1 | 4541 | 3/10 | 1360s 6 t/s | default | 2026-05-25 |
| nemotron-3-nano:30b | 0.445 | n=1 | 5618 | 2/10 | 156s 51 t/s | default | 2026-05-25 |
| Qwen3.6-35B-A3B | 0.278 | n=1 | 195 | 3/10 | 212s 38 t/s | default | 2026-05-25 |
| Mixtral-8x7B | 0.222 | n=1 | 79 | 2/10 | 1s 137 t/s | default | 2026-05-26 |
| Qwen3-Coder-30B | 0.111 | n=1 | 0 | 1/10 | 11s 0 t/s | default | 2026-06-09 |
Blog Temperature Sweep
| Model | Temp | Mean score | Range (min-max) | Mean words | Mean gen |
|---|---|---|---|---|---|
| Claude Opus-4.7 | default (~0.0) | 0.805 | 0.778-0.833 | 2342 | 107.2s |
| Claude Opus-4.7 | default (~0.3) | 0.805 | 0.778-0.833 | 2342 | 107.2s |
| Claude Opus-4.7 | default (~0.7) | 0.805 | 0.778-0.833 | 2342 | 107.2s |
| Claude Opus-4.7 | default (~1.0) | 0.805 | 0.778-0.833 | 2342 | 107.2s |
| DeepSeek-V4-Flash | 0.0 | 0.972 | 0.944-1.000 | 4194 | 78.4s |
| DeepSeek-V4-Flash | 0.3 BEST | 1.000 | 1.000-1.000 | 3774 | 64.2s |
| DeepSeek-V4-Flash | 0.7 | 1.000 | 1.000-1.000 | 3732 | 65.0s |
| DeepSeek-V4-Flash | 1.0 | 0.972 | 0.944-1.000 | 4025 | 72.4s |
| GLM-4.7-Flash | 0.0 BEST | 0.945 | 0.944-0.945 | 2640 | 58.8s |
| GLM-4.7-Flash | 0.3 | 0.805 | 0.722-0.889 | 2210 | 53.1s |
| GLM-4.7-Flash | 0.7 | 0.861 | 0.833-0.889 | 2748 | 57.9s |
| GLM-4.7-Flash | 1.0 | 0.889 | 0.889-0.889 | 2072 | 49.5s |
| Qwen3.6-35B-A3B (FP8) | 0.0 BEST | 0.889 | 0.889-0.889 | 2597 | 37.5s |
| Qwen3.6-35B-A3B (FP8) | 0.3 | 0.806 | 0.667-0.945 | 3204 | 39.8s |
| Qwen3.6-35B-A3B (FP8) | 0.7 | 0.584 | 0.445-0.722 | 2517 | 44.2s |
| Qwen3.6-35B-A3B (FP8) | 1.0 | 0.611 | 0.611-0.611 | 2549 | 41.5s |
| gpt-oss-120b | 0.0 BEST | 0.945 | 0.945-0.945 | 1938 | 131.8s |
| gpt-oss-120b | 0.3 | 0.889 | 0.889-0.889 | 2145 | 137.4s |
| gpt-oss-120b | 0.7 | 0.945 | 0.945-0.945 | 2144 | 127.8s |
| gpt-oss-120b | 1.0 | 0.917 | 0.889-0.945 | 1710 | 124.2s |
How the Scores Are Produced
bash -n, python compile, fish -n) + 0.4 x a blind LM-judge
correctness score (1-5; the judge sees only task and solution, never the model's identity).
Deterministic decoding (temp 0.0). The leaderboard number is the mean across all 22 cases.Findings
What does a 4-way GB10 fabric buy, and what does it take to use it? Four GB10 units on a switched 200G RoCE fabric (two MikroTik CRS812; 196 Gb/s node-to-node RDMA measured through the switch with no PFC/ECN tuning) pool 512 GB. Qwen3.8-Flash-Next NVFP4 (125B-A6B plus a 51B n-gram table) split four ways decodes 39.6 / 57.0 / 76.9 tokens/s on prose / code / math single-stream, 3,550 tokens/s prefill, and 167-184 tokens/s aggregate at six streams with a 2.6M-token KV cache; the same model on a two-unit pair does 35 / 36 / 53 and 118 aggregate. GLM-5.3-Flash NVFP4 (18B active) on the same four units: 26.5 / 38.1 / 46.7, prefill 3,840, 91 aggregate. The platform is memory-bandwidth-bound, so the smaller-active model wins decode and the fabric is never the limiter (11-17 GB of RoCE traffic per run). Three practical findings. Plain 4-way tensor parallelism cannot load the NVFP4 MoE (expert intermediate width 640 does not divide four ways for the 4-bit kernels): expert parallelism is required, and using both 100G halves of each node's link adds 5-15% (EP all-to-all benefits; GLM's all-reduce does not). MTP-4 is the speculative sweet spot: MTP-6 (3.6 of 6 drafts accepted) and piecewise CUDA graphs (25.5 / 37.1 / 44.6, KV cache shrank 6x) both lose, and a DFlash2 block-draft head lifts GLM to 69 tokens/s on math single-stream but halves the 6-stream aggregate (acceptance collapses to about 2). One GB10 can serve the 125B model alone by memory-mapping the n-gram table from NVMe (24.8 / 31.1 / 33.8, 86 aggregate), so four independent units give about 345 tokens/s aggregate versus 184 pooled. Ops caveat that cost a night: a gpu-memory-utilization of 0.85 leaves 1-4% of host memory free on a 121 GB unified-memory unit and an earlyoom daemon will silently kill the worker; 0.65 (4-way) and 0.72-0.78 (single) with a memory guard are the working floors. Throughput only; fabric quality evaluations pending.
GLM-5.3-Flash: what replaced GLM-5.2 on the H200 quad, and its two traps. zai-org/GLM-5.3-Flash (MIT, 320B-A18B, hybrid KDA + NoPE sparse MLA, native FP8, 1M ctx) leads or ties every panel-graded axis: coding 0.9955 (ties #1 of 51 panel rows), research 0.912 (#1 of 21); campaign-judge coding 0.987 (n=3). On 4x NVLink H200: prefill 10.3k -> 20.6k tok/s from 1.5k to 23k context, decode math 339 / prose 246 / code 233 t/s, KV 2.86M tokens (21.8x concurrency at 131k). A 320B model decoding like a 35B. Trap 1: thinking is on by default and gated by reasoning_effort (low | high | max), not a Qwen-style enable_thinking flag; at default effort a long-form request can return empty content with finish_reason=length. Trap 2: free-form summarization fabricates (long_doc_summarization 0.736, judges penalize invented cost bands, dates and regulatory claims) while citation_accuracy is 0.991; use it grounded, never as a free summarizer of compliance material. Launch traps: the vendor's fp8 KV cache is invalid on Hopper (BF16 KV required); --model must be a local path. GLM-5.2 was deleted from the fleet the same day.
Does llama.cpp on a GB10 have a speculative path? Yes, since mainline merged MTP drafting (ggml-org PR #22673: --spec-type draft-mtp --spec-draft-n-max 3). The forum claim of ~27 t/s for a Q4_K_XL GGUF on GB10 was real and used it. Within dense-on-GB10, llama.cpp Q4_K_XL + MTP (20-25 t/s) now beats vLLM FP8 + MTP (12.2 t/s). Trap: --spec-type draft-dspark on a plain GGUF is a silent no-op (healthy server, zero log lines, baseline speed); verify the "creating MTP draft context" log line before believing any spec flag.
Why did a cabled 4-way GB10 cluster get slower, and what is a GB10 actually good at? Direct-attached (two ConnectX-7 ports per node), four units in TP=4 on a 27B dense FP8 model decoded 3.5 t/s and prefilled 539 tok/s versus 8.3 t/s and 1,963 tok/s on ONE unit: all-to-all RDMA is physically impossible with two ports (three links per node needed), so a 4-way cluster needs a switch, not cables. That switch arrived on 2026-08-28 (fabric card above). Also found: one card's firmware had throttled RDMA to 12.74 Gb/s while every status tool reported healthy; only a bandwidth test finds it (111.86 Gb/s after apt full-upgrade + reboot). Re-measure after every GB10 rebuild. The durable number: prefill:decode on one GB10 is 237:1 (1,963 vs 8.3), and decode lands within 12% of the LPDDR5X prediction (273 GB/s / 29 GB = 9.4 t/s). Routing rule: long-prompt, short-output work (RAG, summarization, classification, reranking) to the GB10; long generation to the RTX PRO 6000.
How much does sampling config move SWE-bench? DeepSeek-V4-Flash-0731 at vendor-spec sampling scored 75.3% (113/150) on the stratified SWE-bench Verified set on 2x H200, tying GLM-5.2's 76.0% (itself a temperature-0, off-spec run). The temperature-0 arm of the same model collapsed to 2.7% usable patches. A QA audit found the earlier temp-0 SWE series (Qwen 55.3, Opus 82.0 anchors, DeepSeek 67.3) was accidental and inconsistent across models; treat any agentic-coding number without its sampling config as non-reproducible.
Does speculative decoding cost quality, and how much speed does block-drafting buy? Block-draft speculative decoding (DSpark) took a Nemotron-3.5 MoE on a GB10 node from 75.1 to 184.0 tokens/s (2.45x), beating standard multi-token prediction by 79% on the same hardware; a mixture-of-experts model also beat a dense model of similar quality tier by about 5.5x on this bandwidth-bound platform, so dense models are retired from the GB10 tier entirely. Quality is a separate question and the answer is per-model: against purpose-built controls that change ONLY the speculative flags, Nemotron-3.5 loses about 1.8 points (three-judge panel) under both MTP and DSpark, while Qwen3.6-35B-A3B (0.977 vs its 0.975 baseline, N=2) and Qwen3.8-27B (0.984 DSpark vs 0.979 control, deltas inside run spread) show no measurable cost. Vendor claims of "lossless by construction" should be treated as a hypothesis to test per model, with a control that changes one variable. One measurement caveat that moved every published number: decode speed depends strongly on content class - the same DSpark config measured 184.0 tokens/s on math, 120.6 on code, and 105.0 on prose - so a decode rate without its content class understates or overstates by up to 75%.
Does 4-bit quantization cost quality? It depends on the format, not the bit-width. NVFP4 (W4A4 with lm_head, embeddings, and speculative heads kept in BF16) statistically TIES FP8 on the 22-case coding set under a three-judge frontier panel (0.9795 vs 0.9762 on identical prompts, N=2), while naive 4-bit formats of the same model lose 4-7 points (Q4_K_XL GGUF 0.934, a groupwise-int4 engine 0.909). The earlier rule of thumb "4-bit costs about 4 points of coding quality" is true for naive 4-bit and false for NVFP4. Practical read for 96 GB-class cards: NVFP4 buys FP8 quality at roughly half the weight memory, and the freed VRAM buys context and co-tenancy - decode speed itself is bandwidth-bound and does not improve.
Is Gemma-4-26B-A4B better than Qwen3.6, and can it run in the agent gateway? It depends on task difficulty, judged by a three-model frontier panel (GPT-5.4 + Claude Opus-4.8 + GLM-5.2, the panel in force until DeepSeek-V4 took the third seat in August 2026; median score on identical generations, N=5). On easy/routine coding Qwen wins (dense 27B 0.983, 35B-FP8 0.979 vs Gemma 0.962); on the hard coding set Gemma wins decisively (0.915 vs every Qwen variant 0.78-0.81), corroborated by the deterministic auto-checks. Research and blog are ties (about 0.86 and 0.79). Agentic: served with the correct vLLM tool parser (gemma4, not pythonic - the wrong one produces zero tool calls), Gemma is the strongest tool-caller (55 correct calls vs Qwen's 51) and solves every multi-turn task (pass@max 1.00 vs Qwen-FP8 0.778). Both models were verified end-to-end through the real agent gateway - writing and executing code and returning correct results - which requires serving the model at 131k context (the agent system prompt is about 58k tokens). NVFP4 is near-lossless for Gemma-4 but not for Qwen3.6-35B (a long-form runaway appears at FP4; use FP8).
Does GLM-5.2 lose anything at FP4, and is the six-GPU FP8 layout justified? No measurable loss. GLM-5.2 at FULL FP8 fidelity (Z.ai's own serving, via OpenRouter) scored 0.952 on the v2 hard set (n=3: 0.927/0.964/0.965) and 0.92 on research - versus the FP4 quant on the fleet's 4-way bridge at 0.942 and 0.94. The FP8 delta (+0.010 coding, -0.02 research) is inside the 0.03 tie threshold: FP4 is effectively lossless for GLM-5.2 on these tests, and both configurations score below Ornith-1.0-397B-FP8 (0.971). Implication for fleet layout: an FP8 deployment spanning six GPUs and both NVLink bridges buys no measurable quality over FP4 on four GPUs.
Do two independent judges agree on these scores? Yes. A dual-judge audit re-graded 40 stored blog outputs with DeepSeek-V4-Flash (served on the DGX Spark pair) against their original GPT-4.1 scores using the identical rubric: mean absolute difference 0.45 points on the 1-10 scale, 87.5% of pairs within one point, and the second judge ran slightly STRICTER on average (signed mean -0.3) with its five largest disagreements spread across Anthropic, MiniMax, and Qwen outputs - no family favoritism pattern. This is the periodic cross-judge integrity check the methodology commits to.
How does the fleet compare to the July 2026 cloud frontier on the same tests? Cloud reference rows added: Claude Opus-4.8 v1 0.995 / v2 0.974 / blog 1.000 (GPT-4.1 judge); GPT-5.4 v1 0.970 / v2 0.935 / blog 0.944 and GPT-5.2-Codex v1 0.982 / v2 0.928 (both Haiku-judged per the cross-family gate; OpenAI rows and Anthropic/fleet rows use different judges, so treat cross-vendor deltas under 0.05 with care). Read: fleet-owned Ornith-1.0-397B-FP8 (v2 0.971, zero marginal cost, CUI-capable) sits between GPT-5.4 (0.935) and Opus-4.8 (0.974) on the hard set. Also measured: Claude Fable 5's API safety layer CONTENT-FILTERED 9 of 22 routine infrastructure tasks (finish_reason=content_filter after ~3 tokens: rsync deploy scripts, a phone-number regex, an argparse CLI), making it unusable as a coding baseline here; its blog run scored 1.000. GPT-5.2-Codex is Responses-API-only (the harness gained --use-responses-api).
Do the bridge picks survive a harder test set and multi-topic writing? Yes, and the ordering sharpens. The 20-case v2 coding set (built because the original suite saturated) spreads the leaders decisively: Ornith-1.0-397B-FP8 0.971 (n=3, 20/20 every rep, 13 s per case), GLM-5.2-NVFP4 0.942 (n=3, high variance 0.918-0.968 at 150 s per case), Qwen3-Coder-Next 0.939 (n=3, 2.6 s per case), MiniMax-M3 0.931 (n=3, drops 1-2 cases per rep), incumbent Qwen3.6-35B-A3B 0.892 (n=3). MiniMax-M3's v1 crown (0.984) inverts on v2: the saturated suite was measuring judge quibbles, not capability. Blog re-tested at 3 topics x 3 reps: Ornith-397B a perfect 1.000 on all nine runs; MiniMax-M3 0.975 mean (short only on the CMMC brief); MiniMax-M2.7 0.94 mean (its earlier single-run 1.000 was variance); Qwen3-Coder-Next 0.887 mean. Qwen3-Coder-Next also posted 85.2% on the 88-case tool corpus (53 correct calls, 24 correct declines, 1 hallucinated call): strong for agentic coding, not voice-grade (gpt-oss-20b remains 0% hallucination). Remaining before any production routing change: a local SWE-bench Verified run for Qwen3-Coder-Next.
Which open model is best on each NVLink island of the 6x H200 fleet? First head-to-head since the 4-way bridge install. On the 4-way bridge (575 GB), Ornith-1.0-397B-FP8 (MIT) is the balanced winner: coding 0.980 (N=3), blog 1.000, research 0.91 at 119.7 t/s single-stream and 639 t/s at c=8, roughly 2x the speed of speed-tuned GLM-5.2-NVFP4 at equal or better quality on two of three axes. MiniMax-M3-MXFP8 posted the campaign's best coding rep (0.991; N=3 mean 0.984) with 1M context and multimodal input, held back only by its non-MIT community license. GLM-5.2-NVFP4 keeps the research crown (0.94, citation accuracy 0.992). On the 2-way bridge (287 GB), Qwen3-Coder-Next-FP8 (80B-A3B) is the coder/agentic pick: 0.977 coding with zero variance across three full reps, clean native tool calls, 178 t/s single-stream and 2,789 t/s aggregate at c=32. Ruled out: Kimi-K2.6 (594 GB INT4 exceeds any island), GLM-5.x-FP8 (754 GB needs all six cards), Nemotron-3-Ultra-NVFP4 (numerically broken on Hopper: loads but emits gibberish). Serving notes: Ornith-397B-FP8 requires
VLLM_TEST_FORCE_FP8_MARLIN=1 plus an explicit chat template; MiniMax-M3 requires
--block-size 128; the 2026-07-03 vLLM nightly has broken block-scaled FP8 kernels on
Hopper (pin v0.24.0).Apple Silicon (MLX) Results
These rows come from the Apple Silicon workload matrix measured 2026-05-26 with MLX (mlx_lm.server 0.31.3), one model resident at a time, by unified-memory tier. Coding is cross-judged by Claude Haiku-4.5 (the generator never equals the judge). The first table is the dense-versus-MoE head-to-head on the 64 GB M1 Max at the same quantization; the rest are per-tier matrices for the 16 GB M4 Mac mini, the 32 GB M5 and the 64 GB M1 Max, followed by the macOS GPU wired-memory ceiling rule. Bench-only: nothing here changes production routing.
Dense vs MoE: the decision rule
| Axis | Dense Llama-70B (4-bit) | MoE Qwen3.6-35B-A3B (4-bit) |
|---|---|---|
| Decode | 6.54 t/s | 54.3 t/s (8.3x faster) |
| Coding (Haiku-judged) | 0.899 | 0.952 |
| Memory footprint | 41 GB | 14.2 GB |
| Tool-call accuracy | n/a (no MLX template) | 97% |
16 GB: M4 Mac mini
| Model | Quant | Size | Fits no-swap | Decode t/s | Coding | Tool-call |
|---|---|---|---|---|---|---|
| Llama-3.1-8B-Instruct BEST FIT | MLX 4-bit | 5 GB | yes | 18.6 | 0.843 | 86.4% |
| Qwen2.5-14B-Instruct | MLX 4-bit | 7.7 GB | tight | 10.3 | 0.915 | n/a* |
| gpt-oss-20b | MXFP4 | 10 GB | edge | 34-39 | 0.69** | n/a* |
| Qwen3.6-35B-A3B | MLX 3-bit | 14.5 GB | OOM | n/a | n/a | n/a |
32 GB: M5 (the value sweet spot)
| Model | Quant | Size | Decode t/s | TTFT | Coding | Tool-call |
|---|---|---|---|---|---|---|
| Qwen3.6-35B-A3B DAILY DRIVER | MLX 4-bit MoE | 21 GB | 50.5 | 0.2-0.5s | 0.931 | 97.0% |
| Qwen2.5-7B-Instruct | MLX 8-bit | 8.1 GB | 15.4 | 0.67s | 0.889 | n/a* |
| Llama-3.1-8B-Instruct | MLX 4-bit | 4.8 GB | 26.3 | 0.69s | 0.829 | 86.4% |
| gpt-oss-20b | MXFP4 | 11 GB | 46.6 | 1.21s | 0.694** | n/a* |
64 GB: M1 Max
| Model | Quant | Size | Decode t/s | Coding | Tool-call |
|---|---|---|---|---|---|
| Qwen3.6-35B-A3B DAILY DRIVER | MLX 4-bit MoE | 14.2 GB | 54.3 | 0.952 | 97.0% |
| Llama-3.3-70B-Instruct | MLX 4-bit dense | 41 GB | 6.54 | 0.899 | n/a |
| gpt-oss-120b | MXFP4 | 59 GB | OOM | n/a | n/a |
The ceiling rule (macOS GPU wired cap about 67-75% of RAM)
| Unified RAM | GPU budget | Largest comfortable model | Hard ceiling |
|---|---|---|---|
| 16 GB (M4) | ~10-11 GB | 8B-4bit (14B tight) | ~10 GB / no 20B+ headroom on a shared box |
| 32 GB (M5) | ~20-24 GB | 30-35B-A3B-4bit MoE | no 70B |
| 64 GB (M1 Max) | ~43-48 GB | 35B MoE (or 70B-4bit single-stream) | no 120B |
About the Author
Talk to the team that measured every number on this page
A 30-minute call is enough to tell you which tier fits your users, your data and your budget, and what it will cost to run.
Petronella Technology Group, Inc. · 5540 Centerview Dr., Suite 200, Raleigh, NC 27606 · 919-348-4912 · info@petronellatech.com · Last Updated: August 26, 2026