Skip to content
RustingStrataGitHub

M6 results: v1 gate

RustingStrata pages

v1.0.0 gate run (2026-10-10, 9f4af2a)

Model: Qwen3.6-35B-A3B-UD-Q2_K_XL, RTX 3060 12 GB, Ryzen 5 7600X. 01:00-01:14, same session for both engines, everything under the GPU lock (730-790 MiB VRAM used by others before each step). Another loop’s rustc builds pushed host load to 21-28 during the first Strata runs; the 512/4K rows were re-run at load 2.6-4.1 and those numbers are the ones below. The 32K rows ran at load 21-24 and are lower bounds.

Strata: strata bench --gpu --pp 512,4096 --tg 128 --reps 3 --mtp 0|1 (and --pp 32768 --ctx 33280), mean (min-max) of 3 reps. llama.cpp: llama-bench -ngl 99 -fa 1 -r 3 (CUDA build 5d8ccdf9d), --n-cpu-moe 10 at 512/4K, --n-cpu-moe 14 -d 32768 at 32K, mean +- std, load 4.3-7.5.

TestStrataStrata MTP k = 1llama.cpp
prefill, 512 tokens912.7 (874.8-932.5)910.9 (895.8-932.1)1077.7 +- 13.9
prefill, 4096 tokens2016.3 (2013.0-2019.8)1904.4 (1709.9-2013.1)1059.5 +- 3.4
decode 128 after 512100.7 (100.5-101.0)109.9 (104.5-113.5)61.6 +- 0.5
decode 128 after 409696.0 (95.0-96.6)96.7 (90.7-101.9)—
prefill, 32768 tokens1540.8 (1525.2-1556.1)1536.5 (1474.8-1593.3)808.6 +- 1.2 (pp512 @ d32768)
decode 128 after 3276877.0 (70.4-81.4)79.1 (64.9-86.2)50.4 +- 3.0 (tg128 @ d32768)
  • MTP drafts accepted 62-65%. Short-prompt prefill is still slower than llama.cpp (0.85x): 512 tokens is one chunk and the first rep includes cache warm-up.

Parity

tools/parity.py, 20 prompts x 32 greedy tokens against the token-by-token llama-server reference (target/parity-ref.json, llama.cpp CPU build). For v1 a token counts as agreeing when it matches llama.cpp in either of its own modes: token by token (the reference) or with the context run as one batch (target/llama-batch-argmax.json, tools/selfparity.py). llama.cpp’s two modes agree with each other on only 98.28% of these positions (#33), so a strict 99% gate is above llama.cpp’s own consistency.

ModeStrict (vs reference)Either llama.cpp mode
--gpu --llama-numerics --fixed-experts (parity mode)634 / 633 / 632 of 640 (98.91%)637 / 635 / 635 (99.32%)
--gpu (default, fastest)628 / 630 / 625 (98.07%)633 / 634 / 631 (98.85%)

The default path differs from llama.cpp only in activation rounding and expert placement (the adaptive cache moves experts between GPU int8 and CPU kernels during the run); its misses are near-ties.

Gate

ItemGateResult
Prefill at 4K>= 1017 tok/s2016.3, passed (1.90x llama.cpp)
Decode> 67 tok/s100.7 short / 96.0 after 4K, passed (1.63x llama.cpp); 109.9 with MTP
Parity vs llama.cpp>= 99%99.32% in parity mode (either llama.cpp mode), passed; default path 98.85%

First gate run (2026-10-08, 671050e)

Model: Qwen3.6-35B-A3B-UD-Q2_K_XL, RTX 3060 12 GB, Ryzen 5 7600X (6 cores / 12 threads). 2026-10-08, 21:38-21:46, same session for both engines, nothing else on the GPU (899 MiB used before the run). Host load 6.0 at the start (left over from a llama-server parity job), 1.4-2.4 during the llama-bench runs.

Strata: strata bench --gpu --pp 512,4096 --tg 128 --reps 3 (and --pp 32768 --ctx 33280) at 671050e, mean (min-max) of 3 reps. llama.cpp: llama-bench -ngl 99 -fa 1 -r 3 (CUDA build 5d8ccdf9d), --n-cpu-moe 10 at short and 4K (best setting from docs/baseline.md), --n-cpu-moe 14 -d 32768 at 32K, mean +- std.

TestStrataStrata MTP k = 1llama.cpp
prefill, 512 tokens676.9 (359.6-836.2)853.6 (840.7-861.9)1046.4 +- 23.3
prefill, 4096 tokens1715.6 (1618.4-1765.5)1759.2 (1754.2-1763.4)1048.1 +- 1.9
decode 128 after 51281.7 (80.0-83.3)88.9 (85.7-90.8)69.8 +- 0.2
decode 128 after 409678.3 (76.4-79.4)78.7 (78.4-79.0)—
prefill, 32768 tokens1413.0 (1394.8-1423.1)1418.1 (1411.4-1425.2)775.7 +- 59.1 (pp512 @ d32768)
decode 128 after 3276867.8 (67.5-68.1)64.9 (63.7-65.6)57.9 +- 0.2 (tg128 @ d32768)
  • MTP drafts accepted: 64.1% at 512/4K and at 32K. Expert hit rate 97.1-99.2%.
  • llama.cpp’s 32K numbers are marginal rates at depth 32768 (a 512-token batch and 128 tokens after a filled cache); Strata’s 32K prefill is the average over the whole 32768-token prompt.
  • Short-prompt prefill is slower than llama.cpp: 512 tokens is one chunk, and the first rep includes cache warm-up (359.6 min).

Gate

ItemGateResult
Prefill at 4K>= 1017 tok/s1715.6, passed (1.64x llama.cpp)
Decode> 67 tok/s81.7 short / 78.3 after 4K, passed (1.17x llama.cpp); 88.9 with MTP
Parity vs llama.cpp>= 99%Open: 98.39% mean vs the llama.cpp CPU-build reference (#32), 97.34% vs the CUDA-build reference (#34); llama.cpp scores 98.28% against its own reference (#33)

MTP pays at short context only: +7.2 tok/s after 512 tokens, +0.4 after 4K, -2.9 after 32K.