v1.0.0 gate run (2026-10-10, 9f4af2a)
Model: Qwen3.6-35B-A3B-UD-Q2_K_XL, RTX 3060 12 GB, Ryzen 5 7600X. 01:00-01:14, same session for both engines,
everything under the GPU lock (730-790 MiB VRAM used by others before each step). Another loop’s rustc builds
pushed host load to 21-28 during the first Strata runs; the 512/4K rows were re-run at load 2.6-4.1 and those
numbers are the ones below. The 32K rows ran at load 21-24 and are lower bounds.
Strata: strata bench --gpu --pp 512,4096 --tg 128 --reps 3 --mtp 0|1 (and --pp 32768 --ctx 33280), mean (min-max)
of 3 reps. llama.cpp: llama-bench -ngl 99 -fa 1 -r 3 (CUDA build 5d8ccdf9d), --n-cpu-moe 10 at 512/4K,
--n-cpu-moe 14 -d 32768 at 32K, mean +- std, load 4.3-7.5.
| Test | Strata | Strata MTP k = 1 | llama.cpp |
|---|---|---|---|
| prefill, 512 tokens | 912.7 (874.8-932.5) | 910.9 (895.8-932.1) | 1077.7 +- 13.9 |
| prefill, 4096 tokens | 2016.3 (2013.0-2019.8) | 1904.4 (1709.9-2013.1) | 1059.5 +- 3.4 |
| decode 128 after 512 | 100.7 (100.5-101.0) | 109.9 (104.5-113.5) | 61.6 +- 0.5 |
| decode 128 after 4096 | 96.0 (95.0-96.6) | 96.7 (90.7-101.9) | — |
| prefill, 32768 tokens | 1540.8 (1525.2-1556.1) | 1536.5 (1474.8-1593.3) | 808.6 +- 1.2 (pp512 @ d32768) |
| decode 128 after 32768 | 77.0 (70.4-81.4) | 79.1 (64.9-86.2) | 50.4 +- 3.0 (tg128 @ d32768) |
- MTP drafts accepted 62-65%. Short-prompt prefill is still slower than llama.cpp (0.85x): 512 tokens is one chunk and the first rep includes cache warm-up.
Parity
tools/parity.py, 20 prompts x 32 greedy tokens against the token-by-token llama-server reference
(target/parity-ref.json, llama.cpp CPU build). For v1 a token counts as agreeing when it matches llama.cpp in
either of its own modes: token by token (the reference) or with the context run as one batch
(target/llama-batch-argmax.json, tools/selfparity.py). llama.cpp’s two modes agree with each other on only
98.28% of these positions (#33), so a strict 99% gate is above llama.cpp’s own consistency.
| Mode | Strict (vs reference) | Either llama.cpp mode |
|---|---|---|
--gpu --llama-numerics --fixed-experts (parity mode) | 634 / 633 / 632 of 640 (98.91%) | 637 / 635 / 635 (99.32%) |
--gpu (default, fastest) | 628 / 630 / 625 (98.07%) | 633 / 634 / 631 (98.85%) |
The default path differs from llama.cpp only in activation rounding and expert placement (the adaptive cache moves experts between GPU int8 and CPU kernels during the run); its misses are near-ties.
Gate
| Item | Gate | Result |
|---|---|---|
| Prefill at 4K | >= 1017 tok/s | 2016.3, passed (1.90x llama.cpp) |
| Decode | > 67 tok/s | 100.7 short / 96.0 after 4K, passed (1.63x llama.cpp); 109.9 with MTP |
| Parity vs llama.cpp | >= 99% | 99.32% in parity mode (either llama.cpp mode), passed; default path 98.85% |
First gate run (2026-10-08, 671050e)
Model: Qwen3.6-35B-A3B-UD-Q2_K_XL, RTX 3060 12 GB, Ryzen 5 7600X (6 cores / 12 threads). 2026-10-08, 21:38-21:46, same session for both engines, nothing else on the GPU (899 MiB used before the run). Host load 6.0 at the start (left over from a llama-server parity job), 1.4-2.4 during the llama-bench runs.
Strata: strata bench --gpu --pp 512,4096 --tg 128 --reps 3 (and --pp 32768 --ctx 33280) at 671050e, mean
(min-max) of 3 reps. llama.cpp: llama-bench -ngl 99 -fa 1 -r 3 (CUDA build 5d8ccdf9d), --n-cpu-moe 10 at short
and 4K (best setting from docs/baseline.md), --n-cpu-moe 14 -d 32768 at 32K, mean +- std.
| Test | Strata | Strata MTP k = 1 | llama.cpp |
|---|---|---|---|
| prefill, 512 tokens | 676.9 (359.6-836.2) | 853.6 (840.7-861.9) | 1046.4 +- 23.3 |
| prefill, 4096 tokens | 1715.6 (1618.4-1765.5) | 1759.2 (1754.2-1763.4) | 1048.1 +- 1.9 |
| decode 128 after 512 | 81.7 (80.0-83.3) | 88.9 (85.7-90.8) | 69.8 +- 0.2 |
| decode 128 after 4096 | 78.3 (76.4-79.4) | 78.7 (78.4-79.0) | — |
| prefill, 32768 tokens | 1413.0 (1394.8-1423.1) | 1418.1 (1411.4-1425.2) | 775.7 +- 59.1 (pp512 @ d32768) |
| decode 128 after 32768 | 67.8 (67.5-68.1) | 64.9 (63.7-65.6) | 57.9 +- 0.2 (tg128 @ d32768) |
- MTP drafts accepted: 64.1% at 512/4K and at 32K. Expert hit rate 97.1-99.2%.
- llama.cpp’s 32K numbers are marginal rates at depth 32768 (a 512-token batch and 128 tokens after a filled cache); Strata’s 32K prefill is the average over the whole 32768-token prompt.
- Short-prompt prefill is slower than llama.cpp: 512 tokens is one chunk, and the first rep includes cache warm-up (359.6 min).
Gate
| Item | Gate | Result |
|---|---|---|
| Prefill at 4K | >= 1017 tok/s | 1715.6, passed (1.64x llama.cpp) |
| Decode | > 67 tok/s | 81.7 short / 78.3 after 4K, passed (1.17x llama.cpp); 88.9 with MTP |
| Parity vs llama.cpp | >= 99% | Open: 98.39% mean vs the llama.cpp CPU-build reference (#32), 97.34% vs the CUDA-build reference (#34); llama.cpp scores 98.28% against its own reference (#33) |
MTP pays at short context only: +7.2 tok/s after 512 tokens, +0.4 after 4K, -2.9 after 32K.