RustingStrata
Run Qwen3.6-35B-A3B on a 12 GB GPU faster than llama.cpp. Experts that do not fit in VRAM run on the CPU alongside the GPU, and their placement adapts to the routing it sees. No C++, no ggml — CUDA kernels compile at run time.
$ cargo build --release
$ strata calibrate --model $M
$ strata chat --model $M --gpu
$ strata serve --model $M --gpu --addr 127.0.0.1:8080Read the docs
A pure-Rust inference engine for the Qwen3.5/3.6 hybrid model family (qwen35moe, qwen35) in GGUF. CUDA kernels are compiled at run time with NVRTC through cudarc: no C++ build step, no ggml. Experts that do not fit in VRAM run on the CPU at the same time as the GPU work, and the expert placement adapts to the routing it sees.
Referencellama.cpp baseline (M0)Reference machine: RTX 3060 12 GB (about 1.7 GB used by the desktop), Ryzen 5 7600X, 32 GB RAM, CachyOS. llama.cpp commit 5d8ccdf9d, CUDA build (sm_86), flash attention on, all layers on GPU (-ngl 99). Model: Qwen3.6-35B-A3B-UD-Q2_K_XL.gguf (11.70 GiB).
ReferenceM6 results: v1 gateModel: Qwen3.6-35B-A3B-UD-Q2_K_XL, RTX 3060 12 GB, Ryzen 5 7600X. 01:00-01:14, same session for both engines, everything under the GPU lock (730-790 MiB VRAM used by others before each step). Another loop's rustc builds pushed host load to 21-28 during the first Strata runs; the 512/4K rows were re-run at load 2.6-4.1 and those numbers are the ones below. The 32K rows ran at load 21-24 and are lower bounds.