Skip to content
RustingStrataGitHub
04 / 04
LLM inference / CUDA / Qwen

RustingStrata

Run Qwen3.6-35B-A3B on a 12 GB GPU faster than llama.cpp. Experts that do not fit in VRAM run on the CPU alongside the GPU, and their placement adapts to the routing it sees. No C++, no ggml — CUDA kernels compile at run time.

2,093
tok/s prefill at 4K (llama.cpp: 1,017)
101
tok/s decode (llama.cpp: 65.8)
111
tok/s decode with MTP drafts
Qwen3.6-35B-A3B Q2_K_XL (11.7 GiB) on an RTX 3060 12 GB + Ryzen 5 7600X.
Hybrid CPU/GPU MoEAdaptive expert placementMTP speculative decodingOpenAI-compatible server
runbash
$ cargo build --release
$ strata calibrate --model $M
$ strata chat --model $M --gpu
$ strata serve --model $M --gpu --addr 127.0.0.1:8080
Tutorials

Learn by building

No tutorials yet. Start with the RustingStrata docs.