# gemma-4-12b on the workstation, against the resident Qwen3-1.7B Measured 2026-08-02 on the fixtures as they stand. Dated file: it is not edited after today, and a newer number is a new file. Vikunja #485's first assumption was that a 7-14B measurably beats Qwen3-1.7B on the 77-case RU routing fixture and the 27-case talk fixture. It does, on both, and it is also faster. ## The setup `gemma-4-12B-it-qat-UD-Q4_K_XL` with the `mtp-gemma-4-12B-it-BF16` draft model, served by `llama-server` b10220 on bugmachine (AMD 7900 GRE, 16GB), fronted by `mavgpud` on `192.168.1.105:8080`. Thinking is off through `--chat-template-kwargs '{"enable_thinking":false}'`, speculative decoding is `--spec-type draft-mtp --spec-draft-n-max 2`, context 32768. The exact line is `deploy/mavgpud.json`. Every number below crossed the LAN from homesrv. Note the trap: homesrv's shell exports `HTTP_PROXY`, Go honours it, and the runs need `env -u HTTP_PROXY -u HTTPS_PROXY -u http_proxy -u https_proxy`. ## Routing, 77-case RU fixture | | full | intent-only | p50 | p95 | |---|---|---|---|---| | classifier alone (02-08) | 68.8% | — | 16.6µs | — | | Qwen3-1.7B through the cascade (31-07, 02-08) | 72.7% | 77.9% | 0.80-1.04s | — | | **gemma-4-12b through the cascade** | **84.4%** | **93.5%** | **329ms** | 429ms | | gemma-4-12b alone, no cascade | 55.8% | 85.7% | 335ms | 436ms | The workstation buys 11.7 points of full accuracy over the resident model. It buys 15.6 points of intent-only, at a third of the latency. The router's p50 was never the model's fault, which the 02-08 contention finding already said. A 12B on a free 16GB card answers a routing turn in a third of a second. Two things the table hides. The alone-versus-cascade gap is slots, not intents. gemma reads the intent right 85.7% of the time on its own. It loses full accuracy on seven fact keys (`вода` instead of `water`, `ужин` instead of `meal`) and on six reminder times with no time slot. Stage 0 and the daemon's own extractor repair both, which is why the cascade is 28 points higher. The lesson is that the cascade earns its keep even under a much better model, not that it is scaffolding to remove. `errors: 6` in the alone row are declines on single-token and ambiguous utterances, all of which the cascade caught. The remaining defects through the cascade are three `query→fact` confusions, one `chat→query`, and one false clarify. ## Talk, 27-case conversational fixture | | pass | notes | |---|---|---| | Qwen3-1.7B (31-07) | 20/27 | 11-17/27 for the 0.8B before it | | **gemma-4-12b** | **25/27 (92.6%)** | chat 8/9, knowledge 9/9, query 8/9 | Knowledge is the interesting column: 9/9, in Russian, with real answers about Rayleigh scattering, SSD versus HDD and thunder delay. That is the case the 1.7B cannot do at all and the reason the naming half of the degradation rule exists. Two failures, and one of them is the persona defect the CPT (#122) targets: `query-notes-do-not-answer` wrote `заплатил` where Maven needs the feminine form. The other is `chat-joke`, where the model told a joke without using any of the words the check looks for. Run-to-run variance is about one case: a second run scored 24/27 with `chat-followup-server` also off-topic. ## Nudge phrasing, 15-case fixture 15/15, every check, no errors. `mood`, `lang`, `length`, `feminine`, `hisgender`, `address`, `cringe` and `ontopic` all clean. ## What this settles and what it does not Settled: the size question. A 12B on the workstation beats the resident model on every fixture we have, and it is faster. The offload argument holds. Not settled: how often the card is free. That is #485's second assumption and only the `mavgpud` log answers it, after a week of the owner's normal work. A model that is better whenever it is up is worth little if it is never up.