Both fixtures, run from homesrv across the LAN with the proxy env stripped. Routing: 84.4% full / 93.5% intent-only at p50 329ms through the cascade, against 72.7% / 77.9% at p50 0.80-1.04s for Qwen3-1.7B. Talk: 25/27 against 20/27, with knowledge 9/9. Nudges 15/15. Settles #485's first assumption by measurement. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NJYcaBiuny9UGSpFweQVQ1
3.8 KiB
gemma-4-12b on the workstation, against the resident Qwen3-1.7B
Measured 2026-08-02 on the fixtures as they stand. Dated file: it is not edited after today, and a newer number is a new file.
Vikunja #485's first assumption was that a 7-14B measurably beats Qwen3-1.7B on the 77-case RU routing fixture and the 27-case talk fixture. It does, on both, and it is also faster.
The setup
gemma-4-12B-it-qat-UD-Q4_K_XL with the mtp-gemma-4-12B-it-BF16 draft model,
served by llama-server b10220 on bugmachine (AMD 7900 GRE, 16GB), fronted by
mavgpud on 192.168.1.105:8080. Thinking is off through
--chat-template-kwargs '{"enable_thinking":false}', speculative decoding is
--spec-type draft-mtp --spec-draft-n-max 2, context 32768. The exact line is
deploy/mavgpud.json.
Every number below crossed the LAN from homesrv. Note the trap: homesrv's shell
exports HTTP_PROXY, Go honours it, and the runs need
env -u HTTP_PROXY -u HTTPS_PROXY -u http_proxy -u https_proxy.
Routing, 77-case RU fixture
| full | intent-only | p50 | p95 | |
|---|---|---|---|---|
| classifier alone (02-08) | 68.8% | — | 16.6µs | — |
| Qwen3-1.7B through the cascade (31-07, 02-08) | 72.7% | 77.9% | 0.80-1.04s | — |
| gemma-4-12b through the cascade | 84.4% | 93.5% | 329ms | 429ms |
| gemma-4-12b alone, no cascade | 55.8% | 85.7% | 335ms | 436ms |
The workstation buys 11.7 points of full accuracy over the resident model. It buys 15.6 points of intent-only, at a third of the latency. The router's p50 was never the model's fault, which the 02-08 contention finding already said. A 12B on a free 16GB card answers a routing turn in a third of a second.
Two things the table hides.
The alone-versus-cascade gap is slots, not intents. gemma reads the intent right
85.7% of the time on its own. It loses full accuracy on seven fact keys
(вода instead of water, ужин instead of meal) and on six reminder times
with no time slot. Stage 0 and the daemon's own extractor repair
both, which is why the cascade is 28 points higher. The lesson is that the
cascade earns its keep even under a much better model, not that it is scaffolding
to remove.
errors: 6 in the alone row are declines on single-token and ambiguous
utterances, all of which the cascade caught. The remaining defects through the
cascade are three query→fact confusions, one chat→query, and one false
clarify.
Talk, 27-case conversational fixture
| pass | notes | |
|---|---|---|
| Qwen3-1.7B (31-07) | 20/27 | 11-17/27 for the 0.8B before it |
| gemma-4-12b | 25/27 (92.6%) | chat 8/9, knowledge 9/9, query 8/9 |
Knowledge is the interesting column: 9/9, in Russian, with real answers about Rayleigh scattering, SSD versus HDD and thunder delay. That is the case the 1.7B cannot do at all and the reason the naming half of the degradation rule exists.
Two failures, and one of them is the persona defect the CPT (#122) targets:
query-notes-do-not-answer wrote заплатил where Maven needs the feminine
form. The other is chat-joke, where the model told a joke without using any of
the words the check looks for. Run-to-run variance is about one case: a second
run scored 24/27 with chat-followup-server also off-topic.
Nudge phrasing, 15-case fixture
15/15, every check, no errors. mood, lang, length, feminine,
hisgender, address, cringe and ontopic all clean.
What this settles and what it does not
Settled: the size question. A 12B on the workstation beats the resident model on every fixture we have, and it is faster. The offload argument holds.
Not settled: how often the card is free. That is #485's second assumption and
only the mavgpud log answers it, after a week of the owner's normal work. A
model that is better whenever it is up is worth little if it is never up.