Files
Maven/docs/evals/2026-08-02-workstation-gemma4-12b.md
T
claude 774217199e docs: measure gemma-4-12b on the workstation against the resident model (V-485)
Both fixtures, run from homesrv across the LAN with the proxy env stripped.
Routing: 84.4% full / 93.5% intent-only at p50 329ms through the cascade, against
72.7% / 77.9% at p50 0.80-1.04s for Qwen3-1.7B. Talk: 25/27 against 20/27, with
knowledge 9/9. Nudges 15/15. Settles #485's first assumption by measurement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJYcaBiuny9UGSpFweQVQ1
2026-08-03 10:12:21 +02:00

3.8 KiB

gemma-4-12b on the workstation, against the resident Qwen3-1.7B

Measured 2026-08-02 on the fixtures as they stand. Dated file: it is not edited after today, and a newer number is a new file.

Vikunja #485's first assumption was that a 7-14B measurably beats Qwen3-1.7B on the 77-case RU routing fixture and the 27-case talk fixture. It does, on both, and it is also faster.

The setup

gemma-4-12B-it-qat-UD-Q4_K_XL with the mtp-gemma-4-12B-it-BF16 draft model, served by llama-server b10220 on bugmachine (AMD 7900 GRE, 16GB), fronted by mavgpud on 192.168.1.105:8080. Thinking is off through --chat-template-kwargs '{"enable_thinking":false}', speculative decoding is --spec-type draft-mtp --spec-draft-n-max 2, context 32768. The exact line is deploy/mavgpud.json.

Every number below crossed the LAN from homesrv. Note the trap: homesrv's shell exports HTTP_PROXY, Go honours it, and the runs need env -u HTTP_PROXY -u HTTPS_PROXY -u http_proxy -u https_proxy.

Routing, 77-case RU fixture

full intent-only p50 p95
classifier alone (02-08) 68.8% 16.6µs
Qwen3-1.7B through the cascade (31-07, 02-08) 72.7% 77.9% 0.80-1.04s
gemma-4-12b through the cascade 84.4% 93.5% 329ms 429ms
gemma-4-12b alone, no cascade 55.8% 85.7% 335ms 436ms

The workstation buys 11.7 points of full accuracy over the resident model. It buys 15.6 points of intent-only, at a third of the latency. The router's p50 was never the model's fault, which the 02-08 contention finding already said. A 12B on a free 16GB card answers a routing turn in a third of a second.

Two things the table hides.

The alone-versus-cascade gap is slots, not intents. gemma reads the intent right 85.7% of the time on its own. It loses full accuracy on seven fact keys (вода instead of water, ужин instead of meal) and on six reminder times with no time slot. Stage 0 and the daemon's own extractor repair both, which is why the cascade is 28 points higher. The lesson is that the cascade earns its keep even under a much better model, not that it is scaffolding to remove.

errors: 6 in the alone row are declines on single-token and ambiguous utterances, all of which the cascade caught. The remaining defects through the cascade are three query→fact confusions, one chat→query, and one false clarify.

Talk, 27-case conversational fixture

pass notes
Qwen3-1.7B (31-07) 20/27 11-17/27 for the 0.8B before it
gemma-4-12b 25/27 (92.6%) chat 8/9, knowledge 9/9, query 8/9

Knowledge is the interesting column: 9/9, in Russian, with real answers about Rayleigh scattering, SSD versus HDD and thunder delay. That is the case the 1.7B cannot do at all and the reason the naming half of the degradation rule exists.

Two failures, and one of them is the persona defect the CPT (#122) targets: query-notes-do-not-answer wrote заплатил where Maven needs the feminine form. The other is chat-joke, where the model told a joke without using any of the words the check looks for. Run-to-run variance is about one case: a second run scored 24/27 with chat-followup-server also off-topic.

Nudge phrasing, 15-case fixture

15/15, every check, no errors. mood, lang, length, feminine, hisgender, address, cringe and ontopic all clean.

What this settles and what it does not

Settled: the size question. A 12B on the workstation beats the resident model on every fixture we have, and it is faster. The offload argument holds.

Not settled: how often the card is free. That is #485's second assumption and only the mavgpud log answers it, after a week of the owner's normal work. A model that is better whenever it is up is worth little if it is never up.