docs: measure gemma-4-12b on the workstation against the resident model (V-485)

Both fixtures, run from homesrv across the LAN with the proxy env stripped.
Routing: 84.4% full / 93.5% intent-only at p50 329ms through the cascade, against
72.7% / 77.9% at p50 0.80-1.04s for Qwen3-1.7B. Talk: 25/27 against 20/27, with
knowledge 9/9. Nudges 15/15. Settles #485's first assumption by measurement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJYcaBiuny9UGSpFweQVQ1
This commit is contained in:
2026-08-02 22:51:47 +04:00
committed by kami
parent 2db59d52a7
commit 774217199e
2 changed files with 87 additions and 2 deletions
@@ -0,0 +1,82 @@
# gemma-4-12b on the workstation, against the resident Qwen3-1.7B
Measured 2026-08-02 on the fixtures as they stand. Dated file: it is not edited
after today, and a newer number is a new file.
Vikunja #485's first assumption was that a 7-14B measurably beats Qwen3-1.7B on
the 77-case RU routing fixture and the 27-case talk fixture. It does, on both,
and it is also faster.
## The setup
`gemma-4-12B-it-qat-UD-Q4_K_XL` with the `mtp-gemma-4-12B-it-BF16` draft model,
served by `llama-server` b10220 on bugmachine (AMD 7900 GRE, 16GB), fronted by
`mavgpud` on `192.168.1.105:8080`. Thinking is off through
`--chat-template-kwargs '{"enable_thinking":false}'`, speculative decoding is
`--spec-type draft-mtp --spec-draft-n-max 2`, context 32768. The exact line is
`deploy/mavgpud.json`.
Every number below crossed the LAN from homesrv. Note the trap: homesrv's shell
exports `HTTP_PROXY`, Go honours it, and the runs need
`env -u HTTP_PROXY -u HTTPS_PROXY -u http_proxy -u https_proxy`.
## Routing, 77-case RU fixture
| | full | intent-only | p50 | p95 |
|---|---|---|---|---|
| classifier alone (02-08) | 68.8% | — | 16.6µs | — |
| Qwen3-1.7B through the cascade (31-07, 02-08) | 72.7% | 77.9% | 0.80-1.04s | — |
| **gemma-4-12b through the cascade** | **84.4%** | **93.5%** | **329ms** | 429ms |
| gemma-4-12b alone, no cascade | 55.8% | 85.7% | 335ms | 436ms |
The workstation buys 11.7 points of full accuracy over the resident model. It
buys 15.6 points of intent-only, at a third of the latency. The router's p50 was
never the model's fault, which the 02-08 contention finding already said. A 12B
on a free 16GB card answers a routing turn in a third of a second.
Two things the table hides.
The alone-versus-cascade gap is slots, not intents. gemma reads the intent right
85.7% of the time on its own. It loses full accuracy on seven fact keys
(`вода` instead of `water`, `ужин` instead of `meal`) and on six reminder times
with no time slot. Stage 0 and the daemon's own extractor repair
both, which is why the cascade is 28 points higher. The lesson is that the
cascade earns its keep even under a much better model, not that it is scaffolding
to remove.
`errors: 6` in the alone row are declines on single-token and ambiguous
utterances, all of which the cascade caught. The remaining defects through the
cascade are three `query→fact` confusions, one `chat→query`, and one false
clarify.
## Talk, 27-case conversational fixture
| | pass | notes |
|---|---|---|
| Qwen3-1.7B (31-07) | 20/27 | 11-17/27 for the 0.8B before it |
| **gemma-4-12b** | **25/27 (92.6%)** | chat 8/9, knowledge 9/9, query 8/9 |
Knowledge is the interesting column: 9/9, in Russian, with real answers about
Rayleigh scattering, SSD versus HDD and thunder delay. That is the case the
1.7B cannot do at all and the reason the naming half of the degradation rule
exists.
Two failures, and one of them is the persona defect the CPT (#122) targets:
`query-notes-do-not-answer` wrote `заплатил` where Maven needs the feminine
form. The other is `chat-joke`, where the model told a joke without using any of
the words the check looks for. Run-to-run variance is about one case: a second
run scored 24/27 with `chat-followup-server` also off-topic.
## Nudge phrasing, 15-case fixture
15/15, every check, no errors. `mood`, `lang`, `length`, `feminine`,
`hisgender`, `address`, `cringe` and `ontopic` all clean.
## What this settles and what it does not
Settled: the size question. A 12B on the workstation beats the resident model on
every fixture we have, and it is faster. The offload argument holds.
Not settled: how often the card is free. That is #485's second assumption and
only the `mavgpud` log answers it, after a week of the owner's normal work. A
model that is better whenever it is up is worth little if it is never up.
+5 -2
View File
@@ -129,8 +129,11 @@ Then the embedder above, **whisper.cpp** in `mavsttd`, and **piper** in `mavttsd
builds an `llm.Pair` in `modelSeam` (`cmd/mavend/voicewire.go`), and routing
and replies complete through it. Both are the silent half of the rule. The
naming half is not wired. A world question still goes to the resident model
through `PhraseQuery`, and the fixture measurement has not been run.
Biggest quality delta. A 16GB card runs a 7-14B,
through `PhraseQuery`.
Measured, `docs/evals/2026-08-02-workstation-gemma4-12b.md`: gemma-4-12b
through the cascade scores 84.4% full accuracy at p50 329ms. The resident
model scores 72.7% at p50 0.80-1.04s. On the talk fixture it is 25/27
against 20/27. Biggest quality delta. A 16GB card runs a 7-14B,
which fixes what the 1.7B gets wrong: world knowledge, and the persona the CPT
targets. The degradation path is already written and measured, since the
classifier scores 68.8% full accuracy at p50 16.6µs on its own.