From 774217199e35aeee4b458c26eb9eca10d5e90154 Mon Sep 17 00:00:00 2001 From: claude Date: Sun, 2 Aug 2026 22:51:47 +0400 Subject: [PATCH] docs: measure gemma-4-12b on the workstation against the resident model (V-485) Both fixtures, run from homesrv across the LAN with the proxy env stripped. Routing: 84.4% full / 93.5% intent-only at p50 329ms through the cascade, against 72.7% / 77.9% at p50 0.80-1.04s for Qwen3-1.7B. Talk: 25/27 against 20/27, with knowledge 9/9. Nudges 15/15. Settles #485's first assumption by measurement. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01NJYcaBiuny9UGSpFweQVQ1 --- .../2026-08-02-workstation-gemma4-12b.md | 82 +++++++++++++++++++ docs/offload.md | 7 +- 2 files changed, 87 insertions(+), 2 deletions(-) create mode 100644 docs/evals/2026-08-02-workstation-gemma4-12b.md diff --git a/docs/evals/2026-08-02-workstation-gemma4-12b.md b/docs/evals/2026-08-02-workstation-gemma4-12b.md new file mode 100644 index 0000000..725de3c --- /dev/null +++ b/docs/evals/2026-08-02-workstation-gemma4-12b.md @@ -0,0 +1,82 @@ +# gemma-4-12b on the workstation, against the resident Qwen3-1.7B + +Measured 2026-08-02 on the fixtures as they stand. Dated file: it is not edited +after today, and a newer number is a new file. + +Vikunja #485's first assumption was that a 7-14B measurably beats Qwen3-1.7B on +the 77-case RU routing fixture and the 27-case talk fixture. It does, on both, +and it is also faster. + +## The setup + +`gemma-4-12B-it-qat-UD-Q4_K_XL` with the `mtp-gemma-4-12B-it-BF16` draft model, +served by `llama-server` b10220 on bugmachine (AMD 7900 GRE, 16GB), fronted by +`mavgpud` on `192.168.1.105:8080`. Thinking is off through +`--chat-template-kwargs '{"enable_thinking":false}'`, speculative decoding is +`--spec-type draft-mtp --spec-draft-n-max 2`, context 32768. The exact line is +`deploy/mavgpud.json`. + +Every number below crossed the LAN from homesrv. Note the trap: homesrv's shell +exports `HTTP_PROXY`, Go honours it, and the runs need +`env -u HTTP_PROXY -u HTTPS_PROXY -u http_proxy -u https_proxy`. + +## Routing, 77-case RU fixture + +| | full | intent-only | p50 | p95 | +|---|---|---|---|---| +| classifier alone (02-08) | 68.8% | — | 16.6µs | — | +| Qwen3-1.7B through the cascade (31-07, 02-08) | 72.7% | 77.9% | 0.80-1.04s | — | +| **gemma-4-12b through the cascade** | **84.4%** | **93.5%** | **329ms** | 429ms | +| gemma-4-12b alone, no cascade | 55.8% | 85.7% | 335ms | 436ms | + +The workstation buys 11.7 points of full accuracy over the resident model. It +buys 15.6 points of intent-only, at a third of the latency. The router's p50 was +never the model's fault, which the 02-08 contention finding already said. A 12B +on a free 16GB card answers a routing turn in a third of a second. + +Two things the table hides. + +The alone-versus-cascade gap is slots, not intents. gemma reads the intent right +85.7% of the time on its own. It loses full accuracy on seven fact keys +(`вода` instead of `water`, `ужин` instead of `meal`) and on six reminder times +with no time slot. Stage 0 and the daemon's own extractor repair +both, which is why the cascade is 28 points higher. The lesson is that the +cascade earns its keep even under a much better model, not that it is scaffolding +to remove. + +`errors: 6` in the alone row are declines on single-token and ambiguous +utterances, all of which the cascade caught. The remaining defects through the +cascade are three `query→fact` confusions, one `chat→query`, and one false +clarify. + +## Talk, 27-case conversational fixture + +| | pass | notes | +|---|---|---| +| Qwen3-1.7B (31-07) | 20/27 | 11-17/27 for the 0.8B before it | +| **gemma-4-12b** | **25/27 (92.6%)** | chat 8/9, knowledge 9/9, query 8/9 | + +Knowledge is the interesting column: 9/9, in Russian, with real answers about +Rayleigh scattering, SSD versus HDD and thunder delay. That is the case the +1.7B cannot do at all and the reason the naming half of the degradation rule +exists. + +Two failures, and one of them is the persona defect the CPT (#122) targets: +`query-notes-do-not-answer` wrote `заплатил` where Maven needs the feminine +form. The other is `chat-joke`, where the model told a joke without using any of +the words the check looks for. Run-to-run variance is about one case: a second +run scored 24/27 with `chat-followup-server` also off-topic. + +## Nudge phrasing, 15-case fixture + +15/15, every check, no errors. `mood`, `lang`, `length`, `feminine`, +`hisgender`, `address`, `cringe` and `ontopic` all clean. + +## What this settles and what it does not + +Settled: the size question. A 12B on the workstation beats the resident model on +every fixture we have, and it is faster. The offload argument holds. + +Not settled: how often the card is free. That is #485's second assumption and +only the `mavgpud` log answers it, after a week of the owner's normal work. A +model that is better whenever it is up is worth little if it is never up. diff --git a/docs/offload.md b/docs/offload.md index 32291fe..4971447 100644 --- a/docs/offload.md +++ b/docs/offload.md @@ -129,8 +129,11 @@ Then the embedder above, **whisper.cpp** in `mavsttd`, and **piper** in `mavttsd builds an `llm.Pair` in `modelSeam` (`cmd/mavend/voicewire.go`), and routing and replies complete through it. Both are the silent half of the rule. The naming half is not wired. A world question still goes to the resident model - through `PhraseQuery`, and the fixture measurement has not been run. - Biggest quality delta. A 16GB card runs a 7-14B, + through `PhraseQuery`. + Measured, `docs/evals/2026-08-02-workstation-gemma4-12b.md`: gemma-4-12b + through the cascade scores 84.4% full accuracy at p50 329ms. The resident + model scores 72.7% at p50 0.80-1.04s. On the talk fixture it is 25/27 + against 20/27. Biggest quality delta. A 16GB card runs a 7-14B, which fixes what the 1.7B gets wrong: world knowledge, and the persona the CPT targets. The degradation path is already written and measured, since the classifier scores 68.8% full accuracy at p50 16.6µs on its own.