Seventeen markdown files at the repo root, twelve of them dated one-shot reports sitting next to CLAUDE.md. That is why stale docs read as current: nothing in the path said which was which. Root now keeps CLAUDE.md and AGENTS.md. Living docs move under docs/ and carry a Last verified line. Dated measurements move to docs/evals/ ISO-prefixed, and are never edited after the day, so a newer number is a new file. The senior review moves to docs/archive/. Every reference was rewritten across markdown, Go comments, the Makefile and the recall fixture. The touched Go packages still build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
12 KiB
Resident model bake-off — 31-07-2026
Outcome: the resident model is stock Qwen3-1.7B (UD-Q4_K_XL). Two sweeps ran this
evening and the second one changed the answer — read to the end before acting on any table
here. Second sweep
is the one that holds.
First sweep — LFM2.5-1.2B vs Qwen3.5-0.8B
Verdict, scoped to this pair: keep Qwen3.5-0.8B over LFM2.5-1.2B. LFM2.5-1.2B is worse at routing (52.6% vs 60.5% intent accuracy), and the loss is almost entirely Russian (18/61 vs 22/61 RU, while EN is a wash). It is also 2.4× slower. The Thinking variant is far worse again. This verdict still stands as written — it rejects LFM2.5-1.2B. It is not a recommendation to keep 0.8B as the resident model; the second sweep replaced it with Qwen3-1.7B.
Settles Vikunja #278 / #250.
- Same fixture and scorer as
docs/evals/2026-07-31-routing.md:internal/router/eval/(ru_routing_v1.json, 76 held-out cases). - Reproduce:
MAVEN_LLM_URL=http://127.0.0.1:<port> make eval-router(TestLLMRouterBaseline). (This line used to say there is nomake eval-modelstarget. There is one now — start a server with the gguf you want, thenmake eval-models MAVEN_LLM_URL=http://127.0.0.1:<port>. It runs only the LLM test, since the classifier baselines do not depend on the model.) - All three models served by the same
llama-serverflags —-c 2048 -ngl 99 -t 6, only-mand--portdiffer. One server at a time on an otherwise idle box, so latencies are real and not contention. - Measured on top of the router prompt fix (
origin/overnight/router-promptmerged in), so the Qwen column is directly comparable to the numbers already recorded.
Results
llm-only — the model alone. This is the column that measures the model.
| Qwen3.5-0.8B | LFM2.5-1.2B Instruct | LFM2.5-1.2B Thinking | |
|---|---|---|---|
| intent-only accuracy | 60.5% | 52.6% | 36.8% |
| full accuracy (intent+slots+gate) | 36.8% | 32.9% | 21.1% |
| RU | 22/61 | 18/61 | 10/61 |
| EN | 6/15 | 7/15 | 6/15 |
| route errors | 0 | 0 | 0 |
| p50 / p95 latency | 1.05s / 1.71s | 2.47s / 3.62s | 2.42s / 3.24s |
| missed clarify | 6 / 6 | 6 / 6 | 6 / 6 |
cascade+llm — stage-0 → model → classifier floor, what #320 would actually ship. Same
ordering.
| Qwen3.5-0.8B | LFM2.5-1.2B Instruct | LFM2.5-1.2B Thinking | |
|---|---|---|---|
| intent-only accuracy | 61.8% | 55.3% | 38.2% |
| full accuracy | 46.1% | 42.1% | 30.3% |
| RU / EN | 27/61 / 8/15 | 23/61 / 9/15 | 15/61 / 8/15 |
| route errors | 0 | 0 | 0 |
| p50 / p95 latency | 1.28s / 1.94s | 2.18s / 2.72s | 2.27s / 3.19s |
Full logs: the three runs are archived in the session scratchpad
(qwen08.txt, lfm-instruct.txt, lfm-thinking.txt).
Russian-specific failures — the owner's worry is confirmed
LFM2.5's Russian loss is not spread out. It has one large, specific failure: it hears
almost any Russian imperative or short phrase as reminder.
перезапусти докер→ reminder (want act)включи вытяжку→ reminder (want act)закрой жалюзи→ reminder (want act)заметка: продлить домен в августе→ reminder (want note)запиши что кран на кухне снова капает→ reminder (want note)доброе утро→ reminder (want chat)спасибо тебе→ reminder (want note/chat)переходи в тихий режим→ reminder (want system)
That is note→reminder ×4, act→reminder ×4, chat→reminder ×2 in one run. Qwen's
equivalent failure axis is query→fact ×8, which is a narrower and already-understood bug.
Two more Russian-side problems worth naming:
- Fact keys come back empty or wrong in Russian.
воды попил наконец,поужинал,поспал часов пятьandотметь что я позавтракал овсянкойall returned an empty key.сходил в душandотдохнул минут двадцатьboth returnedwater. Qwen does not do this. - It leaked German.
slept about seven hoursproduced the fact key"7 Stunden geschlafen". Grammar-valid, semantically garbage — a sign the multilingual mix is not anchored where Maven needs it.
The claimed tool-calling advantage did not show up here. act is the closest thing this
fixture has to a tool call, and LFM2.5 got it wrong more often than Qwen, mostly by calling
it a reminder. It also produced no fn slot on any act, same as Qwen.
The Thinking variant
Not viable. 36.8% intent accuracy, 10/61 Russian, and no latency saving over Instruct — the
thinking trace costs time without buying accuracy on a short enum classification. With the
enable_thinking=false diagnostic it collapsed further to 28.9% with 2 route errors
(query→reminder ×12). Do not pursue.
Notes
- Nothing crashed, nothing ignored the GBNF grammar, and no model produced unparseable JSON
in the shippable configurations. Zero route errors for both Instruct and Thinking in
llm-onlyandcascade+llm. The problem with LFM2.5 is what it decides, not whether it can emit the contract. - The
6 / 6missed clarify is unchanged across all three models. No model fixes the missing refusal lane — that isConfidence: 1.0hardcoded inllmrouter.go(Vikunja #359), not a model property. - The report labels every configuration
(0.8B); that string is hardcoded in the test, not a reflection of which gguf was loaded. Model identity was confirmed per run via/v1/models. - No Go code was changed for this measurement, and no bug was found that needed one.
What this does not settle
Routing only. LFM2.5 might still phrase better, and phrasing is the resident model's other job — that needs its own fixture. But routing is the load-bearing path and Maven is Russian-first, so on the evidence here the switch is not worth making.
Second sweep, same evening — five models, and a resident-model change
The sections above compared LFM2.5-1.2B against Qwen3.5-0.8B on routing and concluded "the switch is not worth making". That still holds. This sweep asked a different question — whether a smaller model could work, since LFM2.5's published instruction-following scores beat Qwen3.5-0.8B badly — and answered it, plus found a better resident model by accident.
Outcome: the resident model is now stock Qwen3-1.7B. Sub-500M is a dead end.
Routing — 77 Russian cases, one run each
| model | on disk | llm-only (full) | llm-only (intent) | cascade + fallback |
|---|---|---|---|---|
| LFM2.5-230M-Q8_0 | 246 MB | 23.4% | 33.8% | 36.4% |
| LFM2.5-350M-Q8_0 | 379 MB | 2.6% | 5.2% | 20.8% |
| Qwen3.5-0.8B-Q4_K_M | 527 MB | 36.4% | 59.7% | 61.0% |
| Qwen3.5-2B-UD-Q4_K_XL | 1.34 GB | 42.9% | 62.3% | 63.6% |
| Qwen3-1.7B-UD-Q4_K_XL (stock) | 1.13 GB | 44.2% | 67.5% | 72.7% |
Qwen3-1.7B wins every column, including against a model 20% larger than it.
Talk fixture — 27 cases, three runs each, idle box
| Qwen3.5-0.8B | Qwen3-1.7B stock | |
|---|---|---|
| composite | 13, 11, 8 | 20, 21, 18 |
| address | 21, 18, 18 | 26, 25, 23 |
| feminine | 27, 25, 26 | 26, 27, 26 |
| lang | 27, 27, 26 | 26, 27, 27 |
| ontopic | 16, 19, 19 | 22, 23, 23 |
| canned fallbacks | 8, 5, 6 | 0, 2, 0 |
This also fills the row docs/evals/2026-07-31-talk.md had to void for contamination:
600ch/1024tok on Qwen3.5-0.8B scores 13, 11, 8.
address is the headline. It sat at 18-22 of 27 on the 0.8B no matter how the prompt
was worded — the prompt explicitly forbids "вы" and the model writes вашей,
подождите, делаете anyway. That was read as "prompting is out of levers", and it
was really "0.8B is out of capacity". The 1.7B mostly holds the constraint.
The fallback column matters too: 5-8 of 27 turns on the 0.8B end in a hardcoded
"не знаю.", meaning it failed to emit parseable JSON about a quarter of the time.
The 1.7B does that 0-2 times.
Latency — the long tail is not the Thinking block
Stale, corrected 2026-08-02. The p50 figures in this table are contention on a shared llama-server, not the model's cost. The router measures p50 825ms / p95 1.2s / max 3.0s in
docs/evals/2026-07-31-routing.md, which says so at line 61. Read this table for the shape of the tail only. Take absolute latency from the routing eval.
| p50 | p95 | |
|---|---|---|
| Qwen3.5-0.8B | 2.4s, 2.9s, 2.0s | 17.4s, 17.6s, 17.4s |
| Qwen3-1.7B stock | 2.7s, 2.6s, 2.8s | 16.4s, 6.6s, 3.9s |
p50 is flat across a 2× size difference. The first instinct on seeing the 1.7B's
16s p95 was "that is the reasoning trace, cap it" — wrong. The 0.8B's p95 is a
consistent 17s and the 1.7B beat it in two of three runs. The tail is shared and
lives somewhere else. Do not spend time on /no_think on this evidence.
Sub-500M: not close, and the benchmarks say otherwise for a reason
LFM2.5-350M publishes IFEval 76.96 against Qwen3.5-0.8B's 59.94, and BFCLv3 44.11 against 35.08 — better at instruction-following and structured output, at 2/3 the size. Those numbers are real and they are English. Every benchmark in that table except Multi-IF is English-only.
In Russian, with a 300-token budget and temperature 0:
- 350M, «Столица Франции? Ответь кратко.» → «Сторзит в Париже.» —
Сторзитis not a word; it is invented morphology. - 350M, asked to read back a reminder → a fortune cookie about being attentive and confident. No reminder in it.
- 230M, «Привет, как дела?» → answered in Spanish.
The 230M beating the 350M six-fold on routing (33.8% vs 5.2%) is the other tell: when the larger sibling collapses like that it is format compliance failing, not reasoning.
This is a pretraining gap, not a fine-tuning gap. Teaching Russian to a 350M from near-zero is not an afternoon on a Colab, which was the premise worth checking.
Why this vindicates the 1.7B CPT
Stock Qwen3-1.7B, untrained and unprompted, answers all three probes in fluent
correct Russian. What it gets wrong is the persona: «Привет! Я рад, что ты здесь»
— рад is masculine and Maven needs рада. That is the right kind of remaining
problem, and it is exactly what the CPT (Vikunja #122) is for.
The 1.7B was the correct model choice. What was wrong was treating it as a blocker: stock already beats what was deployed, so it ships now and gets swapped again when the CPT lands.
Caveats
- Routing is one run per model, not three. The gaps between families are far larger than the run-to-run spread seen on the talk fixture, but the 2B-vs-1.7B gap (62.3 vs 67.5) is not safe to call on one run.
The routing numbers only reach production once the LLM router is wired on. It is stillResolved the same evening: the LLM router is wired atnil.voice.go:214behindvoice.llm_router, the default is on, anddeploy/mavend.jsonsets ittrue. These numbers are the production path now. Corrected 2026-08-02: the p50 ≈2.7s in the latency table above WAS a bench artifact. It is contention on the shared llama-server, not the model.docs/evals/2026-07-31-routing.mdline 61 says so, and measures the router at p50 825ms / p95 1.2s / max 3.0s. Cite that file for latency, not this one.Deleted 2026-07-31. The good 1.13 GB copy in/mnt/hdd1/llms/LFM2.5/Qwen3-1.7B-UD-Q4_K_XL.ggufis a 293 MB truncated download in the wrong directory.qwen3/is whatdeploy/mavend.jsonloads.- Harness:
scratchpad/bakeoff.sh, one server at a time, health-checked before each run,/v1/modelsrecorded per run. Never run two LLM consumers at once — see the contamination note indocs/evals/2026-07-31-talk.md.