docs: tier the tree by lifetime, so staleness shows in the path (V-446)
Seventeen markdown files at the repo root, twelve of them dated one-shot reports sitting next to CLAUDE.md. That is why stale docs read as current: nothing in the path said which was which. Root now keeps CLAUDE.md and AGENTS.md. Living docs move under docs/ and carry a Last verified line. Dated measurements move to docs/evals/ ISO-prefixed, and are never edited after the day, so a newer number is a new file. The senior review moves to docs/archive/. Every reference was rewritten across markdown, Go comments, the Makefile and the recall fixture. The touched Go packages still build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,230 @@
|
||||
# Resident model bake-off — 31-07-2026
|
||||
|
||||
**Outcome: the resident model is stock Qwen3-1.7B** (`UD-Q4_K_XL`). Two sweeps ran this
|
||||
evening and the second one changed the answer — read to the end before acting on any table
|
||||
here. [Second sweep](#second-sweep-same-evening--five-models-and-a-resident-model-change)
|
||||
is the one that holds.
|
||||
|
||||
## First sweep — LFM2.5-1.2B vs Qwen3.5-0.8B
|
||||
|
||||
**Verdict, scoped to this pair: keep Qwen3.5-0.8B over LFM2.5-1.2B.** LFM2.5-1.2B is worse
|
||||
at routing (52.6% vs 60.5% intent accuracy), and the loss is almost entirely Russian
|
||||
(18/61 vs 22/61 RU, while EN is a wash). It is also 2.4× slower. The Thinking variant is
|
||||
far worse again. This verdict still stands as written — it rejects LFM2.5-1.2B. It is
|
||||
**not** a recommendation to keep 0.8B as the resident model; the second sweep replaced it
|
||||
with Qwen3-1.7B.
|
||||
|
||||
Settles Vikunja **#278 / #250**.
|
||||
|
||||
- Same fixture and scorer as `docs/evals/2026-07-31-routing.md`: `internal/router/eval/`
|
||||
(`ru_routing_v1.json`, 76 held-out cases).
|
||||
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:<port> make eval-router`
|
||||
(`TestLLMRouterBaseline`). (This line used to say there is no `make eval-models` target.
|
||||
There is one now — start a server with the gguf you want, then
|
||||
`make eval-models MAVEN_LLM_URL=http://127.0.0.1:<port>`. It runs only the LLM test, since
|
||||
the classifier baselines do not depend on the model.)
|
||||
- All three models served by the same `llama-server` flags — `-c 2048 -ngl 99 -t 6`, only
|
||||
`-m` and `--port` differ. One server at a time on an otherwise idle box, so latencies are
|
||||
real and not contention.
|
||||
- Measured on top of the router prompt fix (`origin/overnight/router-prompt` merged in), so
|
||||
the Qwen column is directly comparable to the numbers already recorded.
|
||||
|
||||
## Results
|
||||
|
||||
`llm-only` — the model alone. This is the column that measures the model.
|
||||
|
||||
| | Qwen3.5-0.8B | LFM2.5-1.2B Instruct | LFM2.5-1.2B Thinking |
|
||||
|---|---|---|---|
|
||||
| **intent-only accuracy** | **60.5%** | 52.6% | 36.8% |
|
||||
| full accuracy (intent+slots+gate) | **36.8%** | 32.9% | 21.1% |
|
||||
| **RU** | **22/61** | 18/61 | 10/61 |
|
||||
| EN | 6/15 | **7/15** | 6/15 |
|
||||
| route errors | 0 | 0 | 0 |
|
||||
| **p50 / p95 latency** | **1.05s / 1.71s** | 2.47s / 3.62s | 2.42s / 3.24s |
|
||||
| missed clarify | 6 / 6 | 6 / 6 | 6 / 6 |
|
||||
|
||||
`cascade+llm` — stage-0 → model → classifier floor, what #320 would actually ship. Same
|
||||
ordering.
|
||||
|
||||
| | Qwen3.5-0.8B | LFM2.5-1.2B Instruct | LFM2.5-1.2B Thinking |
|
||||
|---|---|---|---|
|
||||
| intent-only accuracy | **61.8%** | 55.3% | 38.2% |
|
||||
| full accuracy | **46.1%** | 42.1% | 30.3% |
|
||||
| RU / EN | **27/61** / 8/15 | 23/61 / **9/15** | 15/61 / 8/15 |
|
||||
| route errors | 0 | 0 | 0 |
|
||||
| p50 / p95 latency | **1.28s / 1.94s** | 2.18s / 2.72s | 2.27s / 3.19s |
|
||||
|
||||
Full logs: the three runs are archived in the session scratchpad
|
||||
(`qwen08.txt`, `lfm-instruct.txt`, `lfm-thinking.txt`).
|
||||
|
||||
## Russian-specific failures — the owner's worry is confirmed
|
||||
|
||||
LFM2.5's Russian loss is not spread out. It has one large, specific failure: **it hears
|
||||
almost any Russian imperative or short phrase as `reminder`.**
|
||||
|
||||
- `перезапусти докер` → reminder (want act)
|
||||
- `включи вытяжку` → reminder (want act)
|
||||
- `закрой жалюзи` → reminder (want act)
|
||||
- `заметка: продлить домен в августе` → reminder (want note)
|
||||
- `запиши что кран на кухне снова капает` → reminder (want note)
|
||||
- `доброе утро` → reminder (want chat)
|
||||
- `спасибо тебе` → reminder (want note/chat)
|
||||
- `переходи в тихий режим` → reminder (want system)
|
||||
|
||||
That is `note→reminder ×4`, `act→reminder ×4`, `chat→reminder ×2` in one run. Qwen's
|
||||
equivalent failure axis is `query→fact ×8`, which is a narrower and already-understood bug.
|
||||
|
||||
Two more Russian-side problems worth naming:
|
||||
|
||||
1. **Fact keys come back empty or wrong in Russian.** `воды попил наконец`, `поужинал`,
|
||||
`поспал часов пять` and `отметь что я позавтракал овсянкой` all returned an empty key.
|
||||
`сходил в душ` and `отдохнул минут двадцать` both returned `water`. Qwen does not do this.
|
||||
2. **It leaked German.** `slept about seven hours` produced the fact key
|
||||
`"7 Stunden geschlafen"`. Grammar-valid, semantically garbage — a sign the multilingual
|
||||
mix is not anchored where Maven needs it.
|
||||
|
||||
The claimed tool-calling advantage did not show up here. `act` is the closest thing this
|
||||
fixture has to a tool call, and LFM2.5 got it wrong more often than Qwen, mostly by calling
|
||||
it a reminder. It also produced no `fn` slot on any act, same as Qwen.
|
||||
|
||||
## The Thinking variant
|
||||
|
||||
Not viable. 36.8% intent accuracy, 10/61 Russian, and no latency saving over Instruct — the
|
||||
thinking trace costs time without buying accuracy on a short enum classification. With the
|
||||
`enable_thinking=false` diagnostic it collapsed further to 28.9% with 2 route errors
|
||||
(`query→reminder ×12`). Do not pursue.
|
||||
|
||||
## Notes
|
||||
|
||||
- Nothing crashed, nothing ignored the GBNF grammar, and no model produced unparseable JSON
|
||||
in the shippable configurations. Zero route errors for both Instruct and Thinking in
|
||||
`llm-only` and `cascade+llm`. The problem with LFM2.5 is what it decides, not whether it
|
||||
can emit the contract.
|
||||
- The `6 / 6` missed clarify is unchanged across all three models. No model fixes the missing
|
||||
refusal lane — that is `Confidence: 1.0` hardcoded in `llmrouter.go` (Vikunja #359), not a
|
||||
model property.
|
||||
- The report labels every configuration `(0.8B)`; that string is hardcoded in the test, not a
|
||||
reflection of which gguf was loaded. Model identity was confirmed per run via `/v1/models`.
|
||||
- No Go code was changed for this measurement, and no bug was found that needed one.
|
||||
|
||||
## What this does not settle
|
||||
|
||||
Routing only. LFM2.5 might still phrase better, and phrasing is the resident model's other
|
||||
job — that needs its own fixture. But routing is the load-bearing path and Maven is
|
||||
Russian-first, so on the evidence here the switch is not worth making.
|
||||
|
||||
---
|
||||
|
||||
# Second sweep, same evening — five models, and a resident-model change
|
||||
|
||||
The sections above compared LFM2.5-1.2B against Qwen3.5-0.8B on routing and concluded
|
||||
"the switch is not worth making". That still holds. This sweep asked a different
|
||||
question — whether a *smaller* model could work, since LFM2.5's published
|
||||
instruction-following scores beat Qwen3.5-0.8B badly — and answered it, plus found a
|
||||
better resident model by accident.
|
||||
|
||||
**Outcome: the resident model is now stock Qwen3-1.7B.** Sub-500M is a dead end.
|
||||
|
||||
## Routing — 77 Russian cases, one run each
|
||||
|
||||
| model | on disk | llm-only (full) | llm-only (intent) | cascade + fallback |
|
||||
|---|---|---|---|---|
|
||||
| LFM2.5-230M-Q8_0 | 246 MB | 23.4% | 33.8% | 36.4% |
|
||||
| LFM2.5-350M-Q8_0 | 379 MB | 2.6% | **5.2%** | 20.8% |
|
||||
| Qwen3.5-0.8B-Q4_K_M | 527 MB | 36.4% | 59.7% | 61.0% |
|
||||
| Qwen3.5-2B-UD-Q4_K_XL | 1.34 GB | 42.9% | 62.3% | 63.6% |
|
||||
| **Qwen3-1.7B-UD-Q4_K_XL (stock)** | 1.13 GB | **44.2%** | **67.5%** | **72.7%** |
|
||||
|
||||
Qwen3-1.7B wins every column, including against a model 20% larger than it.
|
||||
|
||||
## Talk fixture — 27 cases, three runs each, idle box
|
||||
|
||||
| | Qwen3.5-0.8B | Qwen3-1.7B stock |
|
||||
|---|---|---|
|
||||
| composite | 13, 11, 8 | **20, 21, 18** |
|
||||
| address | 21, 18, 18 | **26, 25, 23** |
|
||||
| feminine | 27, 25, 26 | 26, 27, 26 |
|
||||
| lang | 27, 27, 26 | 26, 27, 27 |
|
||||
| ontopic | 16, 19, 19 | **22, 23, 23** |
|
||||
| canned fallbacks | 8, 5, 6 | **0, 2, 0** |
|
||||
|
||||
This also fills the row `docs/evals/2026-07-31-talk.md` had to void for contamination:
|
||||
**600ch/1024tok on Qwen3.5-0.8B scores 13, 11, 8.**
|
||||
|
||||
`address` is the headline. It sat at 18-22 of 27 on the 0.8B no matter how the prompt
|
||||
was worded — the prompt explicitly forbids "вы" and the model writes `вашей`,
|
||||
`подождите`, `делаете` anyway. That was read as "prompting is out of levers", and it
|
||||
was really "0.8B is out of capacity". The 1.7B mostly holds the constraint.
|
||||
|
||||
The fallback column matters too: 5-8 of 27 turns on the 0.8B end in a hardcoded
|
||||
`"не знаю."`, meaning it failed to emit parseable JSON about a quarter of the time.
|
||||
The 1.7B does that 0-2 times.
|
||||
|
||||
## Latency — the long tail is not the Thinking block
|
||||
|
||||
> **Stale, corrected 2026-08-02.** The p50 figures in this table are contention on a
|
||||
> shared llama-server, not the model's cost. The router measures p50 825ms / p95 1.2s /
|
||||
> max 3.0s in `docs/evals/2026-07-31-routing.md`, which says so at line 61. Read this table for
|
||||
> the shape of the tail only. Take absolute latency from the routing eval.
|
||||
|
||||
| | p50 | p95 |
|
||||
|---|---|---|
|
||||
| Qwen3.5-0.8B | 2.4s, 2.9s, 2.0s | 17.4s, 17.6s, 17.4s |
|
||||
| Qwen3-1.7B stock | 2.7s, 2.6s, 2.8s | 16.4s, 6.6s, 3.9s |
|
||||
|
||||
p50 is flat across a 2× size difference. The first instinct on seeing the 1.7B's
|
||||
16s p95 was "that is the reasoning trace, cap it" — wrong. The 0.8B's p95 is a
|
||||
consistent 17s and the 1.7B beat it in two of three runs. The tail is shared and
|
||||
lives somewhere else. Do not spend time on `/no_think` on this evidence.
|
||||
|
||||
## Sub-500M: not close, and the benchmarks say otherwise for a reason
|
||||
|
||||
LFM2.5-350M publishes IFEval 76.96 against Qwen3.5-0.8B's 59.94, and BFCLv3 44.11
|
||||
against 35.08 — better at instruction-following and structured output, at 2/3 the
|
||||
size. Those numbers are real and they are **English**. Every benchmark in that
|
||||
table except Multi-IF is English-only.
|
||||
|
||||
In Russian, with a 300-token budget and temperature 0:
|
||||
|
||||
- **350M**, «Столица Франции? Ответь кратко.» → *«Сторзит в Париже.»* — `Сторзит` is
|
||||
not a word; it is invented morphology.
|
||||
- **350M**, asked to read back a reminder → a fortune cookie about being attentive
|
||||
and confident. No reminder in it.
|
||||
- **230M**, «Привет, как дела?» → answered **in Spanish**.
|
||||
|
||||
The 230M beating the 350M six-fold on routing (33.8% vs 5.2%) is the other tell:
|
||||
when the larger sibling collapses like that it is format compliance failing, not
|
||||
reasoning.
|
||||
|
||||
This is a pretraining gap, not a fine-tuning gap. Teaching Russian to a 350M from
|
||||
near-zero is not an afternoon on a Colab, which was the premise worth checking.
|
||||
|
||||
## Why this vindicates the 1.7B CPT
|
||||
|
||||
Stock Qwen3-1.7B, untrained and unprompted, answers all three probes in fluent
|
||||
correct Russian. What it gets wrong is the persona: *«Привет! Я рад, что ты здесь»*
|
||||
— `рад` is masculine and Maven needs `рада`. That is the right kind of remaining
|
||||
problem, and it is exactly what the CPT (Vikunja #122) is for.
|
||||
|
||||
The 1.7B was the correct model choice. What was wrong was treating it as a
|
||||
**blocker**: stock already beats what was deployed, so it ships now and gets
|
||||
swapped again when the CPT lands.
|
||||
|
||||
## Caveats
|
||||
|
||||
- Routing is one run per model, not three. The gaps between families are far larger
|
||||
than the run-to-run spread seen on the talk fixture, but the 2B-vs-1.7B gap (62.3
|
||||
vs 67.5) is not safe to call on one run.
|
||||
- ~~The routing numbers only reach production once the LLM router is wired on. It is
|
||||
still `nil`.~~ **Resolved the same evening:** the LLM router is wired at `voice.go:214`
|
||||
behind `voice.llm_router`, the default is on, and `deploy/mavend.json` sets it `true`.
|
||||
These numbers are the production path now. **Corrected 2026-08-02: the p50 ≈2.7s in the
|
||||
latency table above WAS a bench artifact.** It is contention on the shared llama-server,
|
||||
not the model. `docs/evals/2026-07-31-routing.md` line 61 says so, and measures the router at
|
||||
p50 825ms / p95 1.2s / max 3.0s. Cite that file for latency, not this one.
|
||||
- ~~`/mnt/hdd1/llms/LFM2.5/Qwen3-1.7B-UD-Q4_K_XL.gguf` is a 293 MB truncated download
|
||||
in the wrong directory.~~ **Deleted 2026-07-31.** The good 1.13 GB copy in `qwen3/` is
|
||||
what `deploy/mavend.json` loads.
|
||||
- Harness: `scratchpad/bakeoff.sh`, one server at a time, health-checked before each
|
||||
run, `/v1/models` recorded per run. Never run two LLM consumers at once — see the
|
||||
contamination note in `docs/evals/2026-07-31-talk.md`.
|
||||
@@ -0,0 +1,169 @@
|
||||
# Phrasing evaluation — 31-07-2026
|
||||
|
||||
How Maven words a nudge, measured instead of argued. Counterpart to
|
||||
`docs/evals/2026-07-31-routing.md`.
|
||||
|
||||
- Fixture + scorer: `internal/phraser/eval/` (`nudges_v1.json`, 15 cases; `eval.go`, `checks.go`)
|
||||
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing`
|
||||
- Model: Qwen3.5-0.8B Q4_K_M, the resident model. Not swapped.
|
||||
- Commit: `a40bc55` (prompt fix)
|
||||
|
||||
Every check is a string or length test a human can read and disagree with. No model
|
||||
grades another model here.
|
||||
|
||||
## Result
|
||||
|
||||
| | before | after |
|
||||
|---|---|---|
|
||||
| **cases passing every check** | **0/15** | **13/15** |
|
||||
| mood in enum | 6/15 | 15/15 |
|
||||
| Russian | 2/15 | 14/15 |
|
||||
| length (≤120 chars, ≤16 words) | 13/15 | 15/15 |
|
||||
| feminine self-reference | 15/15 | 15/15 |
|
||||
| no cringe | 13/15 | 15/15 |
|
||||
| on topic | 6/15 | 13/15 |
|
||||
| p50 latency | 11.4s | 11.4s |
|
||||
|
||||
Latency did not move and is not good. 11s to word one nudge on this box.
|
||||
|
||||
## The bug reproduced
|
||||
|
||||
Yes, exactly as reported. 7 of 15 messages were the literal string `"..."`, and one was
|
||||
`"full voice message"`. Both are text copied straight out of the prompt.
|
||||
|
||||
The system prompt said:
|
||||
|
||||
```
|
||||
Respond ONLY with valid JSON: {"response": "full voice message", "mood": "neutral"}
|
||||
```
|
||||
|
||||
and the user prompt said:
|
||||
|
||||
```
|
||||
Respond as JSON: {"response": "...", "mood": "..."}
|
||||
```
|
||||
|
||||
A 0.8B does not read `"..."` as "put your answer here". It reads it as the answer. The
|
||||
prompt was a worked example whose worked part was blank, so the model filled the slot by
|
||||
copying. This is the whole of finding 1.
|
||||
|
||||
## What else was wrong
|
||||
|
||||
Four separate faults, all prompt-side:
|
||||
|
||||
1. **Placeholder echo** (7 cases) — above.
|
||||
2. **Wrong language** (13/15 failed the language check). The prompt was entirely English
|
||||
and said "in the user's language (Russian or English)". The model picked English. It is
|
||||
never English: the nudge is spoken by a Russian piper voice.
|
||||
3. **Rule names are English identifiers.** `netdata_critical`, `service_down`, `break` went
|
||||
into the prompt raw. The model cannot nudge about a topic it has not been told in words,
|
||||
so 9/15 were off topic. The daemon knows what its own rules mean; now it says so.
|
||||
4. **Mood invented** (`"warm"`, twice). The enum was listed in a parenthesis at the end of
|
||||
an English sentence. Now it is its own line: "ровно одно из: neutral, happy, thinking,
|
||||
tired, confused."
|
||||
|
||||
Plus two non-prompt faults the run exposed:
|
||||
|
||||
- **The no-parse fallback was English.** When the model returned nothing usable, the body
|
||||
became `fmt.Sprintf("%s — %s", rule, sev)` — `"water — care"` — and that string went to
|
||||
a Russian TTS. Now it falls back to plain Russian.
|
||||
- **Durations were English.** `humanDur` returns "3 hours"; it was landing verbatim inside
|
||||
Russian sentences. Nudges now use a Russian formatter.
|
||||
|
||||
## Three iterations, and what each taught
|
||||
|
||||
| | score | change |
|
||||
|---|---|---|
|
||||
| baseline | 0/15 | — |
|
||||
| iter 1 | 2/15 | Russian prompt, filled-in examples, Russian durations |
|
||||
| iter 2 | 11/15 | required keyword per rule, one example instead of five, Russian fallback |
|
||||
| iter 3 | **13/15** | examples moved to topics that are not rules |
|
||||
|
||||
The interesting step is 1 → 2. Fixing the placeholder did not fix the disease, it moved it:
|
||||
the model stopped copying `"..."` and started copying my first example instead. Five nudges
|
||||
in a row came back as `"Ты не пил воду три часа. Налей стакан."` regardless of the rule.
|
||||
|
||||
**A small model copies the nearest concrete text in its prompt.** That is one failure mode
|
||||
with two symptoms. The fix that stuck was making the examples about laundry and a laptop
|
||||
battery — topics no rule ever produces, so copying them is visible in the score rather than
|
||||
invisibly passing the water cases.
|
||||
|
||||
## Do not oversell 13/15
|
||||
|
||||
Seven of the thirteen passes are the **deterministic fallback**, not the model:
|
||||
`"Напоминаю: таблетки."`, `"Сервис не отвечает."`, `"Критический алярм: проверь диск."`,
|
||||
`"Ты давно не пил воду."`. Those are strings this commit added to Go. The model returned
|
||||
nothing parseable and the fallback scored.
|
||||
|
||||
So the honest reading is roughly **6/15 from the model, 7/15 from a fallback, 2/15 failing**.
|
||||
The prompt fix is real — `"..."` is nearly gone and the language and mood checks are clean —
|
||||
but a large part of the jump is that failure now degrades into Russian instead of into
|
||||
`"water — care"`. That is a genuine improvement for the operator and a weak one for the model.
|
||||
|
||||
The two remaining failures: one `"..."` recurrence (`routine-stretch`) and one meal nudge
|
||||
that never says food.
|
||||
|
||||
## Tried and reverted: an example-led nudge prompt (#393)
|
||||
|
||||
The idea was that a 0.8B copies examples better than it follows rules, so the nudge prompt
|
||||
was rewritten to lead with five on-topic examples (water, break, pills, morning, service) and
|
||||
the prose rules were compressed to pay for the tokens: 1190 chars down to 986.
|
||||
|
||||
It measured **worse**, three runs each side, same llama-server, same fixture:
|
||||
|
||||
| run | before | after |
|
||||
|---|---|---|
|
||||
| 1 | 12/15 (address 14) | 11/15 (address 13) |
|
||||
| 2 | 13/15 (address 15) | 12/15 (address 15) |
|
||||
| 3 | 14/15 (address 15) | 11/15 (address 12) |
|
||||
|
||||
`feminine` and `hisgender` were 15/15 on all six runs, so they measure nothing here. The
|
||||
regression is all in `address`: 44/45 before, 40/45 after. Formal "вы"/"ваше" and plural
|
||||
imperatives came back, and so did `"..."`.
|
||||
|
||||
Two likely causes, both about the same thing — **examples do not carry a prohibition**. The
|
||||
old prompt spent a whole sentence on «говоришь на "ты", в единственном числе»; the new one
|
||||
demoted that to one item in a long "никогда" list, and the model stopped obeying it. And
|
||||
making the examples on-topic let their *wording* leak: a break case came back as
|
||||
«Вы давно не пили воду. Выпей стакан.» — the water example, verbatim, in the wrong slot.
|
||||
That is exactly the failure the laundry/laptop examples were chosen to avoid.
|
||||
|
||||
Change reverted. What survives is the measurement: a rule the model must obey needs its own
|
||||
sentence, and examples must stay off-topic. Also note the before side alone spans 12–14 of
|
||||
15 — this fixture cannot resolve anything smaller than about three cases.
|
||||
|
||||
## Broken, found, not fixed
|
||||
|
||||
1. ~~**`checkFeminine` only catches half the constraint.**~~ **Fixed** (#381). It scanned for
|
||||
masculine self-reference only, so three messages that addressed the *owner* in the feminine
|
||||
("ты давно не отдыхал**а**") scored clean. There is now a second check, `hisgender`: a
|
||||
feminine past-tense verb (-ла/-лась) in a sentence addressed to him ("ты", "тебе", "твой")
|
||||
fails, unless the verb is hers ("я заметила", "напомнила тебе"). It is a suffix rule, not a
|
||||
parser — see the comment in `checks.go` for what it misses. A fresh 15-case run after adding
|
||||
it scored **12/15** with `hisgender` 15/15; the model did not repeat the feminine address in
|
||||
that sample, and the check is pinned by unit tests on the recorded bad strings instead.
|
||||
2. **Grammar is not checked at all, and it is bad.** `"Он не ел 11 дней"` (it was 11 hours),
|
||||
`"Сонуждились 7 дней"` (not a word), `"Они забыли воду"` (wrong person entirely). Every
|
||||
one of these passes all six checks. The fixture measures properties, not fluency, and at
|
||||
0.8B fluency is the binding constraint.
|
||||
3. **Unit confusion.** The model turns hours into days about a third of the time. The
|
||||
prompt now says "11 ч"; it reads it as days.
|
||||
4. **11s p50.** Unchanged and untouched here. A nudge the model takes eleven seconds to
|
||||
word has missed its moment. Worth its own task.
|
||||
5. **The keyword hint is close to teaching to the test.** `ruleKeywords` names the word the
|
||||
on-topic check looks for. It is defensible — the daemon genuinely knows its rule topics
|
||||
and the model genuinely cannot infer them from `netdata_critical` — but the on-topic
|
||||
number is softer than the others because of it.
|
||||
|
||||
## Next steps
|
||||
|
||||
1. ~~**Add a second-person gender check**~~ — done, `hisgender` in `checks.go` (#381).
|
||||
2. **Decide whether the fallback should count as a pass.** Right now `Score` cannot tell a
|
||||
model answer from a fallback. Either mark fallback bodies in `PhrasedNudge` or count them
|
||||
in their own column. Without that, any future prompt change can score well by failing
|
||||
more.
|
||||
3. **Attack the 11s.** Nudge phrasing is short and non-interactive; thinking off is the first
|
||||
thing to try, as it was for routing (#376).
|
||||
4. **Re-measure when #122 lands.** The CPT'd Qwen3-1.7B is the target. 13/15 with seven
|
||||
fallbacks is the floor it has to beat, and the fluency problems above are the ones a
|
||||
bigger, Russian-trained checkpoint should actually fix.
|
||||
@@ -0,0 +1,250 @@
|
||||
# Note recall evaluation — 31-07-2026
|
||||
|
||||
The operator's goal is that Maven "memorize/note things … and know more about me/world". This
|
||||
measures whether the note/recall path delivers that.
|
||||
|
||||
- Fixture + scorer: `internal/memory/recalleval/` (`ru_recall_v1.json`, 30 cases)
|
||||
- Reproduce: `make eval-recall` — hash ratchet always, ONNX when `deps/` is present
|
||||
- Commit: `43470ab` (harness)
|
||||
|
||||
Each case inserts its own 3 notes **plus 12 shared filler notes** into a fresh store, embeds the
|
||||
query, takes the top 3 — the read path `cmd/mavend/voice.go` runs for `IntentQuery`. Filler is
|
||||
load-bearing: with 3 notes and a top-3 search, recall@3 is 100% by construction. 25 answerable
|
||||
cases (paraphrased queries, homelab and preference content, 9 with a plausible second note) and 5
|
||||
that must recall **nothing**. `TestFixtureIsParaphrased` fails the build if a query shares over half
|
||||
its words with its note; equal-score ties count as ties, not recall.
|
||||
|
||||
## Results
|
||||
|
||||
| | recall+hash (CI ratchet) | recall+onnx (deployed) |
|
||||
|---|---|---|
|
||||
| **recall@1** | 36.0% (9/25) | **60.0% (15/25)** |
|
||||
| recall@3 | 76.0% (19/25) | 80.0% (20/25) |
|
||||
| **answered after the 0.55 gate** | **0.0% (0/25)** | **48.0% (12/25)** |
|
||||
| wrong note on top / tie on top | 9 / 7 | 10 / 0 |
|
||||
| ranked first, then silenced by the gate | 9 | 3 |
|
||||
| **false recall** | 0/5 | **1/5 (20%)** |
|
||||
| top-1 score when right, min / median | n/a | 0.559 / 0.678 |
|
||||
| top-1 when it must stay silent, median / max | 0.000 / 0.144 | 0.470 / **0.567** |
|
||||
| RU / EN / `hard` cases passed | 4/24 / 1/6 / 0/11 | 13/24 / 3/6 / 2/11 |
|
||||
| latency p50 / p95 / max | 49µs / 70µs | 59ms / 148ms / 194ms |
|
||||
|
||||
Never compare a hash-embedder number to an ONNX one — the hash floor is lexical and exists only so
|
||||
CI has a deterministic ratchet with no model files.
|
||||
|
||||
## Findings
|
||||
|
||||
### 1. Real recall is 48%, not 60%
|
||||
|
||||
The right note ranks first 60% of the time, but the daemon only *says* it 48% of the time — three
|
||||
more cases rank first and are then silenced by `voice.go:776`'s `queryMinScore`. **Roughly one
|
||||
useful question in two gets "не знаю".** This is not a working memory yet.
|
||||
|
||||
### 2. The gate cannot separate a real recall from a false one — the distributions overlap
|
||||
|
||||
Right-note top-1 scores start at **0.559**. Must-stay-silent top-1 scores reach **0.567**. No
|
||||
threshold keeps every real recall and rejects every false one. From the sweep: gate 0.50 → 13/25
|
||||
answered, 1/5 false; **0.55 (default) → 12/25, 1/5**; **0.60 → 10/25, 0/5**; 0.70 → 5/25, 0/5. What
|
||||
the data says about `DefaultQueryMinScore` (`internal/config/config.go:392`): **0.55 is
|
||||
slightly too loose** — it admits one confident wrong answer ("как зовут сестру моего коллеги"
|
||||
recalls "выучил пару аккордов на гитаре" at 0.567), which the spec ranks as worse than a gap. 0.60
|
||||
silences all five and costs 8 points of real recall. Left alone as instructed; the overlap means
|
||||
the threshold is the wrong dial anyway (finding 3).
|
||||
|
||||
### 3. Filler notes outrank the right answer — the model scores similarity, not relevance
|
||||
|
||||
`models/embedder/` is **paraphrase-multilingual-MiniLM-L12-v2** (`Makefile:119`), a *symmetric*
|
||||
paraphrase model. It scores "do these sentences look alike", not "does this passage answer this
|
||||
question", so question-shaped queries drift to whatever note is stylistically closest. "из-за чего
|
||||
кончилось место" and "откуда берётся токен бота" both return `выучил пару аккордов на гитаре`
|
||||
(0.730, 0.729); "как я восстановил конфиги" returns a bootloader note at 0.703 with the right note
|
||||
not even in the top 3. An unrelated guitar note beating a homelab note at 0.73 is not a tuning
|
||||
problem — an asymmetric retrieval model (`multilingual-e5-small`, with `query:` / `passage:`
|
||||
prefixes) is the targeted fix, and it would move findings 1 and 2 together. Separately:
|
||||
`deploy/mavend.json:39` loads a 470MB fp32 `model.onnx` while `make download-embedder` fetches
|
||||
`model_quantized.onnx` — not the same file.
|
||||
|
||||
`hard` cases score **2/11**: every one is a query where the operator did not reuse his own words.
|
||||
That is the normal case weeks later, and exactly what docs/design.md's "recall when relevant" promises.
|
||||
|
||||
### 4. The memory-store recall branch is dead for notes
|
||||
|
||||
`voice.go:776` only reaches `h.memStore.Search` when the notes-RAG top score is already below
|
||||
`queryMinScore`, and `bestRecall` (`cmd/mavend/recall.go:19`) then applies the **same** gate to the
|
||||
same vector. A note is indexed in both places with the same embedding, so if it failed the gate in
|
||||
`QueryNotes` it fails again here — the branch can only ever return a **fact**. Its comment calls it
|
||||
"additive"; for notes it is not.
|
||||
|
||||
**Fixed (Vikunja #373).** The memory pass now runs *first*, as one search over notes and facts with
|
||||
one gate, so whichever memory is clearly the best match answers — note or fact. The notes-only pass
|
||||
stays behind it for notes the vector index does not hold. No threshold changed, so the set of
|
||||
questions Maven answers is the same; only which memory answers them. The fixture gained two mixed
|
||||
note+fact cases (`ru-mixed-031`, `ru-mixed-032`), which is why the counts below are out of 27
|
||||
answerable cases and not 25: hash recall@1 36.0% (9/25) → 37.0% (10/27), e5 recall@1 72.0% (18/25) →
|
||||
70.4% (19/27) with answered-after-gate 68.0% → 66.7% and false recall unchanged at 1/5.
|
||||
|
||||
### 5. Ranking has no recency or type signal, and the store is not the bottleneck
|
||||
|
||||
`internal/store/notes.go:67` sorts by cosine and uses `ts` only to break an exact float tie, which
|
||||
never happens; `kind` never enters the ranking. Meanwhile `TestPersistentStoreScoresTheSame` scores
|
||||
sqlite-backed `store.MemoryStore` and `memory.InMemoryStore` identically — both full-scan cosine
|
||||
(`internal/store/memory.go:64`) at ~150µs over 42 rows against a ~59ms query embed. An ANN index is
|
||||
not the problem to solve.
|
||||
|
||||
## Re-measured after the embedder swap — 31-07-2026, later the same day
|
||||
|
||||
Changed: `models/embedder/` is now **multilingual-e5-small** (quantized, 118MB), with `query: ` in
|
||||
front of a question and `passage: ` in front of a stored note (Vikunja #371). `deploy/mavend.json`
|
||||
and `make download-embedder` now name the same file, and it is the quantized one — that is what the
|
||||
column below measures (Vikunja #372). Everything else is unchanged: same fixture, same store, same
|
||||
0.55 gate. The old column is the baseline and is left as it was.
|
||||
|
||||
| | recall+onnx, MiniLM (baseline) | recall+onnx, e5-small (new) |
|
||||
|---|---|---|
|
||||
| **recall@1** | 60.0% (15/25) | **72.0% (18/25)** |
|
||||
| recall@3 | 80.0% (20/25) | 84.0% (21/25) |
|
||||
| **answered after the 0.55 gate** | 48.0% (12/25) | **72.0% (18/25)** |
|
||||
| wrong note on top / tie on top | 10 / 0 | 7 / 0 |
|
||||
| ranked first, then silenced by the gate | 3 | 0 |
|
||||
| **false recall** | 1/5 (20%) | **5/5 (100%)** |
|
||||
| top-1 score when right, min / median | 0.559 / 0.678 | 0.791 / 0.857 |
|
||||
| top-1 when it must stay silent, median / max | 0.470 / 0.567 | 0.815 / 0.835 |
|
||||
| RU / EN / `hard` cases passed | 13/24 / 3/6 / 2/11 | 14/24 / 4/6 / 5/11 |
|
||||
| latency p50 / p95 / max | 59ms / 148ms / 194ms | 18ms / 37ms / 49ms |
|
||||
|
||||
### What moved
|
||||
|
||||
Ranking got better and got faster. Half the previously-unwinnable `hard` cases now pass (2/11 →
|
||||
5/11), the guitar note no longer beats the docker-logs note, and the gate stops silencing notes that
|
||||
already ranked first. The quantized e5 is also ~3x quicker than the fp32 MiniLM it replaces.
|
||||
|
||||
### What got worse: the gate is now a no-op
|
||||
|
||||
e5 packs every cosine into a narrow high band. Right-note scores start at 0.791; must-stay-silent
|
||||
scores reach 0.835. **The distributions still overlap, and now they overlap above the gate**, so
|
||||
0.55 admits everything and false recall goes from 1/5 to 5/5. The sweep:
|
||||
|
||||
```
|
||||
gate 0.50–0.70: answered 18/25 (72%) false recall 5/5
|
||||
gate 0.80: answered 17/25 (68%) false recall 4/5
|
||||
gate 0.90: answered 0/25 ( 0%) false recall 0/5
|
||||
```
|
||||
|
||||
There is no value that keeps real recall and rejects made-up questions — same conclusion as before,
|
||||
now with a wider band and no room at all. `query_min_score` was left at 0.55 as instructed. **The
|
||||
recommendation is to leave it there and stop tuning it**: any number under ~0.79 is a no-op and
|
||||
anything above starts cutting real recall long before it stops the false ones. The fix is a margin
|
||||
gate (`top1 − top2 > δ`), next-steps item 3, which is now the top item.
|
||||
|
||||
### The prefixes did not do the work
|
||||
|
||||
A control run with both prefixes set to the empty string scored the **same** recall@1 (72%), a
|
||||
slightly better recall@3 (88%) and the same 5/5 false recall. So on this fixture the gain comes from
|
||||
the model, not from the `query:` / `passage:` split. The prefixes are kept because they are how e5
|
||||
was trained and the split is the right shape for the read path, but they are not worth defending on
|
||||
this evidence — a bigger fixture may say otherwise.
|
||||
|
||||
### Stored vectors from the old model are now junk
|
||||
|
||||
Cosine between a MiniLM vector and an e5 vector means nothing. Every row already in `notes` and in
|
||||
the vector memory table was written by the old model, so after this deploy they will score as noise
|
||||
against a new query. A live database needs every note and fact re-embedded before recall works at
|
||||
all. Filed as its own task.
|
||||
|
||||
## Margin gate — 31-07-2026, third run
|
||||
|
||||
Next-steps item 3, done. The absolute gate is replaced by a **margin gate**: answer only when the
|
||||
top hit beats the runner-up by more than delta (`top1 − top2 > δ`). Same fixture, same e5 embedder,
|
||||
same store as the run above. `internal/memory/gate.go` holds the check; both read paths call it
|
||||
(`cmd/mavend/recall.go` and the notes-RAG branch in `voice.go`). New knob `voice.query_min_margin`
|
||||
in `deploy/mavend.json`, default 0.008.
|
||||
|
||||
### Why the absolute gate could not work, in one line of data
|
||||
|
||||
The harness now prints the margin distributions, and they barely overlap where the raw scores
|
||||
overlap completely:
|
||||
|
||||
| | top-1 score | margin (top1 − top2) |
|
||||
|---|---|---|
|
||||
| right note first (n=18) | min 0.810, median 0.862, max 0.890 | min 0.001, median 0.029, max 0.053 |
|
||||
| must stay silent (n=5) | min 0.795, median 0.815, max 0.835 | min 0.000, median 0.002, **max 0.019** |
|
||||
|
||||
Four of the five must-be-silent cases have a margin at or under 0.002 — when there is nothing to
|
||||
recall, e5 finds several notes equally close and no clear winner. That is the signal the absolute
|
||||
score throws away.
|
||||
|
||||
### The delta sweep
|
||||
|
||||
Absolute gate held at 0.55 throughout.
|
||||
|
||||
```
|
||||
delta 0.000: answered 18/25 (72%) false recall 5/5
|
||||
delta 0.002: answered 17/25 (68%) false recall 3/5
|
||||
delta 0.005: answered 17/25 (68%) false recall 2/5
|
||||
delta 0.008: answered 17/25 (68%) false recall 1/5 <- chosen
|
||||
delta 0.010: answered 15/25 (60%) false recall 1/5
|
||||
delta 0.012: answered 14/25 (56%) false recall 1/5
|
||||
delta 0.015: answered 12/25 (48%) false recall 1/5
|
||||
delta 0.020: answered 11/25 (44%) false recall 0/5
|
||||
delta 0.025: answered 9/25 (36%) false recall 0/5
|
||||
delta 0.030: answered 8/25 (32%) false recall 0/5
|
||||
delta 0.040: answered 4/25 (16%) false recall 0/5
|
||||
delta 0.050: answered 2/25 ( 8%) false recall 0/5
|
||||
delta 0.060: answered 0/25 ( 0%) false recall 0/5
|
||||
```
|
||||
|
||||
### Chosen: δ = 0.008
|
||||
|
||||
It is the best point on the frontier, not a taste call. **0.008 dominates 0.010, 0.012 and 0.015
|
||||
outright** — same 1/5 false recall, 8 to 20 points more real recall. Everything below it buys recall
|
||||
back only by admitting more false recalls (0.005 → 2/5, 0.002 → 3/5). The next real improvement is
|
||||
0.020 at 0/5 false, and it costs 24 points of recall to get there.
|
||||
|
||||
The brief's bar was "recall above 60% with false recall at 1/5 or better". 0.008 clears it with room:
|
||||
68% and 1/5.
|
||||
|
||||
### Before / after
|
||||
|
||||
| | absolute gate 0.55 (previous) | margin gate δ=0.008 |
|
||||
|---|---|---|
|
||||
| recall@1 (ranking, ungated) | 72.0% (18/25) | 72.0% (18/25) — unchanged, the gate does not rank |
|
||||
| **answered after the gate** | 72.0% (18/25) | **68.0% (17/25)** |
|
||||
| **false recall** | **5/5 (100%)** | **1/5 (20%)** |
|
||||
| fixture cases passed | 18/30 | **21/30** |
|
||||
|
||||
Four false recalls removed for one real answer. That is the trade the spec asks for — she is not a
|
||||
guesser-of-truth. The one survivor is `en-pref-025` ("should i be offered wine"), which recalls a
|
||||
filler note at 0.796 with a 0.019 margin: the widest silent-case margin in the fixture, and it sits
|
||||
inside the real-recall range, so no delta removes it without taking real answers with it.
|
||||
|
||||
### Does the absolute cutoff still earn its keep? Marginally — kept
|
||||
|
||||
On this fixture with e5 it is a **no-op**: the lowest right-note score is 0.791, so 0.55 rejects
|
||||
nothing the margin does not already reject. It is kept for two reasons, neither glamorous. It still
|
||||
does real work for the hash embedder (its own sweep shows answers dropping from 16% to 0% between
|
||||
0.30 and 0.50), and it is the only thing standing between the user and a reply built from a store
|
||||
where everything is far away but one row happens to be a little less far — a near-empty database, or
|
||||
the stale-vector case below. Cheap insurance, no measured cost. If a later embedder makes it bite,
|
||||
the sweep is one command.
|
||||
|
||||
### Caveat on the numbers
|
||||
|
||||
Five must-be-silent cases is a thin basis for a 4-point decision. 1/5 and 2/5 differ by one case.
|
||||
The shape of the frontier is trustworthy — margins separate, absolute scores do not — but δ=0.008
|
||||
itself should be re-read off a bigger fixture (next-steps item 6) before anyone defends the third
|
||||
decimal.
|
||||
|
||||
## Next steps — ordered by value-to-risk; nothing here is a decision
|
||||
|
||||
1. **Swap the embedder to `multilingual-e5-small` with `query:`/`passage:` prefixes.** One config
|
||||
change plus a prefix in `onnxembedder.go`, re-measurable in one command.
|
||||
2. **Re-run `make eval-recall`, then set the gate from the sweep** — not before. Any
|
||||
`query_min_score` picked against today's embedder describes a model on its way out.
|
||||
3. ~~**Replace the absolute-score gate with a margin gate**~~ — done, see the section above.
|
||||
δ=0.008, false recall 5/5 → 1/5.
|
||||
4. **Delete or repair the dead `memStore` branch** at `voice.go:776` — search before the gate,
|
||||
gate it separately, or restrict it to facts and say so.
|
||||
5. **Add a mild time decay to ranking** — the newest statement of a preference is the true one.
|
||||
6. **Grow the fixture from real misses.** 30 cases can rank two embedders, not trust 4 points.
|
||||
7. **Re-measure end to end.** Recall is gated twice — the utterance must first route to `query`,
|
||||
which the routing eval puts at ~50%. The product is ~24%, and that is what he experiences.
|
||||
@@ -0,0 +1,300 @@
|
||||
# Routing evaluation — 31-07-2026
|
||||
|
||||
Settles Vikunja **#319** ("measure classifier vs LLM router before flipping"). Everything
|
||||
below is measured against one held-out fixture, not argued from the code.
|
||||
|
||||
- Fixture + scorer: `internal/router/eval/` (`ru_routing_v1.json`, 76 cases; `eval.go`)
|
||||
- Reproduce: `make eval-router` (classifier baselines) and
|
||||
`MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-router` (adds the LLM configurations)
|
||||
- Commits: `c7c4422` (fixture), `d34fdf4` (ONNX baseline), `46259b4` (LLM baseline)
|
||||
|
||||
## Why a new fixture
|
||||
|
||||
`cmd/mavend/eval_scenarios_test.go` could not answer #319: it asserts daemon-side *safety*
|
||||
invariants over already-normalized decisions, so it never exercises routing. And the only
|
||||
utterance corpus that existed — `models/seeds/*.txt` — is the classifier's own training set.
|
||||
Scoring a nearest-centroid classifier there measures memorisation of frozen centroids, which
|
||||
is exactly the illusion behind `voice.go:211`'s "the classifier handles routing reliably".
|
||||
|
||||
`TestFixtureIsHeldOut` fails the build if any fixture utterance appears verbatim in the seed
|
||||
corpus. The fixture is a **contract, not a snapshot**: cases the cascade fails today stay in
|
||||
the file and fail loudly.
|
||||
|
||||
## Results
|
||||
|
||||
| | classifier+hash | classifier+onnx | llm-only (0.8B) | cascade+llm (0.8B) |
|
||||
|---|---|---|---|---|
|
||||
| **intent-only accuracy** | 17.1% | 36.8% | 48.7% | **50.0%** |
|
||||
| full accuracy (intent+slots+gate) | 17.1% | 36.8% | 23.7% | 32.9% |
|
||||
| RU | 10/61 | 25/61 | 13/61 | 18/61 |
|
||||
| EN | 3/15 | 3/15 | 5/15 | 7/15 |
|
||||
| `hard` tag | 0/11 | 4/11 | — | — |
|
||||
| false clarify (asked, shouldn't) | 63 | 21 | 0 | 2 |
|
||||
| **missed clarify (guessed, shouldn't)** | **0 / 6** | **5 / 6** | **6 / 6** | **6 / 6** |
|
||||
| route errors | 0 | 0 | 2 | 0 |
|
||||
| **p50 / p95 / max latency** | 9µs / 14µs | **31ms / 71ms** | 850ms / 1.56s / 3.1s | **825ms / 1.20s / 3.0s** |
|
||||
|
||||
`classifier+hash` is the CI ratchet (deterministic, no model files). `classifier+onnx` is what
|
||||
homesrv runs today. `cascade+llm` is the wiring #320 proposes: stage-0 grammar → resident
|
||||
model → classifier as failure floor.
|
||||
|
||||
Never compare a hash-embedder run to an ONNX one.
|
||||
|
||||
## Re-measured after the prompt fix
|
||||
|
||||
The table above is the **baseline at commit `46259b4`**, kept as-is. The prompt fix (query
|
||||
tested before fact, plus `repeat_penalty` and a bounded grammar string) was then measured on
|
||||
an otherwise idle box — no other eval sharing llama-server, so these latencies are real
|
||||
rather than contention.
|
||||
|
||||
| | llm-only (0.8B) | cascade+llm (0.8B) | llm-only, thinking off |
|
||||
|---|---|---|---|
|
||||
| **intent-only accuracy** | 48.7% → **61.8%** | 50.0% → **63.2%** | **67.1%** |
|
||||
| full accuracy (intent+slots+gate) | 23.7% → **38.2%** | 32.9% → **47.4%** | **42.1%** |
|
||||
| route errors | 2 → **0** | 0 → 0 | **0** |
|
||||
| p50 / p95 latency | **1.08s / 1.55s** | **1.04s / 1.53s** | **0.93s / 1.41s** |
|
||||
|
||||
Three things this run settles:
|
||||
|
||||
1. **The prompt fix holds.** An earlier contended run reported 60.5% / 36.8% for llm-only;
|
||||
the quiet run gives 61.8% / 38.2%. Close enough to call the gain real, and the earlier
|
||||
run's 4-5s latency figures were contention, not the model.
|
||||
2. **`query→fact` fell from ×15 to ×7**, and both unparseable replies are gone. Zero route
|
||||
errors in every LLM configuration.
|
||||
3. **`note→fact ×4` is real, not noise.** It shows up in the quiet run too. The agent that
|
||||
wrote the prompt fix suspected its own change might have caused it by pulling assertive
|
||||
`запиши что…` phrasings toward fact, and that suspicion stands — all five `ru-note-*`
|
||||
cases now land on fact. Tracked as Vikunja #375.
|
||||
|
||||
The `thinking off` column above read as the best configuration measured so far (Vikunja #376).
|
||||
**It was wrong** — see the controlled re-run below. Ignore that column.
|
||||
|
||||
Still `6 / 6` missed clarify — the router has no way to say "I don't know" (Vikunja #359).
|
||||
That is unchanged by anything here.
|
||||
|
||||
## Thinking off — 31-07-2026, controlled re-run (Vikunja #376)
|
||||
|
||||
The "thinking off wins by 6 points" observation above **does not hold**. It was a measurement
|
||||
artefact, and the earlier table's `thinking off` column should be ignored.
|
||||
|
||||
The thinking-off variant was scored by a hand-rolled HTTP client living in the test file
|
||||
instead of `llm.Client`. That copy did not send `repeat_penalty`, which the real router does
|
||||
send (`routeRepeatPenalty = 1.15`). So the two columns differed on two axes at once, and the
|
||||
one that mattered was the penalty, not the thinking mode.
|
||||
|
||||
Re-measured with everything else held equal — same fixture, same prompt, same grammar, same
|
||||
sampling, same idle box, the three configurations run back to back and never concurrently:
|
||||
|
||||
| | llm-only, thinking on | llm-only, thinking off | cascade+llm |
|
||||
|---|---|---|---|
|
||||
| intent-only accuracy | 59.2% (45/76) | 59.2% (45/76) | 61.8% (47/76) |
|
||||
| full accuracy (intent+slots+gate) | 38.2% (29/76) | 38.2% (29/76) | 57.9% (44/76) |
|
||||
| route errors | 3 | 3 | 0 |
|
||||
| grammar violations | 3 (all 3 route errors) | 3 (same 3 cases) | 0 |
|
||||
| missed clarify | 5 / 6 | 5 / 6 | 5 / 6 |
|
||||
| p50 latency | 836ms | 920ms | 810ms |
|
||||
| p95 latency | 1.41s | 2.00s | 1.31s |
|
||||
|
||||
Thinking off is not just a tie on the headline numbers — it is identical case for case, with
|
||||
the same confusion matrix and the same three unparseable replies. The latency difference is
|
||||
run-to-run noise on one box, and it points the wrong way here.
|
||||
|
||||
The reason is simpler than any accuracy argument: **this llama-server build ignores the
|
||||
request-level thinking switch for this model.** Probed directly against the running server
|
||||
with `chat_template_kwargs.enable_thinking = false`, `chat_template_kwargs.thinking = false`
|
||||
and top-level `reasoning_budget = 0` — all three return a byte-identical answer with the
|
||||
thinking trace still in `reasoning_content`, and the server reports the prompt prefix as
|
||||
cached, meaning the rendered template did not change. There was never anything being turned
|
||||
off, which is also why the numbers match exactly.
|
||||
|
||||
Nothing was defaulted. `internal/llm` still has no `chat_template_kwargs` field, `VoiceConfig`
|
||||
has no thinking flag, and `deploy/mavend.json` is unchanged. The misleading third
|
||||
configuration is removed from `internal/router/eval` so the table it produced cannot be quoted
|
||||
again.
|
||||
|
||||
Two caveats worth saying out loud:
|
||||
|
||||
- **The fixture is 76 cases.** A 6-point difference on 76 cases is roughly 4-5 cases and would
|
||||
not have been worth trusting even if it had reproduced. This one was exactly 0 cases, which
|
||||
is a much easier call.
|
||||
- **This is one server build and one checkpoint** (`b9351`, Qwen3.5-0.8B Q4_K_M). If the
|
||||
#122 checkpoint or a newer llama.cpp does honour the switch, the question reopens — but it
|
||||
reopens as an unmeasured question, not as a 6-point win.
|
||||
|
||||
Phrasing was **not** measured. Whether thinking helps there is still open, and now also blocked
|
||||
on the same "can we even turn it off" question.
|
||||
|
||||
## Clock and calendar rule — 31-07-2026 (Vikunja #374)
|
||||
|
||||
`routeSystem` never said whether "который час" or "какое число завтра" are `system` or
|
||||
`query`, and `system→query ×4` showed up in every run. The rule added says: the clock and the
|
||||
calendar date themselves are `system`; what is *written in* the calendar or in memory
|
||||
("что у меня завтра", "какие есть напоминания") stays `query`; and a time named inside a
|
||||
request ("напомни завтра…") is just a detail of the request, not a reason for `system`.
|
||||
|
||||
That split is not a preference. In `cmd/mavend/voice.go` only `replySystem` owns the clock and
|
||||
the date formatter, so a clock question routed to `query` falls into the embedder + note RAG
|
||||
and answers "не знаю". The agenda, on the other hand, is answered by `ParseCalendarDate` +
|
||||
`CalendarEvents` *inside* the `query` branch, so that side has to stay `query`. The rule sits
|
||||
above the question test because every one of these utterances carries a question word and a
|
||||
later rule would never be reached.
|
||||
|
||||
The fixture is now 77 cases: one calendar-agenda case was added
|
||||
(`ru-query-019` "что у меня стоит в календаре на послезавтра", intent `query`) specifically so
|
||||
an over-broad system rule cannot pass unnoticed. The clock/date cases (`ru-sys-001/002/005`,
|
||||
`en-sys-001`) already existed.
|
||||
|
||||
Three runs, same box, back to back, never concurrently:
|
||||
|
||||
| | baseline | first rule (too broad) | rule as committed |
|
||||
|---|---|---|---|
|
||||
| llm-only intent-only | 59.2% (45/76) | 54.5% (42/77) | 59.7% (46/77) |
|
||||
| llm-only full | 38.2% | 35.1% | 39.0% |
|
||||
| llm-only route errors | 3 | 4 | 5 |
|
||||
| llm-only p50 | 1.09s | 0.91s | 0.93s |
|
||||
| cascade+llm intent-only | 61.8% (47/76) | 58.4% | 62.3% (48/77) |
|
||||
| cascade+llm full | 57.9% | 54.5% | 59.7% |
|
||||
| cascade+llm route errors | 0 | 0 | 0 |
|
||||
| cascade+llm p50 | 0.91s | 0.80s | 1.04s |
|
||||
|
||||
**The targeted bug is fixed and the headline number did not move.** `system→query ×4` is gone
|
||||
in both LLM configurations — the `time` and `date` tags go from 0/2 and 0/2 to 2/2 and 2/2 —
|
||||
but the model then over-applies the rule, and `query→system ×5` plus `reminder→system ×2`
|
||||
appear where they did not exist before. Net accuracy is a wash, inside the noise of a 77-case
|
||||
fixture.
|
||||
|
||||
The first attempt is shown because it is the honest history: it said "спрашивает время, дату
|
||||
или день недели → system" with no scope, which swept up reminders, and it cost 3-5 points. It
|
||||
was tightened once, on the reasoning that a rule capturing "напомни завтра в 7" is simply
|
||||
wrong, and not tuned further. The remaining `query/reminder → system` over-trigger is a new,
|
||||
separate weakness of the sub-1B model and deserves its own task rather than more prompt
|
||||
kneading against a held-out fixture.
|
||||
|
||||
The rule is kept. It is correct about what the daemon can answer, and the failure it replaces
|
||||
was silent ("не знаю" to "который час") while the one it introduces is loud.
|
||||
|
||||
## Findings
|
||||
|
||||
### 1. The resident model does route better — 50.0% vs 36.8%
|
||||
|
||||
docs/rearchitecture.md's premise holds; `voice.go:211`'s comment does not. **But the classifier is only
|
||||
~37% correct on held-out utterances, and the model only ~50%.** Neither is "reliable". The
|
||||
gap between them is real but both are far from a system you would describe as working.
|
||||
|
||||
### 2. It costs 27× the latency
|
||||
|
||||
p50 825ms vs 31ms, p95 1.2s, max 3.0s — on the same llama-server the phraser needs, before
|
||||
any phrasing happens. On the CPU/iGPU deploy target this is a trade, not a free win. The
|
||||
review's second-opinion caution was justified.
|
||||
|
||||
### 3. `query→fact ×15` is the dominant LLM failure — and it is a prompt bug
|
||||
|
||||
Four times the classifier's `×4` on the same axis. `routeSystem`'s decision order in
|
||||
`internal/router/llmrouter.go` reads:
|
||||
|
||||
```
|
||||
3. Сообщает или обновляет текущее состояние/событие → fact
|
||||
4. Хочет получить информацию → query
|
||||
```
|
||||
|
||||
Any utterance naming a fact key matches rule 3 first, so a *question about* past state
|
||||
("сколько воды я выпил с утра", "сколько раз я ел вчера") is classified as an *assertion of*
|
||||
that state — and a query becomes a confident wrong write. Reordering query above fact, or
|
||||
adding an explicit interrogative test, is the cheapest accuracy win available and needs no
|
||||
model change.
|
||||
|
||||
### 4. Neither path can refuse — the refusal lane is currently fiction
|
||||
|
||||
| | missed clarify | why |
|
||||
|---|---|---|
|
||||
| classifier+hash | 0 / 6 | cosine never clears 0.55 — refuses by accident |
|
||||
| classifier+onnx | 5 / 6 | better embeddings raise cosine everywhere; the gate stops separating |
|
||||
| LLM (any) | 6 / 6 | `llmrouter.go` hardcodes `Confidence: 1.0`, so stage 3 can never fire |
|
||||
|
||||
The deployed config confidently routes `сделай это` → **act** at 0.847, `ну это` → chat at
|
||||
0.808, `бэкап` → chat at 0.755, `потом` → system at 0.739. `сделай это` → act with unresolved
|
||||
anaphora is the destructive direction; the daemon's confirm gate is the only thing left.
|
||||
|
||||
This is the finding that should block #320. Flipping to the LLM router as-is does not improve
|
||||
the refusal lane — it removes it. Tracked as **#359**.
|
||||
|
||||
### 5. The 50.0% → 32.9% gap is entirely slots
|
||||
|
||||
The LLM path fills neither `Fn` nor `Time`: it returns `Slots.Text` for acts (the verb string,
|
||||
not an allowlist match), and `Extractor.Extract` never runs on an LLM decision at all. Any
|
||||
flip needs the extractor wired onto the LLM branch or every act and reminder arrives without
|
||||
its arguments.
|
||||
|
||||
### 6. The 2 route errors are a missing `RepeatPenalty`, not a grammar flaw
|
||||
|
||||
Both failures (`ru-act-006` "закрой жалюзи", `ru-chat-003` "расскажи анекдот про
|
||||
программистов") are the sub-1B repetition loop *inside* the grammar's `text` field:
|
||||
|
||||
> "Закрывание жалюзи — это действие, которое нужно выполнить. Если это не действие, то это
|
||||
> сообщение пользователя. Если это не действие, то это сообщение пользователя. …"
|
||||
|
||||
It runs to `MaxTokens: 128`, truncates the JSON mid-string, and `parseActions` fails →
|
||||
fallback to the classifier. `llm.Req` already has a `RepeatPenalty` field added for exactly
|
||||
this ("curbs the sub-1B 'тоже тоже тоже' loop") and `LLMRouter.Route` does not set it. Two
|
||||
lines.
|
||||
|
||||
Note the grammar's `string ::= "\"" ([^"\\] | "\\" .)* "\""` is unbounded, so nothing stops a
|
||||
1000-character `text`. Worth a length bound as well.
|
||||
|
||||
### 7. Two hypotheses tested and closed
|
||||
|
||||
- **Thinking mode is a non-issue.** Confirmed twice now, the second time properly — see the
|
||||
controlled re-run section. Grammar-constrained JSON lands in `reasoning_content` with
|
||||
`content` empty and `llm.Client`'s fallback handles it; the request-level switch does
|
||||
nothing on this build. `internal/llm` deliberately does **not** grow a
|
||||
`chat_template_kwargs` field.
|
||||
- **Runaway array repetition does not reproduce.** An isolated smoke test with a stripped
|
||||
grammar emitted `{"intent":"reminder"}` until `MaxTokens`; under the real `routeSystem`
|
||||
prompt the few-shot examples anchor it to one object. 2 errors in 76, not 76.
|
||||
|
||||
### 8. Incidental
|
||||
|
||||
- `ReminderGrammar` deliberately skips the extractor at stage 0; the daemon's `applyAction`
|
||||
parses the time downstream. The scorer counts those as `SlotsDeferred` rather than misses.
|
||||
- A local llama-server must bypass `http_proxy` — this box proxies loopback through a SOCKS
|
||||
bridge that answers 503. `noProxyLoopback` in the test handles it.
|
||||
- The onnxruntime `.so` was already vendored at `deps/onnxruntime-linux-x64-1.26.0`.
|
||||
|
||||
## Next steps
|
||||
|
||||
Ordered by ratio of value to risk. Nothing here is a decision — #320 stays open.
|
||||
|
||||
1. **Fix `routeSystem`'s decision order** (query above fact, or an explicit interrogative
|
||||
test). Largest single accuracy move, no model change, re-measurable in one command.
|
||||
Expected: most of `query→fact ×15`.
|
||||
2. **Set `RepeatPenalty` in `LLMRouter.Route`** and bound the grammar's `string` length.
|
||||
Removes both route errors.
|
||||
3. **Give the router a refusal signal — #359.** Blocks #320.
|
||||
- Classifier: the absolute-cosine gate does not survive a better embedder. A **margin**
|
||||
gate (`top1 − top2 > δ`) is the likely fix — ambiguous utterances should show flat
|
||||
distributions, which absolute cosine cannot see.
|
||||
- LLM: `Confidence: 1.0` must go. Either add an `unclear` intent to the grammar enum, or
|
||||
read logprobs, or gate on the classifier's margin *behind* the LLM decision.
|
||||
- Bar: `MissedClarify ≤ 1` without regressing full accuracy below 28/76.
|
||||
4. **Wire `Extractor.Extract` onto the LLM branch** so acts get `Fn` and reminders get
|
||||
`Time`. Closes the 50.0% → 32.9% slot gap.
|
||||
5. **Re-measure, then decide #320.** At p50 825ms a wholesale swap is probably the wrong
|
||||
shape; the honest candidate is LLM-for-queries with the classifier keeping the fast
|
||||
deterministic paths (stage-0 grammar hits, `system`, exact acts). That hypothesis is
|
||||
testable against this fixture by scoring a per-intent split.
|
||||
6. **Grow the fixture** as failures get understood. 76 cases with ≥5 per intent is enough to
|
||||
rank paths, not enough to trust a 2-point difference. Add cases from real misroutes
|
||||
(`CorrectMisroute` is already the append-only hook).
|
||||
7. **Second checkpoint when #122 lands.** The CPT'd Qwen3-1.7B is the target resident model;
|
||||
the same three configurations should be re-scored against it before it deploys. 0.8B's
|
||||
50.0% is the floor that checkpoint has to beat, and its latency is the number that decides
|
||||
whether the target is affordable at all.
|
||||
|
||||
## Open question worth naming
|
||||
|
||||
Both paths are under 50%. That is low enough that the interesting question may not be
|
||||
"classifier or model" but whether one-shot classification of a bare utterance is the right
|
||||
frame at all — `сделай это`, `потом`, `бэкап` are unanswerable without dialogue context, and
|
||||
`internal/router` currently sees none (`AnaphoraResolver` exists in `slots.go` but the
|
||||
cascade never calls it). A router that could ask one clarifying question and re-route on the
|
||||
answer would beat both numbers here without a better model.
|
||||
@@ -0,0 +1,150 @@
|
||||
# Conversational phrasing eval — 31-07-2026
|
||||
|
||||
Every score measured tonight, on the three paths the nudge eval never touched:
|
||||
chat, query-with-notes, and general knowledge.
|
||||
|
||||
**Short version: the plumbing got fixed and the score barely moved.** Grammar and
|
||||
Russian prompts together took the composite from ~9 to ~14 of 27. Everything
|
||||
still failing is the model not knowing things or not holding a constraint, and
|
||||
prompting is out of levers. Settles the measurement half of Vikunja #395 / #398 /
|
||||
#400.
|
||||
|
||||
## How to reproduce
|
||||
|
||||
```sh
|
||||
# llama-server: -c 4096 -ngl 99 -t 6, model /mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf
|
||||
MAVEN_LLM_URL=http://127.0.0.1:18099 no_proxy=127.0.0.1,localhost \
|
||||
deps/go/go/bin/go test -count=1 -timeout 40m \
|
||||
-run TestLLMTalkBaseline ./internal/phraser/eval/ -v
|
||||
```
|
||||
|
||||
Three runs per configuration, always. The fixture is 27 cases, so one reply
|
||||
changing moves the composite by 3.7 points — a single run cannot tell a real
|
||||
change from sampling noise. This was learned the expensive way: an earlier claim
|
||||
that "one nudge case fails every run" turned out to be three different cases
|
||||
across three runs.
|
||||
|
||||
**Run the box otherwise idle.** See the contamination note at the bottom.
|
||||
|
||||
## Composite, per configuration
|
||||
|
||||
| config | overall /27 | chat /9 | query /9 | knowledge /9 | canned fallbacks |
|
||||
|---|---|---|---|---|---|
|
||||
| baseline, no grammar | 7, 12, 7 | 1, 1, 0 | 2, 4, 2 | 4, 7, 5 | 0, 0, 0 |
|
||||
| + GBNF grammar (#398) | 14, 15, 8 | 1, 3, 0 | 5, 6, 3 | 8, 6, 5 | 0, 0, 0 |
|
||||
| + Russian prompts (#400) | 11, 17, 15 | 1, 5, 3 | 5, 6, 8 | 5, 6, 4 | 0, 0, 0 |
|
||||
| + truncation fix, 1000ch/768tok | 12, 13, 10 | 2, 2, 1 | 7, 7, 5 | 3, 4, 4 | 3, 3, 6 |
|
||||
| + rebalanced, 600ch/1024tok | **void — contaminated** | | | | |
|
||||
|
||||
"Canned fallbacks" counts replies that came back as the hardcoded `"не знаю."`
|
||||
or `"поговорили."`. It is not a check, it is a health signal: those strings mean
|
||||
the phraser gave up, and the eval scores them as ordinary bad replies.
|
||||
|
||||
## Per-check
|
||||
|
||||
| check | no grammar | + grammar | + RU prompts | + truncation fix |
|
||||
|---|---|---|---|---|
|
||||
| nonempty | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 |
|
||||
| ellipsis | 20, 19, 23 | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 |
|
||||
| lang | 13, 16, 15 | 23, 26, 26 | 25, 26, 25 | 26, 27, 27 |
|
||||
| feminine | — | — | 25, 24, 26 | 25, 25, 27 |
|
||||
| address | — | — | 21, 22, 22 | 22, 21, 22 |
|
||||
| ontopic | — | — | 17, 24, 18 | 17, 19, 14 |
|
||||
|
||||
`nonempty` reading 27/27 everywhere is not good news — it was a broken check.
|
||||
It tested for a non-blank string, so replies of literally `{` and `"15-16"`
|
||||
passed it. Fixed on `overnight/fix-truncation`; it needs a letter now.
|
||||
|
||||
## What each change actually bought
|
||||
|
||||
**GBNF grammar (#398) — the biggest single win.** Qwen3.5-0.8B writes
|
||||
`Thinking Process:` as plain text with no tags, `stripThink` only handles
|
||||
`</think>`, so the JSON never closed and the plain-text fallback shipped the
|
||||
literal reasoning. `ellipsis` went 20→27 and `lang` 13→26. The router had been
|
||||
using a grammar for ages; the phraser asking nicely in the prompt was the
|
||||
oversight.
|
||||
|
||||
**Russian prompts (#400) — modest, plus a large latency win.** Chat 1.3→3.0
|
||||
average, query 4.7→6.3, knowledge 6.3→5.0. All inside the run-to-run spread, so
|
||||
"probably better on the paths it targeted, not provable in three runs". p50
|
||||
latency dropped from ~11.5s to ~2.3s and that part is consistent across all
|
||||
three runs — shorter prompts, and she stopped emitting English reasoning first.
|
||||
|
||||
**Truncation fix — necessary, and did not help the score.** Two real bugs
|
||||
(replies of `{`, and a `nonempty` check that passed them), both fixed, and the
|
||||
composite went nowhere. A complete rambling wrong answer fails the same checks a
|
||||
truncated one did. Worth doing anyway: the daemon was shipping `{` to a
|
||||
text-to-speech voice.
|
||||
|
||||
## The truncation bug, since the cause was counter-intuitive
|
||||
|
||||
The grammar's `string ::= ... {0,400}` rule was the cause, not the token cap.
|
||||
Measured against Qwen3.5-0.8B at three caps — 256, 768 and 2048 — the reply came
|
||||
back **exactly 400 characters every time, cut mid-word** (`"Нужно записать и,"`).
|
||||
|
||||
Then I raised the bound to 1000 while the cap was 768 tokens and made it worse:
|
||||
Russian runs ~1.5 characters per token here, so generation died on the *token*
|
||||
cap instead, mid-object, and the new guard correctly refused it and shipped
|
||||
`"не знаю."` — 3, 3 and 6 fallbacks per run, from zero. **The two limits have to
|
||||
agree.** 600 characters needs ~400 tokens; the cap is 1024.
|
||||
|
||||
## Where the remaining failures live
|
||||
|
||||
`address` is stuck at 21-22 of 27 and `ontopic` at 14-19. Both resist prompting.
|
||||
|
||||
**The prompt now explicitly forbids exactly what she does.** It says never "вы",
|
||||
use the singular — and she writes `вашей`, `подождите`, `делаете`, `хотите`,
|
||||
`напишите`. Telling a 0.8B "never do X" does not work. Same for
|
||||
`feminine`: `я готов`, `я понял`, `я нашел`, `я заметил`, `я сказал`.
|
||||
|
||||
**Some of `ontopic` is the fixture, not the model.** `chat-how-are-you` got
|
||||
`"Привет! Я здесь, чтобы поговорить. Как дела сегодня?"` — a fine reply that
|
||||
fails because `want_any` is `[норм, хорош, порядк, тут, работ]`. It fails in
|
||||
every run, so it inflates the count. The `ontopic` column currently measures the
|
||||
fixture as much as the model. Not fixed yet, deliberately: changing it would
|
||||
break comparability with the runs above.
|
||||
|
||||
**Two replies worth reading, because they are not fixable by prompting:**
|
||||
|
||||
- Thunder and lightning: *"Скорость молнии — 8-10 тысяч километров в секунду, но
|
||||
звук — 300 метров в секунду, что делает молнию громче."* Confidently wrong,
|
||||
and it concludes lightning is *louder* rather than sound being *slower*.
|
||||
- "расскажи обо мне": *"Ты — прекрасное существо, с душой и вниманием… Спасибо за
|
||||
твою улыбку… О тебе — заповедь любви."* Sycophantic filler, zero information,
|
||||
and precisely the "not a relationship" non-goal.
|
||||
- Boiling an egg: `"15-16"` one run, `"1"` another. No unit, wrong number.
|
||||
|
||||
The first argues for reading instead of recalling (#403 — Kiwix retrieval scores
|
||||
8/8 on the same questions given English keywords). The second and third argue
|
||||
for templates on the paths where correctness matters (#392).
|
||||
|
||||
## Contamination note — how the last row got voided
|
||||
|
||||
I started the query-rewrite agent against the same llama-server the sweep was
|
||||
using, and assumed contention would only affect latency. It did not. The
|
||||
knowledge path collapsed to 0 of 9 with eight canned `"не знаю."` replies, p95
|
||||
tripled to 23.7s, and **the report still said "0 errors"**.
|
||||
|
||||
That is Vikunja #397, and it is worse than filed: a merely *busy* server
|
||||
produces a clean-looking report with a third of the fixture silently answering
|
||||
`"не знаю."`. `PhraseChat` and `PhraseQuery` swallow every failure and return a
|
||||
hardcoded string, so infrastructure trouble is indistinguishable from bad
|
||||
phrasing in the score. The talk test guards the *start* and *end* of a run with
|
||||
a model check, which catches a dead server but not a loaded one.
|
||||
|
||||
**Until #397 is fixed, treat any run made on a busy box as void.**
|
||||
|
||||
## Next
|
||||
|
||||
- Re-run 600ch/1024tok clean, to fill the void row.
|
||||
- Score `Qwen3.5-2B-UD-Q4_K_XL` (already at `/mnt/hdd1/llms/qwen3.5/`, never
|
||||
measured) on this fixture and the router fixture. Not the 4B — too big for
|
||||
this box, owner's call.
|
||||
- Newer sub-500M candidates (LFM2.5 200M/300M) are worth a run for routing.
|
||||
Note `docs/evals/2026-07-31-model-bakeoff.md` found LFM2.5-**1.2B** worse than
|
||||
Qwen3.5-0.8B at Russian routing and 2.4× slower — but those are a different,
|
||||
older generation, so that result does not predict the small ones.
|
||||
- Fix `chat-how-are-you`'s `want_any`, and re-baseline once, so `ontopic`
|
||||
measures the model.
|
||||
- #397 first if anything, since it decides whether any of the above is
|
||||
trustworthy.
|
||||
Reference in New Issue
Block a user