Files
Maven/docs/evals/2026-07-31-talk.md
T
claude 93987f2dfc docs: tier the tree by lifetime, so staleness shows in the path (V-446)
Seventeen markdown files at the repo root, twelve of them dated one-shot
reports sitting next to CLAUDE.md. That is why stale docs read as
current: nothing in the path said which was which.

Root now keeps CLAUDE.md and AGENTS.md. Living docs move under docs/
and carry a Last verified line. Dated measurements move to docs/evals/
ISO-prefixed, and are never edited after the day, so a newer number is
a new file. The senior review moves to docs/archive/.

Every reference was rewritten across markdown, Go comments, the Makefile
and the recall fixture. The touched Go packages still build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 03:28:49 +04:00

7.9 KiB
Raw Blame History

Conversational phrasing eval — 31-07-2026

Every score measured tonight, on the three paths the nudge eval never touched: chat, query-with-notes, and general knowledge.

Short version: the plumbing got fixed and the score barely moved. Grammar and Russian prompts together took the composite from ~9 to ~14 of 27. Everything still failing is the model not knowing things or not holding a constraint, and prompting is out of levers. Settles the measurement half of Vikunja #395 / #398 / #400.

How to reproduce

# llama-server: -c 4096 -ngl 99 -t 6, model /mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf
MAVEN_LLM_URL=http://127.0.0.1:18099 no_proxy=127.0.0.1,localhost \
  deps/go/go/bin/go test -count=1 -timeout 40m \
  -run TestLLMTalkBaseline ./internal/phraser/eval/ -v

Three runs per configuration, always. The fixture is 27 cases, so one reply changing moves the composite by 3.7 points — a single run cannot tell a real change from sampling noise. This was learned the expensive way: an earlier claim that "one nudge case fails every run" turned out to be three different cases across three runs.

Run the box otherwise idle. See the contamination note at the bottom.

Composite, per configuration

config overall /27 chat /9 query /9 knowledge /9 canned fallbacks
baseline, no grammar 7, 12, 7 1, 1, 0 2, 4, 2 4, 7, 5 0, 0, 0
+ GBNF grammar (#398) 14, 15, 8 1, 3, 0 5, 6, 3 8, 6, 5 0, 0, 0
+ Russian prompts (#400) 11, 17, 15 1, 5, 3 5, 6, 8 5, 6, 4 0, 0, 0
+ truncation fix, 1000ch/768tok 12, 13, 10 2, 2, 1 7, 7, 5 3, 4, 4 3, 3, 6
+ rebalanced, 600ch/1024tok void — contaminated

"Canned fallbacks" counts replies that came back as the hardcoded "не знаю." or "поговорили.". It is not a check, it is a health signal: those strings mean the phraser gave up, and the eval scores them as ordinary bad replies.

Per-check

check no grammar + grammar + RU prompts + truncation fix
nonempty 27, 27, 27 27, 27, 27 27, 27, 27 27, 27, 27
ellipsis 20, 19, 23 27, 27, 27 27, 27, 27 27, 27, 27
lang 13, 16, 15 23, 26, 26 25, 26, 25 26, 27, 27
feminine 25, 24, 26 25, 25, 27
address 21, 22, 22 22, 21, 22
ontopic 17, 24, 18 17, 19, 14

nonempty reading 27/27 everywhere is not good news — it was a broken check. It tested for a non-blank string, so replies of literally { and "15-16" passed it. Fixed on overnight/fix-truncation; it needs a letter now.

What each change actually bought

GBNF grammar (#398) — the biggest single win. Qwen3.5-0.8B writes Thinking Process: as plain text with no tags, stripThink only handles </think>, so the JSON never closed and the plain-text fallback shipped the literal reasoning. ellipsis went 20→27 and lang 13→26. The router had been using a grammar for ages; the phraser asking nicely in the prompt was the oversight.

Russian prompts (#400) — modest, plus a large latency win. Chat 1.3→3.0 average, query 4.7→6.3, knowledge 6.3→5.0. All inside the run-to-run spread, so "probably better on the paths it targeted, not provable in three runs". p50 latency dropped from ~11.5s to ~2.3s and that part is consistent across all three runs — shorter prompts, and she stopped emitting English reasoning first.

Truncation fix — necessary, and did not help the score. Two real bugs (replies of {, and a nonempty check that passed them), both fixed, and the composite went nowhere. A complete rambling wrong answer fails the same checks a truncated one did. Worth doing anyway: the daemon was shipping { to a text-to-speech voice.

The truncation bug, since the cause was counter-intuitive

The grammar's string ::= ... {0,400} rule was the cause, not the token cap. Measured against Qwen3.5-0.8B at three caps — 256, 768 and 2048 — the reply came back exactly 400 characters every time, cut mid-word ("Нужно записать и,").

Then I raised the bound to 1000 while the cap was 768 tokens and made it worse: Russian runs ~1.5 characters per token here, so generation died on the token cap instead, mid-object, and the new guard correctly refused it and shipped "не знаю." — 3, 3 and 6 fallbacks per run, from zero. The two limits have to agree. 600 characters needs ~400 tokens; the cap is 1024.

Where the remaining failures live

address is stuck at 21-22 of 27 and ontopic at 14-19. Both resist prompting.

The prompt now explicitly forbids exactly what she does. It says never "вы", use the singular — and she writes вашей, подождите, делаете, хотите, напишите. Telling a 0.8B "never do X" does not work. Same for feminine: я готов, я понял, я нашел, я заметил, я сказал.

Some of ontopic is the fixture, not the model. chat-how-are-you got "Привет! Я здесь, чтобы поговорить. Как дела сегодня?" — a fine reply that fails because want_any is [норм, хорош, порядк, тут, работ]. It fails in every run, so it inflates the count. The ontopic column currently measures the fixture as much as the model. Not fixed yet, deliberately: changing it would break comparability with the runs above.

Two replies worth reading, because they are not fixable by prompting:

  • Thunder and lightning: "Скорость молнии — 8-10 тысяч километров в секунду, но звук — 300 метров в секунду, что делает молнию громче." Confidently wrong, and it concludes lightning is louder rather than sound being slower.
  • "расскажи обо мне": "Ты — прекрасное существо, с душой и вниманием… Спасибо за твою улыбку… О тебе — заповедь любви." Sycophantic filler, zero information, and precisely the "not a relationship" non-goal.
  • Boiling an egg: "15-16" one run, "1" another. No unit, wrong number.

The first argues for reading instead of recalling (#403 — Kiwix retrieval scores 8/8 on the same questions given English keywords). The second and third argue for templates on the paths where correctness matters (#392).

Contamination note — how the last row got voided

I started the query-rewrite agent against the same llama-server the sweep was using, and assumed contention would only affect latency. It did not. The knowledge path collapsed to 0 of 9 with eight canned "не знаю." replies, p95 tripled to 23.7s, and the report still said "0 errors".

That is Vikunja #397, and it is worse than filed: a merely busy server produces a clean-looking report with a third of the fixture silently answering "не знаю.". PhraseChat and PhraseQuery swallow every failure and return a hardcoded string, so infrastructure trouble is indistinguishable from bad phrasing in the score. The talk test guards the start and end of a run with a model check, which catches a dead server but not a loaded one.

Until #397 is fixed, treat any run made on a busy box as void.

Next

  • Re-run 600ch/1024tok clean, to fill the void row.
  • Score Qwen3.5-2B-UD-Q4_K_XL (already at /mnt/hdd1/llms/qwen3.5/, never measured) on this fixture and the router fixture. Not the 4B — too big for this box, owner's call.
  • Newer sub-500M candidates (LFM2.5 200M/300M) are worth a run for routing. Note docs/evals/2026-07-31-model-bakeoff.md found LFM2.5-1.2B worse than Qwen3.5-0.8B at Russian routing and 2.4× slower — but those are a different, older generation, so that result does not predict the small ones.
  • Fix chat-how-are-you's want_any, and re-baseline once, so ontopic measures the model.
  • #397 first if anything, since it decides whether any of the above is trustworthy.