Files
Maven/docs/evals/2026-07-31-phrasing.md
claude 93987f2dfc docs: tier the tree by lifetime, so staleness shows in the path (V-446)
Seventeen markdown files at the repo root, twelve of them dated one-shot
reports sitting next to CLAUDE.md. That is why stale docs read as
current: nothing in the path said which was which.

Root now keeps CLAUDE.md and AGENTS.md. Living docs move under docs/
and carry a Last verified line. Dated measurements move to docs/evals/
ISO-prefixed, and are never edited after the day, so a newer number is
a new file. The senior review moves to docs/archive/.

Every reference was rewritten across markdown, Go comments, the Makefile
and the recall fixture. The touched Go packages still build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 03:28:49 +04:00

8.7 KiB
Raw Permalink Blame History

Phrasing evaluation — 31-07-2026

How Maven words a nudge, measured instead of argued. Counterpart to docs/evals/2026-07-31-routing.md.

  • Fixture + scorer: internal/phraser/eval/ (nudges_v1.json, 15 cases; eval.go, checks.go)
  • Reproduce: MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing
  • Model: Qwen3.5-0.8B Q4_K_M, the resident model. Not swapped.
  • Commit: a40bc55 (prompt fix)

Every check is a string or length test a human can read and disagree with. No model grades another model here.

Result

before after
cases passing every check 0/15 13/15
mood in enum 6/15 15/15
Russian 2/15 14/15
length (≤120 chars, ≤16 words) 13/15 15/15
feminine self-reference 15/15 15/15
no cringe 13/15 15/15
on topic 6/15 13/15
p50 latency 11.4s 11.4s

Latency did not move and is not good. 11s to word one nudge on this box.

The bug reproduced

Yes, exactly as reported. 7 of 15 messages were the literal string "...", and one was "full voice message". Both are text copied straight out of the prompt.

The system prompt said:

Respond ONLY with valid JSON: {"response": "full voice message", "mood": "neutral"}

and the user prompt said:

Respond as JSON: {"response": "...", "mood": "..."}

A 0.8B does not read "..." as "put your answer here". It reads it as the answer. The prompt was a worked example whose worked part was blank, so the model filled the slot by copying. This is the whole of finding 1.

What else was wrong

Four separate faults, all prompt-side:

  1. Placeholder echo (7 cases) — above.
  2. Wrong language (13/15 failed the language check). The prompt was entirely English and said "in the user's language (Russian or English)". The model picked English. It is never English: the nudge is spoken by a Russian piper voice.
  3. Rule names are English identifiers. netdata_critical, service_down, break went into the prompt raw. The model cannot nudge about a topic it has not been told in words, so 9/15 were off topic. The daemon knows what its own rules mean; now it says so.
  4. Mood invented ("warm", twice). The enum was listed in a parenthesis at the end of an English sentence. Now it is its own line: "ровно одно из: neutral, happy, thinking, tired, confused."

Plus two non-prompt faults the run exposed:

  • The no-parse fallback was English. When the model returned nothing usable, the body became fmt.Sprintf("%s — %s", rule, sev)"water — care" — and that string went to a Russian TTS. Now it falls back to plain Russian.
  • Durations were English. humanDur returns "3 hours"; it was landing verbatim inside Russian sentences. Nudges now use a Russian formatter.

Three iterations, and what each taught

score change
baseline 0/15
iter 1 2/15 Russian prompt, filled-in examples, Russian durations
iter 2 11/15 required keyword per rule, one example instead of five, Russian fallback
iter 3 13/15 examples moved to topics that are not rules

The interesting step is 1 → 2. Fixing the placeholder did not fix the disease, it moved it: the model stopped copying "..." and started copying my first example instead. Five nudges in a row came back as "Ты не пил воду три часа. Налей стакан." regardless of the rule.

A small model copies the nearest concrete text in its prompt. That is one failure mode with two symptoms. The fix that stuck was making the examples about laundry and a laptop battery — topics no rule ever produces, so copying them is visible in the score rather than invisibly passing the water cases.

Do not oversell 13/15

Seven of the thirteen passes are the deterministic fallback, not the model: "Напоминаю: таблетки.", "Сервис не отвечает.", "Критический алярм: проверь диск.", "Ты давно не пил воду.". Those are strings this commit added to Go. The model returned nothing parseable and the fallback scored.

So the honest reading is roughly 6/15 from the model, 7/15 from a fallback, 2/15 failing. The prompt fix is real — "..." is nearly gone and the language and mood checks are clean — but a large part of the jump is that failure now degrades into Russian instead of into "water — care". That is a genuine improvement for the operator and a weak one for the model.

The two remaining failures: one "..." recurrence (routine-stretch) and one meal nudge that never says food.

Tried and reverted: an example-led nudge prompt (#393)

The idea was that a 0.8B copies examples better than it follows rules, so the nudge prompt was rewritten to lead with five on-topic examples (water, break, pills, morning, service) and the prose rules were compressed to pay for the tokens: 1190 chars down to 986.

It measured worse, three runs each side, same llama-server, same fixture:

run before after
1 12/15 (address 14) 11/15 (address 13)
2 13/15 (address 15) 12/15 (address 15)
3 14/15 (address 15) 11/15 (address 12)

feminine and hisgender were 15/15 on all six runs, so they measure nothing here. The regression is all in address: 44/45 before, 40/45 after. Formal "вы"/"ваше" and plural imperatives came back, and so did "...".

Two likely causes, both about the same thing — examples do not carry a prohibition. The old prompt spent a whole sentence on «говоришь на "ты", в единственном числе»; the new one demoted that to one item in a long "никогда" list, and the model stopped obeying it. And making the examples on-topic let their wording leak: a break case came back as «Вы давно не пили воду. Выпей стакан.» — the water example, verbatim, in the wrong slot. That is exactly the failure the laundry/laptop examples were chosen to avoid.

Change reverted. What survives is the measurement: a rule the model must obey needs its own sentence, and examples must stay off-topic. Also note the before side alone spans 1214 of 15 — this fixture cannot resolve anything smaller than about three cases.

Broken, found, not fixed

  1. checkFeminine only catches half the constraint. Fixed (#381). It scanned for masculine self-reference only, so three messages that addressed the owner in the feminine ("ты давно не отдыхала") scored clean. There is now a second check, hisgender: a feminine past-tense verb (-ла/-лась) in a sentence addressed to him ("ты", "тебе", "твой") fails, unless the verb is hers ("я заметила", "напомнила тебе"). It is a suffix rule, not a parser — see the comment in checks.go for what it misses. A fresh 15-case run after adding it scored 12/15 with hisgender 15/15; the model did not repeat the feminine address in that sample, and the check is pinned by unit tests on the recorded bad strings instead.
  2. Grammar is not checked at all, and it is bad. "Он не ел 11 дней" (it was 11 hours), "Сонуждились 7 дней" (not a word), "Они забыли воду" (wrong person entirely). Every one of these passes all six checks. The fixture measures properties, not fluency, and at 0.8B fluency is the binding constraint.
  3. Unit confusion. The model turns hours into days about a third of the time. The prompt now says "11 ч"; it reads it as days.
  4. 11s p50. Unchanged and untouched here. A nudge the model takes eleven seconds to word has missed its moment. Worth its own task.
  5. The keyword hint is close to teaching to the test. ruleKeywords names the word the on-topic check looks for. It is defensible — the daemon genuinely knows its rule topics and the model genuinely cannot infer them from netdata_critical — but the on-topic number is softer than the others because of it.

Next steps

  1. Add a second-person gender check — done, hisgender in checks.go (#381).
  2. Decide whether the fallback should count as a pass. Right now Score cannot tell a model answer from a fallback. Either mark fallback bodies in PhrasedNudge or count them in their own column. Without that, any future prompt change can score well by failing more.
  3. Attack the 11s. Nudge phrasing is short and non-interactive; thinking off is the first thing to try, as it was for routing (#376).
  4. Re-measure when #122 lands. The CPT'd Qwen3-1.7B is the target. 13/15 with seven fallbacks is the floor it has to beat, and the fluency problems above are the ones a bigger, Russian-trained checkpoint should actually fix.