Files
Maven/docs/evals/2026-08-09-e4b-phrasing.md
claude 229890abd7 Measure E4B on phrasing, the half nobody had scored (V-668)
The 2026-08-09 model swap was measured on routing the same day and E4B lost
four destination cases. Phrasing was not measured, and phrasing is the half the
owner hears.

E4B scores nudges 15/15 and the talk fixture 29/36 at p50 516ms, against the
resident model's 25/36 at p50 2.97s on the same 36 cases. lang, feminine and
address are all 36/36, where the resident model loses three on address. Every
failure is ontopic and none is a parse error.

29/36 is one case off the ceiling. The temperature sweep of 2026-08-05 found
two reply cases that fail at every temperature and named a defect in the reply
path, capping the fixture at 30/36. Both are in E4B's failure list. So the swap
costs nothing on phrasing.

One defect no check catches: in chat E4B writes "Я записала несколько идей!"
when nothing was stored. A claim to have saved something is a claim about state.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-09 12:07:13 +04:00

2.8 KiB

gemma-4-E4B on the phrasing and talk fixtures

Date: 2026-08-09. Box: workpc up, E4B loaded on 8080. MAVEN_LLM_URL=http://192.168.1.105:8080 make eval-phrasing.

This was the one unmeasured risk of the 2026-08-09 model swap. Routing was measured the same day and E4B lost four destination cases to the 12B. Phrasing was not measured at all, and phrasing is the half the owner hears.

Result

fixture E4B resident Qwen3-1.7B, 2026-08-05
nudges 15/15 (100%) 15/15 (100%)
talk, passes every check 29/36 (80.6%) 25/36 (69.4%)
lang 36/36
feminine 36/36 36/36
address 36/36 33/36
ontopic 29/36 28/36
p50 latency 516ms 2.97s
p95 latency 921ms
failed generations 0 0

E4B beats the homesrv floor by four cases and answers about six times faster. Persona is clean: lang, feminine and address are perfect, and address is where the resident model still loses three. The 2026-08-05 measurement of the resident model is the comparison, since both ran the same 36-case fixture.

Every failure is ontopic. Nothing failed on persona, nothing failed to parse.

The score is at the ceiling, not below it

The 2026-08-05 temperature sweep found two cases that fail at every temperature in every run: reply-note-router and reply-fact-weight. It named a defect in the reply phrasing path rather than sampling noise. It put the fixture's ceiling at 30/36 before persona is scored. Both cases are in E4B's failure list.

So 29/36 is one case off a ceiling nothing about the model can move. The swap is safe on phrasing. Read this next to the routing result, not instead of it. There E4B costs four destination cases and buys 50ms. Here it costs nothing.

Two findings no check caught

She says she wrote something down when she did not. Asked what to do this evening, E4B writes "Я записала несколько идей!". Asked for a joke, it writes "Я записала одну забавную ситуацию!". Nothing was stored. No check scores it, because ontopic reads the subject and cringe reads pet names. A claim to have saved something is a claim about state, and it is wrong.

Two ontopic failures look like check defects. know-dont-know wants "не зна" or "не мог". It got "Я не умею знать личную информацию о твоих соседях", which declines correctly in words the check does not list. know-hiccups is the same shape. Neither is a model failure and both count against the score.

Not measured here

A 12B control on the same fixture, which would need the card reloaded and is the owner's call. The talk fixture through the daemon rather than through the phraser directly. The CPT'd Qwen3-1.7B, which does not exist yet and is the reason address is a check at all.