The 2026-08-09 model swap was measured on routing the same day and E4B lost four destination cases. Phrasing was not measured, and phrasing is the half the owner hears. E4B scores nudges 15/15 and the talk fixture 29/36 at p50 516ms, against the resident model's 25/36 at p50 2.97s on the same 36 cases. lang, feminine and address are all 36/36, where the resident model loses three on address. Every failure is ontopic and none is a parse error. 29/36 is one case off the ceiling. The temperature sweep of 2026-08-05 found two reply cases that fail at every temperature and named a defect in the reply path, capping the fixture at 30/36. Both are in E4B's failure list. So the swap costs nothing on phrasing. One defect no check catches: in chat E4B writes "Я записала несколько идей!" when nothing was stored. A claim to have saved something is a claim about state. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2.8 KiB
gemma-4-E4B on the phrasing and talk fixtures
Date: 2026-08-09. Box: workpc up, E4B loaded on 8080.
MAVEN_LLM_URL=http://192.168.1.105:8080 make eval-phrasing.
This was the one unmeasured risk of the 2026-08-09 model swap. Routing was measured the same day and E4B lost four destination cases to the 12B. Phrasing was not measured at all, and phrasing is the half the owner hears.
Result
| fixture | E4B | resident Qwen3-1.7B, 2026-08-05 |
|---|---|---|
| nudges | 15/15 (100%) | 15/15 (100%) |
| talk, passes every check | 29/36 (80.6%) | 25/36 (69.4%) |
| lang | 36/36 | — |
| feminine | 36/36 | 36/36 |
| address | 36/36 | 33/36 |
| ontopic | 29/36 | 28/36 |
| p50 latency | 516ms | 2.97s |
| p95 latency | 921ms | — |
| failed generations | 0 | 0 |
E4B beats the homesrv floor by four cases and answers about six times faster.
Persona is clean: lang, feminine and address are perfect, and address
is where the resident model still loses three. The 2026-08-05 measurement of the
resident model is the comparison, since both ran the same 36-case fixture.
Every failure is ontopic. Nothing failed on persona, nothing failed to parse.
The score is at the ceiling, not below it
The 2026-08-05 temperature sweep found two cases that fail at every temperature
in every run: reply-note-router and reply-fact-weight. It named a defect in
the reply phrasing path rather than sampling noise. It put the fixture's ceiling
at 30/36 before persona is scored. Both cases are in E4B's failure list.
So 29/36 is one case off a ceiling nothing about the model can move. The swap is safe on phrasing. Read this next to the routing result, not instead of it. There E4B costs four destination cases and buys 50ms. Here it costs nothing.
Two findings no check caught
She says she wrote something down when she did not. Asked what to do this
evening, E4B writes "Я записала несколько идей!". Asked for a joke, it writes
"Я записала одну забавную ситуацию!". Nothing was stored. No check scores it,
because ontopic reads the subject and cringe reads pet names. A claim to
have saved something is a claim about state, and it is wrong.
Two ontopic failures look like check defects. know-dont-know wants
"не зна" or "не мог". It got "Я не умею знать личную информацию о твоих
соседях", which declines correctly in words the check does not list.
know-hiccups is the same shape. Neither is a model failure and both count
against the score.
Not measured here
A 12B control on the same fixture, which would need the card reloaded and is the
owner's call. The talk fixture through the daemon rather than through the
phraser directly. The CPT'd Qwen3-1.7B, which does not exist yet and is the
reason address is a check at all.