229890abd7
The 2026-08-09 model swap was measured on routing the same day and E4B lost four destination cases. Phrasing was not measured, and phrasing is the half the owner hears. E4B scores nudges 15/15 and the talk fixture 29/36 at p50 516ms, against the resident model's 25/36 at p50 2.97s on the same 36 cases. lang, feminine and address are all 36/36, where the resident model loses three on address. Every failure is ontopic and none is a parse error. 29/36 is one case off the ceiling. The temperature sweep of 2026-08-05 found two reply cases that fail at every temperature and named a defect in the reply path, capping the fixture at 30/36. Both are in E4B's failure list. So the swap costs nothing on phrasing. One defect no check catches: in chat E4B writes "Я записала несколько идей!" when nothing was stored. A claim to have saved something is a claim about state. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
62 lines
2.8 KiB
Markdown
62 lines
2.8 KiB
Markdown
# gemma-4-E4B on the phrasing and talk fixtures
|
|
|
|
Date: 2026-08-09. Box: workpc up, E4B loaded on 8080.
|
|
`MAVEN_LLM_URL=http://192.168.1.105:8080 make eval-phrasing`.
|
|
|
|
This was the one unmeasured risk of the 2026-08-09 model swap. Routing was
|
|
measured the same day and E4B lost four destination cases to the 12B. Phrasing
|
|
was not measured at all, and phrasing is the half the owner hears.
|
|
|
|
## Result
|
|
|
|
| fixture | E4B | resident Qwen3-1.7B, 2026-08-05 |
|
|
|---|---|---|
|
|
| nudges | 15/15 (100%) | 15/15 (100%) |
|
|
| talk, passes every check | **29/36 (80.6%)** | 25/36 (69.4%) |
|
|
| lang | 36/36 | — |
|
|
| feminine | 36/36 | 36/36 |
|
|
| address | **36/36** | 33/36 |
|
|
| ontopic | 29/36 | 28/36 |
|
|
| p50 latency | **516ms** | 2.97s |
|
|
| p95 latency | 921ms | — |
|
|
| failed generations | 0 | 0 |
|
|
|
|
E4B beats the homesrv floor by four cases and answers about six times faster.
|
|
Persona is clean: `lang`, `feminine` and `address` are perfect, and `address`
|
|
is where the resident model still loses three. The 2026-08-05 measurement of the
|
|
resident model is the comparison, since both ran the same 36-case fixture.
|
|
|
|
Every failure is `ontopic`. Nothing failed on persona, nothing failed to parse.
|
|
|
|
## The score is at the ceiling, not below it
|
|
|
|
The 2026-08-05 temperature sweep found two cases that fail at every temperature
|
|
in every run: `reply-note-router` and `reply-fact-weight`. It named a defect in
|
|
the reply phrasing path rather than sampling noise. It put the fixture's ceiling
|
|
at 30/36 before persona is scored. Both cases are in E4B's failure list.
|
|
|
|
So 29/36 is one case off a ceiling nothing about the model can move. The swap is
|
|
safe on phrasing. Read this next to the routing result, not instead of it. There
|
|
E4B costs four destination cases and buys 50ms. Here it costs nothing.
|
|
|
|
## Two findings no check caught
|
|
|
|
**She says she wrote something down when she did not.** Asked what to do this
|
|
evening, E4B writes "Я записала несколько идей!". Asked for a joke, it writes
|
|
"Я записала одну забавную ситуацию!". Nothing was stored. No check scores it,
|
|
because `ontopic` reads the subject and `cringe` reads pet names. A claim to
|
|
have saved something is a claim about state, and it is wrong.
|
|
|
|
**Two `ontopic` failures look like check defects.** `know-dont-know` wants
|
|
"не зна" or "не мог". It got "Я не умею знать личную информацию о твоих
|
|
соседях", which declines correctly in words the check does not list.
|
|
`know-hiccups` is the same shape. Neither is a model failure and both count
|
|
against the score.
|
|
|
|
## Not measured here
|
|
|
|
A 12B control on the same fixture, which would need the card reloaded and is the
|
|
owner's call. The talk fixture through the daemon rather than through the
|
|
phraser directly. The CPT'd Qwen3-1.7B, which does not exist yet and is the
|
|
reason `address` is a check at all.
|