# gemma-4-E4B on the phrasing and talk fixtures Date: 2026-08-09. Box: workpc up, E4B loaded on 8080. `MAVEN_LLM_URL=http://192.168.1.105:8080 make eval-phrasing`. This was the one unmeasured risk of the 2026-08-09 model swap. Routing was measured the same day and E4B lost four destination cases to the 12B. Phrasing was not measured at all, and phrasing is the half the owner hears. ## Result | fixture | E4B | resident Qwen3-1.7B, 2026-08-05 | |---|---|---| | nudges | 15/15 (100%) | 15/15 (100%) | | talk, passes every check | **29/36 (80.6%)** | 25/36 (69.4%) | | lang | 36/36 | — | | feminine | 36/36 | 36/36 | | address | **36/36** | 33/36 | | ontopic | 29/36 | 28/36 | | p50 latency | **516ms** | 2.97s | | p95 latency | 921ms | — | | failed generations | 0 | 0 | E4B beats the homesrv floor by four cases and answers about six times faster. Persona is clean: `lang`, `feminine` and `address` are perfect, and `address` is where the resident model still loses three. The 2026-08-05 measurement of the resident model is the comparison, since both ran the same 36-case fixture. Every failure is `ontopic`. Nothing failed on persona, nothing failed to parse. ## The score is at the ceiling, not below it The 2026-08-05 temperature sweep found two cases that fail at every temperature in every run: `reply-note-router` and `reply-fact-weight`. It named a defect in the reply phrasing path rather than sampling noise. It put the fixture's ceiling at 30/36 before persona is scored. Both cases are in E4B's failure list. So 29/36 is one case off a ceiling nothing about the model can move. The swap is safe on phrasing. Read this next to the routing result, not instead of it. There E4B costs four destination cases and buys 50ms. Here it costs nothing. ## Two findings no check caught **She says she wrote something down when she did not.** Asked what to do this evening, E4B writes "Я записала несколько идей!". Asked for a joke, it writes "Я записала одну забавную ситуацию!". Nothing was stored. No check scores it, because `ontopic` reads the subject and `cringe` reads pet names. A claim to have saved something is a claim about state, and it is wrong. **Two `ontopic` failures look like check defects.** `know-dont-know` wants "не зна" or "не мог". It got "Я не умею знать личную информацию о твоих соседях", which declines correctly in words the check does not list. `know-hiccups` is the same shape. Neither is a model failure and both count against the score. ## Not measured here A 12B control on the same fixture, which would need the card reloaded and is the owner's call. The talk fixture through the daemon rather than through the phraser directly. The CPT'd Qwen3-1.7B, which does not exist yet and is the reason `address` is a check at all.