50ca8c8b5a
The phrasing fixture was 15 nudge cases, so every prompt change we measured only told us about nudges. But the shared context block sits in front of five prompts, and three of them — chat, note query, general knowledge — had no scorer at all. Those are the long free-form replies, where a persona break is most likely and where nothing could see one. 27 cases, nine per path. Nine rather than five because the nudge fixture already cannot resolve a change smaller than about three cases, and a per-path score off five would be worse. Reuses the persona checks instead of copying them. Length, mood and "no questions" are left out on purpose: these paths return no mood, and a follow-up question is a feature in chat, not a fault. The run refuses to score unless the model answers before and after it. PhraseChat and PhraseQuery swallow model errors and return a canned string, so without that guard a dead server produces a full report with zero errors and a bad score — which reads as bad phrasing rather than as nothing measured. Vikunja #397 is the real fix.