c7c44229a2
#319 asks for a measurement before #320 flips the route decider from the classifier cascade to the resident model. There was nothing to measure against: the only routing tests assert single utterances, and the classifier's seed corpus is its own training set — scoring it there measures memorisation of frozen centroids, which is the illusion that hid the weak RU query handling in the first place. internal/router/eval is a separate package so both paths can be scored from outside router (including cmd/mavend, where the real llama-server client lives). The fixture is embedded; the scorer takes a Router interface, so *router.Router and a bare LLM stage both go through the same 76 cases. The fixture is a CONTRACT, not a snapshot: cases the cascade fails today stay in the file and fail loudly. TestFixtureIsHeldOut enforces that no utterance appears verbatim in models/seeds/*.txt. Baseline, hash embedder at the deployed 0.55 gate: 9/76 (11.8%), 63 false clarifies, 0 missed clarifies, p50 9µs. Almost everything falls to the confidence gate — the documented floor behaviour, not a new bug. The number worth comparing is TestONNXBaseline's (skipped without MAVEN_ONNX_LIB); the assertions here are a regression ratchet plus a tight bound on the dangerous direction: ambiguous utterances must not start being routed confidently. Seeding is order-fixed on purpose — a few phrases appear under two intents and map iteration handed them to a different centroid each run, which made the score jitter between 9 and 10. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X5JApcrCRVGmqrxnhynSik