Commit Graph

2 Commits

Author SHA1 Message Date
kami d34fdf40aa Score the routing fixture with the ONNX embedder (Vikunja #319)
The onnxruntime .so was already vendored at deps/onnxruntime-linux-x64-1.26.0
— nothing to download. make eval-router now defaults MAVEN_ONNX_LIB there, so
both baselines run by default and only a fresh clone without deps/ falls back
to the hash ratchet alone.

Prod-representative result, deployed 0.55 gate: 28/76 (36.8%), RU 25/61,
EN 3/15, hard 0/11 → 4/11, p50 31ms / p95 71ms. Versus the hash floor's
13/76 at p50 9µs.

The finding is not the accuracy, it's the refusal lane: missed clarifies went
0 → 5 of 6. Better embeddings raise cosine everywhere, so the 0.55 threshold
that used to hold ambiguous utterances back stops holding — "сделай это"
routes to act at 0.847, "бэкап" to chat at 0.755. The gate was implicitly
tuned to the hash floor's low similarities. That is an argument about the
threshold, not about the embedder, and it lands before #320 rather than after.

Also fixes a fixture-model mismatch: ReminderGrammar deliberately skips the
extractor at stage 0 and the daemon's applyAction parses the time downstream
(stage0.go says so). Charging the router for that slot made 4 exact-match wins
read as misses; they are now counted as SlotsDeferred instead. Hash baseline
moves 13/76, ratchet to 0.15.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X5JApcrCRVGmqrxnhynSik
2026-07-31 00:34:29 +04:00
kami c7c44229a2 Add held-out RU routing fixture and scorer (Vikunja #319)
#319 asks for a measurement before #320 flips the route decider from the
classifier cascade to the resident model. There was nothing to measure
against: the only routing tests assert single utterances, and the
classifier's seed corpus is its own training set — scoring it there
measures memorisation of frozen centroids, which is the illusion that hid
the weak RU query handling in the first place.

internal/router/eval is a separate package so both paths can be scored
from outside router (including cmd/mavend, where the real llama-server
client lives). The fixture is embedded; the scorer takes a Router
interface, so *router.Router and a bare LLM stage both go through the same
76 cases.

The fixture is a CONTRACT, not a snapshot: cases the cascade fails today
stay in the file and fail loudly. TestFixtureIsHeldOut enforces that no
utterance appears verbatim in models/seeds/*.txt.

Baseline, hash embedder at the deployed 0.55 gate: 9/76 (11.8%), 63 false
clarifies, 0 missed clarifies, p50 9µs. Almost everything falls to the
confidence gate — the documented floor behaviour, not a new bug. The
number worth comparing is TestONNXBaseline's (skipped without
MAVEN_ONNX_LIB); the assertions here are a regression ratchet plus a tight
bound on the dangerous direction: ambiguous utterances must not start
being routed confidently.

Seeding is order-fixed on purpose — a few phrases appear under two intents
and map iteration handed them to a different centroid each run, which made
the score jitter between 9 and 10.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X5JApcrCRVGmqrxnhynSik
2026-07-31 00:28:44 +04:00