Both shapes carry no question mark and no interrogative, so the model saw
them with nothing deterministic in front and routed both to fact. The fact
gate caught the write and re-ran the turn as a query, so nothing broke —
what they cost was a full model round trip for a decision two patterns can
make offline.
NarrativeQueryGrammars, wired after the agenda rules so that "расскажи,
что у меня сегодня" stays an agenda question. Two exclusions, both learned
from the fixture: a capture verb in the rest of the utterance means he
asked for a note, and an entertainment noun means chat — "расскажи анекдот
про программистов" is ru-chat-003, and my first pattern took it.
The fixture had no case for either shape, which is why they went unnoticed.
Added as ru-query-020 and ru-query-021: classifier+onnx 53/77 → 55/79
(68.8% → 69.6%), both new cases answered at stage 0, false clarifies
unchanged at 0.
"что у меня сегодня" and "что у меня в календаре сегодня" both routed
IntentSystem on the deployed daemon, and replySystem has no agenda arm,
so both answered "пока не умею". The calendar source that can answer
them lives in the query chain and was never reached. The fixture has
said query since ru-query-019 was written; the daemon disagreed with the
fixture and the daemon was wrong.
AgendaQueryGrammars routes them at stage 0, after the clock rules so
"какой сегодня день" keeps reaching replySystem. Intent only — which
source claims the turn stays the query chain's decision.
This is what made the follow-up continuation look like it only worked
for "what day is it". It did: the query half inherited an intent whose
handler could not answer, so both halves came back "пока не умею".
Measured on the 77-case RU fixture: full accuracy 70.1% → 72.7%,
intent-only 75.3% → 77.9%, calendar 0/2 → 2/2, clarify counts unchanged.
The eval harness wires the new grammars too, or the fixture would stop
being a measurement of the daemon.
Go's \b is ASCII-only and never fires after a Cyrillic letter, which the
first version of the pattern learned the hard way.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
The old model was a symmetric paraphrase model, so it scored "do these
look alike" instead of "does this note answer this question". Also fixes
the file mismatch: the Makefile, the deploy config and both evals now all
name the same quantized file, and the quantized one is what gets measured.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
Completes #319's comparison. Three configurations, because "the LLM router"
was ambiguous: the model alone, the cascade #320 would actually ship (stage-0
grammar → model → classifier floor), and a thinking-off diagnostic.
intent-only full RU missed-clarify p50
classifier+onnx 36.8% 36.8% 25/61 5/6 31ms
llm-only (0.8B) 48.7% 23.7% 13/61 6/6 850ms
cascade+llm (0.8B) 50.0% 32.9% 18/61 6/6 825ms
On the question asked — does the resident model route better? — yes, 50.0%
vs 36.8% intent accuracy. REARCH.md's premise holds. It costs 27x the
latency (p50 825ms vs 31ms, max 3.1s), on the same llama-server the phraser
needs, so it is a trade rather than a free win.
Three things the numbers surface that the headline hides:
query→fact x15 is the dominant failure, four times the classifier's x4 on
the same axis. routeSystem's decision order puts "сообщает или обновляет
состояние" (rule 3) above "хочет получить информацию" (rule 4), so any
utterance naming a fact key matches the earlier rule and a question about
past state reads as an assertion of it. A prompt fix, not a model limit.
The LLM router cannot clarify: llmrouter.go hardcodes Confidence 1.0, so
stage 3's gate can never fire on its decisions — 6/6 missed. With #359's
finding that the classifier's gate is miscalibrated under ONNX, neither path
currently refuses. Flipping #320 as-is removes the refusal lane.
The gap between 50.0% intent and 32.9% full accuracy is entirely slots: the
LLM path fills neither Fn nor Time (it returns Slots.Text for acts, and
Extract never runs on an LLM decision).
Also settles a hypothesis rather than leaving it in the air: thinking mode is
a non-issue under a grammar (identical score), and the grammar's unbounded
("," ws action)* repetition that ran away in an isolated smoke test does not
reproduce under the real prompt — 2 errors in 76, not 76. internal/llm
deliberately does not grow a chat_template_kwargs field.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X5JApcrCRVGmqrxnhynSik
The onnxruntime .so was already vendored at deps/onnxruntime-linux-x64-1.26.0
— nothing to download. make eval-router now defaults MAVEN_ONNX_LIB there, so
both baselines run by default and only a fresh clone without deps/ falls back
to the hash ratchet alone.
Prod-representative result, deployed 0.55 gate: 28/76 (36.8%), RU 25/61,
EN 3/15, hard 0/11 → 4/11, p50 31ms / p95 71ms. Versus the hash floor's
13/76 at p50 9µs.
The finding is not the accuracy, it's the refusal lane: missed clarifies went
0 → 5 of 6. Better embeddings raise cosine everywhere, so the 0.55 threshold
that used to hold ambiguous utterances back stops holding — "сделай это"
routes to act at 0.847, "бэкап" to chat at 0.755. The gate was implicitly
tuned to the hash floor's low similarities. That is an argument about the
threshold, not about the embedder, and it lands before #320 rather than after.
Also fixes a fixture-model mismatch: ReminderGrammar deliberately skips the
extractor at stage 0 and the daemon's applyAction parses the time downstream
(stage0.go says so). Charging the router for that slot made 4 exact-match wins
read as misses; they are now counted as SlotsDeferred instead. Hash baseline
moves 13/76, ratchet to 0.15.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X5JApcrCRVGmqrxnhynSik
#319 asks for a measurement before #320 flips the route decider from the
classifier cascade to the resident model. There was nothing to measure
against: the only routing tests assert single utterances, and the
classifier's seed corpus is its own training set — scoring it there
measures memorisation of frozen centroids, which is the illusion that hid
the weak RU query handling in the first place.
internal/router/eval is a separate package so both paths can be scored
from outside router (including cmd/mavend, where the real llama-server
client lives). The fixture is embedded; the scorer takes a Router
interface, so *router.Router and a bare LLM stage both go through the same
76 cases.
The fixture is a CONTRACT, not a snapshot: cases the cascade fails today
stay in the file and fail loudly. TestFixtureIsHeldOut enforces that no
utterance appears verbatim in models/seeds/*.txt.
Baseline, hash embedder at the deployed 0.55 gate: 9/76 (11.8%), 63 false
clarifies, 0 missed clarifies, p50 9µs. Almost everything falls to the
confidence gate — the documented floor behaviour, not a new bug. The
number worth comparing is TestONNXBaseline's (skipped without
MAVEN_ONNX_LIB); the assertions here are a regression ratchet plus a tight
bound on the dangerous direction: ambiguous utterances must not start
being routed confidently.
Seeding is order-fixed on purpose — a few phrases appear under two intents
and map iteration handed them to a different centroid each run, which made
the score jitter between 9 and 10.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X5JApcrCRVGmqrxnhynSik