Files
Maven/docs/evals/2026-08-08-destination-model-arm.md
claude 00f9239ef9 Record the model arm, and the stage 0 trade it exposed (V-660)
Numbers and the argument in docs/evals/2026-08-08-destination-model-arm.md,
pointer and the short version in CLAUDE.md. The finding worth carrying is
not the 72.7%: it is that stage 0's silence on the possessive agenda rules
used to be free and now costs four destination points, because there is
finally something downstream that would have named the calendar.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-08 18:23:22 +04:00

2.9 KiB

The destination, with a model that can name one

Measured 2026-08-08 against gemma-4-12b on the workstation, the same 96-case fixture V-659 built. Covers V-660.

no_proxy='*' MAVEN_LLM_URL=http://192.168.1.105:8080 \
  make t PKG=./internal/router/eval/ RUN=TestLLMRouterBaseline V=1

The gap was structural

V-659 measured the destination at 12/33 on the classifier cascade, with recall at 0/15. Nothing in routeSystem named a Source and routeGrammar could not emit one, so the resident model had no string to write. That is the shape V-517 measured for Praxis reach at 0/12: not a weak model, an absent contract.

routeGrammar now carries a source rule closed over router.Sources plus the empty floor. The prompt lists the twelve destinations in Russian and says that "" is a normal answer to give often.

Result

run intent destination
classifier + ONNX (V-659) 73/96 (76.0%) 12/33 (36.4%)
gemma-4-12b alone 79/96 intent-only (82.3%) 26/33 (78.8%)
cascade + gemma-4-12b + hash fallback 81/96 (84.4%) 24/33 (72.7%)

Recall is the move: 0/15 to 14/15. Intent did not shift, which was the constraint. The prompt is shared, so a destination rule that costs routing points is not a win.

The eight llm-only errors are the eight want_clarify cases. The model returned unknown on every one, which is correct, and the llm-only harness surfaces a decline as an error by design.

Stage 0 now costs four destination points

The four cases the cascade loses and the model alone wins are all calendar. The possessive agenda rules claim them at stage 0 and deliberately name nothing. "что у меня в списке покупок" matches the same rule. Naming the calendar there would take the list source off the turn (V-655).

So a rule written to be careful about the list now blocks a model that would have named the calendar correctly. Before V-660 that caution was free, because nothing downstream of stage 0 could name anything either.

Three ways out, and each costs something. Split the possessive rule so the calendar-shaped half names its destination. Let a later stage overwrite an empty destination a grammar left behind, which reverses "a matched value always wins". Or leave it, on the argument that four points is cheap next to a wrong destination on a shopping list. This wants the owner's call rather than a quiet edit.

What this does not measure

The resident Qwen3-1.7B, which is what homesrv runs. It binds --port 0 inside the container and no host process can reach it. Scoring it needs a second llama-server on a fixed port. The workstation is never assumed up, so the homesrv number is the one that decides whether this ships on by default.

The fixture is 33 labelled destinations over twelve values. Recall carries 15 of them and five destinations carry none at all. A per-destination number below world, recall, calendar and the floor is not supported by this fixture.