Record the model arm, and the stage 0 trade it exposed (V-660)
Numbers and the argument in docs/evals/2026-08-08-destination-model-arm.md, pointer and the short version in CLAUDE.md. The finding worth carrying is not the 72.7%: it is that stage 0's silence on the possessive agenda rules used to be free and now costs four destination points, because there is finally something downstream that would have named the calendar. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
This commit is contained in:
@@ -0,0 +1,66 @@
|
||||
# The destination, with a model that can name one
|
||||
|
||||
Measured 2026-08-08 against gemma-4-12b on the workstation, the same 96-case
|
||||
fixture V-659 built. Covers V-660.
|
||||
|
||||
```sh
|
||||
no_proxy='*' MAVEN_LLM_URL=http://192.168.1.105:8080 \
|
||||
make t PKG=./internal/router/eval/ RUN=TestLLMRouterBaseline V=1
|
||||
```
|
||||
|
||||
## The gap was structural
|
||||
|
||||
V-659 measured the destination at 12/33 on the classifier cascade, with recall
|
||||
at 0/15. Nothing in `routeSystem` named a `Source` and `routeGrammar` could not
|
||||
emit one, so the resident model had no string to write. That is the shape V-517
|
||||
measured for Praxis reach at 0/12: not a weak model, an absent contract.
|
||||
|
||||
`routeGrammar` now carries a `source` rule closed over `router.Sources` plus the
|
||||
empty floor. The prompt lists the twelve destinations in Russian and says that
|
||||
`""` is a normal answer to give often.
|
||||
|
||||
## Result
|
||||
|
||||
| run | intent | destination |
|
||||
|---|---|---|
|
||||
| classifier + ONNX (V-659) | 73/96 (76.0%) | 12/33 (36.4%) |
|
||||
| gemma-4-12b alone | 79/96 intent-only (82.3%) | 26/33 (78.8%) |
|
||||
| cascade + gemma-4-12b + hash fallback | 81/96 (84.4%) | 24/33 (72.7%) |
|
||||
|
||||
Recall is the move: 0/15 to 14/15. Intent did not shift, which was the
|
||||
constraint. The prompt is shared, so a destination rule that costs routing
|
||||
points is not a win.
|
||||
|
||||
The eight llm-only errors are the eight `want_clarify` cases. The model returned
|
||||
`unknown` on every one, which is correct, and the llm-only harness surfaces a
|
||||
decline as an error by design.
|
||||
|
||||
## Stage 0 now costs four destination points
|
||||
|
||||
The four cases the cascade loses and the model alone wins are all calendar. The
|
||||
possessive agenda rules claim them at stage 0 and deliberately name nothing.
|
||||
"что у меня в списке покупок" matches the same rule. Naming the calendar there
|
||||
would take the list source off the turn (V-655).
|
||||
|
||||
So a rule written to be careful about the list now blocks a model that would
|
||||
have named the calendar correctly. Before V-660 that caution was free, because
|
||||
nothing downstream of stage 0 could name anything either.
|
||||
|
||||
Three ways out, and each costs something. Split the possessive rule so the
|
||||
calendar-shaped half names its destination. Let a later stage overwrite an empty
|
||||
destination a grammar left behind, which reverses "a matched value always wins".
|
||||
Or leave it, on the argument that four points is cheap next to a wrong
|
||||
destination on a shopping list. This wants the owner's call rather than a quiet
|
||||
edit.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
The resident Qwen3-1.7B, which is what homesrv runs. It binds `--port 0` inside
|
||||
the container and no host process can reach it. Scoring it needs a second
|
||||
llama-server on a fixed port. The workstation is never assumed
|
||||
up, so the homesrv number is the one that decides whether this ships on by
|
||||
default.
|
||||
|
||||
The fixture is 33 labelled destinations over twelve values. Recall carries 15 of
|
||||
them and five destinations carry none at all. A per-destination number below
|
||||
world, recall, calendar and the floor is not supported by this fixture.
|
||||
Reference in New Issue
Block a user