Correct the classifier baseline after tokenizer repair (V-704)
This commit is contained in:
@@ -140,10 +140,11 @@ it is deliberate.** It guesses, so naming `SourceWorld` drops it. **Only a stage
|
|||||||
0 grammar may drop it** (owner's call, V-666). `queryWalk` reads
|
0 grammar may drop it** (owner's call, V-666). `queryWalk` reads
|
||||||
`Decision.SourceAnchored` for the source marked `boundary: true` and no other.
|
`Decision.SourceAnchored` for the source marked `boundary: true` and no other.
|
||||||
|
|
||||||
Judge a routing change against the classifier (76.0% intent, 36.4% destination)
|
Judge a routing change against the classifier and the resident model, since
|
||||||
and the resident model (80.2% intent), since those always answer. The fixture
|
those always answer. **Their scores live in `docs/routing.md`, never here.** A
|
||||||
has grown from 77 cases to 96, so a number compares only to another number on
|
pair copied into this file goes stale silently. The fixture has changed size
|
||||||
the same fixture.
|
more than once, so a number compares only to another number on the same
|
||||||
|
fixture.
|
||||||
|
|
||||||
## Language: model output and Russian
|
## Language: model output and Russian
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,71 @@
|
|||||||
|
# The classifier baseline after the tokenizer fix
|
||||||
|
|
||||||
|
Measured 2026-08-11 on master at `2cf8b7e`, the classifier cascade with the ONNX
|
||||||
|
multilingual embedder, the configuration homesrv runs. `make t
|
||||||
|
PKG=./internal/router/eval/ RUN=TestONNXBaseline V=1`. Covers V-704.
|
||||||
|
|
||||||
|
## Result
|
||||||
|
|
||||||
|
Intent is **72/96 (75.0%)**. Destination is **11/33 (33.3%)**.
|
||||||
|
|
||||||
|
`docs/evals/2026-08-08-destination-fixture.md` recorded 73/96 (76.0%) and 12/33
|
||||||
|
(36.4%) three days earlier, and `CLAUDE.md` and `docs/routing.md` both carried
|
||||||
|
that pair until today. Same fixture, same 33 labelled cases, so the two pairs
|
||||||
|
are comparable and one case moved.
|
||||||
|
|
||||||
|
| | 2026-08-08 | today |
|
||||||
|
|---|---|---|
|
||||||
|
| intent | 73/96 (76.0%) | 72/96 (75.0%) |
|
||||||
|
| destination | 12/33 (36.4%) | 11/33 (33.3%) |
|
||||||
|
| world | 5/5 | 5/5 |
|
||||||
|
| the `SourceUnknown` floor | 5/7 | 4/7 |
|
||||||
|
| calendar | 2/6 | 2/6 |
|
||||||
|
| recall | 0/15 | 0/15 |
|
||||||
|
|
||||||
|
## One case moved, and it cost both numbers
|
||||||
|
|
||||||
|
`en-query-002`, "how long since the last backup finished". It routed `query` on
|
||||||
|
08-08 and routes `fact` today, at 0.870.
|
||||||
|
|
||||||
|
The case carries `want_source` present and empty, the floor label. `Score`
|
||||||
|
credits a destination only after the route lands the intent. So the case paid
|
||||||
|
once on the intent and once on the destination. Two numbers moved by one case,
|
||||||
|
which is why the drift read as two regressions.
|
||||||
|
|
||||||
|
## The commit
|
||||||
|
|
||||||
|
`feabf9f`, "The tokenizer read every long word backwards" (V-664), landed
|
||||||
|
2026-08-08 after the destination measurement.
|
||||||
|
|
||||||
|
Bisected by running `TestONNXBaseline` at each commit:
|
||||||
|
|
||||||
|
| commit | intent | destination |
|
||||||
|
|---|---|---|
|
||||||
|
| `b6eaa70` V-659, the 08-08 measurement | 73/96 | 12/33 |
|
||||||
|
| `e470435` V-661 | 73/96 | 12/33 |
|
||||||
|
| `a886217` V-663 | 73/96 | 12/33 |
|
||||||
|
| `50c6637` merge #205, parent of the fix | 73/96 | 12/33 |
|
||||||
|
| **`feabf9f` V-664, the tokenizer fix** | **72/96** | **11/33** |
|
||||||
|
| `c938148` V-666 | 72/96 | 11/33 |
|
||||||
|
| `2cf8b7e` master today | 72/96 | 11/33 |
|
||||||
|
|
||||||
|
Nothing between `a886217` and `feabf9f` touches routing except that commit and
|
||||||
|
one merge, and the merge holds the old pair.
|
||||||
|
|
||||||
|
## What this means for the pair
|
||||||
|
|
||||||
|
The higher pair is not a target to restore. Before `feabf9f`, `encodeWord`
|
||||||
|
emitted the characters of every long word in reverse. So the 76.0% and 36.4%
|
||||||
|
were scored on embeddings from a broken encoder. One case landed right under the
|
||||||
|
defect and lands wrong under the fix.
|
||||||
|
|
||||||
|
So 75.0% and 33.3% is the honest floor, and `en-query-002` is not a regression to
|
||||||
|
chase. It is a case the corrected encoder gets wrong, worth the same attention as
|
||||||
|
the other 23 intent misses and no more.
|
||||||
|
|
||||||
|
## What this does not measure
|
||||||
|
|
||||||
|
The routing heads row and the two gemma rows in `docs/routing.md` were not
|
||||||
|
re-run. The heads commits all land after `feabf9f`, so that row was already
|
||||||
|
measured against the corrected tokenizer. The gemma rows are the model arm and
|
||||||
|
were not checked either way.
|
||||||
+12
-4
@@ -1,6 +1,6 @@
|
|||||||
# Routing
|
# Routing
|
||||||
|
|
||||||
*Last verified: 2026-08-11 @ 25ed201*
|
*Last verified: 2026-08-11 @ 2cf8b7e*
|
||||||
|
|
||||||
How an utterance becomes a `Decision`, why each stage exists, and what every
|
How an utterance becomes a `Decision`, why each stage exists, and what every
|
||||||
stage has measured. `CLAUDE.md` carries the rules an agent must not break. This
|
stage has measured. `CLAUDE.md` carries the rules an agent must not break. This
|
||||||
@@ -60,7 +60,7 @@ classifier and the resident model, since those are what always answer.
|
|||||||
|
|
||||||
| Arm | Intent | Destination | p50 | Measured |
|
| Arm | Intent | Destination | p50 | Measured |
|
||||||
|---|---|---|---|---|
|
|---|---|---|---|---|
|
||||||
| classifier + ONNX | 76.0% (73/96) | 36.4% (12/33) | 16.6µs | 2026-08-08 |
|
| classifier + ONNX | 75.0% (72/96) | 33.3% (11/33) | 20.6ms | 2026-08-11 |
|
||||||
| resident Qwen3-1.7B, cascade | 80.2% | not measured | 1.19s | 2026-08-05 |
|
| resident Qwen3-1.7B, cascade | 80.2% | not measured | 1.19s | 2026-08-05 |
|
||||||
| routing heads, cascade | 96.9% | 75.8% | 27.9ms | 2026-08-08 |
|
| routing heads, cascade | 96.9% | 75.8% | 27.9ms | 2026-08-08 |
|
||||||
| gemma-4-12b, cascade | 84.4% | 72.7% | 329ms | 2026-08-02 |
|
| gemma-4-12b, cascade | 84.4% | 72.7% | 329ms | 2026-08-02 |
|
||||||
@@ -72,7 +72,15 @@ another number on the same fixture. Sources:
|
|||||||
`docs/evals/2026-08-02-workstation-gemma4-12b.md`,
|
`docs/evals/2026-08-02-workstation-gemma4-12b.md`,
|
||||||
`docs/evals/2026-08-09-e4b-vs-12b-routing.md`,
|
`docs/evals/2026-08-09-e4b-vs-12b-routing.md`,
|
||||||
`docs/evals/2026-08-08-routing-heads-in-go.md`,
|
`docs/evals/2026-08-08-routing-heads-in-go.md`,
|
||||||
`docs/evals/2026-08-08-destination-fixture.md`.
|
`docs/evals/2026-08-11-classifier-baseline-after-tokenizer-fix.md`.
|
||||||
|
|
||||||
|
The classifier row read 76.0% and 36.4% until 2026-08-11, from
|
||||||
|
`docs/evals/2026-08-08-destination-fixture.md`. That pair predates `feabf9f`,
|
||||||
|
which fixed `encodeWord` emitting long words in reverse, so it was scored on
|
||||||
|
embeddings from a broken encoder. The fix cost one case, `en-query-002`, and that
|
||||||
|
case carries the floor destination label, so it moved both numbers at once. The
|
||||||
|
row also carried 16.6µs, which came from the classifier measured without the
|
||||||
|
embedder. 20.6ms is the ONNX arm's own p50.
|
||||||
|
|
||||||
The resident model alone scores 37.4% full against 61.5% intent-only. The gap is
|
The resident model alone scores 37.4% full against 61.5% intent-only. The gap is
|
||||||
slots and not routing. It routes `reminder` and leaves the time to the daemon,
|
slots and not routing. It routes `reminder` and leaves the time to the daemon,
|
||||||
@@ -201,7 +209,7 @@ are `docs/evals/2026-08-08-routing-heads-two-head.md`,
|
|||||||
|
|
||||||
Read the best destination run as one seed and not a headline. One case is 3
|
Read the best destination run as one seed and not a headline. One case is 3
|
||||||
points on a fixture this small. Head intent accuracy is **not** comparable to the
|
points on a fixture this small. Head intent accuracy is **not** comparable to the
|
||||||
cascade's 76.0% and 84.4%. A softmax has no clarify class, so the head's fixture
|
cascade's 75.0% and 84.4%. A softmax has no clarify class, so the head's fixture
|
||||||
is 88 cases and not 96.
|
is 88 cases and not 96.
|
||||||
|
|
||||||
Recall is 15/15 and world is 5/5.
|
Recall is 15/15 and world is 5/5.
|
||||||
|
|||||||
Reference in New Issue
Block a user