2.7 KiB
The classifier baseline after the tokenizer fix
Measured 2026-08-11 on master at 2cf8b7e, the classifier cascade with the ONNX
multilingual embedder, the configuration homesrv runs. make t PKG=./internal/router/eval/ RUN=TestONNXBaseline V=1. Covers V-704.
Result
Intent is 72/96 (75.0%). Destination is 11/33 (33.3%).
docs/evals/2026-08-08-destination-fixture.md recorded 73/96 (76.0%) and 12/33
(36.4%) three days earlier, and CLAUDE.md and docs/routing.md both carried
that pair until today. Same fixture, same 33 labelled cases, so the two pairs
are comparable and one case moved.
| 2026-08-08 | today | |
|---|---|---|
| intent | 73/96 (76.0%) | 72/96 (75.0%) |
| destination | 12/33 (36.4%) | 11/33 (33.3%) |
| world | 5/5 | 5/5 |
the SourceUnknown floor |
5/7 | 4/7 |
| calendar | 2/6 | 2/6 |
| recall | 0/15 | 0/15 |
One case moved, and it cost both numbers
en-query-002, "how long since the last backup finished". It routed query on
08-08 and routes fact today, at 0.870.
The case carries want_source present and empty, the floor label. Score
credits a destination only after the route lands the intent. So the case paid
once on the intent and once on the destination. Two numbers moved by one case,
which is why the drift read as two regressions.
The commit
feabf9f, "The tokenizer read every long word backwards" (V-664), landed
2026-08-08 after the destination measurement.
Bisected by running TestONNXBaseline at each commit:
| commit | intent | destination |
|---|---|---|
b6eaa70 V-659, the 08-08 measurement |
73/96 | 12/33 |
e470435 V-661 |
73/96 | 12/33 |
a886217 V-663 |
73/96 | 12/33 |
50c6637 merge #205, parent of the fix |
73/96 | 12/33 |
feabf9f V-664, the tokenizer fix |
72/96 | 11/33 |
c938148 V-666 |
72/96 | 11/33 |
2cf8b7e master today |
72/96 | 11/33 |
Nothing between a886217 and feabf9f touches routing except that commit and
one merge, and the merge holds the old pair.
What this means for the pair
The higher pair is not a target to restore. Before feabf9f, encodeWord
emitted the characters of every long word in reverse. So the 76.0% and 36.4%
were scored on embeddings from a broken encoder. One case landed right under the
defect and lands wrong under the fix.
So 75.0% and 33.3% is the honest floor, and en-query-002 is not a regression to
chase. It is a case the corrected encoder gets wrong, worth the same attention as
the other 23 intent misses and no more.
What this does not measure
The routing heads row and the two gemma rows in docs/routing.md were not
re-run. The heads commits all land after feabf9f, so that row was already
measured against the corrected tokenizer. The gemma rows are the model arm and
were not checked either way.