17 KiB
On the frozen 1652-row five-way residual non-action task the deployed trained cascade (routing heads + e5 centroid classifier, LLM out) already matches a fresh linear head on the same frozen embeddings — 0.715 vs 0.712 macro-F1 — while the deterministic hash-classifier floor resolves only 10.9% of the pool (acc 0.171, 89.1% clarify); cross-family generalization is structurally absent (leave-one-family-out macro-F1 0.150; fact holdout acc 0.018; system/conversation/uncertain holdouts 0.000), the trained cascade calls an executable action on 143 residual non-action rows (8.7%; 91 of them knowledge), and the five-way head swallows 94% of the 766 action OOD rows into non-action buckets at argmax — a new representation or a new head are not justified; the levers are deterministic
Date: 2026-09-08 · Task: slice 22 (brief after the accepted slice 21, task/725) · Box: workpc, Arch, RX 7900 GRE, 32 GB · Build: cmd/semantic-router-experiment/slice22 (Go 1.25.12, -mode legacy/-mode heads) + slice22_emit.py + slice22_main.py in the frozen venv /tmp/mvn-exp-venv (sklearn 1.9.0). ONNX for the heads run via MAVEN_ONNX_LIB pointing at the venv's libonnxruntime.so.1.29.0.
0. Frozen artifacts and population
corpus (3266 → frozen 3025): dev pool 2490, byte-identical to slices 16-21
dev residual (fast_path_resolved=false): 2418
residual non-action pool (this slice): 1652
knowledge 715 / memory_write 553 / system 184 / conversation 93 / uncertain 107
53 family_id = 53 split_group (no row differs); folds {0:317, 1:140, 2:410, 3:408, 4:377}
action rows held out as OOD probes: 766
frozen embeddings: model_quantized@384/tok2, mean-pool + L2, 'query: ' prefix
Artifacts under /tmp/mvn-s22/: pool.json (1652 rows, idx = position in dev
pool dev-order, so embeddings align by index with slice19.load_dev),
ood.json (766), legacy.json (mode legacy), legacy_heads.json (mode
heads), results.json (all sections below). All three JSON outputs are
reproducible: the Go harness is deterministic, and the Python sections carry
random_state=42.
1. What is being asked
When the fast path misses, the execution-frame guard passes, and no executable action is warranted, the utterance is one of five residual non-action semantics: conversation, knowledge, memory_write, system, uncertain. This slice measures whether the deployed e5-small embeddings (384-d, frozen) with a linear softmax head suffice for that five-way decision, against three floors and two deployed baselines. The executable-action population is out of scope for the head (closed in slices 18-21); its rows are OOD probes only.
Decision criteria carried from the brief, judged in §13: e5-linear sufficient / sparse sufficient / corpus-taxonomy problem / new representation justified.
2. Method and baselines
Five-way projection rules (the daemon's behaviour, project() in
legacy_main.go): chat → conversation, query → knowledge,
fact/note → memory_write, system → system; act/reminder on a trusted
non-action row is recorded verbatim as action with
illegal_action_prediction=true and never softly re-mapped to uncertain;
Clarify and route errors become uncertain. In the metrics of §4-§5 the 143
action projections count as errors (predicted uncertain), so the reported
scores already carry the price of the false-action leak.
Three routers measured (all: stage-0 grammars, StubDateTimeParser,
DefaultFactParser, act allowlist of 18 verbs, threshold 0.55):
legacy buildMinimalRouter(): StageZeroGrammars + HashEmbedder(1024)
nearest-centroid classifier, seeds models/seeds (331 examples)
→ the deterministic floor, no model, no ONNX
heads grammars + RouterHeads (router_heads.onnx, 0.6 decline) + e5 ONNX
embedder nearest-centroid classifier (same seeds)
→ the deployed cascade with the LLM stage out (docs/routing.md
pickLLMRouter → classifier); the production offline behaviour
The two modes were smoke-verified to diverge on the same 40-row slice (the
NN heads fire at softmax 0.94-0.98 on residual rows the hash floor clarifies at
0.16-0.24), and the full-population files were checked for internal
consistency (producer/stage/stage-0 counts add to 1652) before use. Any number
in this report comes from results.json, not from a transcript.
3. Stage-0 grammar hit under the live rule set: corpus drift
185 of the 1652 "residual" rows are resolved by the current stage-0
grammars (producer=grammar): by intent query 157 / note 14 / chat 7 / system 6 / reminder 1, by truth-route knowledge 157 / memory_write 12 / system 6 / conversation 2 / uncertain 8. corpus fast_path_resolved was set
by a corpus-side mirror (cmd/corpus-factory/main.go), not by the live rule
set; a grammar landed (or the mirror froze) after the corpus was built. The
pool remains 1652 for the primary metrics — dropping the 185 would not move
any conclusion below — and this drift is reported separately as
grammar_drift: hits 185 / grammar-pure 1467. The mirror should be reconciled
with the live grammar set.
4. Baseline A — the deterministic floor (legacy)
acc 0.1707 macro-F1 0.1202 illegal_action_prediction 3
conversation P 0.000 R 0.000 F1 0.000 n=93
knowledge P 1.000 R 0.220 F1 0.360 n=715
memory_write P 0.857 R 0.022 F1 0.042 n=553
system P 1.000 R 0.033 F1 0.063 n=184
uncertain P 0.073 R 1.000 F1 0.135 n=107
The hash classifier fires above 0.55 on 180 rows (10.9%); the other 1472
(89.1%) are below-threshold uncertain. The three illegal predictions are all
uncertain:incomplete-reminder ("напомни", "напомни, пожалуйста", "я хочу
напомни") — the reminder grammar fires and the daemon would drive an action
probe on a row the corpus trusted as non-action. This is the boundary class
the deterministic floor cannot price, and it is small.
5. Baseline B — the deployed cascade with the LLM out (heads)
acc 0.7724 macro-F1 0.7150 illegal_action_prediction 143
conversation P 0.550 R 0.828 F1 0.661 n=93
knowledge P 0.858 R 0.744 F1 0.797 n=715
memory_write P 0.940 R 0.825 F1 0.879 n=553
system P 0.924 R 0.658 F1 0.768 n=184
uncertain P 0.326 R 0.841 F1 0.470 n=107
Producer mix: NN-heads 1176, e5-classifier 291, grammar 185. The NN heads name
act on 159 rows and reminder on 18; 143 of those land on trusted
non-action rows (illegal), split knowledge 91 / memory_write 46 / system 2 / uncertain 4. By family the biggest offenders are knowledge:capability-ha 45,
knowledge:capability-tool 38 (i.e. 83 capability-question rows — the slice-21
population — called executable), mw-free-remember 35, kq-homelab-disk 5,
mw-note-task 5. Producer split of the illegal 143: classifier 102 (conf
0.89-0.91), NN-heads 40, grammar 1. Example rows: "хватает ли места на диске"
→ act 0.90; "я хочу покажи шаги за неделю" → reminder 0.91; "машинное
обучение опиши" → act 0.89. The trained components cannot refuse: their
largest directional error is toward execution, which is exactly the boundary
the deterministic stage-0 + execution-frame-guard arm owns.
This is the same shape the fresh head shows in §8: confidence 0.89-0.91 is not a refusal signal either.
6. The fresh e5-linear head (the ask)
Out-of-fold, grouped by the pool's split folds, C grid {0.1, 1, 10}:
C=0.1 acc 0.6096 macro-F1 0.2728
C=1.0 acc 0.6949 macro-F1 0.5974
C=10.0 acc 0.7222 macro-F1 0.7117 ← used everywhere below
Per-fold OOF (C=10): fold 0 acc 0.782 / macro 0.783; fold 1 (n=140) 0.736 / 0.605; fold 2 0.705 / 0.576; fold 3 0.713 / 0.710; fold 4 0.695 / 0.592; mean acc 0.726 ± 0.031, mean macro-F1 0.653 ± 0.080.
OOF per class (C=10):
conversation P 0.840 R 0.731 F1 0.782 n=93
knowledge P 0.713 R 0.761 F1 0.736 n=715
memory_write P 0.720 R 0.749 F1 0.734 n=553
system P 0.687 R 0.500 F1 0.579 n=184
uncertain P 0.758 R 0.701 F1 0.728 n=107
The fresh head and the deployed cascade land at the same operating point: OOF 0.712 vs deployed 0.715 macro-F1, 0.722 vs 0.772 accuracy. The linear head on the deployed embeddings does not beat what is already wired, and the accuracy gap is explained by the 143 illegal rows the head cannot make (it has no action class) — a structural, not a representational, difference.
7. Floors
majority (all knowledge) acc 0.4328 macro-F1 0.1208
centroid (cosine, in-fold means) acc 0.6731 macro-F1 0.6494
sparse word+char TF-IDF logistic acc 0.6247 macro-F1 0.4463 vocab 7079
e5-linear (the ask) acc 0.7222 macro-F1 0.7117
Sparse is below e5-linear on both axes despite a 7079-term vocabulary; the lexical signal is present but orders of magnitude weaker than the embedding. Centroid cosine is the closest floor and e5-linear beats it by +0.062 macro-F1.
8. Cross-family generalization is absent
Route-family leave-*-out on the frozen 53 families. Each group below is a single-route family set, so the route is literally absent from training when the family is removed:
family (route) n acc macro-F1 dominant prediction
capability (kno) 126 0.2619 0.0830 knowledge (F1 0.42)
world (kno) 143 0.7622 0.1730 knowledge (F1 0.87)
calendar (kno) 92 0.6304 0.1547 knowledge (F1 0.77)
recall (kno) 114 0.7632 0.1731 knowledge (F1 0.87)
fact (mwr) 389 0.0180 0.0071 memory_write (F1 0.04)
note (mwr) 104 0.2212 0.0724 memory_write (F1 0.36)
remember (mwr) 60 0.2000 0.0667 memory_write (F1 0.33)
system (sys) 184 0.0000 0.0000 (nothing)
conversation (conv) 93 0.0000 0.0000 (nothing)
uncertain (unc) 107 0.0000 0.0000 (nothing)
leave-one-family-out over all 53: acc mean 0.682, macro-F1 mean 0.149
The world/calendar/recall rows read as "good acc" only because knowledge is the majority: the head's macro-F1 collapses (0.15-0.17). On fact, correct predictions fall to 1.8%; on system, conversation and uncertain to exactly zero. Within-family OOF macro-F1 is 0.71 (all routes present), and removing one family drops it to 0.15. The head is a per-family-template memoriser: it transfers no route semantics to an unseen template family. This is the ceiling of the frozen representation, not a hyperparameter problem.
9. knowledge vs memory_write, matched pairs
build_pairs pairs every recall/knowledge row against every write row that
shares a corpus-justified lexeme, and asks whether the model puts the higher
probability on the truth side (memory_write for write, knowledge for recall):
lexeme pairs mw_over_k order_acc mean_margin
"вод" 480 0.863 +0.298
"задач" 720 0.735 +0.122
"dns/сервер/vlan" 180 0.106 -0.228 ← homelab notes vs homelab status
The fact/recall contrast ("воды попил" vs "сколько воды") separates well. The
collision is the note-keep vs report-family: note:homelab and
note:task write rows are ranked below their knowledge:homelab-status and
knowledge:task-check partners 89% of the time, with a negative margin. The
frozen embeddings put both sides of "дом" as one cluster; the head cannot
resolve intent direction there. This is the concrete, smallest instance of the
§8 collapse and the strongest single argument that a deterministic
note-vs-recall disambiguator (subject class, scheduling verbs, report vs
record verbs) would buy more than any representation retrain.
10. uncertain as an explicit class
OOF: uncertain P 0.758 / R 0.701 / F1 0.728 (n=107)
pulled out of uncertain (true uncertain, mislabelled): knowledge 14,
memory_write 11, conversation 7, system 0
pulled into uncertain (not uncertain, predicted uncertain): conversation 9,
knowledge 7, memory_write 5, system 3
uncertain is taught and recalled (R 0.70), and the pulls are small and
mirrored — the "ask instead" bucket behaves like a class, which is what the
daemon needs it to be only if the ask threshold is real (§11).
11. Calibration, abstention, and the action OOD
Confidence (max-softmax of the OOF head): correct-mean 0.766, wrong-mean 0.612, ECE 0.0275, log-loss 0.7365, macro-Brier 0.0781. The head is well calibrated on-pool.
threshold (max-softmax) coverage accuracy macro-F1
0.10 1.000 0.722 0.712
0.28 0.996 0.725 0.715
0.46 0.907 0.756 0.757
0.64 0.676 0.822 0.829
0.82 0.379 0.887 0.913
Action OOD probes — the 766 frozen action rows, scored by the same five-way head, are the wrong-population test the head must refuse:
argmax class: knowledge 370 (48.3%), memory_write 347 (45.3%),
uncertain 38 (5.0%), conversation 11, system 0
conf mean 0.653, conf median 0.644, conf > 0.9 on 46 rows (6.0%)
94.6% of action utterances are named a resolvable non-action route at argmax, and 6% land above the pool's own 0.82 high-confidence line. A five-way head replacing the current cascade would turn action rows into confident knowledge/write decisions. Two consequences: (a) the head must stay behind the deterministic act/reminder + execution-frame-guard arm and never precede it — the current cascade order already does that; (b) it cannot double as an action-refuser by thresholding, so the reminder/act grammar family stays the only refusal mechanism, and its 143-row miss rate (§5) is the real open boundary for the trained arm, priced at 8.7% of residual non-action.
12. Artifact cost
Linear head: 1925 params (5×384 + 5 bias), 7.52 KiB fp32, incremental decision latency mean 190 µs / p50 165 µs (pure head, Python timing on the box; the shared cost is the e5 forward the routing-heads path already pays). The head itself is negligible; the representation forward is not the lever.
13. Verdict on the decision criteria
| criterion | verdict |
|---|---|
| e5-linear suffices | No. In-fold parity with what is deployed (0.712 vs 0.715 macro-F1), but it transfers nothing across template families (LOO macro-F1 0.149) and cannot abstain on action rows (94.6% swallowed). |
| sparse suffices | No. 0.6247 / 0.4463, below the linear head on both axes. |
| corpus-taxonomy problem | Partly. 185 rows the corpus called residual are resolved by the live grammar set today (mirror staleness), and the two largest illegal-action families are capability questions — a corpus-trust boundary, not a representation one. |
| new representation justified | No on this evidence. The ceiling is family-lexicon transfer, not embedding capacity: replacing the frozen e5 changes nothing measured; the floors show the same collapse. A new representation would need to demonstrate the §8 holdouts moving before it earns a retrain. |
14. Operational implications
- Keep the cascade order. Stage-0 grammars and the execution-frame guard (slice 21) are the only two mechanisms that keep the 8.7% false-action leak from reaching tools; neither the linear head nor the deployed NN heads can be promoted to the action boundary.
- Reconcile the corpus fast-path mirror against the live grammar set (185 drifting rows). Purely deterministic, zero-model, and removes a class of "legacy understates the router" arguments.
- Add a deterministic note-vs-recall disambiguator for the homelab/task lexical clusters (the 0.106 ordering in §9): subject class, report verbs vs record verbs, date/scheduling adverbs. Expected lever: most of the memory_write recall gap without touching embeddings.
- Watch the classifier's confident
actcalls (conf 0.89-0.91 on capability-question and free-remember rows, §5). The "хватает ли места на диске" →act0.90 row is a slice-21 capability question that the trained arm still prices as executable; the deterministic guard is the only thing standing between it and an action probe. - Do not threshold the uncertain bucket for action refusals. The 38/766 uncertain picks on OOD action rows show an explicit uncertain class is near useless as a guard; the numbers are confidence, not category.
15. Commit hashes
- Tooling (emit step, Go harness both modes, Python experiment):
29f74dd(router/semantic: slice 22 residual non-action router — emit, harness, experiment). - Report + eval index row: this file with its
docs/evals/CLAUDE.mdentry. - Artifacts under
/tmp/mvn-s22/(not committed; reproducible byslice22_emit.py, thengo run ./cmd/semantic-router-experiment/slice22/ -mode legacyand-mode headswithMAVEN_ONNX_LIB, thenslice22_main.py).