Train the destination head and beat the teacher (V-661)

Step 3 of the routing-heads plan. Intent and destination share one masked mean
pool on e5-small. Destination scores 26/33 against 12/33 for the classifier
cascade and 24/33 for the cascade with gemma-4-12b, which is the teacher these
labels were distilled from. Recall goes 0/15 to 15/15.

Two heads, not four, and both cuts are label problems rather than GPU time.
Mood describes her own reply state and no dataset maps onto it. BIO slot tags
have no Maven-domain corpus.

The MASSIVE warm-start from step 2 is worth nothing here either. Stock ties it
on intent and leads by a third of a case on destination.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
This commit is contained in:
2026-08-08 19:38:23 +04:00
parent e470435cf1
commit f55bedee2e
2 changed files with 201 additions and 0 deletions
@@ -0,0 +1,174 @@
# Two heads on e5-small, and the first destination the router did not need a model for
Measured 2026-08-08 on workpc (Radeon RX 7900 GRE, ROCm). Covers V-661, step 3 of
`docs/plans/18-routing-heads-on-e5-small.md`. Workspace is `~/Programs/embed-training`,
scripts `gen_query_source.py`, `label_source.py`, `build_heads_corpus.py`,
`train_heads.py`, `score_confidence.py`.
## Two heads, not four
Intent is 7 classes and destination is 12 plus the `SourceUnknown` floor, sharing
one masked mean pool over one forward pass. The plan asked for four. Two of them
have no labels and neither is a GPU problem.
**Mood is cut, not deferred.** Maven's enum is `neutral, happy, thinking, tired,
confused` and it describes her own reply state, not the speaker's emotion.
`psytechlab/EmpatheticIntents-ru` was the only candidate and its 32 emotion
labels do not map onto it. There is nothing to train against.
**BIO slot tags stay in the MASSIVE body.** No Maven-domain span corpus exists.
`2026-08-08-massive-warm-start.md` records that no second Russian slot-filling
corpus is reachable at all.
The destination loss is masked with `ignore_index`. Only a query turn reaches
`queryWalk`, so a reminder contributes nothing to it.
## Where the destination labels came from
V-660 taught the router prompt to name a destination. That made gemma-4-12b a
teacher, and this distils it.
Labelling the 300 query rows already in `train_v5.jsonl` gave 277 destinations.
The shape was unusable: the floor 101, calendar 62, recall 47, and `feeds` and
`attention` at zero. A 13-way softmax cannot learn a class with no examples.
`gen_query_source.py` is the destination half of `gen_corpus.py` and runs the
same two passes. Gemma writes questions whose answer lives in one named place.
The daemon's own `routeSystem` prompt then routes each one back. A line survives
only when the intent is `query` **and** the source is the destination it was
generated for. The glosses are copied verbatim out of `route_system.txt`, so the
generator and the labeller work from one definition.
The prompt and the GBNF are dumped from `internal/router/llmrouter.go` by
`TestDumpPrompt`, never retyped. The workspace held its own copies and V-660
changed both.
1229 kept of 2373 generated, 51.8%. Merged corpus is 3664 rows carrying 1727
destinations:
| | rows | | rows |
|---|---|---|---|
| the floor | 220 | world | 132 |
| calendar | 180 | money | 126 |
| tasks | 169 | list | 124 |
| recall | 167 | self, feeds, attention | 120 each |
| weather | 136 | network | 103 |
| | | home | 50 |
`home` is thin because the agreement filter rejected most of what was generated
for it. A question about the house routes `act` more often than `query`. That is
the filter working, and 50 is the finding rather than a shortfall.
The 1229 generated rows carry `intent: null`. Every one is a query by
construction. There are five times as many as the corpus has query rows, so
including them would make query half the intent corpus.
## Result
Three seeds, two bodies, epoch chosen on the intent dev slice and never on a
destination number.
| body | intent mean | destination mean |
|---|---|---|
| warm-started `out/body_massive` | 93.6% | 75.8% |
| stock `multilingual-e5-small` | 93.6% | 76.8% |
Best single run is destination **26/33 (78.8%)**, reached by both bodies at seed
0. Peak 1.68GB of 17.2GB, under four minutes end to end.
Against the two arms already measured on the same 33 labelled cases:
| | destination |
|---|---|
| classifier cascade (V-659) | 12/33 (36.4%) |
| cascade + gemma-4-12b (V-660) | 24/33 (72.7%) |
| two heads on e5-small | 26/33 (78.8%) |
A 118M encoder beats the 12B teacher it was distilled from, on the fixture. The
per-destination split is where it happens: **recall 15/15** and **world 5/5**.
Recall was 0/15 on the cascade and 14/15 through gemma.
Intent is **not** comparable to the 73/96 and 81/96 figures those two arms
scored. A softmax has no clarify class. The head's fixture is the 88 cases that
carry an intent, and the 8 `want_clarify` cases are scored separately below.
## The MASSIVE warm-start is worth nothing here either
Step 2 measured it at +0.4 points of intent accuracy and called that inside seed
noise. Destination was the open question, because MASSIVE has a
`definition_word` slot that looked like a `SourceWorld` signal sitting in a head
already trained.
It is not. The two bodies score the same intent mean to one decimal. Stock is
one point ahead on destination, which is a third of one case. Nothing here argues
for keeping the warm-start step. Dropping it removes a dependency on a corpus
pull that `datasets` 5.0 cannot do.
## What the head gets wrong is the floor
All seven destination misses at seed 0 are the floor and calendar:
```
(floor) 3/7
calendar 3/6
recall 15/15
world 5/5
```
The head names a destination where the fixture says walk the chain, and it is
confident doing it. `"почему сервер тормозит"` reads `world` at 0.80.
`"хватает ли места под новые бэкапы"` reads `network` at 0.82. Those are the six
homelab cases V-659 flagged, where `SourceRecall`, `SourceNetwork` and
`SourceAttention` all overlap because `mavpoll` writes its observations into the
fact store recall reads.
The training floor is generated ambiguous questions. The fixture floor is
homelab operations. Those are not the same distribution and the head learned the
one it was given.
## Max softmax separates, weakly, and the gate stays
The plan argues max softmax is a calibratable confidence where `Confidence: 1.0`
was a hardcode. Measured on the intent head:
| | n | mean confidence |
|---|---|---|
| correct | 80 | 0.897 |
| wrong | 8 | 0.705 |
| `want_clarify` | 8 | 0.685 |
The softest correct answer is 0.66 and 3 of 8 clarify cases sit below it. So a
single cut buys three clarifies at no false-clarify cost, and no more.
The other five explain themselves. `"напомни"` scores 0.94 as `reminder` and
`"сделай это"` scores 0.80 as `act`. Both are intent-certain and slot-empty, and
confidence was never the signal there. `gateLLMDecision` already catches exactly
that shape, an act with no allowlisted fn or a keyless fact, and it keeps doing
so. The head replaces the hardcode. It does not replace the gate.
## An incident worth recording
The first generation run produced zero rows for eight destinations. `mavgpud`
yields the card when another process wants it (V-488) and llama-server answers
503 until the model is back. Every generate call inside that window burned one of
the destination's batches. The run walked its own cap without a single successful
call. The log said `503` 260 times, and the summary line said 64.1% keep rate,
which read as success.
`call()` now retries a 503 with backoff. A generator that treats an unloaded
model as a bad generation is a silent-corpus bug, not a slow one.
## What this does not measure
**Nothing here runs in Go.** The heads are a `heads.pt` and an
`out/body_heads/` on workpc. Reaching the daemon needs an ONNX export and a
caller. The resident e5-small must not be replaced by this copy: recall depends
on that file, and `EmbedQuery`/`EmbedPassage` are its contract.
The destination fixture carries 4 of 13 classes: recall 15, floor 7, calendar 6,
world 5. `tasks`, `money`, `list`, `home`, `network`, `feeds`, `attention` and
`self` have no gold case. So 78.8% is silent on eight destinations that together
hold 800 training rows.
Latency was not measured. A forward pass of a 118M encoder should beat a 1.7B
decoder on a query turn. That is arithmetic, not a number from this box.