Step 3 of the routing-heads plan. Intent and destination share one masked mean pool on e5-small. Destination scores 26/33 against 12/33 for the classifier cascade and 24/33 for the cascade with gemma-4-12b, which is the teacher these labels were distilled from. Recall goes 0/15 to 15/15. Two heads, not four, and both cuts are label problems rather than GPU time. Mood describes her own reply state and no dataset maps onto it. BIO slot tags have no Maven-domain corpus. The MASSIVE warm-start from step 2 is worth nothing here either. Stock ties it on intent and leads by a third of a case on destination. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
7.5 KiB
Two heads on e5-small, and the first destination the router did not need a model for
Measured 2026-08-08 on workpc (Radeon RX 7900 GRE, ROCm). Covers V-661, step 3 of
docs/plans/18-routing-heads-on-e5-small.md. Workspace is ~/Programs/embed-training,
scripts gen_query_source.py, label_source.py, build_heads_corpus.py,
train_heads.py, score_confidence.py.
Two heads, not four
Intent is 7 classes and destination is 12 plus the SourceUnknown floor, sharing
one masked mean pool over one forward pass. The plan asked for four. Two of them
have no labels and neither is a GPU problem.
Mood is cut, not deferred. Maven's enum is neutral, happy, thinking, tired, confused and it describes her own reply state, not the speaker's emotion.
psytechlab/EmpatheticIntents-ru was the only candidate and its 32 emotion
labels do not map onto it. There is nothing to train against.
BIO slot tags stay in the MASSIVE body. No Maven-domain span corpus exists.
2026-08-08-massive-warm-start.md records that no second Russian slot-filling
corpus is reachable at all.
The destination loss is masked with ignore_index. Only a query turn reaches
queryWalk, so a reminder contributes nothing to it.
Where the destination labels came from
V-660 taught the router prompt to name a destination. That made gemma-4-12b a teacher, and this distils it.
Labelling the 300 query rows already in train_v5.jsonl gave 277 destinations.
The shape was unusable: the floor 101, calendar 62, recall 47, and feeds and
attention at zero. A 13-way softmax cannot learn a class with no examples.
gen_query_source.py is the destination half of gen_corpus.py and runs the
same two passes. Gemma writes questions whose answer lives in one named place.
The daemon's own routeSystem prompt then routes each one back. A line survives
only when the intent is query and the source is the destination it was
generated for. The glosses are copied verbatim out of route_system.txt, so the
generator and the labeller work from one definition.
The prompt and the GBNF are dumped from internal/router/llmrouter.go by
TestDumpPrompt, never retyped. The workspace held its own copies and V-660
changed both.
1229 kept of 2373 generated, 51.8%. Merged corpus is 3664 rows carrying 1727 destinations:
| rows | rows | ||
|---|---|---|---|
| the floor | 220 | world | 132 |
| calendar | 180 | money | 126 |
| tasks | 169 | list | 124 |
| recall | 167 | self, feeds, attention | 120 each |
| weather | 136 | network | 103 |
| home | 50 |
home is thin because the agreement filter rejected most of what was generated
for it. A question about the house routes act more often than query. That is
the filter working, and 50 is the finding rather than a shortfall.
The 1229 generated rows carry intent: null. Every one is a query by
construction. There are five times as many as the corpus has query rows, so
including them would make query half the intent corpus.
Result
Three seeds, two bodies, epoch chosen on the intent dev slice and never on a destination number.
| body | intent mean | destination mean |
|---|---|---|
warm-started out/body_massive |
93.6% | 75.8% |
stock multilingual-e5-small |
93.6% | 76.8% |
Best single run is destination 26/33 (78.8%), reached by both bodies at seed 0. Peak 1.68GB of 17.2GB, under four minutes end to end.
Against the two arms already measured on the same 33 labelled cases:
| destination | |
|---|---|
| classifier cascade (V-659) | 12/33 (36.4%) |
| cascade + gemma-4-12b (V-660) | 24/33 (72.7%) |
| two heads on e5-small | 26/33 (78.8%) |
A 118M encoder beats the 12B teacher it was distilled from, on the fixture. The per-destination split is where it happens: recall 15/15 and world 5/5. Recall was 0/15 on the cascade and 14/15 through gemma.
Intent is not comparable to the 73/96 and 81/96 figures those two arms
scored. A softmax has no clarify class. The head's fixture is the 88 cases that
carry an intent, and the 8 want_clarify cases are scored separately below.
The MASSIVE warm-start is worth nothing here either
Step 2 measured it at +0.4 points of intent accuracy and called that inside seed
noise. Destination was the open question, because MASSIVE has a
definition_word slot that looked like a SourceWorld signal sitting in a head
already trained.
It is not. The two bodies score the same intent mean to one decimal. Stock is
one point ahead on destination, which is a third of one case. Nothing here argues
for keeping the warm-start step. Dropping it removes a dependency on a corpus
pull that datasets 5.0 cannot do.
What the head gets wrong is the floor
All seven destination misses at seed 0 are the floor and calendar:
(floor) 3/7
calendar 3/6
recall 15/15
world 5/5
The head names a destination where the fixture says walk the chain, and it is
confident doing it. "почему сервер тормозит" reads world at 0.80.
"хватает ли места под новые бэкапы" reads network at 0.82. Those are the six
homelab cases V-659 flagged, where SourceRecall, SourceNetwork and
SourceAttention all overlap because mavpoll writes its observations into the
fact store recall reads.
The training floor is generated ambiguous questions. The fixture floor is homelab operations. Those are not the same distribution and the head learned the one it was given.
Max softmax separates, weakly, and the gate stays
The plan argues max softmax is a calibratable confidence where Confidence: 1.0
was a hardcode. Measured on the intent head:
| n | mean confidence | |
|---|---|---|
| correct | 80 | 0.897 |
| wrong | 8 | 0.705 |
want_clarify |
8 | 0.685 |
The softest correct answer is 0.66 and 3 of 8 clarify cases sit below it. So a single cut buys three clarifies at no false-clarify cost, and no more.
The other five explain themselves. "напомни" scores 0.94 as reminder and
"сделай это" scores 0.80 as act. Both are intent-certain and slot-empty, and
confidence was never the signal there. gateLLMDecision already catches exactly
that shape, an act with no allowlisted fn or a keyless fact, and it keeps doing
so. The head replaces the hardcode. It does not replace the gate.
An incident worth recording
The first generation run produced zero rows for eight destinations. mavgpud
yields the card when another process wants it (V-488) and llama-server answers
503 until the model is back. Every generate call inside that window burned one of
the destination's batches. The run walked its own cap without a single successful
call. The log said 503 260 times, and the summary line said 64.1% keep rate,
which read as success.
call() now retries a 503 with backoff. A generator that treats an unloaded
model as a bad generation is a silent-corpus bug, not a slow one.
What this does not measure
Nothing here runs in Go. The heads are a heads.pt and an
out/body_heads/ on workpc. Reaching the daemon needs an ONNX export and a
caller. The resident e5-small must not be replaced by this copy: recall depends
on that file, and EmbedQuery/EmbedPassage are its contract.
The destination fixture carries 4 of 13 classes: recall 15, floor 7, calendar 6,
world 5. tasks, money, list, home, network, feeds, attention and
self have no gold case. So 78.8% is silent on eight destinations that together
hold 800 training rows.
Latency was not measured. A forward pass of a 118M encoder should beat a 1.7B decoder on a query turn. That is arithmetic, not a number from this box.