Train the destination head and beat the teacher (V-661)
Step 3 of the routing-heads plan. Intent and destination share one masked mean pool on e5-small. Destination scores 26/33 against 12/33 for the classifier cascade and 24/33 for the cascade with gemma-4-12b, which is the teacher these labels were distilled from. Recall goes 0/15 to 15/15. Two heads, not four, and both cuts are label problems rather than GPU time. Mood describes her own reply state and no dataset maps onto it. BIO slot tags have no Maven-domain corpus. The MASSIVE warm-start from step 2 is worth nothing here either. Stock ties it on intent and leads by a third of a case on destination. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
This commit is contained in:
@@ -246,6 +246,33 @@ was a hardcode. **Fine-tune a copy of the weights.** The resident embedder backs
|
||||
recall. Training it in place couples routing accuracy to recall@1, with nothing in the
|
||||
suite to name the trade.
|
||||
|
||||
**Two of those heads are trained as of 08-08-2026, and they are not the three
|
||||
above** (V-661, `docs/evals/2026-08-08-routing-heads-two-head.md`). Intent and
|
||||
destination share one masked mean pool. Destination scores **26/33 (78.8%)** on
|
||||
the fixture. The classifier cascade scores 12/33 and the cascade with gemma-4-12b
|
||||
scores 24/33, so a 118M encoder beats the 12B teacher it was distilled from.
|
||||
Recall is 15/15 and world is 5/5. Intent is 93.6% mean over three seeds. That is
|
||||
**not** comparable to the 76.0% and 84.4% those two arms scored: a softmax has no
|
||||
clarify class, so the head's fixture is the 88 cases carrying an intent.
|
||||
|
||||
**Mood is cut, not deferred.** The enum describes her own reply state, not the
|
||||
speaker's emotion, and no dataset maps onto it. **BIO slot tags have no
|
||||
Maven-domain corpus**, so they stay in the MASSIVE body from step 2. Both are
|
||||
label problems and neither is a GPU problem: the run is under four minutes.
|
||||
|
||||
The MASSIVE warm-start of step 2 is worth nothing here. Stock e5-small ties it on
|
||||
intent and leads by a third of a case on destination. Nothing argues for keeping
|
||||
that step.
|
||||
|
||||
What the head gets wrong is the floor. It names a destination where the fixture
|
||||
says walk the chain, and it is confident doing it. `"почему сервер тормозит"`
|
||||
reads `world` at 0.80. The training floor is generated ambiguous questions and the fixture floor
|
||||
is homelab operations, which are not the same distribution.
|
||||
|
||||
**Nothing of this runs in Go.** The weights are `heads.pt` and `out/body_heads/`
|
||||
on workpc. Reaching the daemon needs an ONNX export and a caller. The resident
|
||||
e5-small must not be replaced by the copy, because recall depends on that file.
|
||||
|
||||
`Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
|
||||
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
|
||||
#359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with
|
||||
|
||||
@@ -0,0 +1,174 @@
|
||||
# Two heads on e5-small, and the first destination the router did not need a model for
|
||||
|
||||
Measured 2026-08-08 on workpc (Radeon RX 7900 GRE, ROCm). Covers V-661, step 3 of
|
||||
`docs/plans/18-routing-heads-on-e5-small.md`. Workspace is `~/Programs/embed-training`,
|
||||
scripts `gen_query_source.py`, `label_source.py`, `build_heads_corpus.py`,
|
||||
`train_heads.py`, `score_confidence.py`.
|
||||
|
||||
## Two heads, not four
|
||||
|
||||
Intent is 7 classes and destination is 12 plus the `SourceUnknown` floor, sharing
|
||||
one masked mean pool over one forward pass. The plan asked for four. Two of them
|
||||
have no labels and neither is a GPU problem.
|
||||
|
||||
**Mood is cut, not deferred.** Maven's enum is `neutral, happy, thinking, tired,
|
||||
confused` and it describes her own reply state, not the speaker's emotion.
|
||||
`psytechlab/EmpatheticIntents-ru` was the only candidate and its 32 emotion
|
||||
labels do not map onto it. There is nothing to train against.
|
||||
|
||||
**BIO slot tags stay in the MASSIVE body.** No Maven-domain span corpus exists.
|
||||
`2026-08-08-massive-warm-start.md` records that no second Russian slot-filling
|
||||
corpus is reachable at all.
|
||||
|
||||
The destination loss is masked with `ignore_index`. Only a query turn reaches
|
||||
`queryWalk`, so a reminder contributes nothing to it.
|
||||
|
||||
## Where the destination labels came from
|
||||
|
||||
V-660 taught the router prompt to name a destination. That made gemma-4-12b a
|
||||
teacher, and this distils it.
|
||||
|
||||
Labelling the 300 query rows already in `train_v5.jsonl` gave 277 destinations.
|
||||
The shape was unusable: the floor 101, calendar 62, recall 47, and `feeds` and
|
||||
`attention` at zero. A 13-way softmax cannot learn a class with no examples.
|
||||
|
||||
`gen_query_source.py` is the destination half of `gen_corpus.py` and runs the
|
||||
same two passes. Gemma writes questions whose answer lives in one named place.
|
||||
The daemon's own `routeSystem` prompt then routes each one back. A line survives
|
||||
only when the intent is `query` **and** the source is the destination it was
|
||||
generated for. The glosses are copied verbatim out of `route_system.txt`, so the
|
||||
generator and the labeller work from one definition.
|
||||
|
||||
The prompt and the GBNF are dumped from `internal/router/llmrouter.go` by
|
||||
`TestDumpPrompt`, never retyped. The workspace held its own copies and V-660
|
||||
changed both.
|
||||
|
||||
1229 kept of 2373 generated, 51.8%. Merged corpus is 3664 rows carrying 1727
|
||||
destinations:
|
||||
|
||||
| | rows | | rows |
|
||||
|---|---|---|---|
|
||||
| the floor | 220 | world | 132 |
|
||||
| calendar | 180 | money | 126 |
|
||||
| tasks | 169 | list | 124 |
|
||||
| recall | 167 | self, feeds, attention | 120 each |
|
||||
| weather | 136 | network | 103 |
|
||||
| | | home | 50 |
|
||||
|
||||
`home` is thin because the agreement filter rejected most of what was generated
|
||||
for it. A question about the house routes `act` more often than `query`. That is
|
||||
the filter working, and 50 is the finding rather than a shortfall.
|
||||
|
||||
The 1229 generated rows carry `intent: null`. Every one is a query by
|
||||
construction. There are five times as many as the corpus has query rows, so
|
||||
including them would make query half the intent corpus.
|
||||
|
||||
## Result
|
||||
|
||||
Three seeds, two bodies, epoch chosen on the intent dev slice and never on a
|
||||
destination number.
|
||||
|
||||
| body | intent mean | destination mean |
|
||||
|---|---|---|
|
||||
| warm-started `out/body_massive` | 93.6% | 75.8% |
|
||||
| stock `multilingual-e5-small` | 93.6% | 76.8% |
|
||||
|
||||
Best single run is destination **26/33 (78.8%)**, reached by both bodies at seed
|
||||
0. Peak 1.68GB of 17.2GB, under four minutes end to end.
|
||||
|
||||
Against the two arms already measured on the same 33 labelled cases:
|
||||
|
||||
| | destination |
|
||||
|---|---|
|
||||
| classifier cascade (V-659) | 12/33 (36.4%) |
|
||||
| cascade + gemma-4-12b (V-660) | 24/33 (72.7%) |
|
||||
| two heads on e5-small | 26/33 (78.8%) |
|
||||
|
||||
A 118M encoder beats the 12B teacher it was distilled from, on the fixture. The
|
||||
per-destination split is where it happens: **recall 15/15** and **world 5/5**.
|
||||
Recall was 0/15 on the cascade and 14/15 through gemma.
|
||||
|
||||
Intent is **not** comparable to the 73/96 and 81/96 figures those two arms
|
||||
scored. A softmax has no clarify class. The head's fixture is the 88 cases that
|
||||
carry an intent, and the 8 `want_clarify` cases are scored separately below.
|
||||
|
||||
## The MASSIVE warm-start is worth nothing here either
|
||||
|
||||
Step 2 measured it at +0.4 points of intent accuracy and called that inside seed
|
||||
noise. Destination was the open question, because MASSIVE has a
|
||||
`definition_word` slot that looked like a `SourceWorld` signal sitting in a head
|
||||
already trained.
|
||||
|
||||
It is not. The two bodies score the same intent mean to one decimal. Stock is
|
||||
one point ahead on destination, which is a third of one case. Nothing here argues
|
||||
for keeping the warm-start step. Dropping it removes a dependency on a corpus
|
||||
pull that `datasets` 5.0 cannot do.
|
||||
|
||||
## What the head gets wrong is the floor
|
||||
|
||||
All seven destination misses at seed 0 are the floor and calendar:
|
||||
|
||||
```
|
||||
(floor) 3/7
|
||||
calendar 3/6
|
||||
recall 15/15
|
||||
world 5/5
|
||||
```
|
||||
|
||||
The head names a destination where the fixture says walk the chain, and it is
|
||||
confident doing it. `"почему сервер тормозит"` reads `world` at 0.80.
|
||||
`"хватает ли места под новые бэкапы"` reads `network` at 0.82. Those are the six
|
||||
homelab cases V-659 flagged, where `SourceRecall`, `SourceNetwork` and
|
||||
`SourceAttention` all overlap because `mavpoll` writes its observations into the
|
||||
fact store recall reads.
|
||||
|
||||
The training floor is generated ambiguous questions. The fixture floor is
|
||||
homelab operations. Those are not the same distribution and the head learned the
|
||||
one it was given.
|
||||
|
||||
## Max softmax separates, weakly, and the gate stays
|
||||
|
||||
The plan argues max softmax is a calibratable confidence where `Confidence: 1.0`
|
||||
was a hardcode. Measured on the intent head:
|
||||
|
||||
| | n | mean confidence |
|
||||
|---|---|---|
|
||||
| correct | 80 | 0.897 |
|
||||
| wrong | 8 | 0.705 |
|
||||
| `want_clarify` | 8 | 0.685 |
|
||||
|
||||
The softest correct answer is 0.66 and 3 of 8 clarify cases sit below it. So a
|
||||
single cut buys three clarifies at no false-clarify cost, and no more.
|
||||
|
||||
The other five explain themselves. `"напомни"` scores 0.94 as `reminder` and
|
||||
`"сделай это"` scores 0.80 as `act`. Both are intent-certain and slot-empty, and
|
||||
confidence was never the signal there. `gateLLMDecision` already catches exactly
|
||||
that shape, an act with no allowlisted fn or a keyless fact, and it keeps doing
|
||||
so. The head replaces the hardcode. It does not replace the gate.
|
||||
|
||||
## An incident worth recording
|
||||
|
||||
The first generation run produced zero rows for eight destinations. `mavgpud`
|
||||
yields the card when another process wants it (V-488) and llama-server answers
|
||||
503 until the model is back. Every generate call inside that window burned one of
|
||||
the destination's batches. The run walked its own cap without a single successful
|
||||
call. The log said `503` 260 times, and the summary line said 64.1% keep rate,
|
||||
which read as success.
|
||||
|
||||
`call()` now retries a 503 with backoff. A generator that treats an unloaded
|
||||
model as a bad generation is a silent-corpus bug, not a slow one.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
**Nothing here runs in Go.** The heads are a `heads.pt` and an
|
||||
`out/body_heads/` on workpc. Reaching the daemon needs an ONNX export and a
|
||||
caller. The resident e5-small must not be replaced by this copy: recall depends
|
||||
on that file, and `EmbedQuery`/`EmbedPassage` are its contract.
|
||||
|
||||
The destination fixture carries 4 of 13 classes: recall 15, floor 7, calendar 6,
|
||||
world 5. `tasks`, `money`, `list`, `home`, `network`, `feeds`, `attention` and
|
||||
`self` have no gold case. So 78.8% is silent on eight destinations that together
|
||||
hold 800 training rows.
|
||||
|
||||
Latency was not measured. A forward pass of a 118M encoder should beat a 1.7B
|
||||
decoder on a query turn. That is arithmetic, not a number from this box.
|
||||
Reference in New Issue
Block a user