Routing heads: a slot head, a clarify head, and a two-week baseline to diff against #203
@@ -246,6 +246,33 @@ was a hardcode. **Fine-tune a copy of the weights.** The resident embedder backs
|
||||
recall. Training it in place couples routing accuracy to recall@1, with nothing in the
|
||||
suite to name the trade.
|
||||
|
||||
**Two of those heads are trained as of 08-08-2026, and they are not the three
|
||||
above** (V-661, `docs/evals/2026-08-08-routing-heads-two-head.md`). Intent and
|
||||
destination share one masked mean pool. Destination scores **26/33 (78.8%)** on
|
||||
the fixture. The classifier cascade scores 12/33 and the cascade with gemma-4-12b
|
||||
scores 24/33, so a 118M encoder beats the 12B teacher it was distilled from.
|
||||
Recall is 15/15 and world is 5/5. Intent is 93.6% mean over three seeds. That is
|
||||
**not** comparable to the 76.0% and 84.4% those two arms scored: a softmax has no
|
||||
clarify class, so the head's fixture is the 88 cases carrying an intent.
|
||||
|
||||
**Mood is cut, not deferred.** The enum describes her own reply state, not the
|
||||
speaker's emotion, and no dataset maps onto it. **BIO slot tags have no
|
||||
Maven-domain corpus**, so they stay in the MASSIVE body from step 2. Both are
|
||||
label problems and neither is a GPU problem: the run is under four minutes.
|
||||
|
||||
The MASSIVE warm-start of step 2 is worth nothing here. Stock e5-small ties it on
|
||||
intent and leads by a third of a case on destination. Nothing argues for keeping
|
||||
that step.
|
||||
|
||||
What the head gets wrong is the floor. It names a destination where the fixture
|
||||
says walk the chain, and it is confident doing it. `"почему сервер тормозит"`
|
||||
reads `world` at 0.80. The training floor is generated ambiguous questions and the fixture floor
|
||||
is homelab operations, which are not the same distribution.
|
||||
|
||||
**Nothing of this runs in Go.** The weights are `heads.pt` and `out/body_heads/`
|
||||
on workpc. Reaching the daemon needs an ONNX export and a caller. The resident
|
||||
e5-small must not be replaced by the copy, because recall depends on that file.
|
||||
|
||||
`Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
|
||||
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
|
||||
#359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with
|
||||
|
||||
@@ -0,0 +1,174 @@
|
||||
# Two heads on e5-small, and the first destination the router did not need a model for
|
||||
|
||||
Measured 2026-08-08 on workpc (Radeon RX 7900 GRE, ROCm). Covers V-661, step 3 of
|
||||
`docs/plans/18-routing-heads-on-e5-small.md`. Workspace is `~/Programs/embed-training`,
|
||||
scripts `gen_query_source.py`, `label_source.py`, `build_heads_corpus.py`,
|
||||
`train_heads.py`, `score_confidence.py`.
|
||||
|
||||
## Two heads, not four
|
||||
|
||||
Intent is 7 classes and destination is 12 plus the `SourceUnknown` floor, sharing
|
||||
one masked mean pool over one forward pass. The plan asked for four. Two of them
|
||||
have no labels and neither is a GPU problem.
|
||||
|
||||
**Mood is cut, not deferred.** Maven's enum is `neutral, happy, thinking, tired,
|
||||
confused` and it describes her own reply state, not the speaker's emotion.
|
||||
`psytechlab/EmpatheticIntents-ru` was the only candidate and its 32 emotion
|
||||
labels do not map onto it. There is nothing to train against.
|
||||
|
||||
**BIO slot tags stay in the MASSIVE body.** No Maven-domain span corpus exists.
|
||||
`2026-08-08-massive-warm-start.md` records that no second Russian slot-filling
|
||||
corpus is reachable at all.
|
||||
|
||||
The destination loss is masked with `ignore_index`. Only a query turn reaches
|
||||
`queryWalk`, so a reminder contributes nothing to it.
|
||||
|
||||
## Where the destination labels came from
|
||||
|
||||
V-660 taught the router prompt to name a destination. That made gemma-4-12b a
|
||||
teacher, and this distils it.
|
||||
|
||||
Labelling the 300 query rows already in `train_v5.jsonl` gave 277 destinations.
|
||||
The shape was unusable: the floor 101, calendar 62, recall 47, and `feeds` and
|
||||
`attention` at zero. A 13-way softmax cannot learn a class with no examples.
|
||||
|
||||
`gen_query_source.py` is the destination half of `gen_corpus.py` and runs the
|
||||
same two passes. Gemma writes questions whose answer lives in one named place.
|
||||
The daemon's own `routeSystem` prompt then routes each one back. A line survives
|
||||
only when the intent is `query` **and** the source is the destination it was
|
||||
generated for. The glosses are copied verbatim out of `route_system.txt`, so the
|
||||
generator and the labeller work from one definition.
|
||||
|
||||
The prompt and the GBNF are dumped from `internal/router/llmrouter.go` by
|
||||
`TestDumpPrompt`, never retyped. The workspace held its own copies and V-660
|
||||
changed both.
|
||||
|
||||
1229 kept of 2373 generated, 51.8%. Merged corpus is 3664 rows carrying 1727
|
||||
destinations:
|
||||
|
||||
| | rows | | rows |
|
||||
|---|---|---|---|
|
||||
| the floor | 220 | world | 132 |
|
||||
| calendar | 180 | money | 126 |
|
||||
| tasks | 169 | list | 124 |
|
||||
| recall | 167 | self, feeds, attention | 120 each |
|
||||
| weather | 136 | network | 103 |
|
||||
| | | home | 50 |
|
||||
|
||||
`home` is thin because the agreement filter rejected most of what was generated
|
||||
for it. A question about the house routes `act` more often than `query`. That is
|
||||
the filter working, and 50 is the finding rather than a shortfall.
|
||||
|
||||
The 1229 generated rows carry `intent: null`. Every one is a query by
|
||||
construction. There are five times as many as the corpus has query rows, so
|
||||
including them would make query half the intent corpus.
|
||||
|
||||
## Result
|
||||
|
||||
Three seeds, two bodies, epoch chosen on the intent dev slice and never on a
|
||||
destination number.
|
||||
|
||||
| body | intent mean | destination mean |
|
||||
|---|---|---|
|
||||
| warm-started `out/body_massive` | 93.6% | 75.8% |
|
||||
| stock `multilingual-e5-small` | 93.6% | 76.8% |
|
||||
|
||||
Best single run is destination **26/33 (78.8%)**, reached by both bodies at seed
|
||||
0. Peak 1.68GB of 17.2GB, under four minutes end to end.
|
||||
|
||||
Against the two arms already measured on the same 33 labelled cases:
|
||||
|
||||
| | destination |
|
||||
|---|---|
|
||||
| classifier cascade (V-659) | 12/33 (36.4%) |
|
||||
| cascade + gemma-4-12b (V-660) | 24/33 (72.7%) |
|
||||
| two heads on e5-small | 26/33 (78.8%) |
|
||||
|
||||
A 118M encoder beats the 12B teacher it was distilled from, on the fixture. The
|
||||
per-destination split is where it happens: **recall 15/15** and **world 5/5**.
|
||||
Recall was 0/15 on the cascade and 14/15 through gemma.
|
||||
|
||||
Intent is **not** comparable to the 73/96 and 81/96 figures those two arms
|
||||
scored. A softmax has no clarify class. The head's fixture is the 88 cases that
|
||||
carry an intent, and the 8 `want_clarify` cases are scored separately below.
|
||||
|
||||
## The MASSIVE warm-start is worth nothing here either
|
||||
|
||||
Step 2 measured it at +0.4 points of intent accuracy and called that inside seed
|
||||
noise. Destination was the open question, because MASSIVE has a
|
||||
`definition_word` slot that looked like a `SourceWorld` signal sitting in a head
|
||||
already trained.
|
||||
|
||||
It is not. The two bodies score the same intent mean to one decimal. Stock is
|
||||
one point ahead on destination, which is a third of one case. Nothing here argues
|
||||
for keeping the warm-start step. Dropping it removes a dependency on a corpus
|
||||
pull that `datasets` 5.0 cannot do.
|
||||
|
||||
## What the head gets wrong is the floor
|
||||
|
||||
All seven destination misses at seed 0 are the floor and calendar:
|
||||
|
||||
```
|
||||
(floor) 3/7
|
||||
calendar 3/6
|
||||
recall 15/15
|
||||
world 5/5
|
||||
```
|
||||
|
||||
The head names a destination where the fixture says walk the chain, and it is
|
||||
confident doing it. `"почему сервер тормозит"` reads `world` at 0.80.
|
||||
`"хватает ли места под новые бэкапы"` reads `network` at 0.82. Those are the six
|
||||
homelab cases V-659 flagged, where `SourceRecall`, `SourceNetwork` and
|
||||
`SourceAttention` all overlap because `mavpoll` writes its observations into the
|
||||
fact store recall reads.
|
||||
|
||||
The training floor is generated ambiguous questions. The fixture floor is
|
||||
homelab operations. Those are not the same distribution and the head learned the
|
||||
one it was given.
|
||||
|
||||
## Max softmax separates, weakly, and the gate stays
|
||||
|
||||
The plan argues max softmax is a calibratable confidence where `Confidence: 1.0`
|
||||
was a hardcode. Measured on the intent head:
|
||||
|
||||
| | n | mean confidence |
|
||||
|---|---|---|
|
||||
| correct | 80 | 0.897 |
|
||||
| wrong | 8 | 0.705 |
|
||||
| `want_clarify` | 8 | 0.685 |
|
||||
|
||||
The softest correct answer is 0.66 and 3 of 8 clarify cases sit below it. So a
|
||||
single cut buys three clarifies at no false-clarify cost, and no more.
|
||||
|
||||
The other five explain themselves. `"напомни"` scores 0.94 as `reminder` and
|
||||
`"сделай это"` scores 0.80 as `act`. Both are intent-certain and slot-empty, and
|
||||
confidence was never the signal there. `gateLLMDecision` already catches exactly
|
||||
that shape, an act with no allowlisted fn or a keyless fact, and it keeps doing
|
||||
so. The head replaces the hardcode. It does not replace the gate.
|
||||
|
||||
## An incident worth recording
|
||||
|
||||
The first generation run produced zero rows for eight destinations. `mavgpud`
|
||||
yields the card when another process wants it (V-488) and llama-server answers
|
||||
503 until the model is back. Every generate call inside that window burned one of
|
||||
the destination's batches. The run walked its own cap without a single successful
|
||||
call. The log said `503` 260 times, and the summary line said 64.1% keep rate,
|
||||
which read as success.
|
||||
|
||||
`call()` now retries a 503 with backoff. A generator that treats an unloaded
|
||||
model as a bad generation is a silent-corpus bug, not a slow one.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
**Nothing here runs in Go.** The heads are a `heads.pt` and an
|
||||
`out/body_heads/` on workpc. Reaching the daemon needs an ONNX export and a
|
||||
caller. The resident e5-small must not be replaced by this copy: recall depends
|
||||
on that file, and `EmbedQuery`/`EmbedPassage` are its contract.
|
||||
|
||||
The destination fixture carries 4 of 13 classes: recall 15, floor 7, calendar 6,
|
||||
world 5. `tasks`, `money`, `list`, `home`, `network`, `feeds`, `attention` and
|
||||
`self` have no gold case. So 78.8% is silent on eight destinations that together
|
||||
hold 800 training rows.
|
||||
|
||||
Latency was not measured. A forward pass of a 118M encoder should beat a 1.7B
|
||||
decoder on a query turn. That is arithmetic, not a number from this box.
|
||||
Reference in New Issue
Block a user