# A slot head, and the corpus that did not exist this morning Measured 2026-08-08 on workpc, the same day as `2026-08-08-routing-heads-two-head.md` and against the same fixtures. Workspace is `~/Programs/embed-training`, new scripts `slot_grammar.gbnf`, `slot_system.txt`, `label_slots.py`, `merge_slots.py`. ## The corpus was the whole problem The two-head measurement said BIO slot tags stay in the MASSIVE body, because no Maven-domain span corpus exists. That was true of found corpora and false of made ones. Destination had the same shape at breakfast. V-660 gave gemma-4-12b a string to write and the label problem became a generation problem. The same trick applies to spans. `slot_grammar.gbnf` emits a list of `{"slot": ..., "text": ...}` and the enum closes over Maven's own five: `time`, `text`, `key`, `value`, `fn`. `slot_system.txt` demands each span be an exact substring of the utterance. **The agreement filter is free here.** Destination needed a second pass. The daemon's own router prompt had to route each generated line back. A span needs no second call. It either occurs in the utterance or it does not, and `label_slots.py` drops it with `find()`. 1702 rows labelled from `train_v5.jsonl`, the reminder, fact, note, act and query intents. Chat and system carry no slot and were never asked. | slot | spans | |---|---| | text | 1175 | | time | 485 | | fn | 381 | | key | 72 | | value | 65 | 37 spans dropped as not-a-substring, 2.2% of the pile. Nothing failed to parse, which is the grammar doing its job. 409 rows came back with no span at all. Those are kept and tagged all `O`. An utterance carrying no slot teaches the head not to invent one. An empty list is a label and not a miss. `key` and `value` are thin because they come from facts alone. That is the shape of the corpus, not a labeller failure. ## Three heads on one forward pass Intent and destination were already two linear heads over one masked mean pool. Slots is a third head over the per-token states of the same pass, so the marginal cost is one `Linear(384, 11)`. The tag set is `O` plus `B-` and `I-` for each of the five. A softmax cannot emit a tag that does not exist. That is the structural guarantee the GBNF buys for the teacher, and the head gets it for free. Two masking rules, both `ignore_index`. A row with no `spans` key contributes nothing, which covers the 1900 generated destination rows and every chat and system turn. A padding or special-token position contributes nothing either. Scoring is exact-match span F1, not token accuracy. Most tokens are `O`, so a tagger that predicts nothing anywhere scores above 90% on tokens. ## Result Three seeds, 24 epochs, epoch chosen on the intent dev slice alone. | | two heads | three heads | |---|---|---| | intent mean | 93.6% | 92.8% | | destination mean | 80.8% | 82.8% | | destination best | 29/33 (87.9%) | 29/33 (87.9%) | | slot span F1 mean | — | 72.4% | **The slot head costs nothing and adds a third decision.** Intent moves 0.8 points down and destination 2 points up. Both sit inside the seed spread those two numbers already had. Read this as unchanged, not as a trade. The saved checkpoint is seed 1 at epoch 10: intent 93.2%, destination 81.8%, slot F1 75.8%. `heads.pt` now carries three state dicts and the `bio` list beside the intent and source enums. ## Epoch selection is now wrong for one of the three heads Slot F1 was still climbing when the intent-selected epoch stopped it. Seed 0 selects epoch 13 at 70.9% and reaches 76.1% at epoch 20. Seed 1 selects epoch 10 at 75.8% and reaches 80.0% at epoch 24. So the three tasks want different epochs and the harness picks one. Two ways out, and neither was taken here. Select on a joint score, which needs an argument about weights. Or give the slot head its own dev slice and its own early stop, which means the heads stop being one checkpoint. Leaving it costs about 4 points of slot F1 and nothing else. The router contract is graded on intent. Selecting on it is the conservative choice rather than an oversight. ## What this does not measure The same gap as the two-head run. **Nothing of this runs in Go.** Three heads instead of two does not change that. It does change the export. The slot head reads the per-token states, so an ONNX graph exposing only the pooled vector is not enough. There is no held-out span fixture. Span F1 is measured on the dev slice of the same distilled corpus, so it scores agreement with gemma rather than correctness. The V-572 defect utterance, `"напомни в 11:00 позвонить маме"`, is labelled correctly by the teacher, but one utterance is not a fixture. `key` at 72 spans and `value` at 65 are too thin to report a per-slot number.