Files
Maven/docs/evals/2026-08-08-slot-head-three-head.md
claude 01e80fce4a Record a fortnight of usage as a re-runnable baseline (V-661)
The 2026-08-07 week of usage was typed by hand and cannot be replayed, so
it measured a build and not a change. scripts/usage-run.py drives the same
reach from a turns file, which makes the next run a diff.

Baseline is master at beb093a: 140 turns, p50 1.6s, zero errors. Three
defects to move. A parked reminder clarify contaminates 19 later turns and
survives a day boundary. Query sources that guess claim six turns they
cannot answer, which is the class V-655 removes. And one question was read
as a capture.

Also records the slot head: gemma distils 2178 spans, three heads score
intent 92.8%, destination 82.8%, slot span F1 72.4% over three seeds.
2026-08-08 20:27:27 +04:00

4.6 KiB

A slot head, and the corpus that did not exist this morning

Measured 2026-08-08 on workpc, the same day as 2026-08-08-routing-heads-two-head.md and against the same fixtures. Workspace is ~/Programs/embed-training, new scripts slot_grammar.gbnf, slot_system.txt, label_slots.py, merge_slots.py.

The corpus was the whole problem

The two-head measurement said BIO slot tags stay in the MASSIVE body, because no Maven-domain span corpus exists. That was true of found corpora and false of made ones. Destination had the same shape at breakfast. V-660 gave gemma-4-12b a string to write and the label problem became a generation problem.

The same trick applies to spans. slot_grammar.gbnf emits a list of {"slot": ..., "text": ...} and the enum closes over Maven's own five: time, text, key, value, fn. slot_system.txt demands each span be an exact substring of the utterance.

The agreement filter is free here. Destination needed a second pass. The daemon's own router prompt had to route each generated line back. A span needs no second call. It either occurs in the utterance or it does not, and label_slots.py drops it with find().

1702 rows labelled from train_v5.jsonl, the reminder, fact, note, act and query intents. Chat and system carry no slot and were never asked.

slot spans
text 1175
time 485
fn 381
key 72
value 65

37 spans dropped as not-a-substring, 2.2% of the pile. Nothing failed to parse, which is the grammar doing its job. 409 rows came back with no span at all. Those are kept and tagged all O. An utterance carrying no slot teaches the head not to invent one. An empty list is a label and not a miss.

key and value are thin because they come from facts alone. That is the shape of the corpus, not a labeller failure.

Three heads on one forward pass

Intent and destination were already two linear heads over one masked mean pool. Slots is a third head over the per-token states of the same pass, so the marginal cost is one Linear(384, 11).

The tag set is O plus B- and I- for each of the five. A softmax cannot emit a tag that does not exist. That is the structural guarantee the GBNF buys for the teacher, and the head gets it for free.

Two masking rules, both ignore_index. A row with no spans key contributes nothing, which covers the 1900 generated destination rows and every chat and system turn. A padding or special-token position contributes nothing either.

Scoring is exact-match span F1, not token accuracy. Most tokens are O, so a tagger that predicts nothing anywhere scores above 90% on tokens.

Result

Three seeds, 24 epochs, epoch chosen on the intent dev slice alone.

two heads three heads
intent mean 93.6% 92.8%
destination mean 80.8% 82.8%
destination best 29/33 (87.9%) 29/33 (87.9%)
slot span F1 mean 72.4%

The slot head costs nothing and adds a third decision. Intent moves 0.8 points down and destination 2 points up. Both sit inside the seed spread those two numbers already had. Read this as unchanged, not as a trade.

The saved checkpoint is seed 1 at epoch 10: intent 93.2%, destination 81.8%, slot F1 75.8%. heads.pt now carries three state dicts and the bio list beside the intent and source enums.

Epoch selection is now wrong for one of the three heads

Slot F1 was still climbing when the intent-selected epoch stopped it. Seed 0 selects epoch 13 at 70.9% and reaches 76.1% at epoch 20. Seed 1 selects epoch 10 at 75.8% and reaches 80.0% at epoch 24.

So the three tasks want different epochs and the harness picks one. Two ways out, and neither was taken here. Select on a joint score, which needs an argument about weights. Or give the slot head its own dev slice and its own early stop, which means the heads stop being one checkpoint.

Leaving it costs about 4 points of slot F1 and nothing else. The router contract is graded on intent. Selecting on it is the conservative choice rather than an oversight.

What this does not measure

The same gap as the two-head run. Nothing of this runs in Go. Three heads instead of two does not change that. It does change the export. The slot head reads the per-token states, so an ONNX graph exposing only the pooled vector is not enough.

There is no held-out span fixture. Span F1 is measured on the dev slice of the same distilled corpus, so it scores agreement with gemma rather than correctness. The V-572 defect utterance, "напомни в 11:00 позвонить маме", is labelled correctly by the teacher, but one utterance is not a fixture.

key at 72 spans and value at 65 are too thin to report a per-slot number.