A fourth head asks instead of guessing (V-661)
Clarify is not a value of intent. It is a second question over the same pooled vector: can Maven act on this at all. The eight want_clarify fixture cases sat outside every number the heads measured, because a softmax has no clarify class. gen_clarify.py makes the class the corpus lacks. Every existing row was generated FOR an intent, so every one is answerable. The router-prompt agreement filter cannot work here, because routeGrammar has no clarify value and a generated line always agrees with itself. A judge replaces it. The first judge called 24 of 40 answerable rows underspecified. It judged against a generic assistant, one that asks where about lunch. Restating Maven's contract took that to 16 of 60, with all eight fixture cases caught. Three seeds: 7.0 of 8 caught, 2.3 false of 88. The cascade today misses 1 and produces 2. Confidence separates too, 0.851 right against 0.604 wrong.
This commit is contained in:
@@ -0,0 +1,123 @@
|
||||
# A clarify head, and a confidence that is not a hardcode
|
||||
|
||||
Measured 2026-08-08 on workpc, the same day and the same fixtures as
|
||||
`2026-08-08-routing-heads-two-head.md` and `2026-08-08-slot-head-three-head.md`.
|
||||
New script `gen_clarify.py`, new fixture `eval_fixture_clarify.jsonl`.
|
||||
|
||||
## A softmax has no clarify class
|
||||
|
||||
That sentence closed the two-head measurement. It is why the head's fixture was
|
||||
88 cases and not 96. The eight `want_clarify` cases sat outside every number
|
||||
measured, and the head had no way to produce the answer they wanted.
|
||||
|
||||
A fourth head is the answer. Clarify is not a value of intent. It is a second
|
||||
question asked of the same pooled vector: can Maven act on this at all.
|
||||
|
||||
## The corpus had one class
|
||||
|
||||
Every row in `train_heads_slots.jsonl` was generated FOR an intent or a
|
||||
destination. So every row is answerable by construction. A head trained on that
|
||||
alone sees one class and learns to say yes.
|
||||
|
||||
`gen_clarify.py` makes the other class. Five shapes, ten topics. The shapes are
|
||||
the gate's own reasons in `gateLLMDecision` plus the two the fixture carries:
|
||||
bare noun, bare verb, demonstrative, deictic time, dangling reference.
|
||||
|
||||
**The agreement filter that worked for destination cannot work here.**
|
||||
`routeGrammar` has no clarify value. So the router always names an intent, and
|
||||
any generated line always agrees with itself. The second pass is a judge
|
||||
instead. Gemma is asked, without seeing the label, whether Maven would have to
|
||||
ask a question back.
|
||||
|
||||
## The first judge was worthless and the second was measured
|
||||
|
||||
The first judge said "needs clarify" on 24 of 40 plainly answerable corpus
|
||||
rows. It flagged `запиши что я пообедал` and `Покажи расписание поездов на
|
||||
вечер`. It was judging against a generic assistant, one that asks "where?"
|
||||
about lunch. Maven writes that note.
|
||||
|
||||
Rewriting it to state what she can already do took false positives to 16 of 60.
|
||||
It also catches all eight fixture clarifies. So the judge discriminates.
|
||||
|
||||
On the generated pile it removed 7 of 306, a 97.7% keep rate. That is not the
|
||||
judge failing. The generator is aimed at underspecified lines, so there is
|
||||
little for a filter to catch. The 27% false-positive rate is the number to
|
||||
quote, and it is label noise on the positive class.
|
||||
|
||||
**Both passes are gemma-4-12b.** Generation and judging. So the corpus is
|
||||
gemma's opinion of what is underspecified, and the head distills that opinion.
|
||||
What keeps it honest is the fixture. Those eight cases were written by the owner
|
||||
and gemma never saw them.
|
||||
|
||||
299 rows kept, against 3604 answerable. The positive class carries `intent:
|
||||
null`, so it costs the intent head nothing.
|
||||
|
||||
## Result
|
||||
|
||||
Three seeds, 24 epochs, epoch still chosen on the intent dev slice.
|
||||
|
||||
| | two heads | three heads | four heads |
|
||||
|---|---|---|---|
|
||||
| intent mean | 93.6% | 92.8% | 91.7% |
|
||||
| destination mean | 80.8% | 82.8% | 79.8% |
|
||||
| slot span F1 mean | — | 72.4% | 68.3% |
|
||||
| clarify caught | — | — | 7.0 of 8 |
|
||||
| false clarifies | — | — | 2.3 of 88 |
|
||||
|
||||
**The fourth head is not free the way the third was.** Intent, destination and
|
||||
slot F1 all move down. The drop is one to four points, and the seed spread is
|
||||
wide enough to contain it. Seed 2 scores intent 94.3% and destination 84.8%, both above
|
||||
every three-head seed. Read the drop as unproven rather than as absent.
|
||||
|
||||
Accuracy is the wrong number for this head and is reported for completeness at
|
||||
95.8% to 96.9%. Eight of ninety-six cases are positive, so a head that never
|
||||
asks scores 91.7%. Recall on those eight is the number.
|
||||
|
||||
Compare it to what ships. The cascade today misses 1 clarify and produces 2
|
||||
false ones. The head catches 7 of 8 and produces 2.3 false ones. That is
|
||||
parity, from a 118M encoder with no rules in front of it.
|
||||
|
||||
The saved checkpoint is seed 2 at epoch 10. Intent 94.3%, destination 28/33,
|
||||
slot F1 73.6%, clarify 7 of 8 with 3 false. `heads.pt` carries four state dicts.
|
||||
|
||||
## What it gets wrong is consistent across seeds
|
||||
|
||||
`поужинал` is a false clarify on all three seeds. That utterance is already
|
||||
recorded as a real defect. `thinSingleToken` was narrowed on 2026-08-01 to spare
|
||||
a token carrying a Russian verb ending. One word is routinely a whole sentence
|
||||
in Russian. The head relearned the mistake the rule was narrowed to
|
||||
fix.
|
||||
|
||||
`что дальше?` is a false clarify on two seeds. That one is a disagreement rather
|
||||
than an error. The utterance is underspecified, and V-498 decided stage 0 claims
|
||||
it for the calendar on purpose.
|
||||
|
||||
`ну это` is missed on two seeds. `амб-003` is the shortest case in the fixture
|
||||
and the generated demonstratives are longer.
|
||||
|
||||
## Confidence
|
||||
|
||||
`Confidence: 1.0` was a hardcode in `llmrouter.go`, so a correct low confidence
|
||||
could not exist. Max softmax over the intent head is the replacement. It is only worth reading
|
||||
if it is lower where the head is wrong.
|
||||
|
||||
It is. Mean 0.851 where the head is right against 0.604 where it is wrong. It
|
||||
ranks a right case above a wrong one in 83.4% of pairs.
|
||||
|
||||
So there are two signals now and they are not the same signal. Confidence says
|
||||
the head is unsure which intent this is. The clarify head says the utterance
|
||||
does not carry enough to act on. A confident wrong route and an honest "I cannot
|
||||
tell" are different failures, and one number cannot report both.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
The same gap as every head run. **Nothing of this runs in Go.** Four heads
|
||||
instead of three does not change that.
|
||||
|
||||
There is no threshold. Both signals are reported as raw numbers. Turning either
|
||||
into a gate needs a decision about where to cut, and that trades false clarifies
|
||||
against wrong acts. The fixture has 8 positives, which is too few to fit a
|
||||
threshold on.
|
||||
|
||||
The 299 generated rows have no held-out slice of their own. Clarify is scored on
|
||||
the fixture alone.
|
||||
Reference in New Issue
Block a user