Compare commits
8 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 9a333b23d7 | |||
| 663b5c47b9 | |||
| 45c521e1a6 | |||
| e34669a52e | |||
| d434f83c2c | |||
| c310115fd2 | |||
| 3024f76e5f | |||
| 6bc71553ab |
@@ -257,6 +257,31 @@ Recall is 15/15 and world is 5/5. Intent is 93.6% mean over three seeds. That is
|
||||
**not** comparable to the 76.0% and 84.4% those two arms scored: a softmax has no
|
||||
clarify class, so the head's fixture is the 88 cases carrying an intent.
|
||||
|
||||
**A fourth head asks instead of guessing, same day** (V-661,
|
||||
`docs/evals/2026-08-08-clarify-head-four-head.md`). Clarify is not a value
|
||||
of intent, so a softmax cannot emit it. It is a second question over the
|
||||
same pooled vector: can Maven act on this at all. That is why the head's
|
||||
fixture was 88 cases and not 96. Over three seeds it catches **7.0 of the 8
|
||||
`want_clarify` cases and produces 2.3 false clarifies of 88**. The cascade
|
||||
today misses 1 and produces 2, so this is parity with no rules in front of
|
||||
it. Accuracy is the wrong number here and a head that never asks scores
|
||||
91.7%. Confidence is the other half. Max softmax over the intent head reads
|
||||
**0.851 where it is right against 0.604 where it is wrong**, ranking right
|
||||
above wrong in 83.4% of pairs. `Confidence: 1.0` was a hardcode, and this
|
||||
replaces it with a signal. The two are not the same signal: one says which
|
||||
intent is unclear, the other says the utterance carries too little to act
|
||||
on. **The fourth head is not free the way the third was.** Intent,
|
||||
destination and slot F1 each move down one to four points, inside the seed
|
||||
spread. `поужинал` is a false clarify on every seed, which is the same
|
||||
defect `thinSingleToken` was narrowed for on 2026-08-01.
|
||||
|
||||
The corpus for it is generated, because every existing row is answerable by
|
||||
construction. **The router-prompt agreement filter cannot work here**, since
|
||||
`routeGrammar` has no clarify value and a generated line always agrees with
|
||||
itself. A gemma judge replaces it. The first judge called 24 of 40
|
||||
answerable rows underspecified, because it judged against a generic
|
||||
assistant rather than against Maven's contract.
|
||||
|
||||
**Mood is cut, not deferred.** The enum describes her own reply state, not the
|
||||
speaker's emotion, and no dataset maps onto it.
|
||||
|
||||
@@ -417,7 +442,19 @@ queries exactly as it did.
|
||||
That is the safety argument and it is not negotiable. The table's order is
|
||||
load-bearing. Every comment on it argues a reason between two sources, and above all
|
||||
it carries "the owner's data first, then the world". Naming `SourceWorld` does not
|
||||
send the turn outside. His notes, his facts and the personal boundary still run first.
|
||||
send the turn outside on its own. His notes and his facts still run first, because
|
||||
they look rather than guess.
|
||||
|
||||
**The personal boundary is the one exception and it is deliberate.** It guesses,
|
||||
so naming `SourceWorld` drops it. That is what stops it answering "кто такой
|
||||
Линус Торвальдс?" with "не нашла у тебя такой записи", which it did on
|
||||
2026-08-07. The cost is that a destination a model wrote can now take the
|
||||
boundary off a turn. A question about him that the model calls `world` reaches
|
||||
SearXNG, where today the boundary stops it. Only the utterance leaves the box,
|
||||
never his notes or history, so this widens what is asked and not what is sent.
|
||||
`TestNamingRecallKeepsTheBoundary` pins the other half: naming `SourceRecall`
|
||||
keeps the boundary in front of the world. Whether a model may drop it at all is
|
||||
the owner's call and has not been made.
|
||||
|
||||
What comes out is only the sources that **guess**. Those decide a turn is theirs by
|
||||
cosine against frozen seeds, then answer whatever they claimed. They hold no table
|
||||
|
||||
@@ -0,0 +1,123 @@
|
||||
# A clarify head, and a confidence that is not a hardcode
|
||||
|
||||
Measured 2026-08-08 on workpc, the same day and the same fixtures as
|
||||
`2026-08-08-routing-heads-two-head.md` and `2026-08-08-slot-head-three-head.md`.
|
||||
New script `gen_clarify.py`, new fixture `eval_fixture_clarify.jsonl`.
|
||||
|
||||
## A softmax has no clarify class
|
||||
|
||||
That sentence closed the two-head measurement. It is why the head's fixture was
|
||||
88 cases and not 96. The eight `want_clarify` cases sat outside every number
|
||||
measured, and the head had no way to produce the answer they wanted.
|
||||
|
||||
A fourth head is the answer. Clarify is not a value of intent. It is a second
|
||||
question asked of the same pooled vector: can Maven act on this at all.
|
||||
|
||||
## The corpus had one class
|
||||
|
||||
Every row in `train_heads_slots.jsonl` was generated FOR an intent or a
|
||||
destination. So every row is answerable by construction. A head trained on that
|
||||
alone sees one class and learns to say yes.
|
||||
|
||||
`gen_clarify.py` makes the other class. Five shapes, ten topics. The shapes are
|
||||
the gate's own reasons in `gateLLMDecision` plus the two the fixture carries:
|
||||
bare noun, bare verb, demonstrative, deictic time, dangling reference.
|
||||
|
||||
**The agreement filter that worked for destination cannot work here.**
|
||||
`routeGrammar` has no clarify value. So the router always names an intent, and
|
||||
any generated line always agrees with itself. The second pass is a judge
|
||||
instead. Gemma is asked, without seeing the label, whether Maven would have to
|
||||
ask a question back.
|
||||
|
||||
## The first judge was worthless and the second was measured
|
||||
|
||||
The first judge said "needs clarify" on 24 of 40 plainly answerable corpus
|
||||
rows. It flagged `запиши что я пообедал` and `Покажи расписание поездов на
|
||||
вечер`. It was judging against a generic assistant, one that asks "where?"
|
||||
about lunch. Maven writes that note.
|
||||
|
||||
Rewriting it to state what she can already do took false positives to 16 of 60.
|
||||
It also catches all eight fixture clarifies. So the judge discriminates.
|
||||
|
||||
On the generated pile it removed 7 of 306, a 97.7% keep rate. That is not the
|
||||
judge failing. The generator is aimed at underspecified lines, so there is
|
||||
little for a filter to catch. The 27% false-positive rate is the number to
|
||||
quote, and it is label noise on the positive class.
|
||||
|
||||
**Both passes are gemma-4-12b.** Generation and judging. So the corpus is
|
||||
gemma's opinion of what is underspecified, and the head distills that opinion.
|
||||
What keeps it honest is the fixture. Those eight cases were written by the owner
|
||||
and gemma never saw them.
|
||||
|
||||
299 rows kept, against 3604 answerable. The positive class carries `intent:
|
||||
null`, so it costs the intent head nothing.
|
||||
|
||||
## Result
|
||||
|
||||
Three seeds, 24 epochs, epoch still chosen on the intent dev slice.
|
||||
|
||||
| | two heads | three heads | four heads |
|
||||
|---|---|---|---|
|
||||
| intent mean | 93.6% | 92.8% | 91.7% |
|
||||
| destination mean | 80.8% | 82.8% | 79.8% |
|
||||
| slot span F1 mean | — | 72.4% | 68.3% |
|
||||
| clarify caught | — | — | 7.0 of 8 |
|
||||
| false clarifies | — | — | 2.3 of 88 |
|
||||
|
||||
**The fourth head is not free the way the third was.** Intent, destination and
|
||||
slot F1 all move down. The drop is one to four points, and the seed spread is
|
||||
wide enough to contain it. Seed 2 scores intent 94.3% and destination 84.8%, both above
|
||||
every three-head seed. Read the drop as unproven rather than as absent.
|
||||
|
||||
Accuracy is the wrong number for this head and is reported for completeness at
|
||||
95.8% to 96.9%. Eight of ninety-six cases are positive, so a head that never
|
||||
asks scores 91.7%. Recall on those eight is the number.
|
||||
|
||||
Compare it to what ships. The cascade today misses 1 clarify and produces 2
|
||||
false ones. The head catches 7 of 8 and produces 2.3 false ones. That is
|
||||
parity, from a 118M encoder with no rules in front of it.
|
||||
|
||||
The saved checkpoint is seed 2 at epoch 10. Intent 94.3%, destination 28/33,
|
||||
slot F1 73.6%, clarify 7 of 8 with 3 false. `heads.pt` carries four state dicts.
|
||||
|
||||
## What it gets wrong is consistent across seeds
|
||||
|
||||
`поужинал` is a false clarify on all three seeds. That utterance is already
|
||||
recorded as a real defect. `thinSingleToken` was narrowed on 2026-08-01 to spare
|
||||
a token carrying a Russian verb ending. One word is routinely a whole sentence
|
||||
in Russian. The head relearned the mistake the rule was narrowed to
|
||||
fix.
|
||||
|
||||
`что дальше?` is a false clarify on two seeds. That one is a disagreement rather
|
||||
than an error. The utterance is underspecified, and V-498 decided stage 0 claims
|
||||
it for the calendar on purpose.
|
||||
|
||||
`ну это` is missed on two seeds. `amb-003` is the shortest case in the fixture
|
||||
and the generated demonstratives are longer.
|
||||
|
||||
## Confidence
|
||||
|
||||
`Confidence: 1.0` was a hardcode in `llmrouter.go`, so a correct low confidence
|
||||
could not exist. Max softmax over the intent head is the replacement. It is only worth reading
|
||||
if it is lower where the head is wrong.
|
||||
|
||||
It is. Mean 0.851 where the head is right against 0.604 where it is wrong. It
|
||||
ranks a right case above a wrong one in 83.4% of pairs.
|
||||
|
||||
So there are two signals now and they are not the same signal. Confidence says
|
||||
the head is unsure which intent this is. The clarify head says the utterance
|
||||
does not carry enough to act on. A confident wrong route and an honest "I cannot
|
||||
tell" are different failures, and one number cannot report both.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
The same gap as every head run. **Nothing of this runs in Go.** Four heads
|
||||
instead of three does not change that.
|
||||
|
||||
There is no threshold. Both signals are reported as raw numbers. Turning either
|
||||
into a gate needs a decision about where to cut, and that trades false clarifies
|
||||
against wrong acts. The fixture has 8 positives, which is too few to fit a
|
||||
threshold on.
|
||||
|
||||
The 299 generated rows have no held-out slice of their own. Clarify is scored on
|
||||
the fixture alone.
|
||||
@@ -30,7 +30,7 @@ present when a re-run reaches day 1.
|
||||
| p50 | 1.5s | 1.6s |
|
||||
| p95 | 8.0s | 7.1s |
|
||||
| max | 12.3s | 33.7s |
|
||||
| errors | 0 | 0 |
|
||||
| transport errors | 0 | 0 |
|
||||
|
||||
| string in the reply | turns |
|
||||
|---|---|
|
||||
@@ -42,6 +42,9 @@ present when a re-run reaches day 1.
|
||||
| `пока не умею` | 5 |
|
||||
| `Когда?` | 3 |
|
||||
|
||||
**Zero transport errors is not zero wrong answers.** It counts turns that
|
||||
failed to return a reply, and none did. Every quality number is below.
|
||||
|
||||
Those seven strings appear in 49 of 140 turns. Some turns carry two, because a
|
||||
parked clarify appends to whatever else was said.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user