Compare commits

...

8 Commits

Author SHA1 Message Date
claude 9a333b23d7 Merge master after 199-201 landed (V-661) 2026-08-08 21:05:27 +04:00
kami 663b5c47b9 Merge pull request 'The router prompt has no destination, so the model arm of V-655 names nothing' (#201) from task/660-router-prompt-destination into master 2026-08-08 19:03:28 +02:00
kami 45c521e1a6 Merge pull request 'Destination fixture: score Decision.Source, not just the intent' (#200) from task/659-destination-fixture into master 2026-08-08 19:03:24 +02:00
kami e34669a52e Merge pull request 'Query source is a routing decision made outside the router' (#199) from task/655-query-source-is-a-routing-decision-made into master 2026-08-08 19:03:06 +02:00
claude d434f83c2c The personal boundary is a guesser, so say so (V-655)
CLAUDE.md said naming SourceWorld leaves his notes, his facts and the
personal boundary running first. The first two are true and the third is
not. The boundary is marked guesses: true, so queryWalk drops it whenever
the named destination is not recall.

That is deliberate and tested. It is what stops the boundary answering
'кто такой Линус Торвальдс?' with 'не нашла у тебя такой записи'. But it
means a destination a model wrote can take the boundary off a turn about
him, and the doc claimed the opposite.

Flagged as the owner's call rather than changed. Only the utterance leaves
the box either way.
2026-08-08 21:02:41 +04:00
claude c310115fd2 Record the clarify head and the confidence it replaces (V-661) 2026-08-08 20:59:44 +04:00
claude 3024f76e5f A fourth head asks instead of guessing (V-661)
Clarify is not a value of intent. It is a second question over the same
pooled vector: can Maven act on this at all. The eight want_clarify fixture
cases sat outside every number the heads measured, because a softmax has no
clarify class.

gen_clarify.py makes the class the corpus lacks. Every existing row was
generated FOR an intent, so every one is answerable. The router-prompt
agreement filter cannot work here, because routeGrammar has no clarify value
and a generated line always agrees with itself. A judge replaces it.

The first judge called 24 of 40 answerable rows underspecified. It judged
against a generic assistant, one that asks where about lunch. Restating
Maven's contract took that to 16 of 60, with all eight fixture cases caught.

Three seeds: 7.0 of 8 caught, 2.3 false of 88. The cascade today misses 1 and
produces 2. Confidence separates too, 0.851 right against 0.604 wrong.
2026-08-08 20:59:21 +04:00
claude 6bc71553ab Say that a transport error is not a wrong answer (V-661) 2026-08-08 20:47:32 +04:00
3 changed files with 165 additions and 2 deletions
+38 -1
View File
@@ -257,6 +257,31 @@ Recall is 15/15 and world is 5/5. Intent is 93.6% mean over three seeds. That is
**not** comparable to the 76.0% and 84.4% those two arms scored: a softmax has no
clarify class, so the head's fixture is the 88 cases carrying an intent.
**A fourth head asks instead of guessing, same day** (V-661,
`docs/evals/2026-08-08-clarify-head-four-head.md`). Clarify is not a value
of intent, so a softmax cannot emit it. It is a second question over the
same pooled vector: can Maven act on this at all. That is why the head's
fixture was 88 cases and not 96. Over three seeds it catches **7.0 of the 8
`want_clarify` cases and produces 2.3 false clarifies of 88**. The cascade
today misses 1 and produces 2, so this is parity with no rules in front of
it. Accuracy is the wrong number here and a head that never asks scores
91.7%. Confidence is the other half. Max softmax over the intent head reads
**0.851 where it is right against 0.604 where it is wrong**, ranking right
above wrong in 83.4% of pairs. `Confidence: 1.0` was a hardcode, and this
replaces it with a signal. The two are not the same signal: one says which
intent is unclear, the other says the utterance carries too little to act
on. **The fourth head is not free the way the third was.** Intent,
destination and slot F1 each move down one to four points, inside the seed
spread. `поужинал` is a false clarify on every seed, which is the same
defect `thinSingleToken` was narrowed for on 2026-08-01.
The corpus for it is generated, because every existing row is answerable by
construction. **The router-prompt agreement filter cannot work here**, since
`routeGrammar` has no clarify value and a generated line always agrees with
itself. A gemma judge replaces it. The first judge called 24 of 40
answerable rows underspecified, because it judged against a generic
assistant rather than against Maven's contract.
**Mood is cut, not deferred.** The enum describes her own reply state, not the
speaker's emotion, and no dataset maps onto it.
@@ -417,7 +442,19 @@ queries exactly as it did.
That is the safety argument and it is not negotiable. The table's order is
load-bearing. Every comment on it argues a reason between two sources, and above all
it carries "the owner's data first, then the world". Naming `SourceWorld` does not
send the turn outside. His notes, his facts and the personal boundary still run first.
send the turn outside on its own. His notes and his facts still run first, because
they look rather than guess.
**The personal boundary is the one exception and it is deliberate.** It guesses,
so naming `SourceWorld` drops it. That is what stops it answering "кто такой
Линус Торвальдс?" with "не нашла у тебя такой записи", which it did on
2026-08-07. The cost is that a destination a model wrote can now take the
boundary off a turn. A question about him that the model calls `world` reaches
SearXNG, where today the boundary stops it. Only the utterance leaves the box,
never his notes or history, so this widens what is asked and not what is sent.
`TestNamingRecallKeepsTheBoundary` pins the other half: naming `SourceRecall`
keeps the boundary in front of the world. Whether a model may drop it at all is
the owner's call and has not been made.
What comes out is only the sources that **guess**. Those decide a turn is theirs by
cosine against frozen seeds, then answer whatever they claimed. They hold no table
@@ -0,0 +1,123 @@
# A clarify head, and a confidence that is not a hardcode
Measured 2026-08-08 on workpc, the same day and the same fixtures as
`2026-08-08-routing-heads-two-head.md` and `2026-08-08-slot-head-three-head.md`.
New script `gen_clarify.py`, new fixture `eval_fixture_clarify.jsonl`.
## A softmax has no clarify class
That sentence closed the two-head measurement. It is why the head's fixture was
88 cases and not 96. The eight `want_clarify` cases sat outside every number
measured, and the head had no way to produce the answer they wanted.
A fourth head is the answer. Clarify is not a value of intent. It is a second
question asked of the same pooled vector: can Maven act on this at all.
## The corpus had one class
Every row in `train_heads_slots.jsonl` was generated FOR an intent or a
destination. So every row is answerable by construction. A head trained on that
alone sees one class and learns to say yes.
`gen_clarify.py` makes the other class. Five shapes, ten topics. The shapes are
the gate's own reasons in `gateLLMDecision` plus the two the fixture carries:
bare noun, bare verb, demonstrative, deictic time, dangling reference.
**The agreement filter that worked for destination cannot work here.**
`routeGrammar` has no clarify value. So the router always names an intent, and
any generated line always agrees with itself. The second pass is a judge
instead. Gemma is asked, without seeing the label, whether Maven would have to
ask a question back.
## The first judge was worthless and the second was measured
The first judge said "needs clarify" on 24 of 40 plainly answerable corpus
rows. It flagged `запиши что я пообедал` and `Покажи расписание поездов на
вечер`. It was judging against a generic assistant, one that asks "where?"
about lunch. Maven writes that note.
Rewriting it to state what she can already do took false positives to 16 of 60.
It also catches all eight fixture clarifies. So the judge discriminates.
On the generated pile it removed 7 of 306, a 97.7% keep rate. That is not the
judge failing. The generator is aimed at underspecified lines, so there is
little for a filter to catch. The 27% false-positive rate is the number to
quote, and it is label noise on the positive class.
**Both passes are gemma-4-12b.** Generation and judging. So the corpus is
gemma's opinion of what is underspecified, and the head distills that opinion.
What keeps it honest is the fixture. Those eight cases were written by the owner
and gemma never saw them.
299 rows kept, against 3604 answerable. The positive class carries `intent:
null`, so it costs the intent head nothing.
## Result
Three seeds, 24 epochs, epoch still chosen on the intent dev slice.
| | two heads | three heads | four heads |
|---|---|---|---|
| intent mean | 93.6% | 92.8% | 91.7% |
| destination mean | 80.8% | 82.8% | 79.8% |
| slot span F1 mean | — | 72.4% | 68.3% |
| clarify caught | — | — | 7.0 of 8 |
| false clarifies | — | — | 2.3 of 88 |
**The fourth head is not free the way the third was.** Intent, destination and
slot F1 all move down. The drop is one to four points, and the seed spread is
wide enough to contain it. Seed 2 scores intent 94.3% and destination 84.8%, both above
every three-head seed. Read the drop as unproven rather than as absent.
Accuracy is the wrong number for this head and is reported for completeness at
95.8% to 96.9%. Eight of ninety-six cases are positive, so a head that never
asks scores 91.7%. Recall on those eight is the number.
Compare it to what ships. The cascade today misses 1 clarify and produces 2
false ones. The head catches 7 of 8 and produces 2.3 false ones. That is
parity, from a 118M encoder with no rules in front of it.
The saved checkpoint is seed 2 at epoch 10. Intent 94.3%, destination 28/33,
slot F1 73.6%, clarify 7 of 8 with 3 false. `heads.pt` carries four state dicts.
## What it gets wrong is consistent across seeds
`поужинал` is a false clarify on all three seeds. That utterance is already
recorded as a real defect. `thinSingleToken` was narrowed on 2026-08-01 to spare
a token carrying a Russian verb ending. One word is routinely a whole sentence
in Russian. The head relearned the mistake the rule was narrowed to
fix.
`что дальше?` is a false clarify on two seeds. That one is a disagreement rather
than an error. The utterance is underspecified, and V-498 decided stage 0 claims
it for the calendar on purpose.
`ну это` is missed on two seeds. `amb-003` is the shortest case in the fixture
and the generated demonstratives are longer.
## Confidence
`Confidence: 1.0` was a hardcode in `llmrouter.go`, so a correct low confidence
could not exist. Max softmax over the intent head is the replacement. It is only worth reading
if it is lower where the head is wrong.
It is. Mean 0.851 where the head is right against 0.604 where it is wrong. It
ranks a right case above a wrong one in 83.4% of pairs.
So there are two signals now and they are not the same signal. Confidence says
the head is unsure which intent this is. The clarify head says the utterance
does not carry enough to act on. A confident wrong route and an honest "I cannot
tell" are different failures, and one number cannot report both.
## What this does not measure
The same gap as every head run. **Nothing of this runs in Go.** Four heads
instead of three does not change that.
There is no threshold. Both signals are reported as raw numbers. Turning either
into a gate needs a decision about where to cut, and that trades false clarifies
against wrong acts. The fixture has 8 positives, which is too few to fit a
threshold on.
The 299 generated rows have no held-out slice of their own. Clarify is scored on
the fixture alone.
+4 -1
View File
@@ -30,7 +30,7 @@ present when a re-run reaches day 1.
| p50 | 1.5s | 1.6s |
| p95 | 8.0s | 7.1s |
| max | 12.3s | 33.7s |
| errors | 0 | 0 |
| transport errors | 0 | 0 |
| string in the reply | turns |
|---|---|
@@ -42,6 +42,9 @@ present when a re-run reaches day 1.
| `пока не умею` | 5 |
| `Когда?` | 3 |
**Zero transport errors is not zero wrong answers.** It counts turns that
failed to return a reply, and none did. Every quality number is below.
Those seven strings appear in 49 of 140 turns. Some turns carry two, because a
parked clarify appends to whatever else was said.