plan: the claim unit and its four bands, measured (V-565)
Measures what each claimant on an utterance reports across the 91-case RU fixture, then argues an ordinal band set from that distribution. The classifier's cosine is flat against correctness: 62% correct below its median and 62% above, over a spread only 0.083 wide, with every case above the 0.55 gate. Its top-two margin is p50 0.009 and never reaches 0.03. So a calibrated float is not cheaply available and the ledger's assumption holds. Stage 0 is 20/20 on the cases it claims and asserts 1.0 for all of them. The LLM router emits two values, and the lower one is a self-veto with a reason flattened into a number.
This commit is contained in:
@@ -0,0 +1,135 @@
|
||||
# Plan: dialogue arbitration, one channel and many claimants
|
||||
|
||||
Umbrella V-558. This file collects the design for its children.
|
||||
|
||||
Last verified: 06-08-2026 @ b6305f1
|
||||
|
||||
## A common unit for claims on an utterance (V-565)
|
||||
|
||||
**Verdict: four ordinal bands, and the band is the tie-break rather than the decision.
|
||||
Coverage decides first.** The measurement below says no claimant Maven has today can produce
|
||||
a graded confidence. A float would be an invention either way. What is available is the KIND
|
||||
of evidence a claimant holds, and there are exactly four kinds.
|
||||
|
||||
### What the claimants report today
|
||||
|
||||
Measured 06-08-2026 on the 91-case RU fixture (`internal/router/eval`), through the deployed
|
||||
cascade with the quantized multilingual-e5-small embedder. The harness is
|
||||
`TestONNXClaimConfidenceDistribution` and `TestStage0Contention` in
|
||||
`internal/router/eval/claims_test.go`. Correct means the right intent, or a refusal where the
|
||||
fixture wants one. Slots are excluded, because a slot miss is a parser question and would
|
||||
blur what the number is being asked to predict.
|
||||
|
||||
| Claimant | Values it can emit | Distribution on the fixture | Correct |
|
||||
|---|---|---|---|
|
||||
| Stage 0 grammars, 21 of them | `1.0`, always | claimed 20 of 91 cases | 20/20 (100%) |
|
||||
| Classifier, cosine | continuous in principle | observed range 0.859 to 0.942 over 71 cases | 44/71 (62%) |
|
||||
| LLM router | `1.0` or `0.3`, nothing between | not run here, no llama-server | see below |
|
||||
| Query sources, 22 of them | a bool | not routed by the fixture | n/a |
|
||||
| Stateful four | nothing at all | n/a | n/a |
|
||||
|
||||
Four findings, and each one constrains the band set.
|
||||
|
||||
**The classifier's cosine carries no signal about correctness.** It scores 62% below the
|
||||
median and 62% above it. That is 13/21 in 0.8 to 0.9, and 31/50 in 0.9 to 1.0. The spread is
|
||||
0.083 wide. Every case sits above the 0.55 threshold, so the gate never fires here. A number
|
||||
flat against correctness, which never crosses its own gate, is not a confidence.
|
||||
|
||||
**Nor does the margin between its top two intents.** Top1 minus top2 is min 0.000, p50
|
||||
0.009, max 0.025. Sixty-eight of the 71 classified cases sit under 0.02 and score 60%. Three
|
||||
clear 0.02 and score 3/3, which is a sample of three. So the ledger's question is answered:
|
||||
a calibrated float is NOT cheaply available from the classifier alone. Nearest-centroid over
|
||||
frozen seeds ranks intents, and the ranking is decided in the third decimal place. It can say
|
||||
which intent is nearest. It cannot say how near.
|
||||
|
||||
**Stage 0 asserts 1.0 by fiat, and on this fixture the fiat is right.** Twenty of twenty.
|
||||
That is not evidence that a hand-written anchored pattern is always right. It is evidence
|
||||
that anchored and nearest are different kinds of claim, and must not share a scale. The gap
|
||||
is 100% against 62% on the same 91 utterances.
|
||||
|
||||
**Stage 0 contention is rarer than the list order suggests.** Exactly one case of 91 draws
|
||||
two grammars. That is `ru-query-019`, where `calendar-query` and `agenda-query` both match,
|
||||
and `calendar-query` wins because it is earlier in `buildRouter`. Both would route
|
||||
`IntentQuery`, so the ordering costs nothing there. The finding is not that ordering is
|
||||
harmless. It is that the fixture barely exercises what V-558 is about. Part of what a claim
|
||||
object buys is making the contention countable.
|
||||
|
||||
**The LLM router emits two values, and one of them is not a confidence.** `llmFullConfidence`
|
||||
is 1.0 and `llmThinConfidence` is 0.3. `gateLLMDecision` moves a decision to 0.3 through
|
||||
three named arms. A fact with no key, an act with no allowlisted fn, a reminder with no
|
||||
subject. Each is a self-veto with a reason, flattened into a number that then loses the
|
||||
reason. Both values are meaningful only against `config.DefaultRouterThreshold`. 0.3 is below
|
||||
0.55 and 1.0 is above it, and nothing anywhere reads any other property of either.
|
||||
|
||||
### The band set
|
||||
|
||||
Four bands, ordinal, highest first. They name the kind of evidence, because that is the one
|
||||
thing every claimant can report without inventing it.
|
||||
|
||||
**`BandAnchored`.** A literal pattern anchored in the utterance matched, and the matched span
|
||||
is what decides the intent. Stage 0 grammars and query-source matchers. The claimant is
|
||||
certain about the shape of the sentence. That is not the same as being certain about the
|
||||
answer. Measured 20/20.
|
||||
|
||||
**`BandStructural`.** A claimant read the whole sentence and produced a complete route. Every
|
||||
slot the intent requires is filled. The LLM router at `llmFullConfidence` sits here, and so
|
||||
does a stateful claimant holding a pending question. Not anchored, because nothing in the
|
||||
utterance is pointed at.
|
||||
|
||||
**`BandNearest`.** The claim rests only on resemblance to something else. No anchor in the
|
||||
utterance, no structural check behind it. The classifier. One band rather than a graded
|
||||
scale, and the measurement is the argument. 62% at both ends of the cosine range, and a
|
||||
top-two margin that never reaches 0.03.
|
||||
|
||||
**`BandVetoed`.** The claimant will take the turn only if nobody else will, and says why it
|
||||
should not. The three arms of `gateLLMDecision` land here with their reason preserved. A
|
||||
vetoed claim is still a claim. Maven asking "о чём напомнить?" beats silence.
|
||||
|
||||
There is no fifth band, and that is a measurement result rather than a preference. No
|
||||
claimant in the cascade today can report what a fifth band would carry. V-546 lands a softmax
|
||||
head whose max probability is a calibrated number. That one gets read as a number, not
|
||||
squeezed into these four.
|
||||
|
||||
### Coverage decides before the band does
|
||||
|
||||
The band is the tie-break. The first question is how much of the utterance a claim explains,
|
||||
and that is `Consumed` against `Unexplained` on the claim object. Two reasons.
|
||||
|
||||
It is the fix for the failure that opened V-558. "какая сейчас погода в Риме?" arrived while
|
||||
a reminder was pending. The pending claimant ate the whole utterance as a time answer while
|
||||
explaining none of it. Not "погода", not "Риме", not the question mark. A weather claim
|
||||
explains all of it. Coverage-first arbitration prefers the weather claim without knowing that
|
||||
a pending reminder is less trustworthy than a grammar. The pending question then survives to
|
||||
be asked again.
|
||||
|
||||
It also keeps the stateful four out of the top slot without special-casing them. They sit at
|
||||
`BandStructural`, below any anchored claim. That is the whole V-558 complaint about the
|
||||
highest-priority claimants being the least informed, expressed as one rule.
|
||||
|
||||
### The claim object
|
||||
|
||||
```go
|
||||
type Claim struct {
|
||||
Claimant string // who wants the turn
|
||||
Intent string // plain string: internal/dialogue must not import internal/router
|
||||
Filled []string // the slots this claim would fill
|
||||
Consumed []string // utterance tokens this claim explains
|
||||
Unexplained []string // the rest, in order
|
||||
Band Band
|
||||
Veto string // why this claim should NOT win, empty when there is none
|
||||
}
|
||||
```
|
||||
|
||||
`Intent` is a plain `string` rather than `router.Intent` on purpose. `internal/dialogue` must
|
||||
not import `internal/router`, so the claim package must not either, and a shared string costs
|
||||
one conversion at each edge.
|
||||
|
||||
`Unexplained` is carried rather than derived at read time. A claimant can then decline to
|
||||
explain a span it did match.
|
||||
|
||||
### What this task does not do
|
||||
|
||||
`router.Decision.Confidence` stays and keeps its float. `r.threshold` and `gateLLMDecision`
|
||||
read it, and the classifier is the failure floor. A rewire that broke either would trade a
|
||||
measured floor for an unmeasured design. V-565 lands the type and the builder beside the
|
||||
existing path. The arbiter that reads claims is V-560.
|
||||
Reference in New Issue
Block a user