Files
Maven/docs/plans/19-dialogue-arbitration.md
claude da2d11dab6 plan: the claim unit and its four bands, measured (V-565)
Measures what each claimant on an utterance reports across the 91-case RU
fixture, then argues an ordinal band set from that distribution.

The classifier's cosine is flat against correctness: 62% correct below its
median and 62% above, over a spread only 0.083 wide, with every case above
the 0.55 gate. Its top-two margin is p50 0.009 and never reaches 0.03. So a
calibrated float is not cheaply available and the ledger's assumption holds.

Stage 0 is 20/20 on the cases it claims and asserts 1.0 for all of them. The
LLM router emits two values, and the lower one is a self-veto with a reason
flattened into a number.
2026-08-06 00:46:50 +04:00

7.4 KiB
Raw Permalink Blame History

Plan: dialogue arbitration, one channel and many claimants

Umbrella V-558. This file collects the design for its children.

Last verified: 06-08-2026 @ b6305f1

A common unit for claims on an utterance (V-565)

Verdict: four ordinal bands, and the band is the tie-break rather than the decision. Coverage decides first. The measurement below says no claimant Maven has today can produce a graded confidence. A float would be an invention either way. What is available is the KIND of evidence a claimant holds, and there are exactly four kinds.

What the claimants report today

Measured 06-08-2026 on the 91-case RU fixture (internal/router/eval), through the deployed cascade with the quantized multilingual-e5-small embedder. The harness is TestONNXClaimConfidenceDistribution and TestStage0Contention in internal/router/eval/claims_test.go. Correct means the right intent, or a refusal where the fixture wants one. Slots are excluded, because a slot miss is a parser question and would blur what the number is being asked to predict.

Claimant Values it can emit Distribution on the fixture Correct
Stage 0 grammars, 21 of them 1.0, always claimed 20 of 91 cases 20/20 (100%)
Classifier, cosine continuous in principle observed range 0.859 to 0.942 over 71 cases 44/71 (62%)
LLM router 1.0 or 0.3, nothing between not run here, no llama-server see below
Query sources, 22 of them a bool not routed by the fixture n/a
Stateful four nothing at all n/a n/a

Four findings, and each one constrains the band set.

The classifier's cosine carries no signal about correctness. It scores 62% below the median and 62% above it. That is 13/21 in 0.8 to 0.9, and 31/50 in 0.9 to 1.0. The spread is 0.083 wide. Every case sits above the 0.55 threshold, so the gate never fires here. A number flat against correctness, which never crosses its own gate, is not a confidence.

Nor does the margin between its top two intents. Top1 minus top2 is min 0.000, p50 0.009, max 0.025. Sixty-eight of the 71 classified cases sit under 0.02 and score 60%. Three clear 0.02 and score 3/3, which is a sample of three. So the ledger's question is answered: a calibrated float is NOT cheaply available from the classifier alone. Nearest-centroid over frozen seeds ranks intents, and the ranking is decided in the third decimal place. It can say which intent is nearest. It cannot say how near.

Stage 0 asserts 1.0 by fiat, and on this fixture the fiat is right. Twenty of twenty. That is not evidence that a hand-written anchored pattern is always right. It is evidence that anchored and nearest are different kinds of claim, and must not share a scale. The gap is 100% against 62% on the same 91 utterances.

Stage 0 contention is rarer than the list order suggests. Exactly one case of 91 draws two grammars. That is ru-query-019, where calendar-query and agenda-query both match, and calendar-query wins because it is earlier in buildRouter. Both would route IntentQuery, so the ordering costs nothing there. The finding is not that ordering is harmless. It is that the fixture barely exercises what V-558 is about. Part of what a claim object buys is making the contention countable.

The LLM router emits two values, and one of them is not a confidence. llmFullConfidence is 1.0 and llmThinConfidence is 0.3. gateLLMDecision moves a decision to 0.3 through three named arms. A fact with no key, an act with no allowlisted fn, a reminder with no subject. Each is a self-veto with a reason, flattened into a number that then loses the reason. Both values are meaningful only against config.DefaultRouterThreshold. 0.3 is below 0.55 and 1.0 is above it, and nothing anywhere reads any other property of either.

The band set

Four bands, ordinal, highest first. They name the kind of evidence, because that is the one thing every claimant can report without inventing it.

BandAnchored. A literal pattern anchored in the utterance matched, and the matched span is what decides the intent. Stage 0 grammars and query-source matchers. The claimant is certain about the shape of the sentence. That is not the same as being certain about the answer. Measured 20/20.

BandStructural. A claimant read the whole sentence and produced a complete route. Every slot the intent requires is filled. The LLM router at llmFullConfidence sits here, and so does a stateful claimant holding a pending question. Not anchored, because nothing in the utterance is pointed at.

BandNearest. The claim rests only on resemblance to something else. No anchor in the utterance, no structural check behind it. The classifier. One band rather than a graded scale, and the measurement is the argument. 62% at both ends of the cosine range, and a top-two margin that never reaches 0.03.

BandVetoed. The claimant will take the turn only if nobody else will, and says why it should not. The three arms of gateLLMDecision land here with their reason preserved. A vetoed claim is still a claim. Maven asking "о чём напомнить?" beats silence.

There is no fifth band, and that is a measurement result rather than a preference. No claimant in the cascade today can report what a fifth band would carry. V-546 lands a softmax head whose max probability is a calibrated number. That one gets read as a number, not squeezed into these four.

Coverage decides before the band does

The band is the tie-break. The first question is how much of the utterance a claim explains, and that is Consumed against Unexplained on the claim object. Two reasons.

It is the fix for the failure that opened V-558. "какая сейчас погода в Риме?" arrived while a reminder was pending. The pending claimant ate the whole utterance as a time answer while explaining none of it. Not "погода", not "Риме", not the question mark. A weather claim explains all of it. Coverage-first arbitration prefers the weather claim without knowing that a pending reminder is less trustworthy than a grammar. The pending question then survives to be asked again.

It also keeps the stateful four out of the top slot without special-casing them. They sit at BandStructural, below any anchored claim. That is the whole V-558 complaint about the highest-priority claimants being the least informed, expressed as one rule.

The claim object

type Claim struct {
    Claimant    string   // who wants the turn
    Intent      string   // plain string: internal/dialogue must not import internal/router
    Filled      []string // the slots this claim would fill
    Consumed    []string // utterance tokens this claim explains
    Unexplained []string // the rest, in order
    Band        Band
    Veto        string   // why this claim should NOT win, empty when there is none
}

Intent is a plain string rather than router.Intent on purpose. internal/dialogue must not import internal/router, so the claim package must not either, and a shared string costs one conversion at each edge.

Unexplained is carried rather than derived at read time. A claimant can then decline to explain a span it did match.

What this task does not do

router.Decision.Confidence stays and keeps its float. r.threshold and gateLLMDecision read it, and the classifier is the failure floor. A rewire that broke either would trade a measured floor for an unmeasured design. V-565 lands the type and the builder beside the existing path. The arbiter that reads claims is V-560.