Wire the heads between stage 0 and the resident model (V-664)
They run before the model because they are two orders of magnitude faster and score better on both halves of the route. They decline rather than clarify, so a declined turn carries on to the model and then the classifier, which is what a box with no weights file does on every turn. Nil heads are byte-for-byte the cascade that shipped before this. Measured on the 96-case fixture, classifier+ONNX either way: intent 76.0% -> 96.9% destination 36.4% -> 75.8% false clarify 0 -> 1 missed clarify 8 -> 1 p50 24.5ms -> 27.9ms That beats the gemma-4-12b cascade on both halves, 84.4% and 72.7%, at a twelfth of its 329ms. The four remaining destination misses are all calendar, which is the stage 0 trade V-660 flagged and the owner has not called yet. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
This commit is contained in:
@@ -14,12 +14,15 @@ import (
|
||||
"github.com/kami/maven/internal/decision"
|
||||
)
|
||||
|
||||
// The two routing engines, named as claimants. They are one stage and not two,
|
||||
// because only one of them ever runs: the classifier is reached when the model
|
||||
// is absent or errored, never alongside it.
|
||||
// The three routing engines, named as claimants. The model and the classifier
|
||||
// are one stage and not two, because only one of them ever runs: the classifier
|
||||
// is reached when the model is absent or errored, never alongside it. The heads
|
||||
// run before both and decline on low confidence, so they can appear beside
|
||||
// either one in a record.
|
||||
const (
|
||||
claimantLLM = "llm-router"
|
||||
claimantClassifier = "classifier"
|
||||
claimantHeads = "routing-heads"
|
||||
)
|
||||
|
||||
// thinReason names which arm of gateLLMDecision cut the confidence. The gate
|
||||
|
||||
Reference in New Issue
Block a user