docs: session 1 and 2 results, and the classifier baseline was wrong (V-459)
Ran sessions 1 and 2 on the live box. Session 1 steps 1 and 3-6 pass. Steps 2 and 7-9 need a person at the box. POST /api/chat is drivable with form encoding and a cookie jar, so the text half needs no browser. Session 2 confirms the deploy matches the bench at 72.7% full accuracy, and contradicts two recorded numbers. The classifier scores 68.8% at p50 16.6us, not 36.8% at 31ms. Router latency measured under contention again. Also: 319 item 2 point 2 closes on the recall margin sweep, CheckFeminine has a false positive on second-person masculine verbs, and the wake path cannot be checked because mavwaked and mavenclient are deployed nowhere.
This commit is contained in:
@@ -127,11 +127,13 @@ on in deploy** — this section used to say it was wired `nil`, which stopped be
|
||||
Cascade order: `stage0.go` exact-match fast-path → LLM router (when non-nil) → classifier
|
||||
fallback. Any LLM error falls through to the classifier so a turn never breaks on the model.
|
||||
|
||||
Measured on the 77-case RU fixture (`docs/evals/2026-07-31-model-bakeoff.md`): the classifier scores
|
||||
36.8% full accuracy at p50 31ms; Qwen3-1.7B scores 67.5% intent-only / 72.7% through the
|
||||
cascade at p50 ≈825ms. Accuracy roughly doubled, latency is ~27× worse, and that trade was
|
||||
accepted deliberately. **The ≈2.7s figure that stood here until 2026-08-02 was contention,
|
||||
not the model.** See `docs/evals/2026-07-31-routing.md` line 61, which measures the LLM router at
|
||||
Measured on the 77-case RU fixture. **Re-measured 2026-08-02: the classifier scores 68.8%
|
||||
full accuracy at p50 16.6µs**, not the 36.8% at p50 31ms that stood here from
|
||||
`docs/evals/2026-07-31-model-bakeoff.md`. That older figure predates the stage 0 rules and the
|
||||
seed additions, both of which now score inside the classifier baseline. Qwen3-1.7B scores
|
||||
77.9% intent-only / 72.7% through the cascade. So the router buys about 4 points of accuracy,
|
||||
not a doubling, and the trade is worth re-arguing rather than assuming. **The ≈2.7s figure
|
||||
that stood here until 2026-08-02 was contention, not the model.** See `docs/evals/2026-07-31-routing.md` line 61, which measures the LLM router at
|
||||
p50 825ms / p95 1.2s / max 3.0s and the full cascade at p50 0.80-1.04s. Do not plan latency
|
||||
work off the bakeoff table. `Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
|
||||
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
|
||||
|
||||
Reference in New Issue
Block a user