diff --git a/RECALL-EVAL-31-07-2026.md b/RECALL-EVAL-31-07-2026.md index 9330f81..39fb389 100644 --- a/RECALL-EVAL-31-07-2026.md +++ b/RECALL-EVAL-31-07-2026.md @@ -143,14 +143,97 @@ the vector memory table was written by the old model, so after this deploy they against a new query. A live database needs every note and fact re-embedded before recall works at all. Filed as its own task. +## Margin gate — 31-07-2026, third run + +Next-steps item 3, done. The absolute gate is replaced by a **margin gate**: answer only when the +top hit beats the runner-up by more than delta (`top1 − top2 > δ`). Same fixture, same e5 embedder, +same store as the run above. `internal/memory/gate.go` holds the check; both read paths call it +(`cmd/mavend/recall.go` and the notes-RAG branch in `voice.go`). New knob `voice.query_min_margin` +in `deploy/mavend.json`, default 0.008. + +### Why the absolute gate could not work, in one line of data + +The harness now prints the margin distributions, and they barely overlap where the raw scores +overlap completely: + +| | top-1 score | margin (top1 − top2) | +|---|---|---| +| right note first (n=18) | min 0.810, median 0.862, max 0.890 | min 0.001, median 0.029, max 0.053 | +| must stay silent (n=5) | min 0.795, median 0.815, max 0.835 | min 0.000, median 0.002, **max 0.019** | + +Four of the five must-be-silent cases have a margin at or under 0.002 — when there is nothing to +recall, e5 finds several notes equally close and no clear winner. That is the signal the absolute +score throws away. + +### The delta sweep + +Absolute gate held at 0.55 throughout. + +``` +delta 0.000: answered 18/25 (72%) false recall 5/5 +delta 0.002: answered 17/25 (68%) false recall 3/5 +delta 0.005: answered 17/25 (68%) false recall 2/5 +delta 0.008: answered 17/25 (68%) false recall 1/5 <- chosen +delta 0.010: answered 15/25 (60%) false recall 1/5 +delta 0.012: answered 14/25 (56%) false recall 1/5 +delta 0.015: answered 12/25 (48%) false recall 1/5 +delta 0.020: answered 11/25 (44%) false recall 0/5 +delta 0.025: answered 9/25 (36%) false recall 0/5 +delta 0.030: answered 8/25 (32%) false recall 0/5 +delta 0.040: answered 4/25 (16%) false recall 0/5 +delta 0.050: answered 2/25 ( 8%) false recall 0/5 +delta 0.060: answered 0/25 ( 0%) false recall 0/5 +``` + +### Chosen: δ = 0.008 + +It is the best point on the frontier, not a taste call. **0.008 dominates 0.010, 0.012 and 0.015 +outright** — same 1/5 false recall, 8 to 20 points more real recall. Everything below it buys recall +back only by admitting more false recalls (0.005 → 2/5, 0.002 → 3/5). The next real improvement is +0.020 at 0/5 false, and it costs 24 points of recall to get there. + +The brief's bar was "recall above 60% with false recall at 1/5 or better". 0.008 clears it with room: +68% and 1/5. + +### Before / after + +| | absolute gate 0.55 (previous) | margin gate δ=0.008 | +|---|---|---| +| recall@1 (ranking, ungated) | 72.0% (18/25) | 72.0% (18/25) — unchanged, the gate does not rank | +| **answered after the gate** | 72.0% (18/25) | **68.0% (17/25)** | +| **false recall** | **5/5 (100%)** | **1/5 (20%)** | +| fixture cases passed | 18/30 | **21/30** | + +Four false recalls removed for one real answer. That is the trade the spec asks for — she is not a +guesser-of-truth. The one survivor is `en-pref-025` ("should i be offered wine"), which recalls a +filler note at 0.796 with a 0.019 margin: the widest silent-case margin in the fixture, and it sits +inside the real-recall range, so no delta removes it without taking real answers with it. + +### Does the absolute cutoff still earn its keep? Marginally — kept + +On this fixture with e5 it is a **no-op**: the lowest right-note score is 0.791, so 0.55 rejects +nothing the margin does not already reject. It is kept for two reasons, neither glamorous. It still +does real work for the hash embedder (its own sweep shows answers dropping from 16% to 0% between +0.30 and 0.50), and it is the only thing standing between the user and a reply built from a store +where everything is far away but one row happens to be a little less far — a near-empty database, or +the stale-vector case below. Cheap insurance, no measured cost. If a later embedder makes it bite, +the sweep is one command. + +### Caveat on the numbers + +Five must-be-silent cases is a thin basis for a 4-point decision. 1/5 and 2/5 differ by one case. +The shape of the frontier is trustworthy — margins separate, absolute scores do not — but δ=0.008 +itself should be re-read off a bigger fixture (next-steps item 6) before anyone defends the third +decimal. + ## Next steps — ordered by value-to-risk; nothing here is a decision 1. **Swap the embedder to `multilingual-e5-small` with `query:`/`passage:` prefixes.** One config change plus a prefix in `onnxembedder.go`, re-measurable in one command. 2. **Re-run `make eval-recall`, then set the gate from the sweep** — not before. Any `query_min_score` picked against today's embedder describes a model on its way out. -3. **Replace the absolute-score gate with a margin gate** (`top1 − top2 > δ`) — as the routing eval - concluded, absolute cosine cannot see a flat distribution. +3. ~~**Replace the absolute-score gate with a margin gate**~~ — done, see the section above. + δ=0.008, false recall 5/5 → 1/5. 4. **Delete or repair the dead `memStore` branch** at `voice.go:776` — search before the gate, gate it separately, or restrict it to facts and say so. 5. **Add a mild time decay to ranking** — the newest statement of a preference is the true one.