From 1d48755d12fbef00a3637fcb762300569d9d9808 Mon Sep 17 00:00:00 2001 From: kami Date: Fri, 31 Jul 2026 11:38:18 +0400 Subject: [PATCH] Record the recall numbers after the embedder swap recall@1 60% to 72%, answered 48% to 72%, latency 3x better. But false recall went 1/5 to 5/5: e5 packs every score into a narrow high band, so the 0.55 gate now admits everything. Left the gate alone as instructed. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ --- RECALL-EVAL-31-07-2026.md | 60 +++++++++++++++++++++++++++++++++++++++ 1 file changed, 60 insertions(+) diff --git a/RECALL-EVAL-31-07-2026.md b/RECALL-EVAL-31-07-2026.md index 45852b9..9330f81 100644 --- a/RECALL-EVAL-31-07-2026.md +++ b/RECALL-EVAL-31-07-2026.md @@ -83,6 +83,66 @@ sqlite-backed `store.MemoryStore` and `memory.InMemoryStore` identically — bot (`internal/store/memory.go:64`) at ~150µs over 42 rows against a ~59ms query embed. An ANN index is not the problem to solve. +## Re-measured after the embedder swap — 31-07-2026, later the same day + +Changed: `models/embedder/` is now **multilingual-e5-small** (quantized, 118MB), with `query: ` in +front of a question and `passage: ` in front of a stored note (Vikunja #371). `deploy/mavend.json` +and `make download-embedder` now name the same file, and it is the quantized one — that is what the +column below measures (Vikunja #372). Everything else is unchanged: same fixture, same store, same +0.55 gate. The old column is the baseline and is left as it was. + +| | recall+onnx, MiniLM (baseline) | recall+onnx, e5-small (new) | +|---|---|---| +| **recall@1** | 60.0% (15/25) | **72.0% (18/25)** | +| recall@3 | 80.0% (20/25) | 84.0% (21/25) | +| **answered after the 0.55 gate** | 48.0% (12/25) | **72.0% (18/25)** | +| wrong note on top / tie on top | 10 / 0 | 7 / 0 | +| ranked first, then silenced by the gate | 3 | 0 | +| **false recall** | 1/5 (20%) | **5/5 (100%)** | +| top-1 score when right, min / median | 0.559 / 0.678 | 0.791 / 0.857 | +| top-1 when it must stay silent, median / max | 0.470 / 0.567 | 0.815 / 0.835 | +| RU / EN / `hard` cases passed | 13/24 / 3/6 / 2/11 | 14/24 / 4/6 / 5/11 | +| latency p50 / p95 / max | 59ms / 148ms / 194ms | 18ms / 37ms / 49ms | + +### What moved + +Ranking got better and got faster. Half the previously-unwinnable `hard` cases now pass (2/11 → +5/11), the guitar note no longer beats the docker-logs note, and the gate stops silencing notes that +already ranked first. The quantized e5 is also ~3x quicker than the fp32 MiniLM it replaces. + +### What got worse: the gate is now a no-op + +e5 packs every cosine into a narrow high band. Right-note scores start at 0.791; must-stay-silent +scores reach 0.835. **The distributions still overlap, and now they overlap above the gate**, so +0.55 admits everything and false recall goes from 1/5 to 5/5. The sweep: + +``` +gate 0.50–0.70: answered 18/25 (72%) false recall 5/5 +gate 0.80: answered 17/25 (68%) false recall 4/5 +gate 0.90: answered 0/25 ( 0%) false recall 0/5 +``` + +There is no value that keeps real recall and rejects made-up questions — same conclusion as before, +now with a wider band and no room at all. `query_min_score` was left at 0.55 as instructed. **The +recommendation is to leave it there and stop tuning it**: any number under ~0.79 is a no-op and +anything above starts cutting real recall long before it stops the false ones. The fix is a margin +gate (`top1 − top2 > δ`), next-steps item 3, which is now the top item. + +### The prefixes did not do the work + +A control run with both prefixes set to the empty string scored the **same** recall@1 (72%), a +slightly better recall@3 (88%) and the same 5/5 false recall. So on this fixture the gain comes from +the model, not from the `query:` / `passage:` split. The prefixes are kept because they are how e5 +was trained and the split is the right shape for the read path, but they are not worth defending on +this evidence — a bigger fixture may say otherwise. + +### Stored vectors from the old model are now junk + +Cosine between a MiniLM vector and an e5 vector means nothing. Every row already in `notes` and in +the vector memory table was written by the old model, so after this deploy they will score as noise +against a new query. A live database needs every note and fact re-embedded before recall works at +all. Filed as its own task. + ## Next steps — ordered by value-to-risk; nothing here is a decision 1. **Swap the embedder to `multilingual-e5-small` with `query:`/`passage:` prefixes.** One config