Record the recall numbers after the embedder swap
recall@1 60% to 72%, answered 48% to 72%, latency 3x better. But false recall went 1/5 to 5/5: e5 packs every score into a narrow high band, so the 0.55 gate now admits everything. Left the gate alone as instructed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
This commit is contained in:
@@ -83,6 +83,66 @@ sqlite-backed `store.MemoryStore` and `memory.InMemoryStore` identically — bot
|
||||
(`internal/store/memory.go:64`) at ~150µs over 42 rows against a ~59ms query embed. An ANN index is
|
||||
not the problem to solve.
|
||||
|
||||
## Re-measured after the embedder swap — 31-07-2026, later the same day
|
||||
|
||||
Changed: `models/embedder/` is now **multilingual-e5-small** (quantized, 118MB), with `query: ` in
|
||||
front of a question and `passage: ` in front of a stored note (Vikunja #371). `deploy/mavend.json`
|
||||
and `make download-embedder` now name the same file, and it is the quantized one — that is what the
|
||||
column below measures (Vikunja #372). Everything else is unchanged: same fixture, same store, same
|
||||
0.55 gate. The old column is the baseline and is left as it was.
|
||||
|
||||
| | recall+onnx, MiniLM (baseline) | recall+onnx, e5-small (new) |
|
||||
|---|---|---|
|
||||
| **recall@1** | 60.0% (15/25) | **72.0% (18/25)** |
|
||||
| recall@3 | 80.0% (20/25) | 84.0% (21/25) |
|
||||
| **answered after the 0.55 gate** | 48.0% (12/25) | **72.0% (18/25)** |
|
||||
| wrong note on top / tie on top | 10 / 0 | 7 / 0 |
|
||||
| ranked first, then silenced by the gate | 3 | 0 |
|
||||
| **false recall** | 1/5 (20%) | **5/5 (100%)** |
|
||||
| top-1 score when right, min / median | 0.559 / 0.678 | 0.791 / 0.857 |
|
||||
| top-1 when it must stay silent, median / max | 0.470 / 0.567 | 0.815 / 0.835 |
|
||||
| RU / EN / `hard` cases passed | 13/24 / 3/6 / 2/11 | 14/24 / 4/6 / 5/11 |
|
||||
| latency p50 / p95 / max | 59ms / 148ms / 194ms | 18ms / 37ms / 49ms |
|
||||
|
||||
### What moved
|
||||
|
||||
Ranking got better and got faster. Half the previously-unwinnable `hard` cases now pass (2/11 →
|
||||
5/11), the guitar note no longer beats the docker-logs note, and the gate stops silencing notes that
|
||||
already ranked first. The quantized e5 is also ~3x quicker than the fp32 MiniLM it replaces.
|
||||
|
||||
### What got worse: the gate is now a no-op
|
||||
|
||||
e5 packs every cosine into a narrow high band. Right-note scores start at 0.791; must-stay-silent
|
||||
scores reach 0.835. **The distributions still overlap, and now they overlap above the gate**, so
|
||||
0.55 admits everything and false recall goes from 1/5 to 5/5. The sweep:
|
||||
|
||||
```
|
||||
gate 0.50–0.70: answered 18/25 (72%) false recall 5/5
|
||||
gate 0.80: answered 17/25 (68%) false recall 4/5
|
||||
gate 0.90: answered 0/25 ( 0%) false recall 0/5
|
||||
```
|
||||
|
||||
There is no value that keeps real recall and rejects made-up questions — same conclusion as before,
|
||||
now with a wider band and no room at all. `query_min_score` was left at 0.55 as instructed. **The
|
||||
recommendation is to leave it there and stop tuning it**: any number under ~0.79 is a no-op and
|
||||
anything above starts cutting real recall long before it stops the false ones. The fix is a margin
|
||||
gate (`top1 − top2 > δ`), next-steps item 3, which is now the top item.
|
||||
|
||||
### The prefixes did not do the work
|
||||
|
||||
A control run with both prefixes set to the empty string scored the **same** recall@1 (72%), a
|
||||
slightly better recall@3 (88%) and the same 5/5 false recall. So on this fixture the gain comes from
|
||||
the model, not from the `query:` / `passage:` split. The prefixes are kept because they are how e5
|
||||
was trained and the split is the right shape for the read path, but they are not worth defending on
|
||||
this evidence — a bigger fixture may say otherwise.
|
||||
|
||||
### Stored vectors from the old model are now junk
|
||||
|
||||
Cosine between a MiniLM vector and an e5 vector means nothing. Every row already in `notes` and in
|
||||
the vector memory table was written by the old model, so after this deploy they will score as noise
|
||||
against a new query. A live database needs every note and fact re-embedded before recall works at
|
||||
all. Filed as its own task.
|
||||
|
||||
## Next steps — ordered by value-to-risk; nothing here is a decision
|
||||
|
||||
1. **Swap the embedder to `multilingual-e5-small` with `query:`/`passage:` prefixes.** One config
|
||||
|
||||
Reference in New Issue
Block a user