98ee701e03
The memory pass ran only after the notes-only gate had already rejected the same note at the same score. Notes and facts share one vector index, so a note that failed there failed again — the branch could only ever return a fact. Now the memory pass runs first: one search over everything Maven remembers, one gate, and the memory that clearly matches best answers (a note gets phrased, a fact is read back). The notes-only pass stays behind it for notes the vector index does not hold. No threshold moved, so the set of questions answered is unchanged — only which memory answers them. Fixture gained two mixed note+fact cases, so the answerable count goes 25 -> 27: hash recall@1 36.0% -> 37.0% (ratchet 0.32 unchanged, comment updated), e5 recall@1 72.0% -> 70.4%, false recall still 1/5. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
251 lines
14 KiB
Markdown
251 lines
14 KiB
Markdown
# Note recall evaluation — 31-07-2026
|
||
|
||
The operator's goal is that Maven "memorize/note things … and know more about me/world". This
|
||
measures whether the note/recall path delivers that.
|
||
|
||
- Fixture + scorer: `internal/memory/recalleval/` (`ru_recall_v1.json`, 30 cases)
|
||
- Reproduce: `make eval-recall` — hash ratchet always, ONNX when `deps/` is present
|
||
- Commit: `43470ab` (harness)
|
||
|
||
Each case inserts its own 3 notes **plus 12 shared filler notes** into a fresh store, embeds the
|
||
query, takes the top 3 — the read path `cmd/mavend/voice.go` runs for `IntentQuery`. Filler is
|
||
load-bearing: with 3 notes and a top-3 search, recall@3 is 100% by construction. 25 answerable
|
||
cases (paraphrased queries, homelab and preference content, 9 with a plausible second note) and 5
|
||
that must recall **nothing**. `TestFixtureIsParaphrased` fails the build if a query shares over half
|
||
its words with its note; equal-score ties count as ties, not recall.
|
||
|
||
## Results
|
||
|
||
| | recall+hash (CI ratchet) | recall+onnx (deployed) |
|
||
|---|---|---|
|
||
| **recall@1** | 36.0% (9/25) | **60.0% (15/25)** |
|
||
| recall@3 | 76.0% (19/25) | 80.0% (20/25) |
|
||
| **answered after the 0.55 gate** | **0.0% (0/25)** | **48.0% (12/25)** |
|
||
| wrong note on top / tie on top | 9 / 7 | 10 / 0 |
|
||
| ranked first, then silenced by the gate | 9 | 3 |
|
||
| **false recall** | 0/5 | **1/5 (20%)** |
|
||
| top-1 score when right, min / median | n/a | 0.559 / 0.678 |
|
||
| top-1 when it must stay silent, median / max | 0.000 / 0.144 | 0.470 / **0.567** |
|
||
| RU / EN / `hard` cases passed | 4/24 / 1/6 / 0/11 | 13/24 / 3/6 / 2/11 |
|
||
| latency p50 / p95 / max | 49µs / 70µs | 59ms / 148ms / 194ms |
|
||
|
||
Never compare a hash-embedder number to an ONNX one — the hash floor is lexical and exists only so
|
||
CI has a deterministic ratchet with no model files.
|
||
|
||
## Findings
|
||
|
||
### 1. Real recall is 48%, not 60%
|
||
|
||
The right note ranks first 60% of the time, but the daemon only *says* it 48% of the time — three
|
||
more cases rank first and are then silenced by `voice.go:776`'s `queryMinScore`. **Roughly one
|
||
useful question in two gets "не знаю".** This is not a working memory yet.
|
||
|
||
### 2. The gate cannot separate a real recall from a false one — the distributions overlap
|
||
|
||
Right-note top-1 scores start at **0.559**. Must-stay-silent top-1 scores reach **0.567**. No
|
||
threshold keeps every real recall and rejects every false one. From the sweep: gate 0.50 → 13/25
|
||
answered, 1/5 false; **0.55 (default) → 12/25, 1/5**; **0.60 → 10/25, 0/5**; 0.70 → 5/25, 0/5. What
|
||
the data says about `DefaultQueryMinScore` (`internal/config/config.go:392`): **0.55 is
|
||
slightly too loose** — it admits one confident wrong answer ("как зовут сестру моего коллеги"
|
||
recalls "выучил пару аккордов на гитаре" at 0.567), which the spec ranks as worse than a gap. 0.60
|
||
silences all five and costs 8 points of real recall. Left alone as instructed; the overlap means
|
||
the threshold is the wrong dial anyway (finding 3).
|
||
|
||
### 3. Filler notes outrank the right answer — the model scores similarity, not relevance
|
||
|
||
`models/embedder/` is **paraphrase-multilingual-MiniLM-L12-v2** (`Makefile:119`), a *symmetric*
|
||
paraphrase model. It scores "do these sentences look alike", not "does this passage answer this
|
||
question", so question-shaped queries drift to whatever note is stylistically closest. "из-за чего
|
||
кончилось место" and "откуда берётся токен бота" both return `выучил пару аккордов на гитаре`
|
||
(0.730, 0.729); "как я восстановил конфиги" returns a bootloader note at 0.703 with the right note
|
||
not even in the top 3. An unrelated guitar note beating a homelab note at 0.73 is not a tuning
|
||
problem — an asymmetric retrieval model (`multilingual-e5-small`, with `query:` / `passage:`
|
||
prefixes) is the targeted fix, and it would move findings 1 and 2 together. Separately:
|
||
`deploy/mavend.json:39` loads a 470MB fp32 `model.onnx` while `make download-embedder` fetches
|
||
`model_quantized.onnx` — not the same file.
|
||
|
||
`hard` cases score **2/11**: every one is a query where the operator did not reuse his own words.
|
||
That is the normal case weeks later, and exactly what DESIGN.md's "recall when relevant" promises.
|
||
|
||
### 4. The memory-store recall branch is dead for notes
|
||
|
||
`voice.go:776` only reaches `h.memStore.Search` when the notes-RAG top score is already below
|
||
`queryMinScore`, and `bestRecall` (`cmd/mavend/recall.go:19`) then applies the **same** gate to the
|
||
same vector. A note is indexed in both places with the same embedding, so if it failed the gate in
|
||
`QueryNotes` it fails again here — the branch can only ever return a **fact**. Its comment calls it
|
||
"additive"; for notes it is not.
|
||
|
||
**Fixed (Vikunja #373).** The memory pass now runs *first*, as one search over notes and facts with
|
||
one gate, so whichever memory is clearly the best match answers — note or fact. The notes-only pass
|
||
stays behind it for notes the vector index does not hold. No threshold changed, so the set of
|
||
questions Maven answers is the same; only which memory answers them. The fixture gained two mixed
|
||
note+fact cases (`ru-mixed-031`, `ru-mixed-032`), which is why the counts below are out of 27
|
||
answerable cases and not 25: hash recall@1 36.0% (9/25) → 37.0% (10/27), e5 recall@1 72.0% (18/25) →
|
||
70.4% (19/27) with answered-after-gate 68.0% → 66.7% and false recall unchanged at 1/5.
|
||
|
||
### 5. Ranking has no recency or type signal, and the store is not the bottleneck
|
||
|
||
`internal/store/notes.go:67` sorts by cosine and uses `ts` only to break an exact float tie, which
|
||
never happens; `kind` never enters the ranking. Meanwhile `TestPersistentStoreScoresTheSame` scores
|
||
sqlite-backed `store.MemoryStore` and `memory.InMemoryStore` identically — both full-scan cosine
|
||
(`internal/store/memory.go:64`) at ~150µs over 42 rows against a ~59ms query embed. An ANN index is
|
||
not the problem to solve.
|
||
|
||
## Re-measured after the embedder swap — 31-07-2026, later the same day
|
||
|
||
Changed: `models/embedder/` is now **multilingual-e5-small** (quantized, 118MB), with `query: ` in
|
||
front of a question and `passage: ` in front of a stored note (Vikunja #371). `deploy/mavend.json`
|
||
and `make download-embedder` now name the same file, and it is the quantized one — that is what the
|
||
column below measures (Vikunja #372). Everything else is unchanged: same fixture, same store, same
|
||
0.55 gate. The old column is the baseline and is left as it was.
|
||
|
||
| | recall+onnx, MiniLM (baseline) | recall+onnx, e5-small (new) |
|
||
|---|---|---|
|
||
| **recall@1** | 60.0% (15/25) | **72.0% (18/25)** |
|
||
| recall@3 | 80.0% (20/25) | 84.0% (21/25) |
|
||
| **answered after the 0.55 gate** | 48.0% (12/25) | **72.0% (18/25)** |
|
||
| wrong note on top / tie on top | 10 / 0 | 7 / 0 |
|
||
| ranked first, then silenced by the gate | 3 | 0 |
|
||
| **false recall** | 1/5 (20%) | **5/5 (100%)** |
|
||
| top-1 score when right, min / median | 0.559 / 0.678 | 0.791 / 0.857 |
|
||
| top-1 when it must stay silent, median / max | 0.470 / 0.567 | 0.815 / 0.835 |
|
||
| RU / EN / `hard` cases passed | 13/24 / 3/6 / 2/11 | 14/24 / 4/6 / 5/11 |
|
||
| latency p50 / p95 / max | 59ms / 148ms / 194ms | 18ms / 37ms / 49ms |
|
||
|
||
### What moved
|
||
|
||
Ranking got better and got faster. Half the previously-unwinnable `hard` cases now pass (2/11 →
|
||
5/11), the guitar note no longer beats the docker-logs note, and the gate stops silencing notes that
|
||
already ranked first. The quantized e5 is also ~3x quicker than the fp32 MiniLM it replaces.
|
||
|
||
### What got worse: the gate is now a no-op
|
||
|
||
e5 packs every cosine into a narrow high band. Right-note scores start at 0.791; must-stay-silent
|
||
scores reach 0.835. **The distributions still overlap, and now they overlap above the gate**, so
|
||
0.55 admits everything and false recall goes from 1/5 to 5/5. The sweep:
|
||
|
||
```
|
||
gate 0.50–0.70: answered 18/25 (72%) false recall 5/5
|
||
gate 0.80: answered 17/25 (68%) false recall 4/5
|
||
gate 0.90: answered 0/25 ( 0%) false recall 0/5
|
||
```
|
||
|
||
There is no value that keeps real recall and rejects made-up questions — same conclusion as before,
|
||
now with a wider band and no room at all. `query_min_score` was left at 0.55 as instructed. **The
|
||
recommendation is to leave it there and stop tuning it**: any number under ~0.79 is a no-op and
|
||
anything above starts cutting real recall long before it stops the false ones. The fix is a margin
|
||
gate (`top1 − top2 > δ`), next-steps item 3, which is now the top item.
|
||
|
||
### The prefixes did not do the work
|
||
|
||
A control run with both prefixes set to the empty string scored the **same** recall@1 (72%), a
|
||
slightly better recall@3 (88%) and the same 5/5 false recall. So on this fixture the gain comes from
|
||
the model, not from the `query:` / `passage:` split. The prefixes are kept because they are how e5
|
||
was trained and the split is the right shape for the read path, but they are not worth defending on
|
||
this evidence — a bigger fixture may say otherwise.
|
||
|
||
### Stored vectors from the old model are now junk
|
||
|
||
Cosine between a MiniLM vector and an e5 vector means nothing. Every row already in `notes` and in
|
||
the vector memory table was written by the old model, so after this deploy they will score as noise
|
||
against a new query. A live database needs every note and fact re-embedded before recall works at
|
||
all. Filed as its own task.
|
||
|
||
## Margin gate — 31-07-2026, third run
|
||
|
||
Next-steps item 3, done. The absolute gate is replaced by a **margin gate**: answer only when the
|
||
top hit beats the runner-up by more than delta (`top1 − top2 > δ`). Same fixture, same e5 embedder,
|
||
same store as the run above. `internal/memory/gate.go` holds the check; both read paths call it
|
||
(`cmd/mavend/recall.go` and the notes-RAG branch in `voice.go`). New knob `voice.query_min_margin`
|
||
in `deploy/mavend.json`, default 0.008.
|
||
|
||
### Why the absolute gate could not work, in one line of data
|
||
|
||
The harness now prints the margin distributions, and they barely overlap where the raw scores
|
||
overlap completely:
|
||
|
||
| | top-1 score | margin (top1 − top2) |
|
||
|---|---|---|
|
||
| right note first (n=18) | min 0.810, median 0.862, max 0.890 | min 0.001, median 0.029, max 0.053 |
|
||
| must stay silent (n=5) | min 0.795, median 0.815, max 0.835 | min 0.000, median 0.002, **max 0.019** |
|
||
|
||
Four of the five must-be-silent cases have a margin at or under 0.002 — when there is nothing to
|
||
recall, e5 finds several notes equally close and no clear winner. That is the signal the absolute
|
||
score throws away.
|
||
|
||
### The delta sweep
|
||
|
||
Absolute gate held at 0.55 throughout.
|
||
|
||
```
|
||
delta 0.000: answered 18/25 (72%) false recall 5/5
|
||
delta 0.002: answered 17/25 (68%) false recall 3/5
|
||
delta 0.005: answered 17/25 (68%) false recall 2/5
|
||
delta 0.008: answered 17/25 (68%) false recall 1/5 <- chosen
|
||
delta 0.010: answered 15/25 (60%) false recall 1/5
|
||
delta 0.012: answered 14/25 (56%) false recall 1/5
|
||
delta 0.015: answered 12/25 (48%) false recall 1/5
|
||
delta 0.020: answered 11/25 (44%) false recall 0/5
|
||
delta 0.025: answered 9/25 (36%) false recall 0/5
|
||
delta 0.030: answered 8/25 (32%) false recall 0/5
|
||
delta 0.040: answered 4/25 (16%) false recall 0/5
|
||
delta 0.050: answered 2/25 ( 8%) false recall 0/5
|
||
delta 0.060: answered 0/25 ( 0%) false recall 0/5
|
||
```
|
||
|
||
### Chosen: δ = 0.008
|
||
|
||
It is the best point on the frontier, not a taste call. **0.008 dominates 0.010, 0.012 and 0.015
|
||
outright** — same 1/5 false recall, 8 to 20 points more real recall. Everything below it buys recall
|
||
back only by admitting more false recalls (0.005 → 2/5, 0.002 → 3/5). The next real improvement is
|
||
0.020 at 0/5 false, and it costs 24 points of recall to get there.
|
||
|
||
The brief's bar was "recall above 60% with false recall at 1/5 or better". 0.008 clears it with room:
|
||
68% and 1/5.
|
||
|
||
### Before / after
|
||
|
||
| | absolute gate 0.55 (previous) | margin gate δ=0.008 |
|
||
|---|---|---|
|
||
| recall@1 (ranking, ungated) | 72.0% (18/25) | 72.0% (18/25) — unchanged, the gate does not rank |
|
||
| **answered after the gate** | 72.0% (18/25) | **68.0% (17/25)** |
|
||
| **false recall** | **5/5 (100%)** | **1/5 (20%)** |
|
||
| fixture cases passed | 18/30 | **21/30** |
|
||
|
||
Four false recalls removed for one real answer. That is the trade the spec asks for — she is not a
|
||
guesser-of-truth. The one survivor is `en-pref-025` ("should i be offered wine"), which recalls a
|
||
filler note at 0.796 with a 0.019 margin: the widest silent-case margin in the fixture, and it sits
|
||
inside the real-recall range, so no delta removes it without taking real answers with it.
|
||
|
||
### Does the absolute cutoff still earn its keep? Marginally — kept
|
||
|
||
On this fixture with e5 it is a **no-op**: the lowest right-note score is 0.791, so 0.55 rejects
|
||
nothing the margin does not already reject. It is kept for two reasons, neither glamorous. It still
|
||
does real work for the hash embedder (its own sweep shows answers dropping from 16% to 0% between
|
||
0.30 and 0.50), and it is the only thing standing between the user and a reply built from a store
|
||
where everything is far away but one row happens to be a little less far — a near-empty database, or
|
||
the stale-vector case below. Cheap insurance, no measured cost. If a later embedder makes it bite,
|
||
the sweep is one command.
|
||
|
||
### Caveat on the numbers
|
||
|
||
Five must-be-silent cases is a thin basis for a 4-point decision. 1/5 and 2/5 differ by one case.
|
||
The shape of the frontier is trustworthy — margins separate, absolute scores do not — but δ=0.008
|
||
itself should be re-read off a bigger fixture (next-steps item 6) before anyone defends the third
|
||
decimal.
|
||
|
||
## Next steps — ordered by value-to-risk; nothing here is a decision
|
||
|
||
1. **Swap the embedder to `multilingual-e5-small` with `query:`/`passage:` prefixes.** One config
|
||
change plus a prefix in `onnxembedder.go`, re-measurable in one command.
|
||
2. **Re-run `make eval-recall`, then set the gate from the sweep** — not before. Any
|
||
`query_min_score` picked against today's embedder describes a model on its way out.
|
||
3. ~~**Replace the absolute-score gate with a margin gate**~~ — done, see the section above.
|
||
δ=0.008, false recall 5/5 → 1/5.
|
||
4. **Delete or repair the dead `memStore` branch** at `voice.go:776` — search before the gate,
|
||
gate it separately, or restrict it to facts and say so.
|
||
5. **Add a mild time decay to ranking** — the newest statement of a preference is the true one.
|
||
6. **Grow the fixture from real misses.** 30 cases can rank two embedders, not trust 4 points.
|
||
7. **Re-measure end to end.** Recall is gated twice — the utterance must first route to `query`,
|
||
which the routing eval puts at ~50%. The product is ~24%, and that is what he experiences.
|