The memory pass ran only after the notes-only gate had already rejected the same note at the same score. Notes and facts share one vector index, so a note that failed there failed again — the branch could only ever return a fact. Now the memory pass runs first: one search over everything Maven remembers, one gate, and the memory that clearly matches best answers (a note gets phrased, a fact is read back). The notes-only pass stays behind it for notes the vector index does not hold. No threshold moved, so the set of questions answered is unchanged — only which memory answers them. Fixture gained two mixed note+fact cases, so the answerable count goes 25 -> 27: hash recall@1 36.0% -> 37.0% (ratchet 0.32 unchanged, comment updated), e5 recall@1 72.0% -> 70.4%, false recall still 1/5. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
14 KiB
Note recall evaluation — 31-07-2026
The operator's goal is that Maven "memorize/note things … and know more about me/world". This measures whether the note/recall path delivers that.
- Fixture + scorer:
internal/memory/recalleval/(ru_recall_v1.json, 30 cases) - Reproduce:
make eval-recall— hash ratchet always, ONNX whendeps/is present - Commit:
43470ab(harness)
Each case inserts its own 3 notes plus 12 shared filler notes into a fresh store, embeds the
query, takes the top 3 — the read path cmd/mavend/voice.go runs for IntentQuery. Filler is
load-bearing: with 3 notes and a top-3 search, recall@3 is 100% by construction. 25 answerable
cases (paraphrased queries, homelab and preference content, 9 with a plausible second note) and 5
that must recall nothing. TestFixtureIsParaphrased fails the build if a query shares over half
its words with its note; equal-score ties count as ties, not recall.
Results
| recall+hash (CI ratchet) | recall+onnx (deployed) | |
|---|---|---|
| recall@1 | 36.0% (9/25) | 60.0% (15/25) |
| recall@3 | 76.0% (19/25) | 80.0% (20/25) |
| answered after the 0.55 gate | 0.0% (0/25) | 48.0% (12/25) |
| wrong note on top / tie on top | 9 / 7 | 10 / 0 |
| ranked first, then silenced by the gate | 9 | 3 |
| false recall | 0/5 | 1/5 (20%) |
| top-1 score when right, min / median | n/a | 0.559 / 0.678 |
| top-1 when it must stay silent, median / max | 0.000 / 0.144 | 0.470 / 0.567 |
RU / EN / hard cases passed |
4/24 / 1/6 / 0/11 | 13/24 / 3/6 / 2/11 |
| latency p50 / p95 / max | 49µs / 70µs | 59ms / 148ms / 194ms |
Never compare a hash-embedder number to an ONNX one — the hash floor is lexical and exists only so CI has a deterministic ratchet with no model files.
Findings
1. Real recall is 48%, not 60%
The right note ranks first 60% of the time, but the daemon only says it 48% of the time — three
more cases rank first and are then silenced by voice.go:776's queryMinScore. Roughly one
useful question in two gets "не знаю". This is not a working memory yet.
2. The gate cannot separate a real recall from a false one — the distributions overlap
Right-note top-1 scores start at 0.559. Must-stay-silent top-1 scores reach 0.567. No
threshold keeps every real recall and rejects every false one. From the sweep: gate 0.50 → 13/25
answered, 1/5 false; 0.55 (default) → 12/25, 1/5; 0.60 → 10/25, 0/5; 0.70 → 5/25, 0/5. What
the data says about DefaultQueryMinScore (internal/config/config.go:392): 0.55 is
slightly too loose — it admits one confident wrong answer ("как зовут сестру моего коллеги"
recalls "выучил пару аккордов на гитаре" at 0.567), which the spec ranks as worse than a gap. 0.60
silences all five and costs 8 points of real recall. Left alone as instructed; the overlap means
the threshold is the wrong dial anyway (finding 3).
3. Filler notes outrank the right answer — the model scores similarity, not relevance
models/embedder/ is paraphrase-multilingual-MiniLM-L12-v2 (Makefile:119), a symmetric
paraphrase model. It scores "do these sentences look alike", not "does this passage answer this
question", so question-shaped queries drift to whatever note is stylistically closest. "из-за чего
кончилось место" and "откуда берётся токен бота" both return выучил пару аккордов на гитаре
(0.730, 0.729); "как я восстановил конфиги" returns a bootloader note at 0.703 with the right note
not even in the top 3. An unrelated guitar note beating a homelab note at 0.73 is not a tuning
problem — an asymmetric retrieval model (multilingual-e5-small, with query: / passage:
prefixes) is the targeted fix, and it would move findings 1 and 2 together. Separately:
deploy/mavend.json:39 loads a 470MB fp32 model.onnx while make download-embedder fetches
model_quantized.onnx — not the same file.
hard cases score 2/11: every one is a query where the operator did not reuse his own words.
That is the normal case weeks later, and exactly what DESIGN.md's "recall when relevant" promises.
4. The memory-store recall branch is dead for notes
voice.go:776 only reaches h.memStore.Search when the notes-RAG top score is already below
queryMinScore, and bestRecall (cmd/mavend/recall.go:19) then applies the same gate to the
same vector. A note is indexed in both places with the same embedding, so if it failed the gate in
QueryNotes it fails again here — the branch can only ever return a fact. Its comment calls it
"additive"; for notes it is not.
Fixed (Vikunja #373). The memory pass now runs first, as one search over notes and facts with
one gate, so whichever memory is clearly the best match answers — note or fact. The notes-only pass
stays behind it for notes the vector index does not hold. No threshold changed, so the set of
questions Maven answers is the same; only which memory answers them. The fixture gained two mixed
note+fact cases (ru-mixed-031, ru-mixed-032), which is why the counts below are out of 27
answerable cases and not 25: hash recall@1 36.0% (9/25) → 37.0% (10/27), e5 recall@1 72.0% (18/25) →
70.4% (19/27) with answered-after-gate 68.0% → 66.7% and false recall unchanged at 1/5.
5. Ranking has no recency or type signal, and the store is not the bottleneck
internal/store/notes.go:67 sorts by cosine and uses ts only to break an exact float tie, which
never happens; kind never enters the ranking. Meanwhile TestPersistentStoreScoresTheSame scores
sqlite-backed store.MemoryStore and memory.InMemoryStore identically — both full-scan cosine
(internal/store/memory.go:64) at ~150µs over 42 rows against a ~59ms query embed. An ANN index is
not the problem to solve.
Re-measured after the embedder swap — 31-07-2026, later the same day
Changed: models/embedder/ is now multilingual-e5-small (quantized, 118MB), with query: in
front of a question and passage: in front of a stored note (Vikunja #371). deploy/mavend.json
and make download-embedder now name the same file, and it is the quantized one — that is what the
column below measures (Vikunja #372). Everything else is unchanged: same fixture, same store, same
0.55 gate. The old column is the baseline and is left as it was.
| recall+onnx, MiniLM (baseline) | recall+onnx, e5-small (new) | |
|---|---|---|
| recall@1 | 60.0% (15/25) | 72.0% (18/25) |
| recall@3 | 80.0% (20/25) | 84.0% (21/25) |
| answered after the 0.55 gate | 48.0% (12/25) | 72.0% (18/25) |
| wrong note on top / tie on top | 10 / 0 | 7 / 0 |
| ranked first, then silenced by the gate | 3 | 0 |
| false recall | 1/5 (20%) | 5/5 (100%) |
| top-1 score when right, min / median | 0.559 / 0.678 | 0.791 / 0.857 |
| top-1 when it must stay silent, median / max | 0.470 / 0.567 | 0.815 / 0.835 |
RU / EN / hard cases passed |
13/24 / 3/6 / 2/11 | 14/24 / 4/6 / 5/11 |
| latency p50 / p95 / max | 59ms / 148ms / 194ms | 18ms / 37ms / 49ms |
What moved
Ranking got better and got faster. Half the previously-unwinnable hard cases now pass (2/11 →
5/11), the guitar note no longer beats the docker-logs note, and the gate stops silencing notes that
already ranked first. The quantized e5 is also ~3x quicker than the fp32 MiniLM it replaces.
What got worse: the gate is now a no-op
e5 packs every cosine into a narrow high band. Right-note scores start at 0.791; must-stay-silent scores reach 0.835. The distributions still overlap, and now they overlap above the gate, so 0.55 admits everything and false recall goes from 1/5 to 5/5. The sweep:
gate 0.50–0.70: answered 18/25 (72%) false recall 5/5
gate 0.80: answered 17/25 (68%) false recall 4/5
gate 0.90: answered 0/25 ( 0%) false recall 0/5
There is no value that keeps real recall and rejects made-up questions — same conclusion as before,
now with a wider band and no room at all. query_min_score was left at 0.55 as instructed. The
recommendation is to leave it there and stop tuning it: any number under ~0.79 is a no-op and
anything above starts cutting real recall long before it stops the false ones. The fix is a margin
gate (top1 − top2 > δ), next-steps item 3, which is now the top item.
The prefixes did not do the work
A control run with both prefixes set to the empty string scored the same recall@1 (72%), a
slightly better recall@3 (88%) and the same 5/5 false recall. So on this fixture the gain comes from
the model, not from the query: / passage: split. The prefixes are kept because they are how e5
was trained and the split is the right shape for the read path, but they are not worth defending on
this evidence — a bigger fixture may say otherwise.
Stored vectors from the old model are now junk
Cosine between a MiniLM vector and an e5 vector means nothing. Every row already in notes and in
the vector memory table was written by the old model, so after this deploy they will score as noise
against a new query. A live database needs every note and fact re-embedded before recall works at
all. Filed as its own task.
Margin gate — 31-07-2026, third run
Next-steps item 3, done. The absolute gate is replaced by a margin gate: answer only when the
top hit beats the runner-up by more than delta (top1 − top2 > δ). Same fixture, same e5 embedder,
same store as the run above. internal/memory/gate.go holds the check; both read paths call it
(cmd/mavend/recall.go and the notes-RAG branch in voice.go). New knob voice.query_min_margin
in deploy/mavend.json, default 0.008.
Why the absolute gate could not work, in one line of data
The harness now prints the margin distributions, and they barely overlap where the raw scores overlap completely:
| top-1 score | margin (top1 − top2) | |
|---|---|---|
| right note first (n=18) | min 0.810, median 0.862, max 0.890 | min 0.001, median 0.029, max 0.053 |
| must stay silent (n=5) | min 0.795, median 0.815, max 0.835 | min 0.000, median 0.002, max 0.019 |
Four of the five must-be-silent cases have a margin at or under 0.002 — when there is nothing to recall, e5 finds several notes equally close and no clear winner. That is the signal the absolute score throws away.
The delta sweep
Absolute gate held at 0.55 throughout.
delta 0.000: answered 18/25 (72%) false recall 5/5
delta 0.002: answered 17/25 (68%) false recall 3/5
delta 0.005: answered 17/25 (68%) false recall 2/5
delta 0.008: answered 17/25 (68%) false recall 1/5 <- chosen
delta 0.010: answered 15/25 (60%) false recall 1/5
delta 0.012: answered 14/25 (56%) false recall 1/5
delta 0.015: answered 12/25 (48%) false recall 1/5
delta 0.020: answered 11/25 (44%) false recall 0/5
delta 0.025: answered 9/25 (36%) false recall 0/5
delta 0.030: answered 8/25 (32%) false recall 0/5
delta 0.040: answered 4/25 (16%) false recall 0/5
delta 0.050: answered 2/25 ( 8%) false recall 0/5
delta 0.060: answered 0/25 ( 0%) false recall 0/5
Chosen: δ = 0.008
It is the best point on the frontier, not a taste call. 0.008 dominates 0.010, 0.012 and 0.015 outright — same 1/5 false recall, 8 to 20 points more real recall. Everything below it buys recall back only by admitting more false recalls (0.005 → 2/5, 0.002 → 3/5). The next real improvement is 0.020 at 0/5 false, and it costs 24 points of recall to get there.
The brief's bar was "recall above 60% with false recall at 1/5 or better". 0.008 clears it with room: 68% and 1/5.
Before / after
| absolute gate 0.55 (previous) | margin gate δ=0.008 | |
|---|---|---|
| recall@1 (ranking, ungated) | 72.0% (18/25) | 72.0% (18/25) — unchanged, the gate does not rank |
| answered after the gate | 72.0% (18/25) | 68.0% (17/25) |
| false recall | 5/5 (100%) | 1/5 (20%) |
| fixture cases passed | 18/30 | 21/30 |
Four false recalls removed for one real answer. That is the trade the spec asks for — she is not a
guesser-of-truth. The one survivor is en-pref-025 ("should i be offered wine"), which recalls a
filler note at 0.796 with a 0.019 margin: the widest silent-case margin in the fixture, and it sits
inside the real-recall range, so no delta removes it without taking real answers with it.
Does the absolute cutoff still earn its keep? Marginally — kept
On this fixture with e5 it is a no-op: the lowest right-note score is 0.791, so 0.55 rejects nothing the margin does not already reject. It is kept for two reasons, neither glamorous. It still does real work for the hash embedder (its own sweep shows answers dropping from 16% to 0% between 0.30 and 0.50), and it is the only thing standing between the user and a reply built from a store where everything is far away but one row happens to be a little less far — a near-empty database, or the stale-vector case below. Cheap insurance, no measured cost. If a later embedder makes it bite, the sweep is one command.
Caveat on the numbers
Five must-be-silent cases is a thin basis for a 4-point decision. 1/5 and 2/5 differ by one case. The shape of the frontier is trustworthy — margins separate, absolute scores do not — but δ=0.008 itself should be re-read off a bigger fixture (next-steps item 6) before anyone defends the third decimal.
Next steps — ordered by value-to-risk; nothing here is a decision
- Swap the embedder to
multilingual-e5-smallwithquery:/passage:prefixes. One config change plus a prefix inonnxembedder.go, re-measurable in one command. - Re-run
make eval-recall, then set the gate from the sweep — not before. Anyquery_min_scorepicked against today's embedder describes a model on its way out. Replace the absolute-score gate with a margin gate— done, see the section above. δ=0.008, false recall 5/5 → 1/5.- Delete or repair the dead
memStorebranch atvoice.go:776— search before the gate, gate it separately, or restrict it to facts and say so. - Add a mild time decay to ranking — the newest statement of a preference is the true one.
- Grow the fixture from real misses. 30 cases can rank two embedders, not trust 4 points.
- Re-measure end to end. Recall is gated twice — the utterance must first route to
query, which the routing eval puts at ~50%. The product is ~24%, and that is what he experiences.