Files
Maven/RECALL-EVAL-31-07-2026.md
T
kami 98ee701e03 Let a note win a recall, not only a fact (#373)
The memory pass ran only after the notes-only gate had already rejected
the same note at the same score. Notes and facts share one vector index,
so a note that failed there failed again — the branch could only ever
return a fact.

Now the memory pass runs first: one search over everything Maven
remembers, one gate, and the memory that clearly matches best answers
(a note gets phrased, a fact is read back). The notes-only pass stays
behind it for notes the vector index does not hold. No threshold moved,
so the set of questions answered is unchanged — only which memory
answers them.

Fixture gained two mixed note+fact cases, so the answerable count goes
25 -> 27: hash recall@1 36.0% -> 37.0% (ratchet 0.32 unchanged, comment
updated), e5 recall@1 72.0% -> 70.4%, false recall still 1/5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:30:38 +04:00

14 KiB
Raw Blame History

Note recall evaluation — 31-07-2026

The operator's goal is that Maven "memorize/note things … and know more about me/world". This measures whether the note/recall path delivers that.

  • Fixture + scorer: internal/memory/recalleval/ (ru_recall_v1.json, 30 cases)
  • Reproduce: make eval-recall — hash ratchet always, ONNX when deps/ is present
  • Commit: 43470ab (harness)

Each case inserts its own 3 notes plus 12 shared filler notes into a fresh store, embeds the query, takes the top 3 — the read path cmd/mavend/voice.go runs for IntentQuery. Filler is load-bearing: with 3 notes and a top-3 search, recall@3 is 100% by construction. 25 answerable cases (paraphrased queries, homelab and preference content, 9 with a plausible second note) and 5 that must recall nothing. TestFixtureIsParaphrased fails the build if a query shares over half its words with its note; equal-score ties count as ties, not recall.

Results

recall+hash (CI ratchet) recall+onnx (deployed)
recall@1 36.0% (9/25) 60.0% (15/25)
recall@3 76.0% (19/25) 80.0% (20/25)
answered after the 0.55 gate 0.0% (0/25) 48.0% (12/25)
wrong note on top / tie on top 9 / 7 10 / 0
ranked first, then silenced by the gate 9 3
false recall 0/5 1/5 (20%)
top-1 score when right, min / median n/a 0.559 / 0.678
top-1 when it must stay silent, median / max 0.000 / 0.144 0.470 / 0.567
RU / EN / hard cases passed 4/24 / 1/6 / 0/11 13/24 / 3/6 / 2/11
latency p50 / p95 / max 49µs / 70µs 59ms / 148ms / 194ms

Never compare a hash-embedder number to an ONNX one — the hash floor is lexical and exists only so CI has a deterministic ratchet with no model files.

Findings

1. Real recall is 48%, not 60%

The right note ranks first 60% of the time, but the daemon only says it 48% of the time — three more cases rank first and are then silenced by voice.go:776's queryMinScore. Roughly one useful question in two gets "не знаю". This is not a working memory yet.

2. The gate cannot separate a real recall from a false one — the distributions overlap

Right-note top-1 scores start at 0.559. Must-stay-silent top-1 scores reach 0.567. No threshold keeps every real recall and rejects every false one. From the sweep: gate 0.50 → 13/25 answered, 1/5 false; 0.55 (default) → 12/25, 1/5; 0.60 → 10/25, 0/5; 0.70 → 5/25, 0/5. What the data says about DefaultQueryMinScore (internal/config/config.go:392): 0.55 is slightly too loose — it admits one confident wrong answer ("как зовут сестру моего коллеги" recalls "выучил пару аккордов на гитаре" at 0.567), which the spec ranks as worse than a gap. 0.60 silences all five and costs 8 points of real recall. Left alone as instructed; the overlap means the threshold is the wrong dial anyway (finding 3).

3. Filler notes outrank the right answer — the model scores similarity, not relevance

models/embedder/ is paraphrase-multilingual-MiniLM-L12-v2 (Makefile:119), a symmetric paraphrase model. It scores "do these sentences look alike", not "does this passage answer this question", so question-shaped queries drift to whatever note is stylistically closest. "из-за чего кончилось место" and "откуда берётся токен бота" both return выучил пару аккордов на гитаре (0.730, 0.729); "как я восстановил конфиги" returns a bootloader note at 0.703 with the right note not even in the top 3. An unrelated guitar note beating a homelab note at 0.73 is not a tuning problem — an asymmetric retrieval model (multilingual-e5-small, with query: / passage: prefixes) is the targeted fix, and it would move findings 1 and 2 together. Separately: deploy/mavend.json:39 loads a 470MB fp32 model.onnx while make download-embedder fetches model_quantized.onnx — not the same file.

hard cases score 2/11: every one is a query where the operator did not reuse his own words. That is the normal case weeks later, and exactly what DESIGN.md's "recall when relevant" promises.

4. The memory-store recall branch is dead for notes

voice.go:776 only reaches h.memStore.Search when the notes-RAG top score is already below queryMinScore, and bestRecall (cmd/mavend/recall.go:19) then applies the same gate to the same vector. A note is indexed in both places with the same embedding, so if it failed the gate in QueryNotes it fails again here — the branch can only ever return a fact. Its comment calls it "additive"; for notes it is not.

Fixed (Vikunja #373). The memory pass now runs first, as one search over notes and facts with one gate, so whichever memory is clearly the best match answers — note or fact. The notes-only pass stays behind it for notes the vector index does not hold. No threshold changed, so the set of questions Maven answers is the same; only which memory answers them. The fixture gained two mixed note+fact cases (ru-mixed-031, ru-mixed-032), which is why the counts below are out of 27 answerable cases and not 25: hash recall@1 36.0% (9/25) → 37.0% (10/27), e5 recall@1 72.0% (18/25) → 70.4% (19/27) with answered-after-gate 68.0% → 66.7% and false recall unchanged at 1/5.

5. Ranking has no recency or type signal, and the store is not the bottleneck

internal/store/notes.go:67 sorts by cosine and uses ts only to break an exact float tie, which never happens; kind never enters the ranking. Meanwhile TestPersistentStoreScoresTheSame scores sqlite-backed store.MemoryStore and memory.InMemoryStore identically — both full-scan cosine (internal/store/memory.go:64) at ~150µs over 42 rows against a ~59ms query embed. An ANN index is not the problem to solve.

Re-measured after the embedder swap — 31-07-2026, later the same day

Changed: models/embedder/ is now multilingual-e5-small (quantized, 118MB), with query: in front of a question and passage: in front of a stored note (Vikunja #371). deploy/mavend.json and make download-embedder now name the same file, and it is the quantized one — that is what the column below measures (Vikunja #372). Everything else is unchanged: same fixture, same store, same 0.55 gate. The old column is the baseline and is left as it was.

recall+onnx, MiniLM (baseline) recall+onnx, e5-small (new)
recall@1 60.0% (15/25) 72.0% (18/25)
recall@3 80.0% (20/25) 84.0% (21/25)
answered after the 0.55 gate 48.0% (12/25) 72.0% (18/25)
wrong note on top / tie on top 10 / 0 7 / 0
ranked first, then silenced by the gate 3 0
false recall 1/5 (20%) 5/5 (100%)
top-1 score when right, min / median 0.559 / 0.678 0.791 / 0.857
top-1 when it must stay silent, median / max 0.470 / 0.567 0.815 / 0.835
RU / EN / hard cases passed 13/24 / 3/6 / 2/11 14/24 / 4/6 / 5/11
latency p50 / p95 / max 59ms / 148ms / 194ms 18ms / 37ms / 49ms

What moved

Ranking got better and got faster. Half the previously-unwinnable hard cases now pass (2/11 → 5/11), the guitar note no longer beats the docker-logs note, and the gate stops silencing notes that already ranked first. The quantized e5 is also ~3x quicker than the fp32 MiniLM it replaces.

What got worse: the gate is now a no-op

e5 packs every cosine into a narrow high band. Right-note scores start at 0.791; must-stay-silent scores reach 0.835. The distributions still overlap, and now they overlap above the gate, so 0.55 admits everything and false recall goes from 1/5 to 5/5. The sweep:

gate 0.500.70: answered 18/25 (72%)  false recall 5/5
gate 0.80:      answered 17/25 (68%)  false recall 4/5
gate 0.90:      answered  0/25 ( 0%)  false recall 0/5

There is no value that keeps real recall and rejects made-up questions — same conclusion as before, now with a wider band and no room at all. query_min_score was left at 0.55 as instructed. The recommendation is to leave it there and stop tuning it: any number under ~0.79 is a no-op and anything above starts cutting real recall long before it stops the false ones. The fix is a margin gate (top1 top2 > δ), next-steps item 3, which is now the top item.

The prefixes did not do the work

A control run with both prefixes set to the empty string scored the same recall@1 (72%), a slightly better recall@3 (88%) and the same 5/5 false recall. So on this fixture the gain comes from the model, not from the query: / passage: split. The prefixes are kept because they are how e5 was trained and the split is the right shape for the read path, but they are not worth defending on this evidence — a bigger fixture may say otherwise.

Stored vectors from the old model are now junk

Cosine between a MiniLM vector and an e5 vector means nothing. Every row already in notes and in the vector memory table was written by the old model, so after this deploy they will score as noise against a new query. A live database needs every note and fact re-embedded before recall works at all. Filed as its own task.

Margin gate — 31-07-2026, third run

Next-steps item 3, done. The absolute gate is replaced by a margin gate: answer only when the top hit beats the runner-up by more than delta (top1 top2 > δ). Same fixture, same e5 embedder, same store as the run above. internal/memory/gate.go holds the check; both read paths call it (cmd/mavend/recall.go and the notes-RAG branch in voice.go). New knob voice.query_min_margin in deploy/mavend.json, default 0.008.

Why the absolute gate could not work, in one line of data

The harness now prints the margin distributions, and they barely overlap where the raw scores overlap completely:

top-1 score margin (top1 top2)
right note first (n=18) min 0.810, median 0.862, max 0.890 min 0.001, median 0.029, max 0.053
must stay silent (n=5) min 0.795, median 0.815, max 0.835 min 0.000, median 0.002, max 0.019

Four of the five must-be-silent cases have a margin at or under 0.002 — when there is nothing to recall, e5 finds several notes equally close and no clear winner. That is the signal the absolute score throws away.

The delta sweep

Absolute gate held at 0.55 throughout.

delta 0.000: answered 18/25 (72%)  false recall 5/5
delta 0.002: answered 17/25 (68%)  false recall 3/5
delta 0.005: answered 17/25 (68%)  false recall 2/5
delta 0.008: answered 17/25 (68%)  false recall 1/5   <- chosen
delta 0.010: answered 15/25 (60%)  false recall 1/5
delta 0.012: answered 14/25 (56%)  false recall 1/5
delta 0.015: answered 12/25 (48%)  false recall 1/5
delta 0.020: answered 11/25 (44%)  false recall 0/5
delta 0.025: answered  9/25 (36%)  false recall 0/5
delta 0.030: answered  8/25 (32%)  false recall 0/5
delta 0.040: answered  4/25 (16%)  false recall 0/5
delta 0.050: answered  2/25 ( 8%)  false recall 0/5
delta 0.060: answered  0/25 ( 0%)  false recall 0/5

Chosen: δ = 0.008

It is the best point on the frontier, not a taste call. 0.008 dominates 0.010, 0.012 and 0.015 outright — same 1/5 false recall, 8 to 20 points more real recall. Everything below it buys recall back only by admitting more false recalls (0.005 → 2/5, 0.002 → 3/5). The next real improvement is 0.020 at 0/5 false, and it costs 24 points of recall to get there.

The brief's bar was "recall above 60% with false recall at 1/5 or better". 0.008 clears it with room: 68% and 1/5.

Before / after

absolute gate 0.55 (previous) margin gate δ=0.008
recall@1 (ranking, ungated) 72.0% (18/25) 72.0% (18/25) — unchanged, the gate does not rank
answered after the gate 72.0% (18/25) 68.0% (17/25)
false recall 5/5 (100%) 1/5 (20%)
fixture cases passed 18/30 21/30

Four false recalls removed for one real answer. That is the trade the spec asks for — she is not a guesser-of-truth. The one survivor is en-pref-025 ("should i be offered wine"), which recalls a filler note at 0.796 with a 0.019 margin: the widest silent-case margin in the fixture, and it sits inside the real-recall range, so no delta removes it without taking real answers with it.

Does the absolute cutoff still earn its keep? Marginally — kept

On this fixture with e5 it is a no-op: the lowest right-note score is 0.791, so 0.55 rejects nothing the margin does not already reject. It is kept for two reasons, neither glamorous. It still does real work for the hash embedder (its own sweep shows answers dropping from 16% to 0% between 0.30 and 0.50), and it is the only thing standing between the user and a reply built from a store where everything is far away but one row happens to be a little less far — a near-empty database, or the stale-vector case below. Cheap insurance, no measured cost. If a later embedder makes it bite, the sweep is one command.

Caveat on the numbers

Five must-be-silent cases is a thin basis for a 4-point decision. 1/5 and 2/5 differ by one case. The shape of the frontier is trustworthy — margins separate, absolute scores do not — but δ=0.008 itself should be re-read off a bigger fixture (next-steps item 6) before anyone defends the third decimal.

Next steps — ordered by value-to-risk; nothing here is a decision

  1. Swap the embedder to multilingual-e5-small with query:/passage: prefixes. One config change plus a prefix in onnxembedder.go, re-measurable in one command.
  2. Re-run make eval-recall, then set the gate from the sweep — not before. Any query_min_score picked against today's embedder describes a model on its way out.
  3. Replace the absolute-score gate with a margin gate — done, see the section above. δ=0.008, false recall 5/5 → 1/5.
  4. Delete or repair the dead memStore branch at voice.go:776 — search before the gate, gate it separately, or restrict it to facts and say so.
  5. Add a mild time decay to ranking — the newest statement of a preference is the true one.
  6. Grow the fixture from real misses. 30 cases can rank two embedders, not trust 4 points.
  7. Re-measure end to end. Recall is gated twice — the utterance must first route to query, which the routing eval puts at ~50%. The product is ~24%, and that is what he experiences.