Compare commits
2 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 8a174c1c70 | |||
| 43470abc57 |
@@ -16,7 +16,7 @@ PIPER_BIN := $(shell pwd)/deps/piper/piper
|
||||
PIPER_MODEL := $(shell pwd)/models/tts/ru_RU-irina-medium.onnx
|
||||
PIPER_ESPEAK := $(shell pwd)/deps/piper/espeak-ng-data
|
||||
|
||||
.PHONY: all build build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav clean test run-stt run-tts run-web download-embedder deps-go eval-router
|
||||
.PHONY: all build build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav clean test run-stt run-tts run-web download-embedder deps-go eval-router eval-recall
|
||||
|
||||
all: build
|
||||
|
||||
@@ -83,6 +83,13 @@ MAVEN_ONNX_LIB ?= $(shell pwd)/deps/onnxruntime-linux-x64-1.26.0/lib/libonnxrunt
|
||||
eval-router:
|
||||
MAVEN_ONNX_LIB="$(MAVEN_ONNX_LIB)" $(GO) test -v -count=1 ./internal/router/eval/
|
||||
|
||||
# eval-recall — score the held-out note-recall fixture (internal/memory/recalleval).
|
||||
# Answers "can she find the note again when it matters": recall@1, recall@3,
|
||||
# false recall and the query_min_score sweep. Same MAVEN_ONNX_LIB deal as
|
||||
# eval-router; without it only the deterministic hash ratchet runs.
|
||||
eval-recall:
|
||||
MAVEN_ONNX_LIB="$(MAVEN_ONNX_LIB)" $(GO) test -v -count=1 ./internal/memory/recalleval/
|
||||
|
||||
run-stt: build-stt
|
||||
LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
|
||||
./mavsttd -socket /tmp/maven/stt.sock -model $(WHISPER_MODEL)
|
||||
|
||||
@@ -0,0 +1,99 @@
|
||||
# Note recall evaluation — 31-07-2026
|
||||
|
||||
The operator's goal is that Maven "memorize/note things … and know more about me/world". This
|
||||
measures whether the note/recall path delivers that.
|
||||
|
||||
- Fixture + scorer: `internal/memory/recalleval/` (`ru_recall_v1.json`, 30 cases)
|
||||
- Reproduce: `make eval-recall` — hash ratchet always, ONNX when `deps/` is present
|
||||
- Commit: `43470ab` (harness)
|
||||
|
||||
Each case inserts its own 3 notes **plus 12 shared filler notes** into a fresh store, embeds the
|
||||
query, takes the top 3 — the read path `cmd/mavend/voice.go` runs for `IntentQuery`. Filler is
|
||||
load-bearing: with 3 notes and a top-3 search, recall@3 is 100% by construction. 25 answerable
|
||||
cases (paraphrased queries, homelab and preference content, 9 with a plausible second note) and 5
|
||||
that must recall **nothing**. `TestFixtureIsParaphrased` fails the build if a query shares over half
|
||||
its words with its note; equal-score ties count as ties, not recall.
|
||||
|
||||
## Results
|
||||
|
||||
| | recall+hash (CI ratchet) | recall+onnx (deployed) |
|
||||
|---|---|---|
|
||||
| **recall@1** | 36.0% (9/25) | **60.0% (15/25)** |
|
||||
| recall@3 | 76.0% (19/25) | 80.0% (20/25) |
|
||||
| **answered after the 0.55 gate** | **0.0% (0/25)** | **48.0% (12/25)** |
|
||||
| wrong note on top / tie on top | 9 / 7 | 10 / 0 |
|
||||
| ranked first, then silenced by the gate | 9 | 3 |
|
||||
| **false recall** | 0/5 | **1/5 (20%)** |
|
||||
| top-1 score when right, min / median | n/a | 0.559 / 0.678 |
|
||||
| top-1 when it must stay silent, median / max | 0.000 / 0.144 | 0.470 / **0.567** |
|
||||
| RU / EN / `hard` cases passed | 4/24 / 1/6 / 0/11 | 13/24 / 3/6 / 2/11 |
|
||||
| latency p50 / p95 / max | 49µs / 70µs | 59ms / 148ms / 194ms |
|
||||
|
||||
Never compare a hash-embedder number to an ONNX one — the hash floor is lexical and exists only so
|
||||
CI has a deterministic ratchet with no model files.
|
||||
|
||||
## Findings
|
||||
|
||||
### 1. Real recall is 48%, not 60%
|
||||
|
||||
The right note ranks first 60% of the time, but the daemon only *says* it 48% of the time — three
|
||||
more cases rank first and are then silenced by `voice.go:776`'s `queryMinScore`. **Roughly one
|
||||
useful question in two gets "не знаю".** This is not a working memory yet.
|
||||
|
||||
### 2. The gate cannot separate a real recall from a false one — the distributions overlap
|
||||
|
||||
Right-note top-1 scores start at **0.559**. Must-stay-silent top-1 scores reach **0.567**. No
|
||||
threshold keeps every real recall and rejects every false one. From the sweep: gate 0.50 → 13/25
|
||||
answered, 1/5 false; **0.55 (default) → 12/25, 1/5**; **0.60 → 10/25, 0/5**; 0.70 → 5/25, 0/5. What
|
||||
the data says about `DefaultQueryMinScore` (`internal/config/config.go:392`): **0.55 is
|
||||
slightly too loose** — it admits one confident wrong answer ("как зовут сестру моего коллеги"
|
||||
recalls "выучил пару аккордов на гитаре" at 0.567), which the spec ranks as worse than a gap. 0.60
|
||||
silences all five and costs 8 points of real recall. Left alone as instructed; the overlap means
|
||||
the threshold is the wrong dial anyway (finding 3).
|
||||
|
||||
### 3. Filler notes outrank the right answer — the model scores similarity, not relevance
|
||||
|
||||
`models/embedder/` is **paraphrase-multilingual-MiniLM-L12-v2** (`Makefile:119`), a *symmetric*
|
||||
paraphrase model. It scores "do these sentences look alike", not "does this passage answer this
|
||||
question", so question-shaped queries drift to whatever note is stylistically closest. "из-за чего
|
||||
кончилось место" and "откуда берётся токен бота" both return `выучил пару аккордов на гитаре`
|
||||
(0.730, 0.729); "как я восстановил конфиги" returns a bootloader note at 0.703 with the right note
|
||||
not even in the top 3. An unrelated guitar note beating a homelab note at 0.73 is not a tuning
|
||||
problem — an asymmetric retrieval model (`multilingual-e5-small`, with `query:` / `passage:`
|
||||
prefixes) is the targeted fix, and it would move findings 1 and 2 together. Separately:
|
||||
`deploy/mavend.json:39` loads a 470MB fp32 `model.onnx` while `make download-embedder` fetches
|
||||
`model_quantized.onnx` — not the same file.
|
||||
|
||||
`hard` cases score **2/11**: every one is a query where the operator did not reuse his own words.
|
||||
That is the normal case weeks later, and exactly what DESIGN.md's "recall when relevant" promises.
|
||||
|
||||
### 4. The memory-store recall branch is dead for notes
|
||||
|
||||
`voice.go:776` only reaches `h.memStore.Search` when the notes-RAG top score is already below
|
||||
`queryMinScore`, and `bestRecall` (`cmd/mavend/recall.go:19`) then applies the **same** gate to the
|
||||
same vector. A note is indexed in both places with the same embedding, so if it failed the gate in
|
||||
`QueryNotes` it fails again here — the branch can only ever return a **fact**. Its comment calls it
|
||||
"additive"; for notes it is not.
|
||||
|
||||
### 5. Ranking has no recency or type signal, and the store is not the bottleneck
|
||||
|
||||
`internal/store/notes.go:67` sorts by cosine and uses `ts` only to break an exact float tie, which
|
||||
never happens; `kind` never enters the ranking. Meanwhile `TestPersistentStoreScoresTheSame` scores
|
||||
sqlite-backed `store.MemoryStore` and `memory.InMemoryStore` identically — both full-scan cosine
|
||||
(`internal/store/memory.go:64`) at ~150µs over 42 rows against a ~59ms query embed. An ANN index is
|
||||
not the problem to solve.
|
||||
|
||||
## Next steps — ordered by value-to-risk; nothing here is a decision
|
||||
|
||||
1. **Swap the embedder to `multilingual-e5-small` with `query:`/`passage:` prefixes.** One config
|
||||
change plus a prefix in `onnxembedder.go`, re-measurable in one command.
|
||||
2. **Re-run `make eval-recall`, then set the gate from the sweep** — not before. Any
|
||||
`query_min_score` picked against today's embedder describes a model on its way out.
|
||||
3. **Replace the absolute-score gate with a margin gate** (`top1 − top2 > δ`) — as the routing eval
|
||||
concluded, absolute cosine cannot see a flat distribution.
|
||||
4. **Delete or repair the dead `memStore` branch** at `voice.go:776` — search before the gate,
|
||||
gate it separately, or restrict it to facts and say so.
|
||||
5. **Add a mild time decay to ranking** — the newest statement of a preference is the true one.
|
||||
6. **Grow the fixture from real misses.** 30 cases can rank two embedders, not trust 4 points.
|
||||
7. **Re-measure end to end.** Recall is gated twice — the utterance must first route to `query`,
|
||||
which the routing eval puts at ~50%. The product is ~24%, and that is what he experiences.
|
||||
@@ -639,11 +639,11 @@ func TestDispatchRecurringReminderReschedules(t *testing.T) {
|
||||
// ----------------------------- durable outbox --------------------------------
|
||||
|
||||
type outboxAttempt struct {
|
||||
kind, rule string
|
||||
reminderID int64
|
||||
channel, hash string
|
||||
status string
|
||||
begunAt, doneAt time.Time
|
||||
kind, rule string
|
||||
reminderID int64
|
||||
channel, hash string
|
||||
status string
|
||||
begunAt, doneAt time.Time
|
||||
}
|
||||
|
||||
// fakeOutbox — an in-memory Outbox that also lets a test simulate a crash
|
||||
|
||||
@@ -1,247 +0,0 @@
|
||||
package delivery
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/loop"
|
||||
"github.com/kami/maven/internal/store"
|
||||
)
|
||||
|
||||
// panicSink — a sink that dies mid-send. Models the ugly case: the process is
|
||||
// still alive, so startup reconciliation will not run, but the attempt row was
|
||||
// already begun.
|
||||
type panicSink struct{ calls int }
|
||||
|
||||
func (p *panicSink) Send(_ context.Context, _ Sendable) error {
|
||||
p.calls++
|
||||
panic("sink exploded mid-send")
|
||||
}
|
||||
|
||||
// ------------------------- voice fallthrough, per severity -------------------
|
||||
|
||||
// TestVoiceNoSessionFallthroughLeavesOutboxTrail — the fallthrough must be
|
||||
// visible in the ledger too: the voice attempt closes as failed and the away
|
||||
// attempt is a separate row, so an operator can see the reroute happened.
|
||||
func TestVoiceNoSessionFallthroughLeavesOutboxTrail(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
sev loop.Severity
|
||||
wantAt []string // channel per outbox attempt, in order
|
||||
wantEnd []string // status per attempt, in order
|
||||
}{
|
||||
{"sev3 falls through to ntfy", loop.Sev3,
|
||||
[]string{"voice", "ntfy"}, []string{store.DeliveryFailed, store.DeliverySent}},
|
||||
{"sev4 falls through to telegram", loop.Sev4,
|
||||
[]string{"voice", "telegram"}, []string{store.DeliveryFailed, store.DeliverySent}},
|
||||
{"sev1 does not fall through", loop.Sev1,
|
||||
[]string{"voice"}, []string{store.DeliveryFailed}},
|
||||
{"sev2 does not fall through", loop.Sev2,
|
||||
[]string{"voice"}, []string{store.DeliveryFailed}},
|
||||
}
|
||||
for _, c := range cases {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
voice := &fakeSink{err: ErrVoiceNoSession}
|
||||
ntfy, telegram := &fakeSink{}, &fakeSink{}
|
||||
ob := &fakeOutbox{}
|
||||
d := NewDispatcher(Config{
|
||||
Voice: voice, Ntfy: ntfy, Telegram: telegram,
|
||||
Ack: newFakeAck(), Outbox: ob,
|
||||
})
|
||||
|
||||
if _, err := d.DispatchNudge(context.Background(), PhrasedNudge{
|
||||
Candidate: candidate("some_rule", c.sev, store.Present),
|
||||
Body: "detail", Summary: "short",
|
||||
}, refNow()); err != nil {
|
||||
t.Fatalf("dispatch: %v", err)
|
||||
}
|
||||
if len(ob.attempts) != len(c.wantAt) {
|
||||
t.Fatalf("want %d outbox attempts, got %d (%+v)", len(c.wantAt), len(ob.attempts), ob.attempts)
|
||||
}
|
||||
for i, a := range ob.attempts {
|
||||
if a.channel != c.wantAt[i] || a.status != c.wantEnd[i] {
|
||||
t.Fatalf("attempt %d: want %s/%s, got %s/%s", i, c.wantAt[i], c.wantEnd[i], a.channel, a.status)
|
||||
}
|
||||
}
|
||||
// care severities must not reach an away channel — that would
|
||||
// defeat the drop rule.
|
||||
if c.sev <= loop.Sev2 && (len(ntfy.sends) != 0 || len(telegram.sends) != 0) {
|
||||
t.Fatalf("care nudge escaped to an away channel: ntfy=%d telegram=%d",
|
||||
len(ntfy.sends), len(telegram.sends))
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// ------------------------- crash between Begin and Complete ------------------
|
||||
|
||||
// openTestStore — a real store on a temp file. The reconciliation promise is a
|
||||
// SQL promise, so a fake would only test the fake.
|
||||
func openTestStore(t *testing.T) *store.Store {
|
||||
t.Helper()
|
||||
st, err := store.Open(context.Background(), filepath.Join(t.TempDir(), "maven.db"))
|
||||
if err != nil {
|
||||
t.Fatalf("open store: %v", err)
|
||||
}
|
||||
t.Cleanup(func() { _ = st.Close() })
|
||||
return st
|
||||
}
|
||||
|
||||
// attemptStatus reads one attempt row back. Returns ok=false when the row is
|
||||
// gone, which would itself be a broken promise (a dropped attempt).
|
||||
func attemptStatus(t *testing.T, st *store.Store, id int64) (status string, completed bool, ok bool) {
|
||||
t.Helper()
|
||||
tx, err := st.DB(context.Background())
|
||||
if err != nil {
|
||||
t.Fatalf("read tx: %v", err)
|
||||
}
|
||||
defer func() { _ = tx.Rollback() }()
|
||||
var completedTS *int64
|
||||
err = tx.QueryRowContext(context.Background(),
|
||||
`SELECT status, completed_ts FROM delivery_attempts WHERE id = ?`, id).Scan(&status, &completedTS)
|
||||
if err != nil {
|
||||
return "", false, false
|
||||
}
|
||||
return status, completedTS != nil, true
|
||||
}
|
||||
|
||||
// TestCrashBetweenBeginAndCompleteBecomesUnknown — simulate the crash window:
|
||||
// Begin lands, the process dies before Complete. Startup reconciliation must
|
||||
// turn that row into "unknown" — neither resent nor dropped, because Maven
|
||||
// cannot know whether the message left the box.
|
||||
func TestCrashBetweenBeginAndCompleteBecomesUnknown(t *testing.T) {
|
||||
st := openTestStore(t)
|
||||
ctx := context.Background()
|
||||
sink := &fakeSink{}
|
||||
|
||||
// the crash: intent recorded, no completion.
|
||||
id, err := st.BeginDeliveryAttempt(ctx, "nudge", "disk_low", 0, "telegram", "hash", refNow())
|
||||
if err != nil {
|
||||
t.Fatalf("begin: %v", err)
|
||||
}
|
||||
if s, _, ok := attemptStatus(t, st, id); !ok || s != store.DeliveryPending {
|
||||
t.Fatalf("before reconcile: want pending, got %q ok=%v", s, ok)
|
||||
}
|
||||
|
||||
// restart.
|
||||
n, err := st.ReconcileStaleDeliveryAttempts(ctx, refNow().Add(time.Minute))
|
||||
if err != nil {
|
||||
t.Fatalf("reconcile: %v", err)
|
||||
}
|
||||
if n != 1 {
|
||||
t.Fatalf("want 1 row reconciled, got %d", n)
|
||||
}
|
||||
s, completed, ok := attemptStatus(t, st, id)
|
||||
if !ok {
|
||||
t.Fatal("reconciliation dropped the row; the promise is it is never dropped")
|
||||
}
|
||||
if s != store.DeliveryUnknown {
|
||||
t.Fatalf("want status unknown, got %q", s)
|
||||
}
|
||||
if !completed {
|
||||
t.Fatal("reconciled row should carry a completed_ts")
|
||||
}
|
||||
// not resent: reconciliation is bookkeeping only, it must never push.
|
||||
if len(sink.sends) != 0 {
|
||||
t.Fatalf("reconciliation must not resend, got %d sends", len(sink.sends))
|
||||
}
|
||||
|
||||
// idempotent: a second restart must not churn the row again.
|
||||
n2, err := st.ReconcileStaleDeliveryAttempts(ctx, refNow().Add(2*time.Minute))
|
||||
if err != nil {
|
||||
t.Fatalf("reconcile again: %v", err)
|
||||
}
|
||||
if n2 != 0 {
|
||||
t.Fatalf("second reconcile should find nothing, got %d", n2)
|
||||
}
|
||||
if s2, _, _ := attemptStatus(t, st, id); s2 != store.DeliveryUnknown {
|
||||
t.Fatalf("unknown must stay unknown, got %q", s2)
|
||||
}
|
||||
}
|
||||
|
||||
// TestUnknownIsNeverResolvedToSentOrFailed — the "never guess" half of the
|
||||
// promise: nothing may quietly turn an unknown into a definite outcome.
|
||||
func TestUnknownIsNeverResolvedToSentOrFailed(t *testing.T) {
|
||||
st := openTestStore(t)
|
||||
ctx := context.Background()
|
||||
|
||||
id, err := st.BeginDeliveryAttempt(ctx, "nudge", "disk_low", 0, "telegram", "hash", refNow())
|
||||
if err != nil {
|
||||
t.Fatalf("begin: %v", err)
|
||||
}
|
||||
if _, err := st.ReconcileStaleDeliveryAttempts(ctx, refNow()); err != nil {
|
||||
t.Fatalf("reconcile: %v", err)
|
||||
}
|
||||
// a late Complete from the old in-flight send must not win.
|
||||
if err := st.CompleteDeliveryAttempt(ctx, id, store.DeliverySent, refNow().Add(time.Minute)); err != nil {
|
||||
t.Fatalf("late complete: %v", err)
|
||||
}
|
||||
if s, _, _ := attemptStatus(t, st, id); s != store.DeliveryUnknown {
|
||||
t.Fatalf("late complete overwrote an unknown outcome: %q", s)
|
||||
}
|
||||
}
|
||||
|
||||
// ------------------------- the boring failure modes --------------------------
|
||||
|
||||
// TestSendTimeoutResolvesTheAttempt — a send that times out is a definite
|
||||
// failure from Maven's side, so the row must not be left pending.
|
||||
func TestSendTimeoutResolvesTheAttempt(t *testing.T) {
|
||||
ctx, cancel := context.WithCancel(context.Background())
|
||||
cancel() // the deadline already blew
|
||||
ob := &fakeOutbox{}
|
||||
d := NewDispatcher(Config{Ntfy: &fakeSink{err: context.DeadlineExceeded}, Outbox: ob})
|
||||
|
||||
if _, err := d.DispatchNudge(ctx, PhrasedNudge{
|
||||
Candidate: candidate("cert_expiring", loop.Sev3, store.Away),
|
||||
Body: "detail", Summary: "short",
|
||||
}, refNow()); err == nil {
|
||||
t.Fatal("want a timeout error to propagate")
|
||||
}
|
||||
if len(ob.attempts) != 1 || ob.attempts[0].status != store.DeliveryFailed {
|
||||
t.Fatalf("timed-out send must close the attempt as failed, got %+v", ob.attempts)
|
||||
}
|
||||
}
|
||||
|
||||
// TestCompleteFailureLeavesRowPendingForReconciliation — if Complete itself
|
||||
// fails, the row stays pending on purpose. That is the correct ambiguous state
|
||||
// and startup reconciliation is what resolves it.
|
||||
func TestCompleteFailureLeavesRowPendingForReconciliation(t *testing.T) {
|
||||
ob := &fakeOutbox{completeErr: errors.New("db busy")}
|
||||
d := NewDispatcher(Config{Ntfy: &fakeSink{}, Outbox: ob})
|
||||
|
||||
if _, err := d.DispatchNudge(context.Background(), PhrasedNudge{
|
||||
Candidate: candidate("cert_expiring", loop.Sev3, store.Away),
|
||||
Body: "detail", Summary: "short",
|
||||
}, refNow()); err != nil {
|
||||
t.Fatalf("a failed outbox complete must not fail the dispatch: %v", err)
|
||||
}
|
||||
if len(ob.attempts) != 1 || ob.attempts[0].status != store.DeliveryPending {
|
||||
t.Fatalf("want the row left pending, got %+v", ob.attempts)
|
||||
}
|
||||
}
|
||||
|
||||
// TestPanicMidSendResolvesTheAttempt — a sink that panics leaves the attempt
|
||||
// pending forever while the process keeps running: the dispatcher has no
|
||||
// recover, and reconciliation only runs at startup. Written to the promise
|
||||
// ("never silently resent or dropped" implies every attempt gets resolved),
|
||||
// skipped because the code does not keep it.
|
||||
func TestPanicMidSendResolvesTheAttempt(t *testing.T) {
|
||||
t.Skip("real gap: dispatcher.go:168 has no recover around Send, so a panicking sink leaves a permanent pending row (reconciliation only runs at startup, cmd/mavend/main.go:330)")
|
||||
|
||||
ob := &fakeOutbox{}
|
||||
d := NewDispatcher(Config{Ntfy: &panicSink{}, Outbox: ob})
|
||||
|
||||
func() {
|
||||
defer func() { _ = recover() }()
|
||||
_, _ = d.DispatchNudge(context.Background(), PhrasedNudge{
|
||||
Candidate: candidate("cert_expiring", loop.Sev3, store.Away),
|
||||
Body: "detail", Summary: "short",
|
||||
}, refNow())
|
||||
}()
|
||||
if len(ob.attempts) != 1 || ob.attempts[0].status == store.DeliveryPending {
|
||||
t.Fatalf("a panic mid-send must still resolve the attempt, got %+v", ob.attempts)
|
||||
}
|
||||
}
|
||||
@@ -1,264 +0,0 @@
|
||||
package delivery
|
||||
|
||||
import (
|
||||
"context"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/loop"
|
||||
"github.com/kami/maven/internal/store"
|
||||
)
|
||||
|
||||
// This file walks every cell of the DESIGN.md § "Delivery / channel routing"
|
||||
// table, once as the pure table and once through the dispatcher, so a change
|
||||
// to either side has to break a named cell.
|
||||
//
|
||||
// present away
|
||||
// sev1-2 (care) voice drop
|
||||
// sev3 (soft) voice ntfy, once
|
||||
// sev4 (hard) voice + ntfy telegram, repeat til ack
|
||||
|
||||
type tableCell struct {
|
||||
name string
|
||||
sev loop.Severity
|
||||
presence store.Bucket
|
||||
want []Channel
|
||||
}
|
||||
|
||||
func allTableCells() []tableCell {
|
||||
return []tableCell{
|
||||
{"sev1 present", loop.Sev1, store.Present, []Channel{ChannelVoice}},
|
||||
{"sev2 present", loop.Sev2, store.Present, []Channel{ChannelVoice}},
|
||||
{"sev3 present", loop.Sev3, store.Present, []Channel{ChannelVoice}},
|
||||
{"sev4 present", loop.Sev4, store.Present, []Channel{ChannelVoice, ChannelNtfy}},
|
||||
{"sev1 away", loop.Sev1, store.Away, []Channel{ChannelDrop}},
|
||||
{"sev2 away", loop.Sev2, store.Away, []Channel{ChannelDrop}},
|
||||
{"sev3 away", loop.Sev3, store.Away, []Channel{ChannelNtfy}},
|
||||
{"sev4 away", loop.Sev4, store.Away, []Channel{ChannelTelegram}},
|
||||
}
|
||||
}
|
||||
|
||||
func sameChannels(got, want []Channel) bool {
|
||||
if len(got) != len(want) {
|
||||
return false
|
||||
}
|
||||
for i := range got {
|
||||
if got[i] != want[i] {
|
||||
return false
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
func TestChannelsForEveryTableCell(t *testing.T) {
|
||||
for _, c := range allTableCells() {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
got := ChannelsFor(c.sev, c.presence)
|
||||
if !sameChannels(got, c.want) {
|
||||
t.Fatalf("%s: want %v, got %v", c.name, c.want, got)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestDispatchNudgeEveryTableCell — the same eight cells end to end: exactly
|
||||
// the wanted channels get a send, and every other channel gets none.
|
||||
func TestDispatchNudgeEveryTableCell(t *testing.T) {
|
||||
for _, c := range allTableCells() {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
voice, ntfy, telegram := &fakeSink{}, &fakeSink{}, &fakeSink{}
|
||||
rec := &fakeNudgeRecorder{}
|
||||
d := NewDispatcher(Config{
|
||||
Voice: voice, Ntfy: ntfy, Telegram: telegram,
|
||||
Ack: newFakeAck(), Nudges: rec,
|
||||
})
|
||||
|
||||
out, err := d.DispatchNudge(context.Background(), PhrasedNudge{
|
||||
Candidate: candidate("some_rule", c.sev, c.presence),
|
||||
Body: "full detail body",
|
||||
Summary: "short form",
|
||||
}, refNow())
|
||||
if err != nil {
|
||||
t.Fatalf("dispatch: %v", err)
|
||||
}
|
||||
|
||||
sent := map[Channel]int{
|
||||
ChannelVoice: len(voice.sends),
|
||||
ChannelNtfy: len(ntfy.sends),
|
||||
ChannelTelegram: len(telegram.sends),
|
||||
}
|
||||
for ch, n := range sent {
|
||||
want := 0
|
||||
for _, w := range c.want {
|
||||
if w == ch {
|
||||
want = 1
|
||||
}
|
||||
}
|
||||
if n != want {
|
||||
t.Fatalf("%s: channel %s got %d sends, want %d", c.name, ch, n, want)
|
||||
}
|
||||
}
|
||||
|
||||
// one dispatch and one nudge row per real (non-drop) channel.
|
||||
wantDispatches := 0
|
||||
for _, w := range c.want {
|
||||
if w != ChannelDrop {
|
||||
wantDispatches++
|
||||
}
|
||||
}
|
||||
if len(out) != wantDispatches {
|
||||
t.Fatalf("%s: want %d dispatches, got %d", c.name, wantDispatches, len(out))
|
||||
}
|
||||
if len(rec.rows) != wantDispatches {
|
||||
t.Fatalf("%s: want %d nudge rows, got %d", c.name, wantDispatches, len(rec.rows))
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestSev3AwayIsNtfyExactlyOnce — "ntfy, once": one send, and nothing on the
|
||||
// sendable asks for a repeat, so the daemon's repeat driver has no reason to
|
||||
// pick it up.
|
||||
func TestSev3AwayIsNtfyExactlyOnce(t *testing.T) {
|
||||
ntfy := &fakeSink{}
|
||||
ack := newFakeAck()
|
||||
d := NewDispatcher(Config{Ntfy: ntfy, Telegram: &fakeSink{}, Ack: ack})
|
||||
|
||||
out, err := d.DispatchNudge(context.Background(), PhrasedNudge{
|
||||
Candidate: candidate("cert_expiring", loop.Sev3, store.Away),
|
||||
Body: "cert detail", Summary: "cert expiring",
|
||||
}, refNow())
|
||||
if err != nil {
|
||||
t.Fatalf("dispatch: %v", err)
|
||||
}
|
||||
if len(ntfy.sends) != 1 {
|
||||
t.Fatalf("sev3 away: want exactly 1 ntfy send, got %d", len(ntfy.sends))
|
||||
}
|
||||
if out[0].Sendable.RepeatUntilAck {
|
||||
t.Fatalf("sev3 away must not repeat til ack")
|
||||
}
|
||||
if _, ok := ack.lastSent["cert_expiring"]; ok {
|
||||
t.Fatalf("sev3 away must not enter the ack/repeat tracker")
|
||||
}
|
||||
}
|
||||
|
||||
// TestSev4AwayRepeatsUntilAcked — "telegram, repeat til ack": the initial send
|
||||
// arms the ack clock, the repeat driver re-sends while un-acked, and an ack
|
||||
// stops it.
|
||||
func TestSev4AwayRepeatsUntilAcked(t *testing.T) {
|
||||
telegram := &fakeSink{}
|
||||
ack := newFakeAck()
|
||||
d := NewDispatcher(Config{Telegram: telegram, Ack: ack})
|
||||
ctx := context.Background()
|
||||
|
||||
if _, err := d.DispatchNudge(ctx, PhrasedNudge{
|
||||
Candidate: candidate("disk_low", loop.Sev4, store.Away),
|
||||
Body: "disk detail", Summary: "disk low on homesrv",
|
||||
}, refNow()); err != nil {
|
||||
t.Fatalf("dispatch: %v", err)
|
||||
}
|
||||
|
||||
// two intervals pass, still un-acked → two more sends.
|
||||
for i := 1; i <= 2; i++ {
|
||||
at := refNow().Add(time.Duration(i) * 10 * time.Minute)
|
||||
if _, err := d.RepeatUnacked(ctx, []string{"disk_low"}, at, 5*time.Minute, "disk detail", "disk low on homesrv"); err != nil {
|
||||
t.Fatalf("repeat %d: %v", i, err)
|
||||
}
|
||||
}
|
||||
if len(telegram.sends) != 3 {
|
||||
t.Fatalf("want 1 initial + 2 repeats = 3 telegram sends, got %d", len(telegram.sends))
|
||||
}
|
||||
|
||||
// acked → no further sends, however long we wait.
|
||||
_ = ack.MarkAcked(ctx, "disk_low")
|
||||
if _, err := d.RepeatUnacked(ctx, []string{"disk_low"}, refNow().Add(time.Hour), 5*time.Minute, "b", "s"); err != nil {
|
||||
t.Fatalf("repeat after ack: %v", err)
|
||||
}
|
||||
if len(telegram.sends) != 3 {
|
||||
t.Fatalf("ack must stop the repeat; got %d sends", len(telegram.sends))
|
||||
}
|
||||
}
|
||||
|
||||
// TestAwayChannelsGetMinimalBody — what leaves the box is the short form, for
|
||||
// every away cell of the table. messageForChannel is the last-mile choice both
|
||||
// away sinks make too.
|
||||
func TestAwayChannelsGetMinimalBody(t *testing.T) {
|
||||
detail := "disk /mnt/hdd1 on homesrv at 97% — 12GB free, biggest offender /var/lib/docker"
|
||||
short := "disk low on homesrv"
|
||||
|
||||
for _, ch := range []Channel{ChannelNtfy, ChannelTelegram} {
|
||||
t.Run(string(ch), func(t *testing.T) {
|
||||
msg := messageForChannel(Sendable{Channel: ch, Body: detail, Summary: short})
|
||||
if msg != short {
|
||||
t.Fatalf("%s message: want %q, got %q", ch, short, msg)
|
||||
}
|
||||
})
|
||||
}
|
||||
if got := messageForChannel(Sendable{Channel: ChannelVoice, Body: detail, Summary: short}); got != detail {
|
||||
t.Fatalf("voice is local and gets the full body, got %q", got)
|
||||
}
|
||||
}
|
||||
|
||||
// TestSev4AwaySendableCarriesNoDetail — DESIGN.md § Delivery: away channels
|
||||
// leave the box, so a sev4-away message must not carry detail beyond the short
|
||||
// form. Today the dispatcher hands the away sink the FULL Body as well as the
|
||||
// Summary (dispatcher.go:153-162 copies pn.Body into every Sendable) and
|
||||
// trusts each sink to pick Summary. That works for the two sinks in-tree, but
|
||||
// the minimal body is not enforced at the dispatcher, so a new away sink that
|
||||
// reads Body exfils by default.
|
||||
func TestSev4AwaySendableCarriesNoDetail(t *testing.T) {
|
||||
t.Skip("not enforced: dispatcher.go:159 puts the full Body on away sendables; minimal body is only enforced per-sink (ntfysink.go:77, telegramsink.go:148)")
|
||||
|
||||
telegram := &fakeSink{}
|
||||
d := NewDispatcher(Config{Telegram: telegram, Ack: newFakeAck()})
|
||||
detail := "disk /mnt/hdd1 at 97%, biggest offender /var/lib/docker"
|
||||
|
||||
if _, err := d.DispatchNudge(context.Background(), PhrasedNudge{
|
||||
Candidate: candidate("disk_low", loop.Sev4, store.Away),
|
||||
Body: detail, Summary: "disk low on homesrv",
|
||||
}, refNow()); err != nil {
|
||||
t.Fatalf("dispatch: %v", err)
|
||||
}
|
||||
if strings.Contains(telegram.sends[0].Body, "/var/lib/docker") {
|
||||
t.Fatalf("away sendable carries detail: %q", telegram.sends[0].Body)
|
||||
}
|
||||
}
|
||||
|
||||
// TestAwayFallsBackToFullBodyWhenSummaryEmpty — the other half of the same
|
||||
// gap: with no Summary, the full body leaves the box. The code chooses that on
|
||||
// purpose ("a terse full message is better than no message",
|
||||
// dispatcher.go:345-357), which contradicts the spec's minimal-body rule.
|
||||
// Written to the spec, skipped because the code disagrees.
|
||||
func TestAwayFallsBackToFullBodyWhenSummaryEmpty(t *testing.T) {
|
||||
t.Skip("by design today: dispatcher.go:356 and ntfysink.go:79 fall back to the full Body when Summary is empty, so detail can leave the box")
|
||||
|
||||
msg := messageForChannel(Sendable{
|
||||
Channel: ChannelNtfy,
|
||||
Body: "internal detail that should never leave the box",
|
||||
})
|
||||
if msg != "" {
|
||||
t.Fatalf("empty summary must not fall back to body, got %q", msg)
|
||||
}
|
||||
}
|
||||
|
||||
// TestCareAwayDropIsRecorded — DESIGN.md's drop is a decision ("a missed water
|
||||
// nudge is noise, a missed backup failure isn't"), so it should be visible
|
||||
// rather than vanish. Today drop is a bare `continue`: no nudge row, no outbox
|
||||
// attempt, no log — nothing an operator can see afterwards.
|
||||
func TestCareAwayDropIsRecorded(t *testing.T) {
|
||||
t.Skip("not implemented: dispatcher.go:149-151 skips a Drop channel with no record; there is no 'dropped' outcome in store/delivery.go:16-21")
|
||||
|
||||
ob := &fakeOutbox{}
|
||||
d := NewDispatcher(Config{Voice: &fakeSink{}, Nudges: &fakeNudgeRecorder{}, Outbox: ob})
|
||||
|
||||
if _, err := d.DispatchNudge(context.Background(), PhrasedNudge{
|
||||
Candidate: candidate("water", loop.Sev1, store.Away),
|
||||
Body: "drink water", Summary: "water",
|
||||
}, refNow()); err != nil {
|
||||
t.Fatalf("dispatch: %v", err)
|
||||
}
|
||||
if len(ob.attempts) != 1 || ob.attempts[0].channel != string(ChannelDrop) {
|
||||
t.Fatalf("care-away drop should leave a visible record, got %+v", ob.attempts)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,472 @@
|
||||
// Package recalleval is the held-out contract for note recall: can Maven find
|
||||
// the right note again when the user asks for it weeks later?
|
||||
//
|
||||
// Why it sits beside internal/memory rather than inside it: the thing under
|
||||
// test is a whole path, not one function — an embedder (internal/router), a
|
||||
// vector store (internal/memory or internal/store) and the confidence gate the
|
||||
// daemon applies on top (cmd/mavend/recall.go's bestRecall, config's
|
||||
// query_min_score). A _test.go file inside internal/memory could not reach the
|
||||
// persistent store without an import cycle, and testdata is not reachable from
|
||||
// another package's working directory — so the fixture is embedded here and the
|
||||
// scorer takes the store as a factory. Same layout and same reasons as
|
||||
// internal/router/eval.
|
||||
//
|
||||
// The fixture is HELD OUT the same way the routing fixture is: a query never
|
||||
// repeats its note's wording verbatim beyond ordinary shared vocabulary, and
|
||||
// TestFixtureIsParaphrased enforces a floor on how little the two overlap.
|
||||
// Scoring recall on a query that is a copy of the note measures string
|
||||
// matching, not recall.
|
||||
package recalleval
|
||||
|
||||
import (
|
||||
"context"
|
||||
_ "embed"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"sort"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/memory"
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
|
||||
//go:embed ru_recall_v1.json
|
||||
var fixtureJSON []byte
|
||||
|
||||
// SchemaVersion — the version this package understands. The loader refuses any
|
||||
// other version rather than misreading a fixture and reporting a number.
|
||||
const SchemaVersion = 1
|
||||
|
||||
// StoredNote — one thing the user said once, as it lands in the semantic store.
|
||||
// Kind is "note" or "fact"; both share the vector index (see
|
||||
// cmd/mavend/recall.go), so a fact can legitimately win a recall.
|
||||
type StoredNote struct {
|
||||
ID string `json:"id"`
|
||||
Text string `json:"text"`
|
||||
Kind string `json:"kind"`
|
||||
}
|
||||
|
||||
// Case — a small set of notes, one query, and the note that must come back
|
||||
// first. Want is empty exactly when the query should recall NOTHING: that lane
|
||||
// measures false recall, which is the direction the spec cares about ("a
|
||||
// confident wrong fact is worse than a known gap").
|
||||
type Case struct {
|
||||
ID string `json:"id"`
|
||||
Lang string `json:"lang"`
|
||||
Notes []StoredNote `json:"notes"`
|
||||
Query string `json:"query"`
|
||||
Want string `json:"want"`
|
||||
Tags []string `json:"tags"`
|
||||
Note string `json:"note"`
|
||||
}
|
||||
|
||||
// Answerable reports whether the case expects a recall at all.
|
||||
func (c Case) Answerable() bool { return c.Want != "" }
|
||||
|
||||
// Fixture — the versioned envelope, same shape as the routing fixture.
|
||||
//
|
||||
// Filler is inserted into EVERY case's store on top of that case's own notes.
|
||||
// Without it a case with three notes scores recall@3 = 100% by construction,
|
||||
// which measures nothing. A real store holds months of unrelated notes, and the
|
||||
// wanted note has to beat all of them.
|
||||
type Fixture struct {
|
||||
SchemaVersion int `json:"schema_version"`
|
||||
Name string `json:"name"`
|
||||
Notes []string `json:"notes"`
|
||||
Filler []StoredNote `json:"filler"`
|
||||
Cases []Case `json:"cases"`
|
||||
}
|
||||
|
||||
// Load returns the embedded fixture.
|
||||
func Load() (Fixture, error) {
|
||||
var f Fixture
|
||||
if err := json.Unmarshal(fixtureJSON, &f); err != nil {
|
||||
return Fixture{}, fmt.Errorf("parse fixture: %w", err)
|
||||
}
|
||||
if f.SchemaVersion != SchemaVersion {
|
||||
return Fixture{}, fmt.Errorf("fixture schema_version %d, want %d", f.SchemaVersion, SchemaVersion)
|
||||
}
|
||||
if len(f.Cases) == 0 {
|
||||
return Fixture{}, fmt.Errorf("fixture has no cases")
|
||||
}
|
||||
return f, nil
|
||||
}
|
||||
|
||||
// NewStore builds an empty store for one case, plus a function to release it.
|
||||
// A factory rather than a store because every case needs a clean index — notes
|
||||
// from case A must not be visible to case B's query.
|
||||
type NewStore func() (memory.Store, func(), error)
|
||||
|
||||
// InMemory is the NewStore for memory.InMemoryStore — the fallback the daemon
|
||||
// uses when there is no database (cmd/mavend/voice.go:224).
|
||||
func InMemory() (memory.Store, func(), error) {
|
||||
return memory.NewInMemoryStore(), func() {}, nil
|
||||
}
|
||||
|
||||
// Cache wraps an embedder so repeated text is embedded once. The gate sweep
|
||||
// scores the same fixture at nine thresholds, and every case re-inserts the
|
||||
// filler notes — without this the ONNX run spends minutes re-embedding
|
||||
// identical strings. Latency numbers come from the uncached run.
|
||||
func Cache(inner router.Embedder) router.Embedder {
|
||||
return &cachingEmbedder{inner: inner, seen: map[string][]float32{}}
|
||||
}
|
||||
|
||||
type cachingEmbedder struct {
|
||||
inner router.Embedder
|
||||
seen map[string][]float32
|
||||
}
|
||||
|
||||
func (c *cachingEmbedder) Dim() int { return c.inner.Dim() }
|
||||
func (c *cachingEmbedder) Close() error { return nil } // the caller owns inner
|
||||
|
||||
func (c *cachingEmbedder) Embed(ctx context.Context, text string) ([]float32, error) {
|
||||
if v, ok := c.seen[text]; ok {
|
||||
return v, nil
|
||||
}
|
||||
v, err := c.inner.Embed(ctx, text)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
c.seen[text] = v
|
||||
return v, nil
|
||||
}
|
||||
|
||||
// Outcome — one scored case.
|
||||
type Outcome struct {
|
||||
Case Case
|
||||
Hits []memory.Result
|
||||
Err error
|
||||
// Latency is the read path only: embed the query, then Search. Insert time
|
||||
// is excluded because it happens once, weeks earlier.
|
||||
Latency time.Duration
|
||||
// Rank1/Rank3 — the wanted note came back first / in the top three,
|
||||
// ignoring the confidence gate. Ranking is the store's job.
|
||||
Rank1 bool
|
||||
Rank3 bool
|
||||
// Recalled — what the daemon would actually say back: the top hit's text
|
||||
// when it clears the gate. Mirrors bestRecall in cmd/mavend/recall.go.
|
||||
Recalled string
|
||||
// Pass — the wanted note was recalled AND survived the gate; or, for a
|
||||
// no-answer case, nothing was recalled.
|
||||
Pass bool
|
||||
// Tied — the wanted note is on top but shares its score with the next hit,
|
||||
// so the sort decided it, not the embedder. Counted apart from a real hit.
|
||||
Tied bool
|
||||
TopID string
|
||||
TopScor float64
|
||||
Reasons []string
|
||||
}
|
||||
|
||||
// Report — the aggregate. Rank and gate are kept apart on purpose: a note that
|
||||
// ranks first but is silenced by query_min_score is a threshold problem, and a
|
||||
// note that never ranks first is an embedder problem. Those are different fixes.
|
||||
type Report struct {
|
||||
Name string
|
||||
MinScore float64
|
||||
Total int
|
||||
Answerable int
|
||||
Rank1 int
|
||||
Rank3 int
|
||||
// Gated — ranked first but the score was under MinScore, so the daemon
|
||||
// stays silent and answers "не знаю".
|
||||
Gated int
|
||||
// WrongTop — a different note outranked the right one.
|
||||
WrongTop int
|
||||
// Tied — the right note was on top only because of sort order. Not credited
|
||||
// as recall; tracked because it is a distinct failure (the embedder scored
|
||||
// the query and the note the same as everything else).
|
||||
Tied int
|
||||
// NoAnswer / FalseRecall — the cases that must recall nothing, and how many
|
||||
// of them the daemon would answer anyway.
|
||||
NoAnswer int
|
||||
FalseRecall int
|
||||
Errors int
|
||||
Passed int
|
||||
Outcomes []Outcome
|
||||
ByTag map[string]TagStat
|
||||
ByLang map[string]TagStat
|
||||
// CorrectTop / NoAnswerTop — sorted top-1 scores for the answerable cases
|
||||
// where the right note ranked first, and for the no-answer cases. The gap
|
||||
// between these two distributions is what a defensible query_min_score
|
||||
// would have to sit inside; if they overlap, no threshold separates them.
|
||||
CorrectTop []float64
|
||||
NoAnswerTop []float64
|
||||
P50, P95, Max time.Duration
|
||||
}
|
||||
|
||||
// TagStat — passed/total for one slice of the fixture.
|
||||
type TagStat struct{ Passed, Total int }
|
||||
|
||||
// Recall1 — fraction of answerable cases whose wanted note ranked first.
|
||||
func (r Report) Recall1() float64 { return ratio(r.Rank1, r.Answerable) }
|
||||
|
||||
// Recall3 — same, in the top three. The daemon asks for 3 (voice.go), so this
|
||||
// is the ceiling a better gate or a reranker could reach.
|
||||
func (r Report) Recall3() float64 { return ratio(r.Rank3, r.Answerable) }
|
||||
|
||||
// Answered — fraction of answerable cases the daemon would actually answer
|
||||
// correctly, gate included. This is the number the operator experiences.
|
||||
func (r Report) Answered() float64 { return ratio(r.Rank1-r.Gated, r.Answerable) }
|
||||
|
||||
// FalseRecallRate — fraction of the no-answer cases the daemon answers anyway.
|
||||
func (r Report) FalseRecallRate() float64 { return ratio(r.FalseRecall, r.NoAnswer) }
|
||||
|
||||
func ratio(n, d int) float64 {
|
||||
if d == 0 {
|
||||
return 0
|
||||
}
|
||||
return float64(n) / float64(d)
|
||||
}
|
||||
|
||||
// Score runs every case against a fresh store and aggregates. It never fails
|
||||
// the run on an embed or search error: an erroring case scores as a miss and is
|
||||
// counted in Errors, because "the embedder was down" and "the embedder was
|
||||
// wrong" are different numbers.
|
||||
func Score(ctx context.Context, name string, emb router.Embedder, newStore NewStore, minScore float64, f Fixture) (Report, error) {
|
||||
rep := Report{
|
||||
Name: name,
|
||||
MinScore: minScore,
|
||||
Total: len(f.Cases),
|
||||
ByTag: map[string]TagStat{},
|
||||
ByLang: map[string]TagStat{},
|
||||
}
|
||||
lat := make([]time.Duration, 0, len(f.Cases))
|
||||
|
||||
for _, c := range f.Cases {
|
||||
if c.Answerable() {
|
||||
rep.Answerable++
|
||||
} else {
|
||||
rep.NoAnswer++
|
||||
}
|
||||
o, err := scoreCase(ctx, emb, newStore, minScore, c, f.Filler)
|
||||
if err != nil {
|
||||
return Report{}, err
|
||||
}
|
||||
lat = append(lat, o.Latency)
|
||||
|
||||
switch {
|
||||
case o.Err != nil:
|
||||
rep.Errors++
|
||||
case c.Answerable():
|
||||
if o.Rank1 {
|
||||
rep.Rank1++
|
||||
}
|
||||
if o.Rank3 {
|
||||
rep.Rank3++
|
||||
}
|
||||
if o.Rank1 && o.Recalled == "" {
|
||||
rep.Gated++
|
||||
}
|
||||
if o.Tied {
|
||||
rep.Tied++
|
||||
} else if !o.Rank1 {
|
||||
rep.WrongTop++
|
||||
}
|
||||
if o.Rank1 && o.Recalled != "" {
|
||||
rep.CorrectTop = append(rep.CorrectTop, o.TopScor)
|
||||
}
|
||||
default:
|
||||
if o.Recalled != "" {
|
||||
rep.FalseRecall++
|
||||
}
|
||||
rep.NoAnswerTop = append(rep.NoAnswerTop, o.TopScor)
|
||||
}
|
||||
|
||||
if o.Pass {
|
||||
rep.Passed++
|
||||
}
|
||||
bump(rep.ByLang, c.Lang, o.Pass)
|
||||
for _, tag := range c.Tags {
|
||||
bump(rep.ByTag, tag, o.Pass)
|
||||
}
|
||||
rep.Outcomes = append(rep.Outcomes, o)
|
||||
}
|
||||
|
||||
sort.Float64s(rep.CorrectTop)
|
||||
sort.Float64s(rep.NoAnswerTop)
|
||||
sort.Slice(lat, func(i, j int) bool { return lat[i] < lat[j] })
|
||||
rep.P50, rep.P95 = percentile(lat, 0.50), percentile(lat, 0.95)
|
||||
if len(lat) > 0 {
|
||||
rep.Max = lat[len(lat)-1]
|
||||
}
|
||||
return rep, nil
|
||||
}
|
||||
|
||||
// scoreCase inserts the case's notes into a fresh store, then runs the read
|
||||
// path the daemon runs. The returned error is fatal (the harness is broken);
|
||||
// an embedder or store failure on the query lands in Outcome.Err instead.
|
||||
func scoreCase(ctx context.Context, emb router.Embedder, newStore NewStore, minScore float64, c Case, filler []StoredNote) (Outcome, error) {
|
||||
st, release, err := newStore()
|
||||
if err != nil {
|
||||
return Outcome{}, fmt.Errorf("%s: new store: %w", c.ID, err)
|
||||
}
|
||||
defer release()
|
||||
|
||||
all := append(append([]StoredNote(nil), c.Notes...), filler...)
|
||||
for _, n := range all {
|
||||
vec, err := emb.Embed(ctx, n.Text)
|
||||
if err != nil {
|
||||
return Outcome{}, fmt.Errorf("%s: embed note %s: %w", c.ID, n.ID, err)
|
||||
}
|
||||
meta := map[string]string{"text": n.Text, "type": n.Kind}
|
||||
if err := st.Insert(ctx, n.ID, vec, meta); err != nil {
|
||||
return Outcome{}, fmt.Errorf("%s: insert %s: %w", c.ID, n.ID, err)
|
||||
}
|
||||
}
|
||||
|
||||
o := Outcome{Case: c}
|
||||
start := time.Now()
|
||||
qvec, err := emb.Embed(ctx, c.Query)
|
||||
if err != nil {
|
||||
o.Latency = time.Since(start)
|
||||
o.Err = err
|
||||
o.Reasons = []string{fmt.Sprintf("embed query: %v", err)}
|
||||
return o, nil
|
||||
}
|
||||
hits, err := st.Search(ctx, qvec, 3)
|
||||
o.Latency = time.Since(start)
|
||||
if err != nil {
|
||||
o.Err = err
|
||||
o.Reasons = []string{fmt.Sprintf("search: %v", err)}
|
||||
return o, nil
|
||||
}
|
||||
o.Hits = hits
|
||||
|
||||
if len(hits) > 0 {
|
||||
o.TopID, o.TopScor = hits[0].ID, hits[0].Score
|
||||
o.Recalled = bestRecall(hits, minScore)
|
||||
}
|
||||
for i, h := range hits {
|
||||
if h.ID != c.Want {
|
||||
continue
|
||||
}
|
||||
o.Rank3 = true
|
||||
// A tie is not a hit. With a lexical embedder several notes score
|
||||
// exactly 0 against a paraphrased query, and whichever one the sort
|
||||
// happens to leave on top would otherwise be credited as recall.
|
||||
if i == 0 && (len(hits) < 2 || hits[0].Score > hits[1].Score) {
|
||||
o.Rank1 = true
|
||||
}
|
||||
if i == 0 && !o.Rank1 {
|
||||
o.Tied = true
|
||||
}
|
||||
}
|
||||
|
||||
switch {
|
||||
case !c.Answerable():
|
||||
if o.Recalled != "" {
|
||||
o.Reasons = append(o.Reasons, fmt.Sprintf("false recall: %q at %.3f, want silence", o.TopID, o.TopScor))
|
||||
}
|
||||
case o.Tied:
|
||||
o.Reasons = append(o.Reasons, fmt.Sprintf("tie at %.3f — the right note is on top only by sort order", o.TopScor))
|
||||
case !o.Rank1:
|
||||
o.Reasons = append(o.Reasons, fmt.Sprintf("top hit %q (%.3f), want %q%s", o.TopID, o.TopScor, c.Want, rankNote(o.Rank3)))
|
||||
case o.Recalled == "":
|
||||
o.Reasons = append(o.Reasons, fmt.Sprintf("right note ranked first but scored %.3f < gate %.2f — daemon says \"не знаю\"", o.TopScor, minScore))
|
||||
}
|
||||
o.Pass = len(o.Reasons) == 0
|
||||
return o, nil
|
||||
}
|
||||
|
||||
func rankNote(inTop3 bool) string {
|
||||
if inTop3 {
|
||||
return " (wanted note is in the top 3)"
|
||||
}
|
||||
return " (wanted note is not in the top 3)"
|
||||
}
|
||||
|
||||
// bestRecall mirrors cmd/mavend/recall.go — the gate the daemon actually
|
||||
// applies to a memory hit. Duplicated rather than imported because package main
|
||||
// is not importable; recalleval_test.go asserts the two agree in behaviour.
|
||||
func bestRecall(results []memory.Result, min float64) string {
|
||||
if len(results) == 0 || results[0].Score < min {
|
||||
return ""
|
||||
}
|
||||
return results[0].Meta["text"]
|
||||
}
|
||||
|
||||
func bump(m map[string]TagStat, key string, pass bool) {
|
||||
if key == "" {
|
||||
return
|
||||
}
|
||||
s := m[key]
|
||||
s.Total++
|
||||
if pass {
|
||||
s.Passed++
|
||||
}
|
||||
m[key] = s
|
||||
}
|
||||
|
||||
// percentile — nearest-rank on a pre-sorted slice. No interpolation: with ~30
|
||||
// samples an interpolated p95 invents a latency no query actually took.
|
||||
func percentile(sorted []time.Duration, p float64) time.Duration {
|
||||
if len(sorted) == 0 {
|
||||
return 0
|
||||
}
|
||||
i := int(p * float64(len(sorted)))
|
||||
if i >= len(sorted) {
|
||||
i = len(sorted) - 1
|
||||
}
|
||||
return sorted[i]
|
||||
}
|
||||
|
||||
// String renders the report in the routing eval's style — headline first, then
|
||||
// the slices that name where the path is weak.
|
||||
func (r Report) String() string {
|
||||
var b strings.Builder
|
||||
fmt.Fprintf(&b, "%s: %d/%d cases pass (gate %.2f)\n", r.Name, r.Passed, r.Total, r.MinScore)
|
||||
fmt.Fprintf(&b, " recall@1 %.1f%% (%d/%d) recall@3 %.1f%% (%d/%d) answered after gate %.1f%% (%d/%d)\n",
|
||||
100*r.Recall1(), r.Rank1, r.Answerable,
|
||||
100*r.Recall3(), r.Rank3, r.Answerable,
|
||||
100*r.Answered(), r.Rank1-r.Gated, r.Answerable)
|
||||
fmt.Fprintf(&b, " wrong note on top: %d | tie on top (sort order, not recall): %d | silenced by gate: %d | errors: %d\n",
|
||||
r.WrongTop, r.Tied, r.Gated, r.Errors)
|
||||
fmt.Fprintf(&b, " false recall %.1f%% (%d/%d must-be-silent cases answered anyway)\n",
|
||||
100*r.FalseRecallRate(), r.FalseRecall, r.NoAnswer)
|
||||
fmt.Fprintf(&b, " top-1 score, right note first: %s\n", spread(r.CorrectTop))
|
||||
fmt.Fprintf(&b, " top-1 score, must be silent: %s\n", spread(r.NoAnswerTop))
|
||||
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
|
||||
fmt.Fprintf(&b, " by lang: %s\n", renderStats(r.ByLang))
|
||||
fmt.Fprintf(&b, " by tag: %s\n", renderStats(r.ByTag))
|
||||
return b.String()
|
||||
}
|
||||
|
||||
// Failures — per-case detail, sorted by ID so two runs diff cleanly.
|
||||
func (r Report) Failures() string {
|
||||
var b strings.Builder
|
||||
out := append([]Outcome(nil), r.Outcomes...)
|
||||
sort.Slice(out, func(i, j int) bool { return out[i].Case.ID < out[j].Case.ID })
|
||||
for _, o := range out {
|
||||
if o.Pass {
|
||||
continue
|
||||
}
|
||||
fmt.Fprintf(&b, " %s %q: %s\n", o.Case.ID, o.Case.Query, strings.Join(o.Reasons, "; "))
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
|
||||
// spread — min / median / max of a sorted score list. Three numbers is enough
|
||||
// to see whether two distributions overlap, which is the only question a
|
||||
// threshold can answer.
|
||||
func spread(sorted []float64) string {
|
||||
if len(sorted) == 0 {
|
||||
return "n/a"
|
||||
}
|
||||
return fmt.Sprintf("min %.3f median %.3f max %.3f (n=%d)",
|
||||
sorted[0], sorted[len(sorted)/2], sorted[len(sorted)-1], len(sorted))
|
||||
}
|
||||
|
||||
func renderStats(m map[string]TagStat) string {
|
||||
keys := make([]string, 0, len(m))
|
||||
for k := range m {
|
||||
keys = append(keys, k)
|
||||
}
|
||||
sort.Strings(keys)
|
||||
parts := make([]string, 0, len(keys))
|
||||
for _, k := range keys {
|
||||
s := m[k]
|
||||
parts = append(parts, fmt.Sprintf("%s %d/%d", k, s.Passed, s.Total))
|
||||
}
|
||||
return strings.Join(parts, " ")
|
||||
}
|
||||
@@ -0,0 +1,292 @@
|
||||
package recalleval
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
"unicode"
|
||||
|
||||
"github.com/kami/maven/internal/config"
|
||||
"github.com/kami/maven/internal/memory"
|
||||
"github.com/kami/maven/internal/router"
|
||||
"github.com/kami/maven/internal/store"
|
||||
)
|
||||
|
||||
// dim 1024 for the hash embedder: it is bag-of-words, so a narrower space
|
||||
// collides tokens between unrelated notes and would measure the hash.
|
||||
const hashDim = 1024
|
||||
|
||||
func TestLoadFixture(t *testing.T) {
|
||||
f, err := Load()
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
if len(f.Cases) < 25 {
|
||||
t.Errorf("%d cases, want >= 25", len(f.Cases))
|
||||
}
|
||||
// Filler is what stops recall@3 being free: three case notes and a top-3
|
||||
// search would put the wanted note in the top 3 every time.
|
||||
if len(f.Filler) < 10 {
|
||||
t.Errorf("%d filler notes, want >= 10", len(f.Filler))
|
||||
}
|
||||
seen := map[string]bool{}
|
||||
silent, en := 0, 0
|
||||
for _, c := range f.Cases {
|
||||
if c.ID == "" || seen[c.ID] {
|
||||
t.Errorf("case %q: empty or duplicate id", c.ID)
|
||||
}
|
||||
seen[c.ID] = true
|
||||
if c.Lang != "ru" && c.Lang != "en" {
|
||||
t.Errorf("%s: lang %q, want ru|en", c.ID, c.Lang)
|
||||
}
|
||||
if c.Lang == "en" {
|
||||
en++
|
||||
}
|
||||
if strings.TrimSpace(c.Query) == "" {
|
||||
t.Errorf("%s: empty query", c.ID)
|
||||
}
|
||||
// Fewer than three notes and a wrong answer has nowhere to come from,
|
||||
// so recall@1 would be near-free.
|
||||
if len(c.Notes) < 3 {
|
||||
t.Errorf("%s: %d notes, want >= 3", c.ID, len(c.Notes))
|
||||
}
|
||||
ids := map[string]bool{}
|
||||
for _, n := range c.Notes {
|
||||
if n.ID == "" || ids[n.ID] {
|
||||
t.Errorf("%s: note %q empty or duplicate id", c.ID, n.ID)
|
||||
}
|
||||
ids[n.ID] = true
|
||||
if strings.TrimSpace(n.Text) == "" {
|
||||
t.Errorf("%s: note %q empty text", c.ID, n.ID)
|
||||
}
|
||||
if n.Kind != "note" && n.Kind != "fact" {
|
||||
t.Errorf("%s: note %q kind %q, want note|fact", c.ID, n.ID, n.Kind)
|
||||
}
|
||||
}
|
||||
if !c.Answerable() {
|
||||
silent++
|
||||
continue
|
||||
}
|
||||
if !ids[c.Want] {
|
||||
t.Errorf("%s: want %q is not one of the case's notes", c.ID, c.Want)
|
||||
}
|
||||
}
|
||||
// Both lanes need enough cases that a rate means something.
|
||||
if silent < 5 {
|
||||
t.Errorf("%d must-be-silent cases, want >= 5", silent)
|
||||
}
|
||||
if en < 5 {
|
||||
t.Errorf("%d English cases, want >= 5", en)
|
||||
}
|
||||
}
|
||||
|
||||
// TestFixtureIsParaphrased — the fixture's claim to measuring recall at all. If
|
||||
// a query repeats its note's words, cosine over a bag-of-words embedder gets it
|
||||
// for free and the score says nothing about semantic recall. Half the query's
|
||||
// words is the line: some shared vocabulary is natural ("nginx", "чай"), a copy
|
||||
// is the failure.
|
||||
func TestFixtureIsParaphrased(t *testing.T) {
|
||||
f, err := Load()
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
for _, c := range f.Cases {
|
||||
if !c.Answerable() {
|
||||
continue
|
||||
}
|
||||
var want string
|
||||
for _, n := range c.Notes {
|
||||
if n.ID == c.Want {
|
||||
want = n.Text
|
||||
}
|
||||
}
|
||||
q := words(c.Query)
|
||||
if len(q) == 0 {
|
||||
continue
|
||||
}
|
||||
inNote := map[string]bool{}
|
||||
for _, w := range words(want) {
|
||||
inNote[w] = true
|
||||
}
|
||||
shared := 0
|
||||
for _, w := range q {
|
||||
if inNote[w] {
|
||||
shared++
|
||||
}
|
||||
}
|
||||
if frac := float64(shared) / float64(len(q)); frac > 0.5 {
|
||||
t.Errorf("%s: query shares %.0f%% of its words with the note — not a paraphrase\n query: %q\n note: %q",
|
||||
c.ID, 100*frac, c.Query, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// words — lowercased words of two runes or more, matching how the hash
|
||||
// embedder tokenizes.
|
||||
func words(s string) []string {
|
||||
var out []string
|
||||
for _, w := range strings.FieldsFunc(strings.ToLower(s), func(r rune) bool {
|
||||
return !unicode.IsLetter(r) && !unicode.IsDigit(r)
|
||||
}) {
|
||||
if len([]rune(w)) > 1 {
|
||||
out = append(out, w)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// TestBestRecallMatchesDaemon — the harness duplicates bestRecall from
|
||||
// cmd/mavend/recall.go (package main is not importable). This pins the copy to
|
||||
// the original's three rules: no hits, below the gate, or no text ⇒ silence.
|
||||
func TestBestRecallMatchesDaemon(t *testing.T) {
|
||||
if got := bestRecall(nil, 0.55); got != "" {
|
||||
t.Errorf("no hits: got %q, want silence", got)
|
||||
}
|
||||
low := []memory.Result{{ID: "a", Score: 0.4, Meta: map[string]string{"text": "чай"}}}
|
||||
if got := bestRecall(low, 0.55); got != "" {
|
||||
t.Errorf("below gate: got %q, want silence", got)
|
||||
}
|
||||
noText := []memory.Result{{ID: "a", Score: 0.9, Meta: map[string]string{}}}
|
||||
if got := bestRecall(noText, 0.55); got != "" {
|
||||
t.Errorf("no text: got %q, want silence", got)
|
||||
}
|
||||
ok := []memory.Result{{ID: "a", Score: 0.9, Meta: map[string]string{"text": "чай"}}}
|
||||
if got := bestRecall(ok, 0.55); got != "чай" {
|
||||
t.Errorf("above gate: got %q, want %q", got, "чай")
|
||||
}
|
||||
}
|
||||
|
||||
// TestHashRecallBaseline — the CI ratchet. HashEmbedder, so it needs no model
|
||||
// files and is byte-for-byte reproducible.
|
||||
//
|
||||
// It is a floor, not a target. The hash embedder is lexical, so most of this
|
||||
// fixture is unwinnable for it by construction; the number worth moving is
|
||||
// TestONNXRecall's. Never compare a hash-embedder number to an ONNX one.
|
||||
func TestHashRecallBaseline(t *testing.T) {
|
||||
f, err := Load()
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
rep, err := Score(context.Background(), "recall+hash", router.NewHashEmbedder(hashDim), InMemory,
|
||||
config.DefaultQueryMinScore, f)
|
||||
if err != nil {
|
||||
t.Fatalf("Score: %v", err)
|
||||
}
|
||||
t.Log("\n" + rep.String() + rep.Failures())
|
||||
t.Log("\ngate sweep:\n" + sweep(t, router.NewHashEmbedder(hashDim), f))
|
||||
|
||||
// 0.32 sits under the observed 0.360 recall@1.
|
||||
const floorRecall1 = 0.32
|
||||
if rep.Recall1() < floorRecall1 {
|
||||
t.Errorf("recall@1 %.3f below ratchet %.2f — note recall regressed", rep.Recall1(), floorRecall1)
|
||||
}
|
||||
// The dangerous direction, asserted tightly and separately: answering from
|
||||
// the wrong note is worse than a gap. Observed 0 under the hash floor.
|
||||
if rep.FalseRecall > 1 {
|
||||
t.Errorf("%d false recalls, want <= 1:\n%s", rep.FalseRecall, rep.Failures())
|
||||
}
|
||||
}
|
||||
|
||||
// TestPersistentStoreScoresTheSame — the deployed store is sqlite-backed
|
||||
// (store.MemoryStore via st.VectorMemory()), not the in-memory fallback. Its
|
||||
// Search is a separate implementation of the same cosine scan, so it gets its
|
||||
// own run: a divergence here would mean recall quality depends on whether a
|
||||
// database was configured.
|
||||
func TestPersistentStoreScoresTheSame(t *testing.T) {
|
||||
f, err := Load()
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
emb := router.NewHashEmbedder(hashDim)
|
||||
inMem, err := Score(context.Background(), "recall+hash+memory", emb, InMemory, config.DefaultQueryMinScore, f)
|
||||
if err != nil {
|
||||
t.Fatalf("Score in-memory: %v", err)
|
||||
}
|
||||
persistent, err := Score(context.Background(), "recall+hash+sqlite", emb, sqliteStores(t), config.DefaultQueryMinScore, f)
|
||||
if err != nil {
|
||||
t.Fatalf("Score sqlite: %v", err)
|
||||
}
|
||||
t.Log("\n" + persistent.String())
|
||||
if persistent.Rank1 != inMem.Rank1 || persistent.FalseRecall != inMem.FalseRecall {
|
||||
t.Errorf("sqlite recall@1 %d/%d fr %d, in-memory %d/%d fr %d — the two backends disagree",
|
||||
persistent.Rank1, persistent.Answerable, persistent.FalseRecall,
|
||||
inMem.Rank1, inMem.Answerable, inMem.FalseRecall)
|
||||
}
|
||||
}
|
||||
|
||||
// sqliteStores returns a NewStore that hands each case its own plaintext
|
||||
// database file, so cases stay isolated the way they are with InMemory.
|
||||
func sqliteStores(t *testing.T) NewStore {
|
||||
t.Helper()
|
||||
dir := t.TempDir()
|
||||
n := 0
|
||||
return func() (memory.Store, func(), error) {
|
||||
n++
|
||||
st, err := store.Open(context.Background(), filepath.Join(dir, fmt.Sprintf("recall-%d.db", n)))
|
||||
if err != nil {
|
||||
return nil, nil, err
|
||||
}
|
||||
return st.VectorMemory(), func() { _ = st.Close() }, nil
|
||||
}
|
||||
}
|
||||
|
||||
// TestONNXRecall — the number that matters: the multilingual embedder homesrv
|
||||
// actually runs. Opt-in via MAVEN_ONNX_LIB because deps/ is gitignored, exactly
|
||||
// like TestONNXBaseline in internal/router/eval. `make eval-recall` points it at
|
||||
// the vendored runtime.
|
||||
//
|
||||
// Reports rather than asserts. The gate sweep is the point: it prints
|
||||
// answered-vs-false-recall at a range of query_min_score values, so the right
|
||||
// threshold is read off data instead of guessed.
|
||||
func TestONNXRecall(t *testing.T) {
|
||||
lib := os.Getenv("MAVEN_ONNX_LIB")
|
||||
if lib == "" {
|
||||
t.Skip("MAVEN_ONNX_LIB unset — see AGENTS.md § Embedder model for intent routing")
|
||||
}
|
||||
model := filepath.Join("../../..", "models/embedder/model.onnx")
|
||||
tok := filepath.Join("../../..", "models/embedder/tokenizer.json")
|
||||
for _, p := range []string{lib, model, tok} {
|
||||
if _, err := os.Stat(p); err != nil {
|
||||
t.Skipf("missing %s: %v", p, err)
|
||||
}
|
||||
}
|
||||
emb, err := router.NewONNXEmbedder(model, tok, lib)
|
||||
if err != nil {
|
||||
t.Skipf("onnx embedder unavailable: %v", err)
|
||||
}
|
||||
defer emb.Close()
|
||||
|
||||
f, err := Load()
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
rep, err := Score(context.Background(), "recall+onnx", emb, InMemory, config.DefaultQueryMinScore, f)
|
||||
if err != nil {
|
||||
t.Fatalf("Score: %v", err)
|
||||
}
|
||||
t.Log("\n" + rep.String() + rep.Failures())
|
||||
// Cached for the sweep only: the headline run above must pay the real
|
||||
// embedder cost so its latency numbers mean something.
|
||||
t.Log("\ngate sweep:\n" + sweep(t, Cache(emb), f))
|
||||
}
|
||||
|
||||
// sweep scores the fixture at a range of gates and renders one line each. Two
|
||||
// columns matter: how many real questions get answered, and how many made-up
|
||||
// ones get answered anyway. A gate is only defensible if some value keeps the
|
||||
// first high and the second at zero.
|
||||
func sweep(t *testing.T, emb router.Embedder, f Fixture) string {
|
||||
t.Helper()
|
||||
var b strings.Builder
|
||||
for _, gate := range []float64{0.0, 0.30, 0.40, 0.50, 0.55, 0.60, 0.70, 0.80, 0.90} {
|
||||
rep, err := Score(context.Background(), "sweep", emb, InMemory, gate, f)
|
||||
if err != nil {
|
||||
t.Fatalf("sweep at %.2f: %v", gate, err)
|
||||
}
|
||||
fmt.Fprintf(&b, " gate %.2f: answered %d/%d (%.0f%%) false recall %d/%d\n",
|
||||
gate, rep.Rank1-rep.Gated, rep.Answerable, 100*rep.Answered(), rep.FalseRecall, rep.NoAnswer)
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
@@ -0,0 +1,392 @@
|
||||
{
|
||||
"schema_version": 1,
|
||||
"name": "ru_recall_v1",
|
||||
"notes": [
|
||||
"Held-out note-recall fixture. Each case is a fresh semantic store: insert every note, embed the query, take the top 3 — the same read path cmd/mavend/voice.go runs for IntentQuery.",
|
||||
"Queries paraphrase their note on purpose. A query that repeats the note's words measures string matching, not recall. TestFixtureIsParaphrased enforces a ceiling on word overlap.",
|
||||
"want:\"\" means the query must recall NOTHING. Those cases measure false recall — the direction the spec calls out (a confident wrong fact is worse than a known gap).",
|
||||
"The distractor tag marks cases where a second note is plausible and only one is right. The hard tag marks cases with little or no shared vocabulary.",
|
||||
"Content is written for this operator: his preferences, his homelab, things he said once and would expect Maven to remember weeks later."
|
||||
],
|
||||
"filler": [
|
||||
{"id": "f1", "text": "в субботу ходил в баню", "kind": "note"},
|
||||
{"id": "f2", "text": "купил новые кроссовки сорок третьего размера", "kind": "note"},
|
||||
{"id": "f3", "text": "сериал закончился на третьем сезоне", "kind": "note"},
|
||||
{"id": "f4", "text": "сосед сверху делает ремонт", "kind": "note"},
|
||||
{"id": "f5", "text": "билеты в театр брал заранее", "kind": "note"},
|
||||
{"id": "f6", "text": "выучил пару аккордов на гитаре", "kind": "note"},
|
||||
{"id": "f7", "text": "записался к стоматологу", "kind": "note"},
|
||||
{"id": "f8", "text": "поменял лампочку в коридоре", "kind": "note"},
|
||||
{"id": "f9", "text": "погулял вдоль реки", "kind": "note"},
|
||||
{"id": "f10", "text": "the balcony door sticks in winter", "kind": "note"},
|
||||
{"id": "f11", "text": "the neighbour's dog barks at cyclists", "kind": "note"},
|
||||
{"id": "f12", "text": "i finished the book about volcanoes", "kind": "note"}
|
||||
],
|
||||
"cases": [
|
||||
{
|
||||
"id": "ru-pref-001",
|
||||
"lang": "ru",
|
||||
"tags": ["preference", "paraphrase"],
|
||||
"query": "какой кофе мне наливать",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "я пью кофе без сахара", "kind": "note"},
|
||||
{"id": "n2", "text": "по утрам бегаю в парке", "kind": "note"},
|
||||
{"id": "n3", "text": "не люблю громкую музыку", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-pref-002",
|
||||
"lang": "ru",
|
||||
"tags": ["preference", "homelab", "paraphrase", "hard"],
|
||||
"query": "когда запускать резервное копирование",
|
||||
"want": "n1",
|
||||
"note": "The DESIGN.md preference-seam example, phrased as the operator would ask it later.",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "бэкапы лучше делать ночью в три часа", "kind": "note"},
|
||||
{"id": "n2", "text": "обновления ставлю по субботам", "kind": "note"},
|
||||
{"id": "n3", "text": "логи храню месяц", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-home-003",
|
||||
"lang": "ru",
|
||||
"tags": ["homelab", "paraphrase"],
|
||||
"query": "что помогло от мерцания монитора",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "мерцание экрана прошло после обновления драйвера amdgpu", "kind": "note"},
|
||||
{"id": "n2", "text": "вентилятор шумит на полной нагрузке", "kind": "note"},
|
||||
{"id": "n3", "text": "поставил новый ssd в ноутбук", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-home-004",
|
||||
"lang": "ru",
|
||||
"tags": ["homelab", "distractor", "hard"],
|
||||
"query": "адрес домашнего сервера",
|
||||
"want": "n2",
|
||||
"note": "Two notes carry an IP. Only one is the server.",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "роутер живёт на 192.168.1.1", "kind": "note"},
|
||||
{"id": "n2", "text": "домашний сервер на 192.168.1.104", "kind": "note"},
|
||||
{"id": "n3", "text": "принтер подключен по usb", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-home-005",
|
||||
"lang": "ru",
|
||||
"tags": ["homelab"],
|
||||
"query": "где искать настройки nginx",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "конфиг nginx лежит в /etc/nginx/sites-enabled", "kind": "note"},
|
||||
{"id": "n2", "text": "сертификаты обновляет certbot по расписанию", "kind": "note"},
|
||||
{"id": "n3", "text": "порт 8080 занят вебкой", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-pref-006",
|
||||
"lang": "ru",
|
||||
"tags": ["preference", "distractor", "hard"],
|
||||
"query": "что мне нельзя есть",
|
||||
"want": "n1",
|
||||
"note": "The Friday-meat note is a plausible second answer but it is a habit, not a restriction.",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "у меня аллергия на орехи", "kind": "note"},
|
||||
{"id": "n2", "text": "не ем мясо по пятницам", "kind": "note"},
|
||||
{"id": "n3", "text": "люблю острую еду", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-pers-007",
|
||||
"lang": "ru",
|
||||
"tags": ["distractor", "paraphrase"],
|
||||
"query": "когда мамин праздник",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "день рождения мамы четырнадцатого марта", "kind": "note"},
|
||||
{"id": "n2", "text": "у брата день рождения в июле", "kind": "note"},
|
||||
{"id": "n3", "text": "годовщина в сентябре", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-home-008",
|
||||
"lang": "ru",
|
||||
"tags": ["homelab", "hard", "paraphrase"],
|
||||
"query": "чем ускоряется языковая модель",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "модель крутится на встройке через vulkan", "kind": "note"},
|
||||
{"id": "n2", "text": "whisper работает на процессоре", "kind": "note"},
|
||||
{"id": "n3", "text": "голос у piper русский", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-pref-009",
|
||||
"lang": "ru",
|
||||
"tags": ["preference", "distractor"],
|
||||
"query": "во сколько я обычно засыпаю",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "ложусь спать около часа ночи", "kind": "note"},
|
||||
{"id": "n2", "text": "встаю в семь утра", "kind": "note"},
|
||||
{"id": "n3", "text": "днём не сплю", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-home-010",
|
||||
"lang": "ru",
|
||||
"tags": ["homelab", "paraphrase"],
|
||||
"query": "где у меня хранятся пароли",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "пароли держу в keepassxc", "kind": "note"},
|
||||
{"id": "n2", "text": "двухфакторку сделал через totp", "kind": "note"},
|
||||
{"id": "n3", "text": "ssh ключи лежат на юбикее", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-home-011",
|
||||
"lang": "ru",
|
||||
"tags": ["homelab", "hard", "paraphrase"],
|
||||
"query": "из-за чего кончилось место",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "диск забился логами докера в июне", "kind": "note"},
|
||||
{"id": "n2", "text": "рейд собрал из двух дисков", "kind": "note"},
|
||||
{"id": "n3", "text": "бэкап на внешний диск раз в неделю", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-pref-012",
|
||||
"lang": "ru",
|
||||
"tags": ["preference", "distractor"],
|
||||
"query": "какой чай мне нравится",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "чай пью только зелёный", "kind": "note"},
|
||||
{"id": "n2", "text": "кофе пью без сахара", "kind": "note"},
|
||||
{"id": "n3", "text": "воду пью из фильтра", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-silent-013",
|
||||
"lang": "ru",
|
||||
"tags": ["silent"],
|
||||
"query": "какая погода будет в пятницу",
|
||||
"want": "",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "роутер живёт на 192.168.1.1", "kind": "note"},
|
||||
{"id": "n2", "text": "бэкапы лучше делать ночью", "kind": "note"},
|
||||
{"id": "n3", "text": "у меня аллергия на орехи", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-silent-014",
|
||||
"lang": "ru",
|
||||
"tags": ["silent"],
|
||||
"query": "как зовут сестру моего коллеги",
|
||||
"want": "",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "конфиг nginx лежит в /etc/nginx/sites-enabled", "kind": "note"},
|
||||
{"id": "n2", "text": "порт 8080 занят вебкой", "kind": "note"},
|
||||
{"id": "n3", "text": "сертификаты обновляет certbot", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-silent-015",
|
||||
"lang": "ru",
|
||||
"tags": ["silent"],
|
||||
"query": "сколько я заплатил за машину",
|
||||
"want": "",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "чай пью только зелёный", "kind": "note"},
|
||||
{"id": "n2", "text": "ложусь спать около часа ночи", "kind": "note"},
|
||||
{"id": "n3", "text": "не люблю громкую музыку", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-home-016",
|
||||
"lang": "ru",
|
||||
"tags": ["homelab", "paraphrase"],
|
||||
"query": "откуда берётся токен бота",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "токен телеграма лежит в deploy/telegram.env", "kind": "note"},
|
||||
{"id": "n2", "text": "вебхуки не использую, только long-poll", "kind": "note"},
|
||||
{"id": "n3", "text": "уведомления приходят в личку", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-hard-017",
|
||||
"lang": "ru",
|
||||
"tags": ["hard", "paraphrase", "homelab"],
|
||||
"query": "как я восстановил конфиги",
|
||||
"want": "n1",
|
||||
"note": "No shared word between query and note beyond none at all. This is the case a lexical embedder cannot win.",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "после переустановки системы вернул все настройки из git", "kind": "note"},
|
||||
{"id": "n2", "text": "разделы на диске резал вручную", "kind": "note"},
|
||||
{"id": "n3", "text": "загрузчик поставил заново", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-dist-018",
|
||||
"lang": "ru",
|
||||
"tags": ["distractor"],
|
||||
"query": "чем кормить кота",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "кот ест только сухой корм", "kind": "note"},
|
||||
{"id": "n2", "text": "собаке даю мясо", "kind": "note"},
|
||||
{"id": "n3", "text": "рыбок кормлю раз в день", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-home-019",
|
||||
"lang": "ru",
|
||||
"tags": ["homelab", "hard", "paraphrase"],
|
||||
"query": "как контейнер получает доступ к видеокарте",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "docker compose пробрасывает /dev/dri внутрь", "kind": "note"},
|
||||
{"id": "n2", "text": "контейнеры рестартуют сами", "kind": "note"},
|
||||
{"id": "n3", "text": "образы чищу вручную", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-pref-020",
|
||||
"lang": "ru",
|
||||
"tags": ["preference", "distractor"],
|
||||
"query": "когда мне нельзя звонить",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "не звони мне после десяти вечера", "kind": "note"},
|
||||
{"id": "n2", "text": "утром не трогай меня до кофе", "kind": "note"},
|
||||
{"id": "n3", "text": "по выходным не работаю", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "en-pref-021",
|
||||
"lang": "en",
|
||||
"tags": ["preference", "paraphrase"],
|
||||
"query": "which colour scheme do i like",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "i prefer dark theme everywhere", "kind": "note"},
|
||||
{"id": "n2", "text": "font size 14 is fine", "kind": "note"},
|
||||
{"id": "n3", "text": "i use vim keybindings", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "en-home-022",
|
||||
"lang": "en",
|
||||
"tags": ["homelab", "distractor"],
|
||||
"query": "where is the big disk mounted",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "the nas drive is mounted at /mnt/hdd1", "kind": "note"},
|
||||
{"id": "n2", "text": "models live on the ssd", "kind": "note"},
|
||||
{"id": "n3", "text": "backups go to the nas nightly", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "en-silent-023",
|
||||
"lang": "en",
|
||||
"tags": ["silent"],
|
||||
"query": "what is my bank account number",
|
||||
"want": "",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "the nas drive is mounted at /mnt/hdd1", "kind": "note"},
|
||||
{"id": "n2", "text": "i prefer dark theme everywhere", "kind": "note"},
|
||||
{"id": "n3", "text": "the router runs openwrt", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "en-hard-024",
|
||||
"lang": "en",
|
||||
"tags": ["hard", "paraphrase"],
|
||||
"query": "what fixed the screen problem",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "the flicker went away once i swapped the display cable", "kind": "note"},
|
||||
{"id": "n2", "text": "the laptop fan is loud", "kind": "note"},
|
||||
{"id": "n3", "text": "the second monitor is 1440p", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "en-pref-025",
|
||||
"lang": "en",
|
||||
"tags": ["preference", "hard"],
|
||||
"query": "should i be offered wine",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "i do not drink alcohol", "kind": "note"},
|
||||
{"id": "n2", "text": "i skip breakfast", "kind": "note"},
|
||||
{"id": "n3", "text": "i like spicy food", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-home-026",
|
||||
"lang": "ru",
|
||||
"tags": ["homelab", "paraphrase"],
|
||||
"query": "какая модель распознавания речи мне подходит",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "whisper модель small хватает для русского", "kind": "note"},
|
||||
{"id": "n2", "text": "голос ирина звучит лучше остальных", "kind": "note"},
|
||||
{"id": "n3", "text": "слово активации маven", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-fact-027",
|
||||
"lang": "ru",
|
||||
"tags": ["distractor", "hard", "paraphrase"],
|
||||
"query": "когда я последний раз обслуживал машину",
|
||||
"want": "n1",
|
||||
"note": "A fact, not a note — both share the vector index, so a fact can win a recall.",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "последний раз менял масло в мае", "kind": "fact"},
|
||||
{"id": "n2", "text": "шины поменял осенью", "kind": "fact"},
|
||||
{"id": "n3", "text": "страховка до декабря", "kind": "fact"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-pref-028",
|
||||
"lang": "ru",
|
||||
"tags": ["preference", "hard", "paraphrase"],
|
||||
"query": "как мне присылать оповещения",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "терпеть не могу уведомления со звуком", "kind": "note"},
|
||||
{"id": "n2", "text": "вибрацию оставь включённой", "kind": "note"},
|
||||
{"id": "n3", "text": "письма читаю вечером", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-silent-029",
|
||||
"lang": "ru",
|
||||
"tags": ["silent"],
|
||||
"query": "во сколько отходит поезд",
|
||||
"want": "",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "кот ест только сухой корм", "kind": "note"},
|
||||
{"id": "n2", "text": "люблю острую еду", "kind": "note"},
|
||||
{"id": "n3", "text": "пароли держу в keepassxc", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "en-home-030",
|
||||
"lang": "en",
|
||||
"tags": ["homelab", "paraphrase"],
|
||||
"query": "what firmware is on the router",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "the router runs openwrt", "kind": "note"},
|
||||
{"id": "n2", "text": "wifi channel is 6", "kind": "note"},
|
||||
{"id": "n3", "text": "the guest network is off", "kind": "note"}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
Reference in New Issue
Block a user