Compare commits

..

19 Commits

Author SHA1 Message Date
kami 2ad7635501 Merge the address-form eval check 2026-07-31 14:27:16 +04:00
kami 9949b309b1 Don't let a time word blind the third-person check
The check asks whether anyone else was named before "он". Time words
were not stoplisted, so "сегодня он не ел" read "сегодня" as the person
being talked about and passed — which is the recorded break with a word
in front of it, and nudges open with those words constantly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:27:08 +04:00
kami 62d47d28ac Add an eval check for formal and third-person address (#384)
The phrasing run produced two persona breaks that scored clean:
"Приходите… Жду вас" (formal plural) and "Он не ел 11 дней" (talks
about him instead of to him). She is feminine, he is male, and she
speaks to him informally, one to one.

The new `address` check flags the "вы" family, plural imperative
endings, and a third-person "он" with no other subject named earlier in
the message. Like `hisgender` it is a keyword/suffix heuristic, not a
parser, and it prints the word it tripped on so a false alarm is easy to
dismiss. Limits are written out in the comment.

Both recorded strings are pinned as unit tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:25:23 +04:00
kami 3dbf67f8f9 Drop the city time-zone table
The user only ever asks the time in his own zone, so answering other
cities was code kept in step with the weather city list for no gain.
Any named place now gets the honest "local time only" answer that was
already there for unknown cities.

Removes the 22-entry table, the lookup and the embedded tz database.
Closes Vikunja #389 — there is only one city list again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:14:18 +04:00
kami 84ba217892 Say so when the day asked about is out of reach 2026-07-31 14:05:21 +04:00
kami f179ae2fde Merge the system reply fixes 2026-07-31 14:03:42 +04:00
kami d00929ac0b Answer the day the user asked about and the city he named (#388)
replySystem had two arms that PR 30 made reachable, and both answered confidently wrong: the date arm keyword-matched "числ" and always answered today, so "какое число завтра" answered today; the clock arm ignored a named city and answered local time. The date arm now reads the day word through router.ParseCalendarDate (which grew послезавтра/вчера and now cuts the day boundary in the local zone instead of UTC). The clock arm answers the named zone when it resolves offline from the tz database embedded in the binary, and otherwise says plainly that she only knows local time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:02:57 +04:00
kami b6f47fbeb6 Merge the clock and calendar routing rule 2026-07-31 13:53:35 +04:00
kami 2e9b9ec1cf Warn separately when a row has no text to re-embed 2026-07-31 13:52:56 +04:00
kami bfb57c3148 Give the router prompt a rule for clock and calendar questions (#374)
The prompt named seven intents but never said which one a clock or date
question belongs to, so the model guessed: system->query x4 in every eval
run. The rule now says the clock and the calendar date themselves are
system, what is written in the calendar or in memory stays query, and a
time named inside a request is just part of the request.

That split follows what the daemon can answer. Only replySystem owns the
clock and the date formatter, while the agenda is answered from
CalendarEvents inside the query branch.

Also adds one calendar-agenda fixture case so an over-broad system rule
cannot pass unnoticed, and writes up the before/after numbers. The
targeted confusion is gone; the headline accuracy did not move.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:52:00 +04:00
kami d1f6f6355f Merge the vector backfill 2026-07-31 13:51:25 +04:00
kami 92ecb691de Re-embed stored notes and facts after an embedder swap (#378)
The embedder swap left every stored vector in the old model's space, so cosine against a new query vector is noise. Add the one-shot backfill: store.ReembedAll re-embeds every note and fact text with the currently configured embedder (the passage side, which is the side stored text was written with) and rewrites both places a vector lives — the notes table embedding column and the memory_vectors rows.

All of it plus the embedder marker happens in one transaction, so a failure partway changes nothing and writes no marker: re-run it. A run against a DB whose marker already names the current embedder does nothing.

Triggered explicitly with `mavend -reembed`, not automatically on mismatch: ONNX on the laptop CPU makes this minutes of work, and a silent multi-minute stall on boot would look like a hang. The mismatch warning now tells the user to run it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:50:43 +04:00
kami 4282f6b9a9 Warn about vectors written before the marker existed 2026-07-31 13:44:29 +04:00
kami 7bb9f9be06 Merge the embedder marker 2026-07-31 13:42:40 +04:00
kami 1e47eaca5a Record which embedder wrote the stored vectors and warn on a swap (#378)
The embedder moved from paraphrase-multilingual-MiniLM-L12-v2 to
multilingual-e5-small. Both are 384-dimensional, so nothing in the code
noticed: cosine between an old stored vector and a new query vector is
noise, and recall degrades silently.

So the DB now records the embedder that wrote its vectors. One value for
the whole DB (migration #11, a small `meta` key/value table) rather than a
column on every vector row: the backfill re-embeds every note and fact in
one pass, so a per-row marker would hold the same string in every row and
cost a column on two tables for nothing.

The identity comes from the embedder itself via a new optional ID() method
("multilingual-e5-small@384", model file name plus dimension), so pointing
the config at another model changes the string without anyone editing a
constant. mavend logs a loud WARNING at startup naming both the stored and
the configured embedder when they differ.

Detection only — recall behaviour is unchanged. TODO(#378) in
store.CheckEmbedder marks where the backfill will hook in.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:42:07 +04:00
kami 892330eb84 Merge the note recall fix 2026-07-31 13:33:14 +04:00
kami 9a3bcd7c46 Merge the thinking-off measurement 2026-07-31 13:31:31 +04:00
kami 98ee701e03 Let a note win a recall, not only a fact (#373)
The memory pass ran only after the notes-only gate had already rejected
the same note at the same score. Notes and facts share one vector index,
so a note that failed there failed again — the branch could only ever
return a fact.

Now the memory pass runs first: one search over everything Maven
remembers, one gate, and the memory that clearly matches best answers
(a note gets phrased, a fact is read back). The notes-only pass stays
behind it for notes the vector index does not hold. No threshold moved,
so the set of questions answered is unchanged — only which memory
answers them.

Fixture gained two mixed note+fact cases, so the answerable count goes
25 -> 27: hash recall@1 36.0% -> 37.0% (ratchet 0.32 unchanged, comment
updated), e5 recall@1 72.0% -> 70.4%, false recall still 1/5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:30:38 +04:00
kami 04c1088088 Measure thinking off on routing properly — it does not win (#376)
The 67.1% "thinking off" column in ROUTING-EVAL-31-07-2026.md was an
artefact. It came from a hand-rolled HTTP client in the eval test that
did not send repeat_penalty, so it differed from the reference run on two
axes and the penalty was the one that mattered.

Re-scored back to back on an idle box with everything else held equal:
thinking off is identical to thinking on, case for case, same confusion
matrix, same three unparseable replies. A direct probe of the running
llama-server shows enable_thinking, thinking and reasoning_budget are all
ignored for this model on this build, so there was nothing to turn off.

No defaults changed. The misleading third configuration is removed from
internal/router/eval so its table cannot be quoted again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:30:22 +04:00
27 changed files with 1527 additions and 161 deletions
+8
View File
@@ -75,6 +75,14 @@ same vector. A note is indexed in both places with the same embedding, so if it
`QueryNotes` it fails again here — the branch can only ever return a **fact**. Its comment calls it
"additive"; for notes it is not.
**Fixed (Vikunja #373).** The memory pass now runs *first*, as one search over notes and facts with
one gate, so whichever memory is clearly the best match answers — note or fact. The notes-only pass
stays behind it for notes the vector index does not hold. No threshold changed, so the set of
questions Maven answers is the same; only which memory answers them. The fixture gained two mixed
note+fact cases (`ru-mixed-031`, `ru-mixed-032`), which is why the counts below are out of 27
answerable cases and not 25: hash recall@1 36.0% (9/25) → 37.0% (10/27), e5 recall@1 72.0% (18/25) →
70.4% (19/27) with answered-after-gate 68.0% → 66.7% and false recall unchanged at 1/5.
### 5. Ranking has no recency or type signal, and the store is not the bottleneck
`internal/store/notes.go:67` sorts by cosine and uses `ts` only to break an exact float tie, which
+108 -10
View File
@@ -66,15 +66,113 @@ Three things this run settles:
`запиши что…` phrasings toward fact, and that suspicion stands — all five `ru-note-*`
cases now land on fact. Tracked as Vikunja #375.
**Thinking off is the best configuration measured so far**, on both accuracy and latency
(Vikunja #376). That is worth understanding before flipping: routing is a short
classification into a fixed enum with grammar-constrained output, so there is little to
reason about, and the thinking trace mostly gives a small model room to talk itself out of
the right answer. Phrasing is a different job and needs measuring separately.
The `thinking off` column above read as the best configuration measured so far (Vikunja #376).
**It was wrong** — see the controlled re-run below. Ignore that column.
Still `6 / 6` missed clarify — the router has no way to say "I don't know" (Vikunja #359).
That is unchanged by anything here.
## Thinking off — 31-07-2026, controlled re-run (Vikunja #376)
The "thinking off wins by 6 points" observation above **does not hold**. It was a measurement
artefact, and the earlier table's `thinking off` column should be ignored.
The thinking-off variant was scored by a hand-rolled HTTP client living in the test file
instead of `llm.Client`. That copy did not send `repeat_penalty`, which the real router does
send (`routeRepeatPenalty = 1.15`). So the two columns differed on two axes at once, and the
one that mattered was the penalty, not the thinking mode.
Re-measured with everything else held equal — same fixture, same prompt, same grammar, same
sampling, same idle box, the three configurations run back to back and never concurrently:
| | llm-only, thinking on | llm-only, thinking off | cascade+llm |
|---|---|---|---|
| intent-only accuracy | 59.2% (45/76) | 59.2% (45/76) | 61.8% (47/76) |
| full accuracy (intent+slots+gate) | 38.2% (29/76) | 38.2% (29/76) | 57.9% (44/76) |
| route errors | 3 | 3 | 0 |
| grammar violations | 3 (all 3 route errors) | 3 (same 3 cases) | 0 |
| missed clarify | 5 / 6 | 5 / 6 | 5 / 6 |
| p50 latency | 836ms | 920ms | 810ms |
| p95 latency | 1.41s | 2.00s | 1.31s |
Thinking off is not just a tie on the headline numbers — it is identical case for case, with
the same confusion matrix and the same three unparseable replies. The latency difference is
run-to-run noise on one box, and it points the wrong way here.
The reason is simpler than any accuracy argument: **this llama-server build ignores the
request-level thinking switch for this model.** Probed directly against the running server
with `chat_template_kwargs.enable_thinking = false`, `chat_template_kwargs.thinking = false`
and top-level `reasoning_budget = 0` — all three return a byte-identical answer with the
thinking trace still in `reasoning_content`, and the server reports the prompt prefix as
cached, meaning the rendered template did not change. There was never anything being turned
off, which is also why the numbers match exactly.
Nothing was defaulted. `internal/llm` still has no `chat_template_kwargs` field, `VoiceConfig`
has no thinking flag, and `deploy/mavend.json` is unchanged. The misleading third
configuration is removed from `internal/router/eval` so the table it produced cannot be quoted
again.
Two caveats worth saying out loud:
- **The fixture is 76 cases.** A 6-point difference on 76 cases is roughly 4-5 cases and would
not have been worth trusting even if it had reproduced. This one was exactly 0 cases, which
is a much easier call.
- **This is one server build and one checkpoint** (`b9351`, Qwen3.5-0.8B Q4_K_M). If the
#122 checkpoint or a newer llama.cpp does honour the switch, the question reopens — but it
reopens as an unmeasured question, not as a 6-point win.
Phrasing was **not** measured. Whether thinking helps there is still open, and now also blocked
on the same "can we even turn it off" question.
## Clock and calendar rule — 31-07-2026 (Vikunja #374)
`routeSystem` never said whether "который час" or "какое число завтра" are `system` or
`query`, and `system→query ×4` showed up in every run. The rule added says: the clock and the
calendar date themselves are `system`; what is *written in* the calendar or in memory
("что у меня завтра", "какие есть напоминания") stays `query`; and a time named inside a
request ("напомни завтра…") is just a detail of the request, not a reason for `system`.
That split is not a preference. In `cmd/mavend/voice.go` only `replySystem` owns the clock and
the date formatter, so a clock question routed to `query` falls into the embedder + note RAG
and answers "не знаю". The agenda, on the other hand, is answered by `ParseCalendarDate` +
`CalendarEvents` *inside* the `query` branch, so that side has to stay `query`. The rule sits
above the question test because every one of these utterances carries a question word and a
later rule would never be reached.
The fixture is now 77 cases: one calendar-agenda case was added
(`ru-query-019` "что у меня стоит в календаре на послезавтра", intent `query`) specifically so
an over-broad system rule cannot pass unnoticed. The clock/date cases (`ru-sys-001/002/005`,
`en-sys-001`) already existed.
Three runs, same box, back to back, never concurrently:
| | baseline | first rule (too broad) | rule as committed |
|---|---|---|---|
| llm-only intent-only | 59.2% (45/76) | 54.5% (42/77) | 59.7% (46/77) |
| llm-only full | 38.2% | 35.1% | 39.0% |
| llm-only route errors | 3 | 4 | 5 |
| llm-only p50 | 1.09s | 0.91s | 0.93s |
| cascade+llm intent-only | 61.8% (47/76) | 58.4% | 62.3% (48/77) |
| cascade+llm full | 57.9% | 54.5% | 59.7% |
| cascade+llm route errors | 0 | 0 | 0 |
| cascade+llm p50 | 0.91s | 0.80s | 1.04s |
**The targeted bug is fixed and the headline number did not move.** `system→query ×4` is gone
in both LLM configurations — the `time` and `date` tags go from 0/2 and 0/2 to 2/2 and 2/2 —
but the model then over-applies the rule, and `query→system ×5` plus `reminder→system ×2`
appear where they did not exist before. Net accuracy is a wash, inside the noise of a 77-case
fixture.
The first attempt is shown because it is the honest history: it said "спрашивает время, дату
или день недели → system" with no scope, which swept up reminders, and it cost 3-5 points. It
was tightened once, on the reasoning that a rule capturing "напомни завтра в 7" is simply
wrong, and not tuned further. The remaining `query/reminder → system` over-trigger is a new,
separate weakness of the sub-1B model and deserves its own task rather than more prompt
kneading against a held-out fixture.
The rule is kept. It is correct about what the daemon can answer, and the failure it replaces
was silent ("не знаю" to "который час") while the one it introduces is loud.
## Findings
### 1. The resident model does route better — 50.0% vs 36.8%
@@ -145,11 +243,11 @@ Note the grammar's `string ::= "\"" ([^"\\] | "\\" .)* "\""` is unbounded, so no
### 7. Two hypotheses tested and closed
- **Thinking mode is a non-issue.** Qwen3.5's template defaults `thinking = 1`, so
grammar-constrained JSON lands in `reasoning_content` with `content` empty —
`llm.Client`'s fallback handles it. A `thinking off` run scored *identically* (18/76,
48.7%, same p50). `internal/llm` deliberately does **not** grow a `chat_template_kwargs`
field.
- **Thinking mode is a non-issue.** Confirmed twice now, the second time properly — see the
controlled re-run section. Grammar-constrained JSON lands in `reasoning_content` with
`content` empty and `llm.Client`'s fallback handles it; the request-level switch does
nothing on this build. `internal/llm` deliberately does **not** grow a
`chat_template_kwargs` field.
- **Runaway array repetition does not reproduce.** An isolated smoke test with a stripped
grammar emitted `{"intent":"reminder"}` until `MaxTokens`; under the real `routeSystem`
prompt the few-shot examples anchor it to one object. 2 errors in 76, not 76.
+2
View File
@@ -189,7 +189,9 @@ func (l *lockedAPI) MorningStatus(ctx context.Context) ([]ipc.MorningRoutineStat
func run(args []string) error {
cfgPath := flag.String("config", defaultConfigPath(), "path to mavend JSON config")
wrappedKeyPath := flag.String("wrapped-key-file", "", "path to wrapped encryption key blob (enables cold-start unlock)")
reembed := flag.Bool("reembed", false, "re-embed every stored note and fact with the configured embedder, then serve normally (run once after an embedder swap)")
flag.CommandLine.Parse(args)
reembedOnStart = *reembed
cfg, err := config.Load(*cfgPath)
if err != nil {
return err
+158
View File
@@ -0,0 +1,158 @@
package main
import (
"context"
"fmt"
"math"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice"
)
// fixedEmbedder hands back a vector chosen per text, so a test can say exactly
// how close each stored memory is to the question. The real embedders make
// scores that are realistic but not controllable, and this test is about the
// gate, not about the embedder.
type fixedEmbedder struct{ vecs map[string][]float32 }
func (f *fixedEmbedder) Dim() int { return 4 }
func (f *fixedEmbedder) Close() error { return nil }
func (f *fixedEmbedder) Embed(_ context.Context, text string) ([]float32, error) {
v, ok := f.vecs[text]
if !ok {
return nil, fmt.Errorf("fixedEmbedder: no vector for %q", text)
}
return v, nil
}
// scoreVec builds a unit vector whose cosine against the query vector
// (1,0,0,0) is exactly score.
func scoreVec(score float64) []float32 {
rest := math.Sqrt(1 - score*score)
return []float32{float32(score), float32(rest), 0, 0}
}
// recordingPhraser remembers what the query path handed it to phrase, which is
// how the test can tell which pass produced the answer.
type recordingPhraser struct {
*phraser.Stub
notes []string
}
func (r *recordingPhraser) PhraseQuery(ctx context.Context, utterance string, notes []string) (string, error) {
r.notes = notes
return r.Stub.PhraseQuery(ctx, utterance, notes)
}
// recallCase — one stored memory: its text, how close it is to the question,
// whether it is a note or a fact, and whether the notes table holds it too.
type recallCase struct {
text string
score float64
kind string
}
// buildRecallHandler stores the given memories and returns a handler whose
// query path can be run directly. Notes go into BOTH the notes table and the
// vector index, which is what the daemon does (voice.go's IntentNote).
func buildRecallHandler(t *testing.T, question string, mems []recallCase) (*reactiveHandler, *recordingPhraser) {
t.Helper()
ctx := context.Background()
st := newTestStore(t)
emb := &fixedEmbedder{vecs: map[string][]float32{question: {1, 0, 0, 0}}}
mem := memory.NewInMemoryStore()
now := time.Now()
for i, m := range mems {
vec := scoreVec(m.score)
emb.vecs[m.text] = vec
id := fmt.Sprintf("%s:%d", m.kind, i)
if m.kind == "note" {
if _, err := st.WriteNote(ctx, now, m.text, vec, "tap:voice"); err != nil {
t.Fatalf("WriteNote: %v", err)
}
}
if err := mem.Insert(ctx, id, vec, map[string]string{"text": m.text, "type": m.kind}); err != nil {
t.Fatalf("memory insert: %v", err)
}
}
phr := &recordingPhraser{Stub: phraser.NewStub()}
h := &reactiveHandler{
api: ipc.NewStoreAPI(st),
embedder: emb,
replier: voice.NewStubReplier(),
phraser: phr,
now: func() time.Time { return now },
memStore: mem,
dataStore: st,
queryMinScore: 0.55,
queryMinMargin: 0.008,
weatherProvider: nil,
}
return h, phr
}
func askQuery(t *testing.T, h *reactiveHandler, question string) string {
t.Helper()
return h.applyAction(context.Background(), router.Decision{
Intent: router.IntentQuery,
Utterance: question,
})
}
// TestQueryRecallNoteCanWin — the note-recall regression (Vikunja #373). Notes
// and facts share one vector index, and a note that clearly beats everything
// else must be the answer. Before the fix the memory pass only ran after the
// notes-only gate had already rejected the same note at the same score, so only
// a fact could ever come back from it.
func TestQueryRecallNoteCanWin(t *testing.T) {
const q = "где молоко"
t.Run("a clearly best note answers", func(t *testing.T) {
h, phr := buildRecallHandler(t, q, []recallCase{
{text: "молоко стоит в холодильнике", score: 0.90, kind: "note"},
{text: "выучил пару аккордов", score: 0.50, kind: "note"},
})
reply := askQuery(t, h, q)
if want := "вот что я нашла: молоко стоит в холодильнике"; reply != want {
t.Errorf("reply %q, want %q", reply, want)
}
// One text, the winning memory's — the answer came from the memory
// pass, not from handing the phraser every note in the table.
if len(phr.notes) != 1 || phr.notes[0] != "молоко стоит в холодильнике" {
t.Errorf("phraser got %q, want just the recalled note", phr.notes)
}
})
// The other half of "one gate over everything": a fact that matches better
// than the best note now answers, instead of losing to a note that only had
// to beat other notes.
t.Run("the better-matching fact answers", func(t *testing.T) {
h, _ := buildRecallHandler(t, q, []recallCase{
{text: "молоко стоит в холодильнике", score: 0.80, kind: "note"},
{text: "купил молоко в среду", score: 0.95, kind: "fact"},
})
if reply := askQuery(t, h, q); reply != "купил молоко в среду" {
t.Errorf("reply %q, want the fact read back", reply)
}
})
// The gate is untouched: two memories this close mean the embedder cannot
// tell them apart, and silence still beats a coin flip.
t.Run("no clear best stays silent", func(t *testing.T) {
h, _ := buildRecallHandler(t, q, []recallCase{
{text: "молоко стоит в холодильнике", score: 0.860, kind: "note"},
{text: "молоко закончилось", score: 0.858, kind: "note"},
})
if reply := askQuery(t, h, q); reply != "не знаю." {
t.Errorf("reply %q, want silence", reply)
}
})
}
+17 -14
View File
@@ -2,21 +2,24 @@ package main
import "github.com/kami/maven/internal/memory"
// bestRecall is the read side of the long-term memory store: the top hit's
// stored text when it clears the confidence gate. This recalls across BOTH
// notes and facts (facts aren't in the notes table, so this is the only path
// that can answer "when did I last …?" from a captured fact). A note hit here
// is redundant with the notes-RAG path — by design; the two indexes can diverge
// once the backend is swapped for a persistent/external store. ok=false when
// the hit fails the confidence gate (see memory.Confident: an absolute floor
// plus a margin over the runner-up) or carries no text.
func bestRecall(results []memory.Result, minScore, minMargin float64) (string, bool) {
// bestRecall is the read side of the long-term memory store: the top hit when
// it clears the confidence gate. The index holds BOTH notes and facts, and
// either can win — the caller looks at the returned hit's meta["type"] to see
// which. Facts aren't in the notes table, so this is the only path that can
// answer "when did I last …?" from a captured fact.
//
// The whole hit is returned, not just its text, because "which memory answered"
// decides how the answer is said: a note gets phrased in Maven's voice, a fact
// is read back as stored.
//
// ok=false when the hit fails the confidence gate (see memory.Confident: an
// absolute floor plus a margin over the runner-up) or carries no text.
func bestRecall(results []memory.Result, minScore, minMargin float64) (memory.Result, bool) {
if !memory.Confident(results, minScore, minMargin) {
return "", false
return memory.Result{}, false
}
text := results[0].Meta["text"]
if text == "" {
return "", false
if results[0].Meta["text"] == "" {
return memory.Result{}, false
}
return text, true
return results[0], true
}
+21 -2
View File
@@ -39,8 +39,27 @@ func TestBestRecall(t *testing.T) {
if !ok {
t.Fatal("clearing hit not returned")
}
if got != "выпил воды в три часа" {
t.Errorf("wrong text: %q", got)
if got.Meta["text"] != "выпил воды в три часа" {
t.Errorf("wrong text: %q", got.Meta["text"])
}
if got.Meta["type"] != "fact" {
t.Errorf("kind lost: %q", got.Meta["type"])
}
})
// The index holds notes and facts together, so a note has to be able to win
// it — for a long time it could not (Vikunja #373).
t.Run("a note can win", func(t *testing.T) {
res := []memory.Result{
{Score: 0.86, Meta: map[string]string{"text": "молоко в холодильнике", "type": "note"}},
{Score: 0.61, Meta: map[string]string{"text": "выпил воды", "type": "fact"}},
}
got, ok := bestRecall(res, min, margin)
if !ok {
t.Fatal("clearly-best note not returned")
}
if got.Meta["type"] != "note" || got.Meta["text"] != "молоко в холодильнике" {
t.Errorf("got %v, want the note", got.Meta)
}
})
+78
View File
@@ -0,0 +1,78 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/router"
)
// systemHandler — a handler with nothing but a fixed clock, which is all
// replySystem needs.
func systemHandler(now time.Time) *reactiveHandler {
return &reactiveHandler{now: func() time.Time { return now }}
}
// TestReplySystemDateOffset — "какое число завтра" must answer tomorrow's
// date, not today's (Vikunja #388).
func TestReplySystemDateOffset(t *testing.T) {
// Thursday, 30 July 2026.
now := time.Date(2026, 7, 30, 14, 5, 0, 0, time.UTC)
h := systemHandler(now)
cases := []struct{ utterance, want string }{
{"какое сегодня число", "сегодня четверг, 30 июля 2026 года"},
{"какое число", "сегодня четверг, 30 июля 2026 года"},
{"какое число завтра", "завтра пятница, 31 июля 2026 года"},
{"какое число послезавтра", "послезавтра суббота, 1 августа 2026 года"},
{"какое было число вчера", "вчера среда, 29 июля 2026 года"},
}
for _, c := range cases {
got := h.replySystem(context.Background(), router.Decision{Utterance: c.utterance})
if got != c.want {
t.Errorf("replySystem(%q) = %q, want %q", c.utterance, got, c.want)
}
}
}
// A day she cannot work out must not come back as today's date — that is the
// same silent wrong answer #388 was about, one step further out.
func TestReplySystemUnknownDayIsHonest(t *testing.T) {
now := time.Date(2026, 7, 30, 14, 5, 0, 0, time.UTC)
h := systemHandler(now)
for _, u := range []string{
"какое число в пятницу",
"какое число через неделю",
"какое число в понедельник",
} {
got := h.replySystem(context.Background(), router.Decision{Utterance: u})
if got != onlyNearDaysReply {
t.Errorf("replySystem(%q) = %q, want the honest reply", u, got)
}
}
// The days she does know must not be caught by the same guard.
if got := h.replySystem(context.Background(), router.Decision{Utterance: "какое число завтра"}); got == onlyNearDaysReply {
t.Error("завтра was treated as an unknown day")
}
}
// TestReplySystemClockCity — the clock arm must not answer local time for a
// question about another city (Vikunja #388). She keeps one clock, so every
// named place gets the honest "local time only" answer.
func TestReplySystemClockCity(t *testing.T) {
now := time.Date(2026, 7, 30, 12, 0, 0, 0, time.UTC)
h := systemHandler(now)
cases := []struct{ utterance, want string }{
{"который час", "сейчас 12 часов ровно"},
{"который час в киеве", onlyLocalTimeReply},
{"сколько времени в москве", onlyLocalTimeReply},
{"который час в лондоне", onlyLocalTimeReply},
{"который час в бишкеке", onlyLocalTimeReply},
}
for _, c := range cases {
got := h.replySystem(context.Background(), router.Decision{Utterance: c.utterance})
if got != c.want {
t.Errorf("replySystem(%q) = %q, want %q", c.utterance, got, c.want)
}
}
}
+220 -21
View File
@@ -171,6 +171,7 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
emb = router.NewHashEmbedder(1024)
}
w.embedder = emb
checkStoredEmbedder(dataStore, emb)
// ----- tool executor (the enabled act allowlist, store-backed) -----
// Config tools are the declarative bootstrap: seed them into the store as
@@ -751,7 +752,9 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
}
// Calendar questions: "что у меня сегодня?", "планы на завтра?"
if date, ok := router.ParseCalendarDate(dec.Utterance, time.Now()); ok {
// h.now(), not time.Now(): the handler's clock is the injected one, so
// this arm can be tested at a fixed time like the rest.
if date, ok := router.ParseCalendarDate(dec.Utterance, h.now()); ok {
events, err := h.api.CalendarEvents(ctx, date, date.Add(24*time.Hour))
if err != nil {
log.Printf("voice: calendar events: %v", err)
@@ -786,11 +789,43 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
log.Printf("voice: embed query: %v", err)
return "не получилось найти ответ."
}
// Long-term memory first: ONE search over everything Maven remembers
// (notes and facts share this index) and ONE confidence gate, so the
// memory that is clearly the best match answers — a note just as much
// as a fact.
//
// This used to run only after the notes-only gate below had already
// rejected the same note at the same score, which no note could ever
// survive a second time: the branch could only return a fact (#373).
// Order, not the gate, was the bug — the set of questions Maven answers
// is unchanged, only which memory gets to answer them.
if h.memStore != nil {
if hits, herr := h.memStore.Search(ctx, vec, 3); herr == nil {
if hit, ok := bestRecall(hits, h.queryMinScore, h.queryMinMargin); ok {
text := hit.Meta["text"]
// A note is phrased in Maven's voice; a fact is read back
// as it was stored.
if hit.Meta["type"] == "note" {
if reply, perr := h.phraser.PhraseQuery(ctx, dec.Utterance, []string{text}); perr == nil && reply != "" {
return reply
}
}
return text
}
} else {
log.Printf("voice: memory search: %v", herr)
}
}
notes, err := h.api.QueryNotes(ctx, vec, 5)
if err != nil {
log.Printf("voice: query notes: %v", err)
return "не получилось найти ответ."
}
// Notes-only pass, for notes the vector index above does not hold (an
// older note written before it existed). Same gate, notes-only
// candidates.
//
// Confidence gate: below it, say "I don't know" rather than read back
// the least-unrelated note — a confident wrong recall is worse than a
// gap (spec's "not a guesser-of-truth"). Same instinct as the loop's
@@ -802,16 +837,6 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
noteScores[i] = n.Score
}
if !memory.ConfidentScores(noteScores, h.queryMinScore, h.queryMinMargin) {
// Long-term memory recall (notes + facts) before general knowledge:
// the notes table can't answer fact questions, but the memory store
// indexes both. Only runs when notes-RAG already gave up → additive.
if h.memStore != nil {
if hits, herr := h.memStore.Search(ctx, vec, 3); herr == nil {
if text, ok := bestRecall(hits, h.queryMinScore, h.queryMinMargin); ok {
return text
}
}
}
// Try general knowledge from the phraser before giving up
reply, err := h.phraser.PhraseQuery(ctx, dec.Utterance, nil)
if err != nil || reply == "" {
@@ -913,6 +938,101 @@ var ruMonths = []string{
"июля", "августа", "сентября", "октября", "ноября", "декабря",
}
// onlyLocalTimeReply — the honest answer when the user asks the time somewhere
// other than here. She only keeps one clock, and saying so is better than
// naming the wrong city's time.
//
// There used to be a city→time-zone table here. It was removed on purpose: the
// user only ever asks for local time, so the table was a second list of cities
// to keep in step with the weather one for no gain.
const onlyLocalTimeReply = "я знаю только местное время, про другие города пока не скажу."
// notPlaceAfterV — words that follow "в" without naming a place, so
// mentionsUnknownPlace does not mistake them for a city.
var notPlaceAfterV = map[string]bool{
"данный": true, "данную": true, "этот": true, "эту": true,
"котором": true, "какое": true, "какой": true, "который": true,
"общем": true, "точности": true, "курсе": true, "сутках": true,
"часах": true, "минутах": true, "секундах": true, "неделе": true,
}
// mentionsUnknownPlace reports whether the question has a "в <слово>" phrase
// that looks like a place we do not know ("который час в киеве"). Used only to
// pick the honest "local time only" reply instead of answering local time as
// if it were the city's.
func mentionsUnknownPlace(u string) bool {
toks := strings.Fields(u)
for i := 0; i+1 < len(toks); i++ {
if toks[i] != "в" && toks[i] != "во" {
continue
}
next := strings.Trim(toks[i+1], ".,?!")
if next == "" || notPlaceAfterV[next] {
continue
}
// A number after "в" is a clock ("в 5 часов"), not a place.
if _, err := strconv.Atoi(strings.SplitN(next, ":", 2)[0]); err == nil {
continue
}
return true
}
return false
}
// onlyNearDaysReply — she can work out today, tomorrow, the day after and
// yesterday, and nothing further. Said out loud instead of answering today's
// date for a day she did not understand.
const onlyNearDaysReply = "я считаю только сегодня, завтра, послезавтра и вчера — про другие дни пока не скажу."
// dayWords — day references the calendar parser cannot resolve. A weekday name
// or a "через …" phrase means he asked about a specific other day.
var dayWords = []string{
"понедельник", "вторник", "сред", "четверг", "пятниц", "суббот", "воскресен",
"через", "monday", "tuesday", "wednesday", "thursday", "friday", "saturday", "sunday",
}
// mentionsUnknownDay reports whether the question names a day the calendar
// parser could not resolve. Mirror of mentionsUnknownPlace: it exists only to
// pick an honest reply over a confidently wrong one.
//
// Only called after ParseCalendarDate has already failed, so "завтра" and the
// other words it does know never reach here.
func mentionsUnknownDay(u string) bool {
for _, w := range dayWords {
if strings.Contains(u, w) {
return true
}
}
return false
}
// ruClock renders the clock part of the time reply: "15 часов 4 минуты".
func ruClock(t time.Time) string {
h, m := t.Hour(), t.Minute()
hourWord := ruPlural(h, "час", "часа", "часов")
if m == 0 {
return fmt.Sprintf("%d %s ровно", h, hourWord)
}
return fmt.Sprintf("%d %s %d %s", h, hourWord, m, ruPlural(m, "минута", "минуты", "минут"))
}
// dayPrefix names the day relative to now ("завтра", "вчера", …) so the date
// reply opens the way a person would say it.
func dayPrefix(now, day time.Time) string {
base := time.Date(now.Year(), now.Month(), now.Day(), 0, 0, 0, 0, now.Location())
switch int(day.Sub(base).Hours() / 24) {
case -1:
return "вчера"
case 0:
return "сегодня"
case 1:
return "завтра"
case 2:
return "послезавтра"
}
return "это"
}
func ruPlural(n int, one, two, many string) string {
n = n % 100
if n > 10 && n < 20 {
@@ -991,18 +1111,29 @@ func (h *reactiveHandler) replySystem(ctx context.Context, dec router.Decision)
switch {
case strings.Contains(u, "час") || strings.Contains(u, "врем"):
h := now.Hour()
m := now.Minute()
hourWord := ruPlural(h, "час", "часа", "часов")
if m == 0 {
return fmt.Sprintf("сейчас %d %s ровно", h, hourWord)
// "который час в киеве" — she keeps one clock, so any named place gets
// the honest answer. Never local time dressed up as the city's.
if mentionsUnknownPlace(u) {
return onlyLocalTimeReply
}
minWord := ruPlural(m, "минута", "минуты", "минут")
return fmt.Sprintf("сейчас %d %s %d %s", h, hourWord, m, minWord)
return "сейчас " + ruClock(now)
case strings.Contains(u, "день") || strings.Contains(u, "числ"):
dow := ruWeekdays[now.Weekday()]
month := ruMonths[now.Month()-1]
return fmt.Sprintf("сегодня %s, %d %s %d года", dow, now.Day(), month, now.Year())
// "какое число завтра" — answer for the day the user asked about,
// not today. Reuses the router's calendar day-word parser.
day := now
prefix := "сегодня"
if d, ok := router.ParseCalendarDate(u, now); ok {
day = d
prefix = dayPrefix(now, d)
} else if mentionsUnknownDay(u) {
// He named a day she cannot work out ("в пятницу", "через неделю").
// Answering today's date here would be the same silent wrong answer
// this arm was fixed for, so say what she can do instead.
return onlyNearDaysReply
}
dow := ruWeekdays[day.Weekday()]
month := ruMonths[day.Month()-1]
return fmt.Sprintf("%s %s, %d %s %d года", prefix, dow, day.Day(), month, day.Year())
case strings.Contains(u, "кто дома") || strings.Contains(u, "человек дома"):
return "присутствие пока не подключено к голосовому запросу."
case strings.Contains(u, "памят") || strings.Contains(u, "процессор") || strings.Contains(u, "загрузк") || strings.Contains(u, "статус") || strings.Contains(u, "работа") || strings.Contains(u, "сервис") || strings.Contains(u, "диск") || strings.Contains(u, "ip") || strings.Contains(u, "аптайм") || strings.Contains(u, "трафик") || strings.Contains(u, "интернет"):
@@ -1745,3 +1876,71 @@ func jsonStringImpl(s string) string {
b = append(b, '"')
return string(b)
}
// reembedOnStart is the -reembed flag (set in run()). Opt-in on purpose: see
// runReembed.
var reembedOnStart bool
// checkStoredEmbedder compares the embedder we just loaded with the one that
// wrote the vectors already in the DB (Vikunja #378).
//
// The two models we have both make 384-dim vectors, so a size check catches
// nothing: after a swap, recall silently compares vectors from different
// spaces and the scores are noise. So we say it out loud. Recall itself is not
// changed here — the fix is `mavend -reembed`.
func checkStoredEmbedder(dataStore *store.Store, emb router.Embedder) {
if dataStore == nil {
return
}
current := router.EmbedderID(emb)
if reembedOnStart {
runReembed(dataStore, emb, current)
return
}
stored, mismatch, err := dataStore.CheckEmbedder(context.Background(), current)
if err != nil {
log.Printf("voice: embedder marker check failed: %v", err)
return
}
if mismatch {
log.Printf("voice: WARNING embedder MISMATCH — stored vectors were written by %q but the configured embedder is %q; recall scores are noise until the notes and facts are re-embedded — run `mavend -reembed` once (Vikunja #378)", stored, current)
return
}
log.Printf("voice: embedder marker ok (%s)", current)
}
// runReembed is the one-shot backfill behind -reembed.
//
// Why a flag and not automatic on mismatch: the embedder is ONNX on the
// laptop's CPU, so a few thousand notes is minutes of work. Doing that silently
// inside a normal start would look like the daemon hanging on boot. So the user
// runs it once, deliberately, after an embedder swap; the mismatch warning
// above tells them to. It re-embeds, logs what it did, and then the daemon
// carries on serving as usual — no separate binary, no second start needed.
func runReembed(dataStore *store.Store, emb router.Embedder, current string) {
log.Printf("voice: re-embedding stored notes and facts with %s — this can take a few minutes, do not interrupt", current)
res, err := dataStore.ReembedAll(context.Background(), current,
// EmbedPassage, not EmbedQuery: these are stored texts being searched
// FOR, which is the side they were written with.
func(ctx context.Context, text string) ([]float32, error) {
return router.EmbedPassage(ctx, emb, text)
})
if err != nil {
log.Printf("voice: re-embed FAILED, nothing was changed and no marker was written — safe to run again: %v", err)
return
}
if res.Skipped {
log.Printf("voice: re-embed skipped — the stored vectors were already written by %s", current)
return
}
log.Printf("voice: re-embed done — %d notes in the notes table, %d notes and %d facts in the memory index, took %s; stored vectors now belong to %s",
res.Notes, res.MemNotes, res.Facts, res.Took.Round(time.Second), current)
// A row with no text cannot be re-embedded, so its vector is still the old
// model's noise while the marker now says everything is current. Both write
// paths always store the text, so this should be zero — say it loudly
// rather than bury it in the line above if it ever isn't.
if res.NoText > 0 {
log.Printf("voice: WARNING %d stored rows had no text, so their vectors could not be re-embedded and are still noise; they will never match anything useful (Vikunja #378)", res.NoText)
}
}
+2
View File
@@ -422,6 +422,8 @@ func rankNote(inTop3 bool) string {
// bestRecall mirrors cmd/mavend/recall.go — the gate the daemon actually
// applies to a memory hit. Duplicated rather than imported because package main
// is not importable; recalleval_test.go asserts the two agree in behaviour.
// The daemon returns the whole hit (a note and a fact are said differently);
// the harness only scores what came back, so it keeps returning the text.
func bestRecall(results []memory.Result, minScore, minMargin float64) string {
if !memory.Confident(results, minScore, minMargin) {
return ""
@@ -197,7 +197,8 @@ func TestHashRecallBaseline(t *testing.T) {
t.Log("\n" + rep.String() + rep.Failures())
t.Log("\ngate sweep:\n" + sweep(t, router.NewHashEmbedder(hashDim), f))
// 0.32 sits under the observed 0.360 recall@1.
// 0.32 sits under the observed 0.370 recall@1 (was 0.360 over 25 answerable
// cases; the two mixed note+fact cases added with #373 make it 27).
const floorRecall1 = 0.32
if rep.Recall1() < floorRecall1 {
t.Errorf("recall@1 %.3f below ratchet %.2f — note recall regressed", rep.Recall1(), floorRecall1)
@@ -387,6 +387,32 @@
{"id": "n2", "text": "wifi channel is 6", "kind": "note"},
{"id": "n3", "text": "the guest network is off", "kind": "note"}
]
},
{
"id": "ru-mixed-031",
"lang": "ru",
"tags": ["mixed", "paraphrase", "hard"],
"note": "notes and facts in one store and the note is the answer — the daemon indexes both (Vikunja #373)",
"query": "куда я спрятал второй ключ от квартиры",
"want": "n1",
"notes": [
{"id": "n1", "text": "запасной ключ от квартиры лежит в синей коробке на полке", "kind": "note"},
{"id": "x1", "text": "поменял замок в двери двадцатого июня", "kind": "fact"},
{"id": "x2", "text": "отдал ключ соседке в мае", "kind": "fact"}
]
},
{
"id": "ru-mixed-032",
"lang": "ru",
"tags": ["mixed", "distractor"],
"note": "the mirror of ru-mixed-031: the fact answers and the notes are the distractors",
"query": "когда я в последний раз заливал бензин",
"want": "x1",
"notes": [
{"id": "x1", "text": "залил полный бак в четверг вечером", "kind": "fact"},
{"id": "n1", "text": "на заправке у моста дешевле бензин", "kind": "note"},
{"id": "n2", "text": "надо поменять зимние шины", "kind": "note"}
]
}
]
}
@@ -0,0 +1,26 @@
package eval
import "testing"
func TestAddressTimeWordDoesNotBlind(t *testing.T) {
// A nudge that opens with a time word must still be caught. Without the
// time words in the stoplist, "сегодня" was read as the third party.
for _, s := range []string{
"сегодня он не ел 11 дней",
"вчера он не пил воду",
"опять он забыл про таблетки",
} {
if r := checkAddress(s); r.Pass {
t.Errorf("checkAddress(%q) passed, want a third-person failure", s)
}
}
// Still must not fire when a third party really is named.
for _, s := range []string{
"сегодня сервис упал, он не отвечает",
"ты не пил воду четыре часа",
} {
if r := checkAddress(s); !r.Pass {
t.Errorf("checkAddress(%q) failed: %s", s, r.Detail)
}
}
}
+159 -1
View File
@@ -21,10 +21,14 @@ const (
// CheckHisGender — the other half of the persona rule: SHE is feminine, HE
// is male. "ты давно не отдыхала" addresses the operator as a woman.
CheckHisGender = "hisgender"
// CheckAddress — she talks TO him, informally, one to one. Not "вы", not
// "он". See the comment block above checkAddress.
CheckAddress = "address"
)
// CheckNames — report order.
var CheckNames = []string{CheckMood, CheckLang, CheckLength, CheckFeminine, CheckHisGender, CheckCringe, CheckOnTopic}
var CheckNames = []string{CheckMood, CheckLang, CheckLength, CheckFeminine, CheckHisGender, CheckAddress, CheckCringe, CheckOnTopic}
// Result — one check on one message.
type Result struct {
@@ -57,6 +61,7 @@ func RunChecks(c Case, body, mood string) []Result {
checkLength(body),
checkFeminine(body),
checkHisGender(body),
checkAddress(body),
checkCringe(body),
checkOnTopic(c, body),
}
@@ -302,6 +307,159 @@ func prevWord(words []string, i int) string {
return ""
}
// --- how she addresses him ------------------------------------------------
//
// Persona hard constraint: Maven speaks TO him, informally, one to one. The
// phrasing eval produced two breaks of it, and both scored clean:
//
// - "Приходите… Жду вас" — the formal plural. Correct is ты/тебя/тебе and a
// singular imperative ("приходи", "жду тебя").
// - "Он не ел 11 дней" — she talks ABOUT him, in the third person, as if
// reporting to somebody else. Correct is "ты не ел 11 дней".
//
// Like checkHisGender this is a keyword + suffix heuristic, NOT a parser. Every
// hit prints the word it tripped on, so a false alarm is obvious at a glance and
// can be dismissed.
//
// Part 1, formal address. Two signals:
// - the "вы" pronoun family, matched as whole words, so there is nothing to
// exclude — "вы" and "вас" are never anything else.
// - a plural verb ending: -ите/-ете/-йте/-ьте ("приходите", "выпейте",
// "не забудьте", "хотите"). Nouns in the prepositional case share those
// endings ("в интернете", "в свете"), so a word right after a preposition is
// skipped. That is the whole exclusion list, on purpose: a bigger one would
// start swallowing real imperatives.
//
// Part 2, third person. "он" is perfectly fine when the message really is about
// somebody or something else ("сервис упал, он не отвечает"). The way to tell
// them apart: a legitimate third person has an ANTECEDENT — the thing it refers
// to was named earlier in the message. So "он" is only flagged when nothing
// before it in the message could be that thing.
//
// Where this gives up, plainly:
// - it only looks BACKWARD. "Он не отвечает, сервис упал" names the subject
// after the pronoun and is flagged wrongly.
// - any noun earlier in the message counts as an antecedent, even when it is
// not one ("после обеда он не ел" reads as legitimate and is missed). The
// common time words are stoplisted so the usual nudge opening does not
// blind it, but a message with any other noun in front still slips through.
// This is the check's real hole; widening it further would start flagging
// legitimate third-party messages, so it stops here.
// - a message that opens with "ты" and only later slips into "он" is missed,
// because "ты" itself is skipped but the words around it are not.
// - formal address outside these endings (short adjectives, "вашими" style
// forms not listed) is missed.
// addressWordRE also takes Latin words, because "him"/"he" is the same break in
// English.
var addressWordRE = regexp.MustCompile(`[\p{Cyrillic}]+|[a-zA-Z]+|[,.;:!?…—-]`)
// formalPronouns — the "вы" family. Whole-word match, so no false hits.
var formalPronouns = map[string]bool{
"вы": true, "вас": true, "вам": true, "вами": true,
"ваш": true, "ваша": true, "ваше": true, "ваши": true,
"вашего": true, "вашей": true, "вашему": true, "вашим": true,
"вашими": true, "вашу": true,
}
// prepositions — used twice: to skip prepositional-case nouns that look like
// plural verbs, and as words that cannot be what "он" refers to.
var prepositions = map[string]bool{
"в": true, "во": true, "на": true, "о": true, "об": true, "обо": true,
"при": true, "по": true, "за": true, "из": true, "с": true, "со": true,
"к": true, "ко": true, "до": true, "от": true, "у": true, "над": true,
"под": true, "про": true, "без": true, "для": true, "через": true,
}
// pluralVerb reports whether a word looks like a plural/formal verb form:
// "приходите", "выпейте", "забудьте", "хотите".
func pluralVerb(w string) bool {
if len([]rune(w)) < 5 {
return false
}
return strings.HasSuffix(w, "ите") || strings.HasSuffix(w, "ете") ||
strings.HasSuffix(w, "йте") || strings.HasSuffix(w, "ьте")
}
// thirdPersonHim — pronouns that would be talking about him instead of to him.
var thirdPersonHim = map[string]bool{
"он": true, "его": true, "ему": true, "него": true, "нему": true, "ним": true,
"he": true, "him": true, "his": true,
}
// notAnAntecedent — words that cannot be the thing "он" refers to: pronouns,
// particles, conjunctions, adverbs of time. If only these come before "он", the
// message never named a third party and "он" is him.
var notAnAntecedent = map[string]bool{
"не": true, "ни": true, "и": true, "а": true, "но": true, "да": true,
"же": true, "бы": true, "ли": true, "вот": true, "уже": true,
"ещё": true, "еще": true, "тоже": true, "там": true, "тут": true,
"здесь": true, "это": true, "что": true, "как": true, "когда": true,
"чтобы": true, "потому": true, "сейчас": true, "потом": true,
// Time words. A nudge almost always opens with one ("сегодня он не ел"),
// and without them the very next word is read as the person being talked
// about, so the check misses the exact break it was written for.
"сегодня": true, "вчера": true, "завтра": true, "послезавтра": true,
"утром": true, "днём": true, "днем": true, "вечером": true, "ночью": true,
"опять": true, "снова": true, "весь": true, "всю": true, "целый": true,
"я": true, "мне": true, "меня": true, "мной": true, "мы": true, "нас": true,
"ты": true, "тебя": true, "тебе": true, "тобой": true,
"твой": true, "твоя": true, "твоё": true, "твое": true, "твои": true, "твою": true,
}
// looksPastVerb — a past-tense verb needs a subject of its own, so it is not an
// antecedent either. Keeps "сервис упал, он не отвечает" working off "сервис".
func looksPastVerb(w string) bool {
if len([]rune(w)) < 3 {
return false
}
return strings.HasSuffix(w, "л") || strings.HasSuffix(w, "ла") ||
strings.HasSuffix(w, "ло") || strings.HasSuffix(w, "ли")
}
func checkAddress(body string) Result {
words := addressWordRE.FindAllString(strings.ToLower(body), -1)
for i, w := range words {
if formalPronouns[w] {
return Result{CheckAddress, false,
fmt.Sprintf("formal %q — she says ты/тебя/тебе", w)}
}
if pluralVerb(w) && !(i > 0 && prepositions[words[i-1]]) {
return Result{CheckAddress, false,
fmt.Sprintf("plural imperative %q — she uses the singular", w)}
}
}
for i, w := range words {
if !thirdPersonHim[w] {
continue
}
named := false
for j := 0; j < i; j++ {
p := words[j]
if !unicode.Is(unicode.Cyrillic, []rune(p)[0]) && !isLatinWord(p) {
continue // punctuation
}
if notAnAntecedent[p] || prepositions[p] || thirdPersonHim[p] || looksPastVerb(p) {
continue
}
named = true
break
}
if !named {
return Result{CheckAddress, false,
fmt.Sprintf("third person %q with nobody else named — she talks to him, not about him", w)}
}
}
return Result{CheckAddress, true, ""}
}
func isLatinWord(w string) bool {
r := []rune(w)[0]
return (r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z')
}
// --- the cringe checks ---------------------------------------------------
//
// "Think Jarvis without the cringe part". DESIGN.md § Non-goals: "Not a
+35
View File
@@ -75,6 +75,7 @@ func TestStubBaseline(t *testing.T) {
CheckLength: 12,
CheckFeminine: 15,
CheckHisGender: 15,
CheckAddress: 15,
CheckCringe: 15,
CheckOnTopic: 12,
}
@@ -123,6 +124,10 @@ func TestChecksCatchWhatTheyClaim(t *testing.T) {
{"asks how he feels", "как ты себя чувствуешь? попей воды.", CheckCringe},
{"praise", "молодец! теперь попей воды.", CheckCringe},
{"off topic", "пора бы уже что-то сделать.", CheckOnTopic},
// The two recorded persona breaks from the phrasing eval run. Pinned as
// unit tests because an eval run is sampled and may not reproduce them.
{"formal plural", "Приходите… Жду вас", CheckAddress},
{"third person about him", "Он не ел 11 дней", CheckAddress},
}
for _, tc := range cases {
@@ -144,6 +149,36 @@ func TestChecksCatchWhatTheyClaim(t *testing.T) {
}
}
// TestAddressCheck — the address check on its own, so the messages that must NOT
// trip it can be written without also having to satisfy the on-topic check.
func TestAddressCheck(t *testing.T) {
bad := []string{
"Приходите… Жду вас", // the recorded formal-plural break
"Он не ел 11 дней", // the recorded third-person break
"Выпейте воды, пожалуйста.", // plural imperative on its own
"Ваш обед был давно.", // formal possessive
}
for _, body := range bad {
if r := checkAddress(body); r.Pass {
t.Errorf("persona break not caught: %q", body)
} else {
t.Logf("%q -> %s", body, r.Detail)
}
}
good := []string{
"ты не пил воду четыре часа — попей.", // correct informal address
"сервис netdata упал, он не отвечает.", // legitimately about a third party
"я заметила, что зарядка была утром.", // no address at all
"в интернете опять тихо, всё работает.", // "интернете" is a noun, not an imperative
}
for _, body := range good {
if r := checkAddress(body); !r.Pass {
t.Errorf("clean message flagged: %q -> %s", body, r.Detail)
}
}
}
func TestMoodCheckUsesTheEnum(t *testing.T) {
if r := checkMood("cheerful"); r.Pass {
t.Error("mood outside the enum passed")
+23
View File
@@ -2,6 +2,7 @@ package router
import (
"context"
"fmt"
"math"
"unicode"
)
@@ -19,6 +20,24 @@ type Embedder interface {
Close() error
}
// IdentifiedEmbedder — an embedder that can name itself. The name goes into
// the DB next to the vectors it wrote, so a later model swap is caught instead
// of silently returning nonsense scores (Vikunja #378).
type IdentifiedEmbedder interface {
Embedder
ID() string
}
// EmbedderID is the stable string stored alongside the vectors. It comes from
// the embedder itself — nobody hand-types a model name twice — and changes
// whenever the model or its dimension changes.
func EmbedderID(e Embedder) string {
if i, ok := e.(IdentifiedEmbedder); ok {
return i.ID()
}
return fmt.Sprintf("unknown@%d", e.Dim())
}
// AsymmetricEmbedder — an embedder that wants to know whether a text is a
// search query or a stored passage. Recall is asymmetric: a short question
// goes in, a longer note comes out. The e5 family is trained for exactly that
@@ -70,6 +89,10 @@ func NewHashEmbedder(dim int) *HashEmbedder {
func (h *HashEmbedder) Dim() int { return h.dim }
// ID names this embedder for the DB marker. The dimension is part of it
// because a HashEmbedder of another width is a different vector space.
func (h *HashEmbedder) ID() string { return fmt.Sprintf("hash@%d", h.dim) }
func (h *HashEmbedder) Close() error { return nil }
func (h *HashEmbedder) Embed(_ context.Context, text string) ([]float32, error) {
+24
View File
@@ -0,0 +1,24 @@
package router
import "testing"
func TestEmbedderIDFromModelPath(t *testing.T) {
got := modelIDFromPath("/opt/maven/models/embedder/multilingual-e5-small.onnx")
if got != "multilingual-e5-small@384" {
t.Fatalf("modelIDFromPath = %q", got)
}
// A different model file must produce a different id, even at 384 dim.
old := modelIDFromPath("/opt/maven/models/embedder/paraphrase-multilingual-MiniLM-L12-v2.onnx")
if old == got {
t.Fatal("two different models share one id")
}
}
func TestEmbedderIDIncludesDim(t *testing.T) {
if id := EmbedderID(NewHashEmbedder(1024)); id != "hash@1024" {
t.Fatalf("EmbedderID = %q", id)
}
if EmbedderID(NewHashEmbedder(1024)) == EmbedderID(NewHashEmbedder(384)) {
t.Fatal("dimension not part of the id")
}
}
+14 -96
View File
@@ -1,11 +1,8 @@
package eval
import (
"bytes"
"context"
"encoding/json"
"fmt"
"net/http"
"os"
"strings"
"testing"
@@ -29,13 +26,20 @@ import (
// a bake-off across checkpoints (#278, #250) produces tables you can tell
// apart. Point the variable at one server at a time.
//
// Three configurations, because "the LLM router" is ambiguous and the three
// numbers answer different questions:
// Two configurations, because "the LLM router" is ambiguous and the two numbers
// answer different questions:
//
// llm-only — the model alone. Measures the prompt + grammar contract.
// cascade+llm — what #320 would actually ship: stage-0 grammar, then the
// model, then the classifier as the failure floor.
// llm-no-thinking — diagnostic only, not a shippable path (see below).
// llm-only — the model alone. Measures the prompt + grammar contract.
// cascade+llm — what #320 would actually ship: stage-0 grammar, then the
// model, then the classifier as the failure floor.
//
// There used to be a third, "thinking off", which looked 6 points better. It is
// gone: it was measured with a hand-rolled HTTP client that quietly dropped
// repeat_penalty, so the gap was the missing penalty and not the thinking mode.
// Re-measured with everything else held equal, thinking off scores exactly the
// same, case for case — and a direct probe shows this llama-server build ignores
// enable_thinking / reasoning_budget for this model anyway, so there was nothing
// to turn off. Full write-up in ROUTING-EVAL-31-07-2026.md (Vikunja #376).
func TestLLMRouterBaseline(t *testing.T) {
base := os.Getenv("MAVEN_LLM_URL")
if base == "" {
@@ -96,104 +100,18 @@ func TestLLMRouterBaseline(t *testing.T) {
}
t.Log("\n" + repCascade.String() + repCascade.Failures())
// llm-no-thinking: same prompt and grammar with the chat template's
// thinking mode off. Qwen3.5's template defaults thinking=1, so under a
// grammar the constrained JSON lands in reasoning_content with content
// empty — llm.Client's ReasoningContent fallback is what makes the router
// work at all today, by accident rather than design.
//
// MEASURED 2026-07-31: this variant scores identically to as-deployed
// (18/76, 48.7% intent-only, 2 errors, same p50). Thinking mode is a
// non-issue under a grammar — llama.cpp constrains the same token stream
// either way. Kept so the question stays answered instead of being
// re-asked, and so internal/llm does NOT grow a chat_template_kwargs field
// for a problem that does not exist.
repNoThink, err := Score(ctx, "llm-only ("+model+", thinking off) [diagnostic]",
RouterFunc(func(ctx context.Context, u string, now time.Time) (router.Decision, error) {
d, ok, err := router.NewLLMRouter(&noThinkCompleter{base: base, http: &http.Client{Timeout: 60 * time.Second}}).Route(ctx, u, now)
if err != nil {
return d, err
}
if !ok {
return d, fmt.Errorf("llm router declined without an error")
}
return d, nil
}), f)
if err != nil {
t.Fatalf("Score no-thinking: %v", err)
}
t.Log("\n" + repNoThink.String() + repNoThink.Failures())
// Reports rather than asserts — the numbers are inputs to the #320
// decision, and an assertion here would be this test inventing the bar.
// The one thing worth failing on is a harness fault: if every single case
// errors, the run measured infrastructure, not routing, and the report
// must not be mistaken for a score.
for _, rep := range []Report{repLLM, repCascade, repNoThink} {
for _, rep := range []Report{repLLM, repCascade} {
if rep.Errors == rep.Total {
t.Errorf("%s: all %d cases errored — harness fault, not a measurement", rep.Name, rep.Total)
}
}
}
// noThinkCompleter — llm.Client with chat_template_kwargs.enable_thinking
// false. A test-local copy rather than a change to internal/llm: whether the
// daemon should send it is the open question, and answering it here by adding
// the field would prejudge #320.
type noThinkCompleter struct {
base string
http *http.Client
}
func (c *noThinkCompleter) Complete(ctx context.Context, r llm.Req) (string, error) {
payload := map[string]any{
"messages": []map[string]string{
{"role": "system", "content": r.System},
{"role": "user", "content": r.User},
},
"max_tokens": r.MaxTokens,
"temperature": 0,
"grammar": r.Grammar,
"chat_template_kwargs": map[string]any{"enable_thinking": false},
}
b, err := json.Marshal(payload)
if err != nil {
return "", err
}
req, err := http.NewRequestWithContext(ctx, "POST", c.base+"/v1/chat/completions", bytes.NewReader(b))
if err != nil {
return "", err
}
req.Header.Set("Content-Type", "application/json")
resp, err := c.http.Do(req)
if err != nil {
return "", err
}
defer resp.Body.Close()
if resp.StatusCode != 200 {
return "", fmt.Errorf("status %d", resp.StatusCode)
}
var out struct {
Choices []struct {
Message struct {
Content string `json:"content"`
ReasoningContent string `json:"reasoning_content"`
} `json:"message"`
} `json:"choices"`
}
if err := json.NewDecoder(resp.Body).Decode(&out); err != nil {
return "", err
}
if len(out.Choices) == 0 {
return "", fmt.Errorf("no choices")
}
m := out.Choices[0].Message
if m.Content != "" {
return m.Content, nil
}
return m.ReasoningContent, nil
}
func ping(ctx context.Context, c *llm.Client) error {
ctx, cancel := context.WithTimeout(ctx, 90*time.Second)
defer cancel()
+1
View File
@@ -22,6 +22,7 @@
{ "id": "ru-query-011", "utterance": "почему сервер тормозит", "lang": "ru", "intent": "query", "tags": ["homelab", "hard"], "note": "diagnostic question, not a chat opener" },
{ "id": "ru-query-012", "utterance": "какие заметки я оставил про полив", "lang": "ru", "intent": "query", "tags": ["recall"] },
{ "id": "ru-query-013", "utterance": "во сколько у меня встреча", "lang": "ru", "intent": "query", "tags": ["calendar"] },
{ "id": "ru-query-019", "utterance": "что у меня стоит в календаре на послезавтра", "lang": "ru", "intent": "query", "tags": ["calendar", "hard"], "note": "agenda, not the clock: the daemon answers this from CalendarEvents inside the query branch, so the clock/date system rule must not swallow it" },
{ "id": "ru-query-014", "utterance": "я успеваю до дедлайна", "lang": "ru", "intent": "query", "tags": ["hard", "no-question-word"] },
{ "id": "ru-query-015", "utterance": "сколько я прошёл шагов", "lang": "ru", "intent": "query", "tags": ["aggregate"] },
{ "id": "ru-query-016", "utterance": "покажи давление за неделю", "lang": "ru", "intent": "query", "tags": ["hard", "imperative"], "note": "imperative form but a read — must not route to act" },
+23 -8
View File
@@ -44,10 +44,21 @@ ws ::= [ \t\n]*
// Changed again 31-07-2026: added the "unknown" escape hatch so the model can
// admit it cannot route (Vikunja #359).
//
// Changed again 31-07-2026: added the clock/calendar rule (Vikunja #374). The
// prompt never said which side "который час" or "какое число завтра" belong on,
// so the model guessed — `system→query ×4` in every eval run. The rule sits
// above the question test on purpose: these utterances all carry a question
// word, so a later rule would never be reached. The boundary is what the
// daemon can actually answer: only replySystem in cmd/mavend/voice.go owns the
// clock and the calendar formatter, while the agenda ("что у меня завтра") is
// answered inside the query branch, so that side stays query.
//
// The training workspace keeps its own copy of this prompt for relabelling, and
// `llm/check_prompt_parity.py` there compares the two. That copy is in another
// repo and was not touched, so parity will fail until it gets the same edits —
// both the rule reorder and the "unknown" wording (Vikunja #362).
// both the rule reorder and the "unknown" wording (Vikunja #362) — and now the
// clock/calendar rule too. The training workspace is not checked out on this
// box at all, so it could not be updated here; #362 still covers the catch-up.
const routeSystem = `Классифицируй ровно одно сообщение пользователя. Верни ОДИН JSON-массив действий.
Ровно одно намерение: fact, reminder, note, query, act, chat, system.
@@ -56,19 +67,21 @@ const routeSystem = `Классифицируй ровно одно сообще
Классифицируй по цели пользователя. Порядок решения:
1. Хочет напоминание в будущем → reminder
2. Явно просит сохранить информацию → note
3. Задаёт вопрос: есть вопросительное слово (сколько, что, какой, когда, где, кто, почему, как) или знак «?» → query
4. Хочет получить информацию, в том числе о своих же данных → query
5. Утверждает: сообщает или обновляет текущее состояние/событие → fact
6. Просит выполнить работу → act
7. Про ассистента, настройки или память → system
8. Реплика — обрывок или указание на неназванное («это», «то», «потом»), и без него непонятно, что именно нужно сделать → unknown
9. Иначе → chat
3. Спрашивает только «который час» / «какое число» / «какой день недели» — сами часы или календарная дата, без своих данных → system
4. Задаёт вопрос: есть вопросительное слово (сколько, что, какой, когда, где, кто, почему, как) или знак «?» → query
5. Хочет получить информацию, в том числе о своих же данных → query
6. Утверждает: сообщает или обновляет текущее состояние/событиеfact
7. Просит выполнить работу → act
8. Про ассистента, настройки или память → system
9. Реплика — обрывок или указание на неназванное («это», «то», «потом»), и без него непонятно, что именно нужно сделать → unknown
10. Иначе → chat
Различия:
- note — сохранить информацию, без напоминания. text = суть.
- reminder — уведомить позже. text = что напомнить.
- fact — неявное обновление: пользователь сообщает, что что-то в мире изменилось (текущее/изменённое состояние, случившееся событие). key/value.
- unknown — редкий случай. Ставь его, только если в самой реплике нет ни предмета, ни действия. Короткая, простая или незнакомая тема — это не причина для unknown: приветствие и болтовня — это chat, вопрос на любую тему — это query, просьба сделать что-то названное — это act.
- system против query — часы и календарная дата сами по себе (сколько времени, какое число, какой день недели — можно и про завтра, и про другой город) — это system. А что записано в календаре или в памяти («что у меня завтра», «какие есть напоминания») — это query. Если в реплике есть просьба (напомни, запиши, сделай), то названное время — просто деталь просьбы, и это не system.
- query против fact — решает форма реплики, а не тема. Вопрос о состоянии — это query, даже если названо то же самое, что бывает в fact. Только утверждение — это fact.
Примеры:
@@ -82,6 +95,8 @@ const routeSystem = `Классифицируй ровно одно сообще
"что такое docker?" → {"intent":"query","text":"что такое docker"}
"напиши письмо" → {"intent":"act","verb":"написать письмо"}
"очисти память" → {"intent":"system"}
"который час?" → {"intent":"system"}
"какое число завтра?" → {"intent":"system"}
"привет" → {"intent":"chat","text":"привет"}
"сделай это" → {"intent":"unknown"}
"ну это" → {"intent":"unknown"}
+21
View File
@@ -33,6 +33,7 @@ const (
type onnxEmbedder struct {
tokenizer *unigramTokenizer
session *ort.DynamicSession[int64, float32]
id string
}
func NewONNXEmbedder(modelPath, tokenizerPath, libPath string) (*onnxEmbedder, error) {
@@ -58,11 +59,31 @@ func NewONNXEmbedder(modelPath, tokenizerPath, libPath string) (*onnxEmbedder, e
return &onnxEmbedder{
tokenizer: tok,
session: session,
id: modelIDFromPath(modelPath),
}, nil
}
func (e *onnxEmbedder) Dim() int { return embedDim }
// ID names the loaded model for the DB marker (Vikunja #378): the model file's
// own name plus the dimension, so pointing the config at another model changes
// the string on its own.
func (e *onnxEmbedder) ID() string { return e.id }
// modelIDFromPath turns /opt/.../multilingual-e5-small.onnx into
// "multilingual-e5-small@384".
func modelIDFromPath(modelPath string) string {
name := modelPath
if i := strings.LastIndexAny(name, "/\\"); i >= 0 {
name = name[i+1:]
}
name = strings.TrimSuffix(name, ".onnx")
if name == "" {
name = "onnx"
}
return fmt.Sprintf("%s@%d", name, embedDim)
}
// Embed treats the text as a query. The classifier compares one short
// utterance to another short seed phrase, so both sides get the same prefix
// and the comparison stays fair. The recall path must call EmbedQuery and
+22 -8
View File
@@ -449,16 +449,30 @@ func (AnaphoraResolver) Resolve(text string) (ref string, ok bool) {
return "", false
}
// ParseCalendarDate detects RU calendar date words in text and returns the
// resolved time (midnight UTC+0 for "сегодня"/"today", next day for "завтра"/"tomorrow").
// Returns zero time + false if no match.
// ParseCalendarDate detects RU/EN calendar day words in text and returns
// midnight of that day in now's own time zone. Handles "сегодня", "завтра",
// "послезавтра", "вчера" (and the English words). Returns zero time + false
// if no match.
//
// "послезавтра" is checked before "завтра" because it contains it.
func ParseCalendarDate(text string, now time.Time) (time.Time, bool) {
lower := strings.ToLower(text)
if strings.Contains(lower, "сегодня") || strings.Contains(lower, "today") {
return now.Truncate(24 * time.Hour), true
}
if strings.Contains(lower, "завтра") || strings.Contains(lower, "tomorrow") {
return now.Truncate(24 * time.Hour).Add(24 * time.Hour), true
switch {
case strings.Contains(lower, "сегодня") || strings.Contains(lower, "today"):
return midnight(now, 0), true
case strings.Contains(lower, "послезавтра") || strings.Contains(lower, "day after tomorrow"):
return midnight(now, 2), true
case strings.Contains(lower, "завтра") || strings.Contains(lower, "tomorrow"):
return midnight(now, 1), true
case strings.Contains(lower, "вчера") || strings.Contains(lower, "yesterday"):
return midnight(now, -1), true
}
return time.Time{}, false
}
// midnight returns the start of the day that is `days` away from now, in
// now's time zone (now.Truncate(24h) would cut on a UTC boundary instead).
func midnight(now time.Time, days int) time.Time {
y, m, d := now.AddDate(0, 0, days).Date()
return time.Date(y, m, d, 0, 0, 0, 0, now.Location())
}
+3
View File
@@ -45,6 +45,9 @@ func TestParseCalendarDate(t *testing.T) {
{"расписание на завтра", time.Date(2026, 7, 7, 0, 0, 0, 0, time.UTC), true},
{"what's today", time.Date(2026, 7, 6, 0, 0, 0, 0, time.UTC), true},
{"tomorrow plans", time.Date(2026, 7, 7, 0, 0, 0, 0, time.UTC), true},
{"какое число послезавтра", time.Date(2026, 7, 8, 0, 0, 0, 0, time.UTC), true},
{"что было вчера", time.Date(2026, 7, 5, 0, 0, 0, 0, time.UTC), true},
{"yesterday plans", time.Date(2026, 7, 5, 0, 0, 0, 0, time.UTC), true},
{"какая погода", time.Time{}, false},
{"сколько времени", time.Time{}, false},
{"", time.Time{}, false},
+161
View File
@@ -0,0 +1,161 @@
package store
import (
"context"
"encoding/json"
"fmt"
"time"
)
// EmbedFunc embeds one piece of stored text. The caller passes
// router.EmbedPassage — the STORE side of the query/passage asymmetry, which is
// the side every vector in the DB was written with. (Passing the query side
// would put the stored vectors in the wrong half of the space and quietly halve
// recall.) A func instead of an interface keeps this package free of any
// dependency on internal/router.
type EmbedFunc func(ctx context.Context, text string) ([]float32, error)
// BackfillResult is what the re-embed run did, for logging.
type BackfillResult struct {
Skipped bool // marker already matched — nothing to do
Notes int // rows rewritten in the notes table
Facts int // fact rows rewritten in memory_vectors
MemNotes int // note rows rewritten in memory_vectors
NoText int // memory_vectors rows with no text in their meta, left alone
Took time.Duration
}
// ReembedAll rewrites every stored vector with the currently configured
// embedder and then records that embedder as the one that owns the DB.
//
// Both places a vector lives are rewritten in the same pass: the `notes` table
// `embedding` column and the `memory_vectors` rows (notes AND facts). Doing
// only one would leave the two indexes disagreeing, which is worse than leaving
// both stale.
//
// Safe to re-run: if the marker already names the current embedder there is
// nothing to fix, so it returns immediately with Skipped set.
//
// Crash safety: everything — every vector and the marker — happens inside one
// transaction. If anything fails or the process dies partway, the transaction
// rolls back: no vectors changed and no marker written, so the next run does
// the whole job again. The marker is never set unless the full rewrite
// committed.
func (s *Store) ReembedAll(ctx context.Context, currentID string, embed EmbedFunc) (BackfillResult, error) {
start := time.Now()
var res BackfillResult
stored, err := s.Meta(ctx, metaKeyEmbedderID)
if err != nil {
return res, err
}
if stored == currentID {
res.Skipped = true
res.Took = time.Since(start)
return res, nil
}
tx, err := s.db.BeginTx(ctx, nil)
if err != nil {
return res, fmt.Errorf("reembed: begin: %w", err)
}
defer tx.Rollback() // no-op once committed
// ----- notes table -----
type noteRow struct {
id int64
text string
}
var notes []noteRow
rows, err := tx.QueryContext(ctx, `SELECT id, text FROM notes WHERE text != ''`)
if err != nil {
return res, fmt.Errorf("reembed: read notes: %w", err)
}
for rows.Next() {
var n noteRow
if err := rows.Scan(&n.id, &n.text); err != nil {
rows.Close()
return res, fmt.Errorf("reembed: note row: %w", err)
}
notes = append(notes, n)
}
rows.Close()
if err := rows.Err(); err != nil {
return res, fmt.Errorf("reembed: notes: %w", err)
}
for _, n := range notes {
vec, err := embed(ctx, n.text)
if err != nil {
return res, fmt.Errorf("reembed: embed note %d: %w", n.id, err)
}
if _, err := tx.ExecContext(ctx,
`UPDATE notes SET embedding = ? WHERE id = ?`, floatsToBlob(vec), n.id); err != nil {
return res, fmt.Errorf("reembed: write note %d: %w", n.id, err)
}
res.Notes++
}
// ----- memory_vectors (the unified index: notes AND facts) -----
// The text to re-embed is the one carried in the row's meta blob, which is
// exactly the text that was embedded when the row was written.
type vecRow struct {
id, text, kind string
}
var vecs []vecRow
rows, err = tx.QueryContext(ctx, `SELECT id, meta FROM memory_vectors`)
if err != nil {
return res, fmt.Errorf("reembed: read memory vectors: %w", err)
}
for rows.Next() {
var id, metaJSON string
if err := rows.Scan(&id, &metaJSON); err != nil {
rows.Close()
return res, fmt.Errorf("reembed: memory row: %w", err)
}
meta := map[string]string{}
if err := json.Unmarshal([]byte(metaJSON), &meta); err != nil {
rows.Close()
return res, fmt.Errorf("reembed: meta for %q: %w", id, err)
}
if meta["text"] == "" {
res.NoText++
continue
}
vecs = append(vecs, vecRow{id: id, text: meta["text"], kind: meta["type"]})
}
rows.Close()
if err := rows.Err(); err != nil {
return res, fmt.Errorf("reembed: memory vectors: %w", err)
}
for _, v := range vecs {
vec, err := embed(ctx, v.text)
if err != nil {
return res, fmt.Errorf("reembed: embed %q: %w", v.id, err)
}
if _, err := tx.ExecContext(ctx,
`UPDATE memory_vectors SET vec = ? WHERE id = ?`, encodeVec(vec), v.id); err != nil {
return res, fmt.Errorf("reembed: write %q: %w", v.id, err)
}
if v.kind == "fact" {
res.Facts++
} else {
res.MemNotes++
}
}
// Same transaction as the rewrite, on purpose: the marker can only exist if
// every vector above was written.
if _, err := tx.ExecContext(ctx,
`INSERT INTO meta (key, value) VALUES (?,?)
ON CONFLICT(key) DO UPDATE SET value = excluded.value`,
metaKeyEmbedderID, currentID); err != nil {
return res, fmt.Errorf("reembed: write marker: %w", err)
}
if err := tx.Commit(); err != nil {
return res, fmt.Errorf("reembed: commit: %w", err)
}
res.Took = time.Since(start)
return res, nil
}
+171
View File
@@ -0,0 +1,171 @@
package store
import (
"context"
"errors"
"testing"
"time"
)
// markerVec is a recognisable vector: nothing in these tests writes it except
// the backfill, so finding it proves the row really was rewritten.
var markerVec = []float32{9, 9, 9}
func newEmbedder(calls *int) EmbedFunc {
return func(_ context.Context, _ string) ([]float32, error) {
*calls++
return markerVec, nil
}
}
// seedOldVectors puts one note (notes table + unified index) and one fact
// (unified index only) in the DB, both carrying obviously-old vectors.
func seedOldVectors(t *testing.T, s *Store) {
t.Helper()
ctx := context.Background()
old := []float32{0.1, 0.2, 0.3}
id, err := s.WriteNote(ctx, time.Now(), "молоко в холодильнике", old, "voice")
if err != nil {
t.Fatalf("WriteNote: %v", err)
}
mem := s.VectorMemory()
if err := mem.Insert(ctx, "note:1", old, map[string]string{
"type": "note", "text": "молоко в холодильнике",
}); err != nil {
t.Fatalf("Insert note vector: %v", err)
}
if err := mem.Insert(ctx, "fact:water:1", old, map[string]string{
"type": "fact", "text": "я пил воду",
}); err != nil {
t.Fatalf("Insert fact vector: %v", err)
}
_ = id
}
func noteVec(t *testing.T, s *Store) []float32 {
t.Helper()
var blob []byte
if err := s.db.QueryRow(`SELECT embedding FROM notes LIMIT 1`).Scan(&blob); err != nil {
t.Fatalf("read note embedding: %v", err)
}
return blobToFloats(blob)
}
func memVec(t *testing.T, s *Store, id string) []float32 {
t.Helper()
var blob []byte
if err := s.db.QueryRow(`SELECT vec FROM memory_vectors WHERE id = ?`, id).Scan(&blob); err != nil {
t.Fatalf("read memory vector %s: %v", id, err)
}
return decodeVec(blob)
}
func sameVec(a, b []float32) bool {
if len(a) != len(b) {
return false
}
for i := range a {
if a[i] != b[i] {
return false
}
}
return true
}
// The deployed case: old vectors everywhere, no marker. Every vector in both
// places must be rewritten and the marker recorded.
func TestReembedAllRewritesEveryVector(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
seedOldVectors(t, s)
calls := 0
res, err := s.ReembedAll(ctx, "multilingual-e5-small@384", newEmbedder(&calls))
if err != nil {
t.Fatalf("ReembedAll: %v", err)
}
if res.Skipped {
t.Fatal("first run should not skip")
}
if res.Notes != 1 || res.MemNotes != 1 || res.Facts != 1 {
t.Fatalf("counts: notes=%d memNotes=%d facts=%d", res.Notes, res.MemNotes, res.Facts)
}
if calls != 3 {
t.Fatalf("embedder called %d times, want 3", calls)
}
if !sameVec(noteVec(t, s), markerVec) {
t.Fatalf("notes table not rewritten: %v", noteVec(t, s))
}
if !sameVec(memVec(t, s, "note:1"), markerVec) {
t.Fatal("unified index note row not rewritten")
}
if !sameVec(memVec(t, s, "fact:water:1"), markerVec) {
t.Fatal("unified index fact row not rewritten")
}
got, err := s.Meta(ctx, metaKeyEmbedderID)
if err != nil {
t.Fatalf("Meta: %v", err)
}
if got != "multilingual-e5-small@384" {
t.Fatalf("marker = %q", got)
}
}
// Re-running must do nothing at all — not a second pass over the same rows.
func TestReembedAllSecondRunIsNoop(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
seedOldVectors(t, s)
calls := 0
if _, err := s.ReembedAll(ctx, "e5@384", newEmbedder(&calls)); err != nil {
t.Fatalf("first run: %v", err)
}
first := calls
res, err := s.ReembedAll(ctx, "e5@384", newEmbedder(&calls))
if err != nil {
t.Fatalf("second run: %v", err)
}
if !res.Skipped {
t.Fatal("second run should report Skipped")
}
if calls != first {
t.Fatalf("second run embedded %d more rows, want 0", calls-first)
}
}
// A failure partway must leave the DB exactly as it was: no marker, and the old
// vectors still in place (one transaction, rolled back).
func TestReembedAllPartialFailureLeavesMarkerUnset(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
seedOldVectors(t, s)
before := noteVec(t, s)
calls := 0
boom := func(_ context.Context, _ string) ([]float32, error) {
calls++
if calls == 2 {
return nil, errors.New("onnx blew up")
}
return markerVec, nil
}
if _, err := s.ReembedAll(ctx, "e5@384", boom); err == nil {
t.Fatal("expected an error")
}
got, err := s.Meta(ctx, metaKeyEmbedderID)
if err != nil {
t.Fatalf("Meta: %v", err)
}
if got != "" {
t.Fatalf("marker was set to %q after a failed run", got)
}
if !sameVec(noteVec(t, s), before) {
t.Fatal("a failed run left a partially rewritten notes table")
}
// And the mismatch warning must still fire, so the user knows to re-run.
if _, mismatch, err := s.CheckEmbedder(ctx, "e5@384"); err != nil || !mismatch {
t.Fatalf("CheckEmbedder after failed backfill: mismatch=%v err=%v", mismatch, err)
}
}
+97
View File
@@ -0,0 +1,97 @@
package store
import (
"context"
"database/sql"
"errors"
"fmt"
)
// metaKeyEmbedderID names the embedder that wrote the stored vectors.
//
// Why one value for the whole DB and not a column on every vector row: the
// vectors are only ever rewritten all at once (one backfill re-embeds every
// note and fact together), so a per-row marker would hold the same string in
// every row and cost a column on two tables for nothing.
const metaKeyEmbedderID = "embedder_id"
// Meta reads a single value from the meta table. Missing key ⇒ empty string.
func (s *Store) Meta(ctx context.Context, key string) (string, error) {
var v string
err := s.db.QueryRowContext(ctx, `SELECT value FROM meta WHERE key = ?`, key).Scan(&v)
if errors.Is(err, sql.ErrNoRows) {
return "", nil
}
if err != nil {
return "", fmt.Errorf("read meta %s: %w", key, err)
}
return v, nil
}
// SetMeta writes (or overwrites) a single meta value.
func (s *Store) SetMeta(ctx context.Context, key, value string) error {
_, err := s.db.ExecContext(ctx,
`INSERT INTO meta (key, value) VALUES (?,?)
ON CONFLICT(key) DO UPDATE SET value = excluded.value`, key, value)
if err != nil {
return fmt.Errorf("write meta %s: %w", key, err)
}
return nil
}
// EmbedderUnknown is the stored id reported for a DB that already holds
// vectors but never recorded who wrote them.
const EmbedderUnknown = "unknown (written before this marker existed)"
// CheckEmbedder compares the embedder now configured against the one that
// wrote the stored vectors. Returns the stored id and whether it differs.
//
// Vectors from two different models live in different spaces, so cosine
// between them is noise rather than a low score — and both of our models are
// 384-dimensional, so nothing else catches it.
//
// Three cases, and the middle one is the one that actually matters:
//
// - marker present ⇒ compare the two ids.
// - marker absent but vectors already stored ⇒ this is a DB from before the
// marker, so we cannot know who wrote them. Report a mismatch. This is the
// real case on the deployed box: those vectors came from the old embedder,
// and claiming them for the current one would hide the exact problem the
// marker was added to catch.
// - marker absent and no vectors ⇒ fresh DB, claim it, nothing to fix.
//
// On a mismatch the fix is ReembedAll (backfill.go), run explicitly with
// `mavend -reembed`. Nothing is re-embedded here: that work is minutes of CPU
// on the laptop and must not stall a normal start.
func (s *Store) CheckEmbedder(ctx context.Context, currentID string) (stored string, mismatch bool, err error) {
stored, err = s.Meta(ctx, metaKeyEmbedderID)
if err != nil {
return "", false, err
}
if stored != "" {
return stored, stored != currentID, nil
}
n, err := s.countVectors(ctx)
if err != nil {
return "", false, err
}
if n > 0 {
return EmbedderUnknown, true, nil
}
return currentID, false, s.SetMeta(ctx, metaKeyEmbedderID, currentID)
}
// countVectors — how many stored rows carry an embedding. Used only to tell a
// fresh DB apart from one that predates the marker.
func (s *Store) countVectors(ctx context.Context) (int, error) {
var notes, vecs int
if err := s.db.QueryRowContext(ctx,
`SELECT count(*) FROM notes WHERE embedding IS NOT NULL`).Scan(&notes); err != nil {
return 0, fmt.Errorf("count note vectors: %w", err)
}
if err := s.db.QueryRowContext(ctx,
`SELECT count(*) FROM memory_vectors`).Scan(&vecs); err != nil {
return 0, fmt.Errorf("count memory vectors: %w", err)
}
return notes + vecs, nil
}
+100
View File
@@ -0,0 +1,100 @@
package store
import (
"context"
"testing"
"time"
)
// A fresh DB has no marker yet, so the current embedder is recorded and
// nothing is flagged.
func TestCheckEmbedderFreshDBRecords(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
stored, mismatch, err := s.CheckEmbedder(ctx, "multilingual-e5-small@384")
if err != nil {
t.Fatalf("CheckEmbedder: %v", err)
}
if mismatch {
t.Fatal("fresh DB reported a mismatch")
}
if stored != "multilingual-e5-small@384" {
t.Fatalf("stored = %q", stored)
}
got, err := s.Meta(ctx, metaKeyEmbedderID)
if err != nil {
t.Fatalf("Meta: %v", err)
}
if got != "multilingual-e5-small@384" {
t.Fatalf("marker not persisted, got %q", got)
}
}
// The deployed box: notes were written by the old embedder, before the marker
// existed. Claiming them for the current one would hide exactly the problem
// the marker is for, so an unmarked DB that already holds vectors is a
// mismatch.
func TestCheckEmbedderUnmarkedDBWithVectorsIsMismatch(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
if _, err := s.WriteNote(ctx, time.Now(), "молоко в холодильнике", []float32{0.1, 0.2}, "voice"); err != nil {
t.Fatalf("WriteNote: %v", err)
}
stored, mismatch, err := s.CheckEmbedder(ctx, "multilingual-e5-small@384")
if err != nil {
t.Fatalf("CheckEmbedder: %v", err)
}
if !mismatch {
t.Fatal("an unmarked DB with stored vectors should report a mismatch")
}
if stored != EmbedderUnknown {
t.Fatalf("stored = %q, want %q", stored, EmbedderUnknown)
}
// It must NOT claim the DB — that would silence the warning on restart.
got, err := s.Meta(ctx, metaKeyEmbedderID)
if err != nil {
t.Fatalf("Meta: %v", err)
}
if got != "" {
t.Fatalf("marker written despite unknown provenance: %q", got)
}
}
// Both models are 384-dim, so this is the only thing that catches the swap.
func TestCheckEmbedderDifferentModelMismatch(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
if err := s.SetMeta(ctx, metaKeyEmbedderID, "paraphrase-multilingual-MiniLM-L12-v2@384"); err != nil {
t.Fatalf("SetMeta: %v", err)
}
stored, mismatch, err := s.CheckEmbedder(ctx, "multilingual-e5-small@384")
if err != nil {
t.Fatalf("CheckEmbedder: %v", err)
}
if !mismatch {
t.Fatal("different embedder not detected")
}
if stored != "paraphrase-multilingual-MiniLM-L12-v2@384" {
t.Fatalf("stored = %q", stored)
}
}
// The same embedder must never raise a false alarm, including on re-check.
func TestCheckEmbedderSameModelNoAlarm(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
for i := 0; i < 2; i++ {
_, mismatch, err := s.CheckEmbedder(ctx, "multilingual-e5-small@384")
if err != nil {
t.Fatalf("CheckEmbedder: %v", err)
}
if mismatch {
t.Fatalf("false alarm on pass %d", i)
}
}
}
+5
View File
@@ -83,6 +83,11 @@ ALTER TABLE reminders ADD COLUMN next_fire_ts INTEGER;`, // #2
expires_ts INTEGER NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_dialogue_sessions_expires ON dialogue_sessions (expires_ts);`, // #10 — the follow-up session survives a restart (Vikunja #363); small, TTL-pruned table, not a history log
`CREATE TABLE IF NOT EXISTS meta (
key TEXT PRIMARY KEY,
value TEXT NOT NULL
);`, // #11 — small key/value table for facts about the DB itself; first key is embedder_id (Vikunja #378)
}
// migrate applies every migration with a number greater than the DB's current