Compare commits

...

29 Commits

Author SHA1 Message Date
claude c47881106e phraser: say "даже не знаю, что сказать" when there is nothing to say (V-397)
Review of #108: "поговорили." reads as a summary of a conversation that did
not happen. One exported constant now, so the Stub, the LLMPhraser fallback
and the daemon all say the same thing.

internal/voice/replier.go keeps its own copy — that is the separate replier
seam, not this one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 00:50:46 +04:00
claude 9a70f7378b phraser: move errEmptyResponse next to its only caller (V-397)
It sat in world.go, which is about the workstation model; it is a phrasing
error and belongs in llmphraser.go. Also trims the PhraseQuery doc.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 00:47:31 +04:00
claude b18f608594 mavend, eval: use the phrasing errors the phraser now returns (V-397)
Call sites take the fallback text and log the error instead of treating a
canned string as success. phraseSource drops the text entirely — its callers
hold the passage and read it back better than "вот что я нашла: <passage>".

The talk scorer's before-and-after model probe (the #395 workaround) goes;
the run now fails only when every case errored, which is the honest
"nothing was measured" condition. TalkFixture gets its own schema version so
the two fixtures can be versioned apart.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 00:41:16 +04:00
claude d1f8a734c5 phraser: report the failure next to the fallback (V-397)
PhraseChat and PhraseQuery returned canned text with a nil error, so a dead
or OOM-killed server was indistinguishable from bad phrasing — "не знаю." is
also a legitimate answer.

Both now return the fallback text AND the error. The daemon keeps using the
text, so the turn still survives; a measuring caller counts a real failure.
An empty response is its own error: the model is up and said nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 00:41:16 +04:00
kami 71041029e2 Merge pull request 'The reply path can't be tested — llmReplier is stuck in package main' (#107) from task/396-the-reply-path-can-t-be-tested-llmreplie into master
Reviewed-on: #107
2026-08-03 22:35:56 +02:00
claude 35018226ef eval: score the reply path, the fourth phrasing path (V-396)
Nine reply cases and a fourth column in the talk report. The reply path is a
separate object from the phraser in the daemon, so Pair joins a Talker and a
Confirmer for a run that covers everything Maven says.

Cases carry intent/key/value because the replier is phrased from the decision the
router resolved, not from the raw utterance. Three of them are baits the other
paths cannot produce: a masculine verb about himself that she must not copy onto
herself, a polite plural input that must still come back на ты, and an unresolved
note that invites a question a confirmation is not allowed to ask.

Not scored against a model here — this box has no llama-server, and the baseline
test is opt-in on MAVEN_LLM_URL.
2026-08-04 00:33:41 +04:00
claude 8833a9c76b mavend: keep only the stub floor in llmReplier (V-396)
The prompt, the call and the output parsing now live in internal/phraser. What is
left here is the one thing the daemon adds: a clarify, a model error and an
unusable generation all answer from voice.StubReplier, so a turn never breaks on
the model. The duplicated stripThink and parseResponseMood copies are gone;
capture.go uses phraser.StripThink.
2026-08-04 00:33:30 +04:00
claude 6c07409452 phraser: add Replier, the reply path lifted out of package main (V-396)
llmReplier lived in cmd/mavend, so the confirmation he hears after every fact,
note and reminder was the one phrasing path nothing could import or score.

Replier owns the prompt, the call and the parsing, and returns its errors instead
of hiding them — a dead model shows up as an error rather than as bad phrasing.
It has no stub fallback of its own; the daemon keeps that. StripThink is exported
for the daemon's own model callers.
2026-08-04 00:33:30 +04:00
kami 1c2541f7d6 Merge pull request 'llama-server holds 7.9GB RSS for a 1.1GB model, and its startup log goes nowhere' (#105) from task/496-recall-a-cross-language-question-loses-i into master
Reviewed-on: #105
2026-08-03 22:04:29 +02:00
claude 9e25f18a3e memory: record the recall topic veto's real price (V-496)
#496 asked to skip the veto when the question and the hit are in
different scripts, so an English question stops losing a Russian note.
Measured first: the fixture has no cross-language case, and en-hard-024
is an English question against an English note. Both proposed fixes are
no-ops.

What the veto actually does on the fixture, with the real embedder: it
costs en-hard-024 and buys ru-silent-029. Pass count is 22/32 either
way; false recall is 0/5 with it and 1/5 without. The two cases are one
lexical class, so no rule cheap enough for RecallAllowed separates them.

Accepts the loss and pins both sides in a test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 00:01:49 +04:00
kami 197897516e Merge pull request 'Task/495 bug x escapes the personal boundary and' (#104) from task/495-bug-x-escapes-the-personal-boundary-and into master
Reviewed-on: #104
2026-08-03 21:30:56 +02:00
kami 767748720a Merge pull request 'llama-server holds 7.9GB RSS for a 1.1GB model, and its startup log goes nowhere' (#103) from task/499-llama-server-holds-7-9gb-rss-for-a-1-1gb into task/495-bug-x-escapes-the-personal-boundary-and
Reviewed-on: #103
2026-08-03 21:30:37 +02:00
claude 58051b5af1 docs: record the #499 deploy (V-499) 2026-08-03 23:28:36 +04:00
claude f9b2391a8b phraser: cap llama-server's prompt cache at 512 MiB (V-499)
The forwarded log named the cause in one line: the prompt cache limit
defaults to 8192 MiB. llama-server saves the full KV state of every idle
slot it evicts, 112 kiB per token, so RSS climbed about 170MB per
distinct prompt until the deployed server held 7.9GB for a 1.1GB model.

Measured on homesrv today, uncapped versus `--cache-ram 512`: RSS
plateaus at 932MB from the fourth distinct prompt instead of climbing.
The task's leading guess was wrong. `-ngl 99` costs almost no RSS,
because RADV keeps device memory outside the process. Numbers and method
in docs/evals/2026-08-03-llama-prompt-cache.md.

`-c 4096` is untouched. The knob is `phraser.cache_ram_mib`, unset means
512, negative passes no flag for a llama-server too old to know it.

The deploy still runs the old image, so the box keeps its 8 GiB default
until mavend is rebuilt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 23:20:25 +04:00
claude f229795cea phraser: forward llama-server's output to mavend's log (V-499)
mavend scraped the child's stderr for the listen line and threw every
other line away, and never piped its stdout at all. Nothing about the
resident model's memory was diagnosable from a running box: no buffer
sizes, no KV-cache layout, no offload lines, no prompt-cache limit.

Both streams now share one pipe and every line lands in mavend's log
with a `llama:` prefix. The last 12 startup lines are also kept and go
into the error when the server dies before it listens, because bare
"EOF" never named which allocation it choked on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 23:19:03 +04:00
kami 6e5364a0ed Merge pull request 'Bug: "что я говорил про X" escapes the personal boundary and reaches web search' (#102) from task/495-bug-x-escapes-the-personal-boundary-and into master
Reviewed-on: #102
2026-08-03 20:59:34 +02:00
claude 86817d6d06 memory: score the personal boundary on seeds, not word lists (V-495)
"что я говорил про бэкапы?" is his data by definition, and nothing outside the
box has ever heard him say anything. The boundary matched possession words only,
so the question walked past it into SearXNG and came back answered out of a Habr
article about somebody else's backups.

A speech-verb marker class was written first and dropped. Russian gives every
verb a dozen surface forms and the "как я говорил, ..." preamble list has no end,
so each form the lexicon missed was one more question reaching the world, and a
missing verb looks exactly like no bug.

The boundary now embeds two frozen seed sets and scores the turn's own query
vector, already computed upstream, against both. Nearest side wins. The
possession markers stay as the offline floor for a handler with no embedder.

19/19 held-out utterances correct against multilingual-e5-small; see
docs/evals/2026-08-03-personal-boundary.md. The live probe on the deployed box is
not done.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:57:11 +04:00
kami 0fc2e3a18a Merge pull request 'Task/470 bug a question writes invented knowledge' (#101) from task/470-bug-a-question-writes-invented-knowledge into master
Reviewed-on: #101
2026-08-03 20:44:12 +02:00
kami 453919db20 Merge pull request 'Bug: the memory index stores the raw utterance as a fact's recall text, and nothing ever deletes a fact vector' (#100) from task/493-bug-the-memory-index-stores-the-raw-utte into task/470-bug-a-question-writes-invented-knowledge
Reviewed-on: #100
2026-08-03 20:41:09 +02:00
claude ad60e10e95 mavend: run the fact vector repair on start, and test what it does (V-493)
Automatic rather than a flag, unlike -reembed: only voice-tapped facts are in
this index, so it is tens of embeddings rather than thousands of notes. And
waiting for an operator to know the repair exists is the failure being fixed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:36:54 +04:00
claude 1528697287 store: repair fact vectors against the facts they name (V-493)
Every write-path fix leaves the rows already stored wrong, and a box in that
state looks fine: recall answers with the wrong text and nothing logs an error.
That is how the original poison survived four restarts.

RepairFactVectors resolves each fact vector against the fact it names,
re-embeds the ones whose text is stale, and deletes the voided, superseded and
orphaned ones. Marker-guarded and idempotent, so it runs once per box and a run
that dies partway is simply redone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:36:54 +04:00
claude dbdab2d570 store, mavend: a fact is indexed as the fact, not as the utterance (V-493)
queryMemory returns a fact's stored text verbatim, so the text the write path
indexed is what he hears. It was the utterance, which made recall of any
voice-tapped fact answer with the sentence he said: go_version = 1.20 was
indexed as "какая последняя версия языка Go?", and that question came back.

FactRecallText renders the fact instead, and the utterance stays in meta as
provenance. Correcting a value now drops the key's vectors the way voiding one
does, since the superseded value was still answering.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:36:34 +04:00
kami b9371dcac6 Merge pull request 'Bug: a question writes invented knowledge into memory as a self fact, and recall then serves it back for unrelated questions' (#99) from task/470-bug-a-question-writes-invented-knowledge into master
Reviewed-on: #99
2026-08-03 20:11:11 +02:00
claude 62c2e92ec0 mavend, recalleval: wire the topic veto into both recall sources (V-470)
queryMemory and queryNotes both gate on score alone, so both needed it. The
eval keeps its own copy of bestRecall — package main is not importable — and a
fixture that measures a weaker gate than the daemon runs flatters it, so the copy
moves in step and its test pins the new rule.

Measured on the held-out recall fixture with the real embedder: 17/32 cases pass
→ 22/32, false recall 1/5 → 0/5, answered after gate 18/27 → 17/27. The one true
recall lost is en-hard-024, an English question against a Russian note, where no
lexical test can help.
2026-08-03 13:51:04 +04:00
claude aec94eb2e8 memory: a world question must name what the memory mentions (V-470)
The score gate cannot separate the right note from an unrelated one: the
held-out fixture puts the right note at 0.791-0.890 and the must-be-silent cases
at 0.795-0.835, so a note about his slow network answered 'почему небо синее?'.

RecallAllowed adds a topic veto, and applies it only to a question that mentions
nothing of his. That restriction is the whole design: demanding a shared word of
every recall silenced four true recalls on the fixture to kill one false one,
because recall exists to find the note whose words he no longer remembers. A
question about his own life keeps the embedder as its only judge.
2026-08-03 13:50:54 +04:00
claude 4dfe106fe3 mavend: a question is never a fact about him (V-470)
IntentFact used to persist whatever the model invented for a question-shaped
utterance, at confidence 1.00, and index it for recall under the question's own
text. Two such rows then claimed seven unrelated world questions and silently
disabled world answering.

A question now goes down the query chain, which is what he asked for. The second
half is confidence: a value grounded in what he said stays 1.00, a value the model
supplied for words he never said drops to 0.60 and says so in the log. Same
reasoning as 'LLM output is not authorization' on the act path.
2026-08-03 13:40:33 +04:00
claude 2e0e2fd0bb router: a deterministic test for question-shaped text (V-470)
The predicate a fact write needs before it trusts a routing decision. Tokenized,
not substring: 'что' inside 'чтобы' is not a question. Capture verbs win over
every question signal, because 'запиши что я пил воду' contains an interrogative
and is still a capture.
2026-08-03 13:40:33 +04:00
claude f3fa6b353a store: voiding a fact drops its memory vectors (V-470)
Revert voided the fact row and left the vector, so recall kept serving the
voided fact's utterance and the documented repair reported success on a box that
stayed broken. There was no way to repair a poisoned box at all.

DeletePrefix covers every vector for the key, earlier rows included: their values
are superseded, and a superseded value has no business claiming a turn. It is
best-effort — the audit trail is already committed, and a fact that is voided but
still recallable beats a void that failed.
2026-08-03 13:40:13 +04:00
kami 6645f64c3e Merge pull request 'Name the gap: world questions through the workstation model, and the four remaining callers' (#98) from task/490-name-the-gap-world-questions-through-the into master
Reviewed-on: #98
2026-08-03 11:19:56 +02:00
42 changed files with 2316 additions and 276 deletions
+6 -1
View File
@@ -40,6 +40,7 @@ import (
"context"
"log"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
)
@@ -58,10 +59,14 @@ func (h *reactiveHandler) actionChat(ctx context.Context, dec router.Decision) s
// Conversational: build history from dialogue session (prior user turns)
// and let the LLM respond from general knowledge + context.
history := h.chatHistory()
// The phraser hands back its own fallback text alongside the error, so the
// turn survives a dead server and the failure still reaches the log.
reply, err := h.phraser.PhraseChat(ctx, dec.Utterance, history)
if err != nil {
log.Printf("voice: chat: %v", err)
return "поговорили."
}
if reply == "" {
return phraser.ChatFallback
}
return reply
}
+50 -14
View File
@@ -7,6 +7,7 @@ import (
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
)
// actionFact handles router.IntentFact: persist a tapped self-fact, index
@@ -15,14 +16,41 @@ func (h *reactiveHandler) actionFact(ctx context.Context, dec router.Decision) s
if !dec.Slots.HasKey {
return "не разобрала, что записать — попробуй иначе."
}
// A question is never a fact about him (#470). "какая последняя версия
// языка Go?" used to land here, and the value stored was whatever the
// model invented for it, at confidence 1.00, indexed for recall under the
// question's own text. Two such rows then claimed seven unrelated world
// questions through recall and silently disabled world answering.
//
// The routing error itself is not fixed here — the answer is to answer.
// Sending the turn down the query chain is what he asked for anyway, and
// it costs a mis-routed capture nothing: an explicit "запиши ..." is not
// question-shaped, so it never takes this branch.
if router.IsQuestionShaped(dec.Utterance) {
log.Printf("voice: fact write refused, utterance is a question: %q (key %q) — answering as a query",
dec.Utterance, dec.Slots.Key)
q := dec
q.Intent = router.IntentQuery
// The key the model extracted is its guess at what to store, not a
// fact he has. Left in place, queryFactByKey would read it back and
// claim the turn before any real source ran.
q.Slots.Key, q.Slots.HasKey = "", false
q.Slots.Value = ""
return h.actionQuery(ctx, q)
}
now := h.now()
req := ipc.WriteFactReq{
Ts: now,
Kind: "self",
Key: dec.Slots.Key,
Value: dec.Slots.Value,
Source: "tap:voice",
Confidence: 1.0,
Ts: now,
Kind: "self",
Key: dec.Slots.Key,
Value: dec.Slots.Value,
Source: "tap:voice",
// Not 1.00 unconditionally any more (#470). A value he said is
// evidence; a value the model supplied for words he never said is a
// guess, and writing a guess at full confidence is the same mistake
// the act path already refuses under "LLM output is not
// authorization".
Confidence: factConfidence(dec.Utterance, dec.Slots.Value),
// Subject: the key doubles as the entity-resolution candidate —
// a voice-tapped fact's key is usually the thing/person it's
// about ("espresso_machine", "kate"), so queueing it for Nexus
@@ -36,17 +64,25 @@ func (h *reactiveHandler) actionFact(ctx context.Context, dec router.Decision) s
log.Printf("voice: write fact: %v", err)
return "не получилось сохранить факт."
}
// Index the fact utterance in long-term memory (best-effort, must not
// fail the fact write). Facts aren't in the notes table, so this is the
// only recall path for them — "когда я пил воду?" reads back from here.
// Index the fact in long-term memory (best-effort, must not fail the fact
// write). Facts aren't in the notes table, so this is the only recall path
// for them — "когда я пил воду?" reads back from here.
//
// The indexed text is the fact, not the utterance (#493). queryMemory
// returns a fact's stored text verbatim, so what goes in here is what he
// hears; storing the utterance meant recall answered with his own sentence
// rather than the value. The utterance stays alongside as provenance —
// readable on /trace, never the answer and never embedded.
if h.memStore != nil {
if vec, err := router.EmbedPassage(ctx, h.embedder, dec.Utterance); err != nil {
text := store.FactRecallText(dec.Slots.Key, dec.Slots.Value)
if vec, err := router.EmbedPassage(ctx, h.embedder, text); err != nil {
log.Printf("voice: embed fact for memory: %v", err)
} else if err := h.memStore.Insert(ctx, "fact:"+dec.Slots.Key+":"+strconv.FormatInt(now.Unix(), 10), vec, map[string]string{
"source": "voice",
"type": "fact",
"text": dec.Utterance,
"ts": strconv.FormatInt(now.Unix(), 10),
"source": "voice",
"type": "fact",
"text": text,
"utterance": dec.Utterance,
"ts": strconv.FormatInt(now.Unix(), 10),
}); err != nil {
log.Printf("voice: memory insert fact: %v", err)
}
+42 -3
View File
@@ -434,10 +434,24 @@ func (h *reactiveHandler) queryMemory(ctx context.Context, t *queryTurn) (string
return "", false
}
text := hit.Meta["text"]
// The score cleared the gate and the topic still has to match (#470). A
// note about his slow network scored high enough to answer "почему небо
// синее?", because the right-note and must-be-silent score ranges overlap
// and no threshold sits between them.
if !memory.RecallAllowed(t.dec.Utterance, text) {
log.Printf("voice: recall %q rejected for %q: a world question and no shared topic word", text, t.dec.Utterance)
return "", false
}
// A note is phrased in Maven's voice; a fact is read back as it was
// stored.
if hit.Meta["type"] == "note" {
if reply, perr := h.phraser.PhraseQuery(ctx, t.dec.Utterance, []string{text}); perr == nil && reply != "" {
reply, perr := h.phraser.PhraseQuery(ctx, t.dec.Utterance, []string{text})
switch {
case perr != nil:
// Reading the note back verbatim beats the phraser's own fallback,
// which only wraps the same text in "вот что я нашла:".
log.Printf("voice: recall phrase: %v", perr)
case reply != "":
return reply, true
}
}
@@ -469,6 +483,12 @@ func (h *reactiveHandler) queryNotes(ctx context.Context, t *queryTurn) (string,
if !memory.ConfidentScores(noteScores, h.queryMinScore, h.queryMinMargin) {
return "", false
}
// Same topic veto as queryMemory above: the best note must be about what
// he asked, not merely the nearest vector in the index.
if !memory.RecallAllowed(t.dec.Utterance, notes[0].Text) {
log.Printf("voice: note %q rejected for %q: a world question and no shared topic word", notes[0].Text, t.dec.Utterance)
return "", false
}
texts := make([]string, len(notes))
for i, n := range notes {
texts[i] = n.Text
@@ -701,7 +721,7 @@ func (h *reactiveHandler) queryKiwix(ctx context.Context, t *queryTurn) (string,
// not be sent to an upstream engine at all. The guard closes both holes with
// the same test.
func (h *reactiveHandler) queryPersonal(ctx context.Context, t *queryTurn) (string, bool) {
if !isPersonalQuery(t.dec.Utterance) {
if !h.isPersonalTurn(ctx, t) {
return "", false
}
log.Printf("voice: %q is about him and his own data did not answer it; not asking the world", t.dec.Utterance)
@@ -726,7 +746,9 @@ var personalMarkers = []*regexp.Regexp{
regexp.MustCompile(`(?i)\bdid\s+i\b`),
}
// isPersonalQuery reports whether the utterance asks about something of his.
// isPersonalQuery — the offline floor under the boundary. Possession only, and
// deliberately still narrow: it answers when there is no embedder to ask, and a
// broad guess made blind is worse than a narrow one.
func isPersonalQuery(utterance string) bool {
if utterance == "" {
return false
@@ -739,6 +761,23 @@ func isPersonalQuery(utterance string) bool {
return false
}
// isPersonalTurn — the boundary test. The seeds decide when the embedder is
// there, which is every deployed box; the possession markers are the floor
// underneath, for a handler with no embedder or a turn whose vector never got
// computed. Same shape as the cascade: the better test leads, the offline one
// always answers.
func (h *reactiveHandler) isPersonalTurn(ctx context.Context, t *queryTurn) bool {
h.boundary.load(ctx, h.embedder)
if personal, world, ok := h.boundary.score(t.vec); ok {
if personal > world {
log.Printf("voice: %q scores personal %.4f vs world %.4f", t.dec.Utterance, personal, world)
return true
}
return false
}
return isPersonalQuery(t.dec.Utterance)
}
// queryGeneral — general knowledge, the last source before giving up. It always
// claims: either a model answers, or Maven names the gap, or she says she does
// not know.
@@ -19,6 +19,10 @@ func TestIsPersonalQuery(t *testing.T) {
"when is my meeting",
"do i have anything today",
"did i take my vitamins",
// Speech, but only the forms possession already covers ("did i").
// The verb forms the floor cannot see are the seeds' job, scored in
// TestONNXPersonalBoundary.
"what did i say about backups",
} {
if !isPersonalQuery(s) {
t.Errorf("isPersonalQuery(%q) = false, want true", s)
@@ -33,6 +37,10 @@ func TestIsPersonalQuery(t *testing.T) {
"почему небо синее",
"столица франции",
"how do i boil an egg",
// The floor is possession-only by design: a speech verb it cannot see
// passes here and is caught by the seeds instead.
"что я говорил про бэкапы?",
"как я говорил, почему небо синее",
"",
} {
if isPersonalQuery(s) {
+1 -1
View File
@@ -100,7 +100,7 @@ func (l llmCompleter) Complete(ctx context.Context, system, user string) (string
// grammar, or a llama-server too old to honour one, gets the plain text it used
// to get rather than an empty meeting summary.
func unwrapSummary(raw string) string {
s := stripThink(strings.TrimSpace(raw))
s := phraser.StripThink(strings.TrimSpace(raw))
start := strings.Index(s, "{")
end := strings.LastIndex(s, "}")
if start < 0 || end <= start {
+82
View File
@@ -0,0 +1,82 @@
package main
import (
"log"
"strings"
"unicode"
)
// ungroundedConfidence — what a self fact is worth when its value appears
// nowhere in what he said. Below `query_min_score` is not the point (recall
// gates on vector distance, not on this number); the point is that
// `/history` and every future reader can tell a value he said from a value
// the model supplied.
const ungroundedConfidence = 0.6
// factConfidence scores a self fact by whether its value is grounded in the
// utterance it came from. Grounded stays 1.00, which is what a tapped fact
// has always been worth. Ungrounded drops, and says so in the log.
//
// An empty value is grounded by definition: the key alone carries the fact
// ("поужинал"), and there is nothing for the model to have invented.
func factConfidence(utterance, value string) float64 {
if strings.TrimSpace(value) == "" {
return 1.0
}
if valueGrounded(utterance, value) {
return 1.0
}
log.Printf("voice: fact value %q is not in %q — writing at confidence %.2f",
value, utterance, ungroundedConfidence)
return ungroundedConfidence
}
// valueGrounded reports whether every word of value traces back to a word he
// actually said. The comparison is on a 4-rune prefix, so the model's
// normalization survives ("пил воду" → "вода") while an invented value
// ("1.20" for a question about Go) does not.
func valueGrounded(utterance, value string) bool {
said := factTokens(utterance)
words := factTokens(value)
if len(words) == 0 {
return true
}
for _, w := range words {
if !anyTokenMatches(said, w) {
return false
}
}
return true
}
func anyTokenMatches(said []string, w string) bool {
for _, s := range said {
if s == w || sameStem(s, w) {
return true
}
}
return false
}
// sameStem is inflection tolerance and nothing more: it compares all but the
// last rune of the shorter word, and never fewer than three. Russian marks
// case on the ending, so "пил воду" and the stored "вода" are the same word he
// said, while "1.20" and "версия" are not. A word of three runes or fewer must
// match outright, where a shorter prefix would match half the language.
func sameStem(a, b string) bool {
ar, br := []rune(a), []rune(b)
shorter := min(len(ar), len(br))
n := shorter - 1
if n < 3 || len(ar) < n || len(br) < n {
return false
}
return string(ar[:n]) == string(br[:n])
}
// factTokens lowercases and splits on everything that is not a letter or a
// digit, the same shape planTokens uses in the router.
func factTokens(s string) []string {
return strings.FieldsFunc(strings.ToLower(s), func(r rune) bool {
return !unicode.IsLetter(r) && !unicode.IsDigit(r)
})
}
+125
View File
@@ -0,0 +1,125 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/tool"
"github.com/kami/maven/internal/voice"
)
func newFactGateHandler(t *testing.T, now time.Time) (*reactiveHandler, ipc.CoreAPI) {
t.Helper()
st := newTestStore(t)
api := ipc.NewStoreAPI(st)
emb := router.NewHashEmbedder(1024)
h := &reactiveHandler{
api: api,
embedder: emb,
router: buildRouter(emb, tool.NewMatcher(api), 0.55, nil),
replier: voice.NewStubReplier(),
now: func() time.Time { return now },
memStore: memory.NewInMemoryStore(),
dataStore: st,
}
return h, api
}
// The write half of #470: a question routed to IntentFact must not become a
// fact about him, and must not leave a vector behind for recall to serve.
func TestActionFact_QuestionIsNotWritten(t *testing.T) {
ctx := context.Background()
h, api := newFactGateHandler(t, time.Now())
reply := h.actionFact(ctx, router.Decision{
Intent: router.IntentFact,
Utterance: "какая последняя версия языка Go?",
Slots: router.Slots{Key: "go_version", HasKey: true, Value: `"1.20"`},
})
if _, err := api.LatestFact(ctx, "go_version"); err == nil {
t.Fatal("a question was stored as a fact about him")
}
hits, err := h.memStore.Search(ctx, mustEmbedPassage(t, h, "какая последняя версия языка Go?"), 3)
if err != nil {
t.Fatalf("memory search: %v", err)
}
if len(hits) != 0 {
t.Fatalf("the question was indexed for recall: %+v", hits)
}
// It went down the query chain instead. Nothing is configured to answer a
// world question in this harness, so "не знаю." is the honest outcome —
// what matters is that the turn was answered, not stored.
if reply == "" {
t.Fatal("the turn was neither stored nor answered")
}
}
// The capture that must survive the gate: an explicit instruction to record,
// even though it contains an interrogative.
func TestActionFact_ExplicitCaptureStillWrites(t *testing.T) {
ctx := context.Background()
h, api := newFactGateHandler(t, time.Now())
h.actionFact(ctx, router.Decision{
Intent: router.IntentFact,
Utterance: "запиши что я пил воду",
Slots: router.Slots{Key: "water", HasKey: true, Value: `"вода"`},
})
f, err := api.LatestFact(ctx, "water")
if err != nil {
t.Fatalf("an explicit capture was refused: %v", err)
}
if f.Confidence != 1.0 {
t.Errorf("confidence = %v, want 1.0 for a value he said", f.Confidence)
}
// #493: what recall reads back is the fact, not the sentence he said.
// queryMemory returns a fact's text verbatim, so the utterance sitting here
// meant "запиши что я пил воду" was the answer to "когда я пил воду?".
hits, err := h.memStore.Search(ctx, mustEmbedPassage(t, h, "вода"), 3)
if err != nil {
t.Fatalf("memory search: %v", err)
}
if len(hits) != 1 {
t.Fatalf("the fact was not indexed once: %+v", hits)
}
if got := hits[0].Meta["text"]; got != "water — вода" {
t.Errorf("indexed text = %q, want the fact", got)
}
if got := hits[0].Meta["utterance"]; got != "запиши что я пил воду" {
t.Errorf("utterance provenance = %q, want it kept alongside", got)
}
}
func TestFactConfidence(t *testing.T) {
cases := []struct {
utterance, value string
want float64
}{
{"запиши что я пил воду", `"вода"`, 1.0},
{"я выпил кофе", `"кофе"`, 1.0},
{"поужинал", "", 1.0},
{"отметь что я полил кактус", `"полил кактус"`, 1.0},
{"какая последняя версия языка Go", `"1.20"`, ungroundedConfidence},
{"кто премьер Японии", `"Тонио Озаки"`, ungroundedConfidence},
}
for _, c := range cases {
if got := factConfidence(c.utterance, c.value); got != c.want {
t.Errorf("factConfidence(%q, %q) = %v, want %v", c.utterance, c.value, got, c.want)
}
}
}
func mustEmbedPassage(t *testing.T, h *reactiveHandler, text string) []float32 {
t.Helper()
vec, err := router.EmbedQuery(context.Background(), h.embedder, text)
if err != nil {
t.Fatalf("embed %q: %v", text, err)
}
return vec
}
+18
View File
@@ -251,6 +251,7 @@ func run(args []string) error {
Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
CacheRAMMiB: cacheRAMMiB(cfg.Phraser.CacheRAMMiB),
Timeout: time.Duration(cfg.Phraser.Timeout),
LLMNudges: cfg.Phraser.LLMNudges,
ContextBlock: contextBlockFn(cfg, time.Now),
@@ -525,6 +526,7 @@ func run(args []string) error {
Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
CacheRAMMiB: cacheRAMMiB(cfg.Phraser.CacheRAMMiB),
Timeout: time.Duration(cfg.Phraser.Timeout),
LLMNudges: cfg.Phraser.LLMNudges,
ContextBlock: contextBlockFn(cfg, time.Now),
@@ -788,6 +790,22 @@ func personaFacts(cfg *config.Config) persona.Facts {
return f
}
// cacheRAMMiB resolves phraser.cache_ram_mib into the phraser's field. Unset
// means 512 MiB and not "whatever the server does", because the server's own
// default is 8 GiB of prompt cache and that is what put 7.9 GB of RSS and half
// a gigabyte of swap on homesrv for a 1.1 GB model. A negative value is the
// deliberate opt-out: no flag is passed, the server's default applies, and the
// operator owns the consequence.
func cacheRAMMiB(configured int) int {
if configured == 0 {
return 512
}
if configured < 0 {
return 0
}
return configured
}
// contextBlockFn returns the per-turn renderer of the shared context block.
// Per turn, not once at startup, because the block states the current time.
func contextBlockFn(cfg *config.Config, now func() time.Time) func() string {
+154
View File
@@ -0,0 +1,154 @@
package main
import (
"context"
"log"
"math"
"sync"
"github.com/kami/maven/internal/router"
)
// The personal boundary decides one thing: is this question about him. It used
// to decide it by matching possession words, and that was the whole defect
// behind Vikunja #495. "что я говорил про бэкапы?" is his data by definition —
// nothing outside the box has ever heard him say anything — and it carried no
// possession word, so it walked past the boundary into SearXNG and came back
// answered out of a Habr article about somebody else's backups.
//
// The first fix was one more marker class, `я говорил|сказал|писал|…`, plus a
// carve-out so "как я говорил, почему небо синее" stayed a world question. Both
// halves are a lexicon, and a lexicon is the wrong instrument here: Russian
// gives every verb a dozen surface forms, the preamble list has no end, and
// every utterance the list misses is one that reaches the world. It also drifts
// silently — a missing verb looks exactly like no bug.
//
// So the boundary asks the embedder instead. Two frozen seed sets — questions
// about him, questions about the world — are embedded once, and the turn's own
// query vector, already computed by queryEmbed upstream, is scored against
// both. Nearest side wins. Word order, verb form and unseen phrasing stop
// mattering, which is exactly what a lexicon could not do.
//
// Measured 03-08-2026 against multilingual-e5-small on 19 held-out utterances,
// none of them a seed: 19 right (TestONNXPersonalBoundary). A 20th, "as i said,
// what is the population of india", missed by +0.008 during the first pass and
// is a world seed now, which is why it is not in the held-out set. True
// positives clear the world side by +0.014 to +0.089 and the nearest true
// negative sits at -0.005, so the gate is the sign of the difference and
// nothing tighter: the margins are too thin to justify a threshold, and the
// asymmetry favours claiming anyway. A false claim costs one honest "не знаю";
// a false pass sends his life to an upstream engine.
//
// The embedder is the one model CLAUDE.md pins to homesrv permanently, and it
// is what makes this affordable: no llama-server call, no network, one cosine
// per seed against a vector the turn already has.
// personalSeeds — questions about him. Frozen: they are scoring data, so
// editing one moves the boundary and must be re-measured, not eyeballed. Cover
// both classes the boundary owns, possession and first-person speech, in both
// languages.
var personalSeeds = []string{
"что я говорил про это",
"я тебе рассказывал об этом?",
"что я записал про врача",
"я упоминал эту тему?",
"что у меня сегодня",
"когда моя встреча",
"what did i say about this",
"did i mention this to you",
}
// worldSeeds — questions the world can answer, including the two shapes that
// look personal and are not: a first-person preamble on a world question ("как
// я говорил, ..."), and first person without possession ("что я могу
// посмотреть вечером"). Refusing those is the opposite mistake and the older
// comment on personalMarkers already named it.
var worldSeeds = []string{
"почему небо синее",
"какая столица франции",
"как сварить борщ",
"кто написал эту книгу",
"what is the capital of france",
"how do i boil an egg",
"как я говорил, почему небо синее",
"as i said, why is the sky blue",
"as i said, what is the population of india",
"что я могу посмотреть вечером",
"что мне почитать про историю",
"что я должен знать про питон",
"what can i watch tonight",
}
// personalBoundary holds the embedded seeds. Zero value is usable and means
// "not loaded yet"; a handler built without an embedder never loads and the
// boundary falls back to personalMarkers.
type personalBoundary struct {
once sync.Once
personal [][]float32
world [][]float32
loaded bool
}
// load embeds both seed sets, once per process. Seeds are embedded on the QUERY
// side, like the utterance they are compared with — a question against a
// question. Mixing sides would measure the e5 prefix, not the meaning.
func (b *personalBoundary) load(ctx context.Context, emb router.Embedder) {
b.once.Do(func() {
if emb == nil {
return
}
embedAll := func(ss []string) [][]float32 {
out := make([][]float32, 0, len(ss))
for _, s := range ss {
v, err := router.EmbedQuery(ctx, emb, s)
if err != nil {
log.Printf("voice: personal boundary seeds unavailable (%v); falling back to possession markers", err)
return nil
}
out = append(out, v)
}
return out
}
p, w := embedAll(personalSeeds), embedAll(worldSeeds)
if p == nil || w == nil {
return
}
b.personal, b.world, b.loaded = p, w, true
})
}
// score returns the best similarity to each side. ok is false when the seeds
// are not loaded, which is the caller's signal to use the markers instead.
func (b *personalBoundary) score(vec []float32) (personal, world float64, ok bool) {
if !b.loaded || len(vec) == 0 {
return 0, 0, false
}
best := func(seeds [][]float32) float64 {
m := -1.0
for _, s := range seeds {
if c := cosine(vec, s); c > m {
m = c
}
}
return m
}
return best(b.personal), best(b.world), true
}
// cosine — same math as internal/router and internal/memory, small enough that
// importing one of them for it would be the larger coupling.
func cosine(a, b []float32) float64 {
if len(a) != len(b) {
return 0
}
var dot, na, nb float64
for i := range a {
dot += float64(a[i]) * float64(b[i])
na += float64(a[i]) * float64(a[i])
nb += float64(b[i]) * float64(b[i])
}
if na == 0 || nb == 0 {
return 0
}
return dot / (math.Sqrt(na) * math.Sqrt(nb))
}
+94
View File
@@ -0,0 +1,94 @@
package main
import (
"context"
"os"
"path/filepath"
"testing"
"github.com/kami/maven/internal/router"
)
// A handler with no embedder never loads the seeds, so the boundary falls back
// to the possession markers. That is the offline floor and it must keep working
// — an embedder that fails to load must not open the boundary.
func TestBoundaryFallsBackToMarkersWithNoEmbedder(t *testing.T) {
h := personalHandler()
if !h.isPersonalTurn(context.Background(), &queryTurn{
dec: router.Decision{Utterance: "во сколько у меня встреча"},
}) {
t.Error("no embedder: a possession question must still be personal")
}
if h.isPersonalTurn(context.Background(), &queryTurn{
dec: router.Decision{Utterance: "почему небо синее"},
}) {
t.Error("no embedder: a world question must still pass")
}
}
// TestONNXPersonalBoundary — the number that matters, scored against the
// embedder homesrv actually runs. Opt-in via MAVEN_ONNX_LIB, exactly like
// TestONNXRecall in internal/memory/recalleval.
//
// Every case here is held out: none of these strings is a seed. The #495
// regression is the first row — "что я говорил про бэкапы?" reached SearXNG and
// was answered from a Habr article, and no possession word appears in it.
func TestONNXPersonalBoundary(t *testing.T) {
lib := os.Getenv("MAVEN_ONNX_LIB")
if lib == "" {
t.Skip("MAVEN_ONNX_LIB unset — see AGENTS.md § Embedder model for intent routing")
}
dir := filepath.Join("../..", "models/embedder/multilingual-e5-small")
emb, err := router.NewONNXEmbedder(filepath.Join(dir, "model_quantized.onnx"), filepath.Join(dir, "tokenizer.json"), lib)
if err != nil {
t.Skipf("onnx embedder unavailable: %v", err)
}
defer emb.Close()
cases := []struct {
utterance string
personal bool
}{
{"что я говорил про бэкапы?", true},
{"что я сказал вчера про отпуск", true},
{"я писал что-нибудь про сервер", true},
{"я упоминал про конференцию?", true},
{"что я отмечал по поводу переезда", true},
{"я рассказывал тебе про новую работу?", true},
{"во сколько у меня встреча", true},
{"когда мой следующий отпуск", true},
{"what did i say about backups", true},
{"did i tell you about the doctor", true},
{"как я говорил, почему небо синее", false},
{"как уже я говорил, какая столица франции", false},
{"почему трава зелёная", false},
{"столица франции", false},
{"как мне сварить борщ", false},
{"что мне посмотреть вечером", false},
{"я хочу узнать про рим", false},
{"кто такой гагарин", false},
{"how do i boil an egg", false},
}
h := &reactiveHandler{embedder: emb}
ctx := context.Background()
wrong := 0
for _, c := range cases {
vec, err := router.EmbedQuery(ctx, emb, c.utterance)
if err != nil {
t.Fatalf("embed %q: %v", c.utterance, err)
}
turn := &queryTurn{dec: router.Decision{Utterance: c.utterance}, vec: vec}
got := h.isPersonalTurn(ctx, turn)
p, w, ok := h.boundary.score(vec)
if !ok {
t.Fatal("seeds did not load with a working embedder")
}
if got != c.personal {
wrong++
t.Errorf("%q: personal=%v want %v (personal %.4f world %.4f)", c.utterance, got, c.personal, p, w)
}
t.Logf("personal=%-5v personal %.4f world %.4f delta %+.4f %s", got, p, w, p-w, c.utterance)
}
t.Logf("personal boundary: %d/%d held-out utterances correct", len(cases)-wrong, len(cases))
}
+11 -100
View File
@@ -2,122 +2,33 @@ package main
import (
"context"
"encoding/json"
"strings"
"time"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice"
)
// completer is the LLM seam for the replier (subset of router.Completer).
// *llm.Client satisfies it.
type completer interface {
Complete(ctx context.Context, r llm.Req) (string, error)
}
// llmReplier phrases reactive confirmations with the resident model
// (Qwen3-1.7B). Stub is the
// floor on any error (offline-safe). Maven speaks as "she", feminine RU.
// llmReplier is the daemon-side wiring around phraser.Replier: it owns the
// deterministic floor, and nothing else. The phrasing itself, the prompt and the
// output parsing live in internal/phraser so the eval can score them (#396).
type llmReplier struct {
c completer
p *phraser.Replier
stub *voice.StubReplier
// block renders the shared context block per turn (who he is, the time).
// nil ⇒ the prompt stands alone.
block func() string
}
func newLLMReplier(c completer, block func() string) *llmReplier {
return &llmReplier{c: c, stub: voice.NewStubReplier(), block: block}
func newLLMReplier(c phraser.Completer, block func() string) *llmReplier {
return &llmReplier{p: phraser.NewReplier(c, block), stub: voice.NewStubReplier()}
}
const replySystem = `Ты — Maven, домашняя ассистентка (о себе — в женском роде). Владелец — мужчина, говоришь с ним на "ты", в единственном числе; никогда не "вы"/"ваш" и не "он"/"его". Подтверди действие РОВНО ОДНИМ коротким предложением (≤120 символов), по-русски, спокойно и без официальных формулировок. Не задавай вопросов, не повторяй слова, не добавляй ничего после точки. Отвечай ТОЛЬКО одним объектом JSON с полями "response" (текст) и "mood" (ровно одно из: neutral, happy, thinking, tired, confused).
Пример: {"response": "Записала, что ты выпил стакан воды.", "mood": "neutral"}
Никогда не пиши "..." в поле response.`
// Reply never fails: a clarify, a model error and an unusable generation all
// answer from the stub, which is what keeps a turn from breaking on the model.
func (r *llmReplier) Reply(d router.Decision) string {
if d.Clarify {
return r.stub.Reply(d)
}
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
out, err := r.c.Complete(ctx, llm.Req{
System: persona.Prepend(r.block, replySystem),
User: replyContext(d),
Grammar: phraser.ResponseGrammar,
MaxTokens: 512,
RepeatPenalty: 1.3,
})
if err != nil {
out, err := r.p.PhraseReply(context.Background(), d)
if err != nil || out == "" {
return r.stub.Reply(d)
}
out = stripThink(out)
if response, _ := parseResponseMood(out); response != "" {
return response
}
// fallback: try plain-text parsing
if out = firstSentence(out); out != "" {
return out
}
return r.stub.Reply(d)
}
// firstSentence trims the model's output to a single clean confirmation: first
// line, first sentence, whitespace-normalized — the last-line defense against a
// small model that rambles past the first period despite the prompt + stop.
// stripThink removes the <think> block that Thinking-variant models emit.
func stripThink(s string) string {
if i := strings.LastIndex(s, "</think>"); i >= 0 {
s = strings.TrimSpace(s[i+8:])
}
return s
}
func firstSentence(s string) string {
s = strings.TrimSpace(s)
if i := strings.IndexByte(s, '\n'); i >= 0 {
s = s[:i]
}
// keep up to and including the first sentence-ending punctuation.
if i := strings.IndexAny(s, ".!?"); i >= 0 {
s = s[:i+1]
}
return strings.TrimSpace(s)
}
// parseResponseMood extracts {"response","mood"} from LLM output, tolerant
// of thinking tokens and extra text before/after the JSON block.
func parseResponseMood(raw string) (response, mood string) {
cleaned := strings.TrimSpace(raw)
start := strings.Index(cleaned, "{")
end := strings.LastIndex(cleaned, "}")
if start < 0 || end < 0 || end <= start {
return "", ""
}
var parsed struct {
Response string `json:"response"`
Mood string `json:"mood"`
}
if err := json.Unmarshal([]byte(cleaned[start:end+1]), &parsed); err != nil {
return "", ""
}
return parsed.Response, parsed.Mood
}
// replyContext renders the decision into a compact RU description for the model.
func replyContext(d router.Decision) string {
switch d.Intent {
case router.IntentFact:
return "записала факт: " + d.Slots.Key + " " + d.Slots.Value
case router.IntentNote:
return "сохранила заметку: " + d.Slots.Text
case router.IntentReminder:
return "поставила напоминание: " + d.Slots.Text
default:
return string(d.Intent) + ": " + d.Slots.Text
}
return out
}
+20 -50
View File
@@ -5,28 +5,22 @@ import (
"testing"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice"
)
type mockCompleter struct {
// The phrasing itself is tested in internal/phraser. What is left here is the
// only thing the daemon adds: the stub floor, on the three ways a reply can
// fail to arrive.
type stubCompleter struct {
out string
err error
}
func (m mockCompleter) Complete(_ context.Context, _ llm.Req) (string, error) { return m.out, m.err }
func (s stubCompleter) Complete(_ context.Context, _ llm.Req) (string, error) { return s.out, s.err }
func TestLLMReplierReturnsLLMReply(t *testing.T) {
r := newLLMReplier(mockCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`}, nil)
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
if got != "записала, кофе закончился" {
t.Errorf("got %q, want %q", got, "записала, кофе закончился")
}
}
func TestLLMReplierFallsBackToPlainText(t *testing.T) {
r := newLLMReplier(mockCompleter{out: "записала, кофе закончился"}, nil)
func TestLLMReplierPassesTheModelReplyThrough(t *testing.T) {
r := newLLMReplier(stubCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`}, nil)
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
if got != "записала, кофе закончился" {
t.Errorf("got %q, want %q", got, "записала, кофе закончился")
@@ -34,54 +28,30 @@ func TestLLMReplierFallsBackToPlainText(t *testing.T) {
}
func TestLLMReplierFallsBackToStubOnError(t *testing.T) {
r := newLLMReplier(mockCompleter{err: errTestLLMDown}, nil)
noteDec := router.Decision{Intent: router.IntentNote}
got := r.Reply(noteDec)
want := voice.NewStubReplier().Reply(noteDec)
if got != want {
t.Errorf("on llm error: got %q, want stub %q", got, want)
}
r := newLLMReplier(stubCompleter{err: errReplierTest}, nil)
assertStub(t, r, router.Decision{Intent: router.IntentNote}, "llm error")
}
func TestLLMReplierFallsBackToStubOnEmpty(t *testing.T) {
r := newLLMReplier(mockCompleter{out: ""}, nil)
noteDec := router.Decision{Intent: router.IntentNote}
got := r.Reply(noteDec)
want := voice.NewStubReplier().Reply(noteDec)
if got != want {
t.Errorf("on empty llm: got %q, want stub %q", got, want)
}
r := newLLMReplier(stubCompleter{out: ""}, nil)
assertStub(t, r, router.Decision{Intent: router.IntentNote}, "empty llm")
}
func TestLLMReplierClarifyUsesStub(t *testing.T) {
r := newLLMReplier(mockCompleter{out: "я всё поняла"}, nil)
clarifyDec := router.Decision{Clarify: true}
got := r.Reply(clarifyDec)
want := voice.NewStubReplier().Reply(clarifyDec)
r := newLLMReplier(stubCompleter{out: "я всё поняла"}, nil)
assertStub(t, r, router.Decision{Clarify: true}, "clarify")
}
func assertStub(t *testing.T, r *llmReplier, d router.Decision, what string) {
t.Helper()
got, want := r.Reply(d), voice.NewStubReplier().Reply(d)
if got != want {
t.Errorf("on clarify: got %q, want stub %q", got, want)
t.Errorf("on %s: got %q, want stub %q", what, got, want)
}
}
var errTestLLMDown = errTest("llm down")
var errReplierTest = errTest("llm down")
type errTest string
func (e errTest) Error() string { return string(e) }
// grammarRecorder captures the request so the grammar can be asserted on.
type grammarRecorder struct{ req llm.Req }
func (g *grammarRecorder) Complete(_ context.Context, r llm.Req) (string, error) {
g.req = r
return `{"response":"записала","mood":"neutral"}`, nil
}
func TestLLMReplierCarriesTheResponseGrammar(t *testing.T) {
rec := &grammarRecorder{}
r := newLLMReplier(rec, nil)
r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
if rec.req.Grammar != phraser.ResponseGrammar {
t.Errorf("grammar = %q, want phraser.ResponseGrammar", rec.req.Grammar)
}
}
+4
View File
@@ -76,6 +76,10 @@ type reactiveHandler struct {
tts tts.Synthesizer
router *router.Router
embedder router.Embedder // reused for note write/query (same model as the classifier)
// boundary — the embedded seed sets behind the personal boundary
// (personalboundary.go). Zero value is usable and loads on first query;
// with no embedder it never loads and the boundary uses personalMarkers.
boundary personalBoundary
// api — the CoreAPI the handler reads and writes through. Wired with the
// bare store adapter and UPGRADED by main once the daemonAPI exists; see
// upgradeAPI.
+31
View File
@@ -146,6 +146,7 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
emb = router.NewHashEmbedder(1024)
}
w.embedder = emb
repairFactVectors(dataStore, emb)
checkStoredEmbedder(dataStore, emb)
// ----- tool executor (the enabled act allowlist, store-backed) -----
@@ -473,6 +474,36 @@ func seedTools(api ipc.CoreAPI, tools []config.ToolConfig) {
log.Printf("voice: seeded %d act tools from config", n)
}
// repairFactVectors brings stored fact vectors in line with the facts they name
// (#493), once per box, before the embedder marker is even looked at.
//
// Automatic and not a flag, unlike -reembed: only voice-tapped facts are in
// this index, so the work is tens of embeddings rather than the thousands of
// notes that made the backfill a deliberate act. And the box that needs it is
// broken in a way nobody can see — recall answers with the wrong text and
// nothing logs an error — so waiting for an operator to know to run it is how
// the defect survived four restarts in the first place.
func repairFactVectors(dataStore *store.Store, emb router.Embedder) {
if dataStore == nil {
return
}
res, err := dataStore.RepairFactVectors(context.Background(),
// EmbedPassage, the stored side, same as every other writer of these
// vectors.
func(ctx context.Context, text string) ([]float32, error) {
return router.EmbedPassage(ctx, emb, text)
})
if err != nil {
log.Printf("voice: fact vector repair failed, no marker written and nothing half-done — retried next start: %v", err)
return
}
if res.Skipped || res.Rewritten+res.Dropped == 0 {
return
}
log.Printf("voice: fact vector repair — %d re-embedded from the fact they name, %d dropped as voided or superseded, %d already right, took %s (#493)",
res.Rewritten, res.Dropped, res.Kept, res.Took.Round(time.Millisecond))
}
// reembedOnStart is the -reembed flag (set in run()). Opt-in on purpose: see
// runReembed.
var reembedOnStart bool
+4
View File
@@ -54,7 +54,11 @@ func (h *reactiveHandler) phraseSource(ctx context.Context, name, utterance stri
log.Printf("voice: %s: no world model, reading the source back instead", name)
return ""
case err != nil:
// The resident phraser answers this call with its fallback text and the
// error together. Drop the text: these callers hold the passage itself
// and read it back better than "вот что я нашла: <passage>" does.
log.Printf("voice: %s: phrase: %v", name, err)
return ""
}
return reply
}
+1
View File
@@ -20,6 +20,7 @@
"bin_path": "llama-server",
"n_gpu_layers": 99,
"n_ctx": 4096,
"cache_ram_mib": 512,
"timeout": "60s",
"llm_nudges": false
},
@@ -0,0 +1,84 @@
# Where the resident model's 7.9GB of RSS goes (2026-08-03, homesrv)
Measured for Vikunja #499. The deployed llama-server held 7.9GB RSS for a 1.1GB
model file. Half a gigabyte of it was in swap, on a box that also runs
whisper.cpp, piper and the embedder.
## Method
`maven-mavend-1` was stopped for the measurement, with the owner's approval.
Its own binary then ran on the host with the exact deployed command line. That
binary is `/opt/maven/bin/llama-server`, version `1 (4c65955)`, a Vulkan build.
```sh
llama-server -m /mnt/hdd1/llms/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf \
--host 127.0.0.1 --port 18099 -c 4096 -ngl 99 --no-webui
```
RSS was read from `/proc/<pid>/status` after load and after each of 8 distinct
1521-token prompts. `smaps` of the deployed process was read first, from inside
the container, since the host user cannot read another user's maps.
## The cause: the prompt cache, not the weights and not the offload
The startup log says it outright:
```text
srv load_model: prompt cache is enabled, size limit: 8192 MiB
srv llama_server: n_parallel is set to auto, using n_parallel = 4 and kv_unified = true
```
The server saves the full KV state of every idle slot it evicts. It keeps up to
8GiB of those states in host RAM (llama.cpp PR 16391). One saved prompt of 1521
tokens costs 166.377 MiB. That is 112 kiB per token, exactly Qwen3-1.7B's KV
footprint (28 layers x 2 x 1024 dims x 2 bytes).
RSS at rest, and per distinct prompt:
| Prompts served | RSS, default | RSS, `--cache-ram 512` |
|---|---|---|
| 0 (just loaded) | 443 MB | 411 MB |
| 1 | 445 MB | 411 MB |
| 4 | 958 MB | 929 MB |
| 8 | 1641 MB | 932 MB |
Uncapped, RSS climbs about 170MB per distinct prompt and does not stop until
the 8GiB limit. Capped at 512 MiB it plateaus at 932MB from the fourth prompt
on, with the cache holding steady at `3 prompts, 499.132 MiB` and evicting.
The 7.9GB on the running daemon was that climb, weeks of it. Its `smaps` showed
one 6.03GB anonymous mapping at 5.32GB resident plus a 1.45GB mapping at 1.27GB
resident, and only 30MB of file-backed RSS.
## The task's leading guess was wrong
`-ngl 99` on the Vega iGPU costs almost no process RSS. A freshly loaded server
has 95MB of anonymous RSS in total. RADV allocates device memory through the
kernel, outside the process, and the log sees 8202 MiB free on `Vulkan0`. The
weights are mmapped and file-backed, so they are evictable and do not pin RSS. The logit buffer is not visible in the numbers above at all.
## Decision
`--cache-ram 512` is now the default, wired as `phraser.cache_ram_mib` and set
in `deploy/mavend.json`. 512 MiB caps total RSS near 1GB, an eighth of what the
box carried. It still holds three of the 1521-token probes above. Maven's real
routing and phrasing prompts are much shorter, so it holds more of those than
the table suggests. `-c 4096` is untouched, as #499
required. A negative `cache_ram_mib` passes no flag, for a llama-server too old
to know it.
Not changed: `n_parallel = 4`. With `kv_unified = true` the four slots share one
4096-token KV cache, so they do not multiply it.
The other half of #499 was that none of these lines were reachable. mavend
scraped llama-server's stderr for the listen line and discarded it, and never
piped stdout at all. Both streams now go to mavend's log with a `llama:` prefix.
The last 12 startup lines go into the error when the server dies before it
listens.
## Deployed
The `mavenai:latest` image was rebuilt and `maven-mavend-1` recreated the same
day. The daemon's own log now carries the child's startup, it reads
`prompt cache is enabled, size limit: 512 MiB`, and the resident server sat at
439MB RSS after load and 613MB after one served turn.
@@ -0,0 +1,46 @@
# Personal boundary, seed scoring vs possession markers, 2026-08-03
Vikunja #495. `что я говорил про бэкапы?` walked past the personal boundary into
SearXNG and came back answered from a Habr article. The boundary matched
possession words only, so a first-person speech verb was not a personal
question.
## What changed
The boundary now scores the turn's query vector against two frozen seed sets.
It claims the turn when the personal side is nearer than the world side. Seeds
and code are in `cmd/mavend/personalboundary.go`. The possession markers stay as
the offline floor for a handler with no embedder.
A regex speech class was written first and dropped. Russian gives every verb a
dozen surface forms, and the "как я говорил, ..." preamble list has no end. Each
form the lexicon missed was one more question reaching the world.
## Measurement
Embedder: multilingual-e5-small int8, the one homesrv runs. Both sides are
embedded on the query side. Cases are held out, none of them a seed. `make test`
runs the offline part. The scored part is opt-in through `MAVEN_ONNX_LIB`, like
`TestONNXRecall`.
19/19 held-out utterances correct (TestONNXPersonalBoundary)
true positive margins +0.014 to +0.089
nearest true negative -0.005 ("кто такой гагарин")
One case missed during the first pass and is not held out any more: `as i said,
what is the population of india`, +0.008 to the personal side. It is a world seed
now.
The gate is the sign of the difference and nothing tighter. The margins are too
thin for a threshold. The asymmetry favours claiming: a false claim costs one
honest "не знаю", a false pass sends his life to an upstream engine.
`make eval-recall` unchanged, 18/27 answered at gate 0.55. Recall does not touch
this path.
## Not verified
The live probe on the deployed box. The daemon was not rebuilt in this session.
The reply to `что я говорил про бэкапы?` with no matching note is still untested
against a real SearXNG.
@@ -0,0 +1,65 @@
# Recall topic veto, what it costs and what it buys, 2026-08-03
Vikunja #496. The task asked for a cross-language fix. Skip the topic veto in
`memory.RecallAllowed` when the question and the hit are in different scripts.
An English question would then stop losing a Russian note.
No such case exists. No fixture case puts the question and its wanted note in
different scripts. The case the task named is not one either.
en-hard-024
query "what fixed the screen problem"
note "the flicker went away once i swapped the display cable"
Both are English. It is a paraphrase failure, not a language failure. A script
test would not have changed a single case, and neither would a bilingual stem
map.
## What the veto is worth today
Measured with the real embedder, multilingual-e5-small int8, gate 0.55, margin
0.008. The first row is the veto as it ships. The second is `RecallAllowed`
forced to true.
| | cases passing | answered | false recall | silenced by gate |
|---|---|---|---|---|
| veto on | 22/32 | 17/27 | 0/5 | 2 |
| veto off | 22/32 | 18/27 | 1/5 | 1 |
The pass count does not move. The veto trades one true recall for one false one.
It costs `en-hard-024` and it buys `ru-silent-029`:
ru-silent-029
query "во сколько отходит поезд"
note "погулял вдоль реки" 0.835, margin 0.019
The second case counted as silenced by the gate is `ru-home-026` at margin
0.001, which the margin gate stops. The veto has nothing to do with it.
## Why no lexical rule separates the two
`en-hard-024` and `ru-silent-029` are in the same lexical class. Both questions
share zero content words with their hit, and neither carries a first-person
marker. The scores sit on top of each other, 0.826 against 0.835, and so do the
margins, 0.023 against 0.019. Only one thing separates them. A screen problem
and a swapped display cable are the same event. A train and a river walk are
not. The embedder scores that difference at nine thousandths.
So the signal is semantic and the gate is lexical. Any rule cheap enough to sit
in `RecallAllowed` and strong enough to recover `en-hard-024` also re-admits
`ru-silent-029`, which puts false recall back to 1/5.
One near-miss rule was tried on paper and rejected: let the veto pass when the
hit itself is first person. It works on these two, because the English note says
"i swapped" and the Russian note says only "погулял". It is backwards as a
principle. A first-person note is exactly the personal note the veto keeps away
from a world question. The rule would weaken the veto where it was designed to
bite. It survives here only because Russian drops the pronoun.
## Decision
Accept the loss. `en-hard-024` stays silenced and false recall stays 0/5.
The way out is a reranker, not a longer word list. Recall@3 is 85.2% against
recall@1 at 70.4%, so the right note is usually in the returned set and ranked
wrong. That is where the remaining points are, and it is not this task.
+15 -6
View File
@@ -498,12 +498,21 @@ rejects `https://api.openai.com`, and forget really deletes
(`internal/store/memory.go:145` is a real `DELETE`, not a tombstone). Vision is
19/19, speaker 22/22, media 16/16.
**470 got worse.** Both poisoned facts show `voided` on `/history`, and the
defect survives. Re-measured at 15:42, after four restarts: `почему небо синее?`
still answers `какая последняя версия языка Go?` with no `search:` line. What
comes back is the question he typed, not the value the fact held. So the poison
is a vector in the memory index, and `revert` does not remove it. There is
currently no documented way to repair a poisoned box.
**470 got worse, then closed.** Both poisoned facts showed `voided` on
`/history` and the defect survived. Re-measured at 15:42, after four restarts:
`почему небо синее?` still answered `какая последняя версия языка Go?` with no
`search:` line. What came back was the question he typed, not the value the fact
held. So the poison was a vector in the memory index, and `revert` did not
remove it.
Repaired in two parts. 470 stopped the writes: a question is never a fact, and a
void drops the key's vectors. 493 fixed what the index holds. A fact is indexed
as the fact and not as the utterance, and a correction drops its superseded
vector too.
A poisoned box now repairs itself on the next start. `RepairFactVectors`
re-embeds every fact vector from the fact it names, and deletes the voided and
superseded ones. It runs once, guarded by a marker, and logs what it did.
---
+6
View File
@@ -1276,6 +1276,12 @@ type PhraserConfig struct {
NCtx int `json:"n_ctx,omitempty"`
Timeout Duration `json:"timeout,omitempty"`
// CacheRAMMiB bounds llama-server's prompt cache. Omitted ⇒ 512 MiB, which
// is what keeps the resident model near 1 GB of RSS instead of the 7.9 GB
// measured on 2026-08-03. Set it to -1 to pass no flag at all and let the
// server apply its own 8 GiB default. See phraser.Config.CacheRAMMiB.
CacheRAMMiB int `json:"cache_ram_mib,omitempty"`
// LLMNudges — let the model word nudges again. Off by default: nudges are
// worded from hand-written Russian templates now (the model broke the
// persona and invented units). Chat, query and reminder phrasing always go
+134
View File
@@ -0,0 +1,134 @@
package memory
import (
"strings"
"unicode"
)
// stopwords — words that carry no topic. A question and a note that share only
// these share nothing: "почему небо синее" and "сеть какая-то медленная" both
// contain "какая"-shaped filler and are about different worlds.
var stopwords = map[string]bool{
// interrogatives and demonstratives
"что": true, "чего": true, "какой": true, "какая": true, "какое": true,
"какие": true, "каких": true, "кто": true, "кого": true, "кому": true,
"почему": true, "зачем": true, "где": true, "куда": true, "откуда": true,
"когда": true, "сколько": true, "как": true, "то": true, "это": true,
"этот": true, "тот": true, "там": true, "тут": true, "такой": true,
// pronouns — every sentence he says is about him, so "я" is not a topic
"я": true, "меня": true, "мне": true, "мой": true, "моя": true, "мои": true,
"ты": true, "тебя": true, "тебе": true, "твой": true, "он": true, "она": true,
"они": true, "мы": true, "себя": true, "свой": true,
// prepositions, conjunctions, particles, copulas
"в": true, "во": true, "на": true, "с": true, "со": true, "у": true,
"о": true, "об": true, "про": true, "за": true, "из": true, "по": true,
"до": true, "от": true, "для": true, "над": true, "под": true, "при": true,
"и": true, "а": true, "но": true, "или": true, "же": true, "ли": true,
"не": true, "ни": true, "бы": true, "был": true, "была": true, "было": true,
"быть": true, "есть": true, "был-ли": true, "уже": true, "ещё": true,
"еще": true, "так": true, "вот": true, "там-же": true,
// English filler, for the mixed utterances he does say
"the": true, "a": true, "an": true, "is": true, "are": true, "was": true,
"were": true, "be": true, "of": true, "in": true, "on": true, "at": true,
"to": true, "for": true, "about": true, "and": true, "or": true, "not": true,
"what": true, "who": true, "why": true, "when": true, "where": true,
"which": true, "how": true, "i": true, "my": true, "me": true, "it": true,
"this": true, "that": true,
}
// firstPerson — the words that make an utterance a question about his own
// life. Not possession only: "как я восстановил конфиги" owns nothing and is
// still about him.
var firstPerson = map[string]bool{
"я": true, "меня": true, "мне": true, "мной": true, "мой": true,
"моя": true, "моё": true, "мое": true, "мои": true, "моего": true,
"моей": true, "моих": true, "моим": true, "себя": true, "свой": true,
"своя": true, "свои": true, "своего": true, "мною": true,
"i": true, "me": true, "my": true, "mine": true, "myself": true,
}
// RecallAllowed is the second half of the recall gate (#470). A hit that
// cleared the score and margin gate may still be about something else
// entirely: the held-out fixture puts the right note at 0.791-0.890 and the
// must-be-silent cases at 0.795-0.835, so no threshold sits between them, and
// a note about his slow network answered "почему небо синее?".
//
// The veto applies only to a question that mentions nothing of his. That
// restriction is what keeps the fix from costing more than it saves: recall
// exists to find the note whose words he no longer remembers, and demanding a
// shared word of every recall silenced four true recalls on the fixture to
// kill one false one. A question about his own life keeps the embedder alone
// as its judge. A question about the world has to name something the memory
// actually mentions.
//
// The veto's price was re-measured on 2026-08-03 (#496,
// docs/evals/2026-08-03-recall-topic-veto.md). It costs one true recall and
// buys one false one, and the fixture pass count is the same either way. The
// lost case is an English paraphrase, not the cross-language loss it was
// reported as, and the fixture has no cross-language case at all. Do not add a
// script test or a bilingual stem map for it — both are no-ops here. The
// separating signal is semantic and belongs in a reranker, not in this file.
func RecallAllowed(query, text string) bool {
if mentionsHim(query) {
return true
}
return SharesContentWord(query, text)
}
func mentionsHim(query string) bool {
for _, w := range strings.FieldsFunc(strings.ToLower(query), func(r rune) bool {
return !unicode.IsLetter(r) && !unicode.IsDigit(r)
}) {
if firstPerson[w] {
return true
}
}
return false
}
// SharesContentWord reports whether query and text have at least one topic
// word in common, after dropping the words that carry no topic. Stems are
// compared, so the note and the question do not have to inflect alike.
func SharesContentWord(query, text string) bool {
q := contentWords(query)
if len(q) == 0 {
// Nothing to compare — a question made entirely of filler. The score
// gate is then the only judge it can have.
return true
}
t := contentWords(text)
for _, a := range q {
for _, b := range t {
if a == b || sameStem(a, b) {
return true
}
}
}
return false
}
func contentWords(s string) []string {
var out []string
for _, w := range strings.FieldsFunc(strings.ToLower(s), func(r rune) bool {
return !unicode.IsLetter(r) && !unicode.IsDigit(r)
}) {
if !stopwords[w] {
out = append(out, w)
}
}
return out
}
// sameStem is inflection and derivation tolerance: Russian marks case and
// tense on the ending, and the note and the question rarely use the same form.
// "воду" and "вода" are the same water, "кормить" and "корм" the same feeding.
// All but the last rune of the shorter word must match, and never fewer than
// three, which is what keeps "сеть" clear of "сеанс".
func sameStem(a, b string) bool {
ar, br := []rune(a), []rune(b)
n := min(len(ar), len(br)) - 1
if n < 3 {
return false
}
return string(ar[:n]) == string(br[:n])
}
+56
View File
@@ -0,0 +1,56 @@
package memory
import "testing"
func TestRecallAllowed(t *testing.T) {
cases := []struct {
name string
query, text string
want bool
}{
// The #470 shape: a world question and a note about his box.
{"world question, unrelated note", "почему небо синее", "сеть какая-то медленная", false},
{"world question, unrelated fact", "какая столица Франции", "какая последняя версия языка Go", false},
{"silent fixture case", "во сколько отходит поезд", "бэкап запускается в три ночи", false},
// A world question that does name the topic keeps its answer.
{"world question, same topic", "какой поезд идёт в Минск", "поезда в Минск ходят утром", true},
// A question about his own life is judged by the embedder alone,
// because recall exists for words he no longer remembers.
{"about him, no shared word", "во сколько я обычно засыпаю", "ложусь около одиннадцати", true},
{"about him, english", "which colour scheme do i like", "тёмная тема везде", true},
// Inflection must not break a match.
{"inflected", "чем кормить кота", "корм для кота в шкафу", true},
}
for _, c := range cases {
if got := RecallAllowed(c.query, c.text); got != c.want {
t.Errorf("%s: RecallAllowed(%q, %q) = %v, want %v", c.name, c.query, c.text, got, c.want)
}
}
}
// The known cost of the veto and the thing that pays for it, both measured on
// the held-out fixture with the real embedder (#496,
// docs/evals/2026-08-03-recall-topic-veto.md). The two are one lexical class:
// zero shared content words, no first-person marker, scores 0.826 against 0.835
// and margins 0.023 against 0.019. Recovering the first re-admits the second,
// which puts false recall back to 1/5. Anyone loosening the veto has to move
// the first line without moving the second.
func TestRecallVetoTradeIsPinned(t *testing.T) {
if RecallAllowed("what fixed the screen problem", "the flicker went away once i swapped the display cable") {
t.Error("en-hard-024 is expected to stay vetoed — if this passes now, re-measure false recall before celebrating")
}
if RecallAllowed("во сколько отходит поезд", "погулял вдоль реки") {
t.Error("ru-silent-029 must stay vetoed — this is the false recall the veto exists to stop")
}
}
// A question made only of filler has no topic word to match on, and the score
// gate is then the only judge it can have.
func TestRecallAllowedFallsBackWhenNothingToCompare(t *testing.T) {
if !RecallAllowed("что это", "сеть какая-то медленная") {
t.Error("a question with no content word must not be vetoed")
}
}
+11 -3
View File
@@ -378,7 +378,7 @@ func scoreCase(ctx context.Context, emb router.Embedder, newStore NewStore, minS
if len(hits) > 1 {
o.Margin = hits[0].Score - hits[1].Score
}
o.Recalled = bestRecall(hits, minScore, minMargin)
o.Recalled = bestRecall(c.Query, hits, minScore, minMargin)
}
for i, h := range hits {
if h.ID != c.Want {
@@ -424,11 +424,19 @@ func rankNote(inTop3 bool) string {
// is not importable; recalleval_test.go asserts the two agree in behaviour.
// The daemon returns the whole hit (a note and a fact are said differently);
// the harness only scores what came back, so it keeps returning the text.
func bestRecall(results []memory.Result, minScore, minMargin float64) string {
// bestRecall mirrors the daemon's gate in cmd/mavend/recall.go, including the
// topic veto added for #470: a score that clears the gate still has to be
// about what he asked. Keep the two in step — a fixture that measures a
// weaker gate than the daemon runs flatters it.
func bestRecall(query string, results []memory.Result, minScore, minMargin float64) string {
if !memory.Confident(results, minScore, minMargin) {
return ""
}
return results[0].Meta["text"]
text := results[0].Meta["text"]
if !memory.RecallAllowed(query, text) {
return ""
}
return text
}
func bump(m map[string]TagStat, key string, pass bool) {
+15 -8
View File
@@ -140,21 +140,22 @@ func words(s string) []string {
// TestBestRecallMatchesDaemon — the harness duplicates bestRecall from
// cmd/mavend/recall.go (package main is not importable). This pins the copy to
// the original's three rules: no hits, below the gate, or no text ⇒ silence.
// the original's rules: no hits, below the gate, no text, or no shared topic
// word ⇒ silence.
func TestBestRecallMatchesDaemon(t *testing.T) {
if got := bestRecall(nil, 0.55, 0); got != "" {
if got := bestRecall("чай", nil, 0.55, 0); got != "" {
t.Errorf("no hits: got %q, want silence", got)
}
low := []memory.Result{{ID: "a", Score: 0.4, Meta: map[string]string{"text": "чай"}}}
if got := bestRecall(low, 0.55, 0); got != "" {
if got := bestRecall("чай", low, 0.55, 0); got != "" {
t.Errorf("below gate: got %q, want silence", got)
}
noText := []memory.Result{{ID: "a", Score: 0.9, Meta: map[string]string{}}}
if got := bestRecall(noText, 0.55, 0); got != "" {
if got := bestRecall("чай", noText, 0.55, 0); got != "" {
t.Errorf("no text: got %q, want silence", got)
}
ok := []memory.Result{{ID: "a", Score: 0.9, Meta: map[string]string{"text": "чай"}}}
if got := bestRecall(ok, 0.55, 0); got != "чай" {
if got := bestRecall("чай", ok, 0.55, 0); got != "чай" {
t.Errorf("above gate: got %q, want %q", got, "чай")
}
// Margin: a close runner-up means the embedder cannot tell the two apart,
@@ -163,17 +164,23 @@ func TestBestRecallMatchesDaemon(t *testing.T) {
{ID: "a", Score: 0.86, Meta: map[string]string{"text": "чай"}},
{ID: "b", Score: 0.85, Meta: map[string]string{"text": "кофе"}},
}
if got := bestRecall(close, 0.55, 0.03); got != "" {
if got := bestRecall("чай", close, 0.55, 0.03); got != "" {
t.Errorf("thin margin: got %q, want silence", got)
}
if got := bestRecall(close, 0.55, 0); got != "чай" {
if got := bestRecall("чай", close, 0.55, 0); got != "чай" {
t.Errorf("margin off: got %q, want %q", got, "чай")
}
// The topic veto (#470): the score is fine and the note is about
// something else.
offTopic := []memory.Result{{ID: "a", Score: 0.9, Meta: map[string]string{"text": "сеть какая-то медленная"}}}
if got := bestRecall("почему небо синее", offTopic, 0.55, 0); got != "" {
t.Errorf("off topic: got %q, want silence", got)
}
clear := []memory.Result{
{ID: "a", Score: 0.86, Meta: map[string]string{"text": "чай"}},
{ID: "b", Score: 0.70, Meta: map[string]string{"text": "кофе"}},
}
if got := bestRecall(clear, 0.55, 0.03); got != "чай" {
if got := bestRecall("чай", clear, 0.55, 0.03); got != "чай" {
t.Errorf("wide margin: got %q, want %q", got, "чай")
}
}
+55 -6
View File
@@ -26,20 +26,22 @@ import (
"time"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/router"
)
//go:embed talk_v1.json
var talkFixtureJSON []byte
// The three phrasing paths under test. Values match the fixture's "path" field.
// The phrasing paths under test. Values match the fixture's "path" field.
const (
PathChat = "chat" // PhraseChat
PathQuery = "query" // PhraseQuery with notes
PathKnowledge = "knowledge" // PhraseQuery with no notes
PathReply = "reply" // PhraseReply, the reactive confirmation
)
// TalkPaths — report order.
var TalkPaths = []string{PathChat, PathQuery, PathKnowledge}
var TalkPaths = []string{PathChat, PathQuery, PathKnowledge, PathReply}
// TalkCheckNames — the checks that apply to a free-form reply, in report order.
// Deliberately a subset of CheckNames: length, mood and "no questions" are nudge
@@ -58,17 +60,30 @@ var TalkCheckNames = []string{
// WantAny is the on-topic contract: at least one lowercased fragment must appear
// in the reply. Fragments are stems ("пароль" → "парол") so declension does not
// defeat them.
//
// Intent, Key and Value carry the reply path's decision: that path is phrased
// from what the router already resolved, not from the raw utterance. Utterance
// stays filled anyway, because it is what a human reads in the report.
type TalkCase struct {
ID string `json:"id"`
Path string `json:"path"`
Utterance string `json:"utterance"`
History []string `json:"history,omitempty"`
Notes []string `json:"notes,omitempty"`
Intent string `json:"intent,omitempty"`
Key string `json:"key,omitempty"`
Value string `json:"value,omitempty"`
WantAny []string `json:"want_any"`
Tags []string `json:"tags,omitempty"`
Note string `json:"note,omitempty"`
}
// TalkSchemaVersion — the version this loader understands. Separate from the
// nudge fixture's SchemaVersion: the two fixtures have different shapes and
// change on different days, and one shared constant would force a bump on the
// fixture that did not move.
const TalkSchemaVersion = 1
// TalkFixture — the versioned envelope, same gating as Fixture.
type TalkFixture struct {
SchemaVersion int `json:"schema_version"`
@@ -83,8 +98,8 @@ func LoadTalk() (TalkFixture, error) {
if err := json.Unmarshal(talkFixtureJSON, &f); err != nil {
return TalkFixture{}, fmt.Errorf("parse talk fixture: %w", err)
}
if f.SchemaVersion != SchemaVersion {
return TalkFixture{}, fmt.Errorf("talk fixture schema_version %d, want %d", f.SchemaVersion, SchemaVersion)
if f.SchemaVersion != TalkSchemaVersion {
return TalkFixture{}, fmt.Errorf("talk fixture schema_version %d, want %d", f.SchemaVersion, TalkSchemaVersion)
}
if len(f.Cases) == 0 {
return TalkFixture{}, fmt.Errorf("talk fixture has no cases")
@@ -92,13 +107,27 @@ func LoadTalk() (TalkFixture, error) {
return f, nil
}
// Talker — the two methods a conversational path must have to be scorable.
// *phraser.LLMPhraser satisfies it; same trick as Nudger.
// Talker — the methods a conversational path must have to be scorable.
// *phraser.LLMPhraser satisfies the first two; *phraser.Replier satisfies the
// third, so a run that scores all four paths passes a Pair.
type Talker interface {
PhraseChat(ctx context.Context, utterance string, history []dialogue.Turn) (string, error)
PhraseQuery(ctx context.Context, utterance string, notes []string) (string, error)
}
// Confirmer — the reply path. *phraser.Replier satisfies it.
type Confirmer interface {
PhraseReply(ctx context.Context, d router.Decision) (string, error)
}
// Pair joins the two objects the daemon wires separately — the phraser and the
// replier — so one ScoreTalk call covers every path Maven speaks through. A bare
// Talker still works; its reply cases score as errors, which is honest.
type Pair struct {
Talker
Confirmer
}
// TalkOutcome — one scored case.
type TalkOutcome struct {
Case TalkCase
@@ -194,10 +223,30 @@ func (c TalkCase) run(ctx context.Context, t Talker) (string, error) {
return t.PhraseQuery(ctx, c.Utterance, c.Notes)
case PathKnowledge:
return t.PhraseQuery(ctx, c.Utterance, nil)
case PathReply:
conf, ok := t.(Confirmer)
if !ok {
return "", fmt.Errorf("target cannot phrase replies — pass a Pair")
}
return conf.PhraseReply(ctx, c.decision())
}
return "", fmt.Errorf("unknown path %q", c.Path)
}
// decision rebuilds what the router would have handed the replier. Text is the
// utterance for a note or a reminder, which is what the router puts there.
func (c TalkCase) decision() router.Decision {
return router.Decision{
Intent: router.Intent(c.Intent),
Slots: router.Slots{
Key: c.Key,
Value: c.Value,
Text: c.Utterance,
HasKey: c.Key != "",
},
}
}
func (c TalkCase) turns() []dialogue.Turn {
turns := make([]dialogue.Turn, 0, len(c.History))
for _, h := range c.History {
+28 -18
View File
@@ -11,6 +11,7 @@ import (
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
)
// perPathMinimum — the resolution floor. A per-path score built on a handful of
@@ -36,6 +37,10 @@ func TestTalkFixture(t *testing.T) {
switch c.Path {
case PathChat, PathQuery, PathKnowledge:
case PathReply:
if c.Intent == "" {
t.Errorf("%s: reply case has no intent — the replier is phrased from the decision", c.ID)
}
default:
t.Errorf("%s: unknown path %q", c.ID, c.Path)
}
@@ -69,10 +74,15 @@ type fakeTalker struct{ reply string }
func (f fakeTalker) PhraseChat(context.Context, string, []dialogue.Turn) (string, error) {
return f.reply, nil
}
func (f fakeTalker) PhraseQuery(context.Context, string, []string) (string, error) {
return f.reply, nil
}
func (f fakeTalker) PhraseReply(context.Context, router.Decision) (string, error) {
return f.reply, nil
}
// TestScoreTalkCounts — a reply that fails on purpose must be counted on every
// path, so a real run cannot report a hidden zero.
func TestScoreTalkCounts(t *testing.T) {
@@ -104,7 +114,7 @@ func TestScoreTalkCounts(t *testing.T) {
}
}
// TestLLMTalkBaseline — the resident model on the three conversational paths.
// TestLLMTalkBaseline — the resident model on all four phrasing paths.
// Opt-in exactly like TestLLMPhrasingBaseline: CI has no model and a run costs
// minutes on the CPU target.
//
@@ -132,32 +142,32 @@ func TestLLMTalkBaseline(t *testing.T) {
p := phraser.NewLLMPhraserAt(base, cfg)
defer p.Close()
// Unreachable server is fatal here, not a logged warning, and that differs
// from the nudge test on purpose. PhraseNudge returns its errors, so a dead
// server there shows up honestly in the Errors column. PhraseChat and
// PhraseQuery do NOT: they swallow every failure and return a canned string
// ("поговорили.", "не знаю.", "вот что я нашла: …"). So on these three paths
// a dead server produces a full report with 0 errors and a terrible score —
// a number that looks like bad phrasing and is really no phrasing at all.
// Refusing to score without a confirmed model is the only guard available
// until the phraser reports its failures (Vikunja #397).
// The model id names the run in the report. Since Vikunja #397 every path
// returns its errors, so a server that dies mid-run shows up in the Errors
// column instead of scoring as bad phrasing — the before-and-after probe that
// used to stand in for that is gone.
model, err := llm.ModelID(ctx, base)
if err != nil {
t.Fatalf("no model at %s: %v — refusing to score, these paths hide their errors "+
"and would report a plausible-looking result off a dead server", base, err)
t.Fatalf("no model at %s: %v", base, err)
}
t.Logf("scoring model %s at %s", model, base)
rep, err := ScoreTalk(ctx, "llm ("+model+", built-in persona)", p, f)
// The reply path is a separate object in the daemon too: the phraser owns its
// own llama-server, the replier is handed an llm.Client. Pair scores both.
block := func() string { return persona.Facts{}.Block(time.Now()) }
target := Pair{Talker: p, Confirmer: phraser.NewReplier(llm.New(base, cfg.Timeout), block)}
rep, err := ScoreTalk(ctx, "llm ("+model+", built-in persona)", target, f)
if err != nil {
t.Fatalf("ScoreTalk: %v", err)
}
t.Log("\n" + rep.String() + "\nreplies:\n" + rep.Replies() + "\nfailures:\n" + rep.Failures())
// And again afterwards: the run takes minutes, and a server that died or got
// OOM-killed halfway through would leave the first cases scored and the rest
// silently canned. Checking only at the start would not catch that.
if _, err := llm.ModelID(ctx, base); err != nil {
t.Fatalf("model at %s went away during the run: %v — the score above is not trustworthy", base, err)
// A run where nothing was phrased is not a low score, it is no measurement.
if rep.Errors == rep.Total {
t.Fatalf("every case errored — nothing was measured, the score above is not a phrasing result")
}
if rep.Errors > 0 {
t.Logf("%d/%d cases errored — those are model failures, not phrasing failures", rep.Errors, rep.Total)
}
}
+84
View File
@@ -222,6 +222,90 @@
"utterance": "почему гром слышно позже молнии?",
"want_any": ["звук", "све", "быстр", "гром", "молни"],
"tags": ["general"]
},
{
"id": "reply-fact-coffee",
"path": "reply",
"intent": "fact",
"key": "кофе",
"value": "закончился",
"utterance": "кофе закончился",
"want_any": ["коф"],
"tags": ["fact"],
"note": "The plainest confirmation there is, and the sentence he hears most often."
},
{
"id": "reply-fact-weight",
"path": "reply",
"intent": "fact",
"key": "вес",
"value": "82",
"utterance": "мой вес 82",
"want_any": ["вес", "82"],
"tags": ["fact", "number"],
"note": "A number must survive into the confirmation; a paraphrase that drops it is useless."
},
{
"id": "reply-fact-pill",
"path": "reply",
"intent": "fact",
"key": "таблетки",
"value": "выпил",
"utterance": "таблетки выпил",
"want_any": ["таблетк"],
"tags": ["fact", "feminine"],
"note": "He says 'выпил', masculine and about himself. She must not copy the form onto herself."
},
{
"id": "reply-note-router",
"path": "reply",
"intent": "note",
"utterance": "роутер перезагружается сам по ночам",
"want_any": ["роутер"],
"tags": ["note"]
},
{
"id": "reply-note-long",
"path": "reply",
"intent": "note",
"utterance": "если диск снова отвалится, посмотреть кабель, а не контроллер, в прошлый раз был кабель",
"want_any": ["диск", "кабел"],
"tags": ["note", "length"],
"note": "A long note baits a long confirmation. One sentence is the contract."
},
{
"id": "reply-reminder-evening",
"path": "reply",
"intent": "reminder",
"utterance": "напомни вечером полить цветы",
"want_any": ["цвет", "полит", "вечер"],
"tags": ["reminder"]
},
{
"id": "reply-reminder-tomorrow",
"path": "reply",
"intent": "reminder",
"utterance": "напомни завтра позвонить в поликлинику",
"want_any": ["поликлиник", "позвон", "звон"],
"tags": ["reminder"]
},
{
"id": "reply-formality-bait",
"path": "reply",
"intent": "note",
"utterance": "запишите пожалуйста что счётчики я сдал",
"want_any": ["счётчик", "счетчик"],
"tags": ["note", "persona-bait", "address"],
"note": "Polite plural in the input. The confirmation must still be на ты."
},
{
"id": "reply-question-bait",
"path": "reply",
"intent": "note",
"utterance": "надо купить фильтр для воды, не помню какой",
"want_any": ["фильтр"],
"tags": ["note", "no-question"],
"note": "An unresolved note invites her to ask which filter. A confirmation does not ask."
}
]
}
+72
View File
@@ -0,0 +1,72 @@
package phraser
import (
"context"
"net/http"
"net/http/httptest"
"strings"
"testing"
)
// A dead server must be distinguishable from bad phrasing. Both PhraseChat and
// PhraseQuery keep the turn alive with canned text — ChatFallback, "не знаю.",
// "вот что я нашла: …" — and every one of those is also a legitimate reply, so
// the text alone cannot say which happened. The error is the only signal, and
// before Vikunja #397 it was dropped: the talk scorer reported a full run with
// zero errors off a server that answered nothing.
func TestPhrasingReportsTheFailureWithTheFallback(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Error(w, "model not loaded", http.StatusServiceUnavailable)
}))
t.Cleanup(srv.Close)
p := NewLLMPhraserAt(srv.URL, Config{})
cases := []struct {
name string
call func() (string, error)
want string
}{
{"chat", func() (string, error) {
return p.PhraseChat(context.Background(), "как дела", nil)
}, ChatFallback},
{"knowledge", func() (string, error) {
return p.PhraseQuery(context.Background(), "кто написал войну и мир", nil)
}, "не знаю."},
{"evidence", func() (string, error) {
return p.PhraseQuery(context.Background(), "сколько воды я выпил", []string{"два литра"})
}, "вот что я нашла: два литра"},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
got, err := c.call()
if err == nil {
t.Fatalf("no error from a dead server; the scorer would count this as bad phrasing")
}
if got != c.want {
t.Errorf("fallback text = %q, want %q — the daemon still has to say something", got, c.want)
}
})
}
}
// An empty answer is a failure too: the server is up and produced no tokens,
// which is not an answer and must not score as one.
func TestEmptyKnowledgeAnswerIsAnError(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/json")
w.Write([]byte(`{"choices":[{"message":{"content":""}}]}`))
}))
t.Cleanup(srv.Close)
p := NewLLMPhraserAt(srv.URL, Config{})
got, err := p.PhraseQuery(context.Background(), "кто написал войну и мир", nil)
if err == nil {
t.Fatal("an empty response scored as an answer")
}
if got != "не знаю." {
t.Errorf("fallback text = %q, want \"не знаю.\"", got)
}
if !strings.Contains(err.Error(), "empty") {
t.Errorf("error = %v; want it to name the empty response", err)
}
}
+129 -52
View File
@@ -1,13 +1,16 @@
package phraser
import (
"bufio"
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"io"
"log"
"net/http"
"os"
"os/exec"
"regexp"
"strings"
@@ -24,6 +27,11 @@ import (
var listenRE = regexp.MustCompile(`listening on (https?://\S+)`)
// errEmptyResponse — the server answered and said nothing. Separate from a
// transport failure: the model is up and produced no tokens, which is still not
// an answer and must not score as one.
var errEmptyResponse = errors.New("phraser: empty response from the model")
type LLMPhraser struct {
cfg Config
client *http.Client
@@ -83,6 +91,19 @@ type Config struct {
NCtx int
Timeout time.Duration
// CacheRAMMiB bounds llama-server's prompt cache, which is what actually ate
// this box. Measured on homesrv 2026-08-03: the server's own default limit is
// 8192 MiB, it stores the full KV state of every idle slot it evicts (112 kiB
// per token, so 166 MiB for one 1521-token prompt), and RSS climbed by that
// much per distinct prompt until it hit 7.9 GB and half a gigabyte went to
// swap. Weights are only 1.1 GB and mmapped, and -ngl 99 costs almost no RSS
// because RADV keeps device memory outside the process.
//
// 0 ⇒ the flag is not passed and the server's own 8 GiB default applies. That
// is the escape hatch for a llama-server too old to know --cache-ram, not a
// recommendation. See docs/evals/2026-08-03-llama-prompt-cache.md.
CacheRAMMiB int
// ContextBlock renders the shared context block (who he is, how to
// address him, the time) fresh for each turn. See internal/persona.
// nil ⇒ no block, the prompts stand alone.
@@ -116,7 +137,9 @@ func DefaultConfig(modelPath string) Config {
Listen: "127.0.0.1:0",
NGpuLayers: -1,
NCtx: 2048,
Timeout: 30 * time.Second,
// 512 MiB caps total RSS near 1 GB and still holds several recent prompts.
CacheRAMMiB: 512,
Timeout: 30 * time.Second,
}
}
@@ -225,8 +248,10 @@ func spawnLlamaServer(ctx context.Context, cfg Config) (backend, error) {
return p, nil
}
func startLlamaProc(ctx context.Context, cfg Config) (*llamaProc, error) {
p := &llamaProc{}
// llamaArgs is the command line for one resident server. It is a function and
// not an inline literal because kill-maven.sh's orphan sweep matches against
// this exact line, and a test pins the two together.
func llamaArgs(cfg Config) []string {
args := []string{
"-m", cfg.ModelPath,
"--host", "127.0.0.1",
@@ -235,7 +260,15 @@ func startLlamaProc(ctx context.Context, cfg Config) (*llamaProc, error) {
"-ngl", fmt.Sprintf("%d", cfg.NGpuLayers),
"--no-webui",
}
cmd := exec.CommandContext(ctx, cfg.BinPath, args...)
if cfg.CacheRAMMiB > 0 {
args = append(args, "--cache-ram", fmt.Sprintf("%d", cfg.CacheRAMMiB))
}
return args
}
func startLlamaProc(ctx context.Context, cfg Config) (*llamaProc, error) {
p := &llamaProc{}
cmd := exec.CommandContext(ctx, cfg.BinPath, llamaArgs(cfg)...)
// Pdeathsig: the kernel SIGKILLs llama-server the moment mavend dies — by
// ANY means, including SIGKILL/OOM/panic where our Close() never runs. Without
// it a hard-killed mavend orphans its llama-server (reparented to init, keeps
@@ -246,63 +279,104 @@ func startLlamaProc(ctx context.Context, cfg Config) (*llamaProc, error) {
cmd.SysProcAttr = &syscall.SysProcAttr{Setpgid: true, Pdeathsig: syscall.SIGKILL}
p.cmd = cmd
stderr, err := cmd.StderrPipe()
// One pipe for both streams. llama.cpp writes its buffer sizes, KV-cache
// layout and offload lines to stderr and its request log to stdout, and
// stdout used to go nowhere at all — so nothing about the model's memory was
// diagnosable from a running box. Both ends land in mavend's log now.
pr, pw, err := os.Pipe()
if err != nil {
return nil, fmt.Errorf("llm: stderr pipe: %w", err)
return nil, fmt.Errorf("llm: output pipe: %w", err)
}
cmd.Stdout = pw
cmd.Stderr = pw
if err := cmd.Start(); err != nil {
stderr.Close()
pr.Close()
pw.Close()
return nil, fmt.Errorf("llm: start: %w", err)
}
// The child holds the only other reference to the write end. Dropping ours
// is what makes the reader see EOF when the child dies.
pw.Close()
portCh := make(chan string, 1)
errCh := make(chan error, 1)
tail := &lineTail{}
p.wg.Add(1)
go func() {
defer p.wg.Done()
buf := make([]byte, 4096)
var leftover []byte
for {
n, err := stderr.Read(buf)
if n > 0 {
data := append(leftover, buf[:n]...)
lines := bytes.Split(data, []byte("\n"))
for _, line := range lines[:len(lines)-1] {
if m := listenRE.FindSubmatch(line); len(m) > 1 {
addr := string(m[1])
portCh <- addr
close(portCh)
}
defer pr.Close()
sc := bufio.NewScanner(pr)
// llama.cpp prints one prompt per line and a prompt can be long.
sc.Buffer(make([]byte, 0, 64*1024), 1024*1024)
listening := false
for sc.Scan() {
line := sc.Bytes()
log.Printf("llama: %s", line)
if !listening {
tail.add(string(line))
if m := listenRE.FindSubmatch(line); len(m) > 1 {
listening = true
portCh <- string(m[1])
close(portCh)
}
leftover = lines[len(lines)-1]
}
if err != nil {
errCh <- err
return
}
}
err := sc.Err()
if err == nil {
err = io.EOF
}
errCh <- err
}()
fail := func(err error) (*llamaProc, error) {
_ = cmd.Process.Kill()
_ = cmd.Wait()
return nil, err
}
select {
case addr := <-portCh:
p.base = addr
return p, nil
case err := <-errCh:
_ = cmd.Process.Kill()
_ = cmd.Wait()
return nil, fmt.Errorf("llm: server output: %w", err)
// The tail is the whole diagnosis when the server dies during load: bare
// "EOF" never said which layer or which allocation it choked on.
return fail(fmt.Errorf("llm: server output: %w; last output: %s", err, tail.String()))
case <-ctx.Done():
_ = cmd.Process.Kill()
_ = cmd.Wait()
return nil, ctx.Err()
return fail(ctx.Err())
case <-time.After(60 * time.Second):
_ = cmd.Process.Kill()
_ = cmd.Wait()
return nil, fmt.Errorf("llm: server did not start within 60s")
return fail(fmt.Errorf("llm: server did not start within 60s; last output: %s", tail.String()))
}
}
// lineTail keeps the last few startup lines so a server that dies before it
// listens can say why in the error, not just "EOF". Written by the reader
// goroutine and read by whoever gives up on startup, so it takes a lock.
type lineTail struct {
mu sync.Mutex
lines []string
}
const lineTailMax = 12
func (t *lineTail) add(line string) {
t.mu.Lock()
defer t.mu.Unlock()
t.lines = append(t.lines, line)
if len(t.lines) > lineTailMax {
t.lines = t.lines[len(t.lines)-lineTailMax:]
}
}
func (t *lineTail) String() string {
t.mu.Lock()
defer t.mu.Unlock()
if len(t.lines) == 0 {
return "(no output)"
}
return strings.Join(t.lines, " | ")
}
// BaseURL is the llama-server this phraser talks to right now. It changes when
// the model is swapped, so callers that cache it must register an observer
// (OnSwap) rather than keeping the string forever.
@@ -360,8 +434,11 @@ func (p *LLMPhraser) PhraseNudge(ctx context.Context, c loop.Candidate) (deliver
}
// PhraseQuery prompts the LLM with the user's utterance and matching notes to
// compose a natural answer. Falls back to "вот что я нашла: <notes>" on any
// LLM error — better to give the raw data than silence.
// compose a natural answer. On any LLM error it returns the fallback text —
// "вот что я нашла: <notes>", or "не знаю." with no notes — and the error
// together. The daemon uses the text and keeps the turn alive; a caller that is
// measuring counts the failure. Until Vikunja #397 the error was dropped, so a
// dead server scored as bad phrasing.
func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []string) (string, error) {
// Blank sources are no sources. A caller that hands over one empty string —
// a page that fetched to nothing, a snippet trimmed away — used to take the
@@ -371,13 +448,15 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
if len(notes) == 0 {
sys, prompt := p.knowledgePrompt(utterance)
resp, err := p.chatWithSystem(ctx, sys, prompt, 768)
if err != nil || resp == "" {
return "не знаю.", nil
if err != nil {
return "не знаю.", fmt.Errorf("phrase query (knowledge): %w", err)
}
if resp == "" {
return "не знаю.", errEmptyResponse
}
text, _, perr := parseResponseMood(resp)
if perr != nil {
log.Printf("phraser: PhraseQuery: %v", perr)
return "не знаю.", nil
return "не знаю.", fmt.Errorf("phrase query (knowledge): %w", perr)
}
if text != "" {
return text, nil
@@ -389,13 +468,12 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
text, _, perr := parseResponseMood(resp)
if err != nil || perr != nil {
// Read the notes out rather than ship a broken fragment.
if perr != nil {
log.Printf("phraser: PhraseQuery: %v", perr)
cause := err
if cause == nil {
cause = perr
}
if len(notes) == 1 {
return "вот что я нашла: " + notes[0], nil
}
return "вот что я нашла: " + strings.Join(notes, "; "), nil
return "вот что я нашла: " + strings.Join(notes, "; "),
fmt.Errorf("phrase query (evidence): %w", cause)
}
if text != "" {
return text, nil
@@ -404,8 +482,9 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
}
// PhraseChat uses the LLM to respond conversationally, building a multi-turn
// message array from dialogue history + the current user utterance. Falls back
// to a simple greeting on any LLM error — better to say something than nothing.
// message array from dialogue history + the current user utterance. On any LLM
// error it returns both ChatFallback and the error, on the same rule as
// PhraseQuery: the fallback keeps the turn alive, the error stays visible.
func (p *LLMPhraser) PhraseChat(ctx context.Context, utterance string, history []dialogue.Turn) (string, error) {
sys := chatSystemPrompt(p.cfg.ContextBlock)
msgs := []chatMsg{
@@ -422,13 +501,11 @@ func (p *LLMPhraser) PhraseChat(ctx context.Context, utterance string, history [
resp, err := p.chatWithMessages(ctx, msgs, 768)
if err != nil {
log.Printf("phraser: PhraseChat: %v", err)
return "поговорили.", nil
return ChatFallback, fmt.Errorf("phrase chat: %w", err)
}
text, _, perr := parseResponseMood(resp)
if perr != nil {
log.Printf("phraser: PhraseChat: %v", perr)
return "поговорили.", nil
return ChatFallback, fmt.Errorf("phrase chat: %w", perr)
}
if text != "" {
return text, nil
+7 -1
View File
@@ -66,11 +66,17 @@ type Stub struct{}
// NewStub builds the floor phraser. no config — the Stub is stateless.
func NewStub() *Stub { return &Stub{} }
// ChatFallback — what she says on the chat path when the model gave her
// nothing to say. It replaced "поговорили.", which reads as a summary of a
// conversation that did not happen. Said out loud this one is an admission,
// which is what it is.
const ChatFallback = "даже не знаю, что сказать."
// PhraseChat returns a stub reply — the LLMPhraser replaces this with a
// prompted response from the model. The history parameter is accepted but
// ignored at the stub level (the production impl uses it for multi-turn).
func (s *Stub) PhraseChat(_ context.Context, _ string, _ []dialogue.Turn) (string, error) {
return "поговорили.", nil
return ChatFallback, nil
}
// PhraseQuery returns a deterministic summary of the best matching notes.
+112
View File
@@ -0,0 +1,112 @@
// phraser/replier.go — reactive reply phrasing, the confirmation he hears
// after every fact, note and reminder.
//
// It lived in cmd/mavend as package main until Vikunja #396, which meant the
// most frequently heard sentence Maven says was the one path the phrasing eval
// could not import, let alone score. Nothing here talks to the daemon: the
// caller supplies the completer and the context block, and cmd/mavend keeps the
// stub fallback so a model error still answers.
package phraser
import (
"context"
"strings"
"time"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/router"
)
// Completer is the model seam for the replier, a subset of router.Completer.
// *llm.Client satisfies it.
type Completer interface {
Complete(ctx context.Context, r llm.Req) (string, error)
}
// replyTimeout bounds one reply. Generous because the resident model on the CPU
// floor is slow and the caller has a deterministic fallback anyway.
const replyTimeout = 60 * time.Second
// ReplySystemPrompt — the reactive confirmation contract: one short Russian
// sentence, feminine self-reference, informal address, no question.
const ReplySystemPrompt = `Ты — Maven, домашняя ассистентка (о себе — в женском роде). Владелец — мужчина, говоришь с ним на "ты", в единственном числе; никогда не "вы"/"ваш" и не "он"/"его". Подтверди действие РОВНО ОДНИМ коротким предложением (≤120 символов), по-русски, спокойно и без официальных формулировок. Не задавай вопросов, не повторяй слова, не добавляй ничего после точки. Отвечай ТОЛЬКО одним объектом JSON с полями "response" (текст) и "mood" (ровно одно из: neutral, happy, thinking, tired, confused).
Пример: {"response": "Записала, что ты выпил стакан воды.", "mood": "neutral"}
Никогда не пиши "..." в поле response.`
// Replier phrases reactive confirmations with the resident model. It has no
// fallback of its own: an error is returned, and the daemon answers from the
// deterministic stub. That is also what makes it scorable — a dead server shows
// up as an error rather than as bad phrasing.
type Replier struct {
c Completer
// block renders the shared context block per turn (who he is, the time).
// nil ⇒ the prompt stands alone.
block func() string
}
// NewReplier builds a replier over c. block may be nil.
func NewReplier(c Completer, block func() string) *Replier {
return &Replier{c: c, block: block}
}
// PhraseReply returns the confirmation for one decision. An empty string with a
// nil error means the model produced nothing usable, which the caller must
// treat exactly like an error.
func (r *Replier) PhraseReply(ctx context.Context, d router.Decision) (string, error) {
ctx, cancel := context.WithTimeout(ctx, replyTimeout)
defer cancel()
out, err := r.c.Complete(ctx, llm.Req{
System: persona.Prepend(r.block, ReplySystemPrompt),
User: replyContext(d),
Grammar: ResponseGrammar,
MaxTokens: 512,
RepeatPenalty: 1.3,
})
if err != nil {
return "", err
}
out = stripThink(out)
if response, _, perr := parseResponseMood(out); perr != nil {
return "", perr
} else if response != "" {
return response, nil
}
// fallback: the model answered in bare prose, which is fine here.
return firstSentence(out), nil
}
// firstSentence trims the model's output to a single clean confirmation: first
// line, first sentence, whitespace-normalized — the last-line defense against a
// small model that rambles past the first period despite the prompt + stop.
func firstSentence(s string) string {
s = strings.TrimSpace(s)
if i := strings.IndexByte(s, '\n'); i >= 0 {
s = s[:i]
}
// keep up to and including the first sentence-ending punctuation.
if i := strings.IndexAny(s, ".!?"); i >= 0 {
s = s[:i+1]
}
return strings.TrimSpace(s)
}
// replyContext renders the decision into a compact RU description for the model.
func replyContext(d router.Decision) string {
switch d.Intent {
case router.IntentFact:
return "записала факт: " + d.Slots.Key + " " + d.Slots.Value
case router.IntentNote:
return "сохранила заметку: " + d.Slots.Text
case router.IntentReminder:
return "поставила напоминание: " + d.Slots.Text
default:
return string(d.Intent) + ": " + d.Slots.Text
}
}
// StripThink removes the <think> block a Thinking-variant model emits before its
// answer. Exported for the daemon's own model callers, which parse output that
// never passes through a phraser method.
func StripThink(s string) string { return stripThink(s) }
+90
View File
@@ -0,0 +1,90 @@
package phraser
import (
"context"
"testing"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/router"
)
type mockCompleter struct {
out string
err error
}
func (m mockCompleter) Complete(_ context.Context, _ llm.Req) (string, error) { return m.out, m.err }
func TestReplierReturnsLLMReply(t *testing.T) {
r := NewReplier(mockCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`}, nil)
got, err := r.PhraseReply(context.Background(), noteDecision())
if err != nil || got != "записала, кофе закончился" {
t.Errorf("got %q, %v, want %q, nil", got, err, "записала, кофе закончился")
}
}
func TestReplierFallsBackToPlainText(t *testing.T) {
r := NewReplier(mockCompleter{out: "записала, кофе закончился"}, nil)
got, err := r.PhraseReply(context.Background(), noteDecision())
if err != nil || got != "записала, кофе закончился" {
t.Errorf("got %q, %v, want %q, nil", got, err, "записала, кофе закончился")
}
}
func TestReplierReportsTheModelError(t *testing.T) {
r := NewReplier(mockCompleter{err: errTestLLMDown}, nil)
got, err := r.PhraseReply(context.Background(), noteDecision())
if err == nil {
t.Errorf("got %q, nil error — a dead model must be reported, not phrased around", got)
}
}
// A fragment the grammar left half-open is a failed generation. It must come
// back as an error so the daemon reaches its stub, not as a reply.
func TestReplierRejectsBrokenJSON(t *testing.T) {
r := NewReplier(mockCompleter{out: `{"response":"запис`}, nil)
got, err := r.PhraseReply(context.Background(), noteDecision())
if err == nil || got != "" {
t.Errorf("got %q, %v, want empty and an error", got, err)
}
}
func TestReplierEmptyOutputIsEmpty(t *testing.T) {
r := NewReplier(mockCompleter{out: ""}, nil)
got, err := r.PhraseReply(context.Background(), noteDecision())
if err != nil || got != "" {
t.Errorf("got %q, %v, want empty and no error", got, err)
}
}
// grammarRecorder captures the request so the grammar can be asserted on.
type grammarRecorder struct{ req llm.Req }
func (g *grammarRecorder) Complete(_ context.Context, r llm.Req) (string, error) {
g.req = r
return `{"response":"записала","mood":"neutral"}`, nil
}
func TestReplierCarriesTheResponseGrammar(t *testing.T) {
rec := &grammarRecorder{}
r := NewReplier(rec, nil)
if _, err := r.PhraseReply(context.Background(), noteDecision()); err != nil {
t.Fatalf("PhraseReply: %v", err)
}
if rec.req.Grammar != ResponseGrammar {
t.Errorf("grammar = %q, want ResponseGrammar", rec.req.Grammar)
}
if rec.req.System != ReplySystemPrompt {
t.Errorf("system prompt = %q, want ReplySystemPrompt", rec.req.System)
}
}
func noteDecision() router.Decision {
return router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}}
}
var errTestLLMDown = errTest("llm down")
type errTest string
func (e errTest) Error() string { return string(e) }
+64 -10
View File
@@ -1,9 +1,11 @@
package phraser
import (
"bytes"
"context"
"errors"
"fmt"
"log"
"os"
"os/exec"
"path/filepath"
@@ -59,6 +61,19 @@ func TestExtractPort(t *testing.T) {
}
}
// The prompt cache is what ate 6.8GB of the deployed server's RSS, so the cap
// has to reach the command line, and the opt-out has to leave it off.
func TestLlamaArgsCapsPromptCache(t *testing.T) {
cfg := DefaultConfig("/m.gguf")
if got := strings.Join(llamaArgs(cfg), " "); !strings.Contains(got, "--cache-ram 512") {
t.Errorf("default args = %q, want --cache-ram 512", got)
}
cfg.CacheRAMMiB = 0
if got := strings.Join(llamaArgs(cfg), " "); strings.Contains(got, "--cache-ram") {
t.Errorf("args with the cap off = %q, want no --cache-ram flag", got)
}
}
func TestStartLlamaProcScrapesPortAndReaps(t *testing.T) {
bin := fakeLlama(t, listensThenSleeps)
ctx, cancel := context.WithCancel(context.Background())
@@ -86,6 +101,48 @@ func TestStartLlamaProcScrapesPortAndReaps(t *testing.T) {
}
}
// captureLog redirects the standard logger for the duration of a test and
// returns what was written to it.
func captureLog(t *testing.T) *bytes.Buffer {
t.Helper()
var buf bytes.Buffer
old := log.Writer()
flags := log.Flags()
log.SetOutput(&buf)
log.SetFlags(0)
t.Cleanup(func() { log.SetOutput(old); log.SetFlags(flags) })
return &buf
}
// The child's buffer-size, KV-cache and offload lines are the only way to
// account for its memory on a running box, and they used to be dropped: stderr
// was scraped for the listen line and thrown away, stdout was never piped.
func TestStartLlamaProcForwardsChildOutput(t *testing.T) {
buf := captureLog(t)
bin := fakeLlama(t, `echo "load_tensors: Vulkan0 model buffer size = 1053.34 MiB" >&2
echo "llama_context: KV self size = 448.00 MiB"
`+listensThenSleeps)
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
p, err := startLlamaProc(ctx, testCfg(bin))
if err != nil {
t.Fatalf("startLlamaProc: %v", err)
}
p.cancel = cancel
defer p.Close()
got := buf.String()
for _, want := range []string{
"llama: load_tensors: Vulkan0 model buffer size = 1053.34 MiB", // stderr
"llama: llama_context: KV self size = 448.00 MiB", // stdout, previously discarded
} {
if !strings.Contains(got, want) {
t.Errorf("log missing %q\nlog was:\n%s", want, got)
}
}
}
func TestStartLlamaProcFailureArms(t *testing.T) {
t.Run("binary missing", func(t *testing.T) {
cfg := testCfg(filepath.Join(t.TempDir(), "does-not-exist"))
@@ -96,13 +153,18 @@ func TestStartLlamaProcFailureArms(t *testing.T) {
})
t.Run("server exits without listening", func(t *testing.T) {
// stderr closes, so the reader goroutine reports EOF on errCh.
// stderr closes, so the reader goroutine reports EOF on errCh. The error
// must carry the child's last words: bare "EOF" named no cause.
captureLog(t)
bin := fakeLlama(t, `echo "ggml_vulkan: no device" >&2
exit 1`)
_, err := startLlamaProc(context.Background(), testCfg(bin))
if err == nil || !strings.Contains(err.Error(), "llm: server output") {
t.Fatalf("err = %v, want the server-output arm", err)
}
if !strings.Contains(err.Error(), "ggml_vulkan: no device") {
t.Errorf("err = %v, want the child's last output in it", err)
}
})
t.Run("context cancelled during startup", func(t *testing.T) {
@@ -241,15 +303,7 @@ func TestKillMavenScriptMatchesRealCommandLine(t *testing.T) {
// startLlamaProc that breaks the sweep fails here instead of on the box.
cfg := DefaultConfig("/opt/maven/models/llm/Qwen3-1.7B-UD-Q4_K_XL.gguf")
cfg.NCtx, cfg.NGpuLayers = 4096, 99
cmdline := strings.Join([]string{
cfg.BinPath,
"-m", cfg.ModelPath,
"--host", "127.0.0.1",
"--port", extractPort(cfg.Listen),
"-c", fmt.Sprintf("%d", cfg.NCtx),
"-ngl", fmt.Sprintf("%d", cfg.NGpuLayers),
"--no-webui",
}, " ")
cmdline := cfg.BinPath + " " + strings.Join(llamaArgs(cfg), " ")
if !pat.MatchString(cmdline) {
t.Fatalf("kill-maven.sh pattern %q does not match %q — orphans would leak", m[1], cmdline)
}
+5 -3
View File
@@ -215,10 +215,12 @@ func TestSwap_RollbackFailureLeavesNoBackendAndDegrades(t *testing.T) {
if _, _, aerr := p.acquire(); !errors.Is(aerr, ErrNoBackend) {
t.Errorf("acquire error = %v; want ErrNoBackend", aerr)
}
// Phrasing degrades to its fallback instead of failing the turn.
// Phrasing degrades to its fallback instead of failing the turn, and since
// Vikunja #397 it reports the error next to that fallback so a measuring
// caller can tell "no model" from "bad phrasing".
got, err := p.PhraseChat(context.Background(), "привет", nil)
if err != nil {
t.Fatalf("PhraseChat after a total failure returned an error: %v", err)
if !errors.Is(err, ErrNoBackend) {
t.Errorf("PhraseChat error = %v; want ErrNoBackend alongside the fallback", err)
}
if got == "" {
t.Error("PhraseChat returned empty; the fallback must still say something")
+64
View File
@@ -0,0 +1,64 @@
package router
import "strings"
// interrogatives — the question words that mark an utterance as asking rather
// than telling. Tokenized, never substring: "что" inside "чтобы" and "как"
// inside "какао" are not questions.
var interrogatives = []string{
"что", "чего", "какой", "какая", "какое", "какие", "каких",
"кто", "кого", "кому", "чей", "почему", "зачем", "отчего",
"где", "куда", "откуда", "когда", "сколько", "как",
"what", "who", "whom", "why", "when", "where", "which", "how",
}
// narrativeRequests — "tell me about X" asks for knowledge Maven does not
// hold about him. It carries no question mark and no interrogative, which is
// how "расскажи про битву при Ватерлоо" reached the fact store (#470).
var narrativeRequests = []string{
"расскажи", "объясни", "опиши", "перечисли",
"tell", "explain", "describe",
}
// captureVerbs — an explicit instruction to record something. These win over
// every test below, because "запиши что я пил воду" contains an interrogative
// and is still a capture: the word he said is "запиши".
var captureVerbs = []string{
"запиши", "запомни", "отметь", "заметь", "добавь", "сохрани",
"note", "remember", "log", "save",
}
// IsQuestionShaped reports whether text asks for something rather than
// records it. It is a deterministic offline test over tokens, so it costs
// nothing and never depends on the model that produced the routing decision.
//
// It exists because a mis-routed question used to be persisted as a fact
// about the owner, with the model's invented answer as the value (#470). The
// predicate is deliberately blunt: refusing to store a question is cheap and
// reversible, storing an invented fact about him is neither.
func IsQuestionShaped(text string) bool {
t := strings.TrimSpace(text)
if t == "" {
return false
}
toks := planTokens(t)
for _, v := range captureVerbs {
if hasTok(toks, v) {
return false
}
}
if strings.HasSuffix(t, "?") {
return true
}
for _, w := range interrogatives {
if hasTok(toks, w) {
return true
}
}
for _, w := range narrativeRequests {
if hasTok(toks, w) {
return true
}
}
return false
}
+48
View File
@@ -0,0 +1,48 @@
package router
import "testing"
func TestIsQuestionShaped(t *testing.T) {
// The seven utterances #470 recorded, plus the captures that must keep
// working. A capture misread as a question loses a fact; a question
// misread as a capture poisons recall, so the captures are the ones worth
// pinning here.
cases := []struct {
text string
want bool
}{
{"какая последняя версия языка Go?", true},
{"что дальше?", true},
{"расскажи про битву при Ватерлоо", true},
{"почему небо синее?", true},
{"какая столица Австралии?", true},
{"кто такой Никола Тесла?", true},
{"сколько стоит доллар", true},
{"who is the premier of Japan", true},
{"объясни линии Фраунгофера", true},
{"запиши что я пил воду", false},
{"запомни какая у меня машина", false},
{"отметь что я поужинал", false},
{"поужинал", false},
{"я выпил кофе", false},
{"вода", false},
{"привет", false},
{"", false},
}
for _, c := range cases {
if got := IsQuestionShaped(c.text); got != c.want {
t.Errorf("IsQuestionShaped(%q) = %v, want %v", c.text, got, c.want)
}
}
}
// Substring matching is what made the day-plan predicates wrong before, and
// this predicate gates a write, so it gets the same guard.
func TestIsQuestionShapedIsTokenized(t *testing.T) {
for _, text := range []string{"чтобы не забыть, я полил кактус", "какао выпил"} {
if IsQuestionShaped(text) {
t.Errorf("IsQuestionShaped(%q) = true; a question word inside a longer word is not a question", text)
}
}
}
+62
View File
@@ -6,6 +6,7 @@ import (
"encoding/json"
"errors"
"fmt"
"log"
"strings"
"time"
@@ -46,6 +47,39 @@ func (s *Store) WriteFact(ctx context.Context, ts time.Time, kind FactKind, key,
return id, nil
}
// FactRecallText is the text a fact is indexed under and read back as (#493).
//
// It used to be the utterance that wrote the fact, so recall of ANY
// voice-tapped fact answered with the sentence he said instead of the value
// stored: `go_version = 1.20` was indexed as "какая последняя версия языка
// Go?", and that question is what came back. The poisoned rows made the defect
// visible; the shape was wrong for legitimate facts too.
//
// The key is spoken with its underscores dropped, because a key is written for
// the store and this string is read out loud.
func FactRecallText(key, value string) string {
spoken := strings.TrimSpace(strings.ReplaceAll(key, "_", " "))
v := strings.TrimSpace(DecodeFactValue(value))
switch {
case v == "":
return spoken
case spoken == "":
return v
}
return spoken + " — " + v
}
// DecodeFactValue unwraps a stored value for reading. The column holds raw json
// when the writer serialized one (SetValue, CorrectValue) and a plain string
// when it did not (a voice tap), so a reader that wants the text handles both.
func DecodeFactValue(value string) string {
var s string
if err := json.Unmarshal([]byte(value), &s); err == nil {
return s
}
return value
}
// LatestFact returns the latest non-voided fact for key, or ErrNoFact.
// "Non-voided" = no later row has voids_id pointing at it. We resolve this by
// taking the newest row whose id is not referenced by any voids_id.
@@ -290,6 +324,20 @@ func (s *Store) CorrectValue(ctx context.Context, key, source string, value any,
if err != nil {
return 0, fmt.Errorf("last insert id: %w", err)
}
// The same repair a void needs, for the same reason (#493). A correction
// supersedes the value, and the vector still holds the old one, so recall
// kept answering with the value he had just corrected. Dropping it costs
// the key its recall vector until the fact is tapped again: this layer has
// no embedder, and a missing vector loses a question while a stale one
// answers it wrongly.
//
// Best-effort: the corrected row is committed, and a correction that lands
// beats one that fails on cleanup.
if n, derr := s.VectorMemory().DeletePrefix(ctx, "fact:"+key+":"); derr != nil {
log.Printf("store: correct %q: memory vectors survive: %v", key, derr)
} else if n > 0 {
log.Printf("store: correct %q: dropped %d superseded memory vector(s)", key, n)
}
return newID, nil
}
@@ -332,6 +380,20 @@ func (s *Store) VoidLatestFact(ctx context.Context, key, source string, ts time.
if err != nil {
return 0, 0, fmt.Errorf("void: last insert id: %w", err)
}
// The other half of the repair (#470). A fact reaches recall through a
// vector keyed `fact:<key>:<unix>`, holding the utterance that wrote it.
// Voiding the row alone left that vector answering questions, so revert
// reported success on a box that stayed broken. Deleting every vector for
// the key covers the earlier rows too: their values are superseded, and a
// superseded value has no business claiming a turn.
//
// Best-effort by design: the audit trail is already committed, and a fact
// that is voided but still recallable is better than a void that failed.
if n, derr := s.VectorMemory().DeletePrefix(ctx, "fact:"+key+":"); derr != nil {
log.Printf("store: void %q: memory vectors survive: %v", key, derr)
} else if n > 0 {
log.Printf("store: void %q: dropped %d memory vector(s)", key, n)
}
return oldID, newID, nil
}
+177
View File
@@ -0,0 +1,177 @@
package store
import (
"context"
"encoding/json"
"errors"
"fmt"
"strconv"
"strings"
"time"
)
// metaKeyFactVectorShape names the shape the stored fact vectors were written
// in. It exists so the repair below runs once per box instead of on every
// start: the rows it fixes were written by a code path that no longer exists,
// and once fixed nothing writes that shape again.
const metaKeyFactVectorShape = "fact_vector_shape"
// factVectorShapeFact is the shape FactRecallText produces. Anything else in
// the marker (including nothing, which is every box written before #493) means
// the fact vectors still hold utterances.
const factVectorShapeFact = "fact-text (#493)"
// FactVectorRepair is what one repair run did, for logging.
type FactVectorRepair struct {
Skipped bool // marker already matched — nothing to do
Rewritten int // rows re-embedded from the fact they name
Dropped int // rows deleted: voided, superseded, or naming no fact at all
Kept int // rows already holding the right text
Took time.Duration
}
// RepairFactVectors brings the fact rows of memory_vectors in line with the
// facts they name, and is the operator recovery a poisoned box had no path to
// (#470 point 4, #493).
//
// Three defects put wrong text in that index, and all three are write-path
// fixes that do nothing for rows already stored:
//
// - the indexed text was the utterance, so every fact row reads back a
// sentence rather than a value;
// - a void left its vector behind, so retracted junk kept answering;
// - a correction left its vector behind, so the superseded value did.
//
// So each fact row is resolved against the fact store and one of three things
// happens. It is dropped when the key has no fact, when the newest row for the
// key is a void marker, or when a newer vector for the same key exists — a
// superseded value has no business claiming a turn. It is re-embedded when its
// text is not what FactRecallText says the fact is. Otherwise it is left alone.
//
// Idempotent, and safe to interrupt: every step compares before writing and the
// marker is written last, so a run that dies partway is simply redone.
func (s *Store) RepairFactVectors(ctx context.Context, embed EmbedFunc) (FactVectorRepair, error) {
start := time.Now()
var res FactVectorRepair
shape, err := s.Meta(ctx, metaKeyFactVectorShape)
if err != nil {
return res, err
}
if shape == factVectorShapeFact {
res.Skipped = true
res.Took = time.Since(start)
return res, nil
}
rows, err := s.db.QueryContext(ctx, `SELECT id, meta FROM memory_vectors`)
if err != nil {
return res, fmt.Errorf("repair fact vectors: read: %w", err)
}
type factVec struct {
id, key string
meta map[string]string
ts int64
}
var vecs []factVec
newest := map[string]int64{} // key → newest ts seen for it
for rows.Next() {
var id, metaJSON string
if err := rows.Scan(&id, &metaJSON); err != nil {
rows.Close()
return res, fmt.Errorf("repair fact vectors: row: %w", err)
}
meta := map[string]string{}
if err := json.Unmarshal([]byte(metaJSON), &meta); err != nil {
rows.Close()
return res, fmt.Errorf("repair fact vectors: meta for %q: %w", id, err)
}
if meta["type"] != "fact" {
continue
}
key, ts, ok := splitFactVectorID(id)
if !ok {
continue
}
vecs = append(vecs, factVec{id: id, key: key, meta: meta, ts: ts})
if ts > newest[key] {
newest[key] = ts
}
}
rows.Close()
if err := rows.Err(); err != nil {
return res, fmt.Errorf("repair fact vectors: rows: %w", err)
}
for _, v := range vecs {
drop := v.ts < newest[v.key]
var want string
if !drop {
f, ferr := s.LatestFact(ctx, v.key)
switch {
case errors.Is(ferr, ErrNoFact):
drop = true
case ferr != nil:
return res, fmt.Errorf("repair fact vectors: fact %q: %w", v.key, ferr)
case DecodeFactValue(f.Value) == "voided":
drop = true
default:
want = FactRecallText(v.key, f.Value)
}
}
if drop {
if err := s.VectorMemory().Delete(ctx, v.id); err != nil {
return res, err
}
res.Dropped++
continue
}
if v.meta["text"] == want {
res.Kept++
continue
}
vec, err := embed(ctx, want)
if err != nil {
return res, fmt.Errorf("repair fact vectors: embed %q: %w", v.id, err)
}
// The whole meta blob is rewritten in Go rather than patched in SQL,
// because json_set needs the JSON1 extension and this store is opened
// through sqlcipher.
v.meta["text"] = want
metaJSON, err := json.Marshal(v.meta)
if err != nil {
return res, fmt.Errorf("repair fact vectors: meta %q: %w", v.id, err)
}
if _, err := s.db.ExecContext(ctx,
`UPDATE memory_vectors SET vec = ?, meta = ? WHERE id = ?`,
encodeVec(vec), string(metaJSON), v.id); err != nil {
return res, fmt.Errorf("repair fact vectors: write %q: %w", v.id, err)
}
res.Rewritten++
}
if err := s.SetMeta(ctx, metaKeyFactVectorShape, factVectorShapeFact); err != nil {
return res, err
}
res.Took = time.Since(start)
return res, nil
}
// splitFactVectorID reads the key and write time back out of a fact vector's
// id, which the write path builds as `fact:<key>:<unix>`. A key may hold a
// colon, the timestamp may not, so the split is from the right.
func splitFactVectorID(id string) (key string, ts int64, ok bool) {
rest, found := strings.CutPrefix(id, "fact:")
if !found {
return "", 0, false
}
cut := strings.LastIndex(rest, ":")
if cut <= 0 {
return "", 0, false
}
ts, err := strconv.ParseInt(rest[cut+1:], 10, 64)
if err != nil {
return "", 0, false
}
return rest[:cut], ts, true
}
+151
View File
@@ -0,0 +1,151 @@
package store
import (
"context"
"database/sql"
"testing"
"time"
)
// The write-path half of #493: recall of a fact must read back the fact, not
// the sentence he happened to say.
func TestFactRecallText(t *testing.T) {
for _, tc := range []struct {
name, key, value, want string
}{
{"json value", "go_version", `"1.20"`, "go version — 1.20"},
{"plain value", "water", "выпил", "water — выпил"},
{"no value", "shower", "", "shower"},
{"underscores are spoken as spaces", "espresso_machine", `"чистая"`, "espresso machine — чистая"},
} {
t.Run(tc.name, func(t *testing.T) {
if got := FactRecallText(tc.key, tc.value); got != tc.want {
t.Fatalf("FactRecallText(%q, %q) = %q; want %q", tc.key, tc.value, got, tc.want)
}
})
}
}
// A correction left the superseded value in the index, so recall answered with
// the value he had just corrected (#493).
func TestCorrectValueDropsMemoryVectors(t *testing.T) {
ctx := context.Background()
s := newTestStore(t)
now := time.Now()
mem := s.VectorMemory()
if _, err := s.WriteFact(ctx, now, KindSelf, "go_version", `"1.20"`, "tap:voice", 1.0, sql.NullInt64{}); err != nil {
t.Fatalf("WriteFact: %v", err)
}
if err := mem.Insert(ctx, "fact:go_version:1", []float32{1, 0, 0}, map[string]string{
"type": "fact", "text": "go version — 1.20",
}); err != nil {
t.Fatalf("Insert: %v", err)
}
if _, err := s.CorrectValue(ctx, "go_version", "feedback", "1.25", now.Add(time.Minute)); err != nil {
t.Fatalf("CorrectValue: %v", err)
}
got, err := mem.ByPrefix(ctx, "fact:")
if err != nil {
t.Fatalf("ByPrefix: %v", err)
}
if len(got) != 0 {
t.Fatalf("after the correction the index still holds %+v; the superseded value must not answer", got)
}
}
// The recovery path a poisoned box had none of (#470 point 4, #493): rows
// written before the fix hold utterances, voided junk and superseded values,
// and no write-path change reaches any of them.
func TestRepairFactVectors(t *testing.T) {
ctx := context.Background()
s := newTestStore(t)
now := time.Now()
mem := s.VectorMemory()
embed := func(ctx context.Context, text string) ([]float32, error) {
return []float32{float32(len(text)), 1, 0}, nil
}
// A live fact indexed under the question that wrote it — the defect.
if _, err := s.WriteFact(ctx, now, KindSelf, "water", `"выпил"`, "tap:voice", 1.0, sql.NullInt64{}); err != nil {
t.Fatalf("WriteFact water: %v", err)
}
if err := mem.Insert(ctx, "fact:water:100", []float32{9, 9, 9}, map[string]string{
"type": "fact", "source": "voice", "text": "запиши что я пил воду",
}); err != nil {
t.Fatalf("Insert water: %v", err)
}
// A voided fact whose vector survived the void.
if _, err := s.WriteFact(ctx, now, KindSelf, "go_version", `"1.20"`, "tap:voice", 1.0, sql.NullInt64{}); err != nil {
t.Fatalf("WriteFact go_version: %v", err)
}
if _, _, err := s.VoidLatestFact(ctx, "go_version", "feedback", now.Add(time.Minute)); err != nil {
t.Fatalf("VoidLatestFact: %v", err)
}
if err := mem.Insert(ctx, "fact:go_version:100", []float32{9, 9, 9}, map[string]string{
"type": "fact", "text": "какая последняя версия языка Go?",
}); err != nil {
t.Fatalf("Insert go_version: %v", err)
}
// A key with two vectors: only the newest may answer.
if _, err := s.WriteFact(ctx, now, KindSelf, "mood", `"устал"`, "tap:voice", 1.0, sql.NullInt64{}); err != nil {
t.Fatalf("WriteFact mood: %v", err)
}
for _, ts := range []string{"100", "200"} {
if err := mem.Insert(ctx, "fact:mood:"+ts, []float32{9, 9, 9}, map[string]string{
"type": "fact", "text": "мне грустно",
}); err != nil {
t.Fatalf("Insert mood %s: %v", ts, err)
}
}
// A note must be left entirely alone.
if err := mem.Insert(ctx, "note:7", []float32{5, 5, 5}, map[string]string{
"type": "note", "text": "сеть тормозит по вечерам",
}); err != nil {
t.Fatalf("Insert note: %v", err)
}
res, err := s.RepairFactVectors(ctx, embed)
if err != nil {
t.Fatalf("RepairFactVectors: %v", err)
}
if res.Rewritten != 2 || res.Dropped != 2 {
t.Fatalf("repair reported %+v; want 2 rewritten (water, newest mood) and 2 dropped (voided go_version, superseded mood)", res)
}
got, err := mem.ByPrefix(ctx, "fact:")
if err != nil {
t.Fatalf("ByPrefix: %v", err)
}
texts := map[string]string{}
for _, r := range got {
texts[r.ID] = r.Meta["text"]
}
if len(texts) != 2 {
t.Fatalf("the index holds %+v; want only fact:water:100 and fact:mood:200", texts)
}
if texts["fact:water:100"] != "water — выпил" {
t.Fatalf("water reads back %q; want the fact, not the utterance", texts["fact:water:100"])
}
if texts["fact:mood:200"] != "mood — устал" {
t.Fatalf("mood reads back %q", texts["fact:mood:200"])
}
// Provenance the row already carried must survive the rewrite.
for _, r := range got {
if r.ID == "fact:water:100" && r.Meta["source"] != "voice" {
t.Fatalf("water lost its source meta: %+v", r.Meta)
}
}
if notes, err := mem.ByPrefix(ctx, "note:"); err != nil || len(notes) != 1 {
t.Fatalf("the note row was touched: %+v (err %v)", notes, err)
}
// Marker written, so a second run is free and changes nothing.
again, err := s.RepairFactVectors(ctx, embed)
if err != nil {
t.Fatalf("second RepairFactVectors: %v", err)
}
if !again.Skipped {
t.Fatalf("second run did work: %+v; the marker must make it a no-op", again)
}
}
+21
View File
@@ -148,6 +148,27 @@ func (m *MemoryStore) Delete(ctx context.Context, id string) error {
return nil
}
// DeletePrefix removes every vector whose id starts with prefix and returns
// how many went. Same escaping as ByPrefix, so a key containing % or _ cannot
// widen the delete.
//
// It exists for the repair half of a revert (#470). Voiding a fact row left
// its vector in the index, so recall kept serving the voided fact's utterance
// and the documented repair did not repair.
func (m *MemoryStore) DeletePrefix(ctx context.Context, prefix string) (int64, error) {
pattern := escapeLike(prefix) + "%"
res, err := m.db.ExecContext(ctx,
`DELETE FROM memory_vectors WHERE id LIKE ? ESCAPE '\'`, pattern)
if err != nil {
return 0, fmt.Errorf("memory: delete prefix %q: %w", prefix, err)
}
n, err := res.RowsAffected()
if err != nil {
return 0, fmt.Errorf("memory: delete prefix %q: rows affected: %w", prefix, err)
}
return n, nil
}
// escapeLike neutralises the LIKE wildcards in a literal prefix.
func escapeLike(s string) string {
r := strings.NewReplacer(`\`, `\\`, `%`, `\%`, `_`, `\_`)
+64
View File
@@ -0,0 +1,64 @@
package store
import (
"context"
"database/sql"
"testing"
"time"
)
// Stage 3 of #470: reverting a fact reported success and left the vector that
// was answering questions, so the documented repair did not repair.
func TestVoidLatestFactDropsMemoryVectors(t *testing.T) {
ctx := context.Background()
s := newTestStore(t)
now := time.Now()
mem := s.VectorMemory()
if _, err := s.WriteFact(ctx, now, KindSelf, "go_version", `"1.20"`, "tap:voice", 1.0, sql.NullInt64{}); err != nil {
t.Fatalf("WriteFact: %v", err)
}
// The id shape actionFact writes: fact:<key>:<unix>.
if err := mem.Insert(ctx, "fact:go_version:1", []float32{1, 0, 0}, map[string]string{
"type": "fact", "text": "какая последняя версия языка Go?",
}); err != nil {
t.Fatalf("Insert: %v", err)
}
// A vector for another key must survive the void.
if err := mem.Insert(ctx, "fact:water:1", []float32{0, 1, 0}, map[string]string{
"type": "fact", "text": "запиши что я пил воду",
}); err != nil {
t.Fatalf("Insert: %v", err)
}
if _, _, err := s.VoidLatestFact(ctx, "go_version", "feedback", now.Add(time.Minute)); err != nil {
t.Fatalf("VoidLatestFact: %v", err)
}
got, err := mem.ByPrefix(ctx, "fact:")
if err != nil {
t.Fatalf("ByPrefix: %v", err)
}
if len(got) != 1 || got[0].ID != "fact:water:1" {
t.Fatalf("after the void the index holds %+v; want only fact:water:1", got)
}
}
func TestDeletePrefixDoesNotWidenOnWildcards(t *testing.T) {
ctx := context.Background()
s := newTestStore(t)
mem := s.VectorMemory()
for _, id := range []string{"fact:a_b:1", "fact:axb:1"} {
if err := mem.Insert(ctx, id, []float32{1, 0}, map[string]string{"type": "fact"}); err != nil {
t.Fatalf("Insert %q: %v", id, err)
}
}
n, err := mem.DeletePrefix(ctx, "fact:a_b:")
if err != nil {
t.Fatalf("DeletePrefix: %v", err)
}
if n != 1 {
t.Fatalf("deleted %d rows; the _ in the key must not match x", n)
}
}