Add background memory evaluation, off unless configured #55
Closed
claude
wants to merge 1 commits from
overnight/memory-eval into overnight/proactive-proposals
pull from: overnight/memory-eval
merge into: kami:overnight/proactive-proposals
kami:master
kami:task/725-capability-ledger-and-empirical-baseline
kami:task/692-heads-path-may-equal-model-path-and-noth
kami:task/694-staticcheck-and-deadcode-are-still-not-i
kami:task/682-go-1-25-5-and-x-text-0-14-0-carry-20-rea
kami:task/674-caveats
kami:task/673-mavgpud-serves-the-model-to-the-whole-la
kami:task/487-capture-device-doc
kami:task/487-capture-device
kami:task/487-wake-word-deploy
kami:task/487-wake-word-threshold
kami:task/487-wake-word-stage-two
kami:task/671-mavwaked-registers-as-a-voice-consumer-i
kami:task/670-cut-claude-md-to-200-lines
kami:task/515-deploy-mavwaked-workpc
kami:task/669-prune-claude-md
kami:task/668-e4b-phrasing
kami:task/668-title-capital
kami:task/668-kiwix-answers-a-question-it-cannot-answe
kami:task/666-only-a-stage-0-grammar-may-take-the-pers
kami:task/487-mavwaked-has-no-wake-word-only-an-energy
kami:task/486-deploy-the-workstation-transcriber
kami:task/486-move-stt-and-tts-to-the-workstation-wher
kami:task/665-crisperwhisper-2-russian
kami:task/664-routing-heads-in-go
kami:task/662-usage-harness-source-badge
kami:task/661-post-merge-usage-rerun
kami:task/661-routing-heads-step-3-train-the-multi-hea
kami:task/660-router-prompt-destination
kami:task/659-destination-fixture
kami:task/655-query-source-is-a-routing-decision-made
kami:task/654-a-pending-clarify-has-no-way-out-neither
kami:task/654-week-of-usage-eval-docs
kami:task/649-needs-kami-telegram-is-the-only-reach-an
kami:task/643-memorystore-search-decodes-and-unmarshal
kami:task/641-two-maps-grow-for-the-process-lifetime-w
kami:task/644-mavcaldav-is-built-documented-as-running
kami:task/642-the-store-caps-sqlite-at-one-connection
kami:task/647-factenrichmentworker-walks-the-pending-q
kami:task/646-v-637-follow-up-telegram-intake-has-no-d
kami:task/638-no-deadline-survives-the-turn-path-from
kami:task/637-inbound-telegram-turns-and-corrections-f
kami:task/636-correcting-a-turn-from-telegram-and-from
kami:task/634-an-act-alias-resolves-the-verb-but-not-t
kami:task/630-one-gesture-correction-on-chat-v-628
kami:task/629-persist-the-routing-trace-and-record-it
kami:task/631-mode-inventory-written-from-the-handlers
kami:task/586-defaultfactparser-uses-hand-written-russ
kami:task/633-reconcile-the-seed-labels-with-the-handl
kami:task/627-reminder-verbs-has-no-alarm-verb-so-an-a
kami:task/626-the-classifier-seeds-teach-an-older-inte
kami:task/546-route-with-a-fine-tuned-e5-small-instead
kami:task/586-measure-the-fact-parser
kami:fix/gofmt-ecosystem-acts
kami:task/584-media-store-a-failed-write-leaks-its-bud
kami:task/518-no-write-path-for-a-backdated-event-so-t
kami:task/287-qa-voice-session-quality-polish
kami:task/492-qa-plan-reconcile
kami:task/530-sweep-tail-four-files-the-russian-sweep
kami:task/405-score-how-often-a-real-utterance-reaches
kami:task/529-money-and-list-pick-a-mechanism
kami:task/528-sweep-tail-the-three-files-on-467
kami:task/527-embedder-open-set-phrasings-stop-being-r
kami:task/526-morphology-a-dictionary-answers-the-gram
kami:task/525-lexicons-the-finite-russian-sets-move-to
kami:task/524-entity-reference-ask-nexus-about-every-l
kami:task/523-risk-tiers-take-hexis-s-tier-for-a-hexis
kami:task/521-review-pr-111-query-strings-declension-h
kami:task/491-llama-server-core-dumps-on-every-sigterm
kami:task/479-bug-an-unconfigured-capability-does-not
kami:task/467-bug-spoken-task-capture-is-dead-the-rout
kami:task/463-deploy-mavwaked-and-mavenclient-run-nowh
kami:task/480-hearing-no-shipped-client-can-start-a-re
kami:task/432-ambient-calendar-intake-is-fragile-and-p
kami:task/431-board-surface-maven-holds-the-work-board
kami:task/433-reactivehandler-has-30-fields-and-is-pas
kami:task/371-swap-the-embedder-for-an-asymmetric-retr
kami:task/408-review-31-07-split-the-30-method-coreapi
kami:task/410-review-31-07-hand-rolled-string-enums-st
kami:task/423-review-pr50-split-internal-ipc-server-go
kami:task/422-review-pr50-split-cmd-mavend-tick-go-860
kami:task/409-review-31-07-finish-moving-mavweb-markup
kami:task/482-ambient-ingest-reads-a-notification-s-ti
kami:task/444-kuma-a-fact-per-monitor-so-she-can-name
kami:task/452-capability-model-homelab-docker-restart
kami:task/449-destructive-confirm-policy-risk-tiers-no
kami:task/453-grocery-list-items-table-fourth-append-o
kami:task/399-run-the-persona-checks-inside-the-daemon
kami:task/448-bounded-follow-up-state-pending-candidat
kami:task/455-conversation-repair-name-the-misroute-co
kami:task/454-go-mod-tidy
kami:task/458-pronunciation-dictionary-for-piper
kami:task/456-command-history-read-only-query-over-exi
kami:task/457-clarification-templates-for-the-router-s
kami:task/474-query-source-ordering-feeds-and-calendar
kami:task/469-reminders-spelled-out-times-fail-the-bod
kami:task/475-bug-the-praxis-attention-capability-is-u
kami:task/481-bug-a-transient-complaint-is-stored-as-a
kami:task/476-bug-the-router-transliterates-latin-enti
kami:task/385-decide-whether-a-parked-clarify-question
kami:task/377-backfill-routines
kami:task/421-weather-geocoder
kami:task/390-no-read-path-for-delivery-attempts
kami:task/386-recall-fixture-filler-note-ids
kami:task/473-bug-morning-item-has-no-required-flag
kami:task/465-bug-make-simulate-routes-with-an-empty
kami:task/467-bug-spoken-task-capture-is-dead
kami:task/466-bug-a-pending-clarify-is-global-so-one-u
kami:task/468-bug-pattern-detect-has-no-minimum-interv
kami:task/462-bug-checkfeminine-flags-second-person-ma
kami:task/443-safekey-drops-cyrillic-so-russian-calend
kami:task/471-bug-agendaquerygrammars-covers-today-but
kami:task/383-slottext-in-clarify-answer-would-clobber
kami:task/323-qa-phraser-coverage-is-65-3-but-the-llam
kami:task/498-bug-and-x-reach-the-model-with-no-determ
kami:task/506-strings-family-6-summaries-and-reports-i
kami:task/504-strings-family-4-act-and-smart-home-repl
kami:task/503-strings-family-3-query-answers-and-gaps
kami:task/502-strings-family-2-capture-acknowledgement
kami:task/501-strings-family-1-phrasing-fallbacks-into
kami:task/397-phrasechat-and-phrasequery-hide-model-fa
kami:task/396-the-reply-path-can-t-be-tested-llmreplie
kami:task/496-recall-a-cross-language-question-loses-i
kami:task/495-bug-x-escapes-the-personal-boundary-and
kami:task/499-llama-server-holds-7-9gb-rss-for-a-1-1gb
kami:task/470-bug-a-question-writes-invented-knowledge
kami:task/493-bug-the-memory-index-stores-the-raw-utte
kami:task/490-name-the-gap-world-questions-through-the
kami:task/485-run-the-big-model-on-the-workstation-wit
kami:task/489-workstation-deploy-mavgpud-on-workpc-and
kami:task/488-workstation-a-supervisor-that-keeps-llam
kami:task/483-docs-offload-design
kami:task/483-design-offload-ml-to-the-workstation-kee
kami:task/459-docs-refresh-the-qa-plan-against-the-liv
kami:task/446-doc-reorg-tier-the-tree-retire-the-three
kami:fix/367-voice-parks-routine-accept
kami:task/365-dialogue-slots-and-router-slots-are-hand
kami:task/364-snooze-does-nothing-at-runtime-the-gate
kami:task/447-retire-progress-md-the-backlog-and-the-f
kami:task/445-session-workflow
kami:overnight/eco-versioned-traces
kami:overnight/eco-entity-refs
kami:overnight/eco-degraded-suite
kami:overnight/netscan
kami:overnight/smarthome
kami:overnight/replay-simulator
kami:overnight/event-envelope
kami:overnight/coldstart-unlock
kami:overnight/voice-barge-in
kami:overnight/stt-golden-audio
kami:overnight/senses-speaker
kami:overnight/senses-hearing
kami:overnight/senses-media-vision
kami:overnight/mcp-tools
kami:overnight/mcp-client
kami:overnight/self-update
kami:overnight/model-swap
kami:overnight/web-crawler
kami:overnight/rss-feeds
kami:overnight/email-poller
kami:overnight/email-extract
kami:overnight/email-imap
kami:overnight/money-zenmoney
kami:overnight/task-priority
kami:overnight/task-capture
kami:overnight/behavior-profile
kami:overnight/day-plan
kami:overnight/ambient-calendar
kami:overnight/local-calendar
kami:overnight/proactive-proposals
kami:overnight/split-voice-quiet
kami:overnight/nginx-maven-block
kami:overnight/stepup-chat-surface
kami:integration/small-batch
kami:docs/fix-drift
kami:fix/ru-wording
kami:integration/jul31
kami:overnight/resident-1.7b
kami:overnight/nudge-templates
kami:overnight/kiwix-rewrite
kami:overnight/eval-writeup
kami:overnight/fix-truncation
kami:overnight/kiwix-client
kami:overnight/ru-prompts
kami:overnight/external-data
kami:overnight/phrasing-grammar
kami:overnight/talk-eval
kami:overnight/prompt-context
kami:overnight/prompt-address
kami:overnight/eval-label-kill
kami:overnight/delivery-boundary
kami:overnight/address-check
kami:overnight/system-replies-pr
kami:overnight/clock-intent-pr
kami:overnight/embedder-backfill-pr
kami:overnight/embedder-marker-pr
kami:overnight/note-recall-pr
kami:overnight/thinking-off-pr
kami:overnight/dialogue-persist-pr
kami:overnight/persona-2p-pr
kami:overnight/clarify-expiry-pr
kami:overnight/clarify-rework
kami:overnight/phrasing
kami:overnight/bakeoff
kami:overnight/recall-margin
kami:overnight/router-on
kami:overnight/slot-extract
kami:overnight/embedder-e5
kami:overnight/router-refusal
kami:overnight/eval-rerun
kami:overnight/eval-harnesses
kami:overnight/eval-rerun-base
kami:overnight/fmt-gate
kami:overnight/routines-fire
kami:overnight/router-prompt
kami:overnight/away-leak
kami:overnight/recall-eval
kami:overnight/snooze-works
kami:overnight/clarify-wiring
kami:overnight/delivery-tests
kami:overnight/routine-accept
kami:overnight/llm-router-flag
kami:overnight/clarify-data-layer
kami:overnight/loop-rule-tests
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "overnight/memory-eval"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Ships the part of
docs/plans/03-memory-evaluation.mdthat is real, local and testable now. Maven reads back her own recent memory on a slow ticker, asks the resident model what it notices, and records the confident answers as notes.What changed
internal/memeval(new package).Evaluator.Evaluate(ctx, now)readsRecentFacts/RecentNotes/RecentNudges, prompts the resident model under a GBNF grammar bounded to three{observation, confidence, suggested_action}objects, drops anything undermin_confidence, deduplicates against what earlier evaluations wrote, and records the rest as notes with sourceinfer:memory-eval.cmd/mavend/memoryeval.go— the driver: its own goroutine on its own ticker, wired inmain.gobeside the fact-enrichment worker.internal/config— newmemory_evalblock (interval,max_items,min_confidence).Not
internal/memory/eval.goas the plan says:internal/storeimportsinternal/memoryfor the vector-store backend, and an evaluator has to readstore.Fact/Note/Nudge, which would close the import cycle. Hence a sibling package.Visibility comes free —
/dashalready renders notes with their source, so evaluation output shows up with no UI change.Off unless configured
No
memory_evalblock ⇒ the goroutine does not exist.deploy/mavend.jsonis unchanged, so this branch changes nothing about a running deployment until someone adds the block. No LLM phraser also means no loop: there is no template fallback, because a "memory evaluation" assembled from string templates is a fixed sentence pretending to be an observation.What it deliberately cannot do
This is the feature most likely to turn Maven into a nag — an hourly loop with an LLM in it and permission to talk — so the restraint is structural, not conventional:
/dashwhen he wants to. Wiring observations todelivery.Dispatcheris a separate decision with its own opt-in and is not in this branch.suggested_actionis recorded inside the note text and interpreted by nobody. No reminder, routine or fact is created.Deferred (stated at the bottom of the plan doc too)
/evalIPC method and an evaluation-history view./dashcovers reading the output; a trace surface is worth building once there is real output, and should show the prompt.RecentEventsas an input — action/object events already drive pattern proposals (#43).min_confidencedefault are unvalidated until someone reads a week of real output.Verified
make build,make fmt-checkandmake test(go test -race) clean.internal/memeval/eval_test.gocovers: empty store never reaches the model, confidence floor, dedupe across three runs including a whitespace variant, own notes excluded from input, empty array is not an error, LLM failure writes nothing, and prose-wrapped / over-long replies parse and are capped.Vikunja #248
Ships the real, local, testable part of the memory-evaluation plan (docs/plans/03-memory-evaluation.md): Maven reads back her own recent memory on a slow ticker, asks the resident model what it notices, and records the confident answers as notes. internal/memeval — not internal/memory/eval.go as the plan says, because internal/store imports internal/memory for the vector backend and an evaluator has to read store.Fact/Note/Nudge, which would close the cycle. Evaluate() gathers RecentFacts/RecentNotes/RecentNudges, prompts under a GBNF grammar bounded to three {observation, confidence, suggested_action} objects, drops anything under min_confidence, deduplicates against what earlier runs wrote, and writes the rest as notes with source infer:memory-eval. /dash already renders notes with their source, so the output is visible with no UI change. cmd/mavend/memoryeval.go drives it on its own goroutine and ticker, not on the 60s tick: an evaluation is a multi-second round-trip on the same llama-server that answers voice turns, and it runs hourly at most. The memory_eval config block is absent by default and absence means the goroutine does not exist. No llama-server phraser also means no loop — there is no template fallback, because a "memory evaluation" assembled from templates is a fixed sentence pretending to be an observation. What it deliberately cannot do, since this is the feature most likely to turn Maven into a nag: - It cannot speak. No dispatcher reference, no channel, no nudge. An observation is a thought she wrote down and he reads on /dash. Announcing them is a separate decision with its own opt-in. - It cannot act. suggested_action is recorded as text and interpreted by nobody — no reminder, routine or fact is created from it. - It says nothing about an empty store: no memory means no LLM call, so there are no observations invented out of two facts. - Its own notes are excluded from the next evaluation's input, and are written with a nil embedding so they stay out of the recall pool. The plan's remaining items (dispatching observations, an /eval IPC method and trace view, RecentEvents) and the fact that output quality is entirely unmeasured are written up at the bottom of the plan doc.The restraint is the best part of this. No dispatcher reference at all, so the loop cannot reach a channel by accident.
suggested_actionrecorded and interpreted by nobody. Nil embeddings, so generated text stays out of the RAG pool it came from. Empty store means no LLM call, because a 1.7B asked to find a pattern will always find one. The status section indocs/plans/03-memory-evaluation.mdnaming the plan steps skipped on purpose is worth as much as the code.Three things.
1. The dedupe window is a note count, not an eval-note count.
recordedTextsreadsRecentNotes(ctx, 200)and keeps only theinfer:memory-evalrows. Once 200 ordinary notes are newer than an observation, that observation falls out of the window. The next evaluation is then free to write the same sentence again. That is the exact failure the function exists to prevent, arriving quietly after a few months of normal use. A source-filtered read would make the window mean what the comment says.2. Her own notes eat the input budget.
snapshotasks forMaxItemsnotes, then discards theEvalNoteSourceones. After a few weeks of hourly evaluation, most of the 30 most recent notes are hers, so the model sees a handful of real ones. Theowncounter is computed and never read, which suggests this was noticed and left. Either fetch past the discards, or logownso the shrinking window is visible.3. A five-minute timeout on the shared llama-server. The comment says nobody is waiting on the answer. True of the evaluation. Not true of the voice turn that arrives while it runs. There is one resident model and one server, so a long evaluation is a long stall in front of whoever speaks next. The hourly cadence makes the collision rare rather than impossible. Either shorten the timeout to something a voice turn can absorb, or skip the evaluation when a turn ran recently.
Smaller:
formatNoteappends" [action]", andrecordedTextsstrips it withLastIndex(text, " ["). An observation whose own text ends in a bracketed clause loses part of itself before hashing. It only affects the dedupe key, never the stored note, so this is cosmetic.Landed on master. The stack was one linear chain, so #84 carried every commit from #50 up, and master now contains this branch in full. Merging this PR on its own is an empty diff, so it is closed rather than merged. The review findings for it were fixed in the 2026-08-01 pass and are on master as commits on the stack tip, not on this branch.
Pull request closed