Compare commits

...

56 Commits

Author SHA1 Message Date
kami 5a1d465db5 Merge commit '4383844' into overnight-jul31
# Conflicts:
#	deploy/mavend.json
#	internal/config/config.go
2026-07-31 11:52:21 +04:00
kami 43838445ab Write up the margin gate results
Third section: why the absolute gate could not separate the two
distributions, the delta sweep, and the before/after. Marks next-steps
item 3 done.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:51:21 +04:00
kami 11831c6ace Gate recall on the margin over the runner-up, not just the score
The e5 embedder puts every cosine in one narrow band (0.79-0.89), so the
absolute query_min_score gate cannot tell a real hit from a made-up
question: any value under the band answers everything, any value above it
answers nothing. False recall was 5/5.

New gate asks whether one note is clearly the best instead: top1 - top2 >
delta. New query_min_margin config knob, default 0.008, read off the sweep
in the recall harness. The absolute floor stays as a second check.

On the recall fixture with e5: answered 72% -> 68%, false recall 5/5 -> 1/5.

Vikunja #359

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:50:39 +04:00
kami 93c1a41d4a Route with the resident model by default
The two things that made this unsafe are fixed: the router can now
refuse, and slot extraction runs on its decisions.

On the held-out fixture it gets 63.2% of intents right against the
classifier's 50.0%, with no route errors. It costs about a second a
turn instead of 30ms.

The flag is a pointer now, so leaving it out of the config means on
and only writing false turns it off.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:47:57 +04:00
kami bce5ed210c Merge branch 'worktree-agent-a76ce40c73601d90d' into overnight-jul31 2026-07-31 11:44:32 +04:00
kami c31f0d1001 Extract slots for LLM router decisions too
An LLM-routed reminder came back with no parsed time and an act with no
fn, because only the classifier path ran the extractor. Now the router
runs the same extraction after an LLM decision and fills only the empty
slots. No time in the utterance still means no time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:44:06 +04:00
kami 34521c30b8 Merge branch 'worktree-agent-ad5da57e47b822152' into overnight-jul31 2026-07-31 11:39:35 +04:00
kami 1d48755d12 Record the recall numbers after the embedder swap
recall@1 60% to 72%, answered 48% to 72%, latency 3x better. But false
recall went 1/5 to 5/5: e5 packs every score into a narrow high band, so
the 0.55 gate now admits everything. Left the gate alone as instructed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:38:18 +04:00
kami f6d5a2a7a4 Swap the embedder to multilingual-e5-small (Vikunja #371, #372)
The old model was a symmetric paraphrase model, so it scored "do these
look alike" instead of "does this note answer this question". Also fixes
the file mismatch: the Makefile, the deploy config and both evals now all
name the same quantized file, and the quantized one is what gets measured.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:38:18 +04:00
kami 751c2a705f Embed a question and a stored note differently (Vikunja #371)
Note recall is asymmetric: a short question goes in, a longer note comes
out. Adds EmbedQuery/EmbedPassage helpers and the e5 prefixes, and points
the note/fact write path at the passage side and the query path at the
query side. Reviewers: the three call sites in voice.go.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:38:07 +04:00
kami 5d5b0cfd49 Merge branch 'worktree-agent-a0ab2a9b7439296e3' into overnight-jul31 2026-07-31 11:37:53 +04:00
kami bd16ca69e5 Let the LLM router answer "unknown" when it cannot route
Chose an 8th enum value over a confidence number: the model already picks
one enum token, so it costs nothing in the grammar, while a score from a
0.8B model would be uncalibrated noise. A refusal returns "no decision"
with no error, which is the fall-through the caller already uses for a
bad parse, so the classifier and its clarify gate take the turn.

Reviewers: the prompt's counter-examples matter most — a small model will
over-use any easy escape hatch. The training workspace copy of the prompt
still needs the same edit (Vikunja #362).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:35:37 +04:00
kami 0ed386eca6 Re-measure the router on a quiet box and record the numbers
The earlier before/after was taken while another eval shared
llama-server. This run had the box to itself.

Intent accuracy 61.8% llm-only, 63.2% cascade, 67.1% with thinking off.
The prompt fix holds. note→fact shows up here too, so it is real.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:34:57 +04:00
kami 1c4eab2107 Merge commit '94eb92f' into overnight-jul31
# Conflicts:
#	Makefile
2026-07-31 10:11:56 +04:00
kami 0914e0a3d5 Merge commit '4ba9a6f' into overnight-jul31
# Conflicts:
#	Makefile
2026-07-31 10:11:29 +04:00
kami c860808528 Make make test actually gate on gofmt and vet
DESIGN.md has always said `make test` is "gofmt + vet + -race, no
exceptions". It only ever ran the tests, which is how nine files drifted
out of format without anyone noticing.

`test` now depends on `fmt-check` and `vet`. Checked that fmt-check does
fail when a file is unformatted, so the gate is real and not decorative.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 10:10:27 +04:00
kami f7442c3aea Run gofmt over the seven files that had drifted
Formatting only: import order, and statements that were packed onto one
line split out. `git diff -w` shows nothing but that.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 10:09:32 +04:00
kami 75b067ac51 Merge the accepted-routine fix and drop reminder_id from accept
Two merge fixes on top of the branch:

- migrations: keep both new steps, snooze stays #8, the routine columns
  become #9. Both agents had numbered theirs #8.
- accepting no longer takes a reminder id, on the web surface too. The
  web accept path had the same one-shot-reminder bug the voice path did,
  so both now just flip the status and let the tick loop schedule.

The test that asserted "accept creates a reminder and links it" asserted
the bug. It now asserts that accepting creates no reminder.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 10:07:37 +04:00
kami 424d1b3446 Fire accepted routines every interval, not once (Vikunja #366)
The tick loop now reads accepted routines from the store and nudges when
their interval has passed; accepting no longer builds a one-shot reminder.
Look at routine.DueAccepted for the schedule rule (no catch-up backlog) and
at fireAcceptedRoutines for the restraint gate — routines do not bypass it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:45:30 +04:00
kami c47886c2bc Merge branch 'worktree-agent-ab5b5c61a32cac4fe' into overnight-jul31 2026-07-31 02:42:53 +04:00
kami f6236da760 Collapse the duplicate away-detail and panic tests
Two agents wrote the same three test helpers and names for the same two
bugs. Kept the real assertions from dispatcher_test.go and removed the
skipped placeholders they replace. panicSink stays in durability_test.go
since both files use it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:42:21 +04:00
kami 54dc43516b Add accepted-routine timestamps to the store (Vikunja #366)
Data layer only. Migration #8 adds accepted_ts and last_fired_ts to
proposed_routines, plus ListAcceptedRoutines and MarkRoutineFired so the
tick loop can own the schedule. Accepting no longer links a reminder id.
Look at the TODO(vikunja#366) in cmd/mavend/tick.go for the next commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:41:58 +04:00
kami fa51a48958 Merge branch 'worktree-agent-afe3f2ec18b2b8497' into overnight-jul31 2026-07-31 02:40:10 +04:00
kami 5fd25d7ad7 Test the away-channel minimal body and the panicking sink (#368, #369)
The integration branch names one test TestAwayFallsBackToFullBodyWhenSummaryEmpty,
which describes the old bug; it is here as TestAwaySendsGenericLineWhenSummaryEmpty
and asserts the generic line instead of the body. Also covers: a normal summary
goes out unchanged, voice keeps the full body, and one panicking sink does not
eat the other channel for the same nudge. Reformatted one pre-existing struct.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:39:38 +04:00
kami c22fc352fc Merge branch 'worktree-agent-a1610b8c5376eadd6' into overnight-jul31 2026-07-31 02:37:30 +04:00
kami 215aa331c5 Recover from a panicking sink so the attempt is always closed (#369)
A panic in Send used to unwind past completeOutbox and leave the
delivery_attempts row pending forever, since reconciliation only runs at
startup. safeSend turns the panic into an error, logs it loudly, records the
attempt failed, and lets the other channels for the same nudge still go out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:35:54 +04:00
kami 859bbf750f Never send a nudge body off-box when the summary is empty (#368)
Away channels (ntfy, telegram) leave the box, so an empty Summary now sends
a fixed generic line plus the rule name instead of the full Body. The
dispatcher strips detail before any sink sees it, so a sink added later
cannot leak by reading the wrong field. Voice is local and unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:34:54 +04:00
kami 8a174c1c70 Score the recall fixture and write up what it shows
Real recall is 48% after the gate, and one must-be-silent query gets an
answer anyway. Review finding 2 (the score distributions overlap, so no
gate separates a real recall from a false one) and finding 4 (the memStore
branch at voice.go:776 is unreachable for notes). Adds an embedder cache
so the gate sweep does not re-embed the fixture nine times.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:34:28 +04:00
kami 1f55207b58 Merge the snooze read path and wiring
# Conflicts:
#	internal/loop/gate_test.go
2026-07-31 02:33:41 +04:00
kami 4db109346a Ignore the .claude directory
Agent worktrees land in .claude/worktrees, so the directory shows up as
untracked noise in every git status.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:32:44 +04:00
kami 2f00593411 Wire the snooze read into the Gatherer and honour it for reminders (#364)
The Gatherer now fills State.SnoozeUntil from store.SnoozedUntil instead
of nil, so a snooze finally reaches the gate. RemindDecisions gains the
one restraint check that applies to a reminder — quiet hours, presence
and cooldown are still bypassed, so "wake me 7" is unchanged. Reviewer:
the two tests in internal/loop/gate_test.go are the contract.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:32:24 +04:00
kami 784d688b44 Merge the clarify wiring
# Conflicts:
#	internal/dialogue/clarify.go
#	internal/dialogue/clarify_test.go
2026-07-31 02:30:59 +04:00
kami 4ba9a6f422 Add a deterministic scorer for nudge phrasing (Vikunja #323)
Review internal/phraser/eval/checks.go -- it IS the measurement. Each check
names in a comment which DESIGN.md line it defends: length, feminine
self-reference (windowed around "я" so the operator's own masculine
second-person forms are not flagged), the cringe list (pet names, emoji,
"!!", fake concern, apology, emotional support, asking how he feels,
praise), on-topic, mood enum. No send/veto signal anywhere, per
DESIGN.md § "Rules decide, LLM phrases".
Fixture (158 lines) and tests (252) do not count toward the diff ceiling;
the scorer itself is still ~650. Splitting eval.go from checks.go would
give two commits neither of which measures anything.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:30:52 +04:00
kami 32687b3712 Read the recorded snooze outcomes back out of the nudges table (#364)
The gate honours State.SnoozeUntil but nothing ever filled it. New
store.SnoozedUntil returns, per rule, when the newest snooze runs out.
Reviewer: the fixed 2h SnoozeDuration and its reasoning in nudges.go —
nothing upstream can supply a per-nudge length, so no new column.
Expired snoozes are dropped in SQL, so silence can never be permanent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:29:51 +04:00
kami db3e706bdc Merge the delivery routing and durability tests 2026-07-31 02:29:05 +04:00
kami 2f4257e194 Test the clarify round-trip end to end at the daemon level
Covers: a reminder with no time is asked about and completes on the answer; the
same for a fact; an answer past the TTL falls through as a fresh utterance; a
second unclear answer drops the request with no second question; a clarified act
off the allowlist neither runs nor gets enabled; a clarified destructive act
still parks a confirm; noise keeps the canned reply. No model, no network.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:29:04 +04:00
kami 7d8b0af99d Test voice fallthrough and delivery durability
Fallthrough is checked per severity through the outbox trail, so sev3/sev4
reroute and sev1/sev2 still drop. Durability uses a real store on a temp file:
a crash between Begin and Complete becomes unknown, is not resent, is not
dropped, and a late Complete cannot overwrite it. One skipped test marks a real
gap: a panic mid-send leaves a permanent pending row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:27:33 +04:00
kami cfd38d53cf Ask the question, then act on the answer
On a clarify decision with one identifiable gap she now asks instead of saying
"не поняла", and parks the request. The next utterance is parsed as the answer
with the router's own extractor and the completed decision runs through
applyAction like any other — so a clarified act still needs the allowlist and
still hits the destructive confirm gate. An answer that does not fill the gap
drops the request; she never asks twice. Also pulls the session-store block
that HandlePushToTalk and handleText both had into rememberTurn, since the
clarify path needed a third copy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:26:11 +04:00
kami a2835bbdf6 Cover every cell of the delivery routing table
Table-driven tests for all four severity bands crossed with present and away,
both as the pure table and end to end through the dispatcher. Three tests are
written to DESIGN.md and skipped because the code does not keep the claim: the
care-away drop is recorded nowhere, and the minimal body is enforced per-sink
rather than by the dispatcher. Also gofmt'd dispatcher_test.go.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:24:43 +04:00
kami c262c1e4c0 Merge the routine accept path and store gaps 2026-07-31 02:24:07 +04:00
kami a33ad82178 Let the /routines page accept a proposal, gated at step-up (#46)
Accepting a routine gives the trigger loop a new standing reason to speak to
the human, so it is the same authority tier as enabling a tool and shares the
stepUpOK gate; dismiss only ever makes maven quieter, so it is ungated.
Look at handleRoutines and acceptRoutine in cmd/mavweb/main.go: accept creates
the recurring reminder, then links it via the new ipc AcceptProposedRoutine.
The page now says what maven noticed in her own words (pattern.PhraseRoutine).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:23:28 +04:00
kami af35ec3629 Work out which slot is missing and phrase one short question
A table per intent (reminder needs a time, fact needs a key, act needs a fn)
plus one fixed Russian question per slot. Templates, not model output: a 0.8B
would wander and a question that rewords itself is harder to answer. Note,
query, chat and system get no question — for those a clarify decision keeps
the canned reply rather than inventing a question for noise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:22:18 +04:00
kami 9145b83100 Add the clarify data layer: a parked question with one missing slot
The router can already say "I am not sure" (Decision.Clarify) but the daemon
had nowhere to keep the request while it asked. PendingQuestion holds the
original slots, ClarifyStore parks one per dialogue id with a 90s TTL, and
Answer fills only the slots that were missing so an answer can never rewrite
what she already understood. Logic that uses this comes next.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:20:28 +04:00
kami 43470abc57 Add a held-out note-recall harness (fixture + scorer)
Measures whether Maven can find the right note again from a paraphrased
question. Review internal/memory/recalleval/recalleval.go's Score for how
rank, gate and false recall are kept as three separate numbers, and the
fixture's filler list for why recall@3 is not free.
Fixture JSON is generated data and does not count toward the diff limit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:18:58 +04:00
kami 94eb92fb15 Label LLM eval runs with the model llama-server has loaded
The bake-off in #278/#250 needs two models' scores side by side, and the
report names only carried the config, so the rows were indistinguishable.
ModelID reads /v1/models instead of taking a string that goes stale.
New target: make eval-models MAVEN_LLM_URL=...

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:17:20 +04:00
kami 707c3e5040 Ignore deps and models as symlinks, not just directories
.gitignore had deps/ and /models/llm/ with trailing slashes. A trailing slash
only matches a real directory, so a *symlink* with the same name is not ignored
and git add -A commits it as a symlink blob.

That bites anyone working in a git worktree, where deps/ and models/ do not
exist and have to be linked in from the main checkout. It already happened once
tonight.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:17:15 +04:00
kami ace9fbbc06 Merge the llm_router flag and the kill-maven fix 2026-07-31 02:17:00 +04:00
kami 3884db33e9 Give proposed routines a status filter and pin down the dedup rule (#46)
Look at internal/store/proposed_routines.go: status flips in place with an
`AND status = 'proposed'` guard, not append-only like facts/voids_id — a
proposal is a question with one answer, same shape as tools.status. The
UNIQUE(action, object) key is what stops a dismissed routine coming back.
New tests cover re-propose-after-dismiss and listing by status.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:16:39 +04:00
kami dfb8d26b62 Fix kill-maven.sh so it actually kills llama-server
The MODEL default was LFM2, but the deploy runs Qwen3.5-0.8B, so the
pkill pattern matched nothing and the server survived every kill.
Now matches any llama-server serving a .gguf, so changing the model in
deploy/mavend.json cannot break the script again. MODEL still narrows it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:15:30 +04:00
kami bf99fd4192 Add a voice.llm_router flag, default off
Wires cmd/mavend/voice.go to build the LLM router when the operator asks
for it. Default false, so nothing changes on the deploy box.
Look at pickLLMRouter: the flag on with no llama-server logs one line and
keeps the classifier, it never fails a turn.
The default stays off until the router can refuse (#359) and the extractor
runs on LLM decisions — both noted as TODOs in config.go.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:15:30 +04:00
kami 6d9aa83b6e Merge the proactive rule and gate tests 2026-07-31 02:15:15 +04:00
kami f75072175d Merge the pending-question data layer 2026-07-31 02:14:33 +04:00
kami 39d83a33e8 Add the pending-question data layer for slot clarification
PendingQuestion plus ClarifyStore: same shape, locking and expiry as SessionStore. Answer fills only the missing slots and never overwrites a filled one. No wiring yet — TODOs mark the daemon hooks.
Reviewer: MaxAttempts is 1 on purpose (Maven asks once, she is not a nag).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:13:48 +04:00
kami 925ce223a0 Add a Value slot to dialogue.Slots
router.Slots already carries the fact payload; the dialogue copy did not, so a clarifying answer had nowhere to put it. InheritSlots carries it like Key.
Reviewer: check the new inherit block does not overwrite a filled value.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:13:40 +04:00
kami d30618ecb7 Stop the router repetition loop
Route now sets RepeatPenalty on the request, and the grammar's string rule is
capped at 120 characters. Two of 76 fixture cases looped one sentence inside
the text field until MaxTokens, which cut the JSON in half.
Reviewers: the new constant and the grammar string rule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:11:38 +04:00
kami 17b47ce206 Route questions to query, not fact
The router prompt tested "reports current state -> fact" before "wants
information -> query", so a question naming a fact key was written as a fact.
Query now comes first, plus an explicit question test.
Reviewers: the prompt block in llmrouter.go, and the note about the
training-side copy of the prompt that needs the same edit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:10:37 +04:00
70 changed files with 6012 additions and 248 deletions
+10 -1
View File
@@ -10,8 +10,11 @@
# Certs (private keys, don't commit)
certs/
# Dependencies (fetch/build, not vendored)
# Dependencies (fetch/build, not vendored).
# Both forms on purpose: 'deps/' misses a symlink named deps, and agents working
# in a git worktree symlink these in from the main checkout.
deps/
deps
# ML models (large, downloaded separately) — specific dirs, not blanket,
# because models/seeds/*.txt are small, tracked files the classifier needs.
@@ -19,6 +22,9 @@ deps/
/models/stt/
/models/tts/
/models/llm/
# Symlink forms, same reason as deps above.
/models/embedder
/models/llm
# Runtime data
*.db
@@ -36,3 +42,6 @@ opencode.json
# Test coverage output
coverage.out
# Agent worktrees and local agent state
.claude/
+8 -5
View File
@@ -43,14 +43,17 @@ notes. Without it, the floor `HashEmbedder` is used — deterministic but weak
(Russian recall rarely clears the confidence gate, many commands fall to
"clarify").
**Download the embedder** (ONNX, ~90 MB):
**Download the embedder** (ONNX, ~120 MB):
```sh
make download-embedder
```
This fetches `paraphrase-multilingual-MiniLM-L12-v2` (384-dim, 12-layer,
supports 50+ languages including Russian) to `models/embedder/`.
This fetches `multilingual-e5-small` (384-dim, 12-layer, Russian and English)
to `models/embedder/multilingual-e5-small/`. It is an asymmetric retrieval
model: the code puts `query: ` in front of a question and `passage: ` in front
of a stored note, which is how e5 was trained. The quantized file is the one
that is downloaded, deployed and measured.
**Also need ONNX Runtime** (`libonnxruntime.so`):
@@ -64,8 +67,8 @@ sudo cp onnxruntime-linux-x64-1.15.1/lib/libonnxruntime.so* /usr/local/lib/
```json
"voice": {
"embedder": {
"model_path": "models/embedder/model_quantized.onnx",
"tokenizer_path": "models/embedder/tokenizer.json",
"model_path": "models/embedder/multilingual-e5-small/model_quantized.onnx",
"tokenizer_path": "models/embedder/multilingual-e5-small/tokenizer.json",
"lib_path": "/usr/local/lib/libonnxruntime.so"
}
}
+53 -5
View File
@@ -16,7 +16,7 @@ PIPER_BIN := $(shell pwd)/deps/piper/piper
PIPER_MODEL := $(shell pwd)/models/tts/ru_RU-irina-medium.onnx
PIPER_ESPEAK := $(shell pwd)/deps/piper/espeak-ng-data
.PHONY: all build build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav clean test run-stt run-tts run-web download-embedder deps-go eval-router
.PHONY: all build build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav clean test fmt-check vet run-stt run-tts run-web download-embedder deps-go eval-router eval-recall eval-phrasing eval-models
all: build
@@ -69,7 +69,20 @@ deps-go:
done
$(GO) version
test:
# fmt-check fails if any file needs gofmt. DESIGN.md has always said `make
# test` gates on gofmt and vet; it did not, so nine files quietly drifted.
# Run `gofmt -w` on whatever this prints.
fmt-check:
@bad=$$(gofmt -l internal cmd); \
if [ -n "$$bad" ]; then \
echo "these files need gofmt:"; echo "$$bad"; exit 1; \
fi
vet:
CGO_CFLAGS="$(CGO_CFLAGS)" CGO_LDFLAGS="$(CGO_LDFLAGS)" LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
$(GO) vet ./internal/... ./cmd/...
test: fmt-check vet
CGO_CFLAGS="$(CGO_CFLAGS)" CGO_LDFLAGS="$(CGO_LDFLAGS)" LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
$(GO) test -race -coverprofile=coverage.out ./internal/... ./cmd/...
@@ -83,6 +96,37 @@ MAVEN_ONNX_LIB ?= $(shell pwd)/deps/onnxruntime-linux-x64-1.26.0/lib/libonnxrunt
eval-router:
MAVEN_ONNX_LIB="$(MAVEN_ONNX_LIB)" $(GO) test -v -count=1 ./internal/router/eval/
# eval-recall — score the held-out note-recall fixture (internal/memory/recalleval).
# Answers "can she find the note again when it matters": recall@1, recall@3,
# false recall and the query_min_score sweep. Same MAVEN_ONNX_LIB deal as
# eval-router; without it only the deterministic hash ratchet runs.
eval-recall:
MAVEN_ONNX_LIB="$(MAVEN_ONNX_LIB)" $(GO) test -v -count=1 ./internal/memory/recalleval/
# eval-phrasing -- score nudge phrasing (internal/phraser/eval). Verbose so the
# report and every generated message land in the terminal. With no environment
# it scores the deterministic Stub only, which is what CI runs. Set
# MAVEN_LLM_URL to add the resident model:
# MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing
# The model run is slow (minutes) -- the timeout is raised to match.
eval-phrasing:
$(GO) test -v -count=1 -timeout 40m ./internal/phraser/eval/
# eval-models — score ONE llama-server against the same fixture, for the
# resident-model bake-off (#278, #250). Start a server with the gguf you want,
# then:
#
# make eval-models MAVEN_LLM_URL=http://127.0.0.1:18100
#
# The report names carry the model llama-server reports, so runs from two
# checkpoints stay apart. Only the LLM test runs — the classifier baselines do
# not depend on the model and take the ONNX runtime with them.
MAVEN_LLM_URL ?= http://127.0.0.1:18099
eval-models:
MAVEN_LLM_URL="$(MAVEN_LLM_URL)" $(GO) test -v -count=1 -timeout 60m \
-run TestLLMRouterBaseline ./internal/router/eval/
run-stt: build-stt
LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
./mavsttd -socket /tmp/maven/stt.sock -model $(WHISPER_MODEL)
@@ -108,9 +152,13 @@ deps-piper:
-o /tmp/piper.tar.gz
tar -xzf /tmp/piper.tar.gz -C deps/
EMBEDDER_DIR := $(shell pwd)/models/embedder
EMBEDDER_MODEL_URL := https://huggingface.co/Xenova/paraphrase-multilingual-MiniLM-L12-v2/resolve/main/onnx/model_quantized.onnx
EMBEDDER_TOKENIZER_URL := https://huggingface.co/Xenova/paraphrase-multilingual-MiniLM-L12-v2/resolve/main/tokenizer.json
# multilingual-e5-small: an asymmetric retrieval model. It is trained to match
# a short question against a longer passage, which is what note recall is.
# The quantized file is the one we download, deploy and measure — see
# RECALL-EVAL-31-07-2026.md.
EMBEDDER_DIR := $(shell pwd)/models/embedder/multilingual-e5-small
EMBEDDER_MODEL_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/onnx/model_quantized.onnx
EMBEDDER_TOKENIZER_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/tokenizer.json
download-embedder:
mkdir -p $(EMBEDDER_DIR)
+242
View File
@@ -0,0 +1,242 @@
# Note recall evaluation — 31-07-2026
The operator's goal is that Maven "memorize/note things … and know more about me/world". This
measures whether the note/recall path delivers that.
- Fixture + scorer: `internal/memory/recalleval/` (`ru_recall_v1.json`, 30 cases)
- Reproduce: `make eval-recall` — hash ratchet always, ONNX when `deps/` is present
- Commit: `43470ab` (harness)
Each case inserts its own 3 notes **plus 12 shared filler notes** into a fresh store, embeds the
query, takes the top 3 — the read path `cmd/mavend/voice.go` runs for `IntentQuery`. Filler is
load-bearing: with 3 notes and a top-3 search, recall@3 is 100% by construction. 25 answerable
cases (paraphrased queries, homelab and preference content, 9 with a plausible second note) and 5
that must recall **nothing**. `TestFixtureIsParaphrased` fails the build if a query shares over half
its words with its note; equal-score ties count as ties, not recall.
## Results
| | recall+hash (CI ratchet) | recall+onnx (deployed) |
|---|---|---|
| **recall@1** | 36.0% (9/25) | **60.0% (15/25)** |
| recall@3 | 76.0% (19/25) | 80.0% (20/25) |
| **answered after the 0.55 gate** | **0.0% (0/25)** | **48.0% (12/25)** |
| wrong note on top / tie on top | 9 / 7 | 10 / 0 |
| ranked first, then silenced by the gate | 9 | 3 |
| **false recall** | 0/5 | **1/5 (20%)** |
| top-1 score when right, min / median | n/a | 0.559 / 0.678 |
| top-1 when it must stay silent, median / max | 0.000 / 0.144 | 0.470 / **0.567** |
| RU / EN / `hard` cases passed | 4/24 / 1/6 / 0/11 | 13/24 / 3/6 / 2/11 |
| latency p50 / p95 / max | 49µs / 70µs | 59ms / 148ms / 194ms |
Never compare a hash-embedder number to an ONNX one — the hash floor is lexical and exists only so
CI has a deterministic ratchet with no model files.
## Findings
### 1. Real recall is 48%, not 60%
The right note ranks first 60% of the time, but the daemon only *says* it 48% of the time — three
more cases rank first and are then silenced by `voice.go:776`'s `queryMinScore`. **Roughly one
useful question in two gets "не знаю".** This is not a working memory yet.
### 2. The gate cannot separate a real recall from a false one — the distributions overlap
Right-note top-1 scores start at **0.559**. Must-stay-silent top-1 scores reach **0.567**. No
threshold keeps every real recall and rejects every false one. From the sweep: gate 0.50 → 13/25
answered, 1/5 false; **0.55 (default) → 12/25, 1/5**; **0.60 → 10/25, 0/5**; 0.70 → 5/25, 0/5. What
the data says about `DefaultQueryMinScore` (`internal/config/config.go:392`): **0.55 is
slightly too loose** — it admits one confident wrong answer ("как зовут сестру моего коллеги"
recalls "выучил пару аккордов на гитаре" at 0.567), which the spec ranks as worse than a gap. 0.60
silences all five and costs 8 points of real recall. Left alone as instructed; the overlap means
the threshold is the wrong dial anyway (finding 3).
### 3. Filler notes outrank the right answer — the model scores similarity, not relevance
`models/embedder/` is **paraphrase-multilingual-MiniLM-L12-v2** (`Makefile:119`), a *symmetric*
paraphrase model. It scores "do these sentences look alike", not "does this passage answer this
question", so question-shaped queries drift to whatever note is stylistically closest. "из-за чего
кончилось место" and "откуда берётся токен бота" both return `выучил пару аккордов на гитаре`
(0.730, 0.729); "как я восстановил конфиги" returns a bootloader note at 0.703 with the right note
not even in the top 3. An unrelated guitar note beating a homelab note at 0.73 is not a tuning
problem — an asymmetric retrieval model (`multilingual-e5-small`, with `query:` / `passage:`
prefixes) is the targeted fix, and it would move findings 1 and 2 together. Separately:
`deploy/mavend.json:39` loads a 470MB fp32 `model.onnx` while `make download-embedder` fetches
`model_quantized.onnx` — not the same file.
`hard` cases score **2/11**: every one is a query where the operator did not reuse his own words.
That is the normal case weeks later, and exactly what DESIGN.md's "recall when relevant" promises.
### 4. The memory-store recall branch is dead for notes
`voice.go:776` only reaches `h.memStore.Search` when the notes-RAG top score is already below
`queryMinScore`, and `bestRecall` (`cmd/mavend/recall.go:19`) then applies the **same** gate to the
same vector. A note is indexed in both places with the same embedding, so if it failed the gate in
`QueryNotes` it fails again here — the branch can only ever return a **fact**. Its comment calls it
"additive"; for notes it is not.
### 5. Ranking has no recency or type signal, and the store is not the bottleneck
`internal/store/notes.go:67` sorts by cosine and uses `ts` only to break an exact float tie, which
never happens; `kind` never enters the ranking. Meanwhile `TestPersistentStoreScoresTheSame` scores
sqlite-backed `store.MemoryStore` and `memory.InMemoryStore` identically — both full-scan cosine
(`internal/store/memory.go:64`) at ~150µs over 42 rows against a ~59ms query embed. An ANN index is
not the problem to solve.
## Re-measured after the embedder swap — 31-07-2026, later the same day
Changed: `models/embedder/` is now **multilingual-e5-small** (quantized, 118MB), with `query: ` in
front of a question and `passage: ` in front of a stored note (Vikunja #371). `deploy/mavend.json`
and `make download-embedder` now name the same file, and it is the quantized one — that is what the
column below measures (Vikunja #372). Everything else is unchanged: same fixture, same store, same
0.55 gate. The old column is the baseline and is left as it was.
| | recall+onnx, MiniLM (baseline) | recall+onnx, e5-small (new) |
|---|---|---|
| **recall@1** | 60.0% (15/25) | **72.0% (18/25)** |
| recall@3 | 80.0% (20/25) | 84.0% (21/25) |
| **answered after the 0.55 gate** | 48.0% (12/25) | **72.0% (18/25)** |
| wrong note on top / tie on top | 10 / 0 | 7 / 0 |
| ranked first, then silenced by the gate | 3 | 0 |
| **false recall** | 1/5 (20%) | **5/5 (100%)** |
| top-1 score when right, min / median | 0.559 / 0.678 | 0.791 / 0.857 |
| top-1 when it must stay silent, median / max | 0.470 / 0.567 | 0.815 / 0.835 |
| RU / EN / `hard` cases passed | 13/24 / 3/6 / 2/11 | 14/24 / 4/6 / 5/11 |
| latency p50 / p95 / max | 59ms / 148ms / 194ms | 18ms / 37ms / 49ms |
### What moved
Ranking got better and got faster. Half the previously-unwinnable `hard` cases now pass (2/11 →
5/11), the guitar note no longer beats the docker-logs note, and the gate stops silencing notes that
already ranked first. The quantized e5 is also ~3x quicker than the fp32 MiniLM it replaces.
### What got worse: the gate is now a no-op
e5 packs every cosine into a narrow high band. Right-note scores start at 0.791; must-stay-silent
scores reach 0.835. **The distributions still overlap, and now they overlap above the gate**, so
0.55 admits everything and false recall goes from 1/5 to 5/5. The sweep:
```
gate 0.500.70: answered 18/25 (72%) false recall 5/5
gate 0.80: answered 17/25 (68%) false recall 4/5
gate 0.90: answered 0/25 ( 0%) false recall 0/5
```
There is no value that keeps real recall and rejects made-up questions — same conclusion as before,
now with a wider band and no room at all. `query_min_score` was left at 0.55 as instructed. **The
recommendation is to leave it there and stop tuning it**: any number under ~0.79 is a no-op and
anything above starts cutting real recall long before it stops the false ones. The fix is a margin
gate (`top1 top2 > δ`), next-steps item 3, which is now the top item.
### The prefixes did not do the work
A control run with both prefixes set to the empty string scored the **same** recall@1 (72%), a
slightly better recall@3 (88%) and the same 5/5 false recall. So on this fixture the gain comes from
the model, not from the `query:` / `passage:` split. The prefixes are kept because they are how e5
was trained and the split is the right shape for the read path, but they are not worth defending on
this evidence — a bigger fixture may say otherwise.
### Stored vectors from the old model are now junk
Cosine between a MiniLM vector and an e5 vector means nothing. Every row already in `notes` and in
the vector memory table was written by the old model, so after this deploy they will score as noise
against a new query. A live database needs every note and fact re-embedded before recall works at
all. Filed as its own task.
## Margin gate — 31-07-2026, third run
Next-steps item 3, done. The absolute gate is replaced by a **margin gate**: answer only when the
top hit beats the runner-up by more than delta (`top1 top2 > δ`). Same fixture, same e5 embedder,
same store as the run above. `internal/memory/gate.go` holds the check; both read paths call it
(`cmd/mavend/recall.go` and the notes-RAG branch in `voice.go`). New knob `voice.query_min_margin`
in `deploy/mavend.json`, default 0.008.
### Why the absolute gate could not work, in one line of data
The harness now prints the margin distributions, and they barely overlap where the raw scores
overlap completely:
| | top-1 score | margin (top1 top2) |
|---|---|---|
| right note first (n=18) | min 0.810, median 0.862, max 0.890 | min 0.001, median 0.029, max 0.053 |
| must stay silent (n=5) | min 0.795, median 0.815, max 0.835 | min 0.000, median 0.002, **max 0.019** |
Four of the five must-be-silent cases have a margin at or under 0.002 — when there is nothing to
recall, e5 finds several notes equally close and no clear winner. That is the signal the absolute
score throws away.
### The delta sweep
Absolute gate held at 0.55 throughout.
```
delta 0.000: answered 18/25 (72%) false recall 5/5
delta 0.002: answered 17/25 (68%) false recall 3/5
delta 0.005: answered 17/25 (68%) false recall 2/5
delta 0.008: answered 17/25 (68%) false recall 1/5 <- chosen
delta 0.010: answered 15/25 (60%) false recall 1/5
delta 0.012: answered 14/25 (56%) false recall 1/5
delta 0.015: answered 12/25 (48%) false recall 1/5
delta 0.020: answered 11/25 (44%) false recall 0/5
delta 0.025: answered 9/25 (36%) false recall 0/5
delta 0.030: answered 8/25 (32%) false recall 0/5
delta 0.040: answered 4/25 (16%) false recall 0/5
delta 0.050: answered 2/25 ( 8%) false recall 0/5
delta 0.060: answered 0/25 ( 0%) false recall 0/5
```
### Chosen: δ = 0.008
It is the best point on the frontier, not a taste call. **0.008 dominates 0.010, 0.012 and 0.015
outright** — same 1/5 false recall, 8 to 20 points more real recall. Everything below it buys recall
back only by admitting more false recalls (0.005 → 2/5, 0.002 → 3/5). The next real improvement is
0.020 at 0/5 false, and it costs 24 points of recall to get there.
The brief's bar was "recall above 60% with false recall at 1/5 or better". 0.008 clears it with room:
68% and 1/5.
### Before / after
| | absolute gate 0.55 (previous) | margin gate δ=0.008 |
|---|---|---|
| recall@1 (ranking, ungated) | 72.0% (18/25) | 72.0% (18/25) — unchanged, the gate does not rank |
| **answered after the gate** | 72.0% (18/25) | **68.0% (17/25)** |
| **false recall** | **5/5 (100%)** | **1/5 (20%)** |
| fixture cases passed | 18/30 | **21/30** |
Four false recalls removed for one real answer. That is the trade the spec asks for — she is not a
guesser-of-truth. The one survivor is `en-pref-025` ("should i be offered wine"), which recalls a
filler note at 0.796 with a 0.019 margin: the widest silent-case margin in the fixture, and it sits
inside the real-recall range, so no delta removes it without taking real answers with it.
### Does the absolute cutoff still earn its keep? Marginally — kept
On this fixture with e5 it is a **no-op**: the lowest right-note score is 0.791, so 0.55 rejects
nothing the margin does not already reject. It is kept for two reasons, neither glamorous. It still
does real work for the hash embedder (its own sweep shows answers dropping from 16% to 0% between
0.30 and 0.50), and it is the only thing standing between the user and a reply built from a store
where everything is far away but one row happens to be a little less far — a near-empty database, or
the stale-vector case below. Cheap insurance, no measured cost. If a later embedder makes it bite,
the sweep is one command.
### Caveat on the numbers
Five must-be-silent cases is a thin basis for a 4-point decision. 1/5 and 2/5 differ by one case.
The shape of the frontier is trustworthy — margins separate, absolute scores do not — but δ=0.008
itself should be re-read off a bigger fixture (next-steps item 6) before anyone defends the third
decimal.
## Next steps — ordered by value-to-risk; nothing here is a decision
1. **Swap the embedder to `multilingual-e5-small` with `query:`/`passage:` prefixes.** One config
change plus a prefix in `onnxembedder.go`, re-measurable in one command.
2. **Re-run `make eval-recall`, then set the gate from the sweep** — not before. Any
`query_min_score` picked against today's embedder describes a model on its way out.
3. ~~**Replace the absolute-score gate with a margin gate**~~ — done, see the section above.
δ=0.008, false recall 5/5 → 1/5.
4. **Delete or repair the dead `memStore` branch** at `voice.go:776` — search before the gate,
gate it separately, or restrict it to facts and say so.
5. **Add a mild time decay to ranking** — the newest statement of a preference is the true one.
6. **Grow the fixture from real misses.** 30 cases can rank two embedders, not trust 4 points.
7. **Re-measure end to end.** Recall is gated twice — the utterance must first route to `query`,
which the routing eval puts at ~50%. The product is ~24%, and that is what he experiences.
+35
View File
@@ -40,6 +40,41 @@ model → classifier as failure floor.
Never compare a hash-embedder run to an ONNX one.
## Re-measured after the prompt fix
The table above is the **baseline at commit `46259b4`**, kept as-is. The prompt fix (query
tested before fact, plus `repeat_penalty` and a bounded grammar string) was then measured on
an otherwise idle box — no other eval sharing llama-server, so these latencies are real
rather than contention.
| | llm-only (0.8B) | cascade+llm (0.8B) | llm-only, thinking off |
|---|---|---|---|
| **intent-only accuracy** | 48.7% → **61.8%** | 50.0% → **63.2%** | **67.1%** |
| full accuracy (intent+slots+gate) | 23.7% → **38.2%** | 32.9% → **47.4%** | **42.1%** |
| route errors | 2 → **0** | 0 → 0 | **0** |
| p50 / p95 latency | **1.08s / 1.55s** | **1.04s / 1.53s** | **0.93s / 1.41s** |
Three things this run settles:
1. **The prompt fix holds.** An earlier contended run reported 60.5% / 36.8% for llm-only;
the quiet run gives 61.8% / 38.2%. Close enough to call the gain real, and the earlier
run's 4-5s latency figures were contention, not the model.
2. **`query→fact` fell from ×15 to ×7**, and both unparseable replies are gone. Zero route
errors in every LLM configuration.
3. **`note→fact ×4` is real, not noise.** It shows up in the quiet run too. The agent that
wrote the prompt fix suspected its own change might have caused it by pulling assertive
`запиши что…` phrasings toward fact, and that suspicion stands — all five `ru-note-*`
cases now land on fact. Tracked as Vikunja #375.
**Thinking off is the best configuration measured so far**, on both accuracy and latency
(Vikunja #376). That is worth understanding before flipping: routing is a short
classification into a fixed enum with grammar-constrained output, so there is little to
reason about, and the thinking trace mostly gives a small model room to talk itself out of
the right answer. Phrasing is a different job and needs measuring separately.
Still `6 / 6` missed clarify — the router has no way to say "I don't know" (Vikunja #359).
That is unchanged by anything here.
## Findings
### 1. The resident model does route better — 50.0% vs 36.8%
+180
View File
@@ -0,0 +1,180 @@
package main
import (
"context"
"log"
"time"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/router"
)
// clarifyTTL — how long a parked question stays answerable. Same 90s as the
// confirm gate, for the same reason: an answer is a same-breath gesture, and a
// stale question must not eat an unrelated later utterance.
const clarifyTTL = 90 * time.Second
// wantedSlots — what each intent needs before she can act on it. First entry is
// the one she asks about; the rest are only used to decide act-vs-drop.
//
// Intents not listed here are never worth a question: note and query act on the
// raw utterance, chat and system have nothing to fill in. For those a clarify
// decision keeps the canned "не поняла" reply — inventing a question for noise
// is worse than admitting she missed it.
var wantedSlots = map[router.Intent][]dialogue.Slot{
router.IntentReminder: {dialogue.SlotTime},
router.IntentFact: {dialogue.SlotKey},
router.IntentAct: {dialogue.SlotFn},
}
// clarifyQuestions — one short question per missing slot.
//
// These are fixed templates, not model output. The resident model is a 0.8B; it
// would wander, and a question whose wording changes every time is harder to
// answer than a blunt one that always reads the same. They are infinitive
// questions, so there is no gender agreement to get wrong; the feminine
// self-reference lives in the reply she gives when she drops the request.
var clarifyQuestions = map[dialogue.Slot]string{
dialogue.SlotTime: "На когда напомнить?",
dialogue.SlotKey: "Что записать?",
dialogue.SlotFn: "Что сделать?",
}
// clarifyDropped — she asked once, the answer still did not fill the gap, so
// the request is gone. Said plainly, once, with no second question.
const clarifyDropped = "Не разобрала — скажи целиком, пожалуйста."
// missingFor returns the slots a decision still needs, most important first.
// Empty ⇒ there is nothing identifiable to ask about.
func missingFor(dec router.Decision) []dialogue.Slot {
return dialogue.StillMissing(wantedSlots[dec.Intent], toDialogueSlots(dec.Slots))
}
// clarifyQuestion picks the one question to ask for a clarify decision. Returns
// ("", false) when she has no idea what is missing.
//
// One question about one thing: if two slots are missing she asks about the
// first and lets the rest go. Two questions in a row is an interrogation.
func clarifyQuestion(dec router.Decision) (dialogue.Slot, string, bool) {
missing := missingFor(dec)
if len(missing) == 0 {
return "", "", false
}
q, ok := clarifyQuestions[missing[0]]
if !ok {
return "", "", false
}
return missing[0], q, true
}
// askClarify parks the request and returns the question to ask instead of the
// canned "не поняла". Returns ("", false) when there is nothing to ask about, so
// the caller falls back to the canned reply.
func (h *reactiveHandler) askClarify(dec router.Decision) (string, bool) {
if h.clarifyStore == nil {
return "", false
}
slot, question, ok := clarifyQuestion(dec)
if !ok {
return "", false
}
h.clarifyStore.Put(voiceDialogueID, &dialogue.PendingQuestion{
Intent: dialogue.Intent(dec.Intent),
Slots: toDialogueSlots(dec.Slots),
Missing: []dialogue.Slot{slot},
Utterance: dec.Utterance,
Asked: h.now(),
TTL: clarifyTTL,
Attempts: 1, // asked once; MaxAttempts is 1, so there is no second ask
})
log.Printf("voice: clarify — asked about %s for intent=%s", slot, dec.Intent)
return question, true
}
// resolveClarifyAnswer reads an utterance as the answer to a parked question.
// Returns ("", false) when no live question is parked (or it expired), so the
// caller routes the utterance normally as a fresh request. Sibling of
// resolveConfirm and checked in the same place.
//
// The answer is parsed with the same extractor the router uses, for the intent
// she parked — no second parser. If it still does not fill the gap the request
// is dropped: she does not ask again.
func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string) (string, bool) {
if h.clarifyStore == nil {
return "", false
}
q := h.clarifyStore.Get(voiceDialogueID, h.now())
if q == nil {
return "", false
}
// One shot either way: the question is consumed whether or not the answer
// works, so a failed answer can't leave the question armed.
h.clarifyStore.Delete(voiceDialogueID)
intent := router.Intent(q.Intent)
answer := h.extractor.Extract(ctx, intent, text, h.now())
merged := q.Answer(text, toDialogueSlots(answer))
if len(dialogue.StillMissing(q.Missing, merged)) > 0 {
log.Printf("voice: clarify — answer %q did not fill %v, dropping", text, q.Missing)
return clarifyDropped, true
}
// Rebuild the decision as if it had routed cleanly, then run it down the
// normal path. Clarify is deliberately false and the intent is unchanged:
// filling in an argument never grants authority, so the completed decision
// still meets the allowlist and the destructive-act confirm gate in
// applyAction exactly like any other decision.
dec := router.Decision{
Utterance: q.Utterance,
Stage: 2,
Intent: intent,
Slots: applyDialogueSlots(answer, merged),
}
return h.finishClarified(ctx, dec), true
}
// finishClarified runs a completed decision through the same steps a freshly
// routed one takes: remember the turn, act, then phrase.
func (h *reactiveHandler) finishClarified(ctx context.Context, dec router.Decision) string {
if h.dialogueSessions != nil {
now := h.now()
prev := h.dialogueSessions.Get(voiceDialogueID, now)
dec = followUpMerge(prev, dec, now)
h.rememberTurn(prev, dec, now)
}
reply := h.applyAction(ctx, dec)
if reply == "" {
reply = h.replier.Reply(dec)
}
return reply
}
// rememberTurn stores this turn as the dialogue session the next follow-up
// inherits from, carrying up to 4 prior turns of history for anaphora. Capped so
// one long conversation can't grow the session unboundedly.
func (h *reactiveHandler) rememberTurn(prev *dialogue.Session, dec router.Decision, now time.Time) {
var history []dialogue.Turn
if prev != nil {
history = append(history, dialogue.Turn{
Intent: prev.Intent,
Slots: prev.Slots,
Text: prev.Slots.Text,
})
maxHist := len(prev.History)
if maxHist > 3 {
maxHist = 3
}
history = append(history, prev.History[:maxHist]...)
}
ttl := time.Duration(0) // use the store default (2 min)
if dec.Intent == router.IntentChat {
ttl = 15 * time.Minute // conversational turns should last longer
}
h.dialogueSessions.Put(voiceDialogueID, &dialogue.Session{
Intent: dialogue.Intent(dec.Intent),
Slots: toDialogueSlots(dec.Slots),
Timestamp: now,
TTL: ttl,
History: history,
})
}
+238
View File
@@ -0,0 +1,238 @@
package main
import (
"context"
"os"
"path/filepath"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/tool"
"github.com/kami/maven/internal/voice"
)
// newClarifyHandler builds a handler with the clarify path wired and no model:
// stub date parser, the real fact parser, and a matcher over whatever tools the
// test enabled. `now` is fixed so TTL behaviour is testable.
func newClarifyHandler(t *testing.T) (*reactiveHandler, *store.Store, *time.Time) {
t.Helper()
st := newTestStore(t)
api := ipc.NewStoreAPI(st)
now := time.Date(2026, 7, 31, 9, 0, 0, 0, time.UTC)
matcher := tool.NewMatcher(api)
h := &reactiveHandler{
api: api,
dataStore: st,
tools: tool.NewExecutor(api, 2*time.Second),
matcher: matcher,
replier: voice.NewStubReplier(),
now: func() time.Time { return now },
dialogueSessions: dialogue.NewSessionStore(2 * time.Minute),
clarifyStore: dialogue.NewClarifyStore(clarifyTTL),
extractor: router.Extractor{
Time: router.StubDateTimeParser{},
Acts: matcher,
Facts: router.DefaultFactParser{},
},
}
return h, st, &now
}
func clarifyDec(intent router.Intent, slots router.Slots, utterance string) router.Decision {
return router.Decision{Utterance: utterance, Stage: 3, Intent: intent, Slots: slots, Clarify: true}
}
// TestClarifyQuestionForMissingSlot pins which question goes with which gap, and
// which intents get no question at all.
func TestClarifyQuestionForMissingSlot(t *testing.T) {
cases := []struct {
name string
dec router.Decision
want string
asked bool
}{
{"reminder without a time", clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"), "На когда напомнить?", true},
{"fact without a key", clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши"), "Что записать?", true},
{"act without a fn", clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это"), "Что сделать?", true},
{"reminder that already has a time", clarifyDec(router.IntentReminder, router.Slots{HasTime: true}, "напомни в 11"), "", false},
{"chat is never worth a question", clarifyDec(router.IntentChat, router.Slots{Text: "мгм"}, "мгм"), "", false},
{"query is never worth a question", clarifyDec(router.IntentQuery, router.Slots{Text: "а"}, "а"), "", false},
}
for _, tc := range cases {
_, got, asked := clarifyQuestion(tc.dec)
if asked != tc.asked || got != tc.want {
t.Errorf("%s: got (%q, %v), want (%q, %v)", tc.name, got, asked, tc.want, tc.asked)
}
}
}
// TestClarifyReminderCompletesOnAnswer is the whole point of the feature: she
// asks for the missing time and the answer creates the reminder.
func TestClarifyReminderCompletesOnAnswer(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
question, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"))
if !asked || question != "На когда напомнить?" {
t.Fatalf("expected the time question, got %q asked=%v", question, asked)
}
reply, handled := h.resolveClarifyAnswer(ctx, "в 11:00")
if !handled {
t.Fatal("the answer to an open question must be consumed as an answer")
}
if reply == clarifyDropped {
t.Fatalf("a good answer must not drop the request: %q", reply)
}
reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour))
if err != nil || len(reminders) != 1 {
t.Fatalf("clarified reminder was not created: reminders=%v err=%v", reminders, err)
}
if !strings.Contains(reminders[0].Payload, "маме") {
t.Fatalf("the reminder lost the original request: %q", reminders[0].Payload)
}
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
t.Fatal("the question must be cleared once answered")
}
}
// TestClarifyFactCompletesOnAnswer — the fact path, where the answer carries
// both the key and the value.
func TestClarifyFactCompletesOnAnswer(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
if _, asked := h.askClarify(clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши")); !asked {
t.Fatal("a fact with no key should be asked about")
}
if reply, handled := h.resolveClarifyAnswer(ctx, "пил воду"); !handled || reply == clarifyDropped {
t.Fatalf("answer should complete the fact, handled=%v reply=%q", handled, reply)
}
if fact, err := st.LatestFact(ctx, "water"); err != nil || fact.Key != "water" {
t.Fatalf("clarified fact was not written: fact=%+v err=%v", fact, err)
}
}
// TestClarifyAnswerAfterTTLIsANewRequest — a late answer is not an answer.
func TestClarifyAnswerAfterTTLIsANewRequest(t *testing.T) {
ctx := context.Background()
h, st, now := newClarifyHandler(t)
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a question")
}
*now = now.Add(clarifyTTL + time.Second)
if reply, handled := h.resolveClarifyAnswer(ctx, "в 11:00"); handled {
t.Fatalf("an answer past the TTL must fall through to normal routing, got %q", reply)
}
if reminders, err := st.DueReminders(ctx, now.Add(48*time.Hour)); err != nil || len(reminders) != 0 {
t.Fatalf("expired question must not create anything: reminders=%v err=%v", reminders, err)
}
}
// TestClarifyUnclearAnswerDropsWithoutAskingAgain — MaxAttempts is 1.
func TestClarifyUnclearAnswerDropsWithoutAskingAgain(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a question")
}
reply, handled := h.resolveClarifyAnswer(ctx, "ну не знаю")
if !handled || reply != clarifyDropped {
t.Fatalf("an unclear answer should drop the request, handled=%v reply=%q", handled, reply)
}
if strings.Contains(reply, "?") {
t.Fatalf("she must not ask a second question: %q", reply)
}
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
t.Fatal("a dropped request must leave no armed question")
}
if reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour)); err != nil || len(reminders) != 0 {
t.Fatalf("a dropped request must not create anything: reminders=%v err=%v", reminders, err)
}
}
// TestClarifiedActOffAllowlistIsStillRefused — clarification fills in an
// argument, it never grants authority.
func TestClarifiedActOffAllowlistIsStillRefused(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
marker := filepath.Join(t.TempDir(), "not-allowed-ran")
if _, asked := h.askClarify(clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это")); !asked {
t.Fatal("an act with no fn should be asked about")
}
reply, handled := h.resolveClarifyAnswer(ctx, "rm "+marker)
if !handled {
t.Fatal("the answer should be consumed")
}
if strings.Contains(reply, "готово") {
t.Fatalf("an act that is not on the allowlist must not report success: %q", reply)
}
if _, err := os.Stat(marker); !os.IsNotExist(err) {
t.Fatalf("a clarified act off the allowlist ran anyway: %v", err)
}
if tools, err := st.ListTools(ctx, "enabled"); err != nil || len(tools) != 0 {
t.Fatalf("clarify must not enable a tool: tools=%+v err=%v", tools, err)
}
}
// TestClarifiedDestructiveActStillNeedsConfirm — the confirm gate survives the
// clarify path.
func TestClarifiedDestructiveActStillNeedsConfirm(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
marker := filepath.Join(t.TempDir(), "destructive-ran")
if err := st.EnableTool(ctx, "delete_backups", []string{"touch", marker}, true, "test", h.now()); err != nil {
t.Fatal(err)
}
if _, asked := h.askClarify(clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это")); !asked {
t.Fatal("expected a question")
}
reply, handled := h.resolveClarifyAnswer(ctx, "delete_backups")
if !handled {
t.Fatal("the answer should be consumed")
}
if !strings.Contains(reply, "да") || h.pending == nil {
t.Fatalf("a clarified destructive act must still park a confirm: reply=%q pending=%+v", reply, h.pending)
}
if _, err := os.Stat(marker); !os.IsNotExist(err) {
t.Fatalf("a clarified destructive act ran before confirmation: %v", err)
}
}
// TestNoQuestionWhenNothingIsMissing — noise keeps the canned reply, so she
// never invents a question for nothing.
func TestNoQuestionWhenNothingIsMissing(t *testing.T) {
h, _, _ := newClarifyHandler(t)
for _, dec := range []router.Decision{
clarifyDec(router.IntentChat, router.Slots{Text: "эм"}, "эм"),
clarifyDec(router.IntentQuery, router.Slots{Text: "ммм"}, "ммм"),
clarifyDec(router.IntentNote, router.Slots{Text: "..."}, "..."),
} {
if question, asked := h.askClarify(dec); asked {
t.Fatalf("intent %s should keep the canned reply, got %q", dec.Intent, question)
}
}
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
t.Fatal("noise must not park a question")
}
}
// TestNoPendingQuestionFallsThrough — with nothing parked, an utterance routes
// normally.
func TestNoPendingQuestionFallsThrough(t *testing.T) {
h, _, _ := newClarifyHandler(t)
if reply, handled := h.resolveClarifyAnswer(context.Background(), "напомни в 11:00"); handled {
t.Fatalf("no open question ⇒ must not be treated as an answer, got %q", reply)
}
}
+28
View File
@@ -0,0 +1,28 @@
package main
import (
"testing"
"time"
"github.com/kami/maven/internal/llm"
)
func TestPickLLMRouterOff(t *testing.T) {
if r := pickLLMRouter(false, llm.New("http://127.0.0.1:1", time.Second)); r != nil {
t.Error("flag off should give no LLM router")
}
}
// The operator can turn the flag on without an LLM phraser configured. That must
// leave the classifier running, not panic.
func TestPickLLMRouterOnWithoutClient(t *testing.T) {
if r := pickLLMRouter(true, nil); r != nil {
t.Error("no llama-server should give no LLM router")
}
}
func TestPickLLMRouterOn(t *testing.T) {
if r := pickLLMRouter(true, llm.New("http://127.0.0.1:1", time.Second)); r == nil {
t.Error("flag on with a client should give an LLM router")
}
}
+14 -11
View File
@@ -57,10 +57,10 @@ import (
"github.com/kami/maven/internal/delivery/ntfysink"
"github.com/kami/maven/internal/delivery/telegramsink"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/webauthn"
"github.com/kami/maven/internal/loop"
)
var errLocked = errors.New("mavend: daemon locked — complete passkey assertion first")
@@ -161,11 +161,14 @@ func (l *lockedAPI) EnableTool(ctx context.Context, name string, cmd []string, d
return errLocked
}
func (l *lockedAPI) DisableTool(ctx context.Context, name string) error { return errLocked }
func (l *lockedAPI) DeleteTool(ctx context.Context, name string) error { return errLocked }
func (l *lockedAPI) DeleteTool(ctx context.Context, name string) error { return errLocked }
func (l *lockedAPI) ListProposedRoutines(ctx context.Context) ([]ipc.ProposedRoutine, error) {
return nil, errLocked
}
func (l *lockedAPI) DismissProposedRoutine(ctx context.Context, id int64) error { return errLocked }
func (l *lockedAPI) AcceptProposedRoutine(ctx context.Context, id int64) error {
return errLocked
}
func (l *lockedAPI) LookupTool(ctx context.Context, name string) (ipc.Tool, error) {
return ipc.Tool{}, errLocked
}
@@ -244,15 +247,15 @@ func run(args []string) error {
// ----- daemon components (only wired when unlocked) -----
// Pre-declare so the unlock path can wire them later.
var (
gatherer *loop.Gatherer
rules []loop.Rule
phr phraser.Phraser
voiceW *voiceWiring
dispatcher *delivery.Dispatcher
tl *tickLoop
coreAPI ipc.CoreAPI
eco *ecosystemWiring
factWorker *factEnrichmentWorker
gatherer *loop.Gatherer
rules []loop.Rule
phr phraser.Phraser
voiceW *voiceWiring
dispatcher *delivery.Dispatcher
tl *tickLoop
coreAPI ipc.CoreAPI
eco *ecosystemWiring
factWorker *factEnrichmentWorker
)
if !locked {
+5 -8
View File
@@ -8,16 +8,13 @@ import "github.com/kami/maven/internal/memory"
// that can answer "when did I last …?" from a captured fact). A note hit here
// is redundant with the notes-RAG path — by design; the two indexes can diverge
// once the backend is swapped for a persistent/external store. ok=false when
// there's no hit above the threshold or the hit carries no text.
func bestRecall(results []memory.Result, min float64) (string, bool) {
if len(results) == 0 {
// the hit fails the confidence gate (see memory.Confident: an absolute floor
// plus a margin over the runner-up) or carries no text.
func bestRecall(results []memory.Result, minScore, minMargin float64) (string, bool) {
if !memory.Confident(results, minScore, minMargin) {
return "", false
}
top := results[0]
if top.Score < min {
return "", false
}
text := top.Meta["text"]
text := results[0].Meta["text"]
if text == "" {
return "", false
}
+17 -4
View File
@@ -8,23 +8,24 @@ import (
func TestBestRecall(t *testing.T) {
const min = 0.55
const margin = 0.008
t.Run("empty results", func(t *testing.T) {
if _, ok := bestRecall(nil, min); ok {
if _, ok := bestRecall(nil, min, margin); ok {
t.Error("empty results returned ok")
}
})
t.Run("top below threshold", func(t *testing.T) {
res := []memory.Result{{Score: 0.4, Meta: map[string]string{"text": "выпил воды"}}}
if _, ok := bestRecall(res, min); ok {
if _, ok := bestRecall(res, min, margin); ok {
t.Error("below-threshold hit returned ok")
}
})
t.Run("hit without text meta", func(t *testing.T) {
res := []memory.Result{{Score: 0.9, Meta: map[string]string{"type": "fact"}}}
if _, ok := bestRecall(res, min); ok {
if _, ok := bestRecall(res, min, margin); ok {
t.Error("textless hit returned ok")
}
})
@@ -34,7 +35,7 @@ func TestBestRecall(t *testing.T) {
{Score: 0.82, Meta: map[string]string{"text": "выпил воды в три часа", "type": "fact"}},
{Score: 0.60, Meta: map[string]string{"text": "другое"}},
}
got, ok := bestRecall(res, min)
got, ok := bestRecall(res, min, margin)
if !ok {
t.Fatal("clearing hit not returned")
}
@@ -42,4 +43,16 @@ func TestBestRecall(t *testing.T) {
t.Errorf("wrong text: %q", got)
}
})
// The runner-up is almost as close, so the embedder cannot tell the two
// notes apart. Silence beats reading back a coin flip.
t.Run("runner-up too close", func(t *testing.T) {
res := []memory.Result{
{Score: 0.860, Meta: map[string]string{"text": "выпил воды в три часа"}},
{Score: 0.858, Meta: map[string]string{"text": "другое"}},
}
if _, ok := bestRecall(res, min, margin); ok {
t.Error("thin-margin hit returned ok")
}
})
}
+4 -1
View File
@@ -9,7 +9,10 @@ import (
"github.com/kami/maven/internal/voice"
)
type mockCompleter struct{ out string; err error }
type mockCompleter struct {
out string
err error
}
func (m mockCompleter) Complete(_ context.Context, _ llm.Req) (string, error) { return m.out, m.err }
+62
View File
@@ -186,6 +186,10 @@ func (t *tickLoop) tick(ctx context.Context, now time.Time) {
// LLM-phrased — so a routine can't hallucinate. severity comes from config.
t.fireRoutines(ctx, now, state)
// accepted routines: patterns the user confirmed. read straight from the
// store each tick so the schedule survives a restart.
t.fireAcceptedRoutines(ctx, now, state)
// morning routines: daily checklists (medicine/water/pets/...), nagged at
// most once per day per routine, and only for items still unevidenced at
// nudge time. See internal/morning for the "why not four timers" rationale.
@@ -377,6 +381,64 @@ func (t *tickLoop) fireRoutines(ctx context.Context, now time.Time, state loop.S
}
}
// fireAcceptedRoutines nudges about the routines the user accepted, once per
// interval (Vikunja #366). Accepting used to create a single reminder, so a
// non-weekly routine fired once and went quiet forever; the schedule lives in
// the proposed_routines row now and the loop re-reads it every tick.
//
// A routine is a care-class nudge and goes through the restraint gate like any
// other: quiet hours, away presence and snooze all suppress it. Reminders bypass
// that gate; routines must not. A suppressed nudge is NOT marked fired, so it
// goes out on the next tick that the gate allows — one nudge, held, not dropped
// and not repeated.
//
// The body is literal text built from the detected action and object, not
// LLM-phrased, so a routine can't hallucinate. It nudges; it never acts.
func (t *tickLoop) fireAcceptedRoutines(ctx context.Context, now time.Time, state loop.State) {
rows, err := t.store.ListAcceptedRoutines(ctx)
if err != nil {
log.Printf("tick: list accepted routines: %v", err)
return
}
accepted := make([]routine.Accepted, 0, len(rows))
for _, r := range rows {
if r.AcceptedTs == nil {
continue // accepted before the schedule column existed — no clock to start from.
}
accepted = append(accepted, routine.Accepted{
ID: r.ID,
Name: r.Action + " " + r.Object,
IntervalDays: r.IntervalDays,
Accepted: *r.AcceptedTs,
LastFired: r.LastFiredTs,
})
}
for _, a := range routine.DueAccepted(accepted, now) {
rule := loop.Rule{Name: "routine:" + a.Name, Severity: loop.Sev1}
if !loop.Gate(state, rule) {
continue
}
body := "пора: " + a.Name
pn := delivery.PhrasedNudge{
Candidate: loop.Candidate{Rule: rule, Severity: rule.Severity, State: state},
Body: body,
Summary: body,
}
sent, err := t.dispatcher.DispatchNudge(ctx, pn, now)
if err != nil {
log.Printf("tick: dispatch accepted routine %d: %v", a.ID, err)
continue
}
if len(sent) == 0 {
continue // routing dropped it — leave it due.
}
if err := t.store.MarkRoutineFired(ctx, a.ID, now); err != nil {
log.Printf("tick: mark routine %d fired: %v", a.ID, err)
}
}
}
// fireMorningRoutines checks each configured checklist against today's facts
// and dispatches a nag listing exactly what's still missing, at most once per
// routine per calendar day. Fact reads happen here (not in loop.Gatherer)
+109
View File
@@ -96,6 +96,115 @@ func TestTickFiresRoutineWhenScheduleCrosses(t *testing.T) {
}
}
// TestTickFiresAcceptedRoutineEveryInterval — Vikunja #366. An accepted routine
// with a 3-day interval must nudge every 3 days, not once. It also must not
// replay the occurrences it slept through: after a 30-day gap it nudges once.
func TestTickFiresAcceptedRoutineEveryInterval(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
accepted := refNow()
id, err := st.CreateProposedRoutine(ctx, "полить", "цветы", 3.0, accepted)
if err != nil {
t.Fatalf("CreateProposedRoutine: %v", err)
}
if err := st.AcceptProposedRoutine(ctx, id, accepted); err != nil {
t.Fatalf("AcceptProposedRoutine: %v", err)
}
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
const rule = "routine:полить цветы"
// Same day as the accept: not due yet.
markPresent(t, st, ctx, accepted)
tl.tick(ctx, accepted.Add(time.Hour))
if n := countSends(sink, rule); n != 0 {
t.Fatalf("routine fired %d times before its first interval passed, want 0", n)
}
// Three days later: the first nudge.
first := accepted.Add(3 * 24 * time.Hour)
markPresent(t, st, ctx, first)
tl.tick(ctx, first)
if n := countSends(sink, rule); n != 1 {
t.Fatalf("first interval: sends = %d, want 1", n)
}
// Next day: still inside the interval, silent.
sink.sends = nil
markPresent(t, st, ctx, first.Add(24*time.Hour))
tl.tick(ctx, first.Add(24*time.Hour))
if n := countSends(sink, rule); n != 0 {
t.Fatalf("mid-interval: sends = %d, want 0", n)
}
// Three days after the first nudge: it fires again. This is the bug —
// a one-shot reminder would never come back.
second := first.Add(3 * 24 * time.Hour)
markPresent(t, st, ctx, second)
tl.tick(ctx, second)
if n := countSends(sink, rule); n != 1 {
t.Fatalf("second interval: sends = %d, want 1 (a routine repeats)", n)
}
// A long silence must not turn into a backlog of missed nudges.
sink.sends = nil
late := second.Add(30 * 24 * time.Hour)
markPresent(t, st, ctx, late)
tl.tick(ctx, late)
if n := countSends(sink, rule); n != 1 {
t.Fatalf("after a 30-day gap: sends = %d, want exactly 1 (no backlog)", n)
}
}
// TestTickAcceptedRoutineRespectsQuietHours — routines are not reminders: they
// do not inherit the reminder gate bypass. Away presence drops a care-class
// nudge, and the routine stays due so it nudges once the user is back.
func TestTickAcceptedRoutineRespectsGate(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
accepted := refNow()
id, err := st.CreateProposedRoutine(ctx, "полить", "цветы", 3.0, accepted)
if err != nil {
t.Fatalf("CreateProposedRoutine: %v", err)
}
if err := st.AcceptProposedRoutine(ctx, id, accepted); err != nil {
t.Fatalf("AcceptProposedRoutine: %v", err)
}
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
const rule = "routine:полить цветы"
// No presence probes at all ⇒ away ⇒ the care gate blocks the nudge.
due := accepted.Add(3 * 24 * time.Hour)
tl.tick(ctx, due)
if n := countSends(sink, rule); n != 0 {
t.Fatalf("away: sends = %d, want 0 (routine must not bypass the gate)", n)
}
// Back at the desk a minute later: the nudge that was held now goes out.
back := due.Add(time.Minute)
markPresent(t, st, ctx, back)
tl.tick(ctx, back)
if n := countSends(sink, rule); n != 1 {
t.Fatalf("present again: sends = %d, want 1", n)
}
}
// countSends counts captured sends for one rule name.
func countSends(sink *fakeSink, rule string) int {
n := 0
for _, s := range sink.sends {
if s.RuleName == rule {
n++
}
}
return n
}
// refNow — fixed tick time so presence decay + since durations are deterministic.
func refNow() time.Time { return time.Date(2026, 6, 30, 12, 0, 0, 0, time.UTC) }
+87 -82
View File
@@ -199,8 +199,6 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
if lp, ok := phr.(*phraser.LLMPhraser); ok {
llmClient = llm.New(lp.BaseURL(), 60*time.Second)
}
// LLM router disabled — the classifier handles routing reliably.
// ----- router (the cascade; floor examples seed the classifier) -----
// The act matcher's allowlist is exactly the enabled tool names — the
// router only matches acts the executor can run (one source of truth).
@@ -208,7 +206,11 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
if threshold <= 0 {
threshold = config.DefaultRouterThreshold
}
rtr := buildRouter(emb, matcher, threshold, nil) // LLM router disabled
// The resident model routes by default: 63.2% of held-out intents right
// against the classifier's 50.0%, at about 1s a turn instead of 30ms (see
// config.VoiceConfig.LLMRouter). The classifier always stays wired as the
// fallback, so a model error never breaks a turn.
rtr := buildRouter(emb, matcher, threshold, pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient))
// ----- sessions registry (shared with voicesink) -----
sessions := voice.NewSessions()
@@ -226,6 +228,8 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
// ----- dialogue (multi-turn slot carry-over; 2-min follow-up window) -----
dialogueSessions := dialogue.NewSessionStore(2 * time.Minute)
clarifyStore := dialogue.NewClarifyStore(clarifyTTL)
timeParser := router.NewPythonDateParser()
// ----- replier (LLM-backed when the engine is on, Stub floor otherwise) -----
replier := voice.Replier(voice.NewStubReplier())
@@ -250,8 +254,11 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
memStore: memStore,
dataStore: dataStore,
dialogueSessions: dialogueSessions,
clarifyStore: clarifyStore,
extractor: router.Extractor{Time: timeParser, Acts: matcher, Facts: router.DefaultFactParser{}},
queryMinScore: cfg.Voice.QueryMinScore,
timeParser: router.NewPythonDateParser(),
queryMinMargin: cfg.Voice.QueryMinMargin,
timeParser: timeParser,
ecosystem: eco,
}
@@ -294,6 +301,9 @@ type reactiveHandler struct {
// load-bearing math (same posture as the presence thresholds). Set by
// wireVoice from VoiceConfig; default 0.55.
queryMinScore float64
// queryMinMargin — the second half of that gate: how far the top hit must
// beat the runner-up. 0 ⇒ margin off.
queryMinMargin float64
// timeParser — used as a fallback for stage-0 reminder grammar matches
// (where the extractor didn't run). Shared with the router's extractor.
@@ -304,6 +314,14 @@ type reactiveHandler struct {
// box → one session slot, keyed voiceDialogueID). nil ⇒ no carry-over.
dialogueSessions *dialogue.SessionStore
// clarifyStore parks the request behind an open question she asked (see
// clarify.go). nil ⇒ she falls back to the canned "не поняла" reply.
clarifyStore *dialogue.ClarifyStore
// extractor parses the answer to an open question, with the same parsers
// the router's own stage-2 uses.
extractor router.Extractor
// pending destructive-act confirmation. A destructive act replies with a
// "выполнить X? да/нет" prompt and parks here; the NEXT utterance is read as
// the y/n answer. ponytail: single slot, single-user box — a second act
@@ -378,6 +396,13 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
return h.reply(ctx, reply, nil)
}
// 1b2. clarify answer — if she asked a question last turn, this utterance is
// its answer, not a fresh command. After the confirm check: a y/n gate is
// armed by her own prompt and is the narrower claim on the utterance.
if reply, handled := h.resolveClarifyAnswer(ctx, text); handled {
return h.reply(ctx, reply, nil)
}
// 1c. quiet-hours toggle — keyword match, not classifier-dependent.
// "тихий режим" / "quiet on" would route through the classifier
// unreliably (it's a command, not a free-form query), so we match it
@@ -407,34 +432,16 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
prev := h.dialogueSessions.Get(voiceDialogueID, now)
dec = followUpMerge(prev, dec, now)
if !dec.Clarify {
// Build history: carry over up to 4 prior turns for cross-intent
// reference. The most recent prior turn is prepended to history.
var history []dialogue.Turn
if prev != nil {
history = append(history, dialogue.Turn{
Intent: prev.Intent,
Slots: prev.Slots,
Text: prev.Slots.Text, // the prior turn's utterance
})
// Cap history depth so one long conversation can't grow
// the session unboundedly.
maxHist := len(prev.History)
if maxHist > 3 {
maxHist = 3
}
history = append(history, prev.History[:maxHist]...)
}
ttl := time.Duration(0) // use default (2 min)
if dec.Intent == router.IntentChat {
ttl = 15 * time.Minute // conversational turns should last longer
}
h.dialogueSessions.Put(voiceDialogueID, &dialogue.Session{
Intent: dialogue.Intent(dec.Intent),
Slots: toDialogueSlots(dec.Slots),
Timestamp: now,
TTL: ttl,
History: history,
})
h.rememberTurn(prev, dec, now)
}
}
// 2c. clarify — she is not sure. If one named thing is missing, ask about it
// and park the request (clarify.go); otherwise the replier's canned reply
// stands.
if dec.Clarify {
if question, asked := h.askClarify(dec); asked {
return h.reply(ctx, question, nil)
}
}
@@ -465,6 +472,11 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
return reply
}
// 1b2. clarify answer — same check as HandlePushToTalk.
if reply, handled := h.resolveClarifyAnswer(ctx, text); handled {
return reply
}
// 2. router — classify the utterance.
dec, err := h.router.Route(ctx, text, h.now())
if err != nil {
@@ -482,30 +494,14 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
prev := h.dialogueSessions.Get(voiceDialogueID, now)
dec = followUpMerge(prev, dec, now)
if !dec.Clarify {
var history []dialogue.Turn
if prev != nil {
history = append(history, dialogue.Turn{
Intent: prev.Intent,
Slots: prev.Slots,
Text: prev.Slots.Text,
})
maxHist := len(prev.History)
if maxHist > 3 {
maxHist = 3
}
history = append(history, prev.History[:maxHist]...)
}
ttl := time.Duration(0)
if dec.Intent == router.IntentChat {
ttl = 15 * time.Minute
}
h.dialogueSessions.Put(voiceDialogueID, &dialogue.Session{
Intent: dialogue.Intent(dec.Intent),
Slots: toDialogueSlots(dec.Slots),
Timestamp: now,
TTL: ttl,
History: history,
})
h.rememberTurn(prev, dec, now)
}
}
// 2c. clarify — same as HandlePushToTalk: ask about the one missing thing.
if dec.Clarify {
if question, asked := h.askClarify(dec); asked {
return question
}
}
@@ -567,7 +563,7 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
// fail the fact write). Facts aren't in the notes table, so this is the
// only recall path for them — "когда я пил воду?" reads back from here.
if h.memStore != nil {
if vec, err := h.embedder.Embed(ctx, dec.Utterance); err != nil {
if vec, err := router.EmbedPassage(ctx, h.embedder, dec.Utterance); err != nil {
log.Printf("voice: embed fact for memory: %v", err)
} else if err := h.memStore.Insert(ctx, "fact:"+dec.Slots.Key+":"+strconv.FormatInt(now.Unix(), 10), vec, map[string]string{
"source": "voice",
@@ -681,7 +677,7 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
// embed the note text with the same model the classifier uses, persist
// via CoreAPI (source=tap:voice). Semantic recall lives in `notes`, not
// facts — no predicate reads it (spec's two-memory split).
vec, err := h.embedder.Embed(ctx, dec.Utterance)
vec, err := router.EmbedPassage(ctx, h.embedder, dec.Utterance)
if err != nil {
log.Printf("voice: embed note: %v", err)
return "не получилось сохранить заметку."
@@ -758,7 +754,7 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
return fmt.Sprintf("в %s сейчас %.0f градусов, %s.", w.Location, w.Temperature, w.Condition)
}
vec, err := h.embedder.Embed(ctx, dec.Utterance)
vec, err := router.EmbedQuery(ctx, h.embedder, dec.Utterance)
if err != nil {
log.Printf("voice: embed query: %v", err)
return "не получилось найти ответ."
@@ -768,18 +764,23 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
log.Printf("voice: query notes: %v", err)
return "не получилось найти ответ."
}
// Confidence gate: below threshold, say "I don't know" rather than read
// back the least-unrelated note — a confident wrong recall is worse than
// a gap (spec's "not a guesser-of-truth"). Same instinct as the loop's
// since(key)==null → don't fire. Tuned for the ONNX embedder; the Hash
// floor scores lexically and may rarely clear it.
if len(notes) == 0 || notes[0].Score < h.queryMinScore {
// Confidence gate: below it, say "I don't know" rather than read back
// the least-unrelated note — a confident wrong recall is worse than a
// gap (spec's "not a guesser-of-truth"). Same instinct as the loop's
// since(key)==null → don't fire. Two parts: an absolute cosine floor,
// and a margin over the runner-up, which is the part that works with
// the e5 embedder's narrow score band. See memory.Confident.
noteScores := make([]float64, len(notes))
for i, n := range notes {
noteScores[i] = n.Score
}
if !memory.ConfidentScores(noteScores, h.queryMinScore, h.queryMinMargin) {
// Long-term memory recall (notes + facts) before general knowledge:
// the notes table can't answer fact questions, but the memory store
// indexes both. Only runs when notes-RAG already gave up → additive.
if h.memStore != nil {
if hits, herr := h.memStore.Search(ctx, vec, 3); herr == nil {
if text, ok := bestRecall(hits, h.queryMinScore); ok {
if text, ok := bestRecall(hits, h.queryMinScore, h.queryMinMargin); ok {
return text
}
}
@@ -1048,6 +1049,21 @@ func (h *reactiveHandler) reply(ctx context.Context, text string, _ []string) (v
return voice.PushToTalkResp{ReplyText: text, ReplyAudio: audioOut}, nil
}
// pickLLMRouter returns the LLM router when the operator asked for it and there
// is a llama-server to talk to, and nil otherwise. nil is safe: the cascade then
// routes with the classifier, so an unusable setting costs accuracy, not turns.
func pickLLMRouter(enabled bool, c *llm.Client) *router.LLMRouter {
if !enabled {
return nil
}
if c == nil {
log.Printf("voice: voice.llm_router is on but there is no llama-server to route with (the phraser is not an LLM phraser) — using the classifier instead")
return nil
}
log.Printf("voice: LLM router enabled")
return router.NewLLMRouter(c)
}
// buildRouter constructs the reactive-path router with the given embedder
// and confidence threshold.
// - stage-0 grammars from DefaultActMatcher whose fn allowlist is exactly
@@ -1414,23 +1430,12 @@ func (h *reactiveHandler) resolveConfirm(ctx context.Context, text string) (stri
switch classifyConfirm(text) {
case confirmYes:
h.pendingRoutine = nil
// Create a recurring reminder at the detected interval.
// Weekly patterns get a cron expression; arbitrary intervals
// fire once and the detector re-proposes on the next cycle.
intervalDur := time.Duration(pr.interval * 24 * float64(time.Hour))
fire := h.now().Add(intervalDur)
cron := ""
if pr.interval >= 6.5 && pr.interval <= 7.5 {
cron = fmt.Sprintf("0 %d * * %d", fire.Hour(), int(fire.Weekday()))
}
payload := fmt.Sprintf(`{"text":"%s %s"}`, pr.action, pr.object)
remID, err := h.api.CreateReminder(ctx, fire, payload, cron)
if err != nil {
log.Printf("voice: create routine reminder: %v", err)
return "не получилось поставить напоминание.", true
}
if err := h.dataStore.AcceptProposedRoutine(ctx, pr.routineID, remID); err != nil {
// Only record the acceptance. The tick loop reads accepted
// routines and nudges on their own interval. Building a reminder
// here made a routine fire exactly once (Vikunja #366).
if err := h.dataStore.AcceptProposedRoutine(ctx, pr.routineID, h.now()); err != nil {
log.Printf("voice: accept proposed routine: %v", err)
return "не получилось запомнить рутину.", true
}
return "буду напоминать.", true
case confirmNo:
+4 -1
View File
@@ -91,7 +91,10 @@ func handleEcosystem(w http.ResponseWriter, r *http.Request, urls ecoURLs) {
var d ecoData
var wg sync.WaitGroup
wg.Add(3)
go func() { defer wg.Done(); d.Nexus.Err = getEco(ctx, urls.nexus, "/api/v1/entities?limit=50", &d.Nexus.Rows) }()
go func() {
defer wg.Done()
d.Nexus.Err = getEco(ctx, urls.nexus, "/api/v1/entities?limit=50", &d.Nexus.Rows)
}()
go func() {
defer wg.Done()
d.Praxis.Err = getEco(ctx, urls.praxis, "/api/v1/items?limit=50", &d.Praxis.Rows)
+126
View File
@@ -918,3 +918,129 @@ func TestHandleTools_ListToolsError_502(t *testing.T) {
t.Fatalf("status = %d, want 502; body=%s", rr.Code, rr.Body.String())
}
}
// --- handleRoutines ---
// routineCore is a fakeCore that also answers the proposed-routine calls.
type routineCore struct {
fakeCore
routines []ipc.ProposedRoutine
dismissed int64
acceptedID int64
remCron string
}
func (c *routineCore) ListProposedRoutines(_ context.Context) ([]ipc.ProposedRoutine, error) {
return c.routines, nil
}
func (c *routineCore) DismissProposedRoutine(_ context.Context, id int64) error {
c.dismissed = id
return nil
}
func (c *routineCore) AcceptProposedRoutine(_ context.Context, id int64) error {
c.acceptedID = id
return nil
}
func (c *routineCore) CreateReminder(_ context.Context, _ time.Time, _, cron string) (int64, error) {
c.remCron = cron
return 77, nil
}
func weeklyRoutineCore() *routineCore {
return &routineCore{routines: []ipc.ProposedRoutine{{
ID: 3, Action: "refill", Object: "cat_water", IntervalDays: 7,
Status: "proposed", CreatedTs: time.Now().Add(-2 * time.Hour).UnixMilli(),
}}}
}
func postRoutine(action, id string) *http.Request {
return postForm(action, url.Values{"id": {id}})
}
func TestHandleRoutines_GET_ShowsMavensPhrase(t *testing.T) {
rr := httptest.NewRecorder()
handleRoutines(rr, httptest.NewRequest(http.MethodGet, "/routines", nil), weeklyRoutineCore(), nil, false)
if rr.Code != http.StatusOK {
t.Fatalf("status = %d, want 200", rr.Code)
}
body := rr.Body.String()
if !strings.Contains(body, "заправляешь") {
t.Fatalf("want maven's phrasing in the page, got: %s", body)
}
if !strings.Contains(body, "class=scroll") {
t.Fatal("table must be wrapped in <div class=scroll> so it pans on a phone")
}
}
// Accepting hands the loop a new reason to speak, so it needs step-up.
func TestHandleRoutines_Accept_RequiresStepUp(t *testing.T) {
core := weeklyRoutineCore()
rr := httptest.NewRecorder()
handleRoutines(rr, postRoutine("accept", "3"), core, webauthn.NewPasskeySession(5*time.Minute), false)
if rr.Code != http.StatusForbidden {
t.Fatalf("status = %d, want 403", rr.Code)
}
if core.acceptedID != 0 {
t.Fatal("accepted without step-up")
}
}
// Accepting only flips the status. It used to also create a one-shot reminder,
// which is why a non-weekly routine fired once and then went quiet forever
// (Vikunja #366). The tick loop owns the schedule now, so a reminder here would
// be a second, competing schedule.
func TestHandleRoutines_Accept_FlipsStatusAndMakesNoReminder(t *testing.T) {
core := weeklyRoutineCore()
rr := httptest.NewRecorder()
handleRoutines(rr, postRoutine("accept", "3"), core, stepUpSession(), false)
if rr.Code != http.StatusOK {
t.Fatalf("status = %d, want 200; body=%s", rr.Code, rr.Body.String())
}
if core.acceptedID != 3 {
t.Fatalf("accepted id = %d, want 3", core.acceptedID)
}
if core.remCron != "" {
t.Fatalf("accepting must not create a reminder, got cron %q", core.remCron)
}
}
// Dismiss only ever removes a reason to speak, so it is not step-up gated.
func TestHandleRoutines_Dismiss_NoStepUpNeeded(t *testing.T) {
core := weeklyRoutineCore()
rr := httptest.NewRecorder()
handleRoutines(rr, postRoutine("dismiss", "3"), core, webauthn.NewPasskeySession(5*time.Minute), false)
if rr.Code != http.StatusOK {
t.Fatalf("status = %d, want 200; body=%s", rr.Code, rr.Body.String())
}
if core.dismissed != 3 {
t.Fatalf("dismissed = %d, want 3", core.dismissed)
}
}
func TestHandleRoutines_UnknownAction_400(t *testing.T) {
rr := httptest.NewRecorder()
handleRoutines(rr, postRoutine("frobnicate", "3"), weeklyRoutineCore(), stepUpSession(), false)
if rr.Code != http.StatusBadRequest {
t.Fatalf("status = %d, want 400", rr.Code)
}
}
func TestHandleRoutines_BadID_400(t *testing.T) {
rr := httptest.NewRecorder()
handleRoutines(rr, postRoutine("dismiss", "nope"), weeklyRoutineCore(), stepUpSession(), false)
if rr.Code != http.StatusBadRequest {
t.Fatalf("status = %d, want 400", rr.Code)
}
}
func TestHandleRoutines_NilCore_503(t *testing.T) {
rr := httptest.NewRecorder()
handleRoutines(rr, httptest.NewRequest(http.MethodGet, "/routines", nil), nil, nil, false)
if rr.Code != http.StatusServiceUnavailable {
t.Fatalf("status = %d, want 503", rr.Code)
}
}
+92 -15
View File
@@ -24,6 +24,7 @@ import (
"github.com/coder/websocket"
"github.com/kami/maven/internal/audio"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/pattern"
"github.com/kami/maven/internal/voice"
"github.com/kami/maven/internal/webauthn"
)
@@ -395,9 +396,6 @@ func main() {
mux.HandleFunc("/reminders", func(w http.ResponseWriter, r *http.Request) {
handleReminders(w, r, core)
})
mux.HandleFunc("/routines", func(w http.ResponseWriter, r *http.Request) {
handleRoutines(w, r, core)
})
mux.HandleFunc("/morning", func(w http.ResponseWriter, r *http.Request) {
handleMorning(w, r, core)
})
@@ -450,6 +448,11 @@ func main() {
mux.HandleFunc("/tools", func(w http.ResponseWriter, r *http.Request) {
handleTools(w, r, core, stepUpSession, *requireStepUp)
})
// /routines — the authed accept surface. Registered here, next to /tools,
// because accepting shares the same step-up gate.
mux.HandleFunc("/routines", func(w http.ResponseWriter, r *http.Request) {
handleRoutines(w, r, core, stepUpSession, *requireStepUp)
})
// /api/revert voids the latest fact for a key — a store mutation, so it
// sits behind the same passkey step-up as tool enable (nil session ⇒
@@ -680,22 +683,24 @@ const toolsHTML = `{{template "shellTop" "tools"}}
</section>
{{template "shellBottom"}}`
// routinesHTML — proposed routine review surface. Lists detected patterns
// awaiting human confirmation, with accept (→ reminder) and dismiss buttons.
// routinesHTML — proposed routine review surface. One row per thing maven
// noticed, in her words, with at most two actions: accept or dismiss.
const routinesHTML = `{{template "shellTop" "routines"}}
<h1>Routines</h1>
{{if .Msg}}<div class="msg msg-ok">{{.Msg}}</div>{{end}}
<section class=card>
<h2 class=card-title>proposed <span class=badge>{{len .Proposed}}</span></h2>
{{if .Proposed}}<div class=scroll><table><tr><th>action</th><th>object</th><th>every</th><th></th></tr>
<h2 class=card-title>noticed <span class=badge>{{len .Proposed}}</span></h2>
{{if .Proposed}}<div class=scroll><table><tr><th>maven noticed</th><th>when</th><th></th><th></th></tr>
{{range .Proposed}}<tr>
<td><code>{{.Action}}</code></td><td><code>{{.Object}}</code></td><td>{{.IntervalDays}} days</td>
<td>
<form method=post action=/routines class=inline-form>
<td>{{.Phrase}}</td><td class=muted>{{.Noticed}}</td>
<td><form method=post action=/routines class=inline-form>
<input type=hidden name=id value="{{.ID}}">
<input type=hidden name=action value=accept>
<button class=btn>accept</button></form></td>
<td><form method=post action=/routines class=inline-form>
<input type=hidden name=id value="{{.ID}}">
<input type=hidden name=action value=dismiss>
<button class="btn btn-muted">dismiss</button></form>
</td>
<button class="btn btn-muted">dismiss</button></form></td>
</tr>{{end}}</table></div>
{{else}}<div class=empty>
<svg class=icon width="20" height="20"><use href="/ethos-icons.svg#i-wave"/></svg>
@@ -785,7 +790,23 @@ func handleReminders(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
}
}
func handleRoutines(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
// routineRow is one line on the page: what maven noticed, in her words, and
// how long ago she noticed it.
type routineRow struct {
ID int64
Phrase string
Noticed string
}
// handleRoutines serves the routine review surface (GET) and answers a
// proposal (POST id + action=accept|dismiss).
//
// Accept is gated at step-up, the same tier as enabling a tool: saying yes
// hands the trigger loop a new standing reason to speak to the human, so it
// moves the boundary and only an authed surface may do it. Dismiss is not
// gated — it only ever removes a reason to speak, so the worst a weaker caller
// can do is make maven quieter.
func handleRoutines(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI, session *webauthn.PasskeySession, requireStepUp bool) {
if core == nil {
http.Error(w, "routines disabled (no -core)", http.StatusServiceUnavailable)
return
@@ -801,6 +822,17 @@ func handleRoutines(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
return
}
switch action {
case "accept":
if !stepUpOK(session, requireStepUp) {
http.Error(w, "step-up required: assert a passkey first", http.StatusForbidden)
return
}
if err := acceptRoutine(ctx, core, rid); err != nil {
log.Printf("routines: accept %d: %v", rid, err)
http.Error(w, "accept failed: "+err.Error(), http.StatusBadGateway)
return
}
msg = "accepted routine — maven will remind you"
case "dismiss":
if err := core.DismissProposedRoutine(ctx, rid); err != nil {
log.Printf("routines: dismiss %d: %v", rid, err)
@@ -822,12 +854,57 @@ func handleRoutines(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
w.Header().Set("Content-Type", "text/html; charset=utf-8")
if err := routinesTmpl.Execute(w, struct {
Msg string
Proposed []ipc.ProposedRoutine
}{msg, proposed}); err != nil {
Proposed []routineRow
}{msg, routineRows(proposed)}); err != nil {
log.Printf("routines render: %v", err)
}
}
// routineRows turns the wire rows into display rows. The phrase comes from
// pattern.PhraseRoutine so the page says the same thing maven's voice says.
func routineRows(rs []ipc.ProposedRoutine) []routineRow {
out := make([]routineRow, 0, len(rs))
for _, r := range rs {
p := pattern.ProposedRoutine{Action: r.Action, Object: r.Object, IntervalDays: r.IntervalDays}
noticed := "just now"
if r.CreatedTs > 0 {
noticed = time.Since(time.UnixMilli(r.CreatedTs)).Round(time.Minute).String() + " ago"
}
out = append(out, routineRow{ID: r.ID, Phrase: pattern.PhraseRoutine(&p), Noticed: noticed})
}
return out
}
// acceptRoutine creates the recurring reminder for a proposal, then marks the
// proposal accepted and links the reminder to it. Weekly patterns get a cron
// expression; any other interval fires once.
//
// TODO(vikunja#46): this mirrors the voice accept path in cmd/mavend/voice.go.
// When the tick loop learns to read accepted proposals directly, both callers
// should hand off to one place in core instead of each building a reminder.
func acceptRoutine(ctx context.Context, core ipc.CoreAPI, id int64) error {
proposed, err := core.ListProposedRoutines(ctx)
if err != nil {
return err
}
var found *ipc.ProposedRoutine
for i := range proposed {
if proposed[i].ID == id {
found = &proposed[i]
break
}
}
if found == nil {
return errors.New("no such proposed routine")
}
// No reminder is created here. Accepting only flips the status; the tick
// loop reads accepted routines and nudges on the interval (Vikunja #366).
// The old code made a one-shot reminder, so a non-weekly routine fired
// once and then went quiet forever.
return core.AcceptProposedRoutine(ctx, id)
}
func handleTrace(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
if core == nil {
http.Error(w, "trace disabled (no -core)", http.StatusServiceUnavailable)
+5 -2
View File
@@ -36,10 +36,13 @@
"stt": { "socket": "/run/maven/stt.sock", "lang": "ru" },
"tts": { "socket": "/run/maven/tts.sock", "lang": "ru" },
"embedder": {
"model_path": "/opt/maven/models/embedder/model.onnx",
"tokenizer_path": "/opt/maven/models/embedder/tokenizer.json",
"model_path": "/opt/maven/models/embedder/multilingual-e5-small/model_quantized.onnx",
"tokenizer_path": "/opt/maven/models/embedder/multilingual-e5-small/tokenizer.json",
"lib_path": "/opt/maven/lib/libonnxruntime.so"
},
"llm_router": true,
"query_min_score": 0.55,
"query_min_margin": 0.008,
"tool_timeout": "30s",
"tools": [
{ "name": "status", "cmd": ["systemctl", "status"], "scope": "homelab", "destructive": false },
+3
View File
@@ -446,6 +446,9 @@ func (r *recordingAPI) RevertFact(_ context.Context, _ string) (int64, error) {
func (r *recordingAPI) ListProposedRoutines(_ context.Context) ([]ipc.ProposedRoutine, error) {
return nil, nil
}
func (r *recordingAPI) AcceptProposedRoutine(_ context.Context, _ int64) error {
return nil
}
func (r *recordingAPI) DismissProposedRoutine(_ context.Context, _ int64) error {
return nil
}
+58 -1
View File
@@ -257,12 +257,41 @@ type VoiceConfig struct {
// Default 0.35 if unset.
RouterThreshold float64 `json:"router_threshold,omitempty"`
// LLMRouter — route with the resident model instead of the embedding
// classifier. On by default since Vikunja #320.
//
// Measured on the held-out fixture (ROUTING-EVAL-31-07-2026.md): 63.2% of
// intents right against the classifier's 50.0%, and no route errors. It
// costs about 1s per turn instead of 30ms.
//
// It is safe to leave on. The model can refuse — it answers "unknown" when
// it cannot route, and the turn drops to the classifier and its clarify
// gate. Any LLM error does the same, so a turn never breaks on the model.
// Slot extraction runs on LLM decisions too, so acts get their Fn and
// reminders their Time.
//
// Set it false to go back to the classifier, e.g. on a box with no
// llama-server or when 1s a turn is too slow.
//
// It is a pointer so that "missing from the file" and "explicitly false"
// are different things: missing means on, false means off. Read it with
// UseLLMRouter(), not directly.
LLMRouter *bool `json:"llm_router,omitempty"`
// QueryMinScore — the note-recall confidence gate. Top cosine below this
// ⇒ "I don't know" instead of a guess. Tuned for the ONNX embedder (0.55);
// the HashEmbedder floor scores lexically and may never clear it. 0.55
// default if unset.
QueryMinScore float64 `json:"query_min_score,omitempty"`
// QueryMinMargin — the second half of the recall gate: the top hit must
// beat the runner-up by more than this. The absolute score above cannot do
// the job on its own, because the e5 embedder puts every cosine in one
// narrow high band, so a made-up question scores as high as a real one.
// The margin asks whether one note is clearly the best instead.
// Negative ⇒ off. 0 ⇒ the default below.
QueryMinMargin float64 `json:"query_min_margin,omitempty"`
// Persona — optional prompt prefix that tunes maven's character. Prepended
// to every LLM system prompt (nudge phrasing, note queries, general
// knowledge). Empty string ⇒ current hardcoded persona (feminine-gendered
@@ -390,7 +419,14 @@ const (
DefaultAutotuneInterval = 10 * time.Minute
DefaultRouterThreshold = 0.55
DefaultQueryMinScore = 0.55
DefaultToolTimeout = 30 * time.Second
// Read off the margin sweep in internal/memory/recalleval on the e5
// embedder: 0.008 answers 68% of real questions (down from 72%) and cuts
// false recall from 5/5 to 1/5. Every larger delta costs real recall
// without removing that last one until 0.020, which drops recall to 44%.
DefaultQueryMinMargin = 0.008
DefaultToolTimeout = 30 * time.Second
// DefaultLLMRouter — route with the resident model unless told otherwise.
DefaultLLMRouter = true
DefaultFactEnrichmentInterval = 30 * time.Second
)
@@ -472,9 +508,21 @@ func (c *Config) applyDefaults() {
if c.Voice.QueryMinScore <= 0 {
c.Voice.QueryMinScore = DefaultQueryMinScore
}
// Unset ⇒ default. Negative is how you turn the margin off on purpose,
// so it is clamped to 0 rather than replaced by the default.
switch {
case c.Voice.QueryMinMargin == 0:
c.Voice.QueryMinMargin = DefaultQueryMinMargin
case c.Voice.QueryMinMargin < 0:
c.Voice.QueryMinMargin = 0
}
if c.Voice.ToolTimeout <= 0 {
c.Voice.ToolTimeout = Duration(DefaultToolTimeout)
}
if c.Voice.LLMRouter == nil {
on := DefaultLLMRouter
c.Voice.LLMRouter = &on
}
}
// routines: default severity to care-class (1) — the safe floor: a
@@ -493,6 +541,15 @@ func (c *Config) applyDefaults() {
}
}
// UseLLMRouter reports whether to route with the resident model. Unset means
// on; only an explicit false in the config turns it off.
func (v *VoiceConfig) UseLLMRouter() bool {
if v == nil || v.LLMRouter == nil {
return DefaultLLMRouter
}
return *v.LLMRouter
}
func (c *Config) validate() error {
if c.Phraser != nil {
if c.Phraser.ModelPath == "" {
+34
View File
@@ -171,6 +171,40 @@ func TestWeatherConfigNilOK(t *testing.T) {
}
}
func TestLLMRouterDefaultsOn(t *testing.T) {
p := writeConfig(t, `{"voice":{"enabled":true,"bind":"127.0.0.1:9100"}}`)
c, err := Load(p)
if err != nil {
t.Fatalf("Load: %v", err)
}
if !c.Voice.UseLLMRouter() {
t.Error("voice.llm_router absent should mean on")
}
}
// Missing and explicitly false must not mean the same thing.
func TestLLMRouterExplicitFalseTurnsItOff(t *testing.T) {
p := writeConfig(t, `{"voice":{"enabled":true,"bind":"127.0.0.1:9100","llm_router":false}}`)
c, err := Load(p)
if err != nil {
t.Fatalf("Load: %v", err)
}
if c.Voice.UseLLMRouter() {
t.Error("voice.llm_router false should turn it off")
}
}
func TestLLMRouterRead(t *testing.T) {
p := writeConfig(t, `{"voice":{"enabled":true,"bind":"127.0.0.1:9100","llm_router":true}}`)
c, err := Load(p)
if err != nil {
t.Fatalf("Load: %v", err)
}
if !c.Voice.UseLLMRouter() {
t.Error("voice.llm_router true was not read")
}
}
func TestDurationRoundTrip(t *testing.T) {
d := Duration(15 * time.Minute)
b, err := d.MarshalJSON()
+73 -14
View File
@@ -160,12 +160,20 @@ func (d *Dispatcher) DispatchNudge(ctx context.Context, pn PhrasedNudge, now tim
RepeatUntilAck: ch == ChannelTelegram && c.Severity >= loop.Sev4,
Ts: now,
}
s = minimalForAway(s)
sink := d.sinkFor(ch)
if sink == nil {
continue
}
attemptID := d.beginOutbox(ctx, "nudge", c.Rule.Name, 0, ch, messageForChannel(s), now)
if err := sink.Send(ctx, s); err != nil {
if err := safeSend(ctx, sink, s); err != nil {
if errors.Is(err, ErrSinkPanicked) {
// one broken sink must not eat the other channels for this
// nudge (sev4 present is voice + ntfy). the attempt is closed
// as failed and we move on.
d.completeOutbox(ctx, attemptID, store.DeliveryFailed, now)
continue
}
if errors.Is(err, ErrVoiceNoSession) {
// voice was assumed reachable (presence=present) but no live
// session exists — the presence guess was wrong. reroute through
@@ -225,12 +233,17 @@ func (d *Dispatcher) DispatchReminder(ctx context.Context, pr PhrasedReminder, n
Summary: pr.Summary,
Ts: now,
}
s = minimalForAway(s)
sink := d.sinkFor(ch)
if sink == nil {
continue
}
attemptID := d.beginOutbox(ctx, "reminder", "", rd.Reminder.ID, ch, messageForChannel(s), now)
if err := sink.Send(ctx, s); err != nil {
if err := safeSend(ctx, sink, s); err != nil {
if errors.Is(err, ErrSinkPanicked) {
d.completeOutbox(ctx, attemptID, store.DeliveryFailed, now)
continue
}
if errors.Is(err, ErrVoiceNoSession) {
// presence guess was wrong — reroute reminder to the away
// channel (ntfy). voice is the only present channel, so nothing
@@ -315,9 +328,13 @@ func (d *Dispatcher) RepeatUnacked(ctx context.Context, keys []string, now time.
RepeatUntilAck: true,
Ts: now,
}
s = minimalForAway(s)
attemptID := d.beginOutbox(ctx, "nudge", key, 0, ChannelTelegram, messageForChannel(s), now)
if err := d.cfg.Telegram.Send(ctx, s); err != nil {
if err := safeSend(ctx, d.cfg.Telegram, s); err != nil {
d.completeOutbox(ctx, attemptID, store.DeliveryFailed, now)
if errors.Is(err, ErrSinkPanicked) {
continue
}
return out, fmt.Errorf("repeat send telegram %s: %w", key, err)
}
d.completeOutbox(ctx, attemptID, store.DeliverySent, now)
@@ -329,6 +346,24 @@ func (d *Dispatcher) RepeatUnacked(ctx context.Context, keys []string, now time.
return out, nil
}
// ErrSinkPanicked — a sink panicked mid-send. the send did not happen, so the
// attempt is recorded failed and never silently retried as if it had.
var ErrSinkPanicked = errors.New("delivery: sink panicked mid-send")
// safeSend calls a sink and turns a panic into an error. without this a
// panicking sink unwinds past completeOutbox and leaves the delivery_attempts
// row pending forever — reconciliation only runs at daemon startup, and core
// is long-lived, so the row would sit there for weeks.
func safeSend(ctx context.Context, sink Sink, s Sendable) (err error) {
defer func() {
if r := recover(); r != nil {
log.Printf("dispatcher: PANIC in %s sink (this is a bug, fix the sink): %v", s.Channel, r)
err = fmt.Errorf("%w: %s: %v", ErrSinkPanicked, s.Channel, r)
}
}()
return sink.Send(ctx, s)
}
func (d *Dispatcher) sinkFor(ch Channel) Sink {
switch ch {
case ChannelVoice:
@@ -342,20 +377,44 @@ func (d *Dispatcher) sinkFor(ch Channel) Sink {
}
}
// GenericAwayMessage — what an away channel gets when the phraser gave us no
// summary. no gendered forms, so it stays right whoever reads it.
const GenericAwayMessage = "что-то требует внимания"
// isAway — this channel leaves the box, so it only ever gets a minimal body.
func isAway(ch Channel) bool {
return ch == ChannelNtfy || ch == ChannelTelegram
}
// messageForChannel — away channels get the minimal summary (no shoulder-surf
// exfil — "disk low on homesrv," not detail); voice gets the full body (local).
// a missing summary falls back to body — a terse full message is better than
// no message, and the phraser should have produced a summary for away-bound
// severities. this is the "minimal body" rule from the spec, enforced at the
// last mile so a phraser bug can't accidentally exfil via the relay.
// an empty summary must NOT fall back to the body: the resident model is small
// and drops fields often, and the away path crosses the "never phones home"
// boundary. so we send a fixed generic line plus the rule name instead. voice
// is local, so it keeps the full body.
func messageForChannel(s Sendable) string {
switch s.Channel {
case ChannelNtfy, ChannelTelegram:
if s.Summary != "" {
return s.Summary
}
return s.Body
default:
if !isAway(s.Channel) {
return s.Body
}
if s.Summary != "" {
return s.Summary
}
if s.RuleName != "" {
return GenericAwayMessage + ": " + s.RuleName
}
return GenericAwayMessage
}
// minimalForAway — strips detail from a Sendable bound for an away channel
// before any sink sees it. the sinks pick Summary themselves too, but this is
// where the boundary actually is: a sink added later must not be able to leak
// the full body just by reading the wrong field.
func minimalForAway(s Sendable) Sendable {
if !isAway(s.Channel) {
return s
}
msg := messageForChannel(s)
s.Body = msg
s.Summary = msg
return s
}
+219 -5
View File
@@ -639,11 +639,11 @@ func TestDispatchRecurringReminderReschedules(t *testing.T) {
// ----------------------------- durable outbox --------------------------------
type outboxAttempt struct {
kind, rule string
reminderID int64
channel, hash string
status string
begunAt, doneAt time.Time
kind, rule string
reminderID int64
channel, hash string
status string
begunAt, doneAt time.Time
}
// fakeOutbox — an in-memory Outbox that also lets a test simulate a crash
@@ -784,3 +784,217 @@ func TestDispatchNudge_OutboxBeginFailureDoesNotBlockSend(t *testing.T) {
t.Fatalf("send should still happen despite outbox begin failure: sends=%v out=%v", voice.sends, out)
}
}
// ----------------------- away channels carry no detail -----------------------
// panicSink lives in durability_test.go — a second agent wrote the same helper
// for the same #369 case, so this file just uses that one.
// TestAwaySendsGenericLineWhenSummaryEmpty — #368. The phraser is a small
// model and drops fields often. An empty Summary must NOT put the full body
// on a channel that leaves the box; the away sendable gets a fixed generic
// line plus the rule name instead.
func TestAwaySendsGenericLineWhenSummaryEmpty(t *testing.T) {
ntfy := &fakeSink{}
rec := &fakeNudgeRecorder{}
d := NewDispatcher(Config{Ntfy: ntfy, Nudges: rec})
body := "disk /dev/sda1 at 96%, 2.1G free, largest offender /var/lib/docker"
_, err := d.DispatchNudge(context.Background(), PhrasedNudge{
Candidate: candidate("disk-low", loop.Sev3, store.Away),
Body: body,
Summary: "",
}, refNow())
if err != nil {
t.Fatalf("dispatch: %v", err)
}
if len(ntfy.sends) != 1 {
t.Fatalf("want 1 ntfy send, got %d", len(ntfy.sends))
}
want := GenericAwayMessage + ": disk-low"
got := messageForChannel(ntfy.sends[0])
if got != want {
t.Fatalf("away message: want %q, got %q", want, got)
}
if ntfy.sends[0].Body == body {
t.Fatal("away sendable still carries the full body")
}
if len(rec.rows) != 1 || rec.rows[0].message != want {
t.Fatalf("recorded message: want %q, got %+v", want, rec.rows)
}
}
// TestSev4AwaySendableCarriesNoDetail — the rule is enforced at the
// dispatcher, not in each sink. A sink added later must not be able to leak
// the body just by reading the wrong field, so neither field may hold detail.
func TestSev4AwaySendableCarriesNoDetail(t *testing.T) {
tg := &fakeSink{}
d := NewDispatcher(Config{Telegram: tg, Ack: newFakeAck()})
body := "backup job failed: rsync exit 23 on /home/kami, see /var/log/backup.log"
_, err := d.DispatchNudge(context.Background(), PhrasedNudge{
Candidate: candidate("backup-failed", loop.Sev4, store.Away),
Body: body,
Summary: "бэкап не прошёл",
}, refNow())
if err != nil {
t.Fatalf("dispatch: %v", err)
}
if len(tg.sends) != 1 {
t.Fatalf("want 1 telegram send, got %d", len(tg.sends))
}
s := tg.sends[0]
if s.Body != "бэкап не прошёл" || s.Summary != "бэкап не прошёл" {
t.Fatalf("away sendable should hold only the summary, got body=%q summary=%q", s.Body, s.Summary)
}
}
// TestAwayKeepsANonEmptySummary — the normal path is untouched.
func TestAwayKeepsANonEmptySummary(t *testing.T) {
ntfy := &fakeSink{}
d := NewDispatcher(Config{Ntfy: ntfy})
_, err := d.DispatchNudge(context.Background(), PhrasedNudge{
Candidate: candidate("cert", loop.Sev3, store.Away),
Body: "cert for maven.local expires in 3 days, issuer letsencrypt",
Summary: "сертификат истекает",
}, refNow())
if err != nil {
t.Fatalf("dispatch: %v", err)
}
if got := messageForChannel(ntfy.sends[0]); got != "сертификат истекает" {
t.Fatalf("want the summary unchanged, got %q", got)
}
}
// TestVoiceStillGetsTheFullBody — voice never leaves the box, so it keeps the
// full phrased message even when Summary is empty.
func TestVoiceStillGetsTheFullBody(t *testing.T) {
voice := &fakeSink{}
d := NewDispatcher(Config{Voice: voice})
body := "ты не пил воду четыре часа"
_, err := d.DispatchNudge(context.Background(), PhrasedNudge{
Candidate: candidate("water", loop.Sev1, store.Present),
Body: body,
Summary: "",
}, refNow())
if err != nil {
t.Fatalf("dispatch: %v", err)
}
if len(voice.sends) != 1 || voice.sends[0].Body != body {
t.Fatalf("voice should get the full body, got %+v", voice.sends)
}
if got := messageForChannel(voice.sends[0]); got != body {
t.Fatalf("voice message: want %q, got %q", body, got)
}
}
// TestReminderAwaySendsGenericLineWhenSummaryEmpty — the reminder path crosses
// the same boundary and has its own Sendable construction.
func TestReminderAwaySendsGenericLineWhenSummaryEmpty(t *testing.T) {
ntfy := &fakeSink{}
d := NewDispatcher(Config{Ntfy: ntfy})
_, err := d.DispatchReminder(context.Background(), PhrasedReminder{
Decision: loop.ReminderDecision{
Reminder: store.Reminder{ID: 5},
State: loop.State{Now: refNow(), Presence: store.Away},
},
Body: "позвонить в клинику по поводу анализов",
Summary: "",
}, refNow())
if err != nil {
t.Fatalf("dispatch: %v", err)
}
if len(ntfy.sends) != 1 {
t.Fatalf("want 1 ntfy send, got %d", len(ntfy.sends))
}
if got := messageForChannel(ntfy.sends[0]); got != GenericAwayMessage {
t.Fatalf("away reminder message: want %q, got %q", GenericAwayMessage, got)
}
}
// --------------------------- a panicking sink -------------------------------
// TestPanicMidSendResolvesTheAttempt — #369. A panic used to unwind past
// completeOutbox and leave the row pending forever, because reconciliation
// only runs at daemon startup and core is long-lived.
func TestPanicMidSendResolvesTheAttempt(t *testing.T) {
ob := &fakeOutbox{}
d := NewDispatcher(Config{Voice: &panicSink{}, Outbox: ob})
_, err := d.DispatchNudge(context.Background(), PhrasedNudge{
Candidate: candidate("water", loop.Sev1, store.Present),
Body: "body", Summary: "sum",
}, refNow())
if err != nil {
t.Fatalf("a panicking sink must not fail the dispatch: %v", err)
}
if len(ob.attempts) != 1 {
t.Fatalf("want 1 outbox attempt, got %d", len(ob.attempts))
}
if ob.attempts[0].status != store.DeliveryFailed {
t.Fatalf("want status failed after a panic, got %q", ob.attempts[0].status)
}
}
// TestPanicInOneSinkStillDeliversTheOther — sev4 present is voice + ntfy. One
// broken sink must not eat the other channel for the same nudge.
func TestPanicInOneSinkStillDeliversTheOther(t *testing.T) {
bad := &panicSink{}
ntfy := &fakeSink{}
ob := &fakeOutbox{}
d := NewDispatcher(Config{Voice: bad, Ntfy: ntfy, Outbox: ob})
out, err := d.DispatchNudge(context.Background(), PhrasedNudge{
Candidate: candidate("backup-failed", loop.Sev4, store.Present),
Body: "long detailed body", Summary: "бэкап не прошёл",
}, refNow())
if err != nil {
t.Fatalf("dispatch: %v", err)
}
if bad.calls != 1 {
t.Fatalf("want the voice sink called once, got %d", bad.calls)
}
if len(ntfy.sends) != 1 {
t.Fatalf("ntfy should still get the nudge, got %d sends", len(ntfy.sends))
}
if len(out) != 1 || out[0].Sendable.Channel != ChannelNtfy {
t.Fatalf("want only the ntfy dispatch reported, got %+v", out)
}
if len(ob.attempts) != 2 {
t.Fatalf("want 2 outbox attempts, got %d", len(ob.attempts))
}
if ob.attempts[0].status != store.DeliveryFailed || ob.attempts[1].status != store.DeliverySent {
t.Fatalf("want [failed, sent], got %q %q", ob.attempts[0].status, ob.attempts[1].status)
}
}
// TestPanicInReminderSinkResolvesTheAttempt — the reminder path has its own
// send call, and a panic there must not leave the reminder marked fired.
func TestPanicInReminderSinkResolvesTheAttempt(t *testing.T) {
ob := &fakeOutbox{}
rc := &fakeReminderCompleter{}
d := NewDispatcher(Config{Voice: &panicSink{}, Reminders: rc, Outbox: ob})
out, err := d.DispatchReminder(context.Background(), PhrasedReminder{
Decision: loop.ReminderDecision{
Reminder: store.Reminder{ID: 9},
State: loop.State{Now: refNow(), Presence: store.Present},
},
Body: "звонок", Summary: "звонок",
}, refNow())
if err != nil {
t.Fatalf("a panicking sink must not fail the dispatch: %v", err)
}
if len(out) != 0 {
t.Fatalf("nothing was delivered, want no dispatches, got %+v", out)
}
if len(ob.attempts) != 1 || ob.attempts[0].status != store.DeliveryFailed {
t.Fatalf("want 1 attempt with status failed, got %+v", ob.attempts)
}
if len(rc.marked) != 0 {
t.Fatalf("reminder must stay pending after a panic, got %+v", rc.marked)
}
}
+228
View File
@@ -0,0 +1,228 @@
package delivery
import (
"context"
"errors"
"path/filepath"
"testing"
"time"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/store"
)
// panicSink — a sink that dies mid-send. Models the ugly case: the process is
// still alive, so startup reconciliation will not run, but the attempt row was
// already begun.
type panicSink struct{ calls int }
func (p *panicSink) Send(_ context.Context, _ Sendable) error {
p.calls++
panic("sink exploded mid-send")
}
// ------------------------- voice fallthrough, per severity -------------------
// TestVoiceNoSessionFallthroughLeavesOutboxTrail — the fallthrough must be
// visible in the ledger too: the voice attempt closes as failed and the away
// attempt is a separate row, so an operator can see the reroute happened.
func TestVoiceNoSessionFallthroughLeavesOutboxTrail(t *testing.T) {
cases := []struct {
name string
sev loop.Severity
wantAt []string // channel per outbox attempt, in order
wantEnd []string // status per attempt, in order
}{
{"sev3 falls through to ntfy", loop.Sev3,
[]string{"voice", "ntfy"}, []string{store.DeliveryFailed, store.DeliverySent}},
{"sev4 falls through to telegram", loop.Sev4,
[]string{"voice", "telegram"}, []string{store.DeliveryFailed, store.DeliverySent}},
{"sev1 does not fall through", loop.Sev1,
[]string{"voice"}, []string{store.DeliveryFailed}},
{"sev2 does not fall through", loop.Sev2,
[]string{"voice"}, []string{store.DeliveryFailed}},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
voice := &fakeSink{err: ErrVoiceNoSession}
ntfy, telegram := &fakeSink{}, &fakeSink{}
ob := &fakeOutbox{}
d := NewDispatcher(Config{
Voice: voice, Ntfy: ntfy, Telegram: telegram,
Ack: newFakeAck(), Outbox: ob,
})
if _, err := d.DispatchNudge(context.Background(), PhrasedNudge{
Candidate: candidate("some_rule", c.sev, store.Present),
Body: "detail", Summary: "short",
}, refNow()); err != nil {
t.Fatalf("dispatch: %v", err)
}
if len(ob.attempts) != len(c.wantAt) {
t.Fatalf("want %d outbox attempts, got %d (%+v)", len(c.wantAt), len(ob.attempts), ob.attempts)
}
for i, a := range ob.attempts {
if a.channel != c.wantAt[i] || a.status != c.wantEnd[i] {
t.Fatalf("attempt %d: want %s/%s, got %s/%s", i, c.wantAt[i], c.wantEnd[i], a.channel, a.status)
}
}
// care severities must not reach an away channel — that would
// defeat the drop rule.
if c.sev <= loop.Sev2 && (len(ntfy.sends) != 0 || len(telegram.sends) != 0) {
t.Fatalf("care nudge escaped to an away channel: ntfy=%d telegram=%d",
len(ntfy.sends), len(telegram.sends))
}
})
}
}
// ------------------------- crash between Begin and Complete ------------------
// openTestStore — a real store on a temp file. The reconciliation promise is a
// SQL promise, so a fake would only test the fake.
func openTestStore(t *testing.T) *store.Store {
t.Helper()
st, err := store.Open(context.Background(), filepath.Join(t.TempDir(), "maven.db"))
if err != nil {
t.Fatalf("open store: %v", err)
}
t.Cleanup(func() { _ = st.Close() })
return st
}
// attemptStatus reads one attempt row back. Returns ok=false when the row is
// gone, which would itself be a broken promise (a dropped attempt).
func attemptStatus(t *testing.T, st *store.Store, id int64) (status string, completed bool, ok bool) {
t.Helper()
tx, err := st.DB(context.Background())
if err != nil {
t.Fatalf("read tx: %v", err)
}
defer func() { _ = tx.Rollback() }()
var completedTS *int64
err = tx.QueryRowContext(context.Background(),
`SELECT status, completed_ts FROM delivery_attempts WHERE id = ?`, id).Scan(&status, &completedTS)
if err != nil {
return "", false, false
}
return status, completedTS != nil, true
}
// TestCrashBetweenBeginAndCompleteBecomesUnknown — simulate the crash window:
// Begin lands, the process dies before Complete. Startup reconciliation must
// turn that row into "unknown" — neither resent nor dropped, because Maven
// cannot know whether the message left the box.
func TestCrashBetweenBeginAndCompleteBecomesUnknown(t *testing.T) {
st := openTestStore(t)
ctx := context.Background()
sink := &fakeSink{}
// the crash: intent recorded, no completion.
id, err := st.BeginDeliveryAttempt(ctx, "nudge", "disk_low", 0, "telegram", "hash", refNow())
if err != nil {
t.Fatalf("begin: %v", err)
}
if s, _, ok := attemptStatus(t, st, id); !ok || s != store.DeliveryPending {
t.Fatalf("before reconcile: want pending, got %q ok=%v", s, ok)
}
// restart.
n, err := st.ReconcileStaleDeliveryAttempts(ctx, refNow().Add(time.Minute))
if err != nil {
t.Fatalf("reconcile: %v", err)
}
if n != 1 {
t.Fatalf("want 1 row reconciled, got %d", n)
}
s, completed, ok := attemptStatus(t, st, id)
if !ok {
t.Fatal("reconciliation dropped the row; the promise is it is never dropped")
}
if s != store.DeliveryUnknown {
t.Fatalf("want status unknown, got %q", s)
}
if !completed {
t.Fatal("reconciled row should carry a completed_ts")
}
// not resent: reconciliation is bookkeeping only, it must never push.
if len(sink.sends) != 0 {
t.Fatalf("reconciliation must not resend, got %d sends", len(sink.sends))
}
// idempotent: a second restart must not churn the row again.
n2, err := st.ReconcileStaleDeliveryAttempts(ctx, refNow().Add(2*time.Minute))
if err != nil {
t.Fatalf("reconcile again: %v", err)
}
if n2 != 0 {
t.Fatalf("second reconcile should find nothing, got %d", n2)
}
if s2, _, _ := attemptStatus(t, st, id); s2 != store.DeliveryUnknown {
t.Fatalf("unknown must stay unknown, got %q", s2)
}
}
// TestUnknownIsNeverResolvedToSentOrFailed — the "never guess" half of the
// promise: nothing may quietly turn an unknown into a definite outcome.
func TestUnknownIsNeverResolvedToSentOrFailed(t *testing.T) {
st := openTestStore(t)
ctx := context.Background()
id, err := st.BeginDeliveryAttempt(ctx, "nudge", "disk_low", 0, "telegram", "hash", refNow())
if err != nil {
t.Fatalf("begin: %v", err)
}
if _, err := st.ReconcileStaleDeliveryAttempts(ctx, refNow()); err != nil {
t.Fatalf("reconcile: %v", err)
}
// a late Complete from the old in-flight send must not win.
if err := st.CompleteDeliveryAttempt(ctx, id, store.DeliverySent, refNow().Add(time.Minute)); err != nil {
t.Fatalf("late complete: %v", err)
}
if s, _, _ := attemptStatus(t, st, id); s != store.DeliveryUnknown {
t.Fatalf("late complete overwrote an unknown outcome: %q", s)
}
}
// ------------------------- the boring failure modes --------------------------
// TestSendTimeoutResolvesTheAttempt — a send that times out is a definite
// failure from Maven's side, so the row must not be left pending.
func TestSendTimeoutResolvesTheAttempt(t *testing.T) {
ctx, cancel := context.WithCancel(context.Background())
cancel() // the deadline already blew
ob := &fakeOutbox{}
d := NewDispatcher(Config{Ntfy: &fakeSink{err: context.DeadlineExceeded}, Outbox: ob})
if _, err := d.DispatchNudge(ctx, PhrasedNudge{
Candidate: candidate("cert_expiring", loop.Sev3, store.Away),
Body: "detail", Summary: "short",
}, refNow()); err == nil {
t.Fatal("want a timeout error to propagate")
}
if len(ob.attempts) != 1 || ob.attempts[0].status != store.DeliveryFailed {
t.Fatalf("timed-out send must close the attempt as failed, got %+v", ob.attempts)
}
}
// TestCompleteFailureLeavesRowPendingForReconciliation — if Complete itself
// fails, the row stays pending on purpose. That is the correct ambiguous state
// and startup reconciliation is what resolves it.
func TestCompleteFailureLeavesRowPendingForReconciliation(t *testing.T) {
ob := &fakeOutbox{completeErr: errors.New("db busy")}
d := NewDispatcher(Config{Ntfy: &fakeSink{}, Outbox: ob})
if _, err := d.DispatchNudge(context.Background(), PhrasedNudge{
Candidate: candidate("cert_expiring", loop.Sev3, store.Away),
Body: "detail", Summary: "short",
}, refNow()); err != nil {
t.Fatalf("a failed outbox complete must not fail the dispatch: %v", err)
}
if len(ob.attempts) != 1 || ob.attempts[0].status != store.DeliveryPending {
t.Fatalf("want the row left pending, got %+v", ob.attempts)
}
}
// The panic gap this file used to describe as a skipped test is fixed and
// asserted for real in dispatcher_test.go:TestPanicMidSendResolvesTheAttempt.
// panicSink stays here because both files use it.
+227
View File
@@ -0,0 +1,227 @@
package delivery
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/store"
)
// This file walks every cell of the DESIGN.md § "Delivery / channel routing"
// table, once as the pure table and once through the dispatcher, so a change
// to either side has to break a named cell.
//
// present away
// sev1-2 (care) voice drop
// sev3 (soft) voice ntfy, once
// sev4 (hard) voice + ntfy telegram, repeat til ack
type tableCell struct {
name string
sev loop.Severity
presence store.Bucket
want []Channel
}
func allTableCells() []tableCell {
return []tableCell{
{"sev1 present", loop.Sev1, store.Present, []Channel{ChannelVoice}},
{"sev2 present", loop.Sev2, store.Present, []Channel{ChannelVoice}},
{"sev3 present", loop.Sev3, store.Present, []Channel{ChannelVoice}},
{"sev4 present", loop.Sev4, store.Present, []Channel{ChannelVoice, ChannelNtfy}},
{"sev1 away", loop.Sev1, store.Away, []Channel{ChannelDrop}},
{"sev2 away", loop.Sev2, store.Away, []Channel{ChannelDrop}},
{"sev3 away", loop.Sev3, store.Away, []Channel{ChannelNtfy}},
{"sev4 away", loop.Sev4, store.Away, []Channel{ChannelTelegram}},
}
}
func sameChannels(got, want []Channel) bool {
if len(got) != len(want) {
return false
}
for i := range got {
if got[i] != want[i] {
return false
}
}
return true
}
func TestChannelsForEveryTableCell(t *testing.T) {
for _, c := range allTableCells() {
t.Run(c.name, func(t *testing.T) {
got := ChannelsFor(c.sev, c.presence)
if !sameChannels(got, c.want) {
t.Fatalf("%s: want %v, got %v", c.name, c.want, got)
}
})
}
}
// TestDispatchNudgeEveryTableCell — the same eight cells end to end: exactly
// the wanted channels get a send, and every other channel gets none.
func TestDispatchNudgeEveryTableCell(t *testing.T) {
for _, c := range allTableCells() {
t.Run(c.name, func(t *testing.T) {
voice, ntfy, telegram := &fakeSink{}, &fakeSink{}, &fakeSink{}
rec := &fakeNudgeRecorder{}
d := NewDispatcher(Config{
Voice: voice, Ntfy: ntfy, Telegram: telegram,
Ack: newFakeAck(), Nudges: rec,
})
out, err := d.DispatchNudge(context.Background(), PhrasedNudge{
Candidate: candidate("some_rule", c.sev, c.presence),
Body: "full detail body",
Summary: "short form",
}, refNow())
if err != nil {
t.Fatalf("dispatch: %v", err)
}
sent := map[Channel]int{
ChannelVoice: len(voice.sends),
ChannelNtfy: len(ntfy.sends),
ChannelTelegram: len(telegram.sends),
}
for ch, n := range sent {
want := 0
for _, w := range c.want {
if w == ch {
want = 1
}
}
if n != want {
t.Fatalf("%s: channel %s got %d sends, want %d", c.name, ch, n, want)
}
}
// one dispatch and one nudge row per real (non-drop) channel.
wantDispatches := 0
for _, w := range c.want {
if w != ChannelDrop {
wantDispatches++
}
}
if len(out) != wantDispatches {
t.Fatalf("%s: want %d dispatches, got %d", c.name, wantDispatches, len(out))
}
if len(rec.rows) != wantDispatches {
t.Fatalf("%s: want %d nudge rows, got %d", c.name, wantDispatches, len(rec.rows))
}
})
}
}
// TestSev3AwayIsNtfyExactlyOnce — "ntfy, once": one send, and nothing on the
// sendable asks for a repeat, so the daemon's repeat driver has no reason to
// pick it up.
func TestSev3AwayIsNtfyExactlyOnce(t *testing.T) {
ntfy := &fakeSink{}
ack := newFakeAck()
d := NewDispatcher(Config{Ntfy: ntfy, Telegram: &fakeSink{}, Ack: ack})
out, err := d.DispatchNudge(context.Background(), PhrasedNudge{
Candidate: candidate("cert_expiring", loop.Sev3, store.Away),
Body: "cert detail", Summary: "cert expiring",
}, refNow())
if err != nil {
t.Fatalf("dispatch: %v", err)
}
if len(ntfy.sends) != 1 {
t.Fatalf("sev3 away: want exactly 1 ntfy send, got %d", len(ntfy.sends))
}
if out[0].Sendable.RepeatUntilAck {
t.Fatalf("sev3 away must not repeat til ack")
}
if _, ok := ack.lastSent["cert_expiring"]; ok {
t.Fatalf("sev3 away must not enter the ack/repeat tracker")
}
}
// TestSev4AwayRepeatsUntilAcked — "telegram, repeat til ack": the initial send
// arms the ack clock, the repeat driver re-sends while un-acked, and an ack
// stops it.
func TestSev4AwayRepeatsUntilAcked(t *testing.T) {
telegram := &fakeSink{}
ack := newFakeAck()
d := NewDispatcher(Config{Telegram: telegram, Ack: ack})
ctx := context.Background()
if _, err := d.DispatchNudge(ctx, PhrasedNudge{
Candidate: candidate("disk_low", loop.Sev4, store.Away),
Body: "disk detail", Summary: "disk low on homesrv",
}, refNow()); err != nil {
t.Fatalf("dispatch: %v", err)
}
// two intervals pass, still un-acked → two more sends.
for i := 1; i <= 2; i++ {
at := refNow().Add(time.Duration(i) * 10 * time.Minute)
if _, err := d.RepeatUnacked(ctx, []string{"disk_low"}, at, 5*time.Minute, "disk detail", "disk low on homesrv"); err != nil {
t.Fatalf("repeat %d: %v", i, err)
}
}
if len(telegram.sends) != 3 {
t.Fatalf("want 1 initial + 2 repeats = 3 telegram sends, got %d", len(telegram.sends))
}
// acked → no further sends, however long we wait.
_ = ack.MarkAcked(ctx, "disk_low")
if _, err := d.RepeatUnacked(ctx, []string{"disk_low"}, refNow().Add(time.Hour), 5*time.Minute, "b", "s"); err != nil {
t.Fatalf("repeat after ack: %v", err)
}
if len(telegram.sends) != 3 {
t.Fatalf("ack must stop the repeat; got %d sends", len(telegram.sends))
}
}
// TestAwayChannelsGetMinimalBody — what leaves the box is the short form, for
// every away cell of the table. messageForChannel is the last-mile choice both
// away sinks make too.
func TestAwayChannelsGetMinimalBody(t *testing.T) {
detail := "disk /mnt/hdd1 on homesrv at 97% — 12GB free, biggest offender /var/lib/docker"
short := "disk low on homesrv"
for _, ch := range []Channel{ChannelNtfy, ChannelTelegram} {
t.Run(string(ch), func(t *testing.T) {
msg := messageForChannel(Sendable{Channel: ch, Body: detail, Summary: short})
if msg != short {
t.Fatalf("%s message: want %q, got %q", ch, short, msg)
}
})
}
if got := messageForChannel(Sendable{Channel: ChannelVoice, Body: detail, Summary: short}); got != detail {
t.Fatalf("voice is local and gets the full body, got %q", got)
}
}
// Both away-detail gaps this file used to describe as skipped tests are now
// fixed and asserted for real in dispatcher_test.go:
// TestSev4AwaySendableCarriesNoDetail and TestAwaySendsGenericLineWhenSummaryEmpty.
// An empty Summary no longer means "send the whole body" — it means a short
// generic line — so the old expectation here was wrong as well as duplicated.
// TestCareAwayDropIsRecorded — DESIGN.md's drop is a decision ("a missed water
// nudge is noise, a missed backup failure isn't"), so it should be visible
// rather than vanish. Today drop is a bare `continue`: no nudge row, no outbox
// attempt, no log — nothing an operator can see afterwards.
func TestCareAwayDropIsRecorded(t *testing.T) {
t.Skip("not implemented: dispatcher.go:149-151 skips a Drop channel with no record; there is no 'dropped' outcome in store/delivery.go:16-21")
ob := &fakeOutbox{}
d := NewDispatcher(Config{Voice: &fakeSink{}, Nudges: &fakeNudgeRecorder{}, Outbox: ob})
if _, err := d.DispatchNudge(context.Background(), PhrasedNudge{
Candidate: candidate("water", loop.Sev1, store.Away),
Body: "drink water", Summary: "water",
}, refNow()); err != nil {
t.Fatalf("dispatch: %v", err)
}
if len(ob.attempts) != 1 || ob.attempts[0].channel != string(ChannelDrop) {
t.Fatalf("care-away drop should leave a visible record, got %+v", ob.attempts)
}
}
+171
View File
@@ -0,0 +1,171 @@
package dialogue
import (
"sync"
"time"
)
// Slot names one field of Slots. Named type, not a free string, so a missing
// slot cannot be misspelled — the question phrasing switches on these.
type Slot string
const (
SlotTime Slot = "time" // Slots.Time / HasTime
SlotKey Slot = "key" // Slots.Key / HasKey
SlotValue Slot = "value" // Slots.Value (paired with Key)
SlotFn Slot = "fn" // Slots.Fn / HasFn
SlotText Slot = "text" // Slots.Text
)
// MaxAttempts is 1 because Maven is not a nag (DESIGN.md § Non-goals). She asks
// one clarifying question. If the answer still leaves the slot empty she drops
// the request instead of asking again.
const MaxAttempts = 1
// PendingQuestion is what Maven holds while she waits for an answer to an open
// question. Unlike the yes/no confirms in cmd/mavend/voice.go, the answer here
// is free text that fills a missing slot rather than a verdict.
type PendingQuestion struct {
Intent Intent // what the router already guessed
Slots Slots // what it already filled
Missing []Slot // what is still empty, in the order to ask about
Utterance string // the user's original raw words
Asked time.Time
TTL time.Duration
Attempts int // questions already asked; capped by MaxAttempts
}
func (q *PendingQuestion) IsExpired(now time.Time) bool {
return now.After(q.Asked.Add(q.TTL))
}
// CanAsk reports whether Maven may ask another question about this request.
func (q *PendingQuestion) CanAsk() bool {
return q.Attempts < MaxAttempts
}
// TODO: the daemon will phrase the question text from Missing (one short ru
// question per Slot, feminine self-reference) and speak it here.
// ClarifyStore holds the parked questions. Same shape and locking as
// SessionStore: keyed by dialogue id, expired entries dropped on read.
type ClarifyStore struct {
mu sync.RWMutex
questions map[string]*PendingQuestion
defaultTTL time.Duration
}
func NewClarifyStore(defaultTTL time.Duration) *ClarifyStore {
if defaultTTL <= 0 {
// Short, like confirmTTL in voice.go: a clarifying question is a
// same-breath gesture, a stale one should not eat a later utterance.
defaultTTL = 90 * time.Second
}
return &ClarifyStore{
questions: make(map[string]*PendingQuestion),
defaultTTL: defaultTTL,
}
}
// TODO: the daemon will Put a question here when Decision.Clarify fires, in
// place of the flat "не разобрала" reply (cmd/mavend/voice.go).
func (s *ClarifyStore) Put(id string, q *PendingQuestion) {
if q.TTL <= 0 {
q.TTL = s.defaultTTL
}
s.mu.Lock()
s.questions[id] = q
s.mu.Unlock()
}
// TODO: the daemon will Get on the next turn, parse that turn into Slots, call
// Answer, and Delete — the open-question twin of resolveConfirm.
func (s *ClarifyStore) Get(id string, now time.Time) *PendingQuestion {
s.mu.RLock()
q, ok := s.questions[id]
s.mu.RUnlock()
if !ok {
return nil
}
if q.IsExpired(now) {
s.Delete(id)
return nil
}
return q
}
func (s *ClarifyStore) Delete(id string) {
s.mu.Lock()
delete(s.questions, id)
s.mu.Unlock()
}
// Answer merges the slots parsed from the user's answer into the parked ones.
// Only the slots listed in Missing are filled, and an already filled slot is
// never overwritten — the answer completes the original request, it does not
// restate it. Parsing the answer text into `answer` is the caller's job; this
// package must stay free of internal/router.
func (q *PendingQuestion) Answer(text string, answer Slots) Slots {
out := q.Slots
for _, slot := range q.Missing {
switch slot {
case SlotTime:
if !out.HasTime && answer.HasTime {
out.Time = answer.Time
out.HasTime = true
}
case SlotKey:
if !out.HasKey && answer.HasKey {
out.Key = answer.Key
out.HasKey = true
}
case SlotValue:
if out.Value == "" && answer.Value != "" {
out.Value = answer.Value
}
case SlotFn:
if !out.HasFn && answer.HasFn {
out.Fn = answer.Fn
out.HasFn = true
if len(out.Args) == 0 {
out.Args = append([]string(nil), answer.Args...)
}
}
case SlotText:
if out.Text == "" {
if answer.Text != "" {
out.Text = answer.Text
} else {
// No parse for a text slot — the raw answer IS the text.
out.Text = text
}
}
}
}
return out
}
// StillMissing lists the slots that are empty in s, out of the ones asked for.
// The caller uses it to decide between acting and dropping the request.
func StillMissing(want []Slot, s Slots) []Slot {
var out []Slot
for _, slot := range want {
empty := false
switch slot {
case SlotTime:
empty = !s.HasTime
case SlotKey:
empty = !s.HasKey
case SlotValue:
empty = s.Value == ""
case SlotFn:
empty = !s.HasFn
case SlotText:
empty = s.Text == ""
}
if empty {
out = append(out, slot)
}
}
return out
}
+225
View File
@@ -0,0 +1,225 @@
package dialogue
import (
"testing"
"time"
)
var base = time.Date(2026, 7, 31, 12, 0, 0, 0, time.UTC)
func TestPendingQuestionIsExpired(t *testing.T) {
cases := []struct {
name string
ttl time.Duration
now time.Time
want bool
}{
{"fresh", time.Minute, base.Add(10 * time.Second), false},
{"exactly at ttl", time.Minute, base.Add(time.Minute), false},
{"past ttl", time.Minute, base.Add(2 * time.Minute), true},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
q := &PendingQuestion{Asked: base, TTL: tc.ttl}
if got := q.IsExpired(tc.now); got != tc.want {
t.Fatalf("IsExpired = %v, want %v", got, tc.want)
}
})
}
}
func TestClarifyStoreGetPutDelete(t *testing.T) {
s := NewClarifyStore(time.Minute)
if got := s.Get("voice", base); got != nil {
t.Fatalf("empty store returned %+v", got)
}
q := &PendingQuestion{Intent: IntentReminder, Missing: []Slot{SlotTime}, Asked: base}
s.Put("voice", q)
if q.TTL != time.Minute {
t.Fatalf("Put did not apply the default TTL, got %v", q.TTL)
}
if got := s.Get("voice", base.Add(time.Second)); got != q {
t.Fatalf("Get returned %+v, want the parked question", got)
}
// Expired questions are dropped on read, not returned.
if got := s.Get("voice", base.Add(2*time.Minute)); got != nil {
t.Fatalf("expired Get returned %+v", got)
}
if got := s.Get("voice", base); got != nil {
t.Fatalf("expired question was not deleted: %+v", got)
}
s.Put("voice", &PendingQuestion{Asked: base, TTL: time.Hour})
s.Delete("voice")
if got := s.Get("voice", base); got != nil {
t.Fatalf("Delete left %+v", got)
}
}
func TestNewClarifyStoreDefaultTTL(t *testing.T) {
s := NewClarifyStore(0)
q := &PendingQuestion{Asked: base}
s.Put("voice", q)
if q.TTL != 90*time.Second {
t.Fatalf("TTL = %v, want 90s", q.TTL)
}
}
func TestAnswerFillsOnlyMissingSlots(t *testing.T) {
answerTime := base.Add(3 * time.Hour)
other := base.Add(9 * time.Hour)
cases := []struct {
name string
parked Slots
missing []Slot
text string
answer Slots
want Slots
}{
{
name: "fills the missing time",
parked: Slots{Text: "напомни позвонить"},
missing: []Slot{SlotTime},
text: "в три",
answer: Slots{Time: answerTime, HasTime: true},
want: Slots{Text: "напомни позвонить", Time: answerTime, HasTime: true},
},
{
name: "does not overwrite a filled time",
parked: Slots{Time: other, HasTime: true},
missing: []Slot{SlotTime},
text: "в три",
answer: Slots{Time: answerTime, HasTime: true},
want: Slots{Time: other, HasTime: true},
},
{
name: "ignores slots that were not missing",
parked: Slots{Key: "water", HasKey: true},
missing: []Slot{SlotValue},
text: "два литра",
answer: Slots{Key: "sleep", HasKey: true, Value: "2l"},
want: Slots{Key: "water", HasKey: true, Value: "2l"},
},
{
name: "fills key when empty",
parked: Slots{},
missing: []Slot{SlotKey, SlotValue},
text: "воды",
answer: Slots{Key: "water", HasKey: true, Value: `"drank"`},
want: Slots{Key: "water", HasKey: true, Value: `"drank"`},
},
{
name: "fills fn and its args",
parked: Slots{},
missing: []Slot{SlotFn},
text: "перезапусти nginx",
answer: Slots{Fn: "restart", Args: []string{"nginx"}, HasFn: true},
want: Slots{Fn: "restart", Args: []string{"nginx"}, HasFn: true},
},
{
name: "keeps existing args when fn was already known",
parked: Slots{Fn: "restart", Args: []string{"nginx"}, HasFn: true},
missing: []Slot{SlotFn},
text: "останови postgres",
answer: Slots{Fn: "stop", Args: []string{"postgres"}, HasFn: true},
want: Slots{Fn: "restart", Args: []string{"nginx"}, HasFn: true},
},
{
name: "raw answer becomes the text when nothing was parsed",
parked: Slots{},
missing: []Slot{SlotText},
text: "купить хлеб",
answer: Slots{},
want: Slots{Text: "купить хлеб"},
},
{
name: "parsed text wins over the raw answer",
parked: Slots{},
missing: []Slot{SlotText},
text: "запиши купить хлеб",
answer: Slots{Text: "купить хлеб"},
want: Slots{Text: "купить хлеб"},
},
{
name: "empty answer leaves the slot missing",
parked: Slots{Text: "напомни"},
missing: []Slot{SlotTime},
text: "не знаю",
answer: Slots{},
want: Slots{Text: "напомни"},
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
q := &PendingQuestion{Slots: tc.parked, Missing: tc.missing, Asked: base}
got := q.Answer(tc.text, tc.answer)
if got.Time != tc.want.Time || got.HasTime != tc.want.HasTime ||
got.Key != tc.want.Key || got.HasKey != tc.want.HasKey ||
got.Value != tc.want.Value || got.Text != tc.want.Text ||
got.Fn != tc.want.Fn || got.HasFn != tc.want.HasFn {
t.Fatalf("Answer = %+v, want %+v", got, tc.want)
}
if len(got.Args) != len(tc.want.Args) {
t.Fatalf("Args = %v, want %v", got.Args, tc.want.Args)
}
for i := range got.Args {
if got.Args[i] != tc.want.Args[i] {
t.Fatalf("Args = %v, want %v", got.Args, tc.want.Args)
}
}
})
}
}
func TestCanAskCapsAtOneQuestion(t *testing.T) {
if MaxAttempts != 1 {
t.Fatalf("MaxAttempts = %d, want 1 (Maven asks once, she is not a nag)", MaxAttempts)
}
q := &PendingQuestion{Asked: base}
if !q.CanAsk() {
t.Fatal("a fresh question should be askable")
}
q.Attempts = MaxAttempts
if q.CanAsk() {
t.Fatal("the question should not be asked twice")
}
}
func TestStillMissing(t *testing.T) {
want := []Slot{SlotTime, SlotKey, SlotValue, SlotFn, SlotText}
cases := []struct {
name string
slots Slots
want []Slot
}{
{"all empty", Slots{}, want},
{
name: "all filled",
slots: Slots{Time: base, HasTime: true, Key: "water", HasKey: true, Value: "1l", Fn: "restart", HasFn: true, Text: "t"},
want: nil,
},
{
name: "only value left",
slots: Slots{Time: base, HasTime: true, Key: "water", HasKey: true, Fn: "restart", HasFn: true, Text: "t"},
want: []Slot{SlotValue},
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
got := StillMissing(want, tc.slots)
if len(got) != len(tc.want) {
t.Fatalf("StillMissing = %v, want %v", got, tc.want)
}
for i := range got {
if got[i] != tc.want[i] {
t.Fatalf("StillMissing = %v, want %v", got, tc.want)
}
}
})
}
}
+4
View File
@@ -21,6 +21,7 @@ type Slots struct {
Time time.Time
HasTime bool
Key string
Value string // payload for a fact key, mirrors router.Slots.Value
HasKey bool
Text string
Fn string
@@ -104,6 +105,9 @@ func InheritSlots(prev, cur Slots) Slots {
out.Key = prev.Key
out.HasKey = true
}
if out.Value == "" && prev.Value != "" {
out.Value = prev.Value
}
if out.Text == "" && prev.Text != "" {
out.Text = prev.Text
}
+10
View File
@@ -113,4 +113,14 @@ func TestInheritSlots(t *testing.T) {
if inherited6.Text != "какая погода в москве" {
t.Error("should inherit text when current is empty")
}
prevValue := Slots{Key: "water", HasKey: true, Value: `"drank"`}
inherited7 := InheritSlots(prevValue, Slots{})
if inherited7.Value != `"drank"` {
t.Error("should inherit value when current is empty")
}
kept := InheritSlots(prevValue, Slots{Value: "2l"})
if kept.Value != "2l" {
t.Error("should keep current value")
}
}
+7
View File
@@ -237,6 +237,10 @@ type dismissProposedRoutineReq struct {
ID int64 `json:"id"`
}
type acceptProposedRoutineReq struct {
ID int64 `json:"id"`
}
// CoreAPI — what core exposes to modules. One Go interface, satisfied by:
// - the in-process store adapter (server.go storeAPI) — used by the daemon
// for modules that live in-process for now (router, delivery) and by tests,
@@ -286,6 +290,9 @@ type CoreAPI interface {
ListProposedRoutines(ctx context.Context) ([]ProposedRoutine, error)
// DismissProposedRoutine flips a proposed routine to 'dismissed'.
DismissProposedRoutine(ctx context.Context, id int64) error
// AcceptProposedRoutine flips a proposed routine to 'accepted'. The tick
// loop takes the schedule from there — no reminder is created (Vikunja #366).
AcceptProposedRoutine(ctx context.Context, id int64) error
// TickTrace returns the most recent tick's rule trace. The daemon caches
// this after every tick; the store adapter returns an error (trace is not
+4
View File
@@ -430,6 +430,10 @@ func (c *Client) DismissProposedRoutine(ctx context.Context, id int64) error {
return c.call(ctx, MethodDismissProposedRoutine, dismissProposedRoutineReq{ID: id}, nil)
}
func (c *Client) AcceptProposedRoutine(ctx context.Context, id int64) error {
return c.call(ctx, MethodAcceptProposedRoutine, acceptProposedRoutineReq{ID: id}, nil)
}
func (c *Client) Chat(ctx context.Context, text string) (string, error) {
var r chatResp
if err := c.call(ctx, MethodChat, chatReq{Text: text}, &r); err != nil {
+3
View File
@@ -478,6 +478,9 @@ func (a *chatTestAPI) DeleteTool(ctx context.Context, name string) error {
func (a *chatTestAPI) ListProposedRoutines(ctx context.Context) ([]ProposedRoutine, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) AcceptProposedRoutine(ctx context.Context, id int64) error {
return nil
}
func (a *chatTestAPI) DismissProposedRoutine(ctx context.Context, id int64) error {
return ErrUnknownMethod
}
+11
View File
@@ -253,6 +253,10 @@ func (a *storeAPI) DismissProposedRoutine(ctx context.Context, id int64) error {
return mapErr(a.s.DismissProposedRoutine(ctx, id))
}
func (a *storeAPI) AcceptProposedRoutine(ctx context.Context, id int64) error {
return mapErr(a.s.AcceptProposedRoutine(ctx, id, time.Now().UTC()))
}
func toTool(t store.Tool) Tool {
return Tool{
Name: t.Name, Scope: t.Scope, Cmd: t.Cmd, Destructive: t.Destructive,
@@ -774,6 +778,13 @@ func (s *Server) dispatch(ctx context.Context, req Request) (json.RawMessage, er
}
return marshalResult(nil), api.DismissProposedRoutine(ctx, p.ID)
case MethodAcceptProposedRoutine:
var p acceptProposedRoutineReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
return marshalResult(nil), api.AcceptProposedRoutine(ctx, p.ID)
case MethodRevertFact:
var p struct {
Key string `json:"key"`
+28 -27
View File
@@ -13,38 +13,39 @@ import (
type Method string
const (
MethodWriteFact Method = "write_fact"
MethodLatestFact Method = "latest_fact"
MethodLatestFactBySource Method = "latest_fact_by_source"
MethodSince Method = "since"
MethodPresence Method = "presence"
MethodCreateReminder Method = "create_reminder"
MethodMarkReminder Method = "mark_reminder"
MethodListReminders Method = "list_reminders"
MethodRecordNudge Method = "record_nudge"
MethodResolveNudge Method = "resolve_nudge"
MethodRecentOutcomes Method = "recent_outcomes"
MethodRecentFacts Method = "recent_facts"
MethodCalendarEvents Method = "calendar_events"
MethodRecentNudges Method = "recent_nudges"
MethodWriteNote Method = "write_note"
MethodQueryNotes Method = "query_notes"
MethodRecentNotes Method = "recent_notes"
MethodProposeTool Method = "propose_tool"
MethodEnableTool Method = "enable_tool"
MethodDisableTool Method = "disable_tool"
MethodAssertStepUp Method = "assert_stepup"
MethodStoreEncryptionKey Method = "store_encryption_key"
MethodUnlock Method = "unlock"
MethodLookupTool Method = "lookup_tool"
MethodWriteFact Method = "write_fact"
MethodLatestFact Method = "latest_fact"
MethodLatestFactBySource Method = "latest_fact_by_source"
MethodSince Method = "since"
MethodPresence Method = "presence"
MethodCreateReminder Method = "create_reminder"
MethodMarkReminder Method = "mark_reminder"
MethodListReminders Method = "list_reminders"
MethodRecordNudge Method = "record_nudge"
MethodResolveNudge Method = "resolve_nudge"
MethodRecentOutcomes Method = "recent_outcomes"
MethodRecentFacts Method = "recent_facts"
MethodCalendarEvents Method = "calendar_events"
MethodRecentNudges Method = "recent_nudges"
MethodWriteNote Method = "write_note"
MethodQueryNotes Method = "query_notes"
MethodRecentNotes Method = "recent_notes"
MethodProposeTool Method = "propose_tool"
MethodEnableTool Method = "enable_tool"
MethodDisableTool Method = "disable_tool"
MethodAssertStepUp Method = "assert_stepup"
MethodStoreEncryptionKey Method = "store_encryption_key"
MethodUnlock Method = "unlock"
MethodLookupTool Method = "lookup_tool"
MethodListTools Method = "list_tools"
MethodDeleteTool Method = "delete_tool"
MethodListProposedRoutines Method = "list_proposed_routines"
MethodDismissProposedRoutine Method = "dismiss_proposed_routine"
MethodAcceptProposedRoutine Method = "accept_proposed_routine"
MethodRevertFact Method = "revert_fact"
MethodTickTrace Method = "tick_trace"
MethodMorningStatus Method = "morning_status"
MethodChat Method = "chat"
MethodTickTrace Method = "tick_trace"
MethodMorningStatus Method = "morning_status"
MethodChat Method = "chat"
)
// Request — one frame from module to core. Params is the JSON-encoded argument
+1 -1
View File
@@ -19,7 +19,7 @@ func TestComplete(t *testing.T) {
t.Errorf("path = %q, want /v1/chat/completions", r.URL.Path)
}
var reqBody struct {
Messages []struct {
Messages []struct {
Role string `json:"role"`
Content string `json:"content"`
} `json:"messages"`
-2
View File
@@ -273,7 +273,6 @@ func TestRemindersBypassEverySuppressor(t *testing.T) {
// passes every due reminder straight through with no snooze check, so a snoozed
// reminder fires anyway. The test below is what the contract asks for.
func TestRemindersStillHonourSnooze(t *testing.T) {
t.Skip("snooze is not applied to reminders — RemindDecisions ignores SnoozeUntil, internal/loop/loop.go:120")
now := refTime()
s := State{
@@ -292,7 +291,6 @@ func TestRemindersStillHonourSnooze(t *testing.T) {
// unit tests above pass while nothing can ever populate the map. This asserts
// the Gatherer actually produces a snooze map.
func TestGathererPopulatesSnoozeUntil(t *testing.T) {
t.Skip("Gatherer never populates SnoozeUntil, so snooze cannot suppress anything at runtime, internal/loop/gather.go:153")
ctx := context.Background()
st, err := store.Open(ctx, t.TempDir()+"/m.db")
+9 -1
View File
@@ -119,6 +119,14 @@ func (g *Gatherer) GatherState(ctx context.Context, now time.Time) (State, []sto
return State{}, nil, err
}
// live snoozes — "leave me alone until X", per rule. The `snoozed` outcome
// on the nudges table is the whole record; the store turns it into an
// expiry. Absent rules mean "not snoozed", which is what the gate reads.
snoozeUntil, err := g.store.SnoozedUntil(ctx, now)
if err != nil {
return State{}, nil, err
}
// env flags — QuietHours / CalendarBusy as config facts.
// QuietHours: presence != reachability, sleep/quiet-hours handled separately
// in the gate. We read a config `quiet_hours` fact for the boolean.
@@ -150,7 +158,7 @@ func (g *Gatherer) GatherState(ctx context.Context, now time.Time) (State, []sto
PresenceScore: score,
Facts: facts,
LastNudge: lastNudge,
SnoozeUntil: nil, // no snooze persistence yet — daemon wires in
SnoozeUntil: snoozeUntil,
CooldownUntil: cooldownUntil,
QuietHours: quiet,
CalendarBusy: calBusy,
+35 -6
View File
@@ -1,6 +1,7 @@
package loop
import (
"fmt"
"time"
"github.com/kami/maven/internal/store"
@@ -106,25 +107,53 @@ func Tick(s State, rules []Rule) *Candidate {
// ReminderDecision — a due reminder the daemon should deliver now.
// NOT gated by the universal Gate (per spec: "wake me 7" fires in quiet hours;
// that's the point). Snooze still applies — represented by a separate
// snooze-until the gatherer consults; for the scaffold, fired-reminders move
// straight to MarkReminder(fired).
// that's the point). Snooze is the one part of restraint that still applies.
type ReminderDecision struct {
Reminder store.Reminder
State State
}
// RemindDecisions — returns all due reminders (without gating their delivery
// by restraint). Pure: accepts an already-filtered (due) list. The Gatherer
// produces that list from `fire_ts <= now AND pending`.
// ReminderSnoozeKey — the SnoozeUntil key that holds back every due reminder.
// Reminders have no rule name, so they share one key. A snooze aimed at a
// single reminder uses ReminderSnoozeKeyFor instead.
const ReminderSnoozeKey = "reminder"
// ReminderSnoozeKeyFor — the SnoozeUntil key for one reminder by id.
func ReminderSnoozeKeyFor(id int64) string {
return fmt.Sprintf("%s:%d", ReminderSnoozeKey, id)
}
// RemindDecisions — returns the due reminders the daemon should deliver.
// Pure: accepts an already-filtered (due) list. The Gatherer produces that
// list from `fire_ts <= now AND pending`.
//
// Quiet hours, presence and cooldown are deliberately NOT consulted — a
// reminder must wake you at 7 even in the middle of quiet hours. Only snooze
// holds one back. A held reminder stays pending, so it comes back once the
// snooze runs out.
func RemindDecisions(s State, due []store.Reminder) []ReminderDecision {
out := make([]ReminderDecision, 0, len(due))
for _, r := range due {
if reminderSnoozed(s, r) {
continue
}
out = append(out, ReminderDecision{Reminder: r, State: s})
}
return out
}
// reminderSnoozed — true when a snooze on this reminder, or on reminders as a
// class, is still running.
func reminderSnoozed(s State, r store.Reminder) bool {
keys := []string{ReminderSnoozeKey, ReminderSnoozeKeyFor(r.ID)}
for _, k := range keys {
if until, ok := s.SnoozeUntil[k]; ok && s.Now.Before(until) {
return true
}
}
return false
}
// CooldownFor — helper for the Gatherer: given the active cooldown base
// (the rule's static Base, OR the feedback tuner's persisted tuning) and the
// last send ts, compute the wall-clock "cooldown-until" the gate will check.
+38
View File
@@ -0,0 +1,38 @@
package memory
// Confidence gate for a recall. Two checks, both must pass before Maven says a
// note back:
//
// - minScore — an absolute cosine floor.
// - minMargin — the top hit must beat the runner-up by more than this.
//
// The margin is the one that carries the weight. The e5 embedder packs every
// score into a narrow high band (0.79-0.89 on the recall fixture), so an
// absolute floor cannot tell a real hit from a confident-looking miss: every
// value under the band admits everything, every value above it answers nothing.
// A margin asks a different question — "is this note clearly the best one, or
// is the whole shelf equally close?" — and a made-up question has no clear best.
//
// With one hit and no runner-up there is nothing to compare, so only the floor
// applies.
// ConfidentScores reports whether the top score clears both gates. scores must
// be sorted highest first. minMargin <= 0 turns the margin check off.
func ConfidentScores(scores []float64, minScore, minMargin float64) bool {
if len(scores) == 0 || scores[0] < minScore {
return false
}
if minMargin > 0 && len(scores) > 1 && scores[0]-scores[1] <= minMargin {
return false
}
return true
}
// Confident is ConfidentScores for search results.
func Confident(results []Result, minScore, minMargin float64) bool {
scores := make([]float64, len(results))
for i, r := range results {
scores[i] = r.Score
}
return ConfidentScores(scores, minScore, minMargin)
}
+45
View File
@@ -0,0 +1,45 @@
package memory
import "testing"
func TestConfidentScores(t *testing.T) {
cases := []struct {
name string
scores []float64
minScore float64
minMargin float64
want bool
}{
{"no hits", nil, 0.55, 0.008, false},
{"below the floor", []float64{0.40, 0.10}, 0.55, 0.008, false},
{"clear winner", []float64{0.86, 0.70}, 0.55, 0.008, true},
{"runner-up too close", []float64{0.860, 0.858}, 0.55, 0.008, false},
// The rule is "beats the runner-up by MORE than delta". Not testing an
// exactly-equal margin: no pair of these decimals subtracts to exactly
// 0.008 in binary float, so such a test would pin rounding, not the rule.
{"margin just under delta", []float64{0.8079, 0.8}, 0.55, 0.008, false},
{"margin just over delta", []float64{0.8081, 0.8}, 0.55, 0.008, true},
// One hit: nothing to compare against, so only the floor applies.
{"single hit clears", []float64{0.86}, 0.55, 0.008, true},
{"single hit below floor", []float64{0.10}, 0.55, 0.008, false},
// Margin off — the old absolute-only behaviour.
{"margin off admits a tie", []float64{0.86, 0.86}, 0.55, 0, true},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
if got := ConfidentScores(c.scores, c.minScore, c.minMargin); got != c.want {
t.Errorf("got %v, want %v", got, c.want)
}
})
}
}
func TestConfidentReadsResultScores(t *testing.T) {
res := []Result{{ID: "a", Score: 0.86}, {ID: "b", Score: 0.858}}
if Confident(res, 0.55, 0.008) {
t.Error("thin margin passed the gate")
}
if !Confident(res, 0.55, 0) {
t.Error("margin off should fall back to the floor alone")
}
}
+517
View File
@@ -0,0 +1,517 @@
// Package recalleval is the held-out contract for note recall: can Maven find
// the right note again when the user asks for it weeks later?
//
// Why it sits beside internal/memory rather than inside it: the thing under
// test is a whole path, not one function — an embedder (internal/router), a
// vector store (internal/memory or internal/store) and the confidence gate the
// daemon applies on top (cmd/mavend/recall.go's bestRecall, config's
// query_min_score). A _test.go file inside internal/memory could not reach the
// persistent store without an import cycle, and testdata is not reachable from
// another package's working directory — so the fixture is embedded here and the
// scorer takes the store as a factory. Same layout and same reasons as
// internal/router/eval.
//
// The fixture is HELD OUT the same way the routing fixture is: a query never
// repeats its note's wording verbatim beyond ordinary shared vocabulary, and
// TestFixtureIsParaphrased enforces a floor on how little the two overlap.
// Scoring recall on a query that is a copy of the note measures string
// matching, not recall.
package recalleval
import (
"context"
_ "embed"
"encoding/json"
"fmt"
"sort"
"strings"
"time"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/router"
)
//go:embed ru_recall_v1.json
var fixtureJSON []byte
// SchemaVersion — the version this package understands. The loader refuses any
// other version rather than misreading a fixture and reporting a number.
const SchemaVersion = 1
// StoredNote — one thing the user said once, as it lands in the semantic store.
// Kind is "note" or "fact"; both share the vector index (see
// cmd/mavend/recall.go), so a fact can legitimately win a recall.
type StoredNote struct {
ID string `json:"id"`
Text string `json:"text"`
Kind string `json:"kind"`
}
// Case — a small set of notes, one query, and the note that must come back
// first. Want is empty exactly when the query should recall NOTHING: that lane
// measures false recall, which is the direction the spec cares about ("a
// confident wrong fact is worse than a known gap").
type Case struct {
ID string `json:"id"`
Lang string `json:"lang"`
Notes []StoredNote `json:"notes"`
Query string `json:"query"`
Want string `json:"want"`
Tags []string `json:"tags"`
Note string `json:"note"`
}
// Answerable reports whether the case expects a recall at all.
func (c Case) Answerable() bool { return c.Want != "" }
// Fixture — the versioned envelope, same shape as the routing fixture.
//
// Filler is inserted into EVERY case's store on top of that case's own notes.
// Without it a case with three notes scores recall@3 = 100% by construction,
// which measures nothing. A real store holds months of unrelated notes, and the
// wanted note has to beat all of them.
type Fixture struct {
SchemaVersion int `json:"schema_version"`
Name string `json:"name"`
Notes []string `json:"notes"`
Filler []StoredNote `json:"filler"`
Cases []Case `json:"cases"`
}
// Load returns the embedded fixture.
func Load() (Fixture, error) {
var f Fixture
if err := json.Unmarshal(fixtureJSON, &f); err != nil {
return Fixture{}, fmt.Errorf("parse fixture: %w", err)
}
if f.SchemaVersion != SchemaVersion {
return Fixture{}, fmt.Errorf("fixture schema_version %d, want %d", f.SchemaVersion, SchemaVersion)
}
if len(f.Cases) == 0 {
return Fixture{}, fmt.Errorf("fixture has no cases")
}
return f, nil
}
// NewStore builds an empty store for one case, plus a function to release it.
// A factory rather than a store because every case needs a clean index — notes
// from case A must not be visible to case B's query.
type NewStore func() (memory.Store, func(), error)
// InMemory is the NewStore for memory.InMemoryStore — the fallback the daemon
// uses when there is no database (cmd/mavend/voice.go:224).
func InMemory() (memory.Store, func(), error) {
return memory.NewInMemoryStore(), func() {}, nil
}
// Cache wraps an embedder so repeated text is embedded once. The gate sweep
// scores the same fixture at nine thresholds, and every case re-inserts the
// filler notes — without this the ONNX run spends minutes re-embedding
// identical strings. Latency numbers come from the uncached run.
func Cache(inner router.Embedder) router.Embedder {
return &cachingEmbedder{inner: inner, seen: map[string][]float32{}}
}
type cachingEmbedder struct {
inner router.Embedder
seen map[string][]float32
}
var _ router.AsymmetricEmbedder = (*cachingEmbedder)(nil)
func (c *cachingEmbedder) Dim() int { return c.inner.Dim() }
func (c *cachingEmbedder) Close() error { return nil } // the caller owns inner
func (c *cachingEmbedder) Embed(ctx context.Context, text string) ([]float32, error) {
return c.cached(ctx, "embed:"+text, func() ([]float32, error) {
return c.inner.Embed(ctx, text)
})
}
// The two sides of an asymmetric embedder give different vectors for the same
// string, so the cache key has to say which side asked.
func (c *cachingEmbedder) EmbedQuery(ctx context.Context, text string) ([]float32, error) {
return c.cached(ctx, "query:"+text, func() ([]float32, error) {
return router.EmbedQuery(ctx, c.inner, text)
})
}
func (c *cachingEmbedder) EmbedPassage(ctx context.Context, text string) ([]float32, error) {
return c.cached(ctx, "passage:"+text, func() ([]float32, error) {
return router.EmbedPassage(ctx, c.inner, text)
})
}
func (c *cachingEmbedder) cached(_ context.Context, key string, embed func() ([]float32, error)) ([]float32, error) {
if v, ok := c.seen[key]; ok {
return v, nil
}
v, err := embed()
if err != nil {
return nil, err
}
c.seen[key] = v
return v, nil
}
// Outcome — one scored case.
type Outcome struct {
Case Case
Hits []memory.Result
Err error
// Latency is the read path only: embed the query, then Search. Insert time
// is excluded because it happens once, weeks earlier.
Latency time.Duration
// Rank1/Rank3 — the wanted note came back first / in the top three,
// ignoring the confidence gate. Ranking is the store's job.
Rank1 bool
Rank3 bool
// Recalled — what the daemon would actually say back: the top hit's text
// when it clears the gate. Mirrors bestRecall in cmd/mavend/recall.go.
Recalled string
// Pass — the wanted note was recalled AND survived the gate; or, for a
// no-answer case, nothing was recalled.
Pass bool
// Tied — the wanted note is on top but shares its score with the next hit,
// so the sort decided it, not the embedder. Counted apart from a real hit.
Tied bool
TopID string
TopScor float64
// Margin — top1 top2. 0 when fewer than two hits came back.
Margin float64
Reasons []string
}
// Report — the aggregate. Rank and gate are kept apart on purpose: a note that
// ranks first but is silenced by query_min_score is a threshold problem, and a
// note that never ranks first is an embedder problem. Those are different fixes.
type Report struct {
Name string
MinScore float64
// MinMargin — how far the top hit must beat the runner-up. 0 ⇒ off.
MinMargin float64
Total int
Answerable int
Rank1 int
Rank3 int
// Gated — ranked first but the score was under MinScore, so the daemon
// stays silent and answers "не знаю".
Gated int
// WrongTop — a different note outranked the right one.
WrongTop int
// Tied — the right note was on top only because of sort order. Not credited
// as recall; tracked because it is a distinct failure (the embedder scored
// the query and the note the same as everything else).
Tied int
// NoAnswer / FalseRecall — the cases that must recall nothing, and how many
// of them the daemon would answer anyway.
NoAnswer int
FalseRecall int
Errors int
Passed int
Outcomes []Outcome
ByTag map[string]TagStat
ByLang map[string]TagStat
// CorrectTop / NoAnswerTop — sorted top-1 scores for the answerable cases
// where the right note ranked first, and for the no-answer cases. The gap
// between these two distributions is what a defensible query_min_score
// would have to sit inside; if they overlap, no threshold separates them.
CorrectTop []float64
NoAnswerTop []float64
// CorrectMargin / NoAnswerMargin — the same two groups, but top1 top2
// instead of top1. This is the pair the margin gate has to separate, and
// unlike the absolute scores it is what the sweep reads.
CorrectMargin []float64
NoAnswerMargin []float64
P50, P95, Max time.Duration
}
// TagStat — passed/total for one slice of the fixture.
type TagStat struct{ Passed, Total int }
// Recall1 — fraction of answerable cases whose wanted note ranked first.
func (r Report) Recall1() float64 { return ratio(r.Rank1, r.Answerable) }
// Recall3 — same, in the top three. The daemon asks for 3 (voice.go), so this
// is the ceiling a better gate or a reranker could reach.
func (r Report) Recall3() float64 { return ratio(r.Rank3, r.Answerable) }
// Answered — fraction of answerable cases the daemon would actually answer
// correctly, gate included. This is the number the operator experiences.
func (r Report) Answered() float64 { return ratio(r.Rank1-r.Gated, r.Answerable) }
// FalseRecallRate — fraction of the no-answer cases the daemon answers anyway.
func (r Report) FalseRecallRate() float64 { return ratio(r.FalseRecall, r.NoAnswer) }
func ratio(n, d int) float64 {
if d == 0 {
return 0
}
return float64(n) / float64(d)
}
// Score runs every case against a fresh store and aggregates. It never fails
// the run on an embed or search error: an erroring case scores as a miss and is
// counted in Errors, because "the embedder was down" and "the embedder was
// wrong" are different numbers.
func Score(ctx context.Context, name string, emb router.Embedder, newStore NewStore, minScore, minMargin float64, f Fixture) (Report, error) {
rep := Report{
Name: name,
MinScore: minScore,
MinMargin: minMargin,
Total: len(f.Cases),
ByTag: map[string]TagStat{},
ByLang: map[string]TagStat{},
}
lat := make([]time.Duration, 0, len(f.Cases))
for _, c := range f.Cases {
if c.Answerable() {
rep.Answerable++
} else {
rep.NoAnswer++
}
o, err := scoreCase(ctx, emb, newStore, minScore, minMargin, c, f.Filler)
if err != nil {
return Report{}, err
}
lat = append(lat, o.Latency)
switch {
case o.Err != nil:
rep.Errors++
case c.Answerable():
if o.Rank1 {
rep.Rank1++
}
if o.Rank3 {
rep.Rank3++
}
if o.Rank1 && o.Recalled == "" {
rep.Gated++
}
if o.Tied {
rep.Tied++
} else if !o.Rank1 {
rep.WrongTop++
}
if o.Rank1 {
// Margins are collected on rank, not on the gate, so the
// distribution does not move as the sweep changes the gate.
rep.CorrectMargin = append(rep.CorrectMargin, o.Margin)
}
if o.Rank1 && o.Recalled != "" {
rep.CorrectTop = append(rep.CorrectTop, o.TopScor)
}
default:
if o.Recalled != "" {
rep.FalseRecall++
}
rep.NoAnswerTop = append(rep.NoAnswerTop, o.TopScor)
rep.NoAnswerMargin = append(rep.NoAnswerMargin, o.Margin)
}
if o.Pass {
rep.Passed++
}
bump(rep.ByLang, c.Lang, o.Pass)
for _, tag := range c.Tags {
bump(rep.ByTag, tag, o.Pass)
}
rep.Outcomes = append(rep.Outcomes, o)
}
sort.Float64s(rep.CorrectTop)
sort.Float64s(rep.NoAnswerTop)
sort.Float64s(rep.CorrectMargin)
sort.Float64s(rep.NoAnswerMargin)
sort.Slice(lat, func(i, j int) bool { return lat[i] < lat[j] })
rep.P50, rep.P95 = percentile(lat, 0.50), percentile(lat, 0.95)
if len(lat) > 0 {
rep.Max = lat[len(lat)-1]
}
return rep, nil
}
// scoreCase inserts the case's notes into a fresh store, then runs the read
// path the daemon runs. The returned error is fatal (the harness is broken);
// an embedder or store failure on the query lands in Outcome.Err instead.
func scoreCase(ctx context.Context, emb router.Embedder, newStore NewStore, minScore, minMargin float64, c Case, filler []StoredNote) (Outcome, error) {
st, release, err := newStore()
if err != nil {
return Outcome{}, fmt.Errorf("%s: new store: %w", c.ID, err)
}
defer release()
all := append(append([]StoredNote(nil), c.Notes...), filler...)
for _, n := range all {
vec, err := router.EmbedPassage(ctx, emb, n.Text)
if err != nil {
return Outcome{}, fmt.Errorf("%s: embed note %s: %w", c.ID, n.ID, err)
}
meta := map[string]string{"text": n.Text, "type": n.Kind}
if err := st.Insert(ctx, n.ID, vec, meta); err != nil {
return Outcome{}, fmt.Errorf("%s: insert %s: %w", c.ID, n.ID, err)
}
}
o := Outcome{Case: c}
start := time.Now()
qvec, err := router.EmbedQuery(ctx, emb, c.Query)
if err != nil {
o.Latency = time.Since(start)
o.Err = err
o.Reasons = []string{fmt.Sprintf("embed query: %v", err)}
return o, nil
}
hits, err := st.Search(ctx, qvec, 3)
o.Latency = time.Since(start)
if err != nil {
o.Err = err
o.Reasons = []string{fmt.Sprintf("search: %v", err)}
return o, nil
}
o.Hits = hits
if len(hits) > 0 {
o.TopID, o.TopScor = hits[0].ID, hits[0].Score
if len(hits) > 1 {
o.Margin = hits[0].Score - hits[1].Score
}
o.Recalled = bestRecall(hits, minScore, minMargin)
}
for i, h := range hits {
if h.ID != c.Want {
continue
}
o.Rank3 = true
// A tie is not a hit. With a lexical embedder several notes score
// exactly 0 against a paraphrased query, and whichever one the sort
// happens to leave on top would otherwise be credited as recall.
if i == 0 && (len(hits) < 2 || hits[0].Score > hits[1].Score) {
o.Rank1 = true
}
if i == 0 && !o.Rank1 {
o.Tied = true
}
}
switch {
case !c.Answerable():
if o.Recalled != "" {
o.Reasons = append(o.Reasons, fmt.Sprintf("false recall: %q at %.3f (margin %.3f), want silence", o.TopID, o.TopScor, o.Margin))
}
case o.Tied:
o.Reasons = append(o.Reasons, fmt.Sprintf("tie at %.3f — the right note is on top only by sort order", o.TopScor))
case !o.Rank1:
o.Reasons = append(o.Reasons, fmt.Sprintf("top hit %q (%.3f), want %q%s", o.TopID, o.TopScor, c.Want, rankNote(o.Rank3)))
case o.Recalled == "":
o.Reasons = append(o.Reasons, fmt.Sprintf("right note ranked first at %.3f (margin %.3f) but the gate silenced it — daemon says \"не знаю\"", o.TopScor, o.Margin))
}
o.Pass = len(o.Reasons) == 0
return o, nil
}
func rankNote(inTop3 bool) string {
if inTop3 {
return " (wanted note is in the top 3)"
}
return " (wanted note is not in the top 3)"
}
// bestRecall mirrors cmd/mavend/recall.go — the gate the daemon actually
// applies to a memory hit. Duplicated rather than imported because package main
// is not importable; recalleval_test.go asserts the two agree in behaviour.
func bestRecall(results []memory.Result, minScore, minMargin float64) string {
if !memory.Confident(results, minScore, minMargin) {
return ""
}
return results[0].Meta["text"]
}
func bump(m map[string]TagStat, key string, pass bool) {
if key == "" {
return
}
s := m[key]
s.Total++
if pass {
s.Passed++
}
m[key] = s
}
// percentile — nearest-rank on a pre-sorted slice. No interpolation: with ~30
// samples an interpolated p95 invents a latency no query actually took.
func percentile(sorted []time.Duration, p float64) time.Duration {
if len(sorted) == 0 {
return 0
}
i := int(p * float64(len(sorted)))
if i >= len(sorted) {
i = len(sorted) - 1
}
return sorted[i]
}
// String renders the report in the routing eval's style — headline first, then
// the slices that name where the path is weak.
func (r Report) String() string {
var b strings.Builder
fmt.Fprintf(&b, "%s: %d/%d cases pass (gate %.2f, margin %.3f)\n", r.Name, r.Passed, r.Total, r.MinScore, r.MinMargin)
fmt.Fprintf(&b, " recall@1 %.1f%% (%d/%d) recall@3 %.1f%% (%d/%d) answered after gate %.1f%% (%d/%d)\n",
100*r.Recall1(), r.Rank1, r.Answerable,
100*r.Recall3(), r.Rank3, r.Answerable,
100*r.Answered(), r.Rank1-r.Gated, r.Answerable)
fmt.Fprintf(&b, " wrong note on top: %d | tie on top (sort order, not recall): %d | silenced by gate: %d | errors: %d\n",
r.WrongTop, r.Tied, r.Gated, r.Errors)
fmt.Fprintf(&b, " false recall %.1f%% (%d/%d must-be-silent cases answered anyway)\n",
100*r.FalseRecallRate(), r.FalseRecall, r.NoAnswer)
fmt.Fprintf(&b, " top-1 score, right note first: %s\n", spread(r.CorrectTop))
fmt.Fprintf(&b, " top-1 score, must be silent: %s\n", spread(r.NoAnswerTop))
fmt.Fprintf(&b, " margin top1-top2, right note first: %s\n", spread(r.CorrectMargin))
fmt.Fprintf(&b, " margin top1-top2, must be silent: %s\n", spread(r.NoAnswerMargin))
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
fmt.Fprintf(&b, " by lang: %s\n", renderStats(r.ByLang))
fmt.Fprintf(&b, " by tag: %s\n", renderStats(r.ByTag))
return b.String()
}
// Failures — per-case detail, sorted by ID so two runs diff cleanly.
func (r Report) Failures() string {
var b strings.Builder
out := append([]Outcome(nil), r.Outcomes...)
sort.Slice(out, func(i, j int) bool { return out[i].Case.ID < out[j].Case.ID })
for _, o := range out {
if o.Pass {
continue
}
fmt.Fprintf(&b, " %s %q: %s\n", o.Case.ID, o.Case.Query, strings.Join(o.Reasons, "; "))
}
return b.String()
}
// spread — min / median / max of a sorted score list. Three numbers is enough
// to see whether two distributions overlap, which is the only question a
// threshold can answer.
func spread(sorted []float64) string {
if len(sorted) == 0 {
return "n/a"
}
return fmt.Sprintf("min %.3f median %.3f max %.3f (n=%d)",
sorted[0], sorted[len(sorted)/2], sorted[len(sorted)-1], len(sorted))
}
func renderStats(m map[string]TagStat) string {
keys := make([]string, 0, len(m))
for k := range m {
keys = append(keys, k)
}
sort.Strings(keys)
parts := make([]string, 0, len(keys))
for _, k := range keys {
s := m[k]
parts = append(parts, fmt.Sprintf("%s %d/%d", k, s.Passed, s.Total))
}
return strings.Join(parts, " ")
}
@@ -0,0 +1,331 @@
package recalleval
import (
"context"
"fmt"
"os"
"path/filepath"
"strings"
"testing"
"unicode"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
)
// dim 1024 for the hash embedder: it is bag-of-words, so a narrower space
// collides tokens between unrelated notes and would measure the hash.
const hashDim = 1024
func TestLoadFixture(t *testing.T) {
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
if len(f.Cases) < 25 {
t.Errorf("%d cases, want >= 25", len(f.Cases))
}
// Filler is what stops recall@3 being free: three case notes and a top-3
// search would put the wanted note in the top 3 every time.
if len(f.Filler) < 10 {
t.Errorf("%d filler notes, want >= 10", len(f.Filler))
}
seen := map[string]bool{}
silent, en := 0, 0
for _, c := range f.Cases {
if c.ID == "" || seen[c.ID] {
t.Errorf("case %q: empty or duplicate id", c.ID)
}
seen[c.ID] = true
if c.Lang != "ru" && c.Lang != "en" {
t.Errorf("%s: lang %q, want ru|en", c.ID, c.Lang)
}
if c.Lang == "en" {
en++
}
if strings.TrimSpace(c.Query) == "" {
t.Errorf("%s: empty query", c.ID)
}
// Fewer than three notes and a wrong answer has nowhere to come from,
// so recall@1 would be near-free.
if len(c.Notes) < 3 {
t.Errorf("%s: %d notes, want >= 3", c.ID, len(c.Notes))
}
ids := map[string]bool{}
for _, n := range c.Notes {
if n.ID == "" || ids[n.ID] {
t.Errorf("%s: note %q empty or duplicate id", c.ID, n.ID)
}
ids[n.ID] = true
if strings.TrimSpace(n.Text) == "" {
t.Errorf("%s: note %q empty text", c.ID, n.ID)
}
if n.Kind != "note" && n.Kind != "fact" {
t.Errorf("%s: note %q kind %q, want note|fact", c.ID, n.ID, n.Kind)
}
}
if !c.Answerable() {
silent++
continue
}
if !ids[c.Want] {
t.Errorf("%s: want %q is not one of the case's notes", c.ID, c.Want)
}
}
// Both lanes need enough cases that a rate means something.
if silent < 5 {
t.Errorf("%d must-be-silent cases, want >= 5", silent)
}
if en < 5 {
t.Errorf("%d English cases, want >= 5", en)
}
}
// TestFixtureIsParaphrased — the fixture's claim to measuring recall at all. If
// a query repeats its note's words, cosine over a bag-of-words embedder gets it
// for free and the score says nothing about semantic recall. Half the query's
// words is the line: some shared vocabulary is natural ("nginx", "чай"), a copy
// is the failure.
func TestFixtureIsParaphrased(t *testing.T) {
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
for _, c := range f.Cases {
if !c.Answerable() {
continue
}
var want string
for _, n := range c.Notes {
if n.ID == c.Want {
want = n.Text
}
}
q := words(c.Query)
if len(q) == 0 {
continue
}
inNote := map[string]bool{}
for _, w := range words(want) {
inNote[w] = true
}
shared := 0
for _, w := range q {
if inNote[w] {
shared++
}
}
if frac := float64(shared) / float64(len(q)); frac > 0.5 {
t.Errorf("%s: query shares %.0f%% of its words with the note — not a paraphrase\n query: %q\n note: %q",
c.ID, 100*frac, c.Query, want)
}
}
}
// words — lowercased words of two runes or more, matching how the hash
// embedder tokenizes.
func words(s string) []string {
var out []string
for _, w := range strings.FieldsFunc(strings.ToLower(s), func(r rune) bool {
return !unicode.IsLetter(r) && !unicode.IsDigit(r)
}) {
if len([]rune(w)) > 1 {
out = append(out, w)
}
}
return out
}
// TestBestRecallMatchesDaemon — the harness duplicates bestRecall from
// cmd/mavend/recall.go (package main is not importable). This pins the copy to
// the original's three rules: no hits, below the gate, or no text ⇒ silence.
func TestBestRecallMatchesDaemon(t *testing.T) {
if got := bestRecall(nil, 0.55, 0); got != "" {
t.Errorf("no hits: got %q, want silence", got)
}
low := []memory.Result{{ID: "a", Score: 0.4, Meta: map[string]string{"text": "чай"}}}
if got := bestRecall(low, 0.55, 0); got != "" {
t.Errorf("below gate: got %q, want silence", got)
}
noText := []memory.Result{{ID: "a", Score: 0.9, Meta: map[string]string{}}}
if got := bestRecall(noText, 0.55, 0); got != "" {
t.Errorf("no text: got %q, want silence", got)
}
ok := []memory.Result{{ID: "a", Score: 0.9, Meta: map[string]string{"text": "чай"}}}
if got := bestRecall(ok, 0.55, 0); got != "чай" {
t.Errorf("above gate: got %q, want %q", got, "чай")
}
// Margin: a close runner-up means the embedder cannot tell the two apart,
// so Maven stays silent even though both clear the absolute floor.
close := []memory.Result{
{ID: "a", Score: 0.86, Meta: map[string]string{"text": "чай"}},
{ID: "b", Score: 0.85, Meta: map[string]string{"text": "кофе"}},
}
if got := bestRecall(close, 0.55, 0.03); got != "" {
t.Errorf("thin margin: got %q, want silence", got)
}
if got := bestRecall(close, 0.55, 0); got != "чай" {
t.Errorf("margin off: got %q, want %q", got, "чай")
}
clear := []memory.Result{
{ID: "a", Score: 0.86, Meta: map[string]string{"text": "чай"}},
{ID: "b", Score: 0.70, Meta: map[string]string{"text": "кофе"}},
}
if got := bestRecall(clear, 0.55, 0.03); got != "чай" {
t.Errorf("wide margin: got %q, want %q", got, "чай")
}
}
// TestHashRecallBaseline — the CI ratchet. HashEmbedder, so it needs no model
// files and is byte-for-byte reproducible.
//
// It is a floor, not a target. The hash embedder is lexical, so most of this
// fixture is unwinnable for it by construction; the number worth moving is
// TestONNXRecall's. Never compare a hash-embedder number to an ONNX one.
func TestHashRecallBaseline(t *testing.T) {
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
rep, err := Score(context.Background(), "recall+hash", router.NewHashEmbedder(hashDim), InMemory,
config.DefaultQueryMinScore, config.DefaultQueryMinMargin, f)
if err != nil {
t.Fatalf("Score: %v", err)
}
t.Log("\n" + rep.String() + rep.Failures())
t.Log("\ngate sweep:\n" + sweep(t, router.NewHashEmbedder(hashDim), f))
// 0.32 sits under the observed 0.360 recall@1.
const floorRecall1 = 0.32
if rep.Recall1() < floorRecall1 {
t.Errorf("recall@1 %.3f below ratchet %.2f — note recall regressed", rep.Recall1(), floorRecall1)
}
// The dangerous direction, asserted tightly and separately: answering from
// the wrong note is worse than a gap. Observed 0 under the hash floor.
if rep.FalseRecall > 1 {
t.Errorf("%d false recalls, want <= 1:\n%s", rep.FalseRecall, rep.Failures())
}
}
// TestPersistentStoreScoresTheSame — the deployed store is sqlite-backed
// (store.MemoryStore via st.VectorMemory()), not the in-memory fallback. Its
// Search is a separate implementation of the same cosine scan, so it gets its
// own run: a divergence here would mean recall quality depends on whether a
// database was configured.
func TestPersistentStoreScoresTheSame(t *testing.T) {
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
emb := router.NewHashEmbedder(hashDim)
inMem, err := Score(context.Background(), "recall+hash+memory", emb, InMemory, config.DefaultQueryMinScore, config.DefaultQueryMinMargin, f)
if err != nil {
t.Fatalf("Score in-memory: %v", err)
}
persistent, err := Score(context.Background(), "recall+hash+sqlite", emb, sqliteStores(t), config.DefaultQueryMinScore, config.DefaultQueryMinMargin, f)
if err != nil {
t.Fatalf("Score sqlite: %v", err)
}
t.Log("\n" + persistent.String())
if persistent.Rank1 != inMem.Rank1 || persistent.FalseRecall != inMem.FalseRecall {
t.Errorf("sqlite recall@1 %d/%d fr %d, in-memory %d/%d fr %d — the two backends disagree",
persistent.Rank1, persistent.Answerable, persistent.FalseRecall,
inMem.Rank1, inMem.Answerable, inMem.FalseRecall)
}
}
// sqliteStores returns a NewStore that hands each case its own plaintext
// database file, so cases stay isolated the way they are with InMemory.
func sqliteStores(t *testing.T) NewStore {
t.Helper()
dir := t.TempDir()
n := 0
return func() (memory.Store, func(), error) {
n++
st, err := store.Open(context.Background(), filepath.Join(dir, fmt.Sprintf("recall-%d.db", n)))
if err != nil {
return nil, nil, err
}
return st.VectorMemory(), func() { _ = st.Close() }, nil
}
}
// TestONNXRecall — the number that matters: the multilingual embedder homesrv
// actually runs. Opt-in via MAVEN_ONNX_LIB because deps/ is gitignored, exactly
// like TestONNXBaseline in internal/router/eval. `make eval-recall` points it at
// the vendored runtime.
//
// Reports rather than asserts. The gate sweep is the point: it prints
// answered-vs-false-recall at a range of query_min_score values, so the right
// threshold is read off data instead of guessed.
func TestONNXRecall(t *testing.T) {
lib := os.Getenv("MAVEN_ONNX_LIB")
if lib == "" {
t.Skip("MAVEN_ONNX_LIB unset — see AGENTS.md § Embedder model for intent routing")
}
model := filepath.Join("../../..", "models/embedder/multilingual-e5-small/model_quantized.onnx")
tok := filepath.Join("../../..", "models/embedder/multilingual-e5-small/tokenizer.json")
for _, p := range []string{lib, model, tok} {
if _, err := os.Stat(p); err != nil {
t.Skipf("missing %s: %v", p, err)
}
}
emb, err := router.NewONNXEmbedder(model, tok, lib)
if err != nil {
t.Skipf("onnx embedder unavailable: %v", err)
}
defer emb.Close()
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
rep, err := Score(context.Background(), "recall+onnx", emb, InMemory, config.DefaultQueryMinScore, config.DefaultQueryMinMargin, f)
if err != nil {
t.Fatalf("Score: %v", err)
}
t.Log("\n" + rep.String() + rep.Failures())
// Cached for the sweeps only: the headline run above must pay the real
// embedder cost so its latency numbers mean something.
cached := Cache(emb)
t.Log("\ngate sweep (margin off):\n" + sweep(t, cached, f))
t.Log("\nmargin sweep (gate 0.55):\n" + marginSweep(t, cached, f))
}
// sweep scores the fixture at a range of gates and renders one line each. Two
// columns matter: how many real questions get answered, and how many made-up
// ones get answered anyway. A gate is only defensible if some value keeps the
// first high and the second at zero.
func sweep(t *testing.T, emb router.Embedder, f Fixture) string {
t.Helper()
var b strings.Builder
for _, gate := range []float64{0.0, 0.30, 0.40, 0.50, 0.55, 0.60, 0.70, 0.80, 0.90} {
rep, err := Score(context.Background(), "sweep", emb, InMemory, gate, 0, f)
if err != nil {
t.Fatalf("sweep at %.2f: %v", gate, err)
}
fmt.Fprintf(&b, " gate %.2f: answered %d/%d (%.0f%%) false recall %d/%d\n",
gate, rep.Rank1-rep.Gated, rep.Answerable, 100*rep.Answered(), rep.FalseRecall, rep.NoAnswer)
}
return b.String()
}
// marginSweep is the same idea for the margin gate (top1 top2 > delta), with
// the absolute gate held at its default. The absolute score cannot separate a
// real hit from a made-up question under e5 — every score lands in one narrow
// band — so this sweep is the one that picks a number.
func marginSweep(t *testing.T, emb router.Embedder, f Fixture) string {
t.Helper()
var b strings.Builder
for _, d := range []float64{0, 0.002, 0.005, 0.008, 0.01, 0.012, 0.015, 0.02, 0.025, 0.03, 0.04, 0.05, 0.06} {
rep, err := Score(context.Background(), "margin sweep", emb, InMemory, config.DefaultQueryMinScore, d, f)
if err != nil {
t.Fatalf("margin sweep at %.3f: %v", d, err)
}
fmt.Fprintf(&b, " delta %.3f: answered %d/%d (%.0f%%) false recall %d/%d\n",
d, rep.Rank1-rep.Gated, rep.Answerable, 100*rep.Answered(), rep.FalseRecall, rep.NoAnswer)
}
return b.String()
}
@@ -0,0 +1,392 @@
{
"schema_version": 1,
"name": "ru_recall_v1",
"notes": [
"Held-out note-recall fixture. Each case is a fresh semantic store: insert every note, embed the query, take the top 3 — the same read path cmd/mavend/voice.go runs for IntentQuery.",
"Queries paraphrase their note on purpose. A query that repeats the note's words measures string matching, not recall. TestFixtureIsParaphrased enforces a ceiling on word overlap.",
"want:\"\" means the query must recall NOTHING. Those cases measure false recall — the direction the spec calls out (a confident wrong fact is worse than a known gap).",
"The distractor tag marks cases where a second note is plausible and only one is right. The hard tag marks cases with little or no shared vocabulary.",
"Content is written for this operator: his preferences, his homelab, things he said once and would expect Maven to remember weeks later."
],
"filler": [
{"id": "f1", "text": "в субботу ходил в баню", "kind": "note"},
{"id": "f2", "text": "купил новые кроссовки сорок третьего размера", "kind": "note"},
{"id": "f3", "text": "сериал закончился на третьем сезоне", "kind": "note"},
{"id": "f4", "text": "сосед сверху делает ремонт", "kind": "note"},
{"id": "f5", "text": "билеты в театр брал заранее", "kind": "note"},
{"id": "f6", "text": "выучил пару аккордов на гитаре", "kind": "note"},
{"id": "f7", "text": "записался к стоматологу", "kind": "note"},
{"id": "f8", "text": "поменял лампочку в коридоре", "kind": "note"},
{"id": "f9", "text": "погулял вдоль реки", "kind": "note"},
{"id": "f10", "text": "the balcony door sticks in winter", "kind": "note"},
{"id": "f11", "text": "the neighbour's dog barks at cyclists", "kind": "note"},
{"id": "f12", "text": "i finished the book about volcanoes", "kind": "note"}
],
"cases": [
{
"id": "ru-pref-001",
"lang": "ru",
"tags": ["preference", "paraphrase"],
"query": "какой кофе мне наливать",
"want": "n1",
"notes": [
{"id": "n1", "text": "я пью кофе без сахара", "kind": "note"},
{"id": "n2", "text": "по утрам бегаю в парке", "kind": "note"},
{"id": "n3", "text": "не люблю громкую музыку", "kind": "note"}
]
},
{
"id": "ru-pref-002",
"lang": "ru",
"tags": ["preference", "homelab", "paraphrase", "hard"],
"query": "когда запускать резервное копирование",
"want": "n1",
"note": "The DESIGN.md preference-seam example, phrased as the operator would ask it later.",
"notes": [
{"id": "n1", "text": "бэкапы лучше делать ночью в три часа", "kind": "note"},
{"id": "n2", "text": "обновления ставлю по субботам", "kind": "note"},
{"id": "n3", "text": "логи храню месяц", "kind": "note"}
]
},
{
"id": "ru-home-003",
"lang": "ru",
"tags": ["homelab", "paraphrase"],
"query": "что помогло от мерцания монитора",
"want": "n1",
"notes": [
{"id": "n1", "text": "мерцание экрана прошло после обновления драйвера amdgpu", "kind": "note"},
{"id": "n2", "text": "вентилятор шумит на полной нагрузке", "kind": "note"},
{"id": "n3", "text": "поставил новый ssd в ноутбук", "kind": "note"}
]
},
{
"id": "ru-home-004",
"lang": "ru",
"tags": ["homelab", "distractor", "hard"],
"query": "адрес домашнего сервера",
"want": "n2",
"note": "Two notes carry an IP. Only one is the server.",
"notes": [
{"id": "n1", "text": "роутер живёт на 192.168.1.1", "kind": "note"},
{"id": "n2", "text": "домашний сервер на 192.168.1.104", "kind": "note"},
{"id": "n3", "text": "принтер подключен по usb", "kind": "note"}
]
},
{
"id": "ru-home-005",
"lang": "ru",
"tags": ["homelab"],
"query": "где искать настройки nginx",
"want": "n1",
"notes": [
{"id": "n1", "text": "конфиг nginx лежит в /etc/nginx/sites-enabled", "kind": "note"},
{"id": "n2", "text": "сертификаты обновляет certbot по расписанию", "kind": "note"},
{"id": "n3", "text": "порт 8080 занят вебкой", "kind": "note"}
]
},
{
"id": "ru-pref-006",
"lang": "ru",
"tags": ["preference", "distractor", "hard"],
"query": "что мне нельзя есть",
"want": "n1",
"note": "The Friday-meat note is a plausible second answer but it is a habit, not a restriction.",
"notes": [
{"id": "n1", "text": "у меня аллергия на орехи", "kind": "note"},
{"id": "n2", "text": "не ем мясо по пятницам", "kind": "note"},
{"id": "n3", "text": "люблю острую еду", "kind": "note"}
]
},
{
"id": "ru-pers-007",
"lang": "ru",
"tags": ["distractor", "paraphrase"],
"query": "когда мамин праздник",
"want": "n1",
"notes": [
{"id": "n1", "text": "день рождения мамы четырнадцатого марта", "kind": "note"},
{"id": "n2", "text": "у брата день рождения в июле", "kind": "note"},
{"id": "n3", "text": "годовщина в сентябре", "kind": "note"}
]
},
{
"id": "ru-home-008",
"lang": "ru",
"tags": ["homelab", "hard", "paraphrase"],
"query": "чем ускоряется языковая модель",
"want": "n1",
"notes": [
{"id": "n1", "text": "модель крутится на встройке через vulkan", "kind": "note"},
{"id": "n2", "text": "whisper работает на процессоре", "kind": "note"},
{"id": "n3", "text": "голос у piper русский", "kind": "note"}
]
},
{
"id": "ru-pref-009",
"lang": "ru",
"tags": ["preference", "distractor"],
"query": "во сколько я обычно засыпаю",
"want": "n1",
"notes": [
{"id": "n1", "text": "ложусь спать около часа ночи", "kind": "note"},
{"id": "n2", "text": "встаю в семь утра", "kind": "note"},
{"id": "n3", "text": "днём не сплю", "kind": "note"}
]
},
{
"id": "ru-home-010",
"lang": "ru",
"tags": ["homelab", "paraphrase"],
"query": "где у меня хранятся пароли",
"want": "n1",
"notes": [
{"id": "n1", "text": "пароли держу в keepassxc", "kind": "note"},
{"id": "n2", "text": "двухфакторку сделал через totp", "kind": "note"},
{"id": "n3", "text": "ssh ключи лежат на юбикее", "kind": "note"}
]
},
{
"id": "ru-home-011",
"lang": "ru",
"tags": ["homelab", "hard", "paraphrase"],
"query": "из-за чего кончилось место",
"want": "n1",
"notes": [
{"id": "n1", "text": "диск забился логами докера в июне", "kind": "note"},
{"id": "n2", "text": "рейд собрал из двух дисков", "kind": "note"},
{"id": "n3", "text": "бэкап на внешний диск раз в неделю", "kind": "note"}
]
},
{
"id": "ru-pref-012",
"lang": "ru",
"tags": ["preference", "distractor"],
"query": "какой чай мне нравится",
"want": "n1",
"notes": [
{"id": "n1", "text": "чай пью только зелёный", "kind": "note"},
{"id": "n2", "text": "кофе пью без сахара", "kind": "note"},
{"id": "n3", "text": "воду пью из фильтра", "kind": "note"}
]
},
{
"id": "ru-silent-013",
"lang": "ru",
"tags": ["silent"],
"query": "какая погода будет в пятницу",
"want": "",
"notes": [
{"id": "n1", "text": "роутер живёт на 192.168.1.1", "kind": "note"},
{"id": "n2", "text": "бэкапы лучше делать ночью", "kind": "note"},
{"id": "n3", "text": "у меня аллергия на орехи", "kind": "note"}
]
},
{
"id": "ru-silent-014",
"lang": "ru",
"tags": ["silent"],
"query": "как зовут сестру моего коллеги",
"want": "",
"notes": [
{"id": "n1", "text": "конфиг nginx лежит в /etc/nginx/sites-enabled", "kind": "note"},
{"id": "n2", "text": "порт 8080 занят вебкой", "kind": "note"},
{"id": "n3", "text": "сертификаты обновляет certbot", "kind": "note"}
]
},
{
"id": "ru-silent-015",
"lang": "ru",
"tags": ["silent"],
"query": "сколько я заплатил за машину",
"want": "",
"notes": [
{"id": "n1", "text": "чай пью только зелёный", "kind": "note"},
{"id": "n2", "text": "ложусь спать около часа ночи", "kind": "note"},
{"id": "n3", "text": "не люблю громкую музыку", "kind": "note"}
]
},
{
"id": "ru-home-016",
"lang": "ru",
"tags": ["homelab", "paraphrase"],
"query": "откуда берётся токен бота",
"want": "n1",
"notes": [
{"id": "n1", "text": "токен телеграма лежит в deploy/telegram.env", "kind": "note"},
{"id": "n2", "text": "вебхуки не использую, только long-poll", "kind": "note"},
{"id": "n3", "text": "уведомления приходят в личку", "kind": "note"}
]
},
{
"id": "ru-hard-017",
"lang": "ru",
"tags": ["hard", "paraphrase", "homelab"],
"query": "как я восстановил конфиги",
"want": "n1",
"note": "No shared word between query and note beyond none at all. This is the case a lexical embedder cannot win.",
"notes": [
{"id": "n1", "text": "после переустановки системы вернул все настройки из git", "kind": "note"},
{"id": "n2", "text": "разделы на диске резал вручную", "kind": "note"},
{"id": "n3", "text": "загрузчик поставил заново", "kind": "note"}
]
},
{
"id": "ru-dist-018",
"lang": "ru",
"tags": ["distractor"],
"query": "чем кормить кота",
"want": "n1",
"notes": [
{"id": "n1", "text": "кот ест только сухой корм", "kind": "note"},
{"id": "n2", "text": "собаке даю мясо", "kind": "note"},
{"id": "n3", "text": "рыбок кормлю раз в день", "kind": "note"}
]
},
{
"id": "ru-home-019",
"lang": "ru",
"tags": ["homelab", "hard", "paraphrase"],
"query": "как контейнер получает доступ к видеокарте",
"want": "n1",
"notes": [
{"id": "n1", "text": "docker compose пробрасывает /dev/dri внутрь", "kind": "note"},
{"id": "n2", "text": "контейнеры рестартуют сами", "kind": "note"},
{"id": "n3", "text": "образы чищу вручную", "kind": "note"}
]
},
{
"id": "ru-pref-020",
"lang": "ru",
"tags": ["preference", "distractor"],
"query": "когда мне нельзя звонить",
"want": "n1",
"notes": [
{"id": "n1", "text": "не звони мне после десяти вечера", "kind": "note"},
{"id": "n2", "text": "утром не трогай меня до кофе", "kind": "note"},
{"id": "n3", "text": "по выходным не работаю", "kind": "note"}
]
},
{
"id": "en-pref-021",
"lang": "en",
"tags": ["preference", "paraphrase"],
"query": "which colour scheme do i like",
"want": "n1",
"notes": [
{"id": "n1", "text": "i prefer dark theme everywhere", "kind": "note"},
{"id": "n2", "text": "font size 14 is fine", "kind": "note"},
{"id": "n3", "text": "i use vim keybindings", "kind": "note"}
]
},
{
"id": "en-home-022",
"lang": "en",
"tags": ["homelab", "distractor"],
"query": "where is the big disk mounted",
"want": "n1",
"notes": [
{"id": "n1", "text": "the nas drive is mounted at /mnt/hdd1", "kind": "note"},
{"id": "n2", "text": "models live on the ssd", "kind": "note"},
{"id": "n3", "text": "backups go to the nas nightly", "kind": "note"}
]
},
{
"id": "en-silent-023",
"lang": "en",
"tags": ["silent"],
"query": "what is my bank account number",
"want": "",
"notes": [
{"id": "n1", "text": "the nas drive is mounted at /mnt/hdd1", "kind": "note"},
{"id": "n2", "text": "i prefer dark theme everywhere", "kind": "note"},
{"id": "n3", "text": "the router runs openwrt", "kind": "note"}
]
},
{
"id": "en-hard-024",
"lang": "en",
"tags": ["hard", "paraphrase"],
"query": "what fixed the screen problem",
"want": "n1",
"notes": [
{"id": "n1", "text": "the flicker went away once i swapped the display cable", "kind": "note"},
{"id": "n2", "text": "the laptop fan is loud", "kind": "note"},
{"id": "n3", "text": "the second monitor is 1440p", "kind": "note"}
]
},
{
"id": "en-pref-025",
"lang": "en",
"tags": ["preference", "hard"],
"query": "should i be offered wine",
"want": "n1",
"notes": [
{"id": "n1", "text": "i do not drink alcohol", "kind": "note"},
{"id": "n2", "text": "i skip breakfast", "kind": "note"},
{"id": "n3", "text": "i like spicy food", "kind": "note"}
]
},
{
"id": "ru-home-026",
"lang": "ru",
"tags": ["homelab", "paraphrase"],
"query": "какая модель распознавания речи мне подходит",
"want": "n1",
"notes": [
{"id": "n1", "text": "whisper модель small хватает для русского", "kind": "note"},
{"id": "n2", "text": "голос ирина звучит лучше остальных", "kind": "note"},
{"id": "n3", "text": "слово активации маven", "kind": "note"}
]
},
{
"id": "ru-fact-027",
"lang": "ru",
"tags": ["distractor", "hard", "paraphrase"],
"query": "когда я последний раз обслуживал машину",
"want": "n1",
"note": "A fact, not a note — both share the vector index, so a fact can win a recall.",
"notes": [
{"id": "n1", "text": "последний раз менял масло в мае", "kind": "fact"},
{"id": "n2", "text": "шины поменял осенью", "kind": "fact"},
{"id": "n3", "text": "страховка до декабря", "kind": "fact"}
]
},
{
"id": "ru-pref-028",
"lang": "ru",
"tags": ["preference", "hard", "paraphrase"],
"query": "как мне присылать оповещения",
"want": "n1",
"notes": [
{"id": "n1", "text": "терпеть не могу уведомления со звуком", "kind": "note"},
{"id": "n2", "text": "вибрацию оставь включённой", "kind": "note"},
{"id": "n3", "text": "письма читаю вечером", "kind": "note"}
]
},
{
"id": "ru-silent-029",
"lang": "ru",
"tags": ["silent"],
"query": "во сколько отходит поезд",
"want": "",
"notes": [
{"id": "n1", "text": "кот ест только сухой корм", "kind": "note"},
{"id": "n2", "text": "люблю острую еду", "kind": "note"},
{"id": "n3", "text": "пароли держу в keepassxc", "kind": "note"}
]
},
{
"id": "en-home-030",
"lang": "en",
"tags": ["homelab", "paraphrase"],
"query": "what firmware is on the router",
"want": "n1",
"notes": [
{"id": "n1", "text": "the router runs openwrt", "kind": "note"},
{"id": "n2", "text": "wifi channel is 6", "kind": "note"},
{"id": "n3", "text": "the guest network is off", "kind": "note"}
]
}
]
}
+287
View File
@@ -0,0 +1,287 @@
package eval
import (
"fmt"
"regexp"
"strings"
"unicode"
"unicode/utf8"
)
// The check names, in report order. Every check is a string or length test — no
// model grades another model here.
const (
CheckMood = "mood" // mood is in the documented enum
CheckLang = "lang" // the operator's language, not the prompt's
CheckLength = "length" // a nudge is one sentence, not a paragraph
CheckFeminine = "feminine" // her self-reference is feminine (hard constraint)
CheckCringe = "cringe" // DESIGN.md § Non-goals, "not a relationship"
CheckOnTopic = "ontopic" // says the thing the rule is about
)
// CheckNames — report order.
var CheckNames = []string{CheckMood, CheckLang, CheckLength, CheckFeminine, CheckCringe, CheckOnTopic}
// Result — one check on one message.
type Result struct {
Name string
Pass bool
Detail string
}
// Moods — the fixed enum from the LLM output contract. Not extended here; the
// contract lives in the daemon and the harness only reads it.
var Moods = map[string]bool{
"neutral": true, "happy": true, "thinking": true, "tired": true, "confused": true,
}
// Length ceilings. Justification: the nudge is spoken by piper at roughly 14
// characters per second, so 120 characters is about 8 seconds of speech. The
// operator has AuDHD — past one short sentence a nudge stops being a nudge and
// becomes something to tune out, which is exactly the "not a nag" failure. The
// word ceiling catches the same thing for languages that pack more per byte.
const (
MaxChars = 120
MaxWords = 16
)
// RunChecks scores one message. Order matches CheckNames.
func RunChecks(c Case, body, mood string) []Result {
return []Result{
checkMood(mood),
checkLang(body),
checkLength(body),
checkFeminine(body),
checkCringe(body),
checkOnTopic(c, body),
}
}
func checkMood(mood string) Result {
if Moods[mood] {
return Result{CheckMood, true, ""}
}
return Result{CheckMood, false, fmt.Sprintf("mood %q not in the enum", mood)}
}
// checkLang — the operator is Russian-speaking and the nudge is spoken aloud by
// a Russian piper voice. An English nudge is not a tone problem, it is an
// unusable one.
func checkLang(body string) Result {
cyr, lat := 0, 0
for _, r := range body {
switch {
case unicode.Is(unicode.Cyrillic, r):
cyr++
case r >= 'a' && r <= 'z', r >= 'A' && r <= 'Z':
lat++
}
}
if cyr > lat {
return Result{CheckLang, true, ""}
}
return Result{CheckLang, false, fmt.Sprintf("not Russian (%d cyrillic vs %d latin letters)", cyr, lat)}
}
func checkLength(body string) Result {
chars := utf8.RuneCountInString(body)
words := len(strings.Fields(body))
if chars <= MaxChars && words <= MaxWords {
return Result{CheckLength, true, ""}
}
return Result{CheckLength, false,
fmt.Sprintf("%d chars / %d words, ceiling %d / %d", chars, words, MaxChars, MaxWords)}
}
// --- feminine self-reference ---------------------------------------------
//
// The hard constraint (CLAUDE.md, DESIGN.md § Identity): Maven's Russian
// self-reference is feminine. The operator is male, so second-person forms
// addressed to him are MASCULINE and must not be flagged — "ты не пил воду" is
// correct, "я напомнил" is not. Both directions matter, which is why this is a
// windowed scan around "я" and not a bare search for masculine endings.
var wordRE = regexp.MustCompile(`[\p{Cyrillic}]+|[,.;:!?…—-]`)
// secondPerson — pronouns that end the self-reference window. Everything after
// one of these is about him, not about her.
var secondPerson = map[string]bool{
"ты": true, "тебе": true, "тебя": true, "тобой": true,
"вы": true, "вам": true, "вас": true,
"он": true, "она": true, "оно": true, "они": true,
}
// masculinePredicative — short adjectives with no verb ending to key off.
var masculinePredicative = map[string]bool{
"должен": true, "готов": true, "рад": true, "уверен": true,
"обязан": true, "сам": true, "занят": true, "прав": true,
}
// nounsEndingInL — the false positives of "ends in л ⇒ masculine past tense".
// Small on purpose: it only has to cover nouns a nudge might actually use.
var nounsEndingInL = map[string]bool{
"стол": true, "стул": true, "пол": true, "зал": true, "гол": true,
"узел": true, "отдел": true, "файл": true, "канал": true, "угол": true,
"футбол": true, "вокзал": true, "металл": true, "интервал": true,
"уровень": true, "мускул": true, "апрель": true, "июль": true, "рубль": true,
}
// masculinePast reports whether a word looks like a masculine past-tense verb.
// Russian past tense is gendered by suffix: -л (m), -ла (f). A 0.8B with weak
// Russian defaults to the masculine form, which is the exact drift being
// measured.
func masculinePast(w string) bool {
if len([]rune(w)) < 3 || nounsEndingInL[w] {
return false
}
return strings.HasSuffix(w, "л") || strings.HasSuffix(w, "лся")
}
func checkFeminine(body string) Result {
words := wordRE.FindAllString(strings.ToLower(body), -1)
for i, w := range words {
if w != "я" {
continue
}
// Scan the next few words. Stop at punctuation or at a second-person
// pronoun: past that point the sentence is about him and masculine is
// correct.
for j := i + 1; j < len(words) && j <= i+3; j++ {
nw := words[j]
if len(nw) == 1 && !unicode.Is(unicode.Cyrillic, []rune(nw)[0]) {
break
}
if secondPerson[nw] {
break
}
if masculinePast(nw) || masculinePredicative[nw] {
return Result{CheckFeminine, false,
fmt.Sprintf("masculine self-reference %q after \"я\"", nw)}
}
}
}
// Second pass: self-reference with the pronoun dropped — "напомнил тебе",
// "проверил за тебя". A masculine past-tense verb whose object is HIM can
// only be her speaking about herself.
for i, w := range words {
if !masculinePast(w) || i+1 >= len(words) {
continue
}
next := words[i+1]
if next == "тебе" || next == "тебя" || next == "за" {
return Result{CheckFeminine, false,
fmt.Sprintf("masculine self-reference %q before %q", w, next)}
}
}
return Result{CheckFeminine, true, ""}
}
// --- the cringe checks ---------------------------------------------------
//
// "Think Jarvis without the cringe part". DESIGN.md § Non-goals: "Not a
// relationship — mom-tone is a function that makes nudges land, not emotional
// company. Names the drift a warm small model falls into." Each pattern below
// is one shape of that drift. They are deliberately specific: a check that
// flags any warmth at all would make the nudges robotic, which is the other
// failure.
type cringePattern struct {
// what the pattern is defending against, shown in the failure detail.
why string
pat *regexp.Regexp
}
var cringePatterns = []cringePattern{
{
// Endearments. "Not a relationship" — a pet name reframes a nudge as
// intimacy, and the operator asked for mother-like, not girlfriend-like.
why: "pet name / endearment",
// No \b around the Russian alternatives: Go's RE2 \b is ASCII-only and
// never matches at a Cyrillic boundary, so anchoring them would make
// this check silently always pass.
pat: regexp.MustCompile(`(?i)(милый|дорогой|солнышко|солнце моё|солнце мое|зайчик|котик|сладкий|малыш|дружок|родной|любимый|\bhoney\b|\bsweetie\b|\bdarling\b|\bbuddy\b)`),
},
{
// Emoji. The nudge is spoken aloud; an emoji is either silence or a TTS
// artefact. Also the single loudest cringe signal in a small model.
why: "emoji",
pat: nil, // handled by hasEmoji, ranges don't fit a regexp cleanly
},
{
// Exclamation pileup. One "!" is emphasis; two is a cheerleader.
why: "more than one exclamation mark",
pat: regexp.MustCompile(`!.*!|!!`),
},
{
// Fake concern. She has no feelings to report, and reporting them makes
// the nudge about her instead of about the water.
why: "fake concern opener",
pat: regexp.MustCompile(`(?i)(я волну|я беспоко|беспокоюсь|переживаю|я забочусь|я тревож|мне тревожно|i'?m worried)`),
},
{
// Apologising. The rule decided she speaks. Apologising for a greenlit
// nudge undermines the one thing that makes nudges land.
why: "apology",
pat: regexp.MustCompile(`(?i)(извини|прости|сожалею|прошу прощения|не хочу мешать|не хочу отвлекать|sorry|apolog)`),
},
{
// Offering emotional support. The explicit "not emotional company" line.
why: "offer of emotional support",
pat: regexp.MustCompile(`(?i)(я рядом|я здесь для теб|ты не один|всё будет хорошо|все будет хорошо|не переживай|я с тобой|обнимаю|я поддерж|держись)`),
},
{
// Asking how he feels. Turns a one-way nudge into a conversation he now
// owes an answer to — the most reliable way to make him mute it.
why: "asking how he feels",
pat: regexp.MustCompile(`(?i)(как ты\s*[?.!]|как ты себя|как самочувств|как настроение|всё ли в порядке|все ли в порядке|ты в порядке|how are you)`),
},
{
// Praise for compliance. Rewards make the nudge a training exercise;
// "not a relationship" again, from the other side.
why: "praise / reward framing",
pat: regexp.MustCompile(`(?i)(молодец|умница|ты справ|гордюсь|горжусь|отличная работа|так держать|good job|proud of you)`),
},
}
// hasEmoji — the pictographic ranges plus the variation selector. Cyrillic and
// ordinary punctuation are far below all of these.
func hasEmoji(s string) bool {
for _, r := range s {
switch {
case r >= 0x1F000 && r <= 0x1FAFF,
r >= 0x2600 && r <= 0x27BF,
r >= 0x2B00 && r <= 0x2BFF,
r == 0xFE0F, r == 0x203C, r == 0x2049:
return true
}
}
return false
}
func checkCringe(body string) Result {
for _, c := range cringePatterns {
if c.pat == nil {
if hasEmoji(body) {
return Result{CheckCringe, false, c.why}
}
continue
}
if m := c.pat.FindString(body); m != "" {
return Result{CheckCringe, false, fmt.Sprintf("%s (%q)", c.why, m)}
}
}
return Result{CheckCringe, true, ""}
}
// checkOnTopic — the message must name the thing the rule is about. A nudge
// that never mentions water leaves the operator with a chime and no action.
func checkOnTopic(c Case, body string) Result {
low := strings.ToLower(body)
for _, want := range c.WantAny {
if strings.Contains(low, strings.ToLower(want)) {
return Result{CheckOnTopic, true, ""}
}
}
return Result{CheckOnTopic, false,
fmt.Sprintf("mentions none of %v", c.WantAny)}
}
+334
View File
@@ -0,0 +1,334 @@
// Package eval scores nudge phrasing — the sentences the operator actually
// hears. It is the phrasing counterpart to internal/router/eval.
//
// Why a separate package from phraser: the fixture must be scorable by BOTH
// phrasing paths (the deterministic Stub and the resident model) from outside
// the phraser package, and a _test.go file inside phraser cannot be imported.
// So the fixture is embedded here and the scorer takes a Nudger interface.
//
// Why deterministic checks and not model judgement: the resident model is a
// 0.8B. It cannot grade its own tone. Every check in checks.go is a string or
// length test that a human can read and disagree with. A score here is a claim
// about measurable properties, not about whether a sentence is good.
//
// DESIGN.md § "Rules decide, LLM phrases" is why there is no send/veto signal
// anywhere in this package: the rule already decided she speaks. The phraser
// only words it, so a nudge the model refuses to write is a failure, never a
// legitimate outcome.
package eval
import (
"context"
_ "embed"
"encoding/json"
"fmt"
"sort"
"strings"
"time"
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/store"
)
//go:embed nudges_v1.json
var fixtureJSON []byte
// SchemaVersion — the version this package understands.
const SchemaVersion = 1
// Case — one nudge situation, as a real tick would present it. The fields are
// the (rule, severity, context) input DESIGN.md names, flattened to JSON.
//
// WantAny is the on-topic contract: at least one of these lowercased fragments
// must appear in the message. A water nudge that never mentions water is a
// failure however charming it reads. Fragments are stems ("вод") so declension
// does not defeat the check, and they list both languages because the Stub is
// still English (see the writeup).
type Case struct {
ID string `json:"id"`
Rule string `json:"rule"`
Severity int `json:"severity"`
// SinceMinutes — age of the rule's own fact. 0 means "no such fact", which
// is the branch where the phraser has no duration to name.
SinceMinutes int `json:"since_minutes"`
// FactKey/FactValue/FactSource — the aggregate fact behind the ops rules.
// service_down phrasing reads the key for the service name.
FactKey string `json:"fact_key,omitempty"`
FactValue string `json:"fact_value,omitempty"`
FactSource string `json:"fact_source,omitempty"`
// QuietHours/CalendarBusy — the bad moments. The gate already let this
// nudge through (ops outranks quiet hours), so the phrasing still has to be
// short and plain rather than apologetic about the timing.
QuietHours bool `json:"quiet_hours,omitempty"`
CalendarBusy bool `json:"calendar_busy,omitempty"`
WantAny []string `json:"want_any"`
Tags []string `json:"tags,omitempty"`
Note string `json:"note,omitempty"`
}
// Fixture — the versioned envelope. SchemaVersion gates the loader so an older
// binary refuses a fixture it would misread instead of reporting a wrong score.
type Fixture struct {
SchemaVersion int `json:"schema_version"`
Name string `json:"name"`
ReferenceNow string `json:"reference_now"`
Notes []string `json:"notes"`
Cases []Case `json:"cases"`
}
// Load returns the embedded fixture.
func Load() (Fixture, error) {
var f Fixture
if err := json.Unmarshal(fixtureJSON, &f); err != nil {
return Fixture{}, fmt.Errorf("parse fixture: %w", err)
}
if f.SchemaVersion != SchemaVersion {
return Fixture{}, fmt.Errorf("fixture schema_version %d, want %d", f.SchemaVersion, SchemaVersion)
}
if len(f.Cases) == 0 {
return Fixture{}, fmt.Errorf("fixture has no cases")
}
return f, nil
}
// Now — the fixture's reference clock, so fact ages are reproducible.
func (f Fixture) Now() (time.Time, error) {
t, err := time.Parse(time.RFC3339, f.ReferenceNow)
if err != nil {
return time.Time{}, fmt.Errorf("parse reference_now %q: %w", f.ReferenceNow, err)
}
return t, nil
}
// Candidate rebuilds the loop.Candidate a tick would hand the phraser.
func (c Case) Candidate(now time.Time) loop.Candidate {
state := loop.State{
Now: now,
Facts: map[string]store.Fact{},
QuietHours: c.QuietHours,
CalendarBusy: c.CalendarBusy,
}
if c.SinceMinutes > 0 {
state.Facts[c.Rule] = store.Fact{
Key: c.Rule,
Ts: now.Add(-time.Duration(c.SinceMinutes) * time.Minute),
}
}
if c.FactKey != "" {
state.Facts[c.Rule] = store.Fact{
Key: c.FactKey,
Value: c.FactValue,
Source: c.FactSource,
Ts: now.Add(-time.Duration(c.SinceMinutes) * time.Minute),
}
}
sev := loop.Severity(c.Severity)
return loop.Candidate{
Rule: loop.Rule{Name: c.Rule, Severity: sev},
Severity: sev,
State: state,
}
}
// Nudger — the one thing a phrasing path must do to be scorable. Both
// *phraser.Stub and *phraser.LLMPhraser satisfy it.
type Nudger interface {
PhraseNudge(ctx context.Context, c loop.Candidate) (delivery.PhrasedNudge, error)
}
// Outcome — one scored case. Failed lists the check names that did not pass,
// Reasons the human-readable detail. Failed is empty exactly when Pass is true.
type Outcome struct {
Case Case
Body string
Mood string
Err error
Latency time.Duration
Pass bool
Failed []string
Reasons []string
}
// Report — the aggregate. ByCheck is the useful part: one composite percentage
// hides which property broke, and tuning a prompt needs to know.
type Report struct {
Name string
Total int
Passed int
Errors int
ByCheck map[string]int
ByRule map[string]TagStat
Outcomes []Outcome
P50 time.Duration
P95 time.Duration
Max time.Duration
}
// TagStat — passed/total for one slice of the fixture.
type TagStat struct{ Passed, Total int }
// Accuracy — fraction of cases that passed every check.
func (r Report) Accuracy() float64 {
if r.Total == 0 {
return 0
}
return float64(r.Passed) / float64(r.Total)
}
// Score runs every case through p and aggregates. A phrasing error scores as a
// miss and is counted in Errors — "the model was down" and "the model wrote
// something bad" are different numbers and a prompt change must not be able to
// hide behind the first one.
//
// Latency is wall-clock per PhraseNudge call. On the CPU/iGPU target a nudge
// the model takes a minute to word has already missed its moment.
func Score(ctx context.Context, name string, p Nudger, f Fixture) (Report, error) {
now, err := f.Now()
if err != nil {
return Report{}, err
}
rep := Report{
Name: name,
Total: len(f.Cases),
ByCheck: map[string]int{},
ByRule: map[string]TagStat{},
}
for _, name := range CheckNames {
rep.ByCheck[name] = 0
}
lat := make([]time.Duration, 0, len(f.Cases))
for _, c := range f.Cases {
start := time.Now()
pn, err := p.PhraseNudge(ctx, c.Candidate(now))
o := Outcome{Case: c, Body: pn.Body, Mood: pn.Mood, Err: err, Latency: time.Since(start)}
lat = append(lat, o.Latency)
if err != nil {
rep.Errors++
o.Failed = append(o.Failed, "call")
o.Reasons = append(o.Reasons, fmt.Sprintf("phrase error: %v", err))
} else {
for _, res := range RunChecks(c, pn.Body, pn.Mood) {
if res.Pass {
rep.ByCheck[res.Name]++
continue
}
o.Failed = append(o.Failed, res.Name)
o.Reasons = append(o.Reasons, res.Name+": "+res.Detail)
}
}
o.Pass = len(o.Failed) == 0
if o.Pass {
rep.Passed++
}
bump(rep.ByRule, ruleFamily(c.Rule), o.Pass)
rep.Outcomes = append(rep.Outcomes, o)
}
sort.Slice(lat, func(i, j int) bool { return lat[i] < lat[j] })
rep.P50, rep.P95 = percentile(lat, 0.50), percentile(lat, 0.95)
if len(lat) > 0 {
rep.Max = lat[len(lat)-1]
}
return rep, nil
}
// ruleFamily collapses "routine:зарядка" to "routine" so the per-rule table
// stays readable however many routines the operator configures.
func ruleFamily(rule string) string {
if i := strings.IndexByte(rule, ':'); i > 0 {
return rule[:i]
}
return rule
}
func bump(m map[string]TagStat, key string, pass bool) {
if key == "" {
return
}
s := m[key]
s.Total++
if pass {
s.Passed++
}
m[key] = s
}
// percentile — nearest-rank on a pre-sorted slice. No interpolation: with ~15
// samples an interpolated p95 invents a latency no call actually took.
func percentile(sorted []time.Duration, p float64) time.Duration {
if len(sorted) == 0 {
return 0
}
i := int(p * float64(len(sorted)))
if i >= len(sorted) {
i = len(sorted) - 1
}
return sorted[i]
}
// String renders the comparison table — composite score, then per-check so a
// regression names the property it broke, then latency.
func (r Report) String() string {
var b strings.Builder
fmt.Fprintf(&b, "%s: %d/%d cases pass every check (%.1f%%), %d errors\n",
r.Name, r.Passed, r.Total, 100*r.Accuracy(), r.Errors)
for _, name := range CheckNames {
fmt.Fprintf(&b, " %-9s %d/%d\n", name, r.ByCheck[name], r.Total)
}
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
fmt.Fprintf(&b, " by rule: %s\n", renderStats(r.ByRule))
return b.String()
}
// Failures — per-case detail, sorted by ID so two runs diff cleanly.
func (r Report) Failures() string {
var b strings.Builder
for _, o := range r.sorted() {
if o.Pass {
continue
}
fmt.Fprintf(&b, " %s %q\n %s\n", o.Case.ID, o.Body, strings.Join(o.Reasons, "; "))
}
return b.String()
}
// Messages — every generated message verbatim, pass or fail. This is what a
// human reads to judge tone; the score only says which checks fired.
func (r Report) Messages() string {
var b strings.Builder
for _, o := range r.sorted() {
mark := "ok "
if !o.Pass {
mark = "FAIL"
}
fmt.Fprintf(&b, " %s %-22s [%s] %q\n", mark, o.Case.ID, o.Mood, o.Body)
}
return b.String()
}
func (r Report) sorted() []Outcome {
out := append([]Outcome(nil), r.Outcomes...)
sort.Slice(out, func(i, j int) bool { return out[i].Case.ID < out[j].Case.ID })
return out
}
func renderStats(m map[string]TagStat) string {
keys := make([]string, 0, len(m))
for k := range m {
keys = append(keys, k)
}
sort.Strings(keys)
parts := make([]string, 0, len(keys))
for _, k := range keys {
parts = append(parts, fmt.Sprintf("%s %d/%d", k, m[k].Passed, m[k].Total))
}
return strings.Join(parts, " ")
}
+179
View File
@@ -0,0 +1,179 @@
package eval
import (
"context"
"strings"
"testing"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/phraser"
)
func TestLoadFixture(t *testing.T) {
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
if _, err := f.Now(); err != nil {
t.Fatalf("Now: %v", err)
}
seen := map[string]bool{}
for _, c := range f.Cases {
if seen[c.ID] {
t.Errorf("duplicate case id %q", c.ID)
}
seen[c.ID] = true
if c.Rule == "" || c.Severity < 1 || c.Severity > 4 {
t.Errorf("%s: rule %q severity %d", c.ID, c.Rule, c.Severity)
}
if len(c.WantAny) == 0 {
t.Errorf("%s: no want_any, the on-topic check would always pass", c.ID)
}
}
// Coverage floor: all five loop rules plus both minted families, or the
// fixture measures a subset and the score does not mean what it says.
for _, rule := range []string{"water", "meal", "break", "service_down", "netdata_critical", "routine", "morning"} {
found := false
for _, c := range f.Cases {
if ruleFamily(c.Rule) == rule {
found = true
}
}
if !found {
t.Errorf("no case for rule family %q", rule)
}
}
}
// TestStubBaseline is the CI ratchet: the deterministic Stub, no model, no
// network. The floor is low on purpose — the Stub is English template phrasing,
// so it fails `lang` on every case by construction. The point of the ratchet is
// that the checks keep running and the Stub does not get worse, not that the
// Stub is good.
func TestStubBaseline(t *testing.T) {
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
rep, err := Score(context.Background(), "stub (deterministic floor)", phraser.NewStub(), f)
if err != nil {
t.Fatalf("Score: %v", err)
}
t.Log("\n" + rep.String() + rep.Messages())
if rep.Errors != 0 {
t.Errorf("stub returned %d errors — the deterministic path must never fail", rep.Errors)
}
// Per-check ratchets rather than one composite: the Stub's composite is 0
// (it never passes `lang`), so a composite floor would catch nothing.
floors := map[string]int{
CheckMood: 15,
// 12, not 15: the Stub's `break` template genuinely runs past the
// ceiling ("you've been at your desk for 4 hours without a break — step
// away for a bit." is 76 chars but 16 words). Left failing rather than
// raising the ceiling to hide it.
CheckLength: 12,
CheckFeminine: 15,
CheckCringe: 15,
CheckOnTopic: 12,
}
for name, floor := range floors {
if rep.ByCheck[name] < floor {
t.Errorf("check %s: %d/%d, below ratchet %d — phrasing regressed",
name, rep.ByCheck[name], rep.Total, floor)
}
}
}
// TestChecksCatchWhatTheyClaim — the checks are the measurement, so they get
// their own tests. Without these, a bad regexp would silently make every
// phrasing run look clean.
func TestChecksCatchWhatTheyClaim(t *testing.T) {
water := Case{Rule: "water", WantAny: []string{"вод"}}
cases := []struct {
name string
body string
want string // the check that must fail, "" for a clean message
}{
{"clean", "уже четыре часа без воды — попей.", ""},
{"long", "уже четыре часа без воды, а это довольно много, и вообще пить надо регулярно, иначе будет плохо совсем", CheckLength},
{"english", "you haven't had water in 4 hours, drink something", CheckLang},
{"masculine self", "я напомнил про воду.", CheckFeminine},
{"masculine dropped pronoun", "напомнил тебе про воду.", CheckFeminine},
{"masculine predicative", "я должен сказать: попей воды.", CheckFeminine},
// The other direction: HE is male, so second-person masculine is right.
{"second person masculine ok", "ты не пил воду четыре часа.", ""},
{"feminine self ok", "я заметила: воды не было четыре часа.", ""},
{"pet name", "милый, попей воды.", CheckCringe},
{"emoji", "попей воды 💧", CheckCringe},
{"exclamations", "попей воды!!", CheckCringe},
{"fake concern", "я беспокоюсь: воды не было четыре часа.", CheckCringe},
{"apology", "извини, что отвлекаю — попей воды.", CheckCringe},
{"emotional support", "я рядом, ты не один. попей воды.", CheckCringe},
{"asks how he feels", "как ты себя чувствуешь? попей воды.", CheckCringe},
{"praise", "молодец! теперь попей воды.", CheckCringe},
{"off topic", "пора бы уже что-то сделать.", CheckOnTopic},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
var failed []string
for _, r := range RunChecks(water, tc.body, "neutral") {
if !r.Pass {
failed = append(failed, r.Name+"("+r.Detail+")")
}
}
joined := strings.Join(failed, " ")
switch {
case tc.want == "" && len(failed) > 0:
t.Errorf("clean message flagged: %s", joined)
case tc.want != "" && !strings.Contains(joined, tc.want+"("):
t.Errorf("want %s to fail, got %q", tc.want, joined)
}
})
}
}
func TestMoodCheckUsesTheEnum(t *testing.T) {
if r := checkMood("cheerful"); r.Pass {
t.Error("mood outside the enum passed")
}
if r := checkMood(""); r.Pass {
t.Error("empty mood passed")
}
for m := range Moods {
if r := checkMood(m); !r.Pass {
t.Errorf("enum mood %q failed", m)
}
}
}
// TestCandidateCarriesTheContext — the whole harness is worthless if the
// Candidate it builds does not carry the duration the prompt is supposed to
// name.
func TestCandidateCarriesTheContext(t *testing.T) {
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
now, _ := f.Now()
for _, c := range f.Cases {
cand := c.Candidate(now)
if cand.Rule.Name != c.Rule || cand.Severity != loop.Severity(c.Severity) {
t.Errorf("%s: candidate lost rule or severity", c.ID)
}
if c.SinceMinutes > 0 {
d, ok := cand.State.Since(c.Rule)
if !ok || int(d.Minutes()) != c.SinceMinutes {
t.Errorf("%s: since %v ok=%v, want %d minutes", c.ID, d, ok, c.SinceMinutes)
}
}
if c.FactKey != "" {
fact, ok := cand.State.Fact(c.Rule)
if !ok || fact.Key != c.FactKey {
t.Errorf("%s: fact key %q, want %q", c.ID, fact.Key, c.FactKey)
}
}
}
}
+73
View File
@@ -0,0 +1,73 @@
package eval
import (
"context"
"os"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/phraser"
)
// TestLLMPhrasingBaseline — the resident model wording real nudges. Opt-in,
// same shape as internal/router/eval's MAVEN_LLM_URL gate, because CI has no
// model and a phrasing run costs minutes on the CPU target.
//
// llama-server -m /mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf \
// --host 127.0.0.1 --port 18099 -c 2048 -ngl 99
// MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing
//
// It reports and does not assert a quality bar. The numbers are the input to
// tuning the persona prompt; an assertion here would be the test inventing the
// bar rather than measuring against it. The one thing worth failing on is a
// harness fault — every case erroring means the run measured infrastructure.
func TestLLMPhrasingBaseline(t *testing.T) {
base := os.Getenv("MAVEN_LLM_URL")
if base == "" {
t.Skip("MAVEN_LLM_URL unset — point it at a running llama-server (see doc comment)")
}
// A local llama-server must not go through an HTTP proxy. This box proxies
// loopback through a SOCKS bridge that answers 503, which would score every
// case as a phrasing error and read as "the model cannot phrase".
noProxyLoopback(t)
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
cfg := phraser.DefaultConfig("")
// Generous: an unconstrained 0.8B can spend a minute thinking before it
// writes a word, and a timeout would be scored as a model failure.
cfg.Timeout = 5 * time.Minute
p := phraser.NewLLMPhraserAt(base, cfg)
defer p.Close()
rep, err := Score(context.Background(), "llm (0.8B, built-in persona)", p, f)
if err != nil {
t.Fatalf("Score: %v", err)
}
t.Log("\n" + rep.String() + "\nmessages:\n" + rep.Messages() + "\nfailures:\n" + rep.Failures())
if rep.Errors == rep.Total {
t.Errorf("all %d cases errored — harness fault, not a measurement", rep.Total)
}
}
// noProxyLoopback appends the loopback host to no_proxy before any request, so
// http.ProxyFromEnvironment (which caches the environment on first use) sees it.
func noProxyLoopback(t *testing.T) {
t.Helper()
for _, key := range []string{"no_proxy", "NO_PROXY"} {
cur := os.Getenv(key)
if strings.Contains(cur, "127.0.0.1") {
continue
}
if cur == "" {
t.Setenv(key, "127.0.0.1,localhost")
continue
}
t.Setenv(key, cur+",127.0.0.1,localhost")
}
}
+158
View File
@@ -0,0 +1,158 @@
{
"schema_version": 1,
"name": "nudge phrasing v1",
"reference_now": "2026-07-31T21:40:00+03:00",
"notes": [
"The five loop rules (water, meal, break, service_down, netdata_critical) plus one routine: and one morning: case, at the severities they actually ship with.",
"The bad-moment cases (quiet_hours, calendar_busy) are here because the gate already let them through — ops outranks quiet hours. The phrasing must stay short and plain, not apologise for the timing.",
"want_any lists stems, not whole words, so declension does not defeat the on-topic check. English stems are included because the deterministic Stub is still English.",
"There is no expected message. The checks measure properties, not similarity to a reference sentence — a fixed golden string would just freeze one arbitrary phrasing."
],
"cases": [
{
"id": "water-3h",
"rule": "water",
"severity": 1,
"since_minutes": 190,
"want_any": ["вод", "попей", "пить", "выпей", "напит", "water", "drink"],
"tags": ["care"],
"note": "The base case. Just over the 3h predicate."
},
{
"id": "water-7h",
"rule": "water",
"severity": 1,
"since_minutes": 430,
"want_any": ["вод", "попей", "пить", "выпей", "напит", "water", "drink"],
"tags": ["care"],
"note": "Long overdue. Severity is unchanged, so the phrasing must not escalate into alarm."
},
{
"id": "water-busy",
"rule": "water",
"severity": 1,
"since_minutes": 240,
"calendar_busy": true,
"want_any": ["вод", "попей", "пить", "выпей", "напит", "water", "drink"],
"tags": ["care", "bad-moment"],
"note": "Mid-meeting. A bad moment invites an apology, which is the check that should catch it."
},
{
"id": "meal-7h",
"rule": "meal",
"severity": 1,
"since_minutes": 420,
"want_any": ["ешь", "еда", "еды", "поешь", "перекус", "обед", "ужин", "покуш", "food", "eat"],
"tags": ["care"]
},
{
"id": "meal-11h-quiet",
"rule": "meal",
"severity": 1,
"since_minutes": 660,
"quiet_hours": true,
"want_any": ["ешь", "еда", "еды", "поешь", "перекус", "обед", "ужин", "покуш", "food", "eat"],
"tags": ["care", "bad-moment"],
"note": "Quiet hours suppresses care nudges in the gate, so this one only reaches the phraser via an explicit override. Included because it is the shape most likely to draw a hedge."
},
{
"id": "break-90m",
"rule": "break",
"severity": 2,
"since_minutes": 95,
"want_any": ["перерыв", "разомн", "встань", "отдохн", "пауз", "отойд", "размин", "break", "step away"],
"tags": ["care"]
},
{
"id": "break-4h",
"rule": "break",
"severity": 2,
"since_minutes": 240,
"want_any": ["перерыв", "разомн", "встань", "отдохн", "пауз", "отойд", "размин", "break", "step away"],
"tags": ["care"]
},
{
"id": "break-busy",
"rule": "break",
"severity": 2,
"since_minutes": 150,
"calendar_busy": true,
"want_any": ["перерыв", "разомн", "встань", "отдохн", "пауз", "отойд", "размин", "break", "step away"],
"tags": ["care", "bad-moment"]
},
{
"id": "service-down",
"rule": "service_down",
"severity": 4,
"since_minutes": 3,
"fact_key": "vaultwarden",
"fact_value": "\"down\"",
"fact_source": "poll:uptimekuma",
"want_any": ["vaultwarden", "сервис", "упал", "не отвеч", "лежит", "недоступ", "down", "service"],
"tags": ["ops"],
"note": "The service name is in the fact key, not the value. A nudge that says 'a service' without naming it is on-topic but useless — the on-topic check cannot catch that, a human reading Messages() can."
},
{
"id": "service-down-night",
"rule": "service_down",
"severity": 4,
"since_minutes": 2,
"quiet_hours": true,
"fact_key": "nextcloud",
"fact_value": "\"down\"",
"fact_source": "poll:uptimekuma",
"want_any": ["nextcloud", "сервис", "упал", "не отвеч", "лежит", "недоступ", "down", "service"],
"tags": ["ops", "bad-moment"],
"note": "03:00-shaped. Sev4 outranks quiet hours by design, so she speaks — plainly, without softening it into a maybe."
},
{
"id": "netdata-disk",
"rule": "netdata_critical",
"severity": 3,
"since_minutes": 5,
"fact_key": "netdata_alarm",
"fact_value": "\"critical\"",
"fact_source": "poll:netdata",
"want_any": ["netdata", "диск", "критич", "алярм", "аларм", "тревог", "место", "памят", "critical", "alarm", "disk"],
"tags": ["ops"]
},
{
"id": "netdata-busy",
"rule": "netdata_critical",
"severity": 3,
"since_minutes": 12,
"calendar_busy": true,
"fact_key": "netdata_alarm",
"fact_value": "\"critical\"",
"fact_source": "poll:netdata",
"want_any": ["netdata", "диск", "критич", "алярм", "аларм", "тревог", "место", "памят", "critical", "alarm", "disk"],
"tags": ["ops", "bad-moment"]
},
{
"id": "routine-pills",
"rule": "routine:таблетки",
"severity": 2,
"since_minutes": 0,
"want_any": ["таблетк", "приня", "лекарств", "pill"],
"tags": ["routine"],
"note": "A routine: rule has no fact of its own, so there is no duration to name. The rule name is the only context."
},
{
"id": "routine-stretch",
"rule": "routine:зарядка",
"severity": 1,
"since_minutes": 0,
"want_any": ["зарядк", "размин", "упражн", "потянис", "разомн", "stretch", "exercise"],
"tags": ["routine"]
},
{
"id": "morning-checklist",
"rule": "morning:утро",
"severity": 2,
"since_minutes": 0,
"want_any": ["утр", "чеклист", "список", "не сделан", "осталось", "morning"],
"tags": ["routine"],
"note": "In production cmd/mavend/tick.go phrases morning routines deterministically and never calls the LLM. Scored anyway: the phraser is reachable with this rule name, and a fallback that garbles it is still a bug."
}
]
}
+16
View File
@@ -67,6 +67,22 @@ func NewLLMPhraser(ctx context.Context, cfg Config) (*LLMPhraser, error) {
return p, nil
}
// NewLLMPhraserAt wires a phraser to a llama-server that someone else started
// and owns. It spawns nothing, so Close does not kill anything.
//
// This exists for the phrasing scorer (internal/phraser/eval), which must
// measure the phrasing against a shared llama-server without taking the model
// load hit per run or killing a server another process depends on. The daemon
// still uses NewLLMPhraser and still owns its own child process.
func NewLLMPhraserAt(baseURL string, cfg Config) *LLMPhraser {
return &LLMPhraser{
cfg: cfg,
client: &http.Client{Timeout: cfg.Timeout},
port: strings.TrimSuffix(baseURL, "/"),
cancel: func() {},
}
}
func (p *LLMPhraser) start(ctx context.Context) error {
args := []string{
"-m", p.cfg.ModelPath,
+5 -1
View File
@@ -92,7 +92,11 @@ func (s *Stub) Close() error { return nil }
// the predicate fire (the same State the predicate saw).
func (s *Stub) PhraseNudge(_ context.Context, c loop.Candidate) (delivery.PhrasedNudge, error) {
body, summary := phraseNudge(c)
return delivery.PhrasedNudge{Candidate: c, Body: body, Summary: summary}, nil
// "neutral" rather than empty: Mood is part of the documented output
// contract and the Stub is a production fallback, so it must satisfy the
// contract too. Template phrasing has no tone to report, and neutral is the
// enum's own default.
return delivery.PhrasedNudge{Candidate: c, Body: body, Summary: summary, Mood: "neutral"}, nil
}
// PhraseReminder — extracts the user's text from the reminder payload (raw
+4 -4
View File
@@ -26,10 +26,10 @@ func TestPythonDateParser(t *testing.T) {
ctx := context.Background()
tests := []struct {
name string
text string
wantOK bool
checkT func(t *testing.T, got, now time.Time)
name string
text string
wantOK bool
checkT func(t *testing.T, got, now time.Time)
}{
{
name: "ru relative — через час",
+31
View File
@@ -19,6 +19,37 @@ type Embedder interface {
Close() error
}
// AsymmetricEmbedder — an embedder that wants to know whether a text is a
// search query or a stored passage. Recall is asymmetric: a short question
// goes in, a longer note comes out. The e5 family is trained for exactly that
// and needs the side written into the text ("query: " / "passage: ").
//
// Optional on purpose: HashEmbedder has no such notion, so callers go through
// EmbedQuery and EmbedPassage below, which fall back to plain Embed.
type AsymmetricEmbedder interface {
Embedder
EmbedQuery(ctx context.Context, text string) ([]float32, error)
EmbedPassage(ctx context.Context, text string) ([]float32, error)
}
// EmbedQuery embeds text that is being searched WITH — a question.
func EmbedQuery(ctx context.Context, e Embedder, text string) ([]float32, error) {
if a, ok := e.(AsymmetricEmbedder); ok {
return a.EmbedQuery(ctx, text)
}
return e.Embed(ctx, text)
}
// EmbedPassage embeds text that is being searched FOR — a note or a fact on
// its way into the store. Store and lookup must use these two calls, not one
// of them twice, or the asymmetry buys nothing.
func EmbedPassage(ctx context.Context, e Embedder, text string) ([]float32, error) {
if a, ok := e.(AsymmetricEmbedder); ok {
return a.EmbedPassage(ctx, text)
}
return e.Embed(ctx, text)
}
// HashEmbedder — a deterministic bag-of-words embedder used for tests and as a
// non-zero default floor. NOT semantically meaningful across languages; the
// real classifier swaps in the multilingual ONNX model wholesale.
+57
View File
@@ -37,3 +37,60 @@ func TestHashEmbedderCyrillic(t *testing.T) {
t.Fatalf("cosine(shared)=%.3f not > cosine(disjoint)=%.3f", cosine(a, b), cosine(a, c))
}
}
// recordingEmbedder — an asymmetric embedder that only remembers which side
// was asked for. Enough to pin the dispatch; real vectors need the model.
type recordingEmbedder struct{ calls []string }
func (r *recordingEmbedder) Dim() int { return 2 }
func (r *recordingEmbedder) Close() error { return nil }
func (r *recordingEmbedder) Embed(_ context.Context, _ string) ([]float32, error) {
r.calls = append(r.calls, "embed")
return []float32{1, 0}, nil
}
func (r *recordingEmbedder) EmbedQuery(_ context.Context, _ string) ([]float32, error) {
r.calls = append(r.calls, "query")
return []float32{1, 0}, nil
}
func (r *recordingEmbedder) EmbedPassage(_ context.Context, _ string) ([]float32, error) {
r.calls = append(r.calls, "passage")
return []float32{0, 1}, nil
}
// TestEmbedQueryAndPassageSplit — a question and a stored note must not take
// the same path. If both ended up on the same call the asymmetric model buys
// nothing, which is the whole reason for the swap.
func TestEmbedQueryAndPassageSplit(t *testing.T) {
rec := &recordingEmbedder{}
if _, err := EmbedQuery(context.Background(), rec, "где логи?"); err != nil {
t.Fatalf("EmbedQuery: %v", err)
}
if _, err := EmbedPassage(context.Background(), rec, "логи в /var/log"); err != nil {
t.Fatalf("EmbedPassage: %v", err)
}
if len(rec.calls) != 2 || rec.calls[0] != "query" || rec.calls[1] != "passage" {
t.Errorf("calls %v, want [query passage]", rec.calls)
}
}
// TestEmbedFallsBackToPlainEmbed — HashEmbedder has no sides, so both helpers
// must still work and give the same vector.
func TestEmbedFallsBackToPlainEmbed(t *testing.T) {
h := NewHashEmbedder(64)
q, err := EmbedQuery(context.Background(), h, "text")
if err != nil {
t.Fatalf("EmbedQuery: %v", err)
}
p, err := EmbedPassage(context.Background(), h, "text")
if err != nil {
t.Fatalf("EmbedPassage: %v", err)
}
for i := range q {
if q[i] != p[i] {
t.Fatalf("hash embedder gave two different vectors for the same text")
}
}
}
+2 -2
View File
@@ -185,8 +185,8 @@ func TestONNXBaseline(t *testing.T) {
if lib == "" {
t.Skip("MAVEN_ONNX_LIB unset — see AGENTS.md § Embedder model for intent routing")
}
model := filepath.Join("../../..", "models/embedder/model.onnx")
tok := filepath.Join("../../..", "models/embedder/tokenizer.json")
model := filepath.Join("../../..", "models/embedder/multilingual-e5-small/model_quantized.onnx")
tok := filepath.Join("../../..", "models/embedder/multilingual-e5-small/tokenizer.json")
for _, p := range []string{lib, model, tok} {
if _, err := os.Stat(p); err != nil {
t.Skipf("missing %s: %v", p, err)
+16 -4
View File
@@ -23,7 +23,11 @@ import (
//
// llama-server -m /mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf \
// --host 127.0.0.1 --port 18099 -c 2048 -ngl 99
// MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-router
// MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-models
//
// Every report name carries the model llama-server reports over /v1/models, so
// a bake-off across checkpoints (#278, #250) produces tables you can tell
// apart. Point the variable at one server at a time.
//
// Three configurations, because "the LLM router" is ambiguous and the three
// numbers answer different questions:
@@ -53,6 +57,14 @@ func TestLLMRouterBaseline(t *testing.T) {
}
ctx := context.Background()
model, err := ModelID(ctx, base)
if err != nil {
// Not fatal: an unlabelled score is still a score. But say so loudly,
// because an unlabelled row in a bake-off table is worthless.
t.Logf("could not read model id from %s: %v — reports will say %q", base, err, "unknown-model")
model = "unknown-model"
}
t.Logf("scoring model %s at %s", model, base)
lr := router.NewLLMRouter(client)
// llm-only: the LLM stage in isolation. Route returns (Decision, ok, err);
@@ -68,7 +80,7 @@ func TestLLMRouterBaseline(t *testing.T) {
}
return d, nil
})
repLLM, err := Score(ctx, "llm-only (0.8B, as deployed)", llmOnly, f)
repLLM, err := Score(ctx, "llm-only ("+model+", as deployed)", llmOnly, f)
if err != nil {
t.Fatalf("Score llm-only: %v", err)
}
@@ -77,7 +89,7 @@ func TestLLMRouterBaseline(t *testing.T) {
// cascade+llm: stage-0 grammar → LLM → classifier fallback, the wiring #320
// proposes. Hash embedder for the fallback so the classifier contribution is
// the deterministic floor and any lift is attributable to the model.
repCascade, err := Score(ctx, "cascade+llm (0.8B) + hash fallback",
repCascade, err := Score(ctx, "cascade+llm ("+model+") + hash fallback",
newBaselineRouter(t, router.NewHashEmbedder(1024), lr), f)
if err != nil {
t.Fatalf("Score cascade: %v", err)
@@ -96,7 +108,7 @@ func TestLLMRouterBaseline(t *testing.T) {
// either way. Kept so the question stays answered instead of being
// re-asked, and so internal/llm does NOT grow a chat_template_kwargs field
// for a problem that does not exist.
repNoThink, err := Score(ctx, "llm-only (0.8B, thinking off) [diagnostic]",
repNoThink, err := Score(ctx, "llm-only ("+model+", thinking off) [diagnostic]",
RouterFunc(func(ctx context.Context, u string, now time.Time) (router.Decision, error) {
d, ok, err := router.NewLLMRouter(&noThinkCompleter{base: base, http: &http.Client{Timeout: 60 * time.Second}}).Route(ctx, u, now)
if err != nil {
+51
View File
@@ -0,0 +1,51 @@
package eval
import (
"context"
"encoding/json"
"fmt"
"net/http"
"strings"
)
// ModelID asks llama-server which model it has loaded, so a scoring run can
// label itself. Without this a bake-off between two models produces two tables
// that look identical, and the operator has to remember which server was up.
//
// Read from the server rather than passed in on purpose: a hand-typed label
// goes stale the moment someone restarts the server with a different -m.
func ModelID(ctx context.Context, base string) (string, error) {
req, err := http.NewRequestWithContext(ctx, "GET", strings.TrimSuffix(base, "/")+"/v1/models", nil)
if err != nil {
return "", err
}
resp, err := http.DefaultClient.Do(req)
if err != nil {
return "", err
}
defer resp.Body.Close()
if resp.StatusCode != 200 {
return "", fmt.Errorf("models: status %d", resp.StatusCode)
}
var out struct {
Data []struct {
ID string `json:"id"`
} `json:"data"`
}
if err := json.NewDecoder(resp.Body).Decode(&out); err != nil {
return "", err
}
if len(out.Data) == 0 {
return "", fmt.Errorf("models: empty list")
}
return shortModelID(out.Data[0].ID), nil
}
// shortModelID trims the path and the .gguf suffix — llama-server reports the
// file name it was started with, which is too long for a table header.
func shortModelID(id string) string {
if i := strings.LastIndexAny(id, "/\\"); i >= 0 {
id = id[i+1:]
}
return strings.TrimSuffix(id, ".gguf")
}
+64 -9
View File
@@ -23,48 +23,86 @@ func NewLLMRouter(c Completer) *LLMRouter { return &LLMRouter{c: c} }
// routeGrammar — GBNF constraining the model to a JSON ARRAY of fixed-shape
// action objects (one per ask; compound utterances → multiple). Enum + key set
// prevent free-form drift from a sub-1B model.
// prevent free-form drift from a sub-1B model. The string rule is length-bounded
// so a repetition loop cannot fill the whole token budget with one field and
// truncate the JSON.
const routeGrammar = `
root ::= "[" ws action ("," ws action)* ws "]"
action ::= "{" ws "\"intent\"" ws ":" ws intent ("," ws field)* ws "}"
intent ::= "\"fact\"" | "\"reminder\"" | "\"note\"" | "\"query\"" | "\"act\"" | "\"chat\"" | "\"system\""
intent ::= "\"fact\"" | "\"reminder\"" | "\"note\"" | "\"query\"" | "\"act\"" | "\"chat\"" | "\"system\"" | "\"unknown\""
field ::= key ws ":" ws string
key ::= "\"key\"" | "\"value\"" | "\"text\"" | "\"verb\""
string ::= "\"" ([^"\\] | "\\" .)* "\""
string ::= "\"" ([^"\\] | "\\" .){0,120} "\""
ws ::= [ \t\n]*
`
// routeSystem — the router prompt. Changed 31-07-2026: the query test now sits
// above the fact test and there is an explicit question test. Before that, a
// question naming a fact key ("сколько воды я выпил с утра") matched the fact
// rule first and was stored as an assertion — 15 of 76 fixture cases.
//
// Changed again 31-07-2026: added the "unknown" escape hatch so the model can
// admit it cannot route (Vikunja #359).
//
// The training workspace keeps its own copy of this prompt for relabelling, and
// `llm/check_prompt_parity.py` there compares the two. That copy is in another
// repo and was not touched, so parity will fail until it gets the same edits —
// both the rule reorder and the "unknown" wording (Vikunja #362).
const routeSystem = `Классифицируй ровно одно сообщение пользователя. Верни ОДИН JSON-массив действий.
Ровно одно намерение: fact, reminder, note, query, act, chat, system.
Есть восьмое значение unknown только для случаев, когда просьбу невозможно понять.
Классифицируй по цели пользователя. Порядок решения:
1. Хочет напоминание в будущем reminder
2. Явно просит сохранить информацию note
3. Сообщает или обновляет текущее состояние/событие fact
4. Хочет получить информацию query
5. Просит выполнить работу act
6. Про ассистента, настройки или память system
7. Иначе chat
3. Задаёт вопрос: есть вопросительное слово (сколько, что, какой, когда, где, кто, почему, как) или знак «?» query
4. Хочет получить информацию, в том числе о своих же данных query
5. Утверждает: сообщает или обновляет текущее состояние/событие fact
6. Просит выполнить работу act
7. Про ассистента, настройки или память system
8. Реплика обрывок или указание на неназванное («это», «то», «потом»), и без него непонятно, что именно нужно сделать unknown
9. Иначе chat
Различия:
- note сохранить информацию, без напоминания. text = суть.
- reminder уведомить позже. text = что напомнить.
- fact неявное обновление: пользователь сообщает, что что-то в мире изменилось (текущее/изменённое состояние, случившееся событие). key/value.
- unknown редкий случай. Ставь его, только если в самой реплике нет ни предмета, ни действия. Короткая, простая или незнакомая тема это не причина для unknown: приветствие и болтовня это chat, вопрос на любую тему это query, просьба сделать что-то названное это act.
- query против fact решает форма реплики, а не тема. Вопрос о состоянии это query, даже если названо то же самое, что бывает в fact. Только утверждение это fact.
Примеры:
"запиши пароль" {"intent":"note","text":"пароль"}
"напомни купить молоко" {"intent":"reminder","text":"купить молоко"}
"запиши купить молоко" {"intent":"note","text":"купить молоко"}
"я выпил воду" {"intent":"fact","key":"water","value":"выпил"}
"сколько воды я выпил с утра" {"intent":"query","text":"сколько воды я выпил с утра"}
"сколько раз я ел вчера?" {"intent":"query","text":"сколько раз я ел вчера"}
"мой любимый фильм — Интерстеллар" {"intent":"note","text":"любимый фильм — Интерстеллар"}
"что такое docker?" {"intent":"query","text":"что такое docker"}
"напиши письмо" {"intent":"act","verb":"написать письмо"}
"очисти память" {"intent":"system"}
"привет" {"intent":"chat","text":"привет"}
"сделай это" {"intent":"unknown"}
"ну это" {"intent":"unknown"}
"потом" {"intent":"unknown"}
Но не путай здесь unknown не нужен:
"сделай кофе" {"intent":"act","verb":"сделать кофе"}
"что такое кватернион?" {"intent":"query","text":"что такое кватернион"}
"ага" {"intent":"chat","text":"ага"}
Ответ JSON-массив: по одному объекту на каждую просьбу. Обычно один. Если в реплике несколько просьб по объекту на каждую. "напомни купить молоко, и запиши что кофе кончился" [{"intent":"reminder","text":"купить молоко"},{"intent":"note","text":"кофе кончился"}]. Только JSON, без пояснений.`
// routeRepeatPenalty — the sub-1B model loops one sentence inside the text field
// until it runs out of tokens, which truncates the JSON. 1.15 is enough to break
// the loop without hurting short slot values.
const routeRepeatPenalty = 1.15
// routeIntentUnknown — the model's way of saying "I could not route this".
// It is a wire value only: it never becomes a router.Intent, it just makes
// Route return ok=false so the caller drops to the classifier cascade.
const routeIntentUnknown = "unknown"
type routeAction struct {
Intent string `json:"intent"`
Key string `json:"key"`
@@ -73,8 +111,18 @@ type routeAction struct {
Verb string `json:"verb"`
}
// Route asks the model for one decision. The bool is false when there is no
// decision to use: either the model failed (err set) or it refused with the
// "unknown" intent (err nil). Both mean the same thing to the caller — use the
// classifier instead.
func (lr *LLMRouter) Route(ctx context.Context, utterance string, now time.Time) (Decision, bool, error) {
raw, err := lr.c.Complete(ctx, llm.Req{System: routeSystem, User: utterance, Grammar: routeGrammar, MaxTokens: 128})
raw, err := lr.c.Complete(ctx, llm.Req{
System: routeSystem,
User: utterance,
Grammar: routeGrammar,
MaxTokens: 128,
RepeatPenalty: routeRepeatPenalty,
})
if err != nil {
return Decision{}, false, err
}
@@ -90,6 +138,13 @@ func (lr *LLMRouter) Route(ctx context.Context, utterance string, now time.Time)
// with the engine turn-on (Router.Route → []Decision, both voice.go handlers
// loop). Until then only the first ask is honored.
a := acts[0]
// The model refused. Report "no decision" without an error, which is the
// same fall-through the caller already uses for a parse failure — the
// classifier cascade gets the turn and its own confidence gate decides
// whether to ask. Better a slower second opinion than a confident guess.
if a.Intent == routeIntentUnknown {
return Decision{}, false, nil
}
d := Decision{Utterance: utterance, Stage: 1, Confidence: 1.0}
switch Intent(a.Intent) {
case IntentFact:
+190 -3
View File
@@ -3,15 +3,61 @@ package router
import (
"context"
"fmt"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/llm"
)
type mockLLM struct{ out string; err error }
type mockLLM struct {
out string
err error
got *llm.Req // last request, when the test wants to inspect it
}
func (m mockLLM) Complete(_ context.Context, _ llm.Req) (string, error) { return m.out, m.err }
func (m mockLLM) Complete(_ context.Context, r llm.Req) (string, error) {
if m.got != nil {
*m.got = r
}
return m.out, m.err
}
// Without a repeat penalty the model loops inside the text field until MaxTokens
// and the truncated JSON fails to parse.
func TestLLMRouterSetsRepeatPenalty(t *testing.T) {
var got llm.Req
lr := NewLLMRouter(mockLLM{out: `{"intent":"chat","text":"привет"}`, got: &got})
if _, _, err := lr.Route(context.Background(), "привет", time.Now()); err != nil {
t.Fatalf("route: %v", err)
}
if got.RepeatPenalty <= 1 {
t.Fatalf("want repeat penalty above 1, got %v", got.RepeatPenalty)
}
}
// An unbounded string rule lets one field eat the whole token budget.
func TestRouteGrammarBoundsStrings(t *testing.T) {
if !strings.Contains(routeGrammar, `string ::= "\"" ([^"\\] | "\\" .){0,120} "\""`) {
t.Fatal("grammar string rule lost its length bound")
}
}
// A question naming a fact key used to be stored as a fact because the fact rule
// was tested first. Keep the query rule above it.
func TestRoutePromptTestsQueryBeforeFact(t *testing.T) {
query := strings.Index(routeSystem, "→ query")
fact := strings.Index(routeSystem, "состояние/событие → fact")
if query < 0 || fact < 0 {
t.Fatalf("prompt lost a rule: query=%d fact=%d", query, fact)
}
if query > fact {
t.Fatal("query rule must come before the fact rule")
}
if !strings.Contains(routeSystem, "Задаёт вопрос") {
t.Fatal("prompt lost the explicit question test")
}
}
func TestLLMRouterFactMapping(t *testing.T) {
lr := NewLLMRouter(mockLLM{out: `{"intent":"fact","key":"water","value":"выпил"}`})
@@ -54,8 +100,10 @@ func TestLLMRouterReminderMapping(t *testing.T) {
}
}
// An intent name that is not in the contract at all (as opposed to "unknown",
// which is a real refusal) still defaults to chat.
func TestLLMRouterChatFallback(t *testing.T) {
lr := NewLLMRouter(mockLLM{out: `{"intent":"unknown"}`})
lr := NewLLMRouter(mockLLM{out: `{"intent":"banana"}`})
d, ok, err := lr.Route(context.Background(), "как дела?", time.Now())
if err != nil || !ok {
t.Fatalf("ok=%v err=%v", ok, err)
@@ -65,6 +113,61 @@ func TestLLMRouterChatFallback(t *testing.T) {
}
}
// The model must be able to say "I could not route this".
func TestRouteGrammarAllowsUnknown(t *testing.T) {
if !strings.Contains(routeGrammar, `"\"unknown\""`) {
t.Fatal("grammar cannot express a refusal")
}
}
// If the prompt does not tell the model when to refuse, it never will.
func TestRoutePromptExplainsUnknown(t *testing.T) {
if !strings.Contains(routeSystem, "unknown") {
t.Fatal("prompt never mentions the unknown intent")
}
if !strings.Contains(routeSystem, `"сделай это" → {"intent":"unknown"}`) {
t.Fatal("prompt lost its worked refusal example")
}
// A refusal-only router is useless, so the prompt must also show cases that
// look ambiguous but are not.
if !strings.Contains(routeSystem, "здесь unknown не нужен") {
t.Fatal("prompt lost its counter-examples")
}
}
// A refusal is not an error. It reports "no decision" so the cascade moves on.
func TestLLMRouterUnknownRefuses(t *testing.T) {
lr := NewLLMRouter(mockLLM{out: `{"intent":"unknown"}`})
_, ok, err := lr.Route(context.Background(), "сделай это", time.Now())
if ok {
t.Fatal("a refusal must not produce a usable decision")
}
if err != nil {
t.Fatalf("a refusal is not an error, got %v", err)
}
}
// The whole point of the refusal: the turn keeps going on the classifier, the
// same way it does when the model returns garbage.
func TestRouterFallsBackWhenLLMRefuses(t *testing.T) {
c := NewClassifier(NewHashEmbedder(1024))
seedClassifier(t, c)
r := New(Config{
Classifier: c,
Extractor: Extractor{Time: StubDateTimeParser{}, Facts: DefaultFactParser{}},
Threshold: 0.4,
LLM: NewLLMRouter(mockLLM{out: `{"intent":"unknown"}`}),
})
d, err := r.Route(context.Background(), "напомни позвонить маме", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
// Stage 1 is the LLM's own answer; the classifier lands on stage 2 or 3.
if d.Stage < 2 {
t.Fatalf("want the classifier to decide, got stage %d (%+v)", d.Stage, d)
}
}
func TestLLMRouterLLMError(t *testing.T) {
lr := NewLLMRouter(mockLLM{out: "", err: fmt.Errorf("llm down")})
_, ok, err := lr.Route(context.Background(), "x", time.Now())
@@ -72,3 +175,87 @@ func TestLLMRouterLLMError(t *testing.T) {
t.Fatal("want ok=false, err!=nil on llm error")
}
}
// --- slot extraction on top of an LLM decision --------------------------------
// newLLMTestRouter — a router whose route always comes from the mock model.
func newLLMTestRouter(t *testing.T, out string) *Router {
t.Helper()
c := NewClassifier(NewHashEmbedder(1024))
seedClassifier(t, c)
acts := DefaultActMatcher{Fns: []string{"restart", "stop", "run", "backup"}}
return New(Config{
Classifier: c,
Extractor: Extractor{Time: StubDateTimeParser{}, Acts: acts, Facts: DefaultFactParser{}},
Threshold: 0.4,
LLM: NewLLMRouter(mockLLM{out: out}),
})
}
// The model cannot produce a fire time, so without extraction every LLM-routed
// reminder was dropped as "no time".
func TestLLMDecisionGetsReminderTime(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"reminder","text":"позвонить маме"}`)
d, err := r.Route(context.Background(), "напомни позвонить маме через 2 часа", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if d.Intent != IntentReminder {
t.Fatalf("want reminder, got %v", d.Intent)
}
if !d.Slots.HasTime || !d.Slots.Time.Equal(refNow().Add(2*time.Hour)) {
t.Fatalf("want time now+2h, got %+v", d.Slots)
}
if d.Slots.Text != "позвонить маме" {
t.Fatalf("extraction overwrote the model's text: %q", d.Slots.Text)
}
}
// No time in the utterance ⇒ no time in the slots. Do not invent one; the
// daemon says it could not read the time.
func TestLLMReminderWithoutTimeStaysEmpty(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"reminder","text":"позвонить маме"}`)
d, err := r.Route(context.Background(), "напомни позвонить маме", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if d.Slots.HasTime {
t.Fatalf("invented a time: %v", d.Slots.Time)
}
}
// An act decision arrived with no Fn, so the tool never ran.
func TestLLMDecisionGetsActFn(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"act","verb":"restart nginx"}`)
d, err := r.Route(context.Background(), "слушай, restart nginx пожалуйста", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if !d.Slots.HasFn || d.Slots.Fn != "restart" || len(d.Slots.Args) != 1 || d.Slots.Args[0] != "nginx" {
t.Fatalf("want fn=restart args=[nginx], got %+v", d.Slots)
}
}
// The model's own slots win; extraction only fills gaps.
func TestLLMSlotsWinOverExtraction(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"fact","key":"hydration","value":"выпил"}`)
d, err := r.Route(context.Background(), "я выпил воду", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if d.Slots.Key != "hydration" {
t.Fatalf("extraction overwrote the model's key: %q", d.Slots.Key)
}
}
// A fact the model left keyless still gets one from the parser.
func TestLLMFactGetsKeyFromParser(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"fact","text":"я выпил воду"}`)
d, err := r.Route(context.Background(), "я выпил воду", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if !d.Slots.HasKey || d.Slots.Key != "water" {
t.Fatalf("want key=water, got %+v", d.Slots)
}
}
+28 -1
View File
@@ -12,6 +12,15 @@ import (
"golang.org/x/text/unicode/norm"
)
// The deployed model is multilingual-e5-small. e5 was trained with these two
// words glued to the front of every text, and it scores badly without them —
// they are part of the model, not a style choice. Swapping back to a symmetric
// paraphrase model means dropping them again.
const (
queryPrefix = "query: "
passagePrefix = "passage: "
)
const (
padTokenID = 1
unkTokenID = 3
@@ -54,7 +63,25 @@ func NewONNXEmbedder(modelPath, tokenizerPath, libPath string) (*onnxEmbedder, e
func (e *onnxEmbedder) Dim() int { return embedDim }
// Embed treats the text as a query. The classifier compares one short
// utterance to another short seed phrase, so both sides get the same prefix
// and the comparison stays fair. The recall path must call EmbedQuery and
// EmbedPassage instead.
func (e *onnxEmbedder) Embed(ctx context.Context, text string) ([]float32, error) {
return e.embed(ctx, queryPrefix+text)
}
// EmbedQuery — the question the user just asked.
func (e *onnxEmbedder) EmbedQuery(ctx context.Context, text string) ([]float32, error) {
return e.embed(ctx, queryPrefix+text)
}
// EmbedPassage — a note or fact being stored, or re-scored at lookup time.
func (e *onnxEmbedder) EmbedPassage(ctx context.Context, text string) ([]float32, error) {
return e.embed(ctx, passagePrefix+text)
}
func (e *onnxEmbedder) embed(ctx context.Context, text string) ([]float32, error) {
inputIDs, attentionMask, _ := e.tokenizer.Encode(text)
inputShape := ort.NewShape(1, int64(maxLength))
@@ -313,4 +340,4 @@ func preTokenize(text string) []string {
return out
}
var _ Embedder = (*onnxEmbedder)(nil)
var _ AsymmetricEmbedder = (*onnxEmbedder)(nil)
+34
View File
@@ -88,6 +88,7 @@ func (r *Router) Route(ctx context.Context, utterance string, now time.Time) (De
if r.llm != nil {
if d, ok, err := r.llm.Route(ctx, utterance, now); err == nil && ok {
d.Utterance = utterance
r.fillSlots(ctx, &d, now)
return d, nil
} else if err != nil {
log.Printf("router: llm route fell back to classifier: %v", err)
@@ -118,6 +119,39 @@ func (r *Router) Route(ctx context.Context, utterance string, now time.Time) (De
return d, nil
}
// fillSlots — run stage-2 extraction on an LLM decision and fill only the slots
// the model left empty. The LLM wins where it answered: it saw the sentence, the
// parsers are keyword tables. Extraction covers what the model cannot produce at
// all — a parsed reminder time and an allowlist fn.
//
// If a reminder still has no time, leave it missing. The daemon then says it
// could not read the time; inventing one would set a wrong alarm.
func (r *Router) fillSlots(ctx context.Context, d *Decision, now time.Time) {
ex := r.extractor.Extract(ctx, d.Intent, d.Utterance, now)
if !d.Slots.HasTime && ex.HasTime {
d.Slots.Time, d.Slots.HasTime = ex.Time, ex.HasTime
}
if !d.Slots.HasKey && ex.HasKey {
d.Slots.Key, d.Slots.Value, d.Slots.HasKey = ex.Key, ex.Value, ex.HasKey
}
if !d.Slots.HasFn && ex.HasFn {
d.Slots.Fn, d.Slots.Args, d.Slots.HasFn = ex.Fn, ex.Args, ex.HasFn
}
// For an act the model returns the verb in Text ("restart nginx"), which is
// often cleaner than the raw utterance ("maven, could you restart nginx").
// Try it too when the utterance did not match the allowlist.
if d.Intent == IntentAct && !d.Slots.HasFn && r.extractor.Acts != nil &&
d.Slots.Text != "" && d.Slots.Text != d.Utterance {
if fn, args, ok := r.extractor.Acts.Match(d.Slots.Text); ok {
d.Slots.Fn, d.Slots.Args, d.Slots.HasFn = fn, args, true
}
}
if d.Slots.Text == "" {
d.Slots.Text = ex.Text
}
// Stage stays 1: it says who decided the route, and that was the LLM.
}
// CorrectMisroute — the user corrected a bad classification. Appends a new
// example for the corrected intent (append-only — grows the classifier, no
// retrain). Same shape as nudges.outcome tuning cooldowns: more reliable over
+41
View File
@@ -55,6 +55,47 @@ func Validate(routines []Routine) error {
return nil
}
// Accepted — an accepted routine proposal as the tick driver sees it. This is a
// different shape from Routine: the schedule is a plain interval the pattern
// detector measured, not an operator-written cron expression. Accepted is when
// the human said yes; LastFired is nil until the first nudge.
type Accepted struct {
ID int64
Name string
IntervalDays float64
Accepted time.Time
LastFired *time.Time
}
// DueAccepted returns the accepted routines whose interval has passed. It does
// not mutate anything — the caller persists the new last-fired time, because
// that has to survive a restart (unlike Due's in-memory map).
//
// The clock starts at LastFired, or at Accepted for a routine that has never
// nudged. A routine with a non-positive interval never fires: a bad interval
// should mean silence, not a nudge every tick.
//
// One occurrence per call, no catch-up: the caller stamps the fire time as now,
// so a routine that was silent for a month nudges once and then waits a full
// interval. Never a backlog.
func DueAccepted(rs []Accepted, now time.Time) []Accepted {
var out []Accepted
for _, r := range rs {
if r.IntervalDays <= 0 {
continue
}
since := r.Accepted
if r.LastFired != nil {
since = *r.LastFired
}
gap := time.Duration(r.IntervalDays * 24 * float64(time.Hour))
if !now.Before(since.Add(gap)) {
out = append(out, r)
}
}
return out
}
// Due returns the routines whose schedule crossed since their last fire and
// records now as the new last-fire time for each one returned. The caller owns
// `last` (the tick driver holds it across ticks); Due mutates it in place.
+34
View File
@@ -5,6 +5,40 @@ import (
"time"
)
func TestDueAcceptedFiresOncePerInterval(t *testing.T) {
accepted := time.Date(2026, 7, 1, 9, 0, 0, 0, time.UTC)
fired := accepted.Add(3 * 24 * time.Hour)
rs := []Accepted{
{ID: 1, Name: "полить цветы", IntervalDays: 3, Accepted: accepted},
{ID: 2, Name: "покормить рыб", IntervalDays: 3, Accepted: accepted, LastFired: &fired},
{ID: 3, Name: "битый интервал", IntervalDays: 0, Accepted: accepted},
}
// One day in: nothing has waited a full interval.
if got := DueAccepted(rs, accepted.Add(24*time.Hour)); len(got) != 0 {
t.Fatalf("want nothing due after 1 day, got %+v", got)
}
// Three days in: the never-fired one is due. The one that already fired at
// day 3 starts its next three days from there. A zero interval never fires.
got := DueAccepted(rs, fired)
if len(got) != 1 || got[0].ID != 1 {
t.Fatalf("want only routine 1 due at day 3, got %+v", got)
}
// Six days in: both real routines are due.
if got := DueAccepted(rs, accepted.Add(6*24*time.Hour)); len(got) != 2 {
t.Fatalf("want both routines due at day 6, got %+v", got)
}
// A month later the zero-interval routine is still silent.
for _, r := range DueAccepted(rs, accepted.Add(30*24*time.Hour)) {
if r.ID == 3 {
t.Fatal("a routine with a zero interval must never fire")
}
}
}
func TestValidate(t *testing.T) {
ok := []Routine{{Name: "morning", Cron: "0 8 * * *", Body: "доброе утро"}}
if err := Validate(ok); err != nil {
+4
View File
@@ -70,6 +70,10 @@ ALTER TABLE reminders ADD COLUMN next_fire_ts INTEGER;`, // #2
CHECK (resolution_state IN ('none','pending','resolved','ambiguous','not_found'));
CREATE INDEX IF NOT EXISTS idx_facts_entity_id ON facts (entity_id) WHERE entity_id IS NOT NULL;
CREATE INDEX IF NOT EXISTS idx_facts_resolution_pending ON facts (resolution_state) WHERE resolution_state = 'pending';`, // #7 — entity-aware memory (Vikunja #279): facts about a subject get resolved to a Nexus entity_id async
`CREATE INDEX IF NOT EXISTS idx_nudges_snoozed ON nudges (outcome_ts) WHERE outcome = 'snoozed';`, // #8 — SnoozedUntil runs every tick; keep it off a full scan (Vikunja #364)
`ALTER TABLE proposed_routines ADD COLUMN accepted_ts INTEGER;
ALTER TABLE proposed_routines ADD COLUMN last_fired_ts INTEGER;`, // #9 — accepted routines keep firing (Vikunja #366): the tick loop needs to know when a routine was accepted and when it last nudged
}
// migrate applies every migration with a number greater than the DB's current
+45
View File
@@ -28,6 +28,20 @@ const (
NudgeIgnored = "ignored"
)
// SnoozeDuration — how long one `snoozed` outcome keeps its rule quiet.
//
// The nudges table records THAT a snooze happened and when, never for how
// long: nothing upstream can supply a length. ResolveNudge takes only
// (id, outcome, ts), and so do the IPC method and the web/telegram callers
// behind it. So a fixed default it is, rather than a new column no writer
// could fill.
//
// Two hours: longer than every rule's base cooldown (1560m) so a snooze
// actually buys quiet instead of being swallowed by the cooldown, and short
// enough that a snooze the operator forgets about clears the same day. A
// snooze can never outlive this window, so Maven cannot go quiet forever.
const SnoozeDuration = 2 * time.Hour
var (
ErrNudgeNotFound = errors.New("store: nudge not found")
ErrNudgeOutcome = errors.New("store: nudge already resolved")
@@ -138,6 +152,37 @@ func (s *Store) UnackedTelegramRules(ctx context.Context) ([]string, error) {
return out, rows.Err()
}
// SnoozedUntil — per rule, when its most recent snooze runs out. This is the
// read behind the gate's snooze check: the `snoozed` outcome already in the
// nudges table IS the restraint memory, so there is no snooze table.
//
// Rules with no live snooze are absent from the map, which is what the gate
// wants (a missing key means "not snoozed"). Expired snoozes are filtered out
// in SQL, so an old snooze can never come back as a silent forever-mute.
//
// Called every tick (~60s). One indexed lookup over the snoozed rows only.
func (s *Store) SnoozedUntil(ctx context.Context, now time.Time) (map[string]time.Time, error) {
cutoff := now.Add(-SnoozeDuration).UnixMilli()
rows, err := s.db.QueryContext(ctx,
`SELECT rule, MAX(outcome_ts) FROM nudges
WHERE outcome = 'snoozed' AND outcome_ts > ?
GROUP BY rule`, cutoff)
if err != nil {
return nil, fmt.Errorf("snoozed until: %w", err)
}
defer rows.Close()
out := make(map[string]time.Time)
for rows.Next() {
var rule string
var tsMilli int64
if err := rows.Scan(&rule, &tsMilli); err != nil {
return nil, err
}
out[rule] = time.UnixMilli(tsMilli).UTC().Add(SnoozeDuration)
}
return out, rows.Err()
}
// RecentNudges — the newest n nudges across all rules, with outcomes, for the
// monitoring dash. Newest first.
func (s *Store) RecentNudges(ctx context.Context, n int) ([]Nudge, error) {
+104
View File
@@ -0,0 +1,104 @@
package store
import (
"context"
"testing"
"time"
)
// snoozeNudge records a nudge and immediately snoozes it at ts.
func snoozeNudge(t *testing.T, s *Store, rule string, ts time.Time) {
t.Helper()
ctx := context.Background()
id, err := s.RecordNudge(ctx, rule, "voice", "drink water", ts)
if err != nil {
t.Fatalf("RecordNudge: %v", err)
}
if err := s.ResolveNudge(ctx, id, NudgeSnoozed, ts); err != nil {
t.Fatalf("ResolveNudge: %v", err)
}
}
func TestSnoozedUntilPerRule(t *testing.T) {
s := newTestStore(t)
now := time.Now().UTC().Truncate(time.Millisecond)
snoozeNudge(t, s, "water", now.Add(-10*time.Minute))
snoozeNudge(t, s, "break", now.Add(-30*time.Minute))
got, err := s.SnoozedUntil(context.Background(), now)
if err != nil {
t.Fatalf("SnoozedUntil: %v", err)
}
if len(got) != 2 {
t.Fatalf("want 2 snoozed rules, got %v", got)
}
wantWater := now.Add(-10 * time.Minute).Add(SnoozeDuration)
if !got["water"].Equal(wantWater) {
t.Fatalf("water until = %v, want %v", got["water"], wantWater)
}
}
// The map must only ever hold the newest snooze for a rule, so a stale one
// can't shorten (or lengthen) the live one.
func TestSnoozedUntilUsesNewestSnooze(t *testing.T) {
s := newTestStore(t)
now := time.Now().UTC().Truncate(time.Millisecond)
snoozeNudge(t, s, "water", now.Add(-90*time.Minute))
snoozeNudge(t, s, "water", now.Add(-5*time.Minute))
got, err := s.SnoozedUntil(context.Background(), now)
if err != nil {
t.Fatalf("SnoozedUntil: %v", err)
}
want := now.Add(-5 * time.Minute).Add(SnoozeDuration)
if !got["water"].Equal(want) {
t.Fatalf("water until = %v, want %v", got["water"], want)
}
}
// A snooze must expire. If this ever regresses Maven goes quiet forever and
// nobody can tell why.
func TestSnoozedUntilExpires(t *testing.T) {
s := newTestStore(t)
now := time.Now().UTC().Truncate(time.Millisecond)
snoozeNudge(t, s, "water", now.Add(-SnoozeDuration-time.Minute))
got, err := s.SnoozedUntil(context.Background(), now)
if err != nil {
t.Fatalf("SnoozedUntil: %v", err)
}
if _, ok := got["water"]; ok {
t.Fatalf("expired snooze still active: %v", got)
}
}
// Other outcomes are not snoozes.
func TestSnoozedUntilIgnoresOtherOutcomes(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
now := time.Now().UTC().Truncate(time.Millisecond)
for _, outcome := range []string{NudgeActed, NudgeIgnored} {
id, err := s.RecordNudge(ctx, "water", "voice", "drink water", now)
if err != nil {
t.Fatalf("RecordNudge: %v", err)
}
if err := s.ResolveNudge(ctx, id, outcome, now); err != nil {
t.Fatalf("ResolveNudge: %v", err)
}
}
if _, err := s.RecordNudge(ctx, "break", "voice", "stand up", now); err != nil {
t.Fatalf("RecordNudge: %v", err)
}
got, err := s.SnoozedUntil(ctx, now)
if err != nil {
t.Fatalf("SnoozedUntil: %v", err)
}
if len(got) != 0 {
t.Fatalf("want no snoozes, got %v", got)
}
}
+112 -19
View File
@@ -8,10 +8,23 @@ import (
"time"
)
// ProposedRoutine — a detected pattern the system wants to turn into a
// recurring reminder. Status 'proposed' means awaiting human confirmation;
// 'accepted' means the human confirmed and a reminder was created (reminder_id
// set); 'dismissed' means the human declined and we won't re-propose.
// The three states a proposal can be in. A proposal starts 'proposed' and
// moves once, either way, and never moves again.
const (
RoutineProposed = "proposed"
RoutineAccepted = "accepted"
RoutineDismissed = "dismissed"
)
// ProposedRoutine — a detected pattern the system wants to nudge about on a
// repeating interval. Status 'proposed' means awaiting human confirmation;
// 'accepted' means the human confirmed and the tick loop now owns the schedule;
// 'dismissed' means the human declined and we won't re-propose.
//
// AcceptedTs is when the human said yes; it is the clock start for the first
// nudge. LastFiredTs is when the last nudge went out, nil until the first one.
// ReminderID is only set on rows accepted before Vikunja #366, when accepting
// created a one-shot reminder instead.
type ProposedRoutine struct {
ID int64
Action string
@@ -19,7 +32,9 @@ type ProposedRoutine struct {
IntervalDays float64
Status string // proposed | accepted | dismissed
CreatedTs time.Time
ReminderID *int64 // set when accepted
ReminderID *int64
AcceptedTs *time.Time
LastFiredTs *time.Time
}
var (
@@ -30,6 +45,15 @@ var (
// CreateProposedRoutine inserts a new proposed routine. Returns
// ErrProposedRoutineExists if one already exists for this action+object (any
// status) — the pattern detector should only propose once per pair.
//
// action+object is the "same routine" key. It is UNIQUE in the table, so a
// routine the human already dismissed can never come back: the detector will
// keep finding the pattern, and every re-propose is refused here. Maven is not
// a nag.
//
// TODO(vikunja#46): the detector currently only writes here from the voice
// path. Once digestion runs the detector on its own tick, that tick should
// call this too, so a pattern gets noticed even with nobody at the mic.
func (s *Store) CreateProposedRoutine(ctx context.Context, action, object string, intervalDays float64, ts time.Time) (int64, error) {
res, err := s.db.ExecContext(ctx,
`INSERT INTO proposed_routines (action, object, interval_days, status, created_ts)
@@ -57,7 +81,7 @@ func (s *Store) CreateProposedRoutine(ctx context.Context, action, object string
// nil (no error) when no row exists.
func (s *Store) LookupProposedRoutine(ctx context.Context, action, object string) (*ProposedRoutine, error) {
row := s.db.QueryRowContext(ctx, `
SELECT id, action, object, interval_days, status, created_ts, reminder_id
SELECT id, action, object, interval_days, status, created_ts, reminder_id, accepted_ts, last_fired_ts
FROM proposed_routines
WHERE action = ? AND object = ?`, action, object)
r, err := scanProposedRoutine(row)
@@ -70,14 +94,25 @@ func (s *Store) LookupProposedRoutine(ctx context.Context, action, object string
return &r, nil
}
// ListProposedRoutines returns all proposed routines with status='proposed',
// newest first.
// ListProposedRoutines returns the routines still waiting for an answer,
// newest first. This is what the /routines page shows.
func (s *Store) ListProposedRoutines(ctx context.Context) ([]ProposedRoutine, error) {
rows, err := s.db.QueryContext(ctx, `
SELECT id, action, object, interval_days, status, created_ts, reminder_id
FROM proposed_routines
WHERE status = 'proposed'
ORDER BY created_ts DESC, id DESC`)
return s.ListProposedRoutinesByStatus(ctx, RoutineProposed)
}
// ListProposedRoutinesByStatus returns routines in one status, newest first.
// An empty status returns every row.
func (s *Store) ListProposedRoutinesByStatus(ctx context.Context, status string) ([]ProposedRoutine, error) {
q := `SELECT id, action, object, interval_days, status, created_ts, reminder_id, accepted_ts, last_fired_ts
FROM proposed_routines`
var args []any
if status != "" {
q += ` WHERE status = ?`
args = append(args, status)
}
q += ` ORDER BY created_ts DESC, id DESC`
rows, err := s.db.QueryContext(ctx, q, args...)
if err != nil {
return nil, fmt.Errorf("list proposed routines: %w", err)
}
@@ -93,12 +128,22 @@ func (s *Store) ListProposedRoutines(ctx context.Context) ([]ProposedRoutine, er
return out, rows.Err()
}
// AcceptProposedRoutine flips status to 'accepted', links a reminder_id.
// Returns error if not in 'proposed' status.
func (s *Store) AcceptProposedRoutine(ctx context.Context, id, reminderID int64) error {
// Status changes below are an in-place UPDATE, on purpose. Facts are
// append-only (a correction writes a new row and sets voids_id) because a fact
// is a claim about the world and the old claim is still history worth keeping.
// A proposal is not a claim, it is a question with one answer, and the same
// shape already exists for tools (tools.status flips in place). The guard
// `AND status = 'proposed'` makes the move one-way: an answered proposal can
// never be answered again.
//
// AcceptProposedRoutine flips status to 'accepted' and records when. From that
// timestamp the tick loop owns the schedule: it re-reads accepted rows every
// tick and nudges when the interval has passed. Returns an error if the row is
// not in 'proposed' status.
func (s *Store) AcceptProposedRoutine(ctx context.Context, id int64, ts time.Time) error {
res, err := s.db.ExecContext(ctx,
`UPDATE proposed_routines SET status = 'accepted', reminder_id = ? WHERE id = ? AND status = 'proposed'`,
reminderID, id)
`UPDATE proposed_routines SET status = 'accepted', accepted_ts = ? WHERE id = ? AND status = 'proposed'`,
ts.UnixMilli(), id)
if err != nil {
return fmt.Errorf("accept proposed routine: %w", err)
}
@@ -109,6 +154,42 @@ func (s *Store) AcceptProposedRoutine(ctx context.Context, id, reminderID int64)
return nil
}
// ListAcceptedRoutines returns every accepted routine, oldest first. The tick
// loop reads this each tick and decides which ones are due.
func (s *Store) ListAcceptedRoutines(ctx context.Context) ([]ProposedRoutine, error) {
rows, err := s.db.QueryContext(ctx, `
SELECT id, action, object, interval_days, status, created_ts, reminder_id, accepted_ts, last_fired_ts
FROM proposed_routines
WHERE status = 'accepted'
ORDER BY id`)
if err != nil {
return nil, fmt.Errorf("list accepted routines: %w", err)
}
defer rows.Close()
var out []ProposedRoutine
for rows.Next() {
r, err := scanProposedRoutine(rows)
if err != nil {
return nil, err
}
out = append(out, r)
}
return out, rows.Err()
}
// MarkRoutineFired records that a routine just nudged. The stored time is the
// nudge time, not the time it was theoretically due, so a routine that was
// silent for a while starts its next interval from now — missed occurrences are
// dropped, never replayed as a backlog.
func (s *Store) MarkRoutineFired(ctx context.Context, id int64, ts time.Time) error {
if _, err := s.db.ExecContext(ctx,
`UPDATE proposed_routines SET last_fired_ts = ? WHERE id = ?`,
ts.UnixMilli(), id); err != nil {
return fmt.Errorf("mark routine fired: %w", err)
}
return nil
}
// DismissProposedRoutine flips status to 'dismissed'. Idempotent.
func (s *Store) DismissProposedRoutine(ctx context.Context, id int64) error {
_, err := s.db.ExecContext(ctx,
@@ -125,12 +206,24 @@ func scanProposedRoutine(sc scanner) (ProposedRoutine, error) {
var r ProposedRoutine
var created int64
var reminderID sql.NullInt64
if err := sc.Scan(&r.ID, &r.Action, &r.Object, &r.IntervalDays, &r.Status, &created, &reminderID); err != nil {
var accepted, lastFired sql.NullInt64
if err := sc.Scan(&r.ID, &r.Action, &r.Object, &r.IntervalDays, &r.Status, &created, &reminderID, &accepted, &lastFired); err != nil {
return ProposedRoutine{}, err
}
r.CreatedTs = time.UnixMilli(created).UTC()
if reminderID.Valid {
r.ReminderID = &reminderID.Int64
}
r.AcceptedTs = millisToTime(accepted)
r.LastFiredTs = millisToTime(lastFired)
return r, nil
}
// millisToTime turns a nullable unix-millis column into a *time.Time.
func millisToTime(v sql.NullInt64) *time.Time {
if !v.Valid {
return nil
}
t := time.UnixMilli(v.Int64).UTC()
return &t
}
+131 -8
View File
@@ -43,12 +43,7 @@ func TestCreateAndAcceptProposedRoutine(t *testing.T) {
}
// Accept
// First create a reminder to link
remID, err := s.CreateReminder(ctx, now.Add(7*24*time.Hour), `{"text":"refill cat water"}`, "0 10 * * 0")
if err != nil {
t.Fatalf("CreateReminder: %v", err)
}
if err := s.AcceptProposedRoutine(ctx, id, remID); err != nil {
if err := s.AcceptProposedRoutine(ctx, id, now); err != nil {
t.Fatalf("AcceptProposedRoutine: %v", err)
}
@@ -60,8 +55,50 @@ func TestCreateAndAcceptProposedRoutine(t *testing.T) {
if r.Status != "accepted" {
t.Fatalf("want status=accepted, got %s", r.Status)
}
if r.ReminderID == nil || *r.ReminderID != remID {
t.Fatalf("want reminder_id=%d, got %v", remID, r.ReminderID)
if r.AcceptedTs == nil || !r.AcceptedTs.Equal(now.Truncate(time.Millisecond)) {
t.Fatalf("want accepted_ts=%v, got %v", now, r.AcceptedTs)
}
if r.LastFiredTs != nil {
t.Fatalf("a freshly accepted routine has not fired yet, got %v", r.LastFiredTs)
}
}
// TestAcceptedRoutineFiredTimestamp — the tick loop's two reads: the accepted
// list, and the last-fired stamp it writes back after a nudge.
func TestAcceptedRoutineFiredTimestamp(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
now := time.Now().UTC().Truncate(time.Millisecond)
id, err := s.CreateProposedRoutine(ctx, "полить", "цветы", 3.0, now)
if err != nil {
t.Fatalf("CreateProposedRoutine: %v", err)
}
if err := s.AcceptProposedRoutine(ctx, id, now); err != nil {
t.Fatalf("AcceptProposedRoutine: %v", err)
}
list, err := s.ListAcceptedRoutines(ctx)
if err != nil {
t.Fatalf("ListAcceptedRoutines: %v", err)
}
if len(list) != 1 || list[0].ID != id {
t.Fatalf("want the one accepted routine, got %+v", list)
}
if list[0].IntervalDays != 3.0 {
t.Fatalf("want interval_days=3, got %v", list[0].IntervalDays)
}
fired := now.Add(3 * 24 * time.Hour)
if err := s.MarkRoutineFired(ctx, id, fired); err != nil {
t.Fatalf("MarkRoutineFired: %v", err)
}
list, err = s.ListAcceptedRoutines(ctx)
if err != nil {
t.Fatalf("ListAcceptedRoutines: %v", err)
}
if list[0].LastFiredTs == nil || !list[0].LastFiredTs.Equal(fired) {
t.Fatalf("want last_fired_ts=%v, got %v", fired, list[0].LastFiredTs)
}
}
@@ -131,6 +168,92 @@ func TestListProposedRoutines(t *testing.T) {
}
}
// A dismissed routine must never be proposed again. The detector will keep
// finding the same pattern; the store is what stops maven nagging about it.
func TestDismissedProposedRoutineStaysDismissed(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
now := time.Now().UTC()
id, err := s.CreateProposedRoutine(ctx, "clean", "litter_box", 3.0, now)
if err != nil {
t.Fatalf("CreateProposedRoutine: %v", err)
}
if err := s.DismissProposedRoutine(ctx, id); err != nil {
t.Fatalf("DismissProposedRoutine: %v", err)
}
// The detector re-proposes the same pattern.
_, err = s.CreateProposedRoutine(ctx, "clean", "litter_box", 3.0, now.Add(24*time.Hour))
if !errors.Is(err, ErrProposedRoutineExists) {
t.Fatalf("want ErrProposedRoutineExists on re-propose, got %v", err)
}
// And it must not reappear on the review page.
list, err := s.ListProposedRoutines(ctx)
if err != nil {
t.Fatalf("ListProposedRoutines: %v", err)
}
if len(list) != 0 {
t.Fatalf("want 0 proposed, got %d", len(list))
}
// Dismissing again is a no-op, and accepting is refused.
if err := s.DismissProposedRoutine(ctx, id); err != nil {
t.Fatalf("second DismissProposedRoutine: %v", err)
}
if err := s.AcceptProposedRoutine(ctx, id, time.Now().UTC()); !errors.Is(err, ErrProposedRoutineNotFound) {
t.Fatalf("want ErrProposedRoutineNotFound accepting a dismissed routine, got %v", err)
}
r, err := s.LookupProposedRoutine(ctx, "clean", "litter_box")
if err != nil {
t.Fatalf("LookupProposedRoutine: %v", err)
}
if r.Status != RoutineDismissed {
t.Fatalf("want status=dismissed, got %s", r.Status)
}
}
func TestListProposedRoutinesByStatus(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
now := time.Now().UTC()
keep, err := s.CreateProposedRoutine(ctx, "water", "plants", 4.0, now)
if err != nil {
t.Fatalf("CreateProposedRoutine: %v", err)
}
drop, err := s.CreateProposedRoutine(ctx, "walk", "dog", 1.0, now.Add(time.Hour))
if err != nil {
t.Fatalf("CreateProposedRoutine: %v", err)
}
if err := s.AcceptProposedRoutine(ctx, keep, now); err != nil {
t.Fatalf("AcceptProposedRoutine: %v", err)
}
if err := s.DismissProposedRoutine(ctx, drop); err != nil {
t.Fatalf("DismissProposedRoutine: %v", err)
}
cases := []struct {
status string
want int
}{
{RoutineProposed, 0},
{RoutineAccepted, 1},
{RoutineDismissed, 1},
{"", 2}, // empty status ⇒ every row
}
for _, c := range cases {
list, err := s.ListProposedRoutinesByStatus(ctx, c.status)
if err != nil {
t.Fatalf("ListProposedRoutinesByStatus(%q): %v", c.status, err)
}
if len(list) != c.want {
t.Fatalf("status %q: want %d, got %d", c.status, c.want, len(list))
}
}
}
func TestLookupMissingProposedRoutine(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
+6 -1
View File
@@ -55,4 +55,9 @@ func spokenDate(dd, mm, yyyy string) string {
}
func mustInt(s string) int { n, _ := strconv.Atoi(s); return n }
func gap(y string) string { if y == "" { return "" }; return " " + y }
func gap(y string) string {
if y == "" {
return ""
}
return " " + y
}
+14 -4
View File
@@ -7,12 +7,22 @@
set -euo pipefail
# The llama-server the phraser spawns has NO "maven" in its command line (its
# args are `-m /path/to/LFM2.5-...gguf --port ...`), so a `llama-server.*maven`
# args are `-m /path/to/<model>.gguf --port ...`), so a `llama-server.*maven`
# pattern matches nothing and leaks it — the exact bug that let orphans pile up
# and OOM the box. Match the model instead. Override MODEL if you change it.
MODEL="${MODEL:-LFM2}"
# and OOM the box.
#
# We used to match the model name, defaulting to LFM2. The deploy now runs
# Qwen3.5-0.8B, so that default matched nothing and the server survived every
# kill. Match any llama-server serving a .gguf instead, so swapping the model in
# deploy/mavend.json cannot break this script again. Set MODEL to narrow it if
# some other llama-server on this box must be left alone.
MODEL="${MODEL:-}"
PAT='mavend|mavsttd|mavttsd|mavweb|mavpoll|mavenclient'
LLM="llama-server.*${MODEL}"
if [ -n "$MODEL" ]; then
LLM="llama-server.*${MODEL}"
else
LLM='llama-server.*\.gguf'
fi
echo "--- Sending graceful SIGTERM to Maven services ---"
pkill -TERM -f "$PAT" || true