Commit Graph

204 Commits

Author SHA1 Message Date
kami 7bb9f9be06 Merge the embedder marker 2026-07-31 13:42:40 +04:00
kami 1e47eaca5a Record which embedder wrote the stored vectors and warn on a swap (#378)
The embedder moved from paraphrase-multilingual-MiniLM-L12-v2 to
multilingual-e5-small. Both are 384-dimensional, so nothing in the code
noticed: cosine between an old stored vector and a new query vector is
noise, and recall degrades silently.

So the DB now records the embedder that wrote its vectors. One value for
the whole DB (migration #11, a small `meta` key/value table) rather than a
column on every vector row: the backfill re-embeds every note and fact in
one pass, so a per-row marker would hold the same string in every row and
cost a column on two tables for nothing.

The identity comes from the embedder itself via a new optional ID() method
("multilingual-e5-small@384", model file name plus dimension), so pointing
the config at another model changes the string without anyone editing a
constant. mavend logs a loud WARNING at startup naming both the stored and
the configured embedder when they differ.

Detection only — recall behaviour is unchanged. TODO(#378) in
store.CheckEmbedder marks where the backfill will hook in.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:42:07 +04:00
kami 892330eb84 Merge the note recall fix 2026-07-31 13:33:14 +04:00
kami 9a3bcd7c46 Merge the thinking-off measurement 2026-07-31 13:31:31 +04:00
kami 98ee701e03 Let a note win a recall, not only a fact (#373)
The memory pass ran only after the notes-only gate had already rejected
the same note at the same score. Notes and facts share one vector index,
so a note that failed there failed again — the branch could only ever
return a fact.

Now the memory pass runs first: one search over everything Maven
remembers, one gate, and the memory that clearly matches best answers
(a note gets phrased, a fact is read back). The notes-only pass stays
behind it for notes the vector index does not hold. No threshold moved,
so the set of questions answered is unchanged — only which memory
answers them.

Fixture gained two mixed note+fact cases, so the answerable count goes
25 -> 27: hash recall@1 36.0% -> 37.0% (ratchet 0.32 unchanged, comment
updated), e5 recall@1 72.0% -> 70.4%, false recall still 1/5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:30:38 +04:00
kami 04c1088088 Measure thinking off on routing properly — it does not win (#376)
The 67.1% "thinking off" column in ROUTING-EVAL-31-07-2026.md was an
artefact. It came from a hand-rolled HTTP client in the eval test that
did not send repeat_penalty, so it differed from the reference run on two
axes and the penalty was the one that mattered.

Re-scored back to back on an idle box with everything else held equal:
thinking off is identical to thinking on, case for case, same confusion
matrix, same three unparseable replies. A direct probe of the running
llama-server shows enable_thinking, thinking and reasoning_budget are all
ignored for this model on this build, so there was nothing to turn off.

No defaults changed. The misleading third configuration is removed from
internal/router/eval so its table cannot be quoted again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:30:22 +04:00
kami 07c191d8b8 Merge dialogue session persistence 2026-07-31 13:21:03 +04:00
kami c668310b3e Persist the dialogue session so a restart keeps the conversation
Vikunja #363. The follow-up session was a plain in-memory map, so any
mavend restart dropped the thread. It now mirrors to a small TTL-pruned
sqlite table and is loaded on startup; expired sessions are deleted on
load, not revived. Clarify's pending question is untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:20:30 +04:00
kami 1bd2acdc2a Do not exempt Russian words that are both noun and verb 2026-07-31 12:55:45 +04:00
kami 15e5dd8eaa Merge the second-person gender check 2026-07-31 12:54:48 +04:00
kami 10cf6f525c Check that nudges do not address the owner in the feminine
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:54:18 +04:00
kami e2210f6844 Merge the clarify-expiry notice 2026-07-31 12:52:36 +04:00
kami 214a4032cf Tell him when an expired clarify question is dropped
Vikunja #382. A parked clarifying question past its TTL was discarded
silently on read; now she says the old request is gone and the newly
spoken words are still routed as a fresh utterance.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:51:47 +04:00
kami dc70a5a7ab Show clarify_max_attempts in the deployed config
The default is 3 either way. Writing it out means you can see the knob
without reading the Go.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:34:28 +04:00
kami 06aded6ab0 Merge commit 'd2be98e' into overnight-jul31
# Conflicts:
#	cmd/mavend/clarify.go
#	cmd/mavend/clarify_test.go
#	cmd/mavend/voice.go
#	internal/config/config.go
2026-07-31 12:34:00 +04:00
kami d2be98ee2a Say out loud when she gives up instead of dropping the request
An unclear answer used to end the request on the spot. Now she re-asks the same
question while attempts remain, and when they run out she says
"Прости, я не поняла. Скажи, пожалуйста, по-другому." — silence would leave him
thinking it was handled. Same reply when the missing slot has no question to
ask, and as a floor in finishClarified so an empty reply can never ship.

Tests: three questions allowed, the fourth gives up out loud, the cap is
configurable, and a restated time is the one that lands.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:31:55 +04:00
kami 62d320f93a Let her ask three times, and let a restated answer win
MaxAttempts was 1, justified as "not a nag". Wrong reading: "not a nag" is about
interrupting unprompted, and a clarifying question is part of a conversation he
started. Now three, configurable via voice.clarify_max_attempts (default 3).
Three, because after that the likely problem is she misheard the whole request,
not one slot.

Answer used to keep the parked value, so "в три" then "нет, в пять" threw the
five away. Now a value the answer carries wins for the slot she asked about.
Only for the clarify answer — a correction in a fresh turn is followUpMerge.

The eight-field chained assertion in the Answer test is one DeepEqual now, so a
new field in Slots is covered without touching the test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:31:46 +04:00
kami 796e6af3cf Merge commit '74a7088' into overnight-jul31 2026-07-31 12:24:20 +04:00
kami 74a70880a8 Write up the phrasing eval: 0/15 to 13/15, and what the number hides
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:23:54 +04:00
kami a40bc559d5 Fix the nudge phrasing prompt: stop teaching the model to echo the example
The system prompt showed the JSON contract as {"response": "..."} and the
user prompt repeated it. A 0.8B copies whatever sits in the response slot, so
7 of 15 nudges came back as literally "...".

Changes, all prompt-side — the {"response","mood"} contract is unchanged:
- nudge system prompt is Russian, feminine self-reference, with filled-in
  examples on topics that never appear as rules, so copying them is visible
- rule names get a Russian gloss and a required keyword, named last in the
  prompt where a small model weights it hardest
- durations render in Russian, not English
- the no-parse fallback says something Russian instead of "water — care",
  which was going straight to a Russian piper voice
- same "..." placeholder removed from replier_llm.go

Scored on internal/phraser/eval: 0/15 -> 13/15.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:23:03 +04:00
kami 1db0fcfcd0 Merge commit '4ca68d2' into overnight-jul31 2026-07-31 12:07:38 +04:00
kami 4ca68d2f3f Bake off LFM2.5 against Qwen3.5-0.8B on the RU routing fixture
Vikunja #278 / #250. Keep Qwen: LFM2.5-1.2B loses 8 points of intent
accuracy, all of it Russian, and runs 2.4x slower.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:07:06 +04:00
kami 5a1d465db5 Merge commit '4383844' into overnight-jul31
# Conflicts:
#	deploy/mavend.json
#	internal/config/config.go
2026-07-31 11:52:21 +04:00
kami 43838445ab Write up the margin gate results
Third section: why the absolute gate could not separate the two
distributions, the delta sweep, and the before/after. Marks next-steps
item 3 done.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:51:21 +04:00
kami 11831c6ace Gate recall on the margin over the runner-up, not just the score
The e5 embedder puts every cosine in one narrow band (0.79-0.89), so the
absolute query_min_score gate cannot tell a real hit from a made-up
question: any value under the band answers everything, any value above it
answers nothing. False recall was 5/5.

New gate asks whether one note is clearly the best instead: top1 - top2 >
delta. New query_min_margin config knob, default 0.008, read off the sweep
in the recall harness. The absolute floor stays as a second check.

On the recall fixture with e5: answered 72% -> 68%, false recall 5/5 -> 1/5.

Vikunja #359

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:50:39 +04:00
kami 93c1a41d4a Route with the resident model by default
The two things that made this unsafe are fixed: the router can now
refuse, and slot extraction runs on its decisions.

On the held-out fixture it gets 63.2% of intents right against the
classifier's 50.0%, with no route errors. It costs about a second a
turn instead of 30ms.

The flag is a pointer now, so leaving it out of the config means on
and only writing false turns it off.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:47:57 +04:00
kami bce5ed210c Merge branch 'worktree-agent-a76ce40c73601d90d' into overnight-jul31 2026-07-31 11:44:32 +04:00
kami c31f0d1001 Extract slots for LLM router decisions too
An LLM-routed reminder came back with no parsed time and an act with no
fn, because only the classifier path ran the extractor. Now the router
runs the same extraction after an LLM decision and fills only the empty
slots. No time in the utterance still means no time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:44:06 +04:00
kami 34521c30b8 Merge branch 'worktree-agent-ad5da57e47b822152' into overnight-jul31 2026-07-31 11:39:35 +04:00
kami 1d48755d12 Record the recall numbers after the embedder swap
recall@1 60% to 72%, answered 48% to 72%, latency 3x better. But false
recall went 1/5 to 5/5: e5 packs every score into a narrow high band, so
the 0.55 gate now admits everything. Left the gate alone as instructed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:38:18 +04:00
kami f6d5a2a7a4 Swap the embedder to multilingual-e5-small (Vikunja #371, #372)
The old model was a symmetric paraphrase model, so it scored "do these
look alike" instead of "does this note answer this question". Also fixes
the file mismatch: the Makefile, the deploy config and both evals now all
name the same quantized file, and the quantized one is what gets measured.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:38:18 +04:00
kami 751c2a705f Embed a question and a stored note differently (Vikunja #371)
Note recall is asymmetric: a short question goes in, a longer note comes
out. Adds EmbedQuery/EmbedPassage helpers and the e5 prefixes, and points
the note/fact write path at the passage side and the query path at the
query side. Reviewers: the three call sites in voice.go.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:38:07 +04:00
kami 5d5b0cfd49 Merge branch 'worktree-agent-a0ab2a9b7439296e3' into overnight-jul31 2026-07-31 11:37:53 +04:00
kami bd16ca69e5 Let the LLM router answer "unknown" when it cannot route
Chose an 8th enum value over a confidence number: the model already picks
one enum token, so it costs nothing in the grammar, while a score from a
0.8B model would be uncalibrated noise. A refusal returns "no decision"
with no error, which is the fall-through the caller already uses for a
bad parse, so the classifier and its clarify gate take the turn.

Reviewers: the prompt's counter-examples matter most — a small model will
over-use any easy escape hatch. The training workspace copy of the prompt
still needs the same edit (Vikunja #362).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:35:37 +04:00
kami 0ed386eca6 Re-measure the router on a quiet box and record the numbers
The earlier before/after was taken while another eval shared
llama-server. This run had the box to itself.

Intent accuracy 61.8% llm-only, 63.2% cascade, 67.1% with thinking off.
The prompt fix holds. note→fact shows up here too, so it is real.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:34:57 +04:00
kami 1c4eab2107 Merge commit '94eb92f' into overnight-jul31
# Conflicts:
#	Makefile
2026-07-31 10:11:56 +04:00
kami 0914e0a3d5 Merge commit '4ba9a6f' into overnight-jul31
# Conflicts:
#	Makefile
2026-07-31 10:11:29 +04:00
kami c860808528 Make make test actually gate on gofmt and vet
DESIGN.md has always said `make test` is "gofmt + vet + -race, no
exceptions". It only ever ran the tests, which is how nine files drifted
out of format without anyone noticing.

`test` now depends on `fmt-check` and `vet`. Checked that fmt-check does
fail when a file is unformatted, so the gate is real and not decorative.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 10:10:27 +04:00
kami f7442c3aea Run gofmt over the seven files that had drifted
Formatting only: import order, and statements that were packed onto one
line split out. `git diff -w` shows nothing but that.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 10:09:32 +04:00
kami 75b067ac51 Merge the accepted-routine fix and drop reminder_id from accept
Two merge fixes on top of the branch:

- migrations: keep both new steps, snooze stays #8, the routine columns
  become #9. Both agents had numbered theirs #8.
- accepting no longer takes a reminder id, on the web surface too. The
  web accept path had the same one-shot-reminder bug the voice path did,
  so both now just flip the status and let the tick loop schedule.

The test that asserted "accept creates a reminder and links it" asserted
the bug. It now asserts that accepting creates no reminder.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 10:07:37 +04:00
kami 424d1b3446 Fire accepted routines every interval, not once (Vikunja #366)
The tick loop now reads accepted routines from the store and nudges when
their interval has passed; accepting no longer builds a one-shot reminder.
Look at routine.DueAccepted for the schedule rule (no catch-up backlog) and
at fireAcceptedRoutines for the restraint gate — routines do not bypass it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:45:30 +04:00
kami c47886c2bc Merge branch 'worktree-agent-ab5b5c61a32cac4fe' into overnight-jul31 2026-07-31 02:42:53 +04:00
kami f6236da760 Collapse the duplicate away-detail and panic tests
Two agents wrote the same three test helpers and names for the same two
bugs. Kept the real assertions from dispatcher_test.go and removed the
skipped placeholders they replace. panicSink stays in durability_test.go
since both files use it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:42:21 +04:00
kami 54dc43516b Add accepted-routine timestamps to the store (Vikunja #366)
Data layer only. Migration #8 adds accepted_ts and last_fired_ts to
proposed_routines, plus ListAcceptedRoutines and MarkRoutineFired so the
tick loop can own the schedule. Accepting no longer links a reminder id.
Look at the TODO(vikunja#366) in cmd/mavend/tick.go for the next commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:41:58 +04:00
kami fa51a48958 Merge branch 'worktree-agent-afe3f2ec18b2b8497' into overnight-jul31 2026-07-31 02:40:10 +04:00
kami 5fd25d7ad7 Test the away-channel minimal body and the panicking sink (#368, #369)
The integration branch names one test TestAwayFallsBackToFullBodyWhenSummaryEmpty,
which describes the old bug; it is here as TestAwaySendsGenericLineWhenSummaryEmpty
and asserts the generic line instead of the body. Also covers: a normal summary
goes out unchanged, voice keeps the full body, and one panicking sink does not
eat the other channel for the same nudge. Reformatted one pre-existing struct.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:39:38 +04:00
kami c22fc352fc Merge branch 'worktree-agent-a1610b8c5376eadd6' into overnight-jul31 2026-07-31 02:37:30 +04:00
kami 215aa331c5 Recover from a panicking sink so the attempt is always closed (#369)
A panic in Send used to unwind past completeOutbox and leave the
delivery_attempts row pending forever, since reconciliation only runs at
startup. safeSend turns the panic into an error, logs it loudly, records the
attempt failed, and lets the other channels for the same nudge still go out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:35:54 +04:00
kami 859bbf750f Never send a nudge body off-box when the summary is empty (#368)
Away channels (ntfy, telegram) leave the box, so an empty Summary now sends
a fixed generic line plus the rule name instead of the full Body. The
dispatcher strips detail before any sink sees it, so a sink added later
cannot leak by reading the wrong field. Voice is local and unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:34:54 +04:00
kami 8a174c1c70 Score the recall fixture and write up what it shows
Real recall is 48% after the gate, and one must-be-silent query gets an
answer anyway. Review finding 2 (the score distributions overlap, so no
gate separates a real recall from a false one) and finding 4 (the memStore
branch at voice.go:776 is unreachable for notes). Adds an embedder cache
so the gate sweep does not re-embed the fixture nine times.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:34:28 +04:00