Compare commits

...

64 Commits

Author SHA1 Message Date
kami ddb658ffbb Add a Kiwix search client and score retrieval (Vikunja #403)
Step one of letting Maven read instead of recall. No LLM yet.

internal/kiwix/client.go: search a local Kiwix server, parse the RSS
reply, hand back title + path + plain-text snippet + word count. The
snippet is the unit of context; a full article is ~100KB of HTML and
will not fit a 4096 token window.

internal/kiwix/retrieval_eval.go plus knowledge_v1.json: the 9 knowledge
questions from the phrasing fixture, each with hand-written English
keywords, scored on whether a wanted article comes back in the top 5.
Opt-in via MAVEN_KIWIX_URL, since CI has no Kiwix. No pass bar, the
number is the finding.

Result on the live mirror: 8/8 answerable questions hit, 7 of them at
rank 1. Retrieval works. Keywords are written by hand on purpose, since
Kiwix ranks by keyword and not by meaning, so a natural question fails.
A query-rewrite step is the next piece of work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:43:44 +04:00
kami c9d88c152e Drop "never phones home" as a hard rule
The owner's call, 2026-07-31: a 0.8B model does not know enough about the
world to be useful without reading something. So she may now read external
sources to answer world questions.

What replaces the old rule, in all three docs:

- No telemetry, no cloud model, no third-party account. Unchanged.
- Local first: the Kiwix ZIMs on the box before anything on the network.
- External search is allowed but off unless configured, same as weather
  and telegram.
- His notes and facts are never search input. Only the utterance goes out
  — never the persona block, the history, or matched notes.

Docs only, no code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:34:53 +04:00
kami 1890ff5d5d Constrain the phrasing output with a GBNF grammar
The 0.8B answered about one chat turn in three with open reasoning as plain text, so no JSON ever closed and the fallback shipped "Thinking Process:" to the user. A grammar makes that output impossible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:18:06 +04:00
kami 0110e9bc8c Report every address break, and stop -те verbs blinding the check
From a real reply in a nudge eval run: "Смотрите на его потребление
воды" is a plural imperative AND third person about him. Only the plural
printed.

Two separate faults. The check returned on its first hit, so the second
break stayed invisible and the failure read as milder than it was; it now
joins them. And "его" was not detected at all — looksVerb knows the
-й/-йте imperative but not the -те plural, so "смотрите" counted as the
person being talked about, which is what an antecedent means here.
pluralVerb already knows that form, so the antecedent test uses it too.

Third time a verb form has blinded this check. A fourth means it wants a
morphology table rather than another suffix.
2026-07-31 16:52:16 +04:00
kami 50ca8c8b5a Score the chat, query and knowledge phrasing paths (#395)
The phrasing fixture was 15 nudge cases, so every prompt change we
measured only told us about nudges. But the shared context block sits in
front of five prompts, and three of them — chat, note query, general
knowledge — had no scorer at all. Those are the long free-form replies,
where a persona break is most likely and where nothing could see one.

27 cases, nine per path. Nine rather than five because the nudge fixture
already cannot resolve a change smaller than about three cases, and a
per-path score off five would be worse.

Reuses the persona checks instead of copying them. Length, mood and
"no questions" are left out on purpose: these paths return no mood, and
a follow-up question is a feature in chat, not a fault.

The run refuses to score unless the model answers before and after it.
PhraseChat and PhraseQuery swallow model errors and return a canned
string, so without that guard a dead server produces a full report with
zero errors and a bad score — which reads as bad phrasing rather than as
nothing measured. Vikunja #397 is the real fix.
2026-07-31 16:51:52 +04:00
kami de09471421 Merge the shared prompt context block 2026-07-31 16:07:35 +04:00
kami d65c16a567 Don't tell her she can't talk
The block listed what she can do and ended with "nothing else". It sits
in front of the chat and general-knowledge prompts too, so that told her
to refuse the exact thing those prompts are for. Talking is now first in
the list, and the closing line limits ACTIONS rather than everything.

Also dropped the self-introduction from the knowledge prompt. It said
"Мавена, персональный ассистент" — a different name and a masculine
noun, right after the block says she is Maven and feminine. Identity
lives in the block now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 16:07:35 +04:00
kami 062d4252ef Tell her what she can actually do
The context block now lists her real capabilities: reminders, notes and
facts (write and recall), and the calendar — all three are code paths in
mavend today. Weather, telegram and shell acts are listed only when the
config actually has them, because offering something she cannot do is
worse than staying quiet about it.

Also drops the pronouns from the optional name/city line. The block's
own "ты" is Maven, so "тебя зовут" read as her name and "его" would have
shown her the third-person form she must never use about him. They are
plain labels now.

Vikunja #394.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 15:57:58 +04:00
kami 2c27e2ce1f Give every prompt one shared context block
The "address him as ты" rule had only reached two of the five system
prompts. Instead of pasting it into the other three (five copies drift —
that is how this happened), there is now one block, in internal/persona,
prepended to all five: nudges, action replies, chat, note queries and
general knowledge.

The block says who he is and how to address him (a man, always "ты",
never "вы", never "он" about him; Maven stays feminine), plus the
current local date and time. It is rendered fresh each turn because the
time changes, and it is correct with an empty config — the address and
gender rules are defaults in code. Config only adds optional facts:
owner_name, city, and the existing free-text `persona` string, which is
now the static half of the block.

Russian even in front of the English prompts: the rules are Russian
grammar, so they read best stated in Russian, and there is one copy.

Vikunja #394.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 15:55:30 +04:00
kami ccc5cba2a3 Merge the example-led prompt finding 2026-07-31 15:41:01 +04:00
kami 89d83c0b11 Record the example-led nudge prompt experiment (#393) — it made things worse
Tried rewriting the nudge prompt to lead with five on-topic examples instead
of rules. Three eval runs each side: before 12/13/14 of 15, after 11/12/11.
The loss is all in the address check — formal "вы" and plural imperatives came
back once the "говоришь на ты" rule stopped being its own sentence, and the
on-topic examples leaked their wording into the wrong cases.

Prompt reverted. Only the finding is committed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 15:40:18 +04:00
kami a97f554802 Merge the informal address prompt rule 2026-07-31 14:54:20 +04:00
kami f4de2fc5e1 Don't let a verb count as the person being talked about
The third-person check asks whether anyone else was named before "он".
A nudge is mostly verbs, and they were counted as possible people, so
"попробуй встать и отдохнуть — у него есть перерыв" passed. Infinitives
and imperatives now join past tense as words that cannot be a person.

A plain noun before the pronoun still blinds it. That needs a parser,
and the comment says so.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:54:20 +04:00
kami eef5d4da4f Tell the phraser to speak to him informally, singular
The prompts stated the feminine self-reference rule but never said whom she is
speaking to, so the model produced formal plural ("Жду вас") and talked about
him in third person ("Он не ел 11 дней"). Adds the address rule right next to
the feminine one, in the nudge prompt and the confirmation prompt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:52:23 +04:00
kami 09f1696fce Merge the eval label and kill script fixes 2026-07-31 14:32:45 +04:00
kami 80f7322294 Don't fail when docker confirms nothing is running
"Nothing on the host" meant two different things and the script treated
them the same. If docker answers and names no running containers, Maven
really is down and the script should say so and exit 0. Only when docker
cannot be asked is the answer unknown, and that is the case that must
fail loudly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:32:45 +04:00
kami fa5aebfbe4 Merge the delivery boundary fixes 2026-07-31 14:30:54 +04:00
kami 59cec63da1 List the columns in the table rebuild
The migration copied rows with SELECT *, which matches columns by
position. It is correct today, but if the old table's order ever
differed it would shuffle every row instead of failing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:30:54 +04:00
kami 02e8786695 Stop the containers instead of claiming success (#380)
docker-compose.yml has no 'pid: host', so each container has its own PID
namespace and pkill on the host matches nothing inside them. The script
then printed "All services gracefully stopped" while mavend, its
llama-server and the rest were still running.

Now it checks for running compose containers first and stops them with
docker compose. If it cannot ask docker and finds nothing to kill on the
host, or anything survives the kill, it says so and exits non-zero
instead of claiming success. The bare-metal path is unchanged apart from
verifying the SIGKILL actually worked, and no longer risks killing the
shell it was launched from.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:30:51 +04:00
kami 0272dc9d89 Record a suppressed care nudge instead of dropping it silently (#370)
Dropping a sev1-2 care nudge while you're away is right and still happens.
But it was a bare `continue`: no row, no log, so "she dropped it", "the gate
suppressed it" and "the rule never fired" all looked identical afterwards.

Adds a 'dropped' delivery status (migration #12 widens the CHECK constraint;
sqlite can't do that in place, so the table is rebuilt) and records the drop
as one delivery_attempts row plus a log line.

No nudges row for a drop: that table feeds the ignored_rate signal, and a
nudge nobody could see must not count as ignored.

TestVoiceNoSessionFallthroughLeavesOutboxTrail expected exactly one row for
sev1-2 when voice had no session. It now expects the voice failure plus the
drop, which is the point of the change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:27:47 +04:00
kami 2ad7635501 Merge the address-form eval check 2026-07-31 14:27:16 +04:00
kami 9949b309b1 Don't let a time word blind the third-person check
The check asks whether anyone else was named before "он". Time words
were not stoplisted, so "сегодня он не ел" read "сегодня" as the person
being talked about and passed — which is the recorded break with a word
in front of it, and nudges open with those words constantly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:27:08 +04:00
kami a788ca3915 Label eval runs with the model the server actually loaded (#379)
The phrasing eval printed "llm (0.8B, ...)" no matter which gguf
llama-server had loaded, so two runs of two different models came out
named the same and were easy to mix up when comparing.

It now asks llama-server over /v1/models, same as the router eval
already did. The helper moved to internal/llm so both share it, and it
now errors instead of returning a blank name when the id field is
missing — an unreachable server gets labelled "unknown-model", never a
plausible-looking guess.

Both eval paths stay opt-in behind MAVEN_LLM_URL; no server needed for
go test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:25:28 +04:00
kami 62d47d28ac Add an eval check for formal and third-person address (#384)
The phrasing run produced two persona breaks that scored clean:
"Приходите… Жду вас" (formal plural) and "Он не ел 11 дней" (talks
about him instead of to him). She is feminine, he is male, and she
speaks to him informally, one to one.

The new `address` check flags the "вы" family, plural imperative
endings, and a third-person "он" with no other subject named earlier in
the message. Like `hisgender` it is a keyword/suffix heuristic, not a
parser, and it prints the word it tripped on so a false alarm is easy to
dismiss. Limits are written out in the comment.

Both recorded strings are pinned as unit tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:25:23 +04:00
kami e9ff2c4912 Never send the full nudge body off-box (#368)
The away sinks fell back to the whole Body when Summary was empty. ntfy and
telegram leave the box, and the 0.8B phraser drops fields regularly, so that
fallback could push full detail off the machine.

The dispatcher already strips detail from away sendables. This exports that
one rule as delivery.AwayMessage and has both sinks use it, so a sink can't
leak the body on its own either: empty Summary means a generic line plus the
rule name, never the body.

The two sink tests named TestSendFallsBackToBodyWhenSummaryEmpty asserted the
old, wrong behaviour, so they are rewritten to assert the generic line.
TestSendRejectsEmptyMessage is likewise replaced: an away message can no
longer be empty, so the sink has nothing left to reject.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:23:54 +04:00
kami 3dbf67f8f9 Drop the city time-zone table
The user only ever asks the time in his own zone, so answering other
cities was code kept in step with the weather city list for no gain.
Any named place now gets the honest "local time only" answer that was
already there for unknown cities.

Removes the 22-entry table, the lookup and the embedded tz database.
Closes Vikunja #389 — there is only one city list again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:14:18 +04:00
kami 84ba217892 Say so when the day asked about is out of reach 2026-07-31 14:05:21 +04:00
kami f179ae2fde Merge the system reply fixes 2026-07-31 14:03:42 +04:00
kami d00929ac0b Answer the day the user asked about and the city he named (#388)
replySystem had two arms that PR 30 made reachable, and both answered confidently wrong: the date arm keyword-matched "числ" and always answered today, so "какое число завтра" answered today; the clock arm ignored a named city and answered local time. The date arm now reads the day word through router.ParseCalendarDate (which grew послезавтра/вчера and now cuts the day boundary in the local zone instead of UTC). The clock arm answers the named zone when it resolves offline from the tz database embedded in the binary, and otherwise says plainly that she only knows local time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:02:57 +04:00
kami b6f47fbeb6 Merge the clock and calendar routing rule 2026-07-31 13:53:35 +04:00
kami 2e9b9ec1cf Warn separately when a row has no text to re-embed 2026-07-31 13:52:56 +04:00
kami bfb57c3148 Give the router prompt a rule for clock and calendar questions (#374)
The prompt named seven intents but never said which one a clock or date
question belongs to, so the model guessed: system->query x4 in every eval
run. The rule now says the clock and the calendar date themselves are
system, what is written in the calendar or in memory stays query, and a
time named inside a request is just part of the request.

That split follows what the daemon can answer. Only replySystem owns the
clock and the date formatter, while the agenda is answered from
CalendarEvents inside the query branch.

Also adds one calendar-agenda fixture case so an over-broad system rule
cannot pass unnoticed, and writes up the before/after numbers. The
targeted confusion is gone; the headline accuracy did not move.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:52:00 +04:00
kami d1f6f6355f Merge the vector backfill 2026-07-31 13:51:25 +04:00
kami 92ecb691de Re-embed stored notes and facts after an embedder swap (#378)
The embedder swap left every stored vector in the old model's space, so cosine against a new query vector is noise. Add the one-shot backfill: store.ReembedAll re-embeds every note and fact text with the currently configured embedder (the passage side, which is the side stored text was written with) and rewrites both places a vector lives — the notes table embedding column and the memory_vectors rows.

All of it plus the embedder marker happens in one transaction, so a failure partway changes nothing and writes no marker: re-run it. A run against a DB whose marker already names the current embedder does nothing.

Triggered explicitly with `mavend -reembed`, not automatically on mismatch: ONNX on the laptop CPU makes this minutes of work, and a silent multi-minute stall on boot would look like a hang. The mismatch warning now tells the user to run it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:50:43 +04:00
kami 4282f6b9a9 Warn about vectors written before the marker existed 2026-07-31 13:44:29 +04:00
kami 7bb9f9be06 Merge the embedder marker 2026-07-31 13:42:40 +04:00
kami 1e47eaca5a Record which embedder wrote the stored vectors and warn on a swap (#378)
The embedder moved from paraphrase-multilingual-MiniLM-L12-v2 to
multilingual-e5-small. Both are 384-dimensional, so nothing in the code
noticed: cosine between an old stored vector and a new query vector is
noise, and recall degrades silently.

So the DB now records the embedder that wrote its vectors. One value for
the whole DB (migration #11, a small `meta` key/value table) rather than a
column on every vector row: the backfill re-embeds every note and fact in
one pass, so a per-row marker would hold the same string in every row and
cost a column on two tables for nothing.

The identity comes from the embedder itself via a new optional ID() method
("multilingual-e5-small@384", model file name plus dimension), so pointing
the config at another model changes the string without anyone editing a
constant. mavend logs a loud WARNING at startup naming both the stored and
the configured embedder when they differ.

Detection only — recall behaviour is unchanged. TODO(#378) in
store.CheckEmbedder marks where the backfill will hook in.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:42:07 +04:00
kami 892330eb84 Merge the note recall fix 2026-07-31 13:33:14 +04:00
kami 9a3bcd7c46 Merge the thinking-off measurement 2026-07-31 13:31:31 +04:00
kami 98ee701e03 Let a note win a recall, not only a fact (#373)
The memory pass ran only after the notes-only gate had already rejected
the same note at the same score. Notes and facts share one vector index,
so a note that failed there failed again — the branch could only ever
return a fact.

Now the memory pass runs first: one search over everything Maven
remembers, one gate, and the memory that clearly matches best answers
(a note gets phrased, a fact is read back). The notes-only pass stays
behind it for notes the vector index does not hold. No threshold moved,
so the set of questions answered is unchanged — only which memory
answers them.

Fixture gained two mixed note+fact cases, so the answerable count goes
25 -> 27: hash recall@1 36.0% -> 37.0% (ratchet 0.32 unchanged, comment
updated), e5 recall@1 72.0% -> 70.4%, false recall still 1/5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:30:38 +04:00
kami 04c1088088 Measure thinking off on routing properly — it does not win (#376)
The 67.1% "thinking off" column in ROUTING-EVAL-31-07-2026.md was an
artefact. It came from a hand-rolled HTTP client in the eval test that
did not send repeat_penalty, so it differed from the reference run on two
axes and the penalty was the one that mattered.

Re-scored back to back on an idle box with everything else held equal:
thinking off is identical to thinking on, case for case, same confusion
matrix, same three unparseable replies. A direct probe of the running
llama-server shows enable_thinking, thinking and reasoning_budget are all
ignored for this model on this build, so there was nothing to turn off.

No defaults changed. The misleading third configuration is removed from
internal/router/eval so its table cannot be quoted again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:30:22 +04:00
kami 07c191d8b8 Merge dialogue session persistence 2026-07-31 13:21:03 +04:00
kami c668310b3e Persist the dialogue session so a restart keeps the conversation
Vikunja #363. The follow-up session was a plain in-memory map, so any
mavend restart dropped the thread. It now mirrors to a small TTL-pruned
sqlite table and is loaded on startup; expired sessions are deleted on
load, not revived. Clarify's pending question is untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:20:30 +04:00
kami 1bd2acdc2a Do not exempt Russian words that are both noun and verb 2026-07-31 12:55:45 +04:00
kami 15e5dd8eaa Merge the second-person gender check 2026-07-31 12:54:48 +04:00
kami 10cf6f525c Check that nudges do not address the owner in the feminine
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:54:18 +04:00
kami e2210f6844 Merge the clarify-expiry notice 2026-07-31 12:52:36 +04:00
kami 214a4032cf Tell him when an expired clarify question is dropped
Vikunja #382. A parked clarifying question past its TTL was discarded
silently on read; now she says the old request is gone and the newly
spoken words are still routed as a fresh utterance.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:51:47 +04:00
kami dc70a5a7ab Show clarify_max_attempts in the deployed config
The default is 3 either way. Writing it out means you can see the knob
without reading the Go.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:34:28 +04:00
kami 06aded6ab0 Merge commit 'd2be98e' into overnight-jul31
# Conflicts:
#	cmd/mavend/clarify.go
#	cmd/mavend/clarify_test.go
#	cmd/mavend/voice.go
#	internal/config/config.go
2026-07-31 12:34:00 +04:00
kami d2be98ee2a Say out loud when she gives up instead of dropping the request
An unclear answer used to end the request on the spot. Now she re-asks the same
question while attempts remain, and when they run out she says
"Прости, я не поняла. Скажи, пожалуйста, по-другому." — silence would leave him
thinking it was handled. Same reply when the missing slot has no question to
ask, and as a floor in finishClarified so an empty reply can never ship.

Tests: three questions allowed, the fourth gives up out loud, the cap is
configurable, and a restated time is the one that lands.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:31:55 +04:00
kami 62d320f93a Let her ask three times, and let a restated answer win
MaxAttempts was 1, justified as "not a nag". Wrong reading: "not a nag" is about
interrupting unprompted, and a clarifying question is part of a conversation he
started. Now three, configurable via voice.clarify_max_attempts (default 3).
Three, because after that the likely problem is she misheard the whole request,
not one slot.

Answer used to keep the parked value, so "в три" then "нет, в пять" threw the
five away. Now a value the answer carries wins for the slot she asked about.
Only for the clarify answer — a correction in a fresh turn is followUpMerge.

The eight-field chained assertion in the Answer test is one DeepEqual now, so a
new field in Slots is covered without touching the test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:31:46 +04:00
kami 796e6af3cf Merge commit '74a7088' into overnight-jul31 2026-07-31 12:24:20 +04:00
kami 74a70880a8 Write up the phrasing eval: 0/15 to 13/15, and what the number hides
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:23:54 +04:00
kami a40bc559d5 Fix the nudge phrasing prompt: stop teaching the model to echo the example
The system prompt showed the JSON contract as {"response": "..."} and the
user prompt repeated it. A 0.8B copies whatever sits in the response slot, so
7 of 15 nudges came back as literally "...".

Changes, all prompt-side — the {"response","mood"} contract is unchanged:
- nudge system prompt is Russian, feminine self-reference, with filled-in
  examples on topics that never appear as rules, so copying them is visible
- rule names get a Russian gloss and a required keyword, named last in the
  prompt where a small model weights it hardest
- durations render in Russian, not English
- the no-parse fallback says something Russian instead of "water — care",
  which was going straight to a Russian piper voice
- same "..." placeholder removed from replier_llm.go

Scored on internal/phraser/eval: 0/15 -> 13/15.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:23:03 +04:00
kami 1db0fcfcd0 Merge commit '4ca68d2' into overnight-jul31 2026-07-31 12:07:38 +04:00
kami 4ca68d2f3f Bake off LFM2.5 against Qwen3.5-0.8B on the RU routing fixture
Vikunja #278 / #250. Keep Qwen: LFM2.5-1.2B loses 8 points of intent
accuracy, all of it Russian, and runs 2.4x slower.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:07:06 +04:00
kami 5a1d465db5 Merge commit '4383844' into overnight-jul31
# Conflicts:
#	deploy/mavend.json
#	internal/config/config.go
2026-07-31 11:52:21 +04:00
kami 43838445ab Write up the margin gate results
Third section: why the absolute gate could not separate the two
distributions, the delta sweep, and the before/after. Marks next-steps
item 3 done.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:51:21 +04:00
kami 11831c6ace Gate recall on the margin over the runner-up, not just the score
The e5 embedder puts every cosine in one narrow band (0.79-0.89), so the
absolute query_min_score gate cannot tell a real hit from a made-up
question: any value under the band answers everything, any value above it
answers nothing. False recall was 5/5.

New gate asks whether one note is clearly the best instead: top1 - top2 >
delta. New query_min_margin config knob, default 0.008, read off the sweep
in the recall harness. The absolute floor stays as a second check.

On the recall fixture with e5: answered 72% -> 68%, false recall 5/5 -> 1/5.

Vikunja #359

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:50:39 +04:00
kami 93c1a41d4a Route with the resident model by default
The two things that made this unsafe are fixed: the router can now
refuse, and slot extraction runs on its decisions.

On the held-out fixture it gets 63.2% of intents right against the
classifier's 50.0%, with no route errors. It costs about a second a
turn instead of 30ms.

The flag is a pointer now, so leaving it out of the config means on
and only writing false turns it off.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:47:57 +04:00
kami 0b3b8d0a9e Test the clarify round-trip end to end at the daemon level
Covers: a reminder with no time is asked about and completes on the answer; the
same for a fact; an answer past the TTL falls through as a fresh utterance; a
second unclear answer drops the request with no second question; a clarified act
off the allowlist neither runs nor gets enabled; a clarified destructive act
still parks a confirm; noise keeps the canned reply. No model, no network.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:31:07 +04:00
kami fe0e654ab1 Ask the question, then act on the answer
On a clarify decision with one identifiable gap she now asks instead of saying
"не поняла", and parks the request. The next utterance is parsed as the answer
with the router's own extractor and the completed decision runs through
applyAction like any other — so a clarified act still needs the allowlist and
still hits the destructive confirm gate. An answer that does not fill the gap
drops the request; she never asks twice. Also pulls the session-store block
that HandlePushToTalk and handleText both had into rememberTurn, since the
clarify path needed a third copy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:31:07 +04:00
kami a54ebac0cb Work out which slot is missing and phrase one short question
A table per intent (reminder needs a time, fact needs a key, act needs a fn)
plus one fixed Russian question per slot. Templates, not model output: a 0.8B
would wander and a question that rewords itself is harder to answer. Note,
query, chat and system get no question — for those a clarify decision keeps
the canned reply rather than inventing a question for noise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:31:07 +04:00
76 changed files with 5103 additions and 481 deletions
+15 -1
View File
@@ -82,9 +82,23 @@ workspace enforces that the Go and relabelling prompts remain identical.
## Non-goals (hard constraints)
Never phones home. Not a nag, not autonomous. Maven's persona is **feminine** — Russian
Not a nag, not autonomous. Maven's persona is **feminine** — Russian
self-reference must use feminine forms (the user is male; see memory `maven-persona-gender`).
**"Never phones home" is DEPRECATED** (owner's call, 2026-07-31). It used to be a hard
constraint and it is not one any more: a 0.8B — and a 1.7B — does not know enough to answer
world questions, so she needs to read external sources. What replaces it:
- **No telemetry, no cloud model, no third-party account.** That part never changes. Nothing
about Maven is reported to anyone, and inference stays on the box.
- **Local sources first.** Kiwix ZIMs on homesrv (Wikipedia, ifixit) before anything on the
network. Reading beats recalling for a small model, and a local read costs nothing.
- **External search is allowed and off unless configured**, like the weather and telegram
capabilities.
- **His notes and facts are never search input.** Looking up why the sky is blue and sending
his stored personal notes to an upstream engine are different acts. Only the utterance goes
out, never the persona block, history, or matched notes.
## Web UI conventions
Server-rendered pages share `cmd/mavweb/static/ui.css` (served at `/ui.css`) and the `nav`
+10 -4
View File
@@ -15,7 +15,8 @@
**Maven** — self-hosted personal assistant. Manages your day, acts on your
homelab. One daemon on homesrv (always-on, not the workstation), multiple
client surfaces. All local, never phones home.
client surfaces. Inference and data stay on the box; she may READ external
sources (see Non-goals — "never phones home" is deprecated).
Primary name is "Maven", with feminine-gendered Russian self-reference
("она", "меня", "помогла"). Clients may choose their own UI label. Consistent
@@ -35,8 +36,13 @@ Inside boundary — the ones that actually constrain the build:
she records. A confident wrong fact is worse than a known gap.
- **Not a nag** — she'd rather miss a nudge than be mutable. Shuts up when
uncertain. Load-bearing.
- **Not a stranger** — runs on your stuff, your model, your data. Never
phones home.
- **Not a stranger** — runs on your stuff, your model, your data. No
telemetry, no cloud model, no third-party account. She may READ external
sources to answer world questions (Kiwix first, then optional search); she
never reports anything about you to anyone, and your notes and facts are
never used as search input. **"Never phones home" as an absolute is
deprecated** — owner's call, 2026-07-31: a small model does not know enough
to be useful without reading.
- **Not a relationship** — mom-tone is a function that makes nudges land, not
emotional company. Names the drift a warm small model falls into.
@@ -458,7 +464,7 @@ decides *insistence*. Both are needed.
sev ≤ 2 drops on away, sev ≥ 3 holds: a missed water nudge is noise, a missed
backup failure isn't. Away-channels (ntfy/telegram) leave the box — the one
path that crosses "never phones home," through your own relay. **Minimal
path that leaves the box for a person to see, through your own relay. **Minimal
body** — "disk low on homesrv," not detail; don't make notifications a
shoulder-surf exfil surface.
+101
View File
@@ -0,0 +1,101 @@
# Resident model bake-off — 31-07-2026
**Recommendation: keep Qwen3.5-0.8B.** LFM2.5-1.2B is worse at routing (52.6% vs 60.5%
intent accuracy), and the loss is almost entirely Russian (18/61 vs 22/61 RU, while EN is a
wash). It is also 2.4× slower. The Thinking variant is far worse again.
Settles Vikunja **#278 / #250**.
- Same fixture and scorer as `ROUTING-EVAL-31-07-2026.md`: `internal/router/eval/`
(`ru_routing_v1.json`, 76 held-out cases).
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:<port> make eval-router`
(`TestLLMRouterBaseline`). Note: there is no `make eval-models` target.
- All three models served by the same `llama-server` flags — `-c 2048 -ngl 99 -t 6`, only
`-m` and `--port` differ. One server at a time on an otherwise idle box, so latencies are
real and not contention.
- Measured on top of the router prompt fix (`origin/overnight/router-prompt` merged in), so
the Qwen column is directly comparable to the numbers already recorded.
## Results
`llm-only` — the model alone. This is the column that measures the model.
| | Qwen3.5-0.8B | LFM2.5-1.2B Instruct | LFM2.5-1.2B Thinking |
|---|---|---|---|
| **intent-only accuracy** | **60.5%** | 52.6% | 36.8% |
| full accuracy (intent+slots+gate) | **36.8%** | 32.9% | 21.1% |
| **RU** | **22/61** | 18/61 | 10/61 |
| EN | 6/15 | **7/15** | 6/15 |
| route errors | 0 | 0 | 0 |
| **p50 / p95 latency** | **1.05s / 1.71s** | 2.47s / 3.62s | 2.42s / 3.24s |
| missed clarify | 6 / 6 | 6 / 6 | 6 / 6 |
`cascade+llm` — stage-0 → model → classifier floor, what #320 would actually ship. Same
ordering.
| | Qwen3.5-0.8B | LFM2.5-1.2B Instruct | LFM2.5-1.2B Thinking |
|---|---|---|---|
| intent-only accuracy | **61.8%** | 55.3% | 38.2% |
| full accuracy | **46.1%** | 42.1% | 30.3% |
| RU / EN | **27/61** / 8/15 | 23/61 / **9/15** | 15/61 / 8/15 |
| route errors | 0 | 0 | 0 |
| p50 / p95 latency | **1.28s / 1.94s** | 2.18s / 2.72s | 2.27s / 3.19s |
Full logs: the three runs are archived in the session scratchpad
(`qwen08.txt`, `lfm-instruct.txt`, `lfm-thinking.txt`).
## Russian-specific failures — the owner's worry is confirmed
LFM2.5's Russian loss is not spread out. It has one large, specific failure: **it hears
almost any Russian imperative or short phrase as `reminder`.**
- `перезапусти докер` → reminder (want act)
- `включи вытяжку` → reminder (want act)
- `закрой жалюзи` → reminder (want act)
- `заметка: продлить домен в августе` → reminder (want note)
- `запиши что кран на кухне снова капает` → reminder (want note)
- `доброе утро` → reminder (want chat)
- `спасибо тебе` → reminder (want note/chat)
- `переходи в тихий режим` → reminder (want system)
That is `note→reminder ×4`, `act→reminder ×4`, `chat→reminder ×2` in one run. Qwen's
equivalent failure axis is `query→fact ×8`, which is a narrower and already-understood bug.
Two more Russian-side problems worth naming:
1. **Fact keys come back empty or wrong in Russian.** `воды попил наконец`, `поужинал`,
`поспал часов пять` and `отметь что я позавтракал овсянкой` all returned an empty key.
`сходил в душ` and `отдохнул минут двадцать` both returned `water`. Qwen does not do this.
2. **It leaked German.** `slept about seven hours` produced the fact key
`"7 Stunden geschlafen"`. Grammar-valid, semantically garbage — a sign the multilingual
mix is not anchored where Maven needs it.
The claimed tool-calling advantage did not show up here. `act` is the closest thing this
fixture has to a tool call, and LFM2.5 got it wrong more often than Qwen, mostly by calling
it a reminder. It also produced no `fn` slot on any act, same as Qwen.
## The Thinking variant
Not viable. 36.8% intent accuracy, 10/61 Russian, and no latency saving over Instruct — the
thinking trace costs time without buying accuracy on a short enum classification. With the
`enable_thinking=false` diagnostic it collapsed further to 28.9% with 2 route errors
(`query→reminder ×12`). Do not pursue.
## Notes
- Nothing crashed, nothing ignored the GBNF grammar, and no model produced unparseable JSON
in the shippable configurations. Zero route errors for both Instruct and Thinking in
`llm-only` and `cascade+llm`. The problem with LFM2.5 is what it decides, not whether it
can emit the contract.
- The `6 / 6` missed clarify is unchanged across all three models. No model fixes the missing
refusal lane — that is `Confidence: 1.0` hardcoded in `llmrouter.go` (Vikunja #359), not a
model property.
- The report labels every configuration `(0.8B)`; that string is hardcoded in the test, not a
reflection of which gguf was loaded. Model identity was confirmed per run via `/v1/models`.
- No Go code was changed for this measurement, and no bug was found that needed one.
## What this does not settle
Routing only. LFM2.5 might still phrase better, and phrasing is the resident model's other
job — that needs its own fixture. But routing is the load-bearing path and Maven is
Russian-first, so on the evidence here the switch is not worth making.
+6 -3
View File
@@ -103,14 +103,17 @@ eval-router:
eval-recall:
MAVEN_ONNX_LIB="$(MAVEN_ONNX_LIB)" $(GO) test -v -count=1 ./internal/memory/recalleval/
# eval-phrasing -- score nudge phrasing (internal/phraser/eval). Verbose so the
# eval-phrasing -- score nudge phrasing AND the conversational paths (chat,
# query, general knowledge) in internal/phraser/eval. Verbose so the
# report and every generated message land in the terminal. With no environment
# it scores the deterministic Stub only, which is what CI runs. Set
# MAVEN_LLM_URL to add the resident model:
# MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing
# The model run is slow (minutes) -- the timeout is raised to match.
# The model run is slow (minutes) -- the timeout is raised to match. It covers
# two fixtures now (15 nudges + 27 conversational cases, and the chat replies are
# the long ones), hence 90m rather than 40m.
eval-phrasing:
$(GO) test -v -count=1 -timeout 40m ./internal/phraser/eval/
$(GO) test -v -count=1 -timeout 90m ./internal/phraser/eval/
# eval-models — score ONE llama-server against the same fixture, for the
# resident-model bake-off (#278, #250). Start a server with the gguf you want,
+169
View File
@@ -0,0 +1,169 @@
# Phrasing evaluation — 31-07-2026
How Maven words a nudge, measured instead of argued. Counterpart to
`ROUTING-EVAL-31-07-2026.md`.
- Fixture + scorer: `internal/phraser/eval/` (`nudges_v1.json`, 15 cases; `eval.go`, `checks.go`)
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing`
- Model: Qwen3.5-0.8B Q4_K_M, the resident model. Not swapped.
- Commit: `a40bc55` (prompt fix)
Every check is a string or length test a human can read and disagree with. No model
grades another model here.
## Result
| | before | after |
|---|---|---|
| **cases passing every check** | **0/15** | **13/15** |
| mood in enum | 6/15 | 15/15 |
| Russian | 2/15 | 14/15 |
| length (≤120 chars, ≤16 words) | 13/15 | 15/15 |
| feminine self-reference | 15/15 | 15/15 |
| no cringe | 13/15 | 15/15 |
| on topic | 6/15 | 13/15 |
| p50 latency | 11.4s | 11.4s |
Latency did not move and is not good. 11s to word one nudge on this box.
## The bug reproduced
Yes, exactly as reported. 7 of 15 messages were the literal string `"..."`, and one was
`"full voice message"`. Both are text copied straight out of the prompt.
The system prompt said:
```
Respond ONLY with valid JSON: {"response": "full voice message", "mood": "neutral"}
```
and the user prompt said:
```
Respond as JSON: {"response": "...", "mood": "..."}
```
A 0.8B does not read `"..."` as "put your answer here". It reads it as the answer. The
prompt was a worked example whose worked part was blank, so the model filled the slot by
copying. This is the whole of finding 1.
## What else was wrong
Four separate faults, all prompt-side:
1. **Placeholder echo** (7 cases) — above.
2. **Wrong language** (13/15 failed the language check). The prompt was entirely English
and said "in the user's language (Russian or English)". The model picked English. It is
never English: the nudge is spoken by a Russian piper voice.
3. **Rule names are English identifiers.** `netdata_critical`, `service_down`, `break` went
into the prompt raw. The model cannot nudge about a topic it has not been told in words,
so 9/15 were off topic. The daemon knows what its own rules mean; now it says so.
4. **Mood invented** (`"warm"`, twice). The enum was listed in a parenthesis at the end of
an English sentence. Now it is its own line: "ровно одно из: neutral, happy, thinking,
tired, confused."
Plus two non-prompt faults the run exposed:
- **The no-parse fallback was English.** When the model returned nothing usable, the body
became `fmt.Sprintf("%s — %s", rule, sev)``"water — care"` — and that string went to
a Russian TTS. Now it falls back to plain Russian.
- **Durations were English.** `humanDur` returns "3 hours"; it was landing verbatim inside
Russian sentences. Nudges now use a Russian formatter.
## Three iterations, and what each taught
| | score | change |
|---|---|---|
| baseline | 0/15 | — |
| iter 1 | 2/15 | Russian prompt, filled-in examples, Russian durations |
| iter 2 | 11/15 | required keyword per rule, one example instead of five, Russian fallback |
| iter 3 | **13/15** | examples moved to topics that are not rules |
The interesting step is 1 → 2. Fixing the placeholder did not fix the disease, it moved it:
the model stopped copying `"..."` and started copying my first example instead. Five nudges
in a row came back as `"Ты не пил воду три часа. Налей стакан."` regardless of the rule.
**A small model copies the nearest concrete text in its prompt.** That is one failure mode
with two symptoms. The fix that stuck was making the examples about laundry and a laptop
battery — topics no rule ever produces, so copying them is visible in the score rather than
invisibly passing the water cases.
## Do not oversell 13/15
Seven of the thirteen passes are the **deterministic fallback**, not the model:
`"Напоминаю: таблетки."`, `"Сервис не отвечает."`, `"Критический алярм: проверь диск."`,
`"Ты давно не пил воду."`. Those are strings this commit added to Go. The model returned
nothing parseable and the fallback scored.
So the honest reading is roughly **6/15 from the model, 7/15 from a fallback, 2/15 failing**.
The prompt fix is real — `"..."` is nearly gone and the language and mood checks are clean —
but a large part of the jump is that failure now degrades into Russian instead of into
`"water — care"`. That is a genuine improvement for the operator and a weak one for the model.
The two remaining failures: one `"..."` recurrence (`routine-stretch`) and one meal nudge
that never says food.
## Tried and reverted: an example-led nudge prompt (#393)
The idea was that a 0.8B copies examples better than it follows rules, so the nudge prompt
was rewritten to lead with five on-topic examples (water, break, pills, morning, service) and
the prose rules were compressed to pay for the tokens: 1190 chars down to 986.
It measured **worse**, three runs each side, same llama-server, same fixture:
| run | before | after |
|---|---|---|
| 1 | 12/15 (address 14) | 11/15 (address 13) |
| 2 | 13/15 (address 15) | 12/15 (address 15) |
| 3 | 14/15 (address 15) | 11/15 (address 12) |
`feminine` and `hisgender` were 15/15 on all six runs, so they measure nothing here. The
regression is all in `address`: 44/45 before, 40/45 after. Formal "вы"/"ваше" and plural
imperatives came back, and so did `"..."`.
Two likely causes, both about the same thing — **examples do not carry a prohibition**. The
old prompt spent a whole sentence on «говоришь на "ты", в единственном числе»; the new one
demoted that to one item in a long "никогда" list, and the model stopped obeying it. And
making the examples on-topic let their *wording* leak: a break case came back as
«Вы давно не пили воду. Выпей стакан.» — the water example, verbatim, in the wrong slot.
That is exactly the failure the laundry/laptop examples were chosen to avoid.
Change reverted. What survives is the measurement: a rule the model must obey needs its own
sentence, and examples must stay off-topic. Also note the before side alone spans 1214 of
15 — this fixture cannot resolve anything smaller than about three cases.
## Broken, found, not fixed
1. ~~**`checkFeminine` only catches half the constraint.**~~ **Fixed** (#381). It scanned for
masculine self-reference only, so three messages that addressed the *owner* in the feminine
("ты давно не отдыхал**а**") scored clean. There is now a second check, `hisgender`: a
feminine past-tense verb (-ла/-лась) in a sentence addressed to him ("ты", "тебе", "твой")
fails, unless the verb is hers ("я заметила", "напомнила тебе"). It is a suffix rule, not a
parser — see the comment in `checks.go` for what it misses. A fresh 15-case run after adding
it scored **12/15** with `hisgender` 15/15; the model did not repeat the feminine address in
that sample, and the check is pinned by unit tests on the recorded bad strings instead.
2. **Grammar is not checked at all, and it is bad.** `"Он не ел 11 дней"` (it was 11 hours),
`"Сонуждились 7 дней"` (not a word), `"Они забыли воду"` (wrong person entirely). Every
one of these passes all six checks. The fixture measures properties, not fluency, and at
0.8B fluency is the binding constraint.
3. **Unit confusion.** The model turns hours into days about a third of the time. The
prompt now says "11 ч"; it reads it as days.
4. **11s p50.** Unchanged and untouched here. A nudge the model takes eleven seconds to
word has missed its moment. Worth its own task.
5. **The keyword hint is close to teaching to the test.** `ruleKeywords` names the word the
on-topic check looks for. It is defensible — the daemon genuinely knows its rule topics
and the model genuinely cannot infer them from `netdata_critical` — but the on-topic
number is softer than the others because of it.
## Next steps
1. ~~**Add a second-person gender check**~~ — done, `hisgender` in `checks.go` (#381).
2. **Decide whether the fallback should count as a pass.** Right now `Score` cannot tell a
model answer from a fallback. Either mark fallback bodies in `PhrasedNudge` or count them
in their own column. Without that, any future prompt change can score well by failing
more.
3. **Attack the 11s.** Nudge phrasing is short and non-interactive; thinking off is the first
thing to try, as it was for routing (#376).
4. **Re-measure when #122 lands.** The CPT'd Qwen3-1.7B is the target. 13/15 with seven
fallbacks is the floor it has to beat, and the fluency problems above are the ones a
bigger, Russian-trained checkpoint should actually fix.
+4 -1
View File
@@ -90,4 +90,7 @@ later* is the worker + RAG.
4. **Deferred work** — larger reasoner, custom Piper voice and other expansions.
## Non-goals (unchanged)
Never phones home. Not a nag. Not autonomous. Feminine-gendered RU self-ref.
Not a nag. Not autonomous. Feminine-gendered RU self-ref. No telemetry, no
cloud model, no third-party account — but she MAY read external sources to
answer world questions (Kiwix first, search optional). "Never phones home" as
an absolute is deprecated, owner's call 2026-07-31; see CLAUDE.md § Non-goals.
+93 -2
View File
@@ -75,6 +75,14 @@ same vector. A note is indexed in both places with the same embedding, so if it
`QueryNotes` it fails again here — the branch can only ever return a **fact**. Its comment calls it
"additive"; for notes it is not.
**Fixed (Vikunja #373).** The memory pass now runs *first*, as one search over notes and facts with
one gate, so whichever memory is clearly the best match answers — note or fact. The notes-only pass
stays behind it for notes the vector index does not hold. No threshold changed, so the set of
questions Maven answers is the same; only which memory answers them. The fixture gained two mixed
note+fact cases (`ru-mixed-031`, `ru-mixed-032`), which is why the counts below are out of 27
answerable cases and not 25: hash recall@1 36.0% (9/25) → 37.0% (10/27), e5 recall@1 72.0% (18/25) →
70.4% (19/27) with answered-after-gate 68.0% → 66.7% and false recall unchanged at 1/5.
### 5. Ranking has no recency or type signal, and the store is not the bottleneck
`internal/store/notes.go:67` sorts by cosine and uses `ts` only to break an exact float tie, which
@@ -143,14 +151,97 @@ the vector memory table was written by the old model, so after this deploy they
against a new query. A live database needs every note and fact re-embedded before recall works at
all. Filed as its own task.
## Margin gate — 31-07-2026, third run
Next-steps item 3, done. The absolute gate is replaced by a **margin gate**: answer only when the
top hit beats the runner-up by more than delta (`top1 top2 > δ`). Same fixture, same e5 embedder,
same store as the run above. `internal/memory/gate.go` holds the check; both read paths call it
(`cmd/mavend/recall.go` and the notes-RAG branch in `voice.go`). New knob `voice.query_min_margin`
in `deploy/mavend.json`, default 0.008.
### Why the absolute gate could not work, in one line of data
The harness now prints the margin distributions, and they barely overlap where the raw scores
overlap completely:
| | top-1 score | margin (top1 top2) |
|---|---|---|
| right note first (n=18) | min 0.810, median 0.862, max 0.890 | min 0.001, median 0.029, max 0.053 |
| must stay silent (n=5) | min 0.795, median 0.815, max 0.835 | min 0.000, median 0.002, **max 0.019** |
Four of the five must-be-silent cases have a margin at or under 0.002 — when there is nothing to
recall, e5 finds several notes equally close and no clear winner. That is the signal the absolute
score throws away.
### The delta sweep
Absolute gate held at 0.55 throughout.
```
delta 0.000: answered 18/25 (72%) false recall 5/5
delta 0.002: answered 17/25 (68%) false recall 3/5
delta 0.005: answered 17/25 (68%) false recall 2/5
delta 0.008: answered 17/25 (68%) false recall 1/5 <- chosen
delta 0.010: answered 15/25 (60%) false recall 1/5
delta 0.012: answered 14/25 (56%) false recall 1/5
delta 0.015: answered 12/25 (48%) false recall 1/5
delta 0.020: answered 11/25 (44%) false recall 0/5
delta 0.025: answered 9/25 (36%) false recall 0/5
delta 0.030: answered 8/25 (32%) false recall 0/5
delta 0.040: answered 4/25 (16%) false recall 0/5
delta 0.050: answered 2/25 ( 8%) false recall 0/5
delta 0.060: answered 0/25 ( 0%) false recall 0/5
```
### Chosen: δ = 0.008
It is the best point on the frontier, not a taste call. **0.008 dominates 0.010, 0.012 and 0.015
outright** — same 1/5 false recall, 8 to 20 points more real recall. Everything below it buys recall
back only by admitting more false recalls (0.005 → 2/5, 0.002 → 3/5). The next real improvement is
0.020 at 0/5 false, and it costs 24 points of recall to get there.
The brief's bar was "recall above 60% with false recall at 1/5 or better". 0.008 clears it with room:
68% and 1/5.
### Before / after
| | absolute gate 0.55 (previous) | margin gate δ=0.008 |
|---|---|---|
| recall@1 (ranking, ungated) | 72.0% (18/25) | 72.0% (18/25) — unchanged, the gate does not rank |
| **answered after the gate** | 72.0% (18/25) | **68.0% (17/25)** |
| **false recall** | **5/5 (100%)** | **1/5 (20%)** |
| fixture cases passed | 18/30 | **21/30** |
Four false recalls removed for one real answer. That is the trade the spec asks for — she is not a
guesser-of-truth. The one survivor is `en-pref-025` ("should i be offered wine"), which recalls a
filler note at 0.796 with a 0.019 margin: the widest silent-case margin in the fixture, and it sits
inside the real-recall range, so no delta removes it without taking real answers with it.
### Does the absolute cutoff still earn its keep? Marginally — kept
On this fixture with e5 it is a **no-op**: the lowest right-note score is 0.791, so 0.55 rejects
nothing the margin does not already reject. It is kept for two reasons, neither glamorous. It still
does real work for the hash embedder (its own sweep shows answers dropping from 16% to 0% between
0.30 and 0.50), and it is the only thing standing between the user and a reply built from a store
where everything is far away but one row happens to be a little less far — a near-empty database, or
the stale-vector case below. Cheap insurance, no measured cost. If a later embedder makes it bite,
the sweep is one command.
### Caveat on the numbers
Five must-be-silent cases is a thin basis for a 4-point decision. 1/5 and 2/5 differ by one case.
The shape of the frontier is trustworthy — margins separate, absolute scores do not — but δ=0.008
itself should be re-read off a bigger fixture (next-steps item 6) before anyone defends the third
decimal.
## Next steps — ordered by value-to-risk; nothing here is a decision
1. **Swap the embedder to `multilingual-e5-small` with `query:`/`passage:` prefixes.** One config
change plus a prefix in `onnxembedder.go`, re-measurable in one command.
2. **Re-run `make eval-recall`, then set the gate from the sweep** — not before. Any
`query_min_score` picked against today's embedder describes a model on its way out.
3. **Replace the absolute-score gate with a margin gate** (`top1 top2 > δ`) — as the routing eval
concluded, absolute cosine cannot see a flat distribution.
3. ~~**Replace the absolute-score gate with a margin gate**~~ — done, see the section above.
δ=0.008, false recall 5/5 → 1/5.
4. **Delete or repair the dead `memStore` branch** at `voice.go:776` — search before the gate,
gate it separately, or restrict it to facts and say so.
5. **Add a mild time decay to ranking** — the newest statement of a preference is the true one.
+108 -10
View File
@@ -66,15 +66,113 @@ Three things this run settles:
`запиши что…` phrasings toward fact, and that suspicion stands — all five `ru-note-*`
cases now land on fact. Tracked as Vikunja #375.
**Thinking off is the best configuration measured so far**, on both accuracy and latency
(Vikunja #376). That is worth understanding before flipping: routing is a short
classification into a fixed enum with grammar-constrained output, so there is little to
reason about, and the thinking trace mostly gives a small model room to talk itself out of
the right answer. Phrasing is a different job and needs measuring separately.
The `thinking off` column above read as the best configuration measured so far (Vikunja #376).
**It was wrong** — see the controlled re-run below. Ignore that column.
Still `6 / 6` missed clarify — the router has no way to say "I don't know" (Vikunja #359).
That is unchanged by anything here.
## Thinking off — 31-07-2026, controlled re-run (Vikunja #376)
The "thinking off wins by 6 points" observation above **does not hold**. It was a measurement
artefact, and the earlier table's `thinking off` column should be ignored.
The thinking-off variant was scored by a hand-rolled HTTP client living in the test file
instead of `llm.Client`. That copy did not send `repeat_penalty`, which the real router does
send (`routeRepeatPenalty = 1.15`). So the two columns differed on two axes at once, and the
one that mattered was the penalty, not the thinking mode.
Re-measured with everything else held equal — same fixture, same prompt, same grammar, same
sampling, same idle box, the three configurations run back to back and never concurrently:
| | llm-only, thinking on | llm-only, thinking off | cascade+llm |
|---|---|---|---|
| intent-only accuracy | 59.2% (45/76) | 59.2% (45/76) | 61.8% (47/76) |
| full accuracy (intent+slots+gate) | 38.2% (29/76) | 38.2% (29/76) | 57.9% (44/76) |
| route errors | 3 | 3 | 0 |
| grammar violations | 3 (all 3 route errors) | 3 (same 3 cases) | 0 |
| missed clarify | 5 / 6 | 5 / 6 | 5 / 6 |
| p50 latency | 836ms | 920ms | 810ms |
| p95 latency | 1.41s | 2.00s | 1.31s |
Thinking off is not just a tie on the headline numbers — it is identical case for case, with
the same confusion matrix and the same three unparseable replies. The latency difference is
run-to-run noise on one box, and it points the wrong way here.
The reason is simpler than any accuracy argument: **this llama-server build ignores the
request-level thinking switch for this model.** Probed directly against the running server
with `chat_template_kwargs.enable_thinking = false`, `chat_template_kwargs.thinking = false`
and top-level `reasoning_budget = 0` — all three return a byte-identical answer with the
thinking trace still in `reasoning_content`, and the server reports the prompt prefix as
cached, meaning the rendered template did not change. There was never anything being turned
off, which is also why the numbers match exactly.
Nothing was defaulted. `internal/llm` still has no `chat_template_kwargs` field, `VoiceConfig`
has no thinking flag, and `deploy/mavend.json` is unchanged. The misleading third
configuration is removed from `internal/router/eval` so the table it produced cannot be quoted
again.
Two caveats worth saying out loud:
- **The fixture is 76 cases.** A 6-point difference on 76 cases is roughly 4-5 cases and would
not have been worth trusting even if it had reproduced. This one was exactly 0 cases, which
is a much easier call.
- **This is one server build and one checkpoint** (`b9351`, Qwen3.5-0.8B Q4_K_M). If the
#122 checkpoint or a newer llama.cpp does honour the switch, the question reopens — but it
reopens as an unmeasured question, not as a 6-point win.
Phrasing was **not** measured. Whether thinking helps there is still open, and now also blocked
on the same "can we even turn it off" question.
## Clock and calendar rule — 31-07-2026 (Vikunja #374)
`routeSystem` never said whether "который час" or "какое число завтра" are `system` or
`query`, and `system→query ×4` showed up in every run. The rule added says: the clock and the
calendar date themselves are `system`; what is *written in* the calendar or in memory
("что у меня завтра", "какие есть напоминания") stays `query`; and a time named inside a
request ("напомни завтра…") is just a detail of the request, not a reason for `system`.
That split is not a preference. In `cmd/mavend/voice.go` only `replySystem` owns the clock and
the date formatter, so a clock question routed to `query` falls into the embedder + note RAG
and answers "не знаю". The agenda, on the other hand, is answered by `ParseCalendarDate` +
`CalendarEvents` *inside* the `query` branch, so that side has to stay `query`. The rule sits
above the question test because every one of these utterances carries a question word and a
later rule would never be reached.
The fixture is now 77 cases: one calendar-agenda case was added
(`ru-query-019` "что у меня стоит в календаре на послезавтра", intent `query`) specifically so
an over-broad system rule cannot pass unnoticed. The clock/date cases (`ru-sys-001/002/005`,
`en-sys-001`) already existed.
Three runs, same box, back to back, never concurrently:
| | baseline | first rule (too broad) | rule as committed |
|---|---|---|---|
| llm-only intent-only | 59.2% (45/76) | 54.5% (42/77) | 59.7% (46/77) |
| llm-only full | 38.2% | 35.1% | 39.0% |
| llm-only route errors | 3 | 4 | 5 |
| llm-only p50 | 1.09s | 0.91s | 0.93s |
| cascade+llm intent-only | 61.8% (47/76) | 58.4% | 62.3% (48/77) |
| cascade+llm full | 57.9% | 54.5% | 59.7% |
| cascade+llm route errors | 0 | 0 | 0 |
| cascade+llm p50 | 0.91s | 0.80s | 1.04s |
**The targeted bug is fixed and the headline number did not move.** `system→query ×4` is gone
in both LLM configurations — the `time` and `date` tags go from 0/2 and 0/2 to 2/2 and 2/2 —
but the model then over-applies the rule, and `query→system ×5` plus `reminder→system ×2`
appear where they did not exist before. Net accuracy is a wash, inside the noise of a 77-case
fixture.
The first attempt is shown because it is the honest history: it said "спрашивает время, дату
или день недели → system" with no scope, which swept up reminders, and it cost 3-5 points. It
was tightened once, on the reasoning that a rule capturing "напомни завтра в 7" is simply
wrong, and not tuned further. The remaining `query/reminder → system` over-trigger is a new,
separate weakness of the sub-1B model and deserves its own task rather than more prompt
kneading against a held-out fixture.
The rule is kept. It is correct about what the daemon can answer, and the failure it replaces
was silent ("не знаю" to "который час") while the one it introduces is loud.
## Findings
### 1. The resident model does route better — 50.0% vs 36.8%
@@ -145,11 +243,11 @@ Note the grammar's `string ::= "\"" ([^"\\] | "\\" .)* "\""` is unbounded, so no
### 7. Two hypotheses tested and closed
- **Thinking mode is a non-issue.** Qwen3.5's template defaults `thinking = 1`, so
grammar-constrained JSON lands in `reasoning_content` with `content` empty —
`llm.Client`'s fallback handles it. A `thinking off` run scored *identically* (18/76,
48.7%, same p50). `internal/llm` deliberately does **not** grow a `chat_template_kwargs`
field.
- **Thinking mode is a non-issue.** Confirmed twice now, the second time properly — see the
controlled re-run section. Grammar-constrained JSON lands in `reasoning_content` with
`content` empty and `llm.Client`'s fallback handles it; the request-level switch does
nothing on this build. `internal/llm` deliberately does **not** grow a
`chat_template_kwargs` field.
- **Runaway array repetition does not reproduce.** An isolated smoke test with a stripped
grammar emitted `{"intent":"reminder"}` until `MaxTokens`; under the real `routeSystem`
prompt the few-shot examples anchor it to one object. 2 errors in 76, not 76.
+78 -17
View File
@@ -40,9 +40,44 @@ var clarifyQuestions = map[dialogue.Slot]string{
dialogue.SlotFn: "Что сделать?",
}
// clarifyDropped — she asked once, the answer still did not fill the gap, so
// the request is gone. Said plainly, once, with no second question.
const clarifyDropped = "Не разобрала — скажи целиком, пожалуйста."
// clarifyGaveUp — she is out of questions and still does not have the slot. She
// says so out loud: dropping the request in silence would leave him thinking it
// landed. Feminine self-reference ("поняла"), as everywhere.
const clarifyGaveUp = "Прости, я не поняла. Скажи, пожалуйста, по-другому."
// clarifyExpired — his answer came after the TTL, so the parked request is
// already gone. Same tone as clarifyGaveUp, different reason: too much time
// passed, not "I did not understand". Feminine self-reference ("ждала",
// "отпустила"); he is addressed with a plain imperative.
const clarifyExpired = "Прости, я слишком долго ждала ответа и отпустила прошлую просьбу. Если она ещё нужна, скажи заново."
// clarifyExpiredNotice returns that line when a parked question had just timed
// out, and "" when nothing was parked. Call it right after
// resolveClarifyAnswer: a live question is answered there, an expired one is
// only reported here — the words themselves still go on to be routed fresh.
func (h *reactiveHandler) clarifyExpiredNotice() string {
if h.clarifyStore == nil {
return ""
}
if !h.clarifyStore.TakeExpired(voiceDialogueID, h.now()) {
return ""
}
log.Printf("voice: clarify — parked question expired, telling him and routing the words fresh")
return clarifyExpired
}
// withNotice glues the expiry notice in front of this turn's reply. One turn
// carries one reply on the wire, so the notice cannot be a message of its own —
// but neither the notice nor the fresh answer may be dropped.
func withNotice(notice, reply string) string {
if notice == "" {
return reply
}
if reply == "" {
return notice
}
return notice + " " + reply
}
// missingFor returns the slots a decision still needs, most important first.
// Empty ⇒ there is nothing identifiable to ask about.
@@ -79,13 +114,14 @@ func (h *reactiveHandler) askClarify(dec router.Decision) (string, bool) {
return "", false
}
h.clarifyStore.Put(voiceDialogueID, &dialogue.PendingQuestion{
Intent: dialogue.Intent(dec.Intent),
Slots: toDialogueSlots(dec.Slots),
Missing: []dialogue.Slot{slot},
Utterance: dec.Utterance,
Asked: h.now(),
TTL: clarifyTTL,
Attempts: 1, // asked once; MaxAttempts is 1, so there is no second ask
Intent: dialogue.Intent(dec.Intent),
Slots: toDialogueSlots(dec.Slots),
Missing: []dialogue.Slot{slot},
Utterance: dec.Utterance,
Asked: h.now(),
TTL: clarifyTTL,
Attempts: 1, // this ask
MaxAttempts: h.clarifyMaxAttempts,
})
log.Printf("voice: clarify — asked about %s for intent=%s", slot, dec.Intent)
return question, true
@@ -97,8 +133,9 @@ func (h *reactiveHandler) askClarify(dec router.Decision) (string, bool) {
// resolveConfirm and checked in the same place.
//
// The answer is parsed with the same extractor the router uses, for the intent
// she parked — no second parser. If it still does not fill the gap the request
// is dropped: she does not ask again.
// she parked — no second parser. If it still does not fill the gap she asks
// again, up to MaxAttempts; after that she says out loud that she did not
// understand. She never drops the request in silence.
func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string) (string, bool) {
if h.clarifyStore == nil {
return "", false
@@ -107,17 +144,14 @@ func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string)
if q == nil {
return "", false
}
// One shot either way: the question is consumed whether or not the answer
// works, so a failed answer can't leave the question armed.
h.clarifyStore.Delete(voiceDialogueID)
intent := router.Intent(q.Intent)
answer := h.extractor.Extract(ctx, intent, text, h.now())
merged := q.Answer(text, toDialogueSlots(answer))
if len(dialogue.StillMissing(q.Missing, merged)) > 0 {
log.Printf("voice: clarify — answer %q did not fill %v, dropping", text, q.Missing)
return clarifyDropped, true
return h.reaskOrGiveUp(q, merged, text), true
}
h.clarifyStore.Delete(voiceDialogueID)
// Rebuild the decision as if it had routed cleanly, then run it down the
// normal path. Clarify is deliberately false and the intent is unchanged:
@@ -133,6 +167,29 @@ func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string)
return h.finishClarified(ctx, dec), true
}
// reaskOrGiveUp handles an answer that left the gap open: ask the same question
// again while she has attempts left, otherwise say she did not understand and
// let the request go. Never returns "" — a mute give-up reads as "done".
func (h *reactiveHandler) reaskOrGiveUp(q *dialogue.PendingQuestion, merged dialogue.Slots, text string) string {
question := ""
if len(q.Missing) > 0 {
question = clarifyQuestions[q.Missing[0]]
}
if question == "" || !q.CanAsk() {
h.clarifyStore.Delete(voiceDialogueID)
log.Printf("voice: clarify — gave up on %v after %d question(s), answer was %q", q.Missing, q.Attempts, text)
return clarifyGaveUp
}
// Re-park with whatever the answer DID give, the clock restarted and one
// more question spent.
q.Slots = merged
q.Attempts++
q.Asked = h.now()
h.clarifyStore.Put(voiceDialogueID, q)
log.Printf("voice: clarify — answer %q did not fill %v, asking again (attempt %d)", text, q.Missing, q.Attempts)
return question
}
// finishClarified runs a completed decision through the same steps a freshly
// routed one takes: remember the turn, act, then phrase.
func (h *reactiveHandler) finishClarified(ctx context.Context, dec router.Decision) string {
@@ -146,6 +203,10 @@ func (h *reactiveHandler) finishClarified(ctx context.Context, dec router.Decisi
if reply == "" {
reply = h.replier.Reply(dec)
}
if reply == "" {
// Belt: an empty reply here would be a silent drop.
reply = clarifyGaveUp
}
return reply
}
+102 -11
View File
@@ -86,7 +86,7 @@ func TestClarifyReminderCompletesOnAnswer(t *testing.T) {
if !handled {
t.Fatal("the answer to an open question must be consumed as an answer")
}
if reply == clarifyDropped {
if reply == clarifyGaveUp {
t.Fatalf("a good answer must not drop the request: %q", reply)
}
@@ -111,7 +111,7 @@ func TestClarifyFactCompletesOnAnswer(t *testing.T) {
if _, asked := h.askClarify(clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши")); !asked {
t.Fatal("a fact with no key should be asked about")
}
if reply, handled := h.resolveClarifyAnswer(ctx, "пил воду"); !handled || reply == clarifyDropped {
if reply, handled := h.resolveClarifyAnswer(ctx, "пил воду"); !handled || reply == clarifyGaveUp {
t.Fatalf("answer should complete the fact, handled=%v reply=%q", handled, reply)
}
if fact, err := st.LatestFact(ctx, "water"); err != nil || fact.Key != "water" {
@@ -137,26 +137,87 @@ func TestClarifyAnswerAfterTTLIsANewRequest(t *testing.T) {
}
}
// TestClarifyUnclearAnswerDropsWithoutAskingAgain — MaxAttempts is 1.
func TestClarifyUnclearAnswerDropsWithoutAskingAgain(t *testing.T) {
// TestClarifyAsksThreeTimesThenSaysSo — three questions are allowed, the fourth
// is not, and running out is SPOKEN. Silence would read as "handled".
func TestClarifyAsksThreeTimesThenSaysSo(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a question")
t.Fatal("expected a first question")
}
// Two more unclear answers ⇒ two more questions (3 asks in total).
for i := 2; i <= 3; i++ {
reply, handled := h.resolveClarifyAnswer(ctx, "ну не знаю")
if !handled {
t.Fatalf("answer %d must be consumed as an answer", i)
}
if reply != "На когда напомнить?" {
t.Fatalf("attempt %d should ask again, got %q", i, reply)
}
if h.clarifyStore.Get(voiceDialogueID, h.now()) == nil {
t.Fatalf("attempt %d must leave the question armed", i)
}
}
reply, handled := h.resolveClarifyAnswer(ctx, "ну не знаю")
if !handled || reply != clarifyDropped {
t.Fatalf("an unclear answer should drop the request, handled=%v reply=%q", handled, reply)
if !handled || reply != clarifyGaveUp {
t.Fatalf("the fourth try must give up out loud, handled=%v reply=%q", handled, reply)
}
if strings.Contains(reply, "?") {
t.Fatalf("she must not ask a second question: %q", reply)
if reply == "" || strings.Contains(reply, "?") {
t.Fatalf("giving up must be spoken and must not be another question: %q", reply)
}
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
t.Fatal("a dropped request must leave no armed question")
t.Fatal("a given-up request must leave no armed question")
}
if reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour)); err != nil || len(reminders) != 0 {
t.Fatalf("a dropped request must not create anything: reminders=%v err=%v", reminders, err)
t.Fatalf("a given-up request must not create anything: reminders=%v err=%v", reminders, err)
}
}
// TestClarifyMaxAttemptsIsConfigurable — one question when the config says one.
func TestClarifyMaxAttemptsIsConfigurable(t *testing.T) {
ctx := context.Background()
h, _, _ := newClarifyHandler(t)
h.clarifyMaxAttempts = 1
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a question")
}
if reply, handled := h.resolveClarifyAnswer(ctx, "ну не знаю"); !handled || reply != clarifyGaveUp {
t.Fatalf("with max 1 she must give up at once, handled=%v reply=%q", handled, reply)
}
}
// TestClarifyRestatedAnswerWins — «в 11:00», then «нет, в 15:00». The second
// value is the one that lands.
func TestClarifyRestatedAnswerWins(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме")); !asked {
t.Fatal("expected a question")
}
// First answer parses, but re-park it by hand as if she had asked again:
// what matters here is that Answer prefers the newer value over the parked
// one, which is the case the daemon hits on a re-ask.
q := h.clarifyStore.Get(voiceDialogueID, h.now())
if q == nil {
t.Fatal("expected an armed question")
}
first := h.extractor.Extract(ctx, router.IntentReminder, "в 11:00", h.now())
q.Slots = q.Answer("в 11:00", toDialogueSlots(first))
if reply, handled := h.resolveClarifyAnswer(ctx, "нет, в 15:00"); !handled || reply == clarifyGaveUp {
t.Fatalf("the restated answer should complete the request, handled=%v reply=%q", handled, reply)
}
reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour))
if err != nil || len(reminders) != 1 {
t.Fatalf("expected one reminder: %v err=%v", reminders, err)
}
want := h.extractor.Extract(ctx, router.IntentReminder, "в 15:00", h.now())
if !reminders[0].FireTs.Equal(want.Time) {
t.Fatalf("reminder at %v, want the restated %v", reminders[0].FireTs, want.Time)
}
}
@@ -228,6 +289,36 @@ func TestNoQuestionWhenNothingIsMissing(t *testing.T) {
}
}
// TestClarifyExpiryIsAnnouncedAndWordsStillRoute — his answer lands after the
// TTL: she must say the old request is gone AND still answer the new words.
func TestClarifyExpiryIsAnnouncedAndWordsStillRoute(t *testing.T) {
ctx := context.Background()
h, _, now := newClarifyHandler(t)
emb := router.NewHashEmbedder(1024)
h.embedder = emb
h.router = buildRouter(emb, h.matcher, 0.55, nil)
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a question")
}
*now = now.Add(clarifyTTL + time.Second)
reply := h.handleText(ctx, "как дела")
if !strings.HasPrefix(reply, clarifyExpired) {
t.Fatalf("expired question must be announced first, got %q", reply)
}
if strings.TrimSpace(strings.TrimPrefix(reply, clarifyExpired)) == "" {
t.Fatalf("the new words must still be answered, got only the notice: %q", reply)
}
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
t.Fatal("the expired question must be gone")
}
// The notice is said once, not on every later utterance.
if reply := h.handleText(ctx, "как дела"); strings.Contains(reply, clarifyExpired) {
t.Fatalf("notice repeated on a later turn: %q", reply)
}
}
// TestNoPendingQuestionFallsThrough — with nothing parked, an utterance routes
// normally.
func TestNoPendingQuestionFallsThrough(t *testing.T) {
+41 -21
View File
@@ -58,6 +58,7 @@ import (
"github.com/kami/maven/internal/delivery/telegramsink"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/webauthn"
@@ -189,7 +190,9 @@ func (l *lockedAPI) MorningStatus(ctx context.Context) ([]ipc.MorningRoutineStat
func run(args []string) error {
cfgPath := flag.String("config", defaultConfigPath(), "path to mavend JSON config")
wrappedKeyPath := flag.String("wrapped-key-file", "", "path to wrapped encryption key blob (enables cold-start unlock)")
reembed := flag.Bool("reembed", false, "re-embed every stored note and fact with the configured embedder, then serve normally (run once after an embedder swap)")
flag.CommandLine.Parse(args)
reembedOnStart = *reembed
cfg, err := config.Load(*cfgPath)
if err != nil {
return err
@@ -269,13 +272,13 @@ func run(args []string) error {
phr = phraser.NewStub()
if cfg.Phraser != nil {
pc := phraser.Config{
ModelPath: cfg.Phraser.ModelPath,
BinPath: cfg.Phraser.BinPath,
Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
Timeout: time.Duration(cfg.Phraser.Timeout),
Persona: personaFromCfg(cfg),
ModelPath: cfg.Phraser.ModelPath,
BinPath: cfg.Phraser.BinPath,
Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
Timeout: time.Duration(cfg.Phraser.Timeout),
ContextBlock: contextBlockFn(cfg, time.Now),
}
if pc.BinPath == "" {
pc.BinPath = "llama-server"
@@ -439,13 +442,13 @@ func run(args []string) error {
phr = phraser.NewStub()
if cfg.Phraser != nil {
pc := phraser.Config{
ModelPath: cfg.Phraser.ModelPath,
BinPath: cfg.Phraser.BinPath,
Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
Timeout: time.Duration(cfg.Phraser.Timeout),
Persona: personaFromCfg(cfg),
ModelPath: cfg.Phraser.ModelPath,
BinPath: cfg.Phraser.BinPath,
Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
Timeout: time.Duration(cfg.Phraser.Timeout),
ContextBlock: contextBlockFn(cfg, time.Now),
}
if pc.BinPath == "" {
pc.BinPath = "llama-server"
@@ -599,12 +602,29 @@ func run(args []string) error {
return nil
}
// personaFromCfg extracts the voice persona from the config, or returns ""
// when voice isn't configured. Used to pass a character prompt into the
// LLM phraser without requiring voice to be enabled.
func personaFromCfg(cfg *config.Config) string {
if cfg.Voice != nil {
return cfg.Voice.Persona
// personaFacts reads the optional, deployment-specific facts (his name, his
// city, the free-text persona string) out of the config. Everything here may
// be empty — the context block is correct without any of it.
func personaFacts(cfg *config.Config) persona.Facts {
f := persona.Facts{
// Telegram lives outside the voice block, so it counts either way.
Telegram: cfg.Telegram != nil && cfg.Telegram.BotToken != "" && cfg.Telegram.ChatID != "",
}
return ""
if cfg.Voice == nil {
return f
}
f.OwnerName = cfg.Voice.OwnerName
f.City = cfg.Voice.City
f.Static = cfg.Voice.Persona
// Same test wireVoice uses to pick the real provider over the stub.
f.Weather = cfg.Voice.Weather != nil && cfg.Voice.Weather.Provider == "open-meteo"
f.Tools = len(cfg.Voice.Tools) > 0
return f
}
// contextBlockFn returns the per-turn renderer of the shared context block.
// Per turn, not once at startup, because the block states the current time.
func contextBlockFn(cfg *config.Config, now func() time.Time) func() string {
f := personaFacts(cfg)
return func() string { return f.Block(now()) }
}
+158
View File
@@ -0,0 +1,158 @@
package main
import (
"context"
"fmt"
"math"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice"
)
// fixedEmbedder hands back a vector chosen per text, so a test can say exactly
// how close each stored memory is to the question. The real embedders make
// scores that are realistic but not controllable, and this test is about the
// gate, not about the embedder.
type fixedEmbedder struct{ vecs map[string][]float32 }
func (f *fixedEmbedder) Dim() int { return 4 }
func (f *fixedEmbedder) Close() error { return nil }
func (f *fixedEmbedder) Embed(_ context.Context, text string) ([]float32, error) {
v, ok := f.vecs[text]
if !ok {
return nil, fmt.Errorf("fixedEmbedder: no vector for %q", text)
}
return v, nil
}
// scoreVec builds a unit vector whose cosine against the query vector
// (1,0,0,0) is exactly score.
func scoreVec(score float64) []float32 {
rest := math.Sqrt(1 - score*score)
return []float32{float32(score), float32(rest), 0, 0}
}
// recordingPhraser remembers what the query path handed it to phrase, which is
// how the test can tell which pass produced the answer.
type recordingPhraser struct {
*phraser.Stub
notes []string
}
func (r *recordingPhraser) PhraseQuery(ctx context.Context, utterance string, notes []string) (string, error) {
r.notes = notes
return r.Stub.PhraseQuery(ctx, utterance, notes)
}
// recallCase — one stored memory: its text, how close it is to the question,
// whether it is a note or a fact, and whether the notes table holds it too.
type recallCase struct {
text string
score float64
kind string
}
// buildRecallHandler stores the given memories and returns a handler whose
// query path can be run directly. Notes go into BOTH the notes table and the
// vector index, which is what the daemon does (voice.go's IntentNote).
func buildRecallHandler(t *testing.T, question string, mems []recallCase) (*reactiveHandler, *recordingPhraser) {
t.Helper()
ctx := context.Background()
st := newTestStore(t)
emb := &fixedEmbedder{vecs: map[string][]float32{question: {1, 0, 0, 0}}}
mem := memory.NewInMemoryStore()
now := time.Now()
for i, m := range mems {
vec := scoreVec(m.score)
emb.vecs[m.text] = vec
id := fmt.Sprintf("%s:%d", m.kind, i)
if m.kind == "note" {
if _, err := st.WriteNote(ctx, now, m.text, vec, "tap:voice"); err != nil {
t.Fatalf("WriteNote: %v", err)
}
}
if err := mem.Insert(ctx, id, vec, map[string]string{"text": m.text, "type": m.kind}); err != nil {
t.Fatalf("memory insert: %v", err)
}
}
phr := &recordingPhraser{Stub: phraser.NewStub()}
h := &reactiveHandler{
api: ipc.NewStoreAPI(st),
embedder: emb,
replier: voice.NewStubReplier(),
phraser: phr,
now: func() time.Time { return now },
memStore: mem,
dataStore: st,
queryMinScore: 0.55,
queryMinMargin: 0.008,
weatherProvider: nil,
}
return h, phr
}
func askQuery(t *testing.T, h *reactiveHandler, question string) string {
t.Helper()
return h.applyAction(context.Background(), router.Decision{
Intent: router.IntentQuery,
Utterance: question,
})
}
// TestQueryRecallNoteCanWin — the note-recall regression (Vikunja #373). Notes
// and facts share one vector index, and a note that clearly beats everything
// else must be the answer. Before the fix the memory pass only ran after the
// notes-only gate had already rejected the same note at the same score, so only
// a fact could ever come back from it.
func TestQueryRecallNoteCanWin(t *testing.T) {
const q = "где молоко"
t.Run("a clearly best note answers", func(t *testing.T) {
h, phr := buildRecallHandler(t, q, []recallCase{
{text: "молоко стоит в холодильнике", score: 0.90, kind: "note"},
{text: "выучил пару аккордов", score: 0.50, kind: "note"},
})
reply := askQuery(t, h, q)
if want := "вот что я нашла: молоко стоит в холодильнике"; reply != want {
t.Errorf("reply %q, want %q", reply, want)
}
// One text, the winning memory's — the answer came from the memory
// pass, not from handing the phraser every note in the table.
if len(phr.notes) != 1 || phr.notes[0] != "молоко стоит в холодильнике" {
t.Errorf("phraser got %q, want just the recalled note", phr.notes)
}
})
// The other half of "one gate over everything": a fact that matches better
// than the best note now answers, instead of losing to a note that only had
// to beat other notes.
t.Run("the better-matching fact answers", func(t *testing.T) {
h, _ := buildRecallHandler(t, q, []recallCase{
{text: "молоко стоит в холодильнике", score: 0.80, kind: "note"},
{text: "купил молоко в среду", score: 0.95, kind: "fact"},
})
if reply := askQuery(t, h, q); reply != "купил молоко в среду" {
t.Errorf("reply %q, want the fact read back", reply)
}
})
// The gate is untouched: two memories this close mean the embedder cannot
// tell them apart, and silence still beats a coin flip.
t.Run("no clear best stays silent", func(t *testing.T) {
h, _ := buildRecallHandler(t, q, []recallCase{
{text: "молоко стоит в холодильнике", score: 0.860, kind: "note"},
{text: "молоко закончилось", score: 0.858, kind: "note"},
})
if reply := askQuery(t, h, q); reply != "не знаю." {
t.Errorf("reply %q, want silence", reply)
}
})
}
+18 -18
View File
@@ -2,24 +2,24 @@ package main
import "github.com/kami/maven/internal/memory"
// bestRecall is the read side of the long-term memory store: the top hit's
// stored text when it clears the confidence gate. This recalls across BOTH
// notes and facts (facts aren't in the notes table, so this is the only path
// that can answer "when did I last …?" from a captured fact). A note hit here
// is redundant with the notes-RAG path — by design; the two indexes can diverge
// once the backend is swapped for a persistent/external store. ok=false when
// there's no hit above the threshold or the hit carries no text.
func bestRecall(results []memory.Result, min float64) (string, bool) {
if len(results) == 0 {
return "", false
// bestRecall is the read side of the long-term memory store: the top hit when
// it clears the confidence gate. The index holds BOTH notes and facts, and
// either can win — the caller looks at the returned hit's meta["type"] to see
// which. Facts aren't in the notes table, so this is the only path that can
// answer "when did I last …?" from a captured fact.
//
// The whole hit is returned, not just its text, because "which memory answered"
// decides how the answer is said: a note gets phrased in Maven's voice, a fact
// is read back as stored.
//
// ok=false when the hit fails the confidence gate (see memory.Confident: an
// absolute floor plus a margin over the runner-up) or carries no text.
func bestRecall(results []memory.Result, minScore, minMargin float64) (memory.Result, bool) {
if !memory.Confident(results, minScore, minMargin) {
return memory.Result{}, false
}
top := results[0]
if top.Score < min {
return "", false
if results[0].Meta["text"] == "" {
return memory.Result{}, false
}
text := top.Meta["text"]
if text == "" {
return "", false
}
return text, true
return results[0], true
}
+38 -6
View File
@@ -8,23 +8,24 @@ import (
func TestBestRecall(t *testing.T) {
const min = 0.55
const margin = 0.008
t.Run("empty results", func(t *testing.T) {
if _, ok := bestRecall(nil, min); ok {
if _, ok := bestRecall(nil, min, margin); ok {
t.Error("empty results returned ok")
}
})
t.Run("top below threshold", func(t *testing.T) {
res := []memory.Result{{Score: 0.4, Meta: map[string]string{"text": "выпил воды"}}}
if _, ok := bestRecall(res, min); ok {
if _, ok := bestRecall(res, min, margin); ok {
t.Error("below-threshold hit returned ok")
}
})
t.Run("hit without text meta", func(t *testing.T) {
res := []memory.Result{{Score: 0.9, Meta: map[string]string{"type": "fact"}}}
if _, ok := bestRecall(res, min); ok {
if _, ok := bestRecall(res, min, margin); ok {
t.Error("textless hit returned ok")
}
})
@@ -34,12 +35,43 @@ func TestBestRecall(t *testing.T) {
{Score: 0.82, Meta: map[string]string{"text": "выпил воды в три часа", "type": "fact"}},
{Score: 0.60, Meta: map[string]string{"text": "другое"}},
}
got, ok := bestRecall(res, min)
got, ok := bestRecall(res, min, margin)
if !ok {
t.Fatal("clearing hit not returned")
}
if got != "выпил воды в три часа" {
t.Errorf("wrong text: %q", got)
if got.Meta["text"] != "выпил воды в три часа" {
t.Errorf("wrong text: %q", got.Meta["text"])
}
if got.Meta["type"] != "fact" {
t.Errorf("kind lost: %q", got.Meta["type"])
}
})
// The index holds notes and facts together, so a note has to be able to win
// it — for a long time it could not (Vikunja #373).
t.Run("a note can win", func(t *testing.T) {
res := []memory.Result{
{Score: 0.86, Meta: map[string]string{"text": "молоко в холодильнике", "type": "note"}},
{Score: 0.61, Meta: map[string]string{"text": "выпил воды", "type": "fact"}},
}
got, ok := bestRecall(res, min, margin)
if !ok {
t.Fatal("clearly-best note not returned")
}
if got.Meta["type"] != "note" || got.Meta["text"] != "молоко в холодильнике" {
t.Errorf("got %v, want the note", got.Meta)
}
})
// The runner-up is almost as close, so the embedder cannot tell the two
// notes apart. Silence beats reading back a coin flip.
t.Run("runner-up too close", func(t *testing.T) {
res := []memory.Result{
{Score: 0.860, Meta: map[string]string{"text": "выпил воды в три часа"}},
{Score: 0.858, Meta: map[string]string{"text": "другое"}},
}
if _, ok := bestRecall(res, min, margin); ok {
t.Error("thin-margin hit returned ok")
}
})
}
+11 -4
View File
@@ -7,6 +7,7 @@ import (
"time"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice"
)
@@ -23,13 +24,19 @@ type completer interface {
type llmReplier struct {
c completer
stub *voice.StubReplier
// block renders the shared context block per turn (who he is, the time).
// nil ⇒ the prompt stands alone.
block func() string
}
func newLLMReplier(c completer) *llmReplier {
return &llmReplier{c: c, stub: voice.NewStubReplier()}
func newLLMReplier(c completer, block func() string) *llmReplier {
return &llmReplier{c: c, stub: voice.NewStubReplier(), block: block}
}
const replySystem = `Ты — Maven, домашняя ассистентка (о себе — в женском роде). Подтверди действие РОВНО ОДНИМ коротким предложением (≤120 символов), тепло и по-русски. Не задавай вопросов, не повторяй слова, не добавляй ничего после точки. Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}.`
const replySystem = `Ты — Maven, домашняя ассистентка (о себе — в женском роде). Владелец — мужчина, говоришь с ним на "ты", в единственном числе; никогда не "вы"/"ваш" и не "он"/"его". Подтверди действие РОВНО ОДНИМ коротким предложением (≤120 символов), тепло и по-русски. Не задавай вопросов, не повторяй слова, не добавляй ничего после точки. Отвечай ТОЛЬКО одним объектом JSON с полями "response" (текст) и "mood" (ровно одно из: neutral, happy, thinking, tired, confused).
Пример: {"response": "Записала, что ты выпил стакан воды.", "mood": "neutral"}
Никогда не пиши "..." в поле response.`
func (r *llmReplier) Reply(d router.Decision) string {
if d.Clarify {
@@ -38,7 +45,7 @@ func (r *llmReplier) Reply(d router.Decision) string {
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
out, err := r.c.Complete(ctx, llm.Req{
System: replySystem,
System: persona.Prepend(r.block, replySystem),
User: replyContext(d),
MaxTokens: 512,
RepeatPenalty: 1.3,
+5 -5
View File
@@ -17,7 +17,7 @@ type mockCompleter struct {
func (m mockCompleter) Complete(_ context.Context, _ llm.Req) (string, error) { return m.out, m.err }
func TestLLMReplierReturnsLLMReply(t *testing.T) {
r := newLLMReplier(mockCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`})
r := newLLMReplier(mockCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`}, nil)
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
if got != "записала, кофе закончился" {
t.Errorf("got %q, want %q", got, "записала, кофе закончился")
@@ -25,7 +25,7 @@ func TestLLMReplierReturnsLLMReply(t *testing.T) {
}
func TestLLMReplierFallsBackToPlainText(t *testing.T) {
r := newLLMReplier(mockCompleter{out: "записала, кофе закончился"})
r := newLLMReplier(mockCompleter{out: "записала, кофе закончился"}, nil)
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
if got != "записала, кофе закончился" {
t.Errorf("got %q, want %q", got, "записала, кофе закончился")
@@ -33,7 +33,7 @@ func TestLLMReplierFallsBackToPlainText(t *testing.T) {
}
func TestLLMReplierFallsBackToStubOnError(t *testing.T) {
r := newLLMReplier(mockCompleter{err: errTestLLMDown})
r := newLLMReplier(mockCompleter{err: errTestLLMDown}, nil)
noteDec := router.Decision{Intent: router.IntentNote}
got := r.Reply(noteDec)
want := voice.NewStubReplier().Reply(noteDec)
@@ -43,7 +43,7 @@ func TestLLMReplierFallsBackToStubOnError(t *testing.T) {
}
func TestLLMReplierFallsBackToStubOnEmpty(t *testing.T) {
r := newLLMReplier(mockCompleter{out: ""})
r := newLLMReplier(mockCompleter{out: ""}, nil)
noteDec := router.Decision{Intent: router.IntentNote}
got := r.Reply(noteDec)
want := voice.NewStubReplier().Reply(noteDec)
@@ -53,7 +53,7 @@ func TestLLMReplierFallsBackToStubOnEmpty(t *testing.T) {
}
func TestLLMReplierClarifyUsesStub(t *testing.T) {
r := newLLMReplier(mockCompleter{out: "я всё поняла"})
r := newLLMReplier(mockCompleter{out: "я всё поняла"}, nil)
clarifyDec := router.Decision{Clarify: true}
got := r.Reply(clarifyDec)
want := voice.NewStubReplier().Reply(clarifyDec)
+78
View File
@@ -0,0 +1,78 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/router"
)
// systemHandler — a handler with nothing but a fixed clock, which is all
// replySystem needs.
func systemHandler(now time.Time) *reactiveHandler {
return &reactiveHandler{now: func() time.Time { return now }}
}
// TestReplySystemDateOffset — "какое число завтра" must answer tomorrow's
// date, not today's (Vikunja #388).
func TestReplySystemDateOffset(t *testing.T) {
// Thursday, 30 July 2026.
now := time.Date(2026, 7, 30, 14, 5, 0, 0, time.UTC)
h := systemHandler(now)
cases := []struct{ utterance, want string }{
{"какое сегодня число", "сегодня четверг, 30 июля 2026 года"},
{"какое число", "сегодня четверг, 30 июля 2026 года"},
{"какое число завтра", "завтра пятница, 31 июля 2026 года"},
{"какое число послезавтра", "послезавтра суббота, 1 августа 2026 года"},
{"какое было число вчера", "вчера среда, 29 июля 2026 года"},
}
for _, c := range cases {
got := h.replySystem(context.Background(), router.Decision{Utterance: c.utterance})
if got != c.want {
t.Errorf("replySystem(%q) = %q, want %q", c.utterance, got, c.want)
}
}
}
// A day she cannot work out must not come back as today's date — that is the
// same silent wrong answer #388 was about, one step further out.
func TestReplySystemUnknownDayIsHonest(t *testing.T) {
now := time.Date(2026, 7, 30, 14, 5, 0, 0, time.UTC)
h := systemHandler(now)
for _, u := range []string{
"какое число в пятницу",
"какое число через неделю",
"какое число в понедельник",
} {
got := h.replySystem(context.Background(), router.Decision{Utterance: u})
if got != onlyNearDaysReply {
t.Errorf("replySystem(%q) = %q, want the honest reply", u, got)
}
}
// The days she does know must not be caught by the same guard.
if got := h.replySystem(context.Background(), router.Decision{Utterance: "какое число завтра"}); got == onlyNearDaysReply {
t.Error("завтра was treated as an unknown day")
}
}
// TestReplySystemClockCity — the clock arm must not answer local time for a
// question about another city (Vikunja #388). She keeps one clock, so every
// named place gets the honest "local time only" answer.
func TestReplySystemClockCity(t *testing.T) {
now := time.Date(2026, 7, 30, 12, 0, 0, 0, time.UTC)
h := systemHandler(now)
cases := []struct{ utterance, want string }{
{"который час", "сейчас 12 часов ровно"},
{"который час в киеве", onlyLocalTimeReply},
{"сколько времени в москве", onlyLocalTimeReply},
{"который час в лондоне", onlyLocalTimeReply},
{"который час в бишкеке", onlyLocalTimeReply},
}
for _, c := range cases {
got := h.replySystem(context.Background(), router.Decision{Utterance: c.utterance})
if got != c.want {
t.Errorf("replySystem(%q) = %q, want %q", c.utterance, got, c.want)
}
}
}
+286 -51
View File
@@ -171,6 +171,7 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
emb = router.NewHashEmbedder(1024)
}
w.embedder = emb
checkStoredEmbedder(dataStore, emb)
// ----- tool executor (the enabled act allowlist, store-backed) -----
// Config tools are the declarative bootstrap: seed them into the store as
@@ -206,11 +207,11 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
if threshold <= 0 {
threshold = config.DefaultRouterThreshold
}
// Both routing paths are weak on held-out utterances — the classifier gets
// 36.8% of intents right, the resident model 50.0% and much slower. Off by
// default (see config.VoiceConfig.LLMRouter); the classifier always stays
// wired as the fallback, so a model error never breaks a turn.
rtr := buildRouter(emb, matcher, threshold, pickLLMRouter(cfg.Voice.LLMRouter, llmClient))
// The resident model routes by default: 63.2% of held-out intents right
// against the classifier's 50.0%, at about 1s a turn instead of 30ms (see
// config.VoiceConfig.LLMRouter). The classifier always stays wired as the
// fallback, so a model error never breaks a turn.
rtr := buildRouter(emb, matcher, threshold, pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient))
// ----- sessions registry (shared with voicesink) -----
sessions := voice.NewSessions()
@@ -227,14 +228,25 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
}
// ----- dialogue (multi-turn slot carry-over; 2-min follow-up window) -----
dialogueSessions := dialogue.NewSessionStore(2 * time.Minute)
// Store-backed when the daemon passes a store, so a restart mid-conversation
// keeps the thread (Vikunja #363). Sessions past their TTL are dropped on
// load, never revived. Clarify's parked question stays in memory only.
var dialogueSessions *dialogue.SessionStore
if dataStore != nil {
dialogueSessions = dialogue.NewPersistentSessionStore(2*time.Minute, dataStore)
if err := dialogueSessions.Load(context.Background(), time.Now()); err != nil {
log.Printf("dialogue: load saved sessions: %v", err)
}
} else {
dialogueSessions = dialogue.NewSessionStore(2 * time.Minute)
}
clarifyStore := dialogue.NewClarifyStore(clarifyTTL)
timeParser := router.NewPythonDateParser()
// ----- replier (LLM-backed when the engine is on, Stub floor otherwise) -----
replier := voice.Replier(voice.NewStubReplier())
if llmClient != nil {
replier = newLLMReplier(llmClient)
replier = newLLMReplier(llmClient, contextBlockFn(cfg, time.Now))
}
// ----- the handler (the reactive path; closes over stt / tts / router / coreAPI / memory) -----
@@ -255,10 +267,13 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
dataStore: dataStore,
dialogueSessions: dialogueSessions,
clarifyStore: clarifyStore,
extractor: router.Extractor{Time: timeParser, Acts: matcher, Facts: router.DefaultFactParser{}},
queryMinScore: cfg.Voice.QueryMinScore,
timeParser: timeParser,
ecosystem: eco,
// 0 here (unset config) ⇒ the dialogue default.
clarifyMaxAttempts: cfg.Voice.ClarifyMaxAttempts,
extractor: router.Extractor{Time: timeParser, Acts: matcher, Facts: router.DefaultFactParser{}},
queryMinScore: cfg.Voice.QueryMinScore,
queryMinMargin: cfg.Voice.QueryMinMargin,
timeParser: timeParser,
ecosystem: eco,
}
// ----- the server (TCP listener) -----
@@ -300,6 +315,9 @@ type reactiveHandler struct {
// load-bearing math (same posture as the presence thresholds). Set by
// wireVoice from VoiceConfig; default 0.55.
queryMinScore float64
// queryMinMargin — the second half of that gate: how far the top hit must
// beat the runner-up. 0 ⇒ margin off.
queryMinMargin float64
// timeParser — used as a fallback for stage-0 reminder grammar matches
// (where the extractor didn't run). Shared with the router's extractor.
@@ -314,6 +332,10 @@ type reactiveHandler struct {
// clarify.go). nil ⇒ she falls back to the canned "не поняла" reply.
clarifyStore *dialogue.ClarifyStore
// clarifyMaxAttempts — questions per request before she gives up out loud.
// 0 ⇒ dialogue.DefaultMaxAttempts (3). Set from VoiceConfig.
clarifyMaxAttempts int
// extractor parses the answer to an open question, with the same parsers
// the router's own stage-2 uses.
extractor router.Extractor
@@ -392,9 +414,17 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
return h.reply(ctx, reply, nil)
}
// 1b2. clarify answer — if she asked a question last turn, this utterance is
// its answer, not a fresh command. After the confirm check: a y/n gate is
// armed by her own prompt and is the narrower claim on the utterance.
// 1b2. expired clarify — a question was parked but its TTL ran out, so the
// request behind it is gone. Say that out loud (see clarify.go) and carry
// on: these words are still routed as a fresh utterance below, with the
// notice glued in front of whatever the fresh routing answers. Checked
// BEFORE the answer path: reading a parked question drops an expired one.
expiredNotice := h.clarifyExpiredNotice()
// 1b3. clarify answer — if she asked a live question last turn, this
// utterance is its answer, not a fresh command. After the confirm check: a
// y/n gate is armed by her own prompt and is the narrower claim on the
// utterance.
if reply, handled := h.resolveClarifyAnswer(ctx, text); handled {
return h.reply(ctx, reply, nil)
}
@@ -404,7 +434,7 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
// unreliably (it's a command, not a free-form query), so we match it
// before routing. Same pattern as the confirm turn above.
if reply, handled := h.resolveQuietToggle(ctx, text); handled {
return h.reply(ctx, reply, nil)
return h.reply(ctx, withNotice(expiredNotice, reply), nil)
}
// 2. router — classify the utterance.
@@ -413,10 +443,10 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
// ErrNoIntents ⇒ classifier unseeded (cold boot). reply with a
// "still warming up" rather than a wire error.
if errors.Is(err, router.ErrNoIntents) {
return h.reply(ctx, "я ещё не понимаю свободную речь — скоро научусь.", nil)
return h.reply(ctx, withNotice(expiredNotice, "я ещё не понимаю свободную речь — скоро научусь."), nil)
}
log.Printf("voice: router error: %v", err)
return h.reply(ctx, "не получилось разобрать команду.", nil)
return h.reply(ctx, withNotice(expiredNotice, "не получилось разобрать команду."), nil)
}
// 2b. dialogue — fill this turn's missing slots from a prior same-intent
@@ -437,7 +467,7 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
// stands.
if dec.Clarify {
if question, asked := h.askClarify(dec); asked {
return h.reply(ctx, question, nil)
return h.reply(ctx, withNotice(expiredNotice, question), nil)
}
}
@@ -453,7 +483,7 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
// 5. tts — synthesise the reply text; return to the voice server which
// ships it back on the conn.
return h.reply(ctx, replyText, nil)
return h.reply(ctx, withNotice(expiredNotice, replyText), nil)
}
// handleText — the core reactive path without stt/tts: confirm check →
@@ -468,7 +498,9 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
return reply
}
// 1b2. clarify answer — same check as HandlePushToTalk.
// 1b2/1b3. expired clarify then clarify answer — same order and reasons as
// HandlePushToTalk.
expiredNotice := h.clarifyExpiredNotice()
if reply, handled := h.resolveClarifyAnswer(ctx, text); handled {
return reply
}
@@ -477,10 +509,10 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
dec, err := h.router.Route(ctx, text, h.now())
if err != nil {
if errors.Is(err, router.ErrNoIntents) {
return "я ещё не понимаю свободную речь — скоро научусь."
return withNotice(expiredNotice, "я ещё не понимаю свободную речь — скоро научусь.")
}
log.Printf("voice: handleText router error: %v", err)
return "не получилось разобрать команду."
return withNotice(expiredNotice, "не получилось разобрать команду.")
}
log.Printf("voice: route result: intent=%s slots=%+v", dec.Intent, dec.Slots)
@@ -497,7 +529,7 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
// 2c. clarify — same as HandlePushToTalk: ask about the one missing thing.
if dec.Clarify {
if question, asked := h.askClarify(dec); asked {
return question
return withNotice(expiredNotice, question)
}
}
@@ -509,7 +541,7 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
if replyText == "" {
replyText = h.replier.Reply(dec)
}
return replyText
return withNotice(expiredNotice, replyText)
}
// applyAction — executes the router's Decision. Intent-by-intent:
@@ -720,7 +752,9 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
}
// Calendar questions: "что у меня сегодня?", "планы на завтра?"
if date, ok := router.ParseCalendarDate(dec.Utterance, time.Now()); ok {
// h.now(), not time.Now(): the handler's clock is the injected one, so
// this arm can be tested at a fixed time like the rest.
if date, ok := router.ParseCalendarDate(dec.Utterance, h.now()); ok {
events, err := h.api.CalendarEvents(ctx, date, date.Add(24*time.Hour))
if err != nil {
log.Printf("voice: calendar events: %v", err)
@@ -755,27 +789,54 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
log.Printf("voice: embed query: %v", err)
return "не получилось найти ответ."
}
// Long-term memory first: ONE search over everything Maven remembers
// (notes and facts share this index) and ONE confidence gate, so the
// memory that is clearly the best match answers — a note just as much
// as a fact.
//
// This used to run only after the notes-only gate below had already
// rejected the same note at the same score, which no note could ever
// survive a second time: the branch could only return a fact (#373).
// Order, not the gate, was the bug — the set of questions Maven answers
// is unchanged, only which memory gets to answer them.
if h.memStore != nil {
if hits, herr := h.memStore.Search(ctx, vec, 3); herr == nil {
if hit, ok := bestRecall(hits, h.queryMinScore, h.queryMinMargin); ok {
text := hit.Meta["text"]
// A note is phrased in Maven's voice; a fact is read back
// as it was stored.
if hit.Meta["type"] == "note" {
if reply, perr := h.phraser.PhraseQuery(ctx, dec.Utterance, []string{text}); perr == nil && reply != "" {
return reply
}
}
return text
}
} else {
log.Printf("voice: memory search: %v", herr)
}
}
notes, err := h.api.QueryNotes(ctx, vec, 5)
if err != nil {
log.Printf("voice: query notes: %v", err)
return "не получилось найти ответ."
}
// Confidence gate: below threshold, say "I don't know" rather than read
// back the least-unrelated note — a confident wrong recall is worse than
// a gap (spec's "not a guesser-of-truth"). Same instinct as the loop's
// since(key)==null → don't fire. Tuned for the ONNX embedder; the Hash
// floor scores lexically and may rarely clear it.
if len(notes) == 0 || notes[0].Score < h.queryMinScore {
// Long-term memory recall (notes + facts) before general knowledge:
// the notes table can't answer fact questions, but the memory store
// indexes both. Only runs when notes-RAG already gave up → additive.
if h.memStore != nil {
if hits, herr := h.memStore.Search(ctx, vec, 3); herr == nil {
if text, ok := bestRecall(hits, h.queryMinScore); ok {
return text
}
}
}
// Notes-only pass, for notes the vector index above does not hold (an
// older note written before it existed). Same gate, notes-only
// candidates.
//
// Confidence gate: below it, say "I don't know" rather than read back
// the least-unrelated note — a confident wrong recall is worse than a
// gap (spec's "not a guesser-of-truth"). Same instinct as the loop's
// since(key)==null → don't fire. Two parts: an absolute cosine floor,
// and a margin over the runner-up, which is the part that works with
// the e5 embedder's narrow score band. See memory.Confident.
noteScores := make([]float64, len(notes))
for i, n := range notes {
noteScores[i] = n.Score
}
if !memory.ConfidentScores(noteScores, h.queryMinScore, h.queryMinMargin) {
// Try general knowledge from the phraser before giving up
reply, err := h.phraser.PhraseQuery(ctx, dec.Utterance, nil)
if err != nil || reply == "" {
@@ -877,6 +938,101 @@ var ruMonths = []string{
"июля", "августа", "сентября", "октября", "ноября", "декабря",
}
// onlyLocalTimeReply — the honest answer when the user asks the time somewhere
// other than here. She only keeps one clock, and saying so is better than
// naming the wrong city's time.
//
// There used to be a city→time-zone table here. It was removed on purpose: the
// user only ever asks for local time, so the table was a second list of cities
// to keep in step with the weather one for no gain.
const onlyLocalTimeReply = "я знаю только местное время, про другие города пока не скажу."
// notPlaceAfterV — words that follow "в" without naming a place, so
// mentionsUnknownPlace does not mistake them for a city.
var notPlaceAfterV = map[string]bool{
"данный": true, "данную": true, "этот": true, "эту": true,
"котором": true, "какое": true, "какой": true, "который": true,
"общем": true, "точности": true, "курсе": true, "сутках": true,
"часах": true, "минутах": true, "секундах": true, "неделе": true,
}
// mentionsUnknownPlace reports whether the question has a "в <слово>" phrase
// that looks like a place we do not know ("который час в киеве"). Used only to
// pick the honest "local time only" reply instead of answering local time as
// if it were the city's.
func mentionsUnknownPlace(u string) bool {
toks := strings.Fields(u)
for i := 0; i+1 < len(toks); i++ {
if toks[i] != "в" && toks[i] != "во" {
continue
}
next := strings.Trim(toks[i+1], ".,?!")
if next == "" || notPlaceAfterV[next] {
continue
}
// A number after "в" is a clock ("в 5 часов"), not a place.
if _, err := strconv.Atoi(strings.SplitN(next, ":", 2)[0]); err == nil {
continue
}
return true
}
return false
}
// onlyNearDaysReply — she can work out today, tomorrow, the day after and
// yesterday, and nothing further. Said out loud instead of answering today's
// date for a day she did not understand.
const onlyNearDaysReply = "я считаю только сегодня, завтра, послезавтра и вчера — про другие дни пока не скажу."
// dayWords — day references the calendar parser cannot resolve. A weekday name
// or a "через …" phrase means he asked about a specific other day.
var dayWords = []string{
"понедельник", "вторник", "сред", "четверг", "пятниц", "суббот", "воскресен",
"через", "monday", "tuesday", "wednesday", "thursday", "friday", "saturday", "sunday",
}
// mentionsUnknownDay reports whether the question names a day the calendar
// parser could not resolve. Mirror of mentionsUnknownPlace: it exists only to
// pick an honest reply over a confidently wrong one.
//
// Only called after ParseCalendarDate has already failed, so "завтра" and the
// other words it does know never reach here.
func mentionsUnknownDay(u string) bool {
for _, w := range dayWords {
if strings.Contains(u, w) {
return true
}
}
return false
}
// ruClock renders the clock part of the time reply: "15 часов 4 минуты".
func ruClock(t time.Time) string {
h, m := t.Hour(), t.Minute()
hourWord := ruPlural(h, "час", "часа", "часов")
if m == 0 {
return fmt.Sprintf("%d %s ровно", h, hourWord)
}
return fmt.Sprintf("%d %s %d %s", h, hourWord, m, ruPlural(m, "минута", "минуты", "минут"))
}
// dayPrefix names the day relative to now ("завтра", "вчера", …) so the date
// reply opens the way a person would say it.
func dayPrefix(now, day time.Time) string {
base := time.Date(now.Year(), now.Month(), now.Day(), 0, 0, 0, 0, now.Location())
switch int(day.Sub(base).Hours() / 24) {
case -1:
return "вчера"
case 0:
return "сегодня"
case 1:
return "завтра"
case 2:
return "послезавтра"
}
return "это"
}
func ruPlural(n int, one, two, many string) string {
n = n % 100
if n > 10 && n < 20 {
@@ -955,18 +1111,29 @@ func (h *reactiveHandler) replySystem(ctx context.Context, dec router.Decision)
switch {
case strings.Contains(u, "час") || strings.Contains(u, "врем"):
h := now.Hour()
m := now.Minute()
hourWord := ruPlural(h, "час", "часа", "часов")
if m == 0 {
return fmt.Sprintf("сейчас %d %s ровно", h, hourWord)
// "который час в киеве" — she keeps one clock, so any named place gets
// the honest answer. Never local time dressed up as the city's.
if mentionsUnknownPlace(u) {
return onlyLocalTimeReply
}
minWord := ruPlural(m, "минута", "минуты", "минут")
return fmt.Sprintf("сейчас %d %s %d %s", h, hourWord, m, minWord)
return "сейчас " + ruClock(now)
case strings.Contains(u, "день") || strings.Contains(u, "числ"):
dow := ruWeekdays[now.Weekday()]
month := ruMonths[now.Month()-1]
return fmt.Sprintf("сегодня %s, %d %s %d года", dow, now.Day(), month, now.Year())
// "какое число завтра" — answer for the day the user asked about,
// not today. Reuses the router's calendar day-word parser.
day := now
prefix := "сегодня"
if d, ok := router.ParseCalendarDate(u, now); ok {
day = d
prefix = dayPrefix(now, d)
} else if mentionsUnknownDay(u) {
// He named a day she cannot work out ("в пятницу", "через неделю").
// Answering today's date here would be the same silent wrong answer
// this arm was fixed for, so say what she can do instead.
return onlyNearDaysReply
}
dow := ruWeekdays[day.Weekday()]
month := ruMonths[day.Month()-1]
return fmt.Sprintf("%s %s, %d %s %d года", prefix, dow, day.Day(), month, day.Year())
case strings.Contains(u, "кто дома") || strings.Contains(u, "человек дома"):
return "присутствие пока не подключено к голосовому запросу."
case strings.Contains(u, "памят") || strings.Contains(u, "процессор") || strings.Contains(u, "загрузк") || strings.Contains(u, "статус") || strings.Contains(u, "работа") || strings.Contains(u, "сервис") || strings.Contains(u, "диск") || strings.Contains(u, "ip") || strings.Contains(u, "аптайм") || strings.Contains(u, "трафик") || strings.Contains(u, "интернет"):
@@ -1709,3 +1876,71 @@ func jsonStringImpl(s string) string {
b = append(b, '"')
return string(b)
}
// reembedOnStart is the -reembed flag (set in run()). Opt-in on purpose: see
// runReembed.
var reembedOnStart bool
// checkStoredEmbedder compares the embedder we just loaded with the one that
// wrote the vectors already in the DB (Vikunja #378).
//
// The two models we have both make 384-dim vectors, so a size check catches
// nothing: after a swap, recall silently compares vectors from different
// spaces and the scores are noise. So we say it out loud. Recall itself is not
// changed here — the fix is `mavend -reembed`.
func checkStoredEmbedder(dataStore *store.Store, emb router.Embedder) {
if dataStore == nil {
return
}
current := router.EmbedderID(emb)
if reembedOnStart {
runReembed(dataStore, emb, current)
return
}
stored, mismatch, err := dataStore.CheckEmbedder(context.Background(), current)
if err != nil {
log.Printf("voice: embedder marker check failed: %v", err)
return
}
if mismatch {
log.Printf("voice: WARNING embedder MISMATCH — stored vectors were written by %q but the configured embedder is %q; recall scores are noise until the notes and facts are re-embedded — run `mavend -reembed` once (Vikunja #378)", stored, current)
return
}
log.Printf("voice: embedder marker ok (%s)", current)
}
// runReembed is the one-shot backfill behind -reembed.
//
// Why a flag and not automatic on mismatch: the embedder is ONNX on the
// laptop's CPU, so a few thousand notes is minutes of work. Doing that silently
// inside a normal start would look like the daemon hanging on boot. So the user
// runs it once, deliberately, after an embedder swap; the mismatch warning
// above tells them to. It re-embeds, logs what it did, and then the daemon
// carries on serving as usual — no separate binary, no second start needed.
func runReembed(dataStore *store.Store, emb router.Embedder, current string) {
log.Printf("voice: re-embedding stored notes and facts with %s — this can take a few minutes, do not interrupt", current)
res, err := dataStore.ReembedAll(context.Background(), current,
// EmbedPassage, not EmbedQuery: these are stored texts being searched
// FOR, which is the side they were written with.
func(ctx context.Context, text string) ([]float32, error) {
return router.EmbedPassage(ctx, emb, text)
})
if err != nil {
log.Printf("voice: re-embed FAILED, nothing was changed and no marker was written — safe to run again: %v", err)
return
}
if res.Skipped {
log.Printf("voice: re-embed skipped — the stored vectors were already written by %s", current)
return
}
log.Printf("voice: re-embed done — %d notes in the notes table, %d notes and %d facts in the memory index, took %s; stored vectors now belong to %s",
res.Notes, res.MemNotes, res.Facts, res.Took.Round(time.Second), current)
// A row with no text cannot be re-embedded, so its vector is still the old
// model's noise while the marker now says everything is current. Both write
// paths always store the text, so this should be zero — say it loudly
// rather than bury it in the line above if it ever isn't.
if res.NoText > 0 {
log.Printf("voice: WARNING %d stored rows had no text, so their vectors could not be re-embedded and are still noise; they will never match anything useful (Vikunja #378)", res.NoText)
}
}
+4 -1
View File
@@ -40,7 +40,10 @@
"tokenizer_path": "/opt/maven/models/embedder/multilingual-e5-small/tokenizer.json",
"lib_path": "/opt/maven/lib/libonnxruntime.so"
},
"llm_router": false,
"llm_router": true,
"query_min_score": 0.55,
"query_min_margin": 0.008,
"clarify_max_attempts": 3,
"tool_timeout": "30s",
"tools": [
{ "name": "status", "cmd": ["systemctl", "status"], "scope": "homelab", "destructive": false },
+70 -11
View File
@@ -258,18 +258,25 @@ type VoiceConfig struct {
RouterThreshold float64 `json:"router_threshold,omitempty"`
// LLMRouter — route with the resident model instead of the embedding
// classifier. Measured on the held-out fixture (ROUTING-EVAL-31-07-2026.md)
// the model gets 50.0% of intents right against the classifier's 36.8%, but
// it costs about 800ms per turn instead of 30ms.
// classifier. On by default since Vikunja #320.
//
// TODO: the default stays false until this lands.
// Extractor.Extract never runs on an LLM decision, so acts arrive with no
// Fn and reminders with no Time. Turning this on today makes routing more
// accurate and less safe.
// Measured on the held-out fixture (ROUTING-EVAL-31-07-2026.md): 63.2% of
// intents right against the classifier's 50.0%, and no route errors. It
// costs about 1s per turn instead of 30ms.
//
// The router can now refuse: it answers "unknown" when it cannot route, and
// the turn drops to the classifier and its clarify gate (Vikunja #359).
LLMRouter bool `json:"llm_router,omitempty"`
// It is safe to leave on. The model can refuse it answers "unknown" when
// it cannot route, and the turn drops to the classifier and its clarify
// gate. Any LLM error does the same, so a turn never breaks on the model.
// Slot extraction runs on LLM decisions too, so acts get their Fn and
// reminders their Time.
//
// Set it false to go back to the classifier, e.g. on a box with no
// llama-server or when 1s a turn is too slow.
//
// It is a pointer so that "missing from the file" and "explicitly false"
// are different things: missing means on, false means off. Read it with
// UseLLMRouter(), not directly.
LLMRouter *bool `json:"llm_router,omitempty"`
// QueryMinScore — the note-recall confidence gate. Top cosine below this
// ⇒ "I don't know" instead of a guess. Tuned for the ONNX embedder (0.55);
@@ -277,12 +284,31 @@ type VoiceConfig struct {
// default if unset.
QueryMinScore float64 `json:"query_min_score,omitempty"`
// QueryMinMargin — the second half of the recall gate: the top hit must
// beat the runner-up by more than this. The absolute score above cannot do
// the job on its own, because the e5 embedder puts every cosine in one
// narrow high band, so a made-up question scores as high as a real one.
// The margin asks whether one note is clearly the best instead.
// Negative ⇒ off. 0 ⇒ the default below.
QueryMinMargin float64 `json:"query_min_margin,omitempty"`
// ClarifyMaxAttempts — how many clarifying questions she may ask about one
// request before she gives up and says she did not understand. Default 3.
ClarifyMaxAttempts int `json:"clarify_max_attempts,omitempty"`
// Persona — optional prompt prefix that tunes maven's character. Prepended
// to every LLM system prompt (nudge phrasing, note queries, general
// knowledge). Empty string ⇒ current hardcoded persona (feminine-gendered
// Russian self-reference). Example: "Be formal and answer in English only."
Persona string `json:"persona,omitempty"`
// OwnerName / City — optional facts about the owner, added to the shared
// context block (internal/persona). Empty is fine: the block still states
// who he is grammatically (a man, addressed as "ты") and the current time.
// Nothing about correct behaviour may depend on these being filled in.
OwnerName string `json:"owner_name,omitempty"`
City string `json:"city,omitempty"`
// Weather — the weather provider config. nil ⇒ the daemon wires
// the stub provider (returns ErrNotConfigured — "погода не настроена").
// Set provider to "open-meteo" to use the keyless Open-Meteo API.
@@ -404,7 +430,16 @@ const (
DefaultAutotuneInterval = 10 * time.Minute
DefaultRouterThreshold = 0.55
DefaultQueryMinScore = 0.55
DefaultToolTimeout = 30 * time.Second
// Read off the margin sweep in internal/memory/recalleval on the e5
// embedder: 0.008 answers 68% of real questions (down from 72%) and cuts
// false recall from 5/5 to 1/5. Every larger delta costs real recall
// without removing that last one until 0.020, which drops recall to 44%.
DefaultQueryMinMargin = 0.008
// DefaultClarifyMaxAttempts — see dialogue.DefaultMaxAttempts.
DefaultClarifyMaxAttempts = 3
DefaultToolTimeout = 30 * time.Second
// DefaultLLMRouter — route with the resident model unless told otherwise.
DefaultLLMRouter = true
DefaultFactEnrichmentInterval = 30 * time.Second
)
@@ -486,9 +521,24 @@ func (c *Config) applyDefaults() {
if c.Voice.QueryMinScore <= 0 {
c.Voice.QueryMinScore = DefaultQueryMinScore
}
// Unset ⇒ default. Negative is how you turn the margin off on purpose,
// so it is clamped to 0 rather than replaced by the default.
switch {
case c.Voice.QueryMinMargin == 0:
c.Voice.QueryMinMargin = DefaultQueryMinMargin
case c.Voice.QueryMinMargin < 0:
c.Voice.QueryMinMargin = 0
}
if c.Voice.ClarifyMaxAttempts <= 0 {
c.Voice.ClarifyMaxAttempts = DefaultClarifyMaxAttempts
}
if c.Voice.ToolTimeout <= 0 {
c.Voice.ToolTimeout = Duration(DefaultToolTimeout)
}
if c.Voice.LLMRouter == nil {
on := DefaultLLMRouter
c.Voice.LLMRouter = &on
}
}
// routines: default severity to care-class (1) — the safe floor: a
@@ -507,6 +557,15 @@ func (c *Config) applyDefaults() {
}
}
// UseLLMRouter reports whether to route with the resident model. Unset means
// on; only an explicit false in the config turns it off.
func (v *VoiceConfig) UseLLMRouter() bool {
if v == nil || v.LLMRouter == nil {
return DefaultLLMRouter
}
return *v.LLMRouter
}
func (c *Config) validate() error {
if c.Phraser != nil {
if c.Phraser.ModelPath == "" {
+16 -4
View File
@@ -171,14 +171,26 @@ func TestWeatherConfigNilOK(t *testing.T) {
}
}
func TestLLMRouterDefaultsOff(t *testing.T) {
func TestLLMRouterDefaultsOn(t *testing.T) {
p := writeConfig(t, `{"voice":{"enabled":true,"bind":"127.0.0.1:9100"}}`)
c, err := Load(p)
if err != nil {
t.Fatalf("Load: %v", err)
}
if c.Voice.LLMRouter {
t.Error("voice.llm_router absent should mean false")
if !c.Voice.UseLLMRouter() {
t.Error("voice.llm_router absent should mean on")
}
}
// Missing and explicitly false must not mean the same thing.
func TestLLMRouterExplicitFalseTurnsItOff(t *testing.T) {
p := writeConfig(t, `{"voice":{"enabled":true,"bind":"127.0.0.1:9100","llm_router":false}}`)
c, err := Load(p)
if err != nil {
t.Fatalf("Load: %v", err)
}
if c.Voice.UseLLMRouter() {
t.Error("voice.llm_router false should turn it off")
}
}
@@ -188,7 +200,7 @@ func TestLLMRouterRead(t *testing.T) {
if err != nil {
t.Fatalf("Load: %v", err)
}
if !c.Voice.LLMRouter {
if !c.Voice.UseLLMRouter() {
t.Error("voice.llm_router true was not read")
}
}
+21 -3
View File
@@ -137,9 +137,10 @@ func NewDispatcher(cfg Config) *Dispatcher {
// picks for (severity, presence), sends via the matching sink, and records
// one nudge row per successful send. returns the dispatches (one per channel).
//
// a Drop channel = no send, no record (the nudge was suppressed by routing,
// not by a failure — "a missed water nudge is noise"). a nil sink = channel
// not wired, skip silently. a send error stops the dispatch and returns what
// a Drop channel = no send (the nudge was suppressed by routing, not by a
// failure — "a missed water nudge is noise"), but it does leave a 'dropped'
// outbox row so the suppression is visible. a nil sink = channel not wired,
// skip silently. a send error stops the dispatch and returns what
// got through — the daemon decides whether to retry.
func (d *Dispatcher) DispatchNudge(ctx context.Context, pn PhrasedNudge, now time.Time) ([]Dispatch, error) {
c := pn.Candidate
@@ -148,6 +149,16 @@ func (d *Dispatcher) DispatchNudge(ctx context.Context, pn PhrasedNudge, now tim
for i := 0; i < len(channels); i++ {
ch := channels[i]
if ch == ChannelDrop {
// the routing table suppressed this nudge on purpose (a care nudge
// while you're away is noise). that stays — but it must not be
// invisible, or "she dropped it" and "the rule never fired" look
// the same afterwards. no nudges row: that table feeds the
// ignored_rate signal, and a nudge nobody could see must not
// count as ignored.
id := d.beginOutbox(ctx, "nudge", c.Rule.Name, 0, ch, pn.Summary, now)
d.completeOutbox(ctx, id, store.DeliveryDropped, now)
log.Printf("dispatcher: dropped %s (sev%d, presence=%s) — routing table suppressed it",
c.Rule.Name, c.Severity, c.State.Presence)
continue
}
s := Sendable{
@@ -396,6 +407,13 @@ func messageForChannel(s Sendable) string {
if !isAway(s.Channel) {
return s.Body
}
return AwayMessage(s)
}
// AwayMessage — the only text an off-box channel may ever carry. Exported so
// the away sinks share this one rule instead of each inventing a fallback: the
// summary if we have one, otherwise a fixed generic line. Never the body.
func AwayMessage(s Sendable) string {
if s.Summary != "" {
return s.Summary
}
+6 -4
View File
@@ -37,10 +37,12 @@ func TestVoiceNoSessionFallthroughLeavesOutboxTrail(t *testing.T) {
[]string{"voice", "ntfy"}, []string{store.DeliveryFailed, store.DeliverySent}},
{"sev4 falls through to telegram", loop.Sev4,
[]string{"voice", "telegram"}, []string{store.DeliveryFailed, store.DeliverySent}},
{"sev1 does not fall through", loop.Sev1,
[]string{"voice"}, []string{store.DeliveryFailed}},
{"sev2 does not fall through", loop.Sev2,
[]string{"voice"}, []string{store.DeliveryFailed}},
// care severities still don't reach an away channel; since #370 the
// drop itself is a visible row instead of nothing.
{"sev1 drops instead of falling through", loop.Sev1,
[]string{"voice", "drop"}, []string{store.DeliveryFailed, store.DeliveryDropped}},
{"sev2 drops instead of falling through", loop.Sev2,
[]string{"voice", "drop"}, []string{store.DeliveryFailed, store.DeliveryDropped}},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
+9 -13
View File
@@ -2,10 +2,10 @@
//
// ntfy is the away-channel for sev3 (ops soft) nudges, sev4 (ops hard)
// nudges when present (alongside voice), and reminders when away. the
// message body is the Sendable's Summary — the minimal-body rule from the
// message body is delivery.AwayMessage — the minimal-body rule from the
// spec ("disk low on homesrv," not detail; no shoulder-surf exfil through
// the relay). voice gets Body; away channels get Summary, enforced at the
// sink so a phraser bug can't exfil.
// the relay). the dispatcher already strips detail off away sendables; the
// sink uses the same helper so it can't leak the body on its own either.
//
// ntfy runs locally (docker, 127.0.0.1:8085, deny-all auth). maven publishes
// with a dedicated user (write-only to maven-* topics) — the credential is a
@@ -69,18 +69,14 @@ func New(cfg Config) (*Sink, error) {
}, nil
}
// Send publishes one notification to ntfy. the body is the Sendable's Summary
// (minimal body); Title is "maven" (consistent sender identity on the lock
// screen — the content is in the body). Priority maps from severity/kind so
// Send publishes one notification to ntfy. the body is the minimal away
// message (never the full body); Title is "maven" (consistent sender identity
// on the lock screen — the content is in the body). Priority maps from severity/kind so
// the phone client can ring differently for an alarm vs a soft ops nudge.
func (s *Sink) Send(ctx context.Context, d delivery.Sendable) error {
body := d.Summary
if body == "" {
body = d.Body // terse full message beats no message
}
if body == "" {
return fmt.Errorf("ntfysink: empty message for %s", d.Channel)
}
// never fall back to d.Body: ntfy leaves the box, so an empty summary gets
// a generic line instead of the full detail.
body := delivery.AwayMessage(d)
req, err := http.NewRequestWithContext(ctx, http.MethodPost, s.topicURL(), strings.NewReader(body))
if err != nil {
+16 -9
View File
@@ -147,9 +147,9 @@ func TestSendBodyIsSummaryNotFullBody(t *testing.T) {
}
}
func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
// a terse full message is better than no message; the phraser should
// produce a summary for away-bound severities, but don't silently drop.
func TestSendNeverSendsTheBodyWhenSummaryEmpty(t *testing.T) {
// #368: this used to fall back to the full body. ntfy leaves the box, so
// an empty summary gets a fixed generic line plus the rule name instead.
rs := newRecordingServer(t, 200, "")
srv := httptest.NewServer(rs.handler())
defer srv.Close()
@@ -160,12 +160,15 @@ func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
t.Fatalf("Send: %v", err)
}
_, _, body, _, _, _ := rs.snapshot()
if body != s.Body {
t.Fatalf("fallback body: want %q, got %q", s.Body, body)
want := delivery.GenericAwayMessage + ": service_down"
if body != want {
t.Fatalf("body: want %q, got %q", want, body)
}
}
func TestSendRejectsEmptyMessage(t *testing.T) {
func TestSendNeverSendsAnEmptyMessage(t *testing.T) {
// with nothing at all to say we still send the generic line — an away
// channel can never carry detail, but it also never goes out blank.
rs := newRecordingServer(t, 200, "")
srv := httptest.NewServer(rs.handler())
defer srv.Close()
@@ -173,9 +176,13 @@ func TestSendRejectsEmptyMessage(t *testing.T) {
sink, _ := New(Config{BaseURL: srv.URL, Topic: "maven"})
s := nudgeSendable(loop.Sev3, "")
s.Body = ""
err := sink.Send(context.Background(), s)
if err == nil {
t.Fatal("want error for empty message")
s.RuleName = ""
if err := sink.Send(context.Background(), s); err != nil {
t.Fatalf("Send: %v", err)
}
_, _, body, _, _, _ := rs.snapshot()
if body != delivery.GenericAwayMessage {
t.Fatalf("body: want %q, got %q", delivery.GenericAwayMessage, body)
}
}
+2 -3
View File
@@ -208,10 +208,9 @@ func TestAwayChannelsGetMinimalBody(t *testing.T) {
// TestCareAwayDropIsRecorded — DESIGN.md's drop is a decision ("a missed water
// nudge is noise, a missed backup failure isn't"), so it should be visible
// rather than vanish. Today drop is a bare `continue`: no nudge row, no outbox
// attempt, no log — nothing an operator can see afterwards.
// attempt, no log — nothing an operator can see afterwards. now it leaves a
// 'dropped' outbox row.
func TestCareAwayDropIsRecorded(t *testing.T) {
t.Skip("not implemented: dispatcher.go:149-151 skips a Drop channel with no record; there is no 'dropped' outcome in store/delivery.go:16-21")
ob := &fakeOutbox{}
d := NewDispatcher(Config{Voice: &fakeSink{}, Nudges: &fakeNudgeRecorder{}, Outbox: ob})
+12 -16
View File
@@ -2,11 +2,12 @@
//
// telegram is the away-channel for sev4 (ops hard) nudges — "disk-fire alarm
// at 2am routes to telegram, repeat til ack." the message body is the
// Sendable's Summary — the minimal-body rule from the spec ("disk low on
// homesrv," not detail; no shoulder-surf exfil through the relay). voice gets
// Body; away channels get Summary, enforced at the sink so a phraser bug can't
// exfil. additionally, protect_content=true is passed on every send so the
// message can't be forwarded out of the chat — locks the minimal body further.
// delivery.AwayMessage — the minimal-body rule from the spec ("disk low on
// homesrv," not detail; no shoulder-surf exfil through the relay). the
// dispatcher already strips detail off away sendables; the sink uses the same
// helper so it can't leak the body on its own either. additionally,
// protect_content=true is passed on every send so the message can't be
// forwarded out of the chat — locks the minimal body further.
//
// telegram's bot API is region-restricted for this homesrv — direct egress to
// api.telegram.org is unreliable. the spec's "away channels leave the box —
@@ -140,18 +141,13 @@ type telegramResp struct {
}
// Send publishes one message to the configured telegram chat. the body is the
// Sendable's Summary (minimal body); empty Summary falls back to Body (terse
// full message beats no message). protect_content=true so a phraser bug (Body
// leaking detail through Summary) can't be forwarded onward by the user or a
// chat observer — locks the minimal-body rule at the channel's own last mile.
// minimal away message (never the full body). protect_content=true so even
// that can't be forwarded onward by the user or a chat observer — locks the
// minimal-body rule at the channel's own last mile.
func (s *Sink) Send(ctx context.Context, d delivery.Sendable) error {
body := d.Summary
if body == "" {
body = d.Body
}
if body == "" {
return fmt.Errorf("telegramsink: empty message for %s", d.Channel)
}
// never fall back to d.Body: telegram leaves the box, so an empty summary
// gets a generic line instead of the full detail.
body := delivery.AwayMessage(d)
payload := sendMessageReq{
ChatID: s.cfg.ChatID,
@@ -173,9 +173,9 @@ func TestSendBodyIsSummaryNotFullBody(t *testing.T) {
}
}
func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
// terse full message beats none; the phraser should produce a summary for
// away-bound severities, but don't silently drop.
func TestSendNeverSendsTheBodyWhenSummaryEmpty(t *testing.T) {
// #368: this used to fall back to the full body. telegram leaves the box,
// so an empty summary gets a fixed generic line plus the rule name.
rs := newRecordingServer(t, 200, "")
srv := httptest.NewServer(rs.handler())
defer srv.Close()
@@ -188,12 +188,14 @@ func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
_, _, body, _, _ := rs.snapshot()
var req sendMessageReq
_ = json.Unmarshal([]byte(body), &req)
if req.Text != s.Body {
t.Fatalf("fallback text: want %q, got %q", s.Body, req.Text)
want := delivery.GenericAwayMessage + ": service_down"
if req.Text != want {
t.Fatalf("text: want %q, got %q", want, req.Text)
}
}
func TestSendRejectsEmptyMessage(t *testing.T) {
func TestSendNeverSendsAnEmptyMessage(t *testing.T) {
// with nothing at all to say we still send the generic line.
rs := newRecordingServer(t, 200, "")
srv := httptest.NewServer(rs.handler())
defer srv.Close()
@@ -201,9 +203,15 @@ func TestSendRejectsEmptyMessage(t *testing.T) {
sink, _ := New(sinkCfg(srv.URL))
s := nudgeSendable(loop.Sev4, "")
s.Body = ""
err := sink.Send(context.Background(), s)
if err == nil {
t.Fatal("want error for empty message")
s.RuleName = ""
if err := sink.Send(context.Background(), s); err != nil {
t.Fatalf("Send: %v", err)
}
_, _, body, _, _ := rs.snapshot()
var req sendMessageReq
_ = json.Unmarshal([]byte(body), &req)
if req.Text != delivery.GenericAwayMessage {
t.Fatalf("text: want %q, got %q", delivery.GenericAwayMessage, req.Text)
}
}
+53 -31
View File
@@ -17,10 +17,11 @@ const (
SlotText Slot = "text" // Slots.Text
)
// MaxAttempts is 1 because Maven is not a nag (DESIGN.md § Non-goals). She asks
// one clarifying question. If the answer still leaves the slot empty she drops
// the request instead of asking again.
const MaxAttempts = 1
// DefaultMaxAttempts — how many questions she may ask about one request.
// Three, because after three tries the likely problem is that she misheard the
// whole request, not one slot — so another question about that slot won't help.
// Configurable: voice.clarify_max_attempts.
const DefaultMaxAttempts = 3
// PendingQuestion is what Maven holds while she waits for an answer to an open
// question. Unlike the yes/no confirms in cmd/mavend/voice.go, the answer here
@@ -32,7 +33,17 @@ type PendingQuestion struct {
Utterance string // the user's original raw words
Asked time.Time
TTL time.Duration
Attempts int // questions already asked; capped by MaxAttempts
Attempts int // questions already asked
// MaxAttempts caps Attempts. 0 ⇒ DefaultMaxAttempts.
MaxAttempts int
}
// maxAttempts is MaxAttempts with the default filled in.
func (q *PendingQuestion) maxAttempts() int {
if q.MaxAttempts <= 0 {
return DefaultMaxAttempts
}
return q.MaxAttempts
}
func (q *PendingQuestion) IsExpired(now time.Time) bool {
@@ -41,12 +52,9 @@ func (q *PendingQuestion) IsExpired(now time.Time) bool {
// CanAsk reports whether Maven may ask another question about this request.
func (q *PendingQuestion) CanAsk() bool {
return q.Attempts < MaxAttempts
return q.Attempts < q.maxAttempts()
}
// TODO: the daemon will phrase the question text from Missing (one short ru
// question per Slot, feminine self-reference) and speak it here.
// ClarifyStore holds the parked questions. Same shape and locking as
// SessionStore: keyed by dialogue id, expired entries dropped on read.
type ClarifyStore struct {
@@ -67,8 +75,7 @@ func NewClarifyStore(defaultTTL time.Duration) *ClarifyStore {
}
}
// TODO: the daemon will Put a question here when Decision.Clarify fires, in
// place of the flat "не разобрала" reply (cmd/mavend/voice.go).
// Put parks a question. Called on a clarify decision (cmd/mavend/clarify.go).
func (s *ClarifyStore) Put(id string, q *PendingQuestion) {
if q.TTL <= 0 {
q.TTL = s.defaultTTL
@@ -78,8 +85,7 @@ func (s *ClarifyStore) Put(id string, q *PendingQuestion) {
s.mu.Unlock()
}
// TODO: the daemon will Get on the next turn, parse that turn into Slots, call
// Answer, and Delete — the open-question twin of resolveConfirm.
// Get returns the live parked question, or nil when there is none.
func (s *ClarifyStore) Get(id string, now time.Time) *PendingQuestion {
s.mu.RLock()
q, ok := s.questions[id]
@@ -94,6 +100,21 @@ func (s *ClarifyStore) Get(id string, now time.Time) *PendingQuestion {
return q
}
// TakeExpired reports whether a question was parked here but its TTL ran out,
// and drops it. Get drops such a question silently, which leaves the user
// thinking his request is still alive — the caller uses this to tell him it is
// gone before treating his words as a fresh utterance.
func (s *ClarifyStore) TakeExpired(id string, now time.Time) bool {
s.mu.Lock()
defer s.mu.Unlock()
q, ok := s.questions[id]
if !ok || !q.IsExpired(now) {
return false
}
delete(s.questions, id)
return true
}
func (s *ClarifyStore) Delete(id string) {
s.mu.Lock()
delete(s.questions, id)
@@ -101,44 +122,45 @@ func (s *ClarifyStore) Delete(id string) {
}
// Answer merges the slots parsed from the user's answer into the parked ones.
// Only the slots listed in Missing are filled, and an already filled slot is
// never overwritten — the answer completes the original request, it does not
// restate it. Parsing the answer text into `answer` is the caller's job; this
// package must stay free of internal/router.
// Only the slots listed in Missing are touched. Within those, a value the answer
// carries WINS over what was parked: she asked about this slot, so «нет, в пять»
// after «в три» must replace the time, not be thrown away.
//
// This is the clarify answer only. A correction in a fresh turn ("вообще-то
// перенеси на пять") is a different code path (followUpMerge) — not here.
//
// Parsing the answer text into `answer` is the caller's job; this package must
// stay free of internal/router.
func (q *PendingQuestion) Answer(text string, answer Slots) Slots {
out := q.Slots
for _, slot := range q.Missing {
switch slot {
case SlotTime:
if !out.HasTime && answer.HasTime {
if answer.HasTime {
out.Time = answer.Time
out.HasTime = true
}
case SlotKey:
if !out.HasKey && answer.HasKey {
if answer.HasKey {
out.Key = answer.Key
out.HasKey = true
}
case SlotValue:
if out.Value == "" && answer.Value != "" {
if answer.Value != "" {
out.Value = answer.Value
}
case SlotFn:
if !out.HasFn && answer.HasFn {
if answer.HasFn {
out.Fn = answer.Fn
out.HasFn = true
if len(out.Args) == 0 {
out.Args = append([]string(nil), answer.Args...)
}
out.Args = append([]string(nil), answer.Args...)
}
case SlotText:
if out.Text == "" {
if answer.Text != "" {
out.Text = answer.Text
} else {
// No parse for a text slot — the raw answer IS the text.
out.Text = text
}
if answer.Text != "" {
out.Text = answer.Text
} else if out.Text == "" {
// No parse for a text slot — the raw answer IS the text.
out.Text = text
}
}
}
+48 -26
View File
@@ -1,6 +1,7 @@
package dialogue
import (
"reflect"
"testing"
"time"
)
@@ -59,6 +60,28 @@ func TestClarifyStoreGetPutDelete(t *testing.T) {
}
}
// TestClarifyStoreTakeExpired — TakeExpired reports (and drops) only a question
// whose TTL ran out.
func TestClarifyStoreTakeExpired(t *testing.T) {
s := NewClarifyStore(time.Minute)
if s.TakeExpired("voice", base) {
t.Fatal("nothing parked ⇒ nothing expired")
}
s.Put("voice", &PendingQuestion{Missing: []Slot{SlotTime}, Asked: base, TTL: time.Minute})
if s.TakeExpired("voice", base.Add(30*time.Second)) {
t.Fatal("a live question must not report as expired")
}
if s.Get("voice", base.Add(30*time.Second)) == nil {
t.Fatal("a live question must survive TakeExpired")
}
if !s.TakeExpired("voice", base.Add(2*time.Minute)) {
t.Fatal("a stale question must report as expired")
}
if s.TakeExpired("voice", base.Add(2*time.Minute)) {
t.Fatal("TakeExpired must drop the question, so the second call is false")
}
}
func TestNewClarifyStoreDefaultTTL(t *testing.T) {
s := NewClarifyStore(0)
q := &PendingQuestion{Asked: base}
@@ -89,12 +112,13 @@ func TestAnswerFillsOnlyMissingSlots(t *testing.T) {
want: Slots{Text: "напомни позвонить", Time: answerTime, HasTime: true},
},
{
name: "does not overwrite a filled time",
// He restated it: «нет, в пять». The new value wins.
name: "a restated time overwrites the parked one",
parked: Slots{Time: other, HasTime: true},
missing: []Slot{SlotTime},
text: "в три",
text: "нет, в три",
answer: Slots{Time: answerTime, HasTime: true},
want: Slots{Time: other, HasTime: true},
want: Slots{Time: answerTime, HasTime: true},
},
{
name: "ignores slots that were not missing",
@@ -121,12 +145,12 @@ func TestAnswerFillsOnlyMissingSlots(t *testing.T) {
want: Slots{Fn: "restart", Args: []string{"nginx"}, HasFn: true},
},
{
name: "keeps existing args when fn was already known",
name: "a restated fn replaces the fn and its args",
parked: Slots{Fn: "restart", Args: []string{"nginx"}, HasFn: true},
missing: []Slot{SlotFn},
text: "останови postgres",
answer: Slots{Fn: "stop", Args: []string{"postgres"}, HasFn: true},
want: Slots{Fn: "restart", Args: []string{"nginx"}, HasFn: true},
want: Slots{Fn: "stop", Args: []string{"postgres"}, HasFn: true},
},
{
name: "raw answer becomes the text when nothing was parsed",
@@ -157,36 +181,34 @@ func TestAnswerFillsOnlyMissingSlots(t *testing.T) {
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
q := &PendingQuestion{Slots: tc.parked, Missing: tc.missing, Asked: base}
got := q.Answer(tc.text, tc.answer)
if got.Time != tc.want.Time || got.HasTime != tc.want.HasTime ||
got.Key != tc.want.Key || got.HasKey != tc.want.HasKey ||
got.Value != tc.want.Value || got.Text != tc.want.Text ||
got.Fn != tc.want.Fn || got.HasFn != tc.want.HasFn {
// Whole-struct compare: a new field in Slots is covered for free.
if got := q.Answer(tc.text, tc.answer); !reflect.DeepEqual(got, tc.want) {
t.Fatalf("Answer = %+v, want %+v", got, tc.want)
}
if len(got.Args) != len(tc.want.Args) {
t.Fatalf("Args = %v, want %v", got.Args, tc.want.Args)
}
for i := range got.Args {
if got.Args[i] != tc.want.Args[i] {
t.Fatalf("Args = %v, want %v", got.Args, tc.want.Args)
}
}
})
}
}
func TestCanAskCapsAtOneQuestion(t *testing.T) {
if MaxAttempts != 1 {
t.Fatalf("MaxAttempts = %d, want 1 (Maven asks once, she is not a nag)", MaxAttempts)
func TestCanAskAllowsThreeQuestionsByDefault(t *testing.T) {
if DefaultMaxAttempts != 3 {
t.Fatalf("DefaultMaxAttempts = %d, want 3", DefaultMaxAttempts)
}
q := &PendingQuestion{Asked: base}
if !q.CanAsk() {
t.Fatal("a fresh question should be askable")
q := &PendingQuestion{Asked: base} // MaxAttempts unset ⇒ the default
for i := 0; i < 3; i++ {
if !q.CanAsk() {
t.Fatalf("question %d should be allowed", i+1)
}
q.Attempts++
}
q.Attempts = MaxAttempts
if q.CanAsk() {
t.Fatal("the question should not be asked twice")
t.Fatal("a fourth question must not be allowed")
}
}
func TestCanAskHonoursConfiguredMax(t *testing.T) {
q := &PendingQuestion{Asked: base, MaxAttempts: 1, Attempts: 1}
if q.CanAsk() {
t.Fatal("MaxAttempts 1 means one question only")
}
}
+74
View File
@@ -1,8 +1,12 @@
package dialogue
import (
"context"
"encoding/json"
"sync"
"time"
"github.com/kami/maven/internal/store"
)
type Intent string
@@ -50,10 +54,22 @@ func (s *Session) IsExpired(now time.Time) bool {
return now.After(s.Timestamp.Add(s.TTL))
}
// SessionPersister — the bit of the store the session needs, as an interface
// so tests can swap it out. Data is an opaque blob: the store never looks
// inside, we encode the session as JSON here.
type SessionPersister interface {
SaveDialogueSession(ctx context.Context, id string, data []byte, ts time.Time, ttl time.Duration) error
DeleteDialogueSession(ctx context.Context, id string) error
LoadDialogueSessions(ctx context.Context, now time.Time) ([]store.DialogueSessionRow, error)
}
// SessionStore keeps the live sessions in a map (the fast path) and mirrors
// every write to the persister, so a daemon restart can load them back.
type SessionStore struct {
mu sync.RWMutex
sessions map[string]*Session
defaultTTL time.Duration
persist SessionPersister // may be nil: memory only (tests, no-store paths)
}
func NewSessionStore(defaultTTL time.Duration) *SessionStore {
@@ -66,6 +82,43 @@ func NewSessionStore(defaultTTL time.Duration) *SessionStore {
}
}
// NewPersistentSessionStore — same store, but writes also go to the DB.
// Call Load once after this to bring back sessions from a previous run.
func NewPersistentSessionStore(defaultTTL time.Duration, p SessionPersister) *SessionStore {
s := NewSessionStore(defaultTTL)
s.persist = p
return s
}
// Load — read the saved sessions back into memory. Anything past its TTL is
// dropped (and deleted from the DB by the store), never revived.
func (s *SessionStore) Load(ctx context.Context, now time.Time) error {
if s.persist == nil {
return nil
}
rows, err := s.persist.LoadDialogueSessions(ctx, now)
if err != nil {
return err
}
s.mu.Lock()
defer s.mu.Unlock()
for _, r := range rows {
var sess Session
if err := json.Unmarshal(r.Data, &sess); err != nil {
// A blob we can't read is not worth failing a startup over.
continue
}
if sess.TTL <= 0 {
sess.TTL = r.TTL
}
if sess.IsExpired(now) {
continue
}
s.sessions[r.ID] = &sess
}
return nil
}
func (s *SessionStore) Get(id string, now time.Time) *Session {
s.mu.RLock()
sess, ok := s.sessions[id]
@@ -87,12 +140,33 @@ func (s *SessionStore) Put(id string, sess *Session) {
s.mu.Lock()
s.sessions[id] = sess
s.mu.Unlock()
s.save(id, sess)
}
func (s *SessionStore) Delete(id string) {
s.mu.Lock()
delete(s.sessions, id)
s.mu.Unlock()
if s.persist != nil {
_ = s.persist.DeleteDialogueSession(context.Background(), id)
}
}
// save — mirror one session to the DB. Best effort: memory already has it, so
// a write error costs us the restart safety net, not the current turn.
func (s *SessionStore) save(id string, sess *Session) {
if s.persist == nil {
return
}
data, err := json.Marshal(sess)
if err != nil {
return
}
ts := sess.Timestamp
if ts.IsZero() {
ts = time.Now()
}
_ = s.persist.SaveDialogueSession(context.Background(), id, data, ts, sess.TTL)
}
func InheritSlots(prev, cur Slots) Slots {
+118
View File
@@ -0,0 +1,118 @@
package dialogue
import (
"context"
"path/filepath"
"testing"
"time"
"github.com/kami/maven/internal/store"
)
// openStore — a store on disk, so a second handle can reopen the same file.
func openStore(t *testing.T, path string) *store.Store {
t.Helper()
s, err := store.Open(context.Background(), path)
if err != nil {
t.Fatalf("open store: %v", err)
}
t.Cleanup(func() { _ = s.Close() })
return s
}
// A session written before a restart comes back and still merges a follow-up.
func TestSessionSurvivesRestart(t *testing.T) {
ctx := context.Background()
path := filepath.Join(t.TempDir(), "maven_test.db")
now := time.Now().UTC().Truncate(time.Millisecond)
first := openStore(t, path)
before := NewPersistentSessionStore(2*time.Minute, first)
before.Put("voice", &Session{
Intent: IntentReminder,
Slots: Slots{Text: "полить цветы", Time: now.Add(time.Hour), HasTime: true},
Timestamp: now,
TTL: 2 * time.Minute,
})
if err := first.Close(); err != nil {
t.Fatalf("close: %v", err)
}
// fresh handle, fresh in-memory map — as after a daemon restart
after := openStore(t, path)
reloaded := NewPersistentSessionStore(2*time.Minute, after)
if err := reloaded.Load(ctx, now.Add(10*time.Second)); err != nil {
t.Fatalf("Load: %v", err)
}
sess := reloaded.Get("voice", now.Add(10*time.Second))
if sess == nil {
t.Fatal("session did not survive the restart")
}
if sess.Intent != IntentReminder {
t.Fatalf("intent = %q, want reminder", sess.Intent)
}
// the follow-up carries no text of its own; it must inherit the old one
merged := InheritSlots(sess.Slots, Slots{Time: now.Add(2 * time.Hour), HasTime: true})
if merged.Text != "полить цветы" {
t.Fatalf("merged text = %q, want the earlier turn's text", merged.Text)
}
if !merged.Time.Equal(now.Add(2 * time.Hour)) {
t.Fatalf("merged time = %v, want the follow-up's time", merged.Time)
}
}
// A session past its TTL is dead: a restart must not bring it back.
func TestExpiredSessionNotResurrected(t *testing.T) {
ctx := context.Background()
path := filepath.Join(t.TempDir(), "maven_test.db")
now := time.Now().UTC().Truncate(time.Millisecond)
first := openStore(t, path)
before := NewPersistentSessionStore(2*time.Minute, first)
before.Put("voice", &Session{
Intent: IntentReminder,
Slots: Slots{Text: "полить цветы"},
Timestamp: now,
TTL: time.Minute,
})
if err := first.Close(); err != nil {
t.Fatalf("close: %v", err)
}
after := openStore(t, path)
reloaded := NewPersistentSessionStore(2*time.Minute, after)
later := now.Add(5 * time.Minute) // well past the 1-min TTL
if err := reloaded.Load(ctx, later); err != nil {
t.Fatalf("Load: %v", err)
}
if sess := reloaded.Get("voice", later); sess != nil {
t.Fatalf("expired session came back: %+v", sess)
}
// and it is gone from the DB too, not just from memory
rows, err := after.LoadDialogueSessions(ctx, later)
if err != nil {
t.Fatalf("LoadDialogueSessions: %v", err)
}
if len(rows) != 0 {
t.Fatalf("expired row still in the DB: %+v", rows)
}
}
// Delete removes the row as well, so an ended conversation stays ended.
func TestDeleteRemovesPersistedSession(t *testing.T) {
ctx := context.Background()
path := filepath.Join(t.TempDir(), "maven_test.db")
now := time.Now().UTC().Truncate(time.Millisecond)
s := openStore(t, path)
ss := NewPersistentSessionStore(2*time.Minute, s)
ss.Put("voice", &Session{Intent: IntentChat, Timestamp: now, TTL: time.Minute})
ss.Delete("voice")
rows, err := s.LoadDialogueSessions(ctx, now)
if err != nil {
t.Fatalf("LoadDialogueSessions: %v", err)
}
if len(rows) != 0 {
t.Fatalf("row survived Delete: %+v", rows)
}
}
+116
View File
@@ -0,0 +1,116 @@
// Package kiwix reads a local Kiwix server (offline Wikipedia and friends).
//
// Why: the resident model is a 0.8B and invents facts. Letting her read a local
// article snippet beats letting her recall. Nothing here talks to the internet;
// the Kiwix server is on the same box.
//
// This is search only. Full articles are ~100KB of HTML, far too big for a 4096
// token context, so the unit of context is the search snippet (~500 chars).
package kiwix
import (
"context"
"encoding/xml"
"fmt"
"html"
"io"
"net/http"
"net/url"
"regexp"
"strconv"
"strings"
"time"
)
// Result is one search hit.
type Result struct {
Title string // article title, e.g. "Rayleigh scattering"
Path string // e.g. /content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering
Snippet string // plain text, tags stripped, entities decoded
WordCount int // 0 if the server did not say
}
// Client is a Kiwix HTTP client. Boring on purpose: no retries, no cache.
type Client struct {
base string
http *http.Client
}
// New makes a client for a Kiwix base URL like http://127.0.0.1:8034.
func New(baseURL string) *Client {
return &Client{
base: strings.TrimRight(baseURL, "/"),
http: &http.Client{Timeout: 10 * time.Second},
}
}
// Search runs a keyword search in one ZIM (book) and returns up to limit hits.
//
// Ranking is keyword based, not semantic: "Rayleigh scattering" finds the right
// article, "why is the sky blue" finds a TV episode. Pass keywords, not questions.
func (c *Client) Search(ctx context.Context, pattern, book string, limit int) ([]Result, error) {
if limit <= 0 {
limit = 5
}
q := url.Values{}
q.Set("pattern", pattern)
q.Set("books.name", book)
q.Set("format", "xml")
q.Set("pageLength", strconv.Itoa(limit))
req, err := http.NewRequestWithContext(ctx, http.MethodGet, c.base+"/search?"+q.Encode(), nil)
if err != nil {
return nil, err
}
resp, err := c.http.Do(req)
if err != nil {
return nil, err
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return nil, fmt.Errorf("kiwix search: http %d", resp.StatusCode)
}
return ParseSearchRSS(resp.Body)
}
// rss mirrors just the bits of the RSS 2.0 reply we use.
type rss struct {
Items []struct {
Title string `xml:"title"`
Link string `xml:"link"`
// innerxml keeps the <b> match markers so we can strip them ourselves.
Description struct {
Inner string `xml:",innerxml"`
} `xml:"description"`
WordCount string `xml:"wordCount"`
} `xml:"channel>item"`
}
var tagRE = regexp.MustCompile(`<[^>]*>`)
// ParseSearchRSS turns a Kiwix search reply into results. Exported so the parser
// is testable from a captured response, with no server running.
func ParseSearchRSS(r io.Reader) ([]Result, error) {
var doc rss
if err := xml.NewDecoder(r).Decode(&doc); err != nil {
return nil, fmt.Errorf("kiwix search: bad xml: %w", err)
}
out := make([]Result, 0, len(doc.Items))
for _, it := range doc.Items {
n, _ := strconv.Atoi(strings.ReplaceAll(it.WordCount, ",", ""))
out = append(out, Result{
Title: strings.TrimSpace(it.Title),
Path: strings.TrimSpace(it.Link),
Snippet: plainText(it.Description.Inner),
WordCount: n,
})
}
return out, nil
}
// plainText drops markup and decodes entities, leaving text a model can read.
func plainText(s string) string {
s = tagRE.ReplaceAllString(s, "")
s = html.UnescapeString(s)
return strings.TrimSpace(strings.Join(strings.Fields(s), " "))
}
+90
View File
@@ -0,0 +1,90 @@
package kiwix
import (
"context"
"os"
"strings"
"testing"
"time"
)
// A real reply from the live server, trimmed to two items.
const sampleRSS = `<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:opensearch="http://a9.com/-/spec/opensearch/1.1/">
<channel>
<title>Search: Rayleigh scattering</title>
<opensearch:totalResults>800</opensearch:totalResults>
<item>
<title>Rayleigh scattering</title>
<link>/content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering</link>
<description><b>Rayleigh</b> scattering causes the blue color of the sky &amp; yellow colors near the Sun.[1]</description>
<book><title>Wikipedia</title></book>
<wordCount>2,818</wordCount>
</item>
<item>
<title>HyperRayleigh scattering</title>
<link>/content/wikipedia_en_all_maxi_2026-02/Hyper%E2%80%93Rayleigh_scattering</link>
<description>...<b>Rayleigh</b> scattering" is a nonlinear optical counterpart.</description>
<book><title>Wikipedia</title></book>
<wordCount>914</wordCount>
</item>
</channel>
</rss>`
func TestParseSearchRSS(t *testing.T) {
got, err := ParseSearchRSS(strings.NewReader(sampleRSS))
if err != nil {
t.Fatalf("parse: %v", err)
}
if len(got) != 2 {
t.Fatalf("want 2 results, got %d", len(got))
}
if got[0].Title != "Rayleigh scattering" {
t.Errorf("title = %q", got[0].Title)
}
if got[0].Path != "/content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering" {
t.Errorf("path = %q", got[0].Path)
}
if got[0].WordCount != 2818 {
t.Errorf("wordCount = %d", got[0].WordCount)
}
want := "Rayleigh scattering causes the blue color of the sky & yellow colors near the Sun.[1]"
if got[0].Snippet != want {
t.Errorf("snippet = %q, want %q", got[0].Snippet, want)
}
if strings.Contains(got[1].Snippet, "<b>") {
t.Errorf("second snippet still has tags: %q", got[1].Snippet)
}
}
func TestParseSearchRSSBadXML(t *testing.T) {
if _, err := ParseSearchRSS(strings.NewReader("not xml at all")); err == nil {
t.Fatal("want an error on junk input")
}
}
// Opt-in: needs a live Kiwix server. CI has none.
// MAVEN_KIWIX_URL=http://127.0.0.1:8034 no_proxy=127.0.0.1,localhost go test -run Retrieval -v ./internal/kiwix/
func TestRetrievalEval(t *testing.T) {
base := os.Getenv("MAVEN_KIWIX_URL")
if base == "" {
t.Skip("set MAVEN_KIWIX_URL to run the retrieval eval")
}
noProxyLoopback(t)
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
defer cancel()
rep, err := RunRetrievalEval(ctx, New(base), 5)
if err != nil {
t.Fatalf("eval: %v", err)
}
// No pass bar on purpose: the number is the finding.
t.Log("\n" + rep.String() + rep.Detail())
}
// noProxyLoopback stops the box's SOCKS bridge from eating loopback requests.
func noProxyLoopback(t *testing.T) {
t.Setenv("no_proxy", "127.0.0.1,localhost")
t.Setenv("NO_PROXY", "127.0.0.1,localhost")
}
+63
View File
@@ -0,0 +1,63 @@
{
"name": "kiwix-knowledge-v1",
"book": "wikipedia_en_all_maxi_2026-02",
"note": "The 9 knowledge cases from internal/phraser/eval/talk_v1.json. Queries are hand-written English keywords on purpose: Kiwix ranks by keyword, not meaning, so a natural question fails. Writing them by hand separates 'retrieval is broken' from 'the model writes bad queries'.",
"cases": [
{
"id": "know-sky-blue",
"question": "почему небо синее?",
"query": "Rayleigh scattering sky blue",
"want_titles": ["Rayleigh scattering", "Diffuse sky radiation"]
},
{
"id": "know-boil-egg",
"question": "сколько варить яйцо вкрутую?",
"query": "boiled egg cooking",
"want_titles": ["Boiled egg", "Egg as food"]
},
{
"id": "know-ssd-vs-hdd",
"question": "чем ssd отличается от hdd?",
"query": "solid-state drive",
"want_titles": ["Solid-state drive", "Hard disk drive"]
},
{
"id": "know-cat-purr",
"question": "почему кошки мурчат?",
"query": "cat purr",
"want_titles": ["Purr", "Cat communication"]
},
{
"id": "know-hiccups",
"question": "как быстро избавиться от икоты?",
"query": "hiccup",
"want_titles": ["Hiccup"]
},
{
"id": "know-polite-form",
"question": "не могли бы вы объяснить, что такое vpn?",
"query": "virtual private network",
"want_titles": ["Virtual private network"]
},
{
"id": "know-dont-know",
"question": "как зовут моего соседа снизу?",
"query": "name of my downstairs neighbour",
"want_titles": [],
"expect_miss": true,
"note": "Unanswerable by design. Retrieval SHOULD find nothing useful. Counted as a hit only when nothing relevant comes back."
},
{
"id": "know-water-per-day",
"question": "сколько воды в день надо пить?",
"query": "human daily water requirement drinking",
"want_titles": ["Drinking water", "Water", "Dehydration", "Hydration"]
},
{
"id": "know-thunder-delay",
"question": "почему гром слышно позже молнии?",
"query": "thunder speed of sound lightning",
"want_titles": ["Thunder", "Lightning"]
}
]
}
+135
View File
@@ -0,0 +1,135 @@
package kiwix
// This scores retrieval alone: no LLM. For each general-knowledge question we
// hand-write English keywords and ask whether the article that would answer it
// comes back in the top N hits. If this score is low, reading Wikipedia cannot
// help the model no matter how good the prompt is.
//
// The unanswerable case (know-dont-know) is not scored. Whether the junk it
// returns is "nothing useful" is a human judgement, so the report just prints
// the titles and leaves the score to the 8 answerable cases.
import (
"context"
_ "embed"
"encoding/json"
"fmt"
"strings"
)
//go:embed knowledge_v1.json
var knowledgeFixtureJSON []byte
// EvalCase — one question with hand-written keywords.
type EvalCase struct {
ID string `json:"id"`
Question string `json:"question"`
Query string `json:"query"`
WantTitles []string `json:"want_titles"`
ExpectMiss bool `json:"expect_miss"`
}
type fixture struct {
Name string `json:"name"`
Book string `json:"book"`
Cases []EvalCase `json:"cases"`
}
// Outcome — what one case retrieved.
type Outcome struct {
Case EvalCase
Titles []string // titles of the top N hits, in rank order
Rank int // 1-based rank of the first wanted title, 0 if none
Err error
}
// Hit is true when a wanted title came back.
func (o Outcome) Hit() bool { return o.Rank > 0 }
// Report — the score plus per-case detail.
type Report struct {
Name string
Book string
TopN int
Scored int // answerable cases
Hits int
Errors int
Outcomes []Outcome
}
// Accuracy over the answerable cases.
func (r Report) Accuracy() float64 {
if r.Scored == 0 {
return 0
}
return float64(r.Hits) / float64(r.Scored)
}
// RunRetrievalEval searches for every fixture case.
func RunRetrievalEval(ctx context.Context, c *Client, topN int) (Report, error) {
var f fixture
if err := json.Unmarshal(knowledgeFixtureJSON, &f); err != nil {
return Report{}, err
}
rep := Report{Name: f.Name, Book: f.Book, TopN: topN}
for _, cs := range f.Cases {
res, err := c.Search(ctx, cs.Query, f.Book, topN)
o := Outcome{Case: cs, Err: err}
if err != nil {
rep.Errors++
}
for i, hit := range res {
o.Titles = append(o.Titles, hit.Title)
if o.Rank == 0 && matches(cs.WantTitles, hit.Title) {
o.Rank = i + 1
}
}
if !cs.ExpectMiss {
rep.Scored++
if o.Hit() {
rep.Hits++
}
}
rep.Outcomes = append(rep.Outcomes, o)
}
return rep, nil
}
func matches(want []string, title string) bool {
for _, w := range want {
if strings.EqualFold(strings.TrimSpace(title), w) {
return true
}
}
return false
}
// String — the headline number.
func (r Report) String() string {
var b strings.Builder
fmt.Fprintf(&b, "%s: %d/%d answerable questions retrieve a wanted article in top %d (%.1f%%), %d errors\n",
r.Name, r.Hits, r.Scored, r.TopN, 100*r.Accuracy(), r.Errors)
fmt.Fprintf(&b, " book: %s\n", r.Book)
return b.String()
}
// Detail — per case: what was asked, what was searched, what came back.
func (r Report) Detail() string {
var b strings.Builder
for _, o := range r.Outcomes {
mark := "MISS"
switch {
case o.Case.ExpectMiss:
mark = "n/a "
case o.Hit():
mark = fmt.Sprintf("hit@%d", o.Rank)
}
fmt.Fprintf(&b, " %-6s %-20s q=%q\n", mark, o.Case.ID, o.Case.Query)
if o.Err != nil {
fmt.Fprintf(&b, " error: %v\n", o.Err)
continue
}
fmt.Fprintf(&b, " got: %s\n", strings.Join(o.Titles, " | "))
}
return b.String()
}
@@ -1,4 +1,4 @@
package eval
package llm
import (
"context"
@@ -6,8 +6,20 @@ import (
"fmt"
"net/http"
"strings"
"time"
)
// UnknownModel is the label to print when the server would not say what it has
// loaded. Deliberately ugly: an honest "unknown" is fine, a plausible-looking
// but wrong model name is the bug this whole file exists to prevent.
const UnknownModel = "unknown-model"
// llama-server is local, so never send this through a proxy: this box's
// http_proxy answers 503 for loopback, which would look like "server won't say
// which model it has" when the server is right there and fine.
// A Transport with no Proxy set bypasses http_proxy entirely.
var modelHTTP = &http.Client{Timeout: 10 * time.Second, Transport: &http.Transport{}}
// ModelID asks llama-server which model it has loaded, so a scoring run can
// label itself. Without this a bake-off between two models produces two tables
// that look identical, and the operator has to remember which server was up.
@@ -19,7 +31,7 @@ func ModelID(ctx context.Context, base string) (string, error) {
if err != nil {
return "", err
}
resp, err := http.DefaultClient.Do(req)
resp, err := modelHTTP.Do(req)
if err != nil {
return "", err
}
@@ -38,14 +50,21 @@ func ModelID(ctx context.Context, base string) (string, error) {
if len(out.Data) == 0 {
return "", fmt.Errorf("models: empty list")
}
return shortModelID(out.Data[0].ID), nil
short := shortModelID(out.Data[0].ID)
if short == "" {
// Server answered but the id field was missing or blank. Say so
// instead of handing back an empty label that reads as a real name.
return "", fmt.Errorf("models: no id in response")
}
return short, nil
}
// shortModelID trims the path and the .gguf suffix — llama-server reports the
// file name it was started with, which is too long for a table header.
func shortModelID(id string) string {
id = strings.TrimSpace(id)
if i := strings.LastIndexAny(id, "/\\"); i >= 0 {
id = id[i+1:]
}
return strings.TrimSuffix(id, ".gguf")
return strings.TrimSpace(strings.TrimSuffix(id, ".gguf"))
}
+64
View File
@@ -0,0 +1,64 @@
package llm
import (
"context"
"net/http"
"net/http/httptest"
"testing"
)
// The point of these tests: a wrong-but-plausible model label is the bug, so
// every path that cannot learn the real name must return an error instead of a
// guess. No llama-server needed — a stub server stands in.
func TestModelID(t *testing.T) {
cases := []struct {
name string
body string
code int
want string // "" ⇒ expect an error
}{
{"full path", `{"data":[{"id":"/mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf"}]}`, 200, "Qwen3.5-0.8B.Q4_K_M"},
{"bare name", `{"data":[{"id":"LFM2.5-1.2B"}]}`, 200, "LFM2.5-1.2B"},
{"empty list", `{"data":[]}`, 200, ""},
{"id missing", `{"data":[{}]}`, 200, ""},
{"id blank", `{"data":[{"id":" "}]}`, 200, ""},
{"server error", `nope`, 500, ""},
{"not json", `<html>`, 200, ""},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path != "/v1/models" {
t.Errorf("asked for %s, want /v1/models", r.URL.Path)
}
w.WriteHeader(c.code)
_, _ = w.Write([]byte(c.body))
}))
defer srv.Close()
got, err := ModelID(context.Background(), srv.URL+"/")
if c.want == "" {
if err == nil {
t.Fatalf("want an error, got label %q", got)
}
return
}
if err != nil {
t.Fatalf("ModelID: %v", err)
}
if got != c.want {
t.Errorf("got %q, want %q", got, c.want)
}
})
}
}
func TestModelIDUnreachable(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(http.ResponseWriter, *http.Request) {}))
url := srv.URL
srv.Close() // nothing listening now
if got, err := ModelID(context.Background(), url); err == nil {
t.Fatalf("want an error from a dead server, got label %q", got)
}
}
+38
View File
@@ -0,0 +1,38 @@
package memory
// Confidence gate for a recall. Two checks, both must pass before Maven says a
// note back:
//
// - minScore — an absolute cosine floor.
// - minMargin — the top hit must beat the runner-up by more than this.
//
// The margin is the one that carries the weight. The e5 embedder packs every
// score into a narrow high band (0.79-0.89 on the recall fixture), so an
// absolute floor cannot tell a real hit from a confident-looking miss: every
// value under the band admits everything, every value above it answers nothing.
// A margin asks a different question — "is this note clearly the best one, or
// is the whole shelf equally close?" — and a made-up question has no clear best.
//
// With one hit and no runner-up there is nothing to compare, so only the floor
// applies.
// ConfidentScores reports whether the top score clears both gates. scores must
// be sorted highest first. minMargin <= 0 turns the margin check off.
func ConfidentScores(scores []float64, minScore, minMargin float64) bool {
if len(scores) == 0 || scores[0] < minScore {
return false
}
if minMargin > 0 && len(scores) > 1 && scores[0]-scores[1] <= minMargin {
return false
}
return true
}
// Confident is ConfidentScores for search results.
func Confident(results []Result, minScore, minMargin float64) bool {
scores := make([]float64, len(results))
for i, r := range results {
scores[i] = r.Score
}
return ConfidentScores(scores, minScore, minMargin)
}
+45
View File
@@ -0,0 +1,45 @@
package memory
import "testing"
func TestConfidentScores(t *testing.T) {
cases := []struct {
name string
scores []float64
minScore float64
minMargin float64
want bool
}{
{"no hits", nil, 0.55, 0.008, false},
{"below the floor", []float64{0.40, 0.10}, 0.55, 0.008, false},
{"clear winner", []float64{0.86, 0.70}, 0.55, 0.008, true},
{"runner-up too close", []float64{0.860, 0.858}, 0.55, 0.008, false},
// The rule is "beats the runner-up by MORE than delta". Not testing an
// exactly-equal margin: no pair of these decimals subtracts to exactly
// 0.008 in binary float, so such a test would pin rounding, not the rule.
{"margin just under delta", []float64{0.8079, 0.8}, 0.55, 0.008, false},
{"margin just over delta", []float64{0.8081, 0.8}, 0.55, 0.008, true},
// One hit: nothing to compare against, so only the floor applies.
{"single hit clears", []float64{0.86}, 0.55, 0.008, true},
{"single hit below floor", []float64{0.10}, 0.55, 0.008, false},
// Margin off — the old absolute-only behaviour.
{"margin off admits a tie", []float64{0.86, 0.86}, 0.55, 0, true},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
if got := ConfidentScores(c.scores, c.minScore, c.minMargin); got != c.want {
t.Errorf("got %v, want %v", got, c.want)
}
})
}
}
func TestConfidentReadsResultScores(t *testing.T) {
res := []Result{{ID: "a", Score: 0.86}, {ID: "b", Score: 0.858}}
if Confident(res, 0.55, 0.008) {
t.Error("thin margin passed the gate")
}
if !Confident(res, 0.55, 0) {
t.Error("margin off should fall back to the floor alone")
}
}
+44 -19
View File
@@ -177,6 +177,8 @@ type Outcome struct {
Tied bool
TopID string
TopScor float64
// Margin — top1 top2. 0 when fewer than two hits came back.
Margin float64
Reasons []string
}
@@ -184,8 +186,10 @@ type Outcome struct {
// ranks first but is silenced by query_min_score is a threshold problem, and a
// note that never ranks first is an embedder problem. Those are different fixes.
type Report struct {
Name string
MinScore float64
Name string
MinScore float64
// MinMargin — how far the top hit must beat the runner-up. 0 ⇒ off.
MinMargin float64
Total int
Answerable int
Rank1 int
@@ -212,9 +216,14 @@ type Report struct {
// where the right note ranked first, and for the no-answer cases. The gap
// between these two distributions is what a defensible query_min_score
// would have to sit inside; if they overlap, no threshold separates them.
CorrectTop []float64
NoAnswerTop []float64
P50, P95, Max time.Duration
CorrectTop []float64
NoAnswerTop []float64
// CorrectMargin / NoAnswerMargin — the same two groups, but top1 top2
// instead of top1. This is the pair the margin gate has to separate, and
// unlike the absolute scores it is what the sweep reads.
CorrectMargin []float64
NoAnswerMargin []float64
P50, P95, Max time.Duration
}
// TagStat — passed/total for one slice of the fixture.
@@ -245,13 +254,14 @@ func ratio(n, d int) float64 {
// the run on an embed or search error: an erroring case scores as a miss and is
// counted in Errors, because "the embedder was down" and "the embedder was
// wrong" are different numbers.
func Score(ctx context.Context, name string, emb router.Embedder, newStore NewStore, minScore float64, f Fixture) (Report, error) {
func Score(ctx context.Context, name string, emb router.Embedder, newStore NewStore, minScore, minMargin float64, f Fixture) (Report, error) {
rep := Report{
Name: name,
MinScore: minScore,
Total: len(f.Cases),
ByTag: map[string]TagStat{},
ByLang: map[string]TagStat{},
Name: name,
MinScore: minScore,
MinMargin: minMargin,
Total: len(f.Cases),
ByTag: map[string]TagStat{},
ByLang: map[string]TagStat{},
}
lat := make([]time.Duration, 0, len(f.Cases))
@@ -261,7 +271,7 @@ func Score(ctx context.Context, name string, emb router.Embedder, newStore NewSt
} else {
rep.NoAnswer++
}
o, err := scoreCase(ctx, emb, newStore, minScore, c, f.Filler)
o, err := scoreCase(ctx, emb, newStore, minScore, minMargin, c, f.Filler)
if err != nil {
return Report{}, err
}
@@ -285,6 +295,11 @@ func Score(ctx context.Context, name string, emb router.Embedder, newStore NewSt
} else if !o.Rank1 {
rep.WrongTop++
}
if o.Rank1 {
// Margins are collected on rank, not on the gate, so the
// distribution does not move as the sweep changes the gate.
rep.CorrectMargin = append(rep.CorrectMargin, o.Margin)
}
if o.Rank1 && o.Recalled != "" {
rep.CorrectTop = append(rep.CorrectTop, o.TopScor)
}
@@ -293,6 +308,7 @@ func Score(ctx context.Context, name string, emb router.Embedder, newStore NewSt
rep.FalseRecall++
}
rep.NoAnswerTop = append(rep.NoAnswerTop, o.TopScor)
rep.NoAnswerMargin = append(rep.NoAnswerMargin, o.Margin)
}
if o.Pass {
@@ -307,6 +323,8 @@ func Score(ctx context.Context, name string, emb router.Embedder, newStore NewSt
sort.Float64s(rep.CorrectTop)
sort.Float64s(rep.NoAnswerTop)
sort.Float64s(rep.CorrectMargin)
sort.Float64s(rep.NoAnswerMargin)
sort.Slice(lat, func(i, j int) bool { return lat[i] < lat[j] })
rep.P50, rep.P95 = percentile(lat, 0.50), percentile(lat, 0.95)
if len(lat) > 0 {
@@ -318,7 +336,7 @@ func Score(ctx context.Context, name string, emb router.Embedder, newStore NewSt
// scoreCase inserts the case's notes into a fresh store, then runs the read
// path the daemon runs. The returned error is fatal (the harness is broken);
// an embedder or store failure on the query lands in Outcome.Err instead.
func scoreCase(ctx context.Context, emb router.Embedder, newStore NewStore, minScore float64, c Case, filler []StoredNote) (Outcome, error) {
func scoreCase(ctx context.Context, emb router.Embedder, newStore NewStore, minScore, minMargin float64, c Case, filler []StoredNote) (Outcome, error) {
st, release, err := newStore()
if err != nil {
return Outcome{}, fmt.Errorf("%s: new store: %w", c.ID, err)
@@ -357,7 +375,10 @@ func scoreCase(ctx context.Context, emb router.Embedder, newStore NewStore, minS
if len(hits) > 0 {
o.TopID, o.TopScor = hits[0].ID, hits[0].Score
o.Recalled = bestRecall(hits, minScore)
if len(hits) > 1 {
o.Margin = hits[0].Score - hits[1].Score
}
o.Recalled = bestRecall(hits, minScore, minMargin)
}
for i, h := range hits {
if h.ID != c.Want {
@@ -378,14 +399,14 @@ func scoreCase(ctx context.Context, emb router.Embedder, newStore NewStore, minS
switch {
case !c.Answerable():
if o.Recalled != "" {
o.Reasons = append(o.Reasons, fmt.Sprintf("false recall: %q at %.3f, want silence", o.TopID, o.TopScor))
o.Reasons = append(o.Reasons, fmt.Sprintf("false recall: %q at %.3f (margin %.3f), want silence", o.TopID, o.TopScor, o.Margin))
}
case o.Tied:
o.Reasons = append(o.Reasons, fmt.Sprintf("tie at %.3f — the right note is on top only by sort order", o.TopScor))
case !o.Rank1:
o.Reasons = append(o.Reasons, fmt.Sprintf("top hit %q (%.3f), want %q%s", o.TopID, o.TopScor, c.Want, rankNote(o.Rank3)))
case o.Recalled == "":
o.Reasons = append(o.Reasons, fmt.Sprintf("right note ranked first but scored %.3f < gate %.2f — daemon says \"не знаю\"", o.TopScor, minScore))
o.Reasons = append(o.Reasons, fmt.Sprintf("right note ranked first at %.3f (margin %.3f) but the gate silenced it — daemon says \"не знаю\"", o.TopScor, o.Margin))
}
o.Pass = len(o.Reasons) == 0
return o, nil
@@ -401,8 +422,10 @@ func rankNote(inTop3 bool) string {
// bestRecall mirrors cmd/mavend/recall.go — the gate the daemon actually
// applies to a memory hit. Duplicated rather than imported because package main
// is not importable; recalleval_test.go asserts the two agree in behaviour.
func bestRecall(results []memory.Result, min float64) string {
if len(results) == 0 || results[0].Score < min {
// The daemon returns the whole hit (a note and a fact are said differently);
// the harness only scores what came back, so it keeps returning the text.
func bestRecall(results []memory.Result, minScore, minMargin float64) string {
if !memory.Confident(results, minScore, minMargin) {
return ""
}
return results[0].Meta["text"]
@@ -437,7 +460,7 @@ func percentile(sorted []time.Duration, p float64) time.Duration {
// the slices that name where the path is weak.
func (r Report) String() string {
var b strings.Builder
fmt.Fprintf(&b, "%s: %d/%d cases pass (gate %.2f)\n", r.Name, r.Passed, r.Total, r.MinScore)
fmt.Fprintf(&b, "%s: %d/%d cases pass (gate %.2f, margin %.3f)\n", r.Name, r.Passed, r.Total, r.MinScore, r.MinMargin)
fmt.Fprintf(&b, " recall@1 %.1f%% (%d/%d) recall@3 %.1f%% (%d/%d) answered after gate %.1f%% (%d/%d)\n",
100*r.Recall1(), r.Rank1, r.Answerable,
100*r.Recall3(), r.Rank3, r.Answerable,
@@ -448,6 +471,8 @@ func (r Report) String() string {
100*r.FalseRecallRate(), r.FalseRecall, r.NoAnswer)
fmt.Fprintf(&b, " top-1 score, right note first: %s\n", spread(r.CorrectTop))
fmt.Fprintf(&b, " top-1 score, must be silent: %s\n", spread(r.NoAnswerTop))
fmt.Fprintf(&b, " margin top1-top2, right note first: %s\n", spread(r.CorrectMargin))
fmt.Fprintf(&b, " margin top1-top2, must be silent: %s\n", spread(r.NoAnswerMargin))
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
fmt.Fprintf(&b, " by lang: %s\n", renderStats(r.ByLang))
fmt.Fprintf(&b, " by tag: %s\n", renderStats(r.ByTag))
+52 -12
View File
@@ -142,21 +142,40 @@ func words(s string) []string {
// cmd/mavend/recall.go (package main is not importable). This pins the copy to
// the original's three rules: no hits, below the gate, or no text ⇒ silence.
func TestBestRecallMatchesDaemon(t *testing.T) {
if got := bestRecall(nil, 0.55); got != "" {
if got := bestRecall(nil, 0.55, 0); got != "" {
t.Errorf("no hits: got %q, want silence", got)
}
low := []memory.Result{{ID: "a", Score: 0.4, Meta: map[string]string{"text": "чай"}}}
if got := bestRecall(low, 0.55); got != "" {
if got := bestRecall(low, 0.55, 0); got != "" {
t.Errorf("below gate: got %q, want silence", got)
}
noText := []memory.Result{{ID: "a", Score: 0.9, Meta: map[string]string{}}}
if got := bestRecall(noText, 0.55); got != "" {
if got := bestRecall(noText, 0.55, 0); got != "" {
t.Errorf("no text: got %q, want silence", got)
}
ok := []memory.Result{{ID: "a", Score: 0.9, Meta: map[string]string{"text": "чай"}}}
if got := bestRecall(ok, 0.55); got != "чай" {
if got := bestRecall(ok, 0.55, 0); got != "чай" {
t.Errorf("above gate: got %q, want %q", got, "чай")
}
// Margin: a close runner-up means the embedder cannot tell the two apart,
// so Maven stays silent even though both clear the absolute floor.
close := []memory.Result{
{ID: "a", Score: 0.86, Meta: map[string]string{"text": "чай"}},
{ID: "b", Score: 0.85, Meta: map[string]string{"text": "кофе"}},
}
if got := bestRecall(close, 0.55, 0.03); got != "" {
t.Errorf("thin margin: got %q, want silence", got)
}
if got := bestRecall(close, 0.55, 0); got != "чай" {
t.Errorf("margin off: got %q, want %q", got, "чай")
}
clear := []memory.Result{
{ID: "a", Score: 0.86, Meta: map[string]string{"text": "чай"}},
{ID: "b", Score: 0.70, Meta: map[string]string{"text": "кофе"}},
}
if got := bestRecall(clear, 0.55, 0.03); got != "чай" {
t.Errorf("wide margin: got %q, want %q", got, "чай")
}
}
// TestHashRecallBaseline — the CI ratchet. HashEmbedder, so it needs no model
@@ -171,14 +190,15 @@ func TestHashRecallBaseline(t *testing.T) {
t.Fatalf("Load: %v", err)
}
rep, err := Score(context.Background(), "recall+hash", router.NewHashEmbedder(hashDim), InMemory,
config.DefaultQueryMinScore, f)
config.DefaultQueryMinScore, config.DefaultQueryMinMargin, f)
if err != nil {
t.Fatalf("Score: %v", err)
}
t.Log("\n" + rep.String() + rep.Failures())
t.Log("\ngate sweep:\n" + sweep(t, router.NewHashEmbedder(hashDim), f))
// 0.32 sits under the observed 0.360 recall@1.
// 0.32 sits under the observed 0.370 recall@1 (was 0.360 over 25 answerable
// cases; the two mixed note+fact cases added with #373 make it 27).
const floorRecall1 = 0.32
if rep.Recall1() < floorRecall1 {
t.Errorf("recall@1 %.3f below ratchet %.2f — note recall regressed", rep.Recall1(), floorRecall1)
@@ -201,11 +221,11 @@ func TestPersistentStoreScoresTheSame(t *testing.T) {
t.Fatalf("Load: %v", err)
}
emb := router.NewHashEmbedder(hashDim)
inMem, err := Score(context.Background(), "recall+hash+memory", emb, InMemory, config.DefaultQueryMinScore, f)
inMem, err := Score(context.Background(), "recall+hash+memory", emb, InMemory, config.DefaultQueryMinScore, config.DefaultQueryMinMargin, f)
if err != nil {
t.Fatalf("Score in-memory: %v", err)
}
persistent, err := Score(context.Background(), "recall+hash+sqlite", emb, sqliteStores(t), config.DefaultQueryMinScore, f)
persistent, err := Score(context.Background(), "recall+hash+sqlite", emb, sqliteStores(t), config.DefaultQueryMinScore, config.DefaultQueryMinMargin, f)
if err != nil {
t.Fatalf("Score sqlite: %v", err)
}
@@ -263,14 +283,16 @@ func TestONNXRecall(t *testing.T) {
if err != nil {
t.Fatalf("Load: %v", err)
}
rep, err := Score(context.Background(), "recall+onnx", emb, InMemory, config.DefaultQueryMinScore, f)
rep, err := Score(context.Background(), "recall+onnx", emb, InMemory, config.DefaultQueryMinScore, config.DefaultQueryMinMargin, f)
if err != nil {
t.Fatalf("Score: %v", err)
}
t.Log("\n" + rep.String() + rep.Failures())
// Cached for the sweep only: the headline run above must pay the real
// Cached for the sweeps only: the headline run above must pay the real
// embedder cost so its latency numbers mean something.
t.Log("\ngate sweep:\n" + sweep(t, Cache(emb), f))
cached := Cache(emb)
t.Log("\ngate sweep (margin off):\n" + sweep(t, cached, f))
t.Log("\nmargin sweep (gate 0.55):\n" + marginSweep(t, cached, f))
}
// sweep scores the fixture at a range of gates and renders one line each. Two
@@ -281,7 +303,7 @@ func sweep(t *testing.T, emb router.Embedder, f Fixture) string {
t.Helper()
var b strings.Builder
for _, gate := range []float64{0.0, 0.30, 0.40, 0.50, 0.55, 0.60, 0.70, 0.80, 0.90} {
rep, err := Score(context.Background(), "sweep", emb, InMemory, gate, f)
rep, err := Score(context.Background(), "sweep", emb, InMemory, gate, 0, f)
if err != nil {
t.Fatalf("sweep at %.2f: %v", gate, err)
}
@@ -290,3 +312,21 @@ func sweep(t *testing.T, emb router.Embedder, f Fixture) string {
}
return b.String()
}
// marginSweep is the same idea for the margin gate (top1 top2 > delta), with
// the absolute gate held at its default. The absolute score cannot separate a
// real hit from a made-up question under e5 — every score lands in one narrow
// band — so this sweep is the one that picks a number.
func marginSweep(t *testing.T, emb router.Embedder, f Fixture) string {
t.Helper()
var b strings.Builder
for _, d := range []float64{0, 0.002, 0.005, 0.008, 0.01, 0.012, 0.015, 0.02, 0.025, 0.03, 0.04, 0.05, 0.06} {
rep, err := Score(context.Background(), "margin sweep", emb, InMemory, config.DefaultQueryMinScore, d, f)
if err != nil {
t.Fatalf("margin sweep at %.3f: %v", d, err)
}
fmt.Fprintf(&b, " delta %.3f: answered %d/%d (%.0f%%) false recall %d/%d\n",
d, rep.Rank1-rep.Gated, rep.Answerable, 100*rep.Answered(), rep.FalseRecall, rep.NoAnswer)
}
return b.String()
}
@@ -387,6 +387,32 @@
{"id": "n2", "text": "wifi channel is 6", "kind": "note"},
{"id": "n3", "text": "the guest network is off", "kind": "note"}
]
},
{
"id": "ru-mixed-031",
"lang": "ru",
"tags": ["mixed", "paraphrase", "hard"],
"note": "notes and facts in one store and the note is the answer — the daemon indexes both (Vikunja #373)",
"query": "куда я спрятал второй ключ от квартиры",
"want": "n1",
"notes": [
{"id": "n1", "text": "запасной ключ от квартиры лежит в синей коробке на полке", "kind": "note"},
{"id": "x1", "text": "поменял замок в двери двадцатого июня", "kind": "fact"},
{"id": "x2", "text": "отдал ключ соседке в мае", "kind": "fact"}
]
},
{
"id": "ru-mixed-032",
"lang": "ru",
"tags": ["mixed", "distractor"],
"note": "the mirror of ru-mixed-031: the fact answers and the notes are the distractors",
"query": "когда я в последний раз заливал бензин",
"want": "x1",
"notes": [
{"id": "x1", "text": "залил полный бак в четверг вечером", "kind": "fact"},
{"id": "n1", "text": "на заправке у моста дешевле бензин", "kind": "note"},
{"id": "n2", "text": "надо поменять зимние шины", "kind": "note"}
]
}
]
}
+138
View File
@@ -0,0 +1,138 @@
// Package persona builds the one shared context block that goes in front of
// every LLM system prompt: who the owner is, how to address him, and what
// time it is right now.
//
// Why one block and not a line pasted into each prompt: there are five
// prompts (nudges, action replies, chat, note queries, general knowledge) and
// the "address him as ты" rule had only reached two of them. Five copies drift.
// One block cannot.
//
// The rules here are defaults in code, not config. Maven is feminine and the
// owner is a man addressed informally — that is a hard constraint of the
// product, so it must hold with an empty config file. Config only ADDS
// optional facts (his name, his city).
package persona
import (
"fmt"
"strings"
"time"
)
// Facts — the optional, deployment-specific half of the block. All fields may
// be empty; the block is still correct and useful without them.
type Facts struct {
OwnerName string // his name, e.g. "Ками"
City string // where he is, e.g. "Москва"
Static string // the free-text `persona` config string, appended verbatim
// The two config-gated capabilities. They are listed only when this
// deployment actually has them, because a capability she names and cannot
// do is worse than one she never mentions.
Weather bool // an open-meteo provider is configured
Telegram bool // a telegram bot token + chat id are configured
Tools bool // at least one shell act is on the allowlist
}
var ruWeekdays = [...]string{"воскресенье", "понедельник", "вторник", "среда", "четверг", "пятница", "суббота"}
var ruMonths = [...]string{
"января", "февраля", "марта", "апреля", "мая", "июня",
"июля", "августа", "сентября", "октября", "ноября", "декабря",
}
// Block renders the context block for one turn. Russian even in front of the
// English prompts: the rules it states are Russian grammar (ты/тебя, feminine
// verbs), and a Russian rule reads best stated in Russian.
//
// Keep it short. It ships on every turn to a 0.8B on laptop CPU, so every
// line here is latency.
func (f Facts) Block(now time.Time) string {
var b strings.Builder
b.WriteString("Ты — Maven, домашняя ассистентка. О себе говоришь в женском роде: \"я записала\", \"я проверила\".\n")
// The address form gets its own line. It is the thing that kept getting
// lost when it was buried in prose.
b.WriteString("ОБРАЩЕНИЕ: владелец — мужчина, всегда на \"ты\" (ты, тебя, тебе, твой) и в единственном числе (\"выпей\", \"посмотри\"). Никогда \"вы\"/\"вас\"/\"ваш\". Никогда \"он\"/\"его\" о нём — ты говоришь ему, а не о нём. Глаголы о нём — в мужском роде (\"ты забыл\").\n")
if who := f.who(); who != "" {
b.WriteString(who + "\n")
}
b.WriteString(fmt.Sprintf("Сейчас: %s, %d %s %d, %02d:%02d (местное время).\n",
ruWeekdays[int(now.Weekday())], now.Day(), ruMonths[int(now.Month())-1], now.Year(),
now.Hour(), now.Minute()))
b.WriteString("Умеешь: " + strings.Join(f.can(), "; ") +
". Других ДЕЙСТВИЙ не умеешь — если просят такое, скажи прямо.\n")
if s := strings.TrimSpace(f.Static); s != "" {
b.WriteString(s + "\n")
}
return b.String()
}
// can lists what she can really do. Every entry here is a code path that
// exists in the daemon today:
// - reminders: IntentReminder → CoreAPI.CreateReminder, fired by the tick.
// - notes and facts: IntentNote/IntentFact write, IntentQuery reads them back.
// - calendar: IntentQuery answers "что у меня сегодня" from CalendarEvents.
// - weather / telegram / shell acts: only when configured (see Facts).
//
// Nothing speculative goes in this list. A capability she offers and cannot
// perform is worse than one she never mentions.
func (f Facts) can() []string {
c := []string{
// Talking comes first, and the closing line says "действий" rather than
// "ничего", because this same block sits in front of the chat and
// general-knowledge prompts. A flat "you can do nothing else" would
// tell her to refuse the exact thing those two prompts are for.
"разговаривать и отвечать на вопросы",
"ставить напоминания",
"записывать заметки и факты и отвечать по ним",
"смотреть календарь",
}
if f.Weather {
c = append(c, "говорить погоду")
}
if f.Telegram {
c = append(c, "писать в телеграм")
}
if f.Tools {
c = append(c, "запускать разрешённые команды на сервере")
}
return c
}
// who renders the optional name/city line, or "" when neither is configured.
//
// Written as labels ("Имя владельца: ..."), not as a sentence with pronouns:
// the block's own "ты" is Maven, so "тебя зовут" would read as her name and
// "его" would model the third-person form she must never use about him.
func (f Facts) who() string {
name := strings.TrimSpace(f.OwnerName)
city := strings.TrimSpace(f.City)
switch {
case name != "" && city != "":
return "Имя владельца: " + name + ". Город: " + city + "."
case name != "":
return "Имя владельца: " + name + "."
case city != "":
return "Город: " + city + "."
}
return ""
}
// Prepend puts the block in front of a system prompt. Nil-safe: a nil renderer
// (tests, the stub paths) returns the prompt untouched.
func Prepend(block func() string, prompt string) string {
if block == nil {
return prompt
}
s := strings.TrimSpace(block())
if s == "" {
return prompt
}
return s + "\n\n" + prompt
}
+72
View File
@@ -0,0 +1,72 @@
package persona
import (
"strings"
"testing"
"time"
)
var ref = time.Date(2026, 7, 31, 14, 5, 0, 0, time.UTC)
// The block must be correct with an empty config: the address form and the
// gender rules are hard constraints, not preferences.
func TestBlockWorksWithZeroConfig(t *testing.T) {
b := Facts{}.Block(ref)
for _, want := range []string{"женском роде", "ОБРАЩЕНИЕ", "\"ты\"", "31 июля 2026", "пятница", "14:05"} {
if !strings.Contains(b, want) {
t.Errorf("block missing %q:\n%s", want, b)
}
}
}
func TestBlockAddsOptionalFacts(t *testing.T) {
b := Facts{OwnerName: "Ками", City: "Москва", Static: "Будь краткой."}.Block(ref)
for _, want := range []string{"Ками", "Москва", "Будь краткой."} {
if !strings.Contains(b, want) {
t.Errorf("block missing %q:\n%s", want, b)
}
}
}
// The time changes between turns, so two renders must differ.
func TestBlockRendersTimePerTurn(t *testing.T) {
a := Facts{}.Block(ref)
c := Facts{}.Block(ref.Add(time.Hour))
if a == c {
t.Errorf("block did not change with the clock:\n%s", a)
}
}
// She may only offer what this deployment actually has.
func TestCapabilitiesAreConfigGated(t *testing.T) {
bare := Facts{}.Block(ref)
for _, want := range []string{"напоминания", "заметки", "календарь"} {
if !strings.Contains(bare, want) {
t.Errorf("block missing always-on capability %q:\n%s", want, bare)
}
}
for _, unwanted := range []string{"погоду", "телеграм", "команды"} {
if strings.Contains(bare, unwanted) {
t.Errorf("block offers unconfigured %q:\n%s", unwanted, bare)
}
}
full := Facts{Weather: true, Telegram: true, Tools: true}.Block(ref)
for _, want := range []string{"погоду", "телеграм", "команды"} {
if !strings.Contains(full, want) {
t.Errorf("block missing configured capability %q:\n%s", want, full)
}
}
}
func TestPrependNilIsSafe(t *testing.T) {
if got := Prepend(nil, "PROMPT"); got != "PROMPT" {
t.Errorf("Prepend(nil) = %q", got)
}
if got := Prepend(func() string { return " " }, "PROMPT"); got != "PROMPT" {
t.Errorf("Prepend(blank) = %q", got)
}
if got := Prepend(func() string { return "CTX" }, "PROMPT"); got != "CTX\n\nPROMPT" {
t.Errorf("Prepend = %q", got)
}
}
+33
View File
@@ -0,0 +1,33 @@
package phraser
import (
"strings"
"testing"
)
// Every phrasing prompt must carry the shared context block. This is the
// regression guard for the bug that started this: the "ты" rule reached only
// two of the five prompts because each prompt had its own copy of the rules.
func TestEveryPromptCarriesTheContextBlock(t *testing.T) {
block := func() string { return "CTXBLOCK" }
p := &LLMPhraser{cfg: Config{ContextBlock: block}}
prompts := map[string]string{
"nudge": p.systemPrompt(),
"query": p.querySystemPrompt(),
"chat": chatSystemPrompt(block),
}
for name, got := range prompts {
if !strings.HasPrefix(got, "CTXBLOCK\n\n") {
t.Errorf("%s prompt does not start with the context block:\n%s", name, got)
}
}
}
// Without a block the prompts are unchanged — the stub and test paths pass nil.
func TestPromptsWithoutBlockAreUnchanged(t *testing.T) {
p := &LLMPhraser{}
if p.systemPrompt() != nudgeSystem {
t.Errorf("nudge prompt changed with no block set")
}
}
@@ -0,0 +1,33 @@
package eval
import (
"strings"
"testing"
)
// TestAddressReportsEveryBreak — the real reply from a nudge eval run broke in
// two ways at once and the check named only the plural. Both must print: a
// half-reported failure reads as a milder problem than it is.
func TestAddressReportsEveryBreak(t *testing.T) {
body := "Смотрите на его потребление воды."
res := checkAddress(body)
if res.Pass {
t.Fatalf("checkAddress passed %q", body)
}
for _, want := range []string{"смотрите", "его"} {
if !strings.Contains(res.Detail, want) {
t.Errorf("detail %q does not name %q", res.Detail, want)
}
}
}
// One word repeated is one problem, so the detail must not say it twice.
func TestAddressDeduplicates(t *testing.T) {
res := checkAddress("Вам стоит поесть, вам это нужно.")
if res.Pass {
t.Fatal("expected failure")
}
if n := strings.Count(res.Detail, "formal"); n != 1 {
t.Errorf("detail repeats the same break %d times: %q", n, res.Detail)
}
}
@@ -0,0 +1,42 @@
package eval
import "testing"
func TestAddressTimeWordDoesNotBlind(t *testing.T) {
// A nudge that opens with a time word must still be caught. Without the
// time words in the stoplist, "сегодня" was read as the third party.
for _, s := range []string{
"сегодня он не ел 11 дней",
"вчера он не пил воду",
"опять он забыл про таблетки",
} {
if r := checkAddress(s); r.Pass {
t.Errorf("checkAddress(%q) passed, want a third-person failure", s)
}
}
// Still must not fire when a third party really is named.
for _, s := range []string{
"сегодня сервис упал, он не отвечает",
"ты не пил воду четыре часа",
} {
if r := checkAddress(s); !r.Pass {
t.Errorf("checkAddress(%q) failed: %s", s, r.Detail)
}
}
}
// TestAddressVerbIsNotAnAntecedent — a nudge is mostly verbs, and a verb is
// never who "он" refers to. This exact string passed the check before.
func TestAddressVerbIsNotAnAntecedent(t *testing.T) {
s := "попробуй встать и отдохнуть — у него есть перерыв"
if r := checkAddress(s); r.Pass {
t.Errorf("checkAddress(%q) passed, want a third-person failure", s)
}
// Still missed, and this is the documented hole: "выпей воды, он не пил" has
// a real noun ("воды") before the pronoun, so the scan believes somebody
// else was named. Telling that apart needs a parser, not a suffix rule.
// A named third party still wins over the verbs around it.
if r := checkAddress("сервис упал, он не отвечает"); !r.Pass {
t.Errorf("checkAddress on a real third party failed: %s", r.Detail)
}
}
+357 -3
View File
@@ -17,10 +17,18 @@ const (
CheckFeminine = "feminine" // her self-reference is feminine (hard constraint)
CheckCringe = "cringe" // DESIGN.md § Non-goals, "not a relationship"
CheckOnTopic = "ontopic" // says the thing the rule is about
// CheckHisGender — the other half of the persona rule: SHE is feminine, HE
// is male. "ты давно не отдыхала" addresses the operator as a woman.
CheckHisGender = "hisgender"
// CheckAddress — she talks TO him, informally, one to one. Not "вы", not
// "он". See the comment block above checkAddress.
CheckAddress = "address"
)
// CheckNames — report order.
var CheckNames = []string{CheckMood, CheckLang, CheckLength, CheckFeminine, CheckCringe, CheckOnTopic}
var CheckNames = []string{CheckMood, CheckLang, CheckLength, CheckFeminine, CheckHisGender, CheckAddress, CheckCringe, CheckOnTopic}
// Result — one check on one message.
type Result struct {
@@ -52,6 +60,8 @@ func RunChecks(c Case, body, mood string) []Result {
checkLang(body),
checkLength(body),
checkFeminine(body),
checkHisGender(body),
checkAddress(body),
checkCringe(body),
checkOnTopic(c, body),
}
@@ -176,6 +186,315 @@ func checkFeminine(body string) Result {
return Result{CheckFeminine, true, ""}
}
// --- he is male ----------------------------------------------------------
//
// The mirror of checkFeminine, and the failure it was written for: the model
// wrote "ты давно не отдыхала", which addresses the operator as a woman. That
// scored clean, because checkFeminine only ever looks at how SHE speaks about
// herself.
//
// How it works: Russian past tense is gendered by suffix, -л (m) / -ла (f). So
// this looks for feminine past-tense words in a sentence that also talks TO him
// ("ты", "тебя", "тебе", "твой", …). A feminine verb that belongs to her ("я
// заметила", "напомнила тебе") is skipped — that one is correct.
//
// Honest about the limits: this is a suffix rule, not a parser.
// - False positives: a feminine noun can be the subject in the same sentence
// ("зарядка была утром, ты её пропустил"). The guard below skips a verb whose
// previous word looks like a feminine noun, which helps but will not always
// be right.
// - False negatives: gender also shows up outside the past tense (short
// adjectives, "сама"), and none of that is checked here.
//
// That is acceptable for an eval check. It is a signal to read the message, not
// a grammar verdict, and every hit prints the word it tripped on so a human can
// disagree.
// hisMarkers — words that mean the sentence is addressed to him.
var hisMarkers = map[string]bool{
"ты": true, "тебя": true, "тебе": true, "тобой": true, "тобою": true,
"твой": true, "твоя": true, "твоё": true, "твое": true, "твои": true, "твою": true,
}
// notFeminineVerb — ordinary words ending in "-ла" that are not verbs. Small on
// purpose: it only has to cover words a nudge might actually use.
//
// Words that are both a noun and a verb are deliberately NOT here. "села",
// "мыла" and "стекла" are nouns on paper, but in a nudge they are almost always
// verbs ("ты села", "ты мыла"), and listing them would make the check miss the
// exact thing it is for. Missing a real hit is worse than one false alarm.
var notFeminineVerb = map[string]bool{
"школа": true, "скала": true, "игла": true, "метла": true, "смола": true,
"дела": true, "тела": true, "масла": true, "весла": true,
"зола": true, "пчела": true, "числа": true,
}
// femininePast reports whether a word looks like a feminine past-tense verb:
// "отдыхала", "поела", "выспалась".
func femininePast(w string) bool {
if len([]rune(w)) < 3 || notFeminineVerb[w] {
return false
}
return strings.HasSuffix(w, "ла") || strings.HasSuffix(w, "лась")
}
// looksFeminineNoun — a crude guard against "зарядка была": a word right before
// the verb that ends in "а"/"я" and is not itself a verb is probably the subject.
func looksFeminineNoun(w string) bool {
if femininePast(w) || len([]rune(w)) < 3 {
return false
}
return strings.HasSuffix(w, "а") || strings.HasSuffix(w, "я")
}
// sentenceRE splits on sentence-ending punctuation, so a feminine verb in one
// sentence is not blamed on a "ты" in the next.
var sentenceRE = regexp.MustCompile(`[.!?;…]+`)
func checkHisGender(body string) Result {
for _, sentence := range sentenceRE.Split(strings.ToLower(body), -1) {
words := wordRE.FindAllString(sentence, -1)
addressed := false
for _, w := range words {
if hisMarkers[w] {
addressed = true
}
}
if !addressed {
continue
}
for i, w := range words {
if !femininePast(w) || hersNotHis(words, i) {
continue
}
if i > 0 && looksFeminineNoun(prevWord(words, i)) {
continue
}
return Result{CheckHisGender, false,
fmt.Sprintf("feminine %q addressed to him — he is male", w)}
}
}
return Result{CheckHisGender, true, ""}
}
// hersNotHis — the verb is Maven's own if "я" comes shortly before it, or if the
// thing she did was done to him ("напомнила тебе", "проверила за тебя").
func hersNotHis(words []string, i int) bool {
for j := i - 1; j >= 0 && j >= i-3; j-- {
if words[j] == "я" {
return true
}
}
if i+1 < len(words) {
switch words[i+1] {
case "тебе", "тебя", "за", "тобой":
return true
}
}
return false
}
// prevWord — the word before i, skipping "не" and punctuation, so "не отдыхала"
// still sees the subject.
func prevWord(words []string, i int) string {
for j := i - 1; j >= 0; j-- {
w := words[j]
if w == "не" || w == "ни" || !unicode.Is(unicode.Cyrillic, []rune(w)[0]) {
continue
}
return w
}
return ""
}
// --- how she addresses him ------------------------------------------------
//
// Persona hard constraint: Maven speaks TO him, informally, one to one. The
// phrasing eval produced two breaks of it, and both scored clean:
//
// - "Приходите… Жду вас" — the formal plural. Correct is ты/тебя/тебе and a
// singular imperative ("приходи", "жду тебя").
// - "Он не ел 11 дней" — she talks ABOUT him, in the third person, as if
// reporting to somebody else. Correct is "ты не ел 11 дней".
//
// Like checkHisGender this is a keyword + suffix heuristic, NOT a parser. Every
// hit prints the word it tripped on, so a false alarm is obvious at a glance and
// can be dismissed.
//
// Part 1, formal address. Two signals:
// - the "вы" pronoun family, matched as whole words, so there is nothing to
// exclude — "вы" and "вас" are never anything else.
// - a plural verb ending: -ите/-ете/-йте/-ьте ("приходите", "выпейте",
// "не забудьте", "хотите"). Nouns in the prepositional case share those
// endings ("в интернете", "в свете"), so a word right after a preposition is
// skipped. That is the whole exclusion list, on purpose: a bigger one would
// start swallowing real imperatives.
//
// Part 2, third person. "он" is perfectly fine when the message really is about
// somebody or something else ("сервис упал, он не отвечает"). The way to tell
// them apart: a legitimate third person has an ANTECEDENT — the thing it refers
// to was named earlier in the message. So "он" is only flagged when nothing
// before it in the message could be that thing.
//
// Where this gives up, plainly:
// - it only looks BACKWARD. "Он не отвечает, сервис упал" names the subject
// after the pronoun and is flagged wrongly.
// - any noun earlier in the message counts as an antecedent, even when it is
// not one ("после обеда он не ел", "выпей воды, он не пил" — both missed).
// Verbs and time words no longer count, which covers the usual nudge, but a
// plain noun before the pronoun still blinds it. The
// common time words are stoplisted so the usual nudge opening does not
// blind it, but a message with any other noun in front still slips through.
// This is the check's real hole; widening it further would start flagging
// legitimate third-party messages, so it stops here.
// - a message that opens with "ты" and only later slips into "он" is missed,
// because "ты" itself is skipped but the words around it are not.
// - formal address outside these endings (short adjectives, "вашими" style
// forms not listed) is missed.
// addressWordRE also takes Latin words, because "him"/"he" is the same break in
// English.
var addressWordRE = regexp.MustCompile(`[\p{Cyrillic}]+|[a-zA-Z]+|[,.;:!?…—-]`)
// formalPronouns — the "вы" family. Whole-word match, so no false hits.
var formalPronouns = map[string]bool{
"вы": true, "вас": true, "вам": true, "вами": true,
"ваш": true, "ваша": true, "ваше": true, "ваши": true,
"вашего": true, "вашей": true, "вашему": true, "вашим": true,
"вашими": true, "вашу": true,
}
// prepositions — used twice: to skip prepositional-case nouns that look like
// plural verbs, and as words that cannot be what "он" refers to.
var prepositions = map[string]bool{
"в": true, "во": true, "на": true, "о": true, "об": true, "обо": true,
"при": true, "по": true, "за": true, "из": true, "с": true, "со": true,
"к": true, "ко": true, "до": true, "от": true, "у": true, "над": true,
"под": true, "про": true, "без": true, "для": true, "через": true,
}
// pluralVerb reports whether a word looks like a plural/formal verb form:
// "приходите", "выпейте", "забудьте", "хотите".
func pluralVerb(w string) bool {
if len([]rune(w)) < 5 {
return false
}
return strings.HasSuffix(w, "ите") || strings.HasSuffix(w, "ете") ||
strings.HasSuffix(w, "йте") || strings.HasSuffix(w, "ьте")
}
// thirdPersonHim — pronouns that would be talking about him instead of to him.
var thirdPersonHim = map[string]bool{
"он": true, "его": true, "ему": true, "него": true, "нему": true, "ним": true,
"he": true, "him": true, "his": true,
}
// notAnAntecedent — words that cannot be the thing "он" refers to: pronouns,
// particles, conjunctions, adverbs of time. If only these come before "он", the
// message never named a third party and "он" is him.
var notAnAntecedent = map[string]bool{
"не": true, "ни": true, "и": true, "а": true, "но": true, "да": true,
"же": true, "бы": true, "ли": true, "вот": true, "уже": true,
"ещё": true, "еще": true, "тоже": true, "там": true, "тут": true,
"здесь": true, "это": true, "что": true, "как": true, "когда": true,
"чтобы": true, "потому": true, "сейчас": true, "потом": true,
// Time words. A nudge almost always opens with one ("сегодня он не ел"),
// and without them the very next word is read as the person being talked
// about, so the check misses the exact break it was written for.
"сегодня": true, "вчера": true, "завтра": true, "послезавтра": true,
"утром": true, "днём": true, "днем": true, "вечером": true, "ночью": true,
"опять": true, "снова": true, "весь": true, "всю": true, "целый": true,
"я": true, "мне": true, "меня": true, "мной": true, "мы": true, "нас": true,
"ты": true, "тебя": true, "тебе": true, "тобой": true,
"твой": true, "твоя": true, "твоё": true, "твое": true, "твои": true, "твою": true,
}
// looksVerb — a verb is never the thing "он" refers to, so it must not count as
// an antecedent. Past tense keeps "сервис упал, он не отвечает" working off
// "сервис"; the infinitive and imperative endings are here because a nudge is
// mostly made of them ("попробуй встать и отдохнуть — у него есть перерыв"
// slipped through with "попробуй" taken for the person being talked about).
func looksVerb(w string) bool {
r := []rune(w)
if len(r) < 3 {
return false
}
for _, suf := range []string{
"л", "ла", "ло", "ли", // past tense
"ть", "ться", "ти", "чь", // infinitive
"й", "йся", "йте", // imperative
} {
if strings.HasSuffix(w, suf) {
return true
}
}
return false
}
func checkAddress(body string) Result {
words := addressWordRE.FindAllString(strings.ToLower(body), -1)
// Every break, not just the first. A bad reply usually breaks in more than
// one way at once — "Смотрите на его потребление воды" is a plural imperative
// AND third person about him — and reporting only the first hid the second,
// which made the failure look milder than it was.
var breaks []string
seen := map[string]bool{}
add := func(msg string) {
if seen[msg] {
return // the same word twice in one message is one problem, not two
}
seen[msg] = true
breaks = append(breaks, msg)
}
for i, w := range words {
if formalPronouns[w] {
add(fmt.Sprintf("formal %q — she says ты/тебя/тебе", w))
}
if pluralVerb(w) && !(i > 0 && prepositions[words[i-1]]) {
add(fmt.Sprintf("plural imperative %q — she uses the singular", w))
}
}
for i, w := range words {
if !thirdPersonHim[w] {
continue
}
named := false
for j := 0; j < i; j++ {
p := words[j]
if !unicode.Is(unicode.Cyrillic, []rune(p)[0]) && !isLatinWord(p) {
continue // punctuation
}
// pluralVerb as well as looksVerb: looksVerb knows the imperative in
// -й/-йте but not the -те plural ("смотрите"), so "Смотрите на его
// потребление воды" counted "смотрите" as the person being talked
// about and the "его" never printed. Third time a verb form has
// blinded this check — if a fourth turns up, the antecedent test
// wants a real morphology table, not another suffix.
if notAnAntecedent[p] || prepositions[p] || thirdPersonHim[p] || looksVerb(p) || pluralVerb(p) {
continue
}
named = true
break
}
if !named {
add(fmt.Sprintf("third person %q with nobody else named — she talks to him, not about him", w))
}
}
if len(breaks) > 0 {
return Result{CheckAddress, false, strings.Join(breaks, " + ")}
}
return Result{CheckAddress, true, ""}
}
func isLatinWord(w string) bool {
r := []rune(w)[0]
return (r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z')
}
// --- the cringe checks ---------------------------------------------------
//
// "Think Jarvis without the cringe part". DESIGN.md § Non-goals: "Not a
@@ -276,12 +595,47 @@ func checkCringe(body string) Result {
// checkOnTopic — the message must name the thing the rule is about. A nudge
// that never mentions water leaves the operator with a chime and no action.
func checkOnTopic(c Case, body string) Result {
return checkOnTopicAny(c.WantAny, body)
}
// checkOnTopicAny is the same test over a bare want-list, so the talk scorer can
// reuse it without owning a nudge Case.
func checkOnTopicAny(wantAny []string, body string) Result {
low := strings.ToLower(body)
for _, want := range c.WantAny {
for _, want := range wantAny {
if strings.Contains(low, strings.ToLower(want)) {
return Result{CheckOnTopic, true, ""}
}
}
return Result{CheckOnTopic, false,
fmt.Sprintf("mentions none of %v", c.WantAny)}
fmt.Sprintf("mentions none of %v", wantAny)}
}
// --- shape checks for the free-form paths --------------------------------
//
// The nudge checks assume one short sentence. Chat and query replies are longer
// by design, so the only shape worth testing there is that the model produced a
// reply at all and did not trail off. Both are failure modes the fallbacks in
// llmphraser.go hide: a truncated or empty generation still returns nil error.
const (
CheckNonEmpty = "nonempty" // she said something
CheckEllipsis = "ellipsis" // she finished the sentence
)
func checkNonEmpty(body string) Result {
if strings.TrimSpace(body) == "" {
return Result{CheckNonEmpty, false, "empty reply"}
}
return Result{CheckNonEmpty, true, ""}
}
// checkEllipsis — a reply ending in "…" or "..." is a generation that ran out of
// tokens, not a stylistic pause. Mid-sentence ellipses are left alone.
func checkEllipsis(body string) Result {
trimmed := strings.TrimRight(strings.TrimSpace(body), `"'»)`)
if strings.HasSuffix(trimmed, "…") || strings.HasSuffix(trimmed, "...") {
return Result{CheckEllipsis, false, "reply trails off in an ellipsis — likely truncated"}
}
return Result{CheckEllipsis, true, ""}
}
+1 -1
View File
@@ -281,7 +281,7 @@ func (r Report) String() string {
fmt.Fprintf(&b, "%s: %d/%d cases pass every check (%.1f%%), %d errors\n",
r.Name, r.Passed, r.Total, 100*r.Accuracy(), r.Errors)
for _, name := range CheckNames {
fmt.Fprintf(&b, " %-9s %d/%d\n", name, r.ByCheck[name], r.Total)
fmt.Fprintf(&b, " %-10s %d/%d\n", name, r.ByCheck[name], r.Total)
}
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
fmt.Fprintf(&b, " by rule: %s\n", renderStats(r.ByRule))
+48 -4
View File
@@ -72,10 +72,12 @@ func TestStubBaseline(t *testing.T) {
// ceiling ("you've been at your desk for 4 hours without a break — step
// away for a bit." is 76 chars but 16 words). Left failing rather than
// raising the ceiling to hide it.
CheckLength: 12,
CheckFeminine: 15,
CheckCringe: 15,
CheckOnTopic: 12,
CheckLength: 12,
CheckFeminine: 15,
CheckHisGender: 15,
CheckAddress: 15,
CheckCringe: 15,
CheckOnTopic: 12,
}
for name, floor := range floors {
if rep.ByCheck[name] < floor {
@@ -104,6 +106,14 @@ func TestChecksCatchWhatTheyClaim(t *testing.T) {
{"masculine predicative", "я должен сказать: попей воды.", CheckFeminine},
// The other direction: HE is male, so second-person masculine is right.
{"second person masculine ok", "ты не пил воду четыре часа.", ""},
// The real observed failure: she addressed him as a woman.
{"feminine second person", "ты давно не отдыхала — попей воды.", CheckHisGender},
{"feminine second person no dash", "ты пила воду четыре часа назад.", CheckHisGender},
// Her own feminine verb next to "ты" is correct and must not be flagged.
{"her feminine verb near ты", "я заметила, что ты не пил воду.", ""},
{"her feminine verb about him", "напомнила тебе про воду.", ""},
// A feminine noun subject in the same sentence is not him.
{"feminine noun subject ok", "зарядка была утром, ты её пропустил, попей воды.", ""},
{"feminine self ok", "я заметила: воды не было четыре часа.", ""},
{"pet name", "милый, попей воды.", CheckCringe},
{"emoji", "попей воды 💧", CheckCringe},
@@ -114,6 +124,10 @@ func TestChecksCatchWhatTheyClaim(t *testing.T) {
{"asks how he feels", "как ты себя чувствуешь? попей воды.", CheckCringe},
{"praise", "молодец! теперь попей воды.", CheckCringe},
{"off topic", "пора бы уже что-то сделать.", CheckOnTopic},
// The two recorded persona breaks from the phrasing eval run. Pinned as
// unit tests because an eval run is sampled and may not reproduce them.
{"formal plural", "Приходите… Жду вас", CheckAddress},
{"third person about him", "Он не ел 11 дней", CheckAddress},
}
for _, tc := range cases {
@@ -135,6 +149,36 @@ func TestChecksCatchWhatTheyClaim(t *testing.T) {
}
}
// TestAddressCheck — the address check on its own, so the messages that must NOT
// trip it can be written without also having to satisfy the on-topic check.
func TestAddressCheck(t *testing.T) {
bad := []string{
"Приходите… Жду вас", // the recorded formal-plural break
"Он не ел 11 дней", // the recorded third-person break
"Выпейте воды, пожалуйста.", // plural imperative on its own
"Ваш обед был давно.", // formal possessive
}
for _, body := range bad {
if r := checkAddress(body); r.Pass {
t.Errorf("persona break not caught: %q", body)
} else {
t.Logf("%q -> %s", body, r.Detail)
}
}
good := []string{
"ты не пил воду четыре часа — попей.", // correct informal address
"сервис netdata упал, он не отвечает.", // legitimately about a third party
"я заметила, что зарядка была утром.", // no address at all
"в интернете опять тихо, всё работает.", // "интернете" is a noun, not an imperative
}
for _, body := range good {
if r := checkAddress(body); !r.Pass {
t.Errorf("clean message flagged: %q -> %s", body, r.Detail)
}
}
}
func TestMoodCheckUsesTheEnum(t *testing.T) {
if r := checkMood("cheerful"); r.Pass {
t.Error("mood outside the enum passed")
+19 -1
View File
@@ -7,6 +7,8 @@ import (
"testing"
"time"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/phraser"
)
@@ -32,6 +34,7 @@ func TestLLMPhrasingBaseline(t *testing.T) {
// case as a phrasing error and read as "the model cannot phrase".
noProxyLoopback(t)
ctx := context.Background()
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
@@ -41,10 +44,25 @@ func TestLLMPhrasingBaseline(t *testing.T) {
// Generous: an unconstrained 0.8B can spend a minute thinking before it
// writes a word, and a timeout would be scored as a model failure.
cfg.Timeout = 5 * time.Minute
// The same shared context block the daemon prepends (internal/persona),
// with an empty config — that is the deployment we actually ship.
cfg.ContextBlock = func() string { return persona.Facts{}.Block(time.Now()) }
p := phraser.NewLLMPhraserAt(base, cfg)
defer p.Close()
rep, err := Score(context.Background(), "llm (0.8B, built-in persona)", p, f)
// Label the run with whatever gguf the server actually has loaded. It used
// to say "0.8B" no matter what, so two runs of two different models came
// out named the same and were easy to mix up when comparing.
model, err := llm.ModelID(ctx, base)
if err != nil {
// An unlabelled score is still a score, but say so loudly — a made-up
// name in a bake-off table is worse than no name.
t.Logf("could not read model id from %s: %v — report will say %q", base, err, llm.UnknownModel)
model = llm.UnknownModel
}
t.Logf("scoring model %s at %s", model, base)
rep, err := Score(ctx, "llm ("+model+", built-in persona)", p, f)
if err != nil {
t.Fatalf("Score: %v", err)
}
+265
View File
@@ -0,0 +1,265 @@
package eval
// This file scores the CONVERSATIONAL paths, the ones the nudge fixture never
// touches: chat, query-with-notes, and general knowledge. All three now carry
// the shared persona block (internal/persona), and all three produce long
// free-form Russian — which is exactly where a persona break (formality, third
// person, masculine self-reference) is most likely and where, until this file,
// nothing could see one.
//
// Why a second fixture instead of more nudge cases: the checks differ. A nudge
// must be one short sentence with no question in it; a chat reply is allowed
// 1-3 sentences and a follow-up question is a FEATURE there. Mixing them would
// need per-case check masks, and the nudge scorer stays untouched this way.
//
// Why per-path reporting: a chat regression and a knowledge regression have
// different causes (chat prompt vs router.KnowledgePrompt), and one blended
// percentage cannot tell them apart.
import (
"context"
_ "embed"
"encoding/json"
"fmt"
"sort"
"strings"
"time"
"github.com/kami/maven/internal/dialogue"
)
//go:embed talk_v1.json
var talkFixtureJSON []byte
// The three phrasing paths under test. Values match the fixture's "path" field.
const (
PathChat = "chat" // PhraseChat
PathQuery = "query" // PhraseQuery with notes
PathKnowledge = "knowledge" // PhraseQuery with no notes
)
// TalkPaths — report order.
var TalkPaths = []string{PathChat, PathQuery, PathKnowledge}
// TalkCheckNames — the checks that apply to a free-form reply, in report order.
// Deliberately a subset of CheckNames: length, mood and "no questions" are nudge
// properties and would fail a correct chat reply. These paths return no mood at
// all, so there is nothing to check there.
var TalkCheckNames = []string{
CheckNonEmpty, CheckEllipsis, CheckLang, CheckFeminine, CheckAddress, CheckOnTopic,
}
// TalkCase — one turn as the daemon would present it.
//
// History is flat text because that is all PhraseChat uses (it concatenates
// turn texts into one user message); intents and slots would be dead fields.
// Notes are what the store would have matched for a query.
//
// WantAny is the on-topic contract: at least one lowercased fragment must appear
// in the reply. Fragments are stems ("пароль" → "парол") so declension does not
// defeat them.
type TalkCase struct {
ID string `json:"id"`
Path string `json:"path"`
Utterance string `json:"utterance"`
History []string `json:"history,omitempty"`
Notes []string `json:"notes,omitempty"`
WantAny []string `json:"want_any"`
Tags []string `json:"tags,omitempty"`
Note string `json:"note,omitempty"`
}
// TalkFixture — the versioned envelope, same gating as Fixture.
type TalkFixture struct {
SchemaVersion int `json:"schema_version"`
Name string `json:"name"`
Notes []string `json:"notes"`
Cases []TalkCase `json:"cases"`
}
// LoadTalk returns the embedded conversational fixture.
func LoadTalk() (TalkFixture, error) {
var f TalkFixture
if err := json.Unmarshal(talkFixtureJSON, &f); err != nil {
return TalkFixture{}, fmt.Errorf("parse talk fixture: %w", err)
}
if f.SchemaVersion != SchemaVersion {
return TalkFixture{}, fmt.Errorf("talk fixture schema_version %d, want %d", f.SchemaVersion, SchemaVersion)
}
if len(f.Cases) == 0 {
return TalkFixture{}, fmt.Errorf("talk fixture has no cases")
}
return f, nil
}
// Talker — the two methods a conversational path must have to be scorable.
// *phraser.LLMPhraser satisfies it; same trick as Nudger.
type Talker interface {
PhraseChat(ctx context.Context, utterance string, history []dialogue.Turn) (string, error)
PhraseQuery(ctx context.Context, utterance string, notes []string) (string, error)
}
// TalkOutcome — one scored case.
type TalkOutcome struct {
Case TalkCase
Reply string
Err error
Latency time.Duration
Pass bool
Failed []string
Reasons []string
}
// TalkReport — the aggregate. ByPath is the point of this scorer.
type TalkReport struct {
Name string
Total int
Passed int
Errors int
ByCheck map[string]int
ByPath map[string]TagStat
Outcomes []TalkOutcome
P50 time.Duration
P95 time.Duration
Max time.Duration
}
// Accuracy — fraction of cases that passed every check.
func (r TalkReport) Accuracy() float64 {
if r.Total == 0 {
return 0
}
return float64(r.Passed) / float64(r.Total)
}
// ScoreTalk runs every case through t and aggregates. A phrasing error scores as
// a miss and is counted separately: "the model was down" and "the model wrote
// something bad" must not be the same number.
func ScoreTalk(ctx context.Context, name string, t Talker, f TalkFixture) (TalkReport, error) {
rep := TalkReport{
Name: name,
Total: len(f.Cases),
ByCheck: map[string]int{},
ByPath: map[string]TagStat{},
}
for _, n := range TalkCheckNames {
rep.ByCheck[n] = 0
}
lat := make([]time.Duration, 0, len(f.Cases))
for _, c := range f.Cases {
start := time.Now()
reply, err := c.run(ctx, t)
o := TalkOutcome{Case: c, Reply: reply, Err: err, Latency: time.Since(start)}
lat = append(lat, o.Latency)
if err != nil {
rep.Errors++
o.Failed = append(o.Failed, "call")
o.Reasons = append(o.Reasons, fmt.Sprintf("phrase error: %v", err))
} else {
for _, res := range RunTalkChecks(c, reply) {
if res.Pass {
rep.ByCheck[res.Name]++
continue
}
o.Failed = append(o.Failed, res.Name)
o.Reasons = append(o.Reasons, res.Name+": "+res.Detail)
}
}
o.Pass = len(o.Failed) == 0
if o.Pass {
rep.Passed++
}
bump(rep.ByPath, c.Path, o.Pass)
rep.Outcomes = append(rep.Outcomes, o)
}
sort.Slice(lat, func(i, j int) bool { return lat[i] < lat[j] })
rep.P50, rep.P95 = percentile(lat, 0.50), percentile(lat, 0.95)
if len(lat) > 0 {
rep.Max = lat[len(lat)-1]
}
return rep, nil
}
// run dispatches the case to its path. knowledge and query are the same method;
// the empty notes slice is what selects the no-notes branch inside PhraseQuery.
func (c TalkCase) run(ctx context.Context, t Talker) (string, error) {
switch c.Path {
case PathChat:
return t.PhraseChat(ctx, c.Utterance, c.turns())
case PathQuery:
return t.PhraseQuery(ctx, c.Utterance, c.Notes)
case PathKnowledge:
return t.PhraseQuery(ctx, c.Utterance, nil)
}
return "", fmt.Errorf("unknown path %q", c.Path)
}
func (c TalkCase) turns() []dialogue.Turn {
turns := make([]dialogue.Turn, 0, len(c.History))
for _, h := range c.History {
turns = append(turns, dialogue.Turn{Text: h})
}
return turns
}
// RunTalkChecks scores one reply. Order matches TalkCheckNames.
func RunTalkChecks(c TalkCase, reply string) []Result {
return []Result{
checkNonEmpty(reply),
checkEllipsis(reply),
checkLang(reply),
checkFeminine(reply),
checkAddress(reply),
checkOnTopicAny(c.WantAny, reply),
}
}
// String renders the comparison table — composite, then per-check so a
// regression names the property, then per-path so it names the prompt.
func (r TalkReport) String() string {
var b strings.Builder
fmt.Fprintf(&b, "%s: %d/%d cases pass every check (%.1f%%), %d errors\n",
r.Name, r.Passed, r.Total, 100*r.Accuracy(), r.Errors)
for _, name := range TalkCheckNames {
fmt.Fprintf(&b, " %-10s %d/%d\n", name, r.ByCheck[name], r.Total)
}
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
fmt.Fprintf(&b, " by path: %s\n", renderStats(r.ByPath))
return b.String()
}
// Failures — per-case detail, sorted by ID so two runs diff cleanly.
func (r TalkReport) Failures() string {
var b strings.Builder
for _, o := range r.sorted() {
if o.Pass {
continue
}
fmt.Fprintf(&b, " %s %q\n %s\n", o.Case.ID, o.Reply, strings.Join(o.Reasons, "; "))
}
return b.String()
}
// Replies — every generated reply verbatim. This is what a human reads to judge
// tone; the score only says which checks fired.
func (r TalkReport) Replies() string {
var b strings.Builder
for _, o := range r.sorted() {
mark := "ok "
if !o.Pass {
mark = "FAIL"
}
fmt.Fprintf(&b, " %s %-9s %-22s %q\n", mark, o.Case.Path, o.Case.ID, o.Reply)
}
return b.String()
}
func (r TalkReport) sorted() []TalkOutcome {
out := append([]TalkOutcome(nil), r.Outcomes...)
sort.Slice(out, func(i, j int) bool { return out[i].Case.ID < out[j].Case.ID })
return out
}
+163
View File
@@ -0,0 +1,163 @@
package eval
import (
"context"
"os"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/phraser"
)
// perPathMinimum — the resolution floor. A per-path score built on a handful of
// cases moves by 12% when a single reply changes, which cannot distinguish a
// prompt regression from noise.
const perPathMinimum = 8
// TestTalkFixture — the fixture itself has to be sound before any score off it
// means anything.
func TestTalkFixture(t *testing.T) {
f, err := LoadTalk()
if err != nil {
t.Fatalf("LoadTalk: %v", err)
}
seen := map[string]bool{}
byPath := map[string]int{}
for _, c := range f.Cases {
if seen[c.ID] {
t.Errorf("duplicate case id %q", c.ID)
}
seen[c.ID] = true
switch c.Path {
case PathChat, PathQuery, PathKnowledge:
default:
t.Errorf("%s: unknown path %q", c.ID, c.Path)
}
byPath[c.Path]++
if strings.TrimSpace(c.Utterance) == "" {
t.Errorf("%s: empty utterance", c.ID)
}
if len(c.WantAny) == 0 {
t.Errorf("%s: no want_any — the reply cannot be checked for topic", c.ID)
}
// A query case with no notes would silently score the knowledge path.
if c.Path == PathQuery && len(c.Notes) == 0 {
t.Errorf("%s: query case has no notes", c.ID)
}
if c.Path == PathKnowledge && len(c.Notes) > 0 {
t.Errorf("%s: knowledge case must have no notes", c.ID)
}
}
for _, p := range TalkPaths {
if byPath[p] < perPathMinimum {
t.Errorf("path %s has %d cases, want at least %d", p, byPath[p], perPathMinimum)
}
}
}
// fakeTalker — a scripted Talker, so the scorer is testable without a model.
type fakeTalker struct{ reply string }
func (f fakeTalker) PhraseChat(context.Context, string, []dialogue.Turn) (string, error) {
return f.reply, nil
}
func (f fakeTalker) PhraseQuery(context.Context, string, []string) (string, error) {
return f.reply, nil
}
// TestScoreTalkCounts — a reply that fails on purpose must be counted on every
// path, so a real run cannot report a hidden zero.
func TestScoreTalkCounts(t *testing.T) {
f, err := LoadTalk()
if err != nil {
t.Fatalf("LoadTalk: %v", err)
}
// Formal address, off-topic, trailing ellipsis: three checks fail at once.
rep, err := ScoreTalk(context.Background(), "fake", fakeTalker{"Приходите, я вас жду…"}, f)
if err != nil {
t.Fatalf("ScoreTalk: %v", err)
}
if rep.Total != len(f.Cases) || rep.Passed != 0 {
t.Errorf("got %d/%d passing, want 0/%d", rep.Passed, rep.Total, len(f.Cases))
}
if rep.ByCheck[CheckAddress] != 0 {
t.Errorf("formal reply passed the address check %d times", rep.ByCheck[CheckAddress])
}
if rep.ByCheck[CheckEllipsis] != 0 {
t.Errorf("truncated reply passed the ellipsis check %d times", rep.ByCheck[CheckEllipsis])
}
for _, p := range TalkPaths {
if rep.ByPath[p].Total == 0 {
t.Errorf("path %s missing from the report", p)
}
}
if !strings.Contains(rep.String(), "by path") {
t.Error("report does not break down by path")
}
}
// TestLLMTalkBaseline — the resident model on the three conversational paths.
// Opt-in exactly like TestLLMPhrasingBaseline: CI has no model and a run costs
// minutes on the CPU target.
//
// MAVEN_LLM_URL=http://127.0.0.1:18099 \
// go test -run TestLLMTalkBaseline ./internal/phraser/eval/
//
// Reports, does not assert a quality bar — the numbers are the input to tuning
// the persona prompt. The one thing worth failing on is a harness fault.
func TestLLMTalkBaseline(t *testing.T) {
base := os.Getenv("MAVEN_LLM_URL")
if base == "" {
t.Skip("MAVEN_LLM_URL unset — point it at a running llama-server (see doc comment)")
}
noProxyLoopback(t)
ctx := context.Background()
f, err := LoadTalk()
if err != nil {
t.Fatalf("LoadTalk: %v", err)
}
cfg := phraser.DefaultConfig("")
cfg.Timeout = 5 * time.Minute
cfg.ContextBlock = func() string { return persona.Facts{}.Block(time.Now()) }
p := phraser.NewLLMPhraserAt(base, cfg)
defer p.Close()
// Unreachable server is fatal here, not a logged warning, and that differs
// from the nudge test on purpose. PhraseNudge returns its errors, so a dead
// server there shows up honestly in the Errors column. PhraseChat and
// PhraseQuery do NOT: they swallow every failure and return a canned string
// ("поговорили.", "не знаю.", "вот что я нашла: …"). So on these three paths
// a dead server produces a full report with 0 errors and a terrible score —
// a number that looks like bad phrasing and is really no phrasing at all.
// Refusing to score without a confirmed model is the only guard available
// until the phraser reports its failures (Vikunja #397).
model, err := llm.ModelID(ctx, base)
if err != nil {
t.Fatalf("no model at %s: %v — refusing to score, these paths hide their errors "+
"and would report a plausible-looking result off a dead server", base, err)
}
t.Logf("scoring model %s at %s", model, base)
rep, err := ScoreTalk(ctx, "llm ("+model+", built-in persona)", p, f)
if err != nil {
t.Fatalf("ScoreTalk: %v", err)
}
t.Log("\n" + rep.String() + "\nreplies:\n" + rep.Replies() + "\nfailures:\n" + rep.Failures())
// And again afterwards: the run takes minutes, and a server that died or got
// OOM-killed halfway through would leave the first cases scored and the rest
// silently canned. Checking only at the start would not catch that.
if _, err := llm.ModelID(ctx, base); err != nil {
t.Fatalf("model at %s went away during the run: %v — the score above is not trustworthy", base, err)
}
}
+227
View File
@@ -0,0 +1,227 @@
{
"schema_version": 1,
"name": "ru-talk-v1",
"notes": [
"Scores the three conversational phrasing paths: chat (PhraseChat), query (PhraseQuery with notes) and knowledge (PhraseQuery with no notes). The nudge fixture does not cover any of them.",
"Nine cases per path, not five. The nudge fixture is 15 sampled cases and cannot resolve a change smaller than ~3 cases; a per-path score off five cases would be worse still. More cases per path is the point of this fixture.",
"The owner is a man, addressed informally as ty, living alone with a home server. Every utterance is written the way he actually talks to her.",
"chat-formality-bait and chat-about-me exist to provoke the two persona breaks the nudge eval caught: the formal vy/vas plural, and talking about him in the third person.",
"want_any fragments are stems so Russian declension does not defeat the on-topic check. They are lowercased before comparison.",
"want_any is a plain substring test, so a fragment that is too short passes by accident: \"ты\" matches inside \"работы\", \"нет\" inside \"интернет\". Keep every fragment to three or more letters of a real stem.",
"Notes are written as the store would have them: short, first person, no punctuation discipline."
],
"cases": [
{
"id": "chat-how-are-you",
"path": "chat",
"utterance": "привет, как дела?",
"want_any": ["норм", "хорош", "порядк", "тут", "работ"],
"tags": ["greeting"],
"note": "The plainest chat turn there is. If the persona breaks anywhere it breaks here first."
},
{
"id": "chat-formality-bait",
"path": "chat",
"utterance": "не могли бы вы подсказать, чем вы сейчас занимаетесь?",
"want_any": ["сейчас", "ничем", "ничего", "жду", "тут"],
"tags": ["persona-bait", "address"],
"note": "Deliberately polite and plural. A small model mirrors the register and answers with vy/vas — the exact break the address check was written for."
},
{
"id": "chat-about-me",
"path": "chat",
"utterance": "расскажи обо мне",
"want_any": ["теб"],
"tags": ["persona-bait", "third-person"],
"note": "Baits the third person: she should say 'ты живёшь один', not 'он живёт один', as if reporting to somebody else."
},
{
"id": "chat-bored-evening",
"path": "chat",
"utterance": "скучно что-то вечером, посоветуй чем заняться",
"want_any": ["можеш", "попробу", "почита", "прогул", "фильм", "серв"],
"tags": ["open-ended"]
},
{
"id": "chat-followup-server",
"path": "chat",
"utterance": "а стоит его вообще перезагружать?",
"history": ["сервер опять шумит как самолёт", "похоже вентилятор"],
"want_any": ["серв", "перезагру", "вентил", "шум"],
"tags": ["history", "anaphora"],
"note": "The pronoun 'его' only resolves through history. Also the one case where 'он' about the server is legitimate."
},
{
"id": "chat-tired",
"path": "chat",
"utterance": "устал я сегодня, весь день за компом",
"want_any": ["отдохн", "устал", "перерыв", "спат", "день"],
"tags": ["tone"],
"note": "Invites the fake-concern and emotional-support drift; the reply should stay plain."
},
{
"id": "chat-thanks",
"path": "chat",
"utterance": "спасибо, выручила",
"want_any": ["пожалуйст", "не за что", "рада", "обращ"],
"tags": ["persona", "feminine"],
"note": "Feminine self-reference is unavoidable in an answer to thanks: 'рада', not 'рад'."
},
{
"id": "chat-what-can-you-do",
"path": "chat",
"utterance": "что ты вообще умеешь?",
"want_any": ["напомн", "замет", "запис", "могу", "умею"],
"tags": ["self-description", "feminine"]
},
{
"id": "chat-joke",
"path": "chat",
"utterance": "расскажи что-нибудь смешное",
"want_any": ["анекдот", "шутк", "смешн", "истори"],
"tags": ["open-ended"],
"note": "Longest free-form generation in the chat set — the most likely place for a truncated reply."
},
{
"id": "query-router-password",
"path": "query",
"utterance": "что я записывал про пароль от роутера?",
"notes": ["пароль от роутера admin/xxK9tp — на наклейке снизу", "роутер висит в коридоре"],
"want_any": ["парол", "роутер", "наклейк"],
"tags": ["notes", "recall"]
},
{
"id": "query-bedtime-yesterday",
"path": "query",
"utterance": "напомни, во сколько я вчера лёг?",
"notes": ["лёг спать в 02:40", "сегодня встал в 9"],
"want_any": ["02:40", "2:40", "полтрет", "ноч"],
"tags": ["notes", "time"]
},
{
"id": "query-doctor-name",
"path": "query",
"utterance": "как звали того стоматолога, которого мне советовали?",
"notes": ["стоматолог Игорь Валерьевич, клиника на Ленина, советовал Дима"],
"want_any": ["игор", "валерьев", "стоматолог"],
"tags": ["notes", "recall"]
},
{
"id": "query-disk-plan",
"path": "query",
"utterance": "я что-то планировал с диском на сервере, что именно?",
"notes": ["купить второй hdd на 4тб под бэкапы", "перенести медиатеку с системного диска"],
"want_any": ["hdd", "бэкап", "диск", "4тб", "медиатек"],
"tags": ["notes", "homeserver"]
},
{
"id": "query-notes-do-not-answer",
"path": "query",
"utterance": "сколько я заплатил за домен?",
"notes": ["домен продлевается в марте", "хостинг оплачен на год вперёд"],
"want_any": ["домен", "не зна", "не указ"],
"tags": ["notes", "negative"],
"note": "The notes do not contain the price. The prompt tells her to say so; a made-up number is the failure being watched for."
},
{
"id": "query-single-note",
"path": "query",
"utterance": "где лежит запасной ключ?",
"notes": ["запасной ключ у соседа с четвёртого этажа"],
"want_any": ["ключ", "сосед", "четверт"],
"tags": ["notes", "single"],
"note": "One note only — PhraseQuery has a separate branch for len(notes) == 1."
},
{
"id": "query-polite-form",
"path": "query",
"utterance": "подскажите, пожалуйста, что у меня записано по машине?",
"notes": ["замена масла на 92 тысячах", "страховка до 14 сентября"],
"want_any": ["масл", "страховк", "92", "сентябр"],
"tags": ["notes", "persona-bait", "address"],
"note": "Polite plural in the question. The answer must still be ty."
},
{
"id": "query-shopping",
"path": "query",
"utterance": "что мне надо было купить?",
"notes": ["купить кофе и фильтры", "закончилась паста"],
"want_any": ["кофе", "фильтр", "паст"],
"tags": ["notes", "list"]
},
{
"id": "query-wifi-guest",
"path": "query",
"utterance": "я записывал гостевой вайфай?",
"notes": ["гостевая сеть maven-guest, пароль 12345678 меняю раз в месяц"],
"want_any": ["guest", "гостев", "12345678", "парол"],
"tags": ["notes", "recall"]
},
{
"id": "know-sky-blue",
"path": "knowledge",
"utterance": "почему небо синее?",
"want_any": ["све", "рассеи", "атмосфер", "син", "волн"],
"tags": ["general"]
},
{
"id": "know-boil-egg",
"path": "knowledge",
"utterance": "сколько варить яйцо вкрутую?",
"want_any": ["минут", "8", "9", "10", "варит"],
"tags": ["general", "practical"]
},
{
"id": "know-ssd-vs-hdd",
"path": "knowledge",
"utterance": "чем ssd отличается от hdd?",
"want_any": ["ssd", "hdd", "быстр", "диск", "механич"],
"tags": ["general", "tech"]
},
{
"id": "know-cat-purr",
"path": "knowledge",
"utterance": "почему кошки мурчат?",
"want_any": ["кош", "мурч", "вибра", "успока"],
"tags": ["general"]
},
{
"id": "know-hiccups",
"path": "knowledge",
"utterance": "как быстро избавиться от икоты?",
"want_any": ["икот", "дыха", "вод", "задерж"],
"tags": ["general", "practical"]
},
{
"id": "know-polite-form",
"path": "knowledge",
"utterance": "не могли бы вы объяснить, что такое vpn?",
"want_any": ["vpn", "туннел", "трафик", "сет", "шифр"],
"tags": ["general", "persona-bait", "address"],
"note": "Polite plural bait on the knowledge prompt, which is a different system prompt from chat and must hold the same line."
},
{
"id": "know-dont-know",
"path": "knowledge",
"utterance": "как зовут моего соседа снизу?",
"want_any": ["не зна", "не мог"],
"tags": ["general", "negative"],
"note": "Unanswerable without notes. Admitting it beats inventing a name; watching for the invention."
},
{
"id": "know-water-per-day",
"path": "knowledge",
"utterance": "сколько воды в день надо пить?",
"want_any": ["вод", "литр", "стакан", "пит"],
"tags": ["general", "health"],
"note": "Overlaps a nudge rule on purpose: the knowledge answer must not turn into a nudge."
},
{
"id": "know-thunder-delay",
"path": "knowledge",
"utterance": "почему гром слышно позже молнии?",
"want_any": ["звук", "све", "быстр", "гром", "молни"],
"tags": ["general"]
}
]
}
+124
View File
@@ -0,0 +1,124 @@
package phraser
import (
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"strings"
"testing"
"github.com/kami/maven/internal/loop"
)
// grammarSpy stands in for llama-server: it records the grammar field of every
// request and always answers with a contract-shaped reply.
type grammarSpy struct {
srv *httptest.Server
grammars []string
}
func newGrammarSpy(t *testing.T) *grammarSpy {
t.Helper()
s := &grammarSpy{}
s.srv = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
var req chatReq
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
t.Errorf("spy: decode request: %v", err)
}
s.grammars = append(s.grammars, req.Grammar)
w.Header().Set("Content-Type", "application/json")
w.Write([]byte(`{"choices":[{"message":{"content":"{\"response\": \"ага\", \"mood\": \"neutral\"}"}}]}`))
}))
t.Cleanup(s.srv.Close)
return s
}
// callAllPhrasingPaths hits every path that expects the JSON contract.
func callAllPhrasingPaths(t *testing.T, p *LLMPhraser) {
t.Helper()
ctx := context.Background()
if _, err := p.PhraseNudge(ctx, loop.Candidate{Rule: loop.WaterRule(), Severity: loop.Sev1}); err != nil {
t.Fatalf("PhraseNudge: %v", err)
}
if _, err := p.PhraseChat(ctx, "привет", nil); err != nil {
t.Fatalf("PhraseChat: %v", err)
}
// Both branches: no notes (general knowledge) and with notes (grounded).
if _, err := p.PhraseQuery(ctx, "сколько воды я выпил", nil); err != nil {
t.Fatalf("PhraseQuery (no notes): %v", err)
}
if _, err := p.PhraseQuery(ctx, "сколько воды я выпил", []string{"два литра"}); err != nil {
t.Fatalf("PhraseQuery (notes): %v", err)
}
}
func TestGrammarIsAttachedToEveryPhrasingRequest(t *testing.T) {
if strings.TrimSpace(responseGrammar) == "" {
t.Fatal("responseGrammar is empty")
}
spy := newGrammarSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{})
callAllPhrasingPaths(t, p)
if len(spy.grammars) != 4 {
t.Fatalf("expected 4 requests, got %d", len(spy.grammars))
}
for i, g := range spy.grammars {
if g != responseGrammar {
t.Errorf("request %d carries grammar %q, want responseGrammar", i, g)
}
}
}
func TestNoGrammarConfigDisablesIt(t *testing.T) {
spy := newGrammarSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{NoGrammar: true})
callAllPhrasingPaths(t, p)
for i, g := range spy.grammars {
if g != "" {
t.Errorf("request %d still carries a grammar with NoGrammar set: %q", i, g)
}
}
}
// The grammar's string rule must accept any codepoint, not just ASCII. Replies
// are Russian: an ASCII-only class would constrain the model into empty replies.
func TestGrammarStringRuleIsNotASCIIOnly(t *testing.T) {
if !strings.Contains(responseGrammar, `([^"\\] | "\\" ["\\/bfnrt])`) {
t.Error("string rule is not the any-codepoint-except-quote-and-backslash class; Cyrillic replies would be impossible")
}
}
// What the grammar describes must survive the parser that reads it back — a
// Russian body with an escaped quote inside, hand-built to test the contract.
func TestGrammarShapedJSONParses(t *testing.T) {
raw := `{"response": "он сказал \"привет\" и ушёл.\nвот так.", "mood": "confused"}`
text, mood := parseResponseMood(raw)
if want := "он сказал \"привет\" и ушёл.\nвот так."; text != want {
t.Errorf("response = %q, want %q", text, want)
}
if mood != "confused" {
t.Errorf("mood = %q, want confused", mood)
}
}
// Every mood the grammar permits is one the contract knows, and all five are there.
func TestGrammarMoodEnumMatchesTheContract(t *testing.T) {
for _, m := range []string{"neutral", "happy", "thinking", "tired", "confused"} {
if !strings.Contains(responseGrammar, `"\"`+m+`\""`) {
t.Errorf("mood %q missing from the grammar", m)
}
}
// No sixth mood: the enum line lists exactly five alternatives.
for _, line := range strings.Split(responseGrammar, "\n") {
if strings.HasPrefix(line, "mood") {
if n := strings.Count(line, "|") + 1; n != 5 {
t.Errorf("mood rule lists %d alternatives, want 5: %s", n, line)
}
}
}
}
+189 -29
View File
@@ -18,6 +18,7 @@ import (
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/router"
)
@@ -39,7 +40,18 @@ type Config struct {
NGpuLayers int
NCtx int
Timeout time.Duration
Persona string // optional prompt prefix tuning maven's character
// ContextBlock renders the shared context block (who he is, how to
// address him, the time) fresh for each turn. See internal/persona.
// nil ⇒ no block, the prompts stand alone.
ContextBlock func() string
// NoGrammar turns the GBNF constraint off (zero value ⇒ grammar ON).
// The escape hatch exists because the target resident model — the
// locally CPT'd Qwen3-1.7B — does not exist yet: if its chat template
// ever fights the grammar, the fix should be a config flip on the
// deploy box, not a code change and a rebuild.
NoGrammar bool
}
func DefaultConfig(modelPath string) Config {
@@ -184,7 +196,10 @@ func (p *LLMPhraser) PhraseNudge(ctx context.Context, c loop.Candidate) (deliver
body, _ = parsePhrase(resp)
}
if body == "" {
body = fmt.Sprintf("%s — %s", c.Rule.Name, sevLabel(c.Severity))
// The model said nothing usable. Say it in Russian anyway — this text
// goes straight to a Russian piper voice, so the old "water — care"
// fallback was unspeakable.
body = fallbackNudge(c)
}
if mood == "" {
mood = "neutral"
@@ -199,7 +214,7 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
if len(notes) == 0 {
// General knowledge — no notes to ground the answer. The system
// prompt is the single tested source in router.KnowledgePrompt.
sys := router.KnowledgePrompt()
sys := persona.Prepend(p.cfg.ContextBlock, router.KnowledgePrompt())
prompt := fmt.Sprintf("Пользователь спрашивает: \"%s\".", utterance)
resp, err := p.chatWithSystem(ctx, sys, prompt, 256)
if err != nil || resp == "" {
@@ -235,7 +250,7 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
// message array from dialogue history + the current user utterance. Falls back
// to a simple greeting on any LLM error — better to say something than nothing.
func (p *LLMPhraser) PhraseChat(ctx context.Context, utterance string, history []dialogue.Turn) (string, error) {
sys := chatSystemPrompt(p.cfg.Persona)
sys := chatSystemPrompt(p.cfg.ContextBlock)
msgs := []chatMsg{
{Role: "system", Content: sys},
}
@@ -264,17 +279,14 @@ func (p *LLMPhraser) PhraseChat(ctx context.Context, utterance string, history [
}
// chatSystemPrompt returns the system prompt for conversational chat.
// Prepends the configured persona when set.
func chatSystemPrompt(persona string) string {
// Prepends the shared context block when the phraser has one.
func chatSystemPrompt(block func() string) string {
base := `You are maven, a self-hosted personal assistant. You're talking with your owner.
Keep replies brief (1-3 sentences) and natural. You're helpful, curious, and a little warm.
Respond in the user's language (Russian or English, matching their last message).
Never roleplay emotions you don't have, but stay friendly.
Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}. "response" is your reply text; "mood" reflects your tone (neutral/happy/thinking/tired/confused).`
if persona != "" {
base = persona + "\n\n" + base
}
return base
return persona.Prepend(block, base)
}
// chatWithMessages sends a full message array (system + history + current) to
@@ -285,6 +297,7 @@ func (p *LLMPhraser) chatWithMessages(ctx context.Context, msgs []chatMsg, maxTo
Messages: msgs,
Temperature: 0.7,
MaxTokens: maxTokens,
Grammar: p.grammar(),
}
body, err := json.Marshal(req)
if err != nil {
@@ -364,6 +377,37 @@ type chatReq struct {
Messages []chatMsg `json:"messages"`
Temperature float64 `json:"temperature"`
MaxTokens int `json:"max_tokens"`
// Grammar is llama-server's `grammar` field (GBNF). Same wiring as
// internal/llm.Req.Grammar. Empty ⇒ unconstrained sampling.
Grammar string `json:"grammar,omitempty"`
}
// responseGrammar — GBNF constraining the model to the documented phrasing
// contract and nothing else: {"response": "<text>", "mood": "<enum>"}.
//
// Without it a 0.8B answers roughly one chat turn in three with open reasoning
// as plain text ("Thinking Process:" …), which no tag-stripper can remove and
// which eats the token budget before the JSON closes. Modelled on
// routeGrammar in internal/router/llmrouter.go so the two read alike.
//
// text accepts ANY codepoint except the two JSON must escape — the replies are
// Russian, so an ASCII-only rule would make every reply empty. The escape rule
// is what lets the model close a string it opened with a quote inside. Length
// is bounded so a repetition loop truncates the field, not the JSON object.
const responseGrammar = `
root ::= "{" ws "\"response\"" ws ":" ws string ws "," ws "\"mood\"" ws ":" ws mood ws "}"
mood ::= "\"neutral\"" | "\"happy\"" | "\"thinking\"" | "\"tired\"" | "\"confused\""
string ::= "\"" ([^"\\] | "\\" ["\\/bfnrt]){0,400} "\""
ws ::= [ \t\n]*
`
// grammar returns the GBNF to attach to a phrasing request, or "" when the
// operator turned it off.
func (p *LLMPhraser) grammar() string {
if p.cfg.NoGrammar {
return ""
}
return responseGrammar
}
type chatResp struct {
@@ -388,6 +432,7 @@ func (p *LLMPhraser) chatWithSystem(ctx context.Context, system, user string, ma
},
Temperature: 0.7,
MaxTokens: maxTokens,
Grammar: p.grammar(),
}
body, err := json.Marshal(req)
if err != nil {
@@ -428,40 +473,155 @@ func (p *LLMPhraser) chatWithSystem(ctx context.Context, system, user string, ma
return stripThink(content), nil
}
// nudgeSystem — the phrasing contract for nudges.
//
// Written as filled-in examples, not as a schema with "..." in it. A 0.8B
// copies whatever sits in the response slot, so a literal placeholder there
// teaches it to answer with the placeholder. Measured: 7/15 nudges came back
// as "..." before this. See PHRASING-EVAL-31-07-2026.md.
//
// Russian only, feminine self-reference, second person masculine (the owner is
// a man). She talks TO him, informally, singular — never "вы", never "он".
// One short sentence — the nudge is spoken aloud.
const nudgeSystem = `Ты Maven, домашняя ассистентка. О себе говоришь в женском роде ("я проверила", "я записала"). Владелец мужчина, обращайся к нему в мужском роде ("ты пил", "ты забыл").
Говоришь с ним на "ты", в единственном числе ("выпей", "встань"). Никогда не "вы"/"вас"/"ваш" и никогда "он"/"его" ты говоришь ему, а не о нём.
Пиши ОДНО короткое напоминание по-русски: не больше 120 символов и не больше 16 слов. Только по делу.
Запрещено: обращения ("дорогой", "милый"), эмодзи, извинения ("прости", "извини"), вопросы о самочувствии, похвала, больше одного восклицательного знака, английские слова кроме имён сервисов.
Отвечай ТОЛЬКО одним объектом JSON с полями "response" и "mood".
"response" сам текст напоминания.
"mood" ровно одно из: neutral, happy, thinking, tired, confused.
Так выглядит правильный ответ по форме. Темы здесь посторонние их в запросе не будет:
{"response": "Стиральная машина закончила. Развесь бельё.", "mood": "neutral"}
{"response": "Ноутбук на трёх процентах. Я поставила его на зарядку.", "mood": "confused"}
Это примеры ФОРМЫ, а не темы. Пиши только про ту ситуацию, которую тебе дали в запросе. Не копируй примеры и никогда не пиши "..." в поле response.`
func (p *LLMPhraser) systemPrompt() string {
base := `You are maven, a self-hosted personal assistant. Generate brief, natural nudge messages in the user's language (Russian or English). Respond ONLY with valid JSON: {"response": "full voice message", "mood": "neutral"}. "response" is what the user hears; "mood" reflects maven's tone (neutral/happy/thinking/tired/confused).`
if p.cfg.Persona != "" {
base = p.cfg.Persona + "\n\n" + base
}
return base
return persona.Prepend(p.cfg.ContextBlock, nudgeSystem)
}
// querySystemPrompt returns the system prompt for PhraseQuery (notes + general
// knowledge). Prepends the configured persona when set.
func (p *LLMPhraser) querySystemPrompt() string {
base := "You are maven, a self-hosted personal assistant answering from your notes. Answer briefly and naturally in Russian starting with \"вот что я нашла: \". Respond ONLY with valid JSON: {\"response\": \"...\", \"mood\": \"neutral\"}."
if p.cfg.Persona != "" {
base = p.cfg.Persona + "\n\n" + base
return persona.Prepend(p.cfg.ContextBlock, base)
}
// ruleTopics — Russian gloss for each built-in rule name. The rule names are
// English identifiers; a 0.8B asked to nudge about "netdata_critical" writes
// about nothing. The daemon knows what its own rules mean, so it says so.
var ruleTopics = map[string]string{
"water": "он давно не пил воду",
"meal": "он давно не ел",
"break": "он давно без перерыва, пора встать и размяться",
"service_down": "сервис не отвечает, лежит",
"netdata_critical": "критический алярм в netdata, проблема с диском или местом",
}
// ruleKeywords — the word the message must contain. The 0.8B drifts to
// whatever topic it saw last unless the required word is named outright.
var ruleKeywords = map[string]string{
"water": "воду",
"meal": "поешь",
"break": "перерыв",
"service_down": "сервис",
"netdata_critical": "диск",
}
// ruleTopic turns a rule name into a Russian description of the situation.
// "routine:зарядка" and "morning:утро" carry their own Russian suffix.
func ruleTopic(rule string) string {
if t, ok := ruleTopics[rule]; ok {
return t
}
return base
if i := strings.IndexByte(rule, ':'); i > 0 && i+1 < len(rule) {
switch rule[:i] {
case "morning":
return "утро, пора начать день: " + rule[i+1:]
default:
return "пора сделать по распорядку: " + rule[i+1:]
}
}
return rule
}
// ruleKeyword — the word the nudge must contain, or "" when the rule name's
// own Russian suffix already is that word.
func ruleKeyword(rule string) string {
if k, ok := ruleKeywords[rule]; ok {
return k
}
if i := strings.IndexByte(rule, ':'); i > 0 && i+1 < len(rule) {
return rule[i+1:]
}
return ""
}
// ruDur — duration in Russian. humanDur is English and its output was landing
// verbatim in the message.
func ruDur(d time.Duration) string {
if d < 0 {
d = 0
}
h, m := int(d.Hours()), int(d.Minutes())%60
switch {
case h >= 2:
return fmt.Sprintf("%d ч", h)
case h == 1 && m >= 30:
return "полтора часа"
case h == 1:
return "час"
default:
return fmt.Sprintf("%d мин", m)
}
}
// fallbackNudge — plain Russian for when the model returns nothing parseable.
var fallbackNudges = map[string]string{
"water": "Ты давно не пил воду.",
"meal": "Ты давно не ел, поешь.",
"break": "Пора сделать перерыв.",
"service_down": "Сервис не отвечает.",
"netdata_critical": "Критический алярм: проверь диск.",
}
func fallbackNudge(c loop.Candidate) string {
if s, ok := fallbackNudges[c.Rule.Name]; ok {
return s
}
if kw := ruleKeyword(c.Rule.Name); kw != "" {
return "Напоминаю: " + kw + "."
}
return "Напоминаю о деле."
}
func buildNudgePrompt(c loop.Candidate) string {
var ctxParts []string
ctxParts = append(ctxParts, fmt.Sprintf("Rule: %s", c.Rule.Name))
ctxParts = append(ctxParts, fmt.Sprintf("Severity: %s", sevLabel(c.Severity)))
ctxParts = append(ctxParts, "Ситуация: "+ruleTopic(c.Rule.Name))
if f, ok := c.State.Facts[c.Rule.Name]; ok && f.Key != "" && f.Key != c.Rule.Name {
ctxParts = append(ctxParts, "Что именно: "+f.Key)
}
if d, ok := c.State.Since(c.Rule.Name); ok {
ctxParts = append(ctxParts, fmt.Sprintf("Duration since last event: %s", humanDur(d)))
ctxParts = append(ctxParts, "Прошло: "+ruDur(d))
}
switch sevLabel(c.Severity) {
case "alarm":
ctxParts = append(ctxParts, "Срочно, скажи прямо.")
case "ops":
ctxParts = append(ctxParts, "Это про сервер, не про здоровье.")
}
tail := "Напиши напоминание про эту ситуацию. Одно предложение, по-русски, в JSON."
if kw := ruleKeyword(c.Rule.Name); kw != "" {
// Last line on purpose: a 0.8B weights the end of the prompt hardest,
// and without the required word it drifts back to the examples.
tail += " Ответ ДОЛЖЕН содержать слово «" + kw + "»."
}
return fmt.Sprintf(
`Generate a nudge message. Context:
%s
Respond as JSON: {"response": "...", "mood": "..."}`,
strings.Join(ctxParts, "\n"),
)
return strings.Join(ctxParts, "\n") + "\n\n" + tail
}
type responseMood struct {
+23
View File
@@ -2,6 +2,7 @@ package router
import (
"context"
"fmt"
"math"
"unicode"
)
@@ -19,6 +20,24 @@ type Embedder interface {
Close() error
}
// IdentifiedEmbedder — an embedder that can name itself. The name goes into
// the DB next to the vectors it wrote, so a later model swap is caught instead
// of silently returning nonsense scores (Vikunja #378).
type IdentifiedEmbedder interface {
Embedder
ID() string
}
// EmbedderID is the stable string stored alongside the vectors. It comes from
// the embedder itself — nobody hand-types a model name twice — and changes
// whenever the model or its dimension changes.
func EmbedderID(e Embedder) string {
if i, ok := e.(IdentifiedEmbedder); ok {
return i.ID()
}
return fmt.Sprintf("unknown@%d", e.Dim())
}
// AsymmetricEmbedder — an embedder that wants to know whether a text is a
// search query or a stored passage. Recall is asymmetric: a short question
// goes in, a longer note comes out. The e5 family is trained for exactly that
@@ -70,6 +89,10 @@ func NewHashEmbedder(dim int) *HashEmbedder {
func (h *HashEmbedder) Dim() int { return h.dim }
// ID names this embedder for the DB marker. The dimension is part of it
// because a HashEmbedder of another width is a different vector space.
func (h *HashEmbedder) ID() string { return fmt.Sprintf("hash@%d", h.dim) }
func (h *HashEmbedder) Close() error { return nil }
func (h *HashEmbedder) Embed(_ context.Context, text string) ([]float32, error) {
+24
View File
@@ -0,0 +1,24 @@
package router
import "testing"
func TestEmbedderIDFromModelPath(t *testing.T) {
got := modelIDFromPath("/opt/maven/models/embedder/multilingual-e5-small.onnx")
if got != "multilingual-e5-small@384" {
t.Fatalf("modelIDFromPath = %q", got)
}
// A different model file must produce a different id, even at 384 dim.
old := modelIDFromPath("/opt/maven/models/embedder/paraphrase-multilingual-MiniLM-L12-v2.onnx")
if old == got {
t.Fatal("two different models share one id")
}
}
func TestEmbedderIDIncludesDim(t *testing.T) {
if id := EmbedderID(NewHashEmbedder(1024)); id != "hash@1024" {
t.Fatalf("EmbedderID = %q", id)
}
if EmbedderID(NewHashEmbedder(1024)) == EmbedderID(NewHashEmbedder(384)) {
t.Fatal("dimension not part of the id")
}
}
+17 -99
View File
@@ -1,11 +1,8 @@
package eval
import (
"bytes"
"context"
"encoding/json"
"fmt"
"net/http"
"os"
"strings"
"testing"
@@ -29,13 +26,20 @@ import (
// a bake-off across checkpoints (#278, #250) produces tables you can tell
// apart. Point the variable at one server at a time.
//
// Three configurations, because "the LLM router" is ambiguous and the three
// numbers answer different questions:
// Two configurations, because "the LLM router" is ambiguous and the two numbers
// answer different questions:
//
// llm-only — the model alone. Measures the prompt + grammar contract.
// cascade+llm — what #320 would actually ship: stage-0 grammar, then the
// model, then the classifier as the failure floor.
// llm-no-thinking — diagnostic only, not a shippable path (see below).
// llm-only — the model alone. Measures the prompt + grammar contract.
// cascade+llm — what #320 would actually ship: stage-0 grammar, then the
// model, then the classifier as the failure floor.
//
// There used to be a third, "thinking off", which looked 6 points better. It is
// gone: it was measured with a hand-rolled HTTP client that quietly dropped
// repeat_penalty, so the gap was the missing penalty and not the thinking mode.
// Re-measured with everything else held equal, thinking off scores exactly the
// same, case for case — and a direct probe shows this llama-server build ignores
// enable_thinking / reasoning_budget for this model anyway, so there was nothing
// to turn off. Full write-up in ROUTING-EVAL-31-07-2026.md (Vikunja #376).
func TestLLMRouterBaseline(t *testing.T) {
base := os.Getenv("MAVEN_LLM_URL")
if base == "" {
@@ -57,12 +61,12 @@ func TestLLMRouterBaseline(t *testing.T) {
}
ctx := context.Background()
model, err := ModelID(ctx, base)
model, err := llm.ModelID(ctx, base)
if err != nil {
// Not fatal: an unlabelled score is still a score. But say so loudly,
// because an unlabelled row in a bake-off table is worthless.
t.Logf("could not read model id from %s: %v — reports will say %q", base, err, "unknown-model")
model = "unknown-model"
t.Logf("could not read model id from %s: %v — reports will say %q", base, err, llm.UnknownModel)
model = llm.UnknownModel
}
t.Logf("scoring model %s at %s", model, base)
lr := router.NewLLMRouter(client)
@@ -96,104 +100,18 @@ func TestLLMRouterBaseline(t *testing.T) {
}
t.Log("\n" + repCascade.String() + repCascade.Failures())
// llm-no-thinking: same prompt and grammar with the chat template's
// thinking mode off. Qwen3.5's template defaults thinking=1, so under a
// grammar the constrained JSON lands in reasoning_content with content
// empty — llm.Client's ReasoningContent fallback is what makes the router
// work at all today, by accident rather than design.
//
// MEASURED 2026-07-31: this variant scores identically to as-deployed
// (18/76, 48.7% intent-only, 2 errors, same p50). Thinking mode is a
// non-issue under a grammar — llama.cpp constrains the same token stream
// either way. Kept so the question stays answered instead of being
// re-asked, and so internal/llm does NOT grow a chat_template_kwargs field
// for a problem that does not exist.
repNoThink, err := Score(ctx, "llm-only ("+model+", thinking off) [diagnostic]",
RouterFunc(func(ctx context.Context, u string, now time.Time) (router.Decision, error) {
d, ok, err := router.NewLLMRouter(&noThinkCompleter{base: base, http: &http.Client{Timeout: 60 * time.Second}}).Route(ctx, u, now)
if err != nil {
return d, err
}
if !ok {
return d, fmt.Errorf("llm router declined without an error")
}
return d, nil
}), f)
if err != nil {
t.Fatalf("Score no-thinking: %v", err)
}
t.Log("\n" + repNoThink.String() + repNoThink.Failures())
// Reports rather than asserts — the numbers are inputs to the #320
// decision, and an assertion here would be this test inventing the bar.
// The one thing worth failing on is a harness fault: if every single case
// errors, the run measured infrastructure, not routing, and the report
// must not be mistaken for a score.
for _, rep := range []Report{repLLM, repCascade, repNoThink} {
for _, rep := range []Report{repLLM, repCascade} {
if rep.Errors == rep.Total {
t.Errorf("%s: all %d cases errored — harness fault, not a measurement", rep.Name, rep.Total)
}
}
}
// noThinkCompleter — llm.Client with chat_template_kwargs.enable_thinking
// false. A test-local copy rather than a change to internal/llm: whether the
// daemon should send it is the open question, and answering it here by adding
// the field would prejudge #320.
type noThinkCompleter struct {
base string
http *http.Client
}
func (c *noThinkCompleter) Complete(ctx context.Context, r llm.Req) (string, error) {
payload := map[string]any{
"messages": []map[string]string{
{"role": "system", "content": r.System},
{"role": "user", "content": r.User},
},
"max_tokens": r.MaxTokens,
"temperature": 0,
"grammar": r.Grammar,
"chat_template_kwargs": map[string]any{"enable_thinking": false},
}
b, err := json.Marshal(payload)
if err != nil {
return "", err
}
req, err := http.NewRequestWithContext(ctx, "POST", c.base+"/v1/chat/completions", bytes.NewReader(b))
if err != nil {
return "", err
}
req.Header.Set("Content-Type", "application/json")
resp, err := c.http.Do(req)
if err != nil {
return "", err
}
defer resp.Body.Close()
if resp.StatusCode != 200 {
return "", fmt.Errorf("status %d", resp.StatusCode)
}
var out struct {
Choices []struct {
Message struct {
Content string `json:"content"`
ReasoningContent string `json:"reasoning_content"`
} `json:"message"`
} `json:"choices"`
}
if err := json.NewDecoder(resp.Body).Decode(&out); err != nil {
return "", err
}
if len(out.Choices) == 0 {
return "", fmt.Errorf("no choices")
}
m := out.Choices[0].Message
if m.Content != "" {
return m.Content, nil
}
return m.ReasoningContent, nil
}
func ping(ctx context.Context, c *llm.Client) error {
ctx, cancel := context.WithTimeout(ctx, 90*time.Second)
defer cancel()
+1
View File
@@ -22,6 +22,7 @@
{ "id": "ru-query-011", "utterance": "почему сервер тормозит", "lang": "ru", "intent": "query", "tags": ["homelab", "hard"], "note": "diagnostic question, not a chat opener" },
{ "id": "ru-query-012", "utterance": "какие заметки я оставил про полив", "lang": "ru", "intent": "query", "tags": ["recall"] },
{ "id": "ru-query-013", "utterance": "во сколько у меня встреча", "lang": "ru", "intent": "query", "tags": ["calendar"] },
{ "id": "ru-query-019", "utterance": "что у меня стоит в календаре на послезавтра", "lang": "ru", "intent": "query", "tags": ["calendar", "hard"], "note": "agenda, not the clock: the daemon answers this from CalendarEvents inside the query branch, so the clock/date system rule must not swallow it" },
{ "id": "ru-query-014", "utterance": "я успеваю до дедлайна", "lang": "ru", "intent": "query", "tags": ["hard", "no-question-word"] },
{ "id": "ru-query-015", "utterance": "сколько я прошёл шагов", "lang": "ru", "intent": "query", "tags": ["aggregate"] },
{ "id": "ru-query-016", "utterance": "покажи давление за неделю", "lang": "ru", "intent": "query", "tags": ["hard", "imperative"], "note": "imperative form but a read — must not route to act" },
+4 -1
View File
@@ -3,5 +3,8 @@ package router
// KnowledgePrompt returns the system prompt for general knowledge questions
// that the phraser uses when no notes match the query.
func KnowledgePrompt() string {
return `Ты — Мавена, персональный ассистент. Ответь кратко из своих знаний. Если не знаешь — скажи "не знаю". Не выдумывай. Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}.`
// No self-introduction here: the shared persona block already says who she
// is, and this line used to disagree with it — a different name ("Мавена")
// and a masculine noun ("ассистент") in front of a feminine persona.
return `Ответь кратко из своих знаний. Если не знаешь — скажи "не знаю". Не выдумывай. Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}.`
}
+9 -1
View File
@@ -11,10 +11,18 @@ func TestKnowledgePrompt(t *testing.T) {
t.Fatal("KnowledgePrompt returned empty string")
}
// Must contain key instructions
checks := []string{"Мавена", "не знаю", "не выдумывай"}
checks := []string{"не знаю", "не выдумывай"}
for _, c := range checks {
if !strings.Contains(strings.ToLower(prompt), strings.ToLower(c)) {
t.Errorf("KnowledgePrompt should mention %q", c)
}
}
// Who she is comes from the shared persona block now. This prompt used to
// say it too, with a different name and a masculine noun, which is the
// drift the block exists to stop.
for _, w := range []string{"Мавена", "ассистент"} {
if strings.Contains(prompt, w) {
t.Errorf("KnowledgePrompt should not introduce her (%q) — the persona block does", w)
}
}
}
+23 -8
View File
@@ -44,10 +44,21 @@ ws ::= [ \t\n]*
// Changed again 31-07-2026: added the "unknown" escape hatch so the model can
// admit it cannot route (Vikunja #359).
//
// Changed again 31-07-2026: added the clock/calendar rule (Vikunja #374). The
// prompt never said which side "который час" or "какое число завтра" belong on,
// so the model guessed — `system→query ×4` in every eval run. The rule sits
// above the question test on purpose: these utterances all carry a question
// word, so a later rule would never be reached. The boundary is what the
// daemon can actually answer: only replySystem in cmd/mavend/voice.go owns the
// clock and the calendar formatter, while the agenda ("что у меня завтра") is
// answered inside the query branch, so that side stays query.
//
// The training workspace keeps its own copy of this prompt for relabelling, and
// `llm/check_prompt_parity.py` there compares the two. That copy is in another
// repo and was not touched, so parity will fail until it gets the same edits —
// both the rule reorder and the "unknown" wording (Vikunja #362).
// both the rule reorder and the "unknown" wording (Vikunja #362) — and now the
// clock/calendar rule too. The training workspace is not checked out on this
// box at all, so it could not be updated here; #362 still covers the catch-up.
const routeSystem = `Классифицируй ровно одно сообщение пользователя. Верни ОДИН JSON-массив действий.
Ровно одно намерение: fact, reminder, note, query, act, chat, system.
@@ -56,19 +67,21 @@ const routeSystem = `Классифицируй ровно одно сообще
Классифицируй по цели пользователя. Порядок решения:
1. Хочет напоминание в будущем reminder
2. Явно просит сохранить информацию note
3. Задаёт вопрос: есть вопросительное слово (сколько, что, какой, когда, где, кто, почему, как) или знак «?» query
4. Хочет получить информацию, в том числе о своих же данных query
5. Утверждает: сообщает или обновляет текущее состояние/событие fact
6. Просит выполнить работу act
7. Про ассистента, настройки или память system
8. Реплика обрывок или указание на неназванное («это», «то», «потом»), и без него непонятно, что именно нужно сделать unknown
9. Иначе chat
3. Спрашивает только «который час» / «какое число» / «какой день недели» сами часы или календарная дата, без своих данных system
4. Задаёт вопрос: есть вопросительное слово (сколько, что, какой, когда, где, кто, почему, как) или знак «?» query
5. Хочет получить информацию, в том числе о своих же данных query
6. Утверждает: сообщает или обновляет текущее состояние/событие fact
7. Просит выполнить работу act
8. Про ассистента, настройки или память system
9. Реплика обрывок или указание на неназванное («это», «то», «потом»), и без него непонятно, что именно нужно сделать unknown
10. Иначе chat
Различия:
- note сохранить информацию, без напоминания. text = суть.
- reminder уведомить позже. text = что напомнить.
- fact неявное обновление: пользователь сообщает, что что-то в мире изменилось (текущее/изменённое состояние, случившееся событие). key/value.
- unknown редкий случай. Ставь его, только если в самой реплике нет ни предмета, ни действия. Короткая, простая или незнакомая тема это не причина для unknown: приветствие и болтовня это chat, вопрос на любую тему это query, просьба сделать что-то названное это act.
- system против query часы и календарная дата сами по себе (сколько времени, какое число, какой день недели можно и про завтра, и про другой город) это system. А что записано в календаре или в памяти («что у меня завтра», «какие есть напоминания») это query. Если в реплике есть просьба (напомни, запиши, сделай), то названное время просто деталь просьбы, и это не system.
- query против fact решает форма реплики, а не тема. Вопрос о состоянии это query, даже если названо то же самое, что бывает в fact. Только утверждение это fact.
Примеры:
@@ -82,6 +95,8 @@ const routeSystem = `Классифицируй ровно одно сообще
"что такое docker?" {"intent":"query","text":"что такое docker"}
"напиши письмо" {"intent":"act","verb":"написать письмо"}
"очисти память" {"intent":"system"}
"который час?" {"intent":"system"}
"какое число завтра?" {"intent":"system"}
"привет" {"intent":"chat","text":"привет"}
"сделай это" {"intent":"unknown"}
"ну это" {"intent":"unknown"}
+21
View File
@@ -33,6 +33,7 @@ const (
type onnxEmbedder struct {
tokenizer *unigramTokenizer
session *ort.DynamicSession[int64, float32]
id string
}
func NewONNXEmbedder(modelPath, tokenizerPath, libPath string) (*onnxEmbedder, error) {
@@ -58,11 +59,31 @@ func NewONNXEmbedder(modelPath, tokenizerPath, libPath string) (*onnxEmbedder, e
return &onnxEmbedder{
tokenizer: tok,
session: session,
id: modelIDFromPath(modelPath),
}, nil
}
func (e *onnxEmbedder) Dim() int { return embedDim }
// ID names the loaded model for the DB marker (Vikunja #378): the model file's
// own name plus the dimension, so pointing the config at another model changes
// the string on its own.
func (e *onnxEmbedder) ID() string { return e.id }
// modelIDFromPath turns /opt/.../multilingual-e5-small.onnx into
// "multilingual-e5-small@384".
func modelIDFromPath(modelPath string) string {
name := modelPath
if i := strings.LastIndexAny(name, "/\\"); i >= 0 {
name = name[i+1:]
}
name = strings.TrimSuffix(name, ".onnx")
if name == "" {
name = "onnx"
}
return fmt.Sprintf("%s@%d", name, embedDim)
}
// Embed treats the text as a query. The classifier compares one short
// utterance to another short seed phrase, so both sides get the same prefix
// and the comparison stays fair. The recall path must call EmbedQuery and
+22 -8
View File
@@ -449,16 +449,30 @@ func (AnaphoraResolver) Resolve(text string) (ref string, ok bool) {
return "", false
}
// ParseCalendarDate detects RU calendar date words in text and returns the
// resolved time (midnight UTC+0 for "сегодня"/"today", next day for "завтра"/"tomorrow").
// Returns zero time + false if no match.
// ParseCalendarDate detects RU/EN calendar day words in text and returns
// midnight of that day in now's own time zone. Handles "сегодня", "завтра",
// "послезавтра", "вчера" (and the English words). Returns zero time + false
// if no match.
//
// "послезавтра" is checked before "завтра" because it contains it.
func ParseCalendarDate(text string, now time.Time) (time.Time, bool) {
lower := strings.ToLower(text)
if strings.Contains(lower, "сегодня") || strings.Contains(lower, "today") {
return now.Truncate(24 * time.Hour), true
}
if strings.Contains(lower, "завтра") || strings.Contains(lower, "tomorrow") {
return now.Truncate(24 * time.Hour).Add(24 * time.Hour), true
switch {
case strings.Contains(lower, "сегодня") || strings.Contains(lower, "today"):
return midnight(now, 0), true
case strings.Contains(lower, "послезавтра") || strings.Contains(lower, "day after tomorrow"):
return midnight(now, 2), true
case strings.Contains(lower, "завтра") || strings.Contains(lower, "tomorrow"):
return midnight(now, 1), true
case strings.Contains(lower, "вчера") || strings.Contains(lower, "yesterday"):
return midnight(now, -1), true
}
return time.Time{}, false
}
// midnight returns the start of the day that is `days` away from now, in
// now's time zone (now.Truncate(24h) would cut on a UTC boundary instead).
func midnight(now time.Time, days int) time.Time {
y, m, d := now.AddDate(0, 0, days).Date()
return time.Date(y, m, d, 0, 0, 0, 0, now.Location())
}
+3
View File
@@ -45,6 +45,9 @@ func TestParseCalendarDate(t *testing.T) {
{"расписание на завтра", time.Date(2026, 7, 7, 0, 0, 0, 0, time.UTC), true},
{"what's today", time.Date(2026, 7, 6, 0, 0, 0, 0, time.UTC), true},
{"tomorrow plans", time.Date(2026, 7, 7, 0, 0, 0, 0, time.UTC), true},
{"какое число послезавтра", time.Date(2026, 7, 8, 0, 0, 0, 0, time.UTC), true},
{"что было вчера", time.Date(2026, 7, 5, 0, 0, 0, 0, time.UTC), true},
{"yesterday plans", time.Date(2026, 7, 5, 0, 0, 0, 0, time.UTC), true},
{"какая погода", time.Time{}, false},
{"сколько времени", time.Time{}, false},
{"", time.Time{}, false},
+161
View File
@@ -0,0 +1,161 @@
package store
import (
"context"
"encoding/json"
"fmt"
"time"
)
// EmbedFunc embeds one piece of stored text. The caller passes
// router.EmbedPassage — the STORE side of the query/passage asymmetry, which is
// the side every vector in the DB was written with. (Passing the query side
// would put the stored vectors in the wrong half of the space and quietly halve
// recall.) A func instead of an interface keeps this package free of any
// dependency on internal/router.
type EmbedFunc func(ctx context.Context, text string) ([]float32, error)
// BackfillResult is what the re-embed run did, for logging.
type BackfillResult struct {
Skipped bool // marker already matched — nothing to do
Notes int // rows rewritten in the notes table
Facts int // fact rows rewritten in memory_vectors
MemNotes int // note rows rewritten in memory_vectors
NoText int // memory_vectors rows with no text in their meta, left alone
Took time.Duration
}
// ReembedAll rewrites every stored vector with the currently configured
// embedder and then records that embedder as the one that owns the DB.
//
// Both places a vector lives are rewritten in the same pass: the `notes` table
// `embedding` column and the `memory_vectors` rows (notes AND facts). Doing
// only one would leave the two indexes disagreeing, which is worse than leaving
// both stale.
//
// Safe to re-run: if the marker already names the current embedder there is
// nothing to fix, so it returns immediately with Skipped set.
//
// Crash safety: everything — every vector and the marker — happens inside one
// transaction. If anything fails or the process dies partway, the transaction
// rolls back: no vectors changed and no marker written, so the next run does
// the whole job again. The marker is never set unless the full rewrite
// committed.
func (s *Store) ReembedAll(ctx context.Context, currentID string, embed EmbedFunc) (BackfillResult, error) {
start := time.Now()
var res BackfillResult
stored, err := s.Meta(ctx, metaKeyEmbedderID)
if err != nil {
return res, err
}
if stored == currentID {
res.Skipped = true
res.Took = time.Since(start)
return res, nil
}
tx, err := s.db.BeginTx(ctx, nil)
if err != nil {
return res, fmt.Errorf("reembed: begin: %w", err)
}
defer tx.Rollback() // no-op once committed
// ----- notes table -----
type noteRow struct {
id int64
text string
}
var notes []noteRow
rows, err := tx.QueryContext(ctx, `SELECT id, text FROM notes WHERE text != ''`)
if err != nil {
return res, fmt.Errorf("reembed: read notes: %w", err)
}
for rows.Next() {
var n noteRow
if err := rows.Scan(&n.id, &n.text); err != nil {
rows.Close()
return res, fmt.Errorf("reembed: note row: %w", err)
}
notes = append(notes, n)
}
rows.Close()
if err := rows.Err(); err != nil {
return res, fmt.Errorf("reembed: notes: %w", err)
}
for _, n := range notes {
vec, err := embed(ctx, n.text)
if err != nil {
return res, fmt.Errorf("reembed: embed note %d: %w", n.id, err)
}
if _, err := tx.ExecContext(ctx,
`UPDATE notes SET embedding = ? WHERE id = ?`, floatsToBlob(vec), n.id); err != nil {
return res, fmt.Errorf("reembed: write note %d: %w", n.id, err)
}
res.Notes++
}
// ----- memory_vectors (the unified index: notes AND facts) -----
// The text to re-embed is the one carried in the row's meta blob, which is
// exactly the text that was embedded when the row was written.
type vecRow struct {
id, text, kind string
}
var vecs []vecRow
rows, err = tx.QueryContext(ctx, `SELECT id, meta FROM memory_vectors`)
if err != nil {
return res, fmt.Errorf("reembed: read memory vectors: %w", err)
}
for rows.Next() {
var id, metaJSON string
if err := rows.Scan(&id, &metaJSON); err != nil {
rows.Close()
return res, fmt.Errorf("reembed: memory row: %w", err)
}
meta := map[string]string{}
if err := json.Unmarshal([]byte(metaJSON), &meta); err != nil {
rows.Close()
return res, fmt.Errorf("reembed: meta for %q: %w", id, err)
}
if meta["text"] == "" {
res.NoText++
continue
}
vecs = append(vecs, vecRow{id: id, text: meta["text"], kind: meta["type"]})
}
rows.Close()
if err := rows.Err(); err != nil {
return res, fmt.Errorf("reembed: memory vectors: %w", err)
}
for _, v := range vecs {
vec, err := embed(ctx, v.text)
if err != nil {
return res, fmt.Errorf("reembed: embed %q: %w", v.id, err)
}
if _, err := tx.ExecContext(ctx,
`UPDATE memory_vectors SET vec = ? WHERE id = ?`, encodeVec(vec), v.id); err != nil {
return res, fmt.Errorf("reembed: write %q: %w", v.id, err)
}
if v.kind == "fact" {
res.Facts++
} else {
res.MemNotes++
}
}
// Same transaction as the rewrite, on purpose: the marker can only exist if
// every vector above was written.
if _, err := tx.ExecContext(ctx,
`INSERT INTO meta (key, value) VALUES (?,?)
ON CONFLICT(key) DO UPDATE SET value = excluded.value`,
metaKeyEmbedderID, currentID); err != nil {
return res, fmt.Errorf("reembed: write marker: %w", err)
}
if err := tx.Commit(); err != nil {
return res, fmt.Errorf("reembed: commit: %w", err)
}
res.Took = time.Since(start)
return res, nil
}
+171
View File
@@ -0,0 +1,171 @@
package store
import (
"context"
"errors"
"testing"
"time"
)
// markerVec is a recognisable vector: nothing in these tests writes it except
// the backfill, so finding it proves the row really was rewritten.
var markerVec = []float32{9, 9, 9}
func newEmbedder(calls *int) EmbedFunc {
return func(_ context.Context, _ string) ([]float32, error) {
*calls++
return markerVec, nil
}
}
// seedOldVectors puts one note (notes table + unified index) and one fact
// (unified index only) in the DB, both carrying obviously-old vectors.
func seedOldVectors(t *testing.T, s *Store) {
t.Helper()
ctx := context.Background()
old := []float32{0.1, 0.2, 0.3}
id, err := s.WriteNote(ctx, time.Now(), "молоко в холодильнике", old, "voice")
if err != nil {
t.Fatalf("WriteNote: %v", err)
}
mem := s.VectorMemory()
if err := mem.Insert(ctx, "note:1", old, map[string]string{
"type": "note", "text": "молоко в холодильнике",
}); err != nil {
t.Fatalf("Insert note vector: %v", err)
}
if err := mem.Insert(ctx, "fact:water:1", old, map[string]string{
"type": "fact", "text": "я пил воду",
}); err != nil {
t.Fatalf("Insert fact vector: %v", err)
}
_ = id
}
func noteVec(t *testing.T, s *Store) []float32 {
t.Helper()
var blob []byte
if err := s.db.QueryRow(`SELECT embedding FROM notes LIMIT 1`).Scan(&blob); err != nil {
t.Fatalf("read note embedding: %v", err)
}
return blobToFloats(blob)
}
func memVec(t *testing.T, s *Store, id string) []float32 {
t.Helper()
var blob []byte
if err := s.db.QueryRow(`SELECT vec FROM memory_vectors WHERE id = ?`, id).Scan(&blob); err != nil {
t.Fatalf("read memory vector %s: %v", id, err)
}
return decodeVec(blob)
}
func sameVec(a, b []float32) bool {
if len(a) != len(b) {
return false
}
for i := range a {
if a[i] != b[i] {
return false
}
}
return true
}
// The deployed case: old vectors everywhere, no marker. Every vector in both
// places must be rewritten and the marker recorded.
func TestReembedAllRewritesEveryVector(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
seedOldVectors(t, s)
calls := 0
res, err := s.ReembedAll(ctx, "multilingual-e5-small@384", newEmbedder(&calls))
if err != nil {
t.Fatalf("ReembedAll: %v", err)
}
if res.Skipped {
t.Fatal("first run should not skip")
}
if res.Notes != 1 || res.MemNotes != 1 || res.Facts != 1 {
t.Fatalf("counts: notes=%d memNotes=%d facts=%d", res.Notes, res.MemNotes, res.Facts)
}
if calls != 3 {
t.Fatalf("embedder called %d times, want 3", calls)
}
if !sameVec(noteVec(t, s), markerVec) {
t.Fatalf("notes table not rewritten: %v", noteVec(t, s))
}
if !sameVec(memVec(t, s, "note:1"), markerVec) {
t.Fatal("unified index note row not rewritten")
}
if !sameVec(memVec(t, s, "fact:water:1"), markerVec) {
t.Fatal("unified index fact row not rewritten")
}
got, err := s.Meta(ctx, metaKeyEmbedderID)
if err != nil {
t.Fatalf("Meta: %v", err)
}
if got != "multilingual-e5-small@384" {
t.Fatalf("marker = %q", got)
}
}
// Re-running must do nothing at all — not a second pass over the same rows.
func TestReembedAllSecondRunIsNoop(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
seedOldVectors(t, s)
calls := 0
if _, err := s.ReembedAll(ctx, "e5@384", newEmbedder(&calls)); err != nil {
t.Fatalf("first run: %v", err)
}
first := calls
res, err := s.ReembedAll(ctx, "e5@384", newEmbedder(&calls))
if err != nil {
t.Fatalf("second run: %v", err)
}
if !res.Skipped {
t.Fatal("second run should report Skipped")
}
if calls != first {
t.Fatalf("second run embedded %d more rows, want 0", calls-first)
}
}
// A failure partway must leave the DB exactly as it was: no marker, and the old
// vectors still in place (one transaction, rolled back).
func TestReembedAllPartialFailureLeavesMarkerUnset(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
seedOldVectors(t, s)
before := noteVec(t, s)
calls := 0
boom := func(_ context.Context, _ string) ([]float32, error) {
calls++
if calls == 2 {
return nil, errors.New("onnx blew up")
}
return markerVec, nil
}
if _, err := s.ReembedAll(ctx, "e5@384", boom); err == nil {
t.Fatal("expected an error")
}
got, err := s.Meta(ctx, metaKeyEmbedderID)
if err != nil {
t.Fatalf("Meta: %v", err)
}
if got != "" {
t.Fatalf("marker was set to %q after a failed run", got)
}
if !sameVec(noteVec(t, s), before) {
t.Fatal("a failed run left a partially rewritten notes table")
}
// And the mismatch warning must still fire, so the user knows to re-run.
if _, mismatch, err := s.CheckEmbedder(ctx, "e5@384"); err != nil || !mismatch {
t.Fatalf("CheckEmbedder after failed backfill: mismatch=%v err=%v", mismatch, err)
}
}
+8 -3
View File
@@ -13,11 +13,15 @@ import (
// unknown = a pending row found stale at startup: the process that started it
// is gone, and the send may or may not have reached the external channel.
// Never auto-resolved into sent or failed — that would be guessing.
// dropped = the routing table deliberately suppressed this one (a care nudge
// while you're away). Nothing was sent and nothing went wrong; the row exists
// so "she dropped it" and "the rule never fired" don't look the same later.
const (
DeliveryPending = "pending"
DeliverySent = "sent"
DeliveryFailed = "failed"
DeliveryUnknown = "unknown"
DeliveryDropped = "dropped"
)
// BeginDeliveryAttempt durably records intent to send BEFORE the external
@@ -43,10 +47,11 @@ func (s *Store) BeginDeliveryAttempt(ctx context.Context, kind, rule string, rem
}
// CompleteDeliveryAttempt records the sink's outcome for a prior
// BeginDeliveryAttempt. status is "sent" or "failed" — never "pending" or
// "unknown" (those are set only by Begin and reconciliation respectively).
// BeginDeliveryAttempt. status is "sent", "failed" or "dropped" — never
// "pending" or "unknown" (those are set only by Begin and reconciliation
// respectively).
func (s *Store) CompleteDeliveryAttempt(ctx context.Context, id int64, status string, now time.Time) error {
if status != DeliverySent && status != DeliveryFailed {
if status != DeliverySent && status != DeliveryFailed && status != DeliveryDropped {
return fmt.Errorf("store: invalid delivery completion status %q", status)
}
_, err := s.db.ExecContext(ctx,
+34
View File
@@ -0,0 +1,34 @@
package store
import (
"context"
"testing"
"time"
)
// TestDroppedDeliveryAttemptRoundTrips — Vikunja #370. A suppressed nudge is
// recorded as 'dropped'. The status column has a CHECK constraint, so this
// only works if migration #12 widened it; a fake outbox in a unit test would
// not catch that.
func TestDroppedDeliveryAttemptRoundTrips(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
now := time.Now()
id, err := s.BeginDeliveryAttempt(ctx, "nudge", "water", 0, "drop", "abc123", now)
if err != nil {
t.Fatalf("BeginDeliveryAttempt: %v", err)
}
if err := s.CompleteDeliveryAttempt(ctx, id, DeliveryDropped, now); err != nil {
t.Fatalf("CompleteDeliveryAttempt: %v", err)
}
var status string
err = s.db.QueryRowContext(ctx, `SELECT status FROM delivery_attempts WHERE id = ?`, id).Scan(&status)
if err != nil {
t.Fatalf("read back: %v", err)
}
if status != DeliveryDropped {
t.Fatalf("status: want %q, got %q", DeliveryDropped, status)
}
}
+74
View File
@@ -0,0 +1,74 @@
package store
import (
"context"
"fmt"
"time"
)
// DialogueSessionRow — one saved follow-up session. Data is the session
// encoded by the dialogue package; the store does not look inside it.
type DialogueSessionRow struct {
ID string
Data []byte
Ts time.Time
TTL time.Duration
Expires time.Time
}
// SaveDialogueSession — write (or replace) the session for one dialogue id.
// One row per id: a newer turn overwrites the older state.
func (s *Store) SaveDialogueSession(ctx context.Context, id string, data []byte, ts time.Time, ttl time.Duration) error {
expires := ts.Add(ttl)
_, err := s.db.ExecContext(ctx, `
INSERT INTO dialogue_sessions (id, data, ts, ttl_ms, expires_ts) VALUES (?, ?, ?, ?, ?)
ON CONFLICT(id) DO UPDATE SET data = excluded.data,
ts = excluded.ts,
ttl_ms = excluded.ttl_ms,
expires_ts = excluded.expires_ts`,
id, data, ts.UnixMilli(), ttl.Milliseconds(), expires.UnixMilli())
if err != nil {
return fmt.Errorf("save dialogue session: %w", err)
}
return nil
}
// DeleteDialogueSession — drop one session (ended, or expired).
func (s *Store) DeleteDialogueSession(ctx context.Context, id string) error {
if _, err := s.db.ExecContext(ctx, `DELETE FROM dialogue_sessions WHERE id = ?`, id); err != nil {
return fmt.Errorf("delete dialogue session: %w", err)
}
return nil
}
// LoadDialogueSessions — return the sessions still alive at `now` and delete
// the ones that already ran out. An expired session is dead: it never comes
// back after a restart.
func (s *Store) LoadDialogueSessions(ctx context.Context, now time.Time) ([]DialogueSessionRow, error) {
if _, err := s.db.ExecContext(ctx,
`DELETE FROM dialogue_sessions WHERE expires_ts <= ?`, now.UnixMilli()); err != nil {
return nil, fmt.Errorf("prune dialogue sessions: %w", err)
}
rows, err := s.db.QueryContext(ctx,
`SELECT id, data, ts, ttl_ms, expires_ts FROM dialogue_sessions ORDER BY id`)
if err != nil {
return nil, fmt.Errorf("load dialogue sessions: %w", err)
}
defer rows.Close()
var out []DialogueSessionRow
for rows.Next() {
var r DialogueSessionRow
var tsMilli, ttlMilli, expMilli int64
if err := rows.Scan(&r.ID, &r.Data, &tsMilli, &ttlMilli, &expMilli); err != nil {
return nil, fmt.Errorf("scan dialogue session: %w", err)
}
r.Ts = time.UnixMilli(tsMilli).UTC()
r.TTL = time.Duration(ttlMilli) * time.Millisecond
r.Expires = time.UnixMilli(expMilli).UTC()
out = append(out, r)
}
if err := rows.Err(); err != nil {
return nil, fmt.Errorf("load dialogue sessions: %w", err)
}
return out, nil
}
+97
View File
@@ -0,0 +1,97 @@
package store
import (
"context"
"database/sql"
"errors"
"fmt"
)
// metaKeyEmbedderID names the embedder that wrote the stored vectors.
//
// Why one value for the whole DB and not a column on every vector row: the
// vectors are only ever rewritten all at once (one backfill re-embeds every
// note and fact together), so a per-row marker would hold the same string in
// every row and cost a column on two tables for nothing.
const metaKeyEmbedderID = "embedder_id"
// Meta reads a single value from the meta table. Missing key ⇒ empty string.
func (s *Store) Meta(ctx context.Context, key string) (string, error) {
var v string
err := s.db.QueryRowContext(ctx, `SELECT value FROM meta WHERE key = ?`, key).Scan(&v)
if errors.Is(err, sql.ErrNoRows) {
return "", nil
}
if err != nil {
return "", fmt.Errorf("read meta %s: %w", key, err)
}
return v, nil
}
// SetMeta writes (or overwrites) a single meta value.
func (s *Store) SetMeta(ctx context.Context, key, value string) error {
_, err := s.db.ExecContext(ctx,
`INSERT INTO meta (key, value) VALUES (?,?)
ON CONFLICT(key) DO UPDATE SET value = excluded.value`, key, value)
if err != nil {
return fmt.Errorf("write meta %s: %w", key, err)
}
return nil
}
// EmbedderUnknown is the stored id reported for a DB that already holds
// vectors but never recorded who wrote them.
const EmbedderUnknown = "unknown (written before this marker existed)"
// CheckEmbedder compares the embedder now configured against the one that
// wrote the stored vectors. Returns the stored id and whether it differs.
//
// Vectors from two different models live in different spaces, so cosine
// between them is noise rather than a low score — and both of our models are
// 384-dimensional, so nothing else catches it.
//
// Three cases, and the middle one is the one that actually matters:
//
// - marker present ⇒ compare the two ids.
// - marker absent but vectors already stored ⇒ this is a DB from before the
// marker, so we cannot know who wrote them. Report a mismatch. This is the
// real case on the deployed box: those vectors came from the old embedder,
// and claiming them for the current one would hide the exact problem the
// marker was added to catch.
// - marker absent and no vectors ⇒ fresh DB, claim it, nothing to fix.
//
// On a mismatch the fix is ReembedAll (backfill.go), run explicitly with
// `mavend -reembed`. Nothing is re-embedded here: that work is minutes of CPU
// on the laptop and must not stall a normal start.
func (s *Store) CheckEmbedder(ctx context.Context, currentID string) (stored string, mismatch bool, err error) {
stored, err = s.Meta(ctx, metaKeyEmbedderID)
if err != nil {
return "", false, err
}
if stored != "" {
return stored, stored != currentID, nil
}
n, err := s.countVectors(ctx)
if err != nil {
return "", false, err
}
if n > 0 {
return EmbedderUnknown, true, nil
}
return currentID, false, s.SetMeta(ctx, metaKeyEmbedderID, currentID)
}
// countVectors — how many stored rows carry an embedding. Used only to tell a
// fresh DB apart from one that predates the marker.
func (s *Store) countVectors(ctx context.Context) (int, error) {
var notes, vecs int
if err := s.db.QueryRowContext(ctx,
`SELECT count(*) FROM notes WHERE embedding IS NOT NULL`).Scan(&notes); err != nil {
return 0, fmt.Errorf("count note vectors: %w", err)
}
if err := s.db.QueryRowContext(ctx,
`SELECT count(*) FROM memory_vectors`).Scan(&vecs); err != nil {
return 0, fmt.Errorf("count memory vectors: %w", err)
}
return notes + vecs, nil
}
+100
View File
@@ -0,0 +1,100 @@
package store
import (
"context"
"testing"
"time"
)
// A fresh DB has no marker yet, so the current embedder is recorded and
// nothing is flagged.
func TestCheckEmbedderFreshDBRecords(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
stored, mismatch, err := s.CheckEmbedder(ctx, "multilingual-e5-small@384")
if err != nil {
t.Fatalf("CheckEmbedder: %v", err)
}
if mismatch {
t.Fatal("fresh DB reported a mismatch")
}
if stored != "multilingual-e5-small@384" {
t.Fatalf("stored = %q", stored)
}
got, err := s.Meta(ctx, metaKeyEmbedderID)
if err != nil {
t.Fatalf("Meta: %v", err)
}
if got != "multilingual-e5-small@384" {
t.Fatalf("marker not persisted, got %q", got)
}
}
// The deployed box: notes were written by the old embedder, before the marker
// existed. Claiming them for the current one would hide exactly the problem
// the marker is for, so an unmarked DB that already holds vectors is a
// mismatch.
func TestCheckEmbedderUnmarkedDBWithVectorsIsMismatch(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
if _, err := s.WriteNote(ctx, time.Now(), "молоко в холодильнике", []float32{0.1, 0.2}, "voice"); err != nil {
t.Fatalf("WriteNote: %v", err)
}
stored, mismatch, err := s.CheckEmbedder(ctx, "multilingual-e5-small@384")
if err != nil {
t.Fatalf("CheckEmbedder: %v", err)
}
if !mismatch {
t.Fatal("an unmarked DB with stored vectors should report a mismatch")
}
if stored != EmbedderUnknown {
t.Fatalf("stored = %q, want %q", stored, EmbedderUnknown)
}
// It must NOT claim the DB — that would silence the warning on restart.
got, err := s.Meta(ctx, metaKeyEmbedderID)
if err != nil {
t.Fatalf("Meta: %v", err)
}
if got != "" {
t.Fatalf("marker written despite unknown provenance: %q", got)
}
}
// Both models are 384-dim, so this is the only thing that catches the swap.
func TestCheckEmbedderDifferentModelMismatch(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
if err := s.SetMeta(ctx, metaKeyEmbedderID, "paraphrase-multilingual-MiniLM-L12-v2@384"); err != nil {
t.Fatalf("SetMeta: %v", err)
}
stored, mismatch, err := s.CheckEmbedder(ctx, "multilingual-e5-small@384")
if err != nil {
t.Fatalf("CheckEmbedder: %v", err)
}
if !mismatch {
t.Fatal("different embedder not detected")
}
if stored != "paraphrase-multilingual-MiniLM-L12-v2@384" {
t.Fatalf("stored = %q", stored)
}
}
// The same embedder must never raise a false alarm, including on re-check.
func TestCheckEmbedderSameModelNoAlarm(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
for i := 0; i < 2; i++ {
_, mismatch, err := s.CheckEmbedder(ctx, "multilingual-e5-small@384")
if err != nil {
t.Fatalf("CheckEmbedder: %v", err)
}
if mismatch {
t.Fatalf("false alarm on pass %d", i)
}
}
}
+38
View File
@@ -74,6 +74,44 @@ ALTER TABLE reminders ADD COLUMN next_fire_ts INTEGER;`, // #2
`CREATE INDEX IF NOT EXISTS idx_nudges_snoozed ON nudges (outcome_ts) WHERE outcome = 'snoozed';`, // #8 — SnoozedUntil runs every tick; keep it off a full scan (Vikunja #364)
`ALTER TABLE proposed_routines ADD COLUMN accepted_ts INTEGER;
ALTER TABLE proposed_routines ADD COLUMN last_fired_ts INTEGER;`, // #9 — accepted routines keep firing (Vikunja #366): the tick loop needs to know when a routine was accepted and when it last nudged
`CREATE TABLE IF NOT EXISTS dialogue_sessions (
id TEXT PRIMARY KEY,
data BLOB NOT NULL,
ts INTEGER NOT NULL,
ttl_ms INTEGER NOT NULL,
expires_ts INTEGER NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_dialogue_sessions_expires ON dialogue_sessions (expires_ts);`, // #10 — the follow-up session survives a restart (Vikunja #363); small, TTL-pruned table, not a history log
`CREATE TABLE IF NOT EXISTS meta (
key TEXT PRIMARY KEY,
value TEXT NOT NULL
);`, // #11 — small key/value table for facts about the DB itself; first key is embedder_id (Vikunja #378)
// #12 — a suppressed nudge gets a 'dropped' row (Vikunja #370). sqlite
// can't widen a CHECK constraint in place, so the table is rebuilt; the
// index goes with the old table and is recreated. The columns are listed
// out rather than `SELECT *` — copying by position would silently shuffle
// every row if the old table's column order ever differed from this one.
`CREATE TABLE delivery_attempts_v12 (
id INTEGER PRIMARY KEY AUTOINCREMENT,
kind TEXT NOT NULL CHECK (kind IN ('nudge','reminder')),
rule TEXT NOT NULL DEFAULT '',
reminder_id INTEGER NOT NULL DEFAULT 0,
channel TEXT NOT NULL,
body_hash TEXT NOT NULL,
status TEXT NOT NULL DEFAULT 'pending' CHECK (status IN ('pending','sent','failed','unknown','dropped')),
created_ts INTEGER NOT NULL,
completed_ts INTEGER
);
INSERT INTO delivery_attempts_v12
(id, kind, rule, reminder_id, channel, body_hash, status, created_ts, completed_ts)
SELECT id, kind, rule, reminder_id, channel, body_hash, status, created_ts, completed_ts
FROM delivery_attempts;
DROP TABLE delivery_attempts;
ALTER TABLE delivery_attempts_v12 RENAME TO delivery_attempts;
CREATE INDEX IF NOT EXISTS idx_delivery_attempts_status ON delivery_attempts (status);`,
}
// migrate applies every migration with a number greater than the DB's current
+79 -4
View File
@@ -1,8 +1,10 @@
#!/usr/bin/env bash
# Unified script to stop all Maven services.
# Usage: ./kill-maven.sh
# - Graceful SIGTERM is attempted first.
# - If any process lingers, force with SIGKILL.
# - Docker deploy: `docker compose stop` (see why below).
# - Bare-metal / dev run: graceful SIGTERM first, SIGKILL if anything lingers.
# Exits non-zero if it cannot confirm everything is stopped. It must never say
# "stopped" unless it checked.
set -euo pipefail
@@ -24,6 +26,73 @@ else
LLM='llama-server.*\.gguf'
fi
COMPOSE_FILE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/docker-compose.yml"
# --- containerised deploy ------------------------------------------------
# docker-compose.yml does not set `pid: host`, so each container has its own
# PID namespace: pkill on the host sees nothing inside them. This script used
# to print "all stopped" while every daemon was still happily running. Stop the
# containers through compose instead — that actually reaches them.
#
# running_containers prints the ids of the project's running containers, or
# nothing. Empty output plus a non-zero return means "could not ask docker",
# which is different from "nothing is running" and is handled below.
running_containers() {
docker compose -f "$COMPOSE_FILE" ps -q --status running 2>/dev/null
}
DOCKER_OK=0
CONTAINERS=""
if command -v docker >/dev/null 2>&1 && [ -f "$COMPOSE_FILE" ]; then
if CONTAINERS="$(running_containers)"; then
DOCKER_OK=1
fi
fi
if [ "$DOCKER_OK" = 1 ] && [ -n "$CONTAINERS" ]; then
echo "--- Maven is running in containers: stopping via docker compose ---"
if ! docker compose -f "$COMPOSE_FILE" stop; then
echo "ERROR: 'docker compose stop' failed. Containers may still be running." >&2
exit 1
fi
echo "--- Verifying containers are gone ---"
LEFT="$(running_containers || true)"
if [ -n "$LEFT" ]; then
echo "ERROR: containers still running after stop:" >&2
docker compose -f "$COMPOSE_FILE" ps >&2 || true
exit 1
fi
echo "All containers stopped."
exit 0
fi
# --- bare-metal / dev run -----------------------------------------------
# pgrep -f matches whole command lines, so a shell that merely mentions
# "mavend" (this script's own parent, for one) shows up. Drop ourselves and our
# parent, otherwise the SIGKILL sweep can take out the terminal you ran this in.
host_pids() {
pgrep -f "$PAT|$LLM" | grep -v -e "^$$\$" -e "^$PPID\$" | paste -sd, - || true
}
HOST_PIDS=$(host_pids)
if [ -z "$HOST_PIDS" ]; then
# Nothing on the host. Whether that means "already down" depends on whether
# we managed to ask docker, and the two must not read the same.
if [ "$DOCKER_OK" = 1 ]; then
# Docker answered and named no running containers, and there is nothing
# on the host either. That is a real answer: Maven is already stopped.
echo "Nothing to stop: no Maven processes and no running containers."
exit 0
fi
# We could not ask docker, so Maven may be alive in a container we cannot
# see. Saying "stopped" here is the exact false success this script had.
echo "ERROR: no Maven processes on this host, and docker could not be asked." >&2
echo " If this is the container deploy it may still be running:" >&2
echo " docker compose -f $COMPOSE_FILE stop" >&2
echo " Nothing was stopped. Check by hand before assuming Maven is down." >&2
exit 1
fi
echo "--- Sending graceful SIGTERM to Maven services ---"
pkill -TERM -f "$PAT" || true
# mavend's Pdeathsig SIGKILLs its llama-server on exit, but sweep strays too
@@ -32,12 +101,18 @@ pkill -TERM -f "$LLM" || true
echo "--- Verifying processes are gone ---"
sleep 1
PIDS=$(pgrep -d ',' -f "$PAT|$LLM") || PIDS=""
PIDS=$(host_pids)
if [ -n "$PIDS" ]; then
echo "Warning: some processes still alive. PIDs: $PIDS"
echo "--- Force killing with SIGKILL ---"
echo "$PIDS" | tr ',' '\n' | xargs -r kill -9
sleep 1
LEFT=$(host_pids)
if [ -n "$LEFT" ]; then
echo "ERROR: still alive after SIGKILL. PIDs: $LEFT" >&2
exit 1
fi
echo "Done (SIGKILL)."
else
echo "All services gracefully stopped."
fi
fi