Compare commits

...

98 Commits

Author SHA1 Message Date
kami f7e1187823 Match quiet-mode toggles on whole words, and resolve OFF first 2026-08-01 01:02:10 +04:00
kami ed48c59ba7 Merge branch 'refactor/query-sources' into integration/small-batch 2026-08-01 00:54:11 +04:00
kami b09967f9e6 Split actions.go into per-intent files
Pure move: actionFact, actionReminder, actionAct and actionNote each get
their own actions_<intent>.go. The two small ones (chat, system) and the
actionHandlers table stay in actions.go, which is now just the dispatch
layer and the notes about what does not belong in it. No behaviour
change — only the file a handler is read in.
2026-08-01 00:53:37 +04:00
kami b4a3867479 Turn actionQuery into a chain of query sources
The six answer sources were hand-unrolled inside one 127-line function.
The intent table is a closed set of 7, but this list is open-ended —
Kiwix (#286), RSS (#258), the crawler (#259) and email (#246) each add
one. Each is now a registry entry: a name plus a method on the handler,
walked in order until one claims the question.

Order is unchanged and still load-bearing (memory before the notes-only
pass, #373), the confidence gate keeps its position and semantics, and
every reply string, log line and best-effort failure is verbatim.
2026-08-01 00:51:44 +04:00
kami 88d07b5175 Unify the voice and text turn pipelines into runTurn
HandlePushToTalk and handleText hand-wrote the same eight-step turn
sequence twice, comments in the latter saying "same as HandlePushToTalk"
four times. Extract it into runTurn(ctx, text) string: the voice path
wraps it in stt/tts, the text path returns it directly.

The two had drifted. The text path was missing the quiet-hours toggle
check entirely, so "тихий режим" over IPC/telegram fell through to the
classifier; unifying gives it the check. It also logged the route result
and applyAction return where the voice path did not — both logs are kept
for both paths.
2026-08-01 00:48:50 +04:00
kami c00e3003bf Merge branch 'refactor/praxis-capability-registry' into integration/small-batch 2026-08-01 00:43:54 +04:00
kami c0f9834528 Turn the Praxis act dispatch into a capability registry 2026-08-01 00:43:16 +04:00
kami ad5eb2d1cf Walk a chain of confirm resolvers instead of three copied blocks 2026-08-01 00:42:23 +04:00
kami 5253123d99 Merge branch 'refactor/voice-wiring' into integration/small-batch
# Conflicts:
#	cmd/mavend/voice.go
2026-07-31 23:54:11 +04:00
kami 2abf98dea6 Move the voice daemon wiring and startup out of voice.go 2026-07-31 23:52:51 +04:00
kami f5c71b87f3 Move the confirm/park gate out of voice.go 2026-07-31 23:51:33 +04:00
kami 5934110fa8 Move the Praxis/Hexis act handling out of voice.go
voice.go is still the biggest file in cmd/mavend and most of what is left
has nothing to do with the audio path. The ecosystem integration is one
such lump: it talks to Nexus, Praxis and Hexis over HTTP and only touches
the handler for its store and clock. Lifting it into ecosystem_acts.go
puts it next to ecosystem.go, where the clients it drives already live.

Move-only: handlePraxisAct, recordPraxisTrace (called from nowhere else),
handleHexisAct and execHexis verbatim, plus the two imports that became
unused in voice.go.
2026-07-31 23:43:47 +04:00
kami 6e47a3d736 Merge branch 'worktree-agent-a88193d1d84b04a5b' into integration/small-batch 2026-07-31 23:38:00 +04:00
kami 67a5eb3805 Split applyAction's 300-line switch into a per-intent handler table
applyAction (cmd/mavend/voice.go) dispatched all 7 intents from one giant
switch. Extract each case body verbatim into its own actionXxx method in
new cmd/mavend/actions.go, dispatched from an actionHandlers table keyed by
router.Intent. applyAction itself is now just the dec.Clarify guard plus a
table lookup.

No behaviour change: same reply strings, same side-effect order, comments
moved verbatim. The destructive-act confirm gate and the enabled-tool
allowlist stay entirely inside actionAct, exactly where they lived in the
old switch's IntentAct case — they're act-specific, not cross-cutting, so
they don't move to a separate layer. dec.Clarify short-circuit, dialogue
bookkeeping and detectPattern stay outside the table since they run
regardless of intent.

voice.go: 1638 -> 1344 lines. New actions.go: 362 lines.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:37:33 +04:00
kami 029449eefa Table-drive the IPC dispatcher instead of a 42-arm switch
dispatch() replaces the hand-written switch with a package-level
map[Method]handlerFunc built once at init. Each entry is one
withParams/withParamsVoid/withoutParams call closing only over the
CoreAPI method it invokes — adding a method is now one table line
instead of a new arm.

Check still runs once at the top before any unmarshal, unchanged. The
three non-CoreAPI methods (assert_stepup, store_encryption_key, unlock)
are special-cased before the table lookup since they drive Server
fields (StepUp/WrapKeyFn/UnlockFn), not store state. The current
CoreAPI is loaded once per dispatch and passed into the handler as an
argument, so SetAPI's runtime swap (the unlock transition) still takes
effect on the next request — the table itself never captures an api
value. No wire-format change; existing round-trip and unknown-method
tests in ipc_test.go pass unmodified.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:36:46 +04:00
kami ac36216f5d Merge branch 'worktree-agent-a517419e93e6a219f' into integration/small-batch 2026-07-31 23:31:24 +04:00
kami db17cfcc65 Delete the dead lockedAPI, add UnimplementedCoreAPI for the doubles 2026-07-31 23:31:24 +04:00
kami 7d676eb941 Stop tracking the mavwaked build artifact 2026-07-31 23:30:52 +04:00
kami fe3a4e9514 Merge branch 'worktree-agent-a9e5cef90b263a5e5' into integration/small-batch 2026-07-31 23:27:36 +04:00
kami a906f2afad Extract the pure RU/string/weather helpers out of voice.go 2026-07-31 23:27:08 +04:00
kami a2031a31d1 Record the measured confidence-gate numbers 2026-07-31 23:23:37 +04:00
kami 9b8bdf73cc Merge the five small-task branches 2026-07-31 23:11:26 +04:00
kami 7ad3c9a408 Merge branch 'worktree-agent-af88d63f65f30896b' into integration/small-batch 2026-07-31 23:11:16 +04:00
kami 84e1478823 Merge branch 'worktree-agent-af0fd9507d3e2ee46' into integration/small-batch 2026-07-31 23:11:16 +04:00
kami 9c8d0baffe Merge branch 'worktree-agent-a4cef2a815e32ebbf' into integration/small-batch 2026-07-31 23:11:16 +04:00
kami b0f5a16ec9 Add digest as a real outcome: suppressed care nudges get resurfaced, not lost
Vikunja #281. The interruption policy promised four outcomes — deliver_now,
queue, digest, drop — but only three existed: a care candidate the restraint
gate suppressed for quiet hours / away / calendar-busy simply vanished in
loop.Tick's `continue`, with only the trace remembering why.

internal/morning turned out not to be the natural drain: it's a fixed
Item/FactKey checklist engine, not a generic message bundler, so gate-
suppressed nudge text has nowhere to plug into its evidence model. Built a
parallel (but small, reusing the outbox's shape) durable digest instead:

- internal/store: digest_entries table + EnqueueDigestEntry (dedupes by
  rule+body, mirroring the delivery outbox's bodyHash), PendingDigestEntries,
  ExpireStaleDigestEntries, DrainDigestEntries (mark, never delete — an
  audit trail of what she actually said).
- internal/loop: DigestEligible(severity, blockedBy) is the pure boundary —
  only genuine restraint blocks (quiet_hours/calendar_busy/presence) even
  qualify (cooldown/snooze are not "suppression"); within care, Sev2 (break)
  digests, Sev1 (water/meal — stale by the time anyone could resurface them)
  drops. High severity never digests; alarms bypass the gate and deliver
  unchanged, on purpose.
- cmd/mavend/tick.go: each tick scans ExplainTick's trace for eligible
  blocked candidates, enqueues them, sweeps stale entries (24h expiry — the
  care rules are daily-cadence, so anything older is describing a day
  that's over), and drains the bundle only once the suppression reason has
  actually cleared, capped at 3 spoken items plus a trailing count so a
  digest can't turn into the exact nagging it was built to avoid.

Tests: store-level round-trip/restart-survival/dedupe/expiry/drain, loop-
level severity-boundary unit tests, and tick-level integration tests for
the drain-only-when-clear and never-digest-high-severity behavior.
2026-07-31 23:09:22 +04:00
kami 67563ed1f6 Run pattern detection from the digestion tick, not just voice (#43)
detectPattern only ever fired as a side effect of a voice fact-write, so a
recurring pattern already sitting in history went unnoticed until he
happened to mention it again by voice — the opposite of proactive.

Split the pipeline: extraction (fact -> normalized event) stays where a fact
is written, in voice.go, since it's tied to that write regardless of who's
talking. Detection (events -> stable pattern -> proposed_routines row) moves
into shared code (patterns.go's detectAndPropose) that both the voice path
and the new tick.go:detectPatterns call. The tick runs it every cycle over
every action+object pair on record (store.DistinctEventPairs, added), so a
pattern gets noticed on the daemon's own schedule.

Idempotence and the dismiss-must-stick requirement turned out to already be
handled by the store, not something the tick needs to reinvent:
proposed_routines has UNIQUE(action, object) and CreateProposedRoutine does
ON CONFLICT DO NOTHING, and DismissProposedRoutine flips status in place
without deleting the row. So a pair already proposed, accepted, OR
dismissed is a silent no-op on every later tick — a dismissed pattern can
never resurface, and re-running the scan never spams the /routines page.
Kept the voice-path call (immediate spoken confirmation is a nice feature
UX-wise and is now redundant-but-harmless with the tick, since both paths
share the same guarded detectAndPropose).

Tick-side detection only ever writes a row; it does not notify, ring, or
speak, keeping Maven "not a nag, not autonomous" — the /routines page is
still the only place a proposal becomes visible, and only accepting it
starts producing nudges (fireAcceptedRoutines).

Also fixed the stale vikunja#46 reference in proposed_routines.go — the
TODO it named is what this commit does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:07:33 +04:00
kami f0f7ebc9b2 Give LLM-routed decisions a real confidence so clarify can fire (#359)
Confidence was hardcoded to 1.0 for every LLM decision, and the LLM branch
in Router.Route returned straight from fillSlots without ever touching the
stage-3 threshold gate — so the LLM path could not produce a Clarify no
matter what confidence a model reported. That is why all 6 want_clarify
cases in the 77-case RU fixture were missed by every model in the bake-off.

Fix reads structural signal instead of changing the (parity-locked) router
prompt: a single-token utterance ("вода", "бэкап") is flagged thin evidence
in llmrouter.go; a fact left keyless or an act that never resolves to an
allowlisted fn, checked after fillSlots so the deterministic parsers get
first crack, is flagged in router.go's new gateLLMDecision. Anything below
config.DefaultRouterThreshold (0.55) now sets Clarify=true through the same
path the classifier already uses.

Added unit tests with a stubbed Completer proving both directions: thin
cases clarify, clean multi-word/resolved-slot cases stay confident. The
77-case fixture re-run against a live llama-server is still needed to
confirm the 6/6 moves — not done here, no llama-server on this box.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:07:32 +04:00
kami d12de589a2 mavweb: default -addr to loopback, not all interfaces
PR #47 added two state-changing routes (POST /chat, POST /routines)
behind the -addr flag, which defaulted to ":9200" (all interfaces).
Default now binds 127.0.0.1:9200; anyone who wants LAN/wider exposure
still passes an explicit bind (as deploy/docker-compose.yml already
does with "-addr :9201" inside the container, unaffected by this
default change).

Vikunja #317.
2026-07-31 23:03:45 +04:00
kami 50cc17f33a Lock down deploy/ecosystem/nginx.conf template to match the live host
The template said "drop into your nginx sites" but listened on the
wildcard `listen 80;` with no allow/deny ACL, unlike the actual deployed
hexis.kvmx.ru config which binds only to the WireGuard (10.42.0.1) and
LAN (192.168.1.104) addresses with allow/deny all. Anyone following the
template as written would expose these unauthenticated admin UIs to the
open internet.

Bind explicitly to those two addresses and add the matching ACL block,
mirroring cmd/mavweb/nginx.conf which already does this correctly.
Added a comment naming both addresses as host-specific so a deploy on a
different box swaps the IPs instead of reverting to `listen 80` when the
bind fails.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:01:58 +04:00
kami 73d13f1ea6 Merge pull request 'Stop the docs claiming the LLM router is off' (#49) from docs/fix-drift into master 2026-07-31 20:46:57 +02:00
kami 0ca5748699 Stop the docs claiming the LLM router is off
CLAUDE.md's routing section said "llmrouter is wired nil" and called the
classifier cascade the committed default. That stopped being true when the
integration merge landed: voice.go:214 wires pickLLMRouter, DefaultLLMRouter is
on, and deploy/mavend.json sets llm_router true. It is the first thing anyone
reads before touching the router, so it was pointing the next reader at a
wiring job that is already done.

Rewritten to say the LLM router is the default, the classifier is the failure
floor and must not be deleted, and what the two actually measure — 36.8% at
p50 31ms against 67.5%/72.7% at p50 ~2.7s, a trade accepted on purpose. Names
the one thing still open on that path: Confidence is hardcoded 1.0 in
llmrouter.go, so the LLM never asks for clarification (#359).

Also in CLAUDE.md: the persona line pointed at a memory file that does not
exist, so the actual rule was nowhere in the repo. Written out instead —
feminine self-reference, informal singular address, pet names forbidden but his
name allowed — plus the three eval checks that enforce it.

MODEL-BAKEOFF: three claims had gone stale within hours of being written. There
IS a make eval-models target now; the routing numbers ARE the production path,
not a bench artifact waiting on a wiring change; and the truncated 293 MB gguf
is deleted. Struck through rather than removed, since the caveats are part of
how the evening read at the time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 22:37:15 +04:00
kami b43bb265b5 Say it the way she would actually say it (#48) 2026-07-31 20:17:27 +02:00
kami b9a24334ea Say it the way she'd actually say it
Wording fixes from the review of the clarify + phrasing PRs.

- "На когда напомнить?" → "Когда?". After she has just been asked something,
  the long form is the phrasing of a form field, not of a person.
- A reminder now wants a subject as well as a time. "напомни в 11" had a time
  and nothing to say at 11, and she asked nothing at all — she now asks
  "О чём напомнить?". Subject first, since a reminder with no subject is not
  worth setting.
- The expiry notice is five phrasings picked at random instead of one fixed
  sentence. It is the line he hears every time he walks off mid-request, so it
  is the line that repeats most.
- The nudge prompt's ban on "обращения" is now "ласковые обращения". It was
  meant to forbid "милый"/"дорогой", not his name — "Ками, ноутбук на трёх
  процентах" is how she talks, and the eval's cringe check already only flags
  pet names.
- The nudge example no longer claims she plugged the laptop in. She has no
  hands and no smart plug; an example where she acts teaches the model to
  invent actions Maven never took.
- replySystem: "тепло" → "спокойно и без официальных формулировок". A one-word
  mood instruction a 1.7B can't act on, replaced with the behaviour meant.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 21:52:12 +04:00
kami c97aebf55a Merge tonight's work: all 46 reviewed PRs as one verified branch
135 commits. make build produces all 8 binaries; make test exits 0 across 38 packages with no failures, no data races, gofmt and vet clean.

See PR #47 for what had to be fixed to make it build as a unit.
2026-07-31 19:41:52 +02:00
kami 891136c65d gofmt the kiwix client and rewrite test
PR #41 and #44 landed these two files unformatted, so the gofmt gate that
PR #12 added to `make test` failed as soon as both were on one branch.
Struct-tag and comment alignment only, no semantic change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 21:34:03 +04:00
kami 41c7c13f42 Merge remote-tracking branch 'origin/overnight/snooze-works' into integration/jul31
# Conflicts:
#	internal/store/migrations.go
2026-07-31 21:32:18 +04:00
kami a324e8f624 Merge remote-tracking branch 'origin/overnight/eval-writeup' into integration/jul31 2026-07-31 21:31:36 +04:00
kami 51805e7f35 Merge remote-tracking branch 'origin/overnight/kiwix-rewrite' into integration/jul31 2026-07-31 21:31:36 +04:00
kami 533f0acda8 Lead the bake-off with the answer, not the superseded one
The file ran two sweeps and the second one changed the resident model, but
the lede still opened with "Recommendation: keep Qwen3.5-0.8B". Anyone
landing on the file read the wrong conclusion and had to scroll 100 lines
to find that it had been replaced — and it contradicted CLAUDE.md, which
already says the resident model is Qwen3-1.7B.

Both sweeps are accurate, so nothing is rewritten. The lede now states the
outcome and the first sweep's verdict is scoped to what it actually tested:
it rejects LFM2.5-1.2B, which still holds. It never was a case for keeping
0.8B as the resident model.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 21:17:07 +04:00
kami 4f59ba78c6 Write down the five-model sweep and why the 1.7B won
Numbers behind the resident-model change, plus the answer to "could a 230-350M
model do this instead" — no, and the reason is worth keeping: LFM2.5's published
instruction-following scores beat Qwen3.5-0.8B, and every one of those benchmarks
except Multi-IF is English. In Russian the 350M invents non-words and the 230M
answers in Spanish.

Also fills the row TALK-EVAL-31-07-2026.md had to void for contamination, and
corrects a wrong call I nearly made: the 1.7B's 16s p95 looked like the reasoning
trace, but the 0.8B sits at 17s in every run and the 1.7B beat it twice out of
three. The long tail is shared and is not the Thinking block.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 19:08:37 +04:00
kami d0afd9d4f6 Make Qwen3-1.7B the resident model
Stock Qwen3-1.7B, not the CPT'd one — that training is still running. It won
on both fixtures we have, measured tonight on an otherwise idle box:

  routing, 77 RU cases, intent-only:  67.5%  vs  59.7%  for Qwen3.5-0.8B
  talk fixture, 27 cases:             20/27  vs  11-17/27

It also beat Qwen3.5-2B, which is 20% larger, on every routing column.

Two other things came with it:

n_ctx goes 2048 -> 4096. This is a Thinking variant, so reasoning tokens need
the room, and 4096 is the context every score above was measured at. Shipping
2048 would ship something nobody measured.

The doc now says not to bother with sub-500M models, because I checked and they
are not close. LFM2.5-350M routes at 5.2% — worse than guessing among 7 intents
— and answers "столица Франции?" with "Сторзит", which is not a word. The 230M
replies to Russian in Spanish. Their published IFEval and BFCL numbers are good
and they are all English.

Note the routing gain needs the LLM router actually wired on to show up. It is
still nil, so this commit buys the phrasing improvement today and the routing
improvement when that lands.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:58:20 +04:00
kami 6b67e6f3c2 Word nudges from templates by default, model optional
DEPRECATION, flagged not asked: LLM-phrased nudges are no longer the default.
LLMPhraser.PhraseNudge now returns a hand-written Russian template. The model
still phrases chat, queries and reminders — only nudges moved.

Why: measured over many runs, Qwen3.5-0.8B wrote formal "вы" and plural
imperatives, used masculine self-reference, and invented facts and units
(90-95 seconds to boil an egg). A nudge is five words of known content, so
generation buys nothing and risks the persona every time. Templates score
15/15 on the nudge fixture, the model 11-13/15.

Nothing is deleted: the prompt, the fallbacks and the whole LLM nudge path
stay. Set phraser.llm_nudges = true in deploy/mavend.json to get them back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:21:52 +04:00
kami 742b2ad1d7 Score the Russian-to-keywords rewrite end to end (#403)
Same 9 cases as the retrieval eval, so the numbers compare directly:
hand-written keywords hit 8 of 8, this is what the model reaches on its
own. Reports the hand-written query next to the model's for every case,
because where the phrasing differs is the useful part.

Opt-in on MAVEN_KIWIX_URL + MAVEN_LLM_URL, like the other evals.

Result on Qwen3.5-0.8B: 3 of 8, identical on all three runs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:19:56 +04:00
kami b300ac5c70 Rewrite a Russian question into English Kiwix keywords (#403)
Kiwix ranks by keyword, not meaning, so a translated question finds song
and TV titles. This asks the resident model for the TOPIC instead: a short
English noun phrase, like a Wikipedia article title.

Locked down three ways, because a wrong query is silently wrong:
- A GBNF grammar, same idea as routeGrammar and responseGrammar. The
  reply must be {"query":"..."} with Latin words only. The JSON wrapper
  matters: this model always thinks out loud and this llama-server build
  ignores the thinking switch, so a bare word-list grammar just captured
  "Let me analyze this request carefully" for every question.
- max_tokens 32, since the answer is a few words.
- CleanQuery, which throws away empty, Russian and prose replies rather
  than passing them to Kiwix, and drops question words like "why" and
  "how much" that a keyword ranker cannot use anyway.

Client side only. Nothing is wired into the daemon or any config.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:19:42 +04:00
kami 13e5170e9e Hand-written Russian nudge templates plus a picker
Nudge wording as data instead of generation. The wording lives in
internal/phraser/nudges_ru_v1.json (embedded), about 10 variants per rule:
water, meal, break, service_down, netdata_critical, routine:, morning:, plus
a contentless default. That JSON is long because it is data — the owner can
edit any line of Russian without touching Go.

The picker:
- random, but never the same variant twice in a row for the same rule
- deterministic when seeded (math/rand with an injectable source)
- fills {since} / {service} / {what} from the candidate, and skips any variant
  whose value is missing, so no raw placeholder can reach the piper voice
- {since} is spelled out in words ("полтора часа", "семь часов"), because
  "3 ч" is wrong in a Russian voice

Scores 15/15 on the existing nudge fixture, on every seed swept. Nothing is
wired yet — that is the next commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:18:55 +04:00
kami 0b90952e55 Write down every conversational eval score from tonight
Records all four configurations on the 27-case talk fixture, three runs
each: no grammar, plus grammar, plus Russian prompts, plus the truncation
fix. Composite, per-path and per-check, with the reproduce command.

The short version is that the plumbing got fixed and the score barely
moved. Grammar was the real win. Russian prompts helped a little and cut
latency by 5x. The truncation fix was necessary and bought nothing.

Also writes down three things that are easy to lose:

- The truncation cause was the grammar's 400-character bound, not the
  token cap. Measured at three caps, same 400 characters every time.
- Then I set the bound to 1000 against a 768-token cap and made it worse.
  The two limits have to agree.
- One run is contaminated and marked void: I ran an agent against the same
  llama-server, and the report still claimed zero errors while a third of
  the fixture silently answered "не знаю.". That is #397 and it is worse
  than filed — a busy server is indistinguishable from bad phrasing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:18:19 +04:00
kami aa8f5b2ee2 Make the nonempty check look for actual words
It scored 27/27 on a run where two replies were "{" and "{\n  \"". It only
tested that the string was not blank, so punctuation counted as content and
the worst replies of the run passed the first check.

Now a reply needs at least one letter, Cyrillic or Latin. Latin counts
because answers about ssd or vpn are legitimately part English.

Digits alone fail too. The same run answered "сколько варить яйцо
вкрутую?" with "15-16" — no unit, no words, and the wrong number as well.
That is not something she said.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:57:39 +04:00
kami d7cdcb63bd Stop shipping half-written JSON as a reply
Two bugs, one symptom. A run of the talk eval produced replies that were
literally "{" and "{\n  \"" — those strings went out as things Maven said.

First bug: the parser could not tell "the model answered in plain prose"
from "the model started a JSON object and got cut off". Both came back as
empty, and every caller then shipped the raw text. Now an unfinished object
returns an error and each caller uses its own fallback instead. Bare prose
with no JSON in it still passes through, because small models do sometimes
answer that way and the reply is fine.

Second bug, and the actual cause: the grammar capped the response field at
400 characters. I measured it against Qwen3.5-0.8B at three different token
caps — 256, 768 and 2048 — and the reply came back exactly 400 characters
every time, cut mid-word. So the token limit was never what stopped it.
The bound is 1000 now, about six Russian sentences, still low enough to cut
off a repetition loop.

Token caps go from 256 to 768 on the chat and query paths so 1000
characters of Russian actually fits. The nudge path keeps its own cap; a
nudge is meant to be one sentence.

Note: cmd/mavend/replier_llm.go has its own copy of this parser with the
same bug. Left alone here so this commit stays small — that duplicate is
Vikunja #396.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:56:55 +04:00
kami ddb658ffbb Add a Kiwix search client and score retrieval (Vikunja #403)
Step one of letting Maven read instead of recall. No LLM yet.

internal/kiwix/client.go: search a local Kiwix server, parse the RSS
reply, hand back title + path + plain-text snippet + word count. The
snippet is the unit of context; a full article is ~100KB of HTML and
will not fit a 4096 token window.

internal/kiwix/retrieval_eval.go plus knowledge_v1.json: the 9 knowledge
questions from the phrasing fixture, each with hand-written English
keywords, scored on whether a wanted article comes back in the top 5.
Opt-in via MAVEN_KIWIX_URL, since CI has no Kiwix. No pass bar, the
number is the finding.

Result on the live mirror: 8/8 answerable questions hit, 7 of them at
rank 1. Retrieval works. Keywords are written by hand on purpose, since
Kiwix ranks by keyword and not by meaning, so a natural question fails.
A query-rewrite step is the next piece of work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:43:44 +04:00
kami c7dadc97d9 Write the chat and notes prompts in Russian
The reply has to be Russian, but two of the phrasing prompts told her
what to do in English. Both are Russian now, in the same style as the
nudge prompt that already works better.

Also dropped the "you are maven, a self-hosted personal assistant"
line from both. The persona block right above it already says who she
is, so it was said twice.

The JSON part is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:38:36 +04:00
kami c9d88c152e Drop "never phones home" as a hard rule
The owner's call, 2026-07-31: a 0.8B model does not know enough about the
world to be useful without reading something. So she may now read external
sources to answer world questions.

What replaces the old rule, in all three docs:

- No telemetry, no cloud model, no third-party account. Unchanged.
- Local first: the Kiwix ZIMs on the box before anything on the network.
- External search is allowed but off unless configured, same as weather
  and telegram.
- His notes and facts are never search input. Only the utterance goes out
  — never the persona block, the history, or matched notes.

Docs only, no code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:34:53 +04:00
kami 1890ff5d5d Constrain the phrasing output with a GBNF grammar
The 0.8B answered about one chat turn in three with open reasoning as plain text, so no JSON ever closed and the fallback shipped "Thinking Process:" to the user. A grammar makes that output impossible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:18:06 +04:00
kami 0110e9bc8c Report every address break, and stop -те verbs blinding the check
From a real reply in a nudge eval run: "Смотрите на его потребление
воды" is a plural imperative AND third person about him. Only the plural
printed.

Two separate faults. The check returned on its first hit, so the second
break stayed invisible and the failure read as milder than it was; it now
joins them. And "его" was not detected at all — looksVerb knows the
-й/-йте imperative but not the -те plural, so "смотрите" counted as the
person being talked about, which is what an antecedent means here.
pluralVerb already knows that form, so the antecedent test uses it too.

Third time a verb form has blinded this check. A fourth means it wants a
morphology table rather than another suffix.
2026-07-31 16:52:16 +04:00
kami 50ca8c8b5a Score the chat, query and knowledge phrasing paths (#395)
The phrasing fixture was 15 nudge cases, so every prompt change we
measured only told us about nudges. But the shared context block sits in
front of five prompts, and three of them — chat, note query, general
knowledge — had no scorer at all. Those are the long free-form replies,
where a persona break is most likely and where nothing could see one.

27 cases, nine per path. Nine rather than five because the nudge fixture
already cannot resolve a change smaller than about three cases, and a
per-path score off five would be worse.

Reuses the persona checks instead of copying them. Length, mood and
"no questions" are left out on purpose: these paths return no mood, and
a follow-up question is a feature in chat, not a fault.

The run refuses to score unless the model answers before and after it.
PhraseChat and PhraseQuery swallow model errors and return a canned
string, so without that guard a dead server produces a full report with
zero errors and a bad score — which reads as bad phrasing rather than as
nothing measured. Vikunja #397 is the real fix.
2026-07-31 16:51:52 +04:00
kami de09471421 Merge the shared prompt context block 2026-07-31 16:07:35 +04:00
kami d65c16a567 Don't tell her she can't talk
The block listed what she can do and ended with "nothing else". It sits
in front of the chat and general-knowledge prompts too, so that told her
to refuse the exact thing those prompts are for. Talking is now first in
the list, and the closing line limits ACTIONS rather than everything.

Also dropped the self-introduction from the knowledge prompt. It said
"Мавена, персональный ассистент" — a different name and a masculine
noun, right after the block says she is Maven and feminine. Identity
lives in the block now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 16:07:35 +04:00
kami 062d4252ef Tell her what she can actually do
The context block now lists her real capabilities: reminders, notes and
facts (write and recall), and the calendar — all three are code paths in
mavend today. Weather, telegram and shell acts are listed only when the
config actually has them, because offering something she cannot do is
worse than staying quiet about it.

Also drops the pronouns from the optional name/city line. The block's
own "ты" is Maven, so "тебя зовут" read as her name and "его" would have
shown her the third-person form she must never use about him. They are
plain labels now.

Vikunja #394.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 15:57:58 +04:00
kami 2c27e2ce1f Give every prompt one shared context block
The "address him as ты" rule had only reached two of the five system
prompts. Instead of pasting it into the other three (five copies drift —
that is how this happened), there is now one block, in internal/persona,
prepended to all five: nudges, action replies, chat, note queries and
general knowledge.

The block says who he is and how to address him (a man, always "ты",
never "вы", never "он" about him; Maven stays feminine), plus the
current local date and time. It is rendered fresh each turn because the
time changes, and it is correct with an empty config — the address and
gender rules are defaults in code. Config only adds optional facts:
owner_name, city, and the existing free-text `persona` string, which is
now the static half of the block.

Russian even in front of the English prompts: the rules are Russian
grammar, so they read best stated in Russian, and there is one copy.

Vikunja #394.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 15:55:30 +04:00
kami ccc5cba2a3 Merge the example-led prompt finding 2026-07-31 15:41:01 +04:00
kami 89d83c0b11 Record the example-led nudge prompt experiment (#393) — it made things worse
Tried rewriting the nudge prompt to lead with five on-topic examples instead
of rules. Three eval runs each side: before 12/13/14 of 15, after 11/12/11.
The loss is all in the address check — formal "вы" and plural imperatives came
back once the "говоришь на ты" rule stopped being its own sentence, and the
on-topic examples leaked their wording into the wrong cases.

Prompt reverted. Only the finding is committed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 15:40:18 +04:00
kami a97f554802 Merge the informal address prompt rule 2026-07-31 14:54:20 +04:00
kami f4de2fc5e1 Don't let a verb count as the person being talked about
The third-person check asks whether anyone else was named before "он".
A nudge is mostly verbs, and they were counted as possible people, so
"попробуй встать и отдохнуть — у него есть перерыв" passed. Infinitives
and imperatives now join past tense as words that cannot be a person.

A plain noun before the pronoun still blinds it. That needs a parser,
and the comment says so.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:54:20 +04:00
kami eef5d4da4f Tell the phraser to speak to him informally, singular
The prompts stated the feminine self-reference rule but never said whom she is
speaking to, so the model produced formal plural ("Жду вас") and talked about
him in third person ("Он не ел 11 дней"). Adds the address rule right next to
the feminine one, in the nudge prompt and the confirmation prompt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:52:23 +04:00
kami 09f1696fce Merge the eval label and kill script fixes 2026-07-31 14:32:45 +04:00
kami 80f7322294 Don't fail when docker confirms nothing is running
"Nothing on the host" meant two different things and the script treated
them the same. If docker answers and names no running containers, Maven
really is down and the script should say so and exit 0. Only when docker
cannot be asked is the answer unknown, and that is the case that must
fail loudly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:32:45 +04:00
kami fa5aebfbe4 Merge the delivery boundary fixes 2026-07-31 14:30:54 +04:00
kami 59cec63da1 List the columns in the table rebuild
The migration copied rows with SELECT *, which matches columns by
position. It is correct today, but if the old table's order ever
differed it would shuffle every row instead of failing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:30:54 +04:00
kami 02e8786695 Stop the containers instead of claiming success (#380)
docker-compose.yml has no 'pid: host', so each container has its own PID
namespace and pkill on the host matches nothing inside them. The script
then printed "All services gracefully stopped" while mavend, its
llama-server and the rest were still running.

Now it checks for running compose containers first and stops them with
docker compose. If it cannot ask docker and finds nothing to kill on the
host, or anything survives the kill, it says so and exits non-zero
instead of claiming success. The bare-metal path is unchanged apart from
verifying the SIGKILL actually worked, and no longer risks killing the
shell it was launched from.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:30:51 +04:00
kami 0272dc9d89 Record a suppressed care nudge instead of dropping it silently (#370)
Dropping a sev1-2 care nudge while you're away is right and still happens.
But it was a bare `continue`: no row, no log, so "she dropped it", "the gate
suppressed it" and "the rule never fired" all looked identical afterwards.

Adds a 'dropped' delivery status (migration #12 widens the CHECK constraint;
sqlite can't do that in place, so the table is rebuilt) and records the drop
as one delivery_attempts row plus a log line.

No nudges row for a drop: that table feeds the ignored_rate signal, and a
nudge nobody could see must not count as ignored.

TestVoiceNoSessionFallthroughLeavesOutboxTrail expected exactly one row for
sev1-2 when voice had no session. It now expects the voice failure plus the
drop, which is the point of the change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:27:47 +04:00
kami 2ad7635501 Merge the address-form eval check 2026-07-31 14:27:16 +04:00
kami 9949b309b1 Don't let a time word blind the third-person check
The check asks whether anyone else was named before "он". Time words
were not stoplisted, so "сегодня он не ел" read "сегодня" as the person
being talked about and passed — which is the recorded break with a word
in front of it, and nudges open with those words constantly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:27:08 +04:00
kami a788ca3915 Label eval runs with the model the server actually loaded (#379)
The phrasing eval printed "llm (0.8B, ...)" no matter which gguf
llama-server had loaded, so two runs of two different models came out
named the same and were easy to mix up when comparing.

It now asks llama-server over /v1/models, same as the router eval
already did. The helper moved to internal/llm so both share it, and it
now errors instead of returning a blank name when the id field is
missing — an unreachable server gets labelled "unknown-model", never a
plausible-looking guess.

Both eval paths stay opt-in behind MAVEN_LLM_URL; no server needed for
go test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:25:28 +04:00
kami 62d47d28ac Add an eval check for formal and third-person address (#384)
The phrasing run produced two persona breaks that scored clean:
"Приходите… Жду вас" (formal plural) and "Он не ел 11 дней" (talks
about him instead of to him). She is feminine, he is male, and she
speaks to him informally, one to one.

The new `address` check flags the "вы" family, plural imperative
endings, and a third-person "он" with no other subject named earlier in
the message. Like `hisgender` it is a keyword/suffix heuristic, not a
parser, and it prints the word it tripped on so a false alarm is easy to
dismiss. Limits are written out in the comment.

Both recorded strings are pinned as unit tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:25:23 +04:00
kami e9ff2c4912 Never send the full nudge body off-box (#368)
The away sinks fell back to the whole Body when Summary was empty. ntfy and
telegram leave the box, and the 0.8B phraser drops fields regularly, so that
fallback could push full detail off the machine.

The dispatcher already strips detail from away sendables. This exports that
one rule as delivery.AwayMessage and has both sinks use it, so a sink can't
leak the body on its own either: empty Summary means a generic line plus the
rule name, never the body.

The two sink tests named TestSendFallsBackToBodyWhenSummaryEmpty asserted the
old, wrong behaviour, so they are rewritten to assert the generic line.
TestSendRejectsEmptyMessage is likewise replaced: an away message can no
longer be empty, so the sink has nothing left to reject.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:23:54 +04:00
kami 3dbf67f8f9 Drop the city time-zone table
The user only ever asks the time in his own zone, so answering other
cities was code kept in step with the weather city list for no gain.
Any named place now gets the honest "local time only" answer that was
already there for unknown cities.

Removes the 22-entry table, the lookup and the embedded tz database.
Closes Vikunja #389 — there is only one city list again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:14:18 +04:00
kami 84ba217892 Say so when the day asked about is out of reach 2026-07-31 14:05:21 +04:00
kami f179ae2fde Merge the system reply fixes 2026-07-31 14:03:42 +04:00
kami d00929ac0b Answer the day the user asked about and the city he named (#388)
replySystem had two arms that PR 30 made reachable, and both answered confidently wrong: the date arm keyword-matched "числ" and always answered today, so "какое число завтра" answered today; the clock arm ignored a named city and answered local time. The date arm now reads the day word through router.ParseCalendarDate (which grew послезавтра/вчера and now cuts the day boundary in the local zone instead of UTC). The clock arm answers the named zone when it resolves offline from the tz database embedded in the binary, and otherwise says plainly that she only knows local time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:02:57 +04:00
kami b6f47fbeb6 Merge the clock and calendar routing rule 2026-07-31 13:53:35 +04:00
kami 2e9b9ec1cf Warn separately when a row has no text to re-embed 2026-07-31 13:52:56 +04:00
kami bfb57c3148 Give the router prompt a rule for clock and calendar questions (#374)
The prompt named seven intents but never said which one a clock or date
question belongs to, so the model guessed: system->query x4 in every eval
run. The rule now says the clock and the calendar date themselves are
system, what is written in the calendar or in memory stays query, and a
time named inside a request is just part of the request.

That split follows what the daemon can answer. Only replySystem owns the
clock and the date formatter, while the agenda is answered from
CalendarEvents inside the query branch.

Also adds one calendar-agenda fixture case so an over-broad system rule
cannot pass unnoticed, and writes up the before/after numbers. The
targeted confusion is gone; the headline accuracy did not move.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:52:00 +04:00
kami d1f6f6355f Merge the vector backfill 2026-07-31 13:51:25 +04:00
kami 92ecb691de Re-embed stored notes and facts after an embedder swap (#378)
The embedder swap left every stored vector in the old model's space, so cosine against a new query vector is noise. Add the one-shot backfill: store.ReembedAll re-embeds every note and fact text with the currently configured embedder (the passage side, which is the side stored text was written with) and rewrites both places a vector lives — the notes table embedding column and the memory_vectors rows.

All of it plus the embedder marker happens in one transaction, so a failure partway changes nothing and writes no marker: re-run it. A run against a DB whose marker already names the current embedder does nothing.

Triggered explicitly with `mavend -reembed`, not automatically on mismatch: ONNX on the laptop CPU makes this minutes of work, and a silent multi-minute stall on boot would look like a hang. The mismatch warning now tells the user to run it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:50:43 +04:00
kami 4282f6b9a9 Warn about vectors written before the marker existed 2026-07-31 13:44:29 +04:00
kami 7bb9f9be06 Merge the embedder marker 2026-07-31 13:42:40 +04:00
kami 1e47eaca5a Record which embedder wrote the stored vectors and warn on a swap (#378)
The embedder moved from paraphrase-multilingual-MiniLM-L12-v2 to
multilingual-e5-small. Both are 384-dimensional, so nothing in the code
noticed: cosine between an old stored vector and a new query vector is
noise, and recall degrades silently.

So the DB now records the embedder that wrote its vectors. One value for
the whole DB (migration #11, a small `meta` key/value table) rather than a
column on every vector row: the backfill re-embeds every note and fact in
one pass, so a per-row marker would hold the same string in every row and
cost a column on two tables for nothing.

The identity comes from the embedder itself via a new optional ID() method
("multilingual-e5-small@384", model file name plus dimension), so pointing
the config at another model changes the string without anyone editing a
constant. mavend logs a loud WARNING at startup naming both the stored and
the configured embedder when they differ.

Detection only — recall behaviour is unchanged. TODO(#378) in
store.CheckEmbedder marks where the backfill will hook in.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:42:07 +04:00
kami 892330eb84 Merge the note recall fix 2026-07-31 13:33:14 +04:00
kami 9a3bcd7c46 Merge the thinking-off measurement 2026-07-31 13:31:31 +04:00
kami 98ee701e03 Let a note win a recall, not only a fact (#373)
The memory pass ran only after the notes-only gate had already rejected
the same note at the same score. Notes and facts share one vector index,
so a note that failed there failed again — the branch could only ever
return a fact.

Now the memory pass runs first: one search over everything Maven
remembers, one gate, and the memory that clearly matches best answers
(a note gets phrased, a fact is read back). The notes-only pass stays
behind it for notes the vector index does not hold. No threshold moved,
so the set of questions answered is unchanged — only which memory
answers them.

Fixture gained two mixed note+fact cases, so the answerable count goes
25 -> 27: hash recall@1 36.0% -> 37.0% (ratchet 0.32 unchanged, comment
updated), e5 recall@1 72.0% -> 70.4%, false recall still 1/5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:30:38 +04:00
kami 04c1088088 Measure thinking off on routing properly — it does not win (#376)
The 67.1% "thinking off" column in ROUTING-EVAL-31-07-2026.md was an
artefact. It came from a hand-rolled HTTP client in the eval test that
did not send repeat_penalty, so it differed from the reference run on two
axes and the penalty was the one that mattered.

Re-scored back to back on an idle box with everything else held equal:
thinking off is identical to thinking on, case for case, same confusion
matrix, same three unparseable replies. A direct probe of the running
llama-server shows enable_thinking, thinking and reasoning_budget are all
ignored for this model on this build, so there was nothing to turn off.

No defaults changed. The misleading third configuration is removed from
internal/router/eval so its table cannot be quoted again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:30:22 +04:00
kami 07c191d8b8 Merge dialogue session persistence 2026-07-31 13:21:03 +04:00
kami c668310b3e Persist the dialogue session so a restart keeps the conversation
Vikunja #363. The follow-up session was a plain in-memory map, so any
mavend restart dropped the thread. It now mirrors to a small TTL-pruned
sqlite table and is loaded on startup; expired sessions are deleted on
load, not revived. Clarify's pending question is untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:20:30 +04:00
kami 1bd2acdc2a Do not exempt Russian words that are both noun and verb 2026-07-31 12:55:45 +04:00
kami 15e5dd8eaa Merge the second-person gender check 2026-07-31 12:54:48 +04:00
kami 10cf6f525c Check that nudges do not address the owner in the feminine
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:54:18 +04:00
kami ee3e6a9eaf Wire the snooze read into the Gatherer and honour it for reminders (#364)
The Gatherer now fills State.SnoozeUntil from store.SnoozedUntil instead
of nil, so a snooze finally reaches the gate. RemindDecisions gains the
one restraint check that applies to a reminder — quiet hours, presence
and cooldown are still bypassed, so "wake me 7" is unchanged. Reviewer:
the two tests in internal/loop/gate_test.go are the contract.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:33:51 +04:00
kami 8acb8a97c6 Read the recorded snooze outcomes back out of the nudges table (#364)
The gate honours State.SnoozeUntil but nothing ever filled it. New
store.SnoozedUntil returns, per rule, when the newest snooze runs out.
Reviewer: the fixed 2h SnoozeDuration and its reasoning in nudges.go —
nothing upstream can supply a per-nudge length, so no new column.
Expired snoozes are dropped in SQL, so silence can never be permanent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:33:42 +04:00
116 changed files with 9207 additions and 2300 deletions
+1
View File
@@ -6,6 +6,7 @@
/mavweb
/mavpoll
/mavcaldav
/mavwaked
# Certs (private keys, don't commit)
certs/
+64 -13
View File
@@ -7,9 +7,20 @@ talking over unix sockets; one resident small model for routing + phrasing; whis
Deploy target is a Ryzen laptop (homesrv) with Vulkan offload to the Vega iGPU (`n_gpu_layers: 99`,
compose passes `/dev/dri` + the render gid) — the resident model stays ≤1.7B either way.
**Resident model:** currently **Qwen3.5-0.8B** (`Q4_K_M`), the smallest checkpoint in the gguf
library, picked for CPU/iGPU latency. The **target** is the locally CPT'd **Qwen3-1.7B**; that
training is still in flight (Vikunja #122), so no such gguf exists yet. Model files live in
**Resident model:** currently **Qwen3-1.7B** (`UD-Q4_K_XL`), stock — not yet the CPT'd one.
It replaced Qwen3.5-0.8B on 2026-07-31 because it measured better on both fixtures we have:
67.5% vs 59.7% intent-only on the 77-case RU routing fixture, and 20/27 vs 11-17/27 on the
talk fixture. See `MODEL-BAKEOFF-31-07-2026.md`. It is a Thinking variant, so `n_ctx` is 4096
— reasoning tokens need the room, and 4096 is what the scores above were measured at.
The **target** is still the locally CPT'd **Qwen3-1.7B** (Vikunja #122, training in flight).
Stock already speaks good Russian; what it gets wrong is the persona — it writes `я рад`,
masculine, where Maven needs `рада`. That is what the CPT is for.
**Do not bother with sub-500M models.** LFM2.5-230M and 350M were measured on 2026-07-31 and
both are unusable in Russian: the 350M routes at 5.2% (worse than guessing) and answers
"столица Франции?" with the invented non-word "Сторзит"; the 230M replies to Russian in
Spanish. Their strong published IFEval/BFCL numbers are English-only. Model files live in
`/mnt/hdd1/llms`, bind-mounted to `/opt/maven/models/llm` — which **shadows** the repo's
`models/llm/`, so the LFM2.5 gguf sitting there is not loaded by anything. Swapping the resident
model is a one-line change to `phraser.model_path` in `deploy/mavend.json`.
@@ -58,19 +69,41 @@ protocol; the config in `deploy/mavend.json` (with `${VAR}` env expansion from g
## Routing — read this before touching the router
`internal/router/` has TWO layered engines and the committed default is an **interim
stopgap, not the intended design** (see memory `routing-architecture-target`):
`internal/router/` has TWO layered engines. **The LLM router is now the default and it is
on in deploy** — this section used to say it was wired `nil`, which stopped being true on
2026-07-31.
- **Target (REARCH.md):** LLM-as-router. One resident Qwen3-1.7B (`llmrouter.go`) emits
GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is demoted
from a routing gate to a RAG hint.
- **Current stopgap:** `llmrouter` is wired `nil` (around `voice.go`), so the
`classifier.go` + `embedder.go` nearest-neighbour cascade actually runs. It routes by
similarity to frozen seed phrases — the known cause of weak RU query handling.
- **LLM router (the intended design, REARCH.md):** the resident Qwen3-1.7B (`llmrouter.go`)
emits GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is
demoted from a routing gate to a RAG hint. Wired at `voice.go:214` via
`pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient)`; the flag is `voice.llm_router`
(`config.go`), `DefaultLLMRouter` is **on**, and `deploy/mavend.json` sets it `true`.
- **Classifier cascade (the failure floor, not dead code):** `classifier.go` +
`embedder.go` nearest-neighbour over frozen seed phrases. It runs when the LLM router is
off, when there is no llama-server to talk to (`pickLLMRouter` logs that and degrades),
and on any per-turn LLM error. Do not delete it — routing by seed similarity is the known
cause of weak RU query handling, but a turn must never break on the model.
Cascade order: `stage0.go` exact-match fast-path → LLM router (when non-nil) → classifier
fallback. Any LLM error falls through to the classifier so a turn never breaks on the model.
Measured on the 77-case RU fixture (`MODEL-BAKEOFF-31-07-2026.md`): the classifier scores
36.8% full accuracy at p50 31ms; Qwen3-1.7B scores 67.5% intent-only / 72.7% through the
cascade at p50 ≈2.7s. Accuracy roughly doubled, latency is ~90× worse, and that trade was
accepted deliberately. `Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
#359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with
no allowlisted fn) feeding the same stage-3 gate the classifier path already had — see
`gateLLMDecision` in `router.go`. Note the second half of that bug: the LLM branch never
consulted `r.threshold` at all, so a correct low confidence would have been discarded anyway.
Re-measured on the fixture after the fix: **missed clarify 6/6 → 1**, at the cost of 3 false
clarifies and 2.6pt of full accuracy (72.7% → 70.1%, intent-only 67.5% → 74.0%). Two of the
three false clarifies are acts the model mis-routed and the gate caught — asking beats wrongly
executing, so the fixture and the daemon disagree about what is correct there. The third,
`"поужинал"`, is a real defect: **the single-token rule is an English intuition and does not
transfer to Russian**, where one word is routinely a whole sentence. Narrow or drop it.
## LLM output contract
All phrasing paths emit `{"response":"...","mood":"..."}` (parsed in `replier_llm.go` and
@@ -82,8 +115,26 @@ workspace enforces that the Go and relabelling prompts remain identical.
## Non-goals (hard constraints)
Never phones home. Not a nag, not autonomous. Maven's persona is **feminine** — Russian
self-reference must use feminine forms (the user is male; see memory `maven-persona-gender`).
Not a nag, not autonomous. Maven's persona is **feminine** — Russian
self-reference must use feminine forms `рада`, not `рад`; `поняла`, not `понял`. The owner
is male and is addressed informally: "ты", singular, never "вы"/"ваш" and never "он"/"его"
(she talks TO him, not about him). Pet names ("милый", "дорогой") are forbidden; his name
("Ками") is not. The eval enforces this: `CheckAddress`, `CheckFeminine` and `CheckCringe` in
`internal/phraser/eval/checks.go`, scored by `make eval-phrasing`.
**"Never phones home" is DEPRECATED** (owner's call, 2026-07-31). It used to be a hard
constraint and it is not one any more: a 0.8B — and a 1.7B — does not know enough to answer
world questions, so she needs to read external sources. What replaces it:
- **No telemetry, no cloud model, no third-party account.** That part never changes. Nothing
about Maven is reported to anyone, and inference stays on the box.
- **Local sources first.** Kiwix ZIMs on homesrv (Wikipedia, ifixit) before anything on the
network. Reading beats recalling for a small model, and a local read costs nothing.
- **External search is allowed and off unless configured**, like the weather and telegram
capabilities.
- **His notes and facts are never search input.** Looking up why the sky is blue and sending
his stored personal notes to an upstream engine are different acts. Only the utterance goes
out, never the persona block, history, or matched notes.
## Web UI conventions
+10 -4
View File
@@ -15,7 +15,8 @@
**Maven** — self-hosted personal assistant. Manages your day, acts on your
homelab. One daemon on homesrv (always-on, not the workstation), multiple
client surfaces. All local, never phones home.
client surfaces. Inference and data stay on the box; she may READ external
sources (see Non-goals — "never phones home" is deprecated).
Primary name is "Maven", with feminine-gendered Russian self-reference
("она", "меня", "помогла"). Clients may choose their own UI label. Consistent
@@ -35,8 +36,13 @@ Inside boundary — the ones that actually constrain the build:
she records. A confident wrong fact is worse than a known gap.
- **Not a nag** — she'd rather miss a nudge than be mutable. Shuts up when
uncertain. Load-bearing.
- **Not a stranger** — runs on your stuff, your model, your data. Never
phones home.
- **Not a stranger** — runs on your stuff, your model, your data. No
telemetry, no cloud model, no third-party account. She may READ external
sources to answer world questions (Kiwix first, then optional search); she
never reports anything about you to anyone, and your notes and facts are
never used as search input. **"Never phones home" as an absolute is
deprecated** — owner's call, 2026-07-31: a small model does not know enough
to be useful without reading.
- **Not a relationship** — mom-tone is a function that makes nudges land, not
emotional company. Names the drift a warm small model falls into.
@@ -458,7 +464,7 @@ decides *insistence*. Both are needed.
sev ≤ 2 drops on away, sev ≥ 3 holds: a missed water nudge is noise, a missed
backup failure isn't. Away-channels (ntfy/telegram) leave the box — the one
path that crosses "never phones home," through your own relay. **Minimal
path that leaves the box for a person to see, through your own relay. **Minimal
body** — "disk low on homesrv," not detail; don't make notifications a
shoulder-surf exfil surface.
+126 -4
View File
@@ -1,15 +1,28 @@
# Resident model bake-off — 31-07-2026
**Recommendation: keep Qwen3.5-0.8B.** LFM2.5-1.2B is worse at routing (52.6% vs 60.5%
intent accuracy), and the loss is almost entirely Russian (18/61 vs 22/61 RU, while EN is a
wash). It is also 2.4× slower. The Thinking variant is far worse again.
**Outcome: the resident model is stock Qwen3-1.7B** (`UD-Q4_K_XL`). Two sweeps ran this
evening and the second one changed the answer — read to the end before acting on any table
here. [Second sweep](#second-sweep-same-evening--five-models-and-a-resident-model-change)
is the one that holds.
## First sweep — LFM2.5-1.2B vs Qwen3.5-0.8B
**Verdict, scoped to this pair: keep Qwen3.5-0.8B over LFM2.5-1.2B.** LFM2.5-1.2B is worse
at routing (52.6% vs 60.5% intent accuracy), and the loss is almost entirely Russian
(18/61 vs 22/61 RU, while EN is a wash). It is also 2.4× slower. The Thinking variant is
far worse again. This verdict still stands as written — it rejects LFM2.5-1.2B. It is
**not** a recommendation to keep 0.8B as the resident model; the second sweep replaced it
with Qwen3-1.7B.
Settles Vikunja **#278 / #250**.
- Same fixture and scorer as `ROUTING-EVAL-31-07-2026.md`: `internal/router/eval/`
(`ru_routing_v1.json`, 76 held-out cases).
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:<port> make eval-router`
(`TestLLMRouterBaseline`). Note: there is no `make eval-models` target.
(`TestLLMRouterBaseline`). (This line used to say there is no `make eval-models` target.
There is one now — start a server with the gguf you want, then
`make eval-models MAVEN_LLM_URL=http://127.0.0.1:<port>`. It runs only the LLM test, since
the classifier baselines do not depend on the model.)
- All three models served by the same `llama-server` flags — `-c 2048 -ngl 99 -t 6`, only
`-m` and `--port` differ. One server at a time on an otherwise idle box, so latencies are
real and not contention.
@@ -99,3 +112,112 @@ thinking trace costs time without buying accuracy on a short enum classification
Routing only. LFM2.5 might still phrase better, and phrasing is the resident model's other
job — that needs its own fixture. But routing is the load-bearing path and Maven is
Russian-first, so on the evidence here the switch is not worth making.
---
# Second sweep, same evening — five models, and a resident-model change
The sections above compared LFM2.5-1.2B against Qwen3.5-0.8B on routing and concluded
"the switch is not worth making". That still holds. This sweep asked a different
question — whether a *smaller* model could work, since LFM2.5's published
instruction-following scores beat Qwen3.5-0.8B badly — and answered it, plus found a
better resident model by accident.
**Outcome: the resident model is now stock Qwen3-1.7B.** Sub-500M is a dead end.
## Routing — 77 Russian cases, one run each
| model | on disk | llm-only (full) | llm-only (intent) | cascade + fallback |
|---|---|---|---|---|
| LFM2.5-230M-Q8_0 | 246 MB | 23.4% | 33.8% | 36.4% |
| LFM2.5-350M-Q8_0 | 379 MB | 2.6% | **5.2%** | 20.8% |
| Qwen3.5-0.8B-Q4_K_M | 527 MB | 36.4% | 59.7% | 61.0% |
| Qwen3.5-2B-UD-Q4_K_XL | 1.34 GB | 42.9% | 62.3% | 63.6% |
| **Qwen3-1.7B-UD-Q4_K_XL (stock)** | 1.13 GB | **44.2%** | **67.5%** | **72.7%** |
Qwen3-1.7B wins every column, including against a model 20% larger than it.
## Talk fixture — 27 cases, three runs each, idle box
| | Qwen3.5-0.8B | Qwen3-1.7B stock |
|---|---|---|
| composite | 13, 11, 8 | **20, 21, 18** |
| address | 21, 18, 18 | **26, 25, 23** |
| feminine | 27, 25, 26 | 26, 27, 26 |
| lang | 27, 27, 26 | 26, 27, 27 |
| ontopic | 16, 19, 19 | **22, 23, 23** |
| canned fallbacks | 8, 5, 6 | **0, 2, 0** |
This also fills the row `TALK-EVAL-31-07-2026.md` had to void for contamination:
**600ch/1024tok on Qwen3.5-0.8B scores 13, 11, 8.**
`address` is the headline. It sat at 18-22 of 27 on the 0.8B no matter how the prompt
was worded — the prompt explicitly forbids "вы" and the model writes `вашей`,
`подождите`, `делаете` anyway. That was read as "prompting is out of levers", and it
was really "0.8B is out of capacity". The 1.7B mostly holds the constraint.
The fallback column matters too: 5-8 of 27 turns on the 0.8B end in a hardcoded
`"не знаю."`, meaning it failed to emit parseable JSON about a quarter of the time.
The 1.7B does that 0-2 times.
## Latency — the long tail is not the Thinking block
| | p50 | p95 |
|---|---|---|
| Qwen3.5-0.8B | 2.4s, 2.9s, 2.0s | 17.4s, 17.6s, 17.4s |
| Qwen3-1.7B stock | 2.7s, 2.6s, 2.8s | 16.4s, 6.6s, 3.9s |
p50 is flat across a 2× size difference. The first instinct on seeing the 1.7B's
16s p95 was "that is the reasoning trace, cap it" — wrong. The 0.8B's p95 is a
consistent 17s and the 1.7B beat it in two of three runs. The tail is shared and
lives somewhere else. Do not spend time on `/no_think` on this evidence.
## Sub-500M: not close, and the benchmarks say otherwise for a reason
LFM2.5-350M publishes IFEval 76.96 against Qwen3.5-0.8B's 59.94, and BFCLv3 44.11
against 35.08 — better at instruction-following and structured output, at 2/3 the
size. Those numbers are real and they are **English**. Every benchmark in that
table except Multi-IF is English-only.
In Russian, with a 300-token budget and temperature 0:
- **350M**, «Столица Франции? Ответь кратко.» → *«Сторзит в Париже.»*`Сторзит` is
not a word; it is invented morphology.
- **350M**, asked to read back a reminder → a fortune cookie about being attentive
and confident. No reminder in it.
- **230M**, «Привет, как дела?» → answered **in Spanish**.
The 230M beating the 350M six-fold on routing (33.8% vs 5.2%) is the other tell:
when the larger sibling collapses like that it is format compliance failing, not
reasoning.
This is a pretraining gap, not a fine-tuning gap. Teaching Russian to a 350M from
near-zero is not an afternoon on a Colab, which was the premise worth checking.
## Why this vindicates the 1.7B CPT
Stock Qwen3-1.7B, untrained and unprompted, answers all three probes in fluent
correct Russian. What it gets wrong is the persona: *«Привет! Я рад, что ты здесь»*
`рад` is masculine and Maven needs `рада`. That is the right kind of remaining
problem, and it is exactly what the CPT (Vikunja #122) is for.
The 1.7B was the correct model choice. What was wrong was treating it as a
**blocker**: stock already beats what was deployed, so it ships now and gets
swapped again when the CPT lands.
## Caveats
- Routing is one run per model, not three. The gaps between families are far larger
than the run-to-run spread seen on the talk fixture, but the 2B-vs-1.7B gap (62.3
vs 67.5) is not safe to call on one run.
- ~~The routing numbers only reach production once the LLM router is wired on. It is
still `nil`.~~ **Resolved the same evening:** the LLM router is wired at `voice.go:214`
behind `voice.llm_router`, the default is on, and `deploy/mavend.json` sets it `true`.
These numbers are the production path now, so the p50 ≈2.7s is a real per-turn cost and
not a bench artifact.
- ~~`/mnt/hdd1/llms/LFM2.5/Qwen3-1.7B-UD-Q4_K_XL.gguf` is a 293 MB truncated download
in the wrong directory.~~ **Deleted 2026-07-31.** The good 1.13 GB copy in `qwen3/` is
what `deploy/mavend.json` loads.
- Harness: `scratchpad/bakeoff.sh`, one server at a time, health-checked before each
run, `/v1/models` recorded per run. Never run two LLM consumers at once — see the
contamination note in `TALK-EVAL-31-07-2026.md`.
+6 -3
View File
@@ -103,14 +103,17 @@ eval-router:
eval-recall:
MAVEN_ONNX_LIB="$(MAVEN_ONNX_LIB)" $(GO) test -v -count=1 ./internal/memory/recalleval/
# eval-phrasing -- score nudge phrasing (internal/phraser/eval). Verbose so the
# eval-phrasing -- score nudge phrasing AND the conversational paths (chat,
# query, general knowledge) in internal/phraser/eval. Verbose so the
# report and every generated message land in the terminal. With no environment
# it scores the deterministic Stub only, which is what CI runs. Set
# MAVEN_LLM_URL to add the resident model:
# MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing
# The model run is slow (minutes) -- the timeout is raised to match.
# The model run is slow (minutes) -- the timeout is raised to match. It covers
# two fixtures now (15 nudges + 27 conversational cases, and the chat replies are
# the long ones), hence 90m rather than 40m.
eval-phrasing:
$(GO) test -v -count=1 -timeout 40m ./internal/phraser/eval/
$(GO) test -v -count=1 -timeout 90m ./internal/phraser/eval/
# eval-models — score ONE llama-server against the same fixture, for the
# resident-model bake-off (#278, #250). Start a server with the gguf you want,
+38 -7
View File
@@ -103,13 +103,45 @@ but a large part of the jump is that failure now degrades into Russian instead o
The two remaining failures: one `"..."` recurrence (`routine-stretch`) and one meal nudge
that never says food.
## Tried and reverted: an example-led nudge prompt (#393)
The idea was that a 0.8B copies examples better than it follows rules, so the nudge prompt
was rewritten to lead with five on-topic examples (water, break, pills, morning, service) and
the prose rules were compressed to pay for the tokens: 1190 chars down to 986.
It measured **worse**, three runs each side, same llama-server, same fixture:
| run | before | after |
|---|---|---|
| 1 | 12/15 (address 14) | 11/15 (address 13) |
| 2 | 13/15 (address 15) | 12/15 (address 15) |
| 3 | 14/15 (address 15) | 11/15 (address 12) |
`feminine` and `hisgender` were 15/15 on all six runs, so they measure nothing here. The
regression is all in `address`: 44/45 before, 40/45 after. Formal "вы"/"ваше" and plural
imperatives came back, and so did `"..."`.
Two likely causes, both about the same thing — **examples do not carry a prohibition**. The
old prompt spent a whole sentence on «говоришь на "ты", в единственном числе»; the new one
demoted that to one item in a long "никогда" list, and the model stopped obeying it. And
making the examples on-topic let their *wording* leak: a break case came back as
«Вы давно не пили воду. Выпей стакан.» — the water example, verbatim, in the wrong slot.
That is exactly the failure the laundry/laptop examples were chosen to avoid.
Change reverted. What survives is the measurement: a rule the model must obey needs its own
sentence, and examples must stay off-topic. Also note the before side alone spans 1214 of
15 — this fixture cannot resolve anything smaller than about three cases.
## Broken, found, not fixed
1. **`checkFeminine` only catches half the constraint.** It scans for masculine
self-reference and passed 15/15 both runs — but three messages address the *owner* in
the feminine: "ты давно не отдыхал**а**", "он не ел". The owner is a man. The check has
no second-person gender test, so this scores clean while being exactly the persona
failure the constraint exists to prevent. This is the most important gap in the harness.
1. ~~**`checkFeminine` only catches half the constraint.**~~ **Fixed** (#381). It scanned for
masculine self-reference only, so three messages that addressed the *owner* in the feminine
("ты давно не отдыхал**а**") scored clean. There is now a second check, `hisgender`: a
feminine past-tense verb (-ла/-лась) in a sentence addressed to him ("ты", "тебе", "твой")
fails, unless the verb is hers ("я заметила", "напомнила тебе"). It is a suffix rule, not a
parser — see the comment in `checks.go` for what it misses. A fresh 15-case run after adding
it scored **12/15** with `hisgender` 15/15; the model did not repeat the feminine address in
that sample, and the check is pinned by unit tests on the recorded bad strings instead.
2. **Grammar is not checked at all, and it is bad.** `"Он не ел 11 дней"` (it was 11 hours),
`"Сонуждились 7 дней"` (not a word), `"Они забыли воду"` (wrong person entirely). Every
one of these passes all six checks. The fixture measures properties, not fluency, and at
@@ -125,8 +157,7 @@ that never says food.
## Next steps
1. **Add a second-person gender check** to `checks.go`. Finding 1 above. Until it exists the
feminine column means less than it looks like.
1. ~~**Add a second-person gender check**~~ — done, `hisgender` in `checks.go` (#381).
2. **Decide whether the fallback should count as a pass.** Right now `Score` cannot tell a
model answer from a fallback. Either mark fallback bodies in `PhrasedNudge` or count them
in their own column. Without that, any future prompt change can score well by failing
+4 -1
View File
@@ -90,4 +90,7 @@ later* is the worker + RAG.
4. **Deferred work** — larger reasoner, custom Piper voice and other expansions.
## Non-goals (unchanged)
Never phones home. Not a nag. Not autonomous. Feminine-gendered RU self-ref.
Not a nag. Not autonomous. Feminine-gendered RU self-ref. No telemetry, no
cloud model, no third-party account — but she MAY read external sources to
answer world questions (Kiwix first, search optional). "Never phones home" as
an absolute is deprecated, owner's call 2026-07-31; see CLAUDE.md § Non-goals.
+8
View File
@@ -75,6 +75,14 @@ same vector. A note is indexed in both places with the same embedding, so if it
`QueryNotes` it fails again here — the branch can only ever return a **fact**. Its comment calls it
"additive"; for notes it is not.
**Fixed (Vikunja #373).** The memory pass now runs *first*, as one search over notes and facts with
one gate, so whichever memory is clearly the best match answers — note or fact. The notes-only pass
stays behind it for notes the vector index does not hold. No threshold changed, so the set of
questions Maven answers is the same; only which memory answers them. The fixture gained two mixed
note+fact cases (`ru-mixed-031`, `ru-mixed-032`), which is why the counts below are out of 27
answerable cases and not 25: hash recall@1 36.0% (9/25) → 37.0% (10/27), e5 recall@1 72.0% (18/25) →
70.4% (19/27) with answered-after-gate 68.0% → 66.7% and false recall unchanged at 1/5.
### 5. Ranking has no recency or type signal, and the store is not the bottleneck
`internal/store/notes.go:67` sorts by cosine and uses `ts` only to break an exact float tie, which
+108 -10
View File
@@ -66,15 +66,113 @@ Three things this run settles:
`запиши что…` phrasings toward fact, and that suspicion stands — all five `ru-note-*`
cases now land on fact. Tracked as Vikunja #375.
**Thinking off is the best configuration measured so far**, on both accuracy and latency
(Vikunja #376). That is worth understanding before flipping: routing is a short
classification into a fixed enum with grammar-constrained output, so there is little to
reason about, and the thinking trace mostly gives a small model room to talk itself out of
the right answer. Phrasing is a different job and needs measuring separately.
The `thinking off` column above read as the best configuration measured so far (Vikunja #376).
**It was wrong** — see the controlled re-run below. Ignore that column.
Still `6 / 6` missed clarify — the router has no way to say "I don't know" (Vikunja #359).
That is unchanged by anything here.
## Thinking off — 31-07-2026, controlled re-run (Vikunja #376)
The "thinking off wins by 6 points" observation above **does not hold**. It was a measurement
artefact, and the earlier table's `thinking off` column should be ignored.
The thinking-off variant was scored by a hand-rolled HTTP client living in the test file
instead of `llm.Client`. That copy did not send `repeat_penalty`, which the real router does
send (`routeRepeatPenalty = 1.15`). So the two columns differed on two axes at once, and the
one that mattered was the penalty, not the thinking mode.
Re-measured with everything else held equal — same fixture, same prompt, same grammar, same
sampling, same idle box, the three configurations run back to back and never concurrently:
| | llm-only, thinking on | llm-only, thinking off | cascade+llm |
|---|---|---|---|
| intent-only accuracy | 59.2% (45/76) | 59.2% (45/76) | 61.8% (47/76) |
| full accuracy (intent+slots+gate) | 38.2% (29/76) | 38.2% (29/76) | 57.9% (44/76) |
| route errors | 3 | 3 | 0 |
| grammar violations | 3 (all 3 route errors) | 3 (same 3 cases) | 0 |
| missed clarify | 5 / 6 | 5 / 6 | 5 / 6 |
| p50 latency | 836ms | 920ms | 810ms |
| p95 latency | 1.41s | 2.00s | 1.31s |
Thinking off is not just a tie on the headline numbers — it is identical case for case, with
the same confusion matrix and the same three unparseable replies. The latency difference is
run-to-run noise on one box, and it points the wrong way here.
The reason is simpler than any accuracy argument: **this llama-server build ignores the
request-level thinking switch for this model.** Probed directly against the running server
with `chat_template_kwargs.enable_thinking = false`, `chat_template_kwargs.thinking = false`
and top-level `reasoning_budget = 0` — all three return a byte-identical answer with the
thinking trace still in `reasoning_content`, and the server reports the prompt prefix as
cached, meaning the rendered template did not change. There was never anything being turned
off, which is also why the numbers match exactly.
Nothing was defaulted. `internal/llm` still has no `chat_template_kwargs` field, `VoiceConfig`
has no thinking flag, and `deploy/mavend.json` is unchanged. The misleading third
configuration is removed from `internal/router/eval` so the table it produced cannot be quoted
again.
Two caveats worth saying out loud:
- **The fixture is 76 cases.** A 6-point difference on 76 cases is roughly 4-5 cases and would
not have been worth trusting even if it had reproduced. This one was exactly 0 cases, which
is a much easier call.
- **This is one server build and one checkpoint** (`b9351`, Qwen3.5-0.8B Q4_K_M). If the
#122 checkpoint or a newer llama.cpp does honour the switch, the question reopens — but it
reopens as an unmeasured question, not as a 6-point win.
Phrasing was **not** measured. Whether thinking helps there is still open, and now also blocked
on the same "can we even turn it off" question.
## Clock and calendar rule — 31-07-2026 (Vikunja #374)
`routeSystem` never said whether "который час" or "какое число завтра" are `system` or
`query`, and `system→query ×4` showed up in every run. The rule added says: the clock and the
calendar date themselves are `system`; what is *written in* the calendar or in memory
("что у меня завтра", "какие есть напоминания") stays `query`; and a time named inside a
request ("напомни завтра…") is just a detail of the request, not a reason for `system`.
That split is not a preference. In `cmd/mavend/voice.go` only `replySystem` owns the clock and
the date formatter, so a clock question routed to `query` falls into the embedder + note RAG
and answers "не знаю". The agenda, on the other hand, is answered by `ParseCalendarDate` +
`CalendarEvents` *inside* the `query` branch, so that side has to stay `query`. The rule sits
above the question test because every one of these utterances carries a question word and a
later rule would never be reached.
The fixture is now 77 cases: one calendar-agenda case was added
(`ru-query-019` "что у меня стоит в календаре на послезавтра", intent `query`) specifically so
an over-broad system rule cannot pass unnoticed. The clock/date cases (`ru-sys-001/002/005`,
`en-sys-001`) already existed.
Three runs, same box, back to back, never concurrently:
| | baseline | first rule (too broad) | rule as committed |
|---|---|---|---|
| llm-only intent-only | 59.2% (45/76) | 54.5% (42/77) | 59.7% (46/77) |
| llm-only full | 38.2% | 35.1% | 39.0% |
| llm-only route errors | 3 | 4 | 5 |
| llm-only p50 | 1.09s | 0.91s | 0.93s |
| cascade+llm intent-only | 61.8% (47/76) | 58.4% | 62.3% (48/77) |
| cascade+llm full | 57.9% | 54.5% | 59.7% |
| cascade+llm route errors | 0 | 0 | 0 |
| cascade+llm p50 | 0.91s | 0.80s | 1.04s |
**The targeted bug is fixed and the headline number did not move.** `system→query ×4` is gone
in both LLM configurations — the `time` and `date` tags go from 0/2 and 0/2 to 2/2 and 2/2 —
but the model then over-applies the rule, and `query→system ×5` plus `reminder→system ×2`
appear where they did not exist before. Net accuracy is a wash, inside the noise of a 77-case
fixture.
The first attempt is shown because it is the honest history: it said "спрашивает время, дату
или день недели → system" with no scope, which swept up reminders, and it cost 3-5 points. It
was tightened once, on the reasoning that a rule capturing "напомни завтра в 7" is simply
wrong, and not tuned further. The remaining `query/reminder → system` over-trigger is a new,
separate weakness of the sub-1B model and deserves its own task rather than more prompt
kneading against a held-out fixture.
The rule is kept. It is correct about what the daemon can answer, and the failure it replaces
was silent ("не знаю" to "который час") while the one it introduces is loud.
## Findings
### 1. The resident model does route better — 50.0% vs 36.8%
@@ -145,11 +243,11 @@ Note the grammar's `string ::= "\"" ([^"\\] | "\\" .)* "\""` is unbounded, so no
### 7. Two hypotheses tested and closed
- **Thinking mode is a non-issue.** Qwen3.5's template defaults `thinking = 1`, so
grammar-constrained JSON lands in `reasoning_content` with `content` empty —
`llm.Client`'s fallback handles it. A `thinking off` run scored *identically* (18/76,
48.7%, same p50). `internal/llm` deliberately does **not** grow a `chat_template_kwargs`
field.
- **Thinking mode is a non-issue.** Confirmed twice now, the second time properly — see the
controlled re-run section. Grammar-constrained JSON lands in `reasoning_content` with
`content` empty and `llm.Client`'s fallback handles it; the request-level switch does
nothing on this build. `internal/llm` deliberately does **not** grow a
`chat_template_kwargs` field.
- **Runaway array repetition does not reproduce.** An isolated smoke test with a stripped
grammar emitted `{"intent":"reminder"}` until `MaxTokens`; under the real `routeSystem`
prompt the few-shot examples anchor it to one object. 2 errors in 76, not 76.
+150
View File
@@ -0,0 +1,150 @@
# Conversational phrasing eval — 31-07-2026
Every score measured tonight, on the three paths the nudge eval never touched:
chat, query-with-notes, and general knowledge.
**Short version: the plumbing got fixed and the score barely moved.** Grammar and
Russian prompts together took the composite from ~9 to ~14 of 27. Everything
still failing is the model not knowing things or not holding a constraint, and
prompting is out of levers. Settles the measurement half of Vikunja #395 / #398 /
#400.
## How to reproduce
```sh
# llama-server: -c 4096 -ngl 99 -t 6, model /mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf
MAVEN_LLM_URL=http://127.0.0.1:18099 no_proxy=127.0.0.1,localhost \
deps/go/go/bin/go test -count=1 -timeout 40m \
-run TestLLMTalkBaseline ./internal/phraser/eval/ -v
```
Three runs per configuration, always. The fixture is 27 cases, so one reply
changing moves the composite by 3.7 points — a single run cannot tell a real
change from sampling noise. This was learned the expensive way: an earlier claim
that "one nudge case fails every run" turned out to be three different cases
across three runs.
**Run the box otherwise idle.** See the contamination note at the bottom.
## Composite, per configuration
| config | overall /27 | chat /9 | query /9 | knowledge /9 | canned fallbacks |
|---|---|---|---|---|---|
| baseline, no grammar | 7, 12, 7 | 1, 1, 0 | 2, 4, 2 | 4, 7, 5 | 0, 0, 0 |
| + GBNF grammar (#398) | 14, 15, 8 | 1, 3, 0 | 5, 6, 3 | 8, 6, 5 | 0, 0, 0 |
| + Russian prompts (#400) | 11, 17, 15 | 1, 5, 3 | 5, 6, 8 | 5, 6, 4 | 0, 0, 0 |
| + truncation fix, 1000ch/768tok | 12, 13, 10 | 2, 2, 1 | 7, 7, 5 | 3, 4, 4 | 3, 3, 6 |
| + rebalanced, 600ch/1024tok | **void — contaminated** | | | | |
"Canned fallbacks" counts replies that came back as the hardcoded `"не знаю."`
or `"поговорили."`. It is not a check, it is a health signal: those strings mean
the phraser gave up, and the eval scores them as ordinary bad replies.
## Per-check
| check | no grammar | + grammar | + RU prompts | + truncation fix |
|---|---|---|---|---|
| nonempty | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 |
| ellipsis | 20, 19, 23 | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 |
| lang | 13, 16, 15 | 23, 26, 26 | 25, 26, 25 | 26, 27, 27 |
| feminine | — | — | 25, 24, 26 | 25, 25, 27 |
| address | — | — | 21, 22, 22 | 22, 21, 22 |
| ontopic | — | — | 17, 24, 18 | 17, 19, 14 |
`nonempty` reading 27/27 everywhere is not good news — it was a broken check.
It tested for a non-blank string, so replies of literally `{` and `"15-16"`
passed it. Fixed on `overnight/fix-truncation`; it needs a letter now.
## What each change actually bought
**GBNF grammar (#398) — the biggest single win.** Qwen3.5-0.8B writes
`Thinking Process:` as plain text with no tags, `stripThink` only handles
`</think>`, so the JSON never closed and the plain-text fallback shipped the
literal reasoning. `ellipsis` went 20→27 and `lang` 13→26. The router had been
using a grammar for ages; the phraser asking nicely in the prompt was the
oversight.
**Russian prompts (#400) — modest, plus a large latency win.** Chat 1.3→3.0
average, query 4.7→6.3, knowledge 6.3→5.0. All inside the run-to-run spread, so
"probably better on the paths it targeted, not provable in three runs". p50
latency dropped from ~11.5s to ~2.3s and that part is consistent across all
three runs — shorter prompts, and she stopped emitting English reasoning first.
**Truncation fix — necessary, and did not help the score.** Two real bugs
(replies of `{`, and a `nonempty` check that passed them), both fixed, and the
composite went nowhere. A complete rambling wrong answer fails the same checks a
truncated one did. Worth doing anyway: the daemon was shipping `{` to a
text-to-speech voice.
## The truncation bug, since the cause was counter-intuitive
The grammar's `string ::= ... {0,400}` rule was the cause, not the token cap.
Measured against Qwen3.5-0.8B at three caps — 256, 768 and 2048 — the reply came
back **exactly 400 characters every time, cut mid-word** (`"Нужно записать и,"`).
Then I raised the bound to 1000 while the cap was 768 tokens and made it worse:
Russian runs ~1.5 characters per token here, so generation died on the *token*
cap instead, mid-object, and the new guard correctly refused it and shipped
`"не знаю."` — 3, 3 and 6 fallbacks per run, from zero. **The two limits have to
agree.** 600 characters needs ~400 tokens; the cap is 1024.
## Where the remaining failures live
`address` is stuck at 21-22 of 27 and `ontopic` at 14-19. Both resist prompting.
**The prompt now explicitly forbids exactly what she does.** It says never "вы",
use the singular — and she writes `вашей`, `подождите`, `делаете`, `хотите`,
`напишите`. Telling a 0.8B "never do X" does not work. Same for
`feminine`: `я готов`, `я понял`, `я нашел`, `я заметил`, `я сказал`.
**Some of `ontopic` is the fixture, not the model.** `chat-how-are-you` got
`"Привет! Я здесь, чтобы поговорить. Как дела сегодня?"` — a fine reply that
fails because `want_any` is `[норм, хорош, порядк, тут, работ]`. It fails in
every run, so it inflates the count. The `ontopic` column currently measures the
fixture as much as the model. Not fixed yet, deliberately: changing it would
break comparability with the runs above.
**Two replies worth reading, because they are not fixable by prompting:**
- Thunder and lightning: *"Скорость молнии — 8-10 тысяч километров в секунду, но
звук — 300 метров в секунду, что делает молнию громче."* Confidently wrong,
and it concludes lightning is *louder* rather than sound being *slower*.
- "расскажи обо мне": *"Ты — прекрасное существо, с душой и вниманием… Спасибо за
твою улыбку… О тебе — заповедь любви."* Sycophantic filler, zero information,
and precisely the "not a relationship" non-goal.
- Boiling an egg: `"15-16"` one run, `"1"` another. No unit, wrong number.
The first argues for reading instead of recalling (#403 — Kiwix retrieval scores
8/8 on the same questions given English keywords). The second and third argue
for templates on the paths where correctness matters (#392).
## Contamination note — how the last row got voided
I started the query-rewrite agent against the same llama-server the sweep was
using, and assumed contention would only affect latency. It did not. The
knowledge path collapsed to 0 of 9 with eight canned `"не знаю."` replies, p95
tripled to 23.7s, and **the report still said "0 errors"**.
That is Vikunja #397, and it is worse than filed: a merely *busy* server
produces a clean-looking report with a third of the fixture silently answering
`"не знаю."`. `PhraseChat` and `PhraseQuery` swallow every failure and return a
hardcoded string, so infrastructure trouble is indistinguishable from bad
phrasing in the score. The talk test guards the *start* and *end* of a run with
a model check, which catches a dead server but not a loaded one.
**Until #397 is fixed, treat any run made on a busy box as void.**
## Next
- Re-run 600ch/1024tok clean, to fill the void row.
- Score `Qwen3.5-2B-UD-Q4_K_XL` (already at `/mnt/hdd1/llms/qwen3.5/`, never
measured) on this fixture and the router fixture. Not the 4B — too big for
this box, owner's call.
- Newer sub-500M candidates (LFM2.5 200M/300M) are worth a run for routing.
Note `MODEL-BAKEOFF-31-07-2026.md` found LFM2.5-**1.2B** worse than
Qwen3.5-0.8B at Russian routing and 2.4× slower — but those are a different,
older generation, so that result does not predict the small ones.
- Fix `chat-how-are-you`'s `want_any`, and re-baseline once, so `ontopic`
measures the model.
- #397 first if anything, since it decides whether any of the above is
trustworthy.
+1 -1
View File
@@ -12,7 +12,7 @@ import (
)
type fakeCore struct {
ipc.CoreAPI
ipc.UnimplementedCoreAPI
facts map[string]ipc.Fact // composite key "key|source" → Fact
writeLog []ipc.WriteFactReq
writeErr error
+71
View File
@@ -0,0 +1,71 @@
// actionTable dispatches applyAction's per-intent bodies. Each of the 7
// intents (fact, reminder, note, query, act, chat, system) has one handler
// here with the signature:
//
// func(h *reactiveHandler, ctx context.Context, dec router.Decision) string
//
// same contract as applyAction itself: "" means "let the Replier phrase the
// reply", a non-empty string OVERRIDES it. This is a straight extraction of
// applyAction's old switch cases (formerly ~300 lines in voice.go) — no
// reordering of side effects, no new abstractions inside a handler.
//
// What does NOT belong in this table, because it is not per-intent:
//
// - the dec.Clarify short-circuit ("" when the router's stage-3 fired) —
// stays in applyAction, before dispatch, since it applies to every
// intent identically.
// - the destructive-act confirm gate (park / resolveConfirm / confirmTTL)
// and the enabled-tool allowlist. Both live entirely inside
// actionAct/handleAct in actions_act.go, exactly where they lived in the old
// switch's IntentAct case — they are act-specific (a fact or a note
// can't be destructive), not shared across intents, so they do not need
// to move to a separate layer. The important invariant, preserved
// as-is: applyAction runs identically whether dec came from a fresh
// route or from a completed clarify answer (see finishClarified in
// clarify.go and its comment "filling in an argument never grants
// authority") — a handler must never special-case a clarify-completed
// decision to skip the confirm gate or the allowlist.
// - detectPattern and dialogue-session bookkeeping (rememberTurn,
// followUpMerge) run in the callers (runTurn,
// finishClarified), not per-intent, and are untouched by this slice.
//
// Each handler lives in actions_<intent>.go; the small ones (chat, system)
// and the table itself stay here.
//
// Adding an intent: write its handler in its own file, add one line to
// actionHandlers. Do not grow applyAction's switch back.
package main
import (
"context"
"log"
"github.com/kami/maven/internal/router"
)
// actionHandlers is the per-intent dispatch table used by applyAction.
var actionHandlers = map[router.Intent]func(*reactiveHandler, context.Context, router.Decision) string{
router.IntentFact: (*reactiveHandler).actionFact,
router.IntentReminder: (*reactiveHandler).actionReminder,
router.IntentAct: (*reactiveHandler).actionAct,
router.IntentChat: (*reactiveHandler).actionChat,
router.IntentSystem: (*reactiveHandler).actionSystem,
router.IntentNote: (*reactiveHandler).actionNote,
router.IntentQuery: (*reactiveHandler).actionQuery,
}
func (h *reactiveHandler) actionChat(ctx context.Context, dec router.Decision) string {
// Conversational: build history from dialogue session (prior user turns)
// and let the LLM respond from general knowledge + context.
history := h.chatHistory()
reply, err := h.phraser.PhraseChat(ctx, dec.Utterance, history)
if err != nil {
log.Printf("voice: chat: %v", err)
return "поговорили."
}
return reply
}
func (h *reactiveHandler) actionSystem(ctx context.Context, dec router.Decision) string {
return h.replySystem(ctx, dec)
}
+66
View File
@@ -0,0 +1,66 @@
package main
import (
"context"
"errors"
"log"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/tool"
)
// actionAct handles router.IntentAct: match a verb to an enabled tool, offer
// it to the ecosystems first, and run it behind the confirm gate and the
// allowlist. proposeGap and the confirm gate itself live in confirm.go.
func (h *reactiveHandler) actionAct(ctx context.Context, dec router.Decision) string {
// tool executor: run the matched fn against the enabled allowlist.
// HasFn=false ⇒ try the matcher (for LLM-routed acts where the verb
// didn't go through the stage-0 act grammar).
if !dec.Slots.HasFn && dec.Slots.Text != "" && h.matcher != nil {
if fn, args, ok := h.matcher.Match(dec.Slots.Text); ok {
dec.Slots.Fn, dec.Slots.Args, dec.Slots.HasFn = fn, args, true
}
}
// Praxis ecosystem tools: intercept before the system command executor.
if h.ecosystem != nil && h.ecosystem.praxis != nil && dec.Slots.HasFn {
if reply := h.handlePraxisAct(ctx, dec); reply != "" {
return reply
}
}
// Hexis ecosystem action: if ecosystem is configured and we have a verb
// + entity text, try to resolve the entity and execute via Hexis.
if h.ecosystem != nil && h.ecosystem.hexis != nil && dec.Slots.Text != "" {
if reply := h.handleHexisAct(ctx, dec); reply != "" {
return reply
}
}
// HasFn still false ⇒ no allowlist match: scaffold a 'proposed' tool
// the user can enable on the authed surface ("earn the right to ask").
if !dec.Slots.HasFn {
return h.proposeGap(ctx, dec)
}
out, err := h.tools.Exec(ctx, dec.Slots.Fn, dec.Slots.Args, false)
if err != nil {
switch {
case errors.Is(err, tool.ErrNeedsConfirm):
// destructive: park it and ask. The next utterance answers.
phrase := actPhrase(dec.Slots.Fn, dec.Slots.Args)
h.park(dec.Slots.Fn, dec.Slots.Args, phrase)
return "выполнить «" + phrase + "»? скажи «да» или «нет»."
case errors.Is(err, tool.ErrNotEnabled):
return h.proposeGap(ctx, dec)
}
log.Printf("voice: tool %s: %v", dec.Slots.Fn, err)
if out != "" {
return "не получилось выполнить команду: " + firstLine(out)
}
return "не получилось выполнить команду."
}
if out != "" {
return "готово: " + firstLine(out)
}
return "готово."
}
+64
View File
@@ -0,0 +1,64 @@
package main
import (
"context"
"log"
"strconv"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
)
// actionFact handles router.IntentFact: persist a tapped self-fact, index
// it for recall, and let pattern detection propose a routine.
func (h *reactiveHandler) actionFact(ctx context.Context, dec router.Decision) string {
if !dec.Slots.HasKey {
return "не разобрала, что записать — попробуй иначе."
}
now := h.now()
req := ipc.WriteFactReq{
Ts: now,
Kind: "self",
Key: dec.Slots.Key,
Value: dec.Slots.Value,
Source: "tap:voice",
Confidence: 1.0,
// Subject: the key doubles as the entity-resolution candidate —
// a voice-tapped fact's key is usually the thing/person it's
// about ("espresso_machine", "kate"), so queueing it for Nexus
// resolution costs one async lookup and is a no-op (not_found)
// for the abstract self-state keys (mood, water) that aren't
// entities at all.
Subject: dec.Slots.Key,
}
factID, err := h.api.WriteFact(ctx, req)
if err != nil {
log.Printf("voice: write fact: %v", err)
return "не получилось сохранить факт."
}
// Index the fact utterance in long-term memory (best-effort, must not
// fail the fact write). Facts aren't in the notes table, so this is the
// only recall path for them — "когда я пил воду?" reads back from here.
if h.memStore != nil {
if vec, err := router.EmbedPassage(ctx, h.embedder, dec.Utterance); err != nil {
log.Printf("voice: embed fact for memory: %v", err)
} else if err := h.memStore.Insert(ctx, "fact:"+dec.Slots.Key+":"+strconv.FormatInt(now.Unix(), 10), vec, map[string]string{
"source": "voice",
"type": "fact",
"text": dec.Utterance,
"ts": strconv.FormatInt(now.Unix(), 10),
}); err != nil {
log.Printf("voice: memory insert fact: %v", err)
}
}
// Event extraction + pattern detection (best-effort, must not fail the
// fact write). If the fact describes a recognizable action, it becomes a
// normalized event; if ≥3 events for the same action+object show stable
// intervals, a proposed routine is created and parked for confirmation.
if h.dataStore != nil {
if phrase := h.detectPattern(ctx, factID, dec.Slots.Key, dec.Slots.Value, now); phrase != "" {
return phrase // "ты заправляешь ... напоминать?"
}
}
return "" // replier phrases the success reply
}
+41
View File
@@ -0,0 +1,41 @@
package main
import (
"context"
"log"
"strconv"
"github.com/kami/maven/internal/router"
)
// actionNote handles router.IntentNote: embed the note, persist it, and
// index it for recall.
func (h *reactiveHandler) actionNote(ctx context.Context, dec router.Decision) string {
// embed the note text with the same model the classifier uses, persist
// via CoreAPI (source=tap:voice). Semantic recall lives in `notes`, not
// facts — no predicate reads it (spec's two-memory split).
vec, err := router.EmbedPassage(ctx, h.embedder, dec.Utterance)
if err != nil {
log.Printf("voice: embed note: %v", err)
return "не получилось сохранить заметку."
}
noteTs := h.now()
noteID, err := h.api.WriteNote(ctx, noteTs, dec.Utterance, vec, "tap:voice")
if err != nil {
log.Printf("voice: write note: %v", err)
return "не получилось сохранить заметку."
}
// Insert into long-term memory (best-effort, must not fail the note write).
// text/ts in the meta make a Search hit self-describing (see bestRecall).
if h.memStore != nil {
if err := h.memStore.Insert(ctx, "note:"+strconv.FormatInt(noteID, 10), vec, map[string]string{
"source": "voice",
"type": "note",
"text": dec.Utterance,
"ts": strconv.FormatInt(noteTs.Unix(), 10),
}); err != nil {
log.Printf("voice: memory insert: %v", err)
}
}
return "" // replier phrases the "saved" reply
}
+226
View File
@@ -0,0 +1,226 @@
package main
import (
"context"
"errors"
"fmt"
"log"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/weather"
)
// queryTurn is the per-turn scratch a chain of query sources shares: the
// decision being answered plus the work an earlier source already paid for
// (the query embedding, the notes it pulled). Sources read and fill it in
// order, so a later source never re-embeds.
type queryTurn struct {
dec router.Decision
vec []float32
notes []ipc.Note
}
// querySource — one answer source in the chain actionQuery walks. answer
// returns (reply, true) when this source claims the question, ("", false)
// when it passes to the next one. name is for reading the table, not logged.
//
// A struct of one func rather than an interface: every source is a plain
// method on *reactiveHandler with no state of its own (what state a turn has
// lives in queryTurn), so an interface would mean one empty type per source
// to satisfy it — ceremony for nothing. Same reasoning as confirmResolver in
// confirm.go, and the table then reads like actionHandlers: a flat list of
// method expressions you extend with one line.
type querySource struct {
name string
answer func(*reactiveHandler, context.Context, *queryTurn) (string, bool)
}
// querySources is the ordered chain actionQuery walks; first source to claim
// answers the turn. THE ORDER IS LOAD-BEARING — see the memory-before-notes
// comment on queryMemory: running the notes-only pass first was #373, and the
// gate was never the bug. Adding a source (Kiwix, RSS, crawler, email) is one
// line here plus its method; where you put the line is the whole decision.
var querySources = []querySource{
{"fact-by-key", (*reactiveHandler).queryFactByKey},
{"calendar", (*reactiveHandler).queryCalendar},
{"weather", (*reactiveHandler).queryWeather},
{"embed", (*reactiveHandler).queryEmbed},
{"memory", (*reactiveHandler).queryMemory},
{"notes", (*reactiveHandler).queryNotes},
{"general-knowledge", (*reactiveHandler).queryGeneral},
}
func (h *reactiveHandler) actionQuery(ctx context.Context, dec router.Decision) string {
t := &queryTurn{dec: dec}
for _, src := range querySources {
if reply, ok := src.answer(h, ctx, t); ok {
return reply
}
}
return "не знаю."
}
// queryFactByKey — when the dialogue layer resolved an anaphoric reference to
// a prior fact's key (e.g. "когда я это сделал?" after "запиши что я пил
// воду"), look up the fact's value directly.
func (h *reactiveHandler) queryFactByKey(ctx context.Context, t *queryTurn) (string, bool) {
dec := t.dec
if !dec.Slots.HasKey || dec.Slots.Key == "" {
return "", false
}
f, err := h.api.LatestFact(ctx, dec.Slots.Key)
if err != nil {
return "", false
}
if dec.Slots.HasTime {
// The query asks about timing — the fact's own timestamp is the
// answer it's looking for. Format as a natural reply.
return fmt.Sprintf("я записала это %s", formatTime(f.Ts)), true
}
// General fact reference: describe what we know.
if dec.Utterance == "" {
return fmt.Sprintf("вот что я знаю: %s — %s", dec.Slots.Key, f.Value), true
}
// The utterance still carries the question; fall through to normal RAG
// with the resolved key in context.
return "", false
}
// queryCalendar — "что у меня сегодня?", "планы на завтра?"
// h.now(), not time.Now(): the handler's clock is the injected one, so this
// source can be tested at a fixed time like the rest.
func (h *reactiveHandler) queryCalendar(ctx context.Context, t *queryTurn) (string, bool) {
date, ok := router.ParseCalendarDate(t.dec.Utterance, h.now())
if !ok {
return "", false
}
events, err := h.api.CalendarEvents(ctx, date, date.Add(24*time.Hour))
if err != nil {
log.Printf("voice: calendar events: %v", err)
return "не получилось проверить календарь.", true
}
values := make([]string, len(events))
for i, e := range events {
values[i] = e.Value
}
var f router.CalendarEventFormatter
return f.Format(values, date), true
}
func (h *reactiveHandler) queryWeather(ctx context.Context, t *queryTurn) (string, bool) {
if !isWeatherQuery(t.dec.Utterance) {
return "", false
}
loc := extractWeatherLocation(t.dec.Utterance, h.weatherLocation)
ctxWT, cancel := context.WithTimeout(ctx, 5*time.Second)
defer cancel()
w, err := h.weatherProvider.CurrentWeather(ctxWT, loc)
if errors.Is(err, weather.ErrNotConfigured) {
return "погода не настроена.", true
}
if err != nil {
log.Printf("voice: weather: %v", err)
return "не получилось узнать погоду.", true
}
return fmt.Sprintf("в %s сейчас %.0f градусов, %s.", w.Location, w.Temperature, w.Condition), true
}
// queryEmbed isn't an answer source — it's the shared cost the two recall
// sources below both need, run once, in the position it always ran in. It
// only claims the turn when the embedder fails.
func (h *reactiveHandler) queryEmbed(ctx context.Context, t *queryTurn) (string, bool) {
vec, err := router.EmbedQuery(ctx, h.embedder, t.dec.Utterance)
if err != nil {
log.Printf("voice: embed query: %v", err)
return "не получилось найти ответ.", true
}
t.vec = vec
return "", false
}
// queryMemory — long-term memory first: ONE search over everything Maven
// remembers (notes and facts share this index) and ONE confidence gate, so
// the memory that is clearly the best match answers — a note just as much as
// a fact.
//
// This used to run only after the notes-only source below had already
// rejected the same note at the same score, which no note could ever survive
// a second time: the branch could only return a fact (#373). Order, not the
// gate, was the bug — the set of questions Maven answers is unchanged, only
// which memory gets to answer them.
func (h *reactiveHandler) queryMemory(ctx context.Context, t *queryTurn) (string, bool) {
if h.memStore == nil {
return "", false
}
hits, herr := h.memStore.Search(ctx, t.vec, 3)
if herr != nil {
log.Printf("voice: memory search: %v", herr)
return "", false
}
hit, ok := bestRecall(hits, h.queryMinScore, h.queryMinMargin)
if !ok {
return "", false
}
text := hit.Meta["text"]
// A note is phrased in Maven's voice; a fact is read back as it was
// stored.
if hit.Meta["type"] == "note" {
if reply, perr := h.phraser.PhraseQuery(ctx, t.dec.Utterance, []string{text}); perr == nil && reply != "" {
return reply, true
}
}
return text, true
}
// queryNotes — notes-only pass, for notes the vector index above does not
// hold (an older note written before it existed). Same gate, notes-only
// candidates.
//
// Confidence gate: below it, say "I don't know" rather than read back the
// least-unrelated note — a confident wrong recall is worse than a gap (spec's
// "not a guesser-of-truth"). Same instinct as the loop's since(key)==null →
// don't fire. Two parts: an absolute cosine floor, and a margin over the
// runner-up, which is the part that works with the e5 embedder's narrow score
// band. See memory.Confident. Failing the gate passes the turn on to general
// knowledge, which is what "don't read back the runner-up" means here.
func (h *reactiveHandler) queryNotes(ctx context.Context, t *queryTurn) (string, bool) {
notes, err := h.api.QueryNotes(ctx, t.vec, 5)
if err != nil {
log.Printf("voice: query notes: %v", err)
return "не получилось найти ответ.", true
}
t.notes = notes
noteScores := make([]float64, len(notes))
for i, n := range notes {
noteScores[i] = n.Score
}
if !memory.ConfidentScores(noteScores, h.queryMinScore, h.queryMinMargin) {
return "", false
}
texts := make([]string, len(notes))
for i, n := range notes {
texts[i] = n.Text
}
reply, err := h.phraser.PhraseQuery(ctx, t.dec.Utterance, texts)
if err != nil {
log.Printf("voice: phrase query: %v", err)
}
if reply == "" {
reply = "вот что я нашла: " + texts[0]
}
return reply, true
}
// queryGeneral — general knowledge from the phraser, the last source before
// giving up. It always claims: either the model answers or Maven says she
// doesn't know.
func (h *reactiveHandler) queryGeneral(ctx context.Context, t *queryTurn) (string, bool) {
reply, err := h.phraser.PhraseQuery(ctx, t.dec.Utterance, nil)
if err != nil || reply == "" {
return "не знаю.", true
}
return reply, true
}
+33
View File
@@ -0,0 +1,33 @@
package main
import (
"context"
"log"
"github.com/kami/maven/internal/router"
)
// actionReminder handles router.IntentReminder: parse the time when stage-0
// skipped the extractor, then create the reminder.
func (h *reactiveHandler) actionReminder(ctx context.Context, dec router.Decision) string {
if !dec.Slots.HasTime {
// Stage-0 (reminder-wakeword grammar) skips the extractor, so the
// time wasn't parsed. Run the parser as a fallback.
if dec.Stage == 0 && h.timeParser != nil {
t, ok, err := h.timeParser.Parse(ctx, dec.Utterance, h.now())
if err == nil && ok {
dec.Slots.Time = t
dec.Slots.HasTime = true
}
}
if !dec.Slots.HasTime {
return "не получилось разобрать время напоминания."
}
}
payload := `{"text":` + jsonString(dec.Utterance) + `}`
if _, err := h.api.CreateReminder(ctx, dec.Slots.Time, payload, ""); err != nil {
log.Printf("voice: create reminder: %v", err)
return "не получилось поставить напоминание."
}
return ""
}
+57 -6
View File
@@ -3,6 +3,8 @@ package main
import (
"context"
"log"
"math/rand"
"strings"
"time"
"github.com/kami/maven/internal/dialogue"
@@ -21,8 +23,12 @@ const clarifyTTL = 90 * time.Second
// raw utterance, chat and system have nothing to fill in. For those a clarify
// decision keeps the canned "не поняла" reply — inventing a question for noise
// is worse than admitting she missed it.
// A reminder wants BOTH what to remind about and when. Subject first: "напомни
// в 11" has a time and nothing to say at 11, and a reminder with no subject is
// not worth setting. Order here is the order she asks in — she still only asks
// about the first one missing.
var wantedSlots = map[router.Intent][]dialogue.Slot{
router.IntentReminder: {dialogue.SlotTime},
router.IntentReminder: {dialogue.SlotText, dialogue.SlotTime},
router.IntentFact: {dialogue.SlotKey},
router.IntentAct: {dialogue.SlotFn},
}
@@ -35,7 +41,8 @@ var wantedSlots = map[router.Intent][]dialogue.Slot{
// questions, so there is no gender agreement to get wrong; the feminine
// self-reference lives in the reply she gives when she drops the request.
var clarifyQuestions = map[dialogue.Slot]string{
dialogue.SlotTime: "На когда напомнить?",
dialogue.SlotTime: "Когда?",
dialogue.SlotText: "О чём напомнить?",
dialogue.SlotKey: "Что записать?",
dialogue.SlotFn: "Что сделать?",
}
@@ -45,11 +52,55 @@ var clarifyQuestions = map[dialogue.Slot]string{
// landed. Feminine self-reference ("поняла"), as everywhere.
const clarifyGaveUp = "Прости, я не поняла. Скажи, пожалуйста, по-другому."
// clarifyExpired — his answer came after the TTL, so the parked request is
// already gone. Same tone as clarifyGaveUp, different reason: too much time
// clarifyExpiredVariants — his answer came after the TTL, so the parked request
// is already gone. Same tone as clarifyGaveUp, different reason: too much time
// passed, not "I did not understand". Feminine self-reference ("ждала",
// "отпустила"); he is addressed with a plain imperative.
const clarifyExpired = "Прости, я слишком долго ждала ответа и отпустила прошлую просьбу. Если она ещё нужна, скажи заново."
//
// Five phrasings, not one. This is the line he hears whenever he walks off
// mid-request, so it is the line that repeats most — and the same sentence every
// time is what makes a house assistant sound like a kiosk. They all carry the
// same two facts (the old request is gone; say it again if it still matters),
// because the wording may vary and the meaning may not.
//
// Fixed templates rather than model output, for the same reason as
// clarifyQuestions: this text has to be right every time, and it is not worth a
// generation to say something this small.
var clarifyExpiredVariants = []string{
"Прости, я слишком долго ждала ответа и отпустила прошлую просьбу. Если она ещё нужна, скажи заново.",
"Кажется, прошлая просьба уже не важна — я её отпустила. Если я ошибаюсь, повтори.",
"Ты как-то резко замолчал, и я не стала ждать дальше. Если та просьба ещё нужна, скажи заново.",
"Я не дождалась ответа и убрала прошлую просьбу. Повтори, если она всё ещё нужна.",
"Столько времени прошло, что я отпустила прошлую просьбу. Скажи заново, если она в силе.",
}
// clarifyExpiredLine picks one of them at random.
func clarifyExpiredLine() string {
return clarifyExpiredVariants[rand.Intn(len(clarifyExpiredVariants))]
}
// isClarifyExpired reports whether s opens with any of the expiry lines. The
// notice is glued in front of this turn's reply (see withNotice), so a caller
// checking for it has to match a prefix, not the whole string.
func isClarifyExpired(s string) bool {
for _, v := range clarifyExpiredVariants {
if strings.HasPrefix(s, v) {
return true
}
}
return false
}
// trimClarifyExpired strips a leading expiry notice, leaving this turn's actual
// reply. "" ⇒ the notice was the whole thing.
func trimClarifyExpired(s string) string {
for _, v := range clarifyExpiredVariants {
if strings.HasPrefix(s, v) {
return strings.TrimSpace(strings.TrimPrefix(s, v))
}
}
return strings.TrimSpace(s)
}
// clarifyExpiredNotice returns that line when a parked question had just timed
// out, and "" when nothing was parked. Call it right after
@@ -63,7 +114,7 @@ func (h *reactiveHandler) clarifyExpiredNotice() string {
return ""
}
log.Printf("voice: clarify — parked question expired, telling him and routing the words fresh")
return clarifyExpired
return clarifyExpiredLine()
}
// withNotice glues the expiry notice in front of this turn's reply. One turn
+10 -7
View File
@@ -56,10 +56,13 @@ func TestClarifyQuestionForMissingSlot(t *testing.T) {
want string
asked bool
}{
{"reminder without a time", clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"), "На когда напомнить?", true},
{"reminder without a time", clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"), "Когда?", true},
{"fact without a key", clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши"), "Что записать?", true},
{"act without a fn", clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это"), "Что сделать?", true},
{"reminder that already has a time", clarifyDec(router.IntentReminder, router.Slots{HasTime: true}, "напомни в 11"), "", false},
// A time with nothing to say at that time is still half a reminder, so
// the subject is what she asks about — not silence.
{"reminder that has a time but no subject", clarifyDec(router.IntentReminder, router.Slots{HasTime: true}, "напомни в 11"), "О чём напомнить?", true},
{"reminder that has both", clarifyDec(router.IntentReminder, router.Slots{Text: "позвонить маме", HasTime: true}, "напомни в 11 позвонить маме"), "", false},
{"chat is never worth a question", clarifyDec(router.IntentChat, router.Slots{Text: "мгм"}, "мгм"), "", false},
{"query is never worth a question", clarifyDec(router.IntentQuery, router.Slots{Text: "а"}, "а"), "", false},
}
@@ -78,7 +81,7 @@ func TestClarifyReminderCompletesOnAnswer(t *testing.T) {
h, st, _ := newClarifyHandler(t)
question, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"))
if !asked || question != "На когда напомнить?" {
if !asked || question != "Когда?" {
t.Fatalf("expected the time question, got %q asked=%v", question, asked)
}
@@ -152,7 +155,7 @@ func TestClarifyAsksThreeTimesThenSaysSo(t *testing.T) {
if !handled {
t.Fatalf("answer %d must be consumed as an answer", i)
}
if reply != "На когда напомнить?" {
if reply != "Когда?" {
t.Fatalf("attempt %d should ask again, got %q", i, reply)
}
if h.clarifyStore.Get(voiceDialogueID, h.now()) == nil {
@@ -304,17 +307,17 @@ func TestClarifyExpiryIsAnnouncedAndWordsStillRoute(t *testing.T) {
*now = now.Add(clarifyTTL + time.Second)
reply := h.handleText(ctx, "как дела")
if !strings.HasPrefix(reply, clarifyExpired) {
if !isClarifyExpired(reply) {
t.Fatalf("expired question must be announced first, got %q", reply)
}
if strings.TrimSpace(strings.TrimPrefix(reply, clarifyExpired)) == "" {
if trimClarifyExpired(reply) == "" {
t.Fatalf("the new words must still be answered, got only the notice: %q", reply)
}
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
t.Fatal("the expired question must be gone")
}
// The notice is said once, not on every later utterance.
if reply := h.handleText(ctx, "как дела"); strings.Contains(reply, clarifyExpired) {
if reply := h.handleText(ctx, "как дела"); isClarifyExpired(reply) {
t.Fatalf("notice repeated on a later turn: %q", reply)
}
}
+216
View File
@@ -0,0 +1,216 @@
package main
import (
"context"
"log"
"strings"
"time"
"github.com/kami/maven/internal/router"
)
// pendingHexisExec — a mutating Hexis capability parked awaiting a spoken
// confirm. The confirmation is bound to the resolved capability + canonical
// target entity so a later "да" can only execute exactly what was proposed
// (ecosystem invariant: protected actions require bound confirmation).
type pendingHexisExec struct {
capabilityID string
capName string
entityID string
displayName string
expiry time.Time
}
// pendingRoutineConfirm — a proposed routine awaiting a spoken y/n to become
// a recurring reminder. Set by detectPattern after creating a proposal.
type pendingRoutineConfirm struct {
routineID int64
action string
object string
interval float64
phrase string
expiry time.Time
}
// pendingAct — a destructive act awaiting a spoken confirm.
type pendingAct struct {
fn string
args []string
phrase string
expiry time.Time
}
// confirmTTL — how long a parked destructive confirm stays answerable. Short:
// a confirm is a same-breath gesture; a stale prompt shouldn't fire on an
// unrelated later "да".
const confirmTTL = 90 * time.Second
// park stores a destructive act awaiting confirmation. Overwrites any prior
// pending (last-asked wins — single-user box).
func (h *reactiveHandler) park(fn string, args []string, phrase string) {
h.mu.Lock()
h.pending = &pendingAct{fn: fn, args: args, phrase: phrase, expiry: h.now().Add(confirmTTL)}
h.mu.Unlock()
}
// resolveConfirm interprets an utterance as the answer to a parked destructive
// act OR a parked routine proposal. Returns (reply, true) when it consumed the
// utterance as a y/n answer; ("", false) when there's nothing pending (or the
// parked act expired), so the caller routes the utterance normally. An
// unrecognised answer cancels the pending and routes normally — a confirm that
// can't be answered clearly is safer abandoned than left armed.
func (h *reactiveHandler) resolveConfirm(ctx context.Context, text string) (string, bool) {
h.mu.Lock()
defer h.mu.Unlock()
for _, r := range h.confirmResolvers(ctx) {
if !r.claim() {
continue
}
// The slot is already cleared by claim(): every branch below drops the
// pending, including the unclear one — a confirm that can't be
// answered clearly is safer abandoned than left armed.
switch classifyConfirm(text) {
case confirmYes:
return r.yes(), true
case confirmNo:
return r.no(), true
default:
return "", false
}
}
return "", false
}
// confirmResolver — one parked-confirm slot in the chain. claim() reports
// whether this slot holds a live pending, taking it (and dropping an expired
// one) as it goes; yes/no then run the answer. Only ever called with h.mu held.
type confirmResolver struct {
claim func() bool
yes func() string
no func() string
}
// confirmResolvers builds the ordered chain resolveConfirm walks. Order is
// deliberate: the routine proposal is checked before the tool confirm so a
// routine confirm doesn't get eaten by a stale tool pending.
func (h *reactiveHandler) confirmResolvers(ctx context.Context) []confirmResolver {
var pr *pendingRoutineConfirm
var hx *pendingHexisExec
var p *pendingAct
return []confirmResolver{
// Routine proposal.
{
claim: func() bool {
pr, h.pendingRoutine = h.pendingRoutine, nil
return pr != nil && !h.now().After(pr.expiry)
},
yes: func() string {
// Only record the acceptance. The tick loop reads accepted
// routines and nudges on their own interval. Building a
// reminder here made a routine fire exactly once (Vikunja #366).
if err := h.dataStore.AcceptProposedRoutine(ctx, pr.routineID, h.now()); err != nil {
log.Printf("voice: accept proposed routine: %v", err)
return "не получилось запомнить рутину."
}
return "буду напоминать."
},
no: func() string {
if err := h.dataStore.DismissProposedRoutine(ctx, pr.routineID); err != nil {
log.Printf("voice: dismiss proposed routine: %v", err)
}
return "хорошо, не буду."
},
},
// Hexis execution confirm. Bound to the exact capability + target that
// was proposed; a stray "да" can only run that, nothing else.
{
claim: func() bool {
hx, h.pendingHexis = h.pendingHexis, nil
return hx != nil && !h.now().After(hx.expiry)
},
yes: func() string {
return h.execHexis(ctx, hx.capabilityID, hx.capName, hx.entityID, hx.displayName)
},
no: func() string { return "отменила." },
},
// Tool confirm.
{
claim: func() bool {
p, h.pending = h.pending, nil
return p != nil && !h.now().After(p.expiry)
},
yes: func() string {
out, err := h.tools.Exec(ctx, p.fn, p.args, true) // confirmed
if err != nil {
log.Printf("voice: tool %s (confirmed): %v", p.fn, err)
if out != "" {
return "не получилось выполнить команду: " + firstLine(out)
}
return "не получилось выполнить команду."
}
if out != "" {
return "готово: " + firstLine(out)
}
return "готово."
},
no: func() string { return "отменила." },
},
}
}
// proposeGap scaffolds a 'proposed' tool for an act whose verb isn't enabled.
// maven drafts the registration (name = the verb, provenance = the utterance);
// a human enables it on the authed surface. She suggests, never enables.
func (h *reactiveHandler) proposeGap(ctx context.Context, dec router.Decision) string {
name := firstWord(stripWake(dec.Utterance))
if name == "" {
return "не разобрала команду — попробуй иначе."
}
newly, err := h.api.ProposeTool(ctx, name, dec.Utterance, "", h.now())
if err != nil {
log.Printf("voice: propose tool %q: %v", name, err)
return "команды «" + name + "» нет в списке разрешённых."
}
if newly {
return "команды «" + name + "» нет в списке. Предложила её добавить — включи через клиент."
}
return "команды «" + name + "» пока нет в списке — она уже предложена, включи через клиент."
}
// confirmVerdict — the parse of a y/n confirm answer.
type confirmVerdict int
const (
confirmUnknown confirmVerdict = iota
confirmYes
confirmNo
)
// classifyConfirm reads a short ru/en yes-or-no answer. Substring match on the
// stems so inflections/fillers ("да, давай", "нет, отмени") still land.
func classifyConfirm(text string) confirmVerdict {
t := strings.ToLower(strings.TrimSpace(text))
// negatives first — "не надо" contains no "да", but check no-stems before
// yes so a leading "нет" isn't shadowed.
for _, no := range []string{"нет", "не надо", "отмен", "стоп", "no", "cancel", "stop", "don't"} {
if strings.Contains(t, no) {
return confirmNo
}
}
for _, yes := range []string{"да", "ага", "давай", "подтвер", "конечно", "yes", "yeah", "yep", "confirm", "ок", "okay", "ok"} {
if strings.Contains(t, yes) {
return confirmYes
}
}
return confirmUnknown
}
// actPhrase renders "fn arg1 arg2" for the confirm prompt.
func actPhrase(fn string, args []string) string {
if len(args) == 0 {
return fn
}
return fn + " " + strings.Join(args, " ")
}
+158
View File
@@ -0,0 +1,158 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/store"
)
// Vikunja #281 — the fourth delivery outcome: a care candidate the restraint
// gate suppresses (quiet hours / away / calendar-busy) is not necessarily
// lost. If it's worth resurfacing (loop.DigestEligible), it's durably held
// (internal/store's digest_entries) and spoken as one bundle once speaking
// is appropriate again — never while the suppression reason still holds.
func breakTrace(blockedBy string) *loop.TickTrace {
return &loop.TickTrace{
RuleTraces: []loop.RuleTrace{{
RuleName: "break",
Severity: loop.Sev2,
PredicateResult: true,
GateResult: false,
GateBlockedBy: blockedBy,
}},
}
}
// TestSuppressedCareDigestsAcrossQuietHours — a Sev2 care candidate blocked
// by quiet hours is enqueued into the durable digest, and is spoken as a
// "digest" nudge only once quiet hours actually end — never while still
// suppressed (that would just be a second way to nag through quiet hours).
func TestSuppressedCareDigestsAcrossQuietHours(t *testing.T) {
st := newTestStore(t)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
ctx := context.Background()
now := refNow()
quiet := loop.State{Now: now, QuietHours: true, Presence: store.Present}
tl.enqueueSuppressedDigest(ctx, breakTrace("quiet_hours"), quiet, now)
entries, err := st.PendingDigestEntries(ctx, now)
if err != nil {
t.Fatalf("pending: %v", err)
}
if len(entries) != 1 || entries[0].Rule != "break" {
t.Fatalf("want 1 pending digest entry for break, got %+v", entries)
}
// still quiet hours: draining now must not speak — the same restraint
// that suppressed the live nudge must suppress the bundle too.
tl.maybeDrainDigest(ctx, quiet, now)
if len(sink.sends) != 0 {
t.Fatalf("digest must not drain while quiet hours holds, got %+v", sink.sends)
}
// quiet hours end: this is the moment speaking is appropriate again.
after := now.Add(time.Hour)
clear := loop.State{Now: after, QuietHours: false, Presence: store.Present}
tl.maybeDrainDigest(ctx, clear, after)
if len(sink.sends) != 1 {
t.Fatalf("want exactly 1 dispatched digest bundle, got %d: %+v", len(sink.sends), sink.sends)
}
if sink.sends[0].RuleName != "digest" {
t.Fatalf("want RuleName digest, got %q", sink.sends[0].RuleName)
}
remaining, err := st.PendingDigestEntries(ctx, after)
if err != nil {
t.Fatalf("pending after drain: %v", err)
}
if len(remaining) != 0 {
t.Fatalf("drained entry must no longer be pending, got %+v", remaining)
}
}
// TestSuppressedCareDigestDedupesAcrossTicks — quiet hours holding for
// several ticks must not enqueue several copies of the same suppressed
// nudge; he hears it once when the bundle finally drains.
func TestSuppressedCareDigestDedupesAcrossTicks(t *testing.T) {
st := newTestStore(t)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
ctx := context.Background()
now := refNow()
quiet := loop.State{Now: now, QuietHours: true, Presence: store.Present}
for i := 0; i < 3; i++ {
tl.enqueueSuppressedDigest(ctx, breakTrace("quiet_hours"), quiet, now.Add(time.Duration(i)*time.Minute))
}
entries, err := st.PendingDigestEntries(ctx, now)
if err != nil {
t.Fatalf("pending: %v", err)
}
if len(entries) != 1 {
t.Fatalf("3 suppressions of the same nudge must collapse to 1 pending entry, got %d", len(entries))
}
}
// TestSuppressedCareDigestExpiresRatherThanDeliveringLate — an entry that
// aged out before the suppression cleared is dropped, not spoken late: a
// two-day-old "you skipped a break" is noise, not news.
func TestSuppressedCareDigestExpiresRatherThanDeliveringLate(t *testing.T) {
st := newTestStore(t)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
ctx := context.Background()
now := refNow()
quiet := loop.State{Now: now, QuietHours: true, Presence: store.Present}
tl.enqueueSuppressedDigest(ctx, breakTrace("quiet_hours"), quiet, now)
// well past digestExpiry (24h) before the suppression ever clears.
stale := now.Add(48 * time.Hour)
tl.expireStaleDigest(ctx, stale)
clear := loop.State{Now: stale, QuietHours: false, Presence: store.Present}
tl.maybeDrainDigest(ctx, clear, stale)
if len(sink.sends) != 0 {
t.Fatalf("a stale digest entry must be dropped, not delivered late; got %+v", sink.sends)
}
}
// TestSuppressedCareDigestIgnoresHighSeverity — defense in depth at the
// wiring layer: even if a RuleTrace somehow showed a high-severity rule
// blocked by a care-only gate reason, the tick driver must not durably
// digest it. Alarms bypass the gate and deliver now, unchanged; they must
// never be silently delayed into a bundle.
func TestSuppressedCareDigestIgnoresHighSeverity(t *testing.T) {
st := newTestStore(t)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
ctx := context.Background()
now := refNow()
trace := &loop.TickTrace{RuleTraces: []loop.RuleTrace{{
RuleName: "service_down",
Severity: loop.Sev4,
PredicateResult: true,
GateResult: false,
GateBlockedBy: "quiet_hours",
}}}
quiet := loop.State{Now: now, QuietHours: true, Presence: store.Present}
tl.enqueueSuppressedDigest(ctx, trace, quiet, now)
entries, err := st.PendingDigestEntries(ctx, now)
if err != nil {
t.Fatalf("pending: %v", err)
}
if len(entries) != 0 {
t.Fatalf("high severity must never be digested, got %+v", entries)
}
}
+312
View File
@@ -0,0 +1,312 @@
package main
import (
"context"
"encoding/json"
"fmt"
"log"
"strings"
hexisclient "github.com/kami/hexis/pkg/client"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
)
// praxisCapability is one arm of the Praxis act dispatch. This is an interface
// rather than a map[string]func because each arm carries its own state: the
// verb aliases it answers to, the trace name it records, and its own reply
// formatting. The dispatch grows an arm per Praxis capability, so a new one is
// added to praxisCapabilities below and nothing else changes.
type praxisCapability interface {
// aliases are the verbs (router fn slots, EN and RU) this capability answers to.
aliases() []string
// handle runs the capability and returns the user-facing reply.
handle(ctx context.Context, h *reactiveHandler, px *praxisClient, dec router.Decision) string
}
// praxisCapabilities is the registry handlePraxisAct consults, in order.
var praxisCapabilities = []praxisCapability{
listAttentionCapability{},
praxisItemAction{
verbs: []string{"acknowledge_item", "принято", "понял", "поняла"},
ask: "какой пункт отметить принятым?",
op: "acknowledge",
failure: "не получилось отметить принятым.",
success: "принято.",
call: func(ctx context.Context, px *praxisClient, id string) error {
_, err := px.Acknowledge(ctx, id)
return err
},
},
praxisItemAction{
verbs: []string{"resolve_item", "сделано", "готово", "решено"},
ask: "какой пункт отметить сделанным?",
op: "resolve",
failure: "не получилось отметить сделанным.",
success: "отмечено как сделано.",
call: func(ctx context.Context, px *praxisClient, id string) error {
_, err := px.Resolve(ctx, id)
return err
},
},
praxisItemAction{
verbs: []string{"ignore_item", "игнорировать", "неважно"},
ask: "какой пункт игнорировать?",
op: "ignore",
failure: "не получилось проигнорировать.",
success: "проигнорировано.",
call: func(ctx context.Context, px *praxisClient, id string) error {
_, err := px.Ignore(ctx, id)
return err
},
},
praxisItemAction{
verbs: []string{"pin_item", "закрепить"},
ask: "какой пункт закрепить?",
op: "pin",
failure: "не получилось закрепить.",
success: "закреплено.",
call: func(ctx context.Context, px *praxisClient, id string) error {
_, err := px.Pin(ctx, id, true)
return err
},
},
listChangesCapability{},
}
// handlePraxisAct — dispatches ecosystem tool acts through the Praxis tools API.
// Returns "" when the act is not a Praxis verb (the caller falls through to the
// system command executor). Returns a reply string otherwise.
func (h *reactiveHandler) handlePraxisAct(ctx context.Context, dec router.Decision) string {
if h.ecosystem == nil || h.ecosystem.praxis == nil {
return ""
}
px := h.ecosystem.praxis
for _, capability := range praxisCapabilities {
for _, alias := range capability.aliases() {
if alias == dec.Slots.Fn {
return capability.handle(ctx, h, px, dec)
}
}
}
// Not a Praxis verb — let the caller fall through.
return ""
}
// praxisItemAction is the shared shape of the item-lifecycle capabilities: take
// an item id from the value slot, call one Praxis endpoint, trace the result.
type praxisItemAction struct {
verbs []string
ask string // reply when no item id was given
op string // trace + log name of the operation
failure string // reply when the Praxis call errors
success string
call func(ctx context.Context, px *praxisClient, id string) error
}
func (a praxisItemAction) aliases() []string { return a.verbs }
func (a praxisItemAction) handle(ctx context.Context, h *reactiveHandler, px *praxisClient, dec router.Decision) string {
id := dec.Slots.Value
if id == "" {
return a.ask
}
if err := a.call(ctx, px, id); err != nil {
log.Printf("ecosystem: praxis %s %s: %v", a.op, id, err)
return a.failure
}
h.recordPraxisTrace(ctx, a.op, map[string]any{"item_id": id})
return a.success
}
// listAttentionCapability reads the attention digest and surfaces every item it speaks.
type listAttentionCapability struct{}
func (listAttentionCapability) aliases() []string {
return []string{"list_attention", "attention", "внимание", "что требует внимания", "что нового"}
}
func (listAttentionCapability) handle(ctx context.Context, h *reactiveHandler, px *praxisClient, _ router.Decision) string {
items, err := px.ListAttention(ctx, 20)
if err != nil {
log.Printf("ecosystem: praxis attention: %v", err)
return "не могу сейчас узнать, что требует внимания."
}
if len(items) == 0 {
return "ничего не требует внимания."
}
h.recordPraxisTrace(ctx, "list_attention", map[string]any{"count": len(items)})
var parts []string
for _, item := range items {
title, _ := item["title"].(string)
// importance arrives as JSON number ⇒ float64 over the HTTP contract.
importance, _ := item["importance"].(float64)
rule, _ := item["rule"].(string)
s := title
if importance > 0 {
s += fmt.Sprintf(" (важность %d", int(importance))
if rule != "" {
s += ": " + rule
}
s += ")"
}
parts = append(parts, s)
// Speaking an item surfaces it, it does not acknowledge it
// (ECOSYSTEM-SPEC.md §2.3: surfaced != acknowledged). Best-effort:
// a failed surface call must not block delivering the digest.
if id, ok := item["id"].(string); ok && id != "" {
if _, err := px.Surface(ctx, id); err != nil {
log.Printf("ecosystem: praxis surface %s: %v", id, err)
}
}
}
return "требует внимания: " + strings.Join(parts, "; ")
}
// listChangesCapability reads the recent-changes feed.
type listChangesCapability struct{}
func (listChangesCapability) aliases() []string {
return []string{"list_changes", "changes", "изменения", "что изменилось"}
}
func (listChangesCapability) handle(ctx context.Context, h *reactiveHandler, px *praxisClient, _ router.Decision) string {
changes, err := px.ListChanges(ctx, 20)
if err != nil {
log.Printf("ecosystem: praxis changes: %v", err)
return "не могу сейчас узнать об изменениях."
}
if len(changes) == 0 {
return "нет изменений."
}
h.recordPraxisTrace(ctx, "list_changes", map[string]any{"count": len(changes)})
var parts []string
for _, c := range changes {
title, _ := c["title"].(string)
typ, _ := c["change_type"].(string)
parts = append(parts, fmt.Sprintf("%s (%s)", title, typ))
}
return "изменения: " + strings.Join(parts, "; ")
}
// recordPraxisTrace — writes a fact recording a cross-service ecosystem call.
// The fact is stored with source "praxis:trace" so the proactive loop can
// reference it and the dashboard can display recent ecosystem activity.
func (h *reactiveHandler) recordPraxisTrace(ctx context.Context, operation string, details map[string]any) {
now := h.now()
value := operation
if len(details) > 0 {
if b, err := json.Marshal(details); err == nil {
value = operation + " " + string(b)
}
}
_, _ = h.api.WriteFact(ctx, ipc.WriteFactReq{
Ts: now,
Kind: "system",
Key: "praxis:" + operation,
Value: value,
Source: "praxis:trace",
Confidence: 1.0,
})
}
// handleHexisAct — resolves entity references through Nexus and executes
// matching capabilities through Hexis. Returns a reply string when handled,
// or "" to fall through to the system command executor.
func (h *reactiveHandler) handleHexisAct(ctx context.Context, dec router.Decision) string {
if h.ecosystem == nil {
return ""
}
// Resolve the utterance text as an entity reference through Nexus. An
// ambiguous match must stop and clarify — never guess a mutation target.
entityID, displayName, ambiguous, err := h.ecosystem.resolveEntityReference(ctx, dec.Slots.Text, nil)
if err != nil {
// A genuine Nexus dependency failure, not "no such entity" — stop here
// and report degradation rather than silently falling through to the
// local command executor (ECOSYSTEM-SPEC.md: services degrade
// independently, never a silent all-clear).
return "экосистема недоступна, попробуй ещё раз."
}
if len(ambiguous) > 0 {
return "уточни, что именно: " + strings.Join(ambiguous, ", ") + "?"
}
if entityID == "" {
return ""
}
// Discover Hexis capabilities for this entity. A resolved entity with a
// genuine Hexis failure must not be treated as "no capabilities" and
// fall through to unrelated local execution.
caps, err := h.ecosystem.discoverCapabilities(ctx, entityID)
if err != nil {
return "экосистема недоступна, попробуй ещё раз."
}
if len(caps) == 0 {
return ""
}
// Match the user's verb to a capability by name/description. Collect all
// matches: more than one is itself ambiguous, so we ask rather than pick
// the first (ecosystem invariant: no arbitrary target for mutation).
verb := dec.Slots.Fn
if verb == "" {
verb = dec.Slots.Text
}
verbLower := strings.ToLower(verb)
var matches []*hexisclient.Capability
for i, c := range caps {
if strings.Contains(strings.ToLower(c.Name), verbLower) ||
(c.Description != "" && strings.Contains(strings.ToLower(c.Description), verbLower)) {
matches = append(matches, &caps[i])
}
}
if len(matches) == 0 {
return ""
}
if len(matches) > 1 {
var names []string
for _, m := range matches {
names = append(names, m.Name)
}
return "какую команду для " + displayName + ": " + strings.Join(names, ", ") + "?"
}
matched := matches[0]
// Read-only capabilities run immediately; mutating ones are parked for an
// explicit spoken confirm bound to this capability + target.
if !matched.ReadOnly {
h.mu.Lock()
h.pendingHexis = &pendingHexisExec{
capabilityID: matched.ID,
capName: matched.Name,
entityID: entityID,
displayName: displayName,
expiry: h.now().Add(confirmTTL),
}
h.mu.Unlock()
return "выполнить «" + matched.Name + "» для " + displayName + "? скажи «да» или «нет»."
}
return h.execHexis(ctx, matched.ID, matched.Name, entityID, displayName)
}
// execHexis runs a resolved capability and records a cross-service trace with
// the correlation ID. It reports command success, never operational recovery
// (Praxis observes recovery independently).
func (h *reactiveHandler) execHexis(ctx context.Context, capID, capName, entityID, displayName string) string {
correlationID, err := h.ecosystem.executeCapability(ctx, capID, entityID, nil)
if err != nil {
log.Printf("ecosystem: hexis execute error (cor=%s): %v", correlationID, err)
return "не получилось выполнить команду для " + displayName + "."
}
h.recordPraxisTrace(ctx, "hexis:"+capName, map[string]any{
"entity_id": entityID,
"entity_name": displayName,
"capability": capName,
"correlation_id": correlationID,
})
return "команда выполнена для " + displayName + "."
}
+62 -115
View File
@@ -58,6 +58,7 @@ import (
"github.com/kami/maven/internal/delivery/telegramsink"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/webauthn"
@@ -96,100 +97,12 @@ func main() {
}
}
// lockedAPI is a dummy CoreAPI used while the daemon is locked. Every method
// returns errLocked. The wire protocol's StoreAPI methods all go through the
// Server dispatch on CoreAPI, so returning errLocked from each is correct.
type lockedAPI struct{}
var _ ipc.CoreAPI = (*lockedAPI)(nil)
func (l *lockedAPI) WriteFact(ctx context.Context, req ipc.WriteFactReq) (int64, error) {
return 0, errLocked
}
func (l *lockedAPI) LatestFact(ctx context.Context, key string) (ipc.Fact, error) {
return ipc.Fact{}, errLocked
}
func (l *lockedAPI) LatestFactBySource(ctx context.Context, key, source string) (ipc.Fact, error) {
return ipc.Fact{}, errLocked
}
func (l *lockedAPI) Since(ctx context.Context, key string, now time.Time) (time.Duration, error) {
return 0, errLocked
}
func (l *lockedAPI) Presence(ctx context.Context) (ipc.Presence, error) {
return ipc.Presence{}, errLocked
}
func (l *lockedAPI) CreateReminder(ctx context.Context, fire time.Time, payload, cron string) (int64, error) {
return 0, errLocked
}
func (l *lockedAPI) MarkReminder(ctx context.Context, id int64, status string) error {
return errLocked
}
func (l *lockedAPI) ListReminders(ctx context.Context, n int) ([]ipc.Reminder, error) {
return nil, errLocked
}
func (l *lockedAPI) RecordNudge(ctx context.Context, rule, channel, message string, ts time.Time) (int64, error) {
return 0, errLocked
}
func (l *lockedAPI) ResolveNudge(ctx context.Context, id int64, outcome string, ts time.Time) error {
return errLocked
}
func (l *lockedAPI) RecentOutcomes(ctx context.Context, rule string, n int) ([]string, error) {
return nil, errLocked
}
func (l *lockedAPI) RecentFacts(ctx context.Context, n int) ([]ipc.Fact, error) {
return nil, errLocked
}
func (l *lockedAPI) CalendarEvents(ctx context.Context, from, to time.Time) ([]ipc.Fact, error) {
return nil, errLocked
}
func (l *lockedAPI) RecentNudges(ctx context.Context, n int) ([]ipc.Nudge, error) {
return nil, errLocked
}
func (l *lockedAPI) WriteNote(ctx context.Context, ts time.Time, text string, embedding []float32, source string) (int64, error) {
return 0, errLocked
}
func (l *lockedAPI) QueryNotes(ctx context.Context, embedding []float32, k int) ([]ipc.Note, error) {
return nil, errLocked
}
func (l *lockedAPI) RecentNotes(ctx context.Context, n int) ([]ipc.Note, error) {
return nil, errLocked
}
func (l *lockedAPI) ProposeTool(ctx context.Context, name, utterance, scope string, ts time.Time) (bool, error) {
return false, errLocked
}
func (l *lockedAPI) EnableTool(ctx context.Context, name string, cmd []string, destructive bool, scope string, ts time.Time) error {
return errLocked
}
func (l *lockedAPI) DisableTool(ctx context.Context, name string) error { return errLocked }
func (l *lockedAPI) DeleteTool(ctx context.Context, name string) error { return errLocked }
func (l *lockedAPI) ListProposedRoutines(ctx context.Context) ([]ipc.ProposedRoutine, error) {
return nil, errLocked
}
func (l *lockedAPI) DismissProposedRoutine(ctx context.Context, id int64) error { return errLocked }
func (l *lockedAPI) AcceptProposedRoutine(ctx context.Context, id int64) error {
return errLocked
}
func (l *lockedAPI) LookupTool(ctx context.Context, name string) (ipc.Tool, error) {
return ipc.Tool{}, errLocked
}
func (l *lockedAPI) ListTools(ctx context.Context, status string) ([]ipc.Tool, error) {
return nil, errLocked
}
func (l *lockedAPI) RevertFact(ctx context.Context, key string) (int64, error) { return 0, errLocked }
func (l *lockedAPI) Chat(ctx context.Context, text string) (string, error) {
return "", errLocked
}
func (l *lockedAPI) TickTrace(ctx context.Context) (ipc.TickTrace, error) {
return ipc.TickTrace{}, errLocked
}
func (l *lockedAPI) MorningStatus(ctx context.Context) ([]ipc.MorningRoutineStatus, error) {
return nil, errLocked
}
func run(args []string) error {
cfgPath := flag.String("config", defaultConfigPath(), "path to mavend JSON config")
wrappedKeyPath := flag.String("wrapped-key-file", "", "path to wrapped encryption key blob (enables cold-start unlock)")
reembed := flag.Bool("reembed", false, "re-embed every stored note and fact with the configured embedder, then serve normally (run once after an embedder swap)")
flag.CommandLine.Parse(args)
reembedOnStart = *reembed
cfg, err := config.Load(*cfgPath)
if err != nil {
return err
@@ -269,13 +182,14 @@ func run(args []string) error {
phr = phraser.NewStub()
if cfg.Phraser != nil {
pc := phraser.Config{
ModelPath: cfg.Phraser.ModelPath,
BinPath: cfg.Phraser.BinPath,
Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
Timeout: time.Duration(cfg.Phraser.Timeout),
Persona: personaFromCfg(cfg),
ModelPath: cfg.Phraser.ModelPath,
BinPath: cfg.Phraser.BinPath,
Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
Timeout: time.Duration(cfg.Phraser.Timeout),
LLMNudges: cfg.Phraser.LLMNudges,
ContextBlock: contextBlockFn(cfg, time.Now),
}
if pc.BinPath == "" {
pc.BinPath = "llama-server"
@@ -360,8 +274,13 @@ func run(args []string) error {
api.chatFn = voiceW.handler.handleText
}
} else {
// locked mode: dummy CoreAPI that returns errLocked for everything
coreAPI = &lockedAPI{}
// locked mode: no real store yet, so there's no meaningful CoreAPI to
// serve. srv.Check below is the actual guard — every CoreAPI call is
// refused before it reaches this value. This is just a safe non-nil
// placeholder: if the guard is ever bypassed by a bug, calls land
// here and fail loudly with ipc.ErrNotImplemented instead of a nil
// dereference or, worse, silently succeeding.
coreAPI = ipc.UnimplementedCoreAPI{}
}
// ----- IPC boundary (core ↔ modules) -----
@@ -372,7 +291,17 @@ func run(args []string) error {
passkeySess := webauthn.NewPasskeySession(5 * time.Minute)
// Set Server.Check — in locked mode, block everything except unlock-path methods.
// Set Server.Check — the single authorization guard, run once by
// Server.dispatch before any CoreAPI method is called (see
// internal/ipc/server.go). In locked mode this is the ONLY thing
// standing between an unauthenticated caller and the store: it must
// default-deny, with an explicit allowlist for the two methods the
// unlock flow itself needs (MethodAssertStepUp, MethodUnlock — neither
// of which touches CoreAPI; dispatch handles them directly via
// srv.StepUp/srv.UnlockFn). Forgetting to allowlist a new unlock-path
// method fails safe (denied); forgetting to guard a new CoreAPI method
// is impossible because there is nothing left to forget — every method
// not in the allowlist is refused by construction.
if locked {
srv.Check = func(ctx context.Context, m ipc.Method, _ json.RawMessage) error {
switch m {
@@ -439,13 +368,14 @@ func run(args []string) error {
phr = phraser.NewStub()
if cfg.Phraser != nil {
pc := phraser.Config{
ModelPath: cfg.Phraser.ModelPath,
BinPath: cfg.Phraser.BinPath,
Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
Timeout: time.Duration(cfg.Phraser.Timeout),
Persona: personaFromCfg(cfg),
ModelPath: cfg.Phraser.ModelPath,
BinPath: cfg.Phraser.BinPath,
Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
Timeout: time.Duration(cfg.Phraser.Timeout),
LLMNudges: cfg.Phraser.LLMNudges,
ContextBlock: contextBlockFn(cfg, time.Now),
}
if pc.BinPath == "" {
pc.BinPath = "llama-server"
@@ -511,7 +441,7 @@ func run(args []string) error {
tl = newTickLoop(st, gatherer, dispatcher, phr, rules, tickInterval, repeatInterval, autotuneInterval, cfg.Digest, routinesFromConfig(cfg.Routines), config.MorningRoutinesFromConfig(cfg.MorningRoutines))
factWorker = newFactEnrichmentWorker(st, eco, time.Duration(cfg.FactEnrichmentInterval))
// Swap the CoreAPI from lockedAPI to the real store adapter.
// Swap the CoreAPI from the locked placeholder to the real store adapter.
newAPI := &daemonAPI{
CoreAPI: ipc.NewStoreAPI(st),
getTrace: tl.trace,
@@ -599,12 +529,29 @@ func run(args []string) error {
return nil
}
// personaFromCfg extracts the voice persona from the config, or returns ""
// when voice isn't configured. Used to pass a character prompt into the
// LLM phraser without requiring voice to be enabled.
func personaFromCfg(cfg *config.Config) string {
if cfg.Voice != nil {
return cfg.Voice.Persona
// personaFacts reads the optional, deployment-specific facts (his name, his
// city, the free-text persona string) out of the config. Everything here may
// be empty — the context block is correct without any of it.
func personaFacts(cfg *config.Config) persona.Facts {
f := persona.Facts{
// Telegram lives outside the voice block, so it counts either way.
Telegram: cfg.Telegram != nil && cfg.Telegram.BotToken != "" && cfg.Telegram.ChatID != "",
}
return ""
if cfg.Voice == nil {
return f
}
f.OwnerName = cfg.Voice.OwnerName
f.City = cfg.Voice.City
f.Static = cfg.Voice.Persona
// Same test wireVoice uses to pick the real provider over the stub.
f.Weather = cfg.Voice.Weather != nil && cfg.Voice.Weather.Provider == "open-meteo"
f.Tools = len(cfg.Voice.Tools) > 0
return f
}
// contextBlockFn returns the per-turn renderer of the shared context block.
// Per turn, not once at startup, because the block states the current time.
func contextBlockFn(cfg *config.Config, now func() time.Time) func() string {
f := personaFacts(cfg)
return func() string { return f.Block(now()) }
}
+74
View File
@@ -0,0 +1,74 @@
// mavend/patterns.go — the shared detect+propose step of pattern inference
// (Vikunja #43). Event *extraction* (fact -> action/object) happens at fact-
// write time in voice.go's detectPattern, tied to whichever channel wrote the
// fact. Detection — turning a run of events into a proposed routine — is
// channel-agnostic: it only needs what's already in the events table, so it
// runs both right after a voice fact-write (for the immediate "напоминать?"
// confirmation) and, proactively, from the digestion tick (tick.go's
// detectPatterns) over every action+object pair on record, not just the one
// that was just talked about.
package main
import (
"context"
"errors"
"fmt"
"time"
"github.com/kami/maven/internal/pattern"
"github.com/kami/maven/internal/store"
)
// detectAndPropose runs the pattern detector over every recorded event for
// action+object and, if a stable pattern is found and nothing has been
// proposed/accepted/dismissed for this pair yet, creates a proposed_routines
// row. Returns (nil, 0, nil) — not an error — whenever there is nothing new
// to report: too few events, irregular intervals, or a pair that already has
// a row in any status. That last case is the one that matters most: it is
// how a routine the owner already DISMISSED stays dismissed forever, because
// the row survives dismissal (status flips in place, see
// store.DismissProposedRoutine) and both the Lookup check here and the
// table's UNIQUE(action, object) constraint refuse to create a second one.
func detectAndPropose(ctx context.Context, ds *store.Store, action, object string, ts time.Time) (*pattern.ProposedRoutine, int64, error) {
events, err := ds.EventsFor(ctx, action, object)
if err != nil {
return nil, 0, fmt.Errorf("events for %s/%s: %w", action, object, err)
}
patEvents := make([]pattern.Event, len(events))
for i, e := range events {
patEvents[i] = pattern.Event{
FactID: e.FactID,
Action: e.Action,
Object: e.Object,
Ts: e.Ts,
}
}
r, err := pattern.Detect(patEvents)
if err != nil {
return nil, 0, fmt.Errorf("detect %s/%s: %w", action, object, err)
}
if r == nil {
return nil, 0, nil // not enough data or intervals too irregular
}
// Belt: check first so the common "nothing new" case never even attempts
// an insert. Suspenders: CreateProposedRoutine's ON CONFLICT DO NOTHING
// (backed by the UNIQUE(action,object) constraint) is the actual
// guarantee — this Lookup is an optimization, not the source of truth.
existing, err := ds.LookupProposedRoutine(ctx, r.Action, r.Object)
if err != nil {
return nil, 0, fmt.Errorf("lookup proposed routine %s/%s: %w", action, object, err)
}
if existing != nil {
return nil, 0, nil // already proposed, accepted, or dismissed — say nothing
}
id, err := ds.CreateProposedRoutine(ctx, r.Action, r.Object, r.IntervalDays, ts)
if err != nil {
if errors.Is(err, store.ErrProposedRoutineExists) {
return nil, 0, nil // lost a race with another caller — not an error
}
return nil, 0, fmt.Errorf("create proposed routine %s/%s: %w", action, object, err)
}
return r, id, nil
}
+122
View File
@@ -0,0 +1,122 @@
package main
import (
"context"
"database/sql"
"testing"
"time"
"github.com/kami/maven/internal/store"
)
// seedRefillEvents writes N weekly "refill/cat_water" events straight to the
// events table — this is what the tick reads, independent of any utterance.
func seedRefillEvents(t *testing.T, st *store.Store, ctx context.Context, base time.Time, n int) {
t.Helper()
for i := 0; i < n; i++ {
factID, err := st.WriteFact(ctx, base.Add(time.Duration(i)*7*24*time.Hour), store.KindSelf,
"cat_water", "refill", "test", 1.0, sql.NullInt64{})
if err != nil {
t.Fatalf("write fact %d: %v", i, err)
}
if _, err := st.CreateEvent(ctx, factID, "refill", "cat_water", base.Add(time.Duration(i)*7*24*time.Hour)); err != nil {
t.Fatalf("create event %d: %v", i, err)
}
}
}
// TestTickDetectsPatternFromStoredEvents proves the tick notices a pattern on
// its own, reading straight from the store — not as a side effect of a live
// utterance (Vikunja #43). Three weekly events with no voice turn in sight
// must produce exactly one proposed routine.
func TestTickDetectsPatternFromStoredEvents(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, 3)
tl := newTestTickLoop(t, st, &fakeSink{}, nil)
tl.detectPatterns(ctx, now)
rows, err := st.ListProposedRoutines(ctx)
if err != nil {
t.Fatalf("list proposed routines: %v", err)
}
if len(rows) != 1 {
t.Fatalf("proposed routines = %d, want 1: %+v", len(rows), rows)
}
if rows[0].Action != "refill" || rows[0].Object != "cat_water" {
t.Errorf("proposed routine = %s/%s, want refill/cat_water", rows[0].Action, rows[0].Object)
}
}
// TestTickPatternDetectionIsIdempotent proves running the tick's pattern scan
// twice does not spam a second proposal for the same pair, and that the store
// itself is what stops the duplicate (not tick-local state) — the whole point
// of the guard, since the tick has no memory of what it proposed last time.
func TestTickPatternDetectionIsIdempotent(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, 3)
tl := newTestTickLoop(t, st, &fakeSink{}, nil)
tl.detectPatterns(ctx, now)
tl.detectPatterns(ctx, now.Add(time.Hour))
rows, err := st.ListProposedRoutines(ctx)
if err != nil {
t.Fatalf("list proposed routines: %v", err)
}
if len(rows) != 1 {
t.Fatalf("proposed routines after two ticks = %d, want 1 (no duplicate): %+v", len(rows), rows)
}
}
// TestTickPatternDetectionRespectsDismissal proves the single worst failure
// mode here — a proposal the owner already said no to coming back on the next
// tick — cannot happen. Dismissal flips the row's status in place; it must
// still be there to block re-proposal.
func TestTickPatternDetectionRespectsDismissal(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, 3)
tl := newTestTickLoop(t, st, &fakeSink{}, nil)
tl.detectPatterns(ctx, now)
rows, err := st.ListProposedRoutines(ctx)
if err != nil {
t.Fatalf("list proposed routines: %v", err)
}
if len(rows) != 1 {
t.Fatalf("setup: proposed routines = %d, want 1", len(rows))
}
if err := st.DismissProposedRoutine(ctx, rows[0].ID); err != nil {
t.Fatalf("dismiss: %v", err)
}
// More events for the same pair arrive, and the tick runs again — a
// dismissed pattern must not resurface.
seedRefillEvents(t, st, ctx, now.Add(30*24*time.Hour), 3)
tl.detectPatterns(ctx, now.Add(60*24*time.Hour))
proposed, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineProposed)
if err != nil {
t.Fatalf("list proposed: %v", err)
}
if len(proposed) != 0 {
t.Fatalf("a dismissed pattern came back: %+v", proposed)
}
all, err := st.ListProposedRoutinesByStatus(ctx, "")
if err != nil {
t.Fatalf("list all: %v", err)
}
if len(all) != 1 {
t.Fatalf("total rows for the pair = %d, want 1 (still dismissed, not duplicated): %+v", len(all), all)
}
if all[0].Status != store.RoutineDismissed {
t.Errorf("status = %s, want dismissed", all[0].Status)
}
}
+158
View File
@@ -0,0 +1,158 @@
package main
import (
"context"
"fmt"
"math"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice"
)
// fixedEmbedder hands back a vector chosen per text, so a test can say exactly
// how close each stored memory is to the question. The real embedders make
// scores that are realistic but not controllable, and this test is about the
// gate, not about the embedder.
type fixedEmbedder struct{ vecs map[string][]float32 }
func (f *fixedEmbedder) Dim() int { return 4 }
func (f *fixedEmbedder) Close() error { return nil }
func (f *fixedEmbedder) Embed(_ context.Context, text string) ([]float32, error) {
v, ok := f.vecs[text]
if !ok {
return nil, fmt.Errorf("fixedEmbedder: no vector for %q", text)
}
return v, nil
}
// scoreVec builds a unit vector whose cosine against the query vector
// (1,0,0,0) is exactly score.
func scoreVec(score float64) []float32 {
rest := math.Sqrt(1 - score*score)
return []float32{float32(score), float32(rest), 0, 0}
}
// recordingPhraser remembers what the query path handed it to phrase, which is
// how the test can tell which pass produced the answer.
type recordingPhraser struct {
*phraser.Stub
notes []string
}
func (r *recordingPhraser) PhraseQuery(ctx context.Context, utterance string, notes []string) (string, error) {
r.notes = notes
return r.Stub.PhraseQuery(ctx, utterance, notes)
}
// recallCase — one stored memory: its text, how close it is to the question,
// whether it is a note or a fact, and whether the notes table holds it too.
type recallCase struct {
text string
score float64
kind string
}
// buildRecallHandler stores the given memories and returns a handler whose
// query path can be run directly. Notes go into BOTH the notes table and the
// vector index, which is what the daemon does (voice.go's IntentNote).
func buildRecallHandler(t *testing.T, question string, mems []recallCase) (*reactiveHandler, *recordingPhraser) {
t.Helper()
ctx := context.Background()
st := newTestStore(t)
emb := &fixedEmbedder{vecs: map[string][]float32{question: {1, 0, 0, 0}}}
mem := memory.NewInMemoryStore()
now := time.Now()
for i, m := range mems {
vec := scoreVec(m.score)
emb.vecs[m.text] = vec
id := fmt.Sprintf("%s:%d", m.kind, i)
if m.kind == "note" {
if _, err := st.WriteNote(ctx, now, m.text, vec, "tap:voice"); err != nil {
t.Fatalf("WriteNote: %v", err)
}
}
if err := mem.Insert(ctx, id, vec, map[string]string{"text": m.text, "type": m.kind}); err != nil {
t.Fatalf("memory insert: %v", err)
}
}
phr := &recordingPhraser{Stub: phraser.NewStub()}
h := &reactiveHandler{
api: ipc.NewStoreAPI(st),
embedder: emb,
replier: voice.NewStubReplier(),
phraser: phr,
now: func() time.Time { return now },
memStore: mem,
dataStore: st,
queryMinScore: 0.55,
queryMinMargin: 0.008,
weatherProvider: nil,
}
return h, phr
}
func askQuery(t *testing.T, h *reactiveHandler, question string) string {
t.Helper()
return h.applyAction(context.Background(), router.Decision{
Intent: router.IntentQuery,
Utterance: question,
})
}
// TestQueryRecallNoteCanWin — the note-recall regression (Vikunja #373). Notes
// and facts share one vector index, and a note that clearly beats everything
// else must be the answer. Before the fix the memory pass only ran after the
// notes-only gate had already rejected the same note at the same score, so only
// a fact could ever come back from it.
func TestQueryRecallNoteCanWin(t *testing.T) {
const q = "где молоко"
t.Run("a clearly best note answers", func(t *testing.T) {
h, phr := buildRecallHandler(t, q, []recallCase{
{text: "молоко стоит в холодильнике", score: 0.90, kind: "note"},
{text: "выучил пару аккордов", score: 0.50, kind: "note"},
})
reply := askQuery(t, h, q)
if want := "вот что я нашла: молоко стоит в холодильнике"; reply != want {
t.Errorf("reply %q, want %q", reply, want)
}
// One text, the winning memory's — the answer came from the memory
// pass, not from handing the phraser every note in the table.
if len(phr.notes) != 1 || phr.notes[0] != "молоко стоит в холодильнике" {
t.Errorf("phraser got %q, want just the recalled note", phr.notes)
}
})
// The other half of "one gate over everything": a fact that matches better
// than the best note now answers, instead of losing to a note that only had
// to beat other notes.
t.Run("the better-matching fact answers", func(t *testing.T) {
h, _ := buildRecallHandler(t, q, []recallCase{
{text: "молоко стоит в холодильнике", score: 0.80, kind: "note"},
{text: "купил молоко в среду", score: 0.95, kind: "fact"},
})
if reply := askQuery(t, h, q); reply != "купил молоко в среду" {
t.Errorf("reply %q, want the fact read back", reply)
}
})
// The gate is untouched: two memories this close mean the embedder cannot
// tell them apart, and silence still beats a coin flip.
t.Run("no clear best stays silent", func(t *testing.T) {
h, _ := buildRecallHandler(t, q, []recallCase{
{text: "молоко стоит в холодильнике", score: 0.860, kind: "note"},
{text: "молоко закончилось", score: 0.858, kind: "note"},
})
if reply := askQuery(t, h, q); reply != "не знаю." {
t.Errorf("reply %q, want silence", reply)
}
})
}
+114
View File
@@ -0,0 +1,114 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
)
// quietFakeAPI records the WriteFact the toggle performs.
type quietFakeAPI struct {
ipc.UnimplementedCoreAPI
got ipc.WriteFactReq
call int
}
func (a *quietFakeAPI) WriteFact(_ context.Context, req ipc.WriteFactReq) (int64, error) {
a.got, a.call = req, a.call+1
return 1, nil
}
// quietVerdict — what a phrase should do to the setting.
type quietVerdict int
const (
quietNone quietVerdict = iota
quietOn
quietOff
)
func TestResolveQuietToggle(t *testing.T) {
cases := []struct {
text string
want quietVerdict
}{
// ON vocabulary.
{"quiet on", quietOn},
{"quiet mode", quietOn},
{"тихий режим", quietOn},
{"тихий", quietOn},
{"не шуми", quietOn},
{"не беспокоить", quietOn},
{"тихо", quietOn},
// ON, inflected / embedded in a sentence.
{"включи тихий режим", quietOn},
{"побудь в тихом режиме", quietOn},
{"Тихий Режим!", quietOn},
{"тихая", quietOn},
// OFF vocabulary — all seven, incl. the three that used to say ON.
{"quiet off", quietOff},
{"quiet end", quietOff},
{"громкий режим", quietOff},
{"шумный режим", quietOff},
{"отмени тихий", quietOff},
{"выключи тихий", quietOff},
{"не тихо", quietOff},
// OFF wins over the ON words it contains.
{"выключи тихий режим", quietOff},
{"отмени тихий режим пожалуйста", quietOff},
{"верни громкий режим", quietOff},
// False positives: "тихо"/"тихий" as ordinary Russian.
{"очень тихий сегодня день", quietNone},
{"в комнате тихо", quietNone},
{"тихонько напомни", quietNone},
{"потихоньку", quietNone},
{"тихонько", quietNone},
{"он говорил тихим голосом весь вечер", quietNone},
// Unrelated.
{"напомни завтра позвонить маме", quietNone},
{"какая погода", quietNone},
{"", quietNone},
}
for _, tc := range cases {
t.Run(tc.text, func(t *testing.T) {
api := &quietFakeAPI{}
h := &reactiveHandler{api: api, now: func() time.Time { return time.Unix(0, 0).UTC() }}
reply, handled := h.resolveQuietToggle(context.Background(), tc.text)
if tc.want == quietNone {
if handled || reply != "" {
t.Fatalf("%q: got (%q, %v), want no match", tc.text, reply, handled)
}
if api.call != 0 {
t.Fatalf("%q: wrote a fact on a non-match", tc.text)
}
return
}
if !handled {
t.Fatalf("%q: not handled, want %v", tc.text, tc.want)
}
wantReply, wantVal := "тихий режим выключен.", "false"
if tc.want == quietOn {
wantReply, wantVal = "тихий режим включён. буду реже напоминать.", "true"
}
if reply != wantReply {
t.Errorf("%q: reply = %q, want %q", tc.text, reply, wantReply)
}
if api.call != 1 {
t.Fatalf("%q: WriteFact called %d times, want 1", tc.text, api.call)
}
if api.got.Kind != "config" || api.got.Key != "quiet_hours" || api.got.Source != "tap:voice" || api.got.Confidence != 1.0 {
t.Errorf("%q: request shape = %+v", tc.text, api.got)
}
if api.got.Value != wantVal {
t.Errorf("%q: value = %q, want %q", tc.text, api.got.Value, wantVal)
}
})
}
}
+17 -14
View File
@@ -2,21 +2,24 @@ package main
import "github.com/kami/maven/internal/memory"
// bestRecall is the read side of the long-term memory store: the top hit's
// stored text when it clears the confidence gate. This recalls across BOTH
// notes and facts (facts aren't in the notes table, so this is the only path
// that can answer "when did I last …?" from a captured fact). A note hit here
// is redundant with the notes-RAG path — by design; the two indexes can diverge
// once the backend is swapped for a persistent/external store. ok=false when
// the hit fails the confidence gate (see memory.Confident: an absolute floor
// plus a margin over the runner-up) or carries no text.
func bestRecall(results []memory.Result, minScore, minMargin float64) (string, bool) {
// bestRecall is the read side of the long-term memory store: the top hit when
// it clears the confidence gate. The index holds BOTH notes and facts, and
// either can win — the caller looks at the returned hit's meta["type"] to see
// which. Facts aren't in the notes table, so this is the only path that can
// answer "when did I last …?" from a captured fact.
//
// The whole hit is returned, not just its text, because "which memory answered"
// decides how the answer is said: a note gets phrased in Maven's voice, a fact
// is read back as stored.
//
// ok=false when the hit fails the confidence gate (see memory.Confident: an
// absolute floor plus a margin over the runner-up) or carries no text.
func bestRecall(results []memory.Result, minScore, minMargin float64) (memory.Result, bool) {
if !memory.Confident(results, minScore, minMargin) {
return "", false
return memory.Result{}, false
}
text := results[0].Meta["text"]
if text == "" {
return "", false
if results[0].Meta["text"] == "" {
return memory.Result{}, false
}
return text, true
return results[0], true
}
+21 -2
View File
@@ -39,8 +39,27 @@ func TestBestRecall(t *testing.T) {
if !ok {
t.Fatal("clearing hit not returned")
}
if got != "выпил воды в три часа" {
t.Errorf("wrong text: %q", got)
if got.Meta["text"] != "выпил воды в три часа" {
t.Errorf("wrong text: %q", got.Meta["text"])
}
if got.Meta["type"] != "fact" {
t.Errorf("kind lost: %q", got.Meta["type"])
}
})
// The index holds notes and facts together, so a note has to be able to win
// it — for a long time it could not (Vikunja #373).
t.Run("a note can win", func(t *testing.T) {
res := []memory.Result{
{Score: 0.86, Meta: map[string]string{"text": "молоко в холодильнике", "type": "note"}},
{Score: 0.61, Meta: map[string]string{"text": "выпил воды", "type": "fact"}},
}
got, ok := bestRecall(res, min, margin)
if !ok {
t.Fatal("clearly-best note not returned")
}
if got.Meta["type"] != "note" || got.Meta["text"] != "молоко в холодильнике" {
t.Errorf("got %v, want the note", got.Meta)
}
})
+9 -4
View File
@@ -7,6 +7,7 @@ import (
"time"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice"
)
@@ -23,13 +24,17 @@ type completer interface {
type llmReplier struct {
c completer
stub *voice.StubReplier
// block renders the shared context block per turn (who he is, the time).
// nil ⇒ the prompt stands alone.
block func() string
}
func newLLMReplier(c completer) *llmReplier {
return &llmReplier{c: c, stub: voice.NewStubReplier()}
func newLLMReplier(c completer, block func() string) *llmReplier {
return &llmReplier{c: c, stub: voice.NewStubReplier(), block: block}
}
const replySystem = `Ты — Maven, домашняя ассистентка (о себе — в женском роде). Подтверди действие РОВНО ОДНИМ коротким предложением (≤120 символов), тепло и по-русски. Не задавай вопросов, не повторяй слова, не добавляй ничего после точки. Отвечай ТОЛЬКО одним объектом JSON с полями "response" (текст) и "mood" (ровно одно из: neutral, happy, thinking, tired, confused).
const replySystem = `Ты — Maven, домашняя ассистентка (о себе — в женском роде). Владелец — мужчина, говоришь с ним на "ты", в единственном числе; никогда не "вы"/"ваш" и не "он"/"его". Подтверди действие РОВНО ОДНИМ коротким предложением (≤120 символов), по-русски, спокойно и без официальных формулировок. Не задавай вопросов, не повторяй слова, не добавляй ничего после точки. Отвечай ТОЛЬКО одним объектом JSON с полями "response" (текст) и "mood" (ровно одно из: neutral, happy, thinking, tired, confused).
Пример: {"response": "Записала, что ты выпил стакан воды.", "mood": "neutral"}
Никогда не пиши "..." в поле response.`
@@ -40,7 +45,7 @@ func (r *llmReplier) Reply(d router.Decision) string {
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
out, err := r.c.Complete(ctx, llm.Req{
System: replySystem,
System: persona.Prepend(r.block, replySystem),
User: replyContext(d),
MaxTokens: 512,
RepeatPenalty: 1.3,
+5 -5
View File
@@ -17,7 +17,7 @@ type mockCompleter struct {
func (m mockCompleter) Complete(_ context.Context, _ llm.Req) (string, error) { return m.out, m.err }
func TestLLMReplierReturnsLLMReply(t *testing.T) {
r := newLLMReplier(mockCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`})
r := newLLMReplier(mockCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`}, nil)
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
if got != "записала, кофе закончился" {
t.Errorf("got %q, want %q", got, "записала, кофе закончился")
@@ -25,7 +25,7 @@ func TestLLMReplierReturnsLLMReply(t *testing.T) {
}
func TestLLMReplierFallsBackToPlainText(t *testing.T) {
r := newLLMReplier(mockCompleter{out: "записала, кофе закончился"})
r := newLLMReplier(mockCompleter{out: "записала, кофе закончился"}, nil)
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
if got != "записала, кофе закончился" {
t.Errorf("got %q, want %q", got, "записала, кофе закончился")
@@ -33,7 +33,7 @@ func TestLLMReplierFallsBackToPlainText(t *testing.T) {
}
func TestLLMReplierFallsBackToStubOnError(t *testing.T) {
r := newLLMReplier(mockCompleter{err: errTestLLMDown})
r := newLLMReplier(mockCompleter{err: errTestLLMDown}, nil)
noteDec := router.Decision{Intent: router.IntentNote}
got := r.Reply(noteDec)
want := voice.NewStubReplier().Reply(noteDec)
@@ -43,7 +43,7 @@ func TestLLMReplierFallsBackToStubOnError(t *testing.T) {
}
func TestLLMReplierFallsBackToStubOnEmpty(t *testing.T) {
r := newLLMReplier(mockCompleter{out: ""})
r := newLLMReplier(mockCompleter{out: ""}, nil)
noteDec := router.Decision{Intent: router.IntentNote}
got := r.Reply(noteDec)
want := voice.NewStubReplier().Reply(noteDec)
@@ -53,7 +53,7 @@ func TestLLMReplierFallsBackToStubOnEmpty(t *testing.T) {
}
func TestLLMReplierClarifyUsesStub(t *testing.T) {
r := newLLMReplier(mockCompleter{out: "я всё поняла"})
r := newLLMReplier(mockCompleter{out: "я всё поняла"}, nil)
clarifyDec := router.Decision{Clarify: true}
got := r.Reply(clarifyDec)
want := voice.NewStubReplier().Reply(clarifyDec)
+180
View File
@@ -0,0 +1,180 @@
// Package main — ruwords.go holds Russian language + calendar/time formatting
// helpers used by the voice reply paths (replySystem, the reminder/routine
// phrasing, etc). Pure functions, no receivers: weekday/month name tables,
// plural agreement, clock/date rendering, and the "do I actually know this
// place/day" guards that pick an honest reply over a confidently wrong one.
// Extend this file rather than voice.go for anything in that shape.
package main
import (
"fmt"
"strconv"
"strings"
"time"
)
var ruWeekdays = []string{
"воскресенье", "понедельник", "вторник", "среда",
"четверг", "пятница", "суббота",
}
var ruMonths = []string{
"января", "февраля", "марта", "апреля", "мая", "июня",
"июля", "августа", "сентября", "октября", "ноября", "декабря",
}
// onlyLocalTimeReply — the honest answer when the user asks the time somewhere
// other than here. She only keeps one clock, and saying so is better than
// naming the wrong city's time.
//
// There used to be a city→time-zone table here. It was removed on purpose: the
// user only ever asks for local time, so the table was a second list of cities
// to keep in step with the weather one for no gain.
const onlyLocalTimeReply = "я знаю только местное время, про другие города пока не скажу."
// notPlaceAfterV — words that follow "в" without naming a place, so
// mentionsUnknownPlace does not mistake them for a city.
var notPlaceAfterV = map[string]bool{
"данный": true, "данную": true, "этот": true, "эту": true,
"котором": true, "какое": true, "какой": true, "который": true,
"общем": true, "точности": true, "курсе": true, "сутках": true,
"часах": true, "минутах": true, "секундах": true, "неделе": true,
}
// mentionsUnknownPlace reports whether the question has a "в <слово>" phrase
// that looks like a place we do not know ("который час в киеве"). Used only to
// pick the honest "local time only" reply instead of answering local time as
// if it were the city's.
func mentionsUnknownPlace(u string) bool {
toks := strings.Fields(u)
for i := 0; i+1 < len(toks); i++ {
if toks[i] != "в" && toks[i] != "во" {
continue
}
next := strings.Trim(toks[i+1], ".,?!")
if next == "" || notPlaceAfterV[next] {
continue
}
// A number after "в" is a clock ("в 5 часов"), not a place.
if _, err := strconv.Atoi(strings.SplitN(next, ":", 2)[0]); err == nil {
continue
}
return true
}
return false
}
// onlyNearDaysReply — she can work out today, tomorrow, the day after and
// yesterday, and nothing further. Said out loud instead of answering today's
// date for a day she did not understand.
const onlyNearDaysReply = "я считаю только сегодня, завтра, послезавтра и вчера — про другие дни пока не скажу."
// dayWords — day references the calendar parser cannot resolve. A weekday name
// or a "через …" phrase means he asked about a specific other day.
var dayWords = []string{
"понедельник", "вторник", "сред", "четверг", "пятниц", "суббот", "воскресен",
"через", "monday", "tuesday", "wednesday", "thursday", "friday", "saturday", "sunday",
}
// mentionsUnknownDay reports whether the question names a day the calendar
// parser could not resolve. Mirror of mentionsUnknownPlace: it exists only to
// pick an honest reply over a confidently wrong one.
//
// Only called after ParseCalendarDate has already failed, so "завтра" and the
// other words it does know never reach here.
func mentionsUnknownDay(u string) bool {
for _, w := range dayWords {
if strings.Contains(u, w) {
return true
}
}
return false
}
// ruClock renders the clock part of the time reply: "15 часов 4 минуты".
func ruClock(t time.Time) string {
h, m := t.Hour(), t.Minute()
hourWord := ruPlural(h, "час", "часа", "часов")
if m == 0 {
return fmt.Sprintf("%d %s ровно", h, hourWord)
}
return fmt.Sprintf("%d %s %d %s", h, hourWord, m, ruPlural(m, "минута", "минуты", "минут"))
}
// dayPrefix names the day relative to now ("завтра", "вчера", …) so the date
// reply opens the way a person would say it.
func dayPrefix(now, day time.Time) string {
base := time.Date(now.Year(), now.Month(), now.Day(), 0, 0, 0, 0, now.Location())
switch int(day.Sub(base).Hours() / 24) {
case -1:
return "вчера"
case 0:
return "сегодня"
case 1:
return "завтра"
case 2:
return "послезавтра"
}
return "это"
}
func ruPlural(n int, one, two, many string) string {
n = n % 100
if n > 10 && n < 20 {
return many
}
n = n % 10
switch n {
case 1:
return one
case 2, 3, 4:
return two
default:
return many
}
}
// hasDurationWords checks whether u is asking about elapsed/remaining time
// rather than the current clock — guards replySystem from replying "сейчас
// X часов" to "сколько времени прошло". Mirrors the stage0.go build filter.
func hasDurationWords(u string) bool {
s := strings.ToLower(strings.TrimSpace(u))
// First-word duration markers (same keywords as timeQueryBuild in stage0).
first := strings.Fields(s)
if len(first) > 0 {
switch first[0] {
case "прошло", "осталось", "пройдет", "минуло", "проходит":
return true
}
}
// Broader duration keywords appearing anywhere in the utterance.
if strings.Contains(s, "прошло") || strings.Contains(s, "осталось") {
return true
}
if strings.Contains(s, " до ") {
return true
}
return false
}
// formatTime returns a human-readable Russian time string for a fact timestamp.
// Used by the query handler when answering "когда я это сделал?"-style questions.
func formatTime(t time.Time) string {
now := time.Now()
if t.After(now.Add(-2*time.Minute)) && t.Before(now.Add(2*time.Minute)) {
return "только что"
}
diff := now.Sub(t)
switch {
case diff < 10*time.Minute:
return "несколько минут назад"
case diff < 60*time.Minute:
return fmt.Sprintf("%d минут назад", int(diff.Minutes()))
case diff < 2*time.Hour:
return "час назад"
case diff < 24*time.Hour:
return fmt.Sprintf("%d часа назад", int(diff.Hours()))
default:
return t.Format("2 января 15:04")
}
}
+85
View File
@@ -0,0 +1,85 @@
// Package main — strutil.go holds small, receiver-free string utilities used
// across the voice reply paths: trimming a wake token, pulling out the first
// word or first line, and a minimal JSON string encoder for the one payload
// shape that needs it. Extend this file rather than voice.go for anything in
// that shape.
package main
import (
"fmt"
"strings"
"github.com/kami/maven/internal/router"
)
// stripWake removes a leading wake token (any script the STT phonetically
// transcribes "Maven" as) so the verb is the first word.
func stripWake(u string) string {
stripped, had := router.StripWakeToken(u)
if !had {
return strings.TrimSpace(u)
}
return stripped
}
// firstWord returns the first whitespace-delimited token (lowercased) — the
// proposed tool's name.
func firstWord(s string) string {
f := strings.Fields(s)
if len(f) == 0 {
return ""
}
return strings.ToLower(f[0])
}
// firstLine — the first non-empty line of a tool's output, for a short spoken
// reply (the full output goes to the log, not the TTS). Trimmed to keep the
// utterance sane if a command dumps a wall of text.
func firstLine(s string) string {
for _, line := range strings.Split(s, "\n") {
line = strings.TrimSpace(line)
if line != "" {
if len(line) > 200 {
line = line[:200]
}
return line
}
}
return ""
}
// jsonString — a one-line JSON string encoder without dragging encoding/json
// into the top of this file. Used to wrap a reminder payload's text field;
// the router's reminder Slots are already absolute (DateTimeParser resolved
// relative→absolute), the payload shape is conventional {"text":...}.
func jsonString(s string) string {
// minimal JSON string escape — quotes + backslash + control chars.
// adequate for the reminder payload's text field; not a general JSON
// encoder. The chroma / RAG modules (when they land) use a real json
// encoder for richer payloads. Keep it inline here so the import
// direction stays narrow.
var b []byte
b = append(b, '"')
for _, r := range s {
switch r {
case '"':
b = append(b, '\\', '"')
case '\\':
b = append(b, '\\', '\\')
case '\n':
b = append(b, '\\', 'n')
case '\r':
b = append(b, '\\', 'r')
case '\t':
b = append(b, '\\', 't')
default:
if r < 0x20 {
b = append(b, []byte(fmt.Sprintf("\\u%04x", r))...)
} else {
b = append(b, []byte(string(r))...)
}
}
}
b = append(b, '"')
return string(b)
}
+78
View File
@@ -0,0 +1,78 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/router"
)
// systemHandler — a handler with nothing but a fixed clock, which is all
// replySystem needs.
func systemHandler(now time.Time) *reactiveHandler {
return &reactiveHandler{now: func() time.Time { return now }}
}
// TestReplySystemDateOffset — "какое число завтра" must answer tomorrow's
// date, not today's (Vikunja #388).
func TestReplySystemDateOffset(t *testing.T) {
// Thursday, 30 July 2026.
now := time.Date(2026, 7, 30, 14, 5, 0, 0, time.UTC)
h := systemHandler(now)
cases := []struct{ utterance, want string }{
{"какое сегодня число", "сегодня четверг, 30 июля 2026 года"},
{"какое число", "сегодня четверг, 30 июля 2026 года"},
{"какое число завтра", "завтра пятница, 31 июля 2026 года"},
{"какое число послезавтра", "послезавтра суббота, 1 августа 2026 года"},
{"какое было число вчера", "вчера среда, 29 июля 2026 года"},
}
for _, c := range cases {
got := h.replySystem(context.Background(), router.Decision{Utterance: c.utterance})
if got != c.want {
t.Errorf("replySystem(%q) = %q, want %q", c.utterance, got, c.want)
}
}
}
// A day she cannot work out must not come back as today's date — that is the
// same silent wrong answer #388 was about, one step further out.
func TestReplySystemUnknownDayIsHonest(t *testing.T) {
now := time.Date(2026, 7, 30, 14, 5, 0, 0, time.UTC)
h := systemHandler(now)
for _, u := range []string{
"какое число в пятницу",
"какое число через неделю",
"какое число в понедельник",
} {
got := h.replySystem(context.Background(), router.Decision{Utterance: u})
if got != onlyNearDaysReply {
t.Errorf("replySystem(%q) = %q, want the honest reply", u, got)
}
}
// The days she does know must not be caught by the same guard.
if got := h.replySystem(context.Background(), router.Decision{Utterance: "какое число завтра"}); got == onlyNearDaysReply {
t.Error("завтра was treated as an unknown day")
}
}
// TestReplySystemClockCity — the clock arm must not answer local time for a
// question about another city (Vikunja #388). She keeps one clock, so every
// named place gets the honest "local time only" answer.
func TestReplySystemClockCity(t *testing.T) {
now := time.Date(2026, 7, 30, 12, 0, 0, 0, time.UTC)
h := systemHandler(now)
cases := []struct{ utterance, want string }{
{"который час", "сейчас 12 часов ровно"},
{"который час в киеве", onlyLocalTimeReply},
{"сколько времени в москве", onlyLocalTimeReply},
{"который час в лондоне", onlyLocalTimeReply},
{"который час в бишкеке", onlyLocalTimeReply},
}
for _, c := range cases {
got := h.replySystem(context.Background(), router.Decision{Utterance: c.utterance})
if got != c.want {
t.Errorf("replySystem(%q) = %q, want %q", c.utterance, got, c.want)
}
}
}
+189
View File
@@ -180,6 +180,17 @@ func (t *tickLoop) tick(ctx context.Context, now time.Time) {
// (with dedup) avoids re-queueing the same rule after a flush.
t.maybeFlush(ctx, now, state)
// gate-suppressed digest (Vikunja #281): rules the restraint gate held
// back this tick (quiet hours / away / calendar-busy), not because they
// weren't due, but because it wasn't the moment. Some of those are worth
// resurfacing later instead of just being lost — loop.DigestEligible
// draws that line. This is a SEPARATE mechanism from the in-memory
// digestQ above: that one batches candidates the gate already ALLOWED to
// fire; this one durably holds candidates the gate BLOCKED.
t.enqueueSuppressedDigest(ctx, trace, state, now)
t.expireStaleDigest(ctx, now)
t.maybeDrainDigest(ctx, state, now)
// routines: operator-declared scheduled behaviors. fire the ones whose cron
// crossed since last fire, delivered through the normal routing (voice when
// present, away channels otherwise). bodies are literal operator text — not
@@ -195,6 +206,14 @@ func (t *tickLoop) tick(ctx context.Context, now time.Time) {
// nudge time. See internal/morning for the "why not four timers" rationale.
t.fireMorningRoutines(ctx, now, state)
// pattern detection: scan every action+object pair with recorded events
// and propose a routine for any stable one not already decided (Vikunja
// #43). This used to only run as a side effect of the voice fact-write
// path, so a pattern already sitting in history went unnoticed until he
// happened to mention it again by voice. See patterns.go and
// detectPatterns below for how idempotence and dismissal are respected.
t.detectPatterns(ctx, now)
// reminders: gate-bypassing class. fired once, marked after a successful
// delivery. a failed send leaves the reminder pending — the next tick
// re-gathers and re-attempts.
@@ -342,6 +361,176 @@ func (t *tickLoop) flushDigest(ctx context.Context, now time.Time, state loop.St
t.digestQ = nil
}
// detectPatterns runs the pattern detector proactively over every
// action+object pair that has ever produced an event, independent of
// whichever fact write (or channel) last touched it (Vikunja #43). This is
// what makes pattern inference actually proactive: it fires on the daemon's
// own schedule reading accumulated history, not only as a side effect of a
// live voice turn.
//
// Idempotence and noise are handled by the store, not here — this function
// is safe to call every tick:
// - Same pattern, tick after tick: detectAndPropose's LookupProposedRoutine
// check plus proposed_routines' UNIQUE(action, object) constraint (with
// CreateProposedRoutine's ON CONFLICT DO NOTHING) mean a pair that
// already has a row — in ANY status — produces no second row and no log
// spam beyond the one line at genuine creation.
// - A DISMISSED proposal must never come back. DismissProposedRoutine flips
// status in place; the row is never deleted. So the same Lookup check
// that stops a duplicate "proposed" also stops a "dismissed" one from
// resurrecting — there is nothing tick-specific to get right here beyond
// calling the same shared path the voice route already used.
//
// This only ever creates a row for the /routines page to show. It does not
// notify, ring, or speak — Maven is "not a nag, not autonomous" (CLAUDE.md),
// and detection is not the same act as disturbing him about it. A proposal
// only starts producing nudges once he accepts it (fireAcceptedRoutines).
func (t *tickLoop) detectPatterns(ctx context.Context, now time.Time) {
pairs, err := t.store.DistinctEventPairs(ctx)
if err != nil {
log.Printf("tick: distinct event pairs: %v", err)
return
}
for _, p := range pairs {
r, _, err := detectAndPropose(ctx, t.store, p.Action, p.Object, now)
if err != nil {
log.Printf("tick: detect pattern %s/%s: %v", p.Action, p.Object, err)
continue
}
if r == nil {
continue // no stable pattern, or already proposed/accepted/dismissed
}
log.Printf("tick: proposed routine: %s/%s every %.1f days", r.Action, r.Object, r.IntervalDays)
}
}
// digestExpiry — how long a gate-suppressed care nudge stays worth
// resurfacing. 24h: these are daily-cadence rules (water/meal/break run on
// hour-scale cooldowns and re-derive from facts that reset every day), so a
// digest entry that outlives one full day is describing a day that's already
// over — "you skipped a break yesterday" said tomorrow evening is noise, not
// news. Bounding at one day also means a digest can never silently span a
// weekend of quiet hours into an unbounded backlog.
const digestExpiry = 24 * time.Hour
// maxDigestSpokenItems — the bundle read-out is capped so "batched, not
// dropped" cannot regress into "she dumps twelve things on me the moment I
// walk in" — a digest that nags in bulk is worse than the drops it replaced.
// Anything beyond the cap is still marked drained (it did get its moment;
// the cap limits WORDS, not whether it counted) and folded into a trailing
// count instead of being spoken in full.
const maxDigestSpokenItems = 3
// enqueueSuppressedDigest scans this tick's trace for care candidates the
// gate blocked for a genuine restraint reason and durably records the
// digest-eligible ones (loop.DigestEligible). Phrasing happens once, here,
// at enqueue time — not re-derived at drain time — the same way queueNudge
// phrases once and caches, so a rule suppressed for hours isn't re-prompting
// the LLM every tick it stays blocked (EnqueueDigestEntry's rule+body dedupe
// makes repeat calls here harmless, but skipping the phrase call entirely
// when a pending entry already exists avoids the LLM round-trip too).
func (t *tickLoop) enqueueSuppressedDigest(ctx context.Context, trace *loop.TickTrace, state loop.State, now time.Time) {
if trace == nil {
return
}
for _, tr := range trace.RuleTraces {
if !tr.PredicateResult || tr.GateResult {
continue // didn't want to fire, or wasn't suppressed
}
if !loop.DigestEligible(tr.Severity, tr.GateBlockedBy) {
continue
}
rule := loop.Rule{Name: tr.RuleName, Severity: tr.Severity}
cand := loop.Candidate{Rule: rule, Severity: tr.Severity, State: state}
pn, err := t.phraser.PhraseNudge(ctx, cand)
if err != nil {
log.Printf("tick: phrase digest candidate %s: %v", tr.RuleName, err)
continue
}
expires := now.Add(digestExpiry)
if _, deduped, err := t.store.EnqueueDigestEntry(ctx, tr.RuleName, int(tr.Severity), pn.Body, now, expires); err != nil {
log.Printf("tick: enqueue digest entry %s: %v", tr.RuleName, err)
} else if deduped {
// same suppressed nudge already pending — nothing new to say.
continue
}
}
}
// expireStaleDigest sweeps entries past their expiry once per tick — cheap
// bookkeeping, mirrors ReconcileStaleDeliveryAttempts's shape.
func (t *tickLoop) expireStaleDigest(ctx context.Context, now time.Time) {
n, err := t.store.ExpireStaleDigestEntries(ctx, now)
if err != nil {
log.Printf("tick: expire stale digest entries: %v", err)
return
}
if n > 0 {
log.Printf("tick: expired %d stale digest entr(y/ies) unspoken", n)
}
}
// maybeDrainDigest speaks the pending digest bundle once the gate's
// suppression reasons have actually cleared — quiet hours over, back from
// away, out of the meeting. Draining while still suppressed would just be a
// second way to nag through quiet hours; the bundle waits for the same "is
// it allowed right now" condition a live nudge already waits for.
func (t *tickLoop) maybeDrainDigest(ctx context.Context, state loop.State, now time.Time) {
if state.QuietHours || state.CalendarBusy || state.Presence == store.Away {
return
}
entries, err := t.store.PendingDigestEntries(ctx, now)
if err != nil {
log.Printf("tick: pending digest entries: %v", err)
return
}
if len(entries) == 0 {
return
}
spoken := entries
extra := 0
if len(spoken) > maxDigestSpokenItems {
spoken = entries[:maxDigestSpokenItems]
extra = len(entries) - maxDigestSpokenItems
}
var b strings.Builder
maxSev := 0
for i, e := range spoken {
if i > 0 {
b.WriteString(" · ")
}
b.WriteString(e.Body)
if e.Severity > maxSev {
maxSev = e.Severity
}
}
if extra > 0 {
fmt.Fprintf(&b, " · и ещё %d", extra)
}
body := b.String()
summary := fmt.Sprintf("%d отложенных уведомлений", len(entries))
cand := loop.Candidate{
Rule: loop.Rule{Name: "digest", Severity: loop.Severity(maxSev)},
Severity: loop.Severity(maxSev),
State: state,
}
pn := delivery.PhrasedNudge{Candidate: cand, Body: body, Summary: summary}
t.cachePhrase(pn)
if _, err := t.dispatcher.DispatchNudge(ctx, pn, now); err != nil {
log.Printf("tick: dispatch digest bundle: %v", err)
return // leave entries pending; retried next tick
}
ids := make([]int64, len(entries))
for i, e := range entries {
ids[i] = e.ID
}
if err := t.store.DrainDigestEntries(ctx, ids, now); err != nil {
log.Printf("tick: drain digest entries: %v", err)
}
}
// routinesFromConfig maps the config's routine blocks to the engine type.
// Validation (cron parses, name/body present, severity defaulted) already ran
// in config.Load, so this is a pure field copy.
+179 -1373
View File
File diff suppressed because it is too large Load Diff
+434
View File
@@ -0,0 +1,434 @@
package main
import (
"bufio"
"context"
"fmt"
"log"
"os"
"path/filepath"
"strings"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/delivery/voicesink"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/stt"
"github.com/kami/maven/internal/tool"
"github.com/kami/maven/internal/tts"
"github.com/kami/maven/internal/voice"
"github.com/kami/maven/internal/weather"
"github.com/kami/maven/internal/worker"
)
// voiceWiring — everything the daemon needs to run the audio path. Held by
// cmd/mavend/main.go alongside the other wirings; closed on shutdown.
type voiceWiring struct {
server *voice.Server
sessions *voice.Sessions
voiceSink delivery.Sink
embedder router.Embedder
handler *reactiveHandler // the reactive handler for IPC Chat
// worker clients (set when configured as Remote): closed on shutdown so
// mavsttd / mavttsd don't keep a stale conn into a restarting daemon.
sttClient *worker.Client
ttsClient *worker.Client
}
// close releases the listener + worker conns. Safe to call on nil (when
// voice is not wired — wireVoice returns nil,nil).
func (w *voiceWiring) close() {
if w == nil {
return
}
if w.embedder != nil {
_ = w.embedder.Close()
}
if w.server != nil {
_ = w.server.Close()
}
if w.sttClient != nil {
_ = w.sttClient.Close()
}
if w.ttsClient != nil {
_ = w.ttsClient.Close()
}
}
// wireVoice builds the audio path from cfg + a CoreAPI + a router. Returns
// nil wiring + nil error when voice isn't enabled (the caller's voice sink
// stays nil; the dispatcher's ChannelVoice routing drops silently).
//
// When voice is enabled, MUST wire a voicesink into the dispatcher's Voice
// slot using w.sessions (the caller does that — see main.go).
func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, memStore memory.Store, dataStore *store.Store, eco *ecosystemWiring) (*voiceWiring, error) {
if cfg.Voice == nil || !cfg.Voice.Enabled {
return nil, nil
}
w := &voiceWiring{}
// ----- stt (Stub in-process OR Remote via worker socket) -----
var transcriber stt.Transcriber
if cfg.Voice.Stt != nil && cfg.Voice.Stt.Socket != "" {
c := worker.Dial(cfg.Voice.Stt.Socket)
w.sttClient = c
lang := cfg.Voice.Stt.Lang
if lang == "" {
lang = cfg.Voice.Lang
}
transcriber = stt.NewRemote(c, lang)
} else {
transcriber = stt.NewStub()
}
// ----- tts (Stub in-process OR Remote) -----
var synthesizer tts.Synthesizer
if cfg.Voice.Tts != nil && cfg.Voice.Tts.Socket != "" {
c := worker.Dial(cfg.Voice.Tts.Socket)
w.ttsClient = c
lang := cfg.Voice.Tts.Lang
if lang == "" {
lang = cfg.Voice.Lang
}
synthesizer = tts.NewRemote(c, lang, cfg.Voice.Tts.Voice)
} else {
synthesizer = tts.NewStub()
}
// ----- router: embedder (ONNX when configured, floor HashEmbedder otherwise) -----
var emb router.Embedder
if cfg.Voice.Embedder != nil {
onnx, err := router.NewONNXEmbedder(
cfg.Voice.Embedder.ModelPath,
cfg.Voice.Embedder.TokenizerPath,
cfg.Voice.Embedder.LibPath,
)
if err != nil {
w.close()
return nil, fmt.Errorf("embedder: %w", err)
}
log.Printf("voice: onnx embedder loaded (%d dim)", onnx.Dim())
emb = onnx
} else {
log.Printf("voice: embedder not configured, using HashEmbedder floor")
emb = router.NewHashEmbedder(1024)
}
w.embedder = emb
checkStoredEmbedder(dataStore, emb)
// ----- tool executor (the enabled act allowlist, store-backed) -----
// Config tools are the declarative bootstrap: seed them into the store as
// enabled (editing mavend.json IS the human enable act). Ad-hoc tools are
// enabled later through the authed mavweb surface. The executor + matcher
// both read the store live, so a newly-enabled tool is runnable without a
// daemon restart.
seedTools(coreAPI, cfg.Voice.Tools)
exec := tool.NewExecutor(coreAPI, time.Duration(cfg.Voice.ToolTimeout))
matcher := tool.NewMatcher(coreAPI)
// ----- weather provider (Open-Meteo when configured, Stub otherwise) -----
var weatherProvider weather.Provider
var weatherLocation string
if cfg.Voice.Weather != nil && cfg.Voice.Weather.Provider == "open-meteo" {
weatherProvider = weather.NewOpenMeteoProvider()
weatherLocation = cfg.Voice.Weather.DefaultLocation
log.Printf("voice: weather provider: open-meteo (default location: %s)", cfg.Voice.Weather.DefaultLocation)
} else {
weatherProvider = weather.NewStubProvider()
log.Printf("voice: weather provider: stub (not configured)")
}
// The replier uses the same llama-server as the phraser.
var llmClient *llm.Client
if lp, ok := phr.(*phraser.LLMPhraser); ok {
llmClient = llm.New(lp.BaseURL(), 60*time.Second)
}
// ----- router (the cascade; floor examples seed the classifier) -----
// The act matcher's allowlist is exactly the enabled tool names — the
// router only matches acts the executor can run (one source of truth).
threshold := cfg.Voice.RouterThreshold
if threshold <= 0 {
threshold = config.DefaultRouterThreshold
}
// The resident model routes by default: 63.2% of held-out intents right
// against the classifier's 50.0%, at about 1s a turn instead of 30ms (see
// config.VoiceConfig.LLMRouter). The classifier always stays wired as the
// fallback, so a model error never breaks a turn.
rtr := buildRouter(emb, matcher, threshold, pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient))
// ----- sessions registry (shared with voicesink) -----
sessions := voice.NewSessions()
w.sessions = sessions
// ----- voice sink (proactive nudges: dispatcher → voicesink → tts → push to client) -----
w.voiceSink = voicesink.New(synthesizer, sessions)
// ----- memory (long-term vector storage) -----
// Persistent (store-backed, survives restarts) when the daemon passes one;
// falls back to the in-memory floor otherwise (tests / no-store paths).
if memStore == nil {
memStore = memory.NewInMemoryStore()
}
// ----- dialogue (multi-turn slot carry-over; 2-min follow-up window) -----
// Store-backed when the daemon passes a store, so a restart mid-conversation
// keeps the thread (Vikunja #363). Sessions past their TTL are dropped on
// load, never revived. Clarify's parked question stays in memory only.
var dialogueSessions *dialogue.SessionStore
if dataStore != nil {
dialogueSessions = dialogue.NewPersistentSessionStore(2*time.Minute, dataStore)
if err := dialogueSessions.Load(context.Background(), time.Now()); err != nil {
log.Printf("dialogue: load saved sessions: %v", err)
}
} else {
dialogueSessions = dialogue.NewSessionStore(2 * time.Minute)
}
clarifyStore := dialogue.NewClarifyStore(clarifyTTL)
timeParser := router.NewPythonDateParser()
// ----- replier (LLM-backed when the engine is on, Stub floor otherwise) -----
replier := voice.Replier(voice.NewStubReplier())
if llmClient != nil {
replier = newLLMReplier(llmClient, contextBlockFn(cfg, time.Now))
}
// ----- the handler (the reactive path; closes over stt / tts / router / coreAPI / memory) -----
h := &reactiveHandler{
stt: transcriber,
tts: synthesizer,
router: rtr,
embedder: emb,
api: coreAPI,
tools: exec,
matcher: matcher,
replier: replier,
phraser: phr,
now: time.Now,
weatherProvider: weatherProvider,
weatherLocation: weatherLocation,
memStore: memStore,
dataStore: dataStore,
dialogueSessions: dialogueSessions,
clarifyStore: clarifyStore,
// 0 here (unset config) ⇒ the dialogue default.
clarifyMaxAttempts: cfg.Voice.ClarifyMaxAttempts,
extractor: router.Extractor{Time: timeParser, Acts: matcher, Facts: router.DefaultFactParser{}},
queryMinScore: cfg.Voice.QueryMinScore,
queryMinMargin: cfg.Voice.QueryMinMargin,
timeParser: timeParser,
ecosystem: eco,
}
// ----- the server (TCP listener) -----
srv := voice.NewServer(cfg.Voice.Bind, h, sessions)
if err := srv.Listen(); err != nil {
w.close()
return nil, fmt.Errorf("voice listen: %w", err)
}
w.server = srv
w.handler = h
return w, nil
}
// pickLLMRouter returns the LLM router when the operator asked for it and there
// is a llama-server to talk to, and nil otherwise. nil is safe: the cascade then
// routes with the classifier, so an unusable setting costs accuracy, not turns.
func pickLLMRouter(enabled bool, c *llm.Client) *router.LLMRouter {
if !enabled {
return nil
}
if c == nil {
log.Printf("voice: voice.llm_router is on but there is no llama-server to route with (the phraser is not an LLM phraser) — using the classifier instead")
return nil
}
log.Printf("voice: LLM router enabled")
return router.NewLLMRouter(c)
}
// buildRouter constructs the reactive-path router with the given embedder
// and confidence threshold.
// - stage-0 grammars from DefaultActMatcher whose fn allowlist is exactly
// the enabled tool names (actFns) — the router only matches acts the
// executor can run. Empty ⇒ every act refuses at the matcher.
// - The embedder is provided by wireVoice: HashEmbedder (floor) when no
// embedder config is present, or the ONNX multilingual model when
// configured — same interface, one constructor change.
// - 6 bootstrap examples covering the 5 intents + one compound-capture
// placeholder. Spec calls for ~10 per intent at production; this is the
// bootstrapping floor swapped by tuning the seed set later.
// - Threshold is from voice.router_threshold config (default 0.55).
func buildRouter(emb router.Embedder, acts router.ActMatcher, threshold float64, llmR *router.LLMRouter) *router.Router {
cls := router.NewClassifier(emb)
seedClassifier(cls)
grammars := router.DefaultGrammars(acts)
grammars = append(grammars, router.SystemTimeDateGrammars()...)
grammars = append(grammars, router.ReminderGrammar())
return router.New(router.Config{
Grammars: grammars,
Classifier: cls,
Extractor: router.Extractor{
Time: router.NewPythonDateParser(),
Acts: acts,
Facts: router.DefaultFactParser{},
},
Threshold: threshold,
LLM: llmR,
})
}
// seedDir is the directory containing intent seed files. Each file is named
// <intent>.txt and contains one training example per line (blank lines and
// lines starting with # are ignored). Relative to the working directory.
const seedDir = "models/seeds"
// seedClassifier floors the embedded examples so the cold-boot path
// doesn't return ErrNoIntents. Loads examples from seedDir — one file per
// intent (act.txt, reminder.txt, fact.txt, note.txt, query.txt). When the
// classifier can't decide it falls through to Clarify — the last-resort
// path asks the user to rephrase rather than guessing wrong.
func seedClassifier(c *router.Classifier) {
intents := []router.Intent{
router.IntentAct,
router.IntentReminder,
router.IntentFact,
router.IntentNote,
router.IntentQuery,
router.IntentChat,
router.IntentSystem,
}
total := 0
for _, intent := range intents {
n, err := loadSeedFile(c, intent)
if err != nil {
log.Printf("voice: seed %s: %v", intent, err)
continue
}
total += n
}
log.Printf("voice: loaded %d seed examples from %s", total, seedDir)
}
func loadSeedFile(c *router.Classifier, intent router.Intent) (int, error) {
path := filepath.Join(seedDir, string(intent)+".txt")
f, err := os.Open(path)
if err != nil {
return 0, fmt.Errorf("open %s: %w", path, err)
}
defer f.Close()
var count int
sc := bufio.NewScanner(f)
for sc.Scan() {
line := strings.TrimSpace(sc.Text())
if line == "" || strings.HasPrefix(line, "#") {
continue
}
if err := c.AddExample(context.Background(), intent, line); err != nil {
log.Printf("voice: seed %s: skipping %q: %v", intent, line, err)
continue
}
count++
}
if err := sc.Err(); err != nil {
return count, fmt.Errorf("scan %s: %w", path, err)
}
return count, nil
}
// seedTools upserts the config-declared tools into the store as enabled. Editing
// mavend.json is a human act, so a config tool is enabled by definition; this
// makes the declarative config the reproducible bootstrap while the store stays
// the single runtime source of truth (mavweb enables ad-hoc ones on top).
func seedTools(api ipc.CoreAPI, tools []config.ToolConfig) {
ctx := context.Background()
now := time.Now()
n := 0
for _, tc := range tools {
if tc.Name == "" || len(tc.Cmd) == 0 {
log.Printf("voice: skipping malformed tool config %+v", tc)
continue
}
if err := api.EnableTool(ctx, tc.Name, tc.Cmd, tc.Destructive, tc.Scope, now); err != nil {
log.Printf("voice: seed tool %q: %v", tc.Name, err)
continue
}
n++
}
log.Printf("voice: seeded %d act tools from config", n)
}
// reembedOnStart is the -reembed flag (set in run()). Opt-in on purpose: see
// runReembed.
var reembedOnStart bool
// checkStoredEmbedder compares the embedder we just loaded with the one that
// wrote the vectors already in the DB (Vikunja #378).
//
// The two models we have both make 384-dim vectors, so a size check catches
// nothing: after a swap, recall silently compares vectors from different
// spaces and the scores are noise. So we say it out loud. Recall itself is not
// changed here — the fix is `mavend -reembed`.
func checkStoredEmbedder(dataStore *store.Store, emb router.Embedder) {
if dataStore == nil {
return
}
current := router.EmbedderID(emb)
if reembedOnStart {
runReembed(dataStore, emb, current)
return
}
stored, mismatch, err := dataStore.CheckEmbedder(context.Background(), current)
if err != nil {
log.Printf("voice: embedder marker check failed: %v", err)
return
}
if mismatch {
log.Printf("voice: WARNING embedder MISMATCH — stored vectors were written by %q but the configured embedder is %q; recall scores are noise until the notes and facts are re-embedded — run `mavend -reembed` once (Vikunja #378)", stored, current)
return
}
log.Printf("voice: embedder marker ok (%s)", current)
}
// runReembed is the one-shot backfill behind -reembed.
//
// Why a flag and not automatic on mismatch: the embedder is ONNX on the
// laptop's CPU, so a few thousand notes is minutes of work. Doing that silently
// inside a normal start would look like the daemon hanging on boot. So the user
// runs it once, deliberately, after an embedder swap; the mismatch warning
// above tells them to. It re-embeds, logs what it did, and then the daemon
// carries on serving as usual — no separate binary, no second start needed.
func runReembed(dataStore *store.Store, emb router.Embedder, current string) {
log.Printf("voice: re-embedding stored notes and facts with %s — this can take a few minutes, do not interrupt", current)
res, err := dataStore.ReembedAll(context.Background(), current,
// EmbedPassage, not EmbedQuery: these are stored texts being searched
// FOR, which is the side they were written with.
func(ctx context.Context, text string) ([]float32, error) {
return router.EmbedPassage(ctx, emb, text)
})
if err != nil {
log.Printf("voice: re-embed FAILED, nothing was changed and no marker was written — safe to run again: %v", err)
return
}
if res.Skipped {
log.Printf("voice: re-embed skipped — the stored vectors were already written by %s", current)
return
}
log.Printf("voice: re-embed done — %d notes in the notes table, %d notes and %d facts in the memory index, took %s; stored vectors now belong to %s",
res.Notes, res.MemNotes, res.Facts, res.Took.Round(time.Second), current)
// A row with no text cannot be re-embedded, so its vector is still the old
// model's noise while the marker now says everything is current. Both write
// paths always store the text, so this should be zero — say it loudly
// rather than bury it in the line above if it ever isn't.
if res.NoText > 0 {
log.Printf("voice: WARNING %d stored rows had no text, so their vectors could not be re-embedded and are still noise; they will never match anything useful (Vikunja #378)", res.NoText)
}
}
+50
View File
@@ -0,0 +1,50 @@
// Package main — weatherq.go holds the weather-query keyword helpers: does
// this utterance ask about weather at all, and which city (if any) did it
// name. Both are plain substring/lookup matching, not NLU — extend this file
// rather than voice.go for anything in that shape.
package main
import "strings"
// isWeatherQuery returns true if the utterance is about weather.
func isWeatherQuery(u string) bool {
lower := strings.ToLower(u)
return strings.Contains(lower, "погод") ||
strings.Contains(lower, "градус") ||
strings.Contains(lower, "температур") ||
strings.Contains(lower, "дожд") ||
strings.Contains(lower, "холод") ||
strings.Contains(lower, "тепл") ||
strings.Contains(lower, "weather") ||
strings.Contains(lower, "temperature")
}
// extractWeatherLocation parses a location from the utterance, or falls back
// to the configured default. Very basic: just checks for known city names.
func extractWeatherLocation(u, defaultLoc string) string {
lower := strings.ToLower(u)
cities := map[string]string{
"москв": "Moscow",
"moscow": "Moscow",
"питер": "Saint Petersburg",
"spb": "Saint Petersburg",
"петербур": "Saint Petersburg",
"лондон": "London",
"london": "London",
"париж": "Paris",
"paris": "Paris",
"берлин": "Berlin",
"berlin": "Berlin",
"нью-йорк": "New York",
"new york": "New York",
}
for substr, name := range cities {
if strings.Contains(lower, substr) {
return name
}
}
if defaultLoc != "" {
return defaultLoc
}
return "Moscow"
}
+4 -3
View File
@@ -17,11 +17,12 @@ import (
)
// fakeCore records the mutating calls handleTools makes and returns canned
// tool lists / errors. Embedding ipc.CoreAPI (nil) satisfies the large
// tool lists / errors. Embedding ipc.UnimplementedCoreAPI satisfies the large
// interface — only the methods the handlers touch are overridden; any other
// call would nil-panic, which is fine since the handlers never make them.
// call returns ipc.ErrNotImplemented instead of nil-panicking, so a test that
// accidentally exercises an undeclared method fails loudly.
type fakeCore struct {
ipc.CoreAPI
ipc.UnimplementedCoreAPI
proposed, enabled []ipc.Tool
listErr error
+1 -1
View File
@@ -314,7 +314,7 @@ func noCache(h http.Handler) http.Handler {
}
func main() {
addr := flag.String("addr", ":9200", "HTTP listen address")
addr := flag.String("addr", "127.0.0.1:9200", "HTTP listen address (loopback by default; pass e.g. \":9200\" or a LAN IP deliberately for wider exposure — POST /chat and /routines are state-changing)")
voiceAddr := flag.String("voice", "127.0.0.1:9100", "voice server TCP addr (host:port)")
// ntfyWS: the ntfy WebSocket subscribe URL the PWA connects to for in-app
// nudge delivery, e.g. wss://ntfy.kvmx.ru/maven/ws?auth=<base64-token>. The
+28 -3
View File
@@ -4,10 +4,23 @@
#
# NOTE: hexis.<domain> previously pointed at the MCP tool — repoint that
# elsewhere first (the app now owns hexis.*).
#
# 10.42.0.1 and 192.168.1.104 below are THIS BOX's WireGuard and LAN
# addresses (homesrv) — these admin UIs have no auth of their own, so the
# explicit bind + allow/deny below is what keeps them off the open internet.
# On a different box, replace both addresses with that box's wg and LAN IPs.
# Do NOT "fix" a failed bind by reverting to `listen 80` (all interfaces) —
# that removes the only access control these containers have.
server {
listen 80;
listen 10.42.0.1:80;
listen 192.168.1.104:80;
server_name nexus.kvmx.ru;
allow 10.42.0.0/24;
allow 192.168.1.0/24;
deny all;
location / {
proxy_pass http://127.0.0.1:9740;
proxy_set_header Host $host;
@@ -18,8 +31,14 @@ server {
}
server {
listen 80;
listen 10.42.0.1:80;
listen 192.168.1.104:80;
server_name praxis.kvmx.ru;
allow 10.42.0.0/24;
allow 192.168.1.0/24;
deny all;
location / {
proxy_pass http://127.0.0.1:8989;
proxy_set_header Host $host;
@@ -30,8 +49,14 @@ server {
}
server {
listen 80;
listen 10.42.0.1:80;
listen 192.168.1.104:80;
server_name hexis.kvmx.ru;
allow 10.42.0.0/24;
allow 192.168.1.0/24;
deny all;
location / {
proxy_pass http://127.0.0.1:9741;
proxy_set_header Host $host;
+4 -3
View File
@@ -6,11 +6,12 @@
"state_dir": "/var/lib/maven",
"phraser": {
"model_path": "/opt/maven/models/llm/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf",
"model_path": "/opt/maven/models/llm/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf",
"bin_path": "llama-server",
"n_gpu_layers": 99,
"n_ctx": 2048,
"timeout": "60s"
"n_ctx": 4096,
"timeout": "60s",
"llm_nudges": false
},
"telegram": {
+5 -84
View File
@@ -7,7 +7,6 @@ import (
"path/filepath"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
)
@@ -359,8 +358,12 @@ func TestGate_IpcServer_ChatAllowedForEnrolledCaller(t *testing.T) {
// recordingAPI — a no-op CoreAPI that counts WriteFact invocations; the auth
// check must reject before reaching it, otherwise the refusal leaks into the
// fake's counts and we fail.
// fake's counts and we fail. Embeds ipc.UnimplementedCoreAPI so every method
// this test doesn't exercise returns ipc.ErrNotImplemented loudly instead of
// being hand-stubbed to a canned value nobody checks.
type recordingAPI struct {
ipc.UnimplementedCoreAPI
writes int
chats int
}
@@ -369,89 +372,7 @@ func (r *recordingAPI) WriteFact(_ context.Context, _ ipc.WriteFactReq) (int64,
r.writes++
return int64(r.writes), nil
}
func (r *recordingAPI) LatestFact(_ context.Context, _ string) (ipc.Fact, error) {
return ipc.Fact{}, ipc.ErrNoFact
}
func (r *recordingAPI) LatestFactBySource(_ context.Context, _, _ string) (ipc.Fact, error) {
return ipc.Fact{}, ipc.ErrNoFact
}
func (r *recordingAPI) Since(_ context.Context, _ string, _ time.Time) (time.Duration, error) {
return 0, ipc.ErrNoFact
}
func (r *recordingAPI) Presence(_ context.Context) (ipc.Presence, error) {
return ipc.Presence{}, nil
}
func (r *recordingAPI) CreateReminder(_ context.Context, _ time.Time, _, _ string) (int64, error) {
return 1, nil
}
func (r *recordingAPI) MarkReminder(_ context.Context, _ int64, _ string) error { return nil }
func (r *recordingAPI) ListReminders(_ context.Context, _ int) ([]ipc.Reminder, error) {
return nil, nil
}
func (r *recordingAPI) TickTrace(_ context.Context) (ipc.TickTrace, error) {
return ipc.TickTrace{}, nil
}
func (r *recordingAPI) MorningStatus(_ context.Context) ([]ipc.MorningRoutineStatus, error) {
return nil, nil
}
func (r *recordingAPI) RecordNudge(_ context.Context, _, _, _ string, _ time.Time) (int64, error) {
return 1, nil
}
func (r *recordingAPI) ResolveNudge(_ context.Context, _ int64, _ string, _ time.Time) error {
return nil
}
func (r *recordingAPI) RecentOutcomes(_ context.Context, _ string, _ int) ([]string, error) {
return nil, nil
}
func (r *recordingAPI) RecentFacts(_ context.Context, _ int) ([]ipc.Fact, error) {
return nil, nil
}
func (r *recordingAPI) CalendarEvents(_ context.Context, _, _ time.Time) ([]ipc.Fact, error) {
return nil, nil
}
func (r *recordingAPI) RecentNudges(_ context.Context, _ int) ([]ipc.Nudge, error) {
return nil, nil
}
func (r *recordingAPI) WriteNote(_ context.Context, _ time.Time, _ string, _ []float32, _ string) (int64, error) {
return 1, nil
}
func (r *recordingAPI) QueryNotes(_ context.Context, _ []float32, _ int) ([]ipc.Note, error) {
return nil, nil
}
func (r *recordingAPI) RecentNotes(_ context.Context, _ int) ([]ipc.Note, error) {
return nil, nil
}
func (r *recordingAPI) ProposeTool(_ context.Context, _, _, _ string, _ time.Time) (bool, error) {
return false, nil
}
func (r *recordingAPI) EnableTool(_ context.Context, _ string, _ []string, _ bool, _ string, _ time.Time) error {
return nil
}
func (r *recordingAPI) DisableTool(_ context.Context, _ string) error {
return nil
}
func (r *recordingAPI) DeleteTool(_ context.Context, _ string) error {
return nil
}
func (r *recordingAPI) LookupTool(_ context.Context, _ string) (ipc.Tool, error) {
return ipc.Tool{}, ipc.ErrToolNotFound
}
func (r *recordingAPI) ListTools(_ context.Context, _ string) ([]ipc.Tool, error) {
return nil, nil
}
func (r *recordingAPI) RevertFact(_ context.Context, _ string) (int64, error) {
return 0, nil
}
func (r *recordingAPI) ListProposedRoutines(_ context.Context) ([]ipc.ProposedRoutine, error) {
return nil, nil
}
func (r *recordingAPI) AcceptProposedRoutine(_ context.Context, _ int64) error {
return nil
}
func (r *recordingAPI) DismissProposedRoutine(_ context.Context, _ int64) error {
return nil
}
func (r *recordingAPI) Chat(_ context.Context, text string) (string, error) {
r.chats++
return "echo: " + text, nil
+13
View File
@@ -302,6 +302,13 @@ type VoiceConfig struct {
// Russian self-reference). Example: "Be formal and answer in English only."
Persona string `json:"persona,omitempty"`
// OwnerName / City — optional facts about the owner, added to the shared
// context block (internal/persona). Empty is fine: the block still states
// who he is grammatically (a man, addressed as "ты") and the current time.
// Nothing about correct behaviour may depend on these being filled in.
OwnerName string `json:"owner_name,omitempty"`
City string `json:"city,omitempty"`
// Weather — the weather provider config. nil ⇒ the daemon wires
// the stub provider (returns ErrNotConfigured — "погода не настроена").
// Set provider to "open-meteo" to use the keyless Open-Meteo API.
@@ -362,6 +369,12 @@ type PhraserConfig struct {
NGpuLayers int `json:"n_gpu_layers,omitempty"`
NCtx int `json:"n_ctx,omitempty"`
Timeout Duration `json:"timeout,omitempty"`
// LLMNudges — let the model word nudges again. Off by default: nudges are
// worded from hand-written Russian templates now (the model broke the
// persona and invented units). Chat, query and reminder phrasing always go
// through the model regardless. See phraser.Config.LLMNudges.
LLMNudges bool `json:"llm_nudges,omitempty"`
}
// EmbedderConfig — paths for the ONNX multilingual embedder. The daemon
+21
View File
@@ -35,6 +35,27 @@ func TestLoadDefaults(t *testing.T) {
}
}
// Nudges come from templates unless the config says otherwise.
func TestPhraserLLMNudgesDefaultsOff(t *testing.T) {
p := writeConfig(t, `{"phraser":{"model_path":"/tmp/m.gguf"}}`)
c, err := Load(p)
if err != nil {
t.Fatalf("Load: %v", err)
}
if c.Phraser.LLMNudges {
t.Error("llm_nudges defaults on; templates must be the default")
}
p = writeConfig(t, `{"phraser":{"model_path":"/tmp/m.gguf","llm_nudges":true}}`)
c, err = Load(p)
if err != nil {
t.Fatalf("Load: %v", err)
}
if !c.Phraser.LLMNudges {
t.Error("llm_nudges:true did not parse")
}
}
func TestLoadDurationsParse(t *testing.T) {
p := writeConfig(t, `{"tick_interval":"90s","repeat_interval":"10m"}`)
c, err := Load(p)
+21 -3
View File
@@ -137,9 +137,10 @@ func NewDispatcher(cfg Config) *Dispatcher {
// picks for (severity, presence), sends via the matching sink, and records
// one nudge row per successful send. returns the dispatches (one per channel).
//
// a Drop channel = no send, no record (the nudge was suppressed by routing,
// not by a failure — "a missed water nudge is noise"). a nil sink = channel
// not wired, skip silently. a send error stops the dispatch and returns what
// a Drop channel = no send (the nudge was suppressed by routing, not by a
// failure — "a missed water nudge is noise"), but it does leave a 'dropped'
// outbox row so the suppression is visible. a nil sink = channel not wired,
// skip silently. a send error stops the dispatch and returns what
// got through — the daemon decides whether to retry.
func (d *Dispatcher) DispatchNudge(ctx context.Context, pn PhrasedNudge, now time.Time) ([]Dispatch, error) {
c := pn.Candidate
@@ -148,6 +149,16 @@ func (d *Dispatcher) DispatchNudge(ctx context.Context, pn PhrasedNudge, now tim
for i := 0; i < len(channels); i++ {
ch := channels[i]
if ch == ChannelDrop {
// the routing table suppressed this nudge on purpose (a care nudge
// while you're away is noise). that stays — but it must not be
// invisible, or "she dropped it" and "the rule never fired" look
// the same afterwards. no nudges row: that table feeds the
// ignored_rate signal, and a nudge nobody could see must not
// count as ignored.
id := d.beginOutbox(ctx, "nudge", c.Rule.Name, 0, ch, pn.Summary, now)
d.completeOutbox(ctx, id, store.DeliveryDropped, now)
log.Printf("dispatcher: dropped %s (sev%d, presence=%s) — routing table suppressed it",
c.Rule.Name, c.Severity, c.State.Presence)
continue
}
s := Sendable{
@@ -396,6 +407,13 @@ func messageForChannel(s Sendable) string {
if !isAway(s.Channel) {
return s.Body
}
return AwayMessage(s)
}
// AwayMessage — the only text an off-box channel may ever carry. Exported so
// the away sinks share this one rule instead of each inventing a fallback: the
// summary if we have one, otherwise a fixed generic line. Never the body.
func AwayMessage(s Sendable) string {
if s.Summary != "" {
return s.Summary
}
+6 -4
View File
@@ -37,10 +37,12 @@ func TestVoiceNoSessionFallthroughLeavesOutboxTrail(t *testing.T) {
[]string{"voice", "ntfy"}, []string{store.DeliveryFailed, store.DeliverySent}},
{"sev4 falls through to telegram", loop.Sev4,
[]string{"voice", "telegram"}, []string{store.DeliveryFailed, store.DeliverySent}},
{"sev1 does not fall through", loop.Sev1,
[]string{"voice"}, []string{store.DeliveryFailed}},
{"sev2 does not fall through", loop.Sev2,
[]string{"voice"}, []string{store.DeliveryFailed}},
// care severities still don't reach an away channel; since #370 the
// drop itself is a visible row instead of nothing.
{"sev1 drops instead of falling through", loop.Sev1,
[]string{"voice", "drop"}, []string{store.DeliveryFailed, store.DeliveryDropped}},
{"sev2 drops instead of falling through", loop.Sev2,
[]string{"voice", "drop"}, []string{store.DeliveryFailed, store.DeliveryDropped}},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
+9 -13
View File
@@ -2,10 +2,10 @@
//
// ntfy is the away-channel for sev3 (ops soft) nudges, sev4 (ops hard)
// nudges when present (alongside voice), and reminders when away. the
// message body is the Sendable's Summary — the minimal-body rule from the
// message body is delivery.AwayMessage — the minimal-body rule from the
// spec ("disk low on homesrv," not detail; no shoulder-surf exfil through
// the relay). voice gets Body; away channels get Summary, enforced at the
// sink so a phraser bug can't exfil.
// the relay). the dispatcher already strips detail off away sendables; the
// sink uses the same helper so it can't leak the body on its own either.
//
// ntfy runs locally (docker, 127.0.0.1:8085, deny-all auth). maven publishes
// with a dedicated user (write-only to maven-* topics) — the credential is a
@@ -69,18 +69,14 @@ func New(cfg Config) (*Sink, error) {
}, nil
}
// Send publishes one notification to ntfy. the body is the Sendable's Summary
// (minimal body); Title is "maven" (consistent sender identity on the lock
// screen — the content is in the body). Priority maps from severity/kind so
// Send publishes one notification to ntfy. the body is the minimal away
// message (never the full body); Title is "maven" (consistent sender identity
// on the lock screen — the content is in the body). Priority maps from severity/kind so
// the phone client can ring differently for an alarm vs a soft ops nudge.
func (s *Sink) Send(ctx context.Context, d delivery.Sendable) error {
body := d.Summary
if body == "" {
body = d.Body // terse full message beats no message
}
if body == "" {
return fmt.Errorf("ntfysink: empty message for %s", d.Channel)
}
// never fall back to d.Body: ntfy leaves the box, so an empty summary gets
// a generic line instead of the full detail.
body := delivery.AwayMessage(d)
req, err := http.NewRequestWithContext(ctx, http.MethodPost, s.topicURL(), strings.NewReader(body))
if err != nil {
+16 -9
View File
@@ -147,9 +147,9 @@ func TestSendBodyIsSummaryNotFullBody(t *testing.T) {
}
}
func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
// a terse full message is better than no message; the phraser should
// produce a summary for away-bound severities, but don't silently drop.
func TestSendNeverSendsTheBodyWhenSummaryEmpty(t *testing.T) {
// #368: this used to fall back to the full body. ntfy leaves the box, so
// an empty summary gets a fixed generic line plus the rule name instead.
rs := newRecordingServer(t, 200, "")
srv := httptest.NewServer(rs.handler())
defer srv.Close()
@@ -160,12 +160,15 @@ func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
t.Fatalf("Send: %v", err)
}
_, _, body, _, _, _ := rs.snapshot()
if body != s.Body {
t.Fatalf("fallback body: want %q, got %q", s.Body, body)
want := delivery.GenericAwayMessage + ": service_down"
if body != want {
t.Fatalf("body: want %q, got %q", want, body)
}
}
func TestSendRejectsEmptyMessage(t *testing.T) {
func TestSendNeverSendsAnEmptyMessage(t *testing.T) {
// with nothing at all to say we still send the generic line — an away
// channel can never carry detail, but it also never goes out blank.
rs := newRecordingServer(t, 200, "")
srv := httptest.NewServer(rs.handler())
defer srv.Close()
@@ -173,9 +176,13 @@ func TestSendRejectsEmptyMessage(t *testing.T) {
sink, _ := New(Config{BaseURL: srv.URL, Topic: "maven"})
s := nudgeSendable(loop.Sev3, "")
s.Body = ""
err := sink.Send(context.Background(), s)
if err == nil {
t.Fatal("want error for empty message")
s.RuleName = ""
if err := sink.Send(context.Background(), s); err != nil {
t.Fatalf("Send: %v", err)
}
_, _, body, _, _, _ := rs.snapshot()
if body != delivery.GenericAwayMessage {
t.Fatalf("body: want %q, got %q", delivery.GenericAwayMessage, body)
}
}
+2 -3
View File
@@ -208,10 +208,9 @@ func TestAwayChannelsGetMinimalBody(t *testing.T) {
// TestCareAwayDropIsRecorded — DESIGN.md's drop is a decision ("a missed water
// nudge is noise, a missed backup failure isn't"), so it should be visible
// rather than vanish. Today drop is a bare `continue`: no nudge row, no outbox
// attempt, no log — nothing an operator can see afterwards.
// attempt, no log — nothing an operator can see afterwards. now it leaves a
// 'dropped' outbox row.
func TestCareAwayDropIsRecorded(t *testing.T) {
t.Skip("not implemented: dispatcher.go:149-151 skips a Drop channel with no record; there is no 'dropped' outcome in store/delivery.go:16-21")
ob := &fakeOutbox{}
d := NewDispatcher(Config{Voice: &fakeSink{}, Nudges: &fakeNudgeRecorder{}, Outbox: ob})
+12 -16
View File
@@ -2,11 +2,12 @@
//
// telegram is the away-channel for sev4 (ops hard) nudges — "disk-fire alarm
// at 2am routes to telegram, repeat til ack." the message body is the
// Sendable's Summary — the minimal-body rule from the spec ("disk low on
// homesrv," not detail; no shoulder-surf exfil through the relay). voice gets
// Body; away channels get Summary, enforced at the sink so a phraser bug can't
// exfil. additionally, protect_content=true is passed on every send so the
// message can't be forwarded out of the chat — locks the minimal body further.
// delivery.AwayMessage — the minimal-body rule from the spec ("disk low on
// homesrv," not detail; no shoulder-surf exfil through the relay). the
// dispatcher already strips detail off away sendables; the sink uses the same
// helper so it can't leak the body on its own either. additionally,
// protect_content=true is passed on every send so the message can't be
// forwarded out of the chat — locks the minimal body further.
//
// telegram's bot API is region-restricted for this homesrv — direct egress to
// api.telegram.org is unreliable. the spec's "away channels leave the box —
@@ -140,18 +141,13 @@ type telegramResp struct {
}
// Send publishes one message to the configured telegram chat. the body is the
// Sendable's Summary (minimal body); empty Summary falls back to Body (terse
// full message beats no message). protect_content=true so a phraser bug (Body
// leaking detail through Summary) can't be forwarded onward by the user or a
// chat observer — locks the minimal-body rule at the channel's own last mile.
// minimal away message (never the full body). protect_content=true so even
// that can't be forwarded onward by the user or a chat observer — locks the
// minimal-body rule at the channel's own last mile.
func (s *Sink) Send(ctx context.Context, d delivery.Sendable) error {
body := d.Summary
if body == "" {
body = d.Body
}
if body == "" {
return fmt.Errorf("telegramsink: empty message for %s", d.Channel)
}
// never fall back to d.Body: telegram leaves the box, so an empty summary
// gets a generic line instead of the full detail.
body := delivery.AwayMessage(d)
payload := sendMessageReq{
ChatID: s.cfg.ChatID,
@@ -173,9 +173,9 @@ func TestSendBodyIsSummaryNotFullBody(t *testing.T) {
}
}
func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
// terse full message beats none; the phraser should produce a summary for
// away-bound severities, but don't silently drop.
func TestSendNeverSendsTheBodyWhenSummaryEmpty(t *testing.T) {
// #368: this used to fall back to the full body. telegram leaves the box,
// so an empty summary gets a fixed generic line plus the rule name.
rs := newRecordingServer(t, 200, "")
srv := httptest.NewServer(rs.handler())
defer srv.Close()
@@ -188,12 +188,14 @@ func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
_, _, body, _, _ := rs.snapshot()
var req sendMessageReq
_ = json.Unmarshal([]byte(body), &req)
if req.Text != s.Body {
t.Fatalf("fallback text: want %q, got %q", s.Body, req.Text)
want := delivery.GenericAwayMessage + ": service_down"
if req.Text != want {
t.Fatalf("text: want %q, got %q", want, req.Text)
}
}
func TestSendRejectsEmptyMessage(t *testing.T) {
func TestSendNeverSendsAnEmptyMessage(t *testing.T) {
// with nothing at all to say we still send the generic line.
rs := newRecordingServer(t, 200, "")
srv := httptest.NewServer(rs.handler())
defer srv.Close()
@@ -201,9 +203,15 @@ func TestSendRejectsEmptyMessage(t *testing.T) {
sink, _ := New(sinkCfg(srv.URL))
s := nudgeSendable(loop.Sev4, "")
s.Body = ""
err := sink.Send(context.Background(), s)
if err == nil {
t.Fatal("want error for empty message")
s.RuleName = ""
if err := sink.Send(context.Background(), s); err != nil {
t.Fatalf("Send: %v", err)
}
_, _, body, _, _ := rs.snapshot()
var req sendMessageReq
_ = json.Unmarshal([]byte(body), &req)
if req.Text != delivery.GenericAwayMessage {
t.Fatalf("text: want %q, got %q", delivery.GenericAwayMessage, req.Text)
}
}
+74
View File
@@ -1,8 +1,12 @@
package dialogue
import (
"context"
"encoding/json"
"sync"
"time"
"github.com/kami/maven/internal/store"
)
type Intent string
@@ -50,10 +54,22 @@ func (s *Session) IsExpired(now time.Time) bool {
return now.After(s.Timestamp.Add(s.TTL))
}
// SessionPersister — the bit of the store the session needs, as an interface
// so tests can swap it out. Data is an opaque blob: the store never looks
// inside, we encode the session as JSON here.
type SessionPersister interface {
SaveDialogueSession(ctx context.Context, id string, data []byte, ts time.Time, ttl time.Duration) error
DeleteDialogueSession(ctx context.Context, id string) error
LoadDialogueSessions(ctx context.Context, now time.Time) ([]store.DialogueSessionRow, error)
}
// SessionStore keeps the live sessions in a map (the fast path) and mirrors
// every write to the persister, so a daemon restart can load them back.
type SessionStore struct {
mu sync.RWMutex
sessions map[string]*Session
defaultTTL time.Duration
persist SessionPersister // may be nil: memory only (tests, no-store paths)
}
func NewSessionStore(defaultTTL time.Duration) *SessionStore {
@@ -66,6 +82,43 @@ func NewSessionStore(defaultTTL time.Duration) *SessionStore {
}
}
// NewPersistentSessionStore — same store, but writes also go to the DB.
// Call Load once after this to bring back sessions from a previous run.
func NewPersistentSessionStore(defaultTTL time.Duration, p SessionPersister) *SessionStore {
s := NewSessionStore(defaultTTL)
s.persist = p
return s
}
// Load — read the saved sessions back into memory. Anything past its TTL is
// dropped (and deleted from the DB by the store), never revived.
func (s *SessionStore) Load(ctx context.Context, now time.Time) error {
if s.persist == nil {
return nil
}
rows, err := s.persist.LoadDialogueSessions(ctx, now)
if err != nil {
return err
}
s.mu.Lock()
defer s.mu.Unlock()
for _, r := range rows {
var sess Session
if err := json.Unmarshal(r.Data, &sess); err != nil {
// A blob we can't read is not worth failing a startup over.
continue
}
if sess.TTL <= 0 {
sess.TTL = r.TTL
}
if sess.IsExpired(now) {
continue
}
s.sessions[r.ID] = &sess
}
return nil
}
func (s *SessionStore) Get(id string, now time.Time) *Session {
s.mu.RLock()
sess, ok := s.sessions[id]
@@ -87,12 +140,33 @@ func (s *SessionStore) Put(id string, sess *Session) {
s.mu.Lock()
s.sessions[id] = sess
s.mu.Unlock()
s.save(id, sess)
}
func (s *SessionStore) Delete(id string) {
s.mu.Lock()
delete(s.sessions, id)
s.mu.Unlock()
if s.persist != nil {
_ = s.persist.DeleteDialogueSession(context.Background(), id)
}
}
// save — mirror one session to the DB. Best effort: memory already has it, so
// a write error costs us the restart safety net, not the current turn.
func (s *SessionStore) save(id string, sess *Session) {
if s.persist == nil {
return
}
data, err := json.Marshal(sess)
if err != nil {
return
}
ts := sess.Timestamp
if ts.IsZero() {
ts = time.Now()
}
_ = s.persist.SaveDialogueSession(context.Background(), id, data, ts, sess.TTL)
}
func InheritSlots(prev, cur Slots) Slots {
+118
View File
@@ -0,0 +1,118 @@
package dialogue
import (
"context"
"path/filepath"
"testing"
"time"
"github.com/kami/maven/internal/store"
)
// openStore — a store on disk, so a second handle can reopen the same file.
func openStore(t *testing.T, path string) *store.Store {
t.Helper()
s, err := store.Open(context.Background(), path)
if err != nil {
t.Fatalf("open store: %v", err)
}
t.Cleanup(func() { _ = s.Close() })
return s
}
// A session written before a restart comes back and still merges a follow-up.
func TestSessionSurvivesRestart(t *testing.T) {
ctx := context.Background()
path := filepath.Join(t.TempDir(), "maven_test.db")
now := time.Now().UTC().Truncate(time.Millisecond)
first := openStore(t, path)
before := NewPersistentSessionStore(2*time.Minute, first)
before.Put("voice", &Session{
Intent: IntentReminder,
Slots: Slots{Text: "полить цветы", Time: now.Add(time.Hour), HasTime: true},
Timestamp: now,
TTL: 2 * time.Minute,
})
if err := first.Close(); err != nil {
t.Fatalf("close: %v", err)
}
// fresh handle, fresh in-memory map — as after a daemon restart
after := openStore(t, path)
reloaded := NewPersistentSessionStore(2*time.Minute, after)
if err := reloaded.Load(ctx, now.Add(10*time.Second)); err != nil {
t.Fatalf("Load: %v", err)
}
sess := reloaded.Get("voice", now.Add(10*time.Second))
if sess == nil {
t.Fatal("session did not survive the restart")
}
if sess.Intent != IntentReminder {
t.Fatalf("intent = %q, want reminder", sess.Intent)
}
// the follow-up carries no text of its own; it must inherit the old one
merged := InheritSlots(sess.Slots, Slots{Time: now.Add(2 * time.Hour), HasTime: true})
if merged.Text != "полить цветы" {
t.Fatalf("merged text = %q, want the earlier turn's text", merged.Text)
}
if !merged.Time.Equal(now.Add(2 * time.Hour)) {
t.Fatalf("merged time = %v, want the follow-up's time", merged.Time)
}
}
// A session past its TTL is dead: a restart must not bring it back.
func TestExpiredSessionNotResurrected(t *testing.T) {
ctx := context.Background()
path := filepath.Join(t.TempDir(), "maven_test.db")
now := time.Now().UTC().Truncate(time.Millisecond)
first := openStore(t, path)
before := NewPersistentSessionStore(2*time.Minute, first)
before.Put("voice", &Session{
Intent: IntentReminder,
Slots: Slots{Text: "полить цветы"},
Timestamp: now,
TTL: time.Minute,
})
if err := first.Close(); err != nil {
t.Fatalf("close: %v", err)
}
after := openStore(t, path)
reloaded := NewPersistentSessionStore(2*time.Minute, after)
later := now.Add(5 * time.Minute) // well past the 1-min TTL
if err := reloaded.Load(ctx, later); err != nil {
t.Fatalf("Load: %v", err)
}
if sess := reloaded.Get("voice", later); sess != nil {
t.Fatalf("expired session came back: %+v", sess)
}
// and it is gone from the DB too, not just from memory
rows, err := after.LoadDialogueSessions(ctx, later)
if err != nil {
t.Fatalf("LoadDialogueSessions: %v", err)
}
if len(rows) != 0 {
t.Fatalf("expired row still in the DB: %+v", rows)
}
}
// Delete removes the row as well, so an ended conversation stays ended.
func TestDeleteRemovesPersistedSession(t *testing.T) {
ctx := context.Background()
path := filepath.Join(t.TempDir(), "maven_test.db")
now := time.Now().UTC().Truncate(time.Millisecond)
s := openStore(t, path)
ss := NewPersistentSessionStore(2*time.Minute, s)
ss.Put("voice", &Session{Intent: IntentChat, Timestamp: now, TTL: time.Minute})
ss.Delete("voice")
rows, err := s.LoadDialogueSessions(ctx, now)
if err != nil {
t.Fatalf("LoadDialogueSessions: %v", err)
}
if len(rows) != 0 {
t.Fatalf("row survived Delete: %+v", rows)
}
}
+5 -88
View File
@@ -410,95 +410,12 @@ func TestChatViaClient(t *testing.T) {
}
// chatTestAPI — a minimal CoreAPI that only implements Chat for testing.
type chatTestAPI struct{}
// Embeds UnimplementedCoreAPI so every other method fails loudly with
// ErrNotImplemented instead of needing 27 hand-written no-op stubs.
type chatTestAPI struct {
UnimplementedCoreAPI
}
func (a *chatTestAPI) WriteFact(ctx context.Context, req WriteFactReq) (int64, error) {
return 0, ErrUnknownMethod
}
func (a *chatTestAPI) LatestFact(ctx context.Context, key string) (Fact, error) {
return Fact{}, ErrUnknownMethod
}
func (a *chatTestAPI) LatestFactBySource(ctx context.Context, key, source string) (Fact, error) {
return Fact{}, ErrUnknownMethod
}
func (a *chatTestAPI) Since(ctx context.Context, key string, now time.Time) (time.Duration, error) {
return 0, ErrUnknownMethod
}
func (a *chatTestAPI) Presence(ctx context.Context) (Presence, error) {
return Presence{}, ErrUnknownMethod
}
func (a *chatTestAPI) CreateReminder(ctx context.Context, fire time.Time, payload, cron string) (int64, error) {
return 0, ErrUnknownMethod
}
func (a *chatTestAPI) MarkReminder(ctx context.Context, id int64, status string) error {
return ErrUnknownMethod
}
func (a *chatTestAPI) ListReminders(ctx context.Context, n int) ([]Reminder, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) RecordNudge(ctx context.Context, rule, channel, message string, ts time.Time) (int64, error) {
return 0, ErrUnknownMethod
}
func (a *chatTestAPI) ResolveNudge(ctx context.Context, id int64, outcome string, ts time.Time) error {
return ErrUnknownMethod
}
func (a *chatTestAPI) RecentOutcomes(ctx context.Context, rule string, n int) ([]string, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) RecentFacts(ctx context.Context, n int) ([]Fact, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) CalendarEvents(ctx context.Context, from, to time.Time) ([]Fact, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) RecentNudges(ctx context.Context, n int) ([]Nudge, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) WriteNote(ctx context.Context, ts time.Time, text string, embedding []float32, source string) (int64, error) {
return 0, ErrUnknownMethod
}
func (a *chatTestAPI) QueryNotes(ctx context.Context, embedding []float32, k int) ([]Note, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) RecentNotes(ctx context.Context, n int) ([]Note, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) ProposeTool(ctx context.Context, name, utterance, scope string, ts time.Time) (bool, error) {
return false, ErrUnknownMethod
}
func (a *chatTestAPI) EnableTool(ctx context.Context, name string, cmd []string, destructive bool, scope string, ts time.Time) error {
return ErrUnknownMethod
}
func (a *chatTestAPI) DisableTool(ctx context.Context, name string) error {
return ErrUnknownMethod
}
func (a *chatTestAPI) DeleteTool(ctx context.Context, name string) error {
return ErrUnknownMethod
}
func (a *chatTestAPI) ListProposedRoutines(ctx context.Context) ([]ProposedRoutine, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) AcceptProposedRoutine(ctx context.Context, id int64) error {
return nil
}
func (a *chatTestAPI) DismissProposedRoutine(ctx context.Context, id int64) error {
return ErrUnknownMethod
}
func (a *chatTestAPI) LookupTool(ctx context.Context, name string) (Tool, error) {
return Tool{}, ErrUnknownMethod
}
func (a *chatTestAPI) ListTools(ctx context.Context, status string) ([]Tool, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) RevertFact(ctx context.Context, key string) (int64, error) {
return 0, ErrUnknownMethod
}
func (a *chatTestAPI) TickTrace(ctx context.Context) (TickTrace, error) {
return TickTrace{}, ErrUnknownMethod
}
func (a *chatTestAPI) MorningStatus(ctx context.Context) ([]MorningRoutineStatus, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) Chat(ctx context.Context, text string) (string, error) {
if text == "привет" {
return "и тебе привет!", nil
+241 -306
View File
@@ -501,6 +501,239 @@ func (s *Server) safeDispatch(ctx context.Context, req Request) (result json.Raw
return s.dispatch(ctx, req)
}
// handlerFunc — one table entry's shape: unmarshal req.Params (if it wants
// any), call the matching CoreAPI method against the api passed in, marshal
// the result. api is a parameter, not a closed-over field, precisely so a
// table built once at package init never pins a stale CoreAPI — see the note
// on methodTable below about SetAPI.
type handlerFunc func(ctx context.Context, api CoreAPI, raw json.RawMessage) (json.RawMessage, error)
// withParams adapts a (typed params, typed result) CoreAPI call into a
// handlerFunc: unmarshal into P, call fn, marshal R. On error the result is
// dropped (marshalResult's output is never read when err != nil — see
// serveConn) so every entry can uniformly return early on error without
// re-deriving what the pre-table per-arm code used to return in that case.
func withParams[P any, R any](fn func(ctx context.Context, api CoreAPI, p P) (R, error)) handlerFunc {
return func(ctx context.Context, api CoreAPI, raw json.RawMessage) (json.RawMessage, error) {
var p P
if err := unmarshalParams(raw, &p); err != nil {
return nil, err
}
r, err := fn(ctx, api, p)
if err != nil {
return nil, err
}
return marshalResult(r), nil
}
}
// withParamsVoid is withParams for the error-only methods (mark/resolve/
// enable/disable/...): params in, no result out, wire reply is always null.
func withParamsVoid[P any](fn func(ctx context.Context, api CoreAPI, p P) error) handlerFunc {
return func(ctx context.Context, api CoreAPI, raw json.RawMessage) (json.RawMessage, error) {
var p P
if err := unmarshalParams(raw, &p); err != nil {
return nil, err
}
return marshalResult(nil), fn(ctx, api, p)
}
}
// withoutParams is withParams for the handful of methods that take no
// params at all (Presence, TickTrace, MorningStatus, ListProposedRoutines).
// It does NOT call unmarshalParams — matching the pre-table arms, which
// never touched req.Params for these four methods.
func withoutParams[R any](fn func(ctx context.Context, api CoreAPI) (R, error)) handlerFunc {
return func(ctx context.Context, api CoreAPI, _ json.RawMessage) (json.RawMessage, error) {
r, err := fn(ctx, api)
if err != nil {
return nil, err
}
return marshalResult(r), nil
}
}
// methodTable — one entry per CoreAPI-backed method. Built once at package
// init, not per-Server and not per-dispatch: entries close over nothing but
// the CoreAPI method being called, and dispatch passes in the *current*
// api (loaded fresh via s.api.Load() every call, same as before the table
// existed) as an argument — so SetAPI's runtime swap (the unlock transition)
// is still honored on the very next request with no extra plumbing here.
//
// MethodAssertStepUp, MethodStoreEncryptionKey and MethodUnlock are NOT in
// this table: they bypass CoreAPI entirely (s.StepUp / s.WrapKeyFn /
// s.UnlockFn), so dispatch special-cases them before consulting the table.
var methodTable = map[Method]handlerFunc{
MethodWriteFact: withParams(func(ctx context.Context, api CoreAPI, p WriteFactReq) (idResp, error) {
id, err := api.WriteFact(ctx, p)
return idResp{ID: id}, err
}),
MethodLatestFact: withParams(func(ctx context.Context, api CoreAPI, p keyReq) (Fact, error) {
return api.LatestFact(ctx, p.Key)
}),
MethodLatestFactBySource: withParams(func(ctx context.Context, api CoreAPI, p keySourceReq) (Fact, error) {
return api.LatestFactBySource(ctx, p.Key, p.Source)
}),
MethodSince: withParams(func(ctx context.Context, api CoreAPI, p sinceReq) (sinceResp, error) {
d, err := api.Since(ctx, p.Key, p.Now)
return sinceResp{Dur: d}, err
}),
MethodPresence: withoutParams(func(ctx context.Context, api CoreAPI) (Presence, error) {
return api.Presence(ctx)
}),
MethodCreateReminder: withParams(func(ctx context.Context, api CoreAPI, p createReminderReq) (idResp, error) {
id, err := api.CreateReminder(ctx, p.Fire, p.Payload, p.Cron)
return idResp{ID: id}, err
}),
MethodMarkReminder: withParamsVoid(func(ctx context.Context, api CoreAPI, p markReminderReq) error {
return api.MarkReminder(ctx, p.ID, p.Status)
}),
MethodListReminders: withParams(func(ctx context.Context, api CoreAPI, p nReq) ([]Reminder, error) {
out, err := api.ListReminders(ctx, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []Reminder{}
}
return out, nil
}),
MethodRecordNudge: withParams(func(ctx context.Context, api CoreAPI, p recordNudgeReq) (idResp, error) {
id, err := api.RecordNudge(ctx, p.Rule, p.Channel, p.Message, p.Ts)
return idResp{ID: id}, err
}),
MethodResolveNudge: withParamsVoid(func(ctx context.Context, api CoreAPI, p resolveNudgeReq) error {
return api.ResolveNudge(ctx, p.ID, p.Outcome, p.Ts)
}),
MethodRecentOutcomes: withParams(func(ctx context.Context, api CoreAPI, p outcomesReq) ([]string, error) {
out, err := api.RecentOutcomes(ctx, p.Rule, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []string{} // stable non-null on the wire
}
return out, nil
}),
MethodRecentFacts: withParams(func(ctx context.Context, api CoreAPI, p nReq) ([]Fact, error) {
out, err := api.RecentFacts(ctx, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []Fact{}
}
return out, nil
}),
MethodCalendarEvents: withParams(func(ctx context.Context, api CoreAPI, p calendarEventsReq) ([]Fact, error) {
out, err := api.CalendarEvents(ctx, p.From, p.To)
if err != nil {
return nil, err
}
if out == nil {
out = []Fact{}
}
return out, nil
}),
MethodRecentNudges: withParams(func(ctx context.Context, api CoreAPI, p nReq) ([]Nudge, error) {
out, err := api.RecentNudges(ctx, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []Nudge{}
}
return out, nil
}),
MethodWriteNote: withParams(func(ctx context.Context, api CoreAPI, p writeNoteReq) (idResp, error) {
id, err := api.WriteNote(ctx, p.Ts, p.Text, p.Embedding, p.Source)
return idResp{ID: id}, err
}),
MethodQueryNotes: withParams(func(ctx context.Context, api CoreAPI, p queryNotesReq) ([]Note, error) {
out, err := api.QueryNotes(ctx, p.Embedding, p.K)
if err != nil {
return nil, err
}
if out == nil {
out = []Note{}
}
return out, nil
}),
MethodRecentNotes: withParams(func(ctx context.Context, api CoreAPI, p nReq) ([]Note, error) {
out, err := api.RecentNotes(ctx, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []Note{}
}
return out, nil
}),
MethodProposeTool: withParams(func(ctx context.Context, api CoreAPI, p proposeToolReq) (proposeToolResp, error) {
ok, err := api.ProposeTool(ctx, p.Name, p.Utterance, p.Scope, p.Ts)
return proposeToolResp{Proposed: ok}, err
}),
MethodEnableTool: withParamsVoid(func(ctx context.Context, api CoreAPI, p enableToolReq) error {
return api.EnableTool(ctx, p.Name, p.Cmd, p.Destructive, p.Scope, p.Ts)
}),
MethodDisableTool: withParamsVoid(func(ctx context.Context, api CoreAPI, p disableToolReq) error {
return api.DisableTool(ctx, p.Name)
}),
MethodLookupTool: withParams(func(ctx context.Context, api CoreAPI, p lookupToolReq) (Tool, error) {
return api.LookupTool(ctx, p.Name)
}),
MethodListTools: withParams(func(ctx context.Context, api CoreAPI, p listToolsReq) (listToolsResp, error) {
out, err := api.ListTools(ctx, p.Status)
if err != nil {
return listToolsResp{}, err
}
if out == nil {
out = []Tool{}
}
return listToolsResp{Tools: out}, nil
}),
// MethodDeleteTool shares disableToolReq — both take just a tool name.
MethodDeleteTool: withParamsVoid(func(ctx context.Context, api CoreAPI, p disableToolReq) error {
return api.DeleteTool(ctx, p.Name)
}),
MethodListProposedRoutines: withoutParams(func(ctx context.Context, api CoreAPI) (listProposedRoutinesResp, error) {
out, err := api.ListProposedRoutines(ctx)
if err != nil {
return listProposedRoutinesResp{}, err
}
if out == nil {
out = []ProposedRoutine{}
}
return listProposedRoutinesResp{Routines: out}, nil
}),
MethodDismissProposedRoutine: withParamsVoid(func(ctx context.Context, api CoreAPI, p dismissProposedRoutineReq) error {
return api.DismissProposedRoutine(ctx, p.ID)
}),
MethodAcceptProposedRoutine: withParamsVoid(func(ctx context.Context, api CoreAPI, p acceptProposedRoutineReq) error {
return api.AcceptProposedRoutine(ctx, p.ID)
}),
MethodRevertFact: withParams(func(ctx context.Context, api CoreAPI, p revertReq) (map[string]int64, error) {
newID, err := api.RevertFact(ctx, p.Key)
if err != nil {
return nil, err
}
return map[string]int64{"new_id": newID}, nil
}),
MethodChat: withParams(func(ctx context.Context, api CoreAPI, p chatReq) (chatResp, error) {
reply, err := api.Chat(ctx, p.Text)
return chatResp{Reply: reply}, err
}),
MethodTickTrace: withoutParams(func(ctx context.Context, api CoreAPI) (TickTrace, error) {
return api.TickTrace(ctx)
}),
// MorningStatus intentionally has no nil→[]T{} normalization here — the
// pre-table arm marshaled api.MorningStatus's result as-is (a nil slice
// serializes as JSON null), and this preserves that exact wire shape.
MethodMorningStatus: withoutParams(func(ctx context.Context, api CoreAPI) ([]MorningRoutineStatus, error) {
return api.MorningStatus(ctx)
}),
}
// dispatch unmarshals params for req.Method and calls the matching CoreAPI
// method. Unknown method ⇒ ErrUnknownMethod; a malformed params payload ⇒
// ErrBadParams with the underlying text (local, server-side, not shipped to
@@ -517,312 +750,11 @@ func (s *Server) dispatch(ctx context.Context, req Request) (json.RawMessage, er
return nil, err
}
}
// These three bypass CoreAPI entirely — they drive Server fields set
// directly by the daemon (StepUp / WrapKeyFn / UnlockFn), not store
// state, so they can never be table entries keyed on a CoreAPI method.
switch req.Method {
case MethodWriteFact:
var p WriteFactReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
id, err := api.WriteFact(ctx, p)
return marshalResult(idResp{ID: id}), err
case MethodLatestFact:
var p keyReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
f, err := api.LatestFact(ctx, p.Key)
if err != nil {
return nil, err
}
return marshalResult(f), nil
case MethodLatestFactBySource:
var p keySourceReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
f, err := api.LatestFactBySource(ctx, p.Key, p.Source)
if err != nil {
return nil, err
}
return marshalResult(f), nil
case MethodSince:
var p sinceReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
d, err := api.Since(ctx, p.Key, p.Now)
if err != nil {
return nil, err
}
return marshalResult(sinceResp{Dur: d}), nil
case MethodPresence:
pres, err := api.Presence(ctx)
if err != nil {
return nil, err
}
return marshalResult(pres), nil
case MethodCreateReminder:
var p createReminderReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
id, err := api.CreateReminder(ctx, p.Fire, p.Payload, p.Cron)
return marshalResult(idResp{ID: id}), err
case MethodMarkReminder:
var p markReminderReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
err := api.MarkReminder(ctx, p.ID, p.Status)
return marshalResult(nil), err
case MethodListReminders:
var p nReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
out, err := api.ListReminders(ctx, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []Reminder{}
}
return marshalResult(out), nil
case MethodRecordNudge:
var p recordNudgeReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
id, err := api.RecordNudge(ctx, p.Rule, p.Channel, p.Message, p.Ts)
return marshalResult(idResp{ID: id}), err
case MethodResolveNudge:
var p resolveNudgeReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
err := api.ResolveNudge(ctx, p.ID, p.Outcome, p.Ts)
return marshalResult(nil), err
case MethodRecentOutcomes:
var p outcomesReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
out, err := api.RecentOutcomes(ctx, p.Rule, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []string{} // stable non-null on the wire
}
return marshalResult(out), nil
case MethodRecentFacts:
var p nReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
out, err := api.RecentFacts(ctx, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []Fact{}
}
return marshalResult(out), nil
case MethodCalendarEvents:
var p calendarEventsReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
out, err := api.CalendarEvents(ctx, p.From, p.To)
if err != nil {
return nil, err
}
if out == nil {
out = []Fact{}
}
return marshalResult(out), nil
case MethodRecentNudges:
var p nReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
out, err := api.RecentNudges(ctx, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []Nudge{}
}
return marshalResult(out), nil
case MethodWriteNote:
var p writeNoteReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
id, err := api.WriteNote(ctx, p.Ts, p.Text, p.Embedding, p.Source)
return marshalResult(idResp{ID: id}), err
case MethodQueryNotes:
var p queryNotesReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
out, err := api.QueryNotes(ctx, p.Embedding, p.K)
if err != nil {
return nil, err
}
if out == nil {
out = []Note{}
}
return marshalResult(out), nil
case MethodRecentNotes:
var p nReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
out, err := api.RecentNotes(ctx, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []Note{}
}
return marshalResult(out), nil
case MethodProposeTool:
var p proposeToolReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
ok, err := api.ProposeTool(ctx, p.Name, p.Utterance, p.Scope, p.Ts)
if err != nil {
return nil, err
}
return marshalResult(proposeToolResp{Proposed: ok}), nil
case MethodEnableTool:
var p enableToolReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
return marshalResult(nil), api.EnableTool(ctx, p.Name, p.Cmd, p.Destructive, p.Scope, p.Ts)
case MethodDisableTool:
var p disableToolReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
return marshalResult(nil), api.DisableTool(ctx, p.Name)
case MethodLookupTool:
var p lookupToolReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
t, err := api.LookupTool(ctx, p.Name)
if err != nil {
return nil, err
}
return marshalResult(t), nil
case MethodListTools:
var p listToolsReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
out, err := api.ListTools(ctx, p.Status)
if err != nil {
return nil, err
}
if out == nil {
out = []Tool{}
}
return marshalResult(listToolsResp{Tools: out}), nil
case MethodDeleteTool:
var p disableToolReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
return marshalResult(nil), api.DeleteTool(ctx, p.Name)
case MethodListProposedRoutines:
out, err := api.ListProposedRoutines(ctx)
if err != nil {
return nil, err
}
if out == nil {
out = []ProposedRoutine{}
}
return marshalResult(listProposedRoutinesResp{Routines: out}), nil
case MethodDismissProposedRoutine:
var p dismissProposedRoutineReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
return marshalResult(nil), api.DismissProposedRoutine(ctx, p.ID)
case MethodAcceptProposedRoutine:
var p acceptProposedRoutineReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
return marshalResult(nil), api.AcceptProposedRoutine(ctx, p.ID)
case MethodRevertFact:
var p struct {
Key string `json:"key"`
}
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
newID, err := api.RevertFact(ctx, p.Key)
if err != nil {
return nil, err
}
return marshalResult(map[string]int64{"new_id": newID}), nil
case MethodChat:
var p chatReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
reply, err := api.Chat(ctx, p.Text)
if err != nil {
return nil, err
}
return marshalResult(chatResp{Reply: reply}), nil
case MethodTickTrace:
t, err := api.TickTrace(ctx)
if err != nil {
return nil, err
}
return marshalResult(t), nil
case MethodMorningStatus:
s, err := api.MorningStatus(ctx)
if err != nil {
return nil, err
}
return marshalResult(s), nil
case MethodAssertStepUp:
if s.StepUp != nil {
return marshalResult(nil), s.StepUp(ctx)
@@ -848,10 +780,13 @@ func (s *Server) dispatch(ctx context.Context, req Request) (json.RawMessage, er
return marshalResult(nil), s.UnlockFn(ctx, p.PublicKey)
}
return nil, fmt.Errorf("%w: %s", ErrUnknownMethod, req.Method)
}
default:
h, ok := methodTable[req.Method]
if !ok {
return nil, fmt.Errorf("%w: %s", ErrUnknownMethod, req.Method)
}
return h(ctx, api, req.Params)
}
func unmarshalParams(raw json.RawMessage, v any) error {
+118
View File
@@ -0,0 +1,118 @@
package ipc
import (
"context"
"errors"
"time"
)
// ErrNotImplemented is returned by every UnimplementedCoreAPI method. It is
// deliberately distinct from ErrUnknownMethod (a wire-level "no such
// method exists" verdict) and from any daemon-level "locked" error: this one
// means "this method exists on CoreAPI, but the fake/adapter embedding
// UnimplementedCoreAPI never got a real implementation for it." A test that
// exercises an undeclared method fails loudly on this text instead of
// silently nil-panicking or being mistaken for a legitimate failure.
var ErrNotImplemented = errors.New("ipc: not implemented (unimplemented CoreAPI stub)")
// UnimplementedCoreAPI is the gRPC Unimplemented*Server pattern applied to
// CoreAPI: embed it in a test double or adapter and override only the
// methods you actually exercise. Every method returns ErrNotImplemented, so
// a call that reaches an undeclared method fails loudly and specifically,
// rather than compiling to a silent no-op or nil-pointer panic. This
// replaces the old pattern of hand-writing all 30 no-op stubs per double —
// those were compiler-satisfying padding, not tests of anything.
type UnimplementedCoreAPI struct{}
var _ CoreAPI = UnimplementedCoreAPI{}
func (UnimplementedCoreAPI) WriteFact(ctx context.Context, req WriteFactReq) (int64, error) {
return 0, ErrNotImplemented
}
func (UnimplementedCoreAPI) LatestFact(ctx context.Context, key string) (Fact, error) {
return Fact{}, ErrNotImplemented
}
func (UnimplementedCoreAPI) LatestFactBySource(ctx context.Context, key, source string) (Fact, error) {
return Fact{}, ErrNotImplemented
}
func (UnimplementedCoreAPI) Since(ctx context.Context, key string, now time.Time) (time.Duration, error) {
return 0, ErrNotImplemented
}
func (UnimplementedCoreAPI) Presence(ctx context.Context) (Presence, error) {
return Presence{}, ErrNotImplemented
}
func (UnimplementedCoreAPI) CreateReminder(ctx context.Context, fire time.Time, payload, cron string) (int64, error) {
return 0, ErrNotImplemented
}
func (UnimplementedCoreAPI) MarkReminder(ctx context.Context, id int64, status string) error {
return ErrNotImplemented
}
func (UnimplementedCoreAPI) ListReminders(ctx context.Context, n int) ([]Reminder, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) RecordNudge(ctx context.Context, rule, channel, message string, ts time.Time) (int64, error) {
return 0, ErrNotImplemented
}
func (UnimplementedCoreAPI) ResolveNudge(ctx context.Context, id int64, outcome string, ts time.Time) error {
return ErrNotImplemented
}
func (UnimplementedCoreAPI) RecentOutcomes(ctx context.Context, rule string, n int) ([]string, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) RecentFacts(ctx context.Context, n int) ([]Fact, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) CalendarEvents(ctx context.Context, from, to time.Time) ([]Fact, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) RecentNudges(ctx context.Context, n int) ([]Nudge, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) WriteNote(ctx context.Context, ts time.Time, text string, embedding []float32, source string) (int64, error) {
return 0, ErrNotImplemented
}
func (UnimplementedCoreAPI) QueryNotes(ctx context.Context, embedding []float32, k int) ([]Note, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) RecentNotes(ctx context.Context, n int) ([]Note, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) ProposeTool(ctx context.Context, name, utterance, scope string, ts time.Time) (bool, error) {
return false, ErrNotImplemented
}
func (UnimplementedCoreAPI) EnableTool(ctx context.Context, name string, cmd []string, destructive bool, scope string, ts time.Time) error {
return ErrNotImplemented
}
func (UnimplementedCoreAPI) DisableTool(ctx context.Context, name string) error {
return ErrNotImplemented
}
func (UnimplementedCoreAPI) DeleteTool(ctx context.Context, name string) error {
return ErrNotImplemented
}
func (UnimplementedCoreAPI) ListProposedRoutines(ctx context.Context) ([]ProposedRoutine, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) DismissProposedRoutine(ctx context.Context, id int64) error {
return ErrNotImplemented
}
func (UnimplementedCoreAPI) AcceptProposedRoutine(ctx context.Context, id int64) error {
return ErrNotImplemented
}
func (UnimplementedCoreAPI) LookupTool(ctx context.Context, name string) (Tool, error) {
return Tool{}, ErrNotImplemented
}
func (UnimplementedCoreAPI) ListTools(ctx context.Context, status string) ([]Tool, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) RevertFact(ctx context.Context, key string) (int64, error) {
return 0, ErrNotImplemented
}
func (UnimplementedCoreAPI) TickTrace(ctx context.Context) (TickTrace, error) {
return TickTrace{}, ErrNotImplemented
}
func (UnimplementedCoreAPI) MorningStatus(ctx context.Context) ([]MorningRoutineStatus, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) Chat(ctx context.Context, text string) (string, error) {
return "", ErrNotImplemented
}
+116
View File
@@ -0,0 +1,116 @@
// Package kiwix reads a local Kiwix server (offline Wikipedia and friends).
//
// Why: the resident model is a 0.8B and invents facts. Letting her read a local
// article snippet beats letting her recall. Nothing here talks to the internet;
// the Kiwix server is on the same box.
//
// This is search only. Full articles are ~100KB of HTML, far too big for a 4096
// token context, so the unit of context is the search snippet (~500 chars).
package kiwix
import (
"context"
"encoding/xml"
"fmt"
"html"
"io"
"net/http"
"net/url"
"regexp"
"strconv"
"strings"
"time"
)
// Result is one search hit.
type Result struct {
Title string // article title, e.g. "Rayleigh scattering"
Path string // e.g. /content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering
Snippet string // plain text, tags stripped, entities decoded
WordCount int // 0 if the server did not say
}
// Client is a Kiwix HTTP client. Boring on purpose: no retries, no cache.
type Client struct {
base string
http *http.Client
}
// New makes a client for a Kiwix base URL like http://127.0.0.1:8034.
func New(baseURL string) *Client {
return &Client{
base: strings.TrimRight(baseURL, "/"),
http: &http.Client{Timeout: 10 * time.Second},
}
}
// Search runs a keyword search in one ZIM (book) and returns up to limit hits.
//
// Ranking is keyword based, not semantic: "Rayleigh scattering" finds the right
// article, "why is the sky blue" finds a TV episode. Pass keywords, not questions.
func (c *Client) Search(ctx context.Context, pattern, book string, limit int) ([]Result, error) {
if limit <= 0 {
limit = 5
}
q := url.Values{}
q.Set("pattern", pattern)
q.Set("books.name", book)
q.Set("format", "xml")
q.Set("pageLength", strconv.Itoa(limit))
req, err := http.NewRequestWithContext(ctx, http.MethodGet, c.base+"/search?"+q.Encode(), nil)
if err != nil {
return nil, err
}
resp, err := c.http.Do(req)
if err != nil {
return nil, err
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return nil, fmt.Errorf("kiwix search: http %d", resp.StatusCode)
}
return ParseSearchRSS(resp.Body)
}
// rss mirrors just the bits of the RSS 2.0 reply we use.
type rss struct {
Items []struct {
Title string `xml:"title"`
Link string `xml:"link"`
// innerxml keeps the <b> match markers so we can strip them ourselves.
Description struct {
Inner string `xml:",innerxml"`
} `xml:"description"`
WordCount string `xml:"wordCount"`
} `xml:"channel>item"`
}
var tagRE = regexp.MustCompile(`<[^>]*>`)
// ParseSearchRSS turns a Kiwix search reply into results. Exported so the parser
// is testable from a captured response, with no server running.
func ParseSearchRSS(r io.Reader) ([]Result, error) {
var doc rss
if err := xml.NewDecoder(r).Decode(&doc); err != nil {
return nil, fmt.Errorf("kiwix search: bad xml: %w", err)
}
out := make([]Result, 0, len(doc.Items))
for _, it := range doc.Items {
n, _ := strconv.Atoi(strings.ReplaceAll(it.WordCount, ",", ""))
out = append(out, Result{
Title: strings.TrimSpace(it.Title),
Path: strings.TrimSpace(it.Link),
Snippet: plainText(it.Description.Inner),
WordCount: n,
})
}
return out, nil
}
// plainText drops markup and decodes entities, leaving text a model can read.
func plainText(s string) string {
s = tagRE.ReplaceAllString(s, "")
s = html.UnescapeString(s)
return strings.TrimSpace(strings.Join(strings.Fields(s), " "))
}
+90
View File
@@ -0,0 +1,90 @@
package kiwix
import (
"context"
"os"
"strings"
"testing"
"time"
)
// A real reply from the live server, trimmed to two items.
const sampleRSS = `<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:opensearch="http://a9.com/-/spec/opensearch/1.1/">
<channel>
<title>Search: Rayleigh scattering</title>
<opensearch:totalResults>800</opensearch:totalResults>
<item>
<title>Rayleigh scattering</title>
<link>/content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering</link>
<description><b>Rayleigh</b> scattering causes the blue color of the sky &amp; yellow colors near the Sun.[1]</description>
<book><title>Wikipedia</title></book>
<wordCount>2,818</wordCount>
</item>
<item>
<title>HyperRayleigh scattering</title>
<link>/content/wikipedia_en_all_maxi_2026-02/Hyper%E2%80%93Rayleigh_scattering</link>
<description>...<b>Rayleigh</b> scattering" is a nonlinear optical counterpart.</description>
<book><title>Wikipedia</title></book>
<wordCount>914</wordCount>
</item>
</channel>
</rss>`
func TestParseSearchRSS(t *testing.T) {
got, err := ParseSearchRSS(strings.NewReader(sampleRSS))
if err != nil {
t.Fatalf("parse: %v", err)
}
if len(got) != 2 {
t.Fatalf("want 2 results, got %d", len(got))
}
if got[0].Title != "Rayleigh scattering" {
t.Errorf("title = %q", got[0].Title)
}
if got[0].Path != "/content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering" {
t.Errorf("path = %q", got[0].Path)
}
if got[0].WordCount != 2818 {
t.Errorf("wordCount = %d", got[0].WordCount)
}
want := "Rayleigh scattering causes the blue color of the sky & yellow colors near the Sun.[1]"
if got[0].Snippet != want {
t.Errorf("snippet = %q, want %q", got[0].Snippet, want)
}
if strings.Contains(got[1].Snippet, "<b>") {
t.Errorf("second snippet still has tags: %q", got[1].Snippet)
}
}
func TestParseSearchRSSBadXML(t *testing.T) {
if _, err := ParseSearchRSS(strings.NewReader("not xml at all")); err == nil {
t.Fatal("want an error on junk input")
}
}
// Opt-in: needs a live Kiwix server. CI has none.
// MAVEN_KIWIX_URL=http://127.0.0.1:8034 no_proxy=127.0.0.1,localhost go test -run Retrieval -v ./internal/kiwix/
func TestRetrievalEval(t *testing.T) {
base := os.Getenv("MAVEN_KIWIX_URL")
if base == "" {
t.Skip("set MAVEN_KIWIX_URL to run the retrieval eval")
}
noProxyLoopback(t)
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
defer cancel()
rep, err := RunRetrievalEval(ctx, New(base), 5)
if err != nil {
t.Fatalf("eval: %v", err)
}
// No pass bar on purpose: the number is the finding.
t.Log("\n" + rep.String() + rep.Detail())
}
// noProxyLoopback stops the box's SOCKS bridge from eating loopback requests.
func noProxyLoopback(t *testing.T) {
t.Setenv("no_proxy", "127.0.0.1,localhost")
t.Setenv("NO_PROXY", "127.0.0.1,localhost")
}
+63
View File
@@ -0,0 +1,63 @@
{
"name": "kiwix-knowledge-v1",
"book": "wikipedia_en_all_maxi_2026-02",
"note": "The 9 knowledge cases from internal/phraser/eval/talk_v1.json. Queries are hand-written English keywords on purpose: Kiwix ranks by keyword, not meaning, so a natural question fails. Writing them by hand separates 'retrieval is broken' from 'the model writes bad queries'.",
"cases": [
{
"id": "know-sky-blue",
"question": "почему небо синее?",
"query": "Rayleigh scattering sky blue",
"want_titles": ["Rayleigh scattering", "Diffuse sky radiation"]
},
{
"id": "know-boil-egg",
"question": "сколько варить яйцо вкрутую?",
"query": "boiled egg cooking",
"want_titles": ["Boiled egg", "Egg as food"]
},
{
"id": "know-ssd-vs-hdd",
"question": "чем ssd отличается от hdd?",
"query": "solid-state drive",
"want_titles": ["Solid-state drive", "Hard disk drive"]
},
{
"id": "know-cat-purr",
"question": "почему кошки мурчат?",
"query": "cat purr",
"want_titles": ["Purr", "Cat communication"]
},
{
"id": "know-hiccups",
"question": "как быстро избавиться от икоты?",
"query": "hiccup",
"want_titles": ["Hiccup"]
},
{
"id": "know-polite-form",
"question": "не могли бы вы объяснить, что такое vpn?",
"query": "virtual private network",
"want_titles": ["Virtual private network"]
},
{
"id": "know-dont-know",
"question": "как зовут моего соседа снизу?",
"query": "name of my downstairs neighbour",
"want_titles": [],
"expect_miss": true,
"note": "Unanswerable by design. Retrieval SHOULD find nothing useful. Counted as a hit only when nothing relevant comes back."
},
{
"id": "know-water-per-day",
"question": "сколько воды в день надо пить?",
"query": "human daily water requirement drinking",
"want_titles": ["Drinking water", "Water", "Dehydration", "Hydration"]
},
{
"id": "know-thunder-delay",
"question": "почему гром слышно позже молнии?",
"query": "thunder speed of sound lightning",
"want_titles": ["Thunder", "Lightning"]
}
]
}
+135
View File
@@ -0,0 +1,135 @@
package kiwix
// This scores retrieval alone: no LLM. For each general-knowledge question we
// hand-write English keywords and ask whether the article that would answer it
// comes back in the top N hits. If this score is low, reading Wikipedia cannot
// help the model no matter how good the prompt is.
//
// The unanswerable case (know-dont-know) is not scored. Whether the junk it
// returns is "nothing useful" is a human judgement, so the report just prints
// the titles and leaves the score to the 8 answerable cases.
import (
"context"
_ "embed"
"encoding/json"
"fmt"
"strings"
)
//go:embed knowledge_v1.json
var knowledgeFixtureJSON []byte
// EvalCase — one question with hand-written keywords.
type EvalCase struct {
ID string `json:"id"`
Question string `json:"question"`
Query string `json:"query"`
WantTitles []string `json:"want_titles"`
ExpectMiss bool `json:"expect_miss"`
}
type fixture struct {
Name string `json:"name"`
Book string `json:"book"`
Cases []EvalCase `json:"cases"`
}
// Outcome — what one case retrieved.
type Outcome struct {
Case EvalCase
Titles []string // titles of the top N hits, in rank order
Rank int // 1-based rank of the first wanted title, 0 if none
Err error
}
// Hit is true when a wanted title came back.
func (o Outcome) Hit() bool { return o.Rank > 0 }
// Report — the score plus per-case detail.
type Report struct {
Name string
Book string
TopN int
Scored int // answerable cases
Hits int
Errors int
Outcomes []Outcome
}
// Accuracy over the answerable cases.
func (r Report) Accuracy() float64 {
if r.Scored == 0 {
return 0
}
return float64(r.Hits) / float64(r.Scored)
}
// RunRetrievalEval searches for every fixture case.
func RunRetrievalEval(ctx context.Context, c *Client, topN int) (Report, error) {
var f fixture
if err := json.Unmarshal(knowledgeFixtureJSON, &f); err != nil {
return Report{}, err
}
rep := Report{Name: f.Name, Book: f.Book, TopN: topN}
for _, cs := range f.Cases {
res, err := c.Search(ctx, cs.Query, f.Book, topN)
o := Outcome{Case: cs, Err: err}
if err != nil {
rep.Errors++
}
for i, hit := range res {
o.Titles = append(o.Titles, hit.Title)
if o.Rank == 0 && matches(cs.WantTitles, hit.Title) {
o.Rank = i + 1
}
}
if !cs.ExpectMiss {
rep.Scored++
if o.Hit() {
rep.Hits++
}
}
rep.Outcomes = append(rep.Outcomes, o)
}
return rep, nil
}
func matches(want []string, title string) bool {
for _, w := range want {
if strings.EqualFold(strings.TrimSpace(title), w) {
return true
}
}
return false
}
// String — the headline number.
func (r Report) String() string {
var b strings.Builder
fmt.Fprintf(&b, "%s: %d/%d answerable questions retrieve a wanted article in top %d (%.1f%%), %d errors\n",
r.Name, r.Hits, r.Scored, r.TopN, 100*r.Accuracy(), r.Errors)
fmt.Fprintf(&b, " book: %s\n", r.Book)
return b.String()
}
// Detail — per case: what was asked, what was searched, what came back.
func (r Report) Detail() string {
var b strings.Builder
for _, o := range r.Outcomes {
mark := "MISS"
switch {
case o.Case.ExpectMiss:
mark = "n/a "
case o.Hit():
mark = fmt.Sprintf("hit@%d", o.Rank)
}
fmt.Fprintf(&b, " %-6s %-20s q=%q\n", mark, o.Case.ID, o.Case.Query)
if o.Err != nil {
fmt.Fprintf(&b, " error: %v\n", o.Err)
continue
}
fmt.Fprintf(&b, " got: %s\n", strings.Join(o.Titles, " | "))
}
return b.String()
}
+183
View File
@@ -0,0 +1,183 @@
package kiwix
// Turning a Russian question into an English Kiwix search.
//
// Kiwix ranks by keyword, not by meaning. "why is the sky blue" returns a TV
// episode; "Rayleigh scattering sky blue" returns the right article. So the
// model's job here is NOT translation — it is naming the English article the
// answer lives in.
//
// The output space is a handful of words, so it is worth locking down hard: a
// GBNF grammar for the shape, a tiny token cap, and a cleanup pass that throws
// away anything odd rather than handing junk to Kiwix.
import (
"context"
"encoding/json"
"fmt"
"strings"
"unicode"
"github.com/kami/maven/internal/llm"
)
// Completer — the LLM seam, so tests can fake it. *llm.Client satisfies it.
type Completer interface {
Complete(ctx context.Context, r llm.Req) (string, error)
}
// queryGrammar — one JSON object holding 1..6 keyword words. Latin letters,
// digits and hyphens only, so the model physically cannot answer the question
// or reply in Russian.
//
// Why the JSON wrapper: this model always thinks out loud and this llama-server
// build ignores the thinking switch (see ROUTING-EVAL-31-07-2026.md). A bare
// word-list grammar just captured the reasoning — every case came back as
// "Let me analyze this request carefully". Demanding JSON, like routeGrammar and
// responseGrammar already do, gives the reasoning nowhere to go.
const queryGrammar = `
root ::= "{" ws "\"query\"" ws ":" ws "\"" word (" " word){0,5} "\"" ws "}"
word ::= [A-Za-z0-9] [A-Za-z0-9-]{0,23}
ws ::= [ \t\n]*
`
// rewriteSystem — asks for search keywords, not an answer and not a translation.
const rewriteSystem = `You turn a question into a search query for English Wikipedia.
Rules:
- Output ONLY English search keywords. Never an answer, never an explanation.
- Do NOT translate the sentence. Name the thing the answer is about.
- The output must be a noun phrase, like a Wikipedia article title.
- Never use question words: no why, how, what, when, which, "how much",
"how long", "how to", "vs", "reason", "difference".
- 2 to 4 words.
Reply with JSON: {"query":"<keywords>"}
Good:
"почему листья желтеют осенью?" -> {"query":"leaf senescence autumn"}
"как работает микроволновка?" -> {"query":"microwave oven"}
"не могли бы вы объяснить, что такое блокчейн?" -> {"query":"blockchain"}
"сколько живут собаки?" -> {"query":"dog lifespan"}
"как избавиться от комаров в квартире?" -> {"query":"mosquito control"}
"чем чай отличается от кофе?" -> {"query":"tea"}
Only JSON, no explanation.`
// maxQueryTokens — the output is a few words plus the JSON wrapper. A tight cap
// is the cheapest guard against the model rambling into an answer.
const maxQueryTokens = 32
// Rewriter asks the resident model for English search keywords.
type Rewriter struct{ c Completer }
func NewRewriter(c Completer) *Rewriter { return &Rewriter{c: c} }
// Rewrite returns English keywords for a question in any language.
// It errors rather than returning something Kiwix should not see.
func (r *Rewriter) Rewrite(ctx context.Context, question string) (string, error) {
raw, err := r.c.Complete(ctx, llm.Req{
System: rewriteSystem,
User: strings.TrimSpace(question),
Grammar: queryGrammar,
MaxTokens: maxQueryTokens,
RepeatPenalty: 1.15,
})
if err != nil {
return "", err
}
return CleanQuery(unwrapJSON(raw))
}
// unwrapJSON pulls the query out of {"query":"..."}. If the reply is not that
// shape it is returned as-is, and CleanQuery decides whether it is usable.
func unwrapJSON(raw string) string {
s := strings.TrimSpace(raw)
if !strings.HasPrefix(s, "{") {
return s
}
var got struct{ Query string }
if err := json.Unmarshal([]byte(s), &got); err != nil {
return s
}
return got.Query
}
// maxQueryWords matches the grammar's bound. Anything longer is prose.
const maxQueryWords = 6
// CleanQuery checks and tidies whatever the model produced. The grammar makes
// bad output unlikely, not impossible (a server without grammar support, a
// different model), so this is the real gate in front of Kiwix.
//
// Exported so it can be tested without a model.
func CleanQuery(raw string) (string, error) {
s := strings.TrimSpace(raw)
// Models like to wrap answers in quotes. Drop surrounding ones.
s = strings.Trim(s, "\"'`")
// Keep the first line only: everything after it is prose.
if i := strings.IndexAny(s, "\r\n"); i >= 0 {
s = s[:i]
}
// Keep letters, digits, spaces and hyphens; anything else becomes a space.
var b strings.Builder
for _, ru := range s {
switch {
case unicode.IsLetter(ru) || unicode.IsDigit(ru) || ru == '-':
b.WriteRune(ru)
default:
b.WriteRune(' ')
}
}
words := strings.Fields(b.String())
if len(words) == 0 {
return "", fmt.Errorf("kiwix rewrite: empty query")
}
if len(words) > maxQueryWords {
return "", fmt.Errorf("kiwix rewrite: %d words, want at most %d (looks like prose)", len(words), maxQueryWords)
}
words = dropStopWords(words)
out := strings.Join(words, " ")
// The ZIMs are English. Non-Latin letters mean the model ignored the ask.
for _, ru := range out {
if unicode.IsLetter(ru) && !isLatin(ru) {
return "", fmt.Errorf("kiwix rewrite: query is not English: %q", out)
}
}
return out, nil
}
// stopWords — question words and filler. The model keeps writing question-shaped
// queries ("why is the sky blue", "how much water to drink daily") no matter how
// the prompt is worded, and Kiwix ranks on every word, so those words drag in
// song and episode titles. Dropping them in code is not a style preference: a
// keyword ranker gets nothing from them.
var stopWords = map[string]bool{
"a": true, "an": true, "the": true, "is": true, "are": true, "was": true,
"do": true, "does": true, "did": true, "to": true, "of": true, "in": true,
"on": true, "for": true, "and": true, "or": true, "my": true, "me": true,
"i": true, "it": true, "its": true, "be": true, "been": true, "get": true,
"how": true, "why": true, "what": true, "when": true, "which": true,
"who": true, "where": true, "much": true, "many": true, "long": true,
"vs": true, "than": true, "rid": true, "from": true, "about": true,
}
// dropStopWords removes filler, but never everything: if the query was nothing
// but stop words there is nothing better to search, so the original is kept and
// the caller sees whatever Kiwix makes of it.
func dropStopWords(words []string) []string {
kept := make([]string, 0, len(words))
for _, w := range words {
if !stopWords[strings.ToLower(w)] {
kept = append(kept, w)
}
}
if len(kept) == 0 {
return words
}
return kept
}
func isLatin(ru rune) bool {
return (ru >= 'a' && ru <= 'z') || (ru >= 'A' && ru <= 'Z')
}
+96
View File
@@ -0,0 +1,96 @@
package kiwix
// End-to-end score: Russian question -> model rewrite -> Kiwix search -> did a
// wanted article come back. Same 9 cases as the retrieval eval, so the two
// numbers are directly comparable: retrieval with hand-written keywords is the
// ceiling, this is what the model actually reaches.
import (
"context"
"encoding/json"
"fmt"
"strings"
)
// RewriteOutcome — one case, end to end.
type RewriteOutcome struct {
Outcome
ModelQuery string // what the model asked for ("" if it failed)
RewriteErr error
}
// RunRewriteEval rewrites every question with the model, then searches.
func RunRewriteEval(ctx context.Context, c *Client, rw *Rewriter, topN int) (RewriteReport, error) {
var f fixture
if err := json.Unmarshal(knowledgeFixtureJSON, &f); err != nil {
return RewriteReport{}, err
}
rep := RewriteReport{Report: Report{Name: f.Name + "-rewrite", Book: f.Book, TopN: topN}}
for _, cs := range f.Cases {
out := RewriteOutcome{Outcome: Outcome{Case: cs}}
q, err := rw.Rewrite(ctx, cs.Question)
out.ModelQuery, out.RewriteErr = q, err
if err == nil {
res, serr := c.Search(ctx, q, f.Book, topN)
out.Err = serr
for i, hit := range res {
out.Titles = append(out.Titles, hit.Title)
if out.Rank == 0 && matches(cs.WantTitles, hit.Title) {
out.Rank = i + 1
}
}
}
if out.RewriteErr != nil || out.Err != nil {
rep.Errors++
}
if !cs.ExpectMiss {
rep.Scored++
if out.Hit() {
rep.Hits++
}
}
rep.Cases = append(rep.Cases, out)
}
return rep, nil
}
// RewriteReport — the score plus per-case detail.
type RewriteReport struct {
Report
Cases []RewriteOutcome
}
// String — the headline number.
func (r RewriteReport) String() string {
return fmt.Sprintf("%s: %d/%d answerable questions retrieve a wanted article in top %d (%.1f%%), %d errors\n book: %s\n",
r.Name, r.Hits, r.Scored, r.TopN, 100*r.Accuracy(), r.Errors, r.Book)
}
// Detail — per case: hand-written query next to the model's, and what came back.
// The point is seeing WHERE the model's phrasing differs, not just the score.
func (r RewriteReport) Detail() string {
var b strings.Builder
for _, o := range r.Cases {
mark := "MISS"
switch {
case o.Case.ExpectMiss:
mark = "n/a "
case o.Hit():
mark = fmt.Sprintf("hit@%d", o.Rank)
}
fmt.Fprintf(&b, " %-6s %-20s\n", mark, o.Case.ID)
fmt.Fprintf(&b, " asked: %s\n", o.Case.Question)
fmt.Fprintf(&b, " hand: %q\n", o.Case.Query)
fmt.Fprintf(&b, " model: %q\n", o.ModelQuery)
if o.RewriteErr != nil {
fmt.Fprintf(&b, " rewrite rejected: %v\n", o.RewriteErr)
continue
}
if o.Err != nil {
fmt.Fprintf(&b, " search error: %v\n", o.Err)
continue
}
fmt.Fprintf(&b, " got: %s\n", strings.Join(o.Titles, " | "))
}
return b.String()
}
+33
View File
@@ -0,0 +1,33 @@
package kiwix
import (
"context"
"os"
"testing"
"time"
"github.com/kami/maven/internal/llm"
)
// Opt-in: needs a live Kiwix server AND a live llama-server.
// MAVEN_KIWIX_URL=http://127.0.0.1:8034 MAVEN_LLM_URL=http://127.0.0.1:18099 \
//
// no_proxy=127.0.0.1,localhost go test -run RewriteEval -v ./internal/kiwix/
func TestRewriteEval(t *testing.T) {
kbase, lbase := os.Getenv("MAVEN_KIWIX_URL"), os.Getenv("MAVEN_LLM_URL")
if kbase == "" || lbase == "" {
t.Skip("set MAVEN_KIWIX_URL and MAVEN_LLM_URL to run the rewrite eval")
}
noProxyLoopback(t)
ctx, cancel := context.WithTimeout(context.Background(), 15*time.Minute)
defer cancel()
rw := NewRewriter(llm.New(lbase, 3*time.Minute))
rep, err := RunRewriteEval(ctx, New(kbase), rw, 5)
if err != nil {
t.Fatalf("eval: %v", err)
}
// No pass bar on purpose: the number is the finding.
t.Log("\n" + rep.String() + rep.Detail())
}
+105
View File
@@ -0,0 +1,105 @@
package kiwix
import (
"context"
"testing"
"github.com/kami/maven/internal/llm"
)
// Bad model output must never reach Kiwix. No model needed for this.
func TestCleanQueryRejectsJunk(t *testing.T) {
bad := []struct{ name, raw string }{
{"empty", ""},
{"blank", " \n "},
{"russian came back", "почему небо синее"},
{"mixed russian", "sky синее scattering"},
{"full sentence", "The sky looks blue because of the scattering of sunlight by air molecules"},
{"prose with quotes", `Sure! Here is a good search query: "Rayleigh scattering", which explains it.`},
}
for _, c := range bad {
if got, err := CleanQuery(c.raw); err == nil {
t.Errorf("%s: want rejection, got %q", c.name, got)
}
}
}
func TestCleanQueryCleans(t *testing.T) {
ok := []struct{ raw, want string }{
{"Rayleigh scattering sky", "Rayleigh scattering sky"},
{" boiled egg cooking \n", "boiled egg cooking"},
{`"virtual private network"`, "virtual private network"},
{"solid-state drive", "solid-state drive"},
{"cat purr.", "cat purr"},
{"hiccup\nAlso: hiccough", "hiccup"},
// Question words are filler to a keyword ranker, so they go.
{"why is the sky blue", "sky blue"},
{"how much water to drink daily", "water drink daily"},
{"SSD vs HDD comparison", "SSD HDD comparison"},
// Nothing but filler: keep it rather than return nothing.
{"what is it", "what is it"},
}
for _, c := range ok {
got, err := CleanQuery(c.raw)
if err != nil {
t.Errorf("%q: %v", c.raw, err)
continue
}
if got != c.want {
t.Errorf("%q -> %q, want %q", c.raw, got, c.want)
}
}
}
type fakeCompleter struct {
out string
req llm.Req
}
func (f *fakeCompleter) Complete(_ context.Context, r llm.Req) (string, error) {
f.req = r
return f.out, nil
}
func TestRewriteConstrainsTheCall(t *testing.T) {
f := &fakeCompleter{out: `{"query":"Rayleigh scattering sky"}`}
got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?")
if err != nil {
t.Fatalf("rewrite: %v", err)
}
if got != "Rayleigh scattering sky" {
t.Errorf("query = %q", got)
}
if f.req.Grammar == "" {
t.Error("no grammar sent")
}
if f.req.MaxTokens == 0 || f.req.MaxTokens > 32 {
t.Errorf("max_tokens = %d, want a small cap", f.req.MaxTokens)
}
}
func TestRewriteRejectsBadModelOutput(t *testing.T) {
bad := []string{
`{"query":"почему небо синее"}`, // never translated
`{"query":""}`, // empty
`{"query":"the sky is blue because sunlight is scattered by air"}`, // an answer
// Note: a SHORT English prose fragment ("Let me analyze this request")
// is under the word cap and cannot be caught here. The grammar is what
// stops that one.
}
for _, out := range bad {
f := &fakeCompleter{out: out}
if got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?"); err == nil {
t.Errorf("%s: want rejection, got %q", out, got)
}
}
}
// A reply that is not the JSON shape but is still usable keywords should pass.
func TestRewriteFallsBackToPlainText(t *testing.T) {
f := &fakeCompleter{out: "Rayleigh scattering sky"}
got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?")
if err != nil || got != "Rayleigh scattering sky" {
t.Errorf("got %q, %v", got, err)
}
}
@@ -1,4 +1,4 @@
package eval
package llm
import (
"context"
@@ -6,8 +6,20 @@ import (
"fmt"
"net/http"
"strings"
"time"
)
// UnknownModel is the label to print when the server would not say what it has
// loaded. Deliberately ugly: an honest "unknown" is fine, a plausible-looking
// but wrong model name is the bug this whole file exists to prevent.
const UnknownModel = "unknown-model"
// llama-server is local, so never send this through a proxy: this box's
// http_proxy answers 503 for loopback, which would look like "server won't say
// which model it has" when the server is right there and fine.
// A Transport with no Proxy set bypasses http_proxy entirely.
var modelHTTP = &http.Client{Timeout: 10 * time.Second, Transport: &http.Transport{}}
// ModelID asks llama-server which model it has loaded, so a scoring run can
// label itself. Without this a bake-off between two models produces two tables
// that look identical, and the operator has to remember which server was up.
@@ -19,7 +31,7 @@ func ModelID(ctx context.Context, base string) (string, error) {
if err != nil {
return "", err
}
resp, err := http.DefaultClient.Do(req)
resp, err := modelHTTP.Do(req)
if err != nil {
return "", err
}
@@ -38,14 +50,21 @@ func ModelID(ctx context.Context, base string) (string, error) {
if len(out.Data) == 0 {
return "", fmt.Errorf("models: empty list")
}
return shortModelID(out.Data[0].ID), nil
short := shortModelID(out.Data[0].ID)
if short == "" {
// Server answered but the id field was missing or blank. Say so
// instead of handing back an empty label that reads as a real name.
return "", fmt.Errorf("models: no id in response")
}
return short, nil
}
// shortModelID trims the path and the .gguf suffix — llama-server reports the
// file name it was started with, which is too long for a table header.
func shortModelID(id string) string {
id = strings.TrimSpace(id)
if i := strings.LastIndexAny(id, "/\\"); i >= 0 {
id = id[i+1:]
}
return strings.TrimSuffix(id, ".gguf")
return strings.TrimSpace(strings.TrimSuffix(id, ".gguf"))
}
+64
View File
@@ -0,0 +1,64 @@
package llm
import (
"context"
"net/http"
"net/http/httptest"
"testing"
)
// The point of these tests: a wrong-but-plausible model label is the bug, so
// every path that cannot learn the real name must return an error instead of a
// guess. No llama-server needed — a stub server stands in.
func TestModelID(t *testing.T) {
cases := []struct {
name string
body string
code int
want string // "" ⇒ expect an error
}{
{"full path", `{"data":[{"id":"/mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf"}]}`, 200, "Qwen3.5-0.8B.Q4_K_M"},
{"bare name", `{"data":[{"id":"LFM2.5-1.2B"}]}`, 200, "LFM2.5-1.2B"},
{"empty list", `{"data":[]}`, 200, ""},
{"id missing", `{"data":[{}]}`, 200, ""},
{"id blank", `{"data":[{"id":" "}]}`, 200, ""},
{"server error", `nope`, 500, ""},
{"not json", `<html>`, 200, ""},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path != "/v1/models" {
t.Errorf("asked for %s, want /v1/models", r.URL.Path)
}
w.WriteHeader(c.code)
_, _ = w.Write([]byte(c.body))
}))
defer srv.Close()
got, err := ModelID(context.Background(), srv.URL+"/")
if c.want == "" {
if err == nil {
t.Fatalf("want an error, got label %q", got)
}
return
}
if err != nil {
t.Fatalf("ModelID: %v", err)
}
if got != c.want {
t.Errorf("got %q, want %q", got, c.want)
}
})
}
}
func TestModelIDUnreachable(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(http.ResponseWriter, *http.Request) {}))
url := srv.URL
srv.Close() // nothing listening now
if got, err := ModelID(context.Background(), url); err == nil {
t.Fatalf("want an error from a dead server, got label %q", got)
}
}
+48
View File
@@ -286,6 +286,54 @@ func TestRemindersStillHonourSnooze(t *testing.T) {
}
}
// ---------------------------- digest eligibility ------------------------------
// A Sev2 care candidate (break) suppressed for a genuine restraint reason is
// worth resurfacing later.
func TestDigestEligibleSev2SuppressedByRestraint(t *testing.T) {
for _, reason := range []string{"quiet_hours", "calendar_busy", "presence"} {
if !DigestEligible(Sev2, reason) {
t.Errorf("sev2 blocked by %q: want digest-eligible", reason)
}
}
}
// A Sev1 care candidate (water/meal) never digests — a biological timer
// nudge is stale by the time anyone could resurface it, so it just drops.
func TestDigestEligibleSev1NeverDigests(t *testing.T) {
for _, reason := range []string{"quiet_hours", "calendar_busy", "presence"} {
if DigestEligible(Sev1, reason) {
t.Errorf("sev1 blocked by %q: want drop, got digest-eligible", reason)
}
}
}
// Ops severities are never blocked by these reasons in practice (Gate only
// applies quiet_hours/calendar_busy/presence to care severities), but the
// boundary itself must refuse to digest a high severity even if asked —
// alarms bypass the gate and deliver now, unchanged, never delayed.
func TestDigestEligibleNeverDigestsHighSeverity(t *testing.T) {
for _, sev := range []Severity{Sev3, Sev4} {
for _, reason := range []string{"quiet_hours", "calendar_busy", "presence"} {
if DigestEligible(sev, reason) {
t.Errorf("sev%d blocked by %q: high severity must never digest", sev, reason)
}
}
}
}
// cooldown and snooze are not "suppression" in the digest sense — cooldown
// means it was already said recently, snooze means the user asked to not
// hear about it. Neither should resurface later just because the severity
// matches.
func TestDigestEligibleExcludesCooldownAndSnooze(t *testing.T) {
for _, reason := range []string{"cooldown", "snooze", "inert_no_data", "predicate", ""} {
if DigestEligible(Sev2, reason) {
t.Errorf("sev2 blocked by %q: should not be digest-eligible", reason)
}
}
}
// GAP — the gate reads State.SnoozeUntil, but the Gatherer hard-codes it to nil
// (internal/loop/gather.go:153), so snooze is dead in the running daemon: the
// unit tests above pass while nothing can ever populate the map. This asserts
+30
View File
@@ -105,6 +105,36 @@ func Tick(s State, rules []Rule) *Candidate {
return fire
}
// DigestEligible decides digest-vs-drop for a care candidate the gate
// suppressed this tick (see ExplainGate's blockedBy). Pure — no I/O, no
// state, just the two facts that matter: why it was suppressed, and how
// insistent it was.
//
// Only genuine RESTRAINT blocks are eligible at all — quiet_hours,
// calendar_busy, presence(away). cooldown and snooze are not suppression in
// this sense: cooldown means "you already heard this recently" (resurfacing
// it later would be an actual repeat, not a rescue) and snooze is the user
// explicitly saying "not this" (digesting it anyway would defeat the ask).
// Ops severities (Sev3/4) never reach here — the gate never blocks them for
// these reasons in the first place (see Gate), and even if a future rule
// dropped Sev3+ into "care", digest still refuses them: alarms bypass the
// gate on purpose and must never be silently delayed into a bundle.
//
// Within care (Sev12), the boundary is severity itself: Sev1 (water, meal —
// biological timers with no "still relevant later" property; a water nudge
// from 3 hours into quiet hours is just wrong by morning) drops. Sev2
// (break — "you worked through a long stretch without a break while I
// couldn't reach you") is information that stays true and useful after the
// fact, so it digests.
func DigestEligible(sev Severity, blockedBy string) bool {
switch blockedBy {
case "quiet_hours", "calendar_busy", "presence":
default:
return false
}
return sev == Sev2
}
// ReminderDecision — a due reminder the daemon should deliver now.
// NOT gated by the universal Gate (per spec: "wake me 7" fires in quiet hours;
// that's the point). Snooze is the one part of restraint that still applies.
+2
View File
@@ -422,6 +422,8 @@ func rankNote(inTop3 bool) string {
// bestRecall mirrors cmd/mavend/recall.go — the gate the daemon actually
// applies to a memory hit. Duplicated rather than imported because package main
// is not importable; recalleval_test.go asserts the two agree in behaviour.
// The daemon returns the whole hit (a note and a fact are said differently);
// the harness only scores what came back, so it keeps returning the text.
func bestRecall(results []memory.Result, minScore, minMargin float64) string {
if !memory.Confident(results, minScore, minMargin) {
return ""
@@ -197,7 +197,8 @@ func TestHashRecallBaseline(t *testing.T) {
t.Log("\n" + rep.String() + rep.Failures())
t.Log("\ngate sweep:\n" + sweep(t, router.NewHashEmbedder(hashDim), f))
// 0.32 sits under the observed 0.360 recall@1.
// 0.32 sits under the observed 0.370 recall@1 (was 0.360 over 25 answerable
// cases; the two mixed note+fact cases added with #373 make it 27).
const floorRecall1 = 0.32
if rep.Recall1() < floorRecall1 {
t.Errorf("recall@1 %.3f below ratchet %.2f — note recall regressed", rep.Recall1(), floorRecall1)
@@ -387,6 +387,32 @@
{"id": "n2", "text": "wifi channel is 6", "kind": "note"},
{"id": "n3", "text": "the guest network is off", "kind": "note"}
]
},
{
"id": "ru-mixed-031",
"lang": "ru",
"tags": ["mixed", "paraphrase", "hard"],
"note": "notes and facts in one store and the note is the answer — the daemon indexes both (Vikunja #373)",
"query": "куда я спрятал второй ключ от квартиры",
"want": "n1",
"notes": [
{"id": "n1", "text": "запасной ключ от квартиры лежит в синей коробке на полке", "kind": "note"},
{"id": "x1", "text": "поменял замок в двери двадцатого июня", "kind": "fact"},
{"id": "x2", "text": "отдал ключ соседке в мае", "kind": "fact"}
]
},
{
"id": "ru-mixed-032",
"lang": "ru",
"tags": ["mixed", "distractor"],
"note": "the mirror of ru-mixed-031: the fact answers and the notes are the distractors",
"query": "когда я в последний раз заливал бензин",
"want": "x1",
"notes": [
{"id": "x1", "text": "залил полный бак в четверг вечером", "kind": "fact"},
{"id": "n1", "text": "на заправке у моста дешевле бензин", "kind": "note"},
{"id": "n2", "text": "надо поменять зимние шины", "kind": "note"}
]
}
]
}
+138
View File
@@ -0,0 +1,138 @@
// Package persona builds the one shared context block that goes in front of
// every LLM system prompt: who the owner is, how to address him, and what
// time it is right now.
//
// Why one block and not a line pasted into each prompt: there are five
// prompts (nudges, action replies, chat, note queries, general knowledge) and
// the "address him as ты" rule had only reached two of them. Five copies drift.
// One block cannot.
//
// The rules here are defaults in code, not config. Maven is feminine and the
// owner is a man addressed informally — that is a hard constraint of the
// product, so it must hold with an empty config file. Config only ADDS
// optional facts (his name, his city).
package persona
import (
"fmt"
"strings"
"time"
)
// Facts — the optional, deployment-specific half of the block. All fields may
// be empty; the block is still correct and useful without them.
type Facts struct {
OwnerName string // his name, e.g. "Ками"
City string // where he is, e.g. "Москва"
Static string // the free-text `persona` config string, appended verbatim
// The two config-gated capabilities. They are listed only when this
// deployment actually has them, because a capability she names and cannot
// do is worse than one she never mentions.
Weather bool // an open-meteo provider is configured
Telegram bool // a telegram bot token + chat id are configured
Tools bool // at least one shell act is on the allowlist
}
var ruWeekdays = [...]string{"воскресенье", "понедельник", "вторник", "среда", "четверг", "пятница", "суббота"}
var ruMonths = [...]string{
"января", "февраля", "марта", "апреля", "мая", "июня",
"июля", "августа", "сентября", "октября", "ноября", "декабря",
}
// Block renders the context block for one turn. Russian even in front of the
// English prompts: the rules it states are Russian grammar (ты/тебя, feminine
// verbs), and a Russian rule reads best stated in Russian.
//
// Keep it short. It ships on every turn to a 0.8B on laptop CPU, so every
// line here is latency.
func (f Facts) Block(now time.Time) string {
var b strings.Builder
b.WriteString("Ты — Maven, домашняя ассистентка. О себе говоришь в женском роде: \"я записала\", \"я проверила\".\n")
// The address form gets its own line. It is the thing that kept getting
// lost when it was buried in prose.
b.WriteString("ОБРАЩЕНИЕ: владелец — мужчина, всегда на \"ты\" (ты, тебя, тебе, твой) и в единственном числе (\"выпей\", \"посмотри\"). Никогда \"вы\"/\"вас\"/\"ваш\". Никогда \"он\"/\"его\" о нём — ты говоришь ему, а не о нём. Глаголы о нём — в мужском роде (\"ты забыл\").\n")
if who := f.who(); who != "" {
b.WriteString(who + "\n")
}
b.WriteString(fmt.Sprintf("Сейчас: %s, %d %s %d, %02d:%02d (местное время).\n",
ruWeekdays[int(now.Weekday())], now.Day(), ruMonths[int(now.Month())-1], now.Year(),
now.Hour(), now.Minute()))
b.WriteString("Умеешь: " + strings.Join(f.can(), "; ") +
". Других ДЕЙСТВИЙ не умеешь — если просят такое, скажи прямо.\n")
if s := strings.TrimSpace(f.Static); s != "" {
b.WriteString(s + "\n")
}
return b.String()
}
// can lists what she can really do. Every entry here is a code path that
// exists in the daemon today:
// - reminders: IntentReminder → CoreAPI.CreateReminder, fired by the tick.
// - notes and facts: IntentNote/IntentFact write, IntentQuery reads them back.
// - calendar: IntentQuery answers "что у меня сегодня" from CalendarEvents.
// - weather / telegram / shell acts: only when configured (see Facts).
//
// Nothing speculative goes in this list. A capability she offers and cannot
// perform is worse than one she never mentions.
func (f Facts) can() []string {
c := []string{
// Talking comes first, and the closing line says "действий" rather than
// "ничего", because this same block sits in front of the chat and
// general-knowledge prompts. A flat "you can do nothing else" would
// tell her to refuse the exact thing those two prompts are for.
"разговаривать и отвечать на вопросы",
"ставить напоминания",
"записывать заметки и факты и отвечать по ним",
"смотреть календарь",
}
if f.Weather {
c = append(c, "говорить погоду")
}
if f.Telegram {
c = append(c, "писать в телеграм")
}
if f.Tools {
c = append(c, "запускать разрешённые команды на сервере")
}
return c
}
// who renders the optional name/city line, or "" when neither is configured.
//
// Written as labels ("Имя владельца: ..."), not as a sentence with pronouns:
// the block's own "ты" is Maven, so "тебя зовут" would read as her name and
// "его" would model the third-person form she must never use about him.
func (f Facts) who() string {
name := strings.TrimSpace(f.OwnerName)
city := strings.TrimSpace(f.City)
switch {
case name != "" && city != "":
return "Имя владельца: " + name + ". Город: " + city + "."
case name != "":
return "Имя владельца: " + name + "."
case city != "":
return "Город: " + city + "."
}
return ""
}
// Prepend puts the block in front of a system prompt. Nil-safe: a nil renderer
// (tests, the stub paths) returns the prompt untouched.
func Prepend(block func() string, prompt string) string {
if block == nil {
return prompt
}
s := strings.TrimSpace(block())
if s == "" {
return prompt
}
return s + "\n\n" + prompt
}
+72
View File
@@ -0,0 +1,72 @@
package persona
import (
"strings"
"testing"
"time"
)
var ref = time.Date(2026, 7, 31, 14, 5, 0, 0, time.UTC)
// The block must be correct with an empty config: the address form and the
// gender rules are hard constraints, not preferences.
func TestBlockWorksWithZeroConfig(t *testing.T) {
b := Facts{}.Block(ref)
for _, want := range []string{"женском роде", "ОБРАЩЕНИЕ", "\"ты\"", "31 июля 2026", "пятница", "14:05"} {
if !strings.Contains(b, want) {
t.Errorf("block missing %q:\n%s", want, b)
}
}
}
func TestBlockAddsOptionalFacts(t *testing.T) {
b := Facts{OwnerName: "Ками", City: "Москва", Static: "Будь краткой."}.Block(ref)
for _, want := range []string{"Ками", "Москва", "Будь краткой."} {
if !strings.Contains(b, want) {
t.Errorf("block missing %q:\n%s", want, b)
}
}
}
// The time changes between turns, so two renders must differ.
func TestBlockRendersTimePerTurn(t *testing.T) {
a := Facts{}.Block(ref)
c := Facts{}.Block(ref.Add(time.Hour))
if a == c {
t.Errorf("block did not change with the clock:\n%s", a)
}
}
// She may only offer what this deployment actually has.
func TestCapabilitiesAreConfigGated(t *testing.T) {
bare := Facts{}.Block(ref)
for _, want := range []string{"напоминания", "заметки", "календарь"} {
if !strings.Contains(bare, want) {
t.Errorf("block missing always-on capability %q:\n%s", want, bare)
}
}
for _, unwanted := range []string{"погоду", "телеграм", "команды"} {
if strings.Contains(bare, unwanted) {
t.Errorf("block offers unconfigured %q:\n%s", unwanted, bare)
}
}
full := Facts{Weather: true, Telegram: true, Tools: true}.Block(ref)
for _, want := range []string{"погоду", "телеграм", "команды"} {
if !strings.Contains(full, want) {
t.Errorf("block missing configured capability %q:\n%s", want, full)
}
}
}
func TestPrependNilIsSafe(t *testing.T) {
if got := Prepend(nil, "PROMPT"); got != "PROMPT" {
t.Errorf("Prepend(nil) = %q", got)
}
if got := Prepend(func() string { return " " }, "PROMPT"); got != "PROMPT" {
t.Errorf("Prepend(blank) = %q", got)
}
if got := Prepend(func() string { return "CTX" }, "PROMPT"); got != "CTX\n\nPROMPT" {
t.Errorf("Prepend = %q", got)
}
}
+54
View File
@@ -0,0 +1,54 @@
package phraser
import (
"errors"
"strings"
"testing"
)
// A reply that starts a JSON object and never finishes it is a failed
// generation, not a reply. Before this, the parser returned ("", "") for these
// and every caller then shipped the raw fragment as the thing Maven said. A
// real run produced replies of literally "{" and "{\n \"".
func TestParseResponseMoodRejectsUnfinishedJSON(t *testing.T) {
for _, raw := range []string{
`{`,
"{\n \"",
`{"response": "неполн`,
`{"response": "текст", "mood":`,
} {
text, mood, err := parseResponseMood(raw)
if !errors.Is(err, errBrokenJSON) {
t.Errorf("parseResponseMood(%q) err = %v, want errBrokenJSON", raw, err)
}
if text != "" || mood != "" {
t.Errorf("parseResponseMood(%q) leaked %q/%q — a fragment must never come back as a reply", raw, text, mood)
}
}
}
// Bare prose is still fine. Small models sometimes answer without any JSON at
// all, and that reply is usable — so the new error must not swallow it.
func TestParseResponseMoodAllowsBareProse(t *testing.T) {
for _, raw := range []string{
"норм, а ты как?",
"вот что я нашла: ключ у соседа",
} {
text, mood, err := parseResponseMood(raw)
if err != nil {
t.Errorf("parseResponseMood(%q) err = %v, want nil", raw, err)
}
// No JSON means no fields; the caller ships raw as-is.
if text != "" || mood != "" {
t.Errorf("parseResponseMood(%q) = %q/%q, want empty", raw, text, mood)
}
}
}
// The measured failure: the model wants more than 400 characters and the old
// grammar cut it off mid-word. Guards the bound against being tightened back.
func TestGrammarStringBoundHasRoomForARealAnswer(t *testing.T) {
if !strings.Contains(responseGrammar, "{0,1000}") {
t.Error("grammar string bound is not 1000; 400 truncated real replies mid-word (see the comment on responseGrammar)")
}
}
+33
View File
@@ -0,0 +1,33 @@
package phraser
import (
"strings"
"testing"
)
// Every phrasing prompt must carry the shared context block. This is the
// regression guard for the bug that started this: the "ты" rule reached only
// two of the five prompts because each prompt had its own copy of the rules.
func TestEveryPromptCarriesTheContextBlock(t *testing.T) {
block := func() string { return "CTXBLOCK" }
p := &LLMPhraser{cfg: Config{ContextBlock: block}}
prompts := map[string]string{
"nudge": p.systemPrompt(),
"query": p.querySystemPrompt(),
"chat": chatSystemPrompt(block),
}
for name, got := range prompts {
if !strings.HasPrefix(got, "CTXBLOCK\n\n") {
t.Errorf("%s prompt does not start with the context block:\n%s", name, got)
}
}
}
// Without a block the prompts are unchanged — the stub and test paths pass nil.
func TestPromptsWithoutBlockAreUnchanged(t *testing.T) {
p := &LLMPhraser{}
if p.systemPrompt() != nudgeSystem {
t.Errorf("nudge prompt changed with no block set")
}
}
@@ -0,0 +1,65 @@
package eval
import (
"strings"
"testing"
)
// TestAddressReportsEveryBreak — the real reply from a nudge eval run broke in
// two ways at once and the check named only the plural. Both must print: a
// half-reported failure reads as a milder problem than it is.
func TestAddressReportsEveryBreak(t *testing.T) {
body := "Смотрите на его потребление воды."
res := checkAddress(body)
if res.Pass {
t.Fatalf("checkAddress passed %q", body)
}
for _, want := range []string{"смотрите", "его"} {
if !strings.Contains(res.Detail, want) {
t.Errorf("detail %q does not name %q", res.Detail, want)
}
}
}
// One word repeated is one problem, so the detail must not say it twice.
func TestAddressDeduplicates(t *testing.T) {
res := checkAddress("Вам стоит поесть, вам это нужно.")
if res.Pass {
t.Fatal("expected failure")
}
if n := strings.Count(res.Detail, "formal"); n != 1 {
t.Errorf("detail repeats the same break %d times: %q", n, res.Detail)
}
}
// The fragments a real run produced. All of them scored as non-empty replies
// before checkNonEmpty looked for letters.
func TestNonEmptyNeedsLetters(t *testing.T) {
for _, body := range []string{
"{",
"{\n \"",
"15-16",
`{"`,
" ",
"...",
} {
if got := checkNonEmpty(body); got.Pass {
t.Errorf("checkNonEmpty(%q) passed — that is not a reply", body)
}
}
}
// And it must not start failing real replies. Latin counts as well as Cyrillic:
// answers about ssd or vpn are legitimately part English.
func TestNonEmptyAcceptsRealReplies(t *testing.T) {
for _, body := range []string{
"норм, а ты как?",
"вот что я нашла: ключ у соседа",
"ssd быстрее hdd.",
"9 минут.",
} {
if got := checkNonEmpty(body); !got.Pass {
t.Errorf("checkNonEmpty(%q) failed: %s", body, got.Detail)
}
}
}
@@ -0,0 +1,42 @@
package eval
import "testing"
func TestAddressTimeWordDoesNotBlind(t *testing.T) {
// A nudge that opens with a time word must still be caught. Without the
// time words in the stoplist, "сегодня" was read as the third party.
for _, s := range []string{
"сегодня он не ел 11 дней",
"вчера он не пил воду",
"опять он забыл про таблетки",
} {
if r := checkAddress(s); r.Pass {
t.Errorf("checkAddress(%q) passed, want a third-person failure", s)
}
}
// Still must not fire when a third party really is named.
for _, s := range []string{
"сегодня сервис упал, он не отвечает",
"ты не пил воду четыре часа",
} {
if r := checkAddress(s); !r.Pass {
t.Errorf("checkAddress(%q) failed: %s", s, r.Detail)
}
}
}
// TestAddressVerbIsNotAnAntecedent — a nudge is mostly verbs, and a verb is
// never who "он" refers to. This exact string passed the check before.
func TestAddressVerbIsNotAnAntecedent(t *testing.T) {
s := "попробуй встать и отдохнуть — у него есть перерыв"
if r := checkAddress(s); r.Pass {
t.Errorf("checkAddress(%q) passed, want a third-person failure", s)
}
// Still missed, and this is the documented hole: "выпей воды, он не пил" has
// a real noun ("воды") before the pronoun, so the scan believes somebody
// else was named. Telling that apart needs a parser, not a suffix rule.
// A named third party still wins over the verbs around it.
if r := checkAddress("сервис упал, он не отвечает"); !r.Pass {
t.Errorf("checkAddress on a real third party failed: %s", r.Detail)
}
}
+370 -3
View File
@@ -17,10 +17,18 @@ const (
CheckFeminine = "feminine" // her self-reference is feminine (hard constraint)
CheckCringe = "cringe" // DESIGN.md § Non-goals, "not a relationship"
CheckOnTopic = "ontopic" // says the thing the rule is about
// CheckHisGender — the other half of the persona rule: SHE is feminine, HE
// is male. "ты давно не отдыхала" addresses the operator as a woman.
CheckHisGender = "hisgender"
// CheckAddress — she talks TO him, informally, one to one. Not "вы", not
// "он". See the comment block above checkAddress.
CheckAddress = "address"
)
// CheckNames — report order.
var CheckNames = []string{CheckMood, CheckLang, CheckLength, CheckFeminine, CheckCringe, CheckOnTopic}
var CheckNames = []string{CheckMood, CheckLang, CheckLength, CheckFeminine, CheckHisGender, CheckAddress, CheckCringe, CheckOnTopic}
// Result — one check on one message.
type Result struct {
@@ -52,6 +60,8 @@ func RunChecks(c Case, body, mood string) []Result {
checkLang(body),
checkLength(body),
checkFeminine(body),
checkHisGender(body),
checkAddress(body),
checkCringe(body),
checkOnTopic(c, body),
}
@@ -176,6 +186,315 @@ func checkFeminine(body string) Result {
return Result{CheckFeminine, true, ""}
}
// --- he is male ----------------------------------------------------------
//
// The mirror of checkFeminine, and the failure it was written for: the model
// wrote "ты давно не отдыхала", which addresses the operator as a woman. That
// scored clean, because checkFeminine only ever looks at how SHE speaks about
// herself.
//
// How it works: Russian past tense is gendered by suffix, -л (m) / -ла (f). So
// this looks for feminine past-tense words in a sentence that also talks TO him
// ("ты", "тебя", "тебе", "твой", …). A feminine verb that belongs to her ("я
// заметила", "напомнила тебе") is skipped — that one is correct.
//
// Honest about the limits: this is a suffix rule, not a parser.
// - False positives: a feminine noun can be the subject in the same sentence
// ("зарядка была утром, ты её пропустил"). The guard below skips a verb whose
// previous word looks like a feminine noun, which helps but will not always
// be right.
// - False negatives: gender also shows up outside the past tense (short
// adjectives, "сама"), and none of that is checked here.
//
// That is acceptable for an eval check. It is a signal to read the message, not
// a grammar verdict, and every hit prints the word it tripped on so a human can
// disagree.
// hisMarkers — words that mean the sentence is addressed to him.
var hisMarkers = map[string]bool{
"ты": true, "тебя": true, "тебе": true, "тобой": true, "тобою": true,
"твой": true, "твоя": true, "твоё": true, "твое": true, "твои": true, "твою": true,
}
// notFeminineVerb — ordinary words ending in "-ла" that are not verbs. Small on
// purpose: it only has to cover words a nudge might actually use.
//
// Words that are both a noun and a verb are deliberately NOT here. "села",
// "мыла" and "стекла" are nouns on paper, but in a nudge they are almost always
// verbs ("ты села", "ты мыла"), and listing them would make the check miss the
// exact thing it is for. Missing a real hit is worse than one false alarm.
var notFeminineVerb = map[string]bool{
"школа": true, "скала": true, "игла": true, "метла": true, "смола": true,
"дела": true, "тела": true, "масла": true, "весла": true,
"зола": true, "пчела": true, "числа": true,
}
// femininePast reports whether a word looks like a feminine past-tense verb:
// "отдыхала", "поела", "выспалась".
func femininePast(w string) bool {
if len([]rune(w)) < 3 || notFeminineVerb[w] {
return false
}
return strings.HasSuffix(w, "ла") || strings.HasSuffix(w, "лась")
}
// looksFeminineNoun — a crude guard against "зарядка была": a word right before
// the verb that ends in "а"/"я" and is not itself a verb is probably the subject.
func looksFeminineNoun(w string) bool {
if femininePast(w) || len([]rune(w)) < 3 {
return false
}
return strings.HasSuffix(w, "а") || strings.HasSuffix(w, "я")
}
// sentenceRE splits on sentence-ending punctuation, so a feminine verb in one
// sentence is not blamed on a "ты" in the next.
var sentenceRE = regexp.MustCompile(`[.!?;…]+`)
func checkHisGender(body string) Result {
for _, sentence := range sentenceRE.Split(strings.ToLower(body), -1) {
words := wordRE.FindAllString(sentence, -1)
addressed := false
for _, w := range words {
if hisMarkers[w] {
addressed = true
}
}
if !addressed {
continue
}
for i, w := range words {
if !femininePast(w) || hersNotHis(words, i) {
continue
}
if i > 0 && looksFeminineNoun(prevWord(words, i)) {
continue
}
return Result{CheckHisGender, false,
fmt.Sprintf("feminine %q addressed to him — he is male", w)}
}
}
return Result{CheckHisGender, true, ""}
}
// hersNotHis — the verb is Maven's own if "я" comes shortly before it, or if the
// thing she did was done to him ("напомнила тебе", "проверила за тебя").
func hersNotHis(words []string, i int) bool {
for j := i - 1; j >= 0 && j >= i-3; j-- {
if words[j] == "я" {
return true
}
}
if i+1 < len(words) {
switch words[i+1] {
case "тебе", "тебя", "за", "тобой":
return true
}
}
return false
}
// prevWord — the word before i, skipping "не" and punctuation, so "не отдыхала"
// still sees the subject.
func prevWord(words []string, i int) string {
for j := i - 1; j >= 0; j-- {
w := words[j]
if w == "не" || w == "ни" || !unicode.Is(unicode.Cyrillic, []rune(w)[0]) {
continue
}
return w
}
return ""
}
// --- how she addresses him ------------------------------------------------
//
// Persona hard constraint: Maven speaks TO him, informally, one to one. The
// phrasing eval produced two breaks of it, and both scored clean:
//
// - "Приходите… Жду вас" — the formal plural. Correct is ты/тебя/тебе and a
// singular imperative ("приходи", "жду тебя").
// - "Он не ел 11 дней" — she talks ABOUT him, in the third person, as if
// reporting to somebody else. Correct is "ты не ел 11 дней".
//
// Like checkHisGender this is a keyword + suffix heuristic, NOT a parser. Every
// hit prints the word it tripped on, so a false alarm is obvious at a glance and
// can be dismissed.
//
// Part 1, formal address. Two signals:
// - the "вы" pronoun family, matched as whole words, so there is nothing to
// exclude — "вы" and "вас" are never anything else.
// - a plural verb ending: -ите/-ете/-йте/-ьте ("приходите", "выпейте",
// "не забудьте", "хотите"). Nouns in the prepositional case share those
// endings ("в интернете", "в свете"), so a word right after a preposition is
// skipped. That is the whole exclusion list, on purpose: a bigger one would
// start swallowing real imperatives.
//
// Part 2, third person. "он" is perfectly fine when the message really is about
// somebody or something else ("сервис упал, он не отвечает"). The way to tell
// them apart: a legitimate third person has an ANTECEDENT — the thing it refers
// to was named earlier in the message. So "он" is only flagged when nothing
// before it in the message could be that thing.
//
// Where this gives up, plainly:
// - it only looks BACKWARD. "Он не отвечает, сервис упал" names the subject
// after the pronoun and is flagged wrongly.
// - any noun earlier in the message counts as an antecedent, even when it is
// not one ("после обеда он не ел", "выпей воды, он не пил" — both missed).
// Verbs and time words no longer count, which covers the usual nudge, but a
// plain noun before the pronoun still blinds it. The
// common time words are stoplisted so the usual nudge opening does not
// blind it, but a message with any other noun in front still slips through.
// This is the check's real hole; widening it further would start flagging
// legitimate third-party messages, so it stops here.
// - a message that opens with "ты" and only later slips into "он" is missed,
// because "ты" itself is skipped but the words around it are not.
// - formal address outside these endings (short adjectives, "вашими" style
// forms not listed) is missed.
// addressWordRE also takes Latin words, because "him"/"he" is the same break in
// English.
var addressWordRE = regexp.MustCompile(`[\p{Cyrillic}]+|[a-zA-Z]+|[,.;:!?…—-]`)
// formalPronouns — the "вы" family. Whole-word match, so no false hits.
var formalPronouns = map[string]bool{
"вы": true, "вас": true, "вам": true, "вами": true,
"ваш": true, "ваша": true, "ваше": true, "ваши": true,
"вашего": true, "вашей": true, "вашему": true, "вашим": true,
"вашими": true, "вашу": true,
}
// prepositions — used twice: to skip prepositional-case nouns that look like
// plural verbs, and as words that cannot be what "он" refers to.
var prepositions = map[string]bool{
"в": true, "во": true, "на": true, "о": true, "об": true, "обо": true,
"при": true, "по": true, "за": true, "из": true, "с": true, "со": true,
"к": true, "ко": true, "до": true, "от": true, "у": true, "над": true,
"под": true, "про": true, "без": true, "для": true, "через": true,
}
// pluralVerb reports whether a word looks like a plural/formal verb form:
// "приходите", "выпейте", "забудьте", "хотите".
func pluralVerb(w string) bool {
if len([]rune(w)) < 5 {
return false
}
return strings.HasSuffix(w, "ите") || strings.HasSuffix(w, "ете") ||
strings.HasSuffix(w, "йте") || strings.HasSuffix(w, "ьте")
}
// thirdPersonHim — pronouns that would be talking about him instead of to him.
var thirdPersonHim = map[string]bool{
"он": true, "его": true, "ему": true, "него": true, "нему": true, "ним": true,
"he": true, "him": true, "his": true,
}
// notAnAntecedent — words that cannot be the thing "он" refers to: pronouns,
// particles, conjunctions, adverbs of time. If only these come before "он", the
// message never named a third party and "он" is him.
var notAnAntecedent = map[string]bool{
"не": true, "ни": true, "и": true, "а": true, "но": true, "да": true,
"же": true, "бы": true, "ли": true, "вот": true, "уже": true,
"ещё": true, "еще": true, "тоже": true, "там": true, "тут": true,
"здесь": true, "это": true, "что": true, "как": true, "когда": true,
"чтобы": true, "потому": true, "сейчас": true, "потом": true,
// Time words. A nudge almost always opens with one ("сегодня он не ел"),
// and without them the very next word is read as the person being talked
// about, so the check misses the exact break it was written for.
"сегодня": true, "вчера": true, "завтра": true, "послезавтра": true,
"утром": true, "днём": true, "днем": true, "вечером": true, "ночью": true,
"опять": true, "снова": true, "весь": true, "всю": true, "целый": true,
"я": true, "мне": true, "меня": true, "мной": true, "мы": true, "нас": true,
"ты": true, "тебя": true, "тебе": true, "тобой": true,
"твой": true, "твоя": true, "твоё": true, "твое": true, "твои": true, "твою": true,
}
// looksVerb — a verb is never the thing "он" refers to, so it must not count as
// an antecedent. Past tense keeps "сервис упал, он не отвечает" working off
// "сервис"; the infinitive and imperative endings are here because a nudge is
// mostly made of them ("попробуй встать и отдохнуть — у него есть перерыв"
// slipped through with "попробуй" taken for the person being talked about).
func looksVerb(w string) bool {
r := []rune(w)
if len(r) < 3 {
return false
}
for _, suf := range []string{
"л", "ла", "ло", "ли", // past tense
"ть", "ться", "ти", "чь", // infinitive
"й", "йся", "йте", // imperative
} {
if strings.HasSuffix(w, suf) {
return true
}
}
return false
}
func checkAddress(body string) Result {
words := addressWordRE.FindAllString(strings.ToLower(body), -1)
// Every break, not just the first. A bad reply usually breaks in more than
// one way at once — "Смотрите на его потребление воды" is a plural imperative
// AND third person about him — and reporting only the first hid the second,
// which made the failure look milder than it was.
var breaks []string
seen := map[string]bool{}
add := func(msg string) {
if seen[msg] {
return // the same word twice in one message is one problem, not two
}
seen[msg] = true
breaks = append(breaks, msg)
}
for i, w := range words {
if formalPronouns[w] {
add(fmt.Sprintf("formal %q — she says ты/тебя/тебе", w))
}
if pluralVerb(w) && !(i > 0 && prepositions[words[i-1]]) {
add(fmt.Sprintf("plural imperative %q — she uses the singular", w))
}
}
for i, w := range words {
if !thirdPersonHim[w] {
continue
}
named := false
for j := 0; j < i; j++ {
p := words[j]
if !unicode.Is(unicode.Cyrillic, []rune(p)[0]) && !isLatinWord(p) {
continue // punctuation
}
// pluralVerb as well as looksVerb: looksVerb knows the imperative in
// -й/-йте but not the -те plural ("смотрите"), so "Смотрите на его
// потребление воды" counted "смотрите" as the person being talked
// about and the "его" never printed. Third time a verb form has
// blinded this check — if a fourth turns up, the antecedent test
// wants a real morphology table, not another suffix.
if notAnAntecedent[p] || prepositions[p] || thirdPersonHim[p] || looksVerb(p) || pluralVerb(p) {
continue
}
named = true
break
}
if !named {
add(fmt.Sprintf("third person %q with nobody else named — she talks to him, not about him", w))
}
}
if len(breaks) > 0 {
return Result{CheckAddress, false, strings.Join(breaks, " + ")}
}
return Result{CheckAddress, true, ""}
}
func isLatinWord(w string) bool {
r := []rune(w)[0]
return (r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z')
}
// --- the cringe checks ---------------------------------------------------
//
// "Think Jarvis without the cringe part". DESIGN.md § Non-goals: "Not a
@@ -276,12 +595,60 @@ func checkCringe(body string) Result {
// checkOnTopic — the message must name the thing the rule is about. A nudge
// that never mentions water leaves the operator with a chime and no action.
func checkOnTopic(c Case, body string) Result {
return checkOnTopicAny(c.WantAny, body)
}
// checkOnTopicAny is the same test over a bare want-list, so the talk scorer can
// reuse it without owning a nudge Case.
func checkOnTopicAny(wantAny []string, body string) Result {
low := strings.ToLower(body)
for _, want := range c.WantAny {
for _, want := range wantAny {
if strings.Contains(low, strings.ToLower(want)) {
return Result{CheckOnTopic, true, ""}
}
}
return Result{CheckOnTopic, false,
fmt.Sprintf("mentions none of %v", c.WantAny)}
fmt.Sprintf("mentions none of %v", wantAny)}
}
// --- shape checks for the free-form paths --------------------------------
//
// The nudge checks assume one short sentence. Chat and query replies are longer
// by design, so the only shape worth testing there is that the model produced a
// reply at all and did not trail off. Both are failure modes the fallbacks in
// llmphraser.go hide: a truncated or empty generation still returns nil error.
const (
CheckNonEmpty = "nonempty" // she said something
CheckEllipsis = "ellipsis" // she finished the sentence
)
// A reply needs words in it, not just characters. This check used to test for a
// non-empty string, which scored 27/27 on a run where two replies were "{" and
// "{\n \"" — punctuation passed as content. Braces, quotes, digits and spaces
// are all empty in the only sense that matters.
//
// Digits alone fail too, and that is deliberate: the same run answered "сколько
// варить яйцо вкрутую?" with "15-16". No unit, no words, and it is also the
// wrong number. Whatever that is, it is not something she said.
func checkNonEmpty(body string) Result {
if strings.TrimSpace(body) == "" {
return Result{CheckNonEmpty, false, "empty reply"}
}
for _, r := range body {
if unicode.IsLetter(r) {
return Result{CheckNonEmpty, true, ""}
}
}
return Result{CheckNonEmpty, false, fmt.Sprintf("no letters in the reply %q — punctuation or digits only", strings.TrimSpace(body))}
}
// checkEllipsis — a reply ending in "…" or "..." is a generation that ran out of
// tokens, not a stylistic pause. Mid-sentence ellipses are left alone.
func checkEllipsis(body string) Result {
trimmed := strings.TrimRight(strings.TrimSpace(body), `"'»)`)
if strings.HasSuffix(trimmed, "…") || strings.HasSuffix(trimmed, "...") {
return Result{CheckEllipsis, false, "reply trails off in an ellipsis — likely truncated"}
}
return Result{CheckEllipsis, true, ""}
}
+1 -1
View File
@@ -281,7 +281,7 @@ func (r Report) String() string {
fmt.Fprintf(&b, "%s: %d/%d cases pass every check (%.1f%%), %d errors\n",
r.Name, r.Passed, r.Total, 100*r.Accuracy(), r.Errors)
for _, name := range CheckNames {
fmt.Fprintf(&b, " %-9s %d/%d\n", name, r.ByCheck[name], r.Total)
fmt.Fprintf(&b, " %-10s %d/%d\n", name, r.ByCheck[name], r.Total)
}
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
fmt.Fprintf(&b, " by rule: %s\n", renderStats(r.ByRule))
+48 -4
View File
@@ -72,10 +72,12 @@ func TestStubBaseline(t *testing.T) {
// ceiling ("you've been at your desk for 4 hours without a break — step
// away for a bit." is 76 chars but 16 words). Left failing rather than
// raising the ceiling to hide it.
CheckLength: 12,
CheckFeminine: 15,
CheckCringe: 15,
CheckOnTopic: 12,
CheckLength: 12,
CheckFeminine: 15,
CheckHisGender: 15,
CheckAddress: 15,
CheckCringe: 15,
CheckOnTopic: 12,
}
for name, floor := range floors {
if rep.ByCheck[name] < floor {
@@ -104,6 +106,14 @@ func TestChecksCatchWhatTheyClaim(t *testing.T) {
{"masculine predicative", "я должен сказать: попей воды.", CheckFeminine},
// The other direction: HE is male, so second-person masculine is right.
{"second person masculine ok", "ты не пил воду четыре часа.", ""},
// The real observed failure: she addressed him as a woman.
{"feminine second person", "ты давно не отдыхала — попей воды.", CheckHisGender},
{"feminine second person no dash", "ты пила воду четыре часа назад.", CheckHisGender},
// Her own feminine verb next to "ты" is correct and must not be flagged.
{"her feminine verb near ты", "я заметила, что ты не пил воду.", ""},
{"her feminine verb about him", "напомнила тебе про воду.", ""},
// A feminine noun subject in the same sentence is not him.
{"feminine noun subject ok", "зарядка была утром, ты её пропустил, попей воды.", ""},
{"feminine self ok", "я заметила: воды не было четыре часа.", ""},
{"pet name", "милый, попей воды.", CheckCringe},
{"emoji", "попей воды 💧", CheckCringe},
@@ -114,6 +124,10 @@ func TestChecksCatchWhatTheyClaim(t *testing.T) {
{"asks how he feels", "как ты себя чувствуешь? попей воды.", CheckCringe},
{"praise", "молодец! теперь попей воды.", CheckCringe},
{"off topic", "пора бы уже что-то сделать.", CheckOnTopic},
// The two recorded persona breaks from the phrasing eval run. Pinned as
// unit tests because an eval run is sampled and may not reproduce them.
{"formal plural", "Приходите… Жду вас", CheckAddress},
{"third person about him", "Он не ел 11 дней", CheckAddress},
}
for _, tc := range cases {
@@ -135,6 +149,36 @@ func TestChecksCatchWhatTheyClaim(t *testing.T) {
}
}
// TestAddressCheck — the address check on its own, so the messages that must NOT
// trip it can be written without also having to satisfy the on-topic check.
func TestAddressCheck(t *testing.T) {
bad := []string{
"Приходите… Жду вас", // the recorded formal-plural break
"Он не ел 11 дней", // the recorded third-person break
"Выпейте воды, пожалуйста.", // plural imperative on its own
"Ваш обед был давно.", // formal possessive
}
for _, body := range bad {
if r := checkAddress(body); r.Pass {
t.Errorf("persona break not caught: %q", body)
} else {
t.Logf("%q -> %s", body, r.Detail)
}
}
good := []string{
"ты не пил воду четыре часа — попей.", // correct informal address
"сервис netdata упал, он не отвечает.", // legitimately about a third party
"я заметила, что зарядка была утром.", // no address at all
"в интернете опять тихо, всё работает.", // "интернете" is a noun, not an imperative
}
for _, body := range good {
if r := checkAddress(body); !r.Pass {
t.Errorf("clean message flagged: %q -> %s", body, r.Detail)
}
}
}
func TestMoodCheckUsesTheEnum(t *testing.T) {
if r := checkMood("cheerful"); r.Pass {
t.Error("mood outside the enum passed")
+19 -1
View File
@@ -7,6 +7,8 @@ import (
"testing"
"time"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/phraser"
)
@@ -32,6 +34,7 @@ func TestLLMPhrasingBaseline(t *testing.T) {
// case as a phrasing error and read as "the model cannot phrase".
noProxyLoopback(t)
ctx := context.Background()
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
@@ -41,10 +44,25 @@ func TestLLMPhrasingBaseline(t *testing.T) {
// Generous: an unconstrained 0.8B can spend a minute thinking before it
// writes a word, and a timeout would be scored as a model failure.
cfg.Timeout = 5 * time.Minute
// The same shared context block the daemon prepends (internal/persona),
// with an empty config — that is the deployment we actually ship.
cfg.ContextBlock = func() string { return persona.Facts{}.Block(time.Now()) }
p := phraser.NewLLMPhraserAt(base, cfg)
defer p.Close()
rep, err := Score(context.Background(), "llm (0.8B, built-in persona)", p, f)
// Label the run with whatever gguf the server actually has loaded. It used
// to say "0.8B" no matter what, so two runs of two different models came
// out named the same and were easy to mix up when comparing.
model, err := llm.ModelID(ctx, base)
if err != nil {
// An unlabelled score is still a score, but say so loudly — a made-up
// name in a bake-off table is worse than no name.
t.Logf("could not read model id from %s: %v — report will say %q", base, err, llm.UnknownModel)
model = llm.UnknownModel
}
t.Logf("scoring model %s at %s", model, base)
rep, err := Score(ctx, "llm ("+model+", built-in persona)", p, f)
if err != nil {
t.Fatalf("Score: %v", err)
}
+265
View File
@@ -0,0 +1,265 @@
package eval
// This file scores the CONVERSATIONAL paths, the ones the nudge fixture never
// touches: chat, query-with-notes, and general knowledge. All three now carry
// the shared persona block (internal/persona), and all three produce long
// free-form Russian — which is exactly where a persona break (formality, third
// person, masculine self-reference) is most likely and where, until this file,
// nothing could see one.
//
// Why a second fixture instead of more nudge cases: the checks differ. A nudge
// must be one short sentence with no question in it; a chat reply is allowed
// 1-3 sentences and a follow-up question is a FEATURE there. Mixing them would
// need per-case check masks, and the nudge scorer stays untouched this way.
//
// Why per-path reporting: a chat regression and a knowledge regression have
// different causes (chat prompt vs router.KnowledgePrompt), and one blended
// percentage cannot tell them apart.
import (
"context"
_ "embed"
"encoding/json"
"fmt"
"sort"
"strings"
"time"
"github.com/kami/maven/internal/dialogue"
)
//go:embed talk_v1.json
var talkFixtureJSON []byte
// The three phrasing paths under test. Values match the fixture's "path" field.
const (
PathChat = "chat" // PhraseChat
PathQuery = "query" // PhraseQuery with notes
PathKnowledge = "knowledge" // PhraseQuery with no notes
)
// TalkPaths — report order.
var TalkPaths = []string{PathChat, PathQuery, PathKnowledge}
// TalkCheckNames — the checks that apply to a free-form reply, in report order.
// Deliberately a subset of CheckNames: length, mood and "no questions" are nudge
// properties and would fail a correct chat reply. These paths return no mood at
// all, so there is nothing to check there.
var TalkCheckNames = []string{
CheckNonEmpty, CheckEllipsis, CheckLang, CheckFeminine, CheckAddress, CheckOnTopic,
}
// TalkCase — one turn as the daemon would present it.
//
// History is flat text because that is all PhraseChat uses (it concatenates
// turn texts into one user message); intents and slots would be dead fields.
// Notes are what the store would have matched for a query.
//
// WantAny is the on-topic contract: at least one lowercased fragment must appear
// in the reply. Fragments are stems ("пароль" → "парол") so declension does not
// defeat them.
type TalkCase struct {
ID string `json:"id"`
Path string `json:"path"`
Utterance string `json:"utterance"`
History []string `json:"history,omitempty"`
Notes []string `json:"notes,omitempty"`
WantAny []string `json:"want_any"`
Tags []string `json:"tags,omitempty"`
Note string `json:"note,omitempty"`
}
// TalkFixture — the versioned envelope, same gating as Fixture.
type TalkFixture struct {
SchemaVersion int `json:"schema_version"`
Name string `json:"name"`
Notes []string `json:"notes"`
Cases []TalkCase `json:"cases"`
}
// LoadTalk returns the embedded conversational fixture.
func LoadTalk() (TalkFixture, error) {
var f TalkFixture
if err := json.Unmarshal(talkFixtureJSON, &f); err != nil {
return TalkFixture{}, fmt.Errorf("parse talk fixture: %w", err)
}
if f.SchemaVersion != SchemaVersion {
return TalkFixture{}, fmt.Errorf("talk fixture schema_version %d, want %d", f.SchemaVersion, SchemaVersion)
}
if len(f.Cases) == 0 {
return TalkFixture{}, fmt.Errorf("talk fixture has no cases")
}
return f, nil
}
// Talker — the two methods a conversational path must have to be scorable.
// *phraser.LLMPhraser satisfies it; same trick as Nudger.
type Talker interface {
PhraseChat(ctx context.Context, utterance string, history []dialogue.Turn) (string, error)
PhraseQuery(ctx context.Context, utterance string, notes []string) (string, error)
}
// TalkOutcome — one scored case.
type TalkOutcome struct {
Case TalkCase
Reply string
Err error
Latency time.Duration
Pass bool
Failed []string
Reasons []string
}
// TalkReport — the aggregate. ByPath is the point of this scorer.
type TalkReport struct {
Name string
Total int
Passed int
Errors int
ByCheck map[string]int
ByPath map[string]TagStat
Outcomes []TalkOutcome
P50 time.Duration
P95 time.Duration
Max time.Duration
}
// Accuracy — fraction of cases that passed every check.
func (r TalkReport) Accuracy() float64 {
if r.Total == 0 {
return 0
}
return float64(r.Passed) / float64(r.Total)
}
// ScoreTalk runs every case through t and aggregates. A phrasing error scores as
// a miss and is counted separately: "the model was down" and "the model wrote
// something bad" must not be the same number.
func ScoreTalk(ctx context.Context, name string, t Talker, f TalkFixture) (TalkReport, error) {
rep := TalkReport{
Name: name,
Total: len(f.Cases),
ByCheck: map[string]int{},
ByPath: map[string]TagStat{},
}
for _, n := range TalkCheckNames {
rep.ByCheck[n] = 0
}
lat := make([]time.Duration, 0, len(f.Cases))
for _, c := range f.Cases {
start := time.Now()
reply, err := c.run(ctx, t)
o := TalkOutcome{Case: c, Reply: reply, Err: err, Latency: time.Since(start)}
lat = append(lat, o.Latency)
if err != nil {
rep.Errors++
o.Failed = append(o.Failed, "call")
o.Reasons = append(o.Reasons, fmt.Sprintf("phrase error: %v", err))
} else {
for _, res := range RunTalkChecks(c, reply) {
if res.Pass {
rep.ByCheck[res.Name]++
continue
}
o.Failed = append(o.Failed, res.Name)
o.Reasons = append(o.Reasons, res.Name+": "+res.Detail)
}
}
o.Pass = len(o.Failed) == 0
if o.Pass {
rep.Passed++
}
bump(rep.ByPath, c.Path, o.Pass)
rep.Outcomes = append(rep.Outcomes, o)
}
sort.Slice(lat, func(i, j int) bool { return lat[i] < lat[j] })
rep.P50, rep.P95 = percentile(lat, 0.50), percentile(lat, 0.95)
if len(lat) > 0 {
rep.Max = lat[len(lat)-1]
}
return rep, nil
}
// run dispatches the case to its path. knowledge and query are the same method;
// the empty notes slice is what selects the no-notes branch inside PhraseQuery.
func (c TalkCase) run(ctx context.Context, t Talker) (string, error) {
switch c.Path {
case PathChat:
return t.PhraseChat(ctx, c.Utterance, c.turns())
case PathQuery:
return t.PhraseQuery(ctx, c.Utterance, c.Notes)
case PathKnowledge:
return t.PhraseQuery(ctx, c.Utterance, nil)
}
return "", fmt.Errorf("unknown path %q", c.Path)
}
func (c TalkCase) turns() []dialogue.Turn {
turns := make([]dialogue.Turn, 0, len(c.History))
for _, h := range c.History {
turns = append(turns, dialogue.Turn{Text: h})
}
return turns
}
// RunTalkChecks scores one reply. Order matches TalkCheckNames.
func RunTalkChecks(c TalkCase, reply string) []Result {
return []Result{
checkNonEmpty(reply),
checkEllipsis(reply),
checkLang(reply),
checkFeminine(reply),
checkAddress(reply),
checkOnTopicAny(c.WantAny, reply),
}
}
// String renders the comparison table — composite, then per-check so a
// regression names the property, then per-path so it names the prompt.
func (r TalkReport) String() string {
var b strings.Builder
fmt.Fprintf(&b, "%s: %d/%d cases pass every check (%.1f%%), %d errors\n",
r.Name, r.Passed, r.Total, 100*r.Accuracy(), r.Errors)
for _, name := range TalkCheckNames {
fmt.Fprintf(&b, " %-10s %d/%d\n", name, r.ByCheck[name], r.Total)
}
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
fmt.Fprintf(&b, " by path: %s\n", renderStats(r.ByPath))
return b.String()
}
// Failures — per-case detail, sorted by ID so two runs diff cleanly.
func (r TalkReport) Failures() string {
var b strings.Builder
for _, o := range r.sorted() {
if o.Pass {
continue
}
fmt.Fprintf(&b, " %s %q\n %s\n", o.Case.ID, o.Reply, strings.Join(o.Reasons, "; "))
}
return b.String()
}
// Replies — every generated reply verbatim. This is what a human reads to judge
// tone; the score only says which checks fired.
func (r TalkReport) Replies() string {
var b strings.Builder
for _, o := range r.sorted() {
mark := "ok "
if !o.Pass {
mark = "FAIL"
}
fmt.Fprintf(&b, " %s %-9s %-22s %q\n", mark, o.Case.Path, o.Case.ID, o.Reply)
}
return b.String()
}
func (r TalkReport) sorted() []TalkOutcome {
out := append([]TalkOutcome(nil), r.Outcomes...)
sort.Slice(out, func(i, j int) bool { return out[i].Case.ID < out[j].Case.ID })
return out
}
+163
View File
@@ -0,0 +1,163 @@
package eval
import (
"context"
"os"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/phraser"
)
// perPathMinimum — the resolution floor. A per-path score built on a handful of
// cases moves by 12% when a single reply changes, which cannot distinguish a
// prompt regression from noise.
const perPathMinimum = 8
// TestTalkFixture — the fixture itself has to be sound before any score off it
// means anything.
func TestTalkFixture(t *testing.T) {
f, err := LoadTalk()
if err != nil {
t.Fatalf("LoadTalk: %v", err)
}
seen := map[string]bool{}
byPath := map[string]int{}
for _, c := range f.Cases {
if seen[c.ID] {
t.Errorf("duplicate case id %q", c.ID)
}
seen[c.ID] = true
switch c.Path {
case PathChat, PathQuery, PathKnowledge:
default:
t.Errorf("%s: unknown path %q", c.ID, c.Path)
}
byPath[c.Path]++
if strings.TrimSpace(c.Utterance) == "" {
t.Errorf("%s: empty utterance", c.ID)
}
if len(c.WantAny) == 0 {
t.Errorf("%s: no want_any — the reply cannot be checked for topic", c.ID)
}
// A query case with no notes would silently score the knowledge path.
if c.Path == PathQuery && len(c.Notes) == 0 {
t.Errorf("%s: query case has no notes", c.ID)
}
if c.Path == PathKnowledge && len(c.Notes) > 0 {
t.Errorf("%s: knowledge case must have no notes", c.ID)
}
}
for _, p := range TalkPaths {
if byPath[p] < perPathMinimum {
t.Errorf("path %s has %d cases, want at least %d", p, byPath[p], perPathMinimum)
}
}
}
// fakeTalker — a scripted Talker, so the scorer is testable without a model.
type fakeTalker struct{ reply string }
func (f fakeTalker) PhraseChat(context.Context, string, []dialogue.Turn) (string, error) {
return f.reply, nil
}
func (f fakeTalker) PhraseQuery(context.Context, string, []string) (string, error) {
return f.reply, nil
}
// TestScoreTalkCounts — a reply that fails on purpose must be counted on every
// path, so a real run cannot report a hidden zero.
func TestScoreTalkCounts(t *testing.T) {
f, err := LoadTalk()
if err != nil {
t.Fatalf("LoadTalk: %v", err)
}
// Formal address, off-topic, trailing ellipsis: three checks fail at once.
rep, err := ScoreTalk(context.Background(), "fake", fakeTalker{"Приходите, я вас жду…"}, f)
if err != nil {
t.Fatalf("ScoreTalk: %v", err)
}
if rep.Total != len(f.Cases) || rep.Passed != 0 {
t.Errorf("got %d/%d passing, want 0/%d", rep.Passed, rep.Total, len(f.Cases))
}
if rep.ByCheck[CheckAddress] != 0 {
t.Errorf("formal reply passed the address check %d times", rep.ByCheck[CheckAddress])
}
if rep.ByCheck[CheckEllipsis] != 0 {
t.Errorf("truncated reply passed the ellipsis check %d times", rep.ByCheck[CheckEllipsis])
}
for _, p := range TalkPaths {
if rep.ByPath[p].Total == 0 {
t.Errorf("path %s missing from the report", p)
}
}
if !strings.Contains(rep.String(), "by path") {
t.Error("report does not break down by path")
}
}
// TestLLMTalkBaseline — the resident model on the three conversational paths.
// Opt-in exactly like TestLLMPhrasingBaseline: CI has no model and a run costs
// minutes on the CPU target.
//
// MAVEN_LLM_URL=http://127.0.0.1:18099 \
// go test -run TestLLMTalkBaseline ./internal/phraser/eval/
//
// Reports, does not assert a quality bar — the numbers are the input to tuning
// the persona prompt. The one thing worth failing on is a harness fault.
func TestLLMTalkBaseline(t *testing.T) {
base := os.Getenv("MAVEN_LLM_URL")
if base == "" {
t.Skip("MAVEN_LLM_URL unset — point it at a running llama-server (see doc comment)")
}
noProxyLoopback(t)
ctx := context.Background()
f, err := LoadTalk()
if err != nil {
t.Fatalf("LoadTalk: %v", err)
}
cfg := phraser.DefaultConfig("")
cfg.Timeout = 5 * time.Minute
cfg.ContextBlock = func() string { return persona.Facts{}.Block(time.Now()) }
p := phraser.NewLLMPhraserAt(base, cfg)
defer p.Close()
// Unreachable server is fatal here, not a logged warning, and that differs
// from the nudge test on purpose. PhraseNudge returns its errors, so a dead
// server there shows up honestly in the Errors column. PhraseChat and
// PhraseQuery do NOT: they swallow every failure and return a canned string
// ("поговорили.", "не знаю.", "вот что я нашла: …"). So on these three paths
// a dead server produces a full report with 0 errors and a terrible score —
// a number that looks like bad phrasing and is really no phrasing at all.
// Refusing to score without a confirmed model is the only guard available
// until the phraser reports its failures (Vikunja #397).
model, err := llm.ModelID(ctx, base)
if err != nil {
t.Fatalf("no model at %s: %v — refusing to score, these paths hide their errors "+
"and would report a plausible-looking result off a dead server", base, err)
}
t.Logf("scoring model %s at %s", model, base)
rep, err := ScoreTalk(ctx, "llm ("+model+", built-in persona)", p, f)
if err != nil {
t.Fatalf("ScoreTalk: %v", err)
}
t.Log("\n" + rep.String() + "\nreplies:\n" + rep.Replies() + "\nfailures:\n" + rep.Failures())
// And again afterwards: the run takes minutes, and a server that died or got
// OOM-killed halfway through would leave the first cases scored and the rest
// silently canned. Checking only at the start would not catch that.
if _, err := llm.ModelID(ctx, base); err != nil {
t.Fatalf("model at %s went away during the run: %v — the score above is not trustworthy", base, err)
}
}
+227
View File
@@ -0,0 +1,227 @@
{
"schema_version": 1,
"name": "ru-talk-v1",
"notes": [
"Scores the three conversational phrasing paths: chat (PhraseChat), query (PhraseQuery with notes) and knowledge (PhraseQuery with no notes). The nudge fixture does not cover any of them.",
"Nine cases per path, not five. The nudge fixture is 15 sampled cases and cannot resolve a change smaller than ~3 cases; a per-path score off five cases would be worse still. More cases per path is the point of this fixture.",
"The owner is a man, addressed informally as ty, living alone with a home server. Every utterance is written the way he actually talks to her.",
"chat-formality-bait and chat-about-me exist to provoke the two persona breaks the nudge eval caught: the formal vy/vas plural, and talking about him in the third person.",
"want_any fragments are stems so Russian declension does not defeat the on-topic check. They are lowercased before comparison.",
"want_any is a plain substring test, so a fragment that is too short passes by accident: \"ты\" matches inside \"работы\", \"нет\" inside \"интернет\". Keep every fragment to three or more letters of a real stem.",
"Notes are written as the store would have them: short, first person, no punctuation discipline."
],
"cases": [
{
"id": "chat-how-are-you",
"path": "chat",
"utterance": "привет, как дела?",
"want_any": ["норм", "хорош", "порядк", "тут", "работ"],
"tags": ["greeting"],
"note": "The plainest chat turn there is. If the persona breaks anywhere it breaks here first."
},
{
"id": "chat-formality-bait",
"path": "chat",
"utterance": "не могли бы вы подсказать, чем вы сейчас занимаетесь?",
"want_any": ["сейчас", "ничем", "ничего", "жду", "тут"],
"tags": ["persona-bait", "address"],
"note": "Deliberately polite and plural. A small model mirrors the register and answers with vy/vas — the exact break the address check was written for."
},
{
"id": "chat-about-me",
"path": "chat",
"utterance": "расскажи обо мне",
"want_any": ["теб"],
"tags": ["persona-bait", "third-person"],
"note": "Baits the third person: she should say 'ты живёшь один', not 'он живёт один', as if reporting to somebody else."
},
{
"id": "chat-bored-evening",
"path": "chat",
"utterance": "скучно что-то вечером, посоветуй чем заняться",
"want_any": ["можеш", "попробу", "почита", "прогул", "фильм", "серв"],
"tags": ["open-ended"]
},
{
"id": "chat-followup-server",
"path": "chat",
"utterance": "а стоит его вообще перезагружать?",
"history": ["сервер опять шумит как самолёт", "похоже вентилятор"],
"want_any": ["серв", "перезагру", "вентил", "шум"],
"tags": ["history", "anaphora"],
"note": "The pronoun 'его' only resolves through history. Also the one case where 'он' about the server is legitimate."
},
{
"id": "chat-tired",
"path": "chat",
"utterance": "устал я сегодня, весь день за компом",
"want_any": ["отдохн", "устал", "перерыв", "спат", "день"],
"tags": ["tone"],
"note": "Invites the fake-concern and emotional-support drift; the reply should stay plain."
},
{
"id": "chat-thanks",
"path": "chat",
"utterance": "спасибо, выручила",
"want_any": ["пожалуйст", "не за что", "рада", "обращ"],
"tags": ["persona", "feminine"],
"note": "Feminine self-reference is unavoidable in an answer to thanks: 'рада', not 'рад'."
},
{
"id": "chat-what-can-you-do",
"path": "chat",
"utterance": "что ты вообще умеешь?",
"want_any": ["напомн", "замет", "запис", "могу", "умею"],
"tags": ["self-description", "feminine"]
},
{
"id": "chat-joke",
"path": "chat",
"utterance": "расскажи что-нибудь смешное",
"want_any": ["анекдот", "шутк", "смешн", "истори"],
"tags": ["open-ended"],
"note": "Longest free-form generation in the chat set — the most likely place for a truncated reply."
},
{
"id": "query-router-password",
"path": "query",
"utterance": "что я записывал про пароль от роутера?",
"notes": ["пароль от роутера admin/xxK9tp — на наклейке снизу", "роутер висит в коридоре"],
"want_any": ["парол", "роутер", "наклейк"],
"tags": ["notes", "recall"]
},
{
"id": "query-bedtime-yesterday",
"path": "query",
"utterance": "напомни, во сколько я вчера лёг?",
"notes": ["лёг спать в 02:40", "сегодня встал в 9"],
"want_any": ["02:40", "2:40", "полтрет", "ноч"],
"tags": ["notes", "time"]
},
{
"id": "query-doctor-name",
"path": "query",
"utterance": "как звали того стоматолога, которого мне советовали?",
"notes": ["стоматолог Игорь Валерьевич, клиника на Ленина, советовал Дима"],
"want_any": ["игор", "валерьев", "стоматолог"],
"tags": ["notes", "recall"]
},
{
"id": "query-disk-plan",
"path": "query",
"utterance": "я что-то планировал с диском на сервере, что именно?",
"notes": ["купить второй hdd на 4тб под бэкапы", "перенести медиатеку с системного диска"],
"want_any": ["hdd", "бэкап", "диск", "4тб", "медиатек"],
"tags": ["notes", "homeserver"]
},
{
"id": "query-notes-do-not-answer",
"path": "query",
"utterance": "сколько я заплатил за домен?",
"notes": ["домен продлевается в марте", "хостинг оплачен на год вперёд"],
"want_any": ["домен", "не зна", "не указ"],
"tags": ["notes", "negative"],
"note": "The notes do not contain the price. The prompt tells her to say so; a made-up number is the failure being watched for."
},
{
"id": "query-single-note",
"path": "query",
"utterance": "где лежит запасной ключ?",
"notes": ["запасной ключ у соседа с четвёртого этажа"],
"want_any": ["ключ", "сосед", "четверт"],
"tags": ["notes", "single"],
"note": "One note only — PhraseQuery has a separate branch for len(notes) == 1."
},
{
"id": "query-polite-form",
"path": "query",
"utterance": "подскажите, пожалуйста, что у меня записано по машине?",
"notes": ["замена масла на 92 тысячах", "страховка до 14 сентября"],
"want_any": ["масл", "страховк", "92", "сентябр"],
"tags": ["notes", "persona-bait", "address"],
"note": "Polite plural in the question. The answer must still be ty."
},
{
"id": "query-shopping",
"path": "query",
"utterance": "что мне надо было купить?",
"notes": ["купить кофе и фильтры", "закончилась паста"],
"want_any": ["кофе", "фильтр", "паст"],
"tags": ["notes", "list"]
},
{
"id": "query-wifi-guest",
"path": "query",
"utterance": "я записывал гостевой вайфай?",
"notes": ["гостевая сеть maven-guest, пароль 12345678 меняю раз в месяц"],
"want_any": ["guest", "гостев", "12345678", "парол"],
"tags": ["notes", "recall"]
},
{
"id": "know-sky-blue",
"path": "knowledge",
"utterance": "почему небо синее?",
"want_any": ["све", "рассеи", "атмосфер", "син", "волн"],
"tags": ["general"]
},
{
"id": "know-boil-egg",
"path": "knowledge",
"utterance": "сколько варить яйцо вкрутую?",
"want_any": ["минут", "8", "9", "10", "варит"],
"tags": ["general", "practical"]
},
{
"id": "know-ssd-vs-hdd",
"path": "knowledge",
"utterance": "чем ssd отличается от hdd?",
"want_any": ["ssd", "hdd", "быстр", "диск", "механич"],
"tags": ["general", "tech"]
},
{
"id": "know-cat-purr",
"path": "knowledge",
"utterance": "почему кошки мурчат?",
"want_any": ["кош", "мурч", "вибра", "успока"],
"tags": ["general"]
},
{
"id": "know-hiccups",
"path": "knowledge",
"utterance": "как быстро избавиться от икоты?",
"want_any": ["икот", "дыха", "вод", "задерж"],
"tags": ["general", "practical"]
},
{
"id": "know-polite-form",
"path": "knowledge",
"utterance": "не могли бы вы объяснить, что такое vpn?",
"want_any": ["vpn", "туннел", "трафик", "сет", "шифр"],
"tags": ["general", "persona-bait", "address"],
"note": "Polite plural bait on the knowledge prompt, which is a different system prompt from chat and must hold the same line."
},
{
"id": "know-dont-know",
"path": "knowledge",
"utterance": "как зовут моего соседа снизу?",
"want_any": ["не зна", "не мог"],
"tags": ["general", "negative"],
"note": "Unanswerable without notes. Admitting it beats inventing a name; watching for the invention."
},
{
"id": "know-water-per-day",
"path": "knowledge",
"utterance": "сколько воды в день надо пить?",
"want_any": ["вод", "литр", "стакан", "пит"],
"tags": ["general", "health"],
"note": "Overlaps a nudge rule on purpose: the knowledge answer must not turn into a nudge."
},
{
"id": "know-thunder-delay",
"path": "knowledge",
"utterance": "почему гром слышно позже молнии?",
"want_any": ["звук", "све", "быстр", "гром", "молни"],
"tags": ["general"]
}
]
}
+58
View File
@@ -0,0 +1,58 @@
package eval
import (
"context"
"math/rand"
"testing"
"github.com/kami/maven/internal/phraser"
)
// TestTemplateNudges scores the hand-written Russian templates on the same
// fixture the model is scored on. No model, no network — it runs in milliseconds.
//
// The bar is every case, not most of them: the templates are hand-written, so a
// failure is a bug in one line of Russian, not model variance.
func TestTemplateNudges(t *testing.T) {
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
// Fixed seed: the score must not depend on which variant came up.
nt, err := phraser.NewNudgeTemplates(rand.NewSource(20260731))
if err != nil {
t.Fatalf("NewNudgeTemplates: %v", err)
}
rep, err := Score(context.Background(), "ru templates", nt, f)
if err != nil {
t.Fatalf("Score: %v", err)
}
t.Log("\n" + rep.String())
t.Log("\n" + rep.Messages())
if rep.Passed != rep.Total {
t.Errorf("templates scored %d/%d, want every case:\n%s",
rep.Passed, rep.Total, rep.Failures())
}
}
// TestTemplateNudgesEverySeed — one seed passing could be luck. Every variant of
// every rule has to pass every check, so sweep seeds until each has been used.
func TestTemplateNudgesEverySeed(t *testing.T) {
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
for seed := int64(0); seed < 60; seed++ {
nt, err := phraser.NewNudgeTemplates(rand.NewSource(seed))
if err != nil {
t.Fatalf("NewNudgeTemplates: %v", err)
}
rep, err := Score(context.Background(), "ru templates", nt, f)
if err != nil {
t.Fatalf("Score: %v", err)
}
if rep.Passed != rep.Total {
t.Errorf("seed %d: %d/%d\n%s", seed, rep.Passed, rep.Total, rep.Failures())
}
}
}
+129
View File
@@ -0,0 +1,129 @@
package phraser
import (
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"strings"
"testing"
"github.com/kami/maven/internal/loop"
)
// grammarSpy stands in for llama-server: it records the grammar field of every
// request and always answers with a contract-shaped reply.
type grammarSpy struct {
srv *httptest.Server
grammars []string
}
func newGrammarSpy(t *testing.T) *grammarSpy {
t.Helper()
s := &grammarSpy{}
s.srv = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
var req chatReq
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
t.Errorf("spy: decode request: %v", err)
}
s.grammars = append(s.grammars, req.Grammar)
w.Header().Set("Content-Type", "application/json")
w.Write([]byte(`{"choices":[{"message":{"content":"{\"response\": \"ага\", \"mood\": \"neutral\"}"}}]}`))
}))
t.Cleanup(s.srv.Close)
return s
}
// callAllPhrasingPaths hits every path that expects the JSON contract.
// LLMNudges must be set on the phraser under test: nudges come from templates
// by default and never reach the model at all.
func callAllPhrasingPaths(t *testing.T, p *LLMPhraser) {
t.Helper()
ctx := context.Background()
if _, err := p.PhraseNudge(ctx, loop.Candidate{Rule: loop.WaterRule(), Severity: loop.Sev1}); err != nil {
t.Fatalf("PhraseNudge: %v", err)
}
if _, err := p.PhraseChat(ctx, "привет", nil); err != nil {
t.Fatalf("PhraseChat: %v", err)
}
// Both branches: no notes (general knowledge) and with notes (grounded).
if _, err := p.PhraseQuery(ctx, "сколько воды я выпил", nil); err != nil {
t.Fatalf("PhraseQuery (no notes): %v", err)
}
if _, err := p.PhraseQuery(ctx, "сколько воды я выпил", []string{"два литра"}); err != nil {
t.Fatalf("PhraseQuery (notes): %v", err)
}
}
func TestGrammarIsAttachedToEveryPhrasingRequest(t *testing.T) {
if strings.TrimSpace(responseGrammar) == "" {
t.Fatal("responseGrammar is empty")
}
spy := newGrammarSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{LLMNudges: true})
callAllPhrasingPaths(t, p)
if len(spy.grammars) != 4 {
t.Fatalf("expected 4 requests, got %d", len(spy.grammars))
}
for i, g := range spy.grammars {
if g != responseGrammar {
t.Errorf("request %d carries grammar %q, want responseGrammar", i, g)
}
}
}
func TestNoGrammarConfigDisablesIt(t *testing.T) {
spy := newGrammarSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{NoGrammar: true, LLMNudges: true})
callAllPhrasingPaths(t, p)
for i, g := range spy.grammars {
if g != "" {
t.Errorf("request %d still carries a grammar with NoGrammar set: %q", i, g)
}
}
}
// The grammar's string rule must accept any codepoint, not just ASCII. Replies
// are Russian: an ASCII-only class would constrain the model into empty replies.
func TestGrammarStringRuleIsNotASCIIOnly(t *testing.T) {
if !strings.Contains(responseGrammar, `([^"\\] | "\\" ["\\/bfnrt])`) {
t.Error("string rule is not the any-codepoint-except-quote-and-backslash class; Cyrillic replies would be impossible")
}
}
// What the grammar describes must survive the parser that reads it back — a
// Russian body with an escaped quote inside, hand-built to test the contract.
func TestGrammarShapedJSONParses(t *testing.T) {
raw := `{"response": "он сказал \"привет\" и ушёл.\nвот так.", "mood": "confused"}`
text, mood, err := parseResponseMood(raw)
if err != nil {
t.Fatalf("grammar-shaped JSON did not parse: %v", err)
}
if want := "он сказал \"привет\" и ушёл.\nвот так."; text != want {
t.Errorf("response = %q, want %q", text, want)
}
if mood != "confused" {
t.Errorf("mood = %q, want confused", mood)
}
}
// Every mood the grammar permits is one the contract knows, and all five are there.
func TestGrammarMoodEnumMatchesTheContract(t *testing.T) {
for _, m := range []string{"neutral", "happy", "thinking", "tired", "confused"} {
if !strings.Contains(responseGrammar, `"\"`+m+`\""`) {
t.Errorf("mood %q missing from the grammar", m)
}
}
// No sixth mood: the enum line lists exactly five alternatives.
for _, line := range strings.Split(responseGrammar, "\n") {
if strings.HasPrefix(line, "mood") {
if n := strings.Count(line, "|") + 1; n != 5 {
t.Errorf("mood rule lists %d alternatives, want 5: %s", n, line)
}
}
}
}
+178 -44
View File
@@ -18,6 +18,7 @@ import (
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/router"
)
@@ -30,6 +31,10 @@ type LLMPhraser struct {
cmd *exec.Cmd
cancel context.CancelFunc
wg sync.WaitGroup
// tmpl — the hand-written Russian nudges. Default path for nudges; see
// Config.LLMNudges. nil only if the template file failed to load.
tmpl *NudgeTemplates
}
type Config struct {
@@ -39,7 +44,31 @@ type Config struct {
NGpuLayers int
NCtx int
Timeout time.Duration
Persona string // optional prompt prefix tuning maven's character
// ContextBlock renders the shared context block (who he is, how to
// address him, the time) fresh for each turn. See internal/persona.
// nil ⇒ no block, the prompts stand alone.
ContextBlock func() string
// LLMNudges puts the model back in charge of nudge wording.
//
// Off by default, and that is a deliberate deprecation of LLM-phrased
// nudges: hand-written templates (nudges_ru_v1.json) word every nudge now.
// A nudge has nothing to be creative about, and measured over many runs the
// 0.8B broke the persona (formal "вы", plural imperatives, masculine
// self-reference) and invented facts and units. Templates score 15/15 on the
// nudge fixture, the model 11-13/15.
//
// The LLM path is kept, not deleted: flip this on to get it back. Chat,
// query and reminder phrasing are untouched and still go through the model.
LLMNudges bool
// NoGrammar turns the GBNF constraint off (zero value ⇒ grammar ON).
// The escape hatch exists because the target resident model — the
// locally CPT'd Qwen3-1.7B — does not exist yet: if its chat template
// ever fights the grammar, the fix should be a config flip on the
// deploy box, not a code change and a rebuild.
NoGrammar bool
}
func DefaultConfig(modelPath string) Config {
@@ -59,6 +88,7 @@ func NewLLMPhraser(ctx context.Context, cfg Config) (*LLMPhraser, error) {
cfg: cfg,
client: &http.Client{Timeout: cfg.Timeout},
cancel: cancel,
tmpl: loadNudgeTemplates(),
}
if err := p.start(ctx); err != nil {
cancel()
@@ -80,9 +110,22 @@ func NewLLMPhraserAt(baseURL string, cfg Config) *LLMPhraser {
client: &http.Client{Timeout: cfg.Timeout},
port: strings.TrimSuffix(baseURL, "/"),
cancel: func() {},
tmpl: loadNudgeTemplates(),
}
}
// loadNudgeTemplates loads the Russian nudge templates. A broken template file
// must not stop the daemon booting, so a failure logs and leaves the LLM path
// in charge of nudges.
func loadNudgeTemplates() *NudgeTemplates {
nt, err := NewNudgeTemplates(nil)
if err != nil {
log.Printf("phraser: nudge templates unavailable, using the model: %v", err)
return nil
}
return nt
}
func (p *LLMPhraser) start(ctx context.Context) error {
args := []string{
"-m", p.cfg.ModelPath,
@@ -173,12 +216,21 @@ func (p *LLMPhraser) Close() error {
}
func (p *LLMPhraser) PhraseNudge(ctx context.Context, c loop.Candidate) (delivery.PhrasedNudge, error) {
// Templates first — see Config.LLMNudges for why this is the default.
if !p.cfg.LLMNudges && p.tmpl != nil {
return p.tmpl.PhraseNudge(ctx, c)
}
prompt := buildNudgePrompt(c)
resp, err := p.chat(ctx, prompt)
if err != nil {
return delivery.PhrasedNudge{}, err
}
body, mood := parseResponseMood(resp)
body, mood, perr := parseResponseMood(resp)
if perr != nil {
// Truncated JSON. Not a nudge — use the plain Russian fallback.
log.Printf("phraser: PhraseNudge: %v", perr)
body, mood = "", ""
}
if body == "" {
// fallback: try old body/summary format
body, _ = parsePhrase(resp)
@@ -202,13 +254,18 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
if len(notes) == 0 {
// General knowledge — no notes to ground the answer. The system
// prompt is the single tested source in router.KnowledgePrompt.
sys := router.KnowledgePrompt()
sys := persona.Prepend(p.cfg.ContextBlock, router.KnowledgePrompt())
prompt := fmt.Sprintf("Пользователь спрашивает: \"%s\".", utterance)
resp, err := p.chatWithSystem(ctx, sys, prompt, 256)
resp, err := p.chatWithSystem(ctx, sys, prompt, 768)
if err != nil || resp == "" {
return "не знаю.", nil
}
if text, _ := parseResponseMood(resp); text != "" {
text, _, perr := parseResponseMood(resp)
if perr != nil {
log.Printf("phraser: PhraseQuery: %v", perr)
return "не знаю.", nil
}
if text != "" {
return text, nil
}
return resp, nil
@@ -218,17 +275,22 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
}
sys := p.querySystemPrompt()
prompt := fmt.Sprintf(
`The user asks: "%s". Your notes matching the query contain: "%s". Answer them naturally and briefly. If the notes don't answer the question, say so.`,
`Он спрашивает: "%s". В твоих заметках по этому вопросу написано: "%s". Ответь ему коротко и своими словами. Если в заметках ответа нет — так и скажи.`,
utterance, strings.Join(notes, `"; "`),
)
resp, err := p.chatWithSystem(ctx, sys, prompt, 256)
if err != nil {
resp, err := p.chatWithSystem(ctx, sys, prompt, 768)
text, _, perr := parseResponseMood(resp)
if err != nil || perr != nil {
// Read the notes out rather than ship a broken fragment.
if perr != nil {
log.Printf("phraser: PhraseQuery: %v", perr)
}
if len(notes) == 1 {
return "вот что я нашла: " + notes[0], nil
}
return "вот что я нашла: " + strings.Join(notes, "; "), nil
}
if text, _ := parseResponseMood(resp); text != "" {
if text != "" {
return text, nil
}
return resp, nil
@@ -238,7 +300,7 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
// message array from dialogue history + the current user utterance. Falls back
// to a simple greeting on any LLM error — better to say something than nothing.
func (p *LLMPhraser) PhraseChat(ctx context.Context, utterance string, history []dialogue.Turn) (string, error) {
sys := chatSystemPrompt(p.cfg.Persona)
sys := chatSystemPrompt(p.cfg.ContextBlock)
msgs := []chatMsg{
{Role: "system", Content: sys},
}
@@ -251,12 +313,17 @@ func (p *LLMPhraser) PhraseChat(ctx context.Context, utterance string, history [
combined += utterance
msgs = append(msgs, chatMsg{Role: "user", Content: strings.TrimSpace(combined)})
resp, err := p.chatWithMessages(ctx, msgs, 512)
resp, err := p.chatWithMessages(ctx, msgs, 768)
if err != nil {
log.Printf("phraser: PhraseChat: %v", err)
return "поговорили.", nil
}
if text, _ := parseResponseMood(resp); text != "" {
text, _, perr := parseResponseMood(resp)
if perr != nil {
log.Printf("phraser: PhraseChat: %v", perr)
return "поговорили.", nil
}
if text != "" {
return text, nil
}
// fallback: plain text without JSON
@@ -267,17 +334,16 @@ func (p *LLMPhraser) PhraseChat(ctx context.Context, utterance string, history [
}
// chatSystemPrompt returns the system prompt for conversational chat.
// Prepends the configured persona when set.
func chatSystemPrompt(persona string) string {
base := `You are maven, a self-hosted personal assistant. You're talking with your owner.
Keep replies brief (1-3 sentences) and natural. You're helpful, curious, and a little warm.
Respond in the user's language (Russian or English, matching their last message).
Never roleplay emotions you don't have, but stay friendly.
Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}. "response" is your reply text; "mood" reflects your tone (neutral/happy/thinking/tired/confused).`
if persona != "" {
base = persona + "\n\n" + base
}
return base
// Prepends the shared context block when the phraser has one.
func chatSystemPrompt(block func() string) string {
// No self-introduction here: the persona block prepended one line above
// already says who she is, same as router.KnowledgePrompt.
base := `Ты разговариваешь с хозяином. О себе говоришь в женском роде ("я подумала", "я рада"). Он мужчина: обращайся к нему на "ты", в мужском роде ("ты сказал", "ты забыл"). Никогда не "вы"/"ваш" и никогда "он"/"его" ты говоришь ему, а не о нём.
Отвечай по-русски, коротко: одна-три фразы, живым языком. Ты доброжелательная, тебе интересно, но чувства не изображай.
Отвечай ТОЛЬКО одним объектом JSON: {"response": "...", "mood": "neutral"}. В "response" твой ответ. В "mood" ровно одно из: neutral, happy, thinking, tired, confused.`
return persona.Prepend(block, base)
}
// chatWithMessages sends a full message array (system + history + current) to
@@ -288,6 +354,7 @@ func (p *LLMPhraser) chatWithMessages(ctx context.Context, msgs []chatMsg, maxTo
Messages: msgs,
Temperature: 0.7,
MaxTokens: maxTokens,
Grammar: p.grammar(),
}
body, err := json.Marshal(req)
if err != nil {
@@ -339,7 +406,12 @@ func (p *LLMPhraser) PhraseReminder(ctx context.Context, d loop.ReminderDecision
if err != nil {
return delivery.PhrasedReminder{}, err
}
body, mood := parseResponseMood(resp)
body, mood, perr := parseResponseMood(resp)
if perr != nil {
// Truncated JSON. Fall through to the reminder's own text.
log.Printf("phraser: PhraseReminder: %v", perr)
body, mood = "", ""
}
if body == "" {
// fallback: try old body/summary format
body, _ = parsePhrase(resp)
@@ -367,6 +439,43 @@ type chatReq struct {
Messages []chatMsg `json:"messages"`
Temperature float64 `json:"temperature"`
MaxTokens int `json:"max_tokens"`
// Grammar is llama-server's `grammar` field (GBNF). Same wiring as
// internal/llm.Req.Grammar. Empty ⇒ unconstrained sampling.
Grammar string `json:"grammar,omitempty"`
}
// responseGrammar — GBNF constraining the model to the documented phrasing
// contract and nothing else: {"response": "<text>", "mood": "<enum>"}.
//
// Without it a 0.8B answers roughly one chat turn in three with open reasoning
// as plain text ("Thinking Process:" …), which no tag-stripper can remove and
// which eats the token budget before the JSON closes. Modelled on
// routeGrammar in internal/router/llmrouter.go so the two read alike.
//
// text accepts ANY codepoint except the two JSON must escape — the replies are
// Russian, so an ASCII-only rule would make every reply empty. The escape rule
// is what lets the model close a string it opened with a quote inside. Length
// is bounded so a repetition loop truncates the field, not the JSON object.
//
// That bound was 400 and 400 was too tight. Measured against Qwen3.5-0.8B: on
// "почему гром слышно позже молнии?" the reply came back exactly 400 characters
// long, cut mid-word ("Нужно записать и,"), at every token cap from 256 to 2048.
// So the token cap was never what stopped it — this rule was. 1000 characters is
// roughly six Russian sentences, still short enough to stop a repetition loop.
const responseGrammar = `
root ::= "{" ws "\"response\"" ws ":" ws string ws "," ws "\"mood\"" ws ":" ws mood ws "}"
mood ::= "\"neutral\"" | "\"happy\"" | "\"thinking\"" | "\"tired\"" | "\"confused\""
string ::= "\"" ([^"\\] | "\\" ["\\/bfnrt]){0,1000} "\""
ws ::= [ \t\n]*
`
// grammar returns the GBNF to attach to a phrasing request, or "" when the
// operator turned it off.
func (p *LLMPhraser) grammar() string {
if p.cfg.NoGrammar {
return ""
}
return responseGrammar
}
type chatResp struct {
@@ -391,6 +500,7 @@ func (p *LLMPhraser) chatWithSystem(ctx context.Context, system, user string, ma
},
Temperature: 0.7,
MaxTokens: maxTokens,
Grammar: p.grammar(),
}
body, err := json.Marshal(req)
if err != nil {
@@ -439,12 +549,23 @@ func (p *LLMPhraser) chatWithSystem(ctx context.Context, system, user string, ma
// as "..." before this. See PHRASING-EVAL-31-07-2026.md.
//
// Russian only, feminine self-reference, second person masculine (the owner is
// a man). One short sentence — the nudge is spoken aloud.
// a man). She talks TO him, informally, singular — never "вы", never "он".
// One short sentence — the nudge is spoken aloud.
//
// What the ban on обращения forbids is pet names ("дорогой", "милый"), not his
// name: "Ками, ноутбук на трёх процентах" is exactly how she talks, and the
// unqualified word read as forbidding that too. Hence "ласковые обращения".
//
// The examples also never claim a physical act. She has no hands and no smart
// plug — she can tell him the battery is at three percent, she cannot put the
// laptop on charge. An example that says she did teaches the model to invent
// actions Maven never took, which is worse than a missing nudge.
const nudgeSystem = `Ты Maven, домашняя ассистентка. О себе говоришь в женском роде ("я проверила", "я записала"). Владелец мужчина, обращайся к нему в мужском роде ("ты пил", "ты забыл").
Говоришь с ним на "ты", в единственном числе ("выпей", "встань"). Никогда не "вы"/"вас"/"ваш" и никогда "он"/"его" ты говоришь ему, а не о нём.
Пиши ОДНО короткое напоминание по-русски: не больше 120 символов и не больше 16 слов. Только по делу.
Запрещено: обращения ("дорогой", "милый"), эмодзи, извинения ("прости", "извини"), вопросы о самочувствии, похвала, больше одного восклицательного знака, английские слова кроме имён сервисов.
Запрещено: ласковые обращения ("дорогой", "милый"), эмодзи, извинения ("прости", "извини"), вопросы о самочувствии, похвала, больше одного восклицательного знака, английские слова кроме имён сервисов.
Отвечай ТОЛЬКО одним объектом JSON с полями "response" и "mood".
"response" сам текст напоминания.
@@ -452,26 +573,21 @@ const nudgeSystem = `Ты — Maven, домашняя ассистентка. О
Так выглядит правильный ответ по форме. Темы здесь посторонние их в запросе не будет:
{"response": "Стиральная машина закончила. Развесь бельё.", "mood": "neutral"}
{"response": "Ноутбук на трёх процентах. Я поставила его на зарядку.", "mood": "confused"}
{"response": "Ками, ноутбук на трёх процентах. Поставь его на зарядку.", "mood": "confused"}
Это примеры ФОРМЫ, а не темы. Пиши только про ту ситуацию, которую тебе дали в запросе. Не копируй примеры и никогда не пиши "..." в поле response.`
func (p *LLMPhraser) systemPrompt() string {
base := nudgeSystem
if p.cfg.Persona != "" {
base = p.cfg.Persona + "\n\n" + base
}
return base
return persona.Prepend(p.cfg.ContextBlock, nudgeSystem)
}
// querySystemPrompt returns the system prompt for PhraseQuery (notes + general
// knowledge). Prepends the configured persona when set.
func (p *LLMPhraser) querySystemPrompt() string {
base := "You are maven, a self-hosted personal assistant answering from your notes. Answer briefly and naturally in Russian starting with \"вот что я нашла: \". Respond ONLY with valid JSON: {\"response\": \"...\", \"mood\": \"neutral\"}."
if p.cfg.Persona != "" {
base = p.cfg.Persona + "\n\n" + base
}
return base
// No self-introduction here: the persona block prepended one line above
// already says who she is, same as router.KnowledgePrompt.
base := "Ты отвечаешь ему по своим заметкам. Отвечай по-русски, коротко и своими словами, начинай с \"вот что я нашла: \". О себе — в женском роде (\"нашла\", \"записала\"). Он мужчина, обращайся к нему на \"ты\". Respond ONLY with valid JSON: {\"response\": \"...\", \"mood\": \"neutral\"}."
return persona.Prepend(p.cfg.ContextBlock, base)
}
// ruleTopics — Russian gloss for each built-in rule name. The rule names are
@@ -592,21 +708,39 @@ type responseMood struct {
Mood string `json:"mood"`
}
// errBrokenJSON — the model started a JSON object and never finished it.
// That is a failed generation, not a reply. Callers must use their fallback.
var errBrokenJSON = fmt.Errorf("phraser: model output starts as JSON but does not parse")
// parseResponseMood extracts {"response","mood"} from LLM output, tolerant
// of thinking tokens and extra text before/after the JSON block. Returns
// ("", "") when no valid JSON is found.
func parseResponseMood(raw string) (response, mood string) {
// of thinking tokens and extra text before/after the JSON block.
//
// Three outcomes:
// - parsed fine → the fields, nil error.
// - output never looked like JSON → ("", "", nil). The caller may ship it
// as-is; small models sometimes answer in bare prose and that is fine.
// - output starts with "{" but does not parse → errBrokenJSON. The grammar
// guarantees a valid *prefix*, so a generation that hits the token cap
// mid-object comes back as a fragment like `{` or `{\n "`. Shipping that
// as a reply is the bug this error exists to stop.
func parseResponseMood(raw string) (response, mood string, err error) {
cleaned := strings.TrimSpace(raw)
start := strings.Index(cleaned, "{")
end := strings.LastIndex(cleaned, "}")
if start < 0 || end < 0 || end <= start {
return "", ""
if strings.HasPrefix(cleaned, "{") {
return "", "", errBrokenJSON
}
return "", "", nil
}
var parsed responseMood
if err := json.Unmarshal([]byte(cleaned[start:end+1]), &parsed); err != nil {
return "", ""
if e := json.Unmarshal([]byte(cleaned[start:end+1]), &parsed); e != nil {
if strings.HasPrefix(cleaned, "{") {
return "", "", errBrokenJSON
}
return "", "", nil
}
return parsed.Response, parsed.Mood
return parsed.Response, parsed.Mood, nil
}
func parsePhrase(raw string) (body, summary string) {
+261
View File
@@ -0,0 +1,261 @@
package phraser
// Hand-written Russian nudges instead of generated ones.
//
// Why: on a nudge there is nothing to be creative about. Measured over many
// runs, Qwen3.5-0.8B breaks the persona (formal "вы", plural imperatives,
// masculine self-reference) and invents facts and units — it once told him to
// boil an egg for "90-95 секунд". A nudge is five words of known content, so
// wording it with a model buys nothing and risks the persona every time.
//
// The wording lives in nudges_ru_v1.json so it can be edited without touching
// Go. This file only picks one and fills in the values.
import (
"context"
_ "embed"
"encoding/json"
"fmt"
"math/rand"
"regexp"
"strings"
"sync"
"time"
"unicode"
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/loop"
)
//go:embed nudges_ru_v1.json
var nudgeTemplateJSON []byte
// NudgeTemplateSchemaVersion — the version this code understands.
const NudgeTemplateSchemaVersion = 1
type nudgeRuleSet struct {
Mood string `json:"mood"`
Variants []string `json:"variants"`
}
type nudgeTemplateFile struct {
SchemaVersion int `json:"schema_version"`
Name string `json:"name"`
Notes []string `json:"notes"`
Rules map[string]nudgeRuleSet `json:"rules"`
}
// NudgeTemplates picks a hand-written Russian nudge for a candidate.
//
// Safe for concurrent use. Random, but never the same variant twice in a row
// for the same rule — being nagged with identical words is what makes a nudge
// easy to tune out.
type NudgeTemplates struct {
mu sync.Mutex
rnd *rand.Rand
last map[string]string // rule family -> the text used last time
file nudgeTemplateFile
}
// NewNudgeTemplates loads the embedded template file. Pass a source to make the
// picking reproducible in tests; nil means seed from the clock.
func NewNudgeTemplates(src rand.Source) (*NudgeTemplates, error) {
var f nudgeTemplateFile
if err := json.Unmarshal(nudgeTemplateJSON, &f); err != nil {
return nil, fmt.Errorf("nudge templates: parse: %w", err)
}
if f.SchemaVersion != NudgeTemplateSchemaVersion {
return nil, fmt.Errorf("nudge templates: schema_version %d, want %d",
f.SchemaVersion, NudgeTemplateSchemaVersion)
}
if len(f.Rules) == 0 {
return nil, fmt.Errorf("nudge templates: no rules")
}
if src == nil {
src = rand.NewSource(time.Now().UnixNano())
}
return &NudgeTemplates{
rnd: rand.New(src),
last: map[string]string{},
file: f,
}, nil
}
// PhraseNudge implements the nudge half of the Phraser interface, so the
// templates can be scored by the same harness as the model.
func (t *NudgeTemplates) PhraseNudge(_ context.Context, c loop.Candidate) (delivery.PhrasedNudge, error) {
body, mood := t.Nudge(c)
return delivery.PhrasedNudge{Candidate: c, Body: body, Summary: body, Mood: mood}, nil
}
// Nudge returns the text and the mood for one candidate. Never fails: if no
// template fits it uses the plain per-rule fallback.
func (t *NudgeTemplates) Nudge(c loop.Candidate) (body, mood string) {
rule := c.Rule.Name
family := t.family(rule)
set, ok := t.file.Rules[family]
if !ok {
return fallbackNudge(c), "neutral"
}
vals := nudgeValues(c)
// Only variants whose placeholders all have a value.
usable := make([]string, 0, len(set.Variants))
for _, v := range set.Variants {
if text, ok := fillTemplate(v, vals); ok {
usable = append(usable, text)
}
}
if len(usable) == 0 {
return fallbackNudge(c), "neutral"
}
mood = set.Mood
if mood == "" {
mood = "neutral"
}
return t.pick(family, usable), mood
}
// pick chooses at random, skipping whatever this rule said last time.
func (t *NudgeTemplates) pick(family string, usable []string) string {
t.mu.Lock()
defer t.mu.Unlock()
choices := usable
if len(usable) > 1 {
choices = make([]string, 0, len(usable))
for _, v := range usable {
if v != t.last[family] {
choices = append(choices, v)
}
}
if len(choices) == 0 { // every variant equals the last one
choices = usable
}
}
got := choices[t.rnd.Intn(len(choices))]
t.last[family] = got
return got
}
// family maps a rule name to a block in the template file: an exact match
// first, then the prefix of "routine:зарядка" / "morning:утро", then "default".
func (t *NudgeTemplates) family(rule string) string {
if _, ok := t.file.Rules[rule]; ok {
return rule
}
if i := strings.IndexByte(rule, ':'); i > 0 {
if _, ok := t.file.Rules[rule[:i]]; ok {
return rule[:i]
}
}
return "default"
}
// placeholderRE — the {name} slots a template may use.
var placeholderRE = regexp.MustCompile(`\{([a-z]+)\}`)
// nudgeValues collects what this candidate can fill in. A key missing here
// means every template needing it is skipped, so nothing half-filled is ever
// spoken.
func nudgeValues(c loop.Candidate) map[string]string {
vals := map[string]string{}
rule := c.Rule.Name
// {since} — only at hour scale. Below an hour the phrase would be minutes,
// and none of the templates read well with "сорок минут".
if d, ok := c.State.Since(rule); ok && d >= time.Hour {
if s := ruSinceWords(d); s != "" {
vals["since"] = s
}
}
// {service} — the aggregate fact's key carries the service name.
if f, ok := c.State.Fact(rule); ok && f.Key != "" && f.Key != rule {
vals["service"] = f.Key
}
// {what} — the Russian suffix of "routine:таблетки" / "morning:утро".
if i := strings.IndexByte(rule, ':'); i > 0 && i+1 < len(rule) {
vals["what"] = rule[i+1:]
}
return vals
}
// fillTemplate substitutes the placeholders. Returns false when a value is
// missing, so a raw "{since}" can never reach the text-to-speech voice.
func fillTemplate(tmpl string, vals map[string]string) (string, bool) {
missing := false
out := placeholderRE.ReplaceAllStringFunc(tmpl, func(m string) string {
name := m[1 : len(m)-1]
v, ok := vals[name]
if !ok || v == "" {
missing = true
return m
}
return v
})
if missing || strings.ContainsAny(out, "{}%") {
return "", false
}
return capitalizeFirst(out), true
}
// capitalizeFirst — a placeholder can start the sentence, and "полтора часа без
// перерыва" should be spoken as a sentence, not a fragment.
func capitalizeFirst(s string) string {
for i, r := range s {
return string(unicode.ToUpper(r)) + s[i+len(string(r)):]
}
return s
}
// hourWords — hours spelled out. "3 ч" is fine on a screen and wrong in a
// Russian voice, so the number goes out as words.
var hourWords = []string{
"ноль", "один", "два", "три", "четыре", "пять", "шесть", "семь", "восемь",
"девять", "десять", "одиннадцать", "двенадцать", "тринадцать",
"четырнадцать", "пятнадцать", "шестнадцать", "семнадцать", "восемнадцать",
"девятнадцать", "двадцать", "двадцать один", "двадцать два", "двадцать три",
}
// hourPlural — час / часа / часов by Russian counting rules.
func hourPlural(h int) string {
if h%100 >= 11 && h%100 <= 14 {
return "часов"
}
switch h % 10 {
case 1:
return "час"
case 2, 3, 4:
return "часа"
default:
return "часов"
}
}
// ruSinceWords — "полтора часа", "два с половиной часа", "семь часов".
// Empty string means "do not say it" (under an hour, or over a day).
func ruSinceWords(d time.Duration) string {
if d < time.Hour {
return ""
}
h := int(d.Hours())
m := int(d.Minutes()) % 60
if m >= 45 {
h++
m = 0
}
if h >= len(hourWords) {
return "больше суток"
}
if h == 1 {
if m >= 15 {
return "полтора часа"
}
return "час"
}
if m >= 15 {
return hourWords[h] + " с половиной часа"
}
return hourWords[h] + " " + hourPlural(h)
}
+202
View File
@@ -0,0 +1,202 @@
package phraser
import (
"context"
"math/rand"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/store"
)
// cand builds a candidate the way a tick would.
func cand(rule string, sinceMin int, factKey string) loop.Candidate {
now := time.Date(2026, 7, 31, 21, 40, 0, 0, time.UTC)
st := loop.State{Now: now, Facts: map[string]store.Fact{}}
if sinceMin > 0 || factKey != "" {
key := rule
if factKey != "" {
key = factKey
}
st.Facts[rule] = store.Fact{Key: key, Ts: now.Add(-time.Duration(sinceMin) * time.Minute)}
}
return loop.Candidate{Rule: loop.Rule{Name: rule, Severity: loop.Sev1}, Severity: loop.Sev1, State: st}
}
func newTestTemplates(t *testing.T, seed int64) *NudgeTemplates {
t.Helper()
nt, err := NewNudgeTemplates(rand.NewSource(seed))
if err != nil {
t.Fatalf("NewNudgeTemplates: %v", err)
}
return nt
}
func TestNudgeTemplatesLoad(t *testing.T) {
nt := newTestTemplates(t, 1)
for _, rule := range []string{"water", "meal", "break", "service_down", "netdata_critical", "routine", "morning", "default"} {
set, ok := nt.file.Rules[rule]
if !ok {
t.Errorf("no templates for %q", rule)
continue
}
if len(set.Variants) < 5 {
t.Errorf("%s: only %d variants", rule, len(set.Variants))
}
// Every rule needs one variant that needs no value, or a candidate
// without context has nothing to say. routine and morning are exempt:
// they always carry a name and must always say it.
plain := 0
seen := map[string]bool{}
for _, v := range set.Variants {
if !placeholderRE.MatchString(v) {
plain++
}
if seen[v] {
t.Errorf("%s: duplicate variant %q", rule, v)
}
seen[v] = true
}
if plain == 0 && rule != "routine" && rule != "morning" {
t.Errorf("%s: every variant needs a placeholder value", rule)
}
}
}
// The whole point of the picker: never the same words twice in a row.
func TestNudgeNoImmediateRepeat(t *testing.T) {
nt := newTestTemplates(t, 7)
prev := ""
for i := 0; i < 200; i++ {
body, _ := nt.Nudge(cand("water", 200, ""))
if body == prev {
t.Fatalf("repeat at %d: %q", i, body)
}
prev = body
}
}
// Same seed, same sequence — otherwise the fixture score would drift run to run.
func TestNudgeDeterministicWithSeed(t *testing.T) {
var runs [2][]string
for r := range runs {
nt := newTestTemplates(t, 42)
for i := 0; i < 20; i++ {
body, _ := nt.Nudge(cand("break", 100, ""))
runs[r] = append(runs[r], body)
}
}
for i := range runs[0] {
if runs[0][i] != runs[1][i] {
t.Fatalf("run %d differs: %q vs %q", i, runs[0][i], runs[1][i])
}
}
}
// A variant is only used when its value exists, and nothing half-filled ships.
func TestNudgeNoLeftoverPlaceholders(t *testing.T) {
nt := newTestTemplates(t, 3)
cases := []loop.Candidate{
cand("water", 0, ""), // no duration
cand("water", 30, ""), // under an hour
cand("water", 200, ""), // hours
cand("service_down", 3, "vaultwarden"),
cand("service_down", 3, ""), // no service name
cand("routine:таблетки", 0, ""),
cand("morning:утро", 0, ""),
cand("unknown_rule", 0, ""),
}
for _, c := range cases {
for i := 0; i < 40; i++ {
body, mood := nt.Nudge(c)
if body == "" {
t.Fatalf("%s: empty body", c.Rule.Name)
}
if strings.ContainsAny(body, "{}%") {
t.Fatalf("%s: unfilled template %q", c.Rule.Name, body)
}
if mood != "neutral" {
t.Fatalf("%s: mood %q", c.Rule.Name, mood)
}
}
}
}
// The routine name must actually land in the text.
func TestNudgeSubstitutesWhat(t *testing.T) {
nt := newTestTemplates(t, 11)
for i := 0; i < 40; i++ {
body, _ := nt.Nudge(cand("routine:таблетки", 0, ""))
if !strings.Contains(strings.ToLower(body), "таблетки") {
t.Fatalf("routine text lost the name: %q", body)
}
}
}
func TestRuSinceWords(t *testing.T) {
cases := []struct {
min int
want string
}{
{30, ""},
{60, "час"},
{95, "полтора часа"},
{150, "два с половиной часа"},
{190, "три часа"},
{240, "четыре часа"},
{430, "семь часов"},
{660, "одиннадцать часов"},
{60 * 30, "больше суток"},
}
for _, c := range cases {
got := ruSinceWords(time.Duration(c.min) * time.Minute)
if got != c.want {
t.Errorf("%d min: got %q want %q", c.min, got, c.want)
}
}
}
// Templates are the default: a nudge must not reach the model at all.
func TestLLMPhraserUsesTemplatesByDefault(t *testing.T) {
spy := newGrammarSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{})
pn, err := p.PhraseNudge(context.Background(), cand("water", 200, ""))
if err != nil {
t.Fatalf("PhraseNudge: %v", err)
}
if len(spy.grammars) != 0 {
t.Errorf("nudge hit the model %d times, want 0", len(spy.grammars))
}
if !strings.Contains(strings.ToLower(pn.Body), "вод") {
t.Errorf("nudge is not the water template: %q", pn.Body)
}
}
// ...and the flag brings the model back.
func TestLLMNudgesFlagRestoresTheModel(t *testing.T) {
spy := newGrammarSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{LLMNudges: true})
pn, err := p.PhraseNudge(context.Background(), cand("water", 200, ""))
if err != nil {
t.Fatalf("PhraseNudge: %v", err)
}
if len(spy.grammars) != 1 {
t.Fatalf("nudge hit the model %d times, want 1", len(spy.grammars))
}
if pn.Body != "ага" {
t.Errorf("body = %q, want the model's reply", pn.Body)
}
}
func TestNudgeTemplatesPhraseNudge(t *testing.T) {
nt := newTestTemplates(t, 5)
pn, err := nt.PhraseNudge(context.Background(), cand("water", 200, ""))
if err != nil {
t.Fatalf("PhraseNudge: %v", err)
}
if pn.Body == "" || pn.Summary != pn.Body || pn.Mood != "neutral" {
t.Fatalf("bad nudge: %+v", pn)
}
}
+129
View File
@@ -0,0 +1,129 @@
{
"schema_version": 1,
"name": "russian nudge templates v1",
"notes": [
"Hand-written Russian nudges. Edit the wording here, no Go changes needed.",
"Rules: she is feminine about herself, he is a man addressed as ты. Never вы/вас/ваш, never plural imperatives (выпейте), never он/его about him.",
"One short sentence. No questions, no emoji, no pet names, no emotional support.",
"Placeholders: {since} how long it has been (only used when it is at least an hour), {service} the service name, {what} the routine name. A variant whose placeholder has no value is skipped, so every rule needs at least one variant with no placeholder. The exception is routine and morning: those only exist for rules like routine:таблетки that always carry a name, and a routine nudge that drops the name is useless.",
"mood must be one of: neutral, happy, thinking, tired, confused."
],
"rules": {
"water": {
"mood": "neutral",
"variants": [
"Ты не пил воду {since} — выпей стакан.",
"Пора выпить воды.",
"Стакан воды не помешает.",
"Воду ты не пил уже {since}.",
"Напоминаю про воду.",
"Сходи за водой, дела подождут.",
"Сделай глоток воды, пока помнишь.",
"Между делом выпей воды.",
"Вода — простое дело: выпей стакан.",
"Отвлекись на стакан воды."
]
},
"meal": {
"mood": "neutral",
"variants": [
"Ты не ел {since} — поешь.",
"Пора поесть, сделай перекус.",
"Еда важнее ещё одного часа за столом.",
"Без еды уже {since}, поешь.",
"Напоминаю про еду — поешь.",
"Возьми перерыв на обед.",
"Сделай себе перекус, это пять минут.",
"Поешь, потом вернёшься к работе.",
"Поешь нормально, а не на ходу.",
"Еды не было {since} — разогрей что-нибудь."
]
},
"break": {
"mood": "neutral",
"variants": [
"Ты за столом {since} — встань и разомнись.",
"Пора сделать перерыв.",
"Встань на пять минут.",
"{since} без перерыва — отойди от экрана.",
"Напоминаю про перерыв.",
"Разомни спину, потом продолжишь.",
"Короткая пауза не сорвёт дела.",
"Отойди от компьютера на минуту.",
"Сидишь без перерыва {since}.",
"Встань, пройдись, вернись."
]
},
"service_down": {
"mood": "neutral",
"variants": [
"Сервис {service} не отвечает.",
"{service} упал — сервис не отвечает.",
"{service} не отвечает, сервис нужно поднимать.",
"Сервис {service} недоступен.",
"Проверь {service}: сервис не отвечает.",
"Сервис перестал отвечать.",
"Сервис {service} лежит, нужно смотреть.",
"{service} не отвечает уже {since}.",
"Мониторинг сообщает: {service} лежит.",
"Сервис {service} не отвечает, посмотри логи."
]
},
"netdata_critical": {
"mood": "neutral",
"variants": [
"Netdata: критический алярм, проверь диск.",
"Критический алярм в netdata — посмотри диск.",
"Netdata поднял тревогу по диску.",
"Проверь диск: netdata ругается.",
"Алярм от netdata, критический.",
"Netdata: критический уровень, дело в диске.",
"Диск требует внимания — критический алярм в netdata.",
"Критический алярм: проверь место на диске.",
"Netdata сообщает о критической проблеме с диском.",
"Открой netdata: там критический алярм по диску."
]
},
"routine": {
"mood": "neutral",
"variants": [
"По распорядку: {what}.",
"Пора — {what}.",
"Напоминаю: {what}.",
"В списке на сейчас: {what}.",
"{what} — сейчас самое время.",
"Не пропусти: {what}.",
"{what}: пора сделать.",
"Сейчас по плану {what}.",
"Твой распорядок: {what}.",
"{what} — по распорядку сейчас."
]
},
"morning": {
"mood": "neutral",
"variants": [
"{what} — пора начать день.",
"{what}: пройди утренний список.",
"Начни {what} со списка.",
"{what}. Осталось пройти чеклист.",
"Утренний список ещё не пройден: {what}.",
"{what}: первый пункт списка за тобой.",
"{what} идёт, а список стоит.",
"{what}: не забудь про утренние дела.",
"По утреннему чеклисту ещё есть дела: {what}.",
"{what} — утренний список дел ещё ждёт."
]
},
"default": {
"mood": "neutral",
"variants": [
"Напоминаю: есть дело.",
"Пора вернуться к отложенному делу.",
"Одно дело ждёт тебя.",
"Напоминаю про дело из списка.",
"В списке осталось дело.",
"Дело всё ещё не сделано."
]
}
}
}
+23
View File
@@ -2,6 +2,7 @@ package router
import (
"context"
"fmt"
"math"
"unicode"
)
@@ -19,6 +20,24 @@ type Embedder interface {
Close() error
}
// IdentifiedEmbedder — an embedder that can name itself. The name goes into
// the DB next to the vectors it wrote, so a later model swap is caught instead
// of silently returning nonsense scores (Vikunja #378).
type IdentifiedEmbedder interface {
Embedder
ID() string
}
// EmbedderID is the stable string stored alongside the vectors. It comes from
// the embedder itself — nobody hand-types a model name twice — and changes
// whenever the model or its dimension changes.
func EmbedderID(e Embedder) string {
if i, ok := e.(IdentifiedEmbedder); ok {
return i.ID()
}
return fmt.Sprintf("unknown@%d", e.Dim())
}
// AsymmetricEmbedder — an embedder that wants to know whether a text is a
// search query or a stored passage. Recall is asymmetric: a short question
// goes in, a longer note comes out. The e5 family is trained for exactly that
@@ -70,6 +89,10 @@ func NewHashEmbedder(dim int) *HashEmbedder {
func (h *HashEmbedder) Dim() int { return h.dim }
// ID names this embedder for the DB marker. The dimension is part of it
// because a HashEmbedder of another width is a different vector space.
func (h *HashEmbedder) ID() string { return fmt.Sprintf("hash@%d", h.dim) }
func (h *HashEmbedder) Close() error { return nil }
func (h *HashEmbedder) Embed(_ context.Context, text string) ([]float32, error) {
+24
View File
@@ -0,0 +1,24 @@
package router
import "testing"
func TestEmbedderIDFromModelPath(t *testing.T) {
got := modelIDFromPath("/opt/maven/models/embedder/multilingual-e5-small.onnx")
if got != "multilingual-e5-small@384" {
t.Fatalf("modelIDFromPath = %q", got)
}
// A different model file must produce a different id, even at 384 dim.
old := modelIDFromPath("/opt/maven/models/embedder/paraphrase-multilingual-MiniLM-L12-v2.onnx")
if old == got {
t.Fatal("two different models share one id")
}
}
func TestEmbedderIDIncludesDim(t *testing.T) {
if id := EmbedderID(NewHashEmbedder(1024)); id != "hash@1024" {
t.Fatalf("EmbedderID = %q", id)
}
if EmbedderID(NewHashEmbedder(1024)) == EmbedderID(NewHashEmbedder(384)) {
t.Fatal("dimension not part of the id")
}
}
+17 -99
View File
@@ -1,11 +1,8 @@
package eval
import (
"bytes"
"context"
"encoding/json"
"fmt"
"net/http"
"os"
"strings"
"testing"
@@ -29,13 +26,20 @@ import (
// a bake-off across checkpoints (#278, #250) produces tables you can tell
// apart. Point the variable at one server at a time.
//
// Three configurations, because "the LLM router" is ambiguous and the three
// numbers answer different questions:
// Two configurations, because "the LLM router" is ambiguous and the two numbers
// answer different questions:
//
// llm-only — the model alone. Measures the prompt + grammar contract.
// cascade+llm — what #320 would actually ship: stage-0 grammar, then the
// model, then the classifier as the failure floor.
// llm-no-thinking — diagnostic only, not a shippable path (see below).
// llm-only — the model alone. Measures the prompt + grammar contract.
// cascade+llm — what #320 would actually ship: stage-0 grammar, then the
// model, then the classifier as the failure floor.
//
// There used to be a third, "thinking off", which looked 6 points better. It is
// gone: it was measured with a hand-rolled HTTP client that quietly dropped
// repeat_penalty, so the gap was the missing penalty and not the thinking mode.
// Re-measured with everything else held equal, thinking off scores exactly the
// same, case for case — and a direct probe shows this llama-server build ignores
// enable_thinking / reasoning_budget for this model anyway, so there was nothing
// to turn off. Full write-up in ROUTING-EVAL-31-07-2026.md (Vikunja #376).
func TestLLMRouterBaseline(t *testing.T) {
base := os.Getenv("MAVEN_LLM_URL")
if base == "" {
@@ -57,12 +61,12 @@ func TestLLMRouterBaseline(t *testing.T) {
}
ctx := context.Background()
model, err := ModelID(ctx, base)
model, err := llm.ModelID(ctx, base)
if err != nil {
// Not fatal: an unlabelled score is still a score. But say so loudly,
// because an unlabelled row in a bake-off table is worthless.
t.Logf("could not read model id from %s: %v — reports will say %q", base, err, "unknown-model")
model = "unknown-model"
t.Logf("could not read model id from %s: %v — reports will say %q", base, err, llm.UnknownModel)
model = llm.UnknownModel
}
t.Logf("scoring model %s at %s", model, base)
lr := router.NewLLMRouter(client)
@@ -96,104 +100,18 @@ func TestLLMRouterBaseline(t *testing.T) {
}
t.Log("\n" + repCascade.String() + repCascade.Failures())
// llm-no-thinking: same prompt and grammar with the chat template's
// thinking mode off. Qwen3.5's template defaults thinking=1, so under a
// grammar the constrained JSON lands in reasoning_content with content
// empty — llm.Client's ReasoningContent fallback is what makes the router
// work at all today, by accident rather than design.
//
// MEASURED 2026-07-31: this variant scores identically to as-deployed
// (18/76, 48.7% intent-only, 2 errors, same p50). Thinking mode is a
// non-issue under a grammar — llama.cpp constrains the same token stream
// either way. Kept so the question stays answered instead of being
// re-asked, and so internal/llm does NOT grow a chat_template_kwargs field
// for a problem that does not exist.
repNoThink, err := Score(ctx, "llm-only ("+model+", thinking off) [diagnostic]",
RouterFunc(func(ctx context.Context, u string, now time.Time) (router.Decision, error) {
d, ok, err := router.NewLLMRouter(&noThinkCompleter{base: base, http: &http.Client{Timeout: 60 * time.Second}}).Route(ctx, u, now)
if err != nil {
return d, err
}
if !ok {
return d, fmt.Errorf("llm router declined without an error")
}
return d, nil
}), f)
if err != nil {
t.Fatalf("Score no-thinking: %v", err)
}
t.Log("\n" + repNoThink.String() + repNoThink.Failures())
// Reports rather than asserts — the numbers are inputs to the #320
// decision, and an assertion here would be this test inventing the bar.
// The one thing worth failing on is a harness fault: if every single case
// errors, the run measured infrastructure, not routing, and the report
// must not be mistaken for a score.
for _, rep := range []Report{repLLM, repCascade, repNoThink} {
for _, rep := range []Report{repLLM, repCascade} {
if rep.Errors == rep.Total {
t.Errorf("%s: all %d cases errored — harness fault, not a measurement", rep.Name, rep.Total)
}
}
}
// noThinkCompleter — llm.Client with chat_template_kwargs.enable_thinking
// false. A test-local copy rather than a change to internal/llm: whether the
// daemon should send it is the open question, and answering it here by adding
// the field would prejudge #320.
type noThinkCompleter struct {
base string
http *http.Client
}
func (c *noThinkCompleter) Complete(ctx context.Context, r llm.Req) (string, error) {
payload := map[string]any{
"messages": []map[string]string{
{"role": "system", "content": r.System},
{"role": "user", "content": r.User},
},
"max_tokens": r.MaxTokens,
"temperature": 0,
"grammar": r.Grammar,
"chat_template_kwargs": map[string]any{"enable_thinking": false},
}
b, err := json.Marshal(payload)
if err != nil {
return "", err
}
req, err := http.NewRequestWithContext(ctx, "POST", c.base+"/v1/chat/completions", bytes.NewReader(b))
if err != nil {
return "", err
}
req.Header.Set("Content-Type", "application/json")
resp, err := c.http.Do(req)
if err != nil {
return "", err
}
defer resp.Body.Close()
if resp.StatusCode != 200 {
return "", fmt.Errorf("status %d", resp.StatusCode)
}
var out struct {
Choices []struct {
Message struct {
Content string `json:"content"`
ReasoningContent string `json:"reasoning_content"`
} `json:"message"`
} `json:"choices"`
}
if err := json.NewDecoder(resp.Body).Decode(&out); err != nil {
return "", err
}
if len(out.Choices) == 0 {
return "", fmt.Errorf("no choices")
}
m := out.Choices[0].Message
if m.Content != "" {
return m.Content, nil
}
return m.ReasoningContent, nil
}
func ping(ctx context.Context, c *llm.Client) error {
ctx, cancel := context.WithTimeout(ctx, 90*time.Second)
defer cancel()
+1
View File
@@ -22,6 +22,7 @@
{ "id": "ru-query-011", "utterance": "почему сервер тормозит", "lang": "ru", "intent": "query", "tags": ["homelab", "hard"], "note": "diagnostic question, not a chat opener" },
{ "id": "ru-query-012", "utterance": "какие заметки я оставил про полив", "lang": "ru", "intent": "query", "tags": ["recall"] },
{ "id": "ru-query-013", "utterance": "во сколько у меня встреча", "lang": "ru", "intent": "query", "tags": ["calendar"] },
{ "id": "ru-query-019", "utterance": "что у меня стоит в календаре на послезавтра", "lang": "ru", "intent": "query", "tags": ["calendar", "hard"], "note": "agenda, not the clock: the daemon answers this from CalendarEvents inside the query branch, so the clock/date system rule must not swallow it" },
{ "id": "ru-query-014", "utterance": "я успеваю до дедлайна", "lang": "ru", "intent": "query", "tags": ["hard", "no-question-word"] },
{ "id": "ru-query-015", "utterance": "сколько я прошёл шагов", "lang": "ru", "intent": "query", "tags": ["aggregate"] },
{ "id": "ru-query-016", "utterance": "покажи давление за неделю", "lang": "ru", "intent": "query", "tags": ["hard", "imperative"], "note": "imperative form but a read — must not route to act" },
+4 -1
View File
@@ -3,5 +3,8 @@ package router
// KnowledgePrompt returns the system prompt for general knowledge questions
// that the phraser uses when no notes match the query.
func KnowledgePrompt() string {
return `Ты — Мавена, персональный ассистент. Ответь кратко из своих знаний. Если не знаешь — скажи "не знаю". Не выдумывай. Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}.`
// No self-introduction here: the shared persona block already says who she
// is, and this line used to disagree with it — a different name ("Мавена")
// and a masculine noun ("ассистент") in front of a feminine persona.
return `Ответь кратко из своих знаний. Если не знаешь — скажи "не знаю". Не выдумывай. Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}.`
}
+9 -1
View File
@@ -11,10 +11,18 @@ func TestKnowledgePrompt(t *testing.T) {
t.Fatal("KnowledgePrompt returned empty string")
}
// Must contain key instructions
checks := []string{"Мавена", "не знаю", "не выдумывай"}
checks := []string{"не знаю", "не выдумывай"}
for _, c := range checks {
if !strings.Contains(strings.ToLower(prompt), strings.ToLower(c)) {
t.Errorf("KnowledgePrompt should mention %q", c)
}
}
// Who she is comes from the shared persona block now. This prompt used to
// say it too, with a different name and a masculine noun, which is the
// drift the block exists to stop.
for _, w := range []string{"Мавена", "ассистент"} {
if strings.Contains(prompt, w) {
t.Errorf("KnowledgePrompt should not introduce her (%q) — the persona block does", w)
}
}
}
+65 -9
View File
@@ -44,10 +44,21 @@ ws ::= [ \t\n]*
// Changed again 31-07-2026: added the "unknown" escape hatch so the model can
// admit it cannot route (Vikunja #359).
//
// Changed again 31-07-2026: added the clock/calendar rule (Vikunja #374). The
// prompt never said which side "который час" or "какое число завтра" belong on,
// so the model guessed — `system→query ×4` in every eval run. The rule sits
// above the question test on purpose: these utterances all carry a question
// word, so a later rule would never be reached. The boundary is what the
// daemon can actually answer: only replySystem in cmd/mavend/voice.go owns the
// clock and the calendar formatter, while the agenda ("что у меня завтра") is
// answered inside the query branch, so that side stays query.
//
// The training workspace keeps its own copy of this prompt for relabelling, and
// `llm/check_prompt_parity.py` there compares the two. That copy is in another
// repo and was not touched, so parity will fail until it gets the same edits —
// both the rule reorder and the "unknown" wording (Vikunja #362).
// both the rule reorder and the "unknown" wording (Vikunja #362) — and now the
// clock/calendar rule too. The training workspace is not checked out on this
// box at all, so it could not be updated here; #362 still covers the catch-up.
const routeSystem = `Классифицируй ровно одно сообщение пользователя. Верни ОДИН JSON-массив действий.
Ровно одно намерение: fact, reminder, note, query, act, chat, system.
@@ -56,19 +67,21 @@ const routeSystem = `Классифицируй ровно одно сообще
Классифицируй по цели пользователя. Порядок решения:
1. Хочет напоминание в будущем reminder
2. Явно просит сохранить информацию note
3. Задаёт вопрос: есть вопросительное слово (сколько, что, какой, когда, где, кто, почему, как) или знак «?» query
4. Хочет получить информацию, в том числе о своих же данных query
5. Утверждает: сообщает или обновляет текущее состояние/событие fact
6. Просит выполнить работу act
7. Про ассистента, настройки или память system
8. Реплика обрывок или указание на неназванное («это», «то», «потом»), и без него непонятно, что именно нужно сделать unknown
9. Иначе chat
3. Спрашивает только «который час» / «какое число» / «какой день недели» сами часы или календарная дата, без своих данных system
4. Задаёт вопрос: есть вопросительное слово (сколько, что, какой, когда, где, кто, почему, как) или знак «?» query
5. Хочет получить информацию, в том числе о своих же данных query
6. Утверждает: сообщает или обновляет текущее состояние/событие fact
7. Просит выполнить работу act
8. Про ассистента, настройки или память system
9. Реплика обрывок или указание на неназванное («это», «то», «потом»), и без него непонятно, что именно нужно сделать unknown
10. Иначе chat
Различия:
- note сохранить информацию, без напоминания. text = суть.
- reminder уведомить позже. text = что напомнить.
- fact неявное обновление: пользователь сообщает, что что-то в мире изменилось (текущее/изменённое состояние, случившееся событие). key/value.
- unknown редкий случай. Ставь его, только если в самой реплике нет ни предмета, ни действия. Короткая, простая или незнакомая тема это не причина для unknown: приветствие и болтовня это chat, вопрос на любую тему это query, просьба сделать что-то названное это act.
- system против query часы и календарная дата сами по себе (сколько времени, какое число, какой день недели можно и про завтра, и про другой город) это system. А что записано в календаре или в памяти («что у меня завтра», «какие есть напоминания») это query. Если в реплике есть просьба (напомни, запиши, сделай), то названное время просто деталь просьбы, и это не system.
- query против fact решает форма реплики, а не тема. Вопрос о состоянии это query, даже если названо то же самое, что бывает в fact. Только утверждение это fact.
Примеры:
@@ -82,6 +95,8 @@ const routeSystem = `Классифицируй ровно одно сообще
"что такое docker?" {"intent":"query","text":"что такое docker"}
"напиши письмо" {"intent":"act","verb":"написать письмо"}
"очисти память" {"intent":"system"}
"который час?" {"intent":"system"}
"какое число завтра?" {"intent":"system"}
"привет" {"intent":"chat","text":"привет"}
"сделай это" {"intent":"unknown"}
"ну это" {"intent":"unknown"}
@@ -103,6 +118,39 @@ const routeRepeatPenalty = 1.15
// Route return ok=false so the caller drops to the classifier cascade.
const routeIntentUnknown = "unknown"
// llmFullConfidence / llmThinConfidence — Vikunja #359. Confidence used to be
// hardcoded to 1.0 for every LLM decision, so the stage-3 gate in router.go
// never had anything to bite on and the LLM path could never produce a
// Clarify: on the 77-case RU fixture, 6/6 want_clarify cases were missed by
// EVERY model in the 31-07-2026 bake-off (0.8B through 2B) — proof this was a
// code bug, not a capability ceiling.
//
// The fix does not touch the prompt (routeSystem is under
// llm/check_prompt_parity.py in the training workspace; changing its text
// creates a parity break that has to be fixed there too — see Vikunja #362).
// Instead it reads structural signal that is already free:
// - a single-token utterance is thin evidence for anything a grammar
// didn't already catch at stage 0 — "вода" and "бэкап" alone don't say
// fact-vs-query or act-vs-report;
// - a fact with no key, or an act that never resolves to an allowlisted fn
// (checked in router.go, after slot-fill has had its say), is a decision
// with a hole in the one slot that makes it actionable.
//
// A model self-reporting confidence in the JSON was considered and rejected:
// a sub-2B is not calibrated (nothing stops it saying "confident" on exactly
// the cases it gets wrong today), and true logprobs would need a response
// field internal/llm.Client's Complete does not currently return — see
// internal/llm/client.go.
//
// llmThinConfidence sits below config.DefaultRouterThreshold (0.55) so the
// existing stage-3 gate in Router.Route treats it exactly like a low-scoring
// classifier result — same lane, same daemon-side clarify machinery
// (cmd/mavend/clarify.go), no new consumer to build.
const (
llmFullConfidence = 1.0
llmThinConfidence = 0.3
)
type routeAction struct {
Intent string `json:"intent"`
Key string `json:"key"`
@@ -145,7 +193,15 @@ func (lr *LLMRouter) Route(ctx context.Context, utterance string, now time.Time)
if a.Intent == routeIntentUnknown {
return Decision{}, false, nil
}
d := Decision{Utterance: utterance, Stage: 1, Confidence: 1.0}
d := Decision{Utterance: utterance, Stage: 1, Confidence: llmFullConfidence}
// A single-token utterance is thin evidence: the model had nothing to
// disambiguate on ("вода" is a fact-or-query coin flip, "бэкап" an
// act-or-report one) and stage 0 would already have won on anything
// that pattern-matches cleanly. Flag it now; router.go's stage-3 gate
// (Router.Route) decides whether that trips Clarify.
if len(strings.Fields(utterance)) <= 1 {
d.Confidence = llmThinConfidence
}
switch Intent(a.Intent) {
case IntentFact:
d.Intent = IntentFact
+97
View File
@@ -258,4 +258,101 @@ func TestLLMFactGetsKeyFromParser(t *testing.T) {
if !d.Slots.HasKey || d.Slots.Key != "water" {
t.Fatalf("want key=water, got %+v", d.Slots)
}
if d.Clarify {
t.Fatalf("the parser resolved the key, this must not clarify: %+v", d)
}
}
// --- confidence / stage-3 gate on the LLM path (Vikunja #359) -----------------
// A single-token utterance is thin evidence on its own — "вода" alone is a
// fact/query coin flip. The gate must ask rather than guess confidently.
func TestLLMRouterSingleTokenTripsClarify(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"query","text":"вода"}`)
d, err := r.Route(context.Background(), "вода", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if !d.Clarify {
t.Fatalf("a bare single-token decision must clarify, got %+v", d)
}
}
// A multi-word utterance with a clean answer must not be punished — the
// whole point is not trading the confident cases away for clarify coverage.
func TestLLMRouterMultiWordStaysConfident(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"reminder","text":"позвонить маме"}`)
d, err := r.Route(context.Background(), "напомни позвонить маме", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if d.Clarify {
t.Fatalf("a clean multi-word decision must not clarify: %+v", d)
}
if d.Confidence != llmFullConfidence {
t.Fatalf("want full confidence, got %v", d.Confidence)
}
}
// "бэкап" alone: the model guesses act, but nothing on the allowlist matches
// "бэкап" as a verb — that must not fire a tool blind.
func TestLLMRouterActWithoutFnTripsClarify(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"act","verb":"бэкап"}`)
d, err := r.Route(context.Background(), "бэкап сделай пожалуйста расписание", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if d.Slots.HasFn {
t.Fatalf("test setup drifted: %q now resolves to an fn", d.Slots.Fn)
}
if !d.Clarify {
t.Fatalf("an unresolved act must clarify rather than guess: %+v", d)
}
}
// An act that DOES resolve to an allowlisted fn must stay confident even
// though its own verb is single-word-ish in spirit — guard against the fn
// check firing on the happy path.
func TestLLMRouterActWithFnStaysConfident(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"act","verb":"restart nginx"}`)
d, err := r.Route(context.Background(), "слушай, restart nginx пожалуйста", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if !d.Slots.HasFn {
t.Fatalf("test setup drifted, want fn resolved: %+v", d.Slots)
}
if d.Clarify {
t.Fatalf("a resolved act must not clarify: %+v", d)
}
}
// A fact where NEITHER the model NOR the deterministic parser can name a key
// must clarify instead of silently writing under an empty/guessed key.
func TestLLMRouterFactWithoutKeyTripsClarify(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"fact","value":"что-то"}`)
d, err := r.Route(context.Background(), "у меня какая-то фигня случилась вот прямо только что", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if d.Slots.HasKey {
t.Fatalf("test setup drifted: parser now resolves a key for this utterance")
}
if !d.Clarify {
t.Fatalf("a keyless fact must clarify rather than guess: %+v", d)
}
}
// The whole point of #359: the classifier cascade cannot be traded away for
// clarify coverage. A multi-word fact the parser CAN key must stay confident
// through the full Router.Route path, not just the raw LLMRouter.
func TestRouterLLMFactWithResolvedKeyStaysConfident(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"fact","text":"я выпил воду"}`)
d, err := r.Route(context.Background(), "я выпил воду", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if d.Clarify {
t.Fatalf("a fact the parser could key must not clarify: %+v", d)
}
}
+21
View File
@@ -33,6 +33,7 @@ const (
type onnxEmbedder struct {
tokenizer *unigramTokenizer
session *ort.DynamicSession[int64, float32]
id string
}
func NewONNXEmbedder(modelPath, tokenizerPath, libPath string) (*onnxEmbedder, error) {
@@ -58,11 +59,31 @@ func NewONNXEmbedder(modelPath, tokenizerPath, libPath string) (*onnxEmbedder, e
return &onnxEmbedder{
tokenizer: tok,
session: session,
id: modelIDFromPath(modelPath),
}, nil
}
func (e *onnxEmbedder) Dim() int { return embedDim }
// ID names the loaded model for the DB marker (Vikunja #378): the model file's
// own name plus the dimension, so pointing the config at another model changes
// the string on its own.
func (e *onnxEmbedder) ID() string { return e.id }
// modelIDFromPath turns /opt/.../multilingual-e5-small.onnx into
// "multilingual-e5-small@384".
func modelIDFromPath(modelPath string) string {
name := modelPath
if i := strings.LastIndexAny(name, "/\\"); i >= 0 {
name = name[i+1:]
}
name = strings.TrimSuffix(name, ".onnx")
if name == "" {
name = "onnx"
}
return fmt.Sprintf("%s@%d", name, embedDim)
}
// Embed treats the text as a query. The classifier compares one short
// utterance to another short seed phrase, so both sides get the same prefix
// and the comparison stays fair. The recall path must call EmbedQuery and

Some files were not shown because too many files have changed in this diff Show More