Compare commits

...

87 Commits

Author SHA1 Message Date
kami 7b2b96b957 Capture tasks, with one intake seam mail can call later (#130)
A task is not a fact and not a note. A fact is a claim about the world that a
correction supersedes; a note is something to recall by meaning. A task is work
with a lifecycle, and the read that matters is "everything outstanding right
now" — which over an append-only log would mean replaying history on every
question. So: a tasks table, migration #14, statuses candidate/open/done/dropped
that each move forward exactly once.

Dedupe is on normalised text among LIVE rows only, via a partial unique index.
That is the property the mail side needs: an extractor may call CaptureTask for
every message it reads, as often as it likes, without growing the list — while a
weekly errand is still capturable again once the last one is done.

Three ways in, one seam. ipc.CaptureTaskReq is it: the voice path
(router.ParseTaskCapture on an explicit marker — "добавь в задачи …", never
"надо бы поспать"), the /tasks form, and the email reader from #246 when it
exists. Mail-derived items set Source "email:<account>", Status "candidate" and
Evidence to whatever makes the row reviewable; a candidate is inert until he
confirms it on /tasks, and Maven names it as unconfirmed when she recites the
list rather than putting words in his mouth.

No new intent — the router enum is a contract with the relabelling prompt, so
capture rides the note intent and the list rides a query source, both matched
deterministically like the calendar and plan matchers already are.

Nothing here speaks. No tick rule reads tasks; the list is answered when asked
about, which is why /tasks POST is not step-up gated the way /tools and
/routines are — a task write moves no boundary.

Vikunja #130
2026-08-01 02:33:47 +04:00
kami c8444813e2 Answer "что я обычно делаю по вторникам?" by counting, not guessing (#254)
Behavioural memory, narrowed on purpose. internal/memory/behavior.go builds a
profile out of self-facts — distinct days per weekday, median time of day — and
reads it back in RU; router.ParseHabitQuery finds the weekday deterministically;
a `habits` query source answers the question.

Three things the plan doc asks for are deliberately absent, and the doc now
records why:

- The profile is COUNTED, not LLM-generated. A 1.7B asked to summarise a year of
  habits writes fluent claims about the owner's life that no row supports, and a
  wrong claim about him is the most expensive kind of wrong maven can be.
- No cached profile fact, so no "update on fact write" machinery. It is
  recomputed on the question; a cache that can disagree with its own rows is two
  truths.
- No proactive daily plan nudge. A dispatcher proposal at 08:00 every day is the
  definition of a nag. The path from "she noticed a pattern" to "she acts on it"
  already exists in internal/pattern with the proposal queue on /routines, and it
  goes through him.

A one-off is not a habit: an activity needs two distinct days before she will
call it usual, and until then she says she does not know yet. Only self-facts
count — env rows are the world, config rows are her own tuning state. The typical
time is a median so one 03:00 outlier cannot move a morning habit into the night.
An unrecognised fact key is read back verbatim rather than glossed into something
she made up.

The source sits before "calendar" in querySources, and its matcher requires a
habit marker, so "что я делаю в среду?" still reaches the calendar — answering a
question about this coming Wednesday with a statistical average would be
answering a different question.

Verified: make build and make test both exit 0.
2026-08-01 02:21:09 +04:00
kami ed9bdd5e09 Add the day plan she can recite when asked (#128)
The plan answers "какие планы на сегодня?" by putting one day in order:
calendar events (with #126's ambient provenance carried through and hedged),
pending reminders, and one line per morning routine that still has items
outstanding. "что дальше?" trims what has already passed.

It lives in internal/morning, not in a parallel system, because it is the same
question the checklist asks at a different scale — the routine knows what is
missing from a window, the plan knows what the whole day holds, and both read
the same facts and the same idea of "today". BuildPlan is pure; tickLoop.dayPlan
is the impure half that reads the store.

It is not a nag. Nothing here fires, schedules or announces: the plan is built
only when asked, over IPC (day_plan) or on the existing /morning page.
Unprompted delivery stays with the morning nudge and the dispatcher's policy.

The query source sits before "calendar" in querySources because both match
"…на сегодня" and the plan's matcher is the more specific one; IsDayPlanQuery
matches whole words so "планёрка" (a meeting) is not read as a request for the
plan, and refuses any utterance naming another day, since the plan is built for
the clock's own day only.

Verified: make build and make test both exit 0; new tests cover plan ordering,
the checklist-only-what-is-left rule, other-day rejection, the RU rendering
against the persona checks, rest-of-day trimming, the source ordering, and the
matcher's refusals.
2026-08-01 02:15:18 +04:00
kami 49f089d8a6 Read the work calendar as a notification signal, not a mailbox (#126)
Maven does not get a work credential. A corp mail or calendar session living on
the homelab ties the box's blast radius to the employer's data, which is the
thing this task exists to refuse. What she reads instead is the signal: an
Android notification-listener on the phone relays meeting notifications over
wg/LAN to POST /api/ambient, and the ones that clearly describe a meeting become
calendar events at source=ambient:notif, confidence 0.6.

The provenance is the point. A notification is evidence about a meeting, not a
reading of a calendar, so it is never indistinguishable from one: it is stored
below full confidence, store.CalendarEvents keeps the source and confidence on
every row it returns, and the query path hedges — "похоже, Планёрка @ 14:00" for
a relayed event, plain text for a CalDAV read.

The parse is deliberately conservative (internal/calendar/ambient.go). It needs
a real clock reading and a summary that is not just that clock reading;
otherwise it stores nothing at all. A bare hour is not a time, an unread count
is not a time, and "срок 2026.08.15" does not offer 08:15 as a meeting — loose
digits in a notification are far more often a badge or a date, and a mailbox of
noise rendered as invented meetings is worse than a gap.

The ingest is off unless configured: no -ambient-token, no route registered. The
token is a shared secret compared in constant time, because the poster is a
background Android service and WebAuthn has no answer for one. The endpoint is
write-only, accepts one shape of write, and cannot read anything back out.
Reposts of the same notification dedupe against the latest fact for that
key+source, the same append-only discipline cmd/mavcaldav follows.

Not shipped: the Android relay app itself, which is a separate artifact and a
device, not Go in this repo.
2026-08-01 02:04:06 +04:00
kami 3af290152c Render maven's own reminders to a calendar she owns (#127)
Radicale becomes a write-only render target, not a store. sqlite stays
canonical: every poll mavcaldav reads the pending reminders out of core and
publishes each one as a single-event iCal resource, withdrawing the ones that
have fired or been cancelled. Losing the collection costs nothing — the next
tick rebuilds it, and nothing is ever read back from it.

It structurally cannot write to a calendar maven only reads. The render URL and
credential are their own flags, and -render-url is refused at startup when it
names the collection -url reads; the only paths it addresses carry the
maven-reminder- prefix, so even aimed at the wrong collection it can only touch
resources it created. Rendering is off unless -render-url is given.

The calendar data model now lives in one place, internal/calendar: the Event,
the iCal parse it comes from and the render it goes to, the fact key/value
encoding, and the source constants that say which calendars may be written to.
It was a parse inlined in cmd/mavcaldav and a Sprintf in two files; #126 and
#128 both need to agree with it.

Fixes a latent day-boundary bug moved out of that inline parse: it took the day
number off a local clock reading but built the window boundaries in UTC, so on
a box east of Greenwich part of the evening fell outside "today" and the poller
saw an empty calendar after 20:00 UTC. Today is now the owner's day in the
owner's location, which is what the busy gate and the day plan mean.
2026-08-01 01:55:51 +04:00
kami dc7c72a3d7 Add background memory evaluation, off unless configured (#248)
Ships the real, local, testable part of the memory-evaluation plan
(docs/plans/03-memory-evaluation.md): Maven reads back her own recent
memory on a slow ticker, asks the resident model what it notices, and
records the confident answers as notes.

internal/memeval — not internal/memory/eval.go as the plan says, because
internal/store imports internal/memory for the vector backend and an
evaluator has to read store.Fact/Note/Nudge, which would close the
cycle. Evaluate() gathers RecentFacts/RecentNotes/RecentNudges, prompts
under a GBNF grammar bounded to three {observation, confidence,
suggested_action} objects, drops anything under min_confidence,
deduplicates against what earlier runs wrote, and writes the rest as
notes with source infer:memory-eval. /dash already renders notes with
their source, so the output is visible with no UI change.

cmd/mavend/memoryeval.go drives it on its own goroutine and ticker, not
on the 60s tick: an evaluation is a multi-second round-trip on the same
llama-server that answers voice turns, and it runs hourly at most. The
memory_eval config block is absent by default and absence means the
goroutine does not exist. No llama-server phraser also means no loop —
there is no template fallback, because a "memory evaluation" assembled
from templates is a fixed sentence pretending to be an observation.

What it deliberately cannot do, since this is the feature most likely to
turn Maven into a nag:

  - It cannot speak. No dispatcher reference, no channel, no nudge. An
    observation is a thought she wrote down and he reads on /dash.
    Announcing them is a separate decision with its own opt-in.
  - It cannot act. suggested_action is recorded as text and interpreted
    by nobody — no reminder, routine or fact is created from it.
  - It says nothing about an empty store: no memory means no LLM call,
    so there are no observations invented out of two facts.
  - Its own notes are excluded from the next evaluation's input, and are
    written with a nil embedding so they stay out of the recall pool.

The plan's remaining items (dispatching observations, an /eval IPC
method and trace view, RecentEvents) and the fact that output quality is
entirely unmeasured are written up at the bottom of the plan doc.
2026-08-01 01:45:49 +04:00
kami 766ca091a7 Announce tick-inferred routines, opt-in and rate-limited (#247, #43)
The digestion tick already runs the pattern detector over all recorded
events (67563ed) and writes a proposed_routines row. What was missing is
the other half of #247: a proposal that nobody is at the mic for reaches
nothing but the /routines page, so a pattern noticed at 03:00 is only
seen if he goes looking.

This wires the tick's proposals into the existing care-delivery path
rather than a second channel: sev1 nudge, loop.Gate, dispatcher, same
routing table as an accepted routine. Restraints, since a feature that
speaks unprompted is the easiest way to turn Maven into a nag:

  - off unless configured — the new pattern_proposals block, absent by
    default, and deploy/mavend.json ships notify: false;
  - at most one announcement per tick however many patterns surfaced;
  - at most one per cooldown (24h default) across all pairs;
  - sev1, so quiet hours, away and snooze suppress it;
  - suppressed means dropped, not queued — /routines still has it;
  - once per pair for good, since proposed_routines is
    UNIQUE(action, object) and the row survives dismissal.

The body is pattern.PhraseRoutine's literal Russian, not LLM-worded, so
an inferred routine cannot arrive describing something never observed.

Also raises pattern.MinEvents from 3 to 4 — the interval-quality item on
#43. Two intervals with a ±50% band is a coincidence with a mean, not a
pattern, and now that a scan of all history can announce itself the cost
of a false positive is a permanent dismissal of that pair.
2026-08-01 01:37:50 +04:00
kami c5317eb2b4 Move the quiet-toggle and pattern-extraction slices out of voice.go (#321)
Continues the decomposition PR #50 started. voice.go 542 -> 365:

  quiet_toggle.go  144  resolveQuietToggle, quietInflections, quietStem,
                        quietTokens, quietPhrase, quietOn/OffPhrases,
                        classifyQuietToggle  (quiet_toggle_test.go already
                        existed for these)
  patterns.go     +44  detectPattern, next to detectAndPropose which it calls
                        and which patterns.go's own header already pointed at

What is left in voice.go is the handler: reactiveHandler, HandlePushToTalk,
handleText, runTurn, applyAction, replySystem, chatHistory, reply.

Move-only: all 133 distinct non-blank lines removed from voice.go were
matched in the two destination files, zero lines added to voice.go. The only
non-move edits are import lists (log added to patterns.go, unicode and
internal/pattern dropped from voice.go) and two comments that pointed at
voice.go for code that is no longer there.
2026-08-01 01:30:06 +04:00
kami 9190f897a3 Add a locked-down maven.<domain> block to the nginx template (#354)
The template's wildcard `listen 80` with no ACL was fixed in 50cc17f, but it
still only covered nexus/praxis/hexis. mavweb — the one service in the set
that serves an RCE surface (POST /tools defines argv internal/tool executes)
— had no block at all, so anyone wiring it up wrote their own, which is how
the wildcard got there the first time.

Adds a maven.kvmx.ru server with the same wg+LAN bind and allow/deny,
proxying 127.0.0.1:9201, with the WebSocket upgrade /ws needs, a 32m body
limit for push-to-talk PCM, and a 300s read timeout because an LLM turn on
the iGPU is slow.

Also records in deploy/ecosystem/docker-compose.yml that the sibling
`build:` paths pin nothing and ship the sibling working tree, with the
command to check what is about to be deployed. The stale public DNS records
(item 2) are outside the repo.

Verified: nginx -t on the template inside a minimal http{} accepts it.
2026-08-01 01:27:11 +04:00
kami d29e7ba813 Gate POST /api/chat on the same step-up as /tools (#317)
/api/chat reaches the router, the LLM and, through applyAction, the whole
act path, so it is the widest state-changing surface mavweb serves. It was
the only one with no gate. It now goes through stepUpOK like POST /tools,
POST /routines and POST /api/revert: unchanged in the default deploy
(WebAuthn unconfigured, fail-open behind wg+nginx), 403 under
-require-stepup or an unasserted passkey session.

The route table now carries an explicit enumeration of every state-changing
route and its gate, and the two startup SECURITY log lines name /routines
and /api/chat alongside /tools and /api/revert.

The loopback -addr default the task also asked for landed earlier in
d12de58; the compose already publishes mavweb on 127.0.0.1 only.
2026-08-01 01:24:42 +04:00
kami f7e1187823 Match quiet-mode toggles on whole words, and resolve OFF first 2026-08-01 01:02:10 +04:00
kami ed48c59ba7 Merge branch 'refactor/query-sources' into integration/small-batch 2026-08-01 00:54:11 +04:00
kami b09967f9e6 Split actions.go into per-intent files
Pure move: actionFact, actionReminder, actionAct and actionNote each get
their own actions_<intent>.go. The two small ones (chat, system) and the
actionHandlers table stay in actions.go, which is now just the dispatch
layer and the notes about what does not belong in it. No behaviour
change — only the file a handler is read in.
2026-08-01 00:53:37 +04:00
kami b4a3867479 Turn actionQuery into a chain of query sources
The six answer sources were hand-unrolled inside one 127-line function.
The intent table is a closed set of 7, but this list is open-ended —
Kiwix (#286), RSS (#258), the crawler (#259) and email (#246) each add
one. Each is now a registry entry: a name plus a method on the handler,
walked in order until one claims the question.

Order is unchanged and still load-bearing (memory before the notes-only
pass, #373), the confidence gate keeps its position and semantics, and
every reply string, log line and best-effort failure is verbatim.
2026-08-01 00:51:44 +04:00
kami 88d07b5175 Unify the voice and text turn pipelines into runTurn
HandlePushToTalk and handleText hand-wrote the same eight-step turn
sequence twice, comments in the latter saying "same as HandlePushToTalk"
four times. Extract it into runTurn(ctx, text) string: the voice path
wraps it in stt/tts, the text path returns it directly.

The two had drifted. The text path was missing the quiet-hours toggle
check entirely, so "тихий режим" over IPC/telegram fell through to the
classifier; unifying gives it the check. It also logged the route result
and applyAction return where the voice path did not — both logs are kept
for both paths.
2026-08-01 00:48:50 +04:00
kami c00e3003bf Merge branch 'refactor/praxis-capability-registry' into integration/small-batch 2026-08-01 00:43:54 +04:00
kami c0f9834528 Turn the Praxis act dispatch into a capability registry 2026-08-01 00:43:16 +04:00
kami ad5eb2d1cf Walk a chain of confirm resolvers instead of three copied blocks 2026-08-01 00:42:23 +04:00
kami 5253123d99 Merge branch 'refactor/voice-wiring' into integration/small-batch
# Conflicts:
#	cmd/mavend/voice.go
2026-07-31 23:54:11 +04:00
kami 2abf98dea6 Move the voice daemon wiring and startup out of voice.go 2026-07-31 23:52:51 +04:00
kami f5c71b87f3 Move the confirm/park gate out of voice.go 2026-07-31 23:51:33 +04:00
kami 5934110fa8 Move the Praxis/Hexis act handling out of voice.go
voice.go is still the biggest file in cmd/mavend and most of what is left
has nothing to do with the audio path. The ecosystem integration is one
such lump: it talks to Nexus, Praxis and Hexis over HTTP and only touches
the handler for its store and clock. Lifting it into ecosystem_acts.go
puts it next to ecosystem.go, where the clients it drives already live.

Move-only: handlePraxisAct, recordPraxisTrace (called from nowhere else),
handleHexisAct and execHexis verbatim, plus the two imports that became
unused in voice.go.
2026-07-31 23:43:47 +04:00
kami 6e47a3d736 Merge branch 'worktree-agent-a88193d1d84b04a5b' into integration/small-batch 2026-07-31 23:38:00 +04:00
kami 67a5eb3805 Split applyAction's 300-line switch into a per-intent handler table
applyAction (cmd/mavend/voice.go) dispatched all 7 intents from one giant
switch. Extract each case body verbatim into its own actionXxx method in
new cmd/mavend/actions.go, dispatched from an actionHandlers table keyed by
router.Intent. applyAction itself is now just the dec.Clarify guard plus a
table lookup.

No behaviour change: same reply strings, same side-effect order, comments
moved verbatim. The destructive-act confirm gate and the enabled-tool
allowlist stay entirely inside actionAct, exactly where they lived in the
old switch's IntentAct case — they're act-specific, not cross-cutting, so
they don't move to a separate layer. dec.Clarify short-circuit, dialogue
bookkeeping and detectPattern stay outside the table since they run
regardless of intent.

voice.go: 1638 -> 1344 lines. New actions.go: 362 lines.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:37:33 +04:00
kami 029449eefa Table-drive the IPC dispatcher instead of a 42-arm switch
dispatch() replaces the hand-written switch with a package-level
map[Method]handlerFunc built once at init. Each entry is one
withParams/withParamsVoid/withoutParams call closing only over the
CoreAPI method it invokes — adding a method is now one table line
instead of a new arm.

Check still runs once at the top before any unmarshal, unchanged. The
three non-CoreAPI methods (assert_stepup, store_encryption_key, unlock)
are special-cased before the table lookup since they drive Server
fields (StepUp/WrapKeyFn/UnlockFn), not store state. The current
CoreAPI is loaded once per dispatch and passed into the handler as an
argument, so SetAPI's runtime swap (the unlock transition) still takes
effect on the next request — the table itself never captures an api
value. No wire-format change; existing round-trip and unknown-method
tests in ipc_test.go pass unmodified.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:36:46 +04:00
kami ac36216f5d Merge branch 'worktree-agent-a517419e93e6a219f' into integration/small-batch 2026-07-31 23:31:24 +04:00
kami db17cfcc65 Delete the dead lockedAPI, add UnimplementedCoreAPI for the doubles 2026-07-31 23:31:24 +04:00
kami 7d676eb941 Stop tracking the mavwaked build artifact 2026-07-31 23:30:52 +04:00
kami fe3a4e9514 Merge branch 'worktree-agent-a9e5cef90b263a5e5' into integration/small-batch 2026-07-31 23:27:36 +04:00
kami a906f2afad Extract the pure RU/string/weather helpers out of voice.go 2026-07-31 23:27:08 +04:00
kami a2031a31d1 Record the measured confidence-gate numbers 2026-07-31 23:23:37 +04:00
kami 9b8bdf73cc Merge the five small-task branches 2026-07-31 23:11:26 +04:00
kami 7ad3c9a408 Merge branch 'worktree-agent-af88d63f65f30896b' into integration/small-batch 2026-07-31 23:11:16 +04:00
kami 84e1478823 Merge branch 'worktree-agent-af0fd9507d3e2ee46' into integration/small-batch 2026-07-31 23:11:16 +04:00
kami 9c8d0baffe Merge branch 'worktree-agent-a4cef2a815e32ebbf' into integration/small-batch 2026-07-31 23:11:16 +04:00
kami b0f5a16ec9 Add digest as a real outcome: suppressed care nudges get resurfaced, not lost
Vikunja #281. The interruption policy promised four outcomes — deliver_now,
queue, digest, drop — but only three existed: a care candidate the restraint
gate suppressed for quiet hours / away / calendar-busy simply vanished in
loop.Tick's `continue`, with only the trace remembering why.

internal/morning turned out not to be the natural drain: it's a fixed
Item/FactKey checklist engine, not a generic message bundler, so gate-
suppressed nudge text has nowhere to plug into its evidence model. Built a
parallel (but small, reusing the outbox's shape) durable digest instead:

- internal/store: digest_entries table + EnqueueDigestEntry (dedupes by
  rule+body, mirroring the delivery outbox's bodyHash), PendingDigestEntries,
  ExpireStaleDigestEntries, DrainDigestEntries (mark, never delete — an
  audit trail of what she actually said).
- internal/loop: DigestEligible(severity, blockedBy) is the pure boundary —
  only genuine restraint blocks (quiet_hours/calendar_busy/presence) even
  qualify (cooldown/snooze are not "suppression"); within care, Sev2 (break)
  digests, Sev1 (water/meal — stale by the time anyone could resurface them)
  drops. High severity never digests; alarms bypass the gate and deliver
  unchanged, on purpose.
- cmd/mavend/tick.go: each tick scans ExplainTick's trace for eligible
  blocked candidates, enqueues them, sweeps stale entries (24h expiry — the
  care rules are daily-cadence, so anything older is describing a day
  that's over), and drains the bundle only once the suppression reason has
  actually cleared, capped at 3 spoken items plus a trailing count so a
  digest can't turn into the exact nagging it was built to avoid.

Tests: store-level round-trip/restart-survival/dedupe/expiry/drain, loop-
level severity-boundary unit tests, and tick-level integration tests for
the drain-only-when-clear and never-digest-high-severity behavior.
2026-07-31 23:09:22 +04:00
kami 67563ed1f6 Run pattern detection from the digestion tick, not just voice (#43)
detectPattern only ever fired as a side effect of a voice fact-write, so a
recurring pattern already sitting in history went unnoticed until he
happened to mention it again by voice — the opposite of proactive.

Split the pipeline: extraction (fact -> normalized event) stays where a fact
is written, in voice.go, since it's tied to that write regardless of who's
talking. Detection (events -> stable pattern -> proposed_routines row) moves
into shared code (patterns.go's detectAndPropose) that both the voice path
and the new tick.go:detectPatterns call. The tick runs it every cycle over
every action+object pair on record (store.DistinctEventPairs, added), so a
pattern gets noticed on the daemon's own schedule.

Idempotence and the dismiss-must-stick requirement turned out to already be
handled by the store, not something the tick needs to reinvent:
proposed_routines has UNIQUE(action, object) and CreateProposedRoutine does
ON CONFLICT DO NOTHING, and DismissProposedRoutine flips status in place
without deleting the row. So a pair already proposed, accepted, OR
dismissed is a silent no-op on every later tick — a dismissed pattern can
never resurface, and re-running the scan never spams the /routines page.
Kept the voice-path call (immediate spoken confirmation is a nice feature
UX-wise and is now redundant-but-harmless with the tick, since both paths
share the same guarded detectAndPropose).

Tick-side detection only ever writes a row; it does not notify, ring, or
speak, keeping Maven "not a nag, not autonomous" — the /routines page is
still the only place a proposal becomes visible, and only accepting it
starts producing nudges (fireAcceptedRoutines).

Also fixed the stale vikunja#46 reference in proposed_routines.go — the
TODO it named is what this commit does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:07:33 +04:00
kami f0f7ebc9b2 Give LLM-routed decisions a real confidence so clarify can fire (#359)
Confidence was hardcoded to 1.0 for every LLM decision, and the LLM branch
in Router.Route returned straight from fillSlots without ever touching the
stage-3 threshold gate — so the LLM path could not produce a Clarify no
matter what confidence a model reported. That is why all 6 want_clarify
cases in the 77-case RU fixture were missed by every model in the bake-off.

Fix reads structural signal instead of changing the (parity-locked) router
prompt: a single-token utterance ("вода", "бэкап") is flagged thin evidence
in llmrouter.go; a fact left keyless or an act that never resolves to an
allowlisted fn, checked after fillSlots so the deterministic parsers get
first crack, is flagged in router.go's new gateLLMDecision. Anything below
config.DefaultRouterThreshold (0.55) now sets Clarify=true through the same
path the classifier already uses.

Added unit tests with a stubbed Completer proving both directions: thin
cases clarify, clean multi-word/resolved-slot cases stay confident. The
77-case fixture re-run against a live llama-server is still needed to
confirm the 6/6 moves — not done here, no llama-server on this box.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:07:32 +04:00
kami d12de589a2 mavweb: default -addr to loopback, not all interfaces
PR #47 added two state-changing routes (POST /chat, POST /routines)
behind the -addr flag, which defaulted to ":9200" (all interfaces).
Default now binds 127.0.0.1:9200; anyone who wants LAN/wider exposure
still passes an explicit bind (as deploy/docker-compose.yml already
does with "-addr :9201" inside the container, unaffected by this
default change).

Vikunja #317.
2026-07-31 23:03:45 +04:00
kami 50cc17f33a Lock down deploy/ecosystem/nginx.conf template to match the live host
The template said "drop into your nginx sites" but listened on the
wildcard `listen 80;` with no allow/deny ACL, unlike the actual deployed
hexis.kvmx.ru config which binds only to the WireGuard (10.42.0.1) and
LAN (192.168.1.104) addresses with allow/deny all. Anyone following the
template as written would expose these unauthenticated admin UIs to the
open internet.

Bind explicitly to those two addresses and add the matching ACL block,
mirroring cmd/mavweb/nginx.conf which already does this correctly.
Added a comment naming both addresses as host-specific so a deploy on a
different box swaps the IPs instead of reverting to `listen 80` when the
bind fails.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:01:58 +04:00
kami 73d13f1ea6 Merge pull request 'Stop the docs claiming the LLM router is off' (#49) from docs/fix-drift into master 2026-07-31 20:46:57 +02:00
kami 0ca5748699 Stop the docs claiming the LLM router is off
CLAUDE.md's routing section said "llmrouter is wired nil" and called the
classifier cascade the committed default. That stopped being true when the
integration merge landed: voice.go:214 wires pickLLMRouter, DefaultLLMRouter is
on, and deploy/mavend.json sets llm_router true. It is the first thing anyone
reads before touching the router, so it was pointing the next reader at a
wiring job that is already done.

Rewritten to say the LLM router is the default, the classifier is the failure
floor and must not be deleted, and what the two actually measure — 36.8% at
p50 31ms against 67.5%/72.7% at p50 ~2.7s, a trade accepted on purpose. Names
the one thing still open on that path: Confidence is hardcoded 1.0 in
llmrouter.go, so the LLM never asks for clarification (#359).

Also in CLAUDE.md: the persona line pointed at a memory file that does not
exist, so the actual rule was nowhere in the repo. Written out instead —
feminine self-reference, informal singular address, pet names forbidden but his
name allowed — plus the three eval checks that enforce it.

MODEL-BAKEOFF: three claims had gone stale within hours of being written. There
IS a make eval-models target now; the routing numbers ARE the production path,
not a bench artifact waiting on a wiring change; and the truncated 293 MB gguf
is deleted. Struck through rather than removed, since the caveats are part of
how the evening read at the time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 22:37:15 +04:00
kami b43bb265b5 Say it the way she would actually say it (#48) 2026-07-31 20:17:27 +02:00
kami b9a24334ea Say it the way she'd actually say it
Wording fixes from the review of the clarify + phrasing PRs.

- "На когда напомнить?" → "Когда?". After she has just been asked something,
  the long form is the phrasing of a form field, not of a person.
- A reminder now wants a subject as well as a time. "напомни в 11" had a time
  and nothing to say at 11, and she asked nothing at all — she now asks
  "О чём напомнить?". Subject first, since a reminder with no subject is not
  worth setting.
- The expiry notice is five phrasings picked at random instead of one fixed
  sentence. It is the line he hears every time he walks off mid-request, so it
  is the line that repeats most.
- The nudge prompt's ban on "обращения" is now "ласковые обращения". It was
  meant to forbid "милый"/"дорогой", not his name — "Ками, ноутбук на трёх
  процентах" is how she talks, and the eval's cringe check already only flags
  pet names.
- The nudge example no longer claims she plugged the laptop in. She has no
  hands and no smart plug; an example where she acts teaches the model to
  invent actions Maven never took.
- replySystem: "тепло" → "спокойно и без официальных формулировок". A one-word
  mood instruction a 1.7B can't act on, replaced with the behaviour meant.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 21:52:12 +04:00
kami c97aebf55a Merge tonight's work: all 46 reviewed PRs as one verified branch
135 commits. make build produces all 8 binaries; make test exits 0 across 38 packages with no failures, no data races, gofmt and vet clean.

See PR #47 for what had to be fixed to make it build as a unit.
2026-07-31 19:41:52 +02:00
kami 891136c65d gofmt the kiwix client and rewrite test
PR #41 and #44 landed these two files unformatted, so the gofmt gate that
PR #12 added to `make test` failed as soon as both were on one branch.
Struct-tag and comment alignment only, no semantic change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 21:34:03 +04:00
kami 41c7c13f42 Merge remote-tracking branch 'origin/overnight/snooze-works' into integration/jul31
# Conflicts:
#	internal/store/migrations.go
2026-07-31 21:32:18 +04:00
kami a324e8f624 Merge remote-tracking branch 'origin/overnight/eval-writeup' into integration/jul31 2026-07-31 21:31:36 +04:00
kami 51805e7f35 Merge remote-tracking branch 'origin/overnight/kiwix-rewrite' into integration/jul31 2026-07-31 21:31:36 +04:00
kami 533f0acda8 Lead the bake-off with the answer, not the superseded one
The file ran two sweeps and the second one changed the resident model, but
the lede still opened with "Recommendation: keep Qwen3.5-0.8B". Anyone
landing on the file read the wrong conclusion and had to scroll 100 lines
to find that it had been replaced — and it contradicted CLAUDE.md, which
already says the resident model is Qwen3-1.7B.

Both sweeps are accurate, so nothing is rewritten. The lede now states the
outcome and the first sweep's verdict is scoped to what it actually tested:
it rejects LFM2.5-1.2B, which still holds. It never was a case for keeping
0.8B as the resident model.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 21:17:07 +04:00
kami 4f59ba78c6 Write down the five-model sweep and why the 1.7B won
Numbers behind the resident-model change, plus the answer to "could a 230-350M
model do this instead" — no, and the reason is worth keeping: LFM2.5's published
instruction-following scores beat Qwen3.5-0.8B, and every one of those benchmarks
except Multi-IF is English. In Russian the 350M invents non-words and the 230M
answers in Spanish.

Also fills the row TALK-EVAL-31-07-2026.md had to void for contamination, and
corrects a wrong call I nearly made: the 1.7B's 16s p95 looked like the reasoning
trace, but the 0.8B sits at 17s in every run and the 1.7B beat it twice out of
three. The long tail is shared and is not the Thinking block.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 19:08:37 +04:00
kami d0afd9d4f6 Make Qwen3-1.7B the resident model
Stock Qwen3-1.7B, not the CPT'd one — that training is still running. It won
on both fixtures we have, measured tonight on an otherwise idle box:

  routing, 77 RU cases, intent-only:  67.5%  vs  59.7%  for Qwen3.5-0.8B
  talk fixture, 27 cases:             20/27  vs  11-17/27

It also beat Qwen3.5-2B, which is 20% larger, on every routing column.

Two other things came with it:

n_ctx goes 2048 -> 4096. This is a Thinking variant, so reasoning tokens need
the room, and 4096 is the context every score above was measured at. Shipping
2048 would ship something nobody measured.

The doc now says not to bother with sub-500M models, because I checked and they
are not close. LFM2.5-350M routes at 5.2% — worse than guessing among 7 intents
— and answers "столица Франции?" with "Сторзит", which is not a word. The 230M
replies to Russian in Spanish. Their published IFEval and BFCL numbers are good
and they are all English.

Note the routing gain needs the LLM router actually wired on to show up. It is
still nil, so this commit buys the phrasing improvement today and the routing
improvement when that lands.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:58:20 +04:00
kami 6b67e6f3c2 Word nudges from templates by default, model optional
DEPRECATION, flagged not asked: LLM-phrased nudges are no longer the default.
LLMPhraser.PhraseNudge now returns a hand-written Russian template. The model
still phrases chat, queries and reminders — only nudges moved.

Why: measured over many runs, Qwen3.5-0.8B wrote formal "вы" and plural
imperatives, used masculine self-reference, and invented facts and units
(90-95 seconds to boil an egg). A nudge is five words of known content, so
generation buys nothing and risks the persona every time. Templates score
15/15 on the nudge fixture, the model 11-13/15.

Nothing is deleted: the prompt, the fallbacks and the whole LLM nudge path
stay. Set phraser.llm_nudges = true in deploy/mavend.json to get them back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:21:52 +04:00
kami 742b2ad1d7 Score the Russian-to-keywords rewrite end to end (#403)
Same 9 cases as the retrieval eval, so the numbers compare directly:
hand-written keywords hit 8 of 8, this is what the model reaches on its
own. Reports the hand-written query next to the model's for every case,
because where the phrasing differs is the useful part.

Opt-in on MAVEN_KIWIX_URL + MAVEN_LLM_URL, like the other evals.

Result on Qwen3.5-0.8B: 3 of 8, identical on all three runs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:19:56 +04:00
kami b300ac5c70 Rewrite a Russian question into English Kiwix keywords (#403)
Kiwix ranks by keyword, not meaning, so a translated question finds song
and TV titles. This asks the resident model for the TOPIC instead: a short
English noun phrase, like a Wikipedia article title.

Locked down three ways, because a wrong query is silently wrong:
- A GBNF grammar, same idea as routeGrammar and responseGrammar. The
  reply must be {"query":"..."} with Latin words only. The JSON wrapper
  matters: this model always thinks out loud and this llama-server build
  ignores the thinking switch, so a bare word-list grammar just captured
  "Let me analyze this request carefully" for every question.
- max_tokens 32, since the answer is a few words.
- CleanQuery, which throws away empty, Russian and prose replies rather
  than passing them to Kiwix, and drops question words like "why" and
  "how much" that a keyword ranker cannot use anyway.

Client side only. Nothing is wired into the daemon or any config.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:19:42 +04:00
kami 13e5170e9e Hand-written Russian nudge templates plus a picker
Nudge wording as data instead of generation. The wording lives in
internal/phraser/nudges_ru_v1.json (embedded), about 10 variants per rule:
water, meal, break, service_down, netdata_critical, routine:, morning:, plus
a contentless default. That JSON is long because it is data — the owner can
edit any line of Russian without touching Go.

The picker:
- random, but never the same variant twice in a row for the same rule
- deterministic when seeded (math/rand with an injectable source)
- fills {since} / {service} / {what} from the candidate, and skips any variant
  whose value is missing, so no raw placeholder can reach the piper voice
- {since} is spelled out in words ("полтора часа", "семь часов"), because
  "3 ч" is wrong in a Russian voice

Scores 15/15 on the existing nudge fixture, on every seed swept. Nothing is
wired yet — that is the next commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:18:55 +04:00
kami 0b90952e55 Write down every conversational eval score from tonight
Records all four configurations on the 27-case talk fixture, three runs
each: no grammar, plus grammar, plus Russian prompts, plus the truncation
fix. Composite, per-path and per-check, with the reproduce command.

The short version is that the plumbing got fixed and the score barely
moved. Grammar was the real win. Russian prompts helped a little and cut
latency by 5x. The truncation fix was necessary and bought nothing.

Also writes down three things that are easy to lose:

- The truncation cause was the grammar's 400-character bound, not the
  token cap. Measured at three caps, same 400 characters every time.
- Then I set the bound to 1000 against a 768-token cap and made it worse.
  The two limits have to agree.
- One run is contaminated and marked void: I ran an agent against the same
  llama-server, and the report still claimed zero errors while a third of
  the fixture silently answered "не знаю.". That is #397 and it is worse
  than filed — a busy server is indistinguishable from bad phrasing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:18:19 +04:00
kami aa8f5b2ee2 Make the nonempty check look for actual words
It scored 27/27 on a run where two replies were "{" and "{\n  \"". It only
tested that the string was not blank, so punctuation counted as content and
the worst replies of the run passed the first check.

Now a reply needs at least one letter, Cyrillic or Latin. Latin counts
because answers about ssd or vpn are legitimately part English.

Digits alone fail too. The same run answered "сколько варить яйцо
вкрутую?" with "15-16" — no unit, no words, and the wrong number as well.
That is not something she said.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:57:39 +04:00
kami d7cdcb63bd Stop shipping half-written JSON as a reply
Two bugs, one symptom. A run of the talk eval produced replies that were
literally "{" and "{\n  \"" — those strings went out as things Maven said.

First bug: the parser could not tell "the model answered in plain prose"
from "the model started a JSON object and got cut off". Both came back as
empty, and every caller then shipped the raw text. Now an unfinished object
returns an error and each caller uses its own fallback instead. Bare prose
with no JSON in it still passes through, because small models do sometimes
answer that way and the reply is fine.

Second bug, and the actual cause: the grammar capped the response field at
400 characters. I measured it against Qwen3.5-0.8B at three different token
caps — 256, 768 and 2048 — and the reply came back exactly 400 characters
every time, cut mid-word. So the token limit was never what stopped it.
The bound is 1000 now, about six Russian sentences, still low enough to cut
off a repetition loop.

Token caps go from 256 to 768 on the chat and query paths so 1000
characters of Russian actually fits. The nudge path keeps its own cap; a
nudge is meant to be one sentence.

Note: cmd/mavend/replier_llm.go has its own copy of this parser with the
same bug. Left alone here so this commit stays small — that duplicate is
Vikunja #396.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:56:55 +04:00
kami ddb658ffbb Add a Kiwix search client and score retrieval (Vikunja #403)
Step one of letting Maven read instead of recall. No LLM yet.

internal/kiwix/client.go: search a local Kiwix server, parse the RSS
reply, hand back title + path + plain-text snippet + word count. The
snippet is the unit of context; a full article is ~100KB of HTML and
will not fit a 4096 token window.

internal/kiwix/retrieval_eval.go plus knowledge_v1.json: the 9 knowledge
questions from the phrasing fixture, each with hand-written English
keywords, scored on whether a wanted article comes back in the top 5.
Opt-in via MAVEN_KIWIX_URL, since CI has no Kiwix. No pass bar, the
number is the finding.

Result on the live mirror: 8/8 answerable questions hit, 7 of them at
rank 1. Retrieval works. Keywords are written by hand on purpose, since
Kiwix ranks by keyword and not by meaning, so a natural question fails.
A query-rewrite step is the next piece of work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:43:44 +04:00
kami c7dadc97d9 Write the chat and notes prompts in Russian
The reply has to be Russian, but two of the phrasing prompts told her
what to do in English. Both are Russian now, in the same style as the
nudge prompt that already works better.

Also dropped the "you are maven, a self-hosted personal assistant"
line from both. The persona block right above it already says who she
is, so it was said twice.

The JSON part is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:38:36 +04:00
kami c9d88c152e Drop "never phones home" as a hard rule
The owner's call, 2026-07-31: a 0.8B model does not know enough about the
world to be useful without reading something. So she may now read external
sources to answer world questions.

What replaces the old rule, in all three docs:

- No telemetry, no cloud model, no third-party account. Unchanged.
- Local first: the Kiwix ZIMs on the box before anything on the network.
- External search is allowed but off unless configured, same as weather
  and telegram.
- His notes and facts are never search input. Only the utterance goes out
  — never the persona block, the history, or matched notes.

Docs only, no code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:34:53 +04:00
kami 1890ff5d5d Constrain the phrasing output with a GBNF grammar
The 0.8B answered about one chat turn in three with open reasoning as plain text, so no JSON ever closed and the fallback shipped "Thinking Process:" to the user. A grammar makes that output impossible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:18:06 +04:00
kami 0110e9bc8c Report every address break, and stop -те verbs blinding the check
From a real reply in a nudge eval run: "Смотрите на его потребление
воды" is a plural imperative AND third person about him. Only the plural
printed.

Two separate faults. The check returned on its first hit, so the second
break stayed invisible and the failure read as milder than it was; it now
joins them. And "его" was not detected at all — looksVerb knows the
-й/-йте imperative but not the -те plural, so "смотрите" counted as the
person being talked about, which is what an antecedent means here.
pluralVerb already knows that form, so the antecedent test uses it too.

Third time a verb form has blinded this check. A fourth means it wants a
morphology table rather than another suffix.
2026-07-31 16:52:16 +04:00
kami 50ca8c8b5a Score the chat, query and knowledge phrasing paths (#395)
The phrasing fixture was 15 nudge cases, so every prompt change we
measured only told us about nudges. But the shared context block sits in
front of five prompts, and three of them — chat, note query, general
knowledge — had no scorer at all. Those are the long free-form replies,
where a persona break is most likely and where nothing could see one.

27 cases, nine per path. Nine rather than five because the nudge fixture
already cannot resolve a change smaller than about three cases, and a
per-path score off five would be worse.

Reuses the persona checks instead of copying them. Length, mood and
"no questions" are left out on purpose: these paths return no mood, and
a follow-up question is a feature in chat, not a fault.

The run refuses to score unless the model answers before and after it.
PhraseChat and PhraseQuery swallow model errors and return a canned
string, so without that guard a dead server produces a full report with
zero errors and a bad score — which reads as bad phrasing rather than as
nothing measured. Vikunja #397 is the real fix.
2026-07-31 16:51:52 +04:00
kami de09471421 Merge the shared prompt context block 2026-07-31 16:07:35 +04:00
kami d65c16a567 Don't tell her she can't talk
The block listed what she can do and ended with "nothing else". It sits
in front of the chat and general-knowledge prompts too, so that told her
to refuse the exact thing those prompts are for. Talking is now first in
the list, and the closing line limits ACTIONS rather than everything.

Also dropped the self-introduction from the knowledge prompt. It said
"Мавена, персональный ассистент" — a different name and a masculine
noun, right after the block says she is Maven and feminine. Identity
lives in the block now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 16:07:35 +04:00
kami 062d4252ef Tell her what she can actually do
The context block now lists her real capabilities: reminders, notes and
facts (write and recall), and the calendar — all three are code paths in
mavend today. Weather, telegram and shell acts are listed only when the
config actually has them, because offering something she cannot do is
worse than staying quiet about it.

Also drops the pronouns from the optional name/city line. The block's
own "ты" is Maven, so "тебя зовут" read as her name and "его" would have
shown her the third-person form she must never use about him. They are
plain labels now.

Vikunja #394.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 15:57:58 +04:00
kami 2c27e2ce1f Give every prompt one shared context block
The "address him as ты" rule had only reached two of the five system
prompts. Instead of pasting it into the other three (five copies drift —
that is how this happened), there is now one block, in internal/persona,
prepended to all five: nudges, action replies, chat, note queries and
general knowledge.

The block says who he is and how to address him (a man, always "ты",
never "вы", never "он" about him; Maven stays feminine), plus the
current local date and time. It is rendered fresh each turn because the
time changes, and it is correct with an empty config — the address and
gender rules are defaults in code. Config only adds optional facts:
owner_name, city, and the existing free-text `persona` string, which is
now the static half of the block.

Russian even in front of the English prompts: the rules are Russian
grammar, so they read best stated in Russian, and there is one copy.

Vikunja #394.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 15:55:30 +04:00
kami ccc5cba2a3 Merge the example-led prompt finding 2026-07-31 15:41:01 +04:00
kami 89d83c0b11 Record the example-led nudge prompt experiment (#393) — it made things worse
Tried rewriting the nudge prompt to lead with five on-topic examples instead
of rules. Three eval runs each side: before 12/13/14 of 15, after 11/12/11.
The loss is all in the address check — formal "вы" and plural imperatives came
back once the "говоришь на ты" rule stopped being its own sentence, and the
on-topic examples leaked their wording into the wrong cases.

Prompt reverted. Only the finding is committed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 15:40:18 +04:00
kami a97f554802 Merge the informal address prompt rule 2026-07-31 14:54:20 +04:00
kami f4de2fc5e1 Don't let a verb count as the person being talked about
The third-person check asks whether anyone else was named before "он".
A nudge is mostly verbs, and they were counted as possible people, so
"попробуй встать и отдохнуть — у него есть перерыв" passed. Infinitives
and imperatives now join past tense as words that cannot be a person.

A plain noun before the pronoun still blinds it. That needs a parser,
and the comment says so.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:54:20 +04:00
kami eef5d4da4f Tell the phraser to speak to him informally, singular
The prompts stated the feminine self-reference rule but never said whom she is
speaking to, so the model produced formal plural ("Жду вас") and talked about
him in third person ("Он не ел 11 дней"). Adds the address rule right next to
the feminine one, in the nudge prompt and the confirmation prompt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:52:23 +04:00
kami 09f1696fce Merge the eval label and kill script fixes 2026-07-31 14:32:45 +04:00
kami 80f7322294 Don't fail when docker confirms nothing is running
"Nothing on the host" meant two different things and the script treated
them the same. If docker answers and names no running containers, Maven
really is down and the script should say so and exit 0. Only when docker
cannot be asked is the answer unknown, and that is the case that must
fail loudly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:32:45 +04:00
kami fa5aebfbe4 Merge the delivery boundary fixes 2026-07-31 14:30:54 +04:00
kami 59cec63da1 List the columns in the table rebuild
The migration copied rows with SELECT *, which matches columns by
position. It is correct today, but if the old table's order ever
differed it would shuffle every row instead of failing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:30:54 +04:00
kami 02e8786695 Stop the containers instead of claiming success (#380)
docker-compose.yml has no 'pid: host', so each container has its own PID
namespace and pkill on the host matches nothing inside them. The script
then printed "All services gracefully stopped" while mavend, its
llama-server and the rest were still running.

Now it checks for running compose containers first and stops them with
docker compose. If it cannot ask docker and finds nothing to kill on the
host, or anything survives the kill, it says so and exits non-zero
instead of claiming success. The bare-metal path is unchanged apart from
verifying the SIGKILL actually worked, and no longer risks killing the
shell it was launched from.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:30:51 +04:00
kami 0272dc9d89 Record a suppressed care nudge instead of dropping it silently (#370)
Dropping a sev1-2 care nudge while you're away is right and still happens.
But it was a bare `continue`: no row, no log, so "she dropped it", "the gate
suppressed it" and "the rule never fired" all looked identical afterwards.

Adds a 'dropped' delivery status (migration #12 widens the CHECK constraint;
sqlite can't do that in place, so the table is rebuilt) and records the drop
as one delivery_attempts row plus a log line.

No nudges row for a drop: that table feeds the ignored_rate signal, and a
nudge nobody could see must not count as ignored.

TestVoiceNoSessionFallthroughLeavesOutboxTrail expected exactly one row for
sev1-2 when voice had no session. It now expects the voice failure plus the
drop, which is the point of the change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:27:47 +04:00
kami 2ad7635501 Merge the address-form eval check 2026-07-31 14:27:16 +04:00
kami 9949b309b1 Don't let a time word blind the third-person check
The check asks whether anyone else was named before "он". Time words
were not stoplisted, so "сегодня он не ел" read "сегодня" as the person
being talked about and passed — which is the recorded break with a word
in front of it, and nudges open with those words constantly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:27:08 +04:00
kami a788ca3915 Label eval runs with the model the server actually loaded (#379)
The phrasing eval printed "llm (0.8B, ...)" no matter which gguf
llama-server had loaded, so two runs of two different models came out
named the same and were easy to mix up when comparing.

It now asks llama-server over /v1/models, same as the router eval
already did. The helper moved to internal/llm so both share it, and it
now errors instead of returning a blank name when the id field is
missing — an unreachable server gets labelled "unknown-model", never a
plausible-looking guess.

Both eval paths stay opt-in behind MAVEN_LLM_URL; no server needed for
go test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:25:28 +04:00
kami 62d47d28ac Add an eval check for formal and third-person address (#384)
The phrasing run produced two persona breaks that scored clean:
"Приходите… Жду вас" (formal plural) and "Он не ел 11 дней" (talks
about him instead of to him). She is feminine, he is male, and she
speaks to him informally, one to one.

The new `address` check flags the "вы" family, plural imperative
endings, and a third-person "он" with no other subject named earlier in
the message. Like `hisgender` it is a keyword/suffix heuristic, not a
parser, and it prints the word it tripped on so a false alarm is easy to
dismiss. Limits are written out in the comment.

Both recorded strings are pinned as unit tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:25:23 +04:00
kami e9ff2c4912 Never send the full nudge body off-box (#368)
The away sinks fell back to the whole Body when Summary was empty. ntfy and
telegram leave the box, and the 0.8B phraser drops fields regularly, so that
fallback could push full detail off the machine.

The dispatcher already strips detail from away sendables. This exports that
one rule as delivery.AwayMessage and has both sinks use it, so a sink can't
leak the body on its own either: empty Summary means a generic line plus the
rule name, never the body.

The two sink tests named TestSendFallsBackToBodyWhenSummaryEmpty asserted the
old, wrong behaviour, so they are rewritten to assert the generic line.
TestSendRejectsEmptyMessage is likewise replaced: an away message can no
longer be empty, so the sink has nothing left to reject.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:23:54 +04:00
kami ee3e6a9eaf Wire the snooze read into the Gatherer and honour it for reminders (#364)
The Gatherer now fills State.SnoozeUntil from store.SnoozedUntil instead
of nil, so a snooze finally reaches the gate. RemindDecisions gains the
one restraint check that applies to a reminder — quiet hours, presence
and cooldown are still bypassed, so "wake me 7" is unchanged. Reviewer:
the two tests in internal/loop/gate_test.go are the contract.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:33:51 +04:00
kami 8acb8a97c6 Read the recorded snooze outcomes back out of the nudges table (#364)
The gate honours State.SnoozeUntil but nothing ever filled it. New
store.SnoozedUntil returns, per rule, when the newest snooze runs out.
Reviewer: the fixed 2h SnoozeDuration and its reasoning in nudges.go —
nothing upstream can supply a per-nudge length, so no new column.
Expired snoozes are dropped in SQL, so silence can never be permanent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:33:42 +04:00
139 changed files with 14044 additions and 2755 deletions
+1
View File
@@ -6,6 +6,7 @@
/mavweb /mavweb
/mavpoll /mavpoll
/mavcaldav /mavcaldav
/mavwaked
# Certs (private keys, don't commit) # Certs (private keys, don't commit)
certs/ certs/
+64 -13
View File
@@ -7,9 +7,20 @@ talking over unix sockets; one resident small model for routing + phrasing; whis
Deploy target is a Ryzen laptop (homesrv) with Vulkan offload to the Vega iGPU (`n_gpu_layers: 99`, Deploy target is a Ryzen laptop (homesrv) with Vulkan offload to the Vega iGPU (`n_gpu_layers: 99`,
compose passes `/dev/dri` + the render gid) — the resident model stays ≤1.7B either way. compose passes `/dev/dri` + the render gid) — the resident model stays ≤1.7B either way.
**Resident model:** currently **Qwen3.5-0.8B** (`Q4_K_M`), the smallest checkpoint in the gguf **Resident model:** currently **Qwen3-1.7B** (`UD-Q4_K_XL`), stock — not yet the CPT'd one.
library, picked for CPU/iGPU latency. The **target** is the locally CPT'd **Qwen3-1.7B**; that It replaced Qwen3.5-0.8B on 2026-07-31 because it measured better on both fixtures we have:
training is still in flight (Vikunja #122), so no such gguf exists yet. Model files live in 67.5% vs 59.7% intent-only on the 77-case RU routing fixture, and 20/27 vs 11-17/27 on the
talk fixture. See `MODEL-BAKEOFF-31-07-2026.md`. It is a Thinking variant, so `n_ctx` is 4096
— reasoning tokens need the room, and 4096 is what the scores above were measured at.
The **target** is still the locally CPT'd **Qwen3-1.7B** (Vikunja #122, training in flight).
Stock already speaks good Russian; what it gets wrong is the persona — it writes `я рад`,
masculine, where Maven needs `рада`. That is what the CPT is for.
**Do not bother with sub-500M models.** LFM2.5-230M and 350M were measured on 2026-07-31 and
both are unusable in Russian: the 350M routes at 5.2% (worse than guessing) and answers
"столица Франции?" with the invented non-word "Сторзит"; the 230M replies to Russian in
Spanish. Their strong published IFEval/BFCL numbers are English-only. Model files live in
`/mnt/hdd1/llms`, bind-mounted to `/opt/maven/models/llm` — which **shadows** the repo's `/mnt/hdd1/llms`, bind-mounted to `/opt/maven/models/llm` — which **shadows** the repo's
`models/llm/`, so the LFM2.5 gguf sitting there is not loaded by anything. Swapping the resident `models/llm/`, so the LFM2.5 gguf sitting there is not loaded by anything. Swapping the resident
model is a one-line change to `phraser.model_path` in `deploy/mavend.json`. model is a one-line change to `phraser.model_path` in `deploy/mavend.json`.
@@ -58,19 +69,41 @@ protocol; the config in `deploy/mavend.json` (with `${VAR}` env expansion from g
## Routing — read this before touching the router ## Routing — read this before touching the router
`internal/router/` has TWO layered engines and the committed default is an **interim `internal/router/` has TWO layered engines. **The LLM router is now the default and it is
stopgap, not the intended design** (see memory `routing-architecture-target`): on in deploy** — this section used to say it was wired `nil`, which stopped being true on
2026-07-31.
- **Target (REARCH.md):** LLM-as-router. One resident Qwen3-1.7B (`llmrouter.go`) emits - **LLM router (the intended design, REARCH.md):** the resident Qwen3-1.7B (`llmrouter.go`)
GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is demoted emits GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is
from a routing gate to a RAG hint. demoted from a routing gate to a RAG hint. Wired at `voice.go:214` via
- **Current stopgap:** `llmrouter` is wired `nil` (around `voice.go`), so the `pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient)`; the flag is `voice.llm_router`
`classifier.go` + `embedder.go` nearest-neighbour cascade actually runs. It routes by (`config.go`), `DefaultLLMRouter` is **on**, and `deploy/mavend.json` sets it `true`.
similarity to frozen seed phrases — the known cause of weak RU query handling. - **Classifier cascade (the failure floor, not dead code):** `classifier.go` +
`embedder.go` nearest-neighbour over frozen seed phrases. It runs when the LLM router is
off, when there is no llama-server to talk to (`pickLLMRouter` logs that and degrades),
and on any per-turn LLM error. Do not delete it — routing by seed similarity is the known
cause of weak RU query handling, but a turn must never break on the model.
Cascade order: `stage0.go` exact-match fast-path → LLM router (when non-nil) → classifier Cascade order: `stage0.go` exact-match fast-path → LLM router (when non-nil) → classifier
fallback. Any LLM error falls through to the classifier so a turn never breaks on the model. fallback. Any LLM error falls through to the classifier so a turn never breaks on the model.
Measured on the 77-case RU fixture (`MODEL-BAKEOFF-31-07-2026.md`): the classifier scores
36.8% full accuracy at p50 31ms; Qwen3-1.7B scores 67.5% intent-only / 72.7% through the
cascade at p50 ≈2.7s. Accuracy roughly doubled, latency is ~90× worse, and that trade was
accepted deliberately. `Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
#359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with
no allowlisted fn) feeding the same stage-3 gate the classifier path already had — see
`gateLLMDecision` in `router.go`. Note the second half of that bug: the LLM branch never
consulted `r.threshold` at all, so a correct low confidence would have been discarded anyway.
Re-measured on the fixture after the fix: **missed clarify 6/6 → 1**, at the cost of 3 false
clarifies and 2.6pt of full accuracy (72.7% → 70.1%, intent-only 67.5% → 74.0%). Two of the
three false clarifies are acts the model mis-routed and the gate caught — asking beats wrongly
executing, so the fixture and the daemon disagree about what is correct there. The third,
`"поужинал"`, is a real defect: **the single-token rule is an English intuition and does not
transfer to Russian**, where one word is routinely a whole sentence. Narrow or drop it.
## LLM output contract ## LLM output contract
All phrasing paths emit `{"response":"...","mood":"..."}` (parsed in `replier_llm.go` and All phrasing paths emit `{"response":"...","mood":"..."}` (parsed in `replier_llm.go` and
@@ -82,8 +115,26 @@ workspace enforces that the Go and relabelling prompts remain identical.
## Non-goals (hard constraints) ## Non-goals (hard constraints)
Never phones home. Not a nag, not autonomous. Maven's persona is **feminine** — Russian Not a nag, not autonomous. Maven's persona is **feminine** — Russian
self-reference must use feminine forms (the user is male; see memory `maven-persona-gender`). self-reference must use feminine forms `рада`, not `рад`; `поняла`, not `понял`. The owner
is male and is addressed informally: "ты", singular, never "вы"/"ваш" and never "он"/"его"
(she talks TO him, not about him). Pet names ("милый", "дорогой") are forbidden; his name
("Ками") is not. The eval enforces this: `CheckAddress`, `CheckFeminine` and `CheckCringe` in
`internal/phraser/eval/checks.go`, scored by `make eval-phrasing`.
**"Never phones home" is DEPRECATED** (owner's call, 2026-07-31). It used to be a hard
constraint and it is not one any more: a 0.8B — and a 1.7B — does not know enough to answer
world questions, so she needs to read external sources. What replaces it:
- **No telemetry, no cloud model, no third-party account.** That part never changes. Nothing
about Maven is reported to anyone, and inference stays on the box.
- **Local sources first.** Kiwix ZIMs on homesrv (Wikipedia, ifixit) before anything on the
network. Reading beats recalling for a small model, and a local read costs nothing.
- **External search is allowed and off unless configured**, like the weather and telegram
capabilities.
- **His notes and facts are never search input.** Looking up why the sky is blue and sending
his stored personal notes to an upstream engine are different acts. Only the utterance goes
out, never the persona block, history, or matched notes.
## Web UI conventions ## Web UI conventions
+10 -4
View File
@@ -15,7 +15,8 @@
**Maven** — self-hosted personal assistant. Manages your day, acts on your **Maven** — self-hosted personal assistant. Manages your day, acts on your
homelab. One daemon on homesrv (always-on, not the workstation), multiple homelab. One daemon on homesrv (always-on, not the workstation), multiple
client surfaces. All local, never phones home. client surfaces. Inference and data stay on the box; she may READ external
sources (see Non-goals — "never phones home" is deprecated).
Primary name is "Maven", with feminine-gendered Russian self-reference Primary name is "Maven", with feminine-gendered Russian self-reference
("она", "меня", "помогла"). Clients may choose their own UI label. Consistent ("она", "меня", "помогла"). Clients may choose their own UI label. Consistent
@@ -35,8 +36,13 @@ Inside boundary — the ones that actually constrain the build:
she records. A confident wrong fact is worse than a known gap. she records. A confident wrong fact is worse than a known gap.
- **Not a nag** — she'd rather miss a nudge than be mutable. Shuts up when - **Not a nag** — she'd rather miss a nudge than be mutable. Shuts up when
uncertain. Load-bearing. uncertain. Load-bearing.
- **Not a stranger** — runs on your stuff, your model, your data. Never - **Not a stranger** — runs on your stuff, your model, your data. No
phones home. telemetry, no cloud model, no third-party account. She may READ external
sources to answer world questions (Kiwix first, then optional search); she
never reports anything about you to anyone, and your notes and facts are
never used as search input. **"Never phones home" as an absolute is
deprecated** — owner's call, 2026-07-31: a small model does not know enough
to be useful without reading.
- **Not a relationship** — mom-tone is a function that makes nudges land, not - **Not a relationship** — mom-tone is a function that makes nudges land, not
emotional company. Names the drift a warm small model falls into. emotional company. Names the drift a warm small model falls into.
@@ -458,7 +464,7 @@ decides *insistence*. Both are needed.
sev ≤ 2 drops on away, sev ≥ 3 holds: a missed water nudge is noise, a missed sev ≤ 2 drops on away, sev ≥ 3 holds: a missed water nudge is noise, a missed
backup failure isn't. Away-channels (ntfy/telegram) leave the box — the one backup failure isn't. Away-channels (ntfy/telegram) leave the box — the one
path that crosses "never phones home," through your own relay. **Minimal path that leaves the box for a person to see, through your own relay. **Minimal
body** — "disk low on homesrv," not detail; don't make notifications a body** — "disk low on homesrv," not detail; don't make notifications a
shoulder-surf exfil surface. shoulder-surf exfil surface.
+126 -4
View File
@@ -1,15 +1,28 @@
# Resident model bake-off — 31-07-2026 # Resident model bake-off — 31-07-2026
**Recommendation: keep Qwen3.5-0.8B.** LFM2.5-1.2B is worse at routing (52.6% vs 60.5% **Outcome: the resident model is stock Qwen3-1.7B** (`UD-Q4_K_XL`). Two sweeps ran this
intent accuracy), and the loss is almost entirely Russian (18/61 vs 22/61 RU, while EN is a evening and the second one changed the answer — read to the end before acting on any table
wash). It is also 2.4× slower. The Thinking variant is far worse again. here. [Second sweep](#second-sweep-same-evening--five-models-and-a-resident-model-change)
is the one that holds.
## First sweep — LFM2.5-1.2B vs Qwen3.5-0.8B
**Verdict, scoped to this pair: keep Qwen3.5-0.8B over LFM2.5-1.2B.** LFM2.5-1.2B is worse
at routing (52.6% vs 60.5% intent accuracy), and the loss is almost entirely Russian
(18/61 vs 22/61 RU, while EN is a wash). It is also 2.4× slower. The Thinking variant is
far worse again. This verdict still stands as written — it rejects LFM2.5-1.2B. It is
**not** a recommendation to keep 0.8B as the resident model; the second sweep replaced it
with Qwen3-1.7B.
Settles Vikunja **#278 / #250**. Settles Vikunja **#278 / #250**.
- Same fixture and scorer as `ROUTING-EVAL-31-07-2026.md`: `internal/router/eval/` - Same fixture and scorer as `ROUTING-EVAL-31-07-2026.md`: `internal/router/eval/`
(`ru_routing_v1.json`, 76 held-out cases). (`ru_routing_v1.json`, 76 held-out cases).
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:<port> make eval-router` - Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:<port> make eval-router`
(`TestLLMRouterBaseline`). Note: there is no `make eval-models` target. (`TestLLMRouterBaseline`). (This line used to say there is no `make eval-models` target.
There is one now — start a server with the gguf you want, then
`make eval-models MAVEN_LLM_URL=http://127.0.0.1:<port>`. It runs only the LLM test, since
the classifier baselines do not depend on the model.)
- All three models served by the same `llama-server` flags — `-c 2048 -ngl 99 -t 6`, only - All three models served by the same `llama-server` flags — `-c 2048 -ngl 99 -t 6`, only
`-m` and `--port` differ. One server at a time on an otherwise idle box, so latencies are `-m` and `--port` differ. One server at a time on an otherwise idle box, so latencies are
real and not contention. real and not contention.
@@ -99,3 +112,112 @@ thinking trace costs time without buying accuracy on a short enum classification
Routing only. LFM2.5 might still phrase better, and phrasing is the resident model's other Routing only. LFM2.5 might still phrase better, and phrasing is the resident model's other
job — that needs its own fixture. But routing is the load-bearing path and Maven is job — that needs its own fixture. But routing is the load-bearing path and Maven is
Russian-first, so on the evidence here the switch is not worth making. Russian-first, so on the evidence here the switch is not worth making.
---
# Second sweep, same evening — five models, and a resident-model change
The sections above compared LFM2.5-1.2B against Qwen3.5-0.8B on routing and concluded
"the switch is not worth making". That still holds. This sweep asked a different
question — whether a *smaller* model could work, since LFM2.5's published
instruction-following scores beat Qwen3.5-0.8B badly — and answered it, plus found a
better resident model by accident.
**Outcome: the resident model is now stock Qwen3-1.7B.** Sub-500M is a dead end.
## Routing — 77 Russian cases, one run each
| model | on disk | llm-only (full) | llm-only (intent) | cascade + fallback |
|---|---|---|---|---|
| LFM2.5-230M-Q8_0 | 246 MB | 23.4% | 33.8% | 36.4% |
| LFM2.5-350M-Q8_0 | 379 MB | 2.6% | **5.2%** | 20.8% |
| Qwen3.5-0.8B-Q4_K_M | 527 MB | 36.4% | 59.7% | 61.0% |
| Qwen3.5-2B-UD-Q4_K_XL | 1.34 GB | 42.9% | 62.3% | 63.6% |
| **Qwen3-1.7B-UD-Q4_K_XL (stock)** | 1.13 GB | **44.2%** | **67.5%** | **72.7%** |
Qwen3-1.7B wins every column, including against a model 20% larger than it.
## Talk fixture — 27 cases, three runs each, idle box
| | Qwen3.5-0.8B | Qwen3-1.7B stock |
|---|---|---|
| composite | 13, 11, 8 | **20, 21, 18** |
| address | 21, 18, 18 | **26, 25, 23** |
| feminine | 27, 25, 26 | 26, 27, 26 |
| lang | 27, 27, 26 | 26, 27, 27 |
| ontopic | 16, 19, 19 | **22, 23, 23** |
| canned fallbacks | 8, 5, 6 | **0, 2, 0** |
This also fills the row `TALK-EVAL-31-07-2026.md` had to void for contamination:
**600ch/1024tok on Qwen3.5-0.8B scores 13, 11, 8.**
`address` is the headline. It sat at 18-22 of 27 on the 0.8B no matter how the prompt
was worded — the prompt explicitly forbids "вы" and the model writes `вашей`,
`подождите`, `делаете` anyway. That was read as "prompting is out of levers", and it
was really "0.8B is out of capacity". The 1.7B mostly holds the constraint.
The fallback column matters too: 5-8 of 27 turns on the 0.8B end in a hardcoded
`"не знаю."`, meaning it failed to emit parseable JSON about a quarter of the time.
The 1.7B does that 0-2 times.
## Latency — the long tail is not the Thinking block
| | p50 | p95 |
|---|---|---|
| Qwen3.5-0.8B | 2.4s, 2.9s, 2.0s | 17.4s, 17.6s, 17.4s |
| Qwen3-1.7B stock | 2.7s, 2.6s, 2.8s | 16.4s, 6.6s, 3.9s |
p50 is flat across a 2× size difference. The first instinct on seeing the 1.7B's
16s p95 was "that is the reasoning trace, cap it" — wrong. The 0.8B's p95 is a
consistent 17s and the 1.7B beat it in two of three runs. The tail is shared and
lives somewhere else. Do not spend time on `/no_think` on this evidence.
## Sub-500M: not close, and the benchmarks say otherwise for a reason
LFM2.5-350M publishes IFEval 76.96 against Qwen3.5-0.8B's 59.94, and BFCLv3 44.11
against 35.08 — better at instruction-following and structured output, at 2/3 the
size. Those numbers are real and they are **English**. Every benchmark in that
table except Multi-IF is English-only.
In Russian, with a 300-token budget and temperature 0:
- **350M**, «Столица Франции? Ответь кратко.» → *«Сторзит в Париже.»*`Сторзит` is
not a word; it is invented morphology.
- **350M**, asked to read back a reminder → a fortune cookie about being attentive
and confident. No reminder in it.
- **230M**, «Привет, как дела?» → answered **in Spanish**.
The 230M beating the 350M six-fold on routing (33.8% vs 5.2%) is the other tell:
when the larger sibling collapses like that it is format compliance failing, not
reasoning.
This is a pretraining gap, not a fine-tuning gap. Teaching Russian to a 350M from
near-zero is not an afternoon on a Colab, which was the premise worth checking.
## Why this vindicates the 1.7B CPT
Stock Qwen3-1.7B, untrained and unprompted, answers all three probes in fluent
correct Russian. What it gets wrong is the persona: *«Привет! Я рад, что ты здесь»*
`рад` is masculine and Maven needs `рада`. That is the right kind of remaining
problem, and it is exactly what the CPT (Vikunja #122) is for.
The 1.7B was the correct model choice. What was wrong was treating it as a
**blocker**: stock already beats what was deployed, so it ships now and gets
swapped again when the CPT lands.
## Caveats
- Routing is one run per model, not three. The gaps between families are far larger
than the run-to-run spread seen on the talk fixture, but the 2B-vs-1.7B gap (62.3
vs 67.5) is not safe to call on one run.
- ~~The routing numbers only reach production once the LLM router is wired on. It is
still `nil`.~~ **Resolved the same evening:** the LLM router is wired at `voice.go:214`
behind `voice.llm_router`, the default is on, and `deploy/mavend.json` sets it `true`.
These numbers are the production path now, so the p50 ≈2.7s is a real per-turn cost and
not a bench artifact.
- ~~`/mnt/hdd1/llms/LFM2.5/Qwen3-1.7B-UD-Q4_K_XL.gguf` is a 293 MB truncated download
in the wrong directory.~~ **Deleted 2026-07-31.** The good 1.13 GB copy in `qwen3/` is
what `deploy/mavend.json` loads.
- Harness: `scratchpad/bakeoff.sh`, one server at a time, health-checked before each
run, `/v1/models` recorded per run. Never run two LLM consumers at once — see the
contamination note in `TALK-EVAL-31-07-2026.md`.
+6 -3
View File
@@ -103,14 +103,17 @@ eval-router:
eval-recall: eval-recall:
MAVEN_ONNX_LIB="$(MAVEN_ONNX_LIB)" $(GO) test -v -count=1 ./internal/memory/recalleval/ MAVEN_ONNX_LIB="$(MAVEN_ONNX_LIB)" $(GO) test -v -count=1 ./internal/memory/recalleval/
# eval-phrasing -- score nudge phrasing (internal/phraser/eval). Verbose so the # eval-phrasing -- score nudge phrasing AND the conversational paths (chat,
# query, general knowledge) in internal/phraser/eval. Verbose so the
# report and every generated message land in the terminal. With no environment # report and every generated message land in the terminal. With no environment
# it scores the deterministic Stub only, which is what CI runs. Set # it scores the deterministic Stub only, which is what CI runs. Set
# MAVEN_LLM_URL to add the resident model: # MAVEN_LLM_URL to add the resident model:
# MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing # MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing
# The model run is slow (minutes) -- the timeout is raised to match. # The model run is slow (minutes) -- the timeout is raised to match. It covers
# two fixtures now (15 nudges + 27 conversational cases, and the chat replies are
# the long ones), hence 90m rather than 40m.
eval-phrasing: eval-phrasing:
$(GO) test -v -count=1 -timeout 40m ./internal/phraser/eval/ $(GO) test -v -count=1 -timeout 90m ./internal/phraser/eval/
# eval-models — score ONE llama-server against the same fixture, for the # eval-models — score ONE llama-server against the same fixture, for the
# resident-model bake-off (#278, #250). Start a server with the gguf you want, # resident-model bake-off (#278, #250). Start a server with the gguf you want,
+29
View File
@@ -103,6 +103,35 @@ but a large part of the jump is that failure now degrades into Russian instead o
The two remaining failures: one `"..."` recurrence (`routine-stretch`) and one meal nudge The two remaining failures: one `"..."` recurrence (`routine-stretch`) and one meal nudge
that never says food. that never says food.
## Tried and reverted: an example-led nudge prompt (#393)
The idea was that a 0.8B copies examples better than it follows rules, so the nudge prompt
was rewritten to lead with five on-topic examples (water, break, pills, morning, service) and
the prose rules were compressed to pay for the tokens: 1190 chars down to 986.
It measured **worse**, three runs each side, same llama-server, same fixture:
| run | before | after |
|---|---|---|
| 1 | 12/15 (address 14) | 11/15 (address 13) |
| 2 | 13/15 (address 15) | 12/15 (address 15) |
| 3 | 14/15 (address 15) | 11/15 (address 12) |
`feminine` and `hisgender` were 15/15 on all six runs, so they measure nothing here. The
regression is all in `address`: 44/45 before, 40/45 after. Formal "вы"/"ваше" and plural
imperatives came back, and so did `"..."`.
Two likely causes, both about the same thing — **examples do not carry a prohibition**. The
old prompt spent a whole sentence on «говоришь на "ты", в единственном числе»; the new one
demoted that to one item in a long "никогда" list, and the model stopped obeying it. And
making the examples on-topic let their *wording* leak: a break case came back as
«Вы давно не пили воду. Выпей стакан.» — the water example, verbatim, in the wrong slot.
That is exactly the failure the laundry/laptop examples were chosen to avoid.
Change reverted. What survives is the measurement: a rule the model must obey needs its own
sentence, and examples must stay off-topic. Also note the before side alone spans 1214 of
15 — this fixture cannot resolve anything smaller than about three cases.
## Broken, found, not fixed ## Broken, found, not fixed
1. ~~**`checkFeminine` only catches half the constraint.**~~ **Fixed** (#381). It scanned for 1. ~~**`checkFeminine` only catches half the constraint.**~~ **Fixed** (#381). It scanned for
+4 -1
View File
@@ -90,4 +90,7 @@ later* is the worker + RAG.
4. **Deferred work** — larger reasoner, custom Piper voice and other expansions. 4. **Deferred work** — larger reasoner, custom Piper voice and other expansions.
## Non-goals (unchanged) ## Non-goals (unchanged)
Never phones home. Not a nag. Not autonomous. Feminine-gendered RU self-ref. Not a nag. Not autonomous. Feminine-gendered RU self-ref. No telemetry, no
cloud model, no third-party account — but she MAY read external sources to
answer world questions (Kiwix first, search optional). "Never phones home" as
an absolute is deprecated, owner's call 2026-07-31; see CLAUDE.md § Non-goals.
+150
View File
@@ -0,0 +1,150 @@
# Conversational phrasing eval — 31-07-2026
Every score measured tonight, on the three paths the nudge eval never touched:
chat, query-with-notes, and general knowledge.
**Short version: the plumbing got fixed and the score barely moved.** Grammar and
Russian prompts together took the composite from ~9 to ~14 of 27. Everything
still failing is the model not knowing things or not holding a constraint, and
prompting is out of levers. Settles the measurement half of Vikunja #395 / #398 /
#400.
## How to reproduce
```sh
# llama-server: -c 4096 -ngl 99 -t 6, model /mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf
MAVEN_LLM_URL=http://127.0.0.1:18099 no_proxy=127.0.0.1,localhost \
deps/go/go/bin/go test -count=1 -timeout 40m \
-run TestLLMTalkBaseline ./internal/phraser/eval/ -v
```
Three runs per configuration, always. The fixture is 27 cases, so one reply
changing moves the composite by 3.7 points — a single run cannot tell a real
change from sampling noise. This was learned the expensive way: an earlier claim
that "one nudge case fails every run" turned out to be three different cases
across three runs.
**Run the box otherwise idle.** See the contamination note at the bottom.
## Composite, per configuration
| config | overall /27 | chat /9 | query /9 | knowledge /9 | canned fallbacks |
|---|---|---|---|---|---|
| baseline, no grammar | 7, 12, 7 | 1, 1, 0 | 2, 4, 2 | 4, 7, 5 | 0, 0, 0 |
| + GBNF grammar (#398) | 14, 15, 8 | 1, 3, 0 | 5, 6, 3 | 8, 6, 5 | 0, 0, 0 |
| + Russian prompts (#400) | 11, 17, 15 | 1, 5, 3 | 5, 6, 8 | 5, 6, 4 | 0, 0, 0 |
| + truncation fix, 1000ch/768tok | 12, 13, 10 | 2, 2, 1 | 7, 7, 5 | 3, 4, 4 | 3, 3, 6 |
| + rebalanced, 600ch/1024tok | **void — contaminated** | | | | |
"Canned fallbacks" counts replies that came back as the hardcoded `"не знаю."`
or `"поговорили."`. It is not a check, it is a health signal: those strings mean
the phraser gave up, and the eval scores them as ordinary bad replies.
## Per-check
| check | no grammar | + grammar | + RU prompts | + truncation fix |
|---|---|---|---|---|
| nonempty | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 |
| ellipsis | 20, 19, 23 | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 |
| lang | 13, 16, 15 | 23, 26, 26 | 25, 26, 25 | 26, 27, 27 |
| feminine | — | — | 25, 24, 26 | 25, 25, 27 |
| address | — | — | 21, 22, 22 | 22, 21, 22 |
| ontopic | — | — | 17, 24, 18 | 17, 19, 14 |
`nonempty` reading 27/27 everywhere is not good news — it was a broken check.
It tested for a non-blank string, so replies of literally `{` and `"15-16"`
passed it. Fixed on `overnight/fix-truncation`; it needs a letter now.
## What each change actually bought
**GBNF grammar (#398) — the biggest single win.** Qwen3.5-0.8B writes
`Thinking Process:` as plain text with no tags, `stripThink` only handles
`</think>`, so the JSON never closed and the plain-text fallback shipped the
literal reasoning. `ellipsis` went 20→27 and `lang` 13→26. The router had been
using a grammar for ages; the phraser asking nicely in the prompt was the
oversight.
**Russian prompts (#400) — modest, plus a large latency win.** Chat 1.3→3.0
average, query 4.7→6.3, knowledge 6.3→5.0. All inside the run-to-run spread, so
"probably better on the paths it targeted, not provable in three runs". p50
latency dropped from ~11.5s to ~2.3s and that part is consistent across all
three runs — shorter prompts, and she stopped emitting English reasoning first.
**Truncation fix — necessary, and did not help the score.** Two real bugs
(replies of `{`, and a `nonempty` check that passed them), both fixed, and the
composite went nowhere. A complete rambling wrong answer fails the same checks a
truncated one did. Worth doing anyway: the daemon was shipping `{` to a
text-to-speech voice.
## The truncation bug, since the cause was counter-intuitive
The grammar's `string ::= ... {0,400}` rule was the cause, not the token cap.
Measured against Qwen3.5-0.8B at three caps — 256, 768 and 2048 — the reply came
back **exactly 400 characters every time, cut mid-word** (`"Нужно записать и,"`).
Then I raised the bound to 1000 while the cap was 768 tokens and made it worse:
Russian runs ~1.5 characters per token here, so generation died on the *token*
cap instead, mid-object, and the new guard correctly refused it and shipped
`"не знаю."` — 3, 3 and 6 fallbacks per run, from zero. **The two limits have to
agree.** 600 characters needs ~400 tokens; the cap is 1024.
## Where the remaining failures live
`address` is stuck at 21-22 of 27 and `ontopic` at 14-19. Both resist prompting.
**The prompt now explicitly forbids exactly what she does.** It says never "вы",
use the singular — and she writes `вашей`, `подождите`, `делаете`, `хотите`,
`напишите`. Telling a 0.8B "never do X" does not work. Same for
`feminine`: `я готов`, `я понял`, `я нашел`, `я заметил`, `я сказал`.
**Some of `ontopic` is the fixture, not the model.** `chat-how-are-you` got
`"Привет! Я здесь, чтобы поговорить. Как дела сегодня?"` — a fine reply that
fails because `want_any` is `[норм, хорош, порядк, тут, работ]`. It fails in
every run, so it inflates the count. The `ontopic` column currently measures the
fixture as much as the model. Not fixed yet, deliberately: changing it would
break comparability with the runs above.
**Two replies worth reading, because they are not fixable by prompting:**
- Thunder and lightning: *"Скорость молнии — 8-10 тысяч километров в секунду, но
звук — 300 метров в секунду, что делает молнию громче."* Confidently wrong,
and it concludes lightning is *louder* rather than sound being *slower*.
- "расскажи обо мне": *"Ты — прекрасное существо, с душой и вниманием… Спасибо за
твою улыбку… О тебе — заповедь любви."* Sycophantic filler, zero information,
and precisely the "not a relationship" non-goal.
- Boiling an egg: `"15-16"` one run, `"1"` another. No unit, wrong number.
The first argues for reading instead of recalling (#403 — Kiwix retrieval scores
8/8 on the same questions given English keywords). The second and third argue
for templates on the paths where correctness matters (#392).
## Contamination note — how the last row got voided
I started the query-rewrite agent against the same llama-server the sweep was
using, and assumed contention would only affect latency. It did not. The
knowledge path collapsed to 0 of 9 with eight canned `"не знаю."` replies, p95
tripled to 23.7s, and **the report still said "0 errors"**.
That is Vikunja #397, and it is worse than filed: a merely *busy* server
produces a clean-looking report with a third of the fixture silently answering
`"не знаю."`. `PhraseChat` and `PhraseQuery` swallow every failure and return a
hardcoded string, so infrastructure trouble is indistinguishable from bad
phrasing in the score. The talk test guards the *start* and *end* of a run with
a model check, which catches a dead server but not a loaded one.
**Until #397 is fixed, treat any run made on a busy box as void.**
## Next
- Re-run 600ch/1024tok clean, to fill the void row.
- Score `Qwen3.5-2B-UD-Q4_K_XL` (already at `/mnt/hdd1/llms/qwen3.5/`, never
measured) on this fixture and the router fixture. Not the 4B — too big for
this box, owner's call.
- Newer sub-500M candidates (LFM2.5 200M/300M) are worth a run for routing.
Note `MODEL-BAKEOFF-31-07-2026.md` found LFM2.5-**1.2B** worse than
Qwen3.5-0.8B at Russian routing and 2.4× slower — but those are a different,
older generation, so that result does not predict the small ones.
- Fix `chat-how-are-you`'s `want_any`, and re-baseline once, so `ontopic`
measures the model.
- #397 first if anything, since it decides whether any of the above is
trustworthy.
+75 -144
View File
@@ -1,17 +1,24 @@
// mavcaldav — the CalDAV poller module. // mavcaldav — the CalDAV module: reads calendars into facts, and renders
// maven's own reminders back out to a calendar she owns.
// //
// Polls a Radicale (or any CalDAV) server for today's events and writes // READ side (unchanged behaviour): polls a Radicale (or any CalDAV) server for
// `facts (kind=env, source=poll:caldav)` through core's IPC socket. // today's events and writes `facts (kind=env, source=poll:caldav)` through
// Key-free, restart-free, fail-independent — crashes can't touch the // core's IPC socket. Key-free, restart-free, fail-independent — crashes can't
// store key, worst case a stale calendar_busy fact until the next poll. // touch the store key, worst case a stale calendar_busy fact until the next
// poll. Two facts:
// //
// Two facts written:
// - calendar_busy ("true"/"false") — read by the loop gate to suppress // - calendar_busy ("true"/"false") — read by the loop gate to suppress
// nudges during meetings // nudges during meetings
// - calendar_event ("<summary> @ <start>-<end>") — per-event for query // - calendar_event ("<summary> @ <start>-<end>") — per-event for query
// //
// Append-only discipline: a fact is written only when its value CHANGED // Append-only discipline: a fact is written only when its value CHANGED vs the
// vs the latest for that key+source. // latest for that key+source.
//
// RENDER side (Vikunja #127, off unless -render-url is given): publishes each
// pending reminder as a single-event iCal resource in a collection maven owns.
// The calendar is a view, sqlite is the store — see render.go. The render URL
// must differ from the read URL, checked at startup, so the render target can
// never be a calendar maven is only supposed to read.
package main package main
import ( import (
@@ -27,6 +34,7 @@ import (
"syscall" "syscall"
"time" "time"
"github.com/kami/maven/internal/calendar"
"github.com/kami/maven/internal/ipc" "github.com/kami/maven/internal/ipc"
) )
@@ -43,6 +51,10 @@ func run(args []string) error {
url := fs.String("url", "", "CalDAV calendar URL, e.g. http://localhost:5232/kami/personal (required)") url := fs.String("url", "", "CalDAV calendar URL, e.g. http://localhost:5232/kami/personal (required)")
user := fs.String("user", "", "CalDAV basic-auth username (required)") user := fs.String("user", "", "CalDAV basic-auth username (required)")
pass := fs.String("pass", "", "CalDAV basic-auth password (required)") pass := fs.String("pass", "", "CalDAV basic-auth password (required)")
renderURL := fs.String("render-url", "", "CalDAV collection maven publishes her own reminders to; empty disables rendering")
renderUser := fs.String("render-user", "", "basic-auth username for -render-url (defaults to -user)")
renderPass := fs.String("render-pass", "", "basic-auth password for -render-url (defaults to -pass)")
renderDur := fs.Duration("render-duration", calendar.DefaultReminderDuration, "how long a rendered reminder occupies")
interval := fs.Duration("interval", 5*time.Minute, "poll cadence") interval := fs.Duration("interval", 5*time.Minute, "poll cadence")
timeout := fs.Duration("timeout", 10*time.Second, "per-request HTTP timeout") timeout := fs.Duration("timeout", 10*time.Second, "per-request HTTP timeout")
if err := fs.Parse(args); err != nil { if err := fs.Parse(args); err != nil {
@@ -54,6 +66,9 @@ func run(args []string) error {
if *url == "" || *user == "" || *pass == "" { if *url == "" || *user == "" || *pass == "" {
return fmt.Errorf("-url, -user, -pass are required") return fmt.Errorf("-url, -user, -pass are required")
} }
if err := checkRenderTarget(*url, *renderURL); err != nil {
return err
}
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM) ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
defer stop() defer stop()
@@ -64,16 +79,36 @@ func run(args []string) error {
} }
defer core.Close() defer core.Close()
hc := &http.Client{Timeout: *timeout}
p := &poller{ p := &poller{
core: core, core: core,
http: &http.Client{Timeout: *timeout}, http: hc,
url: strings.TrimRight(*url, "/"), url: strings.TrimRight(*url, "/"),
user: *user, user: *user,
pass: *pass, pass: *pass,
} }
var rend *renderer
if *renderURL != "" {
ru, rp := *renderUser, *renderPass
if ru == "" {
ru = *user
}
if rp == "" {
rp = *pass
}
rend = newRenderer(core, hc, *renderURL, ru, rp, *renderDur)
log.Printf("mavcaldav: rendering reminders to %s", *renderURL)
}
log.Printf("mavcaldav: polling %s every %s", *url, *interval) log.Printf("mavcaldav: polling %s every %s", *url, *interval)
p.pollOnce(ctx) // fire immediately tick := func() {
p.pollOnce(ctx)
if rend != nil {
rend.renderOnce(ctx)
}
}
tick() // fire immediately
t := time.NewTicker(*interval) t := time.NewTicker(*interval)
defer t.Stop() defer t.Stop()
for { for {
@@ -82,11 +117,30 @@ func run(args []string) error {
log.Printf("mavcaldav: bye") log.Printf("mavcaldav: bye")
return nil return nil
case <-t.C: case <-t.C:
p.pollOnce(ctx) tick()
} }
} }
} }
// checkRenderTarget refuses a render URL that is also a read URL. This is the
// structural half of #127's "cannot write to your work calendar": the write
// credential and the write URL are separate flags, and the one calendar maven
// is known to only read is rejected as a target at startup rather than trusted
// at runtime.
func checkRenderTarget(readURL, renderURL string) error {
if renderURL == "" {
return nil
}
if sameCollection(readURL, renderURL) {
return fmt.Errorf("-render-url must differ from -url: maven renders into a calendar she owns, never into one she reads")
}
return nil
}
func sameCollection(a, b string) bool {
return strings.EqualFold(strings.TrimRight(a, "/"), strings.TrimRight(b, "/"))
}
type poller struct { type poller struct {
core ipc.CoreAPI core ipc.CoreAPI
http *http.Client http *http.Client
@@ -95,12 +149,6 @@ type poller struct {
pass string pass string
} }
type icalEvent struct {
start time.Time
end time.Time
summary string
}
func (p *poller) pollOnce(ctx context.Context) { func (p *poller) pollOnce(ctx context.Context) {
now := time.Now() now := time.Now()
events, err := p.fetchEvents(ctx, now) events, err := p.fetchEvents(ctx, now)
@@ -109,38 +157,30 @@ func (p *poller) pollOnce(ctx context.Context) {
return return
} }
busy := false
for _, e := range events {
if !now.Before(e.start) && now.Before(e.end) {
busy = true
break
}
}
busyVal := "false" busyVal := "false"
if busy { if calendar.Busy(events, now) {
busyVal = "true" busyVal = "true"
} }
// Write calendar_busy on change. // Write calendar_busy on change.
if err := p.writeIfChanged(ctx, "calendar_busy", "poll:caldav", busyVal, now); err != nil { if err := p.writeIfChanged(ctx, "calendar_busy", calendar.SourcePersonal, busyVal, now, 1.0); err != nil {
log.Printf("mavcaldav: write calendar_busy: %v", err) log.Printf("mavcaldav: write calendar_busy: %v", err)
return return
} }
// Write per-event facts (one per event, keyed by event summary + start). // Write per-event facts (one per event, keyed by day + event summary).
// This lets the note RAG path answer "what's on my calendar" without // This lets the note RAG path answer "what's on my calendar" without
// reaching back to Radicale. // reaching back to Radicale.
for _, e := range events { for _, e := range events {
val := fmt.Sprintf("%s @ %s-%s", e.summary, e.start.Format("15:04"), e.end.Format("15:04")) key := calendar.FactKey(e)
eventKey := fmt.Sprintf("calendar_event_%s_%s", e.start.Format("20060102"), safeKey(e.summary)) if err := p.writeIfChanged(ctx, key, calendar.SourcePersonal, calendar.FactValue(e), e.Start, 1.0); err != nil {
if err := p.writeIfChanged(ctx, eventKey, "poll:caldav", val, e.start); err != nil { log.Printf("mavcaldav: write %s: %v", key, err)
log.Printf("mavcaldav: write %s: %v", eventKey, err)
} }
} }
} }
// fetchEvents GETs the calendar URL and parses VEVENTs from the iCal response. // fetchEvents GETs the calendar URL and parses VEVENTs from the iCal response.
func (p *poller) fetchEvents(ctx context.Context, now time.Time) ([]icalEvent, error) { func (p *poller) fetchEvents(ctx context.Context, now time.Time) ([]calendar.Event, error) {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, p.url, nil) req, err := http.NewRequestWithContext(ctx, http.MethodGet, p.url, nil)
if err != nil { if err != nil {
return nil, err return nil, err
@@ -162,120 +202,11 @@ func (p *poller) fetchEvents(ctx context.Context, now time.Time) ([]icalEvent, e
return nil, fmt.Errorf("GET %s: %s", p.url, resp.Status) return nil, fmt.Errorf("GET %s: %s", p.url, resp.Status)
} }
return parseICal(body, now), nil return calendar.ParseICalDay(body, now), nil
}
// parseICal scans iCal text for VEVENT components. Returns events that overlap
// with today (UTC day boundaries) to keep the response manageable.
func parseICal(body []byte, now time.Time) []icalEvent {
todayStart := time.Date(now.Year(), now.Month(), now.Day(), 0, 0, 0, 0, time.UTC)
todayEnd := todayStart.AddDate(0, 0, 1)
var events []icalEvent
text := string(body)
for {
veventStart := strings.Index(text, "BEGIN:VEVENT")
if veventStart < 0 {
break
}
text = text[veventStart+len("BEGIN:VEVENT"):]
veventEnd := strings.Index(text, "END:VEVENT")
if veventEnd < 0 {
break
}
block := text[:veventEnd]
text = text[veventEnd+len("END:VEVENT"):]
e := parseVEVENT(block)
if e == nil {
continue
}
// Only keep events overlapping today.
if e.end.After(todayStart) && e.start.Before(todayEnd) {
events = append(events, *e)
}
}
return events
}
// parseVEVENT extracts start, end, summary from a VEVENT block.
// Supports both UTC (DTEND:20260703T100000Z) and local (DTSTART;TZID=...:...)
// formats. Returns nil for all-day events (no DTSTART/DTEND time component) or
// parse failures.
func parseVEVENT(block string) *icalEvent {
var e icalEvent
lines := strings.Split(block, "\n")
for _, line := range lines {
line = strings.TrimSpace(line)
switch {
case strings.HasPrefix(line, "DTSTART"):
if t, ok := parseDT(line); ok {
e.start = t
}
case strings.HasPrefix(line, "DTEND"):
if t, ok := parseDT(line); ok {
e.end = t
}
case strings.HasPrefix(line, "SUMMARY"):
if idx := strings.Index(line, ":"); idx >= 0 {
e.summary = strings.TrimSpace(line[idx+1:])
}
}
}
if e.start.IsZero() || e.end.IsZero() {
return nil
}
return &e
}
// parseDT parses a DTSTART/DTEND value. Supports:
// - UTC: DTEND:20260703T100000Z
// - Local: DTSTART;TZID=Europe/Moscow:20260703T130000
// - Value-date (all-day): DTSTART;VALUE=DATE:20260703 (returns zero time)
func parseDT(line string) (time.Time, bool) {
if strings.Contains(line, "VALUE=DATE:") {
return time.Time{}, false // all-day, skip
}
idx := strings.LastIndex(line, ":")
if idx < 0 {
return time.Time{}, false
}
val := line[idx+1:]
val = strings.TrimSuffix(val, "Z")
// Try UTC first (has Z suffix, or ended in Z before TrimSuffix).
if strings.HasSuffix(line, "Z") {
t, err := time.Parse("20060102T150405", val)
if err != nil {
return time.Time{}, false
}
return t.UTC(), true
}
// Local time — treat as UTC for simplicity (CalDAV server and poller
// run in the same timezone; the gate only needs busy/not-busy accuracy).
t, err := time.Parse("20060102T150405", val)
if err != nil {
return time.Time{}, false
}
return t.UTC(), true
}
// safeKey makes an event summary safe to use as a fact key (alphanumeric + dash).
func safeKey(s string) string {
var b strings.Builder
for _, r := range s {
if (r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z') || (r >= '0' && r <= '9') || r == '-' {
b.WriteRune(r)
} else if r == ' ' || r == '_' {
b.WriteRune('-')
}
}
return b.String()
} }
// writeIfChanged writes a fact only when the value differs from the latest. // writeIfChanged writes a fact only when the value differs from the latest.
func (p *poller) writeIfChanged(ctx context.Context, key, source, val string, ts time.Time) error { func (p *poller) writeIfChanged(ctx context.Context, key, source, val string, ts time.Time, confidence float64) error {
prev, err := p.core.LatestFactBySource(ctx, key, source) prev, err := p.core.LatestFactBySource(ctx, key, source)
switch { switch {
case err == nil && prev.Value == val: case err == nil && prev.Value == val:
@@ -289,7 +220,7 @@ func (p *poller) writeIfChanged(ctx context.Context, key, source, val string, ts
Key: key, Key: key,
Value: val, Value: val,
Source: source, Source: source,
Confidence: 1.0, Confidence: confidence,
}) })
if err != nil { if err != nil {
return fmt.Errorf("write %s: %w", key, err) return fmt.Errorf("write %s: %w", key, err)
+6 -165
View File
@@ -12,7 +12,7 @@ import (
) )
type fakeCore struct { type fakeCore struct {
ipc.CoreAPI ipc.UnimplementedCoreAPI
facts map[string]ipc.Fact // composite key "key|source" → Fact facts map[string]ipc.Fact // composite key "key|source" → Fact
writeLog []ipc.WriteFactReq writeLog []ipc.WriteFactReq
writeErr error writeErr error
@@ -51,165 +51,6 @@ func (f *fakeCore) WriteFact(_ context.Context, req ipc.WriteFactReq) (int64, er
return int64(len(f.writeLog)), nil return int64(len(f.writeLog)), nil
} }
// ---------------------------------------------------------------------------
// Parsing tests
// ---------------------------------------------------------------------------
func TestParseICal(t *testing.T) {
now := time.Date(2026, 7, 3, 12, 0, 0, 0, time.UTC)
body := []byte(`BEGIN:VCALENDAR
BEGIN:VEVENT
DTSTART:20260703T090000Z
DTEND:20260703T100000Z
SUMMARY:Morning standup
END:VEVENT
BEGIN:VEVENT
DTSTART:20260703T140000Z
DTEND:20260703T150000Z
SUMMARY:Team sync
END:VEVENT
BEGIN:VEVENT
DTSTART:20260702T140000Z
DTEND:20260702T150000Z
SUMMARY:Yesterday retro
END:VEVENT
BEGIN:VEVENT
DTSTART:20260704T090000Z
DTEND:20260704T100000Z
SUMMARY:Tomorrow standup
END:VEVENT
BEGIN:VEVENT
DTSTART;VALUE=DATE:20260704
DTEND;VALUE=DATE:20260705
SUMMARY:All-day event
END:VEVENT
END:VCALENDAR`)
events := parseICal(body, now)
if len(events) != 2 {
t.Fatalf("got %d events, want 2 (today events, no all-day/past/future)", len(events))
}
// Morning standup — overlaps today.
if events[0].summary != "Morning standup" {
t.Errorf("events[0].summary = %q, want %q", events[0].summary, "Morning standup")
}
wantStart0 := time.Date(2026, 7, 3, 9, 0, 0, 0, time.UTC)
if !events[0].start.Equal(wantStart0) {
t.Errorf("events[0].start = %v, want %v", events[0].start, wantStart0)
}
wantEnd0 := time.Date(2026, 7, 3, 10, 0, 0, 0, time.UTC)
if !events[0].end.Equal(wantEnd0) {
t.Errorf("events[0].end = %v, want %v", events[0].end, wantEnd0)
}
// Team sync — overlaps today.
if events[1].summary != "Team sync" {
t.Errorf("events[1].summary = %q, want %q", events[1].summary, "Team sync")
}
wantStart1 := time.Date(2026, 7, 3, 14, 0, 0, 0, time.UTC)
if !events[1].start.Equal(wantStart1) {
t.Errorf("events[1].start = %v, want %v", events[1].start, wantStart1)
}
wantEnd1 := time.Date(2026, 7, 3, 15, 0, 0, 0, time.UTC)
if !events[1].end.Equal(wantEnd1) {
t.Errorf("events[1].end = %v, want %v", events[1].end, wantEnd1)
}
}
func TestParseVEVENT(t *testing.T) {
// Normal event with TZID in DTSTART and UTC DTEND.
block := "DTSTART;TZID=Europe/Moscow:20260703T130000\nDTEND:20260703T140000Z\nSUMMARY:Stand up meeting"
e := parseVEVENT(block)
if e == nil {
t.Fatal("expected non-nil icalEvent")
}
wantStart := time.Date(2026, 7, 3, 13, 0, 0, 0, time.UTC)
if !e.start.Equal(wantStart) {
t.Errorf("start = %v, want %v", e.start, wantStart)
}
wantEnd := time.Date(2026, 7, 3, 14, 0, 0, 0, time.UTC)
if !e.end.Equal(wantEnd) {
t.Errorf("end = %v, want %v", e.end, wantEnd)
}
if e.summary != "Stand up meeting" {
t.Errorf("summary = %q, want %q", e.summary, "Stand up meeting")
}
// All-day event (VALUE=DATE) → nil.
allDay := "DTSTART;VALUE=DATE:20260703\nDTEND;VALUE=DATE:20260704\nSUMMARY:All-day"
if e2 := parseVEVENT(allDay); e2 != nil {
t.Error("expected nil for all-day event")
}
}
func TestParseDT(t *testing.T) {
tests := []struct {
name string
line string
want time.Time
wantOK bool
}{
{
name: "UTC",
line: "DTEND:20260703T100000Z",
want: time.Date(2026, 7, 3, 10, 0, 0, 0, time.UTC),
wantOK: true,
},
{
name: "local time",
line: "DTSTART;TZID=Europe/Moscow:20260703T130000",
want: time.Date(2026, 7, 3, 13, 0, 0, 0, time.UTC),
wantOK: true,
},
{
name: "all-day",
line: "DTSTART;VALUE=DATE:20260703",
want: time.Time{},
wantOK: false,
},
{
name: "invalid",
line: "DTSTART:garbage",
want: time.Time{},
wantOK: false,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
got, ok := parseDT(tt.line)
if ok != tt.wantOK {
t.Errorf("ok = %v, want %v", ok, tt.wantOK)
}
if !got.Equal(tt.want) {
t.Errorf("got = %v, want %v", got, tt.want)
}
})
}
}
func TestSafeKey(t *testing.T) {
tests := []struct {
input string
want string
}{
{"Stand up meeting", "Stand-up-meeting"},
{"Hello_World", "Hello-World"},
{"special@#$chars!!", "specialchars"},
{"ALL_CAPS_123", "ALL-CAPS-123"},
}
for _, tt := range tests {
got := safeKey(tt.input)
if got != tt.want {
t.Errorf("safeKey(%q) = %q, want %q", tt.input, got, tt.want)
}
}
}
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Core logic tests // Core logic tests
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -221,7 +62,7 @@ func TestWriteIfChanged(t *testing.T) {
t.Run("no previous fact writes", func(t *testing.T) { t.Run("no previous fact writes", func(t *testing.T) {
fc := &fakeCore{} fc := &fakeCore{}
p := &poller{core: fc} p := &poller{core: fc}
err := p.writeIfChanged(ctx, "test_key", "poll:caldav", "hello", now) err := p.writeIfChanged(ctx, "test_key", "poll:caldav", "hello", now, 1.0)
if err != nil { if err != nil {
t.Fatalf("unexpected error: %v", err) t.Fatalf("unexpected error: %v", err)
} }
@@ -249,7 +90,7 @@ func TestWriteIfChanged(t *testing.T) {
}, },
} }
p := &poller{core: fc} p := &poller{core: fc}
err := p.writeIfChanged(ctx, "test_key", "poll:caldav", "hello", now) err := p.writeIfChanged(ctx, "test_key", "poll:caldav", "hello", now, 1.0)
if err != nil { if err != nil {
t.Fatalf("unexpected error: %v", err) t.Fatalf("unexpected error: %v", err)
} }
@@ -265,7 +106,7 @@ func TestWriteIfChanged(t *testing.T) {
}, },
} }
p := &poller{core: fc} p := &poller{core: fc}
err := p.writeIfChanged(ctx, "test_key", "poll:caldav", "new", now) err := p.writeIfChanged(ctx, "test_key", "poll:caldav", "new", now, 1.0)
if err != nil { if err != nil {
t.Fatalf("unexpected error: %v", err) t.Fatalf("unexpected error: %v", err)
} }
@@ -280,7 +121,7 @@ func TestWriteIfChanged(t *testing.T) {
t.Run("read error other than ErrNoFact returns error", func(t *testing.T) { t.Run("read error other than ErrNoFact returns error", func(t *testing.T) {
fc := &fakeCore{readErr: fmt.Errorf("connection refused")} fc := &fakeCore{readErr: fmt.Errorf("connection refused")}
p := &poller{core: fc} p := &poller{core: fc}
err := p.writeIfChanged(ctx, "fail_key", "poll:caldav", "x", now) err := p.writeIfChanged(ctx, "fail_key", "poll:caldav", "x", now, 1.0)
if err == nil { if err == nil {
t.Fatal("expected error, got nil") t.Fatal("expected error, got nil")
} }
@@ -292,7 +133,7 @@ func TestWriteIfChanged(t *testing.T) {
writeErr: fmt.Errorf("disk full"), writeErr: fmt.Errorf("disk full"),
} }
p := &poller{core: fc} p := &poller{core: fc}
err := p.writeIfChanged(ctx, "test_key", "poll:caldav", "hello", now) err := p.writeIfChanged(ctx, "test_key", "poll:caldav", "hello", now, 1.0)
if err == nil { if err == nil {
t.Fatal("expected error, got nil") t.Fatal("expected error, got nil")
} }
+145
View File
@@ -0,0 +1,145 @@
package main
import (
"context"
"fmt"
"io"
"log"
"net/http"
"strings"
"time"
"github.com/kami/maven/internal/calendar"
"github.com/kami/maven/internal/ipc"
)
// renderer is the write half of maven's own local calendar (Vikunja #127).
//
// It is a RENDER TARGET, not a store. sqlite stays canonical: every tick the
// renderer reads the pending reminders out of core and publishes each one as a
// single-event iCal resource in a CalDAV collection maven owns. Nothing is ever
// read back from that collection, and losing it costs nothing — the next tick
// rebuilds it.
//
// It structurally cannot write to a calendar maven only reads. The URL comes
// from its own flag, checked at startup against every read URL (see
// run in main.go), and the only paths it ever addresses carry
// calendar.ReminderUIDPrefix — so even pointed at the wrong collection it can
// only touch resources it created.
type renderer struct {
core ipc.CoreAPI
http *http.Client
url string
user string
pass string
dur time.Duration
// published maps reminder id → the body last successfully PUT, so an
// unchanged reminder costs nothing. Purely an optimisation: a restart
// re-publishes every reminder once, which is idempotent.
published map[int64]string
}
func newRenderer(core ipc.CoreAPI, hc *http.Client, url, user, pass string, dur time.Duration) *renderer {
return &renderer{
core: core,
http: hc,
url: strings.TrimRight(url, "/"),
user: user,
pass: pass,
dur: dur,
published: make(map[int64]string),
}
}
// renderOnce publishes every pending reminder and withdraws the ones that are
// no longer pending. Errors are logged and skipped: a calendar maven cannot
// reach must never break the reminder itself, which lives in sqlite.
func (r *renderer) renderOnce(ctx context.Context) {
reminders, err := r.core.ListReminders(ctx, renderMaxReminders)
if err != nil {
log.Printf("mavcaldav: list reminders: %v", err)
return
}
live := make(map[int64]bool, len(reminders))
for _, rem := range reminders {
if rem.Status != "pending" {
continue
}
live[rem.ID] = true
e := calendar.ReminderEvent(rem.ID, fireTime(rem), rem.Payload, r.dur)
body := calendar.RenderICal([]calendar.Event{e})
if r.published[rem.ID] == body {
continue
}
if err := r.put(ctx, calendar.ReminderPath(rem.ID), body); err != nil {
log.Printf("mavcaldav: render reminder %d: %v", rem.ID, err)
continue
}
r.published[rem.ID] = body
log.Printf("mavcaldav: rendered reminder %d (%s)", rem.ID, e.Summary)
}
for id := range r.published {
if live[id] {
continue
}
if err := r.delete(ctx, calendar.ReminderPath(id)); err != nil {
log.Printf("mavcaldav: withdraw reminder %d: %v", id, err)
continue
}
delete(r.published, id)
log.Printf("mavcaldav: withdrew reminder %d", id)
}
}
// renderMaxReminders bounds the read. Reminders past this count are older than
// anything a calendar view is useful for.
const renderMaxReminders = 200
// fireTime prefers NextFireTs — for a recurring reminder that is the occurrence
// worth showing; FireTs is the original statement.
func fireTime(rem ipc.Reminder) time.Time {
if !rem.NextFireTs.IsZero() {
return rem.NextFireTs
}
return rem.FireTs
}
func (r *renderer) put(ctx context.Context, name, body string) error {
req, err := http.NewRequestWithContext(ctx, http.MethodPut, r.url+"/"+name, strings.NewReader(body))
if err != nil {
return err
}
req.SetBasicAuth(r.user, r.pass)
req.Header.Set("Content-Type", "text/calendar; charset=utf-8")
return r.do(req, name)
}
func (r *renderer) delete(ctx context.Context, name string) error {
req, err := http.NewRequestWithContext(ctx, http.MethodDelete, r.url+"/"+name, nil)
if err != nil {
return err
}
req.SetBasicAuth(r.user, r.pass)
return r.do(req, name)
}
// do runs the request and treats any 2xx, plus 404 on a DELETE, as success —
// a resource that is already gone is the state the caller wanted.
func (r *renderer) do(req *http.Request, name string) error {
resp, err := r.http.Do(req)
if err != nil {
return err
}
defer resp.Body.Close()
io.Copy(io.Discard, io.LimitReader(resp.Body, 1<<16))
switch {
case resp.StatusCode >= 200 && resp.StatusCode < 300:
return nil
case req.Method == http.MethodDelete && resp.StatusCode == http.StatusNotFound:
return nil
}
return fmt.Errorf("%s %s: %s", req.Method, name, resp.Status)
}
+186
View File
@@ -0,0 +1,186 @@
package main
import (
"context"
"io"
"net/http"
"net/http/httptest"
"strings"
"sync"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
)
// reminderCore is a fakeCore that also answers ListReminders.
type reminderCore struct {
fakeCore
reminders []ipc.Reminder
listErr error
}
func (c *reminderCore) ListReminders(context.Context, int) ([]ipc.Reminder, error) {
if c.listErr != nil {
return nil, c.listErr
}
return c.reminders, nil
}
// calSrv records what a CalDAV collection received.
type calSrv struct {
mu sync.Mutex
puts map[string]string
dels []string
status int
*httptest.Server
}
func newCalSrv() *calSrv {
s := &calSrv{puts: map[string]string{}, status: http.StatusCreated}
s.Server = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
body, _ := io.ReadAll(r.Body)
s.mu.Lock()
defer s.mu.Unlock()
switch r.Method {
case http.MethodPut:
s.puts[strings.TrimPrefix(r.URL.Path, "/cal/")] = string(body)
case http.MethodDelete:
s.dels = append(s.dels, strings.TrimPrefix(r.URL.Path, "/cal/"))
}
w.WriteHeader(s.status)
}))
return s
}
func (s *calSrv) putCount() int {
s.mu.Lock()
defer s.mu.Unlock()
return len(s.puts)
}
func TestRenderOncePublishesPendingReminders(t *testing.T) {
fire := time.Date(2026, 8, 1, 18, 30, 0, 0, time.UTC)
srv := newCalSrv()
defer srv.Close()
core := &reminderCore{reminders: []ipc.Reminder{
{ID: 7, FireTs: fire, Payload: "позвонить маме", Status: "pending"},
{ID: 8, FireTs: fire, Payload: "уже сделано", Status: "fired"},
{ID: 9, FireTs: fire, Payload: "отменено", Status: "cancelled"},
}}
r := newRenderer(core, srv.Client(), srv.URL+"/cal/", "u", "p", 0)
r.renderOnce(context.Background())
srv.mu.Lock()
body, ok := srv.puts["maven-reminder-7.ics"]
n := len(srv.puts)
srv.mu.Unlock()
if n != 1 {
t.Fatalf("expected exactly the pending reminder to be published, got %d PUTs", n)
}
if !ok {
t.Fatal("pending reminder 7 was not published")
}
if !strings.Contains(body, "SUMMARY:позвонить маме") {
t.Errorf("payload missing from rendered body:\n%s", body)
}
if !strings.Contains(body, "UID:maven-reminder-7") {
t.Errorf("UID missing from rendered body:\n%s", body)
}
}
func TestRenderOnceSkipsUnchanged(t *testing.T) {
srv := newCalSrv()
defer srv.Close()
core := &reminderCore{reminders: []ipc.Reminder{
{ID: 1, FireTs: time.Date(2026, 8, 1, 9, 0, 0, 0, time.UTC), Payload: "выпить воды", Status: "pending"},
}}
r := newRenderer(core, srv.Client(), srv.URL+"/cal", "u", "p", 0)
r.renderOnce(context.Background())
r.renderOnce(context.Background())
if got := srv.putCount(); got != 1 {
t.Fatalf("an unchanged reminder was re-published: %d distinct PUTs", got)
}
}
func TestRenderOnceWithdrawsResolvedReminders(t *testing.T) {
srv := newCalSrv()
defer srv.Close()
core := &reminderCore{reminders: []ipc.Reminder{
{ID: 5, FireTs: time.Date(2026, 8, 1, 9, 0, 0, 0, time.UTC), Payload: "встреча", Status: "pending"},
}}
r := newRenderer(core, srv.Client(), srv.URL+"/cal", "u", "p", 0)
r.renderOnce(context.Background())
core.reminders[0].Status = "fired"
r.renderOnce(context.Background())
srv.mu.Lock()
dels := append([]string(nil), srv.dels...)
srv.mu.Unlock()
if len(dels) != 1 || dels[0] != "maven-reminder-5.ics" {
t.Fatalf("resolved reminder was not withdrawn: %v", dels)
}
if len(r.published) != 0 {
t.Errorf("published map still holds %v", r.published)
}
}
// A calendar maven cannot reach must never break anything: sqlite is canonical.
func TestRenderOnceSurvivesServerErrors(t *testing.T) {
srv := newCalSrv()
srv.status = http.StatusInternalServerError
defer srv.Close()
core := &reminderCore{reminders: []ipc.Reminder{
{ID: 1, FireTs: time.Date(2026, 8, 1, 9, 0, 0, 0, time.UTC), Payload: "x", Status: "pending"},
}}
r := newRenderer(core, srv.Client(), srv.URL+"/cal", "u", "p", 0)
r.renderOnce(context.Background())
if len(r.published) != 0 {
t.Error("a failed PUT must not be recorded as published, or it never retries")
}
}
func TestRenderOnceUsesNextFireForRecurring(t *testing.T) {
srv := newCalSrv()
defer srv.Close()
next := time.Date(2026, 8, 2, 7, 0, 0, 0, time.UTC)
core := &reminderCore{reminders: []ipc.Reminder{{
ID: 3,
FireTs: time.Date(2026, 8, 1, 7, 0, 0, 0, time.UTC),
NextFireTs: next,
Payload: "зарядка",
Status: "pending",
Cron: "0 7 * * *",
}}}
r := newRenderer(core, srv.Client(), srv.URL+"/cal", "u", "p", 0)
r.renderOnce(context.Background())
srv.mu.Lock()
body := srv.puts["maven-reminder-3.ics"]
srv.mu.Unlock()
if !strings.Contains(body, "DTSTART:20260802T070000Z") {
t.Errorf("recurring reminder should render its next occurrence:\n%s", body)
}
}
func TestCheckRenderTargetRefusesTheCalendarItReads(t *testing.T) {
read := "http://localhost:5232/kami/personal"
if err := checkRenderTarget(read, ""); err != nil {
t.Fatalf("rendering off must be fine: %v", err)
}
if err := checkRenderTarget(read, "http://localhost:5232/kami/maven"); err != nil {
t.Fatalf("a distinct collection must be accepted: %v", err)
}
if err := checkRenderTarget(read, read); err == nil {
t.Error("rendering into the read calendar must be refused")
}
if err := checkRenderTarget(read, read+"/"); err == nil {
t.Error("a trailing slash must not defeat the check")
}
if err := checkRenderTarget(read, strings.ToUpper(read)); err == nil {
t.Error("case must not defeat the check")
}
}
+71
View File
@@ -0,0 +1,71 @@
// actionTable dispatches applyAction's per-intent bodies. Each of the 7
// intents (fact, reminder, note, query, act, chat, system) has one handler
// here with the signature:
//
// func(h *reactiveHandler, ctx context.Context, dec router.Decision) string
//
// same contract as applyAction itself: "" means "let the Replier phrase the
// reply", a non-empty string OVERRIDES it. This is a straight extraction of
// applyAction's old switch cases (formerly ~300 lines in voice.go) — no
// reordering of side effects, no new abstractions inside a handler.
//
// What does NOT belong in this table, because it is not per-intent:
//
// - the dec.Clarify short-circuit ("" when the router's stage-3 fired) —
// stays in applyAction, before dispatch, since it applies to every
// intent identically.
// - the destructive-act confirm gate (park / resolveConfirm / confirmTTL)
// and the enabled-tool allowlist. Both live entirely inside
// actionAct/handleAct in actions_act.go, exactly where they lived in the old
// switch's IntentAct case — they are act-specific (a fact or a note
// can't be destructive), not shared across intents, so they do not need
// to move to a separate layer. The important invariant, preserved
// as-is: applyAction runs identically whether dec came from a fresh
// route or from a completed clarify answer (see finishClarified in
// clarify.go and its comment "filling in an argument never grants
// authority") — a handler must never special-case a clarify-completed
// decision to skip the confirm gate or the allowlist.
// - detectPattern and dialogue-session bookkeeping (rememberTurn,
// followUpMerge) run in the callers (runTurn,
// finishClarified), not per-intent, and are untouched by this slice.
//
// Each handler lives in actions_<intent>.go; the small ones (chat, system)
// and the table itself stay here.
//
// Adding an intent: write its handler in its own file, add one line to
// actionHandlers. Do not grow applyAction's switch back.
package main
import (
"context"
"log"
"github.com/kami/maven/internal/router"
)
// actionHandlers is the per-intent dispatch table used by applyAction.
var actionHandlers = map[router.Intent]func(*reactiveHandler, context.Context, router.Decision) string{
router.IntentFact: (*reactiveHandler).actionFact,
router.IntentReminder: (*reactiveHandler).actionReminder,
router.IntentAct: (*reactiveHandler).actionAct,
router.IntentChat: (*reactiveHandler).actionChat,
router.IntentSystem: (*reactiveHandler).actionSystem,
router.IntentNote: (*reactiveHandler).actionNote,
router.IntentQuery: (*reactiveHandler).actionQuery,
}
func (h *reactiveHandler) actionChat(ctx context.Context, dec router.Decision) string {
// Conversational: build history from dialogue session (prior user turns)
// and let the LLM respond from general knowledge + context.
history := h.chatHistory()
reply, err := h.phraser.PhraseChat(ctx, dec.Utterance, history)
if err != nil {
log.Printf("voice: chat: %v", err)
return "поговорили."
}
return reply
}
func (h *reactiveHandler) actionSystem(ctx context.Context, dec router.Decision) string {
return h.replySystem(ctx, dec)
}
+66
View File
@@ -0,0 +1,66 @@
package main
import (
"context"
"errors"
"log"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/tool"
)
// actionAct handles router.IntentAct: match a verb to an enabled tool, offer
// it to the ecosystems first, and run it behind the confirm gate and the
// allowlist. proposeGap and the confirm gate itself live in confirm.go.
func (h *reactiveHandler) actionAct(ctx context.Context, dec router.Decision) string {
// tool executor: run the matched fn against the enabled allowlist.
// HasFn=false ⇒ try the matcher (for LLM-routed acts where the verb
// didn't go through the stage-0 act grammar).
if !dec.Slots.HasFn && dec.Slots.Text != "" && h.matcher != nil {
if fn, args, ok := h.matcher.Match(dec.Slots.Text); ok {
dec.Slots.Fn, dec.Slots.Args, dec.Slots.HasFn = fn, args, true
}
}
// Praxis ecosystem tools: intercept before the system command executor.
if h.ecosystem != nil && h.ecosystem.praxis != nil && dec.Slots.HasFn {
if reply := h.handlePraxisAct(ctx, dec); reply != "" {
return reply
}
}
// Hexis ecosystem action: if ecosystem is configured and we have a verb
// + entity text, try to resolve the entity and execute via Hexis.
if h.ecosystem != nil && h.ecosystem.hexis != nil && dec.Slots.Text != "" {
if reply := h.handleHexisAct(ctx, dec); reply != "" {
return reply
}
}
// HasFn still false ⇒ no allowlist match: scaffold a 'proposed' tool
// the user can enable on the authed surface ("earn the right to ask").
if !dec.Slots.HasFn {
return h.proposeGap(ctx, dec)
}
out, err := h.tools.Exec(ctx, dec.Slots.Fn, dec.Slots.Args, false)
if err != nil {
switch {
case errors.Is(err, tool.ErrNeedsConfirm):
// destructive: park it and ask. The next utterance answers.
phrase := actPhrase(dec.Slots.Fn, dec.Slots.Args)
h.park(dec.Slots.Fn, dec.Slots.Args, phrase)
return "выполнить «" + phrase + "»? скажи «да» или «нет»."
case errors.Is(err, tool.ErrNotEnabled):
return h.proposeGap(ctx, dec)
}
log.Printf("voice: tool %s: %v", dec.Slots.Fn, err)
if out != "" {
return "не получилось выполнить команду: " + firstLine(out)
}
return "не получилось выполнить команду."
}
if out != "" {
return "готово: " + firstLine(out)
}
return "готово."
}
+64
View File
@@ -0,0 +1,64 @@
package main
import (
"context"
"log"
"strconv"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
)
// actionFact handles router.IntentFact: persist a tapped self-fact, index
// it for recall, and let pattern detection propose a routine.
func (h *reactiveHandler) actionFact(ctx context.Context, dec router.Decision) string {
if !dec.Slots.HasKey {
return "не разобрала, что записать — попробуй иначе."
}
now := h.now()
req := ipc.WriteFactReq{
Ts: now,
Kind: "self",
Key: dec.Slots.Key,
Value: dec.Slots.Value,
Source: "tap:voice",
Confidence: 1.0,
// Subject: the key doubles as the entity-resolution candidate —
// a voice-tapped fact's key is usually the thing/person it's
// about ("espresso_machine", "kate"), so queueing it for Nexus
// resolution costs one async lookup and is a no-op (not_found)
// for the abstract self-state keys (mood, water) that aren't
// entities at all.
Subject: dec.Slots.Key,
}
factID, err := h.api.WriteFact(ctx, req)
if err != nil {
log.Printf("voice: write fact: %v", err)
return "не получилось сохранить факт."
}
// Index the fact utterance in long-term memory (best-effort, must not
// fail the fact write). Facts aren't in the notes table, so this is the
// only recall path for them — "когда я пил воду?" reads back from here.
if h.memStore != nil {
if vec, err := router.EmbedPassage(ctx, h.embedder, dec.Utterance); err != nil {
log.Printf("voice: embed fact for memory: %v", err)
} else if err := h.memStore.Insert(ctx, "fact:"+dec.Slots.Key+":"+strconv.FormatInt(now.Unix(), 10), vec, map[string]string{
"source": "voice",
"type": "fact",
"text": dec.Utterance,
"ts": strconv.FormatInt(now.Unix(), 10),
}); err != nil {
log.Printf("voice: memory insert fact: %v", err)
}
}
// Event extraction + pattern detection (best-effort, must not fail the
// fact write). If the fact describes a recognizable action, it becomes a
// normalized event; if ≥3 events for the same action+object show stable
// intervals, a proposed routine is created and parked for confirmation.
if h.dataStore != nil {
if phrase := h.detectPattern(ctx, factID, dec.Slots.Key, dec.Slots.Value, now); phrase != "" {
return phrase // "ты заправляешь ... напоминать?"
}
}
return "" // replier phrases the success reply
}
+47
View File
@@ -0,0 +1,47 @@
package main
import (
"context"
"log"
"strconv"
"github.com/kami/maven/internal/router"
)
// actionNote handles router.IntentNote: embed the note, persist it, and
// index it for recall.
func (h *reactiveHandler) actionNote(ctx context.Context, dec router.Decision) string {
// An utterance that explicitly files a task is work, not recall, and
// belongs in the task store (Vikunja #130). Checked before the embedding
// is paid for. Everything else is a note, exactly as before.
if reply, ok := h.captureTaskFromNote(ctx, dec); ok {
return reply
}
// embed the note text with the same model the classifier uses, persist
// via CoreAPI (source=tap:voice). Semantic recall lives in `notes`, not
// facts — no predicate reads it (spec's two-memory split).
vec, err := router.EmbedPassage(ctx, h.embedder, dec.Utterance)
if err != nil {
log.Printf("voice: embed note: %v", err)
return "не получилось сохранить заметку."
}
noteTs := h.now()
noteID, err := h.api.WriteNote(ctx, noteTs, dec.Utterance, vec, "tap:voice")
if err != nil {
log.Printf("voice: write note: %v", err)
return "не получилось сохранить заметку."
}
// Insert into long-term memory (best-effort, must not fail the note write).
// text/ts in the meta make a Search hit self-describing (see bestRecall).
if h.memStore != nil {
if err := h.memStore.Insert(ctx, "note:"+strconv.FormatInt(noteID, 10), vec, map[string]string{
"source": "voice",
"type": "note",
"text": dec.Utterance,
"ts": strconv.FormatInt(noteTs.Unix(), 10),
}); err != nil {
log.Printf("voice: memory insert: %v", err)
}
}
return "" // replier phrases the "saved" reply
}
+315
View File
@@ -0,0 +1,315 @@
package main
import (
"context"
"errors"
"fmt"
"log"
"strings"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/morning"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/weather"
)
// queryTurn is the per-turn scratch a chain of query sources shares: the
// decision being answered plus the work an earlier source already paid for
// (the query embedding, the notes it pulled). Sources read and fill it in
// order, so a later source never re-embeds.
type queryTurn struct {
dec router.Decision
vec []float32
notes []ipc.Note
}
// querySource — one answer source in the chain actionQuery walks. answer
// returns (reply, true) when this source claims the question, ("", false)
// when it passes to the next one. name is for reading the table, not logged.
//
// A struct of one func rather than an interface: every source is a plain
// method on *reactiveHandler with no state of its own (what state a turn has
// lives in queryTurn), so an interface would mean one empty type per source
// to satisfy it — ceremony for nothing. Same reasoning as confirmResolver in
// confirm.go, and the table then reads like actionHandlers: a flat list of
// method expressions you extend with one line.
type querySource struct {
name string
answer func(*reactiveHandler, context.Context, *queryTurn) (string, bool)
}
// querySources is the ordered chain actionQuery walks; first source to claim
// answers the turn. THE ORDER IS LOAD-BEARING — see the memory-before-notes
// comment on queryMemory: running the notes-only pass first was #373, and the
// gate was never the bug. Adding a source (Kiwix, RSS, crawler, email) is one
// line here plus its method; where you put the line is the whole decision.
var querySources = []querySource{
{"fact-by-key", (*reactiveHandler).queryFactByKey},
// Before "calendar" on purpose: both match "…на сегодня", and the plan is
// the more specific ask (its matcher requires a plan word), so the calendar
// listing would otherwise swallow it.
{"day-plan", (*reactiveHandler).queryDayPlan},
// Also before "calendar": "что я обычно делаю по средам?" names a weekday,
// and the habit question is the more specific one. Its matcher requires a
// habit marker ("обычно", "каждый", …), so a question about this coming
// Wednesday still reaches the calendar.
{"habits", (*reactiveHandler).queryHabits},
// Before "calendar" and before the recall sources: "что мне нужно
// сделать?" is a question about the task list, and the notes pass would
// otherwise answer it with whatever note happens to be nearest. Its
// matcher requires a task noun or an explicit "что … сделать", so a
// date-bearing question still reaches the calendar.
{"tasks", (*reactiveHandler).queryTasks},
{"calendar", (*reactiveHandler).queryCalendar},
{"weather", (*reactiveHandler).queryWeather},
{"embed", (*reactiveHandler).queryEmbed},
{"memory", (*reactiveHandler).queryMemory},
{"notes", (*reactiveHandler).queryNotes},
{"general-knowledge", (*reactiveHandler).queryGeneral},
}
func (h *reactiveHandler) actionQuery(ctx context.Context, dec router.Decision) string {
t := &queryTurn{dec: dec}
for _, src := range querySources {
if reply, ok := src.answer(h, ctx, t); ok {
return reply
}
}
return "не знаю."
}
// queryFactByKey — when the dialogue layer resolved an anaphoric reference to
// a prior fact's key (e.g. "когда я это сделал?" after "запиши что я пил
// воду"), look up the fact's value directly.
func (h *reactiveHandler) queryFactByKey(ctx context.Context, t *queryTurn) (string, bool) {
dec := t.dec
if !dec.Slots.HasKey || dec.Slots.Key == "" {
return "", false
}
f, err := h.api.LatestFact(ctx, dec.Slots.Key)
if err != nil {
return "", false
}
if dec.Slots.HasTime {
// The query asks about timing — the fact's own timestamp is the
// answer it's looking for. Format as a natural reply.
return fmt.Sprintf("я записала это %s", formatTime(f.Ts)), true
}
// General fact reference: describe what we know.
if dec.Utterance == "" {
return fmt.Sprintf("вот что я знаю: %s — %s", dec.Slots.Key, f.Value), true
}
// The utterance still carries the question; fall through to normal RAG
// with the resolved key in context.
return "", false
}
// queryDayPlan — "какие планы на сегодня?", "что у меня по плану?", "что
// дальше?" (Vikunja #128). Recites the day: calendar events, pending
// reminders, and any morning checklist still outstanding.
//
// Read-only by construction — the plan is assembled and rendered core-side and
// nothing here schedules or announces. "что дальше?" asks for the rest of the
// day, so that phrasing trims what has already passed.
func (h *reactiveHandler) queryDayPlan(ctx context.Context, t *queryTurn) (string, bool) {
if !router.IsDayPlanQuery(t.dec.Utterance) {
return "", false
}
plan, err := h.api.DayPlan(ctx)
if err != nil {
log.Printf("voice: day plan: %v", err)
return "не получилось собрать план.", true
}
if !isRestOfDayQuery(t.dec.Utterance) {
return plan.Spoken, true
}
// Rebuild the pure plan so the rest-of-day rendering is the same code that
// rendered the whole day — one formatter, one persona.
p := morning.Plan{Date: plan.Date}
for _, it := range plan.Items {
p.Items = append(p.Items, morning.PlanEntry{
At: it.At,
Text: it.Text,
Kind: morning.PlanKind(it.Kind),
Uncertain: it.Uncertain,
})
}
return p.After(h.now()).FormatRU(), true
}
// isRestOfDayQuery — "что дальше?" and its English form, the only plan phrasing
// that means "from now on" rather than "the whole day".
func isRestOfDayQuery(text string) bool {
s := strings.ToLower(text)
return strings.Contains(s, "дальше") || strings.Contains(s, "next")
}
// habitFactWindow — how many recent facts the behaviour profile is counted
// over. Enough for a season of habits without scanning the whole store on every
// question; the profile is recomputed on read, so the bound is the cost control.
const habitFactWindow = 2000
// queryHabits — "что я обычно делаю по вторникам?" (Vikunja #254). Counts the
// answer out of the fact log rather than asking the model to summarise a life:
// see internal/memory/behavior.go for why nothing here is generated.
func (h *reactiveHandler) queryHabits(ctx context.Context, t *queryTurn) (string, bool) {
q, ok := router.ParseHabitQuery(t.dec.Utterance)
if !ok {
return "", false
}
facts, err := h.api.RecentFacts(ctx, habitFactWindow)
if err != nil {
log.Printf("voice: habits: recent facts: %v", err)
return "не получилось посмотреть записи.", true
}
obs := make([]memory.Observation, 0, len(facts))
for _, f := range facts {
obs = append(obs, memory.Observation{At: f.Ts, Key: f.Key, Kind: f.Kind})
}
profile := memory.BuildProfile(obs, h.now())
if q.HasWeekday {
return profile.FormatWeekdayRU(q.Weekday), true
}
return profile.FormatOverallRU(), true
}
// queryCalendar — "что у меня сегодня?", "планы на завтра?"
// h.now(), not time.Now(): the handler's clock is the injected one, so this
// source can be tested at a fixed time like the rest.
func (h *reactiveHandler) queryCalendar(ctx context.Context, t *queryTurn) (string, bool) {
date, ok := router.ParseCalendarDate(t.dec.Utterance, h.now())
if !ok {
return "", false
}
events, err := h.api.CalendarEvents(ctx, date, date.Add(24*time.Hour))
if err != nil {
log.Printf("voice: calendar events: %v", err)
return "не получилось проверить календарь.", true
}
// Provenance travels with each event. A work meeting relayed off a phone
// notification (source ambient:notif, #126) is stored below full confidence
// and gets hedged; a CalDAV read is recited plainly.
entries := make([]router.CalendarEntry, len(events))
for i, e := range events {
entries[i] = router.CalendarEntry{Text: e.Value, Uncertain: e.Confidence < 1.0}
}
var f router.CalendarEventFormatter
return f.FormatEntries(entries, date), true
}
func (h *reactiveHandler) queryWeather(ctx context.Context, t *queryTurn) (string, bool) {
if !isWeatherQuery(t.dec.Utterance) {
return "", false
}
loc := extractWeatherLocation(t.dec.Utterance, h.weatherLocation)
ctxWT, cancel := context.WithTimeout(ctx, 5*time.Second)
defer cancel()
w, err := h.weatherProvider.CurrentWeather(ctxWT, loc)
if errors.Is(err, weather.ErrNotConfigured) {
return "погода не настроена.", true
}
if err != nil {
log.Printf("voice: weather: %v", err)
return "не получилось узнать погоду.", true
}
return fmt.Sprintf("в %s сейчас %.0f градусов, %s.", w.Location, w.Temperature, w.Condition), true
}
// queryEmbed isn't an answer source — it's the shared cost the two recall
// sources below both need, run once, in the position it always ran in. It
// only claims the turn when the embedder fails.
func (h *reactiveHandler) queryEmbed(ctx context.Context, t *queryTurn) (string, bool) {
vec, err := router.EmbedQuery(ctx, h.embedder, t.dec.Utterance)
if err != nil {
log.Printf("voice: embed query: %v", err)
return "не получилось найти ответ.", true
}
t.vec = vec
return "", false
}
// queryMemory — long-term memory first: ONE search over everything Maven
// remembers (notes and facts share this index) and ONE confidence gate, so
// the memory that is clearly the best match answers — a note just as much as
// a fact.
//
// This used to run only after the notes-only source below had already
// rejected the same note at the same score, which no note could ever survive
// a second time: the branch could only return a fact (#373). Order, not the
// gate, was the bug — the set of questions Maven answers is unchanged, only
// which memory gets to answer them.
func (h *reactiveHandler) queryMemory(ctx context.Context, t *queryTurn) (string, bool) {
if h.memStore == nil {
return "", false
}
hits, herr := h.memStore.Search(ctx, t.vec, 3)
if herr != nil {
log.Printf("voice: memory search: %v", herr)
return "", false
}
hit, ok := bestRecall(hits, h.queryMinScore, h.queryMinMargin)
if !ok {
return "", false
}
text := hit.Meta["text"]
// A note is phrased in Maven's voice; a fact is read back as it was
// stored.
if hit.Meta["type"] == "note" {
if reply, perr := h.phraser.PhraseQuery(ctx, t.dec.Utterance, []string{text}); perr == nil && reply != "" {
return reply, true
}
}
return text, true
}
// queryNotes — notes-only pass, for notes the vector index above does not
// hold (an older note written before it existed). Same gate, notes-only
// candidates.
//
// Confidence gate: below it, say "I don't know" rather than read back the
// least-unrelated note — a confident wrong recall is worse than a gap (spec's
// "not a guesser-of-truth"). Same instinct as the loop's since(key)==null →
// don't fire. Two parts: an absolute cosine floor, and a margin over the
// runner-up, which is the part that works with the e5 embedder's narrow score
// band. See memory.Confident. Failing the gate passes the turn on to general
// knowledge, which is what "don't read back the runner-up" means here.
func (h *reactiveHandler) queryNotes(ctx context.Context, t *queryTurn) (string, bool) {
notes, err := h.api.QueryNotes(ctx, t.vec, 5)
if err != nil {
log.Printf("voice: query notes: %v", err)
return "не получилось найти ответ.", true
}
t.notes = notes
noteScores := make([]float64, len(notes))
for i, n := range notes {
noteScores[i] = n.Score
}
if !memory.ConfidentScores(noteScores, h.queryMinScore, h.queryMinMargin) {
return "", false
}
texts := make([]string, len(notes))
for i, n := range notes {
texts[i] = n.Text
}
reply, err := h.phraser.PhraseQuery(ctx, t.dec.Utterance, texts)
if err != nil {
log.Printf("voice: phrase query: %v", err)
}
if reply == "" {
reply = "вот что я нашла: " + texts[0]
}
return reply, true
}
// queryGeneral — general knowledge from the phraser, the last source before
// giving up. It always claims: either the model answers or Maven says she
// doesn't know.
func (h *reactiveHandler) queryGeneral(ctx context.Context, t *queryTurn) (string, bool) {
reply, err := h.phraser.PhraseQuery(ctx, t.dec.Utterance, nil)
if err != nil || reply == "" {
return "не знаю.", true
}
return reply, true
}
+33
View File
@@ -0,0 +1,33 @@
package main
import (
"context"
"log"
"github.com/kami/maven/internal/router"
)
// actionReminder handles router.IntentReminder: parse the time when stage-0
// skipped the extractor, then create the reminder.
func (h *reactiveHandler) actionReminder(ctx context.Context, dec router.Decision) string {
if !dec.Slots.HasTime {
// Stage-0 (reminder-wakeword grammar) skips the extractor, so the
// time wasn't parsed. Run the parser as a fallback.
if dec.Stage == 0 && h.timeParser != nil {
t, ok, err := h.timeParser.Parse(ctx, dec.Utterance, h.now())
if err == nil && ok {
dec.Slots.Time = t
dec.Slots.HasTime = true
}
}
if !dec.Slots.HasTime {
return "не получилось разобрать время напоминания."
}
}
payload := `{"text":` + jsonString(dec.Utterance) + `}`
if _, err := h.api.CreateReminder(ctx, dec.Slots.Time, payload, ""); err != nil {
log.Printf("voice: create reminder: %v", err)
return "не получилось поставить напоминание."
}
return ""
}
+99
View File
@@ -0,0 +1,99 @@
package main
import (
"context"
"log"
"strings"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
)
// Task capture on the voice/chat path (Vikunja #130).
//
// Two halves, both deliberately small:
//
// - captureTaskFromNote runs at the top of actionNote. An utterance that
// explicitly files a task ("добавь в задачи купить молоко") goes to the task
// store instead of the note store. Anything without an explicit marker is
// still a note — see router.ParseTaskCapture for why "надо бы поспать" must
// not become a task.
// - queryTasks is a query source that reads the list back.
//
// Nothing here speaks unprompted. Tasks are answered when asked about; no tick
// rule reads the table.
// captureTaskFromNote claims the turn when the utterance explicitly files a
// task, returning the reply. ("", false) hands the turn back to the note path.
func (h *reactiveHandler) captureTaskFromNote(ctx context.Context, dec router.Decision) (string, bool) {
text, ok := router.ParseTaskCapture(dec.Utterance)
if !ok {
return "", false
}
resp, err := h.api.CaptureTask(ctx, ipc.CaptureTaskReq{
Text: text,
Source: "tap:voice",
Status: store.TaskOpen, // he stated it himself — not a candidate
Ts: h.now(),
})
if err != nil {
log.Printf("voice: capture task: %v", err)
return "не получилось записать задачу.", true
}
if !resp.Created {
return "это уже в списке.", true
}
return "записала: " + text, true
}
// queryTasks — "какие у меня задачи?", "что мне нужно сделать?".
//
// Reads the live set and recites it. Newest first, which is the order the store
// returns: this source has no opinion about which task matters more, and
// pretending otherwise would be a guess. Ranking is Vikunja #129.
func (h *reactiveHandler) queryTasks(ctx context.Context, t *queryTurn) (string, bool) {
if !router.IsTaskListQuery(t.dec.Utterance) {
return "", false
}
tasks, err := h.api.ListTasks(ctx, "live")
if err != nil {
log.Printf("voice: list tasks: %v", err)
return "не получилось посмотреть задачи.", true
}
return formatTaskListRU(tasks), true
}
// formatTaskListRU renders the live task list the way Maven says it. Candidates
// are named as candidates — a task she pulled out of his mail is something she
// suggests, and saying it in the same breath as work he actually stated would
// put words in his mouth.
func formatTaskListRU(tasks []ipc.Task) string {
var open, cands []string
for _, t := range tasks {
switch t.Status {
case store.TaskCandidate:
cands = append(cands, t.Text)
default:
open = append(open, t.Text)
}
}
if len(open) == 0 && len(cands) == 0 {
return "задач нет."
}
var b strings.Builder
if len(open) > 0 {
b.WriteString("в списке: ")
b.WriteString(strings.Join(open, "; "))
b.WriteString(".")
}
if len(cands) > 0 {
if b.Len() > 0 {
b.WriteString(" ")
}
b.WriteString("ещё я нашла, но ты не подтвердил: ")
b.WriteString(strings.Join(cands, "; "))
b.WriteString(".")
}
return b.String()
}
+188
View File
@@ -0,0 +1,188 @@
package main
import (
"context"
"errors"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
)
// taskAPI answers only the three task methods; every other call is
// unimplemented, which is the assertion that capture needs nothing else — in
// particular no embedder, so a filed task costs no model call.
type taskAPI struct {
ipc.UnimplementedCoreAPI
captured []ipc.CaptureTaskReq
created bool
capErr error
tasks []ipc.Task
listArg string
listErr error
}
func (a *taskAPI) CaptureTask(_ context.Context, req ipc.CaptureTaskReq) (ipc.CaptureTaskResp, error) {
a.captured = append(a.captured, req)
if a.capErr != nil {
return ipc.CaptureTaskResp{}, a.capErr
}
return ipc.CaptureTaskResp{ID: 1, Created: a.created}, nil
}
func (a *taskAPI) ListTasks(_ context.Context, status string) ([]ipc.Task, error) {
a.listArg = status
return a.tasks, a.listErr
}
func taskNow() time.Time { return time.Date(2026, 8, 1, 9, 0, 0, 0, time.UTC) }
func taskHandler(api ipc.CoreAPI) *reactiveHandler {
return &reactiveHandler{api: api, now: taskNow}
}
func TestCaptureTaskFromNoteFilesTheTask(t *testing.T) {
api := &taskAPI{created: true}
h := taskHandler(api)
reply, ok := h.captureTaskFromNote(context.Background(), router.Decision{
Intent: router.IntentNote, Utterance: "добавь в задачи купить молоко",
})
if !ok {
t.Fatal("an explicit capture must claim the turn")
}
if len(api.captured) != 1 {
t.Fatalf("captured %d, want 1", len(api.captured))
}
got := api.captured[0]
if got.Text != "купить молоко" {
t.Errorf("text = %q, want the marker stripped", got.Text)
}
if got.Source != "tap:voice" {
t.Errorf("source = %q, want tap:voice", got.Source)
}
if got.Status != "open" {
t.Errorf("status = %q — work he stated is open, never a candidate", got.Status)
}
if !got.Ts.Equal(taskNow()) {
t.Errorf("ts = %v, want the handler clock", got.Ts)
}
if !strings.Contains(reply, "купить молоко") {
t.Errorf("reply = %q, want it to read the task back", reply)
}
}
// A note is still a note: capture only fires on an explicit marker, so
// ordinary recall is untouched.
func TestCaptureTaskFromNotePassesOrdinaryNotes(t *testing.T) {
api := &taskAPI{}
h := taskHandler(api)
for _, u := range []string{"надо бы поспать", "мне понравился этот фильм", "запиши что я пил воду"} {
if _, ok := h.captureTaskFromNote(context.Background(), router.Decision{Utterance: u}); ok {
t.Errorf("%q was captured as a task", u)
}
}
if len(api.captured) != 0 {
t.Errorf("captured %d requests, want none", len(api.captured))
}
}
func TestCaptureTaskFromNoteSaysAlreadyOnTheList(t *testing.T) {
h := taskHandler(&taskAPI{created: false})
reply, ok := h.captureTaskFromNote(context.Background(), router.Decision{Utterance: "добавь в задачи купить молоко"})
if !ok {
t.Fatal("expected the capture path to claim it")
}
if !strings.Contains(reply, "уже") {
t.Errorf("reply = %q — a deduped capture must not claim it saved something new", reply)
}
}
func TestCaptureTaskFromNoteReportsStoreFailure(t *testing.T) {
h := taskHandler(&taskAPI{capErr: errors.New("db is on fire")})
reply, ok := h.captureTaskFromNote(context.Background(), router.Decision{Utterance: "добавь задачу починить кран"})
if !ok {
t.Fatal("a failed capture still claims the turn — the note path must not double-write")
}
if !strings.Contains(reply, "не получилось") {
t.Errorf("reply = %q, want an honest failure", reply)
}
}
func TestQueryTasksRecitesTheLiveList(t *testing.T) {
api := &taskAPI{tasks: []ipc.Task{
{ID: 1, Text: "купить молоко", Status: "open"},
{ID: 2, Text: "продлить страховку", Status: "candidate"},
}}
h := taskHandler(api)
reply, ok := h.queryTasks(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "какие у меня задачи?"},
})
if !ok {
t.Fatal("the task source must claim a task-list question")
}
if api.listArg != "live" {
t.Errorf("ListTasks(%q), want \"live\" — a resolved task is not outstanding work", api.listArg)
}
if !strings.Contains(reply, "купить молоко") || !strings.Contains(reply, "продлить страховку") {
t.Errorf("reply = %q, want both tasks", reply)
}
// The candidate must be named as unconfirmed, not recited as his work.
openIdx := strings.Index(reply, "купить молоко")
candIdx := strings.Index(reply, "продлить страховку")
if !(openIdx < candIdx) {
t.Errorf("reply = %q, want confirmed work before candidates", reply)
}
if !strings.Contains(reply, "не подтвердил") {
t.Errorf("reply = %q, want the candidate flagged as unconfirmed", reply)
}
}
func TestQueryTasksEmptyList(t *testing.T) {
h := taskHandler(&taskAPI{})
reply, ok := h.queryTasks(context.Background(), &queryTurn{
dec: router.Decision{Utterance: "что мне нужно сделать?"},
})
if !ok {
t.Fatal("expected the task source to claim it")
}
if reply != "задач нет." {
t.Errorf("reply = %q", reply)
}
}
func TestQueryTasksPassesOtherQuestions(t *testing.T) {
api := &taskAPI{}
h := taskHandler(api)
for _, u := range []string{"как дела?", "какая погода в москве?", "что у меня сегодня?"} {
if _, ok := h.queryTasks(context.Background(), &queryTurn{dec: router.Decision{Utterance: u}}); ok {
t.Errorf("the task source claimed %q", u)
}
}
if api.listArg != "" {
t.Error("a non-task question must not read the task list")
}
}
// The chain must reach the task source before the recall sources, or "что мне
// нужно сделать?" gets answered by whatever note is nearest.
func TestQuerySourcesOrderTasksBeforeRecall(t *testing.T) {
var tasksAt, notesAt = -1, -1
for i, src := range querySources {
switch src.name {
case "tasks":
tasksAt = i
case "notes":
notesAt = i
}
}
if tasksAt < 0 || notesAt < 0 {
t.Fatalf("sources missing: tasks=%d notes=%d", tasksAt, notesAt)
}
if tasksAt > notesAt {
t.Errorf("tasks source at %d, after notes at %d", tasksAt, notesAt)
}
}
+57 -6
View File
@@ -3,6 +3,8 @@ package main
import ( import (
"context" "context"
"log" "log"
"math/rand"
"strings"
"time" "time"
"github.com/kami/maven/internal/dialogue" "github.com/kami/maven/internal/dialogue"
@@ -21,8 +23,12 @@ const clarifyTTL = 90 * time.Second
// raw utterance, chat and system have nothing to fill in. For those a clarify // raw utterance, chat and system have nothing to fill in. For those a clarify
// decision keeps the canned "не поняла" reply — inventing a question for noise // decision keeps the canned "не поняла" reply — inventing a question for noise
// is worse than admitting she missed it. // is worse than admitting she missed it.
// A reminder wants BOTH what to remind about and when. Subject first: "напомни
// в 11" has a time and nothing to say at 11, and a reminder with no subject is
// not worth setting. Order here is the order she asks in — she still only asks
// about the first one missing.
var wantedSlots = map[router.Intent][]dialogue.Slot{ var wantedSlots = map[router.Intent][]dialogue.Slot{
router.IntentReminder: {dialogue.SlotTime}, router.IntentReminder: {dialogue.SlotText, dialogue.SlotTime},
router.IntentFact: {dialogue.SlotKey}, router.IntentFact: {dialogue.SlotKey},
router.IntentAct: {dialogue.SlotFn}, router.IntentAct: {dialogue.SlotFn},
} }
@@ -35,7 +41,8 @@ var wantedSlots = map[router.Intent][]dialogue.Slot{
// questions, so there is no gender agreement to get wrong; the feminine // questions, so there is no gender agreement to get wrong; the feminine
// self-reference lives in the reply she gives when she drops the request. // self-reference lives in the reply she gives when she drops the request.
var clarifyQuestions = map[dialogue.Slot]string{ var clarifyQuestions = map[dialogue.Slot]string{
dialogue.SlotTime: "На когда напомнить?", dialogue.SlotTime: "Когда?",
dialogue.SlotText: "О чём напомнить?",
dialogue.SlotKey: "Что записать?", dialogue.SlotKey: "Что записать?",
dialogue.SlotFn: "Что сделать?", dialogue.SlotFn: "Что сделать?",
} }
@@ -45,11 +52,55 @@ var clarifyQuestions = map[dialogue.Slot]string{
// landed. Feminine self-reference ("поняла"), as everywhere. // landed. Feminine self-reference ("поняла"), as everywhere.
const clarifyGaveUp = "Прости, я не поняла. Скажи, пожалуйста, по-другому." const clarifyGaveUp = "Прости, я не поняла. Скажи, пожалуйста, по-другому."
// clarifyExpired — his answer came after the TTL, so the parked request is // clarifyExpiredVariants — his answer came after the TTL, so the parked request
// already gone. Same tone as clarifyGaveUp, different reason: too much time // is already gone. Same tone as clarifyGaveUp, different reason: too much time
// passed, not "I did not understand". Feminine self-reference ("ждала", // passed, not "I did not understand". Feminine self-reference ("ждала",
// "отпустила"); he is addressed with a plain imperative. // "отпустила"); he is addressed with a plain imperative.
const clarifyExpired = "Прости, я слишком долго ждала ответа и отпустила прошлую просьбу. Если она ещё нужна, скажи заново." //
// Five phrasings, not one. This is the line he hears whenever he walks off
// mid-request, so it is the line that repeats most — and the same sentence every
// time is what makes a house assistant sound like a kiosk. They all carry the
// same two facts (the old request is gone; say it again if it still matters),
// because the wording may vary and the meaning may not.
//
// Fixed templates rather than model output, for the same reason as
// clarifyQuestions: this text has to be right every time, and it is not worth a
// generation to say something this small.
var clarifyExpiredVariants = []string{
"Прости, я слишком долго ждала ответа и отпустила прошлую просьбу. Если она ещё нужна, скажи заново.",
"Кажется, прошлая просьба уже не важна — я её отпустила. Если я ошибаюсь, повтори.",
"Ты как-то резко замолчал, и я не стала ждать дальше. Если та просьба ещё нужна, скажи заново.",
"Я не дождалась ответа и убрала прошлую просьбу. Повтори, если она всё ещё нужна.",
"Столько времени прошло, что я отпустила прошлую просьбу. Скажи заново, если она в силе.",
}
// clarifyExpiredLine picks one of them at random.
func clarifyExpiredLine() string {
return clarifyExpiredVariants[rand.Intn(len(clarifyExpiredVariants))]
}
// isClarifyExpired reports whether s opens with any of the expiry lines. The
// notice is glued in front of this turn's reply (see withNotice), so a caller
// checking for it has to match a prefix, not the whole string.
func isClarifyExpired(s string) bool {
for _, v := range clarifyExpiredVariants {
if strings.HasPrefix(s, v) {
return true
}
}
return false
}
// trimClarifyExpired strips a leading expiry notice, leaving this turn's actual
// reply. "" ⇒ the notice was the whole thing.
func trimClarifyExpired(s string) string {
for _, v := range clarifyExpiredVariants {
if strings.HasPrefix(s, v) {
return strings.TrimSpace(strings.TrimPrefix(s, v))
}
}
return strings.TrimSpace(s)
}
// clarifyExpiredNotice returns that line when a parked question had just timed // clarifyExpiredNotice returns that line when a parked question had just timed
// out, and "" when nothing was parked. Call it right after // out, and "" when nothing was parked. Call it right after
@@ -63,7 +114,7 @@ func (h *reactiveHandler) clarifyExpiredNotice() string {
return "" return ""
} }
log.Printf("voice: clarify — parked question expired, telling him and routing the words fresh") log.Printf("voice: clarify — parked question expired, telling him and routing the words fresh")
return clarifyExpired return clarifyExpiredLine()
} }
// withNotice glues the expiry notice in front of this turn's reply. One turn // withNotice glues the expiry notice in front of this turn's reply. One turn
+10 -7
View File
@@ -56,10 +56,13 @@ func TestClarifyQuestionForMissingSlot(t *testing.T) {
want string want string
asked bool asked bool
}{ }{
{"reminder without a time", clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"), "На когда напомнить?", true}, {"reminder without a time", clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"), "Когда?", true},
{"fact without a key", clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши"), "Что записать?", true}, {"fact without a key", clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши"), "Что записать?", true},
{"act without a fn", clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это"), "Что сделать?", true}, {"act without a fn", clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это"), "Что сделать?", true},
{"reminder that already has a time", clarifyDec(router.IntentReminder, router.Slots{HasTime: true}, "напомни в 11"), "", false}, // A time with nothing to say at that time is still half a reminder, so
// the subject is what she asks about — not silence.
{"reminder that has a time but no subject", clarifyDec(router.IntentReminder, router.Slots{HasTime: true}, "напомни в 11"), "О чём напомнить?", true},
{"reminder that has both", clarifyDec(router.IntentReminder, router.Slots{Text: "позвонить маме", HasTime: true}, "напомни в 11 позвонить маме"), "", false},
{"chat is never worth a question", clarifyDec(router.IntentChat, router.Slots{Text: "мгм"}, "мгм"), "", false}, {"chat is never worth a question", clarifyDec(router.IntentChat, router.Slots{Text: "мгм"}, "мгм"), "", false},
{"query is never worth a question", clarifyDec(router.IntentQuery, router.Slots{Text: "а"}, "а"), "", false}, {"query is never worth a question", clarifyDec(router.IntentQuery, router.Slots{Text: "а"}, "а"), "", false},
} }
@@ -78,7 +81,7 @@ func TestClarifyReminderCompletesOnAnswer(t *testing.T) {
h, st, _ := newClarifyHandler(t) h, st, _ := newClarifyHandler(t)
question, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме")) question, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"))
if !asked || question != "На когда напомнить?" { if !asked || question != "Когда?" {
t.Fatalf("expected the time question, got %q asked=%v", question, asked) t.Fatalf("expected the time question, got %q asked=%v", question, asked)
} }
@@ -152,7 +155,7 @@ func TestClarifyAsksThreeTimesThenSaysSo(t *testing.T) {
if !handled { if !handled {
t.Fatalf("answer %d must be consumed as an answer", i) t.Fatalf("answer %d must be consumed as an answer", i)
} }
if reply != "На когда напомнить?" { if reply != "Когда?" {
t.Fatalf("attempt %d should ask again, got %q", i, reply) t.Fatalf("attempt %d should ask again, got %q", i, reply)
} }
if h.clarifyStore.Get(voiceDialogueID, h.now()) == nil { if h.clarifyStore.Get(voiceDialogueID, h.now()) == nil {
@@ -304,17 +307,17 @@ func TestClarifyExpiryIsAnnouncedAndWordsStillRoute(t *testing.T) {
*now = now.Add(clarifyTTL + time.Second) *now = now.Add(clarifyTTL + time.Second)
reply := h.handleText(ctx, "как дела") reply := h.handleText(ctx, "как дела")
if !strings.HasPrefix(reply, clarifyExpired) { if !isClarifyExpired(reply) {
t.Fatalf("expired question must be announced first, got %q", reply) t.Fatalf("expired question must be announced first, got %q", reply)
} }
if strings.TrimSpace(strings.TrimPrefix(reply, clarifyExpired)) == "" { if trimClarifyExpired(reply) == "" {
t.Fatalf("the new words must still be answered, got only the notice: %q", reply) t.Fatalf("the new words must still be answered, got only the notice: %q", reply)
} }
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil { if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
t.Fatal("the expired question must be gone") t.Fatal("the expired question must be gone")
} }
// The notice is said once, not on every later utterance. // The notice is said once, not on every later utterance.
if reply := h.handleText(ctx, "как дела"); strings.Contains(reply, clarifyExpired) { if reply := h.handleText(ctx, "как дела"); isClarifyExpired(reply) {
t.Fatalf("notice repeated on a later turn: %q", reply) t.Fatalf("notice repeated on a later turn: %q", reply)
} }
} }
+216
View File
@@ -0,0 +1,216 @@
package main
import (
"context"
"log"
"strings"
"time"
"github.com/kami/maven/internal/router"
)
// pendingHexisExec — a mutating Hexis capability parked awaiting a spoken
// confirm. The confirmation is bound to the resolved capability + canonical
// target entity so a later "да" can only execute exactly what was proposed
// (ecosystem invariant: protected actions require bound confirmation).
type pendingHexisExec struct {
capabilityID string
capName string
entityID string
displayName string
expiry time.Time
}
// pendingRoutineConfirm — a proposed routine awaiting a spoken y/n to become
// a recurring reminder. Set by detectPattern after creating a proposal.
type pendingRoutineConfirm struct {
routineID int64
action string
object string
interval float64
phrase string
expiry time.Time
}
// pendingAct — a destructive act awaiting a spoken confirm.
type pendingAct struct {
fn string
args []string
phrase string
expiry time.Time
}
// confirmTTL — how long a parked destructive confirm stays answerable. Short:
// a confirm is a same-breath gesture; a stale prompt shouldn't fire on an
// unrelated later "да".
const confirmTTL = 90 * time.Second
// park stores a destructive act awaiting confirmation. Overwrites any prior
// pending (last-asked wins — single-user box).
func (h *reactiveHandler) park(fn string, args []string, phrase string) {
h.mu.Lock()
h.pending = &pendingAct{fn: fn, args: args, phrase: phrase, expiry: h.now().Add(confirmTTL)}
h.mu.Unlock()
}
// resolveConfirm interprets an utterance as the answer to a parked destructive
// act OR a parked routine proposal. Returns (reply, true) when it consumed the
// utterance as a y/n answer; ("", false) when there's nothing pending (or the
// parked act expired), so the caller routes the utterance normally. An
// unrecognised answer cancels the pending and routes normally — a confirm that
// can't be answered clearly is safer abandoned than left armed.
func (h *reactiveHandler) resolveConfirm(ctx context.Context, text string) (string, bool) {
h.mu.Lock()
defer h.mu.Unlock()
for _, r := range h.confirmResolvers(ctx) {
if !r.claim() {
continue
}
// The slot is already cleared by claim(): every branch below drops the
// pending, including the unclear one — a confirm that can't be
// answered clearly is safer abandoned than left armed.
switch classifyConfirm(text) {
case confirmYes:
return r.yes(), true
case confirmNo:
return r.no(), true
default:
return "", false
}
}
return "", false
}
// confirmResolver — one parked-confirm slot in the chain. claim() reports
// whether this slot holds a live pending, taking it (and dropping an expired
// one) as it goes; yes/no then run the answer. Only ever called with h.mu held.
type confirmResolver struct {
claim func() bool
yes func() string
no func() string
}
// confirmResolvers builds the ordered chain resolveConfirm walks. Order is
// deliberate: the routine proposal is checked before the tool confirm so a
// routine confirm doesn't get eaten by a stale tool pending.
func (h *reactiveHandler) confirmResolvers(ctx context.Context) []confirmResolver {
var pr *pendingRoutineConfirm
var hx *pendingHexisExec
var p *pendingAct
return []confirmResolver{
// Routine proposal.
{
claim: func() bool {
pr, h.pendingRoutine = h.pendingRoutine, nil
return pr != nil && !h.now().After(pr.expiry)
},
yes: func() string {
// Only record the acceptance. The tick loop reads accepted
// routines and nudges on their own interval. Building a
// reminder here made a routine fire exactly once (Vikunja #366).
if err := h.dataStore.AcceptProposedRoutine(ctx, pr.routineID, h.now()); err != nil {
log.Printf("voice: accept proposed routine: %v", err)
return "не получилось запомнить рутину."
}
return "буду напоминать."
},
no: func() string {
if err := h.dataStore.DismissProposedRoutine(ctx, pr.routineID); err != nil {
log.Printf("voice: dismiss proposed routine: %v", err)
}
return "хорошо, не буду."
},
},
// Hexis execution confirm. Bound to the exact capability + target that
// was proposed; a stray "да" can only run that, nothing else.
{
claim: func() bool {
hx, h.pendingHexis = h.pendingHexis, nil
return hx != nil && !h.now().After(hx.expiry)
},
yes: func() string {
return h.execHexis(ctx, hx.capabilityID, hx.capName, hx.entityID, hx.displayName)
},
no: func() string { return "отменила." },
},
// Tool confirm.
{
claim: func() bool {
p, h.pending = h.pending, nil
return p != nil && !h.now().After(p.expiry)
},
yes: func() string {
out, err := h.tools.Exec(ctx, p.fn, p.args, true) // confirmed
if err != nil {
log.Printf("voice: tool %s (confirmed): %v", p.fn, err)
if out != "" {
return "не получилось выполнить команду: " + firstLine(out)
}
return "не получилось выполнить команду."
}
if out != "" {
return "готово: " + firstLine(out)
}
return "готово."
},
no: func() string { return "отменила." },
},
}
}
// proposeGap scaffolds a 'proposed' tool for an act whose verb isn't enabled.
// maven drafts the registration (name = the verb, provenance = the utterance);
// a human enables it on the authed surface. She suggests, never enables.
func (h *reactiveHandler) proposeGap(ctx context.Context, dec router.Decision) string {
name := firstWord(stripWake(dec.Utterance))
if name == "" {
return "не разобрала команду — попробуй иначе."
}
newly, err := h.api.ProposeTool(ctx, name, dec.Utterance, "", h.now())
if err != nil {
log.Printf("voice: propose tool %q: %v", name, err)
return "команды «" + name + "» нет в списке разрешённых."
}
if newly {
return "команды «" + name + "» нет в списке. Предложила её добавить — включи через клиент."
}
return "команды «" + name + "» пока нет в списке — она уже предложена, включи через клиент."
}
// confirmVerdict — the parse of a y/n confirm answer.
type confirmVerdict int
const (
confirmUnknown confirmVerdict = iota
confirmYes
confirmNo
)
// classifyConfirm reads a short ru/en yes-or-no answer. Substring match on the
// stems so inflections/fillers ("да, давай", "нет, отмени") still land.
func classifyConfirm(text string) confirmVerdict {
t := strings.ToLower(strings.TrimSpace(text))
// negatives first — "не надо" contains no "да", but check no-stems before
// yes so a leading "нет" isn't shadowed.
for _, no := range []string{"нет", "не надо", "отмен", "стоп", "no", "cancel", "stop", "don't"} {
if strings.Contains(t, no) {
return confirmNo
}
}
for _, yes := range []string{"да", "ага", "давай", "подтвер", "конечно", "yes", "yeah", "yep", "confirm", "ок", "okay", "ok"} {
if strings.Contains(t, yes) {
return confirmYes
}
}
return confirmUnknown
}
// actPhrase renders "fn arg1 arg2" for the confirm prompt.
func actPhrase(fn string, args []string) string {
if len(args) == 0 {
return fn
}
return fn + " " + strings.Join(args, " ")
}
+227
View File
@@ -0,0 +1,227 @@
package main
import (
"context"
"errors"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
)
// planAPI answers only DayPlan; every other call is unimplemented, which is
// exactly the assertion that the plan source needs nothing else.
type planAPI struct {
ipc.UnimplementedCoreAPI
plan ipc.DayPlan
err error
calls int
}
func (a *planAPI) DayPlan(context.Context) (ipc.DayPlan, error) {
a.calls++
if a.err != nil {
return ipc.DayPlan{}, a.err
}
return a.plan, nil
}
func planDay() time.Time { return time.Date(2026, 8, 3, 12, 0, 0, 0, time.UTC) }
func samplePlan() ipc.DayPlan {
day := planDay()
mid := time.Date(2026, 8, 3, 0, 0, 0, 0, time.UTC)
return ipc.DayPlan{
Date: mid,
Items: []ipc.DayPlanItem{
{At: day.Add(-2 * time.Hour), Text: "Standup @ 10:00-10:30", Kind: "event"},
{At: day.Add(2 * time.Hour), Text: "Планёрка @ 14:00-14:30", Kind: "event", Uncertain: true},
{At: day.Add(6 * time.Hour), Text: "позвонить маме", Kind: "reminder"},
},
Spoken: "план на 03.08.2026: 10:00 — Standup @ 10:00-10:30; " +
"похоже, 14:00 — Планёрка @ 14:00-14:30; 18:00 — позвонить маме.",
}
}
func planHandler(api ipc.CoreAPI) *reactiveHandler {
return &reactiveHandler{api: api, now: planDay}
}
func TestQueryDayPlanRecitesTheDay(t *testing.T) {
api := &planAPI{plan: samplePlan()}
h := planHandler(api)
reply, ok := h.queryDayPlan(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "какие планы на сегодня?"},
})
if !ok {
t.Fatal("the plan source must claim a plan question")
}
if reply != api.plan.Spoken {
t.Errorf("reply = %q, want the core's spoken plan %q", reply, api.plan.Spoken)
}
}
// "что дальше?" is the rest of the day, not the whole day: what has already
// happened is not a plan.
func TestQueryDayPlanTrimsToRestOfDay(t *testing.T) {
h := planHandler(&planAPI{plan: samplePlan()})
reply, ok := h.queryDayPlan(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "что дальше?"},
})
if !ok {
t.Fatal("expected the plan source to claim it")
}
if strings.Contains(reply, "Standup") {
t.Errorf("a passed item must not be read back: %q", reply)
}
if !strings.Contains(reply, "Планёрка") || !strings.Contains(reply, "позвонить маме") {
t.Errorf("the rest of the day is missing: %q", reply)
}
// Provenance survives the trim.
if !strings.Contains(reply, "похоже,") {
t.Errorf("a relayed event must stay hedged: %q", reply)
}
}
// A question that is not about the plan must fall through, or the plan buries
// the calendar listing and the weather behind it.
func TestQueryDayPlanPassesOnEverythingElse(t *testing.T) {
for _, q := range []string{
"что у меня сегодня?",
"какие планы на завтра?",
"когда планёрка?",
"какая погода?",
"",
} {
api := &planAPI{plan: samplePlan()}
reply, ok := planHandler(api).queryDayPlan(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: q},
})
if ok {
t.Errorf("%q was claimed by the plan source (reply %q)", q, reply)
}
if api.calls != 0 {
t.Errorf("%q hit the core for a plan it does not want", q)
}
}
}
func TestQueryDayPlanCoreFailure(t *testing.T) {
h := planHandler(&planAPI{err: errors.New("socket closed")})
reply, ok := h.queryDayPlan(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "план на сегодня"},
})
if !ok {
t.Fatal("a failed plan read must still answer, not fall through to RAG")
}
if reply != "не получилось собрать план." {
t.Errorf("reply = %q", reply)
}
}
// The day plan must sit before the calendar listing: both match "…на сегодня",
// and the more specific matcher has to get first refusal (see #373 for what
// happens when the order is wrong).
func TestDayPlanSourcePrecedesCalendar(t *testing.T) {
plan, cal := -1, -1
for i, s := range querySources {
switch s.name {
case "day-plan":
plan = i
case "calendar":
cal = i
}
}
if plan < 0 || cal < 0 {
t.Fatalf("sources missing: day-plan=%d calendar=%d", plan, cal)
}
if plan > cal {
t.Errorf("day-plan at %d must come before calendar at %d", plan, cal)
}
}
// habitAPI answers only RecentFacts — the whole input the behaviour profile
// needs (Vikunja #254). Nothing is asked of the LLM, so nothing else is wired.
type habitAPI struct {
ipc.UnimplementedCoreAPI
facts []ipc.Fact
err error
calls int
}
func (a *habitAPI) RecentFacts(_ context.Context, _ int) ([]ipc.Fact, error) {
a.calls++
return a.facts, a.err
}
// tuesdayFacts — n weekly Tuesday rows for key, ending before now.
func tuesdayFacts(key string, hh, weeks int, now time.Time) []ipc.Fact {
d := now
for d.Weekday() != time.Tuesday {
d = d.AddDate(0, 0, -1)
}
var out []ipc.Fact
for i := 0; i < weeks; i++ {
day := d.AddDate(0, 0, -7*i)
out = append(out, ipc.Fact{
Ts: time.Date(day.Year(), day.Month(), day.Day(), hh, 0, 0, 0, now.Location()),
Kind: "self",
Key: key,
})
}
return out
}
func TestQueryHabitsAnswersFromCountedFacts(t *testing.T) {
now := planDay() // a Monday
api := &habitAPI{facts: tuesdayFacts("workout", 19, 4, now)}
h := &reactiveHandler{api: api, now: func() time.Time { return now }}
reply, ok := h.queryHabits(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "что я обычно делаю по вторникам?"},
})
if !ok {
t.Fatal("the habit source must claim a habit question")
}
if want := "по вторникам ты обычно тренируешься около 19:00."; reply != want {
t.Errorf("reply = %q, want %q", reply, want)
}
}
func TestQueryHabitsPassesOnEverythingElse(t *testing.T) {
now := planDay()
for _, q := range []string{"что я делаю в среду?", "что у меня сегодня?", "какие планы на сегодня?", ""} {
api := &habitAPI{}
h := &reactiveHandler{api: api, now: func() time.Time { return now }}
if reply, ok := h.queryHabits(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: q},
}); ok {
t.Errorf("%q was claimed by the habit source (reply %q)", q, reply)
}
if api.calls != 0 {
t.Errorf("%q scanned the fact log for a profile it does not want", q)
}
}
}
// Both specific sources must precede the calendar listing, which matches any
// utterance naming a day.
func TestHabitSourcePrecedesCalendar(t *testing.T) {
habits, cal := -1, -1
for i, s := range querySources {
switch s.name {
case "habits":
habits = i
case "calendar":
cal = i
}
}
if habits < 0 || cal < 0 {
t.Fatalf("sources missing: habits=%d calendar=%d", habits, cal)
}
if habits > cal {
t.Errorf("habits at %d must come before calendar at %d", habits, cal)
}
}
+158
View File
@@ -0,0 +1,158 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/store"
)
// Vikunja #281 — the fourth delivery outcome: a care candidate the restraint
// gate suppresses (quiet hours / away / calendar-busy) is not necessarily
// lost. If it's worth resurfacing (loop.DigestEligible), it's durably held
// (internal/store's digest_entries) and spoken as one bundle once speaking
// is appropriate again — never while the suppression reason still holds.
func breakTrace(blockedBy string) *loop.TickTrace {
return &loop.TickTrace{
RuleTraces: []loop.RuleTrace{{
RuleName: "break",
Severity: loop.Sev2,
PredicateResult: true,
GateResult: false,
GateBlockedBy: blockedBy,
}},
}
}
// TestSuppressedCareDigestsAcrossQuietHours — a Sev2 care candidate blocked
// by quiet hours is enqueued into the durable digest, and is spoken as a
// "digest" nudge only once quiet hours actually end — never while still
// suppressed (that would just be a second way to nag through quiet hours).
func TestSuppressedCareDigestsAcrossQuietHours(t *testing.T) {
st := newTestStore(t)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
ctx := context.Background()
now := refNow()
quiet := loop.State{Now: now, QuietHours: true, Presence: store.Present}
tl.enqueueSuppressedDigest(ctx, breakTrace("quiet_hours"), quiet, now)
entries, err := st.PendingDigestEntries(ctx, now)
if err != nil {
t.Fatalf("pending: %v", err)
}
if len(entries) != 1 || entries[0].Rule != "break" {
t.Fatalf("want 1 pending digest entry for break, got %+v", entries)
}
// still quiet hours: draining now must not speak — the same restraint
// that suppressed the live nudge must suppress the bundle too.
tl.maybeDrainDigest(ctx, quiet, now)
if len(sink.sends) != 0 {
t.Fatalf("digest must not drain while quiet hours holds, got %+v", sink.sends)
}
// quiet hours end: this is the moment speaking is appropriate again.
after := now.Add(time.Hour)
clear := loop.State{Now: after, QuietHours: false, Presence: store.Present}
tl.maybeDrainDigest(ctx, clear, after)
if len(sink.sends) != 1 {
t.Fatalf("want exactly 1 dispatched digest bundle, got %d: %+v", len(sink.sends), sink.sends)
}
if sink.sends[0].RuleName != "digest" {
t.Fatalf("want RuleName digest, got %q", sink.sends[0].RuleName)
}
remaining, err := st.PendingDigestEntries(ctx, after)
if err != nil {
t.Fatalf("pending after drain: %v", err)
}
if len(remaining) != 0 {
t.Fatalf("drained entry must no longer be pending, got %+v", remaining)
}
}
// TestSuppressedCareDigestDedupesAcrossTicks — quiet hours holding for
// several ticks must not enqueue several copies of the same suppressed
// nudge; he hears it once when the bundle finally drains.
func TestSuppressedCareDigestDedupesAcrossTicks(t *testing.T) {
st := newTestStore(t)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
ctx := context.Background()
now := refNow()
quiet := loop.State{Now: now, QuietHours: true, Presence: store.Present}
for i := 0; i < 3; i++ {
tl.enqueueSuppressedDigest(ctx, breakTrace("quiet_hours"), quiet, now.Add(time.Duration(i)*time.Minute))
}
entries, err := st.PendingDigestEntries(ctx, now)
if err != nil {
t.Fatalf("pending: %v", err)
}
if len(entries) != 1 {
t.Fatalf("3 suppressions of the same nudge must collapse to 1 pending entry, got %d", len(entries))
}
}
// TestSuppressedCareDigestExpiresRatherThanDeliveringLate — an entry that
// aged out before the suppression cleared is dropped, not spoken late: a
// two-day-old "you skipped a break" is noise, not news.
func TestSuppressedCareDigestExpiresRatherThanDeliveringLate(t *testing.T) {
st := newTestStore(t)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
ctx := context.Background()
now := refNow()
quiet := loop.State{Now: now, QuietHours: true, Presence: store.Present}
tl.enqueueSuppressedDigest(ctx, breakTrace("quiet_hours"), quiet, now)
// well past digestExpiry (24h) before the suppression ever clears.
stale := now.Add(48 * time.Hour)
tl.expireStaleDigest(ctx, stale)
clear := loop.State{Now: stale, QuietHours: false, Presence: store.Present}
tl.maybeDrainDigest(ctx, clear, stale)
if len(sink.sends) != 0 {
t.Fatalf("a stale digest entry must be dropped, not delivered late; got %+v", sink.sends)
}
}
// TestSuppressedCareDigestIgnoresHighSeverity — defense in depth at the
// wiring layer: even if a RuleTrace somehow showed a high-severity rule
// blocked by a care-only gate reason, the tick driver must not durably
// digest it. Alarms bypass the gate and deliver now, unchanged; they must
// never be silently delayed into a bundle.
func TestSuppressedCareDigestIgnoresHighSeverity(t *testing.T) {
st := newTestStore(t)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
ctx := context.Background()
now := refNow()
trace := &loop.TickTrace{RuleTraces: []loop.RuleTrace{{
RuleName: "service_down",
Severity: loop.Sev4,
PredicateResult: true,
GateResult: false,
GateBlockedBy: "quiet_hours",
}}}
quiet := loop.State{Now: now, QuietHours: true, Presence: store.Present}
tl.enqueueSuppressedDigest(ctx, trace, quiet, now)
entries, err := st.PendingDigestEntries(ctx, now)
if err != nil {
t.Fatalf("pending: %v", err)
}
if len(entries) != 0 {
t.Fatalf("high severity must never be digested, got %+v", entries)
}
}
+312
View File
@@ -0,0 +1,312 @@
package main
import (
"context"
"encoding/json"
"fmt"
"log"
"strings"
hexisclient "github.com/kami/hexis/pkg/client"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
)
// praxisCapability is one arm of the Praxis act dispatch. This is an interface
// rather than a map[string]func because each arm carries its own state: the
// verb aliases it answers to, the trace name it records, and its own reply
// formatting. The dispatch grows an arm per Praxis capability, so a new one is
// added to praxisCapabilities below and nothing else changes.
type praxisCapability interface {
// aliases are the verbs (router fn slots, EN and RU) this capability answers to.
aliases() []string
// handle runs the capability and returns the user-facing reply.
handle(ctx context.Context, h *reactiveHandler, px *praxisClient, dec router.Decision) string
}
// praxisCapabilities is the registry handlePraxisAct consults, in order.
var praxisCapabilities = []praxisCapability{
listAttentionCapability{},
praxisItemAction{
verbs: []string{"acknowledge_item", "принято", "понял", "поняла"},
ask: "какой пункт отметить принятым?",
op: "acknowledge",
failure: "не получилось отметить принятым.",
success: "принято.",
call: func(ctx context.Context, px *praxisClient, id string) error {
_, err := px.Acknowledge(ctx, id)
return err
},
},
praxisItemAction{
verbs: []string{"resolve_item", "сделано", "готово", "решено"},
ask: "какой пункт отметить сделанным?",
op: "resolve",
failure: "не получилось отметить сделанным.",
success: "отмечено как сделано.",
call: func(ctx context.Context, px *praxisClient, id string) error {
_, err := px.Resolve(ctx, id)
return err
},
},
praxisItemAction{
verbs: []string{"ignore_item", "игнорировать", "неважно"},
ask: "какой пункт игнорировать?",
op: "ignore",
failure: "не получилось проигнорировать.",
success: "проигнорировано.",
call: func(ctx context.Context, px *praxisClient, id string) error {
_, err := px.Ignore(ctx, id)
return err
},
},
praxisItemAction{
verbs: []string{"pin_item", "закрепить"},
ask: "какой пункт закрепить?",
op: "pin",
failure: "не получилось закрепить.",
success: "закреплено.",
call: func(ctx context.Context, px *praxisClient, id string) error {
_, err := px.Pin(ctx, id, true)
return err
},
},
listChangesCapability{},
}
// handlePraxisAct — dispatches ecosystem tool acts through the Praxis tools API.
// Returns "" when the act is not a Praxis verb (the caller falls through to the
// system command executor). Returns a reply string otherwise.
func (h *reactiveHandler) handlePraxisAct(ctx context.Context, dec router.Decision) string {
if h.ecosystem == nil || h.ecosystem.praxis == nil {
return ""
}
px := h.ecosystem.praxis
for _, capability := range praxisCapabilities {
for _, alias := range capability.aliases() {
if alias == dec.Slots.Fn {
return capability.handle(ctx, h, px, dec)
}
}
}
// Not a Praxis verb — let the caller fall through.
return ""
}
// praxisItemAction is the shared shape of the item-lifecycle capabilities: take
// an item id from the value slot, call one Praxis endpoint, trace the result.
type praxisItemAction struct {
verbs []string
ask string // reply when no item id was given
op string // trace + log name of the operation
failure string // reply when the Praxis call errors
success string
call func(ctx context.Context, px *praxisClient, id string) error
}
func (a praxisItemAction) aliases() []string { return a.verbs }
func (a praxisItemAction) handle(ctx context.Context, h *reactiveHandler, px *praxisClient, dec router.Decision) string {
id := dec.Slots.Value
if id == "" {
return a.ask
}
if err := a.call(ctx, px, id); err != nil {
log.Printf("ecosystem: praxis %s %s: %v", a.op, id, err)
return a.failure
}
h.recordPraxisTrace(ctx, a.op, map[string]any{"item_id": id})
return a.success
}
// listAttentionCapability reads the attention digest and surfaces every item it speaks.
type listAttentionCapability struct{}
func (listAttentionCapability) aliases() []string {
return []string{"list_attention", "attention", "внимание", "что требует внимания", "что нового"}
}
func (listAttentionCapability) handle(ctx context.Context, h *reactiveHandler, px *praxisClient, _ router.Decision) string {
items, err := px.ListAttention(ctx, 20)
if err != nil {
log.Printf("ecosystem: praxis attention: %v", err)
return "не могу сейчас узнать, что требует внимания."
}
if len(items) == 0 {
return "ничего не требует внимания."
}
h.recordPraxisTrace(ctx, "list_attention", map[string]any{"count": len(items)})
var parts []string
for _, item := range items {
title, _ := item["title"].(string)
// importance arrives as JSON number ⇒ float64 over the HTTP contract.
importance, _ := item["importance"].(float64)
rule, _ := item["rule"].(string)
s := title
if importance > 0 {
s += fmt.Sprintf(" (важность %d", int(importance))
if rule != "" {
s += ": " + rule
}
s += ")"
}
parts = append(parts, s)
// Speaking an item surfaces it, it does not acknowledge it
// (ECOSYSTEM-SPEC.md §2.3: surfaced != acknowledged). Best-effort:
// a failed surface call must not block delivering the digest.
if id, ok := item["id"].(string); ok && id != "" {
if _, err := px.Surface(ctx, id); err != nil {
log.Printf("ecosystem: praxis surface %s: %v", id, err)
}
}
}
return "требует внимания: " + strings.Join(parts, "; ")
}
// listChangesCapability reads the recent-changes feed.
type listChangesCapability struct{}
func (listChangesCapability) aliases() []string {
return []string{"list_changes", "changes", "изменения", "что изменилось"}
}
func (listChangesCapability) handle(ctx context.Context, h *reactiveHandler, px *praxisClient, _ router.Decision) string {
changes, err := px.ListChanges(ctx, 20)
if err != nil {
log.Printf("ecosystem: praxis changes: %v", err)
return "не могу сейчас узнать об изменениях."
}
if len(changes) == 0 {
return "нет изменений."
}
h.recordPraxisTrace(ctx, "list_changes", map[string]any{"count": len(changes)})
var parts []string
for _, c := range changes {
title, _ := c["title"].(string)
typ, _ := c["change_type"].(string)
parts = append(parts, fmt.Sprintf("%s (%s)", title, typ))
}
return "изменения: " + strings.Join(parts, "; ")
}
// recordPraxisTrace — writes a fact recording a cross-service ecosystem call.
// The fact is stored with source "praxis:trace" so the proactive loop can
// reference it and the dashboard can display recent ecosystem activity.
func (h *reactiveHandler) recordPraxisTrace(ctx context.Context, operation string, details map[string]any) {
now := h.now()
value := operation
if len(details) > 0 {
if b, err := json.Marshal(details); err == nil {
value = operation + " " + string(b)
}
}
_, _ = h.api.WriteFact(ctx, ipc.WriteFactReq{
Ts: now,
Kind: "system",
Key: "praxis:" + operation,
Value: value,
Source: "praxis:trace",
Confidence: 1.0,
})
}
// handleHexisAct — resolves entity references through Nexus and executes
// matching capabilities through Hexis. Returns a reply string when handled,
// or "" to fall through to the system command executor.
func (h *reactiveHandler) handleHexisAct(ctx context.Context, dec router.Decision) string {
if h.ecosystem == nil {
return ""
}
// Resolve the utterance text as an entity reference through Nexus. An
// ambiguous match must stop and clarify — never guess a mutation target.
entityID, displayName, ambiguous, err := h.ecosystem.resolveEntityReference(ctx, dec.Slots.Text, nil)
if err != nil {
// A genuine Nexus dependency failure, not "no such entity" — stop here
// and report degradation rather than silently falling through to the
// local command executor (ECOSYSTEM-SPEC.md: services degrade
// independently, never a silent all-clear).
return "экосистема недоступна, попробуй ещё раз."
}
if len(ambiguous) > 0 {
return "уточни, что именно: " + strings.Join(ambiguous, ", ") + "?"
}
if entityID == "" {
return ""
}
// Discover Hexis capabilities for this entity. A resolved entity with a
// genuine Hexis failure must not be treated as "no capabilities" and
// fall through to unrelated local execution.
caps, err := h.ecosystem.discoverCapabilities(ctx, entityID)
if err != nil {
return "экосистема недоступна, попробуй ещё раз."
}
if len(caps) == 0 {
return ""
}
// Match the user's verb to a capability by name/description. Collect all
// matches: more than one is itself ambiguous, so we ask rather than pick
// the first (ecosystem invariant: no arbitrary target for mutation).
verb := dec.Slots.Fn
if verb == "" {
verb = dec.Slots.Text
}
verbLower := strings.ToLower(verb)
var matches []*hexisclient.Capability
for i, c := range caps {
if strings.Contains(strings.ToLower(c.Name), verbLower) ||
(c.Description != "" && strings.Contains(strings.ToLower(c.Description), verbLower)) {
matches = append(matches, &caps[i])
}
}
if len(matches) == 0 {
return ""
}
if len(matches) > 1 {
var names []string
for _, m := range matches {
names = append(names, m.Name)
}
return "какую команду для " + displayName + ": " + strings.Join(names, ", ") + "?"
}
matched := matches[0]
// Read-only capabilities run immediately; mutating ones are parked for an
// explicit spoken confirm bound to this capability + target.
if !matched.ReadOnly {
h.mu.Lock()
h.pendingHexis = &pendingHexisExec{
capabilityID: matched.ID,
capName: matched.Name,
entityID: entityID,
displayName: displayName,
expiry: h.now().Add(confirmTTL),
}
h.mu.Unlock()
return "выполнить «" + matched.Name + "» для " + displayName + "? скажи «да» или «нет»."
}
return h.execHexis(ctx, matched.ID, matched.Name, entityID, displayName)
}
// execHexis runs a resolved capability and records a cross-service trace with
// the correlation ID. It reports command success, never operational recovery
// (Praxis observes recovery independently).
func (h *reactiveHandler) execHexis(ctx context.Context, capID, capName, entityID, displayName string) string {
correlationID, err := h.ecosystem.executeCapability(ctx, capID, entityID, nil)
if err != nil {
log.Printf("ecosystem: hexis execute error (cor=%s): %v", correlationID, err)
return "не получилось выполнить команду для " + displayName + "."
}
h.recordPraxisTrace(ctx, "hexis:"+capName, map[string]any{
"entity_id": entityID,
"entity_name": displayName,
"capability": capName,
"correlation_id": correlationID,
})
return "команда выполнена для " + displayName + "."
}
+81 -117
View File
@@ -58,6 +58,7 @@ import (
"github.com/kami/maven/internal/delivery/telegramsink" "github.com/kami/maven/internal/delivery/telegramsink"
"github.com/kami/maven/internal/ipc" "github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/loop" "github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/phraser" "github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/store" "github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/webauthn" "github.com/kami/maven/internal/webauthn"
@@ -96,96 +97,6 @@ func main() {
} }
} }
// lockedAPI is a dummy CoreAPI used while the daemon is locked. Every method
// returns errLocked. The wire protocol's StoreAPI methods all go through the
// Server dispatch on CoreAPI, so returning errLocked from each is correct.
type lockedAPI struct{}
var _ ipc.CoreAPI = (*lockedAPI)(nil)
func (l *lockedAPI) WriteFact(ctx context.Context, req ipc.WriteFactReq) (int64, error) {
return 0, errLocked
}
func (l *lockedAPI) LatestFact(ctx context.Context, key string) (ipc.Fact, error) {
return ipc.Fact{}, errLocked
}
func (l *lockedAPI) LatestFactBySource(ctx context.Context, key, source string) (ipc.Fact, error) {
return ipc.Fact{}, errLocked
}
func (l *lockedAPI) Since(ctx context.Context, key string, now time.Time) (time.Duration, error) {
return 0, errLocked
}
func (l *lockedAPI) Presence(ctx context.Context) (ipc.Presence, error) {
return ipc.Presence{}, errLocked
}
func (l *lockedAPI) CreateReminder(ctx context.Context, fire time.Time, payload, cron string) (int64, error) {
return 0, errLocked
}
func (l *lockedAPI) MarkReminder(ctx context.Context, id int64, status string) error {
return errLocked
}
func (l *lockedAPI) ListReminders(ctx context.Context, n int) ([]ipc.Reminder, error) {
return nil, errLocked
}
func (l *lockedAPI) RecordNudge(ctx context.Context, rule, channel, message string, ts time.Time) (int64, error) {
return 0, errLocked
}
func (l *lockedAPI) ResolveNudge(ctx context.Context, id int64, outcome string, ts time.Time) error {
return errLocked
}
func (l *lockedAPI) RecentOutcomes(ctx context.Context, rule string, n int) ([]string, error) {
return nil, errLocked
}
func (l *lockedAPI) RecentFacts(ctx context.Context, n int) ([]ipc.Fact, error) {
return nil, errLocked
}
func (l *lockedAPI) CalendarEvents(ctx context.Context, from, to time.Time) ([]ipc.Fact, error) {
return nil, errLocked
}
func (l *lockedAPI) RecentNudges(ctx context.Context, n int) ([]ipc.Nudge, error) {
return nil, errLocked
}
func (l *lockedAPI) WriteNote(ctx context.Context, ts time.Time, text string, embedding []float32, source string) (int64, error) {
return 0, errLocked
}
func (l *lockedAPI) QueryNotes(ctx context.Context, embedding []float32, k int) ([]ipc.Note, error) {
return nil, errLocked
}
func (l *lockedAPI) RecentNotes(ctx context.Context, n int) ([]ipc.Note, error) {
return nil, errLocked
}
func (l *lockedAPI) ProposeTool(ctx context.Context, name, utterance, scope string, ts time.Time) (bool, error) {
return false, errLocked
}
func (l *lockedAPI) EnableTool(ctx context.Context, name string, cmd []string, destructive bool, scope string, ts time.Time) error {
return errLocked
}
func (l *lockedAPI) DisableTool(ctx context.Context, name string) error { return errLocked }
func (l *lockedAPI) DeleteTool(ctx context.Context, name string) error { return errLocked }
func (l *lockedAPI) ListProposedRoutines(ctx context.Context) ([]ipc.ProposedRoutine, error) {
return nil, errLocked
}
func (l *lockedAPI) DismissProposedRoutine(ctx context.Context, id int64) error { return errLocked }
func (l *lockedAPI) AcceptProposedRoutine(ctx context.Context, id int64) error {
return errLocked
}
func (l *lockedAPI) LookupTool(ctx context.Context, name string) (ipc.Tool, error) {
return ipc.Tool{}, errLocked
}
func (l *lockedAPI) ListTools(ctx context.Context, status string) ([]ipc.Tool, error) {
return nil, errLocked
}
func (l *lockedAPI) RevertFact(ctx context.Context, key string) (int64, error) { return 0, errLocked }
func (l *lockedAPI) Chat(ctx context.Context, text string) (string, error) {
return "", errLocked
}
func (l *lockedAPI) TickTrace(ctx context.Context) (ipc.TickTrace, error) {
return ipc.TickTrace{}, errLocked
}
func (l *lockedAPI) MorningStatus(ctx context.Context) ([]ipc.MorningRoutineStatus, error) {
return nil, errLocked
}
func run(args []string) error { func run(args []string) error {
cfgPath := flag.String("config", defaultConfigPath(), "path to mavend JSON config") cfgPath := flag.String("config", defaultConfigPath(), "path to mavend JSON config")
wrappedKeyPath := flag.String("wrapped-key-file", "", "path to wrapped encryption key blob (enables cold-start unlock)") wrappedKeyPath := flag.String("wrapped-key-file", "", "path to wrapped encryption key blob (enables cold-start unlock)")
@@ -258,6 +169,7 @@ func run(args []string) error {
coreAPI ipc.CoreAPI coreAPI ipc.CoreAPI
eco *ecosystemWiring eco *ecosystemWiring
factWorker *factEnrichmentWorker factWorker *factEnrichmentWorker
evalWorker *memoryEvalWorker // nil ⇒ memory evaluation off (the default)
) )
if !locked { if !locked {
@@ -271,13 +183,14 @@ func run(args []string) error {
phr = phraser.NewStub() phr = phraser.NewStub()
if cfg.Phraser != nil { if cfg.Phraser != nil {
pc := phraser.Config{ pc := phraser.Config{
ModelPath: cfg.Phraser.ModelPath, ModelPath: cfg.Phraser.ModelPath,
BinPath: cfg.Phraser.BinPath, BinPath: cfg.Phraser.BinPath,
Listen: cfg.Phraser.Listen, Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers, NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx, NCtx: cfg.Phraser.NCtx,
Timeout: time.Duration(cfg.Phraser.Timeout), Timeout: time.Duration(cfg.Phraser.Timeout),
Persona: personaFromCfg(cfg), LLMNudges: cfg.Phraser.LLMNudges,
ContextBlock: contextBlockFn(cfg, time.Now),
} }
if pc.BinPath == "" { if pc.BinPath == "" {
pc.BinPath = "llama-server" pc.BinPath = "llama-server"
@@ -349,21 +262,28 @@ func run(args []string) error {
tickInterval := time.Duration(cfg.TickInterval) tickInterval := time.Duration(cfg.TickInterval)
repeatInterval := time.Duration(cfg.RepeatInterval) repeatInterval := time.Duration(cfg.RepeatInterval)
autotuneInterval := time.Duration(cfg.AutotuneInterval) autotuneInterval := time.Duration(cfg.AutotuneInterval)
tl = newTickLoop(st, gatherer, dispatcher, phr, rules, tickInterval, repeatInterval, autotuneInterval, cfg.Digest, routinesFromConfig(cfg.Routines), config.MorningRoutinesFromConfig(cfg.MorningRoutines)) tl = newTickLoop(st, gatherer, dispatcher, phr, rules, tickInterval, repeatInterval, autotuneInterval, cfg.Digest, routinesFromConfig(cfg.Routines), config.MorningRoutinesFromConfig(cfg.MorningRoutines), cfg.PatternProposals)
factWorker = newFactEnrichmentWorker(st, eco, time.Duration(cfg.FactEnrichmentInterval)) factWorker = newFactEnrichmentWorker(st, eco, time.Duration(cfg.FactEnrichmentInterval))
evalWorker = newMemoryEvalWorker(st, phr, cfg)
coreAPI = &daemonAPI{ coreAPI = &daemonAPI{
CoreAPI: ipc.NewStoreAPI(st), CoreAPI: ipc.NewStoreAPI(st),
getTrace: tl.trace, getTrace: tl.trace,
getMorningStatus: func(ctx context.Context) []ipc.MorningRoutineStatus { return tl.morningStatus(ctx, time.Now()) }, getMorningStatus: func(ctx context.Context) []ipc.MorningRoutineStatus { return tl.morningStatus(ctx, time.Now()) },
getDayPlan: func(ctx context.Context) ipc.DayPlan { return tl.dayPlan(ctx, time.Now()) },
} }
if voiceW != nil && voiceW.handler != nil { if voiceW != nil && voiceW.handler != nil {
api := coreAPI.(*daemonAPI) api := coreAPI.(*daemonAPI)
api.chatFn = voiceW.handler.handleText api.chatFn = voiceW.handler.handleText
} }
} else { } else {
// locked mode: dummy CoreAPI that returns errLocked for everything // locked mode: no real store yet, so there's no meaningful CoreAPI to
coreAPI = &lockedAPI{} // serve. srv.Check below is the actual guard — every CoreAPI call is
// refused before it reaches this value. This is just a safe non-nil
// placeholder: if the guard is ever bypassed by a bug, calls land
// here and fail loudly with ipc.ErrNotImplemented instead of a nil
// dereference or, worse, silently succeeding.
coreAPI = ipc.UnimplementedCoreAPI{}
} }
// ----- IPC boundary (core ↔ modules) ----- // ----- IPC boundary (core ↔ modules) -----
@@ -374,7 +294,17 @@ func run(args []string) error {
passkeySess := webauthn.NewPasskeySession(5 * time.Minute) passkeySess := webauthn.NewPasskeySession(5 * time.Minute)
// Set Server.Check — in locked mode, block everything except unlock-path methods. // Set Server.Check — the single authorization guard, run once by
// Server.dispatch before any CoreAPI method is called (see
// internal/ipc/server.go). In locked mode this is the ONLY thing
// standing between an unauthenticated caller and the store: it must
// default-deny, with an explicit allowlist for the two methods the
// unlock flow itself needs (MethodAssertStepUp, MethodUnlock — neither
// of which touches CoreAPI; dispatch handles them directly via
// srv.StepUp/srv.UnlockFn). Forgetting to allowlist a new unlock-path
// method fails safe (denied); forgetting to guard a new CoreAPI method
// is impossible because there is nothing left to forget — every method
// not in the allowlist is refused by construction.
if locked { if locked {
srv.Check = func(ctx context.Context, m ipc.Method, _ json.RawMessage) error { srv.Check = func(ctx context.Context, m ipc.Method, _ json.RawMessage) error {
switch m { switch m {
@@ -441,13 +371,14 @@ func run(args []string) error {
phr = phraser.NewStub() phr = phraser.NewStub()
if cfg.Phraser != nil { if cfg.Phraser != nil {
pc := phraser.Config{ pc := phraser.Config{
ModelPath: cfg.Phraser.ModelPath, ModelPath: cfg.Phraser.ModelPath,
BinPath: cfg.Phraser.BinPath, BinPath: cfg.Phraser.BinPath,
Listen: cfg.Phraser.Listen, Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers, NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx, NCtx: cfg.Phraser.NCtx,
Timeout: time.Duration(cfg.Phraser.Timeout), Timeout: time.Duration(cfg.Phraser.Timeout),
Persona: personaFromCfg(cfg), LLMNudges: cfg.Phraser.LLMNudges,
ContextBlock: contextBlockFn(cfg, time.Now),
} }
if pc.BinPath == "" { if pc.BinPath == "" {
pc.BinPath = "llama-server" pc.BinPath = "llama-server"
@@ -510,14 +441,16 @@ func run(args []string) error {
tickInterval := time.Duration(cfg.TickInterval) tickInterval := time.Duration(cfg.TickInterval)
repeatInterval := time.Duration(cfg.RepeatInterval) repeatInterval := time.Duration(cfg.RepeatInterval)
autotuneInterval := time.Duration(cfg.AutotuneInterval) autotuneInterval := time.Duration(cfg.AutotuneInterval)
tl = newTickLoop(st, gatherer, dispatcher, phr, rules, tickInterval, repeatInterval, autotuneInterval, cfg.Digest, routinesFromConfig(cfg.Routines), config.MorningRoutinesFromConfig(cfg.MorningRoutines)) tl = newTickLoop(st, gatherer, dispatcher, phr, rules, tickInterval, repeatInterval, autotuneInterval, cfg.Digest, routinesFromConfig(cfg.Routines), config.MorningRoutinesFromConfig(cfg.MorningRoutines), cfg.PatternProposals)
factWorker = newFactEnrichmentWorker(st, eco, time.Duration(cfg.FactEnrichmentInterval)) factWorker = newFactEnrichmentWorker(st, eco, time.Duration(cfg.FactEnrichmentInterval))
evalWorker = newMemoryEvalWorker(st, phr, cfg)
// Swap the CoreAPI from lockedAPI to the real store adapter. // Swap the CoreAPI from the locked placeholder to the real store adapter.
newAPI := &daemonAPI{ newAPI := &daemonAPI{
CoreAPI: ipc.NewStoreAPI(st), CoreAPI: ipc.NewStoreAPI(st),
getTrace: tl.trace, getTrace: tl.trace,
getMorningStatus: func(ctx context.Context) []ipc.MorningRoutineStatus { return tl.morningStatus(ctx, time.Now()) }, getMorningStatus: func(ctx context.Context) []ipc.MorningRoutineStatus { return tl.morningStatus(ctx, time.Now()) },
getDayPlan: func(ctx context.Context) ipc.DayPlan { return tl.dayPlan(ctx, time.Now()) },
} }
if voiceW != nil && voiceW.handler != nil { if voiceW != nil && voiceW.handler != nil {
newAPI.chatFn = voiceW.handler.handleText newAPI.chatFn = voiceW.handler.handleText
@@ -548,6 +481,13 @@ func run(args []string) error {
factWorker.run(ctx) factWorker.run(ctx)
}() }()
// Start background memory evaluation (nil unless configured).
if evalWorker != nil {
go func() {
evalWorker.run(ctx)
}()
}
dl.unlock() dl.unlock()
log.Printf("mavend: unlocked via passkey assertion") log.Printf("mavend: unlocked via passkey assertion")
return nil return nil
@@ -586,6 +526,13 @@ func run(args []string) error {
defer wg.Done() defer wg.Done()
factWorker.run(ctx) factWorker.run(ctx)
}() }()
if evalWorker != nil {
wg.Add(1)
go func() {
defer wg.Done()
evalWorker.run(ctx)
}()
}
} }
<-ctx.Done() <-ctx.Done()
@@ -601,12 +548,29 @@ func run(args []string) error {
return nil return nil
} }
// personaFromCfg extracts the voice persona from the config, or returns "" // personaFacts reads the optional, deployment-specific facts (his name, his
// when voice isn't configured. Used to pass a character prompt into the // city, the free-text persona string) out of the config. Everything here may
// LLM phraser without requiring voice to be enabled. // be empty — the context block is correct without any of it.
func personaFromCfg(cfg *config.Config) string { func personaFacts(cfg *config.Config) persona.Facts {
if cfg.Voice != nil { f := persona.Facts{
return cfg.Voice.Persona // Telegram lives outside the voice block, so it counts either way.
Telegram: cfg.Telegram != nil && cfg.Telegram.BotToken != "" && cfg.Telegram.ChatID != "",
} }
return "" if cfg.Voice == nil {
return f
}
f.OwnerName = cfg.Voice.OwnerName
f.City = cfg.Voice.City
f.Static = cfg.Voice.Persona
// Same test wireVoice uses to pick the real provider over the stub.
f.Weather = cfg.Voice.Weather != nil && cfg.Voice.Weather.Provider == "open-meteo"
f.Tools = len(cfg.Voice.Tools) > 0
return f
}
// contextBlockFn returns the per-turn renderer of the shared context block.
// Per turn, not once at startup, because the block states the current time.
func contextBlockFn(cfg *config.Config, now func() time.Time) func() string {
f := personaFacts(cfg)
return func() string { return f.Block(now()) }
} }
+88
View File
@@ -0,0 +1,88 @@
// mavend/memoryeval.go — the driver for background memory evaluation
// (Vikunja #248). The evaluator itself is pure-ish and lives in
// internal/memeval; this is the one impure part: a ticker, the store, and the
// resident model's base URL.
//
// It is its own goroutine and NOT a step on the main tick, deliberately. The
// tick runs every 60s and has a delivery deadline behind it; an evaluation is
// a multi-second LLM round-trip on the same llama-server that answers voice
// turns, and it happens hourly at most. Bolting it onto the tick would make
// every hour's tick the slow one for no benefit.
package main
import (
"context"
"log"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/memeval"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/store"
)
// memoryEvalWorker — ticker + evaluator.
type memoryEvalWorker struct {
eval *memeval.Evaluator
interval time.Duration
}
// newMemoryEvalWorker wires the evaluation loop, or returns nil when it should
// not run at all. nil is the normal case and every caller must handle it:
//
// - no memory_eval config block ⇒ off (a capability is off unless configured);
// - no LLM phraser ⇒ nothing to evaluate with. There is no template fallback
// here on purpose: a "memory evaluation" assembled from string templates
// would be a fixed sentence pretending to be an observation.
func newMemoryEvalWorker(st *store.Store, phr phraser.Phraser, cfg *config.Config) *memoryEvalWorker {
if cfg.MemoryEval == nil {
return nil
}
lp, ok := phr.(*phraser.LLMPhraser)
if !ok {
log.Printf("memory eval: configured but no llama-server phraser — evaluation disabled")
return nil
}
interval := time.Duration(cfg.MemoryEval.Interval)
if interval <= 0 {
interval = config.DefaultMemoryEvalInterval
}
// A generous per-request timeout: this is a long prompt to a Thinking model
// and nobody is waiting on the answer.
client := llm.New(lp.BaseURL(), 5*time.Minute)
ev := memeval.NewEvaluator(st, st, client, memeval.Config{
MaxItems: cfg.MemoryEval.MaxItems,
MinConfidence: cfg.MemoryEval.MinConfidence,
ContextBlock: contextBlockFn(cfg, time.Now),
})
log.Printf("memory eval: enabled, every %s", interval)
return &memoryEvalWorker{eval: ev, interval: interval}
}
// run evaluates every interval until ctx is canceled.
//
// The first evaluation waits a full interval rather than firing at startup, the
// opposite of the tick loop's cold-start behaviour. A tick that fires late is a
// nudge that arrives late; an evaluation that fires late is nothing at all, and
// the alternative is a heavy LLM call competing with startup — including with
// the first voice turn after a restart.
func (w *memoryEvalWorker) run(ctx context.Context) {
ticker := time.NewTicker(w.interval)
defer ticker.Stop()
for {
select {
case <-ctx.Done():
return
case now := <-ticker.C:
obs, err := w.eval.Evaluate(ctx, now)
if err != nil {
log.Printf("memory eval: %v", err)
continue
}
for _, o := range obs {
log.Printf("memory eval: noted (%.2f, %s): %s", o.Conf, o.Action, o.Text)
}
}
}
}
+120
View File
@@ -0,0 +1,120 @@
// mavend/patterns.go — the shared detect+propose step of pattern inference
// (Vikunja #43). Event *extraction* (fact -> action/object) happens at fact-
// write time in detectPattern below, tied to whichever channel wrote the
// fact. Detection — turning a run of events into a proposed routine — is
// channel-agnostic: it only needs what's already in the events table, so it
// runs both right after a voice fact-write (for the immediate "напоминать?"
// confirmation) and, proactively, from the digestion tick (tick.go's
// detectPatterns) over every action+object pair on record, not just the one
// that was just talked about.
package main
import (
"context"
"errors"
"fmt"
"log"
"time"
"github.com/kami/maven/internal/pattern"
"github.com/kami/maven/internal/store"
)
// detectAndPropose runs the pattern detector over every recorded event for
// action+object and, if a stable pattern is found and nothing has been
// proposed/accepted/dismissed for this pair yet, creates a proposed_routines
// row. Returns (nil, 0, nil) — not an error — whenever there is nothing new
// to report: too few events, irregular intervals, or a pair that already has
// a row in any status. That last case is the one that matters most: it is
// how a routine the owner already DISMISSED stays dismissed forever, because
// the row survives dismissal (status flips in place, see
// store.DismissProposedRoutine) and both the Lookup check here and the
// table's UNIQUE(action, object) constraint refuse to create a second one.
func detectAndPropose(ctx context.Context, ds *store.Store, action, object string, ts time.Time) (*pattern.ProposedRoutine, int64, error) {
events, err := ds.EventsFor(ctx, action, object)
if err != nil {
return nil, 0, fmt.Errorf("events for %s/%s: %w", action, object, err)
}
patEvents := make([]pattern.Event, len(events))
for i, e := range events {
patEvents[i] = pattern.Event{
FactID: e.FactID,
Action: e.Action,
Object: e.Object,
Ts: e.Ts,
}
}
r, err := pattern.Detect(patEvents)
if err != nil {
return nil, 0, fmt.Errorf("detect %s/%s: %w", action, object, err)
}
if r == nil {
return nil, 0, nil // not enough data or intervals too irregular
}
// Belt: check first so the common "nothing new" case never even attempts
// an insert. Suspenders: CreateProposedRoutine's ON CONFLICT DO NOTHING
// (backed by the UNIQUE(action,object) constraint) is the actual
// guarantee — this Lookup is an optimization, not the source of truth.
existing, err := ds.LookupProposedRoutine(ctx, r.Action, r.Object)
if err != nil {
return nil, 0, fmt.Errorf("lookup proposed routine %s/%s: %w", action, object, err)
}
if existing != nil {
return nil, 0, nil // already proposed, accepted, or dismissed — say nothing
}
id, err := ds.CreateProposedRoutine(ctx, r.Action, r.Object, r.IntervalDays, ts)
if err != nil {
if errors.Is(err, store.ErrProposedRoutineExists) {
return nil, 0, nil // lost a race with another caller — not an error
}
return nil, 0, fmt.Errorf("create proposed routine %s/%s: %w", action, object, err)
}
return r, id, nil
}
// detectPattern extracts an event from the written fact and runs the pattern
// detector. If a stable recurring pattern is found and no proposed routine
// exists for this action+object yet, one is created and the user is prompted
// to confirm via the park() mechanism. Returns the suggestion phrase when a
// new proposal was created and parked; "" otherwise.
func (h *reactiveHandler) detectPattern(ctx context.Context, factID int64, key, value string, ts time.Time) string {
ev := pattern.Extract(factID, key, value, ts)
if ev == nil {
return "" // not an actionable event
}
if _, err := h.dataStore.CreateEvent(ctx, factID, ev.Action, ev.Object, ts); err != nil {
log.Printf("voice: create event: %v", err)
return ""
}
// Detect+propose (Vikunja #43) is shared with the digestion tick's
// proactive scan — see detectAndPropose above. Event *extraction* stays
// here, tied to this fact write; detection over the accumulated history does
// not need to happen right now for the voice path to have already done
// its job — it's dedupe-safe to also let the next tick find the same
// pattern independently.
r, id, err := detectAndPropose(ctx, h.dataStore, ev.Action, ev.Object, ts)
if err != nil {
log.Printf("voice: detect pattern %s/%s: %v", ev.Action, ev.Object, err)
return ""
}
if r == nil {
return "" // not enough data, too irregular, or already proposed/decided
}
log.Printf("voice: proposed routine: %s/%s every %.1f days", r.Action, r.Object, r.IntervalDays)
// Park the proposal for voice confirmation.
phrase := pattern.PhraseRoutine(r)
h.mu.Lock()
h.pendingRoutine = &pendingRoutineConfirm{
routineID: id,
action: r.Action,
object: r.Object,
interval: r.IntervalDays,
phrase: phrase,
expiry: ts.Add(confirmTTL),
}
h.mu.Unlock()
return phrase
}
+285
View File
@@ -0,0 +1,285 @@
package main
import (
"context"
"database/sql"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/pattern"
"github.com/kami/maven/internal/store"
)
// seedRefillEvents writes N weekly "refill/cat_water" events straight to the
// events table — this is what the tick reads, independent of any utterance.
func seedRefillEvents(t *testing.T, st *store.Store, ctx context.Context, base time.Time, n int) {
t.Helper()
for i := 0; i < n; i++ {
factID, err := st.WriteFact(ctx, base.Add(time.Duration(i)*7*24*time.Hour), store.KindSelf,
"cat_water", "refill", "test", 1.0, sql.NullInt64{})
if err != nil {
t.Fatalf("write fact %d: %v", i, err)
}
if _, err := st.CreateEvent(ctx, factID, "refill", "cat_water", base.Add(time.Duration(i)*7*24*time.Hour)); err != nil {
t.Fatalf("create event %d: %v", i, err)
}
}
}
// TestTickDetectsPatternFromStoredEvents proves the tick notices a pattern on
// its own, reading straight from the store — not as a side effect of a live
// utterance (Vikunja #43). MinEvents weekly events with no voice turn in
// sight must produce exactly one proposed routine.
func TestTickDetectsPatternFromStoredEvents(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents)
tl := newTestTickLoop(t, st, &fakeSink{}, nil)
tl.detectPatterns(ctx, now, loop.State{})
rows, err := st.ListProposedRoutines(ctx)
if err != nil {
t.Fatalf("list proposed routines: %v", err)
}
if len(rows) != 1 {
t.Fatalf("proposed routines = %d, want 1: %+v", len(rows), rows)
}
if rows[0].Action != "refill" || rows[0].Object != "cat_water" {
t.Errorf("proposed routine = %s/%s, want refill/cat_water", rows[0].Action, rows[0].Object)
}
}
// TestTickPatternDetectionIsIdempotent proves running the tick's pattern scan
// twice does not spam a second proposal for the same pair, and that the store
// itself is what stops the duplicate (not tick-local state) — the whole point
// of the guard, since the tick has no memory of what it proposed last time.
func TestTickPatternDetectionIsIdempotent(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents)
tl := newTestTickLoop(t, st, &fakeSink{}, nil)
tl.detectPatterns(ctx, now, loop.State{})
tl.detectPatterns(ctx, now.Add(time.Hour), loop.State{})
rows, err := st.ListProposedRoutines(ctx)
if err != nil {
t.Fatalf("list proposed routines: %v", err)
}
if len(rows) != 1 {
t.Fatalf("proposed routines after two ticks = %d, want 1 (no duplicate): %+v", len(rows), rows)
}
}
// TestTickPatternDetectionRespectsDismissal proves the single worst failure
// mode here — a proposal the owner already said no to coming back on the next
// tick — cannot happen. Dismissal flips the row's status in place; it must
// still be there to block re-proposal.
func TestTickPatternDetectionRespectsDismissal(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents)
tl := newTestTickLoop(t, st, &fakeSink{}, nil)
tl.detectPatterns(ctx, now, loop.State{})
rows, err := st.ListProposedRoutines(ctx)
if err != nil {
t.Fatalf("list proposed routines: %v", err)
}
if len(rows) != 1 {
t.Fatalf("setup: proposed routines = %d, want 1", len(rows))
}
if err := st.DismissProposedRoutine(ctx, rows[0].ID); err != nil {
t.Fatalf("dismiss: %v", err)
}
// More events for the same pair arrive, and the tick runs again — a
// dismissed pattern must not resurface.
seedRefillEvents(t, st, ctx, now.Add(30*24*time.Hour), pattern.MinEvents)
tl.detectPatterns(ctx, now.Add(60*24*time.Hour), loop.State{})
proposed, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineProposed)
if err != nil {
t.Fatalf("list proposed: %v", err)
}
if len(proposed) != 0 {
t.Fatalf("a dismissed pattern came back: %+v", proposed)
}
all, err := st.ListProposedRoutinesByStatus(ctx, "")
if err != nil {
t.Fatalf("list all: %v", err)
}
if len(all) != 1 {
t.Fatalf("total rows for the pair = %d, want 1 (still dismissed, not duplicated): %+v", len(all), all)
}
if all[0].Status != store.RoutineDismissed {
t.Errorf("status = %s, want dismissed", all[0].Status)
}
}
// proposalRule — the rule name announceProposal uses for the seeded pair.
const proposalRule = "proposal:refill cat_water"
// TestTickProposalSilentByDefault — detection is always on, announcing is not.
// With no pattern_proposals block the tick still records the proposal, and says
// nothing about it: Maven is not autonomous, so a behaviour that speaks without
// being asked stays off until it is configured.
func TestTickProposalSilentByDefault(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents)
markPresent(t, st, ctx, now)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
tl.tick(ctx, now)
if n := countSends(sink, proposalRule); n != 0 {
t.Fatalf("announced %d proposals with no config, want 0", n)
}
rows, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineProposed)
if err != nil {
t.Fatalf("list proposed: %v", err)
}
if len(rows) != 1 {
t.Fatalf("proposed routines = %d, want 1 (silent, but recorded)", len(rows))
}
}
// TestTickAnnouncesProposalWhenConfigured — with notify on, the proposal goes
// out once through the ordinary delivery path, worded by the detector itself.
// Later ticks stay quiet because the pair is already proposed: one pattern is
// one announcement, ever.
func TestTickAnnouncesProposalWhenConfigured(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents)
markPresent(t, st, ctx, now)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
tl.proposalCfg = &config.PatternProposalConfig{Notify: true}
tl.tick(ctx, now)
var got *delivery.Sendable
for i := range sink.sends {
if sink.sends[i].RuleName == proposalRule {
got = &sink.sends[i]
}
}
if got == nil {
t.Fatalf("proposal was not announced; sends=%+v", sink.sends)
}
if !strings.Contains(got.Body, "напоминать?") {
t.Errorf("body = %q, want the detector's own question", got.Body)
}
if got.Channel != delivery.ChannelVoice {
t.Errorf("channel = %v, want voice (sev1, present)", got.Channel)
}
// A month of further ticks: the pair already has a row, so there is
// nothing new to detect and nothing more to say.
sink.sends = nil
later := now.Add(40 * 24 * time.Hour)
markPresent(t, st, ctx, later)
tl.tick(ctx, later)
if n := countSends(sink, proposalRule); n != 0 {
t.Fatalf("re-announced an existing proposal %d times, want 0", n)
}
}
// TestTickProposalRespectsGate — a proposal is the least urgent thing Maven can
// say, so it is sev1 and the restraint gate suppresses it. Away presence means
// it is not announced at all: it is not held, not retried, it just lives on
// /routines. The proposal row is still written — noticing is never gated.
func TestTickProposalRespectsGate(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents)
// no presence probes ⇒ away ⇒ care-class gate blocks.
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
tl.proposalCfg = &config.PatternProposalConfig{Notify: true}
tl.tick(ctx, now)
if n := countSends(sink, proposalRule); n != 0 {
t.Fatalf("away: announced %d proposals, want 0", n)
}
if !tl.lastProposalAt.IsZero() {
t.Error("cooldown clock advanced on a suppressed announcement")
}
rows, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineProposed)
if err != nil {
t.Fatalf("list proposed: %v", err)
}
if len(rows) != 1 {
t.Fatalf("proposed routines = %d, want 1 (detection is never gated)", len(rows))
}
}
// TestTickProposalCooldownSpacesAnnouncements — two patterns detected on the
// same tick must not become two interruptions. The second one waits for the
// cooldown, and is on /routines meanwhile.
func TestTickProposalCooldownSpacesAnnouncements(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents)
for i := 0; i < pattern.MinEvents; i++ {
ts := now.Add(time.Duration(i) * 3 * 24 * time.Hour)
factID, err := st.WriteFact(ctx, ts, store.KindSelf, "litter_box", "clean", "test", 1.0, sql.NullInt64{})
if err != nil {
t.Fatalf("write fact: %v", err)
}
if _, err := st.CreateEvent(ctx, factID, "clean", "litter_box", ts); err != nil {
t.Fatalf("create event: %v", err)
}
}
markPresent(t, st, ctx, now)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
tl.proposalCfg = &config.PatternProposalConfig{Notify: true, Cooldown: config.Duration(24 * time.Hour)}
tl.tick(ctx, now)
announced := 0
for _, s := range sink.sends {
if strings.HasPrefix(s.RuleName, "proposal:") {
announced++
}
}
if announced != 1 {
t.Fatalf("announced %d proposals on one tick, want exactly 1", announced)
}
rows, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineProposed)
if err != nil {
t.Fatalf("list proposed: %v", err)
}
if len(rows) != 2 {
t.Fatalf("proposed routines = %d, want 2 (both recorded, one announced)", len(rows))
}
// Still inside the cooldown: silence, even though a proposal is pending.
sink.sends = nil
soon := now.Add(time.Hour)
markPresent(t, st, ctx, soon)
tl.tick(ctx, soon)
for _, s := range sink.sends {
if strings.HasPrefix(s.RuleName, "proposal:") {
t.Fatalf("announced %q inside the cooldown", s.RuleName)
}
}
}
+144
View File
@@ -0,0 +1,144 @@
// Quiet-mode toggle recognition — the pre-route keyword check that lets
// "тихий режим" flip the daemon-wide quiet_hours config without going through
// the router. Moved out of voice.go unchanged (Vikunja #321); the tests live in
// quiet_toggle_test.go.
package main
import (
"context"
"log"
"strings"
"unicode"
"github.com/kami/maven/internal/ipc"
)
// resolveQuietToggle — pre-route keyword check. Returns (reply, true) when
// the utterance is a quiet-on/off command; ("", false) otherwise. Called from
// runTurn BEFORE the router so a classifier miscue can't drop it — which means
// both the voice path and the text path (mavweb /api/chat, telegram) reach it,
// so a false positive here is a network-reachable way to flip a daemon-wide
// setting. See classifyQuietToggle for the matching rule.
func (h *reactiveHandler) resolveQuietToggle(ctx context.Context, text string) (string, bool) {
on, off := classifyQuietToggle(text)
if !on && !off {
return "", false
}
val := "false"
reply := "тихий режим выключен."
if on {
val = "true"
reply = "тихий режим включён. буду реже напоминать."
}
if _, err := h.api.WriteFact(ctx, ipc.WriteFactReq{
Ts: h.now(),
Kind: "config",
Key: "quiet_hours",
Value: val,
Source: "tap:voice",
Confidence: 1.0,
}); err != nil {
log.Printf("voice: write quiet_hours: %v", err)
return "не получилось переключить тихий режим.", true
}
return reply, true
}
// quietInflections — the inflectional endings a stem may carry and still be
// the same word. Adjective/adverb/noun/verb endings, all ≤3 letters. This is
// what separates "тихий"/"тихом"/"тихо" (stem "тих" + a real ending) from
// "тихонько"/"потихоньку", which are different words: "онько" is not an
// ending, and "потихоньку" doesn't start with the stem at all.
var quietInflections = []string{
"", "а", "е", "и", "й", "о", "у", "ы", "ю", "я",
"ая", "ее", "ей", "ем", "ие", "ий", "им", "их", "ия", "ию", "ое", "ой", "ом", "ую", "ые", "ый", "ым", "ых", "ья",
"ами", "ого", "ому", "ыми", "ать", "ить", "ять",
}
// quietStem reports whether tok is the given stem carrying at most one
// inflectional ending. Word boundaries come from tokenisation (see
// quietTokens), not from a regexp — Go's \b is ASCII-oriented and treats every
// Cyrillic letter as a non-word character, so `\bтих\b` would happily match
// inside "тихонько". Comparing whole tokens sidesteps that entirely.
func quietStem(tok, stem string) bool {
if !strings.HasPrefix(tok, stem) {
return false
}
suffix := tok[len(stem):]
for _, e := range quietInflections {
if suffix == e {
return true
}
}
return false
}
// quietTokens splits an utterance into lowercase word tokens, dropping
// punctuation and spacing. Unicode-aware, so Cyrillic words tokenise the same
// way ASCII ones do.
func quietTokens(text string) []string {
return strings.FieldsFunc(strings.ToLower(strings.TrimSpace(text)), func(r rune) bool {
return !unicode.IsLetter(r) && !unicode.IsDigit(r)
})
}
// quietPhrase matches a pattern (a sequence of stems) against the token list.
// Multi-word patterns match any contiguous run of tokens — "включи тихий
// режим" carries "тихий режим". Single-word patterns match ONLY when they are
// the whole utterance: bare "тихо" is a command, but "в комнате тихо" is a
// remark about the room and must not flip a daemon-wide setting.
func quietPhrase(tokens, pattern []string) bool {
if len(pattern) == 0 || len(tokens) < len(pattern) {
return false
}
if len(pattern) == 1 {
return len(tokens) == 1 && quietStem(tokens[0], pattern[0])
}
for i := 0; i+len(pattern) <= len(tokens); i++ {
hit := true
for j, stem := range pattern {
if !quietStem(tokens[i+j], stem) {
hit = false
break
}
}
if hit {
return true
}
}
return false
}
// quietOffPhrases / quietOnPhrases — the toggle vocabulary, as stem sequences.
var (
quietOffPhrases = [][]string{
{"quiet", "off"}, {"quiet", "end"},
{"громк", "режим"}, {"шумн", "режим"},
{"отмен", "тих"}, {"выключ", "тих"}, {"не", "тих"},
}
quietOnPhrases = [][]string{
{"quiet", "on"}, {"quiet", "mode"},
{"тих", "режим"}, {"не", "шум"}, {"не", "беспоко"},
{"тих"},
}
)
// classifyQuietToggle reads an utterance as a quiet-mode command. OFF is
// resolved before ON for the same reason classifyConfirm checks negatives
// first: the OFF phrases are built out of the ON words ("выключи тихий"
// contains "тихий"), so scanning ON first would shadow them and "выключи
// тихий режим" would turn quiet mode on. Negation wins.
func classifyQuietToggle(text string) (on, off bool) {
tokens := quietTokens(text)
for _, p := range quietOffPhrases {
if quietPhrase(tokens, p) {
return false, true
}
}
for _, p := range quietOnPhrases {
if quietPhrase(tokens, p) {
return true, false
}
}
return false, false
}
+114
View File
@@ -0,0 +1,114 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
)
// quietFakeAPI records the WriteFact the toggle performs.
type quietFakeAPI struct {
ipc.UnimplementedCoreAPI
got ipc.WriteFactReq
call int
}
func (a *quietFakeAPI) WriteFact(_ context.Context, req ipc.WriteFactReq) (int64, error) {
a.got, a.call = req, a.call+1
return 1, nil
}
// quietVerdict — what a phrase should do to the setting.
type quietVerdict int
const (
quietNone quietVerdict = iota
quietOn
quietOff
)
func TestResolveQuietToggle(t *testing.T) {
cases := []struct {
text string
want quietVerdict
}{
// ON vocabulary.
{"quiet on", quietOn},
{"quiet mode", quietOn},
{"тихий режим", quietOn},
{"тихий", quietOn},
{"не шуми", quietOn},
{"не беспокоить", quietOn},
{"тихо", quietOn},
// ON, inflected / embedded in a sentence.
{"включи тихий режим", quietOn},
{"побудь в тихом режиме", quietOn},
{"Тихий Режим!", quietOn},
{"тихая", quietOn},
// OFF vocabulary — all seven, incl. the three that used to say ON.
{"quiet off", quietOff},
{"quiet end", quietOff},
{"громкий режим", quietOff},
{"шумный режим", quietOff},
{"отмени тихий", quietOff},
{"выключи тихий", quietOff},
{"не тихо", quietOff},
// OFF wins over the ON words it contains.
{"выключи тихий режим", quietOff},
{"отмени тихий режим пожалуйста", quietOff},
{"верни громкий режим", quietOff},
// False positives: "тихо"/"тихий" as ordinary Russian.
{"очень тихий сегодня день", quietNone},
{"в комнате тихо", quietNone},
{"тихонько напомни", quietNone},
{"потихоньку", quietNone},
{"тихонько", quietNone},
{"он говорил тихим голосом весь вечер", quietNone},
// Unrelated.
{"напомни завтра позвонить маме", quietNone},
{"какая погода", quietNone},
{"", quietNone},
}
for _, tc := range cases {
t.Run(tc.text, func(t *testing.T) {
api := &quietFakeAPI{}
h := &reactiveHandler{api: api, now: func() time.Time { return time.Unix(0, 0).UTC() }}
reply, handled := h.resolveQuietToggle(context.Background(), tc.text)
if tc.want == quietNone {
if handled || reply != "" {
t.Fatalf("%q: got (%q, %v), want no match", tc.text, reply, handled)
}
if api.call != 0 {
t.Fatalf("%q: wrote a fact on a non-match", tc.text)
}
return
}
if !handled {
t.Fatalf("%q: not handled, want %v", tc.text, tc.want)
}
wantReply, wantVal := "тихий режим выключен.", "false"
if tc.want == quietOn {
wantReply, wantVal = "тихий режим включён. буду реже напоминать.", "true"
}
if reply != wantReply {
t.Errorf("%q: reply = %q, want %q", tc.text, reply, wantReply)
}
if api.call != 1 {
t.Fatalf("%q: WriteFact called %d times, want 1", tc.text, api.call)
}
if api.got.Kind != "config" || api.got.Key != "quiet_hours" || api.got.Source != "tap:voice" || api.got.Confidence != 1.0 {
t.Errorf("%q: request shape = %+v", tc.text, api.got)
}
if api.got.Value != wantVal {
t.Errorf("%q: value = %q, want %q", tc.text, api.got.Value, wantVal)
}
})
}
}
+9 -4
View File
@@ -7,6 +7,7 @@ import (
"time" "time"
"github.com/kami/maven/internal/llm" "github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/router" "github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice" "github.com/kami/maven/internal/voice"
) )
@@ -23,13 +24,17 @@ type completer interface {
type llmReplier struct { type llmReplier struct {
c completer c completer
stub *voice.StubReplier stub *voice.StubReplier
// block renders the shared context block per turn (who he is, the time).
// nil ⇒ the prompt stands alone.
block func() string
} }
func newLLMReplier(c completer) *llmReplier { func newLLMReplier(c completer, block func() string) *llmReplier {
return &llmReplier{c: c, stub: voice.NewStubReplier()} return &llmReplier{c: c, stub: voice.NewStubReplier(), block: block}
} }
const replySystem = `Ты — Maven, домашняя ассистентка (о себе — в женском роде). Подтверди действие РОВНО ОДНИМ коротким предложением (≤120 символов), тепло и по-русски. Не задавай вопросов, не повторяй слова, не добавляй ничего после точки. Отвечай ТОЛЬКО одним объектом JSON с полями "response" (текст) и "mood" (ровно одно из: neutral, happy, thinking, tired, confused). const replySystem = `Ты — Maven, домашняя ассистентка (о себе — в женском роде). Владелец — мужчина, говоришь с ним на "ты", в единственном числе; никогда не "вы"/"ваш" и не "он"/"его". Подтверди действие РОВНО ОДНИМ коротким предложением (≤120 символов), по-русски, спокойно и без официальных формулировок. Не задавай вопросов, не повторяй слова, не добавляй ничего после точки. Отвечай ТОЛЬКО одним объектом JSON с полями "response" (текст) и "mood" (ровно одно из: neutral, happy, thinking, tired, confused).
Пример: {"response": "Записала, что ты выпил стакан воды.", "mood": "neutral"} Пример: {"response": "Записала, что ты выпил стакан воды.", "mood": "neutral"}
Никогда не пиши "..." в поле response.` Никогда не пиши "..." в поле response.`
@@ -40,7 +45,7 @@ func (r *llmReplier) Reply(d router.Decision) string {
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second) ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel() defer cancel()
out, err := r.c.Complete(ctx, llm.Req{ out, err := r.c.Complete(ctx, llm.Req{
System: replySystem, System: persona.Prepend(r.block, replySystem),
User: replyContext(d), User: replyContext(d),
MaxTokens: 512, MaxTokens: 512,
RepeatPenalty: 1.3, RepeatPenalty: 1.3,
+5 -5
View File
@@ -17,7 +17,7 @@ type mockCompleter struct {
func (m mockCompleter) Complete(_ context.Context, _ llm.Req) (string, error) { return m.out, m.err } func (m mockCompleter) Complete(_ context.Context, _ llm.Req) (string, error) { return m.out, m.err }
func TestLLMReplierReturnsLLMReply(t *testing.T) { func TestLLMReplierReturnsLLMReply(t *testing.T) {
r := newLLMReplier(mockCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`}) r := newLLMReplier(mockCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`}, nil)
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}}) got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
if got != "записала, кофе закончился" { if got != "записала, кофе закончился" {
t.Errorf("got %q, want %q", got, "записала, кофе закончился") t.Errorf("got %q, want %q", got, "записала, кофе закончился")
@@ -25,7 +25,7 @@ func TestLLMReplierReturnsLLMReply(t *testing.T) {
} }
func TestLLMReplierFallsBackToPlainText(t *testing.T) { func TestLLMReplierFallsBackToPlainText(t *testing.T) {
r := newLLMReplier(mockCompleter{out: "записала, кофе закончился"}) r := newLLMReplier(mockCompleter{out: "записала, кофе закончился"}, nil)
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}}) got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
if got != "записала, кофе закончился" { if got != "записала, кофе закончился" {
t.Errorf("got %q, want %q", got, "записала, кофе закончился") t.Errorf("got %q, want %q", got, "записала, кофе закончился")
@@ -33,7 +33,7 @@ func TestLLMReplierFallsBackToPlainText(t *testing.T) {
} }
func TestLLMReplierFallsBackToStubOnError(t *testing.T) { func TestLLMReplierFallsBackToStubOnError(t *testing.T) {
r := newLLMReplier(mockCompleter{err: errTestLLMDown}) r := newLLMReplier(mockCompleter{err: errTestLLMDown}, nil)
noteDec := router.Decision{Intent: router.IntentNote} noteDec := router.Decision{Intent: router.IntentNote}
got := r.Reply(noteDec) got := r.Reply(noteDec)
want := voice.NewStubReplier().Reply(noteDec) want := voice.NewStubReplier().Reply(noteDec)
@@ -43,7 +43,7 @@ func TestLLMReplierFallsBackToStubOnError(t *testing.T) {
} }
func TestLLMReplierFallsBackToStubOnEmpty(t *testing.T) { func TestLLMReplierFallsBackToStubOnEmpty(t *testing.T) {
r := newLLMReplier(mockCompleter{out: ""}) r := newLLMReplier(mockCompleter{out: ""}, nil)
noteDec := router.Decision{Intent: router.IntentNote} noteDec := router.Decision{Intent: router.IntentNote}
got := r.Reply(noteDec) got := r.Reply(noteDec)
want := voice.NewStubReplier().Reply(noteDec) want := voice.NewStubReplier().Reply(noteDec)
@@ -53,7 +53,7 @@ func TestLLMReplierFallsBackToStubOnEmpty(t *testing.T) {
} }
func TestLLMReplierClarifyUsesStub(t *testing.T) { func TestLLMReplierClarifyUsesStub(t *testing.T) {
r := newLLMReplier(mockCompleter{out: "я всё поняла"}) r := newLLMReplier(mockCompleter{out: "я всё поняла"}, nil)
clarifyDec := router.Decision{Clarify: true} clarifyDec := router.Decision{Clarify: true}
got := r.Reply(clarifyDec) got := r.Reply(clarifyDec)
want := voice.NewStubReplier().Reply(clarifyDec) want := voice.NewStubReplier().Reply(clarifyDec)
+180
View File
@@ -0,0 +1,180 @@
// Package main — ruwords.go holds Russian language + calendar/time formatting
// helpers used by the voice reply paths (replySystem, the reminder/routine
// phrasing, etc). Pure functions, no receivers: weekday/month name tables,
// plural agreement, clock/date rendering, and the "do I actually know this
// place/day" guards that pick an honest reply over a confidently wrong one.
// Extend this file rather than voice.go for anything in that shape.
package main
import (
"fmt"
"strconv"
"strings"
"time"
)
var ruWeekdays = []string{
"воскресенье", "понедельник", "вторник", "среда",
"четверг", "пятница", "суббота",
}
var ruMonths = []string{
"января", "февраля", "марта", "апреля", "мая", "июня",
"июля", "августа", "сентября", "октября", "ноября", "декабря",
}
// onlyLocalTimeReply — the honest answer when the user asks the time somewhere
// other than here. She only keeps one clock, and saying so is better than
// naming the wrong city's time.
//
// There used to be a city→time-zone table here. It was removed on purpose: the
// user only ever asks for local time, so the table was a second list of cities
// to keep in step with the weather one for no gain.
const onlyLocalTimeReply = "я знаю только местное время, про другие города пока не скажу."
// notPlaceAfterV — words that follow "в" without naming a place, so
// mentionsUnknownPlace does not mistake them for a city.
var notPlaceAfterV = map[string]bool{
"данный": true, "данную": true, "этот": true, "эту": true,
"котором": true, "какое": true, "какой": true, "который": true,
"общем": true, "точности": true, "курсе": true, "сутках": true,
"часах": true, "минутах": true, "секундах": true, "неделе": true,
}
// mentionsUnknownPlace reports whether the question has a "в <слово>" phrase
// that looks like a place we do not know ("который час в киеве"). Used only to
// pick the honest "local time only" reply instead of answering local time as
// if it were the city's.
func mentionsUnknownPlace(u string) bool {
toks := strings.Fields(u)
for i := 0; i+1 < len(toks); i++ {
if toks[i] != "в" && toks[i] != "во" {
continue
}
next := strings.Trim(toks[i+1], ".,?!")
if next == "" || notPlaceAfterV[next] {
continue
}
// A number after "в" is a clock ("в 5 часов"), not a place.
if _, err := strconv.Atoi(strings.SplitN(next, ":", 2)[0]); err == nil {
continue
}
return true
}
return false
}
// onlyNearDaysReply — she can work out today, tomorrow, the day after and
// yesterday, and nothing further. Said out loud instead of answering today's
// date for a day she did not understand.
const onlyNearDaysReply = "я считаю только сегодня, завтра, послезавтра и вчера — про другие дни пока не скажу."
// dayWords — day references the calendar parser cannot resolve. A weekday name
// or a "через …" phrase means he asked about a specific other day.
var dayWords = []string{
"понедельник", "вторник", "сред", "четверг", "пятниц", "суббот", "воскресен",
"через", "monday", "tuesday", "wednesday", "thursday", "friday", "saturday", "sunday",
}
// mentionsUnknownDay reports whether the question names a day the calendar
// parser could not resolve. Mirror of mentionsUnknownPlace: it exists only to
// pick an honest reply over a confidently wrong one.
//
// Only called after ParseCalendarDate has already failed, so "завтра" and the
// other words it does know never reach here.
func mentionsUnknownDay(u string) bool {
for _, w := range dayWords {
if strings.Contains(u, w) {
return true
}
}
return false
}
// ruClock renders the clock part of the time reply: "15 часов 4 минуты".
func ruClock(t time.Time) string {
h, m := t.Hour(), t.Minute()
hourWord := ruPlural(h, "час", "часа", "часов")
if m == 0 {
return fmt.Sprintf("%d %s ровно", h, hourWord)
}
return fmt.Sprintf("%d %s %d %s", h, hourWord, m, ruPlural(m, "минута", "минуты", "минут"))
}
// dayPrefix names the day relative to now ("завтра", "вчера", …) so the date
// reply opens the way a person would say it.
func dayPrefix(now, day time.Time) string {
base := time.Date(now.Year(), now.Month(), now.Day(), 0, 0, 0, 0, now.Location())
switch int(day.Sub(base).Hours() / 24) {
case -1:
return "вчера"
case 0:
return "сегодня"
case 1:
return "завтра"
case 2:
return "послезавтра"
}
return "это"
}
func ruPlural(n int, one, two, many string) string {
n = n % 100
if n > 10 && n < 20 {
return many
}
n = n % 10
switch n {
case 1:
return one
case 2, 3, 4:
return two
default:
return many
}
}
// hasDurationWords checks whether u is asking about elapsed/remaining time
// rather than the current clock — guards replySystem from replying "сейчас
// X часов" to "сколько времени прошло". Mirrors the stage0.go build filter.
func hasDurationWords(u string) bool {
s := strings.ToLower(strings.TrimSpace(u))
// First-word duration markers (same keywords as timeQueryBuild in stage0).
first := strings.Fields(s)
if len(first) > 0 {
switch first[0] {
case "прошло", "осталось", "пройдет", "минуло", "проходит":
return true
}
}
// Broader duration keywords appearing anywhere in the utterance.
if strings.Contains(s, "прошло") || strings.Contains(s, "осталось") {
return true
}
if strings.Contains(s, " до ") {
return true
}
return false
}
// formatTime returns a human-readable Russian time string for a fact timestamp.
// Used by the query handler when answering "когда я это сделал?"-style questions.
func formatTime(t time.Time) string {
now := time.Now()
if t.After(now.Add(-2*time.Minute)) && t.Before(now.Add(2*time.Minute)) {
return "только что"
}
diff := now.Sub(t)
switch {
case diff < 10*time.Minute:
return "несколько минут назад"
case diff < 60*time.Minute:
return fmt.Sprintf("%d минут назад", int(diff.Minutes()))
case diff < 2*time.Hour:
return "час назад"
case diff < 24*time.Hour:
return fmt.Sprintf("%d часа назад", int(diff.Hours()))
default:
return t.Format("2 января 15:04")
}
}
+85
View File
@@ -0,0 +1,85 @@
// Package main — strutil.go holds small, receiver-free string utilities used
// across the voice reply paths: trimming a wake token, pulling out the first
// word or first line, and a minimal JSON string encoder for the one payload
// shape that needs it. Extend this file rather than voice.go for anything in
// that shape.
package main
import (
"fmt"
"strings"
"github.com/kami/maven/internal/router"
)
// stripWake removes a leading wake token (any script the STT phonetically
// transcribes "Maven" as) so the verb is the first word.
func stripWake(u string) string {
stripped, had := router.StripWakeToken(u)
if !had {
return strings.TrimSpace(u)
}
return stripped
}
// firstWord returns the first whitespace-delimited token (lowercased) — the
// proposed tool's name.
func firstWord(s string) string {
f := strings.Fields(s)
if len(f) == 0 {
return ""
}
return strings.ToLower(f[0])
}
// firstLine — the first non-empty line of a tool's output, for a short spoken
// reply (the full output goes to the log, not the TTS). Trimmed to keep the
// utterance sane if a command dumps a wall of text.
func firstLine(s string) string {
for _, line := range strings.Split(s, "\n") {
line = strings.TrimSpace(line)
if line != "" {
if len(line) > 200 {
line = line[:200]
}
return line
}
}
return ""
}
// jsonString — a one-line JSON string encoder without dragging encoding/json
// into the top of this file. Used to wrap a reminder payload's text field;
// the router's reminder Slots are already absolute (DateTimeParser resolved
// relative→absolute), the payload shape is conventional {"text":...}.
func jsonString(s string) string {
// minimal JSON string escape — quotes + backslash + control chars.
// adequate for the reminder payload's text field; not a general JSON
// encoder. The chroma / RAG modules (when they land) use a real json
// encoder for richer payloads. Keep it inline here so the import
// direction stays narrow.
var b []byte
b = append(b, '"')
for _, r := range s {
switch r {
case '"':
b = append(b, '\\', '"')
case '\\':
b = append(b, '\\', '\\')
case '\n':
b = append(b, '\\', 'n')
case '\r':
b = append(b, '\\', 'r')
case '\t':
b = append(b, '\\', 't')
default:
if r < 0x20 {
b = append(b, []byte(fmt.Sprintf("\\u%04x", r))...)
} else {
b = append(b, []byte(string(r))...)
}
}
}
b = append(b, '"')
return string(b)
}
+338
View File
@@ -24,6 +24,7 @@ import (
"github.com/kami/maven/internal/ipc" "github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/loop" "github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/morning" "github.com/kami/maven/internal/morning"
"github.com/kami/maven/internal/pattern"
"github.com/kami/maven/internal/phraser" "github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/routine" "github.com/kami/maven/internal/routine"
"github.com/kami/maven/internal/store" "github.com/kami/maven/internal/store"
@@ -67,6 +68,14 @@ type tickLoop struct {
morningRoutines []morning.Routine morningRoutines []morning.Routine
morningLast map[string]time.Time morningLast map[string]time.Time
// proposalCfg — announcement policy for routines the tick inferred itself.
// nil ⇒ detect silently, never announce (the default). lastProposalAt is
// the cooldown clock, in-memory on purpose: a restart is allowed to permit
// one more announcement, and a restart-per-day loop is a bigger problem
// than a duplicate proposal notice.
proposalCfg *config.PatternProposalConfig
lastProposalAt time.Time
// digestQ — in-memory queue of eligible nudges waiting for batch flush. // digestQ — in-memory queue of eligible nudges waiting for batch flush.
// populated when digestCfg != nil && digestCfg.Enabled. // populated when digestCfg != nil && digestCfg.Enabled.
digestQ []QueuedNudge digestQ []QueuedNudge
@@ -92,6 +101,7 @@ func newTickLoop(
digestCfg *config.DigestConfig, digestCfg *config.DigestConfig,
routines []routine.Routine, routines []routine.Routine,
morningRoutines []morning.Routine, morningRoutines []morning.Routine,
proposalCfg *config.PatternProposalConfig,
) *tickLoop { ) *tickLoop {
return &tickLoop{ return &tickLoop{
store: st, store: st,
@@ -108,6 +118,7 @@ func newTickLoop(
routineLast: make(map[string]time.Time), routineLast: make(map[string]time.Time),
morningRoutines: morningRoutines, morningRoutines: morningRoutines,
morningLast: make(map[string]time.Time), morningLast: make(map[string]time.Time),
proposalCfg: proposalCfg,
lastPhrase: make(map[string]delivery.PhrasedNudge), lastPhrase: make(map[string]delivery.PhrasedNudge),
} }
} }
@@ -180,6 +191,17 @@ func (t *tickLoop) tick(ctx context.Context, now time.Time) {
// (with dedup) avoids re-queueing the same rule after a flush. // (with dedup) avoids re-queueing the same rule after a flush.
t.maybeFlush(ctx, now, state) t.maybeFlush(ctx, now, state)
// gate-suppressed digest (Vikunja #281): rules the restraint gate held
// back this tick (quiet hours / away / calendar-busy), not because they
// weren't due, but because it wasn't the moment. Some of those are worth
// resurfacing later instead of just being lost — loop.DigestEligible
// draws that line. This is a SEPARATE mechanism from the in-memory
// digestQ above: that one batches candidates the gate already ALLOWED to
// fire; this one durably holds candidates the gate BLOCKED.
t.enqueueSuppressedDigest(ctx, trace, state, now)
t.expireStaleDigest(ctx, now)
t.maybeDrainDigest(ctx, state, now)
// routines: operator-declared scheduled behaviors. fire the ones whose cron // routines: operator-declared scheduled behaviors. fire the ones whose cron
// crossed since last fire, delivered through the normal routing (voice when // crossed since last fire, delivered through the normal routing (voice when
// present, away channels otherwise). bodies are literal operator text — not // present, away channels otherwise). bodies are literal operator text — not
@@ -195,6 +217,14 @@ func (t *tickLoop) tick(ctx context.Context, now time.Time) {
// nudge time. See internal/morning for the "why not four timers" rationale. // nudge time. See internal/morning for the "why not four timers" rationale.
t.fireMorningRoutines(ctx, now, state) t.fireMorningRoutines(ctx, now, state)
// pattern detection: scan every action+object pair with recorded events
// and propose a routine for any stable one not already decided (Vikunja
// #43). This used to only run as a side effect of the voice fact-write
// path, so a pattern already sitting in history went unnoticed until he
// happened to mention it again by voice. See patterns.go and
// detectPatterns below for how idempotence and dismissal are respected.
t.detectPatterns(ctx, now, state)
// reminders: gate-bypassing class. fired once, marked after a successful // reminders: gate-bypassing class. fired once, marked after a successful
// delivery. a failed send leaves the reminder pending — the next tick // delivery. a failed send leaves the reminder pending — the next tick
// re-gathers and re-attempts. // re-gathers and re-attempts.
@@ -342,6 +372,235 @@ func (t *tickLoop) flushDigest(ctx context.Context, now time.Time, state loop.St
t.digestQ = nil t.digestQ = nil
} }
// detectPatterns runs the pattern detector proactively over every
// action+object pair that has ever produced an event, independent of
// whichever fact write (or channel) last touched it (Vikunja #43). This is
// what makes pattern inference actually proactive: it fires on the daemon's
// own schedule reading accumulated history, not only as a side effect of a
// live voice turn.
//
// Idempotence and noise are handled by the store, not here — this function
// is safe to call every tick:
// - Same pattern, tick after tick: detectAndPropose's LookupProposedRoutine
// check plus proposed_routines' UNIQUE(action, object) constraint (with
// CreateProposedRoutine's ON CONFLICT DO NOTHING) mean a pair that
// already has a row — in ANY status — produces no second row and no log
// spam beyond the one line at genuine creation.
// - A DISMISSED proposal must never come back. DismissProposedRoutine flips
// status in place; the row is never deleted. So the same Lookup check
// that stops a duplicate "proposed" also stops a "dismissed" one from
// resurrecting — there is nothing tick-specific to get right here beyond
// calling the same shared path the voice route already used.
//
// By default this only creates a row for the /routines page to show: it does
// not notify, ring, or speak. Detection is not the same act as disturbing him
// about it, and Maven is "not a nag, not autonomous" (CLAUDE.md). Announcing
// is opt-in through the pattern_proposals config block — see announceProposal
// for the restraints that apply even then. A proposal only starts producing
// recurring nudges once he accepts it (fireAcceptedRoutines).
func (t *tickLoop) detectPatterns(ctx context.Context, now time.Time, state loop.State) {
pairs, err := t.store.DistinctEventPairs(ctx)
if err != nil {
log.Printf("tick: distinct event pairs: %v", err)
return
}
announced := false
for _, p := range pairs {
r, _, err := detectAndPropose(ctx, t.store, p.Action, p.Object, now)
if err != nil {
log.Printf("tick: detect pattern %s/%s: %v", p.Action, p.Object, err)
continue
}
if r == nil {
continue // no stable pattern, or already proposed/accepted/dismissed
}
log.Printf("tick: proposed routine: %s/%s every %.1f days", r.Action, r.Object, r.IntervalDays)
// One announcement per tick at most, whatever the scan turned up. The
// rest are on /routines; they are not lost, they are just not shouted.
if announced {
continue
}
announced = t.announceProposal(ctx, r, now, state)
}
}
// announceProposal offers a freshly inferred routine through the ordinary
// care-delivery path, if announcing is switched on at all. Returns true when
// something was actually sent.
//
// Everything here is restraint. The feature is off unless configured; when on
// it is sev1 (the lowest severity, so quiet hours, away presence and snooze
// all suppress it via loop.Gate exactly like a care nudge); it is spaced by
// proposalCfg.Cooldown across every pair, not per pair; and a suppressed or
// dropped announcement is NOT retried — the cooldown clock advances only on a
// real send, but the proposal row already exists, so the next tick will not
// re-detect it and nothing queues up behind it. A missed announcement means
// he reads it on /routines instead, which is the whole point of the page.
//
// The body is the detector's own literal Russian phrasing (pattern.PhraseRoutine
// — "ты заправляешь поилку раз в 7 дней — напоминать?"), not LLM-generated, so
// an inferred routine cannot arrive worded as something Maven never observed.
func (t *tickLoop) announceProposal(ctx context.Context, r *pattern.ProposedRoutine, now time.Time, state loop.State) bool {
if !t.proposalCfg.AnnounceProposals() {
return false
}
cooldown := time.Duration(t.proposalCfg.Cooldown)
if cooldown <= 0 {
cooldown = config.DefaultProposalCooldown
}
if !t.lastProposalAt.IsZero() && now.Sub(t.lastProposalAt) < cooldown {
return false
}
rule := loop.Rule{Name: "proposal:" + r.Action + " " + r.Object, Severity: loop.Sev1}
if !loop.Gate(state, rule) {
return false
}
body := pattern.PhraseRoutine(r)
pn := delivery.PhrasedNudge{
Candidate: loop.Candidate{Rule: rule, Severity: rule.Severity, State: state},
Body: body,
Summary: body,
}
sent, err := t.dispatcher.DispatchNudge(ctx, pn, now)
if err != nil {
log.Printf("tick: announce proposal %s/%s: %v", r.Action, r.Object, err)
return false
}
if len(sent) == 0 {
return false // routing dropped it — /routines still has it.
}
t.lastProposalAt = now
return true
}
// digestExpiry — how long a gate-suppressed care nudge stays worth
// resurfacing. 24h: these are daily-cadence rules (water/meal/break run on
// hour-scale cooldowns and re-derive from facts that reset every day), so a
// digest entry that outlives one full day is describing a day that's already
// over — "you skipped a break yesterday" said tomorrow evening is noise, not
// news. Bounding at one day also means a digest can never silently span a
// weekend of quiet hours into an unbounded backlog.
const digestExpiry = 24 * time.Hour
// maxDigestSpokenItems — the bundle read-out is capped so "batched, not
// dropped" cannot regress into "she dumps twelve things on me the moment I
// walk in" — a digest that nags in bulk is worse than the drops it replaced.
// Anything beyond the cap is still marked drained (it did get its moment;
// the cap limits WORDS, not whether it counted) and folded into a trailing
// count instead of being spoken in full.
const maxDigestSpokenItems = 3
// enqueueSuppressedDigest scans this tick's trace for care candidates the
// gate blocked for a genuine restraint reason and durably records the
// digest-eligible ones (loop.DigestEligible). Phrasing happens once, here,
// at enqueue time — not re-derived at drain time — the same way queueNudge
// phrases once and caches, so a rule suppressed for hours isn't re-prompting
// the LLM every tick it stays blocked (EnqueueDigestEntry's rule+body dedupe
// makes repeat calls here harmless, but skipping the phrase call entirely
// when a pending entry already exists avoids the LLM round-trip too).
func (t *tickLoop) enqueueSuppressedDigest(ctx context.Context, trace *loop.TickTrace, state loop.State, now time.Time) {
if trace == nil {
return
}
for _, tr := range trace.RuleTraces {
if !tr.PredicateResult || tr.GateResult {
continue // didn't want to fire, or wasn't suppressed
}
if !loop.DigestEligible(tr.Severity, tr.GateBlockedBy) {
continue
}
rule := loop.Rule{Name: tr.RuleName, Severity: tr.Severity}
cand := loop.Candidate{Rule: rule, Severity: tr.Severity, State: state}
pn, err := t.phraser.PhraseNudge(ctx, cand)
if err != nil {
log.Printf("tick: phrase digest candidate %s: %v", tr.RuleName, err)
continue
}
expires := now.Add(digestExpiry)
if _, deduped, err := t.store.EnqueueDigestEntry(ctx, tr.RuleName, int(tr.Severity), pn.Body, now, expires); err != nil {
log.Printf("tick: enqueue digest entry %s: %v", tr.RuleName, err)
} else if deduped {
// same suppressed nudge already pending — nothing new to say.
continue
}
}
}
// expireStaleDigest sweeps entries past their expiry once per tick — cheap
// bookkeeping, mirrors ReconcileStaleDeliveryAttempts's shape.
func (t *tickLoop) expireStaleDigest(ctx context.Context, now time.Time) {
n, err := t.store.ExpireStaleDigestEntries(ctx, now)
if err != nil {
log.Printf("tick: expire stale digest entries: %v", err)
return
}
if n > 0 {
log.Printf("tick: expired %d stale digest entr(y/ies) unspoken", n)
}
}
// maybeDrainDigest speaks the pending digest bundle once the gate's
// suppression reasons have actually cleared — quiet hours over, back from
// away, out of the meeting. Draining while still suppressed would just be a
// second way to nag through quiet hours; the bundle waits for the same "is
// it allowed right now" condition a live nudge already waits for.
func (t *tickLoop) maybeDrainDigest(ctx context.Context, state loop.State, now time.Time) {
if state.QuietHours || state.CalendarBusy || state.Presence == store.Away {
return
}
entries, err := t.store.PendingDigestEntries(ctx, now)
if err != nil {
log.Printf("tick: pending digest entries: %v", err)
return
}
if len(entries) == 0 {
return
}
spoken := entries
extra := 0
if len(spoken) > maxDigestSpokenItems {
spoken = entries[:maxDigestSpokenItems]
extra = len(entries) - maxDigestSpokenItems
}
var b strings.Builder
maxSev := 0
for i, e := range spoken {
if i > 0 {
b.WriteString(" · ")
}
b.WriteString(e.Body)
if e.Severity > maxSev {
maxSev = e.Severity
}
}
if extra > 0 {
fmt.Fprintf(&b, " · и ещё %d", extra)
}
body := b.String()
summary := fmt.Sprintf("%d отложенных уведомлений", len(entries))
cand := loop.Candidate{
Rule: loop.Rule{Name: "digest", Severity: loop.Severity(maxSev)},
Severity: loop.Severity(maxSev),
State: state,
}
pn := delivery.PhrasedNudge{Candidate: cand, Body: body, Summary: summary}
t.cachePhrase(pn)
if _, err := t.dispatcher.DispatchNudge(ctx, pn, now); err != nil {
log.Printf("tick: dispatch digest bundle: %v", err)
return // leave entries pending; retried next tick
}
ids := make([]int64, len(entries))
for i, e := range entries {
ids[i] = e.ID
}
if err := t.store.DrainDigestEntries(ctx, ids, now); err != nil {
log.Printf("tick: drain digest entries: %v", err)
}
}
// routinesFromConfig maps the config's routine blocks to the engine type. // routinesFromConfig maps the config's routine blocks to the engine type.
// Validation (cron parses, name/body present, severity defaulted) already ran // Validation (cron parses, name/body present, severity defaulted) already ran
// in config.Load, so this is a pure field copy. // in config.Load, so this is a pure field copy.
@@ -530,6 +789,77 @@ func (t *tickLoop) morningStatus(ctx context.Context, now time.Time) []ipc.Morni
return out return out
} }
// dayPlan is the read-only "what does today hold" query (Vikunja #128). It is
// the impure half of morning.BuildPlan: it reads the calendar events, the
// pending reminders and the checklist facts, and the pure builder orders them.
//
// It never dispatches. Asking for the plan is a query like any other; the only
// unprompted delivery in maven stays with the morning nudge and the
// dispatcher's policy.
func (t *tickLoop) dayPlan(ctx context.Context, now time.Time) ipc.DayPlan {
y, m, d := now.Date()
dayStart := time.Date(y, m, d, 0, 0, 0, 0, now.Location())
dayEnd := dayStart.AddDate(0, 0, 1)
var events []morning.PlanEntry
facts, err := t.store.CalendarEvents(ctx, dayStart, dayEnd)
if err != nil {
log.Printf("tick: day plan: calendar events: %v", err)
}
for _, f := range facts {
events = append(events, morning.PlanEntry{
At: f.Ts,
Text: f.Value,
Kind: morning.PlanEvent,
// Provenance below a calendar read (an ambient relay, #126) is
// hedged rather than recited as fact.
Uncertain: f.Confidence < 1.0,
})
}
var reminders []morning.PlanEntry
rems, err := t.store.ListReminders(ctx, dayPlanMaxReminders)
if err != nil {
log.Printf("tick: day plan: list reminders: %v", err)
}
for _, r := range rems {
if r.Status != "pending" {
continue
}
fire := r.NextFireTs
if fire.IsZero() {
fire = r.FireTs
}
reminders = append(reminders, morning.PlanEntry{
At: fire,
Text: strings.TrimSpace(r.Payload),
Kind: morning.PlanReminder,
})
}
var checklistFacts map[string]store.Fact
if len(t.morningRoutines) > 0 {
checklistFacts = t.gatherMorningFacts(ctx)
}
plan := morning.BuildPlan(t.morningRoutines, checklistFacts, events, reminders, now)
out := ipc.DayPlan{Date: plan.Date, Spoken: plan.FormatRU()}
out.Items = make([]ipc.DayPlanItem, len(plan.Items))
for i, it := range plan.Items {
out.Items[i] = ipc.DayPlanItem{
At: it.At,
Text: it.Text,
Kind: string(it.Kind),
Uncertain: it.Uncertain,
}
}
return out
}
// dayPlanMaxReminders bounds the reminder scan. The plan covers one day; a
// pending queue longer than this is a bug elsewhere, not a plan to recite.
const dayPlanMaxReminders = 500
// tune — the feedback auto-tuner's impure step. runs on a slow cadence // tune — the feedback auto-tuner's impure step. runs on a slow cadence
// (autotuneInterval, see run) so it doesn't write a fact every tick. for each // (autotuneInterval, see run) so it doesn't write a fact every tick. for each
// rule: // rule:
@@ -609,6 +939,7 @@ type daemonAPI struct {
ipc.CoreAPI ipc.CoreAPI
getTrace func() *loop.TickTrace getTrace func() *loop.TickTrace
getMorningStatus func(ctx context.Context) []ipc.MorningRoutineStatus getMorningStatus func(ctx context.Context) []ipc.MorningRoutineStatus
getDayPlan func(ctx context.Context) ipc.DayPlan
chatFn func(ctx context.Context, text string) string chatFn func(ctx context.Context, text string) string
} }
@@ -634,6 +965,13 @@ func (d *daemonAPI) MorningStatus(ctx context.Context) ([]ipc.MorningRoutineStat
return d.getMorningStatus(ctx), nil return d.getMorningStatus(ctx), nil
} }
func (d *daemonAPI) DayPlan(ctx context.Context) (ipc.DayPlan, error) {
if d.getDayPlan == nil {
return ipc.DayPlan{}, errors.New("mavend: day plan not available")
}
return d.getDayPlan(ctx), nil
}
func toIPCTickTrace(t loop.TickTrace) ipc.TickTrace { func toIPCTickTrace(t loop.TickTrace) ipc.TickTrace {
rules := make([]ipc.RuleTrace, len(t.RuleTraces)) rules := make([]ipc.RuleTrace, len(t.RuleTraces))
for i, r := range t.RuleTraces { for i, r := range t.RuleTraces {
+2 -2
View File
@@ -46,7 +46,7 @@ func newTestTickLoop(t *testing.T, st *store.Store, sink delivery.Sink, digestCf
Nudges: st, Nudges: st,
Reminders: st, Reminders: st,
}) })
return newTickLoop(st, g, d, phraser.NewStub(), rules, time.Second, 5*time.Minute, 0, digestCfg, nil, nil) return newTickLoop(st, g, d, phraser.NewStub(), rules, time.Second, 5*time.Minute, 0, digestCfg, nil, nil, nil)
} }
func TestTickFiresRoutineWhenScheduleCrosses(t *testing.T) { func TestTickFiresRoutineWhenScheduleCrosses(t *testing.T) {
@@ -63,7 +63,7 @@ func TestTickFiresRoutineWhenScheduleCrosses(t *testing.T) {
sink := &fakeSink{} sink := &fakeSink{}
d := delivery.NewDispatcher(delivery.Config{Voice: sink, Ntfy: sink, Telegram: sink, Nudges: st, Reminders: st}) d := delivery.NewDispatcher(delivery.Config{Voice: sink, Ntfy: sink, Telegram: sink, Nudges: st, Reminders: st})
rs := []routine.Routine{{Name: "morning", Cron: "0 12 * * *", Body: "полдень, время воды", Severity: 1}} rs := []routine.Routine{{Name: "morning", Cron: "0 12 * * *", Body: "полдень, время воды", Severity: 1}}
tl := newTickLoop(st, g, d, phraser.NewStub(), rules, time.Second, 5*time.Minute, 0, nil, rs, nil) tl := newTickLoop(st, g, d, phraser.NewStub(), rules, time.Second, 5*time.Minute, 0, nil, rs, nil, nil)
// first tick: seeds, does not fire the routine. // first tick: seeds, does not fire the routine.
tl.tick(ctx, now) tl.tick(ctx, now)
+44 -1625
View File
File diff suppressed because it is too large Load Diff
+434
View File
@@ -0,0 +1,434 @@
package main
import (
"bufio"
"context"
"fmt"
"log"
"os"
"path/filepath"
"strings"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/delivery/voicesink"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/stt"
"github.com/kami/maven/internal/tool"
"github.com/kami/maven/internal/tts"
"github.com/kami/maven/internal/voice"
"github.com/kami/maven/internal/weather"
"github.com/kami/maven/internal/worker"
)
// voiceWiring — everything the daemon needs to run the audio path. Held by
// cmd/mavend/main.go alongside the other wirings; closed on shutdown.
type voiceWiring struct {
server *voice.Server
sessions *voice.Sessions
voiceSink delivery.Sink
embedder router.Embedder
handler *reactiveHandler // the reactive handler for IPC Chat
// worker clients (set when configured as Remote): closed on shutdown so
// mavsttd / mavttsd don't keep a stale conn into a restarting daemon.
sttClient *worker.Client
ttsClient *worker.Client
}
// close releases the listener + worker conns. Safe to call on nil (when
// voice is not wired — wireVoice returns nil,nil).
func (w *voiceWiring) close() {
if w == nil {
return
}
if w.embedder != nil {
_ = w.embedder.Close()
}
if w.server != nil {
_ = w.server.Close()
}
if w.sttClient != nil {
_ = w.sttClient.Close()
}
if w.ttsClient != nil {
_ = w.ttsClient.Close()
}
}
// wireVoice builds the audio path from cfg + a CoreAPI + a router. Returns
// nil wiring + nil error when voice isn't enabled (the caller's voice sink
// stays nil; the dispatcher's ChannelVoice routing drops silently).
//
// When voice is enabled, MUST wire a voicesink into the dispatcher's Voice
// slot using w.sessions (the caller does that — see main.go).
func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, memStore memory.Store, dataStore *store.Store, eco *ecosystemWiring) (*voiceWiring, error) {
if cfg.Voice == nil || !cfg.Voice.Enabled {
return nil, nil
}
w := &voiceWiring{}
// ----- stt (Stub in-process OR Remote via worker socket) -----
var transcriber stt.Transcriber
if cfg.Voice.Stt != nil && cfg.Voice.Stt.Socket != "" {
c := worker.Dial(cfg.Voice.Stt.Socket)
w.sttClient = c
lang := cfg.Voice.Stt.Lang
if lang == "" {
lang = cfg.Voice.Lang
}
transcriber = stt.NewRemote(c, lang)
} else {
transcriber = stt.NewStub()
}
// ----- tts (Stub in-process OR Remote) -----
var synthesizer tts.Synthesizer
if cfg.Voice.Tts != nil && cfg.Voice.Tts.Socket != "" {
c := worker.Dial(cfg.Voice.Tts.Socket)
w.ttsClient = c
lang := cfg.Voice.Tts.Lang
if lang == "" {
lang = cfg.Voice.Lang
}
synthesizer = tts.NewRemote(c, lang, cfg.Voice.Tts.Voice)
} else {
synthesizer = tts.NewStub()
}
// ----- router: embedder (ONNX when configured, floor HashEmbedder otherwise) -----
var emb router.Embedder
if cfg.Voice.Embedder != nil {
onnx, err := router.NewONNXEmbedder(
cfg.Voice.Embedder.ModelPath,
cfg.Voice.Embedder.TokenizerPath,
cfg.Voice.Embedder.LibPath,
)
if err != nil {
w.close()
return nil, fmt.Errorf("embedder: %w", err)
}
log.Printf("voice: onnx embedder loaded (%d dim)", onnx.Dim())
emb = onnx
} else {
log.Printf("voice: embedder not configured, using HashEmbedder floor")
emb = router.NewHashEmbedder(1024)
}
w.embedder = emb
checkStoredEmbedder(dataStore, emb)
// ----- tool executor (the enabled act allowlist, store-backed) -----
// Config tools are the declarative bootstrap: seed them into the store as
// enabled (editing mavend.json IS the human enable act). Ad-hoc tools are
// enabled later through the authed mavweb surface. The executor + matcher
// both read the store live, so a newly-enabled tool is runnable without a
// daemon restart.
seedTools(coreAPI, cfg.Voice.Tools)
exec := tool.NewExecutor(coreAPI, time.Duration(cfg.Voice.ToolTimeout))
matcher := tool.NewMatcher(coreAPI)
// ----- weather provider (Open-Meteo when configured, Stub otherwise) -----
var weatherProvider weather.Provider
var weatherLocation string
if cfg.Voice.Weather != nil && cfg.Voice.Weather.Provider == "open-meteo" {
weatherProvider = weather.NewOpenMeteoProvider()
weatherLocation = cfg.Voice.Weather.DefaultLocation
log.Printf("voice: weather provider: open-meteo (default location: %s)", cfg.Voice.Weather.DefaultLocation)
} else {
weatherProvider = weather.NewStubProvider()
log.Printf("voice: weather provider: stub (not configured)")
}
// The replier uses the same llama-server as the phraser.
var llmClient *llm.Client
if lp, ok := phr.(*phraser.LLMPhraser); ok {
llmClient = llm.New(lp.BaseURL(), 60*time.Second)
}
// ----- router (the cascade; floor examples seed the classifier) -----
// The act matcher's allowlist is exactly the enabled tool names — the
// router only matches acts the executor can run (one source of truth).
threshold := cfg.Voice.RouterThreshold
if threshold <= 0 {
threshold = config.DefaultRouterThreshold
}
// The resident model routes by default: 63.2% of held-out intents right
// against the classifier's 50.0%, at about 1s a turn instead of 30ms (see
// config.VoiceConfig.LLMRouter). The classifier always stays wired as the
// fallback, so a model error never breaks a turn.
rtr := buildRouter(emb, matcher, threshold, pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient))
// ----- sessions registry (shared with voicesink) -----
sessions := voice.NewSessions()
w.sessions = sessions
// ----- voice sink (proactive nudges: dispatcher → voicesink → tts → push to client) -----
w.voiceSink = voicesink.New(synthesizer, sessions)
// ----- memory (long-term vector storage) -----
// Persistent (store-backed, survives restarts) when the daemon passes one;
// falls back to the in-memory floor otherwise (tests / no-store paths).
if memStore == nil {
memStore = memory.NewInMemoryStore()
}
// ----- dialogue (multi-turn slot carry-over; 2-min follow-up window) -----
// Store-backed when the daemon passes a store, so a restart mid-conversation
// keeps the thread (Vikunja #363). Sessions past their TTL are dropped on
// load, never revived. Clarify's parked question stays in memory only.
var dialogueSessions *dialogue.SessionStore
if dataStore != nil {
dialogueSessions = dialogue.NewPersistentSessionStore(2*time.Minute, dataStore)
if err := dialogueSessions.Load(context.Background(), time.Now()); err != nil {
log.Printf("dialogue: load saved sessions: %v", err)
}
} else {
dialogueSessions = dialogue.NewSessionStore(2 * time.Minute)
}
clarifyStore := dialogue.NewClarifyStore(clarifyTTL)
timeParser := router.NewPythonDateParser()
// ----- replier (LLM-backed when the engine is on, Stub floor otherwise) -----
replier := voice.Replier(voice.NewStubReplier())
if llmClient != nil {
replier = newLLMReplier(llmClient, contextBlockFn(cfg, time.Now))
}
// ----- the handler (the reactive path; closes over stt / tts / router / coreAPI / memory) -----
h := &reactiveHandler{
stt: transcriber,
tts: synthesizer,
router: rtr,
embedder: emb,
api: coreAPI,
tools: exec,
matcher: matcher,
replier: replier,
phraser: phr,
now: time.Now,
weatherProvider: weatherProvider,
weatherLocation: weatherLocation,
memStore: memStore,
dataStore: dataStore,
dialogueSessions: dialogueSessions,
clarifyStore: clarifyStore,
// 0 here (unset config) ⇒ the dialogue default.
clarifyMaxAttempts: cfg.Voice.ClarifyMaxAttempts,
extractor: router.Extractor{Time: timeParser, Acts: matcher, Facts: router.DefaultFactParser{}},
queryMinScore: cfg.Voice.QueryMinScore,
queryMinMargin: cfg.Voice.QueryMinMargin,
timeParser: timeParser,
ecosystem: eco,
}
// ----- the server (TCP listener) -----
srv := voice.NewServer(cfg.Voice.Bind, h, sessions)
if err := srv.Listen(); err != nil {
w.close()
return nil, fmt.Errorf("voice listen: %w", err)
}
w.server = srv
w.handler = h
return w, nil
}
// pickLLMRouter returns the LLM router when the operator asked for it and there
// is a llama-server to talk to, and nil otherwise. nil is safe: the cascade then
// routes with the classifier, so an unusable setting costs accuracy, not turns.
func pickLLMRouter(enabled bool, c *llm.Client) *router.LLMRouter {
if !enabled {
return nil
}
if c == nil {
log.Printf("voice: voice.llm_router is on but there is no llama-server to route with (the phraser is not an LLM phraser) — using the classifier instead")
return nil
}
log.Printf("voice: LLM router enabled")
return router.NewLLMRouter(c)
}
// buildRouter constructs the reactive-path router with the given embedder
// and confidence threshold.
// - stage-0 grammars from DefaultActMatcher whose fn allowlist is exactly
// the enabled tool names (actFns) — the router only matches acts the
// executor can run. Empty ⇒ every act refuses at the matcher.
// - The embedder is provided by wireVoice: HashEmbedder (floor) when no
// embedder config is present, or the ONNX multilingual model when
// configured — same interface, one constructor change.
// - 6 bootstrap examples covering the 5 intents + one compound-capture
// placeholder. Spec calls for ~10 per intent at production; this is the
// bootstrapping floor swapped by tuning the seed set later.
// - Threshold is from voice.router_threshold config (default 0.55).
func buildRouter(emb router.Embedder, acts router.ActMatcher, threshold float64, llmR *router.LLMRouter) *router.Router {
cls := router.NewClassifier(emb)
seedClassifier(cls)
grammars := router.DefaultGrammars(acts)
grammars = append(grammars, router.SystemTimeDateGrammars()...)
grammars = append(grammars, router.ReminderGrammar())
return router.New(router.Config{
Grammars: grammars,
Classifier: cls,
Extractor: router.Extractor{
Time: router.NewPythonDateParser(),
Acts: acts,
Facts: router.DefaultFactParser{},
},
Threshold: threshold,
LLM: llmR,
})
}
// seedDir is the directory containing intent seed files. Each file is named
// <intent>.txt and contains one training example per line (blank lines and
// lines starting with # are ignored). Relative to the working directory.
const seedDir = "models/seeds"
// seedClassifier floors the embedded examples so the cold-boot path
// doesn't return ErrNoIntents. Loads examples from seedDir — one file per
// intent (act.txt, reminder.txt, fact.txt, note.txt, query.txt). When the
// classifier can't decide it falls through to Clarify — the last-resort
// path asks the user to rephrase rather than guessing wrong.
func seedClassifier(c *router.Classifier) {
intents := []router.Intent{
router.IntentAct,
router.IntentReminder,
router.IntentFact,
router.IntentNote,
router.IntentQuery,
router.IntentChat,
router.IntentSystem,
}
total := 0
for _, intent := range intents {
n, err := loadSeedFile(c, intent)
if err != nil {
log.Printf("voice: seed %s: %v", intent, err)
continue
}
total += n
}
log.Printf("voice: loaded %d seed examples from %s", total, seedDir)
}
func loadSeedFile(c *router.Classifier, intent router.Intent) (int, error) {
path := filepath.Join(seedDir, string(intent)+".txt")
f, err := os.Open(path)
if err != nil {
return 0, fmt.Errorf("open %s: %w", path, err)
}
defer f.Close()
var count int
sc := bufio.NewScanner(f)
for sc.Scan() {
line := strings.TrimSpace(sc.Text())
if line == "" || strings.HasPrefix(line, "#") {
continue
}
if err := c.AddExample(context.Background(), intent, line); err != nil {
log.Printf("voice: seed %s: skipping %q: %v", intent, line, err)
continue
}
count++
}
if err := sc.Err(); err != nil {
return count, fmt.Errorf("scan %s: %w", path, err)
}
return count, nil
}
// seedTools upserts the config-declared tools into the store as enabled. Editing
// mavend.json is a human act, so a config tool is enabled by definition; this
// makes the declarative config the reproducible bootstrap while the store stays
// the single runtime source of truth (mavweb enables ad-hoc ones on top).
func seedTools(api ipc.CoreAPI, tools []config.ToolConfig) {
ctx := context.Background()
now := time.Now()
n := 0
for _, tc := range tools {
if tc.Name == "" || len(tc.Cmd) == 0 {
log.Printf("voice: skipping malformed tool config %+v", tc)
continue
}
if err := api.EnableTool(ctx, tc.Name, tc.Cmd, tc.Destructive, tc.Scope, now); err != nil {
log.Printf("voice: seed tool %q: %v", tc.Name, err)
continue
}
n++
}
log.Printf("voice: seeded %d act tools from config", n)
}
// reembedOnStart is the -reembed flag (set in run()). Opt-in on purpose: see
// runReembed.
var reembedOnStart bool
// checkStoredEmbedder compares the embedder we just loaded with the one that
// wrote the vectors already in the DB (Vikunja #378).
//
// The two models we have both make 384-dim vectors, so a size check catches
// nothing: after a swap, recall silently compares vectors from different
// spaces and the scores are noise. So we say it out loud. Recall itself is not
// changed here — the fix is `mavend -reembed`.
func checkStoredEmbedder(dataStore *store.Store, emb router.Embedder) {
if dataStore == nil {
return
}
current := router.EmbedderID(emb)
if reembedOnStart {
runReembed(dataStore, emb, current)
return
}
stored, mismatch, err := dataStore.CheckEmbedder(context.Background(), current)
if err != nil {
log.Printf("voice: embedder marker check failed: %v", err)
return
}
if mismatch {
log.Printf("voice: WARNING embedder MISMATCH — stored vectors were written by %q but the configured embedder is %q; recall scores are noise until the notes and facts are re-embedded — run `mavend -reembed` once (Vikunja #378)", stored, current)
return
}
log.Printf("voice: embedder marker ok (%s)", current)
}
// runReembed is the one-shot backfill behind -reembed.
//
// Why a flag and not automatic on mismatch: the embedder is ONNX on the
// laptop's CPU, so a few thousand notes is minutes of work. Doing that silently
// inside a normal start would look like the daemon hanging on boot. So the user
// runs it once, deliberately, after an embedder swap; the mismatch warning
// above tells them to. It re-embeds, logs what it did, and then the daemon
// carries on serving as usual — no separate binary, no second start needed.
func runReembed(dataStore *store.Store, emb router.Embedder, current string) {
log.Printf("voice: re-embedding stored notes and facts with %s — this can take a few minutes, do not interrupt", current)
res, err := dataStore.ReembedAll(context.Background(), current,
// EmbedPassage, not EmbedQuery: these are stored texts being searched
// FOR, which is the side they were written with.
func(ctx context.Context, text string) ([]float32, error) {
return router.EmbedPassage(ctx, emb, text)
})
if err != nil {
log.Printf("voice: re-embed FAILED, nothing was changed and no marker was written — safe to run again: %v", err)
return
}
if res.Skipped {
log.Printf("voice: re-embed skipped — the stored vectors were already written by %s", current)
return
}
log.Printf("voice: re-embed done — %d notes in the notes table, %d notes and %d facts in the memory index, took %s; stored vectors now belong to %s",
res.Notes, res.MemNotes, res.Facts, res.Took.Round(time.Second), current)
// A row with no text cannot be re-embedded, so its vector is still the old
// model's noise while the marker now says everything is current. Both write
// paths always store the text, so this should be zero — say it loudly
// rather than bury it in the line above if it ever isn't.
if res.NoText > 0 {
log.Printf("voice: WARNING %d stored rows had no text, so their vectors could not be re-embedded and are still noise; they will never match anything useful (Vikunja #378)", res.NoText)
}
}
+50
View File
@@ -0,0 +1,50 @@
// Package main — weatherq.go holds the weather-query keyword helpers: does
// this utterance ask about weather at all, and which city (if any) did it
// name. Both are plain substring/lookup matching, not NLU — extend this file
// rather than voice.go for anything in that shape.
package main
import "strings"
// isWeatherQuery returns true if the utterance is about weather.
func isWeatherQuery(u string) bool {
lower := strings.ToLower(u)
return strings.Contains(lower, "погод") ||
strings.Contains(lower, "градус") ||
strings.Contains(lower, "температур") ||
strings.Contains(lower, "дожд") ||
strings.Contains(lower, "холод") ||
strings.Contains(lower, "тепл") ||
strings.Contains(lower, "weather") ||
strings.Contains(lower, "temperature")
}
// extractWeatherLocation parses a location from the utterance, or falls back
// to the configured default. Very basic: just checks for known city names.
func extractWeatherLocation(u, defaultLoc string) string {
lower := strings.ToLower(u)
cities := map[string]string{
"москв": "Moscow",
"moscow": "Moscow",
"питер": "Saint Petersburg",
"spb": "Saint Petersburg",
"петербур": "Saint Petersburg",
"лондон": "London",
"london": "London",
"париж": "Paris",
"paris": "Paris",
"берлин": "Berlin",
"berlin": "Berlin",
"нью-йорк": "New York",
"new york": "New York",
}
for substr, name := range cities {
if strings.Contains(lower, substr) {
return name
}
}
if defaultLoc != "" {
return defaultLoc
}
return "Moscow"
}
+135
View File
@@ -0,0 +1,135 @@
package main
import (
"crypto/subtle"
"encoding/json"
"errors"
"io"
"log"
"net/http"
"strings"
"github.com/kami/maven/internal/calendar"
"github.com/kami/maven/internal/ipc"
)
// POST /api/ambient — the work calendar read (Vikunja #126).
//
// Maven does not hold a work credential. A corp mail or calendar session on the
// homelab ties the box's blast radius to the employer's data, so the work
// calendar is read as a SIGNAL instead: an Android notification-listener on the
// owner's phone posts meeting notifications here over wg/LAN, and the ones that
// clearly describe a meeting become calendar events at source=ambient:notif,
// confidence below 1.0. Mail as a notification signal, not a mailbox.
//
// Off unless configured: no -ambient-token, no route. The token is a shared
// secret because the poster is a phone service, not a browser — WebAuthn has no
// answer for a background Android service. The endpoint is write-only and
// accepts exactly one shape of write; it cannot read anything back out.
//
// A notification with no recognisable clock reading stores NOTHING. Maven is
// not a guesser-of-truth, and a mailbox of noise rendered as invented meetings
// is worse than a gap.
// ambientMaxBody bounds the request. A notification is two short lines.
const ambientMaxBody = 8 << 10
type ambientResp struct {
Stored bool `json:"stored"`
Key string `json:"key,omitempty"`
Reason string `json:"reason,omitempty"`
}
// handleAmbient ingests one relayed notification. token is the configured
// shared secret; an empty token means the capability is off and the handler is
// never registered, so it is treated as a hard failure here too.
func handleAmbient(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI, token string) {
if r.Method != http.MethodPost {
http.Error(w, "POST only", http.StatusMethodNotAllowed)
return
}
if token == "" {
http.Error(w, "ambient ingest disabled (no -ambient-token)", http.StatusServiceUnavailable)
return
}
if !ambientAuthorized(r, token) {
http.Error(w, "unauthorized", http.StatusUnauthorized)
return
}
if core == nil {
http.Error(w, "ambient ingest disabled (no -core)", http.StatusServiceUnavailable)
return
}
var n calendar.Notification
body, err := io.ReadAll(io.LimitReader(r.Body, ambientMaxBody))
if err != nil {
http.Error(w, "read failed", http.StatusBadRequest)
return
}
if err := json.Unmarshal(body, &n); err != nil {
http.Error(w, "bad json", http.StatusBadRequest)
return
}
if n.Posted.IsZero() {
writeAmbient(w, http.StatusBadRequest, ambientResp{Reason: "posted_at is required"})
return
}
ev, ok := calendar.EventFromNotification(n)
if !ok {
// Not an event. 202: the relay did its job, there is just nothing here
// worth remembering, and it must not retry.
writeAmbient(w, http.StatusAccepted, ambientResp{Reason: "no meeting time in notification"})
return
}
key := calendar.FactKey(ev)
val := calendar.FactValue(ev)
// Append-only discipline, same as cmd/mavcaldav: a phone reposts the same
// notification many times, and each repost is the same event.
if prev, err := core.LatestFactBySource(r.Context(), key, calendar.SourceAmbient); err == nil && prev.Value == val {
writeAmbient(w, http.StatusOK, ambientResp{Stored: false, Key: key, Reason: "unchanged"})
return
} else if err != nil && !errors.Is(err, ipc.ErrNoFact) {
log.Printf("ambient: read %s: %v", key, err)
http.Error(w, "read failed", http.StatusBadGateway)
return
}
// kind=env: an observation about the world, never a self-fact — a passive
// signal does not write truth about the owner. Confidence below 1.0 is the
// honest part: this is a notification about a meeting, not a reading of a
// calendar, and the query path hedges when it recites one.
if _, err := core.WriteFact(r.Context(), ipc.WriteFactReq{
Ts: ev.Start,
Kind: "env",
Key: key,
Value: val,
Source: calendar.SourceAmbient,
Confidence: calendar.AmbientConfidence,
}); err != nil {
log.Printf("ambient: write %s: %v", key, err)
http.Error(w, "write failed", http.StatusBadGateway)
return
}
log.Printf("ambient: %s=%s (%s, pkg=%s)", key, val, calendar.SourceAmbient, n.Package)
writeAmbient(w, http.StatusCreated, ambientResp{Stored: true, Key: key})
}
// ambientAuthorized accepts the token as a bearer header or as an X-Maven-Token
// header, compared in constant time.
func ambientAuthorized(r *http.Request, token string) bool {
got := strings.TrimSpace(strings.TrimPrefix(r.Header.Get("Authorization"), "Bearer"))
if got == "" {
got = strings.TrimSpace(r.Header.Get("X-Maven-Token"))
}
return subtle.ConstantTimeCompare([]byte(got), []byte(token)) == 1
}
func writeAmbient(w http.ResponseWriter, code int, resp ambientResp) {
w.Header().Set("Content-Type", "application/json")
w.WriteHeader(code)
json.NewEncoder(w).Encode(resp)
}
+223
View File
@@ -0,0 +1,223 @@
package main
import (
"context"
"encoding/json"
"fmt"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/calendar"
"github.com/kami/maven/internal/ipc"
)
const ambientTestToken = "s3cret"
// ambientCore adds provenance-scoped reads to fakeCore, which the dedupe path
// needs.
type ambientCore struct {
fakeCore
latest map[string]ipc.Fact // "key|source" → fact
readErr error
}
func (c *ambientCore) LatestFactBySource(_ context.Context, key, source string) (ipc.Fact, error) {
if c.readErr != nil {
return ipc.Fact{}, c.readErr
}
f, ok := c.latest[key+"|"+source]
if !ok {
return ipc.Fact{}, ipc.ErrNoFact
}
return f, nil
}
func postAmbient(t *testing.T, core ipc.CoreAPI, token string, n calendar.Notification) (*httptest.ResponseRecorder, ambientResp) {
t.Helper()
body, err := json.Marshal(n)
if err != nil {
t.Fatal(err)
}
req := httptest.NewRequest(http.MethodPost, "/api/ambient", strings.NewReader(string(body)))
req.Header.Set("Authorization", "Bearer "+ambientTestToken)
rr := httptest.NewRecorder()
handleAmbient(rr, req, core, token)
var resp ambientResp
json.Unmarshal(rr.Body.Bytes(), &resp)
return rr, resp
}
func meetingNotification() calendar.Notification {
return calendar.Notification{
Package: "com.google.android.gm",
Title: "Планёрка",
Text: "10:00-10:30",
Posted: time.Date(2026, 8, 3, 9, 40, 0, 0, time.UTC),
}
}
func TestHandleAmbientStoresMeeting(t *testing.T) {
core := &ambientCore{}
rr, resp := postAmbient(t, core, ambientTestToken, meetingNotification())
if rr.Code != http.StatusCreated {
t.Fatalf("status = %d, want 201: %s", rr.Code, rr.Body)
}
if !resp.Stored {
t.Errorf("resp = %+v, want stored", resp)
}
if len(core.writeLog) != 1 {
t.Fatalf("expected 1 fact write, got %d", len(core.writeLog))
}
got := core.writeLog[0]
if got.Source != calendar.SourceAmbient {
t.Errorf("source = %q, want %q", got.Source, calendar.SourceAmbient)
}
if got.Confidence >= 1.0 {
t.Errorf("confidence = %v — a notification is not a calendar read", got.Confidence)
}
if got.Confidence != calendar.AmbientConfidence {
t.Errorf("confidence = %v, want %v", got.Confidence, calendar.AmbientConfidence)
}
if got.Kind != "env" {
t.Errorf("kind = %q — a passive signal never writes a self-fact", got.Kind)
}
if want := "calendar_event_20260803_"; !strings.HasPrefix(got.Key, want) {
t.Errorf("key = %q, want prefix %q", got.Key, want)
}
if got.Value != "Планёрка @ 10:00-10:30" {
t.Errorf("value = %q", got.Value)
}
}
// A phone reposts the same notification many times. Each repost is the same
// event, and the append-only log must not fill with duplicates.
func TestHandleAmbientDedupesReposts(t *testing.T) {
core := &ambientCore{}
postAmbient(t, core, ambientTestToken, meetingNotification())
if len(core.writeLog) != 1 {
t.Fatalf("first post did not write")
}
w := core.writeLog[0]
core.latest = map[string]ipc.Fact{w.Key + "|" + w.Source: {Value: w.Value}}
rr, resp := postAmbient(t, core, ambientTestToken, meetingNotification())
if rr.Code != http.StatusOK {
t.Errorf("status = %d, want 200 for an unchanged repost", rr.Code)
}
if resp.Stored {
t.Error("a repost must not be stored again")
}
if len(core.writeLog) != 1 {
t.Errorf("wrote %d facts, want 1", len(core.writeLog))
}
}
// The conservative half: noise stores nothing at all.
func TestHandleAmbientIgnoresNonMeetings(t *testing.T) {
core := &ambientCore{}
rr, resp := postAmbient(t, core, ambientTestToken, calendar.Notification{
Package: "com.google.android.gm",
Title: "3 новых письма",
Posted: time.Now(),
})
if rr.Code != http.StatusAccepted {
t.Errorf("status = %d, want 202 (accepted, nothing to store — the relay must not retry)", rr.Code)
}
if resp.Stored {
t.Error("a notification with no meeting time must store nothing")
}
if len(core.writeLog) != 0 {
t.Fatalf("wrote %d facts for a non-meeting", len(core.writeLog))
}
}
func TestHandleAmbientAuth(t *testing.T) {
body := `{"title":"Планёрка 10:00","posted_at":"2026-08-03T09:40:00Z"}`
newReq := func(hdr, val string) *http.Request {
r := httptest.NewRequest(http.MethodPost, "/api/ambient", strings.NewReader(body))
if hdr != "" {
r.Header.Set(hdr, val)
}
return r
}
t.Run("no token rejected", func(t *testing.T) {
core := &ambientCore{}
rr := httptest.NewRecorder()
handleAmbient(rr, newReq("", ""), core, ambientTestToken)
if rr.Code != http.StatusUnauthorized {
t.Errorf("status = %d, want 401", rr.Code)
}
if len(core.writeLog) != 0 {
t.Error("an unauthorized post must not write")
}
})
t.Run("wrong token rejected", func(t *testing.T) {
rr := httptest.NewRecorder()
handleAmbient(rr, newReq("Authorization", "Bearer nope"), &ambientCore{}, ambientTestToken)
if rr.Code != http.StatusUnauthorized {
t.Errorf("status = %d, want 401", rr.Code)
}
})
t.Run("X-Maven-Token accepted", func(t *testing.T) {
rr := httptest.NewRecorder()
handleAmbient(rr, newReq("X-Maven-Token", ambientTestToken), &ambientCore{}, ambientTestToken)
if rr.Code != http.StatusCreated {
t.Errorf("status = %d, want 201: %s", rr.Code, rr.Body)
}
})
t.Run("capability off", func(t *testing.T) {
rr := httptest.NewRecorder()
handleAmbient(rr, newReq("Authorization", "Bearer "+ambientTestToken), &ambientCore{}, "")
if rr.Code != http.StatusServiceUnavailable {
t.Errorf("status = %d, want 503 when no token is configured", rr.Code)
}
})
t.Run("GET rejected", func(t *testing.T) {
rr := httptest.NewRecorder()
r := httptest.NewRequest(http.MethodGet, "/api/ambient", nil)
handleAmbient(rr, r, &ambientCore{}, ambientTestToken)
if rr.Code != http.StatusMethodNotAllowed {
t.Errorf("status = %d, want 405 — the ingest is write-only", rr.Code)
}
})
}
func TestHandleAmbientBadInput(t *testing.T) {
t.Run("bad json", func(t *testing.T) {
req := httptest.NewRequest(http.MethodPost, "/api/ambient", strings.NewReader("{nope"))
req.Header.Set("X-Maven-Token", ambientTestToken)
rr := httptest.NewRecorder()
handleAmbient(rr, req, &ambientCore{}, ambientTestToken)
if rr.Code != http.StatusBadRequest {
t.Errorf("status = %d, want 400", rr.Code)
}
})
t.Run("missing posted_at", func(t *testing.T) {
req := httptest.NewRequest(http.MethodPost, "/api/ambient", strings.NewReader(`{"title":"Планёрка 10:00"}`))
req.Header.Set("X-Maven-Token", ambientTestToken)
rr := httptest.NewRecorder()
handleAmbient(rr, req, &ambientCore{}, ambientTestToken)
if rr.Code != http.StatusBadRequest {
t.Errorf("status = %d, want 400", rr.Code)
}
})
t.Run("read error surfaces", func(t *testing.T) {
core := &ambientCore{readErr: fmt.Errorf("socket closed")}
rr, _ := postAmbient(t, core, ambientTestToken, meetingNotification())
if rr.Code != http.StatusBadGateway {
t.Errorf("status = %d, want 502", rr.Code)
}
})
}
+77 -3
View File
@@ -17,11 +17,12 @@ import (
) )
// fakeCore records the mutating calls handleTools makes and returns canned // fakeCore records the mutating calls handleTools makes and returns canned
// tool lists / errors. Embedding ipc.CoreAPI (nil) satisfies the large // tool lists / errors. Embedding ipc.UnimplementedCoreAPI satisfies the large
// interface — only the methods the handlers touch are overridden; any other // interface — only the methods the handlers touch are overridden; any other
// call would nil-panic, which is fine since the handlers never make them. // call returns ipc.ErrNotImplemented instead of nil-panicking, so a test that
// accidentally exercises an undeclared method fails loudly.
type fakeCore struct { type fakeCore struct {
ipc.CoreAPI ipc.UnimplementedCoreAPI
proposed, enabled []ipc.Tool proposed, enabled []ipc.Tool
listErr error listErr error
@@ -62,6 +63,18 @@ type fakeCore struct {
// for handleTrace tests // for handleTrace tests
tickTrace ipc.TickTrace tickTrace ipc.TickTrace
traceErr error traceErr error
// for handleChatAPI tests
chatText string
chatErr error
}
func (f *fakeCore) Chat(_ context.Context, text string) (string, error) {
f.chatText = text
if f.chatErr != nil {
return "", f.chatErr
}
return "поняла", nil
} }
func (f *fakeCore) EnableTool(_ context.Context, name string, cmd []string, destructive bool, scope string, _ time.Time) error { func (f *fakeCore) EnableTool(_ context.Context, name string, cmd []string, destructive bool, scope string, _ time.Time) error {
@@ -1044,3 +1057,64 @@ func TestHandleRoutines_NilCore_503(t *testing.T) {
t.Fatalf("status = %d, want 503", rr.Code) t.Fatalf("status = %d, want 503", rr.Code)
} }
} }
// --- handleChatAPI step-up gate (Vikunja #317) ---
//
// POST /api/chat reaches the router, the LLM and the act path, so it carries
// the same gate as POST /tools and POST /api/revert.
func postChat(text string) *http.Request {
req := httptest.NewRequest(http.MethodPost, "/api/chat", strings.NewReader("text="+url.QueryEscape(text)))
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
return req
}
func TestHandleChatAPI_RequireStepUp_FailsClosed(t *testing.T) {
core := &fakeCore{}
rr := httptest.NewRecorder()
handleChatAPI(rr, postChat("выключи свет"), core, nil, true)
if rr.Code != http.StatusForbidden {
t.Fatalf("status = %d, want 403; body=%s", rr.Code, rr.Body.String())
}
if core.chatText != "" {
t.Errorf("core.Chat called with %q, but -require-stepup should deny", core.chatText)
}
}
func TestHandleChatAPI_UnassertedSession_Denied(t *testing.T) {
core := &fakeCore{}
rr := httptest.NewRecorder()
handleChatAPI(rr, postChat("выключи свет"), core, webauthn.NewPasskeySession(5*time.Minute), false)
if rr.Code != http.StatusForbidden {
t.Fatalf("status = %d, want 403", rr.Code)
}
if core.chatText != "" {
t.Errorf("core.Chat called with %q despite an unasserted session", core.chatText)
}
}
func TestHandleChatAPI_AssertedSession_PassesGate(t *testing.T) {
core := &fakeCore{}
rr := httptest.NewRecorder()
handleChatAPI(rr, postChat("привет"), core, stepUpSession(), true)
if rr.Code != http.StatusSeeOther {
t.Fatalf("status = %d, want 303; body=%s", rr.Code, rr.Body.String())
}
if core.chatText != "привет" {
t.Errorf("core.Chat text = %q, want %q", core.chatText, "привет")
}
}
// Default deploy: WebAuthn unconfigured and -require-stepup off ⇒ chat keeps
// working, resting on the transport-level auth in front of mavweb.
func TestHandleChatAPI_FailOpenByDefault(t *testing.T) {
core := &fakeCore{}
rr := httptest.NewRecorder()
handleChatAPI(rr, postChat("привет"), core, nil, false)
if rr.Code != http.StatusSeeOther {
t.Fatalf("status = %d, want 303", rr.Code)
}
if core.chatText != "привет" {
t.Errorf("core.Chat text = %q, want %q", core.chatText, "привет")
}
}
+219 -10
View File
@@ -58,6 +58,9 @@ var notificationsHTML string
//go:embed reminders.html //go:embed reminders.html
var remindersHTML string var remindersHTML string
//go:embed tasks.html
var tasksHTML string
//go:embed voice.html //go:embed voice.html
var voiceHTML string var voiceHTML string
@@ -97,6 +100,7 @@ var sidebarSections = []struct {
Pages: []struct{ Label, URL, Key string }{ Pages: []struct{ Label, URL, Key string }{
{Label: "Rule Trace", URL: "/trace", Key: "trace"}, {Label: "Rule Trace", URL: "/trace", Key: "trace"},
{Label: "Notifications", URL: "/notifications", Key: "notifications"}, {Label: "Notifications", URL: "/notifications", Key: "notifications"},
{Label: "Tasks", URL: "/tasks", Key: "tasks"},
{Label: "Reminders", URL: "/reminders", Key: "reminders"}, {Label: "Reminders", URL: "/reminders", Key: "reminders"},
{Label: "Routines", URL: "/routines", Key: "routines"}, {Label: "Routines", URL: "/routines", Key: "routines"},
{Label: "Morning", URL: "/morning", Key: "morning"}, {Label: "Morning", URL: "/morning", Key: "morning"},
@@ -170,6 +174,8 @@ func pageIcon(key string) string {
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-wave"/></svg>` return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-wave"/></svg>`
case "notifications": case "notifications":
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-bell"/></svg>` return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-bell"/></svg>`
case "tasks":
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-grid"/></svg>`
case "reminders": case "reminders":
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-calendar"/></svg>` return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-calendar"/></svg>`
case "routines": case "routines":
@@ -202,6 +208,8 @@ func pageTitle(key string) string {
return "Rule Trace" return "Rule Trace"
case "notifications": case "notifications":
return "Notifications" return "Notifications"
case "tasks":
return "Tasks"
case "reminders": case "reminders":
return "Reminders" return "Reminders"
case "routines": case "routines":
@@ -314,7 +322,7 @@ func noCache(h http.Handler) http.Handler {
} }
func main() { func main() {
addr := flag.String("addr", ":9200", "HTTP listen address") addr := flag.String("addr", "127.0.0.1:9200", "HTTP listen address (loopback by default; pass e.g. \":9200\" or a LAN IP deliberately for wider exposure — POST /chat and /routines are state-changing)")
voiceAddr := flag.String("voice", "127.0.0.1:9100", "voice server TCP addr (host:port)") voiceAddr := flag.String("voice", "127.0.0.1:9100", "voice server TCP addr (host:port)")
// ntfyWS: the ntfy WebSocket subscribe URL the PWA connects to for in-app // ntfyWS: the ntfy WebSocket subscribe URL the PWA connects to for in-app
// nudge delivery, e.g. wss://ntfy.kvmx.ru/maven/ws?auth=<base64-token>. The // nudge delivery, e.g. wss://ntfy.kvmx.ru/maven/ws?auth=<base64-token>. The
@@ -329,11 +337,15 @@ func main() {
coreSock := flag.String("core", "", "mavend IPC socket path for presence-signal ingest (empty = disabled)") coreSock := flag.String("core", "", "mavend IPC socket path for presence-signal ingest (empty = disabled)")
pkOrigin := flag.String("webauthn-origin", "", "WebAuthn origin URL (e.g. https://maven.kvmx.ru)") pkOrigin := flag.String("webauthn-origin", "", "WebAuthn origin URL (e.g. https://maven.kvmx.ru)")
pkRPID := flag.String("webauthn-rpid", "", "WebAuthn RP ID (e.g. maven.kvmx.ru)") pkRPID := flag.String("webauthn-rpid", "", "WebAuthn RP ID (e.g. maven.kvmx.ru)")
requireStepUp := flag.Bool("require-stepup", false, "fail closed on step-up-gated actions (/tools POST, /api/revert) when WebAuthn step-up cannot be asserted; default false preserves the historical fail-open behaviour") requireStepUp := flag.Bool("require-stepup", false, "fail closed on step-up-gated actions (POST /tools, /routines, /api/revert, /api/chat) when WebAuthn step-up cannot be asserted; default false preserves the historical fail-open behaviour")
pkFile := flag.String("passkey-file", "./passkeys.json", "path to WebAuthn credential store (JSON)") pkFile := flag.String("passkey-file", "./passkeys.json", "path to WebAuthn credential store (JSON)")
nexusURL := flag.String("nexus", "", "Nexus base URL for the /ecosystem panel (empty = not configured)") nexusURL := flag.String("nexus", "", "Nexus base URL for the /ecosystem panel (empty = not configured)")
praxisURL := flag.String("praxis", "", "Praxis base URL for the /ecosystem panel (empty = not configured)") praxisURL := flag.String("praxis", "", "Praxis base URL for the /ecosystem panel (empty = not configured)")
hexisURL := flag.String("hexis", "", "Hexis base URL for the /ecosystem panel (empty = not configured)") hexisURL := flag.String("hexis", "", "Hexis base URL for the /ecosystem panel (empty = not configured)")
// Shared secret for POST /api/ambient, the notification-relay ingest that
// reads the work calendar as a signal instead of holding a work credential
// (see ambient.go). Empty ⇒ the route is not registered at all.
ambientToken := flag.String("ambient-token", "", "shared secret for POST /api/ambient notification ingest (empty = ingest disabled, route not registered)")
flag.Parse() flag.Parse()
var core ipc.CoreAPI var core ipc.CoreAPI
@@ -381,6 +393,14 @@ func main() {
mux.HandleFunc("/api/signal", func(w http.ResponseWriter, r *http.Request) { mux.HandleFunc("/api/signal", func(w http.ResponseWriter, r *http.Request) {
handleSignal(w, r, core) handleSignal(w, r, core)
}) })
// Off unless configured: no token, no route — an unconfigured ingest is not
// a 503 waiting to be probed, it does not exist.
if *ambientToken != "" {
mux.HandleFunc("/api/ambient", func(w http.ResponseWriter, r *http.Request) {
handleAmbient(w, r, core, *ambientToken)
})
log.Printf("mavweb: ambient notification ingest enabled at POST /api/ambient")
}
mux.HandleFunc("/dash", func(w http.ResponseWriter, r *http.Request) { mux.HandleFunc("/dash", func(w http.ResponseWriter, r *http.Request) {
handleDash(w, r, core) handleDash(w, r, core)
}) })
@@ -396,6 +416,11 @@ func main() {
mux.HandleFunc("/reminders", func(w http.ResponseWriter, r *http.Request) { mux.HandleFunc("/reminders", func(w http.ResponseWriter, r *http.Request) {
handleReminders(w, r, core) handleReminders(w, r, core)
}) })
// /tasks — capture + review. POST is not step-up gated; see handleTasks for
// why a task write is not in the same class as /tools or /routines.
mux.HandleFunc("/tasks", func(w http.ResponseWriter, r *http.Request) {
handleTasks(w, r, core)
})
mux.HandleFunc("/morning", func(w http.ResponseWriter, r *http.Request) { mux.HandleFunc("/morning", func(w http.ResponseWriter, r *http.Request) {
handleMorning(w, r, core) handleMorning(w, r, core)
}) })
@@ -434,9 +459,9 @@ func main() {
} }
if stepUpSession == nil { if stepUpSession == nil {
if *requireStepUp { if *requireStepUp {
log.Printf("SECURITY: step-up verification is DISABLED (-webauthn-origin/-webauthn-rpid unset) and -require-stepup is set: POST /tools (tool enable/disable/dismiss — defines and executes arbitrary argv) and POST /api/revert will be DENIED (403). Set -webauthn-origin and -webauthn-rpid to enable passkey step-up.") log.Printf("SECURITY: step-up verification is DISABLED (-webauthn-origin/-webauthn-rpid unset) and -require-stepup is set: POST /tools (tool enable/disable/dismiss — defines and executes arbitrary argv), POST /routines (accepting schedules recurring firing), POST /api/revert and POST /api/chat (reaches the router, the LLM and the act path) will be DENIED (403). Set -webauthn-origin and -webauthn-rpid to enable passkey step-up.")
} else { } else {
log.Printf("SECURITY WARNING: step-up verification is DISABLED because -webauthn-origin/-webauthn-rpid are unset. UNGUARDED SURFACES: POST /tools (defines arbitrary argv via name+cmd, which internal/tool then EXECUTES) and POST /api/revert (voids the latest fact for a key). These are protected only by whatever transport-level auth sits in front of mavweb (wg+nginx+auth) — do NOT expose -addr on a public interface. Set -webauthn-origin and -webauthn-rpid to require passkey step-up, or pass -require-stepup to fail closed instead.") log.Printf("SECURITY WARNING: step-up verification is DISABLED because -webauthn-origin/-webauthn-rpid are unset. UNGUARDED SURFACES: POST /tools (defines arbitrary argv via name+cmd, which internal/tool then EXECUTES), POST /routines (accepting schedules recurring firing), POST /api/revert (voids the latest fact for a key) and POST /api/chat (reaches the router, the LLM and, through applyAction, the act path). These are protected only by whatever transport-level auth sits in front of mavweb (wg+nginx+auth) — do NOT expose -addr on a public interface. Set -webauthn-origin and -webauthn-rpid to require passkey step-up, or pass -require-stepup to fail closed instead.")
} }
} }
@@ -454,14 +479,26 @@ func main() {
handleRoutines(w, r, core, stepUpSession, *requireStepUp) handleRoutines(w, r, core, stepUpSession, *requireStepUp)
}) })
// /api/revert voids the latest fact for a key — a store mutation, so it // State-changing routes on this server, and their gate (Vikunja #317):
// sits behind the same passkey step-up as tool enable (nil session ⇒ //
// WebAuthn unconfigured ⇒ transport-level auth only, same as /tools). // POST /tools step-up — defines argv that internal/tool executes
// POST /routines step-up — accepting schedules recurring firing
// POST /api/revert step-up — voids the latest fact for a key
// POST /api/chat step-up — reaches the router, LLM and the act path
// POST /api/signal none — appends a presence fact, no argv, no act
// POST /api/ptt, /ws none — proxy audio to mavend's voice port, which
// is itself only reachable inside the deploy
//
// "step-up" means stepUpOK: asserted passkey when WebAuthn is configured,
// otherwise fail-open unless -require-stepup, which denies.
//
// GET /chat only renders the page and echoes back the q/r query params the
// POST redirect set — nothing to gate.
mux.HandleFunc("/chat", func(w http.ResponseWriter, r *http.Request) { mux.HandleFunc("/chat", func(w http.ResponseWriter, r *http.Request) {
handleChatPage(w, r, core) handleChatPage(w, r, core)
}) })
mux.HandleFunc("/api/chat", func(w http.ResponseWriter, r *http.Request) { mux.HandleFunc("/api/chat", func(w http.ResponseWriter, r *http.Request) {
handleChatAPI(w, r, core) handleChatAPI(w, r, core, stepUpSession, *requireStepUp)
}) })
mux.HandleFunc("/api/revert", func(w http.ResponseWriter, r *http.Request) { mux.HandleFunc("/api/revert", func(w http.ResponseWriter, r *http.Request) {
handleRevert(w, r, core, stepUpSession, *requireStepUp) handleRevert(w, r, core, stepUpSession, *requireStepUp)
@@ -720,6 +757,8 @@ var passkeyTmpl = template.Must(template.New("passkey").Funcs(shellFuncs()).Pars
var voiceTmpl = template.Must(template.New("voice").Funcs(shellFuncs()).Parse(shellTopHTML + voiceHTML + shellBottomHTML)) var voiceTmpl = template.Must(template.New("voice").Funcs(shellFuncs()).Parse(shellTopHTML + voiceHTML + shellBottomHTML))
var tasksTmpl = template.Must(template.New("tasks").Funcs(shellFuncs()).Parse(shellTopHTML + tasksHTML + shellBottomHTML))
var routinesTmpl = template.Must(template.New("routines").Funcs(shellFuncs()).Parse(shellTopHTML + routinesHTML + shellBottomHTML)) var routinesTmpl = template.Must(template.New("routines").Funcs(shellFuncs()).Parse(shellTopHTML + routinesHTML + shellBottomHTML))
var traceTmpl = template.Must(template.New("trace").Funcs(func() template.FuncMap { var traceTmpl = template.Must(template.New("trace").Funcs(func() template.FuncMap {
@@ -790,6 +829,145 @@ func handleReminders(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
} }
} }
// taskRow is one line on /tasks, with every timestamp already formatted so the
// template holds no date logic.
type taskRow struct {
ID int64
Text string
Source string
Evidence string
Status string
Due string
Created string
Resolved string
}
// handleTasks serves the task review surface (GET) and the four writes it
// offers (POST): add, confirm, done, drop.
//
// Not step-up gated, unlike /tools and /routines, and the difference is the
// point: enabling a tool defines argv Maven will execute, and accepting a
// routine hands the tick loop a new standing reason to interrupt him. A task is
// neither — nothing in the tick loop reads the tasks table, so the worst a
// weaker caller can do here is write a line onto a list he reads himself. It
// still sits behind whatever transport auth fronts mavweb, like every other
// page.
//
// "confirm" is the only interesting move: it promotes a candidate Maven derived
// from something she read into work he owns. That review step is why derived
// tasks are captured as candidates in the first place.
func handleTasks(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
if core == nil {
http.Error(w, "tasks disabled (no -core)", http.StatusServiceUnavailable)
return
}
ctx := r.Context()
var msg, errMsg string
if r.Method == http.MethodPost {
var err error
msg, err = applyTaskPost(ctx, core, r)
if err != nil {
log.Printf("tasks: %v", err)
errMsg = err.Error()
}
}
all, err := core.ListTasks(ctx, "")
if err != nil {
log.Printf("tasks: list: %v", err)
http.Error(w, "tasks error: "+err.Error(), http.StatusBadGateway)
return
}
var cands, open, resolved []taskRow
for _, t := range all {
row := taskRow{
ID: t.ID, Text: t.Text, Source: t.Source, Evidence: t.Evidence,
Status: t.Status, Created: fmtTaskTime(&t.CreatedTs),
Due: fmtTaskDate(t.Due), Resolved: fmtTaskTime(t.Resolved),
}
switch t.Status {
case "candidate":
cands = append(cands, row)
case "open":
open = append(open, row)
default:
resolved = append(resolved, row)
}
}
w.Header().Set("Content-Type", "text/html; charset=utf-8")
if err := tasksTmpl.Execute(w, struct {
Msg, Err string
Candidates []taskRow
Open []taskRow
Resolved []taskRow
}{msg, errMsg, cands, open, resolved}); err != nil {
log.Printf("tasks render: %v", err)
}
}
// applyTaskPost performs one write and returns the message to show. A bad
// request returns an error, which the page renders inline rather than as a
// bare 400 — this is a form surface, not an API.
func applyTaskPost(ctx context.Context, core ipc.CoreAPI, r *http.Request) (string, error) {
action := r.FormValue("action")
if action == "add" {
text := strings.TrimSpace(r.FormValue("text"))
if text == "" {
return "", errors.New("empty task text")
}
req := ipc.CaptureTaskReq{Text: text, Source: "tap:web", Status: "open", Ts: time.Now()}
if d := r.FormValue("due"); d != "" {
due, err := time.ParseInLocation("2006-01-02", d, time.Local)
if err != nil {
return "", fmt.Errorf("bad due date %q", d)
}
req.Due = &due
}
resp, err := core.CaptureTask(ctx, req)
if err != nil {
return "", err
}
if !resp.Created {
return "already on the list", nil
}
return "added task", nil
}
var id int64
if n, _ := fmt.Sscanf(r.FormValue("id"), "%d", &id); n != 1 {
return "", errors.New("invalid id")
}
var status, msg string
switch action {
case "confirm":
status, msg = "open", "confirmed task"
case "done":
status, msg = "done", "task done"
case "drop":
status, msg = "dropped", "dropped task"
default:
return "", fmt.Errorf("unknown action %q", action)
}
if err := core.SetTaskStatus(ctx, id, status, time.Now()); err != nil {
return "", err
}
return msg, nil
}
func fmtTaskTime(t *time.Time) string {
if t == nil || t.IsZero() {
return "—"
}
return t.Local().Format("02 Jan 15:04")
}
func fmtTaskDate(t *time.Time) string {
if t == nil || t.IsZero() {
return "—"
}
return t.Local().Format("02 Jan")
}
// routineRow is one line on the page: what maven noticed, in her words, and // routineRow is one line on the page: what maven noticed, in her words, and
// how long ago she noticed it. // how long ago she noticed it.
type routineRow struct { type routineRow struct {
@@ -935,12 +1113,32 @@ func handleMorning(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
http.Error(w, "core read failed", http.StatusBadGateway) http.Error(w, "core read failed", http.StatusBadGateway)
return return
} }
view := morningView{Routines: status}
// The day plan (#128) shows on this page because it is the same question at
// a different scale. A plan read that fails must not take the checklist
// down with it — the page degrades to what it had before.
plan, err := core.DayPlan(ctx)
if err != nil {
log.Printf("morning: day plan: %v", err)
view.PlanErr = err.Error()
} else {
view.Plan = &plan
}
w.Header().Set("Content-Type", "text/html; charset=utf-8") w.Header().Set("Content-Type", "text/html; charset=utf-8")
if err := morningTmpl.Execute(w, status); err != nil { if err := morningTmpl.Execute(w, view); err != nil {
log.Printf("morning render: %v", err) log.Printf("morning render: %v", err)
} }
} }
// morningView — what /morning renders: today's plan on top, the checklist
// state under it. PlanErr is set instead of Plan when the core could not build
// a plan, so the page says so rather than showing an empty day.
type morningView struct {
Plan *ipc.DayPlan
PlanErr string
Routines []ipc.MorningRoutineStatus
}
func handleVoice(w http.ResponseWriter, r *http.Request) { func handleVoice(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "text/html; charset=utf-8") w.Header().Set("Content-Type", "text/html; charset=utf-8")
if err := voiceTmpl.Execute(w, nil); err != nil { if err := voiceTmpl.Execute(w, nil); err != nil {
@@ -1258,7 +1456,14 @@ func handleChatPage(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
} }
// handleChatAPI processes a chat message POST and redirects back to /chat. // handleChatAPI processes a chat message POST and redirects back to /chat.
func handleChatAPI(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) { //
// State-changing, and the widest surface on this server: the text reaches the
// router, the LLM, and through mavend's applyAction the whole action path
// including `act` — so it is gated on the same step-up as POST /tools and
// POST /api/revert (Vikunja #317). With WebAuthn unconfigured the gate is
// fail-open exactly like the others (see stepUpOK); with -require-stepup it
// denies, which is the point of that flag.
func handleChatAPI(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI, session *webauthn.PasskeySession, requireStepUp bool) {
if r.Method != http.MethodPost { if r.Method != http.MethodPost {
http.Error(w, "POST only", http.StatusMethodNotAllowed) http.Error(w, "POST only", http.StatusMethodNotAllowed)
return return
@@ -1267,6 +1472,10 @@ func handleChatAPI(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
http.Error(w, "chat disabled (no -core)", http.StatusServiceUnavailable) http.Error(w, "chat disabled (no -core)", http.StatusServiceUnavailable)
return return
} }
if !stepUpOK(session, requireStepUp) {
http.Error(w, "step-up required: assert a passkey first", http.StatusForbidden)
return
}
text := strings.TrimSpace(r.FormValue("text")) text := strings.TrimSpace(r.FormValue("text"))
if text == "" { if text == "" {
http.Redirect(w, r, "/chat", http.StatusSeeOther) http.Redirect(w, r, "/chat", http.StatusSeeOther)
+20 -2
View File
@@ -1,9 +1,27 @@
{{template "shellTop" "morning"}} {{template "shellTop" "morning"}}
<h1>Today</h1>
{{with .Plan}}
<div class=hint>{{.Date.Format "02.01.2006"}}</div>
{{if not .Items}}
<div class=hint>nothing planned</div>
{{else}}
<div class=scroll><table class=mono>
<tr><th>at<th>kind<th>what</tr>
{{range .Items}}<tr>
<td>{{.At.Format "15:04"}}</td>
<td class=gray>{{.Kind}}</td>
<td>{{if .Uncertain}}<span class=hint title="relayed notification, not a calendar read">похоже,</span> {{end}}{{.Text}}</td>
</tr>{{end}}
</table></div>
{{end}}
{{end}}
{{if .PlanErr}}<div class=hint>plan unavailable: {{.PlanErr}}</div>{{end}}
<h1>Morning Routines</h1> <h1>Morning Routines</h1>
{{if not .}} {{if not .Routines}}
<div class=hint>no morning routines configured</div> <div class=hint>no morning routines configured</div>
{{else}} {{else}}
{{range .}} {{range .Routines}}
<div class="mb-4"> <div class="mb-4">
<div><strong>{{.Name}}</strong> <div><strong>{{.Name}}</strong>
<span class={{if .Active}}green{{else}}gray{{end}}>{{if .Active}}active now{{else}}outside window{{end}}</span> <span class={{if .Active}}green{{else}}gray{{end}}>{{if .Active}}active now{{else}}outside window{{end}}</span>
+76
View File
@@ -0,0 +1,76 @@
{{template "shellTop" "tasks"}}
<h1>Tasks</h1>
{{if .Msg}}<div class="msg msg-ok">{{.Msg}}</div>{{end}}
{{if .Err}}<div class="msg msg-err">{{.Err}}</div>{{end}}
<section class=card>
<h2 class=card-title>add</h2>
<form method=post action=/tasks class=inline-form>
<input type=hidden name=action value=add>
<input type=text name=text placeholder="что нужно сделать" size=44 required>
<input type=date name=due title="due date (optional)">
<button class=btn>add</button>
</form>
</section>
{{if .Candidates}}
<section class=card>
<h2 class=card-title>found, not confirmed <span class=badge>{{len .Candidates}}</span></h2>
<div class=hint>maven derived these from something she read. nothing counts as your work until you confirm it.</div>
<div class=scroll><table>
<tr><th>task</th><th>where from</th><th>due</th><th>captured</th><th></th><th></th></tr>
{{range .Candidates}}<tr>
<td class=text-max>{{.Text}}</td>
<td class=hint>{{.Source}}{{if .Evidence}} — {{.Evidence}}{{end}}</td>
<td>{{.Due}}</td>
<td class=muted>{{.Created}}</td>
<td><form method=post action=/tasks class=inline-form>
<input type=hidden name=id value="{{.ID}}">
<input type=hidden name=action value=confirm>
<button class=btn>confirm</button></form></td>
<td><form method=post action=/tasks class=inline-form>
<input type=hidden name=id value="{{.ID}}">
<input type=hidden name=action value=drop>
<button class="btn btn-muted">drop</button></form></td>
</tr>{{end}}</table></div>
</section>
{{end}}
<section class=card>
<h2 class=card-title>open <span class=badge>{{len .Open}}</span></h2>
{{if .Open}}<div class=scroll><table>
<tr><th>task</th><th>from</th><th>due</th><th>captured</th><th></th><th></th></tr>
{{range .Open}}<tr>
<td class=text-max>{{.Text}}</td>
<td class=hint>{{.Source}}</td>
<td>{{.Due}}</td>
<td class=muted>{{.Created}}</td>
<td><form method=post action=/tasks class=inline-form>
<input type=hidden name=id value="{{.ID}}">
<input type=hidden name=action value=done>
<button class=btn>done</button></form></td>
<td><form method=post action=/tasks class=inline-form>
<input type=hidden name=id value="{{.ID}}">
<input type=hidden name=action value=drop>
<button class="btn btn-muted">drop</button></form></td>
</tr>{{end}}</table></div>
{{else}}<div class=empty>
<svg class=icon width="20" height="20"><use href="/ethos-icons.svg#i-grid"/></svg>
<div>no open tasks</div>
<div class=hint>add one above, or tell maven "добавь в задачи …"</div>
</div>{{end}}
</section>
{{if .Resolved}}
<section class=card>
<h2 class=card-title>resolved <span class=badge>{{len .Resolved}}</span></h2>
<div class=scroll><table>
<tr><th>task</th><th>status</th><th>when</th></tr>
{{range .Resolved}}<tr>
<td class=text-max>{{.Text}}</td>
<td><span class="badge {{.Status}}">{{.Status}}</span></td>
<td class=muted>{{.Resolved}}</td>
</tr>{{end}}</table></div>
</section>
{{end}}
{{template "shellBottom"}}
+166
View File
@@ -0,0 +1,166 @@
package main
import (
"context"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
)
// fakeTaskCore serves the /tasks handler: a canned list plus a log of the
// writes the page made.
type fakeTaskCore struct {
ipc.UnimplementedCoreAPI
tasks []ipc.Task
listErr error
captured []ipc.CaptureTaskReq
created bool
captureErr error
statusID int64
statusVal string
statusErr error
}
func (f *fakeTaskCore) ListTasks(_ context.Context, status string) ([]ipc.Task, error) {
if f.listErr != nil {
return nil, f.listErr
}
return f.tasks, nil
}
func (f *fakeTaskCore) CaptureTask(_ context.Context, req ipc.CaptureTaskReq) (ipc.CaptureTaskResp, error) {
f.captured = append(f.captured, req)
if f.captureErr != nil {
return ipc.CaptureTaskResp{}, f.captureErr
}
return ipc.CaptureTaskResp{ID: 7, Created: f.created}, nil
}
func (f *fakeTaskCore) SetTaskStatus(_ context.Context, id int64, status string, _ time.Time) error {
f.statusID, f.statusVal = id, status
return f.statusErr
}
func TestHandleTasksSplitsCandidatesFromOpen(t *testing.T) {
now := time.Date(2026, 8, 1, 9, 0, 0, 0, time.UTC)
resolved := now.Add(time.Hour)
core := &fakeTaskCore{tasks: []ipc.Task{
{ID: 1, Text: "купить молоко", Source: "tap:voice", Status: "open", CreatedTs: now},
{ID: 2, Text: "продлить страховку", Source: "email:kami", Evidence: "полис истекает", Status: "candidate", CreatedTs: now},
{ID: 3, Text: "полить цветы", Source: "tap:web", Status: "done", CreatedTs: now, Resolved: &resolved},
}}
rec := httptest.NewRecorder()
handleTasks(rec, httptest.NewRequest(http.MethodGet, "/tasks", nil), core)
if rec.Code != http.StatusOK {
t.Fatalf("status = %d", rec.Code)
}
body := rec.Body.String()
for _, want := range []string{
"купить молоко", "продлить страховку", "полить цветы",
"полис истекает", // the evidence trail is visible for review
"found, not confirmed", // candidates get their own section
} {
if !strings.Contains(body, want) {
t.Errorf("body missing %q", want)
}
}
// The candidate must offer confirm, and the open task must not.
if !strings.Contains(body, "value=confirm") {
t.Error("candidate row has no confirm action")
}
}
func TestHandleTasksAddCaptures(t *testing.T) {
core := &fakeTaskCore{created: true}
form := url.Values{"action": {"add"}, "text": {" позвонить в банк "}, "due": {"2026-08-05"}}
req := httptest.NewRequest(http.MethodPost, "/tasks", strings.NewReader(form.Encode()))
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
rec := httptest.NewRecorder()
handleTasks(rec, req, core)
if rec.Code != http.StatusOK {
t.Fatalf("status = %d", rec.Code)
}
if len(core.captured) != 1 {
t.Fatalf("captured %d requests, want 1", len(core.captured))
}
got := core.captured[0]
if got.Text != "позвонить в банк" {
t.Errorf("text = %q, want trimmed", got.Text)
}
if got.Source != "tap:web" {
t.Errorf("source = %q, want tap:web", got.Source)
}
if got.Status != "open" {
t.Errorf("status = %q — a task he typed himself is open, not a candidate", got.Status)
}
if got.Due == nil || got.Due.Format("2006-01-02") != "2026-08-05" {
t.Errorf("due = %v", got.Due)
}
if !strings.Contains(rec.Body.String(), "added task") {
t.Error("no confirmation message")
}
}
func TestHandleTasksAddSaysAlreadyOnTheList(t *testing.T) {
core := &fakeTaskCore{created: false}
form := url.Values{"action": {"add"}, "text": {"купить молоко"}}
req := httptest.NewRequest(http.MethodPost, "/tasks", strings.NewReader(form.Encode()))
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
rec := httptest.NewRecorder()
handleTasks(rec, req, core)
if !strings.Contains(rec.Body.String(), "already on the list") {
t.Error("a deduped capture must not claim it saved something new")
}
}
func TestHandleTasksStatusActions(t *testing.T) {
for _, tc := range []struct{ action, want string }{
{"confirm", "open"},
{"done", "done"},
{"drop", "dropped"},
} {
core := &fakeTaskCore{}
form := url.Values{"action": {tc.action}, "id": {"42"}}
req := httptest.NewRequest(http.MethodPost, "/tasks", strings.NewReader(form.Encode()))
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
handleTasks(httptest.NewRecorder(), req, core)
if core.statusID != 42 || core.statusVal != tc.want {
t.Errorf("%s → SetTaskStatus(%d, %q), want (42, %q)", tc.action, core.statusID, core.statusVal, tc.want)
}
}
}
func TestHandleTasksRejectsBadPost(t *testing.T) {
core := &fakeTaskCore{}
form := url.Values{"action": {"explode"}, "id": {"1"}}
req := httptest.NewRequest(http.MethodPost, "/tasks", strings.NewReader(form.Encode()))
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
rec := httptest.NewRecorder()
handleTasks(rec, req, core)
// The page still renders, with the error inline — and nothing was written.
if rec.Code != http.StatusOK {
t.Fatalf("status = %d", rec.Code)
}
if core.statusVal != "" || len(core.captured) != 0 {
t.Error("an unknown action must write nothing")
}
if !strings.Contains(rec.Body.String(), "unknown action") {
t.Error("error not surfaced on the page")
}
}
func TestHandleTasksNoCore(t *testing.T) {
rec := httptest.NewRecorder()
handleTasks(rec, httptest.NewRequest(http.MethodGet, "/tasks", nil), nil)
if rec.Code != http.StatusServiceUnavailable {
t.Errorf("status = %d, want 503", rec.Code)
}
}
+12
View File
@@ -8,6 +8,18 @@
# #
# Maven's own compose joins this same network (add `ecosystem` as an external # Maven's own compose joins this same network (add `ecosystem` as an external
# network there) to reach nexus:9740 / praxis:8989 / hexis:9741 directly. # network there) to reach nexus:9740 / praxis:8989 / hexis:9741 directly.
#
# NO RELEASE PINNING (Vikunja #354): each `build:` below points at a sibling
# WORKING TREE, so `up --build` ships whatever is checked out there, including
# uncommitted edits. Before bringing this up, check what you are about to
# deploy:
#
# for r in nexus praxis hexis; do git -C ../../../$r status --short; \
# git -C ../../../$r log -1 --oneline; done
#
# The host nginx that fronts these is deploy/ecosystem/nginx.conf — it binds
# the wg and LAN addresses only, with allow/deny. Keep it that way: none of
# these containers has auth of its own.
name: ecosystem name: ecosystem
services: services:
+76 -6
View File
@@ -1,13 +1,71 @@
# Reverse-proxy the three sibling admin UIs. Drop into your nginx sites (or the # Reverse-proxy Maven's own web UI plus the three sibling admin UIs. Drop into
# nginx-panel app) and reload. Assumes the compose publishes each service on # your nginx sites (or the nginx-panel app) and reload. Assumes the compose
# 127.0.0.1:<port>. Add TLS (certbot / your existing cert block) per server. # publishes each service on 127.0.0.1:<port>. Add TLS (certbot / your existing
# cert block) per server.
# #
# NOTE: hexis.<domain> previously pointed at the MCP tool repoint that # NOTE: hexis.<domain> previously pointed at the MCP tool repoint that
# elsewhere first (the app now owns hexis.*). # elsewhere first (the app now owns hexis.*).
#
# 10.42.0.1 and 192.168.1.104 below are THIS BOX's WireGuard and LAN
# addresses (homesrv) these admin UIs have no auth of their own, so the
# explicit bind + allow/deny below is what keeps them off the open internet.
# On a different box, replace both addresses with that box's wg and LAN IPs.
# Do NOT "fix" a failed bind by reverting to `listen 80` (all interfaces)
# that removes the only access control these containers have.
# maven.<domain> mavweb (docker-compose.yml publishes it on 127.0.0.1:9201).
# Same bind + ACL as the siblings, and for a stronger reason: mavweb serves
# POST /tools, which defines argv that internal/tool EXECUTES, plus POST
# /routines, /api/revert and /api/chat (Vikunja #317). Without
# -webauthn-origin/-webauthn-rpid mavweb has no auth of its own, so this block
# is the auth. If you add TLS and a basic-auth/oauth2-proxy layer, keep the
# allow/deny anyway belt and braces on an RCE surface.
#
# WebSocket upgrade matters here: /ws carries push-to-talk audio, so the
# Upgrade/Connection headers below are required, not decoration. The map keeps
# `Connection: upgrade` off plain requests; it sits in the http context, which
# is where sites-available files are included if your nginx already defines
# $connection_upgrade, drop this block.
map $http_upgrade $connection_upgrade {
default upgrade;
'' close;
}
server { server {
listen 80; listen 10.42.0.1:80;
listen 192.168.1.104:80;
server_name maven.kvmx.ru;
allow 10.42.0.0/24;
allow 192.168.1.0/24;
deny all;
# push-to-talk uploads raw PCM; the default 1m is enough for a short
# utterance but not for a long one.
client_max_body_size 32m;
location / {
proxy_pass http://127.0.0.1:9201;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection $connection_upgrade;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_read_timeout 300s; # an LLM turn can take minutes on the iGPU
}
}
server {
listen 10.42.0.1:80;
listen 192.168.1.104:80;
server_name nexus.kvmx.ru; server_name nexus.kvmx.ru;
allow 10.42.0.0/24;
allow 192.168.1.0/24;
deny all;
location / { location / {
proxy_pass http://127.0.0.1:9740; proxy_pass http://127.0.0.1:9740;
proxy_set_header Host $host; proxy_set_header Host $host;
@@ -18,8 +76,14 @@ server {
} }
server { server {
listen 80; listen 10.42.0.1:80;
listen 192.168.1.104:80;
server_name praxis.kvmx.ru; server_name praxis.kvmx.ru;
allow 10.42.0.0/24;
allow 192.168.1.0/24;
deny all;
location / { location / {
proxy_pass http://127.0.0.1:8989; proxy_pass http://127.0.0.1:8989;
proxy_set_header Host $host; proxy_set_header Host $host;
@@ -30,8 +94,14 @@ server {
} }
server { server {
listen 80; listen 10.42.0.1:80;
listen 192.168.1.104:80;
server_name hexis.kvmx.ru; server_name hexis.kvmx.ru;
allow 10.42.0.0/24;
allow 192.168.1.0/24;
deny all;
location / { location / {
proxy_pass http://127.0.0.1:9741; proxy_pass http://127.0.0.1:9741;
proxy_set_header Host $host; proxy_set_header Host $host;
+9 -3
View File
@@ -6,11 +6,12 @@
"state_dir": "/var/lib/maven", "state_dir": "/var/lib/maven",
"phraser": { "phraser": {
"model_path": "/opt/maven/models/llm/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf", "model_path": "/opt/maven/models/llm/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf",
"bin_path": "llama-server", "bin_path": "llama-server",
"n_gpu_layers": 99, "n_gpu_layers": 99,
"n_ctx": 2048, "n_ctx": 4096,
"timeout": "60s" "timeout": "60s",
"llm_nudges": false
}, },
"telegram": { "telegram": {
@@ -25,6 +26,11 @@
"severity_ceiling": 2 "severity_ceiling": 2
}, },
"pattern_proposals": {
"notify": false,
"cooldown": "24h"
},
"nexus": { "url": "http://nexus:9740" }, "nexus": { "url": "http://nexus:9740" },
"praxis": { "url": "http://praxis:8989" }, "praxis": { "url": "http://praxis:8989" },
"hexis": { "url": "http://hexis:9741" }, "hexis": { "url": "http://hexis:9741" },
+43
View File
@@ -25,3 +25,46 @@
6. Add `/eval` API method to `ipc.CoreAPI` (or reuse `Chat` with system context) so mavweb can show evaluation history 6. Add `/eval` API method to `ipc.CoreAPI` (or reuse `Chat` with system context) so mavweb can show evaluation history
7. Add `memory_eval` block to `deploy/mavend.json` 7. Add `memory_eval` block to `deploy/mavend.json`
8. Test with synthetic store state — verify observations match expected patterns 8. Test with synthetic store state — verify observations match expected patterns
---
## Status 2026-08-01 — foundation shipped (Vikunja #248)
**Shipped:** `internal/memeval` (not `internal/memory/eval.go``internal/store`
imports `internal/memory` for the vector backend, so an evaluator that reads
`store.Fact` there would close an import cycle). `Evaluator.Evaluate` reads
`RecentFacts` / `RecentNotes` / `RecentNudges`, prompts the resident model under
a GBNF grammar for at most three `{observation, confidence, suggested_action}`
objects, drops anything under `min_confidence`, deduplicates against what earlier
evaluations wrote, and records the rest as notes with source `infer:memory-eval`.
Driver: `cmd/mavend/memoryeval.go`, its own goroutine on its own ticker. Config:
the `memory_eval` block — **absent ⇒ the loop does not run**. Visibility: `/dash`
already renders notes with their source, so evaluation output is visible with no
UI change.
**Deliberately not shipped — this is policy, not an unfinished edge:**
- *Dispatching observations as care nudges (plan step 4).* An hourly LLM loop
with permission to speak is a machine for generating interruptions, and the
content is model-generated text about his own life. The evaluator has no
dispatcher reference at all, so it cannot reach a channel by accident. Wiring
it to `delivery.Dispatcher` is a separate decision with its own opt-in.
- *Acting on `suggested_action`.* It is recorded inside the note text and
interpreted by nobody. No reminder, routine or fact is created.
- *Writing observation embeddings.* Notes are written with a nil embedding, so
they stay out of the RAG recall pool. Feeding generated text back into the pool
it came from is how a small model starts citing its own guesses as evidence.
**Deferred, wants a decision or another capability:**
- *Plan step 6, the `/eval` IPC method and an evaluation-history view.* `/dash`
covers reading the output; a dedicated trace surface is worth building once
there is real output to look at, and it should probably show the prompt too.
- *`RecentEvents`.* The plan lists it; the evaluator reads facts, notes and
nudges. Detected action/object events already drive pattern proposals (#43), and
duplicating them here would mostly re-derive that.
- *Output quality is unmeasured.* There is no fixture for "did she notice
something true". The tests cover the machinery — empty store, confidence floor,
dedupe, own-notes exclusion, error handling — not the observations. Until
someone reads a week of real output on `/dash`, treat the wording and the
`min_confidence` default as unvalidated.
+33
View File
@@ -27,3 +27,36 @@
6. Wire voice query — `"что я обычно делаю?"` routes to `IntentQuery` → behavior profile lookup → LLM-phrased answer 6. Wire voice query — `"что я обычно делаю?"` routes to `IntentQuery` → behavior profile lookup → LLM-phrased answer
7. Add IPC read method `MethodGetBehaviorProfile` so mavweb can display it on `/dash` 7. Add IPC read method `MethodGetBehaviorProfile` so mavweb can display it on `/dash`
8. Test with synthetic fact history — verify weekly schedule is correctly inferred 8. Test with synthetic fact history — verify weekly schedule is correctly inferred
---
## Status (2026-08-01) — partially shipped, deliberately narrowed
Shipped on `overnight/behavior-profile`:
- `internal/memory/behavior.go``BuildProfile` counts habits per weekday out of
self-facts: distinct-day counts (`MinHabitDays = 2`), a median time-of-day, and
`FormatWeekdayRU` / `FormatOverallRU` for the spoken answer.
- `internal/router/habit.go``ParseHabitQuery`, which requires a habit marker
("обычно", "каждую", "привычки", …) and parses the weekday deterministically.
- `cmd/mavend/actions_query.go` — a `habits` query source, so "что я обычно делаю
по вторникам?" is answered.
**Not shipped, and not to be shipped as written:**
- *Step 3, LLM-generated profile stored as a fact.* The profile is COUNTED, not
generated. A 1.7B asked to summarise a year of habits produces fluent claims
about the owner's life that no row supports, and a wrong claim about him is the
most expensive kind of wrong maven can be. Counting is verifiable and cheap.
- *Step 5, incremental updates on fact write.* There is no cache to keep fresh —
the profile is recomputed on the question, so a new fact is already in the next
answer. A cached profile that can disagree with its own rows is two truths.
- *Step 4, proactive daily plan proposals via the dispatcher.* Maven is not a nag,
and a nudge at 08:00 every day proposing the day is the definition of one. The
sanctioned path from "she noticed a pattern" to "she acts on it" already exists:
`internal/pattern/detector.go` proposes a routine, and the owner accepts it on
`/routines`. It goes through him.
Still open, if wanted later: `MethodGetBehaviorProfile` + a `/dash` panel (step 7).
The counted profile needs no new IPC method to be *asked* about — the query source
reads `RecentFacts` over the existing surface — so this is a display concern only.
+5 -84
View File
@@ -7,7 +7,6 @@ import (
"path/filepath" "path/filepath"
"strings" "strings"
"testing" "testing"
"time"
"github.com/kami/maven/internal/ipc" "github.com/kami/maven/internal/ipc"
) )
@@ -359,8 +358,12 @@ func TestGate_IpcServer_ChatAllowedForEnrolledCaller(t *testing.T) {
// recordingAPI — a no-op CoreAPI that counts WriteFact invocations; the auth // recordingAPI — a no-op CoreAPI that counts WriteFact invocations; the auth
// check must reject before reaching it, otherwise the refusal leaks into the // check must reject before reaching it, otherwise the refusal leaks into the
// fake's counts and we fail. // fake's counts and we fail. Embeds ipc.UnimplementedCoreAPI so every method
// this test doesn't exercise returns ipc.ErrNotImplemented loudly instead of
// being hand-stubbed to a canned value nobody checks.
type recordingAPI struct { type recordingAPI struct {
ipc.UnimplementedCoreAPI
writes int writes int
chats int chats int
} }
@@ -369,89 +372,7 @@ func (r *recordingAPI) WriteFact(_ context.Context, _ ipc.WriteFactReq) (int64,
r.writes++ r.writes++
return int64(r.writes), nil return int64(r.writes), nil
} }
func (r *recordingAPI) LatestFact(_ context.Context, _ string) (ipc.Fact, error) {
return ipc.Fact{}, ipc.ErrNoFact
}
func (r *recordingAPI) LatestFactBySource(_ context.Context, _, _ string) (ipc.Fact, error) {
return ipc.Fact{}, ipc.ErrNoFact
}
func (r *recordingAPI) Since(_ context.Context, _ string, _ time.Time) (time.Duration, error) {
return 0, ipc.ErrNoFact
}
func (r *recordingAPI) Presence(_ context.Context) (ipc.Presence, error) {
return ipc.Presence{}, nil
}
func (r *recordingAPI) CreateReminder(_ context.Context, _ time.Time, _, _ string) (int64, error) {
return 1, nil
}
func (r *recordingAPI) MarkReminder(_ context.Context, _ int64, _ string) error { return nil }
func (r *recordingAPI) ListReminders(_ context.Context, _ int) ([]ipc.Reminder, error) {
return nil, nil
}
func (r *recordingAPI) TickTrace(_ context.Context) (ipc.TickTrace, error) {
return ipc.TickTrace{}, nil
}
func (r *recordingAPI) MorningStatus(_ context.Context) ([]ipc.MorningRoutineStatus, error) {
return nil, nil
}
func (r *recordingAPI) RecordNudge(_ context.Context, _, _, _ string, _ time.Time) (int64, error) {
return 1, nil
}
func (r *recordingAPI) ResolveNudge(_ context.Context, _ int64, _ string, _ time.Time) error {
return nil
}
func (r *recordingAPI) RecentOutcomes(_ context.Context, _ string, _ int) ([]string, error) {
return nil, nil
}
func (r *recordingAPI) RecentFacts(_ context.Context, _ int) ([]ipc.Fact, error) {
return nil, nil
}
func (r *recordingAPI) CalendarEvents(_ context.Context, _, _ time.Time) ([]ipc.Fact, error) {
return nil, nil
}
func (r *recordingAPI) RecentNudges(_ context.Context, _ int) ([]ipc.Nudge, error) {
return nil, nil
}
func (r *recordingAPI) WriteNote(_ context.Context, _ time.Time, _ string, _ []float32, _ string) (int64, error) {
return 1, nil
}
func (r *recordingAPI) QueryNotes(_ context.Context, _ []float32, _ int) ([]ipc.Note, error) {
return nil, nil
}
func (r *recordingAPI) RecentNotes(_ context.Context, _ int) ([]ipc.Note, error) {
return nil, nil
}
func (r *recordingAPI) ProposeTool(_ context.Context, _, _, _ string, _ time.Time) (bool, error) {
return false, nil
}
func (r *recordingAPI) EnableTool(_ context.Context, _ string, _ []string, _ bool, _ string, _ time.Time) error {
return nil
}
func (r *recordingAPI) DisableTool(_ context.Context, _ string) error {
return nil
}
func (r *recordingAPI) DeleteTool(_ context.Context, _ string) error {
return nil
}
func (r *recordingAPI) LookupTool(_ context.Context, _ string) (ipc.Tool, error) {
return ipc.Tool{}, ipc.ErrToolNotFound
}
func (r *recordingAPI) ListTools(_ context.Context, _ string) ([]ipc.Tool, error) {
return nil, nil
}
func (r *recordingAPI) RevertFact(_ context.Context, _ string) (int64, error) {
return 0, nil
}
func (r *recordingAPI) ListProposedRoutines(_ context.Context) ([]ipc.ProposedRoutine, error) {
return nil, nil
}
func (r *recordingAPI) AcceptProposedRoutine(_ context.Context, _ int64) error {
return nil
}
func (r *recordingAPI) DismissProposedRoutine(_ context.Context, _ int64) error {
return nil
}
func (r *recordingAPI) Chat(_ context.Context, text string) (string, error) { func (r *recordingAPI) Chat(_ context.Context, text string) (string, error) {
r.chats++ r.chats++
return "echo: " + text, nil return "echo: " + text, nil
+10 -1
View File
@@ -65,7 +65,16 @@ func Requirement(m ipc.Method) Authority {
ipc.MethodCreateReminder, ipc.MethodCreateReminder,
ipc.MethodMarkReminder, ipc.MethodMarkReminder,
ipc.MethodRecordNudge, ipc.MethodRecordNudge,
ipc.MethodResolveNudge: ipc.MethodResolveNudge,
// Task capture (Vikunja #130). Listed explicitly rather than left to
// the default so the intent is on the record: capturing a task is a
// module write, not an allowlist mutation and not a new standing reason
// for Maven to speak — nothing in the tick loop reads tasks. It stays
// at AuthRead, the same rung as CreateReminder, which is the closest
// existing analogue.
ipc.MethodCaptureTask,
ipc.MethodListTasks,
ipc.MethodSetTaskStatus:
return AuthRead return AuthRead
} }
// Unknown method ⇒ AuthRead, but ipc.dispatch returns ErrUnknownMethod // Unknown method ⇒ AuthRead, but ipc.dispatch returns ErrUnknownMethod
+221
View File
@@ -0,0 +1,221 @@
package calendar
import (
"strings"
"time"
"unicode"
)
// Ambient events — the work calendar read (Vikunja #126).
//
// The work calendar is not read by holding a work credential. A corp mail or
// calendar session living on the homelab ties the box's blast radius to the
// employer's data, which is the thing the task exists to refuse. What maven
// reads instead is the SIGNAL: an Android notification-listener on the owner's
// phone relays meeting notifications over wg/LAN, and maven turns the ones that
// clearly describe a meeting into calendar events.
//
// That makes the provenance honest. A notification is evidence about an event,
// not a reading of the calendar, so it is stored under SourceAmbient at
// AmbientConfidence — never indistinguishable from a real CalDAV read, and the
// query path hedges when it recites one.
//
// The parse is deliberately conservative. A notification with no recognisable
// clock reading produces nothing at all: maven is not a guesser-of-truth, and a
// mailbox full of noise turned into invented events is worse than a gap. Mail
// as a notification signal, not a mailbox.
// Notification — one relayed Android notification. Package is the posting app
// (for the log and for the owner to see where a wrong event came from), Title
// and Text are the notification's two text lines, Posted is when the phone
// showed it. Nothing else off the notification is kept.
type Notification struct {
Package string `json:"package"`
Title string `json:"title"`
Text string `json:"text"`
Posted time.Time `json:"posted_at"`
}
// EventFromNotification turns a notification into the event it describes, or
// reports false when it does not clearly describe one.
//
// It needs two things: a clock reading, and a summary that is not just that
// clock reading. Everything else is defaulted — the date is Posted's day (a
// meeting notification is about today or it would not be firing now), and a
// bare start time gets DefaultReminderDuration.
func EventFromNotification(n Notification) (Event, bool) {
if n.Posted.IsZero() {
return Event{}, false
}
line := strings.TrimSpace(n.Title + " " + n.Text)
start, end, ok := parseTimeRange(line)
if !ok {
return Event{}, false
}
summary := notificationSummary(n)
if summary == "" {
return Event{}, false
}
y, m, d := n.Posted.Date()
loc := n.Posted.Location()
s := time.Date(y, m, d, start.hour, start.min, 0, 0, loc)
var e time.Time
if end != nil {
e = time.Date(y, m, d, end.hour, end.min, 0, 0, loc)
// A range that ends before it starts crossed midnight.
if !e.After(s) {
e = e.AddDate(0, 0, 1)
}
} else {
e = s.Add(DefaultReminderDuration)
}
return Event{Summary: summary, Start: s, End: e}, true
}
// notificationSummary picks the text that names the meeting: the title when it
// carries words, otherwise the body. The clock reading is stripped out — it
// already lives in the times, and FactValue renders it again.
func notificationSummary(n Notification) string {
for _, cand := range []string{n.Title, n.Text} {
s := strings.TrimSpace(stripClock(cand))
s = strings.Trim(s, " \t-–—,;:@|·")
s = strings.Join(strings.Fields(s), " ")
if hasLetters(s) {
return s
}
}
return ""
}
type clock struct{ hour, min int }
// parseTimeRange finds the first clock reading in s, and a second one if the
// text spells a range. Accepted separators between hours and minutes are ":"
// and "."; between the two ends of a range, "-", "", "—" or "до".
//
// Bare hours ("в 14") are NOT accepted. Loose digits in a notification are far
// more often a count, a date or an unread badge than a meeting time, and an
// invented event is worse than no event.
func parseTimeRange(s string) (start clock, end *clock, ok bool) {
first, _, firstEnd, ok := nextClock(s, 0)
if !ok {
return clock{}, nil, false
}
sep := strings.TrimLeft(s[firstEnd:], " \t")
for _, p := range []string{"-", "", "—", "до "} {
if !strings.HasPrefix(sep, p) {
continue
}
if second, _, _, ok2 := nextClock(strings.TrimPrefix(sep, p), 0); ok2 {
return first, &second, true
}
break
}
return first, nil, true
}
// nextClock scans s from byte offset `from` for the first HH:MM (or HH.MM) and
// returns it with the byte range it occupied. Digits and separators are ASCII,
// so byte offsets are safe over Cyrillic text.
func nextClock(s string, from int) (c clock, start, end int, ok bool) {
for i := from; i < len(s); i++ {
if !isDigit(s[i]) {
continue
}
j := i
for j < len(s) && isDigit(s[j]) {
j++
}
// A run longer than two digits is a year, an id or an unread count.
if j-i > 2 {
i = j
continue
}
if j >= len(s) || (s[j] != ':' && s[j] != '.') {
i = j
continue
}
k := j + 1
for k < len(s) && isDigit(s[k]) {
k++
}
if k-(j+1) != 2 {
i = j
continue
}
// Reject a group that is a link in a longer dotted or colon chain:
// "2026.08.15" would otherwise offer "08.15" as 08:15, and a deadline
// date invented as a meeting time is exactly the wrong kind of guess.
// A trailing ":ss" is fine — that is a time with seconds.
if i > 0 && (s[i-1] == '.' || s[i-1] == ':' || isDigit(s[i-1])) {
i = k
continue
}
if k < len(s) && s[k] == '.' && k+1 < len(s) && isDigit(s[k+1]) {
i = k
continue
}
hour, min := atoi(s[i:j]), atoi(s[j+1:k])
if hour > 23 || min > 59 {
i = k
continue
}
return clock{hour, min}, i, k, true
}
return clock{}, 0, 0, false
}
func isDigit(b byte) bool { return b >= '0' && b <= '9' }
func atoi(s string) int {
n := 0
for i := 0; i < len(s); i++ {
n = n*10 + int(s[i]-'0')
}
return n
}
// stripClock removes every clock reading from a summary candidate, along with
// the preposition or separator that introduced it.
func stripClock(s string) string {
for {
_, start, end, ok := nextClock(s, 0)
if !ok {
return s
}
head := trimTrailingPreposition(strings.TrimRight(s[:start], "0123456789:.-–— \t"))
s = strings.TrimSpace(strings.TrimSpace(head) + " " + strings.TrimSpace(s[end:]))
}
}
// trimTrailingPreposition drops the word that introduced a clock reading, so
// "Встреча в 14:00" becomes "Встреча" and "с 11:30 до 12:15 Созвон" does not
// keep a dangling "с". It repeats, because a range has two of them.
func trimTrailingPreposition(s string) string {
preps := []string{"в", "с", "до", "от", "at", "from", "to"}
for again := true; again; {
again = false
s = strings.TrimRight(s, " \t")
for _, p := range preps {
if s == p {
return ""
}
if strings.HasSuffix(s, " "+p) {
s = s[:len(s)-len(p)-1]
again = true
break
}
}
}
return s
}
func hasLetters(s string) bool {
for _, r := range s {
if unicode.IsLetter(r) {
return true
}
}
return false
}
+148
View File
@@ -0,0 +1,148 @@
package calendar
import (
"testing"
"time"
)
func TestEventFromNotification(t *testing.T) {
posted := time.Date(2026, 8, 3, 9, 40, 0, 0, time.FixedZone("+04", 4*3600))
tests := []struct {
name string
title, text string
wantOK bool
wantSummary string
wantStart string // "15:04"
wantEnd string
}{
{
name: "range in the body",
title: "Планёрка",
text: "10:00-10:30",
wantOK: true,
wantSummary: "Планёрка",
wantStart: "10:00", wantEnd: "10:30",
},
{
name: "russian preposition and single time",
title: "Встреча с подрядчиком в 14:00",
wantOK: true,
wantSummary: "Встреча с подрядчиком",
wantStart: "14:00", wantEnd: "14:30",
},
{
name: "en dash range",
title: "Sprint review",
text: "Today 16:00 17:00, Meet",
wantOK: true,
wantSummary: "Sprint review",
wantStart: "16:00", wantEnd: "17:00",
},
{
name: "до as a range separator",
title: "Созвон",
text: "с 11:30 до 12:15",
wantOK: true,
wantSummary: "Созвон",
wantStart: "11:30", wantEnd: "12:15",
},
{
name: "dotted clock",
title: "Обед 13.00",
wantOK: true,
wantSummary: "Обед",
wantStart: "13:00", wantEnd: "13:30",
},
{
name: "range crossing midnight",
title: "Ночной релиз",
text: "23:30-00:30",
wantOK: true,
wantSummary: "Ночной релиз",
wantStart: "23:30", wantEnd: "00:30",
},
// The conservative half: no clock reading, no event.
{name: "no time at all", title: "3 новых письма", wantOK: false},
{name: "bare hour is not a time", title: "Планёрка в 14", wantOK: false},
{name: "unread count", title: "Входящие", text: "12 непрочитанных", wantOK: false},
{name: "a date is not a clock", title: "Отчёт", text: "срок 2026.08.15", wantOK: false},
{name: "time but nothing named", title: "10:00-10:30", wantOK: false},
{name: "impossible clock", title: "Смена 99:99", wantOK: false},
{name: "empty", wantOK: false},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
ev, ok := EventFromNotification(Notification{
Package: "com.google.android.gm",
Title: tt.title,
Text: tt.text,
Posted: posted,
})
if ok != tt.wantOK {
t.Fatalf("ok = %v, want %v (event %+v)", ok, tt.wantOK, ev)
}
if !ok {
return
}
if ev.Summary != tt.wantSummary {
t.Errorf("summary = %q, want %q", ev.Summary, tt.wantSummary)
}
if got := ev.Start.Format("15:04"); got != tt.wantStart {
t.Errorf("start = %s, want %s", got, tt.wantStart)
}
if got := ev.End.Format("15:04"); got != tt.wantEnd {
t.Errorf("end = %s, want %s", got, tt.wantEnd)
}
if !ev.End.After(ev.Start) {
t.Errorf("end %v must be after start %v", ev.End, ev.Start)
}
// The event lands on the day the phone showed it, in the phone's
// location — not shifted into UTC.
if ev.Start.Location() != posted.Location() {
t.Errorf("location = %v, want %v", ev.Start.Location(), posted.Location())
}
if y, m, d := ev.Start.Date(); y != 2026 || m != time.August || d != 3 {
t.Errorf("date = %d-%02d-%02d, want 2026-08-03", y, m, d)
}
})
}
}
func TestEventFromNotificationNeedsPostedAt(t *testing.T) {
if _, ok := EventFromNotification(Notification{Title: "Планёрка 10:00"}); ok {
t.Error("a notification with no posted_at has no date to sit on")
}
}
// An ambient event must never be indistinguishable from a calendar read.
func TestAmbientEventsAreStoredAtReducedConfidence(t *testing.T) {
ev, ok := EventFromNotification(Notification{
Title: "Планёрка 10:00-10:30",
Posted: time.Date(2026, 8, 3, 9, 0, 0, 0, time.UTC),
})
if !ok {
t.Fatal("expected an event")
}
if FactKey(ev) == "" || FactValue(ev) == "" {
t.Fatal("ambient events must use the shared fact encoding")
}
if AmbientConfidence >= 1.0 {
t.Fatal("ambient confidence must be below a calendar read's")
}
}
func TestStripClock(t *testing.T) {
tests := []struct{ in, want string }{
{"Встреча в 14:00", "Встреча"},
{"Планёрка 10:00-10:30", "Планёрка"},
{"с 11:30 до 12:15 Созвон", "Созвон"},
{"Ничего", "Ничего"},
}
for _, tt := range tests {
if got := stripClock(tt.in); got != tt.want {
t.Errorf("stripClock(%q) = %q, want %q", tt.in, got, tt.want)
}
}
}
+129
View File
@@ -0,0 +1,129 @@
// Package calendar is the one calendar data model the rest of maven shares:
// an Event, the iCal text it is parsed from and rendered to, and the fact
// encoding that puts it in the store.
//
// It exists because three separate features read or write the same events and
// must agree on their shape: the CalDAV read side (cmd/mavcaldav, Vikunja
// #126/#127), the write-only render target that publishes maven's own
// reminders as a calendar (#127), and the day plan that recites them (#128).
// Before this package the parse lived inline in cmd/mavcaldav and the fact key
// format was a Sprintf in two places.
//
// The package is pure: no HTTP, no store, no clock of its own. Callers own the
// impurity, the way internal/morning and internal/loop do.
package calendar
import (
"fmt"
"sort"
"strings"
"time"
)
// Fact sources. A calendar event reaches the store as a
// `facts (kind=env, key=calendar_event_..., source=<one of these>)` row, and
// the source is the whole provenance story:
//
// - SourcePersonal — maven's own Radicale, read AND rendered to. Canonical
// state stays in sqlite; the calendar is a render target (#127).
// - SourceWork — a work calendar, read-only by definition (#126). Nothing in
// maven ever writes to it: no code path pairs this source with a PUT.
// - SourceAmbient — inferred from an Android notification-listener relay
// rather than read from a server (#126). Confidence is below 1.0 because a
// notification is a signal about an event, not the event.
const (
SourcePersonal = "poll:caldav"
SourceWork = "poll:caldav:work"
SourceAmbient = "ambient:notif"
)
// AmbientConfidence — the confidence a notification-derived event is stored
// with. A parsed notification line is evidence, not a reading of the calendar,
// so it must never be indistinguishable from one (#126).
const AmbientConfidence = 0.6
// Sources lists every source a calendar event may legitimately carry, for the
// store query that reads the calendar back out. Ordered from most to least
// trusted.
func Sources() []string {
return []string{SourcePersonal, SourceWork, SourceAmbient}
}
// ReadOnlySource reports whether events from this source may never be written
// back. The work calendar is read-only by definition — see #126: maven holding
// a credential that can write to an employer's calendar is the thing the task
// exists to avoid.
func ReadOnlySource(source string) bool {
return source == SourceWork || source == SourceAmbient
}
// Event — one calendar entry. UID is the iCal UID when the event was parsed
// from a server and the identity maven renders under when it publishes one;
// Start/End are instants. All-day events are not modelled: the busy gate and
// the day plan both need a time of day, and an all-day marker answers neither.
type Event struct {
UID string
Summary string
Start time.Time
End time.Time
}
// FactKey is the store key for an event: one key per day per summary, stable
// across polls so re-reading an unchanged calendar rewrites nothing.
//
// The date prefix is load-bearing — store.CalendarEvents selects a day range
// by key prefix, not by a timestamp column.
func FactKey(e Event) string {
return fmt.Sprintf("calendar_event_%s_%s", e.Start.Format("20060102"), safeKey(e.Summary))
}
// FactValue is the human-readable rendering stored as the fact value, and the
// string the day plan and the query path read back.
func FactValue(e Event) string {
return fmt.Sprintf("%s @ %s-%s", e.Summary, e.Start.Format("15:04"), e.End.Format("15:04"))
}
// KeyPrefixForDay is the fact-key prefix covering one calendar day. The store
// range-scans between two of these.
func KeyPrefixForDay(day time.Time) string {
return fmt.Sprintf("calendar_event_%s", day.Format("20060102"))
}
// Busy reports whether any event covers the instant now — the read the loop
// gate uses to suppress nudges during a meeting.
func Busy(events []Event, now time.Time) bool {
for _, e := range events {
if !now.Before(e.Start) && now.Before(e.End) {
return true
}
}
return false
}
// Overlapping returns the events intersecting [from, to), sorted by start.
func Overlapping(events []Event, from, to time.Time) []Event {
var out []Event
for _, e := range events {
if e.End.After(from) && e.Start.Before(to) {
out = append(out, e)
}
}
sort.Slice(out, func(i, j int) bool { return out[i].Start.Before(out[j].Start) })
return out
}
// safeKey makes a summary safe to use inside a fact key (ASCII alphanumerics
// and dashes). Non-Latin summaries collapse to their punctuation, which is why
// the day prefix carries the identity and this only disambiguates within a day.
func safeKey(s string) string {
var b strings.Builder
for _, r := range s {
switch {
case (r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z') || (r >= '0' && r <= '9') || r == '-':
b.WriteRune(r)
case r == ' ' || r == '_':
b.WriteRune('-')
}
}
return b.String()
}
+202
View File
@@ -0,0 +1,202 @@
package calendar
import (
"strings"
"testing"
"time"
)
func TestParseICalDayKeepsOnlyToday(t *testing.T) {
now := time.Date(2026, 7, 3, 12, 0, 0, 0, time.UTC)
body := []byte(`BEGIN:VCALENDAR
BEGIN:VEVENT
UID:a@example
DTSTART:20260703T090000Z
DTEND:20260703T100000Z
SUMMARY:Morning standup
END:VEVENT
BEGIN:VEVENT
DTSTART:20260703T140000Z
DTEND:20260703T150000Z
SUMMARY:Team sync
END:VEVENT
BEGIN:VEVENT
DTSTART:20260702T140000Z
DTEND:20260702T150000Z
SUMMARY:Yesterday retro
END:VEVENT
BEGIN:VEVENT
DTSTART:20260704T090000Z
DTEND:20260704T100000Z
SUMMARY:Tomorrow standup
END:VEVENT
BEGIN:VEVENT
DTSTART;VALUE=DATE:20260704
DTEND;VALUE=DATE:20260705
SUMMARY:All-day event
END:VEVENT
END:VCALENDAR`)
events := ParseICalDay(body, now)
if len(events) != 2 {
t.Fatalf("got %d events, want 2 (today only, no all-day/past/future)", len(events))
}
if events[0].Summary != "Morning standup" || events[0].UID != "a@example" {
t.Errorf("events[0] = %+v", events[0])
}
if !events[0].Start.Equal(time.Date(2026, 7, 3, 9, 0, 0, 0, time.UTC)) {
t.Errorf("events[0].Start = %v", events[0].Start)
}
if !events[0].End.Equal(time.Date(2026, 7, 3, 10, 0, 0, 0, time.UTC)) {
t.Errorf("events[0].End = %v", events[0].End)
}
if events[1].Summary != "Team sync" {
t.Errorf("events[1].Summary = %q", events[1].Summary)
}
}
// Regression: "today" is the owner's day, in the owner's location. Taking the
// day number off a local clock but building the boundaries in UTC made the
// evening fall outside the window on any box east of Greenwich.
func TestParseICalDayUsesOwnersDay(t *testing.T) {
plus4 := time.FixedZone("+04", 4*60*60)
// 01:00 on Aug 1 local is 21:00 on Jul 31 UTC.
now := time.Date(2026, 8, 1, 1, 0, 0, 0, plus4)
body := []byte("BEGIN:VCALENDAR\nBEGIN:VEVENT\n" +
"DTSTART:20260731T195406Z\nDTEND:20260731T235406Z\nSUMMARY:Current meeting\n" +
"END:VEVENT\nEND:VCALENDAR")
events := ParseICalDay(body, now)
if len(events) != 1 {
t.Fatalf("got %d events, want the in-progress one", len(events))
}
if !Busy(events, now.UTC()) {
t.Error("an event in progress right now must read as busy")
}
}
func TestParseVEVENT(t *testing.T) {
block := "DTSTART;TZID=Europe/Moscow:20260703T130000\nDTEND:20260703T140000Z\nSUMMARY:Stand up meeting"
e, ok := parseVEVENT(block)
if !ok {
t.Fatal("expected a parsed event")
}
if !e.Start.Equal(time.Date(2026, 7, 3, 13, 0, 0, 0, time.UTC)) {
t.Errorf("start = %v", e.Start)
}
if !e.End.Equal(time.Date(2026, 7, 3, 14, 0, 0, 0, time.UTC)) {
t.Errorf("end = %v", e.End)
}
if e.Summary != "Stand up meeting" {
t.Errorf("summary = %q", e.Summary)
}
allDay := "DTSTART;VALUE=DATE:20260703\nDTEND;VALUE=DATE:20260704\nSUMMARY:All-day"
if _, ok := parseVEVENT(allDay); ok {
t.Error("all-day event should be rejected")
}
}
func TestParseDT(t *testing.T) {
tests := []struct {
name string
line string
want time.Time
wantOK bool
}{
{"UTC", "DTEND:20260703T100000Z", time.Date(2026, 7, 3, 10, 0, 0, 0, time.UTC), true},
{"local", "DTSTART;TZID=Europe/Moscow:20260703T130000", time.Date(2026, 7, 3, 13, 0, 0, 0, time.UTC), true},
{"all-day", "DTSTART;VALUE=DATE:20260703", time.Time{}, false},
{"garbage", "DTSTART:garbage", time.Time{}, false},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
got, ok := parseDT(tt.line)
if ok != tt.wantOK {
t.Errorf("ok = %v, want %v", ok, tt.wantOK)
}
if !got.Equal(tt.want) {
t.Errorf("got %v, want %v", got, tt.want)
}
})
}
}
func TestSafeKey(t *testing.T) {
tests := []struct{ in, want string }{
{"Stand up meeting", "Stand-up-meeting"},
{"Hello_World", "Hello-World"},
{"special@#$chars!!", "specialchars"},
{"ALL_CAPS_123", "ALL-CAPS-123"},
}
for _, tt := range tests {
if got := safeKey(tt.in); got != tt.want {
t.Errorf("safeKey(%q) = %q, want %q", tt.in, got, tt.want)
}
}
}
func TestFactKeyAndValue(t *testing.T) {
e := Event{
Summary: "Team sync",
Start: time.Date(2026, 7, 3, 14, 0, 0, 0, time.UTC),
End: time.Date(2026, 7, 3, 15, 0, 0, 0, time.UTC),
}
if got, want := FactKey(e), "calendar_event_20260703_Team-sync"; got != want {
t.Errorf("FactKey = %q, want %q", got, want)
}
if got, want := FactValue(e), "Team sync @ 14:00-15:00"; got != want {
t.Errorf("FactValue = %q, want %q", got, want)
}
if got, want := KeyPrefixForDay(e.Start), "calendar_event_20260703"; got != want {
t.Errorf("KeyPrefixForDay = %q, want %q", got, want)
}
if !strings.HasPrefix(FactKey(e), KeyPrefixForDay(e.Start)) {
t.Error("FactKey must start with the day prefix the store range-scans on")
}
}
func TestBusyAndOverlapping(t *testing.T) {
base := time.Date(2026, 7, 3, 0, 0, 0, 0, time.UTC)
events := []Event{
{Summary: "late", Start: base.Add(15 * time.Hour), End: base.Add(16 * time.Hour)},
{Summary: "early", Start: base.Add(9 * time.Hour), End: base.Add(10 * time.Hour)},
}
if !Busy(events, base.Add(9*time.Hour+30*time.Minute)) {
t.Error("should be busy inside the early event")
}
if Busy(events, base.Add(12*time.Hour)) {
t.Error("should be free at noon")
}
// Half-open: the end instant is free.
if Busy(events, base.Add(10*time.Hour)) {
t.Error("the end instant should not count as busy")
}
got := Overlapping(events, base.Add(8*time.Hour), base.Add(11*time.Hour))
if len(got) != 1 || got[0].Summary != "early" {
t.Fatalf("Overlapping = %+v", got)
}
all := Overlapping(events, base, base.AddDate(0, 0, 1))
if len(all) != 2 || all[0].Summary != "early" {
t.Fatalf("Overlapping must sort by start: %+v", all)
}
}
func TestSourceTrust(t *testing.T) {
if ReadOnlySource(SourcePersonal) {
t.Error("the personal calendar is the one maven may render to")
}
if !ReadOnlySource(SourceWork) {
t.Error("the work calendar must be read-only")
}
if !ReadOnlySource(SourceAmbient) {
t.Error("an ambient notification is not a writable calendar")
}
if AmbientConfidence >= 1.0 {
t.Error("ambient events must be less trusted than a calendar read")
}
if len(Sources()) != 3 {
t.Errorf("Sources() = %v", Sources())
}
}
+145
View File
@@ -0,0 +1,145 @@
package calendar
import (
"fmt"
"strings"
"time"
)
// ParseICal scans iCal text for VEVENT components and returns the events
// overlapping [from, to). All-day events are skipped: parseDT reports no time
// for a VALUE=DATE value, and an event with no clock reading answers neither
// the busy gate nor the day plan.
func ParseICal(body []byte, from, to time.Time) []Event {
var events []Event
text := string(body)
for {
i := strings.Index(text, "BEGIN:VEVENT")
if i < 0 {
break
}
text = text[i+len("BEGIN:VEVENT"):]
j := strings.Index(text, "END:VEVENT")
if j < 0 {
break
}
block := text[:j]
text = text[j+len("END:VEVENT"):]
e, ok := parseVEVENT(block)
if !ok {
continue
}
if e.End.After(from) && e.Start.Before(to) {
events = append(events, e)
}
}
return events
}
// ParseICalDay is ParseICal over the calendar day containing now, in now's own
// location — the window cmd/mavcaldav polls.
//
// The location matters. The old inline version took the day number off a local
// clock reading but built the boundaries in UTC, so east of Greenwich the
// window was shifted by the offset and part of the evening fell outside
// "today": on a +04 box after 20:00 UTC the poller saw an empty calendar. The
// owner's day is the day the day plan and the busy gate mean.
func ParseICalDay(body []byte, now time.Time) []Event {
y, m, d := now.Date()
start := time.Date(y, m, d, 0, 0, 0, 0, now.Location())
return ParseICal(body, start, start.AddDate(0, 0, 1))
}
// parseVEVENT extracts UID, start, end and summary from a VEVENT block.
// Reports false for all-day events and parse failures.
func parseVEVENT(block string) (Event, bool) {
var e Event
for _, line := range strings.Split(block, "\n") {
line = strings.TrimSpace(line)
switch {
case strings.HasPrefix(line, "DTSTART"):
if t, ok := parseDT(line); ok {
e.Start = t
}
case strings.HasPrefix(line, "DTEND"):
if t, ok := parseDT(line); ok {
e.End = t
}
case strings.HasPrefix(line, "SUMMARY"):
e.Summary = afterColon(line)
case strings.HasPrefix(line, "UID"):
e.UID = afterColon(line)
}
}
if e.Start.IsZero() || e.End.IsZero() {
return Event{}, false
}
return e, true
}
func afterColon(line string) string {
if i := strings.Index(line, ":"); i >= 0 {
return strings.TrimSpace(line[i+1:])
}
return ""
}
// parseDT parses a DTSTART/DTEND value:
//
// - UTC: DTEND:20260703T100000Z
// - Local: DTSTART;TZID=Europe/Moscow:20260703T130000
// - All-day: DTSTART;VALUE=DATE:20260703 (rejected)
//
// A local time is read as UTC, the behaviour cmd/mavcaldav has always had: the
// CalDAV server and the poller run in the same timezone, and the busy gate only
// needs busy/not-busy to be right.
func parseDT(line string) (time.Time, bool) {
if strings.Contains(line, "VALUE=DATE:") {
return time.Time{}, false
}
i := strings.LastIndex(line, ":")
if i < 0 {
return time.Time{}, false
}
val := strings.TrimSuffix(strings.TrimSpace(line[i+1:]), "Z")
t, err := time.Parse("20060102T150405", val)
if err != nil {
return time.Time{}, false
}
return t.UTC(), true
}
// RenderICal wraps events in a VCALENDAR body suitable for PUTting to a CalDAV
// collection. One event per file is the CalDAV convention, so callers normally
// pass a single event.
//
// This is the write half of #127 and it only ever renders: the canonical state
// is sqlite, the calendar is a view of it. Nothing reads a rendered file back.
func RenderICal(events []Event) string {
var b strings.Builder
b.WriteString("BEGIN:VCALENDAR\r\nVERSION:2.0\r\nPRODID:-//maven//local calendar//RU\r\n")
for _, e := range events {
b.WriteString("BEGIN:VEVENT\r\n")
fmt.Fprintf(&b, "UID:%s\r\n", escapeText(e.UID))
fmt.Fprintf(&b, "DTSTAMP:%s\r\n", e.Start.UTC().Format("20060102T150405Z"))
fmt.Fprintf(&b, "DTSTART:%s\r\n", e.Start.UTC().Format("20060102T150405Z"))
fmt.Fprintf(&b, "DTEND:%s\r\n", e.End.UTC().Format("20060102T150405Z"))
fmt.Fprintf(&b, "SUMMARY:%s\r\n", escapeText(e.Summary))
b.WriteString("END:VEVENT\r\n")
}
b.WriteString("END:VCALENDAR\r\n")
return b.String()
}
// escapeText applies RFC 5545 TEXT escaping and strips the line breaks that
// would otherwise let a reminder payload inject iCal properties.
func escapeText(s string) string {
s = strings.ReplaceAll(s, "\\", "\\\\")
s = strings.ReplaceAll(s, ";", "\\;")
s = strings.ReplaceAll(s, ",", "\\,")
s = strings.ReplaceAll(s, "\r\n", "\\n")
s = strings.ReplaceAll(s, "\n", "\\n")
s = strings.ReplaceAll(s, "\r", "\\n")
return s
}
+69
View File
@@ -0,0 +1,69 @@
package calendar
import (
"strings"
"testing"
"time"
)
func TestRenderICalRoundTrips(t *testing.T) {
e := ReminderEvent(7, time.Date(2026, 8, 1, 18, 30, 0, 0, time.UTC), "позвонить маме", 0)
if e.UID != "maven-reminder-7" {
t.Errorf("UID = %q", e.UID)
}
if got := e.End.Sub(e.Start); got != DefaultReminderDuration {
t.Errorf("duration = %v, want %v", got, DefaultReminderDuration)
}
if got, want := ReminderPath(7), "maven-reminder-7.ics"; got != want {
t.Errorf("ReminderPath = %q, want %q", got, want)
}
body := RenderICal([]Event{e})
if !strings.HasPrefix(body, "BEGIN:VCALENDAR\r\n") || !strings.HasSuffix(body, "END:VCALENDAR\r\n") {
t.Fatalf("not a VCALENDAR body:\n%s", body)
}
back := ParseICal([]byte(body), e.Start.Add(-time.Hour), e.Start.Add(time.Hour))
if len(back) != 1 {
t.Fatalf("got %d events back, want 1:\n%s", len(back), body)
}
if back[0].UID != e.UID || back[0].Summary != e.Summary {
t.Errorf("round trip lost identity: %+v", back[0])
}
if !back[0].Start.Equal(e.Start) || !back[0].End.Equal(e.End) {
t.Errorf("round trip lost times: %+v", back[0])
}
}
func TestRenderICalIsDeterministic(t *testing.T) {
e := ReminderEvent(1, time.Date(2026, 8, 1, 9, 0, 0, 0, time.UTC), "выпить воды", 0)
if RenderICal([]Event{e}) != RenderICal([]Event{e}) {
t.Error("the same reminder must render byte-identically, or every poll re-PUTs it")
}
}
// A reminder payload is owner-supplied text. It must not be able to close the
// VEVENT and inject properties of its own.
func TestRenderICalEscapesInjection(t *testing.T) {
e := ReminderEvent(2, time.Date(2026, 8, 1, 9, 0, 0, 0, time.UTC),
"обед\r\nEND:VEVENT\r\nBEGIN:VEVENT\r\nSUMMARY:injected", 0)
body := RenderICal([]Event{e})
// Count line-initial occurrences: the escaped text still contains the
// characters "BEGIN:VEVENT", it just can no longer start a line.
if n := strings.Count(body, "\r\nBEGIN:VEVENT\r\n"); n != 1 {
t.Fatalf("payload injected a second VEVENT (%d):\n%s", n, body)
}
if n := strings.Count(body, "\r\nEND:VEVENT\r\n"); n != 1 {
t.Fatalf("payload closed the VEVENT early (%d):\n%s", n, body)
}
if !strings.Contains(body, `SUMMARY:обед\nEND:VEVENT`) {
t.Errorf("newlines should be escaped, not dropped:\n%s", body)
}
}
func TestReminderEventEmptyPayload(t *testing.T) {
e := ReminderEvent(3, time.Date(2026, 8, 1, 9, 0, 0, 0, time.UTC), " ", 0)
if e.Summary != "напоминание" {
t.Errorf("Summary = %q, want the neutral RU fallback", e.Summary)
}
}
+45
View File
@@ -0,0 +1,45 @@
package calendar
import (
"fmt"
"strings"
"time"
)
// ReminderUIDPrefix namespaces every event maven publishes. Two reasons it is a
// fixed prefix and not a random UUID: the render is idempotent (the same
// reminder always lands on the same UID, so re-rendering overwrites instead of
// duplicating), and everything maven owns in the target collection is
// identifiable at a glance — she never touches a file she did not create.
const ReminderUIDPrefix = "maven-reminder-"
// DefaultReminderDuration — how long a rendered reminder occupies. A reminder
// is an instant, a calendar entry is a span, so one has to be invented; 30
// minutes reads as a block in a calendar app without swallowing the afternoon.
const DefaultReminderDuration = 30 * time.Minute
// ReminderEvent maps a reminder to the event that represents it. id and fire
// come from the store; payload is the RU text as the owner said it, rendered
// verbatim as the summary — the calendar is a view of sqlite, not a place to
// rephrase.
func ReminderEvent(id int64, fire time.Time, payload string, dur time.Duration) Event {
if dur <= 0 {
dur = DefaultReminderDuration
}
summary := strings.TrimSpace(payload)
if summary == "" {
summary = "напоминание"
}
return Event{
UID: fmt.Sprintf("%s%d", ReminderUIDPrefix, id),
Summary: summary,
Start: fire,
End: fire.Add(dur),
}
}
// ReminderPath is the collection-relative filename for a rendered reminder.
// One event per resource, per the CalDAV convention.
func ReminderPath(id int64) string {
return fmt.Sprintf("%s%d.ics", ReminderUIDPrefix, id)
}
+100
View File
@@ -140,6 +140,16 @@ type Config struct {
// item. See internal/morning for the evaluation engine. Empty ⇒ disabled. // item. See internal/morning for the evaluation engine. Empty ⇒ disabled.
MorningRoutines []MorningRoutineConfig `json:"morning_routines,omitempty"` MorningRoutines []MorningRoutineConfig `json:"morning_routines,omitempty"`
// PatternProposals — whether a routine the digestion tick inferred on its
// own may be announced, and how often. nil / absent ⇒ silent detection
// only: proposals are written for /routines and never announced. See
// PatternProposalConfig.
PatternProposals *PatternProposalConfig `json:"pattern_proposals,omitempty"`
// MemoryEval — background memory evaluation (internal/memeval). nil /
// absent ⇒ no evaluation loop at all. See MemoryEvalConfig.
MemoryEval *MemoryEvalConfig `json:"memory_eval,omitempty"`
// Praxis — the ecosystem attention-state service. When configured, maven // Praxis — the ecosystem attention-state service. When configured, maven
// calls the Praxis HTTP tools API for attention listing and item lifecycle. // calls the Praxis HTTP tools API for attention listing and item lifecycle.
// Maven never touches Praxis's database directly (ecosystem invariant: no // Maven never touches Praxis's database directly (ecosystem invariant: no
@@ -302,6 +312,13 @@ type VoiceConfig struct {
// Russian self-reference). Example: "Be formal and answer in English only." // Russian self-reference). Example: "Be formal and answer in English only."
Persona string `json:"persona,omitempty"` Persona string `json:"persona,omitempty"`
// OwnerName / City — optional facts about the owner, added to the shared
// context block (internal/persona). Empty is fine: the block still states
// who he is grammatically (a man, addressed as "ты") and the current time.
// Nothing about correct behaviour may depend on these being filled in.
OwnerName string `json:"owner_name,omitempty"`
City string `json:"city,omitempty"`
// Weather — the weather provider config. nil ⇒ the daemon wires // Weather — the weather provider config. nil ⇒ the daemon wires
// the stub provider (returns ErrNotConfigured — "погода не настроена"). // the stub provider (returns ErrNotConfigured — "погода не настроена").
// Set provider to "open-meteo" to use the keyless Open-Meteo API. // Set provider to "open-meteo" to use the keyless Open-Meteo API.
@@ -345,6 +362,62 @@ type DigestConfig struct {
SeverityCeiling int `json:"severity_ceiling,omitempty"` // max sev batched SeverityCeiling int `json:"severity_ceiling,omitempty"` // max sev batched
} }
// PatternProposalConfig — announcement policy for routines the digestion tick
// inferred by itself (Vikunja #247, #43).
//
// Detection is always on and always silent by default: the tick writes a
// proposed_routines row and the /routines page shows it. Notify is what turns
// "she noticed" into "she said something", and it is OFF unless configured —
// Maven is not a nag and not autonomous, so a behaviour that speaks without
// being asked has to be switched on deliberately, like weather and telegram.
//
// When Notify is on, the announcement is still heavily restrained:
// - at most one proposal per tick, however many were detected;
// - at most one per Cooldown across all pairs (not per pair), so a batch of
// freshly-detected patterns cannot turn into a queue of interruptions;
// - through the ordinary care-class gate (quiet hours / away / snooze), at
// sev1 — the lowest severity there is. A proposal is the least urgent
// thing Maven can say.
//
// A pair is only ever announced once, because it is only ever proposed once:
// proposed_routines is UNIQUE(action, object) and the row survives dismissal.
type PatternProposalConfig struct {
// Notify — announce newly inferred routines. Default false.
Notify bool `json:"notify,omitempty"`
// Cooldown — minimum spacing between two proposal announcements. 0 ⇒
// DefaultProposalCooldown (24h).
Cooldown Duration `json:"cooldown,omitempty"`
}
// AnnounceProposals reports whether inferred routines may be announced. Safe
// on a nil receiver — an absent config block means silent detection.
func (p *PatternProposalConfig) AnnounceProposals() bool {
return p != nil && p.Notify
}
// MemoryEvalConfig — the background memory-evaluation loop (Vikunja #248).
// Absent ⇒ off, like every other capability that costs something the owner did
// not ask for. Each evaluation is a full LLM round-trip on the one resident
// model, which is the same model answering him; running it hourly by default
// would put a multi-second stall in front of an occasional voice turn for a
// feature he may not want.
//
// The loop only ever writes notes (source infer:memory-eval, visible on
// /dash). It cannot speak — see internal/memeval.
type MemoryEvalConfig struct {
// Interval — how often to evaluate. 0 ⇒ DefaultMemoryEvalInterval.
Interval Duration `json:"interval,omitempty"`
// MaxItems — recent facts / notes / nudges fed into one evaluation.
// 0 ⇒ memeval.DefaultMaxItems.
MaxItems int `json:"max_items,omitempty"`
// MinConfidence — observations the model scores below this are dropped.
// 0 ⇒ memeval.DefaultMinConfidence.
MinConfidence float64 `json:"min_confidence,omitempty"`
}
// PhraserConfig — the LLM-backed phraser seam. The daemon spawns llama-server // PhraserConfig — the LLM-backed phraser seam. The daemon spawns llama-server
// as a managed subprocess and sends chat-completion requests to phrase nudge // as a managed subprocess and sends chat-completion requests to phrase nudge
// and reminder messages. nil ⇒ the template-based Stub is used instead. // and reminder messages. nil ⇒ the template-based Stub is used instead.
@@ -362,6 +435,12 @@ type PhraserConfig struct {
NGpuLayers int `json:"n_gpu_layers,omitempty"` NGpuLayers int `json:"n_gpu_layers,omitempty"`
NCtx int `json:"n_ctx,omitempty"` NCtx int `json:"n_ctx,omitempty"`
Timeout Duration `json:"timeout,omitempty"` Timeout Duration `json:"timeout,omitempty"`
// LLMNudges — let the model word nudges again. Off by default: nudges are
// worded from hand-written Russian templates now (the model broke the
// persona and invented units). Chat, query and reminder phrasing always go
// through the model regardless. See phraser.Config.LLMNudges.
LLMNudges bool `json:"llm_nudges,omitempty"`
} }
// EmbedderConfig — paths for the ONNX multilingual embedder. The daemon // EmbedderConfig — paths for the ONNX multilingual embedder. The daemon
@@ -435,6 +514,15 @@ const (
DefaultLLMRouter = true DefaultLLMRouter = true
DefaultFactEnrichmentInterval = 30 * time.Second DefaultFactEnrichmentInterval = 30 * time.Second
// DefaultProposalCooldown — one inferred-routine announcement per day at
// most. A proposal is never urgent; if two patterns surface in the same
// hour, the second one waits, and the /routines page has it either way.
DefaultProposalCooldown = 24 * time.Hour
// DefaultMemoryEvalInterval — the plan's cadence (1h) for the memory
// evaluation loop, applied only when the block is present at all.
DefaultMemoryEvalInterval = time.Hour
) )
// Load reads the JSON config at path and applies defaults. A missing file is // Load reads the JSON config at path and applies defaults. A missing file is
@@ -507,6 +595,18 @@ func (c *Config) applyDefaults() {
c.Digest.SeverityCeiling = 2 c.Digest.SeverityCeiling = 2
} }
// Absent block stays nil (⇒ silent detection). Present-but-partial gets the
// cooldown default, so `{"notify": true}` is enough to switch it on.
if c.PatternProposals != nil && c.PatternProposals.Cooldown <= 0 {
c.PatternProposals.Cooldown = Duration(DefaultProposalCooldown)
}
// Same rule: absent stays nil (⇒ no evaluation loop), present gets defaults
// so `{}` is a valid "on with the plan's cadence".
if c.MemoryEval != nil && c.MemoryEval.Interval <= 0 {
c.MemoryEval.Interval = Duration(DefaultMemoryEvalInterval)
}
if c.Voice != nil { if c.Voice != nil {
if c.Voice.RouterThreshold <= 0 { if c.Voice.RouterThreshold <= 0 {
c.Voice.RouterThreshold = DefaultRouterThreshold c.Voice.RouterThreshold = DefaultRouterThreshold
+71
View File
@@ -35,6 +35,27 @@ func TestLoadDefaults(t *testing.T) {
} }
} }
// Nudges come from templates unless the config says otherwise.
func TestPhraserLLMNudgesDefaultsOff(t *testing.T) {
p := writeConfig(t, `{"phraser":{"model_path":"/tmp/m.gguf"}}`)
c, err := Load(p)
if err != nil {
t.Fatalf("Load: %v", err)
}
if c.Phraser.LLMNudges {
t.Error("llm_nudges defaults on; templates must be the default")
}
p = writeConfig(t, `{"phraser":{"model_path":"/tmp/m.gguf","llm_nudges":true}}`)
c, err = Load(p)
if err != nil {
t.Fatalf("Load: %v", err)
}
if !c.Phraser.LLMNudges {
t.Error("llm_nudges:true did not parse")
}
}
func TestLoadDurationsParse(t *testing.T) { func TestLoadDurationsParse(t *testing.T) {
p := writeConfig(t, `{"tick_interval":"90s","repeat_interval":"10m"}`) p := writeConfig(t, `{"tick_interval":"90s","repeat_interval":"10m"}`)
c, err := Load(p) c, err := Load(p)
@@ -222,3 +243,53 @@ func TestDurationRoundTrip(t *testing.T) {
t.Errorf("round-trip = %v, want %v", d2, d) t.Errorf("round-trip = %v, want %v", d2, d)
} }
} }
// Both new opt-in capabilities follow the same rule: absent block ⇒ nil ⇒ the
// behaviour does not exist. Presence is the enable act, so a bare `{}` block is
// valid and gets the defaults filled in.
func TestOptInBlocksAbsentStayNil(t *testing.T) {
c, err := Load(writeConfig(t, `{}`))
if err != nil {
t.Fatalf("Load: %v", err)
}
if c.PatternProposals != nil {
t.Errorf("pattern_proposals absent but got %+v", c.PatternProposals)
}
if c.PatternProposals.AnnounceProposals() {
t.Error("AnnounceProposals() true with no config block")
}
if c.MemoryEval != nil {
t.Errorf("memory_eval absent but got %+v", c.MemoryEval)
}
}
func TestOptInBlocksGetDefaultsWhenPresent(t *testing.T) {
c, err := Load(writeConfig(t, `{"pattern_proposals":{"notify":true},"memory_eval":{}}`))
if err != nil {
t.Fatalf("Load: %v", err)
}
if !c.PatternProposals.AnnounceProposals() {
t.Error("notify:true did not enable announcements")
}
if time.Duration(c.PatternProposals.Cooldown) != DefaultProposalCooldown {
t.Errorf("proposal cooldown = %v, want %v", c.PatternProposals.Cooldown, DefaultProposalCooldown)
}
if time.Duration(c.MemoryEval.Interval) != DefaultMemoryEvalInterval {
t.Errorf("memory eval interval = %v, want %v", c.MemoryEval.Interval, DefaultMemoryEvalInterval)
}
}
// Notify is off even when the block exists — the block is where you tune it,
// notify:true is the act that lets her speak.
func TestPatternProposalNotifyDefaultsOff(t *testing.T) {
c, err := Load(writeConfig(t, `{"pattern_proposals":{"cooldown":"6h"}}`))
if err != nil {
t.Fatalf("Load: %v", err)
}
if c.PatternProposals.AnnounceProposals() {
t.Error("notify defaulted to on")
}
if time.Duration(c.PatternProposals.Cooldown) != 6*time.Hour {
t.Errorf("cooldown = %v, want 6h", c.PatternProposals.Cooldown)
}
}
+21 -3
View File
@@ -137,9 +137,10 @@ func NewDispatcher(cfg Config) *Dispatcher {
// picks for (severity, presence), sends via the matching sink, and records // picks for (severity, presence), sends via the matching sink, and records
// one nudge row per successful send. returns the dispatches (one per channel). // one nudge row per successful send. returns the dispatches (one per channel).
// //
// a Drop channel = no send, no record (the nudge was suppressed by routing, // a Drop channel = no send (the nudge was suppressed by routing, not by a
// not by a failure — "a missed water nudge is noise"). a nil sink = channel // failure — "a missed water nudge is noise"), but it does leave a 'dropped'
// not wired, skip silently. a send error stops the dispatch and returns what // outbox row so the suppression is visible. a nil sink = channel not wired,
// skip silently. a send error stops the dispatch and returns what
// got through — the daemon decides whether to retry. // got through — the daemon decides whether to retry.
func (d *Dispatcher) DispatchNudge(ctx context.Context, pn PhrasedNudge, now time.Time) ([]Dispatch, error) { func (d *Dispatcher) DispatchNudge(ctx context.Context, pn PhrasedNudge, now time.Time) ([]Dispatch, error) {
c := pn.Candidate c := pn.Candidate
@@ -148,6 +149,16 @@ func (d *Dispatcher) DispatchNudge(ctx context.Context, pn PhrasedNudge, now tim
for i := 0; i < len(channels); i++ { for i := 0; i < len(channels); i++ {
ch := channels[i] ch := channels[i]
if ch == ChannelDrop { if ch == ChannelDrop {
// the routing table suppressed this nudge on purpose (a care nudge
// while you're away is noise). that stays — but it must not be
// invisible, or "she dropped it" and "the rule never fired" look
// the same afterwards. no nudges row: that table feeds the
// ignored_rate signal, and a nudge nobody could see must not
// count as ignored.
id := d.beginOutbox(ctx, "nudge", c.Rule.Name, 0, ch, pn.Summary, now)
d.completeOutbox(ctx, id, store.DeliveryDropped, now)
log.Printf("dispatcher: dropped %s (sev%d, presence=%s) — routing table suppressed it",
c.Rule.Name, c.Severity, c.State.Presence)
continue continue
} }
s := Sendable{ s := Sendable{
@@ -396,6 +407,13 @@ func messageForChannel(s Sendable) string {
if !isAway(s.Channel) { if !isAway(s.Channel) {
return s.Body return s.Body
} }
return AwayMessage(s)
}
// AwayMessage — the only text an off-box channel may ever carry. Exported so
// the away sinks share this one rule instead of each inventing a fallback: the
// summary if we have one, otherwise a fixed generic line. Never the body.
func AwayMessage(s Sendable) string {
if s.Summary != "" { if s.Summary != "" {
return s.Summary return s.Summary
} }
+6 -4
View File
@@ -37,10 +37,12 @@ func TestVoiceNoSessionFallthroughLeavesOutboxTrail(t *testing.T) {
[]string{"voice", "ntfy"}, []string{store.DeliveryFailed, store.DeliverySent}}, []string{"voice", "ntfy"}, []string{store.DeliveryFailed, store.DeliverySent}},
{"sev4 falls through to telegram", loop.Sev4, {"sev4 falls through to telegram", loop.Sev4,
[]string{"voice", "telegram"}, []string{store.DeliveryFailed, store.DeliverySent}}, []string{"voice", "telegram"}, []string{store.DeliveryFailed, store.DeliverySent}},
{"sev1 does not fall through", loop.Sev1, // care severities still don't reach an away channel; since #370 the
[]string{"voice"}, []string{store.DeliveryFailed}}, // drop itself is a visible row instead of nothing.
{"sev2 does not fall through", loop.Sev2, {"sev1 drops instead of falling through", loop.Sev1,
[]string{"voice"}, []string{store.DeliveryFailed}}, []string{"voice", "drop"}, []string{store.DeliveryFailed, store.DeliveryDropped}},
{"sev2 drops instead of falling through", loop.Sev2,
[]string{"voice", "drop"}, []string{store.DeliveryFailed, store.DeliveryDropped}},
} }
for _, c := range cases { for _, c := range cases {
t.Run(c.name, func(t *testing.T) { t.Run(c.name, func(t *testing.T) {
+9 -13
View File
@@ -2,10 +2,10 @@
// //
// ntfy is the away-channel for sev3 (ops soft) nudges, sev4 (ops hard) // ntfy is the away-channel for sev3 (ops soft) nudges, sev4 (ops hard)
// nudges when present (alongside voice), and reminders when away. the // nudges when present (alongside voice), and reminders when away. the
// message body is the Sendable's Summary — the minimal-body rule from the // message body is delivery.AwayMessage — the minimal-body rule from the
// spec ("disk low on homesrv," not detail; no shoulder-surf exfil through // spec ("disk low on homesrv," not detail; no shoulder-surf exfil through
// the relay). voice gets Body; away channels get Summary, enforced at the // the relay). the dispatcher already strips detail off away sendables; the
// sink so a phraser bug can't exfil. // sink uses the same helper so it can't leak the body on its own either.
// //
// ntfy runs locally (docker, 127.0.0.1:8085, deny-all auth). maven publishes // ntfy runs locally (docker, 127.0.0.1:8085, deny-all auth). maven publishes
// with a dedicated user (write-only to maven-* topics) — the credential is a // with a dedicated user (write-only to maven-* topics) — the credential is a
@@ -69,18 +69,14 @@ func New(cfg Config) (*Sink, error) {
}, nil }, nil
} }
// Send publishes one notification to ntfy. the body is the Sendable's Summary // Send publishes one notification to ntfy. the body is the minimal away
// (minimal body); Title is "maven" (consistent sender identity on the lock // message (never the full body); Title is "maven" (consistent sender identity
// screen — the content is in the body). Priority maps from severity/kind so // on the lock screen — the content is in the body). Priority maps from severity/kind so
// the phone client can ring differently for an alarm vs a soft ops nudge. // the phone client can ring differently for an alarm vs a soft ops nudge.
func (s *Sink) Send(ctx context.Context, d delivery.Sendable) error { func (s *Sink) Send(ctx context.Context, d delivery.Sendable) error {
body := d.Summary // never fall back to d.Body: ntfy leaves the box, so an empty summary gets
if body == "" { // a generic line instead of the full detail.
body = d.Body // terse full message beats no message body := delivery.AwayMessage(d)
}
if body == "" {
return fmt.Errorf("ntfysink: empty message for %s", d.Channel)
}
req, err := http.NewRequestWithContext(ctx, http.MethodPost, s.topicURL(), strings.NewReader(body)) req, err := http.NewRequestWithContext(ctx, http.MethodPost, s.topicURL(), strings.NewReader(body))
if err != nil { if err != nil {
+16 -9
View File
@@ -147,9 +147,9 @@ func TestSendBodyIsSummaryNotFullBody(t *testing.T) {
} }
} }
func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) { func TestSendNeverSendsTheBodyWhenSummaryEmpty(t *testing.T) {
// a terse full message is better than no message; the phraser should // #368: this used to fall back to the full body. ntfy leaves the box, so
// produce a summary for away-bound severities, but don't silently drop. // an empty summary gets a fixed generic line plus the rule name instead.
rs := newRecordingServer(t, 200, "") rs := newRecordingServer(t, 200, "")
srv := httptest.NewServer(rs.handler()) srv := httptest.NewServer(rs.handler())
defer srv.Close() defer srv.Close()
@@ -160,12 +160,15 @@ func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
t.Fatalf("Send: %v", err) t.Fatalf("Send: %v", err)
} }
_, _, body, _, _, _ := rs.snapshot() _, _, body, _, _, _ := rs.snapshot()
if body != s.Body { want := delivery.GenericAwayMessage + ": service_down"
t.Fatalf("fallback body: want %q, got %q", s.Body, body) if body != want {
t.Fatalf("body: want %q, got %q", want, body)
} }
} }
func TestSendRejectsEmptyMessage(t *testing.T) { func TestSendNeverSendsAnEmptyMessage(t *testing.T) {
// with nothing at all to say we still send the generic line — an away
// channel can never carry detail, but it also never goes out blank.
rs := newRecordingServer(t, 200, "") rs := newRecordingServer(t, 200, "")
srv := httptest.NewServer(rs.handler()) srv := httptest.NewServer(rs.handler())
defer srv.Close() defer srv.Close()
@@ -173,9 +176,13 @@ func TestSendRejectsEmptyMessage(t *testing.T) {
sink, _ := New(Config{BaseURL: srv.URL, Topic: "maven"}) sink, _ := New(Config{BaseURL: srv.URL, Topic: "maven"})
s := nudgeSendable(loop.Sev3, "") s := nudgeSendable(loop.Sev3, "")
s.Body = "" s.Body = ""
err := sink.Send(context.Background(), s) s.RuleName = ""
if err == nil { if err := sink.Send(context.Background(), s); err != nil {
t.Fatal("want error for empty message") t.Fatalf("Send: %v", err)
}
_, _, body, _, _, _ := rs.snapshot()
if body != delivery.GenericAwayMessage {
t.Fatalf("body: want %q, got %q", delivery.GenericAwayMessage, body)
} }
} }
+2 -3
View File
@@ -208,10 +208,9 @@ func TestAwayChannelsGetMinimalBody(t *testing.T) {
// TestCareAwayDropIsRecorded — DESIGN.md's drop is a decision ("a missed water // TestCareAwayDropIsRecorded — DESIGN.md's drop is a decision ("a missed water
// nudge is noise, a missed backup failure isn't"), so it should be visible // nudge is noise, a missed backup failure isn't"), so it should be visible
// rather than vanish. Today drop is a bare `continue`: no nudge row, no outbox // rather than vanish. Today drop is a bare `continue`: no nudge row, no outbox
// attempt, no log — nothing an operator can see afterwards. // attempt, no log — nothing an operator can see afterwards. now it leaves a
// 'dropped' outbox row.
func TestCareAwayDropIsRecorded(t *testing.T) { func TestCareAwayDropIsRecorded(t *testing.T) {
t.Skip("not implemented: dispatcher.go:149-151 skips a Drop channel with no record; there is no 'dropped' outcome in store/delivery.go:16-21")
ob := &fakeOutbox{} ob := &fakeOutbox{}
d := NewDispatcher(Config{Voice: &fakeSink{}, Nudges: &fakeNudgeRecorder{}, Outbox: ob}) d := NewDispatcher(Config{Voice: &fakeSink{}, Nudges: &fakeNudgeRecorder{}, Outbox: ob})
+12 -16
View File
@@ -2,11 +2,12 @@
// //
// telegram is the away-channel for sev4 (ops hard) nudges — "disk-fire alarm // telegram is the away-channel for sev4 (ops hard) nudges — "disk-fire alarm
// at 2am routes to telegram, repeat til ack." the message body is the // at 2am routes to telegram, repeat til ack." the message body is the
// Sendable's Summary — the minimal-body rule from the spec ("disk low on // delivery.AwayMessage — the minimal-body rule from the spec ("disk low on
// homesrv," not detail; no shoulder-surf exfil through the relay). voice gets // homesrv," not detail; no shoulder-surf exfil through the relay). the
// Body; away channels get Summary, enforced at the sink so a phraser bug can't // dispatcher already strips detail off away sendables; the sink uses the same
// exfil. additionally, protect_content=true is passed on every send so the // helper so it can't leak the body on its own either. additionally,
// message can't be forwarded out of the chat — locks the minimal body further. // protect_content=true is passed on every send so the message can't be
// forwarded out of the chat — locks the minimal body further.
// //
// telegram's bot API is region-restricted for this homesrv — direct egress to // telegram's bot API is region-restricted for this homesrv — direct egress to
// api.telegram.org is unreliable. the spec's "away channels leave the box — // api.telegram.org is unreliable. the spec's "away channels leave the box —
@@ -140,18 +141,13 @@ type telegramResp struct {
} }
// Send publishes one message to the configured telegram chat. the body is the // Send publishes one message to the configured telegram chat. the body is the
// Sendable's Summary (minimal body); empty Summary falls back to Body (terse // minimal away message (never the full body). protect_content=true so even
// full message beats no message). protect_content=true so a phraser bug (Body // that can't be forwarded onward by the user or a chat observer — locks the
// leaking detail through Summary) can't be forwarded onward by the user or a // minimal-body rule at the channel's own last mile.
// chat observer — locks the minimal-body rule at the channel's own last mile.
func (s *Sink) Send(ctx context.Context, d delivery.Sendable) error { func (s *Sink) Send(ctx context.Context, d delivery.Sendable) error {
body := d.Summary // never fall back to d.Body: telegram leaves the box, so an empty summary
if body == "" { // gets a generic line instead of the full detail.
body = d.Body body := delivery.AwayMessage(d)
}
if body == "" {
return fmt.Errorf("telegramsink: empty message for %s", d.Channel)
}
payload := sendMessageReq{ payload := sendMessageReq{
ChatID: s.cfg.ChatID, ChatID: s.cfg.ChatID,
@@ -173,9 +173,9 @@ func TestSendBodyIsSummaryNotFullBody(t *testing.T) {
} }
} }
func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) { func TestSendNeverSendsTheBodyWhenSummaryEmpty(t *testing.T) {
// terse full message beats none; the phraser should produce a summary for // #368: this used to fall back to the full body. telegram leaves the box,
// away-bound severities, but don't silently drop. // so an empty summary gets a fixed generic line plus the rule name.
rs := newRecordingServer(t, 200, "") rs := newRecordingServer(t, 200, "")
srv := httptest.NewServer(rs.handler()) srv := httptest.NewServer(rs.handler())
defer srv.Close() defer srv.Close()
@@ -188,12 +188,14 @@ func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
_, _, body, _, _ := rs.snapshot() _, _, body, _, _ := rs.snapshot()
var req sendMessageReq var req sendMessageReq
_ = json.Unmarshal([]byte(body), &req) _ = json.Unmarshal([]byte(body), &req)
if req.Text != s.Body { want := delivery.GenericAwayMessage + ": service_down"
t.Fatalf("fallback text: want %q, got %q", s.Body, req.Text) if req.Text != want {
t.Fatalf("text: want %q, got %q", want, req.Text)
} }
} }
func TestSendRejectsEmptyMessage(t *testing.T) { func TestSendNeverSendsAnEmptyMessage(t *testing.T) {
// with nothing at all to say we still send the generic line.
rs := newRecordingServer(t, 200, "") rs := newRecordingServer(t, 200, "")
srv := httptest.NewServer(rs.handler()) srv := httptest.NewServer(rs.handler())
defer srv.Close() defer srv.Close()
@@ -201,9 +203,15 @@ func TestSendRejectsEmptyMessage(t *testing.T) {
sink, _ := New(sinkCfg(srv.URL)) sink, _ := New(sinkCfg(srv.URL))
s := nudgeSendable(loop.Sev4, "") s := nudgeSendable(loop.Sev4, "")
s.Body = "" s.Body = ""
err := sink.Send(context.Background(), s) s.RuleName = ""
if err == nil { if err := sink.Send(context.Background(), s); err != nil {
t.Fatal("want error for empty message") t.Fatalf("Send: %v", err)
}
_, _, body, _, _ := rs.snapshot()
var req sendMessageReq
_ = json.Unmarshal([]byte(body), &req)
if req.Text != delivery.GenericAwayMessage {
t.Fatalf("text: want %q, got %q", delivery.GenericAwayMessage, req.Text)
} }
} }
+97
View File
@@ -93,6 +93,63 @@ type WriteFactReq struct {
Subject string `json:"subject,omitempty"` Subject string `json:"subject,omitempty"`
} }
// Task — one captured piece of work (Vikunja #130). Status is
// "candidate" (Maven derived it and it is unconfirmed), "open" (his work),
// "done" or "dropped". Source is provenance in the facts vocabulary:
// "tap:voice", "tap:web", "email:<account>". Evidence is the trail a derived
// task came from, empty for anything he stated himself.
type Task struct {
ID int64 `json:"id"`
CreatedTs time.Time `json:"created_ts"`
Text string `json:"text"`
Source string `json:"source"`
Evidence string `json:"evidence,omitempty"`
Status string `json:"status"`
Due *time.Time `json:"due,omitempty"`
Weight int `json:"weight,omitempty"`
Resolved *time.Time `json:"resolved,omitempty"`
}
// CaptureTaskReq — THE INTAKE SEAM. Everything that captures a task goes
// through this one shape: the voice path, the web form, and (Vikunja #246) the
// email reader, which has not been built yet.
//
// An extractor that reads mail sets Source "email:<account>", Status
// "candidate", and Evidence to whatever makes the task reviewable (the subject
// line). It must NOT set Status "open" — work Maven inferred from something she
// read is a suggestion until the owner confirms it on the /tasks page. Capture
// is idempotent on normalised text among live tasks, so re-reading the same
// mailbox is free.
type CaptureTaskReq struct {
Text string `json:"text"`
Source string `json:"source"`
Evidence string `json:"evidence,omitempty"`
Status string `json:"status,omitempty"` // "" ⇒ open
Due *time.Time `json:"due,omitempty"`
Weight int `json:"weight,omitempty"`
Ts time.Time `json:"ts"`
}
// CaptureTaskResp — Created is false when the same live task already existed,
// in which case ID is the existing row. A caller tells the owner "уже в
// списке" rather than claiming it saved something new.
type CaptureTaskResp struct {
ID int64 `json:"id"`
Created bool `json:"created"`
}
type listTasksReq struct {
Status string `json:"status"` // "" all | "live" | candidate|open|done|dropped
}
type listTasksResp struct {
Tasks []Task `json:"tasks"`
}
type setTaskStatusReq struct {
ID int64 `json:"id"`
Status string `json:"status"`
Ts time.Time `json:"ts"`
}
// idReq — methods keyed by a single id. // idReq — methods keyed by a single id.
type idReq struct { type idReq struct {
ID int64 `json:"id"` ID int64 `json:"id"`
@@ -294,6 +351,18 @@ type CoreAPI interface {
// loop takes the schedule from there — no reminder is created (Vikunja #366). // loop takes the schedule from there — no reminder is created (Vikunja #366).
AcceptProposedRoutine(ctx context.Context, id int64) error AcceptProposedRoutine(ctx context.Context, id int64) error
// CaptureTask records a task. See CaptureTaskReq — this is the single
// intake seam for the voice path, the web form and the future email
// extractor. Idempotent per live normalised text; the response says
// whether a row was actually created.
CaptureTask(ctx context.Context, req CaptureTaskReq) (CaptureTaskResp, error)
// ListTasks returns tasks in one status, newest first. "" is every row,
// "live" is candidate + open (outstanding work).
ListTasks(ctx context.Context, status string) ([]Task, error)
// SetTaskStatus moves a task forward once: candidate→open|dropped,
// open→done|dropped. Any other move is refused.
SetTaskStatus(ctx context.Context, id int64, status string, ts time.Time) error
// TickTrace returns the most recent tick's rule trace. The daemon caches // TickTrace returns the most recent tick's rule trace. The daemon caches
// this after every tick; the store adapter returns an error (trace is not // this after every tick; the store adapter returns an error (trace is not
// persisted — it's a daemon-level cache). // persisted — it's a daemon-level cache).
@@ -306,6 +375,14 @@ type CoreAPI interface {
// TickTrace. // TickTrace.
MorningStatus(ctx context.Context) ([]MorningRoutineStatus, error) MorningStatus(ctx context.Context) ([]MorningRoutineStatus, error)
// DayPlan returns today's ordered plan — calendar events, pending
// reminders and any morning checklist still outstanding (see
// internal/morning.BuildPlan) — plus the spoken RU rendering of it.
// Read-only: asking for the plan never dispatches or schedules anything.
// The store adapter returns an error (the plan needs the daemon's routine
// config) — same shape as TickTrace and MorningStatus.
DayPlan(ctx context.Context) (DayPlan, error)
// Chat routes a text utterance through the reactive handler's core path // Chat routes a text utterance through the reactive handler's core path
// (router → dialogue → action → replier) and returns the reply text. // (router → dialogue → action → replier) and returns the reply text.
// No audio or stt/tts — for text channels (mavweb, telegram). // No audio or stt/tts — for text channels (mavweb, telegram).
@@ -359,6 +436,26 @@ type MorningRoutineStatus struct {
Items []MorningRoutineItem `json:"items"` Items []MorningRoutineItem `json:"items"`
} }
// DayPlanItem — one line of the day plan. Kind is "event", "reminder" or
// "checklist"; Uncertain marks an item whose provenance is below a full
// calendar read (a meeting relayed off a phone notification), so a UI can hedge
// the same way the spoken form does.
type DayPlanItem struct {
At time.Time `json:"at"`
Text string `json:"text"`
Kind string `json:"kind"`
Uncertain bool `json:"uncertain,omitempty"`
}
// DayPlan — the plan for one calendar day. Spoken is the RU sentence maven
// says when asked, rendered core-side so the voice reply and the web view can
// never drift apart.
type DayPlan struct {
Date time.Time `json:"date"`
Items []DayPlanItem `json:"items"`
Spoken string `json:"spoken"`
}
// storeEncryptionKeyReq — passkey credential public key for wrapping the store // storeEncryptionKeyReq — passkey credential public key for wrapping the store
// encryption key at enrollment time. Called by mavweb after RegisterFinish. // encryption key at enrollment time. Called by mavweb after RegisterFinish.
type storeEncryptionKeyReq struct { type storeEncryptionKeyReq struct {
+30
View File
@@ -69,8 +69,10 @@ var readOnlyMethods = map[Method]bool{
MethodLookupTool: true, MethodLookupTool: true,
MethodListTools: true, MethodListTools: true,
MethodListProposedRoutines: true, MethodListProposedRoutines: true,
MethodListTasks: true,
MethodTickTrace: true, MethodTickTrace: true,
MethodMorningStatus: true, MethodMorningStatus: true,
MethodDayPlan: true,
} }
// Dial connects to a core socket at path and returns a Client. The module // Dial connects to a core socket at path and returns a Client. The module
@@ -426,6 +428,26 @@ func (c *Client) ListProposedRoutines(ctx context.Context) ([]ProposedRoutine, e
return r.Routines, nil return r.Routines, nil
} }
func (c *Client) CaptureTask(ctx context.Context, req CaptureTaskReq) (CaptureTaskResp, error) {
var r CaptureTaskResp
if err := c.call(ctx, MethodCaptureTask, req, &r); err != nil {
return CaptureTaskResp{}, err
}
return r, nil
}
func (c *Client) ListTasks(ctx context.Context, status string) ([]Task, error) {
var r listTasksResp
if err := c.call(ctx, MethodListTasks, listTasksReq{Status: status}, &r); err != nil {
return nil, err
}
return r.Tasks, nil
}
func (c *Client) SetTaskStatus(ctx context.Context, id int64, status string, ts time.Time) error {
return c.call(ctx, MethodSetTaskStatus, setTaskStatusReq{ID: id, Status: status, Ts: ts}, nil)
}
func (c *Client) DismissProposedRoutine(ctx context.Context, id int64) error { func (c *Client) DismissProposedRoutine(ctx context.Context, id int64) error {
return c.call(ctx, MethodDismissProposedRoutine, dismissProposedRoutineReq{ID: id}, nil) return c.call(ctx, MethodDismissProposedRoutine, dismissProposedRoutineReq{ID: id}, nil)
} }
@@ -458,6 +480,14 @@ func (c *Client) MorningStatus(ctx context.Context) ([]MorningRoutineStatus, err
return s, nil return s, nil
} }
func (c *Client) DayPlan(ctx context.Context) (DayPlan, error) {
var p DayPlan
if err := c.call(ctx, MethodDayPlan, nil, &p); err != nil {
return DayPlan{}, err
}
return p, nil
}
func (c *Client) RevertFact(ctx context.Context, key string) (int64, error) { func (c *Client) RevertFact(ctx context.Context, key string) (int64, error) {
var result struct { var result struct {
NewID int64 `json:"new_id"` NewID int64 `json:"new_id"`
+5 -88
View File
@@ -410,95 +410,12 @@ func TestChatViaClient(t *testing.T) {
} }
// chatTestAPI — a minimal CoreAPI that only implements Chat for testing. // chatTestAPI — a minimal CoreAPI that only implements Chat for testing.
type chatTestAPI struct{} // Embeds UnimplementedCoreAPI so every other method fails loudly with
// ErrNotImplemented instead of needing 27 hand-written no-op stubs.
type chatTestAPI struct {
UnimplementedCoreAPI
}
func (a *chatTestAPI) WriteFact(ctx context.Context, req WriteFactReq) (int64, error) {
return 0, ErrUnknownMethod
}
func (a *chatTestAPI) LatestFact(ctx context.Context, key string) (Fact, error) {
return Fact{}, ErrUnknownMethod
}
func (a *chatTestAPI) LatestFactBySource(ctx context.Context, key, source string) (Fact, error) {
return Fact{}, ErrUnknownMethod
}
func (a *chatTestAPI) Since(ctx context.Context, key string, now time.Time) (time.Duration, error) {
return 0, ErrUnknownMethod
}
func (a *chatTestAPI) Presence(ctx context.Context) (Presence, error) {
return Presence{}, ErrUnknownMethod
}
func (a *chatTestAPI) CreateReminder(ctx context.Context, fire time.Time, payload, cron string) (int64, error) {
return 0, ErrUnknownMethod
}
func (a *chatTestAPI) MarkReminder(ctx context.Context, id int64, status string) error {
return ErrUnknownMethod
}
func (a *chatTestAPI) ListReminders(ctx context.Context, n int) ([]Reminder, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) RecordNudge(ctx context.Context, rule, channel, message string, ts time.Time) (int64, error) {
return 0, ErrUnknownMethod
}
func (a *chatTestAPI) ResolveNudge(ctx context.Context, id int64, outcome string, ts time.Time) error {
return ErrUnknownMethod
}
func (a *chatTestAPI) RecentOutcomes(ctx context.Context, rule string, n int) ([]string, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) RecentFacts(ctx context.Context, n int) ([]Fact, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) CalendarEvents(ctx context.Context, from, to time.Time) ([]Fact, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) RecentNudges(ctx context.Context, n int) ([]Nudge, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) WriteNote(ctx context.Context, ts time.Time, text string, embedding []float32, source string) (int64, error) {
return 0, ErrUnknownMethod
}
func (a *chatTestAPI) QueryNotes(ctx context.Context, embedding []float32, k int) ([]Note, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) RecentNotes(ctx context.Context, n int) ([]Note, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) ProposeTool(ctx context.Context, name, utterance, scope string, ts time.Time) (bool, error) {
return false, ErrUnknownMethod
}
func (a *chatTestAPI) EnableTool(ctx context.Context, name string, cmd []string, destructive bool, scope string, ts time.Time) error {
return ErrUnknownMethod
}
func (a *chatTestAPI) DisableTool(ctx context.Context, name string) error {
return ErrUnknownMethod
}
func (a *chatTestAPI) DeleteTool(ctx context.Context, name string) error {
return ErrUnknownMethod
}
func (a *chatTestAPI) ListProposedRoutines(ctx context.Context) ([]ProposedRoutine, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) AcceptProposedRoutine(ctx context.Context, id int64) error {
return nil
}
func (a *chatTestAPI) DismissProposedRoutine(ctx context.Context, id int64) error {
return ErrUnknownMethod
}
func (a *chatTestAPI) LookupTool(ctx context.Context, name string) (Tool, error) {
return Tool{}, ErrUnknownMethod
}
func (a *chatTestAPI) ListTools(ctx context.Context, status string) ([]Tool, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) RevertFact(ctx context.Context, key string) (int64, error) {
return 0, ErrUnknownMethod
}
func (a *chatTestAPI) TickTrace(ctx context.Context) (TickTrace, error) {
return TickTrace{}, ErrUnknownMethod
}
func (a *chatTestAPI) MorningStatus(ctx context.Context) ([]MorningRoutineStatus, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) Chat(ctx context.Context, text string) (string, error) { func (a *chatTestAPI) Chat(ctx context.Context, text string) (string, error) {
if text == "привет" { if text == "привет" {
return "и тебе привет!", nil return "и тебе привет!", nil
+306 -306
View File
@@ -211,6 +211,10 @@ func (a *storeAPI) MorningStatus(ctx context.Context) ([]MorningRoutineStatus, e
return nil, errors.New("store: morning status not available via direct store API") return nil, errors.New("store: morning status not available via direct store API")
} }
func (a *storeAPI) DayPlan(ctx context.Context) (DayPlan, error) {
return DayPlan{}, errors.New("store: day plan not available via direct store API")
}
func (a *storeAPI) ListTools(ctx context.Context, status string) ([]Tool, error) { func (a *storeAPI) ListTools(ctx context.Context, status string) ([]Tool, error) {
ts, err := a.s.ListTools(ctx, status) ts, err := a.s.ListTools(ctx, status)
if err != nil { if err != nil {
@@ -227,6 +231,48 @@ func (a *storeAPI) DeleteTool(ctx context.Context, name string) error {
return mapErr(a.s.DeleteTool(ctx, name)) return mapErr(a.s.DeleteTool(ctx, name))
} }
func (a *storeAPI) CaptureTask(ctx context.Context, req CaptureTaskReq) (CaptureTaskResp, error) {
id, created, err := a.s.CaptureTask(ctx, store.Task{
CreatedTs: req.Ts,
Text: req.Text,
Source: req.Source,
Evidence: req.Evidence,
Status: req.Status,
Due: req.Due,
Weight: req.Weight,
})
if err != nil {
return CaptureTaskResp{}, mapErr(err)
}
return CaptureTaskResp{ID: id, Created: created}, nil
}
func (a *storeAPI) ListTasks(ctx context.Context, status string) ([]Task, error) {
ts, err := a.s.ListTasks(ctx, status)
if err != nil {
return nil, mapErr(err)
}
out := make([]Task, len(ts))
for i, t := range ts {
out[i] = Task{
ID: t.ID,
CreatedTs: t.CreatedTs,
Text: t.Text,
Source: t.Source,
Evidence: t.Evidence,
Status: t.Status,
Due: t.Due,
Weight: t.Weight,
Resolved: t.ResolvedTs,
}
}
return out, nil
}
func (a *storeAPI) SetTaskStatus(ctx context.Context, id int64, status string, ts time.Time) error {
return mapErr(a.s.SetTaskStatus(ctx, id, status, ts))
}
func (a *storeAPI) ListProposedRoutines(ctx context.Context) ([]ProposedRoutine, error) { func (a *storeAPI) ListProposedRoutines(ctx context.Context) ([]ProposedRoutine, error) {
rs, err := a.s.ListProposedRoutines(ctx) rs, err := a.s.ListProposedRoutines(ctx)
if err != nil { if err != nil {
@@ -501,6 +547,258 @@ func (s *Server) safeDispatch(ctx context.Context, req Request) (result json.Raw
return s.dispatch(ctx, req) return s.dispatch(ctx, req)
} }
// handlerFunc — one table entry's shape: unmarshal req.Params (if it wants
// any), call the matching CoreAPI method against the api passed in, marshal
// the result. api is a parameter, not a closed-over field, precisely so a
// table built once at package init never pins a stale CoreAPI — see the note
// on methodTable below about SetAPI.
type handlerFunc func(ctx context.Context, api CoreAPI, raw json.RawMessage) (json.RawMessage, error)
// withParams adapts a (typed params, typed result) CoreAPI call into a
// handlerFunc: unmarshal into P, call fn, marshal R. On error the result is
// dropped (marshalResult's output is never read when err != nil — see
// serveConn) so every entry can uniformly return early on error without
// re-deriving what the pre-table per-arm code used to return in that case.
func withParams[P any, R any](fn func(ctx context.Context, api CoreAPI, p P) (R, error)) handlerFunc {
return func(ctx context.Context, api CoreAPI, raw json.RawMessage) (json.RawMessage, error) {
var p P
if err := unmarshalParams(raw, &p); err != nil {
return nil, err
}
r, err := fn(ctx, api, p)
if err != nil {
return nil, err
}
return marshalResult(r), nil
}
}
// withParamsVoid is withParams for the error-only methods (mark/resolve/
// enable/disable/...): params in, no result out, wire reply is always null.
func withParamsVoid[P any](fn func(ctx context.Context, api CoreAPI, p P) error) handlerFunc {
return func(ctx context.Context, api CoreAPI, raw json.RawMessage) (json.RawMessage, error) {
var p P
if err := unmarshalParams(raw, &p); err != nil {
return nil, err
}
return marshalResult(nil), fn(ctx, api, p)
}
}
// withoutParams is withParams for the handful of methods that take no
// params at all (Presence, TickTrace, MorningStatus, ListProposedRoutines).
// It does NOT call unmarshalParams — matching the pre-table arms, which
// never touched req.Params for these four methods.
func withoutParams[R any](fn func(ctx context.Context, api CoreAPI) (R, error)) handlerFunc {
return func(ctx context.Context, api CoreAPI, _ json.RawMessage) (json.RawMessage, error) {
r, err := fn(ctx, api)
if err != nil {
return nil, err
}
return marshalResult(r), nil
}
}
// methodTable — one entry per CoreAPI-backed method. Built once at package
// init, not per-Server and not per-dispatch: entries close over nothing but
// the CoreAPI method being called, and dispatch passes in the *current*
// api (loaded fresh via s.api.Load() every call, same as before the table
// existed) as an argument — so SetAPI's runtime swap (the unlock transition)
// is still honored on the very next request with no extra plumbing here.
//
// MethodAssertStepUp, MethodStoreEncryptionKey and MethodUnlock are NOT in
// this table: they bypass CoreAPI entirely (s.StepUp / s.WrapKeyFn /
// s.UnlockFn), so dispatch special-cases them before consulting the table.
var methodTable = map[Method]handlerFunc{
MethodWriteFact: withParams(func(ctx context.Context, api CoreAPI, p WriteFactReq) (idResp, error) {
id, err := api.WriteFact(ctx, p)
return idResp{ID: id}, err
}),
MethodLatestFact: withParams(func(ctx context.Context, api CoreAPI, p keyReq) (Fact, error) {
return api.LatestFact(ctx, p.Key)
}),
MethodLatestFactBySource: withParams(func(ctx context.Context, api CoreAPI, p keySourceReq) (Fact, error) {
return api.LatestFactBySource(ctx, p.Key, p.Source)
}),
MethodSince: withParams(func(ctx context.Context, api CoreAPI, p sinceReq) (sinceResp, error) {
d, err := api.Since(ctx, p.Key, p.Now)
return sinceResp{Dur: d}, err
}),
MethodPresence: withoutParams(func(ctx context.Context, api CoreAPI) (Presence, error) {
return api.Presence(ctx)
}),
MethodCreateReminder: withParams(func(ctx context.Context, api CoreAPI, p createReminderReq) (idResp, error) {
id, err := api.CreateReminder(ctx, p.Fire, p.Payload, p.Cron)
return idResp{ID: id}, err
}),
MethodMarkReminder: withParamsVoid(func(ctx context.Context, api CoreAPI, p markReminderReq) error {
return api.MarkReminder(ctx, p.ID, p.Status)
}),
MethodListReminders: withParams(func(ctx context.Context, api CoreAPI, p nReq) ([]Reminder, error) {
out, err := api.ListReminders(ctx, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []Reminder{}
}
return out, nil
}),
MethodRecordNudge: withParams(func(ctx context.Context, api CoreAPI, p recordNudgeReq) (idResp, error) {
id, err := api.RecordNudge(ctx, p.Rule, p.Channel, p.Message, p.Ts)
return idResp{ID: id}, err
}),
MethodResolveNudge: withParamsVoid(func(ctx context.Context, api CoreAPI, p resolveNudgeReq) error {
return api.ResolveNudge(ctx, p.ID, p.Outcome, p.Ts)
}),
MethodRecentOutcomes: withParams(func(ctx context.Context, api CoreAPI, p outcomesReq) ([]string, error) {
out, err := api.RecentOutcomes(ctx, p.Rule, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []string{} // stable non-null on the wire
}
return out, nil
}),
MethodRecentFacts: withParams(func(ctx context.Context, api CoreAPI, p nReq) ([]Fact, error) {
out, err := api.RecentFacts(ctx, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []Fact{}
}
return out, nil
}),
MethodCalendarEvents: withParams(func(ctx context.Context, api CoreAPI, p calendarEventsReq) ([]Fact, error) {
out, err := api.CalendarEvents(ctx, p.From, p.To)
if err != nil {
return nil, err
}
if out == nil {
out = []Fact{}
}
return out, nil
}),
MethodRecentNudges: withParams(func(ctx context.Context, api CoreAPI, p nReq) ([]Nudge, error) {
out, err := api.RecentNudges(ctx, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []Nudge{}
}
return out, nil
}),
MethodWriteNote: withParams(func(ctx context.Context, api CoreAPI, p writeNoteReq) (idResp, error) {
id, err := api.WriteNote(ctx, p.Ts, p.Text, p.Embedding, p.Source)
return idResp{ID: id}, err
}),
MethodQueryNotes: withParams(func(ctx context.Context, api CoreAPI, p queryNotesReq) ([]Note, error) {
out, err := api.QueryNotes(ctx, p.Embedding, p.K)
if err != nil {
return nil, err
}
if out == nil {
out = []Note{}
}
return out, nil
}),
MethodRecentNotes: withParams(func(ctx context.Context, api CoreAPI, p nReq) ([]Note, error) {
out, err := api.RecentNotes(ctx, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []Note{}
}
return out, nil
}),
MethodProposeTool: withParams(func(ctx context.Context, api CoreAPI, p proposeToolReq) (proposeToolResp, error) {
ok, err := api.ProposeTool(ctx, p.Name, p.Utterance, p.Scope, p.Ts)
return proposeToolResp{Proposed: ok}, err
}),
MethodEnableTool: withParamsVoid(func(ctx context.Context, api CoreAPI, p enableToolReq) error {
return api.EnableTool(ctx, p.Name, p.Cmd, p.Destructive, p.Scope, p.Ts)
}),
MethodDisableTool: withParamsVoid(func(ctx context.Context, api CoreAPI, p disableToolReq) error {
return api.DisableTool(ctx, p.Name)
}),
MethodLookupTool: withParams(func(ctx context.Context, api CoreAPI, p lookupToolReq) (Tool, error) {
return api.LookupTool(ctx, p.Name)
}),
MethodListTools: withParams(func(ctx context.Context, api CoreAPI, p listToolsReq) (listToolsResp, error) {
out, err := api.ListTools(ctx, p.Status)
if err != nil {
return listToolsResp{}, err
}
if out == nil {
out = []Tool{}
}
return listToolsResp{Tools: out}, nil
}),
// MethodDeleteTool shares disableToolReq — both take just a tool name.
MethodDeleteTool: withParamsVoid(func(ctx context.Context, api CoreAPI, p disableToolReq) error {
return api.DeleteTool(ctx, p.Name)
}),
MethodCaptureTask: withParams(func(ctx context.Context, api CoreAPI, p CaptureTaskReq) (CaptureTaskResp, error) {
return api.CaptureTask(ctx, p)
}),
MethodListTasks: withParams(func(ctx context.Context, api CoreAPI, p listTasksReq) (listTasksResp, error) {
out, err := api.ListTasks(ctx, p.Status)
if err != nil {
return listTasksResp{}, err
}
if out == nil {
out = []Task{}
}
return listTasksResp{Tasks: out}, nil
}),
MethodSetTaskStatus: withParamsVoid(func(ctx context.Context, api CoreAPI, p setTaskStatusReq) error {
return api.SetTaskStatus(ctx, p.ID, p.Status, p.Ts)
}),
MethodListProposedRoutines: withoutParams(func(ctx context.Context, api CoreAPI) (listProposedRoutinesResp, error) {
out, err := api.ListProposedRoutines(ctx)
if err != nil {
return listProposedRoutinesResp{}, err
}
if out == nil {
out = []ProposedRoutine{}
}
return listProposedRoutinesResp{Routines: out}, nil
}),
MethodDismissProposedRoutine: withParamsVoid(func(ctx context.Context, api CoreAPI, p dismissProposedRoutineReq) error {
return api.DismissProposedRoutine(ctx, p.ID)
}),
MethodAcceptProposedRoutine: withParamsVoid(func(ctx context.Context, api CoreAPI, p acceptProposedRoutineReq) error {
return api.AcceptProposedRoutine(ctx, p.ID)
}),
MethodRevertFact: withParams(func(ctx context.Context, api CoreAPI, p revertReq) (map[string]int64, error) {
newID, err := api.RevertFact(ctx, p.Key)
if err != nil {
return nil, err
}
return map[string]int64{"new_id": newID}, nil
}),
MethodChat: withParams(func(ctx context.Context, api CoreAPI, p chatReq) (chatResp, error) {
reply, err := api.Chat(ctx, p.Text)
return chatResp{Reply: reply}, err
}),
MethodTickTrace: withoutParams(func(ctx context.Context, api CoreAPI) (TickTrace, error) {
return api.TickTrace(ctx)
}),
// MorningStatus intentionally has no nil→[]T{} normalization here — the
// pre-table arm marshaled api.MorningStatus's result as-is (a nil slice
// serializes as JSON null), and this preserves that exact wire shape.
MethodDayPlan: withoutParams(func(ctx context.Context, api CoreAPI) (DayPlan, error) {
return api.DayPlan(ctx)
}),
MethodMorningStatus: withoutParams(func(ctx context.Context, api CoreAPI) ([]MorningRoutineStatus, error) {
return api.MorningStatus(ctx)
}),
}
// dispatch unmarshals params for req.Method and calls the matching CoreAPI // dispatch unmarshals params for req.Method and calls the matching CoreAPI
// method. Unknown method ⇒ ErrUnknownMethod; a malformed params payload ⇒ // method. Unknown method ⇒ ErrUnknownMethod; a malformed params payload ⇒
// ErrBadParams with the underlying text (local, server-side, not shipped to // ErrBadParams with the underlying text (local, server-side, not shipped to
@@ -517,312 +815,11 @@ func (s *Server) dispatch(ctx context.Context, req Request) (json.RawMessage, er
return nil, err return nil, err
} }
} }
// These three bypass CoreAPI entirely — they drive Server fields set
// directly by the daemon (StepUp / WrapKeyFn / UnlockFn), not store
// state, so they can never be table entries keyed on a CoreAPI method.
switch req.Method { switch req.Method {
case MethodWriteFact:
var p WriteFactReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
id, err := api.WriteFact(ctx, p)
return marshalResult(idResp{ID: id}), err
case MethodLatestFact:
var p keyReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
f, err := api.LatestFact(ctx, p.Key)
if err != nil {
return nil, err
}
return marshalResult(f), nil
case MethodLatestFactBySource:
var p keySourceReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
f, err := api.LatestFactBySource(ctx, p.Key, p.Source)
if err != nil {
return nil, err
}
return marshalResult(f), nil
case MethodSince:
var p sinceReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
d, err := api.Since(ctx, p.Key, p.Now)
if err != nil {
return nil, err
}
return marshalResult(sinceResp{Dur: d}), nil
case MethodPresence:
pres, err := api.Presence(ctx)
if err != nil {
return nil, err
}
return marshalResult(pres), nil
case MethodCreateReminder:
var p createReminderReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
id, err := api.CreateReminder(ctx, p.Fire, p.Payload, p.Cron)
return marshalResult(idResp{ID: id}), err
case MethodMarkReminder:
var p markReminderReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
err := api.MarkReminder(ctx, p.ID, p.Status)
return marshalResult(nil), err
case MethodListReminders:
var p nReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
out, err := api.ListReminders(ctx, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []Reminder{}
}
return marshalResult(out), nil
case MethodRecordNudge:
var p recordNudgeReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
id, err := api.RecordNudge(ctx, p.Rule, p.Channel, p.Message, p.Ts)
return marshalResult(idResp{ID: id}), err
case MethodResolveNudge:
var p resolveNudgeReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
err := api.ResolveNudge(ctx, p.ID, p.Outcome, p.Ts)
return marshalResult(nil), err
case MethodRecentOutcomes:
var p outcomesReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
out, err := api.RecentOutcomes(ctx, p.Rule, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []string{} // stable non-null on the wire
}
return marshalResult(out), nil
case MethodRecentFacts:
var p nReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
out, err := api.RecentFacts(ctx, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []Fact{}
}
return marshalResult(out), nil
case MethodCalendarEvents:
var p calendarEventsReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
out, err := api.CalendarEvents(ctx, p.From, p.To)
if err != nil {
return nil, err
}
if out == nil {
out = []Fact{}
}
return marshalResult(out), nil
case MethodRecentNudges:
var p nReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
out, err := api.RecentNudges(ctx, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []Nudge{}
}
return marshalResult(out), nil
case MethodWriteNote:
var p writeNoteReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
id, err := api.WriteNote(ctx, p.Ts, p.Text, p.Embedding, p.Source)
return marshalResult(idResp{ID: id}), err
case MethodQueryNotes:
var p queryNotesReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
out, err := api.QueryNotes(ctx, p.Embedding, p.K)
if err != nil {
return nil, err
}
if out == nil {
out = []Note{}
}
return marshalResult(out), nil
case MethodRecentNotes:
var p nReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
out, err := api.RecentNotes(ctx, p.N)
if err != nil {
return nil, err
}
if out == nil {
out = []Note{}
}
return marshalResult(out), nil
case MethodProposeTool:
var p proposeToolReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
ok, err := api.ProposeTool(ctx, p.Name, p.Utterance, p.Scope, p.Ts)
if err != nil {
return nil, err
}
return marshalResult(proposeToolResp{Proposed: ok}), nil
case MethodEnableTool:
var p enableToolReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
return marshalResult(nil), api.EnableTool(ctx, p.Name, p.Cmd, p.Destructive, p.Scope, p.Ts)
case MethodDisableTool:
var p disableToolReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
return marshalResult(nil), api.DisableTool(ctx, p.Name)
case MethodLookupTool:
var p lookupToolReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
t, err := api.LookupTool(ctx, p.Name)
if err != nil {
return nil, err
}
return marshalResult(t), nil
case MethodListTools:
var p listToolsReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
out, err := api.ListTools(ctx, p.Status)
if err != nil {
return nil, err
}
if out == nil {
out = []Tool{}
}
return marshalResult(listToolsResp{Tools: out}), nil
case MethodDeleteTool:
var p disableToolReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
return marshalResult(nil), api.DeleteTool(ctx, p.Name)
case MethodListProposedRoutines:
out, err := api.ListProposedRoutines(ctx)
if err != nil {
return nil, err
}
if out == nil {
out = []ProposedRoutine{}
}
return marshalResult(listProposedRoutinesResp{Routines: out}), nil
case MethodDismissProposedRoutine:
var p dismissProposedRoutineReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
return marshalResult(nil), api.DismissProposedRoutine(ctx, p.ID)
case MethodAcceptProposedRoutine:
var p acceptProposedRoutineReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
return marshalResult(nil), api.AcceptProposedRoutine(ctx, p.ID)
case MethodRevertFact:
var p struct {
Key string `json:"key"`
}
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
newID, err := api.RevertFact(ctx, p.Key)
if err != nil {
return nil, err
}
return marshalResult(map[string]int64{"new_id": newID}), nil
case MethodChat:
var p chatReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
reply, err := api.Chat(ctx, p.Text)
if err != nil {
return nil, err
}
return marshalResult(chatResp{Reply: reply}), nil
case MethodTickTrace:
t, err := api.TickTrace(ctx)
if err != nil {
return nil, err
}
return marshalResult(t), nil
case MethodMorningStatus:
s, err := api.MorningStatus(ctx)
if err != nil {
return nil, err
}
return marshalResult(s), nil
case MethodAssertStepUp: case MethodAssertStepUp:
if s.StepUp != nil { if s.StepUp != nil {
return marshalResult(nil), s.StepUp(ctx) return marshalResult(nil), s.StepUp(ctx)
@@ -848,10 +845,13 @@ func (s *Server) dispatch(ctx context.Context, req Request) (json.RawMessage, er
return marshalResult(nil), s.UnlockFn(ctx, p.PublicKey) return marshalResult(nil), s.UnlockFn(ctx, p.PublicKey)
} }
return nil, fmt.Errorf("%w: %s", ErrUnknownMethod, req.Method) return nil, fmt.Errorf("%w: %s", ErrUnknownMethod, req.Method)
}
default: h, ok := methodTable[req.Method]
if !ok {
return nil, fmt.Errorf("%w: %s", ErrUnknownMethod, req.Method) return nil, fmt.Errorf("%w: %s", ErrUnknownMethod, req.Method)
} }
return h(ctx, api, req.Params)
} }
func unmarshalParams(raw json.RawMessage, v any) error { func unmarshalParams(raw json.RawMessage, v any) error {
+130
View File
@@ -0,0 +1,130 @@
package ipc
import (
"context"
"errors"
"time"
)
// ErrNotImplemented is returned by every UnimplementedCoreAPI method. It is
// deliberately distinct from ErrUnknownMethod (a wire-level "no such
// method exists" verdict) and from any daemon-level "locked" error: this one
// means "this method exists on CoreAPI, but the fake/adapter embedding
// UnimplementedCoreAPI never got a real implementation for it." A test that
// exercises an undeclared method fails loudly on this text instead of
// silently nil-panicking or being mistaken for a legitimate failure.
var ErrNotImplemented = errors.New("ipc: not implemented (unimplemented CoreAPI stub)")
// UnimplementedCoreAPI is the gRPC Unimplemented*Server pattern applied to
// CoreAPI: embed it in a test double or adapter and override only the
// methods you actually exercise. Every method returns ErrNotImplemented, so
// a call that reaches an undeclared method fails loudly and specifically,
// rather than compiling to a silent no-op or nil-pointer panic. This
// replaces the old pattern of hand-writing all 30 no-op stubs per double —
// those were compiler-satisfying padding, not tests of anything.
type UnimplementedCoreAPI struct{}
var _ CoreAPI = UnimplementedCoreAPI{}
func (UnimplementedCoreAPI) WriteFact(ctx context.Context, req WriteFactReq) (int64, error) {
return 0, ErrNotImplemented
}
func (UnimplementedCoreAPI) LatestFact(ctx context.Context, key string) (Fact, error) {
return Fact{}, ErrNotImplemented
}
func (UnimplementedCoreAPI) LatestFactBySource(ctx context.Context, key, source string) (Fact, error) {
return Fact{}, ErrNotImplemented
}
func (UnimplementedCoreAPI) Since(ctx context.Context, key string, now time.Time) (time.Duration, error) {
return 0, ErrNotImplemented
}
func (UnimplementedCoreAPI) Presence(ctx context.Context) (Presence, error) {
return Presence{}, ErrNotImplemented
}
func (UnimplementedCoreAPI) CreateReminder(ctx context.Context, fire time.Time, payload, cron string) (int64, error) {
return 0, ErrNotImplemented
}
func (UnimplementedCoreAPI) MarkReminder(ctx context.Context, id int64, status string) error {
return ErrNotImplemented
}
func (UnimplementedCoreAPI) ListReminders(ctx context.Context, n int) ([]Reminder, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) RecordNudge(ctx context.Context, rule, channel, message string, ts time.Time) (int64, error) {
return 0, ErrNotImplemented
}
func (UnimplementedCoreAPI) ResolveNudge(ctx context.Context, id int64, outcome string, ts time.Time) error {
return ErrNotImplemented
}
func (UnimplementedCoreAPI) RecentOutcomes(ctx context.Context, rule string, n int) ([]string, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) RecentFacts(ctx context.Context, n int) ([]Fact, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) CalendarEvents(ctx context.Context, from, to time.Time) ([]Fact, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) RecentNudges(ctx context.Context, n int) ([]Nudge, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) WriteNote(ctx context.Context, ts time.Time, text string, embedding []float32, source string) (int64, error) {
return 0, ErrNotImplemented
}
func (UnimplementedCoreAPI) QueryNotes(ctx context.Context, embedding []float32, k int) ([]Note, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) RecentNotes(ctx context.Context, n int) ([]Note, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) ProposeTool(ctx context.Context, name, utterance, scope string, ts time.Time) (bool, error) {
return false, ErrNotImplemented
}
func (UnimplementedCoreAPI) EnableTool(ctx context.Context, name string, cmd []string, destructive bool, scope string, ts time.Time) error {
return ErrNotImplemented
}
func (UnimplementedCoreAPI) DisableTool(ctx context.Context, name string) error {
return ErrNotImplemented
}
func (UnimplementedCoreAPI) DeleteTool(ctx context.Context, name string) error {
return ErrNotImplemented
}
func (UnimplementedCoreAPI) CaptureTask(ctx context.Context, req CaptureTaskReq) (CaptureTaskResp, error) {
return CaptureTaskResp{}, ErrNotImplemented
}
func (UnimplementedCoreAPI) ListTasks(ctx context.Context, status string) ([]Task, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) SetTaskStatus(ctx context.Context, id int64, status string, ts time.Time) error {
return ErrNotImplemented
}
func (UnimplementedCoreAPI) ListProposedRoutines(ctx context.Context) ([]ProposedRoutine, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) DismissProposedRoutine(ctx context.Context, id int64) error {
return ErrNotImplemented
}
func (UnimplementedCoreAPI) AcceptProposedRoutine(ctx context.Context, id int64) error {
return ErrNotImplemented
}
func (UnimplementedCoreAPI) LookupTool(ctx context.Context, name string) (Tool, error) {
return Tool{}, ErrNotImplemented
}
func (UnimplementedCoreAPI) ListTools(ctx context.Context, status string) ([]Tool, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) RevertFact(ctx context.Context, key string) (int64, error) {
return 0, ErrNotImplemented
}
func (UnimplementedCoreAPI) TickTrace(ctx context.Context) (TickTrace, error) {
return TickTrace{}, ErrNotImplemented
}
func (UnimplementedCoreAPI) MorningStatus(ctx context.Context) ([]MorningRoutineStatus, error) {
return nil, ErrNotImplemented
}
func (UnimplementedCoreAPI) DayPlan(ctx context.Context) (DayPlan, error) {
return DayPlan{}, ErrNotImplemented
}
func (UnimplementedCoreAPI) Chat(ctx context.Context, text string) (string, error) {
return "", ErrNotImplemented
}
+4
View File
@@ -45,7 +45,11 @@ const (
MethodRevertFact Method = "revert_fact" MethodRevertFact Method = "revert_fact"
MethodTickTrace Method = "tick_trace" MethodTickTrace Method = "tick_trace"
MethodMorningStatus Method = "morning_status" MethodMorningStatus Method = "morning_status"
MethodDayPlan Method = "day_plan"
MethodChat Method = "chat" MethodChat Method = "chat"
MethodCaptureTask Method = "capture_task"
MethodListTasks Method = "list_tasks"
MethodSetTaskStatus Method = "set_task_status"
) )
// Request — one frame from module to core. Params is the JSON-encoded argument // Request — one frame from module to core. Params is the JSON-encoded argument
+116
View File
@@ -0,0 +1,116 @@
// Package kiwix reads a local Kiwix server (offline Wikipedia and friends).
//
// Why: the resident model is a 0.8B and invents facts. Letting her read a local
// article snippet beats letting her recall. Nothing here talks to the internet;
// the Kiwix server is on the same box.
//
// This is search only. Full articles are ~100KB of HTML, far too big for a 4096
// token context, so the unit of context is the search snippet (~500 chars).
package kiwix
import (
"context"
"encoding/xml"
"fmt"
"html"
"io"
"net/http"
"net/url"
"regexp"
"strconv"
"strings"
"time"
)
// Result is one search hit.
type Result struct {
Title string // article title, e.g. "Rayleigh scattering"
Path string // e.g. /content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering
Snippet string // plain text, tags stripped, entities decoded
WordCount int // 0 if the server did not say
}
// Client is a Kiwix HTTP client. Boring on purpose: no retries, no cache.
type Client struct {
base string
http *http.Client
}
// New makes a client for a Kiwix base URL like http://127.0.0.1:8034.
func New(baseURL string) *Client {
return &Client{
base: strings.TrimRight(baseURL, "/"),
http: &http.Client{Timeout: 10 * time.Second},
}
}
// Search runs a keyword search in one ZIM (book) and returns up to limit hits.
//
// Ranking is keyword based, not semantic: "Rayleigh scattering" finds the right
// article, "why is the sky blue" finds a TV episode. Pass keywords, not questions.
func (c *Client) Search(ctx context.Context, pattern, book string, limit int) ([]Result, error) {
if limit <= 0 {
limit = 5
}
q := url.Values{}
q.Set("pattern", pattern)
q.Set("books.name", book)
q.Set("format", "xml")
q.Set("pageLength", strconv.Itoa(limit))
req, err := http.NewRequestWithContext(ctx, http.MethodGet, c.base+"/search?"+q.Encode(), nil)
if err != nil {
return nil, err
}
resp, err := c.http.Do(req)
if err != nil {
return nil, err
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return nil, fmt.Errorf("kiwix search: http %d", resp.StatusCode)
}
return ParseSearchRSS(resp.Body)
}
// rss mirrors just the bits of the RSS 2.0 reply we use.
type rss struct {
Items []struct {
Title string `xml:"title"`
Link string `xml:"link"`
// innerxml keeps the <b> match markers so we can strip them ourselves.
Description struct {
Inner string `xml:",innerxml"`
} `xml:"description"`
WordCount string `xml:"wordCount"`
} `xml:"channel>item"`
}
var tagRE = regexp.MustCompile(`<[^>]*>`)
// ParseSearchRSS turns a Kiwix search reply into results. Exported so the parser
// is testable from a captured response, with no server running.
func ParseSearchRSS(r io.Reader) ([]Result, error) {
var doc rss
if err := xml.NewDecoder(r).Decode(&doc); err != nil {
return nil, fmt.Errorf("kiwix search: bad xml: %w", err)
}
out := make([]Result, 0, len(doc.Items))
for _, it := range doc.Items {
n, _ := strconv.Atoi(strings.ReplaceAll(it.WordCount, ",", ""))
out = append(out, Result{
Title: strings.TrimSpace(it.Title),
Path: strings.TrimSpace(it.Link),
Snippet: plainText(it.Description.Inner),
WordCount: n,
})
}
return out, nil
}
// plainText drops markup and decodes entities, leaving text a model can read.
func plainText(s string) string {
s = tagRE.ReplaceAllString(s, "")
s = html.UnescapeString(s)
return strings.TrimSpace(strings.Join(strings.Fields(s), " "))
}
+90
View File
@@ -0,0 +1,90 @@
package kiwix
import (
"context"
"os"
"strings"
"testing"
"time"
)
// A real reply from the live server, trimmed to two items.
const sampleRSS = `<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:opensearch="http://a9.com/-/spec/opensearch/1.1/">
<channel>
<title>Search: Rayleigh scattering</title>
<opensearch:totalResults>800</opensearch:totalResults>
<item>
<title>Rayleigh scattering</title>
<link>/content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering</link>
<description><b>Rayleigh</b> scattering causes the blue color of the sky &amp; yellow colors near the Sun.[1]</description>
<book><title>Wikipedia</title></book>
<wordCount>2,818</wordCount>
</item>
<item>
<title>HyperRayleigh scattering</title>
<link>/content/wikipedia_en_all_maxi_2026-02/Hyper%E2%80%93Rayleigh_scattering</link>
<description>...<b>Rayleigh</b> scattering" is a nonlinear optical counterpart.</description>
<book><title>Wikipedia</title></book>
<wordCount>914</wordCount>
</item>
</channel>
</rss>`
func TestParseSearchRSS(t *testing.T) {
got, err := ParseSearchRSS(strings.NewReader(sampleRSS))
if err != nil {
t.Fatalf("parse: %v", err)
}
if len(got) != 2 {
t.Fatalf("want 2 results, got %d", len(got))
}
if got[0].Title != "Rayleigh scattering" {
t.Errorf("title = %q", got[0].Title)
}
if got[0].Path != "/content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering" {
t.Errorf("path = %q", got[0].Path)
}
if got[0].WordCount != 2818 {
t.Errorf("wordCount = %d", got[0].WordCount)
}
want := "Rayleigh scattering causes the blue color of the sky & yellow colors near the Sun.[1]"
if got[0].Snippet != want {
t.Errorf("snippet = %q, want %q", got[0].Snippet, want)
}
if strings.Contains(got[1].Snippet, "<b>") {
t.Errorf("second snippet still has tags: %q", got[1].Snippet)
}
}
func TestParseSearchRSSBadXML(t *testing.T) {
if _, err := ParseSearchRSS(strings.NewReader("not xml at all")); err == nil {
t.Fatal("want an error on junk input")
}
}
// Opt-in: needs a live Kiwix server. CI has none.
// MAVEN_KIWIX_URL=http://127.0.0.1:8034 no_proxy=127.0.0.1,localhost go test -run Retrieval -v ./internal/kiwix/
func TestRetrievalEval(t *testing.T) {
base := os.Getenv("MAVEN_KIWIX_URL")
if base == "" {
t.Skip("set MAVEN_KIWIX_URL to run the retrieval eval")
}
noProxyLoopback(t)
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
defer cancel()
rep, err := RunRetrievalEval(ctx, New(base), 5)
if err != nil {
t.Fatalf("eval: %v", err)
}
// No pass bar on purpose: the number is the finding.
t.Log("\n" + rep.String() + rep.Detail())
}
// noProxyLoopback stops the box's SOCKS bridge from eating loopback requests.
func noProxyLoopback(t *testing.T) {
t.Setenv("no_proxy", "127.0.0.1,localhost")
t.Setenv("NO_PROXY", "127.0.0.1,localhost")
}
+63
View File
@@ -0,0 +1,63 @@
{
"name": "kiwix-knowledge-v1",
"book": "wikipedia_en_all_maxi_2026-02",
"note": "The 9 knowledge cases from internal/phraser/eval/talk_v1.json. Queries are hand-written English keywords on purpose: Kiwix ranks by keyword, not meaning, so a natural question fails. Writing them by hand separates 'retrieval is broken' from 'the model writes bad queries'.",
"cases": [
{
"id": "know-sky-blue",
"question": "почему небо синее?",
"query": "Rayleigh scattering sky blue",
"want_titles": ["Rayleigh scattering", "Diffuse sky radiation"]
},
{
"id": "know-boil-egg",
"question": "сколько варить яйцо вкрутую?",
"query": "boiled egg cooking",
"want_titles": ["Boiled egg", "Egg as food"]
},
{
"id": "know-ssd-vs-hdd",
"question": "чем ssd отличается от hdd?",
"query": "solid-state drive",
"want_titles": ["Solid-state drive", "Hard disk drive"]
},
{
"id": "know-cat-purr",
"question": "почему кошки мурчат?",
"query": "cat purr",
"want_titles": ["Purr", "Cat communication"]
},
{
"id": "know-hiccups",
"question": "как быстро избавиться от икоты?",
"query": "hiccup",
"want_titles": ["Hiccup"]
},
{
"id": "know-polite-form",
"question": "не могли бы вы объяснить, что такое vpn?",
"query": "virtual private network",
"want_titles": ["Virtual private network"]
},
{
"id": "know-dont-know",
"question": "как зовут моего соседа снизу?",
"query": "name of my downstairs neighbour",
"want_titles": [],
"expect_miss": true,
"note": "Unanswerable by design. Retrieval SHOULD find nothing useful. Counted as a hit only when nothing relevant comes back."
},
{
"id": "know-water-per-day",
"question": "сколько воды в день надо пить?",
"query": "human daily water requirement drinking",
"want_titles": ["Drinking water", "Water", "Dehydration", "Hydration"]
},
{
"id": "know-thunder-delay",
"question": "почему гром слышно позже молнии?",
"query": "thunder speed of sound lightning",
"want_titles": ["Thunder", "Lightning"]
}
]
}
+135
View File
@@ -0,0 +1,135 @@
package kiwix
// This scores retrieval alone: no LLM. For each general-knowledge question we
// hand-write English keywords and ask whether the article that would answer it
// comes back in the top N hits. If this score is low, reading Wikipedia cannot
// help the model no matter how good the prompt is.
//
// The unanswerable case (know-dont-know) is not scored. Whether the junk it
// returns is "nothing useful" is a human judgement, so the report just prints
// the titles and leaves the score to the 8 answerable cases.
import (
"context"
_ "embed"
"encoding/json"
"fmt"
"strings"
)
//go:embed knowledge_v1.json
var knowledgeFixtureJSON []byte
// EvalCase — one question with hand-written keywords.
type EvalCase struct {
ID string `json:"id"`
Question string `json:"question"`
Query string `json:"query"`
WantTitles []string `json:"want_titles"`
ExpectMiss bool `json:"expect_miss"`
}
type fixture struct {
Name string `json:"name"`
Book string `json:"book"`
Cases []EvalCase `json:"cases"`
}
// Outcome — what one case retrieved.
type Outcome struct {
Case EvalCase
Titles []string // titles of the top N hits, in rank order
Rank int // 1-based rank of the first wanted title, 0 if none
Err error
}
// Hit is true when a wanted title came back.
func (o Outcome) Hit() bool { return o.Rank > 0 }
// Report — the score plus per-case detail.
type Report struct {
Name string
Book string
TopN int
Scored int // answerable cases
Hits int
Errors int
Outcomes []Outcome
}
// Accuracy over the answerable cases.
func (r Report) Accuracy() float64 {
if r.Scored == 0 {
return 0
}
return float64(r.Hits) / float64(r.Scored)
}
// RunRetrievalEval searches for every fixture case.
func RunRetrievalEval(ctx context.Context, c *Client, topN int) (Report, error) {
var f fixture
if err := json.Unmarshal(knowledgeFixtureJSON, &f); err != nil {
return Report{}, err
}
rep := Report{Name: f.Name, Book: f.Book, TopN: topN}
for _, cs := range f.Cases {
res, err := c.Search(ctx, cs.Query, f.Book, topN)
o := Outcome{Case: cs, Err: err}
if err != nil {
rep.Errors++
}
for i, hit := range res {
o.Titles = append(o.Titles, hit.Title)
if o.Rank == 0 && matches(cs.WantTitles, hit.Title) {
o.Rank = i + 1
}
}
if !cs.ExpectMiss {
rep.Scored++
if o.Hit() {
rep.Hits++
}
}
rep.Outcomes = append(rep.Outcomes, o)
}
return rep, nil
}
func matches(want []string, title string) bool {
for _, w := range want {
if strings.EqualFold(strings.TrimSpace(title), w) {
return true
}
}
return false
}
// String — the headline number.
func (r Report) String() string {
var b strings.Builder
fmt.Fprintf(&b, "%s: %d/%d answerable questions retrieve a wanted article in top %d (%.1f%%), %d errors\n",
r.Name, r.Hits, r.Scored, r.TopN, 100*r.Accuracy(), r.Errors)
fmt.Fprintf(&b, " book: %s\n", r.Book)
return b.String()
}
// Detail — per case: what was asked, what was searched, what came back.
func (r Report) Detail() string {
var b strings.Builder
for _, o := range r.Outcomes {
mark := "MISS"
switch {
case o.Case.ExpectMiss:
mark = "n/a "
case o.Hit():
mark = fmt.Sprintf("hit@%d", o.Rank)
}
fmt.Fprintf(&b, " %-6s %-20s q=%q\n", mark, o.Case.ID, o.Case.Query)
if o.Err != nil {
fmt.Fprintf(&b, " error: %v\n", o.Err)
continue
}
fmt.Fprintf(&b, " got: %s\n", strings.Join(o.Titles, " | "))
}
return b.String()
}
+183
View File
@@ -0,0 +1,183 @@
package kiwix
// Turning a Russian question into an English Kiwix search.
//
// Kiwix ranks by keyword, not by meaning. "why is the sky blue" returns a TV
// episode; "Rayleigh scattering sky blue" returns the right article. So the
// model's job here is NOT translation — it is naming the English article the
// answer lives in.
//
// The output space is a handful of words, so it is worth locking down hard: a
// GBNF grammar for the shape, a tiny token cap, and a cleanup pass that throws
// away anything odd rather than handing junk to Kiwix.
import (
"context"
"encoding/json"
"fmt"
"strings"
"unicode"
"github.com/kami/maven/internal/llm"
)
// Completer — the LLM seam, so tests can fake it. *llm.Client satisfies it.
type Completer interface {
Complete(ctx context.Context, r llm.Req) (string, error)
}
// queryGrammar — one JSON object holding 1..6 keyword words. Latin letters,
// digits and hyphens only, so the model physically cannot answer the question
// or reply in Russian.
//
// Why the JSON wrapper: this model always thinks out loud and this llama-server
// build ignores the thinking switch (see ROUTING-EVAL-31-07-2026.md). A bare
// word-list grammar just captured the reasoning — every case came back as
// "Let me analyze this request carefully". Demanding JSON, like routeGrammar and
// responseGrammar already do, gives the reasoning nowhere to go.
const queryGrammar = `
root ::= "{" ws "\"query\"" ws ":" ws "\"" word (" " word){0,5} "\"" ws "}"
word ::= [A-Za-z0-9] [A-Za-z0-9-]{0,23}
ws ::= [ \t\n]*
`
// rewriteSystem — asks for search keywords, not an answer and not a translation.
const rewriteSystem = `You turn a question into a search query for English Wikipedia.
Rules:
- Output ONLY English search keywords. Never an answer, never an explanation.
- Do NOT translate the sentence. Name the thing the answer is about.
- The output must be a noun phrase, like a Wikipedia article title.
- Never use question words: no why, how, what, when, which, "how much",
"how long", "how to", "vs", "reason", "difference".
- 2 to 4 words.
Reply with JSON: {"query":"<keywords>"}
Good:
"почему листья желтеют осенью?" -> {"query":"leaf senescence autumn"}
"как работает микроволновка?" -> {"query":"microwave oven"}
"не могли бы вы объяснить, что такое блокчейн?" -> {"query":"blockchain"}
"сколько живут собаки?" -> {"query":"dog lifespan"}
"как избавиться от комаров в квартире?" -> {"query":"mosquito control"}
"чем чай отличается от кофе?" -> {"query":"tea"}
Only JSON, no explanation.`
// maxQueryTokens — the output is a few words plus the JSON wrapper. A tight cap
// is the cheapest guard against the model rambling into an answer.
const maxQueryTokens = 32
// Rewriter asks the resident model for English search keywords.
type Rewriter struct{ c Completer }
func NewRewriter(c Completer) *Rewriter { return &Rewriter{c: c} }
// Rewrite returns English keywords for a question in any language.
// It errors rather than returning something Kiwix should not see.
func (r *Rewriter) Rewrite(ctx context.Context, question string) (string, error) {
raw, err := r.c.Complete(ctx, llm.Req{
System: rewriteSystem,
User: strings.TrimSpace(question),
Grammar: queryGrammar,
MaxTokens: maxQueryTokens,
RepeatPenalty: 1.15,
})
if err != nil {
return "", err
}
return CleanQuery(unwrapJSON(raw))
}
// unwrapJSON pulls the query out of {"query":"..."}. If the reply is not that
// shape it is returned as-is, and CleanQuery decides whether it is usable.
func unwrapJSON(raw string) string {
s := strings.TrimSpace(raw)
if !strings.HasPrefix(s, "{") {
return s
}
var got struct{ Query string }
if err := json.Unmarshal([]byte(s), &got); err != nil {
return s
}
return got.Query
}
// maxQueryWords matches the grammar's bound. Anything longer is prose.
const maxQueryWords = 6
// CleanQuery checks and tidies whatever the model produced. The grammar makes
// bad output unlikely, not impossible (a server without grammar support, a
// different model), so this is the real gate in front of Kiwix.
//
// Exported so it can be tested without a model.
func CleanQuery(raw string) (string, error) {
s := strings.TrimSpace(raw)
// Models like to wrap answers in quotes. Drop surrounding ones.
s = strings.Trim(s, "\"'`")
// Keep the first line only: everything after it is prose.
if i := strings.IndexAny(s, "\r\n"); i >= 0 {
s = s[:i]
}
// Keep letters, digits, spaces and hyphens; anything else becomes a space.
var b strings.Builder
for _, ru := range s {
switch {
case unicode.IsLetter(ru) || unicode.IsDigit(ru) || ru == '-':
b.WriteRune(ru)
default:
b.WriteRune(' ')
}
}
words := strings.Fields(b.String())
if len(words) == 0 {
return "", fmt.Errorf("kiwix rewrite: empty query")
}
if len(words) > maxQueryWords {
return "", fmt.Errorf("kiwix rewrite: %d words, want at most %d (looks like prose)", len(words), maxQueryWords)
}
words = dropStopWords(words)
out := strings.Join(words, " ")
// The ZIMs are English. Non-Latin letters mean the model ignored the ask.
for _, ru := range out {
if unicode.IsLetter(ru) && !isLatin(ru) {
return "", fmt.Errorf("kiwix rewrite: query is not English: %q", out)
}
}
return out, nil
}
// stopWords — question words and filler. The model keeps writing question-shaped
// queries ("why is the sky blue", "how much water to drink daily") no matter how
// the prompt is worded, and Kiwix ranks on every word, so those words drag in
// song and episode titles. Dropping them in code is not a style preference: a
// keyword ranker gets nothing from them.
var stopWords = map[string]bool{
"a": true, "an": true, "the": true, "is": true, "are": true, "was": true,
"do": true, "does": true, "did": true, "to": true, "of": true, "in": true,
"on": true, "for": true, "and": true, "or": true, "my": true, "me": true,
"i": true, "it": true, "its": true, "be": true, "been": true, "get": true,
"how": true, "why": true, "what": true, "when": true, "which": true,
"who": true, "where": true, "much": true, "many": true, "long": true,
"vs": true, "than": true, "rid": true, "from": true, "about": true,
}
// dropStopWords removes filler, but never everything: if the query was nothing
// but stop words there is nothing better to search, so the original is kept and
// the caller sees whatever Kiwix makes of it.
func dropStopWords(words []string) []string {
kept := make([]string, 0, len(words))
for _, w := range words {
if !stopWords[strings.ToLower(w)] {
kept = append(kept, w)
}
}
if len(kept) == 0 {
return words
}
return kept
}
func isLatin(ru rune) bool {
return (ru >= 'a' && ru <= 'z') || (ru >= 'A' && ru <= 'Z')
}
+96
View File
@@ -0,0 +1,96 @@
package kiwix
// End-to-end score: Russian question -> model rewrite -> Kiwix search -> did a
// wanted article come back. Same 9 cases as the retrieval eval, so the two
// numbers are directly comparable: retrieval with hand-written keywords is the
// ceiling, this is what the model actually reaches.
import (
"context"
"encoding/json"
"fmt"
"strings"
)
// RewriteOutcome — one case, end to end.
type RewriteOutcome struct {
Outcome
ModelQuery string // what the model asked for ("" if it failed)
RewriteErr error
}
// RunRewriteEval rewrites every question with the model, then searches.
func RunRewriteEval(ctx context.Context, c *Client, rw *Rewriter, topN int) (RewriteReport, error) {
var f fixture
if err := json.Unmarshal(knowledgeFixtureJSON, &f); err != nil {
return RewriteReport{}, err
}
rep := RewriteReport{Report: Report{Name: f.Name + "-rewrite", Book: f.Book, TopN: topN}}
for _, cs := range f.Cases {
out := RewriteOutcome{Outcome: Outcome{Case: cs}}
q, err := rw.Rewrite(ctx, cs.Question)
out.ModelQuery, out.RewriteErr = q, err
if err == nil {
res, serr := c.Search(ctx, q, f.Book, topN)
out.Err = serr
for i, hit := range res {
out.Titles = append(out.Titles, hit.Title)
if out.Rank == 0 && matches(cs.WantTitles, hit.Title) {
out.Rank = i + 1
}
}
}
if out.RewriteErr != nil || out.Err != nil {
rep.Errors++
}
if !cs.ExpectMiss {
rep.Scored++
if out.Hit() {
rep.Hits++
}
}
rep.Cases = append(rep.Cases, out)
}
return rep, nil
}
// RewriteReport — the score plus per-case detail.
type RewriteReport struct {
Report
Cases []RewriteOutcome
}
// String — the headline number.
func (r RewriteReport) String() string {
return fmt.Sprintf("%s: %d/%d answerable questions retrieve a wanted article in top %d (%.1f%%), %d errors\n book: %s\n",
r.Name, r.Hits, r.Scored, r.TopN, 100*r.Accuracy(), r.Errors, r.Book)
}
// Detail — per case: hand-written query next to the model's, and what came back.
// The point is seeing WHERE the model's phrasing differs, not just the score.
func (r RewriteReport) Detail() string {
var b strings.Builder
for _, o := range r.Cases {
mark := "MISS"
switch {
case o.Case.ExpectMiss:
mark = "n/a "
case o.Hit():
mark = fmt.Sprintf("hit@%d", o.Rank)
}
fmt.Fprintf(&b, " %-6s %-20s\n", mark, o.Case.ID)
fmt.Fprintf(&b, " asked: %s\n", o.Case.Question)
fmt.Fprintf(&b, " hand: %q\n", o.Case.Query)
fmt.Fprintf(&b, " model: %q\n", o.ModelQuery)
if o.RewriteErr != nil {
fmt.Fprintf(&b, " rewrite rejected: %v\n", o.RewriteErr)
continue
}
if o.Err != nil {
fmt.Fprintf(&b, " search error: %v\n", o.Err)
continue
}
fmt.Fprintf(&b, " got: %s\n", strings.Join(o.Titles, " | "))
}
return b.String()
}
+33
View File
@@ -0,0 +1,33 @@
package kiwix
import (
"context"
"os"
"testing"
"time"
"github.com/kami/maven/internal/llm"
)
// Opt-in: needs a live Kiwix server AND a live llama-server.
// MAVEN_KIWIX_URL=http://127.0.0.1:8034 MAVEN_LLM_URL=http://127.0.0.1:18099 \
//
// no_proxy=127.0.0.1,localhost go test -run RewriteEval -v ./internal/kiwix/
func TestRewriteEval(t *testing.T) {
kbase, lbase := os.Getenv("MAVEN_KIWIX_URL"), os.Getenv("MAVEN_LLM_URL")
if kbase == "" || lbase == "" {
t.Skip("set MAVEN_KIWIX_URL and MAVEN_LLM_URL to run the rewrite eval")
}
noProxyLoopback(t)
ctx, cancel := context.WithTimeout(context.Background(), 15*time.Minute)
defer cancel()
rw := NewRewriter(llm.New(lbase, 3*time.Minute))
rep, err := RunRewriteEval(ctx, New(kbase), rw, 5)
if err != nil {
t.Fatalf("eval: %v", err)
}
// No pass bar on purpose: the number is the finding.
t.Log("\n" + rep.String() + rep.Detail())
}
+105
View File
@@ -0,0 +1,105 @@
package kiwix
import (
"context"
"testing"
"github.com/kami/maven/internal/llm"
)
// Bad model output must never reach Kiwix. No model needed for this.
func TestCleanQueryRejectsJunk(t *testing.T) {
bad := []struct{ name, raw string }{
{"empty", ""},
{"blank", " \n "},
{"russian came back", "почему небо синее"},
{"mixed russian", "sky синее scattering"},
{"full sentence", "The sky looks blue because of the scattering of sunlight by air molecules"},
{"prose with quotes", `Sure! Here is a good search query: "Rayleigh scattering", which explains it.`},
}
for _, c := range bad {
if got, err := CleanQuery(c.raw); err == nil {
t.Errorf("%s: want rejection, got %q", c.name, got)
}
}
}
func TestCleanQueryCleans(t *testing.T) {
ok := []struct{ raw, want string }{
{"Rayleigh scattering sky", "Rayleigh scattering sky"},
{" boiled egg cooking \n", "boiled egg cooking"},
{`"virtual private network"`, "virtual private network"},
{"solid-state drive", "solid-state drive"},
{"cat purr.", "cat purr"},
{"hiccup\nAlso: hiccough", "hiccup"},
// Question words are filler to a keyword ranker, so they go.
{"why is the sky blue", "sky blue"},
{"how much water to drink daily", "water drink daily"},
{"SSD vs HDD comparison", "SSD HDD comparison"},
// Nothing but filler: keep it rather than return nothing.
{"what is it", "what is it"},
}
for _, c := range ok {
got, err := CleanQuery(c.raw)
if err != nil {
t.Errorf("%q: %v", c.raw, err)
continue
}
if got != c.want {
t.Errorf("%q -> %q, want %q", c.raw, got, c.want)
}
}
}
type fakeCompleter struct {
out string
req llm.Req
}
func (f *fakeCompleter) Complete(_ context.Context, r llm.Req) (string, error) {
f.req = r
return f.out, nil
}
func TestRewriteConstrainsTheCall(t *testing.T) {
f := &fakeCompleter{out: `{"query":"Rayleigh scattering sky"}`}
got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?")
if err != nil {
t.Fatalf("rewrite: %v", err)
}
if got != "Rayleigh scattering sky" {
t.Errorf("query = %q", got)
}
if f.req.Grammar == "" {
t.Error("no grammar sent")
}
if f.req.MaxTokens == 0 || f.req.MaxTokens > 32 {
t.Errorf("max_tokens = %d, want a small cap", f.req.MaxTokens)
}
}
func TestRewriteRejectsBadModelOutput(t *testing.T) {
bad := []string{
`{"query":"почему небо синее"}`, // never translated
`{"query":""}`, // empty
`{"query":"the sky is blue because sunlight is scattered by air"}`, // an answer
// Note: a SHORT English prose fragment ("Let me analyze this request")
// is under the word cap and cannot be caught here. The grammar is what
// stops that one.
}
for _, out := range bad {
f := &fakeCompleter{out: out}
if got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?"); err == nil {
t.Errorf("%s: want rejection, got %q", out, got)
}
}
}
// A reply that is not the JSON shape but is still usable keywords should pass.
func TestRewriteFallsBackToPlainText(t *testing.T) {
f := &fakeCompleter{out: "Rayleigh scattering sky"}
got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?")
if err != nil || got != "Rayleigh scattering sky" {
t.Errorf("got %q, %v", got, err)
}
}
@@ -1,4 +1,4 @@
package eval package llm
import ( import (
"context" "context"
@@ -6,8 +6,20 @@ import (
"fmt" "fmt"
"net/http" "net/http"
"strings" "strings"
"time"
) )
// UnknownModel is the label to print when the server would not say what it has
// loaded. Deliberately ugly: an honest "unknown" is fine, a plausible-looking
// but wrong model name is the bug this whole file exists to prevent.
const UnknownModel = "unknown-model"
// llama-server is local, so never send this through a proxy: this box's
// http_proxy answers 503 for loopback, which would look like "server won't say
// which model it has" when the server is right there and fine.
// A Transport with no Proxy set bypasses http_proxy entirely.
var modelHTTP = &http.Client{Timeout: 10 * time.Second, Transport: &http.Transport{}}
// ModelID asks llama-server which model it has loaded, so a scoring run can // ModelID asks llama-server which model it has loaded, so a scoring run can
// label itself. Without this a bake-off between two models produces two tables // label itself. Without this a bake-off between two models produces two tables
// that look identical, and the operator has to remember which server was up. // that look identical, and the operator has to remember which server was up.
@@ -19,7 +31,7 @@ func ModelID(ctx context.Context, base string) (string, error) {
if err != nil { if err != nil {
return "", err return "", err
} }
resp, err := http.DefaultClient.Do(req) resp, err := modelHTTP.Do(req)
if err != nil { if err != nil {
return "", err return "", err
} }
@@ -38,14 +50,21 @@ func ModelID(ctx context.Context, base string) (string, error) {
if len(out.Data) == 0 { if len(out.Data) == 0 {
return "", fmt.Errorf("models: empty list") return "", fmt.Errorf("models: empty list")
} }
return shortModelID(out.Data[0].ID), nil short := shortModelID(out.Data[0].ID)
if short == "" {
// Server answered but the id field was missing or blank. Say so
// instead of handing back an empty label that reads as a real name.
return "", fmt.Errorf("models: no id in response")
}
return short, nil
} }
// shortModelID trims the path and the .gguf suffix — llama-server reports the // shortModelID trims the path and the .gguf suffix — llama-server reports the
// file name it was started with, which is too long for a table header. // file name it was started with, which is too long for a table header.
func shortModelID(id string) string { func shortModelID(id string) string {
id = strings.TrimSpace(id)
if i := strings.LastIndexAny(id, "/\\"); i >= 0 { if i := strings.LastIndexAny(id, "/\\"); i >= 0 {
id = id[i+1:] id = id[i+1:]
} }
return strings.TrimSuffix(id, ".gguf") return strings.TrimSpace(strings.TrimSuffix(id, ".gguf"))
} }
+64
View File
@@ -0,0 +1,64 @@
package llm
import (
"context"
"net/http"
"net/http/httptest"
"testing"
)
// The point of these tests: a wrong-but-plausible model label is the bug, so
// every path that cannot learn the real name must return an error instead of a
// guess. No llama-server needed — a stub server stands in.
func TestModelID(t *testing.T) {
cases := []struct {
name string
body string
code int
want string // "" ⇒ expect an error
}{
{"full path", `{"data":[{"id":"/mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf"}]}`, 200, "Qwen3.5-0.8B.Q4_K_M"},
{"bare name", `{"data":[{"id":"LFM2.5-1.2B"}]}`, 200, "LFM2.5-1.2B"},
{"empty list", `{"data":[]}`, 200, ""},
{"id missing", `{"data":[{}]}`, 200, ""},
{"id blank", `{"data":[{"id":" "}]}`, 200, ""},
{"server error", `nope`, 500, ""},
{"not json", `<html>`, 200, ""},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path != "/v1/models" {
t.Errorf("asked for %s, want /v1/models", r.URL.Path)
}
w.WriteHeader(c.code)
_, _ = w.Write([]byte(c.body))
}))
defer srv.Close()
got, err := ModelID(context.Background(), srv.URL+"/")
if c.want == "" {
if err == nil {
t.Fatalf("want an error, got label %q", got)
}
return
}
if err != nil {
t.Fatalf("ModelID: %v", err)
}
if got != c.want {
t.Errorf("got %q, want %q", got, c.want)
}
})
}
}
func TestModelIDUnreachable(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(http.ResponseWriter, *http.Request) {}))
url := srv.URL
srv.Close() // nothing listening now
if got, err := ModelID(context.Background(), url); err == nil {
t.Fatalf("want an error from a dead server, got label %q", got)
}
}
+48
View File
@@ -286,6 +286,54 @@ func TestRemindersStillHonourSnooze(t *testing.T) {
} }
} }
// ---------------------------- digest eligibility ------------------------------
// A Sev2 care candidate (break) suppressed for a genuine restraint reason is
// worth resurfacing later.
func TestDigestEligibleSev2SuppressedByRestraint(t *testing.T) {
for _, reason := range []string{"quiet_hours", "calendar_busy", "presence"} {
if !DigestEligible(Sev2, reason) {
t.Errorf("sev2 blocked by %q: want digest-eligible", reason)
}
}
}
// A Sev1 care candidate (water/meal) never digests — a biological timer
// nudge is stale by the time anyone could resurface it, so it just drops.
func TestDigestEligibleSev1NeverDigests(t *testing.T) {
for _, reason := range []string{"quiet_hours", "calendar_busy", "presence"} {
if DigestEligible(Sev1, reason) {
t.Errorf("sev1 blocked by %q: want drop, got digest-eligible", reason)
}
}
}
// Ops severities are never blocked by these reasons in practice (Gate only
// applies quiet_hours/calendar_busy/presence to care severities), but the
// boundary itself must refuse to digest a high severity even if asked —
// alarms bypass the gate and deliver now, unchanged, never delayed.
func TestDigestEligibleNeverDigestsHighSeverity(t *testing.T) {
for _, sev := range []Severity{Sev3, Sev4} {
for _, reason := range []string{"quiet_hours", "calendar_busy", "presence"} {
if DigestEligible(sev, reason) {
t.Errorf("sev%d blocked by %q: high severity must never digest", sev, reason)
}
}
}
}
// cooldown and snooze are not "suppression" in the digest sense — cooldown
// means it was already said recently, snooze means the user asked to not
// hear about it. Neither should resurface later just because the severity
// matches.
func TestDigestEligibleExcludesCooldownAndSnooze(t *testing.T) {
for _, reason := range []string{"cooldown", "snooze", "inert_no_data", "predicate", ""} {
if DigestEligible(Sev2, reason) {
t.Errorf("sev2 blocked by %q: should not be digest-eligible", reason)
}
}
}
// GAP — the gate reads State.SnoozeUntil, but the Gatherer hard-codes it to nil // GAP — the gate reads State.SnoozeUntil, but the Gatherer hard-codes it to nil
// (internal/loop/gather.go:153), so snooze is dead in the running daemon: the // (internal/loop/gather.go:153), so snooze is dead in the running daemon: the
// unit tests above pass while nothing can ever populate the map. This asserts // unit tests above pass while nothing can ever populate the map. This asserts
+30
View File
@@ -105,6 +105,36 @@ func Tick(s State, rules []Rule) *Candidate {
return fire return fire
} }
// DigestEligible decides digest-vs-drop for a care candidate the gate
// suppressed this tick (see ExplainGate's blockedBy). Pure — no I/O, no
// state, just the two facts that matter: why it was suppressed, and how
// insistent it was.
//
// Only genuine RESTRAINT blocks are eligible at all — quiet_hours,
// calendar_busy, presence(away). cooldown and snooze are not suppression in
// this sense: cooldown means "you already heard this recently" (resurfacing
// it later would be an actual repeat, not a rescue) and snooze is the user
// explicitly saying "not this" (digesting it anyway would defeat the ask).
// Ops severities (Sev3/4) never reach here — the gate never blocks them for
// these reasons in the first place (see Gate), and even if a future rule
// dropped Sev3+ into "care", digest still refuses them: alarms bypass the
// gate on purpose and must never be silently delayed into a bundle.
//
// Within care (Sev12), the boundary is severity itself: Sev1 (water, meal —
// biological timers with no "still relevant later" property; a water nudge
// from 3 hours into quiet hours is just wrong by morning) drops. Sev2
// (break — "you worked through a long stretch without a break while I
// couldn't reach you") is information that stays true and useful after the
// fact, so it digests.
func DigestEligible(sev Severity, blockedBy string) bool {
switch blockedBy {
case "quiet_hours", "calendar_busy", "presence":
default:
return false
}
return sev == Sev2
}
// ReminderDecision — a due reminder the daemon should deliver now. // ReminderDecision — a due reminder the daemon should deliver now.
// NOT gated by the universal Gate (per spec: "wake me 7" fires in quiet hours; // NOT gated by the universal Gate (per spec: "wake me 7" fires in quiet hours;
// that's the point). Snooze is the one part of restraint that still applies. // that's the point). Snooze is the one part of restraint that still applies.
+357
View File
@@ -0,0 +1,357 @@
// Package memeval is background memory evaluation (Vikunja #248,
// docs/plans/03-memory-evaluation.md).
//
// It lives beside internal/memory rather than inside it because
// internal/store imports internal/memory for the vector-store backend, and an
// evaluator has to read store.Fact / store.Note / store.Nudge — putting it in
// internal/memory would close that import cycle.
//
// Every so often Maven reads back her own recent memory — facts, notes, the
// nudges she sent — and asks the resident model what it notices: a habit that
// stopped, a gap, something worth saying later. What comes back is written as
// notes with source EvalNoteSource and nothing else happens. That restraint is
// the design, not an unfinished edge:
//
// - She does not speak here. There is no dispatcher, no channel, no nudge.
// An observation is a thought she wrote down; he reads it on /dash when he
// wants to. "Not a nag, not autonomous" (CLAUDE.md) is easy to violate with
// exactly this feature — an hourly loop with an LLM in it and permission to
// talk is a machine for generating interruptions — so the loop has no way
// to reach him at all. Turning observations into nudges is a separate
// decision with a separate opt-in, and it is deliberately NOT in this file.
// - She does not act. No reminder is created, no routine proposed, no fact
// written. The model's suggested_action is recorded as text inside the note
// and interpreted by nobody.
// - She says nothing about an empty store. No memory ⇒ no LLM call ⇒ no
// "observations" invented out of two facts. A 1.7B asked to find a pattern
// will always find one; the defence is not asking.
//
// Everything the evaluator writes is attributable: source is EvalNoteSource, so
// an inferred observation can never be mistaken for something he said, and the
// whole batch is one SQL delete away if the output turns out to be noise.
package memeval
import (
"context"
"encoding/json"
"fmt"
"sort"
"strings"
"time"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/store"
)
// EvalNoteSource — the source stamped on every note the evaluator writes.
// Same infer:* convention as the rest of the derived facts.
const EvalNoteSource = "infer:memory-eval"
// DefaultMinConfidence — an observation below this is dropped. The model is
// asked for its own confidence and small models are badly calibrated, so this
// is a coarse filter, not a probability: it exists to throw away the guesses
// the model itself hedged on.
const DefaultMinConfidence = 0.7
// DefaultMaxItems — how much recent memory goes into one evaluation, per
// store. 30 facts + 30 notes + 30 nudges is a few thousand tokens of the 4096
// context the resident Thinking model runs with, which leaves room for its
// reasoning tokens. Raising this trades reasoning room for history.
const DefaultMaxItems = 30
// MaxObservations — the model may return at most this many observations per
// evaluation, enforced by the grammar. A cap here is also a noise cap: an
// evaluation that "notices" ten things has noticed nothing.
const MaxObservations = 3
// Observation — one thing the evaluator noticed.
type Observation struct {
Text string `json:"observation"`
Conf float64 `json:"confidence"`
// Action — what the model thinks should happen with this. Recorded, never
// executed: see the file comment. One of "note", "propose", "notify".
Action string `json:"suggested_action"`
}
// Completer — the llama-server seam, same shape router.Completer uses so the
// one resident model serves this caller too.
type Completer interface {
Complete(ctx context.Context, r llm.Req) (string, error)
}
// Reader — the slice of the store an evaluation reads. Narrow on purpose: the
// evaluator gets recent memory and nothing else. No entity graph, no presence,
// no config facts.
type Reader interface {
RecentFacts(ctx context.Context, n int) ([]store.Fact, error)
RecentNotes(ctx context.Context, n int) ([]store.Note, error)
RecentNudges(ctx context.Context, n int) ([]store.Nudge, error)
}
// NoteWriter — where observations land. Embeddings are passed nil: an
// observation is written for a human to read on /dash, not to be recalled by
// similarity. Feeding LLM-generated text back into the RAG pool it was
// generated from is how a small model starts citing its own guesses as
// evidence.
type NoteWriter interface {
WriteNote(ctx context.Context, ts time.Time, text string, embedding []float32, source string) (int64, error)
}
// Config — evaluator tuning. Zero values are replaced by the Default*
// constants, so the zero Config is the sane one.
type Config struct {
MaxItems int
MinConfidence float64
// ContextBlock — the shared persona block (internal/persona), re-evaluated
// per call so the clock in it is current. Prepended to the system prompt so
// observations come out in Maven's voice: feminine self-reference, informal
// "ты". nil is allowed; the base prompt still carries the address rules.
ContextBlock func() string
}
// Evaluator reads recent memory and records what the model notices.
type Evaluator struct {
read Reader
write NoteWriter
llm Completer
cfg Config
}
func NewEvaluator(r Reader, w NoteWriter, c Completer, cfg Config) *Evaluator {
if cfg.MaxItems <= 0 {
cfg.MaxItems = DefaultMaxItems
}
if cfg.MinConfidence <= 0 {
cfg.MinConfidence = DefaultMinConfidence
}
return &Evaluator{read: r, write: w, llm: c, cfg: cfg}
}
// evalGrammar — GBNF pinning the reply to a bounded JSON array of fixed-shape
// observations. Same reasoning as the router's routeGrammar: the enum and the
// length bound are what stop a small model from drifting into free text or
// filling the token budget with one repeated field.
const evalGrammar = `
root ::= "[" ws (obs ("," ws obs){0,2})? ws "]"
obs ::= "{" ws "\"observation\"" ws ":" ws text "," ws "\"confidence\"" ws ":" ws conf "," ws "\"suggested_action\"" ws ":" ws act ws "}"
text ::= "\"" ([^"\\] | "\\" .){1,200} "\""
conf ::= "0" "." [0-9]{1,2} | "1" ("." "0")?
act ::= "\"note\"" | "\"propose\"" | "\"notify\""
ws ::= [ \t\n]*
`
// evalSystem — the evaluation prompt. Two things it insists on, both learned
// from the phraser: state the observation as something she noticed rather than
// an instruction, and say nothing when there is nothing (the model is given an
// explicit way to return an empty array, because a model with no exit returns
// filler).
const evalSystem = `Ты просматриваешь свою собственную память: недавние факты, заметки и напоминания, которые ты отправляла.
Найди то, что действительно заметно: привычка, которая прервалась; пробел в записях; повторяющаяся закономерность.
Правила:
- Отвечай ТОЛЬКО массивом JSON. Каждый элемент: {"observation": "...", "confidence": 0.0-1.0, "suggested_action": "note"|"propose"|"notify"}.
- observation короткая фраза по-русски о том, что ты заметила. О себе в женском роде ("я заметила"). К нему на "ты".
- Не выдумывай. Если в памяти нет ничего заметного, верни пустой массив [].
- Не давай советов и не приказывай. Ты замечаешь, а не требуешь.
- confidence насколько ты уверена, что это настоящая закономерность, а не совпадение.
- Максимум три наблюдения. Лучше одно точное, чем три общих.`
// Evaluate runs one evaluation and returns the observations it recorded.
//
// Returns (nil, nil) — not an error — for every ordinary "nothing to say"
// outcome: an empty store, an empty array from the model, everything below the
// confidence floor, or every observation already recorded earlier. Only a real
// read/LLM/write failure is an error, and the caller (a background ticker) logs
// it and waits for the next interval.
func (e *Evaluator) Evaluate(ctx context.Context, now time.Time) ([]Observation, error) {
snap, err := e.snapshot(ctx)
if err != nil {
return nil, err
}
if snap == "" {
return nil, nil // nothing recorded ⇒ nothing to notice, and no LLM call
}
raw, err := e.llm.Complete(ctx, llm.Req{
System: persona.Prepend(e.cfg.ContextBlock, evalSystem),
User: snap,
Grammar: evalGrammar,
MaxTokens: 512,
RepeatPenalty: 1.1,
})
if err != nil {
return nil, fmt.Errorf("memory eval: complete: %w", err)
}
obs, err := parseObservations(raw)
if err != nil {
return nil, fmt.Errorf("memory eval: parse %q: %w", truncate(raw, 120), err)
}
// Dedupe against what earlier evaluations already wrote. Without this an
// hourly loop over a slowly-changing store writes the same sentence every
// hour until /dash is nothing but the evaluator talking to itself.
seen, err := e.recordedTexts(ctx)
if err != nil {
return nil, err
}
var kept []Observation
for _, o := range obs {
o.Text = strings.TrimSpace(o.Text)
if o.Text == "" || o.Conf < e.cfg.MinConfidence {
continue
}
norm := normalizeObservation(o.Text)
if seen[norm] {
continue
}
seen[norm] = true
if _, err := e.write.WriteNote(ctx, now, formatNote(o), nil, EvalNoteSource); err != nil {
return kept, fmt.Errorf("memory eval: write note: %w", err)
}
kept = append(kept, o)
}
return kept, nil
}
// formatNote — the stored text. The suggested action is kept as a visible
// suffix rather than a column: it is the model's opinion about what to do next,
// and the only consumer is a human reading /dash.
func formatNote(o Observation) string {
if o.Action == "" {
return o.Text
}
return fmt.Sprintf("%s [%s]", o.Text, o.Action)
}
// recordedTexts — the normalized text of every observation earlier evaluations
// wrote, for dedupe. Reads a wider window than MaxItems because the point is to
// remember saying it, not to summarize it.
func (e *Evaluator) recordedTexts(ctx context.Context) (map[string]bool, error) {
notes, err := e.read.RecentNotes(ctx, 200)
if err != nil {
return nil, fmt.Errorf("memory eval: recent notes: %w", err)
}
seen := make(map[string]bool, len(notes))
for _, n := range notes {
if n.Source != EvalNoteSource {
continue
}
text := n.Text
// Strip the "[action]" suffix formatNote appended.
if i := strings.LastIndex(text, " ["); i > 0 && strings.HasSuffix(text, "]") {
text = text[:i]
}
seen[normalizeObservation(text)] = true
}
return seen, nil
}
// normalizeObservation — dedupe key. Case- and whitespace-insensitive, which
// catches the realistic repeat (the model re-emitting the same sentence with a
// different comma) without pretending to do semantic dedupe.
func normalizeObservation(s string) string {
return strings.Join(strings.Fields(strings.ToLower(s)), " ")
}
// snapshot renders recent memory as the user turn. Returns "" when there is
// nothing in any store — the caller treats that as "do not ask the model".
//
// Notes written by earlier evaluations are excluded. Feeding her own
// observations back in is how "я заметила, что ты не записывал еду" becomes
// evidence for noticing it again, three evaluations deep.
func (e *Evaluator) snapshot(ctx context.Context) (string, error) {
n := e.cfg.MaxItems
facts, err := e.read.RecentFacts(ctx, n)
if err != nil {
return "", fmt.Errorf("memory eval: recent facts: %w", err)
}
notes, err := e.read.RecentNotes(ctx, n)
if err != nil {
return "", fmt.Errorf("memory eval: recent notes: %w", err)
}
nudges, err := e.read.RecentNudges(ctx, n)
if err != nil {
return "", fmt.Errorf("memory eval: recent nudges: %w", err)
}
var b strings.Builder
wrote := false
if len(facts) > 0 {
b.WriteString("Факты:\n")
for _, f := range facts {
fmt.Fprintf(&b, "- %s %s=%s (%s)\n", f.Ts.Format("2006-01-02 15:04"), f.Key, truncate(f.Value, 80), f.Source)
wrote = true
}
}
own := 0
var noteLines []string
for _, nt := range notes {
if nt.Source == EvalNoteSource {
own++
continue
}
noteLines = append(noteLines, fmt.Sprintf("- %s %s\n", nt.Ts.Format("2006-01-02 15:04"), truncate(nt.Text, 160)))
}
if len(noteLines) > 0 {
b.WriteString("\nЗаметки:\n")
for _, l := range noteLines {
b.WriteString(l)
wrote = true
}
}
if len(nudges) > 0 {
b.WriteString("\nНапоминания, которые ты отправляла:\n")
for _, nd := range nudges {
outcome := nd.Outcome
if outcome == "" {
outcome = "?"
}
fmt.Fprintf(&b, "- %s %s → %s (%s)\n", nd.Ts.Format("2006-01-02 15:04"), nd.Rule, outcome, nd.Channel)
wrote = true
}
}
if !wrote {
// Only her own past observations, or nothing at all. Either way there is
// no new memory to evaluate.
return "", nil
}
b.WriteString("\nЧто ты замечаешь?")
return b.String(), nil
}
// parseObservations reads the model's array. Tolerates the leading/trailing
// prose a Thinking model sometimes emits around JSON by taking the outermost
// bracketed span, the same tolerance the router's parser has.
func parseObservations(raw string) ([]Observation, error) {
s := strings.TrimSpace(raw)
if i := strings.Index(s, "["); i >= 0 {
if j := strings.LastIndex(s, "]"); j > i {
s = s[i : j+1]
}
}
if s == "" {
return nil, nil
}
var obs []Observation
if err := json.Unmarshal([]byte(s), &obs); err != nil {
return nil, err
}
if len(obs) > MaxObservations {
// The grammar bounds this; a grammar-less server or a future prompt
// change must not be able to flood /dash.
sort.SliceStable(obs, func(i, j int) bool { return obs[i].Conf > obs[j].Conf })
obs = obs[:MaxObservations]
}
return obs, nil
}
func truncate(s string, n int) string {
r := []rune(s)
if len(r) <= n {
return s
}
return string(r[:n]) + "…"
}
+271
View File
@@ -0,0 +1,271 @@
package memeval
import (
"context"
"database/sql"
"errors"
"path/filepath"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/store"
)
// fakeLLM — canned replies, one per call, and a record of what it was asked.
type fakeLLM struct {
replies []string
calls []llm.Req
err error
}
func (f *fakeLLM) Complete(_ context.Context, r llm.Req) (string, error) {
f.calls = append(f.calls, r)
if f.err != nil {
return "", f.err
}
if len(f.replies) == 0 {
return "[]", nil
}
out := f.replies[0]
f.replies = f.replies[1:]
return out, nil
}
func newTestStore(t *testing.T) *store.Store {
t.Helper()
st, err := store.Open(context.Background(), filepath.Join(t.TempDir(), "memeval_test.db"))
if err != nil {
t.Fatalf("store.Open: %v", err)
}
t.Cleanup(func() { _ = st.Close() })
return st
}
func refNow() time.Time { return time.Date(2026, 8, 1, 9, 0, 0, 0, time.UTC) }
// seedMemory writes a little of everything the evaluator reads.
func seedMemory(t *testing.T, st *store.Store, ctx context.Context, now time.Time) {
t.Helper()
for i := 0; i < 3; i++ {
ts := now.Add(-time.Duration(i+1) * 24 * time.Hour)
if _, err := st.WriteFact(ctx, ts, store.KindSelf, "water_ml", "500", "tap:desk", 1.0, sql.NullInt64{}); err != nil {
t.Fatalf("write fact: %v", err)
}
}
if _, err := st.WriteNote(ctx, now.Add(-2*time.Hour), "купить корм для кота", nil, "tap:voice"); err != nil {
t.Fatalf("write note: %v", err)
}
if _, err := st.RecordNudge(ctx, "water", "voice", "пора выпить воды", now.Add(-time.Hour)); err != nil {
t.Fatalf("record nudge: %v", err)
}
}
// TestEvaluateEmptyStoreAsksNothing — the "shuts up when uncertain" floor. An
// empty store must not even reach the model: a small model asked to find a
// pattern in nothing will invent one.
func TestEvaluateEmptyStoreAsksNothing(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
f := &fakeLLM{}
ev := NewEvaluator(st, st, f, Config{})
obs, err := ev.Evaluate(ctx, refNow())
if err != nil {
t.Fatalf("Evaluate: %v", err)
}
if len(obs) != 0 {
t.Fatalf("observations on an empty store = %d, want 0", len(obs))
}
if len(f.calls) != 0 {
t.Fatalf("LLM called %d times on an empty store, want 0", len(f.calls))
}
}
// TestEvaluateWritesHighConfidenceObservations — the happy path. Confident
// observations are written as notes stamped infer:memory-eval, and the low
// ones are dropped.
func TestEvaluateWritesHighConfidenceObservations(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedMemory(t, st, ctx, now)
f := &fakeLLM{replies: []string{`[
{"observation":"ты три дня не записывал еду","confidence":0.9,"suggested_action":"notify"},
{"observation":"может быть, ты стал меньше пить воды","confidence":0.3,"suggested_action":"note"}
]`}}
ev := NewEvaluator(st, st, f, Config{})
obs, err := ev.Evaluate(ctx, now)
if err != nil {
t.Fatalf("Evaluate: %v", err)
}
if len(obs) != 1 {
t.Fatalf("kept %d observations, want 1 (the 0.3 one is below the floor): %+v", len(obs), obs)
}
if obs[0].Text != "ты три дня не записывал еду" {
t.Errorf("kept the wrong observation: %q", obs[0].Text)
}
notes, err := st.RecentNotes(ctx, 50)
if err != nil {
t.Fatalf("RecentNotes: %v", err)
}
var written []store.Note
for _, n := range notes {
if n.Source == EvalNoteSource {
written = append(written, n)
}
}
if len(written) != 1 {
t.Fatalf("notes with source %s = %d, want 1", EvalNoteSource, len(written))
}
if !strings.Contains(written[0].Text, "ты три дня не записывал еду") {
t.Errorf("note text = %q", written[0].Text)
}
if !strings.Contains(written[0].Text, "[notify]") {
t.Errorf("note text = %q, want the suggested action recorded", written[0].Text)
}
// The prompt must carry the memory it is evaluating, and must not carry a
// grammar-free request.
if len(f.calls) != 1 {
t.Fatalf("LLM calls = %d, want 1", len(f.calls))
}
if !strings.Contains(f.calls[0].User, "water_ml") {
t.Errorf("prompt does not mention the seeded facts:\n%s", f.calls[0].User)
}
if f.calls[0].Grammar == "" {
t.Error("evaluation ran without a grammar")
}
}
// TestEvaluateDeduplicatesAcrossRuns — the failure mode that would make this
// feature unusable: an hourly loop over a store that barely changes writing the
// same sentence every hour until /dash is nothing but the evaluator.
func TestEvaluateDeduplicatesAcrossRuns(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedMemory(t, st, ctx, now)
same := `[{"observation":"ты три дня не записывал еду","confidence":0.9,"suggested_action":"note"}]`
spaced := `[{"observation":"Ты три дня не записывал еду","confidence":0.95,"suggested_action":"note"}]`
f := &fakeLLM{replies: []string{same, same, spaced}}
ev := NewEvaluator(st, st, f, Config{})
for i := 0; i < 3; i++ {
if _, err := ev.Evaluate(ctx, now.Add(time.Duration(i)*time.Hour)); err != nil {
t.Fatalf("Evaluate %d: %v", i, err)
}
}
notes, err := st.RecentNotes(ctx, 50)
if err != nil {
t.Fatalf("RecentNotes: %v", err)
}
n := 0
for _, nt := range notes {
if nt.Source == EvalNoteSource {
n++
}
}
if n != 1 {
t.Fatalf("eval notes after three identical evaluations = %d, want 1", n)
}
}
// TestEvaluateIgnoresOwnNotes — her own observations must not become input.
// Otherwise "я заметила X" is evidence for noticing X again, three evaluations
// deep. With nothing but eval notes in the store there is no new memory, so the
// model is not asked at all.
func TestEvaluateIgnoresOwnNotes(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
if _, err := st.WriteNote(ctx, now.Add(-time.Hour), "я заметила, что ты мало пьёшь [note]", nil, EvalNoteSource); err != nil {
t.Fatalf("write note: %v", err)
}
f := &fakeLLM{}
ev := NewEvaluator(st, st, f, Config{})
obs, err := ev.Evaluate(ctx, now)
if err != nil {
t.Fatalf("Evaluate: %v", err)
}
if len(obs) != 0 || len(f.calls) != 0 {
t.Fatalf("observations=%d llm calls=%d, want 0/0 — own notes are not memory to evaluate", len(obs), len(f.calls))
}
}
// TestEvaluateEmptyArrayIsNotAnError — "nothing to say" is the expected outcome
// most of the time and must not be logged as a failure.
func TestEvaluateEmptyArrayIsNotAnError(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedMemory(t, st, ctx, now)
ev := NewEvaluator(st, st, &fakeLLM{replies: []string{"[]"}}, Config{})
obs, err := ev.Evaluate(ctx, now)
if err != nil {
t.Fatalf("Evaluate: %v", err)
}
if len(obs) != 0 {
t.Fatalf("observations = %d, want 0", len(obs))
}
}
// TestEvaluateLLMErrorIsReported — a broken llama-server is an error the caller
// logs; it must not silently write anything.
func TestEvaluateLLMErrorIsReported(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedMemory(t, st, ctx, now)
ev := NewEvaluator(st, st, &fakeLLM{err: errors.New("connection refused")}, Config{})
if _, err := ev.Evaluate(ctx, now); err == nil {
t.Fatal("want an error when the model is unreachable")
}
notes, err := st.RecentNotes(ctx, 50)
if err != nil {
t.Fatalf("RecentNotes: %v", err)
}
for _, n := range notes {
if n.Source == EvalNoteSource {
t.Fatalf("wrote a note despite an LLM failure: %q", n.Text)
}
}
}
// TestParseObservationsTolerantAndBounded — Thinking models wrap JSON in prose,
// and no reply may exceed MaxObservations even if the grammar is bypassed.
func TestParseObservationsTolerantAndBounded(t *testing.T) {
obs, err := parseObservations(`<think>hmm</think> вот: [{"observation":"a","confidence":0.9,"suggested_action":"note"}] всё`)
if err != nil {
t.Fatalf("parse: %v", err)
}
if len(obs) != 1 || obs[0].Text != "a" {
t.Fatalf("got %+v, want one observation 'a'", obs)
}
var b strings.Builder
b.WriteString("[")
for i := 0; i < MaxObservations+3; i++ {
if i > 0 {
b.WriteString(",")
}
b.WriteString(`{"observation":"x","confidence":0.5,"suggested_action":"note"}`)
}
b.WriteString("]")
obs, err = parseObservations(b.String())
if err != nil {
t.Fatalf("parse: %v", err)
}
if len(obs) != MaxObservations {
t.Fatalf("parsed %d observations, want the %d cap", len(obs), MaxObservations)
}
}
+254
View File
@@ -0,0 +1,254 @@
package memory
import (
"fmt"
"sort"
"strings"
"time"
)
// Behavioural memory — "what do I usually do?" (Vikunja #254).
//
// The profile is COUNTED, not generated. docs/plans/09-behavioral-memory.md
// asks for an LLM to write a behaviour profile daily and store it as a fact;
// this does not do that, on purpose. A 1.7B asked to summarise a year of habits
// will produce fluent claims about the owner's life that no row in the store
// supports, and a wrong claim about him is the most expensive kind of wrong
// maven can be. Counting distinct days per weekday is verifiable, cheap enough
// to run on the question, and cannot invent a habit he does not have.
//
// Recomputed on read rather than cached as a fact for the same reason the store
// is append-only: a cached profile can disagree with the rows it came from, and
// then there are two truths. The plan's step 5 ("profile updates on fact write")
// exists to keep a cache fresh; there is no cache, so a new fact is already in
// the next answer.
//
// It is also read-only and unprompted-free. The plan's step 4 — a morning
// dispatcher nudge proposing the day — is deliberately NOT here: maven is not a
// nag, and proposing plans at 08:00 every day is the definition of one. Pattern
// inference that leads to a routine the owner accepts already exists in
// internal/pattern with the proposal queue on /routines; that is the sanctioned
// path from "she noticed" to "she acts", and it goes through him.
// Observation — one thing the owner was recorded doing, reduced to what a habit
// needs: when, and what. Facts arrive as store/ipc rows; the caller maps them
// so this package stays free of both.
type Observation struct {
At time.Time
Key string
Kind string // "self" | "env" | "config"
}
// Activity — one recurring thing, as counted. Days is the number of DISTINCT
// days it was observed on, which is the number that decides whether something
// is a habit; Count can be inflated by one busy day.
//
// TypicalAt is the median time of day it happens at, rounded to the minute — a
// median and not a mean, so one 03:00 outlier does not move "он обычно пьёт
// воду утром" into the night.
type Activity struct {
Key string
Days int
Count int
TypicalAt time.Duration
}
// Profile — the counted behaviour model. Weekly holds the activities that
// recur on a given weekday, Overall the ones that recur at all.
type Profile struct {
Since time.Time
Until time.Time
Weekly map[time.Weekday][]Activity
All []Activity
}
// MinHabitDays — how many distinct days an activity must appear on before maven
// will call it usual. Two is the smallest number that can distinguish a habit
// from a one-off; below that she says she does not know yet, which is true.
const MinHabitDays = 2
// nonBehaviouralKeyPrefixes — keys that are machinery or one-shot records, not
// behaviour. Calendar events carry the day in the key so they can never repeat;
// cooldown and quiet rows are maven's own tuning state, not his habits.
var nonBehaviouralKeyPrefixes = []string{
"calendar_event_",
"cooldown:",
"quiet",
"behavior_profile",
}
// BuildProfile counts habits out of observations. now bounds the window's upper
// end and supplies the location every day boundary is taken in — a habit is
// "on Tuesdays" in the owner's timezone or it is nothing.
//
// Only self-facts count. An env row is the world (weather, a relayed meeting),
// and a config row is maven's own state; neither says anything about what he
// usually does.
func BuildProfile(obs []Observation, now time.Time) Profile {
loc := now.Location()
p := Profile{Until: now, Weekly: map[time.Weekday][]Activity{}}
type bucket struct {
days map[string]struct{}
count int
mins []int
}
// key → bucket, and (weekday, key) → bucket.
all := map[string]*bucket{}
weekly := map[time.Weekday]map[string]*bucket{}
for _, o := range obs {
if o.Kind != "self" {
continue
}
key := strings.TrimSpace(o.Key)
if key == "" || nonBehavioural(key) {
continue
}
at := o.At.In(loc)
if at.IsZero() || at.After(now) {
continue
}
if p.Since.IsZero() || at.Before(p.Since) {
p.Since = at
}
day := at.Format("2006-01-02")
minute := at.Hour()*60 + at.Minute()
bump := func(m map[string]*bucket) {
b := m[key]
if b == nil {
b = &bucket{days: map[string]struct{}{}}
m[key] = b
}
b.days[day] = struct{}{}
b.count++
b.mins = append(b.mins, minute)
}
bump(all)
wd := at.Weekday()
if weekly[wd] == nil {
weekly[wd] = map[string]*bucket{}
}
bump(weekly[wd])
}
harvest := func(m map[string]*bucket) []Activity {
var out []Activity
for key, b := range m {
if len(b.days) < MinHabitDays {
continue
}
out = append(out, Activity{
Key: key,
Days: len(b.days),
Count: b.count,
TypicalAt: time.Duration(medianInt(b.mins)) * time.Minute,
})
}
// Most-established first, then earliest in the day, then by key so the
// same history always reads back the same way.
sort.Slice(out, func(i, j int) bool {
if out[i].Days != out[j].Days {
return out[i].Days > out[j].Days
}
if out[i].TypicalAt != out[j].TypicalAt {
return out[i].TypicalAt < out[j].TypicalAt
}
return out[i].Key < out[j].Key
})
return out
}
p.All = harvest(all)
for wd, m := range weekly {
if acts := harvest(m); len(acts) > 0 {
p.Weekly[wd] = acts
}
}
return p
}
func nonBehavioural(key string) bool {
for _, p := range nonBehaviouralKeyPrefixes {
if strings.HasPrefix(key, p) {
return true
}
}
return false
}
// medianInt — the middle value, averaging the two middles on an even count.
func medianInt(xs []int) int {
if len(xs) == 0 {
return 0
}
s := make([]int, len(xs))
copy(s, xs)
sort.Ints(s)
mid := len(s) / 2
if len(s)%2 == 1 {
return s[mid]
}
return (s[mid-1] + s[mid]) / 2
}
// weekdayRU — accusative, as "по вторникам" and "в среду" both need it read
// back. Index is time.Weekday.
var weekdayRU = [...]string{"воскресеньям", "понедельникам", "вторникам", "средам", "четвергам", "пятницам", "субботам"}
// activityRU glosses the loop's known fact keys. An unknown key is read back
// verbatim: it is what the store holds, and inventing a Russian phrase for a key
// maven does not recognise would be putting words in his mouth.
var activityRU = map[string]string{
"water": "пьёшь воду",
"meal": "ешь",
"sleep": "спишь",
"break": "делаешь перерыв",
"shower": "принимаешь душ",
"walk": "гуляешь",
"pills": "пьёшь витамины",
"workout": "тренируешься",
}
// FormatWeekdayRU reads back what he usually does on a given weekday.
// Second person singular and informal, as she speaks TO him.
func (p Profile) FormatWeekdayRU(wd time.Weekday) string {
acts := p.Weekly[wd]
day := weekdayRU[int(wd)%7]
if len(acts) == 0 {
return fmt.Sprintf("по %s у меня пока нет ничего постоянного.", day)
}
return fmt.Sprintf("по %s ты обычно %s.", day, joinActivities(acts))
}
// FormatOverallRU reads back the habits that hold across the whole week.
func (p Profile) FormatOverallRU() string {
if len(p.All) == 0 {
return "я ещё не набрала достаточно записей, чтобы говорить о привычках."
}
return fmt.Sprintf("обычно ты %s.", joinActivities(p.All))
}
// maxRecited bounds a spoken profile. A list of fifteen habits read aloud is
// not an answer; the most established few are.
const maxRecited = 5
func joinActivities(acts []Activity) string {
if len(acts) > maxRecited {
acts = acts[:maxRecited]
}
parts := make([]string, len(acts))
for i, a := range acts {
gloss, ok := activityRU[a.Key]
if !ok {
gloss = a.Key
}
parts[i] = fmt.Sprintf("%s около %02d:%02d", gloss,
int(a.TypicalAt.Hours()), int(a.TypicalAt.Minutes())%60)
}
if len(parts) == 1 {
return parts[0]
}
return strings.Join(parts[:len(parts)-1], ", ") + " и " + parts[len(parts)-1]
}
+163
View File
@@ -0,0 +1,163 @@
package memory
import (
"strings"
"testing"
"time"
"unicode"
)
// habitHistory — n weeks of the same weekday, at the given local time.
func habitHistory(key string, wd time.Weekday, hh, mm, weeks int, from time.Time) []Observation {
var out []Observation
d := from
for d.Weekday() != wd {
d = d.AddDate(0, 0, -1)
}
for i := 0; i < weeks; i++ {
day := d.AddDate(0, 0, -7*i)
out = append(out, Observation{
At: time.Date(day.Year(), day.Month(), day.Day(), hh, mm, 0, 0, from.Location()),
Key: key,
Kind: "self",
})
}
return out
}
func behaviorNow() time.Time {
// A Monday, so "по вторникам" is a past weekday and not today.
return time.Date(2026, 8, 3, 20, 0, 0, 0, time.UTC)
}
func TestBuildProfileCountsWeekdayHabits(t *testing.T) {
now := behaviorNow()
obs := append(
habitHistory("workout", time.Tuesday, 19, 0, 4, now),
habitHistory("water", time.Tuesday, 9, 0, 3, now)...,
)
p := BuildProfile(obs, now)
tue := p.Weekly[time.Tuesday]
if len(tue) != 2 {
t.Fatalf("got %d tuesday activities, want 2: %+v", len(tue), tue)
}
// Most-established first.
if tue[0].Key != "workout" || tue[0].Days != 4 {
t.Errorf("first = %+v, want workout on 4 days", tue[0])
}
if tue[0].TypicalAt != 19*time.Hour {
t.Errorf("typical at %v, want 19:00", tue[0].TypicalAt)
}
if len(p.Weekly[time.Wednesday]) != 0 {
t.Errorf("wednesday must be empty: %+v", p.Weekly[time.Wednesday])
}
if len(p.All) != 2 {
t.Errorf("the week-wide list should hold both: %+v", p.All)
}
}
// A one-off is not a habit. Saying "ты обычно X" off a single row is a
// confidently wrong claim about his life.
func TestBuildProfileNeedsMoreThanOneDay(t *testing.T) {
now := behaviorNow()
obs := habitHistory("workout", time.Tuesday, 19, 0, 1, now)
// Three rows, same day — a busy Tuesday, not a habit.
obs = append(obs, Observation{At: obs[0].At.Add(time.Hour), Key: "workout", Kind: "self"})
obs = append(obs, Observation{At: obs[0].At.Add(2 * time.Hour), Key: "workout", Kind: "self"})
p := BuildProfile(obs, now)
if len(p.All) != 0 || len(p.Weekly) != 0 {
t.Fatalf("one day of rows must produce no habit: %+v / %+v", p.All, p.Weekly)
}
if got := p.FormatOverallRU(); !strings.Contains(got, "не набрала достаточно") {
t.Errorf("empty profile reads %q", got)
}
}
// Only self-facts describe him. Env rows are the world and config rows are
// maven's own tuning state; counting either as a habit would be a category
// error the owner would then be told about.
func TestBuildProfileIgnoresNonSelfAndMachineryKeys(t *testing.T) {
now := behaviorNow()
var obs []Observation
for _, o := range habitHistory("water", time.Tuesday, 9, 0, 3, now) {
o.Kind = "env"
obs = append(obs, o)
}
for _, o := range habitHistory("cooldown:water", time.Tuesday, 9, 0, 3, now) {
obs = append(obs, o) // kind=self, but a machinery key
}
for _, o := range habitHistory("calendar_event_20260804_standup", time.Tuesday, 10, 0, 3, now) {
obs = append(obs, o)
}
if p := BuildProfile(obs, now); len(p.All) != 0 {
t.Fatalf("nothing here is a habit of his: %+v", p.All)
}
}
// The median, not the mean: one 03:00 outlier must not move a morning habit
// into the night.
func TestBuildProfileTypicalTimeIsMedian(t *testing.T) {
now := behaviorNow()
obs := habitHistory("water", time.Tuesday, 9, 0, 4, now)
obs = append(obs, Observation{At: obs[0].At.AddDate(0, 0, -28).Add(-6 * time.Hour), Key: "water", Kind: "self"})
p := BuildProfile(obs, now)
if len(p.All) != 1 {
t.Fatalf("got %+v", p.All)
}
if p.All[0].TypicalAt != 9*time.Hour {
t.Errorf("typical at %v, want 09:00 despite the outlier", p.All[0].TypicalAt)
}
}
func TestProfileFormatRUPersona(t *testing.T) {
now := behaviorNow()
obs := append(
habitHistory("workout", time.Tuesday, 19, 0, 4, now),
habitHistory("water", time.Tuesday, 9, 5, 3, now)...,
)
p := BuildProfile(obs, now)
got := p.FormatWeekdayRU(time.Tuesday)
want := "по вторникам ты обычно тренируешься около 19:00 и пьёшь воду около 09:05."
if got != want {
t.Errorf("got %q\nwant %q", got, want)
}
if empty := p.FormatWeekdayRU(time.Thursday); !strings.Contains(empty, "ничего постоянного") {
t.Errorf("an unknown weekday reads %q", empty)
}
// Persona: she addresses him informally, never in the masculine about
// herself, and never with a pet name.
for _, s := range []string{got, p.FormatOverallRU(), p.FormatWeekdayRU(time.Thursday)} {
// Whole words: "ничего" contains "его", and a substring test would
// call a correct sentence a persona violation.
for _, tok := range strings.FieldsFunc(strings.ToLower(s), func(r rune) bool {
return !unicode.IsLetter(r)
}) {
switch tok {
case "рад", "понял", "вы", "ваш", "ваши", "милый", "дорогой", "он", "его":
t.Errorf("%q uses %q", s, tok)
}
}
}
}
// An unrecognised key is read back verbatim rather than glossed into something
// maven made up.
func TestProfileUnknownKeyReadBackVerbatim(t *testing.T) {
now := behaviorNow()
p := BuildProfile(habitHistory("починил кран", time.Tuesday, 12, 0, 2, now), now)
if got := p.FormatOverallRU(); !strings.Contains(got, "починил кран") {
t.Errorf("got %q", got)
}
}
// A future-dated row is a clock problem, not a habit.
func TestBuildProfileIgnoresFutureRows(t *testing.T) {
now := behaviorNow()
obs := habitHistory("water", time.Tuesday, 9, 0, 3, now.AddDate(0, 2, 0))
if p := BuildProfile(obs, now); len(p.All) != 0 {
t.Fatalf("future rows counted: %+v", p.All)
}
}
+165
View File
@@ -0,0 +1,165 @@
package morning
import (
"fmt"
"sort"
"strings"
"time"
"github.com/kami/maven/internal/store"
)
// The day plan (Vikunja #128).
//
// It lives here, with the morning routine engine, because it is the same
// question asked at a different scale: the routine knows what is still missing
// from a window, the plan knows what the whole day holds. A parallel system
// would have to re-read the same facts and re-decide what "today" means.
//
// It is pure, like the rest of this package: the daemon reads the calendar,
// the reminders and the checklist facts, and BuildPlan puts them in order.
//
// It is also NOT a nag. A plan she can recite when asked is the whole feature;
// nothing here fires, schedules or announces. Unprompted delivery stays with
// the existing morning nudge and the dispatcher's policy.
// PlanKind — where a plan line came from. It survives into the reply and the
// web view because the three read differently: an event is something happening
// to the owner, a reminder is something he asked for, a checklist item is
// something he has not done yet.
type PlanKind string
const (
PlanEvent PlanKind = "event"
PlanReminder PlanKind = "reminder"
PlanChecklist PlanKind = "checklist"
)
// PlanEntry — one timed thing on the day, as the daemon read it out of the
// store. Text is rendered verbatim; the plan does not rephrase.
//
// Uncertain marks provenance below a full-confidence read — a work meeting
// relayed off a phone notification (#126). It travels through to the reply so
// she hedges instead of reciting a guess as fact.
type PlanEntry struct {
At time.Time
Text string
Kind PlanKind
Uncertain bool
}
// Plan — the ordered day. Date is the calendar day it describes.
type Plan struct {
Date time.Time
Items []PlanEntry
}
// BuildPlan orders everything known about the day Now falls on: calendar
// events, pending reminders, and one line per morning routine that still has
// unfinished items.
//
// Entries outside that calendar day are dropped — a plan for today that
// includes tomorrow's meeting is wrong in a way that is worse than terse.
// Ordering is by time, then by kind, then by text, so the same day always reads
// the same way.
func BuildPlan(routines []Routine, facts map[string]store.Fact, events, reminders []PlanEntry, now time.Time) Plan {
y, m, d := now.Date()
dayStart := time.Date(y, m, d, 0, 0, 0, 0, now.Location())
dayEnd := dayStart.AddDate(0, 0, 1)
p := Plan{Date: dayStart}
for _, group := range [][]PlanEntry{events, reminders} {
for _, e := range group {
at := e.At.In(now.Location())
if at.Before(dayStart) || !at.Before(dayEnd) {
continue
}
if strings.TrimSpace(e.Text) == "" {
continue
}
e.At = at
p.Items = append(p.Items, e)
}
}
p.Items = append(p.Items, checklistEntries(routines, facts, now)...)
sort.SliceStable(p.Items, func(i, j int) bool {
a, b := p.Items[i], p.Items[j]
if !a.At.Equal(b.At) {
return a.At.Before(b.At)
}
if a.Kind != b.Kind {
return a.Kind < b.Kind
}
return a.Text < b.Text
})
return p
}
// checklistEntries renders one line per routine with work left in it, placed at
// the routine's nudge time — where the checklist actually matters in the day.
// A routine that does not apply today, is not in its window, or is already
// complete contributes nothing: the plan says what is left, not what was done.
func checklistEntries(routines []Routine, facts map[string]store.Fact, now time.Time) []PlanEntry {
var out []PlanEntry
for _, r := range routines {
st := Evaluate(r, facts, now)
if !st.Active || len(st.Missing) == 0 {
continue
}
labels := make([]string, 0, len(st.Missing))
for _, it := range st.Missing {
label := it.Label
if label == "" {
label = it.Key
}
labels = append(labels, label)
}
at := r.NudgeAt
if at == "" {
at = r.WindowEnd
}
when, ok := todayAt(at, now)
if !ok {
continue
}
out = append(out, PlanEntry{
At: when,
Text: fmt.Sprintf("%s — осталось: %s", r.Name, strings.Join(labels, ", ")),
Kind: PlanChecklist,
})
}
return out
}
// After returns the part of the plan that has not happened yet — the answer to
// "что дальше?" as opposed to "какие планы на сегодня?". The Date is kept, so an
// empty result still knows which day it is empty for.
func (p Plan) After(now time.Time) Plan {
out := Plan{Date: p.Date}
for _, it := range p.Items {
if it.At.Before(now) {
continue
}
out.Items = append(out.Items, it)
}
return out
}
// FormatRU renders the plan as maven says it. Feminine self-reference,
// informal address, no pet names — and no exhortation: she reads the day back,
// she does not tell him to get on with it.
func (p Plan) FormatRU() string {
if len(p.Items) == 0 {
return fmt.Sprintf("на %s ничего не запланировано.", p.Date.Format("02.01.2006"))
}
parts := make([]string, len(p.Items))
for i, it := range p.Items {
line := fmt.Sprintf("%s — %s", it.At.Format("15:04"), it.Text)
if it.Uncertain {
line = "похоже, " + line
}
parts[i] = line
}
return fmt.Sprintf("план на %s: %s.", p.Date.Format("02.01.2006"), strings.Join(parts, "; "))
}
+169
View File
@@ -0,0 +1,169 @@
package morning
import (
"strings"
"testing"
"time"
"github.com/kami/maven/internal/store"
)
// planAt is at() for the plan tests' day (2026-08-03, a Monday); the existing
// at() in morning_test.go is pinned to a different date.
func planAt(now time.Time, hh, mm int) time.Time {
y, m, d := now.Date()
return time.Date(y, m, d, hh, mm, 0, 0, now.Location())
}
func planFixture(t *testing.T) (Plan, time.Time) {
t.Helper()
now := time.Date(2026, 8, 3, 9, 0, 0, 0, time.UTC)
routines := []Routine{{
Name: "утро",
WindowStart: "07:00",
WindowEnd: "11:00",
NudgeAt: "10:30",
Items: []Item{
{Key: "water", FactKey: "drank_water", Label: "выпить воды"},
{Key: "pills", FactKey: "took_pills", Label: "витамины"},
},
}}
facts := map[string]store.Fact{
"drank_water": {Ts: planAt(now, 8, 0)},
}
events := []PlanEntry{
{At: planAt(now, 14, 0), Text: "Планёрка @ 14:00-14:30", Kind: PlanEvent, Uncertain: true},
{At: planAt(now, 10, 0), Text: "Standup @ 10:00-10:30", Kind: PlanEvent},
}
reminders := []PlanEntry{
{At: planAt(now, 18, 30), Text: "позвонить маме", Kind: PlanReminder},
}
return BuildPlan(routines, facts, events, reminders, now), now
}
func TestBuildPlanOrdersTheDay(t *testing.T) {
p, now := planFixture(t)
if !p.Date.Equal(planAt(now, 0, 0)) {
t.Errorf("Date = %v, want midnight of now's day", p.Date)
}
want := []struct {
hhmm string
kind PlanKind
}{
{"10:00", PlanEvent},
{"10:30", PlanChecklist},
{"14:00", PlanEvent},
{"18:30", PlanReminder},
}
if len(p.Items) != len(want) {
t.Fatalf("got %d items, want %d: %+v", len(p.Items), len(want), p.Items)
}
for i, w := range want {
if got := p.Items[i].At.Format("15:04"); got != w.hhmm {
t.Errorf("item %d at %s, want %s", i, got, w.hhmm)
}
if p.Items[i].Kind != w.kind {
t.Errorf("item %d kind %q, want %q", i, p.Items[i].Kind, w.kind)
}
}
}
// The checklist line says what is LEFT. An item already evidenced today must
// not be read back as outstanding.
func TestBuildPlanChecklistListsOnlyMissing(t *testing.T) {
p, _ := planFixture(t)
var line string
for _, it := range p.Items {
if it.Kind == PlanChecklist {
line = it.Text
}
}
if line == "" {
t.Fatal("no checklist line in the plan")
}
if !strings.Contains(line, "витамины") {
t.Errorf("missing item not listed: %q", line)
}
if strings.Contains(line, "выпить воды") {
t.Errorf("a completed item must not be read back as outstanding: %q", line)
}
if !strings.HasPrefix(line, "утро — осталось:") {
t.Errorf("line = %q", line)
}
}
func TestBuildPlanSkipsCompleteAndInactiveRoutines(t *testing.T) {
now := time.Date(2026, 8, 3, 9, 0, 0, 0, time.UTC)
routines := []Routine{
{
Name: "утро", WindowStart: "07:00", WindowEnd: "11:00",
Items: []Item{{Key: "water", FactKey: "drank_water", Label: "выпить воды"}},
},
{
// Not in its window at 09:00.
Name: "вечер", WindowStart: "20:00", WindowEnd: "23:00",
Items: []Item{{Key: "walk", FactKey: "walked", Label: "прогулка"}},
},
}
facts := map[string]store.Fact{"drank_water": {Ts: planAt(now, 8, 0)}}
p := BuildPlan(routines, facts, nil, nil, now)
if len(p.Items) != 0 {
t.Fatalf("a complete routine and an out-of-window one must contribute nothing: %+v", p.Items)
}
if got, want := p.FormatRU(), "на 03.08.2026 ничего не запланировано."; got != want {
t.Errorf("got %q\nwant %q", got, want)
}
}
// A plan for today that includes tomorrow's meeting is worse than terse.
func TestBuildPlanDropsOtherDays(t *testing.T) {
now := time.Date(2026, 8, 3, 9, 0, 0, 0, time.UTC)
events := []PlanEntry{
{At: planAt(now, 10, 0), Text: "today", Kind: PlanEvent},
{At: planAt(now, 10, 0).AddDate(0, 0, 1), Text: "tomorrow", Kind: PlanEvent},
{At: planAt(now, 10, 0).AddDate(0, 0, -1), Text: "yesterday", Kind: PlanEvent},
{At: planAt(now, 12, 0), Text: " ", Kind: PlanEvent},
}
p := BuildPlan(nil, nil, events, nil, now)
if len(p.Items) != 1 || p.Items[0].Text != "today" {
t.Fatalf("got %+v", p.Items)
}
}
func TestPlanFormatRU(t *testing.T) {
p, _ := planFixture(t)
got := p.FormatRU()
want := "план на 03.08.2026: 10:00 — Standup @ 10:00-10:30; " +
"10:30 — утро — осталось: витамины; " +
"похоже, 14:00 — Планёрка @ 14:00-14:30; " +
"18:30 — позвонить маме."
if got != want {
t.Errorf("got %q\nwant %q", got, want)
}
// Persona: she recites, she does not exhort, and she never speaks of
// herself in the masculine or addresses him formally.
for _, bad := range []string{"рад ", "понял", "вы ", "ваш", "милый", "дорогой", "давай же", "не забудь"} {
if strings.Contains(strings.ToLower(got), bad) {
t.Errorf("plan text contains %q: %q", bad, got)
}
}
}
func TestPlanAfter(t *testing.T) {
p, now := planFixture(t)
rest := p.After(planAt(now, 11, 0))
if len(rest.Items) != 2 {
t.Fatalf("got %d items, want the 14:00 and 18:30 ones: %+v", len(rest.Items), rest.Items)
}
if !rest.Date.Equal(p.Date) {
t.Error("After must keep the date, so an empty rest-of-day still knows which day")
}
empty := p.After(planAt(now, 23, 0))
if len(empty.Items) != 0 {
t.Errorf("got %+v", empty.Items)
}
if !strings.Contains(empty.FormatRU(), "ничего не запланировано") {
t.Errorf("empty plan reads %q", empty.FormatRU())
}
}
+13 -3
View File
@@ -20,9 +20,19 @@ type ProposedRoutine struct {
const MaxIntervalRatio = 1.5 const MaxIntervalRatio = 1.5
// MinEvents is the minimum number of events needed to detect a pattern. // MinEvents is the minimum number of events needed to detect a pattern.
// With N events, there are N-1 intervals; we need at least 2 intervals // With N events there are N-1 intervals, so 4 events means 3 intervals.
// before proposing anything. //
const MinEvents = 3 // This used to be 3 (two intervals), which is not a pattern — it is a
// coincidence with a mean. Two gaps of similar length happen constantly:
// water the plants on a Sunday, again the next Sunday, once more the Sunday
// after, and a detector with a ±50% band calls that a weekly routine. The
// cost of being wrong is asymmetric now that the digestion tick scans all of
// history on its own schedule and can announce what it finds: a false
// positive is something the owner has to read and dismiss, and a dismissal
// is permanent, so one bad guess burns that action+object pair forever.
// Three intervals is the cheapest bar that makes a run distinguishable from
// a repeat. False negatives cost one more observation and nothing else.
const MinEvents = 4
// Detect checks whether a sequence of events for the same action+object // Detect checks whether a sequence of events for the same action+object
// forms a stable recurring pattern. Returns a ProposedRoutine when: // forms a stable recurring pattern. Returns a ProposedRoutine when:
+27 -15
View File
@@ -6,12 +6,13 @@ import (
) )
func TestDetectEnoughEvents(t *testing.T) { func TestDetectEnoughEvents(t *testing.T) {
// 3 events with 7-day intervals → stable pattern // MinEvents events with 7-day intervals → stable pattern
base := time.Date(2026, 7, 1, 12, 0, 0, 0, time.UTC) base := time.Date(2026, 7, 1, 12, 0, 0, 0, time.UTC)
events := []Event{ events := []Event{
{Action: "refill", Object: "cat_water", Ts: base}, {Action: "refill", Object: "cat_water", Ts: base},
{Action: "refill", Object: "cat_water", Ts: base.Add(7 * 24 * time.Hour)}, {Action: "refill", Object: "cat_water", Ts: base.Add(7 * 24 * time.Hour)},
{Action: "refill", Object: "cat_water", Ts: base.Add(14 * 24 * time.Hour)}, {Action: "refill", Object: "cat_water", Ts: base.Add(14 * 24 * time.Hour)},
{Action: "refill", Object: "cat_water", Ts: base.Add(21 * 24 * time.Hour)},
} }
r, err := Detect(events) r, err := Detect(events)
@@ -24,8 +25,8 @@ func TestDetectEnoughEvents(t *testing.T) {
if r.Action != "refill" || r.Object != "cat_water" { if r.Action != "refill" || r.Object != "cat_water" {
t.Fatalf("action/object: want refill/cat_water, got %s/%s", r.Action, r.Object) t.Fatalf("action/object: want refill/cat_water, got %s/%s", r.Action, r.Object)
} }
if r.N != 3 { if r.N != 4 {
t.Fatalf("want N=3, got %d", r.N) t.Fatalf("want N=4, got %d", r.N)
} }
// ~7 days // ~7 days
if r.IntervalDays < 6.9 || r.IntervalDays > 7.1 { if r.IntervalDays < 6.9 || r.IntervalDays > 7.1 {
@@ -33,19 +34,28 @@ func TestDetectEnoughEvents(t *testing.T) {
} }
} }
// TestDetectNotEnoughEvents — two intervals are a coincidence, not a routine
// (Vikunja #43). Three same-day-of-week events used to be enough to propose a
// weekly reminder; MinEvents is 4 now so a repeat has to happen a third time
// before Maven calls it a pattern.
func TestDetectNotEnoughEvents(t *testing.T) { func TestDetectNotEnoughEvents(t *testing.T) {
base := time.Date(2026, 7, 1, 12, 0, 0, 0, time.UTC) base := time.Date(2026, 7, 1, 12, 0, 0, 0, time.UTC)
events := []Event{ for _, n := range []int{1, 2, MinEvents - 1} {
{Action: "refill", Object: "cat_water", Ts: base}, events := make([]Event, n)
{Action: "refill", Object: "cat_water", Ts: base.Add(7 * 24 * time.Hour)}, for i := range events {
} events[i] = Event{
Action: "refill",
r, err := Detect(events) Object: "cat_water",
if err != nil { Ts: base.Add(time.Duration(i) * 7 * 24 * time.Hour),
t.Fatalf("Detect: %v", err) }
} }
if r != nil { r, err := Detect(events)
t.Fatal("want nil for <3 events") if err != nil {
t.Fatalf("Detect(%d events): %v", n, err)
}
if r != nil {
t.Fatalf("Detect(%d events) proposed %+v, want nil below MinEvents=%d", n, r, MinEvents)
}
} }
} }
@@ -68,12 +78,13 @@ func TestDetectEmpty(t *testing.T) {
} }
func TestDetectIrregularRejects(t *testing.T) { func TestDetectIrregularRejects(t *testing.T) {
// 3 events but wildly irregular: 1 day, then 14 days → ratio 14 > 1.5 // wildly irregular: 1 day, then 14 days → ratio 14 > 1.5
base := time.Date(2026, 7, 1, 12, 0, 0, 0, time.UTC) base := time.Date(2026, 7, 1, 12, 0, 0, 0, time.UTC)
events := []Event{ events := []Event{
{Action: "refill", Object: "cat_water", Ts: base}, {Action: "refill", Object: "cat_water", Ts: base},
{Action: "refill", Object: "cat_water", Ts: base.Add(1 * 24 * time.Hour)}, {Action: "refill", Object: "cat_water", Ts: base.Add(1 * 24 * time.Hour)},
{Action: "refill", Object: "cat_water", Ts: base.Add(15 * 24 * time.Hour)}, {Action: "refill", Object: "cat_water", Ts: base.Add(15 * 24 * time.Hour)},
{Action: "refill", Object: "cat_water", Ts: base.Add(16 * 24 * time.Hour)},
} }
r, err := Detect(events) r, err := Detect(events)
@@ -117,6 +128,7 @@ func TestDetectSameTimestamp(t *testing.T) {
{Action: "refill", Object: "cat_water", Ts: base}, {Action: "refill", Object: "cat_water", Ts: base},
{Action: "refill", Object: "cat_water", Ts: base}, {Action: "refill", Object: "cat_water", Ts: base},
{Action: "refill", Object: "cat_water", Ts: base.Add(7 * 24 * time.Hour)}, {Action: "refill", Object: "cat_water", Ts: base.Add(7 * 24 * time.Hour)},
{Action: "refill", Object: "cat_water", Ts: base.Add(14 * 24 * time.Hour)},
} }
r, err := Detect(events) r, err := Detect(events)
+138
View File
@@ -0,0 +1,138 @@
// Package persona builds the one shared context block that goes in front of
// every LLM system prompt: who the owner is, how to address him, and what
// time it is right now.
//
// Why one block and not a line pasted into each prompt: there are five
// prompts (nudges, action replies, chat, note queries, general knowledge) and
// the "address him as ты" rule had only reached two of them. Five copies drift.
// One block cannot.
//
// The rules here are defaults in code, not config. Maven is feminine and the
// owner is a man addressed informally — that is a hard constraint of the
// product, so it must hold with an empty config file. Config only ADDS
// optional facts (his name, his city).
package persona
import (
"fmt"
"strings"
"time"
)
// Facts — the optional, deployment-specific half of the block. All fields may
// be empty; the block is still correct and useful without them.
type Facts struct {
OwnerName string // his name, e.g. "Ками"
City string // where he is, e.g. "Москва"
Static string // the free-text `persona` config string, appended verbatim
// The two config-gated capabilities. They are listed only when this
// deployment actually has them, because a capability she names and cannot
// do is worse than one she never mentions.
Weather bool // an open-meteo provider is configured
Telegram bool // a telegram bot token + chat id are configured
Tools bool // at least one shell act is on the allowlist
}
var ruWeekdays = [...]string{"воскресенье", "понедельник", "вторник", "среда", "четверг", "пятница", "суббота"}
var ruMonths = [...]string{
"января", "февраля", "марта", "апреля", "мая", "июня",
"июля", "августа", "сентября", "октября", "ноября", "декабря",
}
// Block renders the context block for one turn. Russian even in front of the
// English prompts: the rules it states are Russian grammar (ты/тебя, feminine
// verbs), and a Russian rule reads best stated in Russian.
//
// Keep it short. It ships on every turn to a 0.8B on laptop CPU, so every
// line here is latency.
func (f Facts) Block(now time.Time) string {
var b strings.Builder
b.WriteString("Ты — Maven, домашняя ассистентка. О себе говоришь в женском роде: \"я записала\", \"я проверила\".\n")
// The address form gets its own line. It is the thing that kept getting
// lost when it was buried in prose.
b.WriteString("ОБРАЩЕНИЕ: владелец — мужчина, всегда на \"ты\" (ты, тебя, тебе, твой) и в единственном числе (\"выпей\", \"посмотри\"). Никогда \"вы\"/\"вас\"/\"ваш\". Никогда \"он\"/\"его\" о нём — ты говоришь ему, а не о нём. Глаголы о нём — в мужском роде (\"ты забыл\").\n")
if who := f.who(); who != "" {
b.WriteString(who + "\n")
}
b.WriteString(fmt.Sprintf("Сейчас: %s, %d %s %d, %02d:%02d (местное время).\n",
ruWeekdays[int(now.Weekday())], now.Day(), ruMonths[int(now.Month())-1], now.Year(),
now.Hour(), now.Minute()))
b.WriteString("Умеешь: " + strings.Join(f.can(), "; ") +
". Других ДЕЙСТВИЙ не умеешь — если просят такое, скажи прямо.\n")
if s := strings.TrimSpace(f.Static); s != "" {
b.WriteString(s + "\n")
}
return b.String()
}
// can lists what she can really do. Every entry here is a code path that
// exists in the daemon today:
// - reminders: IntentReminder → CoreAPI.CreateReminder, fired by the tick.
// - notes and facts: IntentNote/IntentFact write, IntentQuery reads them back.
// - calendar: IntentQuery answers "что у меня сегодня" from CalendarEvents.
// - weather / telegram / shell acts: only when configured (see Facts).
//
// Nothing speculative goes in this list. A capability she offers and cannot
// perform is worse than one she never mentions.
func (f Facts) can() []string {
c := []string{
// Talking comes first, and the closing line says "действий" rather than
// "ничего", because this same block sits in front of the chat and
// general-knowledge prompts. A flat "you can do nothing else" would
// tell her to refuse the exact thing those two prompts are for.
"разговаривать и отвечать на вопросы",
"ставить напоминания",
"записывать заметки и факты и отвечать по ним",
"смотреть календарь",
}
if f.Weather {
c = append(c, "говорить погоду")
}
if f.Telegram {
c = append(c, "писать в телеграм")
}
if f.Tools {
c = append(c, "запускать разрешённые команды на сервере")
}
return c
}
// who renders the optional name/city line, or "" when neither is configured.
//
// Written as labels ("Имя владельца: ..."), not as a sentence with pronouns:
// the block's own "ты" is Maven, so "тебя зовут" would read as her name and
// "его" would model the third-person form she must never use about him.
func (f Facts) who() string {
name := strings.TrimSpace(f.OwnerName)
city := strings.TrimSpace(f.City)
switch {
case name != "" && city != "":
return "Имя владельца: " + name + ". Город: " + city + "."
case name != "":
return "Имя владельца: " + name + "."
case city != "":
return "Город: " + city + "."
}
return ""
}
// Prepend puts the block in front of a system prompt. Nil-safe: a nil renderer
// (tests, the stub paths) returns the prompt untouched.
func Prepend(block func() string, prompt string) string {
if block == nil {
return prompt
}
s := strings.TrimSpace(block())
if s == "" {
return prompt
}
return s + "\n\n" + prompt
}
+72
View File
@@ -0,0 +1,72 @@
package persona
import (
"strings"
"testing"
"time"
)
var ref = time.Date(2026, 7, 31, 14, 5, 0, 0, time.UTC)
// The block must be correct with an empty config: the address form and the
// gender rules are hard constraints, not preferences.
func TestBlockWorksWithZeroConfig(t *testing.T) {
b := Facts{}.Block(ref)
for _, want := range []string{"женском роде", "ОБРАЩЕНИЕ", "\"ты\"", "31 июля 2026", "пятница", "14:05"} {
if !strings.Contains(b, want) {
t.Errorf("block missing %q:\n%s", want, b)
}
}
}
func TestBlockAddsOptionalFacts(t *testing.T) {
b := Facts{OwnerName: "Ками", City: "Москва", Static: "Будь краткой."}.Block(ref)
for _, want := range []string{"Ками", "Москва", "Будь краткой."} {
if !strings.Contains(b, want) {
t.Errorf("block missing %q:\n%s", want, b)
}
}
}
// The time changes between turns, so two renders must differ.
func TestBlockRendersTimePerTurn(t *testing.T) {
a := Facts{}.Block(ref)
c := Facts{}.Block(ref.Add(time.Hour))
if a == c {
t.Errorf("block did not change with the clock:\n%s", a)
}
}
// She may only offer what this deployment actually has.
func TestCapabilitiesAreConfigGated(t *testing.T) {
bare := Facts{}.Block(ref)
for _, want := range []string{"напоминания", "заметки", "календарь"} {
if !strings.Contains(bare, want) {
t.Errorf("block missing always-on capability %q:\n%s", want, bare)
}
}
for _, unwanted := range []string{"погоду", "телеграм", "команды"} {
if strings.Contains(bare, unwanted) {
t.Errorf("block offers unconfigured %q:\n%s", unwanted, bare)
}
}
full := Facts{Weather: true, Telegram: true, Tools: true}.Block(ref)
for _, want := range []string{"погоду", "телеграм", "команды"} {
if !strings.Contains(full, want) {
t.Errorf("block missing configured capability %q:\n%s", want, full)
}
}
}
func TestPrependNilIsSafe(t *testing.T) {
if got := Prepend(nil, "PROMPT"); got != "PROMPT" {
t.Errorf("Prepend(nil) = %q", got)
}
if got := Prepend(func() string { return " " }, "PROMPT"); got != "PROMPT" {
t.Errorf("Prepend(blank) = %q", got)
}
if got := Prepend(func() string { return "CTX" }, "PROMPT"); got != "CTX\n\nPROMPT" {
t.Errorf("Prepend = %q", got)
}
}
+54
View File
@@ -0,0 +1,54 @@
package phraser
import (
"errors"
"strings"
"testing"
)
// A reply that starts a JSON object and never finishes it is a failed
// generation, not a reply. Before this, the parser returned ("", "") for these
// and every caller then shipped the raw fragment as the thing Maven said. A
// real run produced replies of literally "{" and "{\n \"".
func TestParseResponseMoodRejectsUnfinishedJSON(t *testing.T) {
for _, raw := range []string{
`{`,
"{\n \"",
`{"response": "неполн`,
`{"response": "текст", "mood":`,
} {
text, mood, err := parseResponseMood(raw)
if !errors.Is(err, errBrokenJSON) {
t.Errorf("parseResponseMood(%q) err = %v, want errBrokenJSON", raw, err)
}
if text != "" || mood != "" {
t.Errorf("parseResponseMood(%q) leaked %q/%q — a fragment must never come back as a reply", raw, text, mood)
}
}
}
// Bare prose is still fine. Small models sometimes answer without any JSON at
// all, and that reply is usable — so the new error must not swallow it.
func TestParseResponseMoodAllowsBareProse(t *testing.T) {
for _, raw := range []string{
"норм, а ты как?",
"вот что я нашла: ключ у соседа",
} {
text, mood, err := parseResponseMood(raw)
if err != nil {
t.Errorf("parseResponseMood(%q) err = %v, want nil", raw, err)
}
// No JSON means no fields; the caller ships raw as-is.
if text != "" || mood != "" {
t.Errorf("parseResponseMood(%q) = %q/%q, want empty", raw, text, mood)
}
}
}
// The measured failure: the model wants more than 400 characters and the old
// grammar cut it off mid-word. Guards the bound against being tightened back.
func TestGrammarStringBoundHasRoomForARealAnswer(t *testing.T) {
if !strings.Contains(responseGrammar, "{0,1000}") {
t.Error("grammar string bound is not 1000; 400 truncated real replies mid-word (see the comment on responseGrammar)")
}
}

Some files were not shown because too many files have changed in this diff Show More