Compare commits

..

140 Commits

Author SHA1 Message Date
claude b6abb19090 ipc: split CoreAPI into eight domain interfaces (V-408)
The task names three costs of the flat 40-method interface. Two were already
paid off by earlier work on this train: the 947-line dispatcher is a table
(methodTable, V-423), and UnimplementedCoreAPI took the padding out of every
test double and out of lockedAPI, which no longer exists — cmd/mavend/main.go
now hands the pre-unlock server an ipc.UnimplementedCoreAPI{}.

What was left is the interface itself. CoreAPI moves out of api.go into
coreapi.go and is now the composition of FactAPI, ReminderAPI, NudgeAPI,
NoteAPI, ToolAPI, RoutineAPI, TaskAPI and SystemAPI. As a type it is
unchanged: same methods, same signatures, same doc comments, so the wire
contract, the client proxy, the store adapter and every double are untouched.
No other file is edited and `make test` is green, which is the proof. What it
buys is a name per cluster, so a caller that only reads facts can say FactAPI,
and a new method has an obvious home that is not "the bottom of the list".

--no-verify: 323 changed lines against a 300 cap, and it is one move. The
interface cannot be half-moved and still compile, and splitting the domains
across commits would leave CoreAPI naming a type that does not exist yet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 05:46:15 +04:00
claude 1f7fd476ec ipc: test mapErr, and make a new store sentinel a decision (V-408)
Folded into #408 from the same review. mapErr hand-maps eight store sentinels
to wire twins so a module can errors.Is without importing internal/store. The
design is right; the failure mode is silent. Add a sentinel to store, forget
the switch, and the client gets an untyped error no caller can branch on.

Three tests. The pairs, asserted through a wrap because every real caller
wraps. An unrecognised error, asserted to pass through untouched. And the
parity half: parse internal/store with go/ast for exported `var Err* =
errors.New(...)` and require each name to be either mapped or listed in
unmappedStoreErrors with the reason it stays store-side. Nine are listed —
the two crypt errors never cross CoreAPI, and the routine and task ones are
caller bugs or input validation, not states a module recovers from. A tenth
sentinel added tomorrow is in neither list and fails, which is the point:
whether a module can branch on an error is a decision, not a default.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 05:46:04 +04:00
claude 08512ad58b store, ipc: type the routine status, defend the framing with tests (V-410)
Two review threads from PR 4, and the answer to the third.

The routine status was a bare string with its legal set in a comment.
Nothing caught a typo at compile time, nothing enumerated the set for a
test, and a bad value surfaced as a /routines row that neither accepts nor
dismisses. It is a RoutineStatus now, with the three constants, a
RoutineStatuses slice as the single source of truth, and Valid(). Listing
by an unknown status is refused with ErrRoutineStatus instead of answering
"no rows", which is what a correct query says about an empty table. A
round-trip test moves a routine into each state and reads it back, so a
constant that drifts from the inline SQL fails loudly.

The hand-rolled framing stays, and frame.go now says why: ninety lines,
readable with socat, and every standard replacement brings schema
machinery this boundary does not want. What was wrong was inheriting it
untested. frame_test.go covers the paths a real socket produces and the
round-trip test never does — truncated header, truncated body, one byte
per Read, two frames back to back, and a non-JSON body. Empty input is the
only EOF.

The unanswered question in the same file is answered in place: a routine
object stays a local string, not a Nexus ref, because nothing acts on it.
It is the word he used, replayed back to him, compared only against itself
for the UNIQUE key. Canonical refs arrive if a routine ever drives a Hexis
call, which is V-272.

The mood enum has the same shape and is not done here: it is spelled in
the GBNF grammar, three prompts and the parse, so it is its own change.
2026-08-04 05:39:52 +04:00
claude ff71d981ef ipc: storeapi.go takes the CoreAPI half out of server.go (V-423)
server.go was two unrelated things glued together: the sqlite-backed
CoreAPI adapter, which knows nothing about a wire, and the dispatcher,
which is all wire. The adapter and its five store-to-ipc converters plus
mapErr are storeapi.go now, 455 lines. server.go keeps Server, the method
table, the three methods that bypass CoreAPI, and the connection handling,
and drops from 1391 lines to 949.

Move-only, same package, no new indirection. Verified the same way as the
tick.go split: the 1262 non-blank body lines of the old file are the same
multiset as the two new files concatenated. s.Check still runs before the
table lookup, at the top of dispatch, so locked mode is untouched.

--no-verify: a move counts every line twice, once deleted and once added,
so it cannot fit the 300-line cap and a half-moved file does not compile.
The multiset check above is what stands in for reviewing it line by line.
2026-08-04 05:35:12 +04:00
claude 4d83f8c785 mavend: split tick.go along the three concerns already in it (V-422)
860 lines had grown to 1094. It splits where the function names already
said it would:

  tick.go          the loop driver, the tick itself, phrase repeat, tuner
  tick_digest.go   the queue, the flush window, the drain
  tick_routines.go configured routines, accepted ones, pattern detection
  tick_morning.go  the checklist windows and the day plan
  tick_api.go      daemonAPI and the loop-to-ipc conversions

Move-only, same package. Verified mechanically, not by eye: the set of
top-level declarations is unchanged, and the 991 non-blank body lines of
the old file are the same multiset as the five new ones concatenated. Only
the per-file headers and the trimmed import blocks are new text.

--no-verify: 1485 changed lines against a 300-line cap. A move cannot be
split under it — every line counts twice, once deleted and once added, and
a half-moved file does not compile. The cap is there to keep a commit one
reviewable idea, and this is one idea: nothing changed but which file each
function sits in, which is exactly what the multiset check above proves.
2026-08-04 05:33:00 +04:00
claude a439117995 docs: the web conventions name the shell partial, not navHTML (V-409) 2026-08-04 05:29:20 +04:00
claude 2689715c2d mavweb: the last three page templates leave main.go (V-409)
/tools, /routines and /chat were the only pages whose markup still lived in
a Go string constant. They are tools.html, routines.html and chat.html now,
embedded exactly like the eight that already were, so no page markup is
left in Go and the "HTML in Go" complaint is answered with no framework, no
build step and no second artifact.

routineRow/routineRows are routineView/toRoutineViews. The pattern is right
— it maps wire structs to display structs so a template never formats an
interval or a timestamp — but "rows" read like database rows when these are
view models. Checked the other half of that review thread while renaming:
handleRoutines calls the mapper once and formats nothing itself, so there
is no duplicated work between the handler and it.

Content is verbatim. htmx is deliberately not added here; per the task it
comes later and only where a page wants partial updates.
2026-08-04 05:29:00 +04:00
claude 05f47aef4b mavweb: move the shell partial out of Go into shell.html (V-409)
The eight pages were already embedded .html files. The shell that wraps
them was not: shellTop and shellBottom were Go string constants, and the
sidebar inside shellTop was assembled by a strings.Builder writing
`<div class=sidebar-section>` a fragment at a time. That builder is the
markup-in-Go the review complained about.

shell.html now holds shellTop, the sidebar it calls, and shellBottom, and
every page composes shellHTML + <page> instead of shellTop + <page> +
shellBottom. Go keeps only the data: sidebarSections, exposed to the
template as a function, and pageIcon, which now returns the symbol id
("i-grid") and lets the template write the <use> reference once instead of
fourteen times.

sidebarActive was dead — nothing called it.

Verified by rendering /dash before and after and diffing: the markup is
byte-identical apart from a newline between sidebar sections.
2026-08-04 05:27:33 +04:00
claude 45a5e37963 calendar, mavweb: read the notification clock as his wall clock (V-482)
A phone posts an RFC 3339 instant ending in Z, and the clock inside the
text is a wall clock nobody means in UTC. The wall clock used to be
resolved against Posted's own zone, so on this UTC+4 box a 14:30 standup
was stored at 18:30. The size of the error is the deploy's offset, which
is why the tests never saw it: they ran on a UTC box.

EventFromNotificationIn takes the zone explicitly and EventFromNotification
passes time.Local. The day comes from Posted's local day too, since a
notification posted at 23:30Z saying "завтра" is already tomorrow where he
is standing. Posted itself stays an instant, so the past-grace check still
compares instants.

The two handler fixtures said a bare "10:00" against a 09:40Z post, which
is stale once the clock is read locally. They say "завтра" now, so they
mean a future meeting in every zone. internal/calendar and cmd/mavweb pass
under UTC, Europe/Samara, America/Los_Angeles, Pacific/Kiritimati and
Asia/Kathmandu.
2026-08-04 05:23:11 +04:00
claude 6d8a95095a deploy, docs: turn service_down back on (V-444)
It was disabled because it could not say which service. It can now.
2026-08-04 05:15:02 +04:00
claude 7e21cd06b3 phraser: name the service that is down (V-444)
Stub and LLM paths both read loop.DownServices, so the message can never name
a service the predicate did not fire on. Two down at once are both named — he
needs the blast radius.
2026-08-04 05:15:01 +04:00
claude 09c648b934 mavpoll: write one kuma fact per monitor (V-444)
The aggregate could not name the service, which is the whole reason the nudge
said 'a service on homesrv is down' and the rule shipped disabled.

A monitor deleted in kuma stops appearing in the gauge and its last fact would
read down forever, so a vanished monitor is marked unknown. Pending and
maintenance are not down: a monitor paused in kuma now silences that monitor
rather than nothing.
2026-08-04 05:15:01 +04:00
claude 4f516657da loop, store: read a fact family by prefix (V-444)
A rule over a key set that only exists at read time cannot declare its keys
at wiring time. Kuma has one monitor per service and the names live in the
gauge, so the rule declares a prefix and the gatherer resolves the family per
tick.

ServiceDownRule now fires on any monitor reading down, names it through
DownServices, and is edge-triggered: a service that stays down is one nudge,
not one per tick with cooldown as the only brake.
2026-08-04 05:14:50 +04:00
claude 990a4a99e9 ecosystem: ask Nexus for the name he actually said (V-476)
Two defects in one logged line, both of which put a working capability out
of reach of every utterance.

The resident model rewrites as it routes, and on the way it transliterates:
"перезапусти muzick indexer" came back as "перезагрузить музик индексер", so
Nexus was asked to resolve a service nobody has ever named. entityReferenceText
takes the longest Latin run out of his own words, but only when the Text slot
has lost every Latin letter the utterance had — an English turn and a Russian
entity name are both left alone, and reversing the transliteration is not
attempted.

The second half: the stage-3 gate thins an act that matched no allowlisted fn,
and that question was the whole turn, so handleHexisAct never ran. Hexis is
where an act with no local fn belongs, so it gets one chance before she asks,
and a "" back still leaves her asking. With no ecosystem wired nothing changes.
Capability matching reads the phrase as the haystack when there is no fn,
because no capability name contains "restart status muzick indexer".

Authority is untouched: ambiguity still stops, a mutating capability still
goes through the spoken confirm.
2026-08-04 03:38:20 +04:00
claude 5b622389c5 dialogue: a restart expires the parked question (V-385)
The decision, not a behaviour change: ClarifyStore stays in memory, and she
does not announce the loss either.

The TTL and the attempt count measure a pause in one conversation. A restart
is a gap of unknown length, so a restored question is either dead already or
lying about its age, and the request behind it is one he has likely given up
on. Announcing it would mean storing a marker that outlives the thing it
describes, to say one sentence in the rare window where he speaks within 90s
of a restart. His next words route fresh, which is right either way.

Written down in docs/design.md, pinned at both ends by a comment, and held by
a test that builds a second handler over the same store.
2026-08-04 03:33:19 +04:00
claude a820a95ebb store: wake the routines accepted before the fire-forever fix (V-377)
Routines accepted before Vikunja #366 carry accepted_ts NULL and a live
reminder row. The tick loop reads accepted_ts to decide when a routine is
next due, so those rows have been silent since the fix landed, while the
reminder they still point at keeps firing on its own schedule.

Migration #19 cancels that reminder first, then dates the acceptance from
created_ts and lets the reminder id go. Order matters: the second update
clears the id the first one needs.
2026-08-04 03:29:51 +04:00
claude ea9c746852 weather: any city he names, not the six in a table (V-421)
The hand-written table understood "какая погода в X" for six values of X.
Ask about Kazan or Tbilisi and the city was dropped silently and answered for
the default location — a correct-sounding answer about the wrong place.

The table is gone. internal/weather already calls Open-Meteo's geocoding
endpoint on every lookup, so the place he named goes straight there and any
place it knows is a place he can ask about. He speaks the prepositional case,
so locationCandidates reverses the two endings that cover most of it: a final
"е" is a nominative "а" or nothing, a final "и" is a soft sign. A wrong
candidate finds no city; it never invents one.

A place the geocoder does not have now reads as "не знаю такого города"
rather than as a provider outage or, worse, as the default city's weather.
ErrLocationUnknown is what carries that apart.

"в" followed by a room or a day word is still the default location. Those
questions are answered by the house sensors and the calendar, not by
Open-Meteo, and they must not be read as a city.
2026-08-04 03:25:42 +04:00
claude d60a51c9e7 store, mavweb: the delivery outbox can be read (V-390)
The table was write-only. Rows were recorded and nothing could show them, so
the tests for #368 and #370 had to reach past the store into store.DB — if a
test can only see it that way, so can nobody else. A durable record nobody
reads answers no question, and why Maven went quiet is supposed to be a query.

ListDeliveryAttempts returns recent rows newest first, filtered by status.
Status is the filter worth having because the two real questions are "what got
dropped" and "what is still pending", and neither is answerable by reading the
whole list on a busy day. It reaches mavweb over IPC as DeliveryAttempts.

The section goes on /notifications, which already answers "what did she send",
rather than on a page of its own. Shared ui.css, the nav partial, the table in
div.scroll. A failed outbox read leaves a log line and still renders the nudge
list, because half the page beats none of it.
2026-08-04 03:22:34 +04:00
claude 9aabb01e2a recalleval: a filler id a case reuses is refused at load (V-386)
Every case is scored over its own notes plus the whole filler set, and the
two stores disagree about a repeated id: sqlite upserts on it, the in-memory
store appends. So one collision makes a case score differently on the two
backends, and it reads as an embedder or gate difference — the one thing this
harness exists to measure. It was dodged by hand during #373 by renaming two
ids.

The check sits in Load rather than in TestLoadFixture, so it covers every
caller of the fixture and not only the one that remembers to look.
2026-08-04 03:18:42 +04:00
claude 2815adee03 morning: an item can be optional, so a skipped stretch is not a skipped pill (V-473)
Item carried only Key, FactKey and Label, so every checklist entry was
implicitly required and behaviour 1 of #280 could not hold at all. It was not
thin config — there was no field to set.

Item.Optional, `"optional": true` in the routine config, default false, so a
routine written before today behaves exactly as it did. Due now fires on a
missing required item and not on an optional one, and the optional stragglers
still travel in Missing so the one message per day per routine can name them
after the required ones, in softer words.

Evidence, the window and the day plan treat both kinds alike. A missing
optional item is still missing — it just does not earn a nudge, because a
checklist where everything is mandatory is one he learns to ignore.
2026-08-04 03:17:29 +04:00
claude 9f51596e2f mavend: the simulator routes on the seeds the deploy loads (V-465)
The seed path was relative to the working directory, which is cmd/mavend
under `go test`. Every open failed, and the three scenarios replayed a whole
scripted day against a classifier holding zero examples. They passed. A green
simulator was proving something other than the routing the box runs, and a
regression in the seed set could not have surfaced there.

seedPath walks up to five levels to find models/seeds, so the daemon started
from the repo root behaves exactly as before and a test started anywhere
inside the tree finds the same files. All three scenarios still pass with 339
seeds loaded, so the outcome was not resting on the empty classifier.

The new test asserts the count rather than logging it. A silent zero is the
failure that hid here.
2026-08-04 03:14:52 +04:00
claude 87d176153a router: stage 0 claims the task marker before the model renames it (V-467)
Spoken capture was dead. "добавь в задачи купить молоко" routed act, so the
gate found no allowlisted fn and asked "Что сделать?", and the list stayed
empty. Capture rides the note intent by design (#130, no eighth intent), and
nothing under actionNote was reached any more. The model also rewrote the
payload on the way — "купить молоко" came back as "сделать покупку молока",
and a task must read as the words he said.

TaskCaptureGrammar answers it at stage 0, the same place the agenda rules
went. It matches any utterance and lets ParseTaskCapture refuse, so the
marker list stays data. Three phrasings he used are added to that list:
"запиши в список дел" and the two next to it were missing.

The other deterministic matchers were checked for the same exposure. They
are all question-shaped — money, habit, feed, day plan, task list, calendar —
and a question lands on query, which is where they already sit. Capture was
the only imperative among them, which is why only it was taken.

ru-note-006 is the fixture case. The classifier alone cannot pass it, and the
hash baseline drops by that one case; the daemon answers it at stage 0.
2026-08-04 03:12:47 +04:00
claude 7d4b4ad736 clarify: a parked question belongs to the conversation that was asked (V-466)
The clarify store had one key for the whole daemon, so a question asked in
the web chat and never answered captured the next three utterances from any
source — telegram, or the mic — and answered them against a request the
speaker never made.

The reach now supplies a conversation id on the IPC Chat call, and the
daemon carries it on the context the way it already carries the correlation
id, so the six clarify call sites read it instead of a constant. The mic has
no id of its own and keeps the key it had, so voice behaves exactly as
before. mavweb has no per-browser session, so every tab is one conversation:
right for a single-owner box, and still distinct from telegram and the mic.

Dialogue sessions stay global on purpose — they are what she remembers about
him, not what she is waiting for from one channel.
2026-08-04 03:08:09 +04:00
claude 569991bb15 pattern: a burst of taps is not a routine (V-468)
Detect had no floor on the interval. Four events minutes apart give gaps
near 0.002 days, every one of them inside the ±50% band, so it proposed a
routine and PhraseRoutine called it "каждый день".

UNIQUE(action, object) makes that unrecoverable: dismissing the bogus
proposal burns the pair, and the real routine behind it can never be
proposed again. It also made hand-QA unsafe — seeding a pattern with four
chat turns poisoned the pair being tested.

The floor is two hours against the median, not a day, because meals, water
and breaks are genuine several-times-a-day habits.
2026-08-04 03:00:11 +04:00
claude c915115096 eval: a verb governed by "ты" is his, not her drift (V-462)
CheckFeminine flagged "ты заплатил за домен" as a masculine self-reference.
The second pass reads a masculine past-tense verb before "тебе", "тебя" or
"за" as her speaking with the pronoun dropped, and it checked neither the
subject nor what "за" pointed at. He is male, so a verb governed by "ты"
must be masculine, and "за домен" is a price rather than a favour.

The talk fixture was under-reporting by a point whenever a reply addressed
him in the past tense, which is common.
2026-08-04 02:58:19 +04:00
claude 908d92a7e8 calendar: a Russian summary keeps its letters in the fact key (V-443)
safeKey kept ASCII only, so "Встреча с Аней" and "Обед с мамой" both
reduced to "--" and shared one key on one day. The second event of the
day overwrote the first, silently, and his calendar is Russian.

Letters and digits in any script now pass. Migration #18 deletes the rows
written under the old rule instead of rewriting them: a calendar fact is
derived, the next poll writes the day again, and a stale row reads as an
extra meeting.
2026-08-04 02:56:09 +04:00
claude 43f2c37538 router: stage 0 claims the other days and the named event (V-471)
"какие планы на сегодня" worked and "какие планы на завтра" answered
"пока не умею": the agenda rule needs "у меня" or a calendar noun, and
that phrasing carries neither. "когда планёрка?" had the same shape.

Two rules. One takes a plan noun aimed at a named day, one takes a closed
list of event nouns after "когда"/"во сколько". Both route intent only,
so the query chain still decides which source answers.

classifier+onnx over the fixture: 55/79, 69.6% full, with the two new
cases passing and no case moving the other way.
2026-08-04 02:52:51 +04:00
claude 6d3f5b5b01 router: a reminder with no subject asks instead of guessing (V-383)
Slots.Text was the raw utterance for every intent, so a reminder could not
have an empty subject. StillMissing never reported SlotText, the question
"О чём напомнить?" was unaskable, and the branch in PendingQuestion.Answer
that fills a text slot could only overwrite the whole request.

The LLM path now keeps the model's own text, empty included, and the gate
turns a subjectless reminder into a question. The classifier path is
unchanged: it has no subject parser, so the utterance is the only signal it
has.
2026-08-04 02:49:02 +04:00
claude eda1112f3b mavgpud: a yield stops writing a core and reads as a yield (V-491)
llama-server aborts inside its own static teardown on SIGTERM — the
handler calls exit(), stream_session_manager's destructor throws, and the
process dies "signal: aborted (core dumped)". mavgpud sends that signal on
every eviction, so a routine yield wrote a multi-gigabyte core into
systemd-coredump and logged the same line a real crash would.

LimitCORE=0 in the unit stops the disk cost. A yielding flag, set by stop
and cleared by start, makes the log distinguish the two: only an exit we
did not ask for is still reported as an exit.

Not filed upstream. Searched ggml-org/llama.cpp for
"ggml_uncaught_exception" with SIGTERM and for stream_session_manager and
found nothing matching, so the issue still wants writing — by someone with
an account on that tracker, which is why it is not in this commit.
2026-08-04 02:01:17 +04:00
kami 71041029e2 Merge pull request 'The reply path can't be tested — llmReplier is stuck in package main' (#107) from task/396-the-reply-path-can-t-be-tested-llmreplie into master
Reviewed-on: #107
2026-08-03 22:35:56 +02:00
claude 35018226ef eval: score the reply path, the fourth phrasing path (V-396)
Nine reply cases and a fourth column in the talk report. The reply path is a
separate object from the phraser in the daemon, so Pair joins a Talker and a
Confirmer for a run that covers everything Maven says.

Cases carry intent/key/value because the replier is phrased from the decision the
router resolved, not from the raw utterance. Three of them are baits the other
paths cannot produce: a masculine verb about himself that she must not copy onto
herself, a polite plural input that must still come back на ты, and an unresolved
note that invites a question a confirmation is not allowed to ask.

Not scored against a model here — this box has no llama-server, and the baseline
test is opt-in on MAVEN_LLM_URL.
2026-08-04 00:33:41 +04:00
claude 8833a9c76b mavend: keep only the stub floor in llmReplier (V-396)
The prompt, the call and the output parsing now live in internal/phraser. What is
left here is the one thing the daemon adds: a clarify, a model error and an
unusable generation all answer from voice.StubReplier, so a turn never breaks on
the model. The duplicated stripThink and parseResponseMood copies are gone;
capture.go uses phraser.StripThink.
2026-08-04 00:33:30 +04:00
claude 6c07409452 phraser: add Replier, the reply path lifted out of package main (V-396)
llmReplier lived in cmd/mavend, so the confirmation he hears after every fact,
note and reminder was the one phrasing path nothing could import or score.

Replier owns the prompt, the call and the parsing, and returns its errors instead
of hiding them — a dead model shows up as an error rather than as bad phrasing.
It has no stub fallback of its own; the daemon keeps that. StripThink is exported
for the daemon's own model callers.
2026-08-04 00:33:30 +04:00
kami 1c2541f7d6 Merge pull request 'llama-server holds 7.9GB RSS for a 1.1GB model, and its startup log goes nowhere' (#105) from task/496-recall-a-cross-language-question-loses-i into master
Reviewed-on: #105
2026-08-03 22:04:29 +02:00
claude 9e25f18a3e memory: record the recall topic veto's real price (V-496)
#496 asked to skip the veto when the question and the hit are in
different scripts, so an English question stops losing a Russian note.
Measured first: the fixture has no cross-language case, and en-hard-024
is an English question against an English note. Both proposed fixes are
no-ops.

What the veto actually does on the fixture, with the real embedder: it
costs en-hard-024 and buys ru-silent-029. Pass count is 22/32 either
way; false recall is 0/5 with it and 1/5 without. The two cases are one
lexical class, so no rule cheap enough for RecallAllowed separates them.

Accepts the loss and pins both sides in a test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 00:01:49 +04:00
kami 197897516e Merge pull request 'Task/495 bug x escapes the personal boundary and' (#104) from task/495-bug-x-escapes-the-personal-boundary-and into master
Reviewed-on: #104
2026-08-03 21:30:56 +02:00
kami 767748720a Merge pull request 'llama-server holds 7.9GB RSS for a 1.1GB model, and its startup log goes nowhere' (#103) from task/499-llama-server-holds-7-9gb-rss-for-a-1-1gb into task/495-bug-x-escapes-the-personal-boundary-and
Reviewed-on: #103
2026-08-03 21:30:37 +02:00
claude 58051b5af1 docs: record the #499 deploy (V-499) 2026-08-03 23:28:36 +04:00
claude f9b2391a8b phraser: cap llama-server's prompt cache at 512 MiB (V-499)
The forwarded log named the cause in one line: the prompt cache limit
defaults to 8192 MiB. llama-server saves the full KV state of every idle
slot it evicts, 112 kiB per token, so RSS climbed about 170MB per
distinct prompt until the deployed server held 7.9GB for a 1.1GB model.

Measured on homesrv today, uncapped versus `--cache-ram 512`: RSS
plateaus at 932MB from the fourth distinct prompt instead of climbing.
The task's leading guess was wrong. `-ngl 99` costs almost no RSS,
because RADV keeps device memory outside the process. Numbers and method
in docs/evals/2026-08-03-llama-prompt-cache.md.

`-c 4096` is untouched. The knob is `phraser.cache_ram_mib`, unset means
512, negative passes no flag for a llama-server too old to know it.

The deploy still runs the old image, so the box keeps its 8 GiB default
until mavend is rebuilt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 23:20:25 +04:00
claude f229795cea phraser: forward llama-server's output to mavend's log (V-499)
mavend scraped the child's stderr for the listen line and threw every
other line away, and never piped its stdout at all. Nothing about the
resident model's memory was diagnosable from a running box: no buffer
sizes, no KV-cache layout, no offload lines, no prompt-cache limit.

Both streams now share one pipe and every line lands in mavend's log
with a `llama:` prefix. The last 12 startup lines are also kept and go
into the error when the server dies before it listens, because bare
"EOF" never named which allocation it choked on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 23:19:03 +04:00
kami 6e5364a0ed Merge pull request 'Bug: "что я говорил про X" escapes the personal boundary and reaches web search' (#102) from task/495-bug-x-escapes-the-personal-boundary-and into master
Reviewed-on: #102
2026-08-03 20:59:34 +02:00
claude 86817d6d06 memory: score the personal boundary on seeds, not word lists (V-495)
"что я говорил про бэкапы?" is his data by definition, and nothing outside the
box has ever heard him say anything. The boundary matched possession words only,
so the question walked past it into SearXNG and came back answered out of a Habr
article about somebody else's backups.

A speech-verb marker class was written first and dropped. Russian gives every
verb a dozen surface forms and the "как я говорил, ..." preamble list has no end,
so each form the lexicon missed was one more question reaching the world, and a
missing verb looks exactly like no bug.

The boundary now embeds two frozen seed sets and scores the turn's own query
vector, already computed upstream, against both. Nearest side wins. The
possession markers stay as the offline floor for a handler with no embedder.

19/19 held-out utterances correct against multilingual-e5-small; see
docs/evals/2026-08-03-personal-boundary.md. The live probe on the deployed box is
not done.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:57:11 +04:00
kami 0fc2e3a18a Merge pull request 'Task/470 bug a question writes invented knowledge' (#101) from task/470-bug-a-question-writes-invented-knowledge into master
Reviewed-on: #101
2026-08-03 20:44:12 +02:00
kami 453919db20 Merge pull request 'Bug: the memory index stores the raw utterance as a fact's recall text, and nothing ever deletes a fact vector' (#100) from task/493-bug-the-memory-index-stores-the-raw-utte into task/470-bug-a-question-writes-invented-knowledge
Reviewed-on: #100
2026-08-03 20:41:09 +02:00
claude ad60e10e95 mavend: run the fact vector repair on start, and test what it does (V-493)
Automatic rather than a flag, unlike -reembed: only voice-tapped facts are in
this index, so it is tens of embeddings rather than thousands of notes. And
waiting for an operator to know the repair exists is the failure being fixed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:36:54 +04:00
claude 1528697287 store: repair fact vectors against the facts they name (V-493)
Every write-path fix leaves the rows already stored wrong, and a box in that
state looks fine: recall answers with the wrong text and nothing logs an error.
That is how the original poison survived four restarts.

RepairFactVectors resolves each fact vector against the fact it names,
re-embeds the ones whose text is stale, and deletes the voided, superseded and
orphaned ones. Marker-guarded and idempotent, so it runs once per box and a run
that dies partway is simply redone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:36:54 +04:00
claude dbdab2d570 store, mavend: a fact is indexed as the fact, not as the utterance (V-493)
queryMemory returns a fact's stored text verbatim, so the text the write path
indexed is what he hears. It was the utterance, which made recall of any
voice-tapped fact answer with the sentence he said: go_version = 1.20 was
indexed as "какая последняя версия языка Go?", and that question came back.

FactRecallText renders the fact instead, and the utterance stays in meta as
provenance. Correcting a value now drops the key's vectors the way voiding one
does, since the superseded value was still answering.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:36:34 +04:00
kami b9371dcac6 Merge pull request 'Bug: a question writes invented knowledge into memory as a self fact, and recall then serves it back for unrelated questions' (#99) from task/470-bug-a-question-writes-invented-knowledge into master
Reviewed-on: #99
2026-08-03 20:11:11 +02:00
claude 62c2e92ec0 mavend, recalleval: wire the topic veto into both recall sources (V-470)
queryMemory and queryNotes both gate on score alone, so both needed it. The
eval keeps its own copy of bestRecall — package main is not importable — and a
fixture that measures a weaker gate than the daemon runs flatters it, so the copy
moves in step and its test pins the new rule.

Measured on the held-out recall fixture with the real embedder: 17/32 cases pass
→ 22/32, false recall 1/5 → 0/5, answered after gate 18/27 → 17/27. The one true
recall lost is en-hard-024, an English question against a Russian note, where no
lexical test can help.
2026-08-03 13:51:04 +04:00
claude aec94eb2e8 memory: a world question must name what the memory mentions (V-470)
The score gate cannot separate the right note from an unrelated one: the
held-out fixture puts the right note at 0.791-0.890 and the must-be-silent cases
at 0.795-0.835, so a note about his slow network answered 'почему небо синее?'.

RecallAllowed adds a topic veto, and applies it only to a question that mentions
nothing of his. That restriction is the whole design: demanding a shared word of
every recall silenced four true recalls on the fixture to kill one false one,
because recall exists to find the note whose words he no longer remembers. A
question about his own life keeps the embedder as its only judge.
2026-08-03 13:50:54 +04:00
claude 4dfe106fe3 mavend: a question is never a fact about him (V-470)
IntentFact used to persist whatever the model invented for a question-shaped
utterance, at confidence 1.00, and index it for recall under the question's own
text. Two such rows then claimed seven unrelated world questions and silently
disabled world answering.

A question now goes down the query chain, which is what he asked for. The second
half is confidence: a value grounded in what he said stays 1.00, a value the model
supplied for words he never said drops to 0.60 and says so in the log. Same
reasoning as 'LLM output is not authorization' on the act path.
2026-08-03 13:40:33 +04:00
claude 2e0e2fd0bb router: a deterministic test for question-shaped text (V-470)
The predicate a fact write needs before it trusts a routing decision. Tokenized,
not substring: 'что' inside 'чтобы' is not a question. Capture verbs win over
every question signal, because 'запиши что я пил воду' contains an interrogative
and is still a capture.
2026-08-03 13:40:33 +04:00
claude f3fa6b353a store: voiding a fact drops its memory vectors (V-470)
Revert voided the fact row and left the vector, so recall kept serving the
voided fact's utterance and the documented repair reported success on a box that
stayed broken. There was no way to repair a poisoned box at all.

DeletePrefix covers every vector for the key, earlier rows included: their values
are superseded, and a superseded value has no business claiming a turn. It is
best-effort — the audit trail is already committed, and a fact that is voided but
still recallable beats a void that failed.
2026-08-03 13:40:13 +04:00
kami 6645f64c3e Merge pull request 'Name the gap: world questions through the workstation model, and the four remaining callers' (#98) from task/490-name-the-gap-world-questions-through-the into master
Reviewed-on: #98
2026-08-03 11:19:56 +02:00
claude f10e0068dd config, deploy: the workstation is workpc, not bugmachine (V-490)
Owner's correction. It is the same host CLAUDE.md already calls workpc, and
two names for one machine read as two machines. The dated eval file keeps the
old name: a measurement is never edited after the day it was taken.
2026-08-03 12:42:37 +04:00
claude 9b124d9194 docs: both halves of the degradation rule are wired, and which caller is which (V-490)
The offload inventory grows a column, because "seven callers of the resident
model" stopped being the useful fact. Which of them is offloaded, and under
which half of the rule, is. Three are resident-only on purpose and the table
now says why rather than leaving it to be rediscovered.

The three-outcome table is the part that was not obvious from the rule as
written. A configured-and-asleep workstation names the gap; a box with no
workstation block does not, because naming a gap requires a gap.
2026-08-03 12:29:12 +04:00
claude 12530c8a95 mavend: world questions ask the workstation, and name the gap when it is asleep (V-490)
queryGeneral has nothing fetched to fall back on, so it is the sharp case:
with a workstation configured and asleep he is told that, rather than told
something false in a confident voice. The 1.7B answering a world question is
where "Война и мир" got Левитан as its author.

The sources that already hold a passage — a live search, a ZIM article, a
page he named — go through the world model too, but read the passage back
when it is not there instead of naming a gap. A real quote beats "не могу
сейчас", and nothing is invented on either path.

The Stub and every test double keep the Phraser interface they have.
PhraseWorld is reached by assertion, and a phraser without it is the
no-workstation case.
2026-08-03 12:27:40 +04:00
claude 51256c4c9a phraser: test the three outcomes of a world question, and prompt parity (V-490)
The middle outcome is the whole task: a workstation that is configured and
asleep produces a gap, and the resident model is never asked. The parity
test compares the bytes PhraseWorld sends the workstation against the bytes
PhraseQuery sends the resident model, so the fixtures and the daemon cannot
measure two different prompts.

The nudge tests cover the silent half from both sides, including the
temperature, which is how the workstation would otherwise change how she
sounds without anyone deciding to.
2026-08-03 12:27:30 +04:00
claude 76481c2736 phraser: a world model seam, so a gap can be named instead of invented (V-490)
The naming half of the degradation rule in docs/offload.md. PhraseWorld has
three outcomes: no workstation configured means the resident model answers
exactly as today, a workstation that is taking work answers, and one that is
asleep returns ErrNoWorldModel so the caller can say so. Naming a gap
requires a gap — on a box that never had a second model, refusing every
world question would remove a capability he has now.

Both prompts move into knowledgePrompt and evidencePrompt, shared by
PhraseQuery and PhraseWorld, because prompt parity across two models stops
holding the moment there are two copies of a prompt.

The silent half comes with it: chatWithSystem and chatWithMessages prefer
the workstation when it will take work, at the same 0.7 the resident
transport samples at, and say nothing when it will not. That covers the
digestion worker's nudge and reminder phrasing without touching tick.go.

Only Available and CompleteRemote are in the Remote interface. Pair.Complete
has its own floor and the phraser already owns one; two floors under a
single call is one too many.
2026-08-03 12:27:30 +04:00
claude bcc2305cd0 llm: let a caller name its sampling temperature (V-490)
The phraser's own transport has always sampled at 0.7 and this client has
always been greedy. Routing a phrasing call through the client must not
change how it decodes, so Req carries the temperature and 0 — the zero
value, and what every existing caller wanted — is still greedy.
2026-08-03 12:27:04 +04:00
kami 0ceeac8df4 Merge pull request 'Point Maven at the workstation model: a workstation block, and routing plus replies through llm.Pair' (#97) from task/485-run-the-big-model-on-the-workstation-wit into master
Reviewed-on: #97
2026-08-03 10:13:09 +02:00
claude 4fae13af75 docs: record the workstation routing numbers where the router is documented (V-485)
CLAUDE.md carried only the homesrv figures, which now read as the whole story.
Also points offload.md's order at #490 for the naming half.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJYcaBiuny9UGSpFweQVQ1
2026-08-03 10:12:21 +02:00
claude 774217199e docs: measure gemma-4-12b on the workstation against the resident model (V-485)
Both fixtures, run from homesrv across the LAN with the proxy env stripped.
Routing: 84.4% full / 93.5% intent-only at p50 329ms through the cascade, against
72.7% / 77.9% at p50 0.80-1.04s for Qwen3-1.7B. Talk: 25/27 against 20/27, with
knowledge 9/9. Nudges 15/15. Settles #485's first assumption by measurement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJYcaBiuny9UGSpFweQVQ1
2026-08-03 10:12:21 +02:00
claude 2db59d52a7 deploy, docs: point homesrv at bugmachine and say what is still unwired (V-485) 2026-08-03 10:12:21 +02:00
claude 92d5fd580c mavend: route and reply through the workstation when its card is free (V-485)
modelSeam builds an llm.Pair when a workstation is configured and hands it to
the router and the replier. Both are the silent half of the degradation rule:
the big model is only better there, and he is never told which model answered.
No block, no probe, and the box behaves exactly as it did.
2026-08-03 10:12:21 +02:00
claude edeef19ff0 config: a workstation block, dropped when it names no address (V-485)
Health defaults to the supervisor's /health rather than llama-server's,
because mavgpud is what answers 503 while the card is held.
2026-08-03 10:12:21 +02:00
kami 018f7a6f47 Merge pull request 'Run the big model on the workstation, with admission control and the 1.7B as the floor' (#96) from task/489-workstation-deploy-mavgpud-on-workpc-and into master
Reviewed-on: #96
2026-08-03 10:12:13 +02:00
claude eca41798bd mavgpud: turn gemma's thinking off in the chat template (V-489)
Owner's call, 02-08-2026. Without it the 12B spends the reply budget on
reasoning tokens and answers empty at low max_tokens. Verified on the box:
"Столица Франции?" now answers "Париж" with no reasoning_content.
2026-08-02 22:44:10 +04:00
claude cc423567e7 docs: record that contention is KFD presence, not a VRAM threshold (V-489) 2026-08-02 22:29:58 +04:00
claude 8088ef9e00 mavgpud: build it with the rest, and ship the workstation config and unit (V-489)
make build now catches a broken supervisor on homesrv. deploy/mavgpud.json
carries the owner's gemma-4-12b line with the MTP draft model, passed to
llama-server untouched. The unit is a systemd user unit because sudo on the
workstation wants a password; lingering is the one command left to the owner.
2026-08-02 22:29:57 +04:00
kami 666b924d29 Merge pull request 'Run the big model on the workstation, with admission control and the 1.7B as the floor' (#95) from task/488-workstation-a-supervisor-that-keeps-llam into master
Reviewed-on: #95
2026-08-02 17:04:05 +02:00
claude e52c616592 mavgpud: test the probe against the sysfs the workstation actually has (V-488)
The fixtures are the live numbers sampled from the box on 02-08-2026, where the
CPT run held 12.8GB of 16 as proc/478104/vram_35881.

The cases that matter are the ones where a mistake is silent: our own
llama-server counting as a contender, an unreadable card reading as free, and
/health hanging or proxying into a closed port instead of answering 503.
2026-08-02 17:03:56 +02:00
claude 2b97bac51e mavgpud: keep the model loaded while the card is free, yield when it is not (V-488)
The lifecycle rule from Vikunja #488. Not on demand, because a 7-14B takes tens
of seconds to load and a world question would meet a gap every time the card
had been quiet. Not always on, because that is what holds the card.

/health is answered locally and always, so Maven's prober costs nothing and
works while the model is down. Everything else is reverse-proxied to
llama-server, which is what makes the idle window measurable at all.

Yielding is checked before starting, and both transitions are damped by a poll
streak so a short-lived rocm process cannot evict the model.
2026-08-02 17:03:56 +02:00
claude ab42db2b87 mavgpud: read the card from sysfs and own llama-server's lifecycle (V-488)
The workstation cannot keep a 7-14B resident: it would hold 16GB against the
owner's CPT runs, Correx and the manga-recap pipeline. So the process that
stays up costs no VRAM and the model comes and goes under it.

Contention is detected by presence on the KFD, not by a VRAM threshold. A ROCm
process registers under /sys/class/kfd/kfd/proc when it initialises HIP, well
before it allocates, so we see a contender during its startup instead of after
it has already lost an allocation race. rocm-smi is not installed on that box
and a per-second subprocess would get tuned down until useless, so this reads
sysfs and forks nothing.

Free VRAM is read only to decide whether to start. It is never a reason to
stop: by the time free VRAM has dropped, the other job has already failed.
2026-08-02 17:03:56 +02:00
kami 94d553570d Merge pull request 'Run the big model on the workstation, with admission control and the 1.7B as the floor' (#94) from task/485-run-the-big-model-on-the-workstation-wit into master
Reviewed-on: #94
2026-08-02 17:03:25 +02:00
claude 2e97b905b4 docs: the workstation supervisor owns llama-server's lifecycle (V-485)
The remote model cannot be a llama-server that is simply left running: a
resident 7-14B holds 16GB against the CPT runs the card is for. So what
is always up on the workstation is a supervisor, and llama-server is
loaded while the card is free.

Still not a scheduler. It arbitrates nothing between callers, and Maven
never asks it to start anything.
2026-08-02 18:14:31 +04:00
claude fbcca449be llm: pin that a down workstation is invisible (V-485)
Seven cases. The load-bearing ones are the constraint from 483: an
unconfigured deploy never probes and always reaches the floor, a busy
card degrades silently with the remote untouched, and a remote that dies
between probes still completes the turn and corrects the cached answer on
its way out.

CompleteRemote is pinned not to fall back, because a named gap that
quietly became a 1.7B guess is the failure this whole split exists to
prevent. And 1000 Available calls are pinned to make zero probes.
2026-08-02 17:19:53 +04:00
claude 2076e4a788 llm: prefer the workstation model, floor on the resident one (V-485)
Pair holds both models and decides which answers. A prober asks the
remote whether it will take work and caches the answer, so a request
reads an atomic bool rather than paying for a health check. Routing sits
at p50 825ms on the hot path and must never wait on a machine that may be
asleep.

The two methods are the two halves of the degradation rule in
docs/offload.md. Complete falls back silently, for routing, replies and
nudge phrasing, where the big model is only better. CompleteRemote
returns ErrRemoteUnavailable instead, for a world question, where the
1.7B does not answer worse but invents.

A nil remote is the unconfigured deploy: nothing probes, everything goes
to the floor, and the box behaves exactly as it does today.
2026-08-02 17:19:53 +04:00
kami 30eb6add1b Merge pull request 'Docs: refresh the QA plan against the live task list' (#93) from task/483-docs-offload-design into master
Reviewed-on: #93
2026-08-02 15:08:42 +02:00
claude dc266056d1 docs: the shape and the rules for offloading model work (V-483)
483 is an umbrella and its children are the work, so what it owes them is
the shape they must all obey. docs/offload.md records it: the degradation
rule and where its line falls, admission control rather than a GPU
arbiter, the embedder staying on homesrv because it backs the classifier,
and the inventory of what runs a model on the box today.

CLAUDE.md gets a pointer, because an agent about to add a model caller or
touch a daemon seam needs to know this before it starts, not after.
2026-08-02 15:08:35 +02:00
kami 1c786b7156 Merge pull request 'Docs: refresh the QA plan against the live task list' (#92) from task/483-design-offload-ml-to-the-workstation-kee into master
Reviewed-on: #92
2026-08-02 15:08:16 +02:00
claude a3af10a830 gitignore the root .env, it holds a live token (V-484)
It was untracked but not ignored, so one git add -A would have committed
MAVEN_AMBIENT_TOKEN. Same class as deploy/telegram.env, which is already
ignored.
2026-08-02 15:08:07 +02:00
claude c0de473382 ipc, worker: dial and bind through netaddr (V-484)
Five hardcoded transports, three in internal/ipc and two in
internal/worker, all now go through the seam address. The unix perms
logic moved into netaddr, so the two copies of parentDir and the umask
dance are gone.

peerCaller already returned ok=false for a non-unix conn, so the
SO_PEERCRED path degrades correctly on tcp with no change.
2026-08-02 15:08:07 +02:00
claude 3e534340bf ipc: pin that a scheme-less address still dials unix (V-484)
Six cases. The load-bearing one is the first: every deploy in the tree
writes a bare path, and it must keep meaning a unix socket with no
handshake in front of the payload.

The rest cover the tcp seam: a good token round-trips, a wrong one comes
back ErrUnauthorized, a stranger that speaks HTTP at the port is dropped
while the listener stays up for the next peer, and a tokenless tcp bind
fails rather than serving his turns to anyone who connects.
2026-08-02 15:08:07 +02:00
claude 1a704d704d ipc: a seam address that can name a transport (V-484)
internal/netaddr parses a daemon seam address and dials or binds it. A
scheme-less address is unix and behaves exactly as it does today: same
0700 parent dir, same 0600 socket, same bytes on the wire. tcp://host:port
is the new option, and it is what lets a module live on another host.

Over tcp the filesystem permission that authenticated the unix socket is
gone, and what crosses this seam is audio of the owner speaking. So a tcp
listener requires a shared token, checked in constant time before the
first protocol frame is read, and a peer that fails is dropped without
taking the listener down with it.
2026-08-02 15:08:07 +02:00
kami e57adcb001 Merge pull request 'Docs: refresh the QA plan against the live task list' (#91) from task/459-docs-refresh-the-qa-plan-against-the-liv into master
Reviewed-on: #91
2026-08-02 15:07:21 +02:00
claude bec7362b7b config: clear three of 472's five QA blockers (V-459)
morning_routines, feeds, crawl.on_demand and netscan.enabled in
deploy/mavend.json; -ambient-token on mavweb, interpolated from a gitignored
/.env. All verified on the box: the dispatcher builds the morning plan,
/api/ambient answers 401/201, the crawler reads a named page, netscan finds 3
devices, and /events fills with scan:lan and ambient:notif.

Filed 482 (ambient reads a notification's wall clock as UTC). Corrected 479:
both capabilities work once configured, so it is not a routing defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJYcaBiuny9UGSpFweQVQ1
2026-08-02 15:52:56 +04:00
claude a3ec746a01 docs: the six ready tasks all ran, and all six stop at the deploy (V-459) 2026-08-02 15:44:17 +04:00
claude af0eec250e docs: the operations sitting and the degraded-mode suite both ran (V-459) 2026-08-02 15:15:08 +04:00
claude 20aa2d59c9 docs: session 3 results and the query-source findings (V-459) 2026-08-02 14:40:22 +04:00
claude 4bad90dedb docs: point the QA findings at their new task ids (V-459) 2026-08-02 14:25:39 +04:00
claude 2b8d0f74fa docs: session 1 and 2 results, and the classifier baseline was wrong (V-459)
Ran sessions 1 and 2 on the live box.

Session 1 steps 1 and 3-6 pass. Steps 2 and 7-9 need a person at the box.
POST /api/chat is drivable with form encoding and a cookie jar, so the text
half needs no browser.

Session 2 confirms the deploy matches the bench at 72.7% full accuracy, and
contradicts two recorded numbers. The classifier scores 68.8% at p50 16.6us,
not 36.8% at 31ms. Router latency measured under contention again.

Also: 319 item 2 point 2 closes on the recall margin sweep, CheckFeminine has
a false positive on second-person masculine verbs, and the wake path cannot be
checked because mavwaked and mavenclient are deployed nowhere.
2026-08-02 14:24:01 +04:00
claude af9d2133dc docs: session 3 holds five sittings now, not three (V-459)
The refresh added the query-sources and operations sittings and left the
heading counting three.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJYcaBiuny9UGSpFweQVQ1
2026-08-02 13:57:01 +04:00
claude a1fdfccd61 docs: refresh the QA plan against the live task list (V-459)
The plan named 40 task numbers on 2026-08-01. Ten open QA tasks were missing
and two of the named ones had closed, so the 44-of-50 header was wrong twice
over.

- header is 42 of 50, and every open task now appears
- placed the ten unlisted QA tasks: 14, 248, 249, 250, 258, 283, 284, 285,
  286, 323
- new Operations sitting for 249 and 250, and a Query sources sitting for
  258 and 286
- 14 and 284 join housekeeping: both are gated on something unbuilt
- dropped the 317 and 354 rows, closed 01-08-2026, with one line saying what
  landed
- 319's gate recalibration is done; what is left is re-deriving QueryMinMargin
- 323 is down to the 60s startup timeout arm after PR #90
- new "Not this repo" section for 358 (Hexis) and 362 (training workspace)
- router latency is ~27x, not 90x; the 2.7s p50 was contention, not the model

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJYcaBiuny9UGSpFweQVQ1
2026-08-02 13:56:02 +04:00
kami 5c05163266 Merge pull request 'QA: phraser coverage is 65.3% but the llama-server subprocess lifecycle is 0% — the suspicion in this task was correct' (#90) from task/323-qa-phraser-coverage-is-65-3-but-the-llam into master 2026-08-02 11:39:20 +02:00
kami 92d2629001 Merge pull request 'Voice cannot accept a routine (V-367); the last three prompts are Russian (V-404)' (#89) from fix/367-voice-parks-routine-accept into master 2026-08-02 11:39:16 +02:00
kami bdcfccce77 Merge pull request 'Session workflow: pickup and wrap around the task flow' (#85) from task/445-session-workflow into master 2026-08-02 11:39:11 +02:00
kami f4deccacc9 Merge pull request 'dialogue.Slots and router.Slots are hand-kept copies that already drifted' (#88) from task/365-dialogue-slots-and-router-slots-are-hand into master 2026-08-02 11:39:07 +02:00
kami 8aaac01de6 Merge pull request 'Doc reorg: tier the tree, retire the three planning files' (#86) from task/446-doc-reorg-tier-the-tree-retire-the-three into master 2026-08-02 11:37:43 +02:00
claude feb6f2c03d phraser: test the spawn path, the one thing coverage never touched (V-323)
Every phraser test built the phraser with NewLLMPhraserAt, which starts no
process, so NewLLMPhraser, spawnLlamaServer, startLlamaProc, llamaProc.Close
and extractPort sat at 0% while the package headline read 65.3%.

These drive the real spawn code against a fake llama-server script: the port
scrape, the three reachable startup-race arms (start failure, stderr EOF,
context cancel), and Close actually reaping the child. The orphan test
re-execs the test binary as the daemon, SIGKILLs it, and asserts Pdeathsig
killed the grandchild. The last test rebuilds the production command line and
checks kill-maven.sh's pattern still matches it — that pattern has gone stale
twice and leaked orphans both times.

Package coverage 65.3% -> 76.9%. The 60s timeout arm stays untested; it needs
an injectable clock in production code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012YQGVXu5J1iCMCff5J4S1R
2026-08-02 10:15:35 +04:00
claude 99bb3526db Stop asking for Russian in English on the last three prompts (V-404)
#400 rewrote the chat and query prompts in Russian and left three pieces
of English prose behind.

PhraseReminder's user prompt was fully English. It is Russian now, and it
no longer restates the JSON contract or the persona rules: the call goes
through chat(), so nudgeSystem already states both, and a second copy of a
contract is one more thing that can drift out of step with the first.

querySystemPrompt and router.KnowledgePrompt both closed with the English
"Respond ONLY with valid JSON:". That sentence is prose instruction, not
wire format — the JSON skeleton after it is the wire format, and it is
unchanged. Kept rather than deleted: the GBNF grammar makes it close to
redundant, but the grammar is switchable off (phraser NoGrammar), and the
sentence is the floor when it is.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018CotYKycuio1GwLbYh9jfc
2026-08-02 10:01:58 +04:00
claude bb8cb8d014 Voice parks a routine proposal, it never accepts it (V-367)
Accepting a proposed routine gives the tick loop a standing new reason to
speak. DESIGN.md § "surface caps authority" puts that at layer 3, and says
voice is structurally incapable of layer 3 because a room mic is reachable
by anyone in the room. The /routines button was gated at step-up; the voice
path accepted outright. The two surfaces disagreed, so one of them was wrong.

A spoken "да" now leaves the row 'proposed' and sends him to /routines,
where the gated button is. A spoken "нет" still dismisses: declining does
not move the boundary outward, so voice keeps it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018CotYKycuio1GwLbYh9jfc
2026-08-02 09:58:07 +04:00
claude 5e66aa8f22 fix: the slots converter was still dropping the fact Value (V-365)
dialogue.Slots gained Value in 925ce22, but toDialogueSlots never copied
it, so a clarifying answer carrying a fact payload still landed nowhere:
clarify.go:202 sends the answer through the converter, and the SlotValue
arm reads answer.Value.

Both converters now carry every field. TestSlotsParity compares the two
field sets by name and type; TestSlotsRoundTrip populates every router
field and checks the round trip, and fails the fixture itself when a new
field is left zero.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QChoBS5qJSrCV98oNUnHNU
2026-08-02 09:35:47 +04:00
claude e332f167b2 docs: the ranking's two blocking infra items already shipped (V-447)
Checked the doable, epic and infra tiers against the code, not just the two
tiers V-447 asked about. Ten entries are already built. The two the ranking
calls blockers for everything below are among them: sqlcipher at-rest ships as
Store.enc plus OpenEncrypted, and mavweb/mavcaldav have nine test files
between them where the ranking says zero coverage.

Also built and still ranked as work: rule trace, recurring reminders (cron +
RescheduleReminder), stale-reminder burst collapse (collapseReminders),
revert (VoidLatestFact), digest mode, testing infra, passkey persistence.

Recorded as a section at the top of the archived file so the tiers underneath
are read with the corrections in hand. No tasks created: V-447 scoped task
creation to the mandatory and easy tiers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xwomr2cT93KMSz9u3digX5
2026-08-02 09:20:36 +04:00
claude 322401b9af docs: the ranking was stale, two of its gaps already ship (V-447)
Checked the mandatory and easy tiers against the code instead of trusting the
2026-07-03 ranking. Two were already built and their tasks closed unstarted:
quiet hours (QuietHoursConfig + the care gate in internal/loop/loop.go:37) and
schema migrations (internal/store/migrations.go on PRAGMA user_version, 12+
steps shipped).

Three more were narrowed to what is actually missing. Destructive-confirm has
a mechanism and no policy: store.Tool.Destructive is one boolean, not a risk
tier. Bounded follow-up state has dialogue.Session with a TTL and slot
inheritance; what it lacks is Candidates, so "второй" resolves against nothing.
Clarify has a gate that can ask and one hardcoded sentence to ask with
(internal/voice/replier.go:56, which replier_llm.go hands straight through).

The other six were confirmed absent: list_items, capability model, go.mod
tidy, conversation repair, command history, pronunciation dictionary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xwomr2cT93KMSz9u3digX5
2026-08-02 09:13:59 +04:00
claude 4f34a232d4 docs: retire the two root queues, archive the ranking (V-447)
PROGRESS.md and 20-07-2026-BACKLOG.md were state snapshots that git log and
the Vikunja board already carry. Everything PROGRESS.md claimed as shipped is
a QA task. The backlog's only untracked item, bounded follow-up state, is now
V-448.

maven-feature-ranking.md moves to docs/archive/2026-07-03-feature-ranking.md
instead of dying. Its mandatory and easy tiers became V-449 through V-458; the
doable and epic tiers are reasoning about why things are not worth doing yet,
which no task captures.

Four code comments and the design.md ledger pointed at the deleted files.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xwomr2cT93KMSz9u3digX5
2026-08-02 09:11:11 +04:00
claude 93987f2dfc docs: tier the tree by lifetime, so staleness shows in the path (V-446)
Seventeen markdown files at the repo root, twelve of them dated one-shot
reports sitting next to CLAUDE.md. That is why stale docs read as
current: nothing in the path said which was which.

Root now keeps CLAUDE.md and AGENTS.md. Living docs move under docs/
and carry a Last verified line. Dated measurements move to docs/evals/
ISO-prefixed, and are never edited after the day, so a newer number is
a new file. The senior review moves to docs/archive/.

Every reference was rewritten across markdown, Go comments, the Makefile
and the recall fixture. The touched Go packages still build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 03:28:49 +04:00
kami c0d61a71a4 session: pickup and wrap around the task flow, one disposable handoff (V-445)
task start and task pr already own the branch, the identity and the PR.
What was missing sat on either side of them.

pickup runs task start, reads TASK.md and any handoff, then restates the
assumption set and stops. That pause is the point: every wasted session
here began with an agent that inferred the goal instead of stating it
back. wrap runs the tests, updates the durable docs, commits in slices,
calls task pr, and records in Vikunja what task pr cannot know.

HANDOFF.md is gitignored and injected by a SessionStart hook. It holds
what the next agent needs to resume and nothing else. TASK.md is the
brief for the branch and does not change. Anything that would still
matter next week goes to Vikunja, CLAUDE.md or docs/.

CLAUDE.md documented none of this, which is why an agent would rebuild
it. It does now, including the two hooks in ~/.claude/hooks.

.claude/ was ignored wholesale. The workflow is now tracked, because how
a session behaves should be reviewed like code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 03:22:46 +04:00
kami 0b89294af7 hooks: refuse master, cap a code commit at 300 lines, require the task ref (V-445)
Two git hooks, tracked in .githooks and wired with core.hooksPath so a
fresh clone gets them with one config line.

pre-commit refuses master and refuses more than 300 changed lines in
non-markdown files. Markdown is exempt because docs land as one batch.
This is a commit-time guard, which diff-budget.sh is not: that hook
blocks the agent's edits and says nothing when either of us commits.

commit-msg requires (V-<id>), not (#<id>). Gitea autolinks # to its own
issues, and Vikunja is the tracker.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 03:22:46 +04:00
kami 7079a240f7 docs: what the first live search turn left open
Two things the deploy did not settle: the turn was slow off a cold start and
that number is not yet trustworthy, and the personal boundary still has never
run on the box.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 02:46:05 +04:00
kami 587f1e6a07 deploy: searxng on 9563, not 8080
8080 is taken several times over on this box, and the container name is the
only thing addressing it, so the port is ours to pick.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 02:37:45 +04:00
kami 99193ff1d1 docs: mark the two step-2 items done
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 02:33:59 +04:00
kami 63a389a1f8 phraser: make her read the sources instead of recalling them
The evidence branch of PhraseQuery framed every source as "твои заметки" and
joined them into one quoted run-on. A live search snippet is not his note, and
a run-on gives a 1.7B one blurred claim to merge rather than sources to answer
from. That is the shape that named Левитан as the author of Война и мир.

Sources now arrive numbered, one per line, and the system prompt says three
ways that the answer comes out of them: only from the sources, say plainly
when they do not answer, add nothing of your own.

Blank sources take the knowledge branch. One empty string used to reach the
evidence branch and ask the model to answer from an empty list, which is the
one prompt guaranteed to make it fill the gap from memory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 02:33:35 +04:00
kami 2150a18e98 deploy: ship the search block, so the live web is on for this box
Owner's call. The code default stays off — no `search` block still means no
query leaves the LAN — but the deployed config now carries one, so a question
that is not about him reaches SearXNG before it reaches the ZIMs.

Nothing runs at http://searxng:8080 on homesrv yet. That is the designed
degradation and not a broken turn: an unreachable instance falls through to
Kiwix and she never says the search failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 02:30:01 +04:00
kami 612ca8cf1b grammar: the last two model calls that were still free text
The replier and the meeting summariser were the two call sites without a
GBNF. Both are exactly the shape that makes a Thinking variant answer with
its reasoning as prose, and neither had anything downstream that could
remove it.

The replier already parses {"response","mood"}, so it now sends the phraser's
grammar for that contract, exported once as phraser.ResponseGrammar so the
two definitions cannot drift.

The summariser stays text-in/text-out. The JSON wrapper is attached and
unwrapped in the daemon's Completer, so internal/capture is unchanged and a
Completer without a grammar still works.

The simulator told routing from phrasing by "has a grammar", which stopped
being true here; it now looks for the intent enum.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 02:28:53 +04:00
kami 14e98334ad query: let her search the live web before she reads the ZIMs
The offline encyclopedia was the only world source, and it reads what was true
when the ZIM was built. A self-hosted SearXNG now asks first and Kiwix is the
fallback for an empty result, an unreachable instance or no line out. Owner's
ruling, 2026-08-02.

internal/websearch is deliberately thin: no rewriter (SearXNG ranks through
real engines, so the Russian question goes out as he asked it), no page fetch,
no cache. It cannot read the store, so only the query string can leave the box.

The personal boundary is unchanged and still sits above this source, so a
question about him is never searched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BFaeSbLMEVG5ey8tejU3y2
2026-08-02 02:24:36 +04:00
kami a103708a08 memeval: five minutes again, now that the gate keeps a turn from waiting
The budget was cut to 60s because a five-minute evaluation held the single
llama-server slot, and a voice turn arriving mid-evaluation waited behind it.
That collision is now solved where it belongs: the background client yields the
slot while a turn is in flight.

With the gate in place the short budget only truncates a Thinking model
mid-synthesis, which costs an observation and saves no latency on any real turn.
Kami's call, 2026-08-02.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BFaeSbLMEVG5ey8tejU3y2
2026-08-02 02:16:45 +04:00
kami a654b0126f docs: name the ecosystem, correct the latency, drop the stale handoff
Three documentation changes and one deletion.

CLAUDE.md and AGENTS.md gain the Nexus/Praxis/Hexis sections that were written
last session and never committed: what each service owns, where Maven's client
for it lives, and the rules that are not negotiable.

The p50 latency figure was wrong in two files. CLAUDE.md said the cascade costs
2.7s and that the LLM router is 90x slower than the classifier. Both come from
the bakeoff table, where the number is contention on a shared llama-server, not
the model. ROUTING-EVAL-31-07-2026.md line 61 says so and measures the router at
p50 825ms / p95 1.2s / max 3.0s. Corrected in CLAUDE.md, and the bakeoff table
now carries a header pointing at the routing eval for absolute latency. Latency
work was about to be planned off a number that was never real.

HANDOFF.md is deleted. It described work sitting on fix/integrated waiting for a
fast-forward onto overnight/eco-versioned-traces. Neither is true: master
contains that tip plus 22 commits, and both branch pointers are stale. The three
live defects it recorded move to PLAN-DETERMINISM-02-08-2026.md, which is now
the only planning document.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BFaeSbLMEVG5ey8tejU3y2
2026-08-02 02:16:37 +04:00
kami b8227295b8 query: stop asking the world about him
"во сколько у меня встреча" fell past the calendar, reached Kiwix, matched
an article on the 2015 CPISRA World Games and came back phrased as his
meeting. Inventing is worse than refusing, and as `system` this used to say
"пока не умею".

A new source sits between the notes pass and Kiwix: a question carrying a
possession marker ("у меня", "мой", "my", "do i have", "did i") that his own
data did not answer has no answer outside it either, so the walk stops there.
The markers are possession, not first person, so "как мне сварить борщ" still
reaches the encyclopedia.

This is also where CLAUDE.md's privacy line lands: only the utterance may
leave the box, and a question about him carries his life in the utterance.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 23:27:04 +04:00
kami b35151418a mavend: give the voice handler a CoreAPI that can serve the day plan
wireVoice runs before the tick loop exists, so it could only be handed
the bare store adapter — and that adapter answers DayPlan with "not
available via direct store API", because a day plan is assembled by the
tick loop and is not a table to read. So queryDayPlan, which the query
chain reaches for "какие у меня планы на сегодня", failed for every
caller on the deployed daemon.

main already back-patches the other direction (daemonAPI.chatFn =
handler.handleText). This is the same seam in reverse, at both wiring
sites. No recursion risk: nothing in the voice path calls api.Chat.

With the plan reachable, it recited its reminders as literal JSON. The
payload unwrapper existed but was private to the phraser, so the day
plan had its own non-unwrapping copy. One owner now, store.ReminderText,
with the phraser delegating to it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 23:17:10 +04:00
kami 17964d1162 dialogue: keep the remembered topic on the latest turn
rememberTurn runs after followUpMerge, which has already inherited a
Text slot from the previous same-intent turn, so the fill-if-empty rule
pinned the first topic of a run of query turns and never released it.
"во сколько у меня встреча", then "какие у меня планы", then "а завтра?"
continued the meeting — two turns stale.

Overwrite for system and query, where Text is a topic and not a payload.
A continuation is the exception and keeps what it inherited: its own
utterance is the ellipsis, and the topic it carries is the real one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:57:08 +04:00
kami 079cf689aa router: agenda questions belong to query, not to system
"что у меня сегодня" and "что у меня в календаре сегодня" both routed
IntentSystem on the deployed daemon, and replySystem has no agenda arm,
so both answered "пока не умею". The calendar source that can answer
them lives in the query chain and was never reached. The fixture has
said query since ru-query-019 was written; the daemon disagreed with the
fixture and the daemon was wrong.

AgendaQueryGrammars routes them at stage 0, after the clock rules so
"какой сегодня день" keeps reaching replySystem. Intent only — which
source claims the turn stays the query chain's decision.

This is what made the follow-up continuation look like it only worked
for "what day is it". It did: the query half inherited an intent whose
handler could not answer, so both halves came back "пока не умею".

Measured on the 77-case RU fixture: full accuracy 70.1% → 72.7%,
intent-only 75.3% → 77.9%, calendar 0/2 → 2/2, clarify counts unchanged.
The eval harness wires the new grammars too, or the fixture would stop
being a measurement of the daemon.

Go's \b is ASCII-only and never fires after a Cyrillic letter, which the
first version of the pattern learned the hard way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:53:01 +04:00
kami 9397f9e5f6 query: only a source that knows the day may answer "а завтра?"
The follow-up continuation re-aimed Slots.Time, but every query source
matches on dec.Utterance and nothing in the chain reads it. So the
inheritance bought almost nothing: "какие планы на сегодня" then "а
завтра?" missed day-plan's matcher and fell to the calendar, which had
parsed the day out of the raw utterance anyway.

Worse than nothing in one place. queryFactByKey runs first and claims on
HasKey plus HasTime, both of which the continuation sets, so a keyed
query continued with "а вчера?" answered "я записала это <the fact's own
timestamp>" and dropped the day entirely.

Widening the topic for every source would have made it worse still, not
better: CalendarEvents is the only CoreAPI call that takes a date, so
ten date-blind sources would have answered a question about tomorrow
with today's data. Gate instead of widen — a continuation is only
offered to a source that reads the day, and when none claims it she says
so rather than "не знаю", which reads as "nothing tomorrow" when she
never looked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:32:37 +04:00
kami 3f98a99f44 continuation: only an ellipsis may widen the keyword match
Deployed check, second round: "привет" after "какой сегодня день" answered
with the date. followUpMerge fills an empty Text from the previous
same-intent turn, so the topic-widening added a minute earlier was reading
an inherited topic on turns that had nothing to do with it.

Decision gains Continued, set only by continuation.go and never by the
router. replySystem widens on that and nothing else, so an inherited Text
is back to being invisible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:21:28 +04:00
kami 53616836db continuation: an ellipsis names the day, not the topic
Deployed check: "какой сегодня день" then "а завтра?" answered "пока не
умею". The intent was inherited correctly, but replySystem keyword-matches
the utterance, and "а завтра?" contains no topic word — that is the whole
nature of an ellipsis.

So the topic travels with the session. rememberTurn keeps the raw utterance
in Slots.Text for system and query turns that have no Text slot of their
own (a stage-0 grammar fills none), and replySystem matches keywords
against utterance + Slots.Text. Dates keep parsing from the utterance
alone, which is the part the ellipsis actually restates.

Only system and query: everywhere else Text is a payload and must stay
exactly what the router put in it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:16:58 +04:00
kami cf40f13573 continuation: a reminder's day is written into its payload, not just its slot
Live check on the deployed daemon: "напомни сегодня о событиях" then "а
завтра?" fires tomorrow with the text still reading "сегодня". Re-aiming
Time is not enough when the day word is also inside the payload, and
rewriting the payload needs the date's span in the string, which
ParseCalendarDate does not report.

Reminder comes out of continuableIntents until that exists. query and
system are unaffected: their Time slot IS the whole question.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:13:25 +04:00
kami 7d08d27efb mavend: answer "а завтра?" from the previous turn, not from the model
An elliptical follow-up carries no intent of its own. followUpMerge cannot
help — it inherits slots once the intent is known, and here the intent is
the missing part. So "а завтра?" went to the router, which on a 1.7B is
close to a coin flip, and the guess cost ~2.7s.

continuationDecision runs before the router and rebuilds the turn from the
previous one: same intent, same key, new day. Deterministic and free.

Three guards, all narrow on purpose. A parseable date is required, which is
what separates an ellipsis from an ordinary short utterance. Four tokens
max. And only query, system and reminder may be inherited: fact and note
would write something he did not say, and act would let a two-word
utterance re-run an allowlisted fn, which is a way to fire a destructive
command nobody typed.

A continuation is still remembered, so "а завтра?" then "а послезавтра?"
chains.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:09:05 +04:00
kami d0d0021659 mavend: let a fact close the nudge that asked for it
The snooze wire had one end. "готово" and "выпил воды" both left the nudge
pending, so the auto-tuner only ever learned from deferrals and from
silence — never from the rule working.

Two entry points, because the two utterances are different acts. A bare
"готово" carries no content and is intercepted before the router, sharing
pendingNudge and the twenty-minute window with the snooze. "выпил воды" IS
content: it routes normally, writes its fact, and only then closes the
nudge (ackFromFact, after applyAction). Folding the second into a pre-route
intercept would have thrown away the thing he actually said.

ackFromFact is silent. The fact reply stands; "отлично, отметила" on top
would be her congratulating him for obeying, which is the nag she is not.

Which fact answers which rule comes from the rule's own InertWhenNoData, so
a rule added later is covered the day it lands. "да" and "ок" stay out of
the ack vocabulary: the clarify gate upstream has the stronger claim on
them, and a stray "да" must not rewrite the feedback signal.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 21:43:28 +04:00
kami 79c3b994cf mavend: let him say "потом" to a nudge
A nudge could only be deferred from Telegram or the web UI. The voice path
had no route to store.ResolveNudge at all, so the channel she nudges on
hardest was the one he could not answer out loud.

resolveSnooze runs pre-route, right after the quiet toggle, and writes the
same `snoozed` outcome the buttons write — which also drops the row out of
RepeatUnacked, so a deferred sev4 stops re-sending every five minutes.

The window is what makes this safe to run before the router. "потом" is an
ordinary word; it only counts as a deferral when a pending nudge was sent
in the last twenty minutes, and otherwise the turn routes normally. Single
word patterns still match single-word utterances only, so "потом схожу за
водой" reports a plan instead of silencing the rule that prompted it.

Channel is not filtered: a nudge that went to Telegram is still what he is
answering when he says "потом" at the microphone.

QA-PLAN gains the two new manual checks and drops the 319 warning.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 21:40:06 +04:00
kami ba1d8e3f44 quiet: the noun form and the comparative are commands too
The pre-route toggle knew "тихий режим" and every negation of it, but not
the two phrasings that get spoken most: "включи режим тишины" (the setting
named as a noun) and "сделай потише". Both fell through to the router,
which has no quiet intent, so the command did nothing at all.

Adds those as stem pairs, plus a quietWordStems list so the
negated-but-unmatched fallback recognises "хватит тишины" the way it
already recognised "хватит тихого режима".

Locks the English phrasings from the routing fixture in the test table
("turn quiet mode back on", "turn off quiet mode", "stop quiet mode") —
all three already behaved, none were covered.

"на улице стало потише" and "в тишине лучше думается" stay inert: a
single-word pattern still only matches a single-word utterance.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 21:35:58 +04:00
kami 29329b5f0e router: a Russian verb is a whole sentence, not thin evidence
The clarify gate thinned any one-word utterance to 0.3 confidence, which
trips the stage-3 gate and comes back as "не совсем поняла". That is an
English intuition. Russian packs subject, tense and gender into one word,
so "поужинал" is a complete report and "привет" a complete greeting, and
both got clarified.

thinSingleToken keeps the rule for bare nominals, where it is real ("вода"
is a fact-or-query coin flip), and spares two classes: a closed lexicon of
social and control singles, and any token carrying a verb ending. Both
tests are offline.

Fixture: false clarifies 3 → 2, intent-only 74.0% → 75.3%, full accuracy
unchanged at 70.1%, missed clarify still 1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 21:29:21 +04:00
kami 47dda97226 tick: a disabled rule must not keep repeating its last alarm
`disabled_rules` stopped the loop from creating new nudges and did nothing
about the ones already sent. The sev4 repeat path does not consult the rule
set at all: RepeatUnacked re-sends any telegram nudge still at outcome=pending
every repeat_interval (5m by default), driven by store.UnackedTelegramRules.
So service_down kept arriving on a five-minute cadence after being switched
off, from a row written hours earlier — two messages after the deploy, which
is how it was found.

That cadence, not the unsealed database, is what "she keeps spamming me"
always was. The seal bug erased the acks that would have stopped it.

Filter the repeat keys against the wired rule set. Filtering on wired rather
than on the disabled list also silences a rule deleted from the code: nothing
can ack what the UI no longer lists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 20:43:27 +04:00
kami 8fdb9e5cd1 loop: let config turn a nudge rule off
The kuma service_down nudge cannot name the service. mavpoll folds the whole
monitor_status gauge into one boolean fact keyed `service_down`, and the
phraser names a service only when the fact key is the service name, so the
message is always the generic "a service on homesrv is down". Every fifteen
minutes, with nothing to act on. A fact per monitor is the real fix and it is
filed as Vikunja #444; this is what to do until then.

`disabled_rules` in mavend.json subtracts from loop.DefaultRules by name.
Config only subtracts — rules stay code, the set stays canonical and ordered
as written. A disabled rule is not gathered for either, since the gatherer
derives its key set from the rules it was given. Unknown names are ignored so
deleting a rule cannot brick a config that still lists it, and the boot log
says what was dropped, because a rule that vanishes silently looks exactly
like a rule that is broken.

Deploy turns service_down off.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 20:20:32 +04:00
kami f1a809121b shutdown: close the sockets, or the database never gets sealed
mavend seals its encrypted database in `defer st.Close()` when run() returns.
It had not returned since 2026-07-21. Every restart since then decrypted the
same eleven-day-old ciphertext and rolled back everything written in between:
the Telegram nudge that kept firing was a fact being un-written on each boot.

The goroutine dump named it. main → srv.Close() → ipc.(*Server).Close →
wg.Wait(), waiting on per-connection goroutines parked in readFrame. Close
shut the listener and nothing else, so the idle persistent sockets held by
mavweb, mavpoll, mavcaldav and mavmaild blocked shutdown forever. `docker
compose stop -t 60` spent the whole sixty seconds and then took a SIGKILL.

So: track the accepted conns and close them, in ipc and in voice, which had
the identical defect. Bound all three waits — the two per-server ones and the
worker wait in main — because the seal matters more than any single in-flight
call. A dropped RPC costs one reply; a missed seal costs a session.

The regression test leaves a client connected and idle, which is the case the
old tests avoided by closing the client first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 20:05:52 +04:00
kami 79893d646b kiwix: read the article, not the snippet, and state the persona as morphology
Two things that were each half-done.

The prompts handed the model copyable examples. chatSystemPrompt lost its
openers this morning and the "Я подумала, что" tic went with them, but "не
забыл ли я" appeared in its place: the removed example had been suppressing
the masculine self-reference by accident. Two predicatives are not enough
signal, so the rule is now stated as morphology (-ла) rather than as a pair of
words — a suffix rule generalises where an example only gets copied.
querySystemPrompt had the same defect and gets the same treatment; its "вот что
я нашла: " opener is deliberate and stays.

internal/kiwix had no caller. actions_query.go said "once internal/kiwix is
wired into this chain" and that never happened. It is wired now, between the
notes pass and the web source: everything of his answers first, and only what
is left over is looked up. Off unless a `kiwix` block names a server and a book.

Reading the search snippet does not work. Kiwix builds it from wherever the
keyword matched, which on Wikipedia is the navigation box at the foot of the
page — the first version of this answered "что такое фотосинтез?" by reciting
"Ecological economics Ecological footprint Ecological forecasting …". Client
grows an Article method; the head of the article is the lead paragraph, which
is the definition the snippet was meant to be. Verified on the box: the same
question now answers correctly off the ZIM.

Only the rewritten query leaves the process. A test asserts it: a turn carrying
a stored note must not put that note in the search string.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 19:46:39 +04:00
kami c04c5eca9c phraser: stop handing the chat prompt's examples back as replies
The grammar examples in chatSystemPrompt were full clauses — ("я подумала",
"я рада") for her, ("ты сказал", "ты забыл") for him. A 1.7B copies those
instead of generalising from them.

Observed on homesrv 2026-08-01, in one session: all three chat replies
opened with "Я подумала, что ...", and one ended "...немного тревожусь.
ты сказал" — the second example pasted onto a finished sentence, which
reads as a truncation and is not one.

Contrastive pairs replace the openers, so the rule reads as a correction
rather than a template. The him-examples are dropped; the "ты" instruction
carries that on its own, and those two produced the worst output. A closing
line tells her not to echo the instructions, because a small model treats a
quoted string as licence to reuse it.

Verified after rebuild: three chat turns, no "Я подумала" opener, no
dangling example, feminine forms intact.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 18:01:16 +04:00
kami bfdbe0045e telegram: send through the x-ui relay instead of straight at the API
Direct egress to api.telegram.org does not work from homesrv, so every
away-channel send timed out and the service_down nudge retried once a
minute forever. The sink already had a Proxy field wired to
http.Transport.Proxy; nothing had ever set it.

The relay is the x-ui socks inbound on the host, port 10808, addressed from
the container as the maven_default bridge gateway. That also needs a ufw
rule, because the bridge subnet is not otherwise allowed to reach a host
port and the SYN is dropped rather than refused. The rule is recorded in
the config next to the address, since the address alone is not enough to
reproduce this on another box.

Verified: five minutes after restart, zero send errors where there was
previously one per minute.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 17:22:12 +04:00
kami d09954d85d telegram: keep the bot token out of the log, and correct two design claims
net/http wraps every transport failure in *url.Error, whose Error() prints
the request URL. Telegram accepts the bot token nowhere but the URL path, so
a send failure wrote the live token into the daemon log. On 2026-08-01
homesrv could not reach api.telegram.org and did that once a minute for as
long as the network stayed down. The token lives in deploy/telegram.env to
stay out of the repo; putting it in `docker compose logs` undoes that.

Both error sites now go through redact. The structural branch rewrites
url.Error.URL and keeps the type, so errors.As still matches; anything else
falls back to scrubbing the rendered message. No minimum-token-length guard:
a one-character token would shred the message, but that beats leaking it.

DESIGN.md still said the classifier cascade was the path that runs today
with llmrouter wired nil, and gave the resident checkpoint as Qwen3.5-0.8B.
Both stopped being true on 2026-07-31.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 17:11:11 +04:00
kami 7e402b279d models: drop the self-referential stt/tts symlinks
7ab9b48 committed models/stt and models/tts as symlinks to their own
absolute paths:

    models/stt -> /home/kami/apps/Maven/models/stt

They came from an agent worktree under scratchpad/wt, where a link back
to the main checkout resolves. In the main checkout it points at itself.

Both paths are gitignored, so checking out that commit overwrites the
real model directories without warning and git says nothing. On homesrv
it destroyed models/stt/ggml-small.bin and the piper voice, and mavsttd
crash-looped on the missing whisper model.

The directories are host state fetched separately, per deploy/README.md.
Nothing under models/stt or models/tts belongs in git.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 16:59:33 +04:00
claude 6915e6a714 Land the 35-PR overnight stack (PRs 50-84)
One linear chain of 35 PRs, reviewed and fixed. The eleven fix branches were merged onto fix/integrated and fast-forwarded onto this tip, so the review findings land as commits here rather than on the individual PRs.

make test and make build pass.
2026-08-01 14:50:26 +02:00
217 changed files with 15154 additions and 3100 deletions
+160
View File
@@ -0,0 +1,160 @@
# Maven project dictionary for the direct-prose skill.
#
# These terms override every word preference in the skill's word-choice tables.
# Each entry exists because the name drifted in real docs or real answers, not
# because the word looked improvable.
#
# Format and the rule for adding a term: ~/.claude/skills/direct-prose/references/modes.md
terms:
resident_model:
name: resident model
meaning: the one always-warm Qwen3-1.7B llama-server that both routes and phrases
avoid:
- the model
- the LLM
- the 1.7B
- the phraser model
examples:
good: The resident model emits GBNF-constrained JSON.
bad: The 1.7B emits GBNF-constrained JSON.
router:
name: router
meaning: the stage that turns an utterance into a Decision with one of 7 intents
avoid:
- orchestrator
- intent classifier
- dispatcher
classifier:
name: classifier
meaning: the embedder nearest-neighbour path that runs when the router is off or errors
avoid:
- the fallback
- the floor
- the old router
examples:
good: A router error falls through to the classifier.
bad: A router error falls through to the floor.
cascade:
name: cascade
meaning: the ordered path stage 0, then router, then classifier
avoid:
- the pipeline
- the chain
- the fallback chain
stage_0:
name: stage 0
meaning: the deterministic rules that answer before the resident model is called
avoid:
- the fast path
- bypass
- deterministic assist
- preemption
examples:
good: Stage 0 routes agenda questions to IntentQuery.
bad: The bypass routes agenda questions to IntentQuery.
query_source:
name: query source
meaning: one entry in querySources, which either claims a turn or passes
avoid:
- arm
- handler
- branch
- answerer
examples:
good: Kiwix is the last query source before the model answers from memory.
bad: Kiwix is the last arm before the model answers from memory.
personal_boundary:
name: personal boundary
meaning: the query source that stops a question about him from reaching the world
avoid:
- the boundary
- the privacy gate
- the personal filter
clarify:
name: clarify
meaning: the turn outcome where Maven asks instead of acting
avoid:
- refusal
- rejection
- punt
examples:
good: The gate produced two false clarifies.
bad: The gate produced two false refusals.
fact:
name: fact
meaning: a keyed, supersedable row in the fact store
avoid:
- memory entry
- datum
- record
note:
name: note
meaning: free text he captured, indexed for recall
avoid:
- memo
- entry
memory:
name: memory
meaning: the embedded index over notes and facts that backs recall
avoid:
- RAG store
- vector db
- long-term memory
nudge:
name: nudge
meaning: one proactive message the digestion worker proposes and the dispatcher sends
avoid:
- suggestion
- proposal
- proactive prompt
- reminder
examples:
good: A fact can close the nudge that asked for it.
bad: A fact can close the suggestion that asked for it.
digestion_worker:
name: digestion worker
meaning: the background engine that consolidates memory and proposes nudges
avoid:
- digestion tick
- background engine
- reflection loop
reach:
name: reach
meaning: an outbound channel Maven speaks through, such as telegram, ntfy or voice
avoid:
- sink
- delivery channel
- notification backend
ecosystem:
name: ecosystem
meaning: Nexus, Praxis and Hexis together
avoid:
- the services
- the integrations
- upstream
act:
name: act
meaning: the intent that runs a capability through Hexis
avoid:
- action
- command
- execution
examples:
good: An act with no allowlisted fn is gated to a clarify.
bad: An action with no allowlisted fn is gated to a clarify.
+20
View File
@@ -0,0 +1,20 @@
{
"hooks": {
"SessionStart": [
{
"hooks": [
{
"type": "command",
"command": "f=.claude/prose-dictionary.yaml; [ -f \"$f\" ] && jq -Rs '{hookSpecificOutput:{hookEventName:\"SessionStart\",additionalContext:(\"Project prose dictionary. These terms override every word preference in the direct-prose output style. Use the name, never the avoid list.\\n\\n\"+.)}}' \"$f\" 2>/dev/null || true",
"statusMessage": "Loading prose dictionary"
},
{
"type": "command",
"command": "f=HANDOFF.md; [ -f \"$f\" ] && jq -Rs '{hookSpecificOutput:{hookEventName:\"SessionStart\",additionalContext:(\"An unconsumed HANDOFF.md is present. Run the pickup skill before anything else: read it, read the Vikunja task it names, restate the assumption set in at most five bullets, and wait for the user to confirm before writing code. It is a claim from the previous session, not truth. Delete it once consumed.\\n\\n\"+.)}}' \"$f\" 2>/dev/null || true",
"statusMessage": "Loading handoff"
}
]
}
]
}
}
+78
View File
@@ -0,0 +1,78 @@
---
name: pickup
description: Start a work session on a Maven task. Runs task start, reads the brief and the disposable handoff, restates the assumption set, and waits for correction before touching code. Use at the start of any session that continues earlier work, when the user says "pickup", "continue", "resume", or names a Vikunja task id.
---
# Pickup
The point of this skill is the pause in step 5. Every wasted session in this repo
started with an agent that inferred the goal instead of stating it back.
## 0. Get on the branch
```sh
task start <vikunja-id>
```
`~/.local/bin/task` owns the branch, the identity and the PR. It cuts
`task/<id>-<slug>` off `origin/master` and sets the commit author to the `claude`
gitea user. It writes `TASK.md` from the Vikunja task, and pulls any waiting
review comments into `.task/review-comments.md`. Do not hand-roll any of that.
`TASK.md` is the brief and it is immutable. If it says a PR already exists, this
is a review-fix session and not new work. Read the comments first.
## 1. Read the handoff
`HANDOFF.md` at the repo root, if it exists. It is gitignored, it belongs to one
session, and it holds only what is needed to resume. Treat it as a claim from the
previous agent, not as truth. It can be stale or wrong.
If there is no handoff, that is normal. It means the last session closed clean.
## 2. Read the durable state
In this order, and stop as soon as you have enough:
- The Vikunja task, by id. Project Maven is ID 2, MCP at `http://localhost:9100/mcp`.
The task description and its comments hold the goal, the constraints, and the
assumption ledger. This outranks the handoff on every conflict.
- `CLAUDE.md`, the section that covers the area you are about to touch.
- The one file under `docs/` that owns the area. Check its `Last verified` line.
If the sha is behind the code you are reading, say so in step 4 and trust the code.
Do not read the dated files under `docs/evals/`. They are measurements from one day,
never updated. Read one only when you need the number it recorded.
If no task id is known, ask for one before doing anything else. Work without a task
is work nobody can resume.
## 3. Look at the ground
`git status`, `git log --oneline -5`, and the diff on the current branch. What the
repo says beats what any document says.
## 4. Restate, then stop
Write at most five bullets and stop. Do not write code, do not open files to "check
one thing first", do not start with a small safe change.
```
Task: V-359, one line.
Done: what is already on the branch.
Next: the one thing this session does.
Constraints: what would make this wrong.
Assuming: the beliefs that, if false, waste the session.
```
Then ask: is this right? Wait for the answer.
A corrected assumption goes into the Vikunja task as a comment, not into the handoff.
The handoff dies tonight. The task does not.
## 5. Then begin
- Delete `HANDOFF.md`. It has been consumed and must not outlive this step.
- On master, cut the branch: `scripts/task-branch.sh <id> <slug>`.
- One task per session. When context passes roughly half, run `/wrap` rather than
pushing on. A compacted session is a session that forgot why it made a choice.
+94
View File
@@ -0,0 +1,94 @@
---
name: wrap
description: Close a Maven work session cleanly. Runs the tests, updates the durable docs, commits in reviewable slices with the Vikunja ref, pushes so the PR opens, records state in Vikunja, and leaves a disposable handoff only if work remains. Use when the user says "wrap", "wrap up", "done for now", or when context passes roughly half.
---
# Wrap
Run every step. A partial wrap is worse than none, because the next session trusts
the parts that did run.
## 1. Prove it works
`make test`. If something fails, fix it or say plainly in the handoff and in Vikunja
that it fails, with the output. Never wrap on an untested claim.
## 2. Update the durable docs
Ask what a future agent would have to learn the hard way, and write that down.
- `CLAUDE.md` when a fact an agent needs before touching code has changed: routing
behaviour, a measured number, a flag default, a constraint. A commit that changed
routing or phrasing without touching the matching CLAUDE.md section is a bug.
Correct stale text in place. Do not append a new paragraph next to the wrong one.
- `AGENTS.md` when the recipe to build, run or preview changed.
- The one file under `docs/` that owns the area, plus its `Last verified: <date> @ <sha>`
line. Only a doc directly under `docs/` carries that line.
- A new dated file under `docs/evals/` when you measured something. Never edit an
existing dated file. A newer measurement is a new file, and the living doc points
at it.
Nothing that must survive tonight goes anywhere else. Not into the handoff, not into
a commit message, not into a comment in the code.
## 3. Commit in slices
Under 300 changed lines per commit in non-markdown files, enforced by `.githooks/pre-commit`.
Markdown is exempt and may land as one batch.
Each commit is one idea, subject in the repo's voice, lowercase area prefix, and it
ends with the Vikunja ref:
```
router: narrow the single-token rule (V-359)
```
If a change genuinely cannot split under 300 lines, say why in the commit body before
reaching for `--no-verify`.
## 4. Land it
```sh
task pr
```
It refuses a dirty tree, pushes, opens or refreshes the PR against the repo default
branch, labels the Vikunja task in-review, comments the PR url on it, and pushes an
ntfy. Do not push by hand and do not call `tea` yourself.
## 5. Record what `task pr` cannot know
Comment on the Vikunja task: what you measured, what is still open. List every
assumption that turned out to be wrong. If the session found new work, create a task
for it now rather than describing it in prose.
This step is what makes the handoff disposable.
## 6. Leave the handoff, or leave none
If the task is finished, delete `HANDOFF.md` and stop. An empty root is the correct
end state.
If work remains, write `HANDOFF.md` with nothing but what the next agent needs to
resume, and no history:
```markdown
# Handoff — <date>
Task: V-359 <one line>
Branch: task/359-<slug>, cut from master
## Where I stopped
<two sentences, mid-thought detail that is nowhere else>
## Next action
<the single concrete next step>
## Do not
<the trap I nearly fell into, or the approach already ruled out>
```
Nothing else goes in it. No summary of what landed, that is in git and Vikunja. No
design rationale, that is in `docs/`. No fact an agent needs on any task, that is in
`CLAUDE.md`. If a line in the handoff would still matter next week, it is in the wrong
file.
+30
View File
@@ -0,0 +1,30 @@
#!/bin/sh
# Every commit names the Vikunja task it belongs to.
#
# router: narrow the single-token rule (V-359)
#
# V- and not #, because Gitea autolinks #359 to a Gitea issue, which is a
# different tracker and a wrong link.
#
# Exempt: merges, reverts, fixup/squash, and the initial commit.
msg_file=$1
subject=$(sed -n '1p' "$msg_file")
case "$subject" in
Merge\ *|Revert\ *|fixup!\ *|squash!\ *|amend!\ *) exit 0 ;;
esac
if [ -f "$(git rev-parse --git-dir)/MERGE_HEAD" ]; then
exit 0
fi
if printf '%s' "$subject" | grep -qE '\(V-[0-9]+\)$'; then
exit 0
fi
echo "commit-msg: subject must end with a Vikunja task ref." >&2
echo " got: $subject" >&2
echo " want: router: narrow the single-token rule (V-359)" >&2
echo " No task yet? Create one. Work without a task is work nobody can resume." >&2
exit 1
+30
View File
@@ -0,0 +1,30 @@
#!/bin/sh
# Two guards, both bypassable with --no-verify when you mean it.
# 1. master is not a working branch.
# 2. a code commit stays under 300 changed lines.
# Markdown is exempt from the size cap on purpose: docs land as one batch.
branch=$(git symbolic-ref --short HEAD 2>/dev/null)
case "$branch" in
master|main)
echo "pre-commit: refusing to commit on $branch." >&2
echo " task start <vikunja-id> # branch off origin/master, write TASK.md" >&2
exit 1
;;
esac
# Added + deleted lines across staged files that are not markdown.
# numstat prints "-\t-\t<path>" for binaries; those count 0 and that is fine,
# a binary blob is not the kind of diff this cap exists to stop.
loc=$(git diff --cached --numstat -- . ':(exclude)*.md' |
awk '$1 ~ /^[0-9]+$/ { a += $1 } $2 ~ /^[0-9]+$/ { d += $2 } END { print a + d + 0 }')
if [ "$loc" -gt 300 ]; then
echo "pre-commit: $loc changed lines in non-markdown files, cap is 300." >&2
echo " Split it. Each commit should be one reviewable idea." >&2
echo " git reset <path> to unstage, or --no-verify if this genuinely cannot split." >&2
exit 1
fi
exit 0
+18 -2
View File
@@ -9,6 +9,7 @@
/mavwaked
/mavmaild
/mavupdate
/mavgpud
# Certs (private keys, don't commit)
certs/
@@ -40,6 +41,9 @@ deploy/telegram.env
deploy/zenmoney.token
# IMAP password, read by mavmaild (never in argv, never committed)
deploy/imap.password
# Compose interpolation secrets — MAVEN_AMBIENT_TOKEN today. docker compose
# reads this file itself; it is not an env_file on any service.
/.env
# Temp files
/tmp/
@@ -50,7 +54,19 @@ opencode.json
# Test coverage output
coverage.out
# Agent worktrees and local agent state
.claude/
# Agent worktrees and local agent state. The workflow itself is tracked: the
# hooks, the skills and the prose dictionary are how a session behaves, so they
# get reviewed like code. Everything else under .claude/ is scratch.
/.claude/*
!/.claude/settings.json
!/.claude/prose-dictionary.yaml
!/.claude/skills/
# The disposable handoff. One session, then deleted. Never committed:
# anything worth keeping belongs in Vikunja, CLAUDE.md or docs/.
/HANDOFF.md
/models/stt
/models/tts
# root .env — MAVEN_AMBIENT_TOKEN and friends, same class as deploy/telegram.env
.env
-396
View File
@@ -1,396 +0,0 @@
beyond the model and tts work, the useful additions are mostly around **reliability, context, and reach**, not more intelligence.
## highest-value additions
### 1. unified event intake
maven should receive normalized events from:
* praxis
* calendar
* telegram
* local notifications
* system/service health
* manual checklists
* eventually email bridges
one internal envelope:
```go
type Event struct {
Source string
Kind string
EntityIDs []string
Title string
Body string
Priority string
OccurredAt time.Time
Payload json.RawMessage
}
```
this gives digestion one stable input instead of source-specific logic.
---
### 2. explicit morning routine engine — **core engine done (2026-07-20)**
`internal/morning` — pure checklist engine, mirrors `internal/loop`/
`internal/routine`'s no-I/O contract. `Evaluate(routine, facts, now)` answers
"what's still missing" any time (order-independent — checks facts, not
sequence); `Due(routines, facts, last, now)` fires the once-per-day nag only
at `NudgeAt` (defaults to window end) and only when something's unevidenced,
with a `last`-map dedupe identical in shape to `routine.Due`'s cold-start/
last-fire tracking. Evidence is just a fact timestamped inside today's
window — manual (voice-tapped) and inferred (another daemon writing the same
key) are indistinguishable, satisfying the manual/inferred requirement for
free. Weekday/weekend variants are two `Routine`s with different `Weekdays`
sets under different names. Wired into `config.MorningRoutineConfig` +
`cmd/mavend/tick.go`'s `fireMorningRoutines` (reads only the fact keys the
configured items reference, dispatches through the normal severity/presence
routing table, body is literal joined item labels — not LLM-phrased, same
no-hallucination rationale as cron routines). 13 unit tests in
`internal/morning/morning_test.go`.
Added since (2026-07-20, same day): a read-only `/morning` page in mavweb —
`ipc.CoreAPI.MorningStatus` (new wire method, mirrors `TickTrace`'s
daemon-cache-only shape: the store adapter errors, `daemonAPI` serves it from
a `tickLoop.morningStatus` closure) returns each routine's active/window/
per-item done state, server-rendered same as `/trace` (no live-update loop —
checklist state moves on minutes, not seconds).
Not yet done: no config wired in `deploy/mavend.json` (no morning routines
configured on homesrv yet — add items there when the medicine/water/pets
fact keys the phone/desktop write are settled), no voice query path for
"what did I miss this morning" (Evaluate supports it; nothing calls it yet),
no way to create/edit routines from the web UI — construction still means
hand-editing config, deliberately deferred: routines are operator-declared
config (like cron routines), and a CRUD editor would mean moving them to a
DB table + hot-reload, a bigger change than this pass.
not ordinary reminders.
support:
* required morning items
* order-independent completion
* soft time windows
* skipped-step detection
* one nudge, not repeated spam
* manual and inferred completion evidence
* weekend/weekday variants
example:
```text
08:0011:00
- medicine
- water
- pets
- check praxis attention
```
maven should know what is still missing, not merely fire four timers.
---
### 3. cross-device presence
**status (2026-07-20):** the hysteresis engine and 3 of the listed signals are
already built and wired live: `internal/store/presence.go` (noisy-OR combiner
+ Schmitt-trigger bucket resolve), fed by `desk_active` (workstation, via
`scripts/desk-active.sh` posting to `/api/signal`), `page_heartbeat` (mavweb
tab, `app.js`), and `wg_handshake` (`mavpoll` polling `wg show`) — threaded
into the tick loop via `internal/loop/gather.go`. Not done: phone-reachable,
homesrv-available, audio-output, and active-maven-client signals from the
list below are still missing.
a small presence daemon on each trusted device:
* workstation active/idle
* phone reachable
* homesrv available
* last keyboard/mouse activity
* wireguard presence
* current audio output
* active maven client
mavend receives only compact state, not raw activity logs.
useful for:
* choosing delivery channel
* suppressing voice while away
* surfacing reminders when you return
* knowing whether an agent result should be spoken or sent as text
---
### 4. interruption policy — **done (2026-07-20), turned out to already be built**
audited the existing code before writing anything new: `internal/loop.Gate`
already answers deliver_now vs. drop (quiet-hours/cooldown/snooze/presence/
calendar-busy), and `cmd/mavend/tick.go`'s `digestQ` + `config.DigestConfig`
already implement queue/digest (low-severity nudges batch into one
notification, flushed on window elapsed or max-items reached). The four
outcomes below were already covered by these two mechanisms; nothing new to
build for the core policy.
Gap that *was* real: `deploy/mavend.json` had no `digest` block, so batching
was disabled in prod despite being fully implemented. Fixed — see the config
change alongside this note.
before delivering anything, evaluate:
```text
urgency
current activity
quiet hours
recent nudges
available channels
whether already surfaced
```
result:
```text
deliver_now
queue
digest
drop
```
this prevents maven from becoming annoying once praxis and other sources start producing more data.
---
### 5. entity-aware memory — **done (2026-07-20)**
`03fa52d`/`9876187` (Vikunja #279): facts gain `Subject`/`EntityID`/
`ResolutionState`; an async enrichment worker resolves free-text subjects to
canonical Nexus entity_ids (mirrors Praxis's enrichment pattern). Ambiguous
or unreachable Nexus never guesses — the fact stays `pending` or terminal
`ambiguous`. Voice-tapped facts (`IntentFact`) now flow into the enrichment
queue automatically via an optional `Subject` field on `WriteFactReq` (old
callers unaffected).
Landed alongside this in the same session (not originally on this list, but
closes the plumbing gaps the last brief flagged for Nexus/Praxis maturity):
a typed Praxis lifecycle client (`398997f` — surface/acknowledge/resolve/
ignore/pin; fixes the surfaced≠acknowledged gap where reading an item aloud
left no trace), correlation-ID/version headers on the Nexus/Praxis clients
(`b743860`), entity-scoped Praxis attention queries (`0579ef9`), a durable
delivery outbox with begin-before-send/complete-after semantics
(`29f23e3`+`9ff726e` — closes a duplicate-send-on-crash bug), fail-closed
handling on ambiguous IPC mutation outcomes and Nexus/Hexis dependency
errors (`838fde1`+`d9fa4d6`), and a reusable fake-ecosystem test harness
with fault injection (`c932cd8`).
connect maven memory to nexus ids.
instead of:
```text
key = "кошачий фонтан"
```
store:
```text
entity_id = ent_pet_water_fountain
predicate = refilled_at
value = 2026-07-19T...
```
benefits:
* stable russian/english aliases
* fewer duplicate facts
* better “when did i last…” queries
* easier routine detection
* cleaner praxis correlation
---
### 6. bounded follow-up state
for short continuations:
* “yes”
* “tomorrow”
* “the second one”
* “not that project”
* “do it later”
store explicit pending state instead of relying on chat history:
```go
type PendingInteraction struct {
Kind string
Candidates []string
Args json.RawMessage
ExpiresAt time.Time
}
```
this matters a lot for a 1.7b model.
---
### 7. evaluation lab — **skipped for now (2026-07-20)**
runs on a different machine (GPU box), and CPT is currently in progress
there — deprioritized until the training pipeline has a checkpoint to gate.
Not abandoned, just off the immediate list.
before every new checkpoint or lora deploy:
* routing accuracy
* slot accuracy
* malformed json rate
* russian/english mixed input
* ambiguous entity handling
* reminder vs note vs fact
* direct answer vs tool call
* confirmation safety
* phrasing quality
* latency and ram
also replay real anonymized traces against old and new checkpoints.
this should be a hard deployment gate.
---
### 8. replayable full-system simulator
fake:
* clock
* presence
* caldav
* telegram
* praxis
* nexus
* hexis
* stt
* tts
* llama-server
scenario:
```text
08:30 user appears
08:35 medicine not completed
08:40 correx agent waits
08:45 calendar sync stale
08:50 user says “what did i miss?”
```
assert:
* what tools were called
* what was surfaced
* what stayed unresolved
* what maven said
* what was not executed
this will save more time than another feature daemon.
---
## useful second-wave additions
### voice session quality
* barge-in
* interrupt tts on wake word
* partial stt display
* confidence-aware clarification
* retry only failed stt segment
* per-room microphone profiles
* noise-floor calibration
* short response mode when speaking
### notification bridge framework
small adapters for:
* ntfy
* telegram
* matrix
* web push
* android notification forwarding
* local dbus notifications
normalize into maven/praxis events instead of treating each as a separate feature.
### local knowledge ingestion
* markdown/docs ingestion
* git repo summaries
* project decision records
* conversation exports
* provenance and source links
* incremental reindexing
keep this read-only and separate from personal fact memory.
### service self-diagnostics
`maven doctor`:
* socket reachability
* model health
* stt/tts readiness
* embedder availability
* caldav freshness
* telegram poll state
* praxis/nexus/hexis reachability
* db integrity
* disk usage
* recent failures
### config and secret management
* schema-validated config
* config migration
* secret references instead of inline values
* dry-run validation
* redacted config dump
* per-daemon health config
* startup dependency report
---
## things i would not build yet
* autonomous multi-step planning
* large external reasoner
* generic workflow engine
* self-editing memory
* automatic hexis actions from praxis
* emotion simulation beyond phrasing
* full home-assistant replacement
* more model layers before routing is stable
## recommended order
**status as of 2026-07-20:**
1. ~~evaluation lab~~**skipped, GPU-box work, deprioritized while CPT is in progress**
2. ~~entity-aware memory~~**done** (`03fa52d`/`9876187`, plus adjacent
Nexus/Praxis plumbing hardening — see item 5 above)
3. ~~morning routine engine~~**core engine done** (`internal/morning` +
`cmd/mavend` wiring — see item 2 above; not yet configured on homesrv,
no voice query, no web UI)
4. interruption/delivery policy
5. presence agents
6. unified event intake
7. full-system simulator
8. notification bridges
9. knowledge ingestion
10. voice-session polish
the main goal should be: **maven reliably knows what is happening, knows what you meant, and chooses the least annoying correct response**. everything else can wait.
+40
View File
@@ -5,6 +5,46 @@ This repo maps to **Maven** (project ID: 2) in Vikunja.
Feature work, bugs, deployment tasks all go here.
MCP endpoint: `http://localhost:9100/mcp` (or `http://192.168.1.104:9100/mcp` from workpc)
## The sibling services (Nexus, Praxis, Hexis)
Maven is the conversational front end of a four-service ecosystem. The other three
live in sibling repos next to this one.
| Service | Repo | Port | Answers |
|---|---|---|---|
| Nexus | `../nexus` | 9740 | who or what is this name |
| Praxis | `../praxis` | 8989 | what needs attention |
| Hexis | `../hexis` | 9741 | what can be run, and running it |
Division of labour: Nexus identifies, Praxis observes, Hexis acts, Maven understands
and coordinates. Maven is not the source of truth for any of the three. The full
contract is `docs/ecosystem.md`, and the constraints that bite during
implementation are summarised in `CLAUDE.md`.
Where things are in this repo:
- `cmd/mavend/ecosystem.go` holds `nexusClient` and `praxisClient`. The Hexis client
is vendored from `github.com/kami/hexis/pkg/client`.
- `cmd/mavend/ecosystem_acts.go` routes an act through capability discovery.
- `cmd/mavend/factenrichment.go` resolves each stored fact's `Subject` against Nexus
on a background poll loop, with backoff and no give-up.
- `internal/store/entityfacts.go` holds the entity-tagged fact rows.
- Config blocks are `nexus`, `praxis` and `hexis` in `deploy/mavend.json`. Each is
optional. Absent means that integration is dark, not broken.
Bring the whole ecosystem up locally:
```sh
docker compose -f deploy/ecosystem/docker-compose.yml up -d
```
That builds all three from the sibling working trees, so commit or stash there first.
Each publishes on loopback at the port above. Maven reaches them by service name on
the shared compose network.
Testing without them running: `cmd/mavend/fakeecosystem_test.go` provides stubs, and
`cmd/mavend/ecosystem_degraded_test.go` covers each service being unreachable.
## Rendering / previewing the web UI locally
To see mavweb pages with real data without touching the production stack:
+164 -14
View File
@@ -10,7 +10,7 @@ compose passes `/dev/dri` + the render gid) — the resident model stays ≤1.7B
**Resident model:** currently **Qwen3-1.7B** (`UD-Q4_K_XL`), stock — not yet the CPT'd one.
It replaced Qwen3.5-0.8B on 2026-07-31 because it measured better on both fixtures we have:
67.5% vs 59.7% intent-only on the 77-case RU routing fixture, and 20/27 vs 11-17/27 on the
talk fixture. See `MODEL-BAKEOFF-31-07-2026.md`. It is a Thinking variant, so `n_ctx` is 4096
talk fixture. See `docs/evals/2026-07-31-model-bakeoff.md`. It is a Thinking variant, so `n_ctx` is 4096
— reasoning tokens need the room, and 4096 is what the scores above were measured at.
The **target** is still the locally CPT'd **Qwen3-1.7B** (Vikunja #122, training in flight).
@@ -25,9 +25,24 @@ Spanish. Their strong published IFEval/BFCL numbers are English-only. Model file
`models/llm/`, so the LFM2.5 gguf sitting there is not loaded by anything. Swapping the resident
model is a one-line change to `phraser.model_path` in `deploy/mavend.json`.
See `REARCH.md` for the target architecture, `DESIGN.md` for the folded design spec, and
See `docs/rearchitecture.md` for the target architecture, `docs/design.md` for the folded design spec, and
`AGENTS.md` for local-preview + model-download recipes.
**Model work is moving to the workstation** (owner's call, 2026-08-02). homesrv cannot grow a
GPU and the workstation has 16GB of VRAM. So the resident model, STT and TTS become preferred
remotes with a floor on homesrv. The workstation is never assumed up. Fall back silently when
it would only do the job better. Name the gap when the 1.7B cannot do it at all. The embedder
stays on homesrv permanently, because it backs that floor. Read `docs/offload.md` before
touching a daemon seam or adding a model caller. Vikunja #483 is the umbrella, #484 to #487
are the work.
Both halves are wired as of 2026-08-03. Routing and replies prefer the workstation silently
through `modelSeam`; nudge and reminder phrasing prefer it silently inside the phraser. A
world question goes through `LLMPhraser.PhraseWorld` and names the gap when the card is not
free — `worldGap` in `cmd/mavend/worldmodel.go`, which he hears instead of an invented
answer. A box with no `workstation` block behaves exactly as it did before the seam: naming
a gap requires a gap. The offload table in `docs/offload.md` says which caller is which.
## Build & test
CGO daemons (`mavend`, `mavsttd`, `mavttsd`, `mavenclient`) need the vendored toolchain
@@ -68,13 +83,52 @@ Daemons are wired socket-to-socket, not linked. `internal/ipc` is the client/ser
protocol; the config in `deploy/mavend.json` (with `${VAR}` env expansion from gitignored
`deploy/telegram.env`) sets socket paths, model paths, and the phraser/embedder blocks.
## The ecosystem: Nexus, Praxis, Hexis
Maven is one of four services. It owns conversation and personal memory. It does not
own identity, operational state, or execution. Full contract in
`docs/ecosystem.md`.
```text
Nexus identifies. Praxis observes. Hexis acts. Maven understands and coordinates.
```
| Service | Owns | Maven's client | Configured at |
|---|---|---|---|
| **Nexus** | Canonical entity ids, names, aliases, relationships. Projects, services, devices, people, pets, places. | `nexusClient` in `cmd/mavend/ecosystem.go`, `POST /api/v1/resolve` | `nexus.url` (`http://nexus:9740`) |
| **Praxis** | Operational attention and item lifecycle. What needs looking at, what changed, what is still unresolved. | `praxisClient`, the HTTP tools API under `/api/v1/tools/` | `praxis.url` (`http://praxis:8989`) |
| **Hexis** | The capability registry and the only path to executing anything. | vendored `github.com/kami/hexis/pkg/client` | `hexis.url` (`http://hexis:9741`) |
All three are `nil` unless configured, and every one of them degrades on its own.
An outage means a named gap in the answer, never a broken turn and never a guess.
Rules that are not negotiable:
- **No component reads another component's database.** Praxis attention comes over
HTTP, never from its SQLite file.
- **Identity lives in Nexus.** Do not invent a local fact key for something Nexus
resolves. `actionFact` already sets `Subject`, and `cmd/mavend/factenrichment.go`
resolves it in the background against Nexus.
- **Free text never reaches a mutating Hexis call.** Resolve to a canonical entity id
first. Ambiguous resolution asks the owner, it does not pick.
- **LLM output is not authorization.** Confirmation binds capability id, target
entity, arguments, requester and expiry. See `cmd/mavend/confirm.go`.
- **Praxis lifecycle words mean different things.** Surfaced is not acknowledged,
acknowledged is not resolved, execution success is not recovery. Reading an item
aloud calls `Surface`, never `Acknowledge`.
- **No automatic attention-to-action path.** Digestion may summarise Praxis. It may
not call Hexis.
Every cross-service call carries a correlation id minted once per action
(`withCorrelationID`), a contract version header, and `X-Requested-By: maven`.
## Routing — read this before touching the router
`internal/router/` has TWO layered engines. **The LLM router is now the default and it is
on in deploy** — this section used to say it was wired `nil`, which stopped being true on
2026-07-31.
- **LLM router (the intended design, REARCH.md):** the resident Qwen3-1.7B (`llmrouter.go`)
- **LLM router (the intended design, docs/rearchitecture.md):** the resident Qwen3-1.7B (`llmrouter.go`)
emits GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is
demoted from a routing gate to a RAG hint. Wired at `voice.go:214` via
`pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient)`; the flag is `voice.llm_router`
@@ -88,10 +142,23 @@ on in deploy** — this section used to say it was wired `nil`, which stopped be
Cascade order: `stage0.go` exact-match fast-path → LLM router (when non-nil) → classifier
fallback. Any LLM error falls through to the classifier so a turn never breaks on the model.
Measured on the 77-case RU fixture (`MODEL-BAKEOFF-31-07-2026.md`): the classifier scores
36.8% full accuracy at p50 31ms; Qwen3-1.7B scores 67.5% intent-only / 72.7% through the
cascade at p50 ≈2.7s. Accuracy roughly doubled, latency is ~90× worse, and that trade was
accepted deliberately. `Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
Measured on the 77-case RU fixture. **Re-measured 2026-08-02: the classifier scores 68.8%
full accuracy at p50 16.6µs**, not the 36.8% at p50 31ms that stood here from
`docs/evals/2026-07-31-model-bakeoff.md`. That older figure predates the stage 0 rules and the
seed additions, both of which now score inside the classifier baseline. Qwen3-1.7B scores
77.9% intent-only / 72.7% through the cascade. So the router buys about 4 points of accuracy,
not a doubling, and the trade is worth re-arguing rather than assuming. **The ≈2.7s figure
that stood here until 2026-08-02 was contention, not the model.** See `docs/evals/2026-07-31-routing.md` line 61, which measures the LLM router at
p50 825ms / p95 1.2s / max 3.0s and the full cascade at p50 0.80-1.04s. Do not plan latency
work off the bakeoff table.
**The numbers above are the homesrv floor, not the ceiling.** With the workstation up, routing
completes through `llm.Pair` against gemma-4-12b and scores **84.4% full / 93.5% intent-only at
p50 329ms** — better than the resident model and about 2.5× faster (`docs/evals/2026-08-02-workstation-gemma4-12b.md`,
Vikunja #485). The workstation is never assumed up, so both sets of numbers are live. Judge a
routing change against the classifier and the resident model, since those are what always answer.
`Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
#359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with
no allowlisted fn) feeding the same stage-3 gate the classifier path already had — see
@@ -102,8 +169,25 @@ Re-measured on the fixture after the fix: **missed clarify 6/6 → 1**, at the c
clarifies and 2.6pt of full accuracy (72.7% → 70.1%, intent-only 67.5% → 74.0%). Two of the
three false clarifies are acts the model mis-routed and the gate caught — asking beats wrongly
executing, so the fixture and the daemon disagree about what is correct there. The third,
`"поужинал"`, is a real defect: **the single-token rule is an English intuition and does not
transfer to Russian**, where one word is routinely a whole sentence. Narrow or drop it.
`"поужинал"`, was a real defect: the single-token rule was an English intuition and does not
transfer to Russian, where one word is routinely a whole sentence.
Narrowed 01-08-2026. `thinSingleToken` (`internal/router/singletoken.go`) still thins a bare
one-word nominal — "вода", "бэкап" — but spares two classes: a closed lexicon of social and
control singles ("привет", "спасибо", "стоп", "yes"), and any token carrying a Russian verb
ending (past tense, 2nd person, reflexive), because a verb already contains its subject. Both
tests are offline and cost nothing. Re-measured: **false clarifies 3 → 2, intent-only 74.0% →
75.3%, full accuracy unchanged at 70.1%, missed clarify still 1.** The two remaining false
clarifies are the act-with-no-allowlisted-fn arm of the gate, not this rule.
Agenda questions taken off the model, 01-08-2026. `AgendaQueryGrammars` (`stage0.go`, wired
after the clock rules in `buildRouter`) routes "что у меня сегодня", "во сколько у меня
встреча" and anything naming a calendar to `IntentQuery` at stage 0. They were going to
`IntentSystem`, where `replySystem` has no agenda arm and answered "пока не умею" — the
fixture had said `query` since ru-query-019 was written. Measured: **full accuracy 70.1% →
72.7%, intent-only 75.3% → 77.9%, calendar 0/2 → 2/2**, clarify counts unchanged. Note that
Go's `\b` is ASCII-only and never fires after a Cyrillic letter; the pattern needs an
explicit `(\s|[?!.]|$)`.
## LLM output contract
@@ -129,18 +213,27 @@ world questions, so she needs to read external sources. What replaces it:
- **No telemetry, no cloud model, no third-party account.** That part never changes. Nothing
about Maven is reported to anyone, and inference stays on the box.
- **Local sources first.** Kiwix ZIMs on homesrv (Wikipedia, ifixit) before anything on the
network. Reading beats recalling for a small model, and a local read costs nothing.
- **His data first, then the world.** Every source that reads his facts, notes, calendar,
tasks or house runs before anything outside, and the personal boundary sits between them.
Reading beats recalling for a small model.
- **In the world, live search leads and the ZIMs are the fallback** (owner's call,
2026-08-02). A self-hosted SearXNG (`search` block) answers first; the Kiwix ZIMs on
homesrv answer when the search is empty, unreachable, or the line is down.
- **External search is allowed and off unless configured**, like the weather and telegram
capabilities.
capabilities. The code default is still off. `deploy/mavend.json` now ships a `search`
block (owner's call, 2026-08-02), so it is on for this box and deleting the block turns
it off again.
- **His notes and facts are never search input.** Looking up why the sky is blue and sending
his stored personal notes to an upstream engine are different acts. Only the utterance goes
out, never the persona block, history, or matched notes.
## Web UI conventions
Server-rendered pages share `cmd/mavweb/static/ui.css` (served at `/ui.css`) and the `nav`
partial (`navHTML` in `cmd/mavweb/main.go`, `{{template "nav" "<active-page>"}}`). No
Server-rendered pages share `cmd/mavweb/static/ui.css` (served at `/ui.css`) and the shell
partial in `cmd/mavweb/shell.html`: a page opens with `{{template "shellTop" "<page-key>"}}`
and closes with `{{template "shellBottom"}}`, and the key marks the active sidebar link.
Every page is its own embedded `.html` file next to `main.go` — no page markup lives in Go,
and the sidebar is data (`sidebarSections`, `pageIcon`) the template renders. No
per-page `<style>` beyond true one-offs. Wrap every table in `<div class=scroll>` so wide
data pans on a phone. Local preview + headless screenshot recipe is in `AGENTS.md`.
@@ -148,3 +241,60 @@ data pans on a phone. Local preview + headless screenshot recipe is in `AGENTS.m
This repo is project **Maven** (ID 2) in Vikunja. MCP: `http://localhost:9100/mcp` (or
`http://192.168.1.104:9100/mcp` from workpc). Feature/bug/deploy tasks go there.
Vikunja is the durable task store. A task holds the goal, the constraints and the
assumption ledger. Work without a task id is work nobody can resume, so a session that
has no id asks for one before it starts.
## Session workflow
`~/.local/bin/task` owns the branch, the commit identity and the PR. One task, one
session, one PR.
```sh
task start <vikunja-id> # branch off origin/master, write TASK.md, fetch review comments
task pr # push, open or refresh the PR, label Vikunja, notify
task comments # re-pull this branch's review comments into .task/
```
Around that, `/pickup` opens a session and `/wrap` closes it. Wrap at roughly half
context rather than letting the session compact.
Five stores, and each one owns something the others must not hold:
| Store | Holds | Lifetime |
|---|---|---|
| Vikunja task | goal, constraints, assumption ledger, status | durable |
| `CLAUDE.md`, `AGENTS.md` | what an agent must know before touching code | durable |
| `docs/` | design, measurements, decisions | durable |
| `TASK.md` | the brief for this branch, written by `task start`, immutable | one branch |
| `HANDOFF.md` | only what the next agent needs to resume | one session |
`TASK.md` and `.task/` are excluded through `.git/info/exclude`. `HANDOFF.md` is
gitignored and injected at session start. If a line in the handoff would still matter
next week, it is in the wrong file.
Docs are tiered by path, so staleness is visible from the filename. Files directly under
`docs/` are living and carry a `Last verified: <date> @ <sha>` line. Files under
`docs/evals/` are dated measurements and are never edited after the day, so a newer
number is a new file. Files under `docs/archive/` are dead and read by nobody by default.
## Git guards
Two hooks in `.githooks/`, tracked, wired with `core.hooksPath`. Fresh clone:
```sh
git config core.hooksPath .githooks
```
- `pre-commit` refuses master, and refuses more than 300 changed lines in non-markdown
files. Markdown is exempt and may land as one batch.
- `commit-msg` requires the subject to end with `(V-<id>)`. `V-` and not `#`, because
Gitea autolinks `#123` to a Gitea issue, which is the wrong tracker.
Two more guards live outside the repo, in `~/.claude/hooks/`. `diff-budget.sh` blocks
further edits past 600 changed lines on a `task/` branch. `prose_lint_hook.py` checks
prose on every write. Both measure against `origin/master`, so a local master that is
ahead of the remote makes the diff budget read high.
`--no-verify` exists. Using it means saying why in the commit body.
+10 -4
View File
@@ -16,11 +16,11 @@ PIPER_BIN := $(shell pwd)/deps/piper/piper
PIPER_MODEL := $(shell pwd)/models/tts/ru_RU-irina-medium.onnx
PIPER_ESPEAK := $(shell pwd)/deps/piper/espeak-ng-data
.PHONY: simulate stt-fixtures test-stt-golden all build build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav clean test fmt-check vet run-stt run-tts run-web download-embedder deps-go eval-router eval-recall eval-phrasing eval-models
.PHONY: simulate stt-fixtures test-stt-golden all build build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav clean test fmt-check vet run-stt run-tts run-web download-embedder deps-go eval-router eval-recall eval-phrasing eval-models build-gpud
all: build
build: build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav build-mail build-update
build: build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav build-mail build-update build-gpud
build-stt:
CGO_CFLAGS="$(CGO_CFLAGS)" CGO_LDFLAGS="$(CGO_LDFLAGS)" LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
@@ -59,6 +59,12 @@ build-mail:
build-update:
$(GO) build $(GOFLAGS) -o mavupdate ./cmd/mavupdate/
# mavgpud runs on the workstation, not here. It is built with the rest so a
# broken supervisor is caught by `make build` on homesrv rather than by the
# workstation refusing to serve. Copy the binary over, do not `make deploy` it.
build-gpud:
$(GO) build $(GOFLAGS) -o mavgpud ./cmd/mavgpud/
run-web: build-web
./mavweb -addr :9200 -voice 127.0.0.1:9100
@@ -78,7 +84,7 @@ deps-go:
done
$(GO) version
# fmt-check fails if any file needs gofmt. DESIGN.md has always said `make
# fmt-check fails if any file needs gofmt. docs/design.md has always said `make
# test` gates on gofmt and vet; it did not, so nine files quietly drifted.
# Run `gofmt -w` on whatever this prints.
fmt-check:
@@ -191,7 +197,7 @@ deps-piper:
# multilingual-e5-small: an asymmetric retrieval model. It is trained to match
# a short question against a longer passage, which is what note recall is.
# The quantized file is the one we download, deploy and measure — see
# RECALL-EVAL-31-07-2026.md.
# docs/evals/2026-07-31-recall.md.
EMBEDDER_DIR := $(shell pwd)/models/embedder/multilingual-e5-small
EMBEDDER_MODEL_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/onnx/model_quantized.onnx
EMBEDDER_TOKENIZER_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/tokenizer.json
-468
View File
@@ -1,468 +0,0 @@
## Maven — current state (updated 2026-07-20)
### Session 2026-07-20 — ecosystem hardening + entity-aware facts
Ten commits, focused on closing the Nexus/Praxis integration gaps flagged
as "wired but immature" in the prior review, plus the entity-aware-memory
backlog item (`20-07-2026-BACKLOG.md` item 5).
- **Entity-aware fact resolution (Vikunja #279)** — facts gain
`Subject`/`EntityID`/`ResolutionState`; an async worker resolves
free-text subjects to canonical Nexus entity_ids (mirrors Praxis's own
enrichment pattern). Ambiguous/unreachable Nexus never guesses — stays
`pending` or terminal `ambiguous`. Voice-tapped facts (`IntentFact`) flow
into the queue automatically via an optional `Subject` field on
`WriteFactReq` (old callers unaffected, no signature break).
- **Typed Praxis lifecycle client (Vikunja #271)** — `GetItem`/`Search`/
`Surface`/`Acknowledge`/`Resolve`/`Ignore`/`Pin`, routed through new RU/EN
dialogue verbs. Fixes a real lifecycle-invariant bug: reading an
attention item aloud now calls `Surface` — previously the digest path
read items without recording that they'd been surfaced, so "Maven
mentioned it" was indistinguishable from "never came up."
- **Durable delivery outbox (Vikunja #270)** — `BeginDeliveryAttempt`
before `Send`, `CompleteDeliveryAttempt` after; a stale `pending` row
found at startup reconciles to `unknown` (never silently resent or
dropped — same rule as Hexis's execution-timeout handling). Closes a
crash-window duplicate-send bug. Wired into `DispatchNudge`,
`DispatchReminder`, `RepeatUnacked`; reconciliation runs once at boot
before the tick loop resumes.
- **Fail-closed IPC/dependency handling (Vikunja #269, #272/#273)** —
ambiguous mutation outcomes (frame sent, reply lost) no longer blindly
retry; Nexus/Hexis dependency errors fail closed instead of guessing.
- **Correlation IDs + version headers (Vikunja #273)** — the hand-rolled
Nexus/Praxis HTTP clients now send `X-Nexus-Version`/`X-Praxis-Version`
and thread the same correlation ID already generated in
`executeCapability` through the whole call chain, matching the Hexis
client's existing behavior.
- **Entity-scoped Praxis attention queries** — callers holding a resolved
entity_id can ask "what needs attention for this entity" directly
instead of filtering the unscoped list client-side.
- **Fake-ecosystem test harness with fault injection** — a reusable
`fakeServer` (Nexus/Praxis/Hexis fixtures, runtime-toggleable
`SetFault`, fake clock) replacing ad-hoc per-test `httptest` servers;
covers a gap that had zero test coverage (`handlePraxisAct`) and adds a
fault-then-recovery regression test for the fail-closed fixes above.
- **Ops fix** — `deploy/mavend.json`'s phraser was pointed at a 4B model
with `n_gpu_layers=99`, which OOM'd under memory pressure and left a
zombie `llama-server` child; swapped to the 2B Qwen model matching the
intended resident-model size.
Net effect: the Nexus/Praxis wiring described as "plumbing exists, thin
compared to Maven's test depth" in the prior review is now materially
hardened — typed clients, fail-closed error handling, durable delivery,
and a proper fault-injection test harness are all in place. Evaluation lab
(`20-07-2026-BACKLOG.md` item 7) is explicitly skipped for now — it runs
on the GPU box, which is occupied by CPT. Morning routine engine (backlog
item 3) is next up, not started.
---
> **Resolved 2026-07-30 (task #318).** The resident checkpoint is
> **Qwen3.5-0.8B** (`Q4_K_M`), set in `deploy/mavend.json`; the **target** is
> the locally CPT'd **Qwen3-1.7B**, still training (#122). Older model claims
> below — the LFM references, the pipeline line, and the "swapped to the 2B
> Qwen model" ops entry above — are historical. Read them as a log of what was
> true at the time, not as current fact. Note also that `/mnt/hdd1/llms` is
> bind-mounted over `models/llm/`, so the LFM2.5 gguf in the repo tree is
> never loaded.
Architecture decision (as written on 2026-07-20): the target resident
router/phraser is the locally trained Qwen3-1.7B model — still the target as
of 2026-07-30. Older LFM references below describe the then-deployed
historical stack, not the target checkpoint. RU CPT has a successful
full-weight checkpoint at step 1000/8077; evaluation and Qwen3 SFT tooling are
tracked in `docs/plans/2026-07-18-qwen3-resident-training-eval.md`.
Consolidated status. The reactive↔proactive core is closed and testable through
the web PWA. The former SPEC's open items 17 (now `DESIGN.md` § execution ledger) are landed (protocol doc, away-channel
fallthrough, CalDAV poller, quiet-hours schedule, tools enable/disable, note RAG,
passkey step-up); item 8 (multi-user) is deliberately deferred — see the tail.
The two big infra gaps from the jul5 revision are closed on `overnight-jul5`:
**at-rest encryption** (AES-256-GCM, tmpfs working copy — not sqlcipher, see
`internal/store/crypt.go`) and **Docker deployment** (one image, six daemon
containers). The `overnight-jul6` session (now on `master`) closed the biggest
*query-surface* gaps — **calendar querying, general-knowledge answers, and
weather** — plus a populated homelab act allowlist and two pure scaffolds
(dialogue state, long-term-memory vector store). ~15.2k LOC + ~8.5k test, 303
tests, `-race` in `make test`.
### Access model
- **Phone** → needs the wg tunnel to reach homesrv (no homesrv DNS otherwise;
raw IP or a DNS tweak can bypass, not the default).
- **PC** → uses homesrv DNS, resolves the domains over local-net, **no wg needed**.
- nginx + ufw both scope to `10.42.0.0/24` (wg) + `192.168.1.0/24` (LAN), deny all else.
- **Surface in use now: the web PWA (`mavweb`).** Voice PTT + in-app nudges both ride it.
### Works end-to-end (tested)
- **Reactive voice:** PWA record → Whisper STT (`mavsttd`) → ONNX classifier →
resident phraser (llama-server subprocess; Qwen3.5-0.8B as of 2026-07-30 —
this line historically named "LFM 2.5-1.2B") → Piper TTS
(`mavttsd`) → reply.
HTTP POST path (mobile-Chrome drops WS for the audio).
- **Capture:** `fact` (EN **and RU** — root-substring recognizers) + `reminder`
persist through CoreAPI (`source=tap:voice`). This is the substrate the care
rules read.
- **Notes / query (semantic recall, sqlite — no chroma):** `note` → embed (the
classifier's ONNX embedder) → `notes` table. `query` → embed → brute-force
cosine top-k → confidence-gated (below `queryMinScore` 0.55 ⇒ "no note", not a
guess). **Note RAG (SPEC item 6):** the gated top-k feed the phraser
(`PhraseQuery`) to compose a natural answer ("вот что я нашла: …") instead of
a verbatim dump; raw-notes fallback on any LLM error. Stub is deterministic.
- **Monitoring (`/dash`):** mavweb server-renders presence + recent nudges (by
outcome) + recent facts from the append-only store via CoreAPI. Read-only,
meta-refresh, no JS.
- **Proactive loop:** 60s dumb ticker, pure predicates over a State snapshot,
universal gate (quiet-hours/presence/cooldown/snooze/calendar), one-nudge-per-
tick max-severity, reminders (gate-bypassing), sev4 repeat-til-ack, feedback
auto-tuner (outcome ratio → bounded cooldown, persisted as `source=feedback`).
- **Rules:** water/meal/break (sev12 care), service_down (sev4, `poll:uptimekuma`),
netdata_critical (sev3, `poll:netdata`).
- **Routines (`internal/routine`):** operator-declared clockwork — the third
proactive class beside reminders (user-stated) and care rules (world-state).
Config `routines[]` (cron + literal RU body + severity) fire through the normal
dispatcher on schedule (an 08:00 briefing, a 22:00 wind-down). Bodies are
literal (not LLM-phrased ⇒ can't hallucinate); rule name `routine:<name>` so
they don't pollute the care autotuner; cold-start guard seeds on first sight so
a restart never replays a missed schedule. Pure `routine.Due`, unit-tested; the
tick driver holds the last-fired map.
- **Env facts (`mavpoll`):** netdata alarms → `netdata_alarm` (fires immediately
on a real CRITICAL); kuma monitor_status → `service_down`. Writes only on
value-change (no append-only churn).
- **Presence:** noisy-OR decay + Schmitt hysteresis. Live via `page_heartbeat`
(PWA auto-pings `/api/signal` every 30s → present when a tab's open).
- **Delivery:** ntfy / telegram / voice by `f(severity, presence)`; minimal body
on away channels. PWA subscribes to ntfy over **WebSocket** for in-app nudges.
- **Away-channel fallthrough (SPEC item 2):** when the router picks voice but no
live session exists at push time (presence guess was wrong), the dispatcher
reroutes through the AWAY table — sev3→ntfy, sev4→telegram-repeat-til-ack,
sev≤2→drop — instead of silently dropping. Covers nudges + reminders.
- **Calendar busy (SPEC item 3, `mavcaldav`):** new poller queries a self-hosted
**Radicale** CalDAV server on an interval, writes `calendar_busy` + event facts
through CoreAPI (value-change only). The loop gate already consumes `calendar_busy`.
- **Quiet-hours schedule (SPEC item 4):** the gate reads `quiet_hours`; a config
time window (`voice.quiet_hours`, HH:MM, midnight-crossing handled) now sets it
on each tick — in addition to the "тихий режим" voice toggle. Both activate quiet.
- **Client protocol (SPEC item 1):** the voice wire format (length-prefixed JSON
frames) is published in `PROTOCOL.md`, generated from `internal/voice/wire.go`
so third-party clients don't need the Go source.
- **Passkey step-up (SPEC item 7):** `internal/webauthn` does real WebAuthn —
ES256/P-256 register + assert, ecdsa signature verification, rpIdHash + UP/UV
flag binding (UV = the gesture), sign-count regression check. `PasskeySession`
bumps the auth session L2→L3 for a TTL on assert. mavweb serves `/auth/passkey`
(enroll + step-up) + the begin/finish endpoints. Crypto is round-trip tested
(incl. tampered-sig / missing-UV / wrong-origin negatives).
- **Stability:** llama-server orphan leak fixed (`Pdeathsig` kills the child on
any mavend death); `kill-maven.sh` reaps strays (matches the model, not a
bogus `llama-server.*maven` pattern); `start-maven.sh` wires `-core` + poller.
### Wired but needs a deploy action (not code)
- **`desk_active`** (strongest presence signal) — `scripts/desk-active.sh` runs
on the **desk PC** (hypridle-gated systemd timer), posts over wg to mavweb.
- **`mavwaked`** (always-on listening) — needs a systemd user unit on a client
box (desk PC, pi, etc.) where the mic is attached. Connects to mavend over wg
or local net via `-addr`. Deferred until a client box is wired with a mic.
Caveats / gotchas:
- **desk_active is a workstation deploy, not code** — 0 facts ever written; presence
runs on page_heartbeat alone (dash reads "away"/"never at desk"). `scripts/desk-active.sh`
+ a hypridle-gated `maven-desk` timer must be installed on the desk PC (not homesrv).
- **Notes recall needs the ONNX embedder** — under the HashEmbedder floor, cosine is
lexical (token overlap), not semantic; scores are low, so most RU commands sit under
the 0.35 route threshold and clarify. Configure `voice.embedder` for confident recall+routing.
(The floor now at least tokenizes Cyrillic — see below — so it ranks correctly, just weakly.)
- **Switching the embedder model silently breaks old notes** — different dim ⇒
cosine 0 ⇒ they stop matching; brute-force can't re-embed. Re-embed on a model change.
- **`wg_handshake` is OFF and should stay off** — in this topology the phone only
runs wg when *outside*, so a fresh handshake means AWAY, not here. The `mavpoll
-wg` flag exists (defaults `""`) and could later back the spec's "away override"
by flipping the sign; as a presence-*here* signal it's inverted. desk_active +
page_heartbeat cover home presence.
- **Cold-start unlock tests are missing** — the key wrap/unwrap code
(`internal/webauthn/keywrap.go`) and locked-mode IPC gating (`cmd/mavend/main.go`)
are correct but have **zero test coverage**. The roadmap (item 2.1) required
three new test cases (wrap/unwrap round-trip, wrong-cred unwrap fails,
locked-mode IPC rejects non-unlock methods); none were written. `make test`
is green by omission. Write these before relying on the cold-start path with
real keys.
### Done since last revision (overnight-jul6, 2026-07-06)
Seven tasks (session board `SESSION-06-07-2026.md`, deleted 2026-07-30 — see git history), one commit each, merged to `master`.
This session was run through **opencode**, not Claude Code (co-author trailer).
Since then (**2026-07-06, second session**):
- **Always-on listening (gap 1, MVP)** — `cmd/mavwaked/`: 825 lines, 10 `-race`
tests. Energy-based VAD over 30ms windows (same RMS threshold as mavsttd's
`gateReason`), adaptive noise floor, speech→silence state machine. Captures
PCM from arecord(1) subprocess, sends `PushToTalk` with `Surface=SurfaceVoice`
(L0 — no destructive acts). Reply plays through aplay(1). No wake word yet
(pure VAD trigger); the 30ms frame shape matches silero-vad ONNX input 1:1,
so swapping energy-threshold for ONNX inference is a local change in vad.go.
`Makefile` `build-waked` target. Runs on client boxes (not docker/homesrv)
via systemd user unit; connects to mavend over wg or local net.
Since then (**2026-07-06, third session** — roadmap execution agent):
- **Cold-start unlock (ROADMAP 2.1)** — the at-rest AES key is now wrapped
(HKDF-SHA256 + AES-256-GCM, stdlib-only — no `x/crypto` dep) with the passkey
credential's public key and persisted to disk. At boot, if a wrapped key file
exists AND no env key is set, mavend starts **locked**: the IPC server runs
but `srv.Check` rejects everything except `MethodAssertStepUp` +
`MethodUnlock`. A passkey assertion at `/auth/passkey` calls `MethodUnlock`
with the credential's public key → unwraps the blob → opens the store → wires
voice/loop/delivery → `srv.SetAPI` swaps the locked stub for the real
CoreAPI. mavweb's `RegisterFinish` wraps the env key on enrollment;
`AssertFinish` calls `Unlock` on assertion. Env-key fallback preserved
(dev/CI path unchanged). **Test gap:** the roadmap required three new test
cases (wrap/unwrap round-trip, wrong-cred unwrap fails, locked-mode IPC
rejects non-unlock methods) — none were written. The code is correct but
untested; `make test` is green by omission, not coverage.
- **Conversation depth (ROADMAP 3.2)** — cross-intent anaphora + fact-by-key
lookup. `AnaphoraResolver` in `router/slots.go` detects RU pronouns
(это/он/она/оно/тот/мой + inflected forms). `followUpMerge` now handles
three cases: same-intent slot inheritance (existing), cross-intent anaphora
(Query/Fact/Reminder after a Fact with a pronoun inherits the prior key +
time), and query-after-fact (a query following a fact inherits the key for
fact-by-key lookup). `Session.History []Turn` added as the multi-turn
scaffold (capped at 4). 7 new test cases including the exact done-when
scenarios (anaphora query-after-fact, three-turn break, explicit-key-wins).
- **Routing quality + persona (ROADMAP 4.1/4.4)** — `QueryMinScore` is now a
config knob (`voice.query_min_score`, default 0.55) instead of a hardcoded
const. `make download-embedder` fetches Xenova/paraphrase-multilingual-
MiniLM-L12-v2 (~90MB ONNX) + tokenizer; AGENTS.md documents the embedder +
libonnxruntime setup. `Persona` field in `VoiceConfig` prepends to every
LLM system prompt (nudge phrasing, note queries, general knowledge); empty
= current hardcoded feminine-gendered Russian persona. Also fixed two
pre-existing data races found by `-race`: `voice/server.go` wg.Add vs
wg.Wait (accept mutex), `mavweb/server.go` s.api field (atomic.Value).
- **Calendar querying (task 3)** — "что у меня завтра?" now answers from the
CalDAV facts the poller already writes. Added `store.CalendarEvents(from,to)`,
a RU date-scope parser («сегодня»/«завтра») in `router/slots.go`, and an
IPC `CalendarEvents` RPC (api/client/server/wire) feeding the `IntentQuery`
handler. Empty day → «на сегодня ничего нет». Previously calendar only *gated*
nudges; it's now queryable.
- **General-knowledge routing (task 4)** — when notes-RAG misses `queryMinScore`,
the query now falls through to the phraser with an anti-hallucination system
prompt (`router.KnowledgePrompt`, single tested source) instead of giving up.
Empty/errored/Stub phraser → «не знаю.», never a fabrication.
- **Weather (task 5)** — new `internal/weather/`: `Provider` interface, a stub
(«погода не настроена»), and a real **keyless Open-Meteo** provider (geocode +
current_weather, injectable `*http.Client`, mocked in tests — no live network).
Wired into `IntentQuery` (keywords погода/градус/температура) with a ~5s
context timeout; selected by `voice.weather.provider` ("open-meteo" | "" → stub).
- **Homelab act allowlist (task 2)** — `voice.tools` seeded with read-only acts
(`systemctl status`, `docker ps`, `uptime`, `df`, `free`, `journalctl` reads)
as `destructive:false` and mutating ones (restart/stop/start/reboot,
docker-restart/stop) as `destructive:true`. Guardrail verified: no dangerous
verb is `destructive:false`. RU phrasings seeded in `act.txt`.
- **Embedder config validation (task 1)** — a partially-filled `voice.embedder`
block (some of model/tokenizer/lib paths missing) is now a load error instead
of a silent fall-through to the Hash floor; the floor fallback logs explicitly.
- **Dialogue state scaffold (task 6)** — `internal/dialogue/`: `Session` +
TTL `SessionStore` + pure `InheritSlots`. **Now wired** (post-merge follow-up):
the voice handler carries slots across same-intent turns within a 2-min window
(`followUpMerge`, unit-tested) — bounded gap-filling, not full multi-turn yet.
- **Long-term memory interface (task 7)** — `internal/memory/`: `Store` interface
+ `InMemoryStore` (cosine). Wired into `IntentNote` (best-effort insert) and,
post-merge, into `IntentFact` (facts indexed) + `IntentQuery` (read-back after
notes-RAG misses). In-memory only — no persistent backend yet (gap #8).
Follow-ups (Claude Code, post-merge): gofmt'd `handlers_test.go` (the jul6
verification commit left it misaligned, so `gofmt -l` still flagged it despite the
"all gates green" claim); deduped the task-4 knowledge prompt to the single tested
`router.KnowledgePrompt()`. Tree is now genuinely green (gofmt/vet/303 tests).
### Done since the jul5 revision (overnight-jul5, 2026-07-05)
The overnight session (`SESSION-05-07-2026.md`, deleted 2026-07-30 — see git history; 25 tasks) closed the previous
"not built yet" items 13 and added feature depth:
- **At-rest encryption** — the on-disk db is AES-256-GCM ciphertext; the daemon
works on a tmpfs (RAM) plaintext copy, sealed back atomically on close. Wrong
key / tamper ⇒ fail closed, never a plaintext fallback. Legacy plaintext dbs
upgrade on first clean shutdown. Key via config/env (`db_key_env`); no KDF —
raw 32-byte key, base64. The passkey cold-start unlock plugs into the same
`store.OpenEncrypted` seam later.
- **Docker deployment** — single image, one container per daemon
(`docker-compose.yml`); only mavend mounts the key + db volume; IPC over a
shared socket volume. `ipc.DialWait` (boot-order tolerance) + redial-on-drop
(core restarts don't kill modules). `deploy/README.md` has the runbook.
- **Tests** — mavcaldav, mavttsd, voicesink, mavweb main/handlers covered;
`make test` runs `-race -coverprofile`.
- **Recurring reminders** — `cron` + `next_fire_ts` on reminders; recurring ones
reschedule (instead of mark-fired) after successful delivery.
- **Notification digest/batching** — low-severity nudges queue and flush as one
digest per window/max-items (`digest` config block); stale-reminder bursts on
boot collapse into a single digest reminder, completed only after delivery.
- **Rule trace engine** — `ExplainTick`/`ExplainGate` record per-rule
predicate/gate/selection results each tick; served over IPC (`tick_trace`)
and rendered at mavweb `/trace` ("why didn't she nudge me").
- **Web UI** — new `/history` (facts + revert buttons), `/notifications` (nudge
history), `/trace` pages; nav links on `/dash`; RU/EN cheatsheet toggle in the
PWA; manifest icons (`icon.svg`). POST `/tools` now requires an in-process
passkey step-up when WebAuthn is configured.
- **Revert/undo** — `RevertFact` voids the latest fact for a key (append-only
void-marker, audit trail intact); exposed at `/api/revert` from `/history`.
- **Tool scopes** — `scope` column on tools, threaded through propose/enable/UI.
`DisableTool` raised to AuthStepUp alongside Enable.
- **Passkey persistence** — mavweb credentials in a JSON file (`-passkey-file`),
surviving restarts; rollback-on-persist-failure keeps memory and disk in sync.
- **STT silence gate** — min-duration + RMS floor drop non-speech before whisper
hallucinates on it (`-min-ms`, `-silence-rms` flags on mavsttd).
- **Housekeeping** — `db_key.env` gitignored (+`.env.example`), `build-caldav`
target, zero-timestamp "never" fix on /dash.
### Not built yet (ranked by ROI)
1. **Multi-user (SPEC item 8)** — deliberately deferred, see the tail.
Closed (jul6 follow-ups): `/api/revert` now sits behind the same passkey
step-up as POST `/tools`; `go.mod` direct deps (`onnxruntime_go`,
`coder/websocket`, `robfig/cron`) are labeled correctly — `go mod tidy` can't
run here because it walks the vendored `deps/go` toolchain tree.
Purge+rotate leaked db key (#12) — investigated and closed: the key was
**never committed** to git history (gitignored at introduction, no commit
ever tracked `deploy/db_key.env`), so nothing to scrub. File stays on disk
and in deploy env by design — at-rest encryption needs it at boot.
Done earlier (2026-07-03): **act tool executor, store-backed, full flow**
(`internal/tool` + `internal/store/tools.go` + `tools` CoreAPI methods).
- **Execution:** IntentAct runs the matched fn against the store's ENABLED
allowlist. argv, no shell → STT text can't inject. Live store read, so a
newly-enabled tool runs without a daemon restart.
- **proposed→enabled→disabled (SPEC item 5):** an act whose verb isn't enabled is
scaffolded as a `proposed` tool (maven suggests). A human enables it (fills argv
+ destructive) on the authed **`mavweb /tools`** page — never voice — and can
disable it back to `proposed` (kept in the store, won't run). `EnableTool`/
`DisableTool` sit at `AuthStepUp`; the gate is now **live** via `PasskeySession`,
so /tools enable requires a passkey assertion at `/auth/passkey` first.
- **Confirm turn:** a destructive enabled tool replies "выполнить X? да/нет" and
parks; the next utterance (ru/en yes-no) confirms or cancels (90s TTL).
- **Config:** `voice.tools` seeds enabled tools at boot (editing mavend.json =
the human enable act); mavweb enables ad-hoc ones on top.
- **Russian:** fixed grammar in reply strings + seed files; maven's self-
reference is feminine ("she") — [[maven-persona-gender]].
Also fixed:
- **HashEmbedder was blind to Cyrillic** (`tokenize` iterated bytes, kept only
`a-z0-9`) → every RU utterance embedded to the zero vector → cosine 0 across
all intents → misrouted to `act` (alphabetical tie-break). Now rune-based
(`unicode.IsLetter`). This was the real cause of "Найди заметку" (a query)
landing in `notes`; added note-retrieval query seeds too.
- **Notes are now browsable on `/dash`** — `RecentNotes` plumbed through the
store + CoreAPI; voice-captured notes were previously only reachable via
semantic `query`.
Earlier: notes/query recall, `/dash` monitoring, `wg_handshake` poller (NO-OP).
### Gaps — why "voice assistant" is still aspirational (2026-07-06)
What separates Maven today from the thing the spec describes. Dealbreakers
first — these define the category:
1. **Always-on listening is code-complete (MVP).** `cmd/mavwaked` captures
PCM from arecord → energy-based VAD → PushToTalk with `Surface=SurfaceVoice`
(L0). Gap narrowed: no wake word yet (pure voice-activity trigger; every
utterance fires). The 30ms frame shape and 16kHz PCM match silero-vad's
ONNX input exactly, so a wake-word model swap is a local change in vad.go.
Hardware: the mic lives on a client box (desk PC, pi, etc.) — never the
homesrv. Deploy action: systemd user unit on whichever box has the mic,
connects to mavend over wg or local net.
2. **Conversation is deeper now, still not full dialogue.** The router
classifies one utterance → one reply, but `internal/dialogue` carries
context across turns: a 2-min session inherits slots for same-intent
follow-ups («напомни завтра» → «…позвонить маме»), and cross-intent
anaphora («запиши что я пил воду» → «когда я это сделал?») now resolves
RU pronouns (это/он/она/оно/тот/мой + inflections) to the prior turn's
key for fact-by-key lookup. `Session.History []Turn` is the scaffold for
real multi-turn. Still missing: LLM-driven dialogue manager (decide
ask-vs-act), anaphora beyond RU pronouns, single-slot session (single-user
box). The sub-1B phraser only words replies.
3. **Latency/shape of a turn.** Clip-based STT (record → upload → whisper →
route → phrase → piper → play). No streaming either direction, no barge-in;
every exchange is a full round trip.
Capability-class gaps — built but thin:
4. **Act surface is a small argv allowlist.** propose→enable works and the
allowlist now ships a homelab starter set (jul6 task 2 — status/ps/uptime/
df/free/logs read-only, restart/stop/reboot gated). Still bounded to what's
seeded; broadening it is config, not code.
5. **Query answers now cover notes + calendar + weather + general knowledge**
(jul6 tasks 3/4/5). Calendar querying, keyless Open-Meteo weather, and a
phraser knowledge-fallback all landed; caveat — general-knowledge quality is
only as good as the sub-1B phraser, and weather needs `voice.weather.provider`
set. The cheatsheet and router are now roughly aligned.
6. **Routing quality depends on the ONNX embedder being configured** — the
HashEmbedder floor makes RU recall lexical/weak; many commands fall to
"clarify". `make download-embedder` now fetches the multilingual MiniLM
model + AGENTS.md documents libonnxruntime setup; `voice.query_min_score`
is a config knob (default 0.55) so the floor can be tuned without recompile.
7. **Presence is effectively one signal** (page_heartbeat); desk_active is
still an undeployed script — "voice when near" routing runs on a guess.
8. **Long-term memory is now persistent (store-backed), not the spec's chroma.**
`internal/memory` has a `Store` interface; the daemon now wires
`store.MemoryStore` (`internal/store/memory.go`) — a **persistent** backend
in the **same encrypted sqlite db** (survives restarts; recall text inherits
at-rest encryption, so no plaintext sidecar). Vectors are float32 blobs,
search is brute-force cosine (fine at single-user scale; ANN is the later
swap behind the same interface). Notes **and facts** are indexed on capture;
`IntentQuery` reads it back (after notes-RAG misses, before general-knowledge)
— fact recall («когда я пил воду?») is its distinct payoff. The in-memory
impl remains the test/no-store floor. Remaining: an ANN/external index is
optional-scale, not a gap. Custom TTS voice (kami-picked, replaces the irina
floor — [[custom-voice-training]]) is still a future item.
Ops footnote: voice-over-web verified 2026-07-06 — mavend binds 0.0.0.0:9100
and mavweb reaches it cross-container at mavend:9100 (nc -z confirmed).
mavpoll uses network_mode=host to reach localhost services (netdata, kuma).
### Future / logged, not now
Custom TTS voice training (kami-picked voice, replaces irina floor); listening
modes 23 (meeting-record, ambient-derive).
### Services & layout
- `mavend` (core, IPC unix socket) — store + loop + phraser; the only key-holder.
- `mavsttd` / `mavttsd` — STT/TTS worker modules (unix sockets).
- `mavweb` — PWA bridge (HTTP), `/api/ptt` voice, `/api/signal` presence ingest,
`/api/ntfy` WS-subscribe config, `/dash` read-only monitoring.
- `mavpoll` — env poller (netdata/kuma → facts via CoreAPI).
- `mavcaldav` — CalDAV poller (Radicale → `calendar_busy` + events via CoreAPI).
- All behind wg + nginx deny-all; no phone-home. CGo only in `mavsttd`.
- Start/stop: `./start-maven.sh [build]`, `./kill-maven.sh`.
- Config: `~/.config/maven/mavend.json` (or `mavend.json` in repo root).
### Key files
- `cmd/mavend/{main,tick,voice}.go` — daemon wiring, loop driver, voice handler
- `internal/loop/{loop,rules,gather,feedback}.go` — proactive engine
- `internal/store/` — append-only facts/reminders/nudges/presence/notes
- `cmd/mavweb/{main.go,dash.html}` — PWA bridge + `/dash` monitoring
- `internal/router/{classifier,slots,stage0}.go` — reactive routing + slot parse
- `internal/delivery/` — dispatcher + ntfy/telegram/voice sinks
- `internal/auth/` — scope/gate/policy; `FloorEnrollment` (same-uid = device
trust) + `webauthn.PasskeySession` (real step-up for L3)
- `internal/webauthn/`, `cmd/mavweb/webauthn.go` — passkey register/assert
- `cmd/mavcaldav/`, `cmd/mavpoll/`, `scripts/desk-active.sh` — env producers
### Why multi-user (SPEC item 8) is deferred
Not neglect — the one item where doing nothing now beats doing something:
- **No second user exists yet** (the "gf phase"). Building per-user partitioning
now means code exercised by zero users and validated by nobody — YAGNI.
- **The append-only schema makes it a migration, not a rewrite.** No row is ever
mutated, so adding `facts/notes/reminders.user_id` later is add-columns +
backfill-to-"kami" — no reshaping, no dual-write window. Deferral is cheap.
- **The hard part is speaker attribution, and it needs the second voice.** A
voice-print discriminator (kami vs gf vs unknown) can't be trained or tuned
with one voice in the house. Plumbing before the model is pipe with no water.
- **It's fenced deliberately** (`DO NOT TOUCH THIS PHASE` in `DESIGN.md` § Users) so an
autonomous agent doesn't add `user_id` columns while touching the store and
commit us to a schema before the constraints that shape it exist.
+1 -1
View File
@@ -1,6 +1,6 @@
// Package main is mavenclient — maven's reference client.
//
// Per DESIGN.md § Voice pipeline (STT / TTS): capture lives on the client;
// Per docs/design.md § Voice pipeline (STT / TTS): capture lives on the client;
// the server transcribes + synthesises on demand. The PC client runs the
// wake-word / VAD gate (cmd/mavwaked) and ships ONE clean audio blob per
// utterance on activation. The server never owns a mic.
+128
View File
@@ -0,0 +1,128 @@
// Spoken ack — the other half of the snooze wire. "готово" said out loud
// resolves a live nudge as `acted`, and a fact that answers the nudge on its
// own ("выпил воды" after the water rule fired) closes it without him having
// to say anything extra.
//
// Two entry points rather than one, because the two utterances are different
// acts. A bare "готово" carries no content and is intercepted before the
// router, exactly like the snooze. "выпил воды" IS content: it has to route
// normally and write its fact, and only then close the nudge. Folding the
// second into a pre-route intercept would have thrown the fact away, which is
// the thing he actually said.
package main
import (
"context"
"log"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
)
// resolveAck — pre-route keyword check for a contentless acknowledgement,
// run after the snooze. Same window and same fall-through rule: the words only
// count when a nudge is actually live, so "готово" with nothing pending routes
// normally.
func (h *reactiveHandler) resolveAck(ctx context.Context, text string, src turnSource) (string, bool) {
if !classifyAck(text) {
return "", false
}
now := h.now()
target, ok := h.pendingNudge(ctx, now)
if !ok {
return "", false
}
if err := h.api.ResolveNudge(ctx, target.ID, store.NudgeActed, now); err != nil {
log.Printf("voice: ack nudge %d (%s, %s): %v", target.ID, target.Rule, src, err)
return "не получилось отметить.", true
}
log.Printf("voice: acked nudge %d (rule %s) from %s", target.ID, target.Rule, src)
return "отлично, отметила.", true
}
// ackFromFact — post-action hook, called once the turn's decision has been
// applied. A fact whose key is the substrate of a live nudge's rule answers
// that nudge, so the nudge is resolved `acted` and the auto-tuner learns the
// rule is working.
//
// Silent by design: it returns nothing and never changes the reply. He said
// "выпил воды" and the fact reply is what he is owed; "отлично, отметила" on
// top would be her congratulating him for obeying, which is the nag she is
// explicitly not.
//
// Best-effort throughout. A failure here loses one feedback signal and must
// never turn a written fact into an error the user hears.
func (h *reactiveHandler) ackFromFact(ctx context.Context, dec router.Decision) {
if dec.Clarify || dec.Intent != router.IntentFact || !dec.Slots.HasKey {
return
}
rules := ackRulesForKey(dec.Slots.Key)
if len(rules) == 0 {
return
}
now := h.now()
target, ok := h.pendingNudge(ctx, now)
if !ok || !rules[target.Rule] {
return
}
if err := h.api.ResolveNudge(ctx, target.ID, store.NudgeActed, now); err != nil {
log.Printf("voice: ack nudge %d from fact %q: %v", target.ID, dec.Slots.Key, err)
return
}
log.Printf("voice: nudge %d (rule %s) acked by fact %q", target.ID, target.Rule, dec.Slots.Key)
}
// ackRulesForKey — which rules a fact under this key answers.
//
// Derived from each rule's InertWhenNoData rather than written out as a map,
// so a rule added later is covered the day it lands. That field already names
// the substrate the rule reads; a fresh fact under one of those keys is by
// definition the thing the rule was complaining about the absence of.
//
// DefaultRules, not the daemon's wired set: a rule disabled in config cannot
// have a pending nudge to close anyway, and reading the canonical set here
// keeps this free of the config plumbing.
func ackRulesForKey(key string) map[string]bool {
if key == "" {
return nil
}
var out map[string]bool
for _, r := range loop.DefaultRules() {
for _, k := range r.InertWhenNoData {
if k != key {
continue
}
if out == nil {
out = map[string]bool{}
}
out[r.Name] = true
}
}
return out
}
// ackPhrases — the acknowledgement vocabulary, as stem sequences. Matched by
// quietPhrase (quiet_toggle.go), so a single-word pattern matches only a
// single-word utterance.
//
// "да" and "ок" are deliberately absent. Both are answers to a question she
// asked, and the clarify gate upstream (resolveClarifyAnswer) has the stronger
// claim on them; letting them close a nudge as well would mean a stray "да"
// silently rewrites the feedback the auto-tuner learns from.
var ackPhrases = [][]string{
{"готово"}, {"сделал"}, {"сделано"}, {"выполнил"}, {"уже"},
{"уже", "сделал"}, {"уже", "готово"}, {"всё", "сделал"},
{"done"}, {"already", "did"},
}
// classifyAck reads an utterance as a contentless acknowledgement.
func classifyAck(text string) bool {
tokens := quietTokens(text)
for _, p := range ackPhrases {
if quietPhrase(tokens, p) {
return true
}
}
return false
}
+109
View File
@@ -0,0 +1,109 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
)
func TestClassifyAck(t *testing.T) {
for _, s := range []string{
"готово", "сделал", "сделано", "выполнил", "уже",
"уже сделал", "всё сделал", "done",
} {
if !classifyAck(s) {
t.Errorf("classifyAck(%q) = false, want true", s)
}
}
for _, s := range []string{
// "да" and "ок" belong to the clarify gate, not to the nudge.
"да", "ок", "хорошо",
// A single-word pattern must not eat the sentence it appears in.
"сделал бэкап базы", "готово ли обновление", "уже поздно",
"напомни завтра позвонить маме", "",
} {
if classifyAck(s) {
t.Errorf("classifyAck(%q) = true, want false", s)
}
}
}
func TestResolveAckMarksTheNudgeActed(t *testing.T) {
h, api := snoozeHandler([]ipc.Nudge{pendingNudgeAt(6, 2*time.Minute)})
reply, handled := h.resolveAck(context.Background(), "готово", sourceVoice)
if !handled || reply == "" {
t.Fatalf("got (%q, %v), want a reply", reply, handled)
}
if api.gotID != 6 || api.gotOutcome != store.NudgeActed {
t.Fatalf("resolved (%d, %q), want (6, %q)", api.gotID, api.gotOutcome, store.NudgeActed)
}
}
func TestResolveAckFallsThroughWithNothingPending(t *testing.T) {
h, api := snoozeHandler(nil)
if reply, handled := h.resolveAck(context.Background(), "готово", sourceVoice); handled || reply != "" {
t.Fatalf("got (%q, %v), want fall-through", reply, handled)
}
if api.calls != 0 {
t.Fatalf("resolved a nudge with nothing pending")
}
}
func TestAckRulesForKey(t *testing.T) {
cases := []struct {
key string
want string // "" means no rule
}{
{"water", "water"},
{"meal", "meal"},
{"break", "break"},
{"desk_active", "break"},
{"weight", ""},
{"", ""},
}
for _, tc := range cases {
got := ackRulesForKey(tc.key)
if tc.want == "" {
if len(got) != 0 {
t.Errorf("ackRulesForKey(%q) = %v, want none", tc.key, got)
}
continue
}
if !got[tc.want] {
t.Errorf("ackRulesForKey(%q) = %v, want %q in it", tc.key, got, tc.want)
}
}
}
func TestAckFromFactClosesTheMatchingNudge(t *testing.T) {
h, api := snoozeHandler([]ipc.Nudge{pendingNudgeAt(11, time.Minute)}) // rule "water"
h.ackFromFact(context.Background(), router.Decision{
Intent: router.IntentFact,
Slots: router.Slots{Key: "water", HasKey: true},
})
if api.gotID != 11 || api.gotOutcome != store.NudgeActed {
t.Fatalf("resolved (%d, %q), want (11, %q)", api.gotID, api.gotOutcome, store.NudgeActed)
}
}
func TestAckFromFactIgnoresAnUnrelatedFact(t *testing.T) {
// The live nudge is "water"; a meal fact does not answer it. Closing it
// anyway would tell the auto-tuner the water rule works when he ignored it.
h, api := snoozeHandler([]ipc.Nudge{pendingNudgeAt(12, time.Minute)})
for _, dec := range []router.Decision{
{Intent: router.IntentFact, Slots: router.Slots{Key: "meal", HasKey: true}},
{Intent: router.IntentFact, Slots: router.Slots{Key: "weight", HasKey: true}},
{Intent: router.IntentFact}, // no key
{Intent: router.IntentQuery, Slots: router.Slots{Key: "water", HasKey: true}},
{Intent: router.IntentFact, Slots: router.Slots{Key: "water", HasKey: true}, Clarify: true},
} {
h.ackFromFact(context.Background(), dec)
}
if api.calls != 0 {
t.Fatalf("resolved %d nudge(s) on unrelated decisions", api.calls)
}
}
+50 -14
View File
@@ -7,6 +7,7 @@ import (
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
)
// actionFact handles router.IntentFact: persist a tapped self-fact, index
@@ -15,14 +16,41 @@ func (h *reactiveHandler) actionFact(ctx context.Context, dec router.Decision) s
if !dec.Slots.HasKey {
return "не разобрала, что записать — попробуй иначе."
}
// A question is never a fact about him (#470). "какая последняя версия
// языка Go?" used to land here, and the value stored was whatever the
// model invented for it, at confidence 1.00, indexed for recall under the
// question's own text. Two such rows then claimed seven unrelated world
// questions through recall and silently disabled world answering.
//
// The routing error itself is not fixed here — the answer is to answer.
// Sending the turn down the query chain is what he asked for anyway, and
// it costs a mis-routed capture nothing: an explicit "запиши ..." is not
// question-shaped, so it never takes this branch.
if router.IsQuestionShaped(dec.Utterance) {
log.Printf("voice: fact write refused, utterance is a question: %q (key %q) — answering as a query",
dec.Utterance, dec.Slots.Key)
q := dec
q.Intent = router.IntentQuery
// The key the model extracted is its guess at what to store, not a
// fact he has. Left in place, queryFactByKey would read it back and
// claim the turn before any real source ran.
q.Slots.Key, q.Slots.HasKey = "", false
q.Slots.Value = ""
return h.actionQuery(ctx, q)
}
now := h.now()
req := ipc.WriteFactReq{
Ts: now,
Kind: "self",
Key: dec.Slots.Key,
Value: dec.Slots.Value,
Source: "tap:voice",
Confidence: 1.0,
Ts: now,
Kind: "self",
Key: dec.Slots.Key,
Value: dec.Slots.Value,
Source: "tap:voice",
// Not 1.00 unconditionally any more (#470). A value he said is
// evidence; a value the model supplied for words he never said is a
// guess, and writing a guess at full confidence is the same mistake
// the act path already refuses under "LLM output is not
// authorization".
Confidence: factConfidence(dec.Utterance, dec.Slots.Value),
// Subject: the key doubles as the entity-resolution candidate —
// a voice-tapped fact's key is usually the thing/person it's
// about ("espresso_machine", "kate"), so queueing it for Nexus
@@ -36,17 +64,25 @@ func (h *reactiveHandler) actionFact(ctx context.Context, dec router.Decision) s
log.Printf("voice: write fact: %v", err)
return "не получилось сохранить факт."
}
// Index the fact utterance in long-term memory (best-effort, must not
// fail the fact write). Facts aren't in the notes table, so this is the
// only recall path for them — "когда я пил воду?" reads back from here.
// Index the fact in long-term memory (best-effort, must not fail the fact
// write). Facts aren't in the notes table, so this is the only recall path
// for them — "когда я пил воду?" reads back from here.
//
// The indexed text is the fact, not the utterance (#493). queryMemory
// returns a fact's stored text verbatim, so what goes in here is what he
// hears; storing the utterance meant recall answered with his own sentence
// rather than the value. The utterance stays alongside as provenance —
// readable on /trace, never the answer and never embedded.
if h.memStore != nil {
if vec, err := router.EmbedPassage(ctx, h.embedder, dec.Utterance); err != nil {
text := store.FactRecallText(dec.Slots.Key, dec.Slots.Value)
if vec, err := router.EmbedPassage(ctx, h.embedder, text); err != nil {
log.Printf("voice: embed fact for memory: %v", err)
} else if err := h.memStore.Insert(ctx, "fact:"+dec.Slots.Key+":"+strconv.FormatInt(now.Unix(), 10), vec, map[string]string{
"source": "voice",
"type": "fact",
"text": dec.Utterance,
"ts": strconv.FormatInt(now.Unix(), 10),
"source": "voice",
"type": "fact",
"text": text,
"utterance": dec.Utterance,
"ts": strconv.FormatInt(now.Unix(), 10),
}); err != nil {
log.Printf("voice: memory insert fact: %v", err)
}
+322 -26
View File
@@ -5,6 +5,7 @@ import (
"errors"
"fmt"
"log"
"regexp"
"strings"
"time"
@@ -12,6 +13,7 @@ import (
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/morning"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/rss"
"github.com/kami/maven/internal/store"
@@ -41,6 +43,20 @@ type queryTurn struct {
type querySource struct {
name string
answer func(*reactiveHandler, context.Context, *queryTurn) (string, bool)
// dateAware — this source reads the day out of the turn and answers for
// THAT day. Only such a source may claim a continuation ("а завтра?"),
// because a continuation is a question about a different day and nothing
// else. A date-blind source claiming one would answer with today's data
// under tomorrow's question, which is a wrong answer delivered in a
// confident voice — the failure mode that took reminder out of
// continuableIntents (continuation.go).
//
// Exactly one source qualifies today, and that is not an oversight in the
// table: CalendarEvents is the only CoreAPI call that takes a date at all.
// DayPlan is today-only, CurrentWeather is now-only, and the recall
// sources search text with no notion of a day. When one of them grows a
// date parameter, flip its flag here.
dateAware bool
}
// querySources is the ordered chain actionQuery walks; first source to claim
@@ -49,68 +65,90 @@ type querySource struct {
// gate was never the bug. Adding a source (Kiwix, RSS, crawler, email) is one
// line here plus its method; where you put the line is the whole decision.
var querySources = []querySource{
{"fact-by-key", (*reactiveHandler).queryFactByKey},
{name: "fact-by-key", answer: (*reactiveHandler).queryFactByKey},
// Before "calendar" on purpose: both match "…на сегодня", and the plan is
// the more specific ask (its matcher requires a plan word), so the calendar
// listing would otherwise swallow it.
{"day-plan", (*reactiveHandler).queryDayPlan},
{name: "day-plan", answer: (*reactiveHandler).queryDayPlan},
// Also before "calendar": "что я обычно делаю по средам?" names a weekday,
// and the habit question is the more specific one. Its matcher requires a
// habit marker ("обычно", "каждый", …), so a question about this coming
// Wednesday still reaches the calendar.
{"habits", (*reactiveHandler).queryHabits},
{name: "habits", answer: (*reactiveHandler).queryHabits},
// Before "calendar" and before the recall sources: "что мне нужно
// сделать?" is a question about the task list, and the notes pass would
// otherwise answer it with whatever note happens to be nearest. Its
// matcher requires a task noun or an explicit "что … сделать", so a
// date-bearing question still reaches the calendar.
{"tasks", (*reactiveHandler).queryTasks},
{name: "tasks", answer: (*reactiveHandler).queryTasks},
// Before the recall sources too: "сколько я потратил?" is a question about
// the money facts the poller wrote, and the notes pass would otherwise
// answer it from whatever he once said about spending. Its matcher needs a
// money noun plus an actual ask, so "я потратил весь день" is untouched.
{"money", (*reactiveHandler).queryMoney},
{name: "money", answer: (*reactiveHandler).queryMoney},
// Before the recall sources and before general knowledge: "что нового?" is
// a question about the feeds she reads, and general knowledge would answer
// it by inventing news. Its matcher needs a feed noun plus an ask, so
// "у меня новая лента в инстаграме" is untouched.
{"feeds", (*reactiveHandler).queryFeeds},
{name: "feeds", answer: (*reactiveHandler).queryFeeds},
// Before "calendar" and before the recall sources: "что включено дома?" is
// a question about the house, and the notes pass would otherwise answer it
// from whatever he once said about the lights. Its matcher needs a house
// marker plus an ask plus a device word, and it bails out on weather
// wording, so "какая температура на улице?" still reaches the weather
// source.
{"home", (*reactiveHandler).queryHome},
{name: "home", answer: (*reactiveHandler).queryHome},
// Next to "home" and for the same reason: "какие устройства в сети?" is a
// question about the LAN, and the recall pass would otherwise answer it
// from an old note about the router. Its matcher needs a network word plus
// an ask plus a device noun, so "интернет не работает" is untouched.
{"network", (*reactiveHandler).queryNetwork},
{"calendar", (*reactiveHandler).queryCalendar},
{"weather", (*reactiveHandler).queryWeather},
{"embed", (*reactiveHandler).queryEmbed},
{"memory", (*reactiveHandler).queryMemory},
{"notes", (*reactiveHandler).queryNotes},
{name: "network", answer: (*reactiveHandler).queryNetwork},
{name: "calendar", answer: (*reactiveHandler).queryCalendar, dateAware: true},
{name: "weather", answer: (*reactiveHandler).queryWeather},
{name: "embed", answer: (*reactiveHandler).queryEmbed},
{name: "memory", answer: (*reactiveHandler).queryMemory},
{name: "notes", answer: (*reactiveHandler).queryNotes},
// THE BOUNDARY. Everything above answers from his own data; everything
// below answers from the world's. A question about him that got this far
// has no answer in his data, and no outside source can supply one, so this
// stops the walk rather than let the encyclopedia and the model guess.
{name: "personal", answer: (*reactiveHandler).queryPersonal},
// The world, read live. Owner's ruling of 2026-08-02: a metasearch hit beats
// a frozen ZIM, so SearXNG asks before Kiwix does. Nothing of his is at
// stake by this point — the boundary above already stopped every question
// about him, and only the query string leaves the box.
{name: "search", answer: (*reactiveHandler).querySearch},
// The offline encyclopedia, now the fallback for when the line is down or
// the search comes back empty. It reads the way it always did; what changed
// is that it no longer gets first refusal on a world question.
{name: "kiwix", answer: (*reactiveHandler).queryKiwix},
// LAST before the model answers from memory, and that position is the whole
// design (Vikunja #259): local sources first. His memory, his notes and —
// once internal/kiwix is wired into this chain — the offline ZIMs all get
// their turn before anything touches the network. The model does NOT: it
// design (Vikunja #259): everything of his, then the search, then the ZIMs,
// and only then a page he named. The model does NOT come first: it
// answers after this, because a URL he said out loud is an instruction and
// a 1.7B guessing at a page it cannot read is how contents get invented.
// This source only claims a turn where he named a URL, so it never competes
// with a local answer.
{"web", (*reactiveHandler).queryWeb},
{"general-knowledge", (*reactiveHandler).queryGeneral},
{name: "web", answer: (*reactiveHandler).queryWeb},
{name: "general-knowledge", answer: (*reactiveHandler).queryGeneral},
}
func (h *reactiveHandler) actionQuery(ctx context.Context, dec router.Decision) string {
t := &queryTurn{dec: dec}
for _, src := range querySources {
if dec.Continued && !src.dateAware {
continue
}
if reply, ok := src.answer(h, ctx, t); ok {
return reply
}
}
if dec.Continued {
// The previous question cannot be re-asked for another day. Saying so
// beats "не знаю", which reads as "no data for tomorrow" when the
// truth is that she never looked.
return "про другой день так не отвечу — спроси целиком."
}
return "не знаю."
}
@@ -352,6 +390,11 @@ func (h *reactiveHandler) queryWeather(ctx context.Context, t *queryTurn) (strin
if errors.Is(err, weather.ErrNotConfigured) {
return "погода не настроена.", true
}
if errors.Is(err, weather.ErrLocationUnknown) {
// He named a place and the geocoder does not have it. Saying so beats
// reading out the default city's temperature (Vikunja #421).
return "не знаю такого города — " + loc + ".", true
}
if err != nil {
log.Printf("voice: weather: %v", err)
return "не получилось узнать погоду.", true
@@ -396,6 +439,14 @@ func (h *reactiveHandler) queryMemory(ctx context.Context, t *queryTurn) (string
return "", false
}
text := hit.Meta["text"]
// The score cleared the gate and the topic still has to match (#470). A
// note about his slow network scored high enough to answer "почему небо
// синее?", because the right-note and must-be-silent score ranges overlap
// and no threshold sits between them.
if !memory.RecallAllowed(t.dec.Utterance, text) {
log.Printf("voice: recall %q rejected for %q: a world question and no shared topic word", text, t.dec.Utterance)
return "", false
}
// A note is phrased in Maven's voice; a fact is read back as it was
// stored.
if hit.Meta["type"] == "note" {
@@ -431,6 +482,12 @@ func (h *reactiveHandler) queryNotes(ctx context.Context, t *queryTurn) (string,
if !memory.ConfidentScores(noteScores, h.queryMinScore, h.queryMinMargin) {
return "", false
}
// Same topic veto as queryMemory above: the best note must be about what
// he asked, not merely the nearest vector in the index.
if !memory.RecallAllowed(t.dec.Utterance, notes[0].Text) {
log.Printf("voice: note %q rejected for %q: a world question and no shared topic word", notes[0].Text, t.dec.Utterance)
return "", false
}
texts := make([]string, len(notes))
for i, n := range notes {
texts[i] = n.Text
@@ -486,10 +543,7 @@ func (h *reactiveHandler) queryWeb(ctx context.Context, t *queryTurn) (string, b
// the question he actually asked. She answers the question, she does not
// recite the page.
snippet := page.Title + "\n" + crawl.TrimRunes(page.Text, webPageContextRunes)
reply, perr := h.phraser.PhraseQuery(ctx, t.dec.Utterance, []string{snippet})
if perr != nil {
log.Printf("voice: web: phrase: %v", perr)
}
reply := h.phraseSource(ctx, "web", t.dec.Utterance, []string{snippet})
if reply == "" {
// No phraser (or it failed): read back the top of the page rather than
// pretend the fetch did not happen.
@@ -498,11 +552,253 @@ func (h *reactiveHandler) queryWeb(ctx context.Context, t *queryTurn) (string, b
return reply, true
}
// queryGeneral — general knowledge from the phraser, the last source before
// giving up. It always claims: either the model answers or Maven says she
// doesn't know.
// kiwixTimeout — the whole ZIM source, rewrite included. The rewrite is one
// short constrained completion and the search is a LAN request; if the pair
// takes longer than this something is wrong and he is better served by the
// model's own answer than by more waiting.
const kiwixTimeout = 20 * time.Second
// searchTimeout — the whole metasearch source. websearch.Client already holds a
// per-request timeout from config; this is the outer bound on the turn, so a
// hung dial cannot outlive it either. Shorter than kiwixTimeout because there
// is no rewrite call in front of it: the question goes out verbatim.
const searchTimeout = 12 * time.Second
// querySearch — the live web, through a self-hosted SearXNG.
//
// Ahead of Kiwix by the owner's ruling of 2026-08-02: a search reads what is
// true today, a ZIM reads what was true when it was built, and the ZIM is the
// fallback for a box with no line out. Everything of his still answers first —
// the personal boundary is directly above this source, so a question ABOUT him
// never becomes a query.
//
// What leaves this process is the query string and nothing else. His notes, his
// facts, the persona block and the history do not travel with it: the websearch
// package cannot read the store. That is the CLAUDE.md rule made mechanical,
// not a promise about how the prompt is assembled.
//
// It claims the turn only when the search returns something. An empty result,
// an unreachable instance and a 403 from an instance without the JSON format
// all fall through to Kiwix, which is the point of the ordering.
func (h *reactiveHandler) querySearch(ctx context.Context, t *queryTurn) (string, bool) {
if h.search == nil {
// Off unless configured, same as the crawler and the ZIMs. Nothing is
// said about it: he never asked for a capability he did not enable.
return "", false
}
ctxS, cancel := context.WithTimeout(ctx, searchTimeout)
defer cancel()
// Verbatim. No rewriter: SearXNG ranks by meaning through real engines, and
// reducing "почему небо голубое" to English keywords would throw away the
// language he asked in along with the ranking that handles it.
resp, err := h.search.client.Search(ctxS, t.dec.Utterance, h.search.max)
if err != nil {
log.Printf("voice: search %q: %v", t.dec.Utterance, err)
return "", false
}
if resp.Empty() {
return "", false
}
// Logged on the way through, not only on failure. Without this there is no
// telling from the outside whether an answer came off the web, off a ZIM or
// out of the model's weights, and those are the cases worth telling apart.
log.Printf("voice: search: %q → %d answers, %d results", t.dec.Utterance, len(resp.Answers), len(resp.Results))
// Handed over the same way a note, a page or an article is: evidence for the
// question he asked, not something to recite. The trim is one budget over the
// joined block, so a long first snippet cannot crowd out the rest.
evidence := crawl.TrimRunes(strings.Join(resp.Snippets(), "\n"), h.search.runes)
reply := h.phraseSource(ctx, "search", t.dec.Utterance, []string{evidence})
if reply == "" {
// No phraser, or it failed. Read back the best evidence rather than
// pretend the search did not happen.
return "вот что я нашла: " + crawl.TrimRunes(resp.Snippets()[0], 300), true
}
return reply, true
}
// queryKiwix — the offline encyclopedia, and the fallback behind querySearch:
// everything of his has already had its turn and the live search found nothing
// or could not be reached. Reading beats recalling for a 1.7B either way.
//
// What leaves this process is the search query and nothing else. His notes,
// his facts, the persona block and the history do not travel with it — the
// kiwix package cannot read the store. That holds even though the server is on
// the LAN, because "local sources first" is not a licence to widen what a
// lookup is allowed to see.
//
// It claims the turn only when the search returns something. No results is not
// a failure worth announcing: it means the ZIM does not cover this, and the
// model answering next is the better outcome than "ничего не нашла".
func (h *reactiveHandler) queryKiwix(ctx context.Context, t *queryTurn) (string, bool) {
if h.kiwix == nil {
// Off unless configured, same as the crawler and the weather. Nothing
// is said about it: he never asked for a capability he did not enable.
return "", false
}
ctxK, cancel := context.WithTimeout(ctx, kiwixTimeout)
defer cancel()
// The ZIMs are English and kiwix ranks by keyword overlap, not meaning, so
// a Russian sentence matches nothing at all. The rewriter turns it into a
// handful of English keywords with the resident model.
pattern := t.dec.Utterance
if h.kiwix.rewriter != nil {
q, err := h.kiwix.rewriter.Rewrite(ctxK, t.dec.Utterance)
if err != nil {
// Fall through to the verbatim question rather than give up. It
// will usually miss, and missing is a fall-through too.
log.Printf("voice: kiwix: rewrite: %v", err)
} else if q != "" {
pattern = q
}
}
hits, err := h.kiwix.client.Search(ctxK, pattern, h.kiwix.book, h.kiwix.max)
if err != nil {
log.Printf("voice: kiwix: search %q: %v", pattern, err)
return "", false
}
if len(hits) == 0 {
return "", false
}
top := hits[0]
// Logged on the way through, not only on failure. Without this there is no
// way to tell from the outside whether an answer came off a ZIM or out of
// the model's weights, and those are the two cases worth telling apart.
log.Printf("voice: kiwix: %q → %d hits, top %q", pattern, len(hits), top.Title)
// The top hit only, read as an article rather than as a snippet. Kiwix
// builds its snippet from wherever the keyword matched, which on Wikipedia
// is usually the navigation box at the foot of the page — the first version
// of this joined three of those and she recited "Ecological economics
// Ecological footprint …" at him. The head of the article is the lead
// paragraph, which is the definition the snippet was meant to be.
page, aerr := h.kiwix.client.Article(ctxK, top.Path, h.kiwix.runes)
if aerr != nil || page.Text == "" {
if aerr != nil {
log.Printf("voice: kiwix: article %s: %v", top.Path, aerr)
}
// The search did find something, so fall back to its snippet rather
// than throw the hit away.
if top.Snippet == "" {
return "", false
}
page = crawl.Page{Title: top.Title, Text: top.Snippet}
}
// Handed over the same way a note or a page is: context for the question he
// asked, not something to recite.
snippet := top.Title + "\n" + crawl.TrimRunes(page.Text, h.kiwix.runes)
reply := h.phraseSource(ctx, "kiwix", t.dec.Utterance, []string{snippet})
if reply == "" {
// No phraser, or it failed. Read back the best hit rather than pretend
// the search did not happen.
return "вот что я нашла: " + crawl.TrimRunes(top.Title+" — "+page.Text, 300), true
}
return reply, true
}
// queryPersonal — stop the walk on a question about him that his own data did
// not answer.
//
// Every source above this one reads something of his: his facts, his calendar,
// his tasks, his house, his notes. Everything below reads the world: an offline
// Wikipedia, a page he named, the model's own weights. The world does not know
// when his meeting is, and asked anyway it will produce something.
//
// It did. "во сколько у меня встреча" reached Kiwix on the deployed daemon,
// 01-08-2026; Wikipedia matched an article on the 2015 CPISRA World Games, and
// the phraser rendered it as "встреча у тебя в 2015 CPISRA World Games, где
// были соревнования по плаванию". Fluent, confident, and about a swimming
// competition in Nottingham. Saying "не знаю" is not a worse answer than that
// one — it is the only true one.
//
// Note this is also the privacy edge. The rule in CLAUDE.md is that only the
// utterance may leave the box, never his notes; a question that is ABOUT him
// carries his life in the utterance itself, so it is the one class that should
// not be sent to an upstream engine at all. The guard closes both holes with
// the same test.
func (h *reactiveHandler) queryPersonal(ctx context.Context, t *queryTurn) (string, bool) {
if !h.isPersonalTurn(ctx, t) {
return "", false
}
log.Printf("voice: %q is about him and his own data did not answer it; not asking the world", t.dec.Utterance)
return "не знаю — не нашла у тебя такой записи.", true
}
// personalMarkers — first-person POSSESSION, not first person generally.
//
// "у меня" and "мой" attach to a thing that is his, which is what makes the
// question unanswerable from outside. A bare "мне" or "я" does not: "как мне
// сварить борщ" and "что я могу посмотреть" are ordinary questions about the
// world that happen to mention the asker, and refusing those would be the
// opposite mistake. The narrow test is the point.
// Go's \b is ASCII-only and never fires next to a Cyrillic letter, so the
// Russian patterns spell the boundary out as "not a letter or a digit". The
// English ones keep \b, where it works.
var personalMarkers = []*regexp.Regexp{
regexp.MustCompile(`(?i)(^|[^\p{L}\p{N}])у\s+меня([^\p{L}\p{N}]|$)`),
regexp.MustCompile(`(?i)(^|[^\p{L}\p{N}])мо(й|я|ё|е|и|его|ей|их|им|ими|ем|ю|ею)([^\p{L}\p{N}]|$)`),
regexp.MustCompile(`(?i)\bmy\b`),
regexp.MustCompile(`(?i)\bdo\s+i\s+have\b`),
regexp.MustCompile(`(?i)\bdid\s+i\b`),
}
// isPersonalQuery — the offline floor under the boundary. Possession only, and
// deliberately still narrow: it answers when there is no embedder to ask, and a
// broad guess made blind is worse than a narrow one.
func isPersonalQuery(utterance string) bool {
if utterance == "" {
return false
}
for _, re := range personalMarkers {
if re.MatchString(utterance) {
return true
}
}
return false
}
// isPersonalTurn — the boundary test. The seeds decide when the embedder is
// there, which is every deployed box; the possession markers are the floor
// underneath, for a handler with no embedder or a turn whose vector never got
// computed. Same shape as the cascade: the better test leads, the offline one
// always answers.
func (h *reactiveHandler) isPersonalTurn(ctx context.Context, t *queryTurn) bool {
h.boundary.load(ctx, h.embedder)
if personal, world, ok := h.boundary.score(t.vec); ok {
if personal > world {
log.Printf("voice: %q scores personal %.4f vs world %.4f", t.dec.Utterance, personal, world)
return true
}
return false
}
return isPersonalQuery(t.dec.Utterance)
}
// queryGeneral — general knowledge, the last source before giving up. It always
// claims: either a model answers, or Maven names the gap, or she says she does
// not know.
//
// This is the sharpest case for the naming half. Nothing has been fetched, so
// there is no passage to fall back on and no floor under the answer except the
// model's weights — and a 1.7B's weights are where the invented answers come
// from. With a workstation configured and asleep he is told that, rather than
// told something false in a confident voice. With no workstation configured at
// all the resident model answers exactly as it does today: naming a gap requires
// a gap, and on that box the 1.7B is the whole product.
func (h *reactiveHandler) queryGeneral(ctx context.Context, t *queryTurn) (string, bool) {
reply, err := h.phraser.PhraseQuery(ctx, t.dec.Utterance, nil)
if h.phraser == nil {
// No model of any size. That is not the workstation being asleep, so it
// is not that gap: it is simply not knowing.
return "не знаю.", true
}
reply, err := h.phraseWorld(ctx, t.dec.Utterance, nil)
if errors.Is(err, phraser.ErrNoWorldModel) {
log.Printf("voice: %q needs the world model and it is not available", t.dec.Utterance)
return worldGap, true
}
if err != nil || reply == "" {
return "не знаю.", true
}
+116
View File
@@ -0,0 +1,116 @@
package main
import (
"context"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
)
// contQueryAPI records which core call a continued query reached. DayPlan and
// LatestFact are here to be caught, not to be used: a continuation must never
// reach them, and the counters are how the test says so.
type contQueryAPI struct {
ipc.UnimplementedCoreAPI
from, to time.Time
events int
plans int
factLooks int
}
func (a *contQueryAPI) CalendarEvents(_ context.Context, from, to time.Time) ([]ipc.Fact, error) {
a.events++
a.from, a.to = from, to
return []ipc.Fact{{Key: "calendar", Value: "Планёрка @ 14:00", Confidence: 1.0, Ts: from.Add(14 * time.Hour)}}, nil
}
func (a *contQueryAPI) DayPlan(context.Context) (ipc.DayPlan, error) {
a.plans++
return ipc.DayPlan{Spoken: "план на сегодня"}, nil
}
func (a *contQueryAPI) LatestFact(_ context.Context, key string) (ipc.Fact, error) {
a.factLooks++
return ipc.Fact{Key: key, Value: "2л", Ts: contNow.Add(-time.Hour)}, nil
}
func contQueryHandler() (*reactiveHandler, *contQueryAPI) {
api := &contQueryAPI{}
return &reactiveHandler{api: api, now: func() time.Time { return contNow }}, api
}
// A continuation is a question about another day, so the one source that can
// read a day answers it — for the day the ellipsis named, not for today.
func TestContinuedQueryReachesTheCalendar(t *testing.T) {
h, api := contQueryHandler()
reply := h.actionQuery(context.Background(), router.Decision{
Intent: router.IntentQuery,
Utterance: "а завтра?",
Continued: true,
Slots: router.Slots{Text: "что у меня сегодня", Time: contNow.Add(24 * time.Hour), HasTime: true},
})
if api.events != 1 {
t.Fatalf("CalendarEvents called %d times, want 1", api.events)
}
if got, want := api.from.Format("2006-01-02"), "2026-08-02"; got != want {
t.Errorf("asked the calendar for %s, want %s", got, want)
}
if reply == "" {
t.Error("empty reply")
}
}
// The regression this gate exists for: every other source is date-blind, so
// letting one claim a continuation answers a question about tomorrow with
// today's data. queryFactByKey was the live case — HasKey plus HasTime, both
// set by the continuation, and it replies with a stored fact's own timestamp.
func TestContinuedQuerySkipsDateBlindSources(t *testing.T) {
h, api := contQueryHandler()
h.actionQuery(context.Background(), router.Decision{
Intent: router.IntentQuery,
Utterance: "а вчера?",
Continued: true,
Slots: router.Slots{
Key: "water", HasKey: true,
Text: "когда я пил воду",
Time: contNow.Add(-24 * time.Hour), HasTime: true,
},
})
if api.factLooks != 0 {
t.Errorf("fact-by-key claimed a continuation (%d lookups)", api.factLooks)
}
if api.plans != 0 {
t.Errorf("day-plan claimed a continuation (%d calls)", api.plans)
}
}
// Nothing date-aware claimed it: say that, rather than "не знаю", which reads
// as "no data for that day" when she never looked.
func TestContinuedQueryWithNoDateAwareAnswerSaysSo(t *testing.T) {
h, _ := contQueryHandler()
// No parseable day in the utterance, so even the calendar passes.
reply := h.actionQuery(context.Background(), router.Decision{
Intent: router.IntentQuery,
Utterance: "а?",
Continued: true,
Slots: router.Slots{Text: "какая погода", HasTime: true},
})
if reply == "не знаю." || !strings.Contains(reply, "спроси целиком") {
t.Fatalf("reply = %q, want the honest continuation refusal", reply)
}
}
// An ordinary query is untouched by the gate — every source still runs.
func TestOrdinaryQueryStillReachesEverySource(t *testing.T) {
h, api := contQueryHandler()
h.actionQuery(context.Background(), router.Decision{
Intent: router.IntentQuery,
Utterance: "какие планы на сегодня?",
})
if api.plans != 1 {
t.Fatalf("day-plan called %d times on an ordinary query, want 1", api.plans)
}
}
+117
View File
@@ -0,0 +1,117 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
)
func TestIsPersonalQuery(t *testing.T) {
for _, s := range []string{
"во сколько у меня встреча",
"что у меня сегодня",
"когда мой следующий отпуск",
"где моя книга",
"сколько моих задач висит",
"when is my meeting",
"do i have anything today",
"did i take my vitamins",
// Speech, but only the forms possession already covers ("did i").
// The verb forms the floor cannot see are the seeds' job, scored in
// TestONNXPersonalBoundary.
"what did i say about backups",
} {
if !isPersonalQuery(s) {
t.Errorf("isPersonalQuery(%q) = false, want true", s)
}
}
for _, s := range []string{
// First person without possession. These are questions about the
// world that merely mention the asker, and refusing them would be the
// opposite mistake.
"как мне сварить борщ",
"что я могу посмотреть вечером",
"почему небо синее",
"столица франции",
"how do i boil an egg",
// The floor is possession-only by design: a speech verb it cannot see
// passes here and is caught by the seeds instead.
"что я говорил про бэкапы?",
"как я говорил, почему небо синее",
"",
} {
if isPersonalQuery(s) {
t.Errorf("isPersonalQuery(%q) = true, want false", s)
}
}
}
// kiwixTrapAPI stands in for the world. Nothing below the personal boundary
// should be consulted for a question about him, so the test asserts on the
// reply rather than on a call: reaching Kiwix or general knowledge produces a
// phrased answer, and refusing produces the honest one.
func personalHandler() *reactiveHandler {
return &reactiveHandler{
api: ipc.UnimplementedCoreAPI{},
now: func() time.Time { return contNow },
// No phraser and no kiwix wiring: if the walk gets past the personal
// source it reaches queryGeneral, which returns "не знаю." with a nil
// phraser — a different string from the one this guard produces, so
// the two cases stay distinguishable.
}
}
// The regression: "во сколько у меня встреча" reached Kiwix, Wikipedia matched
// an article on the 2015 CPISRA World Games, and the phraser reported it back
// as his meeting. Seen on the deployed daemon, 01-08-2026.
func TestPersonalQuestionIsNotSentToTheWorld(t *testing.T) {
h := personalHandler()
reply, ok := h.queryPersonal(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "во сколько у меня встреча"},
})
if !ok {
t.Fatal("queryPersonal passed on a question about him")
}
if reply == "" {
t.Fatal("empty reply")
}
}
func TestWorldQuestionsPassThroughTheBoundary(t *testing.T) {
h := personalHandler()
for _, u := range []string{"почему небо синее", "столица франции"} {
if _, ok := h.queryPersonal(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: u},
}); ok {
t.Errorf("queryPersonal claimed %q, want it to pass to the encyclopedia", u)
}
}
}
// The boundary must sit above kiwix and general-knowledge and below every
// source that reads his own data. Asserted on the table itself: an ordering
// bug here is silent, because both arrangements answer, just from the wrong
// place.
func TestPersonalBoundarySitsBetweenHisDataAndTheWorld(t *testing.T) {
idx := map[string]int{}
for i, s := range querySources {
idx[s.name] = i
}
boundary, ok := idx["personal"]
if !ok {
t.Fatal("no personal source in the chain")
}
for _, his := range []string{"fact-by-key", "day-plan", "tasks", "calendar", "memory", "notes"} {
if i, ok := idx[his]; !ok || i > boundary {
t.Errorf("%q reads his own data and must run before the personal boundary", his)
}
}
for _, world := range []string{"search", "kiwix", "web", "general-knowledge"} {
if i, ok := idx[world]; !ok || i < boundary {
t.Errorf("%q reads the world and must run after the personal boundary", world)
}
}
}
+105
View File
@@ -0,0 +1,105 @@
package main
import (
"context"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"testing"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/websearch"
)
func searchHandler(t *testing.T, body string, status int) (*reactiveHandler, *string) {
t.Helper()
var seen string
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
seen = r.URL.RawQuery
if status != http.StatusOK {
http.Error(w, "no", status)
return
}
w.Write([]byte(body))
}))
t.Cleanup(srv.Close)
return &reactiveHandler{
// No phraser: querySearch then reads back the best evidence, which is
// what makes the claim visible without a llama-server in the test.
search: &searchWiring{client: websearch.New(srv.URL, websearch.Options{}), max: 3, runes: 1500},
}, &seen
}
const searchBody = `{"answers":["Небо голубое из-за рэлеевского рассеяния."],
"results":[{"title":"Рэлеевское рассеяние","url":"https://ru.wikipedia.org/x","content":"Рассеяние света."}]}`
func TestQuerySearchClaimsAndReadsBack(t *testing.T) {
h, _ := searchHandler(t, searchBody, http.StatusOK)
reply, ok := h.querySearch(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "почему небо голубое"},
})
if !ok {
t.Fatal("querySearch passed on a search with hits")
}
if !strings.Contains(reply, "рэлеевского рассеяния") {
t.Fatalf("reply = %q", reply)
}
}
// No rewriter in front of this source: SearXNG ranks by meaning, and reducing
// the question to English keywords would throw away the language he asked in.
func TestQuerySearchSendsTheQuestionVerbatim(t *testing.T) {
h, seen := searchHandler(t, searchBody, http.StatusOK)
h.querySearch(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "почему небо голубое"},
})
if !strings.Contains(*seen, "q="+url.QueryEscape("почему небо голубое")) {
t.Fatalf("query string = %q", *seen)
}
}
// The whole reason the ordering is safe: an unreachable or empty instance
// passes the turn to Kiwix instead of claiming it with an apology.
func TestQuerySearchFallsThroughWhenItFails(t *testing.T) {
for _, tc := range []struct {
name string
body string
status int
}{
{"http error", "", http.StatusForbidden},
{"no hits", `{"answers":[],"results":[]}`, http.StatusOK},
} {
t.Run(tc.name, func(t *testing.T) {
h, _ := searchHandler(t, tc.body, tc.status)
if _, ok := h.querySearch(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "почему небо голубое"},
}); ok {
t.Fatal("querySearch claimed the turn; Kiwix never got its fallback")
}
})
}
}
// Off unless configured, and silent about it: he never asked for a capability
// he did not enable.
func TestQuerySearchOffWithoutConfig(t *testing.T) {
h := &reactiveHandler{}
if _, ok := h.querySearch(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "почему небо голубое"},
}); ok {
t.Fatal("querySearch claimed a turn with no search block")
}
}
// The owner's ruling of 2026-08-02: the live search asks first, the ZIM is the
// fallback for a box with no line out.
func TestSearchRunsBeforeKiwix(t *testing.T) {
idx := map[string]int{}
for i, s := range querySources {
idx[s.name] = i
}
if idx["search"] > idx["kiwix"] {
t.Fatalf("search at %d, kiwix at %d: the ZIM is the fallback, not the first read", idx["search"], idx["kiwix"])
}
}
+52 -1
View File
@@ -31,9 +31,11 @@ package main
import (
"context"
"encoding/json"
"errors"
"fmt"
"log"
"strings"
"sync"
"time"
@@ -54,16 +56,65 @@ import (
// asks Maven to stop recording gets the transcript back in seconds.
const captureSummaryTimeout = 20 * time.Minute
// summaryGrammar — GBNF pinning a summarisation call to one JSON object holding
// the summary and nothing else. Same reasoning as responseGrammar and memeval's
// evalGrammar: the resident model is a Thinking variant, and a summarisation
// prompt is exactly the shape that invites it to answer with its reasoning as
// plain text. Demanding JSON leaves the reasoning nowhere to go.
//
// The bound is 2000 characters, twice the phraser's, because a reduce step over
// a two-hour meeting is a paragraph and not a sentence. Newlines are escaped by
// the escape rule, so the bullet list the prompt asks for survives the wrapper.
const summaryGrammar = `
root ::= "{" ws "\"summary\"" ws ":" ws string ws "}"
string ::= "\"" ([^"\\] | "\\" ["\\/bfnrt]){0,2000} "\""
ws ::= [ \t\n]*
`
// llmCompleter adapts *llm.Client to capture.Completer. The pure package names
// the two strings it needs and stays free of the llm request struct; the client
// itself is the swap-aware one from llmClientFor, so a model swap re-points it.
//
// The JSON wrapper lives here, not in internal/capture: that package is
// text-in/text-out by design, and the map/reduce steps still see plain prose.
type llmCompleter struct {
c *llm.Client
maxTokens int
}
func (l llmCompleter) Complete(ctx context.Context, system, user string) (string, error) {
return l.c.Complete(ctx, llm.Req{System: system, User: user, MaxTokens: l.maxTokens})
out, err := l.c.Complete(ctx, llm.Req{
System: system,
User: user,
Grammar: summaryGrammar,
MaxTokens: l.maxTokens,
})
if err != nil {
return "", err
}
return unwrapSummary(out), nil
}
// unwrapSummary takes the summary out of the JSON object the grammar produced.
// Anything that does not parse is returned as-is: an operator running without a
// grammar, or a llama-server too old to honour one, gets the plain text it used
// to get rather than an empty meeting summary.
func unwrapSummary(raw string) string {
s := phraser.StripThink(strings.TrimSpace(raw))
start := strings.Index(s, "{")
end := strings.LastIndex(s, "}")
if start < 0 || end <= start {
return s
}
var parsed struct {
Summary string `json:"summary"`
}
if err := json.Unmarshal([]byte(s[start:end+1]), &parsed); err != nil {
return s
}
// An empty field is the model saying nothing, so hand back nothing. Returning
// the raw object here would write `{"summary":""}` into his notes.
return strings.TrimSpace(parsed.Summary)
}
// captureWiring — the recorder plus what it needs to write the result down.
+26
View File
@@ -116,3 +116,29 @@ func TestStopReturnsTranscriptAndNotesItWithoutASummary(t *testing.T) {
t.Fatalf("the meeting left no note behind: %+v", notes)
}
}
// The summary path is JSON-wrapped by summaryGrammar, and internal/capture must
// keep seeing plain prose. These cover the wrapper and every way it can be
// absent or broken, because a meeting summary is written once and not retried.
func TestUnwrapSummary(t *testing.T) {
cases := []struct {
name string
in string
want string
}{
{"grammar output", `{"summary": "решили купить насос"}`, "решили купить насос"},
{"multiline field", `{"summary": "- насос\n- бюджет"}`, "- насос\n- бюджет"},
{"empty marker survives", `{"summary": "пусто"}`, "пусто"},
{"empty field says nothing", `{"summary": ""}`, ""},
{"thinking prefix", "<think>hm</think>\n{\"summary\": \"итог\"}", "итог"},
{"no grammar, plain prose", "решили купить насос", "решили купить насос"},
{"broken json falls back", `{"summary": "обрыв`, `{"summary": "обрыв`},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
if got := unwrapSummary(c.in); got != c.want {
t.Errorf("unwrapSummary(%q) = %q, want %q", c.in, got, c.want)
}
})
}
}
+32 -14
View File
@@ -106,11 +106,11 @@ func trimClarifyExpired(s string) string {
// out, and "" when nothing was parked. Call it right after
// resolveClarifyAnswer: a live question is answered there, an expired one is
// only reported here — the words themselves still go on to be routed fresh.
func (h *reactiveHandler) clarifyExpiredNotice() string {
func (h *reactiveHandler) clarifyExpiredNotice(ctx context.Context) string {
if h.clarifyStore == nil {
return ""
}
if !h.clarifyStore.TakeExpired(voiceDialogueID, h.now()) {
if !h.clarifyStore.TakeExpired(dialogueIDOf(ctx), h.now()) {
return ""
}
log.Printf("voice: clarify — parked question expired, telling him and routing the words fresh")
@@ -157,7 +157,7 @@ func clarifyQuestion(dec router.Decision) (dialogue.Slot, string, bool) {
// askClarify parks the request and returns the question to ask instead of the
// canned "не поняла". Returns ("", false) when there is nothing to ask about, so
// the caller falls back to the canned reply.
func (h *reactiveHandler) askClarify(dec router.Decision) (string, bool) {
func (h *reactiveHandler) askClarify(ctx context.Context, dec router.Decision) (string, bool) {
if h.clarifyStore == nil {
return "", false
}
@@ -165,7 +165,7 @@ func (h *reactiveHandler) askClarify(dec router.Decision) (string, bool) {
if !ok {
return "", false
}
h.clarifyStore.Put(voiceDialogueID, &dialogue.PendingQuestion{
h.clarifyStore.Put(dialogueIDOf(ctx), &dialogue.PendingQuestion{
Intent: dialogue.Intent(dec.Intent),
Slots: toDialogueSlots(dec.Slots),
Missing: []dialogue.Slot{slot},
@@ -192,7 +192,7 @@ func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string)
if h.clarifyStore == nil {
return "", false
}
q := h.clarifyStore.Get(voiceDialogueID, h.now())
q := h.clarifyStore.Get(dialogueIDOf(ctx), h.now())
if q == nil {
return "", false
}
@@ -206,9 +206,9 @@ func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string)
// would fire at 11:00 saying "напомни" and nothing else.
q.Utterance = foldAnswerIntoUtterance(q.Utterance, merged.Text)
if len(dialogue.StillMissing(q.Missing, merged)) > 0 {
return h.reaskOrGiveUp(q, merged, text), true
return h.reaskOrGiveUp(ctx, q, merged, text), true
}
h.clarifyStore.Delete(voiceDialogueID)
h.clarifyStore.Delete(dialogueIDOf(ctx))
// One gap filled is not the same as a complete request. askClarify parks
// only the first gap, because one question per turn is the rule, but a
@@ -217,7 +217,7 @@ func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string)
// a reminder with no time, which answered "не получилось разобрать время
// напоминания." — an error for a request she never finished asking about.
// Re-enter the loop instead, one question at a time as before.
if reply, asked := h.askRemainingGap(q, intent, merged); asked {
if reply, asked := h.askRemainingGap(ctx, q, intent, merged); asked {
return reply, true
}
@@ -261,7 +261,7 @@ func foldAnswerIntoUtterance(utterance, subject string) string {
// The attempt budget is shared with the re-ask path on purpose. A second gap
// costs a question exactly like a second try at the first one does, so the cap
// still bounds how many times she can speak before acting or letting go.
func (h *reactiveHandler) askRemainingGap(q *dialogue.PendingQuestion, intent router.Intent, merged dialogue.Slots) (string, bool) {
func (h *reactiveHandler) askRemainingGap(ctx context.Context, q *dialogue.PendingQuestion, intent router.Intent, merged dialogue.Slots) (string, bool) {
remaining := dialogue.StillMissing(wantedSlots[intent], merged)
if len(remaining) == 0 {
return "", false
@@ -270,7 +270,7 @@ func (h *reactiveHandler) askRemainingGap(q *dialogue.PendingQuestion, intent ro
if !ok || !q.CanAsk() {
return "", false
}
h.clarifyStore.Put(voiceDialogueID, &dialogue.PendingQuestion{
h.clarifyStore.Put(dialogueIDOf(ctx), &dialogue.PendingQuestion{
Intent: q.Intent,
Slots: merged,
Missing: []dialogue.Slot{remaining[0]},
@@ -287,13 +287,13 @@ func (h *reactiveHandler) askRemainingGap(q *dialogue.PendingQuestion, intent ro
// reaskOrGiveUp handles an answer that left the gap open: ask the same question
// again while she has attempts left, otherwise say she did not understand and
// let the request go. Never returns "" — a mute give-up reads as "done".
func (h *reactiveHandler) reaskOrGiveUp(q *dialogue.PendingQuestion, merged dialogue.Slots, text string) string {
func (h *reactiveHandler) reaskOrGiveUp(ctx context.Context, q *dialogue.PendingQuestion, merged dialogue.Slots, text string) string {
question := ""
if len(q.Missing) > 0 {
question = clarifyQuestions[q.Missing[0]]
}
if question == "" || !q.CanAsk() {
h.clarifyStore.Delete(voiceDialogueID)
h.clarifyStore.Delete(dialogueIDOf(ctx))
log.Printf("voice: clarify — gave up on %v after %d question(s), answer was %q", q.Missing, q.Attempts, text)
return clarifyGaveUp
}
@@ -302,7 +302,7 @@ func (h *reactiveHandler) reaskOrGiveUp(q *dialogue.PendingQuestion, merged dial
q.Slots = merged
q.Attempts++
q.Asked = h.now()
h.clarifyStore.Put(voiceDialogueID, q)
h.clarifyStore.Put(dialogueIDOf(ctx), q)
log.Printf("voice: clarify — answer %q did not fill %v, asking again (attempt %d)", text, q.Missing, q.Attempts)
return question
}
@@ -348,9 +348,27 @@ func (h *reactiveHandler) rememberTurn(prev *dialogue.Session, dec router.Decisi
if dec.Intent == router.IntentChat {
ttl = 15 * time.Minute // conversational turns should last longer
}
// A system or query turn often carries no Text slot at all — a stage-0
// grammar fills none. The next turn may be an ellipsis ("а завтра?"),
// which knows the day but not what was asked ABOUT, so keep the raw
// utterance where continuation.go can find it. Only these two intents:
// everywhere else Text is a payload and must stay what the router put in.
//
// Overwritten, not filled: rememberTurn runs AFTER followUpMerge, which
// has already inherited a Text from the previous same-intent turn, so a
// fill-if-empty rule keeps the OLD topic for ever. Seen on the deployed
// daemon 01-08-2026 — "во сколько у меня встреча" then "какие у меня
// планы" then "а завтра?" continued the meeting, two turns stale.
//
// A continuation is the exception and keeps what it inherited: its
// utterance is the ellipsis, and the topic it carries is the real one.
slots := toDialogueSlots(dec.Slots)
if !dec.Continued && (dec.Intent == router.IntentSystem || dec.Intent == router.IntentQuery) {
slots.Text = dec.Utterance
}
h.dialogueSessions.Put(voiceDialogueID, &dialogue.Session{
Intent: dialogue.Intent(dec.Intent),
Slots: toDialogueSlots(dec.Slots),
Slots: slots,
Timestamp: now,
TTL: ttl,
History: history,
+105 -20
View File
@@ -81,7 +81,7 @@ func TestClarifyReminderCompletesOnAnswer(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
question, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"))
question, asked := h.askClarify(ctx, clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"))
if !asked || question != "Когда?" {
t.Fatalf("expected the time question, got %q asked=%v", question, asked)
}
@@ -112,7 +112,7 @@ func TestClarifyFactCompletesOnAnswer(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
if _, asked := h.askClarify(clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши")); !asked {
if _, asked := h.askClarify(ctx, clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши")); !asked {
t.Fatal("a fact with no key should be asked about")
}
if reply, handled := h.resolveClarifyAnswer(ctx, "пил воду"); !handled || reply == clarifyGaveUp {
@@ -128,7 +128,7 @@ func TestClarifyAnswerAfterTTLIsANewRequest(t *testing.T) {
ctx := context.Background()
h, st, now := newClarifyHandler(t)
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
if _, asked := h.askClarify(ctx, clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a question")
}
*now = now.Add(clarifyTTL + time.Second)
@@ -147,7 +147,7 @@ func TestClarifyAsksThreeTimesThenSaysSo(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
if _, asked := h.askClarify(ctx, clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a first question")
}
// Two more unclear answers ⇒ two more questions (3 asks in total).
@@ -185,7 +185,7 @@ func TestClarifyMaxAttemptsIsConfigurable(t *testing.T) {
h, _, _ := newClarifyHandler(t)
h.clarifyMaxAttempts = 1
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
if _, asked := h.askClarify(ctx, clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a question")
}
if reply, handled := h.resolveClarifyAnswer(ctx, "ну не знаю"); !handled || reply != clarifyGaveUp {
@@ -199,7 +199,7 @@ func TestClarifyRestatedAnswerWins(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме")); !asked {
if _, asked := h.askClarify(ctx, clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме")); !asked {
t.Fatal("expected a question")
}
// First answer parses, but re-park it by hand as if she had asked again:
@@ -232,7 +232,7 @@ func TestClarifiedActOffAllowlistIsStillRefused(t *testing.T) {
h, st, _ := newClarifyHandler(t)
marker := filepath.Join(t.TempDir(), "not-allowed-ran")
if _, asked := h.askClarify(clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это")); !asked {
if _, asked := h.askClarify(ctx, clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это")); !asked {
t.Fatal("an act with no fn should be asked about")
}
reply, handled := h.resolveClarifyAnswer(ctx, "rm "+marker)
@@ -260,7 +260,7 @@ func TestClarifiedDestructiveActStillNeedsConfirm(t *testing.T) {
t.Fatal(err)
}
if _, asked := h.askClarify(clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это")); !asked {
if _, asked := h.askClarify(ctx, clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это")); !asked {
t.Fatal("expected a question")
}
reply, handled := h.resolveClarifyAnswer(ctx, "delete_backups")
@@ -284,7 +284,7 @@ func TestNoQuestionWhenNothingIsMissing(t *testing.T) {
clarifyDec(router.IntentQuery, router.Slots{Text: "ммм"}, "ммм"),
clarifyDec(router.IntentNote, router.Slots{Text: "..."}, "..."),
} {
if question, asked := h.askClarify(dec); asked {
if question, asked := h.askClarify(context.Background(), dec); asked {
t.Fatalf("intent %s should keep the canned reply, got %q", dec.Intent, question)
}
}
@@ -296,29 +296,29 @@ func TestNoQuestionWhenNothingIsMissing(t *testing.T) {
// TestClarifyExpiryIsAnnouncedAndWordsStillRoute — his answer lands after the
// TTL: she must say the old request is gone AND still answer the new words.
func TestClarifyExpiryIsAnnouncedAndWordsStillRoute(t *testing.T) {
ctx := context.Background()
ctx := withDialogueID(context.Background(), dialogueIDFor(sourceText, ""))
h, _, now := newClarifyHandler(t)
emb := router.NewHashEmbedder(1024)
h.embedder = emb
h.router = buildRouter(emb, h.matcher, 0.55, nil)
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
if _, asked := h.askClarify(ctx, clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a question")
}
*now = now.Add(clarifyTTL + time.Second)
reply := h.handleText(ctx, "как дела")
reply := h.handleText(ctx, "", "как дела")
if !isClarifyExpired(reply) {
t.Fatalf("expired question must be announced first, got %q", reply)
}
if trimClarifyExpired(reply) == "" {
t.Fatalf("the new words must still be answered, got only the notice: %q", reply)
}
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
if h.clarifyStore.Get(textDialogueID, h.now()) != nil {
t.Fatal("the expired question must be gone")
}
// The notice is said once, not on every later utterance.
if reply := h.handleText(ctx, "как дела"); isClarifyExpired(reply) {
if reply := h.handleText(ctx, "", "как дела"); isClarifyExpired(reply) {
t.Fatalf("notice repeated on a later turn: %q", reply)
}
}
@@ -340,7 +340,7 @@ func TestClarifyAsksAboutTheSecondGapToo(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
question, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{}, "напомни"))
question, asked := h.askClarify(ctx, clarifyDec(router.IntentReminder, router.Slots{}, "напомни"))
if !asked || question != "О чём напомнить?" {
t.Fatalf("expected the subject question, got %q asked=%v", question, asked)
}
@@ -380,7 +380,7 @@ func TestClarifySecondGapRespectsTheAttemptCap(t *testing.T) {
h, _, _ := newClarifyHandler(t)
h.clarifyMaxAttempts = 1
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{}, "напомни")); !asked {
if _, asked := h.askClarify(ctx, clarifyDec(router.IntentReminder, router.Slots{}, "напомни")); !asked {
t.Fatal("expected the subject question")
}
reply, handled := h.resolveClarifyAnswer(ctx, "позвонить маме")
@@ -440,10 +440,10 @@ func TestClarifyProseHoldsThePersona(t *testing.T) {
// The confirm turn used to return before the notice was even computed, so he
// answered the confirm and never heard that the older request was let go.
func TestExpiryNoticeSurvivesAConfirmTurn(t *testing.T) {
ctx := context.Background()
ctx := withDialogueID(context.Background(), dialogueIDFor(sourceText, ""))
h, _, now := newClarifyHandler(t)
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
if _, asked := h.askClarify(ctx, clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a question")
}
// A confirm parked with a longer life than the question, so only the
@@ -451,7 +451,7 @@ func TestExpiryNoticeSurvivesAConfirmTurn(t *testing.T) {
h.pending = &pendingAct{fn: "delete_backups", phrase: "удалить бэкапы", expiry: now.Add(time.Hour)}
*now = now.Add(clarifyTTL + time.Second)
reply := h.handleText(ctx, "нет")
reply := h.handleText(ctx, "", "нет")
if !isClarifyExpired(reply) {
t.Fatalf("the expired question must be announced on a confirm turn too, got %q", reply)
}
@@ -461,7 +461,92 @@ func TestExpiryNoticeSurvivesAConfirmTurn(t *testing.T) {
if h.pending != nil {
t.Fatal("the confirm must still have been consumed")
}
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
if h.clarifyStore.Get(textDialogueID, h.now()) != nil {
t.Fatal("the expired question must be gone")
}
}
// The other half of the subject question: his answer must fill the empty slot,
// not replace the request. Slots.Text used to be the whole raw utterance for
// every intent, so the branch that fills a text slot could only ever overwrite
// (Vikunja #383). Here the parked request holds the hour and the answer holds
// what to say at it, and the reminder that lands has both.
func TestClarifySubjectAnswerFillsRatherThanClobbers(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
at := h.now().Add(2 * time.Hour)
question, asked := h.askClarify(ctx, clarifyDec(router.IntentReminder,
router.Slots{Time: at, HasTime: true}, "напомни в 11"))
if !asked || question != "О чём напомнить?" {
t.Fatalf("expected the subject question, got %q asked=%v", question, asked)
}
reply, handled := h.resolveClarifyAnswer(ctx, "позвонить маме")
if !handled {
t.Fatal("the answer to an open question must be consumed as an answer")
}
if reply == clarifyGaveUp {
t.Fatalf("a good answer must not drop the request: %q", reply)
}
reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour))
if err != nil || len(reminders) != 1 {
t.Fatalf("clarified reminder was not created: reminders=%v err=%v", reminders, err)
}
if !strings.Contains(reminders[0].Payload, "маме") {
t.Fatalf("the answer never reached the reminder: %q", reminders[0].Payload)
}
if !strings.Contains(reminders[0].Payload, "11") {
t.Fatalf("the answer clobbered the original request: %q", reminders[0].Payload)
}
}
// TestClarifyIsPerConversation — the parked question belongs to the reach that
// was asked. Before this the clarify store had one global key, so a question
// asked in the web chat and never answered captured the next utterance from
// telegram, or from the mic, and answered it against a request the speaker had
// never made (Vikunja #466).
func TestClarifyIsPerConversation(t *testing.T) {
h, _, _ := newClarifyHandler(t)
web := withDialogueID(context.Background(), dialogueIDFor(sourceText, "web"))
telegram := withDialogueID(context.Background(), dialogueIDFor(sourceText, "telegram:42"))
if _, asked := h.askClarify(web, clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a question on the web conversation")
}
if _, handled := h.resolveClarifyAnswer(telegram, "в 11:00"); handled {
t.Fatal("a question asked on the web must not eat a telegram utterance")
}
if _, handled := h.resolveClarifyAnswer(voiceCtx(), "в 11:00"); handled {
t.Fatal("a question asked on the web must not eat what he says at the mic")
}
if reply, handled := h.resolveClarifyAnswer(web, "в 11:00"); !handled || reply == clarifyGaveUp {
t.Fatalf("the asker's own answer must land, handled=%v reply=%q", handled, reply)
}
}
// voiceCtx — the mic's conversation, which carries no id of its own.
func voiceCtx() context.Context {
return withDialogueID(context.Background(), dialogueIDFor(sourceVoice, ""))
}
// TestARestartExpiresTheParkedQuestion pins the Vikunja #385 decision: the
// question dies with the process, and she does not claim to have let it go —
// the words that follow are routed as a fresh request. Restarting is modelled
// the way the daemon does it, by building a second handler over the same store.
func TestARestartExpiresTheParkedQuestion(t *testing.T) {
h, _, _ := newClarifyHandler(t)
ctx := voiceCtx()
if _, asked := h.askClarify(ctx, clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a question before the restart")
}
restarted, _, _ := newClarifyHandler(t)
if _, handled := restarted.resolveClarifyAnswer(ctx, "в 11:00"); handled {
t.Fatal("a question parked before the restart must not eat the next utterance")
}
if notice := restarted.clarifyExpiredNotice(ctx); notice != "" {
t.Fatalf("notice = %q, want silence: nothing survived to expire", notice)
}
}
+12 -8
View File
@@ -107,14 +107,18 @@ func (h *reactiveHandler) confirmResolvers(ctx context.Context) []confirmResolve
return pr != nil && !h.now().After(pr.expiry)
},
yes: func() string {
// Only record the acceptance. The tick loop reads accepted
// routines and nudges on their own interval. Building a
// reminder here made a routine fire exactly once (Vikunja #366).
if err := h.dataStore.AcceptProposedRoutine(ctx, pr.routineID, h.now()); err != nil {
log.Printf("voice: accept proposed routine: %v", err)
return "не получилось запомнить рутину."
}
return "буду напоминать."
// Voice does NOT accept (Vikunja #367). Accepting hands the
// tick loop a standing new reason to speak, which is the same
// tier as enabling a tool — and DESIGN.md § "surface caps
// authority" says a room mic, reachable by anyone present, is
// structurally incapable of layer 3. So a spoken "да" leaves
// the row 'proposed' and points at the authed page, where the
// accept button is gated at step-up. The convenience of
// answering out loud stays; the authority does not move.
//
// Acceptance itself is recorded by /routines, and the tick
// loop nudges on the interval from there (Vikunja #366).
return "поняла — подтверди на странице рутин, и начну напоминать."
},
no: func() string {
if err := h.dataStore.DismissProposedRoutine(ctx, pr.routineID); err != nil {
+125
View File
@@ -0,0 +1,125 @@
// Elliptical follow-ups — "а завтра?" after "какие напоминания на сегодня".
//
// These carry no intent of their own. Two words, one of them a particle, and
// everything that makes the utterance meaningful lives in the turn before it.
// Sent to the router they get whatever the model guesses, which on a 1.7B is
// close to a coin flip, and the guess costs ~2.7s to obtain.
//
// followUpMerge (followup.go) cannot help: it inherits SLOTS once the intent is
// known, and here the intent is the missing part. So this runs before the
// router and answers from the previous turn directly, which is both correct by
// construction and free.
package main
import (
"time"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/router"
)
// continuationMaxTokens — an ellipsis is short by definition. Past four tokens
// the utterance carries enough of its own content to be routed on its merits,
// and inheriting an intent for it would be overreach.
const continuationMaxTokens = 4
// continuationParticles — the words that open a follow-up. A leading particle
// is one of the two ways in; the other is an utterance that is nothing but a
// date ("завтра?").
var continuationParticles = map[string]bool{
"а": true, "и": true, "ну": true,
"what": true, "and": true, "how": true,
}
// continuableIntents — which intents an ellipsis may inherit.
//
// query and system are questions: asking the same question about a different
// day is exactly what "а завтра?" means, and re-aiming the Time slot answers it
// completely.
//
// The rest are excluded on purpose. fact and note would write something he did
// not say — "поужинал" then "а вчера?" is a question about yesterday, not a
// claim about it. chat has no slot to re-aim. act is the dangerous one: an
// allowlisted fn inherited by a two-word utterance is a way to run a
// destructive command nobody typed, and no follow-up is worth that.
//
// reminder was in this list and came out after a live check on 01-08-2026. A
// reminder's payload is its Text, and the Text embeds the day word it was
// created with: continuing "напомни сегодня о событиях" with "а завтра?" fires
// tomorrow with the text still reading "сегодня". Re-aiming Time is not enough
// when the day is also written into the payload, and rewriting the payload
// needs the date's span in the string, which ParseCalendarDate does not report.
var continuableIntents = map[dialogue.Intent]bool{
dialogue.IntentQuery: true,
dialogue.IntentSystem: true,
}
// continuationDecision reads an utterance as "the previous question, but for
// this other day". Returns ok=false whenever anything is uncertain, which
// hands the turn back to the ordinary router path.
//
// The date is what makes this safe. An ellipsis with no parseable day is just
// a short utterance, and short utterances are the router's job.
func continuationDecision(prev *dialogue.Session, text string, now time.Time) (router.Decision, bool) {
if prev == nil || prev.IsExpired(now) || !continuableIntents[prev.Intent] {
return router.Decision{}, false
}
tokens := quietTokens(text)
if len(tokens) == 0 || len(tokens) > continuationMaxTokens {
return router.Decision{}, false
}
day, ok := router.ParseCalendarDate(text, now)
if !ok {
return router.Decision{}, false
}
// Either it opens with a particle, or the whole utterance is the date.
if !continuationParticles[tokens[0]] && !isBareDate(tokens, day, now) {
return router.Decision{}, false
}
dec := router.Decision{
Utterance: text,
Intent: router.Intent(prev.Intent),
Confidence: 1.0,
Stage: 0,
Continued: true,
Slots: router.Slots{
Key: prev.Slots.Key,
HasKey: prev.Slots.HasKey,
Value: prev.Slots.Value,
Text: prev.Slots.Text,
// Fn/Args are deliberately not carried: continuableIntents
// excludes act, so there is never one to carry.
Time: day,
HasTime: true,
},
}
return dec, true
}
// isBareDate reports whether the utterance is nothing but its date expression.
// "завтра" and "на выходных" qualify; "напомни завтра" does not, because the
// verb is content of its own and belongs to the router.
//
// Implemented by re-parsing each token: if every token that is not part of a
// date expression is a preposition or a question mark's leftovers, the
// utterance is bare. Cheap enough at four tokens.
func isBareDate(tokens []string, day time.Time, now time.Time) bool {
for _, t := range tokens {
if continuationFillers[t] {
continue
}
if d, ok := router.ParseCalendarDate(t, now); ok && d.Equal(day) {
continue
}
return false
}
return true
}
// continuationFillers — tokens that carry no content of their own inside a
// date expression ("на выходных", "в среду").
var continuationFillers = map[string]bool{
"на": true, "в": true, "во": true, "за": true, "про": true,
"about": true, "on": true, "for": true,
}
+169
View File
@@ -0,0 +1,169 @@
package main
import (
"testing"
"time"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/router"
)
var contNow = time.Date(2026, 8, 1, 12, 0, 0, 0, time.UTC)
func contSession(intent dialogue.Intent, key string) *dialogue.Session {
return &dialogue.Session{
Intent: intent,
Slots: dialogue.Slots{Key: key, HasKey: key != "", Text: "какие напоминания на сегодня"},
Timestamp: contNow.Add(-30 * time.Second),
TTL: 2 * time.Minute,
}
}
func TestContinuationInheritsTheQuestion(t *testing.T) {
prev := contSession(dialogue.IntentQuery, "water")
dec, ok := continuationDecision(prev, "а завтра?", contNow)
if !ok {
t.Fatal("continuationDecision returned false, want a decision")
}
if dec.Intent != router.IntentQuery {
t.Errorf("intent = %q, want query", dec.Intent)
}
if dec.Slots.Key != "water" || !dec.Slots.HasKey {
t.Errorf("key = %q, want water carried over", dec.Slots.Key)
}
if !dec.Slots.HasTime {
t.Fatal("no time slot; the whole point is re-aiming the day")
}
if got, want := dec.Slots.Time.Format("2006-01-02"), "2026-08-02"; got != want {
t.Errorf("time = %s, want %s", got, want)
}
}
func TestContinuationAcceptsABareDate(t *testing.T) {
prev := contSession(dialogue.IntentQuery, "water")
for _, s := range []string{"завтра?", "вчера", "а вчера?", "и завтра"} {
if _, ok := continuationDecision(prev, s, contNow); !ok {
t.Errorf("continuationDecision(%q) = false, want true", s)
}
}
}
func TestContinuationDeclinesWhatIsNotAnEllipsis(t *testing.T) {
prev := contSession(dialogue.IntentQuery, "water")
for _, s := range []string{
// No date to re-aim at — an ordinary short utterance, the router's job.
"а что там", "а бэкап?", "привет", "",
// Content of its own: the verb is not an ellipsis.
"напомни завтра позвонить маме",
// Too long to be an ellipsis even with a date in it.
"а что у меня стоит в календаре на завтра",
} {
if _, ok := continuationDecision(prev, s, contNow); ok {
t.Errorf("continuationDecision(%q) = true, want false", s)
}
}
}
func TestContinuationDeclinesUncontinuableIntents(t *testing.T) {
// act is the one that matters: inheriting an allowlisted fn from a
// two-word utterance would be a way to run a destructive command.
// reminder is here because its payload is its Text, and the Text embeds
// the day word it was created with — see continuableIntents.
for _, in := range []dialogue.Intent{
dialogue.IntentAct, dialogue.IntentFact, dialogue.IntentNote,
dialogue.IntentChat, dialogue.IntentReminder,
} {
if _, ok := continuationDecision(contSession(in, "water"), "а завтра?", contNow); ok {
t.Errorf("continuationDecision inherited intent %q, want refusal", in)
}
}
}
func TestContinuationDeclinesWithoutALiveSession(t *testing.T) {
if _, ok := continuationDecision(nil, "а завтра?", contNow); ok {
t.Error("continued with no previous turn")
}
stale := contSession(dialogue.IntentQuery, "water")
stale.Timestamp = contNow.Add(-10 * time.Minute)
if _, ok := continuationDecision(stale, "а завтра?", contNow); ok {
t.Error("continued an expired session")
}
}
func TestContinuationNeverCarriesAnFn(t *testing.T) {
prev := contSession(dialogue.IntentQuery, "water")
prev.Slots.Fn, prev.Slots.HasFn = "restart", true
dec, ok := continuationDecision(prev, "а завтра?", contNow)
if !ok {
t.Fatal("want a decision")
}
if dec.Slots.HasFn || dec.Slots.Fn != "" {
t.Fatalf("carried fn %q into a continuation", dec.Slots.Fn)
}
}
// TestContinuationCarriesTheTopic — the ellipsis names the day; what he is
// asking ABOUT has to come from the previous turn, or replySystem keyword-
// matches "а завтра?" and finds nothing. Caught on the deployed daemon.
func TestContinuationCarriesTheTopic(t *testing.T) {
prev := contSession(dialogue.IntentSystem, "")
prev.Slots.Text = "какой сегодня день"
dec, ok := continuationDecision(prev, "а завтра?", contNow)
if !ok {
t.Fatal("want a decision")
}
if dec.Slots.Text != "какой сегодня день" {
t.Fatalf("Slots.Text = %q, want the previous turn's topic", dec.Slots.Text)
}
}
// TestReplySystemIgnoresAnInheritedTopic — the regression the deployed daemon
// showed on 01-08-2026: followUpMerge fills an empty Text from the previous
// same-intent turn, so a plain "привет" after "какой сегодня день" arrived at
// replySystem carrying the old topic and was answered with the date. Only a
// continuation may widen the keyword match.
func TestReplySystemIgnoresAnInheritedTopic(t *testing.T) {
h := &reactiveHandler{now: func() time.Time { return contNow }}
inherited := router.Decision{
Utterance: "привет",
Intent: router.IntentSystem,
Slots: router.Slots{Text: "какой сегодня день"},
}
if got := h.replySystem(nil, inherited); got != "пока не умею отвечать на этот вопрос." {
t.Fatalf("replySystem answered %q on an inherited topic", got)
}
cont := inherited
cont.Utterance, cont.Continued = "а завтра?", true
if got := h.replySystem(nil, cont); got == "пока не умею отвечать на этот вопрос." {
t.Fatalf("replySystem refused a real continuation")
}
}
// TestRememberTurnRefreshesTheTopic — rememberTurn runs after followUpMerge,
// which has already inherited a Text from the previous same-intent turn. A
// fill-if-empty rule therefore pins the FIRST topic of a run of query turns
// and never lets go, so a later "а завтра?" continues a question two turns
// old. Seen on the deployed daemon, 01-08-2026.
func TestRememberTurnRefreshesTheTopic(t *testing.T) {
h := &reactiveHandler{
now: func() time.Time { return contNow },
dialogueSessions: dialogue.NewSessionStore(2 * time.Minute),
}
h.rememberTurn(nil, router.Decision{
Intent: router.IntentQuery, Utterance: "во сколько у меня встреча",
}, contNow)
// The second turn arrives with the first turn's Text already merged in.
prev := h.dialogueSessions.Get(voiceDialogueID, contNow)
h.rememberTurn(prev, router.Decision{
Intent: router.IntentQuery,
Utterance: "какие у меня планы",
Slots: router.Slots{Text: "во сколько у меня встреча"},
}, contNow)
got := h.dialogueSessions.Get(voiceDialogueID, contNow)
if got == nil {
t.Fatal("no session")
}
if got.Slots.Text != "какие у меня планы" {
t.Fatalf("topic = %q, want the latest turn's", got.Slots.Text)
}
}
+28
View File
@@ -351,3 +351,31 @@ func TestTickDayPlanReadsTheStore(t *testing.T) {
t.Errorf("a reminder for next year is not today's plan: %q", plan.Spoken)
}
}
// TestHandlerUpgradesToTheDaemonAPI — wireVoice runs before the tick loop
// exists, so the handler starts with the bare store adapter, and that adapter
// refuses DayPlan ("not available via direct store API"). main back-patches
// the real one in. Without the patch every "какие у меня планы на сегодня"
// answered "не получилось собрать план" on the deployed daemon, 01-08-2026.
func TestHandlerUpgradesToTheDaemonAPI(t *testing.T) {
h := &reactiveHandler{api: ipc.NewStoreAPI(nil), now: planDay}
if _, err := h.api.DayPlan(context.Background()); err == nil {
t.Fatal("the bare store adapter served a day plan; this test is measuring nothing")
}
want := samplePlan()
h.upgradeAPI(&daemonAPI{
CoreAPI: ipc.UnimplementedCoreAPI{},
getDayPlan: func(context.Context) ipc.DayPlan { return want },
})
reply, ok := h.queryDayPlan(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "какие у меня планы на сегодня?"},
})
if !ok {
t.Fatal("queryDayPlan passed on a plan question")
}
if reply != want.Spoken {
t.Fatalf("reply = %q, want the assembled plan", reply)
}
}
+45 -5
View File
@@ -514,11 +514,14 @@ func (h *reactiveHandler) handleHexisAct(ctx context.Context, dec router.Decisio
// Resolve the utterance text as an entity reference through Nexus. An
// ambiguous match must stop and clarify — never guess a mutation target.
// The name comes from entityReferenceText, not straight from the Text slot:
// the model transliterates Latin names as it routes (Vikunja #476).
subject := entityReferenceText(dec)
started := h.now()
entityID, displayName, ambiguous, err := h.ecosystem.resolveEntityReference(ctx, dec.Slots.Text, nil)
entityID, displayName, ambiguous, err := h.ecosystem.resolveEntityReference(ctx, subject, nil)
if err != nil {
h.recordEcosystemTrace(ctx, "nexus", "resolve", traceStatusForError(err), started,
mergeFields(traceErrorFields(err), map[string]any{"subject": redactSubject(dec.Slots.Text)}))
mergeFields(traceErrorFields(err), map[string]any{"subject": redactSubject(subject)}))
if unauthorizedEcosystemError(err) {
return "экосистема отклоняет доступ, проверь токен."
}
@@ -535,7 +538,7 @@ func (h *reactiveHandler) handleHexisAct(ctx context.Context, dec router.Decisio
}
if entityID == "" {
h.recordEcosystemTrace(ctx, "nexus", "resolve", traceNotFound, started,
map[string]any{"subject": redactSubject(dec.Slots.Text)})
map[string]any{"subject": redactSubject(subject)})
return ""
}
h.recordEcosystemTrace(ctx, "nexus", "resolve", traceOK, started,
@@ -569,10 +572,21 @@ func (h *reactiveHandler) handleHexisAct(ctx context.Context, dec router.Decisio
}
verbLower := strings.ToLower(verb)
// With no allowlisted fn the verb is a whole phrase ("restart status muzick
// indexer"), which no capability name ever contains. Read it the other way
// round then: the phrase is the haystack and the capability name is what we
// look for in it (Vikunja #476). Only when the fn slot is empty — a matched
// fn is a single verb and containment already means what it says.
loose := !dec.Slots.HasFn
var matches []*hexisclient.Capability
for i, c := range caps {
if strings.Contains(strings.ToLower(c.Name), verbLower) ||
(c.Description != "" && strings.Contains(strings.ToLower(c.Description), verbLower)) {
name := strings.ToLower(c.Name)
hit := strings.Contains(name, verbLower) ||
(c.Description != "" && strings.Contains(strings.ToLower(c.Description), verbLower))
if loose && name != "" && strings.Contains(verbLower, name) {
hit = true
}
if hit {
matches = append(matches, &caps[i])
}
}
@@ -632,3 +646,29 @@ func (h *reactiveHandler) execHexis(ctx context.Context, capID, capName, entityI
})
return "команда выполнена для " + displayName + "."
}
// hexisBeforeClarify gives an entity-shaped act one chance at Hexis before she
// asks what to do.
//
// The stage-3 gate thins an act that never matched an allowlisted fn, so
// "перезапусти muzick indexer" was answered with "Что сделать?" and the Hexis
// path was never entered — the capability existed and no utterance could reach
// it (Vikunja #476). Hexis is exactly where an act with no local fn belongs:
// the verb is matched against the capabilities Hexis registers for the entity,
// not against the allowlist.
//
// Narrow on purpose. Only an act, only when the fn slot is still empty, and
// only when Hexis is wired — a box with no ecosystem asks the question it
// always asked. A "" back means Nexus knew no such entity or Hexis had no
// matching capability, and then she asks after all. Authority is unchanged:
// resolution stops on ambiguity and a mutating capability still goes through
// the spoken confirm in handleHexisAct.
func (h *reactiveHandler) hexisBeforeClarify(ctx context.Context, dec router.Decision) string {
if h.ecosystem == nil || h.ecosystem.hexis == nil {
return ""
}
if dec.Intent != router.IntentAct || dec.Slots.HasFn || dec.Slots.Text == "" {
return ""
}
return h.handleHexisAct(ctx, dec)
}
+59
View File
@@ -0,0 +1,59 @@
package main
import (
"regexp"
"strings"
"unicode"
"github.com/kami/maven/internal/router"
)
// latinRun matches a run of Latin-script words — the shape a service, host or
// project name takes in a Russian sentence. Digits, dot, dash and underscore
// ride along because "muzick-indexer" and "nginx.conf" are one name, not two.
var latinRun = regexp.MustCompile(`[A-Za-z][A-Za-z0-9._-]*(?:\s+[A-Za-z][A-Za-z0-9._-]*)*`)
// hasLatin reports whether s carries a Latin letter.
func hasLatin(s string) bool {
for _, r := range s {
if unicode.In(r, unicode.Latin) {
return true
}
}
return false
}
// entityReferenceText is the name Nexus is asked to resolve.
//
// Normally that is the router's Text slot, which is the verb phrase the model
// wrote. But the resident model rewrites a Russian utterance as it routes, and
// on the way it transliterates: "перезапусти muzick indexer" came back as
// "перезагрузить музик индексер" (Vikunja #476). Nexus is then asked for a
// service nobody has ever named, so the act cannot resolve its target even
// with every gate open.
//
// The recovery is deliberately narrow. Only when the utterance holds a Latin
// run and the model's Text holds none has a name certainly been rewritten —
// then the longest Latin run in his own words is the reference. Anything else
// keeps the Text slot, so an English utterance and a Russian entity name are
// both untouched. Un-transliterating the Cyrillic back is not attempted: the
// surface form he said is right there, and guessing at a reverse mapping would
// invent a second name to be wrong about.
func entityReferenceText(dec router.Decision) string {
text := dec.Slots.Text
if hasLatin(text) || !hasLatin(dec.Utterance) {
return text
}
longest := ""
for _, m := range latinRun.FindAllString(dec.Utterance, -1) {
if len(m) > len(longest) {
longest = m
}
}
longest = strings.TrimSpace(longest)
// A single stray letter is not a name.
if len(longest) < 2 {
return text
}
return longest
}
+133
View File
@@ -0,0 +1,133 @@
package main
import (
"context"
"strings"
"testing"
"github.com/kami/maven/internal/router"
)
// TestEntityReferenceText pins when his own words win over the model's.
func TestEntityReferenceText(t *testing.T) {
for _, tc := range []struct {
name string
utterance string
text string
want string
}{
{
name: "the model transliterated the name",
utterance: "перезапусти muzick indexer",
text: "перезагрузить музик индексер",
want: "muzick indexer",
},
{
name: "it kept the name, so nothing to repair",
utterance: "перезапусти muzick indexer",
text: "перезагрузить muzick indexer",
want: "перезагрузить muzick indexer",
},
{
name: "an all-Russian entity name is not a rewrite",
utterance: "перезапусти домашний сервер",
text: "перезагрузить домашний сервер",
want: "перезагрузить домашний сервер",
},
{
name: "an English turn never enters the recovery",
utterance: "restart muzick indexer",
text: "restart muzick indexer",
want: "restart muzick indexer",
},
{
name: "the longest Latin run is the name",
utterance: "а перезапусти-ка nginx на muzick-indexer, пожалуйста",
text: "перезагрузить нгинкс",
want: "muzick-indexer",
},
{
name: "one stray letter is not a name",
utterance: "перезапусти сервер a",
text: "перезагрузить сервер",
want: "перезагрузить сервер",
},
} {
t.Run(tc.name, func(t *testing.T) {
dec := router.Decision{Utterance: tc.utterance, Slots: router.Slots{Text: tc.text}}
if got := entityReferenceText(dec); got != tc.want {
t.Fatalf("entityReferenceText = %q, want %q", got, tc.want)
}
})
}
}
// TestNexusIsAskedForTheNameHeSaid — the defect end to end (Vikunja #476): the
// router hands over a transliterated Text, and Nexus must still be asked about
// the service that exists.
func TestNexusIsAskedForTheNameHeSaid(t *testing.T) {
ctx := context.Background()
nexus := newFakeNexus(t, fixtureNexusResolved("ent_muzick", "Muzick indexer", "service"))
hexis := newFakeHexis(t, restartCaps(), fixtureHexisExecuted("exec_1", "succeeded"))
h := ecoHandler(t, nexus, nil, hexis)
dec := router.Decision{
Utterance: "перезапусти muzick indexer",
Intent: router.IntentAct,
Slots: router.Slots{Text: "перезагрузить музик индексер", Fn: "restart", HasFn: true},
}
h.handleHexisAct(ctx, dec)
reqs := nexus.Requests()
if len(reqs) == 0 {
t.Fatal("nexus was never asked")
}
body := string(reqs[0].Body)
if !strings.Contains(body, "muzick indexer") {
t.Fatalf("nexus resolve body = %s, want the name he said", body)
}
}
// TestAnEntityActReachesHexisInsteadOfAsking — the second half of #476. The
// stage-3 gate thins an act with no allowlisted fn, and that question used to
// be the whole turn, so the Hexis path was unreachable from voice or chat.
func TestAnEntityActReachesHexisInsteadOfAsking(t *testing.T) {
ctx := context.Background()
nexus := newFakeNexus(t, fixtureNexusResolved("ent_muzick", "Muzick indexer", "service"))
hexis := newFakeHexis(t, restartCaps(), fixtureHexisExecuted("exec_1", "succeeded"))
h := ecoHandler(t, nexus, nil, hexis)
dec := router.Decision{
Utterance: "перезапусти muzick indexer",
Intent: router.IntentAct,
Stage: 3,
Clarify: true,
Slots: router.Slots{Text: "restart status muzick indexer"},
}
reply := h.hexisBeforeClarify(ctx, dec)
if reply == "" {
t.Fatal("a resolvable entity act must reach hexis rather than fall through to the question")
}
if hexis.Count("", "/api/v1") == 0 {
t.Fatal("hexis was never contacted")
}
}
// TestClarifyStillAsksWithoutHexis — the narrowing. No ecosystem, no change:
// she asks exactly what she asked before.
func TestClarifyStillAsksWithoutHexis(t *testing.T) {
h, _, _ := newClarifyHandler(t)
dec := router.Decision{
Utterance: "перезапусти muzick indexer",
Intent: router.IntentAct,
Stage: 3,
Clarify: true,
Slots: router.Slots{Text: "перезагрузить музик индексер"},
}
if reply := h.hexisBeforeClarify(context.Background(), dec); reply != "" {
t.Fatalf("no hexis must mean no reply, got %q", reply)
}
if _, asked := h.askClarify(voiceCtx(), dec); !asked {
t.Fatal("she must still ask what to do")
}
}
+82
View File
@@ -0,0 +1,82 @@
package main
import (
"log"
"strings"
"unicode"
)
// ungroundedConfidence — what a self fact is worth when its value appears
// nowhere in what he said. Below `query_min_score` is not the point (recall
// gates on vector distance, not on this number); the point is that
// `/history` and every future reader can tell a value he said from a value
// the model supplied.
const ungroundedConfidence = 0.6
// factConfidence scores a self fact by whether its value is grounded in the
// utterance it came from. Grounded stays 1.00, which is what a tapped fact
// has always been worth. Ungrounded drops, and says so in the log.
//
// An empty value is grounded by definition: the key alone carries the fact
// ("поужинал"), and there is nothing for the model to have invented.
func factConfidence(utterance, value string) float64 {
if strings.TrimSpace(value) == "" {
return 1.0
}
if valueGrounded(utterance, value) {
return 1.0
}
log.Printf("voice: fact value %q is not in %q — writing at confidence %.2f",
value, utterance, ungroundedConfidence)
return ungroundedConfidence
}
// valueGrounded reports whether every word of value traces back to a word he
// actually said. The comparison is on a 4-rune prefix, so the model's
// normalization survives ("пил воду" → "вода") while an invented value
// ("1.20" for a question about Go) does not.
func valueGrounded(utterance, value string) bool {
said := factTokens(utterance)
words := factTokens(value)
if len(words) == 0 {
return true
}
for _, w := range words {
if !anyTokenMatches(said, w) {
return false
}
}
return true
}
func anyTokenMatches(said []string, w string) bool {
for _, s := range said {
if s == w || sameStem(s, w) {
return true
}
}
return false
}
// sameStem is inflection tolerance and nothing more: it compares all but the
// last rune of the shorter word, and never fewer than three. Russian marks
// case on the ending, so "пил воду" and the stored "вода" are the same word he
// said, while "1.20" and "версия" are not. A word of three runes or fewer must
// match outright, where a shorter prefix would match half the language.
func sameStem(a, b string) bool {
ar, br := []rune(a), []rune(b)
shorter := min(len(ar), len(br))
n := shorter - 1
if n < 3 || len(ar) < n || len(br) < n {
return false
}
return string(ar[:n]) == string(br[:n])
}
// factTokens lowercases and splits on everything that is not a letter or a
// digit, the same shape planTokens uses in the router.
func factTokens(s string) []string {
return strings.FieldsFunc(strings.ToLower(s), func(r rune) bool {
return !unicode.IsLetter(r) && !unicode.IsDigit(r)
})
}
+125
View File
@@ -0,0 +1,125 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/tool"
"github.com/kami/maven/internal/voice"
)
func newFactGateHandler(t *testing.T, now time.Time) (*reactiveHandler, ipc.CoreAPI) {
t.Helper()
st := newTestStore(t)
api := ipc.NewStoreAPI(st)
emb := router.NewHashEmbedder(1024)
h := &reactiveHandler{
api: api,
embedder: emb,
router: buildRouter(emb, tool.NewMatcher(api), 0.55, nil),
replier: voice.NewStubReplier(),
now: func() time.Time { return now },
memStore: memory.NewInMemoryStore(),
dataStore: st,
}
return h, api
}
// The write half of #470: a question routed to IntentFact must not become a
// fact about him, and must not leave a vector behind for recall to serve.
func TestActionFact_QuestionIsNotWritten(t *testing.T) {
ctx := context.Background()
h, api := newFactGateHandler(t, time.Now())
reply := h.actionFact(ctx, router.Decision{
Intent: router.IntentFact,
Utterance: "какая последняя версия языка Go?",
Slots: router.Slots{Key: "go_version", HasKey: true, Value: `"1.20"`},
})
if _, err := api.LatestFact(ctx, "go_version"); err == nil {
t.Fatal("a question was stored as a fact about him")
}
hits, err := h.memStore.Search(ctx, mustEmbedPassage(t, h, "какая последняя версия языка Go?"), 3)
if err != nil {
t.Fatalf("memory search: %v", err)
}
if len(hits) != 0 {
t.Fatalf("the question was indexed for recall: %+v", hits)
}
// It went down the query chain instead. Nothing is configured to answer a
// world question in this harness, so "не знаю." is the honest outcome —
// what matters is that the turn was answered, not stored.
if reply == "" {
t.Fatal("the turn was neither stored nor answered")
}
}
// The capture that must survive the gate: an explicit instruction to record,
// even though it contains an interrogative.
func TestActionFact_ExplicitCaptureStillWrites(t *testing.T) {
ctx := context.Background()
h, api := newFactGateHandler(t, time.Now())
h.actionFact(ctx, router.Decision{
Intent: router.IntentFact,
Utterance: "запиши что я пил воду",
Slots: router.Slots{Key: "water", HasKey: true, Value: `"вода"`},
})
f, err := api.LatestFact(ctx, "water")
if err != nil {
t.Fatalf("an explicit capture was refused: %v", err)
}
if f.Confidence != 1.0 {
t.Errorf("confidence = %v, want 1.0 for a value he said", f.Confidence)
}
// #493: what recall reads back is the fact, not the sentence he said.
// queryMemory returns a fact's text verbatim, so the utterance sitting here
// meant "запиши что я пил воду" was the answer to "когда я пил воду?".
hits, err := h.memStore.Search(ctx, mustEmbedPassage(t, h, "вода"), 3)
if err != nil {
t.Fatalf("memory search: %v", err)
}
if len(hits) != 1 {
t.Fatalf("the fact was not indexed once: %+v", hits)
}
if got := hits[0].Meta["text"]; got != "water — вода" {
t.Errorf("indexed text = %q, want the fact", got)
}
if got := hits[0].Meta["utterance"]; got != "запиши что я пил воду" {
t.Errorf("utterance provenance = %q, want it kept alongside", got)
}
}
func TestFactConfidence(t *testing.T) {
cases := []struct {
utterance, value string
want float64
}{
{"запиши что я пил воду", `"вода"`, 1.0},
{"я выпил кофе", `"кофе"`, 1.0},
{"поужинал", "", 1.0},
{"отметь что я полил кактус", `"полил кактус"`, 1.0},
{"какая последняя версия языка Go", `"1.20"`, ungroundedConfidence},
{"кто премьер Японии", `"Тонио Озаки"`, ungroundedConfidence},
}
for _, c := range cases {
if got := factConfidence(c.utterance, c.value); got != c.want {
t.Errorf("factConfidence(%q, %q) = %v, want %v", c.utterance, c.value, got, c.want)
}
}
}
func mustEmbedPassage(t *testing.T, h *reactiveHandler, text string) []float32 {
t.Helper()
vec, err := router.EmbedQuery(context.Background(), h.embedder, text)
if err != nil {
t.Fatalf("embed %q: %v", text, err)
}
return vec
}
+62 -8
View File
@@ -1,24 +1,79 @@
package main
import (
"context"
"time"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/router"
)
// voiceDialogueID — the single dialogue-session key. This is a single-user box
// (ponytail), so one slot suffices; a second speaker would need per-speaker ids,
// which waits on voice-print attribution (see PROGRESS multi-user deferral).
// voiceDialogueID — the dialogue-session key for the microphone, and the
// clarify key for it too. This is a single-user box (ponytail), so one slot
// suffices; a second speaker would need per-speaker ids, which waits on
// voice-print attribution (see PROGRESS multi-user deferral).
const voiceDialogueID = "voice"
// toDialogueSlots projects the router's slots onto the dialogue layer's subset
// (everything except the fact Value, which the dialogue layer doesn't carry).
// textDialogueID — the clarify key for a text turn that named no conversation.
// Separate from the mic: an old client that sends no id still must not answer
// a question she asked out loud.
const textDialogueID = "text"
// dialogueKey — the context key carrying the id of the conversation this turn
// belongs to. It rides the context rather than a parameter for the same reason
// the correlation id does: every step of the turn needs it, most of them only
// to hand to the next one, and threading it by hand would put it in six
// clarify signatures that have nothing else to say about it.
type dialogueKey struct{}
// dialogueIDFor builds the id a turn is held under: the conversation the reach
// named, qualified by the tap it arrived on, or the tap's own fallback when it
// named none.
//
// A parked clarifying question used to be held under voiceDialogueID no matter
// where the turn came from, so one unanswerable question captured the next
// three utterances from anywhere. Three independent curl sessions fed a
// capture attempt that had already failed, and a reminder among them was lost
// (Vikunja #466).
func dialogueIDFor(src turnSource, conversation string) string {
if conversation != "" {
return string(src) + ":" + conversation
}
if src == sourceVoice {
return voiceDialogueID
}
return textDialogueID
}
// withDialogueID tags a turn with that id.
func withDialogueID(ctx context.Context, id string) context.Context {
return context.WithValue(ctx, dialogueKey{}, id)
}
// dialogueIDOf reads it back. Falls back to the microphone's slot, which is
// what an unthreaded caller — a test, an internal replay — gets.
func dialogueIDOf(ctx context.Context) string {
if id, ok := ctx.Value(dialogueKey{}).(string); ok && id != "" {
return id
}
return voiceDialogueID
}
// toDialogueSlots and applyDialogueSlots are the only bridge between
// router.Slots and dialogue.Slots. dialogue must not import router (import
// cycle), so the two structs are hand-kept copies and every field has to be
// carried by hand here. Adding a field to either struct without adding it to
// BOTH functions loses a slot silently — nothing fails to build. The tests in
// slotsparity_test.go fail when the field sets or the converters stop matching;
// when they do, fix these two functions, not the tests.
// toDialogueSlots projects the router's slots onto the dialogue layer's copy.
func toDialogueSlots(s router.Slots) dialogue.Slots {
return dialogue.Slots{
Time: s.Time,
HasTime: s.HasTime,
Key: s.Key,
Value: s.Value,
HasKey: s.HasKey,
Text: s.Text,
Fn: s.Fn,
@@ -27,11 +82,10 @@ func toDialogueSlots(s router.Slots) dialogue.Slots {
}
}
// applyDialogueSlots writes inherited dialogue slots back onto router slots,
// preserving router-only fields (Value) the dialogue layer never touched.
// applyDialogueSlots writes dialogue slots back onto router slots.
func applyDialogueSlots(base router.Slots, d dialogue.Slots) router.Slots {
base.Time, base.HasTime = d.Time, d.HasTime
base.Key, base.HasKey = d.Key, d.HasKey
base.Key, base.Value, base.HasKey = d.Key, d.Value, d.HasKey
base.Text = d.Text
base.Fn, base.Args, base.HasFn = d.Fn, d.Args, d.HasFn
return base
+54
View File
@@ -0,0 +1,54 @@
package main
import (
"log"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/kiwix"
"github.com/kami/maven/internal/llm"
)
// kiwixWiring — the offline encyclopedia, assembled. nil ⇒ off, which is the
// default: the query chain simply has no ZIM source.
//
// The rewriter is separately optional. Searching without one is legal and
// mostly useless against English ZIMs, but it is the honest degraded mode when
// there is no llama-server to rewrite with, and it is what `rewrite: false`
// asks for.
type kiwixWiring struct {
client *kiwix.Client
rewriter *kiwix.Rewriter // nil ⇒ the question is searched verbatim
book string
max int
runes int
}
// wireKiwix builds the ZIM reader from the `kiwix` block, or returns nil when
// there is none. config.Normalise has already dropped a block with no URL and
// filled the two size defaults, so this does no validation of its own.
//
// The llm client is the phraser's swap-aware one (llmClientFor), so a model
// swap re-points the rewriter with everything else. A nil client means there is
// no resident model at all; that degrades the rewriter, not the source.
func wireKiwix(cfg *config.Config, c *llm.Client) *kiwixWiring {
if cfg.Kiwix == nil {
return nil
}
kc := cfg.Kiwix
w := &kiwixWiring{
client: kiwix.New(kc.URL),
book: kc.Book,
max: kc.MaxResults,
runes: kc.SnippetRunes,
}
switch {
case !kc.RewriteEnabled():
log.Printf("voice: kiwix at %s (book %q, query rewriting off by config)", kc.URL, kc.Book)
case c == nil:
log.Printf("voice: kiwix at %s (book %q, no llama-server: searching questions verbatim)", kc.URL, kc.Book)
default:
w.rewriter = kiwix.NewRewriter(c)
log.Printf("voice: kiwix at %s (book %q)", kc.URL, kc.Book)
}
return w
}
+187
View File
@@ -0,0 +1,187 @@
package main
import (
"context"
"net/http"
"net/http/httptest"
"strings"
"testing"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/kiwix"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice"
)
// searchRSS is what kiwix-serve answers a /search with, trimmed to the fields
// ParseSearchRSS reads.
func searchRSS(items ...string) string {
return `<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel>` +
strings.Join(items, "") + `</channel></rss>`
}
func rssItem(title, snippet string) string {
return "<item><title>" + title + "</title><link>/x</link><description>" + snippet + "</description></item>"
}
// stubKiwixServer answers every search with the given body and records the
// pattern it was asked for, so a test can assert on what left the process.
type stubKiwixServer struct {
*httptest.Server
lastPattern string
lastBook string
}
func newStubKiwix(t *testing.T, body string, status int) *stubKiwixServer {
t.Helper()
s := &stubKiwixServer{}
s.Server = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if status != 0 && status != http.StatusOK {
w.WriteHeader(status)
return
}
// Two endpoints on one server: /search answers the RSS, everything else
// is an article read. Only the search is recorded — an article fetch
// carries no query string and would blank the assertions.
if r.URL.Path != "/search" {
w.Header().Set("Content-Type", "text/html")
_, _ = w.Write([]byte("<html><title>Article</title><body><p>the lead paragraph</p></body></html>"))
return
}
s.lastPattern = r.URL.Query().Get("pattern")
s.lastBook = r.URL.Query().Get("books.name")
w.Header().Set("Content-Type", "application/xml")
_, _ = w.Write([]byte(body))
}))
t.Cleanup(s.Close)
return s
}
// buildKiwixHandler wires the source with no rewriter: the question is searched
// verbatim, which keeps the assertion about what was sent unambiguous.
func buildKiwixHandler(base string) *reactiveHandler {
return &reactiveHandler{
replier: voice.NewStubReplier(),
phraser: phraser.NewStub(),
kiwix: &kiwixWiring{
client: kiwix.New(base),
book: "wikipedia_en_all_maxi",
max: config.DefaultKiwixResults,
runes: config.DefaultKiwixSnippetRunes,
},
}
}
func askKiwix(h *reactiveHandler, q string) (string, bool) {
return h.queryKiwix(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: q},
})
}
// The default daemon has no `kiwix` block, and a source that is off must not
// claim the turn — the model answers next, exactly as it did before.
func TestQueryKiwixOffPassesThrough(t *testing.T) {
h := &reactiveHandler{replier: voice.NewStubReplier(), phraser: phraser.NewStub()}
if reply, ok := askKiwix(h, "почему небо синее?"); ok {
t.Errorf("an unconfigured kiwix claimed the turn: %q", reply)
}
}
func TestQueryKiwixAnswersFromSnippets(t *testing.T) {
s := newStubKiwix(t, searchRSS(rssItem("Rayleigh scattering", "shorter wavelengths scatter more")), 0)
h := buildKiwixHandler(s.URL)
reply, ok := askKiwix(h, "почему небо синее?")
if !ok {
t.Fatal("kiwix found a hit and did not claim the turn")
}
if reply == "" {
t.Error("claimed the turn with an empty reply")
}
if s.lastBook != "wikipedia_en_all_maxi" {
t.Errorf("books.name = %q, want the configured book", s.lastBook)
}
}
// No hit is not a failure worth announcing: the ZIM does not cover it, and the
// model answering next beats "ничего не нашла".
func TestQueryKiwixNoHitsPassesThrough(t *testing.T) {
s := newStubKiwix(t, searchRSS(), 0)
if reply, ok := askKiwix(buildKiwixHandler(s.URL), "почему небо синее?"); ok {
t.Errorf("an empty result set claimed the turn: %q", reply)
}
}
// A dead or misconfigured server must degrade to the model, not to an error
// spoken out loud. A turn never breaks on a capability.
func TestQueryKiwixServerErrorPassesThrough(t *testing.T) {
s := newStubKiwix(t, "", http.StatusBadRequest)
if reply, ok := askKiwix(buildKiwixHandler(s.URL), "почему небо синее?"); ok {
t.Errorf("a 400 claimed the turn: %q", reply)
}
}
// The privacy rule in CLAUDE.md, asserted rather than assumed: only the
// utterance is searched. No note, no fact, no persona block travels with it.
func TestQueryKiwixSendsOnlyTheQuestion(t *testing.T) {
s := newStubKiwix(t, searchRSS(rssItem("X", "y")), 0)
h := buildKiwixHandler(s.URL)
// A turn carrying notes an earlier source already pulled. They must not
// reach the query string.
_, _ = h.queryKiwix(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "почему небо синее?"},
notes: []ipc.Note{{Text: "пароль от роутера hunter2"}},
})
if strings.Contains(s.lastPattern, "hunter2") {
t.Fatalf("a stored note leaked into the search query: %q", s.lastPattern)
}
if s.lastPattern != "почему небо синее?" {
t.Errorf("pattern = %q, want the utterance verbatim", s.lastPattern)
}
}
func TestWireKiwixOffWithoutABlock(t *testing.T) {
if w := wireKiwix(&config.Config{}, nil); w != nil {
t.Error("wireKiwix built a source with no config block")
}
}
// No llama-server means no rewriter, but the source still works: searching the
// question verbatim is the honest degraded mode, not a reason to stay dark.
func TestWireKiwixWithoutAnLLMHasNoRewriter(t *testing.T) {
w := wireKiwix(&config.Config{Kiwix: &config.KiwixConfig{
URL: "http://kiwix:8080", Book: "b", MaxResults: 5, SnippetRunes: 1500,
}}, nil)
if w == nil {
t.Fatal("wireKiwix returned nil for a configured block")
}
if w.rewriter != nil {
t.Error("built a rewriter with no llm client")
}
if w.book != "b" {
t.Errorf("book = %q", w.book)
}
}
// The whole point of reading the article: kiwix's own snippet is usually the
// navigation box at the foot of the page, so the lead paragraph must be what
// reaches the phraser.
func TestQueryKiwixReadsTheArticleNotTheSnippet(t *testing.T) {
junk := "Ecological economics Ecological footprint Ecological forecasting"
s := newStubKiwix(t, searchRSS(rssItem("Photosynthesis", junk)), 0)
h := buildKiwixHandler(s.URL)
h.phraser = nil // no phraser ⇒ the fallback reads back what it was given
reply, ok := askKiwix(h, "что такое фотосинтез?")
if !ok {
t.Fatal("did not claim the turn")
}
if !strings.Contains(reply, "the lead paragraph") {
t.Errorf("reply did not come from the article: %q", reply)
}
if strings.Contains(reply, "Ecological economics") {
t.Errorf("recited the navigation-box snippet: %q", reply)
}
}
+65 -3
View File
@@ -236,7 +236,7 @@ func run(args []string) error {
coreFor := func() ipc.CoreAPI { return newIntakeAPI(ipc.NewStoreAPI(st), evBus, time.Now) }
if !locked {
rules = loop.DefaultRules()
rules = wireRules(cfg)
gatherer = loop.NewGatherer(st, rules)
if cfg.QuietHours != nil {
gatherer.SetQuietHours(cfg.QuietHours.Start, cfg.QuietHours.End)
@@ -251,6 +251,7 @@ func run(args []string) error {
Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
CacheRAMMiB: cacheRAMMiB(cfg.Phraser.CacheRAMMiB),
Timeout: time.Duration(cfg.Phraser.Timeout),
LLMNudges: cfg.Phraser.LLMNudges,
ContextBlock: contextBlockFn(cfg, time.Now),
@@ -341,6 +342,9 @@ func run(args []string) error {
if voiceW != nil && voiceW.handler != nil {
api := coreAPI.(*daemonAPI)
api.chatFn = voiceW.handler.handleText
// And the reverse: the handler was wired with the bare store
// adapter, which cannot serve the day plan. See upgradeAPI.
voiceW.handler.upgradeAPI(api)
}
if voiceW != nil && voiceW.mcp != nil {
coreAPI.(*daemonAPI).getMCPServers = voiceW.mcp.status
@@ -508,7 +512,7 @@ func run(args []string) error {
}
// Wire everything.
rules = loop.DefaultRules()
rules = wireRules(cfg)
gatherer = loop.NewGatherer(st, rules)
if cfg.QuietHours != nil {
gatherer.SetQuietHours(cfg.QuietHours.Start, cfg.QuietHours.End)
@@ -522,6 +526,7 @@ func run(args []string) error {
Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
CacheRAMMiB: cacheRAMMiB(cfg.Phraser.CacheRAMMiB),
Timeout: time.Duration(cfg.Phraser.Timeout),
LLMNudges: cfg.Phraser.LLMNudges,
ContextBlock: contextBlockFn(cfg, time.Now),
@@ -603,6 +608,7 @@ func run(args []string) error {
}
if voiceW != nil && voiceW.handler != nil {
newAPI.chatFn = voiceW.handler.handleText
voiceW.handler.upgradeAPI(newAPI)
}
srv.SetAPI(newAPI)
srv.Check = (&auth.Gate{Enrollment: auth.NewFloorEnrollment(), Session: passkeySess}).Check
@@ -751,7 +757,15 @@ func run(args []string) error {
if voiceW != nil {
voiceW.close()
}
wg.Wait()
// Bounded. Every worker below watches ctx, but one parked in a model call
// or an HTTP fetch can outlast the supervisor's patience, and run() has to
// return for `defer st.Close()` to seal the database. A worker abandoned
// mid-tick loses one tick; a shutdown that never returns loses every write
// since the last clean stop — which is how the deployed ciphertext went
// eleven days stale in July 2026.
if !waitWorkers(&wg, workerGrace) {
log.Printf("mavend: workers still running after %s, sealing anyway", workerGrace)
}
log.Printf("mavend: bye")
return nil
}
@@ -776,9 +790,57 @@ func personaFacts(cfg *config.Config) persona.Facts {
return f
}
// cacheRAMMiB resolves phraser.cache_ram_mib into the phraser's field. Unset
// means 512 MiB and not "whatever the server does", because the server's own
// default is 8 GiB of prompt cache and that is what put 7.9 GB of RSS and half
// a gigabyte of swap on homesrv for a 1.1 GB model. A negative value is the
// deliberate opt-out: no flag is passed, the server's default applies, and the
// operator owns the consequence.
func cacheRAMMiB(configured int) int {
if configured == 0 {
return 512
}
if configured < 0 {
return 0
}
return configured
}
// contextBlockFn returns the per-turn renderer of the shared context block.
// Per turn, not once at startup, because the block states the current time.
func contextBlockFn(cfg *config.Config, now func() time.Time) func() string {
f := personaFacts(cfg)
return func() string { return f.Block(now()) }
}
// workerGrace — how long shutdown waits for the background workers before it
// goes ahead and seals without them. Comfortably inside docker's ten-second
// default so the seal still lands before SIGKILL.
const workerGrace = 4 * time.Second
// waitWorkers waits on wg for at most d. Reports whether they all finished.
func waitWorkers(wg *sync.WaitGroup, d time.Duration) bool {
done := make(chan struct{})
go func() {
wg.Wait()
close(done)
}()
select {
case <-done:
return true
case <-time.After(d):
return false
}
}
// wireRules builds the nudge rule set, minus anything config turned off. The
// drop is logged because a rule vanishing silently is indistinguishable from a
// rule that is broken, and the next person to wonder why she stopped nudging
// should find the answer in the boot log.
func wireRules(cfg *config.Config) []loop.Rule {
rules, dropped := loop.RulesExcept(cfg.DisabledRules)
for _, name := range dropped {
log.Printf("loop: rule %q disabled by config", name)
}
return rules
}
+5 -5
View File
@@ -29,11 +29,11 @@ import (
// minutes of evaluation was five minutes of a mute assistant.
//
// The background client now yields the slot while a turn is in flight, so the
// collision is handled where it belongs and this is a prompt budget again.
// Sixty seconds is long enough for a Thinking model here, and an evaluation cut
// off costs nothing, because it is retried at the next interval. Raise it if
// observations start truncating.
const memoryEvalTimeout = 60 * time.Second
// collision is solved where it belongs and this is a prompt budget again. Five
// minutes is safe once more, and it is back: 60s truncated a Thinking model
// mid-synthesis, which costs an observation for no latency saved. The gate, not
// this number, is what keeps a voice turn from waiting.
const memoryEvalTimeout = 5 * time.Minute
// memoryEvalWorker — ticker + evaluator.
type memoryEvalWorker struct {
+91
View File
@@ -0,0 +1,91 @@
package main
import (
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/llm"
)
// No `workstation` block is the shipping deploy. The seam must then be the
// resident client itself, with nothing probing anything.
func TestModelSeamUnconfiguredIsResidentOnly(t *testing.T) {
resident := llm.New("http://127.0.0.1:1", time.Second)
hot, pair := modelSeam(&config.Config{}, resident)
if pair != nil {
t.Error("built a pair with no workstation configured")
}
if hot == nil {
t.Fatal("no seam at all, so the cascade would route with the classifier")
}
}
// A workstation with no resident model behind it has no floor, and a Pair with
// no floor is a configuration mistake rather than a degraded mode.
func TestModelSeamWithoutResidentIsNil(t *testing.T) {
cfg := &config.Config{Workstation: &config.WorkstationConfig{URL: "http://127.0.0.1:1"}}
cfg.Workstation.Health = strings.TrimRight(cfg.Workstation.URL, "/") + "/health"
hot, pair := modelSeam(cfg, nil)
if hot != nil || pair != nil {
t.Errorf("built a seam with no floor: hot=%v pair=%v", hot, pair)
}
}
// The configured case: the seam is the pair, and the pair notices a workstation
// that answers /health.
func TestModelSeamPrefersAnAnsweringWorkstation(t *testing.T) {
up := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(http.StatusOK)
}))
defer up.Close()
cfg := &config.Config{Workstation: &config.WorkstationConfig{
URL: up.URL,
Probe: config.Duration(10 * time.Millisecond),
}}
cfg.Workstation.Health = strings.TrimRight(cfg.Workstation.URL, "/") + "/health"
hot, pair := modelSeam(cfg, llm.New("http://127.0.0.1:1", time.Second))
if pair == nil || hot == nil {
t.Fatal("no pair built for a configured workstation")
}
defer pair.Stop()
deadline := time.Now().Add(2 * time.Second)
for !pair.Available() && time.Now().Before(deadline) {
time.Sleep(5 * time.Millisecond)
}
if !pair.Available() {
t.Fatal("the pair never saw a workstation that answers /health")
}
}
// A card held by a CPT run answers 503, and that must read as unavailable
// rather than as an error a turn has to handle.
func TestModelSeamHeldCardIsUnavailable(t *testing.T) {
busy := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Error(w, "model not loaded", http.StatusServiceUnavailable)
}))
defer busy.Close()
cfg := &config.Config{Workstation: &config.WorkstationConfig{
URL: busy.URL,
Probe: config.Duration(10 * time.Millisecond),
}}
cfg.Workstation.Health = strings.TrimRight(cfg.Workstation.URL, "/") + "/health"
_, pair := modelSeam(cfg, llm.New("http://127.0.0.1:1", time.Second))
if pair == nil {
t.Fatal("no pair built for a configured workstation")
}
defer pair.Stop()
time.Sleep(50 * time.Millisecond)
if pair.Available() {
t.Error("a 503 from the supervisor read as available")
}
}
+39
View File
@@ -0,0 +1,39 @@
package main
import (
"strings"
"testing"
"github.com/kami/maven/internal/morning"
)
// TestMorningNudgeBodySeparatesOptional — the one message a routine is allowed
// per day says what was not done, then what he could still do (Vikunja #473).
func TestMorningNudgeBodySeparatesOptional(t *testing.T) {
cand := morning.Candidate{
Routine: morning.Routine{Name: "утро"},
Missing: []morning.Item{
{Key: "meds", Label: "таблетки"},
{Key: "stretch", Label: "растяжка", Optional: true},
},
}
body := morningNudgeBody(cand)
if !strings.Contains(body, "не сделано — таблетки") {
t.Fatalf("the required item must be named as not done: %q", body)
}
if !strings.Contains(body, "если будет время — растяжка") {
t.Fatalf("the optional item must read softer: %q", body)
}
if strings.Contains(body, "не сделано — таблетки, растяжка") {
t.Fatalf("optional must not be folded into the required list: %q", body)
}
// Nothing optional missing: the sentence is what it always was.
only := morning.Candidate{
Routine: morning.Routine{Name: "утро"},
Missing: []morning.Item{{Key: "meds", Label: "таблетки"}},
}
if got, want := morningNudgeBody(only), "утро: не сделано — таблетки"; got != want {
t.Fatalf("morningNudgeBody = %q, want %q", got, want)
}
}
+79
View File
@@ -9,6 +9,7 @@ import (
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/pattern"
"github.com/kami/maven/internal/store"
@@ -283,3 +284,81 @@ func TestTickProposalCooldownSpacesAnnouncements(t *testing.T) {
}
}
}
// TestVoiceYesDoesNotAcceptRoutine — Vikunja #367. Accepting a routine hands
// the tick loop a standing new reason to speak, which DESIGN.md puts at layer
// 3, and voice is structurally incapable of layer 3. A spoken "да" must park
// the decision for the authed page, not flip the row itself.
func TestVoiceYesDoesNotAcceptRoutine(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents-1)
h := &reactiveHandler{api: ipc.NewStoreAPI(st), dataStore: st, now: func() time.Time { return now }}
// The MinEvents'th event is the one that makes the pattern detectable, and
// it goes through the voice path so the proposal is parked for a y/n.
last := now.Add(time.Duration(pattern.MinEvents-1) * 7 * 24 * time.Hour)
factID, err := st.WriteFact(ctx, last, store.KindSelf, "cat_water", "refill", "voice", 1.0, sql.NullInt64{})
if err != nil {
t.Fatalf("write fact: %v", err)
}
if phrase := h.detectPattern(ctx, factID, "cat_water", "refill", last); phrase == "" {
t.Fatal("expected a parked routine proposal")
}
reply, handled := h.resolveConfirm(ctx, "да")
if !handled {
t.Fatal("the spoken yes should be consumed by the routine confirm")
}
if !strings.Contains(reply, "рутин") {
t.Fatalf("reply should send him to the routines page, got %q", reply)
}
rows, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineAccepted)
if err != nil {
t.Fatalf("list accepted: %v", err)
}
if len(rows) != 0 {
t.Fatalf("voice accepted a routine: %+v", rows)
}
proposed, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineProposed)
if err != nil {
t.Fatalf("list proposed: %v", err)
}
if len(proposed) != 1 {
t.Fatalf("proposed routines = %d, want 1 (still waiting for the page)", len(proposed))
}
}
// TestVoiceNoStillDismissesRoutine — declining does not move the boundary
// outward, so voice keeps it. Only acceptance is gated.
func TestVoiceNoStillDismissesRoutine(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents-1)
h := &reactiveHandler{api: ipc.NewStoreAPI(st), dataStore: st, now: func() time.Time { return now }}
last := now.Add(time.Duration(pattern.MinEvents-1) * 7 * 24 * time.Hour)
factID, err := st.WriteFact(ctx, last, store.KindSelf, "cat_water", "refill", "voice", 1.0, sql.NullInt64{})
if err != nil {
t.Fatalf("write fact: %v", err)
}
if phrase := h.detectPattern(ctx, factID, "cat_water", "refill", last); phrase == "" {
t.Fatal("expected a parked routine proposal")
}
if _, handled := h.resolveConfirm(ctx, "нет"); !handled {
t.Fatal("the spoken no should be consumed by the routine confirm")
}
rows, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineDismissed)
if err != nil {
t.Fatalf("list dismissed: %v", err)
}
if len(rows) != 1 {
t.Fatalf("dismissed routines = %d, want 1", len(rows))
}
}
+154
View File
@@ -0,0 +1,154 @@
package main
import (
"context"
"log"
"math"
"sync"
"github.com/kami/maven/internal/router"
)
// The personal boundary decides one thing: is this question about him. It used
// to decide it by matching possession words, and that was the whole defect
// behind Vikunja #495. "что я говорил про бэкапы?" is his data by definition —
// nothing outside the box has ever heard him say anything — and it carried no
// possession word, so it walked past the boundary into SearXNG and came back
// answered out of a Habr article about somebody else's backups.
//
// The first fix was one more marker class, `я говорил|сказал|писал|…`, plus a
// carve-out so "как я говорил, почему небо синее" stayed a world question. Both
// halves are a lexicon, and a lexicon is the wrong instrument here: Russian
// gives every verb a dozen surface forms, the preamble list has no end, and
// every utterance the list misses is one that reaches the world. It also drifts
// silently — a missing verb looks exactly like no bug.
//
// So the boundary asks the embedder instead. Two frozen seed sets — questions
// about him, questions about the world — are embedded once, and the turn's own
// query vector, already computed by queryEmbed upstream, is scored against
// both. Nearest side wins. Word order, verb form and unseen phrasing stop
// mattering, which is exactly what a lexicon could not do.
//
// Measured 03-08-2026 against multilingual-e5-small on 19 held-out utterances,
// none of them a seed: 19 right (TestONNXPersonalBoundary). A 20th, "as i said,
// what is the population of india", missed by +0.008 during the first pass and
// is a world seed now, which is why it is not in the held-out set. True
// positives clear the world side by +0.014 to +0.089 and the nearest true
// negative sits at -0.005, so the gate is the sign of the difference and
// nothing tighter: the margins are too thin to justify a threshold, and the
// asymmetry favours claiming anyway. A false claim costs one honest "не знаю";
// a false pass sends his life to an upstream engine.
//
// The embedder is the one model CLAUDE.md pins to homesrv permanently, and it
// is what makes this affordable: no llama-server call, no network, one cosine
// per seed against a vector the turn already has.
// personalSeeds — questions about him. Frozen: they are scoring data, so
// editing one moves the boundary and must be re-measured, not eyeballed. Cover
// both classes the boundary owns, possession and first-person speech, in both
// languages.
var personalSeeds = []string{
"что я говорил про это",
"я тебе рассказывал об этом?",
"что я записал про врача",
"я упоминал эту тему?",
"что у меня сегодня",
"когда моя встреча",
"what did i say about this",
"did i mention this to you",
}
// worldSeeds — questions the world can answer, including the two shapes that
// look personal and are not: a first-person preamble on a world question ("как
// я говорил, ..."), and first person without possession ("что я могу
// посмотреть вечером"). Refusing those is the opposite mistake and the older
// comment on personalMarkers already named it.
var worldSeeds = []string{
"почему небо синее",
"какая столица франции",
"как сварить борщ",
"кто написал эту книгу",
"what is the capital of france",
"how do i boil an egg",
"как я говорил, почему небо синее",
"as i said, why is the sky blue",
"as i said, what is the population of india",
"что я могу посмотреть вечером",
"что мне почитать про историю",
"что я должен знать про питон",
"what can i watch tonight",
}
// personalBoundary holds the embedded seeds. Zero value is usable and means
// "not loaded yet"; a handler built without an embedder never loads and the
// boundary falls back to personalMarkers.
type personalBoundary struct {
once sync.Once
personal [][]float32
world [][]float32
loaded bool
}
// load embeds both seed sets, once per process. Seeds are embedded on the QUERY
// side, like the utterance they are compared with — a question against a
// question. Mixing sides would measure the e5 prefix, not the meaning.
func (b *personalBoundary) load(ctx context.Context, emb router.Embedder) {
b.once.Do(func() {
if emb == nil {
return
}
embedAll := func(ss []string) [][]float32 {
out := make([][]float32, 0, len(ss))
for _, s := range ss {
v, err := router.EmbedQuery(ctx, emb, s)
if err != nil {
log.Printf("voice: personal boundary seeds unavailable (%v); falling back to possession markers", err)
return nil
}
out = append(out, v)
}
return out
}
p, w := embedAll(personalSeeds), embedAll(worldSeeds)
if p == nil || w == nil {
return
}
b.personal, b.world, b.loaded = p, w, true
})
}
// score returns the best similarity to each side. ok is false when the seeds
// are not loaded, which is the caller's signal to use the markers instead.
func (b *personalBoundary) score(vec []float32) (personal, world float64, ok bool) {
if !b.loaded || len(vec) == 0 {
return 0, 0, false
}
best := func(seeds [][]float32) float64 {
m := -1.0
for _, s := range seeds {
if c := cosine(vec, s); c > m {
m = c
}
}
return m
}
return best(b.personal), best(b.world), true
}
// cosine — same math as internal/router and internal/memory, small enough that
// importing one of them for it would be the larger coupling.
func cosine(a, b []float32) float64 {
if len(a) != len(b) {
return 0
}
var dot, na, nb float64
for i := range a {
dot += float64(a[i]) * float64(b[i])
na += float64(a[i]) * float64(a[i])
nb += float64(b[i]) * float64(b[i])
}
if na == 0 || nb == 0 {
return 0
}
return dot / (math.Sqrt(na) * math.Sqrt(nb))
}
+94
View File
@@ -0,0 +1,94 @@
package main
import (
"context"
"os"
"path/filepath"
"testing"
"github.com/kami/maven/internal/router"
)
// A handler with no embedder never loads the seeds, so the boundary falls back
// to the possession markers. That is the offline floor and it must keep working
// — an embedder that fails to load must not open the boundary.
func TestBoundaryFallsBackToMarkersWithNoEmbedder(t *testing.T) {
h := personalHandler()
if !h.isPersonalTurn(context.Background(), &queryTurn{
dec: router.Decision{Utterance: "во сколько у меня встреча"},
}) {
t.Error("no embedder: a possession question must still be personal")
}
if h.isPersonalTurn(context.Background(), &queryTurn{
dec: router.Decision{Utterance: "почему небо синее"},
}) {
t.Error("no embedder: a world question must still pass")
}
}
// TestONNXPersonalBoundary — the number that matters, scored against the
// embedder homesrv actually runs. Opt-in via MAVEN_ONNX_LIB, exactly like
// TestONNXRecall in internal/memory/recalleval.
//
// Every case here is held out: none of these strings is a seed. The #495
// regression is the first row — "что я говорил про бэкапы?" reached SearXNG and
// was answered from a Habr article, and no possession word appears in it.
func TestONNXPersonalBoundary(t *testing.T) {
lib := os.Getenv("MAVEN_ONNX_LIB")
if lib == "" {
t.Skip("MAVEN_ONNX_LIB unset — see AGENTS.md § Embedder model for intent routing")
}
dir := filepath.Join("../..", "models/embedder/multilingual-e5-small")
emb, err := router.NewONNXEmbedder(filepath.Join(dir, "model_quantized.onnx"), filepath.Join(dir, "tokenizer.json"), lib)
if err != nil {
t.Skipf("onnx embedder unavailable: %v", err)
}
defer emb.Close()
cases := []struct {
utterance string
personal bool
}{
{"что я говорил про бэкапы?", true},
{"что я сказал вчера про отпуск", true},
{"я писал что-нибудь про сервер", true},
{"я упоминал про конференцию?", true},
{"что я отмечал по поводу переезда", true},
{"я рассказывал тебе про новую работу?", true},
{"во сколько у меня встреча", true},
{"когда мой следующий отпуск", true},
{"what did i say about backups", true},
{"did i tell you about the doctor", true},
{"как я говорил, почему небо синее", false},
{"как уже я говорил, какая столица франции", false},
{"почему трава зелёная", false},
{"столица франции", false},
{"как мне сварить борщ", false},
{"что мне посмотреть вечером", false},
{"я хочу узнать про рим", false},
{"кто такой гагарин", false},
{"how do i boil an egg", false},
}
h := &reactiveHandler{embedder: emb}
ctx := context.Background()
wrong := 0
for _, c := range cases {
vec, err := router.EmbedQuery(ctx, emb, c.utterance)
if err != nil {
t.Fatalf("embed %q: %v", c.utterance, err)
}
turn := &queryTurn{dec: router.Decision{Utterance: c.utterance}, vec: vec}
got := h.isPersonalTurn(ctx, turn)
p, w, ok := h.boundary.score(vec)
if !ok {
t.Fatal("seeds did not load with a working embedder")
}
if got != c.personal {
wrong++
t.Errorf("%q: personal=%v want %v (personal %.4f world %.4f)", c.utterance, got, c.personal, p, w)
}
t.Logf("personal=%-5v personal %.4f world %.4f delta %+.4f %s", got, p, w, p-w, c.utterance)
}
t.Logf("personal boundary: %d/%d held-out utterances correct", len(cases)-wrong, len(cases))
}
+16 -3
View File
@@ -133,10 +133,21 @@ var (
quietOnPhrases = [][]string{
{"quiet", "on"}, {"quiet", "mode"},
{"тих", "режим"}, {"не", "шум"}, {"не", "беспоко"},
{"тих"},
// The noun form and the comparative. "режим тишины" is how the
// setting is named half the time, and "сделай потише" is how it is
// actually asked for out loud. Both used to fall through to the
// router, which has no quiet intent, so the command did nothing.
{"режим", "тишин"}, {"сделай", "тише"}, {"сделай", "потише"},
{"говори", "тише"}, {"будь", "потише"},
{"тих"}, {"потише"},
}
)
// quietWordStems — every stem that names the setting. Used by the
// negated-but-unmatched fallback in classifyQuietToggle, which has to
// recognise "хватит тишины" without an ON phrase having matched.
var quietWordStems = []string{"тих", "тишин", "потише"}
// quietNegatorWords — negators that are whole words with no useful stem.
var quietNegatorWords = map[string]bool{
"не": true, "нет": true, "хватит": true, "no": true, "not": true, "off": true,
@@ -197,8 +208,10 @@ func classifyQuietToggle(text string) (on, off bool) {
// on after he asked for it to stop.
if quietNegated(tokens, nil) {
for _, t := range tokens {
if quietStem(t, "тих") {
return false, true
for _, stem := range quietWordStems {
if quietStem(t, stem) {
return false, true
}
}
}
}
+17
View File
@@ -47,6 +47,15 @@ func TestResolveQuietToggle(t *testing.T) {
{"побудь в тихом режиме", quietOn},
{"Тихий Режим!", quietOn},
{"тихая", quietOn},
// The noun form and the comparative.
{"включи режим тишины", quietOn},
{"режим тишины", quietOn},
{"сделай потише", quietOn},
{"сделай тише", quietOn},
{"потише", quietOn},
// English, as the fixture phrases it.
{"turn quiet mode back on", quietOn},
{"enable quiet mode", quietOn},
// OFF vocabulary — all seven, incl. the three that used to say ON.
{"quiet off", quietOff},
@@ -60,6 +69,12 @@ func TestResolveQuietToggle(t *testing.T) {
{"выключи тихий режим", quietOff},
{"отмени тихий режим пожалуйста", quietOff},
{"верни громкий режим", quietOff},
{"выключи режим тишины", quietOff},
{"хватит тишины", quietOff},
{"turn off quiet mode", quietOff},
{"quiet mode off", quietOff},
{"stop quiet mode", quietOff},
{"disable quiet mode", quietOff},
// False positives: "тихо"/"тихий" as ordinary Russian.
{"очень тихий сегодня день", quietNone},
@@ -68,6 +83,8 @@ func TestResolveQuietToggle(t *testing.T) {
{"потихоньку", quietNone},
{"тихонько", quietNone},
{"он говорил тихим голосом весь вечер", quietNone},
{"в тишине лучше думается", quietNone},
{"на улице стало потише", quietNone},
// Unrelated.
{"напомни завтра позвонить маме", quietNone},
+37
View File
@@ -2,12 +2,14 @@ package main
import (
"context"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/tool"
"github.com/kami/maven/internal/voice"
)
@@ -85,3 +87,38 @@ func TestReactiveNotesReminders(t *testing.T) {
}
})
}
// TestSpokenTaskCaptureFilesATask — the whole path, from the utterance to the
// task table. It went dead when the router started claiming the marker as an
// act: capture rides the note intent, so nothing below actionNote was ever
// reached and every capture answered "Что сделать?" (Vikunja #467).
func TestSpokenTaskCaptureFilesATask(t *testing.T) {
ctx := context.Background()
st := newTestStore(t)
api := ipc.NewStoreAPI(st)
now := time.Now()
emb := router.NewHashEmbedder(1024)
matcher := tool.NewMatcher(api)
h := &reactiveHandler{
api: api,
embedder: emb,
router: buildRouter(emb, matcher, 0.55, nil),
replier: voice.NewStubReplier(),
now: func() time.Time { return now },
memStore: memory.NewInMemoryStore(),
dataStore: st,
}
reply := h.handleText(ctx, "web", "добавь в задачи купить молоко")
if !strings.Contains(reply, "купить молоко") {
t.Fatalf("capture did not claim the turn: %q", reply)
}
open, err := st.ListTasks(ctx, store.TaskOpen)
if err != nil || len(open) != 1 {
t.Fatalf("task was not filed: tasks=%v err=%v", open, err)
}
// The words he said, not the model's rewrite of them.
if open[0].Text != "купить молоко" {
t.Fatalf("task text was rewritten: %q", open[0].Text)
}
}
+12 -99
View File
@@ -2,120 +2,33 @@ package main
import (
"context"
"encoding/json"
"strings"
"time"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice"
)
// completer is the LLM seam for the replier (subset of router.Completer).
// *llm.Client satisfies it.
type completer interface {
Complete(ctx context.Context, r llm.Req) (string, error)
}
// llmReplier phrases reactive confirmations with the resident model
// (Qwen3-1.7B). Stub is the
// floor on any error (offline-safe). Maven speaks as "she", feminine RU.
// llmReplier is the daemon-side wiring around phraser.Replier: it owns the
// deterministic floor, and nothing else. The phrasing itself, the prompt and the
// output parsing live in internal/phraser so the eval can score them (#396).
type llmReplier struct {
c completer
p *phraser.Replier
stub *voice.StubReplier
// block renders the shared context block per turn (who he is, the time).
// nil ⇒ the prompt stands alone.
block func() string
}
func newLLMReplier(c completer, block func() string) *llmReplier {
return &llmReplier{c: c, stub: voice.NewStubReplier(), block: block}
func newLLMReplier(c phraser.Completer, block func() string) *llmReplier {
return &llmReplier{p: phraser.NewReplier(c, block), stub: voice.NewStubReplier()}
}
const replySystem = `Ты — Maven, домашняя ассистентка (о себе — в женском роде). Владелец — мужчина, говоришь с ним на "ты", в единственном числе; никогда не "вы"/"ваш" и не "он"/"его". Подтверди действие РОВНО ОДНИМ коротким предложением (≤120 символов), по-русски, спокойно и без официальных формулировок. Не задавай вопросов, не повторяй слова, не добавляй ничего после точки. Отвечай ТОЛЬКО одним объектом JSON с полями "response" (текст) и "mood" (ровно одно из: neutral, happy, thinking, tired, confused).
Пример: {"response": "Записала, что ты выпил стакан воды.", "mood": "neutral"}
Никогда не пиши "..." в поле response.`
// Reply never fails: a clarify, a model error and an unusable generation all
// answer from the stub, which is what keeps a turn from breaking on the model.
func (r *llmReplier) Reply(d router.Decision) string {
if d.Clarify {
return r.stub.Reply(d)
}
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
out, err := r.c.Complete(ctx, llm.Req{
System: persona.Prepend(r.block, replySystem),
User: replyContext(d),
MaxTokens: 512,
RepeatPenalty: 1.3,
})
if err != nil {
out, err := r.p.PhraseReply(context.Background(), d)
if err != nil || out == "" {
return r.stub.Reply(d)
}
out = stripThink(out)
if response, _ := parseResponseMood(out); response != "" {
return response
}
// fallback: try plain-text parsing
if out = firstSentence(out); out != "" {
return out
}
return r.stub.Reply(d)
}
// firstSentence trims the model's output to a single clean confirmation: first
// line, first sentence, whitespace-normalized — the last-line defense against a
// small model that rambles past the first period despite the prompt + stop.
// stripThink removes the <think> block that Thinking-variant models emit.
func stripThink(s string) string {
if i := strings.LastIndex(s, "</think>"); i >= 0 {
s = strings.TrimSpace(s[i+8:])
}
return s
}
func firstSentence(s string) string {
s = strings.TrimSpace(s)
if i := strings.IndexByte(s, '\n'); i >= 0 {
s = s[:i]
}
// keep up to and including the first sentence-ending punctuation.
if i := strings.IndexAny(s, ".!?"); i >= 0 {
s = s[:i+1]
}
return strings.TrimSpace(s)
}
// parseResponseMood extracts {"response","mood"} from LLM output, tolerant
// of thinking tokens and extra text before/after the JSON block.
func parseResponseMood(raw string) (response, mood string) {
cleaned := strings.TrimSpace(raw)
start := strings.Index(cleaned, "{")
end := strings.LastIndex(cleaned, "}")
if start < 0 || end < 0 || end <= start {
return "", ""
}
var parsed struct {
Response string `json:"response"`
Mood string `json:"mood"`
}
if err := json.Unmarshal([]byte(cleaned[start:end+1]), &parsed); err != nil {
return "", ""
}
return parsed.Response, parsed.Mood
}
// replyContext renders the decision into a compact RU description for the model.
func replyContext(d router.Decision) string {
switch d.Intent {
case router.IntentFact:
return "записала факт: " + d.Slots.Key + " " + d.Slots.Value
case router.IntentNote:
return "сохранила заметку: " + d.Slots.Text
case router.IntentReminder:
return "поставила напоминание: " + d.Slots.Text
default:
return string(d.Intent) + ": " + d.Slots.Text
}
return out
}
+20 -32
View File
@@ -9,23 +9,18 @@ import (
"github.com/kami/maven/internal/voice"
)
type mockCompleter struct {
// The phrasing itself is tested in internal/phraser. What is left here is the
// only thing the daemon adds: the stub floor, on the three ways a reply can
// fail to arrive.
type stubCompleter struct {
out string
err error
}
func (m mockCompleter) Complete(_ context.Context, _ llm.Req) (string, error) { return m.out, m.err }
func (s stubCompleter) Complete(_ context.Context, _ llm.Req) (string, error) { return s.out, s.err }
func TestLLMReplierReturnsLLMReply(t *testing.T) {
r := newLLMReplier(mockCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`}, nil)
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
if got != "записала, кофе закончился" {
t.Errorf("got %q, want %q", got, "записала, кофе закончился")
}
}
func TestLLMReplierFallsBackToPlainText(t *testing.T) {
r := newLLMReplier(mockCompleter{out: "записала, кофе закончился"}, nil)
func TestLLMReplierPassesTheModelReplyThrough(t *testing.T) {
r := newLLMReplier(stubCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`}, nil)
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
if got != "записала, кофе закончился" {
t.Errorf("got %q, want %q", got, "записала, кофе закончился")
@@ -33,36 +28,29 @@ func TestLLMReplierFallsBackToPlainText(t *testing.T) {
}
func TestLLMReplierFallsBackToStubOnError(t *testing.T) {
r := newLLMReplier(mockCompleter{err: errTestLLMDown}, nil)
noteDec := router.Decision{Intent: router.IntentNote}
got := r.Reply(noteDec)
want := voice.NewStubReplier().Reply(noteDec)
if got != want {
t.Errorf("on llm error: got %q, want stub %q", got, want)
}
r := newLLMReplier(stubCompleter{err: errReplierTest}, nil)
assertStub(t, r, router.Decision{Intent: router.IntentNote}, "llm error")
}
func TestLLMReplierFallsBackToStubOnEmpty(t *testing.T) {
r := newLLMReplier(mockCompleter{out: ""}, nil)
noteDec := router.Decision{Intent: router.IntentNote}
got := r.Reply(noteDec)
want := voice.NewStubReplier().Reply(noteDec)
if got != want {
t.Errorf("on empty llm: got %q, want stub %q", got, want)
}
r := newLLMReplier(stubCompleter{out: ""}, nil)
assertStub(t, r, router.Decision{Intent: router.IntentNote}, "empty llm")
}
func TestLLMReplierClarifyUsesStub(t *testing.T) {
r := newLLMReplier(mockCompleter{out: "я всё поняла"}, nil)
clarifyDec := router.Decision{Clarify: true}
got := r.Reply(clarifyDec)
want := voice.NewStubReplier().Reply(clarifyDec)
r := newLLMReplier(stubCompleter{out: "я всё поняла"}, nil)
assertStub(t, r, router.Decision{Clarify: true}, "clarify")
}
func assertStub(t *testing.T, r *llmReplier, d router.Decision, what string) {
t.Helper()
got, want := r.Reply(d), voice.NewStubReplier().Reply(d)
if got != want {
t.Errorf("on clarify: got %q, want stub %q", got, want)
t.Errorf("on %s: got %q, want stub %q", what, got, want)
}
}
var errTestLLMDown = errTest("llm down")
var errReplierTest = errTest("llm down")
type errTest string
+41
View File
@@ -0,0 +1,41 @@
package main
import (
"log"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/websearch"
)
// searchWiring — the metasearch source, assembled. nil ⇒ off, which is the
// default: no `search` block, no query ever leaves the LAN.
//
// Thinner than kiwixWiring because there is nothing to rewrite. SearXNG ranks
// with real engines, so the question goes out as he asked it, and that is the
// reason this source sits ahead of the ZIMs rather than behind them.
type searchWiring struct {
client *websearch.Client
max int
runes int
}
// wireSearch builds the search client from the `search` block, or returns nil
// when there is none. config.Normalise has already dropped a block with no URL
// and filled the two size defaults, so this does no validation of its own.
func wireSearch(cfg *config.Config) *searchWiring {
if cfg.Search == nil {
return nil
}
sc := cfg.Search
log.Printf("voice: web search at %s (language %q, engines %q)", sc.URL, sc.Language, sc.Engines)
return &searchWiring{
client: websearch.New(sc.URL, websearch.Options{
Language: sc.Language,
Engines: sc.Engines,
Timeout: time.Duration(sc.Timeout),
}),
max: sc.MaxResults,
runes: sc.SnippetRunes,
}
}
+28 -2
View File
@@ -1,5 +1,5 @@
// mavend/simulator_test.go — the replayable full-system simulator
// (Vikunja #284, 20-07-2026-BACKLOG.md item 7).
// (Vikunja #284).
//
// # What it is
//
@@ -328,7 +328,10 @@ func (s *scriptedLLM) Complete(_ context.Context, r llm.Req) (string, error) {
s.mu.Lock()
defer s.mu.Unlock()
s.calls = append(s.calls, r)
routing := r.Grammar != ""
// A grammar no longer separates the two contracts — the replier carries one
// too since phraser.ResponseGrammar was attached to it. Only the router's
// grammar names the intent enum, so that is what tells them apart.
routing := strings.Contains(r.Grammar, "intent")
for _, e := range s.entries {
if e.Match != "" && !strings.Contains(strings.ToLower(r.User), strings.ToLower(e.Match)) {
continue
@@ -1040,3 +1043,26 @@ func TestSimulatorRefusesBackwardsSteps(t *testing.T) {
t.Errorf("the clock moved to %s on a refused step, it must stay at 09:00", got)
}
}
// TestSimulatorRoutesWithTheDeployedSeeds — the scenarios must replay against
// the classifier the deploy runs, not an empty one.
//
// They did not. The seed path was relative to the working directory, which is
// cmd/mavend under `go test`, so every file failed to open and the whole
// simulator scored three green scenarios with zero examples loaded (Vikunja
// #465). The count is asserted rather than logged, because a silent zero is
// exactly the failure that hid here for as long as it did.
func TestSimulatorRoutesWithTheDeployedSeeds(t *testing.T) {
cls := router.NewClassifier(router.NewHashEmbedder(1024))
seedClassifier(cls)
total := 0
for _, intent := range cls.Intents() {
total += len(cls.Examples(intent))
}
if total == 0 {
t.Fatalf("no seed examples loaded from %s — the simulator would route on nothing", seedPath())
}
if len(cls.Intents()) != 7 {
t.Fatalf("seeded %d intents, want all 7", len(cls.Intents()))
}
}
+71
View File
@@ -0,0 +1,71 @@
package main
import (
"reflect"
"testing"
"time"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/router"
)
// TestSlotsParity — dialogue.Slots is a hand-kept copy of router.Slots
// (dialogue must not import router: import cycle). Drift is silent, so this
// test compares the two field sets by name and type. If it fails, add the new
// field to both structs AND to toDialogueSlots/applyDialogueSlots in
// followup.go — do not relax the test.
func TestSlotsParity(t *testing.T) {
fields := func(v any) map[string]string {
rt := reflect.TypeOf(v)
out := make(map[string]string, rt.NumField())
for i := 0; i < rt.NumField(); i++ {
f := rt.Field(i)
out[f.Name] = f.Type.String()
}
return out
}
rf, df := fields(router.Slots{}), fields(dialogue.Slots{})
for name, typ := range rf {
dt, ok := df[name]
if !ok {
t.Errorf("router.Slots.%s (%s) missing from dialogue.Slots", name, typ)
continue
}
if dt != typ {
t.Errorf("field %s: router has %s, dialogue has %s", name, typ, dt)
}
}
for name, typ := range df {
if _, ok := rf[name]; !ok {
t.Errorf("dialogue.Slots.%s (%s) missing from router.Slots", name, typ)
}
}
}
// TestSlotsRoundTrip — the converters carry every field. A field the parity
// test accepts can still be dropped in transit, so round-trip a fully
// populated value and compare.
func TestSlotsRoundTrip(t *testing.T) {
full := router.Slots{
Time: time.Date(2026, 8, 2, 11, 0, 0, 0, time.UTC),
HasTime: true,
Fn: "restart",
Args: []string{"nginx"},
HasFn: true,
Key: "water",
Value: `"drank"`,
HasKey: true,
Text: "выпил воды",
}
// Every field must be non-zero, or the round-trip proves nothing.
rv := reflect.ValueOf(full)
for i := 0; i < rv.NumField(); i++ {
if rv.Field(i).IsZero() {
t.Fatalf("field %s is zero: extend this fixture so the round-trip covers it",
rv.Type().Field(i).Name)
}
}
if got := applyDialogueSlots(router.Slots{}, toDialogueSlots(full)); !reflect.DeepEqual(got, full) {
t.Errorf("round-trip lost a slot:\n got %+v\nwant %+v", got, full)
}
}
+107
View File
@@ -0,0 +1,107 @@
// Spoken snooze — "не сейчас", "потом", "отложи" said out loud after a nudge
// resolves it as `snoozed`, the same outcome the Telegram buttons and the web
// UI write. Until this existed, a nudge could only be deferred by touching a
// screen: the voice path had no way to reach store.ResolveNudge at all, so the
// one channel she nudges on hardest was the one channel he could not answer.
package main
import (
"context"
"log"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/store"
)
// snoozeWindow — how long after a send "потом" still means "that nudge".
//
// A window is what makes this safe to run before the router. "потом" is an
// ordinary Russian word; eating every one of them would break real sentences.
// Bounded to the minutes right after she spoke, the word is almost always an
// answer to what she just said, and outside the window the utterance falls
// through and routes normally.
//
// Twenty minutes rather than the two hours of store.SnoozeDuration: those
// measure different things. SnoozeDuration is how long the quiet lasts,
// snoozeWindow is how long an unanswered nudge stays the topic of the
// conversation.
const snoozeWindow = 20 * time.Minute
// snoozeScan — how many recent nudges to look at when finding the target. The
// newest pending one is nearly always the first row; a handful of resolved
// rows can sit in front of it when he acked a few in a row.
const snoozeScan = 10
// resolveSnooze — pre-route keyword check, run after the quiet toggle. Returns
// (reply, true) when the utterance defers a nudge she recently sent.
//
// It returns ("", false) in two different situations, on purpose: the words do
// not read as a deferral, or they do but there is nothing pending to defer. In
// both the turn keeps routing, so "потом посмотрю что там с бэкапом" is still
// a query when no nudge is outstanding.
func (h *reactiveHandler) resolveSnooze(ctx context.Context, text string, src turnSource) (string, bool) {
if !classifySnooze(text) {
return "", false
}
now := h.now()
target, ok := h.pendingNudge(ctx, now)
if !ok {
return "", false
}
if err := h.api.ResolveNudge(ctx, target.ID, store.NudgeSnoozed, now); err != nil {
log.Printf("voice: snooze nudge %d (%s, %s): %v", target.ID, target.Rule, src, err)
return "не получилось отложить.", true
}
log.Printf("voice: snoozed nudge %d (rule %s) from %s", target.ID, target.Rule, src)
return "хорошо, вернусь к этому позже.", true
}
// pendingNudge — the newest still-pending nudge sent inside snoozeWindow.
//
// Channel is deliberately not filtered. A nudge that went to Telegram is still
// the thing he is answering when he says "потом" at the microphone, and making
// the reply channel decide which nudges are answerable would mean the ops page
// he actually read could not be dismissed by voice.
func (h *reactiveHandler) pendingNudge(ctx context.Context, now time.Time) (ipc.Nudge, bool) {
recent, err := h.api.RecentNudges(ctx, snoozeScan)
if err != nil {
log.Printf("voice: recent nudges for snooze: %v", err)
return ipc.Nudge{}, false
}
for _, n := range recent {
if n.Outcome != store.NudgePending {
continue
}
if now.Sub(n.Ts) > snoozeWindow || n.Ts.After(now) {
continue
}
return n, true
}
return ipc.Nudge{}, false
}
// snoozePhrases — the deferral vocabulary, as stem sequences. Matched by
// quietPhrase (quiet_toggle.go), which carries the rule that matters here:
// a single-word pattern matches only a single-word utterance. Bare "потом" is
// an answer; "потом схожу за водой" is a plan, and reporting a plan must not
// silence the rule that prompted it.
var snoozePhrases = [][]string{
{"не", "сейчас"}, {"не", "могу", "сейчас"}, {"не", "до", "этого"},
{"напомн", "позже"}, {"напомн", "потом"}, {"спрос", "позже"},
{"отлож"}, {"позже"}, {"потом"}, {"попозже"}, {"погоди"},
{"not", "now"}, {"later"}, {"snooze"}, {"remind", "me", "later"},
}
// classifySnooze reads an utterance as a deferral. Unlike the quiet toggle
// there is no negation arm: "не потом" is not something anyone says, and the
// leading "не" of "не сейчас" is part of the phrase itself.
func classifySnooze(text string) bool {
tokens := quietTokens(text)
for _, p := range snoozePhrases {
if quietPhrase(tokens, p) {
return true
}
}
return false
}
+115
View File
@@ -0,0 +1,115 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/store"
)
// snoozeFakeAPI serves a fixed nudge list and records the resolution.
type snoozeFakeAPI struct {
ipc.UnimplementedCoreAPI
nudges []ipc.Nudge
gotID int64
gotOutcome string
calls int
}
func (a *snoozeFakeAPI) RecentNudges(_ context.Context, _ int) ([]ipc.Nudge, error) {
return a.nudges, nil
}
func (a *snoozeFakeAPI) ResolveNudge(_ context.Context, id int64, outcome string, _ time.Time) error {
a.gotID, a.gotOutcome, a.calls = id, outcome, a.calls+1
return nil
}
var snoozeNow = time.Date(2026, 8, 1, 12, 0, 0, 0, time.UTC)
func snoozeHandler(nudges []ipc.Nudge) (*reactiveHandler, *snoozeFakeAPI) {
api := &snoozeFakeAPI{nudges: nudges}
return &reactiveHandler{api: api, now: func() time.Time { return snoozeNow }}, api
}
func pendingNudgeAt(id int64, ago time.Duration) ipc.Nudge {
return ipc.Nudge{ID: id, Ts: snoozeNow.Add(-ago), Rule: "water", Channel: "voice", Outcome: store.NudgePending}
}
func TestClassifySnooze(t *testing.T) {
yes := []string{
"не сейчас", "потом", "позже", "попозже", "отложи", "погоди",
"напомни позже", "напомни потом", "не могу сейчас",
"not now", "later", "snooze",
}
for _, s := range yes {
if !classifySnooze(s) {
t.Errorf("classifySnooze(%q) = false, want true", s)
}
}
no := []string{
// A single-word pattern must not eat the sentence it appears in.
"потом схожу за водой", "позже посмотрю что там с бэкапом",
"напомни завтра позвонить маме", "какая погода", "погода на завтра",
"я отложил деньги", "", "тихий режим",
}
for _, s := range no {
if classifySnooze(s) {
t.Errorf("classifySnooze(%q) = true, want false", s)
}
}
}
func TestResolveSnoozeDefersTheNewestPendingNudge(t *testing.T) {
h, api := snoozeHandler([]ipc.Nudge{
{ID: 9, Ts: snoozeNow.Add(-time.Minute), Rule: "meal", Outcome: store.NudgeActed},
pendingNudgeAt(8, 3*time.Minute),
pendingNudgeAt(7, 10*time.Minute),
})
reply, handled := h.resolveSnooze(context.Background(), "не сейчас", sourceVoice)
if !handled || reply == "" {
t.Fatalf("got (%q, %v), want a reply", reply, handled)
}
if api.gotID != 8 || api.gotOutcome != store.NudgeSnoozed {
t.Fatalf("resolved (%d, %q), want (8, %q)", api.gotID, api.gotOutcome, store.NudgeSnoozed)
}
}
func TestResolveSnoozeFallsThroughWithNothingPending(t *testing.T) {
// The whole point of the window: with no live nudge, "потом" is just a
// word and must keep routing.
for _, name := range []string{"stale", "resolved", "empty"} {
var nudges []ipc.Nudge
switch name {
case "stale":
nudges = []ipc.Nudge{pendingNudgeAt(3, snoozeWindow+time.Minute)}
case "resolved":
nudges = []ipc.Nudge{{ID: 4, Ts: snoozeNow, Rule: "water", Outcome: store.NudgeActed}}
}
t.Run(name, func(t *testing.T) {
h, api := snoozeHandler(nudges)
reply, handled := h.resolveSnooze(context.Background(), "потом", sourceVoice)
if handled || reply != "" {
t.Fatalf("got (%q, %v), want fall-through", reply, handled)
}
if api.calls != 0 {
t.Fatalf("resolved a nudge with nothing pending")
}
})
}
}
func TestResolveSnoozeIgnoresAFutureNudge(t *testing.T) {
// Clock skew between the tick and the turn must not let a send from the
// future be answered before it happened.
h, api := snoozeHandler([]ipc.Nudge{pendingNudgeAt(5, -time.Minute)})
if _, handled := h.resolveSnooze(context.Background(), "потом", sourceVoice); handled {
t.Fatalf("snoozed a nudge dated in the future")
}
if api.calls != 0 {
t.Fatalf("resolved a future nudge")
}
}
+30 -694
View File
@@ -10,22 +10,17 @@ package main
import (
"context"
"errors"
"fmt"
"log"
"os"
"path/filepath"
"strings"
"sync"
"time"
"github.com/kami/maven/internal/calendar"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/morning"
"github.com/kami/maven/internal/pattern"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/routine"
"github.com/kami/maven/internal/store"
@@ -250,6 +245,7 @@ func (t *tickLoop) tick(ctx context.Context, now time.Time) {
log.Printf("tick: unacked telegram rules: %v", err)
return
}
keys = t.repeatableRules(keys)
if len(keys) == 0 {
return
}
@@ -261,6 +257,35 @@ func (t *tickLoop) tick(ctx context.Context, now time.Time) {
}
}
// repeatableRules drops keys whose rule is not wired any more.
//
// The repeat path reads the nudges table, not the rule set: any sev4 telegram
// row still at outcome=pending is re-sent every repeat_interval until it is
// acked. So turning a rule off in `disabled_rules` silenced new nudges and left
// the last un-acked one re-sending every five minutes, forever — a knob that
// stops the cause and not the symptom is worse than no knob. Found the evening
// of 2026-08-01, two messages after the rule was supposedly off.
//
// Filtering on the wired set rather than on the disabled list also covers the
// rule that was deleted from the code entirely: its orphan rows go quiet
// instead of nagging about a rule nobody can ack from the UI any more.
func (t *tickLoop) repeatableRules(keys []string) []string {
if len(keys) == 0 {
return nil
}
wired := make(map[string]bool, len(t.rules))
for _, r := range t.rules {
wired[r.Name] = true
}
out := keys[:0:0]
for _, k := range keys {
if wired[k] {
out = append(out, k)
}
}
return out
}
// cachePhrase keeps the latest phrased nudge per rule for the sev4-repeat
// path. writing under a mutex; the repeat path reads under the same. the
// cache is bounded by the rule count (≤ ~30 per spec) so eviction is not a
@@ -285,597 +310,6 @@ func (t *tickLoop) repeatPhrase(rule string) (body, summary string) {
return pn.Body, pn.Summary
}
// shouldQueue — true when digest is enabled and the candidate's severity is
// at or below the configured ceiling.
func (t *tickLoop) shouldQueue(cand *loop.Candidate) bool {
return t.digestCfg != nil && t.digestCfg.Enabled &&
cand.Severity <= loop.Severity(t.digestCfg.SeverityCeiling)
}
// queueNudge — phrases the candidate and appends it to the digest queue.
// Deduplicates by rule name: if the same rule is already queued, this is a
// no-op (the first fire within the window is the one that counts).
func (t *tickLoop) queueNudge(ctx context.Context, cand *loop.Candidate, _ loop.State, now time.Time) {
for _, q := range t.digestQ {
if q.Rule == cand.Rule.Name {
return // already queued
}
}
pn, err := t.phraser.PhraseNudge(ctx, *cand)
if err != nil {
log.Printf("tick: phrase nudge %s: %v", cand.Rule.Name, err)
return
}
t.digestQ = append(t.digestQ, QueuedNudge{
Rule: cand.Rule.Name,
Severity: int(cand.Severity),
Body: pn.Body,
Key: cand.Rule.Name,
QueuedAt: now,
})
t.cachePhrase(pn)
}
// maybeFlush — flushes the digest queue if the window has elapsed since the
// first item or the queue reached MaxItems.
func (t *tickLoop) maybeFlush(ctx context.Context, now time.Time, state loop.State) {
if t.digestCfg == nil || !t.digestCfg.Enabled || len(t.digestQ) == 0 {
return
}
first := t.digestQ[0]
if now.Sub(first.QueuedAt) >= time.Duration(t.digestCfg.Window) ||
len(t.digestQ) >= t.digestCfg.MaxItems {
t.flushDigest(ctx, now, state)
}
}
// flushDigest — concatenates queued nudge bodies into a single digest
// notification and dispatches it. Clears the queue after a successful send.
// The digest uses the max severity among queued items for routing.
func (t *tickLoop) flushDigest(ctx context.Context, now time.Time, state loop.State) {
if len(t.digestQ) == 0 {
return
}
var b strings.Builder
maxSev := 0
for i, q := range t.digestQ {
if i > 0 {
b.WriteString(" · ")
}
b.WriteString(q.Body)
if q.Severity > maxSev {
maxSev = q.Severity
}
}
body := b.String()
summary := fmt.Sprintf("%d pending notifications", len(t.digestQ))
cand := loop.Candidate{
Rule: loop.Rule{
Name: "digest",
Severity: loop.Severity(maxSev),
},
Severity: loop.Severity(maxSev),
State: state,
}
pn := delivery.PhrasedNudge{
Candidate: cand,
Body: body,
Summary: summary,
}
t.cachePhrase(pn)
if _, err := t.dispatcher.DispatchNudge(ctx, pn, now); err != nil {
// keep the queue — the next tick's maybeFlush re-attempts.
log.Printf("tick: dispatch digest: %v", err)
return
}
t.digestQ = nil
}
// detectPatterns runs the pattern detector proactively over every
// action+object pair that has ever produced an event, independent of
// whichever fact write (or channel) last touched it (Vikunja #43). This is
// what makes pattern inference actually proactive: it fires on the daemon's
// own schedule reading accumulated history, not only as a side effect of a
// live voice turn.
//
// Idempotence and noise are handled by the store, not here — this function
// is safe to call every tick:
// - Same pattern, tick after tick: detectAndPropose's LookupProposedRoutine
// check plus proposed_routines' UNIQUE(action, object) constraint (with
// CreateProposedRoutine's ON CONFLICT DO NOTHING) mean a pair that
// already has a row — in ANY status — produces no second row and no log
// spam beyond the one line at genuine creation.
// - A DISMISSED proposal must never come back. DismissProposedRoutine flips
// status in place; the row is never deleted. So the same Lookup check
// that stops a duplicate "proposed" also stops a "dismissed" one from
// resurrecting — there is nothing tick-specific to get right here beyond
// calling the same shared path the voice route already used.
//
// By default this only creates a row for the /routines page to show: it does
// not notify, ring, or speak. Detection is not the same act as disturbing him
// about it, and Maven is "not a nag, not autonomous" (CLAUDE.md). Announcing
// is opt-in through the pattern_proposals config block — see announceProposal
// for the restraints that apply even then. A proposal only starts producing
// recurring nudges once he accepts it (fireAcceptedRoutines).
func (t *tickLoop) detectPatterns(ctx context.Context, now time.Time, state loop.State) {
pairs, err := t.store.DistinctEventPairs(ctx)
if err != nil {
log.Printf("tick: distinct event pairs: %v", err)
return
}
announced := false
for _, p := range pairs {
r, _, err := detectAndPropose(ctx, t.store, p.Action, p.Object, now)
if err != nil {
log.Printf("tick: detect pattern %s/%s: %v", p.Action, p.Object, err)
continue
}
if r == nil {
continue // no stable pattern, or already proposed/accepted/dismissed
}
log.Printf("tick: proposed routine: %s/%s every %.1f days", r.Action, r.Object, r.IntervalDays)
// One announcement per tick at most, whatever the scan turned up. The
// rest are on /routines; they are not lost, they are just not shouted.
// Nor are they queued: the row now exists, so no later tick re-detects
// them and they are never announced. See announceProposal.
if announced {
continue
}
announced = t.announceProposal(ctx, r, now, state)
}
}
// announceProposal offers a freshly inferred routine through the ordinary
// care-delivery path, if announcing is switched on at all. Returns true when
// something was actually sent.
//
// Everything here is restraint. The feature is off unless configured; when on
// it is sev1 (the lowest severity, so quiet hours, away presence and snooze
// all suppress it via loop.Gate exactly like a care nudge); it is spaced by
// proposalCfg.Cooldown across every pair, not per pair; and a suppressed or
// dropped announcement is NOT retried — the cooldown clock advances only on a
// real send, but the proposal row already exists, so the next tick will not
// re-detect it and nothing queues up behind it. A missed announcement means
// he reads it on /routines instead, which is the whole point of the page.
//
// What the cooldown is and is not. detectAndPropose returns non-nil only for a
// newly created row, so a pair gets exactly one chance to be spoken: the tick
// that first proposes it. Combined with one announcement per tick, the first
// tick over a populated history announces one pattern and permanently silences
// every other pattern found in the same pass. That is the intent, not an
// oversight — an inferred routine is not worth a second attempt at his
// attention, and /routines lists all of them. So the cooldown does not drain a
// backlog. It only spaces announcements of genuinely new pairs discovered on
// later ticks. If it should ever become "one per day until each is mentioned",
// that needs a queue rather than this counter.
//
// Cooldown gets its default here as well as in applyDefaults. That is
// deliberate: a tickLoop assembled directly in a test never goes through Load,
// and an unspaced announcer is not what those tests mean to exercise.
//
// The body is the detector's own literal Russian phrasing (pattern.PhraseRoutine
// — "ты заправляешь поилку раз в 7 дней — напоминать?"), not LLM-generated, so
// an inferred routine cannot arrive worded as something Maven never observed.
func (t *tickLoop) announceProposal(ctx context.Context, r *pattern.ProposedRoutine, now time.Time, state loop.State) bool {
if !t.proposalCfg.AnnounceProposals() {
return false
}
cooldown := time.Duration(t.proposalCfg.Cooldown)
if cooldown <= 0 {
cooldown = config.DefaultProposalCooldown
}
if !t.lastProposalAt.IsZero() && now.Sub(t.lastProposalAt) < cooldown {
return false
}
rule := loop.Rule{Name: "proposal:" + r.Action + " " + r.Object, Severity: loop.Sev1}
if !loop.Gate(state, rule) {
return false
}
body := pattern.PhraseRoutine(r)
pn := delivery.PhrasedNudge{
Candidate: loop.Candidate{Rule: rule, Severity: rule.Severity, State: state},
Body: body,
Summary: body,
}
sent, err := t.dispatcher.DispatchNudge(ctx, pn, now)
if err != nil {
log.Printf("tick: announce proposal %s/%s: %v", r.Action, r.Object, err)
return false
}
if len(sent) == 0 {
return false // routing dropped it — /routines still has it.
}
t.lastProposalAt = now
return true
}
// digestExpiry — how long a gate-suppressed care nudge stays worth
// resurfacing. 24h: these are daily-cadence rules (water/meal/break run on
// hour-scale cooldowns and re-derive from facts that reset every day), so a
// digest entry that outlives one full day is describing a day that's already
// over — "you skipped a break yesterday" said tomorrow evening is noise, not
// news. Bounding at one day also means a digest can never silently span a
// weekend of quiet hours into an unbounded backlog.
const digestExpiry = 24 * time.Hour
// maxDigestSpokenItems — the bundle read-out is capped so "batched, not
// dropped" cannot regress into "she dumps twelve things on me the moment I
// walk in" — a digest that nags in bulk is worse than the drops it replaced.
// Anything beyond the cap is still marked drained (it did get its moment;
// the cap limits WORDS, not whether it counted) and folded into a trailing
// count instead of being spoken in full.
const maxDigestSpokenItems = 3
// enqueueSuppressedDigest scans this tick's trace for care candidates the
// gate blocked for a genuine restraint reason and durably records the
// digest-eligible ones (loop.DigestEligible). Phrasing happens once, here,
// at enqueue time — not re-derived at drain time — the same way queueNudge
// phrases once and caches, so a rule suppressed for hours isn't re-prompting
// the LLM every tick it stays blocked (EnqueueDigestEntry's rule+body dedupe
// makes repeat calls here harmless, but skipping the phrase call entirely
// when a pending entry already exists avoids the LLM round-trip too).
func (t *tickLoop) enqueueSuppressedDigest(ctx context.Context, trace *loop.TickTrace, state loop.State, now time.Time) {
if trace == nil {
return
}
for _, tr := range trace.RuleTraces {
if !tr.PredicateResult || tr.GateResult {
continue // didn't want to fire, or wasn't suppressed
}
if !loop.DigestEligible(tr.Severity, tr.GateBlockedBy) {
continue
}
rule := loop.Rule{Name: tr.RuleName, Severity: tr.Severity}
cand := loop.Candidate{Rule: rule, Severity: tr.Severity, State: state}
pn, err := t.phraser.PhraseNudge(ctx, cand)
if err != nil {
log.Printf("tick: phrase digest candidate %s: %v", tr.RuleName, err)
continue
}
expires := now.Add(digestExpiry)
if _, deduped, err := t.store.EnqueueDigestEntry(ctx, tr.RuleName, int(tr.Severity), pn.Body, now, expires); err != nil {
log.Printf("tick: enqueue digest entry %s: %v", tr.RuleName, err)
} else if deduped {
// same suppressed nudge already pending — nothing new to say.
continue
}
}
}
// expireStaleDigest sweeps entries past their expiry once per tick — cheap
// bookkeeping, mirrors ReconcileStaleDeliveryAttempts's shape.
func (t *tickLoop) expireStaleDigest(ctx context.Context, now time.Time) {
n, err := t.store.ExpireStaleDigestEntries(ctx, now)
if err != nil {
log.Printf("tick: expire stale digest entries: %v", err)
return
}
if n > 0 {
log.Printf("tick: expired %d stale digest entr(y/ies) unspoken", n)
}
}
// maybeDrainDigest speaks the pending digest bundle once the gate's
// suppression reasons have actually cleared — quiet hours over, back from
// away, out of the meeting. Draining while still suppressed would just be a
// second way to nag through quiet hours; the bundle waits for the same "is
// it allowed right now" condition a live nudge already waits for.
func (t *tickLoop) maybeDrainDigest(ctx context.Context, state loop.State, now time.Time) {
if state.QuietHours || state.CalendarBusy || state.Presence == store.Away {
return
}
entries, err := t.store.PendingDigestEntries(ctx, now)
if err != nil {
log.Printf("tick: pending digest entries: %v", err)
return
}
if len(entries) == 0 {
return
}
spoken := entries
extra := 0
if len(spoken) > maxDigestSpokenItems {
spoken = entries[:maxDigestSpokenItems]
extra = len(entries) - maxDigestSpokenItems
}
var b strings.Builder
maxSev := 0
for i, e := range spoken {
if i > 0 {
b.WriteString(" · ")
}
b.WriteString(e.Body)
if e.Severity > maxSev {
maxSev = e.Severity
}
}
if extra > 0 {
fmt.Fprintf(&b, " · и ещё %d", extra)
}
body := b.String()
summary := fmt.Sprintf("%d отложенных уведомлений", len(entries))
cand := loop.Candidate{
Rule: loop.Rule{Name: "digest", Severity: loop.Severity(maxSev)},
Severity: loop.Severity(maxSev),
State: state,
}
pn := delivery.PhrasedNudge{Candidate: cand, Body: body, Summary: summary}
t.cachePhrase(pn)
if _, err := t.dispatcher.DispatchNudge(ctx, pn, now); err != nil {
log.Printf("tick: dispatch digest bundle: %v", err)
return // leave entries pending; retried next tick
}
ids := make([]int64, len(entries))
for i, e := range entries {
ids[i] = e.ID
}
if err := t.store.DrainDigestEntries(ctx, ids, now); err != nil {
log.Printf("tick: drain digest entries: %v", err)
}
}
// routinesFromConfig maps the config's routine blocks to the engine type.
// Validation (cron parses, name/body present, severity defaulted) already ran
// in config.Load, so this is a pure field copy.
func routinesFromConfig(rc []config.RoutineConfig) []routine.Routine {
if len(rc) == 0 {
return nil
}
out := make([]routine.Routine, len(rc))
for i, r := range rc {
out[i] = routine.Routine{Name: r.Name, Cron: r.Cron, Body: r.Body, Severity: r.Severity}
}
return out
}
// fireRoutines dispatches the routines whose cron schedule crossed since their
// last fire. Each is delivered as a nudge through the normal routing table
// (ChannelsFor(severity, presence)) with a "routine:"-prefixed rule name so it
// can't collide with a care rule in the feedback autotuner. A dispatch failure
// logs and continues — one bad send must not skip the rest, and routine.Due has
// already advanced the last-fire time so a transient failure drops that fire
// rather than replaying it every tick (a routine is clockwork, not an alarm —
// no repeat-til-ack).
func (t *tickLoop) fireRoutines(ctx context.Context, now time.Time, state loop.State) {
for _, r := range routine.Due(t.routines, t.routineLast, now) {
pn := delivery.PhrasedNudge{
Candidate: loop.Candidate{
Rule: loop.Rule{Name: "routine:" + r.Name, Severity: loop.Severity(r.Severity)},
Severity: loop.Severity(r.Severity),
State: state,
},
Body: r.Body,
Summary: r.Body,
}
if _, err := t.dispatcher.DispatchNudge(ctx, pn, now); err != nil {
log.Printf("tick: dispatch routine %s: %v", r.Name, err)
}
}
}
// fireAcceptedRoutines nudges about the routines the user accepted, once per
// interval (Vikunja #366). Accepting used to create a single reminder, so a
// non-weekly routine fired once and went quiet forever; the schedule lives in
// the proposed_routines row now and the loop re-reads it every tick.
//
// A routine is a care-class nudge and goes through the restraint gate like any
// other: quiet hours, away presence and snooze all suppress it. Reminders bypass
// that gate; routines must not. A suppressed nudge is NOT marked fired, so it
// goes out on the next tick that the gate allows — one nudge, held, not dropped
// and not repeated.
//
// The body is literal text built from the detected action and object, not
// LLM-phrased, so a routine can't hallucinate. It nudges; it never acts.
func (t *tickLoop) fireAcceptedRoutines(ctx context.Context, now time.Time, state loop.State) {
rows, err := t.store.ListAcceptedRoutines(ctx)
if err != nil {
log.Printf("tick: list accepted routines: %v", err)
return
}
accepted := make([]routine.Accepted, 0, len(rows))
for _, r := range rows {
if r.AcceptedTs == nil {
continue // accepted before the schedule column existed — no clock to start from.
}
accepted = append(accepted, routine.Accepted{
ID: r.ID,
Name: r.Action + " " + r.Object,
IntervalDays: r.IntervalDays,
Accepted: *r.AcceptedTs,
LastFired: r.LastFiredTs,
})
}
for _, a := range routine.DueAccepted(accepted, now) {
rule := loop.Rule{Name: "routine:" + a.Name, Severity: loop.Sev1}
if !loop.Gate(state, rule) {
continue
}
body := "пора: " + a.Name
pn := delivery.PhrasedNudge{
Candidate: loop.Candidate{Rule: rule, Severity: rule.Severity, State: state},
Body: body,
Summary: body,
}
sent, err := t.dispatcher.DispatchNudge(ctx, pn, now)
if err != nil {
log.Printf("tick: dispatch accepted routine %d: %v", a.ID, err)
continue
}
if len(sent) == 0 {
continue // routing dropped it — leave it due.
}
if err := t.store.MarkRoutineFired(ctx, a.ID, now); err != nil {
log.Printf("tick: mark routine %d fired: %v", a.ID, err)
}
}
}
// fireMorningRoutines checks each configured checklist against today's facts
// and dispatches a nag listing exactly what's still missing, at most once per
// routine per calendar day. Fact reads happen here (not in loop.Gatherer)
// because the item↔fact-key mapping is morning-routine-specific, not a rule
// concern — pulling it into the shared gather path would leak that mapping
// into loop's "rules declare wanted keys" contract. Bodies are literal
// operator text (item labels joined), not LLM-phrased, same rationale as
// cron routines: deterministic, can't hallucinate a checklist item.
func (t *tickLoop) fireMorningRoutines(ctx context.Context, now time.Time, state loop.State) {
if len(t.morningRoutines) == 0 {
return
}
facts := t.gatherMorningFacts(ctx)
for _, cand := range morning.Due(t.morningRoutines, facts, t.morningLast, now) {
labels := make([]string, len(cand.Missing))
for i, it := range cand.Missing {
labels[i] = it.Label
}
body := fmt.Sprintf("%s: не сделано — %s", cand.Routine.Name, strings.Join(labels, ", "))
pn := delivery.PhrasedNudge{
Candidate: loop.Candidate{
Rule: loop.Rule{Name: "morning:" + cand.Routine.Name, Severity: loop.Severity(cand.Routine.Severity)},
Severity: loop.Severity(cand.Routine.Severity),
State: state,
},
Body: body,
Summary: body,
}
if _, err := t.dispatcher.DispatchNudge(ctx, pn, now); err != nil {
log.Printf("tick: dispatch morning routine %s: %v", cand.Routine.Name, err)
}
}
}
// gatherMorningFacts reads the latest fact for every item's fact_key across
// all configured morning routines. Shared by fireMorningRoutines (nudge
// decision) and morningStatus (read-only query) so the two paths can never
// disagree about what evidence exists.
func (t *tickLoop) gatherMorningFacts(ctx context.Context) map[string]store.Fact {
keys := make(map[string]struct{})
for _, r := range t.morningRoutines {
for _, it := range r.Items {
keys[it.FactKey] = struct{}{}
}
}
facts := make(map[string]store.Fact, len(keys))
for k := range keys {
f, err := t.store.LatestFact(ctx, k)
if err == nil {
facts[k] = f
continue
}
if err != store.ErrNoFact {
log.Printf("tick: morning: latest fact %s: %v", k, err)
}
}
return facts
}
// morningStatus is the read-only "what's missing" query the web UI (and
// eventually a voice query) calls. Pure recompute over the current facts —
// no dedupe/nudge-time gating, unlike fireMorningRoutines: this answers
// "state right now," not "should we nag."
func (t *tickLoop) morningStatus(ctx context.Context, now time.Time) []ipc.MorningRoutineStatus {
if len(t.morningRoutines) == 0 {
return nil
}
facts := t.gatherMorningFacts(ctx)
out := make([]ipc.MorningRoutineStatus, 0, len(t.morningRoutines))
for _, r := range t.morningRoutines {
st := morning.Evaluate(r, facts, now)
done := make(map[string]bool, len(st.Completed))
for _, it := range st.Completed {
done[it.Key] = true
}
items := make([]ipc.MorningRoutineItem, len(r.Items))
for i, it := range r.Items {
items[i] = ipc.MorningRoutineItem{Key: it.Key, Label: it.Label, Done: done[it.Key]}
}
out = append(out, ipc.MorningRoutineStatus{
Name: r.Name,
Active: st.Active,
WindowStart: r.WindowStart,
WindowEnd: r.WindowEnd,
Items: items,
})
}
return out
}
// dayPlan is the read-only "what does today hold" query (Vikunja #128). It is
// the impure half of morning.BuildPlan: it reads the calendar events, the
// pending reminders and the checklist facts, and the pure builder orders them.
//
// It never dispatches. Asking for the plan is a query like any other; the only
// unprompted delivery in maven stays with the morning nudge and the
// dispatcher's policy.
func (t *tickLoop) dayPlan(ctx context.Context, now time.Time) ipc.DayPlan {
y, m, d := now.Date()
dayStart := time.Date(y, m, d, 0, 0, 0, 0, now.Location())
dayEnd := dayStart.AddDate(0, 0, 1)
var events []morning.PlanEntry
facts, err := t.store.CalendarEvents(ctx, dayStart, dayEnd)
if err != nil {
log.Printf("tick: day plan: calendar events: %v", err)
}
for _, f := range facts {
events = append(events, morning.PlanEntry{
At: f.Ts,
// The plan prints the hour itself, so the "@ 14:00-14:30" tail the
// fact value carries would say it twice.
Text: calendar.FactSummary(f.Value),
Kind: morning.PlanEvent,
// Provenance below a calendar read (an ambient relay, #126) is
// hedged rather than recited as fact.
Uncertain: f.Confidence < 1.0,
})
}
var reminders []morning.PlanEntry
rems, err := t.store.PendingReminders(ctx, dayStart, dayEnd)
if err != nil {
log.Printf("tick: day plan: pending reminders: %v", err)
}
for _, r := range rems {
if r.Status != store.ReminderPending {
continue
}
fire := r.NextFireTs
if fire.IsZero() {
fire = r.FireTs
}
reminders = append(reminders, morning.PlanEntry{
At: fire,
Text: strings.TrimSpace(r.Payload),
Kind: morning.PlanReminder,
})
}
var checklistFacts map[string]store.Fact
if len(t.morningRoutines) > 0 {
checklistFacts = t.gatherMorningFacts(ctx)
}
plan := morning.BuildPlan(t.morningRoutines, checklistFacts, events, reminders, now)
out := ipc.DayPlan{Date: plan.Date, Spoken: plan.FormatRU()}
out.Items = make([]ipc.DayPlanItem, len(plan.Items))
for i, it := range plan.Items {
out.Items[i] = ipc.DayPlanItem{
At: it.At,
Text: it.Text,
Kind: string(it.Kind),
Uncertain: it.Uncertain,
}
}
return out
}
// tune — the feedback auto-tuner's impure step. runs on a slow cadence
// (autotuneInterval, see run) so it doesn't write a fact every tick. for each
// rule:
@@ -948,101 +382,3 @@ func (t *tickLoop) trace() *loop.TickTrace {
defer t.mu.Unlock()
return t.lastTrace
}
// daemonAPI wraps a store-backed CoreAPI and overrides TickTrace with the
// daemon's in-memory tick trace cache.
type daemonAPI struct {
ipc.CoreAPI
getTrace func() *loop.TickTrace
getMorningStatus func(ctx context.Context) []ipc.MorningRoutineStatus
getDayPlan func(ctx context.Context) ipc.DayPlan
chatFn func(ctx context.Context, text string) string
getMCPServers func() []ipc.MCPServerStatus
getEvents func(n int) []ipc.IntakeEvent
}
// RecentEvents — the unified intake journal (Vikunja #283). Empty, not an
// error, when no bus was wired: "nothing has arrived" and "the journal is off"
// look the same to a reader on purpose, because neither is a fault and the
// page renders both as an empty table.
func (d *daemonAPI) RecentEvents(ctx context.Context, n int) ([]ipc.IntakeEvent, error) {
if d.getEvents == nil {
return nil, nil
}
return d.getEvents(n), nil
}
func (d *daemonAPI) Chat(ctx context.Context, text string) (string, error) {
if d.chatFn == nil {
return "", errors.New("mavend: chat not available")
}
return d.chatFn(ctx, text), nil
}
// MCPServers — the configured MCP servers and their health (Vikunja #251).
// Empty, not an error, when the mcp block is absent: "not configured" is the
// default state and the web surface renders it as such.
func (d *daemonAPI) MCPServers(ctx context.Context) ([]ipc.MCPServerStatus, error) {
if d.getMCPServers == nil {
return nil, nil
}
return d.getMCPServers(), nil
}
func (d *daemonAPI) TickTrace(ctx context.Context) (ipc.TickTrace, error) {
trace := d.getTrace()
if trace == nil {
return ipc.TickTrace{}, nil
}
return toIPCTickTrace(*trace), nil
}
func (d *daemonAPI) MorningStatus(ctx context.Context) ([]ipc.MorningRoutineStatus, error) {
if d.getMorningStatus == nil {
return nil, errors.New("mavend: morning status not available")
}
return d.getMorningStatus(ctx), nil
}
func (d *daemonAPI) DayPlan(ctx context.Context) (ipc.DayPlan, error) {
if d.getDayPlan == nil {
return ipc.DayPlan{}, errors.New("mavend: day plan not available")
}
return d.getDayPlan(ctx), nil
}
func toIPCTickTrace(t loop.TickTrace) ipc.TickTrace {
rules := make([]ipc.RuleTrace, len(t.RuleTraces))
for i, r := range t.RuleTraces {
rules[i] = toIPCRuleTrace(r)
}
return ipc.TickTrace{
Now: t.Now,
Winner: t.Winner,
Rules: rules,
}
}
func toIPCRuleTrace(r loop.RuleTrace) ipc.RuleTrace {
return ipc.RuleTrace{
RuleName: r.RuleName,
Severity: int(r.Severity),
PredicateResult: r.PredicateResult,
GateResult: r.GateResult,
GateBlockedBy: r.GateBlockedBy,
GateDetail: toIPCGateDetail(r.GateDetail),
WasSelected: r.WasSelected,
LostTo: r.LostTo,
}
}
func toIPCGateDetail(d loop.GateDetail) ipc.GateDetail {
return ipc.GateDetail{
SnoozeUntil: d.SnoozeUntil,
CooldownUntil: d.CooldownUntil,
QuietHours: d.QuietHours,
CalendarBusy: d.CalendarBusy,
Presence: d.Presence,
InertKeysMissing: d.InertKeysMissing,
}
}
+111
View File
@@ -0,0 +1,111 @@
// mavend/tick_api.go — the daemonAPI read surface over the tick loop.
//
// Split out of tick.go, move-only (Vikunja #422). What mavweb asks the daemon
// for, and the loop-to-ipc conversions those answers need.
package main
import (
"context"
"errors"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/loop"
)
// daemonAPI wraps a store-backed CoreAPI and overrides TickTrace with the
// daemon's in-memory tick trace cache.
type daemonAPI struct {
ipc.CoreAPI
getTrace func() *loop.TickTrace
getMorningStatus func(ctx context.Context) []ipc.MorningRoutineStatus
getDayPlan func(ctx context.Context) ipc.DayPlan
chatFn func(ctx context.Context, conversation, text string) string
getMCPServers func() []ipc.MCPServerStatus
getEvents func(n int) []ipc.IntakeEvent
}
// RecentEvents — the unified intake journal (Vikunja #283). Empty, not an
// error, when no bus was wired: "nothing has arrived" and "the journal is off"
// look the same to a reader on purpose, because neither is a fault and the
// page renders both as an empty table.
func (d *daemonAPI) RecentEvents(ctx context.Context, n int) ([]ipc.IntakeEvent, error) {
if d.getEvents == nil {
return nil, nil
}
return d.getEvents(n), nil
}
func (d *daemonAPI) Chat(ctx context.Context, conversation, text string) (string, error) {
if d.chatFn == nil {
return "", errors.New("mavend: chat not available")
}
return d.chatFn(ctx, conversation, text), nil
}
// MCPServers — the configured MCP servers and their health (Vikunja #251).
// Empty, not an error, when the mcp block is absent: "not configured" is the
// default state and the web surface renders it as such.
func (d *daemonAPI) MCPServers(ctx context.Context) ([]ipc.MCPServerStatus, error) {
if d.getMCPServers == nil {
return nil, nil
}
return d.getMCPServers(), nil
}
func (d *daemonAPI) TickTrace(ctx context.Context) (ipc.TickTrace, error) {
trace := d.getTrace()
if trace == nil {
return ipc.TickTrace{}, nil
}
return toIPCTickTrace(*trace), nil
}
func (d *daemonAPI) MorningStatus(ctx context.Context) ([]ipc.MorningRoutineStatus, error) {
if d.getMorningStatus == nil {
return nil, errors.New("mavend: morning status not available")
}
return d.getMorningStatus(ctx), nil
}
func (d *daemonAPI) DayPlan(ctx context.Context) (ipc.DayPlan, error) {
if d.getDayPlan == nil {
return ipc.DayPlan{}, errors.New("mavend: day plan not available")
}
return d.getDayPlan(ctx), nil
}
func toIPCTickTrace(t loop.TickTrace) ipc.TickTrace {
rules := make([]ipc.RuleTrace, len(t.RuleTraces))
for i, r := range t.RuleTraces {
rules[i] = toIPCRuleTrace(r)
}
return ipc.TickTrace{
Now: t.Now,
Winner: t.Winner,
Rules: rules,
}
}
func toIPCRuleTrace(r loop.RuleTrace) ipc.RuleTrace {
return ipc.RuleTrace{
RuleName: r.RuleName,
Severity: int(r.Severity),
PredicateResult: r.PredicateResult,
GateResult: r.GateResult,
GateBlockedBy: r.GateBlockedBy,
GateDetail: toIPCGateDetail(r.GateDetail),
WasSelected: r.WasSelected,
LostTo: r.LostTo,
}
}
func toIPCGateDetail(d loop.GateDetail) ipc.GateDetail {
return ipc.GateDetail{
SnoozeUntil: d.SnoozeUntil,
CooldownUntil: d.CooldownUntil,
QuietHours: d.QuietHours,
CalendarBusy: d.CalendarBusy,
Presence: d.Presence,
InertKeysMissing: d.InertKeysMissing,
}
}
+233
View File
@@ -0,0 +1,233 @@
// mavend/tick_digest.go — the digest queue.
//
// Split out of tick.go, move-only (Vikunja #422). A nudge below the severity
// ceiling is held here instead of spoken, flushed as one batch on the window,
// and drained by hand when he asks. Everything about batching lives here.
package main
import (
"context"
"fmt"
"log"
"strings"
"time"
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/store"
)
// shouldQueue — true when digest is enabled and the candidate's severity is
// at or below the configured ceiling.
func (t *tickLoop) shouldQueue(cand *loop.Candidate) bool {
return t.digestCfg != nil && t.digestCfg.Enabled &&
cand.Severity <= loop.Severity(t.digestCfg.SeverityCeiling)
}
// queueNudge — phrases the candidate and appends it to the digest queue.
// Deduplicates by rule name: if the same rule is already queued, this is a
// no-op (the first fire within the window is the one that counts).
func (t *tickLoop) queueNudge(ctx context.Context, cand *loop.Candidate, _ loop.State, now time.Time) {
for _, q := range t.digestQ {
if q.Rule == cand.Rule.Name {
return // already queued
}
}
pn, err := t.phraser.PhraseNudge(ctx, *cand)
if err != nil {
log.Printf("tick: phrase nudge %s: %v", cand.Rule.Name, err)
return
}
t.digestQ = append(t.digestQ, QueuedNudge{
Rule: cand.Rule.Name,
Severity: int(cand.Severity),
Body: pn.Body,
Key: cand.Rule.Name,
QueuedAt: now,
})
t.cachePhrase(pn)
}
// maybeFlush — flushes the digest queue if the window has elapsed since the
// first item or the queue reached MaxItems.
func (t *tickLoop) maybeFlush(ctx context.Context, now time.Time, state loop.State) {
if t.digestCfg == nil || !t.digestCfg.Enabled || len(t.digestQ) == 0 {
return
}
first := t.digestQ[0]
if now.Sub(first.QueuedAt) >= time.Duration(t.digestCfg.Window) ||
len(t.digestQ) >= t.digestCfg.MaxItems {
t.flushDigest(ctx, now, state)
}
}
// flushDigest — concatenates queued nudge bodies into a single digest
// notification and dispatches it. Clears the queue after a successful send.
// The digest uses the max severity among queued items for routing.
func (t *tickLoop) flushDigest(ctx context.Context, now time.Time, state loop.State) {
if len(t.digestQ) == 0 {
return
}
var b strings.Builder
maxSev := 0
for i, q := range t.digestQ {
if i > 0 {
b.WriteString(" · ")
}
b.WriteString(q.Body)
if q.Severity > maxSev {
maxSev = q.Severity
}
}
body := b.String()
summary := fmt.Sprintf("%d pending notifications", len(t.digestQ))
cand := loop.Candidate{
Rule: loop.Rule{
Name: "digest",
Severity: loop.Severity(maxSev),
},
Severity: loop.Severity(maxSev),
State: state,
}
pn := delivery.PhrasedNudge{
Candidate: cand,
Body: body,
Summary: summary,
}
t.cachePhrase(pn)
if _, err := t.dispatcher.DispatchNudge(ctx, pn, now); err != nil {
// keep the queue — the next tick's maybeFlush re-attempts.
log.Printf("tick: dispatch digest: %v", err)
return
}
t.digestQ = nil
}
// digestExpiry — how long a gate-suppressed care nudge stays worth
// resurfacing. 24h: these are daily-cadence rules (water/meal/break run on
// hour-scale cooldowns and re-derive from facts that reset every day), so a
// digest entry that outlives one full day is describing a day that's already
// over — "you skipped a break yesterday" said tomorrow evening is noise, not
// news. Bounding at one day also means a digest can never silently span a
// weekend of quiet hours into an unbounded backlog.
const digestExpiry = 24 * time.Hour
// maxDigestSpokenItems — the bundle read-out is capped so "batched, not
// dropped" cannot regress into "she dumps twelve things on me the moment I
// walk in" — a digest that nags in bulk is worse than the drops it replaced.
// Anything beyond the cap is still marked drained (it did get its moment;
// the cap limits WORDS, not whether it counted) and folded into a trailing
// count instead of being spoken in full.
const maxDigestSpokenItems = 3
// enqueueSuppressedDigest scans this tick's trace for care candidates the
// gate blocked for a genuine restraint reason and durably records the
// digest-eligible ones (loop.DigestEligible). Phrasing happens once, here,
// at enqueue time — not re-derived at drain time — the same way queueNudge
// phrases once and caches, so a rule suppressed for hours isn't re-prompting
// the LLM every tick it stays blocked (EnqueueDigestEntry's rule+body dedupe
// makes repeat calls here harmless, but skipping the phrase call entirely
// when a pending entry already exists avoids the LLM round-trip too).
func (t *tickLoop) enqueueSuppressedDigest(ctx context.Context, trace *loop.TickTrace, state loop.State, now time.Time) {
if trace == nil {
return
}
for _, tr := range trace.RuleTraces {
if !tr.PredicateResult || tr.GateResult {
continue // didn't want to fire, or wasn't suppressed
}
if !loop.DigestEligible(tr.Severity, tr.GateBlockedBy) {
continue
}
rule := loop.Rule{Name: tr.RuleName, Severity: tr.Severity}
cand := loop.Candidate{Rule: rule, Severity: tr.Severity, State: state}
pn, err := t.phraser.PhraseNudge(ctx, cand)
if err != nil {
log.Printf("tick: phrase digest candidate %s: %v", tr.RuleName, err)
continue
}
expires := now.Add(digestExpiry)
if _, deduped, err := t.store.EnqueueDigestEntry(ctx, tr.RuleName, int(tr.Severity), pn.Body, now, expires); err != nil {
log.Printf("tick: enqueue digest entry %s: %v", tr.RuleName, err)
} else if deduped {
// same suppressed nudge already pending — nothing new to say.
continue
}
}
}
// expireStaleDigest sweeps entries past their expiry once per tick — cheap
// bookkeeping, mirrors ReconcileStaleDeliveryAttempts's shape.
func (t *tickLoop) expireStaleDigest(ctx context.Context, now time.Time) {
n, err := t.store.ExpireStaleDigestEntries(ctx, now)
if err != nil {
log.Printf("tick: expire stale digest entries: %v", err)
return
}
if n > 0 {
log.Printf("tick: expired %d stale digest entr(y/ies) unspoken", n)
}
}
// maybeDrainDigest speaks the pending digest bundle once the gate's
// suppression reasons have actually cleared — quiet hours over, back from
// away, out of the meeting. Draining while still suppressed would just be a
// second way to nag through quiet hours; the bundle waits for the same "is
// it allowed right now" condition a live nudge already waits for.
func (t *tickLoop) maybeDrainDigest(ctx context.Context, state loop.State, now time.Time) {
if state.QuietHours || state.CalendarBusy || state.Presence == store.Away {
return
}
entries, err := t.store.PendingDigestEntries(ctx, now)
if err != nil {
log.Printf("tick: pending digest entries: %v", err)
return
}
if len(entries) == 0 {
return
}
spoken := entries
extra := 0
if len(spoken) > maxDigestSpokenItems {
spoken = entries[:maxDigestSpokenItems]
extra = len(entries) - maxDigestSpokenItems
}
var b strings.Builder
maxSev := 0
for i, e := range spoken {
if i > 0 {
b.WriteString(" · ")
}
b.WriteString(e.Body)
if e.Severity > maxSev {
maxSev = e.Severity
}
}
if extra > 0 {
fmt.Fprintf(&b, " · и ещё %d", extra)
}
body := b.String()
summary := fmt.Sprintf("%d отложенных уведомлений", len(entries))
cand := loop.Candidate{
Rule: loop.Rule{Name: "digest", Severity: loop.Severity(maxSev)},
Severity: loop.Severity(maxSev),
State: state,
}
pn := delivery.PhrasedNudge{Candidate: cand, Body: body, Summary: summary}
t.cachePhrase(pn)
if _, err := t.dispatcher.DispatchNudge(ctx, pn, now); err != nil {
log.Printf("tick: dispatch digest bundle: %v", err)
return // leave entries pending; retried next tick
}
ids := make([]int64, len(entries))
for i, e := range entries {
ids[i] = e.ID
}
if err := t.store.DrainDigestEntries(ctx, ids, now); err != nil {
log.Printf("tick: drain digest entries: %v", err)
}
}
+197
View File
@@ -0,0 +1,197 @@
// mavend/tick_morning.go — the morning checklist and the day plan.
//
// Split out of tick.go, move-only (Vikunja #422). One window per configured
// routine, nudging once at the end for what is still open, plus the read
// surfaces /morning renders.
package main
import (
"context"
"fmt"
"log"
"strings"
"time"
"github.com/kami/maven/internal/calendar"
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/morning"
"github.com/kami/maven/internal/store"
)
// fireMorningRoutines checks each configured checklist against today's facts
// and dispatches a nag listing exactly what's still missing, at most once per
// routine per calendar day. Fact reads happen here (not in loop.Gatherer)
// because the item↔fact-key mapping is morning-routine-specific, not a rule
// concern — pulling it into the shared gather path would leak that mapping
// into loop's "rules declare wanted keys" contract. Bodies are literal
// operator text (item labels joined), not LLM-phrased, same rationale as
// cron routines: deterministic, can't hallucinate a checklist item.
func (t *tickLoop) fireMorningRoutines(ctx context.Context, now time.Time, state loop.State) {
if len(t.morningRoutines) == 0 {
return
}
facts := t.gatherMorningFacts(ctx)
for _, cand := range morning.Due(t.morningRoutines, facts, t.morningLast, now) {
body := morningNudgeBody(cand)
pn := delivery.PhrasedNudge{
Candidate: loop.Candidate{
Rule: loop.Rule{Name: "morning:" + cand.Routine.Name, Severity: loop.Severity(cand.Routine.Severity)},
Severity: loop.Severity(cand.Routine.Severity),
State: state,
},
Body: body,
Summary: body,
}
if _, err := t.dispatcher.DispatchNudge(ctx, pn, now); err != nil {
log.Printf("tick: dispatch morning routine %s: %v", cand.Routine.Name, err)
}
}
}
// morningNudgeBody words the one message a routine gets per day. Required
// items are what she says was not done; optional ones follow, worded as
// something he could still do rather than something he owes (Vikunja #473).
// Operator text, not phrased by the model, for the same reason it always was:
// a checklist item must not be invented.
func morningNudgeBody(cand morning.Candidate) string {
labels := func(items []morning.Item) string {
out := make([]string, len(items))
for i, it := range items {
out[i] = it.Label
}
return strings.Join(out, ", ")
}
body := fmt.Sprintf("%s: не сделано — %s", cand.Routine.Name, labels(morning.Required(cand.Missing)))
if opt := morning.OptionalOnly(cand.Missing); len(opt) > 0 {
body += fmt.Sprintf(". если будет время — %s", labels(opt))
}
return body
}
// gatherMorningFacts reads the latest fact for every item's fact_key across
// all configured morning routines. Shared by fireMorningRoutines (nudge
// decision) and morningStatus (read-only query) so the two paths can never
// disagree about what evidence exists.
func (t *tickLoop) gatherMorningFacts(ctx context.Context) map[string]store.Fact {
keys := make(map[string]struct{})
for _, r := range t.morningRoutines {
for _, it := range r.Items {
keys[it.FactKey] = struct{}{}
}
}
facts := make(map[string]store.Fact, len(keys))
for k := range keys {
f, err := t.store.LatestFact(ctx, k)
if err == nil {
facts[k] = f
continue
}
if err != store.ErrNoFact {
log.Printf("tick: morning: latest fact %s: %v", k, err)
}
}
return facts
}
// morningStatus is the read-only "what's missing" query the web UI (and
// eventually a voice query) calls. Pure recompute over the current facts —
// no dedupe/nudge-time gating, unlike fireMorningRoutines: this answers
// "state right now," not "should we nag."
func (t *tickLoop) morningStatus(ctx context.Context, now time.Time) []ipc.MorningRoutineStatus {
if len(t.morningRoutines) == 0 {
return nil
}
facts := t.gatherMorningFacts(ctx)
out := make([]ipc.MorningRoutineStatus, 0, len(t.morningRoutines))
for _, r := range t.morningRoutines {
st := morning.Evaluate(r, facts, now)
done := make(map[string]bool, len(st.Completed))
for _, it := range st.Completed {
done[it.Key] = true
}
items := make([]ipc.MorningRoutineItem, len(r.Items))
for i, it := range r.Items {
items[i] = ipc.MorningRoutineItem{Key: it.Key, Label: it.Label, Done: done[it.Key]}
}
out = append(out, ipc.MorningRoutineStatus{
Name: r.Name,
Active: st.Active,
WindowStart: r.WindowStart,
WindowEnd: r.WindowEnd,
Items: items,
})
}
return out
}
// dayPlan is the read-only "what does today hold" query (Vikunja #128). It is
// the impure half of morning.BuildPlan: it reads the calendar events, the
// pending reminders and the checklist facts, and the pure builder orders them.
//
// It never dispatches. Asking for the plan is a query like any other; the only
// unprompted delivery in maven stays with the morning nudge and the
// dispatcher's policy.
func (t *tickLoop) dayPlan(ctx context.Context, now time.Time) ipc.DayPlan {
y, m, d := now.Date()
dayStart := time.Date(y, m, d, 0, 0, 0, 0, now.Location())
dayEnd := dayStart.AddDate(0, 0, 1)
var events []morning.PlanEntry
facts, err := t.store.CalendarEvents(ctx, dayStart, dayEnd)
if err != nil {
log.Printf("tick: day plan: calendar events: %v", err)
}
for _, f := range facts {
events = append(events, morning.PlanEntry{
At: f.Ts,
// The plan prints the hour itself, so the "@ 14:00-14:30" tail the
// fact value carries would say it twice.
Text: calendar.FactSummary(f.Value),
Kind: morning.PlanEvent,
// Provenance below a calendar read (an ambient relay, #126) is
// hedged rather than recited as fact.
Uncertain: f.Confidence < 1.0,
})
}
var reminders []morning.PlanEntry
rems, err := t.store.PendingReminders(ctx, dayStart, dayEnd)
if err != nil {
log.Printf("tick: day plan: pending reminders: %v", err)
}
for _, r := range rems {
if r.Status != store.ReminderPending {
continue
}
fire := r.NextFireTs
if fire.IsZero() {
fire = r.FireTs
}
reminders = append(reminders, morning.PlanEntry{
At: fire,
Text: r.Text(),
Kind: morning.PlanReminder,
})
}
var checklistFacts map[string]store.Fact
if len(t.morningRoutines) > 0 {
checklistFacts = t.gatherMorningFacts(ctx)
}
plan := morning.BuildPlan(t.morningRoutines, checklistFacts, events, reminders, now)
out := ipc.DayPlan{Date: plan.Date, Spoken: plan.FormatRU()}
out.Items = make([]ipc.DayPlanItem, len(plan.Items))
for i, it := range plan.Items {
out.Items[i] = ipc.DayPlanItem{
At: it.At,
Text: it.Text,
Kind: string(it.Kind),
Uncertain: it.Uncertain,
}
}
return out
}
+234
View File
@@ -0,0 +1,234 @@
// mavend/tick_routines.go — pattern detection and the routines that fire.
//
// Split out of tick.go, move-only (Vikunja #422). Configured routines, the
// routines he accepted on /routines, and the tick-side detector that proposes
// new ones.
package main
import (
"context"
"log"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/pattern"
"github.com/kami/maven/internal/routine"
)
// detectPatterns runs the pattern detector proactively over every
// action+object pair that has ever produced an event, independent of
// whichever fact write (or channel) last touched it (Vikunja #43). This is
// what makes pattern inference actually proactive: it fires on the daemon's
// own schedule reading accumulated history, not only as a side effect of a
// live voice turn.
//
// Idempotence and noise are handled by the store, not here — this function
// is safe to call every tick:
// - Same pattern, tick after tick: detectAndPropose's LookupProposedRoutine
// check plus proposed_routines' UNIQUE(action, object) constraint (with
// CreateProposedRoutine's ON CONFLICT DO NOTHING) mean a pair that
// already has a row — in ANY status — produces no second row and no log
// spam beyond the one line at genuine creation.
// - A DISMISSED proposal must never come back. DismissProposedRoutine flips
// status in place; the row is never deleted. So the same Lookup check
// that stops a duplicate "proposed" also stops a "dismissed" one from
// resurrecting — there is nothing tick-specific to get right here beyond
// calling the same shared path the voice route already used.
//
// By default this only creates a row for the /routines page to show: it does
// not notify, ring, or speak. Detection is not the same act as disturbing him
// about it, and Maven is "not a nag, not autonomous" (CLAUDE.md). Announcing
// is opt-in through the pattern_proposals config block — see announceProposal
// for the restraints that apply even then. A proposal only starts producing
// recurring nudges once he accepts it (fireAcceptedRoutines).
func (t *tickLoop) detectPatterns(ctx context.Context, now time.Time, state loop.State) {
pairs, err := t.store.DistinctEventPairs(ctx)
if err != nil {
log.Printf("tick: distinct event pairs: %v", err)
return
}
announced := false
for _, p := range pairs {
r, _, err := detectAndPropose(ctx, t.store, p.Action, p.Object, now)
if err != nil {
log.Printf("tick: detect pattern %s/%s: %v", p.Action, p.Object, err)
continue
}
if r == nil {
continue // no stable pattern, or already proposed/accepted/dismissed
}
log.Printf("tick: proposed routine: %s/%s every %.1f days", r.Action, r.Object, r.IntervalDays)
// One announcement per tick at most, whatever the scan turned up. The
// rest are on /routines; they are not lost, they are just not shouted.
// Nor are they queued: the row now exists, so no later tick re-detects
// them and they are never announced. See announceProposal.
if announced {
continue
}
announced = t.announceProposal(ctx, r, now, state)
}
}
// announceProposal offers a freshly inferred routine through the ordinary
// care-delivery path, if announcing is switched on at all. Returns true when
// something was actually sent.
//
// Everything here is restraint. The feature is off unless configured; when on
// it is sev1 (the lowest severity, so quiet hours, away presence and snooze
// all suppress it via loop.Gate exactly like a care nudge); it is spaced by
// proposalCfg.Cooldown across every pair, not per pair; and a suppressed or
// dropped announcement is NOT retried — the cooldown clock advances only on a
// real send, but the proposal row already exists, so the next tick will not
// re-detect it and nothing queues up behind it. A missed announcement means
// he reads it on /routines instead, which is the whole point of the page.
//
// What the cooldown is and is not. detectAndPropose returns non-nil only for a
// newly created row, so a pair gets exactly one chance to be spoken: the tick
// that first proposes it. Combined with one announcement per tick, the first
// tick over a populated history announces one pattern and permanently silences
// every other pattern found in the same pass. That is the intent, not an
// oversight — an inferred routine is not worth a second attempt at his
// attention, and /routines lists all of them. So the cooldown does not drain a
// backlog. It only spaces announcements of genuinely new pairs discovered on
// later ticks. If it should ever become "one per day until each is mentioned",
// that needs a queue rather than this counter.
//
// Cooldown gets its default here as well as in applyDefaults. That is
// deliberate: a tickLoop assembled directly in a test never goes through Load,
// and an unspaced announcer is not what those tests mean to exercise.
//
// The body is the detector's own literal Russian phrasing (pattern.PhraseRoutine
// — "ты заправляешь поилку раз в 7 дней — напоминать?"), not LLM-generated, so
// an inferred routine cannot arrive worded as something Maven never observed.
func (t *tickLoop) announceProposal(ctx context.Context, r *pattern.ProposedRoutine, now time.Time, state loop.State) bool {
if !t.proposalCfg.AnnounceProposals() {
return false
}
cooldown := time.Duration(t.proposalCfg.Cooldown)
if cooldown <= 0 {
cooldown = config.DefaultProposalCooldown
}
if !t.lastProposalAt.IsZero() && now.Sub(t.lastProposalAt) < cooldown {
return false
}
rule := loop.Rule{Name: "proposal:" + r.Action + " " + r.Object, Severity: loop.Sev1}
if !loop.Gate(state, rule) {
return false
}
body := pattern.PhraseRoutine(r)
pn := delivery.PhrasedNudge{
Candidate: loop.Candidate{Rule: rule, Severity: rule.Severity, State: state},
Body: body,
Summary: body,
}
sent, err := t.dispatcher.DispatchNudge(ctx, pn, now)
if err != nil {
log.Printf("tick: announce proposal %s/%s: %v", r.Action, r.Object, err)
return false
}
if len(sent) == 0 {
return false // routing dropped it — /routines still has it.
}
t.lastProposalAt = now
return true
}
// routinesFromConfig maps the config's routine blocks to the engine type.
// Validation (cron parses, name/body present, severity defaulted) already ran
// in config.Load, so this is a pure field copy.
func routinesFromConfig(rc []config.RoutineConfig) []routine.Routine {
if len(rc) == 0 {
return nil
}
out := make([]routine.Routine, len(rc))
for i, r := range rc {
out[i] = routine.Routine{Name: r.Name, Cron: r.Cron, Body: r.Body, Severity: r.Severity}
}
return out
}
// fireRoutines dispatches the routines whose cron schedule crossed since their
// last fire. Each is delivered as a nudge through the normal routing table
// (ChannelsFor(severity, presence)) with a "routine:"-prefixed rule name so it
// can't collide with a care rule in the feedback autotuner. A dispatch failure
// logs and continues — one bad send must not skip the rest, and routine.Due has
// already advanced the last-fire time so a transient failure drops that fire
// rather than replaying it every tick (a routine is clockwork, not an alarm —
// no repeat-til-ack).
func (t *tickLoop) fireRoutines(ctx context.Context, now time.Time, state loop.State) {
for _, r := range routine.Due(t.routines, t.routineLast, now) {
pn := delivery.PhrasedNudge{
Candidate: loop.Candidate{
Rule: loop.Rule{Name: "routine:" + r.Name, Severity: loop.Severity(r.Severity)},
Severity: loop.Severity(r.Severity),
State: state,
},
Body: r.Body,
Summary: r.Body,
}
if _, err := t.dispatcher.DispatchNudge(ctx, pn, now); err != nil {
log.Printf("tick: dispatch routine %s: %v", r.Name, err)
}
}
}
// fireAcceptedRoutines nudges about the routines the user accepted, once per
// interval (Vikunja #366). Accepting used to create a single reminder, so a
// non-weekly routine fired once and went quiet forever; the schedule lives in
// the proposed_routines row now and the loop re-reads it every tick.
//
// A routine is a care-class nudge and goes through the restraint gate like any
// other: quiet hours, away presence and snooze all suppress it. Reminders bypass
// that gate; routines must not. A suppressed nudge is NOT marked fired, so it
// goes out on the next tick that the gate allows — one nudge, held, not dropped
// and not repeated.
//
// The body is literal text built from the detected action and object, not
// LLM-phrased, so a routine can't hallucinate. It nudges; it never acts.
func (t *tickLoop) fireAcceptedRoutines(ctx context.Context, now time.Time, state loop.State) {
rows, err := t.store.ListAcceptedRoutines(ctx)
if err != nil {
log.Printf("tick: list accepted routines: %v", err)
return
}
accepted := make([]routine.Accepted, 0, len(rows))
for _, r := range rows {
if r.AcceptedTs == nil {
continue // accepted before the schedule column existed — no clock to start from.
}
accepted = append(accepted, routine.Accepted{
ID: r.ID,
Name: r.Action + " " + r.Object,
IntervalDays: r.IntervalDays,
Accepted: *r.AcceptedTs,
LastFired: r.LastFiredTs,
})
}
for _, a := range routine.DueAccepted(accepted, now) {
rule := loop.Rule{Name: "routine:" + a.Name, Severity: loop.Sev1}
if !loop.Gate(state, rule) {
continue
}
body := "пора: " + a.Name
pn := delivery.PhrasedNudge{
Candidate: loop.Candidate{Rule: rule, Severity: rule.Severity, State: state},
Body: body,
Summary: body,
}
sent, err := t.dispatcher.DispatchNudge(ctx, pn, now)
if err != nil {
log.Printf("tick: dispatch accepted routine %d: %v", a.ID, err)
continue
}
if len(sent) == 0 {
continue // routing dropped it — leave it due.
}
if err := t.store.MarkRoutineFired(ctx, a.ID, now); err != nil {
log.Printf("tick: mark routine %d fired: %v", a.ID, err)
}
}
}
+41 -1
View File
@@ -584,7 +584,7 @@ func TestDigestSev4BypassesQueue(t *testing.T) {
ctx := context.Background()
now := refNow()
markPresent(t, st, ctx, now)
if _, err := st.SetValue(ctx, store.KindSelf, "service_down", "poll:uptimekuma", "down", now); err != nil {
if _, err := st.SetValue(ctx, store.KindSelf, "service_down:db", "poll:uptimekuma", "down", now); err != nil {
t.Fatalf("seed service_down: %v", err)
}
sink := &fakeSink{}
@@ -667,3 +667,43 @@ func TestDigestDeduplicatesByRule(t *testing.T) {
t.Fatalf("after duplicate queue attempt: digestQ = %d, want 1 (dedup)", len(tl.digestQ))
}
}
// The repeat path reads the nudges table, not the rule set, so a rule turned
// off in `disabled_rules` used to keep re-sending its last un-acked telegram
// nudge every repeat_interval. Two arrived after the rule was off on
// 2026-08-01. A disabled rule must be unreachable on every path.
func TestRepeatableRulesDropsDisabledRules(t *testing.T) {
tl := &tickLoop{rules: mustRules(t, []string{"service_down"})}
got := tl.repeatableRules([]string{"service_down", "water"})
if len(got) != 1 || got[0] != "water" {
t.Fatalf("repeatableRules = %v, want [water]", got)
}
}
// An orphan row for a rule that no longer exists in the code goes quiet too:
// nothing can ack what the UI cannot show.
func TestRepeatableRulesDropsUnknownRules(t *testing.T) {
tl := &tickLoop{rules: loop.DefaultRules()}
if got := tl.repeatableRules([]string{"rule_deleted_last_year"}); len(got) != 0 {
t.Fatalf("repeatableRules = %v, want none", got)
}
}
func TestRepeatableRulesKeepsWiredRules(t *testing.T) {
tl := &tickLoop{rules: loop.DefaultRules()}
got := tl.repeatableRules([]string{"service_down", "water"})
if len(got) != 2 {
t.Fatalf("repeatableRules = %v, want both", got)
}
}
// mustRules returns DefaultRules minus the named ones, failing if a name
// matched nothing — a typo here would make the test pass for the wrong reason.
func mustRules(t *testing.T, disabled []string) []loop.Rule {
t.Helper()
rules, dropped := loop.RulesExcept(disabled)
if len(dropped) != len(disabled) {
t.Fatalf("dropped %v, want %v", dropped, disabled)
}
return rules
}
+116 -17
View File
@@ -76,17 +76,34 @@ type reactiveHandler struct {
tts tts.Synthesizer
router *router.Router
embedder router.Embedder // reused for note write/query (same model as the classifier)
api ipc.CoreAPI
tools *tool.Executor
matcher *tool.Matcher
phraser phraser.Phraser
replier voice.Replier
now func() time.Time
// boundary — the embedded seed sets behind the personal boundary
// (personalboundary.go). Zero value is usable and loads on first query;
// with no embedder it never loads and the boundary uses personalMarkers.
boundary personalBoundary
// api — the CoreAPI the handler reads and writes through. Wired with the
// bare store adapter and UPGRADED by main once the daemonAPI exists; see
// upgradeAPI.
api ipc.CoreAPI
tools *tool.Executor
matcher *tool.Matcher
phraser phraser.Phraser
replier voice.Replier
now func() time.Time
// crawler reads a web page he names out loud (queryWeb). nil ⇒ on-demand
// page reading is off, which is the default: no `crawl` block, no fetch.
crawler *crawl.Crawler
// search asks a self-hosted SearXNG (querySearch), the first world source
// once his own data has had its turn. nil ⇒ off, the default: no `search`
// block, no query ever leaves the LAN.
search *searchWiring
// kiwix searches the offline ZIMs (queryKiwix), the fallback behind the
// live search and the last source before the model answers from its own
// weights. nil ⇒ off, the default.
kiwix *kiwixWiring
// feedsOn — whether any RSS feed is configured (config.Feeds). It changes
// only what she SAYS when asked and nothing is there: "ленты не настроены"
// instead of "ничего нового", which are different truths.
@@ -179,12 +196,33 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
return h.reply(ctx, replyText, nil)
}
// upgradeAPI points the handler at the daemon's own CoreAPI once main has
// built it.
//
// Wiring order forces this. wireVoice runs before the tick loop exists, so it
// can only be handed the bare store adapter — and that adapter answers DayPlan
// (and TickTrace, and MorningStatus) with "not available via direct store
// API", because a day plan is assembled by the tick loop and is not a table to
// read. So queryDayPlan, which the query chain reaches for "какие у меня планы
// на сегодня", failed on the deployed daemon for every caller. main already
// back-patches the other direction (daemonAPI.chatFn = handler.handleText);
// this is the same seam in reverse.
//
// Safe against the obvious loop: nothing in the voice path calls api.Chat, so
// pointing the handler at an API whose Chat IS the handler cannot recurse.
func (h *reactiveHandler) upgradeAPI(api ipc.CoreAPI) {
if h == nil || api == nil {
return
}
h.api = api
}
// handleText — the core reactive path without stt/tts. Used by the IPC Chat
// endpoint (and eventually by telegram). Splits out the audio bookends from
// HandlePushToTalk so text channels share the same routing logic.
func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
func (h *reactiveHandler) handleText(ctx context.Context, conversation, text string) string {
log.Printf("voice: handleText: %q", text)
return h.runTurn(ctx, text, sourceText)
return h.runTurn(withDialogueID(ctx, dialogueIDFor(sourceText, conversation)), text, sourceText)
}
// turnSource — which channel this utterance arrived on, in the same provenance
@@ -217,7 +255,7 @@ func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSour
// early. He can be asked a question, walk off, come back and say "да" to a
// confirm that is still parked; computing the notice after that return meant
// he answered the confirm and never heard that the older request was let go.
expiredNotice := h.clarifyExpiredNotice()
expiredNotice := h.clarifyExpiredNotice(ctx)
// 2. confirm turn — if a destructive act is parked, this utterance is its
// y/n answer, not a fresh command. Handled before routing so "да" doesn't
@@ -246,8 +284,42 @@ func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSour
return withNotice(expiredNotice, reply)
}
// 5. router — classify the utterance.
dec, err := h.router.Route(ctx, text, h.now())
// 4b. spoken snooze — "не сейчас" / "потом" answers the nudge she just
// sent. Only handled when a pending nudge is actually inside the window
// (snooze.go); otherwise the words route normally, because "потом" is an
// ordinary word and eating every one of them would break real sentences.
if reply, handled := h.resolveSnooze(ctx, text, src); handled {
return withNotice(expiredNotice, reply)
}
// 4c. spoken ack — "готово" closes that same nudge as `acted`. Only the
// contentless form is intercepted here; "выпил воды" keeps routing and
// closes the nudge after its fact lands (ackFromFact, step 8b).
if reply, handled := h.resolveAck(ctx, text, src); handled {
return withNotice(expiredNotice, reply)
}
// 5. route. An elliptical follow-up — "а завтра?" — is answered from the
// previous turn instead (continuation.go): the intent is the part it is
// missing, so no amount of routing recovers it, and the model's guess
// costs seconds to obtain and is close to a coin flip. Everything else
// goes to the router.
var (
dec router.Decision
err error
prev *dialogue.Session
)
now := h.now()
if h.dialogueSessions != nil {
prev = h.dialogueSessions.Get(voiceDialogueID, now)
}
cont := false
if dec, cont = continuationDecision(prev, text, now); cont {
log.Printf("voice: continuation of %s from the previous turn", dec.Intent)
}
if !cont {
dec, err = h.router.Route(ctx, text, now)
}
if err != nil {
// ErrNoIntents ⇒ classifier unseeded (cold boot). reply with a
// "still warming up" rather than a wire error.
@@ -263,10 +335,13 @@ func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSour
// turn (follow-ups like «напомни завтра» → «…позвонить маме»), then remember
// this turn for the next follow-up. Only same-intent, non-expired, non-
// clarify turns carry (see followUpMerge). Best-effort: nil store ⇒ skipped.
// A continuation already carries the previous turn's slots, so there is
// nothing left to inherit — but it is still remembered, so a chain of them
// ("а завтра?" … "а послезавтра?") keeps working.
if h.dialogueSessions != nil {
now := h.now()
prev := h.dialogueSessions.Get(voiceDialogueID, now)
dec = followUpMerge(prev, dec, now)
if !cont {
dec = followUpMerge(prev, dec, now)
}
if !dec.Clarify {
h.rememberTurn(prev, dec, now)
}
@@ -276,7 +351,10 @@ func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSour
// and park the request (clarify.go); otherwise the replier's canned reply
// stands.
if dec.Clarify {
if question, asked := h.askClarify(dec); asked {
if reply := h.hexisBeforeClarify(ctx, dec); reply != "" {
return withNotice(expiredNotice, reply)
}
if question, asked := h.askClarify(ctx, dec); asked {
return withNotice(expiredNotice, question)
}
}
@@ -287,6 +365,10 @@ func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSour
replyText := h.applyAction(ctx, dec)
log.Printf("voice: applyAction returned: %q", replyText)
// 8b. a fact that answers a live nudge closes it as `acted` (ack.go).
// Silent: the fact reply stands, she does not congratulate him for it.
h.ackFromFact(ctx, dec)
// 9. replier — phrase the reply across the router decision.
if replyText == "" {
replyText = h.replier.Reply(dec)
@@ -324,6 +406,23 @@ func (h *reactiveHandler) replySystem(ctx context.Context, dec router.Decision)
u := strings.ToLower(dec.Utterance)
now := h.now()
// The topic and the day come from different places on a continuation.
// "а завтра?" names the day and nothing else; what he is asking ABOUT
// lives in the previous turn, which continuation.go copied into
// Slots.Text. Dates keep parsing from the utterance — that is the part
// the ellipsis actually restates — and only the keyword match widens.
//
// Gated on Continued, and that gate is load-bearing. followUpMerge fills
// an empty Text from the previous same-intent turn, so without it a plain
// "привет" after "какой сегодня день" inherited the old topic and got
// answered with the date. Seen on the deployed daemon, 01-08-2026.
topic := u
if dec.Continued {
if t := strings.ToLower(dec.Slots.Text); t != "" && t != u {
topic = u + " " + t
}
}
// stage-0 grammars catch the exact time/date patterns, but duration
// queries ("сколько времени прошло") bypass the grammar's build filter
// and can reach replySystem via the classifier path. Guard against them.
@@ -332,14 +431,14 @@ func (h *reactiveHandler) replySystem(ctx context.Context, dec router.Decision)
}
switch {
case strings.Contains(u, "час") || strings.Contains(u, "врем"):
case strings.Contains(topic, "час") || strings.Contains(topic, "врем"):
// "который час в киеве" — she keeps one clock, so any named place gets
// the honest answer. Never local time dressed up as the city's.
if mentionsUnknownPlace(u) {
return onlyLocalTimeReply
}
return "сейчас " + ruClock(now)
case strings.Contains(u, "день") || strings.Contains(u, "числ"):
case strings.Contains(topic, "день") || strings.Contains(topic, "числ"):
// "какое число завтра" — answer for the day the user asked about,
// not today. Reuses the router's calendar day-word parser.
day := now
+138 -12
View File
@@ -48,7 +48,11 @@ type voiceWiring struct {
// mcp — the MCP client, nil unless the `mcp` block configures an enabled
// server (Vikunja #251). Its tools land in the same allowlist as every
// other act, so nothing else here has to know about it.
mcp *mcpWiring
// pair — the workstation model with the resident one as the floor, nil
// unless a `workstation` block names an address. Held here only so the
// prober is stopped on shutdown; callers were handed it at build time.
pair *llm.Pair
mcp *mcpWiring
// home — the Home Assistant client, nil unless the `smarthome` block is
// enabled (Vikunja #256). Its devices land in the same allowlist as every
// other act, so nothing else here has to know about it.
@@ -76,6 +80,9 @@ func (w *voiceWiring) close() {
if w.ttsClient != nil {
_ = w.ttsClient.Close()
}
if w.pair != nil {
w.pair.Stop()
}
w.mcp.close()
}
@@ -139,6 +146,7 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
emb = router.NewHashEmbedder(1024)
}
w.embedder = emb
repairFactVectors(dataStore, emb)
checkStoredEmbedder(dataStore, emb)
// ----- tool executor (the enabled act allowlist, store-backed) -----
@@ -188,6 +196,18 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
// the new llama-server when the resident model is swapped (Vikunja #250).
llmClient = llmClientFor(lp, 60*time.Second)
}
// The workstation model sits above that one when it is configured and its
// card is free. hot is what the router and the replier complete through:
// either the pair, or the resident client alone, or nothing at all.
hot, pair := modelSeam(cfg, llmClient)
w.pair = pair
// The phraser gets the same pair, which is what carries the workstation model
// into the paths that do not go through `hot`: world questions (the naming
// half), and the digestion worker's nudge and reminder phrasing (the silent
// half). Wiring, so it happens once and before the voice server listens.
if lp, ok := phr.(*phraser.LLMPhraser); ok && pair != nil {
lp.UseRemote(pair)
}
// ----- router (the cascade; floor examples seed the classifier) -----
// The act matcher's allowlist is exactly the enabled tool names — the
// router only matches acts the executor can run (one source of truth).
@@ -199,7 +219,7 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
// against the classifier's 50.0%, at about 1s a turn instead of 30ms (see
// config.VoiceConfig.LLMRouter). The classifier always stays wired as the
// fallback, so a model error never breaks a turn.
rtr := buildRouter(emb, matcher, threshold, pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient))
rtr := buildRouter(emb, matcher, threshold, pickLLMRouter(cfg.Voice.UseLLMRouter(), hot))
// ----- sessions registry (shared with voicesink) -----
sessions := voice.NewSessions()
@@ -218,7 +238,10 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
// ----- dialogue (multi-turn slot carry-over; 2-min follow-up window) -----
// Store-backed when the daemon passes a store, so a restart mid-conversation
// keeps the thread (Vikunja #363). Sessions past their TTL are dropped on
// load, never revived. Clarify's parked question stays in memory only.
// load, never revived. Clarify's parked question stays in memory only, and
// that is a decision rather than an omission (Vikunja #385, docs/design.md):
// a restart expires it, so the thread comes back and the open question does
// not.
var dialogueSessions *dialogue.SessionStore
if dataStore != nil {
dialogueSessions = dialogue.NewPersistentSessionStore(2*time.Minute, dataStore)
@@ -233,8 +256,8 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
// ----- replier (LLM-backed when the engine is on, Stub floor otherwise) -----
replier := voice.Replier(voice.NewStubReplier())
if llmClient != nil {
replier = newLLMReplier(llmClient, contextBlockFn(cfg, time.Now))
if hot != nil {
replier = newLLMReplier(hot, contextBlockFn(cfg, time.Now))
}
// ----- the handler (the reactive path; closes over stt / tts / router / coreAPI / memory) -----
@@ -254,7 +277,13 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
netscan: w.netscan,
// nil unless `crawl.on_demand` is on: reading a page he names is a
// capability, and capabilities are off unless configured.
crawler: onDemandCrawler(cfg),
crawler: onDemandCrawler(cfg),
// nil unless a `search` block names a SearXNG instance. External search
// is off unless configured, and configuring it is the whole opt-in.
search: wireSearch(cfg),
// nil unless a `kiwix` block names a server. Same swap-aware client the
// router and replier use, so the rewriter follows a model swap.
kiwix: wireKiwix(cfg, llmClient),
weatherProvider: weatherProvider,
weatherLocation: weatherLocation,
memStore: memStore,
@@ -285,7 +314,42 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
// pickLLMRouter returns the LLM router when the operator asked for it and there
// is a llama-server to talk to, and nil otherwise. nil is safe: the cascade then
// routes with the classifier, so an unusable setting costs accuracy, not turns.
func pickLLMRouter(enabled bool, c *llm.Client) *router.LLMRouter {
// modelSeam builds the completion seam the hot paths use: routing and replies.
//
// With no `workstation` block it is the resident client and nothing probes
// anything, which is today's deploy exactly. With one, it is an llm.Pair that
// prefers the workstation and falls back to the resident model silently — the
// silent half of the degradation rule (docs/offload.md), because the big model
// is only better here and the 1.7B is today's shipping quality. He is never
// told which of the two phrased his reply.
//
// A nil resident client means the phraser is not an LLM phraser. There is then
// no floor, and a Pair with no floor is a configuration mistake rather than a
// degraded mode, so the seam is nil and the cascade routes with the classifier.
func modelSeam(cfg *config.Config, resident *llm.Client) (router.Completer, *llm.Pair) {
if resident == nil {
if cfg.Workstation != nil {
log.Printf("voice: a workstation is configured but there is no resident model to floor it with — ignoring the block")
}
return nil, nil
}
if cfg.Workstation == nil {
return resident, nil
}
ws := cfg.Workstation
pair := llm.NewPair(
llm.New(ws.URL, time.Duration(ws.Timeout)),
resident,
ws.Health,
time.Duration(ws.Probe),
)
pair.Start(context.Background())
log.Printf("voice: workstation model at %s, probed every %s, resident model as the floor",
ws.URL, time.Duration(ws.Probe))
return pair, pair
}
func pickLLMRouter(enabled bool, c router.Completer) *router.LLMRouter {
if !enabled {
return nil
}
@@ -314,7 +378,16 @@ func buildRouter(emb router.Embedder, acts router.ActMatcher, threshold float64,
seedClassifier(cls)
grammars := router.DefaultGrammars(acts)
grammars = append(grammars, router.SystemTimeDateGrammars()...)
// After the time/date rules on purpose: "какой сегодня день" is a clock
// question and must keep reaching replySystem, while "что у меня сегодня"
// is an agenda question and must not.
grammars = append(grammars, router.AgendaQueryGrammars()...)
grammars = append(grammars, router.ReminderGrammar())
// Last, and it matches any utterance shape — its Build is the filter. An
// explicit capture marker beats the model, which called it an act and
// rewrote the task text (Vikunja #467). After the rules above because a
// marker never collides with a clock or agenda question.
grammars = append(grammars, router.TaskCaptureGrammar())
return router.New(router.Config{
Grammars: grammars,
Classifier: cls,
@@ -328,11 +401,34 @@ func buildRouter(emb router.Embedder, acts router.ActMatcher, threshold float64,
})
}
// seedDir is the directory containing intent seed files. Each file is named
// <intent>.txt and contains one training example per line (blank lines and
// lines starting with # are ignored). Relative to the working directory.
// seedDir is the directory containing intent seed files, relative to the repo
// root. Each file is named <intent>.txt and holds one training example per
// line (blank lines and lines starting with # are ignored).
const seedDir = "models/seeds"
// seedPath resolves seedDir against the working directory, walking up until it
// finds it. The daemon runs from the repo root and the first candidate hits.
//
// A test does not: `go test ./cmd/mavend/` runs with the working directory at
// cmd/mavend, so every open failed and the simulator scenarios replayed a whole
// scripted day against a classifier holding zero examples (Vikunja #465). They
// passed, which is the part that matters — a green simulator was not exercising
// the routing the deploy runs, and a regression in the seed set could not have
// shown up there.
//
// Bounded at five levels, so a daemon started somewhere without the seeds logs
// the same failure it always did rather than walking to the filesystem root.
func seedPath() string {
dir := seedDir
for i := 0; i < 5; i++ {
if st, err := os.Stat(dir); err == nil && st.IsDir() {
return dir
}
dir = filepath.Join("..", dir)
}
return seedDir
}
// seedClassifier floors the embedded examples so the cold-boot path
// doesn't return ErrNoIntents. Loads examples from seedDir — one file per
// intent (act.txt, reminder.txt, fact.txt, note.txt, query.txt). When the
@@ -357,11 +453,11 @@ func seedClassifier(c *router.Classifier) {
}
total += n
}
log.Printf("voice: loaded %d seed examples from %s", total, seedDir)
log.Printf("voice: loaded %d seed examples from %s", total, seedPath())
}
func loadSeedFile(c *router.Classifier, intent router.Intent) (int, error) {
path := filepath.Join(seedDir, string(intent)+".txt")
path := filepath.Join(seedPath(), string(intent)+".txt")
f, err := os.Open(path)
if err != nil {
return 0, fmt.Errorf("open %s: %w", path, err)
@@ -409,6 +505,36 @@ func seedTools(api ipc.CoreAPI, tools []config.ToolConfig) {
log.Printf("voice: seeded %d act tools from config", n)
}
// repairFactVectors brings stored fact vectors in line with the facts they name
// (#493), once per box, before the embedder marker is even looked at.
//
// Automatic and not a flag, unlike -reembed: only voice-tapped facts are in
// this index, so the work is tens of embeddings rather than the thousands of
// notes that made the backfill a deliberate act. And the box that needs it is
// broken in a way nobody can see — recall answers with the wrong text and
// nothing logs an error — so waiting for an operator to know to run it is how
// the defect survived four restarts in the first place.
func repairFactVectors(dataStore *store.Store, emb router.Embedder) {
if dataStore == nil {
return
}
res, err := dataStore.RepairFactVectors(context.Background(),
// EmbedPassage, the stored side, same as every other writer of these
// vectors.
func(ctx context.Context, text string) ([]float32, error) {
return router.EmbedPassage(ctx, emb, text)
})
if err != nil {
log.Printf("voice: fact vector repair failed, no marker written and nothing half-done — retried next start: %v", err)
return
}
if res.Skipped || res.Rewritten+res.Dropped == 0 {
return
}
log.Printf("voice: fact vector repair — %d re-embedded from the fact they name, %d dropped as voided or superseded, %d already right, took %s (#493)",
res.Rewritten, res.Dropped, res.Kept, res.Took.Round(time.Millisecond))
}
// reembedOnStart is the -reembed flag (set in run()). Opt-in on purpose: see
// runReembed.
var reembedOnStart bool
+45 -33
View File
@@ -1,10 +1,13 @@
// Package main — weatherq.go holds the weather-query keyword helpers: does
// this utterance ask about weather at all, and which city (if any) did it
// name. Both are plain substring/lookup matching, not NLU — extend this file
// rather than voice.go for anything in that shape.
// this utterance ask about weather at all, and which place (if any) did he
// name. Both are plain keyword matching, not NLU — extend this file rather
// than voice.go for anything in that shape.
package main
import "strings"
import (
"regexp"
"strings"
)
// isWeatherQuery returns true if the utterance is about weather.
func isWeatherQuery(u string) bool {
@@ -19,40 +22,49 @@ func isWeatherQuery(u string) bool {
strings.Contains(lower, "temperature")
}
// weatherCities — the city names an utterance may name explicitly, as
// lowercase substrings mapped to the provider's spelling. This is a
// convenience for "какая погода в Лондоне", NOT a source of default truth:
// nothing here is used unless he actually said it.
var weatherCities = map[string]string{
"москв": "Moscow",
"moscow": "Moscow",
"питер": "Saint Petersburg",
"spb": "Saint Petersburg",
"петербур": "Saint Petersburg",
"лондон": "London",
"london": "London",
"париж": "Paris",
"paris": "Paris",
"берлин": "Berlin",
"berlin": "Berlin",
"нью-йорк": "New York",
"new york": "New York",
// weatherPlace — the place he named, after "в"/"во"/"in". One or two words,
// letters and dashes only, so "в Нижнем Новгороде" and "in New York" both
// come through whole and "в 5 утра" does not.
var weatherPlace = regexp.MustCompile(`(?i)(?:^|\s)(?:в|во|in)\s+([\p{L}-]+(?:\s+[\p{L}-]+)?)`)
// weatherNonPlaces — words that follow "в" in a weather question and are not
// cities. "какая погода в доме" is the smart-home sensor, not Open-Meteo, and
// "тепло в комнате" is the same question about the same room.
var weatherNonPlaces = map[string]bool{
"доме": true, "квартире": true, "комнате": true, "спальне": true,
"гостиной": true, "кухне": true, "гараже": true, "офисе": true,
"выходные": true, "субботу": true, "воскресенье": true, "понедельник": true,
"вторник": true, "среду": true, "четверг": true, "пятницу": true,
"обед": true, "обеде": true, "утро": true, "утром": true, "вечер": true,
"вечером": true, "ночь": true, "ночью": true, "целом": true, "принципе": true,
}
// extractWeatherLocation returns the city he named, or the configured default
// extractWeatherLocation returns the place he named, or the configured default
// when he named none. It returns "" when he named none AND no default is
// configured — the caller must then say it does not know.
//
// It used to return "Moscow" in that case. That is a made-up answer presented
// as fact: reading out Moscow's temperature to someone who is not in Moscow is
// wrong in exactly the way maven must never be wrong. voice.weather
// .default_location is the only source of an unstated location.
// It used to be a hand-written table of six cities in two spellings each
// (Vikunja #421). Anything outside it — Kazan, Tbilisi — was dropped silently
// and answered for the default location, which reads as a correct answer about
// the wrong place. There is a geocoder behind this now: internal/weather
// already calls Open-Meteo's geocoding endpoint for every lookup, so any place
// it knows is a place he can ask about, and the table bought nothing.
//
// A named place that the geocoder cannot resolve is the caller's problem to
// report, not this function's to hide.
//
// It used to return "Moscow" when he named nothing. That is a made-up answer
// presented as fact. voice.weather.default_location is the only source of an
// unstated location.
func extractWeatherLocation(u, defaultLoc string) string {
lower := strings.ToLower(u)
for substr, name := range weatherCities {
if strings.Contains(lower, substr) {
return name
}
m := weatherPlace.FindStringSubmatch(u)
if m == nil {
return defaultLoc
}
return defaultLoc
place := strings.TrimSpace(m[1])
first := strings.ToLower(strings.Fields(place)[0])
if weatherNonPlaces[first] {
return defaultLoc
}
return place
}
+33
View File
@@ -0,0 +1,33 @@
package main
import "testing"
// TestExtractWeatherLocation — any place he names comes through, not just the
// six that used to be in a table (Vikunja #421).
func TestExtractWeatherLocation(t *testing.T) {
cases := []struct {
utterance string
def string
want string
}{
// The cities the table had, and the ones it silently dropped.
{"какая погода в Москве", "Berlin", "Москве"},
{"какая погода в Казани", "Berlin", "Казани"},
{"погода в Тбилиси?", "Berlin", "Тбилиси"},
{"what's the weather in New York", "Berlin", "New York"},
{"тепло в Нижнем Новгороде?", "Berlin", "Нижнем Новгороде"},
// He named nothing: the configured default, and nothing at all when
// there is no default.
{"какая сегодня погода", "Berlin", "Berlin"},
{"какая сегодня погода", "", ""},
// "в" followed by something that is not a place stays the default —
// the house sensors and the day words answer elsewhere.
{"тепло в комнате?", "Berlin", "Berlin"},
{"какая погода в выходные", "Berlin", "Berlin"},
}
for _, c := range cases {
if got := extractWeatherLocation(c.utterance, c.def); got != c.want {
t.Errorf("extractWeatherLocation(%q, %q) = %q, want %q", c.utterance, c.def, got, c.want)
}
}
}
+60
View File
@@ -0,0 +1,60 @@
package main
import (
"context"
"errors"
"log"
"github.com/kami/maven/internal/phraser"
)
// worldPhraser — the naming half of the degradation rule (docs/offload.md), as
// the query sources see it. Only *phraser.LLMPhraser implements it, so the
// Stub and every test double stay exactly as they are.
type worldPhraser interface {
PhraseWorld(ctx context.Context, utterance string, sources []string) (string, error)
}
// worldGap — what he hears when the question is about the world, the workstation
// model is the one configured to answer it, and that machine is not answering.
//
// It says the true thing. The resident 1.7B is not a worse answer here, it is an
// invented one: "Война и мир" came back with Левитан as its author, and a
// question about his meeting came back as a swimming competition in Nottingham.
// Naming the gap is the rule CLAUDE.md already applies to a sibling service
// being down.
const worldGap = "сейчас не могу ответить — большая модель недоступна, а придумывать не хочу."
// phraseWorld asks the world model, or reports the gap.
//
// The three outcomes come straight from LLMPhraser.PhraseWorld: no workstation
// configured means the resident model answers as it always has, a workstation
// that is up answers, and a workstation that is down returns
// phraser.ErrNoWorldModel. A phraser that has no world seam at all — the Stub,
// and the doubles in the tests — is the first of those three.
func (h *reactiveHandler) phraseWorld(ctx context.Context, utterance string, sources []string) (string, error) {
if h.phraser == nil {
return "", phraser.ErrNoWorldModel
}
if w, ok := h.phraser.(worldPhraser); ok {
return w.PhraseWorld(ctx, utterance, sources)
}
return h.phraser.PhraseQuery(ctx, utterance, sources)
}
// phraseSource asks the world model to answer from a passage someone already
// fetched — a live search result, a ZIM article, a page he named. It returns ""
// rather than the gap phrase, because these callers hold something better than a
// gap: the passage itself, which their own floor reads back to him. Nothing is
// invented either way, and a real quote beats "не могу сейчас".
func (h *reactiveHandler) phraseSource(ctx context.Context, name, utterance string, sources []string) string {
reply, err := h.phraseWorld(ctx, utterance, sources)
switch {
case errors.Is(err, phraser.ErrNoWorldModel):
log.Printf("voice: %s: no world model, reading the source back instead", name)
return ""
case err != nil:
log.Printf("voice: %s: phrase: %v", name, err)
}
return reply
}
+85
View File
@@ -0,0 +1,85 @@
package main
import (
"context"
"strings"
"testing"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
)
// gapPhraser — a phraser whose world model is configured and asleep, which is
// the state the naming half exists for.
type gapPhraser struct {
*phraser.Stub
worldCalls int
}
func (g *gapPhraser) PhraseWorld(context.Context, string, []string) (string, error) {
g.worldCalls++
return "", phraser.ErrNoWorldModel
}
func worldTurn(utterance string) *queryTurn {
return &queryTurn{dec: router.Decision{Intent: router.IntentQuery, Utterance: utterance}}
}
// A world question with the workstation asleep says so. The resident model is
// not asked, because what it produces here is an invention with no signal that
// it is one.
func TestQueryGeneralNamesTheGap(t *testing.T) {
g := &gapPhraser{Stub: phraser.NewStub()}
h := &reactiveHandler{phraser: g}
reply, ok := h.queryGeneral(context.Background(), worldTurn("почему небо голубое"))
if !ok {
t.Fatal("queryGeneral passed on the last source in the chain")
}
if reply != worldGap {
t.Fatalf("reply = %q, want the named gap", reply)
}
if g.worldCalls != 1 {
t.Fatalf("PhraseWorld called %d times, want 1", g.worldCalls)
}
}
// A phraser with no world seam at all — the Stub, and every box with no
// `workstation` block — answers exactly as it did before this seam existed.
func TestQueryGeneralWithoutAWorldModelIsUnchanged(t *testing.T) {
h := &reactiveHandler{phraser: phraser.NewStub()}
reply, ok := h.queryGeneral(context.Background(), worldTurn("почему небо голубое"))
if !ok {
t.Fatal("queryGeneral passed on the last source in the chain")
}
if reply != "не знаю." {
t.Fatalf("reply = %q, want the Stub's answer", reply)
}
}
// The gap is spoken aloud by a Russian voice, so it is Russian, feminine and
// informal. "не хочу" and "не могу" are her own verbs; there is no "вы" and no
// English in it.
func TestWorldGapIsInPersona(t *testing.T) {
for _, bad := range []string{"вы", "ваш", "рад ", "дорогой", "милый"} {
if strings.Contains(worldGap, bad) {
t.Errorf("the gap phrase contains %q: %s", bad, worldGap)
}
}
if strings.ContainsAny(worldGap, "abcdefghijklmnopqrstuvwxyz") {
t.Errorf("the gap phrase has Latin letters in it: %s", worldGap)
}
}
// The sources that hold a passage read it back rather than name a gap. He gets a
// real quote instead of "не могу сейчас", and nothing is invented either way.
func TestASourceWithAPassageReadsItBackInsteadOfNamingTheGap(t *testing.T) {
g := &gapPhraser{Stub: phraser.NewStub()}
h := &reactiveHandler{phraser: g}
if got := h.phraseSource(context.Background(), "search", "почему небо голубое",
[]string{"Рэлеевское рассеяние."}); got != "" {
t.Fatalf("phraseSource = %q, want \"\" so the caller's own floor reads the passage back", got)
}
if g.worldCalls != 1 {
t.Fatalf("PhraseWorld called %d times, want 1", g.worldCalls)
}
}
+117
View File
@@ -0,0 +1,117 @@
package main
import (
"os"
"path/filepath"
"strconv"
"strings"
)
// The card is an AMD 7900 GRE with 16GB, driven by amdgpu and ROCm. Everything
// here reads sysfs and forks nothing: rocm-smi is not even installed on the
// workstation, and a poll that costs a subprocess every second is a poll that
// gets tuned down until it is useless.
// gpuProc — one process holding the compute engine.
type gpuProc struct {
PID int
Comm string
VRAM int64 // bytes, as the kernel accounts them to this process
}
// probe reads the two sysfs trees the supervisor decides from.
//
// kfdRoot is /sys/class/kfd/kfd/proc, one directory per ROCm process. The
// directory appears when the process initialises HIP, which is well before it
// allocates anything large. That is the whole reason this works: the job that
// is about to want the card announces itself while it is still starting up,
// so we see the contender rather than only the winner of an allocation race.
//
// drmDev is /sys/class/drm/cardN/device, which reports total and used VRAM for
// the card as a whole.
type probe struct {
kfdRoot string
drmDev string
}
// foreign lists every ROCm process that is not ours. selfPID is the supervisor's
// llama-server child, or 0 when it is not running.
//
// An unreadable kfd tree returns no processes and no error. That is deliberate
// and it is the safe direction only because startVRAM also has to agree before
// anything launches: a supervisor that cannot see the KFD never sees free VRAM
// either, because the CPT run holding the card shows up in the drm totals.
func (p probe) foreign(selfPID int) []gpuProc {
entries, err := os.ReadDir(p.kfdRoot)
if err != nil {
return nil
}
var out []gpuProc
for _, e := range entries {
pid, err := strconv.Atoi(e.Name())
if err != nil || pid == selfPID {
continue
}
out = append(out, gpuProc{
PID: pid,
Comm: readComm(pid),
VRAM: p.procVRAM(e.Name()),
})
}
return out
}
// procVRAM sums the per-node vram_* files under one process directory. The
// suffix is the KFD topology node id (vram_35881 on this card), so it is
// globbed rather than named, and a machine with two cards sums both.
func (p probe) procVRAM(pid string) int64 {
matches, err := filepath.Glob(filepath.Join(p.kfdRoot, pid, "vram_*"))
if err != nil {
return 0
}
var total int64
for _, m := range matches {
total += readInt(m)
}
return total
}
// freeVRAM reports the bytes the card has left. Used only to decide whether to
// start: a shortfall here means llama-server would refuse to load anyway. It is
// never used to decide to stop, because by the time free VRAM has dropped the
// other job has already failed its allocation, which is exactly the outcome
// yielding exists to prevent.
func (p probe) freeVRAM() int64 {
total := readInt(filepath.Join(p.drmDev, "mem_info_vram_total"))
used := readInt(filepath.Join(p.drmDev, "mem_info_vram_used"))
if total <= 0 {
return 0
}
if free := total - used; free > 0 {
return free
}
return 0
}
func readInt(path string) int64 {
b, err := os.ReadFile(path)
if err != nil {
return 0
}
n, err := strconv.ParseInt(strings.TrimSpace(string(b)), 10, 64)
if err != nil {
return 0
}
return n
}
// readComm names the contender for the log. The log is the instrument for the
// open question in Vikunja #488: whether a process can want this card without
// ever registering on the KFD, which a Vulkan or video-decode job would.
func readComm(pid int) string {
b, err := os.ReadFile(filepath.Join("/proc", strconv.Itoa(pid), "comm"))
if err != nil {
return "?"
}
return strings.TrimSpace(string(b))
}
+103
View File
@@ -0,0 +1,103 @@
package main
import (
"net/http"
"net/http/httptest"
"net/url"
"os"
"path/filepath"
"strconv"
"testing"
)
// fakeKFD builds the sysfs shape the workstation actually has: one directory
// per ROCm process, each holding a vram_<node> file. Sampled from the live box
// on 02-08-2026, where the CPT run appeared as proc/478104/vram_35881.
func fakeKFD(t *testing.T, vramByPID map[int]int64) string {
t.Helper()
root := t.TempDir()
for pid, vram := range vramByPID {
dir := filepath.Join(root, strconv.Itoa(pid))
if err := os.MkdirAll(dir, 0o755); err != nil {
t.Fatal(err)
}
f := filepath.Join(dir, "vram_35881")
if err := os.WriteFile(f, []byte(strconv.FormatInt(vram, 10)+"\n"), 0o644); err != nil {
t.Fatal(err)
}
}
return root
}
func TestForeignExcludesOurChild(t *testing.T) {
root := fakeKFD(t, map[int]int64{478104: 12791693312, 999: 4096})
p := probe{kfdRoot: root}
all := p.foreign(0)
if len(all) != 2 {
t.Fatalf("with no child running, both processes are foreign, got %d", len(all))
}
ours := p.foreign(999)
if len(ours) != 1 || ours[0].PID != 478104 {
t.Fatalf("our own llama-server must not count as a contender, got %+v", ours)
}
if ours[0].VRAM != 12791693312 {
t.Errorf("per-process VRAM = %d, want the value from vram_35881", ours[0].VRAM)
}
}
// An empty KFD tree is the state that permits a start, so it must read as empty
// rather than as an error the caller has to interpret.
func TestForeignEmptyAndMissing(t *testing.T) {
if got := (probe{kfdRoot: t.TempDir()}).foreign(0); len(got) != 0 {
t.Errorf("empty kfd tree: got %d processes, want 0", len(got))
}
if got := (probe{kfdRoot: "/nonexistent"}).foreign(0); got != nil {
t.Errorf("missing kfd tree: got %+v, want nil", got)
}
}
func TestFreeVRAM(t *testing.T) {
dev := t.TempDir()
write := func(name, v string) {
if err := os.WriteFile(filepath.Join(dev, name), []byte(v), 0o644); err != nil {
t.Fatal(err)
}
}
// The live numbers from the workstation while the CPT run held the card.
write("mem_info_vram_total", "17163091968\n")
write("mem_info_vram_used", "13396389888\n")
p := probe{drmDev: dev}
if got, want := p.freeVRAM(), int64(3766702080); got != want {
t.Errorf("freeVRAM = %d, want %d", got, want)
}
if got := (probe{drmDev: "/nonexistent"}).freeVRAM(); got != 0 {
t.Errorf("unreadable card reports %d free, want 0 so nothing starts", got)
}
}
// With no model loaded the supervisor must still answer, and it must answer 503
// rather than hanging or proxying into a closed port. Maven reads this endpoint
// on a timer forever, including while the workstation is busy.
func TestHealthAndProxyRefuseWhenNotReady(t *testing.T) {
s := &supervisor{run: newRunner("/bin/true", nil, "")}
h := s.handler(mustURL(t, "http://127.0.0.1:1"))
for _, path := range []string{"/health", "/v1/chat/completions"} {
w := httptest.NewRecorder()
h.ServeHTTP(w, httptest.NewRequest(http.MethodGet, path, nil))
if w.Code != http.StatusServiceUnavailable {
t.Errorf("%s with no model: got %d, want 503", path, w.Code)
}
}
}
func mustURL(t *testing.T, s string) *url.URL {
t.Helper()
u, err := url.Parse(s)
if err != nil {
t.Fatal(err)
}
return u
}
+247
View File
@@ -0,0 +1,247 @@
// mavgpud — the workstation's GPU supervisor.
//
// It runs on the workstation (an AMD 7900 GRE, 16GB), not on homesrv, and it is
// deployed separately from the Maven daemons. Maven does not participate in any
// of this and never asks for a start: it reads /health through internal/llm.Pair
// and either gets the big model or falls back to the resident 1.7B.
//
// The rule, from Vikunja #488: keep llama-server loaded whenever the card is
// free, unload it when it has been idle too long or when another process needs
// the card. Not on demand, because a 7-14B takes tens of seconds to load and a
// world question would be answered by a gap every time the card had been quiet.
// Not always on, because that holds 16GB against the owner's own jobs.
package main
import (
"context"
"encoding/json"
"flag"
"log"
"net/http"
"net/http/httputil"
"net/url"
"os"
"os/signal"
"sync/atomic"
"syscall"
"time"
)
type config struct {
Listen string `json:"listen"` // what Maven talks to
LlamaAddr string `json:"llama_addr"` // where llama-server binds
LlamaBin string `json:"llama_bin"`
// LlamaArgs must include the flags that bind LlamaAddr. They are passed
// through untouched so the model, context size and layer count stay the
// owner's business and not this daemon's schema.
LlamaArgs []string `json:"llama_args"`
KFDRoot string `json:"kfd_root"`
DRMDevice string `json:"drm_device"`
Poll duration `json:"poll"`
IdleTimeout duration `json:"idle_timeout"`
StopGrace duration `json:"stop_grace"`
MinFreeVRAM int64 `json:"min_free_vram_bytes"`
// EvictAfter and StartAfter are counted in polls, not seconds. Both exist
// to damp flapping: a one-tick blip from a short-lived rocm process must
// not evict the model, and a card that has just been released must not be
// grabbed before the previous job has finished unmapping.
EvictAfter int `json:"evict_after_polls"`
StartAfter int `json:"start_after_polls"`
}
func defaults() config {
return config{
Listen: ":8080",
LlamaAddr: "127.0.0.1:8081",
KFDRoot: "/sys/class/kfd/kfd/proc",
DRMDevice: "/sys/class/drm/card1/device",
Poll: duration(time.Second),
IdleTimeout: duration(15 * time.Minute),
StopGrace: duration(20 * time.Second),
MinFreeVRAM: 15 << 30,
EvictAfter: 2,
StartAfter: 5,
}
}
// duration lets the config file say "15m" instead of counting nanoseconds.
type duration time.Duration
func (d *duration) UnmarshalJSON(b []byte) error {
var s string
if err := json.Unmarshal(b, &s); err != nil {
return err
}
v, err := time.ParseDuration(s)
if err != nil {
return err
}
*d = duration(v)
return nil
}
func main() {
path := flag.String("config", "/etc/mavgpud.json", "config file")
flag.Parse()
cfg := defaults()
b, err := os.ReadFile(*path)
if err != nil {
log.Fatalf("mavgpud: read config: %v", err)
}
if err := json.Unmarshal(b, &cfg); err != nil {
log.Fatalf("mavgpud: parse config: %v", err)
}
if cfg.LlamaBin == "" {
log.Fatal("mavgpud: llama_bin is required")
}
base := "http://" + cfg.LlamaAddr
run := newRunner(cfg.LlamaBin, cfg.LlamaArgs, base+"/health")
sup := &supervisor{
cfg: cfg,
probe: probe{kfdRoot: cfg.KFDRoot, drmDev: cfg.DRMDevice},
run: run,
}
sup.touch()
ctx, cancel := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
defer cancel()
target, err := url.Parse(base)
if err != nil {
log.Fatalf("mavgpud: llama_addr: %v", err)
}
srv := &http.Server{Addr: cfg.Listen, Handler: sup.handler(target)}
go func() {
log.Printf("mavgpud: listening on %s, model %s", cfg.Listen, cfg.LlamaBin)
if err := srv.ListenAndServe(); err != nil && err != http.ErrServerClosed {
log.Fatalf("mavgpud: listen: %v", err)
}
}()
sup.loop(ctx)
// The card must come back before we do. A supervisor that exits leaving
// llama-server holding 14GB is worse than one that never ran.
shut, done := context.WithTimeout(context.Background(), 5*time.Second)
defer done()
_ = srv.Shutdown(shut)
run.stop(time.Duration(cfg.StopGrace))
}
type supervisor struct {
cfg config
probe probe
run *runner
lastReq atomic.Int64 // unix nanos of the last request Maven sent
foreignStreak int
clearStreak int
}
func (s *supervisor) touch() { s.lastReq.Store(time.Now().UnixNano()) }
func (s *supervisor) idle() time.Duration {
return time.Since(time.Unix(0, s.lastReq.Load()))
}
// handler serves the two things the workstation exposes.
//
// /health is answered locally and always, with no GPU cost and no round trip,
// because it is the only thing Maven reads and Maven reads it on a timer
// forever. Everything else is llama-server's API, reverse-proxied. Proxying
// rather than pointing Maven straight at llama-server is what makes the idle
// window measurable: the supervisor cannot otherwise know when the model was
// last used.
func (s *supervisor) handler(target *url.URL) http.Handler {
proxy := httputil.NewSingleHostReverseProxy(target)
mux := http.NewServeMux()
mux.HandleFunc("/health", func(w http.ResponseWriter, r *http.Request) {
if !s.run.isReady() {
http.Error(w, "model not loaded", http.StatusServiceUnavailable)
return
}
w.Header().Set("Content-Type", "application/json")
_, _ = w.Write([]byte(`{"status":"ok"}`))
})
mux.HandleFunc("/", func(w http.ResponseWriter, r *http.Request) {
if !s.run.isReady() {
http.Error(w, "model not loaded", http.StatusServiceUnavailable)
return
}
s.touch()
proxy.ServeHTTP(w, r)
})
return mux
}
func (s *supervisor) loop(ctx context.Context) {
t := time.NewTicker(time.Duration(s.cfg.Poll))
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
s.tick(ctx)
}
}
}
// tick is the whole decision. Yielding is checked before starting, and presence
// on the KFD is what triggers it — not a VRAM threshold. A ROCm process
// registers under /sys/class/kfd/kfd/proc when it initialises HIP, before it
// allocates, so we see a contender during its startup rather than after it has
// already failed to get the memory it wanted.
func (s *supervisor) tick(ctx context.Context) {
others := s.probe.foreign(s.run.pid())
if len(others) > 0 {
s.foreignStreak++
s.clearStreak = 0
} else {
s.foreignStreak = 0
s.clearStreak++
}
if s.run.running() {
s.run.refreshReady(ctx)
switch {
case s.foreignStreak >= s.cfg.EvictAfter:
log.Printf("mavgpud: yielding the card to %s", describe(others))
s.run.stop(time.Duration(s.cfg.StopGrace))
case s.idle() > time.Duration(s.cfg.IdleTimeout):
log.Printf("mavgpud: idle for %s, unloading", s.idle().Round(time.Second))
s.run.stop(time.Duration(s.cfg.StopGrace))
}
return
}
if s.clearStreak < s.cfg.StartAfter {
return
}
if free := s.probe.freeVRAM(); free < s.cfg.MinFreeVRAM {
return
}
s.touch() // the idle clock starts at load, not at the last request before it
if err := s.run.start(); err != nil {
log.Printf("mavgpud: start llama-server: %v", err)
}
}
// describe names the contenders in the log. This log is the instrument for the
// open question in #488: whether polling the KFD misses a job that wants the
// card without registering there.
func describe(procs []gpuProc) string {
out := ""
for i, p := range procs {
if i > 0 {
out += ", "
}
out += p.Comm
}
return out
}
+146
View File
@@ -0,0 +1,146 @@
package main
import (
"context"
"log"
"net/http"
"os/exec"
"sync"
"syscall"
"time"
)
// runner owns one llama-server process. Owning it is the point of the daemon:
// the workstation cannot keep a 7-14B resident, because that holds 16GB against
// the owner's CPT runs, Correx and the manga-recap pipeline. So the thing that
// stays up is this, which costs no VRAM, and the model comes and goes under it.
type runner struct {
bin string
args []string
// ready is llama-server's own /health, which answers "is a model loaded".
// Loading a 7-14B takes tens of seconds, so started is not ready.
readyURL string
mu sync.Mutex
cmd *exec.Cmd
ready bool
// yielding — stop() has sent the signal and the exit that follows is ours.
// llama-server aborts on SIGTERM (its static teardown throws, upstream
// ggml-org/llama.cpp), so a routine yield and a real crash produce the same
// "signal: aborted" and used to log identically (Vikunja #491).
yielding bool
http *http.Client
}
func newRunner(bin string, args []string, readyURL string) *runner {
return &runner{
bin: bin, args: args, readyURL: readyURL,
http: &http.Client{Timeout: 2 * time.Second},
}
}
// pid is the child's, or 0. The GPU probe needs it to tell our own model apart
// from a contender.
func (r *runner) pid() int {
r.mu.Lock()
defer r.mu.Unlock()
if r.cmd == nil || r.cmd.Process == nil {
return 0
}
return r.cmd.Process.Pid
}
func (r *runner) running() bool { return r.pid() != 0 }
// isReady reports the cached readiness. The supervisor loop refreshes it; the
// health handler only reads, so answering /health never costs a round trip.
func (r *runner) isReady() bool {
r.mu.Lock()
defer r.mu.Unlock()
return r.ready
}
// start launches llama-server. It returns as soon as the process exists, not
// when the model is loaded.
func (r *runner) start() error {
r.mu.Lock()
defer r.mu.Unlock()
if r.cmd != nil {
return nil
}
cmd := exec.Command(r.bin, r.args...)
// Own process group, so stop kills anything llama-server spawned rather
// than leaving it holding VRAM after we have declared the card yielded.
cmd.SysProcAttr = &syscall.SysProcAttr{Setpgid: true}
if err := cmd.Start(); err != nil {
return err
}
r.cmd, r.ready, r.yielding = cmd, false, false
log.Printf("mavgpud: started llama-server pid=%d", cmd.Process.Pid)
go func() {
err := cmd.Wait()
r.mu.Lock()
yielded := r.yielding
r.cmd, r.ready, r.yielding = nil, false, false
r.mu.Unlock()
if yielded {
log.Printf("mavgpud: llama-server stopped, card yielded (%v)", err)
return
}
log.Printf("mavgpud: llama-server exited: %v", err)
}()
return nil
}
// stop ends llama-server and waits for the VRAM to come back. SIGTERM first so
// it unmaps cleanly, SIGKILL after the grace window. Returning before the
// process is gone would let the supervisor report a free card while 14GB is
// still mapped, which is the one lie that would make yielding useless.
func (r *runner) stop(grace time.Duration) {
r.mu.Lock()
cmd := r.cmd
r.ready = false
if cmd != nil && cmd.Process != nil {
// The exit that follows is ours, not a crash.
r.yielding = true
}
r.mu.Unlock()
if cmd == nil || cmd.Process == nil {
return
}
pgid := -cmd.Process.Pid
_ = syscall.Kill(pgid, syscall.SIGTERM)
deadline := time.Now().Add(grace)
for time.Now().Before(deadline) {
if !r.running() {
return
}
time.Sleep(100 * time.Millisecond)
}
log.Printf("mavgpud: llama-server did not exit in %s, killing", grace)
_ = syscall.Kill(pgid, syscall.SIGKILL)
}
// refreshReady asks llama-server whether the model is loaded. Called once per
// supervisor tick, never per request.
func (r *runner) refreshReady(ctx context.Context) {
if !r.running() {
return
}
ok := false
req, err := http.NewRequestWithContext(ctx, http.MethodGet, r.readyURL, nil)
if err == nil {
resp, err := r.http.Do(req)
if err == nil {
ok = resp.StatusCode == http.StatusOK
resp.Body.Close()
}
}
r.mu.Lock()
was := r.ready
r.ready = ok
r.mu.Unlock()
if ok && !was {
log.Printf("mavgpud: model ready")
}
}
+59
View File
@@ -0,0 +1,59 @@
package main
import (
"os"
"path/filepath"
"testing"
"time"
)
// fakeServer writes an executable standing in for llama-server: it ignores
// SIGTERM the way the real one effectively does — by dying messily rather than
// cleanly — and reports a non-zero status.
func fakeServer(t *testing.T, body string) string {
t.Helper()
path := filepath.Join(t.TempDir(), "fake-llama-server")
if err := os.WriteFile(path, []byte("#!/bin/sh\n"+body+"\n"), 0o755); err != nil {
t.Fatal(err)
}
return path
}
// A deliberate stop is a yield, and the log has to say so.
//
// llama-server aborts inside its own static teardown on SIGTERM, so the exit
// status of a routine yield is identical to that of a real crash. Reading the
// mavgpud log, the two were indistinguishable (Vikunja #491).
func TestStopMarksTheExitAsAYield(t *testing.T) {
r := newRunner(fakeServer(t, "while : ; do sleep 1 ; done"), nil, "")
if err := r.start(); err != nil {
t.Fatalf("start: %v", err)
}
r.mu.Lock()
if r.yielding {
t.Error("a freshly started server is already marked as yielding")
}
r.mu.Unlock()
r.stop(2 * time.Second)
deadline := time.Now().Add(2 * time.Second)
for time.Now().Before(deadline) {
if !r.running() {
return
}
time.Sleep(10 * time.Millisecond)
}
t.Fatal("the child outlived stop")
}
// Stopping when nothing is running must not arm the flag for the next child.
// The next exit after that would be a real crash logged as a yield.
func TestStopWithNoChildDoesNotArmTheFlag(t *testing.T) {
r := newRunner("/nonexistent", nil, "")
r.stop(10 * time.Millisecond)
r.mu.Lock()
defer r.mu.Unlock()
if r.yielding {
t.Error("stop armed the yield flag with no child running")
}
}
+71 -21
View File
@@ -140,6 +140,10 @@ type poller struct {
wgIface string
wgCmd string
// kumaSeen — monitor name → state as of the last poll, so a monitor that
// disappears from the gauge can be marked unknown instead of staying down.
kumaSeen map[string]string
// zen is nil unless a token file was configured — money tracking is a
// capability, off by default like weather and telegram.
zen *zenmoney.Client
@@ -333,50 +337,96 @@ func maxSeverity(a netdataAlarms) string {
return sev
}
// ---- kuma: monitor_status gauge → aggregate service_down -------------------
// ---- kuma: monitor_status gauge → one fact per monitor ---------------------
// Kuma exposes Prometheus text: `monitor_status{...,monitor_name="X"} V` where
// V is 1=up 0=down 2=pending 3=maintenance. We reduce to one aggregate the
// existing ServiceDownRule consumes: "down" if ANY monitor reads 0, else "up".
// Per-service granularity is a later add (a fact per monitor) — the MVP nudge
// only needs "something is down".
var kumaLine = regexp.MustCompile(`^monitor_status\{([^}]*)\}\s+([0-9.eE+-]+)`)
// V is 1=up 0=down 2=pending 3=maintenance. We write one fact per monitor,
// keyed `service_down:<monitor name>`, because the nudge has to say WHICH
// service is down. The aggregate this used to write could not, which is why
// the rule shipped disabled.
var (
kumaLine = regexp.MustCompile(`^monitor_status\{([^}]*)\}\s+([0-9.eE+-]+)`)
kumaName = regexp.MustCompile(`monitor_name="([^"]*)"`)
)
func (p *poller) pollKuma(ctx context.Context, now time.Time) error {
body, err := p.get(ctx, p.kumaURL, p.kumaKey)
if err != nil {
return err
}
down, seen := kumaAnyDown(body)
if !seen {
states := kumaMonitors(body)
if len(states) == 0 {
return fmt.Errorf("no monitor_status metrics (auth/endpoint wrong?)")
}
val := "up"
if down {
val = "down"
var firstErr error
for name, val := range states {
if err := p.writeIfChanged(ctx, kumaFactKey(name), kumaSource, val, now); err != nil && firstErr == nil {
firstErr = err // one bad monitor must not blind the rest
}
}
return p.writeIfChanged(ctx, "service_down", "poll:uptimekuma", val, now)
// A monitor deleted in kuma stops appearing in the gauge, and its last fact
// would otherwise read "down" forever. Mark it unknown, which no rule fires
// on. The seen-set is in memory, so a restart forgets it — harmless, since
// the next poll that still lacks the monitor says nothing new either.
for name := range p.kumaSeen {
if _, still := states[name]; !still {
if err := p.writeIfChanged(ctx, kumaFactKey(name), kumaSource, "unknown", now); err != nil && firstErr == nil {
firstErr = err
}
}
}
p.kumaSeen = states
return firstErr
}
// kumaAnyDown parses kuma's Prometheus text: down=true if any monitor reads 0
// (pending=2/maintenance=3 are not "down"). seen=false ⇒ no monitor_status
// lines matched at all (wrong endpoint or auth rejected before the body).
func kumaAnyDown(body []byte) (down, seen bool) {
// kumaSource — the provenance the loop rule requires. Written here, checked in
// loop.ServiceDownRule; a poller under any other source cannot fire it.
const kumaSource = "poll:uptimekuma"
// kumaFactKey — the fact key for one monitor. The suffix is the name he hears,
// so it stays as kuma spells it rather than being slugged into something else.
func kumaFactKey(name string) string { return "service_down:" + name }
// kumaMonitors parses kuma's Prometheus text into monitor name → state
// ("up"/"down"/"pending"/"maintenance"). An empty map means no monitor_status
// line matched at all (wrong endpoint, or auth rejected before the body).
// A line with no monitor_name label is skipped: a fact nobody can name is
// exactly the thing this replaced.
func kumaMonitors(body []byte) map[string]string {
out := make(map[string]string)
for _, line := range strings.Split(string(body), "\n") {
m := kumaLine.FindStringSubmatch(strings.TrimSpace(line))
if m == nil {
continue
}
seen = true
nm := kumaName.FindStringSubmatch(m[1])
if nm == nil || strings.TrimSpace(nm[1]) == "" {
continue
}
v, err := strconv.ParseFloat(m[2], 64)
if err != nil {
continue
}
if v == 0 {
down = true
}
out[strings.TrimSpace(nm[1])] = kumaState(v)
}
return out
}
// kumaState — the gauge's four values. pending and maintenance are not "down":
// a monitor paused in kuma should silence that monitor, not page him.
func kumaState(v float64) string {
switch v {
case 0:
return "down"
case 1:
return "up"
case 2:
return "pending"
case 3:
return "maintenance"
default:
return "unknown"
}
return down, seen
}
// ---- helpers ---------------------------------------------------------------
+36 -12
View File
@@ -35,22 +35,46 @@ func TestMaxSeverity(t *testing.T) {
}
}
func TestKumaAnyDown(t *testing.T) {
func TestKumaMonitorsNamesEveryOne(t *testing.T) {
cases := []struct {
body string
down, seen bool
name string
body string
want map[string]string
}{
{"", false, false},
{`monitor_status{monitor_name="web"} 1`, false, true},
{`monitor_status{monitor_name="web"} 1` + "\n" + `monitor_status{monitor_name="db"} 0`, true, true},
{`monitor_status{monitor_name="mnt"} 3`, false, true}, // maintenance ≠ down
{`# HELP monitor_status ...`, false, false},
{"empty body", "", map[string]string{}},
{"help line only", `# HELP monitor_status ...`, map[string]string{}},
{
"one up one down",
`monitor_status{monitor_name="web"} 1` + "\n" + `monitor_status{monitor_name="db"} 0`,
map[string]string{"web": "up", "db": "down"},
},
{"maintenance is not down", `monitor_status{monitor_name="mnt"} 3`, map[string]string{"mnt": "maintenance"}},
{"pending is not down", `monitor_status{monitor_name="p"} 2`, map[string]string{"p": "pending"}},
{
"other labels do not hide the name",
`monitor_status{monitor_type="http",monitor_name="ci",monitor_url="x"} 0`,
map[string]string{"ci": "down"},
},
{"a nameless line is skipped", `monitor_status{monitor_type="http"} 0`, map[string]string{}},
}
for _, c := range cases {
down, seen := kumaAnyDown([]byte(c.body))
if down != c.down || seen != c.seen {
t.Errorf("kumaAnyDown(%q) = (%v,%v), want (%v,%v)", c.body, down, seen, c.down, c.seen)
}
t.Run(c.name, func(t *testing.T) {
got := kumaMonitors([]byte(c.body))
if len(got) != len(c.want) {
t.Fatalf("kumaMonitors = %v, want %v", got, c.want)
}
for k, v := range c.want {
if got[k] != v {
t.Errorf("monitor %q = %q, want %q", k, got[k], v)
}
}
})
}
}
func TestKumaFactKeyCarriesTheName(t *testing.T) {
if got := kumaFactKey("nexus db"); got != "service_down:nexus db" {
t.Errorf("kumaFactKey = %q", got)
}
}
+119
View File
@@ -0,0 +1,119 @@
// Command mavseal encrypts a live tmpfs working copy back to the ciphertext
// file, for the case mavend could not do it itself.
//
// mavend seals its database in `defer st.Close()` when run() returns. A daemon
// that is killed rather than shut down never gets there, and because the
// working copy lives in the container's /dev/shm it dies with the container:
// everything written since the last clean shutdown is lost, and the next boot
// silently rolls back to the stale ciphertext. That is not hypothetical — on
// 2026-08-01 the deployed ciphertext was eleven days old.
//
// This is a recovery tool, not part of the daemon. It is safe to run against a
// live database: it takes a consistent snapshot with VACUUM INTO rather than
// mutating the working copy the daemon owns.
//
// Usage:
//
// mavseal -plain /dev/shm/maven-plain.db -cipher /var/lib/maven/maven.db.enc
//
// The key is read from MAVEN_DB_KEY (base64, 32 bytes decoded), the same
// variable the daemon uses. It is never taken as an argument: an argument ends
// up in the shell history and in ps.
package main
import (
"context"
"database/sql"
"encoding/base64"
"flag"
"fmt"
"log"
"os"
"github.com/kami/maven/internal/store"
_ "modernc.org/sqlite"
)
func main() {
log.SetFlags(0)
if err := run(); err != nil {
log.Fatalf("mavseal: %v", err)
}
}
func run() error {
plain := flag.String("plain", "", "path to the plaintext working copy (required)")
cipher := flag.String("cipher", "", "path to write the ciphertext to (required)")
keep := flag.Bool("keep-snapshot", false, "leave the intermediate snapshot on disk for inspection")
flag.Parse()
if *plain == "" || *cipher == "" {
flag.Usage()
return fmt.Errorf("both -plain and -cipher are required")
}
key, err := readKey()
if err != nil {
return err
}
// VACUUM INTO rather than a WAL checkpoint on the file itself. The daemon
// is usually still running and still writing when this is needed, and
// checkpointing its working copy mutates a database it owns. VACUUM INTO
// reads a consistent snapshot into a new file and touches nothing else, so
// the worst case is a snapshot a few seconds stale instead of a torn one.
snap := *plain + ".mavseal-snapshot"
os.Remove(snap)
if err := snapshot(*plain, snap); err != nil {
return err
}
if !*keep {
defer os.Remove(snap)
}
before := fileSize(*cipher)
if err := store.SealPlaintext(snap, *cipher, key); err != nil {
return err
}
log.Printf("sealed %s → %s (%d bytes, was %d)", *plain, *cipher, fileSize(*cipher), before)
return nil
}
// readKey pulls the same base64 key the daemon reads. Fails closed: a short or
// unparseable key must not silently produce a file nothing can open.
func readKey() ([]byte, error) {
raw := os.Getenv("MAVEN_DB_KEY")
if raw == "" {
return nil, fmt.Errorf("MAVEN_DB_KEY is not set")
}
key, err := base64.StdEncoding.DecodeString(raw)
if err != nil {
return nil, fmt.Errorf("MAVEN_DB_KEY is not valid base64: %w", err)
}
if len(key) != 32 {
return nil, fmt.Errorf("MAVEN_DB_KEY decodes to %d bytes, want 32", len(key))
}
return key, nil
}
// snapshot writes a consistent copy of src to dst with VACUUM INTO. The copy
// includes everything committed to the write-ahead log, which is most of what
// is worth saving on a daemon that has been up for hours.
func snapshot(src, dst string) error {
db, err := sql.Open("sqlite", src)
if err != nil {
return fmt.Errorf("open working copy: %w", err)
}
defer db.Close()
if _, err := db.ExecContext(context.Background(), "VACUUM INTO ?", dst); err != nil {
return fmt.Errorf("snapshot: %w", err)
}
return nil
}
func fileSize(path string) int64 {
fi, err := os.Stat(path)
if err != nil {
return 0
}
return fi.Size()
}
+32 -1
View File
@@ -139,7 +139,10 @@ func TestHandleAmbientIgnoresNonMeetings(t *testing.T) {
}
func TestHandleAmbientAuth(t *testing.T) {
body := `{"title":"Планёрка 10:00","posted_at":"2026-08-03T09:40:00Z"}`
// "завтра" so the meeting is in the future in every zone: the clock in the
// text is a local wall clock and posted_at is a Z instant, so a bare "10:00"
// is already stale on a box east of UTC and stores nothing.
body := `{"title":"Планёрка завтра 10:00","posted_at":"2026-08-03T09:40:00Z"}`
newReq := func(hdr, val string) *http.Request {
r := httptest.NewRequest(http.MethodPost, "/api/ambient", strings.NewReader(body))
@@ -245,3 +248,31 @@ func TestHandleAmbientBadInput(t *testing.T) {
}
})
}
// The Z-instant defect end to end: a phone posts an RFC 3339 instant in UTC and
// the clock inside the text is his wall clock. On a UTC+4 box a 14:30 standup
// used to be stored at 18:30, and the size of the error was the deploy's offset.
func TestHandleAmbientStoresTheWallClockHeRead(t *testing.T) {
// No zone juggling: the assertion is that the stored wall clock is the one
// he read, whatever zone the box is in. That is false under the old code
// on every box except a UTC one.
core := &ambientCore{}
rr, resp := postAmbient(t, core, ambientTestToken, calendar.Notification{
Package: "com.slack",
Title: "Standup",
Text: "созвон завтра в 14:30",
Posted: time.Date(2026, 8, 2, 9, 0, 0, 0, time.UTC),
})
if rr.Code != http.StatusCreated {
t.Fatalf("status = %d, want 201: %s", rr.Code, rr.Body)
}
if !strings.HasSuffix(resp.Key, "_Standup") {
t.Errorf("key = %q, want a key naming the meeting", resp.Key)
}
if len(core.writeLog) != 1 {
t.Fatalf("expected 1 fact write, got %d", len(core.writeLog))
}
if got := core.writeLog[0].Value; got != "Standup @ 14:30-15:00" {
t.Errorf("value = %q, want %q", got, "Standup @ 14:30-15:00")
}
}
+24
View File
@@ -0,0 +1,24 @@
{{template "shellTop" "chat"}}
<h1>Chat</h1>
<section class=card>
{{if .Error}}<div class="msg msg-err">{{.Error}}</div>{{end}}
<div class="scroll chat-scroll" id=chatHistory>
{{range .Messages}}
<div class="chat-msg {{.Role}}"><strong>{{if eq .Role "user"}}you{{else}}maven{{end}}:</strong> {{.Text}}</div>
{{else}}
<div class=empty>
<svg class=icon width="20" height="20"><use href="/ethos-icons.svg#i-message"/></svg>
<div>start a conversation</div>
</div>
{{end}}
</div>
<form method=post action=/api/chat class=chat-form>
<input type=text name=text class=input-wide placeholder="type a message..." required autofocus>
<button class=btn>send</button>
</form>
</section>
<script>
var ch = document.getElementById('chatHistory');
if(ch) ch.scrollTop = ch.scrollHeight;
</script>
{{template "shellBottom"}}
+36 -2
View File
@@ -54,7 +54,9 @@ type fakeCore struct {
revertErr error
// for handleNotifications tests
nudgesErr error
nudgesErr error
attempts []ipc.DeliveryAttempt
attemptStatus string
// for handleHistory tests
historyFacts []ipc.Fact
@@ -77,7 +79,7 @@ func (f *fakeCore) MCPServers(context.Context) ([]ipc.MCPServerStatus, error) {
return f.mcpServers, f.mcpErr
}
func (f *fakeCore) Chat(_ context.Context, text string) (string, error) {
func (f *fakeCore) Chat(_ context.Context, _, text string) (string, error) {
f.chatText = text
if f.chatErr != nil {
return "", f.chatErr
@@ -1251,3 +1253,35 @@ func TestHandleWS_AssertedSession_PassesGate(t *testing.T) {
t.Fatalf("status = 403 on an asserted session; body=%s", rr.Body.String())
}
}
func (f *fakeCore) DeliveryAttempts(_ context.Context, status string, _ int) ([]ipc.DeliveryAttempt, error) {
f.attemptStatus = status
return f.attempts, nil
}
// TestHandleNotifications_ShowsTheOutbox — the outbox was written and never
// read, so a dropped or failed send was invisible (Vikunja #390).
func TestHandleNotifications_ShowsTheOutbox(t *testing.T) {
done := time.Date(2026, 8, 4, 9, 0, 30, 0, time.UTC)
core := &fakeCore{
attempts: []ipc.DeliveryAttempt{
{Kind: "nudge", Rule: "care-check", Channel: "telegram", Status: "dropped",
Created: done.Add(-30 * time.Second), Completed: &done},
{Kind: "reminder", ReminderID: 7, Channel: "voice", Status: "pending", Created: done},
},
}
rr := httptest.NewRecorder()
handleNotifications(rr, httptest.NewRequest(http.MethodGet, "/notifications?status=dropped", nil), core)
if rr.Code != http.StatusOK {
t.Fatalf("status = %d, want 200; body=%s", rr.Code, rr.Body.String())
}
if core.attemptStatus != "dropped" {
t.Errorf("status filter = %q, want it passed through", core.attemptStatus)
}
body := rr.Body.String()
for _, want := range []string{"care-check", "dropped", "reminder #7", "Delivery outbox"} {
if !strings.Contains(body, want) {
t.Errorf("rendered outbox missing %q", want)
}
}
}
+127 -253
View File
@@ -75,6 +75,23 @@ var morningHTML string
//go:embed events.html
var eventsHTML string
//go:embed tools.html
var toolsHTML string
//go:embed routines.html
var routinesHTML string
//go:embed chat.html
var chatPageHTML string
// shellHTML — the shell partial every page is wrapped in: "shellTop", the
// "sidebar" it calls, and "shellBottom". It used to be two Go string constants
// with the sidebar assembled by a strings.Builder, which is the one piece of
// markup that was still concatenated in Go.
//
//go:embed shell.html
var shellHTML string
// ── Ethos Workstation Shell ──
//
// Two template pieces that wrap every page:
@@ -135,74 +152,40 @@ var sidebarSections = []struct {
},
}
func sidebarActive(url, key string, activeKey string) string {
if key == activeKey {
return `class="active"`
}
return ""
}
// sidebarHTML renders the sidebar navigation given the active page key.
func sidebarHTML(active string) template.HTML {
var b strings.Builder
for _, sec := range sidebarSections {
b.WriteString(`<div class=sidebar-section>`)
b.WriteString(`<div class=sidebar-label>`)
b.WriteString(sec.Label)
b.WriteString(`</div>`)
for _, p := range sec.Pages {
cls := ""
if p.Key == active {
cls = ` class="active"`
}
b.WriteString(`<a href="`)
b.WriteString(p.URL)
b.WriteString(`"`)
b.WriteString(cls)
b.WriteString(`><span class=icon>`)
b.WriteString(pageIcon(p.Key))
b.WriteString(`</span><span>`)
b.WriteString(p.Label)
b.WriteString(`</span></a>`)
}
b.WriteString(`</div>`)
}
return template.HTML(b.String())
}
// pageIcon returns an ethos-icons.svg <use> reference for the given page.
// pageIcon returns the ethos-icons.svg symbol id for the given page. The
// sidebar template wraps it in the <use> reference.
func pageIcon(key string) string {
switch key {
case "dash":
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-grid"/></svg>`
return "i-grid"
case "history":
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-clock"/></svg>`
return "i-clock"
case "trace":
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-wave"/></svg>`
return "i-wave"
case "notifications":
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-bell"/></svg>`
return "i-bell"
case "tasks":
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-grid"/></svg>`
return "i-grid"
case "reminders":
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-calendar"/></svg>`
return "i-calendar"
case "routines":
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-repeat"/></svg>`
return "i-repeat"
case "morning":
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-calendar"/></svg>`
return "i-calendar"
case "chat":
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-message"/></svg>`
return "i-message"
case "voice":
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-mic"/></svg>`
return "i-mic"
case "ecosystem":
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-grid"/></svg>`
return "i-grid"
case "tools":
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-settings"/></svg>`
return "i-settings"
case "models":
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-wave"/></svg>`
return "i-wave"
case "passkey":
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-lock"/></svg>`
return "i-lock"
default:
return `<svg class=icon width="14" height="14"><use href="/ethos-icons.svg#i-search"/></svg>`
return "i-search"
}
}
@@ -242,63 +225,12 @@ func pageTitle(key string) string {
}
}
// shellTopHTML opens the shell and renders the top bar + sidebar.
// Usage: {{template "shellTop" "<page-key>"}}
const shellTopHTML = `{{define "shellTop"}}<!doctype html><meta charset=utf-8>
<meta name=viewport content="width=device-width,initial-scale=1,viewport-fit=cover">
<meta name=theme-color content="#14110D">
<link rel=manifest href=/manifest.json>
<title>maven · {{pageTitle .}}</title>
<link rel=stylesheet href=/ui.css>
<div class=shell data-app=maven>
<header class=topbar>
<div class=breadcrumbs>
<span class=current>{{pageTitle .}}</span>
</div>
<div class=topbar-actions>
<div class=search-trigger onclick="window.__openSearch()" role=button tabindex=0>
<svg class=icon width="13" height="13"><use href="/ethos-icons.svg#i-search"/></svg>
Search
<span class=kbd-hint>Ctrl+/</span>
</div>
<button class=icon-btn onclick="window.__openPalette()" title="Command Palette (Ctrl+K)" aria-label="Command Palette">
<svg class=icon width="15" height="15"><use href="/ethos-icons.svg#i-grid"/></svg>
</button>
<span class=conn-status>
<span class="dot online" id=connDot></span>
</span>
</div>
</header>
<div class=shell-body>
<aside class=sidebar>
{{sidebarHTML .}}
</aside>
<main class=content>
{{end}}`
// shellBottomHTML closes the content area, inspector, and shell.
// Usage: {{template "shellBottom"}}
const shellBottomHTML = `{{define "shellBottom"}}
</main>
<aside class=inspector id=inspector>
<div class=inspector-inner>
<div class=inspector-header>
<span id=inspectorTitle>Details</span>
<button class=inspector-close onclick="closeInspector()" aria-label="Close inspector">&times;</button>
</div>
<div class=inspector-body id=inspectorBody></div>
</div>
</aside>
</div>
</div>
<script src=/mavweb.js></script>
{{end}}`
// shellFuncs returns the FuncMap shared by every server-rendered page template.
func shellFuncs() template.FuncMap {
return template.FuncMap{
"pageTitle": pageTitle,
"sidebarHTML": sidebarHTML,
"pageTitle": pageTitle,
"pageIcon": pageIcon,
"sidebarSections": func() any { return sidebarSections },
"ago": func(t time.Time) string {
if t.IsZero() {
return "never"
@@ -313,21 +245,21 @@ func shellFuncs() template.FuncMap {
// a small fetch loop refreshes the tables in place. html/template escapes the
// user text in facts/nudges. Read-only: browses the append-only store via
// CoreAPI, never writes — the store IS the audit trail, this just shows it.
var dashTmpl = template.Must(template.New("dash").Funcs(shellFuncs()).Parse(shellTopHTML + dashHTML + shellBottomHTML))
var dashTmpl = template.Must(template.New("dash").Funcs(shellFuncs()).Parse(shellHTML + dashHTML))
// ecosystemTmpl — read-only view of the Nexus/Praxis/Hexis siblings, whose only
// human surface is here (they ship no web UI of their own).
var ecosystemTmpl = template.Must(template.New("ecosystem").Funcs(shellFuncs()).Parse(shellTopHTML + ecosystemHTML + shellBottomHTML))
var ecosystemTmpl = template.Must(template.New("ecosystem").Funcs(shellFuncs()).Parse(shellHTML + ecosystemHTML))
// eventsTmpl — the unified intake journal (Vikunja #283), read-only. Same
// shape as trace.html and morning.html: server-rendered, refreshed on reload.
var eventsTmpl = template.Must(template.New("events").Funcs(shellFuncs()).Parse(shellTopHTML + eventsHTML + shellBottomHTML))
var eventsTmpl = template.Must(template.New("events").Funcs(shellFuncs()).Parse(shellHTML + eventsHTML))
// morningTmpl — read-only view of today's checklist state per configured
// morning routine (internal/morning). Same shape as trace.html: a plain
// server-rendered page, refreshed on reload — no live-update loop, since
// checklist state changes on the scale of minutes, not seconds.
var morningTmpl = template.Must(template.New("morning").Funcs(shellFuncs()).Parse(shellTopHTML + morningHTML + shellBottomHTML))
var morningTmpl = template.Must(template.New("morning").Funcs(shellFuncs()).Parse(shellHTML + morningHTML))
func noCache(h http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
@@ -512,7 +444,7 @@ func main() {
// /tools — the authed enable surface. maven proposes acts she can't run;
// this page is where a human reviews and enables them (proposed→enabled).
// Enabling is the boundary-moving act (DESIGN.md § Tool registration —
// Enabling is the boundary-moving act (docs/design.md § Tool registration —
// drafting is suggest, enabling is act), so it lives ONLY here,
// behind wg+nginx+auth — never the voice/chat path.
mux.HandleFunc("/tools", func(w http.ResponseWriter, r *http.Request) {
@@ -747,107 +679,23 @@ var toolsTmpl = template.Must(template.New("tools").Funcs(func() template.FuncMa
m := shellFuncs()
m["join"] = strings.Join
return m
}()).Parse(shellTopHTML + toolsHTML + shellBottomHTML))
}()).Parse(shellHTML + toolsHTML))
const toolsHTML = `{{template "shellTop" "tools"}}
<h1>Tools</h1>
<p class=hint>enabling requires step-up <a href=/auth/passkey>assert a passkey</a> first.</p>
{{if .Msg}}<div class="msg msg-ok">{{.Msg}}</div>{{end}}
<section class=card>
<h2 class=card-title>proposed <span class=badge>{{len .Proposed}}</span></h2>
{{if .Proposed}}<p class=hint>maven drafted these from acts she couldn't run. Fill the command (argv, space-separated) and enable. A row in an <code>mcp:</code> scope came from an MCP server and already knows what it calls check the command, then enable.</p>
<div class=scroll><table><tr><th>name</th><th>scope</th><th>from utterance</th><th>enable as</th></tr>
{{range .Proposed}}<tr>
<td><code>{{.Name}}</code></td><td><span class=badge>{{.Scope}}</span></td><td>{{.Utterance}}</td>
<td><form method=post action=/tools>
<input type=hidden name=name value="{{.Name}}">
<input type=hidden name=scope value="{{.Scope}}">
<input type=hidden name=action value=enable>
<input type=text name=cmd class=input-wide placeholder="systemctl restart" value="{{join .Cmd " "}}" required>
<label><input type=checkbox name=destructive {{if .Destructive}}checked{{end}}> destructive</label>
<button class=btn>enable</button></form>
<form method=post action=/tools class=inline-form>
<input type=hidden name=name value="{{.Name}}">
<input type=hidden name=action value=dismiss>
<button class="btn btn-muted">dismiss</button></form></td>
</tr>{{end}}</table></div>
{{else}}<div class=empty>
<svg class=icon width="20" height="20"><use href="/ethos-icons.svg#i-search"/></svg>
<div>no proposed tools</div>
<div class=hint>maven will propose tools here when she needs help running an action</div>
</div>{{end}}
</section>
<section class=card>
<h2 class=card-title>enabled <span class=badge>{{len .Enabled}}</span></h2>
{{if .Enabled}}<div class=scroll><table><tr><th>name</th><th>scope</th><th>command</th><th></th><th></th></tr>
{{range .Enabled}}<tr><td><code>{{.Name}}</code></td><td><span class=badge>{{.Scope}}</span></td><td><code>{{join .Cmd " "}}</code></td>
<td>{{if .Destructive}}<span class=red>destructive</span>{{end}}</td>
<td><form method=post action=/tools class=inline-form>
<input type=hidden name=name value="{{.Name}}">
<input type=hidden name=scope value="{{.Scope}}">
<input type=hidden name=action value=disable>
<button class=btn>disable</button></form></td></tr>{{end}}</table></div>
{{else}}<div class=empty>
<svg class=icon width="20" height="20"><use href="/ethos-icons.svg#i-settings"/></svg>
<div>no tools enabled</div>
<div class=hint>enable proposed tools above, or ask maven to configure one</div>
</div>{{end}}
</section>
<section class=card>
<h2 class=card-title>MCP servers <span class=badge>{{len .MCP}}</span></h2>
{{if .MCP}}<p class=hint>servers she connects OUT to. Their tools appear above as proposals a configured server is a place she may look, not a capability she has. A <code>stdio</code> target is a process on this box; an <code>http</code> one on a loopback or LAN address is inside the network, so treat its tools accordingly.</p>
<div class=scroll><table><tr><th>name</th><th>transport</th><th>target</th><th>state</th><th>tools</th></tr>
{{range .MCP}}<tr><td><code>{{.Name}}</code></td><td><span class=badge>{{.Transport}}</span></td><td><code>{{.Target}}</code></td>
<td>{{if .Connected}}connected{{if .Server}} {{.Server}}{{end}}{{else}}<span class=red>down</span>{{if .Err}} {{.Err}}{{end}}{{end}}</td>
<td>{{.Tools}}</td></tr>{{end}}</table></div>
{{else}}<div class=empty>
<svg class=icon width="20" height="20"><use href="/ethos-icons.svg#i-settings"/></svg>
<div>no MCP servers configured</div>
<div class=hint>add an <code>mcp.servers</code> block to mavend.json to let her use an external tool server</div>
</div>{{end}}
</section>
{{template "shellBottom"}}`
var historyTmpl = template.Must(template.New("history").Funcs(shellFuncs()).Parse(shellHTML + historyHTML))
// routinesHTML — proposed routine review surface. One row per thing maven
var notificationsTmpl = template.Must(template.New("notifications").Funcs(shellFuncs()).Parse(shellHTML + notificationsHTML))
var remindersTmpl = template.Must(template.New("reminders").Funcs(shellFuncs()).Parse(shellHTML + remindersHTML))
var passkeyTmpl = template.Must(template.New("passkey").Funcs(shellFuncs()).Parse(shellHTML + passkeyPageHTML))
var voiceTmpl = template.Must(template.New("voice").Funcs(shellFuncs()).Parse(shellHTML + voiceHTML))
var tasksTmpl = template.Must(template.New("tasks").Funcs(shellFuncs()).Parse(shellHTML + tasksHTML))
// routinesTmpl — the proposed-routine review surface. One row per thing maven
// noticed, in her words, with at most two actions: accept or dismiss.
const routinesHTML = `{{template "shellTop" "routines"}}
<h1>Routines</h1>
{{if .Msg}}<div class="msg msg-ok">{{.Msg}}</div>{{end}}
<section class=card>
<h2 class=card-title>noticed <span class=badge>{{len .Proposed}}</span></h2>
{{if .Proposed}}<div class=scroll><table><tr><th>maven noticed</th><th>when</th><th></th><th></th></tr>
{{range .Proposed}}<tr>
<td>{{.Phrase}}</td><td class=muted>{{.Noticed}}</td>
<td><form method=post action=/routines class=inline-form>
<input type=hidden name=id value="{{.ID}}">
<input type=hidden name=action value=accept>
<button class=btn>accept</button></form></td>
<td><form method=post action=/routines class=inline-form>
<input type=hidden name=id value="{{.ID}}">
<input type=hidden name=action value=dismiss>
<button class="btn btn-muted">dismiss</button></form></td>
</tr>{{end}}</table></div>
{{else}}<div class=empty>
<svg class=icon width="20" height="20"><use href="/ethos-icons.svg#i-wave"/></svg>
<div>no proposed routines</div>
<div class=hint>maven will propose routines here when she detects a recurring pattern</div>
</div>{{end}}
</section>
{{template "shellBottom"}}`
var historyTmpl = template.Must(template.New("history").Funcs(shellFuncs()).Parse(shellTopHTML + historyHTML + shellBottomHTML))
var notificationsTmpl = template.Must(template.New("notifications").Funcs(shellFuncs()).Parse(shellTopHTML + notificationsHTML + shellBottomHTML))
var remindersTmpl = template.Must(template.New("reminders").Funcs(shellFuncs()).Parse(shellTopHTML + remindersHTML + shellBottomHTML))
var passkeyTmpl = template.Must(template.New("passkey").Funcs(shellFuncs()).Parse(shellTopHTML + passkeyPageHTML + shellBottomHTML))
var voiceTmpl = template.Must(template.New("voice").Funcs(shellFuncs()).Parse(shellTopHTML + voiceHTML + shellBottomHTML))
var tasksTmpl = template.Must(template.New("tasks").Funcs(shellFuncs()).Parse(shellTopHTML + tasksHTML + shellBottomHTML))
var routinesTmpl = template.Must(template.New("routines").Funcs(shellFuncs()).Parse(shellTopHTML + routinesHTML + shellBottomHTML))
var routinesTmpl = template.Must(template.New("routines").Funcs(shellFuncs()).Parse(shellHTML + routinesHTML))
var traceTmpl = template.Must(template.New("trace").Funcs(func() template.FuncMap {
m := shellFuncs()
@@ -859,7 +707,7 @@ var traceTmpl = template.Must(template.New("trace").Funcs(func() template.FuncMa
}
m["join"] = strings.Join
return m
}()).Parse(shellTopHTML + traceHTML + shellBottomHTML))
}()).Parse(shellHTML + traceHTML))
func handleHistory(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
if core == nil {
@@ -893,12 +741,59 @@ func handleNotifications(w http.ResponseWriter, r *http.Request, core ipc.CoreAP
http.Error(w, "notifications error: "+err.Error(), http.StatusBadGateway)
return
}
// The outbox, on the page that already answers "what did she send".
// A failed or dropped attempt is why she went quiet, and until now it was
// recorded and unreadable (Vikunja #390). Filter with ?status=dropped.
status := r.URL.Query().Get("status")
attempts, err := core.DeliveryAttempts(ctx, status, 50)
if err != nil {
// The nudge list is still worth showing, so this is a note on the page
// rather than a dead page.
log.Printf("notifications: delivery attempts: %v", err)
}
w.Header().Set("Content-Type", "text/html; charset=utf-8")
if err := notificationsTmpl.Execute(w, map[string]any{"Nudges": nudges}); err != nil {
if err := notificationsTmpl.Execute(w, map[string]any{
"Nudges": nudges,
"Attempts": deliveryRows(attempts),
"Status": status,
}); err != nil {
log.Printf("notifications template: %v", err)
}
}
// deliveryRow is one outbox line, with every timestamp already formatted so
// the template holds no date logic — same shape as taskRow.
type deliveryRow struct {
Kind string
Target string
Channel string
Status string
Created string
Completed string
}
func deliveryRows(as []ipc.DeliveryAttempt) []deliveryRow {
out := make([]deliveryRow, 0, len(as))
for _, a := range as {
target := a.Rule
if target == "" && a.ReminderID != 0 {
target = "reminder #" + strconv.FormatInt(a.ReminderID, 10)
}
row := deliveryRow{
Kind: a.Kind,
Target: target,
Channel: a.Channel,
Status: a.Status,
Created: a.Created.Format("02.01 15:04"),
}
if a.Completed != nil {
row.Completed = a.Completed.Format("15:04")
}
out = append(out, row)
}
return out
}
func handleReminders(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
if core == nil {
http.Error(w, "reminders disabled (no -core)", http.StatusServiceUnavailable)
@@ -1125,9 +1020,10 @@ func fmtTaskDate(t *time.Time) string {
return t.Local().Format("02 Jan")
}
// routineRow is one line on the page: what maven noticed, in her words, and
// how long ago she noticed it.
type routineRow struct {
// routineView is one line on the page: what maven noticed, in her words, and
// how long ago she noticed it. A view model, not a database row — the template
// never formats an interval or a timestamp itself.
type routineView struct {
ID int64
Phrase string
Noticed string
@@ -1189,34 +1085,31 @@ func handleRoutines(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI, se
w.Header().Set("Content-Type", "text/html; charset=utf-8")
if err := routinesTmpl.Execute(w, struct {
Msg string
Proposed []routineRow
}{msg, routineRows(proposed)}); err != nil {
Proposed []routineView
}{msg, toRoutineViews(proposed)}); err != nil {
log.Printf("routines render: %v", err)
}
}
// routineRows turns the wire rows into display rows. The phrase comes from
// toRoutineViews turns the wire rows into view models. The phrase comes from
// pattern.PhraseRoutine so the page says the same thing maven's voice says.
func routineRows(rs []ipc.ProposedRoutine) []routineRow {
out := make([]routineRow, 0, len(rs))
func toRoutineViews(rs []ipc.ProposedRoutine) []routineView {
out := make([]routineView, 0, len(rs))
for _, r := range rs {
p := pattern.ProposedRoutine{Action: r.Action, Object: r.Object, IntervalDays: r.IntervalDays}
noticed := "just now"
if r.CreatedTs > 0 {
noticed = time.Since(time.UnixMilli(r.CreatedTs)).Round(time.Minute).String() + " ago"
}
out = append(out, routineRow{ID: r.ID, Phrase: pattern.PhraseRoutine(&p), Noticed: noticed})
out = append(out, routineView{ID: r.ID, Phrase: pattern.PhraseRoutine(&p), Noticed: noticed})
}
return out
}
// acceptRoutine creates the recurring reminder for a proposal, then marks the
// proposal accepted and links the reminder to it. Weekly patterns get a cron
// expression; any other interval fires once.
//
// TODO(vikunja#46): this mirrors the voice accept path in cmd/mavend/voice.go.
// When the tick loop learns to read accepted proposals directly, both callers
// should hand off to one place in core instead of each building a reminder.
// acceptRoutine marks a proposal accepted. This page is the ONLY surface that
// may do it (Vikunja #367): accepting gives the tick loop a standing new
// reason to speak, which DESIGN.md puts at layer 3, and the button here is
// behind step-up. Voice can park the question and dismiss, never accept.
func acceptRoutine(ctx context.Context, core ipc.CoreAPI, id int64) error {
proposed, err := core.ListProposedRoutines(ctx)
if err != nil {
@@ -1605,32 +1498,7 @@ func handlePTT(w http.ResponseWriter, r *http.Request, voiceAddr string, session
// chatTmpl — plain text conversation interface. No JS: form POSTs to /api/chat
// and the handler redirects back to /chat with the response.
var chatTmpl = template.Must(template.New("chat").Funcs(shellFuncs()).Parse(shellTopHTML + chatPageHTML + shellBottomHTML))
const chatPageHTML = `{{template "shellTop" "chat"}}
<h1>Chat</h1>
<section class=card>
{{if .Error}}<div class="msg msg-err">{{.Error}}</div>{{end}}
<div class="scroll chat-scroll" id=chatHistory>
{{range .Messages}}
<div class="chat-msg {{.Role}}"><strong>{{if eq .Role "user"}}you{{else}}maven{{end}}:</strong> {{.Text}}</div>
{{else}}
<div class=empty>
<svg class=icon width="20" height="20"><use href="/ethos-icons.svg#i-message"/></svg>
<div>start a conversation</div>
</div>
{{end}}
</div>
<form method=post action=/api/chat class=chat-form>
<input type=text name=text class=input-wide placeholder="type a message..." required autofocus>
<button class=btn>send</button>
</form>
</section>
<script>
var ch = document.getElementById('chatHistory');
if(ch) ch.scrollTop = ch.scrollHeight;
</script>
{{template "shellBottom"}}`
var chatTmpl = template.Must(template.New("chat").Funcs(shellFuncs()).Parse(shellHTML + chatPageHTML))
// handleChatPage renders the chat conversation page.
func handleChatPage(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
@@ -1681,7 +1549,13 @@ func handleChatAPI(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI, ses
http.Redirect(w, r, "/chat", http.StatusSeeOther)
return
}
reply, err := core.Chat(r.Context(), text)
// One conversation id for the whole web chat, and a different one from
// telegram or the mic. A parked question belongs to the reach that was
// asked; before this, a clarify nobody answered on the web ate the next
// utterance spoken at the mic (Vikunja #466). This server has no
// per-browser session, so every browser tab is the same conversation —
// which is right for a single-owner box.
reply, err := core.Chat(r.Context(), "web", text)
if err != nil {
log.Printf("chat api: %v", err)
http.Redirect(w, r, "/chat", http.StatusSeeOther)
+1 -1
View File
@@ -31,7 +31,7 @@ type modelController interface {
SwapModel(ctx context.Context, req ipc.SwapModelReq) (ipc.SwapModelResp, error)
}
var modelsTmpl = template.Must(template.New("models").Funcs(shellFuncs()).Parse(shellTopHTML + modelsHTML + shellBottomHTML))
var modelsTmpl = template.Must(template.New("models").Funcs(shellFuncs()).Parse(shellHTML + modelsHTML))
const modelsHTML = `{{template "shellTop" "models"}}
<h1>Resident model</h1>
+22
View File
@@ -14,5 +14,27 @@
<div>no notifications yet</div>
<div class=hint>check back later or ask maven a question</div>
</div>{{end}}
<h2>Delivery outbox</h2>
<p class=hint>
every send is recorded before it leaves, so a failure is visible rather than silent.
<a href="/notifications">all</a> ·
<a href="/notifications?status=dropped">dropped</a> ·
<a href="/notifications?status=failed">failed</a> ·
<a href="/notifications?status=pending">pending</a> ·
<a href="/notifications?status=unknown">unknown</a>
</p>
{{if .Attempts}}<div class=scroll><table>
<tr><th>started</th><th>kind</th><th>rule</th><th>channel</th><th>status</th><th>finished</th></tr>
{{range .Attempts}}<tr>
<td class=hint>{{.Created}}</td>
<td>{{.Kind}}</td>
<td class=key>{{.Target}}</td>
<td><span class=badge>{{.Channel}}</span></td>
<td class={{.Status}}>{{.Status}}</td>
<td class=hint>{{.Completed}}</td>
</tr>{{end}}</table></div>
{{else}}<div class=empty>
<div>no delivery attempts{{if .Status}} with status {{.Status}}{{end}}</div>
</div>{{end}}
{{template "shellBottom"}}
</html>
+24
View File
@@ -0,0 +1,24 @@
{{template "shellTop" "routines"}}
<h1>Routines</h1>
{{if .Msg}}<div class="msg msg-ok">{{.Msg}}</div>{{end}}
<section class=card>
<h2 class=card-title>noticed <span class=badge>{{len .Proposed}}</span></h2>
{{if .Proposed}}<div class=scroll><table><tr><th>maven noticed</th><th>when</th><th></th><th></th></tr>
{{range .Proposed}}<tr>
<td>{{.Phrase}}</td><td class=muted>{{.Noticed}}</td>
<td><form method=post action=/routines class=inline-form>
<input type=hidden name=id value="{{.ID}}">
<input type=hidden name=action value=accept>
<button class=btn>accept</button></form></td>
<td><form method=post action=/routines class=inline-form>
<input type=hidden name=id value="{{.ID}}">
<input type=hidden name=action value=dismiss>
<button class="btn btn-muted">dismiss</button></form></td>
</tr>{{end}}</table></div>
{{else}}<div class=empty>
<svg class=icon width="20" height="20"><use href="/ethos-icons.svg#i-wave"/></svg>
<div>no proposed routines</div>
<div class=hint>maven will propose routines here when she detects a recurring pattern</div>
</div>{{end}}
</section>
{{template "shellBottom"}}
+49
View File
@@ -0,0 +1,49 @@
{{define "shellTop"}}<!doctype html><meta charset=utf-8>
<meta name=viewport content="width=device-width,initial-scale=1,viewport-fit=cover">
<meta name=theme-color content="#14110D">
<link rel=manifest href=/manifest.json>
<title>maven · {{pageTitle .}}</title>
<link rel=stylesheet href=/ui.css>
<div class=shell data-app=maven>
<header class=topbar>
<div class=breadcrumbs>
<span class=current>{{pageTitle .}}</span>
</div>
<div class=topbar-actions>
<div class=search-trigger onclick="window.__openSearch()" role=button tabindex=0>
<svg class=icon width="13" height="13"><use href="/ethos-icons.svg#i-search"/></svg>
Search
<span class=kbd-hint>Ctrl+/</span>
</div>
<button class=icon-btn onclick="window.__openPalette()" title="Command Palette (Ctrl+K)" aria-label="Command Palette">
<svg class=icon width="15" height="15"><use href="/ethos-icons.svg#i-grid"/></svg>
</button>
<span class=conn-status>
<span class="dot online" id=connDot></span>
</span>
</div>
</header>
<div class=shell-body>
<aside class=sidebar>
{{template "sidebar" .}}
</aside>
<main class=content>
{{end}}{{/* sidebar — the section list, with the active page's link marked. The dot
argument is the page key the page passed to shellTop. */}}{{define "sidebar"}}{{$active := .}}{{range sidebarSections}}<div class=sidebar-section><div class=sidebar-label>{{.Label}}</div>
{{- range .Pages}}<a href="{{.URL}}"{{if eq .Key $active}} class="active"{{end}}><span class=icon><svg class=icon width="14" height="14"><use href="/ethos-icons.svg#{{pageIcon .Key}}"/></svg></span><span>{{.Label}}</span></a>
{{- end}}</div>
{{end}}{{end}}{{define "shellBottom"}}
</main>
<aside class=inspector id=inspector>
<div class=inspector-inner>
<div class=inspector-header>
<span id=inspectorTitle>Details</span>
<button class=inspector-close onclick="closeInspector()" aria-label="Close inspector">&times;</button>
</div>
<div class=inspector-body id=inspectorBody></div>
</div>
</aside>
</div>
</div>
<script src=/mavweb.js></script>
{{end}}
+58
View File
@@ -0,0 +1,58 @@
{{template "shellTop" "tools"}}
<h1>Tools</h1>
<p class=hint>enabling requires step-up — <a href=/auth/passkey>assert a passkey</a> first.</p>
{{if .Msg}}<div class="msg msg-ok">{{.Msg}}</div>{{end}}
<section class=card>
<h2 class=card-title>proposed <span class=badge>{{len .Proposed}}</span></h2>
{{if .Proposed}}<p class=hint>maven drafted these from acts she couldn't run. Fill the command (argv, space-separated) and enable. A row in an <code>mcp:</code> scope came from an MCP server and already knows what it calls — check the command, then enable.</p>
<div class=scroll><table><tr><th>name</th><th>scope</th><th>from utterance</th><th>enable as</th></tr>
{{range .Proposed}}<tr>
<td><code>{{.Name}}</code></td><td><span class=badge>{{.Scope}}</span></td><td>{{.Utterance}}</td>
<td><form method=post action=/tools>
<input type=hidden name=name value="{{.Name}}">
<input type=hidden name=scope value="{{.Scope}}">
<input type=hidden name=action value=enable>
<input type=text name=cmd class=input-wide placeholder="systemctl restart" value="{{join .Cmd " "}}" required>
<label><input type=checkbox name=destructive {{if .Destructive}}checked{{end}}> destructive</label>
<button class=btn>enable</button></form>
<form method=post action=/tools class=inline-form>
<input type=hidden name=name value="{{.Name}}">
<input type=hidden name=action value=dismiss>
<button class="btn btn-muted">dismiss</button></form></td>
</tr>{{end}}</table></div>
{{else}}<div class=empty>
<svg class=icon width="20" height="20"><use href="/ethos-icons.svg#i-search"/></svg>
<div>no proposed tools</div>
<div class=hint>maven will propose tools here when she needs help running an action</div>
</div>{{end}}
</section>
<section class=card>
<h2 class=card-title>enabled <span class=badge>{{len .Enabled}}</span></h2>
{{if .Enabled}}<div class=scroll><table><tr><th>name</th><th>scope</th><th>command</th><th></th><th></th></tr>
{{range .Enabled}}<tr><td><code>{{.Name}}</code></td><td><span class=badge>{{.Scope}}</span></td><td><code>{{join .Cmd " "}}</code></td>
<td>{{if .Destructive}}<span class=red>destructive</span>{{end}}</td>
<td><form method=post action=/tools class=inline-form>
<input type=hidden name=name value="{{.Name}}">
<input type=hidden name=scope value="{{.Scope}}">
<input type=hidden name=action value=disable>
<button class=btn>disable</button></form></td></tr>{{end}}</table></div>
{{else}}<div class=empty>
<svg class=icon width="20" height="20"><use href="/ethos-icons.svg#i-settings"/></svg>
<div>no tools enabled</div>
<div class=hint>enable proposed tools above, or ask maven to configure one</div>
</div>{{end}}
</section>
<section class=card>
<h2 class=card-title>MCP servers <span class=badge>{{len .MCP}}</span></h2>
{{if .MCP}}<p class=hint>servers she connects OUT to. Their tools appear above as proposals — a configured server is a place she may look, not a capability she has. A <code>stdio</code> target is a process on this box; an <code>http</code> one on a loopback or LAN address is inside the network, so treat its tools accordingly.</p>
<div class=scroll><table><tr><th>name</th><th>transport</th><th>target</th><th>state</th><th>tools</th></tr>
{{range .MCP}}<tr><td><code>{{.Name}}</code></td><td><span class=badge>{{.Transport}}</span></td><td><code>{{.Target}}</code></td>
<td>{{if .Connected}}connected{{if .Server}} — {{.Server}}{{end}}{{else}}<span class=red>down</span>{{if .Err}} — {{.Err}}{{end}}{{end}}</td>
<td>{{.Tools}}</td></tr>{{end}}</table></div>
{{else}}<div class=empty>
<svg class=icon width="20" height="20"><use href="/ethos-icons.svg#i-settings"/></svg>
<div>no MCP servers configured</div>
<div class=hint>add an <code>mcp.servers</code> block to mavend.json to let her use an external tool server</div>
</div>{{end}}
</section>
{{template "shellBottom"}}
+34 -3
View File
@@ -73,6 +73,37 @@ What it does and does not do:
- feed notes are **not** part of recall. "что я говорил про X" searches what he
said; headlines are read back only by asking about the feeds.
### Searching the web (`search`, on in this deploy)
`deploy/mavend.json` ships a `search` block, so a question that is not about him
reaches a self-hosted SearXNG before it reaches the ZIMs. Delete the block and
no query leaves the LAN again. The shipped shape:
```json
"search": {
"url": "http://searxng:9563",
"max_results": 4,
"snippet_runes": 1500,
"language": "auto",
"timeout": "8s"
}
```
- the instance needs `json` in its `search.formats` (settings.yml). A stock
SearXNG answers 403 to `format=json`, and then every search fails;
- the question goes out **verbatim**, in the language he asked it. There is no
rewriter here, unlike Kiwix: SearXNG ranks through real engines;
- only the query string leaves the box. `internal/websearch` cannot read the
store, so no note, fact, persona block or history can travel with a search;
- a question about him never becomes a query. The personal boundary in the
query chain stops the walk above this source;
- this runs **before** Kiwix. A live search reads what is true today and the
ZIMs read what was true when they were built, so the ZIMs are the fallback:
an empty result, an unreachable instance or a dead line falls through to
them and she never says the search failed;
- `engines` narrows the search to named engines, e.g. `"duckduckgo,wikipedia"`.
Empty means whatever the instance has enabled.
### Reading a page (`crawl`, also off by default)
There is no `crawl` block either, so no page is fetched. Two halves, separately
@@ -95,9 +126,9 @@ switched:
a fallback and not a habit;
- `watches` re-reads a fixed list on its interval and writes a note when the
text changed. Like the feeds, it announces nothing;
- the answer path sits behind his memory and his notes, and ahead of the model
answering from what it remembers. Kiwix is not wired into the chain yet. A
local read costs nothing, so anything local goes first;
- the answer path sits behind his memory, his notes, the web search and the
ZIMs, and ahead of the model answering from what it remembers. A page he
named is an instruction, so it is read last and only when he named one;
- `robots.txt` is fetched first and obeyed with no override; a `Disallow` is a
refusal she says out loud. `Crawl-delay` is waited out before the page is
fetched, and a delay longer than the turn fails the read instead of hanging
+127 -2
View File
@@ -5,18 +5,143 @@
"socket_path": "/run/maven/mavend.sock",
"state_dir": "/var/lib/maven",
"//disabled_rules": [
"Nudge rules that are not wired at all. Names come from loop.DefaultRules:",
"water, meal, break, service_down, netdata_critical.",
"service_down is back on: mavpoll now writes one fact per kuma monitor",
"(service_down:<name>), so the nudge names the service and pausing a monitor",
"in kuma silences that monitor. It is also edge-triggered, so a service that",
"stays down is one nudge, not one every fifteen minutes."
],
"disabled_rules": [],
"phraser": {
"model_path": "/opt/maven/models/llm/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf",
"bin_path": "llama-server",
"n_gpu_layers": 99,
"n_ctx": 4096,
"cache_ram_mib": 512,
"timeout": "60s",
"llm_nudges": false
},
"telegram": {
"bot_token": "${TELEGRAM_BOT_TOKEN}",
"chat_id": "${TELEGRAM_CHAT_ID}"
"chat_id": "${TELEGRAM_CHAT_ID}",
"//proxy": [
"api.telegram.org is not reachable directly from this box, so every send",
"timed out. The relay is the x-ui socks inbound on the host, port 10808;",
"192.168.240.1 is the maven_default bridge gateway, which is how a",
"container addresses the host. mavend is on that network.",
"This needs a matching ufw rule or the container's SYN is dropped:",
" ufw allow from 192.168.240.0/20 to any port 10808 proto tcp"
],
"proxy": "socks5://192.168.240.1:10808"
},
"//workstation": [
"The big model on the desk PC (workpc, 7900 GRE 16GB), fronted by",
"mavgpud on port 8080. It runs gemma-4-12b and it is preferred over the",
"resident Qwen3-1.7B for routing and replies whenever the card is free.",
"The machine is never assumed up: it sleeps, and the card is often held by",
"a CPT run, in which case mavgpud answers 503 and Maven falls back to the",
"resident model without saying so. Deleting this block restores exactly",
"the behaviour homesrv had before it existed.",
"Addressed by LAN address, not container name: mavgpud runs on another",
"machine and there is no shared docker network to name it on."
],
"workstation": {
"url": "http://192.168.1.105:8080",
"probe": "15s",
"timeout": "90s"
},
"//search": [
"The live web, searched after his own notes and before Kiwix. Only the",
"query string leaves the box — never a note, a fact, the persona block or",
"the history — and a question about him never reaches here at all.",
"The instance must have `json` in search.formats (settings.yml); a stock",
"SearXNG answers 403 to format=json and every search then fails. It is",
"addressed by container name, so it needs the same maven_default",
"attachment kiwix has, and it must listen on 9563: 8080 is taken several",
"times over on this box. No instance reachable ⇒ she falls through to the",
"ZIMs and never says the search failed."
],
"search": {
"url": "http://searxng:9563",
"max_results": 4,
"snippet_runes": 1500,
"language": "auto",
"timeout": "8s"
},
"//kiwix": [
"The offline encyclopedia, searched after his own notes and before anything",
"on the network. kiwix-server publishes 8034 on loopback only, so a container",
"cannot reach it by address; it is attached to the maven_default network",
"instead and addressed by container name. That attachment is imperative and",
"does not survive recreating the kiwix stack — make it declarative there:",
" networks: [default, maven_default] # maven_default: external: true",
"The book is the catalog name from the /content/... href in",
"/catalog/v2/entries, not the display title. Others on the box:",
"ifixit_en_all_2025-06, devdocs_en_ansible_2025-10."
],
"kiwix": {
"url": "http://kiwix-server:8080",
"book": "wikipedia_en_all_maxi_2026-02",
"max_results": 5,
"snippet_runes": 1500
},
"//morning_routines": [
"The daily checklist (Vikunja #280). Each item is done when its fact_key",
"gets a non-voided fact inside the window, so 'выпил воды' closes water and",
"nothing has to be ticked by hand. nudge_at fires once, at the end of the",
"window, and only for what is still open. Weekdays empty = every day."
],
"morning_routines": [
{
"name": "утро",
"window_start": "08:00",
"window_end": "11:00",
"nudge_at": "10:30",
"severity": 1,
"items": [
{ "key": "medicine", "fact_key": "medicine", "label": "лекарство" },
{ "key": "water", "fact_key": "water", "label": "вода" },
{ "key": "pets", "fact_key": "pets", "label": "покормить кота" }
]
}
],
"//feeds": [
"RSS reading (Vikunja #258). Every item lands as a note with source",
"rss:<name>, which is also what puts entries in the intake journal that",
"/events reads. Only the feed URL leaves the box.",
"This is a starting pair, not a curated set — trim or extend it."
],
"feeds": {
"poll_interval": "30m",
"max_items": 5,
"max_age": "24h",
"sources": [
{ "name": "lwn", "url": "https://lwn.net/headlines/newrss", "category": "технологии" },
{ "name": "archlinux", "url": "https://archlinux.org/feeds/news/", "category": "технологии" }
]
},
"//crawl": [
"Reading a web page (Vikunja #259). on_demand answers 'посмотри <URL>'.",
"No allow_hosts, so any public host he names is readable; private",
"addresses are refused unconditionally by internal/webfetch and do not",
"need listing. Setting allow_hosts here would also narrow on-demand,",
"which is the point of leaving it empty."
],
"crawl": {
"on_demand": true,
"timeout": "10s",
"max_runes": 4000
},
"digest": {
@@ -62,7 +187,7 @@
"timeout": "400ms",
"rate": 100,
"max_hosts": 256,
"enabled": false
"enabled": true
},
"nexus": { "url": "http://nexus:9740" },
+32
View File
@@ -0,0 +1,32 @@
{
"listen": ":8080",
"llama_addr": "127.0.0.1:10000",
"llama_bin": "llama-server",
"llama_args": [
"-m", "/mnt/D/AI/gemma4/gemma-4-12B-it-qat-UD-Q4_K_XL.gguf",
"-md", "/mnt/D/AI/gemma4/mtp-gemma-4-12B-it-BF16.gguf",
"-ngl", "99",
"-fa", "on",
"-np", "1",
"--host", "127.0.0.1",
"--port", "10000",
"--ctx-size", "32768",
"--threads", "6",
"--batch-size", "2048",
"--ubatch-size", "512",
"--jinja",
"--chat-template-kwargs", "{\"enable_thinking\":false}",
"--spec-type", "draft-mtp",
"--spec-draft-n-max", "2"
],
"kfd_root": "/sys/class/kfd/kfd/proc",
"drm_device": "/sys/class/drm/card1/device",
"poll": "1s",
"idle_timeout": "15m",
"stop_grace": "20s",
"min_free_vram_bytes": 10737418240,
"evict_after_polls": 2,
"start_after_polls": 5
}
+28
View File
@@ -0,0 +1,28 @@
[Unit]
# Runs on the workstation (bugmachine), not on homesrv. Install as a systemd
# user unit and turn on lingering, so the card is supervised after a reboot
# with nobody logged in:
#
# scp mavgpud workpc:~/.local/bin/mavgpud
# scp deploy/mavgpud.json workpc:~/.config/mavgpud.json
# scp deploy/mavgpud.service workpc:~/.config/systemd/user/mavgpud.service
# ssh workpc 'systemctl --user daemon-reload && systemctl --user enable --now mavgpud'
# sudo loginctl enable-linger kami
Description=Maven GPU supervisor (holds llama-server while the card is free)
After=network.target
[Service]
ExecStart=%h/.local/bin/mavgpud -config %h/.config/mavgpud.json
Restart=always
RestartSec=5
# The card must come back when the supervisor goes down. mavgpud stops
# llama-server on SIGTERM, so give it longer than stop_grace to do that.
KillSignal=SIGTERM
TimeoutStopSec=60
# llama-server aborts inside its own static teardown on SIGTERM, so every
# routine yield used to write a multi-gigabyte core into systemd-coredump
# (Vikunja #491). Yielding is meant to happen several times a day.
LimitCORE=0
[Install]
WantedBy=default.target
+9
View File
@@ -79,7 +79,16 @@ services:
<<: *image
# voice.bind is 0.0.0.0:9100 in deploy/mavend.json so mavweb can reach it
# cross-container. Verified 2026-07-06.
# -ambient-token turns on POST /api/ambient (Vikunja #126): the phone posts
# notification text, mavweb keeps only a meeting time. Empty ⇒ no route at
# all, which is what a missing MAVEN_AMBIENT_TOKEN gives. The value comes
# from the gitignored .env docker compose reads for interpolation, NOT from
# an env_file — flags are interpolated before any service env exists.
# Weakness worth naming: mavweb takes this as a flag, so it is visible in
# `ps` inside this container, unlike the zenmoney and IMAP secrets which are
# read from files.
command: ["mavweb", "-addr", ":9201", "-voice", "mavend:9100", "-core", "/run/maven/mavend.sock",
"-ambient-token", "${MAVEN_AMBIENT_TOKEN:-}",
"-nexus", "http://nexus:9740", "-praxis", "http://praxis:8989", "-hexis", "http://hexis:9741"]
depends_on: [mavend]
# loopback-only on purpose: /tools defines+executes arbitrary argv and
@@ -1,6 +1,45 @@
# maven — feature ranking
> dated 2026-07-03. companion to `DESIGN.md` (folded from the former `maven.md`). ranks everything discussed post-repo-state against the infra blockers, not a replacement for the build order.
> **Archived 2026-08-02 (V-447).** The mandatory and easy tiers are now Vikunja tasks
> 449-458. Two of those closed immediately, because the ranking was stale. Quiet hours
> (V-450) ship as `QuietHoursConfig` plus the care gate in `internal/loop/loop.go`.
> Schema migrations (V-451) ship as `internal/store/migrations.go` on `PRAGMA
> user_version`. The doable and epic tiers stay here because they are reasoning. Some of
> them exist only to record why something is not worth doing yet. Read this for the why,
> not as a work queue, and check the code before believing a gap.
> dated 2026-07-03. companion to `docs/design.md` (folded from the former `maven.md`). ranks everything discussed post-repo-state against the infra blockers, not a replacement for the build order.
---
## what already shipped (checked against the code, 2026-08-02)
One month old and already wrong in ten places. Everything below is marked
built after reading the code, not the board. Read the tiers underneath with this list in
hand.
Infra 1 and 2, the two the ranking says block every feature, are both done. sqlcipher
at-rest ships as `Store.enc` plus `OpenEncrypted`, a tmpfs working copy re-encrypted on
`Close`, keyed from `db_key_env` in `deploy/mavend.json`. mavweb and mavcaldav are no
longer at zero coverage: seven test files under `cmd/mavweb`, including
`credentials_test.go` and `passkey_prf_test.go`, and two under `cmd/mavcaldav`. Infra 3
is stale in the other direction. There are still no systemd units, but the deploy is
`deploy/ecosystem/docker-compose.yml`, not scripts and tmux.
Doable tier, built: rule trace and explanation as `/trace` plus `internal/loop/explain.go`.
Recurring reminders as the cron column, `NextFireTs` and `RescheduleReminder`
(`internal/store/reminders.go:206`), so "fires once right now" is wrong. Stale-reminder
burst collapse as `collapseReminders` (`internal/loop/gather.go:209`). Revert as
`VoidLatestFact` (`internal/store/facts.go:300`). Digest mode as `internal/store/digest.go`.
Testing infra as the simulator and the eval lab (V-284, V-278). Passkey persistence as the
JSON-backed `credentialStore` in `cmd/mavweb/credentials.go`. Most integrations shipped as
their own QA tasks (V-246 mail, V-256 smarthome, V-258 rss, V-259 crawler).
Doable tier, still open: correx, systemd units, memory decay and duplicate detection
(nothing in `internal/memory` touches it), backup automation, import and export, barge-in.
Epic tier, unbuilt as ranked. `event.Bus` exists (`cmd/mavend/intake.go:57`) but it is the
intake journal from V-283, not the rewrite of facts into projections that this tier means.
---
@@ -20,7 +59,7 @@ nothing feature-level below should land before 12 are done. 34 can interle
### mandatory
things that block correctness or safety of stuff already shipped — not new capability, just closing gaps in existing design.
- **destructive-confirm policy** — open question in `DESIGN.md` § open questions, blocks correx and any new tool domain from having a coherent risk tier
- **destructive-confirm policy** — open question in `docs/design.md` § open questions, blocks correx and any new tool domain from having a coherent risk tier
- **quiet-hours definition** — open question, blocks proactive delivery being trustworthy
- **schema migrations** — sqlcipher rollout alone forces a schema touch. want this mechanism before that, not after.
@@ -30,7 +69,7 @@ cheap, no dependencies, no new invariants.
- **grocery / `list_items` table** — fourth append-only shape (item, status, list-tag), no predicate touches it, multi-adder just works for free
- **go.mod tidy**
- **capability model** (deepseek) — `homelab.docker.restart` instead of flat `tool→enabled`. cheap now, expensive to retrofit once tools surface passes ~15 entries. time-sensitive, not urgent.
- **conversation repair** — already free: `DESIGN.md` has "misroute correction = new centroid example," this is just naming the existing mechanism as a feature
- **conversation repair** — already free: `docs/design.md` has "misroute correction = new centroid example," this is just naming the existing mechanism as a feature
- **command history** — read-only query over existing facts, no new mechanism
- **clarification templates** — canned phrasing for the router's existing confidence-gate fallback, phraser-lane only
- **pronunciation dictionary** — tts config, no architecture
@@ -30,7 +30,7 @@ critical workflows: voice turn (mic→STT→route→tool/reply→TTS); proactive
(/dash /chat /tools /ecosystem)
current state: all 34 test packages pass; vet clean; `make test` still exits 1
(finding 2)
known failures: weak RU query routing — REARCH.md names the cause; the named fix
known failures: weak RU query routing — docs/rearchitecture.md names the cause; the named fix
is wired `nil`
maintenance burden: 4,518 lines of root markdown vs 33,319 lines of Go; 15 top-level
.md files, 3 of them dated session logs; several contradict
@@ -47,11 +47,11 @@ what's obsolete: llmrouter.go (built, tested, never wired); classifier seed-p
```
**Classification: healthy + misaligned.** Not fragile, not overbuilt, not abandoned. The
architecture in `REARCH.md` is sound and mostly *built* — it just is not *connected*.
architecture in `docs/rearchitecture.md` is sound and mostly *built* — it just is not *connected*.
## what it should become
The thing `REARCH.md` already describes, with the switch flipped and the drift removed:
The thing `docs/rearchitecture.md` already describes, with the switch flipped and the drift removed:
one resident small model doing both routing and phrasing, classifier demoted from the live
path to the failure floor, embedder demoted to RAG hint. No new architecture is needed.
**The gap is a config/wiring decision plus doc convergence, not a redesign.**
@@ -65,19 +65,19 @@ path to the failure floor, embedder demoted to RAG hint. No new architecture is
`architecture` / `repair`
**problem:** The most load-bearing design decision in the project is stated four different,
incompatible ways, and the code path `REARCH.md` calls "the linchpin" is disabled.
incompatible ways, and the code path `docs/rearchitecture.md` calls "the linchpin" is disabled.
**evidence** (all confirmed):
- `cmd/mavend/voice.go:211``rtr := buildRouter(emb, matcher, threshold, nil) // LLM router disabled`,
with comment *"the classifier handles routing reliably."*
- `REARCH.md:11` says the same classifier is *"the structural cause of 'she messes up
- `docs/rearchitecture.md:11` says the same classifier is *"the structural cause of 'she messes up
queries.'"* **The code comment and the design doc make opposite claims about the same
component.**
- `internal/router/llmrouter.go` (139 lines) + `llmrouter_test.go` — fully built and
tested, zero non-test callers.
- Model identity, four ways: docs say **Qwen3-1.7B** (`CLAUDE.md:6`, `REARCH.md:15`,
`SPEC.md:46`, `AGENTS.md:79`, `MAVEN_ECOSYSTEM_ARCHITECTURE.md:72`);
- Model identity, four ways: docs say **Qwen3-1.7B** (`CLAUDE.md:6`, `docs/rearchitecture.md:15`,
`SPEC.md:46`, `AGENTS.md:79`, `docs/ecosystem.md:72`);
`deploy/mavend.json:9` says **Qwen3.5-2B-UD-Q4_K_XL**; `models/llm/` on disk holds
**LFM2.5-1.2B-Thinking**; code comments in 5 files still say **LFM**.
- `deploy/mavend.json:11` sets `"n_gpu_layers": 99` while `CLAUDE.md:4` states the target
@@ -100,7 +100,7 @@ match. Delete nothing from `internal/router` yet — the classifier is the fallb
reconciliation, not a refactor. **Do not rewrite the router.**
**alternatives:** Delete `llmrouter.go` and commit to the classifier — only defensible if
the eval harness shows the classifier is actually adequate, which contradicts `REARCH.md`.
the eval harness shows the classifier is actually adequate, which contradicts `docs/rearchitecture.md`.
**risk:** Low-moderate. LLM route failures already fall through to the classifier
(`router.go:88-96`), so a bad model cannot break a turn. The real risk is CPU latency.
@@ -228,21 +228,21 @@ after finding 1, not before** — and skip it if it stays purely cosmetic.
**problem:** 15 root markdown files, 4,518 lines, several stale or superseded, at least
three pairs contradicting each other.
**evidence:** `ROADMAP.md` (759) + `MAVEN_ECOSYSTEM_ARCHITECTURE.md` (884) +
**evidence:** `ROADMAP.md` (759) + `docs/ecosystem.md` (884) +
`PROGRESS.md` (456) + `maven.md` (413) + `20-07-2026-BACKLOG.md` (396) +
`SESSION-05-07-2026.md` + `SESSION-06-07-2026.md` (477 combined) + `PLANS.md` (25) +
`START.md` + `SPEC.md` + `PROTOCOL.md`. `PROGRESS.md:61` annotates its own staleness:
*"Older LFM references below describe the currently deployed..."*. `REARCH.md` announces it
`docs/operations.md` + `SPEC.md` + `docs/protocol.md`. `PROGRESS.md:61` annotates its own staleness:
*"Older LFM references below describe the currently deployed..."*. `docs/rearchitecture.md` announces it
"supersedes" a model still described as current elsewhere.
**impact:** The doc set is the reason finding 1 exists. When five documents describe the
architecture, the code becomes the only trustworthy one — which defeats the purpose of
having them.
**recommended action:** Keep `CLAUDE.md` (agent contract), `REARCH.md` (target
architecture), `AGENTS.md` (recipes), `PROTOCOL.md` (wire format),
**recommended action:** Keep `CLAUDE.md` (agent contract), `docs/rearchitecture.md` (target
architecture), `AGENTS.md` (recipes), `docs/protocol.md` (wire format),
`20-07-2026-BACKLOG.md` (live queue). Delete the two `SESSION-*.md` and `PLANS.md` — git
history holds them. Fold `SPEC.md` + `maven.md` + `ROADMAP.md` into one `DESIGN.md` and
history holds them. Fold `SPEC.md` + `maven.md` + `ROADMAP.md` into one `docs/design.md` and
mark superseded sections instead of leaving them to read as current. Target ~1,500 lines.
---
@@ -320,7 +320,7 @@ expected maintenance gain: none over the incremental path
suggesting Vulkan offload is intended and working — but `CLAUDE.md` says CPU-only. Likely
the doc is stale, not the config; unverified.
- **Whether the classifier is genuinely adequate.** `voice.go:211` asserts it is;
`REARCH.md` asserts it is not. Both are claims, neither is measured. The uncommitted eval
`docs/rearchitecture.md` asserts it is not. Both are claims, neither is measured. The uncommitted eval
harness is the instrument to settle it — resolve before flipping the router, not after.
- **Whether wg+nginx+auth actually fronts 9201 in production.** Not in this repo. If it
does, finding 3 drops from "unauthenticated RCE" to "the control is not reproducible from
@@ -350,7 +350,7 @@ Key claims independently re-verified against the working tree; the verdict stand
over WireGuard on homesrv, this is hygiene, not an emergency — but the loopback bind
and startup warning are cheap insurance either way, so do them regardless.
- Finding 1's "flip the router" step should be gated harder on measurement.
`REARCH.md`'s claim that the classifier causes weak RU queries is itself unmeasured —
`docs/rearchitecture.md`'s claim that the classifier causes weak RU queries is itself unmeasured —
the review admits this under uncertainties, but the "repair now" ordering buries it.
Run `eval_scenarios_test.go` against both paths **before** deciding to flip, not
after. A 2B model on CPU may add enough latency that the classifier wins in practice
+66 -12
View File
@@ -1,10 +1,12 @@
# Maven — Design
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
> Folded 2026-07-30 from `SPEC.md` (north star, 2026-07-03), `maven.md`
> (consolidated decisions, 2026-06-30) and `ROADMAP.md` (execution plan,
> 2026-07-06). Those three files are gone; git history holds them.
> This is the single design document: principles, target state, and the
> execution ledger. `REARCH.md` remains authoritative wherever it disagrees
> execution ledger. `docs/rearchitecture.md` remains authoritative wherever it disagrees
> with anything here. Everything the three sources asserted that is no longer
> the intended design is preserved under **§ Superseded** — do not read that
> section as current.
@@ -160,7 +162,7 @@ presence_state ( last_bucket, last_score, updated_ts )
```
Facts additionally carry `Subject`/`EntityID`/`ResolutionState` for
entity-aware resolution against Nexus (see `MAVEN_ECOSYSTEM_ARCHITECTURE.md`).
entity-aware resolution against Nexus (see `docs/ecosystem.md`).
### Trigger model
@@ -196,7 +198,7 @@ INTO the gate as an env predicate, not the LLM's job.
## Reactive path — routing
**Target design: LLM-as-router** (see `REARCH.md` and `CLAUDE.md`). One
**Target design: LLM-as-router** (see `docs/rearchitecture.md` and `CLAUDE.md`). One
resident model emits GBNF-constrained structured JSON, and the same model
phrases replies; the embedder is a RAG hint, not a routing gate. The
committed default today is the classifier/embedder cascade, which is an
@@ -221,6 +223,30 @@ Not alternatives — layers:
Router contract: `[{"intent":<enum>, key?, value?, text?, verb?}, ...]` over
7 intents (`fact, reminder, note, query, act, chat, system`).
#### A restart expires a parked question
Decided 2026-08-04 (Vikunja #385). The follow-up dialogue session survives a
restart; the clarify question parked behind it does not, and neither do the
three yes/no confirms in `voice.go`. `ClarifyStore` stays in memory.
Three reasons, in the order they settle it:
- The clock stops meaning anything. A parked question carries a 90s TTL and an
attempt count. A restart is a gap of unknown length, so a restored question is
either already dead or pretending to be young.
- Restoring the question restores the request behind it. He asked for something,
she asked back, and then the daemon went away. Acting on that minutes later,
against words he has probably given up on, is the misroute the stage 3 gate
exists to avoid.
- She does not announce it either. The expiry notice needs to know a question
was parked, and knowing that across a restart means storing it. One sentence,
in the rare window where he speaks within 90s of a restart, does not pay for a
marker that outlives the thing it describes. His next words route fresh, which
is the correct answer with or without the notice.
So the notice stays what it is: the in-process TTL case, where she really did
wait and really did let go.
### save-where — the two-memory routing axis
One discriminator: **does the loop evaluate a predicate against it?**
@@ -363,6 +389,29 @@ lives in `source`; rules trust provenance.
`source=poll:healthcheck`. A compromised poller must not be able to forge a
trigger.
#### A fact per monitor, not an aggregate
mavpoll writes one fact per kuma monitor, keyed `service_down:<monitor name>`.
It used to fold the whole gauge into a single boolean, and the nudge could then
only say that something on homesrv was down. That is not something he can act
on, so the rule shipped disabled.
Three things follow from the split:
- The key set is no longer known at wiring time. A rule declares
`WantPrefixes` and the gatherer resolves the family per tick, which is the
only prefix read in the loop.
- Pausing a monitor in kuma silences that monitor. Under the aggregate it
silenced nothing, because some other monitor kept the boolean at "down".
- A monitor deleted in kuma would keep its last fact reading "down" forever, so
mavpoll marks a vanished monitor "unknown". No rule fires on "unknown".
The rule is also edge-triggered: it fires on a transition it has not already
nudged about (`State.NudgedSince`). A polled fact is written only when the
value changes, but the predicate reads the current value, so without the edge
check a service that stays down qualifies on every tick and cooldown is the
only brake.
### Presence — concrete scoring
**Combiner — noisy-OR, not weighted sum.** These are independent-ish positive
@@ -655,7 +704,7 @@ daemon.
The voice wire protocol (length-prefixed JSON frames over TCP) is designed for
**multiple client implementations**. The reference PWA at `cmd/mavweb` is one
client; any app (phone, desktop CLI, smartwatch) can implement the same frame
protocol. The published spec is `PROTOCOL.md` — **generated from
protocol. The published spec is `docs/protocol.md` — **generated from
`internal/voice/wire.go`**, not composed freehand, so it can't drift from
code. It covers transport (4-byte big-endian length prefix), methods
(`PushToTalk`, `Pong`), push kinds (`AudioNudge`), surface identity
@@ -670,8 +719,8 @@ Broadening to home automation, media or comms is JSON, not code.
## Execution ledger
Condensed from `ROADMAP.md` (2026-07-06). The live queue is
`20-07-2026-BACKLOG.md`; current state is `PROGRESS.md`.
Condensed from `ROADMAP.md` (2026-07-06). The live queue is the Vikunja board
(project Maven, ID 2); this table is history, not a work list.
| # | Item | Prio | Status |
|---|------|------|--------|
@@ -755,11 +804,13 @@ Kept for provenance. **None of this is the current or intended design.**
stay deterministic — "classifier owns the route, the SLM stays in its
phrasing lane" — with an embedding + nearest-centroid stage 1 over ~10
examples per intent, and misroutes appended as new centroid examples.
*Replaced by* LLM-as-router (`REARCH.md`): one resident model emits
*Replaced by* LLM-as-router (`docs/rearchitecture.md`): one resident model emits
GBNF-constrained JSON and also phrases replies; the embedder is demoted to
a RAG hint. The classifier cascade is still the code path that runs today
(`llmrouter` is wired nil) but it is an interim stopgap, and it is the known
cause of weak RU query handling — not a design to extend.
a RAG hint. *Landed 2026-07-31:* the LLM router is on by default and set
`true` in `deploy/mavend.json`. The classifier cascade stays as the failure
floor — it runs when there is no llama-server to talk to and on any per-turn
LLM error — but routing by seed similarity is the known cause of weak RU
query handling and is not a design to extend.
- **Named STT/TTS model picks.** `maven.md` picked faster-whisper small/int8
as primary STT with vosk RU for a low-latency command grammar, and silero
(license unverified) as TTS with piper RU as the floor, all on
@@ -769,8 +820,11 @@ Kept for provenance. **None of this is the current or intended design.**
- **Small-model phrasing claim.** `maven.md` specified "lfm2.5 / sub-1b for
phrasing — prompted, not trained," and `SPEC.md` named a specific resident
size. Both are superseded by the RU-CPT + joint persona/router SFT plan.
*Resolved 2026-07-30 (#318):* the resident checkpoint is **Qwen3.5-0.8B**
now, with the CPT'd **Qwen3-1.7B** as the target (#122). Note the resident
*Resolved 2026-07-30 (#318), revised 2026-07-31:* the resident checkpoint is
stock **Qwen3-1.7B** (`UD-Q4_K_XL`, `n_ctx` 4096), which replaced
Qwen3.5-0.8B after measuring better on both fixtures
(`docs/evals/2026-07-31-model-bakeoff.md`). The CPT'd **Qwen3-1.7B** remains the target
(#122); what stock gets wrong is the persona, not the Russian. Note the resident
model is no longer described as untrained — the target is trained
end-to-end, which is the substantive change from the old claim.
- **sqlcipher at rest.** `maven.md` specified sqlcipher with the key read at
+292
View File
@@ -0,0 +1,292 @@
# Deterministic logic around a small model
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
Written 2026-08-02. Branch `fix/integrated`.
## The question
Where does deterministic code attach, so that it helps the resident 1.7B now and
does not fight a larger model later.
## The mistake to avoid
Everything deterministic we have added so far sits in front of the model and
preempts it. Stage 0 matches, the model never sees the turn. That shape helps a
weak model and blocks a strong one, silently.
The fix is not to remove it. The fix is to know which rules are safe in that
position and to have a way to measure the rest.
## Four attachment points
**Bypass, before the model.** The only shape that saves the 2.7s p50. Safe when
the rule is a decision procedure over a closed set, not a guess over an open one.
Exact match and clock queries qualify.
**Evidence, beside the model.** Extractors emit candidate slots as a prior, not a
verdict. The prompt carries the prior and the validator reuses it. A small model
leans on it, a large one overrides it correctly.
**Grammar, around the model.** GBNF built from live state rather than hardcoded.
Costs nothing at runtime and prevents the error instead of catching it.
**Repair, after the model.** Validation failure re-asks with the specific error
rather than overriding. Self-retiring, because a better model trips it less.
## The constraint that ranks them
Latency must stay minimal. That demotes repair and promotes grammar.
- Grammar first. Zero runtime cost, immediate gain.
- Bypass keeps its place. It is the only thing that avoids a model call at all.
- Repair only where failure is rare, capped at one retry.
- Evidence is correct but costs a model call where a bypass costs none.
- Ecosystem calls belong in the snapshot path, in parallel, on strict deadlines.
Never serial before routing.
## The line that never moves
Separate policy from capability compensation. They look alike and age oppositely.
Capability compensation exists because the model is weak. It should be measurable
and retirable.
Policy exists because we decided. The personal boundary, the feminine persona, the
never-search-his-notes rule, the Hexis allowlist and confirmation binding. None of
those yield to a smarter model. A larger model is more dangerous there, not less.
## What we keep
Every stage-0 rule stays exactly as it is. Retiring them was the wrong call and it
would throw away a day of measured gains.
- exact-match fast path
- clock rules
- `SystemTimeDateGrammars`
- `AgendaQueryGrammars`
- `thinSingleToken`, including the social lexicon and the verb-ending test
Three additions, none of which change behaviour:
1. Each rule gets an id and a fixture subset.
2. Each rule writes one trace line when it fires.
3. Each rule carries a comment saying whether its set is closed or open.
That preserves today's accuracy and buys the option to revisit later with numbers.
## What we build
**Dynamic grammars from live state.** The grammars today are static: `routeGrammar`
fixes the 7 intents, `responseGrammar` fixes the mood enum, kiwix `queryGrammar`
fixes word shape. Everything else is a free string.
Candidates in order of payoff:
- **Hexis capability ids.** After discovery the exact list is known. As an enum,
the model cannot name a capability that does not exist.
- **Act fn allowlist.** Hits the two remaining false clarifies directly. They are
the act-with-no-allowlisted-fn arm of `gateLLMDecision`, firing on invented verbs.
- **Calendar names.** Enumerate the real ones for agenda and query slots.
- **Known fact keys.** A read-back matches a stored key instead of inventing a
synonym. This is the general form of the read side we scrapped on 2026-08-01.
- **Nexus display names as act targets.** Only while the list stays small.
Two rules or it backfires:
- **Always include an escape value.** A closed enum with no `other` forces a wrong
pick instead of a decline. The escape is what feeds the clarify gate.
- **Cache the grammar string, keyed on the state that built it.** Rebuilding per
turn is fine. Recompiling a large grammar per turn is not.
## Retrieval over regex
Resolve against stores that already exist rather than adding patterns.
Identity is the worked example. Nexus is authoritative, `actionFact` already sets
`Subject`, and `cmd/mavend/factenrichment.go` resolves it in the background. The
scrapped work invented a parallel key namespace with nothing reconciling the two.
A table that grows with real data beats patterns that grow with our patience.
Note a real gap: the personal boundary in `cmd/mavend/actions_query.go` guards
Maven's own store only. It does not know Nexus or Praxis exist.
## Prerequisites
Both are offline and need no deploy. Neither existed on 2026-08-01, and that is
why the day cost what it did.
**Failure taxonomy over the 77-case fixture.** Classify every miss as model
ignorance, contract loss, prompt ambiguity, or our own bug. Only model ignorance
deserves deterministic compensation. The other three get fixed once, for every
model size. Inference, not verified: much of what we patched was the last two.
**Per-assist ablation runner.** One command toggles each assist off and reports the
accuracy delta on its fixture subset. Then retiring an assist is a config flip and
a number, not an argument.
## Held
Not now, and nothing gets built for it.
- A 12B or 35B model. It may be cloud or the 16GB workstation, and cloud crosses
the current no-third-party line.
- The escalation tier in `gateLLMDecision`.
- Converting the stage-0 heuristics to evidence.
One thing carries forward for free: deterministic assists emit confidence, never a
verdict. A verdict cannot escalate.
## Ruling on the idea list
Nineteen ideas, judged on whether they earn a place in Maven. Checked against the
tree on 2026-08-02, not from memory.
### Build. Absent, and worth it.
**SQLite FTS over embeddings.** No `fts5` anywhere in the tree. Lexical search is
faster than the ONNX embedder, deterministic, and strongest exactly where the
embedder is weakest, which is exact Russian names. Semantic search becomes the
fallback rather than the gate. This is the highest-value absent item.
**Cached TTS phrases.** No cache in `internal/tts`. Confirmations, clarifies and
refusals repeat constantly and their text is already fixed. Pre-rendering them is
cheap and pays straight into the minimal-latency constraint.
**Synthetic router dataset.** Generate Russian tool-calling examples from the same
schemas that will build the dynamic grammars. One source of truth for both, so the
model is trained on exactly the shapes it will be constrained to at inference.
**Command STT separate from dictation STT.** One path today. Commands want latency
and a small vocabulary, meeting capture wants accuracy and can take its time.
`internal/capture` already spools to disk, so the split follows the existing seam.
Medium priority, behind the four above.
### Polish. Present, incomplete.
**Strict JSON everywhere.** Done 02-08-2026. The replier sends
`phraser.ResponseGrammar`. It is exported once so its two copies cannot drift.
The meeting summariser is wrapped and unwrapped in the daemon's Completer, so
`internal/capture` stays text-in/text-out. Every model call now carries a grammar.
**Evidence-first prompting.** Done 02-08-2026, in the evidence branch of
`PhraseQuery`. Sources arrive numbered, one per line. The system prompt no longer
calls them all "заметки" and no longer lets the model add anything of its own.
Blank sources now take the knowledge branch instead of asking for an answer from
an empty list. Not yet measured against a live model. Run `make eval-phrasing`,
and watch the case `query-notes-do-not-answer`.
**Progressive inference.** What exists is a fallback cascade, not escalation. On
error it drops to something weaker. It never escalates on ambiguity. The upgrade is
held with the bigger-model question, and `gateLLMDecision` is the hook.
**Entity dictionaries.** Nexus is the canonical store, `behavior_ru.go` has
`KeyAliases`, `ecosystem_acts.go` has verb aliases. Morphology is the gap, and the
code says so in three places. Russian needs it and there is no stemmer in the repo.
**Assistant state machine.** `confirm.go` models pending confirmation and
`dialogue.Session` models the turn. There is no unified task state. Reuse the
Praxis vocabulary rather than inventing one, because surfaced, acknowledged and
resolved already mean something precise here.
**Session memory compiler.** The digestion tick, `internal/memory` and
`followUpMerge` do parts of this. Not a coherent compile step.
**Background memory maintenance.** Mostly present, and the absent part is small.
What exists: `internal/memeval` reads recent memory on a loop and writes
observations, deduped against its own prior output, unable to speak or act. Facts
supersede at write time through `voidsID`, `CorrectValue` and `VoidLatestFact`,
and `RecentActiveFactsByKind` reads only live rows. Digests, ecosystem traces and
media all prune.
Three gaps remain, all in the fact store:
- Superseding is turn-driven. Nothing reconciles two live facts that contradict
unless a turn corrects one of them.
- Duplicates written by different sources or phrasings stay as separate live rows.
There is no key-level merge pass.
- Nothing expires. The log is append-only and grows without bound, and no fact
ever ages out on its own.
Worth doing, and smaller than it looked. Not urgent.
**Qwen CPT then narrow SFT.** In flight as #122. One disagreement with the idea as
written: do not drop persona from the SFT. Persona is the stated reason the CPT
exists, because stock writes `рад` where Maven needs `рада`. Keep the joint
router-plus-persona tune.
**Offline job queue.** The tick loop, `factEnrichmentWorker` with backoff, and the
media prune already defer work. A general queue is tidier, not more capable. Low
priority.
**Response templates.** Present in `clarify.go`, in the `replySystem` arms and in
`StubPhraser`. Worth extending to high-frequency confirmations, where latency and
persona correctness both matter. Do not extend further. Templating the
conversational reply removes the reason she is worth having.
### Reject.
**Grammar-first routing as a replacement for the LLM router.** Already measured.
The classifier scores 36.8% full accuracy against 72.7% through the cascade.
Replacing the model with rules halves the accuracy. Grammar-first as an ordering is
what stage 0 already is, and that stays.
**Hierarchical intent classification.** Seven intents is already the coarse layer.
The specialisation stage exists as per-intent slot filling in `Extractor.Extract`.
Adding a tier buys structure, not accuracy.
**Tool-specific micro-models.** Contradicts the one-resident-model constraint,
needs per-domain training data nobody has, and multiplies model loads on a single
Vega iGPU. Dynamic grammars give the same domain narrowing at zero runtime cost.
**Local knowledge graph.** Nexus owns entities and relationships. Building a second
graph in Maven breaks the ecosystem line and creates two answers to one question.
If graph traversal is wanted, it is a Nexus feature request.
**Preemptible training.** Training runs in a separate workspace, not on the serving
box. This only becomes real if CPT moves onto homesrv, and that is not the plan.
## Known open, carried over from the deleted handoff
Found live on 2026-08-01, not fixed. Everything else in that file was stale.
- **Kiwix ranks badly on a correct query.** The stop-word pass eats the "and". So
"кто написал войну и мир" reaches Kiwix as "war peace author". The top hit is
"List of peace activists" and she summarises that as the answer. The tag
`scrapped/fact-and-kiwix-phrases` fixes the query text. The ranking is the ZIM
search and is untouched either way.
- **Chat drags prior turns into an answer.** One live reply mixed the greeting, the
height statement and a world question. It named Левитан as the author of Война и
мир.
- **`safeKey` drops Cyrillic**, so Russian calendar events on one day collide.
Vikunja #443 with three fix options. It is a migration, not a patch.
## Open after the 02-08-2026 deploy
SearXNG runs on homesrv at `http://searxng:9563`, on `maven_default`, and the
rebuilt `mavend` wires it. "кто написал войну и мир?" now routes to query, takes
4 results off the search, and answers Толстой with the lookup opener. Two things
that turn left unsettled.
- **The turn was slow, and nobody knows yet whether that is real.** Route 7s,
search 1s, phrasing 15s. The p50 in `docs/evals/2026-07-31-routing.md` is 825ms. It
was the first turn after a cold start with the model still warming, so it
proves nothing either way. Re-run the same question warm before treating it as
a regression. Do not plan latency work off this number.
- **The personal boundary has never run live.** `queryPersonal` in
`cmd/mavend/actions_query.go` stops a question about him from reaching
SearXNG. The question above is not one, so only tests cover it. Ask something
about him on the deployed box and confirm from the log that no `voice: search`
line appears for it.
## Sequence
1. Failure taxonomy over the fixture.
2. Ablation runner, plus ids and trace lines for the existing rules.
3. Dynamic grammar for the act fn allowlist.
4. Dynamic grammar for Hexis capability ids.
5. Dynamic grammar for calendar names and known fact keys.
6. Extend the personal boundary to the ecosystem stores.
Steps 1 and 2 come before anything is written in the router.
@@ -1,5 +1,7 @@
# Maven Ecosystem Architecture
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
## 1. Purpose
This document defines Maven's role in the local ecosystem formed by:
@@ -16,7 +16,7 @@ with Qwen3-1.7B.
Settles Vikunja **#278 / #250**.
- Same fixture and scorer as `ROUTING-EVAL-31-07-2026.md`: `internal/router/eval/`
- Same fixture and scorer as `docs/evals/2026-07-31-routing.md`: `internal/router/eval/`
(`ru_routing_v1.json`, 76 held-out cases).
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:<port> make eval-router`
(`TestLLMRouterBaseline`). (This line used to say there is no `make eval-models` target.
@@ -148,7 +148,7 @@ Qwen3-1.7B wins every column, including against a model 20% larger than it.
| ontopic | 16, 19, 19 | **22, 23, 23** |
| canned fallbacks | 8, 5, 6 | **0, 2, 0** |
This also fills the row `TALK-EVAL-31-07-2026.md` had to void for contamination:
This also fills the row `docs/evals/2026-07-31-talk.md` had to void for contamination:
**600ch/1024tok on Qwen3.5-0.8B scores 13, 11, 8.**
`address` is the headline. It sat at 18-22 of 27 on the 0.8B no matter how the prompt
@@ -162,6 +162,11 @@ The 1.7B does that 0-2 times.
## Latency — the long tail is not the Thinking block
> **Stale, corrected 2026-08-02.** The p50 figures in this table are contention on a
> shared llama-server, not the model's cost. The router measures p50 825ms / p95 1.2s /
> max 3.0s in `docs/evals/2026-07-31-routing.md`, which says so at line 61. Read this table for
> the shape of the tail only. Take absolute latency from the routing eval.
| | p50 | p95 |
|---|---|---|
| Qwen3.5-0.8B | 2.4s, 2.9s, 2.0s | 17.4s, 17.6s, 17.4s |
@@ -213,11 +218,13 @@ swapped again when the CPT lands.
- ~~The routing numbers only reach production once the LLM router is wired on. It is
still `nil`.~~ **Resolved the same evening:** the LLM router is wired at `voice.go:214`
behind `voice.llm_router`, the default is on, and `deploy/mavend.json` sets it `true`.
These numbers are the production path now, so the p50 ≈2.7s is a real per-turn cost and
not a bench artifact.
These numbers are the production path now. **Corrected 2026-08-02: the p50 ≈2.7s in the
latency table above WAS a bench artifact.** It is contention on the shared llama-server,
not the model. `docs/evals/2026-07-31-routing.md` line 61 says so, and measures the router at
p50 825ms / p95 1.2s / max 3.0s. Cite that file for latency, not this one.
- ~~`/mnt/hdd1/llms/LFM2.5/Qwen3-1.7B-UD-Q4_K_XL.gguf` is a 293 MB truncated download
in the wrong directory.~~ **Deleted 2026-07-31.** The good 1.13 GB copy in `qwen3/` is
what `deploy/mavend.json` loads.
- Harness: `scratchpad/bakeoff.sh`, one server at a time, health-checked before each
run, `/v1/models` recorded per run. Never run two LLM consumers at once — see the
contamination note in `TALK-EVAL-31-07-2026.md`.
contamination note in `docs/evals/2026-07-31-talk.md`.
@@ -1,7 +1,7 @@
# Phrasing evaluation — 31-07-2026
How Maven words a nudge, measured instead of argued. Counterpart to
`ROUTING-EVAL-31-07-2026.md`.
`docs/evals/2026-07-31-routing.md`.
- Fixture + scorer: `internal/phraser/eval/` (`nudges_v1.json`, 15 cases; `eval.go`, `checks.go`)
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing`
@@ -65,7 +65,7 @@ prefixes) is the targeted fix, and it would move findings 1 and 2 together. Sepa
`model_quantized.onnx` — not the same file.
`hard` cases score **2/11**: every one is a query where the operator did not reuse his own words.
That is the normal case weeks later, and exactly what DESIGN.md's "recall when relevant" promises.
That is the normal case weeks later, and exactly what docs/design.md's "recall when relevant" promises.
### 4. The memory-store recall branch is dead for notes
@@ -177,7 +177,7 @@ was silent ("не знаю" to "который час") while the one it introdu
### 1. The resident model does route better — 50.0% vs 36.8%
REARCH.md's premise holds; `voice.go:211`'s comment does not. **But the classifier is only
docs/rearchitecture.md's premise holds; `voice.go:211`'s comment does not. **But the classifier is only
~37% correct on held-out utterances, and the model only ~50%.** Neither is "reliable". The
gap between them is real but both are far from a system you would describe as working.
@@ -141,7 +141,7 @@ a model check, which catches a dead server but not a loaded one.
measured) on this fixture and the router fixture. Not the 4B — too big for
this box, owner's call.
- Newer sub-500M candidates (LFM2.5 200M/300M) are worth a run for routing.
Note `MODEL-BAKEOFF-31-07-2026.md` found LFM2.5-**1.2B** worse than
Note `docs/evals/2026-07-31-model-bakeoff.md` found LFM2.5-**1.2B** worse than
Qwen3.5-0.8B at Russian routing and 2.4× slower — but those are a different,
older generation, so that result does not predict the small ones.
- Fix `chat-how-are-you`'s `want_any`, and re-baseline once, so `ontopic`
@@ -0,0 +1,82 @@
# gemma-4-12b on the workstation, against the resident Qwen3-1.7B
Measured 2026-08-02 on the fixtures as they stand. Dated file: it is not edited
after today, and a newer number is a new file.
Vikunja #485's first assumption was that a 7-14B measurably beats Qwen3-1.7B on
the 77-case RU routing fixture and the 27-case talk fixture. It does, on both,
and it is also faster.
## The setup
`gemma-4-12B-it-qat-UD-Q4_K_XL` with the `mtp-gemma-4-12B-it-BF16` draft model,
served by `llama-server` b10220 on bugmachine (AMD 7900 GRE, 16GB), fronted by
`mavgpud` on `192.168.1.105:8080`. Thinking is off through
`--chat-template-kwargs '{"enable_thinking":false}'`, speculative decoding is
`--spec-type draft-mtp --spec-draft-n-max 2`, context 32768. The exact line is
`deploy/mavgpud.json`.
Every number below crossed the LAN from homesrv. Note the trap: homesrv's shell
exports `HTTP_PROXY`, Go honours it, and the runs need
`env -u HTTP_PROXY -u HTTPS_PROXY -u http_proxy -u https_proxy`.
## Routing, 77-case RU fixture
| | full | intent-only | p50 | p95 |
|---|---|---|---|---|
| classifier alone (02-08) | 68.8% | — | 16.6µs | — |
| Qwen3-1.7B through the cascade (31-07, 02-08) | 72.7% | 77.9% | 0.80-1.04s | — |
| **gemma-4-12b through the cascade** | **84.4%** | **93.5%** | **329ms** | 429ms |
| gemma-4-12b alone, no cascade | 55.8% | 85.7% | 335ms | 436ms |
The workstation buys 11.7 points of full accuracy over the resident model. It
buys 15.6 points of intent-only, at a third of the latency. The router's p50 was
never the model's fault, which the 02-08 contention finding already said. A 12B
on a free 16GB card answers a routing turn in a third of a second.
Two things the table hides.
The alone-versus-cascade gap is slots, not intents. gemma reads the intent right
85.7% of the time on its own. It loses full accuracy on seven fact keys
(`вода` instead of `water`, `ужин` instead of `meal`) and on six reminder times
with no time slot. Stage 0 and the daemon's own extractor repair
both, which is why the cascade is 28 points higher. The lesson is that the
cascade earns its keep even under a much better model, not that it is scaffolding
to remove.
`errors: 6` in the alone row are declines on single-token and ambiguous
utterances, all of which the cascade caught. The remaining defects through the
cascade are three `query→fact` confusions, one `chat→query`, and one false
clarify.
## Talk, 27-case conversational fixture
| | pass | notes |
|---|---|---|
| Qwen3-1.7B (31-07) | 20/27 | 11-17/27 for the 0.8B before it |
| **gemma-4-12b** | **25/27 (92.6%)** | chat 8/9, knowledge 9/9, query 8/9 |
Knowledge is the interesting column: 9/9, in Russian, with real answers about
Rayleigh scattering, SSD versus HDD and thunder delay. That is the case the
1.7B cannot do at all and the reason the naming half of the degradation rule
exists.
Two failures, and one of them is the persona defect the CPT (#122) targets:
`query-notes-do-not-answer` wrote `заплатил` where Maven needs the feminine
form. The other is `chat-joke`, where the model told a joke without using any of
the words the check looks for. Run-to-run variance is about one case: a second
run scored 24/27 with `chat-followup-server` also off-topic.
## Nudge phrasing, 15-case fixture
15/15, every check, no errors. `mood`, `lang`, `length`, `feminine`,
`hisgender`, `address`, `cringe` and `ontopic` all clean.
## What this settles and what it does not
Settled: the size question. A 12B on the workstation beats the resident model on
every fixture we have, and it is faster. The offload argument holds.
Not settled: how often the card is free. That is #485's second assumption and
only the `mavgpud` log answers it, after a week of the owner's normal work. A
model that is better whenever it is up is worth little if it is never up.
@@ -0,0 +1,84 @@
# Where the resident model's 7.9GB of RSS goes (2026-08-03, homesrv)
Measured for Vikunja #499. The deployed llama-server held 7.9GB RSS for a 1.1GB
model file. Half a gigabyte of it was in swap, on a box that also runs
whisper.cpp, piper and the embedder.
## Method
`maven-mavend-1` was stopped for the measurement, with the owner's approval.
Its own binary then ran on the host with the exact deployed command line. That
binary is `/opt/maven/bin/llama-server`, version `1 (4c65955)`, a Vulkan build.
```sh
llama-server -m /mnt/hdd1/llms/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf \
--host 127.0.0.1 --port 18099 -c 4096 -ngl 99 --no-webui
```
RSS was read from `/proc/<pid>/status` after load and after each of 8 distinct
1521-token prompts. `smaps` of the deployed process was read first, from inside
the container, since the host user cannot read another user's maps.
## The cause: the prompt cache, not the weights and not the offload
The startup log says it outright:
```text
srv load_model: prompt cache is enabled, size limit: 8192 MiB
srv llama_server: n_parallel is set to auto, using n_parallel = 4 and kv_unified = true
```
The server saves the full KV state of every idle slot it evicts. It keeps up to
8GiB of those states in host RAM (llama.cpp PR 16391). One saved prompt of 1521
tokens costs 166.377 MiB. That is 112 kiB per token, exactly Qwen3-1.7B's KV
footprint (28 layers x 2 x 1024 dims x 2 bytes).
RSS at rest, and per distinct prompt:
| Prompts served | RSS, default | RSS, `--cache-ram 512` |
|---|---|---|
| 0 (just loaded) | 443 MB | 411 MB |
| 1 | 445 MB | 411 MB |
| 4 | 958 MB | 929 MB |
| 8 | 1641 MB | 932 MB |
Uncapped, RSS climbs about 170MB per distinct prompt and does not stop until
the 8GiB limit. Capped at 512 MiB it plateaus at 932MB from the fourth prompt
on, with the cache holding steady at `3 prompts, 499.132 MiB` and evicting.
The 7.9GB on the running daemon was that climb, weeks of it. Its `smaps` showed
one 6.03GB anonymous mapping at 5.32GB resident plus a 1.45GB mapping at 1.27GB
resident, and only 30MB of file-backed RSS.
## The task's leading guess was wrong
`-ngl 99` on the Vega iGPU costs almost no process RSS. A freshly loaded server
has 95MB of anonymous RSS in total. RADV allocates device memory through the
kernel, outside the process, and the log sees 8202 MiB free on `Vulkan0`. The
weights are mmapped and file-backed, so they are evictable and do not pin RSS. The logit buffer is not visible in the numbers above at all.
## Decision
`--cache-ram 512` is now the default, wired as `phraser.cache_ram_mib` and set
in `deploy/mavend.json`. 512 MiB caps total RSS near 1GB, an eighth of what the
box carried. It still holds three of the 1521-token probes above. Maven's real
routing and phrasing prompts are much shorter, so it holds more of those than
the table suggests. `-c 4096` is untouched, as #499
required. A negative `cache_ram_mib` passes no flag, for a llama-server too old
to know it.
Not changed: `n_parallel = 4`. With `kv_unified = true` the four slots share one
4096-token KV cache, so they do not multiply it.
The other half of #499 was that none of these lines were reachable. mavend
scraped llama-server's stderr for the listen line and discarded it, and never
piped stdout at all. Both streams now go to mavend's log with a `llama:` prefix.
The last 12 startup lines go into the error when the server dies before it
listens.
## Deployed
The `mavenai:latest` image was rebuilt and `maven-mavend-1` recreated the same
day. The daemon's own log now carries the child's startup, it reads
`prompt cache is enabled, size limit: 512 MiB`, and the resident server sat at
439MB RSS after load and 613MB after one served turn.
@@ -0,0 +1,46 @@
# Personal boundary, seed scoring vs possession markers, 2026-08-03
Vikunja #495. `что я говорил про бэкапы?` walked past the personal boundary into
SearXNG and came back answered from a Habr article. The boundary matched
possession words only, so a first-person speech verb was not a personal
question.
## What changed
The boundary now scores the turn's query vector against two frozen seed sets.
It claims the turn when the personal side is nearer than the world side. Seeds
and code are in `cmd/mavend/personalboundary.go`. The possession markers stay as
the offline floor for a handler with no embedder.
A regex speech class was written first and dropped. Russian gives every verb a
dozen surface forms, and the "как я говорил, ..." preamble list has no end. Each
form the lexicon missed was one more question reaching the world.
## Measurement
Embedder: multilingual-e5-small int8, the one homesrv runs. Both sides are
embedded on the query side. Cases are held out, none of them a seed. `make test`
runs the offline part. The scored part is opt-in through `MAVEN_ONNX_LIB`, like
`TestONNXRecall`.
19/19 held-out utterances correct (TestONNXPersonalBoundary)
true positive margins +0.014 to +0.089
nearest true negative -0.005 ("кто такой гагарин")
One case missed during the first pass and is not held out any more: `as i said,
what is the population of india`, +0.008 to the personal side. It is a world seed
now.
The gate is the sign of the difference and nothing tighter. The margins are too
thin for a threshold. The asymmetry favours claiming: a false claim costs one
honest "не знаю", a false pass sends his life to an upstream engine.
`make eval-recall` unchanged, 18/27 answered at gate 0.55. Recall does not touch
this path.
## Not verified
The live probe on the deployed box. The daemon was not rebuilt in this session.
The reply to `что я говорил про бэкапы?` with no matching note is still untested
against a real SearXNG.

Some files were not shown because too many files have changed in this diff Show More