Compare commits

..

35 Commits

Author SHA1 Message Date
claude e332f167b2 docs: the ranking's two blocking infra items already shipped (V-447)
Checked the doable, epic and infra tiers against the code, not just the two
tiers V-447 asked about. Ten entries are already built. The two the ranking
calls blockers for everything below are among them: sqlcipher at-rest ships as
Store.enc plus OpenEncrypted, and mavweb/mavcaldav have nine test files
between them where the ranking says zero coverage.

Also built and still ranked as work: rule trace, recurring reminders (cron +
RescheduleReminder), stale-reminder burst collapse (collapseReminders),
revert (VoidLatestFact), digest mode, testing infra, passkey persistence.

Recorded as a section at the top of the archived file so the tiers underneath
are read with the corrections in hand. No tasks created: V-447 scoped task
creation to the mandatory and easy tiers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xwomr2cT93KMSz9u3digX5
2026-08-02 09:20:36 +04:00
claude 322401b9af docs: the ranking was stale, two of its gaps already ship (V-447)
Checked the mandatory and easy tiers against the code instead of trusting the
2026-07-03 ranking. Two were already built and their tasks closed unstarted:
quiet hours (QuietHoursConfig + the care gate in internal/loop/loop.go:37) and
schema migrations (internal/store/migrations.go on PRAGMA user_version, 12+
steps shipped).

Three more were narrowed to what is actually missing. Destructive-confirm has
a mechanism and no policy: store.Tool.Destructive is one boolean, not a risk
tier. Bounded follow-up state has dialogue.Session with a TTL and slot
inheritance; what it lacks is Candidates, so "второй" resolves against nothing.
Clarify has a gate that can ask and one hardcoded sentence to ask with
(internal/voice/replier.go:56, which replier_llm.go hands straight through).

The other six were confirmed absent: list_items, capability model, go.mod
tidy, conversation repair, command history, pronunciation dictionary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xwomr2cT93KMSz9u3digX5
2026-08-02 09:13:59 +04:00
claude 4f34a232d4 docs: retire the two root queues, archive the ranking (V-447)
PROGRESS.md and 20-07-2026-BACKLOG.md were state snapshots that git log and
the Vikunja board already carry. Everything PROGRESS.md claimed as shipped is
a QA task. The backlog's only untracked item, bounded follow-up state, is now
V-448.

maven-feature-ranking.md moves to docs/archive/2026-07-03-feature-ranking.md
instead of dying. Its mandatory and easy tiers became V-449 through V-458; the
doable and epic tiers are reasoning about why things are not worth doing yet,
which no task captures.

Four code comments and the design.md ledger pointed at the deleted files.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xwomr2cT93KMSz9u3digX5
2026-08-02 09:11:11 +04:00
claude 93987f2dfc docs: tier the tree by lifetime, so staleness shows in the path (V-446)
Seventeen markdown files at the repo root, twelve of them dated one-shot
reports sitting next to CLAUDE.md. That is why stale docs read as
current: nothing in the path said which was which.

Root now keeps CLAUDE.md and AGENTS.md. Living docs move under docs/
and carry a Last verified line. Dated measurements move to docs/evals/
ISO-prefixed, and are never edited after the day, so a newer number is
a new file. The senior review moves to docs/archive/.

Every reference was rewritten across markdown, Go comments, the Makefile
and the recall fixture. The touched Go packages still build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 03:28:49 +04:00
kami 7079a240f7 docs: what the first live search turn left open
Two things the deploy did not settle: the turn was slow off a cold start and
that number is not yet trustworthy, and the personal boundary still has never
run on the box.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 02:46:05 +04:00
kami 587f1e6a07 deploy: searxng on 9563, not 8080
8080 is taken several times over on this box, and the container name is the
only thing addressing it, so the port is ours to pick.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 02:37:45 +04:00
kami 99193ff1d1 docs: mark the two step-2 items done
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 02:33:59 +04:00
kami 63a389a1f8 phraser: make her read the sources instead of recalling them
The evidence branch of PhraseQuery framed every source as "твои заметки" and
joined them into one quoted run-on. A live search snippet is not his note, and
a run-on gives a 1.7B one blurred claim to merge rather than sources to answer
from. That is the shape that named Левитан as the author of Война и мир.

Sources now arrive numbered, one per line, and the system prompt says three
ways that the answer comes out of them: only from the sources, say plainly
when they do not answer, add nothing of your own.

Blank sources take the knowledge branch. One empty string used to reach the
evidence branch and ask the model to answer from an empty list, which is the
one prompt guaranteed to make it fill the gap from memory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 02:33:35 +04:00
kami 2150a18e98 deploy: ship the search block, so the live web is on for this box
Owner's call. The code default stays off — no `search` block still means no
query leaves the LAN — but the deployed config now carries one, so a question
that is not about him reaches SearXNG before it reaches the ZIMs.

Nothing runs at http://searxng:8080 on homesrv yet. That is the designed
degradation and not a broken turn: an unreachable instance falls through to
Kiwix and she never says the search failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 02:30:01 +04:00
kami 612ca8cf1b grammar: the last two model calls that were still free text
The replier and the meeting summariser were the two call sites without a
GBNF. Both are exactly the shape that makes a Thinking variant answer with
its reasoning as prose, and neither had anything downstream that could
remove it.

The replier already parses {"response","mood"}, so it now sends the phraser's
grammar for that contract, exported once as phraser.ResponseGrammar so the
two definitions cannot drift.

The summariser stays text-in/text-out. The JSON wrapper is attached and
unwrapped in the daemon's Completer, so internal/capture is unchanged and a
Completer without a grammar still works.

The simulator told routing from phrasing by "has a grammar", which stopped
being true here; it now looks for the intent enum.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 02:28:53 +04:00
kami 14e98334ad query: let her search the live web before she reads the ZIMs
The offline encyclopedia was the only world source, and it reads what was true
when the ZIM was built. A self-hosted SearXNG now asks first and Kiwix is the
fallback for an empty result, an unreachable instance or no line out. Owner's
ruling, 2026-08-02.

internal/websearch is deliberately thin: no rewriter (SearXNG ranks through
real engines, so the Russian question goes out as he asked it), no page fetch,
no cache. It cannot read the store, so only the query string can leave the box.

The personal boundary is unchanged and still sits above this source, so a
question about him is never searched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BFaeSbLMEVG5ey8tejU3y2
2026-08-02 02:24:36 +04:00
kami a103708a08 memeval: five minutes again, now that the gate keeps a turn from waiting
The budget was cut to 60s because a five-minute evaluation held the single
llama-server slot, and a voice turn arriving mid-evaluation waited behind it.
That collision is now solved where it belongs: the background client yields the
slot while a turn is in flight.

With the gate in place the short budget only truncates a Thinking model
mid-synthesis, which costs an observation and saves no latency on any real turn.
Kami's call, 2026-08-02.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BFaeSbLMEVG5ey8tejU3y2
2026-08-02 02:16:45 +04:00
kami a654b0126f docs: name the ecosystem, correct the latency, drop the stale handoff
Three documentation changes and one deletion.

CLAUDE.md and AGENTS.md gain the Nexus/Praxis/Hexis sections that were written
last session and never committed: what each service owns, where Maven's client
for it lives, and the rules that are not negotiable.

The p50 latency figure was wrong in two files. CLAUDE.md said the cascade costs
2.7s and that the LLM router is 90x slower than the classifier. Both come from
the bakeoff table, where the number is contention on a shared llama-server, not
the model. ROUTING-EVAL-31-07-2026.md line 61 says so and measures the router at
p50 825ms / p95 1.2s / max 3.0s. Corrected in CLAUDE.md, and the bakeoff table
now carries a header pointing at the routing eval for absolute latency. Latency
work was about to be planned off a number that was never real.

HANDOFF.md is deleted. It described work sitting on fix/integrated waiting for a
fast-forward onto overnight/eco-versioned-traces. Neither is true: master
contains that tip plus 22 commits, and both branch pointers are stale. The three
live defects it recorded move to PLAN-DETERMINISM-02-08-2026.md, which is now
the only planning document.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BFaeSbLMEVG5ey8tejU3y2
2026-08-02 02:16:37 +04:00
kami b8227295b8 query: stop asking the world about him
"во сколько у меня встреча" fell past the calendar, reached Kiwix, matched
an article on the 2015 CPISRA World Games and came back phrased as his
meeting. Inventing is worse than refusing, and as `system` this used to say
"пока не умею".

A new source sits between the notes pass and Kiwix: a question carrying a
possession marker ("у меня", "мой", "my", "do i have", "did i") that his own
data did not answer has no answer outside it either, so the walk stops there.
The markers are possession, not first person, so "как мне сварить борщ" still
reaches the encyclopedia.

This is also where CLAUDE.md's privacy line lands: only the utterance may
leave the box, and a question about him carries his life in the utterance.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 23:27:04 +04:00
kami b35151418a mavend: give the voice handler a CoreAPI that can serve the day plan
wireVoice runs before the tick loop exists, so it could only be handed
the bare store adapter — and that adapter answers DayPlan with "not
available via direct store API", because a day plan is assembled by the
tick loop and is not a table to read. So queryDayPlan, which the query
chain reaches for "какие у меня планы на сегодня", failed for every
caller on the deployed daemon.

main already back-patches the other direction (daemonAPI.chatFn =
handler.handleText). This is the same seam in reverse, at both wiring
sites. No recursion risk: nothing in the voice path calls api.Chat.

With the plan reachable, it recited its reminders as literal JSON. The
payload unwrapper existed but was private to the phraser, so the day
plan had its own non-unwrapping copy. One owner now, store.ReminderText,
with the phraser delegating to it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 23:17:10 +04:00
kami 17964d1162 dialogue: keep the remembered topic on the latest turn
rememberTurn runs after followUpMerge, which has already inherited a
Text slot from the previous same-intent turn, so the fill-if-empty rule
pinned the first topic of a run of query turns and never released it.
"во сколько у меня встреча", then "какие у меня планы", then "а завтра?"
continued the meeting — two turns stale.

Overwrite for system and query, where Text is a topic and not a payload.
A continuation is the exception and keeps what it inherited: its own
utterance is the ellipsis, and the topic it carries is the real one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:57:08 +04:00
kami 079cf689aa router: agenda questions belong to query, not to system
"что у меня сегодня" and "что у меня в календаре сегодня" both routed
IntentSystem on the deployed daemon, and replySystem has no agenda arm,
so both answered "пока не умею". The calendar source that can answer
them lives in the query chain and was never reached. The fixture has
said query since ru-query-019 was written; the daemon disagreed with the
fixture and the daemon was wrong.

AgendaQueryGrammars routes them at stage 0, after the clock rules so
"какой сегодня день" keeps reaching replySystem. Intent only — which
source claims the turn stays the query chain's decision.

This is what made the follow-up continuation look like it only worked
for "what day is it". It did: the query half inherited an intent whose
handler could not answer, so both halves came back "пока не умею".

Measured on the 77-case RU fixture: full accuracy 70.1% → 72.7%,
intent-only 75.3% → 77.9%, calendar 0/2 → 2/2, clarify counts unchanged.
The eval harness wires the new grammars too, or the fixture would stop
being a measurement of the daemon.

Go's \b is ASCII-only and never fires after a Cyrillic letter, which the
first version of the pattern learned the hard way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:53:01 +04:00
kami 9397f9e5f6 query: only a source that knows the day may answer "а завтра?"
The follow-up continuation re-aimed Slots.Time, but every query source
matches on dec.Utterance and nothing in the chain reads it. So the
inheritance bought almost nothing: "какие планы на сегодня" then "а
завтра?" missed day-plan's matcher and fell to the calendar, which had
parsed the day out of the raw utterance anyway.

Worse than nothing in one place. queryFactByKey runs first and claims on
HasKey plus HasTime, both of which the continuation sets, so a keyed
query continued with "а вчера?" answered "я записала это <the fact's own
timestamp>" and dropped the day entirely.

Widening the topic for every source would have made it worse still, not
better: CalendarEvents is the only CoreAPI call that takes a date, so
ten date-blind sources would have answered a question about tomorrow
with today's data. Gate instead of widen — a continuation is only
offered to a source that reads the day, and when none claims it she says
so rather than "не знаю", which reads as "nothing tomorrow" when she
never looked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:32:37 +04:00
kami 3f98a99f44 continuation: only an ellipsis may widen the keyword match
Deployed check, second round: "привет" after "какой сегодня день" answered
with the date. followUpMerge fills an empty Text from the previous
same-intent turn, so the topic-widening added a minute earlier was reading
an inherited topic on turns that had nothing to do with it.

Decision gains Continued, set only by continuation.go and never by the
router. replySystem widens on that and nothing else, so an inherited Text
is back to being invisible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:21:28 +04:00
kami 53616836db continuation: an ellipsis names the day, not the topic
Deployed check: "какой сегодня день" then "а завтра?" answered "пока не
умею". The intent was inherited correctly, but replySystem keyword-matches
the utterance, and "а завтра?" contains no topic word — that is the whole
nature of an ellipsis.

So the topic travels with the session. rememberTurn keeps the raw utterance
in Slots.Text for system and query turns that have no Text slot of their
own (a stage-0 grammar fills none), and replySystem matches keywords
against utterance + Slots.Text. Dates keep parsing from the utterance
alone, which is the part the ellipsis actually restates.

Only system and query: everywhere else Text is a payload and must stay
exactly what the router put in it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:16:58 +04:00
kami cf40f13573 continuation: a reminder's day is written into its payload, not just its slot
Live check on the deployed daemon: "напомни сегодня о событиях" then "а
завтра?" fires tomorrow with the text still reading "сегодня". Re-aiming
Time is not enough when the day word is also inside the payload, and
rewriting the payload needs the date's span in the string, which
ParseCalendarDate does not report.

Reminder comes out of continuableIntents until that exists. query and
system are unaffected: their Time slot IS the whole question.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:13:25 +04:00
kami 7d08d27efb mavend: answer "а завтра?" from the previous turn, not from the model
An elliptical follow-up carries no intent of its own. followUpMerge cannot
help — it inherits slots once the intent is known, and here the intent is
the missing part. So "а завтра?" went to the router, which on a 1.7B is
close to a coin flip, and the guess cost ~2.7s.

continuationDecision runs before the router and rebuilds the turn from the
previous one: same intent, same key, new day. Deterministic and free.

Three guards, all narrow on purpose. A parseable date is required, which is
what separates an ellipsis from an ordinary short utterance. Four tokens
max. And only query, system and reminder may be inherited: fact and note
would write something he did not say, and act would let a two-word
utterance re-run an allowlisted fn, which is a way to fire a destructive
command nobody typed.

A continuation is still remembered, so "а завтра?" then "а послезавтра?"
chains.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:09:05 +04:00
kami d0d0021659 mavend: let a fact close the nudge that asked for it
The snooze wire had one end. "готово" and "выпил воды" both left the nudge
pending, so the auto-tuner only ever learned from deferrals and from
silence — never from the rule working.

Two entry points, because the two utterances are different acts. A bare
"готово" carries no content and is intercepted before the router, sharing
pendingNudge and the twenty-minute window with the snooze. "выпил воды" IS
content: it routes normally, writes its fact, and only then closes the
nudge (ackFromFact, after applyAction). Folding the second into a pre-route
intercept would have thrown away the thing he actually said.

ackFromFact is silent. The fact reply stands; "отлично, отметила" on top
would be her congratulating him for obeying, which is the nag she is not.

Which fact answers which rule comes from the rule's own InertWhenNoData, so
a rule added later is covered the day it lands. "да" and "ок" stay out of
the ack vocabulary: the clarify gate upstream has the stronger claim on
them, and a stray "да" must not rewrite the feedback signal.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 21:43:28 +04:00
kami 79c3b994cf mavend: let him say "потом" to a nudge
A nudge could only be deferred from Telegram or the web UI. The voice path
had no route to store.ResolveNudge at all, so the channel she nudges on
hardest was the one he could not answer out loud.

resolveSnooze runs pre-route, right after the quiet toggle, and writes the
same `snoozed` outcome the buttons write — which also drops the row out of
RepeatUnacked, so a deferred sev4 stops re-sending every five minutes.

The window is what makes this safe to run before the router. "потом" is an
ordinary word; it only counts as a deferral when a pending nudge was sent
in the last twenty minutes, and otherwise the turn routes normally. Single
word patterns still match single-word utterances only, so "потом схожу за
водой" reports a plan instead of silencing the rule that prompted it.

Channel is not filtered: a nudge that went to Telegram is still what he is
answering when he says "потом" at the microphone.

QA-PLAN gains the two new manual checks and drops the 319 warning.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 21:40:06 +04:00
kami ba1d8e3f44 quiet: the noun form and the comparative are commands too
The pre-route toggle knew "тихий режим" and every negation of it, but not
the two phrasings that get spoken most: "включи режим тишины" (the setting
named as a noun) and "сделай потише". Both fell through to the router,
which has no quiet intent, so the command did nothing at all.

Adds those as stem pairs, plus a quietWordStems list so the
negated-but-unmatched fallback recognises "хватит тишины" the way it
already recognised "хватит тихого режима".

Locks the English phrasings from the routing fixture in the test table
("turn quiet mode back on", "turn off quiet mode", "stop quiet mode") —
all three already behaved, none were covered.

"на улице стало потише" and "в тишине лучше думается" stay inert: a
single-word pattern still only matches a single-word utterance.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 21:35:58 +04:00
kami 29329b5f0e router: a Russian verb is a whole sentence, not thin evidence
The clarify gate thinned any one-word utterance to 0.3 confidence, which
trips the stage-3 gate and comes back as "не совсем поняла". That is an
English intuition. Russian packs subject, tense and gender into one word,
so "поужинал" is a complete report and "привет" a complete greeting, and
both got clarified.

thinSingleToken keeps the rule for bare nominals, where it is real ("вода"
is a fact-or-query coin flip), and spares two classes: a closed lexicon of
social and control singles, and any token carrying a verb ending. Both
tests are offline.

Fixture: false clarifies 3 → 2, intent-only 74.0% → 75.3%, full accuracy
unchanged at 70.1%, missed clarify still 1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 21:29:21 +04:00
kami 47dda97226 tick: a disabled rule must not keep repeating its last alarm
`disabled_rules` stopped the loop from creating new nudges and did nothing
about the ones already sent. The sev4 repeat path does not consult the rule
set at all: RepeatUnacked re-sends any telegram nudge still at outcome=pending
every repeat_interval (5m by default), driven by store.UnackedTelegramRules.
So service_down kept arriving on a five-minute cadence after being switched
off, from a row written hours earlier — two messages after the deploy, which
is how it was found.

That cadence, not the unsealed database, is what "she keeps spamming me"
always was. The seal bug erased the acks that would have stopped it.

Filter the repeat keys against the wired rule set. Filtering on wired rather
than on the disabled list also silences a rule deleted from the code: nothing
can ack what the UI no longer lists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 20:43:27 +04:00
kami 8fdb9e5cd1 loop: let config turn a nudge rule off
The kuma service_down nudge cannot name the service. mavpoll folds the whole
monitor_status gauge into one boolean fact keyed `service_down`, and the
phraser names a service only when the fact key is the service name, so the
message is always the generic "a service on homesrv is down". Every fifteen
minutes, with nothing to act on. A fact per monitor is the real fix and it is
filed as Vikunja #444; this is what to do until then.

`disabled_rules` in mavend.json subtracts from loop.DefaultRules by name.
Config only subtracts — rules stay code, the set stays canonical and ordered
as written. A disabled rule is not gathered for either, since the gatherer
derives its key set from the rules it was given. Unknown names are ignored so
deleting a rule cannot brick a config that still lists it, and the boot log
says what was dropped, because a rule that vanishes silently looks exactly
like a rule that is broken.

Deploy turns service_down off.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 20:20:32 +04:00
kami f1a809121b shutdown: close the sockets, or the database never gets sealed
mavend seals its encrypted database in `defer st.Close()` when run() returns.
It had not returned since 2026-07-21. Every restart since then decrypted the
same eleven-day-old ciphertext and rolled back everything written in between:
the Telegram nudge that kept firing was a fact being un-written on each boot.

The goroutine dump named it. main → srv.Close() → ipc.(*Server).Close →
wg.Wait(), waiting on per-connection goroutines parked in readFrame. Close
shut the listener and nothing else, so the idle persistent sockets held by
mavweb, mavpoll, mavcaldav and mavmaild blocked shutdown forever. `docker
compose stop -t 60` spent the whole sixty seconds and then took a SIGKILL.

So: track the accepted conns and close them, in ipc and in voice, which had
the identical defect. Bound all three waits — the two per-server ones and the
worker wait in main — because the seal matters more than any single in-flight
call. A dropped RPC costs one reply; a missed seal costs a session.

The regression test leaves a client connected and idle, which is the case the
old tests avoided by closing the client first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 20:05:52 +04:00
kami 79893d646b kiwix: read the article, not the snippet, and state the persona as morphology
Two things that were each half-done.

The prompts handed the model copyable examples. chatSystemPrompt lost its
openers this morning and the "Я подумала, что" tic went with them, but "не
забыл ли я" appeared in its place: the removed example had been suppressing
the masculine self-reference by accident. Two predicatives are not enough
signal, so the rule is now stated as morphology (-ла) rather than as a pair of
words — a suffix rule generalises where an example only gets copied.
querySystemPrompt had the same defect and gets the same treatment; its "вот что
я нашла: " opener is deliberate and stays.

internal/kiwix had no caller. actions_query.go said "once internal/kiwix is
wired into this chain" and that never happened. It is wired now, between the
notes pass and the web source: everything of his answers first, and only what
is left over is looked up. Off unless a `kiwix` block names a server and a book.

Reading the search snippet does not work. Kiwix builds it from wherever the
keyword matched, which on Wikipedia is the navigation box at the foot of the
page — the first version of this answered "что такое фотосинтез?" by reciting
"Ecological economics Ecological footprint Ecological forecasting …". Client
grows an Article method; the head of the article is the lead paragraph, which
is the definition the snippet was meant to be. Verified on the box: the same
question now answers correctly off the ZIM.

Only the rewritten query leaves the process. A test asserts it: a turn carrying
a stored note must not put that note in the search string.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 19:46:39 +04:00
kami c04c5eca9c phraser: stop handing the chat prompt's examples back as replies
The grammar examples in chatSystemPrompt were full clauses — ("я подумала",
"я рада") for her, ("ты сказал", "ты забыл") for him. A 1.7B copies those
instead of generalising from them.

Observed on homesrv 2026-08-01, in one session: all three chat replies
opened with "Я подумала, что ...", and one ended "...немного тревожусь.
ты сказал" — the second example pasted onto a finished sentence, which
reads as a truncation and is not one.

Contrastive pairs replace the openers, so the rule reads as a correction
rather than a template. The him-examples are dropped; the "ты" instruction
carries that on its own, and those two produced the worst output. A closing
line tells her not to echo the instructions, because a small model treats a
quoted string as licence to reuse it.

Verified after rebuild: three chat turns, no "Я подумала" opener, no
dangling example, feminine forms intact.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 18:01:16 +04:00
kami bfdbe0045e telegram: send through the x-ui relay instead of straight at the API
Direct egress to api.telegram.org does not work from homesrv, so every
away-channel send timed out and the service_down nudge retried once a
minute forever. The sink already had a Proxy field wired to
http.Transport.Proxy; nothing had ever set it.

The relay is the x-ui socks inbound on the host, port 10808, addressed from
the container as the maven_default bridge gateway. That also needs a ufw
rule, because the bridge subnet is not otherwise allowed to reach a host
port and the SYN is dropped rather than refused. The rule is recorded in
the config next to the address, since the address alone is not enough to
reproduce this on another box.

Verified: five minutes after restart, zero send errors where there was
previously one per minute.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 17:22:12 +04:00
kami d09954d85d telegram: keep the bot token out of the log, and correct two design claims
net/http wraps every transport failure in *url.Error, whose Error() prints
the request URL. Telegram accepts the bot token nowhere but the URL path, so
a send failure wrote the live token into the daemon log. On 2026-08-01
homesrv could not reach api.telegram.org and did that once a minute for as
long as the network stayed down. The token lives in deploy/telegram.env to
stay out of the repo; putting it in `docker compose logs` undoes that.

Both error sites now go through redact. The structural branch rewrites
url.Error.URL and keeps the type, so errors.As still matches; anything else
falls back to scrubbing the rendered message. No minimum-token-length guard:
a one-character token would shred the message, but that beats leaking it.

DESIGN.md still said the classifier cascade was the path that runs today
with llmrouter wired nil, and gave the resident checkpoint as Qwen3.5-0.8B.
Both stopped being true on 2026-07-31.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 17:11:11 +04:00
kami 7e402b279d models: drop the self-referential stt/tts symlinks
7ab9b48 committed models/stt and models/tts as symlinks to their own
absolute paths:

    models/stt -> /home/kami/apps/Maven/models/stt

They came from an agent worktree under scratchpad/wt, where a link back
to the main checkout resolves. In the main checkout it points at itself.

Both paths are gitignored, so checking out that commit overwrites the
real model directories without warning and git says nothing. On homesrv
it destroyed models/stt/ggml-small.bin and the piper voice, and mavsttd
crash-looped on the missing whisper model.

The directories are host state fetched separately, per deploy/README.md.
Nothing under models/stt or models/tts belongs in git.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 16:59:33 +04:00
claude 6915e6a714 Land the 35-PR overnight stack (PRs 50-84)
One linear chain of 35 PRs, reviewed and fixed. The eleven fix branches were merged onto fix/integrated and fast-forwarded onto this tip, so the review findings land as commits here rather than on the individual PRs.

make test and make build pass.
2026-08-01 14:50:26 +02:00
95 changed files with 4557 additions and 1057 deletions
-396
View File
@@ -1,396 +0,0 @@
beyond the model and tts work, the useful additions are mostly around **reliability, context, and reach**, not more intelligence.
## highest-value additions
### 1. unified event intake
maven should receive normalized events from:
* praxis
* calendar
* telegram
* local notifications
* system/service health
* manual checklists
* eventually email bridges
one internal envelope:
```go
type Event struct {
Source string
Kind string
EntityIDs []string
Title string
Body string
Priority string
OccurredAt time.Time
Payload json.RawMessage
}
```
this gives digestion one stable input instead of source-specific logic.
---
### 2. explicit morning routine engine — **core engine done (2026-07-20)**
`internal/morning` — pure checklist engine, mirrors `internal/loop`/
`internal/routine`'s no-I/O contract. `Evaluate(routine, facts, now)` answers
"what's still missing" any time (order-independent — checks facts, not
sequence); `Due(routines, facts, last, now)` fires the once-per-day nag only
at `NudgeAt` (defaults to window end) and only when something's unevidenced,
with a `last`-map dedupe identical in shape to `routine.Due`'s cold-start/
last-fire tracking. Evidence is just a fact timestamped inside today's
window — manual (voice-tapped) and inferred (another daemon writing the same
key) are indistinguishable, satisfying the manual/inferred requirement for
free. Weekday/weekend variants are two `Routine`s with different `Weekdays`
sets under different names. Wired into `config.MorningRoutineConfig` +
`cmd/mavend/tick.go`'s `fireMorningRoutines` (reads only the fact keys the
configured items reference, dispatches through the normal severity/presence
routing table, body is literal joined item labels — not LLM-phrased, same
no-hallucination rationale as cron routines). 13 unit tests in
`internal/morning/morning_test.go`.
Added since (2026-07-20, same day): a read-only `/morning` page in mavweb —
`ipc.CoreAPI.MorningStatus` (new wire method, mirrors `TickTrace`'s
daemon-cache-only shape: the store adapter errors, `daemonAPI` serves it from
a `tickLoop.morningStatus` closure) returns each routine's active/window/
per-item done state, server-rendered same as `/trace` (no live-update loop —
checklist state moves on minutes, not seconds).
Not yet done: no config wired in `deploy/mavend.json` (no morning routines
configured on homesrv yet — add items there when the medicine/water/pets
fact keys the phone/desktop write are settled), no voice query path for
"what did I miss this morning" (Evaluate supports it; nothing calls it yet),
no way to create/edit routines from the web UI — construction still means
hand-editing config, deliberately deferred: routines are operator-declared
config (like cron routines), and a CRUD editor would mean moving them to a
DB table + hot-reload, a bigger change than this pass.
not ordinary reminders.
support:
* required morning items
* order-independent completion
* soft time windows
* skipped-step detection
* one nudge, not repeated spam
* manual and inferred completion evidence
* weekend/weekday variants
example:
```text
08:0011:00
- medicine
- water
- pets
- check praxis attention
```
maven should know what is still missing, not merely fire four timers.
---
### 3. cross-device presence
**status (2026-07-20):** the hysteresis engine and 3 of the listed signals are
already built and wired live: `internal/store/presence.go` (noisy-OR combiner
+ Schmitt-trigger bucket resolve), fed by `desk_active` (workstation, via
`scripts/desk-active.sh` posting to `/api/signal`), `page_heartbeat` (mavweb
tab, `app.js`), and `wg_handshake` (`mavpoll` polling `wg show`) — threaded
into the tick loop via `internal/loop/gather.go`. Not done: phone-reachable,
homesrv-available, audio-output, and active-maven-client signals from the
list below are still missing.
a small presence daemon on each trusted device:
* workstation active/idle
* phone reachable
* homesrv available
* last keyboard/mouse activity
* wireguard presence
* current audio output
* active maven client
mavend receives only compact state, not raw activity logs.
useful for:
* choosing delivery channel
* suppressing voice while away
* surfacing reminders when you return
* knowing whether an agent result should be spoken or sent as text
---
### 4. interruption policy — **done (2026-07-20), turned out to already be built**
audited the existing code before writing anything new: `internal/loop.Gate`
already answers deliver_now vs. drop (quiet-hours/cooldown/snooze/presence/
calendar-busy), and `cmd/mavend/tick.go`'s `digestQ` + `config.DigestConfig`
already implement queue/digest (low-severity nudges batch into one
notification, flushed on window elapsed or max-items reached). The four
outcomes below were already covered by these two mechanisms; nothing new to
build for the core policy.
Gap that *was* real: `deploy/mavend.json` had no `digest` block, so batching
was disabled in prod despite being fully implemented. Fixed — see the config
change alongside this note.
before delivering anything, evaluate:
```text
urgency
current activity
quiet hours
recent nudges
available channels
whether already surfaced
```
result:
```text
deliver_now
queue
digest
drop
```
this prevents maven from becoming annoying once praxis and other sources start producing more data.
---
### 5. entity-aware memory — **done (2026-07-20)**
`03fa52d`/`9876187` (Vikunja #279): facts gain `Subject`/`EntityID`/
`ResolutionState`; an async enrichment worker resolves free-text subjects to
canonical Nexus entity_ids (mirrors Praxis's enrichment pattern). Ambiguous
or unreachable Nexus never guesses — the fact stays `pending` or terminal
`ambiguous`. Voice-tapped facts (`IntentFact`) now flow into the enrichment
queue automatically via an optional `Subject` field on `WriteFactReq` (old
callers unaffected).
Landed alongside this in the same session (not originally on this list, but
closes the plumbing gaps the last brief flagged for Nexus/Praxis maturity):
a typed Praxis lifecycle client (`398997f` — surface/acknowledge/resolve/
ignore/pin; fixes the surfaced≠acknowledged gap where reading an item aloud
left no trace), correlation-ID/version headers on the Nexus/Praxis clients
(`b743860`), entity-scoped Praxis attention queries (`0579ef9`), a durable
delivery outbox with begin-before-send/complete-after semantics
(`29f23e3`+`9ff726e` — closes a duplicate-send-on-crash bug), fail-closed
handling on ambiguous IPC mutation outcomes and Nexus/Hexis dependency
errors (`838fde1`+`d9fa4d6`), and a reusable fake-ecosystem test harness
with fault injection (`c932cd8`).
connect maven memory to nexus ids.
instead of:
```text
key = "кошачий фонтан"
```
store:
```text
entity_id = ent_pet_water_fountain
predicate = refilled_at
value = 2026-07-19T...
```
benefits:
* stable russian/english aliases
* fewer duplicate facts
* better “when did i last…” queries
* easier routine detection
* cleaner praxis correlation
---
### 6. bounded follow-up state
for short continuations:
* “yes”
* “tomorrow”
* “the second one”
* “not that project”
* “do it later”
store explicit pending state instead of relying on chat history:
```go
type PendingInteraction struct {
Kind string
Candidates []string
Args json.RawMessage
ExpiresAt time.Time
}
```
this matters a lot for a 1.7b model.
---
### 7. evaluation lab — **skipped for now (2026-07-20)**
runs on a different machine (GPU box), and CPT is currently in progress
there — deprioritized until the training pipeline has a checkpoint to gate.
Not abandoned, just off the immediate list.
before every new checkpoint or lora deploy:
* routing accuracy
* slot accuracy
* malformed json rate
* russian/english mixed input
* ambiguous entity handling
* reminder vs note vs fact
* direct answer vs tool call
* confirmation safety
* phrasing quality
* latency and ram
also replay real anonymized traces against old and new checkpoints.
this should be a hard deployment gate.
---
### 8. replayable full-system simulator
fake:
* clock
* presence
* caldav
* telegram
* praxis
* nexus
* hexis
* stt
* tts
* llama-server
scenario:
```text
08:30 user appears
08:35 medicine not completed
08:40 correx agent waits
08:45 calendar sync stale
08:50 user says “what did i miss?”
```
assert:
* what tools were called
* what was surfaced
* what stayed unresolved
* what maven said
* what was not executed
this will save more time than another feature daemon.
---
## useful second-wave additions
### voice session quality
* barge-in
* interrupt tts on wake word
* partial stt display
* confidence-aware clarification
* retry only failed stt segment
* per-room microphone profiles
* noise-floor calibration
* short response mode when speaking
### notification bridge framework
small adapters for:
* ntfy
* telegram
* matrix
* web push
* android notification forwarding
* local dbus notifications
normalize into maven/praxis events instead of treating each as a separate feature.
### local knowledge ingestion
* markdown/docs ingestion
* git repo summaries
* project decision records
* conversation exports
* provenance and source links
* incremental reindexing
keep this read-only and separate from personal fact memory.
### service self-diagnostics
`maven doctor`:
* socket reachability
* model health
* stt/tts readiness
* embedder availability
* caldav freshness
* telegram poll state
* praxis/nexus/hexis reachability
* db integrity
* disk usage
* recent failures
### config and secret management
* schema-validated config
* config migration
* secret references instead of inline values
* dry-run validation
* redacted config dump
* per-daemon health config
* startup dependency report
---
## things i would not build yet
* autonomous multi-step planning
* large external reasoner
* generic workflow engine
* self-editing memory
* automatic hexis actions from praxis
* emotion simulation beyond phrasing
* full home-assistant replacement
* more model layers before routing is stable
## recommended order
**status as of 2026-07-20:**
1. ~~evaluation lab~~**skipped, GPU-box work, deprioritized while CPT is in progress**
2. ~~entity-aware memory~~**done** (`03fa52d`/`9876187`, plus adjacent
Nexus/Praxis plumbing hardening — see item 5 above)
3. ~~morning routine engine~~**core engine done** (`internal/morning` +
`cmd/mavend` wiring — see item 2 above; not yet configured on homesrv,
no voice query, no web UI)
4. interruption/delivery policy
5. presence agents
6. unified event intake
7. full-system simulator
8. notification bridges
9. knowledge ingestion
10. voice-session polish
the main goal should be: **maven reliably knows what is happening, knows what you meant, and chooses the least annoying correct response**. everything else can wait.
+40
View File
@@ -5,6 +5,46 @@ This repo maps to **Maven** (project ID: 2) in Vikunja.
Feature work, bugs, deployment tasks all go here.
MCP endpoint: `http://localhost:9100/mcp` (or `http://192.168.1.104:9100/mcp` from workpc)
## The sibling services (Nexus, Praxis, Hexis)
Maven is the conversational front end of a four-service ecosystem. The other three
live in sibling repos next to this one.
| Service | Repo | Port | Answers |
|---|---|---|---|
| Nexus | `../nexus` | 9740 | who or what is this name |
| Praxis | `../praxis` | 8989 | what needs attention |
| Hexis | `../hexis` | 9741 | what can be run, and running it |
Division of labour: Nexus identifies, Praxis observes, Hexis acts, Maven understands
and coordinates. Maven is not the source of truth for any of the three. The full
contract is `docs/ecosystem.md`, and the constraints that bite during
implementation are summarised in `CLAUDE.md`.
Where things are in this repo:
- `cmd/mavend/ecosystem.go` holds `nexusClient` and `praxisClient`. The Hexis client
is vendored from `github.com/kami/hexis/pkg/client`.
- `cmd/mavend/ecosystem_acts.go` routes an act through capability discovery.
- `cmd/mavend/factenrichment.go` resolves each stored fact's `Subject` against Nexus
on a background poll loop, with backoff and no give-up.
- `internal/store/entityfacts.go` holds the entity-tagged fact rows.
- Config blocks are `nexus`, `praxis` and `hexis` in `deploy/mavend.json`. Each is
optional. Absent means that integration is dark, not broken.
Bring the whole ecosystem up locally:
```sh
docker compose -f deploy/ecosystem/docker-compose.yml up -d
```
That builds all three from the sibling working trees, so commit or stash there first.
Each publishes on loopback at the port above. Maven reaches them by service name on
the shared compose network.
Testing without them running: `cmd/mavend/fakeecosystem_test.go` provides stubs, and
`cmd/mavend/ecosystem_degraded_test.go` covers each service being unreachable.
## Rendering / previewing the web UI locally
To see mavweb pages with real data without touching the production stack:
+76 -11
View File
@@ -10,7 +10,7 @@ compose passes `/dev/dri` + the render gid) — the resident model stays ≤1.7B
**Resident model:** currently **Qwen3-1.7B** (`UD-Q4_K_XL`), stock — not yet the CPT'd one.
It replaced Qwen3.5-0.8B on 2026-07-31 because it measured better on both fixtures we have:
67.5% vs 59.7% intent-only on the 77-case RU routing fixture, and 20/27 vs 11-17/27 on the
talk fixture. See `MODEL-BAKEOFF-31-07-2026.md`. It is a Thinking variant, so `n_ctx` is 4096
talk fixture. See `docs/evals/2026-07-31-model-bakeoff.md`. It is a Thinking variant, so `n_ctx` is 4096
— reasoning tokens need the room, and 4096 is what the scores above were measured at.
The **target** is still the locally CPT'd **Qwen3-1.7B** (Vikunja #122, training in flight).
@@ -25,7 +25,7 @@ Spanish. Their strong published IFEval/BFCL numbers are English-only. Model file
`models/llm/`, so the LFM2.5 gguf sitting there is not loaded by anything. Swapping the resident
model is a one-line change to `phraser.model_path` in `deploy/mavend.json`.
See `REARCH.md` for the target architecture, `DESIGN.md` for the folded design spec, and
See `docs/rearchitecture.md` for the target architecture, `docs/design.md` for the folded design spec, and
`AGENTS.md` for local-preview + model-download recipes.
## Build & test
@@ -68,13 +68,52 @@ Daemons are wired socket-to-socket, not linked. `internal/ipc` is the client/ser
protocol; the config in `deploy/mavend.json` (with `${VAR}` env expansion from gitignored
`deploy/telegram.env`) sets socket paths, model paths, and the phraser/embedder blocks.
## The ecosystem: Nexus, Praxis, Hexis
Maven is one of four services. It owns conversation and personal memory. It does not
own identity, operational state, or execution. Full contract in
`docs/ecosystem.md`.
```text
Nexus identifies. Praxis observes. Hexis acts. Maven understands and coordinates.
```
| Service | Owns | Maven's client | Configured at |
|---|---|---|---|
| **Nexus** | Canonical entity ids, names, aliases, relationships. Projects, services, devices, people, pets, places. | `nexusClient` in `cmd/mavend/ecosystem.go`, `POST /api/v1/resolve` | `nexus.url` (`http://nexus:9740`) |
| **Praxis** | Operational attention and item lifecycle. What needs looking at, what changed, what is still unresolved. | `praxisClient`, the HTTP tools API under `/api/v1/tools/` | `praxis.url` (`http://praxis:8989`) |
| **Hexis** | The capability registry and the only path to executing anything. | vendored `github.com/kami/hexis/pkg/client` | `hexis.url` (`http://hexis:9741`) |
All three are `nil` unless configured, and every one of them degrades on its own.
An outage means a named gap in the answer, never a broken turn and never a guess.
Rules that are not negotiable:
- **No component reads another component's database.** Praxis attention comes over
HTTP, never from its SQLite file.
- **Identity lives in Nexus.** Do not invent a local fact key for something Nexus
resolves. `actionFact` already sets `Subject`, and `cmd/mavend/factenrichment.go`
resolves it in the background against Nexus.
- **Free text never reaches a mutating Hexis call.** Resolve to a canonical entity id
first. Ambiguous resolution asks the owner, it does not pick.
- **LLM output is not authorization.** Confirmation binds capability id, target
entity, arguments, requester and expiry. See `cmd/mavend/confirm.go`.
- **Praxis lifecycle words mean different things.** Surfaced is not acknowledged,
acknowledged is not resolved, execution success is not recovery. Reading an item
aloud calls `Surface`, never `Acknowledge`.
- **No automatic attention-to-action path.** Digestion may summarise Praxis. It may
not call Hexis.
Every cross-service call carries a correlation id minted once per action
(`withCorrelationID`), a contract version header, and `X-Requested-By: maven`.
## Routing — read this before touching the router
`internal/router/` has TWO layered engines. **The LLM router is now the default and it is
on in deploy** — this section used to say it was wired `nil`, which stopped being true on
2026-07-31.
- **LLM router (the intended design, REARCH.md):** the resident Qwen3-1.7B (`llmrouter.go`)
- **LLM router (the intended design, docs/rearchitecture.md):** the resident Qwen3-1.7B (`llmrouter.go`)
emits GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is
demoted from a routing gate to a RAG hint. Wired at `voice.go:214` via
`pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient)`; the flag is `voice.llm_router`
@@ -88,10 +127,13 @@ on in deploy** — this section used to say it was wired `nil`, which stopped be
Cascade order: `stage0.go` exact-match fast-path → LLM router (when non-nil) → classifier
fallback. Any LLM error falls through to the classifier so a turn never breaks on the model.
Measured on the 77-case RU fixture (`MODEL-BAKEOFF-31-07-2026.md`): the classifier scores
Measured on the 77-case RU fixture (`docs/evals/2026-07-31-model-bakeoff.md`): the classifier scores
36.8% full accuracy at p50 31ms; Qwen3-1.7B scores 67.5% intent-only / 72.7% through the
cascade at p50 ≈2.7s. Accuracy roughly doubled, latency is ~90× worse, and that trade was
accepted deliberately. `Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
cascade at p50 ≈825ms. Accuracy roughly doubled, latency is ~27× worse, and that trade was
accepted deliberately. **The ≈2.7s figure that stood here until 2026-08-02 was contention,
not the model.** See `docs/evals/2026-07-31-routing.md` line 61, which measures the LLM router at
p50 825ms / p95 1.2s / max 3.0s and the full cascade at p50 0.80-1.04s. Do not plan latency
work off the bakeoff table. `Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
#359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with
no allowlisted fn) feeding the same stage-3 gate the classifier path already had — see
@@ -102,8 +144,25 @@ Re-measured on the fixture after the fix: **missed clarify 6/6 → 1**, at the c
clarifies and 2.6pt of full accuracy (72.7% → 70.1%, intent-only 67.5% → 74.0%). Two of the
three false clarifies are acts the model mis-routed and the gate caught — asking beats wrongly
executing, so the fixture and the daemon disagree about what is correct there. The third,
`"поужинал"`, is a real defect: **the single-token rule is an English intuition and does not
transfer to Russian**, where one word is routinely a whole sentence. Narrow or drop it.
`"поужинал"`, was a real defect: the single-token rule was an English intuition and does not
transfer to Russian, where one word is routinely a whole sentence.
Narrowed 01-08-2026. `thinSingleToken` (`internal/router/singletoken.go`) still thins a bare
one-word nominal — "вода", "бэкап" — but spares two classes: a closed lexicon of social and
control singles ("привет", "спасибо", "стоп", "yes"), and any token carrying a Russian verb
ending (past tense, 2nd person, reflexive), because a verb already contains its subject. Both
tests are offline and cost nothing. Re-measured: **false clarifies 3 → 2, intent-only 74.0% →
75.3%, full accuracy unchanged at 70.1%, missed clarify still 1.** The two remaining false
clarifies are the act-with-no-allowlisted-fn arm of the gate, not this rule.
Agenda questions taken off the model, 01-08-2026. `AgendaQueryGrammars` (`stage0.go`, wired
after the clock rules in `buildRouter`) routes "что у меня сегодня", "во сколько у меня
встреча" and anything naming a calendar to `IntentQuery` at stage 0. They were going to
`IntentSystem`, where `replySystem` has no agenda arm and answered "пока не умею" — the
fixture had said `query` since ru-query-019 was written. Measured: **full accuracy 70.1% →
72.7%, intent-only 75.3% → 77.9%, calendar 0/2 → 2/2**, clarify counts unchanged. Note that
Go's `\b` is ASCII-only and never fires after a Cyrillic letter; the pattern needs an
explicit `(\s|[?!.]|$)`.
## LLM output contract
@@ -129,10 +188,16 @@ world questions, so she needs to read external sources. What replaces it:
- **No telemetry, no cloud model, no third-party account.** That part never changes. Nothing
about Maven is reported to anyone, and inference stays on the box.
- **Local sources first.** Kiwix ZIMs on homesrv (Wikipedia, ifixit) before anything on the
network. Reading beats recalling for a small model, and a local read costs nothing.
- **His data first, then the world.** Every source that reads his facts, notes, calendar,
tasks or house runs before anything outside, and the personal boundary sits between them.
Reading beats recalling for a small model.
- **In the world, live search leads and the ZIMs are the fallback** (owner's call,
2026-08-02). A self-hosted SearXNG (`search` block) answers first; the Kiwix ZIMs on
homesrv answer when the search is empty, unreachable, or the line is down.
- **External search is allowed and off unless configured**, like the weather and telegram
capabilities.
capabilities. The code default is still off. `deploy/mavend.json` now ships a `search`
block (owner's call, 2026-08-02), so it is on for this box and deleting the block turns
it off again.
- **His notes and facts are never search input.** Looking up why the sky is blue and sending
his stored personal notes to an upstream engine are different acts. Only the utterance goes
out, never the persona block, history, or matched notes.
+2 -2
View File
@@ -78,7 +78,7 @@ deps-go:
done
$(GO) version
# fmt-check fails if any file needs gofmt. DESIGN.md has always said `make
# fmt-check fails if any file needs gofmt. docs/design.md has always said `make
# test` gates on gofmt and vet; it did not, so nine files quietly drifted.
# Run `gofmt -w` on whatever this prints.
fmt-check:
@@ -191,7 +191,7 @@ deps-piper:
# multilingual-e5-small: an asymmetric retrieval model. It is trained to match
# a short question against a longer passage, which is what note recall is.
# The quantized file is the one we download, deploy and measure — see
# RECALL-EVAL-31-07-2026.md.
# docs/evals/2026-07-31-recall.md.
EMBEDDER_DIR := $(shell pwd)/models/embedder/multilingual-e5-small
EMBEDDER_MODEL_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/onnx/model_quantized.onnx
EMBEDDER_TOKENIZER_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/tokenizer.json
-468
View File
@@ -1,468 +0,0 @@
## Maven — current state (updated 2026-07-20)
### Session 2026-07-20 — ecosystem hardening + entity-aware facts
Ten commits, focused on closing the Nexus/Praxis integration gaps flagged
as "wired but immature" in the prior review, plus the entity-aware-memory
backlog item (`20-07-2026-BACKLOG.md` item 5).
- **Entity-aware fact resolution (Vikunja #279)** — facts gain
`Subject`/`EntityID`/`ResolutionState`; an async worker resolves
free-text subjects to canonical Nexus entity_ids (mirrors Praxis's own
enrichment pattern). Ambiguous/unreachable Nexus never guesses — stays
`pending` or terminal `ambiguous`. Voice-tapped facts (`IntentFact`) flow
into the queue automatically via an optional `Subject` field on
`WriteFactReq` (old callers unaffected, no signature break).
- **Typed Praxis lifecycle client (Vikunja #271)** — `GetItem`/`Search`/
`Surface`/`Acknowledge`/`Resolve`/`Ignore`/`Pin`, routed through new RU/EN
dialogue verbs. Fixes a real lifecycle-invariant bug: reading an
attention item aloud now calls `Surface` — previously the digest path
read items without recording that they'd been surfaced, so "Maven
mentioned it" was indistinguishable from "never came up."
- **Durable delivery outbox (Vikunja #270)** — `BeginDeliveryAttempt`
before `Send`, `CompleteDeliveryAttempt` after; a stale `pending` row
found at startup reconciles to `unknown` (never silently resent or
dropped — same rule as Hexis's execution-timeout handling). Closes a
crash-window duplicate-send bug. Wired into `DispatchNudge`,
`DispatchReminder`, `RepeatUnacked`; reconciliation runs once at boot
before the tick loop resumes.
- **Fail-closed IPC/dependency handling (Vikunja #269, #272/#273)** —
ambiguous mutation outcomes (frame sent, reply lost) no longer blindly
retry; Nexus/Hexis dependency errors fail closed instead of guessing.
- **Correlation IDs + version headers (Vikunja #273)** — the hand-rolled
Nexus/Praxis HTTP clients now send `X-Nexus-Version`/`X-Praxis-Version`
and thread the same correlation ID already generated in
`executeCapability` through the whole call chain, matching the Hexis
client's existing behavior.
- **Entity-scoped Praxis attention queries** — callers holding a resolved
entity_id can ask "what needs attention for this entity" directly
instead of filtering the unscoped list client-side.
- **Fake-ecosystem test harness with fault injection** — a reusable
`fakeServer` (Nexus/Praxis/Hexis fixtures, runtime-toggleable
`SetFault`, fake clock) replacing ad-hoc per-test `httptest` servers;
covers a gap that had zero test coverage (`handlePraxisAct`) and adds a
fault-then-recovery regression test for the fail-closed fixes above.
- **Ops fix** — `deploy/mavend.json`'s phraser was pointed at a 4B model
with `n_gpu_layers=99`, which OOM'd under memory pressure and left a
zombie `llama-server` child; swapped to the 2B Qwen model matching the
intended resident-model size.
Net effect: the Nexus/Praxis wiring described as "plumbing exists, thin
compared to Maven's test depth" in the prior review is now materially
hardened — typed clients, fail-closed error handling, durable delivery,
and a proper fault-injection test harness are all in place. Evaluation lab
(`20-07-2026-BACKLOG.md` item 7) is explicitly skipped for now — it runs
on the GPU box, which is occupied by CPT. Morning routine engine (backlog
item 3) is next up, not started.
---
> **Resolved 2026-07-30 (task #318).** The resident checkpoint is
> **Qwen3.5-0.8B** (`Q4_K_M`), set in `deploy/mavend.json`; the **target** is
> the locally CPT'd **Qwen3-1.7B**, still training (#122). Older model claims
> below — the LFM references, the pipeline line, and the "swapped to the 2B
> Qwen model" ops entry above — are historical. Read them as a log of what was
> true at the time, not as current fact. Note also that `/mnt/hdd1/llms` is
> bind-mounted over `models/llm/`, so the LFM2.5 gguf in the repo tree is
> never loaded.
Architecture decision (as written on 2026-07-20): the target resident
router/phraser is the locally trained Qwen3-1.7B model — still the target as
of 2026-07-30. Older LFM references below describe the then-deployed
historical stack, not the target checkpoint. RU CPT has a successful
full-weight checkpoint at step 1000/8077; evaluation and Qwen3 SFT tooling are
tracked in `docs/plans/2026-07-18-qwen3-resident-training-eval.md`.
Consolidated status. The reactive↔proactive core is closed and testable through
the web PWA. The former SPEC's open items 17 (now `DESIGN.md` § execution ledger) are landed (protocol doc, away-channel
fallthrough, CalDAV poller, quiet-hours schedule, tools enable/disable, note RAG,
passkey step-up); item 8 (multi-user) is deliberately deferred — see the tail.
The two big infra gaps from the jul5 revision are closed on `overnight-jul5`:
**at-rest encryption** (AES-256-GCM, tmpfs working copy — not sqlcipher, see
`internal/store/crypt.go`) and **Docker deployment** (one image, six daemon
containers). The `overnight-jul6` session (now on `master`) closed the biggest
*query-surface* gaps — **calendar querying, general-knowledge answers, and
weather** — plus a populated homelab act allowlist and two pure scaffolds
(dialogue state, long-term-memory vector store). ~15.2k LOC + ~8.5k test, 303
tests, `-race` in `make test`.
### Access model
- **Phone** → needs the wg tunnel to reach homesrv (no homesrv DNS otherwise;
raw IP or a DNS tweak can bypass, not the default).
- **PC** → uses homesrv DNS, resolves the domains over local-net, **no wg needed**.
- nginx + ufw both scope to `10.42.0.0/24` (wg) + `192.168.1.0/24` (LAN), deny all else.
- **Surface in use now: the web PWA (`mavweb`).** Voice PTT + in-app nudges both ride it.
### Works end-to-end (tested)
- **Reactive voice:** PWA record → Whisper STT (`mavsttd`) → ONNX classifier →
resident phraser (llama-server subprocess; Qwen3.5-0.8B as of 2026-07-30 —
this line historically named "LFM 2.5-1.2B") → Piper TTS
(`mavttsd`) → reply.
HTTP POST path (mobile-Chrome drops WS for the audio).
- **Capture:** `fact` (EN **and RU** — root-substring recognizers) + `reminder`
persist through CoreAPI (`source=tap:voice`). This is the substrate the care
rules read.
- **Notes / query (semantic recall, sqlite — no chroma):** `note` → embed (the
classifier's ONNX embedder) → `notes` table. `query` → embed → brute-force
cosine top-k → confidence-gated (below `queryMinScore` 0.55 ⇒ "no note", not a
guess). **Note RAG (SPEC item 6):** the gated top-k feed the phraser
(`PhraseQuery`) to compose a natural answer ("вот что я нашла: …") instead of
a verbatim dump; raw-notes fallback on any LLM error. Stub is deterministic.
- **Monitoring (`/dash`):** mavweb server-renders presence + recent nudges (by
outcome) + recent facts from the append-only store via CoreAPI. Read-only,
meta-refresh, no JS.
- **Proactive loop:** 60s dumb ticker, pure predicates over a State snapshot,
universal gate (quiet-hours/presence/cooldown/snooze/calendar), one-nudge-per-
tick max-severity, reminders (gate-bypassing), sev4 repeat-til-ack, feedback
auto-tuner (outcome ratio → bounded cooldown, persisted as `source=feedback`).
- **Rules:** water/meal/break (sev12 care), service_down (sev4, `poll:uptimekuma`),
netdata_critical (sev3, `poll:netdata`).
- **Routines (`internal/routine`):** operator-declared clockwork — the third
proactive class beside reminders (user-stated) and care rules (world-state).
Config `routines[]` (cron + literal RU body + severity) fire through the normal
dispatcher on schedule (an 08:00 briefing, a 22:00 wind-down). Bodies are
literal (not LLM-phrased ⇒ can't hallucinate); rule name `routine:<name>` so
they don't pollute the care autotuner; cold-start guard seeds on first sight so
a restart never replays a missed schedule. Pure `routine.Due`, unit-tested; the
tick driver holds the last-fired map.
- **Env facts (`mavpoll`):** netdata alarms → `netdata_alarm` (fires immediately
on a real CRITICAL); kuma monitor_status → `service_down`. Writes only on
value-change (no append-only churn).
- **Presence:** noisy-OR decay + Schmitt hysteresis. Live via `page_heartbeat`
(PWA auto-pings `/api/signal` every 30s → present when a tab's open).
- **Delivery:** ntfy / telegram / voice by `f(severity, presence)`; minimal body
on away channels. PWA subscribes to ntfy over **WebSocket** for in-app nudges.
- **Away-channel fallthrough (SPEC item 2):** when the router picks voice but no
live session exists at push time (presence guess was wrong), the dispatcher
reroutes through the AWAY table — sev3→ntfy, sev4→telegram-repeat-til-ack,
sev≤2→drop — instead of silently dropping. Covers nudges + reminders.
- **Calendar busy (SPEC item 3, `mavcaldav`):** new poller queries a self-hosted
**Radicale** CalDAV server on an interval, writes `calendar_busy` + event facts
through CoreAPI (value-change only). The loop gate already consumes `calendar_busy`.
- **Quiet-hours schedule (SPEC item 4):** the gate reads `quiet_hours`; a config
time window (`voice.quiet_hours`, HH:MM, midnight-crossing handled) now sets it
on each tick — in addition to the "тихий режим" voice toggle. Both activate quiet.
- **Client protocol (SPEC item 1):** the voice wire format (length-prefixed JSON
frames) is published in `PROTOCOL.md`, generated from `internal/voice/wire.go`
so third-party clients don't need the Go source.
- **Passkey step-up (SPEC item 7):** `internal/webauthn` does real WebAuthn —
ES256/P-256 register + assert, ecdsa signature verification, rpIdHash + UP/UV
flag binding (UV = the gesture), sign-count regression check. `PasskeySession`
bumps the auth session L2→L3 for a TTL on assert. mavweb serves `/auth/passkey`
(enroll + step-up) + the begin/finish endpoints. Crypto is round-trip tested
(incl. tampered-sig / missing-UV / wrong-origin negatives).
- **Stability:** llama-server orphan leak fixed (`Pdeathsig` kills the child on
any mavend death); `kill-maven.sh` reaps strays (matches the model, not a
bogus `llama-server.*maven` pattern); `start-maven.sh` wires `-core` + poller.
### Wired but needs a deploy action (not code)
- **`desk_active`** (strongest presence signal) — `scripts/desk-active.sh` runs
on the **desk PC** (hypridle-gated systemd timer), posts over wg to mavweb.
- **`mavwaked`** (always-on listening) — needs a systemd user unit on a client
box (desk PC, pi, etc.) where the mic is attached. Connects to mavend over wg
or local net via `-addr`. Deferred until a client box is wired with a mic.
Caveats / gotchas:
- **desk_active is a workstation deploy, not code** — 0 facts ever written; presence
runs on page_heartbeat alone (dash reads "away"/"never at desk"). `scripts/desk-active.sh`
+ a hypridle-gated `maven-desk` timer must be installed on the desk PC (not homesrv).
- **Notes recall needs the ONNX embedder** — under the HashEmbedder floor, cosine is
lexical (token overlap), not semantic; scores are low, so most RU commands sit under
the 0.35 route threshold and clarify. Configure `voice.embedder` for confident recall+routing.
(The floor now at least tokenizes Cyrillic — see below — so it ranks correctly, just weakly.)
- **Switching the embedder model silently breaks old notes** — different dim ⇒
cosine 0 ⇒ they stop matching; brute-force can't re-embed. Re-embed on a model change.
- **`wg_handshake` is OFF and should stay off** — in this topology the phone only
runs wg when *outside*, so a fresh handshake means AWAY, not here. The `mavpoll
-wg` flag exists (defaults `""`) and could later back the spec's "away override"
by flipping the sign; as a presence-*here* signal it's inverted. desk_active +
page_heartbeat cover home presence.
- **Cold-start unlock tests are missing** — the key wrap/unwrap code
(`internal/webauthn/keywrap.go`) and locked-mode IPC gating (`cmd/mavend/main.go`)
are correct but have **zero test coverage**. The roadmap (item 2.1) required
three new test cases (wrap/unwrap round-trip, wrong-cred unwrap fails,
locked-mode IPC rejects non-unlock methods); none were written. `make test`
is green by omission. Write these before relying on the cold-start path with
real keys.
### Done since last revision (overnight-jul6, 2026-07-06)
Seven tasks (session board `SESSION-06-07-2026.md`, deleted 2026-07-30 — see git history), one commit each, merged to `master`.
This session was run through **opencode**, not Claude Code (co-author trailer).
Since then (**2026-07-06, second session**):
- **Always-on listening (gap 1, MVP)** — `cmd/mavwaked/`: 825 lines, 10 `-race`
tests. Energy-based VAD over 30ms windows (same RMS threshold as mavsttd's
`gateReason`), adaptive noise floor, speech→silence state machine. Captures
PCM from arecord(1) subprocess, sends `PushToTalk` with `Surface=SurfaceVoice`
(L0 — no destructive acts). Reply plays through aplay(1). No wake word yet
(pure VAD trigger); the 30ms frame shape matches silero-vad ONNX input 1:1,
so swapping energy-threshold for ONNX inference is a local change in vad.go.
`Makefile` `build-waked` target. Runs on client boxes (not docker/homesrv)
via systemd user unit; connects to mavend over wg or local net.
Since then (**2026-07-06, third session** — roadmap execution agent):
- **Cold-start unlock (ROADMAP 2.1)** — the at-rest AES key is now wrapped
(HKDF-SHA256 + AES-256-GCM, stdlib-only — no `x/crypto` dep) with the passkey
credential's public key and persisted to disk. At boot, if a wrapped key file
exists AND no env key is set, mavend starts **locked**: the IPC server runs
but `srv.Check` rejects everything except `MethodAssertStepUp` +
`MethodUnlock`. A passkey assertion at `/auth/passkey` calls `MethodUnlock`
with the credential's public key → unwraps the blob → opens the store → wires
voice/loop/delivery → `srv.SetAPI` swaps the locked stub for the real
CoreAPI. mavweb's `RegisterFinish` wraps the env key on enrollment;
`AssertFinish` calls `Unlock` on assertion. Env-key fallback preserved
(dev/CI path unchanged). **Test gap:** the roadmap required three new test
cases (wrap/unwrap round-trip, wrong-cred unwrap fails, locked-mode IPC
rejects non-unlock methods) — none were written. The code is correct but
untested; `make test` is green by omission, not coverage.
- **Conversation depth (ROADMAP 3.2)** — cross-intent anaphora + fact-by-key
lookup. `AnaphoraResolver` in `router/slots.go` detects RU pronouns
(это/он/она/оно/тот/мой + inflected forms). `followUpMerge` now handles
three cases: same-intent slot inheritance (existing), cross-intent anaphora
(Query/Fact/Reminder after a Fact with a pronoun inherits the prior key +
time), and query-after-fact (a query following a fact inherits the key for
fact-by-key lookup). `Session.History []Turn` added as the multi-turn
scaffold (capped at 4). 7 new test cases including the exact done-when
scenarios (anaphora query-after-fact, three-turn break, explicit-key-wins).
- **Routing quality + persona (ROADMAP 4.1/4.4)** — `QueryMinScore` is now a
config knob (`voice.query_min_score`, default 0.55) instead of a hardcoded
const. `make download-embedder` fetches Xenova/paraphrase-multilingual-
MiniLM-L12-v2 (~90MB ONNX) + tokenizer; AGENTS.md documents the embedder +
libonnxruntime setup. `Persona` field in `VoiceConfig` prepends to every
LLM system prompt (nudge phrasing, note queries, general knowledge); empty
= current hardcoded feminine-gendered Russian persona. Also fixed two
pre-existing data races found by `-race`: `voice/server.go` wg.Add vs
wg.Wait (accept mutex), `mavweb/server.go` s.api field (atomic.Value).
- **Calendar querying (task 3)** — "что у меня завтра?" now answers from the
CalDAV facts the poller already writes. Added `store.CalendarEvents(from,to)`,
a RU date-scope parser («сегодня»/«завтра») in `router/slots.go`, and an
IPC `CalendarEvents` RPC (api/client/server/wire) feeding the `IntentQuery`
handler. Empty day → «на сегодня ничего нет». Previously calendar only *gated*
nudges; it's now queryable.
- **General-knowledge routing (task 4)** — when notes-RAG misses `queryMinScore`,
the query now falls through to the phraser with an anti-hallucination system
prompt (`router.KnowledgePrompt`, single tested source) instead of giving up.
Empty/errored/Stub phraser → «не знаю.», never a fabrication.
- **Weather (task 5)** — new `internal/weather/`: `Provider` interface, a stub
(«погода не настроена»), and a real **keyless Open-Meteo** provider (geocode +
current_weather, injectable `*http.Client`, mocked in tests — no live network).
Wired into `IntentQuery` (keywords погода/градус/температура) with a ~5s
context timeout; selected by `voice.weather.provider` ("open-meteo" | "" → stub).
- **Homelab act allowlist (task 2)** — `voice.tools` seeded with read-only acts
(`systemctl status`, `docker ps`, `uptime`, `df`, `free`, `journalctl` reads)
as `destructive:false` and mutating ones (restart/stop/start/reboot,
docker-restart/stop) as `destructive:true`. Guardrail verified: no dangerous
verb is `destructive:false`. RU phrasings seeded in `act.txt`.
- **Embedder config validation (task 1)** — a partially-filled `voice.embedder`
block (some of model/tokenizer/lib paths missing) is now a load error instead
of a silent fall-through to the Hash floor; the floor fallback logs explicitly.
- **Dialogue state scaffold (task 6)** — `internal/dialogue/`: `Session` +
TTL `SessionStore` + pure `InheritSlots`. **Now wired** (post-merge follow-up):
the voice handler carries slots across same-intent turns within a 2-min window
(`followUpMerge`, unit-tested) — bounded gap-filling, not full multi-turn yet.
- **Long-term memory interface (task 7)** — `internal/memory/`: `Store` interface
+ `InMemoryStore` (cosine). Wired into `IntentNote` (best-effort insert) and,
post-merge, into `IntentFact` (facts indexed) + `IntentQuery` (read-back after
notes-RAG misses). In-memory only — no persistent backend yet (gap #8).
Follow-ups (Claude Code, post-merge): gofmt'd `handlers_test.go` (the jul6
verification commit left it misaligned, so `gofmt -l` still flagged it despite the
"all gates green" claim); deduped the task-4 knowledge prompt to the single tested
`router.KnowledgePrompt()`. Tree is now genuinely green (gofmt/vet/303 tests).
### Done since the jul5 revision (overnight-jul5, 2026-07-05)
The overnight session (`SESSION-05-07-2026.md`, deleted 2026-07-30 — see git history; 25 tasks) closed the previous
"not built yet" items 13 and added feature depth:
- **At-rest encryption** — the on-disk db is AES-256-GCM ciphertext; the daemon
works on a tmpfs (RAM) plaintext copy, sealed back atomically on close. Wrong
key / tamper ⇒ fail closed, never a plaintext fallback. Legacy plaintext dbs
upgrade on first clean shutdown. Key via config/env (`db_key_env`); no KDF —
raw 32-byte key, base64. The passkey cold-start unlock plugs into the same
`store.OpenEncrypted` seam later.
- **Docker deployment** — single image, one container per daemon
(`docker-compose.yml`); only mavend mounts the key + db volume; IPC over a
shared socket volume. `ipc.DialWait` (boot-order tolerance) + redial-on-drop
(core restarts don't kill modules). `deploy/README.md` has the runbook.
- **Tests** — mavcaldav, mavttsd, voicesink, mavweb main/handlers covered;
`make test` runs `-race -coverprofile`.
- **Recurring reminders** — `cron` + `next_fire_ts` on reminders; recurring ones
reschedule (instead of mark-fired) after successful delivery.
- **Notification digest/batching** — low-severity nudges queue and flush as one
digest per window/max-items (`digest` config block); stale-reminder bursts on
boot collapse into a single digest reminder, completed only after delivery.
- **Rule trace engine** — `ExplainTick`/`ExplainGate` record per-rule
predicate/gate/selection results each tick; served over IPC (`tick_trace`)
and rendered at mavweb `/trace` ("why didn't she nudge me").
- **Web UI** — new `/history` (facts + revert buttons), `/notifications` (nudge
history), `/trace` pages; nav links on `/dash`; RU/EN cheatsheet toggle in the
PWA; manifest icons (`icon.svg`). POST `/tools` now requires an in-process
passkey step-up when WebAuthn is configured.
- **Revert/undo** — `RevertFact` voids the latest fact for a key (append-only
void-marker, audit trail intact); exposed at `/api/revert` from `/history`.
- **Tool scopes** — `scope` column on tools, threaded through propose/enable/UI.
`DisableTool` raised to AuthStepUp alongside Enable.
- **Passkey persistence** — mavweb credentials in a JSON file (`-passkey-file`),
surviving restarts; rollback-on-persist-failure keeps memory and disk in sync.
- **STT silence gate** — min-duration + RMS floor drop non-speech before whisper
hallucinates on it (`-min-ms`, `-silence-rms` flags on mavsttd).
- **Housekeeping** — `db_key.env` gitignored (+`.env.example`), `build-caldav`
target, zero-timestamp "never" fix on /dash.
### Not built yet (ranked by ROI)
1. **Multi-user (SPEC item 8)** — deliberately deferred, see the tail.
Closed (jul6 follow-ups): `/api/revert` now sits behind the same passkey
step-up as POST `/tools`; `go.mod` direct deps (`onnxruntime_go`,
`coder/websocket`, `robfig/cron`) are labeled correctly — `go mod tidy` can't
run here because it walks the vendored `deps/go` toolchain tree.
Purge+rotate leaked db key (#12) — investigated and closed: the key was
**never committed** to git history (gitignored at introduction, no commit
ever tracked `deploy/db_key.env`), so nothing to scrub. File stays on disk
and in deploy env by design — at-rest encryption needs it at boot.
Done earlier (2026-07-03): **act tool executor, store-backed, full flow**
(`internal/tool` + `internal/store/tools.go` + `tools` CoreAPI methods).
- **Execution:** IntentAct runs the matched fn against the store's ENABLED
allowlist. argv, no shell → STT text can't inject. Live store read, so a
newly-enabled tool runs without a daemon restart.
- **proposed→enabled→disabled (SPEC item 5):** an act whose verb isn't enabled is
scaffolded as a `proposed` tool (maven suggests). A human enables it (fills argv
+ destructive) on the authed **`mavweb /tools`** page — never voice — and can
disable it back to `proposed` (kept in the store, won't run). `EnableTool`/
`DisableTool` sit at `AuthStepUp`; the gate is now **live** via `PasskeySession`,
so /tools enable requires a passkey assertion at `/auth/passkey` first.
- **Confirm turn:** a destructive enabled tool replies "выполнить X? да/нет" and
parks; the next utterance (ru/en yes-no) confirms or cancels (90s TTL).
- **Config:** `voice.tools` seeds enabled tools at boot (editing mavend.json =
the human enable act); mavweb enables ad-hoc ones on top.
- **Russian:** fixed grammar in reply strings + seed files; maven's self-
reference is feminine ("she") — [[maven-persona-gender]].
Also fixed:
- **HashEmbedder was blind to Cyrillic** (`tokenize` iterated bytes, kept only
`a-z0-9`) → every RU utterance embedded to the zero vector → cosine 0 across
all intents → misrouted to `act` (alphabetical tie-break). Now rune-based
(`unicode.IsLetter`). This was the real cause of "Найди заметку" (a query)
landing in `notes`; added note-retrieval query seeds too.
- **Notes are now browsable on `/dash`** — `RecentNotes` plumbed through the
store + CoreAPI; voice-captured notes were previously only reachable via
semantic `query`.
Earlier: notes/query recall, `/dash` monitoring, `wg_handshake` poller (NO-OP).
### Gaps — why "voice assistant" is still aspirational (2026-07-06)
What separates Maven today from the thing the spec describes. Dealbreakers
first — these define the category:
1. **Always-on listening is code-complete (MVP).** `cmd/mavwaked` captures
PCM from arecord → energy-based VAD → PushToTalk with `Surface=SurfaceVoice`
(L0). Gap narrowed: no wake word yet (pure voice-activity trigger; every
utterance fires). The 30ms frame shape and 16kHz PCM match silero-vad's
ONNX input exactly, so a wake-word model swap is a local change in vad.go.
Hardware: the mic lives on a client box (desk PC, pi, etc.) — never the
homesrv. Deploy action: systemd user unit on whichever box has the mic,
connects to mavend over wg or local net.
2. **Conversation is deeper now, still not full dialogue.** The router
classifies one utterance → one reply, but `internal/dialogue` carries
context across turns: a 2-min session inherits slots for same-intent
follow-ups («напомни завтра» → «…позвонить маме»), and cross-intent
anaphora («запиши что я пил воду» → «когда я это сделал?») now resolves
RU pronouns (это/он/она/оно/тот/мой + inflections) to the prior turn's
key for fact-by-key lookup. `Session.History []Turn` is the scaffold for
real multi-turn. Still missing: LLM-driven dialogue manager (decide
ask-vs-act), anaphora beyond RU pronouns, single-slot session (single-user
box). The sub-1B phraser only words replies.
3. **Latency/shape of a turn.** Clip-based STT (record → upload → whisper →
route → phrase → piper → play). No streaming either direction, no barge-in;
every exchange is a full round trip.
Capability-class gaps — built but thin:
4. **Act surface is a small argv allowlist.** propose→enable works and the
allowlist now ships a homelab starter set (jul6 task 2 — status/ps/uptime/
df/free/logs read-only, restart/stop/reboot gated). Still bounded to what's
seeded; broadening it is config, not code.
5. **Query answers now cover notes + calendar + weather + general knowledge**
(jul6 tasks 3/4/5). Calendar querying, keyless Open-Meteo weather, and a
phraser knowledge-fallback all landed; caveat — general-knowledge quality is
only as good as the sub-1B phraser, and weather needs `voice.weather.provider`
set. The cheatsheet and router are now roughly aligned.
6. **Routing quality depends on the ONNX embedder being configured** — the
HashEmbedder floor makes RU recall lexical/weak; many commands fall to
"clarify". `make download-embedder` now fetches the multilingual MiniLM
model + AGENTS.md documents libonnxruntime setup; `voice.query_min_score`
is a config knob (default 0.55) so the floor can be tuned without recompile.
7. **Presence is effectively one signal** (page_heartbeat); desk_active is
still an undeployed script — "voice when near" routing runs on a guess.
8. **Long-term memory is now persistent (store-backed), not the spec's chroma.**
`internal/memory` has a `Store` interface; the daemon now wires
`store.MemoryStore` (`internal/store/memory.go`) — a **persistent** backend
in the **same encrypted sqlite db** (survives restarts; recall text inherits
at-rest encryption, so no plaintext sidecar). Vectors are float32 blobs,
search is brute-force cosine (fine at single-user scale; ANN is the later
swap behind the same interface). Notes **and facts** are indexed on capture;
`IntentQuery` reads it back (after notes-RAG misses, before general-knowledge)
— fact recall («когда я пил воду?») is its distinct payoff. The in-memory
impl remains the test/no-store floor. Remaining: an ANN/external index is
optional-scale, not a gap. Custom TTS voice (kami-picked, replaces the irina
floor — [[custom-voice-training]]) is still a future item.
Ops footnote: voice-over-web verified 2026-07-06 — mavend binds 0.0.0.0:9100
and mavweb reaches it cross-container at mavend:9100 (nc -z confirmed).
mavpoll uses network_mode=host to reach localhost services (netdata, kuma).
### Future / logged, not now
Custom TTS voice training (kami-picked voice, replaces irina floor); listening
modes 23 (meeting-record, ambient-derive).
### Services & layout
- `mavend` (core, IPC unix socket) — store + loop + phraser; the only key-holder.
- `mavsttd` / `mavttsd` — STT/TTS worker modules (unix sockets).
- `mavweb` — PWA bridge (HTTP), `/api/ptt` voice, `/api/signal` presence ingest,
`/api/ntfy` WS-subscribe config, `/dash` read-only monitoring.
- `mavpoll` — env poller (netdata/kuma → facts via CoreAPI).
- `mavcaldav` — CalDAV poller (Radicale → `calendar_busy` + events via CoreAPI).
- All behind wg + nginx deny-all; no phone-home. CGo only in `mavsttd`.
- Start/stop: `./start-maven.sh [build]`, `./kill-maven.sh`.
- Config: `~/.config/maven/mavend.json` (or `mavend.json` in repo root).
### Key files
- `cmd/mavend/{main,tick,voice}.go` — daemon wiring, loop driver, voice handler
- `internal/loop/{loop,rules,gather,feedback}.go` — proactive engine
- `internal/store/` — append-only facts/reminders/nudges/presence/notes
- `cmd/mavweb/{main.go,dash.html}` — PWA bridge + `/dash` monitoring
- `internal/router/{classifier,slots,stage0}.go` — reactive routing + slot parse
- `internal/delivery/` — dispatcher + ntfy/telegram/voice sinks
- `internal/auth/` — scope/gate/policy; `FloorEnrollment` (same-uid = device
trust) + `webauthn.PasskeySession` (real step-up for L3)
- `internal/webauthn/`, `cmd/mavweb/webauthn.go` — passkey register/assert
- `cmd/mavcaldav/`, `cmd/mavpoll/`, `scripts/desk-active.sh` — env producers
### Why multi-user (SPEC item 8) is deferred
Not neglect — the one item where doing nothing now beats doing something:
- **No second user exists yet** (the "gf phase"). Building per-user partitioning
now means code exercised by zero users and validated by nobody — YAGNI.
- **The append-only schema makes it a migration, not a rewrite.** No row is ever
mutated, so adding `facts/notes/reminders.user_id` later is add-columns +
backfill-to-"kami" — no reshaping, no dual-write window. Deferral is cheap.
- **The hard part is speaker attribution, and it needs the second voice.** A
voice-print discriminator (kami vs gf vs unknown) can't be trained or tuned
with one voice in the house. Plumbing before the model is pipe with no water.
- **It's fenced deliberately** (`DO NOT TOUCH THIS PHASE` in `DESIGN.md` § Users) so an
autonomous agent doesn't add `user_id` columns while touching the store and
commit us to a schema before the constraints that shape it exist.
+1 -1
View File
@@ -1,6 +1,6 @@
// Package main is mavenclient — maven's reference client.
//
// Per DESIGN.md § Voice pipeline (STT / TTS): capture lives on the client;
// Per docs/design.md § Voice pipeline (STT / TTS): capture lives on the client;
// the server transcribes + synthesises on demand. The PC client runs the
// wake-word / VAD gate (cmd/mavwaked) and ships ONE clean audio blob per
// utterance on activation. The server never owns a mic.
+128
View File
@@ -0,0 +1,128 @@
// Spoken ack — the other half of the snooze wire. "готово" said out loud
// resolves a live nudge as `acted`, and a fact that answers the nudge on its
// own ("выпил воды" after the water rule fired) closes it without him having
// to say anything extra.
//
// Two entry points rather than one, because the two utterances are different
// acts. A bare "готово" carries no content and is intercepted before the
// router, exactly like the snooze. "выпил воды" IS content: it has to route
// normally and write its fact, and only then close the nudge. Folding the
// second into a pre-route intercept would have thrown the fact away, which is
// the thing he actually said.
package main
import (
"context"
"log"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
)
// resolveAck — pre-route keyword check for a contentless acknowledgement,
// run after the snooze. Same window and same fall-through rule: the words only
// count when a nudge is actually live, so "готово" with nothing pending routes
// normally.
func (h *reactiveHandler) resolveAck(ctx context.Context, text string, src turnSource) (string, bool) {
if !classifyAck(text) {
return "", false
}
now := h.now()
target, ok := h.pendingNudge(ctx, now)
if !ok {
return "", false
}
if err := h.api.ResolveNudge(ctx, target.ID, store.NudgeActed, now); err != nil {
log.Printf("voice: ack nudge %d (%s, %s): %v", target.ID, target.Rule, src, err)
return "не получилось отметить.", true
}
log.Printf("voice: acked nudge %d (rule %s) from %s", target.ID, target.Rule, src)
return "отлично, отметила.", true
}
// ackFromFact — post-action hook, called once the turn's decision has been
// applied. A fact whose key is the substrate of a live nudge's rule answers
// that nudge, so the nudge is resolved `acted` and the auto-tuner learns the
// rule is working.
//
// Silent by design: it returns nothing and never changes the reply. He said
// "выпил воды" and the fact reply is what he is owed; "отлично, отметила" on
// top would be her congratulating him for obeying, which is the nag she is
// explicitly not.
//
// Best-effort throughout. A failure here loses one feedback signal and must
// never turn a written fact into an error the user hears.
func (h *reactiveHandler) ackFromFact(ctx context.Context, dec router.Decision) {
if dec.Clarify || dec.Intent != router.IntentFact || !dec.Slots.HasKey {
return
}
rules := ackRulesForKey(dec.Slots.Key)
if len(rules) == 0 {
return
}
now := h.now()
target, ok := h.pendingNudge(ctx, now)
if !ok || !rules[target.Rule] {
return
}
if err := h.api.ResolveNudge(ctx, target.ID, store.NudgeActed, now); err != nil {
log.Printf("voice: ack nudge %d from fact %q: %v", target.ID, dec.Slots.Key, err)
return
}
log.Printf("voice: nudge %d (rule %s) acked by fact %q", target.ID, target.Rule, dec.Slots.Key)
}
// ackRulesForKey — which rules a fact under this key answers.
//
// Derived from each rule's InertWhenNoData rather than written out as a map,
// so a rule added later is covered the day it lands. That field already names
// the substrate the rule reads; a fresh fact under one of those keys is by
// definition the thing the rule was complaining about the absence of.
//
// DefaultRules, not the daemon's wired set: a rule disabled in config cannot
// have a pending nudge to close anyway, and reading the canonical set here
// keeps this free of the config plumbing.
func ackRulesForKey(key string) map[string]bool {
if key == "" {
return nil
}
var out map[string]bool
for _, r := range loop.DefaultRules() {
for _, k := range r.InertWhenNoData {
if k != key {
continue
}
if out == nil {
out = map[string]bool{}
}
out[r.Name] = true
}
}
return out
}
// ackPhrases — the acknowledgement vocabulary, as stem sequences. Matched by
// quietPhrase (quiet_toggle.go), so a single-word pattern matches only a
// single-word utterance.
//
// "да" and "ок" are deliberately absent. Both are answers to a question she
// asked, and the clarify gate upstream (resolveClarifyAnswer) has the stronger
// claim on them; letting them close a nudge as well would mean a stray "да"
// silently rewrites the feedback the auto-tuner learns from.
var ackPhrases = [][]string{
{"готово"}, {"сделал"}, {"сделано"}, {"выполнил"}, {"уже"},
{"уже", "сделал"}, {"уже", "готово"}, {"всё", "сделал"},
{"done"}, {"already", "did"},
}
// classifyAck reads an utterance as a contentless acknowledgement.
func classifyAck(text string) bool {
tokens := quietTokens(text)
for _, p := range ackPhrases {
if quietPhrase(tokens, p) {
return true
}
}
return false
}
+109
View File
@@ -0,0 +1,109 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
)
func TestClassifyAck(t *testing.T) {
for _, s := range []string{
"готово", "сделал", "сделано", "выполнил", "уже",
"уже сделал", "всё сделал", "done",
} {
if !classifyAck(s) {
t.Errorf("classifyAck(%q) = false, want true", s)
}
}
for _, s := range []string{
// "да" and "ок" belong to the clarify gate, not to the nudge.
"да", "ок", "хорошо",
// A single-word pattern must not eat the sentence it appears in.
"сделал бэкап базы", "готово ли обновление", "уже поздно",
"напомни завтра позвонить маме", "",
} {
if classifyAck(s) {
t.Errorf("classifyAck(%q) = true, want false", s)
}
}
}
func TestResolveAckMarksTheNudgeActed(t *testing.T) {
h, api := snoozeHandler([]ipc.Nudge{pendingNudgeAt(6, 2*time.Minute)})
reply, handled := h.resolveAck(context.Background(), "готово", sourceVoice)
if !handled || reply == "" {
t.Fatalf("got (%q, %v), want a reply", reply, handled)
}
if api.gotID != 6 || api.gotOutcome != store.NudgeActed {
t.Fatalf("resolved (%d, %q), want (6, %q)", api.gotID, api.gotOutcome, store.NudgeActed)
}
}
func TestResolveAckFallsThroughWithNothingPending(t *testing.T) {
h, api := snoozeHandler(nil)
if reply, handled := h.resolveAck(context.Background(), "готово", sourceVoice); handled || reply != "" {
t.Fatalf("got (%q, %v), want fall-through", reply, handled)
}
if api.calls != 0 {
t.Fatalf("resolved a nudge with nothing pending")
}
}
func TestAckRulesForKey(t *testing.T) {
cases := []struct {
key string
want string // "" means no rule
}{
{"water", "water"},
{"meal", "meal"},
{"break", "break"},
{"desk_active", "break"},
{"weight", ""},
{"", ""},
}
for _, tc := range cases {
got := ackRulesForKey(tc.key)
if tc.want == "" {
if len(got) != 0 {
t.Errorf("ackRulesForKey(%q) = %v, want none", tc.key, got)
}
continue
}
if !got[tc.want] {
t.Errorf("ackRulesForKey(%q) = %v, want %q in it", tc.key, got, tc.want)
}
}
}
func TestAckFromFactClosesTheMatchingNudge(t *testing.T) {
h, api := snoozeHandler([]ipc.Nudge{pendingNudgeAt(11, time.Minute)}) // rule "water"
h.ackFromFact(context.Background(), router.Decision{
Intent: router.IntentFact,
Slots: router.Slots{Key: "water", HasKey: true},
})
if api.gotID != 11 || api.gotOutcome != store.NudgeActed {
t.Fatalf("resolved (%d, %q), want (11, %q)", api.gotID, api.gotOutcome, store.NudgeActed)
}
}
func TestAckFromFactIgnoresAnUnrelatedFact(t *testing.T) {
// The live nudge is "water"; a meal fact does not answer it. Closing it
// anyway would tell the auto-tuner the water rule works when he ignored it.
h, api := snoozeHandler([]ipc.Nudge{pendingNudgeAt(12, time.Minute)})
for _, dec := range []router.Decision{
{Intent: router.IntentFact, Slots: router.Slots{Key: "meal", HasKey: true}},
{Intent: router.IntentFact, Slots: router.Slots{Key: "weight", HasKey: true}},
{Intent: router.IntentFact}, // no key
{Intent: router.IntentQuery, Slots: router.Slots{Key: "water", HasKey: true}},
{Intent: router.IntentFact, Slots: router.Slots{Key: "water", HasKey: true}, Clarify: true},
} {
h.ackFromFact(context.Background(), dec)
}
if api.calls != 0 {
t.Fatalf("resolved %d nudge(s) on unrelated decisions", api.calls)
}
}
+275 -18
View File
@@ -5,6 +5,7 @@ import (
"errors"
"fmt"
"log"
"regexp"
"strings"
"time"
@@ -41,6 +42,20 @@ type queryTurn struct {
type querySource struct {
name string
answer func(*reactiveHandler, context.Context, *queryTurn) (string, bool)
// dateAware — this source reads the day out of the turn and answers for
// THAT day. Only such a source may claim a continuation ("а завтра?"),
// because a continuation is a question about a different day and nothing
// else. A date-blind source claiming one would answer with today's data
// under tomorrow's question, which is a wrong answer delivered in a
// confident voice — the failure mode that took reminder out of
// continuableIntents (continuation.go).
//
// Exactly one source qualifies today, and that is not an oversight in the
// table: CalendarEvents is the only CoreAPI call that takes a date at all.
// DayPlan is today-only, CurrentWeather is now-only, and the recall
// sources search text with no notion of a day. When one of them grows a
// date parameter, flip its flag here.
dateAware bool
}
// querySources is the ordered chain actionQuery walks; first source to claim
@@ -49,68 +64,90 @@ type querySource struct {
// gate was never the bug. Adding a source (Kiwix, RSS, crawler, email) is one
// line here plus its method; where you put the line is the whole decision.
var querySources = []querySource{
{"fact-by-key", (*reactiveHandler).queryFactByKey},
{name: "fact-by-key", answer: (*reactiveHandler).queryFactByKey},
// Before "calendar" on purpose: both match "…на сегодня", and the plan is
// the more specific ask (its matcher requires a plan word), so the calendar
// listing would otherwise swallow it.
{"day-plan", (*reactiveHandler).queryDayPlan},
{name: "day-plan", answer: (*reactiveHandler).queryDayPlan},
// Also before "calendar": "что я обычно делаю по средам?" names a weekday,
// and the habit question is the more specific one. Its matcher requires a
// habit marker ("обычно", "каждый", …), so a question about this coming
// Wednesday still reaches the calendar.
{"habits", (*reactiveHandler).queryHabits},
{name: "habits", answer: (*reactiveHandler).queryHabits},
// Before "calendar" and before the recall sources: "что мне нужно
// сделать?" is a question about the task list, and the notes pass would
// otherwise answer it with whatever note happens to be nearest. Its
// matcher requires a task noun or an explicit "что … сделать", so a
// date-bearing question still reaches the calendar.
{"tasks", (*reactiveHandler).queryTasks},
{name: "tasks", answer: (*reactiveHandler).queryTasks},
// Before the recall sources too: "сколько я потратил?" is a question about
// the money facts the poller wrote, and the notes pass would otherwise
// answer it from whatever he once said about spending. Its matcher needs a
// money noun plus an actual ask, so "я потратил весь день" is untouched.
{"money", (*reactiveHandler).queryMoney},
{name: "money", answer: (*reactiveHandler).queryMoney},
// Before the recall sources and before general knowledge: "что нового?" is
// a question about the feeds she reads, and general knowledge would answer
// it by inventing news. Its matcher needs a feed noun plus an ask, so
// "у меня новая лента в инстаграме" is untouched.
{"feeds", (*reactiveHandler).queryFeeds},
{name: "feeds", answer: (*reactiveHandler).queryFeeds},
// Before "calendar" and before the recall sources: "что включено дома?" is
// a question about the house, and the notes pass would otherwise answer it
// from whatever he once said about the lights. Its matcher needs a house
// marker plus an ask plus a device word, and it bails out on weather
// wording, so "какая температура на улице?" still reaches the weather
// source.
{"home", (*reactiveHandler).queryHome},
{name: "home", answer: (*reactiveHandler).queryHome},
// Next to "home" and for the same reason: "какие устройства в сети?" is a
// question about the LAN, and the recall pass would otherwise answer it
// from an old note about the router. Its matcher needs a network word plus
// an ask plus a device noun, so "интернет не работает" is untouched.
{"network", (*reactiveHandler).queryNetwork},
{"calendar", (*reactiveHandler).queryCalendar},
{"weather", (*reactiveHandler).queryWeather},
{"embed", (*reactiveHandler).queryEmbed},
{"memory", (*reactiveHandler).queryMemory},
{"notes", (*reactiveHandler).queryNotes},
{name: "network", answer: (*reactiveHandler).queryNetwork},
{name: "calendar", answer: (*reactiveHandler).queryCalendar, dateAware: true},
{name: "weather", answer: (*reactiveHandler).queryWeather},
{name: "embed", answer: (*reactiveHandler).queryEmbed},
{name: "memory", answer: (*reactiveHandler).queryMemory},
{name: "notes", answer: (*reactiveHandler).queryNotes},
// THE BOUNDARY. Everything above answers from his own data; everything
// below answers from the world's. A question about him that got this far
// has no answer in his data, and no outside source can supply one, so this
// stops the walk rather than let the encyclopedia and the model guess.
{name: "personal", answer: (*reactiveHandler).queryPersonal},
// The world, read live. Owner's ruling of 2026-08-02: a metasearch hit beats
// a frozen ZIM, so SearXNG asks before Kiwix does. Nothing of his is at
// stake by this point — the boundary above already stopped every question
// about him, and only the query string leaves the box.
{name: "search", answer: (*reactiveHandler).querySearch},
// The offline encyclopedia, now the fallback for when the line is down or
// the search comes back empty. It reads the way it always did; what changed
// is that it no longer gets first refusal on a world question.
{name: "kiwix", answer: (*reactiveHandler).queryKiwix},
// LAST before the model answers from memory, and that position is the whole
// design (Vikunja #259): local sources first. His memory, his notes and —
// once internal/kiwix is wired into this chain — the offline ZIMs all get
// their turn before anything touches the network. The model does NOT: it
// design (Vikunja #259): everything of his, then the search, then the ZIMs,
// and only then a page he named. The model does NOT come first: it
// answers after this, because a URL he said out loud is an instruction and
// a 1.7B guessing at a page it cannot read is how contents get invented.
// This source only claims a turn where he named a URL, so it never competes
// with a local answer.
{"web", (*reactiveHandler).queryWeb},
{"general-knowledge", (*reactiveHandler).queryGeneral},
{name: "web", answer: (*reactiveHandler).queryWeb},
{name: "general-knowledge", answer: (*reactiveHandler).queryGeneral},
}
func (h *reactiveHandler) actionQuery(ctx context.Context, dec router.Decision) string {
t := &queryTurn{dec: dec}
for _, src := range querySources {
if dec.Continued && !src.dateAware {
continue
}
if reply, ok := src.answer(h, ctx, t); ok {
return reply
}
}
if dec.Continued {
// The previous question cannot be re-asked for another day. Saying so
// beats "не знаю", which reads as "no data for tomorrow" when the
// truth is that she never looked.
return "про другой день так не отвечу — спроси целиком."
}
return "не знаю."
}
@@ -498,6 +535,226 @@ func (h *reactiveHandler) queryWeb(ctx context.Context, t *queryTurn) (string, b
return reply, true
}
// kiwixTimeout — the whole ZIM source, rewrite included. The rewrite is one
// short constrained completion and the search is a LAN request; if the pair
// takes longer than this something is wrong and he is better served by the
// model's own answer than by more waiting.
const kiwixTimeout = 20 * time.Second
// searchTimeout — the whole metasearch source. websearch.Client already holds a
// per-request timeout from config; this is the outer bound on the turn, so a
// hung dial cannot outlive it either. Shorter than kiwixTimeout because there
// is no rewrite call in front of it: the question goes out verbatim.
const searchTimeout = 12 * time.Second
// querySearch — the live web, through a self-hosted SearXNG.
//
// Ahead of Kiwix by the owner's ruling of 2026-08-02: a search reads what is
// true today, a ZIM reads what was true when it was built, and the ZIM is the
// fallback for a box with no line out. Everything of his still answers first —
// the personal boundary is directly above this source, so a question ABOUT him
// never becomes a query.
//
// What leaves this process is the query string and nothing else. His notes, his
// facts, the persona block and the history do not travel with it: the websearch
// package cannot read the store. That is the CLAUDE.md rule made mechanical,
// not a promise about how the prompt is assembled.
//
// It claims the turn only when the search returns something. An empty result,
// an unreachable instance and a 403 from an instance without the JSON format
// all fall through to Kiwix, which is the point of the ordering.
func (h *reactiveHandler) querySearch(ctx context.Context, t *queryTurn) (string, bool) {
if h.search == nil {
// Off unless configured, same as the crawler and the ZIMs. Nothing is
// said about it: he never asked for a capability he did not enable.
return "", false
}
ctxS, cancel := context.WithTimeout(ctx, searchTimeout)
defer cancel()
// Verbatim. No rewriter: SearXNG ranks by meaning through real engines, and
// reducing "почему небо голубое" to English keywords would throw away the
// language he asked in along with the ranking that handles it.
resp, err := h.search.client.Search(ctxS, t.dec.Utterance, h.search.max)
if err != nil {
log.Printf("voice: search %q: %v", t.dec.Utterance, err)
return "", false
}
if resp.Empty() {
return "", false
}
// Logged on the way through, not only on failure. Without this there is no
// telling from the outside whether an answer came off the web, off a ZIM or
// out of the model's weights, and those are the cases worth telling apart.
log.Printf("voice: search: %q → %d answers, %d results", t.dec.Utterance, len(resp.Answers), len(resp.Results))
// Handed over the same way a note, a page or an article is: evidence for the
// question he asked, not something to recite. The trim is one budget over the
// joined block, so a long first snippet cannot crowd out the rest.
evidence := crawl.TrimRunes(strings.Join(resp.Snippets(), "\n"), h.search.runes)
var reply string
if h.phraser != nil {
var perr error
reply, perr = h.phraser.PhraseQuery(ctx, t.dec.Utterance, []string{evidence})
if perr != nil {
log.Printf("voice: search: phrase: %v", perr)
}
}
if reply == "" {
// No phraser, or it failed. Read back the best evidence rather than
// pretend the search did not happen.
return "вот что я нашла: " + crawl.TrimRunes(resp.Snippets()[0], 300), true
}
return reply, true
}
// queryKiwix — the offline encyclopedia, and the fallback behind querySearch:
// everything of his has already had its turn and the live search found nothing
// or could not be reached. Reading beats recalling for a 1.7B either way.
//
// What leaves this process is the search query and nothing else. His notes,
// his facts, the persona block and the history do not travel with it — the
// kiwix package cannot read the store. That holds even though the server is on
// the LAN, because "local sources first" is not a licence to widen what a
// lookup is allowed to see.
//
// It claims the turn only when the search returns something. No results is not
// a failure worth announcing: it means the ZIM does not cover this, and the
// model answering next is the better outcome than "ничего не нашла".
func (h *reactiveHandler) queryKiwix(ctx context.Context, t *queryTurn) (string, bool) {
if h.kiwix == nil {
// Off unless configured, same as the crawler and the weather. Nothing
// is said about it: he never asked for a capability he did not enable.
return "", false
}
ctxK, cancel := context.WithTimeout(ctx, kiwixTimeout)
defer cancel()
// The ZIMs are English and kiwix ranks by keyword overlap, not meaning, so
// a Russian sentence matches nothing at all. The rewriter turns it into a
// handful of English keywords with the resident model.
pattern := t.dec.Utterance
if h.kiwix.rewriter != nil {
q, err := h.kiwix.rewriter.Rewrite(ctxK, t.dec.Utterance)
if err != nil {
// Fall through to the verbatim question rather than give up. It
// will usually miss, and missing is a fall-through too.
log.Printf("voice: kiwix: rewrite: %v", err)
} else if q != "" {
pattern = q
}
}
hits, err := h.kiwix.client.Search(ctxK, pattern, h.kiwix.book, h.kiwix.max)
if err != nil {
log.Printf("voice: kiwix: search %q: %v", pattern, err)
return "", false
}
if len(hits) == 0 {
return "", false
}
top := hits[0]
// Logged on the way through, not only on failure. Without this there is no
// way to tell from the outside whether an answer came off a ZIM or out of
// the model's weights, and those are the two cases worth telling apart.
log.Printf("voice: kiwix: %q → %d hits, top %q", pattern, len(hits), top.Title)
// The top hit only, read as an article rather than as a snippet. Kiwix
// builds its snippet from wherever the keyword matched, which on Wikipedia
// is usually the navigation box at the foot of the page — the first version
// of this joined three of those and she recited "Ecological economics
// Ecological footprint …" at him. The head of the article is the lead
// paragraph, which is the definition the snippet was meant to be.
page, aerr := h.kiwix.client.Article(ctxK, top.Path, h.kiwix.runes)
if aerr != nil || page.Text == "" {
if aerr != nil {
log.Printf("voice: kiwix: article %s: %v", top.Path, aerr)
}
// The search did find something, so fall back to its snippet rather
// than throw the hit away.
if top.Snippet == "" {
return "", false
}
page = crawl.Page{Title: top.Title, Text: top.Snippet}
}
// Handed over the same way a note or a page is: context for the question he
// asked, not something to recite.
snippet := top.Title + "\n" + crawl.TrimRunes(page.Text, h.kiwix.runes)
var reply string
if h.phraser != nil {
var perr error
reply, perr = h.phraser.PhraseQuery(ctx, t.dec.Utterance, []string{snippet})
if perr != nil {
log.Printf("voice: kiwix: phrase: %v", perr)
}
}
if reply == "" {
// No phraser, or it failed. Read back the best hit rather than pretend
// the search did not happen.
return "вот что я нашла: " + crawl.TrimRunes(top.Title+" — "+page.Text, 300), true
}
return reply, true
}
// queryPersonal — stop the walk on a question about him that his own data did
// not answer.
//
// Every source above this one reads something of his: his facts, his calendar,
// his tasks, his house, his notes. Everything below reads the world: an offline
// Wikipedia, a page he named, the model's own weights. The world does not know
// when his meeting is, and asked anyway it will produce something.
//
// It did. "во сколько у меня встреча" reached Kiwix on the deployed daemon,
// 01-08-2026; Wikipedia matched an article on the 2015 CPISRA World Games, and
// the phraser rendered it as "встреча у тебя в 2015 CPISRA World Games, где
// были соревнования по плаванию". Fluent, confident, and about a swimming
// competition in Nottingham. Saying "не знаю" is not a worse answer than that
// one — it is the only true one.
//
// Note this is also the privacy edge. The rule in CLAUDE.md is that only the
// utterance may leave the box, never his notes; a question that is ABOUT him
// carries his life in the utterance itself, so it is the one class that should
// not be sent to an upstream engine at all. The guard closes both holes with
// the same test.
func (h *reactiveHandler) queryPersonal(ctx context.Context, t *queryTurn) (string, bool) {
if !isPersonalQuery(t.dec.Utterance) {
return "", false
}
log.Printf("voice: %q is about him and his own data did not answer it; not asking the world", t.dec.Utterance)
return "не знаю — не нашла у тебя такой записи.", true
}
// personalMarkers — first-person POSSESSION, not first person generally.
//
// "у меня" and "мой" attach to a thing that is his, which is what makes the
// question unanswerable from outside. A bare "мне" or "я" does not: "как мне
// сварить борщ" and "что я могу посмотреть" are ordinary questions about the
// world that happen to mention the asker, and refusing those would be the
// opposite mistake. The narrow test is the point.
// Go's \b is ASCII-only and never fires next to a Cyrillic letter, so the
// Russian patterns spell the boundary out as "not a letter or a digit". The
// English ones keep \b, where it works.
var personalMarkers = []*regexp.Regexp{
regexp.MustCompile(`(?i)(^|[^\p{L}\p{N}])у\s+меня([^\p{L}\p{N}]|$)`),
regexp.MustCompile(`(?i)(^|[^\p{L}\p{N}])мо(й|я|ё|е|и|его|ей|их|им|ими|ем|ю|ею)([^\p{L}\p{N}]|$)`),
regexp.MustCompile(`(?i)\bmy\b`),
regexp.MustCompile(`(?i)\bdo\s+i\s+have\b`),
regexp.MustCompile(`(?i)\bdid\s+i\b`),
}
// isPersonalQuery reports whether the utterance asks about something of his.
func isPersonalQuery(utterance string) bool {
if utterance == "" {
return false
}
for _, re := range personalMarkers {
if re.MatchString(utterance) {
return true
}
}
return false
}
// queryGeneral — general knowledge from the phraser, the last source before
// giving up. It always claims: either the model answers or Maven says she
// doesn't know.
+116
View File
@@ -0,0 +1,116 @@
package main
import (
"context"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
)
// contQueryAPI records which core call a continued query reached. DayPlan and
// LatestFact are here to be caught, not to be used: a continuation must never
// reach them, and the counters are how the test says so.
type contQueryAPI struct {
ipc.UnimplementedCoreAPI
from, to time.Time
events int
plans int
factLooks int
}
func (a *contQueryAPI) CalendarEvents(_ context.Context, from, to time.Time) ([]ipc.Fact, error) {
a.events++
a.from, a.to = from, to
return []ipc.Fact{{Key: "calendar", Value: "Планёрка @ 14:00", Confidence: 1.0, Ts: from.Add(14 * time.Hour)}}, nil
}
func (a *contQueryAPI) DayPlan(context.Context) (ipc.DayPlan, error) {
a.plans++
return ipc.DayPlan{Spoken: "план на сегодня"}, nil
}
func (a *contQueryAPI) LatestFact(_ context.Context, key string) (ipc.Fact, error) {
a.factLooks++
return ipc.Fact{Key: key, Value: "2л", Ts: contNow.Add(-time.Hour)}, nil
}
func contQueryHandler() (*reactiveHandler, *contQueryAPI) {
api := &contQueryAPI{}
return &reactiveHandler{api: api, now: func() time.Time { return contNow }}, api
}
// A continuation is a question about another day, so the one source that can
// read a day answers it — for the day the ellipsis named, not for today.
func TestContinuedQueryReachesTheCalendar(t *testing.T) {
h, api := contQueryHandler()
reply := h.actionQuery(context.Background(), router.Decision{
Intent: router.IntentQuery,
Utterance: "а завтра?",
Continued: true,
Slots: router.Slots{Text: "что у меня сегодня", Time: contNow.Add(24 * time.Hour), HasTime: true},
})
if api.events != 1 {
t.Fatalf("CalendarEvents called %d times, want 1", api.events)
}
if got, want := api.from.Format("2006-01-02"), "2026-08-02"; got != want {
t.Errorf("asked the calendar for %s, want %s", got, want)
}
if reply == "" {
t.Error("empty reply")
}
}
// The regression this gate exists for: every other source is date-blind, so
// letting one claim a continuation answers a question about tomorrow with
// today's data. queryFactByKey was the live case — HasKey plus HasTime, both
// set by the continuation, and it replies with a stored fact's own timestamp.
func TestContinuedQuerySkipsDateBlindSources(t *testing.T) {
h, api := contQueryHandler()
h.actionQuery(context.Background(), router.Decision{
Intent: router.IntentQuery,
Utterance: "а вчера?",
Continued: true,
Slots: router.Slots{
Key: "water", HasKey: true,
Text: "когда я пил воду",
Time: contNow.Add(-24 * time.Hour), HasTime: true,
},
})
if api.factLooks != 0 {
t.Errorf("fact-by-key claimed a continuation (%d lookups)", api.factLooks)
}
if api.plans != 0 {
t.Errorf("day-plan claimed a continuation (%d calls)", api.plans)
}
}
// Nothing date-aware claimed it: say that, rather than "не знаю", which reads
// as "no data for that day" when she never looked.
func TestContinuedQueryWithNoDateAwareAnswerSaysSo(t *testing.T) {
h, _ := contQueryHandler()
// No parseable day in the utterance, so even the calendar passes.
reply := h.actionQuery(context.Background(), router.Decision{
Intent: router.IntentQuery,
Utterance: "а?",
Continued: true,
Slots: router.Slots{Text: "какая погода", HasTime: true},
})
if reply == "не знаю." || !strings.Contains(reply, "спроси целиком") {
t.Fatalf("reply = %q, want the honest continuation refusal", reply)
}
}
// An ordinary query is untouched by the gate — every source still runs.
func TestOrdinaryQueryStillReachesEverySource(t *testing.T) {
h, api := contQueryHandler()
h.actionQuery(context.Background(), router.Decision{
Intent: router.IntentQuery,
Utterance: "какие планы на сегодня?",
})
if api.plans != 1 {
t.Fatalf("day-plan called %d times on an ordinary query, want 1", api.plans)
}
}
+109
View File
@@ -0,0 +1,109 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
)
func TestIsPersonalQuery(t *testing.T) {
for _, s := range []string{
"во сколько у меня встреча",
"что у меня сегодня",
"когда мой следующий отпуск",
"где моя книга",
"сколько моих задач висит",
"when is my meeting",
"do i have anything today",
"did i take my vitamins",
} {
if !isPersonalQuery(s) {
t.Errorf("isPersonalQuery(%q) = false, want true", s)
}
}
for _, s := range []string{
// First person without possession. These are questions about the
// world that merely mention the asker, and refusing them would be the
// opposite mistake.
"как мне сварить борщ",
"что я могу посмотреть вечером",
"почему небо синее",
"столица франции",
"how do i boil an egg",
"",
} {
if isPersonalQuery(s) {
t.Errorf("isPersonalQuery(%q) = true, want false", s)
}
}
}
// kiwixTrapAPI stands in for the world. Nothing below the personal boundary
// should be consulted for a question about him, so the test asserts on the
// reply rather than on a call: reaching Kiwix or general knowledge produces a
// phrased answer, and refusing produces the honest one.
func personalHandler() *reactiveHandler {
return &reactiveHandler{
api: ipc.UnimplementedCoreAPI{},
now: func() time.Time { return contNow },
// No phraser and no kiwix wiring: if the walk gets past the personal
// source it reaches queryGeneral, which returns "не знаю." with a nil
// phraser — a different string from the one this guard produces, so
// the two cases stay distinguishable.
}
}
// The regression: "во сколько у меня встреча" reached Kiwix, Wikipedia matched
// an article on the 2015 CPISRA World Games, and the phraser reported it back
// as his meeting. Seen on the deployed daemon, 01-08-2026.
func TestPersonalQuestionIsNotSentToTheWorld(t *testing.T) {
h := personalHandler()
reply, ok := h.queryPersonal(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "во сколько у меня встреча"},
})
if !ok {
t.Fatal("queryPersonal passed on a question about him")
}
if reply == "" {
t.Fatal("empty reply")
}
}
func TestWorldQuestionsPassThroughTheBoundary(t *testing.T) {
h := personalHandler()
for _, u := range []string{"почему небо синее", "столица франции"} {
if _, ok := h.queryPersonal(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: u},
}); ok {
t.Errorf("queryPersonal claimed %q, want it to pass to the encyclopedia", u)
}
}
}
// The boundary must sit above kiwix and general-knowledge and below every
// source that reads his own data. Asserted on the table itself: an ordering
// bug here is silent, because both arrangements answer, just from the wrong
// place.
func TestPersonalBoundarySitsBetweenHisDataAndTheWorld(t *testing.T) {
idx := map[string]int{}
for i, s := range querySources {
idx[s.name] = i
}
boundary, ok := idx["personal"]
if !ok {
t.Fatal("no personal source in the chain")
}
for _, his := range []string{"fact-by-key", "day-plan", "tasks", "calendar", "memory", "notes"} {
if i, ok := idx[his]; !ok || i > boundary {
t.Errorf("%q reads his own data and must run before the personal boundary", his)
}
}
for _, world := range []string{"search", "kiwix", "web", "general-knowledge"} {
if i, ok := idx[world]; !ok || i < boundary {
t.Errorf("%q reads the world and must run after the personal boundary", world)
}
}
}
+105
View File
@@ -0,0 +1,105 @@
package main
import (
"context"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"testing"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/websearch"
)
func searchHandler(t *testing.T, body string, status int) (*reactiveHandler, *string) {
t.Helper()
var seen string
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
seen = r.URL.RawQuery
if status != http.StatusOK {
http.Error(w, "no", status)
return
}
w.Write([]byte(body))
}))
t.Cleanup(srv.Close)
return &reactiveHandler{
// No phraser: querySearch then reads back the best evidence, which is
// what makes the claim visible without a llama-server in the test.
search: &searchWiring{client: websearch.New(srv.URL, websearch.Options{}), max: 3, runes: 1500},
}, &seen
}
const searchBody = `{"answers":["Небо голубое из-за рэлеевского рассеяния."],
"results":[{"title":"Рэлеевское рассеяние","url":"https://ru.wikipedia.org/x","content":"Рассеяние света."}]}`
func TestQuerySearchClaimsAndReadsBack(t *testing.T) {
h, _ := searchHandler(t, searchBody, http.StatusOK)
reply, ok := h.querySearch(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "почему небо голубое"},
})
if !ok {
t.Fatal("querySearch passed on a search with hits")
}
if !strings.Contains(reply, "рэлеевского рассеяния") {
t.Fatalf("reply = %q", reply)
}
}
// No rewriter in front of this source: SearXNG ranks by meaning, and reducing
// the question to English keywords would throw away the language he asked in.
func TestQuerySearchSendsTheQuestionVerbatim(t *testing.T) {
h, seen := searchHandler(t, searchBody, http.StatusOK)
h.querySearch(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "почему небо голубое"},
})
if !strings.Contains(*seen, "q="+url.QueryEscape("почему небо голубое")) {
t.Fatalf("query string = %q", *seen)
}
}
// The whole reason the ordering is safe: an unreachable or empty instance
// passes the turn to Kiwix instead of claiming it with an apology.
func TestQuerySearchFallsThroughWhenItFails(t *testing.T) {
for _, tc := range []struct {
name string
body string
status int
}{
{"http error", "", http.StatusForbidden},
{"no hits", `{"answers":[],"results":[]}`, http.StatusOK},
} {
t.Run(tc.name, func(t *testing.T) {
h, _ := searchHandler(t, tc.body, tc.status)
if _, ok := h.querySearch(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "почему небо голубое"},
}); ok {
t.Fatal("querySearch claimed the turn; Kiwix never got its fallback")
}
})
}
}
// Off unless configured, and silent about it: he never asked for a capability
// he did not enable.
func TestQuerySearchOffWithoutConfig(t *testing.T) {
h := &reactiveHandler{}
if _, ok := h.querySearch(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "почему небо голубое"},
}); ok {
t.Fatal("querySearch claimed a turn with no search block")
}
}
// The owner's ruling of 2026-08-02: the live search asks first, the ZIM is the
// fallback for a box with no line out.
func TestSearchRunsBeforeKiwix(t *testing.T) {
idx := map[string]int{}
for i, s := range querySources {
idx[s.name] = i
}
if idx["search"] > idx["kiwix"] {
t.Fatalf("search at %d, kiwix at %d: the ZIM is the fallback, not the first read", idx["search"], idx["kiwix"])
}
}
+52 -1
View File
@@ -31,9 +31,11 @@ package main
import (
"context"
"encoding/json"
"errors"
"fmt"
"log"
"strings"
"sync"
"time"
@@ -54,16 +56,65 @@ import (
// asks Maven to stop recording gets the transcript back in seconds.
const captureSummaryTimeout = 20 * time.Minute
// summaryGrammar — GBNF pinning a summarisation call to one JSON object holding
// the summary and nothing else. Same reasoning as responseGrammar and memeval's
// evalGrammar: the resident model is a Thinking variant, and a summarisation
// prompt is exactly the shape that invites it to answer with its reasoning as
// plain text. Demanding JSON leaves the reasoning nowhere to go.
//
// The bound is 2000 characters, twice the phraser's, because a reduce step over
// a two-hour meeting is a paragraph and not a sentence. Newlines are escaped by
// the escape rule, so the bullet list the prompt asks for survives the wrapper.
const summaryGrammar = `
root ::= "{" ws "\"summary\"" ws ":" ws string ws "}"
string ::= "\"" ([^"\\] | "\\" ["\\/bfnrt]){0,2000} "\""
ws ::= [ \t\n]*
`
// llmCompleter adapts *llm.Client to capture.Completer. The pure package names
// the two strings it needs and stays free of the llm request struct; the client
// itself is the swap-aware one from llmClientFor, so a model swap re-points it.
//
// The JSON wrapper lives here, not in internal/capture: that package is
// text-in/text-out by design, and the map/reduce steps still see plain prose.
type llmCompleter struct {
c *llm.Client
maxTokens int
}
func (l llmCompleter) Complete(ctx context.Context, system, user string) (string, error) {
return l.c.Complete(ctx, llm.Req{System: system, User: user, MaxTokens: l.maxTokens})
out, err := l.c.Complete(ctx, llm.Req{
System: system,
User: user,
Grammar: summaryGrammar,
MaxTokens: l.maxTokens,
})
if err != nil {
return "", err
}
return unwrapSummary(out), nil
}
// unwrapSummary takes the summary out of the JSON object the grammar produced.
// Anything that does not parse is returned as-is: an operator running without a
// grammar, or a llama-server too old to honour one, gets the plain text it used
// to get rather than an empty meeting summary.
func unwrapSummary(raw string) string {
s := stripThink(strings.TrimSpace(raw))
start := strings.Index(s, "{")
end := strings.LastIndex(s, "}")
if start < 0 || end <= start {
return s
}
var parsed struct {
Summary string `json:"summary"`
}
if err := json.Unmarshal([]byte(s[start:end+1]), &parsed); err != nil {
return s
}
// An empty field is the model saying nothing, so hand back nothing. Returning
// the raw object here would write `{"summary":""}` into his notes.
return strings.TrimSpace(parsed.Summary)
}
// captureWiring — the recorder plus what it needs to write the result down.
+26
View File
@@ -116,3 +116,29 @@ func TestStopReturnsTranscriptAndNotesItWithoutASummary(t *testing.T) {
t.Fatalf("the meeting left no note behind: %+v", notes)
}
}
// The summary path is JSON-wrapped by summaryGrammar, and internal/capture must
// keep seeing plain prose. These cover the wrapper and every way it can be
// absent or broken, because a meeting summary is written once and not retried.
func TestUnwrapSummary(t *testing.T) {
cases := []struct {
name string
in string
want string
}{
{"grammar output", `{"summary": "решили купить насос"}`, "решили купить насос"},
{"multiline field", `{"summary": "- насос\n- бюджет"}`, "- насос\n- бюджет"},
{"empty marker survives", `{"summary": "пусто"}`, "пусто"},
{"empty field says nothing", `{"summary": ""}`, ""},
{"thinking prefix", "<think>hm</think>\n{\"summary\": \"итог\"}", "итог"},
{"no grammar, plain prose", "решили купить насос", "решили купить насос"},
{"broken json falls back", `{"summary": "обрыв`, `{"summary": "обрыв`},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
if got := unwrapSummary(c.in); got != c.want {
t.Errorf("unwrapSummary(%q) = %q, want %q", c.in, got, c.want)
}
})
}
}
+19 -1
View File
@@ -348,9 +348,27 @@ func (h *reactiveHandler) rememberTurn(prev *dialogue.Session, dec router.Decisi
if dec.Intent == router.IntentChat {
ttl = 15 * time.Minute // conversational turns should last longer
}
// A system or query turn often carries no Text slot at all — a stage-0
// grammar fills none. The next turn may be an ellipsis ("а завтра?"),
// which knows the day but not what was asked ABOUT, so keep the raw
// utterance where continuation.go can find it. Only these two intents:
// everywhere else Text is a payload and must stay what the router put in.
//
// Overwritten, not filled: rememberTurn runs AFTER followUpMerge, which
// has already inherited a Text from the previous same-intent turn, so a
// fill-if-empty rule keeps the OLD topic for ever. Seen on the deployed
// daemon 01-08-2026 — "во сколько у меня встреча" then "какие у меня
// планы" then "а завтра?" continued the meeting, two turns stale.
//
// A continuation is the exception and keeps what it inherited: its
// utterance is the ellipsis, and the topic it carries is the real one.
slots := toDialogueSlots(dec.Slots)
if !dec.Continued && (dec.Intent == router.IntentSystem || dec.Intent == router.IntentQuery) {
slots.Text = dec.Utterance
}
h.dialogueSessions.Put(voiceDialogueID, &dialogue.Session{
Intent: dialogue.Intent(dec.Intent),
Slots: toDialogueSlots(dec.Slots),
Slots: slots,
Timestamp: now,
TTL: ttl,
History: history,
+125
View File
@@ -0,0 +1,125 @@
// Elliptical follow-ups — "а завтра?" after "какие напоминания на сегодня".
//
// These carry no intent of their own. Two words, one of them a particle, and
// everything that makes the utterance meaningful lives in the turn before it.
// Sent to the router they get whatever the model guesses, which on a 1.7B is
// close to a coin flip, and the guess costs ~2.7s to obtain.
//
// followUpMerge (followup.go) cannot help: it inherits SLOTS once the intent is
// known, and here the intent is the missing part. So this runs before the
// router and answers from the previous turn directly, which is both correct by
// construction and free.
package main
import (
"time"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/router"
)
// continuationMaxTokens — an ellipsis is short by definition. Past four tokens
// the utterance carries enough of its own content to be routed on its merits,
// and inheriting an intent for it would be overreach.
const continuationMaxTokens = 4
// continuationParticles — the words that open a follow-up. A leading particle
// is one of the two ways in; the other is an utterance that is nothing but a
// date ("завтра?").
var continuationParticles = map[string]bool{
"а": true, "и": true, "ну": true,
"what": true, "and": true, "how": true,
}
// continuableIntents — which intents an ellipsis may inherit.
//
// query and system are questions: asking the same question about a different
// day is exactly what "а завтра?" means, and re-aiming the Time slot answers it
// completely.
//
// The rest are excluded on purpose. fact and note would write something he did
// not say — "поужинал" then "а вчера?" is a question about yesterday, not a
// claim about it. chat has no slot to re-aim. act is the dangerous one: an
// allowlisted fn inherited by a two-word utterance is a way to run a
// destructive command nobody typed, and no follow-up is worth that.
//
// reminder was in this list and came out after a live check on 01-08-2026. A
// reminder's payload is its Text, and the Text embeds the day word it was
// created with: continuing "напомни сегодня о событиях" with "а завтра?" fires
// tomorrow with the text still reading "сегодня". Re-aiming Time is not enough
// when the day is also written into the payload, and rewriting the payload
// needs the date's span in the string, which ParseCalendarDate does not report.
var continuableIntents = map[dialogue.Intent]bool{
dialogue.IntentQuery: true,
dialogue.IntentSystem: true,
}
// continuationDecision reads an utterance as "the previous question, but for
// this other day". Returns ok=false whenever anything is uncertain, which
// hands the turn back to the ordinary router path.
//
// The date is what makes this safe. An ellipsis with no parseable day is just
// a short utterance, and short utterances are the router's job.
func continuationDecision(prev *dialogue.Session, text string, now time.Time) (router.Decision, bool) {
if prev == nil || prev.IsExpired(now) || !continuableIntents[prev.Intent] {
return router.Decision{}, false
}
tokens := quietTokens(text)
if len(tokens) == 0 || len(tokens) > continuationMaxTokens {
return router.Decision{}, false
}
day, ok := router.ParseCalendarDate(text, now)
if !ok {
return router.Decision{}, false
}
// Either it opens with a particle, or the whole utterance is the date.
if !continuationParticles[tokens[0]] && !isBareDate(tokens, day, now) {
return router.Decision{}, false
}
dec := router.Decision{
Utterance: text,
Intent: router.Intent(prev.Intent),
Confidence: 1.0,
Stage: 0,
Continued: true,
Slots: router.Slots{
Key: prev.Slots.Key,
HasKey: prev.Slots.HasKey,
Value: prev.Slots.Value,
Text: prev.Slots.Text,
// Fn/Args are deliberately not carried: continuableIntents
// excludes act, so there is never one to carry.
Time: day,
HasTime: true,
},
}
return dec, true
}
// isBareDate reports whether the utterance is nothing but its date expression.
// "завтра" and "на выходных" qualify; "напомни завтра" does not, because the
// verb is content of its own and belongs to the router.
//
// Implemented by re-parsing each token: if every token that is not part of a
// date expression is a preposition or a question mark's leftovers, the
// utterance is bare. Cheap enough at four tokens.
func isBareDate(tokens []string, day time.Time, now time.Time) bool {
for _, t := range tokens {
if continuationFillers[t] {
continue
}
if d, ok := router.ParseCalendarDate(t, now); ok && d.Equal(day) {
continue
}
return false
}
return true
}
// continuationFillers — tokens that carry no content of their own inside a
// date expression ("на выходных", "в среду").
var continuationFillers = map[string]bool{
"на": true, "в": true, "во": true, "за": true, "про": true,
"about": true, "on": true, "for": true,
}
+169
View File
@@ -0,0 +1,169 @@
package main
import (
"testing"
"time"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/router"
)
var contNow = time.Date(2026, 8, 1, 12, 0, 0, 0, time.UTC)
func contSession(intent dialogue.Intent, key string) *dialogue.Session {
return &dialogue.Session{
Intent: intent,
Slots: dialogue.Slots{Key: key, HasKey: key != "", Text: "какие напоминания на сегодня"},
Timestamp: contNow.Add(-30 * time.Second),
TTL: 2 * time.Minute,
}
}
func TestContinuationInheritsTheQuestion(t *testing.T) {
prev := contSession(dialogue.IntentQuery, "water")
dec, ok := continuationDecision(prev, "а завтра?", contNow)
if !ok {
t.Fatal("continuationDecision returned false, want a decision")
}
if dec.Intent != router.IntentQuery {
t.Errorf("intent = %q, want query", dec.Intent)
}
if dec.Slots.Key != "water" || !dec.Slots.HasKey {
t.Errorf("key = %q, want water carried over", dec.Slots.Key)
}
if !dec.Slots.HasTime {
t.Fatal("no time slot; the whole point is re-aiming the day")
}
if got, want := dec.Slots.Time.Format("2006-01-02"), "2026-08-02"; got != want {
t.Errorf("time = %s, want %s", got, want)
}
}
func TestContinuationAcceptsABareDate(t *testing.T) {
prev := contSession(dialogue.IntentQuery, "water")
for _, s := range []string{"завтра?", "вчера", "а вчера?", "и завтра"} {
if _, ok := continuationDecision(prev, s, contNow); !ok {
t.Errorf("continuationDecision(%q) = false, want true", s)
}
}
}
func TestContinuationDeclinesWhatIsNotAnEllipsis(t *testing.T) {
prev := contSession(dialogue.IntentQuery, "water")
for _, s := range []string{
// No date to re-aim at — an ordinary short utterance, the router's job.
"а что там", "а бэкап?", "привет", "",
// Content of its own: the verb is not an ellipsis.
"напомни завтра позвонить маме",
// Too long to be an ellipsis even with a date in it.
"а что у меня стоит в календаре на завтра",
} {
if _, ok := continuationDecision(prev, s, contNow); ok {
t.Errorf("continuationDecision(%q) = true, want false", s)
}
}
}
func TestContinuationDeclinesUncontinuableIntents(t *testing.T) {
// act is the one that matters: inheriting an allowlisted fn from a
// two-word utterance would be a way to run a destructive command.
// reminder is here because its payload is its Text, and the Text embeds
// the day word it was created with — see continuableIntents.
for _, in := range []dialogue.Intent{
dialogue.IntentAct, dialogue.IntentFact, dialogue.IntentNote,
dialogue.IntentChat, dialogue.IntentReminder,
} {
if _, ok := continuationDecision(contSession(in, "water"), "а завтра?", contNow); ok {
t.Errorf("continuationDecision inherited intent %q, want refusal", in)
}
}
}
func TestContinuationDeclinesWithoutALiveSession(t *testing.T) {
if _, ok := continuationDecision(nil, "а завтра?", contNow); ok {
t.Error("continued with no previous turn")
}
stale := contSession(dialogue.IntentQuery, "water")
stale.Timestamp = contNow.Add(-10 * time.Minute)
if _, ok := continuationDecision(stale, "а завтра?", contNow); ok {
t.Error("continued an expired session")
}
}
func TestContinuationNeverCarriesAnFn(t *testing.T) {
prev := contSession(dialogue.IntentQuery, "water")
prev.Slots.Fn, prev.Slots.HasFn = "restart", true
dec, ok := continuationDecision(prev, "а завтра?", contNow)
if !ok {
t.Fatal("want a decision")
}
if dec.Slots.HasFn || dec.Slots.Fn != "" {
t.Fatalf("carried fn %q into a continuation", dec.Slots.Fn)
}
}
// TestContinuationCarriesTheTopic — the ellipsis names the day; what he is
// asking ABOUT has to come from the previous turn, or replySystem keyword-
// matches "а завтра?" and finds nothing. Caught on the deployed daemon.
func TestContinuationCarriesTheTopic(t *testing.T) {
prev := contSession(dialogue.IntentSystem, "")
prev.Slots.Text = "какой сегодня день"
dec, ok := continuationDecision(prev, "а завтра?", contNow)
if !ok {
t.Fatal("want a decision")
}
if dec.Slots.Text != "какой сегодня день" {
t.Fatalf("Slots.Text = %q, want the previous turn's topic", dec.Slots.Text)
}
}
// TestReplySystemIgnoresAnInheritedTopic — the regression the deployed daemon
// showed on 01-08-2026: followUpMerge fills an empty Text from the previous
// same-intent turn, so a plain "привет" after "какой сегодня день" arrived at
// replySystem carrying the old topic and was answered with the date. Only a
// continuation may widen the keyword match.
func TestReplySystemIgnoresAnInheritedTopic(t *testing.T) {
h := &reactiveHandler{now: func() time.Time { return contNow }}
inherited := router.Decision{
Utterance: "привет",
Intent: router.IntentSystem,
Slots: router.Slots{Text: "какой сегодня день"},
}
if got := h.replySystem(nil, inherited); got != "пока не умею отвечать на этот вопрос." {
t.Fatalf("replySystem answered %q on an inherited topic", got)
}
cont := inherited
cont.Utterance, cont.Continued = "а завтра?", true
if got := h.replySystem(nil, cont); got == "пока не умею отвечать на этот вопрос." {
t.Fatalf("replySystem refused a real continuation")
}
}
// TestRememberTurnRefreshesTheTopic — rememberTurn runs after followUpMerge,
// which has already inherited a Text from the previous same-intent turn. A
// fill-if-empty rule therefore pins the FIRST topic of a run of query turns
// and never lets go, so a later "а завтра?" continues a question two turns
// old. Seen on the deployed daemon, 01-08-2026.
func TestRememberTurnRefreshesTheTopic(t *testing.T) {
h := &reactiveHandler{
now: func() time.Time { return contNow },
dialogueSessions: dialogue.NewSessionStore(2 * time.Minute),
}
h.rememberTurn(nil, router.Decision{
Intent: router.IntentQuery, Utterance: "во сколько у меня встреча",
}, contNow)
// The second turn arrives with the first turn's Text already merged in.
prev := h.dialogueSessions.Get(voiceDialogueID, contNow)
h.rememberTurn(prev, router.Decision{
Intent: router.IntentQuery,
Utterance: "какие у меня планы",
Slots: router.Slots{Text: "во сколько у меня встреча"},
}, contNow)
got := h.dialogueSessions.Get(voiceDialogueID, contNow)
if got == nil {
t.Fatal("no session")
}
if got.Slots.Text != "какие у меня планы" {
t.Fatalf("topic = %q, want the latest turn's", got.Slots.Text)
}
}
+28
View File
@@ -351,3 +351,31 @@ func TestTickDayPlanReadsTheStore(t *testing.T) {
t.Errorf("a reminder for next year is not today's plan: %q", plan.Spoken)
}
}
// TestHandlerUpgradesToTheDaemonAPI — wireVoice runs before the tick loop
// exists, so the handler starts with the bare store adapter, and that adapter
// refuses DayPlan ("not available via direct store API"). main back-patches
// the real one in. Without the patch every "какие у меня планы на сегодня"
// answered "не получилось собрать план" on the deployed daemon, 01-08-2026.
func TestHandlerUpgradesToTheDaemonAPI(t *testing.T) {
h := &reactiveHandler{api: ipc.NewStoreAPI(nil), now: planDay}
if _, err := h.api.DayPlan(context.Background()); err == nil {
t.Fatal("the bare store adapter served a day plan; this test is measuring nothing")
}
want := samplePlan()
h.upgradeAPI(&daemonAPI{
CoreAPI: ipc.UnimplementedCoreAPI{},
getDayPlan: func(context.Context) ipc.DayPlan { return want },
})
reply, ok := h.queryDayPlan(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "какие у меня планы на сегодня?"},
})
if !ok {
t.Fatal("queryDayPlan passed on a plan question")
}
if reply != want.Spoken {
t.Fatalf("reply = %q, want the assembled plan", reply)
}
}
+54
View File
@@ -0,0 +1,54 @@
package main
import (
"log"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/kiwix"
"github.com/kami/maven/internal/llm"
)
// kiwixWiring — the offline encyclopedia, assembled. nil ⇒ off, which is the
// default: the query chain simply has no ZIM source.
//
// The rewriter is separately optional. Searching without one is legal and
// mostly useless against English ZIMs, but it is the honest degraded mode when
// there is no llama-server to rewrite with, and it is what `rewrite: false`
// asks for.
type kiwixWiring struct {
client *kiwix.Client
rewriter *kiwix.Rewriter // nil ⇒ the question is searched verbatim
book string
max int
runes int
}
// wireKiwix builds the ZIM reader from the `kiwix` block, or returns nil when
// there is none. config.Normalise has already dropped a block with no URL and
// filled the two size defaults, so this does no validation of its own.
//
// The llm client is the phraser's swap-aware one (llmClientFor), so a model
// swap re-points the rewriter with everything else. A nil client means there is
// no resident model at all; that degrades the rewriter, not the source.
func wireKiwix(cfg *config.Config, c *llm.Client) *kiwixWiring {
if cfg.Kiwix == nil {
return nil
}
kc := cfg.Kiwix
w := &kiwixWiring{
client: kiwix.New(kc.URL),
book: kc.Book,
max: kc.MaxResults,
runes: kc.SnippetRunes,
}
switch {
case !kc.RewriteEnabled():
log.Printf("voice: kiwix at %s (book %q, query rewriting off by config)", kc.URL, kc.Book)
case c == nil:
log.Printf("voice: kiwix at %s (book %q, no llama-server: searching questions verbatim)", kc.URL, kc.Book)
default:
w.rewriter = kiwix.NewRewriter(c)
log.Printf("voice: kiwix at %s (book %q)", kc.URL, kc.Book)
}
return w
}
+187
View File
@@ -0,0 +1,187 @@
package main
import (
"context"
"net/http"
"net/http/httptest"
"strings"
"testing"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/kiwix"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice"
)
// searchRSS is what kiwix-serve answers a /search with, trimmed to the fields
// ParseSearchRSS reads.
func searchRSS(items ...string) string {
return `<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel>` +
strings.Join(items, "") + `</channel></rss>`
}
func rssItem(title, snippet string) string {
return "<item><title>" + title + "</title><link>/x</link><description>" + snippet + "</description></item>"
}
// stubKiwixServer answers every search with the given body and records the
// pattern it was asked for, so a test can assert on what left the process.
type stubKiwixServer struct {
*httptest.Server
lastPattern string
lastBook string
}
func newStubKiwix(t *testing.T, body string, status int) *stubKiwixServer {
t.Helper()
s := &stubKiwixServer{}
s.Server = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if status != 0 && status != http.StatusOK {
w.WriteHeader(status)
return
}
// Two endpoints on one server: /search answers the RSS, everything else
// is an article read. Only the search is recorded — an article fetch
// carries no query string and would blank the assertions.
if r.URL.Path != "/search" {
w.Header().Set("Content-Type", "text/html")
_, _ = w.Write([]byte("<html><title>Article</title><body><p>the lead paragraph</p></body></html>"))
return
}
s.lastPattern = r.URL.Query().Get("pattern")
s.lastBook = r.URL.Query().Get("books.name")
w.Header().Set("Content-Type", "application/xml")
_, _ = w.Write([]byte(body))
}))
t.Cleanup(s.Close)
return s
}
// buildKiwixHandler wires the source with no rewriter: the question is searched
// verbatim, which keeps the assertion about what was sent unambiguous.
func buildKiwixHandler(base string) *reactiveHandler {
return &reactiveHandler{
replier: voice.NewStubReplier(),
phraser: phraser.NewStub(),
kiwix: &kiwixWiring{
client: kiwix.New(base),
book: "wikipedia_en_all_maxi",
max: config.DefaultKiwixResults,
runes: config.DefaultKiwixSnippetRunes,
},
}
}
func askKiwix(h *reactiveHandler, q string) (string, bool) {
return h.queryKiwix(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: q},
})
}
// The default daemon has no `kiwix` block, and a source that is off must not
// claim the turn — the model answers next, exactly as it did before.
func TestQueryKiwixOffPassesThrough(t *testing.T) {
h := &reactiveHandler{replier: voice.NewStubReplier(), phraser: phraser.NewStub()}
if reply, ok := askKiwix(h, "почему небо синее?"); ok {
t.Errorf("an unconfigured kiwix claimed the turn: %q", reply)
}
}
func TestQueryKiwixAnswersFromSnippets(t *testing.T) {
s := newStubKiwix(t, searchRSS(rssItem("Rayleigh scattering", "shorter wavelengths scatter more")), 0)
h := buildKiwixHandler(s.URL)
reply, ok := askKiwix(h, "почему небо синее?")
if !ok {
t.Fatal("kiwix found a hit and did not claim the turn")
}
if reply == "" {
t.Error("claimed the turn with an empty reply")
}
if s.lastBook != "wikipedia_en_all_maxi" {
t.Errorf("books.name = %q, want the configured book", s.lastBook)
}
}
// No hit is not a failure worth announcing: the ZIM does not cover it, and the
// model answering next beats "ничего не нашла".
func TestQueryKiwixNoHitsPassesThrough(t *testing.T) {
s := newStubKiwix(t, searchRSS(), 0)
if reply, ok := askKiwix(buildKiwixHandler(s.URL), "почему небо синее?"); ok {
t.Errorf("an empty result set claimed the turn: %q", reply)
}
}
// A dead or misconfigured server must degrade to the model, not to an error
// spoken out loud. A turn never breaks on a capability.
func TestQueryKiwixServerErrorPassesThrough(t *testing.T) {
s := newStubKiwix(t, "", http.StatusBadRequest)
if reply, ok := askKiwix(buildKiwixHandler(s.URL), "почему небо синее?"); ok {
t.Errorf("a 400 claimed the turn: %q", reply)
}
}
// The privacy rule in CLAUDE.md, asserted rather than assumed: only the
// utterance is searched. No note, no fact, no persona block travels with it.
func TestQueryKiwixSendsOnlyTheQuestion(t *testing.T) {
s := newStubKiwix(t, searchRSS(rssItem("X", "y")), 0)
h := buildKiwixHandler(s.URL)
// A turn carrying notes an earlier source already pulled. They must not
// reach the query string.
_, _ = h.queryKiwix(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "почему небо синее?"},
notes: []ipc.Note{{Text: "пароль от роутера hunter2"}},
})
if strings.Contains(s.lastPattern, "hunter2") {
t.Fatalf("a stored note leaked into the search query: %q", s.lastPattern)
}
if s.lastPattern != "почему небо синее?" {
t.Errorf("pattern = %q, want the utterance verbatim", s.lastPattern)
}
}
func TestWireKiwixOffWithoutABlock(t *testing.T) {
if w := wireKiwix(&config.Config{}, nil); w != nil {
t.Error("wireKiwix built a source with no config block")
}
}
// No llama-server means no rewriter, but the source still works: searching the
// question verbatim is the honest degraded mode, not a reason to stay dark.
func TestWireKiwixWithoutAnLLMHasNoRewriter(t *testing.T) {
w := wireKiwix(&config.Config{Kiwix: &config.KiwixConfig{
URL: "http://kiwix:8080", Book: "b", MaxResults: 5, SnippetRunes: 1500,
}}, nil)
if w == nil {
t.Fatal("wireKiwix returned nil for a configured block")
}
if w.rewriter != nil {
t.Error("built a rewriter with no llm client")
}
if w.book != "b" {
t.Errorf("book = %q", w.book)
}
}
// The whole point of reading the article: kiwix's own snippet is usually the
// navigation box at the foot of the page, so the lead paragraph must be what
// reaches the phraser.
func TestQueryKiwixReadsTheArticleNotTheSnippet(t *testing.T) {
junk := "Ecological economics Ecological footprint Ecological forecasting"
s := newStubKiwix(t, searchRSS(rssItem("Photosynthesis", junk)), 0)
h := buildKiwixHandler(s.URL)
h.phraser = nil // no phraser ⇒ the fallback reads back what it was given
reply, ok := askKiwix(h, "что такое фотосинтез?")
if !ok {
t.Fatal("did not claim the turn")
}
if !strings.Contains(reply, "the lead paragraph") {
t.Errorf("reply did not come from the article: %q", reply)
}
if strings.Contains(reply, "Ecological economics") {
t.Errorf("recited the navigation-box snippet: %q", reply)
}
}
+47 -3
View File
@@ -236,7 +236,7 @@ func run(args []string) error {
coreFor := func() ipc.CoreAPI { return newIntakeAPI(ipc.NewStoreAPI(st), evBus, time.Now) }
if !locked {
rules = loop.DefaultRules()
rules = wireRules(cfg)
gatherer = loop.NewGatherer(st, rules)
if cfg.QuietHours != nil {
gatherer.SetQuietHours(cfg.QuietHours.Start, cfg.QuietHours.End)
@@ -341,6 +341,9 @@ func run(args []string) error {
if voiceW != nil && voiceW.handler != nil {
api := coreAPI.(*daemonAPI)
api.chatFn = voiceW.handler.handleText
// And the reverse: the handler was wired with the bare store
// adapter, which cannot serve the day plan. See upgradeAPI.
voiceW.handler.upgradeAPI(api)
}
if voiceW != nil && voiceW.mcp != nil {
coreAPI.(*daemonAPI).getMCPServers = voiceW.mcp.status
@@ -508,7 +511,7 @@ func run(args []string) error {
}
// Wire everything.
rules = loop.DefaultRules()
rules = wireRules(cfg)
gatherer = loop.NewGatherer(st, rules)
if cfg.QuietHours != nil {
gatherer.SetQuietHours(cfg.QuietHours.Start, cfg.QuietHours.End)
@@ -603,6 +606,7 @@ func run(args []string) error {
}
if voiceW != nil && voiceW.handler != nil {
newAPI.chatFn = voiceW.handler.handleText
voiceW.handler.upgradeAPI(newAPI)
}
srv.SetAPI(newAPI)
srv.Check = (&auth.Gate{Enrollment: auth.NewFloorEnrollment(), Session: passkeySess}).Check
@@ -751,7 +755,15 @@ func run(args []string) error {
if voiceW != nil {
voiceW.close()
}
wg.Wait()
// Bounded. Every worker below watches ctx, but one parked in a model call
// or an HTTP fetch can outlast the supervisor's patience, and run() has to
// return for `defer st.Close()` to seal the database. A worker abandoned
// mid-tick loses one tick; a shutdown that never returns loses every write
// since the last clean stop — which is how the deployed ciphertext went
// eleven days stale in July 2026.
if !waitWorkers(&wg, workerGrace) {
log.Printf("mavend: workers still running after %s, sealing anyway", workerGrace)
}
log.Printf("mavend: bye")
return nil
}
@@ -782,3 +794,35 @@ func contextBlockFn(cfg *config.Config, now func() time.Time) func() string {
f := personaFacts(cfg)
return func() string { return f.Block(now()) }
}
// workerGrace — how long shutdown waits for the background workers before it
// goes ahead and seals without them. Comfortably inside docker's ten-second
// default so the seal still lands before SIGKILL.
const workerGrace = 4 * time.Second
// waitWorkers waits on wg for at most d. Reports whether they all finished.
func waitWorkers(wg *sync.WaitGroup, d time.Duration) bool {
done := make(chan struct{})
go func() {
wg.Wait()
close(done)
}()
select {
case <-done:
return true
case <-time.After(d):
return false
}
}
// wireRules builds the nudge rule set, minus anything config turned off. The
// drop is logged because a rule vanishing silently is indistinguishable from a
// rule that is broken, and the next person to wonder why she stopped nudging
// should find the answer in the boot log.
func wireRules(cfg *config.Config) []loop.Rule {
rules, dropped := loop.RulesExcept(cfg.DisabledRules)
for _, name := range dropped {
log.Printf("loop: rule %q disabled by config", name)
}
return rules
}
+5 -5
View File
@@ -29,11 +29,11 @@ import (
// minutes of evaluation was five minutes of a mute assistant.
//
// The background client now yields the slot while a turn is in flight, so the
// collision is handled where it belongs and this is a prompt budget again.
// Sixty seconds is long enough for a Thinking model here, and an evaluation cut
// off costs nothing, because it is retried at the next interval. Raise it if
// observations start truncating.
const memoryEvalTimeout = 60 * time.Second
// collision is solved where it belongs and this is a prompt budget again. Five
// minutes is safe once more, and it is back: 60s truncated a Thinking model
// mid-synthesis, which costs an observation for no latency saved. The gate, not
// this number, is what keeps a voice turn from waiting.
const memoryEvalTimeout = 5 * time.Minute
// memoryEvalWorker — ticker + evaluator.
type memoryEvalWorker struct {
+16 -3
View File
@@ -133,10 +133,21 @@ var (
quietOnPhrases = [][]string{
{"quiet", "on"}, {"quiet", "mode"},
{"тих", "режим"}, {"не", "шум"}, {"не", "беспоко"},
{"тих"},
// The noun form and the comparative. "режим тишины" is how the
// setting is named half the time, and "сделай потише" is how it is
// actually asked for out loud. Both used to fall through to the
// router, which has no quiet intent, so the command did nothing.
{"режим", "тишин"}, {"сделай", "тише"}, {"сделай", "потише"},
{"говори", "тише"}, {"будь", "потише"},
{"тих"}, {"потише"},
}
)
// quietWordStems — every stem that names the setting. Used by the
// negated-but-unmatched fallback in classifyQuietToggle, which has to
// recognise "хватит тишины" without an ON phrase having matched.
var quietWordStems = []string{"тих", "тишин", "потише"}
// quietNegatorWords — negators that are whole words with no useful stem.
var quietNegatorWords = map[string]bool{
"не": true, "нет": true, "хватит": true, "no": true, "not": true, "off": true,
@@ -197,8 +208,10 @@ func classifyQuietToggle(text string) (on, off bool) {
// on after he asked for it to stop.
if quietNegated(tokens, nil) {
for _, t := range tokens {
if quietStem(t, "тих") {
return false, true
for _, stem := range quietWordStems {
if quietStem(t, stem) {
return false, true
}
}
}
}
+17
View File
@@ -47,6 +47,15 @@ func TestResolveQuietToggle(t *testing.T) {
{"побудь в тихом режиме", quietOn},
{"Тихий Режим!", quietOn},
{"тихая", quietOn},
// The noun form and the comparative.
{"включи режим тишины", quietOn},
{"режим тишины", quietOn},
{"сделай потише", quietOn},
{"сделай тише", quietOn},
{"потише", quietOn},
// English, as the fixture phrases it.
{"turn quiet mode back on", quietOn},
{"enable quiet mode", quietOn},
// OFF vocabulary — all seven, incl. the three that used to say ON.
{"quiet off", quietOff},
@@ -60,6 +69,12 @@ func TestResolveQuietToggle(t *testing.T) {
{"выключи тихий режим", quietOff},
{"отмени тихий режим пожалуйста", quietOff},
{"верни громкий режим", quietOff},
{"выключи режим тишины", quietOff},
{"хватит тишины", quietOff},
{"turn off quiet mode", quietOff},
{"quiet mode off", quietOff},
{"stop quiet mode", quietOff},
{"disable quiet mode", quietOff},
// False positives: "тихо"/"тихий" as ordinary Russian.
{"очень тихий сегодня день", quietNone},
@@ -68,6 +83,8 @@ func TestResolveQuietToggle(t *testing.T) {
{"потихоньку", quietNone},
{"тихонько", quietNone},
{"он говорил тихим голосом весь вечер", quietNone},
{"в тишине лучше думается", quietNone},
{"на улице стало потише", quietNone},
// Unrelated.
{"напомни завтра позвонить маме", quietNone},
+2
View File
@@ -8,6 +8,7 @@ import (
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice"
)
@@ -47,6 +48,7 @@ func (r *llmReplier) Reply(d router.Decision) string {
out, err := r.c.Complete(ctx, llm.Req{
System: persona.Prepend(r.block, replySystem),
User: replyContext(d),
Grammar: phraser.ResponseGrammar,
MaxTokens: 512,
RepeatPenalty: 1.3,
})
+18
View File
@@ -5,6 +5,7 @@ import (
"testing"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice"
)
@@ -67,3 +68,20 @@ var errTestLLMDown = errTest("llm down")
type errTest string
func (e errTest) Error() string { return string(e) }
// grammarRecorder captures the request so the grammar can be asserted on.
type grammarRecorder struct{ req llm.Req }
func (g *grammarRecorder) Complete(_ context.Context, r llm.Req) (string, error) {
g.req = r
return `{"response":"записала","mood":"neutral"}`, nil
}
func TestLLMReplierCarriesTheResponseGrammar(t *testing.T) {
rec := &grammarRecorder{}
r := newLLMReplier(rec, nil)
r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
if rec.req.Grammar != phraser.ResponseGrammar {
t.Errorf("grammar = %q, want phraser.ResponseGrammar", rec.req.Grammar)
}
}
+41
View File
@@ -0,0 +1,41 @@
package main
import (
"log"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/websearch"
)
// searchWiring — the metasearch source, assembled. nil ⇒ off, which is the
// default: no `search` block, no query ever leaves the LAN.
//
// Thinner than kiwixWiring because there is nothing to rewrite. SearXNG ranks
// with real engines, so the question goes out as he asked it, and that is the
// reason this source sits ahead of the ZIMs rather than behind them.
type searchWiring struct {
client *websearch.Client
max int
runes int
}
// wireSearch builds the search client from the `search` block, or returns nil
// when there is none. config.Normalise has already dropped a block with no URL
// and filled the two size defaults, so this does no validation of its own.
func wireSearch(cfg *config.Config) *searchWiring {
if cfg.Search == nil {
return nil
}
sc := cfg.Search
log.Printf("voice: web search at %s (language %q, engines %q)", sc.URL, sc.Language, sc.Engines)
return &searchWiring{
client: websearch.New(sc.URL, websearch.Options{
Language: sc.Language,
Engines: sc.Engines,
Timeout: time.Duration(sc.Timeout),
}),
max: sc.MaxResults,
runes: sc.SnippetRunes,
}
}
+5 -2
View File
@@ -1,5 +1,5 @@
// mavend/simulator_test.go — the replayable full-system simulator
// (Vikunja #284, 20-07-2026-BACKLOG.md item 7).
// (Vikunja #284).
//
// # What it is
//
@@ -328,7 +328,10 @@ func (s *scriptedLLM) Complete(_ context.Context, r llm.Req) (string, error) {
s.mu.Lock()
defer s.mu.Unlock()
s.calls = append(s.calls, r)
routing := r.Grammar != ""
// A grammar no longer separates the two contracts — the replier carries one
// too since phraser.ResponseGrammar was attached to it. Only the router's
// grammar names the intent enum, so that is what tells them apart.
routing := strings.Contains(r.Grammar, "intent")
for _, e := range s.entries {
if e.Match != "" && !strings.Contains(strings.ToLower(r.User), strings.ToLower(e.Match)) {
continue
+107
View File
@@ -0,0 +1,107 @@
// Spoken snooze — "не сейчас", "потом", "отложи" said out loud after a nudge
// resolves it as `snoozed`, the same outcome the Telegram buttons and the web
// UI write. Until this existed, a nudge could only be deferred by touching a
// screen: the voice path had no way to reach store.ResolveNudge at all, so the
// one channel she nudges on hardest was the one channel he could not answer.
package main
import (
"context"
"log"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/store"
)
// snoozeWindow — how long after a send "потом" still means "that nudge".
//
// A window is what makes this safe to run before the router. "потом" is an
// ordinary Russian word; eating every one of them would break real sentences.
// Bounded to the minutes right after she spoke, the word is almost always an
// answer to what she just said, and outside the window the utterance falls
// through and routes normally.
//
// Twenty minutes rather than the two hours of store.SnoozeDuration: those
// measure different things. SnoozeDuration is how long the quiet lasts,
// snoozeWindow is how long an unanswered nudge stays the topic of the
// conversation.
const snoozeWindow = 20 * time.Minute
// snoozeScan — how many recent nudges to look at when finding the target. The
// newest pending one is nearly always the first row; a handful of resolved
// rows can sit in front of it when he acked a few in a row.
const snoozeScan = 10
// resolveSnooze — pre-route keyword check, run after the quiet toggle. Returns
// (reply, true) when the utterance defers a nudge she recently sent.
//
// It returns ("", false) in two different situations, on purpose: the words do
// not read as a deferral, or they do but there is nothing pending to defer. In
// both the turn keeps routing, so "потом посмотрю что там с бэкапом" is still
// a query when no nudge is outstanding.
func (h *reactiveHandler) resolveSnooze(ctx context.Context, text string, src turnSource) (string, bool) {
if !classifySnooze(text) {
return "", false
}
now := h.now()
target, ok := h.pendingNudge(ctx, now)
if !ok {
return "", false
}
if err := h.api.ResolveNudge(ctx, target.ID, store.NudgeSnoozed, now); err != nil {
log.Printf("voice: snooze nudge %d (%s, %s): %v", target.ID, target.Rule, src, err)
return "не получилось отложить.", true
}
log.Printf("voice: snoozed nudge %d (rule %s) from %s", target.ID, target.Rule, src)
return "хорошо, вернусь к этому позже.", true
}
// pendingNudge — the newest still-pending nudge sent inside snoozeWindow.
//
// Channel is deliberately not filtered. A nudge that went to Telegram is still
// the thing he is answering when he says "потом" at the microphone, and making
// the reply channel decide which nudges are answerable would mean the ops page
// he actually read could not be dismissed by voice.
func (h *reactiveHandler) pendingNudge(ctx context.Context, now time.Time) (ipc.Nudge, bool) {
recent, err := h.api.RecentNudges(ctx, snoozeScan)
if err != nil {
log.Printf("voice: recent nudges for snooze: %v", err)
return ipc.Nudge{}, false
}
for _, n := range recent {
if n.Outcome != store.NudgePending {
continue
}
if now.Sub(n.Ts) > snoozeWindow || n.Ts.After(now) {
continue
}
return n, true
}
return ipc.Nudge{}, false
}
// snoozePhrases — the deferral vocabulary, as stem sequences. Matched by
// quietPhrase (quiet_toggle.go), which carries the rule that matters here:
// a single-word pattern matches only a single-word utterance. Bare "потом" is
// an answer; "потом схожу за водой" is a plan, and reporting a plan must not
// silence the rule that prompted it.
var snoozePhrases = [][]string{
{"не", "сейчас"}, {"не", "могу", "сейчас"}, {"не", "до", "этого"},
{"напомн", "позже"}, {"напомн", "потом"}, {"спрос", "позже"},
{"отлож"}, {"позже"}, {"потом"}, {"попозже"}, {"погоди"},
{"not", "now"}, {"later"}, {"snooze"}, {"remind", "me", "later"},
}
// classifySnooze reads an utterance as a deferral. Unlike the quiet toggle
// there is no negation arm: "не потом" is not something anyone says, and the
// leading "не" of "не сейчас" is part of the phrase itself.
func classifySnooze(text string) bool {
tokens := quietTokens(text)
for _, p := range snoozePhrases {
if quietPhrase(tokens, p) {
return true
}
}
return false
}
+115
View File
@@ -0,0 +1,115 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/store"
)
// snoozeFakeAPI serves a fixed nudge list and records the resolution.
type snoozeFakeAPI struct {
ipc.UnimplementedCoreAPI
nudges []ipc.Nudge
gotID int64
gotOutcome string
calls int
}
func (a *snoozeFakeAPI) RecentNudges(_ context.Context, _ int) ([]ipc.Nudge, error) {
return a.nudges, nil
}
func (a *snoozeFakeAPI) ResolveNudge(_ context.Context, id int64, outcome string, _ time.Time) error {
a.gotID, a.gotOutcome, a.calls = id, outcome, a.calls+1
return nil
}
var snoozeNow = time.Date(2026, 8, 1, 12, 0, 0, 0, time.UTC)
func snoozeHandler(nudges []ipc.Nudge) (*reactiveHandler, *snoozeFakeAPI) {
api := &snoozeFakeAPI{nudges: nudges}
return &reactiveHandler{api: api, now: func() time.Time { return snoozeNow }}, api
}
func pendingNudgeAt(id int64, ago time.Duration) ipc.Nudge {
return ipc.Nudge{ID: id, Ts: snoozeNow.Add(-ago), Rule: "water", Channel: "voice", Outcome: store.NudgePending}
}
func TestClassifySnooze(t *testing.T) {
yes := []string{
"не сейчас", "потом", "позже", "попозже", "отложи", "погоди",
"напомни позже", "напомни потом", "не могу сейчас",
"not now", "later", "snooze",
}
for _, s := range yes {
if !classifySnooze(s) {
t.Errorf("classifySnooze(%q) = false, want true", s)
}
}
no := []string{
// A single-word pattern must not eat the sentence it appears in.
"потом схожу за водой", "позже посмотрю что там с бэкапом",
"напомни завтра позвонить маме", "какая погода", "погода на завтра",
"я отложил деньги", "", "тихий режим",
}
for _, s := range no {
if classifySnooze(s) {
t.Errorf("classifySnooze(%q) = true, want false", s)
}
}
}
func TestResolveSnoozeDefersTheNewestPendingNudge(t *testing.T) {
h, api := snoozeHandler([]ipc.Nudge{
{ID: 9, Ts: snoozeNow.Add(-time.Minute), Rule: "meal", Outcome: store.NudgeActed},
pendingNudgeAt(8, 3*time.Minute),
pendingNudgeAt(7, 10*time.Minute),
})
reply, handled := h.resolveSnooze(context.Background(), "не сейчас", sourceVoice)
if !handled || reply == "" {
t.Fatalf("got (%q, %v), want a reply", reply, handled)
}
if api.gotID != 8 || api.gotOutcome != store.NudgeSnoozed {
t.Fatalf("resolved (%d, %q), want (8, %q)", api.gotID, api.gotOutcome, store.NudgeSnoozed)
}
}
func TestResolveSnoozeFallsThroughWithNothingPending(t *testing.T) {
// The whole point of the window: with no live nudge, "потом" is just a
// word and must keep routing.
for _, name := range []string{"stale", "resolved", "empty"} {
var nudges []ipc.Nudge
switch name {
case "stale":
nudges = []ipc.Nudge{pendingNudgeAt(3, snoozeWindow+time.Minute)}
case "resolved":
nudges = []ipc.Nudge{{ID: 4, Ts: snoozeNow, Rule: "water", Outcome: store.NudgeActed}}
}
t.Run(name, func(t *testing.T) {
h, api := snoozeHandler(nudges)
reply, handled := h.resolveSnooze(context.Background(), "потом", sourceVoice)
if handled || reply != "" {
t.Fatalf("got (%q, %v), want fall-through", reply, handled)
}
if api.calls != 0 {
t.Fatalf("resolved a nudge with nothing pending")
}
})
}
}
func TestResolveSnoozeIgnoresAFutureNudge(t *testing.T) {
// Clock skew between the tick and the turn must not let a send from the
// future be answered before it happened.
h, api := snoozeHandler([]ipc.Nudge{pendingNudgeAt(5, -time.Minute)})
if _, handled := h.resolveSnooze(context.Background(), "потом", sourceVoice); handled {
t.Fatalf("snoozed a nudge dated in the future")
}
if api.calls != 0 {
t.Fatalf("resolved a future nudge")
}
}
+31 -1
View File
@@ -250,6 +250,7 @@ func (t *tickLoop) tick(ctx context.Context, now time.Time) {
log.Printf("tick: unacked telegram rules: %v", err)
return
}
keys = t.repeatableRules(keys)
if len(keys) == 0 {
return
}
@@ -261,6 +262,35 @@ func (t *tickLoop) tick(ctx context.Context, now time.Time) {
}
}
// repeatableRules drops keys whose rule is not wired any more.
//
// The repeat path reads the nudges table, not the rule set: any sev4 telegram
// row still at outcome=pending is re-sent every repeat_interval until it is
// acked. So turning a rule off in `disabled_rules` silenced new nudges and left
// the last un-acked one re-sending every five minutes, forever — a knob that
// stops the cause and not the symptom is worse than no knob. Found the evening
// of 2026-08-01, two messages after the rule was supposedly off.
//
// Filtering on the wired set rather than on the disabled list also covers the
// rule that was deleted from the code entirely: its orphan rows go quiet
// instead of nagging about a rule nobody can ack from the UI any more.
func (t *tickLoop) repeatableRules(keys []string) []string {
if len(keys) == 0 {
return nil
}
wired := make(map[string]bool, len(t.rules))
for _, r := range t.rules {
wired[r.Name] = true
}
out := keys[:0:0]
for _, k := range keys {
if wired[k] {
out = append(out, k)
}
}
return out
}
// cachePhrase keeps the latest phrased nudge per rule for the sev4-repeat
// path. writing under a mutex; the repeat path reads under the same. the
// cache is bounded by the rule count (≤ ~30 per spec) so eviction is not a
@@ -852,7 +882,7 @@ func (t *tickLoop) dayPlan(ctx context.Context, now time.Time) ipc.DayPlan {
}
reminders = append(reminders, morning.PlanEntry{
At: fire,
Text: strings.TrimSpace(r.Payload),
Text: r.Text(),
Kind: morning.PlanReminder,
})
}
+40
View File
@@ -667,3 +667,43 @@ func TestDigestDeduplicatesByRule(t *testing.T) {
t.Fatalf("after duplicate queue attempt: digestQ = %d, want 1 (dedup)", len(tl.digestQ))
}
}
// The repeat path reads the nudges table, not the rule set, so a rule turned
// off in `disabled_rules` used to keep re-sending its last un-acked telegram
// nudge every repeat_interval. Two arrived after the rule was off on
// 2026-08-01. A disabled rule must be unreachable on every path.
func TestRepeatableRulesDropsDisabledRules(t *testing.T) {
tl := &tickLoop{rules: mustRules(t, []string{"service_down"})}
got := tl.repeatableRules([]string{"service_down", "water"})
if len(got) != 1 || got[0] != "water" {
t.Fatalf("repeatableRules = %v, want [water]", got)
}
}
// An orphan row for a rule that no longer exists in the code goes quiet too:
// nothing can ack what the UI cannot show.
func TestRepeatableRulesDropsUnknownRules(t *testing.T) {
tl := &tickLoop{rules: loop.DefaultRules()}
if got := tl.repeatableRules([]string{"rule_deleted_last_year"}); len(got) != 0 {
t.Fatalf("repeatableRules = %v, want none", got)
}
}
func TestRepeatableRulesKeepsWiredRules(t *testing.T) {
tl := &tickLoop{rules: loop.DefaultRules()}
got := tl.repeatableRules([]string{"service_down", "water"})
if len(got) != 2 {
t.Fatalf("repeatableRules = %v, want both", got)
}
}
// mustRules returns DefaultRules minus the named ones, failing if a name
// matched nothing — a typo here would make the test pass for the wrong reason.
func mustRules(t *testing.T, disabled []string) []loop.Rule {
t.Helper()
rules, dropped := loop.RulesExcept(disabled)
if len(dropped) != len(disabled) {
t.Fatalf("dropped %v, want %v", dropped, disabled)
}
return rules
}
+105 -13
View File
@@ -76,17 +76,30 @@ type reactiveHandler struct {
tts tts.Synthesizer
router *router.Router
embedder router.Embedder // reused for note write/query (same model as the classifier)
api ipc.CoreAPI
tools *tool.Executor
matcher *tool.Matcher
phraser phraser.Phraser
replier voice.Replier
now func() time.Time
// api — the CoreAPI the handler reads and writes through. Wired with the
// bare store adapter and UPGRADED by main once the daemonAPI exists; see
// upgradeAPI.
api ipc.CoreAPI
tools *tool.Executor
matcher *tool.Matcher
phraser phraser.Phraser
replier voice.Replier
now func() time.Time
// crawler reads a web page he names out loud (queryWeb). nil ⇒ on-demand
// page reading is off, which is the default: no `crawl` block, no fetch.
crawler *crawl.Crawler
// search asks a self-hosted SearXNG (querySearch), the first world source
// once his own data has had its turn. nil ⇒ off, the default: no `search`
// block, no query ever leaves the LAN.
search *searchWiring
// kiwix searches the offline ZIMs (queryKiwix), the fallback behind the
// live search and the last source before the model answers from its own
// weights. nil ⇒ off, the default.
kiwix *kiwixWiring
// feedsOn — whether any RSS feed is configured (config.Feeds). It changes
// only what she SAYS when asked and nothing is there: "ленты не настроены"
// instead of "ничего нового", which are different truths.
@@ -179,6 +192,27 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
return h.reply(ctx, replyText, nil)
}
// upgradeAPI points the handler at the daemon's own CoreAPI once main has
// built it.
//
// Wiring order forces this. wireVoice runs before the tick loop exists, so it
// can only be handed the bare store adapter — and that adapter answers DayPlan
// (and TickTrace, and MorningStatus) with "not available via direct store
// API", because a day plan is assembled by the tick loop and is not a table to
// read. So queryDayPlan, which the query chain reaches for "какие у меня планы
// на сегодня", failed on the deployed daemon for every caller. main already
// back-patches the other direction (daemonAPI.chatFn = handler.handleText);
// this is the same seam in reverse.
//
// Safe against the obvious loop: nothing in the voice path calls api.Chat, so
// pointing the handler at an API whose Chat IS the handler cannot recurse.
func (h *reactiveHandler) upgradeAPI(api ipc.CoreAPI) {
if h == nil || api == nil {
return
}
h.api = api
}
// handleText — the core reactive path without stt/tts. Used by the IPC Chat
// endpoint (and eventually by telegram). Splits out the audio bookends from
// HandlePushToTalk so text channels share the same routing logic.
@@ -246,8 +280,42 @@ func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSour
return withNotice(expiredNotice, reply)
}
// 5. router — classify the utterance.
dec, err := h.router.Route(ctx, text, h.now())
// 4b. spoken snooze — "не сейчас" / "потом" answers the nudge she just
// sent. Only handled when a pending nudge is actually inside the window
// (snooze.go); otherwise the words route normally, because "потом" is an
// ordinary word and eating every one of them would break real sentences.
if reply, handled := h.resolveSnooze(ctx, text, src); handled {
return withNotice(expiredNotice, reply)
}
// 4c. spoken ack — "готово" closes that same nudge as `acted`. Only the
// contentless form is intercepted here; "выпил воды" keeps routing and
// closes the nudge after its fact lands (ackFromFact, step 8b).
if reply, handled := h.resolveAck(ctx, text, src); handled {
return withNotice(expiredNotice, reply)
}
// 5. route. An elliptical follow-up — "а завтра?" — is answered from the
// previous turn instead (continuation.go): the intent is the part it is
// missing, so no amount of routing recovers it, and the model's guess
// costs seconds to obtain and is close to a coin flip. Everything else
// goes to the router.
var (
dec router.Decision
err error
prev *dialogue.Session
)
now := h.now()
if h.dialogueSessions != nil {
prev = h.dialogueSessions.Get(voiceDialogueID, now)
}
cont := false
if dec, cont = continuationDecision(prev, text, now); cont {
log.Printf("voice: continuation of %s from the previous turn", dec.Intent)
}
if !cont {
dec, err = h.router.Route(ctx, text, now)
}
if err != nil {
// ErrNoIntents ⇒ classifier unseeded (cold boot). reply with a
// "still warming up" rather than a wire error.
@@ -263,10 +331,13 @@ func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSour
// turn (follow-ups like «напомни завтра» → «…позвонить маме»), then remember
// this turn for the next follow-up. Only same-intent, non-expired, non-
// clarify turns carry (see followUpMerge). Best-effort: nil store ⇒ skipped.
// A continuation already carries the previous turn's slots, so there is
// nothing left to inherit — but it is still remembered, so a chain of them
// ("а завтра?" … "а послезавтра?") keeps working.
if h.dialogueSessions != nil {
now := h.now()
prev := h.dialogueSessions.Get(voiceDialogueID, now)
dec = followUpMerge(prev, dec, now)
if !cont {
dec = followUpMerge(prev, dec, now)
}
if !dec.Clarify {
h.rememberTurn(prev, dec, now)
}
@@ -287,6 +358,10 @@ func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSour
replyText := h.applyAction(ctx, dec)
log.Printf("voice: applyAction returned: %q", replyText)
// 8b. a fact that answers a live nudge closes it as `acted` (ack.go).
// Silent: the fact reply stands, she does not congratulate him for it.
h.ackFromFact(ctx, dec)
// 9. replier — phrase the reply across the router decision.
if replyText == "" {
replyText = h.replier.Reply(dec)
@@ -324,6 +399,23 @@ func (h *reactiveHandler) replySystem(ctx context.Context, dec router.Decision)
u := strings.ToLower(dec.Utterance)
now := h.now()
// The topic and the day come from different places on a continuation.
// "а завтра?" names the day and nothing else; what he is asking ABOUT
// lives in the previous turn, which continuation.go copied into
// Slots.Text. Dates keep parsing from the utterance — that is the part
// the ellipsis actually restates — and only the keyword match widens.
//
// Gated on Continued, and that gate is load-bearing. followUpMerge fills
// an empty Text from the previous same-intent turn, so without it a plain
// "привет" after "какой сегодня день" inherited the old topic and got
// answered with the date. Seen on the deployed daemon, 01-08-2026.
topic := u
if dec.Continued {
if t := strings.ToLower(dec.Slots.Text); t != "" && t != u {
topic = u + " " + t
}
}
// stage-0 grammars catch the exact time/date patterns, but duration
// queries ("сколько времени прошло") bypass the grammar's build filter
// and can reach replySystem via the classifier path. Guard against them.
@@ -332,14 +424,14 @@ func (h *reactiveHandler) replySystem(ctx context.Context, dec router.Decision)
}
switch {
case strings.Contains(u, "час") || strings.Contains(u, "врем"):
case strings.Contains(topic, "час") || strings.Contains(topic, "врем"):
// "который час в киеве" — she keeps one clock, so any named place gets
// the honest answer. Never local time dressed up as the city's.
if mentionsUnknownPlace(u) {
return onlyLocalTimeReply
}
return "сейчас " + ruClock(now)
case strings.Contains(u, "день") || strings.Contains(u, "числ"):
case strings.Contains(topic, "день") || strings.Contains(topic, "числ"):
// "какое число завтра" — answer for the day the user asked about,
// not today. Reuses the router's calendar day-word parser.
day := now
+11 -1
View File
@@ -254,7 +254,13 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
netscan: w.netscan,
// nil unless `crawl.on_demand` is on: reading a page he names is a
// capability, and capabilities are off unless configured.
crawler: onDemandCrawler(cfg),
crawler: onDemandCrawler(cfg),
// nil unless a `search` block names a SearXNG instance. External search
// is off unless configured, and configuring it is the whole opt-in.
search: wireSearch(cfg),
// nil unless a `kiwix` block names a server. Same swap-aware client the
// router and replier use, so the rewriter follows a model swap.
kiwix: wireKiwix(cfg, llmClient),
weatherProvider: weatherProvider,
weatherLocation: weatherLocation,
memStore: memStore,
@@ -314,6 +320,10 @@ func buildRouter(emb router.Embedder, acts router.ActMatcher, threshold float64,
seedClassifier(cls)
grammars := router.DefaultGrammars(acts)
grammars = append(grammars, router.SystemTimeDateGrammars()...)
// After the time/date rules on purpose: "какой сегодня день" is a clock
// question and must keep reaching replySystem, while "что у меня сегодня"
// is an agenda question and must not.
grammars = append(grammars, router.AgendaQueryGrammars()...)
grammars = append(grammars, router.ReminderGrammar())
return router.New(router.Config{
Grammars: grammars,
+119
View File
@@ -0,0 +1,119 @@
// Command mavseal encrypts a live tmpfs working copy back to the ciphertext
// file, for the case mavend could not do it itself.
//
// mavend seals its database in `defer st.Close()` when run() returns. A daemon
// that is killed rather than shut down never gets there, and because the
// working copy lives in the container's /dev/shm it dies with the container:
// everything written since the last clean shutdown is lost, and the next boot
// silently rolls back to the stale ciphertext. That is not hypothetical — on
// 2026-08-01 the deployed ciphertext was eleven days old.
//
// This is a recovery tool, not part of the daemon. It is safe to run against a
// live database: it takes a consistent snapshot with VACUUM INTO rather than
// mutating the working copy the daemon owns.
//
// Usage:
//
// mavseal -plain /dev/shm/maven-plain.db -cipher /var/lib/maven/maven.db.enc
//
// The key is read from MAVEN_DB_KEY (base64, 32 bytes decoded), the same
// variable the daemon uses. It is never taken as an argument: an argument ends
// up in the shell history and in ps.
package main
import (
"context"
"database/sql"
"encoding/base64"
"flag"
"fmt"
"log"
"os"
"github.com/kami/maven/internal/store"
_ "modernc.org/sqlite"
)
func main() {
log.SetFlags(0)
if err := run(); err != nil {
log.Fatalf("mavseal: %v", err)
}
}
func run() error {
plain := flag.String("plain", "", "path to the plaintext working copy (required)")
cipher := flag.String("cipher", "", "path to write the ciphertext to (required)")
keep := flag.Bool("keep-snapshot", false, "leave the intermediate snapshot on disk for inspection")
flag.Parse()
if *plain == "" || *cipher == "" {
flag.Usage()
return fmt.Errorf("both -plain and -cipher are required")
}
key, err := readKey()
if err != nil {
return err
}
// VACUUM INTO rather than a WAL checkpoint on the file itself. The daemon
// is usually still running and still writing when this is needed, and
// checkpointing its working copy mutates a database it owns. VACUUM INTO
// reads a consistent snapshot into a new file and touches nothing else, so
// the worst case is a snapshot a few seconds stale instead of a torn one.
snap := *plain + ".mavseal-snapshot"
os.Remove(snap)
if err := snapshot(*plain, snap); err != nil {
return err
}
if !*keep {
defer os.Remove(snap)
}
before := fileSize(*cipher)
if err := store.SealPlaintext(snap, *cipher, key); err != nil {
return err
}
log.Printf("sealed %s → %s (%d bytes, was %d)", *plain, *cipher, fileSize(*cipher), before)
return nil
}
// readKey pulls the same base64 key the daemon reads. Fails closed: a short or
// unparseable key must not silently produce a file nothing can open.
func readKey() ([]byte, error) {
raw := os.Getenv("MAVEN_DB_KEY")
if raw == "" {
return nil, fmt.Errorf("MAVEN_DB_KEY is not set")
}
key, err := base64.StdEncoding.DecodeString(raw)
if err != nil {
return nil, fmt.Errorf("MAVEN_DB_KEY is not valid base64: %w", err)
}
if len(key) != 32 {
return nil, fmt.Errorf("MAVEN_DB_KEY decodes to %d bytes, want 32", len(key))
}
return key, nil
}
// snapshot writes a consistent copy of src to dst with VACUUM INTO. The copy
// includes everything committed to the write-ahead log, which is most of what
// is worth saving on a daemon that has been up for hours.
func snapshot(src, dst string) error {
db, err := sql.Open("sqlite", src)
if err != nil {
return fmt.Errorf("open working copy: %w", err)
}
defer db.Close()
if _, err := db.ExecContext(context.Background(), "VACUUM INTO ?", dst); err != nil {
return fmt.Errorf("snapshot: %w", err)
}
return nil
}
func fileSize(path string) int64 {
fi, err := os.Stat(path)
if err != nil {
return 0
}
return fi.Size()
}
+1 -1
View File
@@ -512,7 +512,7 @@ func main() {
// /tools — the authed enable surface. maven proposes acts she can't run;
// this page is where a human reviews and enables them (proposed→enabled).
// Enabling is the boundary-moving act (DESIGN.md § Tool registration —
// Enabling is the boundary-moving act (docs/design.md § Tool registration —
// drafting is suggest, enabling is act), so it lives ONLY here,
// behind wg+nginx+auth — never the voice/chat path.
mux.HandleFunc("/tools", func(w http.ResponseWriter, r *http.Request) {
+34 -3
View File
@@ -73,6 +73,37 @@ What it does and does not do:
- feed notes are **not** part of recall. "что я говорил про X" searches what he
said; headlines are read back only by asking about the feeds.
### Searching the web (`search`, on in this deploy)
`deploy/mavend.json` ships a `search` block, so a question that is not about him
reaches a self-hosted SearXNG before it reaches the ZIMs. Delete the block and
no query leaves the LAN again. The shipped shape:
```json
"search": {
"url": "http://searxng:9563",
"max_results": 4,
"snippet_runes": 1500,
"language": "auto",
"timeout": "8s"
}
```
- the instance needs `json` in its `search.formats` (settings.yml). A stock
SearXNG answers 403 to `format=json`, and then every search fails;
- the question goes out **verbatim**, in the language he asked it. There is no
rewriter here, unlike Kiwix: SearXNG ranks through real engines;
- only the query string leaves the box. `internal/websearch` cannot read the
store, so no note, fact, persona block or history can travel with a search;
- a question about him never becomes a query. The personal boundary in the
query chain stops the walk above this source;
- this runs **before** Kiwix. A live search reads what is true today and the
ZIMs read what was true when they were built, so the ZIMs are the fallback:
an empty result, an unreachable instance or a dead line falls through to
them and she never says the search failed;
- `engines` narrows the search to named engines, e.g. `"duckduckgo,wikipedia"`.
Empty means whatever the instance has enabled.
### Reading a page (`crawl`, also off by default)
There is no `crawl` block either, so no page is fetched. Two halves, separately
@@ -95,9 +126,9 @@ switched:
a fallback and not a habit;
- `watches` re-reads a fixed list on its interval and writes a note when the
text changed. Like the feeds, it announces nothing;
- the answer path sits behind his memory and his notes, and ahead of the model
answering from what it remembers. Kiwix is not wired into the chain yet. A
local read costs nothing, so anything local goes first;
- the answer path sits behind his memory, his notes, the web search and the
ZIMs, and ahead of the model answering from what it remembers. A page he
named is an instruction, so it is read last and only when he named one;
- `robots.txt` is fetched first and obeyed with no override; a `Disallow` is a
refusal she says out loud. `Crawl-delay` is waited out before the page is
fetched, and a delay longer than the turn fails the read instead of hanging
+58 -1
View File
@@ -5,6 +5,16 @@
"socket_path": "/run/maven/mavend.sock",
"state_dir": "/var/lib/maven",
"//disabled_rules": [
"Nudge rules that are not wired at all. Names come from loop.DefaultRules:",
"water, meal, break, service_down, netdata_critical.",
"service_down is off because it cannot say WHICH service — mavpoll folds the",
"whole kuma gauge into one boolean, so the nudge is always the generic 'a",
"service on homesrv is down'. Nothing to act on, every fifteen minutes.",
"Turn it back on once Vikunja #444 lands a fact per monitor."
],
"disabled_rules": ["service_down"],
"phraser": {
"model_path": "/opt/maven/models/llm/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf",
"bin_path": "llama-server",
@@ -16,7 +26,54 @@
"telegram": {
"bot_token": "${TELEGRAM_BOT_TOKEN}",
"chat_id": "${TELEGRAM_CHAT_ID}"
"chat_id": "${TELEGRAM_CHAT_ID}",
"//proxy": [
"api.telegram.org is not reachable directly from this box, so every send",
"timed out. The relay is the x-ui socks inbound on the host, port 10808;",
"192.168.240.1 is the maven_default bridge gateway, which is how a",
"container addresses the host. mavend is on that network.",
"This needs a matching ufw rule or the container's SYN is dropped:",
" ufw allow from 192.168.240.0/20 to any port 10808 proto tcp"
],
"proxy": "socks5://192.168.240.1:10808"
},
"//search": [
"The live web, searched after his own notes and before Kiwix. Only the",
"query string leaves the box — never a note, a fact, the persona block or",
"the history — and a question about him never reaches here at all.",
"The instance must have `json` in search.formats (settings.yml); a stock",
"SearXNG answers 403 to format=json and every search then fails. It is",
"addressed by container name, so it needs the same maven_default",
"attachment kiwix has, and it must listen on 9563: 8080 is taken several",
"times over on this box. No instance reachable ⇒ she falls through to the",
"ZIMs and never says the search failed."
],
"search": {
"url": "http://searxng:9563",
"max_results": 4,
"snippet_runes": 1500,
"language": "auto",
"timeout": "8s"
},
"//kiwix": [
"The offline encyclopedia, searched after his own notes and before anything",
"on the network. kiwix-server publishes 8034 on loopback only, so a container",
"cannot reach it by address; it is attached to the maven_default network",
"instead and addressed by container name. That attachment is imperative and",
"does not survive recreating the kiwix stack — make it declarative there:",
" networks: [default, maven_default] # maven_default: external: true",
"The book is the catalog name from the /content/... href in",
"/catalog/v2/entries, not the display title. Others on the box:",
"ifixit_en_all_2025-06, devdocs_en_ansible_2025-10."
],
"kiwix": {
"url": "http://kiwix-server:8080",
"book": "wikipedia_en_all_maxi_2026-02",
"max_results": 5,
"snippet_runes": 1500
},
"digest": {
@@ -1,6 +1,45 @@
# maven — feature ranking
> dated 2026-07-03. companion to `DESIGN.md` (folded from the former `maven.md`). ranks everything discussed post-repo-state against the infra blockers, not a replacement for the build order.
> **Archived 2026-08-02 (V-447).** The mandatory and easy tiers are now Vikunja tasks
> 449-458. Two of those closed immediately, because the ranking was stale. Quiet hours
> (V-450) ship as `QuietHoursConfig` plus the care gate in `internal/loop/loop.go`.
> Schema migrations (V-451) ship as `internal/store/migrations.go` on `PRAGMA
> user_version`. The doable and epic tiers stay here because they are reasoning. Some of
> them exist only to record why something is not worth doing yet. Read this for the why,
> not as a work queue, and check the code before believing a gap.
> dated 2026-07-03. companion to `docs/design.md` (folded from the former `maven.md`). ranks everything discussed post-repo-state against the infra blockers, not a replacement for the build order.
---
## what already shipped (checked against the code, 2026-08-02)
One month old and already wrong in ten places. Everything below is marked
built after reading the code, not the board. Read the tiers underneath with this list in
hand.
Infra 1 and 2, the two the ranking says block every feature, are both done. sqlcipher
at-rest ships as `Store.enc` plus `OpenEncrypted`, a tmpfs working copy re-encrypted on
`Close`, keyed from `db_key_env` in `deploy/mavend.json`. mavweb and mavcaldav are no
longer at zero coverage: seven test files under `cmd/mavweb`, including
`credentials_test.go` and `passkey_prf_test.go`, and two under `cmd/mavcaldav`. Infra 3
is stale in the other direction. There are still no systemd units, but the deploy is
`deploy/ecosystem/docker-compose.yml`, not scripts and tmux.
Doable tier, built: rule trace and explanation as `/trace` plus `internal/loop/explain.go`.
Recurring reminders as the cron column, `NextFireTs` and `RescheduleReminder`
(`internal/store/reminders.go:206`), so "fires once right now" is wrong. Stale-reminder
burst collapse as `collapseReminders` (`internal/loop/gather.go:209`). Revert as
`VoidLatestFact` (`internal/store/facts.go:300`). Digest mode as `internal/store/digest.go`.
Testing infra as the simulator and the eval lab (V-284, V-278). Passkey persistence as the
JSON-backed `credentialStore` in `cmd/mavweb/credentials.go`. Most integrations shipped as
their own QA tasks (V-246 mail, V-256 smarthome, V-258 rss, V-259 crawler).
Doable tier, still open: correx, systemd units, memory decay and duplicate detection
(nothing in `internal/memory` touches it), backup automation, import and export, barge-in.
Epic tier, unbuilt as ranked. `event.Bus` exists (`cmd/mavend/intake.go:57`) but it is the
intake journal from V-283, not the rewrite of facts into projections that this tier means.
---
@@ -20,7 +59,7 @@ nothing feature-level below should land before 12 are done. 34 can interle
### mandatory
things that block correctness or safety of stuff already shipped — not new capability, just closing gaps in existing design.
- **destructive-confirm policy** — open question in `DESIGN.md` § open questions, blocks correx and any new tool domain from having a coherent risk tier
- **destructive-confirm policy** — open question in `docs/design.md` § open questions, blocks correx and any new tool domain from having a coherent risk tier
- **quiet-hours definition** — open question, blocks proactive delivery being trustworthy
- **schema migrations** — sqlcipher rollout alone forces a schema touch. want this mechanism before that, not after.
@@ -30,7 +69,7 @@ cheap, no dependencies, no new invariants.
- **grocery / `list_items` table** — fourth append-only shape (item, status, list-tag), no predicate touches it, multi-adder just works for free
- **go.mod tidy**
- **capability model** (deepseek) — `homelab.docker.restart` instead of flat `tool→enabled`. cheap now, expensive to retrofit once tools surface passes ~15 entries. time-sensitive, not urgent.
- **conversation repair** — already free: `DESIGN.md` has "misroute correction = new centroid example," this is just naming the existing mechanism as a feature
- **conversation repair** — already free: `docs/design.md` has "misroute correction = new centroid example," this is just naming the existing mechanism as a feature
- **command history** — read-only query over existing facts, no new mechanism
- **clarification templates** — canned phrasing for the router's existing confidence-gate fallback, phraser-lane only
- **pronunciation dictionary** — tts config, no architecture
@@ -30,7 +30,7 @@ critical workflows: voice turn (mic→STT→route→tool/reply→TTS); proactive
(/dash /chat /tools /ecosystem)
current state: all 34 test packages pass; vet clean; `make test` still exits 1
(finding 2)
known failures: weak RU query routing — REARCH.md names the cause; the named fix
known failures: weak RU query routing — docs/rearchitecture.md names the cause; the named fix
is wired `nil`
maintenance burden: 4,518 lines of root markdown vs 33,319 lines of Go; 15 top-level
.md files, 3 of them dated session logs; several contradict
@@ -47,11 +47,11 @@ what's obsolete: llmrouter.go (built, tested, never wired); classifier seed-p
```
**Classification: healthy + misaligned.** Not fragile, not overbuilt, not abandoned. The
architecture in `REARCH.md` is sound and mostly *built* — it just is not *connected*.
architecture in `docs/rearchitecture.md` is sound and mostly *built* — it just is not *connected*.
## what it should become
The thing `REARCH.md` already describes, with the switch flipped and the drift removed:
The thing `docs/rearchitecture.md` already describes, with the switch flipped and the drift removed:
one resident small model doing both routing and phrasing, classifier demoted from the live
path to the failure floor, embedder demoted to RAG hint. No new architecture is needed.
**The gap is a config/wiring decision plus doc convergence, not a redesign.**
@@ -65,19 +65,19 @@ path to the failure floor, embedder demoted to RAG hint. No new architecture is
`architecture` / `repair`
**problem:** The most load-bearing design decision in the project is stated four different,
incompatible ways, and the code path `REARCH.md` calls "the linchpin" is disabled.
incompatible ways, and the code path `docs/rearchitecture.md` calls "the linchpin" is disabled.
**evidence** (all confirmed):
- `cmd/mavend/voice.go:211``rtr := buildRouter(emb, matcher, threshold, nil) // LLM router disabled`,
with comment *"the classifier handles routing reliably."*
- `REARCH.md:11` says the same classifier is *"the structural cause of 'she messes up
- `docs/rearchitecture.md:11` says the same classifier is *"the structural cause of 'she messes up
queries.'"* **The code comment and the design doc make opposite claims about the same
component.**
- `internal/router/llmrouter.go` (139 lines) + `llmrouter_test.go` — fully built and
tested, zero non-test callers.
- Model identity, four ways: docs say **Qwen3-1.7B** (`CLAUDE.md:6`, `REARCH.md:15`,
`SPEC.md:46`, `AGENTS.md:79`, `MAVEN_ECOSYSTEM_ARCHITECTURE.md:72`);
- Model identity, four ways: docs say **Qwen3-1.7B** (`CLAUDE.md:6`, `docs/rearchitecture.md:15`,
`SPEC.md:46`, `AGENTS.md:79`, `docs/ecosystem.md:72`);
`deploy/mavend.json:9` says **Qwen3.5-2B-UD-Q4_K_XL**; `models/llm/` on disk holds
**LFM2.5-1.2B-Thinking**; code comments in 5 files still say **LFM**.
- `deploy/mavend.json:11` sets `"n_gpu_layers": 99` while `CLAUDE.md:4` states the target
@@ -100,7 +100,7 @@ match. Delete nothing from `internal/router` yet — the classifier is the fallb
reconciliation, not a refactor. **Do not rewrite the router.**
**alternatives:** Delete `llmrouter.go` and commit to the classifier — only defensible if
the eval harness shows the classifier is actually adequate, which contradicts `REARCH.md`.
the eval harness shows the classifier is actually adequate, which contradicts `docs/rearchitecture.md`.
**risk:** Low-moderate. LLM route failures already fall through to the classifier
(`router.go:88-96`), so a bad model cannot break a turn. The real risk is CPU latency.
@@ -228,21 +228,21 @@ after finding 1, not before** — and skip it if it stays purely cosmetic.
**problem:** 15 root markdown files, 4,518 lines, several stale or superseded, at least
three pairs contradicting each other.
**evidence:** `ROADMAP.md` (759) + `MAVEN_ECOSYSTEM_ARCHITECTURE.md` (884) +
**evidence:** `ROADMAP.md` (759) + `docs/ecosystem.md` (884) +
`PROGRESS.md` (456) + `maven.md` (413) + `20-07-2026-BACKLOG.md` (396) +
`SESSION-05-07-2026.md` + `SESSION-06-07-2026.md` (477 combined) + `PLANS.md` (25) +
`START.md` + `SPEC.md` + `PROTOCOL.md`. `PROGRESS.md:61` annotates its own staleness:
*"Older LFM references below describe the currently deployed..."*. `REARCH.md` announces it
`docs/operations.md` + `SPEC.md` + `docs/protocol.md`. `PROGRESS.md:61` annotates its own staleness:
*"Older LFM references below describe the currently deployed..."*. `docs/rearchitecture.md` announces it
"supersedes" a model still described as current elsewhere.
**impact:** The doc set is the reason finding 1 exists. When five documents describe the
architecture, the code becomes the only trustworthy one — which defeats the purpose of
having them.
**recommended action:** Keep `CLAUDE.md` (agent contract), `REARCH.md` (target
architecture), `AGENTS.md` (recipes), `PROTOCOL.md` (wire format),
**recommended action:** Keep `CLAUDE.md` (agent contract), `docs/rearchitecture.md` (target
architecture), `AGENTS.md` (recipes), `docs/protocol.md` (wire format),
`20-07-2026-BACKLOG.md` (live queue). Delete the two `SESSION-*.md` and `PLANS.md` — git
history holds them. Fold `SPEC.md` + `maven.md` + `ROADMAP.md` into one `DESIGN.md` and
history holds them. Fold `SPEC.md` + `maven.md` + `ROADMAP.md` into one `docs/design.md` and
mark superseded sections instead of leaving them to read as current. Target ~1,500 lines.
---
@@ -320,7 +320,7 @@ expected maintenance gain: none over the incremental path
suggesting Vulkan offload is intended and working — but `CLAUDE.md` says CPU-only. Likely
the doc is stale, not the config; unverified.
- **Whether the classifier is genuinely adequate.** `voice.go:211` asserts it is;
`REARCH.md` asserts it is not. Both are claims, neither is measured. The uncommitted eval
`docs/rearchitecture.md` asserts it is not. Both are claims, neither is measured. The uncommitted eval
harness is the instrument to settle it — resolve before flipping the router, not after.
- **Whether wg+nginx+auth actually fronts 9201 in production.** Not in this repo. If it
does, finding 3 drops from "unauthenticated RCE" to "the control is not reproducible from
@@ -350,7 +350,7 @@ Key claims independently re-verified against the working tree; the verdict stand
over WireGuard on homesrv, this is hygiene, not an emergency — but the loopback bind
and startup warning are cheap insurance either way, so do them regardless.
- Finding 1's "flip the router" step should be gated harder on measurement.
`REARCH.md`'s claim that the classifier causes weak RU queries is itself unmeasured —
`docs/rearchitecture.md`'s claim that the classifier causes weak RU queries is itself unmeasured —
the review admits this under uncertainties, but the "repair now" ordering buries it.
Run `eval_scenarios_test.go` against both paths **before** deciding to flip, not
after. A 2B model on CPU may add enough latency that the classifier wins in practice
+19 -12
View File
@@ -1,10 +1,12 @@
# Maven — Design
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
> Folded 2026-07-30 from `SPEC.md` (north star, 2026-07-03), `maven.md`
> (consolidated decisions, 2026-06-30) and `ROADMAP.md` (execution plan,
> 2026-07-06). Those three files are gone; git history holds them.
> This is the single design document: principles, target state, and the
> execution ledger. `REARCH.md` remains authoritative wherever it disagrees
> execution ledger. `docs/rearchitecture.md` remains authoritative wherever it disagrees
> with anything here. Everything the three sources asserted that is no longer
> the intended design is preserved under **§ Superseded** — do not read that
> section as current.
@@ -160,7 +162,7 @@ presence_state ( last_bucket, last_score, updated_ts )
```
Facts additionally carry `Subject`/`EntityID`/`ResolutionState` for
entity-aware resolution against Nexus (see `MAVEN_ECOSYSTEM_ARCHITECTURE.md`).
entity-aware resolution against Nexus (see `docs/ecosystem.md`).
### Trigger model
@@ -196,7 +198,7 @@ INTO the gate as an env predicate, not the LLM's job.
## Reactive path — routing
**Target design: LLM-as-router** (see `REARCH.md` and `CLAUDE.md`). One
**Target design: LLM-as-router** (see `docs/rearchitecture.md` and `CLAUDE.md`). One
resident model emits GBNF-constrained structured JSON, and the same model
phrases replies; the embedder is a RAG hint, not a routing gate. The
committed default today is the classifier/embedder cascade, which is an
@@ -655,7 +657,7 @@ daemon.
The voice wire protocol (length-prefixed JSON frames over TCP) is designed for
**multiple client implementations**. The reference PWA at `cmd/mavweb` is one
client; any app (phone, desktop CLI, smartwatch) can implement the same frame
protocol. The published spec is `PROTOCOL.md` — **generated from
protocol. The published spec is `docs/protocol.md` — **generated from
`internal/voice/wire.go`**, not composed freehand, so it can't drift from
code. It covers transport (4-byte big-endian length prefix), methods
(`PushToTalk`, `Pong`), push kinds (`AudioNudge`), surface identity
@@ -670,8 +672,8 @@ Broadening to home automation, media or comms is JSON, not code.
## Execution ledger
Condensed from `ROADMAP.md` (2026-07-06). The live queue is
`20-07-2026-BACKLOG.md`; current state is `PROGRESS.md`.
Condensed from `ROADMAP.md` (2026-07-06). The live queue is the Vikunja board
(project Maven, ID 2); this table is history, not a work list.
| # | Item | Prio | Status |
|---|------|------|--------|
@@ -755,11 +757,13 @@ Kept for provenance. **None of this is the current or intended design.**
stay deterministic — "classifier owns the route, the SLM stays in its
phrasing lane" — with an embedding + nearest-centroid stage 1 over ~10
examples per intent, and misroutes appended as new centroid examples.
*Replaced by* LLM-as-router (`REARCH.md`): one resident model emits
*Replaced by* LLM-as-router (`docs/rearchitecture.md`): one resident model emits
GBNF-constrained JSON and also phrases replies; the embedder is demoted to
a RAG hint. The classifier cascade is still the code path that runs today
(`llmrouter` is wired nil) but it is an interim stopgap, and it is the known
cause of weak RU query handling — not a design to extend.
a RAG hint. *Landed 2026-07-31:* the LLM router is on by default and set
`true` in `deploy/mavend.json`. The classifier cascade stays as the failure
floor — it runs when there is no llama-server to talk to and on any per-turn
LLM error — but routing by seed similarity is the known cause of weak RU
query handling and is not a design to extend.
- **Named STT/TTS model picks.** `maven.md` picked faster-whisper small/int8
as primary STT with vosk RU for a low-latency command grammar, and silero
(license unverified) as TTS with piper RU as the floor, all on
@@ -769,8 +773,11 @@ Kept for provenance. **None of this is the current or intended design.**
- **Small-model phrasing claim.** `maven.md` specified "lfm2.5 / sub-1b for
phrasing — prompted, not trained," and `SPEC.md` named a specific resident
size. Both are superseded by the RU-CPT + joint persona/router SFT plan.
*Resolved 2026-07-30 (#318):* the resident checkpoint is **Qwen3.5-0.8B**
now, with the CPT'd **Qwen3-1.7B** as the target (#122). Note the resident
*Resolved 2026-07-30 (#318), revised 2026-07-31:* the resident checkpoint is
stock **Qwen3-1.7B** (`UD-Q4_K_XL`, `n_ctx` 4096), which replaced
Qwen3.5-0.8B after measuring better on both fixtures
(`docs/evals/2026-07-31-model-bakeoff.md`). The CPT'd **Qwen3-1.7B** remains the target
(#122); what stock gets wrong is the persona, not the Russian. Note the resident
model is no longer described as untrained — the target is trained
end-to-end, which is the substantive change from the old claim.
- **sqlcipher at rest.** `maven.md` specified sqlcipher with the key read at
+292
View File
@@ -0,0 +1,292 @@
# Deterministic logic around a small model
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
Written 2026-08-02. Branch `fix/integrated`.
## The question
Where does deterministic code attach, so that it helps the resident 1.7B now and
does not fight a larger model later.
## The mistake to avoid
Everything deterministic we have added so far sits in front of the model and
preempts it. Stage 0 matches, the model never sees the turn. That shape helps a
weak model and blocks a strong one, silently.
The fix is not to remove it. The fix is to know which rules are safe in that
position and to have a way to measure the rest.
## Four attachment points
**Bypass, before the model.** The only shape that saves the 2.7s p50. Safe when
the rule is a decision procedure over a closed set, not a guess over an open one.
Exact match and clock queries qualify.
**Evidence, beside the model.** Extractors emit candidate slots as a prior, not a
verdict. The prompt carries the prior and the validator reuses it. A small model
leans on it, a large one overrides it correctly.
**Grammar, around the model.** GBNF built from live state rather than hardcoded.
Costs nothing at runtime and prevents the error instead of catching it.
**Repair, after the model.** Validation failure re-asks with the specific error
rather than overriding. Self-retiring, because a better model trips it less.
## The constraint that ranks them
Latency must stay minimal. That demotes repair and promotes grammar.
- Grammar first. Zero runtime cost, immediate gain.
- Bypass keeps its place. It is the only thing that avoids a model call at all.
- Repair only where failure is rare, capped at one retry.
- Evidence is correct but costs a model call where a bypass costs none.
- Ecosystem calls belong in the snapshot path, in parallel, on strict deadlines.
Never serial before routing.
## The line that never moves
Separate policy from capability compensation. They look alike and age oppositely.
Capability compensation exists because the model is weak. It should be measurable
and retirable.
Policy exists because we decided. The personal boundary, the feminine persona, the
never-search-his-notes rule, the Hexis allowlist and confirmation binding. None of
those yield to a smarter model. A larger model is more dangerous there, not less.
## What we keep
Every stage-0 rule stays exactly as it is. Retiring them was the wrong call and it
would throw away a day of measured gains.
- exact-match fast path
- clock rules
- `SystemTimeDateGrammars`
- `AgendaQueryGrammars`
- `thinSingleToken`, including the social lexicon and the verb-ending test
Three additions, none of which change behaviour:
1. Each rule gets an id and a fixture subset.
2. Each rule writes one trace line when it fires.
3. Each rule carries a comment saying whether its set is closed or open.
That preserves today's accuracy and buys the option to revisit later with numbers.
## What we build
**Dynamic grammars from live state.** The grammars today are static: `routeGrammar`
fixes the 7 intents, `responseGrammar` fixes the mood enum, kiwix `queryGrammar`
fixes word shape. Everything else is a free string.
Candidates in order of payoff:
- **Hexis capability ids.** After discovery the exact list is known. As an enum,
the model cannot name a capability that does not exist.
- **Act fn allowlist.** Hits the two remaining false clarifies directly. They are
the act-with-no-allowlisted-fn arm of `gateLLMDecision`, firing on invented verbs.
- **Calendar names.** Enumerate the real ones for agenda and query slots.
- **Known fact keys.** A read-back matches a stored key instead of inventing a
synonym. This is the general form of the read side we scrapped on 2026-08-01.
- **Nexus display names as act targets.** Only while the list stays small.
Two rules or it backfires:
- **Always include an escape value.** A closed enum with no `other` forces a wrong
pick instead of a decline. The escape is what feeds the clarify gate.
- **Cache the grammar string, keyed on the state that built it.** Rebuilding per
turn is fine. Recompiling a large grammar per turn is not.
## Retrieval over regex
Resolve against stores that already exist rather than adding patterns.
Identity is the worked example. Nexus is authoritative, `actionFact` already sets
`Subject`, and `cmd/mavend/factenrichment.go` resolves it in the background. The
scrapped work invented a parallel key namespace with nothing reconciling the two.
A table that grows with real data beats patterns that grow with our patience.
Note a real gap: the personal boundary in `cmd/mavend/actions_query.go` guards
Maven's own store only. It does not know Nexus or Praxis exist.
## Prerequisites
Both are offline and need no deploy. Neither existed on 2026-08-01, and that is
why the day cost what it did.
**Failure taxonomy over the 77-case fixture.** Classify every miss as model
ignorance, contract loss, prompt ambiguity, or our own bug. Only model ignorance
deserves deterministic compensation. The other three get fixed once, for every
model size. Inference, not verified: much of what we patched was the last two.
**Per-assist ablation runner.** One command toggles each assist off and reports the
accuracy delta on its fixture subset. Then retiring an assist is a config flip and
a number, not an argument.
## Held
Not now, and nothing gets built for it.
- A 12B or 35B model. It may be cloud or the 16GB workstation, and cloud crosses
the current no-third-party line.
- The escalation tier in `gateLLMDecision`.
- Converting the stage-0 heuristics to evidence.
One thing carries forward for free: deterministic assists emit confidence, never a
verdict. A verdict cannot escalate.
## Ruling on the idea list
Nineteen ideas, judged on whether they earn a place in Maven. Checked against the
tree on 2026-08-02, not from memory.
### Build. Absent, and worth it.
**SQLite FTS over embeddings.** No `fts5` anywhere in the tree. Lexical search is
faster than the ONNX embedder, deterministic, and strongest exactly where the
embedder is weakest, which is exact Russian names. Semantic search becomes the
fallback rather than the gate. This is the highest-value absent item.
**Cached TTS phrases.** No cache in `internal/tts`. Confirmations, clarifies and
refusals repeat constantly and their text is already fixed. Pre-rendering them is
cheap and pays straight into the minimal-latency constraint.
**Synthetic router dataset.** Generate Russian tool-calling examples from the same
schemas that will build the dynamic grammars. One source of truth for both, so the
model is trained on exactly the shapes it will be constrained to at inference.
**Command STT separate from dictation STT.** One path today. Commands want latency
and a small vocabulary, meeting capture wants accuracy and can take its time.
`internal/capture` already spools to disk, so the split follows the existing seam.
Medium priority, behind the four above.
### Polish. Present, incomplete.
**Strict JSON everywhere.** Done 02-08-2026. The replier sends
`phraser.ResponseGrammar`. It is exported once so its two copies cannot drift.
The meeting summariser is wrapped and unwrapped in the daemon's Completer, so
`internal/capture` stays text-in/text-out. Every model call now carries a grammar.
**Evidence-first prompting.** Done 02-08-2026, in the evidence branch of
`PhraseQuery`. Sources arrive numbered, one per line. The system prompt no longer
calls them all "заметки" and no longer lets the model add anything of its own.
Blank sources now take the knowledge branch instead of asking for an answer from
an empty list. Not yet measured against a live model. Run `make eval-phrasing`,
and watch the case `query-notes-do-not-answer`.
**Progressive inference.** What exists is a fallback cascade, not escalation. On
error it drops to something weaker. It never escalates on ambiguity. The upgrade is
held with the bigger-model question, and `gateLLMDecision` is the hook.
**Entity dictionaries.** Nexus is the canonical store, `behavior_ru.go` has
`KeyAliases`, `ecosystem_acts.go` has verb aliases. Morphology is the gap, and the
code says so in three places. Russian needs it and there is no stemmer in the repo.
**Assistant state machine.** `confirm.go` models pending confirmation and
`dialogue.Session` models the turn. There is no unified task state. Reuse the
Praxis vocabulary rather than inventing one, because surfaced, acknowledged and
resolved already mean something precise here.
**Session memory compiler.** The digestion tick, `internal/memory` and
`followUpMerge` do parts of this. Not a coherent compile step.
**Background memory maintenance.** Mostly present, and the absent part is small.
What exists: `internal/memeval` reads recent memory on a loop and writes
observations, deduped against its own prior output, unable to speak or act. Facts
supersede at write time through `voidsID`, `CorrectValue` and `VoidLatestFact`,
and `RecentActiveFactsByKind` reads only live rows. Digests, ecosystem traces and
media all prune.
Three gaps remain, all in the fact store:
- Superseding is turn-driven. Nothing reconciles two live facts that contradict
unless a turn corrects one of them.
- Duplicates written by different sources or phrasings stay as separate live rows.
There is no key-level merge pass.
- Nothing expires. The log is append-only and grows without bound, and no fact
ever ages out on its own.
Worth doing, and smaller than it looked. Not urgent.
**Qwen CPT then narrow SFT.** In flight as #122. One disagreement with the idea as
written: do not drop persona from the SFT. Persona is the stated reason the CPT
exists, because stock writes `рад` where Maven needs `рада`. Keep the joint
router-plus-persona tune.
**Offline job queue.** The tick loop, `factEnrichmentWorker` with backoff, and the
media prune already defer work. A general queue is tidier, not more capable. Low
priority.
**Response templates.** Present in `clarify.go`, in the `replySystem` arms and in
`StubPhraser`. Worth extending to high-frequency confirmations, where latency and
persona correctness both matter. Do not extend further. Templating the
conversational reply removes the reason she is worth having.
### Reject.
**Grammar-first routing as a replacement for the LLM router.** Already measured.
The classifier scores 36.8% full accuracy against 72.7% through the cascade.
Replacing the model with rules halves the accuracy. Grammar-first as an ordering is
what stage 0 already is, and that stays.
**Hierarchical intent classification.** Seven intents is already the coarse layer.
The specialisation stage exists as per-intent slot filling in `Extractor.Extract`.
Adding a tier buys structure, not accuracy.
**Tool-specific micro-models.** Contradicts the one-resident-model constraint,
needs per-domain training data nobody has, and multiplies model loads on a single
Vega iGPU. Dynamic grammars give the same domain narrowing at zero runtime cost.
**Local knowledge graph.** Nexus owns entities and relationships. Building a second
graph in Maven breaks the ecosystem line and creates two answers to one question.
If graph traversal is wanted, it is a Nexus feature request.
**Preemptible training.** Training runs in a separate workspace, not on the serving
box. This only becomes real if CPT moves onto homesrv, and that is not the plan.
## Known open, carried over from the deleted handoff
Found live on 2026-08-01, not fixed. Everything else in that file was stale.
- **Kiwix ranks badly on a correct query.** The stop-word pass eats the "and". So
"кто написал войну и мир" reaches Kiwix as "war peace author". The top hit is
"List of peace activists" and she summarises that as the answer. The tag
`scrapped/fact-and-kiwix-phrases` fixes the query text. The ranking is the ZIM
search and is untouched either way.
- **Chat drags prior turns into an answer.** One live reply mixed the greeting, the
height statement and a world question. It named Левитан as the author of Война и
мир.
- **`safeKey` drops Cyrillic**, so Russian calendar events on one day collide.
Vikunja #443 with three fix options. It is a migration, not a patch.
## Open after the 02-08-2026 deploy
SearXNG runs on homesrv at `http://searxng:9563`, on `maven_default`, and the
rebuilt `mavend` wires it. "кто написал войну и мир?" now routes to query, takes
4 results off the search, and answers Толстой with the lookup opener. Two things
that turn left unsettled.
- **The turn was slow, and nobody knows yet whether that is real.** Route 7s,
search 1s, phrasing 15s. The p50 in `docs/evals/2026-07-31-routing.md` is 825ms. It
was the first turn after a cold start with the model still warming, so it
proves nothing either way. Re-run the same question warm before treating it as
a regression. Do not plan latency work off this number.
- **The personal boundary has never run live.** `queryPersonal` in
`cmd/mavend/actions_query.go` stops a question about him from reaching
SearXNG. The question above is not one, so only tests cover it. Ask something
about him on the deployed box and confirm from the log that no `voice: search`
line appears for it.
## Sequence
1. Failure taxonomy over the fixture.
2. Ablation runner, plus ids and trace lines for the existing rules.
3. Dynamic grammar for the act fn allowlist.
4. Dynamic grammar for Hexis capability ids.
5. Dynamic grammar for calendar names and known fact keys.
6. Extend the personal boundary to the ecosystem stores.
Steps 1 and 2 come before anything is written in the router.
@@ -1,5 +1,7 @@
# Maven Ecosystem Architecture
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
## 1. Purpose
This document defines Maven's role in the local ecosystem formed by:
@@ -16,7 +16,7 @@ with Qwen3-1.7B.
Settles Vikunja **#278 / #250**.
- Same fixture and scorer as `ROUTING-EVAL-31-07-2026.md`: `internal/router/eval/`
- Same fixture and scorer as `docs/evals/2026-07-31-routing.md`: `internal/router/eval/`
(`ru_routing_v1.json`, 76 held-out cases).
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:<port> make eval-router`
(`TestLLMRouterBaseline`). (This line used to say there is no `make eval-models` target.
@@ -148,7 +148,7 @@ Qwen3-1.7B wins every column, including against a model 20% larger than it.
| ontopic | 16, 19, 19 | **22, 23, 23** |
| canned fallbacks | 8, 5, 6 | **0, 2, 0** |
This also fills the row `TALK-EVAL-31-07-2026.md` had to void for contamination:
This also fills the row `docs/evals/2026-07-31-talk.md` had to void for contamination:
**600ch/1024tok on Qwen3.5-0.8B scores 13, 11, 8.**
`address` is the headline. It sat at 18-22 of 27 on the 0.8B no matter how the prompt
@@ -162,6 +162,11 @@ The 1.7B does that 0-2 times.
## Latency — the long tail is not the Thinking block
> **Stale, corrected 2026-08-02.** The p50 figures in this table are contention on a
> shared llama-server, not the model's cost. The router measures p50 825ms / p95 1.2s /
> max 3.0s in `docs/evals/2026-07-31-routing.md`, which says so at line 61. Read this table for
> the shape of the tail only. Take absolute latency from the routing eval.
| | p50 | p95 |
|---|---|---|
| Qwen3.5-0.8B | 2.4s, 2.9s, 2.0s | 17.4s, 17.6s, 17.4s |
@@ -213,11 +218,13 @@ swapped again when the CPT lands.
- ~~The routing numbers only reach production once the LLM router is wired on. It is
still `nil`.~~ **Resolved the same evening:** the LLM router is wired at `voice.go:214`
behind `voice.llm_router`, the default is on, and `deploy/mavend.json` sets it `true`.
These numbers are the production path now, so the p50 ≈2.7s is a real per-turn cost and
not a bench artifact.
These numbers are the production path now. **Corrected 2026-08-02: the p50 ≈2.7s in the
latency table above WAS a bench artifact.** It is contention on the shared llama-server,
not the model. `docs/evals/2026-07-31-routing.md` line 61 says so, and measures the router at
p50 825ms / p95 1.2s / max 3.0s. Cite that file for latency, not this one.
- ~~`/mnt/hdd1/llms/LFM2.5/Qwen3-1.7B-UD-Q4_K_XL.gguf` is a 293 MB truncated download
in the wrong directory.~~ **Deleted 2026-07-31.** The good 1.13 GB copy in `qwen3/` is
what `deploy/mavend.json` loads.
- Harness: `scratchpad/bakeoff.sh`, one server at a time, health-checked before each
run, `/v1/models` recorded per run. Never run two LLM consumers at once — see the
contamination note in `TALK-EVAL-31-07-2026.md`.
contamination note in `docs/evals/2026-07-31-talk.md`.
@@ -1,7 +1,7 @@
# Phrasing evaluation — 31-07-2026
How Maven words a nudge, measured instead of argued. Counterpart to
`ROUTING-EVAL-31-07-2026.md`.
`docs/evals/2026-07-31-routing.md`.
- Fixture + scorer: `internal/phraser/eval/` (`nudges_v1.json`, 15 cases; `eval.go`, `checks.go`)
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing`
@@ -65,7 +65,7 @@ prefixes) is the targeted fix, and it would move findings 1 and 2 together. Sepa
`model_quantized.onnx` — not the same file.
`hard` cases score **2/11**: every one is a query where the operator did not reuse his own words.
That is the normal case weeks later, and exactly what DESIGN.md's "recall when relevant" promises.
That is the normal case weeks later, and exactly what docs/design.md's "recall when relevant" promises.
### 4. The memory-store recall branch is dead for notes
@@ -177,7 +177,7 @@ was silent ("не знаю" to "который час") while the one it introdu
### 1. The resident model does route better — 50.0% vs 36.8%
REARCH.md's premise holds; `voice.go:211`'s comment does not. **But the classifier is only
docs/rearchitecture.md's premise holds; `voice.go:211`'s comment does not. **But the classifier is only
~37% correct on held-out utterances, and the model only ~50%.** Neither is "reliable". The
gap between them is real but both are far from a system you would describe as working.
@@ -141,7 +141,7 @@ a model check, which catches a dead server but not a loaded one.
measured) on this fixture and the router fixture. Not the 4B — too big for
this box, owner's call.
- Newer sub-500M candidates (LFM2.5 200M/300M) are worth a run for routing.
Note `MODEL-BAKEOFF-31-07-2026.md` found LFM2.5-**1.2B** worse than
Note `docs/evals/2026-07-31-model-bakeoff.md` found LFM2.5-**1.2B** worse than
Qwen3.5-0.8B at Russian routing and 2.4× slower — but those are a different,
older generation, so that result does not predict the small ones.
- Fix `chat-how-are-you`'s `want_any`, and re-baseline once, so `ontopic`
+2
View File
@@ -1,5 +1,7 @@
# Start Commands
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
All commands assume `ROOT=/home/kami/apps/Maven` and the local Go toolchain at `$ROOT/deps/go/go/bin/go`.
## Prerequisites
@@ -6,7 +6,7 @@
> use the embedded LFM model paths or old single-object examples as current ops
> guidance; see `2026-07-18-qwen3-resident-training-eval.md`.
> Scope from `REARCH.md`. Make Maven trustworthy: the LFM becomes the router
> Scope from `docs/rearchitecture.md`. Make Maven trustworthy: the LFM becomes the router
> (fixes "messes up queries" / "doesn't take notes"), the engine actually runs
> (fixes stub replies), dates stop being read as "number dot number dot number",
> and telegram becomes a reach channel. NOT in scope: on-demand 4B reasoner,
@@ -693,7 +693,7 @@ ssh kami@192.168.1.104 'curl -s localhost:9201/api/chat -d "{\"text\":\"запо
4. `docker compose up -d mavend && docker logs -f maven-mavend-1` — confirm the
phraser spawns and no `phraser: NewStub` path. Run the two verify curls.
5. Update `AGENTS.md`: LFM model download + note that routing is now LFM-first
with classifier fallback (`REARCH.md` is the design of record).
with classifier fallback (`docs/rearchitecture.md` is the design of record).
---
+2
View File
@@ -1,5 +1,7 @@
# Maven Voice Protocol
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
> Auto-generated from `internal/voice/wire.go`, `internal/voice/errors.go`,
> `internal/voice/frame.go`, `internal/voice/client.go`. If this file and
> those files disagree, the code wins.
+186
View File
@@ -0,0 +1,186 @@
# QA plan: checking Maven properly
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
Written 2026-08-01, after the 35-PR stack landed and the box came back up.
44 of the 50 open Vikunja tasks are `QA:` tasks. They are verification work, not
build work. Most sat unverifiable while Maven was down for 11 days. That
blocker is gone.
This plan orders them by what unblocks what. Do sessions 1 and 2 first. Almost everything
downstream assumes the voice loop works, and nobody has confirmed that since
the redeploy.
---
## Before you start
Two things bite anyone running these checks on homesrv.
**curl needs `--noproxy '*'`.** The shell exports `http_proxy=http://127.0.0.1:18080`.
Without the flag, every local check returns 503 from the proxy and looks like a
dead service. This cost me a false regression report today.
**The database is not readable with sqlite3.** Four older QA steps say
`docker compose exec mavend sqlite3 /data/maven.db "select ..."`. That cannot
work: the container has no `sqlite3` binary, and the store is AES-256-GCM at
rest with a tmpfs working copy. Read state through mavweb instead, at
`/history`, `/trace`, `/routines` and `/dash`.
---
## Session 1: the voice loop (half a day)
Nothing here has been confirmed since the redeploy, and everything else assumes
it works. Do this first.
Closes or advances: **44** (conversation), **45** (text chat), **287** (voice
session quality), **321** steps 3-5 (quiet mode), **288** (STT fixtures).
1. Open `http://127.0.0.1:9201/chat` and hold a short conversation in Russian.
Watch for three things: she answers in feminine forms (`рада`, `поняла`), she
says `ты` and never `вы`, and no pet names appear.
2. Press push-to-talk on `/dash`. Say `привет`. Confirm a spoken reply comes
back. This is the only check that covers mic to STT to core to TTS to
speaker as one path. It is also the path the eleven-day outage most likely
broke.
3. Say `тихий режим`. Expect `тихий режим включён. буду реже напоминать.`
4. Say `выключи тихий режим`. Expect `тихий режим выключен.` Negation must win.
5. Say `в комнате тихо`. Quiet mode must NOT flip. Confirm on `/history` that no
`quiet_hours` fact was written.
6. Say `включи режим тишины`, then `сделай потише`. Both must flip quiet mode
on. These are the noun form and the comparative, added 01-08-2026.
7. Wait for a nudge, then say `потом` within twenty minutes. Expect `хорошо,
вернусь к этому позже.` and the nudge row on `/notifications` reading
`snoozed`. Say `потом` again with nothing pending: it must route as an
ordinary utterance, not be swallowed.
8. Wait for the water nudge, then say `выпил воды`. Expect the ordinary fact
reply and nothing extra — she must not congratulate you. Check
`/notifications`: the row reads `acted`. Then trigger another nudge and say
`готово`; expect `отлично, отметила.` and the same outcome.
9. Note anything where she is slow, cuts off, or talks over herself. That is
287's whole content and it has no written acceptance criteria yet.
**319 is fixed** (01-08-2026). Single-word Russian utterances no longer come
back as `не совсем поняла — можешь переформулировать?`. `привет` and `поужинал`
both pass now: `thinSingleToken` spares social singles and any token carrying a
verb ending, and only thins a bare nominal like `вода`. If a one-word utterance
still gets clarified during the smoke test, that is a new case for the lexicon,
not the old bug.
---
## Session 2: measurement (half a day, mostly waiting)
Closes or advances: **320** items 2-4, **278** (make the eval lab routine),
**319** (gate recalibration).
The resident llama-server cannot be reached by the eval harness. It binds
`--host 127.0.0.1 --port 0` inside the container, so the port is kernel-assigned
and never published. Start a second one on a fixed port instead:
```sh
llama-server -m /mnt/hdd1/llms/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf \
--host 127.0.0.1 --port 18100 -c 4096 -ngl 99 --no-webui
```
`-c 4096` matters. The recorded numbers were measured at that context size, and
a mismatch invalidates the comparison.
Then:
```sh
make eval-models MAVEN_LLM_URL=http://127.0.0.1:18100 # want ~72.7% cascade
make eval-router # classifier baseline
MAVEN_LLM_URL=http://127.0.0.1:18100 make eval-phrasing # persona checks, slow
make eval-recall
```
A large miss against 72.7% means the deploy differs from the bench harness.
Two things to decide while the numbers are in front of you:
- **319's gate recalibration.** The single-token rule needs narrowing or
dropping. This needs your judgement, not a threshold sweep. The fixture and the
daemon disagree about what is correct on two of the three false clarifies.
- **278's real ask** is making the eval lab routine rather than building it. It
is built. Decide whether it runs on a timer, on every merge, or on demand, and
the task can close.
Item 4 of **320** needs a permission I do not have. Kill the `llama-server`
pid under `maven-mavend-1`, post a turn, and confirm it still completes
through the classifier. Either grant it or run it yourself. It is the only
check that the failure floor catches a mid-session model death.
---
## Session 3: the interaction batch (a day, or three sittings)
These need real use rather than a command, grouped by what one sitting covers.
**Morning and delivery** (**280**, **281**, **128**, **282**): open `/morning`,
walk the seven required behaviours, then check the four interruption outcomes
and the digest gap. **282** needs the `desk_active` script enabled on the desk
PC first, which is **15** and needs you at that machine.
**Tasks and calendar** (**129**, **130**, **127**, **126**, **246**): capture a
task by voice, confirm it lands, check prioritisation ordering is not nonsense.
**246** (mail reader) also exercises the `IngestMail` rung that moved to
`AuthWrite` this morning.
**Routines and patterns** (**43**, **46**, **247**, **254**): these need history
to detect against. If the database is thin after the outage, they may have
nothing to propose, which is not a failure. Check `/routines` before
concluding anything.
**Ecosystem** (**272**, **273**, **276**): nexus, hexis and praxis are wired and
logged clean at boot. **276** is the degraded-mode suite, which means taking
siblings down on purpose. Worth doing while you are already in there.
---
## Housekeeping (one sitting, no box needed)
Four QA tasks will not close no matter how long they sit, because they are
gated on something that does not exist:
- **125** zenmoney: needs a token you have not minted.
- **256** Home Assistant: needs HA configured.
- **257** Bluetooth: BLOCKED, no bluez on the box. Says so in the title.
- **288** STT golden audio: needs fixtures generated.
Relabel these so they stop reading as backlog. They are not verification work
that is pending, they are work that has not started.
Same treatment for the five plan-only tasks (**251** MCP, **252** vision,
**253** hearing, **255** speaker recognition, **259** crawler). A `QA:` prefix on
a plan is misleading.
---
## Needs you specifically
Not QA. These are blocked on a decision or a credential only you have.
| # | what |
|---|---|
| 16 | Create the Kuma API key. `-kuma-key uk5_mavpoll-key` in `docker-compose.yml` is still the placeholder. |
| 15 | Deploy `desk_active` on the desk PC. Blocks **282**. |
| 122 | Finish the CPT run for Qwen3-1.7B. The persona fix depends on it. |
| 355 | Deploy the Hexis auth change. Was blocked on Maven being under construction, which it no longer is. The client half is vendored and wired. |
| 357 | Decide whether entity-existence validation is the permanent target guard or whether blessing lands in Nexus. |
| 275 | Hexis native API and MCP parity. |
| — | Decide on `-require-stepup`. Making it the default needs WebAuthn configured first, or it locks you out of your own admin surfaces. See **317**. |
| — | Three nginx sites bind wildcard `:80` (`acme.conf`, `matrix`, `panel`), so the ecosystem's bind-level protection is not in effect and `allow`/`deny` is carrying it alone. See **354**. |
---
## Suggested order
1. Session 1. If the voice loop is broken, nothing else matters.
2. The `-require-stepup` and Kuma decisions. Five minutes, unblocks **317** fully
and **16**.
3. Session 2. The numbers tell you whether the router is worth its 90x latency.
4. Housekeeping. Cheap, and it makes the remaining backlog honest.
5. Session 3, split whichever way suits you.
+2
View File
@@ -1,5 +1,7 @@
# Maven — Re-architecture (Qwen3 resident model, revised 2026-07-18)
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
> Supersedes the classifier-first routing model. Agreed in a design session
> after diagnosing that homesrv deploys with a **stub phraser** (no LLM
> running) and an embedder-classifier that routes by nearest-neighbor between
+1 -1
View File
@@ -1,7 +1,7 @@
// Package auth is maven's authority layer — the 4-layer cascade and the
// "surface caps authority" invariant.
//
// Spec contract (from DESIGN.md § Auth):
// Spec contract (from docs/design.md § Auth):
//
// a cascade, not a pick-one — each layer answers a different question:
//
+11 -2
View File
@@ -71,6 +71,13 @@ func NewSummarizer(llm Completer, chunkRunes, maxChunks int, contextBlock func()
return &Summarizer{llm: llm, chunkRunes: chunkRunes, maxChunks: maxChunks, context: contextBlock}
}
// Both prompts ask for a JSON wrapper because the daemon's Completer attaches a
// grammar of that shape (summaryGrammar in cmd/mavend/capture.go) and unwraps it
// again before the text reaches this package. The wrapper is what keeps a
// Thinking-variant model from answering a summarisation prompt with its
// reasoning. Nothing here parses it: the map and reduce steps see plain prose,
// and a Completer without the grammar still works.
//
// chunkPrompt — the map step. Deliberately plain: this is not Maven speaking to
// him, it is a model condensing text, so there is no first person in it at all
// and therefore nothing for the persona's gender rules to get wrong. The reply
@@ -79,12 +86,14 @@ func NewSummarizer(llm Completer, chunkRunes, maxChunks int, contextBlock func()
const chunkPrompt = `Ты обрабатываешь фрагмент расшифровки разговора.
Сожми его до 2-4 пунктов: о чём говорили, какие решения приняли, какие задачи назвали.
Без вступлений и выводов. Только по тексту не придумывай того, чего в нём нет.
Если во фрагменте нет ничего содержательного, ответь одним словом: пусто.`
Если во фрагменте нет ничего содержательного, напиши одно слово: пусто.
Отвечай ТОЛЬКО объектом JSON с одним полем: {"summary": "..."}.`
// reducePrompt — the reduce step. Same rules, over the chunk summaries.
const reducePrompt = `Ниже конспекты фрагментов одной встречи, по порядку.
Собери из них один короткий итог: о чём была встреча, какие решения приняли, что кому делать.
Не повторяйся, не придумывай, не добавляй вступлений.`
Не повторяйся, не придумывай, не добавляй вступлений.
Отвечай ТОЛЬКО объектом JSON с одним полем: {"summary": "..."}.`
// emptyMarker — what the map step answers for a chunk with nothing in it. Such
// chunks are dropped before the reduce step rather than padding it with noise.
+156 -1
View File
@@ -144,6 +144,22 @@ type Config struct {
// nil ⇒ quiet hours only activate via the voice toggle.
QuietHours *QuietHoursConfig `json:"quiet_hours,omitempty"`
// DisabledRules — nudge rules that are not wired at all, by name
// ("service_down", "netdata_critical", "water", "meal", "break").
//
// Rules are code, not config (see loop.DefaultRules), and that stays true:
// this only subtracts. It exists because a rule can be right in principle
// and useless in practice — kuma's service_down cannot name the service it
// is nudging about (Vikunja #444), so being told "a service on homesrv is
// down" every fifteen minutes is noise with no action attached. Turning it
// off beats learning to ignore her.
//
// A disabled rule is never gathered for, never evaluated, and never
// delivered on any channel. Unknown names are ignored, so removing a rule
// from the code does not break a config that still lists it.
// Empty ⇒ every rule runs, which is the default.
DisabledRules []string `json:"disabled_rules,omitempty"`
// Digest — notification batching / digest mode. nil ⇒ digest disabled
// (every nudge is sent as it fires — legacy behaviour).
Digest *DigestConfig `json:"digest,omitempty"`
@@ -199,6 +215,15 @@ type Config struct {
// fetches a page: not on request, not on a schedule. See CrawlConfig.
Crawl *CrawlConfig `json:"crawl,omitempty"`
// Kiwix — the offline ZIM reader (Vikunja #122 neighbourhood). nil / absent
// / url empty ⇒ the query chain has no ZIM source. See KiwixConfig.
Kiwix *KiwixConfig `json:"kiwix,omitempty"`
// Search — the SearXNG metasearch instance. nil / absent / url empty ⇒ the
// query chain has no web-search source and Kiwix is the only encyclopedia.
// See SearchConfig.
Search *SearchConfig `json:"search,omitempty"`
// Praxis — the ecosystem attention-state service. When configured, maven
// calls the Praxis HTTP tools API for attention listing and item lifecycle.
// Maven never touches Praxis's database directly (ecosystem invariant: no
@@ -624,7 +649,7 @@ type VoiceConfig struct {
// LLMRouter — route with the resident model instead of the embedding
// classifier. On by default since Vikunja #320.
//
// Measured on the held-out fixture (ROUTING-EVAL-31-07-2026.md): 63.2% of
// Measured on the held-out fixture (docs/evals/2026-07-31-routing.md): 63.2% of
// intents right against the classifier's 50.0%, and no route errors. It
// costs about 1s per turn instead of 30ms.
//
@@ -1027,6 +1052,110 @@ type CrawlConfig struct {
MaxRunes int `json:"max_runes,omitempty"`
}
// KiwixConfig — the offline encyclopedia. A kiwix-serve instance holding ZIM
// archives (Wikipedia, ifixit, devdocs) on the LAN, searched before anything
// touches the network. Dark until configured, same as every other reach.
//
// This is the "local sources first" rule in CLAUDE.md made concrete: a 1.7B
// does not know enough to answer a world question, but it can read. A local
// read costs nothing and leaves the box only as far as the LAN.
//
// Only the rewritten search query leaves this process. His notes, facts,
// persona block and history are never part of a request.
type KiwixConfig struct {
// URL — base address of kiwix-serve, e.g. "http://kiwix:8080". Empty ⇒ the
// whole block is normalised to nil and the source stays off.
URL string `json:"url,omitempty"`
// Book — the ZIM to search, by its catalog name, e.g.
// "wikipedia_en_all_maxi_2026-02". Take it from the /content/… href in
// /catalog/v2/entries; the display title is not the name.
//
// Required. kiwix-serve answers 400 to a search with an empty books.name,
// so a block without one is normalised to nil rather than left to fail one
// query at a time.
Book string `json:"book,omitempty"`
// MaxResults — how many hits are asked for. 0 ⇒ DefaultKiwixResults.
// Only the top few reach the phraser regardless; the rest are context the
// snippet ranking throws away.
MaxResults int `json:"max_results,omitempty"`
// SnippetRunes — how much of the joined snippets is handed to the phraser.
// 0 ⇒ DefaultKiwixSnippetRunes. Sized against the 4096-token context, which
// also holds the persona block and the prompt.
SnippetRunes int `json:"snippet_runes,omitempty"`
// Rewrite — turn the Russian question into English keywords with the
// resident model before searching. The ZIMs are English and kiwix ranks by
// keyword, not meaning, so a Russian sentence matches nothing. Costs one
// short LLM call per query. Default true; set false only to measure the
// difference or when the books are Russian.
Rewrite *bool `json:"rewrite,omitempty"`
}
// RewriteEnabled — Rewrite with its default applied. Absent ⇒ on.
func (k *KiwixConfig) RewriteEnabled() bool {
return k.Rewrite == nil || *k.Rewrite
}
// Kiwix defaults, applied in Normalise.
const (
DefaultKiwixResults = 5
DefaultKiwixSnippetRunes = 1500
)
// SearchConfig — the self-hosted SearXNG instance she searches with.
//
// External search is allowed and off unless configured (CLAUDE.md). Configuring
// it is the whole opt-in: no `search` block, no query ever leaves the LAN.
//
// It sits AHEAD of Kiwix in the query chain, and that is the owner's ruling of
// 2026-08-02: a live search answers better than a frozen ZIM, and the ZIM is
// what she falls back to when the line is down. Everything of HIS still comes
// first — the personal boundary runs above both, so a question about him is
// never searched.
//
// Only the query string leaves the box. Notes, facts, the persona block and the
// history are never part of a request; internal/websearch cannot read the store.
type SearchConfig struct {
// URL — base address of the SearXNG instance, e.g. "http://searxng:9563".
// Empty ⇒ the whole block is normalised to nil and the source stays off.
//
// The instance needs `search.formats` to include `json` in its settings.yml.
// A stock install answers 403 to format=json, and then every search fails.
URL string `json:"url,omitempty"`
// MaxResults — how many hits are kept as evidence. 0 ⇒ DefaultSearchResults.
// Small on purpose: the snippets share a 4096-token context with the persona
// block and the prompt.
MaxResults int `json:"max_results,omitempty"`
// SnippetRunes — how much of the joined evidence reaches the phraser.
// 0 ⇒ DefaultSearchSnippetRunes.
SnippetRunes int `json:"snippet_runes,omitempty"`
// Language — SearXNG's `language` parameter, e.g. "ru", "en" or "auto".
// Empty ⇒ the instance default. He asks in Russian and in English, so
// pinning one language here is usually the wrong call.
Language string `json:"language,omitempty"`
// Engines — comma-separated engine names to restrict the search to, e.g.
// "duckduckgo,wikipedia". Empty ⇒ whatever the instance has enabled.
Engines string `json:"engines,omitempty"`
// Timeout — per-search budget. 0 ⇒ websearch.DefaultTimeout. SearXNG waits
// on the slowest upstream engine, so this is the knob that decides how long
// a voice turn can stall on a bad network.
Timeout Duration `json:"timeout,omitempty"`
}
// Search defaults, applied in Normalise.
const (
DefaultSearchResults = 4
DefaultSearchSnippetRunes = 1500
)
// CrawlWatchConfig — one page kept an eye on.
type CrawlWatchConfig struct {
Name string `json:"name"` // note source is "crawl:<name>"
@@ -1345,6 +1474,32 @@ func (c *Config) applyDefaults() {
c.Crawl = nil
}
// Same rule for the ZIM reader: no address or no book, nothing to search.
if c.Kiwix != nil && (strings.TrimSpace(c.Kiwix.URL) == "" || strings.TrimSpace(c.Kiwix.Book) == "") {
c.Kiwix = nil
}
if c.Kiwix != nil {
if c.Kiwix.MaxResults <= 0 {
c.Kiwix.MaxResults = DefaultKiwixResults
}
if c.Kiwix.SnippetRunes <= 0 {
c.Kiwix.SnippetRunes = DefaultKiwixSnippetRunes
}
}
// Same rule for the metasearch instance: no address, nothing to search.
if c.Search != nil && strings.TrimSpace(c.Search.URL) == "" {
c.Search = nil
}
if c.Search != nil {
if c.Search.MaxResults <= 0 {
c.Search.MaxResults = DefaultSearchResults
}
if c.Search.SnippetRunes <= 0 {
c.Search.SnippetRunes = DefaultSearchSnippetRunes
}
}
if c.Voice != nil {
if c.Voice.RouterThreshold <= 0 {
c.Voice.RouterThreshold = DefaultRouterThreshold
+46
View File
@@ -367,3 +367,49 @@ func TestUpdateBlockValidatedAtStartup(t *testing.T) {
t.Error("Load accepted an update block with no health_socket")
}
}
// A kiwix block with no address, or no book, has nothing to search.
// kiwix-serve answers 400 to an empty books.name, so the block is dropped here
// rather than left to fail one query at a time.
func TestNormaliseDropsIncompleteKiwix(t *testing.T) {
for _, tc := range []struct {
name string
in *KiwixConfig
}{
{"no url", &KiwixConfig{Book: "wikipedia_en_all_maxi"}},
{"no book", &KiwixConfig{URL: "http://kiwix:8080"}},
{"blank url", &KiwixConfig{URL: " ", Book: "b"}},
} {
t.Run(tc.name, func(t *testing.T) {
c := &Config{Kiwix: tc.in}
c.applyDefaults()
if c.Kiwix != nil {
t.Errorf("kept an unusable kiwix block: %+v", c.Kiwix)
}
})
}
}
func TestNormaliseFillsKiwixDefaults(t *testing.T) {
c := &Config{Kiwix: &KiwixConfig{URL: "http://kiwix:8080", Book: "b"}}
c.applyDefaults()
if c.Kiwix == nil {
t.Fatal("dropped a complete kiwix block")
}
if c.Kiwix.MaxResults != DefaultKiwixResults {
t.Errorf("MaxResults = %d, want %d", c.Kiwix.MaxResults, DefaultKiwixResults)
}
if c.Kiwix.SnippetRunes != DefaultKiwixSnippetRunes {
t.Errorf("SnippetRunes = %d, want %d", c.Kiwix.SnippetRunes, DefaultKiwixSnippetRunes)
}
// Rewriting is on unless it is turned off: an English ZIM searched with a
// Russian sentence matches nothing, so the useful default is the on one.
if !c.Kiwix.RewriteEnabled() {
t.Error("rewriting defaulted to off")
}
off := false
c.Kiwix.Rewrite = &off
if c.Kiwix.RewriteEnabled() {
t.Error("rewrite: false was not honoured")
}
}
+1 -1
View File
@@ -1,6 +1,6 @@
// Package delivery is maven's channel-routing + dispatch layer.
//
// Spec contract (from DESIGN.md § Delivery / channel routing):
// Spec contract (from docs/design.md § Delivery / channel routing):
//
// - routing = f(severity, presence). presence decides REACHABILITY; severity
// decides INSISTENCE. need both.
+2 -2
View File
@@ -9,7 +9,7 @@ import (
"github.com/kami/maven/internal/store"
)
// This file walks every cell of the DESIGN.md § "Delivery / channel routing"
// This file walks every cell of the docs/design.md § "Delivery / channel routing"
// table, once as the pure table and once through the dispatcher, so a change
// to either side has to break a named cell.
//
@@ -205,7 +205,7 @@ func TestAwayChannelsGetMinimalBody(t *testing.T) {
// An empty Summary no longer means "send the whole body" — it means a short
// generic line — so the old expectation here was wrong as well as duplicated.
// TestCareAwayDropIsRecorded — DESIGN.md's drop is a decision ("a missed water
// TestCareAwayDropIsRecorded — docs/design.md's drop is a decision ("a missed water
// nudge is noise, a missed backup failure isn't"), so it should be visible
// rather than vanish. Today drop is a bare `continue`: no nudge row, no outbox
// attempt, no log — nothing an operator can see afterwards. now it leaves a
+43 -4
View File
@@ -29,6 +29,7 @@ import (
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
@@ -162,13 +163,13 @@ func (s *Sink) Send(ctx context.Context, d delivery.Sendable) error {
req, err := http.NewRequestWithContext(ctx, http.MethodPost, s.sendMessageURL(), bytes.NewReader(pb))
if err != nil {
return fmt.Errorf("telegramsink: build request: %w", err)
return fmt.Errorf("telegramsink: build request: %w", s.redact(err))
}
req.Header.Set("Content-Type", "application/json")
resp, err := s.hc.Do(req)
if err != nil {
return fmt.Errorf("telegramsink: sendMessage: %w", err)
return fmt.Errorf("telegramsink: sendMessage: %w", s.redact(err))
}
defer resp.Body.Close()
rb, _ := io.ReadAll(io.LimitReader(resp.Body, 4096))
@@ -188,8 +189,46 @@ func (s *Sink) Send(ctx context.Context, d delivery.Sendable) error {
// sendMessageURL — the bot API path. the token is in the URL path
// (https://api.telegram.org/bot<token>/sendMessage); telegram does not accept
// it anywhere else. the URL is built per-send from the resolved base — the
// token never leaves the sink, no logging.
// it anywhere else. the URL is built per-send from the resolved base and never
// stored, but it does end up inside transport errors — see redact.
func (s *Sink) sendMessageURL() string {
return s.base + "/bot" + s.cfg.BotToken + "/sendMessage"
}
// tokenPlaceholder — what a redacted token reads as in an error. Recognisable
// on sight, so nobody reads a redacted URL as a malformed one.
const tokenPlaceholder = "<redacted>"
// redact strips the bot token out of a transport error before it becomes a
// returned error, and from there a log line.
//
// This is not hypothetical. net/http wraps every transport failure in
// *url.Error, whose Error() prints the full request URL, and the token is IN
// that URL because telegram accepts it nowhere else. On 2026-08-01 homesrv
// could not reach api.telegram.org, so the retry wrote the whole bot token
// into the daemon log once a minute for as long as the network stayed down.
// The token lives in deploy/telegram.env specifically to stay out of the repo;
// putting it in `docker compose logs` undoes that.
//
// The structural case rewrites url.Error.URL and keeps the error's type, so
// callers matching on *url.Error still work. Anything else falls back to
// scrubbing the rendered message, which loses the type but cannot leak.
//
// There is deliberately no minimum-length guard. A one-character token would
// make this replace every occurrence of that character in the message, which
// is ugly; leaking a short token is worse. New already refuses an empty one.
func (s *Sink) redact(err error) error {
if err == nil {
return err
}
var ue *url.Error
if errors.As(err, &ue) {
clean := *ue
clean.URL = strings.ReplaceAll(clean.URL, s.cfg.BotToken, tokenPlaceholder)
err = &clean
}
if msg := strings.ReplaceAll(err.Error(), s.cfg.BotToken, tokenPlaceholder); msg != err.Error() {
return errors.New(msg)
}
return err
}
@@ -3,6 +3,7 @@ package telegramsink
import (
"context"
"encoding/json"
"errors"
"io"
"net/http"
"net/http/httptest"
@@ -326,6 +327,80 @@ func TestSendConnectionRefusedReturnsError(t *testing.T) {
}
}
// --------------------------- token redaction --------------------------------
// realToken — shaped like a real BotFather token, unlike sinkCfg's "123:abc".
// The redaction tests need something long and distinctive enough that finding
// it in an error message is unambiguous.
const realToken = "7556767480:AAFh0vLU9sg8l7DwXU9y-VZQquSJKW3lsVQ"
// A transport error renders the whole request URL, and telegram accepts the
// token nowhere but the URL path. On 2026-08-01 that put the live bot token in
// `docker compose logs mavend` once a minute while egress was down.
func TestSendTransportErrorRedactsToken(t *testing.T) {
cases := []struct {
name string
run func(*Sink) error
}{
{"connection refused", func(s *Sink) error {
return s.Send(context.Background(), nudgeSendable(loop.Sev4, "down"))
}},
{"context cancel", func(s *Sink) error {
ctx, cancel := context.WithTimeout(context.Background(), 1*time.Nanosecond)
defer cancel()
return s.Send(ctx, nudgeSendable(loop.Sev4, "down"))
}},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
cfg := sinkCfg("http://127.0.0.1:1")
cfg.BotToken = realToken
cfg.Timeout = time.Second
sink, _ := New(cfg)
err := tc.run(sink)
if err == nil {
t.Fatal("want a transport error")
}
if strings.Contains(err.Error(), realToken) {
t.Fatalf("token leaked into error: %v", err)
}
if !strings.Contains(err.Error(), tokenPlaceholder) {
t.Fatalf("want %q in the redacted error, got: %v", tokenPlaceholder, err)
}
})
}
}
// The structural branch keeps the error's type so errors.As still matches.
func TestRedactPreservesURLErrorType(t *testing.T) {
cfg := sinkCfg("http://127.0.0.1:1")
cfg.BotToken = realToken
cfg.Timeout = time.Second
sink, _ := New(cfg)
err := sink.Send(context.Background(), nudgeSendable(loop.Sev4, "down"))
var ue *url.Error
if !errors.As(err, &ue) {
t.Fatalf("want *url.Error to survive redaction, got %T: %v", err, err)
}
if strings.Contains(ue.URL, realToken) {
t.Fatalf("token left in url.Error.URL: %s", ue.URL)
}
}
// Nothing to redact must not disturb the error.
func TestRedactLeavesCleanErrorsAlone(t *testing.T) {
sink, _ := New(sinkCfg("http://127.0.0.1:1"))
in := errors.New("dial tcp: no route to host")
if got := sink.redact(in); got != in {
t.Fatalf("want the same error back, got %v", got)
}
if sink.redact(nil) != nil {
t.Fatal("want nil for nil")
}
}
// ----------------------------- proxy seam -----------------------------------
func TestProxyWiredIntoTransport(t *testing.T) {
+1 -1
View File
@@ -19,7 +19,7 @@
// 3. if no live session exists, Send returns voice.ErrNoSession
// (wrapped). The daemon logs the partial dispatch; an OPEN deferred
// question is whether the dispatcher should reroute to away-channels
// instead of returning partial — listed in PROGRESS.md.
// instead of returning partial.
//
// Import direction: voicesink imports internal/tts (synth seam) and
// internal/voice (Sessions registry). Both are siblings of delivery; the
+1 -2
View File
@@ -1,5 +1,4 @@
// Package event is the unified intake envelope (Vikunja #283,
// 20-07-2026-BACKLOG.md item 1).
// Package event is the unified intake envelope (Vikunja #283).
//
// # The problem it solves
//
+27
View File
@@ -640,3 +640,30 @@ func TestSwapModel_Hook(t *testing.T) {
t.Fatalf("swap to a non-allowlisted path = %v; want ErrForbidden", err)
}
}
// The eleven-day bug: Close shut the listener but not the accepted conns, so
// an idle client left serveConn parked in readFrame and wg.Wait never
// returned. mavend deadlocked before `defer st.Close()` could re-encrypt the
// database, and every write since the last clean stop was rolled back on the
// next boot. The client here stays connected and idle on purpose.
func TestCloseReturnsWithAnIdleClientConnected(t *testing.T) {
_, srv, cli, _ := newServerWithStore(t)
// Prove the conn is live and then leave it alone — no cli.Close().
if _, err := cli.RecentFacts(context.Background(), 1); err != nil {
t.Fatalf("warm-up call: %v", err)
}
done := make(chan error, 1)
go func() { done <- srv.Close() }()
select {
case err := <-done:
if err != nil {
t.Fatalf("close: %v", err)
}
// Deliberately shorter than closeGrace: the grace timer is the backstop,
// not the mechanism. Closing the conns is what makes readFrame return, and
// if that regresses this waits out the full grace and fails here.
case <-time.After(closeGrace / 2):
t.Fatal("Close blocked on an idle connection — the shutdown deadlock is back")
}
}
+87 -1
View File
@@ -6,6 +6,7 @@ import (
"encoding/json"
"errors"
"fmt"
"log"
"net"
"os"
"sync"
@@ -451,6 +452,19 @@ type Server struct {
done chan struct{}
accept sync.Mutex // guards wg.Add vs Close's wg.Wait sequence
// conns — every accepted connection still being served. Close needs these
// because closing the listener does nothing to a connection already
// accepted: serveConn is parked in readFrame waiting for a peer that may
// never say anything again, and the wg.Wait below would block forever.
//
// This was not theoretical. mavweb, mavpoll, mavcaldav and mavmaild all
// hold a long-lived connection open, so on 2026-08-01 mavend deadlocked on
// every single shutdown, never returned from run(), and never reached the
// `defer st.Close()` that seals the database. The deployed ciphertext was
// eleven days stale before anyone noticed.
connMu sync.Mutex
conns map[net.Conn]struct{}
// Check — optional authorization hook. dispatch runs it BEFORE method
// dispatch, with the raw params, so the auth layer can make verdicts
// that depend on the call's shape (e.g. WriteFact's source). A non-nil
@@ -648,8 +662,10 @@ func (s *Server) Serve() error {
s.accept.Lock()
s.wg.Add(1)
s.accept.Unlock()
s.trackConn(c)
go func(c net.Conn) {
defer s.wg.Done()
defer s.untrackConn(c)
defer c.Close()
s.serveConn(c)
}(c)
@@ -1231,18 +1247,88 @@ func (s *Server) Close() error {
close(s.done)
}
err := s.ln.Close()
// Closing the listener stops new connections; it does nothing to the ones
// already accepted. Close those too, or every serveConn parked in readFrame
// waits on a peer that has no reason to hang up and the Wait below never
// returns. See the comment on Server.conns.
s.closeConns()
// Under accept lock: after the listener closes, no new Accept can complete,
// so no new wg.Add will be called. The Wait is safe to observe the wg
// counter because any in-flight Accept that already got a conn either
// already called wg.Add (before releasing the lock) or will see the closed
// listener error and not call wg.Add at all.
s.accept.Lock()
s.wg.Wait()
waited := waitTimeout(&s.wg, closeGrace)
s.accept.Unlock()
if !waited {
// Bounded on purpose. A dispatch can be mid-call into the resident
// model, which has its own timeout measured in tens of seconds, and the
// caller of Close is on its way to sealing the database with whatever
// grace the supervisor allows. Abandoning one in-flight RPC is cheap;
// missing the seal costs every write since the last clean shutdown.
log.Printf("ipc: %d connection(s) still busy after %s, closing anyway", s.liveConns(), closeGrace)
}
_ = os.Remove(s.path)
return err
}
// closeGrace — how long Close waits for in-flight dispatches to finish before
// giving up on them. Well inside the ten seconds docker allows by default, so
// the caller still has time to seal.
const closeGrace = 3 * time.Second
func (s *Server) trackConn(c net.Conn) {
s.connMu.Lock()
defer s.connMu.Unlock()
if s.conns == nil {
s.conns = make(map[net.Conn]struct{})
}
s.conns[c] = struct{}{}
}
func (s *Server) untrackConn(c net.Conn) {
s.connMu.Lock()
defer s.connMu.Unlock()
delete(s.conns, c)
}
func (s *Server) liveConns() int {
s.connMu.Lock()
defer s.connMu.Unlock()
return len(s.conns)
}
// closeConns closes every live connection, which is what unblocks the reads.
// The serveConn goroutines see the resulting error and return.
func (s *Server) closeConns() {
s.connMu.Lock()
live := make([]net.Conn, 0, len(s.conns))
for c := range s.conns {
live = append(live, c)
}
s.connMu.Unlock()
for _, c := range live {
_ = c.Close()
}
}
// waitTimeout waits on wg for at most d, reporting whether it finished. The
// abandoned goroutines are still holding a wg count, so nothing may reuse the
// WaitGroup afterwards — Close is terminal, which is what makes this safe.
func waitTimeout(wg *sync.WaitGroup, d time.Duration) bool {
done := make(chan struct{})
go func() {
wg.Wait()
close(done)
}()
select {
case <-done:
return true
case <-time.After(d):
return false
}
}
// Path returns the filesystem path of the listening socket.
func (s *Server) Path() string { return s.path }
+51 -2
View File
@@ -4,8 +4,9 @@
// article snippet beats letting her recall. Nothing here talks to the internet;
// the Kiwix server is on the same box.
//
// This is search only. Full articles are ~100KB of HTML, far too big for a 4096
// token context, so the unit of context is the search snippet (~500 chars).
// Search finds the article; Article reads it. The snippet a search returns is
// NOT usable context on its own — see the comment on Article — so the unit of
// context is the head of the article, truncated to fit a 4096 token window.
package kiwix
import (
@@ -20,6 +21,8 @@ import (
"strconv"
"strings"
"time"
"github.com/kami/maven/internal/crawl"
)
// Result is one search hit.
@@ -73,6 +76,52 @@ func (c *Client) Search(ctx context.Context, pattern, book string, limit int) ([
return ParseSearchRSS(resp.Body)
}
// articleMaxBytes — how much of an article HTML document is read before the
// rest is discarded. A maxi Wikipedia page is around 100KB; 512KB is slack for
// the long ones and a hard stop against a ZIM entry that is really a binary.
const articleMaxBytes = 512 << 10
// Article fetches one article by the Path a search hit carries and returns it
// as extracted plain text, capped at maxRunes (0 ⇒ crawl.DefaultMaxRunes).
//
// This exists because the search snippet is not usable context. Kiwix builds
// the snippet from wherever the keyword matched, and on a Wikipedia ZIM that is
// routinely the "see also" navigation box at the foot of the page: a search for
// "photosynthesis" comes back with "Ecological economics Ecological footprint
// Ecological forecasting …" and a model handed that writes nothing worth
// hearing. The lead paragraphs are at the top of the document, so truncating an
// article from the front gets the definition the snippet was supposed to be.
//
// Nothing here reaches the internet: the path is resolved against the same
// server the search went to.
func (c *Client) Article(ctx context.Context, path string, maxRunes int) (crawl.Page, error) {
path = strings.TrimSpace(path)
if path == "" {
return crawl.Page{}, fmt.Errorf("kiwix article: empty path")
}
if !strings.HasPrefix(path, "/") {
path = "/" + path
}
u := c.base + path
req, err := http.NewRequestWithContext(ctx, http.MethodGet, u, nil)
if err != nil {
return crawl.Page{}, err
}
resp, err := c.http.Do(req)
if err != nil {
return crawl.Page{}, err
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return crawl.Page{}, fmt.Errorf("kiwix article %s: http %d", path, resp.StatusCode)
}
body, err := io.ReadAll(io.LimitReader(resp.Body, articleMaxBytes))
if err != nil {
return crawl.Page{}, err
}
return crawl.Extract(u, body, maxRunes), nil
}
// rss mirrors just the bits of the RSS 2.0 reply we use.
type rss struct {
Items []struct {
+1 -1
View File
@@ -31,7 +31,7 @@ type Completer interface {
// or reply in Russian.
//
// Why the JSON wrapper: this model always thinks out loud and this llama-server
// build ignores the thinking switch (see ROUTING-EVAL-31-07-2026.md). A bare
// build ignores the thinking switch (see docs/evals/2026-07-31-routing.md). A bare
// word-list grammar just captured the reasoning — every case came back as
// "Let me analyze this request carefully". Demanding JSON, like routeGrammar and
// responseGrammar already do, gives the reasoning nowhere to go.
+5 -5
View File
@@ -10,12 +10,12 @@ import (
// Tests for the universal restraint gate.
//
// DESIGN.md § Trigger model: "the gate is universal, applied by the loop, never
// docs/design.md § Trigger model: "the gate is universal, applied by the loop, never
// per-rule — quiet-hours, presence, cooldown, snooze, calendar-busy all live in
// one fires()." These tests pin the CONSERVATIVE side of that: the cases where
// Maven must stay quiet. They exist so nobody loosens the gate by accident.
//
// Where the code does not yet do what DESIGN.md promises, the test is written to
// Where the code does not yet do what docs/design.md promises, the test is written to
// show the gap and then skipped, with the file and line to fix. Behaviour is not
// changed to make a test pass.
@@ -54,7 +54,7 @@ func TestGateQuietHoursSuppressesCareOnly(t *testing.T) {
// ---------------------------- presence ---------------------------------------
// DESIGN.md § Delivery: "sev <= 2 drops on away, sev >= 3 holds: a missed water
// docs/design.md § Delivery: "sev <= 2 drops on away, sev >= 3 holds: a missed water
// nudge is noise, a missed backup failure isn't."
func TestGateAwayDropsCareHoldsOps(t *testing.T) {
cases := []struct {
@@ -250,7 +250,7 @@ func TestTickOrderOfRulesDoesNotMatter(t *testing.T) {
// ---------------------------- reminders bypass the gate ----------------------
// DESIGN.md § User reminders: "bypasses the restraint gate — 'wake me 7' fires
// docs/design.md § User reminders: "bypasses the restraint gate — 'wake me 7' fires
// in quiet hours; that's the point." Every suppressor set at once, and the
// reminder still comes through.
func TestRemindersBypassEverySuppressor(t *testing.T) {
@@ -269,7 +269,7 @@ func TestRemindersBypassEverySuppressor(t *testing.T) {
}
}
// GAP — DESIGN.md § User reminders ends "Snooze still applies." RemindDecisions
// GAP — docs/design.md § User reminders ends "Snooze still applies." RemindDecisions
// passes every due reminder straight through with no snooze check, so a snoozed
// reminder fires anyway. The test below is what the contract asks for.
func TestRemindersStillHonourSnooze(t *testing.T) {
+35 -1
View File
@@ -1,6 +1,9 @@
package loop
import "time"
import (
"strings"
"time"
)
// Rule — a proactive rule. Rules are CODE, not a DSL config — until ~30 rules
// and you feel the pain (per spec). A Rule has a name (ids it in nudges.outcome
@@ -137,6 +140,37 @@ func NetdataCriticalRule() Rule {
}
}
// RulesExcept returns DefaultRules minus the named ones (config's
// `disabled_rules`). Config subtracts from the canonical set; it never adds to
// it and never reorders it, so the "rules are code" line above still holds.
//
// Names are matched exactly and an unknown one is ignored, on purpose: a
// config that still lists a rule someone deleted must not stop the daemon from
// booting. The logging of what was actually dropped belongs to the caller,
// which knows whether anyone is listening.
func RulesExcept(disabled []string) ([]Rule, []string) {
all := DefaultRules()
if len(disabled) == 0 {
return all, nil
}
off := make(map[string]bool, len(disabled))
for _, n := range disabled {
if n = strings.TrimSpace(n); n != "" {
off[n] = true
}
}
out := make([]Rule, 0, len(all))
var dropped []string
for _, r := range all {
if off[r.Name] {
dropped = append(dropped, r.Name)
continue
}
out = append(out, r)
}
return out, dropped
}
// DefaultRules — the canonical set the daemon wires. Add more as code, not config.
// Order here is NOT load-bearing — the loop picks max severity, ties broken by
// (severity desc, name asc) for deterministic output.
+46 -4
View File
@@ -18,7 +18,7 @@ import (
// - does it stay quiet when it should?
// - is it silent when the key it needs has no data at all?
//
// The last one is load-bearing. DESIGN.md: "since(key)==null → don't fire.
// The last one is load-bearing. docs/design.md: "since(key)==null → don't fire.
// Silence on no-data is 'shuts up when uncertain'."
// stateWith builds a snapshot at refTime() holding just the given facts.
@@ -179,7 +179,7 @@ func TestCareRulePredicates(t *testing.T) {
// ---------------------------- ops rules --------------------------------------
// The two ops rules match on a value AND on which poller wrote it. DESIGN.md:
// The two ops rules match on a value AND on which poller wrote it. docs/design.md:
// "a compromised poller must not be able to forge a trigger." Half of this
// table is forgery attempts; all of them must be refused.
func TestOpsRulePredicates(t *testing.T) {
@@ -331,7 +331,7 @@ func TestNoDefaultRuleFiresOnEmptyState(t *testing.T) {
}
}
// Severities are the delivery contract (DESIGN.md § Delivery / channel
// Severities are the delivery contract (docs/design.md § Delivery / channel
// routing): care is sev1-2 and drops when away, ops is sev3-4 and holds. Pin
// them so a change to a rule's insistence has to be deliberate.
func TestDefaultRuleSeverities(t *testing.T) {
@@ -356,7 +356,7 @@ func TestDefaultRuleSeverities(t *testing.T) {
}
}
// Cooldown bounds keep the feedback tuner honest — DESIGN.md wants
// Cooldown bounds keep the feedback tuner honest — docs/design.md wants
// `cooldown in [min,max]` "so a weird week can't mutate Maven silent or
// stalker". A base outside its own envelope would make that meaningless.
func TestDefaultRuleCooldownsAreBounded(t *testing.T) {
@@ -392,3 +392,45 @@ func TestPredicatesArePure(t *testing.T) {
}
}
}
func ruleNames(rs []Rule) []string {
out := make([]string, len(rs))
for i, r := range rs {
out[i] = r.Name
}
return out
}
func TestRulesExceptDropsOnlyTheNamed(t *testing.T) {
rules, dropped := RulesExcept([]string{"service_down"})
if len(rules) != len(DefaultRules())-1 {
t.Fatalf("got %v", ruleNames(rules))
}
for _, r := range rules {
if r.Name == "service_down" {
t.Error("a disabled rule was still wired")
}
}
if len(dropped) != 1 || dropped[0] != "service_down" {
t.Errorf("dropped = %v", dropped)
}
}
// A config that names a rule nobody wrote must not stop the daemon booting,
// and must not quietly drop a real rule alongside it.
func TestRulesExceptIgnoresUnknownNames(t *testing.T) {
rules, dropped := RulesExcept([]string{"no_such_rule", " ", ""})
if len(rules) != len(DefaultRules()) {
t.Errorf("an unknown name removed something: %v", ruleNames(rules))
}
if len(dropped) != 0 {
t.Errorf("dropped = %v, want nothing", dropped)
}
}
func TestRulesExceptEmptyKeepsEverything(t *testing.T) {
rules, dropped := RulesExcept(nil)
if len(rules) != len(DefaultRules()) || dropped != nil {
t.Errorf("rules = %v, dropped = %v", ruleNames(rules), dropped)
}
}
+1 -1
View File
@@ -41,7 +41,7 @@
"tags": ["preference", "homelab", "paraphrase", "hard"],
"query": "когда запускать резервное копирование",
"want": "n1",
"note": "The DESIGN.md preference-seam example, phrased as the operator would ask it later.",
"note": "The docs/design.md preference-seam example, phrased as the operator would ask it later.",
"notes": [
{"id": "n1", "text": "бэкапы лучше делать ночью в три часа", "kind": "note"},
{"id": "n2", "text": "обновления ставлю по субботам", "kind": "note"},
+1 -1
View File
@@ -1,5 +1,5 @@
// Package morning is maven's morning routine engine — item #3 off the
// 2026-07-20 backlog (see Maven/20-07-2026-BACKLOG.md).
// 2026-07-20 backlog (Vikunja #280).
//
// A Routine is NOT four independent reminder timers. It's a checklist for a
// daily window: several Items, each evidenced by a fact key, completed in
+3 -3
View File
@@ -15,7 +15,7 @@ const (
CheckLang = "lang" // the operator's language, not the prompt's
CheckLength = "length" // a nudge is one sentence, not a paragraph
CheckFeminine = "feminine" // her self-reference is feminine (hard constraint)
CheckCringe = "cringe" // DESIGN.md § Non-goals, "not a relationship"
CheckCringe = "cringe" // docs/design.md § Non-goals, "not a relationship"
CheckOnTopic = "ontopic" // says the thing the rule is about
// CheckHisGender — the other half of the persona rule: SHE is feminine, HE
@@ -105,7 +105,7 @@ func checkLength(body string) Result {
// --- feminine self-reference ---------------------------------------------
//
// The hard constraint (CLAUDE.md, DESIGN.md § Identity): Maven's Russian
// The hard constraint (CLAUDE.md, docs/design.md § Identity): Maven's Russian
// self-reference is feminine. The operator is male, so second-person forms
// addressed to him are MASCULINE and must not be flagged — "ты не пил воду" is
// correct, "я напомнил" is not. Both directions matter, which is why this is a
@@ -497,7 +497,7 @@ func isLatinWord(w string) bool {
// --- the cringe checks ---------------------------------------------------
//
// "Think Jarvis without the cringe part". DESIGN.md § Non-goals: "Not a
// "Think Jarvis without the cringe part". docs/design.md § Non-goals: "Not a
// relationship — mom-tone is a function that makes nudges land, not emotional
// company. Names the drift a warm small model falls into." Each pattern below
// is one shape of that drift. They are deliberately specific: a check that
+2 -2
View File
@@ -11,7 +11,7 @@
// length test that a human can read and disagree with. A score here is a claim
// about measurable properties, not about whether a sentence is good.
//
// DESIGN.md § "Rules decide, LLM phrases" is why there is no send/veto signal
// docs/design.md § "Rules decide, LLM phrases" is why there is no send/veto signal
// anywhere in this package: the rule already decided she speaks. The phraser
// only words it, so a nudge the model refuses to write is a failure, never a
// legitimate outcome.
@@ -38,7 +38,7 @@ var fixtureJSON []byte
const SchemaVersion = 1
// Case — one nudge situation, as a real tick would present it. The fields are
// the (rule, severity, context) input DESIGN.md names, flattened to JSON.
// the (rule, severity, context) input docs/design.md names, flattened to JSON.
//
// WantAny is the on-topic contract: at least one of these lowercased fragments
// must appear in the message. A water nudge that never mentions water is a
+102
View File
@@ -0,0 +1,102 @@
package phraser
import (
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"strings"
"testing"
)
// promptSpy records the system and user strings of every request, which is
// where the evidence-first discipline either exists or does not.
type promptSpy struct {
srv *httptest.Server
system []string
user []string
}
func newPromptSpy(t *testing.T) *promptSpy {
t.Helper()
s := &promptSpy{}
s.srv = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
var req chatReq
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
t.Errorf("spy: decode request: %v", err)
}
for _, m := range req.Messages {
switch m.Role {
case "system":
s.system = append(s.system, m.Content)
case "user":
s.user = append(s.user, m.Content)
}
}
w.Header().Set("Content-Type", "application/json")
w.Write([]byte(`{"choices":[{"message":{"content":"{\"response\": \"вот что я нашла: два литра\", \"mood\": \"neutral\"}"}}]}`))
}))
t.Cleanup(s.srv.Close)
return s
}
func TestEvidenceReachesTheModelAsNumberedSources(t *testing.T) {
spy := newPromptSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{})
if _, err := p.PhraseQuery(context.Background(), "сколько воды я выпил",
[]string{"выпил два литра", "бутылка на 0.7"}); err != nil {
t.Fatal(err)
}
if len(spy.user) != 1 {
t.Fatalf("got %d user messages, want 1", len(spy.user))
}
for _, want := range []string{"[1] выпил два литра", "[2] бутылка на 0.7", "Источники:"} {
if !strings.Contains(spy.user[0], want) {
t.Errorf("user prompt is missing %q:\n%s", want, spy.user[0])
}
}
}
// The system prompt is the whole fix for the Левитан fabrication: answer from
// the sources, say so when they do not answer, add nothing from memory.
func TestEvidencePromptForbidsAnsweringFromMemory(t *testing.T) {
spy := newPromptSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{})
if _, err := p.PhraseQuery(context.Background(), "кто написал войну и мир",
[]string{"Лев Толстой, роман 1869 года"}); err != nil {
t.Fatal(err)
}
sys := spy.system[0]
for _, want := range []string{"ТОЛЬКО по ним", "Если ответа в них нет"} {
if !strings.Contains(sys, want) {
t.Errorf("system prompt is missing %q:\n%s", want, sys)
}
}
if strings.Contains(sys, "заметк") {
t.Errorf("system prompt still calls every source a note:\n%s", sys)
}
}
// A source that trimmed away is not a source. Handing the evidence branch an
// empty list is the one prompt that reliably makes a small model invent.
func TestBlankSourcesTakeTheKnowledgeBranch(t *testing.T) {
spy := newPromptSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{})
if _, err := p.PhraseQuery(context.Background(), "что я записывал", []string{"", " "}); err != nil {
t.Fatal(err)
}
if strings.Contains(spy.user[0], "Источники:") {
t.Errorf("blank sources still took the evidence branch:\n%s", spy.user[0])
}
}
func TestNonEmptyDoesNotMutateTheCallersSlice(t *testing.T) {
in := []string{" один ", "", "два"}
got := nonEmpty(in)
if want := []string{"один", "два"}; len(got) != 2 || got[0] != want[0] || got[1] != want[1] {
t.Errorf("nonEmpty = %q, want %q", got, want)
}
if in[0] != " один " {
t.Errorf("caller's slice was mutated: %q", in)
}
}
+92 -13
View File
@@ -357,6 +357,11 @@ func (p *LLMPhraser) PhraseNudge(ctx context.Context, c loop.Candidate) (deliver
// compose a natural answer. Falls back to "вот что я нашла: <notes>" on any
// LLM error — better to give the raw data than silence.
func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []string) (string, error) {
// Blank sources are no sources. A caller that hands over one empty string —
// a page that fetched to nothing, a snippet trimmed away — used to take the
// evidence branch and be told to answer from an empty list, which is the one
// prompt guaranteed to make a small model fill the gap from memory.
notes = nonEmpty(notes)
if len(notes) == 0 {
// General knowledge — no notes to ground the answer. The system
// prompt is the single tested source in router.KnowledgePrompt.
@@ -376,13 +381,10 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
}
return resp, nil
}
if len(notes) == 1 {
notes[0] = strings.TrimSpace(notes[0])
}
sys := p.querySystemPrompt()
prompt := fmt.Sprintf(
`Он спрашивает: "%s". В твоих заметках по этому вопросу написано: "%s". Ответь ему коротко и своими словами. Если в заметках ответа нет — так и скажи.`,
utterance, strings.Join(notes, `"; "`),
"Он спрашивает: \"%s\"\n\nИсточники:\n%s\nОтветь ему коротко и своими словами, опираясь только на эти источники. Если ответа в них нет — так и скажи.",
utterance, evidenceBlock(notes),
)
resp, err := p.chatWithSystem(ctx, sys, prompt, 768)
text, _, perr := parseResponseMood(resp)
@@ -444,9 +446,32 @@ func (p *LLMPhraser) PhraseChat(ctx context.Context, utterance string, history [
func chatSystemPrompt(block func() string) string {
// No self-introduction here: the persona block prepended one line above
// already says who she is, same as router.KnowledgePrompt.
base := `Ты разговариваешь с хозяином. О себе говоришь в женском роде ("я подумала", "я рада"). Он мужчина: обращайся к нему на "ты", в мужском роде ("ты сказал", "ты забыл"). Никогда не "вы"/"ваш" и никогда "он"/"его" ты говоришь ему, а не о нём.
//
// The grammar examples used to be full clauses: ("я подумала", "я рада")
// for her, ("ты сказал", "ты забыл") for him. A 1.7B copies those rather
// than generalising from them. Observed on the box 2026-08-01: all three
// chat replies in one session opened with "Я подумала, что ...", and one
// ended "...немного тревожусь. ты сказал" — the second example pasted onto
// the end of a finished sentence, which reads as a truncation but is not.
//
// So: contrastive pairs instead of usable openers. "рада, не рад" states
// the rule as a correction, and short predicatives do not hand the model a
// sentence frame to start with. The him-examples are gone entirely; the
// "ты" instruction carries that on its own and those two produced the
// worst output. The last line says outright not to echo the instructions,
// because a small model will otherwise treat any quoted string as licence.
//
// Amended the same day: with the openers gone the tic went with them, but
// "не забыл ли я" appeared — masculine, about herself. The old "я подумала"
// had been suppressing that by accident, being a feminine past tense the
// model could copy. Two short predicatives are not enough signal on their
// own, so the rule is now stated as morphology (-ла) rather than as a pair
// of words. A suffix rule generalises where an example only gets copied.
base := `Ты разговариваешь с хозяином.
Отвечай по-русски, коротко: одна-три фразы, живым языком. Ты доброжелательная, тебе интересно, но чувства не изображай.
О себе в женском роде: "рада", не "рад"; "поняла", не "понял". Все свои глаголы в прошедшем времени оканчивай на -ла: сделала, забыла, записала, подумала. Он мужчина: обращайся к нему на "ты", в мужском роде. Никогда не "вы"/"ваш" и никогда "он"/"его" ты говоришь ему, а не о нём.
Отвечай по-русски, коротко: одна-три фразы, живым языком. Ты доброжелательная, тебе интересно, но чувства не изображай. Не повторяй формулировки из этой инструкции отвечай своими словами.
Отвечай ТОЛЬКО одним объектом JSON: {"response": "...", "mood": "neutral"}. В "response" твой ответ. В "mood" ровно одно из: neutral, happy, thinking, tired, confused.`
return persona.Prepend(block, base)
@@ -580,6 +605,11 @@ string ::= "\"" ([^"\\] | "\\" ["\\/bfnrt]){0,1000} "\""
ws ::= [ \t\n]*
`
// ResponseGrammar exposes responseGrammar to the other callers that emit the
// same {"response","mood"} contract — cmd/mavend's reactive replier, which is
// parsed by the same two fields. One definition, so the two cannot drift.
const ResponseGrammar = responseGrammar
// grammar returns the GBNF to attach to a phrasing request, or "" when the
// operator turned it off.
func (p *LLMPhraser) grammar() string {
@@ -662,7 +692,7 @@ func (p *LLMPhraser) chatWithSystem(ctx context.Context, system, user string, ma
// Written as filled-in examples, not as a schema with "..." in it. A 0.8B
// copies whatever sits in the response slot, so a literal placeholder there
// teaches it to answer with the placeholder. Measured: 7/15 nudges came back
// as "..." before this. See PHRASING-EVAL-31-07-2026.md.
// as "..." before this. See docs/evals/2026-07-31-phrasing.md.
//
// Russian only, feminine self-reference, second person masculine (the owner is
// a man). She talks TO him, informally, singular — never "вы", never "он".
@@ -697,15 +727,64 @@ func (p *LLMPhraser) systemPrompt() string {
return persona.Prepend(p.cfg.ContextBlock, nudgeSystem)
}
// querySystemPrompt returns the system prompt for PhraseQuery (notes + general
// knowledge). Prepends the configured persona when set.
// querySystemPrompt returns the system prompt for the evidence branch of
// PhraseQuery. Prepends the configured persona when set.
//
// Evidence-first, and that is the whole point of this prompt. Every source that
// reaches PhraseQuery with something in hand — his notes, a stored fact, a page,
// a live search, a ZIM article — arrives as numbered sources, and the model's
// job here is to READ them, not to recall. A 1.7B asked a world question
// answers from its weights with total confidence and no signal that it is
// guessing; that is how "Война и мир" got Левитан as its author. The rule that
// prevents it is stated three ways, because one way did not hold: answer from
// the sources, say plainly when they do not answer, add nothing of your own.
//
// It no longer says "заметки". The sources are not always his notes, and
// calling a Wikipedia paragraph his note both misleads him and licenses the
// model to blur where an answer came from.
//
// No self-introduction here: the persona block prepended one line above already
// says who she is, same as router.KnowledgePrompt.
//
// The opener is deliberate and stays: the fixed prefix is what marks the answer
// as a lookup rather than as something she knows. The grammar examples are not
// deliberate — same defect chatSystemPrompt had, where a 1.7B copies a quoted
// word instead of generalising from it. Stated as morphology instead.
func (p *LLMPhraser) querySystemPrompt() string {
// No self-introduction here: the persona block prepended one line above
// already says who she is, same as router.KnowledgePrompt.
base := "Ты отвечаешь ему по своим заметкам. Отвечай по-русски, коротко и своими словами, начинай с \"вот что я нашла: \". О себе — в женском роде (\"нашла\", \"записала\"). Он мужчина, обращайся к нему на \"ты\". Respond ONLY with valid JSON: {\"response\": \"...\", \"mood\": \"neutral\"}."
base := "Ты отвечаешь ему по источникам, которые тебе дали. Отвечай ТОЛЬКО по ним: всё, что ты говоришь, должно быть написано в источниках. " +
"Если ответа в них нет — так и скажи и на этом остановись; не добавляй ничего из своих знаний и не догадывайся. " +
"Не приплетай прошлые реплики разговора. " +
"Отвечай по-русски, коротко и своими словами, начинай с \"вот что я нашла: \". О себе — в женском роде, глаголы в прошедшем времени с окончанием -ла. Он мужчина, обращайся к нему на \"ты\". Respond ONLY with valid JSON: {\"response\": \"...\", \"mood\": \"neutral\"}."
return persona.Prepend(p.cfg.ContextBlock, base)
}
// evidenceBlock renders the sources for the evidence branch of PhraseQuery.
//
// Numbered lines, one source each, rather than the quoted semicolon-joined
// string this used to build. Two reasons, both measured on small models: a
// numbered list survives being long, where a run-on quoted string blurs into
// one claim the model then merges; and the numbering gives it something to
// answer FROM, which is what makes "этого в источниках нет" reachable at all.
func evidenceBlock(sources []string) string {
var b strings.Builder
for i, s := range sources {
fmt.Fprintf(&b, "[%d] %s\n", i+1, s)
}
return b.String()
}
// nonEmpty drops blank sources and trims the rest, without touching the
// caller's slice.
func nonEmpty(sources []string) []string {
out := make([]string, 0, len(sources))
for _, s := range sources {
if s = strings.TrimSpace(s); s != "" {
out = append(out, s)
}
}
return out
}
// ruleTopics — Russian gloss for each built-in rule name. The rule names are
// English identifiers; a 0.8B asked to nudge about "netdata_critical" writes
// about nothing. The daemon knows what its own rules mean, so it says so.
+8 -19
View File
@@ -1,9 +1,9 @@
// Package phraser is maven's "rules decide, llm phrases" seam — the layer
// that turns a loop decision into the body + summary the delivery module ships.
//
// Per DESIGN.md § Resident language model: the phraser is the resident model
// Per docs/design.md § Resident language model: the phraser is the resident model
// (Qwen3-1.7B — RU continued pretraining plus joint persona/router SFT, not a
// sub-1b prompted-only model as the retired spec claimed; see DESIGN.md
// sub-1b prompted-only model as the retired spec claimed; see docs/design.md
// § Superseded, "small-model phrasing claim"). It takes
// (rule, severity, context) and produces Body (full voice message, local — no
// shoulder-surf concern beyond who's in the room) + Summary (minimal body for
@@ -24,7 +24,6 @@ package phraser
import (
"context"
"encoding/json"
"fmt"
"strings"
"time"
@@ -32,6 +31,7 @@ import (
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/store"
)
// Phraser — the seam the daemon wires. one method per delivery path (nudge
@@ -161,22 +161,11 @@ func phraseNudge(c loop.Candidate) (body, summary string) {
}
}
// extractReminderText — the reminder payload is raw JSON; the router's
// reminder slot extraction owns the shape. the conventional field is "text".
// fall back to the raw payload if it isn't JSON or lacks the field — the user
// said it, it's the user's words.
func extractReminderText(payload string) string {
var m map[string]any
if err := json.Unmarshal([]byte(payload), &m); err == nil {
if t, ok := m["text"].(string); ok && t != "" {
return t
}
if t, ok := m["text"]; ok {
return fmt.Sprintf("%v", t)
}
}
return strings.TrimSpace(payload)
}
// extractReminderText — the reminder payload is raw JSON and store.ReminderText
// owns the unwrapping. It used to be a second copy of that logic here, which is
// how the day plan came to recite a reminder as its literal JSON: the copies
// were never going to be kept in step.
func extractReminderText(payload string) string { return store.ReminderText(payload) }
// humanDur — round a duration to the coarsest sensible unit for speech.
// "4h12m" → "4 hours"; "92m" → "1h32m" → "an hour and a half". keep it simple:
+77
View File
@@ -0,0 +1,77 @@
package router
import (
"context"
"testing"
)
// agendaRouter wires both grammar sets in the order the daemon wires them
// (voicewire.go): the clock rules first, the agenda rules after, so a test
// that passes here is a test of the deployed precedence.
func agendaRouter(t *testing.T) *Router {
t.Helper()
r := newTestRouter(t, 0.0)
r.grammars = append(r.grammars, SystemTimeDateGrammars()...)
r.grammars = append(r.grammars, AgendaQueryGrammars()...)
return r
}
// An agenda question is answered from the calendar, which lives in the query
// chain. Routed to system it reaches replySystem, which has no agenda arm and
// says "пока не умею" — seen on the deployed daemon, 01-08-2026.
func TestAgendaQuestionsRouteToQuery(t *testing.T) {
r := agendaRouter(t)
for _, u := range []string{
"что у меня сегодня",
"что у меня в календаре сегодня",
"что у меня стоит в календаре на послезавтра",
"какие у меня встречи завтра",
"покажи расписание на среду",
} {
d, err := r.Route(context.Background(), u, refNow())
if err != nil {
t.Fatalf("route(%q): %v", u, err)
}
if d.Intent != IntentQuery {
t.Errorf("route(%q) = %s, want query", u, d.Intent)
}
}
}
// The clock rules keep their utterances. They are registered first and the
// agenda patterns do not match them, so both statements have to hold.
func TestAgendaGrammarsLeaveTheClockAlone(t *testing.T) {
r := agendaRouter(t)
for _, u := range []string{
"какой сегодня день",
"какое сегодня число",
"который час",
"сколько сейчас времени",
} {
d, err := r.Route(context.Background(), u, refNow())
if err != nil {
t.Fatalf("route(%q): %v", u, err)
}
if d.Intent != IntentSystem {
t.Errorf("route(%q) = %s, want system", u, d.Intent)
}
}
}
// The agenda pattern is anchored and needs the possessive, so an ordinary
// statement that happens to contain "у меня" is not swallowed.
func TestAgendaGrammarSparesStatements(t *testing.T) {
r := agendaRouter(t)
for _, u := range []string{
"у меня кончилась вода",
"напомни мне завтра позвонить маме",
} {
d, err := r.Route(context.Background(), u, refNow())
if err != nil {
t.Fatalf("route(%q): %v", u, err)
}
if d.Stage == 0 && d.Intent == IntentQuery {
t.Errorf("route(%q) was claimed by the agenda grammar", u)
}
}
}
+3
View File
@@ -233,6 +233,9 @@ func newBaselineRouter(t *testing.T, emb router.Embedder, llmR *router.LLMRouter
}
grammars := router.DefaultGrammars(acts)
grammars = append(grammars, router.SystemTimeDateGrammars()...)
// Same order as buildRouter (voicewire.go). The fixture is only worth
// anything while its grammar set is the daemon's grammar set.
grammars = append(grammars, router.AgendaQueryGrammars()...)
grammars = append(grammars, router.ReminderGrammar())
return router.New(router.Config{
Grammars: grammars,
+1 -1
View File
@@ -39,7 +39,7 @@ import (
// Re-measured with everything else held equal, thinking off scores exactly the
// same, case for case — and a direct probe shows this llama-server build ignores
// enable_thinking / reasoning_budget for this model anyway, so there was nothing
// to turn off. Full write-up in ROUTING-EVAL-31-07-2026.md (Vikunja #376).
// to turn off. Full write-up in docs/evals/2026-07-31-routing.md (Vikunja #376).
func TestLLMRouterBaseline(t *testing.T) {
base := os.Getenv("MAVEN_LLM_URL")
if base == "" {
+10 -3
View File
@@ -1,13 +1,13 @@
// Package router is maven's reactive path — the cascade that turns a free-form
// utterance into a deterministic Decision.
//
// Spec contract (from DESIGN.md § Reactive path — routing):
// Spec contract (from docs/design.md § Reactive path — routing):
//
// - the TARGET design is LLM-as-router: the resident model (Qwen3-1.7B)
// emits GBNF-constrained structured JSON for the route, and the same
// model phrases replies; the embedder is a RAG hint, not a routing gate.
// the classifier/embedder cascade below is the committed default today,
// but it is an interim stopgap (DESIGN.md § Superseded, "classifier-owns-
// but it is an interim stopgap (docs/design.md § Superseded, "classifier-owns-
// the-route") and the known cause of weak RU query handling — not a
// design to extend.
// - a CASCADE, not one decider — layers:
@@ -35,7 +35,7 @@ package router
import "time"
// Intent — the seven save-where labels from DESIGN.md's routing table. The
// Intent — the seven save-where labels from docs/design.md's routing table. The
// discriminator is "does the loop evaluate a predicate against it?":
//
// - act: command now, not stored (function call into the allowlist)
@@ -104,4 +104,11 @@ type Decision struct {
Confidence float64 // 1.0 for stage-0; classifier cosine similarity for 1+
Slots Slots
Clarify bool // stage 3: below threshold — ask, don't guess
// Continued — this decision was rebuilt from the previous turn rather
// than routed, because the utterance was an ellipsis ("а завтра?").
// Handlers use it to know that Slots.Text is the PREVIOUS turn's topic
// and not something the current utterance said. Nothing in the router
// sets it; the daemon's continuation path does.
Continued bool
}
+5 -4
View File
@@ -194,12 +194,13 @@ func (lr *LLMRouter) Route(ctx context.Context, utterance string, now time.Time)
return Decision{}, false, nil
}
d := Decision{Utterance: utterance, Stage: 1, Confidence: llmFullConfidence}
// A single-token utterance is thin evidence: the model had nothing to
// A bare one-word nominal is thin evidence: the model had nothing to
// disambiguate on ("вода" is a fact-or-query coin flip, "бэкап" an
// act-or-report one) and stage 0 would already have won on anything
// that pattern-matches cleanly. Flag it now; router.go's stage-3 gate
// (Router.Route) decides whether that trips Clarify.
if len(strings.Fields(utterance)) <= 1 {
// that pattern-matches cleanly. A greeting or an inflected verb is NOT
// thin, however short — see thinSingleToken. Flag it now; router.go's
// stage-3 gate (Router.Route) decides whether that trips Clarify.
if thinSingleToken(utterance) {
d.Confidence = llmThinConfidence
}
switch Intent(a.Intent) {
+90
View File
@@ -0,0 +1,90 @@
package router
import "strings"
// thinSingleToken — is a one-word utterance thin evidence, or is it a whole
// sentence?
//
// The rule this replaces was `len(strings.Fields(u)) <= 1`, an English
// intuition. It does not transfer: Russian packs a subject, a tense and a
// gender into one word, so "поужинал" is a complete report and "привет" a
// complete greeting, yet both got thinned and came back as "не совсем поняла".
// Meanwhile the case the rule exists for is real — a bare noun like "вода" or
// "бэкап" genuinely does not say fact-vs-query or act-vs-report.
//
// So: still one token, but only thin it when the token is a bare nominal.
// Two escapes, both cheap and both offline:
//
// - a closed lexicon of social and command singles, which are complete by
// definition ("привет", "спасибо", "стоп", "yes");
// - a suffix test for an inflected predicate — past tense, 2nd person,
// reflexive. Verbs carry their own subject, so a verb IS a sentence.
//
// The suffix test is deliberately loose about nouns that happen to end the
// same way ("канал" reads as past tense here). That direction of error only
// costs a clarify we would not have asked for; the other direction — treating
// a real report as thin — is the bug being fixed.
func thinSingleToken(utterance string) bool {
f := strings.Fields(utterance)
if len(f) != 1 {
return false
}
w := strings.ToLower(strings.Trim(f[0], ".,!?;:—-\"'«»()"))
if w == "" {
return false
}
if completeSingles[w] {
return false
}
return !looksInflected(w)
}
// completeSingles — one-word utterances that need no second half. Greetings,
// acknowledgements and the control words a voice loop has to honour instantly.
var completeSingles = map[string]bool{
// ru: social
"привет": true, "здравствуй": true, "здравствуйте": true, "здорово": true,
"пока": true, "прощай": true, "спокойной": true, "спасибо": true,
"благодарю": true, "извини": true, "прости": true, "пожалуйста": true,
"да": true, "нет": true, "ага": true, "угу": true, "ок": true, "окей": true,
"хорошо": true, "ладно": true, "конечно": true, "верно": true, "точно": true,
// ru: control
"стоп": true, "отмена": true, "отбой": true, "хватит": true, "тихо": true,
"повтори": true, "продолжай": true, "помоги": true, "помощь": true,
// en
"hi": true, "hello": true, "hey": true, "bye": true, "goodbye": true,
"thanks": true, "thank": true, "sorry": true, "please": true,
"yes": true, "no": true, "yep": true, "nope": true, "ok": true, "okay": true,
"sure": true, "right": true, "stop": true, "cancel": true, "help": true,
"repeat": true, "continue": true,
}
// inflectedSuffixes — endings that mark a finite or past-tense Russian verb.
// Ordered longest-first is unnecessary (any match wins), but each entry is
// chosen to be long enough that common nouns rarely collide.
var inflectedSuffixes = []string{
// reflexive — strongly verbal whatever precedes it
"ся", "сь",
// past tense
"ал", "ял", "ил", "ел", "ыл", "ул", "ёл", "ала", "яла", "ила", "ела",
"ыла", "ула", "али", "яли", "или", "ели",
// 2nd person singular
"ешь", "ишь", "ёшь",
// 1st/2nd person plural, 3rd person plural
"аем", "яем", "уем", "аете", "ите", "ают", "яют", "уют", "ат", "ят",
}
// looksInflected — does the word carry a verb ending? Short words are exempt:
// a three-letter token is not enough stem to trust a two-letter suffix on
// ("газ" would otherwise never match, but "нос" and "лес" would).
func looksInflected(w string) bool {
if len([]rune(w)) < 5 {
return false
}
for _, s := range inflectedSuffixes {
if strings.HasSuffix(w, s) {
return true
}
}
return false
}
+33
View File
@@ -0,0 +1,33 @@
package router
import "testing"
func TestThinSingleTokenThinsBareNominals(t *testing.T) {
// The case the rule exists for: one noun, no way to tell what was asked.
for _, w := range []string{"вода", "бэкап", "нексус", "почта", "backup"} {
if !thinSingleToken(w) {
t.Errorf("thinSingleToken(%q) = false, want true", w)
}
}
}
func TestThinSingleTokenSparesCompleteUtterances(t *testing.T) {
// Regression: every one of these used to be answered with
// "не совсем поняла — можешь переформулировать?".
for _, w := range []string{
"привет", "Привет!", "спасибо", "да", "нет", "стоп", "hello", "yes",
"поужинал", "проснулась", "устал", "выспался", "договорились",
} {
if thinSingleToken(w) {
t.Errorf("thinSingleToken(%q) = true, want false", w)
}
}
}
func TestThinSingleTokenIgnoresMultiWord(t *testing.T) {
for _, s := range []string{"выпил воды", "что там с бэкапом", ""} {
if thinSingleToken(s) {
t.Errorf("thinSingleToken(%q) = true, want false", s)
}
}
}
+56
View File
@@ -140,6 +140,62 @@ func SystemTimeDateGrammars() []Grammar {
}
}
// AgendaQueryGrammars — stage-0 grammars for "what have I got on" questions,
// routed to IntentQuery so they reach the query chain (queryDayPlan,
// queryCalendar) instead of replySystem.
//
// This exists because the model puts them in IntentSystem. Measured on the
// deployed daemon 01-08-2026: "что у меня сегодня" and "что у меня в календаре
// сегодня" both routed system, and replySystem has no agenda arm, so both
// answered "пока не умею". The fixture has said query since ru-query-019 was
// written ("the clock/date system rule must not swallow it"); the daemon
// disagreed with the fixture and the daemon was wrong.
//
// Routing, not answering. These set the intent and nothing else — which source
// in the query chain claims the turn stays the chain's decision, and a
// question with no date still falls through queryCalendar to recall.
//
// Deliberately not folded into SystemTimeDateGrammars: those exist to send
// utterances TO system, these exist to keep utterances OUT of it, and one
// function returning both would read as a list of clock rules.
func AgendaQueryGrammars() []Grammar {
return []Grammar{
{
// An explicit calendar noun is unambiguous wherever it appears:
// "что в календаре на завтра", "покажи расписание на среду".
Name: "calendar-query",
Pattern: regexp.MustCompile(`(?i)(календар|расписани|повестк)`),
Build: agendaQueryBuild,
},
{
// The agenda phrasing with no calendar noun. Anchored at the start
// and requiring the possessive, so it reads as a question about his
// day: "что у меня сегодня", "что у меня стоит на послезавтра".
// "у меня кончилась вода" is a fact and does not match.
Name: "agenda-query",
// (\s|[?!.]|$) rather than \b: Go's \b is ASCII-only, so it does
// not see a boundary after a Cyrillic letter and the pattern
// silently never fires.
// "во сколько у меня встреча" is the same agenda question with a
// clock word in front, and the clock word is what sent it to
// system (fixture ru-query-013).
Pattern: regexp.MustCompile(`(?i)^\s*(что|чего|какие|сколько|во\s+сколько|когда)\s+у\s+меня(\s|[?!.]|$)`),
Build: agendaQueryBuild,
},
}
}
// agendaQueryBuild — shared Build for the agenda grammars. Confidence 1.0 on
// the intent only: the utterance travels intact and the query chain's own
// matchers decide the rest.
func agendaQueryBuild(m []string) (Decision, bool) {
return Decision{
Stage: 0,
Intent: IntentQuery,
Confidence: 1.0,
}, true
}
// timeQueryBuild — Build for the time-query grammar. Returns ok=false for
// elapsed/duration queries ("сколько времени прошло", "сколько времени
// осталось", "сколько времени до") so they fall through to the classifier.
+31
View File
@@ -268,3 +268,34 @@ func zero(b []byte) {
b[i] = 0
}
}
// SealPlaintext encrypts an existing plaintext sqlite file at plainPath and
// writes the ciphertext to cipherPath, atomically. key must be 32 bytes. The
// plaintext file is left alone: this is a recovery path, and deleting the only
// good copy of the data on the strength of a write that just succeeded is not
// a trade worth making here.
//
// It exists for the case closeAndSeal cannot cover: a daemon that was killed
// rather than shut down, leaving a live working copy in tmpfs and a stale
// ciphertext on disk. mavseal folds the WAL in first, so what arrives here is
// a single complete database.
//
// Nothing else should call this. The normal path is Close, which seals and
// then wipes the plaintext and the key.
func SealPlaintext(plainPath, cipherPath string, key []byte) error {
if len(key) != keyLen {
return ErrKeyLen
}
plain, err := os.ReadFile(plainPath)
if err != nil {
return fmt.Errorf("read working copy: %w", err)
}
blob, err := encrypt(key, plain)
if err != nil {
return fmt.Errorf("encrypt: %w", err)
}
if err := atomicWrite(cipherPath, blob); err != nil {
return fmt.Errorf("seal ciphertext: %w", err)
}
return nil
}
+29
View File
@@ -3,8 +3,10 @@ package store
import (
"context"
"database/sql"
"encoding/json"
"errors"
"fmt"
"strings"
"time"
"github.com/robfig/cron/v3"
@@ -26,6 +28,33 @@ type Reminder struct {
Collapsed []Reminder
}
// Text — what the user actually asked for, out of the raw-JSON payload.
//
// The router's reminder slot extraction owns the payload shape and the
// conventional field is "text". A payload that is not JSON, or that lacks the
// field, is returned as-is: he said it, so they are his words, and showing
// them beats showing nothing.
//
// Here rather than in a caller because there is more than one caller and they
// disagreed. The phraser unwrapped the payload; the day plan did not, so
// "какие у меня планы на сегодня" recited a reminder as the literal string
// {"text":"..."} on the deployed daemon, 01-08-2026.
func (r Reminder) Text() string { return ReminderText(r.Payload) }
// ReminderText — Reminder.Text for callers holding a bare payload string.
func ReminderText(payload string) string {
var m map[string]any
if err := json.Unmarshal([]byte(payload), &m); err == nil {
if t, ok := m["text"].(string); ok && t != "" {
return t
}
if t, ok := m["text"]; ok {
return fmt.Sprintf("%v", t)
}
}
return strings.TrimSpace(payload)
}
// Reminder lifecycle states. Named for the same reason DigestStatus is: a
// caller filtering on the string literal "pending" is one typo away from a
// filter that silently matches nothing.
+2 -2
View File
@@ -18,9 +18,9 @@
// returns a canned string the router + action path operate on); with a
// worker socket configured, it wires Remote.
//
// Per DESIGN.md § Voice pipeline (STT / TTS): whisper.cpp (CGo, Vulkan) in
// Per docs/design.md § Voice pipeline (STT / TTS): whisper.cpp (CGo, Vulkan) in
// cmd/mavsttd is the production stt — the older faster-whisper/vosk picks are
// retired (DESIGN.md § Superseded, "named STT/TTS model picks"). The
// retired (docs/design.md § Superseded, "named STT/TTS model picks"). The
// server-side stt module is the heavy multilingual path; the client's
// wake-word + stage-0 command grammar (cmd/mavwaked) hits the router directly
// and never crosses this seam. Today's Remote + Stub both return plain text
+2 -2
View File
@@ -2,7 +2,7 @@
// store's allowlist, and drafts 'proposed' scaffolds for acts that aren't on
// it yet.
//
// Boundary discipline (DESIGN.md § "Tool registration — drafting is suggest,
// Boundary discipline (docs/design.md § "Tool registration — drafting is suggest,
// enabling is act"):
//
// - The store is the allowlist. Only status='enabled' rows run. A verb not
@@ -26,7 +26,7 @@
// - Destructive tools don't run on first hearing: Exec returns ErrNeedsConfirm
// and the handler runs a confirm turn ("выполнить X? да/нет"); only a
// confirmed re-Exec runs them. A gate assumes a fully-formed action, which
// an enabled+matched act is (DESIGN.md § "Confirmation is not one
// an enabled+matched act is (docs/design.md § "Confirmation is not one
// mechanism").
package tool
+2 -2
View File
@@ -5,9 +5,9 @@
// PCM, headerless per the audio package; the voice sink + reference client
// wrap it in a WAV at the disk edge.
//
// Per DESIGN.md § Voice pipeline (STT / TTS): piper is the production tts
// Per docs/design.md § Voice pipeline (STT / TTS): piper is the production tts
// (subprocess + espeak-ng, CPU-only on the ryzen box, driven by cmd/mavttsd);
// the older silero pick is retired (DESIGN.md § Superseded, "named STT/TTS
// the older silero pick is retired (docs/design.md § Superseded, "named STT/TTS
// model picks"). A different voice is a model-file swap, not a code change.
// The daemon wires one impl — Remote pointing at the worker socket if
// configured, Stub otherwise.
+70 -1
View File
@@ -57,8 +57,20 @@ type Server struct {
ln net.Listener
wg sync.WaitGroup
done chan struct{}
// Accepted conns, tracked so Close can shut them. Same defect as
// ipc/server.go had: closing only the listener leaves every idle client
// parked in readFrame, wg.Wait never returns, and the daemon dies to
// SIGKILL without sealing the database.
connMu sync.Mutex
conns map[net.Conn]struct{}
}
// closeGrace — how long Close waits for in-flight dispatches before dropping
// them. A push-to-talk turn can be mid-inference; abandoning one costs a reply,
// hanging costs every write since the last clean shutdown.
const closeGrace = 3 * time.Second
// NewServer builds a Server bound to addr (e.g. "127.0.0.1:9100" for a
// local-only smoke; production: a wg-tunnel address). handler is the
// reactive handler; sessions is shared with the voicesink (the daemon
@@ -111,8 +123,10 @@ func (s *Server) Serve() error {
}
}
s.wg.Add(1)
s.trackConn(c)
go func(c net.Conn) {
defer s.wg.Done()
defer s.untrackConn(c)
s.serveConn(c)
}(c)
}
@@ -212,10 +226,65 @@ func (s *Server) Close() error {
if s.ln != nil {
err = s.ln.Close()
}
s.wg.Wait()
// Close the accepted conns too, or a client that is merely idle keeps
// serveConn blocked in readFrame forever.
s.closeConns()
if !waitTimeout(&s.wg, closeGrace) {
log.Printf("voice: %d connection(s) still busy after %s, closing anyway", s.liveConns(), closeGrace)
}
return err
}
func (s *Server) trackConn(c net.Conn) {
s.connMu.Lock()
defer s.connMu.Unlock()
if s.conns == nil {
s.conns = make(map[net.Conn]struct{})
}
s.conns[c] = struct{}{}
}
func (s *Server) untrackConn(c net.Conn) {
s.connMu.Lock()
defer s.connMu.Unlock()
delete(s.conns, c)
}
func (s *Server) liveConns() int {
s.connMu.Lock()
defer s.connMu.Unlock()
return len(s.conns)
}
// closeConns unblocks every parked reader. serveConn's own defer closes the
// conn again; a second Close on a net.Conn is a harmless error.
func (s *Server) closeConns() {
s.connMu.Lock()
conns := make([]net.Conn, 0, len(s.conns))
for c := range s.conns {
conns = append(conns, c)
}
s.connMu.Unlock()
for _, c := range conns {
_ = c.Close()
}
}
// waitTimeout waits on wg, but not forever. Reports whether it finished.
func waitTimeout(wg *sync.WaitGroup, d time.Duration) bool {
done := make(chan struct{})
go func() {
wg.Wait()
close(done)
}()
select {
case <-done:
return true
case <-time.After(d):
return false
}
}
func unmarshalParams(raw json.RawMessage, v any) error {
if len(raw) == 0 {
raw = []byte("null")
+229
View File
@@ -0,0 +1,229 @@
// Package websearch reads a self-hosted SearXNG instance.
//
// Why this exists at all: "never phones home" stopped being a hard constraint
// on 2026-07-31. A 1.7B does not know enough to answer a world question, and
// reading beats recalling at that size. SearXNG is the reading surface for
// anything the offline ZIMs do not hold, and it is off unless configured.
//
// What is NOT here, on purpose:
//
// - No query rewriting. SearXNG ranks with real engines, so the Russian
// question goes out as he asked it. That is the whole reason it sits ahead
// of Kiwix, whose keyword ranker needs kiwix.Rewriter to see anything.
// - No page fetching. A snippet per result is the evidence; following a link
// is crawl.Crawler's job and carries robots and allowlist rules with it.
// - No cache and no retries. Boring on purpose, same posture as kiwix.Client.
//
// Only the query string leaves this process. This package cannot read the
// store, so his notes, facts, persona block and history cannot travel with a
// search even by accident. The personal boundary in the query chain is what
// keeps a question ABOUT him from becoming a query at all.
package websearch
import (
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"net/url"
"strings"
"time"
)
// Result is one search hit, already reduced to what a phraser can read.
type Result struct {
Title string
URL string
Content string // the engine's snippet, plain text
Engine string // which upstream engine produced it, e.g. "duckduckgo"
}
// Response is one search. Answers comes from SearXNG's answerer plugins and
// from instant answers upstream; it is a direct reply to the question and is
// worth more than any snippet, so it is kept separate rather than mixed in.
type Response struct {
Answers []string
Results []Result
}
// Empty reports whether the search found nothing usable. The caller passes the
// turn on when it does — an empty search is not a failure worth announcing.
func (r Response) Empty() bool { return len(r.Answers) == 0 && len(r.Results) == 0 }
// DefaultTimeout — the whole request. SearXNG fans out to upstream engines and
// waits on the slowest, so this is longer than a LAN call but short enough that
// a dead engine does not hold a voice turn open.
const DefaultTimeout = 8 * time.Second
// maxBodyBytes caps the JSON read. A 20-result reply is tens of kilobytes; this
// is slack for a wide one and a hard stop against a misconfigured endpoint.
const maxBodyBytes = 4 << 20
// Client is a SearXNG HTTP client.
type Client struct {
base string
language string
engines string
http *http.Client
}
// Options are the per-instance knobs, all optional.
type Options struct {
// Language — SearXNG's `language` parameter, e.g. "ru" or "auto". Empty ⇒
// the instance default.
Language string
// Engines — comma-separated engine names to restrict the search to. Empty ⇒
// whatever the instance has enabled.
Engines string
// Timeout — per-request budget. 0 ⇒ DefaultTimeout.
Timeout time.Duration
}
// New makes a client for a SearXNG base URL like http://searxng:9563.
//
// The instance must have the JSON format enabled (`search.formats: [html,
// json]` in its settings.yml); a stock install answers 403 to format=json and
// every search will fail with that status.
func New(baseURL string, opt Options) *Client {
t := opt.Timeout
if t <= 0 {
t = DefaultTimeout
}
return &Client{
base: strings.TrimRight(baseURL, "/"),
language: strings.TrimSpace(opt.Language),
engines: strings.TrimSpace(opt.Engines),
http: &http.Client{Timeout: t},
}
}
// Search runs one query and returns up to limit results plus any instant
// answers. The query goes out verbatim.
func (c *Client) Search(ctx context.Context, query string, limit int) (Response, error) {
query = strings.TrimSpace(query)
if query == "" {
return Response{}, fmt.Errorf("websearch: empty query")
}
q := url.Values{}
q.Set("q", query)
q.Set("format", "json")
if c.language != "" {
q.Set("language", c.language)
}
if c.engines != "" {
q.Set("engines", c.engines)
}
req, err := http.NewRequestWithContext(ctx, http.MethodGet, c.base+"/search?"+q.Encode(), nil)
if err != nil {
return Response{}, err
}
resp, err := c.http.Do(req)
if err != nil {
return Response{}, err
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return Response{}, fmt.Errorf("websearch: http %d (json format enabled in searxng?)", resp.StatusCode)
}
body, err := io.ReadAll(io.LimitReader(resp.Body, maxBodyBytes))
if err != nil {
return Response{}, err
}
return ParseResponse(body, limit)
}
// wire mirrors just the fields of the SearXNG JSON reply we read.
type wire struct {
Answers []json.RawMessage `json:"answers"`
Results []struct {
Title string `json:"title"`
URL string `json:"url"`
Content string `json:"content"`
Engine string `json:"engine"`
} `json:"results"`
}
// ParseResponse turns a SearXNG JSON reply into a Response, keeping at most
// limit results. Exported so the parser is testable from a captured reply with
// no instance running.
func ParseResponse(body []byte, limit int) (Response, error) {
var doc wire
if err := json.Unmarshal(body, &doc); err != nil {
return Response{}, fmt.Errorf("websearch: bad json: %w", err)
}
if limit <= 0 {
limit = 5
}
out := Response{}
for _, raw := range doc.Answers {
if s := answerText(raw); s != "" {
out.Answers = append(out.Answers, s)
}
}
for _, r := range doc.Results {
title := clean(r.Title)
content := clean(r.Content)
if title == "" && content == "" {
// A hit with no text is a link with nothing to read. It cannot be
// evidence, and counting it toward the limit would push a usable
// snippet out of the reply.
continue
}
out.Results = append(out.Results, Result{
Title: title,
URL: strings.TrimSpace(r.URL),
Content: content,
Engine: strings.TrimSpace(r.Engine),
})
if len(out.Results) == limit {
break
}
}
return out, nil
}
// answerText reads one entry of `answers`. SearXNG changed its shape: older
// versions emit a bare string, newer ones an object with an `answer` field.
// Both are in the wild depending on when the instance was pulled, so both are
// read rather than pinning a version we do not control.
func answerText(raw json.RawMessage) string {
var s string
if err := json.Unmarshal(raw, &s); err == nil {
return clean(s)
}
var obj struct {
Answer string `json:"answer"`
}
if err := json.Unmarshal(raw, &obj); err == nil {
return clean(obj.Answer)
}
return ""
}
// Snippets renders the response as evidence lines for a phraser: instant
// answers first, then "Title — snippet" per result.
//
// Answers lead because they are a reply to the question, where a result is a
// page that might contain one. The URL is deliberately left out: it is not
// evidence, and piper reads one out character by character.
func (r Response) Snippets() []string {
out := make([]string, 0, len(r.Answers)+len(r.Results))
out = append(out, r.Answers...)
for _, res := range r.Results {
switch {
case res.Content == "":
out = append(out, res.Title)
case res.Title == "":
out = append(out, res.Content)
default:
out = append(out, res.Title+" — "+res.Content)
}
}
return out
}
// clean collapses whitespace. Snippets arrive with newlines and runs of spaces
// from the upstream page, and piper reads a reply built out of them badly.
func clean(s string) string { return strings.Join(strings.Fields(s), " ") }
+145
View File
@@ -0,0 +1,145 @@
package websearch
import (
"context"
"net/http"
"net/http/httptest"
"strings"
"testing"
)
const sampleJSON = `{
"query": "почему небо голубое",
"answers": ["Rayleigh scattering makes the sky blue."],
"results": [
{"title": "Рэлеевское рассеяние", "url": "https://ru.wikipedia.org/x", "content": "Рассеяние\n света на молекулах.", "engine": "wikipedia"},
{"title": "", "url": "https://example.org/empty", "content": "", "engine": "duckduckgo"},
{"title": "Why is the sky blue", "url": "https://example.org/2", "content": "Short answer.", "engine": "duckduckgo"}
]
}`
func TestParseResponse(t *testing.T) {
got, err := ParseResponse([]byte(sampleJSON), 5)
if err != nil {
t.Fatalf("parse: %v", err)
}
if len(got.Answers) != 1 || got.Answers[0] != "Rayleigh scattering makes the sky blue." {
t.Fatalf("answers = %#v", got.Answers)
}
// The textless middle hit is dropped: it is a link with nothing to read.
if len(got.Results) != 2 {
t.Fatalf("results = %#v", got.Results)
}
if got.Results[0].Content != "Рассеяние света на молекулах." {
t.Fatalf("whitespace not collapsed: %q", got.Results[0].Content)
}
if got.Empty() {
t.Fatal("Empty() on a response with hits")
}
}
// The limit counts usable hits, not raw ones — a textless entry must not push a
// real snippet out of the reply.
func TestParseResponseLimitSkipsEmpty(t *testing.T) {
got, err := ParseResponse([]byte(sampleJSON), 2)
if err != nil {
t.Fatalf("parse: %v", err)
}
if len(got.Results) != 2 {
t.Fatalf("results = %d, want 2", len(got.Results))
}
if got.Results[1].Title != "Why is the sky blue" {
t.Fatalf("second hit = %q", got.Results[1].Title)
}
}
// Newer SearXNG emits answers as objects; older ones as bare strings. Both are
// in the wild and both must read.
func TestParseResponseObjectAnswers(t *testing.T) {
got, err := ParseResponse([]byte(`{"answers":[{"answer":"42","url":"x"}],"results":[]}`), 5)
if err != nil {
t.Fatalf("parse: %v", err)
}
if len(got.Answers) != 1 || got.Answers[0] != "42" {
t.Fatalf("answers = %#v", got.Answers)
}
}
func TestResponseEmpty(t *testing.T) {
got, err := ParseResponse([]byte(`{"answers":[],"results":[]}`), 5)
if err != nil {
t.Fatalf("parse: %v", err)
}
if !got.Empty() {
t.Fatal("Empty() = false on a reply with nothing in it")
}
}
func TestSnippetsAnswersFirst(t *testing.T) {
got, _ := ParseResponse([]byte(sampleJSON), 5)
lines := got.Snippets()
if len(lines) != 3 {
t.Fatalf("lines = %#v", lines)
}
if lines[0] != "Rayleigh scattering makes the sky blue." {
t.Fatalf("answer did not lead: %q", lines[0])
}
if !strings.Contains(lines[1], " — ") {
t.Fatalf("result line = %q", lines[1])
}
// No URL travels into the evidence: piper reads one out character by
// character and it is not evidence anyway.
for _, l := range lines {
if strings.Contains(l, "http") {
t.Fatalf("url leaked into evidence: %q", l)
}
}
}
// The query goes out verbatim, and the JSON format is always asked for.
func TestSearchRequest(t *testing.T) {
var gotQuery, gotFormat, gotLang, gotEngines string
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
gotQuery = r.URL.Query().Get("q")
gotFormat = r.URL.Query().Get("format")
gotLang = r.URL.Query().Get("language")
gotEngines = r.URL.Query().Get("engines")
w.Write([]byte(sampleJSON))
}))
defer srv.Close()
c := New(srv.URL, Options{Language: "ru", Engines: "duckduckgo"})
got, err := c.Search(context.Background(), "почему небо голубое", 3)
if err != nil {
t.Fatalf("search: %v", err)
}
if gotQuery != "почему небо голубое" {
t.Fatalf("query was rewritten: %q", gotQuery)
}
if gotFormat != "json" || gotLang != "ru" || gotEngines != "duckduckgo" {
t.Fatalf("format=%q language=%q engines=%q", gotFormat, gotLang, gotEngines)
}
if len(got.Results) != 2 {
t.Fatalf("results = %#v", got.Results)
}
}
// A stock SearXNG answers 403 to format=json. The error must say so, because
// that is the one misconfiguration this client cannot work around.
func TestSearchHTTPError(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Error(w, "forbidden", http.StatusForbidden)
}))
defer srv.Close()
_, err := New(srv.URL, Options{}).Search(context.Background(), "x", 3)
if err == nil || !strings.Contains(err.Error(), "403") {
t.Fatalf("err = %v", err)
}
}
func TestSearchEmptyQuery(t *testing.T) {
if _, err := New("http://example.invalid", Options{}).Search(context.Background(), " ", 3); err == nil {
t.Fatal("empty query accepted")
}
}
-1
View File
@@ -1 +0,0 @@
/home/kami/apps/Maven/models/stt
-1
View File
@@ -1 +0,0 @@
/home/kami/apps/Maven/models/tts