Compare commits

...

64 Commits

Author SHA1 Message Date
kami 453919db20 Merge pull request 'Bug: the memory index stores the raw utterance as a fact's recall text, and nothing ever deletes a fact vector' (#100) from task/493-bug-the-memory-index-stores-the-raw-utte into task/470-bug-a-question-writes-invented-knowledge
Reviewed-on: #100
2026-08-03 20:41:09 +02:00
claude ad60e10e95 mavend: run the fact vector repair on start, and test what it does (V-493)
Automatic rather than a flag, unlike -reembed: only voice-tapped facts are in
this index, so it is tens of embeddings rather than thousands of notes. And
waiting for an operator to know the repair exists is the failure being fixed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:36:54 +04:00
claude 1528697287 store: repair fact vectors against the facts they name (V-493)
Every write-path fix leaves the rows already stored wrong, and a box in that
state looks fine: recall answers with the wrong text and nothing logs an error.
That is how the original poison survived four restarts.

RepairFactVectors resolves each fact vector against the fact it names,
re-embeds the ones whose text is stale, and deletes the voided, superseded and
orphaned ones. Marker-guarded and idempotent, so it runs once per box and a run
that dies partway is simply redone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:36:54 +04:00
claude dbdab2d570 store, mavend: a fact is indexed as the fact, not as the utterance (V-493)
queryMemory returns a fact's stored text verbatim, so the text the write path
indexed is what he hears. It was the utterance, which made recall of any
voice-tapped fact answer with the sentence he said: go_version = 1.20 was
indexed as "какая последняя версия языка Go?", and that question came back.

FactRecallText renders the fact instead, and the utterance stays in meta as
provenance. Correcting a value now drops the key's vectors the way voiding one
does, since the superseded value was still answering.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:36:34 +04:00
kami b9371dcac6 Merge pull request 'Bug: a question writes invented knowledge into memory as a self fact, and recall then serves it back for unrelated questions' (#99) from task/470-bug-a-question-writes-invented-knowledge into master
Reviewed-on: #99
2026-08-03 20:11:11 +02:00
claude 62c2e92ec0 mavend, recalleval: wire the topic veto into both recall sources (V-470)
queryMemory and queryNotes both gate on score alone, so both needed it. The
eval keeps its own copy of bestRecall — package main is not importable — and a
fixture that measures a weaker gate than the daemon runs flatters it, so the copy
moves in step and its test pins the new rule.

Measured on the held-out recall fixture with the real embedder: 17/32 cases pass
→ 22/32, false recall 1/5 → 0/5, answered after gate 18/27 → 17/27. The one true
recall lost is en-hard-024, an English question against a Russian note, where no
lexical test can help.
2026-08-03 13:51:04 +04:00
claude aec94eb2e8 memory: a world question must name what the memory mentions (V-470)
The score gate cannot separate the right note from an unrelated one: the
held-out fixture puts the right note at 0.791-0.890 and the must-be-silent cases
at 0.795-0.835, so a note about his slow network answered 'почему небо синее?'.

RecallAllowed adds a topic veto, and applies it only to a question that mentions
nothing of his. That restriction is the whole design: demanding a shared word of
every recall silenced four true recalls on the fixture to kill one false one,
because recall exists to find the note whose words he no longer remembers. A
question about his own life keeps the embedder as its only judge.
2026-08-03 13:50:54 +04:00
claude 4dfe106fe3 mavend: a question is never a fact about him (V-470)
IntentFact used to persist whatever the model invented for a question-shaped
utterance, at confidence 1.00, and index it for recall under the question's own
text. Two such rows then claimed seven unrelated world questions and silently
disabled world answering.

A question now goes down the query chain, which is what he asked for. The second
half is confidence: a value grounded in what he said stays 1.00, a value the model
supplied for words he never said drops to 0.60 and says so in the log. Same
reasoning as 'LLM output is not authorization' on the act path.
2026-08-03 13:40:33 +04:00
claude 2e0e2fd0bb router: a deterministic test for question-shaped text (V-470)
The predicate a fact write needs before it trusts a routing decision. Tokenized,
not substring: 'что' inside 'чтобы' is not a question. Capture verbs win over
every question signal, because 'запиши что я пил воду' contains an interrogative
and is still a capture.
2026-08-03 13:40:33 +04:00
claude f3fa6b353a store: voiding a fact drops its memory vectors (V-470)
Revert voided the fact row and left the vector, so recall kept serving the
voided fact's utterance and the documented repair reported success on a box that
stayed broken. There was no way to repair a poisoned box at all.

DeletePrefix covers every vector for the key, earlier rows included: their values
are superseded, and a superseded value has no business claiming a turn. It is
best-effort — the audit trail is already committed, and a fact that is voided but
still recallable beats a void that failed.
2026-08-03 13:40:13 +04:00
kami 6645f64c3e Merge pull request 'Name the gap: world questions through the workstation model, and the four remaining callers' (#98) from task/490-name-the-gap-world-questions-through-the into master
Reviewed-on: #98
2026-08-03 11:19:56 +02:00
claude f10e0068dd config, deploy: the workstation is workpc, not bugmachine (V-490)
Owner's correction. It is the same host CLAUDE.md already calls workpc, and
two names for one machine read as two machines. The dated eval file keeps the
old name: a measurement is never edited after the day it was taken.
2026-08-03 12:42:37 +04:00
claude 9b124d9194 docs: both halves of the degradation rule are wired, and which caller is which (V-490)
The offload inventory grows a column, because "seven callers of the resident
model" stopped being the useful fact. Which of them is offloaded, and under
which half of the rule, is. Three are resident-only on purpose and the table
now says why rather than leaving it to be rediscovered.

The three-outcome table is the part that was not obvious from the rule as
written. A configured-and-asleep workstation names the gap; a box with no
workstation block does not, because naming a gap requires a gap.
2026-08-03 12:29:12 +04:00
claude 12530c8a95 mavend: world questions ask the workstation, and name the gap when it is asleep (V-490)
queryGeneral has nothing fetched to fall back on, so it is the sharp case:
with a workstation configured and asleep he is told that, rather than told
something false in a confident voice. The 1.7B answering a world question is
where "Война и мир" got Левитан as its author.

The sources that already hold a passage — a live search, a ZIM article, a
page he named — go through the world model too, but read the passage back
when it is not there instead of naming a gap. A real quote beats "не могу
сейчас", and nothing is invented on either path.

The Stub and every test double keep the Phraser interface they have.
PhraseWorld is reached by assertion, and a phraser without it is the
no-workstation case.
2026-08-03 12:27:40 +04:00
claude 51256c4c9a phraser: test the three outcomes of a world question, and prompt parity (V-490)
The middle outcome is the whole task: a workstation that is configured and
asleep produces a gap, and the resident model is never asked. The parity
test compares the bytes PhraseWorld sends the workstation against the bytes
PhraseQuery sends the resident model, so the fixtures and the daemon cannot
measure two different prompts.

The nudge tests cover the silent half from both sides, including the
temperature, which is how the workstation would otherwise change how she
sounds without anyone deciding to.
2026-08-03 12:27:30 +04:00
claude 76481c2736 phraser: a world model seam, so a gap can be named instead of invented (V-490)
The naming half of the degradation rule in docs/offload.md. PhraseWorld has
three outcomes: no workstation configured means the resident model answers
exactly as today, a workstation that is taking work answers, and one that is
asleep returns ErrNoWorldModel so the caller can say so. Naming a gap
requires a gap — on a box that never had a second model, refusing every
world question would remove a capability he has now.

Both prompts move into knowledgePrompt and evidencePrompt, shared by
PhraseQuery and PhraseWorld, because prompt parity across two models stops
holding the moment there are two copies of a prompt.

The silent half comes with it: chatWithSystem and chatWithMessages prefer
the workstation when it will take work, at the same 0.7 the resident
transport samples at, and say nothing when it will not. That covers the
digestion worker's nudge and reminder phrasing without touching tick.go.

Only Available and CompleteRemote are in the Remote interface. Pair.Complete
has its own floor and the phraser already owns one; two floors under a
single call is one too many.
2026-08-03 12:27:30 +04:00
claude bcc2305cd0 llm: let a caller name its sampling temperature (V-490)
The phraser's own transport has always sampled at 0.7 and this client has
always been greedy. Routing a phrasing call through the client must not
change how it decodes, so Req carries the temperature and 0 — the zero
value, and what every existing caller wanted — is still greedy.
2026-08-03 12:27:04 +04:00
kami 0ceeac8df4 Merge pull request 'Point Maven at the workstation model: a workstation block, and routing plus replies through llm.Pair' (#97) from task/485-run-the-big-model-on-the-workstation-wit into master
Reviewed-on: #97
2026-08-03 10:13:09 +02:00
claude 4fae13af75 docs: record the workstation routing numbers where the router is documented (V-485)
CLAUDE.md carried only the homesrv figures, which now read as the whole story.
Also points offload.md's order at #490 for the naming half.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJYcaBiuny9UGSpFweQVQ1
2026-08-03 10:12:21 +02:00
claude 774217199e docs: measure gemma-4-12b on the workstation against the resident model (V-485)
Both fixtures, run from homesrv across the LAN with the proxy env stripped.
Routing: 84.4% full / 93.5% intent-only at p50 329ms through the cascade, against
72.7% / 77.9% at p50 0.80-1.04s for Qwen3-1.7B. Talk: 25/27 against 20/27, with
knowledge 9/9. Nudges 15/15. Settles #485's first assumption by measurement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJYcaBiuny9UGSpFweQVQ1
2026-08-03 10:12:21 +02:00
claude 2db59d52a7 deploy, docs: point homesrv at bugmachine and say what is still unwired (V-485) 2026-08-03 10:12:21 +02:00
claude 92d5fd580c mavend: route and reply through the workstation when its card is free (V-485)
modelSeam builds an llm.Pair when a workstation is configured and hands it to
the router and the replier. Both are the silent half of the degradation rule:
the big model is only better there, and he is never told which model answered.
No block, no probe, and the box behaves exactly as it did.
2026-08-03 10:12:21 +02:00
claude edeef19ff0 config: a workstation block, dropped when it names no address (V-485)
Health defaults to the supervisor's /health rather than llama-server's,
because mavgpud is what answers 503 while the card is held.
2026-08-03 10:12:21 +02:00
kami 018f7a6f47 Merge pull request 'Run the big model on the workstation, with admission control and the 1.7B as the floor' (#96) from task/489-workstation-deploy-mavgpud-on-workpc-and into master
Reviewed-on: #96
2026-08-03 10:12:13 +02:00
claude eca41798bd mavgpud: turn gemma's thinking off in the chat template (V-489)
Owner's call, 02-08-2026. Without it the 12B spends the reply budget on
reasoning tokens and answers empty at low max_tokens. Verified on the box:
"Столица Франции?" now answers "Париж" with no reasoning_content.
2026-08-02 22:44:10 +04:00
claude cc423567e7 docs: record that contention is KFD presence, not a VRAM threshold (V-489) 2026-08-02 22:29:58 +04:00
claude 8088ef9e00 mavgpud: build it with the rest, and ship the workstation config and unit (V-489)
make build now catches a broken supervisor on homesrv. deploy/mavgpud.json
carries the owner's gemma-4-12b line with the MTP draft model, passed to
llama-server untouched. The unit is a systemd user unit because sudo on the
workstation wants a password; lingering is the one command left to the owner.
2026-08-02 22:29:57 +04:00
kami 666b924d29 Merge pull request 'Run the big model on the workstation, with admission control and the 1.7B as the floor' (#95) from task/488-workstation-a-supervisor-that-keeps-llam into master
Reviewed-on: #95
2026-08-02 17:04:05 +02:00
claude e52c616592 mavgpud: test the probe against the sysfs the workstation actually has (V-488)
The fixtures are the live numbers sampled from the box on 02-08-2026, where the
CPT run held 12.8GB of 16 as proc/478104/vram_35881.

The cases that matter are the ones where a mistake is silent: our own
llama-server counting as a contender, an unreadable card reading as free, and
/health hanging or proxying into a closed port instead of answering 503.
2026-08-02 17:03:56 +02:00
claude 2b97bac51e mavgpud: keep the model loaded while the card is free, yield when it is not (V-488)
The lifecycle rule from Vikunja #488. Not on demand, because a 7-14B takes tens
of seconds to load and a world question would meet a gap every time the card
had been quiet. Not always on, because that is what holds the card.

/health is answered locally and always, so Maven's prober costs nothing and
works while the model is down. Everything else is reverse-proxied to
llama-server, which is what makes the idle window measurable at all.

Yielding is checked before starting, and both transitions are damped by a poll
streak so a short-lived rocm process cannot evict the model.
2026-08-02 17:03:56 +02:00
claude ab42db2b87 mavgpud: read the card from sysfs and own llama-server's lifecycle (V-488)
The workstation cannot keep a 7-14B resident: it would hold 16GB against the
owner's CPT runs, Correx and the manga-recap pipeline. So the process that
stays up costs no VRAM and the model comes and goes under it.

Contention is detected by presence on the KFD, not by a VRAM threshold. A ROCm
process registers under /sys/class/kfd/kfd/proc when it initialises HIP, well
before it allocates, so we see a contender during its startup instead of after
it has already lost an allocation race. rocm-smi is not installed on that box
and a per-second subprocess would get tuned down until useless, so this reads
sysfs and forks nothing.

Free VRAM is read only to decide whether to start. It is never a reason to
stop: by the time free VRAM has dropped, the other job has already failed.
2026-08-02 17:03:56 +02:00
kami 94d553570d Merge pull request 'Run the big model on the workstation, with admission control and the 1.7B as the floor' (#94) from task/485-run-the-big-model-on-the-workstation-wit into master
Reviewed-on: #94
2026-08-02 17:03:25 +02:00
claude 2e97b905b4 docs: the workstation supervisor owns llama-server's lifecycle (V-485)
The remote model cannot be a llama-server that is simply left running: a
resident 7-14B holds 16GB against the CPT runs the card is for. So what
is always up on the workstation is a supervisor, and llama-server is
loaded while the card is free.

Still not a scheduler. It arbitrates nothing between callers, and Maven
never asks it to start anything.
2026-08-02 18:14:31 +04:00
claude fbcca449be llm: pin that a down workstation is invisible (V-485)
Seven cases. The load-bearing ones are the constraint from 483: an
unconfigured deploy never probes and always reaches the floor, a busy
card degrades silently with the remote untouched, and a remote that dies
between probes still completes the turn and corrects the cached answer on
its way out.

CompleteRemote is pinned not to fall back, because a named gap that
quietly became a 1.7B guess is the failure this whole split exists to
prevent. And 1000 Available calls are pinned to make zero probes.
2026-08-02 17:19:53 +04:00
claude 2076e4a788 llm: prefer the workstation model, floor on the resident one (V-485)
Pair holds both models and decides which answers. A prober asks the
remote whether it will take work and caches the answer, so a request
reads an atomic bool rather than paying for a health check. Routing sits
at p50 825ms on the hot path and must never wait on a machine that may be
asleep.

The two methods are the two halves of the degradation rule in
docs/offload.md. Complete falls back silently, for routing, replies and
nudge phrasing, where the big model is only better. CompleteRemote
returns ErrRemoteUnavailable instead, for a world question, where the
1.7B does not answer worse but invents.

A nil remote is the unconfigured deploy: nothing probes, everything goes
to the floor, and the box behaves exactly as it does today.
2026-08-02 17:19:53 +04:00
kami 30eb6add1b Merge pull request 'Docs: refresh the QA plan against the live task list' (#93) from task/483-docs-offload-design into master
Reviewed-on: #93
2026-08-02 15:08:42 +02:00
claude dc266056d1 docs: the shape and the rules for offloading model work (V-483)
483 is an umbrella and its children are the work, so what it owes them is
the shape they must all obey. docs/offload.md records it: the degradation
rule and where its line falls, admission control rather than a GPU
arbiter, the embedder staying on homesrv because it backs the classifier,
and the inventory of what runs a model on the box today.

CLAUDE.md gets a pointer, because an agent about to add a model caller or
touch a daemon seam needs to know this before it starts, not after.
2026-08-02 15:08:35 +02:00
kami 1c786b7156 Merge pull request 'Docs: refresh the QA plan against the live task list' (#92) from task/483-design-offload-ml-to-the-workstation-kee into master
Reviewed-on: #92
2026-08-02 15:08:16 +02:00
claude a3af10a830 gitignore the root .env, it holds a live token (V-484)
It was untracked but not ignored, so one git add -A would have committed
MAVEN_AMBIENT_TOKEN. Same class as deploy/telegram.env, which is already
ignored.
2026-08-02 15:08:07 +02:00
claude c0de473382 ipc, worker: dial and bind through netaddr (V-484)
Five hardcoded transports, three in internal/ipc and two in
internal/worker, all now go through the seam address. The unix perms
logic moved into netaddr, so the two copies of parentDir and the umask
dance are gone.

peerCaller already returned ok=false for a non-unix conn, so the
SO_PEERCRED path degrades correctly on tcp with no change.
2026-08-02 15:08:07 +02:00
claude 3e534340bf ipc: pin that a scheme-less address still dials unix (V-484)
Six cases. The load-bearing one is the first: every deploy in the tree
writes a bare path, and it must keep meaning a unix socket with no
handshake in front of the payload.

The rest cover the tcp seam: a good token round-trips, a wrong one comes
back ErrUnauthorized, a stranger that speaks HTTP at the port is dropped
while the listener stays up for the next peer, and a tokenless tcp bind
fails rather than serving his turns to anyone who connects.
2026-08-02 15:08:07 +02:00
claude 1a704d704d ipc: a seam address that can name a transport (V-484)
internal/netaddr parses a daemon seam address and dials or binds it. A
scheme-less address is unix and behaves exactly as it does today: same
0700 parent dir, same 0600 socket, same bytes on the wire. tcp://host:port
is the new option, and it is what lets a module live on another host.

Over tcp the filesystem permission that authenticated the unix socket is
gone, and what crosses this seam is audio of the owner speaking. So a tcp
listener requires a shared token, checked in constant time before the
first protocol frame is read, and a peer that fails is dropped without
taking the listener down with it.
2026-08-02 15:08:07 +02:00
kami e57adcb001 Merge pull request 'Docs: refresh the QA plan against the live task list' (#91) from task/459-docs-refresh-the-qa-plan-against-the-liv into master
Reviewed-on: #91
2026-08-02 15:07:21 +02:00
claude bec7362b7b config: clear three of 472's five QA blockers (V-459)
morning_routines, feeds, crawl.on_demand and netscan.enabled in
deploy/mavend.json; -ambient-token on mavweb, interpolated from a gitignored
/.env. All verified on the box: the dispatcher builds the morning plan,
/api/ambient answers 401/201, the crawler reads a named page, netscan finds 3
devices, and /events fills with scan:lan and ambient:notif.

Filed 482 (ambient reads a notification's wall clock as UTC). Corrected 479:
both capabilities work once configured, so it is not a routing defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJYcaBiuny9UGSpFweQVQ1
2026-08-02 15:52:56 +04:00
claude a3ec746a01 docs: the six ready tasks all ran, and all six stop at the deploy (V-459) 2026-08-02 15:44:17 +04:00
claude af0eec250e docs: the operations sitting and the degraded-mode suite both ran (V-459) 2026-08-02 15:15:08 +04:00
claude 20aa2d59c9 docs: session 3 results and the query-source findings (V-459) 2026-08-02 14:40:22 +04:00
claude 4bad90dedb docs: point the QA findings at their new task ids (V-459) 2026-08-02 14:25:39 +04:00
claude 2b8d0f74fa docs: session 1 and 2 results, and the classifier baseline was wrong (V-459)
Ran sessions 1 and 2 on the live box.

Session 1 steps 1 and 3-6 pass. Steps 2 and 7-9 need a person at the box.
POST /api/chat is drivable with form encoding and a cookie jar, so the text
half needs no browser.

Session 2 confirms the deploy matches the bench at 72.7% full accuracy, and
contradicts two recorded numbers. The classifier scores 68.8% at p50 16.6us,
not 36.8% at 31ms. Router latency measured under contention again.

Also: 319 item 2 point 2 closes on the recall margin sweep, CheckFeminine has
a false positive on second-person masculine verbs, and the wake path cannot be
checked because mavwaked and mavenclient are deployed nowhere.
2026-08-02 14:24:01 +04:00
claude af9d2133dc docs: session 3 holds five sittings now, not three (V-459)
The refresh added the query-sources and operations sittings and left the
heading counting three.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJYcaBiuny9UGSpFweQVQ1
2026-08-02 13:57:01 +04:00
claude a1fdfccd61 docs: refresh the QA plan against the live task list (V-459)
The plan named 40 task numbers on 2026-08-01. Ten open QA tasks were missing
and two of the named ones had closed, so the 44-of-50 header was wrong twice
over.

- header is 42 of 50, and every open task now appears
- placed the ten unlisted QA tasks: 14, 248, 249, 250, 258, 283, 284, 285,
  286, 323
- new Operations sitting for 249 and 250, and a Query sources sitting for
  258 and 286
- 14 and 284 join housekeeping: both are gated on something unbuilt
- dropped the 317 and 354 rows, closed 01-08-2026, with one line saying what
  landed
- 319's gate recalibration is done; what is left is re-deriving QueryMinMargin
- 323 is down to the 60s startup timeout arm after PR #90
- new "Not this repo" section for 358 (Hexis) and 362 (training workspace)
- router latency is ~27x, not 90x; the 2.7s p50 was contention, not the model

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJYcaBiuny9UGSpFweQVQ1
2026-08-02 13:56:02 +04:00
kami 5c05163266 Merge pull request 'QA: phraser coverage is 65.3% but the llama-server subprocess lifecycle is 0% — the suspicion in this task was correct' (#90) from task/323-qa-phraser-coverage-is-65-3-but-the-llam into master 2026-08-02 11:39:20 +02:00
kami 92d2629001 Merge pull request 'Voice cannot accept a routine (V-367); the last three prompts are Russian (V-404)' (#89) from fix/367-voice-parks-routine-accept into master 2026-08-02 11:39:16 +02:00
kami bdcfccce77 Merge pull request 'Session workflow: pickup and wrap around the task flow' (#85) from task/445-session-workflow into master 2026-08-02 11:39:11 +02:00
kami f4deccacc9 Merge pull request 'dialogue.Slots and router.Slots are hand-kept copies that already drifted' (#88) from task/365-dialogue-slots-and-router-slots-are-hand into master 2026-08-02 11:39:07 +02:00
kami 8aaac01de6 Merge pull request 'Doc reorg: tier the tree, retire the three planning files' (#86) from task/446-doc-reorg-tier-the-tree-retire-the-three into master 2026-08-02 11:37:43 +02:00
claude feb6f2c03d phraser: test the spawn path, the one thing coverage never touched (V-323)
Every phraser test built the phraser with NewLLMPhraserAt, which starts no
process, so NewLLMPhraser, spawnLlamaServer, startLlamaProc, llamaProc.Close
and extractPort sat at 0% while the package headline read 65.3%.

These drive the real spawn code against a fake llama-server script: the port
scrape, the three reachable startup-race arms (start failure, stderr EOF,
context cancel), and Close actually reaping the child. The orphan test
re-execs the test binary as the daemon, SIGKILLs it, and asserts Pdeathsig
killed the grandchild. The last test rebuilds the production command line and
checks kill-maven.sh's pattern still matches it — that pattern has gone stale
twice and leaked orphans both times.

Package coverage 65.3% -> 76.9%. The 60s timeout arm stays untested; it needs
an injectable clock in production code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012YQGVXu5J1iCMCff5J4S1R
2026-08-02 10:15:35 +04:00
claude 99bb3526db Stop asking for Russian in English on the last three prompts (V-404)
#400 rewrote the chat and query prompts in Russian and left three pieces
of English prose behind.

PhraseReminder's user prompt was fully English. It is Russian now, and it
no longer restates the JSON contract or the persona rules: the call goes
through chat(), so nudgeSystem already states both, and a second copy of a
contract is one more thing that can drift out of step with the first.

querySystemPrompt and router.KnowledgePrompt both closed with the English
"Respond ONLY with valid JSON:". That sentence is prose instruction, not
wire format — the JSON skeleton after it is the wire format, and it is
unchanged. Kept rather than deleted: the GBNF grammar makes it close to
redundant, but the grammar is switchable off (phraser NoGrammar), and the
sentence is the floor when it is.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018CotYKycuio1GwLbYh9jfc
2026-08-02 10:01:58 +04:00
claude bb8cb8d014 Voice parks a routine proposal, it never accepts it (V-367)
Accepting a proposed routine gives the tick loop a standing new reason to
speak. DESIGN.md § "surface caps authority" puts that at layer 3, and says
voice is structurally incapable of layer 3 because a room mic is reachable
by anyone in the room. The /routines button was gated at step-up; the voice
path accepted outright. The two surfaces disagreed, so one of them was wrong.

A spoken "да" now leaves the row 'proposed' and sends him to /routines,
where the gated button is. A spoken "нет" still dismisses: declining does
not move the boundary outward, so voice keeps it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018CotYKycuio1GwLbYh9jfc
2026-08-02 09:58:07 +04:00
claude 5e66aa8f22 fix: the slots converter was still dropping the fact Value (V-365)
dialogue.Slots gained Value in 925ce22, but toDialogueSlots never copied
it, so a clarifying answer carrying a fact payload still landed nowhere:
clarify.go:202 sends the answer through the converter, and the SlotValue
arm reads answer.Value.

Both converters now carry every field. TestSlotsParity compares the two
field sets by name and type; TestSlotsRoundTrip populates every router
field and checks the round trip, and fails the fixture itself when a new
field is left zero.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QChoBS5qJSrCV98oNUnHNU
2026-08-02 09:35:47 +04:00
claude e332f167b2 docs: the ranking's two blocking infra items already shipped (V-447)
Checked the doable, epic and infra tiers against the code, not just the two
tiers V-447 asked about. Ten entries are already built. The two the ranking
calls blockers for everything below are among them: sqlcipher at-rest ships as
Store.enc plus OpenEncrypted, and mavweb/mavcaldav have nine test files
between them where the ranking says zero coverage.

Also built and still ranked as work: rule trace, recurring reminders (cron +
RescheduleReminder), stale-reminder burst collapse (collapseReminders),
revert (VoidLatestFact), digest mode, testing infra, passkey persistence.

Recorded as a section at the top of the archived file so the tiers underneath
are read with the corrections in hand. No tasks created: V-447 scoped task
creation to the mandatory and easy tiers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xwomr2cT93KMSz9u3digX5
2026-08-02 09:20:36 +04:00
claude 322401b9af docs: the ranking was stale, two of its gaps already ship (V-447)
Checked the mandatory and easy tiers against the code instead of trusting the
2026-07-03 ranking. Two were already built and their tasks closed unstarted:
quiet hours (QuietHoursConfig + the care gate in internal/loop/loop.go:37) and
schema migrations (internal/store/migrations.go on PRAGMA user_version, 12+
steps shipped).

Three more were narrowed to what is actually missing. Destructive-confirm has
a mechanism and no policy: store.Tool.Destructive is one boolean, not a risk
tier. Bounded follow-up state has dialogue.Session with a TTL and slot
inheritance; what it lacks is Candidates, so "второй" resolves against nothing.
Clarify has a gate that can ask and one hardcoded sentence to ask with
(internal/voice/replier.go:56, which replier_llm.go hands straight through).

The other six were confirmed absent: list_items, capability model, go.mod
tidy, conversation repair, command history, pronunciation dictionary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xwomr2cT93KMSz9u3digX5
2026-08-02 09:13:59 +04:00
claude 4f34a232d4 docs: retire the two root queues, archive the ranking (V-447)
PROGRESS.md and 20-07-2026-BACKLOG.md were state snapshots that git log and
the Vikunja board already carry. Everything PROGRESS.md claimed as shipped is
a QA task. The backlog's only untracked item, bounded follow-up state, is now
V-448.

maven-feature-ranking.md moves to docs/archive/2026-07-03-feature-ranking.md
instead of dying. Its mandatory and easy tiers became V-449 through V-458; the
doable and epic tiers are reasoning about why things are not worth doing yet,
which no task captures.

Four code comments and the design.md ledger pointed at the deleted files.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xwomr2cT93KMSz9u3digX5
2026-08-02 09:11:11 +04:00
claude 93987f2dfc docs: tier the tree by lifetime, so staleness shows in the path (V-446)
Seventeen markdown files at the repo root, twelve of them dated one-shot
reports sitting next to CLAUDE.md. That is why stale docs read as
current: nothing in the path said which was which.

Root now keeps CLAUDE.md and AGENTS.md. Living docs move under docs/
and carry a Last verified line. Dated measurements move to docs/evals/
ISO-prefixed, and are never edited after the day, so a newer number is
a new file. The senior review moves to docs/archive/.

Every reference was rewritten across markdown, Go comments, the Makefile
and the recall fixture. The touched Go packages still build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 03:28:49 +04:00
92 changed files with 4998 additions and 1320 deletions
+7
View File
@@ -9,6 +9,7 @@
/mavwaked
/mavmaild
/mavupdate
/mavgpud
# Certs (private keys, don't commit)
certs/
@@ -40,6 +41,9 @@ deploy/telegram.env
deploy/zenmoney.token
# IMAP password, read by mavmaild (never in argv, never committed)
deploy/imap.password
# Compose interpolation secrets — MAVEN_AMBIENT_TOKEN today. docker compose
# reads this file itself; it is not an env_file on any service.
/.env
# Temp files
/tmp/
@@ -63,3 +67,6 @@ coverage.out
/HANDOFF.md
/models/stt
/models/tts
# root .env — MAVEN_AMBIENT_TOKEN and friends, same class as deploy/telegram.env
.env
-396
View File
@@ -1,396 +0,0 @@
beyond the model and tts work, the useful additions are mostly around **reliability, context, and reach**, not more intelligence.
## highest-value additions
### 1. unified event intake
maven should receive normalized events from:
* praxis
* calendar
* telegram
* local notifications
* system/service health
* manual checklists
* eventually email bridges
one internal envelope:
```go
type Event struct {
Source string
Kind string
EntityIDs []string
Title string
Body string
Priority string
OccurredAt time.Time
Payload json.RawMessage
}
```
this gives digestion one stable input instead of source-specific logic.
---
### 2. explicit morning routine engine — **core engine done (2026-07-20)**
`internal/morning` — pure checklist engine, mirrors `internal/loop`/
`internal/routine`'s no-I/O contract. `Evaluate(routine, facts, now)` answers
"what's still missing" any time (order-independent — checks facts, not
sequence); `Due(routines, facts, last, now)` fires the once-per-day nag only
at `NudgeAt` (defaults to window end) and only when something's unevidenced,
with a `last`-map dedupe identical in shape to `routine.Due`'s cold-start/
last-fire tracking. Evidence is just a fact timestamped inside today's
window — manual (voice-tapped) and inferred (another daemon writing the same
key) are indistinguishable, satisfying the manual/inferred requirement for
free. Weekday/weekend variants are two `Routine`s with different `Weekdays`
sets under different names. Wired into `config.MorningRoutineConfig` +
`cmd/mavend/tick.go`'s `fireMorningRoutines` (reads only the fact keys the
configured items reference, dispatches through the normal severity/presence
routing table, body is literal joined item labels — not LLM-phrased, same
no-hallucination rationale as cron routines). 13 unit tests in
`internal/morning/morning_test.go`.
Added since (2026-07-20, same day): a read-only `/morning` page in mavweb —
`ipc.CoreAPI.MorningStatus` (new wire method, mirrors `TickTrace`'s
daemon-cache-only shape: the store adapter errors, `daemonAPI` serves it from
a `tickLoop.morningStatus` closure) returns each routine's active/window/
per-item done state, server-rendered same as `/trace` (no live-update loop —
checklist state moves on minutes, not seconds).
Not yet done: no config wired in `deploy/mavend.json` (no morning routines
configured on homesrv yet — add items there when the medicine/water/pets
fact keys the phone/desktop write are settled), no voice query path for
"what did I miss this morning" (Evaluate supports it; nothing calls it yet),
no way to create/edit routines from the web UI — construction still means
hand-editing config, deliberately deferred: routines are operator-declared
config (like cron routines), and a CRUD editor would mean moving them to a
DB table + hot-reload, a bigger change than this pass.
not ordinary reminders.
support:
* required morning items
* order-independent completion
* soft time windows
* skipped-step detection
* one nudge, not repeated spam
* manual and inferred completion evidence
* weekend/weekday variants
example:
```text
08:0011:00
- medicine
- water
- pets
- check praxis attention
```
maven should know what is still missing, not merely fire four timers.
---
### 3. cross-device presence
**status (2026-07-20):** the hysteresis engine and 3 of the listed signals are
already built and wired live: `internal/store/presence.go` (noisy-OR combiner
+ Schmitt-trigger bucket resolve), fed by `desk_active` (workstation, via
`scripts/desk-active.sh` posting to `/api/signal`), `page_heartbeat` (mavweb
tab, `app.js`), and `wg_handshake` (`mavpoll` polling `wg show`) — threaded
into the tick loop via `internal/loop/gather.go`. Not done: phone-reachable,
homesrv-available, audio-output, and active-maven-client signals from the
list below are still missing.
a small presence daemon on each trusted device:
* workstation active/idle
* phone reachable
* homesrv available
* last keyboard/mouse activity
* wireguard presence
* current audio output
* active maven client
mavend receives only compact state, not raw activity logs.
useful for:
* choosing delivery channel
* suppressing voice while away
* surfacing reminders when you return
* knowing whether an agent result should be spoken or sent as text
---
### 4. interruption policy — **done (2026-07-20), turned out to already be built**
audited the existing code before writing anything new: `internal/loop.Gate`
already answers deliver_now vs. drop (quiet-hours/cooldown/snooze/presence/
calendar-busy), and `cmd/mavend/tick.go`'s `digestQ` + `config.DigestConfig`
already implement queue/digest (low-severity nudges batch into one
notification, flushed on window elapsed or max-items reached). The four
outcomes below were already covered by these two mechanisms; nothing new to
build for the core policy.
Gap that *was* real: `deploy/mavend.json` had no `digest` block, so batching
was disabled in prod despite being fully implemented. Fixed — see the config
change alongside this note.
before delivering anything, evaluate:
```text
urgency
current activity
quiet hours
recent nudges
available channels
whether already surfaced
```
result:
```text
deliver_now
queue
digest
drop
```
this prevents maven from becoming annoying once praxis and other sources start producing more data.
---
### 5. entity-aware memory — **done (2026-07-20)**
`03fa52d`/`9876187` (Vikunja #279): facts gain `Subject`/`EntityID`/
`ResolutionState`; an async enrichment worker resolves free-text subjects to
canonical Nexus entity_ids (mirrors Praxis's enrichment pattern). Ambiguous
or unreachable Nexus never guesses — the fact stays `pending` or terminal
`ambiguous`. Voice-tapped facts (`IntentFact`) now flow into the enrichment
queue automatically via an optional `Subject` field on `WriteFactReq` (old
callers unaffected).
Landed alongside this in the same session (not originally on this list, but
closes the plumbing gaps the last brief flagged for Nexus/Praxis maturity):
a typed Praxis lifecycle client (`398997f` — surface/acknowledge/resolve/
ignore/pin; fixes the surfaced≠acknowledged gap where reading an item aloud
left no trace), correlation-ID/version headers on the Nexus/Praxis clients
(`b743860`), entity-scoped Praxis attention queries (`0579ef9`), a durable
delivery outbox with begin-before-send/complete-after semantics
(`29f23e3`+`9ff726e` — closes a duplicate-send-on-crash bug), fail-closed
handling on ambiguous IPC mutation outcomes and Nexus/Hexis dependency
errors (`838fde1`+`d9fa4d6`), and a reusable fake-ecosystem test harness
with fault injection (`c932cd8`).
connect maven memory to nexus ids.
instead of:
```text
key = "кошачий фонтан"
```
store:
```text
entity_id = ent_pet_water_fountain
predicate = refilled_at
value = 2026-07-19T...
```
benefits:
* stable russian/english aliases
* fewer duplicate facts
* better “when did i last…” queries
* easier routine detection
* cleaner praxis correlation
---
### 6. bounded follow-up state
for short continuations:
* “yes”
* “tomorrow”
* “the second one”
* “not that project”
* “do it later”
store explicit pending state instead of relying on chat history:
```go
type PendingInteraction struct {
Kind string
Candidates []string
Args json.RawMessage
ExpiresAt time.Time
}
```
this matters a lot for a 1.7b model.
---
### 7. evaluation lab — **skipped for now (2026-07-20)**
runs on a different machine (GPU box), and CPT is currently in progress
there — deprioritized until the training pipeline has a checkpoint to gate.
Not abandoned, just off the immediate list.
before every new checkpoint or lora deploy:
* routing accuracy
* slot accuracy
* malformed json rate
* russian/english mixed input
* ambiguous entity handling
* reminder vs note vs fact
* direct answer vs tool call
* confirmation safety
* phrasing quality
* latency and ram
also replay real anonymized traces against old and new checkpoints.
this should be a hard deployment gate.
---
### 8. replayable full-system simulator
fake:
* clock
* presence
* caldav
* telegram
* praxis
* nexus
* hexis
* stt
* tts
* llama-server
scenario:
```text
08:30 user appears
08:35 medicine not completed
08:40 correx agent waits
08:45 calendar sync stale
08:50 user says “what did i miss?”
```
assert:
* what tools were called
* what was surfaced
* what stayed unresolved
* what maven said
* what was not executed
this will save more time than another feature daemon.
---
## useful second-wave additions
### voice session quality
* barge-in
* interrupt tts on wake word
* partial stt display
* confidence-aware clarification
* retry only failed stt segment
* per-room microphone profiles
* noise-floor calibration
* short response mode when speaking
### notification bridge framework
small adapters for:
* ntfy
* telegram
* matrix
* web push
* android notification forwarding
* local dbus notifications
normalize into maven/praxis events instead of treating each as a separate feature.
### local knowledge ingestion
* markdown/docs ingestion
* git repo summaries
* project decision records
* conversation exports
* provenance and source links
* incremental reindexing
keep this read-only and separate from personal fact memory.
### service self-diagnostics
`maven doctor`:
* socket reachability
* model health
* stt/tts readiness
* embedder availability
* caldav freshness
* telegram poll state
* praxis/nexus/hexis reachability
* db integrity
* disk usage
* recent failures
### config and secret management
* schema-validated config
* config migration
* secret references instead of inline values
* dry-run validation
* redacted config dump
* per-daemon health config
* startup dependency report
---
## things i would not build yet
* autonomous multi-step planning
* large external reasoner
* generic workflow engine
* self-editing memory
* automatic hexis actions from praxis
* emotion simulation beyond phrasing
* full home-assistant replacement
* more model layers before routing is stable
## recommended order
**status as of 2026-07-20:**
1. ~~evaluation lab~~**skipped, GPU-box work, deprioritized while CPT is in progress**
2. ~~entity-aware memory~~**done** (`03fa52d`/`9876187`, plus adjacent
Nexus/Praxis plumbing hardening — see item 5 above)
3. ~~morning routine engine~~**core engine done** (`internal/morning` +
`cmd/mavend` wiring — see item 2 above; not yet configured on homesrv,
no voice query, no web UI)
4. interruption/delivery policy
5. presence agents
6. unified event intake
7. full-system simulator
8. notification bridges
9. knowledge ingestion
10. voice-session polish
the main goal should be: **maven reliably knows what is happening, knows what you meant, and chooses the least annoying correct response**. everything else can wait.
+1 -1
View File
@@ -18,7 +18,7 @@ live in sibling repos next to this one.
Division of labour: Nexus identifies, Praxis observes, Hexis acts, Maven understands
and coordinates. Maven is not the source of truth for any of the three. The full
contract is `MAVEN_ECOSYSTEM_ARCHITECTURE.md`, and the constraints that bite during
contract is `docs/ecosystem.md`, and the constraints that bite during
implementation are summarised in `CLAUDE.md`.
Where things are in this repo:
+35 -10
View File
@@ -10,7 +10,7 @@ compose passes `/dev/dri` + the render gid) — the resident model stays ≤1.7B
**Resident model:** currently **Qwen3-1.7B** (`UD-Q4_K_XL`), stock — not yet the CPT'd one.
It replaced Qwen3.5-0.8B on 2026-07-31 because it measured better on both fixtures we have:
67.5% vs 59.7% intent-only on the 77-case RU routing fixture, and 20/27 vs 11-17/27 on the
talk fixture. See `MODEL-BAKEOFF-31-07-2026.md`. It is a Thinking variant, so `n_ctx` is 4096
talk fixture. See `docs/evals/2026-07-31-model-bakeoff.md`. It is a Thinking variant, so `n_ctx` is 4096
— reasoning tokens need the room, and 4096 is what the scores above were measured at.
The **target** is still the locally CPT'd **Qwen3-1.7B** (Vikunja #122, training in flight).
@@ -25,9 +25,24 @@ Spanish. Their strong published IFEval/BFCL numbers are English-only. Model file
`models/llm/`, so the LFM2.5 gguf sitting there is not loaded by anything. Swapping the resident
model is a one-line change to `phraser.model_path` in `deploy/mavend.json`.
See `REARCH.md` for the target architecture, `DESIGN.md` for the folded design spec, and
See `docs/rearchitecture.md` for the target architecture, `docs/design.md` for the folded design spec, and
`AGENTS.md` for local-preview + model-download recipes.
**Model work is moving to the workstation** (owner's call, 2026-08-02). homesrv cannot grow a
GPU and the workstation has 16GB of VRAM. So the resident model, STT and TTS become preferred
remotes with a floor on homesrv. The workstation is never assumed up. Fall back silently when
it would only do the job better. Name the gap when the 1.7B cannot do it at all. The embedder
stays on homesrv permanently, because it backs that floor. Read `docs/offload.md` before
touching a daemon seam or adding a model caller. Vikunja #483 is the umbrella, #484 to #487
are the work.
Both halves are wired as of 2026-08-03. Routing and replies prefer the workstation silently
through `modelSeam`; nudge and reminder phrasing prefer it silently inside the phraser. A
world question goes through `LLMPhraser.PhraseWorld` and names the gap when the card is not
free — `worldGap` in `cmd/mavend/worldmodel.go`, which he hears instead of an invented
answer. A box with no `workstation` block behaves exactly as it did before the seam: naming
a gap requires a gap. The offload table in `docs/offload.md` says which caller is which.
## Build & test
CGO daemons (`mavend`, `mavsttd`, `mavttsd`, `mavenclient`) need the vendored toolchain
@@ -72,7 +87,7 @@ protocol; the config in `deploy/mavend.json` (with `${VAR}` env expansion from g
Maven is one of four services. It owns conversation and personal memory. It does not
own identity, operational state, or execution. Full contract in
`MAVEN_ECOSYSTEM_ARCHITECTURE.md`.
`docs/ecosystem.md`.
```text
Nexus identifies. Praxis observes. Hexis acts. Maven understands and coordinates.
@@ -113,7 +128,7 @@ Every cross-service call carries a correlation id minted once per action
on in deploy** — this section used to say it was wired `nil`, which stopped being true on
2026-07-31.
- **LLM router (the intended design, REARCH.md):** the resident Qwen3-1.7B (`llmrouter.go`)
- **LLM router (the intended design, docs/rearchitecture.md):** the resident Qwen3-1.7B (`llmrouter.go`)
emits GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is
demoted from a routing gate to a RAG hint. Wired at `voice.go:214` via
`pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient)`; the flag is `voice.llm_router`
@@ -127,13 +142,23 @@ on in deploy** — this section used to say it was wired `nil`, which stopped be
Cascade order: `stage0.go` exact-match fast-path → LLM router (when non-nil) → classifier
fallback. Any LLM error falls through to the classifier so a turn never breaks on the model.
Measured on the 77-case RU fixture (`MODEL-BAKEOFF-31-07-2026.md`): the classifier scores
36.8% full accuracy at p50 31ms; Qwen3-1.7B scores 67.5% intent-only / 72.7% through the
cascade at p50 ≈825ms. Accuracy roughly doubled, latency is ~27× worse, and that trade was
accepted deliberately. **The ≈2.7s figure that stood here until 2026-08-02 was contention,
not the model.** See `ROUTING-EVAL-31-07-2026.md` line 61, which measures the LLM router at
Measured on the 77-case RU fixture. **Re-measured 2026-08-02: the classifier scores 68.8%
full accuracy at p50 16.6µs**, not the 36.8% at p50 31ms that stood here from
`docs/evals/2026-07-31-model-bakeoff.md`. That older figure predates the stage 0 rules and the
seed additions, both of which now score inside the classifier baseline. Qwen3-1.7B scores
77.9% intent-only / 72.7% through the cascade. So the router buys about 4 points of accuracy,
not a doubling, and the trade is worth re-arguing rather than assuming. **The ≈2.7s figure
that stood here until 2026-08-02 was contention, not the model.** See `docs/evals/2026-07-31-routing.md` line 61, which measures the LLM router at
p50 825ms / p95 1.2s / max 3.0s and the full cascade at p50 0.80-1.04s. Do not plan latency
work off the bakeoff table. `Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
work off the bakeoff table.
**The numbers above are the homesrv floor, not the ceiling.** With the workstation up, routing
completes through `llm.Pair` against gemma-4-12b and scores **84.4% full / 93.5% intent-only at
p50 329ms** — better than the resident model and about 2.5× faster (`docs/evals/2026-08-02-workstation-gemma4-12b.md`,
Vikunja #485). The workstation is never assumed up, so both sets of numbers are live. Judge a
routing change against the classifier and the resident model, since those are what always answer.
`Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
#359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with
no allowlisted fn) feeding the same stage-3 gate the classifier path already had — see
+10 -4
View File
@@ -16,11 +16,11 @@ PIPER_BIN := $(shell pwd)/deps/piper/piper
PIPER_MODEL := $(shell pwd)/models/tts/ru_RU-irina-medium.onnx
PIPER_ESPEAK := $(shell pwd)/deps/piper/espeak-ng-data
.PHONY: simulate stt-fixtures test-stt-golden all build build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav clean test fmt-check vet run-stt run-tts run-web download-embedder deps-go eval-router eval-recall eval-phrasing eval-models
.PHONY: simulate stt-fixtures test-stt-golden all build build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav clean test fmt-check vet run-stt run-tts run-web download-embedder deps-go eval-router eval-recall eval-phrasing eval-models build-gpud
all: build
build: build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav build-mail build-update
build: build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav build-mail build-update build-gpud
build-stt:
CGO_CFLAGS="$(CGO_CFLAGS)" CGO_LDFLAGS="$(CGO_LDFLAGS)" LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
@@ -59,6 +59,12 @@ build-mail:
build-update:
$(GO) build $(GOFLAGS) -o mavupdate ./cmd/mavupdate/
# mavgpud runs on the workstation, not here. It is built with the rest so a
# broken supervisor is caught by `make build` on homesrv rather than by the
# workstation refusing to serve. Copy the binary over, do not `make deploy` it.
build-gpud:
$(GO) build $(GOFLAGS) -o mavgpud ./cmd/mavgpud/
run-web: build-web
./mavweb -addr :9200 -voice 127.0.0.1:9100
@@ -78,7 +84,7 @@ deps-go:
done
$(GO) version
# fmt-check fails if any file needs gofmt. DESIGN.md has always said `make
# fmt-check fails if any file needs gofmt. docs/design.md has always said `make
# test` gates on gofmt and vet; it did not, so nine files quietly drifted.
# Run `gofmt -w` on whatever this prints.
fmt-check:
@@ -191,7 +197,7 @@ deps-piper:
# multilingual-e5-small: an asymmetric retrieval model. It is trained to match
# a short question against a longer passage, which is what note recall is.
# The quantized file is the one we download, deploy and measure — see
# RECALL-EVAL-31-07-2026.md.
# docs/evals/2026-07-31-recall.md.
EMBEDDER_DIR := $(shell pwd)/models/embedder/multilingual-e5-small
EMBEDDER_MODEL_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/onnx/model_quantized.onnx
EMBEDDER_TOKENIZER_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/tokenizer.json
-468
View File
@@ -1,468 +0,0 @@
## Maven — current state (updated 2026-07-20)
### Session 2026-07-20 — ecosystem hardening + entity-aware facts
Ten commits, focused on closing the Nexus/Praxis integration gaps flagged
as "wired but immature" in the prior review, plus the entity-aware-memory
backlog item (`20-07-2026-BACKLOG.md` item 5).
- **Entity-aware fact resolution (Vikunja #279)** — facts gain
`Subject`/`EntityID`/`ResolutionState`; an async worker resolves
free-text subjects to canonical Nexus entity_ids (mirrors Praxis's own
enrichment pattern). Ambiguous/unreachable Nexus never guesses — stays
`pending` or terminal `ambiguous`. Voice-tapped facts (`IntentFact`) flow
into the queue automatically via an optional `Subject` field on
`WriteFactReq` (old callers unaffected, no signature break).
- **Typed Praxis lifecycle client (Vikunja #271)** — `GetItem`/`Search`/
`Surface`/`Acknowledge`/`Resolve`/`Ignore`/`Pin`, routed through new RU/EN
dialogue verbs. Fixes a real lifecycle-invariant bug: reading an
attention item aloud now calls `Surface` — previously the digest path
read items without recording that they'd been surfaced, so "Maven
mentioned it" was indistinguishable from "never came up."
- **Durable delivery outbox (Vikunja #270)** — `BeginDeliveryAttempt`
before `Send`, `CompleteDeliveryAttempt` after; a stale `pending` row
found at startup reconciles to `unknown` (never silently resent or
dropped — same rule as Hexis's execution-timeout handling). Closes a
crash-window duplicate-send bug. Wired into `DispatchNudge`,
`DispatchReminder`, `RepeatUnacked`; reconciliation runs once at boot
before the tick loop resumes.
- **Fail-closed IPC/dependency handling (Vikunja #269, #272/#273)** —
ambiguous mutation outcomes (frame sent, reply lost) no longer blindly
retry; Nexus/Hexis dependency errors fail closed instead of guessing.
- **Correlation IDs + version headers (Vikunja #273)** — the hand-rolled
Nexus/Praxis HTTP clients now send `X-Nexus-Version`/`X-Praxis-Version`
and thread the same correlation ID already generated in
`executeCapability` through the whole call chain, matching the Hexis
client's existing behavior.
- **Entity-scoped Praxis attention queries** — callers holding a resolved
entity_id can ask "what needs attention for this entity" directly
instead of filtering the unscoped list client-side.
- **Fake-ecosystem test harness with fault injection** — a reusable
`fakeServer` (Nexus/Praxis/Hexis fixtures, runtime-toggleable
`SetFault`, fake clock) replacing ad-hoc per-test `httptest` servers;
covers a gap that had zero test coverage (`handlePraxisAct`) and adds a
fault-then-recovery regression test for the fail-closed fixes above.
- **Ops fix** — `deploy/mavend.json`'s phraser was pointed at a 4B model
with `n_gpu_layers=99`, which OOM'd under memory pressure and left a
zombie `llama-server` child; swapped to the 2B Qwen model matching the
intended resident-model size.
Net effect: the Nexus/Praxis wiring described as "plumbing exists, thin
compared to Maven's test depth" in the prior review is now materially
hardened — typed clients, fail-closed error handling, durable delivery,
and a proper fault-injection test harness are all in place. Evaluation lab
(`20-07-2026-BACKLOG.md` item 7) is explicitly skipped for now — it runs
on the GPU box, which is occupied by CPT. Morning routine engine (backlog
item 3) is next up, not started.
---
> **Resolved 2026-07-30 (task #318).** The resident checkpoint is
> **Qwen3.5-0.8B** (`Q4_K_M`), set in `deploy/mavend.json`; the **target** is
> the locally CPT'd **Qwen3-1.7B**, still training (#122). Older model claims
> below — the LFM references, the pipeline line, and the "swapped to the 2B
> Qwen model" ops entry above — are historical. Read them as a log of what was
> true at the time, not as current fact. Note also that `/mnt/hdd1/llms` is
> bind-mounted over `models/llm/`, so the LFM2.5 gguf in the repo tree is
> never loaded.
Architecture decision (as written on 2026-07-20): the target resident
router/phraser is the locally trained Qwen3-1.7B model — still the target as
of 2026-07-30. Older LFM references below describe the then-deployed
historical stack, not the target checkpoint. RU CPT has a successful
full-weight checkpoint at step 1000/8077; evaluation and Qwen3 SFT tooling are
tracked in `docs/plans/2026-07-18-qwen3-resident-training-eval.md`.
Consolidated status. The reactive↔proactive core is closed and testable through
the web PWA. The former SPEC's open items 17 (now `DESIGN.md` § execution ledger) are landed (protocol doc, away-channel
fallthrough, CalDAV poller, quiet-hours schedule, tools enable/disable, note RAG,
passkey step-up); item 8 (multi-user) is deliberately deferred — see the tail.
The two big infra gaps from the jul5 revision are closed on `overnight-jul5`:
**at-rest encryption** (AES-256-GCM, tmpfs working copy — not sqlcipher, see
`internal/store/crypt.go`) and **Docker deployment** (one image, six daemon
containers). The `overnight-jul6` session (now on `master`) closed the biggest
*query-surface* gaps — **calendar querying, general-knowledge answers, and
weather** — plus a populated homelab act allowlist and two pure scaffolds
(dialogue state, long-term-memory vector store). ~15.2k LOC + ~8.5k test, 303
tests, `-race` in `make test`.
### Access model
- **Phone** → needs the wg tunnel to reach homesrv (no homesrv DNS otherwise;
raw IP or a DNS tweak can bypass, not the default).
- **PC** → uses homesrv DNS, resolves the domains over local-net, **no wg needed**.
- nginx + ufw both scope to `10.42.0.0/24` (wg) + `192.168.1.0/24` (LAN), deny all else.
- **Surface in use now: the web PWA (`mavweb`).** Voice PTT + in-app nudges both ride it.
### Works end-to-end (tested)
- **Reactive voice:** PWA record → Whisper STT (`mavsttd`) → ONNX classifier →
resident phraser (llama-server subprocess; Qwen3.5-0.8B as of 2026-07-30 —
this line historically named "LFM 2.5-1.2B") → Piper TTS
(`mavttsd`) → reply.
HTTP POST path (mobile-Chrome drops WS for the audio).
- **Capture:** `fact` (EN **and RU** — root-substring recognizers) + `reminder`
persist through CoreAPI (`source=tap:voice`). This is the substrate the care
rules read.
- **Notes / query (semantic recall, sqlite — no chroma):** `note` → embed (the
classifier's ONNX embedder) → `notes` table. `query` → embed → brute-force
cosine top-k → confidence-gated (below `queryMinScore` 0.55 ⇒ "no note", not a
guess). **Note RAG (SPEC item 6):** the gated top-k feed the phraser
(`PhraseQuery`) to compose a natural answer ("вот что я нашла: …") instead of
a verbatim dump; raw-notes fallback on any LLM error. Stub is deterministic.
- **Monitoring (`/dash`):** mavweb server-renders presence + recent nudges (by
outcome) + recent facts from the append-only store via CoreAPI. Read-only,
meta-refresh, no JS.
- **Proactive loop:** 60s dumb ticker, pure predicates over a State snapshot,
universal gate (quiet-hours/presence/cooldown/snooze/calendar), one-nudge-per-
tick max-severity, reminders (gate-bypassing), sev4 repeat-til-ack, feedback
auto-tuner (outcome ratio → bounded cooldown, persisted as `source=feedback`).
- **Rules:** water/meal/break (sev12 care), service_down (sev4, `poll:uptimekuma`),
netdata_critical (sev3, `poll:netdata`).
- **Routines (`internal/routine`):** operator-declared clockwork — the third
proactive class beside reminders (user-stated) and care rules (world-state).
Config `routines[]` (cron + literal RU body + severity) fire through the normal
dispatcher on schedule (an 08:00 briefing, a 22:00 wind-down). Bodies are
literal (not LLM-phrased ⇒ can't hallucinate); rule name `routine:<name>` so
they don't pollute the care autotuner; cold-start guard seeds on first sight so
a restart never replays a missed schedule. Pure `routine.Due`, unit-tested; the
tick driver holds the last-fired map.
- **Env facts (`mavpoll`):** netdata alarms → `netdata_alarm` (fires immediately
on a real CRITICAL); kuma monitor_status → `service_down`. Writes only on
value-change (no append-only churn).
- **Presence:** noisy-OR decay + Schmitt hysteresis. Live via `page_heartbeat`
(PWA auto-pings `/api/signal` every 30s → present when a tab's open).
- **Delivery:** ntfy / telegram / voice by `f(severity, presence)`; minimal body
on away channels. PWA subscribes to ntfy over **WebSocket** for in-app nudges.
- **Away-channel fallthrough (SPEC item 2):** when the router picks voice but no
live session exists at push time (presence guess was wrong), the dispatcher
reroutes through the AWAY table — sev3→ntfy, sev4→telegram-repeat-til-ack,
sev≤2→drop — instead of silently dropping. Covers nudges + reminders.
- **Calendar busy (SPEC item 3, `mavcaldav`):** new poller queries a self-hosted
**Radicale** CalDAV server on an interval, writes `calendar_busy` + event facts
through CoreAPI (value-change only). The loop gate already consumes `calendar_busy`.
- **Quiet-hours schedule (SPEC item 4):** the gate reads `quiet_hours`; a config
time window (`voice.quiet_hours`, HH:MM, midnight-crossing handled) now sets it
on each tick — in addition to the "тихий режим" voice toggle. Both activate quiet.
- **Client protocol (SPEC item 1):** the voice wire format (length-prefixed JSON
frames) is published in `PROTOCOL.md`, generated from `internal/voice/wire.go`
so third-party clients don't need the Go source.
- **Passkey step-up (SPEC item 7):** `internal/webauthn` does real WebAuthn —
ES256/P-256 register + assert, ecdsa signature verification, rpIdHash + UP/UV
flag binding (UV = the gesture), sign-count regression check. `PasskeySession`
bumps the auth session L2→L3 for a TTL on assert. mavweb serves `/auth/passkey`
(enroll + step-up) + the begin/finish endpoints. Crypto is round-trip tested
(incl. tampered-sig / missing-UV / wrong-origin negatives).
- **Stability:** llama-server orphan leak fixed (`Pdeathsig` kills the child on
any mavend death); `kill-maven.sh` reaps strays (matches the model, not a
bogus `llama-server.*maven` pattern); `start-maven.sh` wires `-core` + poller.
### Wired but needs a deploy action (not code)
- **`desk_active`** (strongest presence signal) — `scripts/desk-active.sh` runs
on the **desk PC** (hypridle-gated systemd timer), posts over wg to mavweb.
- **`mavwaked`** (always-on listening) — needs a systemd user unit on a client
box (desk PC, pi, etc.) where the mic is attached. Connects to mavend over wg
or local net via `-addr`. Deferred until a client box is wired with a mic.
Caveats / gotchas:
- **desk_active is a workstation deploy, not code** — 0 facts ever written; presence
runs on page_heartbeat alone (dash reads "away"/"never at desk"). `scripts/desk-active.sh`
+ a hypridle-gated `maven-desk` timer must be installed on the desk PC (not homesrv).
- **Notes recall needs the ONNX embedder** — under the HashEmbedder floor, cosine is
lexical (token overlap), not semantic; scores are low, so most RU commands sit under
the 0.35 route threshold and clarify. Configure `voice.embedder` for confident recall+routing.
(The floor now at least tokenizes Cyrillic — see below — so it ranks correctly, just weakly.)
- **Switching the embedder model silently breaks old notes** — different dim ⇒
cosine 0 ⇒ they stop matching; brute-force can't re-embed. Re-embed on a model change.
- **`wg_handshake` is OFF and should stay off** — in this topology the phone only
runs wg when *outside*, so a fresh handshake means AWAY, not here. The `mavpoll
-wg` flag exists (defaults `""`) and could later back the spec's "away override"
by flipping the sign; as a presence-*here* signal it's inverted. desk_active +
page_heartbeat cover home presence.
- **Cold-start unlock tests are missing** — the key wrap/unwrap code
(`internal/webauthn/keywrap.go`) and locked-mode IPC gating (`cmd/mavend/main.go`)
are correct but have **zero test coverage**. The roadmap (item 2.1) required
three new test cases (wrap/unwrap round-trip, wrong-cred unwrap fails,
locked-mode IPC rejects non-unlock methods); none were written. `make test`
is green by omission. Write these before relying on the cold-start path with
real keys.
### Done since last revision (overnight-jul6, 2026-07-06)
Seven tasks (session board `SESSION-06-07-2026.md`, deleted 2026-07-30 — see git history), one commit each, merged to `master`.
This session was run through **opencode**, not Claude Code (co-author trailer).
Since then (**2026-07-06, second session**):
- **Always-on listening (gap 1, MVP)** — `cmd/mavwaked/`: 825 lines, 10 `-race`
tests. Energy-based VAD over 30ms windows (same RMS threshold as mavsttd's
`gateReason`), adaptive noise floor, speech→silence state machine. Captures
PCM from arecord(1) subprocess, sends `PushToTalk` with `Surface=SurfaceVoice`
(L0 — no destructive acts). Reply plays through aplay(1). No wake word yet
(pure VAD trigger); the 30ms frame shape matches silero-vad ONNX input 1:1,
so swapping energy-threshold for ONNX inference is a local change in vad.go.
`Makefile` `build-waked` target. Runs on client boxes (not docker/homesrv)
via systemd user unit; connects to mavend over wg or local net.
Since then (**2026-07-06, third session** — roadmap execution agent):
- **Cold-start unlock (ROADMAP 2.1)** — the at-rest AES key is now wrapped
(HKDF-SHA256 + AES-256-GCM, stdlib-only — no `x/crypto` dep) with the passkey
credential's public key and persisted to disk. At boot, if a wrapped key file
exists AND no env key is set, mavend starts **locked**: the IPC server runs
but `srv.Check` rejects everything except `MethodAssertStepUp` +
`MethodUnlock`. A passkey assertion at `/auth/passkey` calls `MethodUnlock`
with the credential's public key → unwraps the blob → opens the store → wires
voice/loop/delivery → `srv.SetAPI` swaps the locked stub for the real
CoreAPI. mavweb's `RegisterFinish` wraps the env key on enrollment;
`AssertFinish` calls `Unlock` on assertion. Env-key fallback preserved
(dev/CI path unchanged). **Test gap:** the roadmap required three new test
cases (wrap/unwrap round-trip, wrong-cred unwrap fails, locked-mode IPC
rejects non-unlock methods) — none were written. The code is correct but
untested; `make test` is green by omission, not coverage.
- **Conversation depth (ROADMAP 3.2)** — cross-intent anaphora + fact-by-key
lookup. `AnaphoraResolver` in `router/slots.go` detects RU pronouns
(это/он/она/оно/тот/мой + inflected forms). `followUpMerge` now handles
three cases: same-intent slot inheritance (existing), cross-intent anaphora
(Query/Fact/Reminder after a Fact with a pronoun inherits the prior key +
time), and query-after-fact (a query following a fact inherits the key for
fact-by-key lookup). `Session.History []Turn` added as the multi-turn
scaffold (capped at 4). 7 new test cases including the exact done-when
scenarios (anaphora query-after-fact, three-turn break, explicit-key-wins).
- **Routing quality + persona (ROADMAP 4.1/4.4)** — `QueryMinScore` is now a
config knob (`voice.query_min_score`, default 0.55) instead of a hardcoded
const. `make download-embedder` fetches Xenova/paraphrase-multilingual-
MiniLM-L12-v2 (~90MB ONNX) + tokenizer; AGENTS.md documents the embedder +
libonnxruntime setup. `Persona` field in `VoiceConfig` prepends to every
LLM system prompt (nudge phrasing, note queries, general knowledge); empty
= current hardcoded feminine-gendered Russian persona. Also fixed two
pre-existing data races found by `-race`: `voice/server.go` wg.Add vs
wg.Wait (accept mutex), `mavweb/server.go` s.api field (atomic.Value).
- **Calendar querying (task 3)** — "что у меня завтра?" now answers from the
CalDAV facts the poller already writes. Added `store.CalendarEvents(from,to)`,
a RU date-scope parser («сегодня»/«завтра») in `router/slots.go`, and an
IPC `CalendarEvents` RPC (api/client/server/wire) feeding the `IntentQuery`
handler. Empty day → «на сегодня ничего нет». Previously calendar only *gated*
nudges; it's now queryable.
- **General-knowledge routing (task 4)** — when notes-RAG misses `queryMinScore`,
the query now falls through to the phraser with an anti-hallucination system
prompt (`router.KnowledgePrompt`, single tested source) instead of giving up.
Empty/errored/Stub phraser → «не знаю.», never a fabrication.
- **Weather (task 5)** — new `internal/weather/`: `Provider` interface, a stub
(«погода не настроена»), and a real **keyless Open-Meteo** provider (geocode +
current_weather, injectable `*http.Client`, mocked in tests — no live network).
Wired into `IntentQuery` (keywords погода/градус/температура) with a ~5s
context timeout; selected by `voice.weather.provider` ("open-meteo" | "" → stub).
- **Homelab act allowlist (task 2)** — `voice.tools` seeded with read-only acts
(`systemctl status`, `docker ps`, `uptime`, `df`, `free`, `journalctl` reads)
as `destructive:false` and mutating ones (restart/stop/start/reboot,
docker-restart/stop) as `destructive:true`. Guardrail verified: no dangerous
verb is `destructive:false`. RU phrasings seeded in `act.txt`.
- **Embedder config validation (task 1)** — a partially-filled `voice.embedder`
block (some of model/tokenizer/lib paths missing) is now a load error instead
of a silent fall-through to the Hash floor; the floor fallback logs explicitly.
- **Dialogue state scaffold (task 6)** — `internal/dialogue/`: `Session` +
TTL `SessionStore` + pure `InheritSlots`. **Now wired** (post-merge follow-up):
the voice handler carries slots across same-intent turns within a 2-min window
(`followUpMerge`, unit-tested) — bounded gap-filling, not full multi-turn yet.
- **Long-term memory interface (task 7)** — `internal/memory/`: `Store` interface
+ `InMemoryStore` (cosine). Wired into `IntentNote` (best-effort insert) and,
post-merge, into `IntentFact` (facts indexed) + `IntentQuery` (read-back after
notes-RAG misses). In-memory only — no persistent backend yet (gap #8).
Follow-ups (Claude Code, post-merge): gofmt'd `handlers_test.go` (the jul6
verification commit left it misaligned, so `gofmt -l` still flagged it despite the
"all gates green" claim); deduped the task-4 knowledge prompt to the single tested
`router.KnowledgePrompt()`. Tree is now genuinely green (gofmt/vet/303 tests).
### Done since the jul5 revision (overnight-jul5, 2026-07-05)
The overnight session (`SESSION-05-07-2026.md`, deleted 2026-07-30 — see git history; 25 tasks) closed the previous
"not built yet" items 13 and added feature depth:
- **At-rest encryption** — the on-disk db is AES-256-GCM ciphertext; the daemon
works on a tmpfs (RAM) plaintext copy, sealed back atomically on close. Wrong
key / tamper ⇒ fail closed, never a plaintext fallback. Legacy plaintext dbs
upgrade on first clean shutdown. Key via config/env (`db_key_env`); no KDF —
raw 32-byte key, base64. The passkey cold-start unlock plugs into the same
`store.OpenEncrypted` seam later.
- **Docker deployment** — single image, one container per daemon
(`docker-compose.yml`); only mavend mounts the key + db volume; IPC over a
shared socket volume. `ipc.DialWait` (boot-order tolerance) + redial-on-drop
(core restarts don't kill modules). `deploy/README.md` has the runbook.
- **Tests** — mavcaldav, mavttsd, voicesink, mavweb main/handlers covered;
`make test` runs `-race -coverprofile`.
- **Recurring reminders** — `cron` + `next_fire_ts` on reminders; recurring ones
reschedule (instead of mark-fired) after successful delivery.
- **Notification digest/batching** — low-severity nudges queue and flush as one
digest per window/max-items (`digest` config block); stale-reminder bursts on
boot collapse into a single digest reminder, completed only after delivery.
- **Rule trace engine** — `ExplainTick`/`ExplainGate` record per-rule
predicate/gate/selection results each tick; served over IPC (`tick_trace`)
and rendered at mavweb `/trace` ("why didn't she nudge me").
- **Web UI** — new `/history` (facts + revert buttons), `/notifications` (nudge
history), `/trace` pages; nav links on `/dash`; RU/EN cheatsheet toggle in the
PWA; manifest icons (`icon.svg`). POST `/tools` now requires an in-process
passkey step-up when WebAuthn is configured.
- **Revert/undo** — `RevertFact` voids the latest fact for a key (append-only
void-marker, audit trail intact); exposed at `/api/revert` from `/history`.
- **Tool scopes** — `scope` column on tools, threaded through propose/enable/UI.
`DisableTool` raised to AuthStepUp alongside Enable.
- **Passkey persistence** — mavweb credentials in a JSON file (`-passkey-file`),
surviving restarts; rollback-on-persist-failure keeps memory and disk in sync.
- **STT silence gate** — min-duration + RMS floor drop non-speech before whisper
hallucinates on it (`-min-ms`, `-silence-rms` flags on mavsttd).
- **Housekeeping** — `db_key.env` gitignored (+`.env.example`), `build-caldav`
target, zero-timestamp "never" fix on /dash.
### Not built yet (ranked by ROI)
1. **Multi-user (SPEC item 8)** — deliberately deferred, see the tail.
Closed (jul6 follow-ups): `/api/revert` now sits behind the same passkey
step-up as POST `/tools`; `go.mod` direct deps (`onnxruntime_go`,
`coder/websocket`, `robfig/cron`) are labeled correctly — `go mod tidy` can't
run here because it walks the vendored `deps/go` toolchain tree.
Purge+rotate leaked db key (#12) — investigated and closed: the key was
**never committed** to git history (gitignored at introduction, no commit
ever tracked `deploy/db_key.env`), so nothing to scrub. File stays on disk
and in deploy env by design — at-rest encryption needs it at boot.
Done earlier (2026-07-03): **act tool executor, store-backed, full flow**
(`internal/tool` + `internal/store/tools.go` + `tools` CoreAPI methods).
- **Execution:** IntentAct runs the matched fn against the store's ENABLED
allowlist. argv, no shell → STT text can't inject. Live store read, so a
newly-enabled tool runs without a daemon restart.
- **proposed→enabled→disabled (SPEC item 5):** an act whose verb isn't enabled is
scaffolded as a `proposed` tool (maven suggests). A human enables it (fills argv
+ destructive) on the authed **`mavweb /tools`** page — never voice — and can
disable it back to `proposed` (kept in the store, won't run). `EnableTool`/
`DisableTool` sit at `AuthStepUp`; the gate is now **live** via `PasskeySession`,
so /tools enable requires a passkey assertion at `/auth/passkey` first.
- **Confirm turn:** a destructive enabled tool replies "выполнить X? да/нет" and
parks; the next utterance (ru/en yes-no) confirms or cancels (90s TTL).
- **Config:** `voice.tools` seeds enabled tools at boot (editing mavend.json =
the human enable act); mavweb enables ad-hoc ones on top.
- **Russian:** fixed grammar in reply strings + seed files; maven's self-
reference is feminine ("she") — [[maven-persona-gender]].
Also fixed:
- **HashEmbedder was blind to Cyrillic** (`tokenize` iterated bytes, kept only
`a-z0-9`) → every RU utterance embedded to the zero vector → cosine 0 across
all intents → misrouted to `act` (alphabetical tie-break). Now rune-based
(`unicode.IsLetter`). This was the real cause of "Найди заметку" (a query)
landing in `notes`; added note-retrieval query seeds too.
- **Notes are now browsable on `/dash`** — `RecentNotes` plumbed through the
store + CoreAPI; voice-captured notes were previously only reachable via
semantic `query`.
Earlier: notes/query recall, `/dash` monitoring, `wg_handshake` poller (NO-OP).
### Gaps — why "voice assistant" is still aspirational (2026-07-06)
What separates Maven today from the thing the spec describes. Dealbreakers
first — these define the category:
1. **Always-on listening is code-complete (MVP).** `cmd/mavwaked` captures
PCM from arecord → energy-based VAD → PushToTalk with `Surface=SurfaceVoice`
(L0). Gap narrowed: no wake word yet (pure voice-activity trigger; every
utterance fires). The 30ms frame shape and 16kHz PCM match silero-vad's
ONNX input exactly, so a wake-word model swap is a local change in vad.go.
Hardware: the mic lives on a client box (desk PC, pi, etc.) — never the
homesrv. Deploy action: systemd user unit on whichever box has the mic,
connects to mavend over wg or local net.
2. **Conversation is deeper now, still not full dialogue.** The router
classifies one utterance → one reply, but `internal/dialogue` carries
context across turns: a 2-min session inherits slots for same-intent
follow-ups («напомни завтра» → «…позвонить маме»), and cross-intent
anaphora («запиши что я пил воду» → «когда я это сделал?») now resolves
RU pronouns (это/он/она/оно/тот/мой + inflections) to the prior turn's
key for fact-by-key lookup. `Session.History []Turn` is the scaffold for
real multi-turn. Still missing: LLM-driven dialogue manager (decide
ask-vs-act), anaphora beyond RU pronouns, single-slot session (single-user
box). The sub-1B phraser only words replies.
3. **Latency/shape of a turn.** Clip-based STT (record → upload → whisper →
route → phrase → piper → play). No streaming either direction, no barge-in;
every exchange is a full round trip.
Capability-class gaps — built but thin:
4. **Act surface is a small argv allowlist.** propose→enable works and the
allowlist now ships a homelab starter set (jul6 task 2 — status/ps/uptime/
df/free/logs read-only, restart/stop/reboot gated). Still bounded to what's
seeded; broadening it is config, not code.
5. **Query answers now cover notes + calendar + weather + general knowledge**
(jul6 tasks 3/4/5). Calendar querying, keyless Open-Meteo weather, and a
phraser knowledge-fallback all landed; caveat — general-knowledge quality is
only as good as the sub-1B phraser, and weather needs `voice.weather.provider`
set. The cheatsheet and router are now roughly aligned.
6. **Routing quality depends on the ONNX embedder being configured** — the
HashEmbedder floor makes RU recall lexical/weak; many commands fall to
"clarify". `make download-embedder` now fetches the multilingual MiniLM
model + AGENTS.md documents libonnxruntime setup; `voice.query_min_score`
is a config knob (default 0.55) so the floor can be tuned without recompile.
7. **Presence is effectively one signal** (page_heartbeat); desk_active is
still an undeployed script — "voice when near" routing runs on a guess.
8. **Long-term memory is now persistent (store-backed), not the spec's chroma.**
`internal/memory` has a `Store` interface; the daemon now wires
`store.MemoryStore` (`internal/store/memory.go`) — a **persistent** backend
in the **same encrypted sqlite db** (survives restarts; recall text inherits
at-rest encryption, so no plaintext sidecar). Vectors are float32 blobs,
search is brute-force cosine (fine at single-user scale; ANN is the later
swap behind the same interface). Notes **and facts** are indexed on capture;
`IntentQuery` reads it back (after notes-RAG misses, before general-knowledge)
— fact recall («когда я пил воду?») is its distinct payoff. The in-memory
impl remains the test/no-store floor. Remaining: an ANN/external index is
optional-scale, not a gap. Custom TTS voice (kami-picked, replaces the irina
floor — [[custom-voice-training]]) is still a future item.
Ops footnote: voice-over-web verified 2026-07-06 — mavend binds 0.0.0.0:9100
and mavweb reaches it cross-container at mavend:9100 (nc -z confirmed).
mavpoll uses network_mode=host to reach localhost services (netdata, kuma).
### Future / logged, not now
Custom TTS voice training (kami-picked voice, replaces irina floor); listening
modes 23 (meeting-record, ambient-derive).
### Services & layout
- `mavend` (core, IPC unix socket) — store + loop + phraser; the only key-holder.
- `mavsttd` / `mavttsd` — STT/TTS worker modules (unix sockets).
- `mavweb` — PWA bridge (HTTP), `/api/ptt` voice, `/api/signal` presence ingest,
`/api/ntfy` WS-subscribe config, `/dash` read-only monitoring.
- `mavpoll` — env poller (netdata/kuma → facts via CoreAPI).
- `mavcaldav` — CalDAV poller (Radicale → `calendar_busy` + events via CoreAPI).
- All behind wg + nginx deny-all; no phone-home. CGo only in `mavsttd`.
- Start/stop: `./start-maven.sh [build]`, `./kill-maven.sh`.
- Config: `~/.config/maven/mavend.json` (or `mavend.json` in repo root).
### Key files
- `cmd/mavend/{main,tick,voice}.go` — daemon wiring, loop driver, voice handler
- `internal/loop/{loop,rules,gather,feedback}.go` — proactive engine
- `internal/store/` — append-only facts/reminders/nudges/presence/notes
- `cmd/mavweb/{main.go,dash.html}` — PWA bridge + `/dash` monitoring
- `internal/router/{classifier,slots,stage0}.go` — reactive routing + slot parse
- `internal/delivery/` — dispatcher + ntfy/telegram/voice sinks
- `internal/auth/` — scope/gate/policy; `FloorEnrollment` (same-uid = device
trust) + `webauthn.PasskeySession` (real step-up for L3)
- `internal/webauthn/`, `cmd/mavweb/webauthn.go` — passkey register/assert
- `cmd/mavcaldav/`, `cmd/mavpoll/`, `scripts/desk-active.sh` — env producers
### Why multi-user (SPEC item 8) is deferred
Not neglect — the one item where doing nothing now beats doing something:
- **No second user exists yet** (the "gf phase"). Building per-user partitioning
now means code exercised by zero users and validated by nobody — YAGNI.
- **The append-only schema makes it a migration, not a rewrite.** No row is ever
mutated, so adding `facts/notes/reminders.user_id` later is add-columns +
backfill-to-"kami" — no reshaping, no dual-write window. Deferral is cheap.
- **The hard part is speaker attribution, and it needs the second voice.** A
voice-print discriminator (kami vs gf vs unknown) can't be trained or tuned
with one voice in the house. Plumbing before the model is pipe with no water.
- **It's fenced deliberately** (`DO NOT TOUCH THIS PHASE` in `DESIGN.md` § Users) so an
autonomous agent doesn't add `user_id` columns while touching the store and
commit us to a schema before the constraints that shape it exist.
-184
View File
@@ -1,184 +0,0 @@
# QA plan: checking Maven properly
Written 2026-08-01, after the 35-PR stack landed and the box came back up.
44 of the 50 open Vikunja tasks are `QA:` tasks. They are verification work, not
build work. Most sat unverifiable while Maven was down for 11 days. That
blocker is gone.
This plan orders them by what unblocks what. Do sessions 1 and 2 first. Almost everything
downstream assumes the voice loop works, and nobody has confirmed that since
the redeploy.
---
## Before you start
Two things bite anyone running these checks on homesrv.
**curl needs `--noproxy '*'`.** The shell exports `http_proxy=http://127.0.0.1:18080`.
Without the flag, every local check returns 503 from the proxy and looks like a
dead service. This cost me a false regression report today.
**The database is not readable with sqlite3.** Four older QA steps say
`docker compose exec mavend sqlite3 /data/maven.db "select ..."`. That cannot
work: the container has no `sqlite3` binary, and the store is AES-256-GCM at
rest with a tmpfs working copy. Read state through mavweb instead, at
`/history`, `/trace`, `/routines` and `/dash`.
---
## Session 1: the voice loop (half a day)
Nothing here has been confirmed since the redeploy, and everything else assumes
it works. Do this first.
Closes or advances: **44** (conversation), **45** (text chat), **287** (voice
session quality), **321** steps 3-5 (quiet mode), **288** (STT fixtures).
1. Open `http://127.0.0.1:9201/chat` and hold a short conversation in Russian.
Watch for three things: she answers in feminine forms (`рада`, `поняла`), she
says `ты` and never `вы`, and no pet names appear.
2. Press push-to-talk on `/dash`. Say `привет`. Confirm a spoken reply comes
back. This is the only check that covers mic to STT to core to TTS to
speaker as one path. It is also the path the eleven-day outage most likely
broke.
3. Say `тихий режим`. Expect `тихий режим включён. буду реже напоминать.`
4. Say `выключи тихий режим`. Expect `тихий режим выключен.` Negation must win.
5. Say `в комнате тихо`. Quiet mode must NOT flip. Confirm on `/history` that no
`quiet_hours` fact was written.
6. Say `включи режим тишины`, then `сделай потише`. Both must flip quiet mode
on. These are the noun form and the comparative, added 01-08-2026.
7. Wait for a nudge, then say `потом` within twenty minutes. Expect `хорошо,
вернусь к этому позже.` and the nudge row on `/notifications` reading
`snoozed`. Say `потом` again with nothing pending: it must route as an
ordinary utterance, not be swallowed.
8. Wait for the water nudge, then say `выпил воды`. Expect the ordinary fact
reply and nothing extra — she must not congratulate you. Check
`/notifications`: the row reads `acted`. Then trigger another nudge and say
`готово`; expect `отлично, отметила.` and the same outcome.
9. Note anything where she is slow, cuts off, or talks over herself. That is
287's whole content and it has no written acceptance criteria yet.
**319 is fixed** (01-08-2026). Single-word Russian utterances no longer come
back as `не совсем поняла — можешь переформулировать?`. `привет` and `поужинал`
both pass now: `thinSingleToken` spares social singles and any token carrying a
verb ending, and only thins a bare nominal like `вода`. If a one-word utterance
still gets clarified during the smoke test, that is a new case for the lexicon,
not the old bug.
---
## Session 2: measurement (half a day, mostly waiting)
Closes or advances: **320** items 2-4, **278** (make the eval lab routine),
**319** (gate recalibration).
The resident llama-server cannot be reached by the eval harness. It binds
`--host 127.0.0.1 --port 0` inside the container, so the port is kernel-assigned
and never published. Start a second one on a fixed port instead:
```sh
llama-server -m /mnt/hdd1/llms/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf \
--host 127.0.0.1 --port 18100 -c 4096 -ngl 99 --no-webui
```
`-c 4096` matters. The recorded numbers were measured at that context size, and
a mismatch invalidates the comparison.
Then:
```sh
make eval-models MAVEN_LLM_URL=http://127.0.0.1:18100 # want ~72.7% cascade
make eval-router # classifier baseline
MAVEN_LLM_URL=http://127.0.0.1:18100 make eval-phrasing # persona checks, slow
make eval-recall
```
A large miss against 72.7% means the deploy differs from the bench harness.
Two things to decide while the numbers are in front of you:
- **319's gate recalibration.** The single-token rule needs narrowing or
dropping. This needs your judgement, not a threshold sweep. The fixture and the
daemon disagree about what is correct on two of the three false clarifies.
- **278's real ask** is making the eval lab routine rather than building it. It
is built. Decide whether it runs on a timer, on every merge, or on demand, and
the task can close.
Item 4 of **320** needs a permission I do not have. Kill the `llama-server`
pid under `maven-mavend-1`, post a turn, and confirm it still completes
through the classifier. Either grant it or run it yourself. It is the only
check that the failure floor catches a mid-session model death.
---
## Session 3: the interaction batch (a day, or three sittings)
These need real use rather than a command, grouped by what one sitting covers.
**Morning and delivery** (**280**, **281**, **128**, **282**): open `/morning`,
walk the seven required behaviours, then check the four interruption outcomes
and the digest gap. **282** needs the `desk_active` script enabled on the desk
PC first, which is **15** and needs you at that machine.
**Tasks and calendar** (**129**, **130**, **127**, **126**, **246**): capture a
task by voice, confirm it lands, check prioritisation ordering is not nonsense.
**246** (mail reader) also exercises the `IngestMail` rung that moved to
`AuthWrite` this morning.
**Routines and patterns** (**43**, **46**, **247**, **254**): these need history
to detect against. If the database is thin after the outage, they may have
nothing to propose, which is not a failure. Check `/routines` before
concluding anything.
**Ecosystem** (**272**, **273**, **276**): nexus, hexis and praxis are wired and
logged clean at boot. **276** is the degraded-mode suite, which means taking
siblings down on purpose. Worth doing while you are already in there.
---
## Housekeeping (one sitting, no box needed)
Four QA tasks will not close no matter how long they sit, because they are
gated on something that does not exist:
- **125** zenmoney: needs a token you have not minted.
- **256** Home Assistant: needs HA configured.
- **257** Bluetooth: BLOCKED, no bluez on the box. Says so in the title.
- **288** STT golden audio: needs fixtures generated.
Relabel these so they stop reading as backlog. They are not verification work
that is pending, they are work that has not started.
Same treatment for the five plan-only tasks (**251** MCP, **252** vision,
**253** hearing, **255** speaker recognition, **259** crawler). A `QA:` prefix on
a plan is misleading.
---
## Needs you specifically
Not QA. These are blocked on a decision or a credential only you have.
| # | what |
|---|---|
| 16 | Create the Kuma API key. `-kuma-key uk5_mavpoll-key` in `docker-compose.yml` is still the placeholder. |
| 15 | Deploy `desk_active` on the desk PC. Blocks **282**. |
| 122 | Finish the CPT run for Qwen3-1.7B. The persona fix depends on it. |
| 355 | Deploy the Hexis auth change. Was blocked on Maven being under construction, which it no longer is. The client half is vendored and wired. |
| 357 | Decide whether entity-existence validation is the permanent target guard or whether blessing lands in Nexus. |
| 275 | Hexis native API and MCP parity. |
| — | Decide on `-require-stepup`. Making it the default needs WebAuthn configured first, or it locks you out of your own admin surfaces. See **317**. |
| — | Three nginx sites bind wildcard `:80` (`acme.conf`, `matrix`, `panel`), so the ecosystem's bind-level protection is not in effect and `allow`/`deny` is carrying it alone. See **354**. |
---
## Suggested order
1. Session 1. If the voice loop is broken, nothing else matters.
2. The `-require-stepup` and Kuma decisions. Five minutes, unblocks **317** fully
and **16**.
3. Session 2. The numbers tell you whether the router is worth its 90x latency.
4. Housekeeping. Cheap, and it makes the remaining backlog honest.
5. Session 3, split whichever way suits you.
+1 -1
View File
@@ -1,6 +1,6 @@
// Package main is mavenclient — maven's reference client.
//
// Per DESIGN.md § Voice pipeline (STT / TTS): capture lives on the client;
// Per docs/design.md § Voice pipeline (STT / TTS): capture lives on the client;
// the server transcribes + synthesises on demand. The PC client runs the
// wake-word / VAD gate (cmd/mavwaked) and ships ONE clean audio blob per
// utterance on activation. The server never owns a mic.
+50 -14
View File
@@ -7,6 +7,7 @@ import (
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
)
// actionFact handles router.IntentFact: persist a tapped self-fact, index
@@ -15,14 +16,41 @@ func (h *reactiveHandler) actionFact(ctx context.Context, dec router.Decision) s
if !dec.Slots.HasKey {
return "не разобрала, что записать — попробуй иначе."
}
// A question is never a fact about him (#470). "какая последняя версия
// языка Go?" used to land here, and the value stored was whatever the
// model invented for it, at confidence 1.00, indexed for recall under the
// question's own text. Two such rows then claimed seven unrelated world
// questions through recall and silently disabled world answering.
//
// The routing error itself is not fixed here — the answer is to answer.
// Sending the turn down the query chain is what he asked for anyway, and
// it costs a mis-routed capture nothing: an explicit "запиши ..." is not
// question-shaped, so it never takes this branch.
if router.IsQuestionShaped(dec.Utterance) {
log.Printf("voice: fact write refused, utterance is a question: %q (key %q) — answering as a query",
dec.Utterance, dec.Slots.Key)
q := dec
q.Intent = router.IntentQuery
// The key the model extracted is its guess at what to store, not a
// fact he has. Left in place, queryFactByKey would read it back and
// claim the turn before any real source ran.
q.Slots.Key, q.Slots.HasKey = "", false
q.Slots.Value = ""
return h.actionQuery(ctx, q)
}
now := h.now()
req := ipc.WriteFactReq{
Ts: now,
Kind: "self",
Key: dec.Slots.Key,
Value: dec.Slots.Value,
Source: "tap:voice",
Confidence: 1.0,
Ts: now,
Kind: "self",
Key: dec.Slots.Key,
Value: dec.Slots.Value,
Source: "tap:voice",
// Not 1.00 unconditionally any more (#470). A value he said is
// evidence; a value the model supplied for words he never said is a
// guess, and writing a guess at full confidence is the same mistake
// the act path already refuses under "LLM output is not
// authorization".
Confidence: factConfidence(dec.Utterance, dec.Slots.Value),
// Subject: the key doubles as the entity-resolution candidate —
// a voice-tapped fact's key is usually the thing/person it's
// about ("espresso_machine", "kate"), so queueing it for Nexus
@@ -36,17 +64,25 @@ func (h *reactiveHandler) actionFact(ctx context.Context, dec router.Decision) s
log.Printf("voice: write fact: %v", err)
return "не получилось сохранить факт."
}
// Index the fact utterance in long-term memory (best-effort, must not
// fail the fact write). Facts aren't in the notes table, so this is the
// only recall path for them — "когда я пил воду?" reads back from here.
// Index the fact in long-term memory (best-effort, must not fail the fact
// write). Facts aren't in the notes table, so this is the only recall path
// for them — "когда я пил воду?" reads back from here.
//
// The indexed text is the fact, not the utterance (#493). queryMemory
// returns a fact's stored text verbatim, so what goes in here is what he
// hears; storing the utterance meant recall answered with his own sentence
// rather than the value. The utterance stays alongside as provenance —
// readable on /trace, never the answer and never embedded.
if h.memStore != nil {
if vec, err := router.EmbedPassage(ctx, h.embedder, dec.Utterance); err != nil {
text := store.FactRecallText(dec.Slots.Key, dec.Slots.Value)
if vec, err := router.EmbedPassage(ctx, h.embedder, text); err != nil {
log.Printf("voice: embed fact for memory: %v", err)
} else if err := h.memStore.Insert(ctx, "fact:"+dec.Slots.Key+":"+strconv.FormatInt(now.Unix(), 10), vec, map[string]string{
"source": "voice",
"type": "fact",
"text": dec.Utterance,
"ts": strconv.FormatInt(now.Unix(), 10),
"source": "voice",
"type": "fact",
"text": text,
"utterance": dec.Utterance,
"ts": strconv.FormatInt(now.Unix(), 10),
}); err != nil {
log.Printf("voice: memory insert fact: %v", err)
}
+39 -24
View File
@@ -13,6 +13,7 @@ import (
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/morning"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/rss"
"github.com/kami/maven/internal/store"
@@ -433,6 +434,14 @@ func (h *reactiveHandler) queryMemory(ctx context.Context, t *queryTurn) (string
return "", false
}
text := hit.Meta["text"]
// The score cleared the gate and the topic still has to match (#470). A
// note about his slow network scored high enough to answer "почему небо
// синее?", because the right-note and must-be-silent score ranges overlap
// and no threshold sits between them.
if !memory.RecallAllowed(t.dec.Utterance, text) {
log.Printf("voice: recall %q rejected for %q: a world question and no shared topic word", text, t.dec.Utterance)
return "", false
}
// A note is phrased in Maven's voice; a fact is read back as it was
// stored.
if hit.Meta["type"] == "note" {
@@ -468,6 +477,12 @@ func (h *reactiveHandler) queryNotes(ctx context.Context, t *queryTurn) (string,
if !memory.ConfidentScores(noteScores, h.queryMinScore, h.queryMinMargin) {
return "", false
}
// Same topic veto as queryMemory above: the best note must be about what
// he asked, not merely the nearest vector in the index.
if !memory.RecallAllowed(t.dec.Utterance, notes[0].Text) {
log.Printf("voice: note %q rejected for %q: a world question and no shared topic word", notes[0].Text, t.dec.Utterance)
return "", false
}
texts := make([]string, len(notes))
for i, n := range notes {
texts[i] = n.Text
@@ -523,10 +538,7 @@ func (h *reactiveHandler) queryWeb(ctx context.Context, t *queryTurn) (string, b
// the question he actually asked. She answers the question, she does not
// recite the page.
snippet := page.Title + "\n" + crawl.TrimRunes(page.Text, webPageContextRunes)
reply, perr := h.phraser.PhraseQuery(ctx, t.dec.Utterance, []string{snippet})
if perr != nil {
log.Printf("voice: web: phrase: %v", perr)
}
reply := h.phraseSource(ctx, "web", t.dec.Utterance, []string{snippet})
if reply == "" {
// No phraser (or it failed): read back the top of the page rather than
// pretend the fetch did not happen.
@@ -592,14 +604,7 @@ func (h *reactiveHandler) querySearch(ctx context.Context, t *queryTurn) (string
// question he asked, not something to recite. The trim is one budget over the
// joined block, so a long first snippet cannot crowd out the rest.
evidence := crawl.TrimRunes(strings.Join(resp.Snippets(), "\n"), h.search.runes)
var reply string
if h.phraser != nil {
var perr error
reply, perr = h.phraser.PhraseQuery(ctx, t.dec.Utterance, []string{evidence})
if perr != nil {
log.Printf("voice: search: phrase: %v", perr)
}
}
reply := h.phraseSource(ctx, "search", t.dec.Utterance, []string{evidence})
if reply == "" {
// No phraser, or it failed. Read back the best evidence rather than
// pretend the search did not happen.
@@ -680,14 +685,7 @@ func (h *reactiveHandler) queryKiwix(ctx context.Context, t *queryTurn) (string,
// Handed over the same way a note or a page is: context for the question he
// asked, not something to recite.
snippet := top.Title + "\n" + crawl.TrimRunes(page.Text, h.kiwix.runes)
var reply string
if h.phraser != nil {
var perr error
reply, perr = h.phraser.PhraseQuery(ctx, t.dec.Utterance, []string{snippet})
if perr != nil {
log.Printf("voice: kiwix: phrase: %v", perr)
}
}
reply := h.phraseSource(ctx, "kiwix", t.dec.Utterance, []string{snippet})
if reply == "" {
// No phraser, or it failed. Read back the best hit rather than pretend
// the search did not happen.
@@ -755,11 +753,28 @@ func isPersonalQuery(utterance string) bool {
return false
}
// queryGeneral — general knowledge from the phraser, the last source before
// giving up. It always claims: either the model answers or Maven says she
// doesn't know.
// queryGeneral — general knowledge, the last source before giving up. It always
// claims: either a model answers, or Maven names the gap, or she says she does
// not know.
//
// This is the sharpest case for the naming half. Nothing has been fetched, so
// there is no passage to fall back on and no floor under the answer except the
// model's weights — and a 1.7B's weights are where the invented answers come
// from. With a workstation configured and asleep he is told that, rather than
// told something false in a confident voice. With no workstation configured at
// all the resident model answers exactly as it does today: naming a gap requires
// a gap, and on that box the 1.7B is the whole product.
func (h *reactiveHandler) queryGeneral(ctx context.Context, t *queryTurn) (string, bool) {
reply, err := h.phraser.PhraseQuery(ctx, t.dec.Utterance, nil)
if h.phraser == nil {
// No model of any size. That is not the workstation being asleep, so it
// is not that gap: it is simply not knowing.
return "не знаю.", true
}
reply, err := h.phraseWorld(ctx, t.dec.Utterance, nil)
if errors.Is(err, phraser.ErrNoWorldModel) {
log.Printf("voice: %q needs the world model and it is not available", t.dec.Utterance)
return worldGap, true
}
if err != nil || reply == "" {
return "не знаю.", true
}
+12 -8
View File
@@ -107,14 +107,18 @@ func (h *reactiveHandler) confirmResolvers(ctx context.Context) []confirmResolve
return pr != nil && !h.now().After(pr.expiry)
},
yes: func() string {
// Only record the acceptance. The tick loop reads accepted
// routines and nudges on their own interval. Building a
// reminder here made a routine fire exactly once (Vikunja #366).
if err := h.dataStore.AcceptProposedRoutine(ctx, pr.routineID, h.now()); err != nil {
log.Printf("voice: accept proposed routine: %v", err)
return "не получилось запомнить рутину."
}
return "буду напоминать."
// Voice does NOT accept (Vikunja #367). Accepting hands the
// tick loop a standing new reason to speak, which is the same
// tier as enabling a tool — and DESIGN.md § "surface caps
// authority" says a room mic, reachable by anyone present, is
// structurally incapable of layer 3. So a spoken "да" leaves
// the row 'proposed' and points at the authed page, where the
// accept button is gated at step-up. The convenience of
// answering out loud stays; the authority does not move.
//
// Acceptance itself is recorded by /routines, and the tick
// loop nudges on the interval from there (Vikunja #366).
return "поняла — подтверди на странице рутин, и начну напоминать."
},
no: func() string {
if err := h.dataStore.DismissProposedRoutine(ctx, pr.routineID); err != nil {
+82
View File
@@ -0,0 +1,82 @@
package main
import (
"log"
"strings"
"unicode"
)
// ungroundedConfidence — what a self fact is worth when its value appears
// nowhere in what he said. Below `query_min_score` is not the point (recall
// gates on vector distance, not on this number); the point is that
// `/history` and every future reader can tell a value he said from a value
// the model supplied.
const ungroundedConfidence = 0.6
// factConfidence scores a self fact by whether its value is grounded in the
// utterance it came from. Grounded stays 1.00, which is what a tapped fact
// has always been worth. Ungrounded drops, and says so in the log.
//
// An empty value is grounded by definition: the key alone carries the fact
// ("поужинал"), and there is nothing for the model to have invented.
func factConfidence(utterance, value string) float64 {
if strings.TrimSpace(value) == "" {
return 1.0
}
if valueGrounded(utterance, value) {
return 1.0
}
log.Printf("voice: fact value %q is not in %q — writing at confidence %.2f",
value, utterance, ungroundedConfidence)
return ungroundedConfidence
}
// valueGrounded reports whether every word of value traces back to a word he
// actually said. The comparison is on a 4-rune prefix, so the model's
// normalization survives ("пил воду" → "вода") while an invented value
// ("1.20" for a question about Go) does not.
func valueGrounded(utterance, value string) bool {
said := factTokens(utterance)
words := factTokens(value)
if len(words) == 0 {
return true
}
for _, w := range words {
if !anyTokenMatches(said, w) {
return false
}
}
return true
}
func anyTokenMatches(said []string, w string) bool {
for _, s := range said {
if s == w || sameStem(s, w) {
return true
}
}
return false
}
// sameStem is inflection tolerance and nothing more: it compares all but the
// last rune of the shorter word, and never fewer than three. Russian marks
// case on the ending, so "пил воду" and the stored "вода" are the same word he
// said, while "1.20" and "версия" are not. A word of three runes or fewer must
// match outright, where a shorter prefix would match half the language.
func sameStem(a, b string) bool {
ar, br := []rune(a), []rune(b)
shorter := min(len(ar), len(br))
n := shorter - 1
if n < 3 || len(ar) < n || len(br) < n {
return false
}
return string(ar[:n]) == string(br[:n])
}
// factTokens lowercases and splits on everything that is not a letter or a
// digit, the same shape planTokens uses in the router.
func factTokens(s string) []string {
return strings.FieldsFunc(strings.ToLower(s), func(r rune) bool {
return !unicode.IsLetter(r) && !unicode.IsDigit(r)
})
}
+125
View File
@@ -0,0 +1,125 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/tool"
"github.com/kami/maven/internal/voice"
)
func newFactGateHandler(t *testing.T, now time.Time) (*reactiveHandler, ipc.CoreAPI) {
t.Helper()
st := newTestStore(t)
api := ipc.NewStoreAPI(st)
emb := router.NewHashEmbedder(1024)
h := &reactiveHandler{
api: api,
embedder: emb,
router: buildRouter(emb, tool.NewMatcher(api), 0.55, nil),
replier: voice.NewStubReplier(),
now: func() time.Time { return now },
memStore: memory.NewInMemoryStore(),
dataStore: st,
}
return h, api
}
// The write half of #470: a question routed to IntentFact must not become a
// fact about him, and must not leave a vector behind for recall to serve.
func TestActionFact_QuestionIsNotWritten(t *testing.T) {
ctx := context.Background()
h, api := newFactGateHandler(t, time.Now())
reply := h.actionFact(ctx, router.Decision{
Intent: router.IntentFact,
Utterance: "какая последняя версия языка Go?",
Slots: router.Slots{Key: "go_version", HasKey: true, Value: `"1.20"`},
})
if _, err := api.LatestFact(ctx, "go_version"); err == nil {
t.Fatal("a question was stored as a fact about him")
}
hits, err := h.memStore.Search(ctx, mustEmbedPassage(t, h, "какая последняя версия языка Go?"), 3)
if err != nil {
t.Fatalf("memory search: %v", err)
}
if len(hits) != 0 {
t.Fatalf("the question was indexed for recall: %+v", hits)
}
// It went down the query chain instead. Nothing is configured to answer a
// world question in this harness, so "не знаю." is the honest outcome —
// what matters is that the turn was answered, not stored.
if reply == "" {
t.Fatal("the turn was neither stored nor answered")
}
}
// The capture that must survive the gate: an explicit instruction to record,
// even though it contains an interrogative.
func TestActionFact_ExplicitCaptureStillWrites(t *testing.T) {
ctx := context.Background()
h, api := newFactGateHandler(t, time.Now())
h.actionFact(ctx, router.Decision{
Intent: router.IntentFact,
Utterance: "запиши что я пил воду",
Slots: router.Slots{Key: "water", HasKey: true, Value: `"вода"`},
})
f, err := api.LatestFact(ctx, "water")
if err != nil {
t.Fatalf("an explicit capture was refused: %v", err)
}
if f.Confidence != 1.0 {
t.Errorf("confidence = %v, want 1.0 for a value he said", f.Confidence)
}
// #493: what recall reads back is the fact, not the sentence he said.
// queryMemory returns a fact's text verbatim, so the utterance sitting here
// meant "запиши что я пил воду" was the answer to "когда я пил воду?".
hits, err := h.memStore.Search(ctx, mustEmbedPassage(t, h, "вода"), 3)
if err != nil {
t.Fatalf("memory search: %v", err)
}
if len(hits) != 1 {
t.Fatalf("the fact was not indexed once: %+v", hits)
}
if got := hits[0].Meta["text"]; got != "water — вода" {
t.Errorf("indexed text = %q, want the fact", got)
}
if got := hits[0].Meta["utterance"]; got != "запиши что я пил воду" {
t.Errorf("utterance provenance = %q, want it kept alongside", got)
}
}
func TestFactConfidence(t *testing.T) {
cases := []struct {
utterance, value string
want float64
}{
{"запиши что я пил воду", `"вода"`, 1.0},
{"я выпил кофе", `"кофе"`, 1.0},
{"поужинал", "", 1.0},
{"отметь что я полил кактус", `"полил кактус"`, 1.0},
{"какая последняя версия языка Go", `"1.20"`, ungroundedConfidence},
{"кто премьер Японии", `"Тонио Озаки"`, ungroundedConfidence},
}
for _, c := range cases {
if got := factConfidence(c.utterance, c.value); got != c.want {
t.Errorf("factConfidence(%q, %q) = %v, want %v", c.utterance, c.value, got, c.want)
}
}
}
func mustEmbedPassage(t *testing.T, h *reactiveHandler, text string) []float32 {
t.Helper()
vec, err := router.EmbedQuery(context.Background(), h.embedder, text)
if err != nil {
t.Fatalf("embed %q: %v", text, err)
}
return vec
}
+12 -5
View File
@@ -12,13 +12,21 @@ import (
// which waits on voice-print attribution (see PROGRESS multi-user deferral).
const voiceDialogueID = "voice"
// toDialogueSlots projects the router's slots onto the dialogue layer's subset
// (everything except the fact Value, which the dialogue layer doesn't carry).
// toDialogueSlots and applyDialogueSlots are the only bridge between
// router.Slots and dialogue.Slots. dialogue must not import router (import
// cycle), so the two structs are hand-kept copies and every field has to be
// carried by hand here. Adding a field to either struct without adding it to
// BOTH functions loses a slot silently — nothing fails to build. The tests in
// slotsparity_test.go fail when the field sets or the converters stop matching;
// when they do, fix these two functions, not the tests.
// toDialogueSlots projects the router's slots onto the dialogue layer's copy.
func toDialogueSlots(s router.Slots) dialogue.Slots {
return dialogue.Slots{
Time: s.Time,
HasTime: s.HasTime,
Key: s.Key,
Value: s.Value,
HasKey: s.HasKey,
Text: s.Text,
Fn: s.Fn,
@@ -27,11 +35,10 @@ func toDialogueSlots(s router.Slots) dialogue.Slots {
}
}
// applyDialogueSlots writes inherited dialogue slots back onto router slots,
// preserving router-only fields (Value) the dialogue layer never touched.
// applyDialogueSlots writes dialogue slots back onto router slots.
func applyDialogueSlots(base router.Slots, d dialogue.Slots) router.Slots {
base.Time, base.HasTime = d.Time, d.HasTime
base.Key, base.HasKey = d.Key, d.HasKey
base.Key, base.Value, base.HasKey = d.Key, d.Value, d.HasKey
base.Text = d.Text
base.Fn, base.Args, base.HasFn = d.Fn, d.Args, d.HasFn
return base
+91
View File
@@ -0,0 +1,91 @@
package main
import (
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/llm"
)
// No `workstation` block is the shipping deploy. The seam must then be the
// resident client itself, with nothing probing anything.
func TestModelSeamUnconfiguredIsResidentOnly(t *testing.T) {
resident := llm.New("http://127.0.0.1:1", time.Second)
hot, pair := modelSeam(&config.Config{}, resident)
if pair != nil {
t.Error("built a pair with no workstation configured")
}
if hot == nil {
t.Fatal("no seam at all, so the cascade would route with the classifier")
}
}
// A workstation with no resident model behind it has no floor, and a Pair with
// no floor is a configuration mistake rather than a degraded mode.
func TestModelSeamWithoutResidentIsNil(t *testing.T) {
cfg := &config.Config{Workstation: &config.WorkstationConfig{URL: "http://127.0.0.1:1"}}
cfg.Workstation.Health = strings.TrimRight(cfg.Workstation.URL, "/") + "/health"
hot, pair := modelSeam(cfg, nil)
if hot != nil || pair != nil {
t.Errorf("built a seam with no floor: hot=%v pair=%v", hot, pair)
}
}
// The configured case: the seam is the pair, and the pair notices a workstation
// that answers /health.
func TestModelSeamPrefersAnAnsweringWorkstation(t *testing.T) {
up := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(http.StatusOK)
}))
defer up.Close()
cfg := &config.Config{Workstation: &config.WorkstationConfig{
URL: up.URL,
Probe: config.Duration(10 * time.Millisecond),
}}
cfg.Workstation.Health = strings.TrimRight(cfg.Workstation.URL, "/") + "/health"
hot, pair := modelSeam(cfg, llm.New("http://127.0.0.1:1", time.Second))
if pair == nil || hot == nil {
t.Fatal("no pair built for a configured workstation")
}
defer pair.Stop()
deadline := time.Now().Add(2 * time.Second)
for !pair.Available() && time.Now().Before(deadline) {
time.Sleep(5 * time.Millisecond)
}
if !pair.Available() {
t.Fatal("the pair never saw a workstation that answers /health")
}
}
// A card held by a CPT run answers 503, and that must read as unavailable
// rather than as an error a turn has to handle.
func TestModelSeamHeldCardIsUnavailable(t *testing.T) {
busy := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Error(w, "model not loaded", http.StatusServiceUnavailable)
}))
defer busy.Close()
cfg := &config.Config{Workstation: &config.WorkstationConfig{
URL: busy.URL,
Probe: config.Duration(10 * time.Millisecond),
}}
cfg.Workstation.Health = strings.TrimRight(cfg.Workstation.URL, "/") + "/health"
_, pair := modelSeam(cfg, llm.New("http://127.0.0.1:1", time.Second))
if pair == nil {
t.Fatal("no pair built for a configured workstation")
}
defer pair.Stop()
time.Sleep(50 * time.Millisecond)
if pair.Available() {
t.Error("a 503 from the supervisor read as available")
}
}
+79
View File
@@ -9,6 +9,7 @@ import (
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/pattern"
"github.com/kami/maven/internal/store"
@@ -283,3 +284,81 @@ func TestTickProposalCooldownSpacesAnnouncements(t *testing.T) {
}
}
}
// TestVoiceYesDoesNotAcceptRoutine — Vikunja #367. Accepting a routine hands
// the tick loop a standing new reason to speak, which DESIGN.md puts at layer
// 3, and voice is structurally incapable of layer 3. A spoken "да" must park
// the decision for the authed page, not flip the row itself.
func TestVoiceYesDoesNotAcceptRoutine(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents-1)
h := &reactiveHandler{api: ipc.NewStoreAPI(st), dataStore: st, now: func() time.Time { return now }}
// The MinEvents'th event is the one that makes the pattern detectable, and
// it goes through the voice path so the proposal is parked for a y/n.
last := now.Add(time.Duration(pattern.MinEvents-1) * 7 * 24 * time.Hour)
factID, err := st.WriteFact(ctx, last, store.KindSelf, "cat_water", "refill", "voice", 1.0, sql.NullInt64{})
if err != nil {
t.Fatalf("write fact: %v", err)
}
if phrase := h.detectPattern(ctx, factID, "cat_water", "refill", last); phrase == "" {
t.Fatal("expected a parked routine proposal")
}
reply, handled := h.resolveConfirm(ctx, "да")
if !handled {
t.Fatal("the spoken yes should be consumed by the routine confirm")
}
if !strings.Contains(reply, "рутин") {
t.Fatalf("reply should send him to the routines page, got %q", reply)
}
rows, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineAccepted)
if err != nil {
t.Fatalf("list accepted: %v", err)
}
if len(rows) != 0 {
t.Fatalf("voice accepted a routine: %+v", rows)
}
proposed, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineProposed)
if err != nil {
t.Fatalf("list proposed: %v", err)
}
if len(proposed) != 1 {
t.Fatalf("proposed routines = %d, want 1 (still waiting for the page)", len(proposed))
}
}
// TestVoiceNoStillDismissesRoutine — declining does not move the boundary
// outward, so voice keeps it. Only acceptance is gated.
func TestVoiceNoStillDismissesRoutine(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents-1)
h := &reactiveHandler{api: ipc.NewStoreAPI(st), dataStore: st, now: func() time.Time { return now }}
last := now.Add(time.Duration(pattern.MinEvents-1) * 7 * 24 * time.Hour)
factID, err := st.WriteFact(ctx, last, store.KindSelf, "cat_water", "refill", "voice", 1.0, sql.NullInt64{})
if err != nil {
t.Fatalf("write fact: %v", err)
}
if phrase := h.detectPattern(ctx, factID, "cat_water", "refill", last); phrase == "" {
t.Fatal("expected a parked routine proposal")
}
if _, handled := h.resolveConfirm(ctx, "нет"); !handled {
t.Fatal("the spoken no should be consumed by the routine confirm")
}
rows, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineDismissed)
if err != nil {
t.Fatalf("list dismissed: %v", err)
}
if len(rows) != 1 {
t.Fatalf("dismissed routines = %d, want 1", len(rows))
}
}
+1 -1
View File
@@ -1,5 +1,5 @@
// mavend/simulator_test.go — the replayable full-system simulator
// (Vikunja #284, 20-07-2026-BACKLOG.md item 7).
// (Vikunja #284).
//
// # What it is
//
+71
View File
@@ -0,0 +1,71 @@
package main
import (
"reflect"
"testing"
"time"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/router"
)
// TestSlotsParity — dialogue.Slots is a hand-kept copy of router.Slots
// (dialogue must not import router: import cycle). Drift is silent, so this
// test compares the two field sets by name and type. If it fails, add the new
// field to both structs AND to toDialogueSlots/applyDialogueSlots in
// followup.go — do not relax the test.
func TestSlotsParity(t *testing.T) {
fields := func(v any) map[string]string {
rt := reflect.TypeOf(v)
out := make(map[string]string, rt.NumField())
for i := 0; i < rt.NumField(); i++ {
f := rt.Field(i)
out[f.Name] = f.Type.String()
}
return out
}
rf, df := fields(router.Slots{}), fields(dialogue.Slots{})
for name, typ := range rf {
dt, ok := df[name]
if !ok {
t.Errorf("router.Slots.%s (%s) missing from dialogue.Slots", name, typ)
continue
}
if dt != typ {
t.Errorf("field %s: router has %s, dialogue has %s", name, typ, dt)
}
}
for name, typ := range df {
if _, ok := rf[name]; !ok {
t.Errorf("dialogue.Slots.%s (%s) missing from router.Slots", name, typ)
}
}
}
// TestSlotsRoundTrip — the converters carry every field. A field the parity
// test accepts can still be dropped in transit, so round-trip a fully
// populated value and compare.
func TestSlotsRoundTrip(t *testing.T) {
full := router.Slots{
Time: time.Date(2026, 8, 2, 11, 0, 0, 0, time.UTC),
HasTime: true,
Fn: "restart",
Args: []string{"nginx"},
HasFn: true,
Key: "water",
Value: `"drank"`,
HasKey: true,
Text: "выпил воды",
}
// Every field must be non-zero, or the round-trip proves nothing.
rv := reflect.ValueOf(full)
for i := 0; i < rv.NumField(); i++ {
if rv.Field(i).IsZero() {
t.Fatalf("field %s is zero: extend this fixture so the round-trip covers it",
rv.Type().Field(i).Name)
}
}
if got := applyDialogueSlots(router.Slots{}, toDialogueSlots(full)); !reflect.DeepEqual(got, full) {
t.Errorf("round-trip lost a slot:\n got %+v\nwant %+v", got, full)
}
}
+90 -5
View File
@@ -48,7 +48,11 @@ type voiceWiring struct {
// mcp — the MCP client, nil unless the `mcp` block configures an enabled
// server (Vikunja #251). Its tools land in the same allowlist as every
// other act, so nothing else here has to know about it.
mcp *mcpWiring
// pair — the workstation model with the resident one as the floor, nil
// unless a `workstation` block names an address. Held here only so the
// prober is stopped on shutdown; callers were handed it at build time.
pair *llm.Pair
mcp *mcpWiring
// home — the Home Assistant client, nil unless the `smarthome` block is
// enabled (Vikunja #256). Its devices land in the same allowlist as every
// other act, so nothing else here has to know about it.
@@ -76,6 +80,9 @@ func (w *voiceWiring) close() {
if w.ttsClient != nil {
_ = w.ttsClient.Close()
}
if w.pair != nil {
w.pair.Stop()
}
w.mcp.close()
}
@@ -139,6 +146,7 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
emb = router.NewHashEmbedder(1024)
}
w.embedder = emb
repairFactVectors(dataStore, emb)
checkStoredEmbedder(dataStore, emb)
// ----- tool executor (the enabled act allowlist, store-backed) -----
@@ -188,6 +196,18 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
// the new llama-server when the resident model is swapped (Vikunja #250).
llmClient = llmClientFor(lp, 60*time.Second)
}
// The workstation model sits above that one when it is configured and its
// card is free. hot is what the router and the replier complete through:
// either the pair, or the resident client alone, or nothing at all.
hot, pair := modelSeam(cfg, llmClient)
w.pair = pair
// The phraser gets the same pair, which is what carries the workstation model
// into the paths that do not go through `hot`: world questions (the naming
// half), and the digestion worker's nudge and reminder phrasing (the silent
// half). Wiring, so it happens once and before the voice server listens.
if lp, ok := phr.(*phraser.LLMPhraser); ok && pair != nil {
lp.UseRemote(pair)
}
// ----- router (the cascade; floor examples seed the classifier) -----
// The act matcher's allowlist is exactly the enabled tool names — the
// router only matches acts the executor can run (one source of truth).
@@ -199,7 +219,7 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
// against the classifier's 50.0%, at about 1s a turn instead of 30ms (see
// config.VoiceConfig.LLMRouter). The classifier always stays wired as the
// fallback, so a model error never breaks a turn.
rtr := buildRouter(emb, matcher, threshold, pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient))
rtr := buildRouter(emb, matcher, threshold, pickLLMRouter(cfg.Voice.UseLLMRouter(), hot))
// ----- sessions registry (shared with voicesink) -----
sessions := voice.NewSessions()
@@ -233,8 +253,8 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
// ----- replier (LLM-backed when the engine is on, Stub floor otherwise) -----
replier := voice.Replier(voice.NewStubReplier())
if llmClient != nil {
replier = newLLMReplier(llmClient, contextBlockFn(cfg, time.Now))
if hot != nil {
replier = newLLMReplier(hot, contextBlockFn(cfg, time.Now))
}
// ----- the handler (the reactive path; closes over stt / tts / router / coreAPI / memory) -----
@@ -291,7 +311,42 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
// pickLLMRouter returns the LLM router when the operator asked for it and there
// is a llama-server to talk to, and nil otherwise. nil is safe: the cascade then
// routes with the classifier, so an unusable setting costs accuracy, not turns.
func pickLLMRouter(enabled bool, c *llm.Client) *router.LLMRouter {
// modelSeam builds the completion seam the hot paths use: routing and replies.
//
// With no `workstation` block it is the resident client and nothing probes
// anything, which is today's deploy exactly. With one, it is an llm.Pair that
// prefers the workstation and falls back to the resident model silently — the
// silent half of the degradation rule (docs/offload.md), because the big model
// is only better here and the 1.7B is today's shipping quality. He is never
// told which of the two phrased his reply.
//
// A nil resident client means the phraser is not an LLM phraser. There is then
// no floor, and a Pair with no floor is a configuration mistake rather than a
// degraded mode, so the seam is nil and the cascade routes with the classifier.
func modelSeam(cfg *config.Config, resident *llm.Client) (router.Completer, *llm.Pair) {
if resident == nil {
if cfg.Workstation != nil {
log.Printf("voice: a workstation is configured but there is no resident model to floor it with — ignoring the block")
}
return nil, nil
}
if cfg.Workstation == nil {
return resident, nil
}
ws := cfg.Workstation
pair := llm.NewPair(
llm.New(ws.URL, time.Duration(ws.Timeout)),
resident,
ws.Health,
time.Duration(ws.Probe),
)
pair.Start(context.Background())
log.Printf("voice: workstation model at %s, probed every %s, resident model as the floor",
ws.URL, time.Duration(ws.Probe))
return pair, pair
}
func pickLLMRouter(enabled bool, c router.Completer) *router.LLMRouter {
if !enabled {
return nil
}
@@ -419,6 +474,36 @@ func seedTools(api ipc.CoreAPI, tools []config.ToolConfig) {
log.Printf("voice: seeded %d act tools from config", n)
}
// repairFactVectors brings stored fact vectors in line with the facts they name
// (#493), once per box, before the embedder marker is even looked at.
//
// Automatic and not a flag, unlike -reembed: only voice-tapped facts are in
// this index, so the work is tens of embeddings rather than the thousands of
// notes that made the backfill a deliberate act. And the box that needs it is
// broken in a way nobody can see — recall answers with the wrong text and
// nothing logs an error — so waiting for an operator to know to run it is how
// the defect survived four restarts in the first place.
func repairFactVectors(dataStore *store.Store, emb router.Embedder) {
if dataStore == nil {
return
}
res, err := dataStore.RepairFactVectors(context.Background(),
// EmbedPassage, the stored side, same as every other writer of these
// vectors.
func(ctx context.Context, text string) ([]float32, error) {
return router.EmbedPassage(ctx, emb, text)
})
if err != nil {
log.Printf("voice: fact vector repair failed, no marker written and nothing half-done — retried next start: %v", err)
return
}
if res.Skipped || res.Rewritten+res.Dropped == 0 {
return
}
log.Printf("voice: fact vector repair — %d re-embedded from the fact they name, %d dropped as voided or superseded, %d already right, took %s (#493)",
res.Rewritten, res.Dropped, res.Kept, res.Took.Round(time.Millisecond))
}
// reembedOnStart is the -reembed flag (set in run()). Opt-in on purpose: see
// runReembed.
var reembedOnStart bool
+60
View File
@@ -0,0 +1,60 @@
package main
import (
"context"
"errors"
"log"
"github.com/kami/maven/internal/phraser"
)
// worldPhraser — the naming half of the degradation rule (docs/offload.md), as
// the query sources see it. Only *phraser.LLMPhraser implements it, so the
// Stub and every test double stay exactly as they are.
type worldPhraser interface {
PhraseWorld(ctx context.Context, utterance string, sources []string) (string, error)
}
// worldGap — what he hears when the question is about the world, the workstation
// model is the one configured to answer it, and that machine is not answering.
//
// It says the true thing. The resident 1.7B is not a worse answer here, it is an
// invented one: "Война и мир" came back with Левитан as its author, and a
// question about his meeting came back as a swimming competition in Nottingham.
// Naming the gap is the rule CLAUDE.md already applies to a sibling service
// being down.
const worldGap = "сейчас не могу ответить — большая модель недоступна, а придумывать не хочу."
// phraseWorld asks the world model, or reports the gap.
//
// The three outcomes come straight from LLMPhraser.PhraseWorld: no workstation
// configured means the resident model answers as it always has, a workstation
// that is up answers, and a workstation that is down returns
// phraser.ErrNoWorldModel. A phraser that has no world seam at all — the Stub,
// and the doubles in the tests — is the first of those three.
func (h *reactiveHandler) phraseWorld(ctx context.Context, utterance string, sources []string) (string, error) {
if h.phraser == nil {
return "", phraser.ErrNoWorldModel
}
if w, ok := h.phraser.(worldPhraser); ok {
return w.PhraseWorld(ctx, utterance, sources)
}
return h.phraser.PhraseQuery(ctx, utterance, sources)
}
// phraseSource asks the world model to answer from a passage someone already
// fetched — a live search result, a ZIM article, a page he named. It returns ""
// rather than the gap phrase, because these callers hold something better than a
// gap: the passage itself, which their own floor reads back to him. Nothing is
// invented either way, and a real quote beats "не могу сейчас".
func (h *reactiveHandler) phraseSource(ctx context.Context, name, utterance string, sources []string) string {
reply, err := h.phraseWorld(ctx, utterance, sources)
switch {
case errors.Is(err, phraser.ErrNoWorldModel):
log.Printf("voice: %s: no world model, reading the source back instead", name)
return ""
case err != nil:
log.Printf("voice: %s: phrase: %v", name, err)
}
return reply
}
+85
View File
@@ -0,0 +1,85 @@
package main
import (
"context"
"strings"
"testing"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
)
// gapPhraser — a phraser whose world model is configured and asleep, which is
// the state the naming half exists for.
type gapPhraser struct {
*phraser.Stub
worldCalls int
}
func (g *gapPhraser) PhraseWorld(context.Context, string, []string) (string, error) {
g.worldCalls++
return "", phraser.ErrNoWorldModel
}
func worldTurn(utterance string) *queryTurn {
return &queryTurn{dec: router.Decision{Intent: router.IntentQuery, Utterance: utterance}}
}
// A world question with the workstation asleep says so. The resident model is
// not asked, because what it produces here is an invention with no signal that
// it is one.
func TestQueryGeneralNamesTheGap(t *testing.T) {
g := &gapPhraser{Stub: phraser.NewStub()}
h := &reactiveHandler{phraser: g}
reply, ok := h.queryGeneral(context.Background(), worldTurn("почему небо голубое"))
if !ok {
t.Fatal("queryGeneral passed on the last source in the chain")
}
if reply != worldGap {
t.Fatalf("reply = %q, want the named gap", reply)
}
if g.worldCalls != 1 {
t.Fatalf("PhraseWorld called %d times, want 1", g.worldCalls)
}
}
// A phraser with no world seam at all — the Stub, and every box with no
// `workstation` block — answers exactly as it did before this seam existed.
func TestQueryGeneralWithoutAWorldModelIsUnchanged(t *testing.T) {
h := &reactiveHandler{phraser: phraser.NewStub()}
reply, ok := h.queryGeneral(context.Background(), worldTurn("почему небо голубое"))
if !ok {
t.Fatal("queryGeneral passed on the last source in the chain")
}
if reply != "не знаю." {
t.Fatalf("reply = %q, want the Stub's answer", reply)
}
}
// The gap is spoken aloud by a Russian voice, so it is Russian, feminine and
// informal. "не хочу" and "не могу" are her own verbs; there is no "вы" and no
// English in it.
func TestWorldGapIsInPersona(t *testing.T) {
for _, bad := range []string{"вы", "ваш", "рад ", "дорогой", "милый"} {
if strings.Contains(worldGap, bad) {
t.Errorf("the gap phrase contains %q: %s", bad, worldGap)
}
}
if strings.ContainsAny(worldGap, "abcdefghijklmnopqrstuvwxyz") {
t.Errorf("the gap phrase has Latin letters in it: %s", worldGap)
}
}
// The sources that hold a passage read it back rather than name a gap. He gets a
// real quote instead of "не могу сейчас", and nothing is invented either way.
func TestASourceWithAPassageReadsItBackInsteadOfNamingTheGap(t *testing.T) {
g := &gapPhraser{Stub: phraser.NewStub()}
h := &reactiveHandler{phraser: g}
if got := h.phraseSource(context.Background(), "search", "почему небо голубое",
[]string{"Рэлеевское рассеяние."}); got != "" {
t.Fatalf("phraseSource = %q, want \"\" so the caller's own floor reads the passage back", got)
}
if g.worldCalls != 1 {
t.Fatalf("PhraseWorld called %d times, want 1", g.worldCalls)
}
}
+117
View File
@@ -0,0 +1,117 @@
package main
import (
"os"
"path/filepath"
"strconv"
"strings"
)
// The card is an AMD 7900 GRE with 16GB, driven by amdgpu and ROCm. Everything
// here reads sysfs and forks nothing: rocm-smi is not even installed on the
// workstation, and a poll that costs a subprocess every second is a poll that
// gets tuned down until it is useless.
// gpuProc — one process holding the compute engine.
type gpuProc struct {
PID int
Comm string
VRAM int64 // bytes, as the kernel accounts them to this process
}
// probe reads the two sysfs trees the supervisor decides from.
//
// kfdRoot is /sys/class/kfd/kfd/proc, one directory per ROCm process. The
// directory appears when the process initialises HIP, which is well before it
// allocates anything large. That is the whole reason this works: the job that
// is about to want the card announces itself while it is still starting up,
// so we see the contender rather than only the winner of an allocation race.
//
// drmDev is /sys/class/drm/cardN/device, which reports total and used VRAM for
// the card as a whole.
type probe struct {
kfdRoot string
drmDev string
}
// foreign lists every ROCm process that is not ours. selfPID is the supervisor's
// llama-server child, or 0 when it is not running.
//
// An unreadable kfd tree returns no processes and no error. That is deliberate
// and it is the safe direction only because startVRAM also has to agree before
// anything launches: a supervisor that cannot see the KFD never sees free VRAM
// either, because the CPT run holding the card shows up in the drm totals.
func (p probe) foreign(selfPID int) []gpuProc {
entries, err := os.ReadDir(p.kfdRoot)
if err != nil {
return nil
}
var out []gpuProc
for _, e := range entries {
pid, err := strconv.Atoi(e.Name())
if err != nil || pid == selfPID {
continue
}
out = append(out, gpuProc{
PID: pid,
Comm: readComm(pid),
VRAM: p.procVRAM(e.Name()),
})
}
return out
}
// procVRAM sums the per-node vram_* files under one process directory. The
// suffix is the KFD topology node id (vram_35881 on this card), so it is
// globbed rather than named, and a machine with two cards sums both.
func (p probe) procVRAM(pid string) int64 {
matches, err := filepath.Glob(filepath.Join(p.kfdRoot, pid, "vram_*"))
if err != nil {
return 0
}
var total int64
for _, m := range matches {
total += readInt(m)
}
return total
}
// freeVRAM reports the bytes the card has left. Used only to decide whether to
// start: a shortfall here means llama-server would refuse to load anyway. It is
// never used to decide to stop, because by the time free VRAM has dropped the
// other job has already failed its allocation, which is exactly the outcome
// yielding exists to prevent.
func (p probe) freeVRAM() int64 {
total := readInt(filepath.Join(p.drmDev, "mem_info_vram_total"))
used := readInt(filepath.Join(p.drmDev, "mem_info_vram_used"))
if total <= 0 {
return 0
}
if free := total - used; free > 0 {
return free
}
return 0
}
func readInt(path string) int64 {
b, err := os.ReadFile(path)
if err != nil {
return 0
}
n, err := strconv.ParseInt(strings.TrimSpace(string(b)), 10, 64)
if err != nil {
return 0
}
return n
}
// readComm names the contender for the log. The log is the instrument for the
// open question in Vikunja #488: whether a process can want this card without
// ever registering on the KFD, which a Vulkan or video-decode job would.
func readComm(pid int) string {
b, err := os.ReadFile(filepath.Join("/proc", strconv.Itoa(pid), "comm"))
if err != nil {
return "?"
}
return strings.TrimSpace(string(b))
}
+103
View File
@@ -0,0 +1,103 @@
package main
import (
"net/http"
"net/http/httptest"
"net/url"
"os"
"path/filepath"
"strconv"
"testing"
)
// fakeKFD builds the sysfs shape the workstation actually has: one directory
// per ROCm process, each holding a vram_<node> file. Sampled from the live box
// on 02-08-2026, where the CPT run appeared as proc/478104/vram_35881.
func fakeKFD(t *testing.T, vramByPID map[int]int64) string {
t.Helper()
root := t.TempDir()
for pid, vram := range vramByPID {
dir := filepath.Join(root, strconv.Itoa(pid))
if err := os.MkdirAll(dir, 0o755); err != nil {
t.Fatal(err)
}
f := filepath.Join(dir, "vram_35881")
if err := os.WriteFile(f, []byte(strconv.FormatInt(vram, 10)+"\n"), 0o644); err != nil {
t.Fatal(err)
}
}
return root
}
func TestForeignExcludesOurChild(t *testing.T) {
root := fakeKFD(t, map[int]int64{478104: 12791693312, 999: 4096})
p := probe{kfdRoot: root}
all := p.foreign(0)
if len(all) != 2 {
t.Fatalf("with no child running, both processes are foreign, got %d", len(all))
}
ours := p.foreign(999)
if len(ours) != 1 || ours[0].PID != 478104 {
t.Fatalf("our own llama-server must not count as a contender, got %+v", ours)
}
if ours[0].VRAM != 12791693312 {
t.Errorf("per-process VRAM = %d, want the value from vram_35881", ours[0].VRAM)
}
}
// An empty KFD tree is the state that permits a start, so it must read as empty
// rather than as an error the caller has to interpret.
func TestForeignEmptyAndMissing(t *testing.T) {
if got := (probe{kfdRoot: t.TempDir()}).foreign(0); len(got) != 0 {
t.Errorf("empty kfd tree: got %d processes, want 0", len(got))
}
if got := (probe{kfdRoot: "/nonexistent"}).foreign(0); got != nil {
t.Errorf("missing kfd tree: got %+v, want nil", got)
}
}
func TestFreeVRAM(t *testing.T) {
dev := t.TempDir()
write := func(name, v string) {
if err := os.WriteFile(filepath.Join(dev, name), []byte(v), 0o644); err != nil {
t.Fatal(err)
}
}
// The live numbers from the workstation while the CPT run held the card.
write("mem_info_vram_total", "17163091968\n")
write("mem_info_vram_used", "13396389888\n")
p := probe{drmDev: dev}
if got, want := p.freeVRAM(), int64(3766702080); got != want {
t.Errorf("freeVRAM = %d, want %d", got, want)
}
if got := (probe{drmDev: "/nonexistent"}).freeVRAM(); got != 0 {
t.Errorf("unreadable card reports %d free, want 0 so nothing starts", got)
}
}
// With no model loaded the supervisor must still answer, and it must answer 503
// rather than hanging or proxying into a closed port. Maven reads this endpoint
// on a timer forever, including while the workstation is busy.
func TestHealthAndProxyRefuseWhenNotReady(t *testing.T) {
s := &supervisor{run: newRunner("/bin/true", nil, "")}
h := s.handler(mustURL(t, "http://127.0.0.1:1"))
for _, path := range []string{"/health", "/v1/chat/completions"} {
w := httptest.NewRecorder()
h.ServeHTTP(w, httptest.NewRequest(http.MethodGet, path, nil))
if w.Code != http.StatusServiceUnavailable {
t.Errorf("%s with no model: got %d, want 503", path, w.Code)
}
}
}
func mustURL(t *testing.T, s string) *url.URL {
t.Helper()
u, err := url.Parse(s)
if err != nil {
t.Fatal(err)
}
return u
}
+247
View File
@@ -0,0 +1,247 @@
// mavgpud — the workstation's GPU supervisor.
//
// It runs on the workstation (an AMD 7900 GRE, 16GB), not on homesrv, and it is
// deployed separately from the Maven daemons. Maven does not participate in any
// of this and never asks for a start: it reads /health through internal/llm.Pair
// and either gets the big model or falls back to the resident 1.7B.
//
// The rule, from Vikunja #488: keep llama-server loaded whenever the card is
// free, unload it when it has been idle too long or when another process needs
// the card. Not on demand, because a 7-14B takes tens of seconds to load and a
// world question would be answered by a gap every time the card had been quiet.
// Not always on, because that holds 16GB against the owner's own jobs.
package main
import (
"context"
"encoding/json"
"flag"
"log"
"net/http"
"net/http/httputil"
"net/url"
"os"
"os/signal"
"sync/atomic"
"syscall"
"time"
)
type config struct {
Listen string `json:"listen"` // what Maven talks to
LlamaAddr string `json:"llama_addr"` // where llama-server binds
LlamaBin string `json:"llama_bin"`
// LlamaArgs must include the flags that bind LlamaAddr. They are passed
// through untouched so the model, context size and layer count stay the
// owner's business and not this daemon's schema.
LlamaArgs []string `json:"llama_args"`
KFDRoot string `json:"kfd_root"`
DRMDevice string `json:"drm_device"`
Poll duration `json:"poll"`
IdleTimeout duration `json:"idle_timeout"`
StopGrace duration `json:"stop_grace"`
MinFreeVRAM int64 `json:"min_free_vram_bytes"`
// EvictAfter and StartAfter are counted in polls, not seconds. Both exist
// to damp flapping: a one-tick blip from a short-lived rocm process must
// not evict the model, and a card that has just been released must not be
// grabbed before the previous job has finished unmapping.
EvictAfter int `json:"evict_after_polls"`
StartAfter int `json:"start_after_polls"`
}
func defaults() config {
return config{
Listen: ":8080",
LlamaAddr: "127.0.0.1:8081",
KFDRoot: "/sys/class/kfd/kfd/proc",
DRMDevice: "/sys/class/drm/card1/device",
Poll: duration(time.Second),
IdleTimeout: duration(15 * time.Minute),
StopGrace: duration(20 * time.Second),
MinFreeVRAM: 15 << 30,
EvictAfter: 2,
StartAfter: 5,
}
}
// duration lets the config file say "15m" instead of counting nanoseconds.
type duration time.Duration
func (d *duration) UnmarshalJSON(b []byte) error {
var s string
if err := json.Unmarshal(b, &s); err != nil {
return err
}
v, err := time.ParseDuration(s)
if err != nil {
return err
}
*d = duration(v)
return nil
}
func main() {
path := flag.String("config", "/etc/mavgpud.json", "config file")
flag.Parse()
cfg := defaults()
b, err := os.ReadFile(*path)
if err != nil {
log.Fatalf("mavgpud: read config: %v", err)
}
if err := json.Unmarshal(b, &cfg); err != nil {
log.Fatalf("mavgpud: parse config: %v", err)
}
if cfg.LlamaBin == "" {
log.Fatal("mavgpud: llama_bin is required")
}
base := "http://" + cfg.LlamaAddr
run := newRunner(cfg.LlamaBin, cfg.LlamaArgs, base+"/health")
sup := &supervisor{
cfg: cfg,
probe: probe{kfdRoot: cfg.KFDRoot, drmDev: cfg.DRMDevice},
run: run,
}
sup.touch()
ctx, cancel := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
defer cancel()
target, err := url.Parse(base)
if err != nil {
log.Fatalf("mavgpud: llama_addr: %v", err)
}
srv := &http.Server{Addr: cfg.Listen, Handler: sup.handler(target)}
go func() {
log.Printf("mavgpud: listening on %s, model %s", cfg.Listen, cfg.LlamaBin)
if err := srv.ListenAndServe(); err != nil && err != http.ErrServerClosed {
log.Fatalf("mavgpud: listen: %v", err)
}
}()
sup.loop(ctx)
// The card must come back before we do. A supervisor that exits leaving
// llama-server holding 14GB is worse than one that never ran.
shut, done := context.WithTimeout(context.Background(), 5*time.Second)
defer done()
_ = srv.Shutdown(shut)
run.stop(time.Duration(cfg.StopGrace))
}
type supervisor struct {
cfg config
probe probe
run *runner
lastReq atomic.Int64 // unix nanos of the last request Maven sent
foreignStreak int
clearStreak int
}
func (s *supervisor) touch() { s.lastReq.Store(time.Now().UnixNano()) }
func (s *supervisor) idle() time.Duration {
return time.Since(time.Unix(0, s.lastReq.Load()))
}
// handler serves the two things the workstation exposes.
//
// /health is answered locally and always, with no GPU cost and no round trip,
// because it is the only thing Maven reads and Maven reads it on a timer
// forever. Everything else is llama-server's API, reverse-proxied. Proxying
// rather than pointing Maven straight at llama-server is what makes the idle
// window measurable: the supervisor cannot otherwise know when the model was
// last used.
func (s *supervisor) handler(target *url.URL) http.Handler {
proxy := httputil.NewSingleHostReverseProxy(target)
mux := http.NewServeMux()
mux.HandleFunc("/health", func(w http.ResponseWriter, r *http.Request) {
if !s.run.isReady() {
http.Error(w, "model not loaded", http.StatusServiceUnavailable)
return
}
w.Header().Set("Content-Type", "application/json")
_, _ = w.Write([]byte(`{"status":"ok"}`))
})
mux.HandleFunc("/", func(w http.ResponseWriter, r *http.Request) {
if !s.run.isReady() {
http.Error(w, "model not loaded", http.StatusServiceUnavailable)
return
}
s.touch()
proxy.ServeHTTP(w, r)
})
return mux
}
func (s *supervisor) loop(ctx context.Context) {
t := time.NewTicker(time.Duration(s.cfg.Poll))
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
s.tick(ctx)
}
}
}
// tick is the whole decision. Yielding is checked before starting, and presence
// on the KFD is what triggers it — not a VRAM threshold. A ROCm process
// registers under /sys/class/kfd/kfd/proc when it initialises HIP, before it
// allocates, so we see a contender during its startup rather than after it has
// already failed to get the memory it wanted.
func (s *supervisor) tick(ctx context.Context) {
others := s.probe.foreign(s.run.pid())
if len(others) > 0 {
s.foreignStreak++
s.clearStreak = 0
} else {
s.foreignStreak = 0
s.clearStreak++
}
if s.run.running() {
s.run.refreshReady(ctx)
switch {
case s.foreignStreak >= s.cfg.EvictAfter:
log.Printf("mavgpud: yielding the card to %s", describe(others))
s.run.stop(time.Duration(s.cfg.StopGrace))
case s.idle() > time.Duration(s.cfg.IdleTimeout):
log.Printf("mavgpud: idle for %s, unloading", s.idle().Round(time.Second))
s.run.stop(time.Duration(s.cfg.StopGrace))
}
return
}
if s.clearStreak < s.cfg.StartAfter {
return
}
if free := s.probe.freeVRAM(); free < s.cfg.MinFreeVRAM {
return
}
s.touch() // the idle clock starts at load, not at the last request before it
if err := s.run.start(); err != nil {
log.Printf("mavgpud: start llama-server: %v", err)
}
}
// describe names the contenders in the log. This log is the instrument for the
// open question in #488: whether polling the KFD misses a job that wants the
// card without registering there.
func describe(procs []gpuProc) string {
out := ""
for i, p := range procs {
if i > 0 {
out += ", "
}
out += p.Comm
}
return out
}
+132
View File
@@ -0,0 +1,132 @@
package main
import (
"context"
"log"
"net/http"
"os/exec"
"sync"
"syscall"
"time"
)
// runner owns one llama-server process. Owning it is the point of the daemon:
// the workstation cannot keep a 7-14B resident, because that holds 16GB against
// the owner's CPT runs, Correx and the manga-recap pipeline. So the thing that
// stays up is this, which costs no VRAM, and the model comes and goes under it.
type runner struct {
bin string
args []string
// ready is llama-server's own /health, which answers "is a model loaded".
// Loading a 7-14B takes tens of seconds, so started is not ready.
readyURL string
mu sync.Mutex
cmd *exec.Cmd
ready bool
http *http.Client
}
func newRunner(bin string, args []string, readyURL string) *runner {
return &runner{
bin: bin, args: args, readyURL: readyURL,
http: &http.Client{Timeout: 2 * time.Second},
}
}
// pid is the child's, or 0. The GPU probe needs it to tell our own model apart
// from a contender.
func (r *runner) pid() int {
r.mu.Lock()
defer r.mu.Unlock()
if r.cmd == nil || r.cmd.Process == nil {
return 0
}
return r.cmd.Process.Pid
}
func (r *runner) running() bool { return r.pid() != 0 }
// isReady reports the cached readiness. The supervisor loop refreshes it; the
// health handler only reads, so answering /health never costs a round trip.
func (r *runner) isReady() bool {
r.mu.Lock()
defer r.mu.Unlock()
return r.ready
}
// start launches llama-server. It returns as soon as the process exists, not
// when the model is loaded.
func (r *runner) start() error {
r.mu.Lock()
defer r.mu.Unlock()
if r.cmd != nil {
return nil
}
cmd := exec.Command(r.bin, r.args...)
// Own process group, so stop kills anything llama-server spawned rather
// than leaving it holding VRAM after we have declared the card yielded.
cmd.SysProcAttr = &syscall.SysProcAttr{Setpgid: true}
if err := cmd.Start(); err != nil {
return err
}
r.cmd, r.ready = cmd, false
log.Printf("mavgpud: started llama-server pid=%d", cmd.Process.Pid)
go func() {
err := cmd.Wait()
r.mu.Lock()
r.cmd, r.ready = nil, false
r.mu.Unlock()
log.Printf("mavgpud: llama-server exited: %v", err)
}()
return nil
}
// stop ends llama-server and waits for the VRAM to come back. SIGTERM first so
// it unmaps cleanly, SIGKILL after the grace window. Returning before the
// process is gone would let the supervisor report a free card while 14GB is
// still mapped, which is the one lie that would make yielding useless.
func (r *runner) stop(grace time.Duration) {
r.mu.Lock()
cmd := r.cmd
r.ready = false
r.mu.Unlock()
if cmd == nil || cmd.Process == nil {
return
}
pgid := -cmd.Process.Pid
_ = syscall.Kill(pgid, syscall.SIGTERM)
deadline := time.Now().Add(grace)
for time.Now().Before(deadline) {
if !r.running() {
return
}
time.Sleep(100 * time.Millisecond)
}
log.Printf("mavgpud: llama-server did not exit in %s, killing", grace)
_ = syscall.Kill(pgid, syscall.SIGKILL)
}
// refreshReady asks llama-server whether the model is loaded. Called once per
// supervisor tick, never per request.
func (r *runner) refreshReady(ctx context.Context) {
if !r.running() {
return
}
ok := false
req, err := http.NewRequestWithContext(ctx, http.MethodGet, r.readyURL, nil)
if err == nil {
resp, err := r.http.Do(req)
if err == nil {
ok = resp.StatusCode == http.StatusOK
resp.Body.Close()
}
}
r.mu.Lock()
was := r.ready
r.ready = ok
r.mu.Unlock()
if ok && !was {
log.Printf("mavgpud: model ready")
}
}
+5 -8
View File
@@ -512,7 +512,7 @@ func main() {
// /tools — the authed enable surface. maven proposes acts she can't run;
// this page is where a human reviews and enables them (proposed→enabled).
// Enabling is the boundary-moving act (DESIGN.md § Tool registration —
// Enabling is the boundary-moving act (docs/design.md § Tool registration —
// drafting is suggest, enabling is act), so it lives ONLY here,
// behind wg+nginx+auth — never the voice/chat path.
mux.HandleFunc("/tools", func(w http.ResponseWriter, r *http.Request) {
@@ -1210,13 +1210,10 @@ func routineRows(rs []ipc.ProposedRoutine) []routineRow {
return out
}
// acceptRoutine creates the recurring reminder for a proposal, then marks the
// proposal accepted and links the reminder to it. Weekly patterns get a cron
// expression; any other interval fires once.
//
// TODO(vikunja#46): this mirrors the voice accept path in cmd/mavend/voice.go.
// When the tick loop learns to read accepted proposals directly, both callers
// should hand off to one place in core instead of each building a reminder.
// acceptRoutine marks a proposal accepted. This page is the ONLY surface that
// may do it (Vikunja #367): accepting gives the tick loop a standing new
// reason to speak, which DESIGN.md puts at layer 3, and the button here is
// behind step-up. Voice can park the question and dismiss, never accept.
func acceptRoutine(ctx context.Context, core ipc.CoreAPI, id int64) error {
proposed, err := core.ListProposedRoutines(ctx)
if err != nil {
+68 -1
View File
@@ -39,6 +39,23 @@
"proxy": "socks5://192.168.240.1:10808"
},
"//workstation": [
"The big model on the desk PC (workpc, 7900 GRE 16GB), fronted by",
"mavgpud on port 8080. It runs gemma-4-12b and it is preferred over the",
"resident Qwen3-1.7B for routing and replies whenever the card is free.",
"The machine is never assumed up: it sleeps, and the card is often held by",
"a CPT run, in which case mavgpud answers 503 and Maven falls back to the",
"resident model without saying so. Deleting this block restores exactly",
"the behaviour homesrv had before it existed.",
"Addressed by LAN address, not container name: mavgpud runs on another",
"machine and there is no shared docker network to name it on."
],
"workstation": {
"url": "http://192.168.1.105:8080",
"probe": "15s",
"timeout": "90s"
},
"//search": [
"The live web, searched after his own notes and before Kiwix. Only the",
"query string leaves the box — never a note, a fact, the persona block or",
@@ -76,6 +93,56 @@
"snippet_runes": 1500
},
"//morning_routines": [
"The daily checklist (Vikunja #280). Each item is done when its fact_key",
"gets a non-voided fact inside the window, so 'выпил воды' closes water and",
"nothing has to be ticked by hand. nudge_at fires once, at the end of the",
"window, and only for what is still open. Weekdays empty = every day."
],
"morning_routines": [
{
"name": "утро",
"window_start": "08:00",
"window_end": "11:00",
"nudge_at": "10:30",
"severity": 1,
"items": [
{ "key": "medicine", "fact_key": "medicine", "label": "лекарство" },
{ "key": "water", "fact_key": "water", "label": "вода" },
{ "key": "pets", "fact_key": "pets", "label": "покормить кота" }
]
}
],
"//feeds": [
"RSS reading (Vikunja #258). Every item lands as a note with source",
"rss:<name>, which is also what puts entries in the intake journal that",
"/events reads. Only the feed URL leaves the box.",
"This is a starting pair, not a curated set — trim or extend it."
],
"feeds": {
"poll_interval": "30m",
"max_items": 5,
"max_age": "24h",
"sources": [
{ "name": "lwn", "url": "https://lwn.net/headlines/newrss", "category": "технологии" },
{ "name": "archlinux", "url": "https://archlinux.org/feeds/news/", "category": "технологии" }
]
},
"//crawl": [
"Reading a web page (Vikunja #259). on_demand answers 'посмотри <URL>'.",
"No allow_hosts, so any public host he names is readable; private",
"addresses are refused unconditionally by internal/webfetch and do not",
"need listing. Setting allow_hosts here would also narrow on-demand,",
"which is the point of leaving it empty."
],
"crawl": {
"on_demand": true,
"timeout": "10s",
"max_runes": 4000
},
"digest": {
"enabled": true,
"window": "30m",
@@ -119,7 +186,7 @@
"timeout": "400ms",
"rate": 100,
"max_hosts": 256,
"enabled": false
"enabled": true
},
"nexus": { "url": "http://nexus:9740" },
+32
View File
@@ -0,0 +1,32 @@
{
"listen": ":8080",
"llama_addr": "127.0.0.1:10000",
"llama_bin": "llama-server",
"llama_args": [
"-m", "/mnt/D/AI/gemma4/gemma-4-12B-it-qat-UD-Q4_K_XL.gguf",
"-md", "/mnt/D/AI/gemma4/mtp-gemma-4-12B-it-BF16.gguf",
"-ngl", "99",
"-fa", "on",
"-np", "1",
"--host", "127.0.0.1",
"--port", "10000",
"--ctx-size", "32768",
"--threads", "6",
"--batch-size", "2048",
"--ubatch-size", "512",
"--jinja",
"--chat-template-kwargs", "{\"enable_thinking\":false}",
"--spec-type", "draft-mtp",
"--spec-draft-n-max", "2"
],
"kfd_root": "/sys/class/kfd/kfd/proc",
"drm_device": "/sys/class/drm/card1/device",
"poll": "1s",
"idle_timeout": "15m",
"stop_grace": "20s",
"min_free_vram_bytes": 10737418240,
"evict_after_polls": 2,
"start_after_polls": 5
}
+24
View File
@@ -0,0 +1,24 @@
[Unit]
# Runs on the workstation (bugmachine), not on homesrv. Install as a systemd
# user unit and turn on lingering, so the card is supervised after a reboot
# with nobody logged in:
#
# scp mavgpud workpc:~/.local/bin/mavgpud
# scp deploy/mavgpud.json workpc:~/.config/mavgpud.json
# scp deploy/mavgpud.service workpc:~/.config/systemd/user/mavgpud.service
# ssh workpc 'systemctl --user daemon-reload && systemctl --user enable --now mavgpud'
# sudo loginctl enable-linger kami
Description=Maven GPU supervisor (holds llama-server while the card is free)
After=network.target
[Service]
ExecStart=%h/.local/bin/mavgpud -config %h/.config/mavgpud.json
Restart=always
RestartSec=5
# The card must come back when the supervisor goes down. mavgpud stops
# llama-server on SIGTERM, so give it longer than stop_grace to do that.
KillSignal=SIGTERM
TimeoutStopSec=60
[Install]
WantedBy=default.target
+9
View File
@@ -79,7 +79,16 @@ services:
<<: *image
# voice.bind is 0.0.0.0:9100 in deploy/mavend.json so mavweb can reach it
# cross-container. Verified 2026-07-06.
# -ambient-token turns on POST /api/ambient (Vikunja #126): the phone posts
# notification text, mavweb keeps only a meeting time. Empty ⇒ no route at
# all, which is what a missing MAVEN_AMBIENT_TOKEN gives. The value comes
# from the gitignored .env docker compose reads for interpolation, NOT from
# an env_file — flags are interpolated before any service env exists.
# Weakness worth naming: mavweb takes this as a flag, so it is visible in
# `ps` inside this container, unlike the zenmoney and IMAP secrets which are
# read from files.
command: ["mavweb", "-addr", ":9201", "-voice", "mavend:9100", "-core", "/run/maven/mavend.sock",
"-ambient-token", "${MAVEN_AMBIENT_TOKEN:-}",
"-nexus", "http://nexus:9740", "-praxis", "http://praxis:8989", "-hexis", "http://hexis:9741"]
depends_on: [mavend]
# loopback-only on purpose: /tools defines+executes arbitrary argv and
@@ -1,6 +1,45 @@
# maven — feature ranking
> dated 2026-07-03. companion to `DESIGN.md` (folded from the former `maven.md`). ranks everything discussed post-repo-state against the infra blockers, not a replacement for the build order.
> **Archived 2026-08-02 (V-447).** The mandatory and easy tiers are now Vikunja tasks
> 449-458. Two of those closed immediately, because the ranking was stale. Quiet hours
> (V-450) ship as `QuietHoursConfig` plus the care gate in `internal/loop/loop.go`.
> Schema migrations (V-451) ship as `internal/store/migrations.go` on `PRAGMA
> user_version`. The doable and epic tiers stay here because they are reasoning. Some of
> them exist only to record why something is not worth doing yet. Read this for the why,
> not as a work queue, and check the code before believing a gap.
> dated 2026-07-03. companion to `docs/design.md` (folded from the former `maven.md`). ranks everything discussed post-repo-state against the infra blockers, not a replacement for the build order.
---
## what already shipped (checked against the code, 2026-08-02)
One month old and already wrong in ten places. Everything below is marked
built after reading the code, not the board. Read the tiers underneath with this list in
hand.
Infra 1 and 2, the two the ranking says block every feature, are both done. sqlcipher
at-rest ships as `Store.enc` plus `OpenEncrypted`, a tmpfs working copy re-encrypted on
`Close`, keyed from `db_key_env` in `deploy/mavend.json`. mavweb and mavcaldav are no
longer at zero coverage: seven test files under `cmd/mavweb`, including
`credentials_test.go` and `passkey_prf_test.go`, and two under `cmd/mavcaldav`. Infra 3
is stale in the other direction. There are still no systemd units, but the deploy is
`deploy/ecosystem/docker-compose.yml`, not scripts and tmux.
Doable tier, built: rule trace and explanation as `/trace` plus `internal/loop/explain.go`.
Recurring reminders as the cron column, `NextFireTs` and `RescheduleReminder`
(`internal/store/reminders.go:206`), so "fires once right now" is wrong. Stale-reminder
burst collapse as `collapseReminders` (`internal/loop/gather.go:209`). Revert as
`VoidLatestFact` (`internal/store/facts.go:300`). Digest mode as `internal/store/digest.go`.
Testing infra as the simulator and the eval lab (V-284, V-278). Passkey persistence as the
JSON-backed `credentialStore` in `cmd/mavweb/credentials.go`. Most integrations shipped as
their own QA tasks (V-246 mail, V-256 smarthome, V-258 rss, V-259 crawler).
Doable tier, still open: correx, systemd units, memory decay and duplicate detection
(nothing in `internal/memory` touches it), backup automation, import and export, barge-in.
Epic tier, unbuilt as ranked. `event.Bus` exists (`cmd/mavend/intake.go:57`) but it is the
intake journal from V-283, not the rewrite of facts into projections that this tier means.
---
@@ -20,7 +59,7 @@ nothing feature-level below should land before 12 are done. 34 can interle
### mandatory
things that block correctness or safety of stuff already shipped — not new capability, just closing gaps in existing design.
- **destructive-confirm policy** — open question in `DESIGN.md` § open questions, blocks correx and any new tool domain from having a coherent risk tier
- **destructive-confirm policy** — open question in `docs/design.md` § open questions, blocks correx and any new tool domain from having a coherent risk tier
- **quiet-hours definition** — open question, blocks proactive delivery being trustworthy
- **schema migrations** — sqlcipher rollout alone forces a schema touch. want this mechanism before that, not after.
@@ -30,7 +69,7 @@ cheap, no dependencies, no new invariants.
- **grocery / `list_items` table** — fourth append-only shape (item, status, list-tag), no predicate touches it, multi-adder just works for free
- **go.mod tidy**
- **capability model** (deepseek) — `homelab.docker.restart` instead of flat `tool→enabled`. cheap now, expensive to retrofit once tools surface passes ~15 entries. time-sensitive, not urgent.
- **conversation repair** — already free: `DESIGN.md` has "misroute correction = new centroid example," this is just naming the existing mechanism as a feature
- **conversation repair** — already free: `docs/design.md` has "misroute correction = new centroid example," this is just naming the existing mechanism as a feature
- **command history** — read-only query over existing facts, no new mechanism
- **clarification templates** — canned phrasing for the router's existing confidence-gate fallback, phraser-lane only
- **pronunciation dictionary** — tts config, no architecture
@@ -30,7 +30,7 @@ critical workflows: voice turn (mic→STT→route→tool/reply→TTS); proactive
(/dash /chat /tools /ecosystem)
current state: all 34 test packages pass; vet clean; `make test` still exits 1
(finding 2)
known failures: weak RU query routing — REARCH.md names the cause; the named fix
known failures: weak RU query routing — docs/rearchitecture.md names the cause; the named fix
is wired `nil`
maintenance burden: 4,518 lines of root markdown vs 33,319 lines of Go; 15 top-level
.md files, 3 of them dated session logs; several contradict
@@ -47,11 +47,11 @@ what's obsolete: llmrouter.go (built, tested, never wired); classifier seed-p
```
**Classification: healthy + misaligned.** Not fragile, not overbuilt, not abandoned. The
architecture in `REARCH.md` is sound and mostly *built* — it just is not *connected*.
architecture in `docs/rearchitecture.md` is sound and mostly *built* — it just is not *connected*.
## what it should become
The thing `REARCH.md` already describes, with the switch flipped and the drift removed:
The thing `docs/rearchitecture.md` already describes, with the switch flipped and the drift removed:
one resident small model doing both routing and phrasing, classifier demoted from the live
path to the failure floor, embedder demoted to RAG hint. No new architecture is needed.
**The gap is a config/wiring decision plus doc convergence, not a redesign.**
@@ -65,19 +65,19 @@ path to the failure floor, embedder demoted to RAG hint. No new architecture is
`architecture` / `repair`
**problem:** The most load-bearing design decision in the project is stated four different,
incompatible ways, and the code path `REARCH.md` calls "the linchpin" is disabled.
incompatible ways, and the code path `docs/rearchitecture.md` calls "the linchpin" is disabled.
**evidence** (all confirmed):
- `cmd/mavend/voice.go:211``rtr := buildRouter(emb, matcher, threshold, nil) // LLM router disabled`,
with comment *"the classifier handles routing reliably."*
- `REARCH.md:11` says the same classifier is *"the structural cause of 'she messes up
- `docs/rearchitecture.md:11` says the same classifier is *"the structural cause of 'she messes up
queries.'"* **The code comment and the design doc make opposite claims about the same
component.**
- `internal/router/llmrouter.go` (139 lines) + `llmrouter_test.go` — fully built and
tested, zero non-test callers.
- Model identity, four ways: docs say **Qwen3-1.7B** (`CLAUDE.md:6`, `REARCH.md:15`,
`SPEC.md:46`, `AGENTS.md:79`, `MAVEN_ECOSYSTEM_ARCHITECTURE.md:72`);
- Model identity, four ways: docs say **Qwen3-1.7B** (`CLAUDE.md:6`, `docs/rearchitecture.md:15`,
`SPEC.md:46`, `AGENTS.md:79`, `docs/ecosystem.md:72`);
`deploy/mavend.json:9` says **Qwen3.5-2B-UD-Q4_K_XL**; `models/llm/` on disk holds
**LFM2.5-1.2B-Thinking**; code comments in 5 files still say **LFM**.
- `deploy/mavend.json:11` sets `"n_gpu_layers": 99` while `CLAUDE.md:4` states the target
@@ -100,7 +100,7 @@ match. Delete nothing from `internal/router` yet — the classifier is the fallb
reconciliation, not a refactor. **Do not rewrite the router.**
**alternatives:** Delete `llmrouter.go` and commit to the classifier — only defensible if
the eval harness shows the classifier is actually adequate, which contradicts `REARCH.md`.
the eval harness shows the classifier is actually adequate, which contradicts `docs/rearchitecture.md`.
**risk:** Low-moderate. LLM route failures already fall through to the classifier
(`router.go:88-96`), so a bad model cannot break a turn. The real risk is CPU latency.
@@ -228,21 +228,21 @@ after finding 1, not before** — and skip it if it stays purely cosmetic.
**problem:** 15 root markdown files, 4,518 lines, several stale or superseded, at least
three pairs contradicting each other.
**evidence:** `ROADMAP.md` (759) + `MAVEN_ECOSYSTEM_ARCHITECTURE.md` (884) +
**evidence:** `ROADMAP.md` (759) + `docs/ecosystem.md` (884) +
`PROGRESS.md` (456) + `maven.md` (413) + `20-07-2026-BACKLOG.md` (396) +
`SESSION-05-07-2026.md` + `SESSION-06-07-2026.md` (477 combined) + `PLANS.md` (25) +
`START.md` + `SPEC.md` + `PROTOCOL.md`. `PROGRESS.md:61` annotates its own staleness:
*"Older LFM references below describe the currently deployed..."*. `REARCH.md` announces it
`docs/operations.md` + `SPEC.md` + `docs/protocol.md`. `PROGRESS.md:61` annotates its own staleness:
*"Older LFM references below describe the currently deployed..."*. `docs/rearchitecture.md` announces it
"supersedes" a model still described as current elsewhere.
**impact:** The doc set is the reason finding 1 exists. When five documents describe the
architecture, the code becomes the only trustworthy one — which defeats the purpose of
having them.
**recommended action:** Keep `CLAUDE.md` (agent contract), `REARCH.md` (target
architecture), `AGENTS.md` (recipes), `PROTOCOL.md` (wire format),
**recommended action:** Keep `CLAUDE.md` (agent contract), `docs/rearchitecture.md` (target
architecture), `AGENTS.md` (recipes), `docs/protocol.md` (wire format),
`20-07-2026-BACKLOG.md` (live queue). Delete the two `SESSION-*.md` and `PLANS.md` — git
history holds them. Fold `SPEC.md` + `maven.md` + `ROADMAP.md` into one `DESIGN.md` and
history holds them. Fold `SPEC.md` + `maven.md` + `ROADMAP.md` into one `docs/design.md` and
mark superseded sections instead of leaving them to read as current. Target ~1,500 lines.
---
@@ -320,7 +320,7 @@ expected maintenance gain: none over the incremental path
suggesting Vulkan offload is intended and working — but `CLAUDE.md` says CPU-only. Likely
the doc is stale, not the config; unverified.
- **Whether the classifier is genuinely adequate.** `voice.go:211` asserts it is;
`REARCH.md` asserts it is not. Both are claims, neither is measured. The uncommitted eval
`docs/rearchitecture.md` asserts it is not. Both are claims, neither is measured. The uncommitted eval
harness is the instrument to settle it — resolve before flipping the router, not after.
- **Whether wg+nginx+auth actually fronts 9201 in production.** Not in this repo. If it
does, finding 3 drops from "unauthenticated RCE" to "the control is not reproducible from
@@ -350,7 +350,7 @@ Key claims independently re-verified against the working tree; the verdict stand
over WireGuard on homesrv, this is hygiene, not an emergency — but the loopback bind
and startup warning are cheap insurance either way, so do them regardless.
- Finding 1's "flip the router" step should be gated harder on measurement.
`REARCH.md`'s claim that the classifier causes weak RU queries is itself unmeasured —
`docs/rearchitecture.md`'s claim that the classifier causes weak RU queries is itself unmeasured —
the review admits this under uncertainties, but the "repair now" ordering buries it.
Run `eval_scenarios_test.go` against both paths **before** deciding to flip, not
after. A 2B model on CPU may add enough latency that the classifier wins in practice
+10 -8
View File
@@ -1,10 +1,12 @@
# Maven — Design
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
> Folded 2026-07-30 from `SPEC.md` (north star, 2026-07-03), `maven.md`
> (consolidated decisions, 2026-06-30) and `ROADMAP.md` (execution plan,
> 2026-07-06). Those three files are gone; git history holds them.
> This is the single design document: principles, target state, and the
> execution ledger. `REARCH.md` remains authoritative wherever it disagrees
> execution ledger. `docs/rearchitecture.md` remains authoritative wherever it disagrees
> with anything here. Everything the three sources asserted that is no longer
> the intended design is preserved under **§ Superseded** — do not read that
> section as current.
@@ -160,7 +162,7 @@ presence_state ( last_bucket, last_score, updated_ts )
```
Facts additionally carry `Subject`/`EntityID`/`ResolutionState` for
entity-aware resolution against Nexus (see `MAVEN_ECOSYSTEM_ARCHITECTURE.md`).
entity-aware resolution against Nexus (see `docs/ecosystem.md`).
### Trigger model
@@ -196,7 +198,7 @@ INTO the gate as an env predicate, not the LLM's job.
## Reactive path — routing
**Target design: LLM-as-router** (see `REARCH.md` and `CLAUDE.md`). One
**Target design: LLM-as-router** (see `docs/rearchitecture.md` and `CLAUDE.md`). One
resident model emits GBNF-constrained structured JSON, and the same model
phrases replies; the embedder is a RAG hint, not a routing gate. The
committed default today is the classifier/embedder cascade, which is an
@@ -655,7 +657,7 @@ daemon.
The voice wire protocol (length-prefixed JSON frames over TCP) is designed for
**multiple client implementations**. The reference PWA at `cmd/mavweb` is one
client; any app (phone, desktop CLI, smartwatch) can implement the same frame
protocol. The published spec is `PROTOCOL.md` — **generated from
protocol. The published spec is `docs/protocol.md` — **generated from
`internal/voice/wire.go`**, not composed freehand, so it can't drift from
code. It covers transport (4-byte big-endian length prefix), methods
(`PushToTalk`, `Pong`), push kinds (`AudioNudge`), surface identity
@@ -670,8 +672,8 @@ Broadening to home automation, media or comms is JSON, not code.
## Execution ledger
Condensed from `ROADMAP.md` (2026-07-06). The live queue is
`20-07-2026-BACKLOG.md`; current state is `PROGRESS.md`.
Condensed from `ROADMAP.md` (2026-07-06). The live queue is the Vikunja board
(project Maven, ID 2); this table is history, not a work list.
| # | Item | Prio | Status |
|---|------|------|--------|
@@ -755,7 +757,7 @@ Kept for provenance. **None of this is the current or intended design.**
stay deterministic — "classifier owns the route, the SLM stays in its
phrasing lane" — with an embedding + nearest-centroid stage 1 over ~10
examples per intent, and misroutes appended as new centroid examples.
*Replaced by* LLM-as-router (`REARCH.md`): one resident model emits
*Replaced by* LLM-as-router (`docs/rearchitecture.md`): one resident model emits
GBNF-constrained JSON and also phrases replies; the embedder is demoted to
a RAG hint. *Landed 2026-07-31:* the LLM router is on by default and set
`true` in `deploy/mavend.json`. The classifier cascade stays as the failure
@@ -774,7 +776,7 @@ Kept for provenance. **None of this is the current or intended design.**
*Resolved 2026-07-30 (#318), revised 2026-07-31:* the resident checkpoint is
stock **Qwen3-1.7B** (`UD-Q4_K_XL`, `n_ctx` 4096), which replaced
Qwen3.5-0.8B after measuring better on both fixtures
(`MODEL-BAKEOFF-31-07-2026.md`). The CPT'd **Qwen3-1.7B** remains the target
(`docs/evals/2026-07-31-model-bakeoff.md`). The CPT'd **Qwen3-1.7B** remains the target
(#122); what stock gets wrong is the persona, not the Russian. Note the resident
model is no longer described as untrained — the target is trained
end-to-end, which is the substantive change from the old claim.
@@ -1,5 +1,7 @@
# Deterministic logic around a small model
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
Written 2026-08-02. Branch `fix/integrated`.
## The question
@@ -268,7 +270,7 @@ rebuilt `mavend` wires it. "кто написал войну и мир?" now rou
that turn left unsettled.
- **The turn was slow, and nobody knows yet whether that is real.** Route 7s,
search 1s, phrasing 15s. The p50 in `ROUTING-EVAL-31-07-2026.md` is 825ms. It
search 1s, phrasing 15s. The p50 in `docs/evals/2026-07-31-routing.md` is 825ms. It
was the first turn after a cold start with the model still warming, so it
proves nothing either way. Re-run the same question warm before treating it as
a regression. Do not plan latency work off this number.
@@ -1,5 +1,7 @@
# Maven Ecosystem Architecture
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
## 1. Purpose
This document defines Maven's role in the local ecosystem formed by:
@@ -16,7 +16,7 @@ with Qwen3-1.7B.
Settles Vikunja **#278 / #250**.
- Same fixture and scorer as `ROUTING-EVAL-31-07-2026.md`: `internal/router/eval/`
- Same fixture and scorer as `docs/evals/2026-07-31-routing.md`: `internal/router/eval/`
(`ru_routing_v1.json`, 76 held-out cases).
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:<port> make eval-router`
(`TestLLMRouterBaseline`). (This line used to say there is no `make eval-models` target.
@@ -148,7 +148,7 @@ Qwen3-1.7B wins every column, including against a model 20% larger than it.
| ontopic | 16, 19, 19 | **22, 23, 23** |
| canned fallbacks | 8, 5, 6 | **0, 2, 0** |
This also fills the row `TALK-EVAL-31-07-2026.md` had to void for contamination:
This also fills the row `docs/evals/2026-07-31-talk.md` had to void for contamination:
**600ch/1024tok on Qwen3.5-0.8B scores 13, 11, 8.**
`address` is the headline. It sat at 18-22 of 27 on the 0.8B no matter how the prompt
@@ -164,7 +164,7 @@ The 1.7B does that 0-2 times.
> **Stale, corrected 2026-08-02.** The p50 figures in this table are contention on a
> shared llama-server, not the model's cost. The router measures p50 825ms / p95 1.2s /
> max 3.0s in `ROUTING-EVAL-31-07-2026.md`, which says so at line 61. Read this table for
> max 3.0s in `docs/evals/2026-07-31-routing.md`, which says so at line 61. Read this table for
> the shape of the tail only. Take absolute latency from the routing eval.
| | p50 | p95 |
@@ -220,11 +220,11 @@ swapped again when the CPT lands.
behind `voice.llm_router`, the default is on, and `deploy/mavend.json` sets it `true`.
These numbers are the production path now. **Corrected 2026-08-02: the p50 ≈2.7s in the
latency table above WAS a bench artifact.** It is contention on the shared llama-server,
not the model. `ROUTING-EVAL-31-07-2026.md` line 61 says so, and measures the router at
not the model. `docs/evals/2026-07-31-routing.md` line 61 says so, and measures the router at
p50 825ms / p95 1.2s / max 3.0s. Cite that file for latency, not this one.
- ~~`/mnt/hdd1/llms/LFM2.5/Qwen3-1.7B-UD-Q4_K_XL.gguf` is a 293 MB truncated download
in the wrong directory.~~ **Deleted 2026-07-31.** The good 1.13 GB copy in `qwen3/` is
what `deploy/mavend.json` loads.
- Harness: `scratchpad/bakeoff.sh`, one server at a time, health-checked before each
run, `/v1/models` recorded per run. Never run two LLM consumers at once — see the
contamination note in `TALK-EVAL-31-07-2026.md`.
contamination note in `docs/evals/2026-07-31-talk.md`.
@@ -1,7 +1,7 @@
# Phrasing evaluation — 31-07-2026
How Maven words a nudge, measured instead of argued. Counterpart to
`ROUTING-EVAL-31-07-2026.md`.
`docs/evals/2026-07-31-routing.md`.
- Fixture + scorer: `internal/phraser/eval/` (`nudges_v1.json`, 15 cases; `eval.go`, `checks.go`)
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing`
@@ -65,7 +65,7 @@ prefixes) is the targeted fix, and it would move findings 1 and 2 together. Sepa
`model_quantized.onnx` — not the same file.
`hard` cases score **2/11**: every one is a query where the operator did not reuse his own words.
That is the normal case weeks later, and exactly what DESIGN.md's "recall when relevant" promises.
That is the normal case weeks later, and exactly what docs/design.md's "recall when relevant" promises.
### 4. The memory-store recall branch is dead for notes
@@ -177,7 +177,7 @@ was silent ("не знаю" to "который час") while the one it introdu
### 1. The resident model does route better — 50.0% vs 36.8%
REARCH.md's premise holds; `voice.go:211`'s comment does not. **But the classifier is only
docs/rearchitecture.md's premise holds; `voice.go:211`'s comment does not. **But the classifier is only
~37% correct on held-out utterances, and the model only ~50%.** Neither is "reliable". The
gap between them is real but both are far from a system you would describe as working.
@@ -141,7 +141,7 @@ a model check, which catches a dead server but not a loaded one.
measured) on this fixture and the router fixture. Not the 4B — too big for
this box, owner's call.
- Newer sub-500M candidates (LFM2.5 200M/300M) are worth a run for routing.
Note `MODEL-BAKEOFF-31-07-2026.md` found LFM2.5-**1.2B** worse than
Note `docs/evals/2026-07-31-model-bakeoff.md` found LFM2.5-**1.2B** worse than
Qwen3.5-0.8B at Russian routing and 2.4× slower — but those are a different,
older generation, so that result does not predict the small ones.
- Fix `chat-how-are-you`'s `want_any`, and re-baseline once, so `ontopic`
@@ -0,0 +1,82 @@
# gemma-4-12b on the workstation, against the resident Qwen3-1.7B
Measured 2026-08-02 on the fixtures as they stand. Dated file: it is not edited
after today, and a newer number is a new file.
Vikunja #485's first assumption was that a 7-14B measurably beats Qwen3-1.7B on
the 77-case RU routing fixture and the 27-case talk fixture. It does, on both,
and it is also faster.
## The setup
`gemma-4-12B-it-qat-UD-Q4_K_XL` with the `mtp-gemma-4-12B-it-BF16` draft model,
served by `llama-server` b10220 on bugmachine (AMD 7900 GRE, 16GB), fronted by
`mavgpud` on `192.168.1.105:8080`. Thinking is off through
`--chat-template-kwargs '{"enable_thinking":false}'`, speculative decoding is
`--spec-type draft-mtp --spec-draft-n-max 2`, context 32768. The exact line is
`deploy/mavgpud.json`.
Every number below crossed the LAN from homesrv. Note the trap: homesrv's shell
exports `HTTP_PROXY`, Go honours it, and the runs need
`env -u HTTP_PROXY -u HTTPS_PROXY -u http_proxy -u https_proxy`.
## Routing, 77-case RU fixture
| | full | intent-only | p50 | p95 |
|---|---|---|---|---|
| classifier alone (02-08) | 68.8% | — | 16.6µs | — |
| Qwen3-1.7B through the cascade (31-07, 02-08) | 72.7% | 77.9% | 0.80-1.04s | — |
| **gemma-4-12b through the cascade** | **84.4%** | **93.5%** | **329ms** | 429ms |
| gemma-4-12b alone, no cascade | 55.8% | 85.7% | 335ms | 436ms |
The workstation buys 11.7 points of full accuracy over the resident model. It
buys 15.6 points of intent-only, at a third of the latency. The router's p50 was
never the model's fault, which the 02-08 contention finding already said. A 12B
on a free 16GB card answers a routing turn in a third of a second.
Two things the table hides.
The alone-versus-cascade gap is slots, not intents. gemma reads the intent right
85.7% of the time on its own. It loses full accuracy on seven fact keys
(`вода` instead of `water`, `ужин` instead of `meal`) and on six reminder times
with no time slot. Stage 0 and the daemon's own extractor repair
both, which is why the cascade is 28 points higher. The lesson is that the
cascade earns its keep even under a much better model, not that it is scaffolding
to remove.
`errors: 6` in the alone row are declines on single-token and ambiguous
utterances, all of which the cascade caught. The remaining defects through the
cascade are three `query→fact` confusions, one `chat→query`, and one false
clarify.
## Talk, 27-case conversational fixture
| | pass | notes |
|---|---|---|
| Qwen3-1.7B (31-07) | 20/27 | 11-17/27 for the 0.8B before it |
| **gemma-4-12b** | **25/27 (92.6%)** | chat 8/9, knowledge 9/9, query 8/9 |
Knowledge is the interesting column: 9/9, in Russian, with real answers about
Rayleigh scattering, SSD versus HDD and thunder delay. That is the case the
1.7B cannot do at all and the reason the naming half of the degradation rule
exists.
Two failures, and one of them is the persona defect the CPT (#122) targets:
`query-notes-do-not-answer` wrote `заплатил` where Maven needs the feminine
form. The other is `chat-joke`, where the model told a joke without using any of
the words the check looks for. Run-to-run variance is about one case: a second
run scored 24/27 with `chat-followup-server` also off-topic.
## Nudge phrasing, 15-case fixture
15/15, every check, no errors. `mood`, `lang`, `length`, `feminine`,
`hisgender`, `address`, `cringe` and `ontopic` all clean.
## What this settles and what it does not
Settled: the size question. A 12B on the workstation beats the resident model on
every fixture we have, and it is faster. The offload argument holds.
Not settled: how often the card is free. That is #485's second assumption and
only the `mavgpud` log answers it, after a week of the owner's normal work. A
model that is better whenever it is up is worth little if it is never up.
+179
View File
@@ -0,0 +1,179 @@
# Offloading model work to the workstation
*Last verified: 2026-08-03 @ 12530c8. Living doc: correct it in place, do not append.*
Owner's call, 2026-08-02. Vikunja #483 is the umbrella. Tasks #484 to #487 are the
work, and this file holds the shape and the rules all four must obey.
## The goal
homesrv cannot grow a GPU. The workstation has 16GB of VRAM. Move the model work
to the workstation and leave homesrv running the logic that must be always-on,
deterministic and cheap.
## Why this is tractable
The split already exists structurally. `mavsttd` and `mavttsd` are separate
daemons that core reaches over a socket, not linked libraries. Moving them off-box
is a transport change, not a redesign.
The microphone is at the workstation, because that is where the owner sits and
homesrv is headless. So speech-to-text and the wake word are already on the
workstation side by construction. Audio never has to cross the LAN. Only the core
turn does.
## The constraint that shapes everything
The workstation's GPU is often busy: CPT runs, experiments, Correx, the manga-recap
pipeline. It also sleeps. homesrv does not.
So an offloaded model is never *the* model. It is the preferred one, with a floor
on homesrv. That is the shape the cascade already has, where a router error falls
through to the classifier.
## The degradation rule
Two cases, and the line between them is sharp.
**Fall back silently** when the workstation model would only do the job *better*:
routing, phrasing, a nudge. Falling back costs nothing that exists today, because
the resident Qwen3-1.7B is today's production quality. The owner should not be told
that his reply was phrased by the smaller model.
**Name the gap** when the resident model cannot do the job *at all*. A world
question that a 1.7B answers by inventing is the case. A wrong answer is worse
than "не могу сейчас". This is the rule CLAUDE.md already states for a sibling
service being down.
Nothing in between. A turn never breaks on the workstation being asleep.
Both halves are wired, 03-08-2026. `LLMPhraser.PhraseWorld`
(`internal/phraser/world.go`) is the naming half and has three outcomes, not two:
| State | What he hears |
|---|---|
| no `workstation` block | the resident model answers, exactly as before the seam existed |
| configured, card free | the workstation answers |
| configured, asleep or busy | the gap, `worldGap` in `cmd/mavend/worldmodel.go` |
The first row is the one worth stating. Naming a gap requires a gap. On a box with
no second model the 1.7B is the whole product. Refusing every world question there
would remove a capability the owner has today.
A source holding a passage is on the naming half too: a live search, a ZIM
article, a page he named. None of them says "не могу сейчас". They read the
passage back, which is what `phraseSource` returning `""` selects. A real quote
beats a gap, and neither path invents.
## Admission control, not a scheduler
There is no GPU arbiter. That is a service with its own failure modes, and nothing
here needs work *distributed*. It needs admission control. The workstation
advertises free VRAM over a health endpoint, and Maven treats it as one more query
source that claims a turn or passes. llama-server also refuses to load when VRAM is
short, so the failure is detectable without cooperation from the owner's other
jobs.
The caller must be able to ask "is this peer usable right now" without a turn
hanging on a timeout. A dead remote is a normal state, not an error state.
`internal/llm.Pair` is that check on the Maven side. A prober caches the answer,
so `Available()` is an atomic read and no turn pays for a health check.
llama-server does not stay up on the workstation. It cannot: a resident 7-14B
would hold 16GB against the owner's CPT runs. So a supervisor there owns its
lifecycle, keeps it loaded while the card is free, and unloads it on idle or
when another process needs the card (owner's call, 2026-08-02, Vikunja #488).
That supervisor is still not a scheduler, and the distinction is worth holding.
It arbitrates nothing between callers. It reports whether it can take work and
manages one process to back that answer. Maven never asks it to start anything
and never learns that it did.
Contention is decided by presence under `/sys/class/kfd/kfd/proc`, not by a VRAM
threshold. A ROCm process registers there when it initialises HIP, before it
allocates anything. So the supervisor sees a contender during that job's startup,
and yields before the job loses the memory it asked for. A
threshold reads the card too late. By the time free VRAM has dropped, the other
job has already lost the allocation race. Free VRAM is still read, but only as a
precondition for loading, never as the eviction signal. One blind spot is known.
A job can take the card without registering on the KFD, as a Vulkan or a
video-decode job would. `describe()` logs every contender's comm, and that log is
how we find out whether the blind spot is real.
`mavgpud` runs from a systemd unit on the workstation with
`deploy/mavgpud.json` as its config, and `llama_args` is passed to llama-server
untouched. The model, the context size, the layer count and the MTP flags are the
owner's business and not this daemon's schema.
## What stays on homesrv, permanently
The **embedder** (multilingual-e5-small, ONNX, CPU). It backs the classifier, which
must answer while the GPU is saturated. It is also cheap enough on CPU that moving
it buys nothing. Four callers:
| Caller | What for |
|---|---|
| `internal/router/classifier.go` | the routing floor |
| `cmd/mavend/actions_query.go` (`queryEmbed`) | memory recall |
| `cmd/mavend/feeds.go` | ingest embedding for every RSS item |
| `internal/crawl/watch.go` | ingest embedding for every crawled page |
`internal/speaker` becomes a fifth once it lands.
## Inventory: what runs a model on homesrv today
The **resident model** is one llama-server with seven callers, and 03-08-2026 is
the date each of them stopped or did not stop being resident-only:
| Caller | What for | Offloaded |
|---|---|---|
| `cmd/mavend/voicewire.go` | routing | silently, through `hot` |
| `cmd/mavend/replier_llm.go` | replies | silently, through `hot` |
| `cmd/mavend/tick.go` | digestion worker: `PhraseNudge`, `PhraseReminder` | silently, inside the phraser |
| `cmd/mavend/actions_query.go` | world questions, and any fetched passage | names the gap |
| `cmd/mavend/capture.go` | capture summarisation (unreachable, see #480) | no, holds its own client |
| `cmd/mavend/mail.go` | mail extraction (off, no IMAP) | no, holds its own client |
| `memoryeval.go`, `modelswap.go` | admin and evals | no, and deliberately |
The last three rows are resident-only on purpose. `memoryeval.go` and
`modelswap.go` measure and swap the resident model, so sending their work
elsewhere would measure the wrong thing. `capture.go` and `mail.go` are
background jobs that hold a gated background client (`llmBackgroundClientFor`),
and that priority has no equivalent on the remote yet. Both are also unreachable
on this deploy, so wiring them would ship an untestable path.
The `tick.go` row needs one caveat. `phraser.llm_nudges` is `false` in deploy, so
nudges come from templates and the seam under them changes nothing until that
flips. It is wired anyway: `PhraseReminder` is on the same transport and is on.
Then the embedder above, **whisper.cpp** in `mavsttd`, and **piper** in `mavttsd`.
`mavwaked` uses no model at all: an energy-threshold VAD over 30ms frames.
## Order
1. **Transport** (#484). Nothing else is possible until a seam can cross a host.
`internal/netaddr` landed in PR #92. A seam address now carries its own scheme,
and a scheme-less one is still unix. A tcp seam requires a shared token, because
the filesystem permission that authenticated the unix socket is gone.
2. **The resident model** (#485, #490). Wired. A `workstation` block builds an
`llm.Pair` in `modelSeam` (`cmd/mavend/voicewire.go`), routing and replies
complete through it, and the phraser holds the same pair (`UseRemote`). Both
halves of the rule are live: see the table above for which caller gets which.
Measured, `docs/evals/2026-08-02-workstation-gemma4-12b.md`: gemma-4-12b
through the cascade scores 84.4% full accuracy at p50 329ms. The resident
model scores 72.7% at p50 0.80-1.04s. On the talk fixture it is 25/27
against 20/27. Biggest quality delta. A 16GB card runs a 7-14B,
which fixes what the 1.7B gets wrong: world knowledge, and the persona the CPT
targets. The degradation path is already written and measured, since the
classifier scores 68.8% full accuracy at p50 16.6µs on its own.
3. **Speech-to-text and text-to-speech** (#486). They gain a real margin, but on
quality alone, and both already work.
4. **The wake word** (#487). Independent of all of the above.
## Assumptions
- The LAN is trusted enough that wireguard is supported but not required (owner's
call). What crosses the wire is still his utterances. That is why the tcp seam
carries its own token instead of assuming a network boundary.
- The workstation is not expected to be up. Every child task must still serve a
turn while it is down.
+2
View File
@@ -1,5 +1,7 @@
# Start Commands
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
All commands assume `ROOT=/home/kami/apps/Maven` and the local Go toolchain at `$ROOT/deps/go/go/bin/go`.
## Prerequisites
@@ -6,7 +6,7 @@
> use the embedded LFM model paths or old single-object examples as current ops
> guidance; see `2026-07-18-qwen3-resident-training-eval.md`.
> Scope from `REARCH.md`. Make Maven trustworthy: the LFM becomes the router
> Scope from `docs/rearchitecture.md`. Make Maven trustworthy: the LFM becomes the router
> (fixes "messes up queries" / "doesn't take notes"), the engine actually runs
> (fixes stub replies), dates stop being read as "number dot number dot number",
> and telegram becomes a reach channel. NOT in scope: on-demand 4B reasoner,
@@ -693,7 +693,7 @@ ssh kami@192.168.1.104 'curl -s localhost:9201/api/chat -d "{\"text\":\"запо
4. `docker compose up -d mavend && docker logs -f maven-mavend-1` — confirm the
phraser spawns and no `phraser: NewStub` path. Run the two verify curls.
5. Update `AGENTS.md`: LFM model download + note that routing is now LFM-first
with classifier fallback (`REARCH.md` is the design of record).
with classifier fallback (`docs/rearchitecture.md` is the design of record).
---
+2
View File
@@ -1,5 +1,7 @@
# Maven Voice Protocol
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
> Auto-generated from `internal/voice/wire.go`, `internal/voice/errors.go`,
> `internal/voice/frame.go`, `internal/voice/client.go`. If this file and
> those files disagree, the code wins.
+566
View File
@@ -0,0 +1,566 @@
# QA plan: checking Maven properly
*Last verified: 2026-08-02 @ 20aa2d5. Living doc: correct it in place, do not append.*
Written 2026-08-01, after the 35-PR stack landed and the box came back up.
Refreshed 2026-08-02 against the live list, after PRs #85-#90.
42 of the 50 open Vikunja tasks are `QA:` tasks. They are verification work, not
build work. Most sat unverifiable while Maven was down for 11 days. That
blocker is gone.
The plan as written on 2026-08-01 named 40 task numbers. Ten open `QA:` tasks were
missing and two of the named ones had closed. Every open task now appears below,
the eight non-QA ones in the last two sections.
This plan orders them by what unblocks what. Do sessions 1 and 2 first. Almost everything
downstream assumes the voice loop works, and nobody has confirmed that since
the redeploy.
---
## What the 02-08-2026 run found
Sessions 1, 2 and 3 all ran. Read these five before picking anything up.
- **470: a question writes invented knowledge into memory.** Recall then serves
it back. `что дальше?` lands on `IntentFact` and stores the model's answer as a
`self` fact at confidence 1.00. Two junk rows then claimed seven unrelated
world questions through recall, outranking the search leg. A question about the
capital of Australia was answered `какая последняя версия языка Go?`. Two bad
writes silently disabled world answering, with nothing logged.
- **466: a pending clarify is global.** One unanswerable clarify swallowed the
next three utterances from three separate sessions. With ntfy, telegram and
voice all live, a clarify raised on web chat eats the next telegram message.
- **467: spoken task capture is dead.** The router calls the capture marker an
`act`, and capture is reachable only from the `note` intent.
- **The classifier baseline in this repo was wrong**, and it flattered the
router. See session 2 and **464**.
- **477: the model swap and the self-update cannot be triggered on this box.**
Both are built and both are correct in test. The swap needs a passkey and
WebAuthn is unconfigured. `mavupdate` needs to reach a socket that only an
in-container uid can open.
- **479: an unconfigured capability lets the question escape to web search.**
Netscan off, asked `какие устройства в сети?`. She answered from the live web
with a general article about network hardware. A question about his LAN went to
an upstream engine. The crawler fails the same way.
Twenty-one defects were filed on 02-08-2026: 462 through 482. Six tasks this plan
had written off as blocked turned out to be ready to check. All six ran. Every
one of them is code-correct and stops at the deploy.
Three of the five config blockers in **472** were then cleared. The morning
routine, ambient ingest, feeds, the crawler and netscan are all live. Two remain,
and both are the owner's call: a token for each ecosystem sibling, and seed data
in Nexus and Praxis.
---
## Before you start
Two things bite anyone running these checks on homesrv.
**curl needs `--noproxy '*'`.** The shell exports `http_proxy=http://127.0.0.1:18080`.
Without the flag, every local check returns 503 from the proxy and looks like a
dead service. This cost me a false regression report today.
**The database is not readable with sqlite3.** Four older QA steps say
`docker compose exec mavend sqlite3 /data/maven.db "select ..."`. That cannot
work: the container has no `sqlite3` binary, and the store is AES-256-GCM at
rest with a tmpfs working copy. Read state through mavweb instead, at
`/history`, `/trace`, `/routines` and `/dash`.
---
## Session 1: the voice loop (half a day)
Nothing here has been confirmed since the redeploy, and everything else assumes
it works. Do this first.
Closes or advances: **44** (conversation), **45** (text chat), **287** (voice
session quality), **321** steps 3-5 (quiet mode), **288** (STT golden audio).
**288 is not blocked.** The fixtures are committed under `cmd/mavsttd/testdata/`
and `make test-stt-golden` runs today. This plan said otherwise until 02-08-2026.
Steps 1 and 3-6 were run on 02-08-2026 and pass. Steps 2 and 7-9 still need a
person at the box, because they need a microphone or a nudge to arrive.
Steps 1 and 3-6 do not need a browser. `POST /api/chat` takes a form-encoded
`text=` field and a cookie jar, and answers with the rendered `/chat` page:
```sh
curl -s --noproxy '*' -c jar -b jar -L -X POST \
http://127.0.0.1:9201/api/chat --data-urlencode 'text=привет'
```
Parse the whole page, not the last text node. The page carries nav and footer
text. A naive tail of the Cyrillic nodes returns the wrong string, which makes
turns look misaligned when they are not.
1. Open `http://127.0.0.1:9201/chat` and hold a short conversation in Russian.
Watch for three things: she answers in feminine forms (`рада`, `поняла`), she
says `ты` and never `вы`, and no pet names appear.
**Passes** (02-08-2026, five turns): `я рада`, `поняла`, `помогла`,
`проверила`, `записала`, `грустна`, `ты` throughout, no pet names.
2. Press push-to-talk on `/dash`. Say `привет`. Confirm a spoken reply comes
back. This is the only check that covers mic to STT to core to TTS to
speaker as one path. It is also the path the eleven-day outage most likely
broke.
3. Say `тихий режим`. Expect `тихий режим включён. буду реже напоминать.` **Passes.**
4. Say `выключи тихий режим`. Expect `тихий режим выключен.` Negation must win. **Passes.**
5. Say `в комнате тихо`. Quiet mode must NOT flip. Confirm on `/history` that no
`quiet_hours` fact was written. **Passes**: no row written. She answers `пока
не умею отвечать на этот вопрос.`, so it lands on `IntentSystem` with no arm.
6. Say `включи режим тишины`, then `сделай потише`. Both must flip quiet mode
on. These are the noun form and the comparative, added 01-08-2026. **Both pass.**
7. Wait for a nudge, then say `потом` within twenty minutes. Expect `хорошо,
вернусь к этому позже.` and the nudge row on `/notifications` reading
`snoozed`. Say `потом` again with nothing pending: it must route as an
ordinary utterance, not be swallowed.
8. Wait for the water nudge, then say `выпил воды`. Expect the ordinary fact
reply and nothing extra. She must not congratulate you. Check
`/notifications`: the row reads `acted`. Then trigger another nudge and say
`готово`. Expect `отлично, отметила.` and the same outcome.
9. Note anything where she is slow, cuts off, or talks over herself. That is
287's whole content and it has no written acceptance criteria yet.
**First evidence, in text** (02-08-2026): nothing breaks, but answers wander
and stitch unrelated topics. Asked whether he should move flats, she opened
with the weather. That is 287, and it is a phrasing problem, not a loop problem.
**The wake path cannot be checked as deployed.** `mavwaked` and `mavenclient`
appear in no compose file and run as no host process. Step 2 covers only
push-to-talk, from `/dash` through mavsttd and mavttsd. Wake word and VAD
are untested by construction. Decide whether they belong in compose or on a
client machine, and say which in the deploy docs. Tracked as **463**.
**319's single-token bug is fixed** (01-08-2026). Single-word Russian utterances no longer come
back as `не совсем поняла — можешь переформулировать?`. `привет` and `поужинал`
both pass now: `thinSingleToken` spares social singles and any token carrying a
verb ending, and only thins a bare nominal like `вода`. A one-word utterance that
still gets clarified in this session is a new case for the lexicon, not the old bug.
---
## Session 2: measurement (half a day, mostly waiting)
Closes or advances: **320** items 2-4, **278** (make the eval lab routine).
Also **248** (memory evaluation), **319** (the margin gate) and **323** (the
startup timeout arm).
The resident llama-server cannot be reached by the eval harness. It binds
`--host 127.0.0.1 --port 0` inside the container, so the port is kernel-assigned
and never published. Start a second one on a fixed port instead:
```sh
llama-server -m /mnt/hdd1/llms/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf \
--host 127.0.0.1 --port 18100 -c 4096 -ngl 99 --no-webui
```
`-c 4096` matters. The recorded numbers were measured at that context size, and
a mismatch invalidates the comparison.
Then:
```sh
make eval-models MAVEN_LLM_URL=http://127.0.0.1:18100 # want ~72.7% cascade
make eval-router # classifier baseline
MAVEN_LLM_URL=http://127.0.0.1:18100 make eval-phrasing # persona checks, slow
make eval-recall
```
A large miss against 72.7% means the deploy differs from the bench harness.
**Run on 02-08-2026 @ af9d213. The deploy matches the bench.** `eval-models`
scored 56 of 77: 72.7% full, 77.9% intent-only, 2 false clarifies and 1 missed.
That is the recorded figure to the decimal, and calendar sat at 2 of 2, so the
stage 0 agenda rules hold. `eval-phrasing` scored 21 of 27 on the talk fixture
against a recorded 20, and the 15 nudge templates passed every check.
Two numbers in this repo were wrong, and both flattered the resident model.
- **The classifier is not 36.8% and not 31ms.** `make eval-router` reports
`classifier+onnx: 53/77 (68.8% full)` at p50 16.6µs. The figure repeated here
and in `CLAUDE.md` predates the stage 0 rules and the seed additions. Both now
score inside that baseline. The accuracy gap the router buys is
roughly 4 points, not 36. Re-argue the trade on the real numbers: **464**.
- **Router latency was measured under contention again.** p50 1.126s, p95 1.58s,
max 3.24s, against a recorded p50 825ms. The resident model was serving the
daemon on the same iGPU throughout. Do not record this as a regression, and do
not record it as a measurement either. Stop the stack before timing the router.
`classifier+hash` scores 19.5%, which is the no-ONNX degraded path and is not the
failure floor the deploy uses. Do not quote it as the classifier baseline.
Then three things to decide while the numbers are in front of you:
- **319 is done.** 359 gave the LLM path a real confidence signal.
`thinSingleToken` was narrowed on 01-08-2026, and agenda questions moved to
stage 0. Missed clarify sits at 1 of 6 and false clarifies at 2. Item 2 point 2
closed on 02-08-2026: the `make eval-recall` margin sweep is the distribution
that was asked for, and `0.008` sits at the knee.
| delta | answered | false recall |
|---|---|---|
| 0.005 | 18/27 | 2/5 |
| **0.008** | **18/27** | **1/5** |
| 0.010 | 16/27 | 1/5 |
It removes four of five false recalls at no cost in answers, and the next step
costs two answers for nothing. The hand-picked value survives on evidence.
- **278's real ask** is making the eval lab routine rather than building it. It
is built. Decide whether it runs on a timer, on every merge, or on demand, and
the task can close.
- **248** is the memory evaluation loop. It ships, it writes notes, and it cannot
speak. `make eval-recall` covers the retrieval half. The open question is whether
a written evaluation nobody reads is worth the tick.
**323 is down to one check.** PR #90 covered the spawn path and took phraser
coverage to 76.9%. Only the 60s startup timeout arm is untested, because testing it
needs a `StartupTimeout` field on `Config` rather than a test-only hack. While you
are on the box, time a cold 1.7B load off spinning disk. If it runs near 60s, the
default is too tight and the field earns itself twice.
Warm, it is nowhere near. A second llama-server answered `/health` 1.8s after
launch at `n_ctx 4096` on 02-08-2026. That is page cache, so it does not settle
the question. A cold read needs a cache drop, which needs root.
**`CheckFeminine` has a false positive.** On 02-08-2026 it failed
`query-notes-do-not-answer` for `ты заплатил`, calling it masculine
self-reference. Masculine second person is correct, because the owner is male.
The check matches a masculine
past-tense verb before `за` without confirming the subject is `я`. Fix it in
`internal/phraser/eval/checks.go` before trusting a phrasing score to the case.
The real talk-fixture score on that run is 22 of 27, not 21. Tracked as **462**.
Item 4 of **320** needs a permission I do not have. Kill the `llama-server`
pid under `maven-mavend-1`, post a turn, and confirm it still completes
through the classifier. Either grant it or run it yourself. It is the only
check that the failure floor catches a mid-session model death.
---
## Session 3: the interaction batch (a day, or five sittings)
These need real use rather than a command, grouped by what one sitting covers.
**Morning and delivery** (**280**, **281**, **128**, **282**, **283**, **285**):
open `/morning`, walk the seven required behaviours, then check the four
interruption outcomes and the digest gap. **282** needs the `desk_active` script
enabled on the desk PC first, which is **15** and needs you at that machine.
**283** is the event intake envelope every reach shares, so a delivery check
exercises it whether you name it or not. **285** is not verification: the bridge
framework works and the remaining ask is more adapters. Decide which reach comes
next, or park it.
Run 02-08-2026. **280 is blocked.** No morning routine is configured (**472**).
`morning.Item` also has no required-versus-optional field, so behaviour 1 cannot
hold whatever you configure (**473**). **281's digest gap is closed**, and
its presence rule passes on inspection. Three of its five items need traffic the
box has not had. **283 is blocked**: nothing feeds the intake journal. **128
found the worst defect of the whole session, see below.**
Three of 472's five blockers were cleared the same day, in `deploy/mavend.json`
and `docker-compose.yml`.
- A `morning_routines` block, one routine `утро` 08:00-11:00 with medicine,
water and pets. It is live: the dispatcher logged `dropped morning:утро (sev1,
presence=away)`, so the plan builds and the nudge is proposed. 280's
behaviours and 128 step 11 are checkable now. 473 still stands.
- `-ambient-token` on mavweb, value in a gitignored `/.env` that docker compose
reads for interpolation. `/api/ambient` answers 401 without the token and 201
with it, storing `calendar_event_20260802_Standup`. 283 step 5 and 128 step 8
are unblocked. The token is a flag, so it shows in `ps` inside that container.
The zenmoney and IMAP secrets are read from files instead. Ingest also
reads the notification's wall clock as UTC and stores a 14:30 meeting at 18:30
(**482**).
- `feeds` (two sources), `crawl.on_demand` and `netscan.enabled`. The intake
journal now fills: `/events` holds `scan:lan` and `ambient:notif` rows.
Two are not mine to clear. No sibling has a `token` in `deploy/mavend.json`, so
273 steps 6 and 8 need a credential decision. Nexus has no entities and Praxis no
attention items, so 272 step 3 needs seed data whose content is the owner's call.
For **285**, two facts bear on the choice. Synapse is already running on this box
and healthy, so a Matrix reach has a live target and needs no new service. And
mavweb is already a PWA with a service worker, which 285 itself calls the highest
value adapter left. Today's reaches are ntfy, telegram and voice.
**Query sources** (**258**, **286**): ask her something the RSS feeds answer and
something only a ZIM answers, with the search block on. Live search leads and the
ZIMs are the fallback since 02-08-2026. **286**'s remaining half is doc and
git ingestion, which is build work, not a check.
**Do not read `/trace` for this.** `/trace` is the nudge-rule trace: rule,
severity, predicate, gate, selected. No query-source field exists anywhere in the
codebase. The only evidence of which query source claimed a turn is the
`voice: search:` and `voice: kiwix:` lines in `docker compose logs mavend`
(`actions_query.go:589` and `:660`).
Run 02-08-2026, 20 turns. **Search leads and the personal boundary holds.** Every
world question that reached the boundary was claimed by search. All three
personal questions produced no search and no kiwix line at all.
The rest of this sitting went badly. **Kiwix has zero live coverage.** SearXNG
returns four results for everything, including two invented nonsense terms. So
`querySearch` always claims, and Kiwix is unreachable code as deployed. The ZIM
half of the 02-08-2026 decision is unverified. A ZIM answer cannot signal a
silent search failure, because a ZIM answer cannot happen.
**Ordering defects** in feeds and calendar, plus 258 step 1's utterance not
working: **474**. And the sitting independently found stage 2 of **470**.
**Tasks and calendar** (**129**, **130**, **127**, **126**, **246**): capture a
task by voice, confirm it lands, check prioritisation ordering is not nonsense.
**246** (mail reader) also exercises the `IngestMail` rung that moved to
`AuthWrite` this morning.
Run 02-08-2026. **129 passes.** The page and the spoken answer agree on ordering.
The undistinguished task carries no invented reason on either surface, which is
the thing 129 asks for. **130 fails outright** and **127 half fails**:
**467**, **469**. **246 cannot be run**: `mavmaild` is commented out in
`docker-compose.yml` and there is no `email` block, so nothing in steps 4-13 is
reachable. The `IngestMail` rung does sit at `AuthWrite`
(`internal/auth/policy.go:96`, asserted in `auth_test.go:421`), verified by
reading only.
**Routines and patterns** (**43**, **46**, **247**, **254**): these need history
to detect against. If the database is thin after the outage, they may have
nothing to propose, which is not a failure. Check `/routines` before
concluding anything.
Run 02-08-2026. The answer is the middle case: **the detector ran and found
nothing.** The tick loop is live, and `detectPatterns` is called unconditionally
at `cmd/mavend/tick.go:227`. It has run about 25 times since the restart. It
finds nothing because the events table is empty upstream of it. Rows land there
only from `pattern.Extract` at fact-write time, and `Extract` requires the fact
value to match a closed 7-action lexicon. All 200 facts on `/history` are
`page_heartbeat`, `netdata_alarm`, `quiet_hours`, `name`, `service_down` and
`рост`. Not one lexicon hit, so no event can exist, let alone the four one pair
needs. **46 step 5 passes**: `/routines` renders `noticed 0` with the empty state
and the hint string.
Two things block this sitting, and both are build work. The seeding recipe on
**43** goes through `sqlite3` and cannot work. And `pattern.Detect` has no
minimum-interval floor, so seeding by hand mints a permanent false routine
(**468**). Do not try to seed a pattern with four fast chat turns.
**Ecosystem** (**272**, **273**, **276**): nexus, hexis and praxis are wired and
logged clean at boot.
Run 02-08-2026, read-only half. All three answer `/health` 200 and `/ecosystem`
lists 18 Hexis capabilities with correct read-only and mutating badges. **272 and
273 are blocked on empty data**, not on code. Nexus holds no entities, Praxis
holds no attention items, and the Calls panel has never recorded a call. See
**472**, and read its warning first. 273's trace fix has never been validated
here. An empty Calls panel is exactly what the old bug looked like. The page is
`/ecosystem`, not `/siblings`.
**276 ran 02-08-2026 and the suite is sound.** 17 `TestEcosystem_` cases pass
under `-race`, not the 10 the task describes. The mutation check bites: patching
the Nexus-error branch of `handleHexisAct` to `return ""` fails
`TestEcosystem_MalformedNexusResponseFailsClosed` on the expected line.
Steps 4 and 6 could not be checked through chat, because no utterance reaches
Praxis (**475**). «что требует внимания» routes to `intent=query` and is answered
by the search leg, identically whether `ecosystem-praxis-1` is up or stopped. The
degraded string never appears because its branch is never entered. Step 5 is
blocked the same way: `перезапусти muzick indexer` clarifies on
`HasFn:false`, and the router had already rewritten the entity name to
`музик индексер` (**476**).
Both steps were checked on `/ecosystem` instead, which reads Praxis directly.
With Praxis stopped the card reads `praxis — unreachable` while Nexus and Hexis
keep rendering. On `docker start` the card returns to `nothing needs attention.`
with no mavend restart. Independent degradation and recovery both hold.
**Operations** (**249**, **250**): both ran 02-08-2026. The code is correct and
neither lever can be pulled on this box. See **477**.
**250** passes steps 1, 2, 3, 9 and 10 on the deploy. The capability announces
itself. `/models` names the model llama-server reports, not the config filename.
Asking her to switch models does nothing. Removing `swap_models` renders `swap
not configured`. Step 4's refusal half passes at HTTP 403, and the 403 comes from
mavend rather than mavweb. WebAuthn is unconfigured, so the web gate fails open
and the wire gate fails closed. Steps 5 to 8 need a passkey assertion nothing on
this box can produce. They pass in test: 13 swap cases and 7 page cases covering
drain, mid-swap refusal, rollback, failed rollback and the not-owned refusal.
**249** passes steps 1 and 2. Step 3 stops it. `mavupdate` health-checks over
`/run/maven/mavend.sock`, which is `srw------- 1 10001 999` inside a docker
volume. The host owner cannot traverse `/var/lib/docker/volumes` and cannot
connect to a socket owned by an in-container uid. `mavupdate` assumes a
host-installed daemon and the deploy is containers. Do not sudo around this.
---
## Housekeeping (done 02-08-2026, and this section was mostly wrong)
This section claimed eleven tasks were not verification work. **Three were not.
The other eight are.** Every one of the eight has shipped, tested code behind it.
The error ran one way: it wrote off work that is ready to check. Do not trust a
"nothing is built" line in this plan without grepping for the package first.
Relabelled to `Blocked:`, claim verified:
- **125** zenmoney. `internal/zenmoney/` ships and is tested against a fixture.
`deploy/zenmoney.token` does not exist and the compose mount is commented out.
One token unblocks it.
- **256** Home Assistant. `internal/smarthome/` ships, the `smarthome` block sits
in `deploy/mavend.json` at `enabled: false`, and 8123 and 1883 are closed.
- **14** cold-start unlock. The seam is real at `cmd/mavend/main.go:128` and
`internal/webauthn/prf.go` is in place. `lockedAPI` is gone, replaced by
`Server.Check` in `internal/ipc/server.go`. Gated on an authenticator that
implements the WebAuthn PRF extension, which is hardware, not code.
Left alone, because the claim here was false:
- **284** simulator. `cmd/mavend/simulator_test.go`, three scenarios under
`cmd/mavend/testdata/scenarios/`, and a `simulate` target at `Makefile:98`.
**Run 02-08-2026: all three scenarios pass**, plus the determinism and
backwards-step guards. One defect found, see below.
- **288** STT golden audio. Four WAVs and `golden_v1.json` are committed under
`cmd/mavsttd/testdata/`, the make targets exist, and `models/stt/ggml-small.bin`
is on the box. Session 1 lists 288 as blocked on fixtures, which is wrong.
**Run 02-08-2026: all four pass**, WER at or under ceiling with no drift.
| fixture | transcript | WER | ceiling |
|---|---|---|---|
| ru_reminder | `Напомни мне через час позвонить маме.` | 0.00 | 0.10 |
| ru_fact | `А отметь, что я выпил воды.` | 0.20 | 0.25 |
| ru_query | `Что у меня сегодня по календарю?` | 0.00 | 0.10 |
| en_act | `Restart the web server and check the disk space.` | 0.00 | 0.10 |
That also settles a session 1 worry indirectly: whisper.cpp works on Vulkan
after the redeploy. Only the mic and the wake path remain unproven.
**The simulator routes with an empty seed set.** Every `make simulate` run logs
`loaded 0 seed examples from models/seeds`, seven times per scenario. The test
runs from `cmd/mavend`, and the seed path is relative to the repo root. The
scenarios still pass, which means they pass without the classifier having any
seeds to match against. Whatever 284 is proving, it is not proving the routing
the deploy runs. Fix the path before trusting a green simulator.
- **257** Bluetooth. The bluez half is genuinely absent. The LAN-scan half shipped
(`internal/netscan/`), and steps 1-9 run today. Only step 10 is Bluetooth, so
relabelling the whole task would bury real pending work.
- **251** MCP, **253** hearing, **259** crawler. All three ship
(`internal/mcp/`, `internal/capture/`, `internal/crawl/`) with no external gate.
Fully checkable. `259`'s step 1 wants no `crawl` block in `deploy/mavend.json`,
and there is none, so it is already set up correctly.
- **252** vision and **255** speaker recognition. Both ship. Each is blocked only
on a model download: a vision gguf with mmproj, and a speaker embedding model.
Neither is present under `/mnt/hdd1`. Their refusal-path steps run today.
So the honest split is three blocked on a credential or hardware, two blocked on
a download, and six ready to check. That is roughly a session of real QA this
plan had written off as backlog.
**All six ran on 02-08-2026.** Every one of them is code-correct and stops at the
deploy. The pattern repeats often enough to be the headline: the packages pass,
and the box cannot reach them.
**251, MCP.** Steps 1, 2, 3, 4 and 13 pass. Package tests green under `-race`.
Off-by-default is clean, and the SSRF refusal is exact: without `allow_private`
the log reads `refusing to connect to a private address: 127.0.0.1` and `/tools`
shows the server down with zero proposals. Steps 5 to 12 are blocked. `ss -lntp`
shows the Vikunja MCP server on `127.0.0.1:9100` only, so no container reaches it
at any address (**478**). `allow_private` does work, measured both ways.
**253, hearing.** Steps 1, 2 and 17 pass. `internal/capture` covers 90.3%. Steps
7 to 16 are blocked on something nobody can work around: no shipped client calls
`CaptureStart`. There is no `cmd/mavheard`, no mavweb route, and `mavenclient`
never calls it (**480**). Two of its QA steps are also stale.
**257, netscan.** Steps 2, 3 and 9 pass at unit level. Step 1 fails. Steps 4 to 8
need the block enabled. Step 10 is Bluetooth and stays skipped.
**259, crawler.** Steps 1 and 15 pass. Step 2 fails. Steps 3 to 14 need a `crawl`
block that nobody has written.
Both were configured later the same day, and both work. `netscan.enabled: true`
answers `какие устройства в сети?` with `нашла 3 устройства, из них 2 с вебом, 2 с
ssh. список записала.` and the scan lands in the intake journal as `scan:lan`.
`crawl.on_demand: true` answers `посмотри https://lwn.net — что там пишут?` from
the real page. So **479** is one defect, not the routing defect it was filed as.
An unconfigured capability declines its own turn instead of naming the gap.
Nothing is wrong with the routing.
257 step 1 and 259 step 2 fail the same way and share a task (**479**). An
unconfigured capability does not name the gap, so the question escapes to web
search. `какие устройства в сети?` was answered with a general article about
network hardware. That is his LAN going to an upstream engine.
**252 vision and 255 speaker.** Both confirmed blocked. The disk claim was
re-verified rather than taken on trust: 16 text-only ggufs under `/mnt/hdd1`, no
mmproj and no speaker embedding model. Everything not needing the model passes,
including the two refusals that matter. `TestNewLocalRefusesNonPrivateEndpoints`
rejects `https://api.openai.com`, and forget really deletes
(`internal/store/memory.go:145` is a real `DELETE`, not a tombstone). Vision is
19/19, speaker 22/22, media 16/16.
**470 got worse, then closed.** Both poisoned facts showed `voided` on
`/history` and the defect survived. Re-measured at 15:42, after four restarts:
`почему небо синее?` still answered `какая последняя версия языка Go?` with no
`search:` line. What came back was the question he typed, not the value the fact
held. So the poison was a vector in the memory index, and `revert` did not
remove it.
Repaired in two parts. 470 stopped the writes: a question is never a fact, and a
void drops the key's vectors. 493 fixed what the index holds. A fact is indexed
as the fact and not as the utterance, and a correction drops its superseded
vector too.
A poisoned box now repairs itself on the next start. `RepairFactVectors`
re-embeds every fact vector from the fact it names, and deletes the voided and
superseded ones. It runs once, guarded by a marker, and logs what it did.
---
## Needs you specifically
Not QA. These are blocked on a decision or a credential only you have.
| # | what |
|---|---|
| 16 | Create the Kuma API key. `-kuma-key uk5_mavpoll-key` in `docker-compose.yml` is still the placeholder. |
| 15 | Deploy `desk_active` on the desk PC. Blocks **282**. |
| 122 | Finish the CPT run for Qwen3-1.7B. The persona fix depends on it. |
| 355 | Deploy the Hexis auth change. Was blocked on Maven being under construction, which it no longer is. The client half is vendored and wired. |
| 357 | Decide whether entity-existence validation is the permanent target guard or whether blessing lands in Nexus. |
| 275 | Hexis native API and MCP parity. |
| — | Decide on `-require-stepup`. Making it the default needs WebAuthn configured first, or it locks you out of your own admin surfaces. |
317 and 354 closed on 01-08-2026. The step-up gate now covers `POST /api/chat` and
`/routines`, and the nginx template is locked down with a `maven.<domain>` block for
mavweb. The `-require-stepup` default is still your call.
---
## Not this repo
Two open tasks sit on the Maven board and are not Maven work. Move them or note
where they land, so the board stops reading as 50 things Maven owes.
- **358** replace the rowid execution cursor with a real seq column. This is Hexis,
and it must land before any execution retention or pruning does.
- **362** mirror the router prompt reorder into the relabelling prompt. This is the
training workspace, enforced by `llm/check_prompt_parity.py` there, not here.
---
## Suggested order
1. Session 1. If the voice loop is broken, nothing else matters.
2. The `-require-stepup` and Kuma decisions. Five minutes, and it unblocks **16**.
3. Session 2. **Run on 02-08-2026.** The numbers came back worse for the router
than the docs claimed. The classifier is 68.8%, not 36.8%, and 16.6µs, not
31ms. The router buys about 4 points of accuracy for four orders of magnitude
of latency. Whether that still earns its place is now an open question.
4. Housekeeping. Cheap, and it makes the remaining backlog honest.
5. Session 3, split whichever way suits you. All five sittings ran on
02-08-2026. Read the per-sitting notes before repeating any of them.
The next thing to fix is not in this plan. Four defects say the same sentence:
a capability is built and no utterance reaches it. **466** (a clarify is global),
**467** (capture is act-routed), **475** (attention is act-routed), **476** (the
router rewrites entity names). Routing is where the work is.
+2
View File
@@ -1,5 +1,7 @@
# Maven — Re-architecture (Qwen3 resident model, revised 2026-07-18)
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
> Supersedes the classifier-first routing model. Agreed in a design session
> after diagnosing that homesrv deploys with a **stub phraser** (no LLM
> running) and an embedder-classifier that routes by nearest-neighbor between
+1 -1
View File
@@ -1,7 +1,7 @@
// Package auth is maven's authority layer — the 4-layer cascade and the
// "surface caps authority" invariant.
//
// Spec contract (from DESIGN.md § Auth):
// Spec contract (from docs/design.md § Auth):
//
// a cascade, not a pick-one — each layer answers a different question:
//
+62 -1
View File
@@ -224,6 +224,11 @@ type Config struct {
// See SearchConfig.
Search *SearchConfig `json:"search,omitempty"`
// Workstation — the big model on the owner's desktop, preferred over the
// resident one when its GPU is free. nil / absent / url empty ⇒ homesrv
// behaves exactly as it does today. See WorkstationConfig.
Workstation *WorkstationConfig `json:"workstation,omitempty"`
// Praxis — the ecosystem attention-state service. When configured, maven
// calls the Praxis HTTP tools API for attention listing and item lifecycle.
// Maven never touches Praxis's database directly (ecosystem invariant: no
@@ -649,7 +654,7 @@ type VoiceConfig struct {
// LLMRouter — route with the resident model instead of the embedding
// classifier. On by default since Vikunja #320.
//
// Measured on the held-out fixture (ROUTING-EVAL-31-07-2026.md): 63.2% of
// Measured on the held-out fixture (docs/evals/2026-07-31-routing.md): 63.2% of
// intents right against the classifier's 50.0%, and no route errors. It
// costs about 1s per turn instead of 30ms.
//
@@ -1105,6 +1110,44 @@ const (
DefaultKiwixSnippetRunes = 1500
)
// WorkstationConfig — the big model on the owner's desktop (workpc, a
// 7900 GRE with 16GB), fronted by mavgpud.
//
// homesrv cannot grow a GPU, so the resident Qwen3-1.7B is the floor and this
// is the preferred model above it (owner's call, 2026-08-02, docs/offload.md).
// The workstation is never assumed up: its card is often held by a CPT run and
// the machine sleeps. No block, or an empty URL, and homesrv behaves exactly as
// it does today.
//
// Only the prompt crosses the LAN, and the workstation is not "the box". The
// rules in CLAUDE.md about what may leave still apply.
type WorkstationConfig struct {
// URL — where mavgpud listens, e.g. "http://192.168.1.105:8080". Empty ⇒
// the whole block is normalised to nil and nothing probes anything.
URL string `json:"url,omitempty"`
// Health — the admission endpoint. Empty ⇒ URL + "/health", which is what
// mavgpud serves. It answers 503 while the card is held, and that is the
// signal, so it must be the supervisor's endpoint and not llama-server's.
Health string `json:"health,omitempty"`
// Probe — how often admission is re-checked. 0 ⇒ DefaultWorkstationProbe.
// Nothing on the hot path waits for it: the answer is cached and read
// atomically, so this only sets how late Maven notices the card came back.
Probe Duration `json:"probe,omitempty"`
// Timeout — the per-request budget for a completion on the workstation.
// 0 ⇒ DefaultWorkstationTimeout. A big model on a LAN host is slower than
// the resident one, and a request that overruns falls back to the floor.
Timeout Duration `json:"timeout,omitempty"`
}
// Workstation defaults, applied in Normalise.
const (
DefaultWorkstationProbe = 15 * time.Second
DefaultWorkstationTimeout = 90 * time.Second
)
// SearchConfig — the self-hosted SearXNG instance she searches with.
//
// External search is allowed and off unless configured (CLAUDE.md). Configuring
@@ -1500,6 +1543,24 @@ func (c *Config) applyDefaults() {
}
}
// No address, no preferred model. An unconfigured workstation is the
// default deploy and must be indistinguishable from today.
if c.Workstation != nil && strings.TrimSpace(c.Workstation.URL) == "" {
c.Workstation = nil
}
if c.Workstation != nil {
w := c.Workstation
if strings.TrimSpace(w.Health) == "" {
w.Health = strings.TrimRight(w.URL, "/") + "/health"
}
if w.Probe <= 0 {
w.Probe = Duration(DefaultWorkstationProbe)
}
if w.Timeout <= 0 {
w.Timeout = Duration(DefaultWorkstationTimeout)
}
}
if c.Voice != nil {
if c.Voice.RouterThreshold <= 0 {
c.Voice.RouterThreshold = DefaultRouterThreshold
+53
View File
@@ -413,3 +413,56 @@ func TestNormaliseFillsKiwixDefaults(t *testing.T) {
t.Error("rewrite: false was not honoured")
}
}
// A workstation with no address is not a workstation. The unconfigured deploy
// must be indistinguishable from today, so the block is dropped rather than
// left to fail one probe at a time.
func TestNormaliseDropsAddresslessWorkstation(t *testing.T) {
for _, tc := range []struct {
name string
in *WorkstationConfig
}{
{"no url", &WorkstationConfig{Probe: Duration(time.Second)}},
{"blank url", &WorkstationConfig{URL: " "}},
} {
t.Run(tc.name, func(t *testing.T) {
c := &Config{Workstation: tc.in}
c.applyDefaults()
if c.Workstation != nil {
t.Errorf("kept an unusable workstation block: %+v", c.Workstation)
}
})
}
}
// The health endpoint defaults to the supervisor's, not llama-server's: mavgpud
// answers 503 while the card is held, and that refusal is the whole signal.
func TestNormaliseFillsWorkstationDefaults(t *testing.T) {
c := &Config{Workstation: &WorkstationConfig{URL: "http://192.168.1.105:8080/"}}
c.applyDefaults()
if c.Workstation == nil {
t.Fatal("dropped a usable workstation block")
}
if got, want := c.Workstation.Health, "http://192.168.1.105:8080/health"; got != want {
t.Errorf("Health = %q, want %q", got, want)
}
if time.Duration(c.Workstation.Probe) != DefaultWorkstationProbe {
t.Errorf("Probe = %s, want %s", time.Duration(c.Workstation.Probe), DefaultWorkstationProbe)
}
if time.Duration(c.Workstation.Timeout) != DefaultWorkstationTimeout {
t.Errorf("Timeout = %s, want %s", time.Duration(c.Workstation.Timeout), DefaultWorkstationTimeout)
}
}
// An explicit health URL is left alone: the supervisor may sit behind something
// that does not put /health at the root.
func TestNormaliseKeepsExplicitWorkstationHealth(t *testing.T) {
c := &Config{Workstation: &WorkstationConfig{
URL: "http://192.168.1.105:8080",
Health: "http://192.168.1.105:9000/ready",
}}
c.applyDefaults()
if got, want := c.Workstation.Health, "http://192.168.1.105:9000/ready"; got != want {
t.Errorf("Health = %q, want %q", got, want)
}
}
+1 -1
View File
@@ -1,6 +1,6 @@
// Package delivery is maven's channel-routing + dispatch layer.
//
// Spec contract (from DESIGN.md § Delivery / channel routing):
// Spec contract (from docs/design.md § Delivery / channel routing):
//
// - routing = f(severity, presence). presence decides REACHABILITY; severity
// decides INSISTENCE. need both.
+2 -2
View File
@@ -9,7 +9,7 @@ import (
"github.com/kami/maven/internal/store"
)
// This file walks every cell of the DESIGN.md § "Delivery / channel routing"
// This file walks every cell of the docs/design.md § "Delivery / channel routing"
// table, once as the pure table and once through the dispatcher, so a change
// to either side has to break a named cell.
//
@@ -205,7 +205,7 @@ func TestAwayChannelsGetMinimalBody(t *testing.T) {
// An empty Summary no longer means "send the whole body" — it means a short
// generic line — so the old expectation here was wrong as well as duplicated.
// TestCareAwayDropIsRecorded — DESIGN.md's drop is a decision ("a missed water
// TestCareAwayDropIsRecorded — docs/design.md's drop is a decision ("a missed water
// nudge is noise, a missed backup failure isn't"), so it should be visible
// rather than vanish. Today drop is a bare `continue`: no nudge row, no outbox
// attempt, no log — nothing an operator can see afterwards. now it leaves a
+1 -1
View File
@@ -19,7 +19,7 @@
// 3. if no live session exists, Send returns voice.ErrNoSession
// (wrapped). The daemon logs the partial dispatch; an OPEN deferred
// question is whether the dispatcher should reroute to away-channels
// instead of returning partial — listed in PROGRESS.md.
// instead of returning partial.
//
// Import direction: voicesink imports internal/tts (synth seam) and
// internal/voice (Sessions registry). Both are siblings of delivery; the
+1 -2
View File
@@ -1,5 +1,4 @@
// Package event is the unified intake envelope (Vikunja #283,
// 20-07-2026-BACKLOG.md item 1).
// Package event is the unified intake envelope (Vikunja #283).
//
// # The problem it solves
//
+24 -13
View File
@@ -8,13 +8,15 @@ import (
"net"
"sync"
"time"
"github.com/kami/maven/internal/netaddr"
)
// Client — the module side of the boundary. Wraps a unix-socket connection
// and satisfies CoreAPI, so a module imports ipc, holds a CoreAPI, and is
// agnostic to whether it's been wired in-process (tests / daemon-embedded)
// or over this socket (full topology). The swappability is the seam auth
// will insert into without touching module code.
// Client — the module side of the boundary. Wraps a connection to core and
// satisfies CoreAPI, so a module imports ipc, holds a CoreAPI, and is
// agnostic to whether it's been wired in-process (tests / daemon-embedded),
// over a local unix socket, or over tcp to another host. The swappability is
// the seam auth will insert into without touching module code.
//
// One Client ⇒ one conn ⇒ one concurrent request at a time. A module that
// wants parallel requests opens one Client per goroutine; the store is the
@@ -22,7 +24,8 @@ import (
// per-Client lock keeps frame interleaving impossible by construction.
type Client struct {
conn net.Conn
path string // kept so a dropped conn can be re-dialed (core restart)
path string // the address as configured, kept for errors and logs
addr netaddr.Addr // parsed, so a dropped conn can be re-dialed (core restart)
mu sync.Mutex
}
@@ -81,14 +84,22 @@ var readOnlyMethods = map[Method]bool{
MethodPing: true,
}
// Dial connects to a core socket at path and returns a Client. The module
// owns its Client lifecycle; Close on shutdown.
// Dial connects to core at path and returns a Client. The module owns its
// Client lifecycle; Close on shutdown.
//
// path is a netaddr seam address: a bare path is the unix socket it has
// always been, and "tcp://host:port?token=..." reaches a core on another
// host. See internal/netaddr.
func Dial(path string) (*Client, error) {
c, err := net.Dial("unix", path)
addr, err := netaddr.Parse(path)
if err != nil {
return nil, fmt.Errorf("ipc: dial %s: %w", path, err)
return nil, err
}
return &Client{conn: c, path: path}, nil
c, err := netaddr.Dial(addr)
if err != nil {
return nil, fmt.Errorf("ipc: dial %s: %w", addr, err)
}
return &Client{conn: c, path: path, addr: addr}, nil
}
func (c *Client) Close() error {
@@ -189,9 +200,9 @@ func (c *Client) call(ctx context.Context, m Method, params, result any) error {
// re-dials clean. Caller holds c.mu.
func (c *Client) roundtrip(m Method, raw json.RawMessage, resp *Response) error {
if c.conn == nil {
conn, err := net.Dial("unix", c.path)
conn, err := netaddr.Dial(c.addr)
if err != nil {
return fmt.Errorf("%w: dial %s: %v", errWriteLost, c.path, err)
return fmt.Errorf("%w: dial %s: %v", errWriteLost, c.addr, err)
}
c.conn = conn
}
+20 -40
View File
@@ -8,11 +8,11 @@ import (
"fmt"
"log"
"net"
"os"
"sync"
"sync/atomic"
"time"
"github.com/kami/maven/internal/netaddr"
"github.com/kami/maven/internal/store"
"golang.org/x/sys/unix"
)
@@ -446,6 +446,7 @@ func mapErr(err error) error {
type Server struct {
api atomic.Value // stores CoreAPI
path string
addr netaddr.Addr
ln net.Listener
wg sync.WaitGroup
@@ -610,31 +611,29 @@ type CheckFunc func(ctx context.Context, m Method, params json.RawMessage) error
// MethodAssertStepUp dispatch calls this instead of going through CoreAPI.
type StepUpFunc func(ctx context.Context) error
// Listen creates a Server bound to path. path's parent dir must exist and be
// 0700 (we chmod it if we own it); the socket file itself is created 0600 so
// only the same unix user can connect — the current "auth floor", same radius
// as wg at the network boundary. Removing a stale socket at path first lets
// the daemon restart cleanly.
// Listen creates a Server bound to path.
//
// A bare path is a unix socket, unchanged: its parent dir is 0700 and the
// socket file itself is 0600, so only the same unix user can connect — the
// current "auth floor", same radius as wg at the network boundary. A stale
// socket is removed first so the daemon restarts cleanly.
//
// A "tcp://host:port?token=..." address binds a network listener instead, for
// a module that lives on another host. There is no filesystem there to be the
// auth floor, so netaddr checks the shared token before this package sees the
// connection and a token is mandatory. See internal/netaddr.
func Listen(path string, api CoreAPI) (*Server, error) {
_ = os.Remove(path) // stale socket from a crashed daemon; ignore missing
if err := os.MkdirAll(parentDir(path), 0o700); err != nil {
return nil, fmt.Errorf("ipc: mkdir socket dir: %w", err)
}
// umask could widen the perms on socket creation; tighten then chmod to
// be explicit. 0600 ⇒ read+write by owner only.
oldMask := unix.Umask(0o077)
ln, err := net.Listen("unix", path)
unix.Umask(oldMask)
addr, err := netaddr.Parse(path)
if err != nil {
return nil, fmt.Errorf("ipc: listen %s: %w", path, err)
return nil, err
}
if err := os.Chmod(path, 0o600); err != nil {
_ = ln.Close()
_ = os.Remove(path)
return nil, fmt.Errorf("ipc: chmod socket: %w", err)
ln, err := netaddr.Listen(addr)
if err != nil {
return nil, err
}
s := &Server{
path: path,
addr: addr,
ln: ln,
done: make(chan struct{}),
}
@@ -1268,7 +1267,7 @@ func (s *Server) Close() error {
// missing the seal costs every write since the last clean shutdown.
log.Printf("ipc: %d connection(s) still busy after %s, closing anyway", s.liveConns(), closeGrace)
}
_ = os.Remove(s.path)
netaddr.Cleanup(s.addr)
return err
}
@@ -1338,25 +1337,6 @@ func (s *Server) Path() string { return s.path }
// while the server is serving (dispatch loads api once per request via atomic).
func (s *Server) SetAPI(api CoreAPI) { s.api.Store(api) }
func parentDir(p string) string {
if i := lastIndexByte(p, '/'); i >= 0 {
if i == 0 {
return "/"
}
return p[:i]
}
return "."
}
func lastIndexByte(s string, b byte) int {
for i := len(s) - 1; i >= 0; i-- {
if s[i] == b {
return i
}
}
return -1
}
// peerCaller — read SO_PEERCRED off a unix conn to identify the connecting
// process. Returns ok=false on a non-unix conn or a platform without
// SO_PEERCRED; the caller then proceeds without a Caller (the socket perms
+1 -1
View File
@@ -31,7 +31,7 @@ type Completer interface {
// or reply in Russian.
//
// Why the JSON wrapper: this model always thinks out loud and this llama-server
// build ignores the thinking switch (see ROUTING-EVAL-31-07-2026.md). A bare
// build ignores the thinking switch (see docs/evals/2026-07-31-routing.md). A bare
// word-list grammar just captured the reasoning — every case came back as
// "Let me analyze this request carefully". Demanding JSON, like routeGrammar and
// responseGrammar already do, gives the reasoning nowhere to go.
+6 -1
View File
@@ -129,6 +129,11 @@ type Req struct {
RepeatPenalty float64
// Stop — sequences that end generation early (e.g. newline for a one-liner).
Stop []string
// Temperature — 0 (the zero value) is greedy decoding, and greedy is what
// every caller here wanted before this field existed. It is set only by the
// phraser, whose own transport has always sampled at 0.7: routing a phrasing
// call through this client must not quietly change how it decodes.
Temperature float64
}
type msg struct {
@@ -176,7 +181,7 @@ func (c *Client) Complete(ctx context.Context, r Req) (string, error) {
Messages: []msg{{Role: "system", Content: r.System}, {Role: "user", Content: r.User}},
MaxTokens: r.MaxTokens,
Grammar: r.Grammar,
Temp: 0,
Temp: r.Temperature,
RepeatPenalty: r.RepeatPenalty,
Stop: r.Stop,
})
+193
View File
@@ -0,0 +1,193 @@
package llm
import (
"context"
"errors"
"log"
"net/http"
"sync/atomic"
"time"
)
// Pair — a preferred model on another host, with the resident one as the floor.
//
// homesrv cannot grow a GPU and the workstation has 16GB of VRAM, so the big
// model runs there and the resident Qwen3-1.7B stays here. See docs/offload.md.
// The workstation is never assumed up: its GPU is often busy with CPT runs and
// the manga-recap pipeline, and the machine sleeps. So the remote is preferred,
// never required, and Pair is what makes "preferred" mean something precise.
//
// This is admission control, not a scheduler. There is no arbiter deciding who
// gets the card. A prober asks the remote whether it will take work, caches the
// answer, and every request reads that cached answer in nanoseconds. Routing
// sits on the hot path at p50 825ms and must never wait on a machine that may
// be asleep, so no request ever pays for a health check itself.
//
// Pair satisfies nothing by itself. Callers pick a method by which half of the
// degradation rule they live under:
//
// - Complete falls back silently. For routing, replies, and nudge phrasing,
// where the big model is only better and the 1.7B is today's shipping
// quality. He is not told which model phrased his reply.
// - CompleteRemote returns ErrRemoteUnavailable instead of falling back. For
// a world question, or a long Kiwix or search passage, where a 1.7B
// confabulates rather than summarises. A named gap beats an invented
// answer.
type Pair struct {
remote *Client
floor *Client
// up — the cached admission answer, written only by the prober goroutine
// and read by every request. Atomic so the read costs nanoseconds and no
// request ever contends with the prober.
up atomic.Bool
health string
interval time.Duration
http *http.Client
stop chan struct{}
}
// ErrRemoteUnavailable — the workstation model was required and is not
// answering. Callers on the naming half of the degradation rule turn this into
// a gap in the reply ("не могу сейчас"), never into a guess from the floor.
var ErrRemoteUnavailable = errors.New("llm: workstation model unavailable")
// ErrNoFloor — a Pair was built with no resident model to fall back to. A
// configuration mistake: the floor is the whole point.
var ErrNoFloor = errors.New("llm: no floor client")
// NewPair builds the two-model arrangement. remote may be nil, which is the
// unconfigured deploy and must behave exactly as the box behaves today: every
// call goes to the floor and nothing probes anything.
//
// health is the URL the prober asks. llama-server's /health answers "is a model
// loaded and ready", which is the useful signal here, because llama-server
// refuses to load at all when VRAM is short. That makes a busy card detectable
// without any cooperation from the owner's other jobs.
func NewPair(remote, floor *Client, health string, interval time.Duration) *Pair {
p := &Pair{
remote: remote,
floor: floor,
health: health,
interval: interval,
http: &http.Client{Timeout: probeTimeout},
stop: make(chan struct{}),
}
return p
}
// probeTimeout — a remote that cannot answer /health this fast is not going to
// serve a turn either. Short on purpose: the prober runs on its own goroutine,
// but a slow probe still delays the moment Maven notices the card came back.
const probeTimeout = 2 * time.Second
// Start begins probing. It returns immediately, and the first probe runs before
// the first tick so a remote that is already up is used on the first turn
// rather than after one interval of falling back. Safe to call with a nil
// remote; it does nothing.
func (p *Pair) Start(ctx context.Context) {
if p.remote == nil || p.health == "" {
return
}
go func() {
p.probe(ctx)
t := time.NewTicker(p.interval)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-p.stop:
return
case <-t.C:
p.probe(ctx)
}
}
}()
}
// Stop ends the prober. Idempotent.
func (p *Pair) Stop() {
select {
case <-p.stop:
default:
close(p.stop)
}
}
// Available reports whether the workstation will take work right now. It reads
// a cached flag, so it is safe to call per turn on the hot path. A false answer
// is never stale in the direction that matters: the worst case is that Maven
// falls back for up to one probe interval after the card frees up.
func (p *Pair) Available() bool {
return p.remote != nil && p.up.Load()
}
func (p *Pair) probe(ctx context.Context) {
ctx, cancel := context.WithTimeout(ctx, probeTimeout)
defer cancel()
req, err := http.NewRequestWithContext(ctx, http.MethodGet, p.health, nil)
if err != nil {
p.set(false)
return
}
resp, err := p.http.Do(req)
if err != nil {
p.set(false)
return
}
defer resp.Body.Close()
p.set(resp.StatusCode == http.StatusOK)
}
// set records the admission answer and logs only the transitions. A machine
// that sleeps every night would otherwise write one line per interval forever.
func (p *Pair) set(up bool) {
if p.up.Swap(up) == up {
return
}
if up {
log.Printf("llm: workstation model available at %s", p.health)
} else {
log.Printf("llm: workstation model unavailable, falling back to the resident model")
}
}
// Complete runs r on the workstation when it will take work, and on the
// resident model otherwise. A remote that fails mid-request falls back too: the
// admission answer is a cache and can be one interval out of date, so an error
// here is expected rather than exceptional.
//
// This is the silent half of the degradation rule. It must be indistinguishable
// from today's behaviour when the workstation is down.
func (p *Pair) Complete(ctx context.Context, r Req) (string, error) {
if p.floor == nil {
return "", ErrNoFloor
}
if p.Available() {
out, err := p.remote.Complete(ctx, r)
if err == nil {
return out, nil
}
// The cached answer was wrong. Correct it now rather than sending the
// next request into the same hole, then fall back.
p.set(false)
}
return p.floor.Complete(ctx, r)
}
// CompleteRemote runs r on the workstation or refuses. It never falls back,
// because for a world question the resident 1.7B does not answer worse, it
// invents. Callers turn ErrRemoteUnavailable into a named gap.
func (p *Pair) CompleteRemote(ctx context.Context, r Req) (string, error) {
if !p.Available() {
return "", ErrRemoteUnavailable
}
out, err := p.remote.Complete(ctx, r)
if err != nil {
p.set(false)
return "", errors.Join(ErrRemoteUnavailable, err)
}
return out, nil
}
+210
View File
@@ -0,0 +1,210 @@
package llm
import (
"context"
"errors"
"net/http"
"net/http/httptest"
"sync/atomic"
"testing"
"time"
)
// completionServer stands in for a llama-server. It counts what reached it, so
// a test can say which of the two models answered.
func completionServer(t *testing.T, reply string, hits *atomic.Int64) *httptest.Server {
t.Helper()
s := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
hits.Add(1)
w.Header().Set("Content-Type", "application/json")
_, _ = w.Write([]byte(`{"choices":[{"message":{"content":"` + reply + `"}}]}`))
}))
t.Cleanup(s.Close)
return s
}
func healthServer(t *testing.T, ok *atomic.Bool) *httptest.Server {
t.Helper()
s := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if !ok.Load() {
w.WriteHeader(http.StatusServiceUnavailable)
return
}
w.WriteHeader(http.StatusOK)
}))
t.Cleanup(s.Close)
return s
}
// waitFor polls until cond holds or the deadline passes. The prober runs on its
// own goroutine, so a test has to wait for it rather than assume it has run.
func waitFor(t *testing.T, cond func() bool) bool {
t.Helper()
deadline := time.Now().Add(2 * time.Second)
for time.Now().Before(deadline) {
if cond() {
return true
}
time.Sleep(5 * time.Millisecond)
}
return false
}
// The unconfigured deploy. No remote, no probing, every call to the floor —
// exactly what the box does today.
func TestNoRemoteGoesToTheFloor(t *testing.T) {
var floorHits atomic.Int64
floor := completionServer(t, "floor", &floorHits)
p := NewPair(nil, New(floor.URL, time.Second), "", time.Second)
p.Start(context.Background())
defer p.Stop()
if p.Available() {
t.Fatal("a Pair with no remote reports available")
}
out, err := p.Complete(context.Background(), Req{User: "привет"})
if err != nil {
t.Fatalf("complete: %v", err)
}
if out != "floor" || floorHits.Load() != 1 {
t.Fatalf("out = %q, floor hits = %d", out, floorHits.Load())
}
}
// The workstation is up, so it answers and the resident model is not touched.
func TestAvailableRemoteAnswers(t *testing.T) {
var remoteHits, floorHits atomic.Int64
remote := completionServer(t, "remote", &remoteHits)
floor := completionServer(t, "floor", &floorHits)
up := &atomic.Bool{}
up.Store(true)
health := healthServer(t, up)
p := NewPair(New(remote.URL, time.Second), New(floor.URL, time.Second), health.URL, 20*time.Millisecond)
p.Start(context.Background())
defer p.Stop()
if !waitFor(t, p.Available) {
t.Fatal("prober never saw the remote come up")
}
out, err := p.Complete(context.Background(), Req{User: "привет"})
if err != nil {
t.Fatalf("complete: %v", err)
}
if out != "remote" || floorHits.Load() != 0 {
t.Fatalf("out = %q, floor hits = %d", out, floorHits.Load())
}
}
// The card is busy, so /health refuses and Complete degrades silently. This is
// the constraint from 483: the workstation being down is indistinguishable from
// today's behaviour.
func TestBusyCardFallsBackSilently(t *testing.T) {
var remoteHits, floorHits atomic.Int64
remote := completionServer(t, "remote", &remoteHits)
floor := completionServer(t, "floor", &floorHits)
health := healthServer(t, &atomic.Bool{}) // never ok
p := NewPair(New(remote.URL, time.Second), New(floor.URL, time.Second), health.URL, 20*time.Millisecond)
p.Start(context.Background())
defer p.Stop()
time.Sleep(60 * time.Millisecond)
out, err := p.Complete(context.Background(), Req{User: "привет"})
if err != nil {
t.Fatalf("complete: %v", err)
}
if out != "floor" || remoteHits.Load() != 0 {
t.Fatalf("out = %q, remote hits = %d", out, remoteHits.Load())
}
}
// The cached admission answer can be one interval out of date, so a remote that
// dies between probes must still not break the turn.
func TestRemoteErrorMidRequestFallsBack(t *testing.T) {
var floorHits atomic.Int64
dead := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(http.StatusInternalServerError)
}))
defer dead.Close()
floor := completionServer(t, "floor", &floorHits)
up := &atomic.Bool{}
up.Store(true)
health := healthServer(t, up)
p := NewPair(New(dead.URL, time.Second), New(floor.URL, time.Second), health.URL, time.Hour)
p.Start(context.Background())
defer p.Stop()
if !waitFor(t, p.Available) {
t.Fatal("prober never saw the remote come up")
}
out, err := p.Complete(context.Background(), Req{User: "привет"})
if err != nil {
t.Fatalf("complete: %v", err)
}
if out != "floor" || floorHits.Load() != 1 {
t.Fatalf("out = %q, floor hits = %d", out, floorHits.Load())
}
// The failed request must have corrected the cached answer, so the next
// one does not walk into the same hole.
if p.Available() {
t.Fatal("a failed remote request left the admission answer up")
}
}
// The naming half of the degradation rule. A world question must not be handed
// to the resident model, because it answers by inventing.
func TestCompleteRemoteNamesTheGap(t *testing.T) {
var floorHits atomic.Int64
floor := completionServer(t, "floor", &floorHits)
health := healthServer(t, &atomic.Bool{}) // never ok
p := NewPair(New("http://127.0.0.1:1", time.Second), New(floor.URL, time.Second), health.URL, 20*time.Millisecond)
p.Start(context.Background())
defer p.Stop()
time.Sleep(60 * time.Millisecond)
if _, err := p.CompleteRemote(context.Background(), Req{User: "почему небо голубое"}); !errors.Is(err, ErrRemoteUnavailable) {
t.Fatalf("err = %v, want ErrRemoteUnavailable", err)
}
if floorHits.Load() != 0 {
t.Fatalf("CompleteRemote fell back to the floor %d times", floorHits.Load())
}
}
// Routing sits on the hot path and must never pay for a health check. Available
// reads a cached flag, so it costs no network at all.
func TestAvailableDoesNotProbe(t *testing.T) {
var probes atomic.Int64
health := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
probes.Add(1)
w.WriteHeader(http.StatusOK)
}))
defer health.Close()
p := NewPair(New("http://127.0.0.1:1", time.Second), New("http://127.0.0.1:1", time.Second), health.URL, time.Hour)
p.Start(context.Background())
defer p.Stop()
if !waitFor(t, p.Available) {
t.Fatal("prober never ran")
}
before := probes.Load()
for range 1000 {
p.Available()
}
if got := probes.Load(); got != before {
t.Fatalf("1000 Available calls made %d probes", got-before)
}
}
// A Pair with no floor is a configuration mistake, and it must say so rather
// than silently having nowhere to degrade to.
func TestNoFloorIsAnError(t *testing.T) {
p := NewPair(nil, nil, "", time.Second)
if _, err := p.Complete(context.Background(), Req{User: "привет"}); !errors.Is(err, ErrNoFloor) {
t.Fatalf("err = %v, want ErrNoFloor", err)
}
}
+5 -5
View File
@@ -10,12 +10,12 @@ import (
// Tests for the universal restraint gate.
//
// DESIGN.md § Trigger model: "the gate is universal, applied by the loop, never
// docs/design.md § Trigger model: "the gate is universal, applied by the loop, never
// per-rule — quiet-hours, presence, cooldown, snooze, calendar-busy all live in
// one fires()." These tests pin the CONSERVATIVE side of that: the cases where
// Maven must stay quiet. They exist so nobody loosens the gate by accident.
//
// Where the code does not yet do what DESIGN.md promises, the test is written to
// Where the code does not yet do what docs/design.md promises, the test is written to
// show the gap and then skipped, with the file and line to fix. Behaviour is not
// changed to make a test pass.
@@ -54,7 +54,7 @@ func TestGateQuietHoursSuppressesCareOnly(t *testing.T) {
// ---------------------------- presence ---------------------------------------
// DESIGN.md § Delivery: "sev <= 2 drops on away, sev >= 3 holds: a missed water
// docs/design.md § Delivery: "sev <= 2 drops on away, sev >= 3 holds: a missed water
// nudge is noise, a missed backup failure isn't."
func TestGateAwayDropsCareHoldsOps(t *testing.T) {
cases := []struct {
@@ -250,7 +250,7 @@ func TestTickOrderOfRulesDoesNotMatter(t *testing.T) {
// ---------------------------- reminders bypass the gate ----------------------
// DESIGN.md § User reminders: "bypasses the restraint gate — 'wake me 7' fires
// docs/design.md § User reminders: "bypasses the restraint gate — 'wake me 7' fires
// in quiet hours; that's the point." Every suppressor set at once, and the
// reminder still comes through.
func TestRemindersBypassEverySuppressor(t *testing.T) {
@@ -269,7 +269,7 @@ func TestRemindersBypassEverySuppressor(t *testing.T) {
}
}
// GAP — DESIGN.md § User reminders ends "Snooze still applies." RemindDecisions
// GAP — docs/design.md § User reminders ends "Snooze still applies." RemindDecisions
// passes every due reminder straight through with no snooze check, so a snoozed
// reminder fires anyway. The test below is what the contract asks for.
func TestRemindersStillHonourSnooze(t *testing.T) {
+4 -4
View File
@@ -18,7 +18,7 @@ import (
// - does it stay quiet when it should?
// - is it silent when the key it needs has no data at all?
//
// The last one is load-bearing. DESIGN.md: "since(key)==null → don't fire.
// The last one is load-bearing. docs/design.md: "since(key)==null → don't fire.
// Silence on no-data is 'shuts up when uncertain'."
// stateWith builds a snapshot at refTime() holding just the given facts.
@@ -179,7 +179,7 @@ func TestCareRulePredicates(t *testing.T) {
// ---------------------------- ops rules --------------------------------------
// The two ops rules match on a value AND on which poller wrote it. DESIGN.md:
// The two ops rules match on a value AND on which poller wrote it. docs/design.md:
// "a compromised poller must not be able to forge a trigger." Half of this
// table is forgery attempts; all of them must be refused.
func TestOpsRulePredicates(t *testing.T) {
@@ -331,7 +331,7 @@ func TestNoDefaultRuleFiresOnEmptyState(t *testing.T) {
}
}
// Severities are the delivery contract (DESIGN.md § Delivery / channel
// Severities are the delivery contract (docs/design.md § Delivery / channel
// routing): care is sev1-2 and drops when away, ops is sev3-4 and holds. Pin
// them so a change to a rule's insistence has to be deliberate.
func TestDefaultRuleSeverities(t *testing.T) {
@@ -356,7 +356,7 @@ func TestDefaultRuleSeverities(t *testing.T) {
}
}
// Cooldown bounds keep the feedback tuner honest — DESIGN.md wants
// Cooldown bounds keep the feedback tuner honest — docs/design.md wants
// `cooldown in [min,max]` "so a weird week can't mutate Maven silent or
// stalker". A base outside its own envelope would make that meaningless.
func TestDefaultRuleCooldownsAreBounded(t *testing.T) {
+126
View File
@@ -0,0 +1,126 @@
package memory
import (
"strings"
"unicode"
)
// stopwords — words that carry no topic. A question and a note that share only
// these share nothing: "почему небо синее" and "сеть какая-то медленная" both
// contain "какая"-shaped filler and are about different worlds.
var stopwords = map[string]bool{
// interrogatives and demonstratives
"что": true, "чего": true, "какой": true, "какая": true, "какое": true,
"какие": true, "каких": true, "кто": true, "кого": true, "кому": true,
"почему": true, "зачем": true, "где": true, "куда": true, "откуда": true,
"когда": true, "сколько": true, "как": true, "то": true, "это": true,
"этот": true, "тот": true, "там": true, "тут": true, "такой": true,
// pronouns — every sentence he says is about him, so "я" is not a topic
"я": true, "меня": true, "мне": true, "мой": true, "моя": true, "мои": true,
"ты": true, "тебя": true, "тебе": true, "твой": true, "он": true, "она": true,
"они": true, "мы": true, "себя": true, "свой": true,
// prepositions, conjunctions, particles, copulas
"в": true, "во": true, "на": true, "с": true, "со": true, "у": true,
"о": true, "об": true, "про": true, "за": true, "из": true, "по": true,
"до": true, "от": true, "для": true, "над": true, "под": true, "при": true,
"и": true, "а": true, "но": true, "или": true, "же": true, "ли": true,
"не": true, "ни": true, "бы": true, "был": true, "была": true, "было": true,
"быть": true, "есть": true, "был-ли": true, "уже": true, "ещё": true,
"еще": true, "так": true, "вот": true, "там-же": true,
// English filler, for the mixed utterances he does say
"the": true, "a": true, "an": true, "is": true, "are": true, "was": true,
"were": true, "be": true, "of": true, "in": true, "on": true, "at": true,
"to": true, "for": true, "about": true, "and": true, "or": true, "not": true,
"what": true, "who": true, "why": true, "when": true, "where": true,
"which": true, "how": true, "i": true, "my": true, "me": true, "it": true,
"this": true, "that": true,
}
// firstPerson — the words that make an utterance a question about his own
// life. Not possession only: "как я восстановил конфиги" owns nothing and is
// still about him.
var firstPerson = map[string]bool{
"я": true, "меня": true, "мне": true, "мной": true, "мой": true,
"моя": true, "моё": true, "мое": true, "мои": true, "моего": true,
"моей": true, "моих": true, "моим": true, "себя": true, "свой": true,
"своя": true, "свои": true, "своего": true, "мною": true,
"i": true, "me": true, "my": true, "mine": true, "myself": true,
}
// RecallAllowed is the second half of the recall gate (#470). A hit that
// cleared the score and margin gate may still be about something else
// entirely: the held-out fixture puts the right note at 0.791-0.890 and the
// must-be-silent cases at 0.795-0.835, so no threshold sits between them, and
// a note about his slow network answered "почему небо синее?".
//
// The veto applies only to a question that mentions nothing of his. That
// restriction is what keeps the fix from costing more than it saves: recall
// exists to find the note whose words he no longer remembers, and demanding a
// shared word of every recall silenced four true recalls on the fixture to
// kill one false one. A question about his own life keeps the embedder alone
// as its judge. A question about the world has to name something the memory
// actually mentions.
func RecallAllowed(query, text string) bool {
if mentionsHim(query) {
return true
}
return SharesContentWord(query, text)
}
func mentionsHim(query string) bool {
for _, w := range strings.FieldsFunc(strings.ToLower(query), func(r rune) bool {
return !unicode.IsLetter(r) && !unicode.IsDigit(r)
}) {
if firstPerson[w] {
return true
}
}
return false
}
// SharesContentWord reports whether query and text have at least one topic
// word in common, after dropping the words that carry no topic. Stems are
// compared, so the note and the question do not have to inflect alike.
func SharesContentWord(query, text string) bool {
q := contentWords(query)
if len(q) == 0 {
// Nothing to compare — a question made entirely of filler. The score
// gate is then the only judge it can have.
return true
}
t := contentWords(text)
for _, a := range q {
for _, b := range t {
if a == b || sameStem(a, b) {
return true
}
}
}
return false
}
func contentWords(s string) []string {
var out []string
for _, w := range strings.FieldsFunc(strings.ToLower(s), func(r rune) bool {
return !unicode.IsLetter(r) && !unicode.IsDigit(r)
}) {
if !stopwords[w] {
out = append(out, w)
}
}
return out
}
// sameStem is inflection and derivation tolerance: Russian marks case and
// tense on the ending, and the note and the question rarely use the same form.
// "воду" and "вода" are the same water, "кормить" and "корм" the same feeding.
// All but the last rune of the shorter word must match, and never fewer than
// three, which is what keeps "сеть" clear of "сеанс".
func sameStem(a, b string) bool {
ar, br := []rune(a), []rune(b)
n := min(len(ar), len(br)) - 1
if n < 3 {
return false
}
return string(ar[:n]) == string(br[:n])
}
+40
View File
@@ -0,0 +1,40 @@
package memory
import "testing"
func TestRecallAllowed(t *testing.T) {
cases := []struct {
name string
query, text string
want bool
}{
// The #470 shape: a world question and a note about his box.
{"world question, unrelated note", "почему небо синее", "сеть какая-то медленная", false},
{"world question, unrelated fact", "какая столица Франции", "какая последняя версия языка Go", false},
{"silent fixture case", "во сколько отходит поезд", "бэкап запускается в три ночи", false},
// A world question that does name the topic keeps its answer.
{"world question, same topic", "какой поезд идёт в Минск", "поезда в Минск ходят утром", true},
// A question about his own life is judged by the embedder alone,
// because recall exists for words he no longer remembers.
{"about him, no shared word", "во сколько я обычно засыпаю", "ложусь около одиннадцати", true},
{"about him, english", "which colour scheme do i like", "тёмная тема везде", true},
// Inflection must not break a match.
{"inflected", "чем кормить кота", "корм для кота в шкафу", true},
}
for _, c := range cases {
if got := RecallAllowed(c.query, c.text); got != c.want {
t.Errorf("%s: RecallAllowed(%q, %q) = %v, want %v", c.name, c.query, c.text, got, c.want)
}
}
}
// A question made only of filler has no topic word to match on, and the score
// gate is then the only judge it can have.
func TestRecallAllowedFallsBackWhenNothingToCompare(t *testing.T) {
if !RecallAllowed("что это", "сеть какая-то медленная") {
t.Error("a question with no content word must not be vetoed")
}
}
+11 -3
View File
@@ -378,7 +378,7 @@ func scoreCase(ctx context.Context, emb router.Embedder, newStore NewStore, minS
if len(hits) > 1 {
o.Margin = hits[0].Score - hits[1].Score
}
o.Recalled = bestRecall(hits, minScore, minMargin)
o.Recalled = bestRecall(c.Query, hits, minScore, minMargin)
}
for i, h := range hits {
if h.ID != c.Want {
@@ -424,11 +424,19 @@ func rankNote(inTop3 bool) string {
// is not importable; recalleval_test.go asserts the two agree in behaviour.
// The daemon returns the whole hit (a note and a fact are said differently);
// the harness only scores what came back, so it keeps returning the text.
func bestRecall(results []memory.Result, minScore, minMargin float64) string {
// bestRecall mirrors the daemon's gate in cmd/mavend/recall.go, including the
// topic veto added for #470: a score that clears the gate still has to be
// about what he asked. Keep the two in step — a fixture that measures a
// weaker gate than the daemon runs flatters it.
func bestRecall(query string, results []memory.Result, minScore, minMargin float64) string {
if !memory.Confident(results, minScore, minMargin) {
return ""
}
return results[0].Meta["text"]
text := results[0].Meta["text"]
if !memory.RecallAllowed(query, text) {
return ""
}
return text
}
func bump(m map[string]TagStat, key string, pass bool) {
+15 -8
View File
@@ -140,21 +140,22 @@ func words(s string) []string {
// TestBestRecallMatchesDaemon — the harness duplicates bestRecall from
// cmd/mavend/recall.go (package main is not importable). This pins the copy to
// the original's three rules: no hits, below the gate, or no text ⇒ silence.
// the original's rules: no hits, below the gate, no text, or no shared topic
// word ⇒ silence.
func TestBestRecallMatchesDaemon(t *testing.T) {
if got := bestRecall(nil, 0.55, 0); got != "" {
if got := bestRecall("чай", nil, 0.55, 0); got != "" {
t.Errorf("no hits: got %q, want silence", got)
}
low := []memory.Result{{ID: "a", Score: 0.4, Meta: map[string]string{"text": "чай"}}}
if got := bestRecall(low, 0.55, 0); got != "" {
if got := bestRecall("чай", low, 0.55, 0); got != "" {
t.Errorf("below gate: got %q, want silence", got)
}
noText := []memory.Result{{ID: "a", Score: 0.9, Meta: map[string]string{}}}
if got := bestRecall(noText, 0.55, 0); got != "" {
if got := bestRecall("чай", noText, 0.55, 0); got != "" {
t.Errorf("no text: got %q, want silence", got)
}
ok := []memory.Result{{ID: "a", Score: 0.9, Meta: map[string]string{"text": "чай"}}}
if got := bestRecall(ok, 0.55, 0); got != "чай" {
if got := bestRecall("чай", ok, 0.55, 0); got != "чай" {
t.Errorf("above gate: got %q, want %q", got, "чай")
}
// Margin: a close runner-up means the embedder cannot tell the two apart,
@@ -163,17 +164,23 @@ func TestBestRecallMatchesDaemon(t *testing.T) {
{ID: "a", Score: 0.86, Meta: map[string]string{"text": "чай"}},
{ID: "b", Score: 0.85, Meta: map[string]string{"text": "кофе"}},
}
if got := bestRecall(close, 0.55, 0.03); got != "" {
if got := bestRecall("чай", close, 0.55, 0.03); got != "" {
t.Errorf("thin margin: got %q, want silence", got)
}
if got := bestRecall(close, 0.55, 0); got != "чай" {
if got := bestRecall("чай", close, 0.55, 0); got != "чай" {
t.Errorf("margin off: got %q, want %q", got, "чай")
}
// The topic veto (#470): the score is fine and the note is about
// something else.
offTopic := []memory.Result{{ID: "a", Score: 0.9, Meta: map[string]string{"text": "сеть какая-то медленная"}}}
if got := bestRecall("почему небо синее", offTopic, 0.55, 0); got != "" {
t.Errorf("off topic: got %q, want silence", got)
}
clear := []memory.Result{
{ID: "a", Score: 0.86, Meta: map[string]string{"text": "чай"}},
{ID: "b", Score: 0.70, Meta: map[string]string{"text": "кофе"}},
}
if got := bestRecall(clear, 0.55, 0.03); got != "чай" {
if got := bestRecall("чай", clear, 0.55, 0.03); got != "чай" {
t.Errorf("wide margin: got %q, want %q", got, "чай")
}
}
+1 -1
View File
@@ -41,7 +41,7 @@
"tags": ["preference", "homelab", "paraphrase", "hard"],
"query": "когда запускать резервное копирование",
"want": "n1",
"note": "The DESIGN.md preference-seam example, phrased as the operator would ask it later.",
"note": "The docs/design.md preference-seam example, phrased as the operator would ask it later.",
"notes": [
{"id": "n1", "text": "бэкапы лучше делать ночью в три часа", "kind": "note"},
{"id": "n2", "text": "обновления ставлю по субботам", "kind": "note"},
+1 -1
View File
@@ -1,5 +1,5 @@
// Package morning is maven's morning routine engine — item #3 off the
// 2026-07-20 backlog (see Maven/20-07-2026-BACKLOG.md).
// 2026-07-20 backlog (Vikunja #280).
//
// A Routine is NOT four independent reminder timers. It's a checklist for a
// daily window: several Items, each evidenced by a fact key, completed in
+277
View File
@@ -0,0 +1,277 @@
// Package netaddr parses a daemon seam address and dials or binds it.
//
// Every seam between Maven's daemons used to be a unix socket with the
// network hardcoded at the call site — two dials in internal/ipc, one listen,
// and the same pair again in internal/worker. That is correct for co-located
// daemons and it is the reason a module cannot live on another host. This
// package moves the choice into the address string so a deploy picks the
// transport, not a recompile:
//
// /run/maven/stt.sock unix (the default, unchanged)
// unix:///run/maven/stt.sock unix (explicit, same thing)
// tcp://workstation:9310?token=hunter2 tcp
//
// A scheme-less address is unix and behaves exactly as it did before this
// package existed: same 0700 parent dir, same 0600 socket, same bytes on the
// wire with no handshake in front of them.
//
// Over TCP the filesystem permission that authenticated the unix socket is
// gone, and what crosses this seam is audio of the owner speaking and the
// text of his turns. So a TCP seam carries a shared token, checked before the
// first protocol frame is read. Wireguard is supported underneath and is not
// required.
package netaddr
import (
"crypto/subtle"
"errors"
"fmt"
"net"
"net/url"
"os"
"path/filepath"
"strings"
"time"
"golang.org/x/sys/unix"
)
// ErrUnauthorized — the peer presented a token the listener does not accept,
// or presented none when one is required.
var ErrUnauthorized = errors.New("netaddr: unauthorized")
// Addr is a parsed seam endpoint.
type Addr struct {
// Network is "unix" or "tcp".
Network string
// Address is the socket path (unix) or host:port (tcp).
Address string
// Token is the shared secret for a tcp seam. Empty for unix, where the
// filesystem does the same job.
Token string
}
// String renders the address for logs and errors. The token is never included.
func (a Addr) String() string {
if a.Network == "unix" {
return a.Address
}
return a.Network + "://" + a.Address
}
// IsUnix reports whether this seam is a unix socket, and so is local, is
// authenticated by file permissions, and needs no handshake.
func (a Addr) IsUnix() bool { return a.Network == "unix" }
// Parse reads a seam address. Anything without a "scheme://" prefix is a unix
// socket path, which keeps every existing config and every default working
// untouched.
func Parse(s string) (Addr, error) {
if !strings.Contains(s, "://") {
return Addr{Network: "unix", Address: s}, nil
}
u, err := url.Parse(s)
if err != nil {
return Addr{}, fmt.Errorf("netaddr: parse %q: %w", s, err)
}
switch u.Scheme {
case "unix":
return Addr{Network: "unix", Address: u.Path}, nil
case "tcp":
if u.Host == "" {
return Addr{}, fmt.Errorf("netaddr: %q has no host:port", s)
}
return Addr{Network: "tcp", Address: u.Host, Token: u.Query().Get("token")}, nil
default:
return Addr{}, fmt.Errorf("netaddr: unsupported scheme %q", u.Scheme)
}
}
// MustParse is Parse for a literal known good at compile time. It panics on a
// bad address, so use it in tests and constants, never on config input.
func MustParse(s string) Addr {
a, err := Parse(s)
if err != nil {
panic(err)
}
return a
}
// handshakeTimeout bounds the token exchange. A peer that cannot write one
// short line in this long is not going to serve a turn either.
const handshakeTimeout = 5 * time.Second
// greeting prefixes the token line. Versioned so a later mTLS seam can be
// told apart from this one on the wire.
const greeting = "MAVEN1 "
// Dial connects to a. On a tcp seam it sends the token and waits for the
// listener to accept it, so a returned conn is already authorized and the
// caller can write its first protocol frame.
func Dial(a Addr) (net.Conn, error) {
return DialTimeout(a, 0)
}
// DialTimeout is Dial with a bound on the connect. Zero means the operating
// system default. The token exchange gets its own timeout either way.
func DialTimeout(a Addr, timeout time.Duration) (net.Conn, error) {
var c net.Conn
var err error
if timeout > 0 {
c, err = net.DialTimeout(a.Network, a.Address, timeout)
} else {
c, err = net.Dial(a.Network, a.Address)
}
if err != nil {
return nil, err
}
if a.IsUnix() {
return c, nil
}
if err := clientHandshake(c, a.Token); err != nil {
_ = c.Close()
return nil, err
}
return c, nil
}
func clientHandshake(c net.Conn, token string) error {
_ = c.SetDeadline(time.Now().Add(handshakeTimeout))
defer c.SetDeadline(time.Time{})
if _, err := c.Write([]byte(greeting + token + "\n")); err != nil {
return fmt.Errorf("netaddr: send token: %w", err)
}
var reply [3]byte
if _, err := readFull(c, reply[:]); err != nil {
return fmt.Errorf("%w: %v", ErrUnauthorized, err)
}
if string(reply[:]) != "ok\n" {
return ErrUnauthorized
}
return nil
}
// Listener wraps a net.Listener so Accept performs the token check for a tcp
// seam. A connection that fails the check is closed and never surfaces, so
// the protocol above this layer only ever sees authorized peers.
type Listener struct {
net.Listener
addr Addr
}
// Accept returns the next authorized connection. Unauthorized peers are
// dropped and Accept keeps waiting: a bad token is a rejected stranger, not a
// reason to stop serving.
func (l *Listener) Accept() (net.Conn, error) {
for {
c, err := l.Listener.Accept()
if err != nil {
return nil, err
}
if l.addr.IsUnix() {
return c, nil
}
if err := serverHandshake(c, l.addr.Token); err != nil {
_ = c.Close()
continue
}
return c, nil
}
}
// Addr reports the parsed seam address this listener was built from.
func (l *Listener) SeamAddr() Addr { return l.addr }
func serverHandshake(c net.Conn, want string) error {
_ = c.SetDeadline(time.Now().Add(handshakeTimeout))
defer c.SetDeadline(time.Time{})
// The line is bounded: greeting, token, newline. Read a byte at a time so
// nothing of the first protocol frame is consumed when the token is short.
line := make([]byte, 0, 128)
var b [1]byte
for {
if _, err := readFull(c, b[:]); err != nil {
return err
}
if b[0] == '\n' {
break
}
line = append(line, b[0])
if len(line) > 512 {
return ErrUnauthorized
}
}
got, ok := strings.CutPrefix(string(line), greeting)
if !ok {
return ErrUnauthorized
}
if subtle.ConstantTimeCompare([]byte(got), []byte(want)) != 1 {
return ErrUnauthorized
}
if _, err := c.Write([]byte("ok\n")); err != nil {
return err
}
return nil
}
func readFull(c net.Conn, p []byte) (int, error) {
n := 0
for n < len(p) {
m, err := c.Read(p[n:])
n += m
if err != nil {
return n, err
}
}
return n, nil
}
// Listen binds a. A unix seam gets the perms it has always had: parent dir
// 0700, socket 0600, and any stale socket from a crashed daemon removed
// first. A tcp seam must carry a token, because there is no filesystem to
// stand in for one.
func Listen(a Addr) (*Listener, error) {
if a.IsUnix() {
ln, err := listenUnix(a.Address)
if err != nil {
return nil, err
}
return &Listener{Listener: ln, addr: a}, nil
}
if a.Token == "" {
return nil, fmt.Errorf("netaddr: listen %s: tcp seam requires a token", a)
}
ln, err := net.Listen("tcp", a.Address)
if err != nil {
return nil, fmt.Errorf("netaddr: listen %s: %w", a, err)
}
return &Listener{Listener: ln, addr: a}, nil
}
func listenUnix(path string) (net.Listener, error) {
_ = os.Remove(path) // stale socket from a crashed daemon; ignore missing
if err := os.MkdirAll(filepath.Dir(path), 0o700); err != nil {
return nil, fmt.Errorf("netaddr: mkdir socket dir: %w", err)
}
// umask could widen the perms on socket creation; tighten then chmod to
// be explicit. 0600 ⇒ read+write by owner only.
oldMask := unix.Umask(0o077)
ln, err := net.Listen("unix", path)
unix.Umask(oldMask)
if err != nil {
return nil, fmt.Errorf("netaddr: listen %s: %w", path, err)
}
if err := os.Chmod(path, 0o600); err != nil {
_ = ln.Close()
_ = os.Remove(path)
return nil, fmt.Errorf("netaddr: chmod socket: %w", err)
}
return ln, nil
}
// Cleanup removes the socket file behind a unix seam. It is a no-op for tcp.
func Cleanup(a Addr) {
if a.IsUnix() && a.Address != "" {
_ = os.Remove(a.Address)
}
}
+185
View File
@@ -0,0 +1,185 @@
package netaddr
import (
"errors"
"net"
"path/filepath"
"testing"
)
// A scheme-less address must stay unix. Every deploy in the tree writes a bare
// path, so this is the test that says the transport change costs them nothing.
func TestParseSchemelessIsUnix(t *testing.T) {
a, err := Parse("/run/maven/stt.sock")
if err != nil {
t.Fatalf("parse: %v", err)
}
if !a.IsUnix() {
t.Fatalf("want unix, got %q", a.Network)
}
if a.Address != "/run/maven/stt.sock" {
t.Fatalf("address = %q", a.Address)
}
if a.Token != "" {
t.Fatalf("unix seam carries a token: %q", a.Token)
}
}
func TestParse(t *testing.T) {
cases := []struct {
in string
net, addr, tk string
wantErr bool
}{
{in: "", net: "unix", addr: ""},
{in: "unix:///run/maven/core.sock", net: "unix", addr: "/run/maven/core.sock"},
{in: "tcp://workstation:9310", net: "tcp", addr: "workstation:9310"},
{in: "tcp://workstation:9310?token=hunter2", net: "tcp", addr: "workstation:9310", tk: "hunter2"},
{in: "tcp://", wantErr: true},
{in: "udp://workstation:9310", wantErr: true},
}
for _, c := range cases {
a, err := Parse(c.in)
if c.wantErr {
if err == nil {
t.Errorf("Parse(%q) = %v, want error", c.in, a)
}
continue
}
if err != nil {
t.Errorf("Parse(%q): %v", c.in, err)
continue
}
if a.Network != c.net || a.Address != c.addr || a.Token != c.tk {
t.Errorf("Parse(%q) = %+v, want %s/%s/%s", c.in, a, c.net, c.addr, c.tk)
}
}
}
// The token must never reach a log line.
func TestStringHidesToken(t *testing.T) {
a := MustParse("tcp://workstation:9310?token=hunter2")
if got := a.String(); got != "tcp://workstation:9310" {
t.Fatalf("String() = %q", got)
}
}
// A unix seam must round-trip with no handshake in front of the payload: the
// first bytes the listener sees are the caller's, exactly as before.
func TestUnixRoundTripHasNoHandshake(t *testing.T) {
a := MustParse(filepath.Join(t.TempDir(), "s.sock"))
ln, err := Listen(a)
if err != nil {
t.Fatalf("listen: %v", err)
}
defer ln.Close()
go echoOnce(ln)
c, err := Dial(a)
if err != nil {
t.Fatalf("dial: %v", err)
}
defer c.Close()
if got := roundTrip(t, c, "hello"); got != "hello" {
t.Fatalf("got %q", got)
}
}
func TestTCPRoundTripWithToken(t *testing.T) {
ln, addr := listenLoopback(t, "s3cret")
defer ln.Close()
go echoOnce(ln)
c, err := Dial(addr)
if err != nil {
t.Fatalf("dial: %v", err)
}
defer c.Close()
if got := roundTrip(t, c, "hello"); got != "hello" {
t.Fatalf("got %q", got)
}
}
func TestTCPWrongTokenIsRejected(t *testing.T) {
ln, addr := listenLoopback(t, "s3cret")
defer ln.Close()
// Accept keeps waiting past the bad peer, so nothing here should ever
// reach the echo. A conn that does means the token was not checked.
go echoOnce(ln)
bad := addr
bad.Token = "wrong"
if _, err := Dial(bad); !errors.Is(err, ErrUnauthorized) {
t.Fatalf("dial with wrong token: err = %v, want ErrUnauthorized", err)
}
}
// A stranger that speaks the protocol instead of the greeting is dropped, and
// the listener stays up for the peer that follows it.
func TestTCPUngreetedPeerDoesNotKillTheListener(t *testing.T) {
ln, addr := listenLoopback(t, "s3cret")
defer ln.Close()
go echoOnce(ln)
raw, err := net.Dial("tcp", addr.Address)
if err != nil {
t.Fatalf("raw dial: %v", err)
}
if _, err := raw.Write([]byte("GET / HTTP/1.1\n")); err != nil {
t.Fatalf("raw write: %v", err)
}
raw.Close()
c, err := Dial(addr)
if err != nil {
t.Fatalf("dial after stranger: %v", err)
}
defer c.Close()
if got := roundTrip(t, c, "still here"); got != "still here" {
t.Fatalf("got %q", got)
}
}
// A tcp seam with no token is a misconfiguration, and it must fail at bind
// rather than serve the owner's turns to anyone who connects.
func TestTCPListenRequiresToken(t *testing.T) {
if _, err := Listen(MustParse("tcp://127.0.0.1:0")); err == nil {
t.Fatal("listen on a tokenless tcp seam succeeded")
}
}
func listenLoopback(t *testing.T, token string) (*Listener, Addr) {
t.Helper()
ln, err := Listen(Addr{Network: "tcp", Address: "127.0.0.1:0", Token: token})
if err != nil {
t.Fatalf("listen: %v", err)
}
return ln, Addr{Network: "tcp", Address: ln.Addr().String(), Token: token}
}
func echoOnce(ln *Listener) {
c, err := ln.Accept()
if err != nil {
return
}
defer c.Close()
buf := make([]byte, 256)
n, err := c.Read(buf)
if err != nil {
return
}
_, _ = c.Write(buf[:n])
}
func roundTrip(t *testing.T, c net.Conn, msg string) string {
t.Helper()
if _, err := c.Write([]byte(msg)); err != nil {
t.Fatalf("write: %v", err)
}
buf := make([]byte, 256)
n, err := c.Read(buf)
if err != nil {
t.Fatalf("read: %v", err)
}
return string(buf[:n])
}
+3 -3
View File
@@ -15,7 +15,7 @@ const (
CheckLang = "lang" // the operator's language, not the prompt's
CheckLength = "length" // a nudge is one sentence, not a paragraph
CheckFeminine = "feminine" // her self-reference is feminine (hard constraint)
CheckCringe = "cringe" // DESIGN.md § Non-goals, "not a relationship"
CheckCringe = "cringe" // docs/design.md § Non-goals, "not a relationship"
CheckOnTopic = "ontopic" // says the thing the rule is about
// CheckHisGender — the other half of the persona rule: SHE is feminine, HE
@@ -105,7 +105,7 @@ func checkLength(body string) Result {
// --- feminine self-reference ---------------------------------------------
//
// The hard constraint (CLAUDE.md, DESIGN.md § Identity): Maven's Russian
// The hard constraint (CLAUDE.md, docs/design.md § Identity): Maven's Russian
// self-reference is feminine. The operator is male, so second-person forms
// addressed to him are MASCULINE and must not be flagged — "ты не пил воду" is
// correct, "я напомнил" is not. Both directions matter, which is why this is a
@@ -497,7 +497,7 @@ func isLatinWord(w string) bool {
// --- the cringe checks ---------------------------------------------------
//
// "Think Jarvis without the cringe part". DESIGN.md § Non-goals: "Not a
// "Think Jarvis without the cringe part". docs/design.md § Non-goals: "Not a
// relationship — mom-tone is a function that makes nudges land, not emotional
// company. Names the drift a warm small model falls into." Each pattern below
// is one shape of that drift. They are deliberately specific: a check that
+2 -2
View File
@@ -11,7 +11,7 @@
// length test that a human can read and disagree with. A score here is a claim
// about measurable properties, not about whether a sentence is good.
//
// DESIGN.md § "Rules decide, LLM phrases" is why there is no send/veto signal
// docs/design.md § "Rules decide, LLM phrases" is why there is no send/veto signal
// anywhere in this package: the rule already decided she speaks. The phraser
// only words it, so a nudge the model refuses to write is a failure, never a
// legitimate outcome.
@@ -38,7 +38,7 @@ var fixtureJSON []byte
const SchemaVersion = 1
// Case — one nudge situation, as a real tick would present it. The fields are
// the (rule, severity, context) input DESIGN.md names, flattened to JSON.
// the (rule, severity, context) input docs/design.md names, flattened to JSON.
//
// WantAny is the on-topic contract: at least one of these lowercased fragments
// must appear in the message. A water nudge that never mentions water is a
+55 -12
View File
@@ -46,6 +46,12 @@ type LLMPhraser struct {
launch func(ctx context.Context, cfg Config) (backend, error)
probe func(ctx context.Context, base string) (string, error)
// remote — the workstation model, when one is configured. Set once at wiring
// time by UseRemote and read on every phrasing call. nil ⇒ every call goes to
// the resident llama-server this phraser owns, which is the whole deploy
// before a `workstation` block exists. See world.go.
remote Remote
// swapMu — single-flight around Swap. Held for the whole swap, including the
// model load, so two concurrent swap requests can never both be loading.
swapMu sync.Mutex
@@ -363,10 +369,7 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
// prompt guaranteed to make a small model fill the gap from memory.
notes = nonEmpty(notes)
if len(notes) == 0 {
// General knowledge — no notes to ground the answer. The system
// prompt is the single tested source in router.KnowledgePrompt.
sys := persona.Prepend(p.cfg.ContextBlock, router.KnowledgePrompt())
prompt := fmt.Sprintf("Пользователь спрашивает: \"%s\".", utterance)
sys, prompt := p.knowledgePrompt(utterance)
resp, err := p.chatWithSystem(ctx, sys, prompt, 768)
if err != nil || resp == "" {
return "не знаю.", nil
@@ -381,11 +384,7 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
}
return resp, nil
}
sys := p.querySystemPrompt()
prompt := fmt.Sprintf(
"Он спрашивает: \"%s\"\n\nИсточники:\n%s\nОтветь ему коротко и своими словами, опираясь только на эти источники. Если ответа в них нет — так и скажи.",
utterance, evidenceBlock(notes),
)
sys, prompt := p.evidencePrompt(utterance, notes)
resp, err := p.chatWithSystem(ctx, sys, prompt, 768)
text, _, perr := parseResponseMood(resp)
if err != nil || perr != nil {
@@ -481,6 +480,17 @@ func chatSystemPrompt(block func() string) string {
// the LLM completion endpoint. Like chatWithSystem but for an arbitrary message
// slice — the caller owns the system prompt placement.
func (p *LLMPhraser) chatWithMessages(ctx context.Context, msgs []chatMsg, maxTokens int) (string, error) {
// Same silent preference as chatWithSystem, when the array is the shape
// llm.Req can carry: one system turn and one user turn. PhraseChat already
// folds the history into a single user message (some chat templates reject
// consecutive user turns), so today that is every call. A longer array goes
// to the resident model rather than get flattened here, because flattening a
// conversation is a decision its owner should make.
if len(msgs) == 2 && msgs[0].Role == "system" && msgs[1].Role == "user" {
if out, ok := p.remoteChat(ctx, msgs[0].Content, msgs[1].Content, maxTokens); ok {
return out, nil
}
}
base, release, err := p.acquire()
if err != nil {
return "", err
@@ -534,8 +544,14 @@ func (p *LLMPhraser) PhraseReminder(ctx context.Context, d loop.ReminderDecision
text = "reminder"
}
// Russian, like the other two prompts (Vikunja #404). Asking a model for a
// Russian reply in English is asking it to switch languages mid-prompt,
// and a 1.7B sometimes answers in the language it was asked in. The
// persona rules and the JSON contract are not repeated here: this call
// goes through chat(), so nudgeSystem already states both, and a second
// statement of the same contract is one more thing that can drift.
prompt := fmt.Sprintf(
`The user set a reminder: "%s". Rephrase it briefly as a gentle nudge. Respond as JSON: {"response": "...", "mood": "..."}`,
`Он поставил напоминание: "%s". Скажи это своими словами, коротко и мягко — одно предложение.`,
text,
)
resp, err := p.chat(ctx, prompt)
@@ -634,6 +650,12 @@ func (p *LLMPhraser) chat(ctx context.Context, userPrompt string) (string, error
}
func (p *LLMPhraser) chatWithSystem(ctx context.Context, system, user string, maxTokens int) (string, error) {
// The workstation model first when it will take work, and silently: every
// caller of this helper is on the silent half of the degradation rule. It
// answering is not news, and it being asleep is not news either.
if out, ok := p.remoteChat(ctx, system, user, maxTokens); ok {
return out, nil
}
base, release, err := p.acquire()
if err != nil {
return "", err
@@ -692,7 +714,7 @@ func (p *LLMPhraser) chatWithSystem(ctx context.Context, system, user string, ma
// Written as filled-in examples, not as a schema with "..." in it. A 0.8B
// copies whatever sits in the response slot, so a literal placeholder there
// teaches it to answer with the placeholder. Measured: 7/15 nudges came back
// as "..." before this. See PHRASING-EVAL-31-07-2026.md.
// as "..." before this. See docs/evals/2026-07-31-phrasing.md.
//
// Russian only, feminine self-reference, second person masculine (the owner is
// a man). She talks TO him, informally, singular — never "вы", never "он".
@@ -727,6 +749,27 @@ func (p *LLMPhraser) systemPrompt() string {
return persona.Prepend(p.cfg.ContextBlock, nudgeSystem)
}
// knowledgePrompt — the no-sources branch: a world question, answered from
// weights alone. The system prompt is the single tested source in
// router.KnowledgePrompt.
//
// Split out of PhraseQuery so PhraseWorld sends the workstation model the same
// bytes the resident model gets. Prompt parity across two models is a stated
// constraint (CLAUDE.md), and two copies of a prompt is how it stops holding.
func (p *LLMPhraser) knowledgePrompt(utterance string) (sys, user string) {
return persona.Prepend(p.cfg.ContextBlock, router.KnowledgePrompt()),
fmt.Sprintf("Пользователь спрашивает: \"%s\".", utterance)
}
// evidencePrompt — the sources branch: read these, add nothing. Shared with
// PhraseWorld for the same reason as knowledgePrompt.
func (p *LLMPhraser) evidencePrompt(utterance string, notes []string) (sys, user string) {
return p.querySystemPrompt(), fmt.Sprintf(
"Он спрашивает: \"%s\"\n\nИсточники:\n%s\nОтветь ему коротко и своими словами, опираясь только на эти источники. Если ответа в них нет — так и скажи.",
utterance, evidenceBlock(notes),
)
}
// querySystemPrompt returns the system prompt for the evidence branch of
// PhraseQuery. Prepends the configured persona when set.
//
@@ -754,7 +797,7 @@ func (p *LLMPhraser) querySystemPrompt() string {
base := "Ты отвечаешь ему по источникам, которые тебе дали. Отвечай ТОЛЬКО по ним: всё, что ты говоришь, должно быть написано в источниках. " +
"Если ответа в них нет — так и скажи и на этом остановись; не добавляй ничего из своих знаний и не догадывайся. " +
"Не приплетай прошлые реплики разговора. " +
"Отвечай по-русски, коротко и своими словами, начинай с \"вот что я нашла: \". О себе — в женском роде, глаголы в прошедшем времени с окончанием -ла. Он мужчина, обращайся к нему на \"ты\". Respond ONLY with valid JSON: {\"response\": \"...\", \"mood\": \"neutral\"}."
"Отвечай по-русски, коротко и своими словами, начинай с \"вот что я нашла: \". О себе — в женском роде, глаголы в прошедшем времени с окончанием -ла. Он мужчина, обращайся к нему на \"ты\". Отвечай ТОЛЬКО одним объектом JSON: {\"response\": \"...\", \"mood\": \"neutral\"}."
return persona.Prepend(p.cfg.ContextBlock, base)
}
+2 -2
View File
@@ -1,9 +1,9 @@
// Package phraser is maven's "rules decide, llm phrases" seam — the layer
// that turns a loop decision into the body + summary the delivery module ships.
//
// Per DESIGN.md § Resident language model: the phraser is the resident model
// Per docs/design.md § Resident language model: the phraser is the resident model
// (Qwen3-1.7B — RU continued pretraining plus joint persona/router SFT, not a
// sub-1b prompted-only model as the retired spec claimed; see DESIGN.md
// sub-1b prompted-only model as the retired spec claimed; see docs/design.md
// § Superseded, "small-model phrasing claim"). It takes
// (rule, severity, context) and produces Body (full voice message, local — no
// shoulder-surf concern beyond who's in the room) + Summary (minimal body for
+256
View File
@@ -0,0 +1,256 @@
package phraser
import (
"context"
"errors"
"fmt"
"os"
"os/exec"
"path/filepath"
"regexp"
"strconv"
"strings"
"syscall"
"testing"
"time"
)
// The spawn path (NewLLMPhraser, spawnLlamaServer, startLlamaProc, llamaProc.Close)
// was at 0% coverage: every test built the phraser with NewLLMPhraserAt, which
// starts no process. These tests drive the real spawn code against a fake
// llama-server script, so the startup race arms and the reaping are exercised
// without a model or a GPU.
// fakeLlama writes an executable script standing in for llama-server and returns
// its path. body runs after the script has recorded its own pid.
func fakeLlama(t *testing.T, body string) string {
t.Helper()
dir := t.TempDir()
path := filepath.Join(dir, "fake-llama-server")
script := "#!/bin/sh\n" + body + "\n"
if err := os.WriteFile(path, []byte(script), 0o755); err != nil {
t.Fatalf("write fake server: %v", err)
}
return path
}
// listensThenSleeps prints the line startLlamaProc scrapes, then stays alive
// until killed — the shape of a real llama-server that came up.
const listensThenSleeps = `echo "srv load_model: listening on http://127.0.0.1:18081" >&2
while : ; do sleep 1 ; done`
func testCfg(bin string) Config {
cfg := DefaultConfig("/nonexistent/model.gguf")
cfg.BinPath = bin
return cfg
}
func TestExtractPort(t *testing.T) {
for _, tc := range []struct{ in, want string }{
{"127.0.0.1:0", "0"},
{"127.0.0.1:8080", "8080"},
{"127.0.0.1:", "0"},
{"", "0"},
{"8080", "0"}, // no colon: Cut yields no port, so the caller gets the "any port" default
} {
if got := extractPort(tc.in); got != tc.want {
t.Errorf("extractPort(%q) = %q, want %q", tc.in, got, tc.want)
}
}
}
func TestStartLlamaProcScrapesPortAndReaps(t *testing.T) {
bin := fakeLlama(t, listensThenSleeps)
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
p, err := startLlamaProc(ctx, testCfg(bin))
if err != nil {
t.Fatalf("startLlamaProc: %v", err)
}
if p.BaseURL() != "http://127.0.0.1:18081" {
t.Fatalf("BaseURL = %q, want the scraped address", p.BaseURL())
}
pid := p.cmd.Process.Pid
p.cancel = cancel
if err := p.Close(); err != nil {
t.Fatalf("Close: %v", err)
}
// Close must Wait, otherwise the child lingers as a zombie.
if p.cmd.ProcessState == nil {
t.Fatal("Close did not reap the child: ProcessState is nil")
}
if err := syscall.Kill(pid, 0); err == nil {
t.Fatalf("child %d still exists after Close", pid)
}
}
func TestStartLlamaProcFailureArms(t *testing.T) {
t.Run("binary missing", func(t *testing.T) {
cfg := testCfg(filepath.Join(t.TempDir(), "does-not-exist"))
_, err := startLlamaProc(context.Background(), cfg)
if err == nil || !strings.Contains(err.Error(), "llm: start") {
t.Fatalf("err = %v, want a start failure", err)
}
})
t.Run("server exits without listening", func(t *testing.T) {
// stderr closes, so the reader goroutine reports EOF on errCh.
bin := fakeLlama(t, `echo "ggml_vulkan: no device" >&2
exit 1`)
_, err := startLlamaProc(context.Background(), testCfg(bin))
if err == nil || !strings.Contains(err.Error(), "llm: server output") {
t.Fatalf("err = %v, want the server-output arm", err)
}
})
t.Run("context cancelled during startup", func(t *testing.T) {
// Never prints the listen line and never exits: only ctx can end this.
bin := fakeLlama(t, `while : ; do sleep 1 ; done`)
ctx, cancel := context.WithCancel(context.Background())
go func() {
time.Sleep(150 * time.Millisecond)
cancel()
}()
defer cancel()
_, err := startLlamaProc(ctx, testCfg(bin))
if !errors.Is(err, context.Canceled) {
t.Fatalf("err = %v, want context.Canceled", err)
}
})
}
func TestNewLLMPhraserSpawns(t *testing.T) {
bin := fakeLlama(t, listensThenSleeps)
p, err := NewLLMPhraser(context.Background(), testCfg(bin))
if err != nil {
t.Fatalf("NewLLMPhraser: %v", err)
}
if p.BaseURL() != "http://127.0.0.1:18081" {
t.Fatalf("BaseURL = %q", p.BaseURL())
}
pid := p.be.(*llamaProc).cmd.Process.Pid
if err := p.Close(); err != nil {
t.Fatalf("Close: %v", err)
}
if p.BaseURL() != "" {
t.Fatalf("BaseURL after Close = %q, want empty", p.BaseURL())
}
if err := syscall.Kill(pid, 0); err == nil {
t.Fatalf("llama-server %d survived Close", pid)
}
}
func TestNewLLMPhraserSpawnFailure(t *testing.T) {
cfg := testCfg(filepath.Join(t.TempDir(), "does-not-exist"))
p, err := NewLLMPhraser(context.Background(), cfg)
if err == nil {
p.Close()
t.Fatal("want an error when the server cannot start")
}
if p != nil {
t.Fatalf("want a nil phraser on failure, got %#v", p)
}
}
// TestPdeathsigKillsOrphan is the orphan test the task asked for. A SIGKILLed
// mavend never runs Close, so nothing but the kernel's Pdeathsig can stop its
// llama-server. Re-exec this test binary as the "daemon", let it spawn the fake
// server, SIGKILL the daemon, and assert the grandchild died with it.
func TestPdeathsigKillsOrphan(t *testing.T) {
bin := fakeLlama(t, listensThenSleeps)
cmd := exec.Command(os.Args[0], "-test.run=TestSpawnHelperProcess", "-test.v=false")
cmd.Env = append(os.Environ(), "MAVEN_SPAWN_HELPER=1", "MAVEN_FAKE_LLAMA="+bin)
out, err := cmd.StdoutPipe()
if err != nil {
t.Fatalf("stdout pipe: %v", err)
}
if err := cmd.Start(); err != nil {
t.Fatalf("start helper: %v", err)
}
defer func() { _ = cmd.Process.Kill(); _ = cmd.Wait() }()
buf := make([]byte, 256)
n, err := out.Read(buf)
if err != nil {
t.Fatalf("read child pid: %v", err)
}
childPID, err := strconv.Atoi(strings.TrimSpace(string(buf[:n])))
if err != nil {
t.Fatalf("helper printed %q, want a pid: %v", string(buf[:n]), err)
}
if err := syscall.Kill(childPID, 0); err != nil {
t.Fatalf("llama-server %d not running before the kill: %v", childPID, err)
}
// SIGKILL: the helper gets no chance to clean up, exactly like an OOM kill.
if err := cmd.Process.Signal(syscall.SIGKILL); err != nil {
t.Fatalf("kill helper: %v", err)
}
_, _ = cmd.Process.Wait()
deadline := time.Now().Add(5 * time.Second)
for time.Now().Before(deadline) {
if err := syscall.Kill(childPID, 0); err != nil {
return // gone: Pdeathsig did its job
}
time.Sleep(20 * time.Millisecond)
}
_ = syscall.Kill(childPID, syscall.SIGKILL)
t.Fatalf("llama-server %d outlived the SIGKILLed parent", childPID)
}
// TestSpawnHelperProcess is not a test. It is the child half of
// TestPdeathsigKillsOrphan: spawn a llama-server, print its pid, then block.
func TestSpawnHelperProcess(t *testing.T) {
if os.Getenv("MAVEN_SPAWN_HELPER") != "1" {
t.Skip("helper for TestPdeathsigKillsOrphan")
}
cfg := testCfg(os.Getenv("MAVEN_FAKE_LLAMA"))
p, err := startLlamaProc(context.Background(), cfg)
if err != nil {
fmt.Println("spawn failed:", err)
os.Exit(1)
}
fmt.Println(p.cmd.Process.Pid)
os.Stdout.Sync()
select {} // wait to be killed
}
// TestKillMavenScriptMatchesRealCommandLine pins kill-maven.sh's fallback
// pattern to the command line startLlamaProc actually builds. The script leaked
// orphans twice already, both times because the pattern stopped matching: first
// `llama-server.*maven`, then a hardcoded model name after the model was swapped.
func TestKillMavenScriptMatchesRealCommandLine(t *testing.T) {
src, err := os.ReadFile("../../kill-maven.sh")
if err != nil {
t.Fatalf("read kill-maven.sh: %v", err)
}
m := regexp.MustCompile(`(?m)^\s*LLM='([^']+)'`).FindSubmatch(src)
if m == nil {
t.Fatal("no default LLM='...' pattern in kill-maven.sh")
}
pat, err := regexp.Compile(string(m[1]))
if err != nil {
t.Fatalf("LLM pattern %q does not compile: %v", m[1], err)
}
// Rebuild the command line from the production arg list, so a change to
// startLlamaProc that breaks the sweep fails here instead of on the box.
cfg := DefaultConfig("/opt/maven/models/llm/Qwen3-1.7B-UD-Q4_K_XL.gguf")
cfg.NCtx, cfg.NGpuLayers = 4096, 99
cmdline := strings.Join([]string{
cfg.BinPath,
"-m", cfg.ModelPath,
"--host", "127.0.0.1",
"--port", extractPort(cfg.Listen),
"-c", fmt.Sprintf("%d", cfg.NCtx),
"-ngl", fmt.Sprintf("%d", cfg.NGpuLayers),
"--no-webui",
}, " ")
if !pat.MatchString(cmdline) {
t.Fatalf("kill-maven.sh pattern %q does not match %q — orphans would leak", m[1], cmdline)
}
}
+137
View File
@@ -0,0 +1,137 @@
package phraser
import (
"context"
"errors"
"log"
"github.com/kami/maven/internal/llm"
)
// Remote — the workstation model, seen from the phraser. `*llm.Pair` satisfies
// it, and a test fake satisfies it in three lines.
//
// Only the refusing half of Pair is here on purpose. Pair.Complete falls back to
// its own floor client, and the phraser already owns a floor: the llama-server it
// spawned. Two floors under one call is one too many, so the phraser asks whether
// the remote will take work, uses it when it will, and otherwise does exactly
// what it did before this file existed.
type Remote interface {
// Available is an atomic read of a cached probe, so it is free to call per
// turn. See llm.Pair.
Available() bool
// CompleteRemote runs on the workstation or returns ErrRemoteUnavailable. It
// never falls back.
CompleteRemote(ctx context.Context, r llm.Req) (string, error)
}
// ErrNoWorldModel — a world question was asked, a workstation model is
// configured to answer it, and that machine is not answering. The caller turns
// this into a gap he is told about ("не могу сейчас"), never into an answer from
// the resident model.
//
// This is the naming half of the degradation rule in docs/offload.md. The
// resident Qwen3-1.7B does not answer a world question worse than the 12B, it
// invents: measured, the workstation model scores knowledge 9/9 on the talk
// fixture against the resident model's confabulations
// (docs/evals/2026-08-02-workstation-gemma4-12b.md).
var ErrNoWorldModel = errors.New("phraser: no world model available")
// chatTemperature — what the phraser's own transport has always sampled at.
// Named so the remote path cannot drift from it silently. Whether 0.7 is right
// at all is Vikunja #402, and answering that here would hide a phrasing change
// inside a routing change.
const chatTemperature = 0.7
// UseRemote points the phraser at the workstation model. Wiring time only, once,
// before anything phrases: the field is read without a lock on every call
// because a per-turn lock to answer a question that changes at deploy time is
// not worth paying for.
//
// A nil remote is the normal state of a box with no `workstation` block, and it
// must behave exactly as the box behaved before this seam existed.
func (p *LLMPhraser) UseRemote(r Remote) {
p.remote = r
}
// PhraseWorld answers a question about the world — either from the model's own
// knowledge (no sources) or from a passage someone fetched (a live search, a ZIM
// article, a page he named). Three outcomes, and the middle one is the point:
//
// - No workstation configured. The resident model answers, exactly as it does
// today. Naming a gap needs a gap: on a box that never had a second model,
// refusing every world question would remove a capability he has now.
// - Workstation configured and taking work. It answers.
// - Workstation configured and down. ErrNoWorldModel, and the caller says so.
//
// The prompts are the ones PhraseQuery uses, built by the same two functions, so
// the two models are asked the same question in the same words.
func (p *LLMPhraser) PhraseWorld(ctx context.Context, utterance string, sources []string) (string, error) {
sources = nonEmpty(sources)
if p.remote == nil {
return p.PhraseQuery(ctx, utterance, sources)
}
var sys, user string
if len(sources) == 0 {
sys, user = p.knowledgePrompt(utterance)
} else {
sys, user = p.evidencePrompt(utterance, sources)
}
if !p.remote.Available() {
return "", ErrNoWorldModel
}
resp, err := p.remote.CompleteRemote(ctx, llm.Req{
System: sys,
User: user,
Grammar: p.grammar(),
MaxTokens: 768,
Temperature: chatTemperature,
})
if err != nil {
// The cached probe was one interval stale, or the card went away
// mid-request. Either way this is the gap, not an error to log and
// paper over with the smaller model.
log.Printf("phraser: world model: %v", err)
return "", errors.Join(ErrNoWorldModel, err)
}
resp = stripThink(resp)
text, _, perr := parseResponseMood(resp)
if perr != nil {
log.Printf("phraser: PhraseWorld: %v", perr)
return "", errors.Join(ErrNoWorldModel, perr)
}
if text != "" {
return text, nil
}
if resp == "" {
return "", ErrNoWorldModel
}
return resp, nil
}
// remoteChat is the silent half, for the phrasing paths where the workstation
// model is only better: a nudge, a reminder, a reply, a question answered from
// his own notes. It reports whether it answered; it never reports why not,
// because the caller's next move is the resident model either way.
//
// He is not told which of the two models phrased his reply. That is the rule.
func (p *LLMPhraser) remoteChat(ctx context.Context, system, user string, maxTokens int) (string, bool) {
if p.remote == nil || !p.remote.Available() {
return "", false
}
out, err := p.remote.CompleteRemote(ctx, llm.Req{
System: system,
User: user,
Grammar: p.grammar(),
MaxTokens: maxTokens,
Temperature: chatTemperature,
})
if err != nil {
log.Printf("phraser: workstation model declined, phrasing here instead: %v", err)
return "", false
}
if out = stripThink(out); out == "" {
return "", false
}
return out, true
}
+164
View File
@@ -0,0 +1,164 @@
package phraser
import (
"context"
"errors"
"strings"
"testing"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/loop"
)
// fakeRemote — a workstation model that is up or down on command, and records
// what it was asked.
type fakeRemote struct {
up bool
reply string
err error
got []llm.Req
}
func (f *fakeRemote) Available() bool { return f.up }
func (f *fakeRemote) CompleteRemote(_ context.Context, r llm.Req) (string, error) {
f.got = append(f.got, r)
if f.err != nil {
return "", f.err
}
return f.reply, nil
}
// The three outcomes of the naming half, in one place. The middle one is the
// whole task: a gap he is told about, not an answer from the smaller model.
func TestPhraseWorldNamesTheGapOnlyWhenThereIsOne(t *testing.T) {
answer := `{"response": "Небо голубое из-за рэлеевского рассеяния.", "mood": "neutral"}`
t.Run("no workstation configured: the resident model answers as today", func(t *testing.T) {
spy := newPromptSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{})
got, err := p.PhraseWorld(context.Background(), "почему небо голубое", nil)
if err != nil {
t.Fatalf("PhraseWorld: %v", err)
}
if got == "" {
t.Fatal("no reply from the resident model")
}
if len(spy.user) != 1 {
t.Fatalf("resident model saw %d requests, want 1", len(spy.user))
}
})
t.Run("workstation up: it answers and the resident model is not asked", func(t *testing.T) {
spy := newPromptSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{})
remote := &fakeRemote{up: true, reply: answer}
p.UseRemote(remote)
got, err := p.PhraseWorld(context.Background(), "почему небо голубое", nil)
if err != nil {
t.Fatalf("PhraseWorld: %v", err)
}
if !strings.Contains(got, "рассеяния") {
t.Errorf("reply is not the workstation's: %q", got)
}
if len(spy.user) != 0 {
t.Errorf("the resident model was asked %d times, want 0", len(spy.user))
}
})
t.Run("workstation down: the gap, and nothing invented", func(t *testing.T) {
spy := newPromptSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{})
p.UseRemote(&fakeRemote{up: false})
got, err := p.PhraseWorld(context.Background(), "почему небо голубое", nil)
if !errors.Is(err, ErrNoWorldModel) {
t.Fatalf("err = %v, want ErrNoWorldModel", err)
}
if got != "" {
t.Errorf("got a reply %q with no world model", got)
}
if len(spy.user) != 0 {
t.Errorf("the resident model answered a world question %d times, want 0", len(spy.user))
}
})
t.Run("workstation errors mid-request: still the gap", func(t *testing.T) {
spy := newPromptSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{})
p.UseRemote(&fakeRemote{up: true, err: errors.New("connection refused")})
if _, err := p.PhraseWorld(context.Background(), "почему небо голубое", nil); !errors.Is(err, ErrNoWorldModel) {
t.Fatalf("err = %v, want ErrNoWorldModel", err)
}
if len(spy.user) != 0 {
t.Errorf("the resident model answered a world question %d times, want 0", len(spy.user))
}
})
}
// Prompt parity: the workstation model is asked the same question in the same
// words, or the fixtures measure one thing and the daemon ships another.
func TestPhraseWorldSendsTheSamePromptsAsPhraseQuery(t *testing.T) {
spy := newPromptSpy(t)
resident := NewLLMPhraserAt(spy.srv.URL, Config{})
if _, err := resident.PhraseQuery(context.Background(), "кто написал войну и мир", []string{"Лев Толстой"}); err != nil {
t.Fatal(err)
}
remote := &fakeRemote{up: true, reply: `{"response": "Толстой.", "mood": "neutral"}`}
offloaded := NewLLMPhraserAt(spy.srv.URL, Config{})
offloaded.UseRemote(remote)
if _, err := offloaded.PhraseWorld(context.Background(), "кто написал войну и мир", []string{"Лев Толстой"}); err != nil {
t.Fatal(err)
}
if len(remote.got) != 1 {
t.Fatalf("the workstation saw %d requests, want 1", len(remote.got))
}
if remote.got[0].System != spy.system[0] {
t.Errorf("system prompts differ:\nremote: %q\nresident: %q", remote.got[0].System, spy.system[0])
}
if remote.got[0].User != spy.user[0] {
t.Errorf("user prompts differ:\nremote: %q\nresident: %q", remote.got[0].User, spy.user[0])
}
}
// The silent half. A nudge phrased on the workstation is not news, and one
// phrased here because the card is busy is not news either — but it must be
// sampled the same way, or the workstation quietly changes how she sounds.
func TestNudgePhrasingPrefersTheWorkstationSilently(t *testing.T) {
spy := newPromptSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{LLMNudges: true})
remote := &fakeRemote{up: true, reply: `{"response": "Выпей воды.", "mood": "neutral"}`}
p.UseRemote(remote)
pn, err := p.PhraseNudge(context.Background(), loop.Candidate{Rule: loop.WaterRule(), Severity: loop.Sev1})
if err != nil {
t.Fatalf("PhraseNudge: %v", err)
}
if pn.Body != "Выпей воды." {
t.Errorf("body = %q, want the workstation's wording", pn.Body)
}
if len(remote.got) != 1 {
t.Fatalf("the workstation saw %d requests, want 1", len(remote.got))
}
if remote.got[0].Temperature != chatTemperature {
t.Errorf("temperature = %v, want %v (what the resident transport samples at)",
remote.got[0].Temperature, chatTemperature)
}
if len(spy.user) != 0 {
t.Errorf("the resident model phrased %d nudges, want 0", len(spy.user))
}
}
func TestNudgePhrasingFallsBackWhenTheCardIsBusy(t *testing.T) {
spy := newPromptSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{LLMNudges: true})
p.UseRemote(&fakeRemote{up: false})
if _, err := p.PhraseNudge(context.Background(), loop.Candidate{Rule: loop.WaterRule(), Severity: loop.Sev1}); err != nil {
t.Fatalf("PhraseNudge: %v", err)
}
if len(spy.user) != 1 {
t.Fatalf("the resident model phrased %d nudges, want 1", len(spy.user))
}
}
+1 -1
View File
@@ -39,7 +39,7 @@ import (
// Re-measured with everything else held equal, thinking off scores exactly the
// same, case for case — and a direct probe shows this llama-server build ignores
// enable_thinking / reasoning_budget for this model anyway, so there was nothing
// to turn off. Full write-up in ROUTING-EVAL-31-07-2026.md (Vikunja #376).
// to turn off. Full write-up in docs/evals/2026-07-31-routing.md (Vikunja #376).
func TestLLMRouterBaseline(t *testing.T) {
base := os.Getenv("MAVEN_LLM_URL")
if base == "" {
+3 -3
View File
@@ -1,13 +1,13 @@
// Package router is maven's reactive path — the cascade that turns a free-form
// utterance into a deterministic Decision.
//
// Spec contract (from DESIGN.md § Reactive path — routing):
// Spec contract (from docs/design.md § Reactive path — routing):
//
// - the TARGET design is LLM-as-router: the resident model (Qwen3-1.7B)
// emits GBNF-constrained structured JSON for the route, and the same
// model phrases replies; the embedder is a RAG hint, not a routing gate.
// the classifier/embedder cascade below is the committed default today,
// but it is an interim stopgap (DESIGN.md § Superseded, "classifier-owns-
// but it is an interim stopgap (docs/design.md § Superseded, "classifier-owns-
// the-route") and the known cause of weak RU query handling — not a
// design to extend.
// - a CASCADE, not one decider — layers:
@@ -35,7 +35,7 @@ package router
import "time"
// Intent — the seven save-where labels from DESIGN.md's routing table. The
// Intent — the seven save-where labels from docs/design.md's routing table. The
// discriminator is "does the loop evaluate a predicate against it?":
//
// - act: command now, not stored (function call into the allowlist)
+1 -1
View File
@@ -6,5 +6,5 @@ func KnowledgePrompt() string {
// No self-introduction here: the shared persona block already says who she
// is, and this line used to disagree with it — a different name ("Мавена")
// and a masculine noun ("ассистент") in front of a feminine persona.
return `Ответь кратко из своих знаний. Если не знаешь — скажи "не знаю". Не выдумывай. Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}.`
return `Ответь кратко из своих знаний. Если не знаешь — скажи "не знаю". Не выдумывай. Отвечай ТОЛЬКО одним объектом JSON: {"response": "...", "mood": "neutral"}.`
}
+64
View File
@@ -0,0 +1,64 @@
package router
import "strings"
// interrogatives — the question words that mark an utterance as asking rather
// than telling. Tokenized, never substring: "что" inside "чтобы" and "как"
// inside "какао" are not questions.
var interrogatives = []string{
"что", "чего", "какой", "какая", "какое", "какие", "каких",
"кто", "кого", "кому", "чей", "почему", "зачем", "отчего",
"где", "куда", "откуда", "когда", "сколько", "как",
"what", "who", "whom", "why", "when", "where", "which", "how",
}
// narrativeRequests — "tell me about X" asks for knowledge Maven does not
// hold about him. It carries no question mark and no interrogative, which is
// how "расскажи про битву при Ватерлоо" reached the fact store (#470).
var narrativeRequests = []string{
"расскажи", "объясни", "опиши", "перечисли",
"tell", "explain", "describe",
}
// captureVerbs — an explicit instruction to record something. These win over
// every test below, because "запиши что я пил воду" contains an interrogative
// and is still a capture: the word he said is "запиши".
var captureVerbs = []string{
"запиши", "запомни", "отметь", "заметь", "добавь", "сохрани",
"note", "remember", "log", "save",
}
// IsQuestionShaped reports whether text asks for something rather than
// records it. It is a deterministic offline test over tokens, so it costs
// nothing and never depends on the model that produced the routing decision.
//
// It exists because a mis-routed question used to be persisted as a fact
// about the owner, with the model's invented answer as the value (#470). The
// predicate is deliberately blunt: refusing to store a question is cheap and
// reversible, storing an invented fact about him is neither.
func IsQuestionShaped(text string) bool {
t := strings.TrimSpace(text)
if t == "" {
return false
}
toks := planTokens(t)
for _, v := range captureVerbs {
if hasTok(toks, v) {
return false
}
}
if strings.HasSuffix(t, "?") {
return true
}
for _, w := range interrogatives {
if hasTok(toks, w) {
return true
}
}
for _, w := range narrativeRequests {
if hasTok(toks, w) {
return true
}
}
return false
}
+48
View File
@@ -0,0 +1,48 @@
package router
import "testing"
func TestIsQuestionShaped(t *testing.T) {
// The seven utterances #470 recorded, plus the captures that must keep
// working. A capture misread as a question loses a fact; a question
// misread as a capture poisons recall, so the captures are the ones worth
// pinning here.
cases := []struct {
text string
want bool
}{
{"какая последняя версия языка Go?", true},
{"что дальше?", true},
{"расскажи про битву при Ватерлоо", true},
{"почему небо синее?", true},
{"какая столица Австралии?", true},
{"кто такой Никола Тесла?", true},
{"сколько стоит доллар", true},
{"who is the premier of Japan", true},
{"объясни линии Фраунгофера", true},
{"запиши что я пил воду", false},
{"запомни какая у меня машина", false},
{"отметь что я поужинал", false},
{"поужинал", false},
{"я выпил кофе", false},
{"вода", false},
{"привет", false},
{"", false},
}
for _, c := range cases {
if got := IsQuestionShaped(c.text); got != c.want {
t.Errorf("IsQuestionShaped(%q) = %v, want %v", c.text, got, c.want)
}
}
}
// Substring matching is what made the day-plan predicates wrong before, and
// this predicate gates a write, so it gets the same guard.
func TestIsQuestionShapedIsTokenized(t *testing.T) {
for _, text := range []string{"чтобы не забыть, я полил кактус", "какао выпил"} {
if IsQuestionShaped(text) {
t.Errorf("IsQuestionShaped(%q) = true; a question word inside a longer word is not a question", text)
}
}
}
+62
View File
@@ -6,6 +6,7 @@ import (
"encoding/json"
"errors"
"fmt"
"log"
"strings"
"time"
@@ -46,6 +47,39 @@ func (s *Store) WriteFact(ctx context.Context, ts time.Time, kind FactKind, key,
return id, nil
}
// FactRecallText is the text a fact is indexed under and read back as (#493).
//
// It used to be the utterance that wrote the fact, so recall of ANY
// voice-tapped fact answered with the sentence he said instead of the value
// stored: `go_version = 1.20` was indexed as "какая последняя версия языка
// Go?", and that question is what came back. The poisoned rows made the defect
// visible; the shape was wrong for legitimate facts too.
//
// The key is spoken with its underscores dropped, because a key is written for
// the store and this string is read out loud.
func FactRecallText(key, value string) string {
spoken := strings.TrimSpace(strings.ReplaceAll(key, "_", " "))
v := strings.TrimSpace(DecodeFactValue(value))
switch {
case v == "":
return spoken
case spoken == "":
return v
}
return spoken + " — " + v
}
// DecodeFactValue unwraps a stored value for reading. The column holds raw json
// when the writer serialized one (SetValue, CorrectValue) and a plain string
// when it did not (a voice tap), so a reader that wants the text handles both.
func DecodeFactValue(value string) string {
var s string
if err := json.Unmarshal([]byte(value), &s); err == nil {
return s
}
return value
}
// LatestFact returns the latest non-voided fact for key, or ErrNoFact.
// "Non-voided" = no later row has voids_id pointing at it. We resolve this by
// taking the newest row whose id is not referenced by any voids_id.
@@ -290,6 +324,20 @@ func (s *Store) CorrectValue(ctx context.Context, key, source string, value any,
if err != nil {
return 0, fmt.Errorf("last insert id: %w", err)
}
// The same repair a void needs, for the same reason (#493). A correction
// supersedes the value, and the vector still holds the old one, so recall
// kept answering with the value he had just corrected. Dropping it costs
// the key its recall vector until the fact is tapped again: this layer has
// no embedder, and a missing vector loses a question while a stale one
// answers it wrongly.
//
// Best-effort: the corrected row is committed, and a correction that lands
// beats one that fails on cleanup.
if n, derr := s.VectorMemory().DeletePrefix(ctx, "fact:"+key+":"); derr != nil {
log.Printf("store: correct %q: memory vectors survive: %v", key, derr)
} else if n > 0 {
log.Printf("store: correct %q: dropped %d superseded memory vector(s)", key, n)
}
return newID, nil
}
@@ -332,6 +380,20 @@ func (s *Store) VoidLatestFact(ctx context.Context, key, source string, ts time.
if err != nil {
return 0, 0, fmt.Errorf("void: last insert id: %w", err)
}
// The other half of the repair (#470). A fact reaches recall through a
// vector keyed `fact:<key>:<unix>`, holding the utterance that wrote it.
// Voiding the row alone left that vector answering questions, so revert
// reported success on a box that stayed broken. Deleting every vector for
// the key covers the earlier rows too: their values are superseded, and a
// superseded value has no business claiming a turn.
//
// Best-effort by design: the audit trail is already committed, and a fact
// that is voided but still recallable is better than a void that failed.
if n, derr := s.VectorMemory().DeletePrefix(ctx, "fact:"+key+":"); derr != nil {
log.Printf("store: void %q: memory vectors survive: %v", key, derr)
} else if n > 0 {
log.Printf("store: void %q: dropped %d memory vector(s)", key, n)
}
return oldID, newID, nil
}
+177
View File
@@ -0,0 +1,177 @@
package store
import (
"context"
"encoding/json"
"errors"
"fmt"
"strconv"
"strings"
"time"
)
// metaKeyFactVectorShape names the shape the stored fact vectors were written
// in. It exists so the repair below runs once per box instead of on every
// start: the rows it fixes were written by a code path that no longer exists,
// and once fixed nothing writes that shape again.
const metaKeyFactVectorShape = "fact_vector_shape"
// factVectorShapeFact is the shape FactRecallText produces. Anything else in
// the marker (including nothing, which is every box written before #493) means
// the fact vectors still hold utterances.
const factVectorShapeFact = "fact-text (#493)"
// FactVectorRepair is what one repair run did, for logging.
type FactVectorRepair struct {
Skipped bool // marker already matched — nothing to do
Rewritten int // rows re-embedded from the fact they name
Dropped int // rows deleted: voided, superseded, or naming no fact at all
Kept int // rows already holding the right text
Took time.Duration
}
// RepairFactVectors brings the fact rows of memory_vectors in line with the
// facts they name, and is the operator recovery a poisoned box had no path to
// (#470 point 4, #493).
//
// Three defects put wrong text in that index, and all three are write-path
// fixes that do nothing for rows already stored:
//
// - the indexed text was the utterance, so every fact row reads back a
// sentence rather than a value;
// - a void left its vector behind, so retracted junk kept answering;
// - a correction left its vector behind, so the superseded value did.
//
// So each fact row is resolved against the fact store and one of three things
// happens. It is dropped when the key has no fact, when the newest row for the
// key is a void marker, or when a newer vector for the same key exists — a
// superseded value has no business claiming a turn. It is re-embedded when its
// text is not what FactRecallText says the fact is. Otherwise it is left alone.
//
// Idempotent, and safe to interrupt: every step compares before writing and the
// marker is written last, so a run that dies partway is simply redone.
func (s *Store) RepairFactVectors(ctx context.Context, embed EmbedFunc) (FactVectorRepair, error) {
start := time.Now()
var res FactVectorRepair
shape, err := s.Meta(ctx, metaKeyFactVectorShape)
if err != nil {
return res, err
}
if shape == factVectorShapeFact {
res.Skipped = true
res.Took = time.Since(start)
return res, nil
}
rows, err := s.db.QueryContext(ctx, `SELECT id, meta FROM memory_vectors`)
if err != nil {
return res, fmt.Errorf("repair fact vectors: read: %w", err)
}
type factVec struct {
id, key string
meta map[string]string
ts int64
}
var vecs []factVec
newest := map[string]int64{} // key → newest ts seen for it
for rows.Next() {
var id, metaJSON string
if err := rows.Scan(&id, &metaJSON); err != nil {
rows.Close()
return res, fmt.Errorf("repair fact vectors: row: %w", err)
}
meta := map[string]string{}
if err := json.Unmarshal([]byte(metaJSON), &meta); err != nil {
rows.Close()
return res, fmt.Errorf("repair fact vectors: meta for %q: %w", id, err)
}
if meta["type"] != "fact" {
continue
}
key, ts, ok := splitFactVectorID(id)
if !ok {
continue
}
vecs = append(vecs, factVec{id: id, key: key, meta: meta, ts: ts})
if ts > newest[key] {
newest[key] = ts
}
}
rows.Close()
if err := rows.Err(); err != nil {
return res, fmt.Errorf("repair fact vectors: rows: %w", err)
}
for _, v := range vecs {
drop := v.ts < newest[v.key]
var want string
if !drop {
f, ferr := s.LatestFact(ctx, v.key)
switch {
case errors.Is(ferr, ErrNoFact):
drop = true
case ferr != nil:
return res, fmt.Errorf("repair fact vectors: fact %q: %w", v.key, ferr)
case DecodeFactValue(f.Value) == "voided":
drop = true
default:
want = FactRecallText(v.key, f.Value)
}
}
if drop {
if err := s.VectorMemory().Delete(ctx, v.id); err != nil {
return res, err
}
res.Dropped++
continue
}
if v.meta["text"] == want {
res.Kept++
continue
}
vec, err := embed(ctx, want)
if err != nil {
return res, fmt.Errorf("repair fact vectors: embed %q: %w", v.id, err)
}
// The whole meta blob is rewritten in Go rather than patched in SQL,
// because json_set needs the JSON1 extension and this store is opened
// through sqlcipher.
v.meta["text"] = want
metaJSON, err := json.Marshal(v.meta)
if err != nil {
return res, fmt.Errorf("repair fact vectors: meta %q: %w", v.id, err)
}
if _, err := s.db.ExecContext(ctx,
`UPDATE memory_vectors SET vec = ?, meta = ? WHERE id = ?`,
encodeVec(vec), string(metaJSON), v.id); err != nil {
return res, fmt.Errorf("repair fact vectors: write %q: %w", v.id, err)
}
res.Rewritten++
}
if err := s.SetMeta(ctx, metaKeyFactVectorShape, factVectorShapeFact); err != nil {
return res, err
}
res.Took = time.Since(start)
return res, nil
}
// splitFactVectorID reads the key and write time back out of a fact vector's
// id, which the write path builds as `fact:<key>:<unix>`. A key may hold a
// colon, the timestamp may not, so the split is from the right.
func splitFactVectorID(id string) (key string, ts int64, ok bool) {
rest, found := strings.CutPrefix(id, "fact:")
if !found {
return "", 0, false
}
cut := strings.LastIndex(rest, ":")
if cut <= 0 {
return "", 0, false
}
ts, err := strconv.ParseInt(rest[cut+1:], 10, 64)
if err != nil {
return "", 0, false
}
return rest[:cut], ts, true
}
+151
View File
@@ -0,0 +1,151 @@
package store
import (
"context"
"database/sql"
"testing"
"time"
)
// The write-path half of #493: recall of a fact must read back the fact, not
// the sentence he happened to say.
func TestFactRecallText(t *testing.T) {
for _, tc := range []struct {
name, key, value, want string
}{
{"json value", "go_version", `"1.20"`, "go version — 1.20"},
{"plain value", "water", "выпил", "water — выпил"},
{"no value", "shower", "", "shower"},
{"underscores are spoken as spaces", "espresso_machine", `"чистая"`, "espresso machine — чистая"},
} {
t.Run(tc.name, func(t *testing.T) {
if got := FactRecallText(tc.key, tc.value); got != tc.want {
t.Fatalf("FactRecallText(%q, %q) = %q; want %q", tc.key, tc.value, got, tc.want)
}
})
}
}
// A correction left the superseded value in the index, so recall answered with
// the value he had just corrected (#493).
func TestCorrectValueDropsMemoryVectors(t *testing.T) {
ctx := context.Background()
s := newTestStore(t)
now := time.Now()
mem := s.VectorMemory()
if _, err := s.WriteFact(ctx, now, KindSelf, "go_version", `"1.20"`, "tap:voice", 1.0, sql.NullInt64{}); err != nil {
t.Fatalf("WriteFact: %v", err)
}
if err := mem.Insert(ctx, "fact:go_version:1", []float32{1, 0, 0}, map[string]string{
"type": "fact", "text": "go version — 1.20",
}); err != nil {
t.Fatalf("Insert: %v", err)
}
if _, err := s.CorrectValue(ctx, "go_version", "feedback", "1.25", now.Add(time.Minute)); err != nil {
t.Fatalf("CorrectValue: %v", err)
}
got, err := mem.ByPrefix(ctx, "fact:")
if err != nil {
t.Fatalf("ByPrefix: %v", err)
}
if len(got) != 0 {
t.Fatalf("after the correction the index still holds %+v; the superseded value must not answer", got)
}
}
// The recovery path a poisoned box had none of (#470 point 4, #493): rows
// written before the fix hold utterances, voided junk and superseded values,
// and no write-path change reaches any of them.
func TestRepairFactVectors(t *testing.T) {
ctx := context.Background()
s := newTestStore(t)
now := time.Now()
mem := s.VectorMemory()
embed := func(ctx context.Context, text string) ([]float32, error) {
return []float32{float32(len(text)), 1, 0}, nil
}
// A live fact indexed under the question that wrote it — the defect.
if _, err := s.WriteFact(ctx, now, KindSelf, "water", `"выпил"`, "tap:voice", 1.0, sql.NullInt64{}); err != nil {
t.Fatalf("WriteFact water: %v", err)
}
if err := mem.Insert(ctx, "fact:water:100", []float32{9, 9, 9}, map[string]string{
"type": "fact", "source": "voice", "text": "запиши что я пил воду",
}); err != nil {
t.Fatalf("Insert water: %v", err)
}
// A voided fact whose vector survived the void.
if _, err := s.WriteFact(ctx, now, KindSelf, "go_version", `"1.20"`, "tap:voice", 1.0, sql.NullInt64{}); err != nil {
t.Fatalf("WriteFact go_version: %v", err)
}
if _, _, err := s.VoidLatestFact(ctx, "go_version", "feedback", now.Add(time.Minute)); err != nil {
t.Fatalf("VoidLatestFact: %v", err)
}
if err := mem.Insert(ctx, "fact:go_version:100", []float32{9, 9, 9}, map[string]string{
"type": "fact", "text": "какая последняя версия языка Go?",
}); err != nil {
t.Fatalf("Insert go_version: %v", err)
}
// A key with two vectors: only the newest may answer.
if _, err := s.WriteFact(ctx, now, KindSelf, "mood", `"устал"`, "tap:voice", 1.0, sql.NullInt64{}); err != nil {
t.Fatalf("WriteFact mood: %v", err)
}
for _, ts := range []string{"100", "200"} {
if err := mem.Insert(ctx, "fact:mood:"+ts, []float32{9, 9, 9}, map[string]string{
"type": "fact", "text": "мне грустно",
}); err != nil {
t.Fatalf("Insert mood %s: %v", ts, err)
}
}
// A note must be left entirely alone.
if err := mem.Insert(ctx, "note:7", []float32{5, 5, 5}, map[string]string{
"type": "note", "text": "сеть тормозит по вечерам",
}); err != nil {
t.Fatalf("Insert note: %v", err)
}
res, err := s.RepairFactVectors(ctx, embed)
if err != nil {
t.Fatalf("RepairFactVectors: %v", err)
}
if res.Rewritten != 2 || res.Dropped != 2 {
t.Fatalf("repair reported %+v; want 2 rewritten (water, newest mood) and 2 dropped (voided go_version, superseded mood)", res)
}
got, err := mem.ByPrefix(ctx, "fact:")
if err != nil {
t.Fatalf("ByPrefix: %v", err)
}
texts := map[string]string{}
for _, r := range got {
texts[r.ID] = r.Meta["text"]
}
if len(texts) != 2 {
t.Fatalf("the index holds %+v; want only fact:water:100 and fact:mood:200", texts)
}
if texts["fact:water:100"] != "water — выпил" {
t.Fatalf("water reads back %q; want the fact, not the utterance", texts["fact:water:100"])
}
if texts["fact:mood:200"] != "mood — устал" {
t.Fatalf("mood reads back %q", texts["fact:mood:200"])
}
// Provenance the row already carried must survive the rewrite.
for _, r := range got {
if r.ID == "fact:water:100" && r.Meta["source"] != "voice" {
t.Fatalf("water lost its source meta: %+v", r.Meta)
}
}
if notes, err := mem.ByPrefix(ctx, "note:"); err != nil || len(notes) != 1 {
t.Fatalf("the note row was touched: %+v (err %v)", notes, err)
}
// Marker written, so a second run is free and changes nothing.
again, err := s.RepairFactVectors(ctx, embed)
if err != nil {
t.Fatalf("second RepairFactVectors: %v", err)
}
if !again.Skipped {
t.Fatalf("second run did work: %+v; the marker must make it a no-op", again)
}
}
+21
View File
@@ -148,6 +148,27 @@ func (m *MemoryStore) Delete(ctx context.Context, id string) error {
return nil
}
// DeletePrefix removes every vector whose id starts with prefix and returns
// how many went. Same escaping as ByPrefix, so a key containing % or _ cannot
// widen the delete.
//
// It exists for the repair half of a revert (#470). Voiding a fact row left
// its vector in the index, so recall kept serving the voided fact's utterance
// and the documented repair did not repair.
func (m *MemoryStore) DeletePrefix(ctx context.Context, prefix string) (int64, error) {
pattern := escapeLike(prefix) + "%"
res, err := m.db.ExecContext(ctx,
`DELETE FROM memory_vectors WHERE id LIKE ? ESCAPE '\'`, pattern)
if err != nil {
return 0, fmt.Errorf("memory: delete prefix %q: %w", prefix, err)
}
n, err := res.RowsAffected()
if err != nil {
return 0, fmt.Errorf("memory: delete prefix %q: rows affected: %w", prefix, err)
}
return n, nil
}
// escapeLike neutralises the LIKE wildcards in a literal prefix.
func escapeLike(s string) string {
r := strings.NewReplacer(`\`, `\\`, `%`, `\%`, `_`, `\_`)
+64
View File
@@ -0,0 +1,64 @@
package store
import (
"context"
"database/sql"
"testing"
"time"
)
// Stage 3 of #470: reverting a fact reported success and left the vector that
// was answering questions, so the documented repair did not repair.
func TestVoidLatestFactDropsMemoryVectors(t *testing.T) {
ctx := context.Background()
s := newTestStore(t)
now := time.Now()
mem := s.VectorMemory()
if _, err := s.WriteFact(ctx, now, KindSelf, "go_version", `"1.20"`, "tap:voice", 1.0, sql.NullInt64{}); err != nil {
t.Fatalf("WriteFact: %v", err)
}
// The id shape actionFact writes: fact:<key>:<unix>.
if err := mem.Insert(ctx, "fact:go_version:1", []float32{1, 0, 0}, map[string]string{
"type": "fact", "text": "какая последняя версия языка Go?",
}); err != nil {
t.Fatalf("Insert: %v", err)
}
// A vector for another key must survive the void.
if err := mem.Insert(ctx, "fact:water:1", []float32{0, 1, 0}, map[string]string{
"type": "fact", "text": "запиши что я пил воду",
}); err != nil {
t.Fatalf("Insert: %v", err)
}
if _, _, err := s.VoidLatestFact(ctx, "go_version", "feedback", now.Add(time.Minute)); err != nil {
t.Fatalf("VoidLatestFact: %v", err)
}
got, err := mem.ByPrefix(ctx, "fact:")
if err != nil {
t.Fatalf("ByPrefix: %v", err)
}
if len(got) != 1 || got[0].ID != "fact:water:1" {
t.Fatalf("after the void the index holds %+v; want only fact:water:1", got)
}
}
func TestDeletePrefixDoesNotWidenOnWildcards(t *testing.T) {
ctx := context.Background()
s := newTestStore(t)
mem := s.VectorMemory()
for _, id := range []string{"fact:a_b:1", "fact:axb:1"} {
if err := mem.Insert(ctx, id, []float32{1, 0}, map[string]string{"type": "fact"}); err != nil {
t.Fatalf("Insert %q: %v", id, err)
}
}
n, err := mem.DeletePrefix(ctx, "fact:a_b:")
if err != nil {
t.Fatalf("DeletePrefix: %v", err)
}
if n != 1 {
t.Fatalf("deleted %d rows; the _ in the key must not match x", n)
}
}
+2 -2
View File
@@ -18,9 +18,9 @@
// returns a canned string the router + action path operate on); with a
// worker socket configured, it wires Remote.
//
// Per DESIGN.md § Voice pipeline (STT / TTS): whisper.cpp (CGo, Vulkan) in
// Per docs/design.md § Voice pipeline (STT / TTS): whisper.cpp (CGo, Vulkan) in
// cmd/mavsttd is the production stt — the older faster-whisper/vosk picks are
// retired (DESIGN.md § Superseded, "named STT/TTS model picks"). The
// retired (docs/design.md § Superseded, "named STT/TTS model picks"). The
// server-side stt module is the heavy multilingual path; the client's
// wake-word + stage-0 command grammar (cmd/mavwaked) hits the router directly
// and never crosses this seam. Today's Remote + Stub both return plain text
+2 -2
View File
@@ -2,7 +2,7 @@
// store's allowlist, and drafts 'proposed' scaffolds for acts that aren't on
// it yet.
//
// Boundary discipline (DESIGN.md § "Tool registration — drafting is suggest,
// Boundary discipline (docs/design.md § "Tool registration — drafting is suggest,
// enabling is act"):
//
// - The store is the allowlist. Only status='enabled' rows run. A verb not
@@ -26,7 +26,7 @@
// - Destructive tools don't run on first hearing: Exec returns ErrNeedsConfirm
// and the handler runs a confirm turn ("выполнить X? да/нет"); only a
// confirmed re-Exec runs them. A gate assumes a fully-formed action, which
// an enabled+matched act is (DESIGN.md § "Confirmation is not one
// an enabled+matched act is (docs/design.md § "Confirmation is not one
// mechanism").
package tool
+2 -2
View File
@@ -5,9 +5,9 @@
// PCM, headerless per the audio package; the voice sink + reference client
// wrap it in a WAV at the disk edge.
//
// Per DESIGN.md § Voice pipeline (STT / TTS): piper is the production tts
// Per docs/design.md § Voice pipeline (STT / TTS): piper is the production tts
// (subprocess + espeak-ng, CPU-only on the ryzen box, driven by cmd/mavttsd);
// the older silero pick is retired (DESIGN.md § Superseded, "named STT/TTS
// the older silero pick is retired (docs/design.md § Superseded, "named STT/TTS
// model picks"). A different voice is a model-file swap, not a code change.
// The daemon wires one impl — Remote pointing at the worker socket if
// configured, Stub otherwise.
+14 -2
View File
@@ -24,6 +24,8 @@ import (
"net"
"sync"
"time"
"github.com/kami/maven/internal/netaddr"
)
// Client — one connection to one worker module. NOT goroutine-safe for
@@ -38,15 +40,25 @@ type Client struct {
dial func() (net.Conn, error)
}
// Dial opens a Client to the worker socket at path. The first call lazily
// Dial opens a Client to the worker module at path. The first call lazily
// dials; subsequent calls reuse the conn (a fresh dial happens on next call
// after a teardown). Lazy dial keeps a worker that's restarting from
// blocking core's startup; core attempts the dial on first use.
//
// path is a netaddr seam address. A bare path is the unix socket it has
// always been; "tcp://workstation:9310?token=..." reaches a module on another
// host, which is how stt and tts move to the machine with the GPU and the
// microphone. A bad address surfaces on the first call, not here, because
// Dial does not fail — see internal/netaddr.
func Dial(path string) *Client {
addr, err := netaddr.Parse(path)
return &Client{
path: path,
dial: func() (net.Conn, error) {
return net.Dial("unix", path)
if err != nil {
return nil, err
}
return netaddr.Dial(addr)
},
}
}
+22 -34
View File
@@ -7,10 +7,11 @@
// which module to dial; mixing the two is a config error caught cleanly by
// the wire, not a runtime goroutine panic). One Server per module process.
//
// Socket perms mirror ipc.Server: dir 0700, socket 0600 ⇒ same unix user.
// The module has no key, so the floor is "same user"; the wg/mTLS layers
// are out of scope here (this socket never crosses the network radius —
// it's local-only, point-to-point between two processes on the box).
// The seam address decides the transport. On the default unix socket the
// perms mirror ipc.Server — dir 0700, socket 0600 ⇒ same unix user — and that
// is the whole auth floor, because the seam never leaves the box. A tcp
// address moves the module to another host and takes that floor away, so
// netaddr checks a shared token before the first frame. See internal/netaddr.
package worker
import (
@@ -18,11 +19,10 @@ import (
"encoding/json"
"fmt"
"net"
"os"
"sync"
"sync/atomic"
"golang.org/x/sys/unix"
"github.com/kami/maven/internal/netaddr"
)
// Server — a worker module process's listener. Wires either a Transcriber,
@@ -34,6 +34,7 @@ type Server struct {
s Synthesizer
path string
addr netaddr.Addr
ln net.Listener
wg sync.WaitGroup
@@ -61,25 +62,24 @@ func NewSynthesizerServer(path string, s Synthesizer) *Server {
// two separate processes per the restart-free / fail-independent invariant).
func (srv *Server) SetSynthesizer(s Synthesizer) { srv.s = s }
// Listen binds the unix socket with 0700 dir + 0600 socket perms (same floor
// as internal/ipc). A stale socket at path is removed first so the worker
// process restarts cleanly after a crash, no manual cleanup needed.
// Listen binds the seam address the Server was built with.
//
// A bare path is a unix socket with 0700 dir + 0600 socket perms, the same
// floor as internal/ipc, and a stale socket is removed first so the worker
// process restarts cleanly after a crash. A "tcp://host:port?token=..."
// address binds a network listener instead, so this module can run on the
// workstation while core stays on homesrv; the token is mandatory there,
// because there is no filesystem to be the auth floor. See internal/netaddr.
func (srv *Server) Listen() error {
_ = os.Remove(srv.path)
if err := os.MkdirAll(parentDir(srv.path), 0o700); err != nil {
return fmt.Errorf("worker: mkdir socket dir: %w", err)
}
oldMask := unix.Umask(0o077)
ln, err := net.Listen("unix", srv.path)
unix.Umask(oldMask)
addr, err := netaddr.Parse(srv.path)
if err != nil {
return fmt.Errorf("worker: listen %s: %w", srv.path, err)
return err
}
if err := os.Chmod(srv.path, 0o600); err != nil {
_ = ln.Close()
_ = os.Remove(srv.path)
return fmt.Errorf("worker: chmod socket: %w", err)
ln, err := netaddr.Listen(addr)
if err != nil {
return err
}
srv.addr = addr
srv.ln = ln
return nil
}
@@ -193,7 +193,7 @@ func (srv *Server) Close() error {
}
err := srv.ln.Close()
srv.wg.Wait()
_ = os.Remove(srv.path)
netaddr.Cleanup(srv.addr)
return err
}
@@ -214,15 +214,3 @@ func marshalResult(v any) json.RawMessage {
b, _ := json.Marshal(v)
return b
}
func parentDir(p string) string {
for i := len(p) - 1; i >= 0; i-- {
if p[i] == '/' {
if i == 0 {
return "/"
}
return p[:i]
}
}
return "."
}