CLAUDE.md was 805 lines and it is loaded into every session, so every line costs. The routing section alone was 412 of them, and it was a chronological log of every measurement since 2026-07-31: four re-measurements of the same fixture, the history of each of the four routing heads, and the reasoning behind every grammar. None of that is a rule. An agent about to edit the router needs to know that the classifier is the floor, that queryWalk only takes sources out, and that heads_path must never point at model_path. It does not need the seed spread of the third head to read the file at all. So docs/routing.md is a living doc under the tier convention, and it carries the reasoning and the numbers. CLAUDE.md keeps the constraints and points at it. 805 lines to 490, with the routing section at 60. The same cut is applied to the header block and to the world chain under non-goals: the current fact and the eval filename stay, the "measured on date D it went from A to B" narrative moves out or is dropped. Nothing was deleted without checking. Every backticked literal in the old file was diffed against the two new ones, and the forty that fell out were reviewed one by one. Nine were facts rather than narrative and are restored: the ecosystem default URLs, the voice.llm_router flag and pickLLMRouter, the four head eval filenames, handlePraxisAct, SourceAccuracy, and the rule that calendar-query names the calendar where the possessive agenda rules do not. A closing section states the file's own contract, so the next agent adds a measurement to docs/evals/ instead of a paragraph here. diff-budget.sh blocked on 1544 changed lines. It counts markdown, which the repo's own pre-commit hook exempts, and this commit touches nothing else.
25 KiB
CLAUDE.md
Guidance for Claude Code (claude.ai/code) working in this repository.
This file is loaded into every session, so it carries rules and not history. A
measurement lives in docs/evals/<date>-<name>.md and is never edited after the
day. A subsystem's reasoning lives in a living doc under docs/. When a line
here says "see X", read X before changing that subsystem.
| Read this | Before |
|---|---|
docs/routing.md |
touching internal/router/ or queryWalk |
docs/offload.md |
touching a daemon seam or adding a model caller |
docs/ecosystem.md |
touching Nexus, Praxis or Hexis |
docs/rearchitecture.md, docs/design.md |
changing the shape of anything |
AGENTS.md |
local preview, screenshots, model downloads |
What Maven is
A self-hosted, privacy-first voice assistant in Russian and English. Go daemons talk over unix sockets. One resident small model routes and phrases. whisper.cpp does speech-to-text and piper does text-to-speech.
The deploy target is a Ryzen laptop (homesrv) with Vulkan offload to the Vega
iGPU (n_gpu_layers: 99, compose passes /dev/dri and the render gid). The
resident model stays at 1.7B or under either way.
The resident model
Qwen3-1.7B (UD-Q4_K_XL), stock, not yet the CPT'd one. It is a Thinking
variant, so n_ctx is 4096. Reasoning tokens need the room, and 4096 is what
every score was measured at.
The target is the locally CPT'd Qwen3-1.7B (V-122, training in flight). Stock
already speaks good Russian. What it gets wrong is the persona. It writes я рад
where Maven needs рада.
Do not bother with sub-500M models. LFM2.5-230M and 350M were measured on
2026-07-31 and both are unusable in Russian
(docs/evals/2026-07-31-model-bakeoff.md). Their published IFEval and BFCL
numbers are English-only.
Model files live in /mnt/hdd1/llms, bind-mounted to /opt/maven/models/llm.
That shadows the repo's models/llm/, so a gguf sitting there is not loaded
by anything. Swapping the resident model is a one-line change to
phraser.model_path in deploy/mavend.json.
The workstation
Model work moved to workpc on 2026-08-02 (owner's call). homesrv cannot grow a GPU and workpc has 16GB of VRAM. So the resident model, speech-to-text and text-to-speech are preferred remotes with a floor on homesrv.
Three rules:
- The workstation is never assumed up.
- Fall back silently when it would only do the job better.
- Name the gap when the resident model cannot do the job at all. A world
question goes through
LLMPhraser.PhraseWorldand returnsworldGap(cmd/mavend/worldmodel.go) rather than an invented answer. A box with noworkstationblock behaves exactly as it did before the seam.
Routing and replies prefer the workstation through modelSeam. Nudge and
reminder phrasing prefer it inside the phraser. docs/offload.md says which
caller is which.
The embedder stays on homesrv permanently, because it backs that floor. It is
multilingual-e5-small, quantized and asymmetric. EmbedQuery and EmbedPassage
apply the query: and passage: prefixes it was trained with. Calling plain
Embed on a note is a bug. See docs/evals/2026-08-04-recall-e5-small.md.
Speech-to-text
sttSeam in cmd/mavend/voicewire.go builds an stt.Pair beside modelSeam.
It prefers CrisperWhisper 2.0 turbo on workpc with mavsttd as the floor. It takes
only the silent half of the rule, because a worse transcript is still a turn. So
stt.Pair has no TranscribeRemote and the fallback is never spoken.
CW2 turbo scores 10.4% WER in Russian against 27.5% for the ggml-small.bin
mavsttd loads, over 200 Golos clips
(docs/evals/2026-08-09-crisperwhisper2-russian-wer.md).
whisper.cpp cannot load CW2 at all. It reads its language count off the
vocabulary size. CW2's 51897 tokens shift seven special token ids. So CW2 is its
own transformers service on port 8081 (deploy/cw2/serve.py).
stt.HTTPTranscriber posts raw PCM to it with a bearer token, because audio is
the most sensitive thing that crosses this seam. The switch is workstation.stt
in deploy/mavend.json, and deleting the block sends every utterance to mavsttd.
mavgpud runs that service as a second child. This is not an optimisation.
CW2 is a ROCm process on the same card, so it registers on the KFD like any
contender. Under its own systemd unit it made mavgpud evict llama-server every
few seconds. That took the model arm down for eight minutes on 2026-08-09. The
card needs one owner. Any GPU service added beside mavgpud goes in
cmd/mavgpud, never in systemd. CW2 is on the yield clock and not the idle
one. At 1.6GB it denies the card to nobody.
Text-to-speech has not moved. piper on homesrv is the only synthesizer.
Build and test
CGO daemons (mavend, mavsttd, mavttsd, mavenclient) need the vendored
toolchain and libs wired through the Makefile. Do not call go build on them
bare, use make.
make build # all 11 binaries
make build-web # one daemon (web/waked/poll/caldav build without CGO)
make test # go test -race across ./internal/... ./cmd/... with CGO env set
make t PKG=./internal/router/
make t PKG=./cmd/mavend/ RUN=TestSimulator
make t PKG=./internal/router/eval/ RUN='TestONNX' V=1 # V=1 for -v, RACE=0 to drop -race
Do not hand-write the CGO preamble. Past sessions pasted it about 390 times,
and that is where the shell-quoting failures came from. This box runs zsh, so an
unquoted -run Test* dies on "no matches found" before go is ever reached.
make t carries -race, so a green make t cannot turn red under make test.
It carries -count=1, so a cached PASS from before your edit is never mistaken
for a result. It sets MAVEN_ONNX_LIB, which the hand-written recipe did not.
The four TestONNX* measurements self-skip when that variable is unset and the
run still prints ok. So every targeted eval done the old way reported the hash
ratchet while reading as a real embedder score.
The daemons (cmd/)
| Binary | Role |
|---|---|
mavend |
Core. Router, phraser, memory, reminders, digestion tick. Owns the DB and IPC socket. |
mavweb |
HTTP UI and PWA (/dash, /history, /trace, /notifications, /tools), WebAuthn auth. |
mavsttd |
Speech-to-text (whisper.cpp, CGO). |
mavttsd |
Text-to-speech (piper subprocess). |
mavwaked |
Wake-word and VAD gate. Not deployed anywhere yet. |
mavenclient |
Voice loop client (mic, stt, core, tts). Not deployed anywhere yet. |
mavpoll |
Environment poller: netdata alarms, uptime-kuma, zenmoney, wireguard presence. Writes facts, sends nothing. Telegram is internal/delivery/telegramsink. |
mavcaldav |
CalDAV calendar sync. |
mavmaild |
Mail reader (IMAP, read-only). Holds the IMAP password, core never sees it. |
mavgpud |
GPU supervisor. Runs on workpc, own unit deploy/mavgpud.service. Keeps llama-server loaded while the card is free (V-488). Maven never asks it for anything and reads /health through llm.Pair. |
mavupdate |
Not a daemon. Operator CLI a human runs on the box to deploy a new build. |
Two binaries have no Makefile target and neither is deployed. mavseal encrypts
a live tmpfs working copy back to the ciphertext file when mavend was killed
before defer st.Close() sealed it. labelgen runs the stage 0 grammars over
utterances and prints JSONL, the training data for the routing heads.
Daemons are wired socket-to-socket, not linked. internal/ipc is the wire
protocol. deploy/mavend.json sets socket paths, model paths and the phraser and
embedder blocks, with ${VAR} expansion from gitignored deploy/telegram.env.
Who is in compose, and who is not
docker-compose.yml runs five: mavend, mavsttd, mavttsd, mavweb,
mavpoll. Count against compose, not against the table above. Four daemons are
absent and each absence has a different reason.
mavmaild and mavcaldav are commented out, each with the reason beside it. The
first needs a mail account and the second a CalDAV account, and this box has
neither. Two things ride on the CalDAV absence (V-644). Agenda questions route to
IntentQuery at stage 0, and the calendar query source then reads a table
nobody writes. And loop.State.CalendarBusy is fed by the same facts, so the
gate's "do not nag mid-meeting" is permanently false.
mavwaked and mavenclient are absent by decision (V-463,
docs/plans/17-where-the-voice-loop-runs.md). homesrv has a microphone, because
it is a laptop, but it is in the wrong room. They belong on a client machine
where the owner is standing, and that machine is workpc. ipc.Dial already takes
tcp://host:port?token=..., so V-515 is a deployment and not a build.
Until then the wake word and the VAD gate are covered by unit tests and nothing
else. Push-to-talk through /dash is what QA covers. Deploying them does not by
itself prove a wake word. mavwaked has no keyword model (V-487 stage two), so
the loop runs open until that lands.
Passwords are read from files, never taken as flag values. mavcaldav uses
-pass-file and -render-pass-file. mavpoll and mavmaild follow the same
rule.
The ecosystem: Nexus, Praxis, Hexis
Maven is one of four services. It owns conversation and personal memory. It does
not own identity, operational state, or execution. Full contract in
docs/ecosystem.md.
Nexus identifies. Praxis observes. Hexis acts. Maven understands and coordinates.
| Service | Owns | Maven's client | Configured at |
|---|---|---|---|
| Nexus | Canonical entity ids, names, aliases, relationships. Projects, services, devices, people, pets, places. | nexusClient in cmd/mavend/ecosystem.go, POST /api/v1/resolve |
nexus.url (http://nexus:9740) |
| Praxis | Operational attention and item lifecycle. What needs looking at, what changed, what is unresolved. | praxisClient, the HTTP tools API under /api/v1/tools/ |
praxis.url (http://praxis:8989) |
| Hexis | The capability registry and the only path to executing anything. | vendored github.com/kami/hexis/pkg/client |
hexis.url (http://hexis:9741) |
All three are nil unless configured and every one degrades on its own. An
outage means a named gap in the answer, never a broken turn and never a guess.
Rules that are not negotiable:
- No component reads another component's database. Praxis attention comes over HTTP, never from its SQLite file.
- Identity lives in Nexus. Do not invent a local fact key for something Nexus
resolves.
actionFactsetsSubject, andcmd/mavend/factenrichment.goresolves it in the background. - Free text never reaches a mutating Hexis call. Resolve to a canonical entity id first. Ambiguous resolution asks the owner, it does not pick.
- LLM output is not authorization. Confirmation binds capability id, target
entity, arguments, requester and expiry. See
cmd/mavend/confirm.go. - Praxis lifecycle words mean different things. Surfaced is not acknowledged,
acknowledged is not resolved, execution success is not recovery. Reading an
item aloud calls
Surface, neverAcknowledge. - No automatic attention-to-action path. Digestion may summarise Praxis. It may not call Hexis.
Every cross-service call carries a correlation id minted once per action
(withCorrelationID), a contract version header, and X-Requested-By: maven.
Routing
Read docs/routing.md before touching internal/router/ or queryWalk. It
carries the stage-by-stage reasoning, every measurement, and why each rule
exists. What follows is only what must not be broken.
A route produces two decisions. Intent is one of seven values. Source is
where the answer lives and is read on IntentQuery alone. They are scored
separately, because one number hides which one moved.
The cascade is stage 0 grammars, then the routing heads, then the resident model, then the classifier. Every stage may decline and the next one answers.
- The classifier is the floor, not dead code. It runs when the resident model is off. It runs when there is no llama-server, and on any error. Deleting it makes a model outage a broken turn.
- Any model error falls through, so a turn never breaks on a model.
baselineGrammarsineval_test.gomirrorsbuildRouter. Add a grammar to one and it belongs in both, or the fixture scores a set nobody runs.- Go's
\bis ASCII-only and never fires after a Cyrillic letter. A Russian pattern needs an explicit(\s|[?!.]|$). PraxisGrammars()is the only path to Praxis, not a faster one. The model reaches Praxis 0/12 alone, because nothing in the router prompt names a Praxis capability.voice.embedder.heads_pathmust never point atmodel_path. The resident e5-small must not be replaced by the fine-tuned copy. Recall depends on that file scoring what it scored. Fine-tune a copy of the weights.- Bump
tokenizerRevon any change to whatencodeWordemits. The embedder id carries the revision. So a tokenizer fix triggersReembedAllthe way swapping the model file does. - A new rung in the
runTurnladder needs its name inpreRouteLadder(cmd/mavend/decisiontrace.go). Otherwise that rung is silently missing from the decision record. - Routing traces are retained 14 days, enforced on write and again on start.
The utterance is stored in clear and nothing reads it outward.
Store.Wipedeletes it with everything else.
queryWalk and the destination
queryWalk in cmd/mavend/actions_query.go takes sources out and moves
none. That is the safety argument and it is not negotiable. The table's order is
load-bearing, and above all it carries "the owner's data first, then the world".
SourceUnknown is a real value and it is the floor. Nothing named a destination,
so the daemon walks the whole chain. Naming SourceWorld does not send the turn
outside on its own.
What comes out is only the sources marked guesses: true. Those decide a turn is
theirs by cosine against frozen seeds, then answer whatever they claimed. A
source that looks rather than guesses is always asked.
The personal boundary is the one exception and it is deliberate. It guesses,
so naming SourceWorld drops it. Only a stage 0 grammar may drop it
(owner's call, V-666). Decision.SourceAnchored carries the provenance, and
queryWalk reads it for the source marked boundary: true and no other. So
every other guesser still comes off the turn, whoever named the destination.
TestOnlyAGrammarMayDropTheBoundary and TestNamingRecallKeepsTheBoundary pin
both directions.
Current numbers
| Arm | Intent | Destination | p50 |
|---|---|---|---|
| classifier + ONNX | 76.0% | 36.4% | 16.6µs |
| resident Qwen3-1.7B, cascade | 80.2% | not measured | 1.19s |
| routing heads, cascade | 96.9% | 75.8% | 27.9ms |
| workstation gemma-4-E4B, cascade | 89.6% | 57.6% | 294ms |
Judge a routing change against the classifier and the resident model, since those are what always answer. The fixture has grown from 77 cases to 96, so a number is comparable only to another number on the same fixture.
LLM output contract
All phrasing paths emit {"response":"...","mood":"..."}, falling back to plain
text when the model skips the JSON. One parser, parseResponseMood in
internal/phraser/parse.go, and every path reaches it: the six LLMPhraser
methods, PhraseWorld, and Replier.PhraseReply. cmd/mavend/replier_llm.go
wraps the last of those, holds the stub fallback, and does no parsing of its own.
Mood is a fixed enum.
The router prompt is a separate contract:
[{"intent":<enum>, key?, value?, text?, verb?}, ...] over 7 intents (fact, reminder, note, query, act, chat, system). llm/check_prompt_parity.py in the
training workspace enforces that the Go and relabelling prompts stay identical.
Russian patterns: three mechanisms, no fourth
Hand-written Russian stem patterns were swept out on 2026-08-04 (owner's call). A regex whose output is a fact or a route is the defect. A regex over structured input, such as HTML, MIME, JSON, a URL or an argv list, is not. Before writing a Russian word list, pick one of these:
internal/lexiconfor closed classes, inlexicon_ru_v1.json. Interrogatives, capture verbs, reminder verbs, cardinals, day offsets, parts of day, weekdays, months, spoken hours. Editing a word is a data change and there is exactly one copy. Months used to live in three files. Cardinals carry the oblique forms, because a spoken time declines andв семьandк семиare one hour.internal/morphfor grammar, from the vendored golem Russian dictionary.IsVerbFormandSameWord. Lemma matching is BROADER than stem-plus-one-ending, so a verb slot meaning the imperative must be matched exactly.говориandговорилare one lemma and only one of them is a command (cmd/mavend/quiet_toggle.go).cmd/mavend/topics.goand the embedder for open sets, where the question is what a turn is ABOUT. Frozen seeds per subject plus a realotherclass, scored against the turn's own query vector. Same shape as the personal boundary inpersonalboundary.go, with one difference. A topic must clear the runner-up bytopicMargin, because a false claim here spends a network scan rather than one honest "не знаю". The old keyword tests stay as the offline floor.- The ecosystem trio when the answer is not in the utterance at all. Identity is Nexus's, never a local pattern.
Seeds are scoring data. Editing one moves a recogniser and must be re-measured
against the TestONNX* tests, not eyeballed.
Non-goals and hard constraints
Not a nag, not autonomous.
The persona is feminine. Russian self-reference uses feminine forms: рада
not рад, поняла not понял. The owner is male and she speaks to him
informally. Use "ты", singular, never "вы" or "ваш", and never "он" or "его". She talks TO
the owner, not about him. Pet names such as "милый" are forbidden. The name
"Ками" is not. CheckAddress, CheckFeminine and CheckCringe in
internal/phraser/eval/checks.go enforce this, scored by make eval-phrasing.
"Never phones home" is deprecated (owner's call, 2026-07-31). A 1.7B does not know enough to answer world questions, so she reads external sources. What replaces it:
- No telemetry, no cloud model, no third-party account. That part never changes. Nothing about Maven is reported to anyone and inference stays on the box.
- The owner's data first, then the world. Every source reading his facts, notes, calendar, tasks or house runs before anything outside. The personal boundary sits between them. Reading beats recalling for a small model.
- The owner's notes and facts are never search input. Only the utterance goes out. Never the persona block, the history, or matched notes.
- External search is allowed and off unless configured, like weather and
telegram. The code default is off.
deploy/mavend.jsonships asearchblock, so it is on for this box and deleting the block turns it off again. - In the world, live search leads and the ZIMs are the fallback (owner's call, 2026-08-02). A self-hosted SearXNG answers first. The Kiwix ZIMs on homesrv answer when the search is empty, unreachable, or the line is down.
The world chain
Response.Empty() is the whole gate and there is no quality threshold in front
of it. Four signals were tried and none separates a real question from an
invented one. Token overlap would cost "столица Франции" its answer, because the
answer is Париж and that word is not in the question
(docs/evals/2026-08-05-search-quality-signals.md). The embedder is not a
fifth signal: query-to-passage cosine measures topic and not whether the
passage answers, and the two sets overlap
(docs/evals/2026-08-09-kiwix-topic-retrieval.md).
The connect phase alone is capped at dialTimeout (1.5s), because a blackholed
host once cost the owner 8 seconds. A slow instance that did connect keeps the
full 8 (docs/evals/2026-08-05-kiwix-offline-fallback.md).
A Russian question reads wikipedia_ru_all_maxi_2026-02 verbatim through
kiwix.book_ru. The RU→EN rewriter is the workaround for an English book and is
skipped there. Kiwix catalog names come from the filename, not the <name>
field.
Kiwix ranks by keyword overlap. Never send it a whole sentence.
kiwix.Topic drops the narrative request, the interrogative and a verb behind
one. kiwix.TitlePath tries the exact article first, since a ZIM is addressable
by title and a wrong title is a 404. TitleCandidates tries the spoken form and
then the capitalized one. Both apply on the verbatim path alone. The rewriter
already reduces a question, and reducing twice takes the topic off its input
(V-668).
Which query source claimed a turn is readable on /chat as a badge beside
the reply. It is carried on ipc.ChatReply.Source and noted by noteQuerySource
in cmd/mavend/querysource.go. It rides the context, so handleText keeps the
one string signature the mic, telegram and the web share.
Web UI conventions
Server-rendered pages share cmd/mavweb/static/ui.css (served at /ui.css) and
the shell partial in cmd/mavweb/shell.html. A page opens with
{{template "shellTop" "<page-key>"}} and closes with {{template "shellBottom"}},
and the key marks the active sidebar link.
Every page is its own embedded .html file next to main.go. No page markup
lives in Go, and the sidebar is data (sidebarSections, pageIcon) the template
renders. No per-page <style> beyond true one-offs. Wrap every table in
<div class=scroll> so wide data pans on a phone. Local preview and headless
screenshot recipes are in AGENTS.md.
Vikunja
This repo is project Maven (ID 2). MCP at http://localhost:9100/mcp, or
http://192.168.1.104:9100/mcp from workpc. Feature, bug and deploy tasks go
there.
Vikunja is the durable task store. A task holds the goal, the constraints and the assumption ledger. Work without a task id is work nobody can resume, so a session with no id asks for one before it starts.
The MCP tool schemas are deferred. Load the four you use in ONE call at the start of a session:
ToolSearch("select:mcp__vikunja__list_tasks,mcp__vikunja__get_task_details,mcp__vikunja__create_task,mcp__vikunja__update_task")
Close a finished task with done: true and nothing else (owner's call,
2026-08-07). Do not write a completion summary into the description on the way
out. It is lost anyway, and the durable record is the commit messages and the
merged PR. update_task carrying a description resets done to false, which
is why a write-up ever took two calls.
Session workflow
~/.local/bin/task owns the branch, the commit identity and the PR. One task,
one session, one PR.
task start <vikunja-id> # branch off origin/master, write TASK.md, fetch review comments
task pr # push, open or refresh the PR, label Vikunja, notify
task comments # re-pull this branch's review comments into .task/
Around that, /pickup opens a session and /wrap closes it. Wrap at roughly
half context rather than letting the session compact.
Five stores, and each one owns something the others must not hold:
| Store | Holds | Lifetime |
|---|---|---|
| Vikunja task | goal, constraints, assumption ledger, status | durable |
CLAUDE.md, AGENTS.md |
what an agent must know before touching code | durable |
docs/ |
design, measurements, decisions | durable |
TASK.md |
the brief for this branch, written by task start, immutable |
one branch |
HANDOFF.md |
only what the next agent needs to resume | one session |
TASK.md and .task/ are excluded through .git/info/exclude. HANDOFF.md is
gitignored and injected at session start. If a line in the handoff would still
matter next week, it is in the wrong file.
Docs are tiered by path, so staleness is visible from the filename. Files
directly under docs/ are living and carry a Last verified: <date> @ <sha>
line. Files under docs/evals/ are dated measurements and are never edited after
the day, so a newer number is a new file. Files under docs/archive/ are dead
and read by nobody by default.
This file is a rules file. A new measurement belongs in docs/evals/. The
reasoning behind a subsystem belongs in its living doc. A line here earns its
place only by changing what an agent does.
Git guards
Two hooks in .githooks/, tracked, wired with core.hooksPath. Fresh clone:
git config core.hooksPath .githooks
pre-commitrefuses master, and refuses more than 300 changed lines in non-markdown files. Markdown is exempt and may land as one batch.commit-msgrequires the subject to end with(V-<id>).V-and not#, because Gitea autolinks#123to a Gitea issue, which is the wrong tracker.
Two more guards live outside the repo, in ~/.claude/hooks/. diff-budget.sh
blocks further edits past 600 changed lines on a task/ branch.
prose_lint_hook.py checks prose on every write. Both measure against
origin/master, so a local master that is ahead of the remote makes the diff
budget read high.
--no-verify exists. Using it means saying why in the commit body.