Three documentation changes and one deletion. CLAUDE.md and AGENTS.md gain the Nexus/Praxis/Hexis sections that were written last session and never committed: what each service owns, where Maven's client for it lives, and the rules that are not negotiable. The p50 latency figure was wrong in two files. CLAUDE.md said the cascade costs 2.7s and that the LLM router is 90x slower than the classifier. Both come from the bakeoff table, where the number is contention on a shared llama-server, not the model. ROUTING-EVAL-31-07-2026.md line 61 says so and measures the router at p50 825ms / p95 1.2s / max 3.0s. Corrected in CLAUDE.md, and the bakeoff table now carries a header pointing at the routing eval for absolute latency. Latency work was about to be planned off a number that was never real. HANDOFF.md is deleted. It described work sitting on fix/integrated waiting for a fast-forward onto overnight/eco-versioned-traces. Neither is true: master contains that tip plus 22 commits, and both branch pointers are stale. The three live defects it recorded move to PLAN-DETERMINISM-02-08-2026.md, which is now the only planning document. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BFaeSbLMEVG5ey8tejU3y2
13 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Maven is a self-hosted, privacy-first voice assistant (Russian + English). Go daemons
talking over unix sockets; one resident small model for routing + phrasing; whisper.cpp STT, piper TTS.
Deploy target is a Ryzen laptop (homesrv) with Vulkan offload to the Vega iGPU (n_gpu_layers: 99,
compose passes /dev/dri + the render gid) — the resident model stays ≤1.7B either way.
Resident model: currently Qwen3-1.7B (UD-Q4_K_XL), stock — not yet the CPT'd one.
It replaced Qwen3.5-0.8B on 2026-07-31 because it measured better on both fixtures we have:
67.5% vs 59.7% intent-only on the 77-case RU routing fixture, and 20/27 vs 11-17/27 on the
talk fixture. See MODEL-BAKEOFF-31-07-2026.md. It is a Thinking variant, so n_ctx is 4096
— reasoning tokens need the room, and 4096 is what the scores above were measured at.
The target is still the locally CPT'd Qwen3-1.7B (Vikunja #122, training in flight).
Stock already speaks good Russian; what it gets wrong is the persona — it writes я рад,
masculine, where Maven needs рада. That is what the CPT is for.
Do not bother with sub-500M models. LFM2.5-230M and 350M were measured on 2026-07-31 and
both are unusable in Russian: the 350M routes at 5.2% (worse than guessing) and answers
"столица Франции?" with the invented non-word "Сторзит"; the 230M replies to Russian in
Spanish. Their strong published IFEval/BFCL numbers are English-only. Model files live in
/mnt/hdd1/llms, bind-mounted to /opt/maven/models/llm — which shadows the repo's
models/llm/, so the LFM2.5 gguf sitting there is not loaded by anything. Swapping the resident
model is a one-line change to phraser.model_path in deploy/mavend.json.
See REARCH.md for the target architecture, DESIGN.md for the folded design spec, and
AGENTS.md for local-preview + model-download recipes.
Build & test
CGO daemons (mavend, mavsttd, mavttsd, mavenclient) need the vendored toolchain
and libs wired through the Makefile — do not call go build on them bare, use make:
make build # all 9 binaries
make build-web # single daemon (pure-Go ones: web/waked/poll/caldav build without CGO)
make test # go test -race across ./internal/... ./cmd/... with CGO env set
Run a single test (must carry the CGO env for packages that touch STT/TTS/voice):
CGO_CFLAGS="-I$(pwd)/deps/include -I$(pwd)/deps/whisper.cpp/ggml/include" \
CGO_LDFLAGS="-L$(pwd)/deps/lib -Wl,-rpath,$(pwd)/deps/lib" \
LD_LIBRARY_PATH="$(pwd)/deps/lib" \
deps/go/go/bin/go test -run TestName ./internal/router/
Pure-Go packages (router, memory, mavweb, …) run under a plain go test ./pkg/.
The daemons (cmd/)
| Binary | Role |
|---|---|
mavend |
Core. Router, phraser, memory, reminders, digestion tick. Owns the DB + IPC socket. |
mavweb |
HTTP UI + PWA (/dash, /history, /trace, /notifications, /tools); WebAuthn auth. Connects to mavend's socket. |
mavsttd |
Speech-to-text (whisper.cpp, CGO). |
mavttsd |
Text-to-speech (piper subprocess). |
mavwaked |
Wake-word / VAD gate. |
mavenclient |
Voice loop client (mic → stt → core → tts). |
mavpoll |
Telegram long-poll reach. |
mavcaldav |
CalDAV calendar sync. |
mavmaild |
Mail reader (IMAP, read-only). Holds the IMAP password; core never sees it. |
Daemons are wired socket-to-socket, not linked. internal/ipc is the client/server wire
protocol; the config in deploy/mavend.json (with ${VAR} env expansion from gitignored
deploy/telegram.env) sets socket paths, model paths, and the phraser/embedder blocks.
The ecosystem: Nexus, Praxis, Hexis
Maven is one of four services. It owns conversation and personal memory. It does not
own identity, operational state, or execution. Full contract in
MAVEN_ECOSYSTEM_ARCHITECTURE.md.
Nexus identifies. Praxis observes. Hexis acts. Maven understands and coordinates.
| Service | Owns | Maven's client | Configured at |
|---|---|---|---|
| Nexus | Canonical entity ids, names, aliases, relationships. Projects, services, devices, people, pets, places. | nexusClient in cmd/mavend/ecosystem.go, POST /api/v1/resolve |
nexus.url (http://nexus:9740) |
| Praxis | Operational attention and item lifecycle. What needs looking at, what changed, what is still unresolved. | praxisClient, the HTTP tools API under /api/v1/tools/ |
praxis.url (http://praxis:8989) |
| Hexis | The capability registry and the only path to executing anything. | vendored github.com/kami/hexis/pkg/client |
hexis.url (http://hexis:9741) |
All three are nil unless configured, and every one of them degrades on its own.
An outage means a named gap in the answer, never a broken turn and never a guess.
Rules that are not negotiable:
- No component reads another component's database. Praxis attention comes over HTTP, never from its SQLite file.
- Identity lives in Nexus. Do not invent a local fact key for something Nexus
resolves.
actionFactalready setsSubject, andcmd/mavend/factenrichment.goresolves it in the background against Nexus. - Free text never reaches a mutating Hexis call. Resolve to a canonical entity id first. Ambiguous resolution asks the owner, it does not pick.
- LLM output is not authorization. Confirmation binds capability id, target
entity, arguments, requester and expiry. See
cmd/mavend/confirm.go. - Praxis lifecycle words mean different things. Surfaced is not acknowledged,
acknowledged is not resolved, execution success is not recovery. Reading an item
aloud calls
Surface, neverAcknowledge. - No automatic attention-to-action path. Digestion may summarise Praxis. It may not call Hexis.
Every cross-service call carries a correlation id minted once per action
(withCorrelationID), a contract version header, and X-Requested-By: maven.
Routing — read this before touching the router
internal/router/ has TWO layered engines. The LLM router is now the default and it is
on in deploy — this section used to say it was wired nil, which stopped being true on
2026-07-31.
- LLM router (the intended design, REARCH.md): the resident Qwen3-1.7B (
llmrouter.go) emits GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is demoted from a routing gate to a RAG hint. Wired atvoice.go:214viapickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient); the flag isvoice.llm_router(config.go),DefaultLLMRouteris on, anddeploy/mavend.jsonsets ittrue. - Classifier cascade (the failure floor, not dead code):
classifier.go+embedder.gonearest-neighbour over frozen seed phrases. It runs when the LLM router is off, when there is no llama-server to talk to (pickLLMRouterlogs that and degrades), and on any per-turn LLM error. Do not delete it — routing by seed similarity is the known cause of weak RU query handling, but a turn must never break on the model.
Cascade order: stage0.go exact-match fast-path → LLM router (when non-nil) → classifier
fallback. Any LLM error falls through to the classifier so a turn never breaks on the model.
Measured on the 77-case RU fixture (MODEL-BAKEOFF-31-07-2026.md): the classifier scores
36.8% full accuracy at p50 31ms; Qwen3-1.7B scores 67.5% intent-only / 72.7% through the
cascade at p50 ≈825ms. Accuracy roughly doubled, latency is ~27× worse, and that trade was
accepted deliberately. The ≈2.7s figure that stood here until 2026-08-02 was contention,
not the model. See ROUTING-EVAL-31-07-2026.md line 61, which measures the LLM router at
p50 825ms / p95 1.2s / max 3.0s and the full cascade at p50 0.80-1.04s. Do not plan latency
work off the bakeoff table. Confidence: 1.0 used to be hardcoded in llmrouter.go, so the LLM
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
#359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with
no allowlisted fn) feeding the same stage-3 gate the classifier path already had — see
gateLLMDecision in router.go. Note the second half of that bug: the LLM branch never
consulted r.threshold at all, so a correct low confidence would have been discarded anyway.
Re-measured on the fixture after the fix: missed clarify 6/6 → 1, at the cost of 3 false
clarifies and 2.6pt of full accuracy (72.7% → 70.1%, intent-only 67.5% → 74.0%). Two of the
three false clarifies are acts the model mis-routed and the gate caught — asking beats wrongly
executing, so the fixture and the daemon disagree about what is correct there. The third,
"поужинал", was a real defect: the single-token rule was an English intuition and does not
transfer to Russian, where one word is routinely a whole sentence.
Narrowed 01-08-2026. thinSingleToken (internal/router/singletoken.go) still thins a bare
one-word nominal — "вода", "бэкап" — but spares two classes: a closed lexicon of social and
control singles ("привет", "спасибо", "стоп", "yes"), and any token carrying a Russian verb
ending (past tense, 2nd person, reflexive), because a verb already contains its subject. Both
tests are offline and cost nothing. Re-measured: false clarifies 3 → 2, intent-only 74.0% →
75.3%, full accuracy unchanged at 70.1%, missed clarify still 1. The two remaining false
clarifies are the act-with-no-allowlisted-fn arm of the gate, not this rule.
Agenda questions taken off the model, 01-08-2026. AgendaQueryGrammars (stage0.go, wired
after the clock rules in buildRouter) routes "что у меня сегодня", "во сколько у меня
встреча" and anything naming a calendar to IntentQuery at stage 0. They were going to
IntentSystem, where replySystem has no agenda arm and answered "пока не умею" — the
fixture had said query since ru-query-019 was written. Measured: full accuracy 70.1% →
72.7%, intent-only 75.3% → 77.9%, calendar 0/2 → 2/2, clarify counts unchanged. Note that
Go's \b is ASCII-only and never fires after a Cyrillic letter; the pattern needs an
explicit (\s|[?!.]|$).
LLM output contract
All phrasing paths emit {"response":"...","mood":"..."} (parsed in replier_llm.go and
internal/phraser/llmphraser.go), with fallback to plain text and the legacy
{"body","summary"}. Mood is a fixed enum. Router prompt is a separate contract:
[{"intent":<enum>, key?, value?, text?, verb?}, ...], 7 intents (fact, reminder, note, query, act, chat, system). llm/check_prompt_parity.py in the training
workspace enforces that the Go and relabelling prompts remain identical.
Non-goals (hard constraints)
Not a nag, not autonomous. Maven's persona is feminine — Russian
self-reference must use feminine forms — рада, not рад; поняла, not понял. The owner
is male and is addressed informally: "ты", singular, never "вы"/"ваш" and never "он"/"его"
(she talks TO him, not about him). Pet names ("милый", "дорогой") are forbidden; his name
("Ками") is not. The eval enforces this: CheckAddress, CheckFeminine and CheckCringe in
internal/phraser/eval/checks.go, scored by make eval-phrasing.
"Never phones home" is DEPRECATED (owner's call, 2026-07-31). It used to be a hard constraint and it is not one any more: a 0.8B — and a 1.7B — does not know enough to answer world questions, so she needs to read external sources. What replaces it:
- No telemetry, no cloud model, no third-party account. That part never changes. Nothing about Maven is reported to anyone, and inference stays on the box.
- Local sources first. Kiwix ZIMs on homesrv (Wikipedia, ifixit) before anything on the network. Reading beats recalling for a small model, and a local read costs nothing.
- External search is allowed and off unless configured, like the weather and telegram capabilities.
- His notes and facts are never search input. Looking up why the sky is blue and sending his stored personal notes to an upstream engine are different acts. Only the utterance goes out, never the persona block, history, or matched notes.
Web UI conventions
Server-rendered pages share cmd/mavweb/static/ui.css (served at /ui.css) and the nav
partial (navHTML in cmd/mavweb/main.go, {{template "nav" "<active-page>"}}). No
per-page <style> beyond true one-offs. Wrap every table in <div class=scroll> so wide
data pans on a phone. Local preview + headless screenshot recipe is in AGENTS.md.
Vikunja
This repo is project Maven (ID 2) in Vikunja. MCP: http://localhost:9100/mcp (or
http://192.168.1.104:9100/mcp from workpc). Feature/bug/deploy tasks go there.