The owner's call, 2026-07-31: a 0.8B model does not know enough about the world to be useful without reading something. So she may now read external sources to answer world questions. What replaces the old rule, in all three docs: - No telemetry, no cloud model, no third-party account. Unchanged. - Local first: the Kiwix ZIMs on the box before anything on the network. - External search is allowed but off unless configured, same as weather and telegram. - His notes and facts are never search input. Only the utterance goes out — never the persona block, the history, or matched notes. Docs only, no code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
5.8 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Maven is a self-hosted, privacy-first voice assistant (Russian + English). Go daemons
talking over unix sockets; one resident small model for routing + phrasing; whisper.cpp STT, piper TTS.
Deploy target is a Ryzen laptop (homesrv) with Vulkan offload to the Vega iGPU (n_gpu_layers: 99,
compose passes /dev/dri + the render gid) — the resident model stays ≤1.7B either way.
Resident model: currently Qwen3.5-0.8B (Q4_K_M), the smallest checkpoint in the gguf
library, picked for CPU/iGPU latency. The target is the locally CPT'd Qwen3-1.7B; that
training is still in flight (Vikunja #122), so no such gguf exists yet. Model files live in
/mnt/hdd1/llms, bind-mounted to /opt/maven/models/llm — which shadows the repo's
models/llm/, so the LFM2.5 gguf sitting there is not loaded by anything. Swapping the resident
model is a one-line change to phraser.model_path in deploy/mavend.json.
See REARCH.md for the target architecture, DESIGN.md for the folded design spec, and
AGENTS.md for local-preview + model-download recipes.
Build & test
CGO daemons (mavend, mavsttd, mavttsd, mavenclient) need the vendored toolchain
and libs wired through the Makefile — do not call go build on them bare, use make:
make build # all 8 binaries
make build-web # single daemon (pure-Go ones: web/waked/poll/caldav build without CGO)
make test # go test -race across ./internal/... ./cmd/... with CGO env set
Run a single test (must carry the CGO env for packages that touch STT/TTS/voice):
CGO_CFLAGS="-I$(pwd)/deps/include -I$(pwd)/deps/whisper.cpp/ggml/include" \
CGO_LDFLAGS="-L$(pwd)/deps/lib -Wl,-rpath,$(pwd)/deps/lib" \
LD_LIBRARY_PATH="$(pwd)/deps/lib" \
deps/go/go/bin/go test -run TestName ./internal/router/
Pure-Go packages (router, memory, mavweb, …) run under a plain go test ./pkg/.
The daemons (cmd/)
| Binary | Role |
|---|---|
mavend |
Core. Router, phraser, memory, reminders, digestion tick. Owns the DB + IPC socket. |
mavweb |
HTTP UI + PWA (/dash, /history, /trace, /notifications, /tools); WebAuthn auth. Connects to mavend's socket. |
mavsttd |
Speech-to-text (whisper.cpp, CGO). |
mavttsd |
Text-to-speech (piper subprocess). |
mavwaked |
Wake-word / VAD gate. |
mavenclient |
Voice loop client (mic → stt → core → tts). |
mavpoll |
Telegram long-poll reach. |
mavcaldav |
CalDAV calendar sync. |
Daemons are wired socket-to-socket, not linked. internal/ipc is the client/server wire
protocol; the config in deploy/mavend.json (with ${VAR} env expansion from gitignored
deploy/telegram.env) sets socket paths, model paths, and the phraser/embedder blocks.
Routing — read this before touching the router
internal/router/ has TWO layered engines and the committed default is an interim
stopgap, not the intended design (see memory routing-architecture-target):
- Target (REARCH.md): LLM-as-router. One resident Qwen3-1.7B (
llmrouter.go) emits GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is demoted from a routing gate to a RAG hint. - Current stopgap:
llmrouteris wirednil(aroundvoice.go), so theclassifier.go+embedder.gonearest-neighbour cascade actually runs. It routes by similarity to frozen seed phrases — the known cause of weak RU query handling.
Cascade order: stage0.go exact-match fast-path → LLM router (when non-nil) → classifier
fallback. Any LLM error falls through to the classifier so a turn never breaks on the model.
LLM output contract
All phrasing paths emit {"response":"...","mood":"..."} (parsed in replier_llm.go and
internal/phraser/llmphraser.go), with fallback to plain text and the legacy
{"body","summary"}. Mood is a fixed enum. Router prompt is a separate contract:
[{"intent":<enum>, key?, value?, text?, verb?}, ...], 7 intents (fact, reminder, note, query, act, chat, system). llm/check_prompt_parity.py in the training
workspace enforces that the Go and relabelling prompts remain identical.
Non-goals (hard constraints)
Not a nag, not autonomous. Maven's persona is feminine — Russian
self-reference must use feminine forms (the user is male; see memory maven-persona-gender).
"Never phones home" is DEPRECATED (owner's call, 2026-07-31). It used to be a hard constraint and it is not one any more: a 0.8B — and a 1.7B — does not know enough to answer world questions, so she needs to read external sources. What replaces it:
- No telemetry, no cloud model, no third-party account. That part never changes. Nothing about Maven is reported to anyone, and inference stays on the box.
- Local sources first. Kiwix ZIMs on homesrv (Wikipedia, ifixit) before anything on the network. Reading beats recalling for a small model, and a local read costs nothing.
- External search is allowed and off unless configured, like the weather and telegram capabilities.
- His notes and facts are never search input. Looking up why the sky is blue and sending his stored personal notes to an upstream engine are different acts. Only the utterance goes out, never the persona block, history, or matched notes.
Web UI conventions
Server-rendered pages share cmd/mavweb/static/ui.css (served at /ui.css) and the nav
partial (navHTML in cmd/mavweb/main.go, {{template "nav" "<active-page>"}}). No
per-page <style> beyond true one-offs. Wrap every table in <div class=scroll> so wide
data pans on a phone. Local preview + headless screenshot recipe is in AGENTS.md.
Vikunja
This repo is project Maven (ID 2) in Vikunja. MCP: http://localhost:9100/mcp (or
http://192.168.1.104:9100/mcp from workpc). Feature/bug/deploy tasks go there.