PROGRESS.md: - remove Kuma from 'Wired but needs deploy' (key wired ineda434f) - remove cold-start unlock from 'Not built yet' (built in b0932a1+15fe7bb) - add third-session block: cold-start unlock, conversation depth, routing+persona — with the cold-start test gap called out - update gap #2: anaphora + cross-intent landed (05236ad) - update gap #6: note download-embedder + query_min_score knob - remove 'Personality prompt' from Future (landed inb7eb53a) - add caveat: cold-start unlock tests missing (3 required cases absent) ROADMAP.md: - add Status line under every item header with commit SHAs + verdict - summary table gains a Status column - replace stale 'Recommended order' with 'Remaining work' priority list
35 KiB
Maven — Roadmap Spec (post-jul6, 2026-07-06)
The north star (
SPEC.md) is settled; this is the execution plan for everything between "works" and "the thing the spec describes." Each item names the exact files, interfaces, and done-when criteria so an agent (or human) can execute without guessing. Priority is ops → security → dealbreakers → depth → deferred. Estimates are wall-clock for one focused session, not calendar time.
Conventions
- Done when = a checkable finish criterion. If it can't be checked, it isn't done.
- Files = the exact paths to touch. No "somewhere in internal/."
- Interface = the seam a new piece plugs into. If the seam doesn't exist yet, the item says so and names where it's added.
- Deps = what must land first. Items with no deps can start now.
- Every code item ends with
make testgreen (gofmt + vet + 303+ tests,-race). No exceptions.
P1 — Ops quick wins (do today, ~2h total)
These are not code problems — they're operator actions or trivial fixes that unlock already-built features. Highest ROI per minute on the list.
1.1 Kuma API key for service_down polling
Status: done (eda434f, 2026-07-06). Key uk5_mavpoll-key created, wired
into mavpoll. Agent caught a real bug: kuma expects the key as the password
field, not username. mavpoll switched to network_mode: host (compose bridge
couldn't reach localhost netdata/kuma).
Vikunja: #16 (prio 2) Est: ~5 min Deps: none
Status now: cmd/mavpoll/main.go:54 defines -kuma-key (basic-auth
username, empty password). pollKuma (line 238) calls
p.get(ctx, p.kumaURL, p.kumaKey) which sets basic-auth. Without the key,
pollKuma gets 401 and logs "no monitor_status metrics" — the whole
service_down sev4 rule (internal/loop/rules.go ServiceDownRule) is
dark. Netdata polling works independently.
The gap: the key string doesn't exist. Kuma's /metrics endpoint
requires an API key (Settings → API Keys).
Steps:
- Uptime Kuma UI → Settings → API Keys → create key (any label, e.g. "mavpoll").
- Add
-kuma http://<kuma-host>:3001/metrics -kuma-key <key>to themavpollservice command indocker-compose.yml:78. docker compose up -d mavpoll && docker logs mavpoll— confirm aservice_down=up (poll:uptimekuma)line, not "no monitor_status metrics."- Stop a monitored service, confirm
service_down=downappears within one poll interval (60s default).
Done when: docker logs mavpoll shows service_down=up on a healthy
stack, and toggling a monitored service flips it to down within 60s.
The ServiceDownRule then fires a sev4 nudge through the normal dispatcher.
1.2 Voice bind verification + stale comment fix
Status: done (eda434f, 2026-07-06). Stale comment replaced with verified
note. nc -z confirmed mavweb reaches mavend:9100 cross-container.
Vikunja: #18 (prio 2) Est: ~30 min Deps: none
Status now: deploy/mavend.json:10 already has
"bind": "0.0.0.0:9100". The config validator (config.go:402-404)
rejects voice.enabled with empty bind — so the config is correct.
BUT docker-compose.yml:66-69 has a stale comment: "currently unset.
Until that's configured, voice-over-web is inert." The comment is wrong;
the config is right. The task is now: deploy, verify, fix the comment.
The gap: unverified on the target host. The bind is set; whether
mavweb can actually reach mavend:9100 cross-container is untested.
Steps:
- Fix the stale comment in
docker-compose.yml:66-69— replace with:# voice.bind is 0.0.0.0:9100 in deploy/mavend.json so mavweb can# reach it cross-container. Verified <date>. docker compose up -d && docker compose logs mavend— confirmvoice listening on 0.0.0.0:9100.- From the mavweb container:
curl -sS telnet://mavend:9100or open the PWA and do a push-to-talk round trip. Confirm a reply comes back. - If voice fails cross-container: check
docker network ls, confirm both containers are on the same compose network (default bridge for the project). Thesocketsvolume is for IPC; voice is TCP.
Done when: a PWA push-to-talk round trip works in the Docker deploy (STT → router → TTS → reply audio plays), and the stale comment is fixed. Update PROGRESS.md "Ops footnote" (line 284-286) to "verified."
1.3 Deploy desk_active presence script on desk PC
Status: not done. Script exists (scripts/desk-active.sh, complete) but
the systemd user timer + hypridle listener were not installed on the desk PC
(linux). This is an operator action on a different machine — the agent
couldn't reach it from the homesrv context. 0 facts ever written; /dash
presence still runs on page_heartbeat alone.
Vikunja: #15 (prio 3) Est: ~1 h Deps: none (the script is complete; this is a workstation deploy)
Status now: scripts/desk-active.sh is a complete one-shot poster
(25 lines). It POSTs to mavweb /api/signal?key=desk_active over wg.
cmd/mavweb/main.go:36-40 allowlists desk_active → source
infer:hyprland. internal/store/presence.go:45 gives desk_active the
strongest weight (0.90, τ=8min). The script's header (lines 12-18)
documents the exact hypridle + systemd timer wiring. Zero facts have
ever been written — the timer isn't installed on the desk PC.
The gap: the systemd user timer + hypridle listener don't exist on the workstation (linux, arch, hyprland).
Steps (on the desk PC, not homesrv):
- Copy
scripts/desk-active.shto~/.local/bin/desk-active.sh,chmod +x. - Create
~/.config/systemd/user/maven-desk.service:[Unit] Description=maven desk-active presence ping [Service] Type=oneshot Environment=MAVEN_URL=https://maven.kvmx.ru:9443 ExecStart=%h/.local/bin/desk-active.sh - Create
~/.config/systemd/user/maven-desk.timer:[Unit] Description=maven desk-active presence (60s) [Timer] OnBootSec=10s OnUnitActiveSec=60s AccuracySec=5s [Install] WantedBy=timers.target - Add to
~/.config/hypridle.conf(the script header lines 14-18 show this exactly):listener { timeout = 120 on-timeout = systemctl --user stop maven-desk.timer on-resume = systemctl --user start maven-desk.timer } systemctl --user daemon-reload && systemctl --user enable --now maven-desk.timer. Reload hypridle (or restart the session).- On homesrv: confirm
desk_activefacts appear —mavweb /dashpresence should flip from "away" to "present" within 60s of activity, and back to "away" ~8min after going idle.
Done when: /dash reads "present" while the desk PC is in use, and
"away" within ~8min of hypridle triggering (2min idle + 6min decay).
Presence is now 2 signals (desk_active + page_heartbeat) instead of 1.
P2 — Security (real code)
2.1 Cold-start unlock (passkey → L3 key seam)
Status: code done, tests missing (b0932a1 + 15fe7bb, 2026-07-06).
internal/webauthn/keywrap.go (HKDF-SHA256 + AES-256-GCM, stdlib-only),
locked-mode boot in cmd/mavend/main.go (lockedAPI stub, srv.Check
allowlist), MethodStoreEncryptionKey/MethodUnlock IPC, mavweb
RegisterFinish wraps + AssertFinish unlocks. Env-key fallback preserved.
Gap: the roadmap's done-when #4 required three new test cases
(wrap/unwrap round-trip, wrong-cred unwrap fails, locked-mode IPC rejects
non-unlock methods) — none were written. make test is green by omission,
not coverage. Write these before relying on the cold-start path with real keys.
Vikunja: #14 (prio 2) Est: ~1 day Deps: none (the seam is documented in code comments, just not wired)
Status now: the at-rest AES key comes from config.DBEncryptionKey()
(config.go:433) which reads db_key_env (env var) or db_key_b64
(config). cmd/mavend/main.go:76-85 calls this, then
store.OpenEncrypted(ctx, cfg.DBPath, cfg.DBTmpfs, key). The key lives
in the container env (docker-compose.yml:28 env_file: [./deploy/db_key.env]). The seam is documented in three places:
config.go:44-46: "the passkey op produces the 32 bytes and calls store.OpenEncrypted directly, bypassing config"store/crypt.go:40-43: "this same []byte seam is where the L3 passkey-derived cold-start key plugs in later (auth/tier.go Layer3)"auth/tier.go:42-44: Layer3 = "core cold-start unlock"
The passkey infrastructure exists: internal/webauthn/ does real
WebAuthn (ES256, sign-count regression), cmd/mavweb/webauthn.go serves
enroll + assert, PasskeySession.Assert bumps to L3 for 5min
(mavend/main.go:194-196). But the assertion only flips the session
tier — it doesn't produce key material. The store is already open by the
time the session exists.
The gap: the daemon opens the store at boot from env, before any
passkey assertion is possible. There's no path where a passkey gesture
produces the 32 bytes that store.OpenEncrypted consumes. The key is
in the container env, which means anyone with docker inspect or root
on homesrv can read it — the encryption protects against disk theft, not
against container compromise.
Design decision — how the passkey produces the key:
The cleanest path that fits the existing seams:
-
The key is wrapped, not derived. A passkey assertion doesn't produce 32 deterministic bytes (WebAuthn signatures are randomized). Instead: at enrollment, generate a random 32-byte AES key, encrypt it with a key derived from the passkey credential, store the wrapped blob on disk. At cold-start, the passkey assertion unwraps it.
-
KDF: HKDF-SHA256. Input: the credential public key bytes (stable across assertions) + a salt stored alongside the wrapped blob. Output: 32 bytes. This avoids the "sha256(passphrase)" trap
crypt.go:43warns about — the input is high-entropy key material, not a passphrase. New dep:golang.org/x/crypto/hkdf(not currently in go.mod — add it). Alternatively, implement HKDF-SHA256 from stdlib (crypto/hmac+crypto/sha256) in ~30 lines to avoid the dep; the spec doesn't care which, just don't use bare sha256. -
Flow:
- Enroll (
/auth/webauthn/register/finish): afterFinishRegistrationsucceeds, generate random 32-byte AES key, derive wrap-key via HKDF(credPublicKey, salt), AES-GCM-wrap the AES key, write{salt, wrapped}to a file (e.g.~/.config/maven/db_key.wrapped). The raw AES key is returned to mavend in-process (not over the wire) and used to open the store. - Cold-start (daemon boot): mavend starts locked — the store
isn't open, the IPC server serves only
MethodAssertStepUp(or a newMethodUnlock). mavweb's passkey assert calls the unlock RPC, which derives the wrap-key from the credential, unwraps the blob, and callsstore.OpenEncrypted. The daemon then wires the rest (loop, voice, delivery) and flips to "unlocked" mode. - Fallback:
db_key_envstill works (dev/CI, or recovery if the wrapped key is lost). The daemon tries wrapped-key unlock first; if no wrapped file exists, falls back to env. This preserves the dev path (no passkey enrolled = env key = plaintext-in-RAM as today).
- Enroll (
Files:
internal/webauthn/— addWrapKey(credPublicKey []byte) ([]byte, error)andUnwrapKey(credPublicKey, blob []byte) ([]byte, error). HKDF-SHA256, salt in the blob.cmd/mavweb/webauthn.go— inRegisterFinish, afterh.store.Save(id, publicKey), callwebauthn.WrapKey(publicKey), write the wrapped blob to a path (new flag-wrapped-key-file, default./db_key.wrapped).cmd/mavend/main.go— restructure boot: if wrapped-key file exists, start in locked mode (IPC serves only unlock); else fall back to env key (current path). AddMethodUnlockto the IPC API (internal/ipc/) that takes the unwrapped key and opens the store.internal/ipc/— new RPCUnlock(key []byte) erroron the CoreAPI interface (or a separateUnlockAPI). mavweb calls it after a successful assert.cmd/mavweb/main.go— afterAssertFinishsucceeds, if a wrapped-key file exists, read it, callcore.Unlock(unwrappedKey).internal/store/crypt.go— no change (the seam already takes[]byte); the unlock path just callsOpenEncryptedwith the unwrapped key instead ofconfig.DBEncryptionKey().
Locked-mode behavior: the daemon starts the IPC server and a minimal
"waiting for unlock" state. The voice server, loop, and delivery don't
start until unlock succeeds. mavweb serves /auth/passkey (so the user
can assert) but /dash, /tools, /api/ptt return 503 with a
"daemon locked" message. This is the one user-visible behavior change —
a cold homesrv now needs a passkey gesture before maven is live.
Done when:
- A fresh deploy with no
db_key.envbut an enrolled passkey: daemon starts locked, mavweb/auth/passkeyassert unlocks it,/dashcomes alive, the store opens with the unwrapped key. - A deploy with
db_key.envset (dev/CI): daemon starts unlocked (fallback path), no passkey needed. - Wrong passkey / deleted wrapped file + no env: daemon stays locked, logs "unlock failed," doesn't crash.
make testgreen. New tests: wrap/unwrap round-trip, wrong-cred unwrap fails, locked-mode IPC rejects non-unlock methods.
Security note: this moves the key from "container env (root-readable)"
to "wrapped on disk, unwrappable only with a passkey gesture." Root on
homesrv can still dump the unwrapped key from mavend's RAM after unlock
— this is disk-theft protection, not RAM-capture protection (same threat
model as crypt.go:33 documents). The win is: a stolen disk or a
docker inspect no longer yields the key.
P3 — "Voice assistant" dealbreakers (big effort, high impact)
These three gaps define why maven is "a dictaphone with a brain" instead of "something you talk to across the kitchen." Each is a multi-day architecture item, not a config tweak.
3.1 Always-on listening (wake word + ambient capture)
Status: MVP (e57647c + 5d1850c, 2026-07-06). cmd/mavwaked/ (main +
vad + vad_test, 10 -race tests) ships as energy-VAD only, no wake-word
model — every utterance fires. SurfaceVoice (L0) caps it. 30ms/16kHz
frame shape matches silero-vad ONNX input 1:1, so the wake-word swap is a
local change in vad.go. Hardware topology settled: client box with mic
(desk PC / pi), not homesrv. mavwaked is a client binary (systemd user unit),
not a docker daemon — the agent first added it to Docker, then reverted
(5d1850c). Remaining: wake-word model (openWakeWord / silero-vad ONNX).
Est: 1-2 weeks + hardware Deps: a capture device (USB mic or a dedicated ESP32-S3 box)
Status now: voice is push-to-talk in the PWA (cmd/mavweb/ serves
the record button). The voice server (internal/voice/server.go) is a
TCP listener that waits for PushToTalk frames — the client decides
when to send audio. There's no wake-word detection, no ambient capture,
no always-on mic. internal/router/stage0.go:24-28 has a wakeword
grammar (a "maven," prefix fast-path) but it's for the text-after-STT,
not for audio-level detection. auth/tier.go:55-58 documents
SurfaceVoice as "a room mic / wake-word path" — the surface exists in
the auth model, the hardware path doesn't.
The gap: no always-on audio capture, no wake-word model. Three sub-problems:
- Wake-word detection — a small model that runs continuously on a mic stream and fires when it hears "maven" (or a chosen phrase).
- Ambient capture — after the wake word, capture N seconds of audio
and send it as a
PushToTalkframe (the existing path). - Hardware — a mic that's always on. Options: (a) a USB mic on homesrv, (b) an ESP32-S3 with I2S mic that streams over wg, (c) a dedicated Pi. (a) is simplest; (b) is the "room device" the spec wants.
Design:
- Wake-word engine: openWakeWord (Python, ONNX, ~10MB models) or
Porcupine (Picovoice, free tier, binary). openWakeWord fits the
self-hosted/no-phone-home invariant better. Run it as a new module
cmd/mavwaked/(mirrorsmavsttd/mavttsdshape): reads audio from a device, runs the wake model, on detection sends a trigger to mavend. - New module
cmd/mavwaked/:- Flag
-device(alsa device, e.g.hw:1,0),-model(ONNX wake model path),-core(mavend IPC socket) or-voice(voice TCP addr). - Reads 16kHz mono PCM from the device (PortAudio or
malgo/miniaudio — CGo, like mavsttd). - Runs the wake model on a sliding window; on detection, captures
~5s of audio (configurable) and sends it as a
PushToTalkframe to the voice server. - The voice server's existing
HandlePushToTalkdoes the rest (STT → router → TTS → reply). The reply goes... where? This is the open question — see below.
- Flag
- The reply problem: push-to-talk replies go back to the PWA that sent the request. An ambient wake-word path has no PWA. Options: (a) play the reply through a speaker on the capture device (the ESP32/USB-mic box needs a speaker), (b) send the reply to the PWA if a session is live (fallback to ntfy if not). (a) is the "room device" path; (b) is the "homesrv with a speaker" path. Pick based on hardware.
- Auth surface:
SurfaceVoice(L0) is already in the auth model. The wake-word path uses it — the voice server'sserveConn(server.go:128-134) has a TODO for the auth handshake populating the surface; today it defaults toSurfacePCClient. The mavwaked module would setSurface=voicein itsPushToTalkReq, capping it at L0 (no destructive acts, no registration — exactly the spec's invariant).
Files:
cmd/mavwaked/— new module (main.go + audio capture + wake model).internal/voice/wire.go— confirmPushToTalkReq.Surfaceis settable tovoice(it is —server.go:184-186readsreq.Surface).internal/voice/server.go:128-134— replace the floorSurfacePCClientdefault with surface-from-handshake (or from the req field, which already wins).docker-compose.yml— addmavwakedservice with/dev/snddevice mapping.deploy/mavend.json— no change (voice server already binds 0.0.0.0:9100).
Done when: saying "maven, ..." across the room (no button press)
triggers a capture → STT → router → TTS → reply, with the reply
audible on the capture device's speaker (or the PWA if one's live). The
auth surface is voice (L0) — destructive acts are refused. make test green (the new module needs unit tests for the wake-detection
logic, mocked audio input).
Open question for the operator: pick the hardware before starting. A USB mic on homesrv is the fast path; an ESP32-S3 room device is the "real" version. The code is the same either way (mavwaked reads a device); the hardware changes the deploy.
3.2 Conversation depth (multi-turn dialogue)
Status: done (05236ad, 2026-07-06). Path 1 (rule-based deepening)
implemented: AnaphoraResolver in router/slots.go (RU pronouns:
это/он/она/оно/тот/мой + inflections), followUpMerge extended for
cross-intent (Query/Fact/Reminder after Fact with anaphora inherits key +
time), Session.History []Turn added, fact-by-key lookup in applyAction.
7 new test cases including the exact done-when scenarios. Path 2 (LLM
dialogue manager) remains future — the sub-1B phraser can't drive it.
Est: 3-5 days Deps: none (the dialogue scaffold is wired; this deepens it)
Status now: internal/dialogue/ has Session + SessionStore +
InheritSlots (pure). cmd/mavend/voice.go:339-349 wires it: a 2-min
session carries slots across same-intent turns
(followUpMerge in cmd/mavend/followup.go). So «напомни завтра» →
«…позвонить маме» works — the second turn inherits the time slot. But:
- Only same-intent turns carry (a different intent is a fresh
session —
followup.go:41). - No anaphora resolution — "она" / "он" / "это" don't refer back to prior entities.
- No LLM-driven dialogue — the sub-1B phraser (
llmphraser.go) only words replies; it doesn't decide what to ask next. - The session is single-slot (one
voiceDialogueID— single-user box,voice.go:264-266).
The gap: real multi-turn needs (a) anaphora resolution, (b) the
router or a dialogue manager deciding "I need to ask for X" vs "I have
enough to act," (c) cross-intent context. The current followUpMerge
is bounded gap-filling, not dialogue.
Design:
This is the item where the sub-1B phraser isn't enough. Two paths:
-
Rule-based deepening (fast, limited): extend
followUpMergeto handle cross-intent slot inheritance for common patterns (e.g.IntentQueryafterIntentFact— "я пил воду?" after "запиши что я пил воду"). Add anaphora resolution for pronouns that reference the prior turn's key entity. This is morefollowup.gologic, no LLM. Covers maybe 60% of real follow-ups. -
LLM dialogue manager (slow, general): add a dialogue turn where the phraser gets the conversation history and decides: act, ask-for- clarification, or ask-for-missing-slot. This needs a bigger model than the 1.2B phraser (or a dedicated dialogue prompt) and a conversation-history buffer in the
Session. TheSessionstruct (internal/dialogue/) would grow aHistory []Turnfield.
Recommended path: start with (1) — it's testable, deterministic, and covers the common cases. (2) is a "when the phraser model is upgraded" item.
Files (path 1):
cmd/mavend/followup.go— extendfollowUpMergeto handle cross-intent patterns. Add anaphora resolution (a pronoun → priorSlots.Keymapping).internal/dialogue/session.go— addHistory []TurntoSession(even if path 1 doesn't use it yet, the field should exist for path 2).cmd/mavend/followup_test.go— new cases: cross-intent inheritance, anaphora resolution.internal/router/slots.go— pronoun detection in the slot extractor (она/он/это/тот/та → reference marker).
Done when: a two-turn exchange like «запиши что я пил воду» → «когда
я это сделал?» answers from the fact just recorded (cross-intent,
anaphora "это" → "пил воду"). A three-turn exchange that should not
carry context («запиши что я пил воду» → «какая погода в москве?» →
«когда я пил воду?») correctly treats the middle turn as a break. make test green with the new cases.
3.3 Latency / streaming
Status: not started. Correctly deferred — the roadmap itself flagged this as "most likely to be deferred" and lowest-ROI of the dealbreakers.
Est: 1-2 weeks Deps: none (architecture rework)
Status now: every voice exchange is a full round trip: record full
clip → upload → whisper (batch) → route → phrase (batch) → piper (batch)
→ play. No streaming either direction. No barge-in (you can't interrupt
maven mid-reply). internal/voice/server.go reads one PushToTalk frame
(one audio blob) and returns one PushToTalkResp (one reply blob). The
wire protocol (internal/voice/wire.go) is request/response, not
streaming.
The gap: three sub-problems:
- Streaming STT — whisper.cpp supports streaming (partial
transcription as audio arrives).
cmd/mavsttdwould need a streaming mode (send partial results, not one final blob). - Streaming TTS — piper can synthesize in chunks.
cmd/mavttsdwould stream audio back as it's generated, not one blob. - Barge-in — the client needs to signal "stop talking, I'm talking
now" mid-reply. The wire protocol needs a new method (e.g.
MethodBargeIn) or a cancel on the stream.
Design:
This is the biggest architecture item. The wire protocol changes from request/response to bidirectional streaming. Two options:
- WebSocket voice — replace the TCP length-prefixed protocol with
WebSocket frames.
coder/websocketis already a dep (mavweb uses it for ntfy). The voice server gets aws.Servepath; the PWA gets aWebSocketclient. Streaming STT/TTS ride the same ws. Barge-in is a control frame. - Keep TCP, add streaming frames — extend the length-prefixed
protocol with
MethodStreamAudio(client → server, chunked) andMethodStreamReply(server → client, chunked). More work, same result.
Recommended: (1) WebSocket — it's the standard, the dep is present,
and the PWA already speaks ws (for ntfy). The TCP path stays for
non-browser clients (the protocol doc PROTOCOL.md would note both).
Files:
internal/voice/wire.go— new streaming methods + frame types.internal/voice/server.go— WebSocket accept path, streaming handler.internal/voice/client.go— WebSocket client.cmd/mavsttd/— streaming transcribe mode (partial results).cmd/mavttsd/— streaming synthesize mode (chunked audio).cmd/mavweb/main.go— PWA ws client for voice (replaces the current fetch-based/api/ptt).PROTOCOL.md— regenerate from the newwire.go.
Done when: a push-to-talk exchange shows partial transcription
within ~500ms of starting to speak (not after the full clip uploads),
and the reply starts playing before the full TTS is generated. Barge-in
(mid-reply speak) stops the TTS and starts a new turn. make test
green. The old TCP path still works for non-browser clients (backward
compat).
Note: this is the item most likely to be deferred — it's a quality-of-experience improvement, not a capability gap. The dealbreaker is always-on listening (3.1); streaming makes it feel better but doesn't change what maven is.
P4 — Capability depth (built but thin)
4.1 Routing quality (dev embedder)
Status: done (b7eb53a, 2026-07-06). make download-embedder fetches
Xenova/paraphrase-multilingual-MiniLM-L12-v2 (~90MB ONNX) + tokenizer.
AGENTS.md documents embedder + libonnxruntime setup. queryMinScore is now
configurable (voice.query_min_score, default 0.55) instead of a hardcoded
const.
Est: 2-4 h Deps: none
Status now: deploy/mavend.json:14-18 configures the ONNX embedder
(production). cmd/mavend/voice.go:150-165 loads it when configured,
falls back to HashEmbedder (1024-dim, rune-based token overlap) when
not. The dev/preview path (AGENTS.md preview instructions) runs without
the embedder → weak RU recall → many commands fall to "clarify." The
queryMinScore gate (voice.go:382) is 0.55, tuned for ONNX; the Hash
floor rarely clears it.
The gap: no dev embedder model is documented or shipped. A developer running the preview has to either (a) download the ONNX model manually, or (b) accept weak routing.
Steps:
- Document the ONNX embedder model download in
AGENTS.md(or a newMODELS.md): which model (multilingual sentence embedder), where to put it (models/embedder/model.onnx+tokenizer.json), where to getlibonnxruntime.so. - Add a
make download-embeddertarget that fetches the model (curl from a pinned URL — HuggingFace, sha256-checked). - Optionally: lower
queryMinScorefor the Hash floor (a config knob, not a code change —voice.router_thresholdexists, butqueryMinScoreis a const atvoice.go:382). Make it configurable: addvoice.query_min_scoretoVoiceConfig, default 0.55.
Done when: a developer running the AGENTS.md preview with the
downloaded embedder gets confident RU routing (most commands route
correctly, not to "clarify"). make test green.
4.2 Act surface broadening
Status: not a code item — operator config. The seeded homelab set
(status/ps/uptime/df/free/logs read-only, restart/stop/reboot gated) ships
in deploy/mavend.json. Broadening to home-automation/media/comms is
editing JSON, not code.
Est: ongoing config Deps: none
Status now: deploy/mavend.json:20-33 seeds 11 tools (6 read-only,
5 destructive). internal/tool/tool.go runs them (argv, no shell).
voice.go:174-176 seeds them at boot. The allowlist is config-driven —
broadening is editing mavend.json, not code.
The gap: the seeded set is homelab-focused. Broadening to home-automation (lights, thermostat), media (play music), or communication (send message) is config + new tool entries.
This is not a code item — it's operator config. The only code
change that might help: a mavweb /tools UI for adding tools without
editing JSON (the page exists, but it enables proposed tools; adding a
new one from scratch is JSON-only). Low priority.
Done when: (operator-defined) — e.g. "lights on/off" works by voice
after adding the tool to mavend.json and the act seed file.
4.3 LTM ANN (approximate nearest neighbor)
Status: deferred (correctly). The memory.Store interface is the swap
point; brute-force cosine is sub-ms at single-user scale. Not started until
note+fact count exceeds ~10k and Search latency shows up in profiles.
Est: ~1 day Deps: none (the interface is the swap point)
Status now: internal/memory/store.go defines Store interface
(Insert, Search). internal/store/memory.go implements it with
brute-force cosine (full scan, Search loads every row). The comment
at memory.go:24-28 says "an ANN index is the swap for later, behind
this same interface." At single-user scale (thousands of rows) a full
scan is sub-millisecond.
The gap: none yet. This is a "when it bites" item. The swap point
is the memory.Store interface — a new implementation (e.g.
internal/memory/ann.go using hnswlib or a sqlite-vec extension) drops
in without touching voice.go or recall.go.
When to do this: when note+fact count exceeds ~10k and Search
latency shows up in profiles. Not now.
Done when: (future) a new memory.Store impl with ANN search
passes the existing memory_test.go suite and shows <1ms latency at
10k+ vectors. Not started until the scale problem is real.
4.4 Persona prompt
Status: done (b7eb53a, 2026-07-06). Persona field in VoiceConfig,
llmphraser prepends to systemPrompt() + querySystemPrompt(). Empty =
current hardcoded feminine-gendered Russian persona (backward compat).
Est: 2-4 h Deps: none
Status now: the phraser has hardcoded system prompts:
llmphraser.go:188— notes query: "You are maven, a self-hosted personal assistant answering from your notes..."llmphraser.go:298-299— nudge: "You are maven, a self-hosted personal assistant. Generate brief, natural nudge messages..."router.KnowledgePrompt()— general knowledge (the deduped single source).
The persona ("feminine-gendered Russian self-reference, she/her") is
baked into these strings, not configurable. internal/voice/replier.go
documents a "personality-prompted nudge tone" vs "chat tone" but the
prompts are inline.
The gap: no configurable persona. Changing maven's character means editing Go strings and recompiling.
Design:
- Add
voice.personatoVoiceConfig(config.go) — a string (or path to a file) holding the persona prompt prefix. llmphraser.goreads it (passed viaConfigor a new field) and prepends to every system prompt. Default = the current hardcoded string (backward compat).- The three prompt sites (notes, nudge, knowledge) all call a
personaPrompt(cfg, base)helper that concatenates.
Files:
internal/config/config.go— addPersona stringtoVoiceConfig.internal/phraser/llmphraser.go— accept persona inConfig, prepend to system prompts.cmd/mavend/voice.go— passcfg.Voice.Personainto the phraser config.deploy/mavend.json— document the field (empty = current behavior).
Done when: setting voice.persona in mavend.json changes maven's
reply character (e.g. more formal, different gender, different name)
without recompiling. Empty = current behavior. make test green.
4.5 Custom TTS voice
Status: not started. Mostly operator work (record ~50-100 clips, train a
piper model). The code already supports it — -model flag takes any piper
voice file, VoiceConfig.Tts.Voice names it.
Est: ~1 day + training time Deps: none (piper supports custom voices)
Status now: cmd/mavttsd/main.go:7 documents the default voice:
models/tts/ru_RU-irina-medium.onnx. docker-compose.yml:57 mounts it.
The -model flag takes any piper voice file. VoiceConfig.Tts.Voice
(config.go:281) allows naming a voice when the worker supports
multiple.
The gap: the voice is the stock irina model. A kami-picked voice
(specific person, specific tone) needs a piper fine-tune: record ~50-100
clips of the target voice, train a piper model, drop the .onnx file
into models/tts/.
Steps:
- Record or source ~50-100 clean clips of the target voice (16kHz mono, ~5-10s each, varied sentences).
- Train a piper voice (
piper train— see piper docs for the dataset format + training script). - Output:
ru_RU-<name>-medium.onnx→models/tts/. - Update
deploy/mavend.json:13tts.voiceor themavttsd -modelflag indocker-compose.yml:57.
This is mostly operator work (recording + training), not maven code. The code already supports it — it's a model-file swap.
Done when: maven's replies use the custom voice. make test green
(tests use the Stub TTS, unaffected).
P5 — Deferred by design
5.1 Multi-user (SPEC item 8)
Status: deferred by design. SPEC fences this explicitly (DO NOT TOUCH THIS PHASE). No second user exists. The append-only schema makes it a
migration (add user_id columns + backfill), not a rewrite. Speaker
attribution needs the second voice to train against.
When to revisit: when a second person is actually in the house and using maven. Not before.
Do not start this without an explicit operator decision. An
autonomous agent that adds user_id columns while touching the store
commits the project to a schema before the constraints that shape it
exist.
Summary table
| # | Item | Prio | Est | Type | Deps | Status |
|---|---|---|---|---|---|---|
| 1.1 | Kuma API key | P1 | 5m | ops | — | done eda434f |
| 1.2 | Voice bind verify + comment fix | P1 | 30m | ops | — | done eda434f |
| 1.3 | desk_active deploy | P1 | 1h | ops | — | not done (operator action on linux) |
| 2.1 | Cold-start unlock | P2 | 1d | code | — | code done b0932a1+15fe7bb, tests missing |
| 3.1 | Always-on listening | P3 | 1-2w | code+hw | hardware decision | MVP e57647c (VAD only, no wake word) |
| 3.2 | Conversation depth | P3 | 3-5d | code | — | done 05236ad |
| 3.3 | Latency/streaming | P3 | 1-2w | code | — | not started (deferred) |
| 4.1 | Routing quality (dev embedder) | P4 | 2-4h | code+docs | — | done b7eb53a |
| 4.2 | Act surface | P4 | ongoing | config | — | not a code item (config) |
| 4.3 | LTM ANN | P4 | 1d | code | scale problem | deferred (scale) |
| 4.4 | Persona prompt | P4 | 2-4h | code | — | done b7eb53a |
| 4.5 | Custom TTS voice | P4 | 1d+train | ops | — | not started (ops) |
| 5.1 | Multi-user | P5 | deferred | — | second user | deferred by design |
Remaining work (in priority order):
- 2.1 tests — write the 3 missing keywrap/locked-mode test cases (~30 min)
- 1.3 desk_active — install systemd timer + hypridle on
linux(~1h, your hands) - 3.1 wake word — swap energy-VAD for silero-vad/openWakeWord ONNX in
vad.go - 3.3 streaming — lowest ROI, defer until 3.1 is real
- 4.5 custom voice — when recording is done