Files
Maven/deploy
kami 76a6a007ef Pin the resident model to Qwen3.5-0.8B and name Qwen3-1.7B as the target
The most load-bearing decision in the project was stated four incompatible
ways: the docs said Qwen3-1.7B, deploy/mavend.json said Qwen3.5-2B, the repo's
models/llm/ held an LFM2.5-1.2B gguf, and five code comments still said LFM.
Answering "which model is deployed" meant re-deriving it from scratch every
time.

Two facts the review missed, found while resolving it:

- /mnt/hdd1/llms is bind-mounted over /opt/maven/models/llm, which shadows the
  repo's models/llm/. The LFM2.5 gguf sitting there was never loaded by
  anything, so it was not evidence of the deployed model at all.
- That library holds Qwen3.5-0.8B, -2B and -4B, and no Qwen3-1.7B. The config
  pointed at a file that does exist; the docs' Qwen3-1.7B was the stale claim,
  the reverse of the assumed direction. Qwen3-1.7B is the CPT target, and that
  training is still in flight (Vikunja #122), so no such gguf exists yet.

phraser.model_path moves to Qwen3.5-0.8B (Q4_K_M) — the smallest checkpoint on
disk, chosen for latency, and relevant to whether the LLM router is affordable
on this box. Docs and comments now say the same thing in one voice: 0.8B
resident now, CPT'd Qwen3-1.7B as the target, and the bind-mount shadowing
written down so the next reader does not mistake models/llm/ for ground truth.
Comments name the model, never a filename, so a swap stays a one-line config
change.

n_gpu_layers: 99 is correct and stays — compose passes /dev/dri and the render
gid for Vulkan offload to the Vega iGPU. CLAUDE.md's "CPU-only" was the stale
half of that contradiction and is corrected.

phraser.go also dropped a wrong "sub-1b, prompted not trained" size claim: the
target is trained end-to-end (RU CPT + joint persona/router SFT).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X5JApcrCRVGmqrxnhynSik
2026-07-30 23:40:33 +04:00
..

Maven — Docker deployment

One image, one container per daemon (docker-compose.yml). Core (mavend) holds the encryption key and the db; the modules mount only the shared socket dir and read-only models.

First run

# 1. generate the at-rest db key (32 bytes, base64) — keep it safe, losing it loses the db
cp deploy/db_key.env.example deploy/db_key.env
printf 'MAVEN_DB_KEY=%s\n' "$(openssl rand 32 | base64 -w0)" > deploy/db_key.env

# 2. build + start
docker compose build
docker compose up -d

# 3. logs
docker compose logs -f mavend

models/ and deps/ are bind-mounted / baked from the host — they are NOT in git (fetched via make deps + downloaded models). The build context needs deps/lib, deps/piper, deps/include, and deps/whisper.cpp/ggml/include present (see .dockerignore).

Layout

Path (in container) What
/opt/maven/bin the six daemons
/opt/maven/lib native .so (whisper+vulkan, onnxruntime)
/opt/maven/piper piper binary + espeak data
/opt/maven/models (ro) bind-mount of ./models
/run/maven (volume) shared IPC sockets
/var/lib/maven (volume) encrypted db at rest
/dev/shm (tmpfs) decrypted db working copy (RAM only)

Not yet verified / host-dependent

This stack is correct-by-construction but has not been build-tested here (no docker in the authoring env; ~1GB context; GPU). Expect a tweak on first build on the target host, most likely in one of these:

  • GPU passthroughmavsttd maps /dev/dri for Vulkan. On an NVIDIA host you'd swap to the nvidia container runtime instead of /dev/dri.
  • onnxruntime lib pathmavend's embedder needs libonnxruntime.so (on LD_LIBRARY_PATH=/opt/maven/lib). If the embedder wants an explicit path, set it in the config's embedder block.
  • cross-container voicemavweb -voice mavend:9100 only works once mavend binds its voice server on 0.0.0.0:9100 (Voice config, currently unset). Until then, voice-over-web is inert; /tools, passkey, and the dash work fine over the core socket.
  • netdatamavpoll reaches it via host.docker.internal; adjust if netdata runs elsewhere.