d0afd9d4f6
Stock Qwen3-1.7B, not the CPT'd one — that training is still running. It won on both fixtures we have, measured tonight on an otherwise idle box: routing, 77 RU cases, intent-only: 67.5% vs 59.7% for Qwen3.5-0.8B talk fixture, 27 cases: 20/27 vs 11-17/27 It also beat Qwen3.5-2B, which is 20% larger, on every routing column. Two other things came with it: n_ctx goes 2048 -> 4096. This is a Thinking variant, so reasoning tokens need the room, and 4096 is the context every score above was measured at. Shipping 2048 would ship something nobody measured. The doc now says not to bother with sub-500M models, because I checked and they are not close. LFM2.5-350M routes at 5.2% — worse than guessing among 7 intents — and answers "столица Франции?" with "Сторзит", which is not a word. The 230M replies to Russian in Spanish. Their published IFEval and BFCL numbers are good and they are all English. Note the routing gain needs the LLM router actually wired on to show up. It is still nil, so this commit buys the phrasing improvement today and the routing improvement when that lands. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
Maven — Docker deployment
One image, one container per daemon (docker-compose.yml). Core (mavend)
holds the encryption key and the db; the modules mount only the shared socket
dir and read-only models.
First run
# 1. generate the at-rest db key (32 bytes, base64) — keep it safe, losing it loses the db
cp deploy/db_key.env.example deploy/db_key.env
printf 'MAVEN_DB_KEY=%s\n' "$(openssl rand 32 | base64 -w0)" > deploy/db_key.env
# 2. build + start
docker compose build
docker compose up -d
# 3. logs
docker compose logs -f mavend
models/ and deps/ are bind-mounted / baked from the host — they are NOT in
git (fetched via make deps + downloaded models). The build context needs
deps/lib, deps/piper, deps/include, and deps/whisper.cpp/ggml/include
present (see .dockerignore).
Layout
| Path (in container) | What |
|---|---|
/opt/maven/bin |
the six daemons |
/opt/maven/lib |
native .so (whisper+vulkan, onnxruntime) |
/opt/maven/piper |
piper binary + espeak data |
/opt/maven/models (ro) |
bind-mount of ./models |
/run/maven (volume) |
shared IPC sockets |
/var/lib/maven (volume) |
encrypted db at rest |
/dev/shm (tmpfs) |
decrypted db working copy (RAM only) |
Not yet verified / host-dependent
This stack is correct-by-construction but has not been build-tested here (no docker in the authoring env; ~1GB context; GPU). Expect a tweak on first build on the target host, most likely in one of these:
- GPU passthrough —
mavsttdmaps/dev/drifor Vulkan. On an NVIDIA host you'd swap to the nvidia container runtime instead of/dev/dri. - onnxruntime lib path —
mavend's embedder needslibonnxruntime.so(onLD_LIBRARY_PATH=/opt/maven/lib). If the embedder wants an explicit path, set it in the config's embedder block. - cross-container voice —
mavweb -voice mavend:9100only works oncemavendbinds its voice server on0.0.0.0:9100(Voice config, currently unset). Until then, voice-over-web is inert;/tools, passkey, and the dash work fine over the core socket. - netdata —
mavpollreaches it viahost.docker.internal; adjust if netdata runs elsewhere.