The old model was a symmetric paraphrase model, so it scored "do these look alike" instead of "does this note answer this question". Also fixes the file mismatch: the Makefile, the deploy config and both evals now all name the same quantized file, and the quantized one is what gets measured. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
4.8 KiB
Maven — Agent Context
Vikunja
This repo maps to Maven (project ID: 2) in Vikunja.
Feature work, bugs, deployment tasks all go here.
MCP endpoint: http://localhost:9100/mcp (or http://192.168.1.104:9100/mcp from workpc)
Rendering / previewing the web UI locally
To see mavweb pages with real data without touching the production stack:
R=/tmp/mvn-preview; mkdir -p $R
go build -o $R/mavend ./cmd/mavend/ && go build -o $R/mavweb ./cmd/mavweb/
cat > $R/mavend.json <<EOF
{ "db_path": "$R/maven.db", "socket_path": "$R/mavend.sock",
"state_dir": "$R", "tick_interval": "10s" }
EOF
$R/mavend -config $R/mavend.json &
$R/mavweb -addr 127.0.0.1:9299 -core $R/mavend.sock &
- No models/voice/phraser config needed — the phraser stub covers it; mavend
runs fine bare. mavweb serves
/,/dash,/history,/trace,/notifications,/tools. - Socket path must be short — unix sockets cap at ~108 chars; a deep tmp
dir fails with
bind: invalid argument. - Seed data through
ipc.Client(internal package — the seeder must live inside the module, e.g. a throwawaycmd/seedtmp/main.go, deleted after):WriteFact,RecordNudge+ResolveNudge,CreateReminder,ProposeTool. /traceis empty until the first tick fires (wait onetick_interval).- Screenshots:
chromium --headless --disable-gpu --screenshot=out.png --window-size=1280,900 --hide-scrollbars --virtual-time-budget=2000 http://127.0.0.1:9299/dash(use--window-size=430,900for the phone/PWA view). Always pass--virtual-time-budget— without it the screenshot can snap mid-layout and silently drop elements (the PWA lang toggle "disappeared" this way).
Embedder model for intent routing
The router uses a multilingual sentence embedder to classify intents and recall
notes. Without it, the floor HashEmbedder is used — deterministic but weak
(Russian recall rarely clears the confidence gate, many commands fall to
"clarify").
Download the embedder (ONNX, ~120 MB):
make download-embedder
This fetches multilingual-e5-small (384-dim, 12-layer, Russian and English)
to models/embedder/multilingual-e5-small/. It is an asymmetric retrieval
model: the code puts query: in front of a question and passage: in front
of a stored note, which is how e5 was trained. The quantized file is the one
that is downloaded, deployed and measured.
Also need ONNX Runtime (libonnxruntime.so):
curl -sL "https://github.com/microsoft/onnxruntime/releases/download/v1.15.1/onnxruntime-linux-x64-1.15.1.tgz" | tar xz
sudo cp onnxruntime-linux-x64-1.15.1/lib/libonnxruntime.so* /usr/local/lib/
Configure in deploy/mavend.json:
"voice": {
"embedder": {
"model_path": "models/embedder/multilingual-e5-small/model_quantized.onnx",
"tokenizer_path": "models/embedder/multilingual-e5-small/tokenizer.json",
"lib_path": "/usr/local/lib/libonnxruntime.so"
}
}
Without the embedder block, the daemon uses HashEmbedder (works, but weak on
Russian recall — you may see many "clarify" responses).
Qwen3 resident model for router + phraser
The target daemon uses the locally trained Qwen3-1.7B checkpoint for both
routing and phrasing. Training is Qwen3 Base → RU CPT → joint persona/router
SFT → merged GGUF, and is still in flight (#122) — until it lands, the deployed
resident model is stock Qwen3.5-0.8B (Q4_K_M), see deploy/mavend.json.
Without a configured model, StubPhraser plus the classifier remain the
deterministic floor.
During training, use the runbook in
docs/plans/2026-07-18-qwen3-resident-training-eval.md. After the decision gate
and SFT pass, copy the merged GGUF into the mounted model directory and set:
"phraser": {
"model_path": "/opt/maven/models/llm/Qwen3-Maven-1.7B-Q8_0.gguf",
"bin_path": "llama-server",
"n_gpu_layers": 99,
"n_ctx": 2048
}
Configure in deploy/mavend.json — the phraser block points at this
model and the daemon spawns llama-server as a subprocess. The router and
replier use the same llama-server via the shared internal/llm client.
Telegram tokens are read from deploy/telegram.env (gitignored), expanded
via ${VAR} in the JSON config.
Routing is Qwen-first with classifier fallback. The LLM router runs after stage-0 (exact-match grammar) and before the classifier cascade. On any error or parse failure, the classifier handles the utterance — the turn never breaks on the model.
Web UI conventions
- All server-rendered pages share
cmd/mavweb/static/ui.css(served at/ui.css) and thenavtemplate partial (navHTMLincmd/mavweb/main.go, invoked as{{template "nav" "<active-page>"}}). New pages must link both — no per-page inline<style>beyond true one-offs. - Wrap every table in
<div class=scroll>so wide data pans on a phone instead of breaking the layout.