Files
Maven/AGENTS.md
kami f6d5a2a7a4 Swap the embedder to multilingual-e5-small (Vikunja #371, #372)
The old model was a symmetric paraphrase model, so it scored "do these
look alike" instead of "does this note answer this question". Also fixes
the file mismatch: the Makefile, the deploy config and both evals now all
name the same quantized file, and the quantized one is what gets measured.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:38:18 +04:00

4.8 KiB

Maven — Agent Context

Vikunja

This repo maps to Maven (project ID: 2) in Vikunja. Feature work, bugs, deployment tasks all go here. MCP endpoint: http://localhost:9100/mcp (or http://192.168.1.104:9100/mcp from workpc)

Rendering / previewing the web UI locally

To see mavweb pages with real data without touching the production stack:

R=/tmp/mvn-preview; mkdir -p $R
go build -o $R/mavend ./cmd/mavend/ && go build -o $R/mavweb ./cmd/mavweb/
cat > $R/mavend.json <<EOF
{ "db_path": "$R/maven.db", "socket_path": "$R/mavend.sock",
  "state_dir": "$R", "tick_interval": "10s" }
EOF
$R/mavend -config $R/mavend.json &
$R/mavweb -addr 127.0.0.1:9299 -core $R/mavend.sock &
  • No models/voice/phraser config needed — the phraser stub covers it; mavend runs fine bare. mavweb serves /, /dash, /history, /trace, /notifications, /tools.
  • Socket path must be short — unix sockets cap at ~108 chars; a deep tmp dir fails with bind: invalid argument.
  • Seed data through ipc.Client (internal package — the seeder must live inside the module, e.g. a throwaway cmd/seedtmp/main.go, deleted after): WriteFact, RecordNudge+ResolveNudge, CreateReminder, ProposeTool.
  • /trace is empty until the first tick fires (wait one tick_interval).
  • Screenshots: chromium --headless --disable-gpu --screenshot=out.png --window-size=1280,900 --hide-scrollbars --virtual-time-budget=2000 http://127.0.0.1:9299/dash (use --window-size=430,900 for the phone/PWA view). Always pass --virtual-time-budget — without it the screenshot can snap mid-layout and silently drop elements (the PWA lang toggle "disappeared" this way).

Embedder model for intent routing

The router uses a multilingual sentence embedder to classify intents and recall notes. Without it, the floor HashEmbedder is used — deterministic but weak (Russian recall rarely clears the confidence gate, many commands fall to "clarify").

Download the embedder (ONNX, ~120 MB):

make download-embedder

This fetches multilingual-e5-small (384-dim, 12-layer, Russian and English) to models/embedder/multilingual-e5-small/. It is an asymmetric retrieval model: the code puts query: in front of a question and passage: in front of a stored note, which is how e5 was trained. The quantized file is the one that is downloaded, deployed and measured.

Also need ONNX Runtime (libonnxruntime.so):

curl -sL "https://github.com/microsoft/onnxruntime/releases/download/v1.15.1/onnxruntime-linux-x64-1.15.1.tgz" | tar xz
sudo cp onnxruntime-linux-x64-1.15.1/lib/libonnxruntime.so* /usr/local/lib/

Configure in deploy/mavend.json:

"voice": {
  "embedder": {
    "model_path": "models/embedder/multilingual-e5-small/model_quantized.onnx",
    "tokenizer_path": "models/embedder/multilingual-e5-small/tokenizer.json",
    "lib_path": "/usr/local/lib/libonnxruntime.so"
  }
}

Without the embedder block, the daemon uses HashEmbedder (works, but weak on Russian recall — you may see many "clarify" responses).

Qwen3 resident model for router + phraser

The target daemon uses the locally trained Qwen3-1.7B checkpoint for both routing and phrasing. Training is Qwen3 Base → RU CPT → joint persona/router SFT → merged GGUF, and is still in flight (#122) — until it lands, the deployed resident model is stock Qwen3.5-0.8B (Q4_K_M), see deploy/mavend.json. Without a configured model, StubPhraser plus the classifier remain the deterministic floor.

During training, use the runbook in docs/plans/2026-07-18-qwen3-resident-training-eval.md. After the decision gate and SFT pass, copy the merged GGUF into the mounted model directory and set:

"phraser": {
  "model_path": "/opt/maven/models/llm/Qwen3-Maven-1.7B-Q8_0.gguf",
  "bin_path": "llama-server",
  "n_gpu_layers": 99,
  "n_ctx": 2048
}

Configure in deploy/mavend.json — the phraser block points at this model and the daemon spawns llama-server as a subprocess. The router and replier use the same llama-server via the shared internal/llm client.

Telegram tokens are read from deploy/telegram.env (gitignored), expanded via ${VAR} in the JSON config.

Routing is Qwen-first with classifier fallback. The LLM router runs after stage-0 (exact-match grammar) and before the classifier cascade. On any error or parse failure, the classifier handles the utterance — the turn never breaks on the model.

Web UI conventions

  • All server-rendered pages share cmd/mavweb/static/ui.css (served at /ui.css) and the nav template partial (navHTML in cmd/mavweb/main.go, invoked as {{template "nav" "<active-page>"}}). New pages must link both — no per-page inline <style> beyond true one-offs.
  • Wrap every table in <div class=scroll> so wide data pans on a phone instead of breaking the layout.