diff --git a/AGENTS.md b/AGENTS.md index bccb2b3..2b53534 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -135,37 +135,41 @@ Russian recall — you may see many "clarify" responses). ## Qwen3 resident model for router + phraser -The target daemon uses the locally trained Qwen3-1.7B checkpoint for both -routing and phrasing. Training is Qwen3 Base → RU CPT → joint persona/router -SFT → merged GGUF, and is still in flight (#122) — until it lands, the deployed -resident model is stock **Qwen3.5-0.8B** (`Q4_K_M`), see `deploy/mavend.json`. +The deployed resident model is stock **Qwen3-1.7B** (`UD-Q4_K_XL`), a Thinking +variant at `n_ctx` 4096. `CLAUDE.md` carries the rule on which models qualify. Without a configured model, `StubPhraser` plus the classifier remain the deterministic floor. -During training, use the runbook in -`docs/plans/2026-07-18-qwen3-resident-training-eval.md`. After the decision gate -and SFT pass, copy the merged GGUF into the mounted model directory and set: +A locally trained Qwen3-1.7B checkpoint is still in flight (V-122). Training +runs Qwen3 Base, then RU CPT, then joint persona and router SFT, then a merged +GGUF. The +runbook is `docs/plans/2026-07-18-qwen3-resident-training-eval.md`. After the +decision gate and SFT pass, copy the merged GGUF into the mounted model +directory and point `model_path` at it. + +**Configure in `deploy/mavend.json`.** This is the deployed `phraser` block: ```json "phraser": { - "model_path": "/opt/maven/models/llm/Qwen3-Maven-1.7B-Q8_0.gguf", + "model_path": "/opt/maven/models/llm/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf", "bin_path": "llama-server", "n_gpu_layers": 99, - "n_ctx": 2048 + "n_ctx": 4096, + "cache_ram_mib": 512, + "timeout": "60s" } ``` -**Configure in `deploy/mavend.json`** — the `phraser` block points at this -model and the daemon spawns `llama-server` as a subprocess. The router and -replier use the same llama-server via the shared `internal/llm` client. +The daemon spawns `llama-server` as a subprocess. The router and replier reach +that one server through the shared `internal/llm` client. Model files live in +`/mnt/hdd1/llms`, bind-mounted over `models/llm/`, so a gguf sitting in the repo +is loaded by nothing. Telegram tokens are read from `deploy/telegram.env` (gitignored), expanded via `${VAR}` in the JSON config. -**Routing is Qwen-first** with classifier fallback. The LLM router runs -after stage-0 (exact-match grammar) and before the classifier cascade. On any -error or parse failure, the classifier handles the utterance — the turn never -breaks on the model. +The cascade order, and which stage may decline to the next, is in +`docs/routing.md`. It is not restated here. ## Web UI conventions