Reconcile the deployed resident model documentation (V-407)

This commit is contained in:
2026-08-13 01:27:13 +04:00
parent 8035a317d2
commit f957a3ad13
+20 -16
View File
@@ -135,37 +135,41 @@ Russian recall — you may see many "clarify" responses).
## Qwen3 resident model for router + phraser
The target daemon uses the locally trained Qwen3-1.7B checkpoint for both
routing and phrasing. Training is Qwen3 Base → RU CPT → joint persona/router
SFT → merged GGUF, and is still in flight (#122) — until it lands, the deployed
resident model is stock **Qwen3.5-0.8B** (`Q4_K_M`), see `deploy/mavend.json`.
The deployed resident model is stock **Qwen3-1.7B** (`UD-Q4_K_XL`), a Thinking
variant at `n_ctx` 4096. `CLAUDE.md` carries the rule on which models qualify.
Without a configured model, `StubPhraser` plus the classifier remain the
deterministic floor.
During training, use the runbook in
`docs/plans/2026-07-18-qwen3-resident-training-eval.md`. After the decision gate
and SFT pass, copy the merged GGUF into the mounted model directory and set:
A locally trained Qwen3-1.7B checkpoint is still in flight (V-122). Training
runs Qwen3 Base, then RU CPT, then joint persona and router SFT, then a merged
GGUF. The
runbook is `docs/plans/2026-07-18-qwen3-resident-training-eval.md`. After the
decision gate and SFT pass, copy the merged GGUF into the mounted model
directory and point `model_path` at it.
**Configure in `deploy/mavend.json`.** This is the deployed `phraser` block:
```json
"phraser": {
"model_path": "/opt/maven/models/llm/Qwen3-Maven-1.7B-Q8_0.gguf",
"model_path": "/opt/maven/models/llm/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf",
"bin_path": "llama-server",
"n_gpu_layers": 99,
"n_ctx": 2048
"n_ctx": 4096,
"cache_ram_mib": 512,
"timeout": "60s"
}
```
**Configure in `deploy/mavend.json`** — the `phraser` block points at this
model and the daemon spawns `llama-server` as a subprocess. The router and
replier use the same llama-server via the shared `internal/llm` client.
The daemon spawns `llama-server` as a subprocess. The router and replier reach
that one server through the shared `internal/llm` client. Model files live in
`/mnt/hdd1/llms`, bind-mounted over `models/llm/`, so a gguf sitting in the repo
is loaded by nothing.
Telegram tokens are read from `deploy/telegram.env` (gitignored), expanded
via `${VAR}` in the JSON config.
**Routing is Qwen-first** with classifier fallback. The LLM router runs
after stage-0 (exact-match grammar) and before the classifier cascade. On any
error or parse failure, the classifier handles the utterance — the turn never
breaks on the model.
The cascade order, and which stage may decline to the next, is in
`docs/routing.md`. It is not restated here.
## Web UI conventions