Reconcile the deployed resident model documentation (V-407)
This commit is contained in:
@@ -135,37 +135,41 @@ Russian recall — you may see many "clarify" responses).
|
||||
|
||||
## Qwen3 resident model for router + phraser
|
||||
|
||||
The target daemon uses the locally trained Qwen3-1.7B checkpoint for both
|
||||
routing and phrasing. Training is Qwen3 Base → RU CPT → joint persona/router
|
||||
SFT → merged GGUF, and is still in flight (#122) — until it lands, the deployed
|
||||
resident model is stock **Qwen3.5-0.8B** (`Q4_K_M`), see `deploy/mavend.json`.
|
||||
The deployed resident model is stock **Qwen3-1.7B** (`UD-Q4_K_XL`), a Thinking
|
||||
variant at `n_ctx` 4096. `CLAUDE.md` carries the rule on which models qualify.
|
||||
Without a configured model, `StubPhraser` plus the classifier remain the
|
||||
deterministic floor.
|
||||
|
||||
During training, use the runbook in
|
||||
`docs/plans/2026-07-18-qwen3-resident-training-eval.md`. After the decision gate
|
||||
and SFT pass, copy the merged GGUF into the mounted model directory and set:
|
||||
A locally trained Qwen3-1.7B checkpoint is still in flight (V-122). Training
|
||||
runs Qwen3 Base, then RU CPT, then joint persona and router SFT, then a merged
|
||||
GGUF. The
|
||||
runbook is `docs/plans/2026-07-18-qwen3-resident-training-eval.md`. After the
|
||||
decision gate and SFT pass, copy the merged GGUF into the mounted model
|
||||
directory and point `model_path` at it.
|
||||
|
||||
**Configure in `deploy/mavend.json`.** This is the deployed `phraser` block:
|
||||
|
||||
```json
|
||||
"phraser": {
|
||||
"model_path": "/opt/maven/models/llm/Qwen3-Maven-1.7B-Q8_0.gguf",
|
||||
"model_path": "/opt/maven/models/llm/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf",
|
||||
"bin_path": "llama-server",
|
||||
"n_gpu_layers": 99,
|
||||
"n_ctx": 2048
|
||||
"n_ctx": 4096,
|
||||
"cache_ram_mib": 512,
|
||||
"timeout": "60s"
|
||||
}
|
||||
```
|
||||
|
||||
**Configure in `deploy/mavend.json`** — the `phraser` block points at this
|
||||
model and the daemon spawns `llama-server` as a subprocess. The router and
|
||||
replier use the same llama-server via the shared `internal/llm` client.
|
||||
The daemon spawns `llama-server` as a subprocess. The router and replier reach
|
||||
that one server through the shared `internal/llm` client. Model files live in
|
||||
`/mnt/hdd1/llms`, bind-mounted over `models/llm/`, so a gguf sitting in the repo
|
||||
is loaded by nothing.
|
||||
|
||||
Telegram tokens are read from `deploy/telegram.env` (gitignored), expanded
|
||||
via `${VAR}` in the JSON config.
|
||||
|
||||
**Routing is Qwen-first** with classifier fallback. The LLM router runs
|
||||
after stage-0 (exact-match grammar) and before the classifier cascade. On any
|
||||
error or parse failure, the classifier handles the utterance — the turn never
|
||||
breaks on the model.
|
||||
The cascade order, and which stage may decline to the next, is in
|
||||
`docs/routing.md`. It is not restated here.
|
||||
|
||||
## Web UI conventions
|
||||
|
||||
|
||||
Reference in New Issue
Block a user