Files
Maven/docs/plans/2026-07-18-qwen3-resident-training-eval.md
T
kami 5fe8f228c1 feat(mavweb): /ecosystem page consuming Nexus/Praxis/Hexis + shell fixes
Add a read-only /ecosystem page that consumes the sibling services'
JSON APIs (Nexus entities, Praxis attention, Hexis capabilities),
fetched concurrently with honest per-panel error states. Siblings stay
headless — mavweb is their human surface (arch §16). Wired via mavweb
-nexus/-praxis/-hexis flags; mavweb joins the ecosystem compose network.

Fix mobile horizontal overflow across all pages: .content is a flex
child with default min-width:auto, so it refused to shrink below the
tables' intrinsic width. min-width:0 lets wide tables pan inside .scroll
instead of dragging the page sideways. Verified via CDP geometry check
(scrollWidth === clientWidth at 430px).

Also includes in-progress Ethos UI redesign, ecosystem deploy compose,
and planning docs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 22:04:23 +04:00

158 lines
5.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Plan — Qwen3-1.7B resident model: audit, evaluation, SFT and gate
> **Canonical model plan as of 2026-07-18.** Qwen3-1.7B is the one resident
> router/phraser. Larger reasoners and custom Piper training are deferred until
> the main Maven features are complete.
## Locked architecture
```
Qwen3-1.7B-Base
→ full-weight Russian CPT
→ raw-vs-CPT decision gate
→ one balanced persona + route LoRA
→ contract evaluation
→ merge → GGUF Q8_0 → homesrv
```
The route system prompt selects a JSON-array action contract. The persona
system prompt selects `{response,mood}`. Metrics are reported independently so
joint-training interference is visible. If one task regresses, use separate
adapters before considering a merged multi-task checkpoint.
## Current evidence
- Packed corpus: 129,221 × 2,048 = **264,644,608 tokens**.
- Materialized character mix: CulturaX 69.407%, Wikipedia 17.991%, books
12.602%. Planned synthetic/log/conversational buckets are absent from this run.
- All 64 sampled Arrow rows had the right length and valid token IDs.
- The audit found 247 CulturaX documents with glued `<phone>` placeholders and
Cyrillic text. This run continues; fix placeholder spacing before any rebuild.
- Full-weight Adafactor training reached checkpoint 1,000/8,077 (12.38%). Loss
and gradient norms are finite, so this checkpoint is the full-weight smoke
test. Do not launch a competing smoke job.
Machine-readable audit:
`/home/kami/Programs/esp32-whisper-fine-tune/llm/CPT_CORPUS_AUDIT.json`.
## Tools
All commands run from `/home/kami/Programs/esp32-whisper-fine-tune` using its
`.venv`.
```bash
# Reproduce the read-only corpus/checkpoint audit.
./.venv/bin/python llm/audit_cpt_corpus.py \
--output llm/CPT_CORPUS_AUDIT.json
# Fetch pinned, human-written UD RU/EN test sets and their hashes.
./.venv/bin/python llm/prepare_eval_data.py
# Prompt parity and held-out route labels.
./.venv/bin/python llm/check_prompt_parity.py
./.venv/bin/python llm/build_route_eval.py
```
## Baseline and post-CPT evaluation
The evaluation is deterministic and uses token-weighted perplexity. Preserve
the raw result; do not regenerate it with different limits when comparing CPT.
```bash
./.venv/bin/python llm/eval_cyrillic.py \
--model Qwen/Qwen3-1.7B-Base --limit 128 --device cuda \
--output llm/data/eval/raw_qwen3_1.7b.json
# After CPT finishes:
./.venv/bin/python llm/eval_cyrillic.py \
--model llm/Qwen3-1.7B-ru-cpt --limit 128 --device cuda \
--output llm/data/eval/cpt_qwen3_1.7b.json
```
For a publication-quality comparison, repeat both with `--limit 0`; the quick
128-sentence pair is the operational gate and must use identical arguments.
## Exact CPT decision gate
```bash
./.venv/bin/python llm/decision_gate.py \
--raw llm/data/eval/raw_qwen3_1.7b.json \
--cpt llm/data/eval/cpt_qwen3_1.7b.json
```
Exit 0 means all conditions passed:
- RU perplexity improves by at least 2%;
- EN perplexity regresses by no more than 10%;
- RU generation validity does not decline;
- deterministic English probes remain English.
- mean repeated 4-gram rate stays below 15%, or within five percentage points
of the raw baseline when the baseline itself is worse.
Exit 1 blocks SFT. Inspect the JSON check map before changing a threshold. A
threshold change is a documented architecture decision, not a convenient rerun.
Checkpoint 1,000 is an informative interim result: RU PPL improved 4.82% and EN
PPL regressed only 2.20%, but repeated 4-grams rose from 8.56% to 25.90%.
Therefore it passes the language-loss checks but **fails the complete gate**.
This is not a stop signal at 12.38% of training; it is a regression to watch at
the final checkpoint and a reason the gate includes generation degeneration.
## Resume CPT
The latest complete checkpoint is selected by Transformers:
```bash
cd /home/kami/Programs/esp32-whisper-fine-tune/llm
HSA_OVERRIDE_GFX_VERSION=11.0.0 ../.venv/bin/python train_cpt.py --resume \
2>&1 | tee -a cpt_run.log
```
Confirm there is only one `train_cpt.py` process before resuming. Completion is
step 8,077 and must produce the final tokenizer/model files at
`llm/Qwen3-1.7B-ru-cpt/`, not only checkpoint directories.
## Joint Qwen3 SFT
Generate route training data first. `route_train.jsonl` is currently a blocker;
the held-out `route_eval.jsonl` already exists and must never be merged into it.
```bash
cd /home/kami/Programs/esp32-whisper-fine-tune
./.venv/bin/python llm/gen_route_data.py
./.venv/bin/python llm/train_rocm.py --check-data
./.venv/bin/python llm/train_rocm.py
```
The rewritten trainer:
- loads the completed CPT base;
- consumes explicit persona and route train/eval files;
- balances the two tasks by oversampling only within the training set;
- renders Qwen3's own chat template with thinking disabled;
- masks loss before the assistant continuation without hard-coded ChatML IDs;
- supports deterministic seeds, resume, early stopping and best-checkpoint load;
- writes a training manifest.
After training:
```bash
./.venv/bin/python llm/eval_contracts.py \
--base llm/Qwen3-1.7B-ru-cpt \
--adapter llm/Qwen3-1.7B-maven-sft \
--output llm/data/eval/qwen3_contracts.json
```
Initial deploy gates: persona JSON validity ≥99%, route JSON validity ≥99%,
route intent sequence exact ≥90%, action intent accuracy ≥95%, slot exact ≥85%.
Mood accuracy is diagnostic until the persona evaluation set is manually
quality-reviewed; the existing examples contain stale technical answers.
## Deferred
- Piper/phoneme training.
- On-demand larger reasoner.
- Model hot-swap and the broad capability plans.
These resume after the main Maven feature set is complete.