This commit is contained in:
2026-07-19 23:52:25 +04:00
commit 99b06b0468
155 changed files with 3808611 additions and 0 deletions
+64
View File
@@ -0,0 +1,64 @@
# RU-CPT corpus + training — runbook
Scripts for building our own RU-native base (continued pretraining of
Qwen3-1.7B-Base). Full plan: `Maven/docs/plans/2026-07-11-ru-cpt-base.md`.
## One-time setup
```bash
pip install datasets ftfy datasketch fasttext openai peft accelerate
wget -O data/lid.176.bin \
https://dl.fbaipublicfiles.com/fasttext/supervised-models/lid.176.bin
huggingface-cli login # accept CulturaX (gated) license
```
`fasttext` build fails → `pip install fasttext-wheel`. Move the model → `export FASTTEXT_LID=/path`.
## Run order
```bash
# Phase 2 — corpus (build first; longest)
python corpus_fetch.py all
ROUTER_BASE_URL="https://your-router/v1" ROUTER_API_KEY="..." \
ROUTER_MODELS="m-a,m-b,m-c" python corpus_synth.py # ≤10% hard cap
MAVEN_LOGS=/path/to/dialogue.jsonl python corpus_logs.py # optional
python corpus_clean.py # ftfy+langID+dedup
python corpus_stats.py # DONE-CHECK, eyeball
python corpus_pack.py # → data/cpt_packed/
# Phase 3 — CPT (days; --resume to continue; CPT_PATH=dora if OOM)
HSA_OVERRIDE_GFX_VERSION=11.0.0 python train_cpt.py 2>&1 | tee cpt_run.log
# Phase 4 — eval + gate
python eval_cyrillic.py --model Qwen/Qwen3-1.7B-Base # baseline
python eval_cyrillic.py --model ./Qwen3-1.7B-ru-cpt # must beat baseline
# Phase 6-7 — persona LoRA (train_rocm.py @ ./Qwen3-1.7B-ru-cpt) → GGUF → Q8_0 → homesrv
```
## Gotchas (a fresh agent WILL hit these)
- **Run `corpus_clean.py` with NO args (all buckets in ONE invocation).** Dedup is
a single in-memory MinHash-LSH shared across buckets. Cleaning buckets in
separate runs = no cross-bucket dedup = duplicated data across sources.
- **Disk space.** `corpus_fetch.py` over-pulls ~1.5× and CulturaX streaming still
caches to `HF_DATASETS_CACHE` (`/mnt/D/.cache/...`). Expect tens of GB of raw +
cache. Check `df -h /mnt/D` before starting; `rm -rf data/cpt_raw` after packing.
- **Strict order.** fetch → clean → pack → train. `corpus_pack.py` reads
`data/cpt_clean/`; `train_cpt.py` reads `data/cpt_packed/`. Skipping a stage
silently trains on nothing (empty dataset → instant "done", garbage model).
- **`corpus_pack.py` prints the final token count — READ IT.** Must be ~300M ±50M.
Way under → a fetch failed or clean dropped too much; way over → over-pulled.
Do not train on a 20M-token corpus and expect CPT to work.
- **Verify the base repo id exists**: `Qwen/Qwen3-1.7B-Base`. If HF 404s, the
naming shifted — find the current Qwen3 ~1.7B *base* (non-Instruct) id.
- **Eval held-out set must NOT be training data.** `--eval-text` should point at
real RU text you set aside before packing, or the PPL number is a lie.
- **Re-running is safe.** fetch/clean/pack/synth all overwrite their outputs; CPT
resumes with `--resume`. No manual cleanup needed between retries.
## Hard rules (full list in the plan §9)
1. CPT corpus = raw text; persona `{response,mood}` is a separate Phase-6 dataset.
2. Synthetic ≤ 10%. 3. lr 1e-5, 1 epoch. 4. Base = Qwen3-1.7B-**Base**.
5. Never quantize before merge. 6. All sources through corpus_clean.py.
7. Don't skip a DONE-CHECK; if CPT loses to raw, stop (Phase 5 gate).