65 lines
3.2 KiB
Markdown
65 lines
3.2 KiB
Markdown
# RU-CPT corpus + training — runbook
|
||
|
||
Scripts for building our own RU-native base (continued pretraining of
|
||
Qwen3-1.7B-Base). Full plan: `Maven/docs/plans/2026-07-11-ru-cpt-base.md`.
|
||
|
||
## One-time setup
|
||
|
||
```bash
|
||
pip install datasets ftfy datasketch fasttext openai peft accelerate
|
||
wget -O data/lid.176.bin \
|
||
https://dl.fbaipublicfiles.com/fasttext/supervised-models/lid.176.bin
|
||
huggingface-cli login # accept CulturaX (gated) license
|
||
```
|
||
`fasttext` build fails → `pip install fasttext-wheel`. Move the model → `export FASTTEXT_LID=/path`.
|
||
|
||
## Run order
|
||
|
||
```bash
|
||
# Phase 2 — corpus (build first; longest)
|
||
python corpus_fetch.py all
|
||
ROUTER_BASE_URL="https://your-router/v1" ROUTER_API_KEY="..." \
|
||
ROUTER_MODELS="m-a,m-b,m-c" python corpus_synth.py # ≤10% hard cap
|
||
MAVEN_LOGS=/path/to/dialogue.jsonl python corpus_logs.py # optional
|
||
python corpus_clean.py # ftfy+langID+dedup
|
||
python corpus_stats.py # DONE-CHECK, eyeball
|
||
python corpus_pack.py # → data/cpt_packed/
|
||
|
||
# Phase 3 — CPT (days; --resume to continue; CPT_PATH=dora if OOM)
|
||
HSA_OVERRIDE_GFX_VERSION=11.0.0 python train_cpt.py 2>&1 | tee cpt_run.log
|
||
|
||
# Phase 4 — eval + gate
|
||
python eval_cyrillic.py --model Qwen/Qwen3-1.7B-Base # baseline
|
||
python eval_cyrillic.py --model ./Qwen3-1.7B-ru-cpt # must beat baseline
|
||
|
||
# Phase 6-7 — persona LoRA (train_rocm.py @ ./Qwen3-1.7B-ru-cpt) → GGUF → Q8_0 → homesrv
|
||
```
|
||
|
||
## Gotchas (a fresh agent WILL hit these)
|
||
|
||
- **Run `corpus_clean.py` with NO args (all buckets in ONE invocation).** Dedup is
|
||
a single in-memory MinHash-LSH shared across buckets. Cleaning buckets in
|
||
separate runs = no cross-bucket dedup = duplicated data across sources.
|
||
- **Disk space.** `corpus_fetch.py` over-pulls ~1.5× and CulturaX streaming still
|
||
caches to `HF_DATASETS_CACHE` (`/mnt/D/.cache/...`). Expect tens of GB of raw +
|
||
cache. Check `df -h /mnt/D` before starting; `rm -rf data/cpt_raw` after packing.
|
||
- **Strict order.** fetch → clean → pack → train. `corpus_pack.py` reads
|
||
`data/cpt_clean/`; `train_cpt.py` reads `data/cpt_packed/`. Skipping a stage
|
||
silently trains on nothing (empty dataset → instant "done", garbage model).
|
||
- **`corpus_pack.py` prints the final token count — READ IT.** Must be ~300M ±50M.
|
||
Way under → a fetch failed or clean dropped too much; way over → over-pulled.
|
||
Do not train on a 20M-token corpus and expect CPT to work.
|
||
- **Verify the base repo id exists**: `Qwen/Qwen3-1.7B-Base`. If HF 404s, the
|
||
naming shifted — find the current Qwen3 ~1.7B *base* (non-Instruct) id.
|
||
- **Eval held-out set must NOT be training data.** `--eval-text` should point at
|
||
real RU text you set aside before packing, or the PPL number is a lie.
|
||
- **Re-running is safe.** fetch/clean/pack/synth all overwrite their outputs; CPT
|
||
resumes with `--resume`. No manual cleanup needed between retries.
|
||
|
||
## Hard rules (full list in the plan §9)
|
||
|
||
1. CPT corpus = raw text; persona `{response,mood}` is a separate Phase-6 dataset.
|
||
2. Synthetic ≤ 10%. 3. lr 1e-5, 1 epoch. 4. Base = Qwen3-1.7B-**Base**.
|
||
5. Never quantize before merge. 6. All sources through corpus_clean.py.
|
||
7. Don't skip a DONE-CHECK; if CPT loses to raw, stop (Phase 5 gate).
|