Files
model-training/llm/RU_CPT_RUNBOOK.md
T
2026-07-19 23:52:25 +04:00

3.2 KiB
Raw Blame History

RU-CPT corpus + training — runbook

Scripts for building our own RU-native base (continued pretraining of Qwen3-1.7B-Base). Full plan: Maven/docs/plans/2026-07-11-ru-cpt-base.md.

One-time setup

pip install datasets ftfy datasketch fasttext openai peft accelerate
wget -O data/lid.176.bin \
  https://dl.fbaipublicfiles.com/fasttext/supervised-models/lid.176.bin
huggingface-cli login          # accept CulturaX (gated) license

fasttext build fails → pip install fasttext-wheel. Move the model → export FASTTEXT_LID=/path.

Run order

# Phase 2 — corpus (build first; longest)
python corpus_fetch.py all
ROUTER_BASE_URL="https://your-router/v1" ROUTER_API_KEY="..." \
  ROUTER_MODELS="m-a,m-b,m-c" python corpus_synth.py     # ≤10% hard cap
MAVEN_LOGS=/path/to/dialogue.jsonl python corpus_logs.py # optional
python corpus_clean.py                                   # ftfy+langID+dedup
python corpus_stats.py                                   # DONE-CHECK, eyeball
python corpus_pack.py                                    # → data/cpt_packed/

# Phase 3 — CPT (days; --resume to continue; CPT_PATH=dora if OOM)
HSA_OVERRIDE_GFX_VERSION=11.0.0 python train_cpt.py 2>&1 | tee cpt_run.log

# Phase 4 — eval + gate
python eval_cyrillic.py --model Qwen/Qwen3-1.7B-Base     # baseline
python eval_cyrillic.py --model ./Qwen3-1.7B-ru-cpt      # must beat baseline

# Phase 6-7 — persona LoRA (train_rocm.py @ ./Qwen3-1.7B-ru-cpt) → GGUF → Q8_0 → homesrv

Gotchas (a fresh agent WILL hit these)

  • Run corpus_clean.py with NO args (all buckets in ONE invocation). Dedup is a single in-memory MinHash-LSH shared across buckets. Cleaning buckets in separate runs = no cross-bucket dedup = duplicated data across sources.
  • Disk space. corpus_fetch.py over-pulls ~1.5× and CulturaX streaming still caches to HF_DATASETS_CACHE (/mnt/D/.cache/...). Expect tens of GB of raw + cache. Check df -h /mnt/D before starting; rm -rf data/cpt_raw after packing.
  • Strict order. fetch → clean → pack → train. corpus_pack.py reads data/cpt_clean/; train_cpt.py reads data/cpt_packed/. Skipping a stage silently trains on nothing (empty dataset → instant "done", garbage model).
  • corpus_pack.py prints the final token count — READ IT. Must be ~300M ±50M. Way under → a fetch failed or clean dropped too much; way over → over-pulled. Do not train on a 20M-token corpus and expect CPT to work.
  • Verify the base repo id exists: Qwen/Qwen3-1.7B-Base. If HF 404s, the naming shifted — find the current Qwen3 ~1.7B base (non-Instruct) id.
  • Eval held-out set must NOT be training data. --eval-text should point at real RU text you set aside before packing, or the PPL number is a lie.
  • Re-running is safe. fetch/clean/pack/synth all overwrite their outputs; CPT resumes with --resume. No manual cleanup needed between retries.

Hard rules (full list in the plan §9)

  1. CPT corpus = raw text; persona {response,mood} is a separate Phase-6 dataset.
  2. Synthetic ≤ 10%. 3. lr 1e-5, 1 epoch. 4. Base = Qwen3-1.7B-Base.
  3. Never quantize before merge. 6. All sources through corpus_clean.py.
  4. Don't skip a DONE-CHECK; if CPT loses to raw, stop (Phase 5 gate).