Files
Maven/docs/plans/2026-07-11-ru-cpt-base.md
T
kami 5fe8f228c1 feat(mavweb): /ecosystem page consuming Nexus/Praxis/Hexis + shell fixes
Add a read-only /ecosystem page that consumes the sibling services'
JSON APIs (Nexus entities, Praxis attention, Hexis capabilities),
fetched concurrently with honest per-panel error states. Siblings stay
headless — mavweb is their human surface (arch §16). Wired via mavweb
-nexus/-praxis/-hexis flags; mavweb joins the ecosystem compose network.

Fix mobile horizontal overflow across all pages: .content is a flex
child with default min-width:auto, so it refused to shrink below the
tables' intrinsic width. min-width:0 lets wide tables pan inside .scroll
instead of dragging the page sideways. Verified via CDP geometry check
(scrollWidth === clientWidth at 430px).

Also includes in-progress Ethos UI redesign, ecosystem deploy compose,
and planning docs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 22:04:23 +04:00

20 KiB
Raw Blame History

Plan — DIY RU-native base via Continued Pretraining (CPT)

Execution update, 2026-07-18: corpus packing and full-weight smoke are complete; training reached step 1000/8077. The materialized corpus is 264.6M tokens from CulturaX/Wikipedia/books only. Evaluation, the exact decision gate and Qwen3 joint SFT are now canonical in 2026-07-18-qwen3-resident-training-eval.md.

Goal: Build our own RU-native base model instead of using Vikhr. Take a clean newer base (Qwen3-1.7B-Base), continue-pretrain it on a curated Russian corpus so it spells Cyrillic natively, then run the existing persona-LoRA → merge → GGUF → deploy pipeline on top. We own the recipe, the corpus, and the checkpoint — reproducible, no dependency on a third-party RU repo.

This is a big project. The corpus is 80% of the work. Read the whole file before starting. Do phases in order. Each phase has a DONE-CHECK — do not advance until it passes.


0. Context an agent picking this up must know

Two machines — never conflate (from CLAUDE.md):

Role Box Specs Constraint
Train this workstation RX 7900 GRE, gfx1100 (RDNA3), 16GB VRAM, ROCm 7.2 / torch 2.10 HIP, cuda.is_available()==True CPT of 1.7B is tight — see Phase 3 VRAM math
Deploy homesrv kami@192.168.1.104 Ryzen 5 5600U CPU, 13GB RAM already swapping (~8.8GB avail), Ubuntu 24.04 CPU-only llama.cpp inference; model must stay ≤1.7B and every 200MB matters

Working directories:

  • Training scripts & data: /home/kami/Programs/esp32-whisper-fine-tune/llm/
    • existing: train_rocm.py (persona LoRA, ChatML), llama-eval-test.py, data/
    • new files from this plan also go here.
  • Daemon (consumes final GGUF): /mnt/server/home/kami/apps/Maven/ (this repo).

Output contract (from CLAUDE.md, Decision B, locked 2026-07-11): the daemon parses {"response":"...","mood":"..."} in all 4 consumers. This only matters for the persona LoRA stage (Phase 6), NOT for CPT. CPT is raw next-token text, no JSON, no chat format.

Why CPT and not just a LoRA (from CLAUDE.md): a LoRA cannot add spelling knowledge the base lacks — it only steers format/persona/mood/tools. CPT changes the base weights themselves, which is the only lever that can add genuine RU language competence. That is exactly why this is more than a LoRA project.

ROCm gotchas (inherit from train_rocm.py):

  • attn_implementation="eager" (flash-attn ROCm wheels flaky).
  • HSA_OVERRIDE_GFX_VERSION=11.0.0 in env (gfx1100).
  • No bitsandbytes — ROCm build is flaky. This forbids 8-bit Adam; see Phase 3.

1. Pipeline overview (the whole journey)

Phase 2  Corpus build        →  clean packed RU token blocks (.arrow/.bin, ~300M tok)
Phase 3  train_cpt.py        →  CPT'd Qwen3-1.7B-Base (bf16 HF checkpoint)
Phase 4  eval                →  Cyrillic + JSON validator; CPT base vs raw base numbers
Phase 5  DECISION GATE       →  did CPT beat raw? if not, stop / re-corpus
Phase 6  persona LoRA        →  existing train_rocm.py on the CPT'd base (ChatML, {response,mood})
Phase 7  merge → GGUF → Q8_0 →  deploy to homesrv

Build order is not the run order: build the corpus first (Phase 2) because it gates everything and takes the longest.


2. Corpus — the core deliverable

Target: ~300M tokens (Qwen3 tokenizer), curated. Rationale: <100M won't move the needle vs Qwen3's already-strong multilingual; >1B is weeks of wall-clock on one 16GB card. 300M is the "nudge, don't re-pretrain" sweet spot for 1 epoch.

Two token sources, combined:

  • A. Open datasets (the bulk, real human RU text).
  • B. Router-generated synthetic (fills domain/persona/conversational gaps the open corpora lack — smart-home commands, assistant dialogue, Maven's domain). The user's router for free inference providers generates this.

2.1 Corpus mix table (sums to ~300M tokens)

# Bucket Share Tokens Source Why
B1 Cleaned web RU 55% ~165M uonlp/CulturaX (ru) — already cleaned + deduped backbone, broad RU
B2 Encyclopedic 15% ~45M wikimedia/wikipedia config 20231101.ru clean factual RU, correct spelling
B3 Books / long-form 10% ~30M IlyaGusev/librusec (or subset) rich morphology, endings
B4 Conversational 8% ~24M Helsinki-NLP/open_subtitles (ru) spoken register — matches voice assistant
B5 Router-synthetic 10% ~30M generated (§2.4) domain/persona/smart-home RU, absent from open sets
B6 Maven real logs 2% ~6M mine daemon dialogue history / logs exact deploy-domain match

Cap B5 (synthetic) at 10%. Synthetic-heavy CPT causes model collapse / degeneration. The bulk MUST be real human text. This cap is a hard rule.

2.2 Exact acquisition — open datasets (Bucket A: B1B4)

All via HuggingFace datasets, streaming to avoid disk blowup. Some are gated (CulturaX, OSCAR) — accept the license on the HF model page while logged in (huggingface-cli login) first.

# corpus_fetch.py  — run per bucket, writes raw .txt.gz shards to data/cpt_raw/<bucket>/
from datasets import load_dataset
# B1 CulturaX (gated — accept license first)
ds = load_dataset("uonlp/CulturaX", "ru", split="train", streaming=True)
# B2 Wikipedia
ds = load_dataset("wikimedia/wikipedia", "20231101.ru", split="train", streaming=True)
# B3 Books
ds = load_dataset("IlyaGusev/librusec", split="train", streaming=True)
# B4 Subtitles
ds = load_dataset("Helsinki-NLP/open_subtitles", lang1="en", lang2="ru", split="train", streaming=True)

Pull enough raw text per bucket to survive filtering (filtering drops ~2040%); over-pull ~1.5× the target token count, then trim after cleaning. Text field differs per dataset (text, content, translation['ru']) — normalize each to a single text string when writing shards.

Fallbacks if a source is unavailable: B1 → allenai/c4 (ru) or oscar-corpus/OSCAR-2301 (ru); B3 → IlyaGusev/gazeta / Taiga corpus; B4 → Helsinki-NLP/tatoeba (ru pairs).

2.3 Cleaning pipeline (mandatory — dirty corpus = dirty base)

One script corpus_clean.py, applied to every raw shard, in this order:

  1. Unicode fixftfy.fix_text (repairs mojibake, normalizes to NFC).
  2. Language ID — fastText lid.176.bin; drop any doc where top lang != ru or p(ru) < 0.65. (Download lid.176.bin from fastText releases.)
  3. Mixed-script rejection — regex: reject any word containing both Cyrillic and Latin letters ([а-яё].*[a-z]|[a-z].*[а-яё] case-insensitive per word). This is the exact homoglyph problem CLAUDE.md names. Keep pure-Latin words (URLs/names) but drop homoglyph-contaminated ones.
  4. Boilerplate/length — drop docs <200 chars or >50k chars; strip HTML tags, nav junk, repeated whitespace; drop docs with >30% non-letter chars.
  5. Dedup — MinHash LSH via datasketch (MinHashLSH, threshold 0.8, 128 perms) across ALL buckets combined. Near-dup docs collapse to one. This is the single biggest quality lever — do not skip.
  6. PII light-scrub (optional) — regex phone/email → placeholder, since B6 has real logs.

Emit cleaned docs as JSONL {"text": "..."} in data/cpt_clean/.

DONE-CHECK 2.3: run corpus_stats.py → per-bucket doc count, char count, and a 100-doc random sample printed for eyeball review. No mixed-script words in the sample. If sample looks dirty, fix the filter before proceeding.

2.4 Router-synthetic generation (Bucket B5 — the user's router)

Purpose: generate ~30M tokens of clean RU text in registers open corpora lack — smart-home/assistant domain, natural conversational RU, Maven-adjacent topics. This is CPT text (raw prose/dialogue), NOT the {response,mood} persona data (that's Phase 6, a different dataset).

Protocol (corpus_synth.py, calls the free-provider router):

  1. Topic seed list — reuse data/topics.txt (already exists, ~46KB of topics)
    • add smart-home / daily-assistant / RU-domestic topics.
  2. Prompt template (per topic), ask the router model to produce long-form natural Russian prose or dialogue on the topic — NOT JSON, NOT persona, NOT English. e.g. "Напиши развёрнутый естественный текст на русском языке на тему: {topic}. 400–800 слов, живой разговорный стиль."
  3. Diversify: vary style knob (formal/casual/dialogue/monologue), vary length, rotate across the router's providers/models so it's not one model's fingerprint (fingerprint uniformity → collapse risk).
  4. Route every generation through the SAME cleaning pipeline (§2.3) — language ID, mixed-script rejection, dedup against the real corpus too. Synthetic that fails Cyrillic cleanliness is exactly what we must not train on.
  5. Stop at ~30M tokens. Enforce the 10% cap.

DONE-CHECK 2.4: synthetic passes the same stats + eyeball sample as 2.3; token count ≤ 10% of total corpus; provider diversity logged.

2.5 Maven real logs (Bucket B6)

Mine daemon dialogue history / logs for real RU user+assistant utterances. Strip to plain text, run through §2.3. Small (~6M tok) but highest domain value. Source location: check mavend logs / any dialogue history store in the repo.

2.6 Tokenize + pack (final corpus artifact)

corpus_pack.py:

  • Load Qwen3-1.7B-Base tokenizer (transformers.AutoTokenizer).
  • Concatenate all cleaned JSONL, tokenize, append EOS between docs, pack into fixed 2048-token blocks (drop the ragged tail). No padding — packing is efficient for CPT.
  • Save as a HF datasets Arrow dir data/cpt_packed/ (columns: input_ids).
  • Print final total token count — must be ~300M ± 50M. Adjust bucket pulls if off.

DONE-CHECK 2.6: data/cpt_packed/ loads; total tokens logged; a decoded random block reads as clean Russian.


3. train_cpt.py — continued pretraining (fits 16GB)

New file next to train_rocm.py. Inherits the ROCm setup from it.

3.1 VRAM reality (why the config is what it is)

Full-weight CPT of 1.7B with plain Adam ≈ 2124GB → does not fit 16GB. We can't use 8-bit Adam (no bitsandbytes on ROCm). Fitting strategy, in preference order — try A first, fall back to B:

  • A. Full-weight + adafactor — Adafactor has no momentum states, cutting the optimizer-state VRAM roughly in half. Full-weight bf16 + adafactor + gradient checkpointing + batch 1 × grad-accum should fit ~16GB. This is the real "change the base weights" path. Preferred.
  • B. High-rank DoRA on all linear layers — if A OOMs or throughput is unbearable. Fits trivially. Weaker for language injection but still adapts more than a small LoRA. Use r=64, α=128, target all linear, use_dora=True.

3.2 Config

# train_cpt.py essentials
import os; os.environ["HSA_OVERRIDE_GFX_VERSION"] = "11.0.0"
from transformers import (AutoModelForCausalLM, AutoTokenizer,
                          TrainingArguments, Trainer,
                          DataCollatorForLanguageModeling)

MODEL = "Qwen/Qwen3-1.7B-Base"          # BASE, not Instruct
model = AutoModelForCausalLM.from_pretrained(
    MODEL, torch_dtype="bfloat16", attn_implementation="eager")
model.gradient_checkpointing_enable()
model.config.use_cache = False

args = TrainingArguments(
    output_dir="./Qwen3-1.7B-ru-cpt",
    per_device_train_batch_size=1,
    gradient_accumulation_steps=16,      # eff. batch 16 × 2048 tok
    num_train_epochs=1,                  # ONE pass — more risks forgetting
    learning_rate=1e-5,                  # LOW — nudging pretrained weights
    lr_scheduler_type="cosine",
    warmup_ratio=0.03,
    optim="adafactor",                   # path A; no bitsandbytes
    bf16=True,
    gradient_checkpointing=True,
    logging_steps=20,
    save_steps=500,                      # checkpoint often — this runs for DAYS
    save_total_limit=3,
    report_to="none",
)
collator = DataCollatorForLanguageModeling(tokenizer, mlm=False)  # causal LM
  • lr 1e-5 is critical. CPT lr must be ~10× lower than LoRA lr. High lr = catastrophic forgetting (model forgets English, reasoning, everything). If the eval (Phase 4) shows English/tool ability collapsed, lr was too high.
  • 1 epoch. Multiple epochs over 300M tokens overfits and forgets.
  • Resume/checkpoint is mandatory — days-long job; support --resume_from_checkpoint.

3.3 Run

cd /home/kami/Programs/esp32-whisper-fine-tune/llm
HSA_OVERRIDE_GFX_VERSION=11.0.0 python train_cpt.py 2>&1 | tee cpt_run.log

DONE-CHECK 3: training completes 1 epoch without OOM; loss decreases and plateaus (not NaN, not flat-from-step-0); checkpoint saved to ./Qwen3-1.7B-ru-cpt/. If OOM → switch to path B (DoRA).


4. Eval — did CPT actually help?

eval_cyrillic.py now writes deterministic machine-readable results using pinned human-written RU and EN Universal Dependencies test sets.

Metrics (run on held-out human-written text — NOT synthetic, NOT training data):

  1. Cyrillic validity % — generate on RU prompts; % of outputs with zero mixed-script words and no homoglyph swaps (reuse §2.3 rule 3 as the checker).
  2. Perplexity on a held-out clean RU text set (lower = better RU fit).
  3. English-retention check — a few EN prompts; confirm the model still answers in coherent English (catches catastrophic forgetting).
  4. JSON-format sanity (light, for later) — can it produce {response,mood} when asked. (Real test is post-persona-LoRA.)

Run it three ways and tabulate:

Model Cyrillic valid % RU PPL EN retained?
raw Qwen3-1.7B-Base (baseline)
our CPT'd base (must beat raw)

DONE-CHECK 4: table filled with real numbers.


5. DECISION GATE

  • CPT base beats raw on Cyrillic %/PPL AND English retained → proceed to Phase 6.
  • No improvement → CPT didn't earn its compute. Options: (a) more/better corpus (usually the fix — dirty or too-small corpus), (b) tune lr, (c) accept raw Qwen3-1.7B as the base and skip CPT. Do NOT proceed to persona LoRA on a CPT base that lost to raw.
  • English collapsed → lr too high or too many epochs; re-run Phase 3 with lower lr.

6. Persona LoRA on the CPT'd base (existing pipeline)

Now the CPT base is just a better base. The rewritten train_rocm.py trains a balanced joint persona/router adapter and uses Qwen3's own chat template. It masks the rendered assistant continuation instead of assuming literal ChatML boundary token IDs.

  • The persona data is the {response,mood} corpus per CLAUDE.md Decision B (separate from the CPT corpus): 45% persona chit-chat, 15% graceful failure (tired/confused), 20% tool calls, 10% real utterances, 10% EN→EN. Mood quotas + validator per CLAUDE.md. This dataset is still an open item — see CLAUDE.md "Target data recipe" and gen_data.py open item.

DONE-CHECK 6: LoRA trains; eval (Phase 4) re-run on merged model shows {response,mood} parses + Cyrillic clean + moods in enum.


7. Merge → GGUF → quantize → deploy

Per CLAUDE.md deploy path, with Q8_0 first (decision this session: start high-precision to remove the Cyrillic-at-Q4 gamble, drop to Q5/Q4 only if homesrv memory chokes):

# 1. merge LoRA → bf16
python -c "from peft import AutoPeftModelForCausalLM; import torch; \
m=AutoPeftModelForCausalLM.from_pretrained('./Qwen3-1.7B-ru-lora',torch_dtype=torch.bfloat16); \
m.merge_and_unload().save_pretrained('./Qwen3-maven-merged'); \
from transformers import AutoTokenizer; AutoTokenizer.from_pretrained('./Qwen3-1.7B-ru-lora').save_pretrained('./Qwen3-maven-merged')"
# 2. HF → GGUF (in llama.cpp)
python convert_hf_to_gguf.py ./Qwen3-maven-merged --outfile qwen3-maven-f16.gguf --outtype f16
# 3. quantize — Q8_0 first (~1.6GB for 1.7B)
./llama-quantize qwen3-maven-f16.gguf Qwen3-Maven-1.7B-Q8_0.gguf Q8_0
# 4. deploy: scp to homesrv; measure tok/s + swap with --mlock.
#    If RAM chokes → re-quant Q5_K_M, then Q4_K_M. Point deploy/mavend.json
#    phraser.model_path at the file; small -ngl on Vulkan for a speedup.

DONE-CHECK 7: model runs on homesrv at acceptable voice tok/s (measure!), Cyrillic clean in live output, {response,mood} parsed by the daemon.


8. File manifest (what this plan creates)

All in /home/kami/Programs/esp32-whisper-fine-tune/llm/:

File Phase Purpose
corpus_fetch.py 2.2 pull open datasets → data/cpt_raw/
corpus_synth.py 2.4 router-generate synthetic RU → data/cpt_raw/synth/
corpus_clean.py 2.3 ftfy + langID + mixed-script + dedup → data/cpt_clean/
corpus_stats.py 2.3 per-bucket counts + sample dump (DONE-CHECKs)
corpus_pack.py 2.6 tokenize + pack 2048 blocks → data/cpt_packed/
train_cpt.py 3 continued pretraining, adafactor/DoRA
eval_cyrillic.py 4 Cyrillic % + PPL + EN-retention (extends llama-eval-test.py)
data/cpt_packed/ 2.6 final corpus artifact (~300M tok)
./Qwen3-1.7B-ru-cpt/ 3 CPT'd base checkpoint

8.1 One-time setup

cd /home/kami/Programs/esp32-whisper-fine-tune/llm

# python deps (torch/transformers already present from train_rocm.py)
pip install datasets ftfy datasketch fasttext openai peft accelerate

# fastText language-id model → data/lid.176.bin (corpus_clean.py requires it)
wget -O data/lid.176.bin \
  https://dl.fbaipublicfiles.com/fasttext/supervised-models/lid.176.bin

# HF login — needed to accept the CulturaX (gated) license
huggingface-cli login

If fasttext wheels fail to build, pip install fasttext-wheel is the prebuilt fallback. lid.176.bin can also live elsewhere via export FASTTEXT_LID=/path.

8.2 Runbook (full run order)

cd /home/kami/Programs/esp32-whisper-fine-tune/llm

# ── Phase 2: corpus ──────────────────────────────────────────────
python corpus_fetch.py all                              # 2.2 open datasets
ROUTER_BASE_URL="https://your-router/v1" \
  ROUTER_API_KEY="..." ROUTER_MODELS="m-a,m-b,m-c" \
  python corpus_synth.py                                # 2.4 synthetic (≤10%)
MAVEN_LOGS=/path/to/dialogue.jsonl python corpus_logs.py  # 2.5 optional
python corpus_clean.py                                  # 2.3 clean+dedup all
python corpus_stats.py                                  # DONE-CHECK 2.3/2.4
python corpus_pack.py                                   # 2.6 → data/cpt_packed/

# ── Phase 3: continued pretraining (days; --resume to continue) ──
HSA_OVERRIDE_GFX_VERSION=11.0.0 python train_cpt.py 2>&1 | tee cpt_run.log
#   OOM? → CPT_PATH=dora HSA_OVERRIDE_GFX_VERSION=11.0.0 python train_cpt.py

# ── Phase 4: eval + gate (Phase 5) ──────────────────────────────
python eval_cyrillic.py --model Qwen/Qwen3-1.7B-Base    # baseline
python eval_cyrillic.py --model ./Qwen3-1.7B-ru-cpt     # must beat baseline

# ── Phase 6-7: persona LoRA → merge → GGUF → Q8_0 → homesrv ──────
#   train_rocm.py pointed at ./Qwen3-1.7B-ru-cpt ; then §7 deploy block.

Third-party deps: datasets, ftfy, datasketch, fasttext (+ lid.176.bin), openai, peft, transformers, accelerate.


9. Hard rules (do not violate)

  1. CPT corpus is raw text, no JSON, no ChatML. Persona {response,mood} data is a separate dataset (Phase 6).
  2. Synthetic ≤ 10% of corpus. Bulk must be real human RU.
  3. CPT lr = 1e-5, 1 epoch. Higher = catastrophic forgetting.
  4. Base = Qwen3-1.7B-Base, not Instruct.
  5. Never quantize before merge. bf16 all the way through, quantize last.
  6. Every corpus source passes the same cleaning pipeline (§2.3), synthetic included.
  7. Don't advance a phase until its DONE-CHECK passes. Phase 5 gate is real — if CPT loses to raw, stop.
  8. Respect the two-machine split: train on the 7900 GRE, deploy CPU-only on homesrv ≤1.7B.