# Fine-tuning a pretrained Russian BERT (rubert-tiny, 11.8 M) makes the in-pool pragmatics boundary the strongest seen so far (pair ordering 0.93-0.97, median margin +0.7) but fires **every one** of the 126 held-out capability-question rows as an action (FA_rate 1.000) on all 18 config/seeds — supervised fine-tuning on this split memorises the seen families' templates, does not learn the boundary, and adds no safety signal anywhere Date: 2026-09-07 · Task: slice 20 (brief after the accepted slice 19, V-726) · Box: homesrv, Ryzen 5 5600U, 13 GB, CPU-only (this is the production machine) · Build: `cmd/semantic-router-experiment/slice20_*.py`, torch 2.14.0+cpu, transformers 5.16.1 in `/tmp/mvn-exp-venv`, no GPU. ## 0. Frozen artifacts ```text development corpus v2 hash: b27fd48f478ca477 original frozen holdout hash: ad297fbdbbea704b (byte-identical, uninspected) text normalization: NormalizeMatchText (NFKC, lowercase, whitespace-collapse; punctuation and ё kept) embedding file: /tmp/mvn-experiment/embeddings.json total examples: 3025 dev pool: 2490 (796 action / 1694 not_action) frozen holdout: 535 router-residual: 2943 capability_question family: 126 rows, all not_action (the exact safety case) family sizes: polite_request 481 / modal_request 223 / first_person_request 833 / reordered_target 353 paired action/capability rows: 2268 (shared object nouns) ``` Identical dev pool to slices 16-19. Training input per brief §9 is the punctuation-stripped text (`re.sub(r"[^\w\s]", " ", c)` + collapse); the `orig` and `nofinal` variants are evaluation-only stress views. Nothing in the frozen holdout was inspected or added; holding the slice-16 corpus line is unchanged. ## 1. What was tested `cointegrated/rubert-tiny` — the tiny Russian BERT — fine-tuned end-to-end for the binary action/not_action label. The model is the 3-layer, 312-hidden, 12-head configuration (same depth/size family as the slice-19 transformers, but with a 11.78 M-parameter pretrained prior on 655 M tokens of bilingual text). One sequence-classification head (a single output neuron) was attached: this is the full fine-tuning door the slice-19 conclusion asked for — "a small modern pretrained encoder already encodes softeners, modal auxiliaries and the trailing interrogative in its weights". | property | value | | --- | --- | | model | `cointegrated/rubert-tiny`, revision `5441c5ea8026d4f6d7505ec004845409f1259fb1` | | license | MIT (2024-02-09 snapshot) | | params | 11,784,481 | | config | BERT base-tiny: hidden 312, layers 3, heads 12, intermediate 600, max_pos 512 | | tokenizer | WordPiece, vocab 29,564, unpadded length-capped | The brief's alternatives were to hold either submission. Nothing here tests capacity beyond the slice-19 range (tiny two-config depth, six configs total). ## 2. Tokenizer audit (gate before training) The slice-19 method note: BPE trained on stripped text silently drops boundary punctuation (`сервис,` loses the comma) and the stress axis becomes a no-op. WordPiece ships pretrained, so this slice audits the frozen tokenizer instead: ```text vocab: 29,564 UNK: 73/28,051 tokens → 0.0026 (all 73 from 5 CJK word-tokens: 记住, 备份策略, 那 — zero Cyrillic/Latin loss in the whole dev pool) tokens/utterance: mean 11.4 tokens/char: 0.426 seq len: p50 11, p90 16, p99 19, max 21 → n above 96 = 0, above 128 = 0 Cyrillic/mixed/Latin fragments intact (перезапусти сервис mavend → пер ##еза ##пус ##ти се ##рви ##с ma ##ven ##d); hashlike ids fine (ha_cam_12 → ha _ cam _ 12) ``` Gate **PASSED**. The UNK floor is a constant 73 (these five CJK words are the same five everywhere; they occur on both sides of the label), the length distribution is far below any practical cap, and no useful token is lost. ## 3. Sequence length `MAX_LEN = 25` (corpus p99 19 + 6). Truncation measured on the dev pool: 0 of 2490 rows. The 128 cap the brief set is respected with a 5× margin. ## 4. Fine-tune setup Fixed hyperparameters, no grid search on them; the three search axes are the learning rate and (implicitly, by regime) the input view. | regime | training input | purpose | | --- | --- | --- | | A | natural text (orig) | brief §9 primary; punctuation seen at train | | B | punctuation-stripped | the slice-19 carry-out; stress-free train | | hyperparameter | value | | --- | --- | | lr grid | 1e-5 · 2e-5 · 5e-5 | | seeds | 42 · 17 · 7 | | epochs | max 4, early stop on val PR-AUC (patience 1) | | batch | 32 | | weight decay | 0.01 | | val split | 0.12 within-train (per fold, not touching test) | | threads | 4 (fp32) — the fast config; 12 threads and bf16 autocast are several × slower | Grouped 5-fold CV reusing the existing `cv_fold` split; the capability-question/question families are genuinely held out per fold. Each train call resamples fresh fold seeds, giving the 3 seed groupings independence (both for grouped CV and for the leave-generator-out runs). 18 fine-tunes × 5 folds grouped + 18 leave-generator-out runs, ~87 min total on the box. ## 5. Grouped CV, binary gate (strip input, OOF at threshold 0.5) The in-pool aggregate boundary. Rows are OOF (each row's label folded out of its train), three seeds (42/17/7), PR-AUC interval is the 3-seed min..max. | config | PR-AUC | FA rate | P@0.5 | R@0.5 | FA | cap-Q in-pool | | --- | --- | --- | --- | --- | --- | --- | | A @ lr1e-5 | 0.750 (0.732..0.768) | 8.2-8.9% | 0.75 | 0.68 | 205-221 | 0.49-0.60 | | **A @ lr2e-5** | **0.831 (0.800..0.856)** | **6.5-7.5%** | 0.81 | 0.80 | 161-186 | **0.17-0.37** | | A @ lr5e-5 | 0.814 (0.795..0.828) | 9.6-11.1% | 0.78 | 0.76 | 238-277 | 0.19-0.41 | | B @ lr1e-5 | 0.730 (0.678..0.786) | 6.9-11.9% | 0.75 | 0.68 | 171-296 | 0.41-0.87 | | B @ lr2e-5 | 0.810 (0.749..0.845) | 7.4-9.0% | 0.77 | 0.77 | 183-223 | 0.18-0.52 | | B @ lr5e-5 | 0.833 (0.808..0.865) | 6.9-8.3% | 0.80 | 0.75 | 172-207 | 0.10-0.38 | The aggregate boundary ~0.83 PR-AUC enters the slice-18 sparse's territory (0.838) but the FA rate (6.5-11.9%) stays well above the sparse gate's 2.2%. The secondary window cap-Q in-pool (0.10-0.60) shows the family is *learnable when present in train*: the boundary exists inside the model's seen distribution. ### Fold variance (A @ lr2e-5) ```text seed val PR-AUC per fold (a 12.4%-frac of each fold) epochs per fold 42 0.995 0.997 0.972 0.999 0.994 4 4 4 4 4 17 0.999 0.994 0.980 0.992 0.984 4 4 4 4 4 ``` Early stopping never fires — every fold hits the epoch ceiling with val PR-AUC already ≥ 0.97, i.e. the capacity is nowhere near exhausted on the train side. The fold-3 (capability/question-heavy) weakness that from-scratch models showed is invisible here on the val split; the OOF @0.5 FA still jitters 161-186 on the worst seeds. This is the slope the LOFO result (§7), not the fold table, reveals. ## 6. Safety operating curve None of the six configs has any threshold with P ≥ 0.95 at R > 0 under the grouped OOF (`ops` empty in the sync of every seed). The best precision on the entire curve of the best seed is 0.934 at R 0.230. The sparse gate's slice-18 operating point (P ≥ 0.95 at R 0.264) does not transfer to any fine-tuned pretrained model either. **A fine-tuned rubert-tiny is not a safe standalone gate out of the box, even in-pool.** ## 7. Leave-generator-out: capability questions (the critical split, §8/§11) Same leave-one-family-out as slice 19 §7, now 3 seeds × 6 configs. Evaluate on the 126 capability-question rows (0/126 positive — every row here must NOT trip the gate). Threshold 0.5: ```text config FA(tot=126) FA_rate mean action proba A @ lr1e-5 (seeds 7/17/42) 126 / 126 / 126 1.000 × 3 0.86-0.87 A @ lr2e-5 (seeds 7/17/42) 126 / 126 / 126 1.000 × 3 0.965-0.972 A @ lr5e-5 (seeds 7/17/42) 126 / 126 / 126 1.000 × 3 0.988-0.991 B @ lr1e-5 (seeds 7/17/42) 126 / 126 / 126 1.000 × 3 0.876-0.893 B @ lr2e-5 (seeds 7/17/42) 126 / 126 / 126 1.000 × 3 0.965-0.967 B @ lr5e-5 (seeds 7/17/42) 126 / 126 / 126 1.000 × 3 0.990-0.992 ``` **Every held-out capability question reads as an action, with near-deterministic confidence** (mean action proba 0.86-0.99, within-family standard deviation ~0.0045). This is not a tuning artifact: it holds at the weakest LR (1e-5), all seeds, both regimes. It is the sharpest failure in the entire slice family: worse than e5 (0.976-1.000), worse than the from-scratch seed-stable transformer-medium (0.48-0.55), worse than the sparse gate (0.667). The fine-tuned pretrained model — the strongest *in-pool* pragmatics reader these slices have produced (pair 0.97, §9) — collapses to 100% fire on the exact unseen-generator case the task exists for. ## 8. The decisive reading (§17 hypothesis test) The contrast that separates the two hypotheses is now sharp: * **In-pool** (family present in train): cap-Q FA down to 0.17 (§5), pair ordering up to 0.97 (§9). The model learns the boundary whenever the generator family is in the training split. * **LOFO** (family absent): 1.000 across every config and seed. A capacity or signal problem would degrade *both* in-pool and LOFO. Instead the fine-tune is at its *best* in-pool and at its absolute *worst* out-of-family. The increase in pretrained representation power over from-scratch does **not** transfer to the unseen generator; it sharpens the seen-family templates until every capability question is swept into "action". The low LR (1e-5) does not rescue it — it already fires 84-89% of the rows at 1e-5. The phenomenon is template/generator specialization, not lack of prior. ## 9. Paired action/capability ordering 2268 pairs (capability question vs executable sibling on the shared object noun), ordering accuracy and median margin on grouped OOF, strip view: | config | ordering | median margin | | --- | --- | --- | | A @ lr1e-5 | 0.740 (0.731..0.755) | +0.20 | | **A @ lr2e-5** | **0.933 (0.869..0.970)** | **+0.56..+0.78** | | A @ lr5e-5 | 0.822 (0.690..0.952) | +0.06..+0.91 | | B @ lr1e-5 | 0.605 (0.482..0.787) | −0.01..+0.26 | | B @ lr2e-5 | 0.901 (0.874..0.933) | +0.53..+0.62 | | B @ lr5e-5 | 0.808 (0.743..0.918) | +0.10..+0.80 | Best seeds put the executable above its sibling 96-97% of the time with a median margin of +0.74-0.81 — the strongest ordering ever measured on this fixture (from-scratch best: bigru-medium 0.866 / +0.598; sparse 0.571). Pair ordering is the one axis on which the pretrained model is clearly best, and it is the version of the task where the two members share surface structure, so a trained-together contrast is exactly what the model generalizes. The pair task is *in-distribution relative to the model's world prior*; the family boundary is not. ## 10. Punctuation / voice stress Models trained on regime A (natural text) vs B (punct-stripped); punctuation is in the pretrained WordPiece vocabulary. Stress views `orig` / `nofinal` / `strip`, evaluated on the same early-stopped checkpoint: ```text config cap-Q in-pool FA (orig/nofinal/strip) pairs (orig/nofinal/strip) A @ lr2e-5 seed17 0.143 / 0.183 / 0.167 0.968 / 0.959 / 0.961 A @ lr2e-5 seed42 0.294 / 0.365 / 0.365 0.893 / 0.878 / 0.869 B @ lr5e-5 seed7 0.175 / 0.183 / 0.183 0.904 / 0.895 / 0.896 ``` Punctuation is not load-bearing: all three views agree within ~2 points on every metric, for both regimes. The model ignores the `?`/`,` under stress, which is the safest possible behaviour under ASR. The earlier slice-19 "stress-robust" was a tokenizer artifact; here it is checked with punctuation genuinely in the vocabulary and holds. ## 11. Runtime, size, ONNX (homesrv CPU, batch 1) | property | value | | --- | --- | | params | 11,784,481 | | fp32 weights | 47.1 MB (onnx initializer raw) | | fp16 weights | 23.6 MB | | int8 weights | 11.8 MB | | tokenizer files | 709,227 B | | latency p50 / p95 (CPU batch 1, incl. tokenize) | 3.66 / 5.25 ms | | tokenizer | 64.7 µs/utterance | | threads | 12, MAX_LEN 25 | **RAM** (fresh process, torch+transformers baseline then model + inference): baseline 408 MB RSS, with rubert-tiny resident 570 MB → incremental ~161 MB. That delta is what a python-side subprocess would add on top of an already imported transformers; a Go daemon embedding the ONNX directly would pay only the 47 MB fp32 (24 MB fp16 / 12 MB int8) + 0.7 MB tokenizer. Production CPU latency ~3.7 ms p50 satisfies the speech budget (slice-19 BERT-based rows were 0.9-2.6 ms batch-1; this is the same order). **ONNX**: export succeeds (`rubert-tiny-gate.onnx`, 411,087 B zip, fp32 raw). **Parity FAILED**: onnnxruntime CPU probabilities differ from torch by up to 0.56 (mean 0.37) — the exported graph is not faithful to the torch forward on the same inputs. The fault is in the export path, not the fine-tune; but since the model did not survive the §8 selection, §15 does not gate the conclusion. Reported so a future candidate is not burned on a broken export again. ## 12. tiny2 ceiling (§13 of the brief — optional, run) The brief allowed `cointegrated/rubert-tiny2` as a ceiling probe only. Run (the same 3-layer/312-hidden family, newer pretraining, vocab 83,828, max_pos 2048): ```text audit: tokens/utterance mean 11.4, p99 13 (newer, richer WordPiece), unk 0.37% cap-Q LOFO (A @ lr2e-5, seeds 7/17/42): FA = 126/126/126, FA_rate 1.000 × 3 in-pool (grouped, single seed): PR-AUC 0.782, P 0.735, R 0.747, FA 215 (8.6%) pairs 0.699, cap-Q in-pool 0.619 ``` tiny2 is **worse in-pool** (pair 0.699 vs 0.933; cap-Q in-pool 0.62 vs 0.17) and **just as catastrophic on the held-out family**. A newer, larger-vocab pretraining of the same family does not move the LOFO boundary at all. This rules out the "tiny1 is somehow unlucky pretraining" explanation: the family prior itself does not carry the boundary. ## 13. Conclusion ### Verdict: the pretrained fine-tune closes in-pool but not the unseen generator; nothing in the supervised family passes the safety door 1. **In-pool, the pretrained model is the strongest pragmatics reader so far.** Pair ordering 0.933-0.970 with median margin up to +0.81 beats every from-scratch config (best 0.866) and the sparse gate (0.571); cap-Q in-pool FA drops to 0.10-0.37 where from-scratch was 0.23-0.54; aggregate PR-AUC 0.83 ties the sparse row. The pretrained prior demonstrably contains the order/politeness/modality signal slice 18 predicted — when the generator is in-distribution. 2. **The exact safety case fails harder than everything measured before.** Leave-generator-out on capability questions: FA_rate **1.000 on all 18 config/seeds** (mean proba 0.86-0.99, within-family σ ~0.0045), vs sparse 0.667, from-scratch seed-stable 0.48-0.55, e5 0.976-1.000. A model that is the *best* in-pool boundary is simultaneously the *worst* out-of-family. More representation power sharpens the seen templates; it does not open a hole to the unseen generator. every supervised route — frozen, from-scratch, pretrained-fine-tuned — caps or collapses on this split (sparse 0.667 FA, from-scratch ~0.5, pretrained 1.000). 3. **The spike is generator/template memorization, not lack of prior.** tiny2 (newer, larger-vocab pretraining) does not move LOFO (1.000) and is worse in-pool. The decisive §11 test says: the capability-question family is not learnable from the in-pool surface; what the fine-tune learns is "when my friends speak, act". The pragmatics distinction these slices chase lives in the *generator structure*, and in the corpus it is visible only as a template; a supervised encoder sees the template, not the generator. ### What this rules out, and what is left * **No supervised sequence-model scaling is worth another slice of LOFO.** Frozen (e5) and fine-tuned (rubert-tiny) both sit at or above 0.976 FA on the holdout; from-scratch straddles chance. Three different inductive biases have now failed the same split the same categorical way. * **The safety floor remains the sparse gate** (0.667 FA, the single best LOFO number on the fixture) — kept, with its known 2.2% aggregate FA cost. * Remaining doors, in the order the brief's decision rule points at: (a) a **template-structure syntactic gate** (modality-verb pattern on top of sparse) that decomposes the generator rather than learning its shadow; (b) a **frozen-judge LLM** that answers "would this text execute a tool?" zero-shot, decoupled from any fine-tune; (c) more data from the capability-question *generator* (not more utterances) so the family is learned by construction rather than by silhouette. * Carry-outs that survive regardless of route: training input stays stripped (free stress robustness); tokenizers/embeddings audited on natural text; never report one-seed LOFO without its neighbours; a broken ONNX export must not be shipped as "parity OK" by an internal self-check. ## 14. Commit hash for tooling `slice20_audit.py`, `slice20_pretrained.py` (tokenize/grouped/lfo/metrics/ runtime/onnx/ceiling subcommands, `MAX_LEN=25`); artifacts under `/tmp/mvn-s20/`.