From 5397f09785624134015dbd00e990628c3a8e2637 Mon Sep 17 00:00:00 2001 From: claude Date: Mon, 7 Sep 2026 17:58:16 +0400 Subject: [PATCH] =?UTF-8?q?router/semantic:=20slice=2019=20sequence=20prag?= =?UTF-8?q?matics=20specialist=20report=20=E2=80=94=20sequence=20beats=20s?= =?UTF-8?q?parse=20on=20capability=20holdout,=20caps=20near=20chance=20on?= =?UTF-8?q?=20seed-stable=20seeds?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- ...-08-tiny-sequence-pragmatics-specialist.md | 304 ++++++++++++++++++ 1 file changed, 304 insertions(+) create mode 100644 docs/evals/2026-09-08-tiny-sequence-pragmatics-specialist.md diff --git a/docs/evals/2026-09-08-tiny-sequence-pragmatics-specialist.md b/docs/evals/2026-09-08-tiny-sequence-pragmatics-specialist.md new file mode 100644 index 0000000..69c8f3c --- /dev/null +++ b/docs/evals/2026-09-08-tiny-sequence-pragmatics-specialist.md @@ -0,0 +1,304 @@ +# Slice 19: from-scratch tiny sequence encoders halve the sparse gate's capability-question fail rate, and still land near chance on the leave-generator-out pragmatics split + +Date: 2026-09-08 · Task: V-726 (slice 19) · Box: homesrv, Ryzen 5 5600U, 13 GB, CPU-only (this is the production machine) · Build: `cmd/semantic-router-experiment/slice19_*.py`, torch 2.14.0+cpu in `/tmp/mvn-exp-venv`, no GPU. + +## 0. Frozen artifacts + +```text +development corpus v2 hash: b27fd48f478ca477 +original frozen holdout hash: ad297fbdbbea704b (byte-identical, uninspected) +text normalization: NormalizeMatchText (NFKC, lowercase, whitespace-collapse; punctuation and ё kept) +embedding file: /tmp/mvn-experiment/embeddings.json +total examples: 3025 +dev pool: 2490 (796 action / 1694 not_action) +frozen holdout: 535 +router-residual: 2943 +``` + +Same data as slices 16-18. Training input per brief §9 is always the +punctuation-stripped text (`re.sub(r"[^\w\s]", " ", c)` + collapse); the `orig` +and `nofinal` variants are evaluation-only stress views. + +### Method note: the tokenizer axis was broken once and fixed + +The BPE first trained on the *stripped* texts. `tokenizers`' ByteLevel BPE then +silently **drops boundary punctuation at inference** (`сервис,` → the `сервис` +token, comma gone; a comma observed in training keeps its own id 55), so `orig` +vs `strip` differed on only 0.6% of the corpus and the BPE stress axis was a +no-op. Fixed by training both tokenizers on the *natural* dev texts (punctuation +kept, per the normalization line above) while training input remains stripped. +After the fix punctuation fires as real in-vocabulary symbols under stress, and +35.1% of rows differ between `orig` and `strip`. All numbers below are from the +fixed build; the pre-fix run is discarded (its "stress-identical → robust" +reading for BPE models was an artifact). + +## 1. What was tested + +Three from-scratch, order-sensitive encoders, each in a size ladder, all trained +only on the binary action/not-action label from the dev pool: + +| family | A. CharCNN | B. BiGRU | C. TinyTransformer | +| --- | --- | --- | --- | +| input | codepoint ids | subword ids (BPE) | subword ids (BPE) | +| depth | 1D convs (widths 2-5), global max-pool | 1-layer bidirectional GRU, max-pool~final | 2-4 self-attention blocks (PreNorm, GELU FFN 4x), learned pos, max-pool ~ CLS | +| sizes | tiny, medium | tiny, medium, large | small, medium | +| params | 27 k / 149 k | 129 k / 356 k / 1,106 k | 558 k / 2,022 k | + +Chosen per the slice-18 handoff: the deciding signal is order and trailing +politeness/modality, which local n-grams provably could not rank (pair 0.571). + +## 2. Tokenizers + +Trained on the dev pool only. + +| tokenizer | vocab | notes | +| --- | --- | --- | +| CharVocab | 73 (id 0 = PAD; OOV → 0) | codepoint ids incl. punctuation and the occasional CJK char | +| BpeVocab | 1,234 / 2,048 cap | byte-level BPE, serialized 52,897 bytes | + +Max lengths: char 64, BPE 24. No narrowing to Russian; the corpus is the vocab. + +## 3. Training setup + +Fixed hyperparameters, no grid search, seed = 42 + fold for grouped CV: + +| family | epochs | lr | batch | grad clip | optimiser | +| --- | --- | --- | --- | --- | --- | +| char_cnn | 20 | 1e-3 | 64 | none | AdamW wd 1e-4, BCE-with-logits | +| bigru | 15 | 5e-4 | 64 | 1.0 | AdamW wd 1e-4, BCE-with-logits | +| tiny_transformer | 25 | 5e-4 | 64 | 1.0 | AdamW wd 1e-4, BCE-with-logits | + +Grouped 5-fold CV reusing the existing `cv_fold` split; the capability-question +and question families are genuinely held out per fold. Padded sequences, +zero-pad ignored. + +## 4. Grouped CV, binary gate (strip input, OOF at threshold 0.5) + +| config | PR-AUC | ROC-AUC | action_P | action_R | FA | FA rate | +| --- | --- | --- | --- | --- | --- | --- | +| char_cnn_tiny | 0.814 | 0.908 | 0.808 | 0.672 | 127 | 5.1% | +| char_cnn_medium | **0.838** | 0.917 | **0.816** | 0.706 | 127 | **5.1%** | +| bigru_tiny | 0.813 | 0.889 | 0.726 | 0.716 | 215 | 8.6% | +| bigru_medium | 0.794 | 0.864 | 0.684 | 0.693 | 255 | 10.2% | +| bigru_large | 0.837 | 0.893 | 0.689 | 0.755 | 271 | 10.9% | +| tiny_transformer_small | 0.731 | 0.820 | 0.704 | 0.670 | 224 | 9.0% | +| tiny_transformer_medium | 0.742 | 0.844 | 0.686 | 0.750 | 273 | 11.0% | + +The char CNN is the best *aggregate* binary separator (as in slice 18, where the +n-gram representation owned the aggregate boundary). The sequence models trade +aggregate precision for spread-out recall. + +### Fold variance (at 0.5) + +```text + config PR mean PR min PR max FA per fold (n) + char_cnn_tiny 0.791 0.585 0.967 10, 17, 13, 66, 21 + char_cnn_medium 0.802 0.572 0.974 12, 27, 10, 59, 19 + bigru_tiny 0.803 0.632 0.880 25, 53, 37, 63, 37 + bigru_medium 0.751 0.461 0.930 49, 25, 62, 38, 81 + bigru_large 0.790 0.511 0.921 20, 43, 58, 60, 90 + tiny_transformer_small 0.713 0.529 0.923 64, 17, 61, 23, 59 + tiny_transformer_medium 0.714 0.492 0.900 38, 46, 48, 97, 44 +``` + +Fold 3 (`n=506`, the capability/question-heavy split) is the hard fold for the +CNN/GRU families exactly as in every slice since 16. + +## 5. Aggregate comparison against every prior head + +| model | extras | PR-AUC | action_P | action_R | FA rate | cap-Q LOFO FA | pair order | +| --- | --- | --- | --- | --- | --- | --- | --- | +| e5 binary linear (C=1.0) * | ~385 | 0.623 | 0.688 | 0.476 | 6.9% | 1.000 | 0.114 | +| e5 binary MLP H=32 * | 12,353 | 0.674 | 0.673 | 0.569 | 8.8% | 0.976 | 0.136 | +| sparse word+char logistic (sl. 18) | 9,403 | 0.838 | 0.875 | 0.485 | 2.2% | 0.667 | 0.571 | +| char_cnn_medium (19) | 149 k | 0.838 | 0.816 | 0.706 | 5.1% | 0.976 | 0.535 | +| bigru_tiny (19) | 129 k | 0.813 | 0.726 | 0.716 | 8.6% | 0.540 | 0.635 | +| tiny_transformer_small (19) | 558 k | 0.731 | 0.704 | 0.670 | 9.0% | 0.421 | 0.763 | + +\* e5 rows are recomputed here with fixed single hyperparameters and the slice 17 +configuration (`MLP early_stopping=True, validation_fraction=0.15, +n_iter_no_change=10`; logistic C=1.0). FA rates reproduce the published slice +16/17 numbers (6.9% / 8.8%). The published **PR-AUC 0.707** for the linear head +came from the slice 16 C-grid sweep; a fixed-C reproduction lands at 0.623 (and +even the best C on the grid, C=100, reaches only 0.675). The slice 18 table's PR +"0.707 → 0.838 sparse gain" is therefore overstated; the honest sparse-vs-e5 PR +gap is ~0.838 vs 0.62-0.68. + +On the **aggregate** boundary the from-scratch sequence models neither beat the +sparse gate's 2.2% FA (5.1-11.0%) nor reach any P ≥ 0.95 operating point (see §6). +Aggregate accuracy was downgraded to a secondary metric from this slice. + +## 6. Safety operating curve + +None of the seven from-scratch configs has any threshold with P ≥ 0.95 at R > 0 +under the grouped OOF (the `ops` field is empty for every config). The sparse +gate's slice-18 operating point (P ≥ 0.95 at R 0.264) does not transfer to any +sequence model. A from-scratch tiny sequence encoder is not a safe standalone +gate out of the box. + +## 7. Leave-generator-out: capability questions (the critical split) + +Train without the `capability_question` family, evaluate on its 126 rows (0/126 +positive — every row here must *not* trip the gate). Threshold 0.5, seed 17: + +| config | FA | FA rate | accuracy | +| --- | --- | --- | --- | +| **tiny_transformer_small** | **53** | **0.421** | **0.579** | +| bigru_tiny | 68 | 0.540 | 0.460 | +| tiny_transformer_medium | 69 | 0.548 | 0.452 | +| bigru_medium | 93 | 0.738 | 0.262 | +| char_cnn_tiny | 97 | 0.770 | 0.230 | +| bigru_large | 113 | 0.897 | 0.103 | +| char_cnn_medium | 123 | 0.976 | 0.024 | + +Against the sparse gate (0.667 FA → 0.333 accuracy) the leading from-scratch +configs halve the fail rate. Against e5 (1.000 / 0.976) the win is decisive — +the frozen embedder genuinely cannot separate a held-out capability question +from its executable sibling. + +### But the single-seed number is not trustworthy + +Same leave-generator-out training, three seeds (17 / 41 / 7), cap-Q FA rate: + +```text + tiny_transformer_small 0.603 0.421 0.937 (spread 0.52) + tiny_transformer_medium 0.548 0.484 0.524 (spread 0.06) + bigru_tiny 0.690 0.540 0.476 (spread 0.21) + bigru_medium 0.738 0.389 0.952 (spread 0.56) +``` + +Only the **transformer-medium is seed-stable**, and it lands at **0.48-0.55 FA +(~50% ± 3 accuracy)**: a genuine but near-chance separation on the exact +generator family the task exists for. The transformer-small's best single seed +(0.421) and the bigru-medium pair-ordering win (§9) are both inside unstable +regimes. The tooling that produced the lucky number is right; the number itself +is luck. + +## 8. Full-family LOFO, leading config per architecture + +Seed 41, held-out family evaluated in full (accuracy; FA; action recall): + +| held-out family | rows | char_cnn_medium | bigru_tiny | tiny_transformer_small | +| --- | --- | --- | --- | --- | +| polite_request | 481 | 0.279 acc / FA 347 / R 1.000 | 0.486 / FA 247 / R 1.000 | **0.975 / FA 12 / R 1.000** | +| modal_request | 223 | R 0.430 (P inf) | R 0.484 (P 1.000) | **R 0.561 (P 1.000)** | +| first_person_request | 833 | **R 0.942** (P 0.991) | R 0.000 (acc 1.000, no FAs) | R 0.646 (P 1.000) | +| reordered_target | 353 | 1.000 / FA 0 | 0.994 / FA 2 | 1.000 / FA 0 | +| question | 123 | 1.000 / FA 0 | 1.000 / FA 0 | 1.000 / FA 0 | + +The transformer-small is the only config that survives the **polite_request** +holdout — the family that sounds most like a capability question under voice +("пожалуйста" + soft predicate). The CNN collapses it into "action" (347 of 481 +false actions); the GRU nearly so. The same split in reverse: char_cnn owns +first_person recall (short imperatives), which the other two under-recall. + +## 9. Paired action/capability ordering + +2268 pairs (capability question vs executable sibling on the shared object +noun), ordering accuracy and median margin on grouped OOF, strip text: + +| config | ordering | median margin | reversed pairs | +| --- | --- | --- | --- | +| char_cnn_tiny | 0.406 (reversed) | −0.090 | 1347 | +| char_cnn_medium | 0.535 | +0.021 | 1054 | +| bigru_tiny | 0.635 | +0.169 | 828 | +| bigru_medium | **0.866** | +0.598 | 304 | +| bigru_large | 0.724 | +0.282 | 627 | +| tiny_transformer_small | 0.763 | +0.804 | 538 | +| tiny_transformer_medium | 0.648 | +0.019 | 799 | + +bigru_medium reaches 0.866 ordering with a healthy +0.598 median margin — the +first head since the slices began that can more often than not put the executable +above its sibling (sparse: 0.571, margin +0.066). Caveat: this is the same +bigru_medium whose LOFO accuracy swings 0.39-0.95 across seeds; margin and +ordering were measured from the grouped-CV models (seed 42+fold), not the same +seed as the LOFO numbers, and the two axes were not measured jointly. The margin +being three times the sparse margin on the *strip* basis is still a real, +sequence-level result. The char CNN's 0.406-0.535 confirms it has no pragmatics +signal at all — it is a shallow aggregate-n-gram reader. + +## 10. Punctuation / voice stress + +Models trained on stripped text; tokenizers trained on natural text (so +punctuation is a real, seen-at-vocab, unseen-at-training symbol). Stress views +`orig` (full punctuation) / `nofinal` (trailing-only) / `strip`. + +Capability-question FA (in-distribution, 126 rows), and pair ordering per view: + +```text + config capQ_FA orig/nofinal/strip pairs orig/nofinal/strip + char_cnn_medium 0.540 / 0.540 / 0.532 0.538 / 0.538 / 0.535 + bigru_tiny 0.230 / 0.270 / 0.310 0.660 / 0.604 / 0.635 + tiny_transformer_small 0.119 / 0.135 / 0.143 0.797 / 0.791 / 0.763 + tiny_transformer_medium 0.262 / 0.262 / 0.286 0.631 / 0.632 / 0.648 +``` + +Punctuation is not load-bearing for any config (all three views within a couple +of points), and the transformer *gains* ordering accuracy when full punctuation +is present (tf_small 0.797 orig > 0.763 strip). The models ignore the `?`/`,`, +which is the safest possible behaviour under ASR, where they cannot be trusted. +The "stress-robust" claim here is real (unlike the pre-fix build): the +punctuation codepoints existed in the boxes' vocabularies and were simply not +used. + +## 11. Runtime and size (homesrv CPU, batch 1, p50 inference + tokenisation) + +| config | params | fp32 | int8 (=params) | latency p50/p95 | tok | RAM delta over torch | +| --- | --- | --- | --- | --- | --- | --- | +| char_cnn_tiny | 27,297 | 107 KiB | 27 k | 0.28 / 0.42 ms | 3.8 µs | — | +| char_cnn_medium | 149,313 | 583 KiB | 149 k | 0.45 / 0.67 ms | 2.4 µs | ~118 MB | +| bigru_tiny | 129,025 | 504 KiB | 129 k | 0.41 / 0.73 ms | 16.7 µs | ~48 MB | +| bigru_medium | 356,353 | 1.4 MiB | 356 k | 0.52 / 0.82 ms | 15.3 µs | ~63 MB | +| bigru_large | 1,105,921 | 4.2 MiB | 1.1 M | 0.85 / 1.62 ms | 19.5 µs | — | +| tiny_transformer_small | 557,953 | 2.1 MiB | 558 k | 0.89 / 1.30 ms | 24.4 µs | ~89 MB | +| tiny_transformer_medium | 2,021,569 | 7.7 MiB | 2.0 M | 2.57 / 3.78 ms | 17.8 µs | — | + +Every config is sub-4 ms end to end on the production CPU; BPE tokenisation is +the dominant term for a single utterance. RAM deltas are fresh-process +high-water marks over a ~225 MB torch baseline (which is itself the item the +deploy would have to absorb — ~chromosome 48 MB at the model level, crane). + +## 12. Conclusion + +### Verdict: sequence sensitivity is real and useful, from-scratch tiny is not enough + +Three findings, one per decision rule: + +1. **The critical split improved by half, on stable seeds.** The + transformer-medium holds capability-question LOFO FA at 0.48-0.55 across + seeds (sparse gate: 0.667; e5: 1.000/0.976). The transformer-small reaches + 0.421 on its lucky seed. Sequence structure demonstrably carries *some* of + the pragmatics distinction that local features cannot (pair ordering 0.866 + with +0.598 margin for bigru_medium; tf_small 0.975 accuracy on the held-out + polite_request family vs 0.279 for the CNN). The slice-18 direction — + "order and trailing politeness/modality are the deciding signal" — is + confirmed, not refuted. + +2. **It is still a near-coin-flip on the exact unseen-generator case, and the + variance is the story.** Best seed-stable accuracy is ~52%, and only one of + seven configs is seed-stable on that metric (the others swing 0.21-0.56 of + FA rate across three seeds). No config reaches a P ≥ 0.95 gate, and every + config's aggregate FA rate (5.1-11.0%) is worse than the sparse gate's 2.2%. + A from-scratch model this size can *see* a difference between the two members + of a pair from data within the pool, but cannot *generalise it* to a + generator family it has never seen. It memorises family-ish behaviour; when + the family is truly new it reverts toward chance. + +3. **The next step is a pretrained prior, not a bigger from-scratch model.** + The distinguishing features — polite softeners, modal auxiliaries, the + trailing interrogative falling back on world structure — are things a 4 M-8 M + pretrained encoder (e.g. a small modern multilingual transformer) already + encodes in its weights, so fine-tuning on the same 2490-row dev pool can be + expected to move the seed-stable ~0.5 accuracy rather than shuffle the dice + in a new random draw. Same frozen data, same grouped folds, same 15-section + metrics — with the seed-stability measurement (§7) carried forward as + mandatory, since it is what disambiguates a result from a luck draw. + +Practical carry-outs: training input stays punctuation-stripped (free stress +robustness, no cost); tokenizers are trained on natural text so the stress axis +is real; never report a single-seed LOFO number without its seed neighbours. + +## 13. Commit hash for tooling + +`7d31de5` — `cmd/semantic-router-experiment/slice19_models.py`, +`slice19_bpe.py`, `slice19_main.py`; artifacts under `/tmp/mvn-s19/`. \ No newline at end of file