router/semantic: slice 19 sequence pragmatics specialist report — sequence beats sparse on capability holdout, caps near chance on seed-stable seeds

This commit is contained in:
2026-09-07 17:58:16 +04:00
parent 7d31de5a6b
commit 5397f09785
@@ -0,0 +1,304 @@
# Slice 19: from-scratch tiny sequence encoders halve the sparse gate's capability-question fail rate, and still land near chance on the leave-generator-out pragmatics split
Date: 2026-09-08 · Task: V-726 (slice 19) · Box: homesrv, Ryzen 5 5600U, 13 GB, CPU-only (this is the production machine) · Build: `cmd/semantic-router-experiment/slice19_*.py`, torch 2.14.0+cpu in `/tmp/mvn-exp-venv`, no GPU.
## 0. Frozen artifacts
```text
development corpus v2 hash: b27fd48f478ca477
original frozen holdout hash: ad297fbdbbea704b (byte-identical, uninspected)
text normalization: NormalizeMatchText (NFKC, lowercase, whitespace-collapse; punctuation and ё kept)
embedding file: /tmp/mvn-experiment/embeddings.json
total examples: 3025
dev pool: 2490 (796 action / 1694 not_action)
frozen holdout: 535
router-residual: 2943
```
Same data as slices 16-18. Training input per brief §9 is always the
punctuation-stripped text (`re.sub(r"[^\w\s]", " ", c)` + collapse); the `orig`
and `nofinal` variants are evaluation-only stress views.
### Method note: the tokenizer axis was broken once and fixed
The BPE first trained on the *stripped* texts. `tokenizers`' ByteLevel BPE then
silently **drops boundary punctuation at inference** (`сервис,` → the `сервис`
token, comma gone; a comma observed in training keeps its own id 55), so `orig`
vs `strip` differed on only 0.6% of the corpus and the BPE stress axis was a
no-op. Fixed by training both tokenizers on the *natural* dev texts (punctuation
kept, per the normalization line above) while training input remains stripped.
After the fix punctuation fires as real in-vocabulary symbols under stress, and
35.1% of rows differ between `orig` and `strip`. All numbers below are from the
fixed build; the pre-fix run is discarded (its "stress-identical → robust"
reading for BPE models was an artifact).
## 1. What was tested
Three from-scratch, order-sensitive encoders, each in a size ladder, all trained
only on the binary action/not-action label from the dev pool:
| family | A. CharCNN | B. BiGRU | C. TinyTransformer |
| --- | --- | --- | --- |
| input | codepoint ids | subword ids (BPE) | subword ids (BPE) |
| depth | 1D convs (widths 2-5), global max-pool | 1-layer bidirectional GRU, max-pool~final | 2-4 self-attention blocks (PreNorm, GELU FFN 4x), learned pos, max-pool ~ CLS |
| sizes | tiny, medium | tiny, medium, large | small, medium |
| params | 27 k / 149 k | 129 k / 356 k / 1,106 k | 558 k / 2,022 k |
Chosen per the slice-18 handoff: the deciding signal is order and trailing
politeness/modality, which local n-grams provably could not rank (pair 0.571).
## 2. Tokenizers
Trained on the dev pool only.
| tokenizer | vocab | notes |
| --- | --- | --- |
| CharVocab | 73 (id 0 = PAD; OOV → 0) | codepoint ids incl. punctuation and the occasional CJK char |
| BpeVocab | 1,234 / 2,048 cap | byte-level BPE, serialized 52,897 bytes |
Max lengths: char 64, BPE 24. No narrowing to Russian; the corpus is the vocab.
## 3. Training setup
Fixed hyperparameters, no grid search, seed = 42 + fold for grouped CV:
| family | epochs | lr | batch | grad clip | optimiser |
| --- | --- | --- | --- | --- | --- |
| char_cnn | 20 | 1e-3 | 64 | none | AdamW wd 1e-4, BCE-with-logits |
| bigru | 15 | 5e-4 | 64 | 1.0 | AdamW wd 1e-4, BCE-with-logits |
| tiny_transformer | 25 | 5e-4 | 64 | 1.0 | AdamW wd 1e-4, BCE-with-logits |
Grouped 5-fold CV reusing the existing `cv_fold` split; the capability-question
and question families are genuinely held out per fold. Padded sequences,
zero-pad ignored.
## 4. Grouped CV, binary gate (strip input, OOF at threshold 0.5)
| config | PR-AUC | ROC-AUC | action_P | action_R | FA | FA rate |
| --- | --- | --- | --- | --- | --- | --- |
| char_cnn_tiny | 0.814 | 0.908 | 0.808 | 0.672 | 127 | 5.1% |
| char_cnn_medium | **0.838** | 0.917 | **0.816** | 0.706 | 127 | **5.1%** |
| bigru_tiny | 0.813 | 0.889 | 0.726 | 0.716 | 215 | 8.6% |
| bigru_medium | 0.794 | 0.864 | 0.684 | 0.693 | 255 | 10.2% |
| bigru_large | 0.837 | 0.893 | 0.689 | 0.755 | 271 | 10.9% |
| tiny_transformer_small | 0.731 | 0.820 | 0.704 | 0.670 | 224 | 9.0% |
| tiny_transformer_medium | 0.742 | 0.844 | 0.686 | 0.750 | 273 | 11.0% |
The char CNN is the best *aggregate* binary separator (as in slice 18, where the
n-gram representation owned the aggregate boundary). The sequence models trade
aggregate precision for spread-out recall.
### Fold variance (at 0.5)
```text
config PR mean PR min PR max FA per fold (n)
char_cnn_tiny 0.791 0.585 0.967 10, 17, 13, 66, 21
char_cnn_medium 0.802 0.572 0.974 12, 27, 10, 59, 19
bigru_tiny 0.803 0.632 0.880 25, 53, 37, 63, 37
bigru_medium 0.751 0.461 0.930 49, 25, 62, 38, 81
bigru_large 0.790 0.511 0.921 20, 43, 58, 60, 90
tiny_transformer_small 0.713 0.529 0.923 64, 17, 61, 23, 59
tiny_transformer_medium 0.714 0.492 0.900 38, 46, 48, 97, 44
```
Fold 3 (`n=506`, the capability/question-heavy split) is the hard fold for the
CNN/GRU families exactly as in every slice since 16.
## 5. Aggregate comparison against every prior head
| model | extras | PR-AUC | action_P | action_R | FA rate | cap-Q LOFO FA | pair order |
| --- | --- | --- | --- | --- | --- | --- | --- |
| e5 binary linear (C=1.0) * | ~385 | 0.623 | 0.688 | 0.476 | 6.9% | 1.000 | 0.114 |
| e5 binary MLP H=32 * | 12,353 | 0.674 | 0.673 | 0.569 | 8.8% | 0.976 | 0.136 |
| sparse word+char logistic (sl. 18) | 9,403 | 0.838 | 0.875 | 0.485 | 2.2% | 0.667 | 0.571 |
| char_cnn_medium (19) | 149 k | 0.838 | 0.816 | 0.706 | 5.1% | 0.976 | 0.535 |
| bigru_tiny (19) | 129 k | 0.813 | 0.726 | 0.716 | 8.6% | 0.540 | 0.635 |
| tiny_transformer_small (19) | 558 k | 0.731 | 0.704 | 0.670 | 9.0% | 0.421 | 0.763 |
\* e5 rows are recomputed here with fixed single hyperparameters and the slice 17
configuration (`MLP early_stopping=True, validation_fraction=0.15,
n_iter_no_change=10`; logistic C=1.0). FA rates reproduce the published slice
16/17 numbers (6.9% / 8.8%). The published **PR-AUC 0.707** for the linear head
came from the slice 16 C-grid sweep; a fixed-C reproduction lands at 0.623 (and
even the best C on the grid, C=100, reaches only 0.675). The slice 18 table's PR
"0.707 → 0.838 sparse gain" is therefore overstated; the honest sparse-vs-e5 PR
gap is ~0.838 vs 0.62-0.68.
On the **aggregate** boundary the from-scratch sequence models neither beat the
sparse gate's 2.2% FA (5.1-11.0%) nor reach any P ≥ 0.95 operating point (see §6).
Aggregate accuracy was downgraded to a secondary metric from this slice.
## 6. Safety operating curve
None of the seven from-scratch configs has any threshold with P ≥ 0.95 at R > 0
under the grouped OOF (the `ops` field is empty for every config). The sparse
gate's slice-18 operating point (P ≥ 0.95 at R 0.264) does not transfer to any
sequence model. A from-scratch tiny sequence encoder is not a safe standalone
gate out of the box.
## 7. Leave-generator-out: capability questions (the critical split)
Train without the `capability_question` family, evaluate on its 126 rows (0/126
positive — every row here must *not* trip the gate). Threshold 0.5, seed 17:
| config | FA | FA rate | accuracy |
| --- | --- | --- | --- |
| **tiny_transformer_small** | **53** | **0.421** | **0.579** |
| bigru_tiny | 68 | 0.540 | 0.460 |
| tiny_transformer_medium | 69 | 0.548 | 0.452 |
| bigru_medium | 93 | 0.738 | 0.262 |
| char_cnn_tiny | 97 | 0.770 | 0.230 |
| bigru_large | 113 | 0.897 | 0.103 |
| char_cnn_medium | 123 | 0.976 | 0.024 |
Against the sparse gate (0.667 FA → 0.333 accuracy) the leading from-scratch
configs halve the fail rate. Against e5 (1.000 / 0.976) the win is decisive —
the frozen embedder genuinely cannot separate a held-out capability question
from its executable sibling.
### But the single-seed number is not trustworthy
Same leave-generator-out training, three seeds (17 / 41 / 7), cap-Q FA rate:
```text
tiny_transformer_small 0.603 0.421 0.937 (spread 0.52)
tiny_transformer_medium 0.548 0.484 0.524 (spread 0.06)
bigru_tiny 0.690 0.540 0.476 (spread 0.21)
bigru_medium 0.738 0.389 0.952 (spread 0.56)
```
Only the **transformer-medium is seed-stable**, and it lands at **0.48-0.55 FA
(~50% ± 3 accuracy)**: a genuine but near-chance separation on the exact
generator family the task exists for. The transformer-small's best single seed
(0.421) and the bigru-medium pair-ordering win (§9) are both inside unstable
regimes. The tooling that produced the lucky number is right; the number itself
is luck.
## 8. Full-family LOFO, leading config per architecture
Seed 41, held-out family evaluated in full (accuracy; FA; action recall):
| held-out family | rows | char_cnn_medium | bigru_tiny | tiny_transformer_small |
| --- | --- | --- | --- | --- |
| polite_request | 481 | 0.279 acc / FA 347 / R 1.000 | 0.486 / FA 247 / R 1.000 | **0.975 / FA 12 / R 1.000** |
| modal_request | 223 | R 0.430 (P inf) | R 0.484 (P 1.000) | **R 0.561 (P 1.000)** |
| first_person_request | 833 | **R 0.942** (P 0.991) | R 0.000 (acc 1.000, no FAs) | R 0.646 (P 1.000) |
| reordered_target | 353 | 1.000 / FA 0 | 0.994 / FA 2 | 1.000 / FA 0 |
| question | 123 | 1.000 / FA 0 | 1.000 / FA 0 | 1.000 / FA 0 |
The transformer-small is the only config that survives the **polite_request**
holdout — the family that sounds most like a capability question under voice
("пожалуйста" + soft predicate). The CNN collapses it into "action" (347 of 481
false actions); the GRU nearly so. The same split in reverse: char_cnn owns
first_person recall (short imperatives), which the other two under-recall.
## 9. Paired action/capability ordering
2268 pairs (capability question vs executable sibling on the shared object
noun), ordering accuracy and median margin on grouped OOF, strip text:
| config | ordering | median margin | reversed pairs |
| --- | --- | --- | --- |
| char_cnn_tiny | 0.406 (reversed) | 0.090 | 1347 |
| char_cnn_medium | 0.535 | +0.021 | 1054 |
| bigru_tiny | 0.635 | +0.169 | 828 |
| bigru_medium | **0.866** | +0.598 | 304 |
| bigru_large | 0.724 | +0.282 | 627 |
| tiny_transformer_small | 0.763 | +0.804 | 538 |
| tiny_transformer_medium | 0.648 | +0.019 | 799 |
bigru_medium reaches 0.866 ordering with a healthy +0.598 median margin — the
first head since the slices began that can more often than not put the executable
above its sibling (sparse: 0.571, margin +0.066). Caveat: this is the same
bigru_medium whose LOFO accuracy swings 0.39-0.95 across seeds; margin and
ordering were measured from the grouped-CV models (seed 42+fold), not the same
seed as the LOFO numbers, and the two axes were not measured jointly. The margin
being three times the sparse margin on the *strip* basis is still a real,
sequence-level result. The char CNN's 0.406-0.535 confirms it has no pragmatics
signal at all — it is a shallow aggregate-n-gram reader.
## 10. Punctuation / voice stress
Models trained on stripped text; tokenizers trained on natural text (so
punctuation is a real, seen-at-vocab, unseen-at-training symbol). Stress views
`orig` (full punctuation) / `nofinal` (trailing-only) / `strip`.
Capability-question FA (in-distribution, 126 rows), and pair ordering per view:
```text
config capQ_FA orig/nofinal/strip pairs orig/nofinal/strip
char_cnn_medium 0.540 / 0.540 / 0.532 0.538 / 0.538 / 0.535
bigru_tiny 0.230 / 0.270 / 0.310 0.660 / 0.604 / 0.635
tiny_transformer_small 0.119 / 0.135 / 0.143 0.797 / 0.791 / 0.763
tiny_transformer_medium 0.262 / 0.262 / 0.286 0.631 / 0.632 / 0.648
```
Punctuation is not load-bearing for any config (all three views within a couple
of points), and the transformer *gains* ordering accuracy when full punctuation
is present (tf_small 0.797 orig > 0.763 strip). The models ignore the `?`/`,`,
which is the safest possible behaviour under ASR, where they cannot be trusted.
The "stress-robust" claim here is real (unlike the pre-fix build): the
punctuation codepoints existed in the boxes' vocabularies and were simply not
used.
## 11. Runtime and size (homesrv CPU, batch 1, p50 inference + tokenisation)
| config | params | fp32 | int8 (=params) | latency p50/p95 | tok | RAM delta over torch |
| --- | --- | --- | --- | --- | --- | --- |
| char_cnn_tiny | 27,297 | 107 KiB | 27 k | 0.28 / 0.42 ms | 3.8 µs | — |
| char_cnn_medium | 149,313 | 583 KiB | 149 k | 0.45 / 0.67 ms | 2.4 µs | ~118 MB |
| bigru_tiny | 129,025 | 504 KiB | 129 k | 0.41 / 0.73 ms | 16.7 µs | ~48 MB |
| bigru_medium | 356,353 | 1.4 MiB | 356 k | 0.52 / 0.82 ms | 15.3 µs | ~63 MB |
| bigru_large | 1,105,921 | 4.2 MiB | 1.1 M | 0.85 / 1.62 ms | 19.5 µs | — |
| tiny_transformer_small | 557,953 | 2.1 MiB | 558 k | 0.89 / 1.30 ms | 24.4 µs | ~89 MB |
| tiny_transformer_medium | 2,021,569 | 7.7 MiB | 2.0 M | 2.57 / 3.78 ms | 17.8 µs | — |
Every config is sub-4 ms end to end on the production CPU; BPE tokenisation is
the dominant term for a single utterance. RAM deltas are fresh-process
high-water marks over a ~225 MB torch baseline (which is itself the item the
deploy would have to absorb — ~chromosome 48 MB at the model level, crane).
## 12. Conclusion
### Verdict: sequence sensitivity is real and useful, from-scratch tiny is not enough
Three findings, one per decision rule:
1. **The critical split improved by half, on stable seeds.** The
transformer-medium holds capability-question LOFO FA at 0.48-0.55 across
seeds (sparse gate: 0.667; e5: 1.000/0.976). The transformer-small reaches
0.421 on its lucky seed. Sequence structure demonstrably carries *some* of
the pragmatics distinction that local features cannot (pair ordering 0.866
with +0.598 margin for bigru_medium; tf_small 0.975 accuracy on the held-out
polite_request family vs 0.279 for the CNN). The slice-18 direction —
"order and trailing politeness/modality are the deciding signal" — is
confirmed, not refuted.
2. **It is still a near-coin-flip on the exact unseen-generator case, and the
variance is the story.** Best seed-stable accuracy is ~52%, and only one of
seven configs is seed-stable on that metric (the others swing 0.21-0.56 of
FA rate across three seeds). No config reaches a P ≥ 0.95 gate, and every
config's aggregate FA rate (5.1-11.0%) is worse than the sparse gate's 2.2%.
A from-scratch model this size can *see* a difference between the two members
of a pair from data within the pool, but cannot *generalise it* to a
generator family it has never seen. It memorises family-ish behaviour; when
the family is truly new it reverts toward chance.
3. **The next step is a pretrained prior, not a bigger from-scratch model.**
The distinguishing features — polite softeners, modal auxiliaries, the
trailing interrogative falling back on world structure — are things a 4 M-8 M
pretrained encoder (e.g. a small modern multilingual transformer) already
encodes in its weights, so fine-tuning on the same 2490-row dev pool can be
expected to move the seed-stable ~0.5 accuracy rather than shuffle the dice
in a new random draw. Same frozen data, same grouped folds, same 15-section
metrics — with the seed-stability measurement (§7) carried forward as
mandatory, since it is what disambiguates a result from a luck draw.
Practical carry-outs: training input stays punctuation-stripped (free stress
robustness, no cost); tokenizers are trained on natural text so the stress axis
is real; never report a single-seed LOFO number without its seed neighbours.
## 13. Commit hash for tooling
`7d31de5``cmd/semantic-router-experiment/slice19_models.py`,
`slice19_bpe.py`, `slice19_main.py`; artifacts under `/tmp/mvn-s19/`.