Files
Maven/docs/evals/2026-09-07-slice16-diagnostic.md
T

8.5 KiB

Slice 16 Diagnostic: Action/Non-Action Boundary Analysis

0. Slice 15 Commit Hashes

semantic seed type:       07bfcea
merge-corpus tool:        6397ea1
e5 embedding cache:       f8ec77d
sklearn experiment:       9df2239
corpus-factory:           4bb7555
expanded corpus:          87411e5
contract tests:           ad3f2d6
experiment reports:       cbac8b9

development corpus v2 hash: b27fd48f478ca477
original frozen holdout hash: ad297fbdbbea704b (byte-identical, uninspected)

1. Router-Residual OOF Safety Metrics

action precision:    0.620
action recall:       0.826
false-action count:  388
false-action rate:   16.0%
uncertain F1:        0.587
total examples:      2418
best C:              10.0

2. False-Action Decomposition by Semantic Family

Total false actions: 388

By semantic family

family count rate

memory_write 133 34.3% capability_question 112 28.9% knowledge_general 77 19.8% system 29 7.5% uncertain 26 6.7% question 9 2.3% conversation 2 0.5%

By held-out split group (top 20)

split_group count

knowledge:capability-ha 70 knowledge:world-def 46 free:remember 45 knowledge:capability-tool 42 note:task 28 system:self-version 27 fact:meal 27 knowledge:homelab-status 15 uncertain:ambiguous-noun 15 note:homelab 11 fact:water 10 note:idea 8 knowledge:deadline 7 uncertain:single-word-verb 7 knowledge:world-explain 6 knowledge:homelab-disk 5 knowledge:task-check 5 uncertain:multi-ambiguous 3 fact:pills 2 system:quiet-on 2

By fold

fold count

0     28
1     38
2     92
3    104
4    126

Diagnosis

  • Distinct semantic families contributing false actions: 7
  • Distinct held-out split groups: 24
  • Largest single family: memory_write (133 false actions)
  • Fold false-action count CV (std/mean): 0.49
  • Broad distribution across families suggests systematic action/non-action overlap

3. Paired E5 Geometry

cosine(action seed, positive action): mean=0.907 std=0.039 min=0.765 max=0.997 n=38222 cosine(action seed, capability question): mean=0.897 std=0.035 min=0.782 max=0.979 n=796 cosine(action seed, other negative contrast): mean=0.893 std=0.031 min=0.822 max=0.975 n=796

Separation gap (positive - capability): 0.010 WARNING: e5 maps action seeds and capability questions nearly on top of each other. The representation itself may not preserve useful separation for this boundary.

4. Binary Action Probe (Linear Logistic)

ROC-AUC:           0.863
PR-AUC:            0.707
precision:         0.726
recall:            0.538
false-positive:    172
false-negative:    417
total:             2490
best C:            1.0

Per-fold: Fold 0: ROC=0.967 PR=0.937 P=0.978 R=0.618 FP=2 FN=55 n=476 Fold 1: ROC=0.914 PR=0.940 P=0.919 R=0.682 FP=12 FN=64 n=341 Fold 2: ROC=0.962 PR=0.952 P=1.000 R=0.195 FP=0 FN=239 n=707 Fold 3: ROC=0.832 PR=0.472 P=0.482 R=0.767 FP=71 FN=20 n=506 Fold 4: ROC=0.640 PR=0.235 P=0.250 R=0.426 FP=87 FN=39 n=460

5. Cost-Sensitive Linear Action Classification

cost action_P action_R FA count FA rate

1      0.601      0.804        390      15.7%
2      0.562      0.847        481      19.3%
4      0.531      0.891        589      23.5%
8      0.494      0.919        711      28.6%

16 0.461 0.953 847 34.2%

6. Expanded Action-Threshold Sweep (30 action groups)

threshold action_P action_R FA count coverage

  0.50      0.664      0.736        285      92.9%
  0.60      0.687      0.655        229      88.0%
  0.70      0.698      0.537        178      82.1%
  0.80      0.722      0.407        120      75.6%
  0.85      0.714      0.322         99      72.1%
  0.90      0.690      0.227         78      68.2%
  0.95      0.727      0.115         33      62.8%

7. Voice-Like Punctuation Stress Evaluation

Total stress-testable pairs: 965

Sample stress pairs: act-stop-091: original: "останови сервис, пожалуйста" stressed: "останови сервис пожалуйста" route: action act-stop-092: original: "останови nginx, пожалуйста" stressed: "останови nginx пожалуйста" route: action act-stop-093: original: "останови bотоbackup, пожалуйста" stressed: "останови bотоbackup пожалуйста" route: action act-stop-094: original: "выключи сервис, пожалуйста" stressed: "выключи сервис пожалуйста" route: action act-stop-095: original: "выключи nginx, пожалуйста" stressed: "выключи nginx пожалуйста" route: action act-stop-096: original: "выключи bотоbackup, пожалуйста" stressed: "выключи bотоbackup пожалуйста" route: action act-stop-097: original: "можешь останови сервис, пожалуйста" stressed: "можешь останови сервис пожалуйста" route: action act-stop-098: original: "can you останови сервис, please" stressed: "can you останови сервис please" route: action act-stop-099: original: "можешь останови nginx, пожалуйста" stressed: "можешь останови nginx пожалуйста" route: action act-stop-100: original: "can you останови nginx, please" stressed: "can you останови nginx please" route: action

Impact assessment: Removal of punctuation changes:

  • trailing '?' removal eliminates the strongest question signal
  • lowercase normalization removes proper-noun casing cues
  • whitespace collapse has minimal effect on e5 (subword tokenizer) Production voice punctuation is unreliable; the model must not depend on it.

Capability-question pairs total in stress set: 42 Capability-question pairs with trailing '?': 21 Capability-question pairs with comma only: 21

8. E5 + Tiny Structural Features

Features: [e5(384) ; IsQuestion(1) ; trailing_?(1) ; prohibition(1)] = 387 dims
Macro F1:          0.622
False-action rate: 15.1%
False-action count: 373
best C:            10.0

Comparison with pure e5: pure e5: macro_f1=0.616 FA_rate=15.7% e5 + structural bits: macro_f1=0.622 FA_rate=15.1%

9. Four-Hypothesis Comparison Table

experiment action_P action_R FA rate macro F1
six-way e5 linear (residual) 0.620 0.826 16.0% 0.663
six-way + action threshold 0.727 0.115 1.4%
binary linear action probe 0.726 0.538 6.9%
e5 + tiny structural features 0.610 0.805 15.1% 0.622

10. Conservative Interpretation

Action/capability-question embeddings are nearly indistinguishable

The representation itself is suspect for this boundary. Do not claim the boundary is merely nonlinear. A fundamentally different representation or encoder may be needed.

E5 geometry verdict: action seeds and capability questions are nearly indistinguishable in embedding space. The representation intentionally maps pragmatically different but semantically similar sentences close together. This is a fundamental limitation of the frozen e5 representation for this boundary.

Softmax formulation verdict: the six-way head produces 388 false actions at 16.0% rate. The binary probe produces 172 false actions. The binary probe has 0.4x the false-action count of the six-way head. This confirms the six-way softmax competition is a significant contributor.

11. Commit hash for diagnostic tooling

774c066