Files
Maven/docs/evals/2026-09-07-nonlinear-e5-mlp-probe.md
T

16 KiB

Slice 17: Nonlinear MLP Probe — Action Gate Experiment

0. Frozen Artifacts from Slice 16

development corpus v2 hash: b27fd48f478ca477
original frozen holdout hash: ad297fbdbbea704b (byte-identical, uninspected)
embedder: model_quantized@384/tok2
dimension: 384
pooling: mean-pool + L2-normalize
input template: query: <text>
embedding file: /tmp/mvn-experiment/embeddings.json
total examples: 3025
dev pool: 2490
frozen holdout: 535
router-residual: 2943

1. Hidden Sizes and Exact Parameter Counts

Binary action gate: e5[384] → Linear(384→H) → ReLU → Linear(H→1)

H params fp32 bytes int8 bytes

8 3,089 12,356 3,089 16 6,177 24,708 6,177 32 12,353 49,412 12,353 64 24,705 98,820 24,705

Six-way MLP: e5[384] → Linear(384→H) → ReLU → Linear(H→6)

H params fp32 bytes int8 bytes

8 3,134 12,536 3,134 16 6,262 25,048 6,262 32 12,518 50,072 12,518 64 25,030 100,120 25,030

2. Selected Regularization

Best binary MLP: H=32, weight_decay=0.01 Selected by grouped development CV PR-AUC.

Grid results (binary MLP):

H wd PR-AUC ROC-AUC action_P action_R FA count FA rate

16 0.0 0.422 0.628 0.139 0.215 155 6.2% 16 0.0001 0.422 0.628 0.139 0.215 155 6.2% 16 0.001 0.423 0.628 0.139 0.215 155 6.2% 16 0.01 0.423 0.628 0.139 0.215 155 6.2% 32 0.0 0.691 0.854 0.686 0.610 219 8.8% 32 0.0001 0.691 0.854 0.686 0.610 219 8.8% 32 0.001 0.691 0.854 0.686 0.611 219 8.8% 32 0.01 0.692 0.854 0.686 0.605 220 8.8% * 64 0.0 0.547 0.714 0.309 0.381 195 7.8% 64 0.0001 0.547 0.714 0.309 0.381 195 7.8% 64 0.001 0.547 0.714 0.309 0.381 195 7.8% 64 0.01 0.545 0.713 0.308 0.368 194 7.8%

3. Binary MLP OOF Metrics (best: H=32, wd=0.01)

ROC-AUC:           0.854
PR-AUC:            0.692
action precision:  0.686
action recall:     0.605
false-positive:    220
false-negative:    343
false-action rate: 8.8%
total:             2490
mean iters:        64

4. Fold Variance

Binary MLP (H=32, wd=0.01):

fold ROC-AUC PR-AUC action_P action_R FP FN n

0    0.960    0.915      0.935      0.701    7   43   476
1    0.896    0.917      0.840      0.861   33   28   341
2    0.953    0.939      0.988      0.286    1  212   707
3    0.825    0.454      0.440      0.767   84   20   506
4    0.638    0.236      0.228      0.412   95   40   460

Fold variance comparison with linear binary probe: fold linear FP MLP FP linear FA% MLP FA%

0          2          7        0.4%       1.5%
1         12         33        3.5%       9.7%
2          0          1        0.0%       0.1%
3         71         84       14.0%      16.6%
4         87         95       18.9%      20.7%

5. Safety Operating Curve (best binary MLP)

threshold action_P action_R FA count FA rate coverage

 0.300     0.6224     0.7827        378     0.1518     0.9305
 0.325     0.6324     0.7563        350     0.1406     0.9221
 0.350     0.6400     0.7349        329     0.1321     0.9153
 0.375     0.6462     0.6997        305     0.1225     0.9040
 0.400     0.6536     0.6709        283     0.1137     0.8948
 0.425     0.6540     0.6482        273     0.1096     0.8876
 0.450     0.6631     0.6231        252     0.1012     0.8795
 0.475     0.6714     0.5930        231     0.0928     0.8699
 0.500     0.6731     0.5691        220     0.0884     0.8622
 0.525     0.6801     0.5503        206     0.0827     0.8562
 0.550     0.6866     0.5339        194     0.0779     0.8510
 0.575     0.6887     0.5113        184     0.0739     0.8438
 0.600     0.6968     0.4849        168     0.0675     0.8353
 0.625     0.6910     0.4523        161     0.0647     0.8249
 0.650     0.6914     0.4334        154     0.0618     0.8189
 0.675     0.6930     0.4083        144     0.0578     0.8108
 0.700     0.6856     0.3781        138     0.0554     0.8012
 0.725     0.6887     0.3530        127     0.0510     0.7932
 0.750     0.6870     0.3254        118     0.0474     0.7843
 0.775     0.7071     0.3003         99     0.0398     0.7763
 0.800     0.7204     0.2751         85     0.0341     0.7683
 0.825     0.7159     0.2374         75     0.0301     0.7562
 0.850     0.7422     0.2098         58     0.0233     0.7474
 0.875     0.7391     0.1709         48     0.0193     0.7349
 0.900     0.7518     0.1332         35     0.0141     0.7229
 0.925     0.7684     0.0917         22     0.0088     0.7096
 0.950     0.7778     0.0528         12     0.0048     0.6972

No threshold achieves action_P ≥ 0.95.

6. Capability-Question Boundary Diagnostic

capability_question → false action rate: 91.3% (115/126) positive modal request → action recall: 0.569 (796 examples)

Paired capability-question vs action examples

These are the critical diagnostic pairs:

text route predicted action_P FA?

ты можешь выключить свет knowledge action 0.8159 YES ты можешь выключить жалюзи knowledge action 0.9752 YES ты можешь выключить вытяжку knowledge action 0.9720 YES ты можешь выключить вентилятор knowledge action 0.9041 YES умеешь ли включить свет knowledge action 0.5618 YES умеешь ли включить жалюзи knowledge action 0.9195 YES умеешь ли включить вытяжку knowledge action 0.8956 YES умеешь ли включить вентилятор knowledge action 0.6122 YES сможешь открыть свет knowledge action 0.7012 YES сможешь открыть жалюзи knowledge action 0.9633 YES сможешь открыть вытяжку knowledge action 0.9596 YES сможешь открыть вентилятор knowledge action 0.8491 YES ты можешь выключить свет, пожалуйста knowledge action 0.8682 YES ты можешь выключить жалюзи, пожалуйста knowledge action 0.9778 YES ты можешь выключить вытяжку, пожалуйста knowledge action 0.9718 YES ... перезапусти контейнер maven action action 0.8106 перезапусти контейнер nexus action action 0.8002 перезапусти контейнер redis action action 0.6867 перезапусти контейнер postgres action action 0.6016 перезагрузи контейнер maven action action 0.6810 перезагрузи контейнер nexus action action 0.6192 перезагрузи контейнер redis action action 0.5051 перезагрузи контейнер postgres action not_action 0.4449 перезапусти контейнер maven, пожалуйста action action 0.8442 перезапусти контейнер nexus, пожалуйста action action 0.8672

7. Voice-Like Stress Results

Total stress-testable pairs: 965
question with ?:              123
question without ?:           242
positive polite/modal action:  350
capability-question pairs:     42

Sample affected pairs:

act-stop-091: original: "останови сервис, пожалуйста" stressed: "останови сервис пожалуйста" route: action act-stop-092: original: "останови nginx, пожалуйста" stressed: "останови nginx пожалуйста" route: action act-stop-093: original: "останови bотоbackup, пожалуйста" stressed: "останови bотоbackup пожалуйста" route: action act-stop-094: original: "выключи сервис, пожалуйста" stressed: "выключи сервис пожалуйста" route: action act-stop-095: original: "выключи nginx, пожалуйста" stressed: "выключи nginx пожалуйста" route: action act-stop-096: original: "выключи bотоbackup, пожалуйста" stressed: "выключи bотоbackup пожалуйста" route: action act-stop-097: original: "можешь останови сервис, пожалуйста" stressed: "можешь останови сервис пожалуйста" route: action act-stop-098: original: "can you останови сервис, please" stressed: "can you останови сервис please" route: action act-stop-099: original: "можешь останови nginx, пожалуйста" stressed: "можешь останови nginx пожалуйста" route: action act-stop-100: original: "can you останови nginx, please" stressed: "can you останови nginx please" route: action

Impact assessment: The MLP's decision boundary must not depend on punctuation artifacts. If MLP success depends on '?' presence, it will fail under voice input.

8. Six-Way MLP Results

Best six-way MLP: H=64, wd=0.001

accuracy:          68.7%
macro F1:          0.681
action precision:  0.639
action recall:     0.830
false-action rate: 15.0%
false-action count: 374

Per-route F1:

action P=0.639 R=0.830 F1=0.722 (n=796) conversation P=0.893 R=0.720 F1=0.798 (n=93) knowledge P=0.735 R=0.634 F1=0.681 (n=715) memory_write P=0.743 R=0.591 F1=0.659 (n=553) system P=0.646 R=0.597 F1=0.621 (n=226) uncertain P=0.583 R=0.626 F1=0.604 (n=107)

Confusion matrix (rows=expected, cols=predicted): action conversation knowledge memory_write system uncertain action 661 3 67 43 7 15 conversation 0 67 9 3 5 9 knowledge 181 0 453 47 33 1 memory_write 132 0 48 327 29 17 system 37 3 34 11 135 6 uncertain 24 2 5 9 0 67

Fold variance (six-way MLP):

fold accuracy macro_f1 action_P action_R FA

0      82.1%      0.774      0.871      0.938    20
1      81.5%      0.625      0.838      0.925    36
2      64.2%      0.580      0.722      0.771    88
3      69.0%      0.687      0.425      0.895   104
4      51.7%      0.452      0.212      0.500   126

9. Comparison Against All Linear Baselines

model params beyond e5 action P action R FA rate macro F1
6-way linear 2,310 0.629 0.832 15.7% 0.675
binary linear ~385 0.688 0.476 6.9%
structural linear ~2,313 0.610 0.805 15.1% 0.622
binary MLP (H=32) 12,353 0.673 0.569 8.8%
6-way MLP (H=64) 25,030 0.639 0.830 15.0% 0.681

10. Conclusion

Verdict: e5 representation inadequate for the action-pragmatics boundary

The tiny nonlinear MLP provides only marginal improvement over linear baselines and fails the capability-question diagnostic and the safety operating goal.

  1. Binary MLP does not beat binary linear. PR-AUC 0.692 (MLP, H=32) vs 0.707 (linear). The MLP trades precision for recall (P=0.673/R=0.569) and produces more false actions (220 vs 172) than the linear probe. The nonlinearity does not unlock a separability that the linear surface was missing — both are pinned by the same representation.

  2. Capability-question boundary is catastrophically broken for both. The MLP fires action on 91.3% (115/126) of capability questions — the primary diagnostic fails outright. The linear probe reaches the same place (112/126). A nonlinear head cannot separate "ты можешь выключить свет?" (knowledge) from "можешь выключить свет, пожалуйста" (action) because the frozen mean-pooled e5 vector places them on top of each other.

  3. No safe operating region exists. The MLP never reaches action precision ≥ 0.95 with materially non-zero recall. The best it can do is P=0.778 at R=0.053 (essentially zero recall). At any threshold that keeps FA count low, it becomes useless; at any threshold that keeps recall useful, FA rate climbs past 10%.

  4. Six-way MLP is essentially flat. F1=0.681 (H=64) vs 0.675 (linear six-way); FA rate 15.0% vs 15.7%. Per-route F1 is within noise of the linear head. The bad folds from slice 16 (fold 3 P=0.425, fold 4 P=0.212) are not repaired.

Per the brief's decision rule this is the third case:

mean-pooled e5 representation is inadequate for Maven's action-pragmatics boundary. at that point stop probing e5.

Stop probing the frozen mean-pooled e5-small representation. The signal is not present and not merely nonlinear. A fundamentally different representation or encoder is required for the action/capability-question pragmatics boundary.

Capability-question diagnostic

capability_question false-action rate (MLP): 91.3% (115/126) capability_question false-action rate (linear): 88.9% (112/126) positive modal action recall (MLP): 0.569 positive modal action recall (linear): 0.476

The MLP does not materially improve the capability-question boundary.

Six-way MLP vs binary gate

Six-way MLP (F1=0.681) and six-way linear (F1=0.675) are statistically indistinguishable. Neither the six-way formulation nor the nonlinearity fixes the boundary. The architecture evidence does not support routing through a single learned head for Maven's action pragmatics as currently embedded.

11. Commit hash for diagnostic tooling

e80d45f (slice 17 MLP probe tooling: cmd/semantic-router-experiment/slice17_mlp.py)