# A deterministic structural execution-frame gate blocks every one of the 126 capability-question rows on all three stress views (dangerous pass 0/126), never blocks a true action (0/796), declines 196 malformed action rows as ambiguous, and removes the sparse gate's last residual capability-question pass and three of its nine false actions (guard→sparse P 96.7 %, R 22.2 %, FA 6) — execution eligibility is a deterministic policy boundary, separate from semantic routing and capability selection Date: 2026-09-07 · Task: slice 21 (brief after the accepted slice 20, task/725) · Box: workpc, Arch, RX 7900 GRE, 32 GB (deterministic Go; no model, no ONNX, no LLM in the evaluated path) · Build: `cmd/semantic-router-experiment/slice21` (Go 1.25.12, golem ru vendored via `internal/morph`) + `cmd/semantic-router-experiment/slice21_emit.py` in the frozen venv `/tmp/mvn-exp-venv`. ## 0. Frozen artifacts and population ```text dev pool: 2490 rows, byte-identical to slices 16-20 routes: action 796 / knowledge 715 / memory_write 553 / system 226 / conversation 93 / uncertain 107 exclusive families (family_of priority): direct_imperative 456 / polite_request 460 / modal_request 223 / first_person_request 791 / reordered_target 332 / question 102 / capability_question 126 paired capability/action rows: 2268 (shared object nouns) ``` The emit step writes `/tmp/mvn-s21/{pool.json, pairs.json, sparse_oof.json}`. `pool.json` is the slice-20 dev pool verbatim (all 2490 rows, fast-path included, not residual-only). `sparse_oof.json` is the slice-18 grouped-CV "both" OOF proba, **re-derived, not loaded**: `/tmp/mvn-s18` no longer exists locally, so the emit step re-runs the slice-18 grouped-CV header in Python (`slice18_sparse.load_data/filter_dev_pool/build_features` reused verbatim) and reproduces the accepted slice-18 operating points exactly — P 0.959 / R 0.264 / FA 9 at threshold 0.715, and P 0.875 / R 0.485 / FA 55 at 0.5. The composition numbers in §10 therefore rest on the identical OOF the accepted slice-18 report cited. ## 1. The `ExecutionFrameGuard` contract The guard answers one question: **may this utterance become an executable action at all?** It never decides what the utterance *is* — that stays with the route classifier. The three-way verdict (`Eligibility`) is policy, and the policy is asymmetric on purpose: ```text blocked enough structural evidence that execution must not happen (negation, capability question, reported command, quotation, hypothetical) ambiguous insufficient evidence to grant execution authority (no request evidence, uncertain modal, trailing "?") permissive no blocking frame detected NOT equivalent to "this is an action" ``` Blocked and ambiguous must never execute. Permissive only means the downstream action/semantic selection may look; it is not a label and not a recommendation to act. `Evaluate(text)` is a closed pipeline of nine stages in fixed precedence order, each either returning a final frame or declining to the next: ```text 1 quotation quoted span carries command polarity 2 reported speech past/third-person report frame governing a command clause 3 hypothetical if/если scope with infinitive or subjunctive "бы", no real condition 4 prohibition router.IsCommandProhibition (shipped parser) 5 advisory negation "не надо / не нужно / не стоит / …" 6 capability modal the measured matrix (addressed vs bare, can vs ability, поэтому+ли) 7 trailing "?" question posture with no modal at play → ambiguous 8 request evidence positive trigger → permissive (explicit_request) 9 fallback no evidence at all → ambiguous (no_request_evidence) ``` The `Reason` vocabulary is a closed set of nine values (`command_prohibition`, `capability_question`, `reported_speech`, `quotation`, `hypothetical`, `negated_command`, `explicit_request`, `ambiguous_modal`, `no_request_evidence`). A reason vocabulary addition is a design decision and must land in the report, not a silent new branch. Interface shape (experiment-only, single package): ```go type Eligibility int // Permissive, Blocked, Ambiguous type Reason string // closed set above type Frame struct{ Eligibility; Reasons []Reason } func Evaluate(text string) Frame // stages above, first decisive stage wins ``` ## 2. Reused parsers and helpers No new morphological machinery was written. The guard drives the shipped deterministic routers: | piece | home | role in the guard | | --- | --- | --- | | `router.IsCommandProhibition` | `internal/router/commandframe.go` | stage 4: direct negative commands; also re-scored inside quoted spans and inside `stageReport`'s governed clause | | `router.NormalizeMatchText` | `internal/router/matchtext.go` | the corpus normalization (NFKC, lowercase, whitespace collapse; punctuation and ё kept) applied once per utterance | | `morph.IsVerbForm` / `morph.Lemma` | `internal/morph` (golem ru) | finiteness of the head verb: `finite = IsVerbForm && Lemma != tok`, `infinitive = IsVerbForm && Lemma == tok` | | `lexicon.IsFillerParticle` | `internal/lexicon` | leading politeness/particle stripping before command-position detection | | `lexicon.FirstPerson()` / `lexicon.ReminderVerbs()` | `internal/lexicon` | first-person illocution evidence; the parser's own reminder-exemption scope | `tokens()` mirrors the router's `planTokens` discipline: lowercase, split on everything that is not a letter or digit, so em-dashes and CJK-width punctuation do not create spurious tokens. Multi-token closed expressions ("мог бы", "будь добр") are matched as contiguous lowercased token runs over the reconstructed token text, never over raw text. ## 3. New structural rules (measured discriminators) The six rules below are the encoding of numbers measured on the slice-20 dev pool, not guesses: * **Addressed Russian can/ability forms are capability questions — always blocked.** "ты можешь выключить свет" and "ты можешь выключить свет, пожалуйста" are 42/42 capability-question in dev. Any `ты/вы/тебе/…` + can-form (present or ability) → `blocked/сapability_question`. * **Bare ability forms are capability even polite.** "сможешь открыть окно, пожалуйста", "умеешь ли ты …", "мог(ла) бы …" are 7/7 non-action → blocked. Exception: the leading politeness construction "не мог бы ты …, пожалуйста" (prohibition-parser-exempted modal politeness) stays permissive. * **Bare Russian present can-form needs politeness.** "можешь выключить свет, пожалуйста" is 127/127 action → permissive. Bare "можешь …" with no politeness has no dev rows and reads as a capability offer → ambiguous. * **English can/could needs politeness to be a request.** "can you …, please" (English modal frame around a Russian imperative) is 96/96 action → permissive; bare "can you …" → ambiguous. * **Polar "… ли" after a can-form is a capability question → blocked.** "могу ли я …", "можешь ли ты …", "умеешь ли ты …", "можно ли …". * **Trailing "?" with no modal at play → ambiguous.** In dev: 123/123 bare question rows carry no action label; zero bare-можешь rows end in "?"; only 21 of the 126 capability-question rows end in "?" (the rest are declarative templates like "ты можешь выключить свет"). The two rules never collide. Blocked-stage extras measured to zero action rows: quotation of a command, reported/third-person command clauses, hypothetical (non-real) conditionals, "не надо/не нужно/не стоит" commands, and the reported-speech / hypothetical / negation families have **no** dev-pool rows at all — those rules are covered only by the brief fixtures, and the brief families are where a new generator would over-block or over-pass tomorrow. ## 4. Rule count and code size ```text cmd/semantic-router-experiment/slice21/guard.go 659 lines cmd/semantic-router-experiment/slice21/main.go 566 lines cmd/semantic-router-experiment/slice21/guard_test.go 139 lines cmd/semantic-router-experiment/slice21/fixtures.go 95 lines cmd/semantic-router-experiment/slice21_emit.py 135 lines total 1594 lines ``` Rule engine: **9 stages**, **9 closed reasons**, **13 closed evidence tables** (≈159 lexicon entries + 7 quote-pair delimiters): `wakeAddr` 8, `ruAddress` 9, `enAddress` 3, `ruCanForms` 4, `ruAbilityForms` 26 (18 single + 8 multi), `politeNegativeModal` 6, `enCanForms` 2, `politeness` 7 (3 + 4), `reportVerbs` 39, `reportNouns` 19, `hypothesisMarkers` 5, `illocutionVerbs` 24 (18 + 6). One `maybeWord` matcher handles single/multi closed expressions. No file reads, no network, no model: `Evaluate` is CPU-pure (µs-class; not separately timed because it cannot matter). The engine is exercised by 48 brief fixtures (48/48 passing in `TestEvaluateFixtures`) plus three corpus tests: the 126 capability rows must all be blocked with zero permissive (`TestDevCapabilityProhibition`), the 127 bare можешь+пожалуйста rows must all be permissive (`TestDevBareCanPoliteIsAction`), and the 96 can-you-please rows must all be permissive (`TestDevEnglishCanPoliteIsAction`). ## 5. Full residual confusion table Three-way eligibility × the six routes, orig view, all 2490 dev-pool rows (`guard_results.json → rows.frames.orig`): ```text route perm block ambig n action 600 0 196 796 knowledge 274 126 315 715 memory_write 368 0 185 553 system 106 0 120 226 conversation 49 0 44 93 uncertain 64 3 40 107 total 1461 129 900 2490 ``` Residuals, read off the table: * **permissive non-action: 861** (FA at the binary gate, §6/§10). The guard is deliberately permissive: these are overwhelmingly genuine calendar/keyword queries and "расскажи про …" asks that the sparse selector must decline, not the guard. * **blocked: 129**, of which 126 are the capability-question family and 3 are the junk/fragment conditional forms (§7, §12). **Blocked action = 0.** * **ambiguous: 900**, of which 196 are true-action rows (§12). Ambiguous never executes; these are the guard's recall cost. ## 6. Capability-question dangerous pass The safety case, exactly as the brief frames it: 126 capability-question rows, all non-action, must never clear the gate. ```text capability-question dangerous pass = 0/126 (rate 0.0 %) capability-question blocked 126 capability-question ambiguous 0 ``` Identical on all three stress variants (`orig` / `nofinal` / `strip`): 0 / 126 / 0, 0 / 126 / 0, 0 / 126 / 0. This is the first result in the slice family on which the held-out safety family collapses to zero false fires — the sparse gate's best LOFO number was 0.667 FA; every supervised encoder (e5, slice-19 from-scratch, slice-20 rubert-tiny) sat at or above that, rubert-tiny at 1.000 across all 18 config/seeds. ## 7. Paired capability-question / executable-request results 2268 pairs, a capability question and its executable sibling on the shared object noun (`"ты можешь выключить свет"` vs `"выключи свет"`). ```text pairs 2268: cap cleared 0 (rate 0.0 %) action permissive 1872 action ambiguous 396 action blocked 0 ``` The capability row never clears; 82.5 % of the paired executable siblings stay permissive, and none is ever blocked. The 396 declined siblings are the object-first / nonstandard-verb shapes of §12 and are the same rows the binary gate declines. ### The three blocked non-capability non-actions Of the 129 blocked rows, three are not capability questions: ```text idx 2456 "а если" (family direct_imperative, route uncertain, junk fragment) idx 2459 "а если, пожалуйста" (family polite_request, route uncertain, junk fragment) idx 2462 "я хочу а если" (family first_person_request,route uncertain, junk fragment) ``` These are corpus fragments (a conditional opener with nothing after it). The hypothetical stage blocks them, which is the correct conservative call — a lone "если" opener never grants execution. ## 8. Per-generator-family results The exclusive family view (`family_of` priority; counts are ≤ the slice-20 tag frequencies because one tag wins per row): ```text family n perm block ambig direct_imperative 456 180 1 275 polite_request 460 265 1 194 modal_request 223 223 0 0 first_person_request 791 790 1 0 reordered_target 332 3 0 329 capability_question 126 0 126 0 question 102 0 0 102 ``` * **modal_request / first_person_request**: essentially 100 % permissive — "можно …", "я хочу …", "надо …", "можешь …, пожалуйста" all carry request evidence by construction. * **question**: 100 % ambiguous — trailing "?" or no request evidence; 0 permissive (a question never gains execution authority structurally). * **capability_question**: 100 % blocked. * **direct_imperative / polite_request**: split between permissive (leading finite verb) and ambiguous (object-first, nonstandard head verbs, §12). * **reordered_target**: 1 % permissive. The 3 permissive rows are exactly the wake-word-address imperative rows the corpus generates ("maven останови контейнер", "maven покажи статус", "maven проверь статус"); the 329 ambiguous rows are the object-first classifier rows ("сервис останови"). ## 9. Punctuation stress across all three variants Each row is scored three ways: `orig` (as stored), `nofinal` (trailing `[?.!,;:]+` stripped), and `strip` (all punctuation stripped). The guard's stress axis is the same one slices 18-20 trained on (strip) and evaluated on. ```text dangerous-pass blocked ambiguous orig 0 126 0 nofinal 0 126 0 strip 0 126 0 ``` The 126-row safety family is frozen across the three views, and composition on strip (§10) matches composition on orig exactly (P 96.7 % / R 22.2 % / FA 6). Unlike the sequence models, the guard does not need its training view to match its serving view: tokens are punctuation-exposed by construction and no tokenizer can drop a boundary comma, so the stress axis is structurally free. ## 10. Sparse-alone vs guard-alone vs guard→sparse The composition is a one-way valve: sparse may act only when the guard already said permissive, at the accepted slice-18 strict operating point (@0.715). ```text policy P R FA FA rate capQ capQ rate guard alone 41.1% 75.4% 861 34.6% 0 0.0% sparse alone 95.9% 26.4% 9 0.4% 1 0.8% guard → sparse 96.7% 22.2% 6 0.2% 0 0.0% ``` Stress: guard→sparse on `strip` is identical to `orig` (P 96.7 %, R 22.2 %, FA 6, capQ 0). ### Composition conclusion The guard does **not** replace action selection: ```text guard alone: deliberately permissive P 41.1 % R 75.4 % FA 861 sparse alone: P 95.9 % R 26.4 % FA 9 guard → sparse: P 96.7 % R 22.2 % FA 6 ``` The guard runs *before* the sparse selector and removes structurally unsafe speech acts (the capability-question family — including the one residual capability-question pass sparse alone still made — and three of sparse's nine false action passes) at the cost of 4.2 recall points. Precision 96.7 % is the best operating point measured anywhere on this pool. The guard is a policy boundary that the selector sits behind, not a rival boundary. ## 11. Remaining dangerous-pass categories The in-pool dangerous pass is zero, and the composition removes even the sparse gate's single residual capability-question pass. The structural risk that remains is not in this pool: * **Deliberate permissive expedites.** Real-condition hypotheticals ("если будет дождь, выключи полив") and English "can you …, please" are permissive by design because those frames carry request evidence; a *non*-command utterance built from the same frame (a causality statement, a paraphrase) would read permissive too. No such rows exist in dev. * **Closed-list blind spots.** `reportVerbs` (39) and every closed table are finite; a new report verb, a new politeness form, a new modal not in the lexicons will fall through to request-evidence and be judged on finiteness alone. The brief fixtures carry the families that have no dev rows for this reason. * **New generators.** All rules were measured on the frozen corpus. A new utterance generator with an unseen template is unmeasured by definition; the guard's offence is bounded by the four e5/rubert/scanner-decided columns (the corpus templates are closed-class), but that bound is a corpus fact, not yet a system fact. * The guard is **not** a model and leaks nothing via confidence: its verdicts are deterministic, so any residual pass is replicable and classifiable, not a probability tail. None of these categories currently produce a dangerous pass in-pool, and each is a report-table row if a future corpus adds the generator. ## 12. True-action overblocking / ambiguity categories 196 of 796 true-action rows (24.6 %) are declined as ambiguous (0 are blocked). The head-token distribution of the 196 is concentrated in malformed / grammar-broken shapes, not ordinary well-formed requests: ```text denied actions by family: reordered_target 93 / direct_imperative 67 / polite_request 36 head-token top: в 19 · перезагрузи 12 · nginx 10 · купить 9 · оплатить 9 · статус 8 · на 7 · старт 6 · docker 6 · что 6 · логи 6 · лог 6 · … ``` Three structural categories account for the 196: * **Object-first / preposition-led classifier rows** (~100 of 196): the head token is a noun or preposition, so no command-position verb exists — "сервис останови", "nginx останови", "в лог посмотри", "статус покажи". These are the classifier's reordered_target training shapes. No request evidence → ambiguous. * **Nonstandard service-verb heads** (26): "старт сервис", "старт nginx", "перезапусти контейнер maven", "перезагрузи контейнер maven", "close …". These heads are absent from the golem verb lexicon or are identity-lemma entries (перезапусти/перезагрузи), so the finiteness test fails even on well-formed syntax ("перезапусти контейнер maven" is a grammatical command and is still declined). * **Dictionary-infinitive headword used as a brisk command** (26): "купить …", "оплатить …", "починить …", "выключить …". The lemma of an infinitive headword equals the token, which the guard reads as non-finite. All 196 are refusable, non-executing outcomes — the asymmetry holds (0 action rows end up blocked). The recall cost is real but bounded: at the composition operating point it costs the sparse selector 4.2 recall points (§10), and the categories are structural, so each is addressable by lexicon entry or by an explicit object-first word-order arm if recall becomes the binding constraint. ## 13. Final assessment The slice's acceptance criteria, restated and closed: ```text capability-question dangerous pass 0/126 ✓ (on all three stress views) capability row in every paired pair 0/2268 ✓ true actions blocked 0/796 ✓ reuse-first (no new parsers) yes ✓ only closed evidence sets added deterministic, reportable reasons yes ✓ closed 9-value Reason set 48/48 brief fixtures pass ✓ fixtures invertible to code ``` * The deterministic structural path solves the safety split that every supervised route (frozen e5, from-scratch slice-19, pretrained slice-20) failed categorically, and it does so with zero parameters, µs-class latency, and no training-view/serving-view stress gap. * The asymmetric policy is the correct shape for the task: permissive means "no blocking frame", never "this is an action"; the coarse route classifier and the later action selector still decide what the utterance is and whether hardware moves. * The guard's own precision is deliberately low (P 41.1 % alone); it is a filter, not a selector, and its value only appears in the composition (FA 9 → 6 at P 96.7 %). * Nothing here records the sparse gate as production-ready. Its remaining recall (R 22.2 % under the guard, 26.4 % alone) is still low, and its generator-generalization problem (LOFO capability FA 0.667, aggregate FA pattern) remains a known, unresolved fact of the pool. ## 14. Operational implications / next architectural step If the numbers stand (they are re-derivable from the committed runner), the decision recorded for the routing architecture is: ```text execution eligibility is a deterministic policy boundary, separate from semantic routing and capability selection. TryFastPath ↓ miss ExecutionFrameGuard ├─ blocked/ambiguous → action unavailable └─ permissive ↓ action/semantic selection ``` Operational notes for whoever wires this later: * Blocked and ambiguous are both non-executing; the phraser may answer both, but no action engine may fire on either. The trigger wiring (`cmd/mavend/ecosystem_acts.go`) is out of scope for this slice. * The guard is CPU-pure and deterministic; it belongs in the daemon process, not in a subprocess. * A report-logic change must come with a fixture and a row in this file's pattern, not a silent branch. **Next slice: stop the action-pragmatics model branch.** Four inductive biases (frozen e5, from-scratch sequence, pretrained fine-tune, now deterministic structure) have been measured on the same 126-row holdout; the supervised ones cap at 0.48-1.000 FA and the deterministic one closes it to 0. There is no fifth model experiment left that the three failures do not already rule out. Recommended work moves back to the coarse non-action router — `conversation / knowledge / memory_write / system / uncertain` — with executable action eligibility handled separately by `deterministic fast path + ExecutionFrameGuard + later action selector`, and with recall uplift for the selector (the §12 lexicon/word-order arms are the cheap lever) rather than more pragmatics training. ## 15. Commit hashes * Tooling (engine, fixtures, tests, runner, emit step): `fa98e47` (`router/semantic: slice 21 deterministic execution-frame guard engine, fixtures, runner and emit step`). * Report + eval index: this file paired with its `docs/evals/CLAUDE.md` row. * Re-derived artifacts under `/tmp/mvn-s21/` (not committed; reproducible by `slice21_emit.py` then `go run ./cmd/semantic-router-experiment/slice21`).