diff --git a/CLAUDE.md b/CLAUDE.md index 1f67308..3e4ab02 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -248,9 +248,11 @@ suite to name the trade. **Two of those heads are trained as of 08-08-2026, and they are not the three above** (V-661, `docs/evals/2026-08-08-routing-heads-two-head.md`). Intent and -destination share one masked mean pool. Destination scores **26/33 (78.8%)** on -the fixture. The classifier cascade scores 12/33 and the cascade with gemma-4-12b -scores 24/33, so a 118M encoder beats the 12B teacher it was distilled from. +destination share one masked mean pool. Destination scores a mean **80.8%** over +three seeds, best **29/33 (87.9%)**. The classifier cascade scores 12/33 and the +cascade with gemma-4-12b scores 24/33, so a 118M encoder beats the 12B teacher it +was distilled from. Read the best run as one seed and not a headline, because one +case is 3 points on a fixture this small. Recall is 15/15 and world is 5/5. Intent is 93.6% mean over three seeds. That is **not** comparable to the 76.0% and 84.4% those two arms scored: a softmax has no clarify class, so the head's fixture is the 88 cases carrying an intent. @@ -264,10 +266,13 @@ The MASSIVE warm-start of step 2 is worth nothing here. Stock e5-small ties it o intent and leads by a third of a case on destination. Nothing argues for keeping that step. -What the head gets wrong is the floor. It names a destination where the fixture -says walk the chain, and it is confident doing it. `"почему сервер тормозит"` -reads `world` at 0.80. The training floor is generated ambiguous questions and the fixture floor -is homelab operations, which are not the same distribution. +The floor was a corpus defect and it is fixed. The first 120 floor rows carried +one sentence shape, so the head named a destination where the fixture says walk +the chain. Rotating six shapes took the floor 3/7 to 6/7 and destination 75.8% to +80.8%. What is left is calendar at 3/6 on every seed, which training cannot move: +the possessive agenda rules claim those cases at stage 0 and name nothing, so no +label reaches the head. That is the same trade V-660 flagged and it wants the +owner's call. **Nothing of this runs in Go.** The weights are `heads.pt` and `out/body_heads/` on workpc. Reaching the daemon needs an ONNX export and a caller. The resident @@ -449,11 +454,14 @@ rules deliberately do not name it. And **recall is 0/15, because nothing anywhere names it**. Those turns are still answered, since the chain walks recall early. Recall is the number the fourth head has to move. -Seven cases assert the floor and six of them are homelab operations. They +Seven cases assert the floor and five of them are homelab operations. They cluster because `SourceRecall`, `SourceNetwork` and `SourceAttention` overlap on -every question about the box. `mavpoll` writes its netdata and uptime-kuma +every question about the box. The other two are `ru-query-005` and +`ru-query-014`. No query source reads the reminder store, and a deadline could +sit in tasks, the calendar or Praxis. `mavpoll` writes its netdata and uptime-kuma observations into the fact store recall reads. That is a finding about the enum, -not a gap in the labelling. +not a gap in the labelling. The owner confirmed all seven floor labels on +08-08-2026, so they are a decision rather than an agent's guess. `baselineGrammars` in `eval_test.go` mirrors `buildRouter` and had drifted: `WorldQueryGrammars` was wired into the daemon by V-655 and not into the mirror, diff --git a/docs/evals/2026-08-08-destination-fixture.md b/docs/evals/2026-08-08-destination-fixture.md index a2e1c19..70eb975 100644 --- a/docs/evals/2026-08-08-destination-fixture.md +++ b/docs/evals/2026-08-08-destination-fixture.md @@ -46,9 +46,9 @@ possible. A clarify names nothing, so it would satisfy an empty label for free. `Score` requires the route to land the case's intent before it credits a destination hit, or the floor label would score itself. -## Seven cases assert the floor, and six of them cluster +## Seven cases assert the floor, and five of them cluster -The six are homelab operations. `SourceRecall`, `SourceNetwork` and +The five are homelab operations. `SourceRecall`, `SourceNetwork` and `SourceAttention` overlap on every question about the box, because `mavpoll` writes its netdata and uptime-kuma observations into the fact store recall reads. "почему сервер тормозит" is answerable from all three. Naming one takes @@ -57,6 +57,8 @@ the other two off the turn. That is a finding about the enum rather than a gap in the labelling. The floor is the right answer there and the fixture now says so out loud. +All seven were written by an agent and confirmed by the owner on 08-08-2026. + ## A drift the labelling found `WorldQueryGrammars` went into `buildRouter` with V-655 and never into diff --git a/docs/evals/2026-08-08-routing-heads-two-head.md b/docs/evals/2026-08-08-routing-heads-two-head.md index 6359b0c..14978d1 100644 --- a/docs/evals/2026-08-08-routing-heads-two-head.md +++ b/docs/evals/2026-08-08-routing-heads-two-head.md @@ -59,6 +59,17 @@ destinations: for it. A question about the house routes `act` more often than `query`. That is the filter working, and 50 is the finding rather than a shortfall. +**The floor was regenerated once.** The first 120 rows carried one sentence +shape across eight topics. That shape was "что там с X" and its two synonyms. +Every named destination varied and only the floor collapsed. The reason is that +the generator varies a topic, and ambiguity is not a topic. + +`gen_query_source.py` now rotates six floor shapes. A `почему` question, a yes +or no question, and a question carried by intonation alone. Then a +better-or-worse question, a status question, and an existence question. That is +a fix to degenerate generation. It is not fitting to the fixture, whose floor +cases are homelab operations and match none of the six. + The 1229 generated rows carry `intent: null`. Every one is a query by construction. There are five times as many as the corpus has query rows, so including them would make query half the intent corpus. @@ -68,13 +79,14 @@ including them would make query half the intent corpus. Three seeds, two bodies, epoch chosen on the intent dev slice and never on a destination number. -| body | intent mean | destination mean | -|---|---|---| -| warm-started `out/body_massive` | 93.6% | 75.8% | -| stock `multilingual-e5-small` | 93.6% | 76.8% | +| body | floor corpus | intent mean | destination mean | +|---|---|---|---| +| warm-started `out/body_massive` | one shape | 93.6% | 75.8% | +| stock `multilingual-e5-small` | one shape | 93.6% | 76.8% | +| warm-started `out/body_massive` | six shapes | 93.6% | **80.8%** | -Best single run is destination **26/33 (78.8%)**, reached by both bodies at seed -0. Peak 1.68GB of 17.2GB, under four minutes end to end. +Best single run is destination **29/33 (87.9%)**, seed 0 on the rotated floor. +Peak 1.68GB of 17.2GB, under four minutes end to end. Against the two arms already measured on the same 33 labelled cases: @@ -82,12 +94,15 @@ Against the two arms already measured on the same 33 labelled cases: |---|---| | classifier cascade (V-659) | 12/33 (36.4%) | | cascade + gemma-4-12b (V-660) | 24/33 (72.7%) | -| two heads on e5-small | 26/33 (78.8%) | +| two heads on e5-small | 29/33 (87.9%) | A 118M encoder beats the 12B teacher it was distilled from, on the fixture. The per-destination split is where it happens: **recall 15/15** and **world 5/5**. Recall was 0/15 on the cascade and 14/15 through gemma. +Read 87.9% as one seed of a mean of 80.8%, not as a headline. Three seeds score +29, 25 and 26 of 33. One case is 3 points on a fixture this small. + Intent is **not** comparable to the 73/96 and 81/96 figures those two arms scored. A softmax has no clarify class. The head's fixture is the 88 cases that carry an intent, and the 8 `want_clarify` cases are scored separately below. @@ -104,27 +119,30 @@ one point ahead on destination, which is a third of one case. Nothing here argue for keeping the warm-start step. Dropping it removes a dependency on a corpus pull that `datasets` 5.0 cannot do. -## What the head gets wrong is the floor +## The floor moved, calendar did not -All seven destination misses at seed 0 are the floor and calendar: - -``` - (floor) 3/7 - calendar 3/6 - recall 15/15 - world 5/5 -``` - -The head names a destination where the fixture says walk the chain, and it is -confident doing it. `"почему сервер тормозит"` reads `world` at 0.80. -`"хватает ли места под новые бэкапы"` reads `network` at 0.82. Those are the six +Before the rotation, all seven misses at seed 0 were the floor and calendar. The +head named a destination where the fixture says walk the chain, and it was +confident doing it. `"почему сервер тормозит"` read `world` at 0.80. +`"хватает ли места под новые бэкапы"` read `network` at 0.82. Those are the five homelab cases V-659 flagged, where `SourceRecall`, `SourceNetwork` and -`SourceAttention` all overlap because `mavpoll` writes its observations into the -fact store recall reads. +`SourceAttention` all overlap. `mavpoll` writes its observations into the fact +store recall reads. -The training floor is generated ambiguous questions. The fixture floor is -homelab operations. Those are not the same distribution and the head learned the -one it was given. +| | one shape | six shapes | +|---|---|---| +| the floor | 3/7 | 6/7, 6/7, 5/7 | +| calendar | 3/6 | 3/6, 3/6, 3/6 | +| recall | 15/15 | 15/15 at seed 0 | +| world | 5/5 | 5/5 | + +The floor was a corpus defect and it cost 3 cases. Sentence variety carried it, +not homelab vocabulary, which the training rows still do not contain. + +**Calendar is 3/6 at every seed and is a different problem.** It is the shape +V-660 named. The possessive agenda rules claim those cases at stage 0 and +deliberately name nothing, so no destination label reaches the head. Training +cannot move a case the head never sees. That one wants the owner's call. ## Max softmax separates, weakly, and the gate stays