Record the destination number and two stuck measurements (V-659)

CLAUDE.md said the destination had no fixture and no accuracy number. It
has both now: intent 73/96 and destination 12/33 on the classifier cascade,
with the per-destination split, the floor cases and the grammar drift the
labelling turned up. Anyone adding a grammar now reads that baselineGrammars
mirrors buildRouter and drifts silently when it does not.

docs/evals/2026-08-08-massive-warm-start.md was written on the V-655 branch
and parked in .task/, which git excludes, so it was one `task start` away
from being lost. It is a dated measurement and it belongs under docs/evals
whatever branch produced it. Its "destination has no fixture at all" line is
now a pointer to the file beside it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
This commit is contained in:
2026-08-08 18:07:41 +04:00
parent b6eaa704a2
commit c15c2b7bd2
3 changed files with 262 additions and 5 deletions
+38 -5
View File
@@ -398,11 +398,44 @@ deliberately do not. "что у меня в списке покупок" matches
the calendar there would take the list source off the turn.
Fixture unchanged at **69/91 classifier+ONNX**, measured both sides. That is the
expected result, because it scores intent and no case here changes intent. **The
destination has no fixture yet, so it has no accuracy number.** That and the model arm
are the follow-ups. The field is designed so a decider naming nothing costs nothing.
It lands on V-546. Intent, mood and BIO slot tags were already three heads on one
forward pass of the resident e5-small. Destination is a fourth head on the same pass.
expected result, because it scores intent and no case here changes intent.
**The destination has its own fixture and its own number as of 08-08-2026**
(V-659, `docs/evals/2026-08-08-destination-fixture.md`). This section used to say
it had neither. `want_source` on `eval.Case` is a pointer, because the destination
has three states and a bare string has two. Absent is every intent but query,
which never reaches `queryWalk`. Present and empty is the `SourceUnknown`
contract: name nothing and walk the chain. Present and named is a destination the
route must produce. Thirty-three of ninety-six cases carry one.
A destination miss does **not** fail the case. It lands in `Outcome.SourceReason`
and never in `Reasons`, so `Accuracy` and `IntentAccuracy` mean what they meant
and `SourceAccuracy` is a second number over the labelled cases only. Intent and
destination are two decisions, and one number hides which one moved. A route that
lost its intent scores no destination hit, or a clarify would satisfy an empty
label for free.
Measured classifier+ONNX: intent **73/96 (76.0%)**, destination **12/33 (36.4%)**.
The split is the finding. World is 5/5, because a stage 0 rule names it. The
`SourceUnknown` floor is 5/7. Calendar is 2/6, because the possessive agenda
rules deliberately do not name it. And **recall is 0/15, because nothing
anywhere names it**. Those turns are still answered, since the chain walks
recall early. Recall is the number the fourth head has to move.
Seven cases assert the floor and six of them are homelab operations. They
cluster because `SourceRecall`, `SourceNetwork` and `SourceAttention` overlap on
every question about the box. `mavpoll` writes its netdata and uptime-kuma
observations into the fact store recall reads. That is a finding about the enum,
not a gap in the labelling.
`baselineGrammars` in `eval_test.go` mirrors `buildRouter` and had drifted:
`WorldQueryGrammars` was wired into the daemon by V-655 and not into the mirror,
so the fixture scored a grammar set nobody runs. Fixed by V-659, worth 3 points of
destination and nothing else. Check that function when adding a grammar.
The model arm is still the follow-up. It lands on V-546. Intent, mood and BIO slot
tags were already three heads on one forward pass of the resident e5-small.
Destination is a fourth head on the same pass.
## LLM output contract
@@ -0,0 +1,82 @@
# The first destination number
Measured 2026-08-08 on the classifier cascade with the ONNX multilingual
embedder, the configuration homesrv runs. `make t PKG=./internal/router/eval/
RUN=TestONNXBaseline V=1`. Covers V-659, the follow-up V-655 named.
## What was measured
V-655 split a routing decision in two. The cascade sorts an utterance into one
of seven intents, and `Decision.Source` then says where the answer lives. The
first half had a fixture. The second half arrived with none, so twelve
destinations shipped with no accuracy number.
`want_source` is now a field on `eval.Case`. It is a pointer, because the
destination has three states and a bare string has two. Absent is every intent
but query, which never reaches `queryWalk`. Present and empty is the
`SourceUnknown` contract: name nothing and let the daemon walk the chain.
Present and named is a destination the route must produce.
Thirty-three of the ninety-six cases carry one. A destination miss does not
fail the case, so `Accuracy` and `IntentAccuracy` mean what they meant.
`SourceAccuracy` is a second number over the labelled cases only.
## Result
Intent is **73/96 (76.0%)**, against 69/91 (75.8%) before. Four of the five new
cases pass and no existing case moved.
Destination is **12/33 (36.4%)**, and the split is the whole finding.
| destination | scored | note |
|---|---|---|
| world | 5/5 | `WorldQueryGrammars` names it at stage 0 |
| the `SourceUnknown` floor | 5/7 | the two misses lost the intent first |
| calendar | 2/6 | `calendar-query` names it, the possessive agenda rules do not |
| recall | 0/15 | nothing anywhere names it |
Recall is the number to move. Fifteen cases ask about his own words and his own
facts. The route lands `query` on eleven of them and the destination comes back
empty every time. Those turns are answered today, because the daemon walks the
chain in order and the three recall passes are early in it. What is missing is a
decider that says so, and that is the fourth head on V-546.
Two cases labelled the floor lost their intent before a destination was
possible. A clarify names nothing, so it would satisfy an empty label for free.
`Score` requires the route to land the case's intent before it credits a
destination hit, or the floor label would score itself.
## Seven cases assert the floor, and six of them cluster
The six are homelab operations. `SourceRecall`, `SourceNetwork` and
`SourceAttention` overlap on every question about the box, because `mavpoll`
writes its netdata and uptime-kuma observations into the fact store recall
reads. "почему сервер тормозит" is answerable from all three. Naming one takes
the other two off the turn.
That is a finding about the enum rather than a gap in the labelling. The floor
is the right answer there and the fixture now says so out loud.
## A drift the labelling found
`WorldQueryGrammars` went into `buildRouter` with V-655 and never into
`baselineGrammars`, the fixture's mirror of it. So the fixture was scoring a
grammar set the daemon does not run. The comment above that function forbids
exactly that. Adding it moved the destination number from 9/33 to 12/33 and
moved nothing else.
The three cases it recovered are `что такое TCP?`, `сколько будет 17 на 23?`
and `кто такой Линус Торвальдс?`. All three already routed `query` through
`NarrativeQueryGrammars`. So the drift was invisible to every number this
fixture reported, until the destination had one of its own.
## What this does not measure
The model arm. This is the classifier cascade, which names a destination only
where a stage 0 rule filled one in. The resident model has no destination in
its router prompt yet, so 36.4% is a floor and not a comparison.
Two pairs of cases are the same utterance. `ru-query-020` and `ru-query-024`
are both "что дальше?", and `ru-query-021` and `ru-query-025` are both
"расскажи про битву при Ватерлоо". They differ in tags and note only, so both
pairs are counted twice here and in every earlier number this fixture reported.
+142
View File
@@ -0,0 +1,142 @@
# MASSIVE Russian warm-start for the routing heads
Measured 2026-08-08 on workpc (Radeon RX 7900 GRE, ROCm). Covers V-546 step 2.
Workspace is `~/Programs/embed-training` on workpc, scripts `train_massive.py`,
`ab_run.py`, `ab.sh`, `probe_time.py`.
## What was trained
Two heads on a copy of multilingual-e5-small: `Linear(384, 60)` for MASSIVE's
own intents over a masked mean pool, `Linear(384, 111)` per token for BIO slot
tags. MASSIVE's label sets verbatim, no alignment to Maven's 7 intents. The
intent head is an auxiliary loss that shapes the pooled vector and is thrown
away.
Data is `amazon-massive-dataset-1.1` pulled from S3. The Hugging Face repo is
script-only and `datasets` 5.0 refuses those, so `load_dataset` cannot fetch it.
`ru-RU` is 11,514 train, 2,033 dev, 2,974 test, 60 intents, 55 slots, 111 BIO
labels. All 16,521 rows survived span alignment: `annot_utt` re-tokenised to its
own `utt` on every one.
Hyperparameters match `train_intent.py`, so the two runs differ in data only.
Frozen XLM-R vocabulary, body 2e-5, heads 1e-3, batch 32, sequence 64, 10
epochs. MASSIVE's own dev partition selects the epoch, on slot F1 with intent
accuracy as tiebreak. Selecting on 60-class intent accuracy would optimise a
head that gets deleted.
## Result
Epoch 9 of 10 by dev slot F1. Held-out MASSIVE test: intent 86.2%, slot span
F1 71.5% (P 68.5, R 74.8). Peak 1.70GB of 17.2GB, about 22 seconds an epoch,
under 4 minutes end to end. Dev slot F1 climbed monotonically to epoch 9 and
fell at 10, so 10 epochs was the right budget.
Ten slot types sit at 0% test recall. Every one of them has 1 to 7 test
instances: `alarm_type` has 3, `drink_type` has 1. That is support in MASSIVE's
Russian split, not a tagger failure. `playlist_name` at 6% of 16 is the first
real miss.
## The intent A/B, and why it settles nothing
`train_intent.py` was run against both bodies, three seeds by two smoothing
settings, on `train_v4.jsonl`. It is v4 and not v5 because v4 is what
`sweep2.log` measured. `ab_run.py` strips a `--base` flag onto the module global, so
`train_intent.py` is unmodified and its baseline stays reproducible. The stock
arm reproduced `sweep2.log` line for line.
Fixture accuracy, 91 cases, one case is 1.1 points:
| seed / smooth | stock | warm-started |
|---|---|---|
| 0 / 0.0 | 94.0% | 92.8% |
| 0 / 0.1 | 95.2% | 92.8% |
| 1 / 0.0 | 95.2% | 94.0% |
| 1 / 0.1 | 95.2% | 97.6% |
| 2 / 0.0 | 92.8% | 94.0% |
| 2 / 0.1 | 92.8% | 96.4% |
Mean 94.2% against 94.6%. That is +0.4 points, about a third of one case, and
inside seed noise. Spread widened. Stock lands in a 2.4-point band and
warm-started in a 4.8-point one. The warm-started arm holds both the best result
of the sweep and a tie for the worst. Seed 0 is the bad arm and it fails in a
specific way. Its dev peaks at epoch 2 and 3 and never improves, where stock
peaks around 7. The dev slice is a quarter of the seed rows. That is small
enough that early stopping is fragile when the body arrives already fitted.
**The A/B was never the test.** Intent had at most 4.8 points of headroom here.
MASSIVE was not trained for Maven's intents. Read it as "the warm-start does not
cost intent accuracy", nothing more.
## The measurement that does mean something
`want_time` is the one slot Maven's fixture scores, and MASSIVE has `time` and
`date`. Restricted to those two slot types, F1 is 74.9% over 609 gold spans on the
MASSIVE ru test split. Precision is 71.5 and recall 78.7. That beats the 71.5%
all-slot figure. Of the 530 test utterances carrying a time or a date, 73.4% get
every such span exactly right.
Out of domain matters more, because Maven's traffic is not this corpus. Ten
Maven-shaped utterances, none of them in MASSIVE:
| utterance | tagged |
|---|---|
| `напомни в 11:00 позвонить маме` | `time='11:00'`, `relation='маме'` |
| `напомни завтра в семь утра выпить таблетки` | `date='завтра'`, `time='семь утра'` |
| `поставь будильник на полседьмого` | `time='полседьмого'` |
| `через двадцать минут напомни про чайник` | `time='двадцать минут'` |
| `напомни в пятницу вечером забрать посылку` | `date='пятницу'`, `timeofday='вечером'` |
| `что у меня сегодня после обеда` | `date='сегодня'`, `time='после'`, `timeofday='обеда'` |
| `запиши что кофе закончился` | nothing |
| `что такое TCP` | `definition_word='TCP'` |
The first row is the V-572 defect utterance. `ReminderGrammar` handed the daemon
`HasTime: false` there, and the daemon asked "Когда?" at a sentence that had
already said when. `полседьмого` is a colloquial half-past that no digit pattern
catches. `запиши что кофе закончился` correctly carries nothing, because a note
has no time.
Two errors. `после обеда` split into `time='после'` plus `timeofday='обеда'`
when it is one span, and `через двадцать минут` dropped its `через`. Both are
boundary errors on spans the tagger did find.
Unplanned: `что такое TCP` returned `definition_word='TCP'`. MASSIVE has a slot
for the thing being asked about, which is a `SourceWorld` signal sitting in a
head already trained.
Ten hand-picked utterances are evidence, not a fixture.
## What this does not measure
Maven has no span fixture. `want_time` and `want_fn` are presence booleans and
`want_fact_key` is an exact string match, so nothing in the repo can score a
71.5% span tagger. Destination got one the same day, at 12/33 on the classifier
cascade: see `2026-08-08-destination-fixture.md`.
The missing span fixture is why the warm-start stays unjudged against Maven
rather than against MASSIVE.
## Datasets ruled out
Checked on 2026-08-08 and rejected as label sources:
- **MASSIVE's other 50 locales** ship in the same tarball and are parallel by id.
Co-training on them is free and unmeasured. English was ruled out by the owner
on 2026-08-08.
- **CLINC150** is reachable as parquet, 150 intents and 1,200 explicit
out-of-scope queries, English only. Its value is the labeled out-of-scope set
for fitting the energy threshold, not intent labels.
- **`d0rj/dolphin-ru`**, roughly 2.8M rows of FLAN-style tasks translated to
Russian. No intent, no slots, and not utterances anyone says to an assistant.
- **`psytechlab/EmpatheticIntents-ru`**, 24,856 rows of translated
EmpatheticDialogues with 32 emotion labels. Maven's mood enum is `neutral,
happy, thinking, tired, confused` and it describes her own reply, not the
speaker's emotion. No mapping exists.
- **`ai-forever/MERA`** and **`RussianNLP/russian_super_glue`**, benchmark
harnesses. Rows are prompt templates with `{toxic_comment}` placeholders.
- **`ZeroAgency/ru-big-russian-dataset`**, an LLM-judge quality corpus. Its
`question` and `classified_topic` columns are a usable Russian out-of-scope
pool for threshold fitting. That is the one thing CLINC150 can only supply in
English. The questions are long and written, so they belong in the negative
set, never in the in-scope `query` training set.
- No second Russian slot-filling corpus exists. The xSID mirrors are 404,
MultiATIS++ has no Russian, SLURP is not on the Hub.