Compare commits
18 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 9a333b23d7 | |||
| 663b5c47b9 | |||
| 45c521e1a6 | |||
| e34669a52e | |||
| d434f83c2c | |||
| c310115fd2 | |||
| 3024f76e5f | |||
| 6bc71553ab | |||
| ed1730431c | |||
| 01e80fce4a | |||
| e69f1bd0cf | |||
| f55bedee2e | |||
| e470435cf1 | |||
| 00f9239ef9 | |||
| 3513e508b7 | |||
| c15c2b7bd2 | |||
| b6eaa704a2 | |||
| 2597a7b34a |
@@ -246,6 +246,72 @@ was a hardcode. **Fine-tune a copy of the weights.** The resident embedder backs
|
||||
recall. Training it in place couples routing accuracy to recall@1, with nothing in the
|
||||
suite to name the trade.
|
||||
|
||||
**Two of those heads are trained as of 08-08-2026, and they are not the three
|
||||
above** (V-661, `docs/evals/2026-08-08-routing-heads-two-head.md`). Intent and
|
||||
destination share one masked mean pool. Destination scores a mean **80.8%** over
|
||||
three seeds, best **29/33 (87.9%)**. The classifier cascade scores 12/33 and the
|
||||
cascade with gemma-4-12b scores 24/33, so a 118M encoder beats the 12B teacher it
|
||||
was distilled from. Read the best run as one seed and not a headline, because one
|
||||
case is 3 points on a fixture this small.
|
||||
Recall is 15/15 and world is 5/5. Intent is 93.6% mean over three seeds. That is
|
||||
**not** comparable to the 76.0% and 84.4% those two arms scored: a softmax has no
|
||||
clarify class, so the head's fixture is the 88 cases carrying an intent.
|
||||
|
||||
**A fourth head asks instead of guessing, same day** (V-661,
|
||||
`docs/evals/2026-08-08-clarify-head-four-head.md`). Clarify is not a value
|
||||
of intent, so a softmax cannot emit it. It is a second question over the
|
||||
same pooled vector: can Maven act on this at all. That is why the head's
|
||||
fixture was 88 cases and not 96. Over three seeds it catches **7.0 of the 8
|
||||
`want_clarify` cases and produces 2.3 false clarifies of 88**. The cascade
|
||||
today misses 1 and produces 2, so this is parity with no rules in front of
|
||||
it. Accuracy is the wrong number here and a head that never asks scores
|
||||
91.7%. Confidence is the other half. Max softmax over the intent head reads
|
||||
**0.851 where it is right against 0.604 where it is wrong**, ranking right
|
||||
above wrong in 83.4% of pairs. `Confidence: 1.0` was a hardcode, and this
|
||||
replaces it with a signal. The two are not the same signal: one says which
|
||||
intent is unclear, the other says the utterance carries too little to act
|
||||
on. **The fourth head is not free the way the third was.** Intent,
|
||||
destination and slot F1 each move down one to four points, inside the seed
|
||||
spread. `поужинал` is a false clarify on every seed, which is the same
|
||||
defect `thinSingleToken` was narrowed for on 2026-08-01.
|
||||
|
||||
The corpus for it is generated, because every existing row is answerable by
|
||||
construction. **The router-prompt agreement filter cannot work here**, since
|
||||
`routeGrammar` has no clarify value and a generated line always agrees with
|
||||
itself. A gemma judge replaces it. The first judge called 24 of 40
|
||||
answerable rows underspecified, because it judged against a generic
|
||||
assistant rather than against Maven's contract.
|
||||
|
||||
**Mood is cut, not deferred.** The enum describes her own reply state, not the
|
||||
speaker's emotion, and no dataset maps onto it.
|
||||
|
||||
**A third head landed the same day** (`docs/evals/2026-08-08-slot-head-three-head.md`).
|
||||
BIO slot tags had no Maven-domain corpus, which was true of found corpora and
|
||||
false of made ones. `label_slots.py` distils spans out of gemma-4-12b under a
|
||||
GBNF closed over Maven's own five slots. A span survives only when it is a
|
||||
literal substring of the utterance, so the agreement filter costs no second
|
||||
call. 2178 spans over 1702 rows, 37 dropped, nothing unparsed. Three heads score
|
||||
intent **92.8%**, destination **82.8%** and slot span F1 **72.4%** over three
|
||||
seeds. The slot head is free: both other numbers move less than their own seed
|
||||
spread. Epoch selection reads the intent dev slice alone. Slot F1 is still
|
||||
climbing when it stops, which costs about 4 points.
|
||||
|
||||
The MASSIVE warm-start of step 2 is worth nothing here. Stock e5-small ties it on
|
||||
intent and leads by a third of a case on destination. Nothing argues for keeping
|
||||
that step.
|
||||
|
||||
The floor was a corpus defect and it is fixed. The first 120 floor rows carried
|
||||
one sentence shape, so the head named a destination where the fixture says walk
|
||||
the chain. Rotating six shapes took the floor 3/7 to 6/7 and destination 75.8% to
|
||||
80.8%. What is left is calendar at 3/6 on every seed, which training cannot move:
|
||||
the possessive agenda rules claim those cases at stage 0 and name nothing, so no
|
||||
label reaches the head. That is the same trade V-660 flagged and it wants the
|
||||
owner's call.
|
||||
|
||||
**Nothing of this runs in Go.** The weights are `heads.pt` and `out/body_heads/`
|
||||
on workpc. Reaching the daemon needs an ONNX export and a caller. The resident
|
||||
e5-small must not be replaced by the copy, because recall depends on that file.
|
||||
|
||||
`Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
|
||||
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
|
||||
#359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with
|
||||
@@ -376,7 +442,19 @@ queries exactly as it did.
|
||||
That is the safety argument and it is not negotiable. The table's order is
|
||||
load-bearing. Every comment on it argues a reason between two sources, and above all
|
||||
it carries "the owner's data first, then the world". Naming `SourceWorld` does not
|
||||
send the turn outside. His notes, his facts and the personal boundary still run first.
|
||||
send the turn outside on its own. His notes and his facts still run first, because
|
||||
they look rather than guess.
|
||||
|
||||
**The personal boundary is the one exception and it is deliberate.** It guesses,
|
||||
so naming `SourceWorld` drops it. That is what stops it answering "кто такой
|
||||
Линус Торвальдс?" with "не нашла у тебя такой записи", which it did on
|
||||
2026-08-07. The cost is that a destination a model wrote can now take the
|
||||
boundary off a turn. A question about him that the model calls `world` reaches
|
||||
SearXNG, where today the boundary stops it. Only the utterance leaves the box,
|
||||
never his notes or history, so this widens what is asked and not what is sent.
|
||||
`TestNamingRecallKeepsTheBoundary` pins the other half: naming `SourceRecall`
|
||||
keeps the boundary in front of the world. Whether a model may drop it at all is
|
||||
the owner's call and has not been made.
|
||||
|
||||
What comes out is only the sources that **guess**. Those decide a turn is theirs by
|
||||
cosine against frozen seeds, then answer whatever they claimed. They hold no table
|
||||
@@ -398,11 +476,63 @@ deliberately do not. "что у меня в списке покупок" matches
|
||||
the calendar there would take the list source off the turn.
|
||||
|
||||
Fixture unchanged at **69/91 classifier+ONNX**, measured both sides. That is the
|
||||
expected result, because it scores intent and no case here changes intent. **The
|
||||
destination has no fixture yet, so it has no accuracy number.** That and the model arm
|
||||
are the follow-ups. The field is designed so a decider naming nothing costs nothing.
|
||||
It lands on V-546. Intent, mood and BIO slot tags were already three heads on one
|
||||
forward pass of the resident e5-small. Destination is a fourth head on the same pass.
|
||||
expected result, because it scores intent and no case here changes intent.
|
||||
|
||||
**The destination has its own fixture and its own number as of 08-08-2026**
|
||||
(V-659, `docs/evals/2026-08-08-destination-fixture.md`). This section used to say
|
||||
it had neither. `want_source` on `eval.Case` is a pointer, because the destination
|
||||
has three states and a bare string has two. Absent is every intent but query,
|
||||
which never reaches `queryWalk`. Present and empty is the `SourceUnknown`
|
||||
contract: name nothing and walk the chain. Present and named is a destination the
|
||||
route must produce. Thirty-three of ninety-six cases carry one.
|
||||
|
||||
A destination miss does **not** fail the case. It lands in `Outcome.SourceReason`
|
||||
and never in `Reasons`, so `Accuracy` and `IntentAccuracy` mean what they meant
|
||||
and `SourceAccuracy` is a second number over the labelled cases only. Intent and
|
||||
destination are two decisions, and one number hides which one moved. A route that
|
||||
lost its intent scores no destination hit, or a clarify would satisfy an empty
|
||||
label for free.
|
||||
|
||||
Measured classifier+ONNX: intent **73/96 (76.0%)**, destination **12/33 (36.4%)**.
|
||||
The split is the finding. World is 5/5, because a stage 0 rule names it. The
|
||||
`SourceUnknown` floor is 5/7. Calendar is 2/6, because the possessive agenda
|
||||
rules deliberately do not name it. And **recall is 0/15, because nothing
|
||||
anywhere names it**. Those turns are still answered, since the chain walks
|
||||
recall early. Recall is the number the fourth head has to move.
|
||||
|
||||
Seven cases assert the floor and five of them are homelab operations. They
|
||||
cluster because `SourceRecall`, `SourceNetwork` and `SourceAttention` overlap on
|
||||
every question about the box. The other two are `ru-query-005` and
|
||||
`ru-query-014`. No query source reads the reminder store, and a deadline could
|
||||
sit in tasks, the calendar or Praxis. `mavpoll` writes its netdata and uptime-kuma
|
||||
observations into the fact store recall reads. That is a finding about the enum,
|
||||
not a gap in the labelling. The owner confirmed all seven floor labels on
|
||||
08-08-2026, so they are a decision rather than an agent's guess.
|
||||
|
||||
`baselineGrammars` in `eval_test.go` mirrors `buildRouter` and had drifted:
|
||||
`WorldQueryGrammars` was wired into the daemon by V-655 and not into the mirror,
|
||||
so the fixture scored a grammar set nobody runs. Fixed by V-659, worth 3 points of
|
||||
destination and nothing else. Check that function when adding a grammar.
|
||||
|
||||
**The model arm landed the same day** (V-660,
|
||||
`docs/evals/2026-08-08-destination-model-arm.md`). `routeGrammar` carries a
|
||||
`source` rule closed over `router.Sources` plus the empty floor, so the model
|
||||
cannot emit a destination that does not exist. The prompt lists the twelve in
|
||||
Russian and says `""` is a normal answer to give often. `LLMRouter.Route` reads it
|
||||
back through `ValidSource` and on `IntentQuery` alone. Against gemma-4-12b on the
|
||||
workstation the cascade scores destination **24/33 (72.7%)** with intent unmoved
|
||||
at 84.4%, and **recall goes 0/15 to 14/15**. The resident Qwen3-1.7B is
|
||||
unmeasured, because it binds `--port 0` inside the container.
|
||||
|
||||
**Stage 0 now costs four destination points.** It did not before. The four cases
|
||||
the cascade loses and the model alone wins are all calendar. The possessive
|
||||
agenda rules claim them first and name nothing on purpose. That caution was free
|
||||
while nothing downstream could name anything either. It is not free now, and the
|
||||
fix is the owner's call rather than a quiet edit.
|
||||
|
||||
The last arm is V-546. Intent, mood and BIO slot tags were already three heads on
|
||||
one forward pass of the resident e5-small. Destination is a fourth head on the
|
||||
same pass, and 72.7% from a 12B teacher is the label source for training it.
|
||||
|
||||
## LLM output contract
|
||||
|
||||
|
||||
@@ -0,0 +1,123 @@
|
||||
# A clarify head, and a confidence that is not a hardcode
|
||||
|
||||
Measured 2026-08-08 on workpc, the same day and the same fixtures as
|
||||
`2026-08-08-routing-heads-two-head.md` and `2026-08-08-slot-head-three-head.md`.
|
||||
New script `gen_clarify.py`, new fixture `eval_fixture_clarify.jsonl`.
|
||||
|
||||
## A softmax has no clarify class
|
||||
|
||||
That sentence closed the two-head measurement. It is why the head's fixture was
|
||||
88 cases and not 96. The eight `want_clarify` cases sat outside every number
|
||||
measured, and the head had no way to produce the answer they wanted.
|
||||
|
||||
A fourth head is the answer. Clarify is not a value of intent. It is a second
|
||||
question asked of the same pooled vector: can Maven act on this at all.
|
||||
|
||||
## The corpus had one class
|
||||
|
||||
Every row in `train_heads_slots.jsonl` was generated FOR an intent or a
|
||||
destination. So every row is answerable by construction. A head trained on that
|
||||
alone sees one class and learns to say yes.
|
||||
|
||||
`gen_clarify.py` makes the other class. Five shapes, ten topics. The shapes are
|
||||
the gate's own reasons in `gateLLMDecision` plus the two the fixture carries:
|
||||
bare noun, bare verb, demonstrative, deictic time, dangling reference.
|
||||
|
||||
**The agreement filter that worked for destination cannot work here.**
|
||||
`routeGrammar` has no clarify value. So the router always names an intent, and
|
||||
any generated line always agrees with itself. The second pass is a judge
|
||||
instead. Gemma is asked, without seeing the label, whether Maven would have to
|
||||
ask a question back.
|
||||
|
||||
## The first judge was worthless and the second was measured
|
||||
|
||||
The first judge said "needs clarify" on 24 of 40 plainly answerable corpus
|
||||
rows. It flagged `запиши что я пообедал` and `Покажи расписание поездов на
|
||||
вечер`. It was judging against a generic assistant, one that asks "where?"
|
||||
about lunch. Maven writes that note.
|
||||
|
||||
Rewriting it to state what she can already do took false positives to 16 of 60.
|
||||
It also catches all eight fixture clarifies. So the judge discriminates.
|
||||
|
||||
On the generated pile it removed 7 of 306, a 97.7% keep rate. That is not the
|
||||
judge failing. The generator is aimed at underspecified lines, so there is
|
||||
little for a filter to catch. The 27% false-positive rate is the number to
|
||||
quote, and it is label noise on the positive class.
|
||||
|
||||
**Both passes are gemma-4-12b.** Generation and judging. So the corpus is
|
||||
gemma's opinion of what is underspecified, and the head distills that opinion.
|
||||
What keeps it honest is the fixture. Those eight cases were written by the owner
|
||||
and gemma never saw them.
|
||||
|
||||
299 rows kept, against 3604 answerable. The positive class carries `intent:
|
||||
null`, so it costs the intent head nothing.
|
||||
|
||||
## Result
|
||||
|
||||
Three seeds, 24 epochs, epoch still chosen on the intent dev slice.
|
||||
|
||||
| | two heads | three heads | four heads |
|
||||
|---|---|---|---|
|
||||
| intent mean | 93.6% | 92.8% | 91.7% |
|
||||
| destination mean | 80.8% | 82.8% | 79.8% |
|
||||
| slot span F1 mean | — | 72.4% | 68.3% |
|
||||
| clarify caught | — | — | 7.0 of 8 |
|
||||
| false clarifies | — | — | 2.3 of 88 |
|
||||
|
||||
**The fourth head is not free the way the third was.** Intent, destination and
|
||||
slot F1 all move down. The drop is one to four points, and the seed spread is
|
||||
wide enough to contain it. Seed 2 scores intent 94.3% and destination 84.8%, both above
|
||||
every three-head seed. Read the drop as unproven rather than as absent.
|
||||
|
||||
Accuracy is the wrong number for this head and is reported for completeness at
|
||||
95.8% to 96.9%. Eight of ninety-six cases are positive, so a head that never
|
||||
asks scores 91.7%. Recall on those eight is the number.
|
||||
|
||||
Compare it to what ships. The cascade today misses 1 clarify and produces 2
|
||||
false ones. The head catches 7 of 8 and produces 2.3 false ones. That is
|
||||
parity, from a 118M encoder with no rules in front of it.
|
||||
|
||||
The saved checkpoint is seed 2 at epoch 10. Intent 94.3%, destination 28/33,
|
||||
slot F1 73.6%, clarify 7 of 8 with 3 false. `heads.pt` carries four state dicts.
|
||||
|
||||
## What it gets wrong is consistent across seeds
|
||||
|
||||
`поужинал` is a false clarify on all three seeds. That utterance is already
|
||||
recorded as a real defect. `thinSingleToken` was narrowed on 2026-08-01 to spare
|
||||
a token carrying a Russian verb ending. One word is routinely a whole sentence
|
||||
in Russian. The head relearned the mistake the rule was narrowed to
|
||||
fix.
|
||||
|
||||
`что дальше?` is a false clarify on two seeds. That one is a disagreement rather
|
||||
than an error. The utterance is underspecified, and V-498 decided stage 0 claims
|
||||
it for the calendar on purpose.
|
||||
|
||||
`ну это` is missed on two seeds. `amb-003` is the shortest case in the fixture
|
||||
and the generated demonstratives are longer.
|
||||
|
||||
## Confidence
|
||||
|
||||
`Confidence: 1.0` was a hardcode in `llmrouter.go`, so a correct low confidence
|
||||
could not exist. Max softmax over the intent head is the replacement. It is only worth reading
|
||||
if it is lower where the head is wrong.
|
||||
|
||||
It is. Mean 0.851 where the head is right against 0.604 where it is wrong. It
|
||||
ranks a right case above a wrong one in 83.4% of pairs.
|
||||
|
||||
So there are two signals now and they are not the same signal. Confidence says
|
||||
the head is unsure which intent this is. The clarify head says the utterance
|
||||
does not carry enough to act on. A confident wrong route and an honest "I cannot
|
||||
tell" are different failures, and one number cannot report both.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
The same gap as every head run. **Nothing of this runs in Go.** Four heads
|
||||
instead of three does not change that.
|
||||
|
||||
There is no threshold. Both signals are reported as raw numbers. Turning either
|
||||
into a gate needs a decision about where to cut, and that trades false clarifies
|
||||
against wrong acts. The fixture has 8 positives, which is too few to fit a
|
||||
threshold on.
|
||||
|
||||
The 299 generated rows have no held-out slice of their own. Clarify is scored on
|
||||
the fixture alone.
|
||||
@@ -0,0 +1,84 @@
|
||||
# The first destination number
|
||||
|
||||
Measured 2026-08-08 on the classifier cascade with the ONNX multilingual
|
||||
embedder, the configuration homesrv runs. `make t PKG=./internal/router/eval/
|
||||
RUN=TestONNXBaseline V=1`. Covers V-659, the follow-up V-655 named.
|
||||
|
||||
## What was measured
|
||||
|
||||
V-655 split a routing decision in two. The cascade sorts an utterance into one
|
||||
of seven intents, and `Decision.Source` then says where the answer lives. The
|
||||
first half had a fixture. The second half arrived with none, so twelve
|
||||
destinations shipped with no accuracy number.
|
||||
|
||||
`want_source` is now a field on `eval.Case`. It is a pointer, because the
|
||||
destination has three states and a bare string has two. Absent is every intent
|
||||
but query, which never reaches `queryWalk`. Present and empty is the
|
||||
`SourceUnknown` contract: name nothing and let the daemon walk the chain.
|
||||
Present and named is a destination the route must produce.
|
||||
|
||||
Thirty-three of the ninety-six cases carry one. A destination miss does not
|
||||
fail the case, so `Accuracy` and `IntentAccuracy` mean what they meant.
|
||||
`SourceAccuracy` is a second number over the labelled cases only.
|
||||
|
||||
## Result
|
||||
|
||||
Intent is **73/96 (76.0%)**, against 69/91 (75.8%) before. Four of the five new
|
||||
cases pass and no existing case moved.
|
||||
|
||||
Destination is **12/33 (36.4%)**, and the split is the whole finding.
|
||||
|
||||
| destination | scored | note |
|
||||
|---|---|---|
|
||||
| world | 5/5 | `WorldQueryGrammars` names it at stage 0 |
|
||||
| the `SourceUnknown` floor | 5/7 | the two misses lost the intent first |
|
||||
| calendar | 2/6 | `calendar-query` names it, the possessive agenda rules do not |
|
||||
| recall | 0/15 | nothing anywhere names it |
|
||||
|
||||
Recall is the number to move. Fifteen cases ask about his own words and his own
|
||||
facts. The route lands `query` on eleven of them and the destination comes back
|
||||
empty every time. Those turns are answered today, because the daemon walks the
|
||||
chain in order and the three recall passes are early in it. What is missing is a
|
||||
decider that says so, and that is the fourth head on V-546.
|
||||
|
||||
Two cases labelled the floor lost their intent before a destination was
|
||||
possible. A clarify names nothing, so it would satisfy an empty label for free.
|
||||
`Score` requires the route to land the case's intent before it credits a
|
||||
destination hit, or the floor label would score itself.
|
||||
|
||||
## Seven cases assert the floor, and five of them cluster
|
||||
|
||||
The five are homelab operations. `SourceRecall`, `SourceNetwork` and
|
||||
`SourceAttention` overlap on every question about the box, because `mavpoll`
|
||||
writes its netdata and uptime-kuma observations into the fact store recall
|
||||
reads. "почему сервер тормозит" is answerable from all three. Naming one takes
|
||||
the other two off the turn.
|
||||
|
||||
That is a finding about the enum rather than a gap in the labelling. The floor
|
||||
is the right answer there and the fixture now says so out loud.
|
||||
|
||||
All seven were written by an agent and confirmed by the owner on 08-08-2026.
|
||||
|
||||
## A drift the labelling found
|
||||
|
||||
`WorldQueryGrammars` went into `buildRouter` with V-655 and never into
|
||||
`baselineGrammars`, the fixture's mirror of it. So the fixture was scoring a
|
||||
grammar set the daemon does not run. The comment above that function forbids
|
||||
exactly that. Adding it moved the destination number from 9/33 to 12/33 and
|
||||
moved nothing else.
|
||||
|
||||
The three cases it recovered are `что такое TCP?`, `сколько будет 17 на 23?`
|
||||
and `кто такой Линус Торвальдс?`. All three already routed `query` through
|
||||
`NarrativeQueryGrammars`. So the drift was invisible to every number this
|
||||
fixture reported, until the destination had one of its own.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
The model arm. This is the classifier cascade, which names a destination only
|
||||
where a stage 0 rule filled one in. The resident model has no destination in
|
||||
its router prompt yet, so 36.4% is a floor and not a comparison.
|
||||
|
||||
Two pairs of cases are the same utterance. `ru-query-020` and `ru-query-024`
|
||||
are both "что дальше?", and `ru-query-021` and `ru-query-025` are both
|
||||
"расскажи про битву при Ватерлоо". They differ in tags and note only, so both
|
||||
pairs are counted twice here and in every earlier number this fixture reported.
|
||||
@@ -0,0 +1,66 @@
|
||||
# The destination, with a model that can name one
|
||||
|
||||
Measured 2026-08-08 against gemma-4-12b on the workstation, the same 96-case
|
||||
fixture V-659 built. Covers V-660.
|
||||
|
||||
```sh
|
||||
no_proxy='*' MAVEN_LLM_URL=http://192.168.1.105:8080 \
|
||||
make t PKG=./internal/router/eval/ RUN=TestLLMRouterBaseline V=1
|
||||
```
|
||||
|
||||
## The gap was structural
|
||||
|
||||
V-659 measured the destination at 12/33 on the classifier cascade, with recall
|
||||
at 0/15. Nothing in `routeSystem` named a `Source` and `routeGrammar` could not
|
||||
emit one, so the resident model had no string to write. That is the shape V-517
|
||||
measured for Praxis reach at 0/12: not a weak model, an absent contract.
|
||||
|
||||
`routeGrammar` now carries a `source` rule closed over `router.Sources` plus the
|
||||
empty floor. The prompt lists the twelve destinations in Russian and says that
|
||||
`""` is a normal answer to give often.
|
||||
|
||||
## Result
|
||||
|
||||
| run | intent | destination |
|
||||
|---|---|---|
|
||||
| classifier + ONNX (V-659) | 73/96 (76.0%) | 12/33 (36.4%) |
|
||||
| gemma-4-12b alone | 79/96 intent-only (82.3%) | 26/33 (78.8%) |
|
||||
| cascade + gemma-4-12b + hash fallback | 81/96 (84.4%) | 24/33 (72.7%) |
|
||||
|
||||
Recall is the move: 0/15 to 14/15. Intent did not shift, which was the
|
||||
constraint. The prompt is shared, so a destination rule that costs routing
|
||||
points is not a win.
|
||||
|
||||
The eight llm-only errors are the eight `want_clarify` cases. The model returned
|
||||
`unknown` on every one, which is correct, and the llm-only harness surfaces a
|
||||
decline as an error by design.
|
||||
|
||||
## Stage 0 now costs four destination points
|
||||
|
||||
The four cases the cascade loses and the model alone wins are all calendar. The
|
||||
possessive agenda rules claim them at stage 0 and deliberately name nothing.
|
||||
"что у меня в списке покупок" matches the same rule. Naming the calendar there
|
||||
would take the list source off the turn (V-655).
|
||||
|
||||
So a rule written to be careful about the list now blocks a model that would
|
||||
have named the calendar correctly. Before V-660 that caution was free, because
|
||||
nothing downstream of stage 0 could name anything either.
|
||||
|
||||
Three ways out, and each costs something. Split the possessive rule so the
|
||||
calendar-shaped half names its destination. Let a later stage overwrite an empty
|
||||
destination a grammar left behind, which reverses "a matched value always wins".
|
||||
Or leave it, on the argument that four points is cheap next to a wrong
|
||||
destination on a shopping list. This wants the owner's call rather than a quiet
|
||||
edit.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
The resident Qwen3-1.7B, which is what homesrv runs. It binds `--port 0` inside
|
||||
the container and no host process can reach it. Scoring it needs a second
|
||||
llama-server on a fixed port. The workstation is never assumed
|
||||
up, so the homesrv number is the one that decides whether this ships on by
|
||||
default.
|
||||
|
||||
The fixture is 33 labelled destinations over twelve values. Recall carries 15 of
|
||||
them and five destinations carry none at all. A per-destination number below
|
||||
world, recall, calendar and the floor is not supported by this fixture.
|
||||
@@ -0,0 +1,142 @@
|
||||
# MASSIVE Russian warm-start for the routing heads
|
||||
|
||||
Measured 2026-08-08 on workpc (Radeon RX 7900 GRE, ROCm). Covers V-546 step 2.
|
||||
Workspace is `~/Programs/embed-training` on workpc, scripts `train_massive.py`,
|
||||
`ab_run.py`, `ab.sh`, `probe_time.py`.
|
||||
|
||||
## What was trained
|
||||
|
||||
Two heads on a copy of multilingual-e5-small: `Linear(384, 60)` for MASSIVE's
|
||||
own intents over a masked mean pool, `Linear(384, 111)` per token for BIO slot
|
||||
tags. MASSIVE's label sets verbatim, no alignment to Maven's 7 intents. The
|
||||
intent head is an auxiliary loss that shapes the pooled vector and is thrown
|
||||
away.
|
||||
|
||||
Data is `amazon-massive-dataset-1.1` pulled from S3. The Hugging Face repo is
|
||||
script-only and `datasets` 5.0 refuses those, so `load_dataset` cannot fetch it.
|
||||
`ru-RU` is 11,514 train, 2,033 dev, 2,974 test, 60 intents, 55 slots, 111 BIO
|
||||
labels. All 16,521 rows survived span alignment: `annot_utt` re-tokenised to its
|
||||
own `utt` on every one.
|
||||
|
||||
Hyperparameters match `train_intent.py`, so the two runs differ in data only.
|
||||
Frozen XLM-R vocabulary, body 2e-5, heads 1e-3, batch 32, sequence 64, 10
|
||||
epochs. MASSIVE's own dev partition selects the epoch, on slot F1 with intent
|
||||
accuracy as tiebreak. Selecting on 60-class intent accuracy would optimise a
|
||||
head that gets deleted.
|
||||
|
||||
## Result
|
||||
|
||||
Epoch 9 of 10 by dev slot F1. Held-out MASSIVE test: intent 86.2%, slot span
|
||||
F1 71.5% (P 68.5, R 74.8). Peak 1.70GB of 17.2GB, about 22 seconds an epoch,
|
||||
under 4 minutes end to end. Dev slot F1 climbed monotonically to epoch 9 and
|
||||
fell at 10, so 10 epochs was the right budget.
|
||||
|
||||
Ten slot types sit at 0% test recall. Every one of them has 1 to 7 test
|
||||
instances: `alarm_type` has 3, `drink_type` has 1. That is support in MASSIVE's
|
||||
Russian split, not a tagger failure. `playlist_name` at 6% of 16 is the first
|
||||
real miss.
|
||||
|
||||
## The intent A/B, and why it settles nothing
|
||||
|
||||
`train_intent.py` was run against both bodies, three seeds by two smoothing
|
||||
settings, on `train_v4.jsonl`. It is v4 and not v5 because v4 is what
|
||||
`sweep2.log` measured. `ab_run.py` strips a `--base` flag onto the module global, so
|
||||
`train_intent.py` is unmodified and its baseline stays reproducible. The stock
|
||||
arm reproduced `sweep2.log` line for line.
|
||||
|
||||
Fixture accuracy, 91 cases, one case is 1.1 points:
|
||||
|
||||
| seed / smooth | stock | warm-started |
|
||||
|---|---|---|
|
||||
| 0 / 0.0 | 94.0% | 92.8% |
|
||||
| 0 / 0.1 | 95.2% | 92.8% |
|
||||
| 1 / 0.0 | 95.2% | 94.0% |
|
||||
| 1 / 0.1 | 95.2% | 97.6% |
|
||||
| 2 / 0.0 | 92.8% | 94.0% |
|
||||
| 2 / 0.1 | 92.8% | 96.4% |
|
||||
|
||||
Mean 94.2% against 94.6%. That is +0.4 points, about a third of one case, and
|
||||
inside seed noise. Spread widened. Stock lands in a 2.4-point band and
|
||||
warm-started in a 4.8-point one. The warm-started arm holds both the best result
|
||||
of the sweep and a tie for the worst. Seed 0 is the bad arm and it fails in a
|
||||
specific way. Its dev peaks at epoch 2 and 3 and never improves, where stock
|
||||
peaks around 7. The dev slice is a quarter of the seed rows. That is small
|
||||
enough that early stopping is fragile when the body arrives already fitted.
|
||||
|
||||
**The A/B was never the test.** Intent had at most 4.8 points of headroom here.
|
||||
MASSIVE was not trained for Maven's intents. Read it as "the warm-start does not
|
||||
cost intent accuracy", nothing more.
|
||||
|
||||
## The measurement that does mean something
|
||||
|
||||
`want_time` is the one slot Maven's fixture scores, and MASSIVE has `time` and
|
||||
`date`. Restricted to those two slot types, F1 is 74.9% over 609 gold spans on the
|
||||
MASSIVE ru test split. Precision is 71.5 and recall 78.7. That beats the 71.5%
|
||||
all-slot figure. Of the 530 test utterances carrying a time or a date, 73.4% get
|
||||
every such span exactly right.
|
||||
|
||||
Out of domain matters more, because Maven's traffic is not this corpus. Ten
|
||||
Maven-shaped utterances, none of them in MASSIVE:
|
||||
|
||||
| utterance | tagged |
|
||||
|---|---|
|
||||
| `напомни в 11:00 позвонить маме` | `time='11:00'`, `relation='маме'` |
|
||||
| `напомни завтра в семь утра выпить таблетки` | `date='завтра'`, `time='семь утра'` |
|
||||
| `поставь будильник на полседьмого` | `time='полседьмого'` |
|
||||
| `через двадцать минут напомни про чайник` | `time='двадцать минут'` |
|
||||
| `напомни в пятницу вечером забрать посылку` | `date='пятницу'`, `timeofday='вечером'` |
|
||||
| `что у меня сегодня после обеда` | `date='сегодня'`, `time='после'`, `timeofday='обеда'` |
|
||||
| `запиши что кофе закончился` | nothing |
|
||||
| `что такое TCP` | `definition_word='TCP'` |
|
||||
|
||||
The first row is the V-572 defect utterance. `ReminderGrammar` handed the daemon
|
||||
`HasTime: false` there, and the daemon asked "Когда?" at a sentence that had
|
||||
already said when. `полседьмого` is a colloquial half-past that no digit pattern
|
||||
catches. `запиши что кофе закончился` correctly carries nothing, because a note
|
||||
has no time.
|
||||
|
||||
Two errors. `после обеда` split into `time='после'` plus `timeofday='обеда'`
|
||||
when it is one span, and `через двадцать минут` dropped its `через`. Both are
|
||||
boundary errors on spans the tagger did find.
|
||||
|
||||
Unplanned: `что такое TCP` returned `definition_word='TCP'`. MASSIVE has a slot
|
||||
for the thing being asked about, which is a `SourceWorld` signal sitting in a
|
||||
head already trained.
|
||||
|
||||
Ten hand-picked utterances are evidence, not a fixture.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
Maven has no span fixture. `want_time` and `want_fn` are presence booleans and
|
||||
`want_fact_key` is an exact string match, so nothing in the repo can score a
|
||||
71.5% span tagger. Destination got one the same day, at 12/33 on the classifier
|
||||
cascade: see `2026-08-08-destination-fixture.md`.
|
||||
|
||||
The missing span fixture is why the warm-start stays unjudged against Maven
|
||||
rather than against MASSIVE.
|
||||
|
||||
## Datasets ruled out
|
||||
|
||||
Checked on 2026-08-08 and rejected as label sources:
|
||||
|
||||
- **MASSIVE's other 50 locales** ship in the same tarball and are parallel by id.
|
||||
Co-training on them is free and unmeasured. English was ruled out by the owner
|
||||
on 2026-08-08.
|
||||
- **CLINC150** is reachable as parquet, 150 intents and 1,200 explicit
|
||||
out-of-scope queries, English only. Its value is the labeled out-of-scope set
|
||||
for fitting the energy threshold, not intent labels.
|
||||
- **`d0rj/dolphin-ru`**, roughly 2.8M rows of FLAN-style tasks translated to
|
||||
Russian. No intent, no slots, and not utterances anyone says to an assistant.
|
||||
- **`psytechlab/EmpatheticIntents-ru`**, 24,856 rows of translated
|
||||
EmpatheticDialogues with 32 emotion labels. Maven's mood enum is `neutral,
|
||||
happy, thinking, tired, confused` and it describes her own reply, not the
|
||||
speaker's emotion. No mapping exists.
|
||||
- **`ai-forever/MERA`** and **`RussianNLP/russian_super_glue`**, benchmark
|
||||
harnesses. Rows are prompt templates with `{toxic_comment}` placeholders.
|
||||
- **`ZeroAgency/ru-big-russian-dataset`**, an LLM-judge quality corpus. Its
|
||||
`question` and `classified_topic` columns are a usable Russian out-of-scope
|
||||
pool for threshold fitting. That is the one thing CLINC150 can only supply in
|
||||
English. The questions are long and written, so they belong in the negative
|
||||
set, never in the in-scope `query` training set.
|
||||
- No second Russian slot-filling corpus exists. The xSID mirrors are 404,
|
||||
MultiATIS++ has no Russian, SLURP is not on the Hub.
|
||||
@@ -0,0 +1,192 @@
|
||||
# Two heads on e5-small, and the first destination the router did not need a model for
|
||||
|
||||
Measured 2026-08-08 on workpc (Radeon RX 7900 GRE, ROCm). Covers V-661, step 3 of
|
||||
`docs/plans/18-routing-heads-on-e5-small.md`. Workspace is `~/Programs/embed-training`,
|
||||
scripts `gen_query_source.py`, `label_source.py`, `build_heads_corpus.py`,
|
||||
`train_heads.py`, `score_confidence.py`.
|
||||
|
||||
## Two heads, not four
|
||||
|
||||
Intent is 7 classes and destination is 12 plus the `SourceUnknown` floor, sharing
|
||||
one masked mean pool over one forward pass. The plan asked for four. Two of them
|
||||
have no labels and neither is a GPU problem.
|
||||
|
||||
**Mood is cut, not deferred.** Maven's enum is `neutral, happy, thinking, tired,
|
||||
confused` and it describes her own reply state, not the speaker's emotion.
|
||||
`psytechlab/EmpatheticIntents-ru` was the only candidate and its 32 emotion
|
||||
labels do not map onto it. There is nothing to train against.
|
||||
|
||||
**BIO slot tags stay in the MASSIVE body.** No Maven-domain span corpus exists.
|
||||
`2026-08-08-massive-warm-start.md` records that no second Russian slot-filling
|
||||
corpus is reachable at all.
|
||||
|
||||
The destination loss is masked with `ignore_index`. Only a query turn reaches
|
||||
`queryWalk`, so a reminder contributes nothing to it.
|
||||
|
||||
## Where the destination labels came from
|
||||
|
||||
V-660 taught the router prompt to name a destination. That made gemma-4-12b a
|
||||
teacher, and this distils it.
|
||||
|
||||
Labelling the 300 query rows already in `train_v5.jsonl` gave 277 destinations.
|
||||
The shape was unusable: the floor 101, calendar 62, recall 47, and `feeds` and
|
||||
`attention` at zero. A 13-way softmax cannot learn a class with no examples.
|
||||
|
||||
`gen_query_source.py` is the destination half of `gen_corpus.py` and runs the
|
||||
same two passes. Gemma writes questions whose answer lives in one named place.
|
||||
The daemon's own `routeSystem` prompt then routes each one back. A line survives
|
||||
only when the intent is `query` **and** the source is the destination it was
|
||||
generated for. The glosses are copied verbatim out of `route_system.txt`, so the
|
||||
generator and the labeller work from one definition.
|
||||
|
||||
The prompt and the GBNF are dumped from `internal/router/llmrouter.go` by
|
||||
`TestDumpPrompt`, never retyped. The workspace held its own copies and V-660
|
||||
changed both.
|
||||
|
||||
1229 kept of 2373 generated, 51.8%. Merged corpus is 3664 rows carrying 1727
|
||||
destinations:
|
||||
|
||||
| | rows | | rows |
|
||||
|---|---|---|---|
|
||||
| the floor | 220 | world | 132 |
|
||||
| calendar | 180 | money | 126 |
|
||||
| tasks | 169 | list | 124 |
|
||||
| recall | 167 | self, feeds, attention | 120 each |
|
||||
| weather | 136 | network | 103 |
|
||||
| | | home | 50 |
|
||||
|
||||
`home` is thin because the agreement filter rejected most of what was generated
|
||||
for it. A question about the house routes `act` more often than `query`. That is
|
||||
the filter working, and 50 is the finding rather than a shortfall.
|
||||
|
||||
**The floor was regenerated once.** The first 120 rows carried one sentence
|
||||
shape across eight topics. That shape was "что там с X" and its two synonyms.
|
||||
Every named destination varied and only the floor collapsed. The reason is that
|
||||
the generator varies a topic, and ambiguity is not a topic.
|
||||
|
||||
`gen_query_source.py` now rotates six floor shapes. A `почему` question, a yes
|
||||
or no question, and a question carried by intonation alone. Then a
|
||||
better-or-worse question, a status question, and an existence question. That is
|
||||
a fix to degenerate generation. It is not fitting to the fixture, whose floor
|
||||
cases are homelab operations and match none of the six.
|
||||
|
||||
The 1229 generated rows carry `intent: null`. Every one is a query by
|
||||
construction. There are five times as many as the corpus has query rows, so
|
||||
including them would make query half the intent corpus.
|
||||
|
||||
## Result
|
||||
|
||||
Three seeds, two bodies, epoch chosen on the intent dev slice and never on a
|
||||
destination number.
|
||||
|
||||
| body | floor corpus | intent mean | destination mean |
|
||||
|---|---|---|---|
|
||||
| warm-started `out/body_massive` | one shape | 93.6% | 75.8% |
|
||||
| stock `multilingual-e5-small` | one shape | 93.6% | 76.8% |
|
||||
| warm-started `out/body_massive` | six shapes | 93.6% | **80.8%** |
|
||||
|
||||
Best single run is destination **29/33 (87.9%)**, seed 0 on the rotated floor.
|
||||
Peak 1.68GB of 17.2GB, under four minutes end to end.
|
||||
|
||||
Against the two arms already measured on the same 33 labelled cases:
|
||||
|
||||
| | destination |
|
||||
|---|---|
|
||||
| classifier cascade (V-659) | 12/33 (36.4%) |
|
||||
| cascade + gemma-4-12b (V-660) | 24/33 (72.7%) |
|
||||
| two heads on e5-small | 29/33 (87.9%) |
|
||||
|
||||
A 118M encoder beats the 12B teacher it was distilled from, on the fixture. The
|
||||
per-destination split is where it happens: **recall 15/15** and **world 5/5**.
|
||||
Recall was 0/15 on the cascade and 14/15 through gemma.
|
||||
|
||||
Read 87.9% as one seed of a mean of 80.8%, not as a headline. Three seeds score
|
||||
29, 25 and 26 of 33. One case is 3 points on a fixture this small.
|
||||
|
||||
Intent is **not** comparable to the 73/96 and 81/96 figures those two arms
|
||||
scored. A softmax has no clarify class. The head's fixture is the 88 cases that
|
||||
carry an intent, and the 8 `want_clarify` cases are scored separately below.
|
||||
|
||||
## The MASSIVE warm-start is worth nothing here either
|
||||
|
||||
Step 2 measured it at +0.4 points of intent accuracy and called that inside seed
|
||||
noise. Destination was the open question, because MASSIVE has a
|
||||
`definition_word` slot that looked like a `SourceWorld` signal sitting in a head
|
||||
already trained.
|
||||
|
||||
It is not. The two bodies score the same intent mean to one decimal. Stock is
|
||||
one point ahead on destination, which is a third of one case. Nothing here argues
|
||||
for keeping the warm-start step. Dropping it removes a dependency on a corpus
|
||||
pull that `datasets` 5.0 cannot do.
|
||||
|
||||
## The floor moved, calendar did not
|
||||
|
||||
Before the rotation, all seven misses at seed 0 were the floor and calendar. The
|
||||
head named a destination where the fixture says walk the chain, and it was
|
||||
confident doing it. `"почему сервер тормозит"` read `world` at 0.80.
|
||||
`"хватает ли места под новые бэкапы"` read `network` at 0.82. Those are the five
|
||||
homelab cases V-659 flagged, where `SourceRecall`, `SourceNetwork` and
|
||||
`SourceAttention` all overlap. `mavpoll` writes its observations into the fact
|
||||
store recall reads.
|
||||
|
||||
| | one shape | six shapes |
|
||||
|---|---|---|
|
||||
| the floor | 3/7 | 6/7, 6/7, 5/7 |
|
||||
| calendar | 3/6 | 3/6, 3/6, 3/6 |
|
||||
| recall | 15/15 | 15/15 at seed 0 |
|
||||
| world | 5/5 | 5/5 |
|
||||
|
||||
The floor was a corpus defect and it cost 3 cases. Sentence variety carried it,
|
||||
not homelab vocabulary, which the training rows still do not contain.
|
||||
|
||||
**Calendar is 3/6 at every seed and is a different problem.** It is the shape
|
||||
V-660 named. The possessive agenda rules claim those cases at stage 0 and
|
||||
deliberately name nothing, so no destination label reaches the head. Training
|
||||
cannot move a case the head never sees. That one wants the owner's call.
|
||||
|
||||
## Max softmax separates, weakly, and the gate stays
|
||||
|
||||
The plan argues max softmax is a calibratable confidence where `Confidence: 1.0`
|
||||
was a hardcode. Measured on the intent head:
|
||||
|
||||
| | n | mean confidence |
|
||||
|---|---|---|
|
||||
| correct | 80 | 0.897 |
|
||||
| wrong | 8 | 0.705 |
|
||||
| `want_clarify` | 8 | 0.685 |
|
||||
|
||||
The softest correct answer is 0.66 and 3 of 8 clarify cases sit below it. So a
|
||||
single cut buys three clarifies at no false-clarify cost, and no more.
|
||||
|
||||
The other five explain themselves. `"напомни"` scores 0.94 as `reminder` and
|
||||
`"сделай это"` scores 0.80 as `act`. Both are intent-certain and slot-empty, and
|
||||
confidence was never the signal there. `gateLLMDecision` already catches exactly
|
||||
that shape, an act with no allowlisted fn or a keyless fact, and it keeps doing
|
||||
so. The head replaces the hardcode. It does not replace the gate.
|
||||
|
||||
## An incident worth recording
|
||||
|
||||
The first generation run produced zero rows for eight destinations. `mavgpud`
|
||||
yields the card when another process wants it (V-488) and llama-server answers
|
||||
503 until the model is back. Every generate call inside that window burned one of
|
||||
the destination's batches. The run walked its own cap without a single successful
|
||||
call. The log said `503` 260 times, and the summary line said 64.1% keep rate,
|
||||
which read as success.
|
||||
|
||||
`call()` now retries a 503 with backoff. A generator that treats an unloaded
|
||||
model as a bad generation is a silent-corpus bug, not a slow one.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
**Nothing here runs in Go.** The heads are a `heads.pt` and an
|
||||
`out/body_heads/` on workpc. Reaching the daemon needs an ONNX export and a
|
||||
caller. The resident e5-small must not be replaced by this copy: recall depends
|
||||
on that file, and `EmbedQuery`/`EmbedPassage` are its contract.
|
||||
|
||||
The destination fixture carries 4 of 13 classes: recall 15, floor 7, calendar 6,
|
||||
world 5. `tasks`, `money`, `list`, `home`, `network`, `feeds`, `attention` and
|
||||
`self` have no gold case. So 78.8% is silent on eight destinations that together
|
||||
hold 800 training rows.
|
||||
|
||||
Latency was not measured. A forward pass of a 118M encoder should beat a 1.7B
|
||||
decoder on a query turn. That is arithmetic, not a number from this box.
|
||||
@@ -0,0 +1,106 @@
|
||||
# A slot head, and the corpus that did not exist this morning
|
||||
|
||||
Measured 2026-08-08 on workpc, the same day as `2026-08-08-routing-heads-two-head.md`
|
||||
and against the same fixtures. Workspace is `~/Programs/embed-training`, new
|
||||
scripts `slot_grammar.gbnf`, `slot_system.txt`, `label_slots.py`, `merge_slots.py`.
|
||||
|
||||
## The corpus was the whole problem
|
||||
|
||||
The two-head measurement said BIO slot tags stay in the MASSIVE body, because no
|
||||
Maven-domain span corpus exists. That was true of found corpora and false of
|
||||
made ones. Destination had the same shape at breakfast. V-660 gave gemma-4-12b a
|
||||
string to write and the label problem became a generation problem.
|
||||
|
||||
The same trick applies to spans. `slot_grammar.gbnf` emits a list of
|
||||
`{"slot": ..., "text": ...}` and the enum closes over Maven's own five: `time`,
|
||||
`text`, `key`, `value`, `fn`. `slot_system.txt` demands each span be an exact
|
||||
substring of the utterance.
|
||||
|
||||
**The agreement filter is free here.** Destination needed a second pass. The
|
||||
daemon's own router prompt had to route each generated line back. A span needs
|
||||
no second call. It either occurs in the utterance or it does not, and
|
||||
`label_slots.py` drops it with `find()`.
|
||||
|
||||
1702 rows labelled from `train_v5.jsonl`, the reminder, fact, note, act and
|
||||
query intents. Chat and system carry no slot and were never asked.
|
||||
|
||||
| slot | spans |
|
||||
|---|---|
|
||||
| text | 1175 |
|
||||
| time | 485 |
|
||||
| fn | 381 |
|
||||
| key | 72 |
|
||||
| value | 65 |
|
||||
|
||||
37 spans dropped as not-a-substring, 2.2% of the pile. Nothing failed to parse,
|
||||
which is the grammar doing its job. 409 rows came back with no span at all.
|
||||
Those are kept and tagged all `O`. An utterance carrying no slot teaches the
|
||||
head not to invent one. An empty list is a label and not a miss.
|
||||
|
||||
`key` and `value` are thin because they come from facts alone. That is the
|
||||
shape of the corpus, not a labeller failure.
|
||||
|
||||
## Three heads on one forward pass
|
||||
|
||||
Intent and destination were already two linear heads over one masked mean pool.
|
||||
Slots is a third head over the per-token states of the same pass, so the marginal
|
||||
cost is one `Linear(384, 11)`.
|
||||
|
||||
The tag set is `O` plus `B-` and `I-` for each of the five. A softmax cannot
|
||||
emit a tag that does not exist. That is the structural guarantee the GBNF buys
|
||||
for the teacher, and the head gets it for free.
|
||||
|
||||
Two masking rules, both `ignore_index`. A row with no `spans` key contributes
|
||||
nothing, which covers the 1900 generated destination rows and every chat and
|
||||
system turn. A padding or special-token position contributes nothing either.
|
||||
|
||||
Scoring is exact-match span F1, not token accuracy. Most tokens are `O`, so a
|
||||
tagger that predicts nothing anywhere scores above 90% on tokens.
|
||||
|
||||
## Result
|
||||
|
||||
Three seeds, 24 epochs, epoch chosen on the intent dev slice alone.
|
||||
|
||||
| | two heads | three heads |
|
||||
|---|---|---|
|
||||
| intent mean | 93.6% | 92.8% |
|
||||
| destination mean | 80.8% | 82.8% |
|
||||
| destination best | 29/33 (87.9%) | 29/33 (87.9%) |
|
||||
| slot span F1 mean | — | 72.4% |
|
||||
|
||||
**The slot head costs nothing and adds a third decision.** Intent moves 0.8
|
||||
points down and destination 2 points up. Both sit inside the seed spread those
|
||||
two numbers already had. Read this as unchanged, not as a trade.
|
||||
|
||||
The saved checkpoint is seed 1 at epoch 10: intent 93.2%, destination 81.8%,
|
||||
slot F1 75.8%. `heads.pt` now carries three state dicts and the `bio` list
|
||||
beside the intent and source enums.
|
||||
|
||||
## Epoch selection is now wrong for one of the three heads
|
||||
|
||||
Slot F1 was still climbing when the intent-selected epoch stopped it. Seed 0
|
||||
selects epoch 13 at 70.9% and reaches 76.1% at epoch 20. Seed 1 selects epoch 10
|
||||
at 75.8% and reaches 80.0% at epoch 24.
|
||||
|
||||
So the three tasks want different epochs and the harness picks one. Two ways
|
||||
out, and neither was taken here. Select on a joint score, which needs an
|
||||
argument about weights. Or give the slot head its own dev slice and its own
|
||||
early stop, which means the heads stop being one checkpoint.
|
||||
|
||||
Leaving it costs about 4 points of slot F1 and nothing else. The router
|
||||
contract is graded on intent. Selecting on it is the conservative choice rather
|
||||
than an oversight.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
The same gap as the two-head run. **Nothing of this runs in Go.** Three heads
|
||||
instead of two does not change that. It does change the export. The slot head
|
||||
reads the per-token states, so an ONNX graph exposing only the pooled vector is
|
||||
not enough.
|
||||
|
||||
There is no held-out span fixture. Span F1 is measured on the dev slice of the
|
||||
same distilled corpus, so it scores agreement with gemma rather than
|
||||
correctness. The V-572 defect utterance, `"напомни в 11:00 позвонить маме"`, is
|
||||
labelled correctly by the teacher, but one utterance is not a fixture.
|
||||
|
||||
`key` at 72 spans and `value` at 65 are too thin to report a per-slot number.
|
||||
@@ -0,0 +1,636 @@
|
||||
# Raw transcript, two weeks of usage
|
||||
|
||||
Companion to `2026-08-08-two-weeks.md`. 140 turns through `POST /api/chat`,
|
||||
driven by `scripts/usage-run.py` over `scripts/testdata/usage-turns.txt`.
|
||||
Days are simulated. The wall clock was 20:18-20:27 throughout, which is why
|
||||
every clock answer says evening.
|
||||
|
||||
## Day 1
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: Сейчас уже вечер, поэтому я не могу сказать тебе «доброе утро».
|
||||
[1.5s t=78]
|
||||
|
||||
YOU: какой сегодня день?
|
||||
MAVEN: сегодня суббота, 8 августа 2026 года
|
||||
[0.0s t=79]
|
||||
|
||||
YOU: сколько времени?
|
||||
MAVEN: сейчас 20 часов 18 минут
|
||||
[0.0s t=80]
|
||||
|
||||
YOU: запиши что я пью кофе без сахара
|
||||
MAVEN: Я записала твою заметку про кофе。
|
||||
[2.0s t=81]
|
||||
|
||||
YOU: мой любимый язык программирования go
|
||||
MAVEN: Я записала твою заметку про Go.
|
||||
[1.4s t=82]
|
||||
|
||||
YOU: напомни в 11:00 позвонить маме
|
||||
MAVEN: Сейчас 20:18. В какой день?
|
||||
[0.4s t=83]
|
||||
|
||||
YOU: что у меня сегодня?
|
||||
MAVEN: на 08.08.2026 ничего нет, на какое время поставить напоминание?
|
||||
[0.3s t=84]
|
||||
|
||||
YOU: что такое TCP?
|
||||
MAVEN: для какого города? На какое время поставить напоминание?
|
||||
[1.0s t=85]
|
||||
|
||||
YOU: сколько будет 17 на 23?
|
||||
MAVEN: для какого города? На какое время поставить напоминание?
|
||||
[0.9s t=86]
|
||||
|
||||
YOU: спасибо
|
||||
MAVEN: Сейчас 20:18. В какой день?
|
||||
[0.9s t=87]
|
||||
|
||||
```
|
||||
|
||||
## Day 2
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: Сейчас 20:18. В какой день?
|
||||
[0.9s t=88]
|
||||
|
||||
YOU: что нового?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи, на какое время поставить напоминание?
|
||||
[1.2s t=89]
|
||||
|
||||
YOU: какая погода?
|
||||
MAVEN: для какого города? На какое время поставить напоминание?
|
||||
[0.6s t=90]
|
||||
|
||||
YOU: запиши что пароль от вайфая лежит в ящике стола
|
||||
MAVEN: Прошлую просьбу отпускаю. Хорошо, сохранил.
|
||||
[33.7s t=91]
|
||||
|
||||
YOU: где лежит вайфай пароль?
|
||||
MAVEN: вот что я нашла: пароль от вайфая лежит в ящике стола
|
||||
[7.0s t=92]
|
||||
|
||||
YOU: добавь молоко в список покупок
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[1.4s t=93]
|
||||
|
||||
YOU: что у меня в списке покупок?
|
||||
MAVEN: в списке покупок пусто.
|
||||
[0.0s t=94]
|
||||
|
||||
YOU: кто такой Линус Торвальдс?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[1.6s t=95]
|
||||
|
||||
YOU: какой у меня любимый язык?
|
||||
MAVEN: вот что нового: [$] Changes in shadow-utils password-expiration features; The Software Stewardship Lab launches; LightDM lives: version 1.33.0 released
|
||||
[1.3s t=96]
|
||||
|
||||
YOU: сколько у меня задач?
|
||||
MAVEN: сначала: проверить интейк форму на боксе (завтра); купить хлеб; купить молоко; полить цветы; починить кран
|
||||
[0.0s t=97]
|
||||
|
||||
```
|
||||
|
||||
## Day 3
|
||||
|
||||
```
|
||||
YOU: как дела?
|
||||
MAVEN: рада, как дела?
|
||||
[4.1s t=98]
|
||||
|
||||
YOU: напомни завтра в 9 утра купить хлеб
|
||||
MAVEN: хорошо, напомню завтра в 09:00.
|
||||
[0.3s t=99]
|
||||
|
||||
YOU: что у меня завтра?
|
||||
MAVEN: на 09.08.2026 ничего нет.
|
||||
[0.0s t=100]
|
||||
|
||||
YOU: отмени напоминание про хлеб
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[1.6s t=101]
|
||||
|
||||
YOU: какие у меня напоминания?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[0.0s t=102]
|
||||
|
||||
YOU: сохрани мне адрес гостиницы в Сочи
|
||||
MAVEN: Хорошо, сохраню.
|
||||
[5.6s t=103]
|
||||
|
||||
YOU: что я сохранил про Сочи?
|
||||
MAVEN: Хорошо, сохраню.
|
||||
[2.7s t=104]
|
||||
|
||||
YOU: почему сервер тормозит?
|
||||
MAVEN: есть такое: Почему тормозит сервер сайта: причины медленной работы и способы решения — Почему сервер работает медленно: основные причины торможения, диагностика проблем, способы оптимизации. Практические советы по ускорению работы сервера и сайта.
|
||||
[3.1s t=105]
|
||||
|
||||
YOU: хватает ли места под новые бэкапы?
|
||||
MAVEN: нового вот что: [$] Changes in shadow-utils password-expiration features; The Software Stewardship Lab launches; LightDM lives: version 1.33.0 released
|
||||
[1.5s t=106]
|
||||
|
||||
YOU: выключи свет в спальне
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[9.0s t=107]
|
||||
|
||||
```
|
||||
|
||||
## Day 4
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: рада
|
||||
[5.0s t=108]
|
||||
|
||||
YOU: что я пропустил?
|
||||
MAVEN: я не могу ответить на вопрос о том, что ты пропустил. мне нужно больше информации.
|
||||
[5.7s t=109]
|
||||
|
||||
YOU: о чём мы вчера говорили?
|
||||
MAVEN: пока не умею отвечать на этот вопрос.
|
||||
[1.1s t=110]
|
||||
|
||||
YOU: запиши что я записался к врачу на четверг
|
||||
MAVEN: Хорошо, сохранил.
|
||||
[6.0s t=111]
|
||||
|
||||
YOU: когда я иду к врачу?
|
||||
MAVEN: пока не умею отвечать на этот вопрос.
|
||||
[1.1s t=112]
|
||||
|
||||
YOU: что такое ZFS?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[1.0s t=113]
|
||||
|
||||
YOU: столица Франции?
|
||||
MAVEN: пока не умею отвечать на этот вопрос.
|
||||
[1.2s t=114]
|
||||
|
||||
YOU: переведи слово ремонт на английский
|
||||
MAVEN: записала: переведи слово ремонт на английский
|
||||
[1.1s t=115]
|
||||
|
||||
YOU: сколько я потратил в этом месяце?
|
||||
MAVEN: я не отслеживаю траты — не подключено.
|
||||
[1.4s t=116]
|
||||
|
||||
YOU: спокойной ночи
|
||||
MAVEN: спокойной ночи
|
||||
[4.3s t=117]
|
||||
|
||||
```
|
||||
|
||||
## Day 5
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: Привет!
|
||||
[2.6s t=118]
|
||||
|
||||
YOU: какая погода в Москве?
|
||||
MAVEN: погода не настроена.
|
||||
[1.2s t=119]
|
||||
|
||||
YOU: что там с бэкапами?
|
||||
MAVEN: нашла вот это: Бэкап - что такое резервное копирование и как его сделать — Бэкап (backup) — это резервная копия важных данных, сохраненная в специальный файл. Ею пользуются как подстраховкой: если что-то сломается, с помощью бэкапа можно будет все вернуть. Простейший пример бэкапа — сохранение файла в облаке или о…
|
||||
[4.0s t=120]
|
||||
|
||||
YOU: покажи что требует внимания
|
||||
MAVEN: я заметила, что у тебя есть запрос о том, что требуется внимание. давай разберёмся вместе.
|
||||
[5.2s t=121]
|
||||
|
||||
YOU: отметь это как сделанное
|
||||
MAVEN: команды «отметь» нет в списке. Предложила её добавить — включи через клиент.
|
||||
[0.0s t=122]
|
||||
|
||||
YOU: запиши что я купил новые наушники
|
||||
MAVEN: Заметка сохранена.
|
||||
[3.7s t=123]
|
||||
|
||||
YOU: какие у меня заметки за неделю?
|
||||
MAVEN: week summary — подведи итоги недели
|
||||
[0.0s t=124]
|
||||
|
||||
YOU: расскажи про Kubernetes
|
||||
MAVEN: вот что я нашла: Что такое Kubernetes — Сервисы, поддержка и инструменты Kubernetes широко доступны. Название Kubernetes происходит от греческого, что означает рулевой или штурман. Google открыл исходный код Kubernetes в 2014 году.
|
||||
[1.0s t=125]
|
||||
|
||||
YOU: кто я?
|
||||
MAVEN: name — Ками
|
||||
[1.0s t=126]
|
||||
|
||||
YOU: пока
|
||||
MAVEN: Пока, спасибо за разговор.
|
||||
[4.1s t=127]
|
||||
|
||||
```
|
||||
|
||||
## Day 6
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: Добро пожаловать! Я Maven — твой домашний ассистент. Как я могу помочь сегодня?
|
||||
[10.1s t=128]
|
||||
|
||||
YOU: сколько времени?
|
||||
MAVEN: сейчас 20 часов 21 минута
|
||||
[0.0s t=129]
|
||||
|
||||
YOU: напомни в 18:30 позвонить в банк
|
||||
MAVEN: Сейчас 20:21. В какой день?
|
||||
[0.3s t=130]
|
||||
|
||||
YOU: поставь чайник
|
||||
MAVEN: Сейчас 20:21. В какой день?
|
||||
[5.3s t=131]
|
||||
|
||||
YOU: включи музыку
|
||||
MAVEN: Сейчас 20:21. В какой день?
|
||||
[1.7s t=132]
|
||||
|
||||
YOU: что у меня в календаре на пятницу?
|
||||
MAVEN: вот что я нашла: ты записался к врачу на четверг, лежит вайфай пароль в ящике стола, сохранил про Сочи, пью кофе без сахара, на какое время поставить напоминание?
|
||||
[5.5s t=133]
|
||||
|
||||
YOU: во сколько у меня встреча?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи, на какое время поставить напоминание?
|
||||
[0.3s t=134]
|
||||
|
||||
YOU: запиши что дедлайн по проекту в понедельник
|
||||
MAVEN: Прошлую просьбу отпускаю. Хорошо, записала.
|
||||
[6.3s t=135]
|
||||
|
||||
YOU: успею ли я до дедлайна?
|
||||
MAVEN: вот что я нашла: дедлайн по проекту указан как понедельник.
|
||||
[5.3s t=136]
|
||||
|
||||
YOU: спасибо
|
||||
MAVEN: спасиба
|
||||
[3.8s t=137]
|
||||
|
||||
```
|
||||
|
||||
## Day 7
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: Привет!
|
||||
[2.6s t=138]
|
||||
|
||||
YOU: как ты?
|
||||
MAVEN: рада
|
||||
[4.2s t=139]
|
||||
|
||||
YOU: расскажи анекдот
|
||||
MAVEN: Я могу рассказать интересный анекдот для тебя. Вот - в одном из фильмов был такой сценарий: главному герою дали уйти от команды и он остался один, а потом появился его друг, который сказал ему: 'Тебе не нужно было остаться'. Но это просто игра слов.
|
||||
[7.1s t=140]
|
||||
|
||||
YOU: что ты умеешь?
|
||||
MAVEN: Я могу помочь с задачами, напоминать о важных датах и событиях. Также я помогу записывать заметки или отвечать по ним.
|
||||
[4.1s t=141]
|
||||
|
||||
YOU: запиши что я начал бегать по утрам
|
||||
MAVEN: Заметка сохранена.
|
||||
[9.2s t=142]
|
||||
|
||||
YOU: я бегаю по утрам уже неделю
|
||||
MAVEN: записала: я бегаю по утрам уже неделю
|
||||
[1.9s t=143]
|
||||
|
||||
YOU: как часто я бегаю?
|
||||
MAVEN: вот что я нашла: ты бегал по утрам
|
||||
[5.3s t=144]
|
||||
|
||||
YOU: сколько стоит биткоин?
|
||||
MAVEN: есть такое: Цена Bitcoin сегодня | Данные о цене, рынке и графики BTC в ... — В настоящее время 1 Bitcoin стоит 65 тыс $, что на 0,09% выше за последние 24 ч. Отслеживайте цену BTC, рыночную капитализацию, исторический график и рыночную аналитику в реальном времени.
|
||||
[2.1s t=145]
|
||||
|
||||
YOU: какие новости?
|
||||
MAVEN: вот что нового: [$] Changes in shadow-utils password-expiration features; The Software Stewardship Lab launches; LightDM lives: version 1.33.0 released
|
||||
[0.8s t=146]
|
||||
|
||||
YOU: хорошего дня
|
||||
MAVEN: Хорошого дня!
|
||||
[4.1s t=147]
|
||||
|
||||
```
|
||||
|
||||
## Day 8
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: Добро пожаловать! Как могу помочь сегодня?
|
||||
[2.9s t=148]
|
||||
|
||||
YOU: что у меня сегодня?
|
||||
MAVEN: на 08.08.2026 ничего нет.
|
||||
[0.0s t=149]
|
||||
|
||||
YOU: напомни через час выпить воды
|
||||
MAVEN: хорошо, напомню сегодня в 21:22.
|
||||
[0.3s t=150]
|
||||
|
||||
YOU: я выпил воды
|
||||
MAVEN: записала: я выпил воды
|
||||
[1.2s t=151]
|
||||
|
||||
YOU: запиши что кот ест только сухой корм
|
||||
MAVEN: Хорошо, записала.
|
||||
[5.6s t=152]
|
||||
|
||||
YOU: чем питается кот?
|
||||
MAVEN: вот что я нашла: кот питается только сухим кормом.
|
||||
[4.5s t=153]
|
||||
|
||||
YOU: что такое DNS?
|
||||
MAVEN: нашла вот это: Dns — Википедия — В случае рекурсивного запроса DNS-сервер опрашивает серверы (в порядке убывания уровня зон в имени), пока не найдёт ответ или не обнаружит, что домена не существует (на практике поиск ...
|
||||
[2.0s t=154]
|
||||
|
||||
YOU: проверь статус uptime kuma
|
||||
MAVEN: записала: проверь статус uptime kuma
|
||||
[1.6s t=155]
|
||||
|
||||
YOU: всё ли в порядке с сервером?
|
||||
MAVEN: пока не умею отвечать на этот вопрос.
|
||||
[1.3s t=156]
|
||||
|
||||
YOU: спасибо
|
||||
MAVEN: спасибо за пожелание
|
||||
[4.0s t=157]
|
||||
|
||||
```
|
||||
|
||||
## Day 9
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: рада,
|
||||
[2.1s t=158]
|
||||
|
||||
YOU: какой сегодня день недели?
|
||||
MAVEN: сегодня суббота, 8 августа 2026 года
|
||||
[0.0s t=159]
|
||||
|
||||
YOU: добавь хлеб и сыр в список покупок
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[1.3s t=160]
|
||||
|
||||
YOU: что в списке покупок?
|
||||
MAVEN: в списке покупок пусто.
|
||||
[0.0s t=161]
|
||||
|
||||
YOU: удали молоко из списка
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[1.2s t=162]
|
||||
|
||||
YOU: напомни завтра утром вынести мусор
|
||||
MAVEN: Сейчас 20:22. Когда?
|
||||
[0.3s t=163]
|
||||
|
||||
YOU: запиши что я поменял масло в машине
|
||||
MAVEN: Прошлую просьбу отпускаю. Такую команду я не знаю.
|
||||
[1.3s t=164]
|
||||
|
||||
YOU: когда я менял масло?
|
||||
MAVEN: вот что я нашла: по данным из last update 2019 год, заправки Сочи не могли предоставить масло для замены, поэтому это было невозможно.
|
||||
[7.1s t=165]
|
||||
|
||||
YOU: сколько будет 144 делить на 12?
|
||||
MAVEN: есть такое: Разделить в столбик 144 на 12 - Calc-Best.ru — 144 : 12 = 12 Где 144 - делимое, 12 - делитель, 12 - частное ... Разделить 770 на 7 столбиком Разделить 441 на 9 столбиком Разделить 891 на 9 столбиком Разделить 60 на 5 столбиком Разделить 60 на 6 столбиком
|
||||
[9.6s t=166]
|
||||
|
||||
YOU: пока
|
||||
MAVEN: я записала
|
||||
[4.8s t=167]
|
||||
|
||||
```
|
||||
|
||||
## Day 10
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: рада
|
||||
[4.1s t=168]
|
||||
|
||||
YOU: что нового за ночь?
|
||||
MAVEN: нашла вот это: Главные новости сегодня в России, а также самые последние новости со ... — Полное бессилие: киевская ПВО вновь пропустила все российские ракеты. Вооруженные силы РФ в ночь на 8 августа нанесли второй за последнее время массированный ...
|
||||
[2.1s t=169]
|
||||
|
||||
YOU: почему интернет медленный?
|
||||
MAVEN: вот что я нашла: Почему интернет стал медленнее: причины и способы решения - АО «Газпром ... — На скорость интернета сильно влияют торрент-клиенты, VPN, потоковое видео, фоновая загрузка системных обновлений. Посмотреть, куда уходит трафик, можно в настройках ПК или смартфона.
|
||||
[2.4s t=170]
|
||||
|
||||
YOU: какая скорость у меня сейчас?
|
||||
MAVEN: для какого города?
|
||||
[1.2s t=171]
|
||||
|
||||
YOU: запиши что новый роутер стоит 8000 рублей
|
||||
MAVEN: Я записала стоимость нового роутера.
|
||||
[2.5s t=172]
|
||||
|
||||
YOU: сколько стоил роутер?
|
||||
MAVEN: вот что я нашла: новый роутер стоит 8000 рублей.
|
||||
[2.0s t=173]
|
||||
|
||||
YOU: что такое NAT?
|
||||
MAVEN: NAT (Network Address Translation) — это технология, которая позволяет устройствам в локальной сети выходить в интернет, используя один общий внешний IP-адрес. Я отвечаю на этот вопрос, используя поиск в сети.
|
||||
[1.8s t=174]
|
||||
|
||||
YOU: напомни в субботу позвонить бабушке
|
||||
MAVEN: Сейчас 20:23. Когда?
|
||||
[0.3s t=175]
|
||||
|
||||
YOU: покажи мои напоминания
|
||||
MAVEN: вот что я нашла: в твоих источниках нет напоминаний, только запись о том, что ты купил новые наушники, на какое время поставить напоминание?
|
||||
[1.7s t=176]
|
||||
|
||||
YOU: спасибо
|
||||
MAVEN: Сейчас 20:23. Во сколько напомнить?
|
||||
[1.0s t=177]
|
||||
|
||||
```
|
||||
|
||||
## Day 11
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: Сейчас 20:23. Скажи время — например, «в семь вечера» или «через час».
|
||||
[0.9s t=178]
|
||||
|
||||
YOU: как погода на выходных?
|
||||
MAVEN: для какого города? На какое время поставить напоминание?
|
||||
[0.6s t=179]
|
||||
|
||||
YOU: что у меня на этой неделе?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи, на какое время поставить напоминание?
|
||||
[0.3s t=180]
|
||||
|
||||
YOU: запиши что я хочу прочитать книгу про Go
|
||||
MAVEN: Прошлую просьбу отпускаю. Я записала твою заметку о книге по Go.
|
||||
[1.4s t=181]
|
||||
|
||||
YOU: что я хотел прочитать?
|
||||
MAVEN: вот что я нашла: ты хотел прочитать книгу про Go.
|
||||
[1.1s t=182]
|
||||
|
||||
YOU: объясни что такое горутина
|
||||
MAVEN: вот что я нашла: я узнала, что горутины — это легковесные потоки выполнения или функции в Go, которые позволяют параллельно выполнять код. Они управляются планировщиком Go и работают как независимые нити выполнения.
|
||||
[1.9s t=183]
|
||||
|
||||
YOU: кто написал Войну и мир?
|
||||
MAVEN: вот что я нашла: роман «Война и мир» написал Лев Николаевич Толстой.
|
||||
[2.4s t=184]
|
||||
|
||||
YOU: включи свет на кухне
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[0.5s t=185]
|
||||
|
||||
YOU: закрой шторы в комнате
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[0.3s t=186]
|
||||
|
||||
YOU: спокойной ночи
|
||||
MAVEN: Спокойной ночи. Отдыхай, я на связи, если что-то понадобится.
|
||||
[1.2s t=187]
|
||||
|
||||
```
|
||||
|
||||
## Day 12
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: Я не могу сказать тебе «доброе утро», потому что сейчас уже вечер.
|
||||
[1.0s t=188]
|
||||
|
||||
YOU: сколько сейчас времени?
|
||||
MAVEN: сейчас 20 часов 23 минуты
|
||||
[0.0s t=189]
|
||||
|
||||
YOU: я не то имел в виду
|
||||
MAVEN: Я поняла, исправлюсь. Расскажи, что именно ты хотел сказать?
|
||||
[1.4s t=190]
|
||||
|
||||
YOU: о чём мы говорили?
|
||||
MAVEN: Я помню всё, что ты мне говоришь, и сейчас отвечаю на твои вопросы на основе этих данных.
|
||||
[1.4s t=191]
|
||||
|
||||
YOU: напомни
|
||||
MAVEN: Сейчас 20:23. Когда?
|
||||
[0.7s t=192]
|
||||
|
||||
YOU: сделай это
|
||||
MAVEN: пока не умею отвечать на этот вопрос, на какое время поставить напоминание?
|
||||
[0.5s t=193]
|
||||
|
||||
YOU: запиши что я перешёл на новый тариф
|
||||
MAVEN: Прошлую просьбу отпускаю. Я записала твою новую информацию о тарифе.
|
||||
[1.5s t=194]
|
||||
|
||||
YOU: какой у меня тариф?
|
||||
MAVEN: вот что я нашла: ты перешел на новый тариф
|
||||
[24.4s t=195]
|
||||
|
||||
YOU: сколько я плачу за интернет?
|
||||
MAVEN: есть такое: Как узнать, сколько и когда платить за домашний интернет и ТВ — 3 февр. 2026 г. ... Как узнать, сколько и когда платить за домашний интернет и ТВ · Для тарифов только с домашним интернетом · Была ли статья полезной?
|
||||
[2.1s t=196]
|
||||
|
||||
YOU: спасибо
|
||||
MAVEN: Спасибо, что обратилась за помощью.
|
||||
[4.8s t=197]
|
||||
|
||||
```
|
||||
|
||||
## Day 13
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: Привет!
|
||||
[2.5s t=198]
|
||||
|
||||
YOU: что там с задачами?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[1.2s t=199]
|
||||
|
||||
YOU: закрывай
|
||||
MAVEN: какой пункт отметить сделанным?
|
||||
[0.0s t=200]
|
||||
|
||||
YOU: отметь задачу про бэкапы как сделанную
|
||||
MAVEN: Заметка сохранена.
|
||||
[5.8s t=201]
|
||||
|
||||
YOU: что осталось нерешённым?
|
||||
MAVEN: вот что я нашла: Российское регулирование криптовалют: что осталось нерешённым к 1 июля — 2 июл. 2026 г. ... Российское регулирование криптовалют: что осталось нерешённым к 1 июля. Пакет законопроектов о регулировании цифровых активов в России не ...
|
||||
[3.7s t=202]
|
||||
|
||||
YOU: запиши что я договорился о встрече в среду
|
||||
MAVEN: Заметка сохранена.
|
||||
[2.7s t=203]
|
||||
|
||||
YOU: когда у меня встреча?
|
||||
MAVEN: мне пока нечего смотреть — у Praxis нет источников.
|
||||
[0.1s t=204]
|
||||
|
||||
YOU: какая температура на улице?
|
||||
MAVEN: для какого города?
|
||||
[1.7s t=205]
|
||||
|
||||
YOU: что такое RAID 5?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[1.2s t=206]
|
||||
|
||||
YOU: пока
|
||||
MAVEN: Я записала, что мы продолжаем разговор.
|
||||
[4.1s t=207]
|
||||
|
||||
```
|
||||
|
||||
## Day 14
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: рада, доброго утра!
|
||||
[4.5s t=208]
|
||||
|
||||
YOU: подведи итоги недели
|
||||
MAVEN: записала: подведи итоги недели
|
||||
[1.7s t=209]
|
||||
|
||||
YOU: что я делал за последние две недели?
|
||||
MAVEN: вот что я нашла: ты не делал ничего за последние две недели.
|
||||
[6.6s t=210]
|
||||
|
||||
YOU: какие заметки я сохранил?
|
||||
MAVEN: Ты сочти, что сохранила заметку с текстом 'Проверь календарь' и датой 8 августа 2026 г.,
|
||||
[5.4s t=211]
|
||||
|
||||
YOU: о чём я чаще всего спрашиваю?
|
||||
MAVEN: обычно ты пьёшь воду около 01:12, отмечаешь «name» около 23:56 и отмечаешь «week summary» около 23:59 — по записям за последние 7 дней.
|
||||
[4.4s t=212]
|
||||
|
||||
YOU: напомни в понедельник в 10 проверить бэкапы
|
||||
MAVEN: Сейчас 20:25. Это утра или вечера?
|
||||
[0.7s t=213]
|
||||
|
||||
YOU: что у меня в понедельник?
|
||||
MAVEN: вот что я нашла: у тебя в понедельник дедлайн по проекту, на какое время поставить напоминание?
|
||||
[4.6s t=214]
|
||||
|
||||
YOU: ты меня понимаешь?
|
||||
MAVEN: Да, я тебя прекрасно понимаю. Я готова помогать тебе с заметками и задачами, на какое время поставить напоминание?
|
||||
[2.4s t=215]
|
||||
|
||||
YOU: спасибо тебе
|
||||
MAVEN: Сейчас 20:25. Это утра или вечера?
|
||||
[1.0s t=216]
|
||||
|
||||
YOU: спокойной ночи
|
||||
MAVEN: хорошо, напомню послезавтра в 10:00.
|
||||
[0.9s t=217]
|
||||
|
||||
```
|
||||
@@ -0,0 +1,105 @@
|
||||
# Two weeks of talking to Maven, as a baseline to re-run
|
||||
|
||||
Date: 2026-08-08.
|
||||
Build: `beb093a` on master, the five compose services as deployed, 41 hours up.
|
||||
Reach: `POST /api/chat` on mavweb, 140 turns over fourteen simulated days.
|
||||
Turn source is `tap:text`, so this exercises the path the mic and telegram take.
|
||||
|
||||
This exists to be compared against. `scripts/usage-run.py` and
|
||||
`scripts/testdata/usage-turns.txt` are in the repo, so a re-run after a routing
|
||||
change is a diff rather than a new opinion. The 2026-08-07 week of usage was
|
||||
typed by hand and cannot be replayed.
|
||||
|
||||
**It measures master, not the branch.** V-655, V-659 and V-660 are unmerged.
|
||||
Every query source that guesses is still in the chain. That is the change this
|
||||
baseline is for.
|
||||
|
||||
## What re-runs and what does not
|
||||
|
||||
The turns file, the driver and the routing behaviour replay. Three things do
|
||||
not. The wall clock was 20:18 to 20:27 throughout, so every clock and agenda
|
||||
answer reads evening. Live search and the feed return different text each day.
|
||||
And the store carries over between runs. A fact written on day 2 is already
|
||||
present when a re-run reaches day 1.
|
||||
|
||||
## Numbers
|
||||
|
||||
| | week (2026-08-07) | fortnight (2026-08-08) |
|
||||
|---|---|---|
|
||||
| turns | 74 | 140 |
|
||||
| p50 | 1.5s | 1.6s |
|
||||
| p95 | 8.0s | 7.1s |
|
||||
| max | 12.3s | 33.7s |
|
||||
| transport errors | 0 | 0 |
|
||||
|
||||
| string in the reply | turns |
|
||||
|---|---|
|
||||
| `на какое время поставить напоминание` | 13 |
|
||||
| `не нашла у тебя такой записи` | 8 |
|
||||
| `Такую команду я не знаю` | 8 |
|
||||
| `для какого города` | 6 |
|
||||
| `В какой день` | 6 |
|
||||
| `пока не умею` | 5 |
|
||||
| `Когда?` | 3 |
|
||||
|
||||
**Zero transport errors is not zero wrong answers.** It counts turns that
|
||||
failed to return a reply, and none did. Every quality number is below.
|
||||
|
||||
Those seven strings appear in 49 of 140 turns. Some turns carry two, because a
|
||||
parked clarify appends to whatever else was said.
|
||||
|
||||
The 33.7s outlier is one note write on day 2. p95 improved against the week
|
||||
despite it.
|
||||
|
||||
## The three defects worth diffing against
|
||||
|
||||
### 1. A parked reminder clarify still contaminates later turns
|
||||
|
||||
The week test called this the single worst thing to talk to and it is unchanged.
|
||||
Nineteen turns carry a clarify tail. The worst run is day 1, turns 7 to 13,
|
||||
which spans a day boundary:
|
||||
|
||||
```
|
||||
что такое TCP? -> для какого города? На какое время поставить напоминание?
|
||||
сколько будет 17 на 23? -> для какого города? На какое время поставить напоминание?
|
||||
спасибо -> Сейчас 20:18. В какой день?
|
||||
привет -> Сейчас 20:18. В какой день?
|
||||
```
|
||||
|
||||
Note that `привет` and `спасибо` do not clear it, and neither does a new day.
|
||||
|
||||
### 2. Query sources that guess still claim turns they cannot answer
|
||||
|
||||
Weather took `сколько будет 17 на 23?`, `что такое TCP?` and `какая скорость у
|
||||
меня сейчас?`, answering `для какого города?` to all three. The feed took
|
||||
`какой у меня любимый язык?` and `хватает ли места под новые бэкапы?` and
|
||||
answered with kernel headlines.
|
||||
|
||||
This is the exact class V-655 removes by marking a source `guesses: true` and
|
||||
taking it out of `queryWalk`. Six turns here, so the re-run has a number to move.
|
||||
|
||||
### 3. A question can still be read as a capture
|
||||
|
||||
`что я сохранил про Сочи?` answered `Хорошо, сохраню.` The utterance is
|
||||
interrogative and was routed to a write. `IsQuestionShaped` catches this
|
||||
downstream on some paths and did not catch it here.
|
||||
|
||||
## What did work
|
||||
|
||||
Reminders with a spoken time land correctly, which is V-572 holding:
|
||||
`напомни завтра в 9 утра купить хлеб` returned `хорошо, напомню завтра в 09:00.`
|
||||
|
||||
Facts round-trip. `запиши что новый роутер стоит 8000 рублей` then `сколько
|
||||
стоил роутер?` returned the stored value. So did the wifi password and the
|
||||
doctor's appointment.
|
||||
|
||||
World questions answer when no local source claims them first. `что такое NAT?`
|
||||
returned a real definition.
|
||||
|
||||
Stage 0 answers land at 0.0 to 0.4s, unchanged.
|
||||
|
||||
## What this does not cover
|
||||
|
||||
The voice loop, because `mavwaked` and `mavenclient` are not deployed. Reminder
|
||||
delivery, because nothing fired inside the run window. Telegram intake. And the
|
||||
three-head routing model, which does not run in Go at all.
|
||||
@@ -0,0 +1,22 @@
|
||||
package router
|
||||
|
||||
import (
|
||||
"os"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// TestDumpPrompt writes the router prompt and grammar to disk so the training
|
||||
// workspace labels with the daemon's own contract rather than a retyped copy.
|
||||
// It is inert unless MAVEN_DUMP_PROMPT names a directory.
|
||||
func TestDumpPrompt(t *testing.T) {
|
||||
dir := os.Getenv("MAVEN_DUMP_PROMPT")
|
||||
if dir == "" {
|
||||
t.Skip("MAVEN_DUMP_PROMPT unset")
|
||||
}
|
||||
if err := os.WriteFile(dir+"/route_system.txt", []byte(routeSystem), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(dir+"/route_grammar.gbnf", []byte(routeGrammar), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
@@ -36,17 +36,27 @@ var fixtureJSON []byte
|
||||
//
|
||||
// Intent is empty exactly when WantClarify is set: the contract there is that
|
||||
// the router refuses instead of guessing.
|
||||
//
|
||||
// WantSource is a pointer because the destination has three states and a bare
|
||||
// string only has two (V-659). Absent means the case does not score a
|
||||
// destination at all, which is every intent but query: a fact, a reminder, a
|
||||
// note, an act, a chat or a system turn never reaches queryWalk. Present and
|
||||
// empty is the SourceUnknown contract — the decider must name nothing and let
|
||||
// the daemon walk the whole chain, which is the right answer whenever two
|
||||
// destinations can both answer and the utterance does not choose. Present and
|
||||
// named is a destination the route must produce.
|
||||
type Case struct {
|
||||
ID string `json:"id"`
|
||||
Utterance string `json:"utterance"`
|
||||
Lang string `json:"lang"`
|
||||
Intent router.Intent `json:"intent"`
|
||||
WantTime bool `json:"want_time"`
|
||||
WantFn bool `json:"want_fn"`
|
||||
WantFactKey string `json:"want_fact_key"`
|
||||
WantClarify bool `json:"want_clarify"`
|
||||
Tags []string `json:"tags"`
|
||||
Note string `json:"note"`
|
||||
ID string `json:"id"`
|
||||
Utterance string `json:"utterance"`
|
||||
Lang string `json:"lang"`
|
||||
Intent router.Intent `json:"intent"`
|
||||
WantTime bool `json:"want_time"`
|
||||
WantFn bool `json:"want_fn"`
|
||||
WantFactKey string `json:"want_fact_key"`
|
||||
WantClarify bool `json:"want_clarify"`
|
||||
WantSource *router.Source `json:"want_source,omitempty"`
|
||||
Tags []string `json:"tags"`
|
||||
Note string `json:"note"`
|
||||
}
|
||||
|
||||
// Fixture — the versioned envelope, same shape as
|
||||
@@ -118,6 +128,11 @@ type Outcome struct {
|
||||
// (a slot gap is a parser fix; a wrong intent is a router fix).
|
||||
IntentOK bool
|
||||
Reasons []string
|
||||
// SourceReason is set when the case labelled a destination and the route
|
||||
// named a different one. It is kept out of Reasons on purpose: the
|
||||
// destination is the second half of a route and it is scored separately,
|
||||
// so a wrong destination must not move the intent number (V-659).
|
||||
SourceReason string
|
||||
}
|
||||
|
||||
// Report — the aggregate. Accuracy is the headline; the rest exists so a
|
||||
@@ -139,7 +154,15 @@ type Report struct {
|
||||
// (reminder grammar → applyAction's time parser). Not a miss, but not a
|
||||
// full router-level win either; tracked so the two aren't conflated.
|
||||
SlotsDeferred int
|
||||
Outcomes []Outcome
|
||||
// SourceTotal counts the cases carrying a want_source, and SourceHit the
|
||||
// ones whose route named it. Reported apart from Passed because intent and
|
||||
// destination are two decisions, and one number hides which one moved.
|
||||
SourceTotal int
|
||||
SourceHit int
|
||||
// SourceConfusion counts want→got destination pairs. "" reads as the
|
||||
// SourceUnknown floor on either side.
|
||||
SourceConfusion map[string]int
|
||||
Outcomes []Outcome
|
||||
// Confusion counts want→got intent pairs, decided cases only.
|
||||
Confusion map[string]int
|
||||
// ByTag accuracy for the fixture's tags ("hard", "homelab", …).
|
||||
@@ -172,6 +195,17 @@ func (r Report) IntentAccuracy() float64 {
|
||||
return float64(r.IntentHit) / float64(r.Total)
|
||||
}
|
||||
|
||||
// SourceAccuracy — fraction of the labelled cases whose route named the right
|
||||
// destination. Denominator is SourceTotal and not Total, because most of the
|
||||
// fixture never reaches a query source and scoring those would report a
|
||||
// percentage of nothing.
|
||||
func (r Report) SourceAccuracy() float64 {
|
||||
if r.SourceTotal == 0 {
|
||||
return 0
|
||||
}
|
||||
return float64(r.SourceHit) / float64(r.SourceTotal)
|
||||
}
|
||||
|
||||
// Score runs every case through r and aggregates. It never fails the run on a
|
||||
// route error — an erroring case scores as a miss and is counted in Errors,
|
||||
// because "the model was down" and "the model was wrong" are different numbers
|
||||
@@ -186,11 +220,12 @@ func Score(ctx context.Context, name string, r Router, f Fixture) (Report, error
|
||||
return Report{}, err
|
||||
}
|
||||
rep := Report{
|
||||
Name: name,
|
||||
Total: len(f.Cases),
|
||||
Confusion: map[string]int{},
|
||||
ByTag: map[string]TagStat{},
|
||||
ByLang: map[string]TagStat{},
|
||||
Name: name,
|
||||
Total: len(f.Cases),
|
||||
Confusion: map[string]int{},
|
||||
SourceConfusion: map[string]int{},
|
||||
ByTag: map[string]TagStat{},
|
||||
ByLang: map[string]TagStat{},
|
||||
}
|
||||
lat := make([]time.Duration, 0, len(f.Cases))
|
||||
|
||||
@@ -242,6 +277,29 @@ func Score(ctx context.Context, name string, r Router, f Fixture) (Report, error
|
||||
}
|
||||
}
|
||||
|
||||
// The destination is scored outside the switch and outside Pass. A case
|
||||
// that clarified or landed the wrong intent named no destination, and
|
||||
// that is a real miss rather than a case to skip — otherwise the
|
||||
// denominator quietly drops every turn the route already lost. Only a
|
||||
// route error is skipped, because "the model was down" is the Errors
|
||||
// number and not a destination result.
|
||||
if c.WantSource != nil && err == nil {
|
||||
rep.SourceTotal++
|
||||
switch {
|
||||
case !o.IntentOK:
|
||||
// The route never got to a destination, so a match on the
|
||||
// SourceUnknown floor here would be a coincidence scored as a
|
||||
// win: a clarify names nothing and would satisfy "" for free.
|
||||
o.SourceReason = fmt.Sprintf("no destination, route missed %q", c.Intent)
|
||||
rep.SourceConfusion[string(*c.WantSource)+"→(no route)"]++
|
||||
case d.Source == *c.WantSource:
|
||||
rep.SourceHit++
|
||||
default:
|
||||
rep.SourceConfusion[string(*c.WantSource)+"→"+string(d.Source)]++
|
||||
o.SourceReason = fmt.Sprintf("source %q, want %q", d.Source, *c.WantSource)
|
||||
}
|
||||
}
|
||||
|
||||
o.Pass = len(o.Reasons) == 0
|
||||
if o.Pass {
|
||||
rep.Passed++
|
||||
@@ -298,25 +356,40 @@ func (r Report) String() string {
|
||||
r.Name, r.Passed, r.Total, 100*r.Accuracy(), 100*r.IntentAccuracy())
|
||||
fmt.Fprintf(&b, " clarify: %d false (asked, shouldn't) / %d missed (guessed, shouldn't) | errors: %d | slots deferred to daemon: %d\n",
|
||||
r.FalseClarify, r.MissedClarify, r.Errors, r.SlotsDeferred)
|
||||
if r.SourceTotal > 0 {
|
||||
fmt.Fprintf(&b, " destination: %d/%d labelled cases (%.1f%%)\n",
|
||||
r.SourceHit, r.SourceTotal, 100*r.SourceAccuracy())
|
||||
}
|
||||
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
|
||||
fmt.Fprintf(&b, " by lang: %s\n", renderStats(r.ByLang))
|
||||
fmt.Fprintf(&b, " by tag: %s\n", renderStats(r.ByTag))
|
||||
if len(r.Confusion) > 0 {
|
||||
fmt.Fprintf(&b, " confusion: %s\n", renderCounts(r.Confusion))
|
||||
}
|
||||
if len(r.SourceConfusion) > 0 {
|
||||
fmt.Fprintf(&b, " destination confusion: %s\n", renderCounts(r.SourceConfusion))
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
|
||||
// Failures — the per-case detail, sorted by ID so two runs diff cleanly.
|
||||
// Failures — the per-case detail, sorted by ID so two runs diff cleanly. A case
|
||||
// that landed its intent and missed its destination is listed too, marked, so
|
||||
// the half that moved is readable without diffing two percentages.
|
||||
func (r Report) Failures() string {
|
||||
var b strings.Builder
|
||||
out := append([]Outcome(nil), r.Outcomes...)
|
||||
sort.Slice(out, func(i, j int) bool { return out[i].Case.ID < out[j].Case.ID })
|
||||
for _, o := range out {
|
||||
if o.Pass {
|
||||
continue
|
||||
switch {
|
||||
case !o.Pass:
|
||||
reasons := o.Reasons
|
||||
if o.SourceReason != "" {
|
||||
reasons = append(append([]string(nil), reasons...), o.SourceReason)
|
||||
}
|
||||
fmt.Fprintf(&b, " %s %q: %s\n", o.Case.ID, o.Case.Utterance, strings.Join(reasons, "; "))
|
||||
case o.SourceReason != "":
|
||||
fmt.Fprintf(&b, " %s %q: route ok, %s\n", o.Case.ID, o.Case.Utterance, o.SourceReason)
|
||||
}
|
||||
fmt.Fprintf(&b, " %s %q: %s\n", o.Case.ID, o.Case.Utterance, strings.Join(o.Reasons, "; "))
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
|
||||
@@ -266,6 +266,11 @@ func baselineGrammars(acts router.ActMatcher) []router.Grammar {
|
||||
// Same order as buildRouter (voicewire.go). The fixture is only worth
|
||||
// anything while its grammar set is the daemon's grammar set.
|
||||
grammars = append(grammars, router.AgendaQueryGrammars()...)
|
||||
// After the agenda rules and before the feed and list rules, same as
|
||||
// voicewire.go: "что такое лента" is a definition question and the feed
|
||||
// rule would claim it on the noun alone (V-655). Missing here until V-659,
|
||||
// so the fixture was scoring a grammar set the daemon does not run.
|
||||
grammars = append(grammars, router.WorldQueryGrammars()...)
|
||||
grammars = append(grammars, router.FeedQueryGrammar())
|
||||
// The list side of the same exposure: a phrasing with no possessive in it
|
||||
// ("список дел") routed system and never reached queryTasks (Vikunja #467).
|
||||
|
||||
@@ -6,37 +6,41 @@
|
||||
"Held-out routing contract. Every utterance here is absent from models/seeds/*.txt (TestFixtureIsHeldOut enforces it verbatim) — scoring a classifier on its own seed phrases measures memorisation, not routing.",
|
||||
"This is a CONTRACT, not a snapshot of current behaviour. Cases the classifier cascade fails today are expected to stay in the file and fail loudly; that failure count is the number Vikunja #319 compares against the LLM router before #320 flips the default.",
|
||||
"Slot expectations are deployment-independent on purpose. want_fn is a boolean (the act must resolve to SOME allowlisted fn) because the allowlist lives in deploy config, not here. want_fact_key names the loop's rule keys (water/meal/sleep/break/shower) — a fact that lands under the wrong key silently starves the predicate that reads it.",
|
||||
"want_clarify cases carry intent \"\": the contract is that the router refuses rather than guesses. A confident answer there is a worse failure than a miss."
|
||||
"want_clarify cases carry intent \"\": the contract is that the router refuses rather than guesses. A confident answer there is a worse failure than a miss.",
|
||||
"want_source is the second half of a route (V-655). It is present only on query cases, because no other intent reaches queryWalk, and absent there means absent rather than SourceUnknown. Empty is a label and not a gap: it asserts that the decider must name nothing and let the daemon walk the whole chain in order, his data first.",
|
||||
"Seven cases assert that floor and six of them are homelab operations. They cluster because SourceRecall, SourceNetwork and SourceAttention overlap on every question about the box: mavpoll writes its observations into the fact store recall reads. That is a finding about the enum, not a gap in the labelling.",
|
||||
"ru-query-020 and ru-query-024 are the same utterance, as are ru-query-021 and ru-query-025. Both pairs differ in tags and note only, so both pairs are counted twice in every number this fixture reports.",
|
||||
"Every want_source is the destination that SHOULD claim the turn, which on ru-query-026 through 030 is not the one that did. Those five were observed failing on the box on 2026-08-07 (docs/evals/2026-08-07-week-of-usage.md). A fixture that passes on the day it is written measures nothing."
|
||||
],
|
||||
"cases": [
|
||||
{ "id": "ru-query-001", "utterance": "сколько воды я выпил с утра", "lang": "ru", "intent": "query", "tags": ["aggregate"] },
|
||||
{ "id": "ru-query-002", "utterance": "я сегодня вообще пил воду", "lang": "ru", "intent": "query", "tags": ["hard", "fact-shaped"], "note": "past-tense fact lexicon in a question — the classifier's fact centroid pulls this hard" },
|
||||
{ "id": "ru-query-003", "utterance": "во сколько я лёг вчера", "lang": "ru", "intent": "query", "tags": ["temporal"] },
|
||||
{ "id": "ru-query-004", "utterance": "давно я не тренировался", "lang": "ru", "intent": "query", "tags": ["hard", "no-question-word"] },
|
||||
{ "id": "ru-query-005", "utterance": "напоминания на завтра есть", "lang": "ru", "intent": "query", "tags": ["hard", "reminder-shaped"], "note": "asks about reminders, does not create one" },
|
||||
{ "id": "ru-query-006", "utterance": "что я записывал про кота", "lang": "ru", "intent": "query", "tags": ["recall"] },
|
||||
{ "id": "ru-query-007", "utterance": "сколько раз я ел вчера", "lang": "ru", "intent": "query", "tags": ["aggregate", "hard"] },
|
||||
{ "id": "ru-query-008", "utterance": "мой вес за последний месяц", "lang": "ru", "intent": "query", "tags": ["no-verb"] },
|
||||
{ "id": "ru-query-009", "utterance": "когда я в последний раз принимал витамины", "lang": "ru", "intent": "query", "tags": ["temporal"] },
|
||||
{ "id": "ru-query-010", "utterance": "есть новости по бэкапу базы", "lang": "ru", "intent": "query", "tags": ["homelab"] },
|
||||
{ "id": "ru-query-011", "utterance": "почему сервер тормозит", "lang": "ru", "intent": "query", "tags": ["homelab", "hard"], "note": "diagnostic question, not a chat opener" },
|
||||
{ "id": "ru-query-012", "utterance": "какие заметки я оставил про полив", "lang": "ru", "intent": "query", "tags": ["recall"] },
|
||||
{ "id": "ru-query-013", "utterance": "во сколько у меня встреча", "lang": "ru", "intent": "query", "tags": ["calendar"] },
|
||||
{ "id": "ru-query-019", "utterance": "что у меня стоит в календаре на послезавтра", "lang": "ru", "intent": "query", "tags": ["calendar", "hard"], "note": "agenda, not the clock: the daemon answers this from CalendarEvents inside the query branch, so the clock/date system rule must not swallow it" },
|
||||
{ "id": "ru-query-022", "utterance": "какие планы на завтра?", "lang": "ru", "intent": "query", "tags": ["calendar"], "note": "the same agenda question as ru-query-019 aimed at another day; it answered \u043f\u043e\u043a\u0430 \u043d\u0435 \u0443\u043c\u0435\u044e on the deployed daemon while the today form worked (Vikunja #471)" },
|
||||
{ "id": "ru-query-023", "utterance": "\u043a\u043e\u0433\u0434\u0430 \u043f\u043b\u0430\u043d\u0451\u0440\u043a\u0430?", "lang": "ru", "intent": "query", "tags": ["calendar", "hard"], "note": "a named event with no calendar word — the noun is the only signal that this is a question about his day" },
|
||||
{ "id": "ru-query-024", "utterance": "что дальше?", "lang": "ru", "intent": "query", "tags": ["calendar", "no-question-word"], "note": "the rest of the day, with no possessive and no plan word to anchor on; the model called it a fact and the write had to be caught downstream (Vikunja #498)" },
|
||||
{ "id": "ru-query-025", "utterance": "расскажи про битву при Ватерлоо", "lang": "ru", "intent": "query", "tags": ["world", "no-question-word"], "note": "a narrative request carries no question mark and no interrogative, so it routed fact; contrast ru-chat-003, where the same verb asks for a joke" },
|
||||
{ "id": "ru-query-014", "utterance": "я успеваю до дедлайна", "lang": "ru", "intent": "query", "tags": ["hard", "no-question-word"] },
|
||||
{ "id": "ru-query-015", "utterance": "сколько я прошёл шагов", "lang": "ru", "intent": "query", "tags": ["aggregate"] },
|
||||
{ "id": "ru-query-016", "utterance": "покажи давление за неделю", "lang": "ru", "intent": "query", "tags": ["hard", "imperative"], "note": "imperative form but a read — must not route to act" },
|
||||
{ "id": "ru-query-017", "utterance": "чем я занимался в среду", "lang": "ru", "intent": "query", "tags": ["hard", "chat-shaped"] },
|
||||
{ "id": "ru-query-018", "utterance": "хватает ли места под новые бэкапы", "lang": "ru", "intent": "query", "tags": ["homelab"] },
|
||||
{ "id": "ru-query-020", "utterance": "что дальше?", "lang": "ru", "intent": "query", "tags": ["agenda", "hard"], "note": "the rest of the day, with no interrogative the model can read as a question — it routed fact until a stage 0 rule claimed it (V-498)" },
|
||||
{ "id": "ru-query-021", "utterance": "расскажи про битву при Ватерлоо", "lang": "ru", "intent": "query", "tags": ["world", "hard"], "note": "a world question phrased as an instruction. It routed fact, and the fact gate had to catch the write (V-498)" },
|
||||
{ "id": "en-query-001", "utterance": "did I take my vitamins today", "lang": "en", "intent": "query", "tags": ["fact-shaped"] },
|
||||
{ "id": "en-query-002", "utterance": "how long since the last backup finished", "lang": "en", "intent": "query", "tags": ["temporal"] },
|
||||
{ "id": "en-query-003", "utterance": "show me this week's weight", "lang": "en", "intent": "query", "tags": ["imperative"] },
|
||||
{ "id": "ru-query-001", "utterance": "сколько воды я выпил с утра", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["aggregate"] },
|
||||
{ "id": "ru-query-002", "utterance": "я сегодня вообще пил воду", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["hard", "fact-shaped"], "note": "past-tense fact lexicon in a question — the classifier's fact centroid pulls this hard" },
|
||||
{ "id": "ru-query-003", "utterance": "во сколько я лёг вчера", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["temporal"] },
|
||||
{ "id": "ru-query-004", "utterance": "давно я не тренировался", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["hard", "no-question-word"] },
|
||||
{ "id": "ru-query-005", "utterance": "напоминания на завтра есть", "lang": "ru", "intent": "query", "want_source": "", "tags": ["hard", "reminder-shaped"], "note": "asks about reminders, does not create one. want_source is the floor on purpose: no query source reads the reminder store, and day-plan is SourceCalendar over a table this box does not write." },
|
||||
{ "id": "ru-query-006", "utterance": "что я записывал про кота", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["recall"] },
|
||||
{ "id": "ru-query-007", "utterance": "сколько раз я ел вчера", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["aggregate", "hard"] },
|
||||
{ "id": "ru-query-008", "utterance": "мой вес за последний месяц", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["no-verb"] },
|
||||
{ "id": "ru-query-009", "utterance": "когда я в последний раз принимал витамины", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["temporal"] },
|
||||
{ "id": "ru-query-010", "utterance": "есть новости по бэкапу базы", "lang": "ru", "intent": "query", "want_source": "", "tags": ["homelab"], "note": "recall, attention and network can each answer it, because mavpoll writes its netdata and uptime-kuma observations into the fact store recall reads. Naming one takes the other two off the turn." },
|
||||
{ "id": "ru-query-011", "utterance": "почему сервер тормозит", "lang": "ru", "intent": "query", "want_source": "", "tags": ["homelab", "hard"], "note": "diagnostic question, not a chat opener. network holds the box and attention holds the alarm about the box. The utterance does not choose, so neither does the label." },
|
||||
{ "id": "ru-query-012", "utterance": "какие заметки я оставил про полив", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["recall"] },
|
||||
{ "id": "ru-query-013", "utterance": "во сколько у меня встреча", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar"] },
|
||||
{ "id": "ru-query-019", "utterance": "что у меня стоит в календаре на послезавтра", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar", "hard"], "note": "agenda, not the clock: the daemon answers this from CalendarEvents inside the query branch, so the clock/date system rule must not swallow it" },
|
||||
{ "id": "ru-query-022", "utterance": "какие планы на завтра?", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar"], "note": "the same agenda question as ru-query-019 aimed at another day; it answered \u043f\u043e\u043a\u0430 \u043d\u0435 \u0443\u043c\u0435\u044e on the deployed daemon while the today form worked (Vikunja #471)" },
|
||||
{ "id": "ru-query-023", "utterance": "\u043a\u043e\u0433\u0434\u0430 \u043f\u043b\u0430\u043d\u0451\u0440\u043a\u0430?", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar", "hard"], "note": "a named event with no calendar word — the noun is the only signal that this is a question about his day" },
|
||||
{ "id": "ru-query-024", "utterance": "что дальше?", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar", "no-question-word"], "note": "the rest of the day, with no possessive and no plan word to anchor on; the model called it a fact and the write had to be caught downstream (Vikunja #498)" },
|
||||
{ "id": "ru-query-025", "utterance": "расскажи про битву при Ватерлоо", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "no-question-word"], "note": "a narrative request carries no question mark and no interrogative, so it routed fact; contrast ru-chat-003, where the same verb asks for a joke" },
|
||||
{ "id": "ru-query-014", "utterance": "я успеваю до дедлайна", "lang": "ru", "intent": "query", "want_source": "", "tags": ["hard", "no-question-word"], "note": "a deadline lives in the task list, the calendar or Praxis depending on where he put it. The destination depends on his data, not on his words." },
|
||||
{ "id": "ru-query-015", "utterance": "сколько я прошёл шагов", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["aggregate"] },
|
||||
{ "id": "ru-query-016", "utterance": "покажи давление за неделю", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["hard", "imperative"], "note": "imperative form but a read — must not route to act" },
|
||||
{ "id": "ru-query-017", "utterance": "чем я занимался в среду", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["hard", "chat-shaped"] },
|
||||
{ "id": "ru-query-018", "utterance": "хватает ли места под новые бэкапы", "lang": "ru", "intent": "query", "want_source": "", "tags": ["homelab"], "note": "disk headroom. network is the only source that reads the box, but the phrasing is a capacity question and not a LAN one." },
|
||||
{ "id": "ru-query-020", "utterance": "что дальше?", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["agenda", "hard"], "note": "the rest of the day, with no interrogative the model can read as a question — it routed fact until a stage 0 rule claimed it (V-498)" },
|
||||
{ "id": "ru-query-021", "utterance": "расскажи про битву при Ватерлоо", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "hard"], "note": "a world question phrased as an instruction. It routed fact, and the fact gate had to catch the write (V-498)" },
|
||||
{ "id": "en-query-001", "utterance": "did I take my vitamins today", "lang": "en", "intent": "query", "want_source": "recall", "tags": ["fact-shaped"] },
|
||||
{ "id": "en-query-002", "utterance": "how long since the last backup finished", "lang": "en", "intent": "query", "want_source": "", "tags": ["temporal"], "note": "the completion time is a fact the poller wrote, so recall answers it. A person asking this wants the operational answer. Both are true." },
|
||||
{ "id": "en-query-003", "utterance": "show me this week's weight", "lang": "en", "intent": "query", "want_source": "recall", "tags": ["imperative"] },
|
||||
|
||||
{ "id": "ru-fact-001", "utterance": "только что выпил кружку воды", "lang": "ru", "intent": "fact", "want_fact_key": "water" },
|
||||
{ "id": "ru-fact-002", "utterance": "воды попил наконец", "lang": "ru", "intent": "fact", "want_fact_key": "water", "tags": ["inverted"] },
|
||||
@@ -106,6 +110,11 @@
|
||||
{ "id": "amb-005", "utterance": "потом", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "filler"] },
|
||||
{ "id": "amb-006", "utterance": "the thing from earlier", "lang": "en", "want_clarify": true, "tags": ["ambiguous", "anaphora"] },
|
||||
{ "id": "amb-007", "utterance": "напомни", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "reminder"], "note": "the reminder verb and nothing else — she knows the shape of the request and not one thing about it. Answered 'не получилось разобрать время напоминания' on the box until V-548: the subjectless-reminder gate tested Slots.Text == \"\", and fillSlots had put the verb in that slot" },
|
||||
{ "id": "amb-008", "utterance": "ну напомни же", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "reminder", "filler"], "note": "the same request wrapped in particles, which is why filler_particles is a lexicon set — without it the particles read as the subject" }
|
||||
{ "id": "amb-008", "utterance": "ну напомни же", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "reminder", "filler"], "note": "the same request wrapped in particles, which is why filler_particles is a lexicon set — without it the particles read as the subject" },
|
||||
{ "id": "ru-query-026", "utterance": "что такое TCP?", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "regression"], "note": "weather claimed it on 2026-08-07 and answered \"для какого города?\", because it read one percent closer than the leftover seeds. WorldQueryGrammars claims it at stage 0 now." },
|
||||
{ "id": "ru-query-027", "utterance": "сколько будет 17 на 23?", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "arithmetic", "regression"], "note": "same source, same day, same answer about a city. Arithmetic is not a place." },
|
||||
{ "id": "ru-query-028", "utterance": "какой у меня любимый язык?", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["recall", "possessive", "regression"], "note": "the feed answered it with kernel headlines. \"у меня\" is the whole signal and it points inward." },
|
||||
{ "id": "ru-query-029", "utterance": "кто такой Линус Торвальдс?", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "person", "regression"], "note": "the personal boundary answered \"не нашла у тебя такой записи\". A named public person is not his data." },
|
||||
{ "id": "ru-query-030", "utterance": "что там с бэкапами?", "lang": "ru", "intent": "query", "want_source": "", "tags": ["homelab", "regression"], "note": "search claimed it, which inverts the boundary outward. The fix is the chain order and not a destination: recall, attention and network all answer it, same as ru-query-010." }
|
||||
]
|
||||
}
|
||||
|
||||
@@ -45,12 +45,19 @@ const routeGrammar = `
|
||||
root ::= "[" ws action ("," ws action)* ws "]"
|
||||
action ::= "{" ws "\"intent\"" ws ":" ws intent ("," ws field)* ws "}"
|
||||
intent ::= "\"fact\"" | "\"reminder\"" | "\"note\"" | "\"query\"" | "\"act\"" | "\"chat\"" | "\"system\"" | "\"unknown\""
|
||||
field ::= key ws ":" ws string
|
||||
field ::= (key ws ":" ws string) | ("\"source\"" ws ":" ws source)
|
||||
key ::= "\"key\"" | "\"value\"" | "\"text\"" | "\"verb\""
|
||||
source ::= "\"recall\"" | "\"calendar\"" | "\"tasks\"" | "\"list\"" | "\"money\"" | "\"weather\"" | "\"home\"" | "\"network\"" | "\"feeds\"" | "\"attention\"" | "\"self\"" | "\"world\"" | "\"\""
|
||||
string ::= "\"" ([^"\\\x00-\x1F] | "\\" ["\\/bfnrt] | "\\u" [0-9a-fA-F]{4}){0,120} "\""
|
||||
ws ::= [ \t\n]{0,4}
|
||||
`
|
||||
|
||||
// TestRouteGrammarCoversSources holds the source rule above to router.Sources.
|
||||
// The enum is the point: a grammar cannot emit a destination that does not
|
||||
// exist, which is the guarantee V-546 wants from a softmax and gets here for
|
||||
// free. Empty is the thirteenth alternative and it is not an oversight — it is
|
||||
// the SourceUnknown floor, and the model must be able to decline.
|
||||
|
||||
// routeSystem — the router prompt. Changed 31-07-2026: the query test now sits
|
||||
// above the fact test and there is an explicit question test. Before that, a
|
||||
// question naming a fact key ("сколько воды я выпил с утра") matched the fact
|
||||
@@ -121,6 +128,31 @@ const routeSystem = `Классифицируй ровно одно сообще
|
||||
"что такое кватернион?" → {"intent":"query","text":"что такое кватернион"}
|
||||
"ага" → {"intent":"chat","text":"ага"}
|
||||
|
||||
Только для query добавь поле source — где лежит ответ:
|
||||
- recall — его заметки, факты и то, что он раньше говорил
|
||||
- calendar — встречи и события
|
||||
- tasks — список задач
|
||||
- list — списки покупок и другие именованные списки
|
||||
- money — траты
|
||||
- weather — погода
|
||||
- home — свет, устройства, дом
|
||||
- network — локальная сеть, сервер, диски
|
||||
- feeds — новостные ленты
|
||||
- attention — что требует внимания сейчас
|
||||
- self — вопрос про самого ассистента
|
||||
- world — всё остальное: определения, счёт, люди, факты о мире
|
||||
|
||||
Пустое значение "" — нормальный ответ и его надо ставить часто. Ставь "", если ответ могут дать сразу два источника или если не уверен: тогда проверяются все по порядку, и это правильно. Никогда не угадывай.
|
||||
|
||||
"сколько воды я выпил с утра" → {"intent":"query","text":"сколько воды я выпил с утра","source":"recall"}
|
||||
"что я записывал про кота" → {"intent":"query","text":"что я записывал про кота","source":"recall"}
|
||||
"во сколько у меня встреча" → {"intent":"query","text":"во сколько у меня встреча","source":"calendar"}
|
||||
"что такое docker?" → {"intent":"query","text":"что такое docker","source":"world"}
|
||||
"кто такой Линус Торвальдс?" → {"intent":"query","text":"кто такой Линус Торвальдс","source":"world"}
|
||||
"сколько будет 17 на 23?" → {"intent":"query","text":"сколько будет 17 на 23","source":"world"}
|
||||
"почему сервер тормозит" → {"intent":"query","text":"почему сервер тормозит","source":""}
|
||||
"есть новости по бэкапу базы" → {"intent":"query","text":"есть новости по бэкапу базы","source":""}
|
||||
|
||||
Ответ — JSON-массив: по одному объекту на каждую просьбу. Обычно один. Если в реплике несколько просьб — по объекту на каждую. "напомни купить молоко, и запиши что кофе кончился" → [{"intent":"reminder","text":"купить молоко"},{"intent":"note","text":"кофе кончился"}]. Только JSON, без пояснений.`
|
||||
|
||||
// routeRepeatPenalty — the sub-1B model loops one sentence inside the text field
|
||||
@@ -172,6 +204,7 @@ type routeAction struct {
|
||||
Value string `json:"value"`
|
||||
Text string `json:"text"`
|
||||
Verb string `json:"verb"`
|
||||
Source string `json:"source"`
|
||||
}
|
||||
|
||||
// Route asks the model for one decision. The bool is false when there is no
|
||||
@@ -240,6 +273,14 @@ func (lr *LLMRouter) Route(ctx context.Context, utterance string, now time.Time)
|
||||
case IntentQuery:
|
||||
d.Intent = IntentQuery
|
||||
d.Slots.Text = firstNonEmpty(a.Text, utterance)
|
||||
// Through ValidSource, and on query alone. The grammar already bounds
|
||||
// the enum, but the grammar is a request to a server that may be
|
||||
// running a different build, and a destination this binary does not
|
||||
// know would take real query sources off the turn. Anything unknown
|
||||
// drops to SourceUnknown, which is the floor and costs nothing.
|
||||
if ValidSource(Source(a.Source)) {
|
||||
d.Source = Source(a.Source)
|
||||
}
|
||||
case IntentAct:
|
||||
d.Intent = IntentAct
|
||||
d.Slots.Text = firstNonEmpty(a.Verb, utterance)
|
||||
|
||||
@@ -398,3 +398,70 @@ func TestLLMReminderWithSubjectIsNotGated(t *testing.T) {
|
||||
t.Fatalf("a complete reminder was sent back as a question: %+v", d.Slots)
|
||||
}
|
||||
}
|
||||
|
||||
// TestRouteGrammarCoversSources — the grammar enum and router.Sources are two
|
||||
// hand-written lists of the same twelve destinations, and nothing else notices
|
||||
// when one grows. A destination missing from the grammar is a destination the
|
||||
// model is structurally unable to name, which is the exact defect V-517
|
||||
// measured for Praxis: not a weak model, an absent string.
|
||||
func TestRouteGrammarCoversSources(t *testing.T) {
|
||||
for _, s := range Sources {
|
||||
if !strings.Contains(routeGrammar, `"\"`+string(s)+`\""`) {
|
||||
t.Errorf("routeGrammar cannot emit %q — the model can never name it", s)
|
||||
}
|
||||
}
|
||||
// The floor has to be reachable too, or the model is forced to pick one.
|
||||
if !strings.Contains(routeGrammar, `"\"\""`) {
|
||||
t.Error(`routeGrammar cannot emit "" — the model cannot decline a destination`)
|
||||
}
|
||||
// Count the alternatives on the source rule: an extra one is a destination
|
||||
// the daemon would drop to SourceUnknown after the model spent tokens on it.
|
||||
for _, line := range strings.Split(routeGrammar, "\n") {
|
||||
if !strings.HasPrefix(line, "source ") {
|
||||
continue
|
||||
}
|
||||
if got, want := strings.Count(line, "|")+1, len(Sources)+1; got != want {
|
||||
t.Errorf("source rule has %d alternatives, want %d (Sources plus the floor)", got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The destination is read back only through ValidSource. A model on an older or
|
||||
// newer build can write a string this binary does not know, and trusting it
|
||||
// would take real query sources off the turn for a name nothing answers.
|
||||
func TestLLMUnknownSourceFallsToTheFloor(t *testing.T) {
|
||||
r := newLLMTestRouter(t, `{"intent":"query","text":"что там с бэкапами","source":"praxis"}`)
|
||||
d, err := r.Route(context.Background(), "что там с бэкапами", refNow())
|
||||
if err != nil {
|
||||
t.Fatalf("route: %v", err)
|
||||
}
|
||||
if d.Source != SourceUnknown {
|
||||
t.Fatalf("invented destination %q was trusted, want the floor", d.Source)
|
||||
}
|
||||
}
|
||||
|
||||
// And a known one survives, or the read-back is just a filter.
|
||||
func TestLLMNamedSourceSurvives(t *testing.T) {
|
||||
r := newLLMTestRouter(t, `{"intent":"query","text":"кто такой Линус Торвальдс","source":"world"}`)
|
||||
d, err := r.Route(context.Background(), "кто такой Линус Торвальдс?", refNow())
|
||||
if err != nil {
|
||||
t.Fatalf("route: %v", err)
|
||||
}
|
||||
if d.Source != SourceWorld {
|
||||
t.Fatalf("source %q, want %q", d.Source, SourceWorld)
|
||||
}
|
||||
}
|
||||
|
||||
// A destination on anything but a query is dropped. Only IntentQuery reaches
|
||||
// queryWalk, so a source elsewhere is a field nobody reads and a claim nobody
|
||||
// checks.
|
||||
func TestLLMSourceIsQueryOnly(t *testing.T) {
|
||||
r := newLLMTestRouter(t, `{"intent":"note","text":"кофе кончился","source":"recall"}`)
|
||||
d, err := r.Route(context.Background(), "запиши что кофе кончился", refNow())
|
||||
if err != nil {
|
||||
t.Fatalf("route: %v", err)
|
||||
}
|
||||
if d.Source != SourceUnknown {
|
||||
t.Fatalf("a note carried destination %q", d.Source)
|
||||
}
|
||||
}
|
||||
|
||||
Vendored
+167
@@ -0,0 +1,167 @@
|
||||
# Day 1
|
||||
доброе утро
|
||||
какой сегодня день?
|
||||
сколько времени?
|
||||
запиши что я пью кофе без сахара
|
||||
мой любимый язык программирования go
|
||||
напомни в 11:00 позвонить маме
|
||||
что у меня сегодня?
|
||||
что такое TCP?
|
||||
сколько будет 17 на 23?
|
||||
спасибо
|
||||
|
||||
# Day 2
|
||||
привет
|
||||
что нового?
|
||||
какая погода?
|
||||
запиши что пароль от вайфая лежит в ящике стола
|
||||
где лежит вайфай пароль?
|
||||
добавь молоко в список покупок
|
||||
что у меня в списке покупок?
|
||||
кто такой Линус Торвальдс?
|
||||
какой у меня любимый язык?
|
||||
сколько у меня задач?
|
||||
|
||||
# Day 3
|
||||
как дела?
|
||||
напомни завтра в 9 утра купить хлеб
|
||||
что у меня завтра?
|
||||
отмени напоминание про хлеб
|
||||
какие у меня напоминания?
|
||||
сохрани мне адрес гостиницы в Сочи
|
||||
что я сохранил про Сочи?
|
||||
почему сервер тормозит?
|
||||
хватает ли места под новые бэкапы?
|
||||
выключи свет в спальне
|
||||
|
||||
# Day 4
|
||||
доброе утро
|
||||
что я пропустил?
|
||||
о чём мы вчера говорили?
|
||||
запиши что я записался к врачу на четверг
|
||||
когда я иду к врачу?
|
||||
что такое ZFS?
|
||||
столица Франции?
|
||||
переведи слово ремонт на английский
|
||||
сколько я потратил в этом месяце?
|
||||
спокойной ночи
|
||||
|
||||
# Day 5
|
||||
привет
|
||||
какая погода в Москве?
|
||||
что там с бэкапами?
|
||||
покажи что требует внимания
|
||||
отметь это как сделанное
|
||||
запиши что я купил новые наушники
|
||||
какие у меня заметки за неделю?
|
||||
расскажи про Kubernetes
|
||||
кто я?
|
||||
пока
|
||||
|
||||
# Day 6
|
||||
доброе утро
|
||||
сколько времени?
|
||||
напомни в 18:30 позвонить в банк
|
||||
поставь чайник
|
||||
включи музыку
|
||||
что у меня в календаре на пятницу?
|
||||
во сколько у меня встреча?
|
||||
запиши что дедлайн по проекту в понедельник
|
||||
успею ли я до дедлайна?
|
||||
спасибо
|
||||
|
||||
# Day 7
|
||||
привет
|
||||
как ты?
|
||||
расскажи анекдот
|
||||
что ты умеешь?
|
||||
запиши что я начал бегать по утрам
|
||||
я бегаю по утрам уже неделю
|
||||
как часто я бегаю?
|
||||
сколько стоит биткоин?
|
||||
какие новости?
|
||||
хорошего дня
|
||||
|
||||
# Day 8
|
||||
доброе утро
|
||||
что у меня сегодня?
|
||||
напомни через час выпить воды
|
||||
я выпил воды
|
||||
запиши что кот ест только сухой корм
|
||||
чем питается кот?
|
||||
что такое DNS?
|
||||
проверь статус uptime kuma
|
||||
всё ли в порядке с сервером?
|
||||
спасибо
|
||||
|
||||
# Day 9
|
||||
привет
|
||||
какой сегодня день недели?
|
||||
добавь хлеб и сыр в список покупок
|
||||
что в списке покупок?
|
||||
удали молоко из списка
|
||||
напомни завтра утром вынести мусор
|
||||
запиши что я поменял масло в машине
|
||||
когда я менял масло?
|
||||
сколько будет 144 делить на 12?
|
||||
пока
|
||||
|
||||
# Day 10
|
||||
доброе утро
|
||||
что нового за ночь?
|
||||
почему интернет медленный?
|
||||
какая скорость у меня сейчас?
|
||||
запиши что новый роутер стоит 8000 рублей
|
||||
сколько стоил роутер?
|
||||
что такое NAT?
|
||||
напомни в субботу позвонить бабушке
|
||||
покажи мои напоминания
|
||||
спасибо
|
||||
|
||||
# Day 11
|
||||
привет
|
||||
как погода на выходных?
|
||||
что у меня на этой неделе?
|
||||
запиши что я хочу прочитать книгу про Go
|
||||
что я хотел прочитать?
|
||||
объясни что такое горутина
|
||||
кто написал Войну и мир?
|
||||
включи свет на кухне
|
||||
закрой шторы в комнате
|
||||
спокойной ночи
|
||||
|
||||
# Day 12
|
||||
доброе утро
|
||||
сколько сейчас времени?
|
||||
я не то имел в виду
|
||||
о чём мы говорили?
|
||||
напомни
|
||||
сделай это
|
||||
запиши что я перешёл на новый тариф
|
||||
какой у меня тариф?
|
||||
сколько я плачу за интернет?
|
||||
спасибо
|
||||
|
||||
# Day 13
|
||||
привет
|
||||
что там с задачами?
|
||||
закрывай
|
||||
отметь задачу про бэкапы как сделанную
|
||||
что осталось нерешённым?
|
||||
запиши что я договорился о встрече в среду
|
||||
когда у меня встреча?
|
||||
какая температура на улице?
|
||||
что такое RAID 5?
|
||||
пока
|
||||
|
||||
# Day 14
|
||||
доброе утро
|
||||
подведи итоги недели
|
||||
что я делал за последние две недели?
|
||||
какие заметки я сохранил?
|
||||
о чём я чаще всего спрашиваю?
|
||||
напомни в понедельник в 10 проверить бэкапы
|
||||
что у меня в понедельник?
|
||||
ты меня понимаешь?
|
||||
спасибо тебе
|
||||
спокойной ночи
|
||||
@@ -0,0 +1,106 @@
|
||||
"""Drive a fortnight of conversation through POST /api/chat and record it.
|
||||
|
||||
The 2026-08-07 week of usage was typed by hand. This is the same reach and the
|
||||
same turn source, tap:text, so it exercises the path the mic and telegram take.
|
||||
|
||||
The endpoint is a form POST that redirects to /chat with the reply in the query
|
||||
string. Reading the Location header is the whole protocol, so nothing here
|
||||
parses HTML.
|
||||
|
||||
This exists to be re-run. The baseline is 2026-08-08 against master at beb093a,
|
||||
in docs/evals/2026-08-08-two-weeks.md. Re-running the same turns after a routing
|
||||
change is the comparison, so edit the turns file by adding, never by rewriting.
|
||||
|
||||
python3 scripts/usage-run.py scripts/testdata/usage-turns.txt out-prefix
|
||||
|
||||
Input is one utterance per line. A line starting with "# " opens a day. A blank
|
||||
line is ignored. Output is a markdown transcript and a jsonl log beside it.
|
||||
"""
|
||||
|
||||
import json
|
||||
import sys
|
||||
import time
|
||||
import urllib.error
|
||||
import urllib.parse
|
||||
import urllib.request
|
||||
|
||||
URL = "http://127.0.0.1:9201/api/chat"
|
||||
TIMEOUT = 90
|
||||
|
||||
|
||||
class NoRedirect(urllib.request.HTTPRedirectHandler):
|
||||
"""A 303 carries the reply. Following it would throw the reply away."""
|
||||
|
||||
def redirect_request(self, *a, **kw):
|
||||
return None
|
||||
|
||||
|
||||
# ProxyHandler({}) is not optional. This box exports http_proxy, urllib honours
|
||||
# it, and the proxy answers 503 for a loopback address.
|
||||
OPENER = urllib.request.build_opener(NoRedirect, urllib.request.ProxyHandler({}))
|
||||
|
||||
|
||||
def turn(text):
|
||||
body = urllib.parse.urlencode({"text": text}).encode()
|
||||
t0 = time.perf_counter()
|
||||
try:
|
||||
OPENER.open(urllib.request.Request(URL, data=body), timeout=TIMEOUT)
|
||||
return {"reply": "", "error": "no redirect", "secs": time.perf_counter() - t0}
|
||||
except urllib.error.HTTPError as e:
|
||||
dt = time.perf_counter() - t0
|
||||
if e.code != 303:
|
||||
return {"reply": "", "error": f"HTTP {e.code}", "secs": dt}
|
||||
loc = e.headers.get("Location", "")
|
||||
q = urllib.parse.parse_qs(urllib.parse.urlparse(loc).query)
|
||||
return {
|
||||
"reply": q.get("r", [""])[0],
|
||||
"source": q.get("src", q.get("source", [""]))[0],
|
||||
"trace": q.get("t", [""])[0],
|
||||
"secs": dt,
|
||||
}
|
||||
except Exception as e: # a dead box must not lose the turns already done
|
||||
return {"reply": "", "error": str(e), "secs": time.perf_counter() - t0}
|
||||
|
||||
|
||||
def main():
|
||||
lines = [l.rstrip("\n") for l in open(sys.argv[1])]
|
||||
prefix = sys.argv[2]
|
||||
md = open(prefix + "-transcript.md", "w")
|
||||
log = open(prefix + ".jsonl", "w")
|
||||
|
||||
day = 0
|
||||
n = 0
|
||||
print(f"# Raw transcript, two weeks of usage\n", file=md)
|
||||
for line in lines:
|
||||
if not line.strip():
|
||||
continue
|
||||
if line.startswith("# "):
|
||||
if day:
|
||||
print("```\n", file=md)
|
||||
day += 1
|
||||
print(f"## {line[2:]}\n\n```", file=md)
|
||||
continue
|
||||
n += 1
|
||||
r = turn(line)
|
||||
r["day"] = day
|
||||
r["n"] = n
|
||||
r["utterance"] = line
|
||||
log.write(json.dumps(r, ensure_ascii=False) + "\n")
|
||||
log.flush()
|
||||
reply = r.get("error") or r["reply"]
|
||||
print(f"YOU: {line}", file=md)
|
||||
print(f"MAVEN: {reply}", file=md)
|
||||
tag = f"[{r['secs']:.1f}s"
|
||||
if r.get("source"):
|
||||
tag += f" src={r['source']}"
|
||||
print(f" {tag} t={r.get('trace', '')}]\n", file=md)
|
||||
md.flush()
|
||||
print(f"{n:3} d{day} {r['secs']:5.1f}s {line[:40]:40s} -> {reply[:60]}",
|
||||
flush=True)
|
||||
print("```", file=md)
|
||||
md.close()
|
||||
log.close()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Reference in New Issue
Block a user