Compare commits

...

4 Commits

Author SHA1 Message Date
claude f55bedee2e Train the destination head and beat the teacher (V-661)
Step 3 of the routing-heads plan. Intent and destination share one masked mean
pool on e5-small. Destination scores 26/33 against 12/33 for the classifier
cascade and 24/33 for the cascade with gemma-4-12b, which is the teacher these
labels were distilled from. Recall goes 0/15 to 15/15.

Two heads, not four, and both cuts are label problems rather than GPU time.
Mood describes her own reply state and no dataset maps onto it. BIO slot tags
have no Maven-domain corpus.

The MASSIVE warm-start from step 2 is worth nothing here either. Stock ties it
on intent and leads by a third of a case on destination.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-08 19:38:23 +04:00
claude e470435cf1 Dump the router prompt where the labeler can read it (V-661)
The training workspace labels with routeSystem and routeGrammar, and it held
its own copies. V-660 changed both. A retyped prompt drifts silently, which is
the problem llm/check_prompt_parity.py exists for on the other side.

Inert unless MAVEN_DUMP_PROMPT names a directory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-08 18:44:47 +04:00
claude 00f9239ef9 Record the model arm, and the stage 0 trade it exposed (V-660)
Numbers and the argument in docs/evals/2026-08-08-destination-model-arm.md,
pointer and the short version in CLAUDE.md. The finding worth carrying is
not the 72.7%: it is that stage 0's silence on the possessive agenda rules
used to be free and now costs four destination points, because there is
finally something downstream that would have named the calendar.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-08 18:23:22 +04:00
claude 3513e508b7 Give the router prompt a destination to write (V-660)
V-659 measured the destination at 12/33 on the classifier cascade and named
the gap: recall 0/15, because nothing anywhere names it. The model could not
help, for a structural reason rather than a capability one. Nothing in
routeSystem mentioned a Source and routeGrammar could not emit one, so there
was no string for it to write. Same shape as the Praxis reach V-517
measured at 0/12.

routeGrammar grows a source rule, closed over router.Sources plus the empty
floor. A grammar cannot emit a destination that does not exist, which is the
guarantee V-546 wants from a softmax and gets here for free. The prompt
lists the twelve in Russian, one line each, and says plainly that "" is a
normal answer to give often: two sources that can both answer means the
chain walks, and guessing is the failure mode this whole field exists to
stop.

The read-back goes through ValidSource and runs on IntentQuery alone. The
grammar already bounds the enum, but it is a request to a server that may be
running another build, and only a query reaches queryWalk.

Measured against gemma-4-12b on the workstation, same fixture, cascade with
a hash fallback: destination 24/33 (72.7%) against the classifier's 12/33,
and intent 81/96 (84.4%) which is where it already was. Recall is the whole
move, 0/15 to 14/15. The model alone scores 26/33.

Four cases the cascade loses and llm-only wins are calendar. The possessive
agenda rules claim them at stage 0 and deliberately name nothing, because
"что у меня в списке покупок" matches the same rule and naming the calendar
would take the list source off the turn. So stage 0's caution now costs four
destination points it did not cost before. That is a real trade and it wants
its own argument, not a quiet edit here.

The resident Qwen3-1.7B is unmeasured: it binds --port 0 inside the
container and no host process can reach it.

llm/check_prompt_parity.py in the training workspace compares its copy of
routeSystem to this one and will fail until that copy gets the same edit.
V-362 covers the catch-up.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-08 18:22:00 +04:00
6 changed files with 417 additions and 4 deletions
+46 -3
View File
@@ -246,6 +246,33 @@ was a hardcode. **Fine-tune a copy of the weights.** The resident embedder backs
recall. Training it in place couples routing accuracy to recall@1, with nothing in the
suite to name the trade.
**Two of those heads are trained as of 08-08-2026, and they are not the three
above** (V-661, `docs/evals/2026-08-08-routing-heads-two-head.md`). Intent and
destination share one masked mean pool. Destination scores **26/33 (78.8%)** on
the fixture. The classifier cascade scores 12/33 and the cascade with gemma-4-12b
scores 24/33, so a 118M encoder beats the 12B teacher it was distilled from.
Recall is 15/15 and world is 5/5. Intent is 93.6% mean over three seeds. That is
**not** comparable to the 76.0% and 84.4% those two arms scored: a softmax has no
clarify class, so the head's fixture is the 88 cases carrying an intent.
**Mood is cut, not deferred.** The enum describes her own reply state, not the
speaker's emotion, and no dataset maps onto it. **BIO slot tags have no
Maven-domain corpus**, so they stay in the MASSIVE body from step 2. Both are
label problems and neither is a GPU problem: the run is under four minutes.
The MASSIVE warm-start of step 2 is worth nothing here. Stock e5-small ties it on
intent and leads by a third of a case on destination. Nothing argues for keeping
that step.
What the head gets wrong is the floor. It names a destination where the fixture
says walk the chain, and it is confident doing it. `"почему сервер тормозит"`
reads `world` at 0.80. The training floor is generated ambiguous questions and the fixture floor
is homelab operations, which are not the same distribution.
**Nothing of this runs in Go.** The weights are `heads.pt` and `out/body_heads/`
on workpc. Reaching the daemon needs an ONNX export and a caller. The resident
e5-small must not be replaced by the copy, because recall depends on that file.
`Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
#359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with
@@ -433,9 +460,25 @@ not a gap in the labelling.
so the fixture scored a grammar set nobody runs. Fixed by V-659, worth 3 points of
destination and nothing else. Check that function when adding a grammar.
The model arm is still the follow-up. It lands on V-546. Intent, mood and BIO slot
tags were already three heads on one forward pass of the resident e5-small.
Destination is a fourth head on the same pass.
**The model arm landed the same day** (V-660,
`docs/evals/2026-08-08-destination-model-arm.md`). `routeGrammar` carries a
`source` rule closed over `router.Sources` plus the empty floor, so the model
cannot emit a destination that does not exist. The prompt lists the twelve in
Russian and says `""` is a normal answer to give often. `LLMRouter.Route` reads it
back through `ValidSource` and on `IntentQuery` alone. Against gemma-4-12b on the
workstation the cascade scores destination **24/33 (72.7%)** with intent unmoved
at 84.4%, and **recall goes 0/15 to 14/15**. The resident Qwen3-1.7B is
unmeasured, because it binds `--port 0` inside the container.
**Stage 0 now costs four destination points.** It did not before. The four cases
the cascade loses and the model alone wins are all calendar. The possessive
agenda rules claim them first and name nothing on purpose. That caution was free
while nothing downstream could name anything either. It is not free now, and the
fix is the owner's call rather than a quiet edit.
The last arm is V-546. Intent, mood and BIO slot tags were already three heads on
one forward pass of the resident e5-small. Destination is a fourth head on the
same pass, and 72.7% from a 12B teacher is the label source for training it.
## LLM output contract
@@ -0,0 +1,66 @@
# The destination, with a model that can name one
Measured 2026-08-08 against gemma-4-12b on the workstation, the same 96-case
fixture V-659 built. Covers V-660.
```sh
no_proxy='*' MAVEN_LLM_URL=http://192.168.1.105:8080 \
make t PKG=./internal/router/eval/ RUN=TestLLMRouterBaseline V=1
```
## The gap was structural
V-659 measured the destination at 12/33 on the classifier cascade, with recall
at 0/15. Nothing in `routeSystem` named a `Source` and `routeGrammar` could not
emit one, so the resident model had no string to write. That is the shape V-517
measured for Praxis reach at 0/12: not a weak model, an absent contract.
`routeGrammar` now carries a `source` rule closed over `router.Sources` plus the
empty floor. The prompt lists the twelve destinations in Russian and says that
`""` is a normal answer to give often.
## Result
| run | intent | destination |
|---|---|---|
| classifier + ONNX (V-659) | 73/96 (76.0%) | 12/33 (36.4%) |
| gemma-4-12b alone | 79/96 intent-only (82.3%) | 26/33 (78.8%) |
| cascade + gemma-4-12b + hash fallback | 81/96 (84.4%) | 24/33 (72.7%) |
Recall is the move: 0/15 to 14/15. Intent did not shift, which was the
constraint. The prompt is shared, so a destination rule that costs routing
points is not a win.
The eight llm-only errors are the eight `want_clarify` cases. The model returned
`unknown` on every one, which is correct, and the llm-only harness surfaces a
decline as an error by design.
## Stage 0 now costs four destination points
The four cases the cascade loses and the model alone wins are all calendar. The
possessive agenda rules claim them at stage 0 and deliberately name nothing.
"что у меня в списке покупок" matches the same rule. Naming the calendar there
would take the list source off the turn (V-655).
So a rule written to be careful about the list now blocks a model that would
have named the calendar correctly. Before V-660 that caution was free, because
nothing downstream of stage 0 could name anything either.
Three ways out, and each costs something. Split the possessive rule so the
calendar-shaped half names its destination. Let a later stage overwrite an empty
destination a grammar left behind, which reverses "a matched value always wins".
Or leave it, on the argument that four points is cheap next to a wrong
destination on a shopping list. This wants the owner's call rather than a quiet
edit.
## What this does not measure
The resident Qwen3-1.7B, which is what homesrv runs. It binds `--port 0` inside
the container and no host process can reach it. Scoring it needs a second
llama-server on a fixed port. The workstation is never assumed
up, so the homesrv number is the one that decides whether this ships on by
default.
The fixture is 33 labelled destinations over twelve values. Recall carries 15 of
them and five destinations carry none at all. A per-destination number below
world, recall, calendar and the floor is not supported by this fixture.
@@ -0,0 +1,174 @@
# Two heads on e5-small, and the first destination the router did not need a model for
Measured 2026-08-08 on workpc (Radeon RX 7900 GRE, ROCm). Covers V-661, step 3 of
`docs/plans/18-routing-heads-on-e5-small.md`. Workspace is `~/Programs/embed-training`,
scripts `gen_query_source.py`, `label_source.py`, `build_heads_corpus.py`,
`train_heads.py`, `score_confidence.py`.
## Two heads, not four
Intent is 7 classes and destination is 12 plus the `SourceUnknown` floor, sharing
one masked mean pool over one forward pass. The plan asked for four. Two of them
have no labels and neither is a GPU problem.
**Mood is cut, not deferred.** Maven's enum is `neutral, happy, thinking, tired,
confused` and it describes her own reply state, not the speaker's emotion.
`psytechlab/EmpatheticIntents-ru` was the only candidate and its 32 emotion
labels do not map onto it. There is nothing to train against.
**BIO slot tags stay in the MASSIVE body.** No Maven-domain span corpus exists.
`2026-08-08-massive-warm-start.md` records that no second Russian slot-filling
corpus is reachable at all.
The destination loss is masked with `ignore_index`. Only a query turn reaches
`queryWalk`, so a reminder contributes nothing to it.
## Where the destination labels came from
V-660 taught the router prompt to name a destination. That made gemma-4-12b a
teacher, and this distils it.
Labelling the 300 query rows already in `train_v5.jsonl` gave 277 destinations.
The shape was unusable: the floor 101, calendar 62, recall 47, and `feeds` and
`attention` at zero. A 13-way softmax cannot learn a class with no examples.
`gen_query_source.py` is the destination half of `gen_corpus.py` and runs the
same two passes. Gemma writes questions whose answer lives in one named place.
The daemon's own `routeSystem` prompt then routes each one back. A line survives
only when the intent is `query` **and** the source is the destination it was
generated for. The glosses are copied verbatim out of `route_system.txt`, so the
generator and the labeller work from one definition.
The prompt and the GBNF are dumped from `internal/router/llmrouter.go` by
`TestDumpPrompt`, never retyped. The workspace held its own copies and V-660
changed both.
1229 kept of 2373 generated, 51.8%. Merged corpus is 3664 rows carrying 1727
destinations:
| | rows | | rows |
|---|---|---|---|
| the floor | 220 | world | 132 |
| calendar | 180 | money | 126 |
| tasks | 169 | list | 124 |
| recall | 167 | self, feeds, attention | 120 each |
| weather | 136 | network | 103 |
| | | home | 50 |
`home` is thin because the agreement filter rejected most of what was generated
for it. A question about the house routes `act` more often than `query`. That is
the filter working, and 50 is the finding rather than a shortfall.
The 1229 generated rows carry `intent: null`. Every one is a query by
construction. There are five times as many as the corpus has query rows, so
including them would make query half the intent corpus.
## Result
Three seeds, two bodies, epoch chosen on the intent dev slice and never on a
destination number.
| body | intent mean | destination mean |
|---|---|---|
| warm-started `out/body_massive` | 93.6% | 75.8% |
| stock `multilingual-e5-small` | 93.6% | 76.8% |
Best single run is destination **26/33 (78.8%)**, reached by both bodies at seed
0. Peak 1.68GB of 17.2GB, under four minutes end to end.
Against the two arms already measured on the same 33 labelled cases:
| | destination |
|---|---|
| classifier cascade (V-659) | 12/33 (36.4%) |
| cascade + gemma-4-12b (V-660) | 24/33 (72.7%) |
| two heads on e5-small | 26/33 (78.8%) |
A 118M encoder beats the 12B teacher it was distilled from, on the fixture. The
per-destination split is where it happens: **recall 15/15** and **world 5/5**.
Recall was 0/15 on the cascade and 14/15 through gemma.
Intent is **not** comparable to the 73/96 and 81/96 figures those two arms
scored. A softmax has no clarify class. The head's fixture is the 88 cases that
carry an intent, and the 8 `want_clarify` cases are scored separately below.
## The MASSIVE warm-start is worth nothing here either
Step 2 measured it at +0.4 points of intent accuracy and called that inside seed
noise. Destination was the open question, because MASSIVE has a
`definition_word` slot that looked like a `SourceWorld` signal sitting in a head
already trained.
It is not. The two bodies score the same intent mean to one decimal. Stock is
one point ahead on destination, which is a third of one case. Nothing here argues
for keeping the warm-start step. Dropping it removes a dependency on a corpus
pull that `datasets` 5.0 cannot do.
## What the head gets wrong is the floor
All seven destination misses at seed 0 are the floor and calendar:
```
(floor) 3/7
calendar 3/6
recall 15/15
world 5/5
```
The head names a destination where the fixture says walk the chain, and it is
confident doing it. `"почему сервер тормозит"` reads `world` at 0.80.
`"хватает ли места под новые бэкапы"` reads `network` at 0.82. Those are the six
homelab cases V-659 flagged, where `SourceRecall`, `SourceNetwork` and
`SourceAttention` all overlap because `mavpoll` writes its observations into the
fact store recall reads.
The training floor is generated ambiguous questions. The fixture floor is
homelab operations. Those are not the same distribution and the head learned the
one it was given.
## Max softmax separates, weakly, and the gate stays
The plan argues max softmax is a calibratable confidence where `Confidence: 1.0`
was a hardcode. Measured on the intent head:
| | n | mean confidence |
|---|---|---|
| correct | 80 | 0.897 |
| wrong | 8 | 0.705 |
| `want_clarify` | 8 | 0.685 |
The softest correct answer is 0.66 and 3 of 8 clarify cases sit below it. So a
single cut buys three clarifies at no false-clarify cost, and no more.
The other five explain themselves. `"напомни"` scores 0.94 as `reminder` and
`"сделай это"` scores 0.80 as `act`. Both are intent-certain and slot-empty, and
confidence was never the signal there. `gateLLMDecision` already catches exactly
that shape, an act with no allowlisted fn or a keyless fact, and it keeps doing
so. The head replaces the hardcode. It does not replace the gate.
## An incident worth recording
The first generation run produced zero rows for eight destinations. `mavgpud`
yields the card when another process wants it (V-488) and llama-server answers
503 until the model is back. Every generate call inside that window burned one of
the destination's batches. The run walked its own cap without a single successful
call. The log said `503` 260 times, and the summary line said 64.1% keep rate,
which read as success.
`call()` now retries a 503 with backoff. A generator that treats an unloaded
model as a bad generation is a silent-corpus bug, not a slow one.
## What this does not measure
**Nothing here runs in Go.** The heads are a `heads.pt` and an
`out/body_heads/` on workpc. Reaching the daemon needs an ONNX export and a
caller. The resident e5-small must not be replaced by this copy: recall depends
on that file, and `EmbedQuery`/`EmbedPassage` are its contract.
The destination fixture carries 4 of 13 classes: recall 15, floor 7, calendar 6,
world 5. `tasks`, `money`, `list`, `home`, `network`, `feeds`, `attention` and
`self` have no gold case. So 78.8% is silent on eight destinations that together
hold 800 training rows.
Latency was not measured. A forward pass of a 118M encoder should beat a 1.7B
decoder on a query turn. That is arithmetic, not a number from this box.
+22
View File
@@ -0,0 +1,22 @@
package router
import (
"os"
"testing"
)
// TestDumpPrompt writes the router prompt and grammar to disk so the training
// workspace labels with the daemon's own contract rather than a retyped copy.
// It is inert unless MAVEN_DUMP_PROMPT names a directory.
func TestDumpPrompt(t *testing.T) {
dir := os.Getenv("MAVEN_DUMP_PROMPT")
if dir == "" {
t.Skip("MAVEN_DUMP_PROMPT unset")
}
if err := os.WriteFile(dir+"/route_system.txt", []byte(routeSystem), 0o644); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(dir+"/route_grammar.gbnf", []byte(routeGrammar), 0o644); err != nil {
t.Fatal(err)
}
}
+42 -1
View File
@@ -45,12 +45,19 @@ const routeGrammar = `
root ::= "[" ws action ("," ws action)* ws "]"
action ::= "{" ws "\"intent\"" ws ":" ws intent ("," ws field)* ws "}"
intent ::= "\"fact\"" | "\"reminder\"" | "\"note\"" | "\"query\"" | "\"act\"" | "\"chat\"" | "\"system\"" | "\"unknown\""
field ::= key ws ":" ws string
field ::= (key ws ":" ws string) | ("\"source\"" ws ":" ws source)
key ::= "\"key\"" | "\"value\"" | "\"text\"" | "\"verb\""
source ::= "\"recall\"" | "\"calendar\"" | "\"tasks\"" | "\"list\"" | "\"money\"" | "\"weather\"" | "\"home\"" | "\"network\"" | "\"feeds\"" | "\"attention\"" | "\"self\"" | "\"world\"" | "\"\""
string ::= "\"" ([^"\\\x00-\x1F] | "\\" ["\\/bfnrt] | "\\u" [0-9a-fA-F]{4}){0,120} "\""
ws ::= [ \t\n]{0,4}
`
// TestRouteGrammarCoversSources holds the source rule above to router.Sources.
// The enum is the point: a grammar cannot emit a destination that does not
// exist, which is the guarantee V-546 wants from a softmax and gets here for
// free. Empty is the thirteenth alternative and it is not an oversight — it is
// the SourceUnknown floor, and the model must be able to decline.
// routeSystem — the router prompt. Changed 31-07-2026: the query test now sits
// above the fact test and there is an explicit question test. Before that, a
// question naming a fact key ("сколько воды я выпил с утра") matched the fact
@@ -121,6 +128,31 @@ const routeSystem = `Классифицируй ровно одно сообще
"что такое кватернион?" → {"intent":"query","text":"что такое кватернион"}
"ага" → {"intent":"chat","text":"ага"}
Только для query добавь поле source — где лежит ответ:
- recall — его заметки, факты и то, что он раньше говорил
- calendar — встречи и события
- tasks — список задач
- list — списки покупок и другие именованные списки
- money — траты
- weather — погода
- home — свет, устройства, дом
- network — локальная сеть, сервер, диски
- feeds — новостные ленты
- attention — что требует внимания сейчас
- self — вопрос про самого ассистента
- world — всё остальное: определения, счёт, люди, факты о мире
Пустое значение "" — нормальный ответ и его надо ставить часто. Ставь "", если ответ могут дать сразу два источника или если не уверен: тогда проверяются все по порядку, и это правильно. Никогда не угадывай.
"сколько воды я выпил с утра" → {"intent":"query","text":"сколько воды я выпил с утра","source":"recall"}
"что я записывал про кота" → {"intent":"query","text":"что я записывал про кота","source":"recall"}
"во сколько у меня встреча" → {"intent":"query","text":"во сколько у меня встреча","source":"calendar"}
"что такое docker?" → {"intent":"query","text":"что такое docker","source":"world"}
"кто такой Линус Торвальдс?" → {"intent":"query","text":"кто такой Линус Торвальдс","source":"world"}
"сколько будет 17 на 23?" → {"intent":"query","text":"сколько будет 17 на 23","source":"world"}
"почему сервер тормозит" → {"intent":"query","text":"почему сервер тормозит","source":""}
"есть новости по бэкапу базы" → {"intent":"query","text":"есть новости по бэкапу базы","source":""}
Ответ — JSON-массив: по одному объекту на каждую просьбу. Обычно один. Если в реплике несколько просьб — по объекту на каждую. "напомни купить молоко, и запиши что кофе кончился" → [{"intent":"reminder","text":"купить молоко"},{"intent":"note","text":"кофе кончился"}]. Только JSON, без пояснений.`
// routeRepeatPenalty — the sub-1B model loops one sentence inside the text field
@@ -172,6 +204,7 @@ type routeAction struct {
Value string `json:"value"`
Text string `json:"text"`
Verb string `json:"verb"`
Source string `json:"source"`
}
// Route asks the model for one decision. The bool is false when there is no
@@ -240,6 +273,14 @@ func (lr *LLMRouter) Route(ctx context.Context, utterance string, now time.Time)
case IntentQuery:
d.Intent = IntentQuery
d.Slots.Text = firstNonEmpty(a.Text, utterance)
// Through ValidSource, and on query alone. The grammar already bounds
// the enum, but the grammar is a request to a server that may be
// running a different build, and a destination this binary does not
// know would take real query sources off the turn. Anything unknown
// drops to SourceUnknown, which is the floor and costs nothing.
if ValidSource(Source(a.Source)) {
d.Source = Source(a.Source)
}
case IntentAct:
d.Intent = IntentAct
d.Slots.Text = firstNonEmpty(a.Verb, utterance)
+67
View File
@@ -398,3 +398,70 @@ func TestLLMReminderWithSubjectIsNotGated(t *testing.T) {
t.Fatalf("a complete reminder was sent back as a question: %+v", d.Slots)
}
}
// TestRouteGrammarCoversSources — the grammar enum and router.Sources are two
// hand-written lists of the same twelve destinations, and nothing else notices
// when one grows. A destination missing from the grammar is a destination the
// model is structurally unable to name, which is the exact defect V-517
// measured for Praxis: not a weak model, an absent string.
func TestRouteGrammarCoversSources(t *testing.T) {
for _, s := range Sources {
if !strings.Contains(routeGrammar, `"\"`+string(s)+`\""`) {
t.Errorf("routeGrammar cannot emit %q — the model can never name it", s)
}
}
// The floor has to be reachable too, or the model is forced to pick one.
if !strings.Contains(routeGrammar, `"\"\""`) {
t.Error(`routeGrammar cannot emit "" — the model cannot decline a destination`)
}
// Count the alternatives on the source rule: an extra one is a destination
// the daemon would drop to SourceUnknown after the model spent tokens on it.
for _, line := range strings.Split(routeGrammar, "\n") {
if !strings.HasPrefix(line, "source ") {
continue
}
if got, want := strings.Count(line, "|")+1, len(Sources)+1; got != want {
t.Errorf("source rule has %d alternatives, want %d (Sources plus the floor)", got, want)
}
}
}
// The destination is read back only through ValidSource. A model on an older or
// newer build can write a string this binary does not know, and trusting it
// would take real query sources off the turn for a name nothing answers.
func TestLLMUnknownSourceFallsToTheFloor(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"query","text":"что там с бэкапами","source":"praxis"}`)
d, err := r.Route(context.Background(), "что там с бэкапами", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if d.Source != SourceUnknown {
t.Fatalf("invented destination %q was trusted, want the floor", d.Source)
}
}
// And a known one survives, or the read-back is just a filter.
func TestLLMNamedSourceSurvives(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"query","text":"кто такой Линус Торвальдс","source":"world"}`)
d, err := r.Route(context.Background(), "кто такой Линус Торвальдс?", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if d.Source != SourceWorld {
t.Fatalf("source %q, want %q", d.Source, SourceWorld)
}
}
// A destination on anything but a query is dropped. Only IntentQuery reaches
// queryWalk, so a source elsewhere is a field nobody reads and a claim nobody
// checks.
func TestLLMSourceIsQueryOnly(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"note","text":"кофе кончился","source":"recall"}`)
d, err := r.Route(context.Background(), "запиши что кофе кончился", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if d.Source != SourceUnknown {
t.Fatalf("a note carried destination %q", d.Source)
}
}