Compare commits

..

21 Commits

Author SHA1 Message Date
claude ee9d55ca95 Measure what the two clarify bounds bought (V-663)
Tail turns 21 to 17, the longest ride 8 turns to 4, and the two worst
replies in the corpus are gone: "спасибо" and "привет" are no longer
answered with "Сейчас 21:25. В какой день?".

MaxRides is not what fired. With the pleasantry counted as an aside the run
of asides is unbroken, so MaxSuspends reached three and ended it. Rides is
the backstop for the shape where an answer really does break the run, and
no turn in this corpus reaches it. Said so rather than crediting the new
bound.

Four rides did not move. They are asides against a question the owner never
answers, which MaxSuspends already bounds at four turns each.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-08 21:49:32 +04:00
claude a886217223 A greeting is not a failed answer (V-663)
classifyTurnRole read "спасибо" and "привет" as answers to whatever was
parked, so she re-asked "В какой день?" at a man saying thank you and
spent one of three attempts doing it. That attempt is a bound meant to end
the ride, so the pleasantry both produced the worst reply in the corpus and
paid for the privilege.

They are asides now: answered as themselves, the question resumed on the
tail, no attempt spent, one ride counted.

The set is a new closed lexicon entry, matched as WHOLE utterances. Every
token rule tried was wrong on something. "вечер" answers "это утра или
вечера?" and "нет" answers a confirm, so anything that could fill a slot
stays out. The control words stay out too, because isCancel owns them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-08 21:43:15 +04:00
claude de9884e063 Count the rides a question takes, without the reset (V-663)
MaxSuspends did not move the number it was written for. Twenty-six of 140
turns carried a parked clarify tail before it landed and twenty-six after.

Two bounds rearm each other. An aside spends no attempt, so MaxAttempts
never reaches it. A turn reading as a failed answer zeroes Suspends, so
MaxSuspends never reaches the asides. Alternating them restores each bound
with the other's traffic. Measured on 2026-08-08: one question about a
reminder's day rode turns 7 to 13.

PendingQuestion.Rides is the same event counted without the resets. Set
once, incremented only in noteSuspended, carried across the re-park in
askRemainingGap, read by nothing that could lower it. MaxRides is 4, one
looser than MaxSuspends so the tighter statement about a run stays
reachable.

It ends the measured ride one turn early and no more. Most of that ride is
attempts, spent because classifyTurnRole reads "спасибо" and "привет" as
failed answers. Said so in the constant and in the design doc rather than
claiming a fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-08 21:38:32 +04:00
claude bbefda66e2 Read the source column off the badge, not off the wording (V-662)
The third run of the same 140 turns, with the harness fix in. Sixty-eight
turns name a source.

Two findings the wording could not carry. The unfixed homelab turns are
claimed by weather and by feeds, which the destination fixture predicted.
And agenda questions are claimed by the personal boundary and by Praxis,
not by the calendar: 3 of 6, the same 3 of 6 the destination fixture and
every routing-head seed score.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-08 21:31:22 +04:00
claude a37c4138a1 Read the source badge under the name the server writes (V-662)
scripts/usage-run.py read the redirect parameter "src". cmd/mavweb/chat.go
writes it as "s". So Source came back empty on all 140 turns of both
fortnight runs, and every finding in those two docs is read off the reply
wording instead of off the badge.

Re-run confirms the column now arrives: 68 of 140 turns name a source.
The two homelab misses are now direct evidence rather than inference.
"какая скорость у меня сейчас?" is claimed by weather and
"хватает ли места под новые бэкапы?" by feeds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-08 21:30:54 +04:00
kami d6f391430f Merge pull request 'Re-run the fortnight against merged master' (#204) from task/661-post-merge-usage-rerun into master 2026-08-08 19:15:51 +02:00
claude 68b2aa9137 Re-run the fortnight against merged master (V-661)
V-655 fixed four of the six turns a guessing query source claimed. The
clean win is 'что такое TCP?', which stage 0 names world and search now
answers instead of weather asking for a city. 'что я сохранил про Сочи?'
is no longer read as a capture.

The two that did not move are both homelab questions, which is the enum
and not the walk. SourceRecall, SourceNetwork and SourceAttention overlap
on every question about the box, and the destination fixture already
flagged that cluster.

The parked clarify is unchanged at 26 turns. It is dialogue state and no
query source could have touched it.

Latency is reported and not attributed. The workstation was up for the
re-run and its state during the baseline was never recorded.

Also corrects the baseline's '49 of 140 turns'. That was the sum of
occurrences. It is 41 turns.
2026-08-08 21:14:52 +04:00
kami f8fa0d1b44 Merge pull request 'Routing heads: a slot head, a clarify head, and a two-week baseline to diff against' (#203) from task/661-routing-heads-step-3-train-the-multi-hea into master 2026-08-08 19:06:55 +02:00
claude 9a333b23d7 Merge master after 199-201 landed (V-661) 2026-08-08 21:05:27 +04:00
kami 663b5c47b9 Merge pull request 'The router prompt has no destination, so the model arm of V-655 names nothing' (#201) from task/660-router-prompt-destination into master 2026-08-08 19:03:28 +02:00
kami 45c521e1a6 Merge pull request 'Destination fixture: score Decision.Source, not just the intent' (#200) from task/659-destination-fixture into master 2026-08-08 19:03:24 +02:00
kami e34669a52e Merge pull request 'Query source is a routing decision made outside the router' (#199) from task/655-query-source-is-a-routing-decision-made into master 2026-08-08 19:03:06 +02:00
claude d434f83c2c The personal boundary is a guesser, so say so (V-655)
CLAUDE.md said naming SourceWorld leaves his notes, his facts and the
personal boundary running first. The first two are true and the third is
not. The boundary is marked guesses: true, so queryWalk drops it whenever
the named destination is not recall.

That is deliberate and tested. It is what stops the boundary answering
'кто такой Линус Торвальдс?' with 'не нашла у тебя такой записи'. But it
means a destination a model wrote can take the boundary off a turn about
him, and the doc claimed the opposite.

Flagged as the owner's call rather than changed. Only the utterance leaves
the box either way.
2026-08-08 21:02:41 +04:00
claude c310115fd2 Record the clarify head and the confidence it replaces (V-661) 2026-08-08 20:59:44 +04:00
claude 3024f76e5f A fourth head asks instead of guessing (V-661)
Clarify is not a value of intent. It is a second question over the same
pooled vector: can Maven act on this at all. The eight want_clarify fixture
cases sat outside every number the heads measured, because a softmax has no
clarify class.

gen_clarify.py makes the class the corpus lacks. Every existing row was
generated FOR an intent, so every one is answerable. The router-prompt
agreement filter cannot work here, because routeGrammar has no clarify value
and a generated line always agrees with itself. A judge replaces it.

The first judge called 24 of 40 answerable rows underspecified. It judged
against a generic assistant, one that asks where about lunch. Restating
Maven's contract took that to 16 of 60, with all eight fixture cases caught.

Three seeds: 7.0 of 8 caught, 2.3 false of 88. The cascade today misses 1 and
produces 2. Confidence separates too, 0.851 right against 0.604 wrong.
2026-08-08 20:59:21 +04:00
claude 6bc71553ab Say that a transport error is not a wrong answer (V-661) 2026-08-08 20:47:32 +04:00
claude ed1730431c Distil a slot head and record it beside the other two (V-661)
BIO tags had no Maven-domain corpus, which was true of found corpora and
false of made ones. A GBNF closed over Maven's five slots plus a
substring check gives 2178 spans out of gemma-4-12b at no second call.

Three heads over one forward pass: intent 92.8%, destination 82.8%, slot
span F1 72.4% over three seeds. The slot head is free.

Epoch selection reads the intent dev slice, so it stops the slot head
about 4 points early. Recorded rather than fixed.
2026-08-08 20:27:58 +04:00
claude 01e80fce4a Record a fortnight of usage as a re-runnable baseline (V-661)
The 2026-08-07 week of usage was typed by hand and cannot be replayed, so
it measured a build and not a change. scripts/usage-run.py drives the same
reach from a turns file, which makes the next run a diff.

Baseline is master at beb093a: 140 turns, p50 1.6s, zero errors. Three
defects to move. A parked reminder clarify contaminates 19 later turns and
survives a day boundary. Query sources that guess claim six turns they
cannot answer, which is the class V-655 removes. And one question was read
as a capture.

Also records the slot head: gemma distils 2178 spans, three heads score
intent 92.8%, destination 82.8%, slot span F1 72.4% over three seeds.
2026-08-08 20:27:27 +04:00
claude e69f1bd0cf Fix the floor corpus and re-measure the destination head (V-661)
The first 120 floor rows carried one sentence shape, because the generator
varies a topic and ambiguity is not a topic. Rotating six shapes takes the
floor 3/7 to 6/7 and the destination mean 75.8% to 80.8%.

Calendar stays 3/6 at every seed. The possessive agenda rules claim those
cases at stage 0 and name nothing, so no label reaches the head.

Also corrects the floor-case count in three files: five of the seven are
homelab, not six.
2026-08-08 20:10:14 +04:00
claude f55bedee2e Train the destination head and beat the teacher (V-661)
Step 3 of the routing-heads plan. Intent and destination share one masked mean
pool on e5-small. Destination scores 26/33 against 12/33 for the classifier
cascade and 24/33 for the cascade with gemma-4-12b, which is the teacher these
labels were distilled from. Recall goes 0/15 to 15/15.

Two heads, not four, and both cuts are label problems rather than GPU time.
Mood describes her own reply state and no dataset maps onto it. BIO slot tags
have no Maven-domain corpus.

The MASSIVE warm-start from step 2 is worth nothing here either. Stock ties it
on intent and leads by a third of a case on destination.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-08 19:38:23 +04:00
claude e470435cf1 Dump the router prompt where the labeler can read it (V-661)
The training workspace labels with routeSystem and routeGrammar, and it held
its own copies. V-660 changed both. A retyped prompt drifts silently, which is
the problem llm/check_prompt_parity.py exists for on the other side.

Inert unless MAVEN_DUMP_PROMPT names a directory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-08 18:44:47 +04:00
20 changed files with 2599 additions and 9 deletions
+85 -4
View File
@@ -246,6 +246,72 @@ was a hardcode. **Fine-tune a copy of the weights.** The resident embedder backs
recall. Training it in place couples routing accuracy to recall@1, with nothing in the
suite to name the trade.
**Two of those heads are trained as of 08-08-2026, and they are not the three
above** (V-661, `docs/evals/2026-08-08-routing-heads-two-head.md`). Intent and
destination share one masked mean pool. Destination scores a mean **80.8%** over
three seeds, best **29/33 (87.9%)**. The classifier cascade scores 12/33 and the
cascade with gemma-4-12b scores 24/33, so a 118M encoder beats the 12B teacher it
was distilled from. Read the best run as one seed and not a headline, because one
case is 3 points on a fixture this small.
Recall is 15/15 and world is 5/5. Intent is 93.6% mean over three seeds. That is
**not** comparable to the 76.0% and 84.4% those two arms scored: a softmax has no
clarify class, so the head's fixture is the 88 cases carrying an intent.
**A fourth head asks instead of guessing, same day** (V-661,
`docs/evals/2026-08-08-clarify-head-four-head.md`). Clarify is not a value
of intent, so a softmax cannot emit it. It is a second question over the
same pooled vector: can Maven act on this at all. That is why the head's
fixture was 88 cases and not 96. Over three seeds it catches **7.0 of the 8
`want_clarify` cases and produces 2.3 false clarifies of 88**. The cascade
today misses 1 and produces 2, so this is parity with no rules in front of
it. Accuracy is the wrong number here and a head that never asks scores
91.7%. Confidence is the other half. Max softmax over the intent head reads
**0.851 where it is right against 0.604 where it is wrong**, ranking right
above wrong in 83.4% of pairs. `Confidence: 1.0` was a hardcode, and this
replaces it with a signal. The two are not the same signal: one says which
intent is unclear, the other says the utterance carries too little to act
on. **The fourth head is not free the way the third was.** Intent,
destination and slot F1 each move down one to four points, inside the seed
spread. `поужинал` is a false clarify on every seed, which is the same
defect `thinSingleToken` was narrowed for on 2026-08-01.
The corpus for it is generated, because every existing row is answerable by
construction. **The router-prompt agreement filter cannot work here**, since
`routeGrammar` has no clarify value and a generated line always agrees with
itself. A gemma judge replaces it. The first judge called 24 of 40
answerable rows underspecified, because it judged against a generic
assistant rather than against Maven's contract.
**Mood is cut, not deferred.** The enum describes her own reply state, not the
speaker's emotion, and no dataset maps onto it.
**A third head landed the same day** (`docs/evals/2026-08-08-slot-head-three-head.md`).
BIO slot tags had no Maven-domain corpus, which was true of found corpora and
false of made ones. `label_slots.py` distils spans out of gemma-4-12b under a
GBNF closed over Maven's own five slots. A span survives only when it is a
literal substring of the utterance, so the agreement filter costs no second
call. 2178 spans over 1702 rows, 37 dropped, nothing unparsed. Three heads score
intent **92.8%**, destination **82.8%** and slot span F1 **72.4%** over three
seeds. The slot head is free: both other numbers move less than their own seed
spread. Epoch selection reads the intent dev slice alone. Slot F1 is still
climbing when it stops, which costs about 4 points.
The MASSIVE warm-start of step 2 is worth nothing here. Stock e5-small ties it on
intent and leads by a third of a case on destination. Nothing argues for keeping
that step.
The floor was a corpus defect and it is fixed. The first 120 floor rows carried
one sentence shape, so the head named a destination where the fixture says walk
the chain. Rotating six shapes took the floor 3/7 to 6/7 and destination 75.8% to
80.8%. What is left is calendar at 3/6 on every seed, which training cannot move:
the possessive agenda rules claim those cases at stage 0 and name nothing, so no
label reaches the head. That is the same trade V-660 flagged and it wants the
owner's call.
**Nothing of this runs in Go.** The weights are `heads.pt` and `out/body_heads/`
on workpc. Reaching the daemon needs an ONNX export and a caller. The resident
e5-small must not be replaced by the copy, because recall depends on that file.
`Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
#359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with
@@ -376,7 +442,19 @@ queries exactly as it did.
That is the safety argument and it is not negotiable. The table's order is
load-bearing. Every comment on it argues a reason between two sources, and above all
it carries "the owner's data first, then the world". Naming `SourceWorld` does not
send the turn outside. His notes, his facts and the personal boundary still run first.
send the turn outside on its own. His notes and his facts still run first, because
they look rather than guess.
**The personal boundary is the one exception and it is deliberate.** It guesses,
so naming `SourceWorld` drops it. That is what stops it answering "кто такой
Линус Торвальдс?" with "не нашла у тебя такой записи", which it did on
2026-08-07. The cost is that a destination a model wrote can now take the
boundary off a turn. A question about him that the model calls `world` reaches
SearXNG, where today the boundary stops it. Only the utterance leaves the box,
never his notes or history, so this widens what is asked and not what is sent.
`TestNamingRecallKeepsTheBoundary` pins the other half: naming `SourceRecall`
keeps the boundary in front of the world. Whether a model may drop it at all is
the owner's call and has not been made.
What comes out is only the sources that **guess**. Those decide a turn is theirs by
cosine against frozen seeds, then answer whatever they claimed. They hold no table
@@ -422,11 +500,14 @@ rules deliberately do not name it. And **recall is 0/15, because nothing
anywhere names it**. Those turns are still answered, since the chain walks
recall early. Recall is the number the fourth head has to move.
Seven cases assert the floor and six of them are homelab operations. They
Seven cases assert the floor and five of them are homelab operations. They
cluster because `SourceRecall`, `SourceNetwork` and `SourceAttention` overlap on
every question about the box. `mavpoll` writes its netdata and uptime-kuma
every question about the box. The other two are `ru-query-005` and
`ru-query-014`. No query source reads the reminder store, and a deadline could
sit in tasks, the calendar or Praxis. `mavpoll` writes its netdata and uptime-kuma
observations into the fact store recall reads. That is a finding about the enum,
not a gap in the labelling.
not a gap in the labelling. The owner confirmed all seven floor labels on
08-08-2026, so they are a decision rather than an agent's guess.
`baselineGrammars` in `eval_test.go` mirrors `buildRouter` and had drifted:
`WorldQueryGrammars` was wired into the daemon by V-655 and not into the mirror,
+12 -2
View File
@@ -489,15 +489,19 @@ func (h *reactiveHandler) noteSuspended(ctx context.Context, q *dialogue.Pending
if !q.CanResume() {
h.clarifyStore.Delete(dialogueIDOf(ctx))
h.noteDropped(ctx)
log.Printf("voice: clarify — the question about %s stepped aside %d times; letting the request go", q.Missing[0], q.Suspends)
log.Printf("voice: clarify — letting the question about %s go: %d asides in a row, %d rides in all", q.Missing[0], q.Suspends, q.Rides)
return
}
q.Suspends++
// Rides is the same event counted without the reset (V-663). Incremented
// beside Suspends and never anywhere else, so the two cannot disagree about
// what happened, only about how much of it they remember.
q.Rides++
q.Asked = h.now()
h.clarifyStore.Put(dialogueIDOf(ctx), q)
rt.resume = question
rt.suspended = true
log.Printf("voice: clarify — is its own request; suspending the question about %s and resuming it in the same reply (suspend %d of %d)", q.Missing[0], q.Suspends, dialogue.MaxSuspends)
log.Printf("voice: clarify — is its own request; suspending the question about %s and resuming it in the same reply (suspend %d of %d, ride %d of %d)", q.Missing[0], q.Suspends, dialogue.MaxSuspends, q.Rides, dialogue.MaxRides)
}
// foldAnswerIntoUtterance appends an answered subject to the original words,
@@ -540,6 +544,11 @@ func (h *reactiveHandler) askRemainingGap(ctx context.Context, q *dialogue.Pendi
// Suspends is not carried, and by this point it is already zero: the answer
// path resets it (V-654). Left off the literal so the zero is stated where
// the struct is built, rather than inherited from a field nobody names.
//
// Rides IS carried, and that is the whole point of it (V-663). This is the
// same request under a second question, not a new one, so the turns it has
// already ridden still count against it. Dropping the field here is exactly
// the re-basing that let one question ride twenty-six replies.
h.clarifyStore.Put(dialogueIDOf(ctx), &dialogue.PendingQuestion{
Intent: q.Intent,
Slots: merged,
@@ -550,6 +559,7 @@ func (h *reactiveHandler) askRemainingGap(ctx context.Context, q *dialogue.Pendi
TTL: clarifyTTL,
Attempts: q.Attempts + 1,
MaxAttempts: q.MaxAttempts,
Rides: q.Rides,
})
log.Printf("voice: clarify — one gap filled, still missing %s for intent=%s, asking again (attempt %d)", remaining[0], intent, q.Attempts+1)
return question, true
+32
View File
@@ -186,6 +186,28 @@ func carriesReminderVerb(text string) bool {
return false
}
// isPleasantry matches the WHOLE utterance against lexicon.Pleasantries, after
// lowercasing and dropping the punctuation a greeting carries.
//
// Whole utterance and not tokens. Every token rule tried here was wrong on
// something: "вечер" answers "это утра или вечера?", "нет" answers a confirm,
// and "спокойной" alone is not an utterance at all. A greeting is a fixed
// phrase, so matching it as one costs nothing and claims nothing else.
func isPleasantry(text string) bool {
t := strings.ToLower(strings.TrimSpace(text))
t = strings.Trim(t, " .,!?…")
t = strings.Join(strings.Fields(t), " ")
if t == "" {
return false
}
for _, p := range lexicon.Pleasantries() {
if t == p {
return true
}
}
return false
}
// offlineOwnRequest is the shape half of the evidence: the offline token tests,
// which cost nothing and never depend on the model that produced the routing.
// It is also the whole answer when there is no route to read — the classifier
@@ -225,6 +247,16 @@ func classifyTurnRole(q *dialogue.PendingQuestion, text string, answer dialogue.
// hour, and no route saying "question" changes that. It works because the
// extractor no longer reads a day word as the current clock, so a sentence
// that names no hour now fills nothing to weigh.
// A pleasantry is neither (V-663). "спасибо" and "привет" fell through to
// roleAnswer, so a question about a reminder's DAY was re-asked at a man
// saying thank you, and the retry it spent was one of the three bounds
// meant to end the ride. It is an aside: answered as itself, the question
// resumed on the tail, no attempt spent, one ride counted. Placed above the
// content gate because "доброе утро" has content and states nothing, so
// neither half of the evidence below can reach it.
if q != nil && isPleasantry(text) {
return roleAside
}
own := false
if len(ownContent(text)) > 0 {
own = offlineOwnRequest(text) || (ok && carriesOwnRequest(routed, text))
+65
View File
@@ -372,4 +372,69 @@ func TestAnAnsweredGapResetsTheSuspendBudget(t *testing.T) {
if q.Suspends != 0 {
t.Fatalf("answering a gap must reset the suspend budget: suspends = %d", q.Suspends)
}
// The ride it already took is carried across the re-park (V-663). Resetting
// both counters here is what let one question ride twenty-six replies.
if q.Rides != 1 {
t.Fatalf("the aside it already took was forgotten: rides = %d", q.Rides)
}
}
// TestTwoBoundsCannotRearmEachOther — V-663.
//
// MaxSuspends landed and the measurement did not move: twenty-six of 140 turns
// carried a tail before it and twenty-six after. This is the shape it misses,
// taken from the 2026-08-08 run, where one question rode turns 7 to 13.
//
// An aside spends no attempt, so MaxAttempts never reaches it. A turn that
// reads as a failed answer zeroes Suspends, so MaxSuspends never reaches the
// asides either. Alternating the two rearms each bound with the other's
// traffic. Rides counts both kinds and is never reset, so it is what ends this.
func TestTwoBoundsCannotRearmEachOther(t *testing.T) {
ctx := context.Background()
h, _ := newRoutingClarifyHandler(t)
id := dialogueIDFor(sourceText, "web")
resumed, _ := clarifyResumedFor(dialogue.SlotTime)
if reply := h.handleText(ctx, "web", "напомни позвонить маме"); !strings.Contains(reply, "?") {
t.Fatalf("expected the time question, got %q", reply)
}
// Two asides. Each one rides and neither spends an attempt.
for i := 0; i < 2; i++ {
reply := h.handleText(ctx, "web", "какие у меня напоминания?")
if !strings.HasSuffix(reply, resumed) {
t.Fatalf("aside %d: the question must come back, got %q", i+1, reply)
}
}
q := h.clarifyStore.Get(id, h.now())
if q == nil || q.Rides != 2 || q.Suspends != 2 {
t.Fatalf("after two asides: %+v", q)
}
// A pleasantry. It used to read as a failed answer, so she re-asked the
// question at a man saying thank you and spent an attempt doing it. Now it
// is an aside: answered as itself, question on the tail, one more ride.
reply := h.handleText(ctx, "web", "спасибо")
if !strings.HasSuffix(reply, resumed) {
t.Fatalf("a pleasantry lost the parked question: %q", reply)
}
q = h.clarifyStore.Get(id, h.now())
if q == nil || q.Attempts != 1 {
t.Fatalf("a pleasantry spent an attempt: %+v", q)
}
if q.Rides != 3 {
t.Fatalf("a pleasantry rode free: %+v", q)
}
// One more ride of any kind and the request goes, out loud.
reply = h.handleText(ctx, "web", "какие у меня напоминания?")
if !strings.Contains(reply, clarifyDropped) {
t.Fatalf("the question rode four asides and was let go in silence: %q", reply)
}
if strings.HasSuffix(reply, resumed) {
t.Fatalf("a question she has let go must not be asked again: %q", reply)
}
if h.clarifyStore.Get(id, h.now()) != nil {
t.Fatal("the question must be gone once she has said she let it go")
}
}
+23
View File
@@ -308,6 +308,29 @@ The count is of CONSECUTIVE step-asides. It resets the moment he answers, in
too. "Позвонить маме" against a question about the time is still him in the
exchange. The retry it costs is bound enough on its own.
#### And it may ride four turns in all
Decided 2026-08-08 (V-663), because the bound above did not move the number it
was written for. Twenty-six of 140 turns carried a tail before it landed and
twenty-six carried one after.
Two bounds rearm each other. An aside spends no attempt, so `MaxAttempts` never
reaches it. A turn that reads as a failed answer zeroes `Suspends`, so
`MaxSuspends` never reaches the asides. Alternating them, each bound is restored
by the other's traffic. Measured on 2026-08-08: one question about a reminder's
day rode turns 7 to 13. It ended only because turn 14 was a new request.
`PendingQuestion.Rides` counts the same event as `Suspends` with the resets
taken out. It is set once, incremented only in `noteSuspended`, carried across
the re-park in `askRemainingGap`, and read by nothing that could lower it.
`MaxRides` is 4, one looser than `MaxSuspends` so that the tighter statement
about a run stays reachable.
This is a bound, not a cure. It ends the measured ride one turn early. Most of
that ride's length is attempts, spent because `classifyTurnRole` reads "спасибо"
and "привет" as failed answers to a question about a day. That is the next
thing to fix and it is not a bound.
The re-ask is also two sentences rather than one. It used to be spliced onto the
answer with a comma. On a real answer that buries the question in the tail of
one run-on thought:
@@ -0,0 +1,123 @@
# A clarify head, and a confidence that is not a hardcode
Measured 2026-08-08 on workpc, the same day and the same fixtures as
`2026-08-08-routing-heads-two-head.md` and `2026-08-08-slot-head-three-head.md`.
New script `gen_clarify.py`, new fixture `eval_fixture_clarify.jsonl`.
## A softmax has no clarify class
That sentence closed the two-head measurement. It is why the head's fixture was
88 cases and not 96. The eight `want_clarify` cases sat outside every number
measured, and the head had no way to produce the answer they wanted.
A fourth head is the answer. Clarify is not a value of intent. It is a second
question asked of the same pooled vector: can Maven act on this at all.
## The corpus had one class
Every row in `train_heads_slots.jsonl` was generated FOR an intent or a
destination. So every row is answerable by construction. A head trained on that
alone sees one class and learns to say yes.
`gen_clarify.py` makes the other class. Five shapes, ten topics. The shapes are
the gate's own reasons in `gateLLMDecision` plus the two the fixture carries:
bare noun, bare verb, demonstrative, deictic time, dangling reference.
**The agreement filter that worked for destination cannot work here.**
`routeGrammar` has no clarify value. So the router always names an intent, and
any generated line always agrees with itself. The second pass is a judge
instead. Gemma is asked, without seeing the label, whether Maven would have to
ask a question back.
## The first judge was worthless and the second was measured
The first judge said "needs clarify" on 24 of 40 plainly answerable corpus
rows. It flagged `запиши что я пообедал` and `Покажи расписание поездов на
вечер`. It was judging against a generic assistant, one that asks "where?"
about lunch. Maven writes that note.
Rewriting it to state what she can already do took false positives to 16 of 60.
It also catches all eight fixture clarifies. So the judge discriminates.
On the generated pile it removed 7 of 306, a 97.7% keep rate. That is not the
judge failing. The generator is aimed at underspecified lines, so there is
little for a filter to catch. The 27% false-positive rate is the number to
quote, and it is label noise on the positive class.
**Both passes are gemma-4-12b.** Generation and judging. So the corpus is
gemma's opinion of what is underspecified, and the head distills that opinion.
What keeps it honest is the fixture. Those eight cases were written by the owner
and gemma never saw them.
299 rows kept, against 3604 answerable. The positive class carries `intent:
null`, so it costs the intent head nothing.
## Result
Three seeds, 24 epochs, epoch still chosen on the intent dev slice.
| | two heads | three heads | four heads |
|---|---|---|---|
| intent mean | 93.6% | 92.8% | 91.7% |
| destination mean | 80.8% | 82.8% | 79.8% |
| slot span F1 mean | — | 72.4% | 68.3% |
| clarify caught | — | — | 7.0 of 8 |
| false clarifies | — | — | 2.3 of 88 |
**The fourth head is not free the way the third was.** Intent, destination and
slot F1 all move down. The drop is one to four points, and the seed spread is
wide enough to contain it. Seed 2 scores intent 94.3% and destination 84.8%, both above
every three-head seed. Read the drop as unproven rather than as absent.
Accuracy is the wrong number for this head and is reported for completeness at
95.8% to 96.9%. Eight of ninety-six cases are positive, so a head that never
asks scores 91.7%. Recall on those eight is the number.
Compare it to what ships. The cascade today misses 1 clarify and produces 2
false ones. The head catches 7 of 8 and produces 2.3 false ones. That is
parity, from a 118M encoder with no rules in front of it.
The saved checkpoint is seed 2 at epoch 10. Intent 94.3%, destination 28/33,
slot F1 73.6%, clarify 7 of 8 with 3 false. `heads.pt` carries four state dicts.
## What it gets wrong is consistent across seeds
`поужинал` is a false clarify on all three seeds. That utterance is already
recorded as a real defect. `thinSingleToken` was narrowed on 2026-08-01 to spare
a token carrying a Russian verb ending. One word is routinely a whole sentence
in Russian. The head relearned the mistake the rule was narrowed to
fix.
`что дальше?` is a false clarify on two seeds. That one is a disagreement rather
than an error. The utterance is underspecified, and V-498 decided stage 0 claims
it for the calendar on purpose.
`ну это` is missed on two seeds. `amb-003` is the shortest case in the fixture
and the generated demonstratives are longer.
## Confidence
`Confidence: 1.0` was a hardcode in `llmrouter.go`, so a correct low confidence
could not exist. Max softmax over the intent head is the replacement. It is only worth reading
if it is lower where the head is wrong.
It is. Mean 0.851 where the head is right against 0.604 where it is wrong. It
ranks a right case above a wrong one in 83.4% of pairs.
So there are two signals now and they are not the same signal. Confidence says
the head is unsure which intent this is. The clarify head says the utterance
does not carry enough to act on. A confident wrong route and an honest "I cannot
tell" are different failures, and one number cannot report both.
## What this does not measure
The same gap as every head run. **Nothing of this runs in Go.** Four heads
instead of three does not change that.
There is no threshold. Both signals are reported as raw numbers. Turning either
into a gate needs a decision about where to cut, and that trades false clarifies
against wrong acts. The fixture has 8 positives, which is too few to fit a
threshold on.
The 299 generated rows have no held-out slice of their own. Clarify is scored on
the fixture alone.
+4 -2
View File
@@ -46,9 +46,9 @@ possible. A clarify names nothing, so it would satisfy an empty label for free.
`Score` requires the route to land the case's intent before it credits a
destination hit, or the floor label would score itself.
## Seven cases assert the floor, and six of them cluster
## Seven cases assert the floor, and five of them cluster
The six are homelab operations. `SourceRecall`, `SourceNetwork` and
The five are homelab operations. `SourceRecall`, `SourceNetwork` and
`SourceAttention` overlap on every question about the box, because `mavpoll`
writes its netdata and uptime-kuma observations into the fact store recall
reads. "почему сервер тормозит" is answerable from all three. Naming one takes
@@ -57,6 +57,8 @@ the other two off the turn.
That is a finding about the enum rather than a gap in the labelling. The floor
is the right answer there and the fixture now says so out loud.
All seven were written by an agent and confirmed by the owner on 08-08-2026.
## A drift the labelling found
`WorldQueryGrammars` went into `buildRouter` with V-655 and never into
@@ -0,0 +1,88 @@
# The parked clarify ride, bounded and re-measured
Date: 2026-08-08, V-663. Same 140 turns, same driver, third and fourth runs of
the day. Before is `d6f3914`, after is that plus two changes.
## What was measured before
One question about a reminder's day rode turns 6 to 13. It ended only because
turn 14 was a new request. Two of those turns are the worst replies in the
corpus:
```text
спасибо -> Сейчас 21:25. В какой день?
привет -> Сейчас 21:25. В какой день?
```
V-654 had already added `MaxSuspends` and the tail count had not moved.
## Why three bounds let it happen
The TTL, `MaxAttempts` and `MaxSuspends` all exist and all were rearmed.
An aside spends no attempt, so `MaxAttempts` never reaches it. A turn reading as
a failed answer zeroes `Suspends`, so `MaxSuspends` never reaches the asides.
`спасибо` read as a failed answer. So the pleasantry both produced the reply
above and paid for it out of the budget meant to end the ride.
## The two changes
`PendingQuestion.Rides` counts the same event as `Suspends` with the resets
taken out. Set once, incremented only in `noteSuspended`, carried across the
re-park in `askRemainingGap`. `MaxRides` is 4, one looser than `MaxSuspends` so
that the tighter statement about a run stays reachable.
A pleasantry is an aside, not a failed answer. The set is a new closed lexicon
entry, matched as whole utterances rather than tokens. `вечер` answers `это утра
или вечера?` and `нет` answers a confirm. Anything that could fill a slot stays
out.
## Result
| | before | after |
|---|---|---|
| turns carrying a clarify tail | 21 | 17 |
| turns carrying any failure string | 31 | 29 |
| the longest ride | 8 turns | 4 turns |
The turns carrying a tail, by number:
```text
before 6 7 8 9 10 11 12 13 53 54 55 56 57 99 102 103 116 136 137 138 139
after 6 7 8 9 53 54 55 56 57 99 100 101 116 136 137 138 139
```
Turn 10 is the change. It now reads:
```text
спасибо -> Прошлую просьбу отпускаю. Пожалуйста, я всегда готова помочь тебе.
```
She lets the request go, says so, and answers the man. Turns 11 to 13 are clean.
**`MaxRides` is not what fired.** The pleasantry is an aside now, so it no
longer breaks the run. `MaxSuspends` reached three on turn 10 and ended it.
`Rides` is the backstop for the shape where an answer really does break the run.
No turn in this corpus reaches it.
## What did not move
Four rides are untouched. Turns 53 to 57 are five consecutive asides against a
reminder missing its day. Turn 58 is a new request that drops it. Nothing
pleasant appears in that run, so neither change applies. Turns 99 to 101 shifted
by one, and 116 and 136 to 139 are unchanged.
So the fix is worth four turns of twenty-one. What is left is asides against a
question the owner never answers. `MaxSuspends` was written for that shape and
does bound it, at four turns each.
## Not attributable
Latency moved p50 1.1s to 1.5s and p95 2.8s to 3.0s, and the 34.3s outlier in
the earlier run is gone. Both runs had the workstation up. Read none of it as
caused by this change.
One unrelated defect appeared in the after run and is recorded here because it
is visible in the transcript. Turn 4 answered `Я записала твою привычкуRegarding
coffee without sugar.` That is English leaking into a Russian reply with no
space in front of it. It is a phrasing defect and it has no task yet.
@@ -0,0 +1,192 @@
# Two heads on e5-small, and the first destination the router did not need a model for
Measured 2026-08-08 on workpc (Radeon RX 7900 GRE, ROCm). Covers V-661, step 3 of
`docs/plans/18-routing-heads-on-e5-small.md`. Workspace is `~/Programs/embed-training`,
scripts `gen_query_source.py`, `label_source.py`, `build_heads_corpus.py`,
`train_heads.py`, `score_confidence.py`.
## Two heads, not four
Intent is 7 classes and destination is 12 plus the `SourceUnknown` floor, sharing
one masked mean pool over one forward pass. The plan asked for four. Two of them
have no labels and neither is a GPU problem.
**Mood is cut, not deferred.** Maven's enum is `neutral, happy, thinking, tired,
confused` and it describes her own reply state, not the speaker's emotion.
`psytechlab/EmpatheticIntents-ru` was the only candidate and its 32 emotion
labels do not map onto it. There is nothing to train against.
**BIO slot tags stay in the MASSIVE body.** No Maven-domain span corpus exists.
`2026-08-08-massive-warm-start.md` records that no second Russian slot-filling
corpus is reachable at all.
The destination loss is masked with `ignore_index`. Only a query turn reaches
`queryWalk`, so a reminder contributes nothing to it.
## Where the destination labels came from
V-660 taught the router prompt to name a destination. That made gemma-4-12b a
teacher, and this distils it.
Labelling the 300 query rows already in `train_v5.jsonl` gave 277 destinations.
The shape was unusable: the floor 101, calendar 62, recall 47, and `feeds` and
`attention` at zero. A 13-way softmax cannot learn a class with no examples.
`gen_query_source.py` is the destination half of `gen_corpus.py` and runs the
same two passes. Gemma writes questions whose answer lives in one named place.
The daemon's own `routeSystem` prompt then routes each one back. A line survives
only when the intent is `query` **and** the source is the destination it was
generated for. The glosses are copied verbatim out of `route_system.txt`, so the
generator and the labeller work from one definition.
The prompt and the GBNF are dumped from `internal/router/llmrouter.go` by
`TestDumpPrompt`, never retyped. The workspace held its own copies and V-660
changed both.
1229 kept of 2373 generated, 51.8%. Merged corpus is 3664 rows carrying 1727
destinations:
| | rows | | rows |
|---|---|---|---|
| the floor | 220 | world | 132 |
| calendar | 180 | money | 126 |
| tasks | 169 | list | 124 |
| recall | 167 | self, feeds, attention | 120 each |
| weather | 136 | network | 103 |
| | | home | 50 |
`home` is thin because the agreement filter rejected most of what was generated
for it. A question about the house routes `act` more often than `query`. That is
the filter working, and 50 is the finding rather than a shortfall.
**The floor was regenerated once.** The first 120 rows carried one sentence
shape across eight topics. That shape was "что там с X" and its two synonyms.
Every named destination varied and only the floor collapsed. The reason is that
the generator varies a topic, and ambiguity is not a topic.
`gen_query_source.py` now rotates six floor shapes. A `почему` question, a yes
or no question, and a question carried by intonation alone. Then a
better-or-worse question, a status question, and an existence question. That is
a fix to degenerate generation. It is not fitting to the fixture, whose floor
cases are homelab operations and match none of the six.
The 1229 generated rows carry `intent: null`. Every one is a query by
construction. There are five times as many as the corpus has query rows, so
including them would make query half the intent corpus.
## Result
Three seeds, two bodies, epoch chosen on the intent dev slice and never on a
destination number.
| body | floor corpus | intent mean | destination mean |
|---|---|---|---|
| warm-started `out/body_massive` | one shape | 93.6% | 75.8% |
| stock `multilingual-e5-small` | one shape | 93.6% | 76.8% |
| warm-started `out/body_massive` | six shapes | 93.6% | **80.8%** |
Best single run is destination **29/33 (87.9%)**, seed 0 on the rotated floor.
Peak 1.68GB of 17.2GB, under four minutes end to end.
Against the two arms already measured on the same 33 labelled cases:
| | destination |
|---|---|
| classifier cascade (V-659) | 12/33 (36.4%) |
| cascade + gemma-4-12b (V-660) | 24/33 (72.7%) |
| two heads on e5-small | 29/33 (87.9%) |
A 118M encoder beats the 12B teacher it was distilled from, on the fixture. The
per-destination split is where it happens: **recall 15/15** and **world 5/5**.
Recall was 0/15 on the cascade and 14/15 through gemma.
Read 87.9% as one seed of a mean of 80.8%, not as a headline. Three seeds score
29, 25 and 26 of 33. One case is 3 points on a fixture this small.
Intent is **not** comparable to the 73/96 and 81/96 figures those two arms
scored. A softmax has no clarify class. The head's fixture is the 88 cases that
carry an intent, and the 8 `want_clarify` cases are scored separately below.
## The MASSIVE warm-start is worth nothing here either
Step 2 measured it at +0.4 points of intent accuracy and called that inside seed
noise. Destination was the open question, because MASSIVE has a
`definition_word` slot that looked like a `SourceWorld` signal sitting in a head
already trained.
It is not. The two bodies score the same intent mean to one decimal. Stock is
one point ahead on destination, which is a third of one case. Nothing here argues
for keeping the warm-start step. Dropping it removes a dependency on a corpus
pull that `datasets` 5.0 cannot do.
## The floor moved, calendar did not
Before the rotation, all seven misses at seed 0 were the floor and calendar. The
head named a destination where the fixture says walk the chain, and it was
confident doing it. `"почему сервер тормозит"` read `world` at 0.80.
`"хватает ли места под новые бэкапы"` read `network` at 0.82. Those are the five
homelab cases V-659 flagged, where `SourceRecall`, `SourceNetwork` and
`SourceAttention` all overlap. `mavpoll` writes its observations into the fact
store recall reads.
| | one shape | six shapes |
|---|---|---|
| the floor | 3/7 | 6/7, 6/7, 5/7 |
| calendar | 3/6 | 3/6, 3/6, 3/6 |
| recall | 15/15 | 15/15 at seed 0 |
| world | 5/5 | 5/5 |
The floor was a corpus defect and it cost 3 cases. Sentence variety carried it,
not homelab vocabulary, which the training rows still do not contain.
**Calendar is 3/6 at every seed and is a different problem.** It is the shape
V-660 named. The possessive agenda rules claim those cases at stage 0 and
deliberately name nothing, so no destination label reaches the head. Training
cannot move a case the head never sees. That one wants the owner's call.
## Max softmax separates, weakly, and the gate stays
The plan argues max softmax is a calibratable confidence where `Confidence: 1.0`
was a hardcode. Measured on the intent head:
| | n | mean confidence |
|---|---|---|
| correct | 80 | 0.897 |
| wrong | 8 | 0.705 |
| `want_clarify` | 8 | 0.685 |
The softest correct answer is 0.66 and 3 of 8 clarify cases sit below it. So a
single cut buys three clarifies at no false-clarify cost, and no more.
The other five explain themselves. `"напомни"` scores 0.94 as `reminder` and
`"сделай это"` scores 0.80 as `act`. Both are intent-certain and slot-empty, and
confidence was never the signal there. `gateLLMDecision` already catches exactly
that shape, an act with no allowlisted fn or a keyless fact, and it keeps doing
so. The head replaces the hardcode. It does not replace the gate.
## An incident worth recording
The first generation run produced zero rows for eight destinations. `mavgpud`
yields the card when another process wants it (V-488) and llama-server answers
503 until the model is back. Every generate call inside that window burned one of
the destination's batches. The run walked its own cap without a single successful
call. The log said `503` 260 times, and the summary line said 64.1% keep rate,
which read as success.
`call()` now retries a 503 with backoff. A generator that treats an unloaded
model as a bad generation is a silent-corpus bug, not a slow one.
## What this does not measure
**Nothing here runs in Go.** The heads are a `heads.pt` and an
`out/body_heads/` on workpc. Reaching the daemon needs an ONNX export and a
caller. The resident e5-small must not be replaced by this copy: recall depends
on that file, and `EmbedQuery`/`EmbedPassage` are its contract.
The destination fixture carries 4 of 13 classes: recall 15, floor 7, calendar 6,
world 5. `tasks`, `money`, `list`, `home`, `network`, `feeds`, `attention` and
`self` have no gold case. So 78.8% is silent on eight destinations that together
hold 800 training rows.
Latency was not measured. A forward pass of a 118M encoder should beat a 1.7B
decoder on a query turn. That is arithmetic, not a number from this box.
@@ -0,0 +1,106 @@
# A slot head, and the corpus that did not exist this morning
Measured 2026-08-08 on workpc, the same day as `2026-08-08-routing-heads-two-head.md`
and against the same fixtures. Workspace is `~/Programs/embed-training`, new
scripts `slot_grammar.gbnf`, `slot_system.txt`, `label_slots.py`, `merge_slots.py`.
## The corpus was the whole problem
The two-head measurement said BIO slot tags stay in the MASSIVE body, because no
Maven-domain span corpus exists. That was true of found corpora and false of
made ones. Destination had the same shape at breakfast. V-660 gave gemma-4-12b a
string to write and the label problem became a generation problem.
The same trick applies to spans. `slot_grammar.gbnf` emits a list of
`{"slot": ..., "text": ...}` and the enum closes over Maven's own five: `time`,
`text`, `key`, `value`, `fn`. `slot_system.txt` demands each span be an exact
substring of the utterance.
**The agreement filter is free here.** Destination needed a second pass. The
daemon's own router prompt had to route each generated line back. A span needs
no second call. It either occurs in the utterance or it does not, and
`label_slots.py` drops it with `find()`.
1702 rows labelled from `train_v5.jsonl`, the reminder, fact, note, act and
query intents. Chat and system carry no slot and were never asked.
| slot | spans |
|---|---|
| text | 1175 |
| time | 485 |
| fn | 381 |
| key | 72 |
| value | 65 |
37 spans dropped as not-a-substring, 2.2% of the pile. Nothing failed to parse,
which is the grammar doing its job. 409 rows came back with no span at all.
Those are kept and tagged all `O`. An utterance carrying no slot teaches the
head not to invent one. An empty list is a label and not a miss.
`key` and `value` are thin because they come from facts alone. That is the
shape of the corpus, not a labeller failure.
## Three heads on one forward pass
Intent and destination were already two linear heads over one masked mean pool.
Slots is a third head over the per-token states of the same pass, so the marginal
cost is one `Linear(384, 11)`.
The tag set is `O` plus `B-` and `I-` for each of the five. A softmax cannot
emit a tag that does not exist. That is the structural guarantee the GBNF buys
for the teacher, and the head gets it for free.
Two masking rules, both `ignore_index`. A row with no `spans` key contributes
nothing, which covers the 1900 generated destination rows and every chat and
system turn. A padding or special-token position contributes nothing either.
Scoring is exact-match span F1, not token accuracy. Most tokens are `O`, so a
tagger that predicts nothing anywhere scores above 90% on tokens.
## Result
Three seeds, 24 epochs, epoch chosen on the intent dev slice alone.
| | two heads | three heads |
|---|---|---|
| intent mean | 93.6% | 92.8% |
| destination mean | 80.8% | 82.8% |
| destination best | 29/33 (87.9%) | 29/33 (87.9%) |
| slot span F1 mean | — | 72.4% |
**The slot head costs nothing and adds a third decision.** Intent moves 0.8
points down and destination 2 points up. Both sit inside the seed spread those
two numbers already had. Read this as unchanged, not as a trade.
The saved checkpoint is seed 1 at epoch 10: intent 93.2%, destination 81.8%,
slot F1 75.8%. `heads.pt` now carries three state dicts and the `bio` list
beside the intent and source enums.
## Epoch selection is now wrong for one of the three heads
Slot F1 was still climbing when the intent-selected epoch stopped it. Seed 0
selects epoch 13 at 70.9% and reaches 76.1% at epoch 20. Seed 1 selects epoch 10
at 75.8% and reaches 80.0% at epoch 24.
So the three tasks want different epochs and the harness picks one. Two ways
out, and neither was taken here. Select on a joint score, which needs an
argument about weights. Or give the slot head its own dev slice and its own
early stop, which means the heads stop being one checkpoint.
Leaving it costs about 4 points of slot F1 and nothing else. The router
contract is graded on intent. Selecting on it is the conservative choice rather
than an oversight.
## What this does not measure
The same gap as the two-head run. **Nothing of this runs in Go.** Three heads
instead of two does not change that. It does change the export. The slot head
reads the per-token states, so an ONNX graph exposing only the pooled vector is
not enough.
There is no held-out span fixture. Span F1 is measured on the dev slice of the
same distilled corpus, so it scores agreement with gemma rather than
correctness. The V-572 defect utterance, `"напомни в 11:00 позвонить маме"`, is
labelled correctly by the teacher, but one utterance is not a fixture.
`key` at 72 spans and `value` at 65 are too thin to report a per-slot number.
@@ -0,0 +1,631 @@
# Raw transcript, two weeks of usage
## Day 1
```
YOU: доброе утро
MAVEN: Я не говорю так сейчас, ведь на улице уже вечер.
[2.8s t=218]
YOU: какой сегодня день?
MAVEN: сегодня суббота, 8 августа 2026 года
[0.0s t=219]
YOU: сколько времени?
MAVEN: сейчас 21 час 9 минут
[0.0s t=220]
YOU: запиши что я пью кофе без сахара
MAVEN: Я записала твою заметку про кофе。
[1.5s t=221]
YOU: мой любимый язык программирования go
MAVEN: Я записала твою заметку про Go.
[1.5s t=222]
YOU: напомни в 11:00 позвонить маме
MAVEN: Сейчас 21:10. В какой день?
[0.4s t=223]
YOU: что у меня сегодня?
MAVEN: на 08.08.2026 ничего нет. На какое время поставить напоминание?
[2.5s t=224]
YOU: что такое TCP?
MAVEN: вот что я нашла: TCP (Transmission Control Protocol) — это транспортный протокол для передачи данных. Он обеспечивает доставку всех пакетов и позволяет собрать их в правильном порядке. Также TCP является частью сетевой модели TCP/IP, которая описывает способ передачи цифровых данных от источника к получателю. На какое время поставить напоминание?
[2.8s t=225]
YOU: сколько будет 17 на 23?
MAVEN: вот что я нашла: в источниках нет результата умножения 17 на 23. На какое время поставить напоминание?
[1.6s t=226]
YOU: спасибо
MAVEN: Сейчас 21:10. В какой день?
[1.1s t=227]
```
## Day 2
```
YOU: привет
MAVEN: Сейчас 21:10. В какой день?
[0.9s t=228]
YOU: что нового?
MAVEN: не знаю — не нашла у тебя такой записи. На какое время поставить напоминание?
[1.3s t=229]
YOU: какая погода?
MAVEN: для какого города? На какое время поставить напоминание?
[0.7s t=230]
YOU: запиши что пароль от вайфая лежит в ящике стола
MAVEN: Прошлую просьбу отпускаю. Я записала, что пароль от вай-фай лежит в ящике стола.
[1.5s t=231]
YOU: где лежит вайфай пароль?
MAVEN: не знаю — не нашла у тебя такой записи.
[0.5s t=232]
YOU: добавь молоко в список покупок
MAVEN: Такую команду я не знаю.
[1.0s t=233]
YOU: что у меня в списке покупок?
MAVEN: в списке покупок пусто.
[0.0s t=234]
YOU: кто такой Линус Торвальдс?
MAVEN: вот что я нашла: Линус Торвальдс — финско-американский программист из Хельсинки, создатель ядра Linux. Он родился 28 декабря 1969 года и окончил факультет компьютерных наук Хельсинкского университета.
[2.6s t=235]
YOU: какой у меня любимый язык?
MAVEN: не знаю — не нашла у тебя такой записи.
[0.6s t=236]
YOU: сколько у меня задач?
MAVEN: сначала: проверить интейк форму на боксе (завтра); купить хлеб; купить молоко; полить цветы; починить кран
[0.0s t=237]
```
## Day 3
```
YOU: как дела?
MAVEN: У меня всё хорошо, я готова помогать тебе. Чем могу быть полезна?
[1.2s t=238]
YOU: напомни завтра в 9 утра купить хлеб
MAVEN: хорошо, напомню завтра в 09:00.
[0.3s t=239]
YOU: что у меня завтра?
MAVEN: на 09.08.2026 ничего нет.
[0.0s t=240]
YOU: отмени напоминание про хлеб
MAVEN: Такую команду я не знаю.
[0.5s t=241]
YOU: какие у меня напоминания?
MAVEN: вот что я нашла: в источниках нет информации о твоих напоминаниях.
[0.7s t=242]
YOU: сохрани мне адрес гостиницы в Сочи
MAVEN: Я записала твой адрес отеля в Сочи.
[2.2s t=243]
YOU: что я сохранил про Сочи?
MAVEN: вот что я нашла: в источниках нет информации о том, что ты сохранил про Сочи.
[1.5s t=244]
YOU: почему сервер тормозит?
MAVEN: вот что я нашла: сервер может тормозить из-за неудачной конфигурации или неоптимальных настроек, отсутствия кэширования, неоптимизированного программного обеспечения, которое нужно обновлять, а также из-за режима энергосбережения процессора, когда он не успевает повысить частоту при нагрузке.
[3.7s t=245]
YOU: хватает ли места под новые бэкапы?
MAVEN: нового вот что: [$] Changes in shadow-utils password-expiration features; The Software Stewardship Lab launches; LightDM lives: version 1.33.0 released
[0.7s t=246]
YOU: выключи свет в спальне
MAVEN: Такую команду я не знаю.
[0.4s t=247]
```
## Day 4
```
YOU: доброе утро
MAVEN: Сейчас уже вечер, поэтому я не могу сказать тебе «доброе утро».
[1.4s t=248]
YOU: что я пропустил?
MAVEN: вот что я нашла: в источниках нет информации о том, что ты пропустил.
[1.3s t=249]
YOU: о чём мы вчера говорили?
MAVEN: не знаю — не нашла у тебя такой записи.
[1.0s t=250]
YOU: запиши что я записался к врачу на четверг
MAVEN: Я записала твою запись к врачу.
[1.2s t=251]
YOU: когда я иду к врачу?
MAVEN: вот что я нашла: в источниках нет информации о том, когда именно ты идешь к врачу.
[2.1s t=252]
YOU: что такое ZFS?
MAVEN: вот что я нашла: ZFS — это современная и мощная файловая система, в которой сочетаются безопасность, масштабируемость и удобство управления данными. Она была разработана корпорацией Sun Microsystems и объединяет в себе файловую систему с менеджером томов. Изначально она создавалась для того, чтобы преодолеть недостатки традиционных RAID-массивов, такие как проблема «дыры записи».
[2.1s t=253]
YOU: столица Франции?
MAVEN: вот что я нашла: столица Франции — Париж.
[3.7s t=254]
YOU: переведи слово ремонт на английский
MAVEN: вот что я нашла: слово «ремонт» на английский можно перевести как repair (например, ремонт дороги или крыши) или fix (общее слово для починки техники и мебели).
[2.7s t=255]
YOU: сколько я потратил в этом месяце?
MAVEN: я не отслеживаю траты — не подключено.
[0.7s t=256]
YOU: спокойной ночи
MAVEN: Спокойной ночи. Я уже подготовила всё к твоему отдыху.
[1.2s t=257]
```
## Day 5
```
YOU: привет
MAVEN: Привет! Я готова помогать тебе сегодня.
[1.0s t=258]
YOU: какая погода в Москве?
MAVEN: погода не настроена.
[0.9s t=259]
YOU: что там с бэкапами?
MAVEN: вот что я нашла: бэкап — это резервная копия данных, созданная в определенный момент времени и сохраненная на локальном или облачном носителе. Его делают, чтобы быстро восстановить информацию в случае необходимости.
[2.3s t=260]
YOU: покажи что требует внимания
MAVEN: мне пока нечего смотреть — у Praxis нет источников.
[0.5s t=261]
YOU: отметь это как сделанное
MAVEN: команды «отметь» пока нет в списке — она уже предложена, включи через клиент.
[0.0s t=262]
YOU: запиши что я купил новые наушники
MAVEN: Я записала твою покупку новых наушников.
[1.8s t=263]
YOU: какие у меня заметки за неделю?
MAVEN: не знаю — не нашла у тебя такой записи.
[0.0s t=264]
YOU: расскажи про Kubernetes
MAVEN: вот что я нашла: Kubernetes — это портативная расширяемая платформа с открытым исходным кодом для управления сервисами и контейнеризованными рабочими нагрузками. Управлять им можно через интерфейс в командной строке kubectl. Я нашла эти данные в предоставленных источниках.
[1.8s t=265]
YOU: кто я?
MAVEN: вот что я нашла: в источниках нет информации о том, кто ты.
[1.4s t=266]
YOU: пока
MAVEN: До свидания, пока!
[1.4s t=267]
```
## Day 6
```
YOU: доброе утро
MAVEN: Сейчас уже вечер, так что доброе утро я тебе сказать не могу. Я готова помочь тебе с чем-нибудь другим.
[1.4s t=268]
YOU: сколько времени?
MAVEN: сейчас 21 час 11 минут
[0.0s t=269]
YOU: напомни в 18:30 позвонить в банк
MAVEN: Сейчас 21:11. В какой день?
[0.3s t=270]
YOU: поставь чайник
MAVEN: Сейчас 21:11. В какой день?
[1.1s t=271]
YOU: включи музыку
MAVEN: Сейчас 21:11. В какой день?
[0.9s t=272]
YOU: что у меня в календаре на пятницу?
MAVEN: вот что я нашла: в источниках нет информации о твоих личных событиях в календаре на пятницу. В одном из текстов упоминается пятница 24 июля 2026 года, а в другом — что сегодня пятница 7 августа 2026 года. На какое время поставить напоминание?
[2.6s t=273]
YOU: во сколько у меня встреча?
MAVEN: не знаю — не нашла у тебя такой записи. На какое время поставить напоминание?
[0.3s t=274]
YOU: запиши что дедлайн по проекту в понедельник
MAVEN: Прошлую просьбу отпускаю. Я записала твой дедлайн на понедельник.
[1.7s t=275]
YOU: успею ли я до дедлайна?
MAVEN: вот что я нашла: в источниках нет информации о твоем дедлайне, поэтому я не могу ответить, успеешь ты или нет.
[2.0s t=276]
YOU: спасибо
MAVEN: Пожалуйста, я всегда готова помочь тебе.
[1.4s t=277]
```
## Day 7
```
YOU: привет
MAVEN: Привет! Чем я могу тебе помочь?
[1.3s t=278]
YOU: как ты?
MAVEN: Я в порядке, спасибо, что спросил. Готова помогать тебе!
[1.4s t=279]
YOU: расскажи анекдот
MAVEN: вот что я нашла: в источниках есть ссылки на сборники и сайты с анекдотами, но самих текстов анекдотов там нет, кроме одного: ударилась лягушка о землю головой и стала девицей прекрасной, а второй раз — месивом кровавым.
[3.9s t=280]
YOU: что ты умеешь?
MAVEN: Я Мэйвен, твоя помощница. Я живу на твоём сервере и отвечаю на твои вопросы. Я запоминаю то, что ты мне говоришь, веду заметки, ставлю напоминания, а также читаю твой календарь и задачи. Я отвечаю на вопросы о мире, используя поиск или офлайн-энциклопедию. Если ты разрешишь, я могу управлять домом, локальной сетью, лентами, списком покупок и погодой, а также писать в телеграм. Я говорю по-русски и по-английски.
[2.9s t=281]
YOU: запиши что я начал бегать по утрам
MAVEN: Я записала твою новую привычку
[1.5s t=282]
YOU: я бегаю по утрам уже неделю
MAVEN: записала: я бегаю по утрам уже неделю
[0.6s t=283]
YOU: как часто я бегаю?
MAVEN: вот что я нашла: в источниках нет информации о том, как часто бегаешь ты. Там указано, что для похудения частоту тренировок можно доводить до 5–6 раз в неделю, а новичкам ежедневные тренировки без подготовки могут привести к травмам.
[3.3s t=284]
YOU: сколько стоит биткоин?
MAVEN: вот что я нашла: в источниках нет информации о стоимости биткоина.
[1.2s t=285]
YOU: какие новости?
MAVEN: вот что нового: [$] Changes in shadow-utils password-expiration features; The Software Stewardship Lab launches; LightDM lives: version 1.33.0 released
[0.8s t=286]
YOU: хорошего дня
MAVEN: Я не могу сказать «хорошего дня», так как сейчас уже вечер. Хорошего отдыха!
[1.3s t=287]
```
## Day 8
```
YOU: доброе утро
MAVEN: Сейчас вечер, так что доброе утро я не скажу. Чем я могу тебе помочь?
[1.2s t=288]
YOU: что у меня сегодня?
MAVEN: на 08.08.2026 ничего нет.
[0.0s t=289]
YOU: напомни через час выпить воды
MAVEN: хорошо, напомню сегодня в 22:11.
[0.3s t=290]
YOU: я выпил воды
MAVEN: записала: я выпил воды
[0.9s t=291]
YOU: запиши что кот ест только сухой корм
MAVEN: Я записала твою заметку про кота.
[1.2s t=292]
YOU: чем питается кот?
MAVEN: вот что я нашла: в природе кошки едят мышей, птиц и ящериц. Дома их правильно кормить мясом, мясными субпродуктами, яйцами и нежирной морской рыбой. Из-за особенностей пищеварения им не подходят каши, макароны и мучные изделия.
[3.4s t=293]
YOU: что такое DNS?
MAVEN: вот что я нашла: DNS (Domain Name System) — это распределенная база данных, которая сопоставляет домены с IP-адресами. Это не единая база на одном сервере, а иерархическая система серверов, работающих совместно. В ней хранятся записи, которые обеспечивают маршрутизацию трафика и управление электронной почтой.
[2.0s t=294]
YOU: проверь статус uptime kuma
MAVEN: uptime kuma — Проверяю статус Uptime Kuma...
[1.1s t=295]
YOU: всё ли в порядке с сервером?
MAVEN: мне пока нечего смотреть — у Praxis нет источников.
[0.5s t=296]
YOU: спасибо
MAVEN: Пожалуйста, я всегда рада помочь тебе.
[2.7s t=297]
```
## Day 9
```
YOU: привет
MAVEN: Привет! Чем я могу тебе помочь сегодня?
[1.9s t=298]
YOU: какой сегодня день недели?
MAVEN: сегодня суббота, 8 августа 2026 года
[0.0s t=299]
YOU: добавь хлеб и сыр в список покупок
MAVEN: Такую команду я не знаю.
[0.6s t=300]
YOU: что в списке покупок?
MAVEN: в списке покупок пусто.
[0.0s t=301]
YOU: удали молоко из списка
MAVEN: Такую команду я не знаю.
[0.4s t=302]
YOU: напомни завтра утром вынести мусор
MAVEN: Сейчас 21:11. Когда?
[0.3s t=303]
YOU: запиши что я поменял масло в машине
MAVEN: Прошлую просьбу отпускаю. записала: запиши что я поменял масло в машине
[0.7s t=304]
YOU: когда я менял масло?
MAVEN: масло в машине — поменял
[0.4s t=305]
YOU: сколько будет 144 делить на 12?
MAVEN: вот что я нашла: в источниках указано решение деления 144 на 12 в столбик, но итоговый результат деления не прописан.
[2.3s t=306]
YOU: пока
MAVEN: До свидания! Если понадобится что-то записать или проверить, я всегда на связи.
[1.2s t=307]
```
## Day 10
```
YOU: доброе утро
MAVEN: Я не могу сказать «доброе утро», так как сейчас уже вечер. Я записала твои предыдущие вопросы.
[1.8s t=308]
YOU: что нового за ночь?
MAVEN: вот что я нашла: в ночь на 8 августа вооруженные силы РФ нанесли второй за последнее время массированный удар, при этом киевская ПВО пропустила все российские ракеты.
[3.3s t=309]
YOU: почему интернет медленный?
MAVEN: вот что я нашла: интернет может работать медленно из-за плохого сигнала из-за удаленности вышки, исчерпанного лимита трафика по тарифу, нахождения вне зоны покрытия 4G/LTE, устаревшей или неисправной SIM-карты, а также из-за сбоя в настройках роутера. Чтобы исправить ситуацию, можно попробовать перезагрузить роутер или перенастроить устройство.
[3.0s t=310]
YOU: какая скорость у меня сейчас?
MAVEN: для какого города?
[0.6s t=311]
YOU: запиши что новый роутер стоит 8000 рублей
MAVEN: Я записала стоимость нового роутера.
[1.2s t=312]
YOU: сколько стоил роутер?
MAVEN: вот что я нашла: цена зависит от модели. Роутер от Ростелеком стоит 5 900 рублей, на OZON есть варианты от 4 052 до 3 251 рублей, а в Билайне цены начинаются от 2 149 рублей.
[2.6s t=313]
YOU: что такое NAT?
MAVEN: вот что я нашла: NAT (Network Address Translation) — это технология в сетях TCP/IP, которая позволяет изменять IP-адрес в заголовке пакета при его прохождении через устройство. Она перенаправляет трафик между локальной и глобальной сетями, позволяет объединять устройства в локальной сети под одним публичным IP-адресом, а также скрывает сеть и экономит IP.
[2.1s t=314]
YOU: напомни в субботу позвонить бабушке
MAVEN: Сейчас 21:12. Когда?
[0.3s t=315]
YOU: покажи мои напоминания
MAVEN: вот что я нашла: в источниках нет информации о твоих конкретных напоминаниях, там только описано, как их создавать и где искать в приложениях. На какое время поставить напоминание?
[3.0s t=316]
YOU: спасибо
MAVEN: Сейчас 21:12. Во сколько напомнить?
[1.0s t=317]
```
## Day 11
```
YOU: привет
MAVEN: Сейчас 21:12. Скажи время — например, «в семь вечера» или «через час».
[0.8s t=318]
YOU: как погода на выходных?
MAVEN: для какого города? На какое время поставить напоминание?
[0.7s t=319]
YOU: что у меня на этой неделе?
MAVEN: вот что я нашла: в источниках нет информации о том, что у тебя на этой неделе. На какое время поставить напоминание?
[1.0s t=320]
YOU: запиши что я хочу прочитать книгу про Go
MAVEN: Прошлую просьбу отпускаю. Я записала твою заметку о книге по Go.
[1.6s t=321]
YOU: что я хотел прочитать?
MAVEN: не знаю — не нашла у тебя такой записи.
[0.9s t=322]
YOU: объясни что такое горутина
MAVEN: вот что я нашла: я проверила источники и узнала, что горутины — это легковесные потоки выполнения или функции, которые позволяют параллельно выполнять код внутри программы. Они являются легкими нитями, которые создаются и планируются для независимого выполнения планировщиком Go.
[2.5s t=323]
YOU: кто написал Войну и мир?
MAVEN: вот что я нашла: роман «Война и мир» написал Лев Николаевич Толстой.
[2.2s t=324]
YOU: включи свет на кухне
MAVEN: Такую команду я не знаю.
[0.5s t=325]
YOU: закрой шторы в комнате
MAVEN: Такую команду я не знаю.
[0.4s t=326]
YOU: спокойной ночи
MAVEN: Спокойной ночи! Я уже подготовила всё к твоему отдыху.
[1.1s t=327]
```
## Day 12
```
YOU: доброе утро
MAVEN: Сейчас уже вечер, поэтому я не могу сказать тебе «доброе утро».
[1.1s t=328]
YOU: сколько сейчас времени?
MAVEN: сейчас 21 час 12 минут
[0.0s t=329]
YOU: я не то имел в виду
MAVEN: Я поняла тебя, исправь, пожалуйста, свою мысль. Что именно ты хотел сказать?
[1.8s t=330]
YOU: о чём мы говорили?
MAVEN: не знаю — не нашла у тебя такой записи.
[0.6s t=331]
YOU: напомни
MAVEN: Сейчас 21:12. Когда?
[0.5s t=332]
YOU: сделай это
MAVEN: пока не умею отвечать на этот вопрос. На какое время поставить напоминание?
[0.5s t=333]
YOU: запиши что я перешёл на новый тариф
MAVEN: Прошлую просьбу отпускаю. Я записала твою новую информацию о тарифе.
[1.5s t=334]
YOU: какой у меня тариф?
MAVEN: вот что я нашла: в источниках нет информации о твоем тарифе.
[1.2s t=335]
YOU: сколько я плачу за интернет?
MAVEN: вот что я нашла: в источниках нет информации о том, сколько именно ты платишь за интернет. Я проверила документы, но точной суммы там нет.
[2.6s t=336]
YOU: спасибо
MAVEN: Я поняла тебя. Что именно ты хочешь, чтобы я сделала сейчас?
[1.5s t=337]
```
## Day 13
```
YOU: привет
MAVEN: Привет! Я готова помогать тебе. Чем займемся сегодня?
[1.3s t=338]
YOU: что там с задачами?
MAVEN: вот что я нашла: я отметила задачу про бэкапы как сделанную.
[1.5s t=339]
YOU: закрывай
MAVEN: какой пункт отметить сделанным?
[0.0s t=340]
YOU: отметь задачу про бэкапы как сделанную
MAVEN: Такую команду я не знаю.
[0.6s t=341]
YOU: что осталось нерешённым?
MAVEN: вот что я нашла: в твоей жизни всё, что остается нерешенным или незавершенным, будет продолжать существовать и требовать внимания. Также в источниках упоминается пример с скобками, который может остаться нерешенным, и вопрос о том, что остается нерешенным в краснодарских перевозках.
[4.2s t=342]
YOU: запиши что я договорился о встрече в среду
MAVEN: Я записала твою договоренность о встрече.
[2.3s t=343]
YOU: когда у меня встреча?
MAVEN: мне пока нечего смотреть — у Praxis нет источников.
[0.0s t=344]
YOU: какая температура на улице?
MAVEN: для какого города?
[0.9s t=345]
YOU: что такое RAID 5?
MAVEN: вот что я нашла: RAID 5 — это массив из независимых дисков, где данные и контрольные суммы (четность) распределяются по кругу на все диски. Этот уровень обеспечивает отказоустойчивость и повышенную производительность, при этом для его работы требуется не менее трех дисков.
[2.3s t=346]
YOU: пока
MAVEN: До встречи!
[1.4s t=347]
```
## Day 14
```
YOU: доброе утро
MAVEN: Я не могу сказать «доброе утро», потому что сейчас вечер. Я уже записала твой вопрос и готова помочь с чем-то другим.
[1.7s t=348]
YOU: подведи итоги недели
MAVEN: week summary — подведи итоги недели
[0.6s t=349]
YOU: что я делал за последние две недели?
MAVEN: вот что я нашла: в источниках нет информации о том, что ты делал за последние две недели.
[1.3s t=350]
YOU: какие заметки я сохранил?
MAVEN: вот что я нашла: в источниках нет информации о том, какие именно заметки ты сохранил.
[1.3s t=351]
YOU: о чём я чаще всего спрашиваю?
MAVEN: обычно ты пьёшь воду около 00:58, отмечаешь «name» около 23:56 и отмечаешь «week summary» около 23:59 — по записям за последние 7 дней.
[0.9s t=352]
YOU: напомни в понедельник в 10 проверить бэкапы
MAVEN: Сейчас 21:12. Это утра или вечера?
[0.3s t=353]
YOU: что у меня в понедельник?
MAVEN: не знаю — не нашла у тебя такой записи. На какое время поставить напоминание?
[0.3s t=354]
YOU: ты меня понимаешь?
MAVEN: Да, я тебя прекрасно понимаю. Чем я могу тебе помочь? На какое время поставить напоминание?
[1.3s t=355]
YOU: спасибо тебе
MAVEN: Сейчас 21:12. Это утра или вечера?
[1.0s t=356]
YOU: спокойной ночи
MAVEN: хорошо, напомню послезавтра в 10:00.
[1.2s t=357]
```
@@ -0,0 +1,132 @@
# The same fortnight, after V-655 merged
Date: 2026-08-08, a few hours after `2026-08-08-two-weeks.md`.
Build: `f8fa0d1` on master, the five compose services rebuilt and recreated.
Same driver, same 140 turns, same reach. This is the diff that baseline was for.
Master now carries V-655, V-659, V-660 and V-661. The change under test is V-655. A query source that decides by seed similarity
is marked `guesses: true`. It is dropped when the cascade names a different
destination.
## Two things confound the comparison and one of them matters
**The workstation was up for the re-run.** `llama-server` on 192.168.1.105
answered a health probe with 200. So routing completed through `llm.Pair`
against gemma-4-12b, which is the arm that names a destination. Its state
during the baseline was not recorded. So a difference here may be the merge, or
may be the better router, and this run cannot separate them.
**The store carried over**, as the baseline said it would. Facts written by the
first run were present from turn 1 of the second.
## Numbers
| | baseline `beb093a` | after `f8fa0d1` |
|---|---|---|
| turns | 140 | 140 |
| p50 | 1.5s | 1.2s |
| p95 | 7.1s | 3.0s |
| max | 33.7s | 4.2s |
| transport errors | 0 | 0 |
| turns carrying a failure string | 41 | 38 |
| string in the reply | before | after |
|---|---|---|
| `на какое время поставить напоминание` | 13 | 13 |
| `не нашла у тебя такой записи` | 8 | 9 |
| `Такую команду я не знаю` | 8 | 8 |
| `для какого города` | 6 | 4 |
| `В какой день` | 6 | 6 |
| `пока не умею` | 5 | 1 |
| `Когда?` | 3 | 3 |
Read the latency as unattributed. The workstation confound covers all of it.
## Defect 2 is the one this was for: four of six fixed
| utterance | before | after |
|---|---|---|
| `что такое TCP?` | `для какого города?` | a real definition |
| `сколько будет 17 на 23?` | `для какого города?` | search, which has no answer |
| `какой у меня любимый язык?` | kernel headlines | `не нашла у тебя такой записи` |
| `что я сохранил про Сочи?` | `Хорошо, сохраню.` | answered as a question |
| `какая скорость у меня сейчас?` | `для какого города?` | `для какого города?` |
| `хватает ли места под новые бэкапы?` | kernel headlines | kernel headlines |
`что такое TCP?` is the clean win. `WorldQueryGrammars` names `world` at stage
0, weather is dropped, and search answers.
`сколько будет 17 на 23?` moved source and not outcome. Weather no longer claims it. Search
cannot do arithmetic, so the reply says the sources have no product of 17 and
23. That is an honest gap where it used to be a wrong
question. Arithmetic has no destination in the enum.
`что я сохранил про Сочи?` was defect 3 and it is gone. The utterance is no
longer read as a capture.
**The two that did not move are both homelab questions.** They are exactly the
cluster the destination fixture flagged. `SourceRecall`, `SourceNetwork` and
`SourceAttention` overlap on every question about the box. Five of the seven
floor cases in that fixture are homelab operations for the same reason. So this
is the enum, not the walk.
## Defect 1 did not move at all
Twenty-six turns still carry a parked clarify tail, the same count as the
baseline. `спасибо тебе` answers `Сейчас 21:12. Это утра или вечера?` and
`спокойной ночи` answers `хорошо, напомню послезавтра в 10:00.`
V-655 was never going to touch this. A parked clarify is dialogue state and not
a query source. It remains the single worst thing about talking to her. The week test, the
fortnight test and this re-run all report it unchanged.
## A gap in the harness, fixed and re-run the same day
`ipc.ChatReply.Source` came back empty on all 140 turns, in both runs. The
driver read the redirect parameter `src` and `cmd/mavweb/chat.go` writes `s`.
So every finding above is read off the reply text instead of off the badge.
Fixed in V-662 and the 140 turns were driven a third time. Sixty-eight of them
name a source. The rest are not query turns and never reach `queryWalk`.
| source | turns |
|---|---|
| search | 27 |
| memory | 13 |
| personal | 9 |
| weather | 5 |
| calendar | 3 |
| attention | 3 |
| list | 2 |
| feeds | 2 |
| tasks, money, self, habits | 1 each |
## What the badge shows that the wording did not
The two unfixed homelab turns are now direct evidence.
```text
какая скорость у меня сейчас? -> weather
хватает ли места под новые бэкапы? -> feeds
```
Both are guessing sources claiming a turn about the box, exactly as the
destination fixture predicted.
The badge also names a defect the wording hid. **Agenda questions are being
claimed by the personal boundary and by Praxis, not by the calendar.**
```text
во сколько у меня встреча? -> personal не нашла у тебя такой записи
когда у меня встреча? -> attention у Praxis нет источников
что у меня в понедельник? -> personal не нашла у тебя такой записи
```
Calendar claimed 3 turns of the 6 that asked about the calendar. That is the
same 3/6 the destination fixture scores and the same 3/6 every seed of the
routing head scores. Three measurements agree. The cause is the one V-660 named. The possessive
agenda rules claim these at stage 0 and name no destination, so the walk
reaches `personal` and `attention` first.
This is the third independent confirmation that the possessive agenda rules
should name the calendar. That call is still the owner's.
@@ -0,0 +1,636 @@
# Raw transcript, two weeks of usage
Companion to `2026-08-08-two-weeks.md`. 140 turns through `POST /api/chat`,
driven by `scripts/usage-run.py` over `scripts/testdata/usage-turns.txt`.
Days are simulated. The wall clock was 20:18-20:27 throughout, which is why
every clock answer says evening.
## Day 1
```
YOU: доброе утро
MAVEN: Сейчас уже вечер, поэтому я не могу сказать тебе «доброе утро».
[1.5s t=78]
YOU: какой сегодня день?
MAVEN: сегодня суббота, 8 августа 2026 года
[0.0s t=79]
YOU: сколько времени?
MAVEN: сейчас 20 часов 18 минут
[0.0s t=80]
YOU: запиши что я пью кофе без сахара
MAVEN: Я записала твою заметку про кофе。
[2.0s t=81]
YOU: мой любимый язык программирования go
MAVEN: Я записала твою заметку про Go.
[1.4s t=82]
YOU: напомни в 11:00 позвонить маме
MAVEN: Сейчас 20:18. В какой день?
[0.4s t=83]
YOU: что у меня сегодня?
MAVEN: на 08.08.2026 ничего нет, на какое время поставить напоминание?
[0.3s t=84]
YOU: что такое TCP?
MAVEN: для какого города? На какое время поставить напоминание?
[1.0s t=85]
YOU: сколько будет 17 на 23?
MAVEN: для какого города? На какое время поставить напоминание?
[0.9s t=86]
YOU: спасибо
MAVEN: Сейчас 20:18. В какой день?
[0.9s t=87]
```
## Day 2
```
YOU: привет
MAVEN: Сейчас 20:18. В какой день?
[0.9s t=88]
YOU: что нового?
MAVEN: не знаю — не нашла у тебя такой записи, на какое время поставить напоминание?
[1.2s t=89]
YOU: какая погода?
MAVEN: для какого города? На какое время поставить напоминание?
[0.6s t=90]
YOU: запиши что пароль от вайфая лежит в ящике стола
MAVEN: Прошлую просьбу отпускаю. Хорошо, сохранил.
[33.7s t=91]
YOU: где лежит вайфай пароль?
MAVEN: вот что я нашла: пароль от вайфая лежит в ящике стола
[7.0s t=92]
YOU: добавь молоко в список покупок
MAVEN: Такую команду я не знаю.
[1.4s t=93]
YOU: что у меня в списке покупок?
MAVEN: в списке покупок пусто.
[0.0s t=94]
YOU: кто такой Линус Торвальдс?
MAVEN: не знаю — не нашла у тебя такой записи.
[1.6s t=95]
YOU: какой у меня любимый язык?
MAVEN: вот что нового: [$] Changes in shadow-utils password-expiration features; The Software Stewardship Lab launches; LightDM lives: version 1.33.0 released
[1.3s t=96]
YOU: сколько у меня задач?
MAVEN: сначала: проверить интейк форму на боксе (завтра); купить хлеб; купить молоко; полить цветы; починить кран
[0.0s t=97]
```
## Day 3
```
YOU: как дела?
MAVEN: рада, как дела?
[4.1s t=98]
YOU: напомни завтра в 9 утра купить хлеб
MAVEN: хорошо, напомню завтра в 09:00.
[0.3s t=99]
YOU: что у меня завтра?
MAVEN: на 09.08.2026 ничего нет.
[0.0s t=100]
YOU: отмени напоминание про хлеб
MAVEN: Такую команду я не знаю.
[1.6s t=101]
YOU: какие у меня напоминания?
MAVEN: не знаю — не нашла у тебя такой записи.
[0.0s t=102]
YOU: сохрани мне адрес гостиницы в Сочи
MAVEN: Хорошо, сохраню.
[5.6s t=103]
YOU: что я сохранил про Сочи?
MAVEN: Хорошо, сохраню.
[2.7s t=104]
YOU: почему сервер тормозит?
MAVEN: есть такое: Почему тормозит сервер сайта: причины медленной работы и способы решения — Почему сервер работает медленно: основные причины торможения, диагностика проблем, способы оптимизации. Практические советы по ускорению работы сервера и сайта.
[3.1s t=105]
YOU: хватает ли места под новые бэкапы?
MAVEN: нового вот что: [$] Changes in shadow-utils password-expiration features; The Software Stewardship Lab launches; LightDM lives: version 1.33.0 released
[1.5s t=106]
YOU: выключи свет в спальне
MAVEN: Такую команду я не знаю.
[9.0s t=107]
```
## Day 4
```
YOU: доброе утро
MAVEN: рада
[5.0s t=108]
YOU: что я пропустил?
MAVEN: я не могу ответить на вопрос о том, что ты пропустил. мне нужно больше информации.
[5.7s t=109]
YOU: о чём мы вчера говорили?
MAVEN: пока не умею отвечать на этот вопрос.
[1.1s t=110]
YOU: запиши что я записался к врачу на четверг
MAVEN: Хорошо, сохранил.
[6.0s t=111]
YOU: когда я иду к врачу?
MAVEN: пока не умею отвечать на этот вопрос.
[1.1s t=112]
YOU: что такое ZFS?
MAVEN: не знаю — не нашла у тебя такой записи.
[1.0s t=113]
YOU: столица Франции?
MAVEN: пока не умею отвечать на этот вопрос.
[1.2s t=114]
YOU: переведи слово ремонт на английский
MAVEN: записала: переведи слово ремонт на английский
[1.1s t=115]
YOU: сколько я потратил в этом месяце?
MAVEN: я не отслеживаю траты — не подключено.
[1.4s t=116]
YOU: спокойной ночи
MAVEN: спокойной ночи
[4.3s t=117]
```
## Day 5
```
YOU: привет
MAVEN: Привет!
[2.6s t=118]
YOU: какая погода в Москве?
MAVEN: погода не настроена.
[1.2s t=119]
YOU: что там с бэкапами?
MAVEN: нашла вот это: Бэкап - что такое резервное копирование и как его сделать — Бэкап (backup) — это резервная копия важных данных, сохраненная в специальный файл. Ею пользуются как подстраховкой: если что-то сломается, с помощью бэкапа можно будет все вернуть. Простейший пример бэкапа — сохранение файла в облаке или о…
[4.0s t=120]
YOU: покажи что требует внимания
MAVEN: я заметила, что у тебя есть запрос о том, что требуется внимание. давай разберёмся вместе.
[5.2s t=121]
YOU: отметь это как сделанное
MAVEN: команды «отметь» нет в списке. Предложила её добавить — включи через клиент.
[0.0s t=122]
YOU: запиши что я купил новые наушники
MAVEN: Заметка сохранена.
[3.7s t=123]
YOU: какие у меня заметки за неделю?
MAVEN: week summary — подведи итоги недели
[0.0s t=124]
YOU: расскажи про Kubernetes
MAVEN: вот что я нашла: Что такое Kubernetes — Сервисы, поддержка и инструменты Kubernetes широко доступны. Название Kubernetes происходит от греческого, что означает рулевой или штурман. Google открыл исходный код Kubernetes в 2014 году.
[1.0s t=125]
YOU: кто я?
MAVEN: name — Ками
[1.0s t=126]
YOU: пока
MAVEN: Пока, спасибо за разговор.
[4.1s t=127]
```
## Day 6
```
YOU: доброе утро
MAVEN: Добро пожаловать! Я Maven — твой домашний ассистент. Как я могу помочь сегодня?
[10.1s t=128]
YOU: сколько времени?
MAVEN: сейчас 20 часов 21 минута
[0.0s t=129]
YOU: напомни в 18:30 позвонить в банк
MAVEN: Сейчас 20:21. В какой день?
[0.3s t=130]
YOU: поставь чайник
MAVEN: Сейчас 20:21. В какой день?
[5.3s t=131]
YOU: включи музыку
MAVEN: Сейчас 20:21. В какой день?
[1.7s t=132]
YOU: что у меня в календаре на пятницу?
MAVEN: вот что я нашла: ты записался к врачу на четверг, лежит вайфай пароль в ящике стола, сохранил про Сочи, пью кофе без сахара, на какое время поставить напоминание?
[5.5s t=133]
YOU: во сколько у меня встреча?
MAVEN: не знаю — не нашла у тебя такой записи, на какое время поставить напоминание?
[0.3s t=134]
YOU: запиши что дедлайн по проекту в понедельник
MAVEN: Прошлую просьбу отпускаю. Хорошо, записала.
[6.3s t=135]
YOU: успею ли я до дедлайна?
MAVEN: вот что я нашла: дедлайн по проекту указан как понедельник.
[5.3s t=136]
YOU: спасибо
MAVEN: спасиба
[3.8s t=137]
```
## Day 7
```
YOU: привет
MAVEN: Привет!
[2.6s t=138]
YOU: как ты?
MAVEN: рада
[4.2s t=139]
YOU: расскажи анекдот
MAVEN: Я могу рассказать интересный анекдот для тебя. Вот - в одном из фильмов был такой сценарий: главному герою дали уйти от команды и он остался один, а потом появился его друг, который сказал ему: 'Тебе не нужно было остаться'. Но это просто игра слов.
[7.1s t=140]
YOU: что ты умеешь?
MAVEN: Я могу помочь с задачами, напоминать о важных датах и событиях. Также я помогу записывать заметки или отвечать по ним.
[4.1s t=141]
YOU: запиши что я начал бегать по утрам
MAVEN: Заметка сохранена.
[9.2s t=142]
YOU: я бегаю по утрам уже неделю
MAVEN: записала: я бегаю по утрам уже неделю
[1.9s t=143]
YOU: как часто я бегаю?
MAVEN: вот что я нашла: ты бегал по утрам
[5.3s t=144]
YOU: сколько стоит биткоин?
MAVEN: есть такое: Цена Bitcoin сегодня | Данные о цене, рынке и графики BTC в ... — В настоящее время 1 Bitcoin стоит 65 тыс $, что на 0,09% выше за последние 24 ч. Отслеживайте цену BTC, рыночную капитализацию, исторический график и рыночную аналитику в реальном времени.
[2.1s t=145]
YOU: какие новости?
MAVEN: вот что нового: [$] Changes in shadow-utils password-expiration features; The Software Stewardship Lab launches; LightDM lives: version 1.33.0 released
[0.8s t=146]
YOU: хорошего дня
MAVEN: Хорошого дня!
[4.1s t=147]
```
## Day 8
```
YOU: доброе утро
MAVEN: Добро пожаловать! Как могу помочь сегодня?
[2.9s t=148]
YOU: что у меня сегодня?
MAVEN: на 08.08.2026 ничего нет.
[0.0s t=149]
YOU: напомни через час выпить воды
MAVEN: хорошо, напомню сегодня в 21:22.
[0.3s t=150]
YOU: я выпил воды
MAVEN: записала: я выпил воды
[1.2s t=151]
YOU: запиши что кот ест только сухой корм
MAVEN: Хорошо, записала.
[5.6s t=152]
YOU: чем питается кот?
MAVEN: вот что я нашла: кот питается только сухим кормом.
[4.5s t=153]
YOU: что такое DNS?
MAVEN: нашла вот это: Dns — Википедия — В случае рекурсивного запроса DNS-сервер опрашивает серверы (в порядке убывания уровня зон в имени), пока не найдёт ответ или не обнаружит, что домена не существует (на практике поиск ...
[2.0s t=154]
YOU: проверь статус uptime kuma
MAVEN: записала: проверь статус uptime kuma
[1.6s t=155]
YOU: всё ли в порядке с сервером?
MAVEN: пока не умею отвечать на этот вопрос.
[1.3s t=156]
YOU: спасибо
MAVEN: спасибо за пожелание
[4.0s t=157]
```
## Day 9
```
YOU: привет
MAVEN: рада,
[2.1s t=158]
YOU: какой сегодня день недели?
MAVEN: сегодня суббота, 8 августа 2026 года
[0.0s t=159]
YOU: добавь хлеб и сыр в список покупок
MAVEN: Такую команду я не знаю.
[1.3s t=160]
YOU: что в списке покупок?
MAVEN: в списке покупок пусто.
[0.0s t=161]
YOU: удали молоко из списка
MAVEN: Такую команду я не знаю.
[1.2s t=162]
YOU: напомни завтра утром вынести мусор
MAVEN: Сейчас 20:22. Когда?
[0.3s t=163]
YOU: запиши что я поменял масло в машине
MAVEN: Прошлую просьбу отпускаю. Такую команду я не знаю.
[1.3s t=164]
YOU: когда я менял масло?
MAVEN: вот что я нашла: по данным из last update 2019 год, заправки Сочи не могли предоставить масло для замены, поэтому это было невозможно.
[7.1s t=165]
YOU: сколько будет 144 делить на 12?
MAVEN: есть такое: Разделить в столбик 144 на 12 - Calc-Best.ru — 144 : 12 = 12 Где 144 - делимое, 12 - делитель, 12 - частное ... Разделить 770 на 7 столбиком Разделить 441 на 9 столбиком Разделить 891 на 9 столбиком Разделить 60 на 5 столбиком Разделить 60 на 6 столбиком
[9.6s t=166]
YOU: пока
MAVEN: я записала
[4.8s t=167]
```
## Day 10
```
YOU: доброе утро
MAVEN: рада
[4.1s t=168]
YOU: что нового за ночь?
MAVEN: нашла вот это: Главные новости сегодня в России, а также самые последние новости со ... — Полное бессилие: киевская ПВО вновь пропустила все российские ракеты. Вооруженные силы РФ в ночь на 8 августа нанесли второй за последнее время массированный ...
[2.1s t=169]
YOU: почему интернет медленный?
MAVEN: вот что я нашла: Почему интернет стал медленнее: причины и способы решения - АО «Газпром ... — На скорость интернета сильно влияют торрент-клиенты, VPN, потоковое видео, фоновая загрузка системных обновлений. Посмотреть, куда уходит трафик, можно в настройках ПК или смартфона.
[2.4s t=170]
YOU: какая скорость у меня сейчас?
MAVEN: для какого города?
[1.2s t=171]
YOU: запиши что новый роутер стоит 8000 рублей
MAVEN: Я записала стоимость нового роутера.
[2.5s t=172]
YOU: сколько стоил роутер?
MAVEN: вот что я нашла: новый роутер стоит 8000 рублей.
[2.0s t=173]
YOU: что такое NAT?
MAVEN: NAT (Network Address Translation) — это технология, которая позволяет устройствам в локальной сети выходить в интернет, используя один общий внешний IP-адрес. Я отвечаю на этот вопрос, используя поиск в сети.
[1.8s t=174]
YOU: напомни в субботу позвонить бабушке
MAVEN: Сейчас 20:23. Когда?
[0.3s t=175]
YOU: покажи мои напоминания
MAVEN: вот что я нашла: в твоих источниках нет напоминаний, только запись о том, что ты купил новые наушники, на какое время поставить напоминание?
[1.7s t=176]
YOU: спасибо
MAVEN: Сейчас 20:23. Во сколько напомнить?
[1.0s t=177]
```
## Day 11
```
YOU: привет
MAVEN: Сейчас 20:23. Скажи время — например, «в семь вечера» или «через час».
[0.9s t=178]
YOU: как погода на выходных?
MAVEN: для какого города? На какое время поставить напоминание?
[0.6s t=179]
YOU: что у меня на этой неделе?
MAVEN: не знаю — не нашла у тебя такой записи, на какое время поставить напоминание?
[0.3s t=180]
YOU: запиши что я хочу прочитать книгу про Go
MAVEN: Прошлую просьбу отпускаю. Я записала твою заметку о книге по Go.
[1.4s t=181]
YOU: что я хотел прочитать?
MAVEN: вот что я нашла: ты хотел прочитать книгу про Go.
[1.1s t=182]
YOU: объясни что такое горутина
MAVEN: вот что я нашла: я узнала, что горутины — это легковесные потоки выполнения или функции в Go, которые позволяют параллельно выполнять код. Они управляются планировщиком Go и работают как независимые нити выполнения.
[1.9s t=183]
YOU: кто написал Войну и мир?
MAVEN: вот что я нашла: роман «Война и мир» написал Лев Николаевич Толстой.
[2.4s t=184]
YOU: включи свет на кухне
MAVEN: Такую команду я не знаю.
[0.5s t=185]
YOU: закрой шторы в комнате
MAVEN: Такую команду я не знаю.
[0.3s t=186]
YOU: спокойной ночи
MAVEN: Спокойной ночи. Отдыхай, я на связи, если что-то понадобится.
[1.2s t=187]
```
## Day 12
```
YOU: доброе утро
MAVEN: Я не могу сказать тебе «доброе утро», потому что сейчас уже вечер.
[1.0s t=188]
YOU: сколько сейчас времени?
MAVEN: сейчас 20 часов 23 минуты
[0.0s t=189]
YOU: я не то имел в виду
MAVEN: Я поняла, исправлюсь. Расскажи, что именно ты хотел сказать?
[1.4s t=190]
YOU: о чём мы говорили?
MAVEN: Я помню всё, что ты мне говоришь, и сейчас отвечаю на твои вопросы на основе этих данных.
[1.4s t=191]
YOU: напомни
MAVEN: Сейчас 20:23. Когда?
[0.7s t=192]
YOU: сделай это
MAVEN: пока не умею отвечать на этот вопрос, на какое время поставить напоминание?
[0.5s t=193]
YOU: запиши что я перешёл на новый тариф
MAVEN: Прошлую просьбу отпускаю. Я записала твою новую информацию о тарифе.
[1.5s t=194]
YOU: какой у меня тариф?
MAVEN: вот что я нашла: ты перешел на новый тариф
[24.4s t=195]
YOU: сколько я плачу за интернет?
MAVEN: есть такое: Как узнать, сколько и когда платить за домашний интернет и ТВ — 3 февр. 2026 г. ... Как узнать, сколько и когда платить за домашний интернет и ТВ · Для тарифов только с домашним интернетом · Была ли статья полезной?
[2.1s t=196]
YOU: спасибо
MAVEN: Спасибо, что обратилась за помощью.
[4.8s t=197]
```
## Day 13
```
YOU: привет
MAVEN: Привет!
[2.5s t=198]
YOU: что там с задачами?
MAVEN: не знаю — не нашла у тебя такой записи.
[1.2s t=199]
YOU: закрывай
MAVEN: какой пункт отметить сделанным?
[0.0s t=200]
YOU: отметь задачу про бэкапы как сделанную
MAVEN: Заметка сохранена.
[5.8s t=201]
YOU: что осталось нерешённым?
MAVEN: вот что я нашла: Российское регулирование криптовалют: что осталось нерешённым к 1 июля — 2 июл. 2026 г. ... Российское регулирование криптовалют: что осталось нерешённым к 1 июля. Пакет законопроектов о регулировании цифровых активов в России не ...
[3.7s t=202]
YOU: запиши что я договорился о встрече в среду
MAVEN: Заметка сохранена.
[2.7s t=203]
YOU: когда у меня встреча?
MAVEN: мне пока нечего смотреть — у Praxis нет источников.
[0.1s t=204]
YOU: какая температура на улице?
MAVEN: для какого города?
[1.7s t=205]
YOU: что такое RAID 5?
MAVEN: не знаю — не нашла у тебя такой записи.
[1.2s t=206]
YOU: пока
MAVEN: Я записала, что мы продолжаем разговор.
[4.1s t=207]
```
## Day 14
```
YOU: доброе утро
MAVEN: рада, доброго утра!
[4.5s t=208]
YOU: подведи итоги недели
MAVEN: записала: подведи итоги недели
[1.7s t=209]
YOU: что я делал за последние две недели?
MAVEN: вот что я нашла: ты не делал ничего за последние две недели.
[6.6s t=210]
YOU: какие заметки я сохранил?
MAVEN: Ты сочти, что сохранила заметку с текстом 'Проверь календарь' и датой 8 августа 2026 г.,
[5.4s t=211]
YOU: о чём я чаще всего спрашиваю?
MAVEN: обычно ты пьёшь воду около 01:12, отмечаешь «name» около 23:56 и отмечаешь «week summary» около 23:59 — по записям за последние 7 дней.
[4.4s t=212]
YOU: напомни в понедельник в 10 проверить бэкапы
MAVEN: Сейчас 20:25. Это утра или вечера?
[0.7s t=213]
YOU: что у меня в понедельник?
MAVEN: вот что я нашла: у тебя в понедельник дедлайн по проекту, на какое время поставить напоминание?
[4.6s t=214]
YOU: ты меня понимаешь?
MAVEN: Да, я тебя прекрасно понимаю. Я готова помогать тебе с заметками и задачами, на какое время поставить напоминание?
[2.4s t=215]
YOU: спасибо тебе
MAVEN: Сейчас 20:25. Это утра или вечера?
[1.0s t=216]
YOU: спокойной ночи
MAVEN: хорошо, напомню послезавтра в 10:00.
[0.9s t=217]
```
+105
View File
@@ -0,0 +1,105 @@
# Two weeks of talking to Maven, as a baseline to re-run
Date: 2026-08-08.
Build: `beb093a` on master, the five compose services as deployed, 41 hours up.
Reach: `POST /api/chat` on mavweb, 140 turns over fourteen simulated days.
Turn source is `tap:text`, so this exercises the path the mic and telegram take.
This exists to be compared against. `scripts/usage-run.py` and
`scripts/testdata/usage-turns.txt` are in the repo, so a re-run after a routing
change is a diff rather than a new opinion. The 2026-08-07 week of usage was
typed by hand and cannot be replayed.
**It measures master, not the branch.** V-655, V-659 and V-660 are unmerged.
Every query source that guesses is still in the chain. That is the change this
baseline is for.
## What re-runs and what does not
The turns file, the driver and the routing behaviour replay. Three things do
not. The wall clock was 20:18 to 20:27 throughout, so every clock and agenda
answer reads evening. Live search and the feed return different text each day.
And the store carries over between runs. A fact written on day 2 is already
present when a re-run reaches day 1.
## Numbers
| | week (2026-08-07) | fortnight (2026-08-08) |
|---|---|---|
| turns | 74 | 140 |
| p50 | 1.5s | 1.6s |
| p95 | 8.0s | 7.1s |
| max | 12.3s | 33.7s |
| transport errors | 0 | 0 |
| string in the reply | turns |
|---|---|
| `на какое время поставить напоминание` | 13 |
| `не нашла у тебя такой записи` | 8 |
| `Такую команду я не знаю` | 8 |
| `для какого города` | 6 |
| `В какой день` | 6 |
| `пока не умею` | 5 |
| `Когда?` | 3 |
**Zero transport errors is not zero wrong answers.** It counts turns that
failed to return a reply, and none did. Every quality number is below.
Those seven strings appear 49 times across 41 of 140 turns. Some turns carry
two, because a parked clarify appends to whatever else was said.
The 33.7s outlier is one note write on day 2. p95 improved against the week
despite it.
## The three defects worth diffing against
### 1. A parked reminder clarify still contaminates later turns
The week test called this the single worst thing to talk to and it is unchanged.
Nineteen turns carry a clarify tail. The worst run is day 1, turns 7 to 13,
which spans a day boundary:
```
что такое TCP? -> для какого города? На какое время поставить напоминание?
сколько будет 17 на 23? -> для какого города? На какое время поставить напоминание?
спасибо -> Сейчас 20:18. В какой день?
привет -> Сейчас 20:18. В какой день?
```
Note that `привет` and `спасибо` do not clear it, and neither does a new day.
### 2. Query sources that guess still claim turns they cannot answer
Weather took `сколько будет 17 на 23?`, `что такое TCP?` and `какая скорость у
меня сейчас?`, answering `для какого города?` to all three. The feed took
`какой у меня любимый язык?` and `хватает ли места под новые бэкапы?` and
answered with kernel headlines.
This is the exact class V-655 removes by marking a source `guesses: true` and
taking it out of `queryWalk`. Six turns here, so the re-run has a number to move.
### 3. A question can still be read as a capture
`что я сохранил про Сочи?` answered `Хорошо, сохраню.` The utterance is
interrogative and was routed to a write. `IsQuestionShaped` catches this
downstream on some paths and did not catch it here.
## What did work
Reminders with a spoken time land correctly, which is V-572 holding:
`напомни завтра в 9 утра купить хлеб` returned `хорошо, напомню завтра в 09:00.`
Facts round-trip. `запиши что новый роутер стоит 8000 рублей` then `сколько
стоил роутер?` returned the stored value. So did the wifi password and the
doctor's appointment.
World questions answer when no local source claims them first. `что такое NAT?`
returned a real definition.
Stage 0 answers land at 0.0 to 0.4s, unchanged.
## What this does not cover
The voice loop, because `mavwaked` and `mavenclient` are not deployed. Reminder
delivery, because nothing fired inside the run window. Telegram intake. And the
three-head routing model, which does not run in Go at all.
+48 -1
View File
@@ -47,6 +47,11 @@ type PendingQuestion struct {
// charging it a retry is the V-554 shape. See CanResume for why it is
// counted at all.
Suspends int
// Rides counts every turn this question has ridden out on the end of
// someone else's reply, over the whole life of the request. Unlike Suspends
// it is never reset and never re-based, which is the only property that
// matters about it (V-663).
Rides int
}
// MaxSuspends — how many times one question may step aside and come back before
@@ -63,10 +68,52 @@ type PendingQuestion struct {
// is that he has moved on and has not said so.
const MaxSuspends = 3
// MaxRides — how many turns one question may ride out on the end of an
// unrelated reply, counted over its whole life (V-663).
//
// MaxSuspends did not move the measurement it was written for. Twenty-six of
// 140 turns carried a tail before it landed and twenty-six carried one after.
// Every bound on this question is rearmed by something ordinary:
//
// - The TTL is an inactivity timer, and both noteSuspended and reaskOrGiveUp
// restart it, so it cannot arrive while he keeps talking.
// - Suspends is zeroed by any turn that reads as an answer, which is where
// "спасибо" and "привет" land. It resets before anything is known to have
// been filled.
// - askRemainingGap builds a fresh question for the second gap, so a reminder
// with two gaps gets a new allowance halfway through.
//
// So Suspends only bites on four strictly consecutive side queries with nothing
// chat-like between them, which is not the shape real conversation has. Rides is
// the same idea with the resets taken out: set once, incremented, carried
// across a re-park, and read by nothing that could lower it.
//
// The shape it is aimed at is measured, not imagined. In the 2026-08-08 run one
// question about a reminder's day rode turns 7 to 13 and ended only because
// turn 14 was a new request. Three asides, then two turns that read as failed
// answers, then two more asides. The asides spend no attempt and the answers
// reset Suspends, so the two bounds take turns being rearmed by the other's
// traffic.
//
// Four, not three. It has to be looser than MaxSuspends or that bound is dead
// code, because Rides is never lower than Suspends and would always fire first.
//
// Do not read this as a fix for the whole ride. It ends the measured one a turn
// early and no more. Most of that ride's length is attempts, spent by turns
// like "спасибо" and "привет" being read as failed answers to a question about
// a day. That is a defect in classifyTurnRole and not in any bound here.
const MaxRides = 4
// CanResume reports whether this question may step aside once more. False ⇒ the
// caller lets the request go and says so; it must never simply stop resuming,
// because a question dropped in silence reads as one that was answered.
func (q *PendingQuestion) CanResume() bool { return q.Suspends < MaxSuspends }
//
// Two bounds, and they answer different questions. Suspends asks whether he has
// walked away from this exchange in the last few turns. Rides asks whether this
// question has been riding long enough that the answer is no regardless.
func (q *PendingQuestion) CanResume() bool {
return q.Suspends < MaxSuspends && q.Rides < MaxRides
}
// Action reads the parked question as the typed action it is assembling
// (pending.go). Derived rather than stored: the question's fields stay the one
+5
View File
@@ -119,6 +119,11 @@ func PartsOfDay() []string { return words("parts_of_day") }
// ReminderVerbs returns the imperatives that open a reminder.
func ReminderVerbs() []string { return words("reminder_verbs") }
// Pleasantries returns the whole utterances that greet, thank or say goodbye.
// Whole utterances and not tokens: see the set's own note for why the tokens
// are unsafe alone.
func Pleasantries() []string { return words("pleasantries") }
// TaskDoneWords returns the words that finish a task, and TaskDropWords the
// words that abandon one. Two sets rather than one with a value, because the
// store records which of the two happened and the caller has to say so.
+13
View File
@@ -176,6 +176,19 @@
"morning", "afternoon", "evening", "night"
]
},
"pleasantries": {
"note": "Whole utterances that greet, thank or say goodbye. They ask for nothing and answer nothing, so a parked question must neither consume them as a failed answer nor be dropped by them (V-663). Matched as WHOLE utterances and never as tokens, because the tokens are not safe alone: \"вечер\" answers \"это утра или вечера?\" and \"нет\" answers a confirm. Anything that could fill a slot stays out. The control words (\"стоп\", \"отмена\") stay out too, because isCancel already owns them and they mean something stronger.",
"words": [
"привет", "приветик", "здравствуй", "здравствуйте",
"доброе утро", "добрый день", "добрый вечер",
"пока", "прощай", "до свидания", "спокойной ночи",
"спасибо", "спасибо тебе", "большое спасибо", "благодарю",
"извини", "извините", "прости", "простите",
"hi", "hello", "hey", "bye", "goodbye",
"good morning", "good evening", "good night",
"thanks", "thank you", "thanks a lot", "sorry"
]
},
"reminder_verbs": {
"note": "The imperatives that mean \"remind me\", in the forms he speaks. The same kind of set as capture_verbs and decided the same way: it is her vocabulary, not a discovery about Russian (Vikunja #530). The alarm verbs joined them in V-627. \"разбуди меня в 6:30\" is a reminder that fires at the hour he gets up, and the set knew no form of it, so an alarm reached IntentReminder only by resembling one to the embedder.",
"words": [
+22
View File
@@ -0,0 +1,22 @@
package router
import (
"os"
"testing"
)
// TestDumpPrompt writes the router prompt and grammar to disk so the training
// workspace labels with the daemon's own contract rather than a retyped copy.
// It is inert unless MAVEN_DUMP_PROMPT names a directory.
func TestDumpPrompt(t *testing.T) {
dir := os.Getenv("MAVEN_DUMP_PROMPT")
if dir == "" {
t.Skip("MAVEN_DUMP_PROMPT unset")
}
if err := os.WriteFile(dir+"/route_system.txt", []byte(routeSystem), 0o644); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(dir+"/route_grammar.gbnf", []byte(routeGrammar), 0o644); err != nil {
t.Fatal(err)
}
}
+167
View File
@@ -0,0 +1,167 @@
# Day 1
доброе утро
какой сегодня день?
сколько времени?
запиши что я пью кофе без сахара
мой любимый язык программирования go
напомни в 11:00 позвонить маме
что у меня сегодня?
что такое TCP?
сколько будет 17 на 23?
спасибо
# Day 2
привет
что нового?
какая погода?
запиши что пароль от вайфая лежит в ящике стола
где лежит вайфай пароль?
добавь молоко в список покупок
что у меня в списке покупок?
кто такой Линус Торвальдс?
какой у меня любимый язык?
сколько у меня задач?
# Day 3
как дела?
напомни завтра в 9 утра купить хлеб
что у меня завтра?
отмени напоминание про хлеб
какие у меня напоминания?
сохрани мне адрес гостиницы в Сочи
что я сохранил про Сочи?
почему сервер тормозит?
хватает ли места под новые бэкапы?
выключи свет в спальне
# Day 4
доброе утро
что я пропустил?
о чём мы вчера говорили?
запиши что я записался к врачу на четверг
когда я иду к врачу?
что такое ZFS?
столица Франции?
переведи слово ремонт на английский
сколько я потратил в этом месяце?
спокойной ночи
# Day 5
привет
какая погода в Москве?
что там с бэкапами?
покажи что требует внимания
отметь это как сделанное
запиши что я купил новые наушники
какие у меня заметки за неделю?
расскажи про Kubernetes
кто я?
пока
# Day 6
доброе утро
сколько времени?
напомни в 18:30 позвонить в банк
поставь чайник
включи музыку
что у меня в календаре на пятницу?
во сколько у меня встреча?
запиши что дедлайн по проекту в понедельник
успею ли я до дедлайна?
спасибо
# Day 7
привет
как ты?
расскажи анекдот
что ты умеешь?
запиши что я начал бегать по утрам
я бегаю по утрам уже неделю
как часто я бегаю?
сколько стоит биткоин?
какие новости?
хорошего дня
# Day 8
доброе утро
что у меня сегодня?
напомни через час выпить воды
я выпил воды
запиши что кот ест только сухой корм
чем питается кот?
что такое DNS?
проверь статус uptime kuma
всё ли в порядке с сервером?
спасибо
# Day 9
привет
какой сегодня день недели?
добавь хлеб и сыр в список покупок
что в списке покупок?
удали молоко из списка
напомни завтра утром вынести мусор
запиши что я поменял масло в машине
когда я менял масло?
сколько будет 144 делить на 12?
пока
# Day 10
доброе утро
что нового за ночь?
почему интернет медленный?
какая скорость у меня сейчас?
запиши что новый роутер стоит 8000 рублей
сколько стоил роутер?
что такое NAT?
напомни в субботу позвонить бабушке
покажи мои напоминания
спасибо
# Day 11
привет
как погода на выходных?
что у меня на этой неделе?
запиши что я хочу прочитать книгу про Go
что я хотел прочитать?
объясни что такое горутина
кто написал Войну и мир?
включи свет на кухне
закрой шторы в комнате
спокойной ночи
# Day 12
доброе утро
сколько сейчас времени?
я не то имел в виду
о чём мы говорили?
напомни
сделай это
запиши что я перешёл на новый тариф
какой у меня тариф?
сколько я плачу за интернет?
спасибо
# Day 13
привет
что там с задачами?
закрывай
отметь задачу про бэкапы как сделанную
что осталось нерешённым?
запиши что я договорился о встрече в среду
когда у меня встреча?
какая температура на улице?
что такое RAID 5?
пока
# Day 14
доброе утро
подведи итоги недели
что я делал за последние две недели?
какие заметки я сохранил?
о чём я чаще всего спрашиваю?
напомни в понедельник в 10 проверить бэкапы
что у меня в понедельник?
ты меня понимаешь?
спасибо тебе
спокойной ночи
+110
View File
@@ -0,0 +1,110 @@
"""Drive a fortnight of conversation through POST /api/chat and record it.
The 2026-08-07 week of usage was typed by hand. This is the same reach and the
same turn source, tap:text, so it exercises the path the mic and telegram take.
The endpoint is a form POST that redirects to /chat with the reply in the query
string. Reading the Location header is the whole protocol, so nothing here
parses HTML.
This exists to be re-run. The baseline is 2026-08-08 against master at beb093a,
in docs/evals/2026-08-08-two-weeks.md. Re-running the same turns after a routing
change is the comparison, so edit the turns file by adding, never by rewriting.
python3 scripts/usage-run.py scripts/testdata/usage-turns.txt out-prefix
Input is one utterance per line. A line starting with "# " opens a day. A blank
line is ignored. Output is a markdown transcript and a jsonl log beside it.
"""
import json
import sys
import time
import urllib.error
import urllib.parse
import urllib.request
URL = "http://127.0.0.1:9201/api/chat"
TIMEOUT = 90
class NoRedirect(urllib.request.HTTPRedirectHandler):
"""A 303 carries the reply. Following it would throw the reply away."""
def redirect_request(self, *a, **kw):
return None
# ProxyHandler({}) is not optional. This box exports http_proxy, urllib honours
# it, and the proxy answers 503 for a loopback address.
OPENER = urllib.request.build_opener(NoRedirect, urllib.request.ProxyHandler({}))
def turn(text):
body = urllib.parse.urlencode({"text": text}).encode()
t0 = time.perf_counter()
try:
OPENER.open(urllib.request.Request(URL, data=body), timeout=TIMEOUT)
return {"reply": "", "error": "no redirect", "secs": time.perf_counter() - t0}
except urllib.error.HTTPError as e:
dt = time.perf_counter() - t0
if e.code != 303:
return {"reply": "", "error": f"HTTP {e.code}", "secs": dt}
loc = e.headers.get("Location", "")
q = urllib.parse.parse_qs(urllib.parse.urlparse(loc).query)
return {
"reply": q.get("r", [""])[0],
# "s", not "src". cmd/mavweb/chat.go writes the badge under that
# name, and reading the wrong one cost both fortnight runs their
# source column: every finding in those docs is inferred from the
# reply wording instead.
"source": q.get("s", [""])[0],
"trace": q.get("t", [""])[0],
"secs": dt,
}
except Exception as e: # a dead box must not lose the turns already done
return {"reply": "", "error": str(e), "secs": time.perf_counter() - t0}
def main():
lines = [l.rstrip("\n") for l in open(sys.argv[1])]
prefix = sys.argv[2]
md = open(prefix + "-transcript.md", "w")
log = open(prefix + ".jsonl", "w")
day = 0
n = 0
print(f"# Raw transcript, two weeks of usage\n", file=md)
for line in lines:
if not line.strip():
continue
if line.startswith("# "):
if day:
print("```\n", file=md)
day += 1
print(f"## {line[2:]}\n\n```", file=md)
continue
n += 1
r = turn(line)
r["day"] = day
r["n"] = n
r["utterance"] = line
log.write(json.dumps(r, ensure_ascii=False) + "\n")
log.flush()
reply = r.get("error") or r["reply"]
print(f"YOU: {line}", file=md)
print(f"MAVEN: {reply}", file=md)
tag = f"[{r['secs']:.1f}s"
if r.get("source"):
tag += f" src={r['source']}"
print(f" {tag} t={r.get('trace', '')}]\n", file=md)
md.flush()
print(f"{n:3} d{day} {r['secs']:5.1f}s {line[:40]:40s} -> {reply[:60]}",
flush=True)
print("```", file=md)
md.close()
log.close()
if __name__ == "__main__":
main()