Files
Maven/docs/evals/2026-08-08-two-weeks-after-merge.md
T
claude 68b2aa9137 Re-run the fortnight against merged master (V-661)
V-655 fixed four of the six turns a guessing query source claimed. The
clean win is 'что такое TCP?', which stage 0 names world and search now
answers instead of weather asking for a city. 'что я сохранил про Сочи?'
is no longer read as a capture.

The two that did not move are both homelab questions, which is the enum
and not the walk. SourceRecall, SourceNetwork and SourceAttention overlap
on every question about the box, and the destination fixture already
flagged that cluster.

The parked clarify is unchanged at 26 turns. It is dialogue state and no
query source could have touched it.

Latency is reported and not attributed. The workstation was up for the
re-run and its state during the baseline was never recorded.

Also corrects the baseline's '49 of 140 turns'. That was the sum of
occurrences. It is 41 turns.
2026-08-08 21:14:52 +04:00

4.3 KiB
Raw Blame History

The same fortnight, after V-655 merged

Date: 2026-08-08, a few hours after 2026-08-08-two-weeks.md. Build: f8fa0d1 on master, the five compose services rebuilt and recreated. Same driver, same 140 turns, same reach. This is the diff that baseline was for.

Master now carries V-655, V-659, V-660 and V-661. The change under test is V-655. A query source that decides by seed similarity is marked guesses: true. It is dropped when the cascade names a different destination.

Two things confound the comparison and one of them matters

The workstation was up for the re-run. llama-server on 192.168.1.105 answered a health probe with 200. So routing completed through llm.Pair against gemma-4-12b, which is the arm that names a destination. Its state during the baseline was not recorded. So a difference here may be the merge, or may be the better router, and this run cannot separate them.

The store carried over, as the baseline said it would. Facts written by the first run were present from turn 1 of the second.

Numbers

baseline beb093a after f8fa0d1
turns 140 140
p50 1.5s 1.2s
p95 7.1s 3.0s
max 33.7s 4.2s
transport errors 0 0
turns carrying a failure string 41 38
string in the reply before after
на какое время поставить напоминание 13 13
не нашла у тебя такой записи 8 9
Такую команду я не знаю 8 8
для какого города 6 4
В какой день 6 6
пока не умею 5 1
Когда? 3 3

Read the latency as unattributed. The workstation confound covers all of it.

Defect 2 is the one this was for: four of six fixed

utterance before after
что такое TCP? для какого города? a real definition
сколько будет 17 на 23? для какого города? search, which has no answer
какой у меня любимый язык? kernel headlines не нашла у тебя такой записи
что я сохранил про Сочи? Хорошо, сохраню. answered as a question
какая скорость у меня сейчас? для какого города? для какого города?
хватает ли места под новые бэкапы? kernel headlines kernel headlines

что такое TCP? is the clean win. WorldQueryGrammars names world at stage 0, weather is dropped, and search answers.

сколько будет 17 на 23? moved source and not outcome. Weather no longer claims it. Search cannot do arithmetic, so the reply says the sources have no product of 17 and 23. That is an honest gap where it used to be a wrong question. Arithmetic has no destination in the enum.

что я сохранил про Сочи? was defect 3 and it is gone. The utterance is no longer read as a capture.

The two that did not move are both homelab questions. They are exactly the cluster the destination fixture flagged. SourceRecall, SourceNetwork and SourceAttention overlap on every question about the box. Five of the seven floor cases in that fixture are homelab operations for the same reason. So this is the enum, not the walk.

Defect 1 did not move at all

Twenty-six turns still carry a parked clarify tail, the same count as the baseline. спасибо тебе answers Сейчас 21:12. Это утра или вечера? and спокойной ночи answers хорошо, напомню послезавтра в 10:00.

V-655 was never going to touch this. A parked clarify is dialogue state and not a query source. It remains the single worst thing about talking to her. The week test, the fortnight test and this re-run all report it unchanged.

A gap in the harness

ipc.ChatReply.Source came back empty on all 140 turns, in both runs. The driver reads it from the redirect query string and there is nothing there. So the badge that says which query source claimed a turn is invisible to the harness, and every finding above is read off the reply text instead.

That is worth fixing before the next re-run. Reading the source directly would have shown the two homelab misses without inferring them from the wording.