# Two weeks of talking to Maven, as a baseline to re-run Date: 2026-08-08. Build: `beb093a` on master, the five compose services as deployed, 41 hours up. Reach: `POST /api/chat` on mavweb, 140 turns over fourteen simulated days. Turn source is `tap:text`, so this exercises the path the mic and telegram take. This exists to be compared against. `scripts/usage-run.py` and `scripts/testdata/usage-turns.txt` are in the repo, so a re-run after a routing change is a diff rather than a new opinion. The 2026-08-07 week of usage was typed by hand and cannot be replayed. **It measures master, not the branch.** V-655, V-659 and V-660 are unmerged. Every query source that guesses is still in the chain. That is the change this baseline is for. ## What re-runs and what does not The turns file, the driver and the routing behaviour replay. Three things do not. The wall clock was 20:18 to 20:27 throughout, so every clock and agenda answer reads evening. Live search and the feed return different text each day. And the store carries over between runs. A fact written on day 2 is already present when a re-run reaches day 1. ## Numbers | | week (2026-08-07) | fortnight (2026-08-08) | |---|---|---| | turns | 74 | 140 | | p50 | 1.5s | 1.6s | | p95 | 8.0s | 7.1s | | max | 12.3s | 33.7s | | transport errors | 0 | 0 | | string in the reply | turns | |---|---| | `на какое время поставить напоминание` | 13 | | `не нашла у тебя такой записи` | 8 | | `Такую команду я не знаю` | 8 | | `для какого города` | 6 | | `В какой день` | 6 | | `пока не умею` | 5 | | `Когда?` | 3 | **Zero transport errors is not zero wrong answers.** It counts turns that failed to return a reply, and none did. Every quality number is below. Those seven strings appear 49 times across 41 of 140 turns. Some turns carry two, because a parked clarify appends to whatever else was said. The 33.7s outlier is one note write on day 2. p95 improved against the week despite it. ## The three defects worth diffing against ### 1. A parked reminder clarify still contaminates later turns The week test called this the single worst thing to talk to and it is unchanged. Nineteen turns carry a clarify tail. The worst run is day 1, turns 7 to 13, which spans a day boundary: ``` что такое TCP? -> для какого города? На какое время поставить напоминание? сколько будет 17 на 23? -> для какого города? На какое время поставить напоминание? спасибо -> Сейчас 20:18. В какой день? привет -> Сейчас 20:18. В какой день? ``` Note that `привет` and `спасибо` do not clear it, and neither does a new day. ### 2. Query sources that guess still claim turns they cannot answer Weather took `сколько будет 17 на 23?`, `что такое TCP?` and `какая скорость у меня сейчас?`, answering `для какого города?` to all three. The feed took `какой у меня любимый язык?` and `хватает ли места под новые бэкапы?` and answered with kernel headlines. This is the exact class V-655 removes by marking a source `guesses: true` and taking it out of `queryWalk`. Six turns here, so the re-run has a number to move. ### 3. A question can still be read as a capture `что я сохранил про Сочи?` answered `Хорошо, сохраню.` The utterance is interrogative and was routed to a write. `IsQuestionShaped` catches this downstream on some paths and did not catch it here. ## What did work Reminders with a spoken time land correctly, which is V-572 holding: `напомни завтра в 9 утра купить хлеб` returned `хорошо, напомню завтра в 09:00.` Facts round-trip. `запиши что новый роутер стоит 8000 рублей` then `сколько стоил роутер?` returned the stored value. So did the wifi password and the doctor's appointment. World questions answer when no local source claims them first. `что такое NAT?` returned a real definition. Stage 0 answers land at 0.0 to 0.4s, unchanged. ## What this does not cover The voice loop, because `mavwaked` and `mavenclient` are not deployed. Reminder delivery, because nothing fired inside the run window. Telegram intake. And the three-head routing model, which does not run in Go at all.