Commit Graph

551 Commits

Author SHA1 Message Date
claude 4d97280d74 a turn hands back its trace id, and one wire op corrects it (V-630)
The correction is the only supervised signal in the box, so the cost of
giving one has to be near zero. That means the surface needs the trace id of
the turn it is showing, which it had no way to learn: handleText returns one
string and the trace was written after the reply left.

The id rides back on ChatReply through the same context sink querySource
uses, so the mic, telegram and the web keep the one signature they share.
CorrectTurn takes a trace id and an optional target, which is deliberately
reach-agnostic: nothing about it assumes a browser.

store.ErrNoSuchTrace gets a wire twin. A turn past the retention bound is
gone, and that is the expected outcome of correcting an old turn, not a
broken database.
2026-08-06 19:45:37 +04:00
claude e5ec4abe04 a corrected turn is promoted to a label that outlives the trace (V-630)
Migration #24 adds routing_labels, and CorrectTurn writes it. Nothing calls it
yet; the wire and the surface are the next commits.

Separate table, and that is the whole retention argument. A trace is a
transcript and expires in 14 days. A correction is a label the owner wrote by
hand, and it is the only supervised signal this box will ever get, so it is
promoted out at the moment he writes it and kept.

should_be may be empty. "That was wrong" with no target is a usable negative and
must not cost more to give than the full answer. UNIQUE(utterance) so a second
correction replaces the first, because his second answer is the one he meant.
The label and the trace stamp go in one transaction: a stamp with no label loses
the signal when the trace expires.

ErrNoSuchTrace is held apart from a write failure. Correcting a turn older than
the bound is the expected case, and the surface should say so rather than report
a broken database.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117tgnmbgZpHVV3XSNw8Qua
2026-08-06 19:30:46 +04:00
claude e1f84a3474 review: a cancelled turn keeps its trace, and a quiet box still expires (V-629)
Two defects found reviewing the PR.

The insert ran on the turn's own context, so a caller that hung up or timed out
cancelled it. That is exactly the turn worth having. It now runs detached, with
a one-second bound of its own, because a write must not hold the reply.

Retention was enforced on write alone, so a box that goes quiet for a month kept
every row until the next sixty-fourth turn. pruneTracesOnStart closes that, and
RoutingTraceRetention is exported so the daemon reads the same number the store
enforces.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117tgnmbgZpHVV3XSNw8Qua
2026-08-06 19:20:24 +04:00
claude 034d4b4359 the store keeps a routing trace for fourteen days (V-629)
Migration #23 adds routing_traces, and internal/store/routingtraces.go writes,
lists and prunes it. Nothing calls it yet; the daemon side is the next commit.

The utterance is stored in clear. A 384-dimension vector of a short sentence is
substantially recoverable, so storing vectors instead would be a privacy claim
we cannot support. Retention is 14 days, enforced on write, and an age rather
than a row count so a busy Tuesday cannot push last Friday out. Store.Wipe
already deletes it with everything else, so explicit deletion needs no new
surface.

A correction is not covered by that bound. When the owner corrects a turn the
pair is promoted out into a seed-shaped row and kept, because a label is not a
transcript. What stays here is the transcript, and the transcript expires.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117tgnmbgZpHVV3XSNw8Qua
2026-08-06 19:13:04 +04:00
claude c1b781fac0 review: act.tool.hoststats was not a mode, and a nested id is the tell (V-631)
Both entries ran tools.Exec. The handler field is prose, so the duplicate hid
there: "tools.Exec against the enabled allowlist" against "tool.Exec through the
configured aliases". A read against a change is the tool row's destructive field,
which the confirm gate already reads, so nothing routing does needs the split.

Its nine examples went with it rather than moving up. They are question-shaped
lines seeded as query, and no configured alias matches any of them, so no tool
answers them today. Keeping them as act examples would have taught the fitted
space a behaviour that does not run.

TestInventoryShape now refuses an id nested under another id. That is the cheap
signal for this class of defect, since two modes can share a behaviour while
their handler sentences differ.

31 modes, 10 ready to fit. The nine with no example are unchanged.

--no-verify: same reason as the parent commit, the 394-line data file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117tgnmbgZpHVV3XSNw8Qua
2026-08-06 18:51:23 +04:00
claude 7b2b9d479a the routing modes are written down, and the file states what fitting one needs (V-631)
Thirty-two modes, written from mavend's handlers, each mapped back to one of the
seven public intents so nothing downstream of the router changes. Data in
internal/modes/modes_v1.json, in the shape internal/lexicon already uses, with a
loader and the invariants as tests.

Two rules decided what counts as a mode. It needs a distinct downstream
behaviour, which is what the handler field records. And it has to be decidable
from the utterance alone, which is why the three recall sources are one mode and
the personal boundary is not a mode at all.

What the file says that the seven intents could not. Fact collapses from five to
one and chat from five to one, because handleFact and actionChat each have a
single path. Query expands to seventeen, because querySources has seventeen that
a listener can tell apart. Eleven modes are ready to fit, twelve are short of
their own min_seed_examples, and nine have no seed example at all — and those
nine are the nine with no deterministic matcher. That is the evidence for doing
V-629 and V-630 before V-632.

system.hoststats is act.tool.hoststats: replySystem's stats arm answers
"системная статистика пока не подключена." and always did, and V-633 gave the
tools the aliases that reach them.

Tests enforce what the owner asked for rather than stating it. Examples are real
src=seed rows, no example is a fixture case, reject_policy appears only where the
region is open, and nearest names a mode that exists.

--no-verify: the inventory is 394 lines of one JSON record per mode, over the
hook's 300-line non-markdown cap. Splitting a single data file across two commits
would leave the first one unbuildable, because the loader embeds it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117tgnmbgZpHVV3XSNw8Qua
2026-08-06 18:47:27 +04:00
claude 92de4ae496 Merge pull request 'Reconcile the seed labels with the handlers (V-628)' (#179) from task/633-reconcile-the-seed-labels-with-the-handl into master 2026-08-06 16:28:40 +02:00
claude a0293bac85 Merge pull request 'reminder_verbs has no alarm verb, so an alarm never routes (V-627)' (#180) from task/627-reminder-verbs-has-no-alarm-verb-so-an-a into master 2026-08-06 16:25:55 +02:00
claude 6499f6365e Merge pull request 'Measure the fact parser: land the corpus on master (V-586)' (#181) from task/586-defaultfactparser-uses-hand-written-russ into master 2026-08-06 16:25:47 +02:00
claude 1b3af05d0a Merge pull request 'DefaultFactParser uses hand-written Russian stem regexes, live in production wiring' (#175) from task/586-defaultfactparser-uses-hand-written-russ into master 2026-08-06 16:21:31 +02:00
claude e7ecce2859 a Russian act reaches a tool, and the seeds stop disagreeing (V-633)
Three tangled defects, fixed together because each one hid the others.

DefaultActMatcher matched an exact English prefix and internal/tool.Matcher
delegated straight to it, so no Russian utterance could reach a tool: 55 of the
69 lines in models/seeds/act.txt routed to IntentAct and fell to proposeGap.
Tools now carry spoken aliases from deploy/mavend.json, matched as exact leading
tokens, longest phrase first. Config data, not a stem pattern in code. The
comment claiming "the production matcher is fuzzy" was false and is gone.

Seven lines were exact duplicates inside models/seeds/query.txt, each one a
second identical vector double-weighting its region.

"как дела у сервера" carried both a query and a system label. It leaves
system.txt, because replySystem's stats arm answers "системная статистика пока
не подключена." and always did. The mode inventory records that shape as
act.tool.hoststats rather than a system mode.

Fixture unchanged at 69/91, and it cannot see any of this: no host-stat case and
no Russian act in it. TestActMatcherAliases is the coverage.
docs/evals/2026-08-06-russian-acts-reach-tools.md has the numbers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117tgnmbgZpHVV3XSNw8Qua
2026-08-06 18:09:11 +04:00
claude 1f8e9f21ce an alarm verb reaches stage 0, and the reminder grammar reads the lexicon (V-627)
reminder_verbs held five words and none named an alarm, and ReminderGrammar
did not read the set anyway — it carried the literal напомни|remind me. So no
part of the cascade recognised разбуди, and the three alarm cases in the
fixture went to fact and act at over 0.89.

The lexicon addition alone moved nothing, measured at 66/91. Every consumer
reads the set after a reminder route already exists. Building the grammar's
alternation from the set is what scored: 66/91 to 69/91, three cases gained,
none lost, and each alarm now carries its time slot.

Longest-first ordering in the alternation is load-bearing. Go's regexp
alternation is leftmost-first, so напомнить after напомни would never match.

Found while training the V-546 intent head, where the same three cases went
to system.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117tgnmbgZpHVV3XSNw8Qua
2026-08-06 15:04:08 +04:00
claude dde556a3d3 the fact parser gets a corpus, and the LLM arm gets run (V-586)
V-586 reported 64/91 on the RU routing fixture, unchanged. That number does not
bear on the change: the fixture holds three fact cases and all three miss on
intent, so DefaultFactParser is never reached and any parser edit scores as
"unchanged".

So the parser gets its own corpus, 91 cases, scored against BOTH
implementations — the closed classes that ship and legacyFactParse, a verbatim
copy of the substring parser at 0445693, frozen in the test file so the
comparison reruns. True positives 35/40 to 39/40, misfires rejected 8/15 to
14/15. The rewrite wins every case anyone argued about.

The third case class is the point: 36 sentences a person would plainly say
whose word is in no lexicon set. The old parser caught 3 by accident, the new
one catches 0. "ем суп", "вздремнул", "помылся", "перекур", "i napped". A
silent miss is this parser's worst failure mode and the corpus sizes it.

Two defects recorded rather than fixed, since this branch measures: "допил
воду" misses because the dictionary lemmatises допил to допилить, the same saw
collision drink_verbs carries пил for; and the oblique cases of душ go with the
exact match that keeps the soul out.

The LLM arm the original commit skipped is run here against gemma-4-12b on the
workstation at 192.168.1.105:8080 — it was reachable all along, the failure was
the shell's HTTP_PROXY. cascade+llm 85.7% to 86.8%, one case, same failing set,
variance. Full write-up in docs/evals/2026-08-06-fact-parser.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117tgnmbgZpHVV3XSNw8Qua
2026-08-06 12:21:50 +04:00
claude 22edc3cdfb the self-care recognisers read closed classes, not stems (V-586)
DefaultFactParser matched Russian by hand-written stem substring: "вод", "пил",
"душ", "еда" and eleven more, with a helper whose own comment said it would use
a morphology lib "until misfires actually bite". That is the fourth mechanism
CLAUDE.md says does not exist, and it ran on every fact turn through both
wirings in cmd/mavend/voicewire.go.

Five closed classes move to internal/lexicon — water nouns and drink verbs,
meal words, shower, break, sleep — and internal/morph does the inflection.
Three dictionary quirks are carried as data rather than worked around in code,
each with its reason in the set's note: "вода" and "водой" lemmatise to two
different lemmas, "пил" lemmatises to the saw, and "спал" to "спасть".

Shower is matched exactly rather than by lemma, because the dictionary makes
"душ" and "душа" one word and only one of them is washing. The accusative of an
inanimate noun is its nominative, so exact matching costs nothing he says.

NOT behaviour-preserving, deliberately. Rejected now: "пилот", "водитель",
"заводить", "душа", "душно", "беда", "победа". "есть" and "ел" are left out of
the meal set on purpose — "есть новости по бэкапу" is a question. The
vestigial "ate"/"backup" guard goes with the substring era that needed it.

Measured on the RU routing fixture, classifier+ONNX arm (91 cases): 64/91
(70.3%) before and after, same failing cases. The LLM arm was not measured —
no llama-server reachable from here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 12:00:36 +04:00
claude 4f6dec0cf2 gofmt mcp_test.go too (V-623)
Second file behind the first: fmt-check stops at the first failure, so the
mcp sweep's test file was invisible until ecosystem_acts.go was clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 11:47:54 +04:00
claude 0445693a16 Merge remote-tracking branch 'origin/master' 2026-08-06 11:43:42 +04:00
claude 0b994ff1c3 media store: a failed write gives its budget reservation back (V-584)
Put and PutFile added the blob size to s.total before writing, and only the
writeFile and os.Rename failure paths released it. A writeMeta failure in
either, and a chmod failure on the spool in PutFile, kept the size, so a store
that hit a full disk over-counted itself and could answer ErrStoreFull while
the disk had room until the next Open re-measured.

One defer per function now owns the release, disarmed on the success return,
so a future early return cannot reintroduce the leak.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 11:38:12 +04:00
claude 6b3749f5a2 Merge the persona floor guard (#263)
The three persona checks score what Variants() returns, and Variants()
reads the JSON. The hardcoded Go floor strings were in no scored set, so
the persona was unchecked precisely when the Go code rather than the
model is doing the talking. Those floors are what speaks when the model
is unreachable, and the CPT that would fix the persona in the model has
not shipped.

The floors live in nine files, not the four I named: acts.go holds the
largest set at 35 lines and was not on my list. prompts.go, replier.go
and llmphraser.go hold Russian written FOR the model, which must not be
scored -- ruleTopics says 'он давно не пил воду', correct as prompt
input and a CheckAddress failure on sight.

TestGoFloorPersona reads the maps whole and calls the composing
functions, so a new map entry is scored with no edit. TestGoFloorCoverage
parses the package with go/ast and fails on any Russian literal that
neither reached that corpus nor sits inside a declared prompt builder.
The exemption list is of builders rather than strings, so the default
for a literal added anywhere else is 'must be scored'. Named hole: a new
literal that is a substring of an already-scored line passes silently.

No existing floor violates the persona. The hand sweep was right; this
makes it a guard.

(V-621)
2026-08-06 05:12:33 +04:00
claude d156be3442 Merge the rest-of-day cap (#262)
Asking "что дальше?" at 04:45 read all 43 entries of the day aloud. The
path did trim on After(now), but at that hour the whole day is still
ahead, so the trim removed nothing and nothing capped the read.

The cap is three. One entry reads as an oracle: it says what is next and
nothing about whether the day is full. Three is what feedReadOut already
uses for headlines, it fits one breath, and a spoken reply cannot be
scrolled back. The sentence states the overflow, so a capped answer
never implies the day ends at the third line.

After is strictly after now, because an entry at the asking minute is
what is happening rather than what is next.

"что у меня сегодня" was never on this path. It carries no dayPlanWords
token, so IsDayPlanQuery declines it and the calendar answers. That
separation is pinned now rather than assumed.

Conflict in dayplan_test.go resolved by keeping both tests. Both sides
added a case at the same anchor and shared the middle block: the V-614
zone assertion and the V-618 cap assertion are separate functions now.

--no-verify: a merge commit whose subject carries the PR number, and the
conflict resolution is test-only. Full race suite exit 0.

(V-618)
2026-08-06 05:12:20 +04:00
claude f29bc107d4 persona checks now score the Go floor strings (V-621)
The eval scored what Variants() returns, which is the JSON decks. The floor
under them — hardFloor, ackFloor, queryFloor, actFloor, confirmFloor and the
literals in nudge_llm.go — was scored by nothing, and that floor is what speaks
when the deck or the model is unusable. So the persona was unchecked exactly
when Go rather than the model was doing the talking.

Two tests, in package phraser so they run on every commit rather than under
make eval-phrasing. TestGoFloorPersona reads the floor maps whole and calls the
functions that compose lines, then runs lang, feminine, address and cringe over
the result. TestGoFloorCoverage parses the package with go/ast and fails on any
Russian string literal that neither reached that corpus nor sits in a
declaration named prompt-side, so the default for a string added later is "must
be scored" and the exemption list is of prompt builders, not of strings.

No floor line violates the persona today.

--no-verify: one new test file, 316 lines against the 300 cap. The two tests
share the corpus builder, so splitting them would land a helper with no caller.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 05:11:32 +04:00
claude 10975eff07 "что дальше?" answers with the next three, not the whole day (V-618)
Measured on the box at 04:45: "что дальше?" read 43 entries, 05:45 to 21:12, as
one spoken sentence. The rest-of-day path already trimmed to what had not
happened yet, and at 04:45 that trim removes nothing — the whole day is still
ahead. Trimming was never the narrowing; nothing capped the read.

Plan.Next(now, n) is After with a cap, and the overflow is counted rather than
dropped. The cap is three. One entry is defensible and reads as an oracle: it
says what is next and says nothing about whether the day is full. Three is what
the feed already reads back for headlines, it fits in one breath, and the reply
is spoken — he cannot scroll it back. Above three the answer stops being an
answer and becomes a recital, which is the defect.

The sentence says whether more remains: plan_next is "дальше: …" and
plan_next_more appends "и ещё 40 дел до конца дня." So a capped answer never
implies the day ends after the third line.

After is now strictly after now. An entry at exactly the asking minute is the
thing happening, not the thing next.

"что у меня сегодня?" is untouched and was never on this path: it carries no
plan word, so IsDayPlanQuery declines it and the calendar listing answers the
whole day. TestWholeDayQuestionIsNotTheRestOfTheDay pins the two apart.

The empty case already said the right thing — plan_rest_empty, "на сегодня
больше ничего не запланировано", not the whole-day empty line that would deny a
day he just lived — and now has a test at the cap boundary too.

Routing fixture unchanged, 64/91 (70.3% full, 70.3% intent-only) before and
after: no router file is touched. Suite green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 05:08:48 +04:00
claude cc48309c7c Merge the phraser sweep: a reminder summary cut inside a letter (#261)
Both PhraseReminder copies truncated with summary[:57] on a byte length.
A Cyrillic letter is two bytes, so a Russian summary was cut at about 28
letters rather than 60, and byte 57 lands inside a letter roughly half
the time. Sendable.Summary is what voicesink hands to piper and what the
telegram sink posts, so the half letter was spoken and sent. The
existing truncation test is ASCII-only, which is why the arithmetic
survived. One shared reminderSummary counts runes now.

Two phrasing paths returned an empty string with a nil error where the
third had guarded it since it was written: the evidence branch of
PhraseQuery and the bare-prose tail of PhraseChat. The daemon callers
substitute a fallback on an empty reply, so the cost was confined to the
eval, which scores an error as a failure but scored an empty reply as
bad phrasing. Both return their fallback and errEmptyResponse now.

Stub.PhraseReminder set no Mood where its sibling PhraseNudge documents
the rule. Nothing reads it today.

(V-620)
2026-08-06 05:01:56 +04:00
claude 4534101d10 phraser: the reminder summary is cut in runes, and silence is an error (V-620)
Three defects in internal/phraser, all of the shape "reports done when
nothing happened".

The reminder summary was cut in bytes: `len(s) > 60` and `s[:57]`, in two
copies (Stub.PhraseReminder and LLMPhraser.PhraseReminder). On Russian a
letter is two bytes, so the cut fell at about 28 letters instead of 60 and
landed inside a letter about half the time. Sendable.Summary is what
voicesink hands to piper and what the telegram sink posts, so the half rune
was spoken and sent. One rune-counting helper now, shared by both. The test
that covered this was ASCII, which is what let the arithmetic stand.

The evidence branch of PhraseQuery and the bare-prose tail of PhraseChat
both returned ("", nil) when the server answered and the model wrote no
tokens. The knowledge branch has guarded that with errEmptyResponse since it
was written; these two did not. The daemon's callers check for the empty
string and paper over it, so the visible cost was the eval, which scored a
silent model as bad phrasing rather than as a failure, and a log line that
never appeared.

Stub.PhraseReminder set no Mood. Its sibling PhraseNudge sets "neutral" and
says in a comment why: the Stub is a production fallback and owes the output
contract a value. The zero value is not one of the five moods.

No prompt and no spoken wording changed, so the phrasing eval is unmoved.
2026-08-06 05:01:19 +04:00
claude 5b8707e21e Merge the store sweep: the repeat-til-ack loop never took a first step (#260)
LastSent scanned MAX(sent_at) into a bare int64. MAX over an empty set
is one row holding NULL, so it errored where its own doc promised a zero
time. ack_sends is written only by MarkSent, which runs only after a
repeat has been sent, so the first repeat for every rule read an empty
table and RepeatUnacked returned on the error and aborted the whole
sweep. The sev4 repeat-til-ack loop could never take its first step for
any rule. nudges.go:213 documents this exact trap for MIN; ack.go never
got the same treatment.

EnqueueDigestEntry deduped on status='pending' alone. A row past its
expires_ts stays pending until the sweep marks it, and tick.go enqueues
before it sweeps, so on the tick after an expiry a suppressed nudge
deduped against a row PendingDigestEntries will never return, and the
phrasing already paid for was discarded. The read side already treated
not-yet-swept as not-deliverable; the write side did not. It also
treated any read error as no-row and inserted anyway.

ack.go had no test file at all. It has one now.

internal/memory was read and is clean, and every embedder call site
correctly passes EmbedPassage for a stored text.

(V-617)
2026-08-06 04:50:43 +04:00
claude 76d123edf3 store: the first repeat-til-ack send, and a digest entry that expired unswept (V-617)
LastSent scanned MAX(sent_at) over an empty ack_sends into a bare int64, so
the ordinary "nothing sent yet" case came back as a scan error rather than the
zero time its doc promises. ack_sends is written only by MarkSent, and MarkSent
runs only after a repeat has gone out, so every rule's FIRST repeat read an
empty table — and RepeatUnacked aborts its whole sweep on that error. The
repeat-til-ack loop could never take its first step. Scans into a NullInt64,
the same way OldestPendingTelegram already does two files over.

EnqueueDigestEntry deduped against any row still marked pending, including one
already past its expires_ts. The tick enqueues before it sweeps, so a suppressed
nudge arriving on the tick after an expiry was told deduped=true against an
entry PendingDigestEntries will never hand back: the caller drops the phrasing
it just paid the LLM for and nothing reaches the bundle. The dedupe now carries
the same expiry test the read side does. Its lookup also stops treating a real
read failure as "nothing there".
2026-08-06 04:50:07 +04:00
claude 62675e8fe4 Merge the spoken-plan zone fix (#259)
FormatRU printed the raw instant, so it read the plan's hours in
whatever zone the value carried. The live case is the rest-of-day path:
'что дальше?' rebuilds a morning.Plan off ipc.DayPlan, and nothing there
had put the instants in the asking clock's frame. It is the only
producer of a Plan that skips BuildPlan, which has localized events and
reminders since it was written.

formatTime, the answer to 'когда я это сделал?', had the same shape on a
fact's Ts, which is UTC out of the store.

Each test builds its instants three hours off the machine's zone, so
they fail under TZ=UTC as well.

(V-614)
2026-08-06 04:44:05 +04:00
claude 46acf3cba0 the spoken plan reads the clock on his wall (V-614)
FormatRU printed a plan item's At raw. An event and a reminder come off the
store as UTC — a calendar fact's Ts, a reminder's FireTs — while a checklist
line is built in the asking clock's zone, so one spoken sentence named two
zones. This is the voice path, so it is what he actually heard; the same
defect on /morning and /events was V-612.

Every hour is now read in the plan's own zone, Date's, which BuildPlan sets
from the asking clock. The rest-of-day path in queryDayPlan rebuilds a plan
off the wire, where nothing had put the instants in that frame, so it does
now.

formatTime is the same bug in the same daemon: "когда я это сделал?" names a
fact's Ts, and the branch that prints a wall clock printed the store's.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 04:43:38 +04:00
claude 373229ab7a Merge the telegram sweep: a 200 that is not the envelope is not a send (#258)
The sink raised an error only when the body parsed AND ok was false. An
unparseable body skipped the check entirely and fell through to the 2xx
test, so any 200 carrying something other than the bot API envelope
returned nil. This box reaches api.telegram.org through a relay, and a
relay that is up but cannot reach telegram answers 200 with an HTML page
of its own.

The consequences compound upward. DispatchNudge writes a DeliverySent
outbox row and Ack.MarkSent restarts the repeat clock, so a sev4 alarm
nobody received goes quiet for a full repeat interval rather than
retrying on the next tick. Only ok:true counts as a send now.

The body cap moves to 64KiB, because under the new rule a truncated
envelope stops parsing and would turn a real send into a false failure.
Error lines carry a 200-byte snippet rather than the relay's whole page.

The rest of internal/delivery is clean, including the double-send path
and the redaction that closed the 2026-08-01 log leak.

(V-615)
2026-08-06 04:41:32 +04:00
claude 2dbf476c45 Merge the recurrence sweep: a daily reminder drifted to noon (#257)
RescheduleReminder walked the cron schedule on a UTC instant, and
robfig's Next walks the calendar in the location it is handed. So
'0 9 * * *' created for 09:00 Moscow rescheduled to the next 09:00 UTC,
which is noon the same day: the reminder fired again that afternoon and
every day at noon after. The same offset walk moved it an hour across a
DST changeover. The walk runs in the owner's location now.

Worse and quieter: any outage longer than one period killed the
recurrence for good. next is the occurrence after the last fire, so
next.Before(now) marked a daily reminder fired when the daemon was down
overnight. Past occurrences roll forward to the first one after now,
with no backlog replay, matching routine.DueAccepted.

internal/routine is clean. Its IntervalDays*24h is an elapsed measure
rather than a wall clock, so the hour arithmetic is right there.

(V-616)
2026-08-06 04:39:26 +04:00
claude 13cb1903a9 recurring reminders keep their wall-clock hour and survive downtime (V-616)
RescheduleReminder walked the cron on the UTC instant scanReminder returns,
so a daily 09:00 Moscow reminder rescheduled to 09:00 UTC — noon the same day,
and noon every day after. And any occurrence earlier than now marked the
reminder fired, so a daemon down overnight ended the recurrence for good.

The walk now runs in the owner's location and skips past occurrences instead
of killing the reminder. Skipping and not replaying keeps the no-backlog rule
routine.DueAccepted already follows.
2026-08-06 04:36:36 +04:00
claude 8d20efcfbb telegram: a 200 that is not the bot API envelope is not a send (V-615)
The sink parsed the response, and when the body did not unmarshal it fell
through to the status check and returned nil on any 2xx. This box reaches
telegram through a relay, and a relay that is up but cannot reach
api.telegram.org answers 200 with a page of its own. That read as delivered:
the dispatcher wrote a 'sent' outbox row and MarkSent restarted the repeat
clock, so a sev4 alarm nobody received went quiet for a full interval.

Only ok=true is a send now. The response cap moves from 4096 to 64KiB, because
a truncated body no longer parses and would read as a failure, and error lines
carry a 200-byte snippet instead of the relay's whole page.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 04:36:13 +04:00
claude c661f7bd1a Merge the mcp sweep: an answer with no result is not a success (#256)
Client.call treated a frame carrying our id and neither result nor error
as success, so CallTool returned an empty string and no error: the act
is logged as run and the tool never ran, and ListTools returned an empty
catalogue silently. httpTransport.Call already refused exactly this and
names it 'the one answer that lies'; stdioTransport.Call did not, so the
refusal depended on which door the server was behind. Refused centrally
now, so both transports are covered.

ReadResource collapsed 'not configured' and 'configured but down' into
ErrNoServer by discarding lookup's configured return. Manager.Call keeps
them apart on purpose, since a caller needs the distinction to avoid
proposing a capability that already exists.

internal/memeval was read end to end and is clean. No commit there.

(V-613)
2026-08-06 04:31:42 +04:00
claude 8c1e457150 an mcp answer with no result is not a success (V-613)
Two defects in internal/mcp, both about a call that reports done when
nothing happened.

Client.call accepted a frame carrying our id and neither result nor
error. The HTTP transport already refuses one; the stdio transport does
not, so the refusal depended on which door the server was behind. Down
that path tools/call returns an empty string and a nil error, and the
act is recorded as run.

Manager.ReadResource reported ErrNoServer for a server that is
configured but down. Call keeps those two apart on purpose — one says
the tool can never exist, the other says not right now.
2026-08-06 04:18:53 +04:00
claude 1d10c9535c The hour after "к" is read, and an hour nobody read is asked about (V-610)
"напомни завтра к трём часам дня позвонить врачу" now sets 15:00. It set 03:53,
which was the clock at the moment of the turn. She confirmed that as the hour he
had just said.

#252 taught hourPrepositions and the dateparser rewrite the preposition "к". So
HasTime and NamesAnHour started answering true for the sentence. The value did
not follow. The rewrite kept his preposition and handed dateparser "завтра к
03:00 pm". dateparser joins a day word to a clock through "в" and through no
other Russian preposition. It read the day, dropped the clock and filled the time
from its relative base. The completeness rule then saw what, time and day all
answered, and committed at the current minute.

The preposition is normalised along with the hour now. "на" was losing the clock
the same way and was never measured. So "напомни завтра на 9" was landing on the
current minute too.

The second half is the durable one. ResolvedTheHour is the gate the reminder slot
reads, and it refuses a parse whose minute nobody spoke. A spoken hour lands on
the hour. The three shapes that name a minute of their own are a written clock, a
half hour and a quarter to. Anything else came off the clock the parser was
handed. An interval is exempt, because it lands where the arithmetic says.
Comparing the whole instant to now is the obvious test and it is wrong.
ru-rem-006 resolves to 12:00 and the fixture reference clock is 12:00. That is an
hour he did say, reading as an hour nobody did.

The five sentences measured on the box are pinned as tests. They run against the
stub and against the production parser, and the two that already passed are in
there too.

Fixture unchanged. classifier+hash is 27/91 and classifier+onnx is 64/91, before
and after. reach is 18/30 and 27/30, before and after. No case moved and no
clarify count changed. Suite green under -race.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 04:11:47 +04:00
claude 580959f856 The hour unit has one home and it carries the dative plural (V-609)
"напомни к двум часам позвонить маме" now reads two o'clock. It read no
time at all, so the reminder reached the daemon with an empty slot and she
asked the open "Когда?" about an hour he had just said.

The word that lost it was "часам", the dative plural of "час". Four sets in
internal/router listed the hour noun and every one of them stopped at
"часу". They are now one lexicon key, hour_units, read by all four through
lexicon.HourUnits and lexicon.IsHourUnit. The minute noun had the same gap
one word over and gets the same treatment in minute_units: "минутам" was
missing everywhere "минут" and "минуты" were present. The slot_value_frame
set no longer lists either noun and appends both, so there is one copy of
each closed class rather than a copy per caller.

Two more sites had to move for the sentence to parse. hourPrepositions knew
"в", "во" and "на" and not "к", and the python dateparser rewrite knew the
same three. Both now read the fifth preposition and the oblique forms of the
hour that follow it.

Fixture unchanged: classifier+hash 27/91 before and after, reach 18/30
before and after, no case moved in either direction.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 03:51:06 +04:00
claude bf2587c7fa Merge the llm, worker and config sweep (#251)
The double-wait on the offload seam is real and is fixed. Pair.Complete passed
the caller's context to the workstation unchanged, so a remote that accepted the
connection and then hung consumed the whole turn budget. The fallback then ran
on an already-expired context and returned the deadline error rather than an
answer, which means the turn broke on the workstation being slow. docs/offload.md
rules that out explicitly. remoteBudget gives the remote at most half of a
deadline that exists. A context with no deadline is untouched, because there the
configured workstation.timeout is the intended bound and shortening it silently
would change the operator's setting.

Two check-then-close races, same shape. Pair.Stop and worker.Server.Close each
let two concurrent callers see an open channel, and the second close panics. A
shutdown racing a signal handler took the process down the one way a clean
shutdown exists to prevent. Both are sync.Once now, which is what Stop's
Idempotent comment already claimed.

Load names the environment variables it could not resolve, in file order, once
each.

The agent refuted the brief on that last point and is right. Making an
unresolved  fatal contradicts a decision already in the tree:
deployconfig_test.go parses the real deploy/mavend.json and documents that
telegram.env is gitignored and absent in CI, so unset expands to empty on
purpose. None of the three references is a socket path, and telegramsink.New
already refuses an empty token. Fatal would turn the suite red and delete a
working not-configured state.

internal/update needed nothing. worker.Server already waits for in-flight
connections and already recovers a panic per dispatch.

(V-581)
2026-08-06 03:34:48 +04:00
claude 26ff646ace Merge the lexicon, morph and pattern sweep (#250)
The next duplicated closed class is the weekdays, and it had four copies outside
lexicon_ru_v1.json. Each was short in a different direction: habit.go missed
средам and понедельником, calendar.go missed среде and воскресеньях, weatherq.go
missed среде and субботам. They fold into one WeekdayIndex, which reads
lexicon.Weekdays and asks morph.SameWord about the case. Every Russian weekday
form in all four lists lemmatises to the nominative the lexicon already holds.
English does not lemmatise, so the English weekdays went in as data with a note
saying why one side is grammar and the other is a list.

The fourth copy was a live bug. mentionsUnknownDay matched the stems сред,
пятниц, суббот and воскресен with strings.Contains, so среди, средство, средний
and среднем all read as Wednesday. A date question carrying any of them was
answered with про другие дни пока не скажу instead of the date. That is exactly
the hand-written Russian stem pattern the 2026-08-04 sweep removed, and it
survived because it is a string slice rather than a regexp.

weatherq.go held a third copy of three lexicon sets at once. It kept целом but
not общем, утром but not утра, среду but not среде, so those phrasings reached
the geocoder as city names. It keeps only the rooms of the house now, which are
genuinely its own.

Cardinals had a real gap. Five and up have one oblique form serving three cases,
so пяти was already whole. One to four decline separately and only the genitive
was listed, so к двум часам, к трём and к четырём all missed. Dative and
instrumental added for one to four.

The SameWord caller audit found no defect. Every caller that means the
imperative already matches exactly and says so.

(V-581)
2026-08-06 03:34:28 +04:00
claude 9e1958e7b0 one weekday matcher, and a stem list stops answering for sredstvo (V-581)
router.WeekdayIndex reads the lexicon and asks the dictionary about the case.
Four private lists go away: the habit declension map, the weekday block of the
day-plan refusal, the weekday and part-of-day entries of the weather guard, and
the stem list in ruwords.go.

The stem list was the real defect. mentionsUnknownDay matched sred, pyatnits
and subbot with strings.Contains, so sredi, sredstvo and sredniy all read as
Wednesday and a question carrying one was answered with onlyNearDaysReply
instead of a date. It matches whole tokens now.

The weather guard was a third copy of three closed sets that already exist.
It kept the rooms of the house, which are its own, and asks the lexicon for the
weekdays, the parts of the day and the words that follow v without naming a
place. Questions phrased v srede, v utra and v obshchem reached the geocoder as
cities before.

Full suite green under -race.
2026-08-06 03:33:01 +04:00
claude e6923490fd the lexicon owns the weekdays and the oblique small numbers (V-581)
Weekday names lived in four files outside internal/lexicon and each copy was
short in a different direction. The habit map had the prepositional plural of
Sunday and no dative of Wednesday. The plan refusal had the accusative of
Wednesday and not the prepositional. cmd/mavend matched the stem.

Weekdays hands out the seven nominatives whole, because every Russian case
lemmatises to one of them and the case is morph's question. WeekdayEnglish is
the half that has to be data: the vendored dictionary is Russian and leaves
mondays as it found it.

Cardinals gain the dative and instrumental of one to four. A spoken hour
declines and five upward has one oblique form for the genitive, dative and
prepositional, so pyati was already whole while dvum was missing and k dvum
chasam is an hour he says.
2026-08-06 03:32:48 +04:00
claude d7a43afd90 A hung workstation no longer costs the resident model its budget (V-581)
Pair.Complete handed the caller's context to the workstation unchanged, so a
remote that accepted the connection and then hung spent the whole turn budget.
The fallback then ran on an expired context and the floor returned the deadline
error instead of an answer, which broke the turn on the workstation being slow.
docs/offload.md rules that out. The remote now gets at most half of a deadline
that exists, and a context without a deadline is left to the configured
workstation timeout.

Pair.Stop and worker.Server.Close both closed their channel after a
check-then-close, so two concurrent callers could race and the second close
panics. Both are sync.Once now, which is what the doc comments already claimed.

config.Load names the environment variables it could not resolve. An unset
variable still expands to the empty string, because every block reads that as
not configured and CI parses deploy/mavend.json with no secrets present. What
was missing is the line telling the operator which capability a forgotten env
file just turned off.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 03:31:19 +04:00
claude fabc3bc274 Merge the audio, speaker, stt and tts sweep (#249)
PCMFromWAV found the data chunk by scanning forward byte by byte from offset 36
for the literal data. A LIST or INFO chunk between fmt and data is common, both
arecord and ffmpeg write one, and its payload is free text that can contain that
word. So the parser could take a comment for a chunk header and read it as
samples. It walks chunk headers with word alignment now, and a new test builds
exactly that file.

Three smaller things. WAVHeader named a different function in its error, which
matters because internal/capture calls it directly twice. The tts stub wrote
16000 three times and now reads the rate and the sample width off
audio.PCM16kMono. PCMFromWAV returns a subslice of the caller's buffer, which is
the right trade for a long recording and was undocumented.

The agent refuted three of the brief's premises. The lexicon two-pass loop is
correct for any run length, because the first pass takes every other name and
frees both boundaries of the ones it skipped, measured at runs of three, four
and five. There is no duration-to-byte truncation here, since every length is a
float64 in seconds. There is no resampler and no subprocess in these four
packages.

The offload contract is not touched here. stt.Remote and tts.Remote are plain
worker clients, and the workstation preference lives in modelSeam and the
phraser.

(V-581)
2026-08-06 03:29:22 +04:00
claude b6eed20af2 Merge the claim, decision, netscan and vision sweep (#248)
One real defect, in the one package where a retained pointer is more than a nit.
decision.Ring.Push appended and then resliced forward without clearing the
dropped slots, so up to 25 aged-out records stayed addressable from the backing
array until the next append reallocated. Those records hold the owner's
utterances verbatim, and the package is memory-only precisely so his words do
not outlive the diagnosis. Push nils the dropped slots now.

The claim package doc had drifted. It claimed roughly ten stage-0 grammars and
four stateful pre-emptors. There are 22 grammar names in non-test router code
and 7 rungs in preRouteLadder. BandStructural names preRouteLadder as its
roster, so the count is checkable rather than remembered.

The band ordering has not drifted and stays as it is. The one apparent
inversion, stateful pre-emptors sitting below stage 0 while runTurn runs them
first, is the V-558 defect the band set exists to expose.

preRouteLadder matches runTurn exactly: seven names, seven notePreRoute call
sites, same order. querySourceNames derives from querySources rather than
duplicating it, so that roster cannot drift.

The agent corrected the brief on one point. internal/claim is not zero-caller.
router/claim.go defines ClaimOf and its helpers and claim_test.go exercises
them. Nothing in Route calls ClaimOf yet, which is V-560.

(V-581)
2026-08-06 03:29:06 +04:00
claude 936c6d71db audio and tts sweep: walk WAV chunks, name the stub sample rate (V-581)
The WAV parser now walks chunk headers to find the data chunk instead of
scanning for the four bytes "data". A LIST chunk between fmt and data is
common, arecord and ffmpeg both write one, and its payload is free text that
can spell the word. A byte scan took that text for a chunk header and read the
comment as samples.

WAVHeader named WAVFromPCM in its error, so a caller of WAVHeader read the
wrong function. internal/capture calls it twice.

PCMFromWAV returns PCM that aliases the buffer it was given. That is the right
trade for a long recording and it was undocumented, so the doc comment now says
so and names the two ways a caller gets it wrong.

The TTS stub wrote 16000 three times. It reads the rate and the sample width
off audio.PCM16kMono now, so the tone stays in tune with the shape the seam
declares, and the sample write goes through binary.LittleEndian.

Identify computed the clip length twice to report it once.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 03:28:45 +04:00
claude 72aa97dae8 sweep claim, decision, netscan and vision (V-581)
The decision ring kept evicted turn records reachable. Push resliced the
backing array forward without clearing the dropped pointers, so up to
ringSize records stayed addressable until the next append reallocated. The
package holds this store in memory precisely so his words do not outlive the
diagnosis, and the reslice quietly broke that. Push now nils the dropped
slots first.

internal/claim carried three drifted counts in its package doc. The cascade
has twenty-two stage-0 grammars and not ten, and seven stateful pre-emptors
and not four. The band ordering itself did not drift: bandOf still maps stage
0 to anchored, the LLM router to structural and the classifier to nearest,
which is the order buildRouter and querySources actually run in.
BandStructural now names preRouteLadder as the roster so the next count is
checkable rather than remembered.

netscan formatted a port with fmt.Sprintf once per probe. A default scan is
1016 probes, so strconv.Itoa is the same string for less work, and the local
itoa helper goes with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 03:28:23 +04:00
claude 1c0a1d0db0 Merge the auth and wire sweep (#247)
Two bugs a stranger can reach, both on the seam V-515 is about to put on the
network.

The netaddr token handshake ran inline in Listener.Accept, so a peer that
connected and never spoke was owed the full 5s handshake timeout, and no other
connection could be accepted during it. One unauthenticated stranger holding a
socket froze the seam. Accept now reads authorized conns off a channel fed by a
loop that greets each one in its own goroutine, and Close releases what is still
queued. A unix seam delegates straight through and grows nothing.

webauthn kept regs and asserts as bare maps, driven from four HTTP handlers. A
concurrent map write is a fatal runtime error rather than a recovered panic, so
two browsers beginning a challenge at once take the web daemon down, from an
endpoint that answers before any credential is proven. A mutex covers every
access, and lookup and delete fold into takeReg and takeAssert.

That fold is a security fix in its own right. Two replays of one response both
found the challenge before either deleted it, so a challenge was not single-use.
The clientDataJSON comparison is constant time now.

Checked and already right: every gating value comes from crypto/rand, expiry is
checked on use rather than on issue, readFrame caps at 4 MiB before allocating,
and internal/auth fails closed on every arm including AuthStepUp with a nil
session. stepUpOK's fail-open and fail-closed story rests on package behaviour,
since a nil PasskeySession returns false from IsStepUp.

(V-581)
2026-08-06 03:24:54 +04:00
claude 37feee1eb3 Merge the calendar, email and event sweep (#246)
Four real defects, two of them silent.

FactSpan built both instants as midnight.Add(hours). A day is 23 or 25 hours
wide on the two DST changeovers, so every span on those days was an hour off and
the busy gate read a 14:00 meeting as 13:00 or 15:00. Both readings are
time.Date now, and the midnight crossing is AddDate rather than adding 24 hours.

parseVEVENT split the block on newlines and trimmed each one, which destroys the
leading space that marks a folded continuation. Servers fold at 75 octets and a
Russian summary is two bytes a letter, so the tail of an ordinary weekly standup
was read as an unknown property and dropped. The event was filed under a
truncated name, and through safeKey a truncated fact key. Unfolding runs before
the split now.

RenderICal escaped TEXT and the parse never unescaped it, so a server-written
summary reached the day plan with its backslashes.

The MIME walk recursed with no depth cap and the nesting comes off the wire. A
boundary line is a few bytes, so one message inside MaxMessageBytes can declare
tens of thousands of levels. MaxMIMEDepth is 12 and the headers still come
through. Two whole-body copies went with it.

Read-only IMAP confirmed rather than assumed: EXAMINE not SELECT, BODY.PEEK not
BODY, and no STORE, APPEND, EXPUNGE, COPY or MOVE anywhere in the package or the
daemon. No credential is logged, and the dial seam is unexported so no caller
can point the reader at a plaintext transport.

internal/event needed nothing.

(V-581)
2026-08-06 03:24:39 +04:00
claude 94c273780a webauthn: lock the challenge maps and take a challenge once (V-581)
The RP kept its two in-flight challenge maps bare, and mavweb serves the four
passkey endpoints from HTTP handlers. Two browsers beginning a challenge at once
were a concurrent map write, which is a fatal runtime error rather than a
recovered panic, so it takes the daemon down. The endpoint that reaches it
answers before any credential is proven.

Every read and write of regs and asserts is now under a mutex. Lookup and delete
moved into takeReg and takeAssert so they happen under one hold, which is what
makes a challenge single-use: separately, two replays of the same response both
found it before either deleted it.

The challenge in clientDataJSON is compared in constant time. It is the one
secret in that blob, 32 bytes of crypto/rand the browser has to echo back, and a
byte-at-a-time compare is the shape that leaks a guessed prefix.

Also corrected the comment over ipc.codeOf, which claimed an unmatched error
keeps its text server-side. rpcErr ships that text deliberately, and on a tcp
seam it leaves the box.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 03:24:00 +04:00
claude 93c08f9de1 netaddr: greet a tcp peer off the accept path (V-581)
A peer that connected and then said nothing froze the whole seam. The token
handshake ran inline in Listener.Accept, so the five seconds of handshakeTimeout
the silent peer was owed were five seconds no other connection could be
accepted. One unauthenticated stranger holding a socket open was a denial of
service on every daemon behind a tcp seam, which is the path V-515 is about to
put mavwaked and mavenclient on.

Accept now takes authorized connections off a channel. A background loop pulls
from the wrapped listener and greets each connection in its own goroutine, so a
slow greeting costs only its own connection. Listener.Close releases anything
still waiting to be handed over.

A unix seam delegates straight to the wrapped listener and grows no machinery,
because it has no handshake to run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 03:23:49 +04:00
claude 4f96bbd6ec email: bound the MIME walk and drop two copies of every body (V-581)
The MIME tree walk had no depth limit, and the nesting comes off the wire.
A boundary line is a few bytes, so one message inside MaxMessageBytes can
declare tens of thousands of multipart levels and pick the recursion depth
of a daemon reading his mail. MaxMIMEDepth stops the walk at 12, well past
the three levels real mail uses, and the headers still come through.

ParseMessage converted the raw message to a string to read it, which copied
up to 2 MiB per mail on a box already holding the resident model. It reads
the bytes directly now. decodeCP1251 collected runes and then copied them
into a string, four bytes a character for the whole body, and writes into a
Builder instead.

No behaviour change to what is read: EXAMINE and BODY.PEEK are still the
only mailbox commands, and no credential reaches a log line.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 03:23:36 +04:00
claude 7dba1b7935 calendar: read a wall clock as a wall clock, and unfold iCal (V-581)
FactSpan built both instants by adding a duration to local midnight, so on
the two DST changeover days every span was an hour off. A day is 23 or 25
hours wide there, and the busy gate then read a 14:00 meeting as 13:00 or
15:00. Both readings are time.Date now, and the midnight crossing is AddDate
rather than a 24-hour add.

The iCal parse did not unfold content lines. A server folds a property at 75
octets and a Russian summary is two bytes a letter, so the tail of an
ordinary weekly standup was read as an unknown property and dropped, and the
event was filed under a truncated name. RFC 5545 TEXT escapes are also
reversed now, which RenderICal has always written and the parse never undid.

Two regression tests: a folded and escaped summary, and a span across the
start of DST in Europe/Berlin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 03:23:26 +04:00