Commit Graph

277 Commits

Author SHA1 Message Date
claude db0752223f Merge pull request 'Run the persona checks inside the daemon before she speaks' (#143) from task/399-run-the-persona-checks-inside-the-daemon into master 2026-08-04 18:24:36 +02:00
claude a2081d8227 router: stage 0 claims "что дальше" and "расскажи про X" (V-498)
Neither utterance carries a question mark or an interrogative, so nothing at
stage 0 claimed them and the model called both facts. The write is contained —
actions_fact refuses a question-shaped fact and re-runs the turn as a query —
but every one of these paid a full model round trip to reach a decision two
regexes can make, and the fixture scored the routing as wrong.

rest-of-day-query joins the agenda grammars: the predicate for the utterance
already existed as IsRestOfDayQuery, one layer down in the query chain, and
this is what gets the turn there. NarrativeQueryGrammar reads the same
narrativeRequests lexicon IsQuestionShaped reads, and declines the topics that
are chat rather than world questions — a joke, a bedtime story, herself. It is
wired last, so an explicit capture marker still wins.

Fixture: ru-query-024 and ru-query-025, both passing. Classifier + ONNX
baseline 56/80 (70.0%) → 58/82 (70.7%), no case regressed and no new false
clarify. The LLM arm is unmeasured here — no llama-server in this run.

The mavweb auth test posted its instant as "Z", which the #482 fix now reads in
the daemon's zone, making the clock inside the text stale by the test box's own
offset. It carries the local offset now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 11:45:02 +04:00
claude afac8fb670 mavend: run the persona checks before she speaks (V-399)
The checks stay in the eval package and the daemon calls three of them:
feminine, address, and a new leaked-reasoning test. No retry — it doubles
the latency on the turn that is already going badly, and on the nudge path
the moment has passed. A failure falls back to the deterministic floor and
is logged with the whole rejected text and counted by check name.

hisgender is deliberately not run: the simulator showed it rejecting
"записала, что ты выпил воды", which is her own correct self-reference.
2026-08-04 04:35:42 +04:00
claude 77206f298e router: answer the task-list ask at stage 0 too, and mirror it in the fixture (V-467)
The capture half landed with the grammar in 87d1761. This is the exposure
the task asked to check for: IsTaskListQuery is a deterministic lookup that
only runs once the turn is already a query, so a phrasing the model calls
system never reaches it. The eval fixture was also missing both grammars,
which is only worth having while it is the daemon's grammar set.
2026-08-04 04:29:37 +04:00
claude a1f811d4c8 mavend: he can pick one by position (V-448)
Read before routing and only when a list is bound: with nothing offered,
"второй" is an ordinary word and keeps routing. No verb reads it back
rather than guessing what to do with it.
2026-08-04 04:26:10 +04:00
claude 6a85e71077 mavend: wire the correction into the turn, before routing (V-455)
Read next to the confirm and clarify turns, because a correction routed as
a fresh utterance files the correction itself. Only turns she acted on are
remembered: a clarify asked instead of acting.
2026-08-04 04:21:03 +04:00
claude bf6c2bf1a6 mavend: read a spoken correction of the previous turn (V-455)
CorrectMisroute has been in the router since it was written with no caller
outside a test. repair.go is the half that reads the words: a marker saying
she was wrong plus the intent it should have been, with the negated half
skipped, and it teaches the classifier and redoes the request under the
corrected intent.
2026-08-04 04:21:03 +04:00
claude 5afa2dfb38 mavttsd: a pronunciation dictionary, so she says the names right (V-458)
piper reads a Russian sentence with a Russian voice, and a Latin service id
inside it comes out spelled, mangled or read as if it were a Russian word:
"Vikunja", "SearXNG", "homesrv". The lever available is the text, so the
dictionary maps a name to how it should be spelled for the voice to say it,
and mavttsd applies it at the last edge before piper — every caller's text
passes through that one point, and nothing upstream has to know how a name
sounds.

Data, not code. deploy/tts-lexicon.json ships 29 names; adding one needs a
restart of mavttsd and no rebuild of the daemon that produced the text. Off
unless -lexicon is set, like every other optional capability, and a path that
is set and unreadable stops startup — saying names wrong in silence is the
failure it exists to remove.

Two details worth keeping: the alternation is sorted longest-first, or "Home
Assistant" reads as "Хоум Assistant"; and the boundaries are written out
rather than left to \b, which is ASCII-only and never fires next to a
Cyrillic letter.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 04:11:10 +04:00
claude 0f361d2034 mavend: answer what he told her, from the facts he already tapped in (V-456)
Read-only over rows that exist. No new mechanism and no new storage: every
fact he tapped in already carries a source and a timestamp, and the history
source only reads them back.

Only "tap:" sources, and only the last day. A fact written by a poller, an
inference or the ambient relay is a thing she learned rather than a thing he
said, and reading those back under "что я тебе говорил?" would put words in
his mouth. Five at a time, which is what fits in one spoken breath — the rest
are on /history, which is the surface for reading a list.

Above the recall sources, with the others that read his own rows: the notes
pass would otherwise answer this from whatever note is nearest, which reads
as an answer and is not one. The matcher wants both halves of a history
phrase and steps aside when he names a topic, so "что я говорил про сервер"
stays a recall question.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 04:07:43 +04:00
claude 19242ef73b mavend: a clarify asks differently the second time (V-457)
Every clarify turn said one sentence per gap, and a re-ask repeated it word
for word. A question he already failed to answer is the worst one to ask
again unchanged: the second wording is what tells him which part she missed.

clarifytemplates.go holds three wordings per slot, picked by attempt rather
than at random — short first, then naming the gap, then spelling it out with
an example. Past the end she keeps the most explicit one instead of wrapping
back to the short question he has already not answered.

The intents with nothing identifiable to ask about (note, query, chat,
system) kept the stub's single "не совсем поняла — можешь переформулировать?",
which is the line he hears whenever she misses him completely. Four wordings
now, picked by a hash of the utterance so one question asked twice reads the
same and two different misses do not.

Still no model call on this path: the resident model would wander, and this
text has to be right every time. No schema_version either, unlike the nudge
templates — these are Go constants, so no file can drift out of step with the
code that reads it. The persona test already in clarify_test.go covers the
new lines.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 04:04:59 +04:00
claude 26d6d71588 mavend, router: stop three sources claiming turns the world should answer (V-474)
The querySources order predates the 2026-08-02 ruling that live search leads.

An unconfigured feeds source claimed every news question and answered with a
configuration status, so "что происходит сейчас в новостях про искусственный
интеллект?" never reached the search sitting one source below. It now claims
only when neither SearXNG nor the ZIMs are configured, which is the case the
"не читаю ленты" line was written for — general knowledge would otherwise
invent a bulletin.

The calendar matches on a day word alone and sits above the weather, so
"какая сегодня погода в Москве?" answered "на 02.08.2026 ничего нет." It now
steps aside on weather wording, the same bail-out queryHome already does.

"что нового в лентах?" routed system and answered "пока не умею", while the
same question worded with "новостях" worked. FeedQueryGrammar routes it to
query at stage 0, requiring an ask word and a feed noun so the bare greeting
"что нового?" stays a greeting. Wired in the eval too, since the fixture is
only worth anything while its grammar set is the daemon's.

Also: the claiming source is now logged. /trace is the nudge-rule trace and
carries no query-source field, so a wrong answer could not be told apart from
a wrongly-ordered chain.

Kiwix having no live coverage is filed separately as V-508 — it is a decision
about search quality, not an ordering fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 04:00:58 +04:00
claude df58710141 reminders: store what she says, show it in his clock (V-469)
Two of the four defects on the task.

The stored payload was the whole utterance, so /reminders and the agenda
recited "напомни завтра в 9 утра выпить таблетки" where the reminder is
"выпить таблетки". The marker is an instruction that was already carried out
and the hour is already a column, so reminderBody strips both, and falls back
to the unstripped body whenever stripping would leave nothing — a reminder
that fires and says nothing is worse than a wordy one.

The page rendered the raw {"text":...} envelope and the UTC instant. Both are
now done in mavweb: reminderRows unwraps the payload and formats through
Local(). The unwrap is a copy of store.ReminderText rather than a call to it,
because mavweb builds without CGO and internal/store carries the sqlite
driver — the ipc DTOs are decoupled from the store on purpose.

TestClarifySubjectAnswerFillsRatherThanClobbers asserted the hour survived as
a word in the payload. It now asserts the fire time, which is where the hour
lives.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 03:53:20 +04:00
claude e481ad4930 mavend: a question about attention reaches Praxis (V-475)
The capability was built, wired and degrading correctly, and no utterance
could reach it. Its aliases sit on the act dispatch, "что требует внимания"
routes to a query, and every query source passed — so the turn fell to the
web search and came back with an article about the concept of attention.
That reads as an answer, which is worse than silence.

queryAttention sits next to "tasks", above the recall sources and well above
the personal boundary: it is operational state about his things, and a notes
pass would otherwise answer from whatever he once wrote about a server. It
calls the same handler the act path calls, so the outage string comes free.

An absent or unconfigured Praxis falls through instead of claiming the turn,
like queryHome and queryNetwork. A configured Praxis that is down claims it
and names the gap. "что нового" is left to the feeds source.
2026-08-04 03:44:48 +04:00
claude 3d8224fb04 router, mavend: a complaint about a thing is not a fact about him (V-481)
"сеть какая-то медленная" and "интернет не работает" were written as `self`
rows at confidence 1.00. Recall reads a self row back later as if it were
still true, and that is the class of row that outranked live search in #470 —
so a slow afternoon becomes a standing belief about his network.

IsTransientComplaint is the same shape as IsQuestionShaped: deterministic,
offline, and off by default in the two cases where losing a real capture would
cost more than keeping a complaint. An explicit "запомни ..." wins, because he
asked. A first-person marker wins, because "я сломал руку" is durable and the
test is meant for sentences about things.

She answers the turn as chat instead of storing it. actionChat now has the
same nil-phraser floor the other model callers have.

The second defect filed here — a reply body of literally "{" — was closed by
the errBrokenJSON path in V-397 and needs nothing further.
2026-08-04 03:42:07 +04:00
claude 990a4a99e9 ecosystem: ask Nexus for the name he actually said (V-476)
Two defects in one logged line, both of which put a working capability out
of reach of every utterance.

The resident model rewrites as it routes, and on the way it transliterates:
"перезапусти muzick indexer" came back as "перезагрузить музик индексер", so
Nexus was asked to resolve a service nobody has ever named. entityReferenceText
takes the longest Latin run out of his own words, but only when the Text slot
has lost every Latin letter the utterance had — an English turn and a Russian
entity name are both left alone, and reversing the transliteration is not
attempted.

The second half: the stage-3 gate thins an act that matched no allowlisted fn,
and that question was the whole turn, so handleHexisAct never ran. Hexis is
where an act with no local fn belongs, so it gets one chance before she asks,
and a "" back still leaves her asking. With no ecosystem wired nothing changes.
Capability matching reads the phrase as the haystack when there is no fn,
because no capability name contains "restart status muzick indexer".

Authority is untouched: ambiguity still stops, a mutating capability still
goes through the spoken confirm.
2026-08-04 03:38:20 +04:00
claude 5b622389c5 dialogue: a restart expires the parked question (V-385)
The decision, not a behaviour change: ClarifyStore stays in memory, and she
does not announce the loss either.

The TTL and the attempt count measure a pause in one conversation. A restart
is a gap of unknown length, so a restored question is either dead already or
lying about its age, and the request behind it is one he has likely given up
on. Announcing it would mean storing a marker that outlives the thing it
describes, to say one sentence in the rare window where he speaks within 90s
of a restart. His next words route fresh, which is right either way.

Written down in docs/design.md, pinned at both ends by a comment, and held by
a test that builds a second handler over the same store.
2026-08-04 03:33:19 +04:00
claude ea9c746852 weather: any city he names, not the six in a table (V-421)
The hand-written table understood "какая погода в X" for six values of X.
Ask about Kazan or Tbilisi and the city was dropped silently and answered for
the default location — a correct-sounding answer about the wrong place.

The table is gone. internal/weather already calls Open-Meteo's geocoding
endpoint on every lookup, so the place he named goes straight there and any
place it knows is a place he can ask about. He speaks the prepositional case,
so locationCandidates reverses the two endings that cover most of it: a final
"е" is a nominative "а" or nothing, a final "и" is a soft sign. A wrong
candidate finds no city; it never invents one.

A place the geocoder does not have now reads as "не знаю такого города"
rather than as a provider outage or, worse, as the default city's weather.
ErrLocationUnknown is what carries that apart.

"в" followed by a room or a day word is still the default location. Those
questions are answered by the house sensors and the calendar, not by
Open-Meteo, and they must not be read as a city.
2026-08-04 03:25:42 +04:00
claude d60a51c9e7 store, mavweb: the delivery outbox can be read (V-390)
The table was write-only. Rows were recorded and nothing could show them, so
the tests for #368 and #370 had to reach past the store into store.DB — if a
test can only see it that way, so can nobody else. A durable record nobody
reads answers no question, and why Maven went quiet is supposed to be a query.

ListDeliveryAttempts returns recent rows newest first, filtered by status.
Status is the filter worth having because the two real questions are "what got
dropped" and "what is still pending", and neither is answerable by reading the
whole list on a busy day. It reaches mavweb over IPC as DeliveryAttempts.

The section goes on /notifications, which already answers "what did she send",
rather than on a page of its own. Shared ui.css, the nav partial, the table in
div.scroll. A failed outbox read leaves a log line and still renders the nudge
list, because half the page beats none of it.
2026-08-04 03:22:34 +04:00
claude 2815adee03 morning: an item can be optional, so a skipped stretch is not a skipped pill (V-473)
Item carried only Key, FactKey and Label, so every checklist entry was
implicitly required and behaviour 1 of #280 could not hold at all. It was not
thin config — there was no field to set.

Item.Optional, `"optional": true` in the routine config, default false, so a
routine written before today behaves exactly as it did. Due now fires on a
missing required item and not on an optional one, and the optional stragglers
still travel in Missing so the one message per day per routine can name them
after the required ones, in softer words.

Evidence, the window and the day plan treat both kinds alike. A missing
optional item is still missing — it just does not earn a nudge, because a
checklist where everything is mandatory is one he learns to ignore.
2026-08-04 03:17:29 +04:00
claude 9f51596e2f mavend: the simulator routes on the seeds the deploy loads (V-465)
The seed path was relative to the working directory, which is cmd/mavend
under `go test`. Every open failed, and the three scenarios replayed a whole
scripted day against a classifier holding zero examples. They passed. A green
simulator was proving something other than the routing the box runs, and a
regression in the seed set could not have surfaced there.

seedPath walks up to five levels to find models/seeds, so the daemon started
from the repo root behaves exactly as before and a test started anywhere
inside the tree finds the same files. All three scenarios still pass with 339
seeds loaded, so the outcome was not resting on the empty classifier.

The new test asserts the count rather than logging it. A silent zero is the
failure that hid here.
2026-08-04 03:14:52 +04:00
claude 87d176153a router: stage 0 claims the task marker before the model renames it (V-467)
Spoken capture was dead. "добавь в задачи купить молоко" routed act, so the
gate found no allowlisted fn and asked "Что сделать?", and the list stayed
empty. Capture rides the note intent by design (#130, no eighth intent), and
nothing under actionNote was reached any more. The model also rewrote the
payload on the way — "купить молоко" came back as "сделать покупку молока",
and a task must read as the words he said.

TaskCaptureGrammar answers it at stage 0, the same place the agenda rules
went. It matches any utterance and lets ParseTaskCapture refuse, so the
marker list stays data. Three phrasings he used are added to that list:
"запиши в список дел" and the two next to it were missing.

The other deterministic matchers were checked for the same exposure. They
are all question-shaped — money, habit, feed, day plan, task list, calendar —
and a question lands on query, which is where they already sit. Capture was
the only imperative among them, which is why only it was taken.

ru-note-006 is the fixture case. The classifier alone cannot pass it, and the
hash baseline drops by that one case; the daemon answers it at stage 0.
2026-08-04 03:12:47 +04:00
claude 7d4b4ad736 clarify: a parked question belongs to the conversation that was asked (V-466)
The clarify store had one key for the whole daemon, so a question asked in
the web chat and never answered captured the next three utterances from any
source — telegram, or the mic — and answered them against a request the
speaker never made.

The reach now supplies a conversation id on the IPC Chat call, and the
daemon carries it on the context the way it already carries the correlation
id, so the six clarify call sites read it instead of a constant. The mic has
no id of its own and keeps the key it had, so voice behaves exactly as
before. mavweb has no per-browser session, so every tab is one conversation:
right for a single-owner box, and still distinct from telegram and the mic.

Dialogue sessions stay global on purpose — they are what she remembers about
him, not what she is waiting for from one channel.
2026-08-04 03:08:09 +04:00
claude 6d3f5b5b01 router: a reminder with no subject asks instead of guessing (V-383)
Slots.Text was the raw utterance for every intent, so a reminder could not
have an empty subject. StillMissing never reported SlotText, the question
"О чём напомнить?" was unaskable, and the branch in PendingQuestion.Answer
that fills a text slot could only overwrite the whole request.

The LLM path now keeps the model's own text, empty included, and the gate
turns a subjectless reminder into a question. The classifier path is
unchanged: it has no subject parser, so the utterance is the only signal it
has.
2026-08-04 02:49:02 +04:00
claude eda1112f3b mavgpud: a yield stops writing a core and reads as a yield (V-491)
llama-server aborts inside its own static teardown on SIGTERM — the
handler calls exit(), stream_session_manager's destructor throws, and the
process dies "signal: aborted (core dumped)". mavgpud sends that signal on
every eviction, so a routine yield wrote a multi-gigabyte core into
systemd-coredump and logged the same line a real crash would.

LimitCORE=0 in the unit stops the disk cost. A yielding flag, set by stop
and cleared by start, makes the log distinguish the two: only an exit we
did not ask for is still reported as an exit.

Not filed upstream. Searched ggml-org/llama.cpp for
"ggml_uncaught_exception" with SIGTERM and for stream_session_manager and
found nothing matching, so the issue still wants writing — by someone with
an account on that tracker, which is why it is not in this commit.
2026-08-04 02:01:17 +04:00
claude 8833a9c76b mavend: keep only the stub floor in llmReplier (V-396)
The prompt, the call and the output parsing now live in internal/phraser. What is
left here is the one thing the daemon adds: a clarify, a model error and an
unusable generation all answer from voice.StubReplier, so a turn never breaks on
the model. The duplicated stripThink and parseResponseMood copies are gone;
capture.go uses phraser.StripThink.
2026-08-04 00:33:30 +04:00
claude f9b2391a8b phraser: cap llama-server's prompt cache at 512 MiB (V-499)
The forwarded log named the cause in one line: the prompt cache limit
defaults to 8192 MiB. llama-server saves the full KV state of every idle
slot it evicts, 112 kiB per token, so RSS climbed about 170MB per
distinct prompt until the deployed server held 7.9GB for a 1.1GB model.

Measured on homesrv today, uncapped versus `--cache-ram 512`: RSS
plateaus at 932MB from the fourth distinct prompt instead of climbing.
The task's leading guess was wrong. `-ngl 99` costs almost no RSS,
because RADV keeps device memory outside the process. Numbers and method
in docs/evals/2026-08-03-llama-prompt-cache.md.

`-c 4096` is untouched. The knob is `phraser.cache_ram_mib`, unset means
512, negative passes no flag for a llama-server too old to know it.

The deploy still runs the old image, so the box keeps its 8 GiB default
until mavend is rebuilt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 23:20:25 +04:00
claude 86817d6d06 memory: score the personal boundary on seeds, not word lists (V-495)
"что я говорил про бэкапы?" is his data by definition, and nothing outside the
box has ever heard him say anything. The boundary matched possession words only,
so the question walked past it into SearXNG and came back answered out of a Habr
article about somebody else's backups.

A speech-verb marker class was written first and dropped. Russian gives every
verb a dozen surface forms and the "как я говорил, ..." preamble list has no end,
so each form the lexicon missed was one more question reaching the world, and a
missing verb looks exactly like no bug.

The boundary now embeds two frozen seed sets and scores the turn's own query
vector, already computed upstream, against both. Nearest side wins. The
possession markers stay as the offline floor for a handler with no embedder.

19/19 held-out utterances correct against multilingual-e5-small; see
docs/evals/2026-08-03-personal-boundary.md. The live probe on the deployed box is
not done.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:57:11 +04:00
claude ad60e10e95 mavend: run the fact vector repair on start, and test what it does (V-493)
Automatic rather than a flag, unlike -reembed: only voice-tapped facts are in
this index, so it is tens of embeddings rather than thousands of notes. And
waiting for an operator to know the repair exists is the failure being fixed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:36:54 +04:00
claude dbdab2d570 store, mavend: a fact is indexed as the fact, not as the utterance (V-493)
queryMemory returns a fact's stored text verbatim, so the text the write path
indexed is what he hears. It was the utterance, which made recall of any
voice-tapped fact answer with the sentence he said: go_version = 1.20 was
indexed as "какая последняя версия языка Go?", and that question came back.

FactRecallText renders the fact instead, and the utterance stays in meta as
provenance. Correcting a value now drops the key's vectors the way voiding one
does, since the superseded value was still answering.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:36:34 +04:00
claude 62c2e92ec0 mavend, recalleval: wire the topic veto into both recall sources (V-470)
queryMemory and queryNotes both gate on score alone, so both needed it. The
eval keeps its own copy of bestRecall — package main is not importable — and a
fixture that measures a weaker gate than the daemon runs flatters it, so the copy
moves in step and its test pins the new rule.

Measured on the held-out recall fixture with the real embedder: 17/32 cases pass
→ 22/32, false recall 1/5 → 0/5, answered after gate 18/27 → 17/27. The one true
recall lost is en-hard-024, an English question against a Russian note, where no
lexical test can help.
2026-08-03 13:51:04 +04:00
claude 4dfe106fe3 mavend: a question is never a fact about him (V-470)
IntentFact used to persist whatever the model invented for a question-shaped
utterance, at confidence 1.00, and index it for recall under the question's own
text. Two such rows then claimed seven unrelated world questions and silently
disabled world answering.

A question now goes down the query chain, which is what he asked for. The second
half is confidence: a value grounded in what he said stays 1.00, a value the model
supplied for words he never said drops to 0.60 and says so in the log. Same
reasoning as 'LLM output is not authorization' on the act path.
2026-08-03 13:40:33 +04:00
claude 12530c8a95 mavend: world questions ask the workstation, and name the gap when it is asleep (V-490)
queryGeneral has nothing fetched to fall back on, so it is the sharp case:
with a workstation configured and asleep he is told that, rather than told
something false in a confident voice. The 1.7B answering a world question is
where "Война и мир" got Левитан as its author.

The sources that already hold a passage — a live search, a ZIM article, a
page he named — go through the world model too, but read the passage back
when it is not there instead of naming a gap. A real quote beats "не могу
сейчас", and nothing is invented on either path.

The Stub and every test double keep the Phraser interface they have.
PhraseWorld is reached by assertion, and a phraser without it is the
no-workstation case.
2026-08-03 12:27:40 +04:00
claude 92d5fd580c mavend: route and reply through the workstation when its card is free (V-485)
modelSeam builds an llm.Pair when a workstation is configured and hands it to
the router and the replier. Both are the silent half of the degradation rule:
the big model is only better there, and he is never told which model answered.
No block, no probe, and the box behaves exactly as it did.
2026-08-03 10:12:21 +02:00
claude e52c616592 mavgpud: test the probe against the sysfs the workstation actually has (V-488)
The fixtures are the live numbers sampled from the box on 02-08-2026, where the
CPT run held 12.8GB of 16 as proc/478104/vram_35881.

The cases that matter are the ones where a mistake is silent: our own
llama-server counting as a contender, an unreadable card reading as free, and
/health hanging or proxying into a closed port instead of answering 503.
2026-08-02 17:03:56 +02:00
claude 2b97bac51e mavgpud: keep the model loaded while the card is free, yield when it is not (V-488)
The lifecycle rule from Vikunja #488. Not on demand, because a 7-14B takes tens
of seconds to load and a world question would meet a gap every time the card
had been quiet. Not always on, because that is what holds the card.

/health is answered locally and always, so Maven's prober costs nothing and
works while the model is down. Everything else is reverse-proxied to
llama-server, which is what makes the idle window measurable at all.

Yielding is checked before starting, and both transitions are damped by a poll
streak so a short-lived rocm process cannot evict the model.
2026-08-02 17:03:56 +02:00
claude ab42db2b87 mavgpud: read the card from sysfs and own llama-server's lifecycle (V-488)
The workstation cannot keep a 7-14B resident: it would hold 16GB against the
owner's CPT runs, Correx and the manga-recap pipeline. So the process that
stays up costs no VRAM and the model comes and goes under it.

Contention is detected by presence on the KFD, not by a VRAM threshold. A ROCm
process registers under /sys/class/kfd/kfd/proc when it initialises HIP, well
before it allocates, so we see a contender during its startup instead of after
it has already lost an allocation race. rocm-smi is not installed on that box
and a per-second subprocess would get tuned down until useless, so this reads
sysfs and forks nothing.

Free VRAM is read only to decide whether to start. It is never a reason to
stop: by the time free VRAM has dropped, the other job has already failed.
2026-08-02 17:03:56 +02:00
kami 92d2629001 Merge pull request 'Voice cannot accept a routine (V-367); the last three prompts are Russian (V-404)' (#89) from fix/367-voice-parks-routine-accept into master 2026-08-02 11:39:16 +02:00
claude bb8cb8d014 Voice parks a routine proposal, it never accepts it (V-367)
Accepting a proposed routine gives the tick loop a standing new reason to
speak. DESIGN.md § "surface caps authority" puts that at layer 3, and says
voice is structurally incapable of layer 3 because a room mic is reachable
by anyone in the room. The /routines button was gated at step-up; the voice
path accepted outright. The two surfaces disagreed, so one of them was wrong.

A spoken "да" now leaves the row 'proposed' and sends him to /routines,
where the gated button is. A spoken "нет" still dismisses: declining does
not move the boundary outward, so voice keeps it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018CotYKycuio1GwLbYh9jfc
2026-08-02 09:58:07 +04:00
claude 5e66aa8f22 fix: the slots converter was still dropping the fact Value (V-365)
dialogue.Slots gained Value in 925ce22, but toDialogueSlots never copied
it, so a clarifying answer carrying a fact payload still landed nowhere:
clarify.go:202 sends the answer through the converter, and the SlotValue
arm reads answer.Value.

Both converters now carry every field. TestSlotsParity compares the two
field sets by name and type; TestSlotsRoundTrip populates every router
field and checks the round trip, and fails the fixture itself when a new
field is left zero.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QChoBS5qJSrCV98oNUnHNU
2026-08-02 09:35:47 +04:00
claude 4f34a232d4 docs: retire the two root queues, archive the ranking (V-447)
PROGRESS.md and 20-07-2026-BACKLOG.md were state snapshots that git log and
the Vikunja board already carry. Everything PROGRESS.md claimed as shipped is
a QA task. The backlog's only untracked item, bounded follow-up state, is now
V-448.

maven-feature-ranking.md moves to docs/archive/2026-07-03-feature-ranking.md
instead of dying. Its mandatory and easy tiers became V-449 through V-458; the
doable and epic tiers are reasoning about why things are not worth doing yet,
which no task captures.

Four code comments and the design.md ledger pointed at the deleted files.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xwomr2cT93KMSz9u3digX5
2026-08-02 09:11:11 +04:00
claude 93987f2dfc docs: tier the tree by lifetime, so staleness shows in the path (V-446)
Seventeen markdown files at the repo root, twelve of them dated one-shot
reports sitting next to CLAUDE.md. That is why stale docs read as
current: nothing in the path said which was which.

Root now keeps CLAUDE.md and AGENTS.md. Living docs move under docs/
and carry a Last verified line. Dated measurements move to docs/evals/
ISO-prefixed, and are never edited after the day, so a newer number is
a new file. The senior review moves to docs/archive/.

Every reference was rewritten across markdown, Go comments, the Makefile
and the recall fixture. The touched Go packages still build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 03:28:49 +04:00
kami 612ca8cf1b grammar: the last two model calls that were still free text
The replier and the meeting summariser were the two call sites without a
GBNF. Both are exactly the shape that makes a Thinking variant answer with
its reasoning as prose, and neither had anything downstream that could
remove it.

The replier already parses {"response","mood"}, so it now sends the phraser's
grammar for that contract, exported once as phraser.ResponseGrammar so the
two definitions cannot drift.

The summariser stays text-in/text-out. The JSON wrapper is attached and
unwrapped in the daemon's Completer, so internal/capture is unchanged and a
Completer without a grammar still works.

The simulator told routing from phrasing by "has a grammar", which stopped
being true here; it now looks for the intent enum.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 02:28:53 +04:00
kami 14e98334ad query: let her search the live web before she reads the ZIMs
The offline encyclopedia was the only world source, and it reads what was true
when the ZIM was built. A self-hosted SearXNG now asks first and Kiwix is the
fallback for an empty result, an unreachable instance or no line out. Owner's
ruling, 2026-08-02.

internal/websearch is deliberately thin: no rewriter (SearXNG ranks through
real engines, so the Russian question goes out as he asked it), no page fetch,
no cache. It cannot read the store, so only the query string can leave the box.

The personal boundary is unchanged and still sits above this source, so a
question about him is never searched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BFaeSbLMEVG5ey8tejU3y2
2026-08-02 02:24:36 +04:00
kami a103708a08 memeval: five minutes again, now that the gate keeps a turn from waiting
The budget was cut to 60s because a five-minute evaluation held the single
llama-server slot, and a voice turn arriving mid-evaluation waited behind it.
That collision is now solved where it belongs: the background client yields the
slot while a turn is in flight.

With the gate in place the short budget only truncates a Thinking model
mid-synthesis, which costs an observation and saves no latency on any real turn.
Kami's call, 2026-08-02.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BFaeSbLMEVG5ey8tejU3y2
2026-08-02 02:16:45 +04:00
kami b8227295b8 query: stop asking the world about him
"во сколько у меня встреча" fell past the calendar, reached Kiwix, matched
an article on the 2015 CPISRA World Games and came back phrased as his
meeting. Inventing is worse than refusing, and as `system` this used to say
"пока не умею".

A new source sits between the notes pass and Kiwix: a question carrying a
possession marker ("у меня", "мой", "my", "do i have", "did i") that his own
data did not answer has no answer outside it either, so the walk stops there.
The markers are possession, not first person, so "как мне сварить борщ" still
reaches the encyclopedia.

This is also where CLAUDE.md's privacy line lands: only the utterance may
leave the box, and a question about him carries his life in the utterance.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 23:27:04 +04:00
kami b35151418a mavend: give the voice handler a CoreAPI that can serve the day plan
wireVoice runs before the tick loop exists, so it could only be handed
the bare store adapter — and that adapter answers DayPlan with "not
available via direct store API", because a day plan is assembled by the
tick loop and is not a table to read. So queryDayPlan, which the query
chain reaches for "какие у меня планы на сегодня", failed for every
caller on the deployed daemon.

main already back-patches the other direction (daemonAPI.chatFn =
handler.handleText). This is the same seam in reverse, at both wiring
sites. No recursion risk: nothing in the voice path calls api.Chat.

With the plan reachable, it recited its reminders as literal JSON. The
payload unwrapper existed but was private to the phraser, so the day
plan had its own non-unwrapping copy. One owner now, store.ReminderText,
with the phraser delegating to it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 23:17:10 +04:00
kami 17964d1162 dialogue: keep the remembered topic on the latest turn
rememberTurn runs after followUpMerge, which has already inherited a
Text slot from the previous same-intent turn, so the fill-if-empty rule
pinned the first topic of a run of query turns and never released it.
"во сколько у меня встреча", then "какие у меня планы", then "а завтра?"
continued the meeting — two turns stale.

Overwrite for system and query, where Text is a topic and not a payload.
A continuation is the exception and keeps what it inherited: its own
utterance is the ellipsis, and the topic it carries is the real one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:57:08 +04:00
kami 079cf689aa router: agenda questions belong to query, not to system
"что у меня сегодня" and "что у меня в календаре сегодня" both routed
IntentSystem on the deployed daemon, and replySystem has no agenda arm,
so both answered "пока не умею". The calendar source that can answer
them lives in the query chain and was never reached. The fixture has
said query since ru-query-019 was written; the daemon disagreed with the
fixture and the daemon was wrong.

AgendaQueryGrammars routes them at stage 0, after the clock rules so
"какой сегодня день" keeps reaching replySystem. Intent only — which
source claims the turn stays the query chain's decision.

This is what made the follow-up continuation look like it only worked
for "what day is it". It did: the query half inherited an intent whose
handler could not answer, so both halves came back "пока не умею".

Measured on the 77-case RU fixture: full accuracy 70.1% → 72.7%,
intent-only 75.3% → 77.9%, calendar 0/2 → 2/2, clarify counts unchanged.
The eval harness wires the new grammars too, or the fixture would stop
being a measurement of the daemon.

Go's \b is ASCII-only and never fires after a Cyrillic letter, which the
first version of the pattern learned the hard way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:53:01 +04:00
kami 9397f9e5f6 query: only a source that knows the day may answer "а завтра?"
The follow-up continuation re-aimed Slots.Time, but every query source
matches on dec.Utterance and nothing in the chain reads it. So the
inheritance bought almost nothing: "какие планы на сегодня" then "а
завтра?" missed day-plan's matcher and fell to the calendar, which had
parsed the day out of the raw utterance anyway.

Worse than nothing in one place. queryFactByKey runs first and claims on
HasKey plus HasTime, both of which the continuation sets, so a keyed
query continued with "а вчера?" answered "я записала это <the fact's own
timestamp>" and dropped the day entirely.

Widening the topic for every source would have made it worse still, not
better: CalendarEvents is the only CoreAPI call that takes a date, so
ten date-blind sources would have answered a question about tomorrow
with today's data. Gate instead of widen — a continuation is only
offered to a source that reads the day, and when none claims it she says
so rather than "не знаю", which reads as "nothing tomorrow" when she
never looked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:32:37 +04:00
kami 3f98a99f44 continuation: only an ellipsis may widen the keyword match
Deployed check, second round: "привет" after "какой сегодня день" answered
with the date. followUpMerge fills an empty Text from the previous
same-intent turn, so the topic-widening added a minute earlier was reading
an inherited topic on turns that had nothing to do with it.

Decision gains Continued, set only by continuation.go and never by the
router. replySystem widens on that and nothing else, so an inherited Text
is back to being invisible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:21:28 +04:00