Commit Graph

520 Commits

Author SHA1 Message Date
claude 0987dabfc4 tool: risk tiers decide the confirm, not one boolean (V-449)
The Destructive column was a mechanism with no policy behind it: nothing said
which acts are destructive, whether a confirmed act stays confirmed, or what a
new tool domain inherits, so each domain answered for itself.

Three tiers, derived from the row rather than stored, so the answer can be
argued with in one place instead of being whatever the last person to tick the
checkbox believed. Safe runs. Destructive costs a confirm turn, every time —
a confirmation binds one capability, one target and one argument list, and it
dies with the parked turn. Irreversible is refused: a confirm turn there would
be theatre, because the STT, the router and the fuzzy allowlist match are all
guesses and a spoken "да" checks none of them. She names the gap; the row
stays enabled.

An unrecognised dispatch shape inherits destructive, not safe. A domain argues
its way down to running freely, never up to being gated.
2026-08-04 04:50:56 +04:00
claude 947506c7b8 docs: a list is the fourth append-only shape (V-453) 2026-08-04 04:45:41 +04:00
claude 6c67e61962 mavend: cover the spoken list path (V-453) 2026-08-04 04:45:21 +04:00
claude 0990f32808 mavend: the list is reachable from voice (V-453)
An add and a crossing-off run at the top of actionNote, next to task
capture and before the embedding is paid for; the read-back is a query
source sitting beside "tasks", so the recall pass cannot answer "что мне
купить?" from an old note about the shop.

Crossing off one item claims the turn only when the list actually holds
that item, which is what keeps "купил новый ноутбук" a note.

These read h.dataStore rather than the CoreAPI: a list is local to the core
and nothing outside it writes one. The ipc seam is what it grows through
when something outside mavend needs to add to a list.
2026-08-04 04:45:21 +04:00
claude d41878c2b1 router: cover the list parsers and wire the grammars (V-453) 2026-08-04 04:45:13 +04:00
claude e023638135 router: parse list capture, read-back and crossing off (V-453)
Same posture as task capture and for the same reason: the intent enum is a
contract shared with the relabelling prompt, so a list is not an eighth
intent. It is a note-shaped or query-shaped utterance carrying an explicit
marker, and the marker is a lookup.

The markers are deliberately explicit — "молоко закончилось" is an
observation and stays a note. The list tag is matched by stem, because
Russian declines it: "список покупок", "в покупки" and "в покупках" are one
list. ListGrammars puts both halves at stage 0, so an add and a read-back
never depend on the model having a good turn.
2026-08-04 04:45:13 +04:00
claude 0d52344d27 store: cover the list_items shape with tests (V-453) 2026-08-04 04:39:24 +04:00
claude 5bd303788b store: add list_items, the fourth append-only shape (V-453)
A list is a standing set of short strings under a tag. Not a task, because
milk is not work and the prioritiser must not count it as an errand; not a
fact, because it claims nothing. Nothing predicates over it, so two people
adding to the same list at once costs nothing.

Migration #19, plus AddListItem, ListItems, SetListItemStatus and ClearList.
The live-only unique index is the tasks one, per list: молоко twice before
the shop is one row, молоко again after it was crossed off is a new one.
2026-08-04 04:39:24 +04:00
claude afac8fb670 mavend: run the persona checks before she speaks (V-399)
The checks stay in the eval package and the daemon calls three of them:
feminine, address, and a new leaked-reasoning test. No retry — it doubles
the latency on the turn that is already going badly, and on the nudge path
the moment has passed. A failure falls back to the deterministic floor and
is logged with the whole rejected text and counted by check name.

hisgender is deliberately not run: the simulator showed it rejecting
"записала, что ты выпил воды", which is her own correct self-reference.
2026-08-04 04:35:42 +04:00
claude 908d92a7e8 calendar: a Russian summary keeps its letters in the fact key (V-443)
safeKey kept ASCII only, so "Встреча с Аней" and "Обед с мамой" both
reduced to "--" and shared one key on one day. The second event of the
day overwrote the first, silently, and his calendar is Russian.

Letters and digits in any script now pass. Migration #18 deletes the rows
written under the old rule instead of rewriting them: a calendar fact is
derived, the next poll writes the day again, and a stale row reads as an
extra meeting.
2026-08-04 02:56:09 +04:00
claude 43f2c37538 router: stage 0 claims the other days and the named event (V-471)
"какие планы на сегодня" worked and "какие планы на завтра" answered
"пока не умею": the agenda rule needs "у меня" or a calendar noun, and
that phrasing carries neither. "когда планёрка?" had the same shape.

Two rules. One takes a plan noun aimed at a named day, one takes a closed
list of event nouns after "когда"/"во сколько". Both route intent only,
so the query chain still decides which source answers.

classifier+onnx over the fixture: 55/79, 69.6% full, with the two new
cases passing and no case moving the other way.
2026-08-04 02:52:51 +04:00
claude 6d3f5b5b01 router: a reminder with no subject asks instead of guessing (V-383)
Slots.Text was the raw utterance for every intent, so a reminder could not
have an empty subject. StillMissing never reported SlotText, the question
"О чём напомнить?" was unaskable, and the branch in PendingQuestion.Answer
that fills a text slot could only overwrite the whole request.

The LLM path now keeps the model's own text, empty included, and the gate
turns a subjectless reminder into a question. The classifier path is
unchanged: it has no subject parser, so the utterance is the only signal it
has.
2026-08-04 02:49:02 +04:00
claude eda1112f3b mavgpud: a yield stops writing a core and reads as a yield (V-491)
llama-server aborts inside its own static teardown on SIGTERM — the
handler calls exit(), stream_session_manager's destructor throws, and the
process dies "signal: aborted (core dumped)". mavgpud sends that signal on
every eviction, so a routine yield wrote a multi-gigabyte core into
systemd-coredump and logged the same line a real crash would.

LimitCORE=0 in the unit stops the disk cost. A yielding flag, set by stop
and cleared by start, makes the log distinguish the two: only an exit we
did not ask for is still reported as an exit.

Not filed upstream. Searched ggml-org/llama.cpp for
"ggml_uncaught_exception" with SIGTERM and for stream_session_manager and
found nothing matching, so the issue still wants writing — by someone with
an account on that tracker, which is why it is not in this commit.
2026-08-04 02:01:17 +04:00
kami 71041029e2 Merge pull request 'The reply path can't be tested — llmReplier is stuck in package main' (#107) from task/396-the-reply-path-can-t-be-tested-llmreplie into master
Reviewed-on: #107
2026-08-03 22:35:56 +02:00
claude 35018226ef eval: score the reply path, the fourth phrasing path (V-396)
Nine reply cases and a fourth column in the talk report. The reply path is a
separate object from the phraser in the daemon, so Pair joins a Talker and a
Confirmer for a run that covers everything Maven says.

Cases carry intent/key/value because the replier is phrased from the decision the
router resolved, not from the raw utterance. Three of them are baits the other
paths cannot produce: a masculine verb about himself that she must not copy onto
herself, a polite plural input that must still come back на ты, and an unresolved
note that invites a question a confirmation is not allowed to ask.

Not scored against a model here — this box has no llama-server, and the baseline
test is opt-in on MAVEN_LLM_URL.
2026-08-04 00:33:41 +04:00
claude 8833a9c76b mavend: keep only the stub floor in llmReplier (V-396)
The prompt, the call and the output parsing now live in internal/phraser. What is
left here is the one thing the daemon adds: a clarify, a model error and an
unusable generation all answer from voice.StubReplier, so a turn never breaks on
the model. The duplicated stripThink and parseResponseMood copies are gone;
capture.go uses phraser.StripThink.
2026-08-04 00:33:30 +04:00
claude 6c07409452 phraser: add Replier, the reply path lifted out of package main (V-396)
llmReplier lived in cmd/mavend, so the confirmation he hears after every fact,
note and reminder was the one phrasing path nothing could import or score.

Replier owns the prompt, the call and the parsing, and returns its errors instead
of hiding them — a dead model shows up as an error rather than as bad phrasing.
It has no stub fallback of its own; the daemon keeps that. StripThink is exported
for the daemon's own model callers.
2026-08-04 00:33:30 +04:00
kami 1c2541f7d6 Merge pull request 'llama-server holds 7.9GB RSS for a 1.1GB model, and its startup log goes nowhere' (#105) from task/496-recall-a-cross-language-question-loses-i into master
Reviewed-on: #105
2026-08-03 22:04:29 +02:00
claude 9e25f18a3e memory: record the recall topic veto's real price (V-496)
#496 asked to skip the veto when the question and the hit are in
different scripts, so an English question stops losing a Russian note.
Measured first: the fixture has no cross-language case, and en-hard-024
is an English question against an English note. Both proposed fixes are
no-ops.

What the veto actually does on the fixture, with the real embedder: it
costs en-hard-024 and buys ru-silent-029. Pass count is 22/32 either
way; false recall is 0/5 with it and 1/5 without. The two cases are one
lexical class, so no rule cheap enough for RecallAllowed separates them.

Accepts the loss and pins both sides in a test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 00:01:49 +04:00
kami 197897516e Merge pull request 'Task/495 bug x escapes the personal boundary and' (#104) from task/495-bug-x-escapes-the-personal-boundary-and into master
Reviewed-on: #104
2026-08-03 21:30:56 +02:00
kami 767748720a Merge pull request 'llama-server holds 7.9GB RSS for a 1.1GB model, and its startup log goes nowhere' (#103) from task/499-llama-server-holds-7-9gb-rss-for-a-1-1gb into task/495-bug-x-escapes-the-personal-boundary-and
Reviewed-on: #103
2026-08-03 21:30:37 +02:00
claude 58051b5af1 docs: record the #499 deploy (V-499) 2026-08-03 23:28:36 +04:00
claude f9b2391a8b phraser: cap llama-server's prompt cache at 512 MiB (V-499)
The forwarded log named the cause in one line: the prompt cache limit
defaults to 8192 MiB. llama-server saves the full KV state of every idle
slot it evicts, 112 kiB per token, so RSS climbed about 170MB per
distinct prompt until the deployed server held 7.9GB for a 1.1GB model.

Measured on homesrv today, uncapped versus `--cache-ram 512`: RSS
plateaus at 932MB from the fourth distinct prompt instead of climbing.
The task's leading guess was wrong. `-ngl 99` costs almost no RSS,
because RADV keeps device memory outside the process. Numbers and method
in docs/evals/2026-08-03-llama-prompt-cache.md.

`-c 4096` is untouched. The knob is `phraser.cache_ram_mib`, unset means
512, negative passes no flag for a llama-server too old to know it.

The deploy still runs the old image, so the box keeps its 8 GiB default
until mavend is rebuilt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 23:20:25 +04:00
claude f229795cea phraser: forward llama-server's output to mavend's log (V-499)
mavend scraped the child's stderr for the listen line and threw every
other line away, and never piped its stdout at all. Nothing about the
resident model's memory was diagnosable from a running box: no buffer
sizes, no KV-cache layout, no offload lines, no prompt-cache limit.

Both streams now share one pipe and every line lands in mavend's log
with a `llama:` prefix. The last 12 startup lines are also kept and go
into the error when the server dies before it listens, because bare
"EOF" never named which allocation it choked on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 23:19:03 +04:00
kami 6e5364a0ed Merge pull request 'Bug: "что я говорил про X" escapes the personal boundary and reaches web search' (#102) from task/495-bug-x-escapes-the-personal-boundary-and into master
Reviewed-on: #102
2026-08-03 20:59:34 +02:00
claude 86817d6d06 memory: score the personal boundary on seeds, not word lists (V-495)
"что я говорил про бэкапы?" is his data by definition, and nothing outside the
box has ever heard him say anything. The boundary matched possession words only,
so the question walked past it into SearXNG and came back answered out of a Habr
article about somebody else's backups.

A speech-verb marker class was written first and dropped. Russian gives every
verb a dozen surface forms and the "как я говорил, ..." preamble list has no end,
so each form the lexicon missed was one more question reaching the world, and a
missing verb looks exactly like no bug.

The boundary now embeds two frozen seed sets and scores the turn's own query
vector, already computed upstream, against both. Nearest side wins. The
possession markers stay as the offline floor for a handler with no embedder.

19/19 held-out utterances correct against multilingual-e5-small; see
docs/evals/2026-08-03-personal-boundary.md. The live probe on the deployed box is
not done.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:57:11 +04:00
kami 0fc2e3a18a Merge pull request 'Task/470 bug a question writes invented knowledge' (#101) from task/470-bug-a-question-writes-invented-knowledge into master
Reviewed-on: #101
2026-08-03 20:44:12 +02:00
kami 453919db20 Merge pull request 'Bug: the memory index stores the raw utterance as a fact's recall text, and nothing ever deletes a fact vector' (#100) from task/493-bug-the-memory-index-stores-the-raw-utte into task/470-bug-a-question-writes-invented-knowledge
Reviewed-on: #100
2026-08-03 20:41:09 +02:00
claude ad60e10e95 mavend: run the fact vector repair on start, and test what it does (V-493)
Automatic rather than a flag, unlike -reembed: only voice-tapped facts are in
this index, so it is tens of embeddings rather than thousands of notes. And
waiting for an operator to know the repair exists is the failure being fixed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:36:54 +04:00
claude 1528697287 store: repair fact vectors against the facts they name (V-493)
Every write-path fix leaves the rows already stored wrong, and a box in that
state looks fine: recall answers with the wrong text and nothing logs an error.
That is how the original poison survived four restarts.

RepairFactVectors resolves each fact vector against the fact it names,
re-embeds the ones whose text is stale, and deletes the voided, superseded and
orphaned ones. Marker-guarded and idempotent, so it runs once per box and a run
that dies partway is simply redone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:36:54 +04:00
claude dbdab2d570 store, mavend: a fact is indexed as the fact, not as the utterance (V-493)
queryMemory returns a fact's stored text verbatim, so the text the write path
indexed is what he hears. It was the utterance, which made recall of any
voice-tapped fact answer with the sentence he said: go_version = 1.20 was
indexed as "какая последняя версия языка Go?", and that question came back.

FactRecallText renders the fact instead, and the utterance stays in meta as
provenance. Correcting a value now drops the key's vectors the way voiding one
does, since the superseded value was still answering.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 22:36:34 +04:00
kami b9371dcac6 Merge pull request 'Bug: a question writes invented knowledge into memory as a self fact, and recall then serves it back for unrelated questions' (#99) from task/470-bug-a-question-writes-invented-knowledge into master
Reviewed-on: #99
2026-08-03 20:11:11 +02:00
claude 62c2e92ec0 mavend, recalleval: wire the topic veto into both recall sources (V-470)
queryMemory and queryNotes both gate on score alone, so both needed it. The
eval keeps its own copy of bestRecall — package main is not importable — and a
fixture that measures a weaker gate than the daemon runs flatters it, so the copy
moves in step and its test pins the new rule.

Measured on the held-out recall fixture with the real embedder: 17/32 cases pass
→ 22/32, false recall 1/5 → 0/5, answered after gate 18/27 → 17/27. The one true
recall lost is en-hard-024, an English question against a Russian note, where no
lexical test can help.
2026-08-03 13:51:04 +04:00
claude aec94eb2e8 memory: a world question must name what the memory mentions (V-470)
The score gate cannot separate the right note from an unrelated one: the
held-out fixture puts the right note at 0.791-0.890 and the must-be-silent cases
at 0.795-0.835, so a note about his slow network answered 'почему небо синее?'.

RecallAllowed adds a topic veto, and applies it only to a question that mentions
nothing of his. That restriction is the whole design: demanding a shared word of
every recall silenced four true recalls on the fixture to kill one false one,
because recall exists to find the note whose words he no longer remembers. A
question about his own life keeps the embedder as its only judge.
2026-08-03 13:50:54 +04:00
claude 4dfe106fe3 mavend: a question is never a fact about him (V-470)
IntentFact used to persist whatever the model invented for a question-shaped
utterance, at confidence 1.00, and index it for recall under the question's own
text. Two such rows then claimed seven unrelated world questions and silently
disabled world answering.

A question now goes down the query chain, which is what he asked for. The second
half is confidence: a value grounded in what he said stays 1.00, a value the model
supplied for words he never said drops to 0.60 and says so in the log. Same
reasoning as 'LLM output is not authorization' on the act path.
2026-08-03 13:40:33 +04:00
claude 2e0e2fd0bb router: a deterministic test for question-shaped text (V-470)
The predicate a fact write needs before it trusts a routing decision. Tokenized,
not substring: 'что' inside 'чтобы' is not a question. Capture verbs win over
every question signal, because 'запиши что я пил воду' contains an interrogative
and is still a capture.
2026-08-03 13:40:33 +04:00
claude f3fa6b353a store: voiding a fact drops its memory vectors (V-470)
Revert voided the fact row and left the vector, so recall kept serving the
voided fact's utterance and the documented repair reported success on a box that
stayed broken. There was no way to repair a poisoned box at all.

DeletePrefix covers every vector for the key, earlier rows included: their values
are superseded, and a superseded value has no business claiming a turn. It is
best-effort — the audit trail is already committed, and a fact that is voided but
still recallable beats a void that failed.
2026-08-03 13:40:13 +04:00
kami 6645f64c3e Merge pull request 'Name the gap: world questions through the workstation model, and the four remaining callers' (#98) from task/490-name-the-gap-world-questions-through-the into master
Reviewed-on: #98
2026-08-03 11:19:56 +02:00
claude f10e0068dd config, deploy: the workstation is workpc, not bugmachine (V-490)
Owner's correction. It is the same host CLAUDE.md already calls workpc, and
two names for one machine read as two machines. The dated eval file keeps the
old name: a measurement is never edited after the day it was taken.
2026-08-03 12:42:37 +04:00
claude 9b124d9194 docs: both halves of the degradation rule are wired, and which caller is which (V-490)
The offload inventory grows a column, because "seven callers of the resident
model" stopped being the useful fact. Which of them is offloaded, and under
which half of the rule, is. Three are resident-only on purpose and the table
now says why rather than leaving it to be rediscovered.

The three-outcome table is the part that was not obvious from the rule as
written. A configured-and-asleep workstation names the gap; a box with no
workstation block does not, because naming a gap requires a gap.
2026-08-03 12:29:12 +04:00
claude 12530c8a95 mavend: world questions ask the workstation, and name the gap when it is asleep (V-490)
queryGeneral has nothing fetched to fall back on, so it is the sharp case:
with a workstation configured and asleep he is told that, rather than told
something false in a confident voice. The 1.7B answering a world question is
where "Война и мир" got Левитан as its author.

The sources that already hold a passage — a live search, a ZIM article, a
page he named — go through the world model too, but read the passage back
when it is not there instead of naming a gap. A real quote beats "не могу
сейчас", and nothing is invented on either path.

The Stub and every test double keep the Phraser interface they have.
PhraseWorld is reached by assertion, and a phraser without it is the
no-workstation case.
2026-08-03 12:27:40 +04:00
claude 51256c4c9a phraser: test the three outcomes of a world question, and prompt parity (V-490)
The middle outcome is the whole task: a workstation that is configured and
asleep produces a gap, and the resident model is never asked. The parity
test compares the bytes PhraseWorld sends the workstation against the bytes
PhraseQuery sends the resident model, so the fixtures and the daemon cannot
measure two different prompts.

The nudge tests cover the silent half from both sides, including the
temperature, which is how the workstation would otherwise change how she
sounds without anyone deciding to.
2026-08-03 12:27:30 +04:00
claude 76481c2736 phraser: a world model seam, so a gap can be named instead of invented (V-490)
The naming half of the degradation rule in docs/offload.md. PhraseWorld has
three outcomes: no workstation configured means the resident model answers
exactly as today, a workstation that is taking work answers, and one that is
asleep returns ErrNoWorldModel so the caller can say so. Naming a gap
requires a gap — on a box that never had a second model, refusing every
world question would remove a capability he has now.

Both prompts move into knowledgePrompt and evidencePrompt, shared by
PhraseQuery and PhraseWorld, because prompt parity across two models stops
holding the moment there are two copies of a prompt.

The silent half comes with it: chatWithSystem and chatWithMessages prefer
the workstation when it will take work, at the same 0.7 the resident
transport samples at, and say nothing when it will not. That covers the
digestion worker's nudge and reminder phrasing without touching tick.go.

Only Available and CompleteRemote are in the Remote interface. Pair.Complete
has its own floor and the phraser already owns one; two floors under a
single call is one too many.
2026-08-03 12:27:30 +04:00
claude bcc2305cd0 llm: let a caller name its sampling temperature (V-490)
The phraser's own transport has always sampled at 0.7 and this client has
always been greedy. Routing a phrasing call through the client must not
change how it decodes, so Req carries the temperature and 0 — the zero
value, and what every existing caller wanted — is still greedy.
2026-08-03 12:27:04 +04:00
kami 0ceeac8df4 Merge pull request 'Point Maven at the workstation model: a workstation block, and routing plus replies through llm.Pair' (#97) from task/485-run-the-big-model-on-the-workstation-wit into master
Reviewed-on: #97
2026-08-03 10:13:09 +02:00
claude 4fae13af75 docs: record the workstation routing numbers where the router is documented (V-485)
CLAUDE.md carried only the homesrv figures, which now read as the whole story.
Also points offload.md's order at #490 for the naming half.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJYcaBiuny9UGSpFweQVQ1
2026-08-03 10:12:21 +02:00
claude 774217199e docs: measure gemma-4-12b on the workstation against the resident model (V-485)
Both fixtures, run from homesrv across the LAN with the proxy env stripped.
Routing: 84.4% full / 93.5% intent-only at p50 329ms through the cascade, against
72.7% / 77.9% at p50 0.80-1.04s for Qwen3-1.7B. Talk: 25/27 against 20/27, with
knowledge 9/9. Nudges 15/15. Settles #485's first assumption by measurement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJYcaBiuny9UGSpFweQVQ1
2026-08-03 10:12:21 +02:00
claude 2db59d52a7 deploy, docs: point homesrv at bugmachine and say what is still unwired (V-485) 2026-08-03 10:12:21 +02:00
claude 92d5fd580c mavend: route and reply through the workstation when its card is free (V-485)
modelSeam builds an llm.Pair when a workstation is configured and hands it to
the router and the replier. Both are the silent half of the degradation rule:
the big model is only better there, and he is never told which model answered.
No block, no probe, and the box behaves exactly as it did.
2026-08-03 10:12:21 +02:00
claude edeef19ff0 config: a workstation block, dropped when it names no address (V-485)
Health defaults to the supervisor's /health rather than llama-server's,
because mavgpud is what answers 503 while the card is held.
2026-08-03 10:12:21 +02:00