Commit Graph

280 Commits

Author SHA1 Message Date
kami 5934110fa8 Move the Praxis/Hexis act handling out of voice.go
voice.go is still the biggest file in cmd/mavend and most of what is left
has nothing to do with the audio path. The ecosystem integration is one
such lump: it talks to Nexus, Praxis and Hexis over HTTP and only touches
the handler for its store and clock. Lifting it into ecosystem_acts.go
puts it next to ecosystem.go, where the clients it drives already live.

Move-only: handlePraxisAct, recordPraxisTrace (called from nowhere else),
handleHexisAct and execHexis verbatim, plus the two imports that became
unused in voice.go.
2026-07-31 23:43:47 +04:00
kami 6e47a3d736 Merge branch 'worktree-agent-a88193d1d84b04a5b' into integration/small-batch 2026-07-31 23:38:00 +04:00
kami 67a5eb3805 Split applyAction's 300-line switch into a per-intent handler table
applyAction (cmd/mavend/voice.go) dispatched all 7 intents from one giant
switch. Extract each case body verbatim into its own actionXxx method in
new cmd/mavend/actions.go, dispatched from an actionHandlers table keyed by
router.Intent. applyAction itself is now just the dec.Clarify guard plus a
table lookup.

No behaviour change: same reply strings, same side-effect order, comments
moved verbatim. The destructive-act confirm gate and the enabled-tool
allowlist stay entirely inside actionAct, exactly where they lived in the
old switch's IntentAct case — they're act-specific, not cross-cutting, so
they don't move to a separate layer. dec.Clarify short-circuit, dialogue
bookkeeping and detectPattern stay outside the table since they run
regardless of intent.

voice.go: 1638 -> 1344 lines. New actions.go: 362 lines.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:37:33 +04:00
kami 029449eefa Table-drive the IPC dispatcher instead of a 42-arm switch
dispatch() replaces the hand-written switch with a package-level
map[Method]handlerFunc built once at init. Each entry is one
withParams/withParamsVoid/withoutParams call closing only over the
CoreAPI method it invokes — adding a method is now one table line
instead of a new arm.

Check still runs once at the top before any unmarshal, unchanged. The
three non-CoreAPI methods (assert_stepup, store_encryption_key, unlock)
are special-cased before the table lookup since they drive Server
fields (StepUp/WrapKeyFn/UnlockFn), not store state. The current
CoreAPI is loaded once per dispatch and passed into the handler as an
argument, so SetAPI's runtime swap (the unlock transition) still takes
effect on the next request — the table itself never captures an api
value. No wire-format change; existing round-trip and unknown-method
tests in ipc_test.go pass unmodified.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:36:46 +04:00
kami ac36216f5d Merge branch 'worktree-agent-a517419e93e6a219f' into integration/small-batch 2026-07-31 23:31:24 +04:00
kami db17cfcc65 Delete the dead lockedAPI, add UnimplementedCoreAPI for the doubles 2026-07-31 23:31:24 +04:00
kami 7d676eb941 Stop tracking the mavwaked build artifact 2026-07-31 23:30:52 +04:00
kami fe3a4e9514 Merge branch 'worktree-agent-a9e5cef90b263a5e5' into integration/small-batch 2026-07-31 23:27:36 +04:00
kami a906f2afad Extract the pure RU/string/weather helpers out of voice.go 2026-07-31 23:27:08 +04:00
kami a2031a31d1 Record the measured confidence-gate numbers 2026-07-31 23:23:37 +04:00
kami 9b8bdf73cc Merge the five small-task branches 2026-07-31 23:11:26 +04:00
kami 7ad3c9a408 Merge branch 'worktree-agent-af88d63f65f30896b' into integration/small-batch 2026-07-31 23:11:16 +04:00
kami 84e1478823 Merge branch 'worktree-agent-af0fd9507d3e2ee46' into integration/small-batch 2026-07-31 23:11:16 +04:00
kami 9c8d0baffe Merge branch 'worktree-agent-a4cef2a815e32ebbf' into integration/small-batch 2026-07-31 23:11:16 +04:00
kami b0f5a16ec9 Add digest as a real outcome: suppressed care nudges get resurfaced, not lost
Vikunja #281. The interruption policy promised four outcomes — deliver_now,
queue, digest, drop — but only three existed: a care candidate the restraint
gate suppressed for quiet hours / away / calendar-busy simply vanished in
loop.Tick's `continue`, with only the trace remembering why.

internal/morning turned out not to be the natural drain: it's a fixed
Item/FactKey checklist engine, not a generic message bundler, so gate-
suppressed nudge text has nowhere to plug into its evidence model. Built a
parallel (but small, reusing the outbox's shape) durable digest instead:

- internal/store: digest_entries table + EnqueueDigestEntry (dedupes by
  rule+body, mirroring the delivery outbox's bodyHash), PendingDigestEntries,
  ExpireStaleDigestEntries, DrainDigestEntries (mark, never delete — an
  audit trail of what she actually said).
- internal/loop: DigestEligible(severity, blockedBy) is the pure boundary —
  only genuine restraint blocks (quiet_hours/calendar_busy/presence) even
  qualify (cooldown/snooze are not "suppression"); within care, Sev2 (break)
  digests, Sev1 (water/meal — stale by the time anyone could resurface them)
  drops. High severity never digests; alarms bypass the gate and deliver
  unchanged, on purpose.
- cmd/mavend/tick.go: each tick scans ExplainTick's trace for eligible
  blocked candidates, enqueues them, sweeps stale entries (24h expiry — the
  care rules are daily-cadence, so anything older is describing a day
  that's over), and drains the bundle only once the suppression reason has
  actually cleared, capped at 3 spoken items plus a trailing count so a
  digest can't turn into the exact nagging it was built to avoid.

Tests: store-level round-trip/restart-survival/dedupe/expiry/drain, loop-
level severity-boundary unit tests, and tick-level integration tests for
the drain-only-when-clear and never-digest-high-severity behavior.
2026-07-31 23:09:22 +04:00
kami 67563ed1f6 Run pattern detection from the digestion tick, not just voice (#43)
detectPattern only ever fired as a side effect of a voice fact-write, so a
recurring pattern already sitting in history went unnoticed until he
happened to mention it again by voice — the opposite of proactive.

Split the pipeline: extraction (fact -> normalized event) stays where a fact
is written, in voice.go, since it's tied to that write regardless of who's
talking. Detection (events -> stable pattern -> proposed_routines row) moves
into shared code (patterns.go's detectAndPropose) that both the voice path
and the new tick.go:detectPatterns call. The tick runs it every cycle over
every action+object pair on record (store.DistinctEventPairs, added), so a
pattern gets noticed on the daemon's own schedule.

Idempotence and the dismiss-must-stick requirement turned out to already be
handled by the store, not something the tick needs to reinvent:
proposed_routines has UNIQUE(action, object) and CreateProposedRoutine does
ON CONFLICT DO NOTHING, and DismissProposedRoutine flips status in place
without deleting the row. So a pair already proposed, accepted, OR
dismissed is a silent no-op on every later tick — a dismissed pattern can
never resurface, and re-running the scan never spams the /routines page.
Kept the voice-path call (immediate spoken confirmation is a nice feature
UX-wise and is now redundant-but-harmless with the tick, since both paths
share the same guarded detectAndPropose).

Tick-side detection only ever writes a row; it does not notify, ring, or
speak, keeping Maven "not a nag, not autonomous" — the /routines page is
still the only place a proposal becomes visible, and only accepting it
starts producing nudges (fireAcceptedRoutines).

Also fixed the stale vikunja#46 reference in proposed_routines.go — the
TODO it named is what this commit does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:07:33 +04:00
kami f0f7ebc9b2 Give LLM-routed decisions a real confidence so clarify can fire (#359)
Confidence was hardcoded to 1.0 for every LLM decision, and the LLM branch
in Router.Route returned straight from fillSlots without ever touching the
stage-3 threshold gate — so the LLM path could not produce a Clarify no
matter what confidence a model reported. That is why all 6 want_clarify
cases in the 77-case RU fixture were missed by every model in the bake-off.

Fix reads structural signal instead of changing the (parity-locked) router
prompt: a single-token utterance ("вода", "бэкап") is flagged thin evidence
in llmrouter.go; a fact left keyless or an act that never resolves to an
allowlisted fn, checked after fillSlots so the deterministic parsers get
first crack, is flagged in router.go's new gateLLMDecision. Anything below
config.DefaultRouterThreshold (0.55) now sets Clarify=true through the same
path the classifier already uses.

Added unit tests with a stubbed Completer proving both directions: thin
cases clarify, clean multi-word/resolved-slot cases stay confident. The
77-case fixture re-run against a live llama-server is still needed to
confirm the 6/6 moves — not done here, no llama-server on this box.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:07:32 +04:00
kami d12de589a2 mavweb: default -addr to loopback, not all interfaces
PR #47 added two state-changing routes (POST /chat, POST /routines)
behind the -addr flag, which defaulted to ":9200" (all interfaces).
Default now binds 127.0.0.1:9200; anyone who wants LAN/wider exposure
still passes an explicit bind (as deploy/docker-compose.yml already
does with "-addr :9201" inside the container, unaffected by this
default change).

Vikunja #317.
2026-07-31 23:03:45 +04:00
kami 50cc17f33a Lock down deploy/ecosystem/nginx.conf template to match the live host
The template said "drop into your nginx sites" but listened on the
wildcard `listen 80;` with no allow/deny ACL, unlike the actual deployed
hexis.kvmx.ru config which binds only to the WireGuard (10.42.0.1) and
LAN (192.168.1.104) addresses with allow/deny all. Anyone following the
template as written would expose these unauthenticated admin UIs to the
open internet.

Bind explicitly to those two addresses and add the matching ACL block,
mirroring cmd/mavweb/nginx.conf which already does this correctly.
Added a comment naming both addresses as host-specific so a deploy on a
different box swaps the IPs instead of reverting to `listen 80` when the
bind fails.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:01:58 +04:00
kami 73d13f1ea6 Merge pull request 'Stop the docs claiming the LLM router is off' (#49) from docs/fix-drift into master 2026-07-31 20:46:57 +02:00
kami 0ca5748699 Stop the docs claiming the LLM router is off
CLAUDE.md's routing section said "llmrouter is wired nil" and called the
classifier cascade the committed default. That stopped being true when the
integration merge landed: voice.go:214 wires pickLLMRouter, DefaultLLMRouter is
on, and deploy/mavend.json sets llm_router true. It is the first thing anyone
reads before touching the router, so it was pointing the next reader at a
wiring job that is already done.

Rewritten to say the LLM router is the default, the classifier is the failure
floor and must not be deleted, and what the two actually measure — 36.8% at
p50 31ms against 67.5%/72.7% at p50 ~2.7s, a trade accepted on purpose. Names
the one thing still open on that path: Confidence is hardcoded 1.0 in
llmrouter.go, so the LLM never asks for clarification (#359).

Also in CLAUDE.md: the persona line pointed at a memory file that does not
exist, so the actual rule was nowhere in the repo. Written out instead —
feminine self-reference, informal singular address, pet names forbidden but his
name allowed — plus the three eval checks that enforce it.

MODEL-BAKEOFF: three claims had gone stale within hours of being written. There
IS a make eval-models target now; the routing numbers ARE the production path,
not a bench artifact waiting on a wiring change; and the truncated 293 MB gguf
is deleted. Struck through rather than removed, since the caveats are part of
how the evening read at the time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 22:37:15 +04:00
kami b43bb265b5 Say it the way she would actually say it (#48) 2026-07-31 20:17:27 +02:00
kami b9a24334ea Say it the way she'd actually say it
Wording fixes from the review of the clarify + phrasing PRs.

- "На когда напомнить?" → "Когда?". After she has just been asked something,
  the long form is the phrasing of a form field, not of a person.
- A reminder now wants a subject as well as a time. "напомни в 11" had a time
  and nothing to say at 11, and she asked nothing at all — she now asks
  "О чём напомнить?". Subject first, since a reminder with no subject is not
  worth setting.
- The expiry notice is five phrasings picked at random instead of one fixed
  sentence. It is the line he hears every time he walks off mid-request, so it
  is the line that repeats most.
- The nudge prompt's ban on "обращения" is now "ласковые обращения". It was
  meant to forbid "милый"/"дорогой", not his name — "Ками, ноутбук на трёх
  процентах" is how she talks, and the eval's cringe check already only flags
  pet names.
- The nudge example no longer claims she plugged the laptop in. She has no
  hands and no smart plug; an example where she acts teaches the model to
  invent actions Maven never took.
- replySystem: "тепло" → "спокойно и без официальных формулировок". A one-word
  mood instruction a 1.7B can't act on, replaced with the behaviour meant.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 21:52:12 +04:00
kami c97aebf55a Merge tonight's work: all 46 reviewed PRs as one verified branch
135 commits. make build produces all 8 binaries; make test exits 0 across 38 packages with no failures, no data races, gofmt and vet clean.

See PR #47 for what had to be fixed to make it build as a unit.
2026-07-31 19:41:52 +02:00
kami 891136c65d gofmt the kiwix client and rewrite test
PR #41 and #44 landed these two files unformatted, so the gofmt gate that
PR #12 added to `make test` failed as soon as both were on one branch.
Struct-tag and comment alignment only, no semantic change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 21:34:03 +04:00
kami 41c7c13f42 Merge remote-tracking branch 'origin/overnight/snooze-works' into integration/jul31
# Conflicts:
#	internal/store/migrations.go
2026-07-31 21:32:18 +04:00
kami a324e8f624 Merge remote-tracking branch 'origin/overnight/eval-writeup' into integration/jul31 2026-07-31 21:31:36 +04:00
kami 51805e7f35 Merge remote-tracking branch 'origin/overnight/kiwix-rewrite' into integration/jul31 2026-07-31 21:31:36 +04:00
kami 533f0acda8 Lead the bake-off with the answer, not the superseded one
The file ran two sweeps and the second one changed the resident model, but
the lede still opened with "Recommendation: keep Qwen3.5-0.8B". Anyone
landing on the file read the wrong conclusion and had to scroll 100 lines
to find that it had been replaced — and it contradicted CLAUDE.md, which
already says the resident model is Qwen3-1.7B.

Both sweeps are accurate, so nothing is rewritten. The lede now states the
outcome and the first sweep's verdict is scoped to what it actually tested:
it rejects LFM2.5-1.2B, which still holds. It never was a case for keeping
0.8B as the resident model.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 21:17:07 +04:00
kami 4f59ba78c6 Write down the five-model sweep and why the 1.7B won
Numbers behind the resident-model change, plus the answer to "could a 230-350M
model do this instead" — no, and the reason is worth keeping: LFM2.5's published
instruction-following scores beat Qwen3.5-0.8B, and every one of those benchmarks
except Multi-IF is English. In Russian the 350M invents non-words and the 230M
answers in Spanish.

Also fills the row TALK-EVAL-31-07-2026.md had to void for contamination, and
corrects a wrong call I nearly made: the 1.7B's 16s p95 looked like the reasoning
trace, but the 0.8B sits at 17s in every run and the 1.7B beat it twice out of
three. The long tail is shared and is not the Thinking block.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 19:08:37 +04:00
kami d0afd9d4f6 Make Qwen3-1.7B the resident model
Stock Qwen3-1.7B, not the CPT'd one — that training is still running. It won
on both fixtures we have, measured tonight on an otherwise idle box:

  routing, 77 RU cases, intent-only:  67.5%  vs  59.7%  for Qwen3.5-0.8B
  talk fixture, 27 cases:             20/27  vs  11-17/27

It also beat Qwen3.5-2B, which is 20% larger, on every routing column.

Two other things came with it:

n_ctx goes 2048 -> 4096. This is a Thinking variant, so reasoning tokens need
the room, and 4096 is the context every score above was measured at. Shipping
2048 would ship something nobody measured.

The doc now says not to bother with sub-500M models, because I checked and they
are not close. LFM2.5-350M routes at 5.2% — worse than guessing among 7 intents
— and answers "столица Франции?" with "Сторзит", which is not a word. The 230M
replies to Russian in Spanish. Their published IFEval and BFCL numbers are good
and they are all English.

Note the routing gain needs the LLM router actually wired on to show up. It is
still nil, so this commit buys the phrasing improvement today and the routing
improvement when that lands.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:58:20 +04:00
kami 6b67e6f3c2 Word nudges from templates by default, model optional
DEPRECATION, flagged not asked: LLM-phrased nudges are no longer the default.
LLMPhraser.PhraseNudge now returns a hand-written Russian template. The model
still phrases chat, queries and reminders — only nudges moved.

Why: measured over many runs, Qwen3.5-0.8B wrote formal "вы" and plural
imperatives, used masculine self-reference, and invented facts and units
(90-95 seconds to boil an egg). A nudge is five words of known content, so
generation buys nothing and risks the persona every time. Templates score
15/15 on the nudge fixture, the model 11-13/15.

Nothing is deleted: the prompt, the fallbacks and the whole LLM nudge path
stay. Set phraser.llm_nudges = true in deploy/mavend.json to get them back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:21:52 +04:00
kami 742b2ad1d7 Score the Russian-to-keywords rewrite end to end (#403)
Same 9 cases as the retrieval eval, so the numbers compare directly:
hand-written keywords hit 8 of 8, this is what the model reaches on its
own. Reports the hand-written query next to the model's for every case,
because where the phrasing differs is the useful part.

Opt-in on MAVEN_KIWIX_URL + MAVEN_LLM_URL, like the other evals.

Result on Qwen3.5-0.8B: 3 of 8, identical on all three runs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:19:56 +04:00
kami b300ac5c70 Rewrite a Russian question into English Kiwix keywords (#403)
Kiwix ranks by keyword, not meaning, so a translated question finds song
and TV titles. This asks the resident model for the TOPIC instead: a short
English noun phrase, like a Wikipedia article title.

Locked down three ways, because a wrong query is silently wrong:
- A GBNF grammar, same idea as routeGrammar and responseGrammar. The
  reply must be {"query":"..."} with Latin words only. The JSON wrapper
  matters: this model always thinks out loud and this llama-server build
  ignores the thinking switch, so a bare word-list grammar just captured
  "Let me analyze this request carefully" for every question.
- max_tokens 32, since the answer is a few words.
- CleanQuery, which throws away empty, Russian and prose replies rather
  than passing them to Kiwix, and drops question words like "why" and
  "how much" that a keyword ranker cannot use anyway.

Client side only. Nothing is wired into the daemon or any config.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:19:42 +04:00
kami 13e5170e9e Hand-written Russian nudge templates plus a picker
Nudge wording as data instead of generation. The wording lives in
internal/phraser/nudges_ru_v1.json (embedded), about 10 variants per rule:
water, meal, break, service_down, netdata_critical, routine:, morning:, plus
a contentless default. That JSON is long because it is data — the owner can
edit any line of Russian without touching Go.

The picker:
- random, but never the same variant twice in a row for the same rule
- deterministic when seeded (math/rand with an injectable source)
- fills {since} / {service} / {what} from the candidate, and skips any variant
  whose value is missing, so no raw placeholder can reach the piper voice
- {since} is spelled out in words ("полтора часа", "семь часов"), because
  "3 ч" is wrong in a Russian voice

Scores 15/15 on the existing nudge fixture, on every seed swept. Nothing is
wired yet — that is the next commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:18:55 +04:00
kami 0b90952e55 Write down every conversational eval score from tonight
Records all four configurations on the 27-case talk fixture, three runs
each: no grammar, plus grammar, plus Russian prompts, plus the truncation
fix. Composite, per-path and per-check, with the reproduce command.

The short version is that the plumbing got fixed and the score barely
moved. Grammar was the real win. Russian prompts helped a little and cut
latency by 5x. The truncation fix was necessary and bought nothing.

Also writes down three things that are easy to lose:

- The truncation cause was the grammar's 400-character bound, not the
  token cap. Measured at three caps, same 400 characters every time.
- Then I set the bound to 1000 against a 768-token cap and made it worse.
  The two limits have to agree.
- One run is contaminated and marked void: I ran an agent against the same
  llama-server, and the report still claimed zero errors while a third of
  the fixture silently answered "не знаю.". That is #397 and it is worse
  than filed — a busy server is indistinguishable from bad phrasing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:18:19 +04:00
kami aa8f5b2ee2 Make the nonempty check look for actual words
It scored 27/27 on a run where two replies were "{" and "{\n  \"". It only
tested that the string was not blank, so punctuation counted as content and
the worst replies of the run passed the first check.

Now a reply needs at least one letter, Cyrillic or Latin. Latin counts
because answers about ssd or vpn are legitimately part English.

Digits alone fail too. The same run answered "сколько варить яйцо
вкрутую?" with "15-16" — no unit, no words, and the wrong number as well.
That is not something she said.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:57:39 +04:00
kami d7cdcb63bd Stop shipping half-written JSON as a reply
Two bugs, one symptom. A run of the talk eval produced replies that were
literally "{" and "{\n  \"" — those strings went out as things Maven said.

First bug: the parser could not tell "the model answered in plain prose"
from "the model started a JSON object and got cut off". Both came back as
empty, and every caller then shipped the raw text. Now an unfinished object
returns an error and each caller uses its own fallback instead. Bare prose
with no JSON in it still passes through, because small models do sometimes
answer that way and the reply is fine.

Second bug, and the actual cause: the grammar capped the response field at
400 characters. I measured it against Qwen3.5-0.8B at three different token
caps — 256, 768 and 2048 — and the reply came back exactly 400 characters
every time, cut mid-word. So the token limit was never what stopped it.
The bound is 1000 now, about six Russian sentences, still low enough to cut
off a repetition loop.

Token caps go from 256 to 768 on the chat and query paths so 1000
characters of Russian actually fits. The nudge path keeps its own cap; a
nudge is meant to be one sentence.

Note: cmd/mavend/replier_llm.go has its own copy of this parser with the
same bug. Left alone here so this commit stays small — that duplicate is
Vikunja #396.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:56:55 +04:00
kami ddb658ffbb Add a Kiwix search client and score retrieval (Vikunja #403)
Step one of letting Maven read instead of recall. No LLM yet.

internal/kiwix/client.go: search a local Kiwix server, parse the RSS
reply, hand back title + path + plain-text snippet + word count. The
snippet is the unit of context; a full article is ~100KB of HTML and
will not fit a 4096 token window.

internal/kiwix/retrieval_eval.go plus knowledge_v1.json: the 9 knowledge
questions from the phrasing fixture, each with hand-written English
keywords, scored on whether a wanted article comes back in the top 5.
Opt-in via MAVEN_KIWIX_URL, since CI has no Kiwix. No pass bar, the
number is the finding.

Result on the live mirror: 8/8 answerable questions hit, 7 of them at
rank 1. Retrieval works. Keywords are written by hand on purpose, since
Kiwix ranks by keyword and not by meaning, so a natural question fails.
A query-rewrite step is the next piece of work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:43:44 +04:00
kami c7dadc97d9 Write the chat and notes prompts in Russian
The reply has to be Russian, but two of the phrasing prompts told her
what to do in English. Both are Russian now, in the same style as the
nudge prompt that already works better.

Also dropped the "you are maven, a self-hosted personal assistant"
line from both. The persona block right above it already says who she
is, so it was said twice.

The JSON part is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:38:36 +04:00
kami c9d88c152e Drop "never phones home" as a hard rule
The owner's call, 2026-07-31: a 0.8B model does not know enough about the
world to be useful without reading something. So she may now read external
sources to answer world questions.

What replaces the old rule, in all three docs:

- No telemetry, no cloud model, no third-party account. Unchanged.
- Local first: the Kiwix ZIMs on the box before anything on the network.
- External search is allowed but off unless configured, same as weather
  and telegram.
- His notes and facts are never search input. Only the utterance goes out
  — never the persona block, the history, or matched notes.

Docs only, no code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:34:53 +04:00
kami 1890ff5d5d Constrain the phrasing output with a GBNF grammar
The 0.8B answered about one chat turn in three with open reasoning as plain text, so no JSON ever closed and the fallback shipped "Thinking Process:" to the user. A grammar makes that output impossible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:18:06 +04:00
kami 0110e9bc8c Report every address break, and stop -те verbs blinding the check
From a real reply in a nudge eval run: "Смотрите на его потребление
воды" is a plural imperative AND third person about him. Only the plural
printed.

Two separate faults. The check returned on its first hit, so the second
break stayed invisible and the failure read as milder than it was; it now
joins them. And "его" was not detected at all — looksVerb knows the
-й/-йте imperative but not the -те plural, so "смотрите" counted as the
person being talked about, which is what an antecedent means here.
pluralVerb already knows that form, so the antecedent test uses it too.

Third time a verb form has blinded this check. A fourth means it wants a
morphology table rather than another suffix.
2026-07-31 16:52:16 +04:00
kami 50ca8c8b5a Score the chat, query and knowledge phrasing paths (#395)
The phrasing fixture was 15 nudge cases, so every prompt change we
measured only told us about nudges. But the shared context block sits in
front of five prompts, and three of them — chat, note query, general
knowledge — had no scorer at all. Those are the long free-form replies,
where a persona break is most likely and where nothing could see one.

27 cases, nine per path. Nine rather than five because the nudge fixture
already cannot resolve a change smaller than about three cases, and a
per-path score off five would be worse.

Reuses the persona checks instead of copying them. Length, mood and
"no questions" are left out on purpose: these paths return no mood, and
a follow-up question is a feature in chat, not a fault.

The run refuses to score unless the model answers before and after it.
PhraseChat and PhraseQuery swallow model errors and return a canned
string, so without that guard a dead server produces a full report with
zero errors and a bad score — which reads as bad phrasing rather than as
nothing measured. Vikunja #397 is the real fix.
2026-07-31 16:51:52 +04:00
kami de09471421 Merge the shared prompt context block 2026-07-31 16:07:35 +04:00
kami d65c16a567 Don't tell her she can't talk
The block listed what she can do and ended with "nothing else". It sits
in front of the chat and general-knowledge prompts too, so that told her
to refuse the exact thing those prompts are for. Talking is now first in
the list, and the closing line limits ACTIONS rather than everything.

Also dropped the self-introduction from the knowledge prompt. It said
"Мавена, персональный ассистент" — a different name and a masculine
noun, right after the block says she is Maven and feminine. Identity
lives in the block now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 16:07:35 +04:00
kami 062d4252ef Tell her what she can actually do
The context block now lists her real capabilities: reminders, notes and
facts (write and recall), and the calendar — all three are code paths in
mavend today. Weather, telegram and shell acts are listed only when the
config actually has them, because offering something she cannot do is
worse than staying quiet about it.

Also drops the pronouns from the optional name/city line. The block's
own "ты" is Maven, so "тебя зовут" read as her name and "его" would have
shown her the third-person form she must never use about him. They are
plain labels now.

Vikunja #394.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 15:57:58 +04:00
kami 2c27e2ce1f Give every prompt one shared context block
The "address him as ты" rule had only reached two of the five system
prompts. Instead of pasting it into the other three (five copies drift —
that is how this happened), there is now one block, in internal/persona,
prepended to all five: nudges, action replies, chat, note queries and
general knowledge.

The block says who he is and how to address him (a man, always "ты",
never "вы", never "он" about him; Maven stays feminine), plus the
current local date and time. It is rendered fresh each turn because the
time changes, and it is correct with an empty config — the address and
gender rules are defaults in code. Config only adds optional facts:
owner_name, city, and the existing free-text `persona` string, which is
now the static half of the block.

Russian even in front of the English prompts: the rules are Russian
grammar, so they read best stated in Russian, and there is one copy.

Vikunja #394.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 15:55:30 +04:00
kami ccc5cba2a3 Merge the example-led prompt finding 2026-07-31 15:41:01 +04:00
kami 89d83c0b11 Record the example-led nudge prompt experiment (#393) — it made things worse
Tried rewriting the nudge prompt to lead with five on-topic examples instead
of rules. Three eval runs each side: before 12/13/14 of 15, after 11/12/11.
The loss is all in the address check — formal "вы" and plural imperatives came
back once the "говоришь на ты" rule stopped being its own sentence, and the
on-topic examples leaked their wording into the wrong cases.

Prompt reverted. Only the finding is committed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 15:40:18 +04:00