Commit Graph

167 Commits

Author SHA1 Message Date
kami a40bc559d5 Fix the nudge phrasing prompt: stop teaching the model to echo the example
The system prompt showed the JSON contract as {"response": "..."} and the
user prompt repeated it. A 0.8B copies whatever sits in the response slot, so
7 of 15 nudges came back as literally "...".

Changes, all prompt-side — the {"response","mood"} contract is unchanged:
- nudge system prompt is Russian, feminine self-reference, with filled-in
  examples on topics that never appear as rules, so copying them is visible
- rule names get a Russian gloss and a required keyword, named last in the
  prompt where a small model weights it hardest
- durations render in Russian, not English
- the no-parse fallback says something Russian instead of "water — care",
  which was going straight to a Russian piper voice
- same "..." placeholder removed from replier_llm.go

Scored on internal/phraser/eval: 0/15 -> 13/15.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:23:03 +04:00
kami 1c4eab2107 Merge commit '94eb92f' into overnight-jul31
# Conflicts:
#	Makefile
2026-07-31 10:11:56 +04:00
kami 0914e0a3d5 Merge commit '4ba9a6f' into overnight-jul31
# Conflicts:
#	Makefile
2026-07-31 10:11:29 +04:00
kami c860808528 Make make test actually gate on gofmt and vet
DESIGN.md has always said `make test` is "gofmt + vet + -race, no
exceptions". It only ever ran the tests, which is how nine files drifted
out of format without anyone noticing.

`test` now depends on `fmt-check` and `vet`. Checked that fmt-check does
fail when a file is unformatted, so the gate is real and not decorative.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 10:10:27 +04:00
kami f7442c3aea Run gofmt over the seven files that had drifted
Formatting only: import order, and statements that were packed onto one
line split out. `git diff -w` shows nothing but that.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 10:09:32 +04:00
kami 75b067ac51 Merge the accepted-routine fix and drop reminder_id from accept
Two merge fixes on top of the branch:

- migrations: keep both new steps, snooze stays #8, the routine columns
  become #9. Both agents had numbered theirs #8.
- accepting no longer takes a reminder id, on the web surface too. The
  web accept path had the same one-shot-reminder bug the voice path did,
  so both now just flip the status and let the tick loop schedule.

The test that asserted "accept creates a reminder and links it" asserted
the bug. It now asserts that accepting creates no reminder.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 10:07:37 +04:00
kami 424d1b3446 Fire accepted routines every interval, not once (Vikunja #366)
The tick loop now reads accepted routines from the store and nudges when
their interval has passed; accepting no longer builds a one-shot reminder.
Look at routine.DueAccepted for the schedule rule (no catch-up backlog) and
at fireAcceptedRoutines for the restraint gate — routines do not bypass it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:45:30 +04:00
kami c47886c2bc Merge branch 'worktree-agent-ab5b5c61a32cac4fe' into overnight-jul31 2026-07-31 02:42:53 +04:00
kami f6236da760 Collapse the duplicate away-detail and panic tests
Two agents wrote the same three test helpers and names for the same two
bugs. Kept the real assertions from dispatcher_test.go and removed the
skipped placeholders they replace. panicSink stays in durability_test.go
since both files use it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:42:21 +04:00
kami 54dc43516b Add accepted-routine timestamps to the store (Vikunja #366)
Data layer only. Migration #8 adds accepted_ts and last_fired_ts to
proposed_routines, plus ListAcceptedRoutines and MarkRoutineFired so the
tick loop can own the schedule. Accepting no longer links a reminder id.
Look at the TODO(vikunja#366) in cmd/mavend/tick.go for the next commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:41:58 +04:00
kami fa51a48958 Merge branch 'worktree-agent-afe3f2ec18b2b8497' into overnight-jul31 2026-07-31 02:40:10 +04:00
kami 5fd25d7ad7 Test the away-channel minimal body and the panicking sink (#368, #369)
The integration branch names one test TestAwayFallsBackToFullBodyWhenSummaryEmpty,
which describes the old bug; it is here as TestAwaySendsGenericLineWhenSummaryEmpty
and asserts the generic line instead of the body. Also covers: a normal summary
goes out unchanged, voice keeps the full body, and one panicking sink does not
eat the other channel for the same nudge. Reformatted one pre-existing struct.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:39:38 +04:00
kami c22fc352fc Merge branch 'worktree-agent-a1610b8c5376eadd6' into overnight-jul31 2026-07-31 02:37:30 +04:00
kami 215aa331c5 Recover from a panicking sink so the attempt is always closed (#369)
A panic in Send used to unwind past completeOutbox and leave the
delivery_attempts row pending forever, since reconciliation only runs at
startup. safeSend turns the panic into an error, logs it loudly, records the
attempt failed, and lets the other channels for the same nudge still go out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:35:54 +04:00
kami 859bbf750f Never send a nudge body off-box when the summary is empty (#368)
Away channels (ntfy, telegram) leave the box, so an empty Summary now sends
a fixed generic line plus the rule name instead of the full Body. The
dispatcher strips detail before any sink sees it, so a sink added later
cannot leak by reading the wrong field. Voice is local and unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:34:54 +04:00
kami 8a174c1c70 Score the recall fixture and write up what it shows
Real recall is 48% after the gate, and one must-be-silent query gets an
answer anyway. Review finding 2 (the score distributions overlap, so no
gate separates a real recall from a false one) and finding 4 (the memStore
branch at voice.go:776 is unreachable for notes). Adds an embedder cache
so the gate sweep does not re-embed the fixture nine times.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:34:28 +04:00
kami 1f55207b58 Merge the snooze read path and wiring
# Conflicts:
#	internal/loop/gate_test.go
2026-07-31 02:33:41 +04:00
kami 4db109346a Ignore the .claude directory
Agent worktrees land in .claude/worktrees, so the directory shows up as
untracked noise in every git status.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:32:44 +04:00
kami 2f00593411 Wire the snooze read into the Gatherer and honour it for reminders (#364)
The Gatherer now fills State.SnoozeUntil from store.SnoozedUntil instead
of nil, so a snooze finally reaches the gate. RemindDecisions gains the
one restraint check that applies to a reminder — quiet hours, presence
and cooldown are still bypassed, so "wake me 7" is unchanged. Reviewer:
the two tests in internal/loop/gate_test.go are the contract.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:32:24 +04:00
kami 784d688b44 Merge the clarify wiring
# Conflicts:
#	internal/dialogue/clarify.go
#	internal/dialogue/clarify_test.go
2026-07-31 02:30:59 +04:00
kami 4ba9a6f422 Add a deterministic scorer for nudge phrasing (Vikunja #323)
Review internal/phraser/eval/checks.go -- it IS the measurement. Each check
names in a comment which DESIGN.md line it defends: length, feminine
self-reference (windowed around "я" so the operator's own masculine
second-person forms are not flagged), the cringe list (pet names, emoji,
"!!", fake concern, apology, emotional support, asking how he feels,
praise), on-topic, mood enum. No send/veto signal anywhere, per
DESIGN.md § "Rules decide, LLM phrases".
Fixture (158 lines) and tests (252) do not count toward the diff ceiling;
the scorer itself is still ~650. Splitting eval.go from checks.go would
give two commits neither of which measures anything.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:30:52 +04:00
kami 32687b3712 Read the recorded snooze outcomes back out of the nudges table (#364)
The gate honours State.SnoozeUntil but nothing ever filled it. New
store.SnoozedUntil returns, per rule, when the newest snooze runs out.
Reviewer: the fixed 2h SnoozeDuration and its reasoning in nudges.go —
nothing upstream can supply a per-nudge length, so no new column.
Expired snoozes are dropped in SQL, so silence can never be permanent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:29:51 +04:00
kami db3e706bdc Merge the delivery routing and durability tests 2026-07-31 02:29:05 +04:00
kami 2f4257e194 Test the clarify round-trip end to end at the daemon level
Covers: a reminder with no time is asked about and completes on the answer; the
same for a fact; an answer past the TTL falls through as a fresh utterance; a
second unclear answer drops the request with no second question; a clarified act
off the allowlist neither runs nor gets enabled; a clarified destructive act
still parks a confirm; noise keeps the canned reply. No model, no network.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:29:04 +04:00
kami 7d8b0af99d Test voice fallthrough and delivery durability
Fallthrough is checked per severity through the outbox trail, so sev3/sev4
reroute and sev1/sev2 still drop. Durability uses a real store on a temp file:
a crash between Begin and Complete becomes unknown, is not resent, is not
dropped, and a late Complete cannot overwrite it. One skipped test marks a real
gap: a panic mid-send leaves a permanent pending row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:27:33 +04:00
kami cfd38d53cf Ask the question, then act on the answer
On a clarify decision with one identifiable gap she now asks instead of saying
"не поняла", and parks the request. The next utterance is parsed as the answer
with the router's own extractor and the completed decision runs through
applyAction like any other — so a clarified act still needs the allowlist and
still hits the destructive confirm gate. An answer that does not fill the gap
drops the request; she never asks twice. Also pulls the session-store block
that HandlePushToTalk and handleText both had into rememberTurn, since the
clarify path needed a third copy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:26:11 +04:00
kami a2835bbdf6 Cover every cell of the delivery routing table
Table-driven tests for all four severity bands crossed with present and away,
both as the pure table and end to end through the dispatcher. Three tests are
written to DESIGN.md and skipped because the code does not keep the claim: the
care-away drop is recorded nowhere, and the minimal body is enforced per-sink
rather than by the dispatcher. Also gofmt'd dispatcher_test.go.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:24:43 +04:00
kami c262c1e4c0 Merge the routine accept path and store gaps 2026-07-31 02:24:07 +04:00
kami a33ad82178 Let the /routines page accept a proposal, gated at step-up (#46)
Accepting a routine gives the trigger loop a new standing reason to speak to
the human, so it is the same authority tier as enabling a tool and shares the
stepUpOK gate; dismiss only ever makes maven quieter, so it is ungated.
Look at handleRoutines and acceptRoutine in cmd/mavweb/main.go: accept creates
the recurring reminder, then links it via the new ipc AcceptProposedRoutine.
The page now says what maven noticed in her own words (pattern.PhraseRoutine).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:23:28 +04:00
kami af35ec3629 Work out which slot is missing and phrase one short question
A table per intent (reminder needs a time, fact needs a key, act needs a fn)
plus one fixed Russian question per slot. Templates, not model output: a 0.8B
would wander and a question that rewords itself is harder to answer. Note,
query, chat and system get no question — for those a clarify decision keeps
the canned reply rather than inventing a question for noise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:22:18 +04:00
kami 9145b83100 Add the clarify data layer: a parked question with one missing slot
The router can already say "I am not sure" (Decision.Clarify) but the daemon
had nowhere to keep the request while it asked. PendingQuestion holds the
original slots, ClarifyStore parks one per dialogue id with a 90s TTL, and
Answer fills only the slots that were missing so an answer can never rewrite
what she already understood. Logic that uses this comes next.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:20:28 +04:00
kami 43470abc57 Add a held-out note-recall harness (fixture + scorer)
Measures whether Maven can find the right note again from a paraphrased
question. Review internal/memory/recalleval/recalleval.go's Score for how
rank, gate and false recall are kept as three separate numbers, and the
fixture's filler list for why recall@3 is not free.
Fixture JSON is generated data and does not count toward the diff limit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:18:58 +04:00
kami 94eb92fb15 Label LLM eval runs with the model llama-server has loaded
The bake-off in #278/#250 needs two models' scores side by side, and the
report names only carried the config, so the rows were indistinguishable.
ModelID reads /v1/models instead of taking a string that goes stale.
New target: make eval-models MAVEN_LLM_URL=...

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:17:20 +04:00
kami 707c3e5040 Ignore deps and models as symlinks, not just directories
.gitignore had deps/ and /models/llm/ with trailing slashes. A trailing slash
only matches a real directory, so a *symlink* with the same name is not ignored
and git add -A commits it as a symlink blob.

That bites anyone working in a git worktree, where deps/ and models/ do not
exist and have to be linked in from the main checkout. It already happened once
tonight.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:17:15 +04:00
kami ace9fbbc06 Merge the llm_router flag and the kill-maven fix 2026-07-31 02:17:00 +04:00
kami 3884db33e9 Give proposed routines a status filter and pin down the dedup rule (#46)
Look at internal/store/proposed_routines.go: status flips in place with an
`AND status = 'proposed'` guard, not append-only like facts/voids_id — a
proposal is a question with one answer, same shape as tools.status. The
UNIQUE(action, object) key is what stops a dismissed routine coming back.
New tests cover re-propose-after-dismiss and listing by status.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:16:39 +04:00
kami dfb8d26b62 Fix kill-maven.sh so it actually kills llama-server
The MODEL default was LFM2, but the deploy runs Qwen3.5-0.8B, so the
pkill pattern matched nothing and the server survived every kill.
Now matches any llama-server serving a .gguf, so changing the model in
deploy/mavend.json cannot break the script again. MODEL still narrows it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:15:30 +04:00
kami bf99fd4192 Add a voice.llm_router flag, default off
Wires cmd/mavend/voice.go to build the LLM router when the operator asks
for it. Default false, so nothing changes on the deploy box.
Look at pickLLMRouter: the flag on with no llama-server logs one line and
keeps the classifier, it never fails a turn.
The default stays off until the router can refuse (#359) and the extractor
runs on LLM decisions — both noted as TODOs in config.go.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:15:30 +04:00
kami 6d9aa83b6e Merge the proactive rule and gate tests 2026-07-31 02:15:15 +04:00
kami f75072175d Merge the pending-question data layer 2026-07-31 02:14:33 +04:00
kami 39d83a33e8 Add the pending-question data layer for slot clarification
PendingQuestion plus ClarifyStore: same shape, locking and expiry as SessionStore. Answer fills only the missing slots and never overwrites a filled one. No wiring yet — TODOs mark the daemon hooks.
Reviewer: MaxAttempts is 1 on purpose (Maven asks once, she is not a nag).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:13:48 +04:00
kami 9d8fcf42f3 Test the universal restraint gate, including two gaps
Pins the conservative side of Gate(): quiet hours, away, calendar-busy,
cooldown and snooze, plus one nudge per tick at max severity. Reviewers: the
two skipped tests at the bottom are real gaps, not flakes. Reminders ignore
snooze (loop.go:120) and the Gatherer never fills SnoozeUntil (gather.go:153),
so snooze does nothing at runtime. No behaviour was changed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:13:46 +04:00
kami 925ce223a0 Add a Value slot to dialogue.Slots
router.Slots already carries the fact payload; the dialogue copy did not, so a clarifying answer had nowhere to put it. InheritSlots carries it like Key.
Reviewer: check the new inherit block does not overwrite a filled value.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:13:40 +04:00
kami 9e2af3286b Unit-test all five proactive rule predicates
Table-driven tests for water, meal, break, service_down and netdata_critical,
straight against the predicate with a fake State. Reviewers: the no-data rows
(every rule must stay quiet when its key is missing) and the ops forgery rows,
where a fact with the right value but the wrong source must be refused.
No rule fired on missing data, so no fix was needed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:12:25 +04:00
kami d30618ecb7 Stop the router repetition loop
Route now sets RepeatPenalty on the request, and the grammar's string rule is
capped at 120 characters. Two of 76 fixture cases looped one sentence inside
the text field until MaxTokens, which cut the JSON in half.
Reviewers: the new constant and the grammar string rule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:11:38 +04:00
kami 17b47ce206 Route questions to query, not fact
The router prompt tested "reports current state -> fact" before "wants
information -> query", so a question naming a fact key was written as a fact.
Query now comes first, plus an explicit question test.
Reviewers: the prompt block in llmrouter.go, and the note about the
training-side copy of the prompt that needs the same edit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:10:37 +04:00
kami abe9b28719 Write up the routing evaluation findings and next steps
Settles the measurement half of Vikunja #319: the resident 0.8B routes
better than the deployed classifier (50.0% vs 36.8% intent-only) at ~27x
the latency, and neither path can refuse an ambiguous utterance.

Records the four-configuration comparison, the per-case evidence behind
each finding (query->fact x15 traced to routeSystem's rule order, the
five missed-clarify cosines, the two text-field repetition truncations),
the two hypotheses that were tested and closed (thinking mode, grammar
array runaway), and a next-steps list mapped to #319/#320/#359.

#320 should not flip as-is: it would remove the refusal lane rather than
improve it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X5JApcrCRVGmqrxnhynSik
2026-07-31 00:55:32 +04:00
kami 46259b4571 Score the LLM router on Qwen3.5-0.8B against the routing fixture (#319)
Completes #319's comparison. Three configurations, because "the LLM router"
was ambiguous: the model alone, the cascade #320 would actually ship (stage-0
grammar → model → classifier floor), and a thinking-off diagnostic.

                      intent-only  full    RU     missed-clarify  p50
  classifier+onnx     36.8%        36.8%   25/61  5/6             31ms
  llm-only (0.8B)     48.7%        23.7%   13/61  6/6             850ms
  cascade+llm (0.8B)  50.0%        32.9%   18/61  6/6             825ms

On the question asked — does the resident model route better? — yes, 50.0%
vs 36.8% intent accuracy. REARCH.md's premise holds. It costs 27x the
latency (p50 825ms vs 31ms, max 3.1s), on the same llama-server the phraser
needs, so it is a trade rather than a free win.

Three things the numbers surface that the headline hides:

query→fact x15 is the dominant failure, four times the classifier's x4 on
the same axis. routeSystem's decision order puts "сообщает или обновляет
состояние" (rule 3) above "хочет получить информацию" (rule 4), so any
utterance naming a fact key matches the earlier rule and a question about
past state reads as an assertion of it. A prompt fix, not a model limit.

The LLM router cannot clarify: llmrouter.go hardcodes Confidence 1.0, so
stage 3's gate can never fire on its decisions — 6/6 missed. With #359's
finding that the classifier's gate is miscalibrated under ONNX, neither path
currently refuses. Flipping #320 as-is removes the refusal lane.

The gap between 50.0% intent and 32.9% full accuracy is entirely slots: the
LLM path fills neither Fn nor Time (it returns Slots.Text for acts, and
Extract never runs on an LLM decision).

Also settles a hypothesis rather than leaving it in the air: thinking mode is
a non-issue under a grammar (identical score), and the grammar's unbounded
("," ws action)* repetition that ran away in an isolated smoke test does not
reproduce under the real prompt — 2 errors in 76, not 76. internal/llm
deliberately does not grow a chat_template_kwargs field.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X5JApcrCRVGmqrxnhynSik
2026-07-31 00:50:20 +04:00
kami d34fdf40aa Score the routing fixture with the ONNX embedder (Vikunja #319)
The onnxruntime .so was already vendored at deps/onnxruntime-linux-x64-1.26.0
— nothing to download. make eval-router now defaults MAVEN_ONNX_LIB there, so
both baselines run by default and only a fresh clone without deps/ falls back
to the hash ratchet alone.

Prod-representative result, deployed 0.55 gate: 28/76 (36.8%), RU 25/61,
EN 3/15, hard 0/11 → 4/11, p50 31ms / p95 71ms. Versus the hash floor's
13/76 at p50 9µs.

The finding is not the accuracy, it's the refusal lane: missed clarifies went
0 → 5 of 6. Better embeddings raise cosine everywhere, so the 0.55 threshold
that used to hold ambiguous utterances back stops holding — "сделай это"
routes to act at 0.847, "бэкап" to chat at 0.755. The gate was implicitly
tuned to the hash floor's low similarities. That is an argument about the
threshold, not about the embedder, and it lands before #320 rather than after.

Also fixes a fixture-model mismatch: ReminderGrammar deliberately skips the
extractor at stage 0 and the daemon's applyAction parses the time downstream
(stage0.go says so). Charging the router for that slot made 4 exact-match wins
read as misses; they are now counted as SlotsDeferred instead. Hash baseline
moves 13/76, ratchet to 0.15.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X5JApcrCRVGmqrxnhynSik
2026-07-31 00:34:29 +04:00
kami c7c44229a2 Add held-out RU routing fixture and scorer (Vikunja #319)
#319 asks for a measurement before #320 flips the route decider from the
classifier cascade to the resident model. There was nothing to measure
against: the only routing tests assert single utterances, and the
classifier's seed corpus is its own training set — scoring it there
measures memorisation of frozen centroids, which is the illusion that hid
the weak RU query handling in the first place.

internal/router/eval is a separate package so both paths can be scored
from outside router (including cmd/mavend, where the real llama-server
client lives). The fixture is embedded; the scorer takes a Router
interface, so *router.Router and a bare LLM stage both go through the same
76 cases.

The fixture is a CONTRACT, not a snapshot: cases the cascade fails today
stay in the file and fail loudly. TestFixtureIsHeldOut enforces that no
utterance appears verbatim in models/seeds/*.txt.

Baseline, hash embedder at the deployed 0.55 gate: 9/76 (11.8%), 63 false
clarifies, 0 missed clarifies, p50 9µs. Almost everything falls to the
confidence gate — the documented floor behaviour, not a new bug. The
number worth comparing is TestONNXBaseline's (skipped without
MAVEN_ONNX_LIB); the assertions here are a regression ratchet plus a tight
bound on the dangerous direction: ambiguous utterances must not start
being routed confidently.

Seeding is order-fixed on purpose — a few phrases appear under two intents
and map iteration handed them to a different centroid each run, which made
the score jitter between 9 and 10.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X5JApcrCRVGmqrxnhynSik
2026-07-31 00:28:44 +04:00