Commit Graph

12 Commits

Author SHA1 Message Date
kami 079cf689aa router: agenda questions belong to query, not to system
"что у меня сегодня" and "что у меня в календаре сегодня" both routed
IntentSystem on the deployed daemon, and replySystem has no agenda arm,
so both answered "пока не умею". The calendar source that can answer
them lives in the query chain and was never reached. The fixture has
said query since ru-query-019 was written; the daemon disagreed with the
fixture and the daemon was wrong.

AgendaQueryGrammars routes them at stage 0, after the clock rules so
"какой сегодня день" keeps reaching replySystem. Intent only — which
source claims the turn stays the query chain's decision.

This is what made the follow-up continuation look like it only worked
for "what day is it". It did: the query half inherited an intent whose
handler could not answer, so both halves came back "пока не умею".

Measured on the 77-case RU fixture: full accuracy 70.1% → 72.7%,
intent-only 75.3% → 77.9%, calendar 0/2 → 2/2, clarify counts unchanged.
The eval harness wires the new grammars too, or the fixture would stop
being a measurement of the daemon.

Go's \b is ASCII-only and never fires after a Cyrillic letter, which the
first version of the pattern learned the hard way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 22:53:01 +04:00
kami 29329b5f0e router: a Russian verb is a whole sentence, not thin evidence
The clarify gate thinned any one-word utterance to 0.3 confidence, which
trips the stage-3 gate and comes back as "не совсем поняла". That is an
English intuition. Russian packs subject, tense and gender into one word,
so "поужинал" is a complete report and "привет" a complete greeting, and
both got clarified.

thinSingleToken keeps the rule for bare nominals, where it is real ("вода"
is a fact-or-query coin flip), and spares two classes: a closed lexicon of
social and control singles, and any token carrying a verb ending. Both
tests are offline.

Fixture: false clarifies 3 → 2, intent-only 74.0% → 75.3%, full accuracy
unchanged at 70.1%, missed clarify still 1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 21:29:21 +04:00
kami ee7bec11e3 Add mavmaild, the read-only IMAP poller that feeds mail intake (#246)
The extraction seam landed on the previous branch but nothing fed it. This
adds the daemon that does: every interval it opens one mailbox read-only
(EXAMINE + BODY.PEEK, so reading leaves no \Seen behind), fetches the UIDs
it has not handed over yet, and posts each message to core over
ingest_mail. Core runs the model and writes task candidates; this daemon
writes nothing and cannot create a reminder.

It is a separate daemon because of the credential. mavpoll set the
precedent with the zenmoney token (#125): the module talking to the third
party holds the secret, reads it from a file so it never lands in argv, in
docker-compose.yml or in shell history, and core never sees it. There is
deliberately no -password flag, and a test asserts that.

Off unless configured at both ends: without -password-file the daemon
refuses to start, and if core has no email block the first ingest returns
ErrUnknownMethod, which disables the reader instead of hammering a socket
that will keep refusing. A seen-UID state file (0600, atomic write) keeps a
restart from re-extracting the whole lookback window; correctness does not
depend on it, since capture dedupes on normalised text. Logs are counts and
UIDs — no subject, sender or body.

Verified with an in-process IMAP server and a fake core: bulk mail is
filtered before core is asked, seen UIDs are not re-fetched, a failed
ingest is retried next poll, ErrUnknownMethod stops at the first message,
and state survives a restart. The live half is untested by design — no IMAP
credential exists on this box; setup is written up as QA steps.

Vikunja #246
2026-08-01 03:13:18 +04:00
kami a2031a31d1 Record the measured confidence-gate numbers 2026-07-31 23:23:37 +04:00
kami f0f7ebc9b2 Give LLM-routed decisions a real confidence so clarify can fire (#359)
Confidence was hardcoded to 1.0 for every LLM decision, and the LLM branch
in Router.Route returned straight from fillSlots without ever touching the
stage-3 threshold gate — so the LLM path could not produce a Clarify no
matter what confidence a model reported. That is why all 6 want_clarify
cases in the 77-case RU fixture were missed by every model in the bake-off.

Fix reads structural signal instead of changing the (parity-locked) router
prompt: a single-token utterance ("вода", "бэкап") is flagged thin evidence
in llmrouter.go; a fact left keyless or an act that never resolves to an
allowlisted fn, checked after fillSlots so the deterministic parsers get
first crack, is flagged in router.go's new gateLLMDecision. Anything below
config.DefaultRouterThreshold (0.55) now sets Clarify=true through the same
path the classifier already uses.

Added unit tests with a stubbed Completer proving both directions: thin
cases clarify, clean multi-word/resolved-slot cases stay confident. The
77-case fixture re-run against a live llama-server is still needed to
confirm the 6/6 moves — not done here, no llama-server on this box.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:07:32 +04:00
kami 0ca5748699 Stop the docs claiming the LLM router is off
CLAUDE.md's routing section said "llmrouter is wired nil" and called the
classifier cascade the committed default. That stopped being true when the
integration merge landed: voice.go:214 wires pickLLMRouter, DefaultLLMRouter is
on, and deploy/mavend.json sets llm_router true. It is the first thing anyone
reads before touching the router, so it was pointing the next reader at a
wiring job that is already done.

Rewritten to say the LLM router is the default, the classifier is the failure
floor and must not be deleted, and what the two actually measure — 36.8% at
p50 31ms against 67.5%/72.7% at p50 ~2.7s, a trade accepted on purpose. Names
the one thing still open on that path: Confidence is hardcoded 1.0 in
llmrouter.go, so the LLM never asks for clarification (#359).

Also in CLAUDE.md: the persona line pointed at a memory file that does not
exist, so the actual rule was nowhere in the repo. Written out instead —
feminine self-reference, informal singular address, pet names forbidden but his
name allowed — plus the three eval checks that enforce it.

MODEL-BAKEOFF: three claims had gone stale within hours of being written. There
IS a make eval-models target now; the routing numbers ARE the production path,
not a bench artifact waiting on a wiring change; and the truncated 293 MB gguf
is deleted. Struck through rather than removed, since the caveats are part of
how the evening read at the time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 22:37:15 +04:00
kami d0afd9d4f6 Make Qwen3-1.7B the resident model
Stock Qwen3-1.7B, not the CPT'd one — that training is still running. It won
on both fixtures we have, measured tonight on an otherwise idle box:

  routing, 77 RU cases, intent-only:  67.5%  vs  59.7%  for Qwen3.5-0.8B
  talk fixture, 27 cases:             20/27  vs  11-17/27

It also beat Qwen3.5-2B, which is 20% larger, on every routing column.

Two other things came with it:

n_ctx goes 2048 -> 4096. This is a Thinking variant, so reasoning tokens need
the room, and 4096 is the context every score above was measured at. Shipping
2048 would ship something nobody measured.

The doc now says not to bother with sub-500M models, because I checked and they
are not close. LFM2.5-350M routes at 5.2% — worse than guessing among 7 intents
— and answers "столица Франции?" with "Сторзит", which is not a word. The 230M
replies to Russian in Spanish. Their published IFEval and BFCL numbers are good
and they are all English.

Note the routing gain needs the LLM router actually wired on to show up. It is
still nil, so this commit buys the phrasing improvement today and the routing
improvement when that lands.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:58:20 +04:00
kami c9d88c152e Drop "never phones home" as a hard rule
The owner's call, 2026-07-31: a 0.8B model does not know enough about the
world to be useful without reading something. So she may now read external
sources to answer world questions.

What replaces the old rule, in all three docs:

- No telemetry, no cloud model, no third-party account. Unchanged.
- Local first: the Kiwix ZIMs on the box before anything on the network.
- External search is allowed but off unless configured, same as weather
  and telegram.
- His notes and facts are never search input. Only the utterance goes out
  — never the persona block, the history, or matched notes.

Docs only, no code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:34:53 +04:00
kami 76a6a007ef Pin the resident model to Qwen3.5-0.8B and name Qwen3-1.7B as the target
The most load-bearing decision in the project was stated four incompatible
ways: the docs said Qwen3-1.7B, deploy/mavend.json said Qwen3.5-2B, the repo's
models/llm/ held an LFM2.5-1.2B gguf, and five code comments still said LFM.
Answering "which model is deployed" meant re-deriving it from scratch every
time.

Two facts the review missed, found while resolving it:

- /mnt/hdd1/llms is bind-mounted over /opt/maven/models/llm, which shadows the
  repo's models/llm/. The LFM2.5 gguf sitting there was never loaded by
  anything, so it was not evidence of the deployed model at all.
- That library holds Qwen3.5-0.8B, -2B and -4B, and no Qwen3-1.7B. The config
  pointed at a file that does exist; the docs' Qwen3-1.7B was the stale claim,
  the reverse of the assumed direction. Qwen3-1.7B is the CPT target, and that
  training is still in flight (Vikunja #122), so no such gguf exists yet.

phraser.model_path moves to Qwen3.5-0.8B (Q4_K_M) — the smallest checkpoint on
disk, chosen for latency, and relevant to whether the LLM router is affordable
on this box. Docs and comments now say the same thing in one voice: 0.8B
resident now, CPT'd Qwen3-1.7B as the target, and the bind-mount shadowing
written down so the next reader does not mistake models/llm/ for ground truth.
Comments name the model, never a filename, so a swap stays a one-line config
change.

n_gpu_layers: 99 is correct and stays — compose passes /dev/dri and the render
gid for Vulkan offload to the Vega iGPU. CLAUDE.md's "CPU-only" was the stale
half of that contradiction and is corrected.

phraser.go also dropped a wrong "sub-1b, prompted not trained" size claim: the
target is trained end-to-end (RU CPT + joint persona/router SFT).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X5JApcrCRVGmqrxnhynSik
2026-07-30 23:40:33 +04:00
kami 5fe8f228c1 feat(mavweb): /ecosystem page consuming Nexus/Praxis/Hexis + shell fixes
Add a read-only /ecosystem page that consumes the sibling services'
JSON APIs (Nexus entities, Praxis attention, Hexis capabilities),
fetched concurrently with honest per-panel error states. Siblings stay
headless — mavweb is their human surface (arch §16). Wired via mavweb
-nexus/-praxis/-hexis flags; mavweb joins the ecosystem compose network.

Fix mobile horizontal overflow across all pages: .content is a flex
child with default min-width:auto, so it refused to shrink below the
tables' intrinsic width. min-width:0 lets wide tables pan inside .scroll
instead of dragging the page sideways. Verified via CDP geometry check
(scrollWidth === clientWidth at 430px).

Also includes in-progress Ethos UI redesign, ecosystem deploy compose,
and planning docs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 22:04:23 +04:00
kami 0c65387a5f feat(router): array route contract + shorter RU router prompt
Route contract is now a JSON array of action objects (one per ask) so
compound utterances route all their intents, not just the first. Grammar
root emits `[{intent...},...]`; parseActions tolerates a bare object.
Cascade still returns one Decision — full N-action dispatch lands with the
engine turn-on (marked in-code).

Router prompt rewritten shorter + decision-ordered (prompt-guy feedback),
fact redefined as "implicit update" not "trackable state", kept in Russian
to match the CPT base + phraser. "интент" → "намерение".

CLAUDE.md: routing-architecture section + refreshed open items.
docs/plans: route-data generation plan.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017GMrVfuYN3nE4L1vEiFYC9
2026-07-11 23:27:44 +04:00
kami 6a5121657a feat: {response,mood} output contract + router removal, TTS piper plan
Daemon side of Decision B: parse {"response","mood"} across the 4 consumers
(replier, nudges, reminders, chat), fall back to legacy formats. Drop the
LLM router — the classifier handles routing; replier/phraser share one
llm.Client (timeout 20s->60s). llm.Client reads reasoning_content when
content is empty (thinking models).

Docs: TTS piper-student plan (OmniVoice teacher -> piper student, from
scratch, phoneme-first). CLAUDE.md training guide.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 22:51:50 +04:00