Compare commits

...

4 Commits

Author SHA1 Message Date
claude e332f167b2 docs: the ranking's two blocking infra items already shipped (V-447)
Checked the doable, epic and infra tiers against the code, not just the two
tiers V-447 asked about. Ten entries are already built. The two the ranking
calls blockers for everything below are among them: sqlcipher at-rest ships as
Store.enc plus OpenEncrypted, and mavweb/mavcaldav have nine test files
between them where the ranking says zero coverage.

Also built and still ranked as work: rule trace, recurring reminders (cron +
RescheduleReminder), stale-reminder burst collapse (collapseReminders),
revert (VoidLatestFact), digest mode, testing infra, passkey persistence.

Recorded as a section at the top of the archived file so the tiers underneath
are read with the corrections in hand. No tasks created: V-447 scoped task
creation to the mandatory and easy tiers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xwomr2cT93KMSz9u3digX5
2026-08-02 09:20:36 +04:00
claude 322401b9af docs: the ranking was stale, two of its gaps already ship (V-447)
Checked the mandatory and easy tiers against the code instead of trusting the
2026-07-03 ranking. Two were already built and their tasks closed unstarted:
quiet hours (QuietHoursConfig + the care gate in internal/loop/loop.go:37) and
schema migrations (internal/store/migrations.go on PRAGMA user_version, 12+
steps shipped).

Three more were narrowed to what is actually missing. Destructive-confirm has
a mechanism and no policy: store.Tool.Destructive is one boolean, not a risk
tier. Bounded follow-up state has dialogue.Session with a TTL and slot
inheritance; what it lacks is Candidates, so "второй" resolves against nothing.
Clarify has a gate that can ask and one hardcoded sentence to ask with
(internal/voice/replier.go:56, which replier_llm.go hands straight through).

The other six were confirmed absent: list_items, capability model, go.mod
tidy, conversation repair, command history, pronunciation dictionary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xwomr2cT93KMSz9u3digX5
2026-08-02 09:13:59 +04:00
claude 4f34a232d4 docs: retire the two root queues, archive the ranking (V-447)
PROGRESS.md and 20-07-2026-BACKLOG.md were state snapshots that git log and
the Vikunja board already carry. Everything PROGRESS.md claimed as shipped is
a QA task. The backlog's only untracked item, bounded follow-up state, is now
V-448.

maven-feature-ranking.md moves to docs/archive/2026-07-03-feature-ranking.md
instead of dying. Its mandatory and easy tiers became V-449 through V-458; the
doable and epic tiers are reasoning about why things are not worth doing yet,
which no task captures.

Four code comments and the design.md ledger pointed at the deleted files.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xwomr2cT93KMSz9u3digX5
2026-08-02 09:11:11 +04:00
claude 93987f2dfc docs: tier the tree by lifetime, so staleness shows in the path (V-446)
Seventeen markdown files at the repo root, twelve of them dated one-shot
reports sitting next to CLAUDE.md. That is why stale docs read as
current: nothing in the path said which was which.

Root now keeps CLAUDE.md and AGENTS.md. Living docs move under docs/
and carry a Last verified line. Dated measurements move to docs/evals/
ISO-prefixed, and are never edited after the day, so a newer number is
a new file. The senior review moves to docs/archive/.

Every reference was rewritten across markdown, Go comments, the Makefile
and the recall fixture. The touched Go packages still build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 03:28:49 +04:00
43 changed files with 141 additions and 953 deletions
-396
View File
@@ -1,396 +0,0 @@
beyond the model and tts work, the useful additions are mostly around **reliability, context, and reach**, not more intelligence.
## highest-value additions
### 1. unified event intake
maven should receive normalized events from:
* praxis
* calendar
* telegram
* local notifications
* system/service health
* manual checklists
* eventually email bridges
one internal envelope:
```go
type Event struct {
Source string
Kind string
EntityIDs []string
Title string
Body string
Priority string
OccurredAt time.Time
Payload json.RawMessage
}
```
this gives digestion one stable input instead of source-specific logic.
---
### 2. explicit morning routine engine — **core engine done (2026-07-20)**
`internal/morning` — pure checklist engine, mirrors `internal/loop`/
`internal/routine`'s no-I/O contract. `Evaluate(routine, facts, now)` answers
"what's still missing" any time (order-independent — checks facts, not
sequence); `Due(routines, facts, last, now)` fires the once-per-day nag only
at `NudgeAt` (defaults to window end) and only when something's unevidenced,
with a `last`-map dedupe identical in shape to `routine.Due`'s cold-start/
last-fire tracking. Evidence is just a fact timestamped inside today's
window — manual (voice-tapped) and inferred (another daemon writing the same
key) are indistinguishable, satisfying the manual/inferred requirement for
free. Weekday/weekend variants are two `Routine`s with different `Weekdays`
sets under different names. Wired into `config.MorningRoutineConfig` +
`cmd/mavend/tick.go`'s `fireMorningRoutines` (reads only the fact keys the
configured items reference, dispatches through the normal severity/presence
routing table, body is literal joined item labels — not LLM-phrased, same
no-hallucination rationale as cron routines). 13 unit tests in
`internal/morning/morning_test.go`.
Added since (2026-07-20, same day): a read-only `/morning` page in mavweb —
`ipc.CoreAPI.MorningStatus` (new wire method, mirrors `TickTrace`'s
daemon-cache-only shape: the store adapter errors, `daemonAPI` serves it from
a `tickLoop.morningStatus` closure) returns each routine's active/window/
per-item done state, server-rendered same as `/trace` (no live-update loop —
checklist state moves on minutes, not seconds).
Not yet done: no config wired in `deploy/mavend.json` (no morning routines
configured on homesrv yet — add items there when the medicine/water/pets
fact keys the phone/desktop write are settled), no voice query path for
"what did I miss this morning" (Evaluate supports it; nothing calls it yet),
no way to create/edit routines from the web UI — construction still means
hand-editing config, deliberately deferred: routines are operator-declared
config (like cron routines), and a CRUD editor would mean moving them to a
DB table + hot-reload, a bigger change than this pass.
not ordinary reminders.
support:
* required morning items
* order-independent completion
* soft time windows
* skipped-step detection
* one nudge, not repeated spam
* manual and inferred completion evidence
* weekend/weekday variants
example:
```text
08:0011:00
- medicine
- water
- pets
- check praxis attention
```
maven should know what is still missing, not merely fire four timers.
---
### 3. cross-device presence
**status (2026-07-20):** the hysteresis engine and 3 of the listed signals are
already built and wired live: `internal/store/presence.go` (noisy-OR combiner
+ Schmitt-trigger bucket resolve), fed by `desk_active` (workstation, via
`scripts/desk-active.sh` posting to `/api/signal`), `page_heartbeat` (mavweb
tab, `app.js`), and `wg_handshake` (`mavpoll` polling `wg show`) — threaded
into the tick loop via `internal/loop/gather.go`. Not done: phone-reachable,
homesrv-available, audio-output, and active-maven-client signals from the
list below are still missing.
a small presence daemon on each trusted device:
* workstation active/idle
* phone reachable
* homesrv available
* last keyboard/mouse activity
* wireguard presence
* current audio output
* active maven client
mavend receives only compact state, not raw activity logs.
useful for:
* choosing delivery channel
* suppressing voice while away
* surfacing reminders when you return
* knowing whether an agent result should be spoken or sent as text
---
### 4. interruption policy — **done (2026-07-20), turned out to already be built**
audited the existing code before writing anything new: `internal/loop.Gate`
already answers deliver_now vs. drop (quiet-hours/cooldown/snooze/presence/
calendar-busy), and `cmd/mavend/tick.go`'s `digestQ` + `config.DigestConfig`
already implement queue/digest (low-severity nudges batch into one
notification, flushed on window elapsed or max-items reached). The four
outcomes below were already covered by these two mechanisms; nothing new to
build for the core policy.
Gap that *was* real: `deploy/mavend.json` had no `digest` block, so batching
was disabled in prod despite being fully implemented. Fixed — see the config
change alongside this note.
before delivering anything, evaluate:
```text
urgency
current activity
quiet hours
recent nudges
available channels
whether already surfaced
```
result:
```text
deliver_now
queue
digest
drop
```
this prevents maven from becoming annoying once praxis and other sources start producing more data.
---
### 5. entity-aware memory — **done (2026-07-20)**
`03fa52d`/`9876187` (Vikunja #279): facts gain `Subject`/`EntityID`/
`ResolutionState`; an async enrichment worker resolves free-text subjects to
canonical Nexus entity_ids (mirrors Praxis's enrichment pattern). Ambiguous
or unreachable Nexus never guesses — the fact stays `pending` or terminal
`ambiguous`. Voice-tapped facts (`IntentFact`) now flow into the enrichment
queue automatically via an optional `Subject` field on `WriteFactReq` (old
callers unaffected).
Landed alongside this in the same session (not originally on this list, but
closes the plumbing gaps the last brief flagged for Nexus/Praxis maturity):
a typed Praxis lifecycle client (`398997f` — surface/acknowledge/resolve/
ignore/pin; fixes the surfaced≠acknowledged gap where reading an item aloud
left no trace), correlation-ID/version headers on the Nexus/Praxis clients
(`b743860`), entity-scoped Praxis attention queries (`0579ef9`), a durable
delivery outbox with begin-before-send/complete-after semantics
(`29f23e3`+`9ff726e` — closes a duplicate-send-on-crash bug), fail-closed
handling on ambiguous IPC mutation outcomes and Nexus/Hexis dependency
errors (`838fde1`+`d9fa4d6`), and a reusable fake-ecosystem test harness
with fault injection (`c932cd8`).
connect maven memory to nexus ids.
instead of:
```text
key = "кошачий фонтан"
```
store:
```text
entity_id = ent_pet_water_fountain
predicate = refilled_at
value = 2026-07-19T...
```
benefits:
* stable russian/english aliases
* fewer duplicate facts
* better “when did i last…” queries
* easier routine detection
* cleaner praxis correlation
---
### 6. bounded follow-up state
for short continuations:
* “yes”
* “tomorrow”
* “the second one”
* “not that project”
* “do it later”
store explicit pending state instead of relying on chat history:
```go
type PendingInteraction struct {
Kind string
Candidates []string
Args json.RawMessage
ExpiresAt time.Time
}
```
this matters a lot for a 1.7b model.
---
### 7. evaluation lab — **skipped for now (2026-07-20)**
runs on a different machine (GPU box), and CPT is currently in progress
there — deprioritized until the training pipeline has a checkpoint to gate.
Not abandoned, just off the immediate list.
before every new checkpoint or lora deploy:
* routing accuracy
* slot accuracy
* malformed json rate
* russian/english mixed input
* ambiguous entity handling
* reminder vs note vs fact
* direct answer vs tool call
* confirmation safety
* phrasing quality
* latency and ram
also replay real anonymized traces against old and new checkpoints.
this should be a hard deployment gate.
---
### 8. replayable full-system simulator
fake:
* clock
* presence
* caldav
* telegram
* praxis
* nexus
* hexis
* stt
* tts
* llama-server
scenario:
```text
08:30 user appears
08:35 medicine not completed
08:40 correx agent waits
08:45 calendar sync stale
08:50 user says “what did i miss?”
```
assert:
* what tools were called
* what was surfaced
* what stayed unresolved
* what maven said
* what was not executed
this will save more time than another feature daemon.
---
## useful second-wave additions
### voice session quality
* barge-in
* interrupt tts on wake word
* partial stt display
* confidence-aware clarification
* retry only failed stt segment
* per-room microphone profiles
* noise-floor calibration
* short response mode when speaking
### notification bridge framework
small adapters for:
* ntfy
* telegram
* matrix
* web push
* android notification forwarding
* local dbus notifications
normalize into maven/praxis events instead of treating each as a separate feature.
### local knowledge ingestion
* markdown/docs ingestion
* git repo summaries
* project decision records
* conversation exports
* provenance and source links
* incremental reindexing
keep this read-only and separate from personal fact memory.
### service self-diagnostics
`maven doctor`:
* socket reachability
* model health
* stt/tts readiness
* embedder availability
* caldav freshness
* telegram poll state
* praxis/nexus/hexis reachability
* db integrity
* disk usage
* recent failures
### config and secret management
* schema-validated config
* config migration
* secret references instead of inline values
* dry-run validation
* redacted config dump
* per-daemon health config
* startup dependency report
---
## things i would not build yet
* autonomous multi-step planning
* large external reasoner
* generic workflow engine
* self-editing memory
* automatic hexis actions from praxis
* emotion simulation beyond phrasing
* full home-assistant replacement
* more model layers before routing is stable
## recommended order
**status as of 2026-07-20:**
1. ~~evaluation lab~~**skipped, GPU-box work, deprioritized while CPT is in progress**
2. ~~entity-aware memory~~**done** (`03fa52d`/`9876187`, plus adjacent
Nexus/Praxis plumbing hardening — see item 5 above)
3. ~~morning routine engine~~**core engine done** (`internal/morning` +
`cmd/mavend` wiring — see item 2 above; not yet configured on homesrv,
no voice query, no web UI)
4. interruption/delivery policy
5. presence agents
6. unified event intake
7. full-system simulator
8. notification bridges
9. knowledge ingestion
10. voice-session polish
the main goal should be: **maven reliably knows what is happening, knows what you meant, and chooses the least annoying correct response**. everything else can wait.
+1 -1
View File
@@ -18,7 +18,7 @@ live in sibling repos next to this one.
Division of labour: Nexus identifies, Praxis observes, Hexis acts, Maven understands Division of labour: Nexus identifies, Praxis observes, Hexis acts, Maven understands
and coordinates. Maven is not the source of truth for any of the three. The full and coordinates. Maven is not the source of truth for any of the three. The full
contract is `MAVEN_ECOSYSTEM_ARCHITECTURE.md`, and the constraints that bite during contract is `docs/ecosystem.md`, and the constraints that bite during
implementation are summarised in `CLAUDE.md`. implementation are summarised in `CLAUDE.md`.
Where things are in this repo: Where things are in this repo:
+6 -6
View File
@@ -10,7 +10,7 @@ compose passes `/dev/dri` + the render gid) — the resident model stays ≤1.7B
**Resident model:** currently **Qwen3-1.7B** (`UD-Q4_K_XL`), stock — not yet the CPT'd one. **Resident model:** currently **Qwen3-1.7B** (`UD-Q4_K_XL`), stock — not yet the CPT'd one.
It replaced Qwen3.5-0.8B on 2026-07-31 because it measured better on both fixtures we have: It replaced Qwen3.5-0.8B on 2026-07-31 because it measured better on both fixtures we have:
67.5% vs 59.7% intent-only on the 77-case RU routing fixture, and 20/27 vs 11-17/27 on the 67.5% vs 59.7% intent-only on the 77-case RU routing fixture, and 20/27 vs 11-17/27 on the
talk fixture. See `MODEL-BAKEOFF-31-07-2026.md`. It is a Thinking variant, so `n_ctx` is 4096 talk fixture. See `docs/evals/2026-07-31-model-bakeoff.md`. It is a Thinking variant, so `n_ctx` is 4096
— reasoning tokens need the room, and 4096 is what the scores above were measured at. — reasoning tokens need the room, and 4096 is what the scores above were measured at.
The **target** is still the locally CPT'd **Qwen3-1.7B** (Vikunja #122, training in flight). The **target** is still the locally CPT'd **Qwen3-1.7B** (Vikunja #122, training in flight).
@@ -25,7 +25,7 @@ Spanish. Their strong published IFEval/BFCL numbers are English-only. Model file
`models/llm/`, so the LFM2.5 gguf sitting there is not loaded by anything. Swapping the resident `models/llm/`, so the LFM2.5 gguf sitting there is not loaded by anything. Swapping the resident
model is a one-line change to `phraser.model_path` in `deploy/mavend.json`. model is a one-line change to `phraser.model_path` in `deploy/mavend.json`.
See `REARCH.md` for the target architecture, `DESIGN.md` for the folded design spec, and See `docs/rearchitecture.md` for the target architecture, `docs/design.md` for the folded design spec, and
`AGENTS.md` for local-preview + model-download recipes. `AGENTS.md` for local-preview + model-download recipes.
## Build & test ## Build & test
@@ -72,7 +72,7 @@ protocol; the config in `deploy/mavend.json` (with `${VAR}` env expansion from g
Maven is one of four services. It owns conversation and personal memory. It does not Maven is one of four services. It owns conversation and personal memory. It does not
own identity, operational state, or execution. Full contract in own identity, operational state, or execution. Full contract in
`MAVEN_ECOSYSTEM_ARCHITECTURE.md`. `docs/ecosystem.md`.
```text ```text
Nexus identifies. Praxis observes. Hexis acts. Maven understands and coordinates. Nexus identifies. Praxis observes. Hexis acts. Maven understands and coordinates.
@@ -113,7 +113,7 @@ Every cross-service call carries a correlation id minted once per action
on in deploy** — this section used to say it was wired `nil`, which stopped being true on on in deploy** — this section used to say it was wired `nil`, which stopped being true on
2026-07-31. 2026-07-31.
- **LLM router (the intended design, REARCH.md):** the resident Qwen3-1.7B (`llmrouter.go`) - **LLM router (the intended design, docs/rearchitecture.md):** the resident Qwen3-1.7B (`llmrouter.go`)
emits GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is emits GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is
demoted from a routing gate to a RAG hint. Wired at `voice.go:214` via demoted from a routing gate to a RAG hint. Wired at `voice.go:214` via
`pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient)`; the flag is `voice.llm_router` `pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient)`; the flag is `voice.llm_router`
@@ -127,11 +127,11 @@ on in deploy** — this section used to say it was wired `nil`, which stopped be
Cascade order: `stage0.go` exact-match fast-path → LLM router (when non-nil) → classifier Cascade order: `stage0.go` exact-match fast-path → LLM router (when non-nil) → classifier
fallback. Any LLM error falls through to the classifier so a turn never breaks on the model. fallback. Any LLM error falls through to the classifier so a turn never breaks on the model.
Measured on the 77-case RU fixture (`MODEL-BAKEOFF-31-07-2026.md`): the classifier scores Measured on the 77-case RU fixture (`docs/evals/2026-07-31-model-bakeoff.md`): the classifier scores
36.8% full accuracy at p50 31ms; Qwen3-1.7B scores 67.5% intent-only / 72.7% through the 36.8% full accuracy at p50 31ms; Qwen3-1.7B scores 67.5% intent-only / 72.7% through the
cascade at p50 ≈825ms. Accuracy roughly doubled, latency is ~27× worse, and that trade was cascade at p50 ≈825ms. Accuracy roughly doubled, latency is ~27× worse, and that trade was
accepted deliberately. **The ≈2.7s figure that stood here until 2026-08-02 was contention, accepted deliberately. **The ≈2.7s figure that stood here until 2026-08-02 was contention,
not the model.** See `ROUTING-EVAL-31-07-2026.md` line 61, which measures the LLM router at not the model.** See `docs/evals/2026-07-31-routing.md` line 61, which measures the LLM router at
p50 825ms / p95 1.2s / max 3.0s and the full cascade at p50 0.80-1.04s. Do not plan latency p50 825ms / p95 1.2s / max 3.0s and the full cascade at p50 0.80-1.04s. Do not plan latency
work off the bakeoff table. `Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM work off the bakeoff table. `Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
+2 -2
View File
@@ -78,7 +78,7 @@ deps-go:
done done
$(GO) version $(GO) version
# fmt-check fails if any file needs gofmt. DESIGN.md has always said `make # fmt-check fails if any file needs gofmt. docs/design.md has always said `make
# test` gates on gofmt and vet; it did not, so nine files quietly drifted. # test` gates on gofmt and vet; it did not, so nine files quietly drifted.
# Run `gofmt -w` on whatever this prints. # Run `gofmt -w` on whatever this prints.
fmt-check: fmt-check:
@@ -191,7 +191,7 @@ deps-piper:
# multilingual-e5-small: an asymmetric retrieval model. It is trained to match # multilingual-e5-small: an asymmetric retrieval model. It is trained to match
# a short question against a longer passage, which is what note recall is. # a short question against a longer passage, which is what note recall is.
# The quantized file is the one we download, deploy and measure — see # The quantized file is the one we download, deploy and measure — see
# RECALL-EVAL-31-07-2026.md. # docs/evals/2026-07-31-recall.md.
EMBEDDER_DIR := $(shell pwd)/models/embedder/multilingual-e5-small EMBEDDER_DIR := $(shell pwd)/models/embedder/multilingual-e5-small
EMBEDDER_MODEL_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/onnx/model_quantized.onnx EMBEDDER_MODEL_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/onnx/model_quantized.onnx
EMBEDDER_TOKENIZER_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/tokenizer.json EMBEDDER_TOKENIZER_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/tokenizer.json
-468
View File
@@ -1,468 +0,0 @@
## Maven — current state (updated 2026-07-20)
### Session 2026-07-20 — ecosystem hardening + entity-aware facts
Ten commits, focused on closing the Nexus/Praxis integration gaps flagged
as "wired but immature" in the prior review, plus the entity-aware-memory
backlog item (`20-07-2026-BACKLOG.md` item 5).
- **Entity-aware fact resolution (Vikunja #279)** — facts gain
`Subject`/`EntityID`/`ResolutionState`; an async worker resolves
free-text subjects to canonical Nexus entity_ids (mirrors Praxis's own
enrichment pattern). Ambiguous/unreachable Nexus never guesses — stays
`pending` or terminal `ambiguous`. Voice-tapped facts (`IntentFact`) flow
into the queue automatically via an optional `Subject` field on
`WriteFactReq` (old callers unaffected, no signature break).
- **Typed Praxis lifecycle client (Vikunja #271)** — `GetItem`/`Search`/
`Surface`/`Acknowledge`/`Resolve`/`Ignore`/`Pin`, routed through new RU/EN
dialogue verbs. Fixes a real lifecycle-invariant bug: reading an
attention item aloud now calls `Surface` — previously the digest path
read items without recording that they'd been surfaced, so "Maven
mentioned it" was indistinguishable from "never came up."
- **Durable delivery outbox (Vikunja #270)** — `BeginDeliveryAttempt`
before `Send`, `CompleteDeliveryAttempt` after; a stale `pending` row
found at startup reconciles to `unknown` (never silently resent or
dropped — same rule as Hexis's execution-timeout handling). Closes a
crash-window duplicate-send bug. Wired into `DispatchNudge`,
`DispatchReminder`, `RepeatUnacked`; reconciliation runs once at boot
before the tick loop resumes.
- **Fail-closed IPC/dependency handling (Vikunja #269, #272/#273)** —
ambiguous mutation outcomes (frame sent, reply lost) no longer blindly
retry; Nexus/Hexis dependency errors fail closed instead of guessing.
- **Correlation IDs + version headers (Vikunja #273)** — the hand-rolled
Nexus/Praxis HTTP clients now send `X-Nexus-Version`/`X-Praxis-Version`
and thread the same correlation ID already generated in
`executeCapability` through the whole call chain, matching the Hexis
client's existing behavior.
- **Entity-scoped Praxis attention queries** — callers holding a resolved
entity_id can ask "what needs attention for this entity" directly
instead of filtering the unscoped list client-side.
- **Fake-ecosystem test harness with fault injection** — a reusable
`fakeServer` (Nexus/Praxis/Hexis fixtures, runtime-toggleable
`SetFault`, fake clock) replacing ad-hoc per-test `httptest` servers;
covers a gap that had zero test coverage (`handlePraxisAct`) and adds a
fault-then-recovery regression test for the fail-closed fixes above.
- **Ops fix** — `deploy/mavend.json`'s phraser was pointed at a 4B model
with `n_gpu_layers=99`, which OOM'd under memory pressure and left a
zombie `llama-server` child; swapped to the 2B Qwen model matching the
intended resident-model size.
Net effect: the Nexus/Praxis wiring described as "plumbing exists, thin
compared to Maven's test depth" in the prior review is now materially
hardened — typed clients, fail-closed error handling, durable delivery,
and a proper fault-injection test harness are all in place. Evaluation lab
(`20-07-2026-BACKLOG.md` item 7) is explicitly skipped for now — it runs
on the GPU box, which is occupied by CPT. Morning routine engine (backlog
item 3) is next up, not started.
---
> **Resolved 2026-07-30 (task #318).** The resident checkpoint is
> **Qwen3.5-0.8B** (`Q4_K_M`), set in `deploy/mavend.json`; the **target** is
> the locally CPT'd **Qwen3-1.7B**, still training (#122). Older model claims
> below — the LFM references, the pipeline line, and the "swapped to the 2B
> Qwen model" ops entry above — are historical. Read them as a log of what was
> true at the time, not as current fact. Note also that `/mnt/hdd1/llms` is
> bind-mounted over `models/llm/`, so the LFM2.5 gguf in the repo tree is
> never loaded.
Architecture decision (as written on 2026-07-20): the target resident
router/phraser is the locally trained Qwen3-1.7B model — still the target as
of 2026-07-30. Older LFM references below describe the then-deployed
historical stack, not the target checkpoint. RU CPT has a successful
full-weight checkpoint at step 1000/8077; evaluation and Qwen3 SFT tooling are
tracked in `docs/plans/2026-07-18-qwen3-resident-training-eval.md`.
Consolidated status. The reactive↔proactive core is closed and testable through
the web PWA. The former SPEC's open items 17 (now `DESIGN.md` § execution ledger) are landed (protocol doc, away-channel
fallthrough, CalDAV poller, quiet-hours schedule, tools enable/disable, note RAG,
passkey step-up); item 8 (multi-user) is deliberately deferred — see the tail.
The two big infra gaps from the jul5 revision are closed on `overnight-jul5`:
**at-rest encryption** (AES-256-GCM, tmpfs working copy — not sqlcipher, see
`internal/store/crypt.go`) and **Docker deployment** (one image, six daemon
containers). The `overnight-jul6` session (now on `master`) closed the biggest
*query-surface* gaps — **calendar querying, general-knowledge answers, and
weather** — plus a populated homelab act allowlist and two pure scaffolds
(dialogue state, long-term-memory vector store). ~15.2k LOC + ~8.5k test, 303
tests, `-race` in `make test`.
### Access model
- **Phone** → needs the wg tunnel to reach homesrv (no homesrv DNS otherwise;
raw IP or a DNS tweak can bypass, not the default).
- **PC** → uses homesrv DNS, resolves the domains over local-net, **no wg needed**.
- nginx + ufw both scope to `10.42.0.0/24` (wg) + `192.168.1.0/24` (LAN), deny all else.
- **Surface in use now: the web PWA (`mavweb`).** Voice PTT + in-app nudges both ride it.
### Works end-to-end (tested)
- **Reactive voice:** PWA record → Whisper STT (`mavsttd`) → ONNX classifier →
resident phraser (llama-server subprocess; Qwen3.5-0.8B as of 2026-07-30 —
this line historically named "LFM 2.5-1.2B") → Piper TTS
(`mavttsd`) → reply.
HTTP POST path (mobile-Chrome drops WS for the audio).
- **Capture:** `fact` (EN **and RU** — root-substring recognizers) + `reminder`
persist through CoreAPI (`source=tap:voice`). This is the substrate the care
rules read.
- **Notes / query (semantic recall, sqlite — no chroma):** `note` → embed (the
classifier's ONNX embedder) → `notes` table. `query` → embed → brute-force
cosine top-k → confidence-gated (below `queryMinScore` 0.55 ⇒ "no note", not a
guess). **Note RAG (SPEC item 6):** the gated top-k feed the phraser
(`PhraseQuery`) to compose a natural answer ("вот что я нашла: …") instead of
a verbatim dump; raw-notes fallback on any LLM error. Stub is deterministic.
- **Monitoring (`/dash`):** mavweb server-renders presence + recent nudges (by
outcome) + recent facts from the append-only store via CoreAPI. Read-only,
meta-refresh, no JS.
- **Proactive loop:** 60s dumb ticker, pure predicates over a State snapshot,
universal gate (quiet-hours/presence/cooldown/snooze/calendar), one-nudge-per-
tick max-severity, reminders (gate-bypassing), sev4 repeat-til-ack, feedback
auto-tuner (outcome ratio → bounded cooldown, persisted as `source=feedback`).
- **Rules:** water/meal/break (sev12 care), service_down (sev4, `poll:uptimekuma`),
netdata_critical (sev3, `poll:netdata`).
- **Routines (`internal/routine`):** operator-declared clockwork — the third
proactive class beside reminders (user-stated) and care rules (world-state).
Config `routines[]` (cron + literal RU body + severity) fire through the normal
dispatcher on schedule (an 08:00 briefing, a 22:00 wind-down). Bodies are
literal (not LLM-phrased ⇒ can't hallucinate); rule name `routine:<name>` so
they don't pollute the care autotuner; cold-start guard seeds on first sight so
a restart never replays a missed schedule. Pure `routine.Due`, unit-tested; the
tick driver holds the last-fired map.
- **Env facts (`mavpoll`):** netdata alarms → `netdata_alarm` (fires immediately
on a real CRITICAL); kuma monitor_status → `service_down`. Writes only on
value-change (no append-only churn).
- **Presence:** noisy-OR decay + Schmitt hysteresis. Live via `page_heartbeat`
(PWA auto-pings `/api/signal` every 30s → present when a tab's open).
- **Delivery:** ntfy / telegram / voice by `f(severity, presence)`; minimal body
on away channels. PWA subscribes to ntfy over **WebSocket** for in-app nudges.
- **Away-channel fallthrough (SPEC item 2):** when the router picks voice but no
live session exists at push time (presence guess was wrong), the dispatcher
reroutes through the AWAY table — sev3→ntfy, sev4→telegram-repeat-til-ack,
sev≤2→drop — instead of silently dropping. Covers nudges + reminders.
- **Calendar busy (SPEC item 3, `mavcaldav`):** new poller queries a self-hosted
**Radicale** CalDAV server on an interval, writes `calendar_busy` + event facts
through CoreAPI (value-change only). The loop gate already consumes `calendar_busy`.
- **Quiet-hours schedule (SPEC item 4):** the gate reads `quiet_hours`; a config
time window (`voice.quiet_hours`, HH:MM, midnight-crossing handled) now sets it
on each tick — in addition to the "тихий режим" voice toggle. Both activate quiet.
- **Client protocol (SPEC item 1):** the voice wire format (length-prefixed JSON
frames) is published in `PROTOCOL.md`, generated from `internal/voice/wire.go`
so third-party clients don't need the Go source.
- **Passkey step-up (SPEC item 7):** `internal/webauthn` does real WebAuthn —
ES256/P-256 register + assert, ecdsa signature verification, rpIdHash + UP/UV
flag binding (UV = the gesture), sign-count regression check. `PasskeySession`
bumps the auth session L2→L3 for a TTL on assert. mavweb serves `/auth/passkey`
(enroll + step-up) + the begin/finish endpoints. Crypto is round-trip tested
(incl. tampered-sig / missing-UV / wrong-origin negatives).
- **Stability:** llama-server orphan leak fixed (`Pdeathsig` kills the child on
any mavend death); `kill-maven.sh` reaps strays (matches the model, not a
bogus `llama-server.*maven` pattern); `start-maven.sh` wires `-core` + poller.
### Wired but needs a deploy action (not code)
- **`desk_active`** (strongest presence signal) — `scripts/desk-active.sh` runs
on the **desk PC** (hypridle-gated systemd timer), posts over wg to mavweb.
- **`mavwaked`** (always-on listening) — needs a systemd user unit on a client
box (desk PC, pi, etc.) where the mic is attached. Connects to mavend over wg
or local net via `-addr`. Deferred until a client box is wired with a mic.
Caveats / gotchas:
- **desk_active is a workstation deploy, not code** — 0 facts ever written; presence
runs on page_heartbeat alone (dash reads "away"/"never at desk"). `scripts/desk-active.sh`
+ a hypridle-gated `maven-desk` timer must be installed on the desk PC (not homesrv).
- **Notes recall needs the ONNX embedder** — under the HashEmbedder floor, cosine is
lexical (token overlap), not semantic; scores are low, so most RU commands sit under
the 0.35 route threshold and clarify. Configure `voice.embedder` for confident recall+routing.
(The floor now at least tokenizes Cyrillic — see below — so it ranks correctly, just weakly.)
- **Switching the embedder model silently breaks old notes** — different dim ⇒
cosine 0 ⇒ they stop matching; brute-force can't re-embed. Re-embed on a model change.
- **`wg_handshake` is OFF and should stay off** — in this topology the phone only
runs wg when *outside*, so a fresh handshake means AWAY, not here. The `mavpoll
-wg` flag exists (defaults `""`) and could later back the spec's "away override"
by flipping the sign; as a presence-*here* signal it's inverted. desk_active +
page_heartbeat cover home presence.
- **Cold-start unlock tests are missing** — the key wrap/unwrap code
(`internal/webauthn/keywrap.go`) and locked-mode IPC gating (`cmd/mavend/main.go`)
are correct but have **zero test coverage**. The roadmap (item 2.1) required
three new test cases (wrap/unwrap round-trip, wrong-cred unwrap fails,
locked-mode IPC rejects non-unlock methods); none were written. `make test`
is green by omission. Write these before relying on the cold-start path with
real keys.
### Done since last revision (overnight-jul6, 2026-07-06)
Seven tasks (session board `SESSION-06-07-2026.md`, deleted 2026-07-30 — see git history), one commit each, merged to `master`.
This session was run through **opencode**, not Claude Code (co-author trailer).
Since then (**2026-07-06, second session**):
- **Always-on listening (gap 1, MVP)** — `cmd/mavwaked/`: 825 lines, 10 `-race`
tests. Energy-based VAD over 30ms windows (same RMS threshold as mavsttd's
`gateReason`), adaptive noise floor, speech→silence state machine. Captures
PCM from arecord(1) subprocess, sends `PushToTalk` with `Surface=SurfaceVoice`
(L0 — no destructive acts). Reply plays through aplay(1). No wake word yet
(pure VAD trigger); the 30ms frame shape matches silero-vad ONNX input 1:1,
so swapping energy-threshold for ONNX inference is a local change in vad.go.
`Makefile` `build-waked` target. Runs on client boxes (not docker/homesrv)
via systemd user unit; connects to mavend over wg or local net.
Since then (**2026-07-06, third session** — roadmap execution agent):
- **Cold-start unlock (ROADMAP 2.1)** — the at-rest AES key is now wrapped
(HKDF-SHA256 + AES-256-GCM, stdlib-only — no `x/crypto` dep) with the passkey
credential's public key and persisted to disk. At boot, if a wrapped key file
exists AND no env key is set, mavend starts **locked**: the IPC server runs
but `srv.Check` rejects everything except `MethodAssertStepUp` +
`MethodUnlock`. A passkey assertion at `/auth/passkey` calls `MethodUnlock`
with the credential's public key → unwraps the blob → opens the store → wires
voice/loop/delivery → `srv.SetAPI` swaps the locked stub for the real
CoreAPI. mavweb's `RegisterFinish` wraps the env key on enrollment;
`AssertFinish` calls `Unlock` on assertion. Env-key fallback preserved
(dev/CI path unchanged). **Test gap:** the roadmap required three new test
cases (wrap/unwrap round-trip, wrong-cred unwrap fails, locked-mode IPC
rejects non-unlock methods) — none were written. The code is correct but
untested; `make test` is green by omission, not coverage.
- **Conversation depth (ROADMAP 3.2)** — cross-intent anaphora + fact-by-key
lookup. `AnaphoraResolver` in `router/slots.go` detects RU pronouns
(это/он/она/оно/тот/мой + inflected forms). `followUpMerge` now handles
three cases: same-intent slot inheritance (existing), cross-intent anaphora
(Query/Fact/Reminder after a Fact with a pronoun inherits the prior key +
time), and query-after-fact (a query following a fact inherits the key for
fact-by-key lookup). `Session.History []Turn` added as the multi-turn
scaffold (capped at 4). 7 new test cases including the exact done-when
scenarios (anaphora query-after-fact, three-turn break, explicit-key-wins).
- **Routing quality + persona (ROADMAP 4.1/4.4)** — `QueryMinScore` is now a
config knob (`voice.query_min_score`, default 0.55) instead of a hardcoded
const. `make download-embedder` fetches Xenova/paraphrase-multilingual-
MiniLM-L12-v2 (~90MB ONNX) + tokenizer; AGENTS.md documents the embedder +
libonnxruntime setup. `Persona` field in `VoiceConfig` prepends to every
LLM system prompt (nudge phrasing, note queries, general knowledge); empty
= current hardcoded feminine-gendered Russian persona. Also fixed two
pre-existing data races found by `-race`: `voice/server.go` wg.Add vs
wg.Wait (accept mutex), `mavweb/server.go` s.api field (atomic.Value).
- **Calendar querying (task 3)** — "что у меня завтра?" now answers from the
CalDAV facts the poller already writes. Added `store.CalendarEvents(from,to)`,
a RU date-scope parser («сегодня»/«завтра») in `router/slots.go`, and an
IPC `CalendarEvents` RPC (api/client/server/wire) feeding the `IntentQuery`
handler. Empty day → «на сегодня ничего нет». Previously calendar only *gated*
nudges; it's now queryable.
- **General-knowledge routing (task 4)** — when notes-RAG misses `queryMinScore`,
the query now falls through to the phraser with an anti-hallucination system
prompt (`router.KnowledgePrompt`, single tested source) instead of giving up.
Empty/errored/Stub phraser → «не знаю.», never a fabrication.
- **Weather (task 5)** — new `internal/weather/`: `Provider` interface, a stub
(«погода не настроена»), and a real **keyless Open-Meteo** provider (geocode +
current_weather, injectable `*http.Client`, mocked in tests — no live network).
Wired into `IntentQuery` (keywords погода/градус/температура) with a ~5s
context timeout; selected by `voice.weather.provider` ("open-meteo" | "" → stub).
- **Homelab act allowlist (task 2)** — `voice.tools` seeded with read-only acts
(`systemctl status`, `docker ps`, `uptime`, `df`, `free`, `journalctl` reads)
as `destructive:false` and mutating ones (restart/stop/start/reboot,
docker-restart/stop) as `destructive:true`. Guardrail verified: no dangerous
verb is `destructive:false`. RU phrasings seeded in `act.txt`.
- **Embedder config validation (task 1)** — a partially-filled `voice.embedder`
block (some of model/tokenizer/lib paths missing) is now a load error instead
of a silent fall-through to the Hash floor; the floor fallback logs explicitly.
- **Dialogue state scaffold (task 6)** — `internal/dialogue/`: `Session` +
TTL `SessionStore` + pure `InheritSlots`. **Now wired** (post-merge follow-up):
the voice handler carries slots across same-intent turns within a 2-min window
(`followUpMerge`, unit-tested) — bounded gap-filling, not full multi-turn yet.
- **Long-term memory interface (task 7)** — `internal/memory/`: `Store` interface
+ `InMemoryStore` (cosine). Wired into `IntentNote` (best-effort insert) and,
post-merge, into `IntentFact` (facts indexed) + `IntentQuery` (read-back after
notes-RAG misses). In-memory only — no persistent backend yet (gap #8).
Follow-ups (Claude Code, post-merge): gofmt'd `handlers_test.go` (the jul6
verification commit left it misaligned, so `gofmt -l` still flagged it despite the
"all gates green" claim); deduped the task-4 knowledge prompt to the single tested
`router.KnowledgePrompt()`. Tree is now genuinely green (gofmt/vet/303 tests).
### Done since the jul5 revision (overnight-jul5, 2026-07-05)
The overnight session (`SESSION-05-07-2026.md`, deleted 2026-07-30 — see git history; 25 tasks) closed the previous
"not built yet" items 13 and added feature depth:
- **At-rest encryption** — the on-disk db is AES-256-GCM ciphertext; the daemon
works on a tmpfs (RAM) plaintext copy, sealed back atomically on close. Wrong
key / tamper ⇒ fail closed, never a plaintext fallback. Legacy plaintext dbs
upgrade on first clean shutdown. Key via config/env (`db_key_env`); no KDF —
raw 32-byte key, base64. The passkey cold-start unlock plugs into the same
`store.OpenEncrypted` seam later.
- **Docker deployment** — single image, one container per daemon
(`docker-compose.yml`); only mavend mounts the key + db volume; IPC over a
shared socket volume. `ipc.DialWait` (boot-order tolerance) + redial-on-drop
(core restarts don't kill modules). `deploy/README.md` has the runbook.
- **Tests** — mavcaldav, mavttsd, voicesink, mavweb main/handlers covered;
`make test` runs `-race -coverprofile`.
- **Recurring reminders** — `cron` + `next_fire_ts` on reminders; recurring ones
reschedule (instead of mark-fired) after successful delivery.
- **Notification digest/batching** — low-severity nudges queue and flush as one
digest per window/max-items (`digest` config block); stale-reminder bursts on
boot collapse into a single digest reminder, completed only after delivery.
- **Rule trace engine** — `ExplainTick`/`ExplainGate` record per-rule
predicate/gate/selection results each tick; served over IPC (`tick_trace`)
and rendered at mavweb `/trace` ("why didn't she nudge me").
- **Web UI** — new `/history` (facts + revert buttons), `/notifications` (nudge
history), `/trace` pages; nav links on `/dash`; RU/EN cheatsheet toggle in the
PWA; manifest icons (`icon.svg`). POST `/tools` now requires an in-process
passkey step-up when WebAuthn is configured.
- **Revert/undo** — `RevertFact` voids the latest fact for a key (append-only
void-marker, audit trail intact); exposed at `/api/revert` from `/history`.
- **Tool scopes** — `scope` column on tools, threaded through propose/enable/UI.
`DisableTool` raised to AuthStepUp alongside Enable.
- **Passkey persistence** — mavweb credentials in a JSON file (`-passkey-file`),
surviving restarts; rollback-on-persist-failure keeps memory and disk in sync.
- **STT silence gate** — min-duration + RMS floor drop non-speech before whisper
hallucinates on it (`-min-ms`, `-silence-rms` flags on mavsttd).
- **Housekeeping** — `db_key.env` gitignored (+`.env.example`), `build-caldav`
target, zero-timestamp "never" fix on /dash.
### Not built yet (ranked by ROI)
1. **Multi-user (SPEC item 8)** — deliberately deferred, see the tail.
Closed (jul6 follow-ups): `/api/revert` now sits behind the same passkey
step-up as POST `/tools`; `go.mod` direct deps (`onnxruntime_go`,
`coder/websocket`, `robfig/cron`) are labeled correctly — `go mod tidy` can't
run here because it walks the vendored `deps/go` toolchain tree.
Purge+rotate leaked db key (#12) — investigated and closed: the key was
**never committed** to git history (gitignored at introduction, no commit
ever tracked `deploy/db_key.env`), so nothing to scrub. File stays on disk
and in deploy env by design — at-rest encryption needs it at boot.
Done earlier (2026-07-03): **act tool executor, store-backed, full flow**
(`internal/tool` + `internal/store/tools.go` + `tools` CoreAPI methods).
- **Execution:** IntentAct runs the matched fn against the store's ENABLED
allowlist. argv, no shell → STT text can't inject. Live store read, so a
newly-enabled tool runs without a daemon restart.
- **proposed→enabled→disabled (SPEC item 5):** an act whose verb isn't enabled is
scaffolded as a `proposed` tool (maven suggests). A human enables it (fills argv
+ destructive) on the authed **`mavweb /tools`** page — never voice — and can
disable it back to `proposed` (kept in the store, won't run). `EnableTool`/
`DisableTool` sit at `AuthStepUp`; the gate is now **live** via `PasskeySession`,
so /tools enable requires a passkey assertion at `/auth/passkey` first.
- **Confirm turn:** a destructive enabled tool replies "выполнить X? да/нет" and
parks; the next utterance (ru/en yes-no) confirms or cancels (90s TTL).
- **Config:** `voice.tools` seeds enabled tools at boot (editing mavend.json =
the human enable act); mavweb enables ad-hoc ones on top.
- **Russian:** fixed grammar in reply strings + seed files; maven's self-
reference is feminine ("she") — [[maven-persona-gender]].
Also fixed:
- **HashEmbedder was blind to Cyrillic** (`tokenize` iterated bytes, kept only
`a-z0-9`) → every RU utterance embedded to the zero vector → cosine 0 across
all intents → misrouted to `act` (alphabetical tie-break). Now rune-based
(`unicode.IsLetter`). This was the real cause of "Найди заметку" (a query)
landing in `notes`; added note-retrieval query seeds too.
- **Notes are now browsable on `/dash`** — `RecentNotes` plumbed through the
store + CoreAPI; voice-captured notes were previously only reachable via
semantic `query`.
Earlier: notes/query recall, `/dash` monitoring, `wg_handshake` poller (NO-OP).
### Gaps — why "voice assistant" is still aspirational (2026-07-06)
What separates Maven today from the thing the spec describes. Dealbreakers
first — these define the category:
1. **Always-on listening is code-complete (MVP).** `cmd/mavwaked` captures
PCM from arecord → energy-based VAD → PushToTalk with `Surface=SurfaceVoice`
(L0). Gap narrowed: no wake word yet (pure voice-activity trigger; every
utterance fires). The 30ms frame shape and 16kHz PCM match silero-vad's
ONNX input exactly, so a wake-word model swap is a local change in vad.go.
Hardware: the mic lives on a client box (desk PC, pi, etc.) — never the
homesrv. Deploy action: systemd user unit on whichever box has the mic,
connects to mavend over wg or local net.
2. **Conversation is deeper now, still not full dialogue.** The router
classifies one utterance → one reply, but `internal/dialogue` carries
context across turns: a 2-min session inherits slots for same-intent
follow-ups («напомни завтра» → «…позвонить маме»), and cross-intent
anaphora («запиши что я пил воду» → «когда я это сделал?») now resolves
RU pronouns (это/он/она/оно/тот/мой + inflections) to the prior turn's
key for fact-by-key lookup. `Session.History []Turn` is the scaffold for
real multi-turn. Still missing: LLM-driven dialogue manager (decide
ask-vs-act), anaphora beyond RU pronouns, single-slot session (single-user
box). The sub-1B phraser only words replies.
3. **Latency/shape of a turn.** Clip-based STT (record → upload → whisper →
route → phrase → piper → play). No streaming either direction, no barge-in;
every exchange is a full round trip.
Capability-class gaps — built but thin:
4. **Act surface is a small argv allowlist.** propose→enable works and the
allowlist now ships a homelab starter set (jul6 task 2 — status/ps/uptime/
df/free/logs read-only, restart/stop/reboot gated). Still bounded to what's
seeded; broadening it is config, not code.
5. **Query answers now cover notes + calendar + weather + general knowledge**
(jul6 tasks 3/4/5). Calendar querying, keyless Open-Meteo weather, and a
phraser knowledge-fallback all landed; caveat — general-knowledge quality is
only as good as the sub-1B phraser, and weather needs `voice.weather.provider`
set. The cheatsheet and router are now roughly aligned.
6. **Routing quality depends on the ONNX embedder being configured** — the
HashEmbedder floor makes RU recall lexical/weak; many commands fall to
"clarify". `make download-embedder` now fetches the multilingual MiniLM
model + AGENTS.md documents libonnxruntime setup; `voice.query_min_score`
is a config knob (default 0.55) so the floor can be tuned without recompile.
7. **Presence is effectively one signal** (page_heartbeat); desk_active is
still an undeployed script — "voice when near" routing runs on a guess.
8. **Long-term memory is now persistent (store-backed), not the spec's chroma.**
`internal/memory` has a `Store` interface; the daemon now wires
`store.MemoryStore` (`internal/store/memory.go`) — a **persistent** backend
in the **same encrypted sqlite db** (survives restarts; recall text inherits
at-rest encryption, so no plaintext sidecar). Vectors are float32 blobs,
search is brute-force cosine (fine at single-user scale; ANN is the later
swap behind the same interface). Notes **and facts** are indexed on capture;
`IntentQuery` reads it back (after notes-RAG misses, before general-knowledge)
— fact recall («когда я пил воду?») is its distinct payoff. The in-memory
impl remains the test/no-store floor. Remaining: an ANN/external index is
optional-scale, not a gap. Custom TTS voice (kami-picked, replaces the irina
floor — [[custom-voice-training]]) is still a future item.
Ops footnote: voice-over-web verified 2026-07-06 — mavend binds 0.0.0.0:9100
and mavweb reaches it cross-container at mavend:9100 (nc -z confirmed).
mavpoll uses network_mode=host to reach localhost services (netdata, kuma).
### Future / logged, not now
Custom TTS voice training (kami-picked voice, replaces irina floor); listening
modes 23 (meeting-record, ambient-derive).
### Services & layout
- `mavend` (core, IPC unix socket) — store + loop + phraser; the only key-holder.
- `mavsttd` / `mavttsd` — STT/TTS worker modules (unix sockets).
- `mavweb` — PWA bridge (HTTP), `/api/ptt` voice, `/api/signal` presence ingest,
`/api/ntfy` WS-subscribe config, `/dash` read-only monitoring.
- `mavpoll` — env poller (netdata/kuma → facts via CoreAPI).
- `mavcaldav` — CalDAV poller (Radicale → `calendar_busy` + events via CoreAPI).
- All behind wg + nginx deny-all; no phone-home. CGo only in `mavsttd`.
- Start/stop: `./start-maven.sh [build]`, `./kill-maven.sh`.
- Config: `~/.config/maven/mavend.json` (or `mavend.json` in repo root).
### Key files
- `cmd/mavend/{main,tick,voice}.go` — daemon wiring, loop driver, voice handler
- `internal/loop/{loop,rules,gather,feedback}.go` — proactive engine
- `internal/store/` — append-only facts/reminders/nudges/presence/notes
- `cmd/mavweb/{main.go,dash.html}` — PWA bridge + `/dash` monitoring
- `internal/router/{classifier,slots,stage0}.go` — reactive routing + slot parse
- `internal/delivery/` — dispatcher + ntfy/telegram/voice sinks
- `internal/auth/` — scope/gate/policy; `FloorEnrollment` (same-uid = device
trust) + `webauthn.PasskeySession` (real step-up for L3)
- `internal/webauthn/`, `cmd/mavweb/webauthn.go` — passkey register/assert
- `cmd/mavcaldav/`, `cmd/mavpoll/`, `scripts/desk-active.sh` — env producers
### Why multi-user (SPEC item 8) is deferred
Not neglect — the one item where doing nothing now beats doing something:
- **No second user exists yet** (the "gf phase"). Building per-user partitioning
now means code exercised by zero users and validated by nobody — YAGNI.
- **The append-only schema makes it a migration, not a rewrite.** No row is ever
mutated, so adding `facts/notes/reminders.user_id` later is add-columns +
backfill-to-"kami" — no reshaping, no dual-write window. Deferral is cheap.
- **The hard part is speaker attribution, and it needs the second voice.** A
voice-print discriminator (kami vs gf vs unknown) can't be trained or tuned
with one voice in the house. Plumbing before the model is pipe with no water.
- **It's fenced deliberately** (`DO NOT TOUCH THIS PHASE` in `DESIGN.md` § Users) so an
autonomous agent doesn't add `user_id` columns while touching the store and
commit us to a schema before the constraints that shape it exist.
+1 -1
View File
@@ -1,6 +1,6 @@
// Package main is mavenclient — maven's reference client. // Package main is mavenclient — maven's reference client.
// //
// Per DESIGN.md § Voice pipeline (STT / TTS): capture lives on the client; // Per docs/design.md § Voice pipeline (STT / TTS): capture lives on the client;
// the server transcribes + synthesises on demand. The PC client runs the // the server transcribes + synthesises on demand. The PC client runs the
// wake-word / VAD gate (cmd/mavwaked) and ships ONE clean audio blob per // wake-word / VAD gate (cmd/mavwaked) and ships ONE clean audio blob per
// utterance on activation. The server never owns a mic. // utterance on activation. The server never owns a mic.
+1 -1
View File
@@ -1,5 +1,5 @@
// mavend/simulator_test.go — the replayable full-system simulator // mavend/simulator_test.go — the replayable full-system simulator
// (Vikunja #284, 20-07-2026-BACKLOG.md item 7). // (Vikunja #284).
// //
// # What it is // # What it is
// //
+1 -1
View File
@@ -512,7 +512,7 @@ func main() {
// /tools — the authed enable surface. maven proposes acts she can't run; // /tools — the authed enable surface. maven proposes acts she can't run;
// this page is where a human reviews and enables them (proposed→enabled). // this page is where a human reviews and enables them (proposed→enabled).
// Enabling is the boundary-moving act (DESIGN.md § Tool registration — // Enabling is the boundary-moving act (docs/design.md § Tool registration —
// drafting is suggest, enabling is act), so it lives ONLY here, // drafting is suggest, enabling is act), so it lives ONLY here,
// behind wg+nginx+auth — never the voice/chat path. // behind wg+nginx+auth — never the voice/chat path.
mux.HandleFunc("/tools", func(w http.ResponseWriter, r *http.Request) { mux.HandleFunc("/tools", func(w http.ResponseWriter, r *http.Request) {
@@ -1,6 +1,45 @@
# maven — feature ranking # maven — feature ranking
> dated 2026-07-03. companion to `DESIGN.md` (folded from the former `maven.md`). ranks everything discussed post-repo-state against the infra blockers, not a replacement for the build order. > **Archived 2026-08-02 (V-447).** The mandatory and easy tiers are now Vikunja tasks
> 449-458. Two of those closed immediately, because the ranking was stale. Quiet hours
> (V-450) ship as `QuietHoursConfig` plus the care gate in `internal/loop/loop.go`.
> Schema migrations (V-451) ship as `internal/store/migrations.go` on `PRAGMA
> user_version`. The doable and epic tiers stay here because they are reasoning. Some of
> them exist only to record why something is not worth doing yet. Read this for the why,
> not as a work queue, and check the code before believing a gap.
> dated 2026-07-03. companion to `docs/design.md` (folded from the former `maven.md`). ranks everything discussed post-repo-state against the infra blockers, not a replacement for the build order.
---
## what already shipped (checked against the code, 2026-08-02)
One month old and already wrong in ten places. Everything below is marked
built after reading the code, not the board. Read the tiers underneath with this list in
hand.
Infra 1 and 2, the two the ranking says block every feature, are both done. sqlcipher
at-rest ships as `Store.enc` plus `OpenEncrypted`, a tmpfs working copy re-encrypted on
`Close`, keyed from `db_key_env` in `deploy/mavend.json`. mavweb and mavcaldav are no
longer at zero coverage: seven test files under `cmd/mavweb`, including
`credentials_test.go` and `passkey_prf_test.go`, and two under `cmd/mavcaldav`. Infra 3
is stale in the other direction. There are still no systemd units, but the deploy is
`deploy/ecosystem/docker-compose.yml`, not scripts and tmux.
Doable tier, built: rule trace and explanation as `/trace` plus `internal/loop/explain.go`.
Recurring reminders as the cron column, `NextFireTs` and `RescheduleReminder`
(`internal/store/reminders.go:206`), so "fires once right now" is wrong. Stale-reminder
burst collapse as `collapseReminders` (`internal/loop/gather.go:209`). Revert as
`VoidLatestFact` (`internal/store/facts.go:300`). Digest mode as `internal/store/digest.go`.
Testing infra as the simulator and the eval lab (V-284, V-278). Passkey persistence as the
JSON-backed `credentialStore` in `cmd/mavweb/credentials.go`. Most integrations shipped as
their own QA tasks (V-246 mail, V-256 smarthome, V-258 rss, V-259 crawler).
Doable tier, still open: correx, systemd units, memory decay and duplicate detection
(nothing in `internal/memory` touches it), backup automation, import and export, barge-in.
Epic tier, unbuilt as ranked. `event.Bus` exists (`cmd/mavend/intake.go:57`) but it is the
intake journal from V-283, not the rewrite of facts into projections that this tier means.
--- ---
@@ -20,7 +59,7 @@ nothing feature-level below should land before 12 are done. 34 can interle
### mandatory ### mandatory
things that block correctness or safety of stuff already shipped — not new capability, just closing gaps in existing design. things that block correctness or safety of stuff already shipped — not new capability, just closing gaps in existing design.
- **destructive-confirm policy** — open question in `DESIGN.md` § open questions, blocks correx and any new tool domain from having a coherent risk tier - **destructive-confirm policy** — open question in `docs/design.md` § open questions, blocks correx and any new tool domain from having a coherent risk tier
- **quiet-hours definition** — open question, blocks proactive delivery being trustworthy - **quiet-hours definition** — open question, blocks proactive delivery being trustworthy
- **schema migrations** — sqlcipher rollout alone forces a schema touch. want this mechanism before that, not after. - **schema migrations** — sqlcipher rollout alone forces a schema touch. want this mechanism before that, not after.
@@ -30,7 +69,7 @@ cheap, no dependencies, no new invariants.
- **grocery / `list_items` table** — fourth append-only shape (item, status, list-tag), no predicate touches it, multi-adder just works for free - **grocery / `list_items` table** — fourth append-only shape (item, status, list-tag), no predicate touches it, multi-adder just works for free
- **go.mod tidy** - **go.mod tidy**
- **capability model** (deepseek) — `homelab.docker.restart` instead of flat `tool→enabled`. cheap now, expensive to retrofit once tools surface passes ~15 entries. time-sensitive, not urgent. - **capability model** (deepseek) — `homelab.docker.restart` instead of flat `tool→enabled`. cheap now, expensive to retrofit once tools surface passes ~15 entries. time-sensitive, not urgent.
- **conversation repair** — already free: `DESIGN.md` has "misroute correction = new centroid example," this is just naming the existing mechanism as a feature - **conversation repair** — already free: `docs/design.md` has "misroute correction = new centroid example," this is just naming the existing mechanism as a feature
- **command history** — read-only query over existing facts, no new mechanism - **command history** — read-only query over existing facts, no new mechanism
- **clarification templates** — canned phrasing for the router's existing confidence-gate fallback, phraser-lane only - **clarification templates** — canned phrasing for the router's existing confidence-gate fallback, phraser-lane only
- **pronunciation dictionary** — tts config, no architecture - **pronunciation dictionary** — tts config, no architecture
@@ -30,7 +30,7 @@ critical workflows: voice turn (mic→STT→route→tool/reply→TTS); proactive
(/dash /chat /tools /ecosystem) (/dash /chat /tools /ecosystem)
current state: all 34 test packages pass; vet clean; `make test` still exits 1 current state: all 34 test packages pass; vet clean; `make test` still exits 1
(finding 2) (finding 2)
known failures: weak RU query routing — REARCH.md names the cause; the named fix known failures: weak RU query routing — docs/rearchitecture.md names the cause; the named fix
is wired `nil` is wired `nil`
maintenance burden: 4,518 lines of root markdown vs 33,319 lines of Go; 15 top-level maintenance burden: 4,518 lines of root markdown vs 33,319 lines of Go; 15 top-level
.md files, 3 of them dated session logs; several contradict .md files, 3 of them dated session logs; several contradict
@@ -47,11 +47,11 @@ what's obsolete: llmrouter.go (built, tested, never wired); classifier seed-p
``` ```
**Classification: healthy + misaligned.** Not fragile, not overbuilt, not abandoned. The **Classification: healthy + misaligned.** Not fragile, not overbuilt, not abandoned. The
architecture in `REARCH.md` is sound and mostly *built* — it just is not *connected*. architecture in `docs/rearchitecture.md` is sound and mostly *built* — it just is not *connected*.
## what it should become ## what it should become
The thing `REARCH.md` already describes, with the switch flipped and the drift removed: The thing `docs/rearchitecture.md` already describes, with the switch flipped and the drift removed:
one resident small model doing both routing and phrasing, classifier demoted from the live one resident small model doing both routing and phrasing, classifier demoted from the live
path to the failure floor, embedder demoted to RAG hint. No new architecture is needed. path to the failure floor, embedder demoted to RAG hint. No new architecture is needed.
**The gap is a config/wiring decision plus doc convergence, not a redesign.** **The gap is a config/wiring decision plus doc convergence, not a redesign.**
@@ -65,19 +65,19 @@ path to the failure floor, embedder demoted to RAG hint. No new architecture is
`architecture` / `repair` `architecture` / `repair`
**problem:** The most load-bearing design decision in the project is stated four different, **problem:** The most load-bearing design decision in the project is stated four different,
incompatible ways, and the code path `REARCH.md` calls "the linchpin" is disabled. incompatible ways, and the code path `docs/rearchitecture.md` calls "the linchpin" is disabled.
**evidence** (all confirmed): **evidence** (all confirmed):
- `cmd/mavend/voice.go:211``rtr := buildRouter(emb, matcher, threshold, nil) // LLM router disabled`, - `cmd/mavend/voice.go:211``rtr := buildRouter(emb, matcher, threshold, nil) // LLM router disabled`,
with comment *"the classifier handles routing reliably."* with comment *"the classifier handles routing reliably."*
- `REARCH.md:11` says the same classifier is *"the structural cause of 'she messes up - `docs/rearchitecture.md:11` says the same classifier is *"the structural cause of 'she messes up
queries.'"* **The code comment and the design doc make opposite claims about the same queries.'"* **The code comment and the design doc make opposite claims about the same
component.** component.**
- `internal/router/llmrouter.go` (139 lines) + `llmrouter_test.go` — fully built and - `internal/router/llmrouter.go` (139 lines) + `llmrouter_test.go` — fully built and
tested, zero non-test callers. tested, zero non-test callers.
- Model identity, four ways: docs say **Qwen3-1.7B** (`CLAUDE.md:6`, `REARCH.md:15`, - Model identity, four ways: docs say **Qwen3-1.7B** (`CLAUDE.md:6`, `docs/rearchitecture.md:15`,
`SPEC.md:46`, `AGENTS.md:79`, `MAVEN_ECOSYSTEM_ARCHITECTURE.md:72`); `SPEC.md:46`, `AGENTS.md:79`, `docs/ecosystem.md:72`);
`deploy/mavend.json:9` says **Qwen3.5-2B-UD-Q4_K_XL**; `models/llm/` on disk holds `deploy/mavend.json:9` says **Qwen3.5-2B-UD-Q4_K_XL**; `models/llm/` on disk holds
**LFM2.5-1.2B-Thinking**; code comments in 5 files still say **LFM**. **LFM2.5-1.2B-Thinking**; code comments in 5 files still say **LFM**.
- `deploy/mavend.json:11` sets `"n_gpu_layers": 99` while `CLAUDE.md:4` states the target - `deploy/mavend.json:11` sets `"n_gpu_layers": 99` while `CLAUDE.md:4` states the target
@@ -100,7 +100,7 @@ match. Delete nothing from `internal/router` yet — the classifier is the fallb
reconciliation, not a refactor. **Do not rewrite the router.** reconciliation, not a refactor. **Do not rewrite the router.**
**alternatives:** Delete `llmrouter.go` and commit to the classifier — only defensible if **alternatives:** Delete `llmrouter.go` and commit to the classifier — only defensible if
the eval harness shows the classifier is actually adequate, which contradicts `REARCH.md`. the eval harness shows the classifier is actually adequate, which contradicts `docs/rearchitecture.md`.
**risk:** Low-moderate. LLM route failures already fall through to the classifier **risk:** Low-moderate. LLM route failures already fall through to the classifier
(`router.go:88-96`), so a bad model cannot break a turn. The real risk is CPU latency. (`router.go:88-96`), so a bad model cannot break a turn. The real risk is CPU latency.
@@ -228,21 +228,21 @@ after finding 1, not before** — and skip it if it stays purely cosmetic.
**problem:** 15 root markdown files, 4,518 lines, several stale or superseded, at least **problem:** 15 root markdown files, 4,518 lines, several stale or superseded, at least
three pairs contradicting each other. three pairs contradicting each other.
**evidence:** `ROADMAP.md` (759) + `MAVEN_ECOSYSTEM_ARCHITECTURE.md` (884) + **evidence:** `ROADMAP.md` (759) + `docs/ecosystem.md` (884) +
`PROGRESS.md` (456) + `maven.md` (413) + `20-07-2026-BACKLOG.md` (396) + `PROGRESS.md` (456) + `maven.md` (413) + `20-07-2026-BACKLOG.md` (396) +
`SESSION-05-07-2026.md` + `SESSION-06-07-2026.md` (477 combined) + `PLANS.md` (25) + `SESSION-05-07-2026.md` + `SESSION-06-07-2026.md` (477 combined) + `PLANS.md` (25) +
`START.md` + `SPEC.md` + `PROTOCOL.md`. `PROGRESS.md:61` annotates its own staleness: `docs/operations.md` + `SPEC.md` + `docs/protocol.md`. `PROGRESS.md:61` annotates its own staleness:
*"Older LFM references below describe the currently deployed..."*. `REARCH.md` announces it *"Older LFM references below describe the currently deployed..."*. `docs/rearchitecture.md` announces it
"supersedes" a model still described as current elsewhere. "supersedes" a model still described as current elsewhere.
**impact:** The doc set is the reason finding 1 exists. When five documents describe the **impact:** The doc set is the reason finding 1 exists. When five documents describe the
architecture, the code becomes the only trustworthy one — which defeats the purpose of architecture, the code becomes the only trustworthy one — which defeats the purpose of
having them. having them.
**recommended action:** Keep `CLAUDE.md` (agent contract), `REARCH.md` (target **recommended action:** Keep `CLAUDE.md` (agent contract), `docs/rearchitecture.md` (target
architecture), `AGENTS.md` (recipes), `PROTOCOL.md` (wire format), architecture), `AGENTS.md` (recipes), `docs/protocol.md` (wire format),
`20-07-2026-BACKLOG.md` (live queue). Delete the two `SESSION-*.md` and `PLANS.md` — git `20-07-2026-BACKLOG.md` (live queue). Delete the two `SESSION-*.md` and `PLANS.md` — git
history holds them. Fold `SPEC.md` + `maven.md` + `ROADMAP.md` into one `DESIGN.md` and history holds them. Fold `SPEC.md` + `maven.md` + `ROADMAP.md` into one `docs/design.md` and
mark superseded sections instead of leaving them to read as current. Target ~1,500 lines. mark superseded sections instead of leaving them to read as current. Target ~1,500 lines.
--- ---
@@ -320,7 +320,7 @@ expected maintenance gain: none over the incremental path
suggesting Vulkan offload is intended and working — but `CLAUDE.md` says CPU-only. Likely suggesting Vulkan offload is intended and working — but `CLAUDE.md` says CPU-only. Likely
the doc is stale, not the config; unverified. the doc is stale, not the config; unverified.
- **Whether the classifier is genuinely adequate.** `voice.go:211` asserts it is; - **Whether the classifier is genuinely adequate.** `voice.go:211` asserts it is;
`REARCH.md` asserts it is not. Both are claims, neither is measured. The uncommitted eval `docs/rearchitecture.md` asserts it is not. Both are claims, neither is measured. The uncommitted eval
harness is the instrument to settle it — resolve before flipping the router, not after. harness is the instrument to settle it — resolve before flipping the router, not after.
- **Whether wg+nginx+auth actually fronts 9201 in production.** Not in this repo. If it - **Whether wg+nginx+auth actually fronts 9201 in production.** Not in this repo. If it
does, finding 3 drops from "unauthenticated RCE" to "the control is not reproducible from does, finding 3 drops from "unauthenticated RCE" to "the control is not reproducible from
@@ -350,7 +350,7 @@ Key claims independently re-verified against the working tree; the verdict stand
over WireGuard on homesrv, this is hygiene, not an emergency — but the loopback bind over WireGuard on homesrv, this is hygiene, not an emergency — but the loopback bind
and startup warning are cheap insurance either way, so do them regardless. and startup warning are cheap insurance either way, so do them regardless.
- Finding 1's "flip the router" step should be gated harder on measurement. - Finding 1's "flip the router" step should be gated harder on measurement.
`REARCH.md`'s claim that the classifier causes weak RU queries is itself unmeasured — `docs/rearchitecture.md`'s claim that the classifier causes weak RU queries is itself unmeasured —
the review admits this under uncertainties, but the "repair now" ordering buries it. the review admits this under uncertainties, but the "repair now" ordering buries it.
Run `eval_scenarios_test.go` against both paths **before** deciding to flip, not Run `eval_scenarios_test.go` against both paths **before** deciding to flip, not
after. A 2B model on CPU may add enough latency that the classifier wins in practice after. A 2B model on CPU may add enough latency that the classifier wins in practice
+10 -8
View File
@@ -1,10 +1,12 @@
# Maven — Design # Maven — Design
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
> Folded 2026-07-30 from `SPEC.md` (north star, 2026-07-03), `maven.md` > Folded 2026-07-30 from `SPEC.md` (north star, 2026-07-03), `maven.md`
> (consolidated decisions, 2026-06-30) and `ROADMAP.md` (execution plan, > (consolidated decisions, 2026-06-30) and `ROADMAP.md` (execution plan,
> 2026-07-06). Those three files are gone; git history holds them. > 2026-07-06). Those three files are gone; git history holds them.
> This is the single design document: principles, target state, and the > This is the single design document: principles, target state, and the
> execution ledger. `REARCH.md` remains authoritative wherever it disagrees > execution ledger. `docs/rearchitecture.md` remains authoritative wherever it disagrees
> with anything here. Everything the three sources asserted that is no longer > with anything here. Everything the three sources asserted that is no longer
> the intended design is preserved under **§ Superseded** — do not read that > the intended design is preserved under **§ Superseded** — do not read that
> section as current. > section as current.
@@ -160,7 +162,7 @@ presence_state ( last_bucket, last_score, updated_ts )
``` ```
Facts additionally carry `Subject`/`EntityID`/`ResolutionState` for Facts additionally carry `Subject`/`EntityID`/`ResolutionState` for
entity-aware resolution against Nexus (see `MAVEN_ECOSYSTEM_ARCHITECTURE.md`). entity-aware resolution against Nexus (see `docs/ecosystem.md`).
### Trigger model ### Trigger model
@@ -196,7 +198,7 @@ INTO the gate as an env predicate, not the LLM's job.
## Reactive path — routing ## Reactive path — routing
**Target design: LLM-as-router** (see `REARCH.md` and `CLAUDE.md`). One **Target design: LLM-as-router** (see `docs/rearchitecture.md` and `CLAUDE.md`). One
resident model emits GBNF-constrained structured JSON, and the same model resident model emits GBNF-constrained structured JSON, and the same model
phrases replies; the embedder is a RAG hint, not a routing gate. The phrases replies; the embedder is a RAG hint, not a routing gate. The
committed default today is the classifier/embedder cascade, which is an committed default today is the classifier/embedder cascade, which is an
@@ -655,7 +657,7 @@ daemon.
The voice wire protocol (length-prefixed JSON frames over TCP) is designed for The voice wire protocol (length-prefixed JSON frames over TCP) is designed for
**multiple client implementations**. The reference PWA at `cmd/mavweb` is one **multiple client implementations**. The reference PWA at `cmd/mavweb` is one
client; any app (phone, desktop CLI, smartwatch) can implement the same frame client; any app (phone, desktop CLI, smartwatch) can implement the same frame
protocol. The published spec is `PROTOCOL.md` — **generated from protocol. The published spec is `docs/protocol.md` — **generated from
`internal/voice/wire.go`**, not composed freehand, so it can't drift from `internal/voice/wire.go`**, not composed freehand, so it can't drift from
code. It covers transport (4-byte big-endian length prefix), methods code. It covers transport (4-byte big-endian length prefix), methods
(`PushToTalk`, `Pong`), push kinds (`AudioNudge`), surface identity (`PushToTalk`, `Pong`), push kinds (`AudioNudge`), surface identity
@@ -670,8 +672,8 @@ Broadening to home automation, media or comms is JSON, not code.
## Execution ledger ## Execution ledger
Condensed from `ROADMAP.md` (2026-07-06). The live queue is Condensed from `ROADMAP.md` (2026-07-06). The live queue is the Vikunja board
`20-07-2026-BACKLOG.md`; current state is `PROGRESS.md`. (project Maven, ID 2); this table is history, not a work list.
| # | Item | Prio | Status | | # | Item | Prio | Status |
|---|------|------|--------| |---|------|------|--------|
@@ -755,7 +757,7 @@ Kept for provenance. **None of this is the current or intended design.**
stay deterministic — "classifier owns the route, the SLM stays in its stay deterministic — "classifier owns the route, the SLM stays in its
phrasing lane" — with an embedding + nearest-centroid stage 1 over ~10 phrasing lane" — with an embedding + nearest-centroid stage 1 over ~10
examples per intent, and misroutes appended as new centroid examples. examples per intent, and misroutes appended as new centroid examples.
*Replaced by* LLM-as-router (`REARCH.md`): one resident model emits *Replaced by* LLM-as-router (`docs/rearchitecture.md`): one resident model emits
GBNF-constrained JSON and also phrases replies; the embedder is demoted to GBNF-constrained JSON and also phrases replies; the embedder is demoted to
a RAG hint. *Landed 2026-07-31:* the LLM router is on by default and set a RAG hint. *Landed 2026-07-31:* the LLM router is on by default and set
`true` in `deploy/mavend.json`. The classifier cascade stays as the failure `true` in `deploy/mavend.json`. The classifier cascade stays as the failure
@@ -774,7 +776,7 @@ Kept for provenance. **None of this is the current or intended design.**
*Resolved 2026-07-30 (#318), revised 2026-07-31:* the resident checkpoint is *Resolved 2026-07-30 (#318), revised 2026-07-31:* the resident checkpoint is
stock **Qwen3-1.7B** (`UD-Q4_K_XL`, `n_ctx` 4096), which replaced stock **Qwen3-1.7B** (`UD-Q4_K_XL`, `n_ctx` 4096), which replaced
Qwen3.5-0.8B after measuring better on both fixtures Qwen3.5-0.8B after measuring better on both fixtures
(`MODEL-BAKEOFF-31-07-2026.md`). The CPT'd **Qwen3-1.7B** remains the target (`docs/evals/2026-07-31-model-bakeoff.md`). The CPT'd **Qwen3-1.7B** remains the target
(#122); what stock gets wrong is the persona, not the Russian. Note the resident (#122); what stock gets wrong is the persona, not the Russian. Note the resident
model is no longer described as untrained — the target is trained model is no longer described as untrained — the target is trained
end-to-end, which is the substantive change from the old claim. end-to-end, which is the substantive change from the old claim.
@@ -1,5 +1,7 @@
# Deterministic logic around a small model # Deterministic logic around a small model
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
Written 2026-08-02. Branch `fix/integrated`. Written 2026-08-02. Branch `fix/integrated`.
## The question ## The question
@@ -268,7 +270,7 @@ rebuilt `mavend` wires it. "кто написал войну и мир?" now rou
that turn left unsettled. that turn left unsettled.
- **The turn was slow, and nobody knows yet whether that is real.** Route 7s, - **The turn was slow, and nobody knows yet whether that is real.** Route 7s,
search 1s, phrasing 15s. The p50 in `ROUTING-EVAL-31-07-2026.md` is 825ms. It search 1s, phrasing 15s. The p50 in `docs/evals/2026-07-31-routing.md` is 825ms. It
was the first turn after a cold start with the model still warming, so it was the first turn after a cold start with the model still warming, so it
proves nothing either way. Re-run the same question warm before treating it as proves nothing either way. Re-run the same question warm before treating it as
a regression. Do not plan latency work off this number. a regression. Do not plan latency work off this number.
@@ -1,5 +1,7 @@
# Maven Ecosystem Architecture # Maven Ecosystem Architecture
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
## 1. Purpose ## 1. Purpose
This document defines Maven's role in the local ecosystem formed by: This document defines Maven's role in the local ecosystem formed by:
@@ -16,7 +16,7 @@ with Qwen3-1.7B.
Settles Vikunja **#278 / #250**. Settles Vikunja **#278 / #250**.
- Same fixture and scorer as `ROUTING-EVAL-31-07-2026.md`: `internal/router/eval/` - Same fixture and scorer as `docs/evals/2026-07-31-routing.md`: `internal/router/eval/`
(`ru_routing_v1.json`, 76 held-out cases). (`ru_routing_v1.json`, 76 held-out cases).
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:<port> make eval-router` - Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:<port> make eval-router`
(`TestLLMRouterBaseline`). (This line used to say there is no `make eval-models` target. (`TestLLMRouterBaseline`). (This line used to say there is no `make eval-models` target.
@@ -148,7 +148,7 @@ Qwen3-1.7B wins every column, including against a model 20% larger than it.
| ontopic | 16, 19, 19 | **22, 23, 23** | | ontopic | 16, 19, 19 | **22, 23, 23** |
| canned fallbacks | 8, 5, 6 | **0, 2, 0** | | canned fallbacks | 8, 5, 6 | **0, 2, 0** |
This also fills the row `TALK-EVAL-31-07-2026.md` had to void for contamination: This also fills the row `docs/evals/2026-07-31-talk.md` had to void for contamination:
**600ch/1024tok on Qwen3.5-0.8B scores 13, 11, 8.** **600ch/1024tok on Qwen3.5-0.8B scores 13, 11, 8.**
`address` is the headline. It sat at 18-22 of 27 on the 0.8B no matter how the prompt `address` is the headline. It sat at 18-22 of 27 on the 0.8B no matter how the prompt
@@ -164,7 +164,7 @@ The 1.7B does that 0-2 times.
> **Stale, corrected 2026-08-02.** The p50 figures in this table are contention on a > **Stale, corrected 2026-08-02.** The p50 figures in this table are contention on a
> shared llama-server, not the model's cost. The router measures p50 825ms / p95 1.2s / > shared llama-server, not the model's cost. The router measures p50 825ms / p95 1.2s /
> max 3.0s in `ROUTING-EVAL-31-07-2026.md`, which says so at line 61. Read this table for > max 3.0s in `docs/evals/2026-07-31-routing.md`, which says so at line 61. Read this table for
> the shape of the tail only. Take absolute latency from the routing eval. > the shape of the tail only. Take absolute latency from the routing eval.
| | p50 | p95 | | | p50 | p95 |
@@ -220,11 +220,11 @@ swapped again when the CPT lands.
behind `voice.llm_router`, the default is on, and `deploy/mavend.json` sets it `true`. behind `voice.llm_router`, the default is on, and `deploy/mavend.json` sets it `true`.
These numbers are the production path now. **Corrected 2026-08-02: the p50 ≈2.7s in the These numbers are the production path now. **Corrected 2026-08-02: the p50 ≈2.7s in the
latency table above WAS a bench artifact.** It is contention on the shared llama-server, latency table above WAS a bench artifact.** It is contention on the shared llama-server,
not the model. `ROUTING-EVAL-31-07-2026.md` line 61 says so, and measures the router at not the model. `docs/evals/2026-07-31-routing.md` line 61 says so, and measures the router at
p50 825ms / p95 1.2s / max 3.0s. Cite that file for latency, not this one. p50 825ms / p95 1.2s / max 3.0s. Cite that file for latency, not this one.
- ~~`/mnt/hdd1/llms/LFM2.5/Qwen3-1.7B-UD-Q4_K_XL.gguf` is a 293 MB truncated download - ~~`/mnt/hdd1/llms/LFM2.5/Qwen3-1.7B-UD-Q4_K_XL.gguf` is a 293 MB truncated download
in the wrong directory.~~ **Deleted 2026-07-31.** The good 1.13 GB copy in `qwen3/` is in the wrong directory.~~ **Deleted 2026-07-31.** The good 1.13 GB copy in `qwen3/` is
what `deploy/mavend.json` loads. what `deploy/mavend.json` loads.
- Harness: `scratchpad/bakeoff.sh`, one server at a time, health-checked before each - Harness: `scratchpad/bakeoff.sh`, one server at a time, health-checked before each
run, `/v1/models` recorded per run. Never run two LLM consumers at once — see the run, `/v1/models` recorded per run. Never run two LLM consumers at once — see the
contamination note in `TALK-EVAL-31-07-2026.md`. contamination note in `docs/evals/2026-07-31-talk.md`.
@@ -1,7 +1,7 @@
# Phrasing evaluation — 31-07-2026 # Phrasing evaluation — 31-07-2026
How Maven words a nudge, measured instead of argued. Counterpart to How Maven words a nudge, measured instead of argued. Counterpart to
`ROUTING-EVAL-31-07-2026.md`. `docs/evals/2026-07-31-routing.md`.
- Fixture + scorer: `internal/phraser/eval/` (`nudges_v1.json`, 15 cases; `eval.go`, `checks.go`) - Fixture + scorer: `internal/phraser/eval/` (`nudges_v1.json`, 15 cases; `eval.go`, `checks.go`)
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing` - Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing`
@@ -65,7 +65,7 @@ prefixes) is the targeted fix, and it would move findings 1 and 2 together. Sepa
`model_quantized.onnx` — not the same file. `model_quantized.onnx` — not the same file.
`hard` cases score **2/11**: every one is a query where the operator did not reuse his own words. `hard` cases score **2/11**: every one is a query where the operator did not reuse his own words.
That is the normal case weeks later, and exactly what DESIGN.md's "recall when relevant" promises. That is the normal case weeks later, and exactly what docs/design.md's "recall when relevant" promises.
### 4. The memory-store recall branch is dead for notes ### 4. The memory-store recall branch is dead for notes
@@ -177,7 +177,7 @@ was silent ("не знаю" to "который час") while the one it introdu
### 1. The resident model does route better — 50.0% vs 36.8% ### 1. The resident model does route better — 50.0% vs 36.8%
REARCH.md's premise holds; `voice.go:211`'s comment does not. **But the classifier is only docs/rearchitecture.md's premise holds; `voice.go:211`'s comment does not. **But the classifier is only
~37% correct on held-out utterances, and the model only ~50%.** Neither is "reliable". The ~37% correct on held-out utterances, and the model only ~50%.** Neither is "reliable". The
gap between them is real but both are far from a system you would describe as working. gap between them is real but both are far from a system you would describe as working.
@@ -141,7 +141,7 @@ a model check, which catches a dead server but not a loaded one.
measured) on this fixture and the router fixture. Not the 4B — too big for measured) on this fixture and the router fixture. Not the 4B — too big for
this box, owner's call. this box, owner's call.
- Newer sub-500M candidates (LFM2.5 200M/300M) are worth a run for routing. - Newer sub-500M candidates (LFM2.5 200M/300M) are worth a run for routing.
Note `MODEL-BAKEOFF-31-07-2026.md` found LFM2.5-**1.2B** worse than Note `docs/evals/2026-07-31-model-bakeoff.md` found LFM2.5-**1.2B** worse than
Qwen3.5-0.8B at Russian routing and 2.4× slower — but those are a different, Qwen3.5-0.8B at Russian routing and 2.4× slower — but those are a different,
older generation, so that result does not predict the small ones. older generation, so that result does not predict the small ones.
- Fix `chat-how-are-you`'s `want_any`, and re-baseline once, so `ontopic` - Fix `chat-how-are-you`'s `want_any`, and re-baseline once, so `ontopic`
+2
View File
@@ -1,5 +1,7 @@
# Start Commands # Start Commands
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
All commands assume `ROOT=/home/kami/apps/Maven` and the local Go toolchain at `$ROOT/deps/go/go/bin/go`. All commands assume `ROOT=/home/kami/apps/Maven` and the local Go toolchain at `$ROOT/deps/go/go/bin/go`.
## Prerequisites ## Prerequisites
@@ -6,7 +6,7 @@
> use the embedded LFM model paths or old single-object examples as current ops > use the embedded LFM model paths or old single-object examples as current ops
> guidance; see `2026-07-18-qwen3-resident-training-eval.md`. > guidance; see `2026-07-18-qwen3-resident-training-eval.md`.
> Scope from `REARCH.md`. Make Maven trustworthy: the LFM becomes the router > Scope from `docs/rearchitecture.md`. Make Maven trustworthy: the LFM becomes the router
> (fixes "messes up queries" / "doesn't take notes"), the engine actually runs > (fixes "messes up queries" / "doesn't take notes"), the engine actually runs
> (fixes stub replies), dates stop being read as "number dot number dot number", > (fixes stub replies), dates stop being read as "number dot number dot number",
> and telegram becomes a reach channel. NOT in scope: on-demand 4B reasoner, > and telegram becomes a reach channel. NOT in scope: on-demand 4B reasoner,
@@ -693,7 +693,7 @@ ssh kami@192.168.1.104 'curl -s localhost:9201/api/chat -d "{\"text\":\"запо
4. `docker compose up -d mavend && docker logs -f maven-mavend-1` — confirm the 4. `docker compose up -d mavend && docker logs -f maven-mavend-1` — confirm the
phraser spawns and no `phraser: NewStub` path. Run the two verify curls. phraser spawns and no `phraser: NewStub` path. Run the two verify curls.
5. Update `AGENTS.md`: LFM model download + note that routing is now LFM-first 5. Update `AGENTS.md`: LFM model download + note that routing is now LFM-first
with classifier fallback (`REARCH.md` is the design of record). with classifier fallback (`docs/rearchitecture.md` is the design of record).
--- ---
+2
View File
@@ -1,5 +1,7 @@
# Maven Voice Protocol # Maven Voice Protocol
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
> Auto-generated from `internal/voice/wire.go`, `internal/voice/errors.go`, > Auto-generated from `internal/voice/wire.go`, `internal/voice/errors.go`,
> `internal/voice/frame.go`, `internal/voice/client.go`. If this file and > `internal/voice/frame.go`, `internal/voice/client.go`. If this file and
> those files disagree, the code wins. > those files disagree, the code wins.
+2
View File
@@ -1,5 +1,7 @@
# QA plan: checking Maven properly # QA plan: checking Maven properly
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
Written 2026-08-01, after the 35-PR stack landed and the box came back up. Written 2026-08-01, after the 35-PR stack landed and the box came back up.
44 of the 50 open Vikunja tasks are `QA:` tasks. They are verification work, not 44 of the 50 open Vikunja tasks are `QA:` tasks. They are verification work, not
+2
View File
@@ -1,5 +1,7 @@
# Maven — Re-architecture (Qwen3 resident model, revised 2026-07-18) # Maven — Re-architecture (Qwen3 resident model, revised 2026-07-18)
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
> Supersedes the classifier-first routing model. Agreed in a design session > Supersedes the classifier-first routing model. Agreed in a design session
> after diagnosing that homesrv deploys with a **stub phraser** (no LLM > after diagnosing that homesrv deploys with a **stub phraser** (no LLM
> running) and an embedder-classifier that routes by nearest-neighbor between > running) and an embedder-classifier that routes by nearest-neighbor between
+1 -1
View File
@@ -1,7 +1,7 @@
// Package auth is maven's authority layer — the 4-layer cascade and the // Package auth is maven's authority layer — the 4-layer cascade and the
// "surface caps authority" invariant. // "surface caps authority" invariant.
// //
// Spec contract (from DESIGN.md § Auth): // Spec contract (from docs/design.md § Auth):
// //
// a cascade, not a pick-one — each layer answers a different question: // a cascade, not a pick-one — each layer answers a different question:
// //
+1 -1
View File
@@ -649,7 +649,7 @@ type VoiceConfig struct {
// LLMRouter — route with the resident model instead of the embedding // LLMRouter — route with the resident model instead of the embedding
// classifier. On by default since Vikunja #320. // classifier. On by default since Vikunja #320.
// //
// Measured on the held-out fixture (ROUTING-EVAL-31-07-2026.md): 63.2% of // Measured on the held-out fixture (docs/evals/2026-07-31-routing.md): 63.2% of
// intents right against the classifier's 50.0%, and no route errors. It // intents right against the classifier's 50.0%, and no route errors. It
// costs about 1s per turn instead of 30ms. // costs about 1s per turn instead of 30ms.
// //
+1 -1
View File
@@ -1,6 +1,6 @@
// Package delivery is maven's channel-routing + dispatch layer. // Package delivery is maven's channel-routing + dispatch layer.
// //
// Spec contract (from DESIGN.md § Delivery / channel routing): // Spec contract (from docs/design.md § Delivery / channel routing):
// //
// - routing = f(severity, presence). presence decides REACHABILITY; severity // - routing = f(severity, presence). presence decides REACHABILITY; severity
// decides INSISTENCE. need both. // decides INSISTENCE. need both.
+2 -2
View File
@@ -9,7 +9,7 @@ import (
"github.com/kami/maven/internal/store" "github.com/kami/maven/internal/store"
) )
// This file walks every cell of the DESIGN.md § "Delivery / channel routing" // This file walks every cell of the docs/design.md § "Delivery / channel routing"
// table, once as the pure table and once through the dispatcher, so a change // table, once as the pure table and once through the dispatcher, so a change
// to either side has to break a named cell. // to either side has to break a named cell.
// //
@@ -205,7 +205,7 @@ func TestAwayChannelsGetMinimalBody(t *testing.T) {
// An empty Summary no longer means "send the whole body" — it means a short // An empty Summary no longer means "send the whole body" — it means a short
// generic line — so the old expectation here was wrong as well as duplicated. // generic line — so the old expectation here was wrong as well as duplicated.
// TestCareAwayDropIsRecorded — DESIGN.md's drop is a decision ("a missed water // TestCareAwayDropIsRecorded — docs/design.md's drop is a decision ("a missed water
// nudge is noise, a missed backup failure isn't"), so it should be visible // nudge is noise, a missed backup failure isn't"), so it should be visible
// rather than vanish. Today drop is a bare `continue`: no nudge row, no outbox // rather than vanish. Today drop is a bare `continue`: no nudge row, no outbox
// attempt, no log — nothing an operator can see afterwards. now it leaves a // attempt, no log — nothing an operator can see afterwards. now it leaves a
+1 -1
View File
@@ -19,7 +19,7 @@
// 3. if no live session exists, Send returns voice.ErrNoSession // 3. if no live session exists, Send returns voice.ErrNoSession
// (wrapped). The daemon logs the partial dispatch; an OPEN deferred // (wrapped). The daemon logs the partial dispatch; an OPEN deferred
// question is whether the dispatcher should reroute to away-channels // question is whether the dispatcher should reroute to away-channels
// instead of returning partial — listed in PROGRESS.md. // instead of returning partial.
// //
// Import direction: voicesink imports internal/tts (synth seam) and // Import direction: voicesink imports internal/tts (synth seam) and
// internal/voice (Sessions registry). Both are siblings of delivery; the // internal/voice (Sessions registry). Both are siblings of delivery; the
+1 -2
View File
@@ -1,5 +1,4 @@
// Package event is the unified intake envelope (Vikunja #283, // Package event is the unified intake envelope (Vikunja #283).
// 20-07-2026-BACKLOG.md item 1).
// //
// # The problem it solves // # The problem it solves
// //
+1 -1
View File
@@ -31,7 +31,7 @@ type Completer interface {
// or reply in Russian. // or reply in Russian.
// //
// Why the JSON wrapper: this model always thinks out loud and this llama-server // Why the JSON wrapper: this model always thinks out loud and this llama-server
// build ignores the thinking switch (see ROUTING-EVAL-31-07-2026.md). A bare // build ignores the thinking switch (see docs/evals/2026-07-31-routing.md). A bare
// word-list grammar just captured the reasoning — every case came back as // word-list grammar just captured the reasoning — every case came back as
// "Let me analyze this request carefully". Demanding JSON, like routeGrammar and // "Let me analyze this request carefully". Demanding JSON, like routeGrammar and
// responseGrammar already do, gives the reasoning nowhere to go. // responseGrammar already do, gives the reasoning nowhere to go.
+5 -5
View File
@@ -10,12 +10,12 @@ import (
// Tests for the universal restraint gate. // Tests for the universal restraint gate.
// //
// DESIGN.md § Trigger model: "the gate is universal, applied by the loop, never // docs/design.md § Trigger model: "the gate is universal, applied by the loop, never
// per-rule — quiet-hours, presence, cooldown, snooze, calendar-busy all live in // per-rule — quiet-hours, presence, cooldown, snooze, calendar-busy all live in
// one fires()." These tests pin the CONSERVATIVE side of that: the cases where // one fires()." These tests pin the CONSERVATIVE side of that: the cases where
// Maven must stay quiet. They exist so nobody loosens the gate by accident. // Maven must stay quiet. They exist so nobody loosens the gate by accident.
// //
// Where the code does not yet do what DESIGN.md promises, the test is written to // Where the code does not yet do what docs/design.md promises, the test is written to
// show the gap and then skipped, with the file and line to fix. Behaviour is not // show the gap and then skipped, with the file and line to fix. Behaviour is not
// changed to make a test pass. // changed to make a test pass.
@@ -54,7 +54,7 @@ func TestGateQuietHoursSuppressesCareOnly(t *testing.T) {
// ---------------------------- presence --------------------------------------- // ---------------------------- presence ---------------------------------------
// DESIGN.md § Delivery: "sev <= 2 drops on away, sev >= 3 holds: a missed water // docs/design.md § Delivery: "sev <= 2 drops on away, sev >= 3 holds: a missed water
// nudge is noise, a missed backup failure isn't." // nudge is noise, a missed backup failure isn't."
func TestGateAwayDropsCareHoldsOps(t *testing.T) { func TestGateAwayDropsCareHoldsOps(t *testing.T) {
cases := []struct { cases := []struct {
@@ -250,7 +250,7 @@ func TestTickOrderOfRulesDoesNotMatter(t *testing.T) {
// ---------------------------- reminders bypass the gate ---------------------- // ---------------------------- reminders bypass the gate ----------------------
// DESIGN.md § User reminders: "bypasses the restraint gate — 'wake me 7' fires // docs/design.md § User reminders: "bypasses the restraint gate — 'wake me 7' fires
// in quiet hours; that's the point." Every suppressor set at once, and the // in quiet hours; that's the point." Every suppressor set at once, and the
// reminder still comes through. // reminder still comes through.
func TestRemindersBypassEverySuppressor(t *testing.T) { func TestRemindersBypassEverySuppressor(t *testing.T) {
@@ -269,7 +269,7 @@ func TestRemindersBypassEverySuppressor(t *testing.T) {
} }
} }
// GAP — DESIGN.md § User reminders ends "Snooze still applies." RemindDecisions // GAP — docs/design.md § User reminders ends "Snooze still applies." RemindDecisions
// passes every due reminder straight through with no snooze check, so a snoozed // passes every due reminder straight through with no snooze check, so a snoozed
// reminder fires anyway. The test below is what the contract asks for. // reminder fires anyway. The test below is what the contract asks for.
func TestRemindersStillHonourSnooze(t *testing.T) { func TestRemindersStillHonourSnooze(t *testing.T) {
+4 -4
View File
@@ -18,7 +18,7 @@ import (
// - does it stay quiet when it should? // - does it stay quiet when it should?
// - is it silent when the key it needs has no data at all? // - is it silent when the key it needs has no data at all?
// //
// The last one is load-bearing. DESIGN.md: "since(key)==null → don't fire. // The last one is load-bearing. docs/design.md: "since(key)==null → don't fire.
// Silence on no-data is 'shuts up when uncertain'." // Silence on no-data is 'shuts up when uncertain'."
// stateWith builds a snapshot at refTime() holding just the given facts. // stateWith builds a snapshot at refTime() holding just the given facts.
@@ -179,7 +179,7 @@ func TestCareRulePredicates(t *testing.T) {
// ---------------------------- ops rules -------------------------------------- // ---------------------------- ops rules --------------------------------------
// The two ops rules match on a value AND on which poller wrote it. DESIGN.md: // The two ops rules match on a value AND on which poller wrote it. docs/design.md:
// "a compromised poller must not be able to forge a trigger." Half of this // "a compromised poller must not be able to forge a trigger." Half of this
// table is forgery attempts; all of them must be refused. // table is forgery attempts; all of them must be refused.
func TestOpsRulePredicates(t *testing.T) { func TestOpsRulePredicates(t *testing.T) {
@@ -331,7 +331,7 @@ func TestNoDefaultRuleFiresOnEmptyState(t *testing.T) {
} }
} }
// Severities are the delivery contract (DESIGN.md § Delivery / channel // Severities are the delivery contract (docs/design.md § Delivery / channel
// routing): care is sev1-2 and drops when away, ops is sev3-4 and holds. Pin // routing): care is sev1-2 and drops when away, ops is sev3-4 and holds. Pin
// them so a change to a rule's insistence has to be deliberate. // them so a change to a rule's insistence has to be deliberate.
func TestDefaultRuleSeverities(t *testing.T) { func TestDefaultRuleSeverities(t *testing.T) {
@@ -356,7 +356,7 @@ func TestDefaultRuleSeverities(t *testing.T) {
} }
} }
// Cooldown bounds keep the feedback tuner honest — DESIGN.md wants // Cooldown bounds keep the feedback tuner honest — docs/design.md wants
// `cooldown in [min,max]` "so a weird week can't mutate Maven silent or // `cooldown in [min,max]` "so a weird week can't mutate Maven silent or
// stalker". A base outside its own envelope would make that meaningless. // stalker". A base outside its own envelope would make that meaningless.
func TestDefaultRuleCooldownsAreBounded(t *testing.T) { func TestDefaultRuleCooldownsAreBounded(t *testing.T) {
+1 -1
View File
@@ -41,7 +41,7 @@
"tags": ["preference", "homelab", "paraphrase", "hard"], "tags": ["preference", "homelab", "paraphrase", "hard"],
"query": "когда запускать резервное копирование", "query": "когда запускать резервное копирование",
"want": "n1", "want": "n1",
"note": "The DESIGN.md preference-seam example, phrased as the operator would ask it later.", "note": "The docs/design.md preference-seam example, phrased as the operator would ask it later.",
"notes": [ "notes": [
{"id": "n1", "text": "бэкапы лучше делать ночью в три часа", "kind": "note"}, {"id": "n1", "text": "бэкапы лучше делать ночью в три часа", "kind": "note"},
{"id": "n2", "text": "обновления ставлю по субботам", "kind": "note"}, {"id": "n2", "text": "обновления ставлю по субботам", "kind": "note"},
+1 -1
View File
@@ -1,5 +1,5 @@
// Package morning is maven's morning routine engine — item #3 off the // Package morning is maven's morning routine engine — item #3 off the
// 2026-07-20 backlog (see Maven/20-07-2026-BACKLOG.md). // 2026-07-20 backlog (Vikunja #280).
// //
// A Routine is NOT four independent reminder timers. It's a checklist for a // A Routine is NOT four independent reminder timers. It's a checklist for a
// daily window: several Items, each evidenced by a fact key, completed in // daily window: several Items, each evidenced by a fact key, completed in
+3 -3
View File
@@ -15,7 +15,7 @@ const (
CheckLang = "lang" // the operator's language, not the prompt's CheckLang = "lang" // the operator's language, not the prompt's
CheckLength = "length" // a nudge is one sentence, not a paragraph CheckLength = "length" // a nudge is one sentence, not a paragraph
CheckFeminine = "feminine" // her self-reference is feminine (hard constraint) CheckFeminine = "feminine" // her self-reference is feminine (hard constraint)
CheckCringe = "cringe" // DESIGN.md § Non-goals, "not a relationship" CheckCringe = "cringe" // docs/design.md § Non-goals, "not a relationship"
CheckOnTopic = "ontopic" // says the thing the rule is about CheckOnTopic = "ontopic" // says the thing the rule is about
// CheckHisGender — the other half of the persona rule: SHE is feminine, HE // CheckHisGender — the other half of the persona rule: SHE is feminine, HE
@@ -105,7 +105,7 @@ func checkLength(body string) Result {
// --- feminine self-reference --------------------------------------------- // --- feminine self-reference ---------------------------------------------
// //
// The hard constraint (CLAUDE.md, DESIGN.md § Identity): Maven's Russian // The hard constraint (CLAUDE.md, docs/design.md § Identity): Maven's Russian
// self-reference is feminine. The operator is male, so second-person forms // self-reference is feminine. The operator is male, so second-person forms
// addressed to him are MASCULINE and must not be flagged — "ты не пил воду" is // addressed to him are MASCULINE and must not be flagged — "ты не пил воду" is
// correct, "я напомнил" is not. Both directions matter, which is why this is a // correct, "я напомнил" is not. Both directions matter, which is why this is a
@@ -497,7 +497,7 @@ func isLatinWord(w string) bool {
// --- the cringe checks --------------------------------------------------- // --- the cringe checks ---------------------------------------------------
// //
// "Think Jarvis without the cringe part". DESIGN.md § Non-goals: "Not a // "Think Jarvis without the cringe part". docs/design.md § Non-goals: "Not a
// relationship — mom-tone is a function that makes nudges land, not emotional // relationship — mom-tone is a function that makes nudges land, not emotional
// company. Names the drift a warm small model falls into." Each pattern below // company. Names the drift a warm small model falls into." Each pattern below
// is one shape of that drift. They are deliberately specific: a check that // is one shape of that drift. They are deliberately specific: a check that
+2 -2
View File
@@ -11,7 +11,7 @@
// length test that a human can read and disagree with. A score here is a claim // length test that a human can read and disagree with. A score here is a claim
// about measurable properties, not about whether a sentence is good. // about measurable properties, not about whether a sentence is good.
// //
// DESIGN.md § "Rules decide, LLM phrases" is why there is no send/veto signal // docs/design.md § "Rules decide, LLM phrases" is why there is no send/veto signal
// anywhere in this package: the rule already decided she speaks. The phraser // anywhere in this package: the rule already decided she speaks. The phraser
// only words it, so a nudge the model refuses to write is a failure, never a // only words it, so a nudge the model refuses to write is a failure, never a
// legitimate outcome. // legitimate outcome.
@@ -38,7 +38,7 @@ var fixtureJSON []byte
const SchemaVersion = 1 const SchemaVersion = 1
// Case — one nudge situation, as a real tick would present it. The fields are // Case — one nudge situation, as a real tick would present it. The fields are
// the (rule, severity, context) input DESIGN.md names, flattened to JSON. // the (rule, severity, context) input docs/design.md names, flattened to JSON.
// //
// WantAny is the on-topic contract: at least one of these lowercased fragments // WantAny is the on-topic contract: at least one of these lowercased fragments
// must appear in the message. A water nudge that never mentions water is a // must appear in the message. A water nudge that never mentions water is a
+1 -1
View File
@@ -692,7 +692,7 @@ func (p *LLMPhraser) chatWithSystem(ctx context.Context, system, user string, ma
// Written as filled-in examples, not as a schema with "..." in it. A 0.8B // Written as filled-in examples, not as a schema with "..." in it. A 0.8B
// copies whatever sits in the response slot, so a literal placeholder there // copies whatever sits in the response slot, so a literal placeholder there
// teaches it to answer with the placeholder. Measured: 7/15 nudges came back // teaches it to answer with the placeholder. Measured: 7/15 nudges came back
// as "..." before this. See PHRASING-EVAL-31-07-2026.md. // as "..." before this. See docs/evals/2026-07-31-phrasing.md.
// //
// Russian only, feminine self-reference, second person masculine (the owner is // Russian only, feminine self-reference, second person masculine (the owner is
// a man). She talks TO him, informally, singular — never "вы", never "он". // a man). She talks TO him, informally, singular — never "вы", never "он".
+2 -2
View File
@@ -1,9 +1,9 @@
// Package phraser is maven's "rules decide, llm phrases" seam — the layer // Package phraser is maven's "rules decide, llm phrases" seam — the layer
// that turns a loop decision into the body + summary the delivery module ships. // that turns a loop decision into the body + summary the delivery module ships.
// //
// Per DESIGN.md § Resident language model: the phraser is the resident model // Per docs/design.md § Resident language model: the phraser is the resident model
// (Qwen3-1.7B — RU continued pretraining plus joint persona/router SFT, not a // (Qwen3-1.7B — RU continued pretraining plus joint persona/router SFT, not a
// sub-1b prompted-only model as the retired spec claimed; see DESIGN.md // sub-1b prompted-only model as the retired spec claimed; see docs/design.md
// § Superseded, "small-model phrasing claim"). It takes // § Superseded, "small-model phrasing claim"). It takes
// (rule, severity, context) and produces Body (full voice message, local — no // (rule, severity, context) and produces Body (full voice message, local — no
// shoulder-surf concern beyond who's in the room) + Summary (minimal body for // shoulder-surf concern beyond who's in the room) + Summary (minimal body for
+1 -1
View File
@@ -39,7 +39,7 @@ import (
// Re-measured with everything else held equal, thinking off scores exactly the // Re-measured with everything else held equal, thinking off scores exactly the
// same, case for case — and a direct probe shows this llama-server build ignores // same, case for case — and a direct probe shows this llama-server build ignores
// enable_thinking / reasoning_budget for this model anyway, so there was nothing // enable_thinking / reasoning_budget for this model anyway, so there was nothing
// to turn off. Full write-up in ROUTING-EVAL-31-07-2026.md (Vikunja #376). // to turn off. Full write-up in docs/evals/2026-07-31-routing.md (Vikunja #376).
func TestLLMRouterBaseline(t *testing.T) { func TestLLMRouterBaseline(t *testing.T) {
base := os.Getenv("MAVEN_LLM_URL") base := os.Getenv("MAVEN_LLM_URL")
if base == "" { if base == "" {
+3 -3
View File
@@ -1,13 +1,13 @@
// Package router is maven's reactive path — the cascade that turns a free-form // Package router is maven's reactive path — the cascade that turns a free-form
// utterance into a deterministic Decision. // utterance into a deterministic Decision.
// //
// Spec contract (from DESIGN.md § Reactive path — routing): // Spec contract (from docs/design.md § Reactive path — routing):
// //
// - the TARGET design is LLM-as-router: the resident model (Qwen3-1.7B) // - the TARGET design is LLM-as-router: the resident model (Qwen3-1.7B)
// emits GBNF-constrained structured JSON for the route, and the same // emits GBNF-constrained structured JSON for the route, and the same
// model phrases replies; the embedder is a RAG hint, not a routing gate. // model phrases replies; the embedder is a RAG hint, not a routing gate.
// the classifier/embedder cascade below is the committed default today, // the classifier/embedder cascade below is the committed default today,
// but it is an interim stopgap (DESIGN.md § Superseded, "classifier-owns- // but it is an interim stopgap (docs/design.md § Superseded, "classifier-owns-
// the-route") and the known cause of weak RU query handling — not a // the-route") and the known cause of weak RU query handling — not a
// design to extend. // design to extend.
// - a CASCADE, not one decider — layers: // - a CASCADE, not one decider — layers:
@@ -35,7 +35,7 @@ package router
import "time" import "time"
// Intent — the seven save-where labels from DESIGN.md's routing table. The // Intent — the seven save-where labels from docs/design.md's routing table. The
// discriminator is "does the loop evaluate a predicate against it?": // discriminator is "does the loop evaluate a predicate against it?":
// //
// - act: command now, not stored (function call into the allowlist) // - act: command now, not stored (function call into the allowlist)
+2 -2
View File
@@ -18,9 +18,9 @@
// returns a canned string the router + action path operate on); with a // returns a canned string the router + action path operate on); with a
// worker socket configured, it wires Remote. // worker socket configured, it wires Remote.
// //
// Per DESIGN.md § Voice pipeline (STT / TTS): whisper.cpp (CGo, Vulkan) in // Per docs/design.md § Voice pipeline (STT / TTS): whisper.cpp (CGo, Vulkan) in
// cmd/mavsttd is the production stt — the older faster-whisper/vosk picks are // cmd/mavsttd is the production stt — the older faster-whisper/vosk picks are
// retired (DESIGN.md § Superseded, "named STT/TTS model picks"). The // retired (docs/design.md § Superseded, "named STT/TTS model picks"). The
// server-side stt module is the heavy multilingual path; the client's // server-side stt module is the heavy multilingual path; the client's
// wake-word + stage-0 command grammar (cmd/mavwaked) hits the router directly // wake-word + stage-0 command grammar (cmd/mavwaked) hits the router directly
// and never crosses this seam. Today's Remote + Stub both return plain text // and never crosses this seam. Today's Remote + Stub both return plain text
+2 -2
View File
@@ -2,7 +2,7 @@
// store's allowlist, and drafts 'proposed' scaffolds for acts that aren't on // store's allowlist, and drafts 'proposed' scaffolds for acts that aren't on
// it yet. // it yet.
// //
// Boundary discipline (DESIGN.md § "Tool registration — drafting is suggest, // Boundary discipline (docs/design.md § "Tool registration — drafting is suggest,
// enabling is act"): // enabling is act"):
// //
// - The store is the allowlist. Only status='enabled' rows run. A verb not // - The store is the allowlist. Only status='enabled' rows run. A verb not
@@ -26,7 +26,7 @@
// - Destructive tools don't run on first hearing: Exec returns ErrNeedsConfirm // - Destructive tools don't run on first hearing: Exec returns ErrNeedsConfirm
// and the handler runs a confirm turn ("выполнить X? да/нет"); only a // and the handler runs a confirm turn ("выполнить X? да/нет"); only a
// confirmed re-Exec runs them. A gate assumes a fully-formed action, which // confirmed re-Exec runs them. A gate assumes a fully-formed action, which
// an enabled+matched act is (DESIGN.md § "Confirmation is not one // an enabled+matched act is (docs/design.md § "Confirmation is not one
// mechanism"). // mechanism").
package tool package tool
+2 -2
View File
@@ -5,9 +5,9 @@
// PCM, headerless per the audio package; the voice sink + reference client // PCM, headerless per the audio package; the voice sink + reference client
// wrap it in a WAV at the disk edge. // wrap it in a WAV at the disk edge.
// //
// Per DESIGN.md § Voice pipeline (STT / TTS): piper is the production tts // Per docs/design.md § Voice pipeline (STT / TTS): piper is the production tts
// (subprocess + espeak-ng, CPU-only on the ryzen box, driven by cmd/mavttsd); // (subprocess + espeak-ng, CPU-only on the ryzen box, driven by cmd/mavttsd);
// the older silero pick is retired (DESIGN.md § Superseded, "named STT/TTS // the older silero pick is retired (docs/design.md § Superseded, "named STT/TTS
// model picks"). A different voice is a model-file swap, not a code change. // model picks"). A different voice is a model-file swap, not a code change.
// The daemon wires one impl — Remote pointing at the worker socket if // The daemon wires one impl — Remote pointing at the worker socket if
// configured, Stub otherwise. // configured, Stub otherwise.