Compare commits
64 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 453919db20 | |||
| ad60e10e95 | |||
| 1528697287 | |||
| dbdab2d570 | |||
| b9371dcac6 | |||
| 62c2e92ec0 | |||
| aec94eb2e8 | |||
| 4dfe106fe3 | |||
| 2e0e2fd0bb | |||
| f3fa6b353a | |||
| 6645f64c3e | |||
| f10e0068dd | |||
| 9b124d9194 | |||
| 12530c8a95 | |||
| 51256c4c9a | |||
| 76481c2736 | |||
| bcc2305cd0 | |||
| 0ceeac8df4 | |||
| 4fae13af75 | |||
| 774217199e | |||
| 2db59d52a7 | |||
| 92d5fd580c | |||
| edeef19ff0 | |||
| 018f7a6f47 | |||
| eca41798bd | |||
| cc423567e7 | |||
| 8088ef9e00 | |||
| 666b924d29 | |||
| e52c616592 | |||
| 2b97bac51e | |||
| ab42db2b87 | |||
| 94d553570d | |||
| 2e97b905b4 | |||
| fbcca449be | |||
| 2076e4a788 | |||
| 30eb6add1b | |||
| dc266056d1 | |||
| 1c786b7156 | |||
| a3af10a830 | |||
| c0de473382 | |||
| 3e534340bf | |||
| 1a704d704d | |||
| e57adcb001 | |||
| bec7362b7b | |||
| a3ec746a01 | |||
| af0eec250e | |||
| 20aa2d59c9 | |||
| 4bad90dedb | |||
| 2b8d0f74fa | |||
| af9d2133dc | |||
| a1fdfccd61 | |||
| 5c05163266 | |||
| 92d2629001 | |||
| bdcfccce77 | |||
| f4deccacc9 | |||
| 8aaac01de6 | |||
| feb6f2c03d | |||
| 99bb3526db | |||
| bb8cb8d014 | |||
| 5e66aa8f22 | |||
| e332f167b2 | |||
| 322401b9af | |||
| 4f34a232d4 | |||
| 93987f2dfc |
@@ -9,6 +9,7 @@
|
||||
/mavwaked
|
||||
/mavmaild
|
||||
/mavupdate
|
||||
/mavgpud
|
||||
|
||||
# Certs (private keys, don't commit)
|
||||
certs/
|
||||
@@ -40,6 +41,9 @@ deploy/telegram.env
|
||||
deploy/zenmoney.token
|
||||
# IMAP password, read by mavmaild (never in argv, never committed)
|
||||
deploy/imap.password
|
||||
# Compose interpolation secrets — MAVEN_AMBIENT_TOKEN today. docker compose
|
||||
# reads this file itself; it is not an env_file on any service.
|
||||
/.env
|
||||
|
||||
# Temp files
|
||||
/tmp/
|
||||
@@ -63,3 +67,6 @@ coverage.out
|
||||
/HANDOFF.md
|
||||
/models/stt
|
||||
/models/tts
|
||||
|
||||
# root .env — MAVEN_AMBIENT_TOKEN and friends, same class as deploy/telegram.env
|
||||
.env
|
||||
|
||||
@@ -1,396 +0,0 @@
|
||||
beyond the model and tts work, the useful additions are mostly around **reliability, context, and reach**, not more intelligence.
|
||||
|
||||
## highest-value additions
|
||||
|
||||
### 1. unified event intake
|
||||
|
||||
maven should receive normalized events from:
|
||||
|
||||
* praxis
|
||||
* calendar
|
||||
* telegram
|
||||
* local notifications
|
||||
* system/service health
|
||||
* manual checklists
|
||||
* eventually email bridges
|
||||
|
||||
one internal envelope:
|
||||
|
||||
```go
|
||||
type Event struct {
|
||||
Source string
|
||||
Kind string
|
||||
EntityIDs []string
|
||||
Title string
|
||||
Body string
|
||||
Priority string
|
||||
OccurredAt time.Time
|
||||
Payload json.RawMessage
|
||||
}
|
||||
```
|
||||
|
||||
this gives digestion one stable input instead of source-specific logic.
|
||||
|
||||
---
|
||||
|
||||
### 2. explicit morning routine engine — **core engine done (2026-07-20)**
|
||||
|
||||
`internal/morning` — pure checklist engine, mirrors `internal/loop`/
|
||||
`internal/routine`'s no-I/O contract. `Evaluate(routine, facts, now)` answers
|
||||
"what's still missing" any time (order-independent — checks facts, not
|
||||
sequence); `Due(routines, facts, last, now)` fires the once-per-day nag only
|
||||
at `NudgeAt` (defaults to window end) and only when something's unevidenced,
|
||||
with a `last`-map dedupe identical in shape to `routine.Due`'s cold-start/
|
||||
last-fire tracking. Evidence is just a fact timestamped inside today's
|
||||
window — manual (voice-tapped) and inferred (another daemon writing the same
|
||||
key) are indistinguishable, satisfying the manual/inferred requirement for
|
||||
free. Weekday/weekend variants are two `Routine`s with different `Weekdays`
|
||||
sets under different names. Wired into `config.MorningRoutineConfig` +
|
||||
`cmd/mavend/tick.go`'s `fireMorningRoutines` (reads only the fact keys the
|
||||
configured items reference, dispatches through the normal severity/presence
|
||||
routing table, body is literal joined item labels — not LLM-phrased, same
|
||||
no-hallucination rationale as cron routines). 13 unit tests in
|
||||
`internal/morning/morning_test.go`.
|
||||
|
||||
Added since (2026-07-20, same day): a read-only `/morning` page in mavweb —
|
||||
`ipc.CoreAPI.MorningStatus` (new wire method, mirrors `TickTrace`'s
|
||||
daemon-cache-only shape: the store adapter errors, `daemonAPI` serves it from
|
||||
a `tickLoop.morningStatus` closure) returns each routine's active/window/
|
||||
per-item done state, server-rendered same as `/trace` (no live-update loop —
|
||||
checklist state moves on minutes, not seconds).
|
||||
|
||||
Not yet done: no config wired in `deploy/mavend.json` (no morning routines
|
||||
configured on homesrv yet — add items there when the medicine/water/pets
|
||||
fact keys the phone/desktop write are settled), no voice query path for
|
||||
"what did I miss this morning" (Evaluate supports it; nothing calls it yet),
|
||||
no way to create/edit routines from the web UI — construction still means
|
||||
hand-editing config, deliberately deferred: routines are operator-declared
|
||||
config (like cron routines), and a CRUD editor would mean moving them to a
|
||||
DB table + hot-reload, a bigger change than this pass.
|
||||
|
||||
not ordinary reminders.
|
||||
|
||||
support:
|
||||
|
||||
* required morning items
|
||||
* order-independent completion
|
||||
* soft time windows
|
||||
* skipped-step detection
|
||||
* one nudge, not repeated spam
|
||||
* manual and inferred completion evidence
|
||||
* weekend/weekday variants
|
||||
|
||||
example:
|
||||
|
||||
```text
|
||||
08:00–11:00
|
||||
- medicine
|
||||
- water
|
||||
- pets
|
||||
- check praxis attention
|
||||
```
|
||||
|
||||
maven should know what is still missing, not merely fire four timers.
|
||||
|
||||
---
|
||||
|
||||
### 3. cross-device presence
|
||||
|
||||
**status (2026-07-20):** the hysteresis engine and 3 of the listed signals are
|
||||
already built and wired live: `internal/store/presence.go` (noisy-OR combiner
|
||||
+ Schmitt-trigger bucket resolve), fed by `desk_active` (workstation, via
|
||||
`scripts/desk-active.sh` posting to `/api/signal`), `page_heartbeat` (mavweb
|
||||
tab, `app.js`), and `wg_handshake` (`mavpoll` polling `wg show`) — threaded
|
||||
into the tick loop via `internal/loop/gather.go`. Not done: phone-reachable,
|
||||
homesrv-available, audio-output, and active-maven-client signals from the
|
||||
list below are still missing.
|
||||
|
||||
a small presence daemon on each trusted device:
|
||||
|
||||
* workstation active/idle
|
||||
* phone reachable
|
||||
* homesrv available
|
||||
* last keyboard/mouse activity
|
||||
* wireguard presence
|
||||
* current audio output
|
||||
* active maven client
|
||||
|
||||
mavend receives only compact state, not raw activity logs.
|
||||
|
||||
useful for:
|
||||
|
||||
* choosing delivery channel
|
||||
* suppressing voice while away
|
||||
* surfacing reminders when you return
|
||||
* knowing whether an agent result should be spoken or sent as text
|
||||
|
||||
---
|
||||
|
||||
### 4. interruption policy — **done (2026-07-20), turned out to already be built**
|
||||
|
||||
audited the existing code before writing anything new: `internal/loop.Gate`
|
||||
already answers deliver_now vs. drop (quiet-hours/cooldown/snooze/presence/
|
||||
calendar-busy), and `cmd/mavend/tick.go`'s `digestQ` + `config.DigestConfig`
|
||||
already implement queue/digest (low-severity nudges batch into one
|
||||
notification, flushed on window elapsed or max-items reached). The four
|
||||
outcomes below were already covered by these two mechanisms; nothing new to
|
||||
build for the core policy.
|
||||
|
||||
Gap that *was* real: `deploy/mavend.json` had no `digest` block, so batching
|
||||
was disabled in prod despite being fully implemented. Fixed — see the config
|
||||
change alongside this note.
|
||||
|
||||
before delivering anything, evaluate:
|
||||
|
||||
```text
|
||||
urgency
|
||||
current activity
|
||||
quiet hours
|
||||
recent nudges
|
||||
available channels
|
||||
whether already surfaced
|
||||
```
|
||||
|
||||
result:
|
||||
|
||||
```text
|
||||
deliver_now
|
||||
queue
|
||||
digest
|
||||
drop
|
||||
```
|
||||
|
||||
this prevents maven from becoming annoying once praxis and other sources start producing more data.
|
||||
|
||||
---
|
||||
|
||||
### 5. entity-aware memory — **done (2026-07-20)**
|
||||
|
||||
`03fa52d`/`9876187` (Vikunja #279): facts gain `Subject`/`EntityID`/
|
||||
`ResolutionState`; an async enrichment worker resolves free-text subjects to
|
||||
canonical Nexus entity_ids (mirrors Praxis's enrichment pattern). Ambiguous
|
||||
or unreachable Nexus never guesses — the fact stays `pending` or terminal
|
||||
`ambiguous`. Voice-tapped facts (`IntentFact`) now flow into the enrichment
|
||||
queue automatically via an optional `Subject` field on `WriteFactReq` (old
|
||||
callers unaffected).
|
||||
|
||||
Landed alongside this in the same session (not originally on this list, but
|
||||
closes the plumbing gaps the last brief flagged for Nexus/Praxis maturity):
|
||||
a typed Praxis lifecycle client (`398997f` — surface/acknowledge/resolve/
|
||||
ignore/pin; fixes the surfaced≠acknowledged gap where reading an item aloud
|
||||
left no trace), correlation-ID/version headers on the Nexus/Praxis clients
|
||||
(`b743860`), entity-scoped Praxis attention queries (`0579ef9`), a durable
|
||||
delivery outbox with begin-before-send/complete-after semantics
|
||||
(`29f23e3`+`9ff726e` — closes a duplicate-send-on-crash bug), fail-closed
|
||||
handling on ambiguous IPC mutation outcomes and Nexus/Hexis dependency
|
||||
errors (`838fde1`+`d9fa4d6`), and a reusable fake-ecosystem test harness
|
||||
with fault injection (`c932cd8`).
|
||||
|
||||
connect maven memory to nexus ids.
|
||||
|
||||
instead of:
|
||||
|
||||
```text
|
||||
key = "кошачий фонтан"
|
||||
```
|
||||
|
||||
store:
|
||||
|
||||
```text
|
||||
entity_id = ent_pet_water_fountain
|
||||
predicate = refilled_at
|
||||
value = 2026-07-19T...
|
||||
```
|
||||
|
||||
benefits:
|
||||
|
||||
* stable russian/english aliases
|
||||
* fewer duplicate facts
|
||||
* better “when did i last…” queries
|
||||
* easier routine detection
|
||||
* cleaner praxis correlation
|
||||
|
||||
---
|
||||
|
||||
### 6. bounded follow-up state
|
||||
|
||||
for short continuations:
|
||||
|
||||
* “yes”
|
||||
* “tomorrow”
|
||||
* “the second one”
|
||||
* “not that project”
|
||||
* “do it later”
|
||||
|
||||
store explicit pending state instead of relying on chat history:
|
||||
|
||||
```go
|
||||
type PendingInteraction struct {
|
||||
Kind string
|
||||
Candidates []string
|
||||
Args json.RawMessage
|
||||
ExpiresAt time.Time
|
||||
}
|
||||
```
|
||||
|
||||
this matters a lot for a 1.7b model.
|
||||
|
||||
---
|
||||
|
||||
### 7. evaluation lab — **skipped for now (2026-07-20)**
|
||||
|
||||
runs on a different machine (GPU box), and CPT is currently in progress
|
||||
there — deprioritized until the training pipeline has a checkpoint to gate.
|
||||
Not abandoned, just off the immediate list.
|
||||
|
||||
before every new checkpoint or lora deploy:
|
||||
|
||||
* routing accuracy
|
||||
* slot accuracy
|
||||
* malformed json rate
|
||||
* russian/english mixed input
|
||||
* ambiguous entity handling
|
||||
* reminder vs note vs fact
|
||||
* direct answer vs tool call
|
||||
* confirmation safety
|
||||
* phrasing quality
|
||||
* latency and ram
|
||||
|
||||
also replay real anonymized traces against old and new checkpoints.
|
||||
|
||||
this should be a hard deployment gate.
|
||||
|
||||
---
|
||||
|
||||
### 8. replayable full-system simulator
|
||||
|
||||
fake:
|
||||
|
||||
* clock
|
||||
* presence
|
||||
* caldav
|
||||
* telegram
|
||||
* praxis
|
||||
* nexus
|
||||
* hexis
|
||||
* stt
|
||||
* tts
|
||||
* llama-server
|
||||
|
||||
scenario:
|
||||
|
||||
```text
|
||||
08:30 user appears
|
||||
08:35 medicine not completed
|
||||
08:40 correx agent waits
|
||||
08:45 calendar sync stale
|
||||
08:50 user says “what did i miss?”
|
||||
```
|
||||
|
||||
assert:
|
||||
|
||||
* what tools were called
|
||||
* what was surfaced
|
||||
* what stayed unresolved
|
||||
* what maven said
|
||||
* what was not executed
|
||||
|
||||
this will save more time than another feature daemon.
|
||||
|
||||
---
|
||||
|
||||
## useful second-wave additions
|
||||
|
||||
### voice session quality
|
||||
|
||||
* barge-in
|
||||
* interrupt tts on wake word
|
||||
* partial stt display
|
||||
* confidence-aware clarification
|
||||
* retry only failed stt segment
|
||||
* per-room microphone profiles
|
||||
* noise-floor calibration
|
||||
* short response mode when speaking
|
||||
|
||||
### notification bridge framework
|
||||
|
||||
small adapters for:
|
||||
|
||||
* ntfy
|
||||
* telegram
|
||||
* matrix
|
||||
* web push
|
||||
* android notification forwarding
|
||||
* local dbus notifications
|
||||
|
||||
normalize into maven/praxis events instead of treating each as a separate feature.
|
||||
|
||||
### local knowledge ingestion
|
||||
|
||||
* markdown/docs ingestion
|
||||
* git repo summaries
|
||||
* project decision records
|
||||
* conversation exports
|
||||
* provenance and source links
|
||||
* incremental reindexing
|
||||
|
||||
keep this read-only and separate from personal fact memory.
|
||||
|
||||
### service self-diagnostics
|
||||
|
||||
`maven doctor`:
|
||||
|
||||
* socket reachability
|
||||
* model health
|
||||
* stt/tts readiness
|
||||
* embedder availability
|
||||
* caldav freshness
|
||||
* telegram poll state
|
||||
* praxis/nexus/hexis reachability
|
||||
* db integrity
|
||||
* disk usage
|
||||
* recent failures
|
||||
|
||||
### config and secret management
|
||||
|
||||
* schema-validated config
|
||||
* config migration
|
||||
* secret references instead of inline values
|
||||
* dry-run validation
|
||||
* redacted config dump
|
||||
* per-daemon health config
|
||||
* startup dependency report
|
||||
|
||||
---
|
||||
|
||||
## things i would not build yet
|
||||
|
||||
* autonomous multi-step planning
|
||||
* large external reasoner
|
||||
* generic workflow engine
|
||||
* self-editing memory
|
||||
* automatic hexis actions from praxis
|
||||
* emotion simulation beyond phrasing
|
||||
* full home-assistant replacement
|
||||
* more model layers before routing is stable
|
||||
|
||||
## recommended order
|
||||
|
||||
**status as of 2026-07-20:**
|
||||
|
||||
1. ~~evaluation lab~~ — **skipped, GPU-box work, deprioritized while CPT is in progress**
|
||||
2. ~~entity-aware memory~~ — **done** (`03fa52d`/`9876187`, plus adjacent
|
||||
Nexus/Praxis plumbing hardening — see item 5 above)
|
||||
3. ~~morning routine engine~~ — **core engine done** (`internal/morning` +
|
||||
`cmd/mavend` wiring — see item 2 above; not yet configured on homesrv,
|
||||
no voice query, no web UI)
|
||||
4. interruption/delivery policy
|
||||
5. presence agents
|
||||
6. unified event intake
|
||||
7. full-system simulator
|
||||
8. notification bridges
|
||||
9. knowledge ingestion
|
||||
10. voice-session polish
|
||||
|
||||
the main goal should be: **maven reliably knows what is happening, knows what you meant, and chooses the least annoying correct response**. everything else can wait.
|
||||
|
||||
@@ -18,7 +18,7 @@ live in sibling repos next to this one.
|
||||
|
||||
Division of labour: Nexus identifies, Praxis observes, Hexis acts, Maven understands
|
||||
and coordinates. Maven is not the source of truth for any of the three. The full
|
||||
contract is `MAVEN_ECOSYSTEM_ARCHITECTURE.md`, and the constraints that bite during
|
||||
contract is `docs/ecosystem.md`, and the constraints that bite during
|
||||
implementation are summarised in `CLAUDE.md`.
|
||||
|
||||
Where things are in this repo:
|
||||
|
||||
@@ -10,7 +10,7 @@ compose passes `/dev/dri` + the render gid) — the resident model stays ≤1.7B
|
||||
**Resident model:** currently **Qwen3-1.7B** (`UD-Q4_K_XL`), stock — not yet the CPT'd one.
|
||||
It replaced Qwen3.5-0.8B on 2026-07-31 because it measured better on both fixtures we have:
|
||||
67.5% vs 59.7% intent-only on the 77-case RU routing fixture, and 20/27 vs 11-17/27 on the
|
||||
talk fixture. See `MODEL-BAKEOFF-31-07-2026.md`. It is a Thinking variant, so `n_ctx` is 4096
|
||||
talk fixture. See `docs/evals/2026-07-31-model-bakeoff.md`. It is a Thinking variant, so `n_ctx` is 4096
|
||||
— reasoning tokens need the room, and 4096 is what the scores above were measured at.
|
||||
|
||||
The **target** is still the locally CPT'd **Qwen3-1.7B** (Vikunja #122, training in flight).
|
||||
@@ -25,9 +25,24 @@ Spanish. Their strong published IFEval/BFCL numbers are English-only. Model file
|
||||
`models/llm/`, so the LFM2.5 gguf sitting there is not loaded by anything. Swapping the resident
|
||||
model is a one-line change to `phraser.model_path` in `deploy/mavend.json`.
|
||||
|
||||
See `REARCH.md` for the target architecture, `DESIGN.md` for the folded design spec, and
|
||||
See `docs/rearchitecture.md` for the target architecture, `docs/design.md` for the folded design spec, and
|
||||
`AGENTS.md` for local-preview + model-download recipes.
|
||||
|
||||
**Model work is moving to the workstation** (owner's call, 2026-08-02). homesrv cannot grow a
|
||||
GPU and the workstation has 16GB of VRAM. So the resident model, STT and TTS become preferred
|
||||
remotes with a floor on homesrv. The workstation is never assumed up. Fall back silently when
|
||||
it would only do the job better. Name the gap when the 1.7B cannot do it at all. The embedder
|
||||
stays on homesrv permanently, because it backs that floor. Read `docs/offload.md` before
|
||||
touching a daemon seam or adding a model caller. Vikunja #483 is the umbrella, #484 to #487
|
||||
are the work.
|
||||
|
||||
Both halves are wired as of 2026-08-03. Routing and replies prefer the workstation silently
|
||||
through `modelSeam`; nudge and reminder phrasing prefer it silently inside the phraser. A
|
||||
world question goes through `LLMPhraser.PhraseWorld` and names the gap when the card is not
|
||||
free — `worldGap` in `cmd/mavend/worldmodel.go`, which he hears instead of an invented
|
||||
answer. A box with no `workstation` block behaves exactly as it did before the seam: naming
|
||||
a gap requires a gap. The offload table in `docs/offload.md` says which caller is which.
|
||||
|
||||
## Build & test
|
||||
|
||||
CGO daemons (`mavend`, `mavsttd`, `mavttsd`, `mavenclient`) need the vendored toolchain
|
||||
@@ -72,7 +87,7 @@ protocol; the config in `deploy/mavend.json` (with `${VAR}` env expansion from g
|
||||
|
||||
Maven is one of four services. It owns conversation and personal memory. It does not
|
||||
own identity, operational state, or execution. Full contract in
|
||||
`MAVEN_ECOSYSTEM_ARCHITECTURE.md`.
|
||||
`docs/ecosystem.md`.
|
||||
|
||||
```text
|
||||
Nexus identifies. Praxis observes. Hexis acts. Maven understands and coordinates.
|
||||
@@ -113,7 +128,7 @@ Every cross-service call carries a correlation id minted once per action
|
||||
on in deploy** — this section used to say it was wired `nil`, which stopped being true on
|
||||
2026-07-31.
|
||||
|
||||
- **LLM router (the intended design, REARCH.md):** the resident Qwen3-1.7B (`llmrouter.go`)
|
||||
- **LLM router (the intended design, docs/rearchitecture.md):** the resident Qwen3-1.7B (`llmrouter.go`)
|
||||
emits GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is
|
||||
demoted from a routing gate to a RAG hint. Wired at `voice.go:214` via
|
||||
`pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient)`; the flag is `voice.llm_router`
|
||||
@@ -127,13 +142,23 @@ on in deploy** — this section used to say it was wired `nil`, which stopped be
|
||||
Cascade order: `stage0.go` exact-match fast-path → LLM router (when non-nil) → classifier
|
||||
fallback. Any LLM error falls through to the classifier so a turn never breaks on the model.
|
||||
|
||||
Measured on the 77-case RU fixture (`MODEL-BAKEOFF-31-07-2026.md`): the classifier scores
|
||||
36.8% full accuracy at p50 31ms; Qwen3-1.7B scores 67.5% intent-only / 72.7% through the
|
||||
cascade at p50 ≈825ms. Accuracy roughly doubled, latency is ~27× worse, and that trade was
|
||||
accepted deliberately. **The ≈2.7s figure that stood here until 2026-08-02 was contention,
|
||||
not the model.** See `ROUTING-EVAL-31-07-2026.md` line 61, which measures the LLM router at
|
||||
Measured on the 77-case RU fixture. **Re-measured 2026-08-02: the classifier scores 68.8%
|
||||
full accuracy at p50 16.6µs**, not the 36.8% at p50 31ms that stood here from
|
||||
`docs/evals/2026-07-31-model-bakeoff.md`. That older figure predates the stage 0 rules and the
|
||||
seed additions, both of which now score inside the classifier baseline. Qwen3-1.7B scores
|
||||
77.9% intent-only / 72.7% through the cascade. So the router buys about 4 points of accuracy,
|
||||
not a doubling, and the trade is worth re-arguing rather than assuming. **The ≈2.7s figure
|
||||
that stood here until 2026-08-02 was contention, not the model.** See `docs/evals/2026-07-31-routing.md` line 61, which measures the LLM router at
|
||||
p50 825ms / p95 1.2s / max 3.0s and the full cascade at p50 0.80-1.04s. Do not plan latency
|
||||
work off the bakeoff table. `Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
|
||||
work off the bakeoff table.
|
||||
|
||||
**The numbers above are the homesrv floor, not the ceiling.** With the workstation up, routing
|
||||
completes through `llm.Pair` against gemma-4-12b and scores **84.4% full / 93.5% intent-only at
|
||||
p50 329ms** — better than the resident model and about 2.5× faster (`docs/evals/2026-08-02-workstation-gemma4-12b.md`,
|
||||
Vikunja #485). The workstation is never assumed up, so both sets of numbers are live. Judge a
|
||||
routing change against the classifier and the resident model, since those are what always answer.
|
||||
|
||||
`Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
|
||||
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
|
||||
#359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with
|
||||
no allowlisted fn) feeding the same stage-3 gate the classifier path already had — see
|
||||
|
||||
@@ -16,11 +16,11 @@ PIPER_BIN := $(shell pwd)/deps/piper/piper
|
||||
PIPER_MODEL := $(shell pwd)/models/tts/ru_RU-irina-medium.onnx
|
||||
PIPER_ESPEAK := $(shell pwd)/deps/piper/espeak-ng-data
|
||||
|
||||
.PHONY: simulate stt-fixtures test-stt-golden all build build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav clean test fmt-check vet run-stt run-tts run-web download-embedder deps-go eval-router eval-recall eval-phrasing eval-models
|
||||
.PHONY: simulate stt-fixtures test-stt-golden all build build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav clean test fmt-check vet run-stt run-tts run-web download-embedder deps-go eval-router eval-recall eval-phrasing eval-models build-gpud
|
||||
|
||||
all: build
|
||||
|
||||
build: build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav build-mail build-update
|
||||
build: build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav build-mail build-update build-gpud
|
||||
|
||||
build-stt:
|
||||
CGO_CFLAGS="$(CGO_CFLAGS)" CGO_LDFLAGS="$(CGO_LDFLAGS)" LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
|
||||
@@ -59,6 +59,12 @@ build-mail:
|
||||
build-update:
|
||||
$(GO) build $(GOFLAGS) -o mavupdate ./cmd/mavupdate/
|
||||
|
||||
# mavgpud runs on the workstation, not here. It is built with the rest so a
|
||||
# broken supervisor is caught by `make build` on homesrv rather than by the
|
||||
# workstation refusing to serve. Copy the binary over, do not `make deploy` it.
|
||||
build-gpud:
|
||||
$(GO) build $(GOFLAGS) -o mavgpud ./cmd/mavgpud/
|
||||
|
||||
run-web: build-web
|
||||
./mavweb -addr :9200 -voice 127.0.0.1:9100
|
||||
|
||||
@@ -78,7 +84,7 @@ deps-go:
|
||||
done
|
||||
$(GO) version
|
||||
|
||||
# fmt-check fails if any file needs gofmt. DESIGN.md has always said `make
|
||||
# fmt-check fails if any file needs gofmt. docs/design.md has always said `make
|
||||
# test` gates on gofmt and vet; it did not, so nine files quietly drifted.
|
||||
# Run `gofmt -w` on whatever this prints.
|
||||
fmt-check:
|
||||
@@ -191,7 +197,7 @@ deps-piper:
|
||||
# multilingual-e5-small: an asymmetric retrieval model. It is trained to match
|
||||
# a short question against a longer passage, which is what note recall is.
|
||||
# The quantized file is the one we download, deploy and measure — see
|
||||
# RECALL-EVAL-31-07-2026.md.
|
||||
# docs/evals/2026-07-31-recall.md.
|
||||
EMBEDDER_DIR := $(shell pwd)/models/embedder/multilingual-e5-small
|
||||
EMBEDDER_MODEL_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/onnx/model_quantized.onnx
|
||||
EMBEDDER_TOKENIZER_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/tokenizer.json
|
||||
|
||||
-468
@@ -1,468 +0,0 @@
|
||||
## Maven — current state (updated 2026-07-20)
|
||||
|
||||
### Session 2026-07-20 — ecosystem hardening + entity-aware facts
|
||||
|
||||
Ten commits, focused on closing the Nexus/Praxis integration gaps flagged
|
||||
as "wired but immature" in the prior review, plus the entity-aware-memory
|
||||
backlog item (`20-07-2026-BACKLOG.md` item 5).
|
||||
|
||||
- **Entity-aware fact resolution (Vikunja #279)** — facts gain
|
||||
`Subject`/`EntityID`/`ResolutionState`; an async worker resolves
|
||||
free-text subjects to canonical Nexus entity_ids (mirrors Praxis's own
|
||||
enrichment pattern). Ambiguous/unreachable Nexus never guesses — stays
|
||||
`pending` or terminal `ambiguous`. Voice-tapped facts (`IntentFact`) flow
|
||||
into the queue automatically via an optional `Subject` field on
|
||||
`WriteFactReq` (old callers unaffected, no signature break).
|
||||
- **Typed Praxis lifecycle client (Vikunja #271)** — `GetItem`/`Search`/
|
||||
`Surface`/`Acknowledge`/`Resolve`/`Ignore`/`Pin`, routed through new RU/EN
|
||||
dialogue verbs. Fixes a real lifecycle-invariant bug: reading an
|
||||
attention item aloud now calls `Surface` — previously the digest path
|
||||
read items without recording that they'd been surfaced, so "Maven
|
||||
mentioned it" was indistinguishable from "never came up."
|
||||
- **Durable delivery outbox (Vikunja #270)** — `BeginDeliveryAttempt`
|
||||
before `Send`, `CompleteDeliveryAttempt` after; a stale `pending` row
|
||||
found at startup reconciles to `unknown` (never silently resent or
|
||||
dropped — same rule as Hexis's execution-timeout handling). Closes a
|
||||
crash-window duplicate-send bug. Wired into `DispatchNudge`,
|
||||
`DispatchReminder`, `RepeatUnacked`; reconciliation runs once at boot
|
||||
before the tick loop resumes.
|
||||
- **Fail-closed IPC/dependency handling (Vikunja #269, #272/#273)** —
|
||||
ambiguous mutation outcomes (frame sent, reply lost) no longer blindly
|
||||
retry; Nexus/Hexis dependency errors fail closed instead of guessing.
|
||||
- **Correlation IDs + version headers (Vikunja #273)** — the hand-rolled
|
||||
Nexus/Praxis HTTP clients now send `X-Nexus-Version`/`X-Praxis-Version`
|
||||
and thread the same correlation ID already generated in
|
||||
`executeCapability` through the whole call chain, matching the Hexis
|
||||
client's existing behavior.
|
||||
- **Entity-scoped Praxis attention queries** — callers holding a resolved
|
||||
entity_id can ask "what needs attention for this entity" directly
|
||||
instead of filtering the unscoped list client-side.
|
||||
- **Fake-ecosystem test harness with fault injection** — a reusable
|
||||
`fakeServer` (Nexus/Praxis/Hexis fixtures, runtime-toggleable
|
||||
`SetFault`, fake clock) replacing ad-hoc per-test `httptest` servers;
|
||||
covers a gap that had zero test coverage (`handlePraxisAct`) and adds a
|
||||
fault-then-recovery regression test for the fail-closed fixes above.
|
||||
- **Ops fix** — `deploy/mavend.json`'s phraser was pointed at a 4B model
|
||||
with `n_gpu_layers=99`, which OOM'd under memory pressure and left a
|
||||
zombie `llama-server` child; swapped to the 2B Qwen model matching the
|
||||
intended resident-model size.
|
||||
|
||||
Net effect: the Nexus/Praxis wiring described as "plumbing exists, thin
|
||||
compared to Maven's test depth" in the prior review is now materially
|
||||
hardened — typed clients, fail-closed error handling, durable delivery,
|
||||
and a proper fault-injection test harness are all in place. Evaluation lab
|
||||
(`20-07-2026-BACKLOG.md` item 7) is explicitly skipped for now — it runs
|
||||
on the GPU box, which is occupied by CPT. Morning routine engine (backlog
|
||||
item 3) is next up, not started.
|
||||
|
||||
---
|
||||
|
||||
> **Resolved 2026-07-30 (task #318).** The resident checkpoint is
|
||||
> **Qwen3.5-0.8B** (`Q4_K_M`), set in `deploy/mavend.json`; the **target** is
|
||||
> the locally CPT'd **Qwen3-1.7B**, still training (#122). Older model claims
|
||||
> below — the LFM references, the pipeline line, and the "swapped to the 2B
|
||||
> Qwen model" ops entry above — are historical. Read them as a log of what was
|
||||
> true at the time, not as current fact. Note also that `/mnt/hdd1/llms` is
|
||||
> bind-mounted over `models/llm/`, so the LFM2.5 gguf in the repo tree is
|
||||
> never loaded.
|
||||
|
||||
Architecture decision (as written on 2026-07-20): the target resident
|
||||
router/phraser is the locally trained Qwen3-1.7B model — still the target as
|
||||
of 2026-07-30. Older LFM references below describe the then-deployed
|
||||
historical stack, not the target checkpoint. RU CPT has a successful
|
||||
full-weight checkpoint at step 1000/8077; evaluation and Qwen3 SFT tooling are
|
||||
tracked in `docs/plans/2026-07-18-qwen3-resident-training-eval.md`.
|
||||
|
||||
Consolidated status. The reactive↔proactive core is closed and testable through
|
||||
the web PWA. The former SPEC's open items 1–7 (now `DESIGN.md` § execution ledger) are landed (protocol doc, away-channel
|
||||
fallthrough, CalDAV poller, quiet-hours schedule, tools enable/disable, note RAG,
|
||||
passkey step-up); item 8 (multi-user) is deliberately deferred — see the tail.
|
||||
The two big infra gaps from the jul5 revision are closed on `overnight-jul5`:
|
||||
**at-rest encryption** (AES-256-GCM, tmpfs working copy — not sqlcipher, see
|
||||
`internal/store/crypt.go`) and **Docker deployment** (one image, six daemon
|
||||
containers). The `overnight-jul6` session (now on `master`) closed the biggest
|
||||
*query-surface* gaps — **calendar querying, general-knowledge answers, and
|
||||
weather** — plus a populated homelab act allowlist and two pure scaffolds
|
||||
(dialogue state, long-term-memory vector store). ~15.2k LOC + ~8.5k test, 303
|
||||
tests, `-race` in `make test`.
|
||||
|
||||
### Access model
|
||||
|
||||
- **Phone** → needs the wg tunnel to reach homesrv (no homesrv DNS otherwise;
|
||||
raw IP or a DNS tweak can bypass, not the default).
|
||||
- **PC** → uses homesrv DNS, resolves the domains over local-net, **no wg needed**.
|
||||
- nginx + ufw both scope to `10.42.0.0/24` (wg) + `192.168.1.0/24` (LAN), deny all else.
|
||||
- **Surface in use now: the web PWA (`mavweb`).** Voice PTT + in-app nudges both ride it.
|
||||
|
||||
### Works end-to-end (tested)
|
||||
|
||||
- **Reactive voice:** PWA record → Whisper STT (`mavsttd`) → ONNX classifier →
|
||||
resident phraser (llama-server subprocess; Qwen3.5-0.8B as of 2026-07-30 —
|
||||
this line historically named "LFM 2.5-1.2B") → Piper TTS
|
||||
(`mavttsd`) → reply.
|
||||
HTTP POST path (mobile-Chrome drops WS for the audio).
|
||||
- **Capture:** `fact` (EN **and RU** — root-substring recognizers) + `reminder`
|
||||
persist through CoreAPI (`source=tap:voice`). This is the substrate the care
|
||||
rules read.
|
||||
- **Notes / query (semantic recall, sqlite — no chroma):** `note` → embed (the
|
||||
classifier's ONNX embedder) → `notes` table. `query` → embed → brute-force
|
||||
cosine top-k → confidence-gated (below `queryMinScore` 0.55 ⇒ "no note", not a
|
||||
guess). **Note RAG (SPEC item 6):** the gated top-k feed the phraser
|
||||
(`PhraseQuery`) to compose a natural answer ("вот что я нашла: …") instead of
|
||||
a verbatim dump; raw-notes fallback on any LLM error. Stub is deterministic.
|
||||
- **Monitoring (`/dash`):** mavweb server-renders presence + recent nudges (by
|
||||
outcome) + recent facts from the append-only store via CoreAPI. Read-only,
|
||||
meta-refresh, no JS.
|
||||
- **Proactive loop:** 60s dumb ticker, pure predicates over a State snapshot,
|
||||
universal gate (quiet-hours/presence/cooldown/snooze/calendar), one-nudge-per-
|
||||
tick max-severity, reminders (gate-bypassing), sev4 repeat-til-ack, feedback
|
||||
auto-tuner (outcome ratio → bounded cooldown, persisted as `source=feedback`).
|
||||
- **Rules:** water/meal/break (sev1–2 care), service_down (sev4, `poll:uptimekuma`),
|
||||
netdata_critical (sev3, `poll:netdata`).
|
||||
- **Routines (`internal/routine`):** operator-declared clockwork — the third
|
||||
proactive class beside reminders (user-stated) and care rules (world-state).
|
||||
Config `routines[]` (cron + literal RU body + severity) fire through the normal
|
||||
dispatcher on schedule (an 08:00 briefing, a 22:00 wind-down). Bodies are
|
||||
literal (not LLM-phrased ⇒ can't hallucinate); rule name `routine:<name>` so
|
||||
they don't pollute the care autotuner; cold-start guard seeds on first sight so
|
||||
a restart never replays a missed schedule. Pure `routine.Due`, unit-tested; the
|
||||
tick driver holds the last-fired map.
|
||||
- **Env facts (`mavpoll`):** netdata alarms → `netdata_alarm` (fires immediately
|
||||
on a real CRITICAL); kuma monitor_status → `service_down`. Writes only on
|
||||
value-change (no append-only churn).
|
||||
- **Presence:** noisy-OR decay + Schmitt hysteresis. Live via `page_heartbeat`
|
||||
(PWA auto-pings `/api/signal` every 30s → present when a tab's open).
|
||||
- **Delivery:** ntfy / telegram / voice by `f(severity, presence)`; minimal body
|
||||
on away channels. PWA subscribes to ntfy over **WebSocket** for in-app nudges.
|
||||
- **Away-channel fallthrough (SPEC item 2):** when the router picks voice but no
|
||||
live session exists at push time (presence guess was wrong), the dispatcher
|
||||
reroutes through the AWAY table — sev3→ntfy, sev4→telegram-repeat-til-ack,
|
||||
sev≤2→drop — instead of silently dropping. Covers nudges + reminders.
|
||||
- **Calendar busy (SPEC item 3, `mavcaldav`):** new poller queries a self-hosted
|
||||
**Radicale** CalDAV server on an interval, writes `calendar_busy` + event facts
|
||||
through CoreAPI (value-change only). The loop gate already consumes `calendar_busy`.
|
||||
- **Quiet-hours schedule (SPEC item 4):** the gate reads `quiet_hours`; a config
|
||||
time window (`voice.quiet_hours`, HH:MM, midnight-crossing handled) now sets it
|
||||
on each tick — in addition to the "тихий режим" voice toggle. Both activate quiet.
|
||||
- **Client protocol (SPEC item 1):** the voice wire format (length-prefixed JSON
|
||||
frames) is published in `PROTOCOL.md`, generated from `internal/voice/wire.go`
|
||||
so third-party clients don't need the Go source.
|
||||
- **Passkey step-up (SPEC item 7):** `internal/webauthn` does real WebAuthn —
|
||||
ES256/P-256 register + assert, ecdsa signature verification, rpIdHash + UP/UV
|
||||
flag binding (UV = the gesture), sign-count regression check. `PasskeySession`
|
||||
bumps the auth session L2→L3 for a TTL on assert. mavweb serves `/auth/passkey`
|
||||
(enroll + step-up) + the begin/finish endpoints. Crypto is round-trip tested
|
||||
(incl. tampered-sig / missing-UV / wrong-origin negatives).
|
||||
- **Stability:** llama-server orphan leak fixed (`Pdeathsig` kills the child on
|
||||
any mavend death); `kill-maven.sh` reaps strays (matches the model, not a
|
||||
bogus `llama-server.*maven` pattern); `start-maven.sh` wires `-core` + poller.
|
||||
|
||||
### Wired but needs a deploy action (not code)
|
||||
|
||||
- **`desk_active`** (strongest presence signal) — `scripts/desk-active.sh` runs
|
||||
on the **desk PC** (hypridle-gated systemd timer), posts over wg to mavweb.
|
||||
- **`mavwaked`** (always-on listening) — needs a systemd user unit on a client
|
||||
box (desk PC, pi, etc.) where the mic is attached. Connects to mavend over wg
|
||||
or local net via `-addr`. Deferred until a client box is wired with a mic.
|
||||
|
||||
Caveats / gotchas:
|
||||
- **desk_active is a workstation deploy, not code** — 0 facts ever written; presence
|
||||
runs on page_heartbeat alone (dash reads "away"/"never at desk"). `scripts/desk-active.sh`
|
||||
+ a hypridle-gated `maven-desk` timer must be installed on the desk PC (not homesrv).
|
||||
- **Notes recall needs the ONNX embedder** — under the HashEmbedder floor, cosine is
|
||||
lexical (token overlap), not semantic; scores are low, so most RU commands sit under
|
||||
the 0.35 route threshold and clarify. Configure `voice.embedder` for confident recall+routing.
|
||||
(The floor now at least tokenizes Cyrillic — see below — so it ranks correctly, just weakly.)
|
||||
- **Switching the embedder model silently breaks old notes** — different dim ⇒
|
||||
cosine 0 ⇒ they stop matching; brute-force can't re-embed. Re-embed on a model change.
|
||||
- **`wg_handshake` is OFF and should stay off** — in this topology the phone only
|
||||
runs wg when *outside*, so a fresh handshake means AWAY, not here. The `mavpoll
|
||||
-wg` flag exists (defaults `""`) and could later back the spec's "away override"
|
||||
by flipping the sign; as a presence-*here* signal it's inverted. desk_active +
|
||||
page_heartbeat cover home presence.
|
||||
- **Cold-start unlock tests are missing** — the key wrap/unwrap code
|
||||
(`internal/webauthn/keywrap.go`) and locked-mode IPC gating (`cmd/mavend/main.go`)
|
||||
are correct but have **zero test coverage**. The roadmap (item 2.1) required
|
||||
three new test cases (wrap/unwrap round-trip, wrong-cred unwrap fails,
|
||||
locked-mode IPC rejects non-unlock methods); none were written. `make test`
|
||||
is green by omission. Write these before relying on the cold-start path with
|
||||
real keys.
|
||||
|
||||
### Done since last revision (overnight-jul6, 2026-07-06)
|
||||
|
||||
Seven tasks (session board `SESSION-06-07-2026.md`, deleted 2026-07-30 — see git history), one commit each, merged to `master`.
|
||||
This session was run through **opencode**, not Claude Code (co-author trailer).
|
||||
|
||||
Since then (**2026-07-06, second session**):
|
||||
|
||||
- **Always-on listening (gap 1, MVP)** — `cmd/mavwaked/`: 825 lines, 10 `-race`
|
||||
tests. Energy-based VAD over 30ms windows (same RMS threshold as mavsttd's
|
||||
`gateReason`), adaptive noise floor, speech→silence state machine. Captures
|
||||
PCM from arecord(1) subprocess, sends `PushToTalk` with `Surface=SurfaceVoice`
|
||||
(L0 — no destructive acts). Reply plays through aplay(1). No wake word yet
|
||||
(pure VAD trigger); the 30ms frame shape matches silero-vad ONNX input 1:1,
|
||||
so swapping energy-threshold for ONNX inference is a local change in vad.go.
|
||||
`Makefile` `build-waked` target. Runs on client boxes (not docker/homesrv)
|
||||
via systemd user unit; connects to mavend over wg or local net.
|
||||
|
||||
Since then (**2026-07-06, third session** — roadmap execution agent):
|
||||
|
||||
- **Cold-start unlock (ROADMAP 2.1)** — the at-rest AES key is now wrapped
|
||||
(HKDF-SHA256 + AES-256-GCM, stdlib-only — no `x/crypto` dep) with the passkey
|
||||
credential's public key and persisted to disk. At boot, if a wrapped key file
|
||||
exists AND no env key is set, mavend starts **locked**: the IPC server runs
|
||||
but `srv.Check` rejects everything except `MethodAssertStepUp` +
|
||||
`MethodUnlock`. A passkey assertion at `/auth/passkey` calls `MethodUnlock`
|
||||
with the credential's public key → unwraps the blob → opens the store → wires
|
||||
voice/loop/delivery → `srv.SetAPI` swaps the locked stub for the real
|
||||
CoreAPI. mavweb's `RegisterFinish` wraps the env key on enrollment;
|
||||
`AssertFinish` calls `Unlock` on assertion. Env-key fallback preserved
|
||||
(dev/CI path unchanged). **Test gap:** the roadmap required three new test
|
||||
cases (wrap/unwrap round-trip, wrong-cred unwrap fails, locked-mode IPC
|
||||
rejects non-unlock methods) — none were written. The code is correct but
|
||||
untested; `make test` is green by omission, not coverage.
|
||||
- **Conversation depth (ROADMAP 3.2)** — cross-intent anaphora + fact-by-key
|
||||
lookup. `AnaphoraResolver` in `router/slots.go` detects RU pronouns
|
||||
(это/он/она/оно/тот/мой + inflected forms). `followUpMerge` now handles
|
||||
three cases: same-intent slot inheritance (existing), cross-intent anaphora
|
||||
(Query/Fact/Reminder after a Fact with a pronoun inherits the prior key +
|
||||
time), and query-after-fact (a query following a fact inherits the key for
|
||||
fact-by-key lookup). `Session.History []Turn` added as the multi-turn
|
||||
scaffold (capped at 4). 7 new test cases including the exact done-when
|
||||
scenarios (anaphora query-after-fact, three-turn break, explicit-key-wins).
|
||||
- **Routing quality + persona (ROADMAP 4.1/4.4)** — `QueryMinScore` is now a
|
||||
config knob (`voice.query_min_score`, default 0.55) instead of a hardcoded
|
||||
const. `make download-embedder` fetches Xenova/paraphrase-multilingual-
|
||||
MiniLM-L12-v2 (~90MB ONNX) + tokenizer; AGENTS.md documents the embedder +
|
||||
libonnxruntime setup. `Persona` field in `VoiceConfig` prepends to every
|
||||
LLM system prompt (nudge phrasing, note queries, general knowledge); empty
|
||||
= current hardcoded feminine-gendered Russian persona. Also fixed two
|
||||
pre-existing data races found by `-race`: `voice/server.go` wg.Add vs
|
||||
wg.Wait (accept mutex), `mavweb/server.go` s.api field (atomic.Value).
|
||||
|
||||
- **Calendar querying (task 3)** — "что у меня завтра?" now answers from the
|
||||
CalDAV facts the poller already writes. Added `store.CalendarEvents(from,to)`,
|
||||
a RU date-scope parser («сегодня»/«завтра») in `router/slots.go`, and an
|
||||
IPC `CalendarEvents` RPC (api/client/server/wire) feeding the `IntentQuery`
|
||||
handler. Empty day → «на сегодня ничего нет». Previously calendar only *gated*
|
||||
nudges; it's now queryable.
|
||||
- **General-knowledge routing (task 4)** — when notes-RAG misses `queryMinScore`,
|
||||
the query now falls through to the phraser with an anti-hallucination system
|
||||
prompt (`router.KnowledgePrompt`, single tested source) instead of giving up.
|
||||
Empty/errored/Stub phraser → «не знаю.», never a fabrication.
|
||||
- **Weather (task 5)** — new `internal/weather/`: `Provider` interface, a stub
|
||||
(«погода не настроена»), and a real **keyless Open-Meteo** provider (geocode +
|
||||
current_weather, injectable `*http.Client`, mocked in tests — no live network).
|
||||
Wired into `IntentQuery` (keywords погода/градус/температура) with a ~5s
|
||||
context timeout; selected by `voice.weather.provider` ("open-meteo" | "" → stub).
|
||||
- **Homelab act allowlist (task 2)** — `voice.tools` seeded with read-only acts
|
||||
(`systemctl status`, `docker ps`, `uptime`, `df`, `free`, `journalctl` reads)
|
||||
as `destructive:false` and mutating ones (restart/stop/start/reboot,
|
||||
docker-restart/stop) as `destructive:true`. Guardrail verified: no dangerous
|
||||
verb is `destructive:false`. RU phrasings seeded in `act.txt`.
|
||||
- **Embedder config validation (task 1)** — a partially-filled `voice.embedder`
|
||||
block (some of model/tokenizer/lib paths missing) is now a load error instead
|
||||
of a silent fall-through to the Hash floor; the floor fallback logs explicitly.
|
||||
- **Dialogue state scaffold (task 6)** — `internal/dialogue/`: `Session` +
|
||||
TTL `SessionStore` + pure `InheritSlots`. **Now wired** (post-merge follow-up):
|
||||
the voice handler carries slots across same-intent turns within a 2-min window
|
||||
(`followUpMerge`, unit-tested) — bounded gap-filling, not full multi-turn yet.
|
||||
- **Long-term memory interface (task 7)** — `internal/memory/`: `Store` interface
|
||||
+ `InMemoryStore` (cosine). Wired into `IntentNote` (best-effort insert) and,
|
||||
post-merge, into `IntentFact` (facts indexed) + `IntentQuery` (read-back after
|
||||
notes-RAG misses). In-memory only — no persistent backend yet (gap #8).
|
||||
|
||||
Follow-ups (Claude Code, post-merge): gofmt'd `handlers_test.go` (the jul6
|
||||
verification commit left it misaligned, so `gofmt -l` still flagged it despite the
|
||||
"all gates green" claim); deduped the task-4 knowledge prompt to the single tested
|
||||
`router.KnowledgePrompt()`. Tree is now genuinely green (gofmt/vet/303 tests).
|
||||
|
||||
### Done since the jul5 revision (overnight-jul5, 2026-07-05)
|
||||
|
||||
The overnight session (`SESSION-05-07-2026.md`, deleted 2026-07-30 — see git history; 25 tasks) closed the previous
|
||||
"not built yet" items 1–3 and added feature depth:
|
||||
|
||||
- **At-rest encryption** — the on-disk db is AES-256-GCM ciphertext; the daemon
|
||||
works on a tmpfs (RAM) plaintext copy, sealed back atomically on close. Wrong
|
||||
key / tamper ⇒ fail closed, never a plaintext fallback. Legacy plaintext dbs
|
||||
upgrade on first clean shutdown. Key via config/env (`db_key_env`); no KDF —
|
||||
raw 32-byte key, base64. The passkey cold-start unlock plugs into the same
|
||||
`store.OpenEncrypted` seam later.
|
||||
- **Docker deployment** — single image, one container per daemon
|
||||
(`docker-compose.yml`); only mavend mounts the key + db volume; IPC over a
|
||||
shared socket volume. `ipc.DialWait` (boot-order tolerance) + redial-on-drop
|
||||
(core restarts don't kill modules). `deploy/README.md` has the runbook.
|
||||
- **Tests** — mavcaldav, mavttsd, voicesink, mavweb main/handlers covered;
|
||||
`make test` runs `-race -coverprofile`.
|
||||
- **Recurring reminders** — `cron` + `next_fire_ts` on reminders; recurring ones
|
||||
reschedule (instead of mark-fired) after successful delivery.
|
||||
- **Notification digest/batching** — low-severity nudges queue and flush as one
|
||||
digest per window/max-items (`digest` config block); stale-reminder bursts on
|
||||
boot collapse into a single digest reminder, completed only after delivery.
|
||||
- **Rule trace engine** — `ExplainTick`/`ExplainGate` record per-rule
|
||||
predicate/gate/selection results each tick; served over IPC (`tick_trace`)
|
||||
and rendered at mavweb `/trace` ("why didn't she nudge me").
|
||||
- **Web UI** — new `/history` (facts + revert buttons), `/notifications` (nudge
|
||||
history), `/trace` pages; nav links on `/dash`; RU/EN cheatsheet toggle in the
|
||||
PWA; manifest icons (`icon.svg`). POST `/tools` now requires an in-process
|
||||
passkey step-up when WebAuthn is configured.
|
||||
- **Revert/undo** — `RevertFact` voids the latest fact for a key (append-only
|
||||
void-marker, audit trail intact); exposed at `/api/revert` from `/history`.
|
||||
- **Tool scopes** — `scope` column on tools, threaded through propose/enable/UI.
|
||||
`DisableTool` raised to AuthStepUp alongside Enable.
|
||||
- **Passkey persistence** — mavweb credentials in a JSON file (`-passkey-file`),
|
||||
surviving restarts; rollback-on-persist-failure keeps memory and disk in sync.
|
||||
- **STT silence gate** — min-duration + RMS floor drop non-speech before whisper
|
||||
hallucinates on it (`-min-ms`, `-silence-rms` flags on mavsttd).
|
||||
- **Housekeeping** — `db_key.env` gitignored (+`.env.example`), `build-caldav`
|
||||
target, zero-timestamp "never" fix on /dash.
|
||||
|
||||
### Not built yet (ranked by ROI)
|
||||
|
||||
1. **Multi-user (SPEC item 8)** — deliberately deferred, see the tail.
|
||||
|
||||
Closed (jul6 follow-ups): `/api/revert` now sits behind the same passkey
|
||||
step-up as POST `/tools`; `go.mod` direct deps (`onnxruntime_go`,
|
||||
`coder/websocket`, `robfig/cron`) are labeled correctly — `go mod tidy` can't
|
||||
run here because it walks the vendored `deps/go` toolchain tree.
|
||||
Purge+rotate leaked db key (#12) — investigated and closed: the key was
|
||||
**never committed** to git history (gitignored at introduction, no commit
|
||||
ever tracked `deploy/db_key.env`), so nothing to scrub. File stays on disk
|
||||
and in deploy env by design — at-rest encryption needs it at boot.
|
||||
|
||||
Done earlier (2026-07-03): **act tool executor, store-backed, full flow**
|
||||
(`internal/tool` + `internal/store/tools.go` + `tools` CoreAPI methods).
|
||||
- **Execution:** IntentAct runs the matched fn against the store's ENABLED
|
||||
allowlist. argv, no shell → STT text can't inject. Live store read, so a
|
||||
newly-enabled tool runs without a daemon restart.
|
||||
- **proposed→enabled→disabled (SPEC item 5):** an act whose verb isn't enabled is
|
||||
scaffolded as a `proposed` tool (maven suggests). A human enables it (fills argv
|
||||
+ destructive) on the authed **`mavweb /tools`** page — never voice — and can
|
||||
disable it back to `proposed` (kept in the store, won't run). `EnableTool`/
|
||||
`DisableTool` sit at `AuthStepUp`; the gate is now **live** via `PasskeySession`,
|
||||
so /tools enable requires a passkey assertion at `/auth/passkey` first.
|
||||
- **Confirm turn:** a destructive enabled tool replies "выполнить X? да/нет" and
|
||||
parks; the next utterance (ru/en yes-no) confirms or cancels (90s TTL).
|
||||
- **Config:** `voice.tools` seeds enabled tools at boot (editing mavend.json =
|
||||
the human enable act); mavweb enables ad-hoc ones on top.
|
||||
- **Russian:** fixed grammar in reply strings + seed files; maven's self-
|
||||
reference is feminine ("she") — [[maven-persona-gender]].
|
||||
|
||||
Also fixed:
|
||||
- **HashEmbedder was blind to Cyrillic** (`tokenize` iterated bytes, kept only
|
||||
`a-z0-9`) → every RU utterance embedded to the zero vector → cosine 0 across
|
||||
all intents → misrouted to `act` (alphabetical tie-break). Now rune-based
|
||||
(`unicode.IsLetter`). This was the real cause of "Найди заметку" (a query)
|
||||
landing in `notes`; added note-retrieval query seeds too.
|
||||
- **Notes are now browsable on `/dash`** — `RecentNotes` plumbed through the
|
||||
store + CoreAPI; voice-captured notes were previously only reachable via
|
||||
semantic `query`.
|
||||
Earlier: notes/query recall, `/dash` monitoring, `wg_handshake` poller (NO-OP).
|
||||
|
||||
### Gaps — why "voice assistant" is still aspirational (2026-07-06)
|
||||
|
||||
What separates Maven today from the thing the spec describes. Dealbreakers
|
||||
first — these define the category:
|
||||
|
||||
1. **Always-on listening is code-complete (MVP).** `cmd/mavwaked` captures
|
||||
PCM from arecord → energy-based VAD → PushToTalk with `Surface=SurfaceVoice`
|
||||
(L0). Gap narrowed: no wake word yet (pure voice-activity trigger; every
|
||||
utterance fires). The 30ms frame shape and 16kHz PCM match silero-vad's
|
||||
ONNX input exactly, so a wake-word model swap is a local change in vad.go.
|
||||
Hardware: the mic lives on a client box (desk PC, pi, etc.) — never the
|
||||
homesrv. Deploy action: systemd user unit on whichever box has the mic,
|
||||
connects to mavend over wg or local net.
|
||||
2. **Conversation is deeper now, still not full dialogue.** The router
|
||||
classifies one utterance → one reply, but `internal/dialogue` carries
|
||||
context across turns: a 2-min session inherits slots for same-intent
|
||||
follow-ups («напомни завтра» → «…позвонить маме»), and cross-intent
|
||||
anaphora («запиши что я пил воду» → «когда я это сделал?») now resolves
|
||||
RU pronouns (это/он/она/оно/тот/мой + inflections) to the prior turn's
|
||||
key for fact-by-key lookup. `Session.History []Turn` is the scaffold for
|
||||
real multi-turn. Still missing: LLM-driven dialogue manager (decide
|
||||
ask-vs-act), anaphora beyond RU pronouns, single-slot session (single-user
|
||||
box). The sub-1B phraser only words replies.
|
||||
3. **Latency/shape of a turn.** Clip-based STT (record → upload → whisper →
|
||||
route → phrase → piper → play). No streaming either direction, no barge-in;
|
||||
every exchange is a full round trip.
|
||||
|
||||
Capability-class gaps — built but thin:
|
||||
|
||||
4. **Act surface is a small argv allowlist.** propose→enable works and the
|
||||
allowlist now ships a homelab starter set (jul6 task 2 — status/ps/uptime/
|
||||
df/free/logs read-only, restart/stop/reboot gated). Still bounded to what's
|
||||
seeded; broadening it is config, not code.
|
||||
5. **Query answers now cover notes + calendar + weather + general knowledge**
|
||||
(jul6 tasks 3/4/5). Calendar querying, keyless Open-Meteo weather, and a
|
||||
phraser knowledge-fallback all landed; caveat — general-knowledge quality is
|
||||
only as good as the sub-1B phraser, and weather needs `voice.weather.provider`
|
||||
set. The cheatsheet and router are now roughly aligned.
|
||||
6. **Routing quality depends on the ONNX embedder being configured** — the
|
||||
HashEmbedder floor makes RU recall lexical/weak; many commands fall to
|
||||
"clarify". `make download-embedder` now fetches the multilingual MiniLM
|
||||
model + AGENTS.md documents libonnxruntime setup; `voice.query_min_score`
|
||||
is a config knob (default 0.55) so the floor can be tuned without recompile.
|
||||
7. **Presence is effectively one signal** (page_heartbeat); desk_active is
|
||||
still an undeployed script — "voice when near" routing runs on a guess.
|
||||
8. **Long-term memory is now persistent (store-backed), not the spec's chroma.**
|
||||
`internal/memory` has a `Store` interface; the daemon now wires
|
||||
`store.MemoryStore` (`internal/store/memory.go`) — a **persistent** backend
|
||||
in the **same encrypted sqlite db** (survives restarts; recall text inherits
|
||||
at-rest encryption, so no plaintext sidecar). Vectors are float32 blobs,
|
||||
search is brute-force cosine (fine at single-user scale; ANN is the later
|
||||
swap behind the same interface). Notes **and facts** are indexed on capture;
|
||||
`IntentQuery` reads it back (after notes-RAG misses, before general-knowledge)
|
||||
— fact recall («когда я пил воду?») is its distinct payoff. The in-memory
|
||||
impl remains the test/no-store floor. Remaining: an ANN/external index is
|
||||
optional-scale, not a gap. Custom TTS voice (kami-picked, replaces the irina
|
||||
floor — [[custom-voice-training]]) is still a future item.
|
||||
|
||||
Ops footnote: voice-over-web verified 2026-07-06 — mavend binds 0.0.0.0:9100
|
||||
and mavweb reaches it cross-container at mavend:9100 (nc -z confirmed).
|
||||
mavpoll uses network_mode=host to reach localhost services (netdata, kuma).
|
||||
|
||||
### Future / logged, not now
|
||||
|
||||
Custom TTS voice training (kami-picked voice, replaces irina floor); listening
|
||||
modes 2–3 (meeting-record, ambient-derive).
|
||||
|
||||
### Services & layout
|
||||
|
||||
- `mavend` (core, IPC unix socket) — store + loop + phraser; the only key-holder.
|
||||
- `mavsttd` / `mavttsd` — STT/TTS worker modules (unix sockets).
|
||||
- `mavweb` — PWA bridge (HTTP), `/api/ptt` voice, `/api/signal` presence ingest,
|
||||
`/api/ntfy` WS-subscribe config, `/dash` read-only monitoring.
|
||||
- `mavpoll` — env poller (netdata/kuma → facts via CoreAPI).
|
||||
- `mavcaldav` — CalDAV poller (Radicale → `calendar_busy` + events via CoreAPI).
|
||||
- All behind wg + nginx deny-all; no phone-home. CGo only in `mavsttd`.
|
||||
- Start/stop: `./start-maven.sh [build]`, `./kill-maven.sh`.
|
||||
- Config: `~/.config/maven/mavend.json` (or `mavend.json` in repo root).
|
||||
|
||||
### Key files
|
||||
|
||||
- `cmd/mavend/{main,tick,voice}.go` — daemon wiring, loop driver, voice handler
|
||||
- `internal/loop/{loop,rules,gather,feedback}.go` — proactive engine
|
||||
- `internal/store/` — append-only facts/reminders/nudges/presence/notes
|
||||
- `cmd/mavweb/{main.go,dash.html}` — PWA bridge + `/dash` monitoring
|
||||
- `internal/router/{classifier,slots,stage0}.go` — reactive routing + slot parse
|
||||
- `internal/delivery/` — dispatcher + ntfy/telegram/voice sinks
|
||||
- `internal/auth/` — scope/gate/policy; `FloorEnrollment` (same-uid = device
|
||||
trust) + `webauthn.PasskeySession` (real step-up for L3)
|
||||
- `internal/webauthn/`, `cmd/mavweb/webauthn.go` — passkey register/assert
|
||||
- `cmd/mavcaldav/`, `cmd/mavpoll/`, `scripts/desk-active.sh` — env producers
|
||||
|
||||
### Why multi-user (SPEC item 8) is deferred
|
||||
|
||||
Not neglect — the one item where doing nothing now beats doing something:
|
||||
|
||||
- **No second user exists yet** (the "gf phase"). Building per-user partitioning
|
||||
now means code exercised by zero users and validated by nobody — YAGNI.
|
||||
- **The append-only schema makes it a migration, not a rewrite.** No row is ever
|
||||
mutated, so adding `facts/notes/reminders.user_id` later is add-columns +
|
||||
backfill-to-"kami" — no reshaping, no dual-write window. Deferral is cheap.
|
||||
- **The hard part is speaker attribution, and it needs the second voice.** A
|
||||
voice-print discriminator (kami vs gf vs unknown) can't be trained or tuned
|
||||
with one voice in the house. Plumbing before the model is pipe with no water.
|
||||
- **It's fenced deliberately** (`DO NOT TOUCH THIS PHASE` in `DESIGN.md` § Users) so an
|
||||
autonomous agent doesn't add `user_id` columns while touching the store and
|
||||
commit us to a schema before the constraints that shape it exist.
|
||||
-184
@@ -1,184 +0,0 @@
|
||||
# QA plan: checking Maven properly
|
||||
|
||||
Written 2026-08-01, after the 35-PR stack landed and the box came back up.
|
||||
|
||||
44 of the 50 open Vikunja tasks are `QA:` tasks. They are verification work, not
|
||||
build work. Most sat unverifiable while Maven was down for 11 days. That
|
||||
blocker is gone.
|
||||
|
||||
This plan orders them by what unblocks what. Do sessions 1 and 2 first. Almost everything
|
||||
downstream assumes the voice loop works, and nobody has confirmed that since
|
||||
the redeploy.
|
||||
|
||||
---
|
||||
|
||||
## Before you start
|
||||
|
||||
Two things bite anyone running these checks on homesrv.
|
||||
|
||||
**curl needs `--noproxy '*'`.** The shell exports `http_proxy=http://127.0.0.1:18080`.
|
||||
Without the flag, every local check returns 503 from the proxy and looks like a
|
||||
dead service. This cost me a false regression report today.
|
||||
|
||||
**The database is not readable with sqlite3.** Four older QA steps say
|
||||
`docker compose exec mavend sqlite3 /data/maven.db "select ..."`. That cannot
|
||||
work: the container has no `sqlite3` binary, and the store is AES-256-GCM at
|
||||
rest with a tmpfs working copy. Read state through mavweb instead, at
|
||||
`/history`, `/trace`, `/routines` and `/dash`.
|
||||
|
||||
---
|
||||
|
||||
## Session 1: the voice loop (half a day)
|
||||
|
||||
Nothing here has been confirmed since the redeploy, and everything else assumes
|
||||
it works. Do this first.
|
||||
|
||||
Closes or advances: **44** (conversation), **45** (text chat), **287** (voice
|
||||
session quality), **321** steps 3-5 (quiet mode), **288** (STT fixtures).
|
||||
|
||||
1. Open `http://127.0.0.1:9201/chat` and hold a short conversation in Russian.
|
||||
Watch for three things: she answers in feminine forms (`рада`, `поняла`), she
|
||||
says `ты` and never `вы`, and no pet names appear.
|
||||
2. Press push-to-talk on `/dash`. Say `привет`. Confirm a spoken reply comes
|
||||
back. This is the only check that covers mic to STT to core to TTS to
|
||||
speaker as one path. It is also the path the eleven-day outage most likely
|
||||
broke.
|
||||
3. Say `тихий режим`. Expect `тихий режим включён. буду реже напоминать.`
|
||||
4. Say `выключи тихий режим`. Expect `тихий режим выключен.` Negation must win.
|
||||
5. Say `в комнате тихо`. Quiet mode must NOT flip. Confirm on `/history` that no
|
||||
`quiet_hours` fact was written.
|
||||
6. Say `включи режим тишины`, then `сделай потише`. Both must flip quiet mode
|
||||
on. These are the noun form and the comparative, added 01-08-2026.
|
||||
7. Wait for a nudge, then say `потом` within twenty minutes. Expect `хорошо,
|
||||
вернусь к этому позже.` and the nudge row on `/notifications` reading
|
||||
`snoozed`. Say `потом` again with nothing pending: it must route as an
|
||||
ordinary utterance, not be swallowed.
|
||||
8. Wait for the water nudge, then say `выпил воды`. Expect the ordinary fact
|
||||
reply and nothing extra — she must not congratulate you. Check
|
||||
`/notifications`: the row reads `acted`. Then trigger another nudge and say
|
||||
`готово`; expect `отлично, отметила.` and the same outcome.
|
||||
9. Note anything where she is slow, cuts off, or talks over herself. That is
|
||||
287's whole content and it has no written acceptance criteria yet.
|
||||
|
||||
**319 is fixed** (01-08-2026). Single-word Russian utterances no longer come
|
||||
back as `не совсем поняла — можешь переформулировать?`. `привет` and `поужинал`
|
||||
both pass now: `thinSingleToken` spares social singles and any token carrying a
|
||||
verb ending, and only thins a bare nominal like `вода`. If a one-word utterance
|
||||
still gets clarified during the smoke test, that is a new case for the lexicon,
|
||||
not the old bug.
|
||||
|
||||
---
|
||||
|
||||
## Session 2: measurement (half a day, mostly waiting)
|
||||
|
||||
Closes or advances: **320** items 2-4, **278** (make the eval lab routine),
|
||||
**319** (gate recalibration).
|
||||
|
||||
The resident llama-server cannot be reached by the eval harness. It binds
|
||||
`--host 127.0.0.1 --port 0` inside the container, so the port is kernel-assigned
|
||||
and never published. Start a second one on a fixed port instead:
|
||||
|
||||
```sh
|
||||
llama-server -m /mnt/hdd1/llms/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf \
|
||||
--host 127.0.0.1 --port 18100 -c 4096 -ngl 99 --no-webui
|
||||
```
|
||||
|
||||
`-c 4096` matters. The recorded numbers were measured at that context size, and
|
||||
a mismatch invalidates the comparison.
|
||||
|
||||
Then:
|
||||
|
||||
```sh
|
||||
make eval-models MAVEN_LLM_URL=http://127.0.0.1:18100 # want ~72.7% cascade
|
||||
make eval-router # classifier baseline
|
||||
MAVEN_LLM_URL=http://127.0.0.1:18100 make eval-phrasing # persona checks, slow
|
||||
make eval-recall
|
||||
```
|
||||
|
||||
A large miss against 72.7% means the deploy differs from the bench harness.
|
||||
|
||||
Two things to decide while the numbers are in front of you:
|
||||
|
||||
- **319's gate recalibration.** The single-token rule needs narrowing or
|
||||
dropping. This needs your judgement, not a threshold sweep. The fixture and the
|
||||
daemon disagree about what is correct on two of the three false clarifies.
|
||||
- **278's real ask** is making the eval lab routine rather than building it. It
|
||||
is built. Decide whether it runs on a timer, on every merge, or on demand, and
|
||||
the task can close.
|
||||
|
||||
Item 4 of **320** needs a permission I do not have. Kill the `llama-server`
|
||||
pid under `maven-mavend-1`, post a turn, and confirm it still completes
|
||||
through the classifier. Either grant it or run it yourself. It is the only
|
||||
check that the failure floor catches a mid-session model death.
|
||||
|
||||
---
|
||||
|
||||
## Session 3: the interaction batch (a day, or three sittings)
|
||||
|
||||
These need real use rather than a command, grouped by what one sitting covers.
|
||||
|
||||
**Morning and delivery** (**280**, **281**, **128**, **282**): open `/morning`,
|
||||
walk the seven required behaviours, then check the four interruption outcomes
|
||||
and the digest gap. **282** needs the `desk_active` script enabled on the desk
|
||||
PC first, which is **15** and needs you at that machine.
|
||||
|
||||
**Tasks and calendar** (**129**, **130**, **127**, **126**, **246**): capture a
|
||||
task by voice, confirm it lands, check prioritisation ordering is not nonsense.
|
||||
**246** (mail reader) also exercises the `IngestMail` rung that moved to
|
||||
`AuthWrite` this morning.
|
||||
|
||||
**Routines and patterns** (**43**, **46**, **247**, **254**): these need history
|
||||
to detect against. If the database is thin after the outage, they may have
|
||||
nothing to propose, which is not a failure. Check `/routines` before
|
||||
concluding anything.
|
||||
|
||||
**Ecosystem** (**272**, **273**, **276**): nexus, hexis and praxis are wired and
|
||||
logged clean at boot. **276** is the degraded-mode suite, which means taking
|
||||
siblings down on purpose. Worth doing while you are already in there.
|
||||
|
||||
---
|
||||
|
||||
## Housekeeping (one sitting, no box needed)
|
||||
|
||||
Four QA tasks will not close no matter how long they sit, because they are
|
||||
gated on something that does not exist:
|
||||
|
||||
- **125** zenmoney: needs a token you have not minted.
|
||||
- **256** Home Assistant: needs HA configured.
|
||||
- **257** Bluetooth: BLOCKED, no bluez on the box. Says so in the title.
|
||||
- **288** STT golden audio: needs fixtures generated.
|
||||
|
||||
Relabel these so they stop reading as backlog. They are not verification work
|
||||
that is pending, they are work that has not started.
|
||||
|
||||
Same treatment for the five plan-only tasks (**251** MCP, **252** vision,
|
||||
**253** hearing, **255** speaker recognition, **259** crawler). A `QA:` prefix on
|
||||
a plan is misleading.
|
||||
|
||||
---
|
||||
|
||||
## Needs you specifically
|
||||
|
||||
Not QA. These are blocked on a decision or a credential only you have.
|
||||
|
||||
| # | what |
|
||||
|---|---|
|
||||
| 16 | Create the Kuma API key. `-kuma-key uk5_mavpoll-key` in `docker-compose.yml` is still the placeholder. |
|
||||
| 15 | Deploy `desk_active` on the desk PC. Blocks **282**. |
|
||||
| 122 | Finish the CPT run for Qwen3-1.7B. The persona fix depends on it. |
|
||||
| 355 | Deploy the Hexis auth change. Was blocked on Maven being under construction, which it no longer is. The client half is vendored and wired. |
|
||||
| 357 | Decide whether entity-existence validation is the permanent target guard or whether blessing lands in Nexus. |
|
||||
| 275 | Hexis native API and MCP parity. |
|
||||
| — | Decide on `-require-stepup`. Making it the default needs WebAuthn configured first, or it locks you out of your own admin surfaces. See **317**. |
|
||||
| — | Three nginx sites bind wildcard `:80` (`acme.conf`, `matrix`, `panel`), so the ecosystem's bind-level protection is not in effect and `allow`/`deny` is carrying it alone. See **354**. |
|
||||
|
||||
---
|
||||
|
||||
## Suggested order
|
||||
|
||||
1. Session 1. If the voice loop is broken, nothing else matters.
|
||||
2. The `-require-stepup` and Kuma decisions. Five minutes, unblocks **317** fully
|
||||
and **16**.
|
||||
3. Session 2. The numbers tell you whether the router is worth its 90x latency.
|
||||
4. Housekeeping. Cheap, and it makes the remaining backlog honest.
|
||||
5. Session 3, split whichever way suits you.
|
||||
@@ -1,6 +1,6 @@
|
||||
// Package main is mavenclient — maven's reference client.
|
||||
//
|
||||
// Per DESIGN.md § Voice pipeline (STT / TTS): capture lives on the client;
|
||||
// Per docs/design.md § Voice pipeline (STT / TTS): capture lives on the client;
|
||||
// the server transcribes + synthesises on demand. The PC client runs the
|
||||
// wake-word / VAD gate (cmd/mavwaked) and ships ONE clean audio blob per
|
||||
// utterance on activation. The server never owns a mic.
|
||||
|
||||
+50
-14
@@ -7,6 +7,7 @@ import (
|
||||
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/router"
|
||||
"github.com/kami/maven/internal/store"
|
||||
)
|
||||
|
||||
// actionFact handles router.IntentFact: persist a tapped self-fact, index
|
||||
@@ -15,14 +16,41 @@ func (h *reactiveHandler) actionFact(ctx context.Context, dec router.Decision) s
|
||||
if !dec.Slots.HasKey {
|
||||
return "не разобрала, что записать — попробуй иначе."
|
||||
}
|
||||
// A question is never a fact about him (#470). "какая последняя версия
|
||||
// языка Go?" used to land here, and the value stored was whatever the
|
||||
// model invented for it, at confidence 1.00, indexed for recall under the
|
||||
// question's own text. Two such rows then claimed seven unrelated world
|
||||
// questions through recall and silently disabled world answering.
|
||||
//
|
||||
// The routing error itself is not fixed here — the answer is to answer.
|
||||
// Sending the turn down the query chain is what he asked for anyway, and
|
||||
// it costs a mis-routed capture nothing: an explicit "запиши ..." is not
|
||||
// question-shaped, so it never takes this branch.
|
||||
if router.IsQuestionShaped(dec.Utterance) {
|
||||
log.Printf("voice: fact write refused, utterance is a question: %q (key %q) — answering as a query",
|
||||
dec.Utterance, dec.Slots.Key)
|
||||
q := dec
|
||||
q.Intent = router.IntentQuery
|
||||
// The key the model extracted is its guess at what to store, not a
|
||||
// fact he has. Left in place, queryFactByKey would read it back and
|
||||
// claim the turn before any real source ran.
|
||||
q.Slots.Key, q.Slots.HasKey = "", false
|
||||
q.Slots.Value = ""
|
||||
return h.actionQuery(ctx, q)
|
||||
}
|
||||
now := h.now()
|
||||
req := ipc.WriteFactReq{
|
||||
Ts: now,
|
||||
Kind: "self",
|
||||
Key: dec.Slots.Key,
|
||||
Value: dec.Slots.Value,
|
||||
Source: "tap:voice",
|
||||
Confidence: 1.0,
|
||||
Ts: now,
|
||||
Kind: "self",
|
||||
Key: dec.Slots.Key,
|
||||
Value: dec.Slots.Value,
|
||||
Source: "tap:voice",
|
||||
// Not 1.00 unconditionally any more (#470). A value he said is
|
||||
// evidence; a value the model supplied for words he never said is a
|
||||
// guess, and writing a guess at full confidence is the same mistake
|
||||
// the act path already refuses under "LLM output is not
|
||||
// authorization".
|
||||
Confidence: factConfidence(dec.Utterance, dec.Slots.Value),
|
||||
// Subject: the key doubles as the entity-resolution candidate —
|
||||
// a voice-tapped fact's key is usually the thing/person it's
|
||||
// about ("espresso_machine", "kate"), so queueing it for Nexus
|
||||
@@ -36,17 +64,25 @@ func (h *reactiveHandler) actionFact(ctx context.Context, dec router.Decision) s
|
||||
log.Printf("voice: write fact: %v", err)
|
||||
return "не получилось сохранить факт."
|
||||
}
|
||||
// Index the fact utterance in long-term memory (best-effort, must not
|
||||
// fail the fact write). Facts aren't in the notes table, so this is the
|
||||
// only recall path for them — "когда я пил воду?" reads back from here.
|
||||
// Index the fact in long-term memory (best-effort, must not fail the fact
|
||||
// write). Facts aren't in the notes table, so this is the only recall path
|
||||
// for them — "когда я пил воду?" reads back from here.
|
||||
//
|
||||
// The indexed text is the fact, not the utterance (#493). queryMemory
|
||||
// returns a fact's stored text verbatim, so what goes in here is what he
|
||||
// hears; storing the utterance meant recall answered with his own sentence
|
||||
// rather than the value. The utterance stays alongside as provenance —
|
||||
// readable on /trace, never the answer and never embedded.
|
||||
if h.memStore != nil {
|
||||
if vec, err := router.EmbedPassage(ctx, h.embedder, dec.Utterance); err != nil {
|
||||
text := store.FactRecallText(dec.Slots.Key, dec.Slots.Value)
|
||||
if vec, err := router.EmbedPassage(ctx, h.embedder, text); err != nil {
|
||||
log.Printf("voice: embed fact for memory: %v", err)
|
||||
} else if err := h.memStore.Insert(ctx, "fact:"+dec.Slots.Key+":"+strconv.FormatInt(now.Unix(), 10), vec, map[string]string{
|
||||
"source": "voice",
|
||||
"type": "fact",
|
||||
"text": dec.Utterance,
|
||||
"ts": strconv.FormatInt(now.Unix(), 10),
|
||||
"source": "voice",
|
||||
"type": "fact",
|
||||
"text": text,
|
||||
"utterance": dec.Utterance,
|
||||
"ts": strconv.FormatInt(now.Unix(), 10),
|
||||
}); err != nil {
|
||||
log.Printf("voice: memory insert fact: %v", err)
|
||||
}
|
||||
|
||||
+39
-24
@@ -13,6 +13,7 @@ import (
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/memory"
|
||||
"github.com/kami/maven/internal/morning"
|
||||
"github.com/kami/maven/internal/phraser"
|
||||
"github.com/kami/maven/internal/router"
|
||||
"github.com/kami/maven/internal/rss"
|
||||
"github.com/kami/maven/internal/store"
|
||||
@@ -433,6 +434,14 @@ func (h *reactiveHandler) queryMemory(ctx context.Context, t *queryTurn) (string
|
||||
return "", false
|
||||
}
|
||||
text := hit.Meta["text"]
|
||||
// The score cleared the gate and the topic still has to match (#470). A
|
||||
// note about his slow network scored high enough to answer "почему небо
|
||||
// синее?", because the right-note and must-be-silent score ranges overlap
|
||||
// and no threshold sits between them.
|
||||
if !memory.RecallAllowed(t.dec.Utterance, text) {
|
||||
log.Printf("voice: recall %q rejected for %q: a world question and no shared topic word", text, t.dec.Utterance)
|
||||
return "", false
|
||||
}
|
||||
// A note is phrased in Maven's voice; a fact is read back as it was
|
||||
// stored.
|
||||
if hit.Meta["type"] == "note" {
|
||||
@@ -468,6 +477,12 @@ func (h *reactiveHandler) queryNotes(ctx context.Context, t *queryTurn) (string,
|
||||
if !memory.ConfidentScores(noteScores, h.queryMinScore, h.queryMinMargin) {
|
||||
return "", false
|
||||
}
|
||||
// Same topic veto as queryMemory above: the best note must be about what
|
||||
// he asked, not merely the nearest vector in the index.
|
||||
if !memory.RecallAllowed(t.dec.Utterance, notes[0].Text) {
|
||||
log.Printf("voice: note %q rejected for %q: a world question and no shared topic word", notes[0].Text, t.dec.Utterance)
|
||||
return "", false
|
||||
}
|
||||
texts := make([]string, len(notes))
|
||||
for i, n := range notes {
|
||||
texts[i] = n.Text
|
||||
@@ -523,10 +538,7 @@ func (h *reactiveHandler) queryWeb(ctx context.Context, t *queryTurn) (string, b
|
||||
// the question he actually asked. She answers the question, she does not
|
||||
// recite the page.
|
||||
snippet := page.Title + "\n" + crawl.TrimRunes(page.Text, webPageContextRunes)
|
||||
reply, perr := h.phraser.PhraseQuery(ctx, t.dec.Utterance, []string{snippet})
|
||||
if perr != nil {
|
||||
log.Printf("voice: web: phrase: %v", perr)
|
||||
}
|
||||
reply := h.phraseSource(ctx, "web", t.dec.Utterance, []string{snippet})
|
||||
if reply == "" {
|
||||
// No phraser (or it failed): read back the top of the page rather than
|
||||
// pretend the fetch did not happen.
|
||||
@@ -592,14 +604,7 @@ func (h *reactiveHandler) querySearch(ctx context.Context, t *queryTurn) (string
|
||||
// question he asked, not something to recite. The trim is one budget over the
|
||||
// joined block, so a long first snippet cannot crowd out the rest.
|
||||
evidence := crawl.TrimRunes(strings.Join(resp.Snippets(), "\n"), h.search.runes)
|
||||
var reply string
|
||||
if h.phraser != nil {
|
||||
var perr error
|
||||
reply, perr = h.phraser.PhraseQuery(ctx, t.dec.Utterance, []string{evidence})
|
||||
if perr != nil {
|
||||
log.Printf("voice: search: phrase: %v", perr)
|
||||
}
|
||||
}
|
||||
reply := h.phraseSource(ctx, "search", t.dec.Utterance, []string{evidence})
|
||||
if reply == "" {
|
||||
// No phraser, or it failed. Read back the best evidence rather than
|
||||
// pretend the search did not happen.
|
||||
@@ -680,14 +685,7 @@ func (h *reactiveHandler) queryKiwix(ctx context.Context, t *queryTurn) (string,
|
||||
// Handed over the same way a note or a page is: context for the question he
|
||||
// asked, not something to recite.
|
||||
snippet := top.Title + "\n" + crawl.TrimRunes(page.Text, h.kiwix.runes)
|
||||
var reply string
|
||||
if h.phraser != nil {
|
||||
var perr error
|
||||
reply, perr = h.phraser.PhraseQuery(ctx, t.dec.Utterance, []string{snippet})
|
||||
if perr != nil {
|
||||
log.Printf("voice: kiwix: phrase: %v", perr)
|
||||
}
|
||||
}
|
||||
reply := h.phraseSource(ctx, "kiwix", t.dec.Utterance, []string{snippet})
|
||||
if reply == "" {
|
||||
// No phraser, or it failed. Read back the best hit rather than pretend
|
||||
// the search did not happen.
|
||||
@@ -755,11 +753,28 @@ func isPersonalQuery(utterance string) bool {
|
||||
return false
|
||||
}
|
||||
|
||||
// queryGeneral — general knowledge from the phraser, the last source before
|
||||
// giving up. It always claims: either the model answers or Maven says she
|
||||
// doesn't know.
|
||||
// queryGeneral — general knowledge, the last source before giving up. It always
|
||||
// claims: either a model answers, or Maven names the gap, or she says she does
|
||||
// not know.
|
||||
//
|
||||
// This is the sharpest case for the naming half. Nothing has been fetched, so
|
||||
// there is no passage to fall back on and no floor under the answer except the
|
||||
// model's weights — and a 1.7B's weights are where the invented answers come
|
||||
// from. With a workstation configured and asleep he is told that, rather than
|
||||
// told something false in a confident voice. With no workstation configured at
|
||||
// all the resident model answers exactly as it does today: naming a gap requires
|
||||
// a gap, and on that box the 1.7B is the whole product.
|
||||
func (h *reactiveHandler) queryGeneral(ctx context.Context, t *queryTurn) (string, bool) {
|
||||
reply, err := h.phraser.PhraseQuery(ctx, t.dec.Utterance, nil)
|
||||
if h.phraser == nil {
|
||||
// No model of any size. That is not the workstation being asleep, so it
|
||||
// is not that gap: it is simply not knowing.
|
||||
return "не знаю.", true
|
||||
}
|
||||
reply, err := h.phraseWorld(ctx, t.dec.Utterance, nil)
|
||||
if errors.Is(err, phraser.ErrNoWorldModel) {
|
||||
log.Printf("voice: %q needs the world model and it is not available", t.dec.Utterance)
|
||||
return worldGap, true
|
||||
}
|
||||
if err != nil || reply == "" {
|
||||
return "не знаю.", true
|
||||
}
|
||||
|
||||
+12
-8
@@ -107,14 +107,18 @@ func (h *reactiveHandler) confirmResolvers(ctx context.Context) []confirmResolve
|
||||
return pr != nil && !h.now().After(pr.expiry)
|
||||
},
|
||||
yes: func() string {
|
||||
// Only record the acceptance. The tick loop reads accepted
|
||||
// routines and nudges on their own interval. Building a
|
||||
// reminder here made a routine fire exactly once (Vikunja #366).
|
||||
if err := h.dataStore.AcceptProposedRoutine(ctx, pr.routineID, h.now()); err != nil {
|
||||
log.Printf("voice: accept proposed routine: %v", err)
|
||||
return "не получилось запомнить рутину."
|
||||
}
|
||||
return "буду напоминать."
|
||||
// Voice does NOT accept (Vikunja #367). Accepting hands the
|
||||
// tick loop a standing new reason to speak, which is the same
|
||||
// tier as enabling a tool — and DESIGN.md § "surface caps
|
||||
// authority" says a room mic, reachable by anyone present, is
|
||||
// structurally incapable of layer 3. So a spoken "да" leaves
|
||||
// the row 'proposed' and points at the authed page, where the
|
||||
// accept button is gated at step-up. The convenience of
|
||||
// answering out loud stays; the authority does not move.
|
||||
//
|
||||
// Acceptance itself is recorded by /routines, and the tick
|
||||
// loop nudges on the interval from there (Vikunja #366).
|
||||
return "поняла — подтверди на странице рутин, и начну напоминать."
|
||||
},
|
||||
no: func() string {
|
||||
if err := h.dataStore.DismissProposedRoutine(ctx, pr.routineID); err != nil {
|
||||
|
||||
@@ -0,0 +1,82 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"log"
|
||||
"strings"
|
||||
"unicode"
|
||||
)
|
||||
|
||||
// ungroundedConfidence — what a self fact is worth when its value appears
|
||||
// nowhere in what he said. Below `query_min_score` is not the point (recall
|
||||
// gates on vector distance, not on this number); the point is that
|
||||
// `/history` and every future reader can tell a value he said from a value
|
||||
// the model supplied.
|
||||
const ungroundedConfidence = 0.6
|
||||
|
||||
// factConfidence scores a self fact by whether its value is grounded in the
|
||||
// utterance it came from. Grounded stays 1.00, which is what a tapped fact
|
||||
// has always been worth. Ungrounded drops, and says so in the log.
|
||||
//
|
||||
// An empty value is grounded by definition: the key alone carries the fact
|
||||
// ("поужинал"), and there is nothing for the model to have invented.
|
||||
func factConfidence(utterance, value string) float64 {
|
||||
if strings.TrimSpace(value) == "" {
|
||||
return 1.0
|
||||
}
|
||||
if valueGrounded(utterance, value) {
|
||||
return 1.0
|
||||
}
|
||||
log.Printf("voice: fact value %q is not in %q — writing at confidence %.2f",
|
||||
value, utterance, ungroundedConfidence)
|
||||
return ungroundedConfidence
|
||||
}
|
||||
|
||||
// valueGrounded reports whether every word of value traces back to a word he
|
||||
// actually said. The comparison is on a 4-rune prefix, so the model's
|
||||
// normalization survives ("пил воду" → "вода") while an invented value
|
||||
// ("1.20" for a question about Go) does not.
|
||||
func valueGrounded(utterance, value string) bool {
|
||||
said := factTokens(utterance)
|
||||
words := factTokens(value)
|
||||
if len(words) == 0 {
|
||||
return true
|
||||
}
|
||||
for _, w := range words {
|
||||
if !anyTokenMatches(said, w) {
|
||||
return false
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
func anyTokenMatches(said []string, w string) bool {
|
||||
for _, s := range said {
|
||||
if s == w || sameStem(s, w) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// sameStem is inflection tolerance and nothing more: it compares all but the
|
||||
// last rune of the shorter word, and never fewer than three. Russian marks
|
||||
// case on the ending, so "пил воду" and the stored "вода" are the same word he
|
||||
// said, while "1.20" and "версия" are not. A word of three runes or fewer must
|
||||
// match outright, where a shorter prefix would match half the language.
|
||||
func sameStem(a, b string) bool {
|
||||
ar, br := []rune(a), []rune(b)
|
||||
shorter := min(len(ar), len(br))
|
||||
n := shorter - 1
|
||||
if n < 3 || len(ar) < n || len(br) < n {
|
||||
return false
|
||||
}
|
||||
return string(ar[:n]) == string(br[:n])
|
||||
}
|
||||
|
||||
// factTokens lowercases and splits on everything that is not a letter or a
|
||||
// digit, the same shape planTokens uses in the router.
|
||||
func factTokens(s string) []string {
|
||||
return strings.FieldsFunc(strings.ToLower(s), func(r rune) bool {
|
||||
return !unicode.IsLetter(r) && !unicode.IsDigit(r)
|
||||
})
|
||||
}
|
||||
@@ -0,0 +1,125 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/memory"
|
||||
"github.com/kami/maven/internal/router"
|
||||
"github.com/kami/maven/internal/tool"
|
||||
"github.com/kami/maven/internal/voice"
|
||||
)
|
||||
|
||||
func newFactGateHandler(t *testing.T, now time.Time) (*reactiveHandler, ipc.CoreAPI) {
|
||||
t.Helper()
|
||||
st := newTestStore(t)
|
||||
api := ipc.NewStoreAPI(st)
|
||||
emb := router.NewHashEmbedder(1024)
|
||||
h := &reactiveHandler{
|
||||
api: api,
|
||||
embedder: emb,
|
||||
router: buildRouter(emb, tool.NewMatcher(api), 0.55, nil),
|
||||
replier: voice.NewStubReplier(),
|
||||
now: func() time.Time { return now },
|
||||
memStore: memory.NewInMemoryStore(),
|
||||
dataStore: st,
|
||||
}
|
||||
return h, api
|
||||
}
|
||||
|
||||
// The write half of #470: a question routed to IntentFact must not become a
|
||||
// fact about him, and must not leave a vector behind for recall to serve.
|
||||
func TestActionFact_QuestionIsNotWritten(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, api := newFactGateHandler(t, time.Now())
|
||||
|
||||
reply := h.actionFact(ctx, router.Decision{
|
||||
Intent: router.IntentFact,
|
||||
Utterance: "какая последняя версия языка Go?",
|
||||
Slots: router.Slots{Key: "go_version", HasKey: true, Value: `"1.20"`},
|
||||
})
|
||||
|
||||
if _, err := api.LatestFact(ctx, "go_version"); err == nil {
|
||||
t.Fatal("a question was stored as a fact about him")
|
||||
}
|
||||
hits, err := h.memStore.Search(ctx, mustEmbedPassage(t, h, "какая последняя версия языка Go?"), 3)
|
||||
if err != nil {
|
||||
t.Fatalf("memory search: %v", err)
|
||||
}
|
||||
if len(hits) != 0 {
|
||||
t.Fatalf("the question was indexed for recall: %+v", hits)
|
||||
}
|
||||
// It went down the query chain instead. Nothing is configured to answer a
|
||||
// world question in this harness, so "не знаю." is the honest outcome —
|
||||
// what matters is that the turn was answered, not stored.
|
||||
if reply == "" {
|
||||
t.Fatal("the turn was neither stored nor answered")
|
||||
}
|
||||
}
|
||||
|
||||
// The capture that must survive the gate: an explicit instruction to record,
|
||||
// even though it contains an interrogative.
|
||||
func TestActionFact_ExplicitCaptureStillWrites(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, api := newFactGateHandler(t, time.Now())
|
||||
|
||||
h.actionFact(ctx, router.Decision{
|
||||
Intent: router.IntentFact,
|
||||
Utterance: "запиши что я пил воду",
|
||||
Slots: router.Slots{Key: "water", HasKey: true, Value: `"вода"`},
|
||||
})
|
||||
|
||||
f, err := api.LatestFact(ctx, "water")
|
||||
if err != nil {
|
||||
t.Fatalf("an explicit capture was refused: %v", err)
|
||||
}
|
||||
if f.Confidence != 1.0 {
|
||||
t.Errorf("confidence = %v, want 1.0 for a value he said", f.Confidence)
|
||||
}
|
||||
// #493: what recall reads back is the fact, not the sentence he said.
|
||||
// queryMemory returns a fact's text verbatim, so the utterance sitting here
|
||||
// meant "запиши что я пил воду" was the answer to "когда я пил воду?".
|
||||
hits, err := h.memStore.Search(ctx, mustEmbedPassage(t, h, "вода"), 3)
|
||||
if err != nil {
|
||||
t.Fatalf("memory search: %v", err)
|
||||
}
|
||||
if len(hits) != 1 {
|
||||
t.Fatalf("the fact was not indexed once: %+v", hits)
|
||||
}
|
||||
if got := hits[0].Meta["text"]; got != "water — вода" {
|
||||
t.Errorf("indexed text = %q, want the fact", got)
|
||||
}
|
||||
if got := hits[0].Meta["utterance"]; got != "запиши что я пил воду" {
|
||||
t.Errorf("utterance provenance = %q, want it kept alongside", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestFactConfidence(t *testing.T) {
|
||||
cases := []struct {
|
||||
utterance, value string
|
||||
want float64
|
||||
}{
|
||||
{"запиши что я пил воду", `"вода"`, 1.0},
|
||||
{"я выпил кофе", `"кофе"`, 1.0},
|
||||
{"поужинал", "", 1.0},
|
||||
{"отметь что я полил кактус", `"полил кактус"`, 1.0},
|
||||
{"какая последняя версия языка Go", `"1.20"`, ungroundedConfidence},
|
||||
{"кто премьер Японии", `"Тонио Озаки"`, ungroundedConfidence},
|
||||
}
|
||||
for _, c := range cases {
|
||||
if got := factConfidence(c.utterance, c.value); got != c.want {
|
||||
t.Errorf("factConfidence(%q, %q) = %v, want %v", c.utterance, c.value, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func mustEmbedPassage(t *testing.T, h *reactiveHandler, text string) []float32 {
|
||||
t.Helper()
|
||||
vec, err := router.EmbedQuery(context.Background(), h.embedder, text)
|
||||
if err != nil {
|
||||
t.Fatalf("embed %q: %v", text, err)
|
||||
}
|
||||
return vec
|
||||
}
|
||||
+12
-5
@@ -12,13 +12,21 @@ import (
|
||||
// which waits on voice-print attribution (see PROGRESS multi-user deferral).
|
||||
const voiceDialogueID = "voice"
|
||||
|
||||
// toDialogueSlots projects the router's slots onto the dialogue layer's subset
|
||||
// (everything except the fact Value, which the dialogue layer doesn't carry).
|
||||
// toDialogueSlots and applyDialogueSlots are the only bridge between
|
||||
// router.Slots and dialogue.Slots. dialogue must not import router (import
|
||||
// cycle), so the two structs are hand-kept copies and every field has to be
|
||||
// carried by hand here. Adding a field to either struct without adding it to
|
||||
// BOTH functions loses a slot silently — nothing fails to build. The tests in
|
||||
// slotsparity_test.go fail when the field sets or the converters stop matching;
|
||||
// when they do, fix these two functions, not the tests.
|
||||
|
||||
// toDialogueSlots projects the router's slots onto the dialogue layer's copy.
|
||||
func toDialogueSlots(s router.Slots) dialogue.Slots {
|
||||
return dialogue.Slots{
|
||||
Time: s.Time,
|
||||
HasTime: s.HasTime,
|
||||
Key: s.Key,
|
||||
Value: s.Value,
|
||||
HasKey: s.HasKey,
|
||||
Text: s.Text,
|
||||
Fn: s.Fn,
|
||||
@@ -27,11 +35,10 @@ func toDialogueSlots(s router.Slots) dialogue.Slots {
|
||||
}
|
||||
}
|
||||
|
||||
// applyDialogueSlots writes inherited dialogue slots back onto router slots,
|
||||
// preserving router-only fields (Value) the dialogue layer never touched.
|
||||
// applyDialogueSlots writes dialogue slots back onto router slots.
|
||||
func applyDialogueSlots(base router.Slots, d dialogue.Slots) router.Slots {
|
||||
base.Time, base.HasTime = d.Time, d.HasTime
|
||||
base.Key, base.HasKey = d.Key, d.HasKey
|
||||
base.Key, base.Value, base.HasKey = d.Key, d.Value, d.HasKey
|
||||
base.Text = d.Text
|
||||
base.Fn, base.Args, base.HasFn = d.Fn, d.Args, d.HasFn
|
||||
return base
|
||||
|
||||
@@ -0,0 +1,91 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/config"
|
||||
"github.com/kami/maven/internal/llm"
|
||||
)
|
||||
|
||||
// No `workstation` block is the shipping deploy. The seam must then be the
|
||||
// resident client itself, with nothing probing anything.
|
||||
func TestModelSeamUnconfiguredIsResidentOnly(t *testing.T) {
|
||||
resident := llm.New("http://127.0.0.1:1", time.Second)
|
||||
hot, pair := modelSeam(&config.Config{}, resident)
|
||||
if pair != nil {
|
||||
t.Error("built a pair with no workstation configured")
|
||||
}
|
||||
if hot == nil {
|
||||
t.Fatal("no seam at all, so the cascade would route with the classifier")
|
||||
}
|
||||
}
|
||||
|
||||
// A workstation with no resident model behind it has no floor, and a Pair with
|
||||
// no floor is a configuration mistake rather than a degraded mode.
|
||||
func TestModelSeamWithoutResidentIsNil(t *testing.T) {
|
||||
cfg := &config.Config{Workstation: &config.WorkstationConfig{URL: "http://127.0.0.1:1"}}
|
||||
cfg.Workstation.Health = strings.TrimRight(cfg.Workstation.URL, "/") + "/health"
|
||||
hot, pair := modelSeam(cfg, nil)
|
||||
if hot != nil || pair != nil {
|
||||
t.Errorf("built a seam with no floor: hot=%v pair=%v", hot, pair)
|
||||
}
|
||||
}
|
||||
|
||||
// The configured case: the seam is the pair, and the pair notices a workstation
|
||||
// that answers /health.
|
||||
func TestModelSeamPrefersAnAnsweringWorkstation(t *testing.T) {
|
||||
up := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
w.WriteHeader(http.StatusOK)
|
||||
}))
|
||||
defer up.Close()
|
||||
|
||||
cfg := &config.Config{Workstation: &config.WorkstationConfig{
|
||||
URL: up.URL,
|
||||
Probe: config.Duration(10 * time.Millisecond),
|
||||
}}
|
||||
cfg.Workstation.Health = strings.TrimRight(cfg.Workstation.URL, "/") + "/health"
|
||||
|
||||
hot, pair := modelSeam(cfg, llm.New("http://127.0.0.1:1", time.Second))
|
||||
if pair == nil || hot == nil {
|
||||
t.Fatal("no pair built for a configured workstation")
|
||||
}
|
||||
defer pair.Stop()
|
||||
|
||||
deadline := time.Now().Add(2 * time.Second)
|
||||
for !pair.Available() && time.Now().Before(deadline) {
|
||||
time.Sleep(5 * time.Millisecond)
|
||||
}
|
||||
if !pair.Available() {
|
||||
t.Fatal("the pair never saw a workstation that answers /health")
|
||||
}
|
||||
}
|
||||
|
||||
// A card held by a CPT run answers 503, and that must read as unavailable
|
||||
// rather than as an error a turn has to handle.
|
||||
func TestModelSeamHeldCardIsUnavailable(t *testing.T) {
|
||||
busy := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
http.Error(w, "model not loaded", http.StatusServiceUnavailable)
|
||||
}))
|
||||
defer busy.Close()
|
||||
|
||||
cfg := &config.Config{Workstation: &config.WorkstationConfig{
|
||||
URL: busy.URL,
|
||||
Probe: config.Duration(10 * time.Millisecond),
|
||||
}}
|
||||
cfg.Workstation.Health = strings.TrimRight(cfg.Workstation.URL, "/") + "/health"
|
||||
|
||||
_, pair := modelSeam(cfg, llm.New("http://127.0.0.1:1", time.Second))
|
||||
if pair == nil {
|
||||
t.Fatal("no pair built for a configured workstation")
|
||||
}
|
||||
defer pair.Stop()
|
||||
|
||||
time.Sleep(50 * time.Millisecond)
|
||||
if pair.Available() {
|
||||
t.Error("a 503 from the supervisor read as available")
|
||||
}
|
||||
}
|
||||
@@ -9,6 +9,7 @@ import (
|
||||
|
||||
"github.com/kami/maven/internal/config"
|
||||
"github.com/kami/maven/internal/delivery"
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/loop"
|
||||
"github.com/kami/maven/internal/pattern"
|
||||
"github.com/kami/maven/internal/store"
|
||||
@@ -283,3 +284,81 @@ func TestTickProposalCooldownSpacesAnnouncements(t *testing.T) {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestVoiceYesDoesNotAcceptRoutine — Vikunja #367. Accepting a routine hands
|
||||
// the tick loop a standing new reason to speak, which DESIGN.md puts at layer
|
||||
// 3, and voice is structurally incapable of layer 3. A spoken "да" must park
|
||||
// the decision for the authed page, not flip the row itself.
|
||||
func TestVoiceYesDoesNotAcceptRoutine(t *testing.T) {
|
||||
st := newTestStore(t)
|
||||
ctx := context.Background()
|
||||
now := refNow()
|
||||
seedRefillEvents(t, st, ctx, now, pattern.MinEvents-1)
|
||||
|
||||
h := &reactiveHandler{api: ipc.NewStoreAPI(st), dataStore: st, now: func() time.Time { return now }}
|
||||
|
||||
// The MinEvents'th event is the one that makes the pattern detectable, and
|
||||
// it goes through the voice path so the proposal is parked for a y/n.
|
||||
last := now.Add(time.Duration(pattern.MinEvents-1) * 7 * 24 * time.Hour)
|
||||
factID, err := st.WriteFact(ctx, last, store.KindSelf, "cat_water", "refill", "voice", 1.0, sql.NullInt64{})
|
||||
if err != nil {
|
||||
t.Fatalf("write fact: %v", err)
|
||||
}
|
||||
if phrase := h.detectPattern(ctx, factID, "cat_water", "refill", last); phrase == "" {
|
||||
t.Fatal("expected a parked routine proposal")
|
||||
}
|
||||
|
||||
reply, handled := h.resolveConfirm(ctx, "да")
|
||||
if !handled {
|
||||
t.Fatal("the spoken yes should be consumed by the routine confirm")
|
||||
}
|
||||
if !strings.Contains(reply, "рутин") {
|
||||
t.Fatalf("reply should send him to the routines page, got %q", reply)
|
||||
}
|
||||
|
||||
rows, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineAccepted)
|
||||
if err != nil {
|
||||
t.Fatalf("list accepted: %v", err)
|
||||
}
|
||||
if len(rows) != 0 {
|
||||
t.Fatalf("voice accepted a routine: %+v", rows)
|
||||
}
|
||||
proposed, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineProposed)
|
||||
if err != nil {
|
||||
t.Fatalf("list proposed: %v", err)
|
||||
}
|
||||
if len(proposed) != 1 {
|
||||
t.Fatalf("proposed routines = %d, want 1 (still waiting for the page)", len(proposed))
|
||||
}
|
||||
}
|
||||
|
||||
// TestVoiceNoStillDismissesRoutine — declining does not move the boundary
|
||||
// outward, so voice keeps it. Only acceptance is gated.
|
||||
func TestVoiceNoStillDismissesRoutine(t *testing.T) {
|
||||
st := newTestStore(t)
|
||||
ctx := context.Background()
|
||||
now := refNow()
|
||||
seedRefillEvents(t, st, ctx, now, pattern.MinEvents-1)
|
||||
|
||||
h := &reactiveHandler{api: ipc.NewStoreAPI(st), dataStore: st, now: func() time.Time { return now }}
|
||||
|
||||
last := now.Add(time.Duration(pattern.MinEvents-1) * 7 * 24 * time.Hour)
|
||||
factID, err := st.WriteFact(ctx, last, store.KindSelf, "cat_water", "refill", "voice", 1.0, sql.NullInt64{})
|
||||
if err != nil {
|
||||
t.Fatalf("write fact: %v", err)
|
||||
}
|
||||
if phrase := h.detectPattern(ctx, factID, "cat_water", "refill", last); phrase == "" {
|
||||
t.Fatal("expected a parked routine proposal")
|
||||
}
|
||||
|
||||
if _, handled := h.resolveConfirm(ctx, "нет"); !handled {
|
||||
t.Fatal("the spoken no should be consumed by the routine confirm")
|
||||
}
|
||||
rows, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineDismissed)
|
||||
if err != nil {
|
||||
t.Fatalf("list dismissed: %v", err)
|
||||
}
|
||||
if len(rows) != 1 {
|
||||
t.Fatalf("dismissed routines = %d, want 1", len(rows))
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
// mavend/simulator_test.go — the replayable full-system simulator
|
||||
// (Vikunja #284, 20-07-2026-BACKLOG.md item 7).
|
||||
// (Vikunja #284).
|
||||
//
|
||||
// # What it is
|
||||
//
|
||||
|
||||
@@ -0,0 +1,71 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"reflect"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/dialogue"
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
|
||||
// TestSlotsParity — dialogue.Slots is a hand-kept copy of router.Slots
|
||||
// (dialogue must not import router: import cycle). Drift is silent, so this
|
||||
// test compares the two field sets by name and type. If it fails, add the new
|
||||
// field to both structs AND to toDialogueSlots/applyDialogueSlots in
|
||||
// followup.go — do not relax the test.
|
||||
func TestSlotsParity(t *testing.T) {
|
||||
fields := func(v any) map[string]string {
|
||||
rt := reflect.TypeOf(v)
|
||||
out := make(map[string]string, rt.NumField())
|
||||
for i := 0; i < rt.NumField(); i++ {
|
||||
f := rt.Field(i)
|
||||
out[f.Name] = f.Type.String()
|
||||
}
|
||||
return out
|
||||
}
|
||||
rf, df := fields(router.Slots{}), fields(dialogue.Slots{})
|
||||
for name, typ := range rf {
|
||||
dt, ok := df[name]
|
||||
if !ok {
|
||||
t.Errorf("router.Slots.%s (%s) missing from dialogue.Slots", name, typ)
|
||||
continue
|
||||
}
|
||||
if dt != typ {
|
||||
t.Errorf("field %s: router has %s, dialogue has %s", name, typ, dt)
|
||||
}
|
||||
}
|
||||
for name, typ := range df {
|
||||
if _, ok := rf[name]; !ok {
|
||||
t.Errorf("dialogue.Slots.%s (%s) missing from router.Slots", name, typ)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestSlotsRoundTrip — the converters carry every field. A field the parity
|
||||
// test accepts can still be dropped in transit, so round-trip a fully
|
||||
// populated value and compare.
|
||||
func TestSlotsRoundTrip(t *testing.T) {
|
||||
full := router.Slots{
|
||||
Time: time.Date(2026, 8, 2, 11, 0, 0, 0, time.UTC),
|
||||
HasTime: true,
|
||||
Fn: "restart",
|
||||
Args: []string{"nginx"},
|
||||
HasFn: true,
|
||||
Key: "water",
|
||||
Value: `"drank"`,
|
||||
HasKey: true,
|
||||
Text: "выпил воды",
|
||||
}
|
||||
// Every field must be non-zero, or the round-trip proves nothing.
|
||||
rv := reflect.ValueOf(full)
|
||||
for i := 0; i < rv.NumField(); i++ {
|
||||
if rv.Field(i).IsZero() {
|
||||
t.Fatalf("field %s is zero: extend this fixture so the round-trip covers it",
|
||||
rv.Type().Field(i).Name)
|
||||
}
|
||||
}
|
||||
if got := applyDialogueSlots(router.Slots{}, toDialogueSlots(full)); !reflect.DeepEqual(got, full) {
|
||||
t.Errorf("round-trip lost a slot:\n got %+v\nwant %+v", got, full)
|
||||
}
|
||||
}
|
||||
+90
-5
@@ -48,7 +48,11 @@ type voiceWiring struct {
|
||||
// mcp — the MCP client, nil unless the `mcp` block configures an enabled
|
||||
// server (Vikunja #251). Its tools land in the same allowlist as every
|
||||
// other act, so nothing else here has to know about it.
|
||||
mcp *mcpWiring
|
||||
// pair — the workstation model with the resident one as the floor, nil
|
||||
// unless a `workstation` block names an address. Held here only so the
|
||||
// prober is stopped on shutdown; callers were handed it at build time.
|
||||
pair *llm.Pair
|
||||
mcp *mcpWiring
|
||||
// home — the Home Assistant client, nil unless the `smarthome` block is
|
||||
// enabled (Vikunja #256). Its devices land in the same allowlist as every
|
||||
// other act, so nothing else here has to know about it.
|
||||
@@ -76,6 +80,9 @@ func (w *voiceWiring) close() {
|
||||
if w.ttsClient != nil {
|
||||
_ = w.ttsClient.Close()
|
||||
}
|
||||
if w.pair != nil {
|
||||
w.pair.Stop()
|
||||
}
|
||||
w.mcp.close()
|
||||
}
|
||||
|
||||
@@ -139,6 +146,7 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
|
||||
emb = router.NewHashEmbedder(1024)
|
||||
}
|
||||
w.embedder = emb
|
||||
repairFactVectors(dataStore, emb)
|
||||
checkStoredEmbedder(dataStore, emb)
|
||||
|
||||
// ----- tool executor (the enabled act allowlist, store-backed) -----
|
||||
@@ -188,6 +196,18 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
|
||||
// the new llama-server when the resident model is swapped (Vikunja #250).
|
||||
llmClient = llmClientFor(lp, 60*time.Second)
|
||||
}
|
||||
// The workstation model sits above that one when it is configured and its
|
||||
// card is free. hot is what the router and the replier complete through:
|
||||
// either the pair, or the resident client alone, or nothing at all.
|
||||
hot, pair := modelSeam(cfg, llmClient)
|
||||
w.pair = pair
|
||||
// The phraser gets the same pair, which is what carries the workstation model
|
||||
// into the paths that do not go through `hot`: world questions (the naming
|
||||
// half), and the digestion worker's nudge and reminder phrasing (the silent
|
||||
// half). Wiring, so it happens once and before the voice server listens.
|
||||
if lp, ok := phr.(*phraser.LLMPhraser); ok && pair != nil {
|
||||
lp.UseRemote(pair)
|
||||
}
|
||||
// ----- router (the cascade; floor examples seed the classifier) -----
|
||||
// The act matcher's allowlist is exactly the enabled tool names — the
|
||||
// router only matches acts the executor can run (one source of truth).
|
||||
@@ -199,7 +219,7 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
|
||||
// against the classifier's 50.0%, at about 1s a turn instead of 30ms (see
|
||||
// config.VoiceConfig.LLMRouter). The classifier always stays wired as the
|
||||
// fallback, so a model error never breaks a turn.
|
||||
rtr := buildRouter(emb, matcher, threshold, pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient))
|
||||
rtr := buildRouter(emb, matcher, threshold, pickLLMRouter(cfg.Voice.UseLLMRouter(), hot))
|
||||
|
||||
// ----- sessions registry (shared with voicesink) -----
|
||||
sessions := voice.NewSessions()
|
||||
@@ -233,8 +253,8 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
|
||||
|
||||
// ----- replier (LLM-backed when the engine is on, Stub floor otherwise) -----
|
||||
replier := voice.Replier(voice.NewStubReplier())
|
||||
if llmClient != nil {
|
||||
replier = newLLMReplier(llmClient, contextBlockFn(cfg, time.Now))
|
||||
if hot != nil {
|
||||
replier = newLLMReplier(hot, contextBlockFn(cfg, time.Now))
|
||||
}
|
||||
|
||||
// ----- the handler (the reactive path; closes over stt / tts / router / coreAPI / memory) -----
|
||||
@@ -291,7 +311,42 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
|
||||
// pickLLMRouter returns the LLM router when the operator asked for it and there
|
||||
// is a llama-server to talk to, and nil otherwise. nil is safe: the cascade then
|
||||
// routes with the classifier, so an unusable setting costs accuracy, not turns.
|
||||
func pickLLMRouter(enabled bool, c *llm.Client) *router.LLMRouter {
|
||||
// modelSeam builds the completion seam the hot paths use: routing and replies.
|
||||
//
|
||||
// With no `workstation` block it is the resident client and nothing probes
|
||||
// anything, which is today's deploy exactly. With one, it is an llm.Pair that
|
||||
// prefers the workstation and falls back to the resident model silently — the
|
||||
// silent half of the degradation rule (docs/offload.md), because the big model
|
||||
// is only better here and the 1.7B is today's shipping quality. He is never
|
||||
// told which of the two phrased his reply.
|
||||
//
|
||||
// A nil resident client means the phraser is not an LLM phraser. There is then
|
||||
// no floor, and a Pair with no floor is a configuration mistake rather than a
|
||||
// degraded mode, so the seam is nil and the cascade routes with the classifier.
|
||||
func modelSeam(cfg *config.Config, resident *llm.Client) (router.Completer, *llm.Pair) {
|
||||
if resident == nil {
|
||||
if cfg.Workstation != nil {
|
||||
log.Printf("voice: a workstation is configured but there is no resident model to floor it with — ignoring the block")
|
||||
}
|
||||
return nil, nil
|
||||
}
|
||||
if cfg.Workstation == nil {
|
||||
return resident, nil
|
||||
}
|
||||
ws := cfg.Workstation
|
||||
pair := llm.NewPair(
|
||||
llm.New(ws.URL, time.Duration(ws.Timeout)),
|
||||
resident,
|
||||
ws.Health,
|
||||
time.Duration(ws.Probe),
|
||||
)
|
||||
pair.Start(context.Background())
|
||||
log.Printf("voice: workstation model at %s, probed every %s, resident model as the floor",
|
||||
ws.URL, time.Duration(ws.Probe))
|
||||
return pair, pair
|
||||
}
|
||||
|
||||
func pickLLMRouter(enabled bool, c router.Completer) *router.LLMRouter {
|
||||
if !enabled {
|
||||
return nil
|
||||
}
|
||||
@@ -419,6 +474,36 @@ func seedTools(api ipc.CoreAPI, tools []config.ToolConfig) {
|
||||
log.Printf("voice: seeded %d act tools from config", n)
|
||||
}
|
||||
|
||||
// repairFactVectors brings stored fact vectors in line with the facts they name
|
||||
// (#493), once per box, before the embedder marker is even looked at.
|
||||
//
|
||||
// Automatic and not a flag, unlike -reembed: only voice-tapped facts are in
|
||||
// this index, so the work is tens of embeddings rather than the thousands of
|
||||
// notes that made the backfill a deliberate act. And the box that needs it is
|
||||
// broken in a way nobody can see — recall answers with the wrong text and
|
||||
// nothing logs an error — so waiting for an operator to know to run it is how
|
||||
// the defect survived four restarts in the first place.
|
||||
func repairFactVectors(dataStore *store.Store, emb router.Embedder) {
|
||||
if dataStore == nil {
|
||||
return
|
||||
}
|
||||
res, err := dataStore.RepairFactVectors(context.Background(),
|
||||
// EmbedPassage, the stored side, same as every other writer of these
|
||||
// vectors.
|
||||
func(ctx context.Context, text string) ([]float32, error) {
|
||||
return router.EmbedPassage(ctx, emb, text)
|
||||
})
|
||||
if err != nil {
|
||||
log.Printf("voice: fact vector repair failed, no marker written and nothing half-done — retried next start: %v", err)
|
||||
return
|
||||
}
|
||||
if res.Skipped || res.Rewritten+res.Dropped == 0 {
|
||||
return
|
||||
}
|
||||
log.Printf("voice: fact vector repair — %d re-embedded from the fact they name, %d dropped as voided or superseded, %d already right, took %s (#493)",
|
||||
res.Rewritten, res.Dropped, res.Kept, res.Took.Round(time.Millisecond))
|
||||
}
|
||||
|
||||
// reembedOnStart is the -reembed flag (set in run()). Opt-in on purpose: see
|
||||
// runReembed.
|
||||
var reembedOnStart bool
|
||||
|
||||
@@ -0,0 +1,60 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"log"
|
||||
|
||||
"github.com/kami/maven/internal/phraser"
|
||||
)
|
||||
|
||||
// worldPhraser — the naming half of the degradation rule (docs/offload.md), as
|
||||
// the query sources see it. Only *phraser.LLMPhraser implements it, so the
|
||||
// Stub and every test double stay exactly as they are.
|
||||
type worldPhraser interface {
|
||||
PhraseWorld(ctx context.Context, utterance string, sources []string) (string, error)
|
||||
}
|
||||
|
||||
// worldGap — what he hears when the question is about the world, the workstation
|
||||
// model is the one configured to answer it, and that machine is not answering.
|
||||
//
|
||||
// It says the true thing. The resident 1.7B is not a worse answer here, it is an
|
||||
// invented one: "Война и мир" came back with Левитан as its author, and a
|
||||
// question about his meeting came back as a swimming competition in Nottingham.
|
||||
// Naming the gap is the rule CLAUDE.md already applies to a sibling service
|
||||
// being down.
|
||||
const worldGap = "сейчас не могу ответить — большая модель недоступна, а придумывать не хочу."
|
||||
|
||||
// phraseWorld asks the world model, or reports the gap.
|
||||
//
|
||||
// The three outcomes come straight from LLMPhraser.PhraseWorld: no workstation
|
||||
// configured means the resident model answers as it always has, a workstation
|
||||
// that is up answers, and a workstation that is down returns
|
||||
// phraser.ErrNoWorldModel. A phraser that has no world seam at all — the Stub,
|
||||
// and the doubles in the tests — is the first of those three.
|
||||
func (h *reactiveHandler) phraseWorld(ctx context.Context, utterance string, sources []string) (string, error) {
|
||||
if h.phraser == nil {
|
||||
return "", phraser.ErrNoWorldModel
|
||||
}
|
||||
if w, ok := h.phraser.(worldPhraser); ok {
|
||||
return w.PhraseWorld(ctx, utterance, sources)
|
||||
}
|
||||
return h.phraser.PhraseQuery(ctx, utterance, sources)
|
||||
}
|
||||
|
||||
// phraseSource asks the world model to answer from a passage someone already
|
||||
// fetched — a live search result, a ZIM article, a page he named. It returns ""
|
||||
// rather than the gap phrase, because these callers hold something better than a
|
||||
// gap: the passage itself, which their own floor reads back to him. Nothing is
|
||||
// invented either way, and a real quote beats "не могу сейчас".
|
||||
func (h *reactiveHandler) phraseSource(ctx context.Context, name, utterance string, sources []string) string {
|
||||
reply, err := h.phraseWorld(ctx, utterance, sources)
|
||||
switch {
|
||||
case errors.Is(err, phraser.ErrNoWorldModel):
|
||||
log.Printf("voice: %s: no world model, reading the source back instead", name)
|
||||
return ""
|
||||
case err != nil:
|
||||
log.Printf("voice: %s: phrase: %v", name, err)
|
||||
}
|
||||
return reply
|
||||
}
|
||||
@@ -0,0 +1,85 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/kami/maven/internal/phraser"
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
|
||||
// gapPhraser — a phraser whose world model is configured and asleep, which is
|
||||
// the state the naming half exists for.
|
||||
type gapPhraser struct {
|
||||
*phraser.Stub
|
||||
worldCalls int
|
||||
}
|
||||
|
||||
func (g *gapPhraser) PhraseWorld(context.Context, string, []string) (string, error) {
|
||||
g.worldCalls++
|
||||
return "", phraser.ErrNoWorldModel
|
||||
}
|
||||
|
||||
func worldTurn(utterance string) *queryTurn {
|
||||
return &queryTurn{dec: router.Decision{Intent: router.IntentQuery, Utterance: utterance}}
|
||||
}
|
||||
|
||||
// A world question with the workstation asleep says so. The resident model is
|
||||
// not asked, because what it produces here is an invention with no signal that
|
||||
// it is one.
|
||||
func TestQueryGeneralNamesTheGap(t *testing.T) {
|
||||
g := &gapPhraser{Stub: phraser.NewStub()}
|
||||
h := &reactiveHandler{phraser: g}
|
||||
reply, ok := h.queryGeneral(context.Background(), worldTurn("почему небо голубое"))
|
||||
if !ok {
|
||||
t.Fatal("queryGeneral passed on the last source in the chain")
|
||||
}
|
||||
if reply != worldGap {
|
||||
t.Fatalf("reply = %q, want the named gap", reply)
|
||||
}
|
||||
if g.worldCalls != 1 {
|
||||
t.Fatalf("PhraseWorld called %d times, want 1", g.worldCalls)
|
||||
}
|
||||
}
|
||||
|
||||
// A phraser with no world seam at all — the Stub, and every box with no
|
||||
// `workstation` block — answers exactly as it did before this seam existed.
|
||||
func TestQueryGeneralWithoutAWorldModelIsUnchanged(t *testing.T) {
|
||||
h := &reactiveHandler{phraser: phraser.NewStub()}
|
||||
reply, ok := h.queryGeneral(context.Background(), worldTurn("почему небо голубое"))
|
||||
if !ok {
|
||||
t.Fatal("queryGeneral passed on the last source in the chain")
|
||||
}
|
||||
if reply != "не знаю." {
|
||||
t.Fatalf("reply = %q, want the Stub's answer", reply)
|
||||
}
|
||||
}
|
||||
|
||||
// The gap is spoken aloud by a Russian voice, so it is Russian, feminine and
|
||||
// informal. "не хочу" and "не могу" are her own verbs; there is no "вы" and no
|
||||
// English in it.
|
||||
func TestWorldGapIsInPersona(t *testing.T) {
|
||||
for _, bad := range []string{"вы", "ваш", "рад ", "дорогой", "милый"} {
|
||||
if strings.Contains(worldGap, bad) {
|
||||
t.Errorf("the gap phrase contains %q: %s", bad, worldGap)
|
||||
}
|
||||
}
|
||||
if strings.ContainsAny(worldGap, "abcdefghijklmnopqrstuvwxyz") {
|
||||
t.Errorf("the gap phrase has Latin letters in it: %s", worldGap)
|
||||
}
|
||||
}
|
||||
|
||||
// The sources that hold a passage read it back rather than name a gap. He gets a
|
||||
// real quote instead of "не могу сейчас", and nothing is invented either way.
|
||||
func TestASourceWithAPassageReadsItBackInsteadOfNamingTheGap(t *testing.T) {
|
||||
g := &gapPhraser{Stub: phraser.NewStub()}
|
||||
h := &reactiveHandler{phraser: g}
|
||||
if got := h.phraseSource(context.Background(), "search", "почему небо голубое",
|
||||
[]string{"Рэлеевское рассеяние."}); got != "" {
|
||||
t.Fatalf("phraseSource = %q, want \"\" so the caller's own floor reads the passage back", got)
|
||||
}
|
||||
if g.worldCalls != 1 {
|
||||
t.Fatalf("PhraseWorld called %d times, want 1", g.worldCalls)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,117 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// The card is an AMD 7900 GRE with 16GB, driven by amdgpu and ROCm. Everything
|
||||
// here reads sysfs and forks nothing: rocm-smi is not even installed on the
|
||||
// workstation, and a poll that costs a subprocess every second is a poll that
|
||||
// gets tuned down until it is useless.
|
||||
|
||||
// gpuProc — one process holding the compute engine.
|
||||
type gpuProc struct {
|
||||
PID int
|
||||
Comm string
|
||||
VRAM int64 // bytes, as the kernel accounts them to this process
|
||||
}
|
||||
|
||||
// probe reads the two sysfs trees the supervisor decides from.
|
||||
//
|
||||
// kfdRoot is /sys/class/kfd/kfd/proc, one directory per ROCm process. The
|
||||
// directory appears when the process initialises HIP, which is well before it
|
||||
// allocates anything large. That is the whole reason this works: the job that
|
||||
// is about to want the card announces itself while it is still starting up,
|
||||
// so we see the contender rather than only the winner of an allocation race.
|
||||
//
|
||||
// drmDev is /sys/class/drm/cardN/device, which reports total and used VRAM for
|
||||
// the card as a whole.
|
||||
type probe struct {
|
||||
kfdRoot string
|
||||
drmDev string
|
||||
}
|
||||
|
||||
// foreign lists every ROCm process that is not ours. selfPID is the supervisor's
|
||||
// llama-server child, or 0 when it is not running.
|
||||
//
|
||||
// An unreadable kfd tree returns no processes and no error. That is deliberate
|
||||
// and it is the safe direction only because startVRAM also has to agree before
|
||||
// anything launches: a supervisor that cannot see the KFD never sees free VRAM
|
||||
// either, because the CPT run holding the card shows up in the drm totals.
|
||||
func (p probe) foreign(selfPID int) []gpuProc {
|
||||
entries, err := os.ReadDir(p.kfdRoot)
|
||||
if err != nil {
|
||||
return nil
|
||||
}
|
||||
var out []gpuProc
|
||||
for _, e := range entries {
|
||||
pid, err := strconv.Atoi(e.Name())
|
||||
if err != nil || pid == selfPID {
|
||||
continue
|
||||
}
|
||||
out = append(out, gpuProc{
|
||||
PID: pid,
|
||||
Comm: readComm(pid),
|
||||
VRAM: p.procVRAM(e.Name()),
|
||||
})
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// procVRAM sums the per-node vram_* files under one process directory. The
|
||||
// suffix is the KFD topology node id (vram_35881 on this card), so it is
|
||||
// globbed rather than named, and a machine with two cards sums both.
|
||||
func (p probe) procVRAM(pid string) int64 {
|
||||
matches, err := filepath.Glob(filepath.Join(p.kfdRoot, pid, "vram_*"))
|
||||
if err != nil {
|
||||
return 0
|
||||
}
|
||||
var total int64
|
||||
for _, m := range matches {
|
||||
total += readInt(m)
|
||||
}
|
||||
return total
|
||||
}
|
||||
|
||||
// freeVRAM reports the bytes the card has left. Used only to decide whether to
|
||||
// start: a shortfall here means llama-server would refuse to load anyway. It is
|
||||
// never used to decide to stop, because by the time free VRAM has dropped the
|
||||
// other job has already failed its allocation, which is exactly the outcome
|
||||
// yielding exists to prevent.
|
||||
func (p probe) freeVRAM() int64 {
|
||||
total := readInt(filepath.Join(p.drmDev, "mem_info_vram_total"))
|
||||
used := readInt(filepath.Join(p.drmDev, "mem_info_vram_used"))
|
||||
if total <= 0 {
|
||||
return 0
|
||||
}
|
||||
if free := total - used; free > 0 {
|
||||
return free
|
||||
}
|
||||
return 0
|
||||
}
|
||||
|
||||
func readInt(path string) int64 {
|
||||
b, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return 0
|
||||
}
|
||||
n, err := strconv.ParseInt(strings.TrimSpace(string(b)), 10, 64)
|
||||
if err != nil {
|
||||
return 0
|
||||
}
|
||||
return n
|
||||
}
|
||||
|
||||
// readComm names the contender for the log. The log is the instrument for the
|
||||
// open question in Vikunja #488: whether a process can want this card without
|
||||
// ever registering on the KFD, which a Vulkan or video-decode job would.
|
||||
func readComm(pid int) string {
|
||||
b, err := os.ReadFile(filepath.Join("/proc", strconv.Itoa(pid), "comm"))
|
||||
if err != nil {
|
||||
return "?"
|
||||
}
|
||||
return strings.TrimSpace(string(b))
|
||||
}
|
||||
@@ -0,0 +1,103 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"net/url"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// fakeKFD builds the sysfs shape the workstation actually has: one directory
|
||||
// per ROCm process, each holding a vram_<node> file. Sampled from the live box
|
||||
// on 02-08-2026, where the CPT run appeared as proc/478104/vram_35881.
|
||||
func fakeKFD(t *testing.T, vramByPID map[int]int64) string {
|
||||
t.Helper()
|
||||
root := t.TempDir()
|
||||
for pid, vram := range vramByPID {
|
||||
dir := filepath.Join(root, strconv.Itoa(pid))
|
||||
if err := os.MkdirAll(dir, 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
f := filepath.Join(dir, "vram_35881")
|
||||
if err := os.WriteFile(f, []byte(strconv.FormatInt(vram, 10)+"\n"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
return root
|
||||
}
|
||||
|
||||
func TestForeignExcludesOurChild(t *testing.T) {
|
||||
root := fakeKFD(t, map[int]int64{478104: 12791693312, 999: 4096})
|
||||
p := probe{kfdRoot: root}
|
||||
|
||||
all := p.foreign(0)
|
||||
if len(all) != 2 {
|
||||
t.Fatalf("with no child running, both processes are foreign, got %d", len(all))
|
||||
}
|
||||
|
||||
ours := p.foreign(999)
|
||||
if len(ours) != 1 || ours[0].PID != 478104 {
|
||||
t.Fatalf("our own llama-server must not count as a contender, got %+v", ours)
|
||||
}
|
||||
if ours[0].VRAM != 12791693312 {
|
||||
t.Errorf("per-process VRAM = %d, want the value from vram_35881", ours[0].VRAM)
|
||||
}
|
||||
}
|
||||
|
||||
// An empty KFD tree is the state that permits a start, so it must read as empty
|
||||
// rather than as an error the caller has to interpret.
|
||||
func TestForeignEmptyAndMissing(t *testing.T) {
|
||||
if got := (probe{kfdRoot: t.TempDir()}).foreign(0); len(got) != 0 {
|
||||
t.Errorf("empty kfd tree: got %d processes, want 0", len(got))
|
||||
}
|
||||
if got := (probe{kfdRoot: "/nonexistent"}).foreign(0); got != nil {
|
||||
t.Errorf("missing kfd tree: got %+v, want nil", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestFreeVRAM(t *testing.T) {
|
||||
dev := t.TempDir()
|
||||
write := func(name, v string) {
|
||||
if err := os.WriteFile(filepath.Join(dev, name), []byte(v), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
// The live numbers from the workstation while the CPT run held the card.
|
||||
write("mem_info_vram_total", "17163091968\n")
|
||||
write("mem_info_vram_used", "13396389888\n")
|
||||
p := probe{drmDev: dev}
|
||||
if got, want := p.freeVRAM(), int64(3766702080); got != want {
|
||||
t.Errorf("freeVRAM = %d, want %d", got, want)
|
||||
}
|
||||
if got := (probe{drmDev: "/nonexistent"}).freeVRAM(); got != 0 {
|
||||
t.Errorf("unreadable card reports %d free, want 0 so nothing starts", got)
|
||||
}
|
||||
}
|
||||
|
||||
// With no model loaded the supervisor must still answer, and it must answer 503
|
||||
// rather than hanging or proxying into a closed port. Maven reads this endpoint
|
||||
// on a timer forever, including while the workstation is busy.
|
||||
func TestHealthAndProxyRefuseWhenNotReady(t *testing.T) {
|
||||
s := &supervisor{run: newRunner("/bin/true", nil, "")}
|
||||
h := s.handler(mustURL(t, "http://127.0.0.1:1"))
|
||||
|
||||
for _, path := range []string{"/health", "/v1/chat/completions"} {
|
||||
w := httptest.NewRecorder()
|
||||
h.ServeHTTP(w, httptest.NewRequest(http.MethodGet, path, nil))
|
||||
if w.Code != http.StatusServiceUnavailable {
|
||||
t.Errorf("%s with no model: got %d, want 503", path, w.Code)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func mustURL(t *testing.T, s string) *url.URL {
|
||||
t.Helper()
|
||||
u, err := url.Parse(s)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return u
|
||||
}
|
||||
@@ -0,0 +1,247 @@
|
||||
// mavgpud — the workstation's GPU supervisor.
|
||||
//
|
||||
// It runs on the workstation (an AMD 7900 GRE, 16GB), not on homesrv, and it is
|
||||
// deployed separately from the Maven daemons. Maven does not participate in any
|
||||
// of this and never asks for a start: it reads /health through internal/llm.Pair
|
||||
// and either gets the big model or falls back to the resident 1.7B.
|
||||
//
|
||||
// The rule, from Vikunja #488: keep llama-server loaded whenever the card is
|
||||
// free, unload it when it has been idle too long or when another process needs
|
||||
// the card. Not on demand, because a 7-14B takes tens of seconds to load and a
|
||||
// world question would be answered by a gap every time the card had been quiet.
|
||||
// Not always on, because that holds 16GB against the owner's own jobs.
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"flag"
|
||||
"log"
|
||||
"net/http"
|
||||
"net/http/httputil"
|
||||
"net/url"
|
||||
"os"
|
||||
"os/signal"
|
||||
"sync/atomic"
|
||||
"syscall"
|
||||
"time"
|
||||
)
|
||||
|
||||
type config struct {
|
||||
Listen string `json:"listen"` // what Maven talks to
|
||||
LlamaAddr string `json:"llama_addr"` // where llama-server binds
|
||||
LlamaBin string `json:"llama_bin"`
|
||||
// LlamaArgs must include the flags that bind LlamaAddr. They are passed
|
||||
// through untouched so the model, context size and layer count stay the
|
||||
// owner's business and not this daemon's schema.
|
||||
LlamaArgs []string `json:"llama_args"`
|
||||
|
||||
KFDRoot string `json:"kfd_root"`
|
||||
DRMDevice string `json:"drm_device"`
|
||||
|
||||
Poll duration `json:"poll"`
|
||||
IdleTimeout duration `json:"idle_timeout"`
|
||||
StopGrace duration `json:"stop_grace"`
|
||||
MinFreeVRAM int64 `json:"min_free_vram_bytes"`
|
||||
// EvictAfter and StartAfter are counted in polls, not seconds. Both exist
|
||||
// to damp flapping: a one-tick blip from a short-lived rocm process must
|
||||
// not evict the model, and a card that has just been released must not be
|
||||
// grabbed before the previous job has finished unmapping.
|
||||
EvictAfter int `json:"evict_after_polls"`
|
||||
StartAfter int `json:"start_after_polls"`
|
||||
}
|
||||
|
||||
func defaults() config {
|
||||
return config{
|
||||
Listen: ":8080",
|
||||
LlamaAddr: "127.0.0.1:8081",
|
||||
KFDRoot: "/sys/class/kfd/kfd/proc",
|
||||
DRMDevice: "/sys/class/drm/card1/device",
|
||||
Poll: duration(time.Second),
|
||||
IdleTimeout: duration(15 * time.Minute),
|
||||
StopGrace: duration(20 * time.Second),
|
||||
MinFreeVRAM: 15 << 30,
|
||||
EvictAfter: 2,
|
||||
StartAfter: 5,
|
||||
}
|
||||
}
|
||||
|
||||
// duration lets the config file say "15m" instead of counting nanoseconds.
|
||||
type duration time.Duration
|
||||
|
||||
func (d *duration) UnmarshalJSON(b []byte) error {
|
||||
var s string
|
||||
if err := json.Unmarshal(b, &s); err != nil {
|
||||
return err
|
||||
}
|
||||
v, err := time.ParseDuration(s)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
*d = duration(v)
|
||||
return nil
|
||||
}
|
||||
|
||||
func main() {
|
||||
path := flag.String("config", "/etc/mavgpud.json", "config file")
|
||||
flag.Parse()
|
||||
|
||||
cfg := defaults()
|
||||
b, err := os.ReadFile(*path)
|
||||
if err != nil {
|
||||
log.Fatalf("mavgpud: read config: %v", err)
|
||||
}
|
||||
if err := json.Unmarshal(b, &cfg); err != nil {
|
||||
log.Fatalf("mavgpud: parse config: %v", err)
|
||||
}
|
||||
if cfg.LlamaBin == "" {
|
||||
log.Fatal("mavgpud: llama_bin is required")
|
||||
}
|
||||
|
||||
base := "http://" + cfg.LlamaAddr
|
||||
run := newRunner(cfg.LlamaBin, cfg.LlamaArgs, base+"/health")
|
||||
sup := &supervisor{
|
||||
cfg: cfg,
|
||||
probe: probe{kfdRoot: cfg.KFDRoot, drmDev: cfg.DRMDevice},
|
||||
run: run,
|
||||
}
|
||||
sup.touch()
|
||||
|
||||
ctx, cancel := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
|
||||
defer cancel()
|
||||
|
||||
target, err := url.Parse(base)
|
||||
if err != nil {
|
||||
log.Fatalf("mavgpud: llama_addr: %v", err)
|
||||
}
|
||||
srv := &http.Server{Addr: cfg.Listen, Handler: sup.handler(target)}
|
||||
go func() {
|
||||
log.Printf("mavgpud: listening on %s, model %s", cfg.Listen, cfg.LlamaBin)
|
||||
if err := srv.ListenAndServe(); err != nil && err != http.ErrServerClosed {
|
||||
log.Fatalf("mavgpud: listen: %v", err)
|
||||
}
|
||||
}()
|
||||
|
||||
sup.loop(ctx)
|
||||
|
||||
// The card must come back before we do. A supervisor that exits leaving
|
||||
// llama-server holding 14GB is worse than one that never ran.
|
||||
shut, done := context.WithTimeout(context.Background(), 5*time.Second)
|
||||
defer done()
|
||||
_ = srv.Shutdown(shut)
|
||||
run.stop(time.Duration(cfg.StopGrace))
|
||||
}
|
||||
|
||||
type supervisor struct {
|
||||
cfg config
|
||||
probe probe
|
||||
run *runner
|
||||
|
||||
lastReq atomic.Int64 // unix nanos of the last request Maven sent
|
||||
|
||||
foreignStreak int
|
||||
clearStreak int
|
||||
}
|
||||
|
||||
func (s *supervisor) touch() { s.lastReq.Store(time.Now().UnixNano()) }
|
||||
|
||||
func (s *supervisor) idle() time.Duration {
|
||||
return time.Since(time.Unix(0, s.lastReq.Load()))
|
||||
}
|
||||
|
||||
// handler serves the two things the workstation exposes.
|
||||
//
|
||||
// /health is answered locally and always, with no GPU cost and no round trip,
|
||||
// because it is the only thing Maven reads and Maven reads it on a timer
|
||||
// forever. Everything else is llama-server's API, reverse-proxied. Proxying
|
||||
// rather than pointing Maven straight at llama-server is what makes the idle
|
||||
// window measurable: the supervisor cannot otherwise know when the model was
|
||||
// last used.
|
||||
func (s *supervisor) handler(target *url.URL) http.Handler {
|
||||
proxy := httputil.NewSingleHostReverseProxy(target)
|
||||
mux := http.NewServeMux()
|
||||
mux.HandleFunc("/health", func(w http.ResponseWriter, r *http.Request) {
|
||||
if !s.run.isReady() {
|
||||
http.Error(w, "model not loaded", http.StatusServiceUnavailable)
|
||||
return
|
||||
}
|
||||
w.Header().Set("Content-Type", "application/json")
|
||||
_, _ = w.Write([]byte(`{"status":"ok"}`))
|
||||
})
|
||||
mux.HandleFunc("/", func(w http.ResponseWriter, r *http.Request) {
|
||||
if !s.run.isReady() {
|
||||
http.Error(w, "model not loaded", http.StatusServiceUnavailable)
|
||||
return
|
||||
}
|
||||
s.touch()
|
||||
proxy.ServeHTTP(w, r)
|
||||
})
|
||||
return mux
|
||||
}
|
||||
|
||||
func (s *supervisor) loop(ctx context.Context) {
|
||||
t := time.NewTicker(time.Duration(s.cfg.Poll))
|
||||
defer t.Stop()
|
||||
for {
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return
|
||||
case <-t.C:
|
||||
s.tick(ctx)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// tick is the whole decision. Yielding is checked before starting, and presence
|
||||
// on the KFD is what triggers it — not a VRAM threshold. A ROCm process
|
||||
// registers under /sys/class/kfd/kfd/proc when it initialises HIP, before it
|
||||
// allocates, so we see a contender during its startup rather than after it has
|
||||
// already failed to get the memory it wanted.
|
||||
func (s *supervisor) tick(ctx context.Context) {
|
||||
others := s.probe.foreign(s.run.pid())
|
||||
if len(others) > 0 {
|
||||
s.foreignStreak++
|
||||
s.clearStreak = 0
|
||||
} else {
|
||||
s.foreignStreak = 0
|
||||
s.clearStreak++
|
||||
}
|
||||
|
||||
if s.run.running() {
|
||||
s.run.refreshReady(ctx)
|
||||
switch {
|
||||
case s.foreignStreak >= s.cfg.EvictAfter:
|
||||
log.Printf("mavgpud: yielding the card to %s", describe(others))
|
||||
s.run.stop(time.Duration(s.cfg.StopGrace))
|
||||
case s.idle() > time.Duration(s.cfg.IdleTimeout):
|
||||
log.Printf("mavgpud: idle for %s, unloading", s.idle().Round(time.Second))
|
||||
s.run.stop(time.Duration(s.cfg.StopGrace))
|
||||
}
|
||||
return
|
||||
}
|
||||
|
||||
if s.clearStreak < s.cfg.StartAfter {
|
||||
return
|
||||
}
|
||||
if free := s.probe.freeVRAM(); free < s.cfg.MinFreeVRAM {
|
||||
return
|
||||
}
|
||||
s.touch() // the idle clock starts at load, not at the last request before it
|
||||
if err := s.run.start(); err != nil {
|
||||
log.Printf("mavgpud: start llama-server: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// describe names the contenders in the log. This log is the instrument for the
|
||||
// open question in #488: whether polling the KFD misses a job that wants the
|
||||
// card without registering there.
|
||||
func describe(procs []gpuProc) string {
|
||||
out := ""
|
||||
for i, p := range procs {
|
||||
if i > 0 {
|
||||
out += ", "
|
||||
}
|
||||
out += p.Comm
|
||||
}
|
||||
return out
|
||||
}
|
||||
@@ -0,0 +1,132 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"log"
|
||||
"net/http"
|
||||
"os/exec"
|
||||
"sync"
|
||||
"syscall"
|
||||
"time"
|
||||
)
|
||||
|
||||
// runner owns one llama-server process. Owning it is the point of the daemon:
|
||||
// the workstation cannot keep a 7-14B resident, because that holds 16GB against
|
||||
// the owner's CPT runs, Correx and the manga-recap pipeline. So the thing that
|
||||
// stays up is this, which costs no VRAM, and the model comes and goes under it.
|
||||
type runner struct {
|
||||
bin string
|
||||
args []string
|
||||
// ready is llama-server's own /health, which answers "is a model loaded".
|
||||
// Loading a 7-14B takes tens of seconds, so started is not ready.
|
||||
readyURL string
|
||||
|
||||
mu sync.Mutex
|
||||
cmd *exec.Cmd
|
||||
ready bool
|
||||
http *http.Client
|
||||
}
|
||||
|
||||
func newRunner(bin string, args []string, readyURL string) *runner {
|
||||
return &runner{
|
||||
bin: bin, args: args, readyURL: readyURL,
|
||||
http: &http.Client{Timeout: 2 * time.Second},
|
||||
}
|
||||
}
|
||||
|
||||
// pid is the child's, or 0. The GPU probe needs it to tell our own model apart
|
||||
// from a contender.
|
||||
func (r *runner) pid() int {
|
||||
r.mu.Lock()
|
||||
defer r.mu.Unlock()
|
||||
if r.cmd == nil || r.cmd.Process == nil {
|
||||
return 0
|
||||
}
|
||||
return r.cmd.Process.Pid
|
||||
}
|
||||
|
||||
func (r *runner) running() bool { return r.pid() != 0 }
|
||||
|
||||
// isReady reports the cached readiness. The supervisor loop refreshes it; the
|
||||
// health handler only reads, so answering /health never costs a round trip.
|
||||
func (r *runner) isReady() bool {
|
||||
r.mu.Lock()
|
||||
defer r.mu.Unlock()
|
||||
return r.ready
|
||||
}
|
||||
|
||||
// start launches llama-server. It returns as soon as the process exists, not
|
||||
// when the model is loaded.
|
||||
func (r *runner) start() error {
|
||||
r.mu.Lock()
|
||||
defer r.mu.Unlock()
|
||||
if r.cmd != nil {
|
||||
return nil
|
||||
}
|
||||
cmd := exec.Command(r.bin, r.args...)
|
||||
// Own process group, so stop kills anything llama-server spawned rather
|
||||
// than leaving it holding VRAM after we have declared the card yielded.
|
||||
cmd.SysProcAttr = &syscall.SysProcAttr{Setpgid: true}
|
||||
if err := cmd.Start(); err != nil {
|
||||
return err
|
||||
}
|
||||
r.cmd, r.ready = cmd, false
|
||||
log.Printf("mavgpud: started llama-server pid=%d", cmd.Process.Pid)
|
||||
go func() {
|
||||
err := cmd.Wait()
|
||||
r.mu.Lock()
|
||||
r.cmd, r.ready = nil, false
|
||||
r.mu.Unlock()
|
||||
log.Printf("mavgpud: llama-server exited: %v", err)
|
||||
}()
|
||||
return nil
|
||||
}
|
||||
|
||||
// stop ends llama-server and waits for the VRAM to come back. SIGTERM first so
|
||||
// it unmaps cleanly, SIGKILL after the grace window. Returning before the
|
||||
// process is gone would let the supervisor report a free card while 14GB is
|
||||
// still mapped, which is the one lie that would make yielding useless.
|
||||
func (r *runner) stop(grace time.Duration) {
|
||||
r.mu.Lock()
|
||||
cmd := r.cmd
|
||||
r.ready = false
|
||||
r.mu.Unlock()
|
||||
if cmd == nil || cmd.Process == nil {
|
||||
return
|
||||
}
|
||||
pgid := -cmd.Process.Pid
|
||||
_ = syscall.Kill(pgid, syscall.SIGTERM)
|
||||
deadline := time.Now().Add(grace)
|
||||
for time.Now().Before(deadline) {
|
||||
if !r.running() {
|
||||
return
|
||||
}
|
||||
time.Sleep(100 * time.Millisecond)
|
||||
}
|
||||
log.Printf("mavgpud: llama-server did not exit in %s, killing", grace)
|
||||
_ = syscall.Kill(pgid, syscall.SIGKILL)
|
||||
}
|
||||
|
||||
// refreshReady asks llama-server whether the model is loaded. Called once per
|
||||
// supervisor tick, never per request.
|
||||
func (r *runner) refreshReady(ctx context.Context) {
|
||||
if !r.running() {
|
||||
return
|
||||
}
|
||||
ok := false
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodGet, r.readyURL, nil)
|
||||
if err == nil {
|
||||
resp, err := r.http.Do(req)
|
||||
if err == nil {
|
||||
ok = resp.StatusCode == http.StatusOK
|
||||
resp.Body.Close()
|
||||
}
|
||||
}
|
||||
r.mu.Lock()
|
||||
was := r.ready
|
||||
r.ready = ok
|
||||
r.mu.Unlock()
|
||||
if ok && !was {
|
||||
log.Printf("mavgpud: model ready")
|
||||
}
|
||||
}
|
||||
+5
-8
@@ -512,7 +512,7 @@ func main() {
|
||||
|
||||
// /tools — the authed enable surface. maven proposes acts she can't run;
|
||||
// this page is where a human reviews and enables them (proposed→enabled).
|
||||
// Enabling is the boundary-moving act (DESIGN.md § Tool registration —
|
||||
// Enabling is the boundary-moving act (docs/design.md § Tool registration —
|
||||
// drafting is suggest, enabling is act), so it lives ONLY here,
|
||||
// behind wg+nginx+auth — never the voice/chat path.
|
||||
mux.HandleFunc("/tools", func(w http.ResponseWriter, r *http.Request) {
|
||||
@@ -1210,13 +1210,10 @@ func routineRows(rs []ipc.ProposedRoutine) []routineRow {
|
||||
return out
|
||||
}
|
||||
|
||||
// acceptRoutine creates the recurring reminder for a proposal, then marks the
|
||||
// proposal accepted and links the reminder to it. Weekly patterns get a cron
|
||||
// expression; any other interval fires once.
|
||||
//
|
||||
// TODO(vikunja#46): this mirrors the voice accept path in cmd/mavend/voice.go.
|
||||
// When the tick loop learns to read accepted proposals directly, both callers
|
||||
// should hand off to one place in core instead of each building a reminder.
|
||||
// acceptRoutine marks a proposal accepted. This page is the ONLY surface that
|
||||
// may do it (Vikunja #367): accepting gives the tick loop a standing new
|
||||
// reason to speak, which DESIGN.md puts at layer 3, and the button here is
|
||||
// behind step-up. Voice can park the question and dismiss, never accept.
|
||||
func acceptRoutine(ctx context.Context, core ipc.CoreAPI, id int64) error {
|
||||
proposed, err := core.ListProposedRoutines(ctx)
|
||||
if err != nil {
|
||||
|
||||
+68
-1
@@ -39,6 +39,23 @@
|
||||
"proxy": "socks5://192.168.240.1:10808"
|
||||
},
|
||||
|
||||
"//workstation": [
|
||||
"The big model on the desk PC (workpc, 7900 GRE 16GB), fronted by",
|
||||
"mavgpud on port 8080. It runs gemma-4-12b and it is preferred over the",
|
||||
"resident Qwen3-1.7B for routing and replies whenever the card is free.",
|
||||
"The machine is never assumed up: it sleeps, and the card is often held by",
|
||||
"a CPT run, in which case mavgpud answers 503 and Maven falls back to the",
|
||||
"resident model without saying so. Deleting this block restores exactly",
|
||||
"the behaviour homesrv had before it existed.",
|
||||
"Addressed by LAN address, not container name: mavgpud runs on another",
|
||||
"machine and there is no shared docker network to name it on."
|
||||
],
|
||||
"workstation": {
|
||||
"url": "http://192.168.1.105:8080",
|
||||
"probe": "15s",
|
||||
"timeout": "90s"
|
||||
},
|
||||
|
||||
"//search": [
|
||||
"The live web, searched after his own notes and before Kiwix. Only the",
|
||||
"query string leaves the box — never a note, a fact, the persona block or",
|
||||
@@ -76,6 +93,56 @@
|
||||
"snippet_runes": 1500
|
||||
},
|
||||
|
||||
"//morning_routines": [
|
||||
"The daily checklist (Vikunja #280). Each item is done when its fact_key",
|
||||
"gets a non-voided fact inside the window, so 'выпил воды' closes water and",
|
||||
"nothing has to be ticked by hand. nudge_at fires once, at the end of the",
|
||||
"window, and only for what is still open. Weekdays empty = every day."
|
||||
],
|
||||
"morning_routines": [
|
||||
{
|
||||
"name": "утро",
|
||||
"window_start": "08:00",
|
||||
"window_end": "11:00",
|
||||
"nudge_at": "10:30",
|
||||
"severity": 1,
|
||||
"items": [
|
||||
{ "key": "medicine", "fact_key": "medicine", "label": "лекарство" },
|
||||
{ "key": "water", "fact_key": "water", "label": "вода" },
|
||||
{ "key": "pets", "fact_key": "pets", "label": "покормить кота" }
|
||||
]
|
||||
}
|
||||
],
|
||||
|
||||
"//feeds": [
|
||||
"RSS reading (Vikunja #258). Every item lands as a note with source",
|
||||
"rss:<name>, which is also what puts entries in the intake journal that",
|
||||
"/events reads. Only the feed URL leaves the box.",
|
||||
"This is a starting pair, not a curated set — trim or extend it."
|
||||
],
|
||||
"feeds": {
|
||||
"poll_interval": "30m",
|
||||
"max_items": 5,
|
||||
"max_age": "24h",
|
||||
"sources": [
|
||||
{ "name": "lwn", "url": "https://lwn.net/headlines/newrss", "category": "технологии" },
|
||||
{ "name": "archlinux", "url": "https://archlinux.org/feeds/news/", "category": "технологии" }
|
||||
]
|
||||
},
|
||||
|
||||
"//crawl": [
|
||||
"Reading a web page (Vikunja #259). on_demand answers 'посмотри <URL>'.",
|
||||
"No allow_hosts, so any public host he names is readable; private",
|
||||
"addresses are refused unconditionally by internal/webfetch and do not",
|
||||
"need listing. Setting allow_hosts here would also narrow on-demand,",
|
||||
"which is the point of leaving it empty."
|
||||
],
|
||||
"crawl": {
|
||||
"on_demand": true,
|
||||
"timeout": "10s",
|
||||
"max_runes": 4000
|
||||
},
|
||||
|
||||
"digest": {
|
||||
"enabled": true,
|
||||
"window": "30m",
|
||||
@@ -119,7 +186,7 @@
|
||||
"timeout": "400ms",
|
||||
"rate": 100,
|
||||
"max_hosts": 256,
|
||||
"enabled": false
|
||||
"enabled": true
|
||||
},
|
||||
|
||||
"nexus": { "url": "http://nexus:9740" },
|
||||
|
||||
@@ -0,0 +1,32 @@
|
||||
{
|
||||
"listen": ":8080",
|
||||
"llama_addr": "127.0.0.1:10000",
|
||||
"llama_bin": "llama-server",
|
||||
"llama_args": [
|
||||
"-m", "/mnt/D/AI/gemma4/gemma-4-12B-it-qat-UD-Q4_K_XL.gguf",
|
||||
"-md", "/mnt/D/AI/gemma4/mtp-gemma-4-12B-it-BF16.gguf",
|
||||
"-ngl", "99",
|
||||
"-fa", "on",
|
||||
"-np", "1",
|
||||
"--host", "127.0.0.1",
|
||||
"--port", "10000",
|
||||
"--ctx-size", "32768",
|
||||
"--threads", "6",
|
||||
"--batch-size", "2048",
|
||||
"--ubatch-size", "512",
|
||||
"--jinja",
|
||||
"--chat-template-kwargs", "{\"enable_thinking\":false}",
|
||||
"--spec-type", "draft-mtp",
|
||||
"--spec-draft-n-max", "2"
|
||||
],
|
||||
|
||||
"kfd_root": "/sys/class/kfd/kfd/proc",
|
||||
"drm_device": "/sys/class/drm/card1/device",
|
||||
|
||||
"poll": "1s",
|
||||
"idle_timeout": "15m",
|
||||
"stop_grace": "20s",
|
||||
"min_free_vram_bytes": 10737418240,
|
||||
"evict_after_polls": 2,
|
||||
"start_after_polls": 5
|
||||
}
|
||||
@@ -0,0 +1,24 @@
|
||||
[Unit]
|
||||
# Runs on the workstation (bugmachine), not on homesrv. Install as a systemd
|
||||
# user unit and turn on lingering, so the card is supervised after a reboot
|
||||
# with nobody logged in:
|
||||
#
|
||||
# scp mavgpud workpc:~/.local/bin/mavgpud
|
||||
# scp deploy/mavgpud.json workpc:~/.config/mavgpud.json
|
||||
# scp deploy/mavgpud.service workpc:~/.config/systemd/user/mavgpud.service
|
||||
# ssh workpc 'systemctl --user daemon-reload && systemctl --user enable --now mavgpud'
|
||||
# sudo loginctl enable-linger kami
|
||||
Description=Maven GPU supervisor (holds llama-server while the card is free)
|
||||
After=network.target
|
||||
|
||||
[Service]
|
||||
ExecStart=%h/.local/bin/mavgpud -config %h/.config/mavgpud.json
|
||||
Restart=always
|
||||
RestartSec=5
|
||||
# The card must come back when the supervisor goes down. mavgpud stops
|
||||
# llama-server on SIGTERM, so give it longer than stop_grace to do that.
|
||||
KillSignal=SIGTERM
|
||||
TimeoutStopSec=60
|
||||
|
||||
[Install]
|
||||
WantedBy=default.target
|
||||
@@ -79,7 +79,16 @@ services:
|
||||
<<: *image
|
||||
# voice.bind is 0.0.0.0:9100 in deploy/mavend.json so mavweb can reach it
|
||||
# cross-container. Verified 2026-07-06.
|
||||
# -ambient-token turns on POST /api/ambient (Vikunja #126): the phone posts
|
||||
# notification text, mavweb keeps only a meeting time. Empty ⇒ no route at
|
||||
# all, which is what a missing MAVEN_AMBIENT_TOKEN gives. The value comes
|
||||
# from the gitignored .env docker compose reads for interpolation, NOT from
|
||||
# an env_file — flags are interpolated before any service env exists.
|
||||
# Weakness worth naming: mavweb takes this as a flag, so it is visible in
|
||||
# `ps` inside this container, unlike the zenmoney and IMAP secrets which are
|
||||
# read from files.
|
||||
command: ["mavweb", "-addr", ":9201", "-voice", "mavend:9100", "-core", "/run/maven/mavend.sock",
|
||||
"-ambient-token", "${MAVEN_AMBIENT_TOKEN:-}",
|
||||
"-nexus", "http://nexus:9740", "-praxis", "http://praxis:8989", "-hexis", "http://hexis:9741"]
|
||||
depends_on: [mavend]
|
||||
# loopback-only on purpose: /tools defines+executes arbitrary argv and
|
||||
|
||||
@@ -1,6 +1,45 @@
|
||||
# maven — feature ranking
|
||||
|
||||
> dated 2026-07-03. companion to `DESIGN.md` (folded from the former `maven.md`). ranks everything discussed post-repo-state against the infra blockers, not a replacement for the build order.
|
||||
> **Archived 2026-08-02 (V-447).** The mandatory and easy tiers are now Vikunja tasks
|
||||
> 449-458. Two of those closed immediately, because the ranking was stale. Quiet hours
|
||||
> (V-450) ship as `QuietHoursConfig` plus the care gate in `internal/loop/loop.go`.
|
||||
> Schema migrations (V-451) ship as `internal/store/migrations.go` on `PRAGMA
|
||||
> user_version`. The doable and epic tiers stay here because they are reasoning. Some of
|
||||
> them exist only to record why something is not worth doing yet. Read this for the why,
|
||||
> not as a work queue, and check the code before believing a gap.
|
||||
|
||||
> dated 2026-07-03. companion to `docs/design.md` (folded from the former `maven.md`). ranks everything discussed post-repo-state against the infra blockers, not a replacement for the build order.
|
||||
|
||||
---
|
||||
|
||||
## what already shipped (checked against the code, 2026-08-02)
|
||||
|
||||
One month old and already wrong in ten places. Everything below is marked
|
||||
built after reading the code, not the board. Read the tiers underneath with this list in
|
||||
hand.
|
||||
|
||||
Infra 1 and 2, the two the ranking says block every feature, are both done. sqlcipher
|
||||
at-rest ships as `Store.enc` plus `OpenEncrypted`, a tmpfs working copy re-encrypted on
|
||||
`Close`, keyed from `db_key_env` in `deploy/mavend.json`. mavweb and mavcaldav are no
|
||||
longer at zero coverage: seven test files under `cmd/mavweb`, including
|
||||
`credentials_test.go` and `passkey_prf_test.go`, and two under `cmd/mavcaldav`. Infra 3
|
||||
is stale in the other direction. There are still no systemd units, but the deploy is
|
||||
`deploy/ecosystem/docker-compose.yml`, not scripts and tmux.
|
||||
|
||||
Doable tier, built: rule trace and explanation as `/trace` plus `internal/loop/explain.go`.
|
||||
Recurring reminders as the cron column, `NextFireTs` and `RescheduleReminder`
|
||||
(`internal/store/reminders.go:206`), so "fires once right now" is wrong. Stale-reminder
|
||||
burst collapse as `collapseReminders` (`internal/loop/gather.go:209`). Revert as
|
||||
`VoidLatestFact` (`internal/store/facts.go:300`). Digest mode as `internal/store/digest.go`.
|
||||
Testing infra as the simulator and the eval lab (V-284, V-278). Passkey persistence as the
|
||||
JSON-backed `credentialStore` in `cmd/mavweb/credentials.go`. Most integrations shipped as
|
||||
their own QA tasks (V-246 mail, V-256 smarthome, V-258 rss, V-259 crawler).
|
||||
|
||||
Doable tier, still open: correx, systemd units, memory decay and duplicate detection
|
||||
(nothing in `internal/memory` touches it), backup automation, import and export, barge-in.
|
||||
|
||||
Epic tier, unbuilt as ranked. `event.Bus` exists (`cmd/mavend/intake.go:57`) but it is the
|
||||
intake journal from V-283, not the rewrite of facts into projections that this tier means.
|
||||
|
||||
---
|
||||
|
||||
@@ -20,7 +59,7 @@ nothing feature-level below should land before 1–2 are done. 3–4 can interle
|
||||
### mandatory
|
||||
things that block correctness or safety of stuff already shipped — not new capability, just closing gaps in existing design.
|
||||
|
||||
- **destructive-confirm policy** — open question in `DESIGN.md` § open questions, blocks correx and any new tool domain from having a coherent risk tier
|
||||
- **destructive-confirm policy** — open question in `docs/design.md` § open questions, blocks correx and any new tool domain from having a coherent risk tier
|
||||
- **quiet-hours definition** — open question, blocks proactive delivery being trustworthy
|
||||
- **schema migrations** — sqlcipher rollout alone forces a schema touch. want this mechanism before that, not after.
|
||||
|
||||
@@ -30,7 +69,7 @@ cheap, no dependencies, no new invariants.
|
||||
- **grocery / `list_items` table** — fourth append-only shape (item, status, list-tag), no predicate touches it, multi-adder just works for free
|
||||
- **go.mod tidy**
|
||||
- **capability model** (deepseek) — `homelab.docker.restart` instead of flat `tool→enabled`. cheap now, expensive to retrofit once tools surface passes ~15 entries. time-sensitive, not urgent.
|
||||
- **conversation repair** — already free: `DESIGN.md` has "misroute correction = new centroid example," this is just naming the existing mechanism as a feature
|
||||
- **conversation repair** — already free: `docs/design.md` has "misroute correction = new centroid example," this is just naming the existing mechanism as a feature
|
||||
- **command history** — read-only query over existing facts, no new mechanism
|
||||
- **clarification templates** — canned phrasing for the router's existing confidence-gate fallback, phraser-lane only
|
||||
- **pronunciation dictionary** — tts config, no architecture
|
||||
@@ -30,7 +30,7 @@ critical workflows: voice turn (mic→STT→route→tool/reply→TTS); proactive
|
||||
(/dash /chat /tools /ecosystem)
|
||||
current state: all 34 test packages pass; vet clean; `make test` still exits 1
|
||||
(finding 2)
|
||||
known failures: weak RU query routing — REARCH.md names the cause; the named fix
|
||||
known failures: weak RU query routing — docs/rearchitecture.md names the cause; the named fix
|
||||
is wired `nil`
|
||||
maintenance burden: 4,518 lines of root markdown vs 33,319 lines of Go; 15 top-level
|
||||
.md files, 3 of them dated session logs; several contradict
|
||||
@@ -47,11 +47,11 @@ what's obsolete: llmrouter.go (built, tested, never wired); classifier seed-p
|
||||
```
|
||||
|
||||
**Classification: healthy + misaligned.** Not fragile, not overbuilt, not abandoned. The
|
||||
architecture in `REARCH.md` is sound and mostly *built* — it just is not *connected*.
|
||||
architecture in `docs/rearchitecture.md` is sound and mostly *built* — it just is not *connected*.
|
||||
|
||||
## what it should become
|
||||
|
||||
The thing `REARCH.md` already describes, with the switch flipped and the drift removed:
|
||||
The thing `docs/rearchitecture.md` already describes, with the switch flipped and the drift removed:
|
||||
one resident small model doing both routing and phrasing, classifier demoted from the live
|
||||
path to the failure floor, embedder demoted to RAG hint. No new architecture is needed.
|
||||
**The gap is a config/wiring decision plus doc convergence, not a redesign.**
|
||||
@@ -65,19 +65,19 @@ path to the failure floor, embedder demoted to RAG hint. No new architecture is
|
||||
`architecture` / `repair`
|
||||
|
||||
**problem:** The most load-bearing design decision in the project is stated four different,
|
||||
incompatible ways, and the code path `REARCH.md` calls "the linchpin" is disabled.
|
||||
incompatible ways, and the code path `docs/rearchitecture.md` calls "the linchpin" is disabled.
|
||||
|
||||
**evidence** (all confirmed):
|
||||
|
||||
- `cmd/mavend/voice.go:211` — `rtr := buildRouter(emb, matcher, threshold, nil) // LLM router disabled`,
|
||||
with comment *"the classifier handles routing reliably."*
|
||||
- `REARCH.md:11` says the same classifier is *"the structural cause of 'she messes up
|
||||
- `docs/rearchitecture.md:11` says the same classifier is *"the structural cause of 'she messes up
|
||||
queries.'"* **The code comment and the design doc make opposite claims about the same
|
||||
component.**
|
||||
- `internal/router/llmrouter.go` (139 lines) + `llmrouter_test.go` — fully built and
|
||||
tested, zero non-test callers.
|
||||
- Model identity, four ways: docs say **Qwen3-1.7B** (`CLAUDE.md:6`, `REARCH.md:15`,
|
||||
`SPEC.md:46`, `AGENTS.md:79`, `MAVEN_ECOSYSTEM_ARCHITECTURE.md:72`);
|
||||
- Model identity, four ways: docs say **Qwen3-1.7B** (`CLAUDE.md:6`, `docs/rearchitecture.md:15`,
|
||||
`SPEC.md:46`, `AGENTS.md:79`, `docs/ecosystem.md:72`);
|
||||
`deploy/mavend.json:9` says **Qwen3.5-2B-UD-Q4_K_XL**; `models/llm/` on disk holds
|
||||
**LFM2.5-1.2B-Thinking**; code comments in 5 files still say **LFM**.
|
||||
- `deploy/mavend.json:11` sets `"n_gpu_layers": 99` while `CLAUDE.md:4` states the target
|
||||
@@ -100,7 +100,7 @@ match. Delete nothing from `internal/router` yet — the classifier is the fallb
|
||||
reconciliation, not a refactor. **Do not rewrite the router.**
|
||||
|
||||
**alternatives:** Delete `llmrouter.go` and commit to the classifier — only defensible if
|
||||
the eval harness shows the classifier is actually adequate, which contradicts `REARCH.md`.
|
||||
the eval harness shows the classifier is actually adequate, which contradicts `docs/rearchitecture.md`.
|
||||
|
||||
**risk:** Low-moderate. LLM route failures already fall through to the classifier
|
||||
(`router.go:88-96`), so a bad model cannot break a turn. The real risk is CPU latency.
|
||||
@@ -228,21 +228,21 @@ after finding 1, not before** — and skip it if it stays purely cosmetic.
|
||||
**problem:** 15 root markdown files, 4,518 lines, several stale or superseded, at least
|
||||
three pairs contradicting each other.
|
||||
|
||||
**evidence:** `ROADMAP.md` (759) + `MAVEN_ECOSYSTEM_ARCHITECTURE.md` (884) +
|
||||
**evidence:** `ROADMAP.md` (759) + `docs/ecosystem.md` (884) +
|
||||
`PROGRESS.md` (456) + `maven.md` (413) + `20-07-2026-BACKLOG.md` (396) +
|
||||
`SESSION-05-07-2026.md` + `SESSION-06-07-2026.md` (477 combined) + `PLANS.md` (25) +
|
||||
`START.md` + `SPEC.md` + `PROTOCOL.md`. `PROGRESS.md:61` annotates its own staleness:
|
||||
*"Older LFM references below describe the currently deployed..."*. `REARCH.md` announces it
|
||||
`docs/operations.md` + `SPEC.md` + `docs/protocol.md`. `PROGRESS.md:61` annotates its own staleness:
|
||||
*"Older LFM references below describe the currently deployed..."*. `docs/rearchitecture.md` announces it
|
||||
"supersedes" a model still described as current elsewhere.
|
||||
|
||||
**impact:** The doc set is the reason finding 1 exists. When five documents describe the
|
||||
architecture, the code becomes the only trustworthy one — which defeats the purpose of
|
||||
having them.
|
||||
|
||||
**recommended action:** Keep `CLAUDE.md` (agent contract), `REARCH.md` (target
|
||||
architecture), `AGENTS.md` (recipes), `PROTOCOL.md` (wire format),
|
||||
**recommended action:** Keep `CLAUDE.md` (agent contract), `docs/rearchitecture.md` (target
|
||||
architecture), `AGENTS.md` (recipes), `docs/protocol.md` (wire format),
|
||||
`20-07-2026-BACKLOG.md` (live queue). Delete the two `SESSION-*.md` and `PLANS.md` — git
|
||||
history holds them. Fold `SPEC.md` + `maven.md` + `ROADMAP.md` into one `DESIGN.md` and
|
||||
history holds them. Fold `SPEC.md` + `maven.md` + `ROADMAP.md` into one `docs/design.md` and
|
||||
mark superseded sections instead of leaving them to read as current. Target ~1,500 lines.
|
||||
|
||||
---
|
||||
@@ -320,7 +320,7 @@ expected maintenance gain: none over the incremental path
|
||||
suggesting Vulkan offload is intended and working — but `CLAUDE.md` says CPU-only. Likely
|
||||
the doc is stale, not the config; unverified.
|
||||
- **Whether the classifier is genuinely adequate.** `voice.go:211` asserts it is;
|
||||
`REARCH.md` asserts it is not. Both are claims, neither is measured. The uncommitted eval
|
||||
`docs/rearchitecture.md` asserts it is not. Both are claims, neither is measured. The uncommitted eval
|
||||
harness is the instrument to settle it — resolve before flipping the router, not after.
|
||||
- **Whether wg+nginx+auth actually fronts 9201 in production.** Not in this repo. If it
|
||||
does, finding 3 drops from "unauthenticated RCE" to "the control is not reproducible from
|
||||
@@ -350,7 +350,7 @@ Key claims independently re-verified against the working tree; the verdict stand
|
||||
over WireGuard on homesrv, this is hygiene, not an emergency — but the loopback bind
|
||||
and startup warning are cheap insurance either way, so do them regardless.
|
||||
- Finding 1's "flip the router" step should be gated harder on measurement.
|
||||
`REARCH.md`'s claim that the classifier causes weak RU queries is itself unmeasured —
|
||||
`docs/rearchitecture.md`'s claim that the classifier causes weak RU queries is itself unmeasured —
|
||||
the review admits this under uncertainties, but the "repair now" ordering buries it.
|
||||
Run `eval_scenarios_test.go` against both paths **before** deciding to flip, not
|
||||
after. A 2B model on CPU may add enough latency that the classifier wins in practice
|
||||
@@ -1,10 +1,12 @@
|
||||
# Maven — Design
|
||||
|
||||
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
|
||||
|
||||
> Folded 2026-07-30 from `SPEC.md` (north star, 2026-07-03), `maven.md`
|
||||
> (consolidated decisions, 2026-06-30) and `ROADMAP.md` (execution plan,
|
||||
> 2026-07-06). Those three files are gone; git history holds them.
|
||||
> This is the single design document: principles, target state, and the
|
||||
> execution ledger. `REARCH.md` remains authoritative wherever it disagrees
|
||||
> execution ledger. `docs/rearchitecture.md` remains authoritative wherever it disagrees
|
||||
> with anything here. Everything the three sources asserted that is no longer
|
||||
> the intended design is preserved under **§ Superseded** — do not read that
|
||||
> section as current.
|
||||
@@ -160,7 +162,7 @@ presence_state ( last_bucket, last_score, updated_ts )
|
||||
```
|
||||
|
||||
Facts additionally carry `Subject`/`EntityID`/`ResolutionState` for
|
||||
entity-aware resolution against Nexus (see `MAVEN_ECOSYSTEM_ARCHITECTURE.md`).
|
||||
entity-aware resolution against Nexus (see `docs/ecosystem.md`).
|
||||
|
||||
### Trigger model
|
||||
|
||||
@@ -196,7 +198,7 @@ INTO the gate as an env predicate, not the LLM's job.
|
||||
|
||||
## Reactive path — routing
|
||||
|
||||
**Target design: LLM-as-router** (see `REARCH.md` and `CLAUDE.md`). One
|
||||
**Target design: LLM-as-router** (see `docs/rearchitecture.md` and `CLAUDE.md`). One
|
||||
resident model emits GBNF-constrained structured JSON, and the same model
|
||||
phrases replies; the embedder is a RAG hint, not a routing gate. The
|
||||
committed default today is the classifier/embedder cascade, which is an
|
||||
@@ -655,7 +657,7 @@ daemon.
|
||||
The voice wire protocol (length-prefixed JSON frames over TCP) is designed for
|
||||
**multiple client implementations**. The reference PWA at `cmd/mavweb` is one
|
||||
client; any app (phone, desktop CLI, smartwatch) can implement the same frame
|
||||
protocol. The published spec is `PROTOCOL.md` — **generated from
|
||||
protocol. The published spec is `docs/protocol.md` — **generated from
|
||||
`internal/voice/wire.go`**, not composed freehand, so it can't drift from
|
||||
code. It covers transport (4-byte big-endian length prefix), methods
|
||||
(`PushToTalk`, `Pong`), push kinds (`AudioNudge`), surface identity
|
||||
@@ -670,8 +672,8 @@ Broadening to home automation, media or comms is JSON, not code.
|
||||
|
||||
## Execution ledger
|
||||
|
||||
Condensed from `ROADMAP.md` (2026-07-06). The live queue is
|
||||
`20-07-2026-BACKLOG.md`; current state is `PROGRESS.md`.
|
||||
Condensed from `ROADMAP.md` (2026-07-06). The live queue is the Vikunja board
|
||||
(project Maven, ID 2); this table is history, not a work list.
|
||||
|
||||
| # | Item | Prio | Status |
|
||||
|---|------|------|--------|
|
||||
@@ -755,7 +757,7 @@ Kept for provenance. **None of this is the current or intended design.**
|
||||
stay deterministic — "classifier owns the route, the SLM stays in its
|
||||
phrasing lane" — with an embedding + nearest-centroid stage 1 over ~10
|
||||
examples per intent, and misroutes appended as new centroid examples.
|
||||
*Replaced by* LLM-as-router (`REARCH.md`): one resident model emits
|
||||
*Replaced by* LLM-as-router (`docs/rearchitecture.md`): one resident model emits
|
||||
GBNF-constrained JSON and also phrases replies; the embedder is demoted to
|
||||
a RAG hint. *Landed 2026-07-31:* the LLM router is on by default and set
|
||||
`true` in `deploy/mavend.json`. The classifier cascade stays as the failure
|
||||
@@ -774,7 +776,7 @@ Kept for provenance. **None of this is the current or intended design.**
|
||||
*Resolved 2026-07-30 (#318), revised 2026-07-31:* the resident checkpoint is
|
||||
stock **Qwen3-1.7B** (`UD-Q4_K_XL`, `n_ctx` 4096), which replaced
|
||||
Qwen3.5-0.8B after measuring better on both fixtures
|
||||
(`MODEL-BAKEOFF-31-07-2026.md`). The CPT'd **Qwen3-1.7B** remains the target
|
||||
(`docs/evals/2026-07-31-model-bakeoff.md`). The CPT'd **Qwen3-1.7B** remains the target
|
||||
(#122); what stock gets wrong is the persona, not the Russian. Note the resident
|
||||
model is no longer described as untrained — the target is trained
|
||||
end-to-end, which is the substantive change from the old claim.
|
||||
@@ -1,5 +1,7 @@
|
||||
# Deterministic logic around a small model
|
||||
|
||||
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
|
||||
|
||||
Written 2026-08-02. Branch `fix/integrated`.
|
||||
|
||||
## The question
|
||||
@@ -268,7 +270,7 @@ rebuilt `mavend` wires it. "кто написал войну и мир?" now rou
|
||||
that turn left unsettled.
|
||||
|
||||
- **The turn was slow, and nobody knows yet whether that is real.** Route 7s,
|
||||
search 1s, phrasing 15s. The p50 in `ROUTING-EVAL-31-07-2026.md` is 825ms. It
|
||||
search 1s, phrasing 15s. The p50 in `docs/evals/2026-07-31-routing.md` is 825ms. It
|
||||
was the first turn after a cold start with the model still warming, so it
|
||||
proves nothing either way. Re-run the same question warm before treating it as
|
||||
a regression. Do not plan latency work off this number.
|
||||
@@ -1,5 +1,7 @@
|
||||
# Maven Ecosystem Architecture
|
||||
|
||||
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
|
||||
|
||||
## 1. Purpose
|
||||
|
||||
This document defines Maven's role in the local ecosystem formed by:
|
||||
@@ -16,7 +16,7 @@ with Qwen3-1.7B.
|
||||
|
||||
Settles Vikunja **#278 / #250**.
|
||||
|
||||
- Same fixture and scorer as `ROUTING-EVAL-31-07-2026.md`: `internal/router/eval/`
|
||||
- Same fixture and scorer as `docs/evals/2026-07-31-routing.md`: `internal/router/eval/`
|
||||
(`ru_routing_v1.json`, 76 held-out cases).
|
||||
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:<port> make eval-router`
|
||||
(`TestLLMRouterBaseline`). (This line used to say there is no `make eval-models` target.
|
||||
@@ -148,7 +148,7 @@ Qwen3-1.7B wins every column, including against a model 20% larger than it.
|
||||
| ontopic | 16, 19, 19 | **22, 23, 23** |
|
||||
| canned fallbacks | 8, 5, 6 | **0, 2, 0** |
|
||||
|
||||
This also fills the row `TALK-EVAL-31-07-2026.md` had to void for contamination:
|
||||
This also fills the row `docs/evals/2026-07-31-talk.md` had to void for contamination:
|
||||
**600ch/1024tok on Qwen3.5-0.8B scores 13, 11, 8.**
|
||||
|
||||
`address` is the headline. It sat at 18-22 of 27 on the 0.8B no matter how the prompt
|
||||
@@ -164,7 +164,7 @@ The 1.7B does that 0-2 times.
|
||||
|
||||
> **Stale, corrected 2026-08-02.** The p50 figures in this table are contention on a
|
||||
> shared llama-server, not the model's cost. The router measures p50 825ms / p95 1.2s /
|
||||
> max 3.0s in `ROUTING-EVAL-31-07-2026.md`, which says so at line 61. Read this table for
|
||||
> max 3.0s in `docs/evals/2026-07-31-routing.md`, which says so at line 61. Read this table for
|
||||
> the shape of the tail only. Take absolute latency from the routing eval.
|
||||
|
||||
| | p50 | p95 |
|
||||
@@ -220,11 +220,11 @@ swapped again when the CPT lands.
|
||||
behind `voice.llm_router`, the default is on, and `deploy/mavend.json` sets it `true`.
|
||||
These numbers are the production path now. **Corrected 2026-08-02: the p50 ≈2.7s in the
|
||||
latency table above WAS a bench artifact.** It is contention on the shared llama-server,
|
||||
not the model. `ROUTING-EVAL-31-07-2026.md` line 61 says so, and measures the router at
|
||||
not the model. `docs/evals/2026-07-31-routing.md` line 61 says so, and measures the router at
|
||||
p50 825ms / p95 1.2s / max 3.0s. Cite that file for latency, not this one.
|
||||
- ~~`/mnt/hdd1/llms/LFM2.5/Qwen3-1.7B-UD-Q4_K_XL.gguf` is a 293 MB truncated download
|
||||
in the wrong directory.~~ **Deleted 2026-07-31.** The good 1.13 GB copy in `qwen3/` is
|
||||
what `deploy/mavend.json` loads.
|
||||
- Harness: `scratchpad/bakeoff.sh`, one server at a time, health-checked before each
|
||||
run, `/v1/models` recorded per run. Never run two LLM consumers at once — see the
|
||||
contamination note in `TALK-EVAL-31-07-2026.md`.
|
||||
contamination note in `docs/evals/2026-07-31-talk.md`.
|
||||
@@ -1,7 +1,7 @@
|
||||
# Phrasing evaluation — 31-07-2026
|
||||
|
||||
How Maven words a nudge, measured instead of argued. Counterpart to
|
||||
`ROUTING-EVAL-31-07-2026.md`.
|
||||
`docs/evals/2026-07-31-routing.md`.
|
||||
|
||||
- Fixture + scorer: `internal/phraser/eval/` (`nudges_v1.json`, 15 cases; `eval.go`, `checks.go`)
|
||||
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing`
|
||||
@@ -65,7 +65,7 @@ prefixes) is the targeted fix, and it would move findings 1 and 2 together. Sepa
|
||||
`model_quantized.onnx` — not the same file.
|
||||
|
||||
`hard` cases score **2/11**: every one is a query where the operator did not reuse his own words.
|
||||
That is the normal case weeks later, and exactly what DESIGN.md's "recall when relevant" promises.
|
||||
That is the normal case weeks later, and exactly what docs/design.md's "recall when relevant" promises.
|
||||
|
||||
### 4. The memory-store recall branch is dead for notes
|
||||
|
||||
@@ -177,7 +177,7 @@ was silent ("не знаю" to "который час") while the one it introdu
|
||||
|
||||
### 1. The resident model does route better — 50.0% vs 36.8%
|
||||
|
||||
REARCH.md's premise holds; `voice.go:211`'s comment does not. **But the classifier is only
|
||||
docs/rearchitecture.md's premise holds; `voice.go:211`'s comment does not. **But the classifier is only
|
||||
~37% correct on held-out utterances, and the model only ~50%.** Neither is "reliable". The
|
||||
gap between them is real but both are far from a system you would describe as working.
|
||||
|
||||
@@ -141,7 +141,7 @@ a model check, which catches a dead server but not a loaded one.
|
||||
measured) on this fixture and the router fixture. Not the 4B — too big for
|
||||
this box, owner's call.
|
||||
- Newer sub-500M candidates (LFM2.5 200M/300M) are worth a run for routing.
|
||||
Note `MODEL-BAKEOFF-31-07-2026.md` found LFM2.5-**1.2B** worse than
|
||||
Note `docs/evals/2026-07-31-model-bakeoff.md` found LFM2.5-**1.2B** worse than
|
||||
Qwen3.5-0.8B at Russian routing and 2.4× slower — but those are a different,
|
||||
older generation, so that result does not predict the small ones.
|
||||
- Fix `chat-how-are-you`'s `want_any`, and re-baseline once, so `ontopic`
|
||||
@@ -0,0 +1,82 @@
|
||||
# gemma-4-12b on the workstation, against the resident Qwen3-1.7B
|
||||
|
||||
Measured 2026-08-02 on the fixtures as they stand. Dated file: it is not edited
|
||||
after today, and a newer number is a new file.
|
||||
|
||||
Vikunja #485's first assumption was that a 7-14B measurably beats Qwen3-1.7B on
|
||||
the 77-case RU routing fixture and the 27-case talk fixture. It does, on both,
|
||||
and it is also faster.
|
||||
|
||||
## The setup
|
||||
|
||||
`gemma-4-12B-it-qat-UD-Q4_K_XL` with the `mtp-gemma-4-12B-it-BF16` draft model,
|
||||
served by `llama-server` b10220 on bugmachine (AMD 7900 GRE, 16GB), fronted by
|
||||
`mavgpud` on `192.168.1.105:8080`. Thinking is off through
|
||||
`--chat-template-kwargs '{"enable_thinking":false}'`, speculative decoding is
|
||||
`--spec-type draft-mtp --spec-draft-n-max 2`, context 32768. The exact line is
|
||||
`deploy/mavgpud.json`.
|
||||
|
||||
Every number below crossed the LAN from homesrv. Note the trap: homesrv's shell
|
||||
exports `HTTP_PROXY`, Go honours it, and the runs need
|
||||
`env -u HTTP_PROXY -u HTTPS_PROXY -u http_proxy -u https_proxy`.
|
||||
|
||||
## Routing, 77-case RU fixture
|
||||
|
||||
| | full | intent-only | p50 | p95 |
|
||||
|---|---|---|---|---|
|
||||
| classifier alone (02-08) | 68.8% | — | 16.6µs | — |
|
||||
| Qwen3-1.7B through the cascade (31-07, 02-08) | 72.7% | 77.9% | 0.80-1.04s | — |
|
||||
| **gemma-4-12b through the cascade** | **84.4%** | **93.5%** | **329ms** | 429ms |
|
||||
| gemma-4-12b alone, no cascade | 55.8% | 85.7% | 335ms | 436ms |
|
||||
|
||||
The workstation buys 11.7 points of full accuracy over the resident model. It
|
||||
buys 15.6 points of intent-only, at a third of the latency. The router's p50 was
|
||||
never the model's fault, which the 02-08 contention finding already said. A 12B
|
||||
on a free 16GB card answers a routing turn in a third of a second.
|
||||
|
||||
Two things the table hides.
|
||||
|
||||
The alone-versus-cascade gap is slots, not intents. gemma reads the intent right
|
||||
85.7% of the time on its own. It loses full accuracy on seven fact keys
|
||||
(`вода` instead of `water`, `ужин` instead of `meal`) and on six reminder times
|
||||
with no time slot. Stage 0 and the daemon's own extractor repair
|
||||
both, which is why the cascade is 28 points higher. The lesson is that the
|
||||
cascade earns its keep even under a much better model, not that it is scaffolding
|
||||
to remove.
|
||||
|
||||
`errors: 6` in the alone row are declines on single-token and ambiguous
|
||||
utterances, all of which the cascade caught. The remaining defects through the
|
||||
cascade are three `query→fact` confusions, one `chat→query`, and one false
|
||||
clarify.
|
||||
|
||||
## Talk, 27-case conversational fixture
|
||||
|
||||
| | pass | notes |
|
||||
|---|---|---|
|
||||
| Qwen3-1.7B (31-07) | 20/27 | 11-17/27 for the 0.8B before it |
|
||||
| **gemma-4-12b** | **25/27 (92.6%)** | chat 8/9, knowledge 9/9, query 8/9 |
|
||||
|
||||
Knowledge is the interesting column: 9/9, in Russian, with real answers about
|
||||
Rayleigh scattering, SSD versus HDD and thunder delay. That is the case the
|
||||
1.7B cannot do at all and the reason the naming half of the degradation rule
|
||||
exists.
|
||||
|
||||
Two failures, and one of them is the persona defect the CPT (#122) targets:
|
||||
`query-notes-do-not-answer` wrote `заплатил` where Maven needs the feminine
|
||||
form. The other is `chat-joke`, where the model told a joke without using any of
|
||||
the words the check looks for. Run-to-run variance is about one case: a second
|
||||
run scored 24/27 with `chat-followup-server` also off-topic.
|
||||
|
||||
## Nudge phrasing, 15-case fixture
|
||||
|
||||
15/15, every check, no errors. `mood`, `lang`, `length`, `feminine`,
|
||||
`hisgender`, `address`, `cringe` and `ontopic` all clean.
|
||||
|
||||
## What this settles and what it does not
|
||||
|
||||
Settled: the size question. A 12B on the workstation beats the resident model on
|
||||
every fixture we have, and it is faster. The offload argument holds.
|
||||
|
||||
Not settled: how often the card is free. That is #485's second assumption and
|
||||
only the `mavgpud` log answers it, after a week of the owner's normal work. A
|
||||
model that is better whenever it is up is worth little if it is never up.
|
||||
+179
@@ -0,0 +1,179 @@
|
||||
# Offloading model work to the workstation
|
||||
|
||||
*Last verified: 2026-08-03 @ 12530c8. Living doc: correct it in place, do not append.*
|
||||
|
||||
Owner's call, 2026-08-02. Vikunja #483 is the umbrella. Tasks #484 to #487 are the
|
||||
work, and this file holds the shape and the rules all four must obey.
|
||||
|
||||
## The goal
|
||||
|
||||
homesrv cannot grow a GPU. The workstation has 16GB of VRAM. Move the model work
|
||||
to the workstation and leave homesrv running the logic that must be always-on,
|
||||
deterministic and cheap.
|
||||
|
||||
## Why this is tractable
|
||||
|
||||
The split already exists structurally. `mavsttd` and `mavttsd` are separate
|
||||
daemons that core reaches over a socket, not linked libraries. Moving them off-box
|
||||
is a transport change, not a redesign.
|
||||
|
||||
The microphone is at the workstation, because that is where the owner sits and
|
||||
homesrv is headless. So speech-to-text and the wake word are already on the
|
||||
workstation side by construction. Audio never has to cross the LAN. Only the core
|
||||
turn does.
|
||||
|
||||
## The constraint that shapes everything
|
||||
|
||||
The workstation's GPU is often busy: CPT runs, experiments, Correx, the manga-recap
|
||||
pipeline. It also sleeps. homesrv does not.
|
||||
|
||||
So an offloaded model is never *the* model. It is the preferred one, with a floor
|
||||
on homesrv. That is the shape the cascade already has, where a router error falls
|
||||
through to the classifier.
|
||||
|
||||
## The degradation rule
|
||||
|
||||
Two cases, and the line between them is sharp.
|
||||
|
||||
**Fall back silently** when the workstation model would only do the job *better*:
|
||||
routing, phrasing, a nudge. Falling back costs nothing that exists today, because
|
||||
the resident Qwen3-1.7B is today's production quality. The owner should not be told
|
||||
that his reply was phrased by the smaller model.
|
||||
|
||||
**Name the gap** when the resident model cannot do the job *at all*. A world
|
||||
question that a 1.7B answers by inventing is the case. A wrong answer is worse
|
||||
than "не могу сейчас". This is the rule CLAUDE.md already states for a sibling
|
||||
service being down.
|
||||
|
||||
Nothing in between. A turn never breaks on the workstation being asleep.
|
||||
|
||||
Both halves are wired, 03-08-2026. `LLMPhraser.PhraseWorld`
|
||||
(`internal/phraser/world.go`) is the naming half and has three outcomes, not two:
|
||||
|
||||
| State | What he hears |
|
||||
|---|---|
|
||||
| no `workstation` block | the resident model answers, exactly as before the seam existed |
|
||||
| configured, card free | the workstation answers |
|
||||
| configured, asleep or busy | the gap, `worldGap` in `cmd/mavend/worldmodel.go` |
|
||||
|
||||
The first row is the one worth stating. Naming a gap requires a gap. On a box with
|
||||
no second model the 1.7B is the whole product. Refusing every world question there
|
||||
would remove a capability the owner has today.
|
||||
|
||||
A source holding a passage is on the naming half too: a live search, a ZIM
|
||||
article, a page he named. None of them says "не могу сейчас". They read the
|
||||
passage back, which is what `phraseSource` returning `""` selects. A real quote
|
||||
beats a gap, and neither path invents.
|
||||
|
||||
## Admission control, not a scheduler
|
||||
|
||||
There is no GPU arbiter. That is a service with its own failure modes, and nothing
|
||||
here needs work *distributed*. It needs admission control. The workstation
|
||||
advertises free VRAM over a health endpoint, and Maven treats it as one more query
|
||||
source that claims a turn or passes. llama-server also refuses to load when VRAM is
|
||||
short, so the failure is detectable without cooperation from the owner's other
|
||||
jobs.
|
||||
|
||||
The caller must be able to ask "is this peer usable right now" without a turn
|
||||
hanging on a timeout. A dead remote is a normal state, not an error state.
|
||||
`internal/llm.Pair` is that check on the Maven side. A prober caches the answer,
|
||||
so `Available()` is an atomic read and no turn pays for a health check.
|
||||
|
||||
llama-server does not stay up on the workstation. It cannot: a resident 7-14B
|
||||
would hold 16GB against the owner's CPT runs. So a supervisor there owns its
|
||||
lifecycle, keeps it loaded while the card is free, and unloads it on idle or
|
||||
when another process needs the card (owner's call, 2026-08-02, Vikunja #488).
|
||||
|
||||
That supervisor is still not a scheduler, and the distinction is worth holding.
|
||||
It arbitrates nothing between callers. It reports whether it can take work and
|
||||
manages one process to back that answer. Maven never asks it to start anything
|
||||
and never learns that it did.
|
||||
|
||||
Contention is decided by presence under `/sys/class/kfd/kfd/proc`, not by a VRAM
|
||||
threshold. A ROCm process registers there when it initialises HIP, before it
|
||||
allocates anything. So the supervisor sees a contender during that job's startup,
|
||||
and yields before the job loses the memory it asked for. A
|
||||
threshold reads the card too late. By the time free VRAM has dropped, the other
|
||||
job has already lost the allocation race. Free VRAM is still read, but only as a
|
||||
precondition for loading, never as the eviction signal. One blind spot is known.
|
||||
A job can take the card without registering on the KFD, as a Vulkan or a
|
||||
video-decode job would. `describe()` logs every contender's comm, and that log is
|
||||
how we find out whether the blind spot is real.
|
||||
|
||||
`mavgpud` runs from a systemd unit on the workstation with
|
||||
`deploy/mavgpud.json` as its config, and `llama_args` is passed to llama-server
|
||||
untouched. The model, the context size, the layer count and the MTP flags are the
|
||||
owner's business and not this daemon's schema.
|
||||
|
||||
## What stays on homesrv, permanently
|
||||
|
||||
The **embedder** (multilingual-e5-small, ONNX, CPU). It backs the classifier, which
|
||||
must answer while the GPU is saturated. It is also cheap enough on CPU that moving
|
||||
it buys nothing. Four callers:
|
||||
|
||||
| Caller | What for |
|
||||
|---|---|
|
||||
| `internal/router/classifier.go` | the routing floor |
|
||||
| `cmd/mavend/actions_query.go` (`queryEmbed`) | memory recall |
|
||||
| `cmd/mavend/feeds.go` | ingest embedding for every RSS item |
|
||||
| `internal/crawl/watch.go` | ingest embedding for every crawled page |
|
||||
|
||||
`internal/speaker` becomes a fifth once it lands.
|
||||
|
||||
## Inventory: what runs a model on homesrv today
|
||||
|
||||
The **resident model** is one llama-server with seven callers, and 03-08-2026 is
|
||||
the date each of them stopped or did not stop being resident-only:
|
||||
|
||||
| Caller | What for | Offloaded |
|
||||
|---|---|---|
|
||||
| `cmd/mavend/voicewire.go` | routing | silently, through `hot` |
|
||||
| `cmd/mavend/replier_llm.go` | replies | silently, through `hot` |
|
||||
| `cmd/mavend/tick.go` | digestion worker: `PhraseNudge`, `PhraseReminder` | silently, inside the phraser |
|
||||
| `cmd/mavend/actions_query.go` | world questions, and any fetched passage | names the gap |
|
||||
| `cmd/mavend/capture.go` | capture summarisation (unreachable, see #480) | no, holds its own client |
|
||||
| `cmd/mavend/mail.go` | mail extraction (off, no IMAP) | no, holds its own client |
|
||||
| `memoryeval.go`, `modelswap.go` | admin and evals | no, and deliberately |
|
||||
|
||||
The last three rows are resident-only on purpose. `memoryeval.go` and
|
||||
`modelswap.go` measure and swap the resident model, so sending their work
|
||||
elsewhere would measure the wrong thing. `capture.go` and `mail.go` are
|
||||
background jobs that hold a gated background client (`llmBackgroundClientFor`),
|
||||
and that priority has no equivalent on the remote yet. Both are also unreachable
|
||||
on this deploy, so wiring them would ship an untestable path.
|
||||
|
||||
The `tick.go` row needs one caveat. `phraser.llm_nudges` is `false` in deploy, so
|
||||
nudges come from templates and the seam under them changes nothing until that
|
||||
flips. It is wired anyway: `PhraseReminder` is on the same transport and is on.
|
||||
|
||||
Then the embedder above, **whisper.cpp** in `mavsttd`, and **piper** in `mavttsd`.
|
||||
`mavwaked` uses no model at all: an energy-threshold VAD over 30ms frames.
|
||||
|
||||
## Order
|
||||
|
||||
1. **Transport** (#484). Nothing else is possible until a seam can cross a host.
|
||||
`internal/netaddr` landed in PR #92. A seam address now carries its own scheme,
|
||||
and a scheme-less one is still unix. A tcp seam requires a shared token, because
|
||||
the filesystem permission that authenticated the unix socket is gone.
|
||||
2. **The resident model** (#485, #490). Wired. A `workstation` block builds an
|
||||
`llm.Pair` in `modelSeam` (`cmd/mavend/voicewire.go`), routing and replies
|
||||
complete through it, and the phraser holds the same pair (`UseRemote`). Both
|
||||
halves of the rule are live: see the table above for which caller gets which.
|
||||
Measured, `docs/evals/2026-08-02-workstation-gemma4-12b.md`: gemma-4-12b
|
||||
through the cascade scores 84.4% full accuracy at p50 329ms. The resident
|
||||
model scores 72.7% at p50 0.80-1.04s. On the talk fixture it is 25/27
|
||||
against 20/27. Biggest quality delta. A 16GB card runs a 7-14B,
|
||||
which fixes what the 1.7B gets wrong: world knowledge, and the persona the CPT
|
||||
targets. The degradation path is already written and measured, since the
|
||||
classifier scores 68.8% full accuracy at p50 16.6µs on its own.
|
||||
3. **Speech-to-text and text-to-speech** (#486). They gain a real margin, but on
|
||||
quality alone, and both already work.
|
||||
4. **The wake word** (#487). Independent of all of the above.
|
||||
|
||||
## Assumptions
|
||||
|
||||
- The LAN is trusted enough that wireguard is supported but not required (owner's
|
||||
call). What crosses the wire is still his utterances. That is why the tcp seam
|
||||
carries its own token instead of assuming a network boundary.
|
||||
- The workstation is not expected to be up. Every child task must still serve a
|
||||
turn while it is down.
|
||||
@@ -1,5 +1,7 @@
|
||||
# Start Commands
|
||||
|
||||
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
|
||||
|
||||
All commands assume `ROOT=/home/kami/apps/Maven` and the local Go toolchain at `$ROOT/deps/go/go/bin/go`.
|
||||
|
||||
## Prerequisites
|
||||
@@ -6,7 +6,7 @@
|
||||
> use the embedded LFM model paths or old single-object examples as current ops
|
||||
> guidance; see `2026-07-18-qwen3-resident-training-eval.md`.
|
||||
|
||||
> Scope from `REARCH.md`. Make Maven trustworthy: the LFM becomes the router
|
||||
> Scope from `docs/rearchitecture.md`. Make Maven trustworthy: the LFM becomes the router
|
||||
> (fixes "messes up queries" / "doesn't take notes"), the engine actually runs
|
||||
> (fixes stub replies), dates stop being read as "number dot number dot number",
|
||||
> and telegram becomes a reach channel. NOT in scope: on-demand 4B reasoner,
|
||||
@@ -693,7 +693,7 @@ ssh kami@192.168.1.104 'curl -s localhost:9201/api/chat -d "{\"text\":\"запо
|
||||
4. `docker compose up -d mavend && docker logs -f maven-mavend-1` — confirm the
|
||||
phraser spawns and no `phraser: NewStub` path. Run the two verify curls.
|
||||
5. Update `AGENTS.md`: LFM model download + note that routing is now LFM-first
|
||||
with classifier fallback (`REARCH.md` is the design of record).
|
||||
with classifier fallback (`docs/rearchitecture.md` is the design of record).
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -1,5 +1,7 @@
|
||||
# Maven Voice Protocol
|
||||
|
||||
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
|
||||
|
||||
> Auto-generated from `internal/voice/wire.go`, `internal/voice/errors.go`,
|
||||
> `internal/voice/frame.go`, `internal/voice/client.go`. If this file and
|
||||
> those files disagree, the code wins.
|
||||
+566
@@ -0,0 +1,566 @@
|
||||
# QA plan: checking Maven properly
|
||||
|
||||
*Last verified: 2026-08-02 @ 20aa2d5. Living doc: correct it in place, do not append.*
|
||||
|
||||
Written 2026-08-01, after the 35-PR stack landed and the box came back up.
|
||||
Refreshed 2026-08-02 against the live list, after PRs #85-#90.
|
||||
|
||||
42 of the 50 open Vikunja tasks are `QA:` tasks. They are verification work, not
|
||||
build work. Most sat unverifiable while Maven was down for 11 days. That
|
||||
blocker is gone.
|
||||
|
||||
The plan as written on 2026-08-01 named 40 task numbers. Ten open `QA:` tasks were
|
||||
missing and two of the named ones had closed. Every open task now appears below,
|
||||
the eight non-QA ones in the last two sections.
|
||||
|
||||
This plan orders them by what unblocks what. Do sessions 1 and 2 first. Almost everything
|
||||
downstream assumes the voice loop works, and nobody has confirmed that since
|
||||
the redeploy.
|
||||
|
||||
---
|
||||
|
||||
## What the 02-08-2026 run found
|
||||
|
||||
Sessions 1, 2 and 3 all ran. Read these five before picking anything up.
|
||||
|
||||
- **470: a question writes invented knowledge into memory.** Recall then serves
|
||||
it back. `что дальше?` lands on `IntentFact` and stores the model's answer as a
|
||||
`self` fact at confidence 1.00. Two junk rows then claimed seven unrelated
|
||||
world questions through recall, outranking the search leg. A question about the
|
||||
capital of Australia was answered `какая последняя версия языка Go?`. Two bad
|
||||
writes silently disabled world answering, with nothing logged.
|
||||
- **466: a pending clarify is global.** One unanswerable clarify swallowed the
|
||||
next three utterances from three separate sessions. With ntfy, telegram and
|
||||
voice all live, a clarify raised on web chat eats the next telegram message.
|
||||
- **467: spoken task capture is dead.** The router calls the capture marker an
|
||||
`act`, and capture is reachable only from the `note` intent.
|
||||
- **The classifier baseline in this repo was wrong**, and it flattered the
|
||||
router. See session 2 and **464**.
|
||||
- **477: the model swap and the self-update cannot be triggered on this box.**
|
||||
Both are built and both are correct in test. The swap needs a passkey and
|
||||
WebAuthn is unconfigured. `mavupdate` needs to reach a socket that only an
|
||||
in-container uid can open.
|
||||
|
||||
- **479: an unconfigured capability lets the question escape to web search.**
|
||||
Netscan off, asked `какие устройства в сети?`. She answered from the live web
|
||||
with a general article about network hardware. A question about his LAN went to
|
||||
an upstream engine. The crawler fails the same way.
|
||||
|
||||
Twenty-one defects were filed on 02-08-2026: 462 through 482. Six tasks this plan
|
||||
had written off as blocked turned out to be ready to check. All six ran. Every
|
||||
one of them is code-correct and stops at the deploy.
|
||||
|
||||
Three of the five config blockers in **472** were then cleared. The morning
|
||||
routine, ambient ingest, feeds, the crawler and netscan are all live. Two remain,
|
||||
and both are the owner's call: a token for each ecosystem sibling, and seed data
|
||||
in Nexus and Praxis.
|
||||
|
||||
---
|
||||
|
||||
## Before you start
|
||||
|
||||
Two things bite anyone running these checks on homesrv.
|
||||
|
||||
**curl needs `--noproxy '*'`.** The shell exports `http_proxy=http://127.0.0.1:18080`.
|
||||
Without the flag, every local check returns 503 from the proxy and looks like a
|
||||
dead service. This cost me a false regression report today.
|
||||
|
||||
**The database is not readable with sqlite3.** Four older QA steps say
|
||||
`docker compose exec mavend sqlite3 /data/maven.db "select ..."`. That cannot
|
||||
work: the container has no `sqlite3` binary, and the store is AES-256-GCM at
|
||||
rest with a tmpfs working copy. Read state through mavweb instead, at
|
||||
`/history`, `/trace`, `/routines` and `/dash`.
|
||||
|
||||
---
|
||||
|
||||
## Session 1: the voice loop (half a day)
|
||||
|
||||
Nothing here has been confirmed since the redeploy, and everything else assumes
|
||||
it works. Do this first.
|
||||
|
||||
Closes or advances: **44** (conversation), **45** (text chat), **287** (voice
|
||||
session quality), **321** steps 3-5 (quiet mode), **288** (STT golden audio).
|
||||
|
||||
**288 is not blocked.** The fixtures are committed under `cmd/mavsttd/testdata/`
|
||||
and `make test-stt-golden` runs today. This plan said otherwise until 02-08-2026.
|
||||
|
||||
Steps 1 and 3-6 were run on 02-08-2026 and pass. Steps 2 and 7-9 still need a
|
||||
person at the box, because they need a microphone or a nudge to arrive.
|
||||
|
||||
Steps 1 and 3-6 do not need a browser. `POST /api/chat` takes a form-encoded
|
||||
`text=` field and a cookie jar, and answers with the rendered `/chat` page:
|
||||
|
||||
```sh
|
||||
curl -s --noproxy '*' -c jar -b jar -L -X POST \
|
||||
http://127.0.0.1:9201/api/chat --data-urlencode 'text=привет'
|
||||
```
|
||||
|
||||
Parse the whole page, not the last text node. The page carries nav and footer
|
||||
text. A naive tail of the Cyrillic nodes returns the wrong string, which makes
|
||||
turns look misaligned when they are not.
|
||||
|
||||
1. Open `http://127.0.0.1:9201/chat` and hold a short conversation in Russian.
|
||||
Watch for three things: she answers in feminine forms (`рада`, `поняла`), she
|
||||
says `ты` and never `вы`, and no pet names appear.
|
||||
**Passes** (02-08-2026, five turns): `я рада`, `поняла`, `помогла`,
|
||||
`проверила`, `записала`, `грустна`, `ты` throughout, no pet names.
|
||||
2. Press push-to-talk on `/dash`. Say `привет`. Confirm a spoken reply comes
|
||||
back. This is the only check that covers mic to STT to core to TTS to
|
||||
speaker as one path. It is also the path the eleven-day outage most likely
|
||||
broke.
|
||||
3. Say `тихий режим`. Expect `тихий режим включён. буду реже напоминать.` **Passes.**
|
||||
4. Say `выключи тихий режим`. Expect `тихий режим выключен.` Negation must win. **Passes.**
|
||||
5. Say `в комнате тихо`. Quiet mode must NOT flip. Confirm on `/history` that no
|
||||
`quiet_hours` fact was written. **Passes**: no row written. She answers `пока
|
||||
не умею отвечать на этот вопрос.`, so it lands on `IntentSystem` with no arm.
|
||||
6. Say `включи режим тишины`, then `сделай потише`. Both must flip quiet mode
|
||||
on. These are the noun form and the comparative, added 01-08-2026. **Both pass.**
|
||||
7. Wait for a nudge, then say `потом` within twenty minutes. Expect `хорошо,
|
||||
вернусь к этому позже.` and the nudge row on `/notifications` reading
|
||||
`snoozed`. Say `потом` again with nothing pending: it must route as an
|
||||
ordinary utterance, not be swallowed.
|
||||
8. Wait for the water nudge, then say `выпил воды`. Expect the ordinary fact
|
||||
reply and nothing extra. She must not congratulate you. Check
|
||||
`/notifications`: the row reads `acted`. Then trigger another nudge and say
|
||||
`готово`. Expect `отлично, отметила.` and the same outcome.
|
||||
9. Note anything where she is slow, cuts off, or talks over herself. That is
|
||||
287's whole content and it has no written acceptance criteria yet.
|
||||
**First evidence, in text** (02-08-2026): nothing breaks, but answers wander
|
||||
and stitch unrelated topics. Asked whether he should move flats, she opened
|
||||
with the weather. That is 287, and it is a phrasing problem, not a loop problem.
|
||||
|
||||
**The wake path cannot be checked as deployed.** `mavwaked` and `mavenclient`
|
||||
appear in no compose file and run as no host process. Step 2 covers only
|
||||
push-to-talk, from `/dash` through mavsttd and mavttsd. Wake word and VAD
|
||||
are untested by construction. Decide whether they belong in compose or on a
|
||||
client machine, and say which in the deploy docs. Tracked as **463**.
|
||||
|
||||
**319's single-token bug is fixed** (01-08-2026). Single-word Russian utterances no longer come
|
||||
back as `не совсем поняла — можешь переформулировать?`. `привет` and `поужинал`
|
||||
both pass now: `thinSingleToken` spares social singles and any token carrying a
|
||||
verb ending, and only thins a bare nominal like `вода`. A one-word utterance that
|
||||
still gets clarified in this session is a new case for the lexicon, not the old bug.
|
||||
|
||||
---
|
||||
|
||||
## Session 2: measurement (half a day, mostly waiting)
|
||||
|
||||
Closes or advances: **320** items 2-4, **278** (make the eval lab routine).
|
||||
Also **248** (memory evaluation), **319** (the margin gate) and **323** (the
|
||||
startup timeout arm).
|
||||
|
||||
The resident llama-server cannot be reached by the eval harness. It binds
|
||||
`--host 127.0.0.1 --port 0` inside the container, so the port is kernel-assigned
|
||||
and never published. Start a second one on a fixed port instead:
|
||||
|
||||
```sh
|
||||
llama-server -m /mnt/hdd1/llms/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf \
|
||||
--host 127.0.0.1 --port 18100 -c 4096 -ngl 99 --no-webui
|
||||
```
|
||||
|
||||
`-c 4096` matters. The recorded numbers were measured at that context size, and
|
||||
a mismatch invalidates the comparison.
|
||||
|
||||
Then:
|
||||
|
||||
```sh
|
||||
make eval-models MAVEN_LLM_URL=http://127.0.0.1:18100 # want ~72.7% cascade
|
||||
make eval-router # classifier baseline
|
||||
MAVEN_LLM_URL=http://127.0.0.1:18100 make eval-phrasing # persona checks, slow
|
||||
make eval-recall
|
||||
```
|
||||
|
||||
A large miss against 72.7% means the deploy differs from the bench harness.
|
||||
|
||||
**Run on 02-08-2026 @ af9d213. The deploy matches the bench.** `eval-models`
|
||||
scored 56 of 77: 72.7% full, 77.9% intent-only, 2 false clarifies and 1 missed.
|
||||
That is the recorded figure to the decimal, and calendar sat at 2 of 2, so the
|
||||
stage 0 agenda rules hold. `eval-phrasing` scored 21 of 27 on the talk fixture
|
||||
against a recorded 20, and the 15 nudge templates passed every check.
|
||||
|
||||
Two numbers in this repo were wrong, and both flattered the resident model.
|
||||
|
||||
- **The classifier is not 36.8% and not 31ms.** `make eval-router` reports
|
||||
`classifier+onnx: 53/77 (68.8% full)` at p50 16.6µs. The figure repeated here
|
||||
and in `CLAUDE.md` predates the stage 0 rules and the seed additions. Both now
|
||||
score inside that baseline. The accuracy gap the router buys is
|
||||
roughly 4 points, not 36. Re-argue the trade on the real numbers: **464**.
|
||||
- **Router latency was measured under contention again.** p50 1.126s, p95 1.58s,
|
||||
max 3.24s, against a recorded p50 825ms. The resident model was serving the
|
||||
daemon on the same iGPU throughout. Do not record this as a regression, and do
|
||||
not record it as a measurement either. Stop the stack before timing the router.
|
||||
|
||||
`classifier+hash` scores 19.5%, which is the no-ONNX degraded path and is not the
|
||||
failure floor the deploy uses. Do not quote it as the classifier baseline.
|
||||
|
||||
Then three things to decide while the numbers are in front of you:
|
||||
|
||||
- **319 is done.** 359 gave the LLM path a real confidence signal.
|
||||
`thinSingleToken` was narrowed on 01-08-2026, and agenda questions moved to
|
||||
stage 0. Missed clarify sits at 1 of 6 and false clarifies at 2. Item 2 point 2
|
||||
closed on 02-08-2026: the `make eval-recall` margin sweep is the distribution
|
||||
that was asked for, and `0.008` sits at the knee.
|
||||
|
||||
| delta | answered | false recall |
|
||||
|---|---|---|
|
||||
| 0.005 | 18/27 | 2/5 |
|
||||
| **0.008** | **18/27** | **1/5** |
|
||||
| 0.010 | 16/27 | 1/5 |
|
||||
|
||||
It removes four of five false recalls at no cost in answers, and the next step
|
||||
costs two answers for nothing. The hand-picked value survives on evidence.
|
||||
- **278's real ask** is making the eval lab routine rather than building it. It
|
||||
is built. Decide whether it runs on a timer, on every merge, or on demand, and
|
||||
the task can close.
|
||||
- **248** is the memory evaluation loop. It ships, it writes notes, and it cannot
|
||||
speak. `make eval-recall` covers the retrieval half. The open question is whether
|
||||
a written evaluation nobody reads is worth the tick.
|
||||
|
||||
**323 is down to one check.** PR #90 covered the spawn path and took phraser
|
||||
coverage to 76.9%. Only the 60s startup timeout arm is untested, because testing it
|
||||
needs a `StartupTimeout` field on `Config` rather than a test-only hack. While you
|
||||
are on the box, time a cold 1.7B load off spinning disk. If it runs near 60s, the
|
||||
default is too tight and the field earns itself twice.
|
||||
|
||||
Warm, it is nowhere near. A second llama-server answered `/health` 1.8s after
|
||||
launch at `n_ctx 4096` on 02-08-2026. That is page cache, so it does not settle
|
||||
the question. A cold read needs a cache drop, which needs root.
|
||||
|
||||
**`CheckFeminine` has a false positive.** On 02-08-2026 it failed
|
||||
`query-notes-do-not-answer` for `ты заплатил`, calling it masculine
|
||||
self-reference. Masculine second person is correct, because the owner is male.
|
||||
The check matches a masculine
|
||||
past-tense verb before `за` without confirming the subject is `я`. Fix it in
|
||||
`internal/phraser/eval/checks.go` before trusting a phrasing score to the case.
|
||||
The real talk-fixture score on that run is 22 of 27, not 21. Tracked as **462**.
|
||||
|
||||
Item 4 of **320** needs a permission I do not have. Kill the `llama-server`
|
||||
pid under `maven-mavend-1`, post a turn, and confirm it still completes
|
||||
through the classifier. Either grant it or run it yourself. It is the only
|
||||
check that the failure floor catches a mid-session model death.
|
||||
|
||||
---
|
||||
|
||||
## Session 3: the interaction batch (a day, or five sittings)
|
||||
|
||||
These need real use rather than a command, grouped by what one sitting covers.
|
||||
|
||||
**Morning and delivery** (**280**, **281**, **128**, **282**, **283**, **285**):
|
||||
open `/morning`, walk the seven required behaviours, then check the four
|
||||
interruption outcomes and the digest gap. **282** needs the `desk_active` script
|
||||
enabled on the desk PC first, which is **15** and needs you at that machine.
|
||||
**283** is the event intake envelope every reach shares, so a delivery check
|
||||
exercises it whether you name it or not. **285** is not verification: the bridge
|
||||
framework works and the remaining ask is more adapters. Decide which reach comes
|
||||
next, or park it.
|
||||
|
||||
Run 02-08-2026. **280 is blocked.** No morning routine is configured (**472**).
|
||||
`morning.Item` also has no required-versus-optional field, so behaviour 1 cannot
|
||||
hold whatever you configure (**473**). **281's digest gap is closed**, and
|
||||
its presence rule passes on inspection. Three of its five items need traffic the
|
||||
box has not had. **283 is blocked**: nothing feeds the intake journal. **128
|
||||
found the worst defect of the whole session, see below.**
|
||||
|
||||
Three of 472's five blockers were cleared the same day, in `deploy/mavend.json`
|
||||
and `docker-compose.yml`.
|
||||
|
||||
- A `morning_routines` block, one routine `утро` 08:00-11:00 with medicine,
|
||||
water and pets. It is live: the dispatcher logged `dropped morning:утро (sev1,
|
||||
presence=away)`, so the plan builds and the nudge is proposed. 280's
|
||||
behaviours and 128 step 11 are checkable now. 473 still stands.
|
||||
- `-ambient-token` on mavweb, value in a gitignored `/.env` that docker compose
|
||||
reads for interpolation. `/api/ambient` answers 401 without the token and 201
|
||||
with it, storing `calendar_event_20260802_Standup`. 283 step 5 and 128 step 8
|
||||
are unblocked. The token is a flag, so it shows in `ps` inside that container.
|
||||
The zenmoney and IMAP secrets are read from files instead. Ingest also
|
||||
reads the notification's wall clock as UTC and stores a 14:30 meeting at 18:30
|
||||
(**482**).
|
||||
- `feeds` (two sources), `crawl.on_demand` and `netscan.enabled`. The intake
|
||||
journal now fills: `/events` holds `scan:lan` and `ambient:notif` rows.
|
||||
|
||||
Two are not mine to clear. No sibling has a `token` in `deploy/mavend.json`, so
|
||||
273 steps 6 and 8 need a credential decision. Nexus has no entities and Praxis no
|
||||
attention items, so 272 step 3 needs seed data whose content is the owner's call.
|
||||
|
||||
For **285**, two facts bear on the choice. Synapse is already running on this box
|
||||
and healthy, so a Matrix reach has a live target and needs no new service. And
|
||||
mavweb is already a PWA with a service worker, which 285 itself calls the highest
|
||||
value adapter left. Today's reaches are ntfy, telegram and voice.
|
||||
|
||||
**Query sources** (**258**, **286**): ask her something the RSS feeds answer and
|
||||
something only a ZIM answers, with the search block on. Live search leads and the
|
||||
ZIMs are the fallback since 02-08-2026. **286**'s remaining half is doc and
|
||||
git ingestion, which is build work, not a check.
|
||||
|
||||
**Do not read `/trace` for this.** `/trace` is the nudge-rule trace: rule,
|
||||
severity, predicate, gate, selected. No query-source field exists anywhere in the
|
||||
codebase. The only evidence of which query source claimed a turn is the
|
||||
`voice: search:` and `voice: kiwix:` lines in `docker compose logs mavend`
|
||||
(`actions_query.go:589` and `:660`).
|
||||
|
||||
Run 02-08-2026, 20 turns. **Search leads and the personal boundary holds.** Every
|
||||
world question that reached the boundary was claimed by search. All three
|
||||
personal questions produced no search and no kiwix line at all.
|
||||
|
||||
The rest of this sitting went badly. **Kiwix has zero live coverage.** SearXNG
|
||||
returns four results for everything, including two invented nonsense terms. So
|
||||
`querySearch` always claims, and Kiwix is unreachable code as deployed. The ZIM
|
||||
half of the 02-08-2026 decision is unverified. A ZIM answer cannot signal a
|
||||
silent search failure, because a ZIM answer cannot happen.
|
||||
**Ordering defects** in feeds and calendar, plus 258 step 1's utterance not
|
||||
working: **474**. And the sitting independently found stage 2 of **470**.
|
||||
|
||||
**Tasks and calendar** (**129**, **130**, **127**, **126**, **246**): capture a
|
||||
task by voice, confirm it lands, check prioritisation ordering is not nonsense.
|
||||
**246** (mail reader) also exercises the `IngestMail` rung that moved to
|
||||
`AuthWrite` this morning.
|
||||
|
||||
Run 02-08-2026. **129 passes.** The page and the spoken answer agree on ordering.
|
||||
The undistinguished task carries no invented reason on either surface, which is
|
||||
the thing 129 asks for. **130 fails outright** and **127 half fails**:
|
||||
**467**, **469**. **246 cannot be run**: `mavmaild` is commented out in
|
||||
`docker-compose.yml` and there is no `email` block, so nothing in steps 4-13 is
|
||||
reachable. The `IngestMail` rung does sit at `AuthWrite`
|
||||
(`internal/auth/policy.go:96`, asserted in `auth_test.go:421`), verified by
|
||||
reading only.
|
||||
|
||||
**Routines and patterns** (**43**, **46**, **247**, **254**): these need history
|
||||
to detect against. If the database is thin after the outage, they may have
|
||||
nothing to propose, which is not a failure. Check `/routines` before
|
||||
concluding anything.
|
||||
|
||||
Run 02-08-2026. The answer is the middle case: **the detector ran and found
|
||||
nothing.** The tick loop is live, and `detectPatterns` is called unconditionally
|
||||
at `cmd/mavend/tick.go:227`. It has run about 25 times since the restart. It
|
||||
finds nothing because the events table is empty upstream of it. Rows land there
|
||||
only from `pattern.Extract` at fact-write time, and `Extract` requires the fact
|
||||
value to match a closed 7-action lexicon. All 200 facts on `/history` are
|
||||
`page_heartbeat`, `netdata_alarm`, `quiet_hours`, `name`, `service_down` and
|
||||
`рост`. Not one lexicon hit, so no event can exist, let alone the four one pair
|
||||
needs. **46 step 5 passes**: `/routines` renders `noticed 0` with the empty state
|
||||
and the hint string.
|
||||
|
||||
Two things block this sitting, and both are build work. The seeding recipe on
|
||||
**43** goes through `sqlite3` and cannot work. And `pattern.Detect` has no
|
||||
minimum-interval floor, so seeding by hand mints a permanent false routine
|
||||
(**468**). Do not try to seed a pattern with four fast chat turns.
|
||||
|
||||
**Ecosystem** (**272**, **273**, **276**): nexus, hexis and praxis are wired and
|
||||
logged clean at boot.
|
||||
|
||||
Run 02-08-2026, read-only half. All three answer `/health` 200 and `/ecosystem`
|
||||
lists 18 Hexis capabilities with correct read-only and mutating badges. **272 and
|
||||
273 are blocked on empty data**, not on code. Nexus holds no entities, Praxis
|
||||
holds no attention items, and the Calls panel has never recorded a call. See
|
||||
**472**, and read its warning first. 273's trace fix has never been validated
|
||||
here. An empty Calls panel is exactly what the old bug looked like. The page is
|
||||
`/ecosystem`, not `/siblings`.
|
||||
|
||||
**276 ran 02-08-2026 and the suite is sound.** 17 `TestEcosystem_` cases pass
|
||||
under `-race`, not the 10 the task describes. The mutation check bites: patching
|
||||
the Nexus-error branch of `handleHexisAct` to `return ""` fails
|
||||
`TestEcosystem_MalformedNexusResponseFailsClosed` on the expected line.
|
||||
|
||||
Steps 4 and 6 could not be checked through chat, because no utterance reaches
|
||||
Praxis (**475**). «что требует внимания» routes to `intent=query` and is answered
|
||||
by the search leg, identically whether `ecosystem-praxis-1` is up or stopped. The
|
||||
degraded string never appears because its branch is never entered. Step 5 is
|
||||
blocked the same way: `перезапусти muzick indexer` clarifies on
|
||||
`HasFn:false`, and the router had already rewritten the entity name to
|
||||
`музик индексер` (**476**).
|
||||
|
||||
Both steps were checked on `/ecosystem` instead, which reads Praxis directly.
|
||||
With Praxis stopped the card reads `praxis — unreachable` while Nexus and Hexis
|
||||
keep rendering. On `docker start` the card returns to `nothing needs attention.`
|
||||
with no mavend restart. Independent degradation and recovery both hold.
|
||||
|
||||
**Operations** (**249**, **250**): both ran 02-08-2026. The code is correct and
|
||||
neither lever can be pulled on this box. See **477**.
|
||||
|
||||
**250** passes steps 1, 2, 3, 9 and 10 on the deploy. The capability announces
|
||||
itself. `/models` names the model llama-server reports, not the config filename.
|
||||
Asking her to switch models does nothing. Removing `swap_models` renders `swap
|
||||
not configured`. Step 4's refusal half passes at HTTP 403, and the 403 comes from
|
||||
mavend rather than mavweb. WebAuthn is unconfigured, so the web gate fails open
|
||||
and the wire gate fails closed. Steps 5 to 8 need a passkey assertion nothing on
|
||||
this box can produce. They pass in test: 13 swap cases and 7 page cases covering
|
||||
drain, mid-swap refusal, rollback, failed rollback and the not-owned refusal.
|
||||
|
||||
**249** passes steps 1 and 2. Step 3 stops it. `mavupdate` health-checks over
|
||||
`/run/maven/mavend.sock`, which is `srw------- 1 10001 999` inside a docker
|
||||
volume. The host owner cannot traverse `/var/lib/docker/volumes` and cannot
|
||||
connect to a socket owned by an in-container uid. `mavupdate` assumes a
|
||||
host-installed daemon and the deploy is containers. Do not sudo around this.
|
||||
|
||||
---
|
||||
|
||||
## Housekeeping (done 02-08-2026, and this section was mostly wrong)
|
||||
|
||||
This section claimed eleven tasks were not verification work. **Three were not.
|
||||
The other eight are.** Every one of the eight has shipped, tested code behind it.
|
||||
The error ran one way: it wrote off work that is ready to check. Do not trust a
|
||||
"nothing is built" line in this plan without grepping for the package first.
|
||||
|
||||
Relabelled to `Blocked:`, claim verified:
|
||||
|
||||
- **125** zenmoney. `internal/zenmoney/` ships and is tested against a fixture.
|
||||
`deploy/zenmoney.token` does not exist and the compose mount is commented out.
|
||||
One token unblocks it.
|
||||
- **256** Home Assistant. `internal/smarthome/` ships, the `smarthome` block sits
|
||||
in `deploy/mavend.json` at `enabled: false`, and 8123 and 1883 are closed.
|
||||
- **14** cold-start unlock. The seam is real at `cmd/mavend/main.go:128` and
|
||||
`internal/webauthn/prf.go` is in place. `lockedAPI` is gone, replaced by
|
||||
`Server.Check` in `internal/ipc/server.go`. Gated on an authenticator that
|
||||
implements the WebAuthn PRF extension, which is hardware, not code.
|
||||
|
||||
Left alone, because the claim here was false:
|
||||
|
||||
- **284** simulator. `cmd/mavend/simulator_test.go`, three scenarios under
|
||||
`cmd/mavend/testdata/scenarios/`, and a `simulate` target at `Makefile:98`.
|
||||
**Run 02-08-2026: all three scenarios pass**, plus the determinism and
|
||||
backwards-step guards. One defect found, see below.
|
||||
- **288** STT golden audio. Four WAVs and `golden_v1.json` are committed under
|
||||
`cmd/mavsttd/testdata/`, the make targets exist, and `models/stt/ggml-small.bin`
|
||||
is on the box. Session 1 lists 288 as blocked on fixtures, which is wrong.
|
||||
**Run 02-08-2026: all four pass**, WER at or under ceiling with no drift.
|
||||
|
||||
| fixture | transcript | WER | ceiling |
|
||||
|---|---|---|---|
|
||||
| ru_reminder | `Напомни мне через час позвонить маме.` | 0.00 | 0.10 |
|
||||
| ru_fact | `А отметь, что я выпил воды.` | 0.20 | 0.25 |
|
||||
| ru_query | `Что у меня сегодня по календарю?` | 0.00 | 0.10 |
|
||||
| en_act | `Restart the web server and check the disk space.` | 0.00 | 0.10 |
|
||||
|
||||
That also settles a session 1 worry indirectly: whisper.cpp works on Vulkan
|
||||
after the redeploy. Only the mic and the wake path remain unproven.
|
||||
|
||||
**The simulator routes with an empty seed set.** Every `make simulate` run logs
|
||||
`loaded 0 seed examples from models/seeds`, seven times per scenario. The test
|
||||
runs from `cmd/mavend`, and the seed path is relative to the repo root. The
|
||||
scenarios still pass, which means they pass without the classifier having any
|
||||
seeds to match against. Whatever 284 is proving, it is not proving the routing
|
||||
the deploy runs. Fix the path before trusting a green simulator.
|
||||
- **257** Bluetooth. The bluez half is genuinely absent. The LAN-scan half shipped
|
||||
(`internal/netscan/`), and steps 1-9 run today. Only step 10 is Bluetooth, so
|
||||
relabelling the whole task would bury real pending work.
|
||||
- **251** MCP, **253** hearing, **259** crawler. All three ship
|
||||
(`internal/mcp/`, `internal/capture/`, `internal/crawl/`) with no external gate.
|
||||
Fully checkable. `259`'s step 1 wants no `crawl` block in `deploy/mavend.json`,
|
||||
and there is none, so it is already set up correctly.
|
||||
- **252** vision and **255** speaker recognition. Both ship. Each is blocked only
|
||||
on a model download: a vision gguf with mmproj, and a speaker embedding model.
|
||||
Neither is present under `/mnt/hdd1`. Their refusal-path steps run today.
|
||||
|
||||
So the honest split is three blocked on a credential or hardware, two blocked on
|
||||
a download, and six ready to check. That is roughly a session of real QA this
|
||||
plan had written off as backlog.
|
||||
|
||||
**All six ran on 02-08-2026.** Every one of them is code-correct and stops at the
|
||||
deploy. The pattern repeats often enough to be the headline: the packages pass,
|
||||
and the box cannot reach them.
|
||||
|
||||
**251, MCP.** Steps 1, 2, 3, 4 and 13 pass. Package tests green under `-race`.
|
||||
Off-by-default is clean, and the SSRF refusal is exact: without `allow_private`
|
||||
the log reads `refusing to connect to a private address: 127.0.0.1` and `/tools`
|
||||
shows the server down with zero proposals. Steps 5 to 12 are blocked. `ss -lntp`
|
||||
shows the Vikunja MCP server on `127.0.0.1:9100` only, so no container reaches it
|
||||
at any address (**478**). `allow_private` does work, measured both ways.
|
||||
|
||||
**253, hearing.** Steps 1, 2 and 17 pass. `internal/capture` covers 90.3%. Steps
|
||||
7 to 16 are blocked on something nobody can work around: no shipped client calls
|
||||
`CaptureStart`. There is no `cmd/mavheard`, no mavweb route, and `mavenclient`
|
||||
never calls it (**480**). Two of its QA steps are also stale.
|
||||
|
||||
**257, netscan.** Steps 2, 3 and 9 pass at unit level. Step 1 fails. Steps 4 to 8
|
||||
need the block enabled. Step 10 is Bluetooth and stays skipped.
|
||||
|
||||
**259, crawler.** Steps 1 and 15 pass. Step 2 fails. Steps 3 to 14 need a `crawl`
|
||||
block that nobody has written.
|
||||
|
||||
Both were configured later the same day, and both work. `netscan.enabled: true`
|
||||
answers `какие устройства в сети?` with `нашла 3 устройства, из них 2 с вебом, 2 с
|
||||
ssh. список записала.` and the scan lands in the intake journal as `scan:lan`.
|
||||
`crawl.on_demand: true` answers `посмотри https://lwn.net — что там пишут?` from
|
||||
the real page. So **479** is one defect, not the routing defect it was filed as.
|
||||
An unconfigured capability declines its own turn instead of naming the gap.
|
||||
Nothing is wrong with the routing.
|
||||
|
||||
257 step 1 and 259 step 2 fail the same way and share a task (**479**). An
|
||||
unconfigured capability does not name the gap, so the question escapes to web
|
||||
search. `какие устройства в сети?` was answered with a general article about
|
||||
network hardware. That is his LAN going to an upstream engine.
|
||||
|
||||
**252 vision and 255 speaker.** Both confirmed blocked. The disk claim was
|
||||
re-verified rather than taken on trust: 16 text-only ggufs under `/mnt/hdd1`, no
|
||||
mmproj and no speaker embedding model. Everything not needing the model passes,
|
||||
including the two refusals that matter. `TestNewLocalRefusesNonPrivateEndpoints`
|
||||
rejects `https://api.openai.com`, and forget really deletes
|
||||
(`internal/store/memory.go:145` is a real `DELETE`, not a tombstone). Vision is
|
||||
19/19, speaker 22/22, media 16/16.
|
||||
|
||||
**470 got worse, then closed.** Both poisoned facts showed `voided` on
|
||||
`/history` and the defect survived. Re-measured at 15:42, after four restarts:
|
||||
`почему небо синее?` still answered `какая последняя версия языка Go?` with no
|
||||
`search:` line. What came back was the question he typed, not the value the fact
|
||||
held. So the poison was a vector in the memory index, and `revert` did not
|
||||
remove it.
|
||||
|
||||
Repaired in two parts. 470 stopped the writes: a question is never a fact, and a
|
||||
void drops the key's vectors. 493 fixed what the index holds. A fact is indexed
|
||||
as the fact and not as the utterance, and a correction drops its superseded
|
||||
vector too.
|
||||
|
||||
A poisoned box now repairs itself on the next start. `RepairFactVectors`
|
||||
re-embeds every fact vector from the fact it names, and deletes the voided and
|
||||
superseded ones. It runs once, guarded by a marker, and logs what it did.
|
||||
|
||||
---
|
||||
|
||||
## Needs you specifically
|
||||
|
||||
Not QA. These are blocked on a decision or a credential only you have.
|
||||
|
||||
| # | what |
|
||||
|---|---|
|
||||
| 16 | Create the Kuma API key. `-kuma-key uk5_mavpoll-key` in `docker-compose.yml` is still the placeholder. |
|
||||
| 15 | Deploy `desk_active` on the desk PC. Blocks **282**. |
|
||||
| 122 | Finish the CPT run for Qwen3-1.7B. The persona fix depends on it. |
|
||||
| 355 | Deploy the Hexis auth change. Was blocked on Maven being under construction, which it no longer is. The client half is vendored and wired. |
|
||||
| 357 | Decide whether entity-existence validation is the permanent target guard or whether blessing lands in Nexus. |
|
||||
| 275 | Hexis native API and MCP parity. |
|
||||
| — | Decide on `-require-stepup`. Making it the default needs WebAuthn configured first, or it locks you out of your own admin surfaces. |
|
||||
|
||||
317 and 354 closed on 01-08-2026. The step-up gate now covers `POST /api/chat` and
|
||||
`/routines`, and the nginx template is locked down with a `maven.<domain>` block for
|
||||
mavweb. The `-require-stepup` default is still your call.
|
||||
|
||||
---
|
||||
|
||||
## Not this repo
|
||||
|
||||
Two open tasks sit on the Maven board and are not Maven work. Move them or note
|
||||
where they land, so the board stops reading as 50 things Maven owes.
|
||||
|
||||
- **358** replace the rowid execution cursor with a real seq column. This is Hexis,
|
||||
and it must land before any execution retention or pruning does.
|
||||
- **362** mirror the router prompt reorder into the relabelling prompt. This is the
|
||||
training workspace, enforced by `llm/check_prompt_parity.py` there, not here.
|
||||
|
||||
---
|
||||
|
||||
## Suggested order
|
||||
|
||||
1. Session 1. If the voice loop is broken, nothing else matters.
|
||||
2. The `-require-stepup` and Kuma decisions. Five minutes, and it unblocks **16**.
|
||||
3. Session 2. **Run on 02-08-2026.** The numbers came back worse for the router
|
||||
than the docs claimed. The classifier is 68.8%, not 36.8%, and 16.6µs, not
|
||||
31ms. The router buys about 4 points of accuracy for four orders of magnitude
|
||||
of latency. Whether that still earns its place is now an open question.
|
||||
4. Housekeeping. Cheap, and it makes the remaining backlog honest.
|
||||
5. Session 3, split whichever way suits you. All five sittings ran on
|
||||
02-08-2026. Read the per-sitting notes before repeating any of them.
|
||||
|
||||
The next thing to fix is not in this plan. Four defects say the same sentence:
|
||||
a capability is built and no utterance reaches it. **466** (a clarify is global),
|
||||
**467** (capture is act-routed), **475** (attention is act-routed), **476** (the
|
||||
router rewrites entity names). Routing is where the work is.
|
||||
@@ -1,5 +1,7 @@
|
||||
# Maven — Re-architecture (Qwen3 resident model, revised 2026-07-18)
|
||||
|
||||
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
|
||||
|
||||
> Supersedes the classifier-first routing model. Agreed in a design session
|
||||
> after diagnosing that homesrv deploys with a **stub phraser** (no LLM
|
||||
> running) and an embedder-classifier that routes by nearest-neighbor between
|
||||
@@ -1,7 +1,7 @@
|
||||
// Package auth is maven's authority layer — the 4-layer cascade and the
|
||||
// "surface caps authority" invariant.
|
||||
//
|
||||
// Spec contract (from DESIGN.md § Auth):
|
||||
// Spec contract (from docs/design.md § Auth):
|
||||
//
|
||||
// a cascade, not a pick-one — each layer answers a different question:
|
||||
//
|
||||
|
||||
@@ -224,6 +224,11 @@ type Config struct {
|
||||
// See SearchConfig.
|
||||
Search *SearchConfig `json:"search,omitempty"`
|
||||
|
||||
// Workstation — the big model on the owner's desktop, preferred over the
|
||||
// resident one when its GPU is free. nil / absent / url empty ⇒ homesrv
|
||||
// behaves exactly as it does today. See WorkstationConfig.
|
||||
Workstation *WorkstationConfig `json:"workstation,omitempty"`
|
||||
|
||||
// Praxis — the ecosystem attention-state service. When configured, maven
|
||||
// calls the Praxis HTTP tools API for attention listing and item lifecycle.
|
||||
// Maven never touches Praxis's database directly (ecosystem invariant: no
|
||||
@@ -649,7 +654,7 @@ type VoiceConfig struct {
|
||||
// LLMRouter — route with the resident model instead of the embedding
|
||||
// classifier. On by default since Vikunja #320.
|
||||
//
|
||||
// Measured on the held-out fixture (ROUTING-EVAL-31-07-2026.md): 63.2% of
|
||||
// Measured on the held-out fixture (docs/evals/2026-07-31-routing.md): 63.2% of
|
||||
// intents right against the classifier's 50.0%, and no route errors. It
|
||||
// costs about 1s per turn instead of 30ms.
|
||||
//
|
||||
@@ -1105,6 +1110,44 @@ const (
|
||||
DefaultKiwixSnippetRunes = 1500
|
||||
)
|
||||
|
||||
// WorkstationConfig — the big model on the owner's desktop (workpc, a
|
||||
// 7900 GRE with 16GB), fronted by mavgpud.
|
||||
//
|
||||
// homesrv cannot grow a GPU, so the resident Qwen3-1.7B is the floor and this
|
||||
// is the preferred model above it (owner's call, 2026-08-02, docs/offload.md).
|
||||
// The workstation is never assumed up: its card is often held by a CPT run and
|
||||
// the machine sleeps. No block, or an empty URL, and homesrv behaves exactly as
|
||||
// it does today.
|
||||
//
|
||||
// Only the prompt crosses the LAN, and the workstation is not "the box". The
|
||||
// rules in CLAUDE.md about what may leave still apply.
|
||||
type WorkstationConfig struct {
|
||||
// URL — where mavgpud listens, e.g. "http://192.168.1.105:8080". Empty ⇒
|
||||
// the whole block is normalised to nil and nothing probes anything.
|
||||
URL string `json:"url,omitempty"`
|
||||
|
||||
// Health — the admission endpoint. Empty ⇒ URL + "/health", which is what
|
||||
// mavgpud serves. It answers 503 while the card is held, and that is the
|
||||
// signal, so it must be the supervisor's endpoint and not llama-server's.
|
||||
Health string `json:"health,omitempty"`
|
||||
|
||||
// Probe — how often admission is re-checked. 0 ⇒ DefaultWorkstationProbe.
|
||||
// Nothing on the hot path waits for it: the answer is cached and read
|
||||
// atomically, so this only sets how late Maven notices the card came back.
|
||||
Probe Duration `json:"probe,omitempty"`
|
||||
|
||||
// Timeout — the per-request budget for a completion on the workstation.
|
||||
// 0 ⇒ DefaultWorkstationTimeout. A big model on a LAN host is slower than
|
||||
// the resident one, and a request that overruns falls back to the floor.
|
||||
Timeout Duration `json:"timeout,omitempty"`
|
||||
}
|
||||
|
||||
// Workstation defaults, applied in Normalise.
|
||||
const (
|
||||
DefaultWorkstationProbe = 15 * time.Second
|
||||
DefaultWorkstationTimeout = 90 * time.Second
|
||||
)
|
||||
|
||||
// SearchConfig — the self-hosted SearXNG instance she searches with.
|
||||
//
|
||||
// External search is allowed and off unless configured (CLAUDE.md). Configuring
|
||||
@@ -1500,6 +1543,24 @@ func (c *Config) applyDefaults() {
|
||||
}
|
||||
}
|
||||
|
||||
// No address, no preferred model. An unconfigured workstation is the
|
||||
// default deploy and must be indistinguishable from today.
|
||||
if c.Workstation != nil && strings.TrimSpace(c.Workstation.URL) == "" {
|
||||
c.Workstation = nil
|
||||
}
|
||||
if c.Workstation != nil {
|
||||
w := c.Workstation
|
||||
if strings.TrimSpace(w.Health) == "" {
|
||||
w.Health = strings.TrimRight(w.URL, "/") + "/health"
|
||||
}
|
||||
if w.Probe <= 0 {
|
||||
w.Probe = Duration(DefaultWorkstationProbe)
|
||||
}
|
||||
if w.Timeout <= 0 {
|
||||
w.Timeout = Duration(DefaultWorkstationTimeout)
|
||||
}
|
||||
}
|
||||
|
||||
if c.Voice != nil {
|
||||
if c.Voice.RouterThreshold <= 0 {
|
||||
c.Voice.RouterThreshold = DefaultRouterThreshold
|
||||
|
||||
@@ -413,3 +413,56 @@ func TestNormaliseFillsKiwixDefaults(t *testing.T) {
|
||||
t.Error("rewrite: false was not honoured")
|
||||
}
|
||||
}
|
||||
|
||||
// A workstation with no address is not a workstation. The unconfigured deploy
|
||||
// must be indistinguishable from today, so the block is dropped rather than
|
||||
// left to fail one probe at a time.
|
||||
func TestNormaliseDropsAddresslessWorkstation(t *testing.T) {
|
||||
for _, tc := range []struct {
|
||||
name string
|
||||
in *WorkstationConfig
|
||||
}{
|
||||
{"no url", &WorkstationConfig{Probe: Duration(time.Second)}},
|
||||
{"blank url", &WorkstationConfig{URL: " "}},
|
||||
} {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
c := &Config{Workstation: tc.in}
|
||||
c.applyDefaults()
|
||||
if c.Workstation != nil {
|
||||
t.Errorf("kept an unusable workstation block: %+v", c.Workstation)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// The health endpoint defaults to the supervisor's, not llama-server's: mavgpud
|
||||
// answers 503 while the card is held, and that refusal is the whole signal.
|
||||
func TestNormaliseFillsWorkstationDefaults(t *testing.T) {
|
||||
c := &Config{Workstation: &WorkstationConfig{URL: "http://192.168.1.105:8080/"}}
|
||||
c.applyDefaults()
|
||||
if c.Workstation == nil {
|
||||
t.Fatal("dropped a usable workstation block")
|
||||
}
|
||||
if got, want := c.Workstation.Health, "http://192.168.1.105:8080/health"; got != want {
|
||||
t.Errorf("Health = %q, want %q", got, want)
|
||||
}
|
||||
if time.Duration(c.Workstation.Probe) != DefaultWorkstationProbe {
|
||||
t.Errorf("Probe = %s, want %s", time.Duration(c.Workstation.Probe), DefaultWorkstationProbe)
|
||||
}
|
||||
if time.Duration(c.Workstation.Timeout) != DefaultWorkstationTimeout {
|
||||
t.Errorf("Timeout = %s, want %s", time.Duration(c.Workstation.Timeout), DefaultWorkstationTimeout)
|
||||
}
|
||||
}
|
||||
|
||||
// An explicit health URL is left alone: the supervisor may sit behind something
|
||||
// that does not put /health at the root.
|
||||
func TestNormaliseKeepsExplicitWorkstationHealth(t *testing.T) {
|
||||
c := &Config{Workstation: &WorkstationConfig{
|
||||
URL: "http://192.168.1.105:8080",
|
||||
Health: "http://192.168.1.105:9000/ready",
|
||||
}}
|
||||
c.applyDefaults()
|
||||
if got, want := c.Workstation.Health, "http://192.168.1.105:9000/ready"; got != want {
|
||||
t.Errorf("Health = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
// Package delivery is maven's channel-routing + dispatch layer.
|
||||
//
|
||||
// Spec contract (from DESIGN.md § Delivery / channel routing):
|
||||
// Spec contract (from docs/design.md § Delivery / channel routing):
|
||||
//
|
||||
// - routing = f(severity, presence). presence decides REACHABILITY; severity
|
||||
// decides INSISTENCE. need both.
|
||||
|
||||
@@ -9,7 +9,7 @@ import (
|
||||
"github.com/kami/maven/internal/store"
|
||||
)
|
||||
|
||||
// This file walks every cell of the DESIGN.md § "Delivery / channel routing"
|
||||
// This file walks every cell of the docs/design.md § "Delivery / channel routing"
|
||||
// table, once as the pure table and once through the dispatcher, so a change
|
||||
// to either side has to break a named cell.
|
||||
//
|
||||
@@ -205,7 +205,7 @@ func TestAwayChannelsGetMinimalBody(t *testing.T) {
|
||||
// An empty Summary no longer means "send the whole body" — it means a short
|
||||
// generic line — so the old expectation here was wrong as well as duplicated.
|
||||
|
||||
// TestCareAwayDropIsRecorded — DESIGN.md's drop is a decision ("a missed water
|
||||
// TestCareAwayDropIsRecorded — docs/design.md's drop is a decision ("a missed water
|
||||
// nudge is noise, a missed backup failure isn't"), so it should be visible
|
||||
// rather than vanish. Today drop is a bare `continue`: no nudge row, no outbox
|
||||
// attempt, no log — nothing an operator can see afterwards. now it leaves a
|
||||
|
||||
@@ -19,7 +19,7 @@
|
||||
// 3. if no live session exists, Send returns voice.ErrNoSession
|
||||
// (wrapped). The daemon logs the partial dispatch; an OPEN deferred
|
||||
// question is whether the dispatcher should reroute to away-channels
|
||||
// instead of returning partial — listed in PROGRESS.md.
|
||||
// instead of returning partial.
|
||||
//
|
||||
// Import direction: voicesink imports internal/tts (synth seam) and
|
||||
// internal/voice (Sessions registry). Both are siblings of delivery; the
|
||||
|
||||
@@ -1,5 +1,4 @@
|
||||
// Package event is the unified intake envelope (Vikunja #283,
|
||||
// 20-07-2026-BACKLOG.md item 1).
|
||||
// Package event is the unified intake envelope (Vikunja #283).
|
||||
//
|
||||
// # The problem it solves
|
||||
//
|
||||
|
||||
+24
-13
@@ -8,13 +8,15 @@ import (
|
||||
"net"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/netaddr"
|
||||
)
|
||||
|
||||
// Client — the module side of the boundary. Wraps a unix-socket connection
|
||||
// and satisfies CoreAPI, so a module imports ipc, holds a CoreAPI, and is
|
||||
// agnostic to whether it's been wired in-process (tests / daemon-embedded)
|
||||
// or over this socket (full topology). The swappability is the seam auth
|
||||
// will insert into without touching module code.
|
||||
// Client — the module side of the boundary. Wraps a connection to core and
|
||||
// satisfies CoreAPI, so a module imports ipc, holds a CoreAPI, and is
|
||||
// agnostic to whether it's been wired in-process (tests / daemon-embedded),
|
||||
// over a local unix socket, or over tcp to another host. The swappability is
|
||||
// the seam auth will insert into without touching module code.
|
||||
//
|
||||
// One Client ⇒ one conn ⇒ one concurrent request at a time. A module that
|
||||
// wants parallel requests opens one Client per goroutine; the store is the
|
||||
@@ -22,7 +24,8 @@ import (
|
||||
// per-Client lock keeps frame interleaving impossible by construction.
|
||||
type Client struct {
|
||||
conn net.Conn
|
||||
path string // kept so a dropped conn can be re-dialed (core restart)
|
||||
path string // the address as configured, kept for errors and logs
|
||||
addr netaddr.Addr // parsed, so a dropped conn can be re-dialed (core restart)
|
||||
mu sync.Mutex
|
||||
}
|
||||
|
||||
@@ -81,14 +84,22 @@ var readOnlyMethods = map[Method]bool{
|
||||
MethodPing: true,
|
||||
}
|
||||
|
||||
// Dial connects to a core socket at path and returns a Client. The module
|
||||
// owns its Client lifecycle; Close on shutdown.
|
||||
// Dial connects to core at path and returns a Client. The module owns its
|
||||
// Client lifecycle; Close on shutdown.
|
||||
//
|
||||
// path is a netaddr seam address: a bare path is the unix socket it has
|
||||
// always been, and "tcp://host:port?token=..." reaches a core on another
|
||||
// host. See internal/netaddr.
|
||||
func Dial(path string) (*Client, error) {
|
||||
c, err := net.Dial("unix", path)
|
||||
addr, err := netaddr.Parse(path)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("ipc: dial %s: %w", path, err)
|
||||
return nil, err
|
||||
}
|
||||
return &Client{conn: c, path: path}, nil
|
||||
c, err := netaddr.Dial(addr)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("ipc: dial %s: %w", addr, err)
|
||||
}
|
||||
return &Client{conn: c, path: path, addr: addr}, nil
|
||||
}
|
||||
|
||||
func (c *Client) Close() error {
|
||||
@@ -189,9 +200,9 @@ func (c *Client) call(ctx context.Context, m Method, params, result any) error {
|
||||
// re-dials clean. Caller holds c.mu.
|
||||
func (c *Client) roundtrip(m Method, raw json.RawMessage, resp *Response) error {
|
||||
if c.conn == nil {
|
||||
conn, err := net.Dial("unix", c.path)
|
||||
conn, err := netaddr.Dial(c.addr)
|
||||
if err != nil {
|
||||
return fmt.Errorf("%w: dial %s: %v", errWriteLost, c.path, err)
|
||||
return fmt.Errorf("%w: dial %s: %v", errWriteLost, c.addr, err)
|
||||
}
|
||||
c.conn = conn
|
||||
}
|
||||
|
||||
+20
-40
@@ -8,11 +8,11 @@ import (
|
||||
"fmt"
|
||||
"log"
|
||||
"net"
|
||||
"os"
|
||||
"sync"
|
||||
"sync/atomic"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/netaddr"
|
||||
"github.com/kami/maven/internal/store"
|
||||
"golang.org/x/sys/unix"
|
||||
)
|
||||
@@ -446,6 +446,7 @@ func mapErr(err error) error {
|
||||
type Server struct {
|
||||
api atomic.Value // stores CoreAPI
|
||||
path string
|
||||
addr netaddr.Addr
|
||||
|
||||
ln net.Listener
|
||||
wg sync.WaitGroup
|
||||
@@ -610,31 +611,29 @@ type CheckFunc func(ctx context.Context, m Method, params json.RawMessage) error
|
||||
// MethodAssertStepUp dispatch calls this instead of going through CoreAPI.
|
||||
type StepUpFunc func(ctx context.Context) error
|
||||
|
||||
// Listen creates a Server bound to path. path's parent dir must exist and be
|
||||
// 0700 (we chmod it if we own it); the socket file itself is created 0600 so
|
||||
// only the same unix user can connect — the current "auth floor", same radius
|
||||
// as wg at the network boundary. Removing a stale socket at path first lets
|
||||
// the daemon restart cleanly.
|
||||
// Listen creates a Server bound to path.
|
||||
//
|
||||
// A bare path is a unix socket, unchanged: its parent dir is 0700 and the
|
||||
// socket file itself is 0600, so only the same unix user can connect — the
|
||||
// current "auth floor", same radius as wg at the network boundary. A stale
|
||||
// socket is removed first so the daemon restarts cleanly.
|
||||
//
|
||||
// A "tcp://host:port?token=..." address binds a network listener instead, for
|
||||
// a module that lives on another host. There is no filesystem there to be the
|
||||
// auth floor, so netaddr checks the shared token before this package sees the
|
||||
// connection and a token is mandatory. See internal/netaddr.
|
||||
func Listen(path string, api CoreAPI) (*Server, error) {
|
||||
_ = os.Remove(path) // stale socket from a crashed daemon; ignore missing
|
||||
if err := os.MkdirAll(parentDir(path), 0o700); err != nil {
|
||||
return nil, fmt.Errorf("ipc: mkdir socket dir: %w", err)
|
||||
}
|
||||
// umask could widen the perms on socket creation; tighten then chmod to
|
||||
// be explicit. 0600 ⇒ read+write by owner only.
|
||||
oldMask := unix.Umask(0o077)
|
||||
ln, err := net.Listen("unix", path)
|
||||
unix.Umask(oldMask)
|
||||
addr, err := netaddr.Parse(path)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("ipc: listen %s: %w", path, err)
|
||||
return nil, err
|
||||
}
|
||||
if err := os.Chmod(path, 0o600); err != nil {
|
||||
_ = ln.Close()
|
||||
_ = os.Remove(path)
|
||||
return nil, fmt.Errorf("ipc: chmod socket: %w", err)
|
||||
ln, err := netaddr.Listen(addr)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
s := &Server{
|
||||
path: path,
|
||||
addr: addr,
|
||||
ln: ln,
|
||||
done: make(chan struct{}),
|
||||
}
|
||||
@@ -1268,7 +1267,7 @@ func (s *Server) Close() error {
|
||||
// missing the seal costs every write since the last clean shutdown.
|
||||
log.Printf("ipc: %d connection(s) still busy after %s, closing anyway", s.liveConns(), closeGrace)
|
||||
}
|
||||
_ = os.Remove(s.path)
|
||||
netaddr.Cleanup(s.addr)
|
||||
return err
|
||||
}
|
||||
|
||||
@@ -1338,25 +1337,6 @@ func (s *Server) Path() string { return s.path }
|
||||
// while the server is serving (dispatch loads api once per request via atomic).
|
||||
func (s *Server) SetAPI(api CoreAPI) { s.api.Store(api) }
|
||||
|
||||
func parentDir(p string) string {
|
||||
if i := lastIndexByte(p, '/'); i >= 0 {
|
||||
if i == 0 {
|
||||
return "/"
|
||||
}
|
||||
return p[:i]
|
||||
}
|
||||
return "."
|
||||
}
|
||||
|
||||
func lastIndexByte(s string, b byte) int {
|
||||
for i := len(s) - 1; i >= 0; i-- {
|
||||
if s[i] == b {
|
||||
return i
|
||||
}
|
||||
}
|
||||
return -1
|
||||
}
|
||||
|
||||
// peerCaller — read SO_PEERCRED off a unix conn to identify the connecting
|
||||
// process. Returns ok=false on a non-unix conn or a platform without
|
||||
// SO_PEERCRED; the caller then proceeds without a Caller (the socket perms
|
||||
|
||||
@@ -31,7 +31,7 @@ type Completer interface {
|
||||
// or reply in Russian.
|
||||
//
|
||||
// Why the JSON wrapper: this model always thinks out loud and this llama-server
|
||||
// build ignores the thinking switch (see ROUTING-EVAL-31-07-2026.md). A bare
|
||||
// build ignores the thinking switch (see docs/evals/2026-07-31-routing.md). A bare
|
||||
// word-list grammar just captured the reasoning — every case came back as
|
||||
// "Let me analyze this request carefully". Demanding JSON, like routeGrammar and
|
||||
// responseGrammar already do, gives the reasoning nowhere to go.
|
||||
|
||||
@@ -129,6 +129,11 @@ type Req struct {
|
||||
RepeatPenalty float64
|
||||
// Stop — sequences that end generation early (e.g. newline for a one-liner).
|
||||
Stop []string
|
||||
// Temperature — 0 (the zero value) is greedy decoding, and greedy is what
|
||||
// every caller here wanted before this field existed. It is set only by the
|
||||
// phraser, whose own transport has always sampled at 0.7: routing a phrasing
|
||||
// call through this client must not quietly change how it decodes.
|
||||
Temperature float64
|
||||
}
|
||||
|
||||
type msg struct {
|
||||
@@ -176,7 +181,7 @@ func (c *Client) Complete(ctx context.Context, r Req) (string, error) {
|
||||
Messages: []msg{{Role: "system", Content: r.System}, {Role: "user", Content: r.User}},
|
||||
MaxTokens: r.MaxTokens,
|
||||
Grammar: r.Grammar,
|
||||
Temp: 0,
|
||||
Temp: r.Temperature,
|
||||
RepeatPenalty: r.RepeatPenalty,
|
||||
Stop: r.Stop,
|
||||
})
|
||||
|
||||
@@ -0,0 +1,193 @@
|
||||
package llm
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"log"
|
||||
"net/http"
|
||||
"sync/atomic"
|
||||
"time"
|
||||
)
|
||||
|
||||
// Pair — a preferred model on another host, with the resident one as the floor.
|
||||
//
|
||||
// homesrv cannot grow a GPU and the workstation has 16GB of VRAM, so the big
|
||||
// model runs there and the resident Qwen3-1.7B stays here. See docs/offload.md.
|
||||
// The workstation is never assumed up: its GPU is often busy with CPT runs and
|
||||
// the manga-recap pipeline, and the machine sleeps. So the remote is preferred,
|
||||
// never required, and Pair is what makes "preferred" mean something precise.
|
||||
//
|
||||
// This is admission control, not a scheduler. There is no arbiter deciding who
|
||||
// gets the card. A prober asks the remote whether it will take work, caches the
|
||||
// answer, and every request reads that cached answer in nanoseconds. Routing
|
||||
// sits on the hot path at p50 825ms and must never wait on a machine that may
|
||||
// be asleep, so no request ever pays for a health check itself.
|
||||
//
|
||||
// Pair satisfies nothing by itself. Callers pick a method by which half of the
|
||||
// degradation rule they live under:
|
||||
//
|
||||
// - Complete falls back silently. For routing, replies, and nudge phrasing,
|
||||
// where the big model is only better and the 1.7B is today's shipping
|
||||
// quality. He is not told which model phrased his reply.
|
||||
// - CompleteRemote returns ErrRemoteUnavailable instead of falling back. For
|
||||
// a world question, or a long Kiwix or search passage, where a 1.7B
|
||||
// confabulates rather than summarises. A named gap beats an invented
|
||||
// answer.
|
||||
type Pair struct {
|
||||
remote *Client
|
||||
floor *Client
|
||||
|
||||
// up — the cached admission answer, written only by the prober goroutine
|
||||
// and read by every request. Atomic so the read costs nanoseconds and no
|
||||
// request ever contends with the prober.
|
||||
up atomic.Bool
|
||||
|
||||
health string
|
||||
interval time.Duration
|
||||
http *http.Client
|
||||
stop chan struct{}
|
||||
}
|
||||
|
||||
// ErrRemoteUnavailable — the workstation model was required and is not
|
||||
// answering. Callers on the naming half of the degradation rule turn this into
|
||||
// a gap in the reply ("не могу сейчас"), never into a guess from the floor.
|
||||
var ErrRemoteUnavailable = errors.New("llm: workstation model unavailable")
|
||||
|
||||
// ErrNoFloor — a Pair was built with no resident model to fall back to. A
|
||||
// configuration mistake: the floor is the whole point.
|
||||
var ErrNoFloor = errors.New("llm: no floor client")
|
||||
|
||||
// NewPair builds the two-model arrangement. remote may be nil, which is the
|
||||
// unconfigured deploy and must behave exactly as the box behaves today: every
|
||||
// call goes to the floor and nothing probes anything.
|
||||
//
|
||||
// health is the URL the prober asks. llama-server's /health answers "is a model
|
||||
// loaded and ready", which is the useful signal here, because llama-server
|
||||
// refuses to load at all when VRAM is short. That makes a busy card detectable
|
||||
// without any cooperation from the owner's other jobs.
|
||||
func NewPair(remote, floor *Client, health string, interval time.Duration) *Pair {
|
||||
p := &Pair{
|
||||
remote: remote,
|
||||
floor: floor,
|
||||
health: health,
|
||||
interval: interval,
|
||||
http: &http.Client{Timeout: probeTimeout},
|
||||
stop: make(chan struct{}),
|
||||
}
|
||||
return p
|
||||
}
|
||||
|
||||
// probeTimeout — a remote that cannot answer /health this fast is not going to
|
||||
// serve a turn either. Short on purpose: the prober runs on its own goroutine,
|
||||
// but a slow probe still delays the moment Maven notices the card came back.
|
||||
const probeTimeout = 2 * time.Second
|
||||
|
||||
// Start begins probing. It returns immediately, and the first probe runs before
|
||||
// the first tick so a remote that is already up is used on the first turn
|
||||
// rather than after one interval of falling back. Safe to call with a nil
|
||||
// remote; it does nothing.
|
||||
func (p *Pair) Start(ctx context.Context) {
|
||||
if p.remote == nil || p.health == "" {
|
||||
return
|
||||
}
|
||||
go func() {
|
||||
p.probe(ctx)
|
||||
t := time.NewTicker(p.interval)
|
||||
defer t.Stop()
|
||||
for {
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return
|
||||
case <-p.stop:
|
||||
return
|
||||
case <-t.C:
|
||||
p.probe(ctx)
|
||||
}
|
||||
}
|
||||
}()
|
||||
}
|
||||
|
||||
// Stop ends the prober. Idempotent.
|
||||
func (p *Pair) Stop() {
|
||||
select {
|
||||
case <-p.stop:
|
||||
default:
|
||||
close(p.stop)
|
||||
}
|
||||
}
|
||||
|
||||
// Available reports whether the workstation will take work right now. It reads
|
||||
// a cached flag, so it is safe to call per turn on the hot path. A false answer
|
||||
// is never stale in the direction that matters: the worst case is that Maven
|
||||
// falls back for up to one probe interval after the card frees up.
|
||||
func (p *Pair) Available() bool {
|
||||
return p.remote != nil && p.up.Load()
|
||||
}
|
||||
|
||||
func (p *Pair) probe(ctx context.Context) {
|
||||
ctx, cancel := context.WithTimeout(ctx, probeTimeout)
|
||||
defer cancel()
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodGet, p.health, nil)
|
||||
if err != nil {
|
||||
p.set(false)
|
||||
return
|
||||
}
|
||||
resp, err := p.http.Do(req)
|
||||
if err != nil {
|
||||
p.set(false)
|
||||
return
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
p.set(resp.StatusCode == http.StatusOK)
|
||||
}
|
||||
|
||||
// set records the admission answer and logs only the transitions. A machine
|
||||
// that sleeps every night would otherwise write one line per interval forever.
|
||||
func (p *Pair) set(up bool) {
|
||||
if p.up.Swap(up) == up {
|
||||
return
|
||||
}
|
||||
if up {
|
||||
log.Printf("llm: workstation model available at %s", p.health)
|
||||
} else {
|
||||
log.Printf("llm: workstation model unavailable, falling back to the resident model")
|
||||
}
|
||||
}
|
||||
|
||||
// Complete runs r on the workstation when it will take work, and on the
|
||||
// resident model otherwise. A remote that fails mid-request falls back too: the
|
||||
// admission answer is a cache and can be one interval out of date, so an error
|
||||
// here is expected rather than exceptional.
|
||||
//
|
||||
// This is the silent half of the degradation rule. It must be indistinguishable
|
||||
// from today's behaviour when the workstation is down.
|
||||
func (p *Pair) Complete(ctx context.Context, r Req) (string, error) {
|
||||
if p.floor == nil {
|
||||
return "", ErrNoFloor
|
||||
}
|
||||
if p.Available() {
|
||||
out, err := p.remote.Complete(ctx, r)
|
||||
if err == nil {
|
||||
return out, nil
|
||||
}
|
||||
// The cached answer was wrong. Correct it now rather than sending the
|
||||
// next request into the same hole, then fall back.
|
||||
p.set(false)
|
||||
}
|
||||
return p.floor.Complete(ctx, r)
|
||||
}
|
||||
|
||||
// CompleteRemote runs r on the workstation or refuses. It never falls back,
|
||||
// because for a world question the resident 1.7B does not answer worse, it
|
||||
// invents. Callers turn ErrRemoteUnavailable into a named gap.
|
||||
func (p *Pair) CompleteRemote(ctx context.Context, r Req) (string, error) {
|
||||
if !p.Available() {
|
||||
return "", ErrRemoteUnavailable
|
||||
}
|
||||
out, err := p.remote.Complete(ctx, r)
|
||||
if err != nil {
|
||||
p.set(false)
|
||||
return "", errors.Join(ErrRemoteUnavailable, err)
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
@@ -0,0 +1,210 @@
|
||||
package llm
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"sync/atomic"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// completionServer stands in for a llama-server. It counts what reached it, so
|
||||
// a test can say which of the two models answered.
|
||||
func completionServer(t *testing.T, reply string, hits *atomic.Int64) *httptest.Server {
|
||||
t.Helper()
|
||||
s := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
hits.Add(1)
|
||||
w.Header().Set("Content-Type", "application/json")
|
||||
_, _ = w.Write([]byte(`{"choices":[{"message":{"content":"` + reply + `"}}]}`))
|
||||
}))
|
||||
t.Cleanup(s.Close)
|
||||
return s
|
||||
}
|
||||
|
||||
func healthServer(t *testing.T, ok *atomic.Bool) *httptest.Server {
|
||||
t.Helper()
|
||||
s := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
if !ok.Load() {
|
||||
w.WriteHeader(http.StatusServiceUnavailable)
|
||||
return
|
||||
}
|
||||
w.WriteHeader(http.StatusOK)
|
||||
}))
|
||||
t.Cleanup(s.Close)
|
||||
return s
|
||||
}
|
||||
|
||||
// waitFor polls until cond holds or the deadline passes. The prober runs on its
|
||||
// own goroutine, so a test has to wait for it rather than assume it has run.
|
||||
func waitFor(t *testing.T, cond func() bool) bool {
|
||||
t.Helper()
|
||||
deadline := time.Now().Add(2 * time.Second)
|
||||
for time.Now().Before(deadline) {
|
||||
if cond() {
|
||||
return true
|
||||
}
|
||||
time.Sleep(5 * time.Millisecond)
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// The unconfigured deploy. No remote, no probing, every call to the floor —
|
||||
// exactly what the box does today.
|
||||
func TestNoRemoteGoesToTheFloor(t *testing.T) {
|
||||
var floorHits atomic.Int64
|
||||
floor := completionServer(t, "floor", &floorHits)
|
||||
|
||||
p := NewPair(nil, New(floor.URL, time.Second), "", time.Second)
|
||||
p.Start(context.Background())
|
||||
defer p.Stop()
|
||||
|
||||
if p.Available() {
|
||||
t.Fatal("a Pair with no remote reports available")
|
||||
}
|
||||
out, err := p.Complete(context.Background(), Req{User: "привет"})
|
||||
if err != nil {
|
||||
t.Fatalf("complete: %v", err)
|
||||
}
|
||||
if out != "floor" || floorHits.Load() != 1 {
|
||||
t.Fatalf("out = %q, floor hits = %d", out, floorHits.Load())
|
||||
}
|
||||
}
|
||||
|
||||
// The workstation is up, so it answers and the resident model is not touched.
|
||||
func TestAvailableRemoteAnswers(t *testing.T) {
|
||||
var remoteHits, floorHits atomic.Int64
|
||||
remote := completionServer(t, "remote", &remoteHits)
|
||||
floor := completionServer(t, "floor", &floorHits)
|
||||
up := &atomic.Bool{}
|
||||
up.Store(true)
|
||||
health := healthServer(t, up)
|
||||
|
||||
p := NewPair(New(remote.URL, time.Second), New(floor.URL, time.Second), health.URL, 20*time.Millisecond)
|
||||
p.Start(context.Background())
|
||||
defer p.Stop()
|
||||
if !waitFor(t, p.Available) {
|
||||
t.Fatal("prober never saw the remote come up")
|
||||
}
|
||||
|
||||
out, err := p.Complete(context.Background(), Req{User: "привет"})
|
||||
if err != nil {
|
||||
t.Fatalf("complete: %v", err)
|
||||
}
|
||||
if out != "remote" || floorHits.Load() != 0 {
|
||||
t.Fatalf("out = %q, floor hits = %d", out, floorHits.Load())
|
||||
}
|
||||
}
|
||||
|
||||
// The card is busy, so /health refuses and Complete degrades silently. This is
|
||||
// the constraint from 483: the workstation being down is indistinguishable from
|
||||
// today's behaviour.
|
||||
func TestBusyCardFallsBackSilently(t *testing.T) {
|
||||
var remoteHits, floorHits atomic.Int64
|
||||
remote := completionServer(t, "remote", &remoteHits)
|
||||
floor := completionServer(t, "floor", &floorHits)
|
||||
health := healthServer(t, &atomic.Bool{}) // never ok
|
||||
|
||||
p := NewPair(New(remote.URL, time.Second), New(floor.URL, time.Second), health.URL, 20*time.Millisecond)
|
||||
p.Start(context.Background())
|
||||
defer p.Stop()
|
||||
time.Sleep(60 * time.Millisecond)
|
||||
|
||||
out, err := p.Complete(context.Background(), Req{User: "привет"})
|
||||
if err != nil {
|
||||
t.Fatalf("complete: %v", err)
|
||||
}
|
||||
if out != "floor" || remoteHits.Load() != 0 {
|
||||
t.Fatalf("out = %q, remote hits = %d", out, remoteHits.Load())
|
||||
}
|
||||
}
|
||||
|
||||
// The cached admission answer can be one interval out of date, so a remote that
|
||||
// dies between probes must still not break the turn.
|
||||
func TestRemoteErrorMidRequestFallsBack(t *testing.T) {
|
||||
var floorHits atomic.Int64
|
||||
dead := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
w.WriteHeader(http.StatusInternalServerError)
|
||||
}))
|
||||
defer dead.Close()
|
||||
floor := completionServer(t, "floor", &floorHits)
|
||||
up := &atomic.Bool{}
|
||||
up.Store(true)
|
||||
health := healthServer(t, up)
|
||||
|
||||
p := NewPair(New(dead.URL, time.Second), New(floor.URL, time.Second), health.URL, time.Hour)
|
||||
p.Start(context.Background())
|
||||
defer p.Stop()
|
||||
if !waitFor(t, p.Available) {
|
||||
t.Fatal("prober never saw the remote come up")
|
||||
}
|
||||
|
||||
out, err := p.Complete(context.Background(), Req{User: "привет"})
|
||||
if err != nil {
|
||||
t.Fatalf("complete: %v", err)
|
||||
}
|
||||
if out != "floor" || floorHits.Load() != 1 {
|
||||
t.Fatalf("out = %q, floor hits = %d", out, floorHits.Load())
|
||||
}
|
||||
// The failed request must have corrected the cached answer, so the next
|
||||
// one does not walk into the same hole.
|
||||
if p.Available() {
|
||||
t.Fatal("a failed remote request left the admission answer up")
|
||||
}
|
||||
}
|
||||
|
||||
// The naming half of the degradation rule. A world question must not be handed
|
||||
// to the resident model, because it answers by inventing.
|
||||
func TestCompleteRemoteNamesTheGap(t *testing.T) {
|
||||
var floorHits atomic.Int64
|
||||
floor := completionServer(t, "floor", &floorHits)
|
||||
health := healthServer(t, &atomic.Bool{}) // never ok
|
||||
|
||||
p := NewPair(New("http://127.0.0.1:1", time.Second), New(floor.URL, time.Second), health.URL, 20*time.Millisecond)
|
||||
p.Start(context.Background())
|
||||
defer p.Stop()
|
||||
time.Sleep(60 * time.Millisecond)
|
||||
|
||||
if _, err := p.CompleteRemote(context.Background(), Req{User: "почему небо голубое"}); !errors.Is(err, ErrRemoteUnavailable) {
|
||||
t.Fatalf("err = %v, want ErrRemoteUnavailable", err)
|
||||
}
|
||||
if floorHits.Load() != 0 {
|
||||
t.Fatalf("CompleteRemote fell back to the floor %d times", floorHits.Load())
|
||||
}
|
||||
}
|
||||
|
||||
// Routing sits on the hot path and must never pay for a health check. Available
|
||||
// reads a cached flag, so it costs no network at all.
|
||||
func TestAvailableDoesNotProbe(t *testing.T) {
|
||||
var probes atomic.Int64
|
||||
health := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
probes.Add(1)
|
||||
w.WriteHeader(http.StatusOK)
|
||||
}))
|
||||
defer health.Close()
|
||||
|
||||
p := NewPair(New("http://127.0.0.1:1", time.Second), New("http://127.0.0.1:1", time.Second), health.URL, time.Hour)
|
||||
p.Start(context.Background())
|
||||
defer p.Stop()
|
||||
if !waitFor(t, p.Available) {
|
||||
t.Fatal("prober never ran")
|
||||
}
|
||||
|
||||
before := probes.Load()
|
||||
for range 1000 {
|
||||
p.Available()
|
||||
}
|
||||
if got := probes.Load(); got != before {
|
||||
t.Fatalf("1000 Available calls made %d probes", got-before)
|
||||
}
|
||||
}
|
||||
|
||||
// A Pair with no floor is a configuration mistake, and it must say so rather
|
||||
// than silently having nowhere to degrade to.
|
||||
func TestNoFloorIsAnError(t *testing.T) {
|
||||
p := NewPair(nil, nil, "", time.Second)
|
||||
if _, err := p.Complete(context.Background(), Req{User: "привет"}); !errors.Is(err, ErrNoFloor) {
|
||||
t.Fatalf("err = %v, want ErrNoFloor", err)
|
||||
}
|
||||
}
|
||||
@@ -10,12 +10,12 @@ import (
|
||||
|
||||
// Tests for the universal restraint gate.
|
||||
//
|
||||
// DESIGN.md § Trigger model: "the gate is universal, applied by the loop, never
|
||||
// docs/design.md § Trigger model: "the gate is universal, applied by the loop, never
|
||||
// per-rule — quiet-hours, presence, cooldown, snooze, calendar-busy all live in
|
||||
// one fires()." These tests pin the CONSERVATIVE side of that: the cases where
|
||||
// Maven must stay quiet. They exist so nobody loosens the gate by accident.
|
||||
//
|
||||
// Where the code does not yet do what DESIGN.md promises, the test is written to
|
||||
// Where the code does not yet do what docs/design.md promises, the test is written to
|
||||
// show the gap and then skipped, with the file and line to fix. Behaviour is not
|
||||
// changed to make a test pass.
|
||||
|
||||
@@ -54,7 +54,7 @@ func TestGateQuietHoursSuppressesCareOnly(t *testing.T) {
|
||||
|
||||
// ---------------------------- presence ---------------------------------------
|
||||
|
||||
// DESIGN.md § Delivery: "sev <= 2 drops on away, sev >= 3 holds: a missed water
|
||||
// docs/design.md § Delivery: "sev <= 2 drops on away, sev >= 3 holds: a missed water
|
||||
// nudge is noise, a missed backup failure isn't."
|
||||
func TestGateAwayDropsCareHoldsOps(t *testing.T) {
|
||||
cases := []struct {
|
||||
@@ -250,7 +250,7 @@ func TestTickOrderOfRulesDoesNotMatter(t *testing.T) {
|
||||
|
||||
// ---------------------------- reminders bypass the gate ----------------------
|
||||
|
||||
// DESIGN.md § User reminders: "bypasses the restraint gate — 'wake me 7' fires
|
||||
// docs/design.md § User reminders: "bypasses the restraint gate — 'wake me 7' fires
|
||||
// in quiet hours; that's the point." Every suppressor set at once, and the
|
||||
// reminder still comes through.
|
||||
func TestRemindersBypassEverySuppressor(t *testing.T) {
|
||||
@@ -269,7 +269,7 @@ func TestRemindersBypassEverySuppressor(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// GAP — DESIGN.md § User reminders ends "Snooze still applies." RemindDecisions
|
||||
// GAP — docs/design.md § User reminders ends "Snooze still applies." RemindDecisions
|
||||
// passes every due reminder straight through with no snooze check, so a snoozed
|
||||
// reminder fires anyway. The test below is what the contract asks for.
|
||||
func TestRemindersStillHonourSnooze(t *testing.T) {
|
||||
|
||||
@@ -18,7 +18,7 @@ import (
|
||||
// - does it stay quiet when it should?
|
||||
// - is it silent when the key it needs has no data at all?
|
||||
//
|
||||
// The last one is load-bearing. DESIGN.md: "since(key)==null → don't fire.
|
||||
// The last one is load-bearing. docs/design.md: "since(key)==null → don't fire.
|
||||
// Silence on no-data is 'shuts up when uncertain'."
|
||||
|
||||
// stateWith builds a snapshot at refTime() holding just the given facts.
|
||||
@@ -179,7 +179,7 @@ func TestCareRulePredicates(t *testing.T) {
|
||||
|
||||
// ---------------------------- ops rules --------------------------------------
|
||||
|
||||
// The two ops rules match on a value AND on which poller wrote it. DESIGN.md:
|
||||
// The two ops rules match on a value AND on which poller wrote it. docs/design.md:
|
||||
// "a compromised poller must not be able to forge a trigger." Half of this
|
||||
// table is forgery attempts; all of them must be refused.
|
||||
func TestOpsRulePredicates(t *testing.T) {
|
||||
@@ -331,7 +331,7 @@ func TestNoDefaultRuleFiresOnEmptyState(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// Severities are the delivery contract (DESIGN.md § Delivery / channel
|
||||
// Severities are the delivery contract (docs/design.md § Delivery / channel
|
||||
// routing): care is sev1-2 and drops when away, ops is sev3-4 and holds. Pin
|
||||
// them so a change to a rule's insistence has to be deliberate.
|
||||
func TestDefaultRuleSeverities(t *testing.T) {
|
||||
@@ -356,7 +356,7 @@ func TestDefaultRuleSeverities(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// Cooldown bounds keep the feedback tuner honest — DESIGN.md wants
|
||||
// Cooldown bounds keep the feedback tuner honest — docs/design.md wants
|
||||
// `cooldown in [min,max]` "so a weird week can't mutate Maven silent or
|
||||
// stalker". A base outside its own envelope would make that meaningless.
|
||||
func TestDefaultRuleCooldownsAreBounded(t *testing.T) {
|
||||
|
||||
@@ -0,0 +1,126 @@
|
||||
package memory
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"unicode"
|
||||
)
|
||||
|
||||
// stopwords — words that carry no topic. A question and a note that share only
|
||||
// these share nothing: "почему небо синее" and "сеть какая-то медленная" both
|
||||
// contain "какая"-shaped filler and are about different worlds.
|
||||
var stopwords = map[string]bool{
|
||||
// interrogatives and demonstratives
|
||||
"что": true, "чего": true, "какой": true, "какая": true, "какое": true,
|
||||
"какие": true, "каких": true, "кто": true, "кого": true, "кому": true,
|
||||
"почему": true, "зачем": true, "где": true, "куда": true, "откуда": true,
|
||||
"когда": true, "сколько": true, "как": true, "то": true, "это": true,
|
||||
"этот": true, "тот": true, "там": true, "тут": true, "такой": true,
|
||||
// pronouns — every sentence he says is about him, so "я" is not a topic
|
||||
"я": true, "меня": true, "мне": true, "мой": true, "моя": true, "мои": true,
|
||||
"ты": true, "тебя": true, "тебе": true, "твой": true, "он": true, "она": true,
|
||||
"они": true, "мы": true, "себя": true, "свой": true,
|
||||
// prepositions, conjunctions, particles, copulas
|
||||
"в": true, "во": true, "на": true, "с": true, "со": true, "у": true,
|
||||
"о": true, "об": true, "про": true, "за": true, "из": true, "по": true,
|
||||
"до": true, "от": true, "для": true, "над": true, "под": true, "при": true,
|
||||
"и": true, "а": true, "но": true, "или": true, "же": true, "ли": true,
|
||||
"не": true, "ни": true, "бы": true, "был": true, "была": true, "было": true,
|
||||
"быть": true, "есть": true, "был-ли": true, "уже": true, "ещё": true,
|
||||
"еще": true, "так": true, "вот": true, "там-же": true,
|
||||
// English filler, for the mixed utterances he does say
|
||||
"the": true, "a": true, "an": true, "is": true, "are": true, "was": true,
|
||||
"were": true, "be": true, "of": true, "in": true, "on": true, "at": true,
|
||||
"to": true, "for": true, "about": true, "and": true, "or": true, "not": true,
|
||||
"what": true, "who": true, "why": true, "when": true, "where": true,
|
||||
"which": true, "how": true, "i": true, "my": true, "me": true, "it": true,
|
||||
"this": true, "that": true,
|
||||
}
|
||||
|
||||
// firstPerson — the words that make an utterance a question about his own
|
||||
// life. Not possession only: "как я восстановил конфиги" owns nothing and is
|
||||
// still about him.
|
||||
var firstPerson = map[string]bool{
|
||||
"я": true, "меня": true, "мне": true, "мной": true, "мой": true,
|
||||
"моя": true, "моё": true, "мое": true, "мои": true, "моего": true,
|
||||
"моей": true, "моих": true, "моим": true, "себя": true, "свой": true,
|
||||
"своя": true, "свои": true, "своего": true, "мною": true,
|
||||
"i": true, "me": true, "my": true, "mine": true, "myself": true,
|
||||
}
|
||||
|
||||
// RecallAllowed is the second half of the recall gate (#470). A hit that
|
||||
// cleared the score and margin gate may still be about something else
|
||||
// entirely: the held-out fixture puts the right note at 0.791-0.890 and the
|
||||
// must-be-silent cases at 0.795-0.835, so no threshold sits between them, and
|
||||
// a note about his slow network answered "почему небо синее?".
|
||||
//
|
||||
// The veto applies only to a question that mentions nothing of his. That
|
||||
// restriction is what keeps the fix from costing more than it saves: recall
|
||||
// exists to find the note whose words he no longer remembers, and demanding a
|
||||
// shared word of every recall silenced four true recalls on the fixture to
|
||||
// kill one false one. A question about his own life keeps the embedder alone
|
||||
// as its judge. A question about the world has to name something the memory
|
||||
// actually mentions.
|
||||
func RecallAllowed(query, text string) bool {
|
||||
if mentionsHim(query) {
|
||||
return true
|
||||
}
|
||||
return SharesContentWord(query, text)
|
||||
}
|
||||
|
||||
func mentionsHim(query string) bool {
|
||||
for _, w := range strings.FieldsFunc(strings.ToLower(query), func(r rune) bool {
|
||||
return !unicode.IsLetter(r) && !unicode.IsDigit(r)
|
||||
}) {
|
||||
if firstPerson[w] {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// SharesContentWord reports whether query and text have at least one topic
|
||||
// word in common, after dropping the words that carry no topic. Stems are
|
||||
// compared, so the note and the question do not have to inflect alike.
|
||||
func SharesContentWord(query, text string) bool {
|
||||
q := contentWords(query)
|
||||
if len(q) == 0 {
|
||||
// Nothing to compare — a question made entirely of filler. The score
|
||||
// gate is then the only judge it can have.
|
||||
return true
|
||||
}
|
||||
t := contentWords(text)
|
||||
for _, a := range q {
|
||||
for _, b := range t {
|
||||
if a == b || sameStem(a, b) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
func contentWords(s string) []string {
|
||||
var out []string
|
||||
for _, w := range strings.FieldsFunc(strings.ToLower(s), func(r rune) bool {
|
||||
return !unicode.IsLetter(r) && !unicode.IsDigit(r)
|
||||
}) {
|
||||
if !stopwords[w] {
|
||||
out = append(out, w)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// sameStem is inflection and derivation tolerance: Russian marks case and
|
||||
// tense on the ending, and the note and the question rarely use the same form.
|
||||
// "воду" and "вода" are the same water, "кормить" and "корм" the same feeding.
|
||||
// All but the last rune of the shorter word must match, and never fewer than
|
||||
// three, which is what keeps "сеть" clear of "сеанс".
|
||||
func sameStem(a, b string) bool {
|
||||
ar, br := []rune(a), []rune(b)
|
||||
n := min(len(ar), len(br)) - 1
|
||||
if n < 3 {
|
||||
return false
|
||||
}
|
||||
return string(ar[:n]) == string(br[:n])
|
||||
}
|
||||
@@ -0,0 +1,40 @@
|
||||
package memory
|
||||
|
||||
import "testing"
|
||||
|
||||
func TestRecallAllowed(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
query, text string
|
||||
want bool
|
||||
}{
|
||||
// The #470 shape: a world question and a note about his box.
|
||||
{"world question, unrelated note", "почему небо синее", "сеть какая-то медленная", false},
|
||||
{"world question, unrelated fact", "какая столица Франции", "какая последняя версия языка Go", false},
|
||||
{"silent fixture case", "во сколько отходит поезд", "бэкап запускается в три ночи", false},
|
||||
|
||||
// A world question that does name the topic keeps its answer.
|
||||
{"world question, same topic", "какой поезд идёт в Минск", "поезда в Минск ходят утром", true},
|
||||
|
||||
// A question about his own life is judged by the embedder alone,
|
||||
// because recall exists for words he no longer remembers.
|
||||
{"about him, no shared word", "во сколько я обычно засыпаю", "ложусь около одиннадцати", true},
|
||||
{"about him, english", "which colour scheme do i like", "тёмная тема везде", true},
|
||||
|
||||
// Inflection must not break a match.
|
||||
{"inflected", "чем кормить кота", "корм для кота в шкафу", true},
|
||||
}
|
||||
for _, c := range cases {
|
||||
if got := RecallAllowed(c.query, c.text); got != c.want {
|
||||
t.Errorf("%s: RecallAllowed(%q, %q) = %v, want %v", c.name, c.query, c.text, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A question made only of filler has no topic word to match on, and the score
|
||||
// gate is then the only judge it can have.
|
||||
func TestRecallAllowedFallsBackWhenNothingToCompare(t *testing.T) {
|
||||
if !RecallAllowed("что это", "сеть какая-то медленная") {
|
||||
t.Error("a question with no content word must not be vetoed")
|
||||
}
|
||||
}
|
||||
@@ -378,7 +378,7 @@ func scoreCase(ctx context.Context, emb router.Embedder, newStore NewStore, minS
|
||||
if len(hits) > 1 {
|
||||
o.Margin = hits[0].Score - hits[1].Score
|
||||
}
|
||||
o.Recalled = bestRecall(hits, minScore, minMargin)
|
||||
o.Recalled = bestRecall(c.Query, hits, minScore, minMargin)
|
||||
}
|
||||
for i, h := range hits {
|
||||
if h.ID != c.Want {
|
||||
@@ -424,11 +424,19 @@ func rankNote(inTop3 bool) string {
|
||||
// is not importable; recalleval_test.go asserts the two agree in behaviour.
|
||||
// The daemon returns the whole hit (a note and a fact are said differently);
|
||||
// the harness only scores what came back, so it keeps returning the text.
|
||||
func bestRecall(results []memory.Result, minScore, minMargin float64) string {
|
||||
// bestRecall mirrors the daemon's gate in cmd/mavend/recall.go, including the
|
||||
// topic veto added for #470: a score that clears the gate still has to be
|
||||
// about what he asked. Keep the two in step — a fixture that measures a
|
||||
// weaker gate than the daemon runs flatters it.
|
||||
func bestRecall(query string, results []memory.Result, minScore, minMargin float64) string {
|
||||
if !memory.Confident(results, minScore, minMargin) {
|
||||
return ""
|
||||
}
|
||||
return results[0].Meta["text"]
|
||||
text := results[0].Meta["text"]
|
||||
if !memory.RecallAllowed(query, text) {
|
||||
return ""
|
||||
}
|
||||
return text
|
||||
}
|
||||
|
||||
func bump(m map[string]TagStat, key string, pass bool) {
|
||||
|
||||
@@ -140,21 +140,22 @@ func words(s string) []string {
|
||||
|
||||
// TestBestRecallMatchesDaemon — the harness duplicates bestRecall from
|
||||
// cmd/mavend/recall.go (package main is not importable). This pins the copy to
|
||||
// the original's three rules: no hits, below the gate, or no text ⇒ silence.
|
||||
// the original's rules: no hits, below the gate, no text, or no shared topic
|
||||
// word ⇒ silence.
|
||||
func TestBestRecallMatchesDaemon(t *testing.T) {
|
||||
if got := bestRecall(nil, 0.55, 0); got != "" {
|
||||
if got := bestRecall("чай", nil, 0.55, 0); got != "" {
|
||||
t.Errorf("no hits: got %q, want silence", got)
|
||||
}
|
||||
low := []memory.Result{{ID: "a", Score: 0.4, Meta: map[string]string{"text": "чай"}}}
|
||||
if got := bestRecall(low, 0.55, 0); got != "" {
|
||||
if got := bestRecall("чай", low, 0.55, 0); got != "" {
|
||||
t.Errorf("below gate: got %q, want silence", got)
|
||||
}
|
||||
noText := []memory.Result{{ID: "a", Score: 0.9, Meta: map[string]string{}}}
|
||||
if got := bestRecall(noText, 0.55, 0); got != "" {
|
||||
if got := bestRecall("чай", noText, 0.55, 0); got != "" {
|
||||
t.Errorf("no text: got %q, want silence", got)
|
||||
}
|
||||
ok := []memory.Result{{ID: "a", Score: 0.9, Meta: map[string]string{"text": "чай"}}}
|
||||
if got := bestRecall(ok, 0.55, 0); got != "чай" {
|
||||
if got := bestRecall("чай", ok, 0.55, 0); got != "чай" {
|
||||
t.Errorf("above gate: got %q, want %q", got, "чай")
|
||||
}
|
||||
// Margin: a close runner-up means the embedder cannot tell the two apart,
|
||||
@@ -163,17 +164,23 @@ func TestBestRecallMatchesDaemon(t *testing.T) {
|
||||
{ID: "a", Score: 0.86, Meta: map[string]string{"text": "чай"}},
|
||||
{ID: "b", Score: 0.85, Meta: map[string]string{"text": "кофе"}},
|
||||
}
|
||||
if got := bestRecall(close, 0.55, 0.03); got != "" {
|
||||
if got := bestRecall("чай", close, 0.55, 0.03); got != "" {
|
||||
t.Errorf("thin margin: got %q, want silence", got)
|
||||
}
|
||||
if got := bestRecall(close, 0.55, 0); got != "чай" {
|
||||
if got := bestRecall("чай", close, 0.55, 0); got != "чай" {
|
||||
t.Errorf("margin off: got %q, want %q", got, "чай")
|
||||
}
|
||||
// The topic veto (#470): the score is fine and the note is about
|
||||
// something else.
|
||||
offTopic := []memory.Result{{ID: "a", Score: 0.9, Meta: map[string]string{"text": "сеть какая-то медленная"}}}
|
||||
if got := bestRecall("почему небо синее", offTopic, 0.55, 0); got != "" {
|
||||
t.Errorf("off topic: got %q, want silence", got)
|
||||
}
|
||||
clear := []memory.Result{
|
||||
{ID: "a", Score: 0.86, Meta: map[string]string{"text": "чай"}},
|
||||
{ID: "b", Score: 0.70, Meta: map[string]string{"text": "кофе"}},
|
||||
}
|
||||
if got := bestRecall(clear, 0.55, 0.03); got != "чай" {
|
||||
if got := bestRecall("чай", clear, 0.55, 0.03); got != "чай" {
|
||||
t.Errorf("wide margin: got %q, want %q", got, "чай")
|
||||
}
|
||||
}
|
||||
|
||||
@@ -41,7 +41,7 @@
|
||||
"tags": ["preference", "homelab", "paraphrase", "hard"],
|
||||
"query": "когда запускать резервное копирование",
|
||||
"want": "n1",
|
||||
"note": "The DESIGN.md preference-seam example, phrased as the operator would ask it later.",
|
||||
"note": "The docs/design.md preference-seam example, phrased as the operator would ask it later.",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "бэкапы лучше делать ночью в три часа", "kind": "note"},
|
||||
{"id": "n2", "text": "обновления ставлю по субботам", "kind": "note"},
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
// Package morning is maven's morning routine engine — item #3 off the
|
||||
// 2026-07-20 backlog (see Maven/20-07-2026-BACKLOG.md).
|
||||
// 2026-07-20 backlog (Vikunja #280).
|
||||
//
|
||||
// A Routine is NOT four independent reminder timers. It's a checklist for a
|
||||
// daily window: several Items, each evidenced by a fact key, completed in
|
||||
|
||||
@@ -0,0 +1,277 @@
|
||||
// Package netaddr parses a daemon seam address and dials or binds it.
|
||||
//
|
||||
// Every seam between Maven's daemons used to be a unix socket with the
|
||||
// network hardcoded at the call site — two dials in internal/ipc, one listen,
|
||||
// and the same pair again in internal/worker. That is correct for co-located
|
||||
// daemons and it is the reason a module cannot live on another host. This
|
||||
// package moves the choice into the address string so a deploy picks the
|
||||
// transport, not a recompile:
|
||||
//
|
||||
// /run/maven/stt.sock unix (the default, unchanged)
|
||||
// unix:///run/maven/stt.sock unix (explicit, same thing)
|
||||
// tcp://workstation:9310?token=hunter2 tcp
|
||||
//
|
||||
// A scheme-less address is unix and behaves exactly as it did before this
|
||||
// package existed: same 0700 parent dir, same 0600 socket, same bytes on the
|
||||
// wire with no handshake in front of them.
|
||||
//
|
||||
// Over TCP the filesystem permission that authenticated the unix socket is
|
||||
// gone, and what crosses this seam is audio of the owner speaking and the
|
||||
// text of his turns. So a TCP seam carries a shared token, checked before the
|
||||
// first protocol frame is read. Wireguard is supported underneath and is not
|
||||
// required.
|
||||
package netaddr
|
||||
|
||||
import (
|
||||
"crypto/subtle"
|
||||
"errors"
|
||||
"fmt"
|
||||
"net"
|
||||
"net/url"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"golang.org/x/sys/unix"
|
||||
)
|
||||
|
||||
// ErrUnauthorized — the peer presented a token the listener does not accept,
|
||||
// or presented none when one is required.
|
||||
var ErrUnauthorized = errors.New("netaddr: unauthorized")
|
||||
|
||||
// Addr is a parsed seam endpoint.
|
||||
type Addr struct {
|
||||
// Network is "unix" or "tcp".
|
||||
Network string
|
||||
// Address is the socket path (unix) or host:port (tcp).
|
||||
Address string
|
||||
// Token is the shared secret for a tcp seam. Empty for unix, where the
|
||||
// filesystem does the same job.
|
||||
Token string
|
||||
}
|
||||
|
||||
// String renders the address for logs and errors. The token is never included.
|
||||
func (a Addr) String() string {
|
||||
if a.Network == "unix" {
|
||||
return a.Address
|
||||
}
|
||||
return a.Network + "://" + a.Address
|
||||
}
|
||||
|
||||
// IsUnix reports whether this seam is a unix socket, and so is local, is
|
||||
// authenticated by file permissions, and needs no handshake.
|
||||
func (a Addr) IsUnix() bool { return a.Network == "unix" }
|
||||
|
||||
// Parse reads a seam address. Anything without a "scheme://" prefix is a unix
|
||||
// socket path, which keeps every existing config and every default working
|
||||
// untouched.
|
||||
func Parse(s string) (Addr, error) {
|
||||
if !strings.Contains(s, "://") {
|
||||
return Addr{Network: "unix", Address: s}, nil
|
||||
}
|
||||
u, err := url.Parse(s)
|
||||
if err != nil {
|
||||
return Addr{}, fmt.Errorf("netaddr: parse %q: %w", s, err)
|
||||
}
|
||||
switch u.Scheme {
|
||||
case "unix":
|
||||
return Addr{Network: "unix", Address: u.Path}, nil
|
||||
case "tcp":
|
||||
if u.Host == "" {
|
||||
return Addr{}, fmt.Errorf("netaddr: %q has no host:port", s)
|
||||
}
|
||||
return Addr{Network: "tcp", Address: u.Host, Token: u.Query().Get("token")}, nil
|
||||
default:
|
||||
return Addr{}, fmt.Errorf("netaddr: unsupported scheme %q", u.Scheme)
|
||||
}
|
||||
}
|
||||
|
||||
// MustParse is Parse for a literal known good at compile time. It panics on a
|
||||
// bad address, so use it in tests and constants, never on config input.
|
||||
func MustParse(s string) Addr {
|
||||
a, err := Parse(s)
|
||||
if err != nil {
|
||||
panic(err)
|
||||
}
|
||||
return a
|
||||
}
|
||||
|
||||
// handshakeTimeout bounds the token exchange. A peer that cannot write one
|
||||
// short line in this long is not going to serve a turn either.
|
||||
const handshakeTimeout = 5 * time.Second
|
||||
|
||||
// greeting prefixes the token line. Versioned so a later mTLS seam can be
|
||||
// told apart from this one on the wire.
|
||||
const greeting = "MAVEN1 "
|
||||
|
||||
// Dial connects to a. On a tcp seam it sends the token and waits for the
|
||||
// listener to accept it, so a returned conn is already authorized and the
|
||||
// caller can write its first protocol frame.
|
||||
func Dial(a Addr) (net.Conn, error) {
|
||||
return DialTimeout(a, 0)
|
||||
}
|
||||
|
||||
// DialTimeout is Dial with a bound on the connect. Zero means the operating
|
||||
// system default. The token exchange gets its own timeout either way.
|
||||
func DialTimeout(a Addr, timeout time.Duration) (net.Conn, error) {
|
||||
var c net.Conn
|
||||
var err error
|
||||
if timeout > 0 {
|
||||
c, err = net.DialTimeout(a.Network, a.Address, timeout)
|
||||
} else {
|
||||
c, err = net.Dial(a.Network, a.Address)
|
||||
}
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if a.IsUnix() {
|
||||
return c, nil
|
||||
}
|
||||
if err := clientHandshake(c, a.Token); err != nil {
|
||||
_ = c.Close()
|
||||
return nil, err
|
||||
}
|
||||
return c, nil
|
||||
}
|
||||
|
||||
func clientHandshake(c net.Conn, token string) error {
|
||||
_ = c.SetDeadline(time.Now().Add(handshakeTimeout))
|
||||
defer c.SetDeadline(time.Time{})
|
||||
if _, err := c.Write([]byte(greeting + token + "\n")); err != nil {
|
||||
return fmt.Errorf("netaddr: send token: %w", err)
|
||||
}
|
||||
var reply [3]byte
|
||||
if _, err := readFull(c, reply[:]); err != nil {
|
||||
return fmt.Errorf("%w: %v", ErrUnauthorized, err)
|
||||
}
|
||||
if string(reply[:]) != "ok\n" {
|
||||
return ErrUnauthorized
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// Listener wraps a net.Listener so Accept performs the token check for a tcp
|
||||
// seam. A connection that fails the check is closed and never surfaces, so
|
||||
// the protocol above this layer only ever sees authorized peers.
|
||||
type Listener struct {
|
||||
net.Listener
|
||||
addr Addr
|
||||
}
|
||||
|
||||
// Accept returns the next authorized connection. Unauthorized peers are
|
||||
// dropped and Accept keeps waiting: a bad token is a rejected stranger, not a
|
||||
// reason to stop serving.
|
||||
func (l *Listener) Accept() (net.Conn, error) {
|
||||
for {
|
||||
c, err := l.Listener.Accept()
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if l.addr.IsUnix() {
|
||||
return c, nil
|
||||
}
|
||||
if err := serverHandshake(c, l.addr.Token); err != nil {
|
||||
_ = c.Close()
|
||||
continue
|
||||
}
|
||||
return c, nil
|
||||
}
|
||||
}
|
||||
|
||||
// Addr reports the parsed seam address this listener was built from.
|
||||
func (l *Listener) SeamAddr() Addr { return l.addr }
|
||||
|
||||
func serverHandshake(c net.Conn, want string) error {
|
||||
_ = c.SetDeadline(time.Now().Add(handshakeTimeout))
|
||||
defer c.SetDeadline(time.Time{})
|
||||
// The line is bounded: greeting, token, newline. Read a byte at a time so
|
||||
// nothing of the first protocol frame is consumed when the token is short.
|
||||
line := make([]byte, 0, 128)
|
||||
var b [1]byte
|
||||
for {
|
||||
if _, err := readFull(c, b[:]); err != nil {
|
||||
return err
|
||||
}
|
||||
if b[0] == '\n' {
|
||||
break
|
||||
}
|
||||
line = append(line, b[0])
|
||||
if len(line) > 512 {
|
||||
return ErrUnauthorized
|
||||
}
|
||||
}
|
||||
got, ok := strings.CutPrefix(string(line), greeting)
|
||||
if !ok {
|
||||
return ErrUnauthorized
|
||||
}
|
||||
if subtle.ConstantTimeCompare([]byte(got), []byte(want)) != 1 {
|
||||
return ErrUnauthorized
|
||||
}
|
||||
if _, err := c.Write([]byte("ok\n")); err != nil {
|
||||
return err
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func readFull(c net.Conn, p []byte) (int, error) {
|
||||
n := 0
|
||||
for n < len(p) {
|
||||
m, err := c.Read(p[n:])
|
||||
n += m
|
||||
if err != nil {
|
||||
return n, err
|
||||
}
|
||||
}
|
||||
return n, nil
|
||||
}
|
||||
|
||||
// Listen binds a. A unix seam gets the perms it has always had: parent dir
|
||||
// 0700, socket 0600, and any stale socket from a crashed daemon removed
|
||||
// first. A tcp seam must carry a token, because there is no filesystem to
|
||||
// stand in for one.
|
||||
func Listen(a Addr) (*Listener, error) {
|
||||
if a.IsUnix() {
|
||||
ln, err := listenUnix(a.Address)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return &Listener{Listener: ln, addr: a}, nil
|
||||
}
|
||||
if a.Token == "" {
|
||||
return nil, fmt.Errorf("netaddr: listen %s: tcp seam requires a token", a)
|
||||
}
|
||||
ln, err := net.Listen("tcp", a.Address)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("netaddr: listen %s: %w", a, err)
|
||||
}
|
||||
return &Listener{Listener: ln, addr: a}, nil
|
||||
}
|
||||
|
||||
func listenUnix(path string) (net.Listener, error) {
|
||||
_ = os.Remove(path) // stale socket from a crashed daemon; ignore missing
|
||||
if err := os.MkdirAll(filepath.Dir(path), 0o700); err != nil {
|
||||
return nil, fmt.Errorf("netaddr: mkdir socket dir: %w", err)
|
||||
}
|
||||
// umask could widen the perms on socket creation; tighten then chmod to
|
||||
// be explicit. 0600 ⇒ read+write by owner only.
|
||||
oldMask := unix.Umask(0o077)
|
||||
ln, err := net.Listen("unix", path)
|
||||
unix.Umask(oldMask)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("netaddr: listen %s: %w", path, err)
|
||||
}
|
||||
if err := os.Chmod(path, 0o600); err != nil {
|
||||
_ = ln.Close()
|
||||
_ = os.Remove(path)
|
||||
return nil, fmt.Errorf("netaddr: chmod socket: %w", err)
|
||||
}
|
||||
return ln, nil
|
||||
}
|
||||
|
||||
// Cleanup removes the socket file behind a unix seam. It is a no-op for tcp.
|
||||
func Cleanup(a Addr) {
|
||||
if a.IsUnix() && a.Address != "" {
|
||||
_ = os.Remove(a.Address)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,185 @@
|
||||
package netaddr
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"net"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// A scheme-less address must stay unix. Every deploy in the tree writes a bare
|
||||
// path, so this is the test that says the transport change costs them nothing.
|
||||
func TestParseSchemelessIsUnix(t *testing.T) {
|
||||
a, err := Parse("/run/maven/stt.sock")
|
||||
if err != nil {
|
||||
t.Fatalf("parse: %v", err)
|
||||
}
|
||||
if !a.IsUnix() {
|
||||
t.Fatalf("want unix, got %q", a.Network)
|
||||
}
|
||||
if a.Address != "/run/maven/stt.sock" {
|
||||
t.Fatalf("address = %q", a.Address)
|
||||
}
|
||||
if a.Token != "" {
|
||||
t.Fatalf("unix seam carries a token: %q", a.Token)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParse(t *testing.T) {
|
||||
cases := []struct {
|
||||
in string
|
||||
net, addr, tk string
|
||||
wantErr bool
|
||||
}{
|
||||
{in: "", net: "unix", addr: ""},
|
||||
{in: "unix:///run/maven/core.sock", net: "unix", addr: "/run/maven/core.sock"},
|
||||
{in: "tcp://workstation:9310", net: "tcp", addr: "workstation:9310"},
|
||||
{in: "tcp://workstation:9310?token=hunter2", net: "tcp", addr: "workstation:9310", tk: "hunter2"},
|
||||
{in: "tcp://", wantErr: true},
|
||||
{in: "udp://workstation:9310", wantErr: true},
|
||||
}
|
||||
for _, c := range cases {
|
||||
a, err := Parse(c.in)
|
||||
if c.wantErr {
|
||||
if err == nil {
|
||||
t.Errorf("Parse(%q) = %v, want error", c.in, a)
|
||||
}
|
||||
continue
|
||||
}
|
||||
if err != nil {
|
||||
t.Errorf("Parse(%q): %v", c.in, err)
|
||||
continue
|
||||
}
|
||||
if a.Network != c.net || a.Address != c.addr || a.Token != c.tk {
|
||||
t.Errorf("Parse(%q) = %+v, want %s/%s/%s", c.in, a, c.net, c.addr, c.tk)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The token must never reach a log line.
|
||||
func TestStringHidesToken(t *testing.T) {
|
||||
a := MustParse("tcp://workstation:9310?token=hunter2")
|
||||
if got := a.String(); got != "tcp://workstation:9310" {
|
||||
t.Fatalf("String() = %q", got)
|
||||
}
|
||||
}
|
||||
|
||||
// A unix seam must round-trip with no handshake in front of the payload: the
|
||||
// first bytes the listener sees are the caller's, exactly as before.
|
||||
func TestUnixRoundTripHasNoHandshake(t *testing.T) {
|
||||
a := MustParse(filepath.Join(t.TempDir(), "s.sock"))
|
||||
ln, err := Listen(a)
|
||||
if err != nil {
|
||||
t.Fatalf("listen: %v", err)
|
||||
}
|
||||
defer ln.Close()
|
||||
go echoOnce(ln)
|
||||
|
||||
c, err := Dial(a)
|
||||
if err != nil {
|
||||
t.Fatalf("dial: %v", err)
|
||||
}
|
||||
defer c.Close()
|
||||
if got := roundTrip(t, c, "hello"); got != "hello" {
|
||||
t.Fatalf("got %q", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTCPRoundTripWithToken(t *testing.T) {
|
||||
ln, addr := listenLoopback(t, "s3cret")
|
||||
defer ln.Close()
|
||||
go echoOnce(ln)
|
||||
|
||||
c, err := Dial(addr)
|
||||
if err != nil {
|
||||
t.Fatalf("dial: %v", err)
|
||||
}
|
||||
defer c.Close()
|
||||
if got := roundTrip(t, c, "hello"); got != "hello" {
|
||||
t.Fatalf("got %q", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTCPWrongTokenIsRejected(t *testing.T) {
|
||||
ln, addr := listenLoopback(t, "s3cret")
|
||||
defer ln.Close()
|
||||
// Accept keeps waiting past the bad peer, so nothing here should ever
|
||||
// reach the echo. A conn that does means the token was not checked.
|
||||
go echoOnce(ln)
|
||||
|
||||
bad := addr
|
||||
bad.Token = "wrong"
|
||||
if _, err := Dial(bad); !errors.Is(err, ErrUnauthorized) {
|
||||
t.Fatalf("dial with wrong token: err = %v, want ErrUnauthorized", err)
|
||||
}
|
||||
}
|
||||
|
||||
// A stranger that speaks the protocol instead of the greeting is dropped, and
|
||||
// the listener stays up for the peer that follows it.
|
||||
func TestTCPUngreetedPeerDoesNotKillTheListener(t *testing.T) {
|
||||
ln, addr := listenLoopback(t, "s3cret")
|
||||
defer ln.Close()
|
||||
go echoOnce(ln)
|
||||
|
||||
raw, err := net.Dial("tcp", addr.Address)
|
||||
if err != nil {
|
||||
t.Fatalf("raw dial: %v", err)
|
||||
}
|
||||
if _, err := raw.Write([]byte("GET / HTTP/1.1\n")); err != nil {
|
||||
t.Fatalf("raw write: %v", err)
|
||||
}
|
||||
raw.Close()
|
||||
|
||||
c, err := Dial(addr)
|
||||
if err != nil {
|
||||
t.Fatalf("dial after stranger: %v", err)
|
||||
}
|
||||
defer c.Close()
|
||||
if got := roundTrip(t, c, "still here"); got != "still here" {
|
||||
t.Fatalf("got %q", got)
|
||||
}
|
||||
}
|
||||
|
||||
// A tcp seam with no token is a misconfiguration, and it must fail at bind
|
||||
// rather than serve the owner's turns to anyone who connects.
|
||||
func TestTCPListenRequiresToken(t *testing.T) {
|
||||
if _, err := Listen(MustParse("tcp://127.0.0.1:0")); err == nil {
|
||||
t.Fatal("listen on a tokenless tcp seam succeeded")
|
||||
}
|
||||
}
|
||||
|
||||
func listenLoopback(t *testing.T, token string) (*Listener, Addr) {
|
||||
t.Helper()
|
||||
ln, err := Listen(Addr{Network: "tcp", Address: "127.0.0.1:0", Token: token})
|
||||
if err != nil {
|
||||
t.Fatalf("listen: %v", err)
|
||||
}
|
||||
return ln, Addr{Network: "tcp", Address: ln.Addr().String(), Token: token}
|
||||
}
|
||||
|
||||
func echoOnce(ln *Listener) {
|
||||
c, err := ln.Accept()
|
||||
if err != nil {
|
||||
return
|
||||
}
|
||||
defer c.Close()
|
||||
buf := make([]byte, 256)
|
||||
n, err := c.Read(buf)
|
||||
if err != nil {
|
||||
return
|
||||
}
|
||||
_, _ = c.Write(buf[:n])
|
||||
}
|
||||
|
||||
func roundTrip(t *testing.T, c net.Conn, msg string) string {
|
||||
t.Helper()
|
||||
if _, err := c.Write([]byte(msg)); err != nil {
|
||||
t.Fatalf("write: %v", err)
|
||||
}
|
||||
buf := make([]byte, 256)
|
||||
n, err := c.Read(buf)
|
||||
if err != nil {
|
||||
t.Fatalf("read: %v", err)
|
||||
}
|
||||
return string(buf[:n])
|
||||
}
|
||||
@@ -15,7 +15,7 @@ const (
|
||||
CheckLang = "lang" // the operator's language, not the prompt's
|
||||
CheckLength = "length" // a nudge is one sentence, not a paragraph
|
||||
CheckFeminine = "feminine" // her self-reference is feminine (hard constraint)
|
||||
CheckCringe = "cringe" // DESIGN.md § Non-goals, "not a relationship"
|
||||
CheckCringe = "cringe" // docs/design.md § Non-goals, "not a relationship"
|
||||
CheckOnTopic = "ontopic" // says the thing the rule is about
|
||||
|
||||
// CheckHisGender — the other half of the persona rule: SHE is feminine, HE
|
||||
@@ -105,7 +105,7 @@ func checkLength(body string) Result {
|
||||
|
||||
// --- feminine self-reference ---------------------------------------------
|
||||
//
|
||||
// The hard constraint (CLAUDE.md, DESIGN.md § Identity): Maven's Russian
|
||||
// The hard constraint (CLAUDE.md, docs/design.md § Identity): Maven's Russian
|
||||
// self-reference is feminine. The operator is male, so second-person forms
|
||||
// addressed to him are MASCULINE and must not be flagged — "ты не пил воду" is
|
||||
// correct, "я напомнил" is not. Both directions matter, which is why this is a
|
||||
@@ -497,7 +497,7 @@ func isLatinWord(w string) bool {
|
||||
|
||||
// --- the cringe checks ---------------------------------------------------
|
||||
//
|
||||
// "Think Jarvis without the cringe part". DESIGN.md § Non-goals: "Not a
|
||||
// "Think Jarvis without the cringe part". docs/design.md § Non-goals: "Not a
|
||||
// relationship — mom-tone is a function that makes nudges land, not emotional
|
||||
// company. Names the drift a warm small model falls into." Each pattern below
|
||||
// is one shape of that drift. They are deliberately specific: a check that
|
||||
|
||||
@@ -11,7 +11,7 @@
|
||||
// length test that a human can read and disagree with. A score here is a claim
|
||||
// about measurable properties, not about whether a sentence is good.
|
||||
//
|
||||
// DESIGN.md § "Rules decide, LLM phrases" is why there is no send/veto signal
|
||||
// docs/design.md § "Rules decide, LLM phrases" is why there is no send/veto signal
|
||||
// anywhere in this package: the rule already decided she speaks. The phraser
|
||||
// only words it, so a nudge the model refuses to write is a failure, never a
|
||||
// legitimate outcome.
|
||||
@@ -38,7 +38,7 @@ var fixtureJSON []byte
|
||||
const SchemaVersion = 1
|
||||
|
||||
// Case — one nudge situation, as a real tick would present it. The fields are
|
||||
// the (rule, severity, context) input DESIGN.md names, flattened to JSON.
|
||||
// the (rule, severity, context) input docs/design.md names, flattened to JSON.
|
||||
//
|
||||
// WantAny is the on-topic contract: at least one of these lowercased fragments
|
||||
// must appear in the message. A water nudge that never mentions water is a
|
||||
|
||||
@@ -46,6 +46,12 @@ type LLMPhraser struct {
|
||||
launch func(ctx context.Context, cfg Config) (backend, error)
|
||||
probe func(ctx context.Context, base string) (string, error)
|
||||
|
||||
// remote — the workstation model, when one is configured. Set once at wiring
|
||||
// time by UseRemote and read on every phrasing call. nil ⇒ every call goes to
|
||||
// the resident llama-server this phraser owns, which is the whole deploy
|
||||
// before a `workstation` block exists. See world.go.
|
||||
remote Remote
|
||||
|
||||
// swapMu — single-flight around Swap. Held for the whole swap, including the
|
||||
// model load, so two concurrent swap requests can never both be loading.
|
||||
swapMu sync.Mutex
|
||||
@@ -363,10 +369,7 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
|
||||
// prompt guaranteed to make a small model fill the gap from memory.
|
||||
notes = nonEmpty(notes)
|
||||
if len(notes) == 0 {
|
||||
// General knowledge — no notes to ground the answer. The system
|
||||
// prompt is the single tested source in router.KnowledgePrompt.
|
||||
sys := persona.Prepend(p.cfg.ContextBlock, router.KnowledgePrompt())
|
||||
prompt := fmt.Sprintf("Пользователь спрашивает: \"%s\".", utterance)
|
||||
sys, prompt := p.knowledgePrompt(utterance)
|
||||
resp, err := p.chatWithSystem(ctx, sys, prompt, 768)
|
||||
if err != nil || resp == "" {
|
||||
return "не знаю.", nil
|
||||
@@ -381,11 +384,7 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
|
||||
}
|
||||
return resp, nil
|
||||
}
|
||||
sys := p.querySystemPrompt()
|
||||
prompt := fmt.Sprintf(
|
||||
"Он спрашивает: \"%s\"\n\nИсточники:\n%s\nОтветь ему коротко и своими словами, опираясь только на эти источники. Если ответа в них нет — так и скажи.",
|
||||
utterance, evidenceBlock(notes),
|
||||
)
|
||||
sys, prompt := p.evidencePrompt(utterance, notes)
|
||||
resp, err := p.chatWithSystem(ctx, sys, prompt, 768)
|
||||
text, _, perr := parseResponseMood(resp)
|
||||
if err != nil || perr != nil {
|
||||
@@ -481,6 +480,17 @@ func chatSystemPrompt(block func() string) string {
|
||||
// the LLM completion endpoint. Like chatWithSystem but for an arbitrary message
|
||||
// slice — the caller owns the system prompt placement.
|
||||
func (p *LLMPhraser) chatWithMessages(ctx context.Context, msgs []chatMsg, maxTokens int) (string, error) {
|
||||
// Same silent preference as chatWithSystem, when the array is the shape
|
||||
// llm.Req can carry: one system turn and one user turn. PhraseChat already
|
||||
// folds the history into a single user message (some chat templates reject
|
||||
// consecutive user turns), so today that is every call. A longer array goes
|
||||
// to the resident model rather than get flattened here, because flattening a
|
||||
// conversation is a decision its owner should make.
|
||||
if len(msgs) == 2 && msgs[0].Role == "system" && msgs[1].Role == "user" {
|
||||
if out, ok := p.remoteChat(ctx, msgs[0].Content, msgs[1].Content, maxTokens); ok {
|
||||
return out, nil
|
||||
}
|
||||
}
|
||||
base, release, err := p.acquire()
|
||||
if err != nil {
|
||||
return "", err
|
||||
@@ -534,8 +544,14 @@ func (p *LLMPhraser) PhraseReminder(ctx context.Context, d loop.ReminderDecision
|
||||
text = "reminder"
|
||||
}
|
||||
|
||||
// Russian, like the other two prompts (Vikunja #404). Asking a model for a
|
||||
// Russian reply in English is asking it to switch languages mid-prompt,
|
||||
// and a 1.7B sometimes answers in the language it was asked in. The
|
||||
// persona rules and the JSON contract are not repeated here: this call
|
||||
// goes through chat(), so nudgeSystem already states both, and a second
|
||||
// statement of the same contract is one more thing that can drift.
|
||||
prompt := fmt.Sprintf(
|
||||
`The user set a reminder: "%s". Rephrase it briefly as a gentle nudge. Respond as JSON: {"response": "...", "mood": "..."}`,
|
||||
`Он поставил напоминание: "%s". Скажи это своими словами, коротко и мягко — одно предложение.`,
|
||||
text,
|
||||
)
|
||||
resp, err := p.chat(ctx, prompt)
|
||||
@@ -634,6 +650,12 @@ func (p *LLMPhraser) chat(ctx context.Context, userPrompt string) (string, error
|
||||
}
|
||||
|
||||
func (p *LLMPhraser) chatWithSystem(ctx context.Context, system, user string, maxTokens int) (string, error) {
|
||||
// The workstation model first when it will take work, and silently: every
|
||||
// caller of this helper is on the silent half of the degradation rule. It
|
||||
// answering is not news, and it being asleep is not news either.
|
||||
if out, ok := p.remoteChat(ctx, system, user, maxTokens); ok {
|
||||
return out, nil
|
||||
}
|
||||
base, release, err := p.acquire()
|
||||
if err != nil {
|
||||
return "", err
|
||||
@@ -692,7 +714,7 @@ func (p *LLMPhraser) chatWithSystem(ctx context.Context, system, user string, ma
|
||||
// Written as filled-in examples, not as a schema with "..." in it. A 0.8B
|
||||
// copies whatever sits in the response slot, so a literal placeholder there
|
||||
// teaches it to answer with the placeholder. Measured: 7/15 nudges came back
|
||||
// as "..." before this. See PHRASING-EVAL-31-07-2026.md.
|
||||
// as "..." before this. See docs/evals/2026-07-31-phrasing.md.
|
||||
//
|
||||
// Russian only, feminine self-reference, second person masculine (the owner is
|
||||
// a man). She talks TO him, informally, singular — never "вы", never "он".
|
||||
@@ -727,6 +749,27 @@ func (p *LLMPhraser) systemPrompt() string {
|
||||
return persona.Prepend(p.cfg.ContextBlock, nudgeSystem)
|
||||
}
|
||||
|
||||
// knowledgePrompt — the no-sources branch: a world question, answered from
|
||||
// weights alone. The system prompt is the single tested source in
|
||||
// router.KnowledgePrompt.
|
||||
//
|
||||
// Split out of PhraseQuery so PhraseWorld sends the workstation model the same
|
||||
// bytes the resident model gets. Prompt parity across two models is a stated
|
||||
// constraint (CLAUDE.md), and two copies of a prompt is how it stops holding.
|
||||
func (p *LLMPhraser) knowledgePrompt(utterance string) (sys, user string) {
|
||||
return persona.Prepend(p.cfg.ContextBlock, router.KnowledgePrompt()),
|
||||
fmt.Sprintf("Пользователь спрашивает: \"%s\".", utterance)
|
||||
}
|
||||
|
||||
// evidencePrompt — the sources branch: read these, add nothing. Shared with
|
||||
// PhraseWorld for the same reason as knowledgePrompt.
|
||||
func (p *LLMPhraser) evidencePrompt(utterance string, notes []string) (sys, user string) {
|
||||
return p.querySystemPrompt(), fmt.Sprintf(
|
||||
"Он спрашивает: \"%s\"\n\nИсточники:\n%s\nОтветь ему коротко и своими словами, опираясь только на эти источники. Если ответа в них нет — так и скажи.",
|
||||
utterance, evidenceBlock(notes),
|
||||
)
|
||||
}
|
||||
|
||||
// querySystemPrompt returns the system prompt for the evidence branch of
|
||||
// PhraseQuery. Prepends the configured persona when set.
|
||||
//
|
||||
@@ -754,7 +797,7 @@ func (p *LLMPhraser) querySystemPrompt() string {
|
||||
base := "Ты отвечаешь ему по источникам, которые тебе дали. Отвечай ТОЛЬКО по ним: всё, что ты говоришь, должно быть написано в источниках. " +
|
||||
"Если ответа в них нет — так и скажи и на этом остановись; не добавляй ничего из своих знаний и не догадывайся. " +
|
||||
"Не приплетай прошлые реплики разговора. " +
|
||||
"Отвечай по-русски, коротко и своими словами, начинай с \"вот что я нашла: \". О себе — в женском роде, глаголы в прошедшем времени с окончанием -ла. Он мужчина, обращайся к нему на \"ты\". Respond ONLY with valid JSON: {\"response\": \"...\", \"mood\": \"neutral\"}."
|
||||
"Отвечай по-русски, коротко и своими словами, начинай с \"вот что я нашла: \". О себе — в женском роде, глаголы в прошедшем времени с окончанием -ла. Он мужчина, обращайся к нему на \"ты\". Отвечай ТОЛЬКО одним объектом JSON: {\"response\": \"...\", \"mood\": \"neutral\"}."
|
||||
return persona.Prepend(p.cfg.ContextBlock, base)
|
||||
}
|
||||
|
||||
|
||||
@@ -1,9 +1,9 @@
|
||||
// Package phraser is maven's "rules decide, llm phrases" seam — the layer
|
||||
// that turns a loop decision into the body + summary the delivery module ships.
|
||||
//
|
||||
// Per DESIGN.md § Resident language model: the phraser is the resident model
|
||||
// Per docs/design.md § Resident language model: the phraser is the resident model
|
||||
// (Qwen3-1.7B — RU continued pretraining plus joint persona/router SFT, not a
|
||||
// sub-1b prompted-only model as the retired spec claimed; see DESIGN.md
|
||||
// sub-1b prompted-only model as the retired spec claimed; see docs/design.md
|
||||
// § Superseded, "small-model phrasing claim"). It takes
|
||||
// (rule, severity, context) and produces Body (full voice message, local — no
|
||||
// shoulder-surf concern beyond who's in the room) + Summary (minimal body for
|
||||
|
||||
@@ -0,0 +1,256 @@
|
||||
package phraser
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"regexp"
|
||||
"strconv"
|
||||
"strings"
|
||||
"syscall"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// The spawn path (NewLLMPhraser, spawnLlamaServer, startLlamaProc, llamaProc.Close)
|
||||
// was at 0% coverage: every test built the phraser with NewLLMPhraserAt, which
|
||||
// starts no process. These tests drive the real spawn code against a fake
|
||||
// llama-server script, so the startup race arms and the reaping are exercised
|
||||
// without a model or a GPU.
|
||||
|
||||
// fakeLlama writes an executable script standing in for llama-server and returns
|
||||
// its path. body runs after the script has recorded its own pid.
|
||||
func fakeLlama(t *testing.T, body string) string {
|
||||
t.Helper()
|
||||
dir := t.TempDir()
|
||||
path := filepath.Join(dir, "fake-llama-server")
|
||||
script := "#!/bin/sh\n" + body + "\n"
|
||||
if err := os.WriteFile(path, []byte(script), 0o755); err != nil {
|
||||
t.Fatalf("write fake server: %v", err)
|
||||
}
|
||||
return path
|
||||
}
|
||||
|
||||
// listensThenSleeps prints the line startLlamaProc scrapes, then stays alive
|
||||
// until killed — the shape of a real llama-server that came up.
|
||||
const listensThenSleeps = `echo "srv load_model: listening on http://127.0.0.1:18081" >&2
|
||||
while : ; do sleep 1 ; done`
|
||||
|
||||
func testCfg(bin string) Config {
|
||||
cfg := DefaultConfig("/nonexistent/model.gguf")
|
||||
cfg.BinPath = bin
|
||||
return cfg
|
||||
}
|
||||
|
||||
func TestExtractPort(t *testing.T) {
|
||||
for _, tc := range []struct{ in, want string }{
|
||||
{"127.0.0.1:0", "0"},
|
||||
{"127.0.0.1:8080", "8080"},
|
||||
{"127.0.0.1:", "0"},
|
||||
{"", "0"},
|
||||
{"8080", "0"}, // no colon: Cut yields no port, so the caller gets the "any port" default
|
||||
} {
|
||||
if got := extractPort(tc.in); got != tc.want {
|
||||
t.Errorf("extractPort(%q) = %q, want %q", tc.in, got, tc.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestStartLlamaProcScrapesPortAndReaps(t *testing.T) {
|
||||
bin := fakeLlama(t, listensThenSleeps)
|
||||
ctx, cancel := context.WithCancel(context.Background())
|
||||
defer cancel()
|
||||
|
||||
p, err := startLlamaProc(ctx, testCfg(bin))
|
||||
if err != nil {
|
||||
t.Fatalf("startLlamaProc: %v", err)
|
||||
}
|
||||
if p.BaseURL() != "http://127.0.0.1:18081" {
|
||||
t.Fatalf("BaseURL = %q, want the scraped address", p.BaseURL())
|
||||
}
|
||||
pid := p.cmd.Process.Pid
|
||||
|
||||
p.cancel = cancel
|
||||
if err := p.Close(); err != nil {
|
||||
t.Fatalf("Close: %v", err)
|
||||
}
|
||||
// Close must Wait, otherwise the child lingers as a zombie.
|
||||
if p.cmd.ProcessState == nil {
|
||||
t.Fatal("Close did not reap the child: ProcessState is nil")
|
||||
}
|
||||
if err := syscall.Kill(pid, 0); err == nil {
|
||||
t.Fatalf("child %d still exists after Close", pid)
|
||||
}
|
||||
}
|
||||
|
||||
func TestStartLlamaProcFailureArms(t *testing.T) {
|
||||
t.Run("binary missing", func(t *testing.T) {
|
||||
cfg := testCfg(filepath.Join(t.TempDir(), "does-not-exist"))
|
||||
_, err := startLlamaProc(context.Background(), cfg)
|
||||
if err == nil || !strings.Contains(err.Error(), "llm: start") {
|
||||
t.Fatalf("err = %v, want a start failure", err)
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("server exits without listening", func(t *testing.T) {
|
||||
// stderr closes, so the reader goroutine reports EOF on errCh.
|
||||
bin := fakeLlama(t, `echo "ggml_vulkan: no device" >&2
|
||||
exit 1`)
|
||||
_, err := startLlamaProc(context.Background(), testCfg(bin))
|
||||
if err == nil || !strings.Contains(err.Error(), "llm: server output") {
|
||||
t.Fatalf("err = %v, want the server-output arm", err)
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("context cancelled during startup", func(t *testing.T) {
|
||||
// Never prints the listen line and never exits: only ctx can end this.
|
||||
bin := fakeLlama(t, `while : ; do sleep 1 ; done`)
|
||||
ctx, cancel := context.WithCancel(context.Background())
|
||||
go func() {
|
||||
time.Sleep(150 * time.Millisecond)
|
||||
cancel()
|
||||
}()
|
||||
defer cancel()
|
||||
_, err := startLlamaProc(ctx, testCfg(bin))
|
||||
if !errors.Is(err, context.Canceled) {
|
||||
t.Fatalf("err = %v, want context.Canceled", err)
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
func TestNewLLMPhraserSpawns(t *testing.T) {
|
||||
bin := fakeLlama(t, listensThenSleeps)
|
||||
p, err := NewLLMPhraser(context.Background(), testCfg(bin))
|
||||
if err != nil {
|
||||
t.Fatalf("NewLLMPhraser: %v", err)
|
||||
}
|
||||
if p.BaseURL() != "http://127.0.0.1:18081" {
|
||||
t.Fatalf("BaseURL = %q", p.BaseURL())
|
||||
}
|
||||
pid := p.be.(*llamaProc).cmd.Process.Pid
|
||||
if err := p.Close(); err != nil {
|
||||
t.Fatalf("Close: %v", err)
|
||||
}
|
||||
if p.BaseURL() != "" {
|
||||
t.Fatalf("BaseURL after Close = %q, want empty", p.BaseURL())
|
||||
}
|
||||
if err := syscall.Kill(pid, 0); err == nil {
|
||||
t.Fatalf("llama-server %d survived Close", pid)
|
||||
}
|
||||
}
|
||||
|
||||
func TestNewLLMPhraserSpawnFailure(t *testing.T) {
|
||||
cfg := testCfg(filepath.Join(t.TempDir(), "does-not-exist"))
|
||||
p, err := NewLLMPhraser(context.Background(), cfg)
|
||||
if err == nil {
|
||||
p.Close()
|
||||
t.Fatal("want an error when the server cannot start")
|
||||
}
|
||||
if p != nil {
|
||||
t.Fatalf("want a nil phraser on failure, got %#v", p)
|
||||
}
|
||||
}
|
||||
|
||||
// TestPdeathsigKillsOrphan is the orphan test the task asked for. A SIGKILLed
|
||||
// mavend never runs Close, so nothing but the kernel's Pdeathsig can stop its
|
||||
// llama-server. Re-exec this test binary as the "daemon", let it spawn the fake
|
||||
// server, SIGKILL the daemon, and assert the grandchild died with it.
|
||||
func TestPdeathsigKillsOrphan(t *testing.T) {
|
||||
bin := fakeLlama(t, listensThenSleeps)
|
||||
|
||||
cmd := exec.Command(os.Args[0], "-test.run=TestSpawnHelperProcess", "-test.v=false")
|
||||
cmd.Env = append(os.Environ(), "MAVEN_SPAWN_HELPER=1", "MAVEN_FAKE_LLAMA="+bin)
|
||||
out, err := cmd.StdoutPipe()
|
||||
if err != nil {
|
||||
t.Fatalf("stdout pipe: %v", err)
|
||||
}
|
||||
if err := cmd.Start(); err != nil {
|
||||
t.Fatalf("start helper: %v", err)
|
||||
}
|
||||
defer func() { _ = cmd.Process.Kill(); _ = cmd.Wait() }()
|
||||
|
||||
buf := make([]byte, 256)
|
||||
n, err := out.Read(buf)
|
||||
if err != nil {
|
||||
t.Fatalf("read child pid: %v", err)
|
||||
}
|
||||
childPID, err := strconv.Atoi(strings.TrimSpace(string(buf[:n])))
|
||||
if err != nil {
|
||||
t.Fatalf("helper printed %q, want a pid: %v", string(buf[:n]), err)
|
||||
}
|
||||
if err := syscall.Kill(childPID, 0); err != nil {
|
||||
t.Fatalf("llama-server %d not running before the kill: %v", childPID, err)
|
||||
}
|
||||
|
||||
// SIGKILL: the helper gets no chance to clean up, exactly like an OOM kill.
|
||||
if err := cmd.Process.Signal(syscall.SIGKILL); err != nil {
|
||||
t.Fatalf("kill helper: %v", err)
|
||||
}
|
||||
_, _ = cmd.Process.Wait()
|
||||
|
||||
deadline := time.Now().Add(5 * time.Second)
|
||||
for time.Now().Before(deadline) {
|
||||
if err := syscall.Kill(childPID, 0); err != nil {
|
||||
return // gone: Pdeathsig did its job
|
||||
}
|
||||
time.Sleep(20 * time.Millisecond)
|
||||
}
|
||||
_ = syscall.Kill(childPID, syscall.SIGKILL)
|
||||
t.Fatalf("llama-server %d outlived the SIGKILLed parent", childPID)
|
||||
}
|
||||
|
||||
// TestSpawnHelperProcess is not a test. It is the child half of
|
||||
// TestPdeathsigKillsOrphan: spawn a llama-server, print its pid, then block.
|
||||
func TestSpawnHelperProcess(t *testing.T) {
|
||||
if os.Getenv("MAVEN_SPAWN_HELPER") != "1" {
|
||||
t.Skip("helper for TestPdeathsigKillsOrphan")
|
||||
}
|
||||
cfg := testCfg(os.Getenv("MAVEN_FAKE_LLAMA"))
|
||||
p, err := startLlamaProc(context.Background(), cfg)
|
||||
if err != nil {
|
||||
fmt.Println("spawn failed:", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
fmt.Println(p.cmd.Process.Pid)
|
||||
os.Stdout.Sync()
|
||||
select {} // wait to be killed
|
||||
}
|
||||
|
||||
// TestKillMavenScriptMatchesRealCommandLine pins kill-maven.sh's fallback
|
||||
// pattern to the command line startLlamaProc actually builds. The script leaked
|
||||
// orphans twice already, both times because the pattern stopped matching: first
|
||||
// `llama-server.*maven`, then a hardcoded model name after the model was swapped.
|
||||
func TestKillMavenScriptMatchesRealCommandLine(t *testing.T) {
|
||||
src, err := os.ReadFile("../../kill-maven.sh")
|
||||
if err != nil {
|
||||
t.Fatalf("read kill-maven.sh: %v", err)
|
||||
}
|
||||
m := regexp.MustCompile(`(?m)^\s*LLM='([^']+)'`).FindSubmatch(src)
|
||||
if m == nil {
|
||||
t.Fatal("no default LLM='...' pattern in kill-maven.sh")
|
||||
}
|
||||
pat, err := regexp.Compile(string(m[1]))
|
||||
if err != nil {
|
||||
t.Fatalf("LLM pattern %q does not compile: %v", m[1], err)
|
||||
}
|
||||
|
||||
// Rebuild the command line from the production arg list, so a change to
|
||||
// startLlamaProc that breaks the sweep fails here instead of on the box.
|
||||
cfg := DefaultConfig("/opt/maven/models/llm/Qwen3-1.7B-UD-Q4_K_XL.gguf")
|
||||
cfg.NCtx, cfg.NGpuLayers = 4096, 99
|
||||
cmdline := strings.Join([]string{
|
||||
cfg.BinPath,
|
||||
"-m", cfg.ModelPath,
|
||||
"--host", "127.0.0.1",
|
||||
"--port", extractPort(cfg.Listen),
|
||||
"-c", fmt.Sprintf("%d", cfg.NCtx),
|
||||
"-ngl", fmt.Sprintf("%d", cfg.NGpuLayers),
|
||||
"--no-webui",
|
||||
}, " ")
|
||||
if !pat.MatchString(cmdline) {
|
||||
t.Fatalf("kill-maven.sh pattern %q does not match %q — orphans would leak", m[1], cmdline)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,137 @@
|
||||
package phraser
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"log"
|
||||
|
||||
"github.com/kami/maven/internal/llm"
|
||||
)
|
||||
|
||||
// Remote — the workstation model, seen from the phraser. `*llm.Pair` satisfies
|
||||
// it, and a test fake satisfies it in three lines.
|
||||
//
|
||||
// Only the refusing half of Pair is here on purpose. Pair.Complete falls back to
|
||||
// its own floor client, and the phraser already owns a floor: the llama-server it
|
||||
// spawned. Two floors under one call is one too many, so the phraser asks whether
|
||||
// the remote will take work, uses it when it will, and otherwise does exactly
|
||||
// what it did before this file existed.
|
||||
type Remote interface {
|
||||
// Available is an atomic read of a cached probe, so it is free to call per
|
||||
// turn. See llm.Pair.
|
||||
Available() bool
|
||||
// CompleteRemote runs on the workstation or returns ErrRemoteUnavailable. It
|
||||
// never falls back.
|
||||
CompleteRemote(ctx context.Context, r llm.Req) (string, error)
|
||||
}
|
||||
|
||||
// ErrNoWorldModel — a world question was asked, a workstation model is
|
||||
// configured to answer it, and that machine is not answering. The caller turns
|
||||
// this into a gap he is told about ("не могу сейчас"), never into an answer from
|
||||
// the resident model.
|
||||
//
|
||||
// This is the naming half of the degradation rule in docs/offload.md. The
|
||||
// resident Qwen3-1.7B does not answer a world question worse than the 12B, it
|
||||
// invents: measured, the workstation model scores knowledge 9/9 on the talk
|
||||
// fixture against the resident model's confabulations
|
||||
// (docs/evals/2026-08-02-workstation-gemma4-12b.md).
|
||||
var ErrNoWorldModel = errors.New("phraser: no world model available")
|
||||
|
||||
// chatTemperature — what the phraser's own transport has always sampled at.
|
||||
// Named so the remote path cannot drift from it silently. Whether 0.7 is right
|
||||
// at all is Vikunja #402, and answering that here would hide a phrasing change
|
||||
// inside a routing change.
|
||||
const chatTemperature = 0.7
|
||||
|
||||
// UseRemote points the phraser at the workstation model. Wiring time only, once,
|
||||
// before anything phrases: the field is read without a lock on every call
|
||||
// because a per-turn lock to answer a question that changes at deploy time is
|
||||
// not worth paying for.
|
||||
//
|
||||
// A nil remote is the normal state of a box with no `workstation` block, and it
|
||||
// must behave exactly as the box behaved before this seam existed.
|
||||
func (p *LLMPhraser) UseRemote(r Remote) {
|
||||
p.remote = r
|
||||
}
|
||||
|
||||
// PhraseWorld answers a question about the world — either from the model's own
|
||||
// knowledge (no sources) or from a passage someone fetched (a live search, a ZIM
|
||||
// article, a page he named). Three outcomes, and the middle one is the point:
|
||||
//
|
||||
// - No workstation configured. The resident model answers, exactly as it does
|
||||
// today. Naming a gap needs a gap: on a box that never had a second model,
|
||||
// refusing every world question would remove a capability he has now.
|
||||
// - Workstation configured and taking work. It answers.
|
||||
// - Workstation configured and down. ErrNoWorldModel, and the caller says so.
|
||||
//
|
||||
// The prompts are the ones PhraseQuery uses, built by the same two functions, so
|
||||
// the two models are asked the same question in the same words.
|
||||
func (p *LLMPhraser) PhraseWorld(ctx context.Context, utterance string, sources []string) (string, error) {
|
||||
sources = nonEmpty(sources)
|
||||
if p.remote == nil {
|
||||
return p.PhraseQuery(ctx, utterance, sources)
|
||||
}
|
||||
var sys, user string
|
||||
if len(sources) == 0 {
|
||||
sys, user = p.knowledgePrompt(utterance)
|
||||
} else {
|
||||
sys, user = p.evidencePrompt(utterance, sources)
|
||||
}
|
||||
if !p.remote.Available() {
|
||||
return "", ErrNoWorldModel
|
||||
}
|
||||
resp, err := p.remote.CompleteRemote(ctx, llm.Req{
|
||||
System: sys,
|
||||
User: user,
|
||||
Grammar: p.grammar(),
|
||||
MaxTokens: 768,
|
||||
Temperature: chatTemperature,
|
||||
})
|
||||
if err != nil {
|
||||
// The cached probe was one interval stale, or the card went away
|
||||
// mid-request. Either way this is the gap, not an error to log and
|
||||
// paper over with the smaller model.
|
||||
log.Printf("phraser: world model: %v", err)
|
||||
return "", errors.Join(ErrNoWorldModel, err)
|
||||
}
|
||||
resp = stripThink(resp)
|
||||
text, _, perr := parseResponseMood(resp)
|
||||
if perr != nil {
|
||||
log.Printf("phraser: PhraseWorld: %v", perr)
|
||||
return "", errors.Join(ErrNoWorldModel, perr)
|
||||
}
|
||||
if text != "" {
|
||||
return text, nil
|
||||
}
|
||||
if resp == "" {
|
||||
return "", ErrNoWorldModel
|
||||
}
|
||||
return resp, nil
|
||||
}
|
||||
|
||||
// remoteChat is the silent half, for the phrasing paths where the workstation
|
||||
// model is only better: a nudge, a reminder, a reply, a question answered from
|
||||
// his own notes. It reports whether it answered; it never reports why not,
|
||||
// because the caller's next move is the resident model either way.
|
||||
//
|
||||
// He is not told which of the two models phrased his reply. That is the rule.
|
||||
func (p *LLMPhraser) remoteChat(ctx context.Context, system, user string, maxTokens int) (string, bool) {
|
||||
if p.remote == nil || !p.remote.Available() {
|
||||
return "", false
|
||||
}
|
||||
out, err := p.remote.CompleteRemote(ctx, llm.Req{
|
||||
System: system,
|
||||
User: user,
|
||||
Grammar: p.grammar(),
|
||||
MaxTokens: maxTokens,
|
||||
Temperature: chatTemperature,
|
||||
})
|
||||
if err != nil {
|
||||
log.Printf("phraser: workstation model declined, phrasing here instead: %v", err)
|
||||
return "", false
|
||||
}
|
||||
if out = stripThink(out); out == "" {
|
||||
return "", false
|
||||
}
|
||||
return out, true
|
||||
}
|
||||
@@ -0,0 +1,164 @@
|
||||
package phraser
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/kami/maven/internal/llm"
|
||||
"github.com/kami/maven/internal/loop"
|
||||
)
|
||||
|
||||
// fakeRemote — a workstation model that is up or down on command, and records
|
||||
// what it was asked.
|
||||
type fakeRemote struct {
|
||||
up bool
|
||||
reply string
|
||||
err error
|
||||
got []llm.Req
|
||||
}
|
||||
|
||||
func (f *fakeRemote) Available() bool { return f.up }
|
||||
|
||||
func (f *fakeRemote) CompleteRemote(_ context.Context, r llm.Req) (string, error) {
|
||||
f.got = append(f.got, r)
|
||||
if f.err != nil {
|
||||
return "", f.err
|
||||
}
|
||||
return f.reply, nil
|
||||
}
|
||||
|
||||
// The three outcomes of the naming half, in one place. The middle one is the
|
||||
// whole task: a gap he is told about, not an answer from the smaller model.
|
||||
func TestPhraseWorldNamesTheGapOnlyWhenThereIsOne(t *testing.T) {
|
||||
answer := `{"response": "Небо голубое из-за рэлеевского рассеяния.", "mood": "neutral"}`
|
||||
|
||||
t.Run("no workstation configured: the resident model answers as today", func(t *testing.T) {
|
||||
spy := newPromptSpy(t)
|
||||
p := NewLLMPhraserAt(spy.srv.URL, Config{})
|
||||
got, err := p.PhraseWorld(context.Background(), "почему небо голубое", nil)
|
||||
if err != nil {
|
||||
t.Fatalf("PhraseWorld: %v", err)
|
||||
}
|
||||
if got == "" {
|
||||
t.Fatal("no reply from the resident model")
|
||||
}
|
||||
if len(spy.user) != 1 {
|
||||
t.Fatalf("resident model saw %d requests, want 1", len(spy.user))
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("workstation up: it answers and the resident model is not asked", func(t *testing.T) {
|
||||
spy := newPromptSpy(t)
|
||||
p := NewLLMPhraserAt(spy.srv.URL, Config{})
|
||||
remote := &fakeRemote{up: true, reply: answer}
|
||||
p.UseRemote(remote)
|
||||
got, err := p.PhraseWorld(context.Background(), "почему небо голубое", nil)
|
||||
if err != nil {
|
||||
t.Fatalf("PhraseWorld: %v", err)
|
||||
}
|
||||
if !strings.Contains(got, "рассеяния") {
|
||||
t.Errorf("reply is not the workstation's: %q", got)
|
||||
}
|
||||
if len(spy.user) != 0 {
|
||||
t.Errorf("the resident model was asked %d times, want 0", len(spy.user))
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("workstation down: the gap, and nothing invented", func(t *testing.T) {
|
||||
spy := newPromptSpy(t)
|
||||
p := NewLLMPhraserAt(spy.srv.URL, Config{})
|
||||
p.UseRemote(&fakeRemote{up: false})
|
||||
got, err := p.PhraseWorld(context.Background(), "почему небо голубое", nil)
|
||||
if !errors.Is(err, ErrNoWorldModel) {
|
||||
t.Fatalf("err = %v, want ErrNoWorldModel", err)
|
||||
}
|
||||
if got != "" {
|
||||
t.Errorf("got a reply %q with no world model", got)
|
||||
}
|
||||
if len(spy.user) != 0 {
|
||||
t.Errorf("the resident model answered a world question %d times, want 0", len(spy.user))
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("workstation errors mid-request: still the gap", func(t *testing.T) {
|
||||
spy := newPromptSpy(t)
|
||||
p := NewLLMPhraserAt(spy.srv.URL, Config{})
|
||||
p.UseRemote(&fakeRemote{up: true, err: errors.New("connection refused")})
|
||||
if _, err := p.PhraseWorld(context.Background(), "почему небо голубое", nil); !errors.Is(err, ErrNoWorldModel) {
|
||||
t.Fatalf("err = %v, want ErrNoWorldModel", err)
|
||||
}
|
||||
if len(spy.user) != 0 {
|
||||
t.Errorf("the resident model answered a world question %d times, want 0", len(spy.user))
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
// Prompt parity: the workstation model is asked the same question in the same
|
||||
// words, or the fixtures measure one thing and the daemon ships another.
|
||||
func TestPhraseWorldSendsTheSamePromptsAsPhraseQuery(t *testing.T) {
|
||||
spy := newPromptSpy(t)
|
||||
resident := NewLLMPhraserAt(spy.srv.URL, Config{})
|
||||
if _, err := resident.PhraseQuery(context.Background(), "кто написал войну и мир", []string{"Лев Толстой"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
remote := &fakeRemote{up: true, reply: `{"response": "Толстой.", "mood": "neutral"}`}
|
||||
offloaded := NewLLMPhraserAt(spy.srv.URL, Config{})
|
||||
offloaded.UseRemote(remote)
|
||||
if _, err := offloaded.PhraseWorld(context.Background(), "кто написал войну и мир", []string{"Лев Толстой"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
if len(remote.got) != 1 {
|
||||
t.Fatalf("the workstation saw %d requests, want 1", len(remote.got))
|
||||
}
|
||||
if remote.got[0].System != spy.system[0] {
|
||||
t.Errorf("system prompts differ:\nremote: %q\nresident: %q", remote.got[0].System, spy.system[0])
|
||||
}
|
||||
if remote.got[0].User != spy.user[0] {
|
||||
t.Errorf("user prompts differ:\nremote: %q\nresident: %q", remote.got[0].User, spy.user[0])
|
||||
}
|
||||
}
|
||||
|
||||
// The silent half. A nudge phrased on the workstation is not news, and one
|
||||
// phrased here because the card is busy is not news either — but it must be
|
||||
// sampled the same way, or the workstation quietly changes how she sounds.
|
||||
func TestNudgePhrasingPrefersTheWorkstationSilently(t *testing.T) {
|
||||
spy := newPromptSpy(t)
|
||||
p := NewLLMPhraserAt(spy.srv.URL, Config{LLMNudges: true})
|
||||
remote := &fakeRemote{up: true, reply: `{"response": "Выпей воды.", "mood": "neutral"}`}
|
||||
p.UseRemote(remote)
|
||||
|
||||
pn, err := p.PhraseNudge(context.Background(), loop.Candidate{Rule: loop.WaterRule(), Severity: loop.Sev1})
|
||||
if err != nil {
|
||||
t.Fatalf("PhraseNudge: %v", err)
|
||||
}
|
||||
if pn.Body != "Выпей воды." {
|
||||
t.Errorf("body = %q, want the workstation's wording", pn.Body)
|
||||
}
|
||||
if len(remote.got) != 1 {
|
||||
t.Fatalf("the workstation saw %d requests, want 1", len(remote.got))
|
||||
}
|
||||
if remote.got[0].Temperature != chatTemperature {
|
||||
t.Errorf("temperature = %v, want %v (what the resident transport samples at)",
|
||||
remote.got[0].Temperature, chatTemperature)
|
||||
}
|
||||
if len(spy.user) != 0 {
|
||||
t.Errorf("the resident model phrased %d nudges, want 0", len(spy.user))
|
||||
}
|
||||
}
|
||||
|
||||
func TestNudgePhrasingFallsBackWhenTheCardIsBusy(t *testing.T) {
|
||||
spy := newPromptSpy(t)
|
||||
p := NewLLMPhraserAt(spy.srv.URL, Config{LLMNudges: true})
|
||||
p.UseRemote(&fakeRemote{up: false})
|
||||
|
||||
if _, err := p.PhraseNudge(context.Background(), loop.Candidate{Rule: loop.WaterRule(), Severity: loop.Sev1}); err != nil {
|
||||
t.Fatalf("PhraseNudge: %v", err)
|
||||
}
|
||||
if len(spy.user) != 1 {
|
||||
t.Fatalf("the resident model phrased %d nudges, want 1", len(spy.user))
|
||||
}
|
||||
}
|
||||
@@ -39,7 +39,7 @@ import (
|
||||
// Re-measured with everything else held equal, thinking off scores exactly the
|
||||
// same, case for case — and a direct probe shows this llama-server build ignores
|
||||
// enable_thinking / reasoning_budget for this model anyway, so there was nothing
|
||||
// to turn off. Full write-up in ROUTING-EVAL-31-07-2026.md (Vikunja #376).
|
||||
// to turn off. Full write-up in docs/evals/2026-07-31-routing.md (Vikunja #376).
|
||||
func TestLLMRouterBaseline(t *testing.T) {
|
||||
base := os.Getenv("MAVEN_LLM_URL")
|
||||
if base == "" {
|
||||
|
||||
@@ -1,13 +1,13 @@
|
||||
// Package router is maven's reactive path — the cascade that turns a free-form
|
||||
// utterance into a deterministic Decision.
|
||||
//
|
||||
// Spec contract (from DESIGN.md § Reactive path — routing):
|
||||
// Spec contract (from docs/design.md § Reactive path — routing):
|
||||
//
|
||||
// - the TARGET design is LLM-as-router: the resident model (Qwen3-1.7B)
|
||||
// emits GBNF-constrained structured JSON for the route, and the same
|
||||
// model phrases replies; the embedder is a RAG hint, not a routing gate.
|
||||
// the classifier/embedder cascade below is the committed default today,
|
||||
// but it is an interim stopgap (DESIGN.md § Superseded, "classifier-owns-
|
||||
// but it is an interim stopgap (docs/design.md § Superseded, "classifier-owns-
|
||||
// the-route") and the known cause of weak RU query handling — not a
|
||||
// design to extend.
|
||||
// - a CASCADE, not one decider — layers:
|
||||
@@ -35,7 +35,7 @@ package router
|
||||
|
||||
import "time"
|
||||
|
||||
// Intent — the seven save-where labels from DESIGN.md's routing table. The
|
||||
// Intent — the seven save-where labels from docs/design.md's routing table. The
|
||||
// discriminator is "does the loop evaluate a predicate against it?":
|
||||
//
|
||||
// - act: command now, not stored (function call into the allowlist)
|
||||
|
||||
@@ -6,5 +6,5 @@ func KnowledgePrompt() string {
|
||||
// No self-introduction here: the shared persona block already says who she
|
||||
// is, and this line used to disagree with it — a different name ("Мавена")
|
||||
// and a masculine noun ("ассистент") in front of a feminine persona.
|
||||
return `Ответь кратко из своих знаний. Если не знаешь — скажи "не знаю". Не выдумывай. Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}.`
|
||||
return `Ответь кратко из своих знаний. Если не знаешь — скажи "не знаю". Не выдумывай. Отвечай ТОЛЬКО одним объектом JSON: {"response": "...", "mood": "neutral"}.`
|
||||
}
|
||||
|
||||
@@ -0,0 +1,64 @@
|
||||
package router
|
||||
|
||||
import "strings"
|
||||
|
||||
// interrogatives — the question words that mark an utterance as asking rather
|
||||
// than telling. Tokenized, never substring: "что" inside "чтобы" and "как"
|
||||
// inside "какао" are not questions.
|
||||
var interrogatives = []string{
|
||||
"что", "чего", "какой", "какая", "какое", "какие", "каких",
|
||||
"кто", "кого", "кому", "чей", "почему", "зачем", "отчего",
|
||||
"где", "куда", "откуда", "когда", "сколько", "как",
|
||||
"what", "who", "whom", "why", "when", "where", "which", "how",
|
||||
}
|
||||
|
||||
// narrativeRequests — "tell me about X" asks for knowledge Maven does not
|
||||
// hold about him. It carries no question mark and no interrogative, which is
|
||||
// how "расскажи про битву при Ватерлоо" reached the fact store (#470).
|
||||
var narrativeRequests = []string{
|
||||
"расскажи", "объясни", "опиши", "перечисли",
|
||||
"tell", "explain", "describe",
|
||||
}
|
||||
|
||||
// captureVerbs — an explicit instruction to record something. These win over
|
||||
// every test below, because "запиши что я пил воду" contains an interrogative
|
||||
// and is still a capture: the word he said is "запиши".
|
||||
var captureVerbs = []string{
|
||||
"запиши", "запомни", "отметь", "заметь", "добавь", "сохрани",
|
||||
"note", "remember", "log", "save",
|
||||
}
|
||||
|
||||
// IsQuestionShaped reports whether text asks for something rather than
|
||||
// records it. It is a deterministic offline test over tokens, so it costs
|
||||
// nothing and never depends on the model that produced the routing decision.
|
||||
//
|
||||
// It exists because a mis-routed question used to be persisted as a fact
|
||||
// about the owner, with the model's invented answer as the value (#470). The
|
||||
// predicate is deliberately blunt: refusing to store a question is cheap and
|
||||
// reversible, storing an invented fact about him is neither.
|
||||
func IsQuestionShaped(text string) bool {
|
||||
t := strings.TrimSpace(text)
|
||||
if t == "" {
|
||||
return false
|
||||
}
|
||||
toks := planTokens(t)
|
||||
for _, v := range captureVerbs {
|
||||
if hasTok(toks, v) {
|
||||
return false
|
||||
}
|
||||
}
|
||||
if strings.HasSuffix(t, "?") {
|
||||
return true
|
||||
}
|
||||
for _, w := range interrogatives {
|
||||
if hasTok(toks, w) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
for _, w := range narrativeRequests {
|
||||
if hasTok(toks, w) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
@@ -0,0 +1,48 @@
|
||||
package router
|
||||
|
||||
import "testing"
|
||||
|
||||
func TestIsQuestionShaped(t *testing.T) {
|
||||
// The seven utterances #470 recorded, plus the captures that must keep
|
||||
// working. A capture misread as a question loses a fact; a question
|
||||
// misread as a capture poisons recall, so the captures are the ones worth
|
||||
// pinning here.
|
||||
cases := []struct {
|
||||
text string
|
||||
want bool
|
||||
}{
|
||||
{"какая последняя версия языка Go?", true},
|
||||
{"что дальше?", true},
|
||||
{"расскажи про битву при Ватерлоо", true},
|
||||
{"почему небо синее?", true},
|
||||
{"какая столица Австралии?", true},
|
||||
{"кто такой Никола Тесла?", true},
|
||||
{"сколько стоит доллар", true},
|
||||
{"who is the premier of Japan", true},
|
||||
{"объясни линии Фраунгофера", true},
|
||||
|
||||
{"запиши что я пил воду", false},
|
||||
{"запомни какая у меня машина", false},
|
||||
{"отметь что я поужинал", false},
|
||||
{"поужинал", false},
|
||||
{"я выпил кофе", false},
|
||||
{"вода", false},
|
||||
{"привет", false},
|
||||
{"", false},
|
||||
}
|
||||
for _, c := range cases {
|
||||
if got := IsQuestionShaped(c.text); got != c.want {
|
||||
t.Errorf("IsQuestionShaped(%q) = %v, want %v", c.text, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Substring matching is what made the day-plan predicates wrong before, and
|
||||
// this predicate gates a write, so it gets the same guard.
|
||||
func TestIsQuestionShapedIsTokenized(t *testing.T) {
|
||||
for _, text := range []string{"чтобы не забыть, я полил кактус", "какао выпил"} {
|
||||
if IsQuestionShaped(text) {
|
||||
t.Errorf("IsQuestionShaped(%q) = true; a question word inside a longer word is not a question", text)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -6,6 +6,7 @@ import (
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"fmt"
|
||||
"log"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
@@ -46,6 +47,39 @@ func (s *Store) WriteFact(ctx context.Context, ts time.Time, kind FactKind, key,
|
||||
return id, nil
|
||||
}
|
||||
|
||||
// FactRecallText is the text a fact is indexed under and read back as (#493).
|
||||
//
|
||||
// It used to be the utterance that wrote the fact, so recall of ANY
|
||||
// voice-tapped fact answered with the sentence he said instead of the value
|
||||
// stored: `go_version = 1.20` was indexed as "какая последняя версия языка
|
||||
// Go?", and that question is what came back. The poisoned rows made the defect
|
||||
// visible; the shape was wrong for legitimate facts too.
|
||||
//
|
||||
// The key is spoken with its underscores dropped, because a key is written for
|
||||
// the store and this string is read out loud.
|
||||
func FactRecallText(key, value string) string {
|
||||
spoken := strings.TrimSpace(strings.ReplaceAll(key, "_", " "))
|
||||
v := strings.TrimSpace(DecodeFactValue(value))
|
||||
switch {
|
||||
case v == "":
|
||||
return spoken
|
||||
case spoken == "":
|
||||
return v
|
||||
}
|
||||
return spoken + " — " + v
|
||||
}
|
||||
|
||||
// DecodeFactValue unwraps a stored value for reading. The column holds raw json
|
||||
// when the writer serialized one (SetValue, CorrectValue) and a plain string
|
||||
// when it did not (a voice tap), so a reader that wants the text handles both.
|
||||
func DecodeFactValue(value string) string {
|
||||
var s string
|
||||
if err := json.Unmarshal([]byte(value), &s); err == nil {
|
||||
return s
|
||||
}
|
||||
return value
|
||||
}
|
||||
|
||||
// LatestFact returns the latest non-voided fact for key, or ErrNoFact.
|
||||
// "Non-voided" = no later row has voids_id pointing at it. We resolve this by
|
||||
// taking the newest row whose id is not referenced by any voids_id.
|
||||
@@ -290,6 +324,20 @@ func (s *Store) CorrectValue(ctx context.Context, key, source string, value any,
|
||||
if err != nil {
|
||||
return 0, fmt.Errorf("last insert id: %w", err)
|
||||
}
|
||||
// The same repair a void needs, for the same reason (#493). A correction
|
||||
// supersedes the value, and the vector still holds the old one, so recall
|
||||
// kept answering with the value he had just corrected. Dropping it costs
|
||||
// the key its recall vector until the fact is tapped again: this layer has
|
||||
// no embedder, and a missing vector loses a question while a stale one
|
||||
// answers it wrongly.
|
||||
//
|
||||
// Best-effort: the corrected row is committed, and a correction that lands
|
||||
// beats one that fails on cleanup.
|
||||
if n, derr := s.VectorMemory().DeletePrefix(ctx, "fact:"+key+":"); derr != nil {
|
||||
log.Printf("store: correct %q: memory vectors survive: %v", key, derr)
|
||||
} else if n > 0 {
|
||||
log.Printf("store: correct %q: dropped %d superseded memory vector(s)", key, n)
|
||||
}
|
||||
return newID, nil
|
||||
}
|
||||
|
||||
@@ -332,6 +380,20 @@ func (s *Store) VoidLatestFact(ctx context.Context, key, source string, ts time.
|
||||
if err != nil {
|
||||
return 0, 0, fmt.Errorf("void: last insert id: %w", err)
|
||||
}
|
||||
// The other half of the repair (#470). A fact reaches recall through a
|
||||
// vector keyed `fact:<key>:<unix>`, holding the utterance that wrote it.
|
||||
// Voiding the row alone left that vector answering questions, so revert
|
||||
// reported success on a box that stayed broken. Deleting every vector for
|
||||
// the key covers the earlier rows too: their values are superseded, and a
|
||||
// superseded value has no business claiming a turn.
|
||||
//
|
||||
// Best-effort by design: the audit trail is already committed, and a fact
|
||||
// that is voided but still recallable is better than a void that failed.
|
||||
if n, derr := s.VectorMemory().DeletePrefix(ctx, "fact:"+key+":"); derr != nil {
|
||||
log.Printf("store: void %q: memory vectors survive: %v", key, derr)
|
||||
} else if n > 0 {
|
||||
log.Printf("store: void %q: dropped %d memory vector(s)", key, n)
|
||||
}
|
||||
return oldID, newID, nil
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,177 @@
|
||||
package store
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"fmt"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
// metaKeyFactVectorShape names the shape the stored fact vectors were written
|
||||
// in. It exists so the repair below runs once per box instead of on every
|
||||
// start: the rows it fixes were written by a code path that no longer exists,
|
||||
// and once fixed nothing writes that shape again.
|
||||
const metaKeyFactVectorShape = "fact_vector_shape"
|
||||
|
||||
// factVectorShapeFact is the shape FactRecallText produces. Anything else in
|
||||
// the marker (including nothing, which is every box written before #493) means
|
||||
// the fact vectors still hold utterances.
|
||||
const factVectorShapeFact = "fact-text (#493)"
|
||||
|
||||
// FactVectorRepair is what one repair run did, for logging.
|
||||
type FactVectorRepair struct {
|
||||
Skipped bool // marker already matched — nothing to do
|
||||
Rewritten int // rows re-embedded from the fact they name
|
||||
Dropped int // rows deleted: voided, superseded, or naming no fact at all
|
||||
Kept int // rows already holding the right text
|
||||
Took time.Duration
|
||||
}
|
||||
|
||||
// RepairFactVectors brings the fact rows of memory_vectors in line with the
|
||||
// facts they name, and is the operator recovery a poisoned box had no path to
|
||||
// (#470 point 4, #493).
|
||||
//
|
||||
// Three defects put wrong text in that index, and all three are write-path
|
||||
// fixes that do nothing for rows already stored:
|
||||
//
|
||||
// - the indexed text was the utterance, so every fact row reads back a
|
||||
// sentence rather than a value;
|
||||
// - a void left its vector behind, so retracted junk kept answering;
|
||||
// - a correction left its vector behind, so the superseded value did.
|
||||
//
|
||||
// So each fact row is resolved against the fact store and one of three things
|
||||
// happens. It is dropped when the key has no fact, when the newest row for the
|
||||
// key is a void marker, or when a newer vector for the same key exists — a
|
||||
// superseded value has no business claiming a turn. It is re-embedded when its
|
||||
// text is not what FactRecallText says the fact is. Otherwise it is left alone.
|
||||
//
|
||||
// Idempotent, and safe to interrupt: every step compares before writing and the
|
||||
// marker is written last, so a run that dies partway is simply redone.
|
||||
func (s *Store) RepairFactVectors(ctx context.Context, embed EmbedFunc) (FactVectorRepair, error) {
|
||||
start := time.Now()
|
||||
var res FactVectorRepair
|
||||
|
||||
shape, err := s.Meta(ctx, metaKeyFactVectorShape)
|
||||
if err != nil {
|
||||
return res, err
|
||||
}
|
||||
if shape == factVectorShapeFact {
|
||||
res.Skipped = true
|
||||
res.Took = time.Since(start)
|
||||
return res, nil
|
||||
}
|
||||
|
||||
rows, err := s.db.QueryContext(ctx, `SELECT id, meta FROM memory_vectors`)
|
||||
if err != nil {
|
||||
return res, fmt.Errorf("repair fact vectors: read: %w", err)
|
||||
}
|
||||
type factVec struct {
|
||||
id, key string
|
||||
meta map[string]string
|
||||
ts int64
|
||||
}
|
||||
var vecs []factVec
|
||||
newest := map[string]int64{} // key → newest ts seen for it
|
||||
for rows.Next() {
|
||||
var id, metaJSON string
|
||||
if err := rows.Scan(&id, &metaJSON); err != nil {
|
||||
rows.Close()
|
||||
return res, fmt.Errorf("repair fact vectors: row: %w", err)
|
||||
}
|
||||
meta := map[string]string{}
|
||||
if err := json.Unmarshal([]byte(metaJSON), &meta); err != nil {
|
||||
rows.Close()
|
||||
return res, fmt.Errorf("repair fact vectors: meta for %q: %w", id, err)
|
||||
}
|
||||
if meta["type"] != "fact" {
|
||||
continue
|
||||
}
|
||||
key, ts, ok := splitFactVectorID(id)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
vecs = append(vecs, factVec{id: id, key: key, meta: meta, ts: ts})
|
||||
if ts > newest[key] {
|
||||
newest[key] = ts
|
||||
}
|
||||
}
|
||||
rows.Close()
|
||||
if err := rows.Err(); err != nil {
|
||||
return res, fmt.Errorf("repair fact vectors: rows: %w", err)
|
||||
}
|
||||
|
||||
for _, v := range vecs {
|
||||
drop := v.ts < newest[v.key]
|
||||
var want string
|
||||
if !drop {
|
||||
f, ferr := s.LatestFact(ctx, v.key)
|
||||
switch {
|
||||
case errors.Is(ferr, ErrNoFact):
|
||||
drop = true
|
||||
case ferr != nil:
|
||||
return res, fmt.Errorf("repair fact vectors: fact %q: %w", v.key, ferr)
|
||||
case DecodeFactValue(f.Value) == "voided":
|
||||
drop = true
|
||||
default:
|
||||
want = FactRecallText(v.key, f.Value)
|
||||
}
|
||||
}
|
||||
if drop {
|
||||
if err := s.VectorMemory().Delete(ctx, v.id); err != nil {
|
||||
return res, err
|
||||
}
|
||||
res.Dropped++
|
||||
continue
|
||||
}
|
||||
if v.meta["text"] == want {
|
||||
res.Kept++
|
||||
continue
|
||||
}
|
||||
vec, err := embed(ctx, want)
|
||||
if err != nil {
|
||||
return res, fmt.Errorf("repair fact vectors: embed %q: %w", v.id, err)
|
||||
}
|
||||
// The whole meta blob is rewritten in Go rather than patched in SQL,
|
||||
// because json_set needs the JSON1 extension and this store is opened
|
||||
// through sqlcipher.
|
||||
v.meta["text"] = want
|
||||
metaJSON, err := json.Marshal(v.meta)
|
||||
if err != nil {
|
||||
return res, fmt.Errorf("repair fact vectors: meta %q: %w", v.id, err)
|
||||
}
|
||||
if _, err := s.db.ExecContext(ctx,
|
||||
`UPDATE memory_vectors SET vec = ?, meta = ? WHERE id = ?`,
|
||||
encodeVec(vec), string(metaJSON), v.id); err != nil {
|
||||
return res, fmt.Errorf("repair fact vectors: write %q: %w", v.id, err)
|
||||
}
|
||||
res.Rewritten++
|
||||
}
|
||||
|
||||
if err := s.SetMeta(ctx, metaKeyFactVectorShape, factVectorShapeFact); err != nil {
|
||||
return res, err
|
||||
}
|
||||
res.Took = time.Since(start)
|
||||
return res, nil
|
||||
}
|
||||
|
||||
// splitFactVectorID reads the key and write time back out of a fact vector's
|
||||
// id, which the write path builds as `fact:<key>:<unix>`. A key may hold a
|
||||
// colon, the timestamp may not, so the split is from the right.
|
||||
func splitFactVectorID(id string) (key string, ts int64, ok bool) {
|
||||
rest, found := strings.CutPrefix(id, "fact:")
|
||||
if !found {
|
||||
return "", 0, false
|
||||
}
|
||||
cut := strings.LastIndex(rest, ":")
|
||||
if cut <= 0 {
|
||||
return "", 0, false
|
||||
}
|
||||
ts, err := strconv.ParseInt(rest[cut+1:], 10, 64)
|
||||
if err != nil {
|
||||
return "", 0, false
|
||||
}
|
||||
return rest[:cut], ts, true
|
||||
}
|
||||
@@ -0,0 +1,151 @@
|
||||
package store
|
||||
|
||||
import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// The write-path half of #493: recall of a fact must read back the fact, not
|
||||
// the sentence he happened to say.
|
||||
func TestFactRecallText(t *testing.T) {
|
||||
for _, tc := range []struct {
|
||||
name, key, value, want string
|
||||
}{
|
||||
{"json value", "go_version", `"1.20"`, "go version — 1.20"},
|
||||
{"plain value", "water", "выпил", "water — выпил"},
|
||||
{"no value", "shower", "", "shower"},
|
||||
{"underscores are spoken as spaces", "espresso_machine", `"чистая"`, "espresso machine — чистая"},
|
||||
} {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
if got := FactRecallText(tc.key, tc.value); got != tc.want {
|
||||
t.Fatalf("FactRecallText(%q, %q) = %q; want %q", tc.key, tc.value, got, tc.want)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// A correction left the superseded value in the index, so recall answered with
|
||||
// the value he had just corrected (#493).
|
||||
func TestCorrectValueDropsMemoryVectors(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
s := newTestStore(t)
|
||||
now := time.Now()
|
||||
mem := s.VectorMemory()
|
||||
|
||||
if _, err := s.WriteFact(ctx, now, KindSelf, "go_version", `"1.20"`, "tap:voice", 1.0, sql.NullInt64{}); err != nil {
|
||||
t.Fatalf("WriteFact: %v", err)
|
||||
}
|
||||
if err := mem.Insert(ctx, "fact:go_version:1", []float32{1, 0, 0}, map[string]string{
|
||||
"type": "fact", "text": "go version — 1.20",
|
||||
}); err != nil {
|
||||
t.Fatalf("Insert: %v", err)
|
||||
}
|
||||
if _, err := s.CorrectValue(ctx, "go_version", "feedback", "1.25", now.Add(time.Minute)); err != nil {
|
||||
t.Fatalf("CorrectValue: %v", err)
|
||||
}
|
||||
got, err := mem.ByPrefix(ctx, "fact:")
|
||||
if err != nil {
|
||||
t.Fatalf("ByPrefix: %v", err)
|
||||
}
|
||||
if len(got) != 0 {
|
||||
t.Fatalf("after the correction the index still holds %+v; the superseded value must not answer", got)
|
||||
}
|
||||
}
|
||||
|
||||
// The recovery path a poisoned box had none of (#470 point 4, #493): rows
|
||||
// written before the fix hold utterances, voided junk and superseded values,
|
||||
// and no write-path change reaches any of them.
|
||||
func TestRepairFactVectors(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
s := newTestStore(t)
|
||||
now := time.Now()
|
||||
mem := s.VectorMemory()
|
||||
embed := func(ctx context.Context, text string) ([]float32, error) {
|
||||
return []float32{float32(len(text)), 1, 0}, nil
|
||||
}
|
||||
|
||||
// A live fact indexed under the question that wrote it — the defect.
|
||||
if _, err := s.WriteFact(ctx, now, KindSelf, "water", `"выпил"`, "tap:voice", 1.0, sql.NullInt64{}); err != nil {
|
||||
t.Fatalf("WriteFact water: %v", err)
|
||||
}
|
||||
if err := mem.Insert(ctx, "fact:water:100", []float32{9, 9, 9}, map[string]string{
|
||||
"type": "fact", "source": "voice", "text": "запиши что я пил воду",
|
||||
}); err != nil {
|
||||
t.Fatalf("Insert water: %v", err)
|
||||
}
|
||||
// A voided fact whose vector survived the void.
|
||||
if _, err := s.WriteFact(ctx, now, KindSelf, "go_version", `"1.20"`, "tap:voice", 1.0, sql.NullInt64{}); err != nil {
|
||||
t.Fatalf("WriteFact go_version: %v", err)
|
||||
}
|
||||
if _, _, err := s.VoidLatestFact(ctx, "go_version", "feedback", now.Add(time.Minute)); err != nil {
|
||||
t.Fatalf("VoidLatestFact: %v", err)
|
||||
}
|
||||
if err := mem.Insert(ctx, "fact:go_version:100", []float32{9, 9, 9}, map[string]string{
|
||||
"type": "fact", "text": "какая последняя версия языка Go?",
|
||||
}); err != nil {
|
||||
t.Fatalf("Insert go_version: %v", err)
|
||||
}
|
||||
// A key with two vectors: only the newest may answer.
|
||||
if _, err := s.WriteFact(ctx, now, KindSelf, "mood", `"устал"`, "tap:voice", 1.0, sql.NullInt64{}); err != nil {
|
||||
t.Fatalf("WriteFact mood: %v", err)
|
||||
}
|
||||
for _, ts := range []string{"100", "200"} {
|
||||
if err := mem.Insert(ctx, "fact:mood:"+ts, []float32{9, 9, 9}, map[string]string{
|
||||
"type": "fact", "text": "мне грустно",
|
||||
}); err != nil {
|
||||
t.Fatalf("Insert mood %s: %v", ts, err)
|
||||
}
|
||||
}
|
||||
// A note must be left entirely alone.
|
||||
if err := mem.Insert(ctx, "note:7", []float32{5, 5, 5}, map[string]string{
|
||||
"type": "note", "text": "сеть тормозит по вечерам",
|
||||
}); err != nil {
|
||||
t.Fatalf("Insert note: %v", err)
|
||||
}
|
||||
|
||||
res, err := s.RepairFactVectors(ctx, embed)
|
||||
if err != nil {
|
||||
t.Fatalf("RepairFactVectors: %v", err)
|
||||
}
|
||||
if res.Rewritten != 2 || res.Dropped != 2 {
|
||||
t.Fatalf("repair reported %+v; want 2 rewritten (water, newest mood) and 2 dropped (voided go_version, superseded mood)", res)
|
||||
}
|
||||
|
||||
got, err := mem.ByPrefix(ctx, "fact:")
|
||||
if err != nil {
|
||||
t.Fatalf("ByPrefix: %v", err)
|
||||
}
|
||||
texts := map[string]string{}
|
||||
for _, r := range got {
|
||||
texts[r.ID] = r.Meta["text"]
|
||||
}
|
||||
if len(texts) != 2 {
|
||||
t.Fatalf("the index holds %+v; want only fact:water:100 and fact:mood:200", texts)
|
||||
}
|
||||
if texts["fact:water:100"] != "water — выпил" {
|
||||
t.Fatalf("water reads back %q; want the fact, not the utterance", texts["fact:water:100"])
|
||||
}
|
||||
if texts["fact:mood:200"] != "mood — устал" {
|
||||
t.Fatalf("mood reads back %q", texts["fact:mood:200"])
|
||||
}
|
||||
// Provenance the row already carried must survive the rewrite.
|
||||
for _, r := range got {
|
||||
if r.ID == "fact:water:100" && r.Meta["source"] != "voice" {
|
||||
t.Fatalf("water lost its source meta: %+v", r.Meta)
|
||||
}
|
||||
}
|
||||
if notes, err := mem.ByPrefix(ctx, "note:"); err != nil || len(notes) != 1 {
|
||||
t.Fatalf("the note row was touched: %+v (err %v)", notes, err)
|
||||
}
|
||||
|
||||
// Marker written, so a second run is free and changes nothing.
|
||||
again, err := s.RepairFactVectors(ctx, embed)
|
||||
if err != nil {
|
||||
t.Fatalf("second RepairFactVectors: %v", err)
|
||||
}
|
||||
if !again.Skipped {
|
||||
t.Fatalf("second run did work: %+v; the marker must make it a no-op", again)
|
||||
}
|
||||
}
|
||||
@@ -148,6 +148,27 @@ func (m *MemoryStore) Delete(ctx context.Context, id string) error {
|
||||
return nil
|
||||
}
|
||||
|
||||
// DeletePrefix removes every vector whose id starts with prefix and returns
|
||||
// how many went. Same escaping as ByPrefix, so a key containing % or _ cannot
|
||||
// widen the delete.
|
||||
//
|
||||
// It exists for the repair half of a revert (#470). Voiding a fact row left
|
||||
// its vector in the index, so recall kept serving the voided fact's utterance
|
||||
// and the documented repair did not repair.
|
||||
func (m *MemoryStore) DeletePrefix(ctx context.Context, prefix string) (int64, error) {
|
||||
pattern := escapeLike(prefix) + "%"
|
||||
res, err := m.db.ExecContext(ctx,
|
||||
`DELETE FROM memory_vectors WHERE id LIKE ? ESCAPE '\'`, pattern)
|
||||
if err != nil {
|
||||
return 0, fmt.Errorf("memory: delete prefix %q: %w", prefix, err)
|
||||
}
|
||||
n, err := res.RowsAffected()
|
||||
if err != nil {
|
||||
return 0, fmt.Errorf("memory: delete prefix %q: rows affected: %w", prefix, err)
|
||||
}
|
||||
return n, nil
|
||||
}
|
||||
|
||||
// escapeLike neutralises the LIKE wildcards in a literal prefix.
|
||||
func escapeLike(s string) string {
|
||||
r := strings.NewReplacer(`\`, `\\`, `%`, `\%`, `_`, `\_`)
|
||||
|
||||
@@ -0,0 +1,64 @@
|
||||
package store
|
||||
|
||||
import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// Stage 3 of #470: reverting a fact reported success and left the vector that
|
||||
// was answering questions, so the documented repair did not repair.
|
||||
func TestVoidLatestFactDropsMemoryVectors(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
s := newTestStore(t)
|
||||
now := time.Now()
|
||||
mem := s.VectorMemory()
|
||||
|
||||
if _, err := s.WriteFact(ctx, now, KindSelf, "go_version", `"1.20"`, "tap:voice", 1.0, sql.NullInt64{}); err != nil {
|
||||
t.Fatalf("WriteFact: %v", err)
|
||||
}
|
||||
// The id shape actionFact writes: fact:<key>:<unix>.
|
||||
if err := mem.Insert(ctx, "fact:go_version:1", []float32{1, 0, 0}, map[string]string{
|
||||
"type": "fact", "text": "какая последняя версия языка Go?",
|
||||
}); err != nil {
|
||||
t.Fatalf("Insert: %v", err)
|
||||
}
|
||||
// A vector for another key must survive the void.
|
||||
if err := mem.Insert(ctx, "fact:water:1", []float32{0, 1, 0}, map[string]string{
|
||||
"type": "fact", "text": "запиши что я пил воду",
|
||||
}); err != nil {
|
||||
t.Fatalf("Insert: %v", err)
|
||||
}
|
||||
|
||||
if _, _, err := s.VoidLatestFact(ctx, "go_version", "feedback", now.Add(time.Minute)); err != nil {
|
||||
t.Fatalf("VoidLatestFact: %v", err)
|
||||
}
|
||||
|
||||
got, err := mem.ByPrefix(ctx, "fact:")
|
||||
if err != nil {
|
||||
t.Fatalf("ByPrefix: %v", err)
|
||||
}
|
||||
if len(got) != 1 || got[0].ID != "fact:water:1" {
|
||||
t.Fatalf("after the void the index holds %+v; want only fact:water:1", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestDeletePrefixDoesNotWidenOnWildcards(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
s := newTestStore(t)
|
||||
mem := s.VectorMemory()
|
||||
|
||||
for _, id := range []string{"fact:a_b:1", "fact:axb:1"} {
|
||||
if err := mem.Insert(ctx, id, []float32{1, 0}, map[string]string{"type": "fact"}); err != nil {
|
||||
t.Fatalf("Insert %q: %v", id, err)
|
||||
}
|
||||
}
|
||||
n, err := mem.DeletePrefix(ctx, "fact:a_b:")
|
||||
if err != nil {
|
||||
t.Fatalf("DeletePrefix: %v", err)
|
||||
}
|
||||
if n != 1 {
|
||||
t.Fatalf("deleted %d rows; the _ in the key must not match x", n)
|
||||
}
|
||||
}
|
||||
+2
-2
@@ -18,9 +18,9 @@
|
||||
// returns a canned string the router + action path operate on); with a
|
||||
// worker socket configured, it wires Remote.
|
||||
//
|
||||
// Per DESIGN.md § Voice pipeline (STT / TTS): whisper.cpp (CGo, Vulkan) in
|
||||
// Per docs/design.md § Voice pipeline (STT / TTS): whisper.cpp (CGo, Vulkan) in
|
||||
// cmd/mavsttd is the production stt — the older faster-whisper/vosk picks are
|
||||
// retired (DESIGN.md § Superseded, "named STT/TTS model picks"). The
|
||||
// retired (docs/design.md § Superseded, "named STT/TTS model picks"). The
|
||||
// server-side stt module is the heavy multilingual path; the client's
|
||||
// wake-word + stage-0 command grammar (cmd/mavwaked) hits the router directly
|
||||
// and never crosses this seam. Today's Remote + Stub both return plain text
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
// store's allowlist, and drafts 'proposed' scaffolds for acts that aren't on
|
||||
// it yet.
|
||||
//
|
||||
// Boundary discipline (DESIGN.md § "Tool registration — drafting is suggest,
|
||||
// Boundary discipline (docs/design.md § "Tool registration — drafting is suggest,
|
||||
// enabling is act"):
|
||||
//
|
||||
// - The store is the allowlist. Only status='enabled' rows run. A verb not
|
||||
@@ -26,7 +26,7 @@
|
||||
// - Destructive tools don't run on first hearing: Exec returns ErrNeedsConfirm
|
||||
// and the handler runs a confirm turn ("выполнить X? да/нет"); only a
|
||||
// confirmed re-Exec runs them. A gate assumes a fully-formed action, which
|
||||
// an enabled+matched act is (DESIGN.md § "Confirmation is not one
|
||||
// an enabled+matched act is (docs/design.md § "Confirmation is not one
|
||||
// mechanism").
|
||||
package tool
|
||||
|
||||
|
||||
+2
-2
@@ -5,9 +5,9 @@
|
||||
// PCM, headerless per the audio package; the voice sink + reference client
|
||||
// wrap it in a WAV at the disk edge.
|
||||
//
|
||||
// Per DESIGN.md § Voice pipeline (STT / TTS): piper is the production tts
|
||||
// Per docs/design.md § Voice pipeline (STT / TTS): piper is the production tts
|
||||
// (subprocess + espeak-ng, CPU-only on the ryzen box, driven by cmd/mavttsd);
|
||||
// the older silero pick is retired (DESIGN.md § Superseded, "named STT/TTS
|
||||
// the older silero pick is retired (docs/design.md § Superseded, "named STT/TTS
|
||||
// model picks"). A different voice is a model-file swap, not a code change.
|
||||
// The daemon wires one impl — Remote pointing at the worker socket if
|
||||
// configured, Stub otherwise.
|
||||
|
||||
@@ -24,6 +24,8 @@ import (
|
||||
"net"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/netaddr"
|
||||
)
|
||||
|
||||
// Client — one connection to one worker module. NOT goroutine-safe for
|
||||
@@ -38,15 +40,25 @@ type Client struct {
|
||||
dial func() (net.Conn, error)
|
||||
}
|
||||
|
||||
// Dial opens a Client to the worker socket at path. The first call lazily
|
||||
// Dial opens a Client to the worker module at path. The first call lazily
|
||||
// dials; subsequent calls reuse the conn (a fresh dial happens on next call
|
||||
// after a teardown). Lazy dial keeps a worker that's restarting from
|
||||
// blocking core's startup; core attempts the dial on first use.
|
||||
//
|
||||
// path is a netaddr seam address. A bare path is the unix socket it has
|
||||
// always been; "tcp://workstation:9310?token=..." reaches a module on another
|
||||
// host, which is how stt and tts move to the machine with the GPU and the
|
||||
// microphone. A bad address surfaces on the first call, not here, because
|
||||
// Dial does not fail — see internal/netaddr.
|
||||
func Dial(path string) *Client {
|
||||
addr, err := netaddr.Parse(path)
|
||||
return &Client{
|
||||
path: path,
|
||||
dial: func() (net.Conn, error) {
|
||||
return net.Dial("unix", path)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return netaddr.Dial(addr)
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
+22
-34
@@ -7,10 +7,11 @@
|
||||
// which module to dial; mixing the two is a config error caught cleanly by
|
||||
// the wire, not a runtime goroutine panic). One Server per module process.
|
||||
//
|
||||
// Socket perms mirror ipc.Server: dir 0700, socket 0600 ⇒ same unix user.
|
||||
// The module has no key, so the floor is "same user"; the wg/mTLS layers
|
||||
// are out of scope here (this socket never crosses the network radius —
|
||||
// it's local-only, point-to-point between two processes on the box).
|
||||
// The seam address decides the transport. On the default unix socket the
|
||||
// perms mirror ipc.Server — dir 0700, socket 0600 ⇒ same unix user — and that
|
||||
// is the whole auth floor, because the seam never leaves the box. A tcp
|
||||
// address moves the module to another host and takes that floor away, so
|
||||
// netaddr checks a shared token before the first frame. See internal/netaddr.
|
||||
package worker
|
||||
|
||||
import (
|
||||
@@ -18,11 +19,10 @@ import (
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"net"
|
||||
"os"
|
||||
"sync"
|
||||
"sync/atomic"
|
||||
|
||||
"golang.org/x/sys/unix"
|
||||
"github.com/kami/maven/internal/netaddr"
|
||||
)
|
||||
|
||||
// Server — a worker module process's listener. Wires either a Transcriber,
|
||||
@@ -34,6 +34,7 @@ type Server struct {
|
||||
s Synthesizer
|
||||
|
||||
path string
|
||||
addr netaddr.Addr
|
||||
ln net.Listener
|
||||
|
||||
wg sync.WaitGroup
|
||||
@@ -61,25 +62,24 @@ func NewSynthesizerServer(path string, s Synthesizer) *Server {
|
||||
// two separate processes per the restart-free / fail-independent invariant).
|
||||
func (srv *Server) SetSynthesizer(s Synthesizer) { srv.s = s }
|
||||
|
||||
// Listen binds the unix socket with 0700 dir + 0600 socket perms (same floor
|
||||
// as internal/ipc). A stale socket at path is removed first so the worker
|
||||
// process restarts cleanly after a crash, no manual cleanup needed.
|
||||
// Listen binds the seam address the Server was built with.
|
||||
//
|
||||
// A bare path is a unix socket with 0700 dir + 0600 socket perms, the same
|
||||
// floor as internal/ipc, and a stale socket is removed first so the worker
|
||||
// process restarts cleanly after a crash. A "tcp://host:port?token=..."
|
||||
// address binds a network listener instead, so this module can run on the
|
||||
// workstation while core stays on homesrv; the token is mandatory there,
|
||||
// because there is no filesystem to be the auth floor. See internal/netaddr.
|
||||
func (srv *Server) Listen() error {
|
||||
_ = os.Remove(srv.path)
|
||||
if err := os.MkdirAll(parentDir(srv.path), 0o700); err != nil {
|
||||
return fmt.Errorf("worker: mkdir socket dir: %w", err)
|
||||
}
|
||||
oldMask := unix.Umask(0o077)
|
||||
ln, err := net.Listen("unix", srv.path)
|
||||
unix.Umask(oldMask)
|
||||
addr, err := netaddr.Parse(srv.path)
|
||||
if err != nil {
|
||||
return fmt.Errorf("worker: listen %s: %w", srv.path, err)
|
||||
return err
|
||||
}
|
||||
if err := os.Chmod(srv.path, 0o600); err != nil {
|
||||
_ = ln.Close()
|
||||
_ = os.Remove(srv.path)
|
||||
return fmt.Errorf("worker: chmod socket: %w", err)
|
||||
ln, err := netaddr.Listen(addr)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
srv.addr = addr
|
||||
srv.ln = ln
|
||||
return nil
|
||||
}
|
||||
@@ -193,7 +193,7 @@ func (srv *Server) Close() error {
|
||||
}
|
||||
err := srv.ln.Close()
|
||||
srv.wg.Wait()
|
||||
_ = os.Remove(srv.path)
|
||||
netaddr.Cleanup(srv.addr)
|
||||
return err
|
||||
}
|
||||
|
||||
@@ -214,15 +214,3 @@ func marshalResult(v any) json.RawMessage {
|
||||
b, _ := json.Marshal(v)
|
||||
return b
|
||||
}
|
||||
|
||||
func parentDir(p string) string {
|
||||
for i := len(p) - 1; i >= 0; i-- {
|
||||
if p[i] == '/' {
|
||||
if i == 0 {
|
||||
return "/"
|
||||
}
|
||||
return p[:i]
|
||||
}
|
||||
}
|
||||
return "."
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user