From dc266056d193730b239024f695b7a1adbc80bd22 Mon Sep 17 00:00:00 2001 From: claude Date: Sun, 2 Aug 2026 16:56:20 +0400 Subject: [PATCH] docs: the shape and the rules for offloading model work (V-483) 483 is an umbrella and its children are the work, so what it owes them is the shape they must all obey. docs/offload.md records it: the degradation rule and where its line falls, admission control rather than a GPU arbiter, the embedder staying on homesrv because it backs the classifier, and the inventory of what runs a model on the box today. CLAUDE.md gets a pointer, because an agent about to add a model caller or touch a daemon seam needs to know this before it starts, not after. --- CLAUDE.md | 8 ++++ docs/offload.md | 114 ++++++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 122 insertions(+) create mode 100644 docs/offload.md diff --git a/CLAUDE.md b/CLAUDE.md index fb22297..51952b5 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -28,6 +28,14 @@ model is a one-line change to `phraser.model_path` in `deploy/mavend.json`. See `docs/rearchitecture.md` for the target architecture, `docs/design.md` for the folded design spec, and `AGENTS.md` for local-preview + model-download recipes. +**Model work is moving to the workstation** (owner's call, 2026-08-02). homesrv cannot grow a +GPU and the workstation has 16GB of VRAM. So the resident model, STT and TTS become preferred +remotes with a floor on homesrv. The workstation is never assumed up. Fall back silently when +it would only do the job better. Name the gap when the 1.7B cannot do it at all. The embedder +stays on homesrv permanently, because it backs that floor. Read `docs/offload.md` before +touching a daemon seam or adding a model caller. Vikunja #483 is the umbrella, #484 to #487 +are the work. + ## Build & test CGO daemons (`mavend`, `mavsttd`, `mavttsd`, `mavenclient`) need the vendored toolchain diff --git a/docs/offload.md b/docs/offload.md new file mode 100644 index 0000000..e6bf7c4 --- /dev/null +++ b/docs/offload.md @@ -0,0 +1,114 @@ +# Offloading model work to the workstation + +*Last verified: 2026-08-02 @ 5c05163. Living doc: correct it in place, do not append.* + +Owner's call, 2026-08-02. Vikunja #483 is the umbrella. Tasks #484 to #487 are the +work, and this file holds the shape and the rules all four must obey. + +## The goal + +homesrv cannot grow a GPU. The workstation has 16GB of VRAM. Move the model work +to the workstation and leave homesrv running the logic that must be always-on, +deterministic and cheap. + +## Why this is tractable + +The split already exists structurally. `mavsttd` and `mavttsd` are separate +daemons that core reaches over a socket, not linked libraries. Moving them off-box +is a transport change, not a redesign. + +The microphone is at the workstation, because that is where the owner sits and +homesrv is headless. So speech-to-text and the wake word are already on the +workstation side by construction. Audio never has to cross the LAN. Only the core +turn does. + +## The constraint that shapes everything + +The workstation's GPU is often busy: CPT runs, experiments, Correx, the manga-recap +pipeline. It also sleeps. homesrv does not. + +So an offloaded model is never *the* model. It is the preferred one, with a floor +on homesrv. That is the shape the cascade already has, where a router error falls +through to the classifier. + +## The degradation rule + +Two cases, and the line between them is sharp. + +**Fall back silently** when the workstation model would only do the job *better*: +routing, phrasing, a nudge. Falling back costs nothing that exists today, because +the resident Qwen3-1.7B is today's production quality. The owner should not be told +that his reply was phrased by the smaller model. + +**Name the gap** when the resident model cannot do the job *at all*. A world +question that a 1.7B answers by inventing is the case. A wrong answer is worse +than "не могу сейчас". This is the rule CLAUDE.md already states for a sibling +service being down. + +Nothing in between. A turn never breaks on the workstation being asleep. + +## Admission control, not a scheduler + +There is no GPU arbiter. That is a service with its own failure modes, and nothing +here needs work *distributed*. It needs admission control. The workstation +advertises free VRAM over a health endpoint, and Maven treats it as one more query +source that claims a turn or passes. llama-server also refuses to load when VRAM is +short, so the failure is detectable without cooperation from the owner's other +jobs. + +The caller must be able to ask "is this peer usable right now" without a turn +hanging on a timeout. A dead remote is a normal state, not an error state. + +## What stays on homesrv, permanently + +The **embedder** (multilingual-e5-small, ONNX, CPU). It backs the classifier, which +must answer while the GPU is saturated. It is also cheap enough on CPU that moving +it buys nothing. Four callers: + +| Caller | What for | +|---|---| +| `internal/router/classifier.go` | the routing floor | +| `cmd/mavend/actions_query.go` (`queryEmbed`) | memory recall | +| `cmd/mavend/feeds.go` | ingest embedding for every RSS item | +| `internal/crawl/watch.go` | ingest embedding for every crawled page | + +`internal/speaker` becomes a fifth once it lands. + +## Inventory: what runs a model on homesrv today + +The **resident model** is one llama-server with seven callers: + +| Caller | What for | +|---|---| +| `cmd/mavend/voicewire.go` | routing | +| `cmd/mavend/replier_llm.go` | replies | +| `cmd/mavend/tick.go` | digestion worker: `PhraseNudge`, `PhraseReminder` | +| `cmd/mavend/capture.go` | capture summarisation (unreachable, see #480) | +| `cmd/mavend/mail.go` | mail extraction (off, no IMAP) | +| `cmd/mavend/kiwixwire.go` | answering from a Kiwix, search or crawl passage | +| `memoryeval.go`, `modelswap.go` | admin and evals | + +Then the embedder above, **whisper.cpp** in `mavsttd`, and **piper** in `mavttsd`. +`mavwaked` uses no model at all: an energy-threshold VAD over 30ms frames. + +## Order + +1. **Transport** (#484). Nothing else is possible until a seam can cross a host. + `internal/netaddr` landed in PR #92. A seam address now carries its own scheme, + and a scheme-less one is still unix. A tcp seam requires a shared token, because + the filesystem permission that authenticated the unix socket is gone. +2. **The resident model** (#485). Biggest quality delta. A 16GB card runs a 7-14B, + which fixes what the 1.7B gets wrong: world knowledge, and the persona the CPT + targets. The degradation path is already written and measured, since the + classifier scores 68.8% full accuracy at p50 16.6µs on its own. +3. **Speech-to-text and text-to-speech** (#486). They gain a real margin, but on + quality alone, and both already work. +4. **The wake word** (#487). Independent of all of the above. + +## Assumptions + +- The LAN is trusted enough that wireguard is supported but not required (owner's + call). What crosses the wire is still his utterances. That is why the tcp seam + carries its own token instead of assuming a network boundary. +- The workstation is not expected to be up. Every child task must still serve a + turn while it is down. -- 2.52.0