143 lines
6.7 KiB
Markdown
143 lines
6.7 KiB
Markdown
# Offloading model work to the workstation
|
|
|
|
*Last verified: 2026-08-02 @ 5c05163. Living doc: correct it in place, do not append.*
|
|
|
|
Owner's call, 2026-08-02. Vikunja #483 is the umbrella. Tasks #484 to #487 are the
|
|
work, and this file holds the shape and the rules all four must obey.
|
|
|
|
## The goal
|
|
|
|
homesrv cannot grow a GPU. The workstation has 16GB of VRAM. Move the model work
|
|
to the workstation and leave homesrv running the logic that must be always-on,
|
|
deterministic and cheap.
|
|
|
|
## Why this is tractable
|
|
|
|
The split already exists structurally. `mavsttd` and `mavttsd` are separate
|
|
daemons that core reaches over a socket, not linked libraries. Moving them off-box
|
|
is a transport change, not a redesign.
|
|
|
|
The microphone is at the workstation, because that is where the owner sits and
|
|
homesrv is headless. So speech-to-text and the wake word are already on the
|
|
workstation side by construction. Audio never has to cross the LAN. Only the core
|
|
turn does.
|
|
|
|
## The constraint that shapes everything
|
|
|
|
The workstation's GPU is often busy: CPT runs, experiments, Correx, the manga-recap
|
|
pipeline. It also sleeps. homesrv does not.
|
|
|
|
So an offloaded model is never *the* model. It is the preferred one, with a floor
|
|
on homesrv. That is the shape the cascade already has, where a router error falls
|
|
through to the classifier.
|
|
|
|
## The degradation rule
|
|
|
|
Two cases, and the line between them is sharp.
|
|
|
|
**Fall back silently** when the workstation model would only do the job *better*:
|
|
routing, phrasing, a nudge. Falling back costs nothing that exists today, because
|
|
the resident Qwen3-1.7B is today's production quality. The owner should not be told
|
|
that his reply was phrased by the smaller model.
|
|
|
|
**Name the gap** when the resident model cannot do the job *at all*. A world
|
|
question that a 1.7B answers by inventing is the case. A wrong answer is worse
|
|
than "не могу сейчас". This is the rule CLAUDE.md already states for a sibling
|
|
service being down.
|
|
|
|
Nothing in between. A turn never breaks on the workstation being asleep.
|
|
|
|
## Admission control, not a scheduler
|
|
|
|
There is no GPU arbiter. That is a service with its own failure modes, and nothing
|
|
here needs work *distributed*. It needs admission control. The workstation
|
|
advertises free VRAM over a health endpoint, and Maven treats it as one more query
|
|
source that claims a turn or passes. llama-server also refuses to load when VRAM is
|
|
short, so the failure is detectable without cooperation from the owner's other
|
|
jobs.
|
|
|
|
The caller must be able to ask "is this peer usable right now" without a turn
|
|
hanging on a timeout. A dead remote is a normal state, not an error state.
|
|
`internal/llm.Pair` is that check on the Maven side. A prober caches the answer,
|
|
so `Available()` is an atomic read and no turn pays for a health check.
|
|
|
|
llama-server does not stay up on the workstation. It cannot: a resident 7-14B
|
|
would hold 16GB against the owner's CPT runs. So a supervisor there owns its
|
|
lifecycle, keeps it loaded while the card is free, and unloads it on idle or
|
|
when another process needs the card (owner's call, 2026-08-02, Vikunja #488).
|
|
|
|
That supervisor is still not a scheduler, and the distinction is worth holding.
|
|
It arbitrates nothing between callers. It reports whether it can take work and
|
|
manages one process to back that answer. Maven never asks it to start anything
|
|
and never learns that it did.
|
|
|
|
Contention is decided by presence under `/sys/class/kfd/kfd/proc`, not by a VRAM
|
|
threshold. A ROCm process registers there when it initialises HIP, before it
|
|
allocates anything. So the supervisor sees a contender during that job's startup,
|
|
and yields before the job loses the memory it asked for. A
|
|
threshold reads the card too late. By the time free VRAM has dropped, the other
|
|
job has already lost the allocation race. Free VRAM is still read, but only as a
|
|
precondition for loading, never as the eviction signal. One blind spot is known.
|
|
A job can take the card without registering on the KFD, as a Vulkan or a
|
|
video-decode job would. `describe()` logs every contender's comm, and that log is
|
|
how we find out whether the blind spot is real.
|
|
|
|
`mavgpud` runs from a systemd unit on the workstation with
|
|
`deploy/mavgpud.json` as its config, and `llama_args` is passed to llama-server
|
|
untouched. The model, the context size, the layer count and the MTP flags are the
|
|
owner's business and not this daemon's schema.
|
|
|
|
## What stays on homesrv, permanently
|
|
|
|
The **embedder** (multilingual-e5-small, ONNX, CPU). It backs the classifier, which
|
|
must answer while the GPU is saturated. It is also cheap enough on CPU that moving
|
|
it buys nothing. Four callers:
|
|
|
|
| Caller | What for |
|
|
|---|---|
|
|
| `internal/router/classifier.go` | the routing floor |
|
|
| `cmd/mavend/actions_query.go` (`queryEmbed`) | memory recall |
|
|
| `cmd/mavend/feeds.go` | ingest embedding for every RSS item |
|
|
| `internal/crawl/watch.go` | ingest embedding for every crawled page |
|
|
|
|
`internal/speaker` becomes a fifth once it lands.
|
|
|
|
## Inventory: what runs a model on homesrv today
|
|
|
|
The **resident model** is one llama-server with seven callers:
|
|
|
|
| Caller | What for |
|
|
|---|---|
|
|
| `cmd/mavend/voicewire.go` | routing |
|
|
| `cmd/mavend/replier_llm.go` | replies |
|
|
| `cmd/mavend/tick.go` | digestion worker: `PhraseNudge`, `PhraseReminder` |
|
|
| `cmd/mavend/capture.go` | capture summarisation (unreachable, see #480) |
|
|
| `cmd/mavend/mail.go` | mail extraction (off, no IMAP) |
|
|
| `cmd/mavend/kiwixwire.go` | answering from a Kiwix, search or crawl passage |
|
|
| `memoryeval.go`, `modelswap.go` | admin and evals |
|
|
|
|
Then the embedder above, **whisper.cpp** in `mavsttd`, and **piper** in `mavttsd`.
|
|
`mavwaked` uses no model at all: an energy-threshold VAD over 30ms frames.
|
|
|
|
## Order
|
|
|
|
1. **Transport** (#484). Nothing else is possible until a seam can cross a host.
|
|
`internal/netaddr` landed in PR #92. A seam address now carries its own scheme,
|
|
and a scheme-less one is still unix. A tcp seam requires a shared token, because
|
|
the filesystem permission that authenticated the unix socket is gone.
|
|
2. **The resident model** (#485). Biggest quality delta. A 16GB card runs a 7-14B,
|
|
which fixes what the 1.7B gets wrong: world knowledge, and the persona the CPT
|
|
targets. The degradation path is already written and measured, since the
|
|
classifier scores 68.8% full accuracy at p50 16.6µs on its own.
|
|
3. **Speech-to-text and text-to-speech** (#486). They gain a real margin, but on
|
|
quality alone, and both already work.
|
|
4. **The wake word** (#487). Independent of all of the above.
|
|
|
|
## Assumptions
|
|
|
|
- The LAN is trusted enough that wireguard is supported but not required (owner's
|
|
call). What crosses the wire is still his utterances. That is why the tcp seam
|
|
carries its own token instead of assuming a network boundary.
|
|
- The workstation is not expected to be up. Every child task must still serve a
|
|
turn while it is down.
|