docs: both halves of the degradation rule are wired, and which caller is which (V-490)

The offload inventory grows a column, because "seven callers of the resident
model" stopped being the useful fact. Which of them is offloaded, and under
which half of the rule, is. Three are resident-only on purpose and the table
now says why rather than leaving it to be rediscovered.

The three-outcome table is the part that was not obvious from the rule as
written. A configured-and-asleep workstation names the gap; a box with no
workstation block does not, because naming a gap requires a gap.
This commit is contained in:
2026-08-03 12:29:12 +04:00
parent 12530c8a95
commit 9b124d9194
2 changed files with 52 additions and 16 deletions
+7
View File
@@ -36,6 +36,13 @@ stays on homesrv permanently, because it backs that floor. Read `docs/offload.md
touching a daemon seam or adding a model caller. Vikunja #483 is the umbrella, #484 to #487 touching a daemon seam or adding a model caller. Vikunja #483 is the umbrella, #484 to #487
are the work. are the work.
Both halves are wired as of 2026-08-03. Routing and replies prefer the workstation silently
through `modelSeam`; nudge and reminder phrasing prefer it silently inside the phraser. A
world question goes through `LLMPhraser.PhraseWorld` and names the gap when the card is not
free — `worldGap` in `cmd/mavend/worldmodel.go`, which he hears instead of an invented
answer. A box with no `workstation` block behaves exactly as it did before the seam: naming
a gap requires a gap. The offload table in `docs/offload.md` says which caller is which.
## Build & test ## Build & test
CGO daemons (`mavend`, `mavsttd`, `mavttsd`, `mavenclient`) need the vendored toolchain CGO daemons (`mavend`, `mavsttd`, `mavttsd`, `mavenclient`) need the vendored toolchain
+45 -16
View File
@@ -1,6 +1,6 @@
# Offloading model work to the workstation # Offloading model work to the workstation
*Last verified: 2026-08-02 @ 5c05163. Living doc: correct it in place, do not append.* *Last verified: 2026-08-03 @ 12530c8. Living doc: correct it in place, do not append.*
Owner's call, 2026-08-02. Vikunja #483 is the umbrella. Tasks #484 to #487 are the Owner's call, 2026-08-02. Vikunja #483 is the umbrella. Tasks #484 to #487 are the
work, and this file holds the shape and the rules all four must obey. work, and this file holds the shape and the rules all four must obey.
@@ -47,6 +47,24 @@ service being down.
Nothing in between. A turn never breaks on the workstation being asleep. Nothing in between. A turn never breaks on the workstation being asleep.
Both halves are wired, 03-08-2026. `LLMPhraser.PhraseWorld`
(`internal/phraser/world.go`) is the naming half and has three outcomes, not two:
| State | What he hears |
|---|---|
| no `workstation` block | the resident model answers, exactly as before the seam existed |
| configured, card free | the workstation answers |
| configured, asleep or busy | the gap, `worldGap` in `cmd/mavend/worldmodel.go` |
The first row is the one worth stating. Naming a gap requires a gap. On a box with
no second model the 1.7B is the whole product. Refusing every world question there
would remove a capability the owner has today.
A source holding a passage is on the naming half too: a live search, a ZIM
article, a page he named. None of them says "не могу сейчас". They read the
passage back, which is what `phraseSource` returning `""` selects. A real quote
beats a gap, and neither path invents.
## Admission control, not a scheduler ## Admission control, not a scheduler
There is no GPU arbiter. That is a service with its own failure modes, and nothing There is no GPU arbiter. That is a service with its own failure modes, and nothing
@@ -104,17 +122,29 @@ it buys nothing. Four callers:
## Inventory: what runs a model on homesrv today ## Inventory: what runs a model on homesrv today
The **resident model** is one llama-server with seven callers: The **resident model** is one llama-server with seven callers, and 03-08-2026 is
the date each of them stopped or did not stop being resident-only:
| Caller | What for | | Caller | What for | Offloaded |
|---|---| |---|---|---|
| `cmd/mavend/voicewire.go` | routing | | `cmd/mavend/voicewire.go` | routing | silently, through `hot` |
| `cmd/mavend/replier_llm.go` | replies | | `cmd/mavend/replier_llm.go` | replies | silently, through `hot` |
| `cmd/mavend/tick.go` | digestion worker: `PhraseNudge`, `PhraseReminder` | | `cmd/mavend/tick.go` | digestion worker: `PhraseNudge`, `PhraseReminder` | silently, inside the phraser |
| `cmd/mavend/capture.go` | capture summarisation (unreachable, see #480) | | `cmd/mavend/actions_query.go` | world questions, and any fetched passage | names the gap |
| `cmd/mavend/mail.go` | mail extraction (off, no IMAP) | | `cmd/mavend/capture.go` | capture summarisation (unreachable, see #480) | no, holds its own client |
| `cmd/mavend/kiwixwire.go` | answering from a Kiwix, search or crawl passage | | `cmd/mavend/mail.go` | mail extraction (off, no IMAP) | no, holds its own client |
| `memoryeval.go`, `modelswap.go` | admin and evals | | `memoryeval.go`, `modelswap.go` | admin and evals | no, and deliberately |
The last three rows are resident-only on purpose. `memoryeval.go` and
`modelswap.go` measure and swap the resident model, so sending their work
elsewhere would measure the wrong thing. `capture.go` and `mail.go` are
background jobs that hold a gated background client (`llmBackgroundClientFor`),
and that priority has no equivalent on the remote yet. Both are also unreachable
on this deploy, so wiring them would ship an untestable path.
The `tick.go` row needs one caveat. `phraser.llm_nudges` is `false` in deploy, so
nudges come from templates and the seam under them changes nothing until that
flips. It is wired anyway: `PhraseReminder` is on the same transport and is on.
Then the embedder above, **whisper.cpp** in `mavsttd`, and **piper** in `mavttsd`. Then the embedder above, **whisper.cpp** in `mavsttd`, and **piper** in `mavttsd`.
`mavwaked` uses no model at all: an energy-threshold VAD over 30ms frames. `mavwaked` uses no model at all: an energy-threshold VAD over 30ms frames.
@@ -125,11 +155,10 @@ Then the embedder above, **whisper.cpp** in `mavsttd`, and **piper** in `mavttsd
`internal/netaddr` landed in PR #92. A seam address now carries its own scheme, `internal/netaddr` landed in PR #92. A seam address now carries its own scheme,
and a scheme-less one is still unix. A tcp seam requires a shared token, because and a scheme-less one is still unix. A tcp seam requires a shared token, because
the filesystem permission that authenticated the unix socket is gone. the filesystem permission that authenticated the unix socket is gone.
2. **The resident model** (#485). Half wired, 02-08-2026. A `workstation` block 2. **The resident model** (#485, #490). Wired. A `workstation` block builds an
builds an `llm.Pair` in `modelSeam` (`cmd/mavend/voicewire.go`), and routing `llm.Pair` in `modelSeam` (`cmd/mavend/voicewire.go`), routing and replies
and replies complete through it. Both are the silent half of the rule. The complete through it, and the phraser holds the same pair (`UseRemote`). Both
naming half is not wired. A world question still goes to the resident model halves of the rule are live: see the table above for which caller gets which.
through `PhraseQuery`. That, and the four callers 485 did not reach, are #490.
Measured, `docs/evals/2026-08-02-workstation-gemma4-12b.md`: gemma-4-12b Measured, `docs/evals/2026-08-02-workstation-gemma4-12b.md`: gemma-4-12b
through the cascade scores 84.4% full accuracy at p50 329ms. The resident through the cascade scores 84.4% full accuracy at p50 329ms. The resident
model scores 72.7% at p50 0.80-1.04s. On the talk fixture it is 25/27 model scores 72.7% at p50 0.80-1.04s. On the talk fixture it is 25/27