Swap the resident model without restarting mavend (#250)

Loading a different gguf was a one-line edit to phraser.model_path plus a
restart. It is now an owner-triggered IPC call, off unless configured.

internal/phraser/swap.go holds the safety properties as code:

  - Never two models resident. The old llama-server is killed and reaped
    before the new one is launched. One 1.7B fits the Vega iGPU; a
    blue/green overlap would OOM the box, so it is not offered.
  - Atomic from a turn's point of view. Swap drains the in-flight turns
    (they finish on the old model), then refuses arrivals with ErrSwapping
    until the new server has answered /v1/models. No turn ever sees half a
    swap; refused turns fall back to the classifier cascade.
  - A failed load rolls back. If the new model does not start or does not
    probe, the previous one is reloaded and the call returns RolledBack
    with the error. If the rollback also fails the daemon says so and
    degrades to the classifier rather than pretending to serve.

Holders of the completion client are re-pointed, not rebuilt: llm.Client
guards its base URL and LLMPhraser.OnSwap re-points it, so the router, the
replier, the mail extractor and the memory evaluator follow the new port
without knowing a swap happened.

Reach is deliberately narrow. phraser.swap_models is an exact-match
allowlist of absolute paths a human wrote, rejected at startup otherwise,
so "swap the model" can never mean "load any file on my disk"; the running
model is always swappable back to. MethodSwapModel is AuthStepUp, the same
rung as mutating the tool allowlist, and /models gates POST through the
same stepUpOK the tools page uses. Nothing calls Swap on a timer and no
act, intent or utterance reaches it.

Vikunja #250
This commit is contained in:
kami
2026-08-01 03:59:08 +04:00
parent 2c1b0eede0
commit ad074cea31
22 changed files with 1523 additions and 43 deletions
+42
View File
@@ -6,6 +6,7 @@ import (
"encoding/binary"
"encoding/json"
"errors"
"fmt"
"io"
"net"
"os"
@@ -598,3 +599,44 @@ func TestIngestMail_Hook(t *testing.T) {
t.Errorf("req across the wire = %+v", got)
}
}
// TestSwapModel_OffUnlessConfigured — no allowlist in the config means the
// daemon never sets the hook, and the method does not exist. That is what "off
// unless configured" looks like at the wire for the model swap (Vikunja #250).
func TestSwapModel_OffUnlessConfigured(t *testing.T) {
_, _, cli, _ := newServerWithStore(t)
if _, err := cli.SwapModel(context.Background(), SwapModelReq{ModelPath: "/m/x.gguf"}); !errors.Is(err, ErrUnknownMethod) {
t.Fatalf("SwapModel error = %v, want ErrUnknownMethod", err)
}
if _, err := cli.ModelStatus(context.Background()); !errors.Is(err, ErrUnknownMethod) {
t.Fatalf("ModelStatus error = %v, want ErrUnknownMethod", err)
}
}
// TestSwapModel_Hook — the request crosses the boundary intact and the reported
// identity comes back. A refusal from the daemon's allowlist arrives as
// ErrForbidden, which is what a caller keys its error message off.
func TestSwapModel_Hook(t *testing.T) {
_, srv, cli, _ := newServerWithStore(t)
var got SwapModelReq
srv.SwapModelFn = func(_ context.Context, req SwapModelReq) (SwapModelResp, error) {
got = req
if req.ModelPath != "/m/allowed.gguf" {
return SwapModelResp{}, fmt.Errorf("%w: not allowlisted", ErrForbidden)
}
return SwapModelResp{Model: "allowed", ModelPath: req.ModelPath, BaseURL: "http://127.0.0.1:9", TookMs: 12}, nil
}
resp, err := cli.SwapModel(context.Background(), SwapModelReq{ModelPath: "/m/allowed.gguf", NCtx: 4096})
if err != nil {
t.Fatalf("SwapModel: %v", err)
}
if resp.Model != "allowed" || resp.TookMs != 12 {
t.Errorf("resp = %+v", resp)
}
if got.NCtx != 4096 {
t.Errorf("req across the wire = %+v", got)
}
if _, err := cli.SwapModel(context.Background(), SwapModelReq{ModelPath: "/etc/shadow"}); !errors.Is(err, ErrForbidden) {
t.Fatalf("swap to a non-allowlisted path = %v; want ErrForbidden", err)
}
}