Files
Maven/internal/phraser/phraser.go
T
kami 76a6a007ef Pin the resident model to Qwen3.5-0.8B and name Qwen3-1.7B as the target
The most load-bearing decision in the project was stated four incompatible
ways: the docs said Qwen3-1.7B, deploy/mavend.json said Qwen3.5-2B, the repo's
models/llm/ held an LFM2.5-1.2B gguf, and five code comments still said LFM.
Answering "which model is deployed" meant re-deriving it from scratch every
time.

Two facts the review missed, found while resolving it:

- /mnt/hdd1/llms is bind-mounted over /opt/maven/models/llm, which shadows the
  repo's models/llm/. The LFM2.5 gguf sitting there was never loaded by
  anything, so it was not evidence of the deployed model at all.
- That library holds Qwen3.5-0.8B, -2B and -4B, and no Qwen3-1.7B. The config
  pointed at a file that does exist; the docs' Qwen3-1.7B was the stale claim,
  the reverse of the assumed direction. Qwen3-1.7B is the CPT target, and that
  training is still in flight (Vikunja #122), so no such gguf exists yet.

phraser.model_path moves to Qwen3.5-0.8B (Q4_K_M) — the smallest checkpoint on
disk, chosen for latency, and relevant to whether the LLM router is affordable
on this box. Docs and comments now say the same thing in one voice: 0.8B
resident now, CPT'd Qwen3-1.7B as the target, and the bind-mount shadowing
written down so the next reader does not mistake models/llm/ for ground truth.
Comments name the model, never a filename, so a swap stays a one-line config
change.

n_gpu_layers: 99 is correct and stays — compose passes /dev/dri and the render
gid for Vulkan offload to the Vega iGPU. CLAUDE.md's "CPU-only" was the stale
half of that contradiction and is corrected.

phraser.go also dropped a wrong "sub-1b, prompted not trained" size claim: the
target is trained end-to-end (RU CPT + joint persona/router SFT).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X5JApcrCRVGmqrxnhynSik
2026-07-30 23:40:33 +04:00

219 lines
8.5 KiB
Go

// Package phraser is maven's "rules decide, llm phrases" seam — the layer
// that turns a loop decision into the body + summary the delivery module ships.
//
// Per DESIGN.md § Resident language model: the phraser is the resident model
// (Qwen3-1.7B — RU continued pretraining plus joint persona/router SFT, not a
// sub-1b prompted-only model as the retired spec claimed; see DESIGN.md
// § Superseded, "small-model phrasing claim"). It takes
// (rule, severity, context) and produces Body (full voice message, local — no
// shoulder-surf concern beyond who's in the room) + Summary (minimal body for
// away channels — "disk low on homesrv," not detail; no exfil through the
// relay). the phraser NEVER owns the route — it phrases what the loop decided.
//
// This package defines the interface + a deterministic Stub (the floor). the
// Stub is template-based, no model — it exists so the daemon can be wired
// end-to-end before the LLM-backed impl lands. the LLM impl is a single new
// type satisfying the same interface; the daemon swaps one for the other at
// the construction seam, no CoreAPI or delivery change.
//
// Architecture: the phraser is impure (the LLM impl makes RPC calls). the
// Stub is pure (templates over State) and is the test floor. both produce
// delivery.PhrasedNudge / delivery.PhrasedReminder, which the dispatcher
// consumes unchanged.
package phraser
import (
"context"
"encoding/json"
"fmt"
"strings"
"time"
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/loop"
)
// Phraser — the seam the daemon wires. one method per delivery path (nudge
// = loop-derived, reminder = user-stated). both return the
// delivery.Phrased* structs the dispatcher consumes, so the phraser owns the
// full output contract: Body (voice) + Summary (away channels).
//
// the daemon calls PhraseNudge with the loop's *Candidate (Rule + Severity +
// the State snapshot at evaluation time — exactly the (rule, severity,
// context) input the spec names). PhraseReminder with the ReminderDecision
// (Reminder + State). PhraseChat with a conversational utterance + dialogue
// history. the phraser reads the State for context ("you haven't had water in
// 4h, you're at your desk, it's 2pm") — never touches the store.
type Phraser interface {
PhraseNudge(ctx context.Context, c loop.Candidate) (delivery.PhrasedNudge, error)
PhraseReminder(ctx context.Context, d loop.ReminderDecision) (delivery.PhrasedReminder, error)
PhraseQuery(ctx context.Context, utterance string, notes []string) (string, error)
PhraseChat(ctx context.Context, utterance string, history []dialogue.Turn) (string, error)
Close() error
}
// Stub — the deterministic, no-model floor. template-based, reads context
// from the Candidate/Decision State. produces a terse Body (voice) + an even
// terser Summary (away channels). warm-but-functional tone; the LLM impl
// carries the personality-prompt character spec, the Stub does not.
//
// the Stub is the production path until the LLM-backed impl lands, and the
// test path afterward (deterministic phrasing makes the daemon + delivery
// unit-testable without a model in the loop).
type Stub struct{}
// NewStub builds the floor phraser. no config — the Stub is stateless.
func NewStub() *Stub { return &Stub{} }
// PhraseChat returns a stub reply — the LLMPhraser replaces this with a
// prompted response from the model. The history parameter is accepted but
// ignored at the stub level (the production impl uses it for multi-turn).
func (s *Stub) PhraseChat(_ context.Context, _ string, _ []dialogue.Turn) (string, error) {
return "поговорили.", nil
}
// PhraseQuery returns a deterministic summary of the best matching notes.
func (s *Stub) PhraseQuery(_ context.Context, _ string, notes []string) (string, error) {
if len(notes) == 0 {
return "не знаю.", nil
}
if len(notes) == 1 {
return "вот что я нашла: " + notes[0], nil
}
return "вот что я нашла: " + strings.Join(notes, "; "), nil
}
// Close implements Phraser.Close (no-op for the stub).
func (s *Stub) Close() error { return nil }
// PhraseNudge — dispatches on the rule name to a per-rule template, falls
// back to a generic shape. reads the State for the durations/values that made
// the predicate fire (the same State the predicate saw).
func (s *Stub) PhraseNudge(_ context.Context, c loop.Candidate) (delivery.PhrasedNudge, error) {
body, summary := phraseNudge(c)
return delivery.PhrasedNudge{Candidate: c, Body: body, Summary: summary}, nil
}
// PhraseReminder — extracts the user's text from the reminder payload (raw
// JSON, shape owned by the router's reminder slot extraction) and renders it
// as both Body and a short Summary. the reminder's payload is the user's own
// words — the phraser just unwraps it, doesn't editorialize.
func (s *Stub) PhraseReminder(_ context.Context, d loop.ReminderDecision) (delivery.PhrasedReminder, error) {
text := extractReminderText(d.Reminder.Payload)
if text == "" {
text = "reminder"
}
summary := text
if len(summary) > 60 {
summary = summary[:57] + "..."
}
return delivery.PhrasedReminder{Decision: d, Body: text, Summary: summary}, nil
}
// phraseNudge — the per-rule templates. each reads the context the predicate
// used to decide, so the phrased message names WHY the rule fired ("you
// haven't had water in 4h") rather than just THAT it fired.
func phraseNudge(c loop.Candidate) (body, summary string) {
switch c.Rule.Name {
case "water":
if d, ok := c.State.Since("water"); ok {
body = fmt.Sprintf("you haven't had water in %s — drink something.", humanDur(d))
} else {
body = "drink some water."
}
return body, "drink water"
case "meal":
if d, ok := c.State.Since("meal"); ok {
body = fmt.Sprintf("it's been %s since you ate — get some food.", humanDur(d))
} else {
body = "you should eat something."
}
return body, "eat something"
case "break":
if d, ok := c.State.Since("break"); ok {
body = fmt.Sprintf("you've been at your desk for %s without a break — step away for a bit.", humanDur(d))
} else {
body = "take a break."
}
return body, "take a break"
case "service_down":
// the fact value is json `"down"`; the key carries the service name.
body = "a service on homesrv is down — check journalctl."
summary = "service down on homesrv"
if f, ok := c.State.Fact("service_down"); ok {
if f.Key != "" && f.Key != "service_down" {
body = fmt.Sprintf("%s on homesrv is down — check journalctl.", f.Key)
summary = fmt.Sprintf("%s down on homesrv", f.Key)
}
}
return body, summary
default:
// generic: name the rule + severity; the LLM impl replaces this with
// a prompted phrase. the Stub never editorializes beyond the rule name.
body = fmt.Sprintf("%s — %s", c.Rule.Name, sevLabel(c.Severity))
summary = c.Rule.Name
return body, summary
}
}
// extractReminderText — the reminder payload is raw JSON; the router's
// reminder slot extraction owns the shape. the conventional field is "text".
// fall back to the raw payload if it isn't JSON or lacks the field — the user
// said it, it's the user's words.
func extractReminderText(payload string) string {
var m map[string]any
if err := json.Unmarshal([]byte(payload), &m); err == nil {
if t, ok := m["text"].(string); ok && t != "" {
return t
}
if t, ok := m["text"]; ok {
return fmt.Sprintf("%v", t)
}
}
return strings.TrimSpace(payload)
}
// humanDur — round a duration to the coarsest sensible unit for speech.
// "4h12m" → "4 hours"; "92m" → "1h32m" → "an hour and a half". keep it simple:
// hours, then minutes, rounded. this is FLOOR phrasing — the LLM impl can
// natural-language it; the Stub sticks to readable.
func humanDur(d time.Duration) string {
if d < 0 {
d = 0
}
h := int(d.Hours())
m := int(d.Minutes()) % 60
switch {
case h >= 2:
return fmt.Sprintf("%d hours", h)
case h == 1:
if m >= 30 {
return "an hour and a half"
}
return "an hour"
default:
if m >= 45 {
return "an hour"
}
return fmt.Sprintf("%d minutes", m)
}
}
// sevLabel — a one-word gist of severity for the generic fallback. the
// per-rule templates don't use this; it's only for rules without a dedicated
// template (i.e. rules added to DefaultRules after the phrasers ship, before
// they get a template).
func sevLabel(s loop.Severity) string {
switch {
case s <= loop.Sev1:
return "care"
case s == loop.Sev2:
return "care"
case s == loop.Sev3:
return "ops"
default:
return "alarm"
}
}