Files
Maven/internal/router/router.go
T
kami 76a6a007ef Pin the resident model to Qwen3.5-0.8B and name Qwen3-1.7B as the target
The most load-bearing decision in the project was stated four incompatible
ways: the docs said Qwen3-1.7B, deploy/mavend.json said Qwen3.5-2B, the repo's
models/llm/ held an LFM2.5-1.2B gguf, and five code comments still said LFM.
Answering "which model is deployed" meant re-deriving it from scratch every
time.

Two facts the review missed, found while resolving it:

- /mnt/hdd1/llms is bind-mounted over /opt/maven/models/llm, which shadows the
  repo's models/llm/. The LFM2.5 gguf sitting there was never loaded by
  anything, so it was not evidence of the deployed model at all.
- That library holds Qwen3.5-0.8B, -2B and -4B, and no Qwen3-1.7B. The config
  pointed at a file that does exist; the docs' Qwen3-1.7B was the stale claim,
  the reverse of the assumed direction. Qwen3-1.7B is the CPT target, and that
  training is still in flight (Vikunja #122), so no such gguf exists yet.

phraser.model_path moves to Qwen3.5-0.8B (Q4_K_M) — the smallest checkpoint on
disk, chosen for latency, and relevant to whether the LLM router is affordable
on this box. Docs and comments now say the same thing in one voice: 0.8B
resident now, CPT'd Qwen3-1.7B as the target, and the bind-mount shadowing
written down so the next reader does not mistake models/llm/ for ground truth.
Comments name the model, never a filename, so a swap stays a one-line config
change.

n_gpu_layers: 99 is correct and stays — compose passes /dev/dri and the render
gid for Vulkan offload to the Vega iGPU. CLAUDE.md's "CPU-only" was the stale
half of that contradiction and is corrected.

phraser.go also dropped a wrong "sub-1b, prompted not trained" size claim: the
target is trained end-to-end (RU CPT + joint persona/router SFT).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X5JApcrCRVGmqrxnhynSik
2026-07-30 23:40:33 +04:00

128 lines
4.5 KiB
Go

package router
import (
"context"
"log"
"time"
)
// Config — wires the cascade. Build via New; a zero-value Router is unusable.
type Config struct {
// Grammars — stage-0 exact-match rules. DefaultGrammars(actMatcher) wires
// the wake-word act fast path; the daemon may append more.
Grammars []Grammar
// Classifier — stage-1 nearest-centroid classifier. Must be seeded with
// ~10 examples/intent at bootstrap (per spec) before free-form routing
// is trustworthy; until then Route returns ErrNoIntents on free-form input.
Classifier *Classifier
// Extractor — stage-2 per-intent slot extraction. Any nil sub-parser just
// leaves the corresponding Has* flag false for that intent.
Extractor Extractor
// Threshold — stage-3 confidence gate. Below ⇒ Clarify, don't guess. The
// spec leaves this open (defines how often maven asks vs guesses on free-
// form input; the whole reactive mvp feel rides on it). The daemon sets it.
Threshold float64
// LLM — optional agentic router. When set, Route consults it after stage-0
// and before the classifier cascade, classifying the utterance via a
// grammar-constrained call to the resident model (Qwen3-1.7B). On any
// error/parse failure, falls through to the classifier (never fails the
// turn on the model).
LLM *LLMRouter
}
// Router — the deterministic cascade. Route never guesses: stage 0 wins
// outright, stage 1 scores, stage 2 extracts, stage 3 gates. The SLM only
// phrases the reply — it never owns the route.
type Router struct {
grammars []Grammar
classifier *Classifier
extractor Extractor
threshold float64
llm *LLMRouter
}
func New(cfg Config) *Router {
return &Router{
grammars: cfg.Grammars,
classifier: cfg.Classifier,
extractor: cfg.Extractor,
threshold: cfg.Threshold,
llm: cfg.LLM,
}
}
// Route — the cascade: stage 0 (exact match) → 1 (classify) → 2 (extract) →
// 3 (confidence gate).
//
// Stage 0 wins outright: returns at confidence 1.0, no classifier.
// Otherwise the classifier scores every intent; the best wins; slots are
// extracted for that intent. If the winning score < threshold the Decision
// is flagged Clarify (the daemon asks rather than guesses — same shape as
// since(key)==null → don't fire: a misrouted fact is a confident wrong write,
// worse than a gap).
func (r *Router) Route(ctx context.Context, utterance string, now time.Time) (Decision, error) {
// stage 0 — exact match / grammar. First match wins; grammars are ordered.
// Grammars like time/date/reminder don't expect a wake-word prefix, but
// the STT often includes one (transcribed phonetically, any script) — try
// the wake-stripped utterance too so those grammars still fire.
stripped, hadWake := StripWakeToken(utterance)
for _, g := range r.grammars {
m := g.Pattern.FindStringSubmatch(utterance)
if m == nil && hadWake {
m = g.Pattern.FindStringSubmatch(stripped)
}
if m == nil {
continue
}
d, ok := g.Build(m)
if !ok {
continue // grammar matched shape but not content → fall through
}
d.Utterance = utterance
return d, nil
}
// stage 1a — LLM router (when wired). It reasons over the utterance instead
// of nearest-centroid guessing. On any error/parse-fail, fall through to the
// classifier cascade (never fail the turn on the model).
if r.llm != nil {
if d, ok, err := r.llm.Route(ctx, utterance, now); err == nil && ok {
d.Utterance = utterance
return d, nil
} else if err != nil {
log.Printf("router: llm route fell back to classifier: %v", err)
}
}
// stage 1 — intent classifier.
results, err := r.classifier.Classify(ctx, utterance)
if err != nil {
return Decision{}, err
}
best := results[0]
// stage 2 — slot extraction for the winning intent.
d := Decision{
Utterance: utterance,
Stage: 2,
Intent: best.Intent,
Confidence: best.Score,
Slots: r.extractor.Extract(ctx, best.Intent, utterance, now),
}
// stage 3 — confidence gate. Below threshold ⇒ clarify, don't guess.
if d.Confidence < r.threshold {
d.Stage = 3
d.Clarify = true
}
return d, nil
}
// CorrectMisroute — the user corrected a bad classification. Appends a new
// example for the corrected intent (append-only — grows the classifier, no
// retrain). Same shape as nudges.outcome tuning cooldowns: more reliable over
// time, introspectable, no model surgery.
func (r *Router) CorrectMisroute(ctx context.Context, utterance string, corrected Intent) error {
return r.classifier.AddExample(ctx, corrected, utterance)
}