Add a read-only /ecosystem page that consumes the sibling services' JSON APIs (Nexus entities, Praxis attention, Hexis capabilities), fetched concurrently with honest per-panel error states. Siblings stay headless — mavweb is their human surface (arch §16). Wired via mavweb -nexus/-praxis/-hexis flags; mavweb joins the ecosystem compose network. Fix mobile horizontal overflow across all pages: .content is a flex child with default min-width:auto, so it refused to shrink below the tables' intrinsic width. min-width:0 lets wide tables pan inside .scroll instead of dragging the page sideways. Verified via CDP geometry check (scrollWidth === clientWidth at 430px). Also includes in-progress Ethos UI redesign, ecosystem deploy compose, and planning docs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
29 KiB
Plan — Sub-project 1: Router-as-LFM + Foundation
Historical/completed foundation. The shared llama-server client, router, fallback and replier described here were implemented. The resident-model decision changed on 2026-07-18 from LFM to locally trained Qwen3-1.7B. Do not use the embedded LFM model paths or old single-object examples as current ops guidance; see
2026-07-18-qwen3-resident-training-eval.md.
Scope from
REARCH.md. Make Maven trustworthy: the LFM becomes the router (fixes "messes up queries" / "doesn't take notes"), the engine actually runs (fixes stub replies), dates stop being read as "number dot number dot number", and telegram becomes a reach channel. NOT in scope: on-demand 4B reasoner, digestion worker, voice-quality/custom-TTS.
Findings that shape the plan
- Notes + reminders already work in code (
applyAction→WriteNote/CreateReminder,cmd/mavend/voice.go:613,548). The user's "she doesn't take notes / remind" is (a) the embedder-classifier misrouting the utterance away fromIntentNote/IntentReminder, and (b)StubRepliergiving canned confirmations. The LFM router + LLM replier fix both. - Telegram sink already wired (
cmd/mavend/main.go:306-321,delivery.Config.Telegram). Only the deployed config lacks atelegramblock. - Phraser runs llama-server (
internal/phraser/llmphraser.go) over/v1/chat/completionsat base URLp.port. Deployedmavend.jsonhas nophraserblock →phraser.NewStub()(main.go:261). Engine is off. - Import direction:
phraser → router. So the LLM router lives inrouter, using a new dependency-freeinternal/llmclient that both the phraser and router share (one llama-server, two callers).router → llmis acyclic;phraser → llmis acyclic.
File structure
| File | Create/Modify | Responsibility |
|---|---|---|
internal/ttsnorm/ttsnorm.go |
create | RU date/number → speakable text, before TTS |
internal/ttsnorm/ttsnorm_test.go |
create | table tests for date/time/number forms |
cmd/mavend/voice.go |
modify | call ttsnorm.Speakable in reply(); wire LLM router + LLM replier |
internal/delivery/voicesink/voicesink.go |
modify | normalize nudge text before synth |
cmd/mavend/reactive_notes_test.go |
create | notes+reminders round-trip integration test |
internal/llm/client.go |
create | llama-server /v1/chat/completions client w/ GBNF grammar |
internal/llm/client_test.go |
create | httptest server round-trip + grammar passthrough |
internal/phraser/llmphraser.go |
modify | add BaseURL() accessor |
internal/router/llmrouter.go |
create | Action, GBNF grammar, prompt, parse, LLMRouter |
internal/router/llmrouter_test.go |
create | mock client: intent+slot mapping, fallback |
internal/router/router.go |
modify | Config.LLM; Route consults LLM after stage-0 miss |
cmd/mavend/replier_llm.go |
create | llmReplier (voice.Replier) — natural reactive replies |
cmd/mavend/replier_llm_test.go |
create | mock client: reply per intent, stub fallback |
deploy/mavend.json |
modify | add phraser + telegram blocks |
docker-compose.yml |
modify | mount LFM model + telegram env |
AGENTS.md |
modify | document LFM model download + router-LFM |
Task 1 — RU date/number TTS normalizer
Goal: dates/times/bare number-dot-number strings are spoken as words, never "number dot number dot number".
Files: create internal/ttsnorm/ttsnorm.go, internal/ttsnorm/ttsnorm_test.go;
modify cmd/mavend/voice.go (reply), internal/delivery/voicesink/voicesink.go.
Acceptance criteria:
Speakable("напомню 10.07.2026")→"напомню 10 июля 2026".Speakable("встреча в 14:00")→"встреча в 14 часов 00 минут".Speakable("это 3.2.1 версия")→ digits joined with " точка " ("3 точка 2 точка 1") — the generic fallback so no.is voiced as a symbol.- Idempotent:
Speakable(Speakable(x)) == Speakable(x). - Pure, no deps beyond stdlib; RU month table.
Verify: go test ./internal/ttsnorm/ → ok.
Steps (TDD):
- Write
ttsnorm_test.go:
package ttsnorm
import "testing"
func TestSpeakable(t *testing.T) {
cases := []struct{ in, want string }{
{"напомню 10.07.2026", "напомню 10 июля 2026"},
{"срок 01.01", "срок 1 января"},
{"встреча в 14:00", "встреча в 14 часов 00 минут"},
{"в 9:05 подъём", "в 9 часов 05 минут подъём"},
{"это 3.2.1 версия", "это 3 точка 2 точка 1 версия"},
{"без чисел", "без чисел"},
}
for _, c := range cases {
if got := Speakable(c.in); got != c.want {
t.Errorf("Speakable(%q) = %q, want %q", c.in, got, c.want)
}
if got := Speakable(Speakable(c.in)); got != Speakable(c.in) {
t.Errorf("not idempotent for %q: %q", c.in, got)
}
}
}
- Implement
ttsnorm.go. Order matters: dates (with year, then without) → times → generic numeric-dot fallback.
// Package ttsnorm rewrites machine-formatted dates/times/numbers into RU text
// a TTS voice speaks naturally — so "10.07.2026" is not read as "number dot
// number dot number". Pure, deterministic; runs on reply/nudge text before synth.
package ttsnorm
import (
"regexp"
"strconv"
"strings"
)
var months = [...]string{"", "января", "февраля", "марта", "апреля", "мая",
"июня", "июля", "августа", "сентября", "октября", "ноября", "декабря"}
var (
reDateY = regexp.MustCompile(`\b(\d{1,2})\.(\d{1,2})\.(\d{4})\b`)
reDate = regexp.MustCompile(`\b(\d{1,2})\.(\d{1,2})\b`)
reTime = regexp.MustCompile(`\b(\d{1,2}):(\d{2})\b`)
reDots = regexp.MustCompile(`\b\d+(?:\.\d+)+\b`)
)
// Speakable rewrites d.m.y, d.m, h:mm, and residual dotted-number runs.
func Speakable(s string) string {
s = reDateY.ReplaceAllStringFunc(s, func(m string) string {
p := reDateY.FindStringSubmatch(m)
return spokenDate(p[1], p[2], p[3])
})
s = reDate.ReplaceAllStringFunc(s, func(m string) string {
p := reDate.FindStringSubmatch(m)
return spokenDate(p[1], p[2], "")
})
s = reTime.ReplaceAllStringFunc(s, func(m string) string {
p := reTime.FindStringSubmatch(m)
return p[1] + " часов " + p[2] + " минут"
})
s = reDots.ReplaceAllStringFunc(s, func(m string) string {
return strings.Join(strings.Split(m, "."), " точка ")
})
return s
}
func spokenDate(dd, mm, yyyy string) string {
mi, _ := strconv.Atoi(mm)
if mi < 1 || mi > 12 {
return dd + " " + mm + gap(yyyy) // out-of-range: leave numbers, drop dot
}
day := strconv.Itoa(mustInt(dd)) // strip leading zero: "01" → "1"
out := day + " " + months[mi]
if yyyy != "" {
out += " " + yyyy
}
return out
}
func mustInt(s string) int { n, _ := strconv.Atoi(s); return n }
func gap(y string) string { if y == "" { return "" }; return " " + y }
- Wire into
reply()(cmd/mavend/voice.go:972), before synth:
func (h *reactiveHandler) reply(ctx context.Context, text string, _ []string) (voice.PushToTalkResp, error) {
spoken := ttsnorm.Speakable(text)
log.Printf("voice: reply → %q", text)
audioOut, err := h.tts.Synthesize(ctx, spoken)
if err != nil {
log.Printf("voice: tts error: %v", err)
return voice.PushToTalkResp{ReplyText: text, ReplyAudio: audio.Audio{}}, nil
}
return voice.PushToTalkResp{ReplyText: text, ReplyAudio: audioOut}, nil
}
(ReplyText stays raw for on-screen display; only audio is normalized.) Add the
ttsnorm import.
-
Wire into
voicesink.gobefore itsSynthesizecall (same one-linettsnorm.Speakableon the nudge body). Add import. -
go test ./internal/ttsnorm/ ./cmd/mavend/ ./internal/delivery/...green.
Task 2 — Notes + reminders round-trip test (+ fix any real break)
Goal: prove — and lock with a test — that a note utterance persists a note and a reminder utterance persists a reminder through the reactive handler.
Files: create cmd/mavend/reactive_notes_test.go. Modify applyAction /
seeds ONLY if the test exposes a real break.
Acceptance criteria:
- Driving
handleTextwith a note utterance results in anotesrow whose text matches; with a reminder utterance, aremindersrow with a futurefire_ts. - Test uses the real
store.Store(in-tmpfs or file),ipc.NewStoreAPI(st), aHashEmbedder, andStubReplier— no LLM required. - If routing (not the write) is what fails, the test forces intent by using a
stage-0 grammar utterance (
"напомни завтра позвонить маме") and a directly constructedIntentNotedecision path, isolating capture from routing.
Verify: go test ./cmd/mavend/ -run TestReactiveNotesReminders -v → PASS.
Steps (TDD):
- Write the test: open a temp
store.Store, build a minimalreactiveHandler(fields:api: ipc.NewStoreAPI(st),embedder: router.NewHashEmbedder(1024),router: buildRouter(emb, matcher, thr),replier: voice.NewStubReplier(),now: time.Now,memStore: memory.NewInMemoryStore(),dataStore: st,timeParser: router.NewPythonDateParser()). Callh.handleText(ctx, "напомни завтра в 9:00 позвонить маме"); assertst.DueReminders/ListRemindersreturns one row withfire_ts > now. Then a note: constructdec := router.Decision{Intent: router.IntentNote, Utterance: "запомни что кофе закончился"}, callh.applyAction(ctx, dec); assertst.RecentNotes(ctx, 10)contains the text. - Run it. If it fails at the write, fix the write in
applyAction. If it fails only because routing sent the utterance elsewhere, that is Task 4/5's job — this test pins the capture contract, so force the intent as above and leave a// routing covered by llmrouter_testnote. - Keep the test deterministic (no network:
NewPythonDateParsershells out — if unavailable in CI, use a fixedrouter.Slots{Time: now.Add(24h), HasTime:true}decision for the reminder capture assertion instead of the parser path).
Task 3 — internal/llm shared llama-server client
Goal: one client type both the phraser and the LFM router use to call the already-running llama-server, with optional GBNF grammar for constrained output.
Files: create internal/llm/client.go, internal/llm/client_test.go;
modify internal/phraser/llmphraser.go (add BaseURL()).
Acceptance criteria:
Client.Complete(ctx, Req{System, User, Grammar, MaxTokens})POSTs to<base>/v1/chat/completions, returns the assistant message content.- When
Req.Grammar != "", the request body carries a top-level"grammar"field (llama.cpp extension) verbatim. LLMPhraser.BaseURL()returns the spawned server's base URL (p.port).
Verify: go test ./internal/llm/ → ok.
Steps (TDD):
client_test.go: stand uphttptest.NewServerthat (a) asserts the decoded body hasmessages[0].role=="system",messages[1].content=="hi", andgrammar=="root ::= \"x\"", and (b) replies{"choices":[{"message":{"content":"ok"}}]}. AssertCompletereturns"ok".- Implement
client.go:
// Package llm is the shared llama-server completion client — one seam both the
// phraser (talking back) and the router (routing) call. It does NOT spawn the
// server; the daemon owns one llama-server (spawned by the phraser) and hands
// its base URL here, so a single resident model serves both callers.
package llm
import (
"bytes"
"context"
"encoding/json"
"fmt"
"net/http"
"time"
)
type Client struct {
base string
http *http.Client
}
func New(baseURL string, timeout time.Duration) *Client {
return &Client{base: baseURL, http: &http.Client{Timeout: timeout}}
}
type Req struct {
System string
User string
Grammar string // GBNF; empty ⇒ unconstrained
MaxTokens int
}
type msg struct {
Role string `json:"role"`
Content string `json:"content"`
}
type body struct {
Messages []msg `json:"messages"`
MaxTokens int `json:"max_tokens,omitempty"`
Grammar string `json:"grammar,omitempty"`
Temp float64 `json:"temperature"`
}
type resp struct {
Choices []struct {
Message msg `json:"message"`
} `json:"choices"`
}
func (c *Client) Complete(ctx context.Context, r Req) (string, error) {
b, _ := json.Marshal(body{
Messages: []msg{{"system", r.System}, {"user", r.User}},
MaxTokens: r.MaxTokens,
Grammar: r.Grammar,
Temp: 0,
})
req, err := http.NewRequestWithContext(ctx, "POST", c.base+"/v1/chat/completions", bytes.NewReader(b))
if err != nil {
return "", err
}
req.Header.Set("Content-Type", "application/json")
httpResp, err := c.http.Do(req)
if err != nil {
return "", err
}
defer httpResp.Body.Close()
if httpResp.StatusCode != 200 {
return "", fmt.Errorf("llm: status %d", httpResp.StatusCode)
}
var out resp
if err := json.NewDecoder(httpResp.Body).Decode(&out); err != nil {
return "", err
}
if len(out.Choices) == 0 {
return "", fmt.Errorf("llm: no choices")
}
return out.Choices[0].Message.Content, nil
}
- Add to
llmphraser.go:func (p *LLMPhraser) BaseURL() string { return p.port }(the field already holds the full base URL incl. scheme+host+port). go test ./internal/llm/ ./internal/phraser/green.
Task 4 — LFM router core (in internal/router)
Goal: an agentic router that turns an utterance into a structured Action
(intent + raw slots + escalate flag) via one grammar-constrained LFM call, with
deterministic mapping to a Decision. Classifier stays as the fallback.
Files: create internal/router/llmrouter.go, internal/router/llmrouter_test.go;
modify internal/router/router.go.
Acceptance criteria:
LLMRouter.Route(ctx, utterance, now)returns arouter.DecisionwithIntent∈ the existing enum,Slotspopulated (Key/Valuefor fact,Textfor note/query/reminder,Fnleft for the act matcher), andStage = 1.- Output is constrained by a GBNF grammar to a JSON object with an
intentenum and optional string fields — malformed model output cannot escape the enum. - On any error (LLM down, parse fail)
Routereturns(Decision{}, false, err)and the caller falls back to the classifier cascade. - Time is NOT parsed here (kept deterministic): reminder actions carry the raw
text in
Slots.Text; the daemon's existingtimeParserfallback (voice.go:552) resolves it.
Verify: go test ./internal/router/ -run TestLLMRouter -v → PASS.
Steps (TDD):
llmrouter_test.gowith a mock completer (interface, not*llm.Client):
type mockLLM struct{ out string; err error }
func (m mockLLM) Complete(ctx context.Context, r llm.Req) (string, error) { return m.out, m.err }
func TestLLMRouterFactMapping(t *testing.T) {
lr := NewLLMRouter(mockLLM{out: `{"intent":"fact","key":"water","value":"выпил"}`})
d, ok, err := lr.Route(context.Background(), "я выпил воду", time.Now())
if err != nil || !ok { t.Fatalf("ok=%v err=%v", ok, err) }
if d.Intent != IntentFact || d.Slots.Key != "water" || !d.Slots.HasKey {
t.Fatalf("bad decision %+v", d)
}
}
func TestLLMRouterNoteMapping(t *testing.T) { /* intent:"note" → Slots.Text set */ }
func TestLLMRouterBadJSONFallsBack(t *testing.T) {
lr := NewLLMRouter(mockLLM{out: `garbage`})
if _, ok, err := lr.Route(context.Background(), "x", time.Now()); ok || err == nil {
t.Fatal("want ok=false, err!=nil on bad json")
}
}
llmrouter.go:
package router
import (
"context"
"encoding/json"
"fmt"
"strings"
"time"
"github.com/kami/maven/internal/llm"
)
// Completer — the LLM seam (mockable). *llm.Client satisfies it.
type Completer interface {
Complete(ctx context.Context, r llm.Req) (string, error)
}
// LLMRouter — the agentic router. One grammar-constrained call classifies the
// utterance and pulls raw slots; deterministic parsers (time) refine downstream.
type LLMRouter struct{ c Completer }
func NewLLMRouter(c Completer) *LLMRouter { return &LLMRouter{c: c} }
// routeGrammar — GBNF constraining the model to a fixed-shape JSON object with
// an intent enum. Prevents free-form drift from a sub-1B model.
const routeGrammar = `
root ::= "{" ws "\"intent\"" ws ":" ws intent ("," ws field)* ws "}"
intent ::= "\"fact\"" | "\"reminder\"" | "\"note\"" | "\"query\"" | "\"act\"" | "\"chat\"" | "\"system\""
field ::= key ws ":" ws string
key ::= "\"key\"" | "\"value\"" | "\"text\"" | "\"verb\""
string ::= "\"" ([^"\\] | "\\" .)* "\""
ws ::= [ \t\n]*
`
const routeSystem = `Ты — маршрутизатор Maven. По реплике пользователя верни ОДИН JSON-объект:
{"intent": one of fact|reminder|note|query|act|chat|system, ...}
Правила: fact — состояние которое надо запомнить и отслеживать ("я выпил воду" → key=water, value=выпил); reminder — просьба напомнить в будущем (text = что напомнить); note — заметка/предпочтение без отслеживания (text = текст); query — вопрос (text = вопрос); act — команда выполнить действие на сервере (verb = глагол); chat — свободный разговор (text); system — время/дата/день. Только JSON, без пояснений.`
type routeAction struct {
Intent string `json:"intent"`
Key string `json:"key"`
Value string `json:"value"`
Text string `json:"text"`
Verb string `json:"verb"`
}
func (lr *LLMRouter) Route(ctx context.Context, utterance string, now time.Time) (Decision, bool, error) {
raw, err := lr.c.Complete(ctx, llm.Req{System: routeSystem, User: utterance, Grammar: routeGrammar, MaxTokens: 128})
if err != nil {
return Decision{}, false, err
}
var a routeAction
if err := json.Unmarshal([]byte(strings.TrimSpace(raw)), &a); err != nil {
return Decision{}, false, fmt.Errorf("llmrouter: parse %q: %w", raw, err)
}
d := Decision{Utterance: utterance, Stage: 1, Confidence: 1.0}
switch Intent(a.Intent) {
case IntentFact:
d.Intent = IntentFact
d.Slots.Key, d.Slots.Value = a.Key, a.Value
d.Slots.HasKey = a.Key != ""
case IntentReminder:
d.Intent = IntentReminder
d.Slots.Text = firstNonEmpty(a.Text, utterance) // time parsed downstream
case IntentNote:
d.Intent = IntentNote
d.Slots.Text = firstNonEmpty(a.Text, utterance)
case IntentQuery:
d.Intent = IntentQuery
d.Slots.Text = firstNonEmpty(a.Text, utterance)
case IntentAct:
d.Intent = IntentAct // Fn resolved by the act matcher in the daemon
d.Slots.Text = firstNonEmpty(a.Verb, utterance)
case IntentSystem:
d.Intent = IntentSystem
default:
d.Intent = IntentChat
d.Slots.Text = firstNonEmpty(a.Text, utterance)
}
return d, true, nil
}
func firstNonEmpty(a, b string) string { if strings.TrimSpace(a) != "" { return a }; return b }
router.go: addLLM *LLMRoutertoConfigand theRouterstruct; inRoute, after the stage-0 loop and beforer.classifier.Classify:
// stage 1a — LLM router (when wired). It reasons over the utterance instead
// of nearest-centroid guessing. On any error/parse-fail, fall through to the
// classifier cascade (never fail the turn on the model).
if r.llm != nil {
if d, ok, err := r.llm.Route(ctx, utterance, now); err == nil && ok {
d.Utterance = utterance
return d, nil
} else if err != nil {
log.Printf("router: llm route fell back to classifier: %v", err)
}
}
Add "log" import; set llm: cfg.LLM in New.
go test ./internal/router/green (existing classifier tests still pass —LLMnil by default preserves the cascade).
Task 5 — Wire the LFM router into the daemon
Goal: the daemon builds one llm.Client from the running phraser and injects
it into the router, so production routing goes through the LFM (with classifier
fallback). Act intents still resolve Fn via the existing matcher.
Files: modify cmd/mavend/voice.go (wireVoice, buildRouter), and pass
the phraser through (already a wireVoice param phr phraser.Phraser).
Acceptance criteria:
- When
phris an*phraser.LLMPhraser,buildRouterreceives a non-nilCompleterand the router'sConfig.LLMis set; whenphris the Stub,LLMis nil and routing uses the classifier (unchanged behavior / tests). - Act
Fnresolution: after an LLMIntentActdecision, the handler runs the utterance/verb through the existingmatcher.Match(as stage-0 does) soHasFn/Argsare filled beforeapplyAction.
Verify: go build ./cmd/mavend/ && go test ./cmd/mavend/ → ok; manual:
handleText on a paraphrased note/reminder (not matching any seed) routes
correctly (covered by an added cmd/mavend test using a mock completer).
Steps:
- In
wireVoice, after the phraser is known, build the optional client:
var llmRouter *router.LLMRouter
if lp, ok := phr.(*phraser.LLMPhraser); ok {
cli := llm.New(lp.BaseURL(), 20*time.Second)
llmRouter = router.NewLLMRouter(cli)
}
rtr := buildRouter(emb, matcher, threshold, llmRouter)
Add internal/llm import.
-
Extend
buildRoutersignature withllmR *router.LLMRouterand setLLM: llmRinrouter.Config. -
Act
Fnfill: inapplyAction'sIntentActbranch, when!dec.Slots.HasFnanddec.Slots.Text != "", tryh.tools/matcher resolution beforeproposeGap(mirror stage-0'sactMatcher.Match). Add amatcherfield toreactiveHandler(set inwireVoice) and:
case router.IntentAct:
if !dec.Slots.HasFn && dec.Slots.Text != "" && h.matcher != nil {
if fn, args, ok := h.matcher.Match(dec.Slots.Text); ok {
dec.Slots.Fn, dec.Slots.Args, dec.Slots.HasFn = fn, args, true
}
}
if !dec.Slots.HasFn { return h.proposeGap(ctx, dec) }
// ... unchanged executor path
-
Add a
cmd/mavendtest: build a handler with a mock-completer-backedrouter.LLMRouter, drivehandleText("надо не забыть купить молоко"), assert it routes toIntentNote/IntentReminder(per mock) and persists. -
go test ./cmd/mavend/ ./internal/router/green.
Task 6 — LLM-backed reactive Replier (talking back)
Goal: reactive confirmations are phrased by the LFM ("записала: кофе
закончился" in Maven's voice), not StubReplier canned strings — the visible
payoff of the engine being on. Falls back to the Stub on any LLM error.
Files: create cmd/mavend/replier_llm.go, cmd/mavend/replier_llm_test.go;
modify cmd/mavend/voice.go (set replier to the LLM impl when the engine is on).
Acceptance criteria:
llmReplier.Reply(dec)returns a short RU confirmation generated from the decision (intent + slots), feminine self-reference, ≤ ~120 chars.- On LLM error/empty, returns
StubReplier.Reply(dec)(never blank, never a wire error). - Clarify decisions still yield a clarify prompt.
Verify: go test ./cmd/mavend/ -run TestLLMReplier -v → PASS.
Steps (TDD):
replier_llm_test.go: mock completer returning"записала, кофе закончился"→ assertReply(noteDecision)returns it; mock returning error → assert it equalsvoice.NewStubReplier().Reply(dec).replier_llm.go:
package main
import (
"context"
"time"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice"
)
type completer interface {
Complete(ctx context.Context, r llm.Req) (string, error)
}
// llmReplier phrases reactive confirmations with the resident LFM. Stub is the
// floor on any error (offline-safe). Maven speaks as "she", feminine RU.
type llmReplier struct {
c completer
stub *voice.StubReplier
}
func newLLMReplier(c completer) *llmReplier { return &llmReplier{c: c, stub: voice.NewStubReplier()} }
const replySystem = `Ты — Maven, домашняя ассистентка (о себе — в женском роде). Одной короткой фразой (≤120 симв) подтверди действие пользователю тепло и по-русски. Без кавычек и пояснений.`
func (r *llmReplier) Reply(d router.Decision) string {
if d.Clarify {
return r.stub.Reply(d)
}
ctx, cancel := context.WithTimeout(context.Background(), 8*time.Second)
defer cancel()
out, err := r.c.Complete(ctx, llm.Req{System: replySystem, User: replyContext(d), MaxTokens: 64})
if err != nil || out == "" {
return r.stub.Reply(d)
}
return out
}
// replyContext renders the decision into a compact RU description for the model.
func replyContext(d router.Decision) string {
switch d.Intent {
case router.IntentFact:
return "записала факт: " + d.Slots.Key + " " + d.Slots.Value
case router.IntentNote:
return "сохранила заметку: " + d.Slots.Text
case router.IntentReminder:
return "поставила напоминание: " + d.Slots.Text
default:
return string(d.Intent) + ": " + d.Slots.Text
}
}
Note: Reply satisfies voice.Replier (sync). The 8s timeout bounds a slow
sub-1B turn; the Stub floor keeps the round-trip alive if it lapses.
3. In wireVoice, when llmRouter != nil (engine on), set
replier: newLLMReplier(cli) instead of voice.NewStubReplier() (reuse the
same cli). Keep Stub when the engine is off.
4. go test ./cmd/mavend/ green.
Task 7 — Turn the engine on + telegram reach (deploy config)
Goal: homesrv actually runs the LFM (router + phraser + replier live) and can reach the user over telegram. This is the commit that flips Maven from stub to real on the deploy box.
Files: modify deploy/mavend.json, docker-compose.yml, AGENTS.md.
Ops: place the LFM gguf on homesrv; docker compose up -d.
Acceptance criteria:
deploy/mavend.jsonhas aphraserblock pointing at the LFM gguf (LFM2.5-1.2B-Instruct-Q4_K_M.gguf) withbin_path: llama-server,n_gpu_layerstuned for the Vega iGPU (start0= CPU; the 1.2B is real-time on CPU, avoids fragile ROCm),n_ctx: 2048.deploy/mavend.jsonhas atelegramblock (bot_token,chat_id) — token from env, not committed.- On homesrv:
docker logs maven-mavend-1shows a phraser/llama line andvoice listening, and no fall-through to Stub. - A push-to-talk (or
/api/chat) "запомни что кофе закончился" returns an LFM phrased confirmation AND persists a note (check/dashrecent notes). - A sev4 nudge with no live voice session lands in telegram.
Verify:
ssh kami@192.168.1.104 'docker logs maven-mavend-1 2>&1 | grep -iE "llama|phraser|listening"'
# expect a llama-server spawn line + "voice listening on ..."
ssh kami@192.168.1.104 'curl -s localhost:9201/api/chat -d "{\"text\":\"запомни что кофе закончился\"}"'
# expect a non-stub RU confirmation; then /dash shows the note
Steps:
- Add to
deploy/mavend.json:
"phraser": {
"model_path": "/opt/maven/models/llm/LFM2.5-1.2B-Instruct-Q4_K_M.gguf",
"bin_path": "llama-server",
"n_gpu_layers": 0,
"n_ctx": 2048,
"timeout": "20s"
},
"telegram": { "bot_token": "${TELEGRAM_BOT_TOKEN}", "chat_id": "${TELEGRAM_CHAT_ID}" }
(Match the existing config's env-substitution convention; if it doesn't
support ${...}, wire the two telegram values via compose env + a small
config env-read, consistent with db_key_env.)
2. docker-compose.yml: mount ./models/llm:/opt/maven/models/llm:ro on
mavend; add TELEGRAM_BOT_TOKEN / TELEGRAM_CHAT_ID to its env (from a
gitignored deploy/telegram.env, mirroring deploy/db_key.env). Ensure
llama-server is on PATH in the image (it's already used by design; confirm
the Dockerfile installs it — if not, add it).
3. Ops: copy models/llm/LFM2.5-1.2B-Instruct-Q4_K_M.gguf to
homesrv /opt/maven/models/llm/ (the box already has /usr/local/bin/llama-server).
If the LFM gguf isn't on disk yet, download it (HF, sha-checked) — document in
AGENTS.md next to the embedder section.
4. docker compose up -d mavend && docker logs -f maven-mavend-1 — confirm the
phraser spawns and no phraser: NewStub path. Run the two verify curls.
5. Update AGENTS.md: LFM model download + note that routing is now LFM-first
with classifier fallback (REARCH.md is the design of record).
Self-review
- Placeholders: none — every code step carries real code. Deploy steps (Task 7) are concrete commands, not "configure somehow".
- Type consistency:
Completerinterface (Complete(ctx, llm.Req) (string, error)) is used identically in Task 4 (router), Task 5 (wiring), Task 6 (replier);*llm.Clientsatisfies it.router.Decision/Slotsfield names match the real structs (HasKey,Text,Fn,Args).LLMRouter.Routereturns(Decision, bool, error)consistently across Task 4 def and Task 4 router.go call site. - Fallback discipline: every LLM seam (router, replier) falls back to the existing deterministic floor on error — a slow/absent model degrades to today's behavior, never a broken turn. Engine-off path is byte-for-byte unchanged (LLM nil, Stub replier), so all existing tests hold.
- Scope: no 4B reasoner, no digestion worker, no custom TTS.
escalateis intentionally omitted from the grammar this sub-project (the reasoner it would escalate to doesn't exist yet) — added in sub-project 2.
Ordering & deps
1 (ttsnorm) and 2 (notes/reminders test) are independent — do first, any order. 3 (llm client) → 4 (llm router) → 5 (wire) is the critical chain. 6 (replier) depends on 3. 7 (deploy) depends on 3–6 landed. Each task = one commit.