Compare commits

..

4 Commits

Author SHA1 Message Date
kami 9a3bcd7c46 Merge the thinking-off measurement 2026-07-31 13:31:31 +04:00
kami 04c1088088 Measure thinking off on routing properly — it does not win (#376)
The 67.1% "thinking off" column in ROUTING-EVAL-31-07-2026.md was an
artefact. It came from a hand-rolled HTTP client in the eval test that
did not send repeat_penalty, so it differed from the reference run on two
axes and the penalty was the one that mattered.

Re-scored back to back on an idle box with everything else held equal:
thinking off is identical to thinking on, case for case, same confusion
matrix, same three unparseable replies. A direct probe of the running
llama-server shows enable_thinking, thinking and reasoning_budget are all
ignored for this model on this build, so there was nothing to turn off.

No defaults changed. The misleading third configuration is removed from
internal/router/eval so its table cannot be quoted again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:30:22 +04:00
kami 07c191d8b8 Merge dialogue session persistence 2026-07-31 13:21:03 +04:00
kami c668310b3e Persist the dialogue session so a restart keeps the conversation
Vikunja #363. The follow-up session was a plain in-memory map, so any
mavend restart dropped the thread. It now mirrors to a small TTL-pruned
sqlite table and is loaded on startup; expired sessions are deleted on
load, not revived. Clarify's pending question is untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:20:30 +04:00
7 changed files with 360 additions and 107 deletions
+59 -10
View File
@@ -66,15 +66,64 @@ Three things this run settles:
`запиши что…` phrasings toward fact, and that suspicion stands — all five `ru-note-*`
cases now land on fact. Tracked as Vikunja #375.
**Thinking off is the best configuration measured so far**, on both accuracy and latency
(Vikunja #376). That is worth understanding before flipping: routing is a short
classification into a fixed enum with grammar-constrained output, so there is little to
reason about, and the thinking trace mostly gives a small model room to talk itself out of
the right answer. Phrasing is a different job and needs measuring separately.
The `thinking off` column above read as the best configuration measured so far (Vikunja #376).
**It was wrong** — see the controlled re-run below. Ignore that column.
Still `6 / 6` missed clarify — the router has no way to say "I don't know" (Vikunja #359).
That is unchanged by anything here.
## Thinking off — 31-07-2026, controlled re-run (Vikunja #376)
The "thinking off wins by 6 points" observation above **does not hold**. It was a measurement
artefact, and the earlier table's `thinking off` column should be ignored.
The thinking-off variant was scored by a hand-rolled HTTP client living in the test file
instead of `llm.Client`. That copy did not send `repeat_penalty`, which the real router does
send (`routeRepeatPenalty = 1.15`). So the two columns differed on two axes at once, and the
one that mattered was the penalty, not the thinking mode.
Re-measured with everything else held equal — same fixture, same prompt, same grammar, same
sampling, same idle box, the three configurations run back to back and never concurrently:
| | llm-only, thinking on | llm-only, thinking off | cascade+llm |
|---|---|---|---|
| intent-only accuracy | 59.2% (45/76) | 59.2% (45/76) | 61.8% (47/76) |
| full accuracy (intent+slots+gate) | 38.2% (29/76) | 38.2% (29/76) | 57.9% (44/76) |
| route errors | 3 | 3 | 0 |
| grammar violations | 3 (all 3 route errors) | 3 (same 3 cases) | 0 |
| missed clarify | 5 / 6 | 5 / 6 | 5 / 6 |
| p50 latency | 836ms | 920ms | 810ms |
| p95 latency | 1.41s | 2.00s | 1.31s |
Thinking off is not just a tie on the headline numbers — it is identical case for case, with
the same confusion matrix and the same three unparseable replies. The latency difference is
run-to-run noise on one box, and it points the wrong way here.
The reason is simpler than any accuracy argument: **this llama-server build ignores the
request-level thinking switch for this model.** Probed directly against the running server
with `chat_template_kwargs.enable_thinking = false`, `chat_template_kwargs.thinking = false`
and top-level `reasoning_budget = 0` — all three return a byte-identical answer with the
thinking trace still in `reasoning_content`, and the server reports the prompt prefix as
cached, meaning the rendered template did not change. There was never anything being turned
off, which is also why the numbers match exactly.
Nothing was defaulted. `internal/llm` still has no `chat_template_kwargs` field, `VoiceConfig`
has no thinking flag, and `deploy/mavend.json` is unchanged. The misleading third
configuration is removed from `internal/router/eval` so the table it produced cannot be quoted
again.
Two caveats worth saying out loud:
- **The fixture is 76 cases.** A 6-point difference on 76 cases is roughly 4-5 cases and would
not have been worth trusting even if it had reproduced. This one was exactly 0 cases, which
is a much easier call.
- **This is one server build and one checkpoint** (`b9351`, Qwen3.5-0.8B Q4_K_M). If the
#122 checkpoint or a newer llama.cpp does honour the switch, the question reopens — but it
reopens as an unmeasured question, not as a 6-point win.
Phrasing was **not** measured. Whether thinking helps there is still open, and now also blocked
on the same "can we even turn it off" question.
## Findings
### 1. The resident model does route better — 50.0% vs 36.8%
@@ -145,11 +194,11 @@ Note the grammar's `string ::= "\"" ([^"\\] | "\\" .)* "\""` is unbounded, so no
### 7. Two hypotheses tested and closed
- **Thinking mode is a non-issue.** Qwen3.5's template defaults `thinking = 1`, so
grammar-constrained JSON lands in `reasoning_content` with `content` empty —
`llm.Client`'s fallback handles it. A `thinking off` run scored *identically* (18/76,
48.7%, same p50). `internal/llm` deliberately does **not** grow a `chat_template_kwargs`
field.
- **Thinking mode is a non-issue.** Confirmed twice now, the second time properly — see the
controlled re-run section. Grammar-constrained JSON lands in `reasoning_content` with
`content` empty and `llm.Client`'s fallback handles it; the request-level switch does
nothing on this build. `internal/llm` deliberately does **not** grow a
`chat_template_kwargs` field.
- **Runaway array repetition does not reproduce.** An isolated smoke test with a stripped
grammar emitted `{"intent":"reminder"}` until `MaxTokens`; under the real `routeSystem`
prompt the few-shot examples anchor it to one object. 2 errors in 76, not 76.
+12 -1
View File
@@ -227,7 +227,18 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
}
// ----- dialogue (multi-turn slot carry-over; 2-min follow-up window) -----
dialogueSessions := dialogue.NewSessionStore(2 * time.Minute)
// Store-backed when the daemon passes a store, so a restart mid-conversation
// keeps the thread (Vikunja #363). Sessions past their TTL are dropped on
// load, never revived. Clarify's parked question stays in memory only.
var dialogueSessions *dialogue.SessionStore
if dataStore != nil {
dialogueSessions = dialogue.NewPersistentSessionStore(2*time.Minute, dataStore)
if err := dialogueSessions.Load(context.Background(), time.Now()); err != nil {
log.Printf("dialogue: load saved sessions: %v", err)
}
} else {
dialogueSessions = dialogue.NewSessionStore(2 * time.Minute)
}
clarifyStore := dialogue.NewClarifyStore(clarifyTTL)
timeParser := router.NewPythonDateParser()
+74
View File
@@ -1,8 +1,12 @@
package dialogue
import (
"context"
"encoding/json"
"sync"
"time"
"github.com/kami/maven/internal/store"
)
type Intent string
@@ -50,10 +54,22 @@ func (s *Session) IsExpired(now time.Time) bool {
return now.After(s.Timestamp.Add(s.TTL))
}
// SessionPersister — the bit of the store the session needs, as an interface
// so tests can swap it out. Data is an opaque blob: the store never looks
// inside, we encode the session as JSON here.
type SessionPersister interface {
SaveDialogueSession(ctx context.Context, id string, data []byte, ts time.Time, ttl time.Duration) error
DeleteDialogueSession(ctx context.Context, id string) error
LoadDialogueSessions(ctx context.Context, now time.Time) ([]store.DialogueSessionRow, error)
}
// SessionStore keeps the live sessions in a map (the fast path) and mirrors
// every write to the persister, so a daemon restart can load them back.
type SessionStore struct {
mu sync.RWMutex
sessions map[string]*Session
defaultTTL time.Duration
persist SessionPersister // may be nil: memory only (tests, no-store paths)
}
func NewSessionStore(defaultTTL time.Duration) *SessionStore {
@@ -66,6 +82,43 @@ func NewSessionStore(defaultTTL time.Duration) *SessionStore {
}
}
// NewPersistentSessionStore — same store, but writes also go to the DB.
// Call Load once after this to bring back sessions from a previous run.
func NewPersistentSessionStore(defaultTTL time.Duration, p SessionPersister) *SessionStore {
s := NewSessionStore(defaultTTL)
s.persist = p
return s
}
// Load — read the saved sessions back into memory. Anything past its TTL is
// dropped (and deleted from the DB by the store), never revived.
func (s *SessionStore) Load(ctx context.Context, now time.Time) error {
if s.persist == nil {
return nil
}
rows, err := s.persist.LoadDialogueSessions(ctx, now)
if err != nil {
return err
}
s.mu.Lock()
defer s.mu.Unlock()
for _, r := range rows {
var sess Session
if err := json.Unmarshal(r.Data, &sess); err != nil {
// A blob we can't read is not worth failing a startup over.
continue
}
if sess.TTL <= 0 {
sess.TTL = r.TTL
}
if sess.IsExpired(now) {
continue
}
s.sessions[r.ID] = &sess
}
return nil
}
func (s *SessionStore) Get(id string, now time.Time) *Session {
s.mu.RLock()
sess, ok := s.sessions[id]
@@ -87,12 +140,33 @@ func (s *SessionStore) Put(id string, sess *Session) {
s.mu.Lock()
s.sessions[id] = sess
s.mu.Unlock()
s.save(id, sess)
}
func (s *SessionStore) Delete(id string) {
s.mu.Lock()
delete(s.sessions, id)
s.mu.Unlock()
if s.persist != nil {
_ = s.persist.DeleteDialogueSession(context.Background(), id)
}
}
// save — mirror one session to the DB. Best effort: memory already has it, so
// a write error costs us the restart safety net, not the current turn.
func (s *SessionStore) save(id string, sess *Session) {
if s.persist == nil {
return
}
data, err := json.Marshal(sess)
if err != nil {
return
}
ts := sess.Timestamp
if ts.IsZero() {
ts = time.Now()
}
_ = s.persist.SaveDialogueSession(context.Background(), id, data, ts, sess.TTL)
}
func InheritSlots(prev, cur Slots) Slots {
+118
View File
@@ -0,0 +1,118 @@
package dialogue
import (
"context"
"path/filepath"
"testing"
"time"
"github.com/kami/maven/internal/store"
)
// openStore — a store on disk, so a second handle can reopen the same file.
func openStore(t *testing.T, path string) *store.Store {
t.Helper()
s, err := store.Open(context.Background(), path)
if err != nil {
t.Fatalf("open store: %v", err)
}
t.Cleanup(func() { _ = s.Close() })
return s
}
// A session written before a restart comes back and still merges a follow-up.
func TestSessionSurvivesRestart(t *testing.T) {
ctx := context.Background()
path := filepath.Join(t.TempDir(), "maven_test.db")
now := time.Now().UTC().Truncate(time.Millisecond)
first := openStore(t, path)
before := NewPersistentSessionStore(2*time.Minute, first)
before.Put("voice", &Session{
Intent: IntentReminder,
Slots: Slots{Text: "полить цветы", Time: now.Add(time.Hour), HasTime: true},
Timestamp: now,
TTL: 2 * time.Minute,
})
if err := first.Close(); err != nil {
t.Fatalf("close: %v", err)
}
// fresh handle, fresh in-memory map — as after a daemon restart
after := openStore(t, path)
reloaded := NewPersistentSessionStore(2*time.Minute, after)
if err := reloaded.Load(ctx, now.Add(10*time.Second)); err != nil {
t.Fatalf("Load: %v", err)
}
sess := reloaded.Get("voice", now.Add(10*time.Second))
if sess == nil {
t.Fatal("session did not survive the restart")
}
if sess.Intent != IntentReminder {
t.Fatalf("intent = %q, want reminder", sess.Intent)
}
// the follow-up carries no text of its own; it must inherit the old one
merged := InheritSlots(sess.Slots, Slots{Time: now.Add(2 * time.Hour), HasTime: true})
if merged.Text != "полить цветы" {
t.Fatalf("merged text = %q, want the earlier turn's text", merged.Text)
}
if !merged.Time.Equal(now.Add(2 * time.Hour)) {
t.Fatalf("merged time = %v, want the follow-up's time", merged.Time)
}
}
// A session past its TTL is dead: a restart must not bring it back.
func TestExpiredSessionNotResurrected(t *testing.T) {
ctx := context.Background()
path := filepath.Join(t.TempDir(), "maven_test.db")
now := time.Now().UTC().Truncate(time.Millisecond)
first := openStore(t, path)
before := NewPersistentSessionStore(2*time.Minute, first)
before.Put("voice", &Session{
Intent: IntentReminder,
Slots: Slots{Text: "полить цветы"},
Timestamp: now,
TTL: time.Minute,
})
if err := first.Close(); err != nil {
t.Fatalf("close: %v", err)
}
after := openStore(t, path)
reloaded := NewPersistentSessionStore(2*time.Minute, after)
later := now.Add(5 * time.Minute) // well past the 1-min TTL
if err := reloaded.Load(ctx, later); err != nil {
t.Fatalf("Load: %v", err)
}
if sess := reloaded.Get("voice", later); sess != nil {
t.Fatalf("expired session came back: %+v", sess)
}
// and it is gone from the DB too, not just from memory
rows, err := after.LoadDialogueSessions(ctx, later)
if err != nil {
t.Fatalf("LoadDialogueSessions: %v", err)
}
if len(rows) != 0 {
t.Fatalf("expired row still in the DB: %+v", rows)
}
}
// Delete removes the row as well, so an ended conversation stays ended.
func TestDeleteRemovesPersistedSession(t *testing.T) {
ctx := context.Background()
path := filepath.Join(t.TempDir(), "maven_test.db")
now := time.Now().UTC().Truncate(time.Millisecond)
s := openStore(t, path)
ss := NewPersistentSessionStore(2*time.Minute, s)
ss.Put("voice", &Session{Intent: IntentChat, Timestamp: now, TTL: time.Minute})
ss.Delete("voice")
rows, err := s.LoadDialogueSessions(ctx, now)
if err != nil {
t.Fatalf("LoadDialogueSessions: %v", err)
}
if len(rows) != 0 {
t.Fatalf("row survived Delete: %+v", rows)
}
}
+14 -96
View File
@@ -1,11 +1,8 @@
package eval
import (
"bytes"
"context"
"encoding/json"
"fmt"
"net/http"
"os"
"strings"
"testing"
@@ -29,13 +26,20 @@ import (
// a bake-off across checkpoints (#278, #250) produces tables you can tell
// apart. Point the variable at one server at a time.
//
// Three configurations, because "the LLM router" is ambiguous and the three
// numbers answer different questions:
// Two configurations, because "the LLM router" is ambiguous and the two numbers
// answer different questions:
//
// llm-only — the model alone. Measures the prompt + grammar contract.
// cascade+llm — what #320 would actually ship: stage-0 grammar, then the
// model, then the classifier as the failure floor.
// llm-no-thinking — diagnostic only, not a shippable path (see below).
// llm-only — the model alone. Measures the prompt + grammar contract.
// cascade+llm — what #320 would actually ship: stage-0 grammar, then the
// model, then the classifier as the failure floor.
//
// There used to be a third, "thinking off", which looked 6 points better. It is
// gone: it was measured with a hand-rolled HTTP client that quietly dropped
// repeat_penalty, so the gap was the missing penalty and not the thinking mode.
// Re-measured with everything else held equal, thinking off scores exactly the
// same, case for case — and a direct probe shows this llama-server build ignores
// enable_thinking / reasoning_budget for this model anyway, so there was nothing
// to turn off. Full write-up in ROUTING-EVAL-31-07-2026.md (Vikunja #376).
func TestLLMRouterBaseline(t *testing.T) {
base := os.Getenv("MAVEN_LLM_URL")
if base == "" {
@@ -96,104 +100,18 @@ func TestLLMRouterBaseline(t *testing.T) {
}
t.Log("\n" + repCascade.String() + repCascade.Failures())
// llm-no-thinking: same prompt and grammar with the chat template's
// thinking mode off. Qwen3.5's template defaults thinking=1, so under a
// grammar the constrained JSON lands in reasoning_content with content
// empty — llm.Client's ReasoningContent fallback is what makes the router
// work at all today, by accident rather than design.
//
// MEASURED 2026-07-31: this variant scores identically to as-deployed
// (18/76, 48.7% intent-only, 2 errors, same p50). Thinking mode is a
// non-issue under a grammar — llama.cpp constrains the same token stream
// either way. Kept so the question stays answered instead of being
// re-asked, and so internal/llm does NOT grow a chat_template_kwargs field
// for a problem that does not exist.
repNoThink, err := Score(ctx, "llm-only ("+model+", thinking off) [diagnostic]",
RouterFunc(func(ctx context.Context, u string, now time.Time) (router.Decision, error) {
d, ok, err := router.NewLLMRouter(&noThinkCompleter{base: base, http: &http.Client{Timeout: 60 * time.Second}}).Route(ctx, u, now)
if err != nil {
return d, err
}
if !ok {
return d, fmt.Errorf("llm router declined without an error")
}
return d, nil
}), f)
if err != nil {
t.Fatalf("Score no-thinking: %v", err)
}
t.Log("\n" + repNoThink.String() + repNoThink.Failures())
// Reports rather than asserts — the numbers are inputs to the #320
// decision, and an assertion here would be this test inventing the bar.
// The one thing worth failing on is a harness fault: if every single case
// errors, the run measured infrastructure, not routing, and the report
// must not be mistaken for a score.
for _, rep := range []Report{repLLM, repCascade, repNoThink} {
for _, rep := range []Report{repLLM, repCascade} {
if rep.Errors == rep.Total {
t.Errorf("%s: all %d cases errored — harness fault, not a measurement", rep.Name, rep.Total)
}
}
}
// noThinkCompleter — llm.Client with chat_template_kwargs.enable_thinking
// false. A test-local copy rather than a change to internal/llm: whether the
// daemon should send it is the open question, and answering it here by adding
// the field would prejudge #320.
type noThinkCompleter struct {
base string
http *http.Client
}
func (c *noThinkCompleter) Complete(ctx context.Context, r llm.Req) (string, error) {
payload := map[string]any{
"messages": []map[string]string{
{"role": "system", "content": r.System},
{"role": "user", "content": r.User},
},
"max_tokens": r.MaxTokens,
"temperature": 0,
"grammar": r.Grammar,
"chat_template_kwargs": map[string]any{"enable_thinking": false},
}
b, err := json.Marshal(payload)
if err != nil {
return "", err
}
req, err := http.NewRequestWithContext(ctx, "POST", c.base+"/v1/chat/completions", bytes.NewReader(b))
if err != nil {
return "", err
}
req.Header.Set("Content-Type", "application/json")
resp, err := c.http.Do(req)
if err != nil {
return "", err
}
defer resp.Body.Close()
if resp.StatusCode != 200 {
return "", fmt.Errorf("status %d", resp.StatusCode)
}
var out struct {
Choices []struct {
Message struct {
Content string `json:"content"`
ReasoningContent string `json:"reasoning_content"`
} `json:"message"`
} `json:"choices"`
}
if err := json.NewDecoder(resp.Body).Decode(&out); err != nil {
return "", err
}
if len(out.Choices) == 0 {
return "", fmt.Errorf("no choices")
}
m := out.Choices[0].Message
if m.Content != "" {
return m.Content, nil
}
return m.ReasoningContent, nil
}
func ping(ctx context.Context, c *llm.Client) error {
ctx, cancel := context.WithTimeout(ctx, 90*time.Second)
defer cancel()
+74
View File
@@ -0,0 +1,74 @@
package store
import (
"context"
"fmt"
"time"
)
// DialogueSessionRow — one saved follow-up session. Data is the session
// encoded by the dialogue package; the store does not look inside it.
type DialogueSessionRow struct {
ID string
Data []byte
Ts time.Time
TTL time.Duration
Expires time.Time
}
// SaveDialogueSession — write (or replace) the session for one dialogue id.
// One row per id: a newer turn overwrites the older state.
func (s *Store) SaveDialogueSession(ctx context.Context, id string, data []byte, ts time.Time, ttl time.Duration) error {
expires := ts.Add(ttl)
_, err := s.db.ExecContext(ctx, `
INSERT INTO dialogue_sessions (id, data, ts, ttl_ms, expires_ts) VALUES (?, ?, ?, ?, ?)
ON CONFLICT(id) DO UPDATE SET data = excluded.data,
ts = excluded.ts,
ttl_ms = excluded.ttl_ms,
expires_ts = excluded.expires_ts`,
id, data, ts.UnixMilli(), ttl.Milliseconds(), expires.UnixMilli())
if err != nil {
return fmt.Errorf("save dialogue session: %w", err)
}
return nil
}
// DeleteDialogueSession — drop one session (ended, or expired).
func (s *Store) DeleteDialogueSession(ctx context.Context, id string) error {
if _, err := s.db.ExecContext(ctx, `DELETE FROM dialogue_sessions WHERE id = ?`, id); err != nil {
return fmt.Errorf("delete dialogue session: %w", err)
}
return nil
}
// LoadDialogueSessions — return the sessions still alive at `now` and delete
// the ones that already ran out. An expired session is dead: it never comes
// back after a restart.
func (s *Store) LoadDialogueSessions(ctx context.Context, now time.Time) ([]DialogueSessionRow, error) {
if _, err := s.db.ExecContext(ctx,
`DELETE FROM dialogue_sessions WHERE expires_ts <= ?`, now.UnixMilli()); err != nil {
return nil, fmt.Errorf("prune dialogue sessions: %w", err)
}
rows, err := s.db.QueryContext(ctx,
`SELECT id, data, ts, ttl_ms, expires_ts FROM dialogue_sessions ORDER BY id`)
if err != nil {
return nil, fmt.Errorf("load dialogue sessions: %w", err)
}
defer rows.Close()
var out []DialogueSessionRow
for rows.Next() {
var r DialogueSessionRow
var tsMilli, ttlMilli, expMilli int64
if err := rows.Scan(&r.ID, &r.Data, &tsMilli, &ttlMilli, &expMilli); err != nil {
return nil, fmt.Errorf("scan dialogue session: %w", err)
}
r.Ts = time.UnixMilli(tsMilli).UTC()
r.TTL = time.Duration(ttlMilli) * time.Millisecond
r.Expires = time.UnixMilli(expMilli).UTC()
out = append(out, r)
}
if err := rows.Err(); err != nil {
return nil, fmt.Errorf("load dialogue sessions: %w", err)
}
return out, nil
}
+9
View File
@@ -74,6 +74,15 @@ ALTER TABLE reminders ADD COLUMN next_fire_ts INTEGER;`, // #2
`CREATE INDEX IF NOT EXISTS idx_nudges_snoozed ON nudges (outcome_ts) WHERE outcome = 'snoozed';`, // #8 — SnoozedUntil runs every tick; keep it off a full scan (Vikunja #364)
`ALTER TABLE proposed_routines ADD COLUMN accepted_ts INTEGER;
ALTER TABLE proposed_routines ADD COLUMN last_fired_ts INTEGER;`, // #9 — accepted routines keep firing (Vikunja #366): the tick loop needs to know when a routine was accepted and when it last nudged
`CREATE TABLE IF NOT EXISTS dialogue_sessions (
id TEXT PRIMARY KEY,
data BLOB NOT NULL,
ts INTEGER NOT NULL,
ttl_ms INTEGER NOT NULL,
expires_ts INTEGER NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_dialogue_sessions_expires ON dialogue_sessions (expires_ts);`, // #10 — the follow-up session survives a restart (Vikunja #363); small, TTL-pruned table, not a history log
}
// migrate applies every migration with a number greater than the DB's current