Compare commits

...

16 Commits

Author SHA1 Message Date
kami 9a3bcd7c46 Merge the thinking-off measurement 2026-07-31 13:31:31 +04:00
kami 04c1088088 Measure thinking off on routing properly — it does not win (#376)
The 67.1% "thinking off" column in ROUTING-EVAL-31-07-2026.md was an
artefact. It came from a hand-rolled HTTP client in the eval test that
did not send repeat_penalty, so it differed from the reference run on two
axes and the penalty was the one that mattered.

Re-scored back to back on an idle box with everything else held equal:
thinking off is identical to thinking on, case for case, same confusion
matrix, same three unparseable replies. A direct probe of the running
llama-server shows enable_thinking, thinking and reasoning_budget are all
ignored for this model on this build, so there was nothing to turn off.

No defaults changed. The misleading third configuration is removed from
internal/router/eval so its table cannot be quoted again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:30:22 +04:00
kami 07c191d8b8 Merge dialogue session persistence 2026-07-31 13:21:03 +04:00
kami c668310b3e Persist the dialogue session so a restart keeps the conversation
Vikunja #363. The follow-up session was a plain in-memory map, so any
mavend restart dropped the thread. It now mirrors to a small TTL-pruned
sqlite table and is loaded on startup; expired sessions are deleted on
load, not revived. Clarify's pending question is untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:20:30 +04:00
kami 1bd2acdc2a Do not exempt Russian words that are both noun and verb 2026-07-31 12:55:45 +04:00
kami 15e5dd8eaa Merge the second-person gender check 2026-07-31 12:54:48 +04:00
kami 10cf6f525c Check that nudges do not address the owner in the feminine
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:54:18 +04:00
kami e2210f6844 Merge the clarify-expiry notice 2026-07-31 12:52:36 +04:00
kami 214a4032cf Tell him when an expired clarify question is dropped
Vikunja #382. A parked clarifying question past its TTL was discarded
silently on read; now she says the old request is gone and the newly
spoken words are still routed as a fresh utterance.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:51:47 +04:00
kami dc70a5a7ab Show clarify_max_attempts in the deployed config
The default is 3 either way. Writing it out means you can see the knob
without reading the Go.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:34:28 +04:00
kami 06aded6ab0 Merge commit 'd2be98e' into overnight-jul31
# Conflicts:
#	cmd/mavend/clarify.go
#	cmd/mavend/clarify_test.go
#	cmd/mavend/voice.go
#	internal/config/config.go
2026-07-31 12:34:00 +04:00
kami d2be98ee2a Say out loud when she gives up instead of dropping the request
An unclear answer used to end the request on the spot. Now she re-asks the same
question while attempts remain, and when they run out she says
"Прости, я не поняла. Скажи, пожалуйста, по-другому." — silence would leave him
thinking it was handled. Same reply when the missing slot has no question to
ask, and as a floor in finishClarified so an empty reply can never ship.

Tests: three questions allowed, the fourth gives up out loud, the cap is
configurable, and a restated time is the one that lands.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:31:55 +04:00
kami 62d320f93a Let her ask three times, and let a restated answer win
MaxAttempts was 1, justified as "not a nag". Wrong reading: "not a nag" is about
interrupting unprompted, and a clarifying question is part of a conversation he
started. Now three, configurable via voice.clarify_max_attempts (default 3).
Three, because after that the likely problem is she misheard the whole request,
not one slot.

Answer used to keep the parked value, so "в три" then "нет, в пять" threw the
five away. Now a value the answer carries wins for the slot she asked about.
Only for the clarify answer — a correction in a fresh turn is followUpMerge.

The eight-field chained assertion in the Answer test is one DeepEqual now, so a
new field in Slots is covered without touching the test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:31:46 +04:00
kami 0b3b8d0a9e Test the clarify round-trip end to end at the daemon level
Covers: a reminder with no time is asked about and completes on the answer; the
same for a fact; an answer past the TTL falls through as a fresh utterance; a
second unclear answer drops the request with no second question; a clarified act
off the allowlist neither runs nor gets enabled; a clarified destructive act
still parks a confirm; noise keeps the canned reply. No model, no network.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:31:07 +04:00
kami fe0e654ab1 Ask the question, then act on the answer
On a clarify decision with one identifiable gap she now asks instead of saying
"не поняла", and parks the request. The next utterance is parsed as the answer
with the router's own extractor and the completed decision runs through
applyAction like any other — so a clarified act still needs the allowlist and
still hits the destructive confirm gate. An answer that does not fill the gap
drops the request; she never asks twice. Also pulls the session-store block
that HandlePushToTalk and handleText both had into rememberTurn, since the
clarify path needed a third copy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:31:07 +04:00
kami a54ebac0cb Work out which slot is missing and phrase one short question
A table per intent (reminder needs a time, fact needs a key, act needs a fn)
plus one fixed Russian question per slot. Templates, not model output: a 0.8B
would wander and a question that rewords itself is harder to answer. Note,
query, chat and system get no question — for those a clarify decision keeps
the canned reply rather than inventing a question for noise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:31:07 +04:00
17 changed files with 836 additions and 224 deletions
+9 -7
View File
@@ -105,11 +105,14 @@ that never says food.
## Broken, found, not fixed ## Broken, found, not fixed
1. **`checkFeminine` only catches half the constraint.** It scans for masculine 1. ~~**`checkFeminine` only catches half the constraint.**~~ **Fixed** (#381). It scanned for
self-reference and passed 15/15 both runs — but three messages address the *owner* in masculine self-reference only, so three messages that addressed the *owner* in the feminine
the feminine: "ты давно не отдыхал**а**", "он не ел". The owner is a man. The check has ("ты давно не отдыхал**а**") scored clean. There is now a second check, `hisgender`: a
no second-person gender test, so this scores clean while being exactly the persona feminine past-tense verb (-ла/-лась) in a sentence addressed to him ("ты", "тебе", "твой")
failure the constraint exists to prevent. This is the most important gap in the harness. fails, unless the verb is hers ("я заметила", "напомнила тебе"). It is a suffix rule, not a
parser — see the comment in `checks.go` for what it misses. A fresh 15-case run after adding
it scored **12/15** with `hisgender` 15/15; the model did not repeat the feminine address in
that sample, and the check is pinned by unit tests on the recorded bad strings instead.
2. **Grammar is not checked at all, and it is bad.** `"Он не ел 11 дней"` (it was 11 hours), 2. **Grammar is not checked at all, and it is bad.** `"Он не ел 11 дней"` (it was 11 hours),
`"Сонуждились 7 дней"` (not a word), `"Они забыли воду"` (wrong person entirely). Every `"Сонуждились 7 дней"` (not a word), `"Они забыли воду"` (wrong person entirely). Every
one of these passes all six checks. The fixture measures properties, not fluency, and at one of these passes all six checks. The fixture measures properties, not fluency, and at
@@ -125,8 +128,7 @@ that never says food.
## Next steps ## Next steps
1. **Add a second-person gender check** to `checks.go`. Finding 1 above. Until it exists the 1. ~~**Add a second-person gender check**~~ — done, `hisgender` in `checks.go` (#381).
feminine column means less than it looks like.
2. **Decide whether the fallback should count as a pass.** Right now `Score` cannot tell a 2. **Decide whether the fallback should count as a pass.** Right now `Score` cannot tell a
model answer from a fallback. Either mark fallback bodies in `PhrasedNudge` or count them model answer from a fallback. Either mark fallback bodies in `PhrasedNudge` or count them
in their own column. Without that, any future prompt change can score well by failing in their own column. Without that, any future prompt change can score well by failing
+59 -10
View File
@@ -66,15 +66,64 @@ Three things this run settles:
`запиши что…` phrasings toward fact, and that suspicion stands — all five `ru-note-*` `запиши что…` phrasings toward fact, and that suspicion stands — all five `ru-note-*`
cases now land on fact. Tracked as Vikunja #375. cases now land on fact. Tracked as Vikunja #375.
**Thinking off is the best configuration measured so far**, on both accuracy and latency The `thinking off` column above read as the best configuration measured so far (Vikunja #376).
(Vikunja #376). That is worth understanding before flipping: routing is a short **It was wrong** — see the controlled re-run below. Ignore that column.
classification into a fixed enum with grammar-constrained output, so there is little to
reason about, and the thinking trace mostly gives a small model room to talk itself out of
the right answer. Phrasing is a different job and needs measuring separately.
Still `6 / 6` missed clarify — the router has no way to say "I don't know" (Vikunja #359). Still `6 / 6` missed clarify — the router has no way to say "I don't know" (Vikunja #359).
That is unchanged by anything here. That is unchanged by anything here.
## Thinking off — 31-07-2026, controlled re-run (Vikunja #376)
The "thinking off wins by 6 points" observation above **does not hold**. It was a measurement
artefact, and the earlier table's `thinking off` column should be ignored.
The thinking-off variant was scored by a hand-rolled HTTP client living in the test file
instead of `llm.Client`. That copy did not send `repeat_penalty`, which the real router does
send (`routeRepeatPenalty = 1.15`). So the two columns differed on two axes at once, and the
one that mattered was the penalty, not the thinking mode.
Re-measured with everything else held equal — same fixture, same prompt, same grammar, same
sampling, same idle box, the three configurations run back to back and never concurrently:
| | llm-only, thinking on | llm-only, thinking off | cascade+llm |
|---|---|---|---|
| intent-only accuracy | 59.2% (45/76) | 59.2% (45/76) | 61.8% (47/76) |
| full accuracy (intent+slots+gate) | 38.2% (29/76) | 38.2% (29/76) | 57.9% (44/76) |
| route errors | 3 | 3 | 0 |
| grammar violations | 3 (all 3 route errors) | 3 (same 3 cases) | 0 |
| missed clarify | 5 / 6 | 5 / 6 | 5 / 6 |
| p50 latency | 836ms | 920ms | 810ms |
| p95 latency | 1.41s | 2.00s | 1.31s |
Thinking off is not just a tie on the headline numbers — it is identical case for case, with
the same confusion matrix and the same three unparseable replies. The latency difference is
run-to-run noise on one box, and it points the wrong way here.
The reason is simpler than any accuracy argument: **this llama-server build ignores the
request-level thinking switch for this model.** Probed directly against the running server
with `chat_template_kwargs.enable_thinking = false`, `chat_template_kwargs.thinking = false`
and top-level `reasoning_budget = 0` — all three return a byte-identical answer with the
thinking trace still in `reasoning_content`, and the server reports the prompt prefix as
cached, meaning the rendered template did not change. There was never anything being turned
off, which is also why the numbers match exactly.
Nothing was defaulted. `internal/llm` still has no `chat_template_kwargs` field, `VoiceConfig`
has no thinking flag, and `deploy/mavend.json` is unchanged. The misleading third
configuration is removed from `internal/router/eval` so the table it produced cannot be quoted
again.
Two caveats worth saying out loud:
- **The fixture is 76 cases.** A 6-point difference on 76 cases is roughly 4-5 cases and would
not have been worth trusting even if it had reproduced. This one was exactly 0 cases, which
is a much easier call.
- **This is one server build and one checkpoint** (`b9351`, Qwen3.5-0.8B Q4_K_M). If the
#122 checkpoint or a newer llama.cpp does honour the switch, the question reopens — but it
reopens as an unmeasured question, not as a 6-point win.
Phrasing was **not** measured. Whether thinking helps there is still open, and now also blocked
on the same "can we even turn it off" question.
## Findings ## Findings
### 1. The resident model does route better — 50.0% vs 36.8% ### 1. The resident model does route better — 50.0% vs 36.8%
@@ -145,11 +194,11 @@ Note the grammar's `string ::= "\"" ([^"\\] | "\\" .)* "\""` is unbounded, so no
### 7. Two hypotheses tested and closed ### 7. Two hypotheses tested and closed
- **Thinking mode is a non-issue.** Qwen3.5's template defaults `thinking = 1`, so - **Thinking mode is a non-issue.** Confirmed twice now, the second time properly — see the
grammar-constrained JSON lands in `reasoning_content` with `content` empty — controlled re-run section. Grammar-constrained JSON lands in `reasoning_content` with
`llm.Client`'s fallback handles it. A `thinking off` run scored *identically* (18/76, `content` empty and `llm.Client`'s fallback handles it; the request-level switch does
48.7%, same p50). `internal/llm` deliberately does **not** grow a `chat_template_kwargs` nothing on this build. `internal/llm` deliberately does **not** grow a
field. `chat_template_kwargs` field.
- **Runaway array repetition does not reproduce.** An isolated smoke test with a stripped - **Runaway array repetition does not reproduce.** An isolated smoke test with a stripped
grammar emitted `{"intent":"reminder"}` until `MaxTokens`; under the real `routeSystem` grammar emitted `{"intent":"reminder"}` until `MaxTokens`; under the real `routeSystem`
prompt the few-shot examples anchor it to one object. 2 errors in 76, not 76. prompt the few-shot examples anchor it to one object. 2 errors in 76, not 76.
+78 -17
View File
@@ -40,9 +40,44 @@ var clarifyQuestions = map[dialogue.Slot]string{
dialogue.SlotFn: "Что сделать?", dialogue.SlotFn: "Что сделать?",
} }
// clarifyDropped — she asked once, the answer still did not fill the gap, so // clarifyGaveUp — she is out of questions and still does not have the slot. She
// the request is gone. Said plainly, once, with no second question. // says so out loud: dropping the request in silence would leave him thinking it
const clarifyDropped = "Не разобрала — скажи целиком, пожалуйста." // landed. Feminine self-reference ("поняла"), as everywhere.
const clarifyGaveUp = "Прости, я не поняла. Скажи, пожалуйста, по-другому."
// clarifyExpired — his answer came after the TTL, so the parked request is
// already gone. Same tone as clarifyGaveUp, different reason: too much time
// passed, not "I did not understand". Feminine self-reference ("ждала",
// "отпустила"); he is addressed with a plain imperative.
const clarifyExpired = "Прости, я слишком долго ждала ответа и отпустила прошлую просьбу. Если она ещё нужна, скажи заново."
// clarifyExpiredNotice returns that line when a parked question had just timed
// out, and "" when nothing was parked. Call it right after
// resolveClarifyAnswer: a live question is answered there, an expired one is
// only reported here — the words themselves still go on to be routed fresh.
func (h *reactiveHandler) clarifyExpiredNotice() string {
if h.clarifyStore == nil {
return ""
}
if !h.clarifyStore.TakeExpired(voiceDialogueID, h.now()) {
return ""
}
log.Printf("voice: clarify — parked question expired, telling him and routing the words fresh")
return clarifyExpired
}
// withNotice glues the expiry notice in front of this turn's reply. One turn
// carries one reply on the wire, so the notice cannot be a message of its own —
// but neither the notice nor the fresh answer may be dropped.
func withNotice(notice, reply string) string {
if notice == "" {
return reply
}
if reply == "" {
return notice
}
return notice + " " + reply
}
// missingFor returns the slots a decision still needs, most important first. // missingFor returns the slots a decision still needs, most important first.
// Empty ⇒ there is nothing identifiable to ask about. // Empty ⇒ there is nothing identifiable to ask about.
@@ -79,13 +114,14 @@ func (h *reactiveHandler) askClarify(dec router.Decision) (string, bool) {
return "", false return "", false
} }
h.clarifyStore.Put(voiceDialogueID, &dialogue.PendingQuestion{ h.clarifyStore.Put(voiceDialogueID, &dialogue.PendingQuestion{
Intent: dialogue.Intent(dec.Intent), Intent: dialogue.Intent(dec.Intent),
Slots: toDialogueSlots(dec.Slots), Slots: toDialogueSlots(dec.Slots),
Missing: []dialogue.Slot{slot}, Missing: []dialogue.Slot{slot},
Utterance: dec.Utterance, Utterance: dec.Utterance,
Asked: h.now(), Asked: h.now(),
TTL: clarifyTTL, TTL: clarifyTTL,
Attempts: 1, // asked once; MaxAttempts is 1, so there is no second ask Attempts: 1, // this ask
MaxAttempts: h.clarifyMaxAttempts,
}) })
log.Printf("voice: clarify — asked about %s for intent=%s", slot, dec.Intent) log.Printf("voice: clarify — asked about %s for intent=%s", slot, dec.Intent)
return question, true return question, true
@@ -97,8 +133,9 @@ func (h *reactiveHandler) askClarify(dec router.Decision) (string, bool) {
// resolveConfirm and checked in the same place. // resolveConfirm and checked in the same place.
// //
// The answer is parsed with the same extractor the router uses, for the intent // The answer is parsed with the same extractor the router uses, for the intent
// she parked — no second parser. If it still does not fill the gap the request // she parked — no second parser. If it still does not fill the gap she asks
// is dropped: she does not ask again. // again, up to MaxAttempts; after that she says out loud that she did not
// understand. She never drops the request in silence.
func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string) (string, bool) { func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string) (string, bool) {
if h.clarifyStore == nil { if h.clarifyStore == nil {
return "", false return "", false
@@ -107,17 +144,14 @@ func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string)
if q == nil { if q == nil {
return "", false return "", false
} }
// One shot either way: the question is consumed whether or not the answer
// works, so a failed answer can't leave the question armed.
h.clarifyStore.Delete(voiceDialogueID)
intent := router.Intent(q.Intent) intent := router.Intent(q.Intent)
answer := h.extractor.Extract(ctx, intent, text, h.now()) answer := h.extractor.Extract(ctx, intent, text, h.now())
merged := q.Answer(text, toDialogueSlots(answer)) merged := q.Answer(text, toDialogueSlots(answer))
if len(dialogue.StillMissing(q.Missing, merged)) > 0 { if len(dialogue.StillMissing(q.Missing, merged)) > 0 {
log.Printf("voice: clarify — answer %q did not fill %v, dropping", text, q.Missing) return h.reaskOrGiveUp(q, merged, text), true
return clarifyDropped, true
} }
h.clarifyStore.Delete(voiceDialogueID)
// Rebuild the decision as if it had routed cleanly, then run it down the // Rebuild the decision as if it had routed cleanly, then run it down the
// normal path. Clarify is deliberately false and the intent is unchanged: // normal path. Clarify is deliberately false and the intent is unchanged:
@@ -133,6 +167,29 @@ func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string)
return h.finishClarified(ctx, dec), true return h.finishClarified(ctx, dec), true
} }
// reaskOrGiveUp handles an answer that left the gap open: ask the same question
// again while she has attempts left, otherwise say she did not understand and
// let the request go. Never returns "" — a mute give-up reads as "done".
func (h *reactiveHandler) reaskOrGiveUp(q *dialogue.PendingQuestion, merged dialogue.Slots, text string) string {
question := ""
if len(q.Missing) > 0 {
question = clarifyQuestions[q.Missing[0]]
}
if question == "" || !q.CanAsk() {
h.clarifyStore.Delete(voiceDialogueID)
log.Printf("voice: clarify — gave up on %v after %d question(s), answer was %q", q.Missing, q.Attempts, text)
return clarifyGaveUp
}
// Re-park with whatever the answer DID give, the clock restarted and one
// more question spent.
q.Slots = merged
q.Attempts++
q.Asked = h.now()
h.clarifyStore.Put(voiceDialogueID, q)
log.Printf("voice: clarify — answer %q did not fill %v, asking again (attempt %d)", text, q.Missing, q.Attempts)
return question
}
// finishClarified runs a completed decision through the same steps a freshly // finishClarified runs a completed decision through the same steps a freshly
// routed one takes: remember the turn, act, then phrase. // routed one takes: remember the turn, act, then phrase.
func (h *reactiveHandler) finishClarified(ctx context.Context, dec router.Decision) string { func (h *reactiveHandler) finishClarified(ctx context.Context, dec router.Decision) string {
@@ -146,6 +203,10 @@ func (h *reactiveHandler) finishClarified(ctx context.Context, dec router.Decisi
if reply == "" { if reply == "" {
reply = h.replier.Reply(dec) reply = h.replier.Reply(dec)
} }
if reply == "" {
// Belt: an empty reply here would be a silent drop.
reply = clarifyGaveUp
}
return reply return reply
} }
+102 -11
View File
@@ -86,7 +86,7 @@ func TestClarifyReminderCompletesOnAnswer(t *testing.T) {
if !handled { if !handled {
t.Fatal("the answer to an open question must be consumed as an answer") t.Fatal("the answer to an open question must be consumed as an answer")
} }
if reply == clarifyDropped { if reply == clarifyGaveUp {
t.Fatalf("a good answer must not drop the request: %q", reply) t.Fatalf("a good answer must not drop the request: %q", reply)
} }
@@ -111,7 +111,7 @@ func TestClarifyFactCompletesOnAnswer(t *testing.T) {
if _, asked := h.askClarify(clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши")); !asked { if _, asked := h.askClarify(clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши")); !asked {
t.Fatal("a fact with no key should be asked about") t.Fatal("a fact with no key should be asked about")
} }
if reply, handled := h.resolveClarifyAnswer(ctx, "пил воду"); !handled || reply == clarifyDropped { if reply, handled := h.resolveClarifyAnswer(ctx, "пил воду"); !handled || reply == clarifyGaveUp {
t.Fatalf("answer should complete the fact, handled=%v reply=%q", handled, reply) t.Fatalf("answer should complete the fact, handled=%v reply=%q", handled, reply)
} }
if fact, err := st.LatestFact(ctx, "water"); err != nil || fact.Key != "water" { if fact, err := st.LatestFact(ctx, "water"); err != nil || fact.Key != "water" {
@@ -137,26 +137,87 @@ func TestClarifyAnswerAfterTTLIsANewRequest(t *testing.T) {
} }
} }
// TestClarifyUnclearAnswerDropsWithoutAskingAgain — MaxAttempts is 1. // TestClarifyAsksThreeTimesThenSaysSo — three questions are allowed, the fourth
func TestClarifyUnclearAnswerDropsWithoutAskingAgain(t *testing.T) { // is not, and running out is SPOKEN. Silence would read as "handled".
func TestClarifyAsksThreeTimesThenSaysSo(t *testing.T) {
ctx := context.Background() ctx := context.Background()
h, st, _ := newClarifyHandler(t) h, st, _ := newClarifyHandler(t)
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked { if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a question") t.Fatal("expected a first question")
} }
// Two more unclear answers ⇒ two more questions (3 asks in total).
for i := 2; i <= 3; i++ {
reply, handled := h.resolveClarifyAnswer(ctx, "ну не знаю")
if !handled {
t.Fatalf("answer %d must be consumed as an answer", i)
}
if reply != "На когда напомнить?" {
t.Fatalf("attempt %d should ask again, got %q", i, reply)
}
if h.clarifyStore.Get(voiceDialogueID, h.now()) == nil {
t.Fatalf("attempt %d must leave the question armed", i)
}
}
reply, handled := h.resolveClarifyAnswer(ctx, "ну не знаю") reply, handled := h.resolveClarifyAnswer(ctx, "ну не знаю")
if !handled || reply != clarifyDropped { if !handled || reply != clarifyGaveUp {
t.Fatalf("an unclear answer should drop the request, handled=%v reply=%q", handled, reply) t.Fatalf("the fourth try must give up out loud, handled=%v reply=%q", handled, reply)
} }
if strings.Contains(reply, "?") { if reply == "" || strings.Contains(reply, "?") {
t.Fatalf("she must not ask a second question: %q", reply) t.Fatalf("giving up must be spoken and must not be another question: %q", reply)
} }
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil { if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
t.Fatal("a dropped request must leave no armed question") t.Fatal("a given-up request must leave no armed question")
} }
if reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour)); err != nil || len(reminders) != 0 { if reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour)); err != nil || len(reminders) != 0 {
t.Fatalf("a dropped request must not create anything: reminders=%v err=%v", reminders, err) t.Fatalf("a given-up request must not create anything: reminders=%v err=%v", reminders, err)
}
}
// TestClarifyMaxAttemptsIsConfigurable — one question when the config says one.
func TestClarifyMaxAttemptsIsConfigurable(t *testing.T) {
ctx := context.Background()
h, _, _ := newClarifyHandler(t)
h.clarifyMaxAttempts = 1
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a question")
}
if reply, handled := h.resolveClarifyAnswer(ctx, "ну не знаю"); !handled || reply != clarifyGaveUp {
t.Fatalf("with max 1 she must give up at once, handled=%v reply=%q", handled, reply)
}
}
// TestClarifyRestatedAnswerWins — «в 11:00», then «нет, в 15:00». The second
// value is the one that lands.
func TestClarifyRestatedAnswerWins(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме")); !asked {
t.Fatal("expected a question")
}
// First answer parses, but re-park it by hand as if she had asked again:
// what matters here is that Answer prefers the newer value over the parked
// one, which is the case the daemon hits on a re-ask.
q := h.clarifyStore.Get(voiceDialogueID, h.now())
if q == nil {
t.Fatal("expected an armed question")
}
first := h.extractor.Extract(ctx, router.IntentReminder, "в 11:00", h.now())
q.Slots = q.Answer("в 11:00", toDialogueSlots(first))
if reply, handled := h.resolveClarifyAnswer(ctx, "нет, в 15:00"); !handled || reply == clarifyGaveUp {
t.Fatalf("the restated answer should complete the request, handled=%v reply=%q", handled, reply)
}
reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour))
if err != nil || len(reminders) != 1 {
t.Fatalf("expected one reminder: %v err=%v", reminders, err)
}
want := h.extractor.Extract(ctx, router.IntentReminder, "в 15:00", h.now())
if !reminders[0].FireTs.Equal(want.Time) {
t.Fatalf("reminder at %v, want the restated %v", reminders[0].FireTs, want.Time)
} }
} }
@@ -228,6 +289,36 @@ func TestNoQuestionWhenNothingIsMissing(t *testing.T) {
} }
} }
// TestClarifyExpiryIsAnnouncedAndWordsStillRoute — his answer lands after the
// TTL: she must say the old request is gone AND still answer the new words.
func TestClarifyExpiryIsAnnouncedAndWordsStillRoute(t *testing.T) {
ctx := context.Background()
h, _, now := newClarifyHandler(t)
emb := router.NewHashEmbedder(1024)
h.embedder = emb
h.router = buildRouter(emb, h.matcher, 0.55, nil)
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a question")
}
*now = now.Add(clarifyTTL + time.Second)
reply := h.handleText(ctx, "как дела")
if !strings.HasPrefix(reply, clarifyExpired) {
t.Fatalf("expired question must be announced first, got %q", reply)
}
if strings.TrimSpace(strings.TrimPrefix(reply, clarifyExpired)) == "" {
t.Fatalf("the new words must still be answered, got only the notice: %q", reply)
}
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
t.Fatal("the expired question must be gone")
}
// The notice is said once, not on every later utterance.
if reply := h.handleText(ctx, "как дела"); strings.Contains(reply, clarifyExpired) {
t.Fatalf("notice repeated on a later turn: %q", reply)
}
}
// TestNoPendingQuestionFallsThrough — with nothing parked, an utterance routes // TestNoPendingQuestionFallsThrough — with nothing parked, an utterance routes
// normally. // normally.
func TestNoPendingQuestionFallsThrough(t *testing.T) { func TestNoPendingQuestionFallsThrough(t *testing.T) {
+46 -19
View File
@@ -227,7 +227,18 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
} }
// ----- dialogue (multi-turn slot carry-over; 2-min follow-up window) ----- // ----- dialogue (multi-turn slot carry-over; 2-min follow-up window) -----
dialogueSessions := dialogue.NewSessionStore(2 * time.Minute) // Store-backed when the daemon passes a store, so a restart mid-conversation
// keeps the thread (Vikunja #363). Sessions past their TTL are dropped on
// load, never revived. Clarify's parked question stays in memory only.
var dialogueSessions *dialogue.SessionStore
if dataStore != nil {
dialogueSessions = dialogue.NewPersistentSessionStore(2*time.Minute, dataStore)
if err := dialogueSessions.Load(context.Background(), time.Now()); err != nil {
log.Printf("dialogue: load saved sessions: %v", err)
}
} else {
dialogueSessions = dialogue.NewSessionStore(2 * time.Minute)
}
clarifyStore := dialogue.NewClarifyStore(clarifyTTL) clarifyStore := dialogue.NewClarifyStore(clarifyTTL)
timeParser := router.NewPythonDateParser() timeParser := router.NewPythonDateParser()
@@ -255,11 +266,13 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
dataStore: dataStore, dataStore: dataStore,
dialogueSessions: dialogueSessions, dialogueSessions: dialogueSessions,
clarifyStore: clarifyStore, clarifyStore: clarifyStore,
extractor: router.Extractor{Time: timeParser, Acts: matcher, Facts: router.DefaultFactParser{}}, // 0 here (unset config) ⇒ the dialogue default.
queryMinScore: cfg.Voice.QueryMinScore, clarifyMaxAttempts: cfg.Voice.ClarifyMaxAttempts,
queryMinMargin: cfg.Voice.QueryMinMargin, extractor: router.Extractor{Time: timeParser, Acts: matcher, Facts: router.DefaultFactParser{}},
timeParser: timeParser, queryMinScore: cfg.Voice.QueryMinScore,
ecosystem: eco, queryMinMargin: cfg.Voice.QueryMinMargin,
timeParser: timeParser,
ecosystem: eco,
} }
// ----- the server (TCP listener) ----- // ----- the server (TCP listener) -----
@@ -318,6 +331,10 @@ type reactiveHandler struct {
// clarify.go). nil ⇒ she falls back to the canned "не поняла" reply. // clarify.go). nil ⇒ she falls back to the canned "не поняла" reply.
clarifyStore *dialogue.ClarifyStore clarifyStore *dialogue.ClarifyStore
// clarifyMaxAttempts — questions per request before she gives up out loud.
// 0 ⇒ dialogue.DefaultMaxAttempts (3). Set from VoiceConfig.
clarifyMaxAttempts int
// extractor parses the answer to an open question, with the same parsers // extractor parses the answer to an open question, with the same parsers
// the router's own stage-2 uses. // the router's own stage-2 uses.
extractor router.Extractor extractor router.Extractor
@@ -396,9 +413,17 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
return h.reply(ctx, reply, nil) return h.reply(ctx, reply, nil)
} }
// 1b2. clarify answer — if she asked a question last turn, this utterance is // 1b2. expired clarify — a question was parked but its TTL ran out, so the
// its answer, not a fresh command. After the confirm check: a y/n gate is // request behind it is gone. Say that out loud (see clarify.go) and carry
// armed by her own prompt and is the narrower claim on the utterance. // on: these words are still routed as a fresh utterance below, with the
// notice glued in front of whatever the fresh routing answers. Checked
// BEFORE the answer path: reading a parked question drops an expired one.
expiredNotice := h.clarifyExpiredNotice()
// 1b3. clarify answer — if she asked a live question last turn, this
// utterance is its answer, not a fresh command. After the confirm check: a
// y/n gate is armed by her own prompt and is the narrower claim on the
// utterance.
if reply, handled := h.resolveClarifyAnswer(ctx, text); handled { if reply, handled := h.resolveClarifyAnswer(ctx, text); handled {
return h.reply(ctx, reply, nil) return h.reply(ctx, reply, nil)
} }
@@ -408,7 +433,7 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
// unreliably (it's a command, not a free-form query), so we match it // unreliably (it's a command, not a free-form query), so we match it
// before routing. Same pattern as the confirm turn above. // before routing. Same pattern as the confirm turn above.
if reply, handled := h.resolveQuietToggle(ctx, text); handled { if reply, handled := h.resolveQuietToggle(ctx, text); handled {
return h.reply(ctx, reply, nil) return h.reply(ctx, withNotice(expiredNotice, reply), nil)
} }
// 2. router — classify the utterance. // 2. router — classify the utterance.
@@ -417,10 +442,10 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
// ErrNoIntents ⇒ classifier unseeded (cold boot). reply with a // ErrNoIntents ⇒ classifier unseeded (cold boot). reply with a
// "still warming up" rather than a wire error. // "still warming up" rather than a wire error.
if errors.Is(err, router.ErrNoIntents) { if errors.Is(err, router.ErrNoIntents) {
return h.reply(ctx, "я ещё не понимаю свободную речь — скоро научусь.", nil) return h.reply(ctx, withNotice(expiredNotice, "я ещё не понимаю свободную речь — скоро научусь."), nil)
} }
log.Printf("voice: router error: %v", err) log.Printf("voice: router error: %v", err)
return h.reply(ctx, "не получилось разобрать команду.", nil) return h.reply(ctx, withNotice(expiredNotice, "не получилось разобрать команду."), nil)
} }
// 2b. dialogue — fill this turn's missing slots from a prior same-intent // 2b. dialogue — fill this turn's missing slots from a prior same-intent
@@ -441,7 +466,7 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
// stands. // stands.
if dec.Clarify { if dec.Clarify {
if question, asked := h.askClarify(dec); asked { if question, asked := h.askClarify(dec); asked {
return h.reply(ctx, question, nil) return h.reply(ctx, withNotice(expiredNotice, question), nil)
} }
} }
@@ -457,7 +482,7 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
// 5. tts — synthesise the reply text; return to the voice server which // 5. tts — synthesise the reply text; return to the voice server which
// ships it back on the conn. // ships it back on the conn.
return h.reply(ctx, replyText, nil) return h.reply(ctx, withNotice(expiredNotice, replyText), nil)
} }
// handleText — the core reactive path without stt/tts: confirm check → // handleText — the core reactive path without stt/tts: confirm check →
@@ -472,7 +497,9 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
return reply return reply
} }
// 1b2. clarify answer — same check as HandlePushToTalk. // 1b2/1b3. expired clarify then clarify answer — same order and reasons as
// HandlePushToTalk.
expiredNotice := h.clarifyExpiredNotice()
if reply, handled := h.resolveClarifyAnswer(ctx, text); handled { if reply, handled := h.resolveClarifyAnswer(ctx, text); handled {
return reply return reply
} }
@@ -481,10 +508,10 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
dec, err := h.router.Route(ctx, text, h.now()) dec, err := h.router.Route(ctx, text, h.now())
if err != nil { if err != nil {
if errors.Is(err, router.ErrNoIntents) { if errors.Is(err, router.ErrNoIntents) {
return "я ещё не понимаю свободную речь — скоро научусь." return withNotice(expiredNotice, "я ещё не понимаю свободную речь — скоро научусь.")
} }
log.Printf("voice: handleText router error: %v", err) log.Printf("voice: handleText router error: %v", err)
return "не получилось разобрать команду." return withNotice(expiredNotice, "не получилось разобрать команду.")
} }
log.Printf("voice: route result: intent=%s slots=%+v", dec.Intent, dec.Slots) log.Printf("voice: route result: intent=%s slots=%+v", dec.Intent, dec.Slots)
@@ -501,7 +528,7 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
// 2c. clarify — same as HandlePushToTalk: ask about the one missing thing. // 2c. clarify — same as HandlePushToTalk: ask about the one missing thing.
if dec.Clarify { if dec.Clarify {
if question, asked := h.askClarify(dec); asked { if question, asked := h.askClarify(dec); asked {
return question return withNotice(expiredNotice, question)
} }
} }
@@ -513,7 +540,7 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
if replyText == "" { if replyText == "" {
replyText = h.replier.Reply(dec) replyText = h.replier.Reply(dec)
} }
return replyText return withNotice(expiredNotice, replyText)
} }
// applyAction — executes the router's Decision. Intent-by-intent: // applyAction — executes the router's Decision. Intent-by-intent:
+1
View File
@@ -43,6 +43,7 @@
"llm_router": true, "llm_router": true,
"query_min_score": 0.55, "query_min_score": 0.55,
"query_min_margin": 0.008, "query_min_margin": 0.008,
"clarify_max_attempts": 3,
"tool_timeout": "30s", "tool_timeout": "30s",
"tools": [ "tools": [
{ "name": "status", "cmd": ["systemctl", "status"], "scope": "homelab", "destructive": false }, { "name": "status", "cmd": ["systemctl", "status"], "scope": "homelab", "destructive": false },
+10 -1
View File
@@ -292,6 +292,10 @@ type VoiceConfig struct {
// Negative ⇒ off. 0 ⇒ the default below. // Negative ⇒ off. 0 ⇒ the default below.
QueryMinMargin float64 `json:"query_min_margin,omitempty"` QueryMinMargin float64 `json:"query_min_margin,omitempty"`
// ClarifyMaxAttempts — how many clarifying questions she may ask about one
// request before she gives up and says she did not understand. Default 3.
ClarifyMaxAttempts int `json:"clarify_max_attempts,omitempty"`
// Persona — optional prompt prefix that tunes maven's character. Prepended // Persona — optional prompt prefix that tunes maven's character. Prepended
// to every LLM system prompt (nudge phrasing, note queries, general // to every LLM system prompt (nudge phrasing, note queries, general
// knowledge). Empty string ⇒ current hardcoded persona (feminine-gendered // knowledge). Empty string ⇒ current hardcoded persona (feminine-gendered
@@ -424,7 +428,9 @@ const (
// false recall from 5/5 to 1/5. Every larger delta costs real recall // false recall from 5/5 to 1/5. Every larger delta costs real recall
// without removing that last one until 0.020, which drops recall to 44%. // without removing that last one until 0.020, which drops recall to 44%.
DefaultQueryMinMargin = 0.008 DefaultQueryMinMargin = 0.008
DefaultToolTimeout = 30 * time.Second // DefaultClarifyMaxAttempts — see dialogue.DefaultMaxAttempts.
DefaultClarifyMaxAttempts = 3
DefaultToolTimeout = 30 * time.Second
// DefaultLLMRouter — route with the resident model unless told otherwise. // DefaultLLMRouter — route with the resident model unless told otherwise.
DefaultLLMRouter = true DefaultLLMRouter = true
@@ -516,6 +522,9 @@ func (c *Config) applyDefaults() {
case c.Voice.QueryMinMargin < 0: case c.Voice.QueryMinMargin < 0:
c.Voice.QueryMinMargin = 0 c.Voice.QueryMinMargin = 0
} }
if c.Voice.ClarifyMaxAttempts <= 0 {
c.Voice.ClarifyMaxAttempts = DefaultClarifyMaxAttempts
}
if c.Voice.ToolTimeout <= 0 { if c.Voice.ToolTimeout <= 0 {
c.Voice.ToolTimeout = Duration(DefaultToolTimeout) c.Voice.ToolTimeout = Duration(DefaultToolTimeout)
} }
+53 -31
View File
@@ -17,10 +17,11 @@ const (
SlotText Slot = "text" // Slots.Text SlotText Slot = "text" // Slots.Text
) )
// MaxAttempts is 1 because Maven is not a nag (DESIGN.md § Non-goals). She asks // DefaultMaxAttempts — how many questions she may ask about one request.
// one clarifying question. If the answer still leaves the slot empty she drops // Three, because after three tries the likely problem is that she misheard the
// the request instead of asking again. // whole request, not one slot — so another question about that slot won't help.
const MaxAttempts = 1 // Configurable: voice.clarify_max_attempts.
const DefaultMaxAttempts = 3
// PendingQuestion is what Maven holds while she waits for an answer to an open // PendingQuestion is what Maven holds while she waits for an answer to an open
// question. Unlike the yes/no confirms in cmd/mavend/voice.go, the answer here // question. Unlike the yes/no confirms in cmd/mavend/voice.go, the answer here
@@ -32,7 +33,17 @@ type PendingQuestion struct {
Utterance string // the user's original raw words Utterance string // the user's original raw words
Asked time.Time Asked time.Time
TTL time.Duration TTL time.Duration
Attempts int // questions already asked; capped by MaxAttempts Attempts int // questions already asked
// MaxAttempts caps Attempts. 0 ⇒ DefaultMaxAttempts.
MaxAttempts int
}
// maxAttempts is MaxAttempts with the default filled in.
func (q *PendingQuestion) maxAttempts() int {
if q.MaxAttempts <= 0 {
return DefaultMaxAttempts
}
return q.MaxAttempts
} }
func (q *PendingQuestion) IsExpired(now time.Time) bool { func (q *PendingQuestion) IsExpired(now time.Time) bool {
@@ -41,12 +52,9 @@ func (q *PendingQuestion) IsExpired(now time.Time) bool {
// CanAsk reports whether Maven may ask another question about this request. // CanAsk reports whether Maven may ask another question about this request.
func (q *PendingQuestion) CanAsk() bool { func (q *PendingQuestion) CanAsk() bool {
return q.Attempts < MaxAttempts return q.Attempts < q.maxAttempts()
} }
// TODO: the daemon will phrase the question text from Missing (one short ru
// question per Slot, feminine self-reference) and speak it here.
// ClarifyStore holds the parked questions. Same shape and locking as // ClarifyStore holds the parked questions. Same shape and locking as
// SessionStore: keyed by dialogue id, expired entries dropped on read. // SessionStore: keyed by dialogue id, expired entries dropped on read.
type ClarifyStore struct { type ClarifyStore struct {
@@ -67,8 +75,7 @@ func NewClarifyStore(defaultTTL time.Duration) *ClarifyStore {
} }
} }
// TODO: the daemon will Put a question here when Decision.Clarify fires, in // Put parks a question. Called on a clarify decision (cmd/mavend/clarify.go).
// place of the flat "не разобрала" reply (cmd/mavend/voice.go).
func (s *ClarifyStore) Put(id string, q *PendingQuestion) { func (s *ClarifyStore) Put(id string, q *PendingQuestion) {
if q.TTL <= 0 { if q.TTL <= 0 {
q.TTL = s.defaultTTL q.TTL = s.defaultTTL
@@ -78,8 +85,7 @@ func (s *ClarifyStore) Put(id string, q *PendingQuestion) {
s.mu.Unlock() s.mu.Unlock()
} }
// TODO: the daemon will Get on the next turn, parse that turn into Slots, call // Get returns the live parked question, or nil when there is none.
// Answer, and Delete — the open-question twin of resolveConfirm.
func (s *ClarifyStore) Get(id string, now time.Time) *PendingQuestion { func (s *ClarifyStore) Get(id string, now time.Time) *PendingQuestion {
s.mu.RLock() s.mu.RLock()
q, ok := s.questions[id] q, ok := s.questions[id]
@@ -94,6 +100,21 @@ func (s *ClarifyStore) Get(id string, now time.Time) *PendingQuestion {
return q return q
} }
// TakeExpired reports whether a question was parked here but its TTL ran out,
// and drops it. Get drops such a question silently, which leaves the user
// thinking his request is still alive — the caller uses this to tell him it is
// gone before treating his words as a fresh utterance.
func (s *ClarifyStore) TakeExpired(id string, now time.Time) bool {
s.mu.Lock()
defer s.mu.Unlock()
q, ok := s.questions[id]
if !ok || !q.IsExpired(now) {
return false
}
delete(s.questions, id)
return true
}
func (s *ClarifyStore) Delete(id string) { func (s *ClarifyStore) Delete(id string) {
s.mu.Lock() s.mu.Lock()
delete(s.questions, id) delete(s.questions, id)
@@ -101,44 +122,45 @@ func (s *ClarifyStore) Delete(id string) {
} }
// Answer merges the slots parsed from the user's answer into the parked ones. // Answer merges the slots parsed from the user's answer into the parked ones.
// Only the slots listed in Missing are filled, and an already filled slot is // Only the slots listed in Missing are touched. Within those, a value the answer
// never overwritten — the answer completes the original request, it does not // carries WINS over what was parked: she asked about this slot, so «нет, в пять»
// restate it. Parsing the answer text into `answer` is the caller's job; this // after «в три» must replace the time, not be thrown away.
// package must stay free of internal/router. //
// This is the clarify answer only. A correction in a fresh turn ("вообще-то
// перенеси на пять") is a different code path (followUpMerge) — not here.
//
// Parsing the answer text into `answer` is the caller's job; this package must
// stay free of internal/router.
func (q *PendingQuestion) Answer(text string, answer Slots) Slots { func (q *PendingQuestion) Answer(text string, answer Slots) Slots {
out := q.Slots out := q.Slots
for _, slot := range q.Missing { for _, slot := range q.Missing {
switch slot { switch slot {
case SlotTime: case SlotTime:
if !out.HasTime && answer.HasTime { if answer.HasTime {
out.Time = answer.Time out.Time = answer.Time
out.HasTime = true out.HasTime = true
} }
case SlotKey: case SlotKey:
if !out.HasKey && answer.HasKey { if answer.HasKey {
out.Key = answer.Key out.Key = answer.Key
out.HasKey = true out.HasKey = true
} }
case SlotValue: case SlotValue:
if out.Value == "" && answer.Value != "" { if answer.Value != "" {
out.Value = answer.Value out.Value = answer.Value
} }
case SlotFn: case SlotFn:
if !out.HasFn && answer.HasFn { if answer.HasFn {
out.Fn = answer.Fn out.Fn = answer.Fn
out.HasFn = true out.HasFn = true
if len(out.Args) == 0 { out.Args = append([]string(nil), answer.Args...)
out.Args = append([]string(nil), answer.Args...)
}
} }
case SlotText: case SlotText:
if out.Text == "" { if answer.Text != "" {
if answer.Text != "" { out.Text = answer.Text
out.Text = answer.Text } else if out.Text == "" {
} else { // No parse for a text slot — the raw answer IS the text.
// No parse for a text slot — the raw answer IS the text. out.Text = text
out.Text = text
}
} }
} }
} }
+48 -26
View File
@@ -1,6 +1,7 @@
package dialogue package dialogue
import ( import (
"reflect"
"testing" "testing"
"time" "time"
) )
@@ -59,6 +60,28 @@ func TestClarifyStoreGetPutDelete(t *testing.T) {
} }
} }
// TestClarifyStoreTakeExpired — TakeExpired reports (and drops) only a question
// whose TTL ran out.
func TestClarifyStoreTakeExpired(t *testing.T) {
s := NewClarifyStore(time.Minute)
if s.TakeExpired("voice", base) {
t.Fatal("nothing parked ⇒ nothing expired")
}
s.Put("voice", &PendingQuestion{Missing: []Slot{SlotTime}, Asked: base, TTL: time.Minute})
if s.TakeExpired("voice", base.Add(30*time.Second)) {
t.Fatal("a live question must not report as expired")
}
if s.Get("voice", base.Add(30*time.Second)) == nil {
t.Fatal("a live question must survive TakeExpired")
}
if !s.TakeExpired("voice", base.Add(2*time.Minute)) {
t.Fatal("a stale question must report as expired")
}
if s.TakeExpired("voice", base.Add(2*time.Minute)) {
t.Fatal("TakeExpired must drop the question, so the second call is false")
}
}
func TestNewClarifyStoreDefaultTTL(t *testing.T) { func TestNewClarifyStoreDefaultTTL(t *testing.T) {
s := NewClarifyStore(0) s := NewClarifyStore(0)
q := &PendingQuestion{Asked: base} q := &PendingQuestion{Asked: base}
@@ -89,12 +112,13 @@ func TestAnswerFillsOnlyMissingSlots(t *testing.T) {
want: Slots{Text: "напомни позвонить", Time: answerTime, HasTime: true}, want: Slots{Text: "напомни позвонить", Time: answerTime, HasTime: true},
}, },
{ {
name: "does not overwrite a filled time", // He restated it: «нет, в пять». The new value wins.
name: "a restated time overwrites the parked one",
parked: Slots{Time: other, HasTime: true}, parked: Slots{Time: other, HasTime: true},
missing: []Slot{SlotTime}, missing: []Slot{SlotTime},
text: "в три", text: "нет, в три",
answer: Slots{Time: answerTime, HasTime: true}, answer: Slots{Time: answerTime, HasTime: true},
want: Slots{Time: other, HasTime: true}, want: Slots{Time: answerTime, HasTime: true},
}, },
{ {
name: "ignores slots that were not missing", name: "ignores slots that were not missing",
@@ -121,12 +145,12 @@ func TestAnswerFillsOnlyMissingSlots(t *testing.T) {
want: Slots{Fn: "restart", Args: []string{"nginx"}, HasFn: true}, want: Slots{Fn: "restart", Args: []string{"nginx"}, HasFn: true},
}, },
{ {
name: "keeps existing args when fn was already known", name: "a restated fn replaces the fn and its args",
parked: Slots{Fn: "restart", Args: []string{"nginx"}, HasFn: true}, parked: Slots{Fn: "restart", Args: []string{"nginx"}, HasFn: true},
missing: []Slot{SlotFn}, missing: []Slot{SlotFn},
text: "останови postgres", text: "останови postgres",
answer: Slots{Fn: "stop", Args: []string{"postgres"}, HasFn: true}, answer: Slots{Fn: "stop", Args: []string{"postgres"}, HasFn: true},
want: Slots{Fn: "restart", Args: []string{"nginx"}, HasFn: true}, want: Slots{Fn: "stop", Args: []string{"postgres"}, HasFn: true},
}, },
{ {
name: "raw answer becomes the text when nothing was parsed", name: "raw answer becomes the text when nothing was parsed",
@@ -157,36 +181,34 @@ func TestAnswerFillsOnlyMissingSlots(t *testing.T) {
for _, tc := range cases { for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) { t.Run(tc.name, func(t *testing.T) {
q := &PendingQuestion{Slots: tc.parked, Missing: tc.missing, Asked: base} q := &PendingQuestion{Slots: tc.parked, Missing: tc.missing, Asked: base}
got := q.Answer(tc.text, tc.answer) // Whole-struct compare: a new field in Slots is covered for free.
if got.Time != tc.want.Time || got.HasTime != tc.want.HasTime || if got := q.Answer(tc.text, tc.answer); !reflect.DeepEqual(got, tc.want) {
got.Key != tc.want.Key || got.HasKey != tc.want.HasKey ||
got.Value != tc.want.Value || got.Text != tc.want.Text ||
got.Fn != tc.want.Fn || got.HasFn != tc.want.HasFn {
t.Fatalf("Answer = %+v, want %+v", got, tc.want) t.Fatalf("Answer = %+v, want %+v", got, tc.want)
} }
if len(got.Args) != len(tc.want.Args) {
t.Fatalf("Args = %v, want %v", got.Args, tc.want.Args)
}
for i := range got.Args {
if got.Args[i] != tc.want.Args[i] {
t.Fatalf("Args = %v, want %v", got.Args, tc.want.Args)
}
}
}) })
} }
} }
func TestCanAskCapsAtOneQuestion(t *testing.T) { func TestCanAskAllowsThreeQuestionsByDefault(t *testing.T) {
if MaxAttempts != 1 { if DefaultMaxAttempts != 3 {
t.Fatalf("MaxAttempts = %d, want 1 (Maven asks once, she is not a nag)", MaxAttempts) t.Fatalf("DefaultMaxAttempts = %d, want 3", DefaultMaxAttempts)
} }
q := &PendingQuestion{Asked: base} q := &PendingQuestion{Asked: base} // MaxAttempts unset ⇒ the default
if !q.CanAsk() { for i := 0; i < 3; i++ {
t.Fatal("a fresh question should be askable") if !q.CanAsk() {
t.Fatalf("question %d should be allowed", i+1)
}
q.Attempts++
} }
q.Attempts = MaxAttempts
if q.CanAsk() { if q.CanAsk() {
t.Fatal("the question should not be asked twice") t.Fatal("a fourth question must not be allowed")
}
}
func TestCanAskHonoursConfiguredMax(t *testing.T) {
q := &PendingQuestion{Asked: base, MaxAttempts: 1, Attempts: 1}
if q.CanAsk() {
t.Fatal("MaxAttempts 1 means one question only")
} }
} }
+74
View File
@@ -1,8 +1,12 @@
package dialogue package dialogue
import ( import (
"context"
"encoding/json"
"sync" "sync"
"time" "time"
"github.com/kami/maven/internal/store"
) )
type Intent string type Intent string
@@ -50,10 +54,22 @@ func (s *Session) IsExpired(now time.Time) bool {
return now.After(s.Timestamp.Add(s.TTL)) return now.After(s.Timestamp.Add(s.TTL))
} }
// SessionPersister — the bit of the store the session needs, as an interface
// so tests can swap it out. Data is an opaque blob: the store never looks
// inside, we encode the session as JSON here.
type SessionPersister interface {
SaveDialogueSession(ctx context.Context, id string, data []byte, ts time.Time, ttl time.Duration) error
DeleteDialogueSession(ctx context.Context, id string) error
LoadDialogueSessions(ctx context.Context, now time.Time) ([]store.DialogueSessionRow, error)
}
// SessionStore keeps the live sessions in a map (the fast path) and mirrors
// every write to the persister, so a daemon restart can load them back.
type SessionStore struct { type SessionStore struct {
mu sync.RWMutex mu sync.RWMutex
sessions map[string]*Session sessions map[string]*Session
defaultTTL time.Duration defaultTTL time.Duration
persist SessionPersister // may be nil: memory only (tests, no-store paths)
} }
func NewSessionStore(defaultTTL time.Duration) *SessionStore { func NewSessionStore(defaultTTL time.Duration) *SessionStore {
@@ -66,6 +82,43 @@ func NewSessionStore(defaultTTL time.Duration) *SessionStore {
} }
} }
// NewPersistentSessionStore — same store, but writes also go to the DB.
// Call Load once after this to bring back sessions from a previous run.
func NewPersistentSessionStore(defaultTTL time.Duration, p SessionPersister) *SessionStore {
s := NewSessionStore(defaultTTL)
s.persist = p
return s
}
// Load — read the saved sessions back into memory. Anything past its TTL is
// dropped (and deleted from the DB by the store), never revived.
func (s *SessionStore) Load(ctx context.Context, now time.Time) error {
if s.persist == nil {
return nil
}
rows, err := s.persist.LoadDialogueSessions(ctx, now)
if err != nil {
return err
}
s.mu.Lock()
defer s.mu.Unlock()
for _, r := range rows {
var sess Session
if err := json.Unmarshal(r.Data, &sess); err != nil {
// A blob we can't read is not worth failing a startup over.
continue
}
if sess.TTL <= 0 {
sess.TTL = r.TTL
}
if sess.IsExpired(now) {
continue
}
s.sessions[r.ID] = &sess
}
return nil
}
func (s *SessionStore) Get(id string, now time.Time) *Session { func (s *SessionStore) Get(id string, now time.Time) *Session {
s.mu.RLock() s.mu.RLock()
sess, ok := s.sessions[id] sess, ok := s.sessions[id]
@@ -87,12 +140,33 @@ func (s *SessionStore) Put(id string, sess *Session) {
s.mu.Lock() s.mu.Lock()
s.sessions[id] = sess s.sessions[id] = sess
s.mu.Unlock() s.mu.Unlock()
s.save(id, sess)
} }
func (s *SessionStore) Delete(id string) { func (s *SessionStore) Delete(id string) {
s.mu.Lock() s.mu.Lock()
delete(s.sessions, id) delete(s.sessions, id)
s.mu.Unlock() s.mu.Unlock()
if s.persist != nil {
_ = s.persist.DeleteDialogueSession(context.Background(), id)
}
}
// save — mirror one session to the DB. Best effort: memory already has it, so
// a write error costs us the restart safety net, not the current turn.
func (s *SessionStore) save(id string, sess *Session) {
if s.persist == nil {
return
}
data, err := json.Marshal(sess)
if err != nil {
return
}
ts := sess.Timestamp
if ts.IsZero() {
ts = time.Now()
}
_ = s.persist.SaveDialogueSession(context.Background(), id, data, ts, sess.TTL)
} }
func InheritSlots(prev, cur Slots) Slots { func InheritSlots(prev, cur Slots) Slots {
+118
View File
@@ -0,0 +1,118 @@
package dialogue
import (
"context"
"path/filepath"
"testing"
"time"
"github.com/kami/maven/internal/store"
)
// openStore — a store on disk, so a second handle can reopen the same file.
func openStore(t *testing.T, path string) *store.Store {
t.Helper()
s, err := store.Open(context.Background(), path)
if err != nil {
t.Fatalf("open store: %v", err)
}
t.Cleanup(func() { _ = s.Close() })
return s
}
// A session written before a restart comes back and still merges a follow-up.
func TestSessionSurvivesRestart(t *testing.T) {
ctx := context.Background()
path := filepath.Join(t.TempDir(), "maven_test.db")
now := time.Now().UTC().Truncate(time.Millisecond)
first := openStore(t, path)
before := NewPersistentSessionStore(2*time.Minute, first)
before.Put("voice", &Session{
Intent: IntentReminder,
Slots: Slots{Text: "полить цветы", Time: now.Add(time.Hour), HasTime: true},
Timestamp: now,
TTL: 2 * time.Minute,
})
if err := first.Close(); err != nil {
t.Fatalf("close: %v", err)
}
// fresh handle, fresh in-memory map — as after a daemon restart
after := openStore(t, path)
reloaded := NewPersistentSessionStore(2*time.Minute, after)
if err := reloaded.Load(ctx, now.Add(10*time.Second)); err != nil {
t.Fatalf("Load: %v", err)
}
sess := reloaded.Get("voice", now.Add(10*time.Second))
if sess == nil {
t.Fatal("session did not survive the restart")
}
if sess.Intent != IntentReminder {
t.Fatalf("intent = %q, want reminder", sess.Intent)
}
// the follow-up carries no text of its own; it must inherit the old one
merged := InheritSlots(sess.Slots, Slots{Time: now.Add(2 * time.Hour), HasTime: true})
if merged.Text != "полить цветы" {
t.Fatalf("merged text = %q, want the earlier turn's text", merged.Text)
}
if !merged.Time.Equal(now.Add(2 * time.Hour)) {
t.Fatalf("merged time = %v, want the follow-up's time", merged.Time)
}
}
// A session past its TTL is dead: a restart must not bring it back.
func TestExpiredSessionNotResurrected(t *testing.T) {
ctx := context.Background()
path := filepath.Join(t.TempDir(), "maven_test.db")
now := time.Now().UTC().Truncate(time.Millisecond)
first := openStore(t, path)
before := NewPersistentSessionStore(2*time.Minute, first)
before.Put("voice", &Session{
Intent: IntentReminder,
Slots: Slots{Text: "полить цветы"},
Timestamp: now,
TTL: time.Minute,
})
if err := first.Close(); err != nil {
t.Fatalf("close: %v", err)
}
after := openStore(t, path)
reloaded := NewPersistentSessionStore(2*time.Minute, after)
later := now.Add(5 * time.Minute) // well past the 1-min TTL
if err := reloaded.Load(ctx, later); err != nil {
t.Fatalf("Load: %v", err)
}
if sess := reloaded.Get("voice", later); sess != nil {
t.Fatalf("expired session came back: %+v", sess)
}
// and it is gone from the DB too, not just from memory
rows, err := after.LoadDialogueSessions(ctx, later)
if err != nil {
t.Fatalf("LoadDialogueSessions: %v", err)
}
if len(rows) != 0 {
t.Fatalf("expired row still in the DB: %+v", rows)
}
}
// Delete removes the row as well, so an ended conversation stays ended.
func TestDeleteRemovesPersistedSession(t *testing.T) {
ctx := context.Background()
path := filepath.Join(t.TempDir(), "maven_test.db")
now := time.Now().UTC().Truncate(time.Millisecond)
s := openStore(t, path)
ss := NewPersistentSessionStore(2*time.Minute, s)
ss.Put("voice", &Session{Intent: IntentChat, Timestamp: now, TTL: time.Minute})
ss.Delete("voice")
rows, err := s.LoadDialogueSessions(ctx, now)
if err != nil {
t.Fatalf("LoadDialogueSessions: %v", err)
}
if len(rows) != 0 {
t.Fatalf("row survived Delete: %+v", rows)
}
}
+127 -1
View File
@@ -17,10 +17,14 @@ const (
CheckFeminine = "feminine" // her self-reference is feminine (hard constraint) CheckFeminine = "feminine" // her self-reference is feminine (hard constraint)
CheckCringe = "cringe" // DESIGN.md § Non-goals, "not a relationship" CheckCringe = "cringe" // DESIGN.md § Non-goals, "not a relationship"
CheckOnTopic = "ontopic" // says the thing the rule is about CheckOnTopic = "ontopic" // says the thing the rule is about
// CheckHisGender — the other half of the persona rule: SHE is feminine, HE
// is male. "ты давно не отдыхала" addresses the operator as a woman.
CheckHisGender = "hisgender"
) )
// CheckNames — report order. // CheckNames — report order.
var CheckNames = []string{CheckMood, CheckLang, CheckLength, CheckFeminine, CheckCringe, CheckOnTopic} var CheckNames = []string{CheckMood, CheckLang, CheckLength, CheckFeminine, CheckHisGender, CheckCringe, CheckOnTopic}
// Result — one check on one message. // Result — one check on one message.
type Result struct { type Result struct {
@@ -52,6 +56,7 @@ func RunChecks(c Case, body, mood string) []Result {
checkLang(body), checkLang(body),
checkLength(body), checkLength(body),
checkFeminine(body), checkFeminine(body),
checkHisGender(body),
checkCringe(body), checkCringe(body),
checkOnTopic(c, body), checkOnTopic(c, body),
} }
@@ -176,6 +181,127 @@ func checkFeminine(body string) Result {
return Result{CheckFeminine, true, ""} return Result{CheckFeminine, true, ""}
} }
// --- he is male ----------------------------------------------------------
//
// The mirror of checkFeminine, and the failure it was written for: the model
// wrote "ты давно не отдыхала", which addresses the operator as a woman. That
// scored clean, because checkFeminine only ever looks at how SHE speaks about
// herself.
//
// How it works: Russian past tense is gendered by suffix, -л (m) / -ла (f). So
// this looks for feminine past-tense words in a sentence that also talks TO him
// ("ты", "тебя", "тебе", "твой", …). A feminine verb that belongs to her ("я
// заметила", "напомнила тебе") is skipped — that one is correct.
//
// Honest about the limits: this is a suffix rule, not a parser.
// - False positives: a feminine noun can be the subject in the same sentence
// ("зарядка была утром, ты её пропустил"). The guard below skips a verb whose
// previous word looks like a feminine noun, which helps but will not always
// be right.
// - False negatives: gender also shows up outside the past tense (short
// adjectives, "сама"), and none of that is checked here.
//
// That is acceptable for an eval check. It is a signal to read the message, not
// a grammar verdict, and every hit prints the word it tripped on so a human can
// disagree.
// hisMarkers — words that mean the sentence is addressed to him.
var hisMarkers = map[string]bool{
"ты": true, "тебя": true, "тебе": true, "тобой": true, "тобою": true,
"твой": true, "твоя": true, "твоё": true, "твое": true, "твои": true, "твою": true,
}
// notFeminineVerb — ordinary words ending in "-ла" that are not verbs. Small on
// purpose: it only has to cover words a nudge might actually use.
//
// Words that are both a noun and a verb are deliberately NOT here. "села",
// "мыла" and "стекла" are nouns on paper, but in a nudge they are almost always
// verbs ("ты села", "ты мыла"), and listing them would make the check miss the
// exact thing it is for. Missing a real hit is worse than one false alarm.
var notFeminineVerb = map[string]bool{
"школа": true, "скала": true, "игла": true, "метла": true, "смола": true,
"дела": true, "тела": true, "масла": true, "весла": true,
"зола": true, "пчела": true, "числа": true,
}
// femininePast reports whether a word looks like a feminine past-tense verb:
// "отдыхала", "поела", "выспалась".
func femininePast(w string) bool {
if len([]rune(w)) < 3 || notFeminineVerb[w] {
return false
}
return strings.HasSuffix(w, "ла") || strings.HasSuffix(w, "лась")
}
// looksFeminineNoun — a crude guard against "зарядка была": a word right before
// the verb that ends in "а"/"я" and is not itself a verb is probably the subject.
func looksFeminineNoun(w string) bool {
if femininePast(w) || len([]rune(w)) < 3 {
return false
}
return strings.HasSuffix(w, "а") || strings.HasSuffix(w, "я")
}
// sentenceRE splits on sentence-ending punctuation, so a feminine verb in one
// sentence is not blamed on a "ты" in the next.
var sentenceRE = regexp.MustCompile(`[.!?;…]+`)
func checkHisGender(body string) Result {
for _, sentence := range sentenceRE.Split(strings.ToLower(body), -1) {
words := wordRE.FindAllString(sentence, -1)
addressed := false
for _, w := range words {
if hisMarkers[w] {
addressed = true
}
}
if !addressed {
continue
}
for i, w := range words {
if !femininePast(w) || hersNotHis(words, i) {
continue
}
if i > 0 && looksFeminineNoun(prevWord(words, i)) {
continue
}
return Result{CheckHisGender, false,
fmt.Sprintf("feminine %q addressed to him — he is male", w)}
}
}
return Result{CheckHisGender, true, ""}
}
// hersNotHis — the verb is Maven's own if "я" comes shortly before it, or if the
// thing she did was done to him ("напомнила тебе", "проверила за тебя").
func hersNotHis(words []string, i int) bool {
for j := i - 1; j >= 0 && j >= i-3; j-- {
if words[j] == "я" {
return true
}
}
if i+1 < len(words) {
switch words[i+1] {
case "тебе", "тебя", "за", "тобой":
return true
}
}
return false
}
// prevWord — the word before i, skipping "не" and punctuation, so "не отдыхала"
// still sees the subject.
func prevWord(words []string, i int) string {
for j := i - 1; j >= 0; j-- {
w := words[j]
if w == "не" || w == "ни" || !unicode.Is(unicode.Cyrillic, []rune(w)[0]) {
continue
}
return w
}
return ""
}
// --- the cringe checks --------------------------------------------------- // --- the cringe checks ---------------------------------------------------
// //
// "Think Jarvis without the cringe part". DESIGN.md § Non-goals: "Not a // "Think Jarvis without the cringe part". DESIGN.md § Non-goals: "Not a
+1 -1
View File
@@ -281,7 +281,7 @@ func (r Report) String() string {
fmt.Fprintf(&b, "%s: %d/%d cases pass every check (%.1f%%), %d errors\n", fmt.Fprintf(&b, "%s: %d/%d cases pass every check (%.1f%%), %d errors\n",
r.Name, r.Passed, r.Total, 100*r.Accuracy(), r.Errors) r.Name, r.Passed, r.Total, 100*r.Accuracy(), r.Errors)
for _, name := range CheckNames { for _, name := range CheckNames {
fmt.Fprintf(&b, " %-9s %d/%d\n", name, r.ByCheck[name], r.Total) fmt.Fprintf(&b, " %-10s %d/%d\n", name, r.ByCheck[name], r.Total)
} }
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max) fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
fmt.Fprintf(&b, " by rule: %s\n", renderStats(r.ByRule)) fmt.Fprintf(&b, " by rule: %s\n", renderStats(r.ByRule))
+13 -4
View File
@@ -72,10 +72,11 @@ func TestStubBaseline(t *testing.T) {
// ceiling ("you've been at your desk for 4 hours without a break — step // ceiling ("you've been at your desk for 4 hours without a break — step
// away for a bit." is 76 chars but 16 words). Left failing rather than // away for a bit." is 76 chars but 16 words). Left failing rather than
// raising the ceiling to hide it. // raising the ceiling to hide it.
CheckLength: 12, CheckLength: 12,
CheckFeminine: 15, CheckFeminine: 15,
CheckCringe: 15, CheckHisGender: 15,
CheckOnTopic: 12, CheckCringe: 15,
CheckOnTopic: 12,
} }
for name, floor := range floors { for name, floor := range floors {
if rep.ByCheck[name] < floor { if rep.ByCheck[name] < floor {
@@ -104,6 +105,14 @@ func TestChecksCatchWhatTheyClaim(t *testing.T) {
{"masculine predicative", "я должен сказать: попей воды.", CheckFeminine}, {"masculine predicative", "я должен сказать: попей воды.", CheckFeminine},
// The other direction: HE is male, so second-person masculine is right. // The other direction: HE is male, so second-person masculine is right.
{"second person masculine ok", "ты не пил воду четыре часа.", ""}, {"second person masculine ok", "ты не пил воду четыре часа.", ""},
// The real observed failure: she addressed him as a woman.
{"feminine second person", "ты давно не отдыхала — попей воды.", CheckHisGender},
{"feminine second person no dash", "ты пила воду четыре часа назад.", CheckHisGender},
// Her own feminine verb next to "ты" is correct and must not be flagged.
{"her feminine verb near ты", "я заметила, что ты не пил воду.", ""},
{"her feminine verb about him", "напомнила тебе про воду.", ""},
// A feminine noun subject in the same sentence is not him.
{"feminine noun subject ok", "зарядка была утром, ты её пропустил, попей воды.", ""},
{"feminine self ok", "я заметила: воды не было четыре часа.", ""}, {"feminine self ok", "я заметила: воды не было четыре часа.", ""},
{"pet name", "милый, попей воды.", CheckCringe}, {"pet name", "милый, попей воды.", CheckCringe},
{"emoji", "попей воды 💧", CheckCringe}, {"emoji", "попей воды 💧", CheckCringe},
+14 -96
View File
@@ -1,11 +1,8 @@
package eval package eval
import ( import (
"bytes"
"context" "context"
"encoding/json"
"fmt" "fmt"
"net/http"
"os" "os"
"strings" "strings"
"testing" "testing"
@@ -29,13 +26,20 @@ import (
// a bake-off across checkpoints (#278, #250) produces tables you can tell // a bake-off across checkpoints (#278, #250) produces tables you can tell
// apart. Point the variable at one server at a time. // apart. Point the variable at one server at a time.
// //
// Three configurations, because "the LLM router" is ambiguous and the three // Two configurations, because "the LLM router" is ambiguous and the two numbers
// numbers answer different questions: // answer different questions:
// //
// llm-only — the model alone. Measures the prompt + grammar contract. // llm-only — the model alone. Measures the prompt + grammar contract.
// cascade+llm — what #320 would actually ship: stage-0 grammar, then the // cascade+llm — what #320 would actually ship: stage-0 grammar, then the
// model, then the classifier as the failure floor. // model, then the classifier as the failure floor.
// llm-no-thinking — diagnostic only, not a shippable path (see below). //
// There used to be a third, "thinking off", which looked 6 points better. It is
// gone: it was measured with a hand-rolled HTTP client that quietly dropped
// repeat_penalty, so the gap was the missing penalty and not the thinking mode.
// Re-measured with everything else held equal, thinking off scores exactly the
// same, case for case — and a direct probe shows this llama-server build ignores
// enable_thinking / reasoning_budget for this model anyway, so there was nothing
// to turn off. Full write-up in ROUTING-EVAL-31-07-2026.md (Vikunja #376).
func TestLLMRouterBaseline(t *testing.T) { func TestLLMRouterBaseline(t *testing.T) {
base := os.Getenv("MAVEN_LLM_URL") base := os.Getenv("MAVEN_LLM_URL")
if base == "" { if base == "" {
@@ -96,104 +100,18 @@ func TestLLMRouterBaseline(t *testing.T) {
} }
t.Log("\n" + repCascade.String() + repCascade.Failures()) t.Log("\n" + repCascade.String() + repCascade.Failures())
// llm-no-thinking: same prompt and grammar with the chat template's
// thinking mode off. Qwen3.5's template defaults thinking=1, so under a
// grammar the constrained JSON lands in reasoning_content with content
// empty — llm.Client's ReasoningContent fallback is what makes the router
// work at all today, by accident rather than design.
//
// MEASURED 2026-07-31: this variant scores identically to as-deployed
// (18/76, 48.7% intent-only, 2 errors, same p50). Thinking mode is a
// non-issue under a grammar — llama.cpp constrains the same token stream
// either way. Kept so the question stays answered instead of being
// re-asked, and so internal/llm does NOT grow a chat_template_kwargs field
// for a problem that does not exist.
repNoThink, err := Score(ctx, "llm-only ("+model+", thinking off) [diagnostic]",
RouterFunc(func(ctx context.Context, u string, now time.Time) (router.Decision, error) {
d, ok, err := router.NewLLMRouter(&noThinkCompleter{base: base, http: &http.Client{Timeout: 60 * time.Second}}).Route(ctx, u, now)
if err != nil {
return d, err
}
if !ok {
return d, fmt.Errorf("llm router declined without an error")
}
return d, nil
}), f)
if err != nil {
t.Fatalf("Score no-thinking: %v", err)
}
t.Log("\n" + repNoThink.String() + repNoThink.Failures())
// Reports rather than asserts — the numbers are inputs to the #320 // Reports rather than asserts — the numbers are inputs to the #320
// decision, and an assertion here would be this test inventing the bar. // decision, and an assertion here would be this test inventing the bar.
// The one thing worth failing on is a harness fault: if every single case // The one thing worth failing on is a harness fault: if every single case
// errors, the run measured infrastructure, not routing, and the report // errors, the run measured infrastructure, not routing, and the report
// must not be mistaken for a score. // must not be mistaken for a score.
for _, rep := range []Report{repLLM, repCascade, repNoThink} { for _, rep := range []Report{repLLM, repCascade} {
if rep.Errors == rep.Total { if rep.Errors == rep.Total {
t.Errorf("%s: all %d cases errored — harness fault, not a measurement", rep.Name, rep.Total) t.Errorf("%s: all %d cases errored — harness fault, not a measurement", rep.Name, rep.Total)
} }
} }
} }
// noThinkCompleter — llm.Client with chat_template_kwargs.enable_thinking
// false. A test-local copy rather than a change to internal/llm: whether the
// daemon should send it is the open question, and answering it here by adding
// the field would prejudge #320.
type noThinkCompleter struct {
base string
http *http.Client
}
func (c *noThinkCompleter) Complete(ctx context.Context, r llm.Req) (string, error) {
payload := map[string]any{
"messages": []map[string]string{
{"role": "system", "content": r.System},
{"role": "user", "content": r.User},
},
"max_tokens": r.MaxTokens,
"temperature": 0,
"grammar": r.Grammar,
"chat_template_kwargs": map[string]any{"enable_thinking": false},
}
b, err := json.Marshal(payload)
if err != nil {
return "", err
}
req, err := http.NewRequestWithContext(ctx, "POST", c.base+"/v1/chat/completions", bytes.NewReader(b))
if err != nil {
return "", err
}
req.Header.Set("Content-Type", "application/json")
resp, err := c.http.Do(req)
if err != nil {
return "", err
}
defer resp.Body.Close()
if resp.StatusCode != 200 {
return "", fmt.Errorf("status %d", resp.StatusCode)
}
var out struct {
Choices []struct {
Message struct {
Content string `json:"content"`
ReasoningContent string `json:"reasoning_content"`
} `json:"message"`
} `json:"choices"`
}
if err := json.NewDecoder(resp.Body).Decode(&out); err != nil {
return "", err
}
if len(out.Choices) == 0 {
return "", fmt.Errorf("no choices")
}
m := out.Choices[0].Message
if m.Content != "" {
return m.Content, nil
}
return m.ReasoningContent, nil
}
func ping(ctx context.Context, c *llm.Client) error { func ping(ctx context.Context, c *llm.Client) error {
ctx, cancel := context.WithTimeout(ctx, 90*time.Second) ctx, cancel := context.WithTimeout(ctx, 90*time.Second)
defer cancel() defer cancel()
+74
View File
@@ -0,0 +1,74 @@
package store
import (
"context"
"fmt"
"time"
)
// DialogueSessionRow — one saved follow-up session. Data is the session
// encoded by the dialogue package; the store does not look inside it.
type DialogueSessionRow struct {
ID string
Data []byte
Ts time.Time
TTL time.Duration
Expires time.Time
}
// SaveDialogueSession — write (or replace) the session for one dialogue id.
// One row per id: a newer turn overwrites the older state.
func (s *Store) SaveDialogueSession(ctx context.Context, id string, data []byte, ts time.Time, ttl time.Duration) error {
expires := ts.Add(ttl)
_, err := s.db.ExecContext(ctx, `
INSERT INTO dialogue_sessions (id, data, ts, ttl_ms, expires_ts) VALUES (?, ?, ?, ?, ?)
ON CONFLICT(id) DO UPDATE SET data = excluded.data,
ts = excluded.ts,
ttl_ms = excluded.ttl_ms,
expires_ts = excluded.expires_ts`,
id, data, ts.UnixMilli(), ttl.Milliseconds(), expires.UnixMilli())
if err != nil {
return fmt.Errorf("save dialogue session: %w", err)
}
return nil
}
// DeleteDialogueSession — drop one session (ended, or expired).
func (s *Store) DeleteDialogueSession(ctx context.Context, id string) error {
if _, err := s.db.ExecContext(ctx, `DELETE FROM dialogue_sessions WHERE id = ?`, id); err != nil {
return fmt.Errorf("delete dialogue session: %w", err)
}
return nil
}
// LoadDialogueSessions — return the sessions still alive at `now` and delete
// the ones that already ran out. An expired session is dead: it never comes
// back after a restart.
func (s *Store) LoadDialogueSessions(ctx context.Context, now time.Time) ([]DialogueSessionRow, error) {
if _, err := s.db.ExecContext(ctx,
`DELETE FROM dialogue_sessions WHERE expires_ts <= ?`, now.UnixMilli()); err != nil {
return nil, fmt.Errorf("prune dialogue sessions: %w", err)
}
rows, err := s.db.QueryContext(ctx,
`SELECT id, data, ts, ttl_ms, expires_ts FROM dialogue_sessions ORDER BY id`)
if err != nil {
return nil, fmt.Errorf("load dialogue sessions: %w", err)
}
defer rows.Close()
var out []DialogueSessionRow
for rows.Next() {
var r DialogueSessionRow
var tsMilli, ttlMilli, expMilli int64
if err := rows.Scan(&r.ID, &r.Data, &tsMilli, &ttlMilli, &expMilli); err != nil {
return nil, fmt.Errorf("scan dialogue session: %w", err)
}
r.Ts = time.UnixMilli(tsMilli).UTC()
r.TTL = time.Duration(ttlMilli) * time.Millisecond
r.Expires = time.UnixMilli(expMilli).UTC()
out = append(out, r)
}
if err := rows.Err(); err != nil {
return nil, fmt.Errorf("load dialogue sessions: %w", err)
}
return out, nil
}
+9
View File
@@ -74,6 +74,15 @@ ALTER TABLE reminders ADD COLUMN next_fire_ts INTEGER;`, // #2
`CREATE INDEX IF NOT EXISTS idx_nudges_snoozed ON nudges (outcome_ts) WHERE outcome = 'snoozed';`, // #8 — SnoozedUntil runs every tick; keep it off a full scan (Vikunja #364) `CREATE INDEX IF NOT EXISTS idx_nudges_snoozed ON nudges (outcome_ts) WHERE outcome = 'snoozed';`, // #8 — SnoozedUntil runs every tick; keep it off a full scan (Vikunja #364)
`ALTER TABLE proposed_routines ADD COLUMN accepted_ts INTEGER; `ALTER TABLE proposed_routines ADD COLUMN accepted_ts INTEGER;
ALTER TABLE proposed_routines ADD COLUMN last_fired_ts INTEGER;`, // #9 — accepted routines keep firing (Vikunja #366): the tick loop needs to know when a routine was accepted and when it last nudged ALTER TABLE proposed_routines ADD COLUMN last_fired_ts INTEGER;`, // #9 — accepted routines keep firing (Vikunja #366): the tick loop needs to know when a routine was accepted and when it last nudged
`CREATE TABLE IF NOT EXISTS dialogue_sessions (
id TEXT PRIMARY KEY,
data BLOB NOT NULL,
ts INTEGER NOT NULL,
ttl_ms INTEGER NOT NULL,
expires_ts INTEGER NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_dialogue_sessions_expires ON dialogue_sessions (expires_ts);`, // #10 — the follow-up session survives a restart (Vikunja #363); small, TTL-pruned table, not a history log
} }
// migrate applies every migration with a number greater than the DB's current // migrate applies every migration with a number greater than the DB's current