Compare commits
12 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| b9a24334ea | |||
| c97aebf55a | |||
| 891136c65d | |||
| 41c7c13f42 | |||
| a324e8f624 | |||
| 51805e7f35 | |||
| 742b2ad1d7 | |||
| b300ac5c70 | |||
| 0b90952e55 | |||
| ddb658ffbb | |||
| ee3e6a9eaf | |||
| 8acb8a97c6 |
@@ -0,0 +1,150 @@
|
||||
# Conversational phrasing eval — 31-07-2026
|
||||
|
||||
Every score measured tonight, on the three paths the nudge eval never touched:
|
||||
chat, query-with-notes, and general knowledge.
|
||||
|
||||
**Short version: the plumbing got fixed and the score barely moved.** Grammar and
|
||||
Russian prompts together took the composite from ~9 to ~14 of 27. Everything
|
||||
still failing is the model not knowing things or not holding a constraint, and
|
||||
prompting is out of levers. Settles the measurement half of Vikunja #395 / #398 /
|
||||
#400.
|
||||
|
||||
## How to reproduce
|
||||
|
||||
```sh
|
||||
# llama-server: -c 4096 -ngl 99 -t 6, model /mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf
|
||||
MAVEN_LLM_URL=http://127.0.0.1:18099 no_proxy=127.0.0.1,localhost \
|
||||
deps/go/go/bin/go test -count=1 -timeout 40m \
|
||||
-run TestLLMTalkBaseline ./internal/phraser/eval/ -v
|
||||
```
|
||||
|
||||
Three runs per configuration, always. The fixture is 27 cases, so one reply
|
||||
changing moves the composite by 3.7 points — a single run cannot tell a real
|
||||
change from sampling noise. This was learned the expensive way: an earlier claim
|
||||
that "one nudge case fails every run" turned out to be three different cases
|
||||
across three runs.
|
||||
|
||||
**Run the box otherwise idle.** See the contamination note at the bottom.
|
||||
|
||||
## Composite, per configuration
|
||||
|
||||
| config | overall /27 | chat /9 | query /9 | knowledge /9 | canned fallbacks |
|
||||
|---|---|---|---|---|---|
|
||||
| baseline, no grammar | 7, 12, 7 | 1, 1, 0 | 2, 4, 2 | 4, 7, 5 | 0, 0, 0 |
|
||||
| + GBNF grammar (#398) | 14, 15, 8 | 1, 3, 0 | 5, 6, 3 | 8, 6, 5 | 0, 0, 0 |
|
||||
| + Russian prompts (#400) | 11, 17, 15 | 1, 5, 3 | 5, 6, 8 | 5, 6, 4 | 0, 0, 0 |
|
||||
| + truncation fix, 1000ch/768tok | 12, 13, 10 | 2, 2, 1 | 7, 7, 5 | 3, 4, 4 | 3, 3, 6 |
|
||||
| + rebalanced, 600ch/1024tok | **void — contaminated** | | | | |
|
||||
|
||||
"Canned fallbacks" counts replies that came back as the hardcoded `"не знаю."`
|
||||
or `"поговорили."`. It is not a check, it is a health signal: those strings mean
|
||||
the phraser gave up, and the eval scores them as ordinary bad replies.
|
||||
|
||||
## Per-check
|
||||
|
||||
| check | no grammar | + grammar | + RU prompts | + truncation fix |
|
||||
|---|---|---|---|---|
|
||||
| nonempty | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 |
|
||||
| ellipsis | 20, 19, 23 | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 |
|
||||
| lang | 13, 16, 15 | 23, 26, 26 | 25, 26, 25 | 26, 27, 27 |
|
||||
| feminine | — | — | 25, 24, 26 | 25, 25, 27 |
|
||||
| address | — | — | 21, 22, 22 | 22, 21, 22 |
|
||||
| ontopic | — | — | 17, 24, 18 | 17, 19, 14 |
|
||||
|
||||
`nonempty` reading 27/27 everywhere is not good news — it was a broken check.
|
||||
It tested for a non-blank string, so replies of literally `{` and `"15-16"`
|
||||
passed it. Fixed on `overnight/fix-truncation`; it needs a letter now.
|
||||
|
||||
## What each change actually bought
|
||||
|
||||
**GBNF grammar (#398) — the biggest single win.** Qwen3.5-0.8B writes
|
||||
`Thinking Process:` as plain text with no tags, `stripThink` only handles
|
||||
`</think>`, so the JSON never closed and the plain-text fallback shipped the
|
||||
literal reasoning. `ellipsis` went 20→27 and `lang` 13→26. The router had been
|
||||
using a grammar for ages; the phraser asking nicely in the prompt was the
|
||||
oversight.
|
||||
|
||||
**Russian prompts (#400) — modest, plus a large latency win.** Chat 1.3→3.0
|
||||
average, query 4.7→6.3, knowledge 6.3→5.0. All inside the run-to-run spread, so
|
||||
"probably better on the paths it targeted, not provable in three runs". p50
|
||||
latency dropped from ~11.5s to ~2.3s and that part is consistent across all
|
||||
three runs — shorter prompts, and she stopped emitting English reasoning first.
|
||||
|
||||
**Truncation fix — necessary, and did not help the score.** Two real bugs
|
||||
(replies of `{`, and a `nonempty` check that passed them), both fixed, and the
|
||||
composite went nowhere. A complete rambling wrong answer fails the same checks a
|
||||
truncated one did. Worth doing anyway: the daemon was shipping `{` to a
|
||||
text-to-speech voice.
|
||||
|
||||
## The truncation bug, since the cause was counter-intuitive
|
||||
|
||||
The grammar's `string ::= ... {0,400}` rule was the cause, not the token cap.
|
||||
Measured against Qwen3.5-0.8B at three caps — 256, 768 and 2048 — the reply came
|
||||
back **exactly 400 characters every time, cut mid-word** (`"Нужно записать и,"`).
|
||||
|
||||
Then I raised the bound to 1000 while the cap was 768 tokens and made it worse:
|
||||
Russian runs ~1.5 characters per token here, so generation died on the *token*
|
||||
cap instead, mid-object, and the new guard correctly refused it and shipped
|
||||
`"не знаю."` — 3, 3 and 6 fallbacks per run, from zero. **The two limits have to
|
||||
agree.** 600 characters needs ~400 tokens; the cap is 1024.
|
||||
|
||||
## Where the remaining failures live
|
||||
|
||||
`address` is stuck at 21-22 of 27 and `ontopic` at 14-19. Both resist prompting.
|
||||
|
||||
**The prompt now explicitly forbids exactly what she does.** It says never "вы",
|
||||
use the singular — and she writes `вашей`, `подождите`, `делаете`, `хотите`,
|
||||
`напишите`. Telling a 0.8B "never do X" does not work. Same for
|
||||
`feminine`: `я готов`, `я понял`, `я нашел`, `я заметил`, `я сказал`.
|
||||
|
||||
**Some of `ontopic` is the fixture, not the model.** `chat-how-are-you` got
|
||||
`"Привет! Я здесь, чтобы поговорить. Как дела сегодня?"` — a fine reply that
|
||||
fails because `want_any` is `[норм, хорош, порядк, тут, работ]`. It fails in
|
||||
every run, so it inflates the count. The `ontopic` column currently measures the
|
||||
fixture as much as the model. Not fixed yet, deliberately: changing it would
|
||||
break comparability with the runs above.
|
||||
|
||||
**Two replies worth reading, because they are not fixable by prompting:**
|
||||
|
||||
- Thunder and lightning: *"Скорость молнии — 8-10 тысяч километров в секунду, но
|
||||
звук — 300 метров в секунду, что делает молнию громче."* Confidently wrong,
|
||||
and it concludes lightning is *louder* rather than sound being *slower*.
|
||||
- "расскажи обо мне": *"Ты — прекрасное существо, с душой и вниманием… Спасибо за
|
||||
твою улыбку… О тебе — заповедь любви."* Sycophantic filler, zero information,
|
||||
and precisely the "not a relationship" non-goal.
|
||||
- Boiling an egg: `"15-16"` one run, `"1"` another. No unit, wrong number.
|
||||
|
||||
The first argues for reading instead of recalling (#403 — Kiwix retrieval scores
|
||||
8/8 on the same questions given English keywords). The second and third argue
|
||||
for templates on the paths where correctness matters (#392).
|
||||
|
||||
## Contamination note — how the last row got voided
|
||||
|
||||
I started the query-rewrite agent against the same llama-server the sweep was
|
||||
using, and assumed contention would only affect latency. It did not. The
|
||||
knowledge path collapsed to 0 of 9 with eight canned `"не знаю."` replies, p95
|
||||
tripled to 23.7s, and **the report still said "0 errors"**.
|
||||
|
||||
That is Vikunja #397, and it is worse than filed: a merely *busy* server
|
||||
produces a clean-looking report with a third of the fixture silently answering
|
||||
`"не знаю."`. `PhraseChat` and `PhraseQuery` swallow every failure and return a
|
||||
hardcoded string, so infrastructure trouble is indistinguishable from bad
|
||||
phrasing in the score. The talk test guards the *start* and *end* of a run with
|
||||
a model check, which catches a dead server but not a loaded one.
|
||||
|
||||
**Until #397 is fixed, treat any run made on a busy box as void.**
|
||||
|
||||
## Next
|
||||
|
||||
- Re-run 600ch/1024tok clean, to fill the void row.
|
||||
- Score `Qwen3.5-2B-UD-Q4_K_XL` (already at `/mnt/hdd1/llms/qwen3.5/`, never
|
||||
measured) on this fixture and the router fixture. Not the 4B — too big for
|
||||
this box, owner's call.
|
||||
- Newer sub-500M candidates (LFM2.5 200M/300M) are worth a run for routing.
|
||||
Note `MODEL-BAKEOFF-31-07-2026.md` found LFM2.5-**1.2B** worse than
|
||||
Qwen3.5-0.8B at Russian routing and 2.4× slower — but those are a different,
|
||||
older generation, so that result does not predict the small ones.
|
||||
- Fix `chat-how-are-you`'s `want_any`, and re-baseline once, so `ontopic`
|
||||
measures the model.
|
||||
- #397 first if anything, since it decides whether any of the above is
|
||||
trustworthy.
|
||||
+57
-6
@@ -3,6 +3,8 @@ package main
|
||||
import (
|
||||
"context"
|
||||
"log"
|
||||
"math/rand"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/dialogue"
|
||||
@@ -21,8 +23,12 @@ const clarifyTTL = 90 * time.Second
|
||||
// raw utterance, chat and system have nothing to fill in. For those a clarify
|
||||
// decision keeps the canned "не поняла" reply — inventing a question for noise
|
||||
// is worse than admitting she missed it.
|
||||
// A reminder wants BOTH what to remind about and when. Subject first: "напомни
|
||||
// в 11" has a time and nothing to say at 11, and a reminder with no subject is
|
||||
// not worth setting. Order here is the order she asks in — she still only asks
|
||||
// about the first one missing.
|
||||
var wantedSlots = map[router.Intent][]dialogue.Slot{
|
||||
router.IntentReminder: {dialogue.SlotTime},
|
||||
router.IntentReminder: {dialogue.SlotText, dialogue.SlotTime},
|
||||
router.IntentFact: {dialogue.SlotKey},
|
||||
router.IntentAct: {dialogue.SlotFn},
|
||||
}
|
||||
@@ -35,7 +41,8 @@ var wantedSlots = map[router.Intent][]dialogue.Slot{
|
||||
// questions, so there is no gender agreement to get wrong; the feminine
|
||||
// self-reference lives in the reply she gives when she drops the request.
|
||||
var clarifyQuestions = map[dialogue.Slot]string{
|
||||
dialogue.SlotTime: "На когда напомнить?",
|
||||
dialogue.SlotTime: "Когда?",
|
||||
dialogue.SlotText: "О чём напомнить?",
|
||||
dialogue.SlotKey: "Что записать?",
|
||||
dialogue.SlotFn: "Что сделать?",
|
||||
}
|
||||
@@ -45,11 +52,55 @@ var clarifyQuestions = map[dialogue.Slot]string{
|
||||
// landed. Feminine self-reference ("поняла"), as everywhere.
|
||||
const clarifyGaveUp = "Прости, я не поняла. Скажи, пожалуйста, по-другому."
|
||||
|
||||
// clarifyExpired — his answer came after the TTL, so the parked request is
|
||||
// already gone. Same tone as clarifyGaveUp, different reason: too much time
|
||||
// clarifyExpiredVariants — his answer came after the TTL, so the parked request
|
||||
// is already gone. Same tone as clarifyGaveUp, different reason: too much time
|
||||
// passed, not "I did not understand". Feminine self-reference ("ждала",
|
||||
// "отпустила"); he is addressed with a plain imperative.
|
||||
const clarifyExpired = "Прости, я слишком долго ждала ответа и отпустила прошлую просьбу. Если она ещё нужна, скажи заново."
|
||||
//
|
||||
// Five phrasings, not one. This is the line he hears whenever he walks off
|
||||
// mid-request, so it is the line that repeats most — and the same sentence every
|
||||
// time is what makes a house assistant sound like a kiosk. They all carry the
|
||||
// same two facts (the old request is gone; say it again if it still matters),
|
||||
// because the wording may vary and the meaning may not.
|
||||
//
|
||||
// Fixed templates rather than model output, for the same reason as
|
||||
// clarifyQuestions: this text has to be right every time, and it is not worth a
|
||||
// generation to say something this small.
|
||||
var clarifyExpiredVariants = []string{
|
||||
"Прости, я слишком долго ждала ответа и отпустила прошлую просьбу. Если она ещё нужна, скажи заново.",
|
||||
"Кажется, прошлая просьба уже не важна — я её отпустила. Если я ошибаюсь, повтори.",
|
||||
"Ты как-то резко замолчал, и я не стала ждать дальше. Если та просьба ещё нужна, скажи заново.",
|
||||
"Я не дождалась ответа и убрала прошлую просьбу. Повтори, если она всё ещё нужна.",
|
||||
"Столько времени прошло, что я отпустила прошлую просьбу. Скажи заново, если она в силе.",
|
||||
}
|
||||
|
||||
// clarifyExpiredLine picks one of them at random.
|
||||
func clarifyExpiredLine() string {
|
||||
return clarifyExpiredVariants[rand.Intn(len(clarifyExpiredVariants))]
|
||||
}
|
||||
|
||||
// isClarifyExpired reports whether s opens with any of the expiry lines. The
|
||||
// notice is glued in front of this turn's reply (see withNotice), so a caller
|
||||
// checking for it has to match a prefix, not the whole string.
|
||||
func isClarifyExpired(s string) bool {
|
||||
for _, v := range clarifyExpiredVariants {
|
||||
if strings.HasPrefix(s, v) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// trimClarifyExpired strips a leading expiry notice, leaving this turn's actual
|
||||
// reply. "" ⇒ the notice was the whole thing.
|
||||
func trimClarifyExpired(s string) string {
|
||||
for _, v := range clarifyExpiredVariants {
|
||||
if strings.HasPrefix(s, v) {
|
||||
return strings.TrimSpace(strings.TrimPrefix(s, v))
|
||||
}
|
||||
}
|
||||
return strings.TrimSpace(s)
|
||||
}
|
||||
|
||||
// clarifyExpiredNotice returns that line when a parked question had just timed
|
||||
// out, and "" when nothing was parked. Call it right after
|
||||
@@ -63,7 +114,7 @@ func (h *reactiveHandler) clarifyExpiredNotice() string {
|
||||
return ""
|
||||
}
|
||||
log.Printf("voice: clarify — parked question expired, telling him and routing the words fresh")
|
||||
return clarifyExpired
|
||||
return clarifyExpiredLine()
|
||||
}
|
||||
|
||||
// withNotice glues the expiry notice in front of this turn's reply. One turn
|
||||
|
||||
@@ -56,10 +56,13 @@ func TestClarifyQuestionForMissingSlot(t *testing.T) {
|
||||
want string
|
||||
asked bool
|
||||
}{
|
||||
{"reminder without a time", clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"), "На когда напомнить?", true},
|
||||
{"reminder without a time", clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"), "Когда?", true},
|
||||
{"fact without a key", clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши"), "Что записать?", true},
|
||||
{"act without a fn", clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это"), "Что сделать?", true},
|
||||
{"reminder that already has a time", clarifyDec(router.IntentReminder, router.Slots{HasTime: true}, "напомни в 11"), "", false},
|
||||
// A time with nothing to say at that time is still half a reminder, so
|
||||
// the subject is what she asks about — not silence.
|
||||
{"reminder that has a time but no subject", clarifyDec(router.IntentReminder, router.Slots{HasTime: true}, "напомни в 11"), "О чём напомнить?", true},
|
||||
{"reminder that has both", clarifyDec(router.IntentReminder, router.Slots{Text: "позвонить маме", HasTime: true}, "напомни в 11 позвонить маме"), "", false},
|
||||
{"chat is never worth a question", clarifyDec(router.IntentChat, router.Slots{Text: "мгм"}, "мгм"), "", false},
|
||||
{"query is never worth a question", clarifyDec(router.IntentQuery, router.Slots{Text: "а"}, "а"), "", false},
|
||||
}
|
||||
@@ -78,7 +81,7 @@ func TestClarifyReminderCompletesOnAnswer(t *testing.T) {
|
||||
h, st, _ := newClarifyHandler(t)
|
||||
|
||||
question, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"))
|
||||
if !asked || question != "На когда напомнить?" {
|
||||
if !asked || question != "Когда?" {
|
||||
t.Fatalf("expected the time question, got %q asked=%v", question, asked)
|
||||
}
|
||||
|
||||
@@ -152,7 +155,7 @@ func TestClarifyAsksThreeTimesThenSaysSo(t *testing.T) {
|
||||
if !handled {
|
||||
t.Fatalf("answer %d must be consumed as an answer", i)
|
||||
}
|
||||
if reply != "На когда напомнить?" {
|
||||
if reply != "Когда?" {
|
||||
t.Fatalf("attempt %d should ask again, got %q", i, reply)
|
||||
}
|
||||
if h.clarifyStore.Get(voiceDialogueID, h.now()) == nil {
|
||||
@@ -304,17 +307,17 @@ func TestClarifyExpiryIsAnnouncedAndWordsStillRoute(t *testing.T) {
|
||||
*now = now.Add(clarifyTTL + time.Second)
|
||||
|
||||
reply := h.handleText(ctx, "как дела")
|
||||
if !strings.HasPrefix(reply, clarifyExpired) {
|
||||
if !isClarifyExpired(reply) {
|
||||
t.Fatalf("expired question must be announced first, got %q", reply)
|
||||
}
|
||||
if strings.TrimSpace(strings.TrimPrefix(reply, clarifyExpired)) == "" {
|
||||
if trimClarifyExpired(reply) == "" {
|
||||
t.Fatalf("the new words must still be answered, got only the notice: %q", reply)
|
||||
}
|
||||
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
|
||||
t.Fatal("the expired question must be gone")
|
||||
}
|
||||
// The notice is said once, not on every later utterance.
|
||||
if reply := h.handleText(ctx, "как дела"); strings.Contains(reply, clarifyExpired) {
|
||||
if reply := h.handleText(ctx, "как дела"); isClarifyExpired(reply) {
|
||||
t.Fatalf("notice repeated on a later turn: %q", reply)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -34,7 +34,7 @@ func newLLMReplier(c completer, block func() string) *llmReplier {
|
||||
return &llmReplier{c: c, stub: voice.NewStubReplier(), block: block}
|
||||
}
|
||||
|
||||
const replySystem = `Ты — Maven, домашняя ассистентка (о себе — в женском роде). Владелец — мужчина, говоришь с ним на "ты", в единственном числе; никогда не "вы"/"ваш" и не "он"/"его". Подтверди действие РОВНО ОДНИМ коротким предложением (≤120 символов), тепло и по-русски. Не задавай вопросов, не повторяй слова, не добавляй ничего после точки. Отвечай ТОЛЬКО одним объектом JSON с полями "response" (текст) и "mood" (ровно одно из: neutral, happy, thinking, tired, confused).
|
||||
const replySystem = `Ты — Maven, домашняя ассистентка (о себе — в женском роде). Владелец — мужчина, говоришь с ним на "ты", в единственном числе; никогда не "вы"/"ваш" и не "он"/"его". Подтверди действие РОВНО ОДНИМ коротким предложением (≤120 символов), по-русски, спокойно и без официальных формулировок. Не задавай вопросов, не повторяй слова, не добавляй ничего после точки. Отвечай ТОЛЬКО одним объектом JSON с полями "response" (текст) и "mood" (ровно одно из: neutral, happy, thinking, tired, confused).
|
||||
Пример: {"response": "Записала, что ты выпил стакан воды.", "mood": "neutral"}
|
||||
Никогда не пиши "..." в поле response.`
|
||||
|
||||
|
||||
@@ -0,0 +1,116 @@
|
||||
// Package kiwix reads a local Kiwix server (offline Wikipedia and friends).
|
||||
//
|
||||
// Why: the resident model is a 0.8B and invents facts. Letting her read a local
|
||||
// article snippet beats letting her recall. Nothing here talks to the internet;
|
||||
// the Kiwix server is on the same box.
|
||||
//
|
||||
// This is search only. Full articles are ~100KB of HTML, far too big for a 4096
|
||||
// token context, so the unit of context is the search snippet (~500 chars).
|
||||
package kiwix
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/xml"
|
||||
"fmt"
|
||||
"html"
|
||||
"io"
|
||||
"net/http"
|
||||
"net/url"
|
||||
"regexp"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
// Result is one search hit.
|
||||
type Result struct {
|
||||
Title string // article title, e.g. "Rayleigh scattering"
|
||||
Path string // e.g. /content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering
|
||||
Snippet string // plain text, tags stripped, entities decoded
|
||||
WordCount int // 0 if the server did not say
|
||||
}
|
||||
|
||||
// Client is a Kiwix HTTP client. Boring on purpose: no retries, no cache.
|
||||
type Client struct {
|
||||
base string
|
||||
http *http.Client
|
||||
}
|
||||
|
||||
// New makes a client for a Kiwix base URL like http://127.0.0.1:8034.
|
||||
func New(baseURL string) *Client {
|
||||
return &Client{
|
||||
base: strings.TrimRight(baseURL, "/"),
|
||||
http: &http.Client{Timeout: 10 * time.Second},
|
||||
}
|
||||
}
|
||||
|
||||
// Search runs a keyword search in one ZIM (book) and returns up to limit hits.
|
||||
//
|
||||
// Ranking is keyword based, not semantic: "Rayleigh scattering" finds the right
|
||||
// article, "why is the sky blue" finds a TV episode. Pass keywords, not questions.
|
||||
func (c *Client) Search(ctx context.Context, pattern, book string, limit int) ([]Result, error) {
|
||||
if limit <= 0 {
|
||||
limit = 5
|
||||
}
|
||||
q := url.Values{}
|
||||
q.Set("pattern", pattern)
|
||||
q.Set("books.name", book)
|
||||
q.Set("format", "xml")
|
||||
q.Set("pageLength", strconv.Itoa(limit))
|
||||
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodGet, c.base+"/search?"+q.Encode(), nil)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
resp, err := c.http.Do(req)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
return nil, fmt.Errorf("kiwix search: http %d", resp.StatusCode)
|
||||
}
|
||||
return ParseSearchRSS(resp.Body)
|
||||
}
|
||||
|
||||
// rss mirrors just the bits of the RSS 2.0 reply we use.
|
||||
type rss struct {
|
||||
Items []struct {
|
||||
Title string `xml:"title"`
|
||||
Link string `xml:"link"`
|
||||
// innerxml keeps the <b> match markers so we can strip them ourselves.
|
||||
Description struct {
|
||||
Inner string `xml:",innerxml"`
|
||||
} `xml:"description"`
|
||||
WordCount string `xml:"wordCount"`
|
||||
} `xml:"channel>item"`
|
||||
}
|
||||
|
||||
var tagRE = regexp.MustCompile(`<[^>]*>`)
|
||||
|
||||
// ParseSearchRSS turns a Kiwix search reply into results. Exported so the parser
|
||||
// is testable from a captured response, with no server running.
|
||||
func ParseSearchRSS(r io.Reader) ([]Result, error) {
|
||||
var doc rss
|
||||
if err := xml.NewDecoder(r).Decode(&doc); err != nil {
|
||||
return nil, fmt.Errorf("kiwix search: bad xml: %w", err)
|
||||
}
|
||||
out := make([]Result, 0, len(doc.Items))
|
||||
for _, it := range doc.Items {
|
||||
n, _ := strconv.Atoi(strings.ReplaceAll(it.WordCount, ",", ""))
|
||||
out = append(out, Result{
|
||||
Title: strings.TrimSpace(it.Title),
|
||||
Path: strings.TrimSpace(it.Link),
|
||||
Snippet: plainText(it.Description.Inner),
|
||||
WordCount: n,
|
||||
})
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
// plainText drops markup and decodes entities, leaving text a model can read.
|
||||
func plainText(s string) string {
|
||||
s = tagRE.ReplaceAllString(s, "")
|
||||
s = html.UnescapeString(s)
|
||||
return strings.TrimSpace(strings.Join(strings.Fields(s), " "))
|
||||
}
|
||||
@@ -0,0 +1,90 @@
|
||||
package kiwix
|
||||
|
||||
import (
|
||||
"context"
|
||||
"os"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// A real reply from the live server, trimmed to two items.
|
||||
const sampleRSS = `<?xml version="1.0" encoding="UTF-8"?>
|
||||
<rss version="2.0" xmlns:opensearch="http://a9.com/-/spec/opensearch/1.1/">
|
||||
<channel>
|
||||
<title>Search: Rayleigh scattering</title>
|
||||
<opensearch:totalResults>800</opensearch:totalResults>
|
||||
<item>
|
||||
<title>Rayleigh scattering</title>
|
||||
<link>/content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering</link>
|
||||
<description><b>Rayleigh</b> scattering causes the blue color of the sky & yellow colors near the Sun.[1]</description>
|
||||
<book><title>Wikipedia</title></book>
|
||||
<wordCount>2,818</wordCount>
|
||||
</item>
|
||||
<item>
|
||||
<title>Hyper–Rayleigh scattering</title>
|
||||
<link>/content/wikipedia_en_all_maxi_2026-02/Hyper%E2%80%93Rayleigh_scattering</link>
|
||||
<description>...<b>Rayleigh</b> scattering" is a nonlinear optical counterpart.</description>
|
||||
<book><title>Wikipedia</title></book>
|
||||
<wordCount>914</wordCount>
|
||||
</item>
|
||||
</channel>
|
||||
</rss>`
|
||||
|
||||
func TestParseSearchRSS(t *testing.T) {
|
||||
got, err := ParseSearchRSS(strings.NewReader(sampleRSS))
|
||||
if err != nil {
|
||||
t.Fatalf("parse: %v", err)
|
||||
}
|
||||
if len(got) != 2 {
|
||||
t.Fatalf("want 2 results, got %d", len(got))
|
||||
}
|
||||
if got[0].Title != "Rayleigh scattering" {
|
||||
t.Errorf("title = %q", got[0].Title)
|
||||
}
|
||||
if got[0].Path != "/content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering" {
|
||||
t.Errorf("path = %q", got[0].Path)
|
||||
}
|
||||
if got[0].WordCount != 2818 {
|
||||
t.Errorf("wordCount = %d", got[0].WordCount)
|
||||
}
|
||||
want := "Rayleigh scattering causes the blue color of the sky & yellow colors near the Sun.[1]"
|
||||
if got[0].Snippet != want {
|
||||
t.Errorf("snippet = %q, want %q", got[0].Snippet, want)
|
||||
}
|
||||
if strings.Contains(got[1].Snippet, "<b>") {
|
||||
t.Errorf("second snippet still has tags: %q", got[1].Snippet)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseSearchRSSBadXML(t *testing.T) {
|
||||
if _, err := ParseSearchRSS(strings.NewReader("not xml at all")); err == nil {
|
||||
t.Fatal("want an error on junk input")
|
||||
}
|
||||
}
|
||||
|
||||
// Opt-in: needs a live Kiwix server. CI has none.
|
||||
// MAVEN_KIWIX_URL=http://127.0.0.1:8034 no_proxy=127.0.0.1,localhost go test -run Retrieval -v ./internal/kiwix/
|
||||
func TestRetrievalEval(t *testing.T) {
|
||||
base := os.Getenv("MAVEN_KIWIX_URL")
|
||||
if base == "" {
|
||||
t.Skip("set MAVEN_KIWIX_URL to run the retrieval eval")
|
||||
}
|
||||
noProxyLoopback(t)
|
||||
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
|
||||
defer cancel()
|
||||
|
||||
rep, err := RunRetrievalEval(ctx, New(base), 5)
|
||||
if err != nil {
|
||||
t.Fatalf("eval: %v", err)
|
||||
}
|
||||
// No pass bar on purpose: the number is the finding.
|
||||
t.Log("\n" + rep.String() + rep.Detail())
|
||||
}
|
||||
|
||||
// noProxyLoopback stops the box's SOCKS bridge from eating loopback requests.
|
||||
func noProxyLoopback(t *testing.T) {
|
||||
t.Setenv("no_proxy", "127.0.0.1,localhost")
|
||||
t.Setenv("NO_PROXY", "127.0.0.1,localhost")
|
||||
}
|
||||
@@ -0,0 +1,63 @@
|
||||
{
|
||||
"name": "kiwix-knowledge-v1",
|
||||
"book": "wikipedia_en_all_maxi_2026-02",
|
||||
"note": "The 9 knowledge cases from internal/phraser/eval/talk_v1.json. Queries are hand-written English keywords on purpose: Kiwix ranks by keyword, not meaning, so a natural question fails. Writing them by hand separates 'retrieval is broken' from 'the model writes bad queries'.",
|
||||
"cases": [
|
||||
{
|
||||
"id": "know-sky-blue",
|
||||
"question": "почему небо синее?",
|
||||
"query": "Rayleigh scattering sky blue",
|
||||
"want_titles": ["Rayleigh scattering", "Diffuse sky radiation"]
|
||||
},
|
||||
{
|
||||
"id": "know-boil-egg",
|
||||
"question": "сколько варить яйцо вкрутую?",
|
||||
"query": "boiled egg cooking",
|
||||
"want_titles": ["Boiled egg", "Egg as food"]
|
||||
},
|
||||
{
|
||||
"id": "know-ssd-vs-hdd",
|
||||
"question": "чем ssd отличается от hdd?",
|
||||
"query": "solid-state drive",
|
||||
"want_titles": ["Solid-state drive", "Hard disk drive"]
|
||||
},
|
||||
{
|
||||
"id": "know-cat-purr",
|
||||
"question": "почему кошки мурчат?",
|
||||
"query": "cat purr",
|
||||
"want_titles": ["Purr", "Cat communication"]
|
||||
},
|
||||
{
|
||||
"id": "know-hiccups",
|
||||
"question": "как быстро избавиться от икоты?",
|
||||
"query": "hiccup",
|
||||
"want_titles": ["Hiccup"]
|
||||
},
|
||||
{
|
||||
"id": "know-polite-form",
|
||||
"question": "не могли бы вы объяснить, что такое vpn?",
|
||||
"query": "virtual private network",
|
||||
"want_titles": ["Virtual private network"]
|
||||
},
|
||||
{
|
||||
"id": "know-dont-know",
|
||||
"question": "как зовут моего соседа снизу?",
|
||||
"query": "name of my downstairs neighbour",
|
||||
"want_titles": [],
|
||||
"expect_miss": true,
|
||||
"note": "Unanswerable by design. Retrieval SHOULD find nothing useful. Counted as a hit only when nothing relevant comes back."
|
||||
},
|
||||
{
|
||||
"id": "know-water-per-day",
|
||||
"question": "сколько воды в день надо пить?",
|
||||
"query": "human daily water requirement drinking",
|
||||
"want_titles": ["Drinking water", "Water", "Dehydration", "Hydration"]
|
||||
},
|
||||
{
|
||||
"id": "know-thunder-delay",
|
||||
"question": "почему гром слышно позже молнии?",
|
||||
"query": "thunder speed of sound lightning",
|
||||
"want_titles": ["Thunder", "Lightning"]
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,135 @@
|
||||
package kiwix
|
||||
|
||||
// This scores retrieval alone: no LLM. For each general-knowledge question we
|
||||
// hand-write English keywords and ask whether the article that would answer it
|
||||
// comes back in the top N hits. If this score is low, reading Wikipedia cannot
|
||||
// help the model no matter how good the prompt is.
|
||||
//
|
||||
// The unanswerable case (know-dont-know) is not scored. Whether the junk it
|
||||
// returns is "nothing useful" is a human judgement, so the report just prints
|
||||
// the titles and leaves the score to the 8 answerable cases.
|
||||
|
||||
import (
|
||||
"context"
|
||||
_ "embed"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"strings"
|
||||
)
|
||||
|
||||
//go:embed knowledge_v1.json
|
||||
var knowledgeFixtureJSON []byte
|
||||
|
||||
// EvalCase — one question with hand-written keywords.
|
||||
type EvalCase struct {
|
||||
ID string `json:"id"`
|
||||
Question string `json:"question"`
|
||||
Query string `json:"query"`
|
||||
WantTitles []string `json:"want_titles"`
|
||||
ExpectMiss bool `json:"expect_miss"`
|
||||
}
|
||||
|
||||
type fixture struct {
|
||||
Name string `json:"name"`
|
||||
Book string `json:"book"`
|
||||
Cases []EvalCase `json:"cases"`
|
||||
}
|
||||
|
||||
// Outcome — what one case retrieved.
|
||||
type Outcome struct {
|
||||
Case EvalCase
|
||||
Titles []string // titles of the top N hits, in rank order
|
||||
Rank int // 1-based rank of the first wanted title, 0 if none
|
||||
Err error
|
||||
}
|
||||
|
||||
// Hit is true when a wanted title came back.
|
||||
func (o Outcome) Hit() bool { return o.Rank > 0 }
|
||||
|
||||
// Report — the score plus per-case detail.
|
||||
type Report struct {
|
||||
Name string
|
||||
Book string
|
||||
TopN int
|
||||
Scored int // answerable cases
|
||||
Hits int
|
||||
Errors int
|
||||
Outcomes []Outcome
|
||||
}
|
||||
|
||||
// Accuracy over the answerable cases.
|
||||
func (r Report) Accuracy() float64 {
|
||||
if r.Scored == 0 {
|
||||
return 0
|
||||
}
|
||||
return float64(r.Hits) / float64(r.Scored)
|
||||
}
|
||||
|
||||
// RunRetrievalEval searches for every fixture case.
|
||||
func RunRetrievalEval(ctx context.Context, c *Client, topN int) (Report, error) {
|
||||
var f fixture
|
||||
if err := json.Unmarshal(knowledgeFixtureJSON, &f); err != nil {
|
||||
return Report{}, err
|
||||
}
|
||||
rep := Report{Name: f.Name, Book: f.Book, TopN: topN}
|
||||
for _, cs := range f.Cases {
|
||||
res, err := c.Search(ctx, cs.Query, f.Book, topN)
|
||||
o := Outcome{Case: cs, Err: err}
|
||||
if err != nil {
|
||||
rep.Errors++
|
||||
}
|
||||
for i, hit := range res {
|
||||
o.Titles = append(o.Titles, hit.Title)
|
||||
if o.Rank == 0 && matches(cs.WantTitles, hit.Title) {
|
||||
o.Rank = i + 1
|
||||
}
|
||||
}
|
||||
if !cs.ExpectMiss {
|
||||
rep.Scored++
|
||||
if o.Hit() {
|
||||
rep.Hits++
|
||||
}
|
||||
}
|
||||
rep.Outcomes = append(rep.Outcomes, o)
|
||||
}
|
||||
return rep, nil
|
||||
}
|
||||
|
||||
func matches(want []string, title string) bool {
|
||||
for _, w := range want {
|
||||
if strings.EqualFold(strings.TrimSpace(title), w) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// String — the headline number.
|
||||
func (r Report) String() string {
|
||||
var b strings.Builder
|
||||
fmt.Fprintf(&b, "%s: %d/%d answerable questions retrieve a wanted article in top %d (%.1f%%), %d errors\n",
|
||||
r.Name, r.Hits, r.Scored, r.TopN, 100*r.Accuracy(), r.Errors)
|
||||
fmt.Fprintf(&b, " book: %s\n", r.Book)
|
||||
return b.String()
|
||||
}
|
||||
|
||||
// Detail — per case: what was asked, what was searched, what came back.
|
||||
func (r Report) Detail() string {
|
||||
var b strings.Builder
|
||||
for _, o := range r.Outcomes {
|
||||
mark := "MISS"
|
||||
switch {
|
||||
case o.Case.ExpectMiss:
|
||||
mark = "n/a "
|
||||
case o.Hit():
|
||||
mark = fmt.Sprintf("hit@%d", o.Rank)
|
||||
}
|
||||
fmt.Fprintf(&b, " %-6s %-20s q=%q\n", mark, o.Case.ID, o.Case.Query)
|
||||
if o.Err != nil {
|
||||
fmt.Fprintf(&b, " error: %v\n", o.Err)
|
||||
continue
|
||||
}
|
||||
fmt.Fprintf(&b, " got: %s\n", strings.Join(o.Titles, " | "))
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
@@ -0,0 +1,183 @@
|
||||
package kiwix
|
||||
|
||||
// Turning a Russian question into an English Kiwix search.
|
||||
//
|
||||
// Kiwix ranks by keyword, not by meaning. "why is the sky blue" returns a TV
|
||||
// episode; "Rayleigh scattering sky blue" returns the right article. So the
|
||||
// model's job here is NOT translation — it is naming the English article the
|
||||
// answer lives in.
|
||||
//
|
||||
// The output space is a handful of words, so it is worth locking down hard: a
|
||||
// GBNF grammar for the shape, a tiny token cap, and a cleanup pass that throws
|
||||
// away anything odd rather than handing junk to Kiwix.
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"strings"
|
||||
"unicode"
|
||||
|
||||
"github.com/kami/maven/internal/llm"
|
||||
)
|
||||
|
||||
// Completer — the LLM seam, so tests can fake it. *llm.Client satisfies it.
|
||||
type Completer interface {
|
||||
Complete(ctx context.Context, r llm.Req) (string, error)
|
||||
}
|
||||
|
||||
// queryGrammar — one JSON object holding 1..6 keyword words. Latin letters,
|
||||
// digits and hyphens only, so the model physically cannot answer the question
|
||||
// or reply in Russian.
|
||||
//
|
||||
// Why the JSON wrapper: this model always thinks out loud and this llama-server
|
||||
// build ignores the thinking switch (see ROUTING-EVAL-31-07-2026.md). A bare
|
||||
// word-list grammar just captured the reasoning — every case came back as
|
||||
// "Let me analyze this request carefully". Demanding JSON, like routeGrammar and
|
||||
// responseGrammar already do, gives the reasoning nowhere to go.
|
||||
const queryGrammar = `
|
||||
root ::= "{" ws "\"query\"" ws ":" ws "\"" word (" " word){0,5} "\"" ws "}"
|
||||
word ::= [A-Za-z0-9] [A-Za-z0-9-]{0,23}
|
||||
ws ::= [ \t\n]*
|
||||
`
|
||||
|
||||
// rewriteSystem — asks for search keywords, not an answer and not a translation.
|
||||
const rewriteSystem = `You turn a question into a search query for English Wikipedia.
|
||||
|
||||
Rules:
|
||||
- Output ONLY English search keywords. Never an answer, never an explanation.
|
||||
- Do NOT translate the sentence. Name the thing the answer is about.
|
||||
- The output must be a noun phrase, like a Wikipedia article title.
|
||||
- Never use question words: no why, how, what, when, which, "how much",
|
||||
"how long", "how to", "vs", "reason", "difference".
|
||||
- 2 to 4 words.
|
||||
|
||||
Reply with JSON: {"query":"<keywords>"}
|
||||
|
||||
Good:
|
||||
"почему листья желтеют осенью?" -> {"query":"leaf senescence autumn"}
|
||||
"как работает микроволновка?" -> {"query":"microwave oven"}
|
||||
"не могли бы вы объяснить, что такое блокчейн?" -> {"query":"blockchain"}
|
||||
"сколько живут собаки?" -> {"query":"dog lifespan"}
|
||||
"как избавиться от комаров в квартире?" -> {"query":"mosquito control"}
|
||||
"чем чай отличается от кофе?" -> {"query":"tea"}
|
||||
|
||||
Only JSON, no explanation.`
|
||||
|
||||
// maxQueryTokens — the output is a few words plus the JSON wrapper. A tight cap
|
||||
// is the cheapest guard against the model rambling into an answer.
|
||||
const maxQueryTokens = 32
|
||||
|
||||
// Rewriter asks the resident model for English search keywords.
|
||||
type Rewriter struct{ c Completer }
|
||||
|
||||
func NewRewriter(c Completer) *Rewriter { return &Rewriter{c: c} }
|
||||
|
||||
// Rewrite returns English keywords for a question in any language.
|
||||
// It errors rather than returning something Kiwix should not see.
|
||||
func (r *Rewriter) Rewrite(ctx context.Context, question string) (string, error) {
|
||||
raw, err := r.c.Complete(ctx, llm.Req{
|
||||
System: rewriteSystem,
|
||||
User: strings.TrimSpace(question),
|
||||
Grammar: queryGrammar,
|
||||
MaxTokens: maxQueryTokens,
|
||||
RepeatPenalty: 1.15,
|
||||
})
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
return CleanQuery(unwrapJSON(raw))
|
||||
}
|
||||
|
||||
// unwrapJSON pulls the query out of {"query":"..."}. If the reply is not that
|
||||
// shape it is returned as-is, and CleanQuery decides whether it is usable.
|
||||
func unwrapJSON(raw string) string {
|
||||
s := strings.TrimSpace(raw)
|
||||
if !strings.HasPrefix(s, "{") {
|
||||
return s
|
||||
}
|
||||
var got struct{ Query string }
|
||||
if err := json.Unmarshal([]byte(s), &got); err != nil {
|
||||
return s
|
||||
}
|
||||
return got.Query
|
||||
}
|
||||
|
||||
// maxQueryWords matches the grammar's bound. Anything longer is prose.
|
||||
const maxQueryWords = 6
|
||||
|
||||
// CleanQuery checks and tidies whatever the model produced. The grammar makes
|
||||
// bad output unlikely, not impossible (a server without grammar support, a
|
||||
// different model), so this is the real gate in front of Kiwix.
|
||||
//
|
||||
// Exported so it can be tested without a model.
|
||||
func CleanQuery(raw string) (string, error) {
|
||||
s := strings.TrimSpace(raw)
|
||||
// Models like to wrap answers in quotes. Drop surrounding ones.
|
||||
s = strings.Trim(s, "\"'`")
|
||||
// Keep the first line only: everything after it is prose.
|
||||
if i := strings.IndexAny(s, "\r\n"); i >= 0 {
|
||||
s = s[:i]
|
||||
}
|
||||
// Keep letters, digits, spaces and hyphens; anything else becomes a space.
|
||||
var b strings.Builder
|
||||
for _, ru := range s {
|
||||
switch {
|
||||
case unicode.IsLetter(ru) || unicode.IsDigit(ru) || ru == '-':
|
||||
b.WriteRune(ru)
|
||||
default:
|
||||
b.WriteRune(' ')
|
||||
}
|
||||
}
|
||||
words := strings.Fields(b.String())
|
||||
if len(words) == 0 {
|
||||
return "", fmt.Errorf("kiwix rewrite: empty query")
|
||||
}
|
||||
if len(words) > maxQueryWords {
|
||||
return "", fmt.Errorf("kiwix rewrite: %d words, want at most %d (looks like prose)", len(words), maxQueryWords)
|
||||
}
|
||||
words = dropStopWords(words)
|
||||
out := strings.Join(words, " ")
|
||||
// The ZIMs are English. Non-Latin letters mean the model ignored the ask.
|
||||
for _, ru := range out {
|
||||
if unicode.IsLetter(ru) && !isLatin(ru) {
|
||||
return "", fmt.Errorf("kiwix rewrite: query is not English: %q", out)
|
||||
}
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
// stopWords — question words and filler. The model keeps writing question-shaped
|
||||
// queries ("why is the sky blue", "how much water to drink daily") no matter how
|
||||
// the prompt is worded, and Kiwix ranks on every word, so those words drag in
|
||||
// song and episode titles. Dropping them in code is not a style preference: a
|
||||
// keyword ranker gets nothing from them.
|
||||
var stopWords = map[string]bool{
|
||||
"a": true, "an": true, "the": true, "is": true, "are": true, "was": true,
|
||||
"do": true, "does": true, "did": true, "to": true, "of": true, "in": true,
|
||||
"on": true, "for": true, "and": true, "or": true, "my": true, "me": true,
|
||||
"i": true, "it": true, "its": true, "be": true, "been": true, "get": true,
|
||||
"how": true, "why": true, "what": true, "when": true, "which": true,
|
||||
"who": true, "where": true, "much": true, "many": true, "long": true,
|
||||
"vs": true, "than": true, "rid": true, "from": true, "about": true,
|
||||
}
|
||||
|
||||
// dropStopWords removes filler, but never everything: if the query was nothing
|
||||
// but stop words there is nothing better to search, so the original is kept and
|
||||
// the caller sees whatever Kiwix makes of it.
|
||||
func dropStopWords(words []string) []string {
|
||||
kept := make([]string, 0, len(words))
|
||||
for _, w := range words {
|
||||
if !stopWords[strings.ToLower(w)] {
|
||||
kept = append(kept, w)
|
||||
}
|
||||
}
|
||||
if len(kept) == 0 {
|
||||
return words
|
||||
}
|
||||
return kept
|
||||
}
|
||||
|
||||
func isLatin(ru rune) bool {
|
||||
return (ru >= 'a' && ru <= 'z') || (ru >= 'A' && ru <= 'Z')
|
||||
}
|
||||
@@ -0,0 +1,96 @@
|
||||
package kiwix
|
||||
|
||||
// End-to-end score: Russian question -> model rewrite -> Kiwix search -> did a
|
||||
// wanted article come back. Same 9 cases as the retrieval eval, so the two
|
||||
// numbers are directly comparable: retrieval with hand-written keywords is the
|
||||
// ceiling, this is what the model actually reaches.
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// RewriteOutcome — one case, end to end.
|
||||
type RewriteOutcome struct {
|
||||
Outcome
|
||||
ModelQuery string // what the model asked for ("" if it failed)
|
||||
RewriteErr error
|
||||
}
|
||||
|
||||
// RunRewriteEval rewrites every question with the model, then searches.
|
||||
func RunRewriteEval(ctx context.Context, c *Client, rw *Rewriter, topN int) (RewriteReport, error) {
|
||||
var f fixture
|
||||
if err := json.Unmarshal(knowledgeFixtureJSON, &f); err != nil {
|
||||
return RewriteReport{}, err
|
||||
}
|
||||
rep := RewriteReport{Report: Report{Name: f.Name + "-rewrite", Book: f.Book, TopN: topN}}
|
||||
for _, cs := range f.Cases {
|
||||
out := RewriteOutcome{Outcome: Outcome{Case: cs}}
|
||||
q, err := rw.Rewrite(ctx, cs.Question)
|
||||
out.ModelQuery, out.RewriteErr = q, err
|
||||
if err == nil {
|
||||
res, serr := c.Search(ctx, q, f.Book, topN)
|
||||
out.Err = serr
|
||||
for i, hit := range res {
|
||||
out.Titles = append(out.Titles, hit.Title)
|
||||
if out.Rank == 0 && matches(cs.WantTitles, hit.Title) {
|
||||
out.Rank = i + 1
|
||||
}
|
||||
}
|
||||
}
|
||||
if out.RewriteErr != nil || out.Err != nil {
|
||||
rep.Errors++
|
||||
}
|
||||
if !cs.ExpectMiss {
|
||||
rep.Scored++
|
||||
if out.Hit() {
|
||||
rep.Hits++
|
||||
}
|
||||
}
|
||||
rep.Cases = append(rep.Cases, out)
|
||||
}
|
||||
return rep, nil
|
||||
}
|
||||
|
||||
// RewriteReport — the score plus per-case detail.
|
||||
type RewriteReport struct {
|
||||
Report
|
||||
Cases []RewriteOutcome
|
||||
}
|
||||
|
||||
// String — the headline number.
|
||||
func (r RewriteReport) String() string {
|
||||
return fmt.Sprintf("%s: %d/%d answerable questions retrieve a wanted article in top %d (%.1f%%), %d errors\n book: %s\n",
|
||||
r.Name, r.Hits, r.Scored, r.TopN, 100*r.Accuracy(), r.Errors, r.Book)
|
||||
}
|
||||
|
||||
// Detail — per case: hand-written query next to the model's, and what came back.
|
||||
// The point is seeing WHERE the model's phrasing differs, not just the score.
|
||||
func (r RewriteReport) Detail() string {
|
||||
var b strings.Builder
|
||||
for _, o := range r.Cases {
|
||||
mark := "MISS"
|
||||
switch {
|
||||
case o.Case.ExpectMiss:
|
||||
mark = "n/a "
|
||||
case o.Hit():
|
||||
mark = fmt.Sprintf("hit@%d", o.Rank)
|
||||
}
|
||||
fmt.Fprintf(&b, " %-6s %-20s\n", mark, o.Case.ID)
|
||||
fmt.Fprintf(&b, " asked: %s\n", o.Case.Question)
|
||||
fmt.Fprintf(&b, " hand: %q\n", o.Case.Query)
|
||||
fmt.Fprintf(&b, " model: %q\n", o.ModelQuery)
|
||||
if o.RewriteErr != nil {
|
||||
fmt.Fprintf(&b, " rewrite rejected: %v\n", o.RewriteErr)
|
||||
continue
|
||||
}
|
||||
if o.Err != nil {
|
||||
fmt.Fprintf(&b, " search error: %v\n", o.Err)
|
||||
continue
|
||||
}
|
||||
fmt.Fprintf(&b, " got: %s\n", strings.Join(o.Titles, " | "))
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
@@ -0,0 +1,33 @@
|
||||
package kiwix
|
||||
|
||||
import (
|
||||
"context"
|
||||
"os"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/llm"
|
||||
)
|
||||
|
||||
// Opt-in: needs a live Kiwix server AND a live llama-server.
|
||||
// MAVEN_KIWIX_URL=http://127.0.0.1:8034 MAVEN_LLM_URL=http://127.0.0.1:18099 \
|
||||
//
|
||||
// no_proxy=127.0.0.1,localhost go test -run RewriteEval -v ./internal/kiwix/
|
||||
func TestRewriteEval(t *testing.T) {
|
||||
kbase, lbase := os.Getenv("MAVEN_KIWIX_URL"), os.Getenv("MAVEN_LLM_URL")
|
||||
if kbase == "" || lbase == "" {
|
||||
t.Skip("set MAVEN_KIWIX_URL and MAVEN_LLM_URL to run the rewrite eval")
|
||||
}
|
||||
noProxyLoopback(t)
|
||||
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 15*time.Minute)
|
||||
defer cancel()
|
||||
|
||||
rw := NewRewriter(llm.New(lbase, 3*time.Minute))
|
||||
rep, err := RunRewriteEval(ctx, New(kbase), rw, 5)
|
||||
if err != nil {
|
||||
t.Fatalf("eval: %v", err)
|
||||
}
|
||||
// No pass bar on purpose: the number is the finding.
|
||||
t.Log("\n" + rep.String() + rep.Detail())
|
||||
}
|
||||
@@ -0,0 +1,105 @@
|
||||
package kiwix
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
|
||||
"github.com/kami/maven/internal/llm"
|
||||
)
|
||||
|
||||
// Bad model output must never reach Kiwix. No model needed for this.
|
||||
func TestCleanQueryRejectsJunk(t *testing.T) {
|
||||
bad := []struct{ name, raw string }{
|
||||
{"empty", ""},
|
||||
{"blank", " \n "},
|
||||
{"russian came back", "почему небо синее"},
|
||||
{"mixed russian", "sky синее scattering"},
|
||||
{"full sentence", "The sky looks blue because of the scattering of sunlight by air molecules"},
|
||||
{"prose with quotes", `Sure! Here is a good search query: "Rayleigh scattering", which explains it.`},
|
||||
}
|
||||
for _, c := range bad {
|
||||
if got, err := CleanQuery(c.raw); err == nil {
|
||||
t.Errorf("%s: want rejection, got %q", c.name, got)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestCleanQueryCleans(t *testing.T) {
|
||||
ok := []struct{ raw, want string }{
|
||||
{"Rayleigh scattering sky", "Rayleigh scattering sky"},
|
||||
{" boiled egg cooking \n", "boiled egg cooking"},
|
||||
{`"virtual private network"`, "virtual private network"},
|
||||
{"solid-state drive", "solid-state drive"},
|
||||
{"cat purr.", "cat purr"},
|
||||
{"hiccup\nAlso: hiccough", "hiccup"},
|
||||
// Question words are filler to a keyword ranker, so they go.
|
||||
{"why is the sky blue", "sky blue"},
|
||||
{"how much water to drink daily", "water drink daily"},
|
||||
{"SSD vs HDD comparison", "SSD HDD comparison"},
|
||||
// Nothing but filler: keep it rather than return nothing.
|
||||
{"what is it", "what is it"},
|
||||
}
|
||||
for _, c := range ok {
|
||||
got, err := CleanQuery(c.raw)
|
||||
if err != nil {
|
||||
t.Errorf("%q: %v", c.raw, err)
|
||||
continue
|
||||
}
|
||||
if got != c.want {
|
||||
t.Errorf("%q -> %q, want %q", c.raw, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
type fakeCompleter struct {
|
||||
out string
|
||||
req llm.Req
|
||||
}
|
||||
|
||||
func (f *fakeCompleter) Complete(_ context.Context, r llm.Req) (string, error) {
|
||||
f.req = r
|
||||
return f.out, nil
|
||||
}
|
||||
|
||||
func TestRewriteConstrainsTheCall(t *testing.T) {
|
||||
f := &fakeCompleter{out: `{"query":"Rayleigh scattering sky"}`}
|
||||
got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?")
|
||||
if err != nil {
|
||||
t.Fatalf("rewrite: %v", err)
|
||||
}
|
||||
if got != "Rayleigh scattering sky" {
|
||||
t.Errorf("query = %q", got)
|
||||
}
|
||||
if f.req.Grammar == "" {
|
||||
t.Error("no grammar sent")
|
||||
}
|
||||
if f.req.MaxTokens == 0 || f.req.MaxTokens > 32 {
|
||||
t.Errorf("max_tokens = %d, want a small cap", f.req.MaxTokens)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRewriteRejectsBadModelOutput(t *testing.T) {
|
||||
bad := []string{
|
||||
`{"query":"почему небо синее"}`, // never translated
|
||||
`{"query":""}`, // empty
|
||||
`{"query":"the sky is blue because sunlight is scattered by air"}`, // an answer
|
||||
// Note: a SHORT English prose fragment ("Let me analyze this request")
|
||||
// is under the word cap and cannot be caught here. The grammar is what
|
||||
// stops that one.
|
||||
}
|
||||
for _, out := range bad {
|
||||
f := &fakeCompleter{out: out}
|
||||
if got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?"); err == nil {
|
||||
t.Errorf("%s: want rejection, got %q", out, got)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A reply that is not the JSON shape but is still usable keywords should pass.
|
||||
func TestRewriteFallsBackToPlainText(t *testing.T) {
|
||||
f := &fakeCompleter{out: "Rayleigh scattering sky"}
|
||||
got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?")
|
||||
if err != nil || got != "Rayleigh scattering sky" {
|
||||
t.Errorf("got %q, %v", got, err)
|
||||
}
|
||||
}
|
||||
@@ -551,12 +551,21 @@ func (p *LLMPhraser) chatWithSystem(ctx context.Context, system, user string, ma
|
||||
// Russian only, feminine self-reference, second person masculine (the owner is
|
||||
// a man). She talks TO him, informally, singular — never "вы", never "он".
|
||||
// One short sentence — the nudge is spoken aloud.
|
||||
//
|
||||
// What the ban on обращения forbids is pet names ("дорогой", "милый"), not his
|
||||
// name: "Ками, ноутбук на трёх процентах" is exactly how she talks, and the
|
||||
// unqualified word read as forbidding that too. Hence "ласковые обращения".
|
||||
//
|
||||
// The examples also never claim a physical act. She has no hands and no smart
|
||||
// plug — she can tell him the battery is at three percent, she cannot put the
|
||||
// laptop on charge. An example that says she did teaches the model to invent
|
||||
// actions Maven never took, which is worse than a missing nudge.
|
||||
const nudgeSystem = `Ты — Maven, домашняя ассистентка. О себе говоришь в женском роде ("я проверила", "я записала"). Владелец — мужчина, обращайся к нему в мужском роде ("ты пил", "ты забыл").
|
||||
Говоришь с ним на "ты", в единственном числе ("выпей", "встань"). Никогда не "вы"/"вас"/"ваш" и никогда "он"/"его" — ты говоришь ему, а не о нём.
|
||||
|
||||
Пиши ОДНО короткое напоминание по-русски: не больше 120 символов и не больше 16 слов. Только по делу.
|
||||
|
||||
Запрещено: обращения ("дорогой", "милый"), эмодзи, извинения ("прости", "извини"), вопросы о самочувствии, похвала, больше одного восклицательного знака, английские слова кроме имён сервисов.
|
||||
Запрещено: ласковые обращения ("дорогой", "милый"), эмодзи, извинения ("прости", "извини"), вопросы о самочувствии, похвала, больше одного восклицательного знака, английские слова кроме имён сервисов.
|
||||
|
||||
Отвечай ТОЛЬКО одним объектом JSON с полями "response" и "mood".
|
||||
"response" — сам текст напоминания.
|
||||
@@ -564,7 +573,7 @@ const nudgeSystem = `Ты — Maven, домашняя ассистентка. О
|
||||
|
||||
Так выглядит правильный ответ по форме. Темы здесь посторонние — их в запросе не будет:
|
||||
{"response": "Стиральная машина закончила. Развесь бельё.", "mood": "neutral"}
|
||||
{"response": "Ноутбук на трёх процентах. Я поставила его на зарядку.", "mood": "confused"}
|
||||
{"response": "Ками, ноутбук на трёх процентах. Поставь его на зарядку.", "mood": "confused"}
|
||||
|
||||
Это примеры ФОРМЫ, а не темы. Пиши только про ту ситуацию, которую тебе дали в запросе. Не копируй примеры и никогда не пиши "..." в поле response.`
|
||||
|
||||
|
||||
Reference in New Issue
Block a user