diff --git a/ROUTING-EVAL-31-07-2026.md b/ROUTING-EVAL-31-07-2026.md index 7bf81fa..1ddef82 100644 --- a/ROUTING-EVAL-31-07-2026.md +++ b/ROUTING-EVAL-31-07-2026.md @@ -124,6 +124,55 @@ Two caveats worth saying out loud: Phrasing was **not** measured. Whether thinking helps there is still open, and now also blocked on the same "can we even turn it off" question. +## Clock and calendar rule — 31-07-2026 (Vikunja #374) + +`routeSystem` never said whether "который час" or "какое число завтра" are `system` or +`query`, and `system→query ×4` showed up in every run. The rule added says: the clock and the +calendar date themselves are `system`; what is *written in* the calendar or in memory +("что у меня завтра", "какие есть напоминания") stays `query`; and a time named inside a +request ("напомни завтра…") is just a detail of the request, not a reason for `system`. + +That split is not a preference. In `cmd/mavend/voice.go` only `replySystem` owns the clock and +the date formatter, so a clock question routed to `query` falls into the embedder + note RAG +and answers "не знаю". The agenda, on the other hand, is answered by `ParseCalendarDate` + +`CalendarEvents` *inside* the `query` branch, so that side has to stay `query`. The rule sits +above the question test because every one of these utterances carries a question word and a +later rule would never be reached. + +The fixture is now 77 cases: one calendar-agenda case was added +(`ru-query-019` "что у меня стоит в календаре на послезавтра", intent `query`) specifically so +an over-broad system rule cannot pass unnoticed. The clock/date cases (`ru-sys-001/002/005`, +`en-sys-001`) already existed. + +Three runs, same box, back to back, never concurrently: + +| | baseline | first rule (too broad) | rule as committed | +|---|---|---|---| +| llm-only intent-only | 59.2% (45/76) | 54.5% (42/77) | 59.7% (46/77) | +| llm-only full | 38.2% | 35.1% | 39.0% | +| llm-only route errors | 3 | 4 | 5 | +| llm-only p50 | 1.09s | 0.91s | 0.93s | +| cascade+llm intent-only | 61.8% (47/76) | 58.4% | 62.3% (48/77) | +| cascade+llm full | 57.9% | 54.5% | 59.7% | +| cascade+llm route errors | 0 | 0 | 0 | +| cascade+llm p50 | 0.91s | 0.80s | 1.04s | + +**The targeted bug is fixed and the headline number did not move.** `system→query ×4` is gone +in both LLM configurations — the `time` and `date` tags go from 0/2 and 0/2 to 2/2 and 2/2 — +but the model then over-applies the rule, and `query→system ×5` plus `reminder→system ×2` +appear where they did not exist before. Net accuracy is a wash, inside the noise of a 77-case +fixture. + +The first attempt is shown because it is the honest history: it said "спрашивает время, дату +или день недели → system" with no scope, which swept up reminders, and it cost 3-5 points. It +was tightened once, on the reasoning that a rule capturing "напомни завтра в 7" is simply +wrong, and not tuned further. The remaining `query/reminder → system` over-trigger is a new, +separate weakness of the sub-1B model and deserves its own task rather than more prompt +kneading against a held-out fixture. + +The rule is kept. It is correct about what the daemon can answer, and the failure it replaces +was silent ("не знаю" to "который час") while the one it introduces is loud. + ## Findings ### 1. The resident model does route better — 50.0% vs 36.8% diff --git a/internal/router/eval/ru_routing_v1.json b/internal/router/eval/ru_routing_v1.json index 3cfb6a0..6c765f0 100644 --- a/internal/router/eval/ru_routing_v1.json +++ b/internal/router/eval/ru_routing_v1.json @@ -22,6 +22,7 @@ { "id": "ru-query-011", "utterance": "почему сервер тормозит", "lang": "ru", "intent": "query", "tags": ["homelab", "hard"], "note": "diagnostic question, not a chat opener" }, { "id": "ru-query-012", "utterance": "какие заметки я оставил про полив", "lang": "ru", "intent": "query", "tags": ["recall"] }, { "id": "ru-query-013", "utterance": "во сколько у меня встреча", "lang": "ru", "intent": "query", "tags": ["calendar"] }, + { "id": "ru-query-019", "utterance": "что у меня стоит в календаре на послезавтра", "lang": "ru", "intent": "query", "tags": ["calendar", "hard"], "note": "agenda, not the clock: the daemon answers this from CalendarEvents inside the query branch, so the clock/date system rule must not swallow it" }, { "id": "ru-query-014", "utterance": "я успеваю до дедлайна", "lang": "ru", "intent": "query", "tags": ["hard", "no-question-word"] }, { "id": "ru-query-015", "utterance": "сколько я прошёл шагов", "lang": "ru", "intent": "query", "tags": ["aggregate"] }, { "id": "ru-query-016", "utterance": "покажи давление за неделю", "lang": "ru", "intent": "query", "tags": ["hard", "imperative"], "note": "imperative form but a read — must not route to act" }, diff --git a/internal/router/llmrouter.go b/internal/router/llmrouter.go index 36a921a..33c47c5 100644 --- a/internal/router/llmrouter.go +++ b/internal/router/llmrouter.go @@ -44,10 +44,21 @@ ws ::= [ \t\n]* // Changed again 31-07-2026: added the "unknown" escape hatch so the model can // admit it cannot route (Vikunja #359). // +// Changed again 31-07-2026: added the clock/calendar rule (Vikunja #374). The +// prompt never said which side "который час" or "какое число завтра" belong on, +// so the model guessed — `system→query ×4` in every eval run. The rule sits +// above the question test on purpose: these utterances all carry a question +// word, so a later rule would never be reached. The boundary is what the +// daemon can actually answer: only replySystem in cmd/mavend/voice.go owns the +// clock and the calendar formatter, while the agenda ("что у меня завтра") is +// answered inside the query branch, so that side stays query. +// // The training workspace keeps its own copy of this prompt for relabelling, and // `llm/check_prompt_parity.py` there compares the two. That copy is in another // repo and was not touched, so parity will fail until it gets the same edits — -// both the rule reorder and the "unknown" wording (Vikunja #362). +// both the rule reorder and the "unknown" wording (Vikunja #362) — and now the +// clock/calendar rule too. The training workspace is not checked out on this +// box at all, so it could not be updated here; #362 still covers the catch-up. const routeSystem = `Классифицируй ровно одно сообщение пользователя. Верни ОДИН JSON-массив действий. Ровно одно намерение: fact, reminder, note, query, act, chat, system. @@ -56,19 +67,21 @@ const routeSystem = `Классифицируй ровно одно сообще Классифицируй по цели пользователя. Порядок решения: 1. Хочет напоминание в будущем → reminder 2. Явно просит сохранить информацию → note -3. Задаёт вопрос: есть вопросительное слово (сколько, что, какой, когда, где, кто, почему, как) или знак «?» → query -4. Хочет получить информацию, в том числе о своих же данных → query -5. Утверждает: сообщает или обновляет текущее состояние/событие → fact -6. Просит выполнить работу → act -7. Про ассистента, настройки или память → system -8. Реплика — обрывок или указание на неназванное («это», «то», «потом»), и без него непонятно, что именно нужно сделать → unknown -9. Иначе → chat +3. Спрашивает только «который час» / «какое число» / «какой день недели» — сами часы или календарная дата, без своих данных → system +4. Задаёт вопрос: есть вопросительное слово (сколько, что, какой, когда, где, кто, почему, как) или знак «?» → query +5. Хочет получить информацию, в том числе о своих же данных → query +6. Утверждает: сообщает или обновляет текущее состояние/событие → fact +7. Просит выполнить работу → act +8. Про ассистента, настройки или память → system +9. Реплика — обрывок или указание на неназванное («это», «то», «потом»), и без него непонятно, что именно нужно сделать → unknown +10. Иначе → chat Различия: - note — сохранить информацию, без напоминания. text = суть. - reminder — уведомить позже. text = что напомнить. - fact — неявное обновление: пользователь сообщает, что что-то в мире изменилось (текущее/изменённое состояние, случившееся событие). key/value. - unknown — редкий случай. Ставь его, только если в самой реплике нет ни предмета, ни действия. Короткая, простая или незнакомая тема — это не причина для unknown: приветствие и болтовня — это chat, вопрос на любую тему — это query, просьба сделать что-то названное — это act. +- system против query — часы и календарная дата сами по себе (сколько времени, какое число, какой день недели — можно и про завтра, и про другой город) — это system. А что записано в календаре или в памяти («что у меня завтра», «какие есть напоминания») — это query. Если в реплике есть просьба (напомни, запиши, сделай), то названное время — просто деталь просьбы, и это не system. - query против fact — решает форма реплики, а не тема. Вопрос о состоянии — это query, даже если названо то же самое, что бывает в fact. Только утверждение — это fact. Примеры: @@ -82,6 +95,8 @@ const routeSystem = `Классифицируй ровно одно сообще "что такое docker?" → {"intent":"query","text":"что такое docker"} "напиши письмо" → {"intent":"act","verb":"написать письмо"} "очисти память" → {"intent":"system"} +"который час?" → {"intent":"system"} +"какое число завтра?" → {"intent":"system"} "привет" → {"intent":"chat","text":"привет"} "сделай это" → {"intent":"unknown"} "ну это" → {"intent":"unknown"}