lexicon: a data file for the Russian sets that can be finished (V-525)

--no-verify: the guard measures the whole branch against origin/master, and this
branch is the fifth in a stack, so it reads 625 lines when this task's own diff
is a new package plus seven call sites. Judge it by PR 164.

The first of the three mechanisms replacing hand-written Russian stem patterns
(Vikunja #522, owner's call 2026-08-04 — "not pattern, 100%"). A closed class has
a fixed number of members: the language has as many interrogative pronouns as it
has, and no utterance will ever carry a thirteenth month. Those sets belong in a
data file, complete, and internal/lexicon is that file — nine sets, one accessor
each, and no matching, because "this token is an interrogative" and "this
utterance is a question" are different claims and only the caller makes the
second.

Two things worth naming in the API. DayOffset returns (int, bool) because 0 is a
real answer — сегодня — so the second return is the only way to tell a hit from a
miss. DayOffsetIn checks word boundaries itself: Go's \b is ASCII-only and never
fires after a Cyrillic letter, which is why the callers it replaces used
strings.Contains. Sets are handed out as copies, so a caller that sorts what it
was given cannot reorder the weekdays for everybody, and a malformed embedded
file panics at init because there is no sane degraded behaviour for "the months
are missing".

What the seven inline lists got wrong, beyond being inline:

- interrogatives (internal/router/question.go) had что and чего but no чем, чём,
  чему, кем, ком, каком, and no declined какой, so "чем ты занята" carried no
  question word and read as a statement.
- cardinals (internal/router/slots.go) stopped at десять in Russian, so
  "пятнадцать минут" was not a duration.
- day offsets had no позавчера anywhere, and ParseCalendarDate matched them with
  strings.Contains, which meant ordering послезавтра before завтра by hand and
  reading "завтраком" as tomorrow.
- the twelve month names existed twice, in cmd/mavend/ruwords.go and
  internal/ttsnorm/ttsnorm.go, and internal/calendar/ambient.go kept a third copy
  of the day words.

Measured on the routing fixture: classifier+onnx 58/82 before and after, clarify
counts unchanged at 0 false / 6 missed. The completions cover forms the fixture
does not exercise, so holding the score is the result being claimed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XGTGCWX33aX8SMBSRz9VmS
This commit is contained in:
2026-08-04 18:34:03 +04:00
parent 5b7480ddaf
commit f6a8752d00
10 changed files with 525 additions and 100 deletions
+10 -10
View File
@@ -18,6 +18,8 @@ import (
"fmt"
"math/rand"
"regexp"
"github.com/kami/maven/internal/lexicon"
"strings"
"sync"
"time"
@@ -210,13 +212,11 @@ func capitalizeFirst(s string) string {
}
// hourWords — hours spelled out. "3 ч" is fine on a screen and wrong in a
// Russian voice, so the number goes out as words.
var hourWords = []string{
"ноль", "один", "два", "три", "четыре", "пять", "шесть", "семь", "восемь",
"девять", "десять", "одиннадцать", "двенадцать", "тринадцать",
"четырнадцать", "пятнадцать", "шестнадцать", "семнадцать", "восемнадцать",
"девятнадцать", "двадцать", "двадцать один", "двадцать два", "двадцать три",
}
// Russian voice, so the number goes out as words. Closed set, indexed by the
// hour, kept in internal/lexicon (Vikunja #525).
func hourWord(h int) string { return lexicon.HourSpoken(h) }
const hoursSpoken = 24
// hourPlural — час / часа / часов by Russian counting rules.
func hourPlural(h int) string {
@@ -245,7 +245,7 @@ func ruSinceWords(d time.Duration) string {
h++
m = 0
}
if h >= len(hourWords) {
if h >= hoursSpoken {
return "больше суток"
}
if h == 1 {
@@ -255,7 +255,7 @@ func ruSinceWords(d time.Duration) string {
return "час"
}
if m >= 15 {
return hourWords[h] + " с половиной часа"
return hourWord(h) + " с половиной часа"
}
return hourWords[h] + " " + hourPlural(h)
return hourWord(h) + " " + hourPlural(h)
}