Compare commits
164 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 70b32af8a7 | |||
| fffd0cb5fa | |||
| 77a7c994d7 | |||
| 63812af920 | |||
| 869580c913 | |||
| 31deb7d565 | |||
| 1338ec6e2a | |||
| 7843728174 | |||
| a3ad9b5040 | |||
| cb0a3a4f20 | |||
| 6abd2768e8 | |||
| 6e6f73da35 | |||
| 1a64c30427 | |||
| 5753f90752 | |||
| 4d94277836 | |||
| 530c3ff395 | |||
| e031f8f5f5 | |||
| 6c24e19b83 | |||
| 934a28d67c | |||
| cf28f6fdf0 | |||
| eec3d9bed2 | |||
| a5b245dbf5 | |||
| 3e6a427e85 | |||
| 5ac7347c38 | |||
| 5417692566 | |||
| b56e0e6248 | |||
| 0558dfed0f | |||
| 13a5ef0100 | |||
| ac78f83406 | |||
| 40c59aa275 | |||
| 84a75274bf | |||
| da2d11dab6 | |||
| de3f2b5fc2 | |||
| e8f4baf407 | |||
| 92eb6cf6e1 | |||
| 6759ff6003 | |||
| 39284cd851 | |||
| ea0eb167fd | |||
| b6305f1b6e | |||
| 5bd1406c7a | |||
| d988154063 | |||
| 1b76fa8205 | |||
| e87088afb8 | |||
| 1b8d2c60d3 | |||
| 12667fd3b8 | |||
| e4fd6140a9 | |||
| 2ba5d0a60e | |||
| c6b11a6d1d | |||
| 45622eff3d | |||
| 5815f0b8f3 | |||
| f68d49d9e2 | |||
| 36bc603f52 | |||
| 59214b4fdd | |||
| e94c868160 | |||
| de4c47459a | |||
| 27bb9119fb | |||
| 0ab5dc1482 | |||
| 35ae1f41da | |||
| a9db82b04c | |||
| 278eeeffdf | |||
| b23596f54f | |||
| 6e3bb3be97 | |||
| aa7ef33bbf | |||
| f2851b3729 | |||
| d49067f7dd | |||
| b528a8f5c9 | |||
| 44320ee496 | |||
| 337a777d2e | |||
| 576dfd8b4c | |||
| ed9db8dc44 | |||
| a8710c859b | |||
| 1bacbb7952 | |||
| 82d9a3324e | |||
| 03d48ab789 | |||
| 570b204571 | |||
| cb350efb19 | |||
| 43b0c32154 | |||
| b95a0278a4 | |||
| a6b17ada8b | |||
| 92949e886f | |||
| 6b2667b7af | |||
| 494a7721e0 | |||
| 46b58f0278 | |||
| a88c984d16 | |||
| 496559c9dd | |||
| d21b4a65da | |||
| be62660be9 | |||
| cd6549fa51 | |||
| c259ed6c73 | |||
| 8f2377d27d | |||
| d850f1f5fd | |||
| 4ed951040f | |||
| 8d2c1b6f99 | |||
| ee4c26f13c | |||
| 49cadcf1b7 | |||
| 988c2ae981 | |||
| 982c25118a | |||
| de10d7664f | |||
| a08403087f | |||
| 3e87ad7eb6 | |||
| 47128bb1ca | |||
| 8e3b288858 | |||
| 52a4772962 | |||
| 52d80394ec | |||
| 42a7bd88b2 | |||
| b789676244 | |||
| b621c477a0 | |||
| fa67dd82fe | |||
| 52fd218c70 | |||
| 33032b859a | |||
| 0d8cbaec01 | |||
| 1f38e71d1a | |||
| 2de5a339fb | |||
| 1ee9a930d3 | |||
| f48c2280bd | |||
| 888c1c6768 | |||
| 7cacbc8b21 | |||
| dc3cda666e | |||
| 1b354e9b39 | |||
| c07722266a | |||
| be758d9a59 | |||
| 8bbdcd2727 | |||
| f8947bef5a | |||
| b752ec037e | |||
| 4dbeca5a2e | |||
| 9e15ff36aa | |||
| 7955a41105 | |||
| a93a16d7b3 | |||
| dd63180e44 | |||
| 9d80a39a30 | |||
| b5b599e287 | |||
| bb51c28a19 | |||
| 549d4c8380 | |||
| 4766167c3a | |||
| f6f9e75eac | |||
| 1524991adc | |||
| 9da468810e | |||
| 7b4fb6229a | |||
| cdd81e2ad5 | |||
| 32d5f68710 | |||
| fdc18edd87 | |||
| a98d25ecac | |||
| 113508eaac | |||
| 6dc2622596 | |||
| dd91c6961c | |||
| 0560684b35 | |||
| c4cf06d610 | |||
| acd985323e | |||
| 598f4fc011 | |||
| 69db1cf849 | |||
| f5480e281b | |||
| cc72f69769 | |||
| b954e0cea6 | |||
| c82dbd1e65 | |||
| 69ecea19d5 | |||
| 8a13d189bb | |||
| 8e7aa0d451 | |||
| 00595c2211 | |||
| b54ccccd0a | |||
| 4cfef41541 | |||
| 0793955896 | |||
| 71a9a59403 | |||
| d3c63e6493 | |||
| c586346a60 |
@@ -43,7 +43,7 @@ are the work.
|
||||
Both halves are wired as of 2026-08-03. Routing and replies prefer the workstation silently
|
||||
through `modelSeam`; nudge and reminder phrasing prefer it silently inside the phraser. A
|
||||
world question goes through `LLMPhraser.PhraseWorld` and names the gap when the card is not
|
||||
free — `worldGap` in `cmd/mavend/worldmodel.go`, which he hears instead of an invented
|
||||
free — `worldGap` in `cmd/mavend/worldmodel.go`, which the owner hears instead of an invented
|
||||
answer. A box with no `workstation` block behaves exactly as it did before the seam: naming
|
||||
a gap requires a gap. The offload table in `docs/offload.md` says which caller is which.
|
||||
|
||||
@@ -90,11 +90,18 @@ protocol; the config in `deploy/mavend.json` (with `${VAR}` env expansion from g
|
||||
**Seven of the nine run on homesrv. `mavwaked` and `mavenclient` do not, and that is the
|
||||
decision, not an oversight** (Vikunja #463, `docs/plans/17-where-the-voice-loop-runs.md`).
|
||||
homesrv has a microphone — it is a laptop — but it is in the wrong room, so a wake-word
|
||||
daemon there listens to nobody. They belong on a client machine where he is standing.
|
||||
`ipc.Dial` already takes `tcp://host:port?token=...` through the netaddr seam, so nothing
|
||||
needs building to allow it, but no such machine exists yet. **The consequence: the wake
|
||||
word and the VAD gate are covered by unit tests and by nothing else, and no amount of
|
||||
sitting at the box changes that.** Push-to-talk through `/dash` is what QA actually covers.
|
||||
daemon there listens to nobody. They belong on a client machine where the owner is standing.
|
||||
|
||||
**That machine is workpc** (owner's correction, 2026-08-05). This section used to say no
|
||||
such machine existed, which was written when the workstation was only a model host. It is
|
||||
where he sits most of the day and it has the microphone. `ipc.Dial` already takes
|
||||
`tcp://host:port?token=...` through the netaddr seam, so the two daemons need deploying,
|
||||
not building. V-515 is that deployment.
|
||||
|
||||
Until they are deployed, **the wake word and the VAD gate are covered by unit tests and by
|
||||
nothing else**, and push-to-talk through `/dash` is what QA actually covers. Note that
|
||||
deploying them does not by itself prove a wake word: `mavwaked` gates on energy and has no
|
||||
keyword model (V-487), so the loop runs open until that lands.
|
||||
|
||||
## The ecosystem: Nexus, Praxis, Hexis
|
||||
|
||||
@@ -155,6 +162,19 @@ on in deploy** — this section used to say it was wired `nil`, which stopped be
|
||||
Cascade order: `stage0.go` exact-match fast-path → LLM router (when non-nil) → classifier
|
||||
fallback. Any LLM error falls through to the classifier so a turn never breaks on the model.
|
||||
|
||||
**A stage-0 decision is slot-extracted too, since 06-08-2026** (V-572). `fillMatchedSlots`
|
||||
in `router.go` runs the stage-2 extractor over whatever a grammar built and fills only the
|
||||
slots it left empty — a matched value always wins, because the rule read a literal pattern
|
||||
and the extractor guesses. It did not run before, so `ReminderGrammar` handed the daemon
|
||||
`HasTime: false` for "напомни в 11:00 позвонить маме" and `missingFor` read the silence as
|
||||
absence and asked "Когда?". It applies to every grammar and is inert for all but the
|
||||
reminder: `Extract` fills Time, Fn and Key and nothing else, and the query, clock, agenda,
|
||||
feed, list, task and narrative rules all emit intents with no such slot. Benchmarked at
|
||||
20000x, a stage-0 query costs 3.7µs against 3.9µs before. **`Slots.Text` is deliberately not
|
||||
filled** — a grammar that left it empty meant it, and `agendaQueryBuild` hands the query
|
||||
chain the utterance itself. Fixture unchanged at 64/91, with "slots deferred to daemon"
|
||||
6 → 0.
|
||||
|
||||
Measured on the 77-case RU fixture. **Re-measured 2026-08-02: the classifier scores 68.8%
|
||||
full accuracy at p50 16.6µs**, not the 36.8% at p50 31ms that stood here from
|
||||
`docs/evals/2026-07-31-model-bakeoff.md`. That older figure predates the stage 0 rules and the
|
||||
@@ -165,12 +185,32 @@ that stood here until 2026-08-02 was contention, not the model.** See `docs/eval
|
||||
p50 825ms / p95 1.2s / max 3.0s and the full cascade at p50 0.80-1.04s. Do not plan latency
|
||||
work off the bakeoff table.
|
||||
|
||||
**Re-measured 2026-08-05 on the fixture as it now stands, 91 cases** (V-320 item 2,
|
||||
`docs/evals/2026-08-05-routing-resident-model.md`): cascade + resident model scores
|
||||
**75.8% full / 80.2% intent-only at p50 1.19s / p95 1.65s**. That is a new baseline and not
|
||||
a movement, because 14 cases were added since the 77-case number above. The model alone
|
||||
scores 37.4% full against 61.5% intent-only, and the gap is slots rather than routing: it
|
||||
routes `reminder` and leaves the time to the daemon, which is what the contract asks. To
|
||||
re-run it, start a **second** llama-server on a fixed host port — the resident one binds
|
||||
`--port 0` inside the container and no host process can reach it.
|
||||
|
||||
**The numbers above are the homesrv floor, not the ceiling.** With the workstation up, routing
|
||||
completes through `llm.Pair` against gemma-4-12b and scores **84.4% full / 93.5% intent-only at
|
||||
p50 329ms** — better than the resident model and about 2.5× faster (`docs/evals/2026-08-02-workstation-gemma4-12b.md`,
|
||||
Vikunja #485). The workstation is never assumed up, so both sets of numbers are live. Judge a
|
||||
routing change against the classifier and the resident model, since those are what always answer.
|
||||
|
||||
**The intended third engine is not a generative model** (owner's call, 05-08-2026, V-546,
|
||||
`docs/plans/18-routing-heads-on-e5-small.md`). Routing has a bounded output space, so it is
|
||||
classification, and the 118M multilingual-e5-small is already resident. Three heads on one
|
||||
forward pass: intent, mood, and BIO slot tags. Roughly 5e15 FLOPs to train, so 10 to 30
|
||||
minutes on the workstation. A 100M decoder from scratch is 10 to 20 GPU hours. Two things
|
||||
it buys that a decoder cannot. No grammar is needed, because a softmax cannot emit a value
|
||||
that does not exist. And max softmax is a calibratable confidence, where `Confidence: 1.0`
|
||||
was a hardcode. **Fine-tune a copy of the weights.** The resident embedder backs memory
|
||||
recall. Training it in place couples routing accuracy to recall@1, with nothing in the
|
||||
suite to name the trade.
|
||||
|
||||
`Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
|
||||
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
|
||||
#359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with
|
||||
@@ -214,6 +254,42 @@ New fixture cases ru-query-024 and ru-query-025. Classifier + ONNX baseline **56
|
||||
58/82 (70.7%)**, no case regressed, no new false clarify. The LLM arm was not measured (no
|
||||
llama-server in that run), so judge it again before quoting a cascade number.
|
||||
|
||||
Praxis taken off the model, 05-08-2026 (V-516). `PraxisGrammars()`
|
||||
(`internal/router/praxis.go`, wired in `buildRouter` before the capture marker because
|
||||
"отметь" is a capture verb) fills `Slots.Fn` with a Praxis capability name.
|
||||
**These grammars are the only path to Praxis, not a faster one.** Measured
|
||||
2026-08-05 with the resident model as router (V-517,
|
||||
`docs/evals/2026-08-05-reach-llm-router.md`): the model alone reaches Praxis
|
||||
**0/12**, the same as the classifier alone, because nothing in the router
|
||||
prompt names a Praxis capability and there is no string for it to write.
|
||||
Through the cascade it is 11/12. Deleting these rules costs every point. Praxis reach
|
||||
was **0/12 and structurally so**: `handlePraxisAct` compares `Slots.Fn` to a capability
|
||||
alias, and that slot is filled from the deployment's enabled tool names, which no Praxis
|
||||
alias is on. Measured **16/30 → 27/30 overall, praxis 0/12 → 11/12, lifecycle 0/5 → 5/5**
|
||||
(`docs/evals/2026-08-05-praxis-reach.md`). Two rules to know before editing: a **stative**
|
||||
lifecycle word ("готово", "принято") needs an item named beside it, while a bare
|
||||
**imperative** ("закрывай") may ask which one. The bare arm additionally requires that
|
||||
the sentence name no object of its own, or "закрой шторы в комнате" goes to Praxis instead
|
||||
of the house. A demonstrative ("отметь это как сделанное") resolves against
|
||||
`h.surfacedItems` only when exactly one item was spoken. Otherwise the turn goes back to
|
||||
the cascade rather than transitioning the wrong item.
|
||||
|
||||
**Who claimed a turn is now recorded, and so is who did not** (V-564, umbrella
|
||||
V-558). Arbitration between the claimants on the utterance stream is order,
|
||||
hardcoded in the pre-route resolver ladder, in `buildRouter` and in
|
||||
`querySources`. `internal/decision` records one `Record` per turn: every
|
||||
claimant, what it would have made the turn, the score it reported, and whether
|
||||
it won, declined, lost on score, was thinned by a gate or was **never asked**.
|
||||
The record rides the context, the same seam `querysource.go` uses, so a claim
|
||||
site cannot change a route and a context with no record costs nothing. It is
|
||||
installed in `runTurn`, so the mic, telegram and the web all leave the same
|
||||
trail. Storage is a 25-turn in-memory ring on the handler (`decision.Ring`),
|
||||
read over `ipc.TurnDecisions` and rendered as the second table on `/trace`.
|
||||
Nothing persists: a turn record is read minutes later or never, and his words do
|
||||
not belong in a table that outlives the diagnosis. Adding a rung to the ladder
|
||||
in `runTurn` means adding its name to `preRouteLadder` in
|
||||
`cmd/mavend/decisiontrace.go`, or that rung is silently missing from the record.
|
||||
|
||||
## LLM output contract
|
||||
|
||||
All phrasing paths emit `{"response":"...","mood":"..."}` (parsed in `replier_llm.go` and
|
||||
@@ -256,7 +332,7 @@ Seeds are scoring data. Editing one moves a recogniser and must be re-measured a
|
||||
Not a nag, not autonomous. Maven's persona is **feminine** — Russian
|
||||
self-reference must use feminine forms — `рада`, not `рад`; `поняла`, not `понял`. The owner
|
||||
is male and is addressed informally: "ты", singular, never "вы"/"ваш" and never "он"/"его"
|
||||
(she talks TO him, not about him). Pet names ("милый", "дорогой") are forbidden; his name
|
||||
(she talks TO the owner, not about the owner). Pet names ("милый", "дорогой") are forbidden; the name
|
||||
("Ками") is not. The eval enforces this: `CheckAddress`, `CheckFeminine` and `CheckCringe` in
|
||||
`internal/phraser/eval/checks.go`, scored by `make eval-phrasing`.
|
||||
|
||||
@@ -266,18 +342,35 @@ world questions, so she needs to read external sources. What replaces it:
|
||||
|
||||
- **No telemetry, no cloud model, no third-party account.** That part never changes. Nothing
|
||||
about Maven is reported to anyone, and inference stays on the box.
|
||||
- **His data first, then the world.** Every source that reads his facts, notes, calendar,
|
||||
- **The owner's data first, then the world.** Every source that reads the owner's facts, notes, calendar,
|
||||
tasks or house runs before anything outside, and the personal boundary sits between them.
|
||||
Reading beats recalling for a small model.
|
||||
- **In the world, live search leads and the ZIMs are the fallback** (owner's call,
|
||||
2026-08-02). A self-hosted SearXNG (`search` block) answers first; the Kiwix ZIMs on
|
||||
homesrv answer when the search is empty, unreachable, or the line is down.
|
||||
**Verified with the line down on 2026-08-05** (V-508,
|
||||
`docs/evals/2026-08-05-kiwix-offline-fallback.md`): a stopped SearXNG costs nothing,
|
||||
the ZIM answers in the same turn budget. A blackholed host cost 8 seconds the owner waited
|
||||
through. So the connect phase alone is capped at `dialTimeout` (1.5s), while a slow
|
||||
instance that did connect keeps the full 8. **A Russian question reads
|
||||
`wikipedia_ru_all_maxi_2026-02` verbatim** through `kiwix.book_ru`. The RU→EN rewriter
|
||||
is the workaround for an English book and is skipped there. Kiwix catalog names come
|
||||
from the filename, not the `<name>` field.
|
||||
`Response.Empty()` is the whole gate and there is no quality threshold in front of it:
|
||||
the three signals one could read were measured on 2026-08-05 and none of them separate a
|
||||
real question from an invented one. Token overlap would cost "столица Франции" its
|
||||
answer, because the answer is Париж and that word is not in the question. See
|
||||
`docs/evals/2026-08-05-search-quality-signals.md` (V-539). **Which query source claimed
|
||||
a turn is readable on `/chat`** as a badge beside the reply, carried on
|
||||
`ipc.ChatReply.Source` and noted by `noteQuerySource` in `cmd/mavend/querysource.go`. It
|
||||
rides the context, so `handleText` keeps the one string signature the mic, telegram and
|
||||
the web share.
|
||||
- **External search is allowed and off unless configured**, like the weather and telegram
|
||||
capabilities. The code default is still off. `deploy/mavend.json` now ships a `search`
|
||||
block (owner's call, 2026-08-02), so it is on for this box and deleting the block turns
|
||||
it off again.
|
||||
- **His notes and facts are never search input.** Looking up why the sky is blue and sending
|
||||
his stored personal notes to an upstream engine are different acts. Only the utterance goes
|
||||
- **The owner's notes and facts are never search input.** Looking up why the sky is blue and
|
||||
sending the owner's stored personal notes to an upstream engine are different acts. Only the utterance goes
|
||||
out, never the persona block, history, or matched notes.
|
||||
|
||||
## Web UI conventions
|
||||
@@ -299,6 +392,16 @@ Vikunja is the durable task store. A task holds the goal, the constraints and th
|
||||
assumption ledger. Work without a task id is work nobody can resume, so a session that
|
||||
has no id asks for one before it starts.
|
||||
|
||||
The MCP tool schemas are deferred, so load the four you actually use in ONE call at the
|
||||
start of a session rather than one lookup per first use:
|
||||
|
||||
```text
|
||||
ToolSearch("select:mcp__vikunja__list_tasks,mcp__vikunja__get_task_details,mcp__vikunja__create_task,mcp__vikunja__update_task")
|
||||
```
|
||||
|
||||
`update_task` carrying a `description` resets `done` to false, so closing a task with a
|
||||
write-up takes two calls: the description, then `done: true`.
|
||||
|
||||
## Session workflow
|
||||
|
||||
`~/.local/bin/task` owns the branch, the commit identity and the PR. One task, one
|
||||
|
||||
+11
@@ -82,11 +82,22 @@ FROM debian:trixie-slim AS runtime
|
||||
# tzdata so the TZ env (set in compose) resolves — otherwise Go can't load the
|
||||
# zone and time.Now() stays UTC, and mavend answers clock/date queries and
|
||||
# evaluates quiet-hours in UTC.
|
||||
#
|
||||
# TZ is a build arg as well as an env because the image was self-inconsistent
|
||||
# without it (V-545): compose set TZ=Europe/Samara and Go read it, but
|
||||
# /etc/localtime still pointed at Etc/UTC, so anything asking the system zone
|
||||
# instead of the environment answered UTC. The reminder path shells out to
|
||||
# python dateparser, which is exactly such a caller. Compose passes the same
|
||||
# zone it already declares, so the zone is written in one place.
|
||||
ARG TZ=Etc/UTC
|
||||
ENV TZ=$TZ
|
||||
RUN apt-get update && apt-get install -y --no-install-recommends \
|
||||
ca-certificates libvulkan1 mesa-vulkan-drivers libgomp1 tzdata \
|
||||
python3 python3-pip && \
|
||||
pip3 install --no-cache-dir --break-system-packages 'dateparser==1.4.1' && \
|
||||
apt-get purge -y --auto-remove python3-pip && \
|
||||
ln -snf "/usr/share/zoneinfo/$TZ" /etc/localtime && \
|
||||
echo "$TZ" > /etc/timezone && \
|
||||
rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# runtime native libs: whisper/ggml (incl. vulkan) are real files in deps/lib.
|
||||
|
||||
@@ -61,7 +61,7 @@ func (h *reactiveHandler) actionChat(ctx context.Context, dec router.Decision) s
|
||||
if h.phraser == nil {
|
||||
return "поговорили."
|
||||
}
|
||||
history := h.chatHistory()
|
||||
history := h.chatHistory(ctx)
|
||||
// The phraser hands back its own fallback text alongside the error, so the
|
||||
// turn survives a dead server and the failure still reaches the log.
|
||||
reply, err := h.phraser.PhraseChat(ctx, dec.Utterance, history)
|
||||
|
||||
@@ -24,6 +24,14 @@ func (h *reactiveHandler) actionAct(ctx context.Context, dec router.Decision) st
|
||||
}
|
||||
}
|
||||
|
||||
// The board is Maven's own store, so a spoken status change is answered here
|
||||
// and never offered to an ecosystem client (Vikunja #512). First, because
|
||||
// task_status is on no allowlist and no capability registry: reaching either
|
||||
// of them would answer a turn about his own task list with a gap.
|
||||
if dec.Slots.Fn == router.TaskStatusFn {
|
||||
return h.resolveTaskStatus(ctx, dec)
|
||||
}
|
||||
|
||||
// Praxis ecosystem tools: intercept before the system command executor.
|
||||
if h.ecosystem != nil && h.ecosystem.praxis != nil && dec.Slots.HasFn {
|
||||
if reply := h.handlePraxisAct(ctx, dec); reply != "" {
|
||||
|
||||
@@ -6,6 +6,7 @@ import (
|
||||
"strconv"
|
||||
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/memory"
|
||||
"github.com/kami/maven/internal/phraser"
|
||||
"github.com/kami/maven/internal/router"
|
||||
"github.com/kami/maven/internal/store"
|
||||
@@ -89,7 +90,15 @@ func (h *reactiveHandler) actionFact(ctx context.Context, dec router.Decision) s
|
||||
// hears; storing the utterance meant recall answered with his own sentence
|
||||
// rather than the value. The utterance stays alongside as provenance —
|
||||
// readable on /trace, never the answer and never embedded.
|
||||
//
|
||||
// The vector id carries a timestamp, so a second tap of the same key adds a
|
||||
// row rather than replacing one, and recall then scores the superseded
|
||||
// value against the current one. CorrectValue and VoidLatestFact already
|
||||
// drop the key's vectors; an ordinary re-tap is the third way a value is
|
||||
// superseded and it did not (#493). Dropping first keeps exactly one vector
|
||||
// per key, which is what "recall answers with the current value" means.
|
||||
if h.recall.memStore != nil {
|
||||
pruneFactVectors(ctx, h.recall.memStore, dec.Slots.Key)
|
||||
text := store.FactRecallText(dec.Slots.Key, dec.Slots.Value)
|
||||
if vec, err := router.EmbedPassage(ctx, h.recall.embedder, text); err != nil {
|
||||
log.Printf("voice: embed fact for memory: %v", err)
|
||||
@@ -114,3 +123,29 @@ func (h *reactiveHandler) actionFact(ctx context.Context, dec router.Decision) s
|
||||
}
|
||||
return "" // replier phrases the success reply
|
||||
}
|
||||
|
||||
// vectorPruner — the part of the vector index this file needs and memory.Store
|
||||
// does not carry. store.MemoryStore implements it; the in-memory test double
|
||||
// may not, and a double that cannot prune is not a reason to fail a fact write.
|
||||
type vectorPruner interface {
|
||||
DeletePrefix(ctx context.Context, prefix string) (int64, error)
|
||||
}
|
||||
|
||||
// pruneFactVectors drops every vector for one fact key, so the insert that
|
||||
// follows is the only one left. Best-effort and silent on a store that cannot
|
||||
// prune: the fact row is the truth, and a stale vector costs a wrong recall,
|
||||
// not a lost fact.
|
||||
func pruneFactVectors(ctx context.Context, ms memory.Store, key string) {
|
||||
p, ok := ms.(vectorPruner)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
n, err := p.DeletePrefix(ctx, "fact:"+key+":")
|
||||
if err != nil {
|
||||
log.Printf("voice: prune memory vectors for %q: %v", key, err)
|
||||
return
|
||||
}
|
||||
if n > 0 {
|
||||
log.Printf("voice: %q superseded, dropped %d stale memory vector(s)", key, n)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -95,6 +95,12 @@ func (h *reactiveHandler) removeListItem(ctx context.Context, cap router.ListCap
|
||||
return "", false
|
||||
}
|
||||
|
||||
// listFloor — the keyword test behind topicList, in the shape turnIsAbout takes.
|
||||
func listFloor(u string) bool {
|
||||
_, ok := router.ParseListQuery(u)
|
||||
return ok
|
||||
}
|
||||
|
||||
// queryList — "что в списке покупок?", "что мне купить?".
|
||||
//
|
||||
// A query source, so it sits in querySources and either claims the turn or
|
||||
@@ -102,10 +108,15 @@ func (h *reactiveHandler) removeListItem(ctx context.Context, cap router.ListCap
|
||||
// source is: the notes pass would otherwise answer a list question with
|
||||
// whatever note is nearest.
|
||||
func (h *reactiveHandler) queryList(ctx context.Context, t *queryTurn) (string, bool) {
|
||||
list, ok := router.ParseListQuery(t.dec.Utterance)
|
||||
if !ok || h.dataStore == nil {
|
||||
if h.dataStore == nil {
|
||||
return "", false
|
||||
}
|
||||
// The seeds decide the subject and listQueryPrefixes is the floor behind
|
||||
// them (V-522). Which list he named is a noun lookup either way.
|
||||
if !h.turnIsAbout(ctx, t, topicList, listFloor) {
|
||||
return "", false
|
||||
}
|
||||
list := router.ListNamedIn(t.dec.Utterance)
|
||||
items, err := h.dataStore.ListItems(ctx, list, "")
|
||||
if err != nil {
|
||||
log.Printf("voice: list items: %v", err)
|
||||
|
||||
+114
-14
@@ -8,8 +8,10 @@ import (
|
||||
"regexp"
|
||||
"strings"
|
||||
"time"
|
||||
"unicode"
|
||||
|
||||
"github.com/kami/maven/internal/crawl"
|
||||
"github.com/kami/maven/internal/decision"
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/memory"
|
||||
"github.com/kami/maven/internal/morning"
|
||||
@@ -121,6 +123,12 @@ var querySources = []querySource{
|
||||
{name: "network", answer: (*reactiveHandler).queryNetwork},
|
||||
{name: "calendar", answer: (*reactiveHandler).queryCalendar, dateAware: true},
|
||||
{name: "weather", answer: (*reactiveHandler).queryWeather},
|
||||
// A question about her, above the three sources that search his own data
|
||||
// (Vikunja #555). It has no answer anywhere else: below the boundary
|
||||
// SearXNG answers about somebody else's assistant, and above it his notes
|
||||
// answer by proximity — "кто ты" came back from a note of his, measured on
|
||||
// the box, because the recall index has no idea the subject is her.
|
||||
{name: "self", answer: (*reactiveHandler).querySelf},
|
||||
{name: "embed", answer: (*reactiveHandler).queryEmbed},
|
||||
{name: "memory", answer: (*reactiveHandler).queryMemory},
|
||||
{name: "notes", answer: (*reactiveHandler).queryNotes},
|
||||
@@ -151,8 +159,17 @@ var querySources = []querySource{
|
||||
|
||||
func (h *reactiveHandler) actionQuery(ctx context.Context, dec router.Decision) string {
|
||||
t := &queryTurn{dec: dec}
|
||||
// The roster, so the record can say which sources were never reached rather
|
||||
// than leaving them out and letting a reader assume they looked and passed
|
||||
// (V-564). Finish names everyone below the winner.
|
||||
decision.Expect(ctx, decision.StageQuery, querySourceNames())
|
||||
rec := decision.From(ctx)
|
||||
for _, src := range querySources {
|
||||
if dec.Continued && !src.dateAware {
|
||||
rec.Note(decision.Claim{
|
||||
Stage: decision.StageQuery, Claimant: src.name, Outcome: decision.NeverAsked,
|
||||
Reason: "a continuation turn only asks the date-aware sources",
|
||||
})
|
||||
continue
|
||||
}
|
||||
if reply, ok := src.answer(h, ctx, t); ok {
|
||||
@@ -161,10 +178,20 @@ func (h *reactiveHandler) actionQuery(ctx context.Context, dec router.Decision)
|
||||
// no query-source field, so a wrong answer could not be told from a
|
||||
// wrongly-ordered chain (Vikunja #474). Only the name is logged —
|
||||
// the utterance and the answer are already on the voice lines above
|
||||
// and below this one.
|
||||
// and below this one. The same name goes to the turn's sink when the
|
||||
// caller asked for one, so /chat can show it (V-539).
|
||||
log.Printf("voice: query claimed by source %q", src.name)
|
||||
noteQuerySource(ctx, src.name)
|
||||
rec.Note(decision.Claim{
|
||||
Stage: decision.StageQuery, Claimant: src.name,
|
||||
Intent: string(dec.Intent), Outcome: decision.Won,
|
||||
})
|
||||
return reply
|
||||
}
|
||||
rec.Note(decision.Claim{
|
||||
Stage: decision.StageQuery, Claimant: src.name, Outcome: decision.Declined,
|
||||
Reason: "it had no answer for this turn",
|
||||
})
|
||||
}
|
||||
if dec.Continued {
|
||||
// The previous question cannot be re-asked for another day. Saying so
|
||||
@@ -289,6 +316,14 @@ const (
|
||||
feedReadOut = 3
|
||||
)
|
||||
|
||||
// feedFloor — the keyword test behind topicFeed, in the one-string shape
|
||||
// turnIsAbout takes. router.ParseFeedQuery returns the category too, which the
|
||||
// gate has no use for; the caller reads it separately.
|
||||
func feedFloor(u string) bool {
|
||||
_, ok := router.ParseFeedQuery(u)
|
||||
return ok
|
||||
}
|
||||
|
||||
// queryFeeds — "что нового в лентах?", "что нового по технологиям?"
|
||||
// (Vikunja #258).
|
||||
//
|
||||
@@ -296,10 +331,13 @@ const (
|
||||
// never speaks; asking is the trigger. If that ever changes, the thing that
|
||||
// changed is "Maven is not a nag", not a detail of this file.
|
||||
func (h *reactiveHandler) queryFeeds(ctx context.Context, t *queryTurn) (string, bool) {
|
||||
q, ok := router.ParseFeedQuery(t.dec.Utterance)
|
||||
if !ok {
|
||||
// The seeds decide the subject; router.ParseFeedQuery is the floor behind
|
||||
// them (V-522). The category still comes from the utterance either way,
|
||||
// because a topic is marked by a preposition and needs no recogniser.
|
||||
if !h.turnIsAbout(ctx, t, topicFeed, feedFloor) {
|
||||
return "", false
|
||||
}
|
||||
category := router.FeedCategoryOf(t.dec.Utterance)
|
||||
if !h.feedsOn {
|
||||
// Claim only when nothing below can read the world. The reason this
|
||||
// source used to claim unconditionally was that general knowledge would
|
||||
@@ -324,7 +362,7 @@ func (h *reactiveHandler) queryFeeds(ctx context.Context, t *queryTurn) (string,
|
||||
}
|
||||
var picked []string
|
||||
for _, n := range notes {
|
||||
if !router.CategoryMatches(rss.NoteCategory(n.Text), q.Category) {
|
||||
if !router.CategoryMatches(rss.NoteCategory(n.Text), category) {
|
||||
continue
|
||||
}
|
||||
// The note carries title, summary, category tag and link; she reads the
|
||||
@@ -336,7 +374,7 @@ func (h *reactiveHandler) queryFeeds(ctx context.Context, t *queryTurn) (string,
|
||||
}
|
||||
}
|
||||
if len(picked) == 0 {
|
||||
if q.Category != "" {
|
||||
if category != "" {
|
||||
return phraser.Q(phraser.QueryFeedsTopic, nil), true
|
||||
}
|
||||
return phraser.Q(phraser.QueryFeedsEmpty, nil), true
|
||||
@@ -357,6 +395,19 @@ func (h *reactiveHandler) queryCalendar(ctx context.Context, t *queryTurn) (stri
|
||||
if isWeatherQuery(t.dec.Utterance) {
|
||||
return "", false
|
||||
}
|
||||
// Weather was one instance of a wider class (Vikunja #552). Naming a day
|
||||
// does not make a question his agenda: "какой сегодня курс доллара" and
|
||||
// "во сколько закат сегодня" both answered "ничего нет", which reads as an
|
||||
// answer about a subject she never looked at. All of them have an answer
|
||||
// in search, and search sits below this source. So the question must ask
|
||||
// about his schedule, not merely name a day.
|
||||
//
|
||||
// A continuation is exempt. "а завтра?" names no agenda and cannot: the
|
||||
// subject was in the turn before it, and this is the only date-aware
|
||||
// source there is.
|
||||
if !t.dec.Continued && !router.IsAgendaQuestion(t.dec.Utterance) {
|
||||
return "", false
|
||||
}
|
||||
date, ok := router.ParseCalendarDate(t.dec.Utterance, h.now())
|
||||
if !ok {
|
||||
return "", false
|
||||
@@ -452,12 +503,27 @@ func (h *reactiveHandler) queryWeather(ctx context.Context, t *queryTurn) (strin
|
||||
|
||||
// queryEmbed isn't an answer source — it's the shared cost the two recall
|
||||
// sources below both need, run once, in the position it always ran in. It
|
||||
// only claims the turn when the embedder fails.
|
||||
// never claims the turn.
|
||||
//
|
||||
// It used to claim on an embedder error, and that made a RAG hint a hard gate
|
||||
// over everything below it (V-568): one failing EmbedQuery and the memory, the
|
||||
// notes, the boundary, the search, the ZIMs, the named page and the model all
|
||||
// answered "не смогла ответить", including the questions search and Kiwix
|
||||
// would have answered without ever touching the embedder. A failed embed means
|
||||
// this source cannot claim, not that the turn is over — same shape as
|
||||
// turnVector in topics.go, which had it right.
|
||||
func (h *reactiveHandler) queryEmbed(ctx context.Context, t *queryTurn) (string, bool) {
|
||||
// A topic source above already paid for this one; see turnVector.
|
||||
if len(t.vec) > 0 {
|
||||
return "", false
|
||||
}
|
||||
vec, err := router.EmbedQuery(ctx, h.recall.embedder, t.dec.Utterance)
|
||||
if err != nil {
|
||||
// Logged once, here, and the chain walks on. The two recall sources
|
||||
// below read the empty vector and pass; the boundary drops to its
|
||||
// offline floor.
|
||||
log.Printf("voice: embed query: %v", err)
|
||||
return phraser.Q(phraser.QueryFailAnswer, nil), true
|
||||
return "", false
|
||||
}
|
||||
t.vec = vec
|
||||
return "", false
|
||||
@@ -477,6 +543,12 @@ func (h *reactiveHandler) queryMemory(ctx context.Context, t *queryTurn) (string
|
||||
if h.recall.memStore == nil {
|
||||
return "", false
|
||||
}
|
||||
if len(t.vec) == 0 {
|
||||
// No query vector: the embed above failed or there is no embedder.
|
||||
// Searching on an empty vector is not a search, and its scores are not
|
||||
// a "there is nothing" answer — pass rather than gate the chain.
|
||||
return "", false
|
||||
}
|
||||
hits, herr := h.recall.memStore.Search(ctx, t.vec, 3)
|
||||
if herr != nil {
|
||||
log.Printf("voice: memory search: %v", herr)
|
||||
@@ -523,10 +595,18 @@ func (h *reactiveHandler) queryMemory(ctx context.Context, t *queryTurn) (string
|
||||
// band. See memory.Confident. Failing the gate passes the turn on to general
|
||||
// knowledge, which is what "don't read back the runner-up" means here.
|
||||
func (h *reactiveHandler) queryNotes(ctx context.Context, t *queryTurn) (string, bool) {
|
||||
if len(t.vec) == 0 {
|
||||
// Same reason as queryMemory above (V-568): with no query vector this
|
||||
// source could not look, and could-not-look passes.
|
||||
return "", false
|
||||
}
|
||||
notes, err := h.api.QueryNotes(ctx, t.vec, 5)
|
||||
if err != nil {
|
||||
// The store failed, so this source could not look either. It used to
|
||||
// claim here, which stopped the search, the ZIMs and the model from
|
||||
// answering a question that never needed a note (V-568).
|
||||
log.Printf("voice: query notes: %v", err)
|
||||
return phraser.Q(phraser.QueryFailAnswer, nil), true
|
||||
return "", false
|
||||
}
|
||||
t.notes = notes
|
||||
noteScores := make([]float64, len(notes))
|
||||
@@ -693,11 +773,19 @@ func (h *reactiveHandler) queryKiwix(ctx context.Context, t *queryTurn) (string,
|
||||
ctxK, cancel := context.WithTimeout(ctx, kiwixTimeout)
|
||||
defer cancel()
|
||||
|
||||
// The ZIMs are English and kiwix ranks by keyword overlap, not meaning, so
|
||||
// a Russian sentence matches nothing at all. The rewriter turns it into a
|
||||
// handful of English keywords with the resident model.
|
||||
// A Russian question reads the Russian ZIM verbatim when there is one
|
||||
// (V-508). Kiwix ranks by keyword overlap rather than meaning, so an English
|
||||
// book matches a Russian sentence not at all, and the rewriter exists to
|
||||
// turn the question into English keywords with the resident model. Against a
|
||||
// Russian book that is a translation of his own words back at him: it costs
|
||||
// a model call and drops whatever the keywords do not carry.
|
||||
book, verbatim := h.kiwix.book, false
|
||||
if h.kiwix.bookRU != "" && hasCyrillic(t.dec.Utterance) {
|
||||
book, verbatim = h.kiwix.bookRU, true
|
||||
}
|
||||
|
||||
pattern := t.dec.Utterance
|
||||
if h.kiwix.rewriter != nil {
|
||||
if h.kiwix.rewriter != nil && !verbatim {
|
||||
q, err := h.kiwix.rewriter.Rewrite(ctxK, t.dec.Utterance)
|
||||
if err != nil {
|
||||
// Fall through to the verbatim question rather than give up. It
|
||||
@@ -708,7 +796,7 @@ func (h *reactiveHandler) queryKiwix(ctx context.Context, t *queryTurn) (string,
|
||||
}
|
||||
}
|
||||
|
||||
hits, err := h.kiwix.client.Search(ctxK, pattern, h.kiwix.book, h.kiwix.max)
|
||||
hits, err := h.kiwix.client.Search(ctxK, pattern, book, h.kiwix.max)
|
||||
if err != nil {
|
||||
log.Printf("voice: kiwix: search %q: %v", pattern, err)
|
||||
return "", false
|
||||
@@ -720,7 +808,7 @@ func (h *reactiveHandler) queryKiwix(ctx context.Context, t *queryTurn) (string,
|
||||
// Logged on the way through, not only on failure. Without this there is no
|
||||
// way to tell from the outside whether an answer came off a ZIM or out of
|
||||
// the model's weights, and those are the two cases worth telling apart.
|
||||
log.Printf("voice: kiwix: %q → %d hits, top %q", pattern, len(hits), top.Title)
|
||||
log.Printf("voice: kiwix: %q in %q → %d hits, top %q", pattern, book, len(hits), top.Title)
|
||||
|
||||
// The top hit only, read as an article rather than as a snippet. Kiwix
|
||||
// builds its snippet from wherever the keyword matched, which on Wikipedia
|
||||
@@ -752,6 +840,18 @@ func (h *reactiveHandler) queryKiwix(ctx context.Context, t *queryTurn) (string,
|
||||
return reply, true
|
||||
}
|
||||
|
||||
// hasCyrillic reports whether the text carries a Cyrillic letter, which is the
|
||||
// whole test for "he asked this in Russian". A question mixing a Latin proper
|
||||
// noun into a Russian sentence is still Russian, so one letter is enough.
|
||||
func hasCyrillic(s string) bool {
|
||||
for _, r := range s {
|
||||
if unicode.Is(unicode.Cyrillic, r) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// queryPersonal — stop the walk on a question about him that his own data did
|
||||
// not answer.
|
||||
//
|
||||
|
||||
@@ -2,8 +2,11 @@ package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"log"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/lexicon"
|
||||
"github.com/kami/maven/internal/phraser"
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
@@ -33,5 +36,24 @@ func (h *reactiveHandler) actionReminder(ctx context.Context, dec router.Decisio
|
||||
log.Printf("voice: create reminder: %v", err)
|
||||
return phraser.Ack(phraser.FailReminder, nil)
|
||||
}
|
||||
return ""
|
||||
// Phrased from the row, never from the utterance (Vikunja #507). The
|
||||
// replier only ever saw Slots.Text, so it named whatever hour the sentence
|
||||
// contained — including one the parser had rejected or read differently.
|
||||
// A confirmation naming an hour no row holds is worse than a clarify,
|
||||
// because he stops thinking about it.
|
||||
return reminderConfirm(dec.Slots.Time, h.now())
|
||||
}
|
||||
|
||||
// reminderConfirm — the confirmation for a reminder that exists, naming the
|
||||
// stored fire time. Deterministic on purpose: the one sentence that must match
|
||||
// a database row is not one to hand to a 1.7B.
|
||||
func reminderConfirm(fire, now time.Time) string {
|
||||
when := dayPrefix(now, fire)
|
||||
if when == "это" {
|
||||
// Further out than the day words reach — say the date instead of a
|
||||
// word that would be wrong.
|
||||
date := fmt.Sprintf("%d %s", fire.Day(), lexicon.MonthGenitive(int(fire.Month())))
|
||||
return "хорошо, напомню " + date + " в " + fire.Format("15:04") + "."
|
||||
}
|
||||
return "хорошо, напомню " + when + " в " + fire.Format("15:04") + "."
|
||||
}
|
||||
|
||||
@@ -3,6 +3,7 @@ package main
|
||||
import (
|
||||
"context"
|
||||
"log"
|
||||
"strings"
|
||||
|
||||
"github.com/kami/maven/internal/dialogue"
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
@@ -80,8 +81,96 @@ func (h *reactiveHandler) queryTasks(ctx context.Context, t *queryTurn) (string,
|
||||
for _, r := range spoken {
|
||||
cands = append(cands, dialogue.Candidate{Kind: "task", Ref: r.ID, Label: r.Text})
|
||||
}
|
||||
h.offerCandidates(cands)
|
||||
return tasks.FormatRU(ranked), true
|
||||
h.offerCandidates(ctx, cands)
|
||||
reply := tasks.FormatRU(ranked)
|
||||
// The counted shapes, after the list and only when there are any (V-512).
|
||||
// They answer "what is going wrong with this list" without assessing any of
|
||||
// it, and they are said here rather than announced: no tick rule reads them.
|
||||
if stalls := tasks.StallsRU(tasks.Stalls(taskItems(live), h.now())); stalls != "" {
|
||||
if !strings.HasSuffix(reply, ".") {
|
||||
reply += "."
|
||||
}
|
||||
reply += " " + stalls
|
||||
}
|
||||
return reply, true
|
||||
}
|
||||
|
||||
// resolveTaskStatus moves a task he named out loud (Vikunja #512).
|
||||
//
|
||||
// The position path already worked: resolveCandidate answers "первую сделал"
|
||||
// against the list she just read. This is the other half — naming the task
|
||||
// instead of its position, which reached no code at all before the stage-0 rule
|
||||
// in internal/router/taskstatus.go filled the fn slot.
|
||||
//
|
||||
// Three answers besides the move, and none of them guesses. No match says so. A
|
||||
// match on more than one asks which, because closing the wrong task is work he
|
||||
// never finished being marked done. No task named asks which too, since the
|
||||
// router claims the turn without the referent and the list lives here.
|
||||
func (h *reactiveHandler) resolveTaskStatus(ctx context.Context, dec router.Decision) string {
|
||||
live, err := h.api.ListTasks(ctx, "live")
|
||||
if err != nil {
|
||||
log.Printf("voice: task status: list: %v", err)
|
||||
return "не получилось посмотреть задачи."
|
||||
}
|
||||
if dec.Slots.Text == "" {
|
||||
return "какую задачу?"
|
||||
}
|
||||
match := matchTaskText(live, dec.Slots.Text)
|
||||
switch len(match) {
|
||||
case 0:
|
||||
return "не нашла такой задачи."
|
||||
case 1:
|
||||
default:
|
||||
return "у тебя несколько подходящих — какую именно?"
|
||||
}
|
||||
pick := match[0]
|
||||
status := dec.Slots.Value
|
||||
// A candidate is work Maven proposed and he never confirmed, and the store
|
||||
// refuses candidate → done: the legal move is to open it first. Saying it is
|
||||
// done IS the confirmation, so both writes happen rather than the turn
|
||||
// naming a gap about a distinction he did not make.
|
||||
if pick.Status == store.TaskCandidate && status == store.TaskDone {
|
||||
if err := h.api.SetTaskStatus(ctx, pick.ID, store.TaskOpen, h.now(), string(sourceVoice)); err != nil {
|
||||
log.Printf("voice: task status: promote %d: %v", pick.ID, err)
|
||||
return "не получилось изменить задачу."
|
||||
}
|
||||
}
|
||||
if err := h.api.SetTaskStatus(ctx, pick.ID, status, h.now(), string(sourceVoice)); err != nil {
|
||||
log.Printf("voice: task status: %d → %s: %v", pick.ID, status, err)
|
||||
return "не получилось изменить задачу."
|
||||
}
|
||||
log.Printf("voice: task %d (%q) → %s", pick.ID, pick.Text, status)
|
||||
if status == store.TaskDropped {
|
||||
return "убрала: " + pick.Text
|
||||
}
|
||||
return "закрыла: " + pick.Text
|
||||
}
|
||||
|
||||
// matchTaskText finds the live tasks he could have meant.
|
||||
//
|
||||
// Normalised containment, either direction, over store.NormalizeTaskText — the
|
||||
// same key capture dedupes on, so a task he can file twice is a task he can name
|
||||
// twice. Either direction because he shortens what he said ("молоко" for
|
||||
// "купить молоко") as often as he pads it.
|
||||
//
|
||||
// Deliberately not fuzzy. A ranked best guess would always return exactly one
|
||||
// answer, and the one thing this must be able to say is that it is not sure.
|
||||
func matchTaskText(live []ipc.Task, named string) []ipc.Task {
|
||||
want := store.NormalizeTaskText(named)
|
||||
if want == "" {
|
||||
return nil
|
||||
}
|
||||
var out []ipc.Task
|
||||
for _, t := range live {
|
||||
have := store.NormalizeTaskText(t.Text)
|
||||
if have == "" {
|
||||
continue
|
||||
}
|
||||
if strings.Contains(have, want) || strings.Contains(want, have) {
|
||||
out = append(out, t)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// taskItems maps wire rows onto the ranker's input. Written here rather than in
|
||||
|
||||
@@ -26,6 +26,9 @@ type taskAPI struct {
|
||||
tasks []ipc.Task
|
||||
listArg string
|
||||
listErr error
|
||||
|
||||
moved []setStatusCall
|
||||
moveErr error
|
||||
}
|
||||
|
||||
func (a *taskAPI) CaptureTask(_ context.Context, req ipc.CaptureTaskReq) (ipc.CaptureTaskResp, error) {
|
||||
@@ -253,3 +256,87 @@ func TestCaptureTaskFromNoteAcknowledgesAPromotion(t *testing.T) {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// setStatusCall — one SetTaskStatus the arm made, in order, so a candidate he
|
||||
// says is done can be shown to take both legal moves.
|
||||
type setStatusCall struct {
|
||||
id int64
|
||||
status string
|
||||
by string
|
||||
}
|
||||
|
||||
func (a *taskAPI) SetTaskStatus(_ context.Context, id int64, status string, _ time.Time, by string) error {
|
||||
a.moved = append(a.moved, setStatusCall{id: id, status: status, by: by})
|
||||
return a.moveErr
|
||||
}
|
||||
|
||||
func TestResolveTaskStatusMovesTheNamedTask(t *testing.T) {
|
||||
api := &taskAPI{tasks: []ipc.Task{
|
||||
{ID: 7, Text: "купить молоко", Status: "open"},
|
||||
{ID: 8, Text: "оплатить интернет", Status: "open"},
|
||||
}}
|
||||
h := taskHandler(api)
|
||||
reply := h.resolveTaskStatus(context.Background(), router.Decision{
|
||||
Intent: router.IntentAct,
|
||||
Slots: router.Slots{Fn: router.TaskStatusFn, HasFn: true, Value: "done", Text: "молоко"},
|
||||
})
|
||||
if api.listArg != "live" {
|
||||
t.Errorf("listed %q, want live — a resolved task cannot be resolved again", api.listArg)
|
||||
}
|
||||
if len(api.moved) != 1 {
|
||||
t.Fatalf("moved %d tasks, want 1: %+v", len(api.moved), api.moved)
|
||||
}
|
||||
if api.moved[0].id != 7 || api.moved[0].status != "done" {
|
||||
t.Errorf("moved %+v, want id 7 → done", api.moved[0])
|
||||
}
|
||||
if !strings.Contains(reply, "купить молоко") {
|
||||
t.Errorf("reply = %q, want the task named back", reply)
|
||||
}
|
||||
}
|
||||
|
||||
func TestResolveTaskStatusRefusesToGuess(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
tasks []ipc.Task
|
||||
named string
|
||||
want string
|
||||
}{
|
||||
{"no match", []ipc.Task{{ID: 7, Text: "купить молоко", Status: "open"}}, "позвонить маме", "не нашла"},
|
||||
{"two matches", []ipc.Task{
|
||||
{ID: 7, Text: "купить молоко", Status: "open"},
|
||||
{ID: 8, Text: "купить молоко и хлеб", Status: "open"},
|
||||
}, "купить молоко", "несколько"},
|
||||
{"none named", []ipc.Task{{ID: 7, Text: "купить молоко", Status: "open"}}, "", "какую"},
|
||||
}
|
||||
for _, c := range cases {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
api := &taskAPI{tasks: c.tasks}
|
||||
h := taskHandler(api)
|
||||
reply := h.resolveTaskStatus(context.Background(), router.Decision{
|
||||
Slots: router.Slots{Fn: router.TaskStatusFn, HasFn: true, Value: "done", Text: c.named},
|
||||
})
|
||||
if len(api.moved) != 0 {
|
||||
t.Errorf("moved %+v — closing the wrong task is the failure this arm exists to avoid", api.moved)
|
||||
}
|
||||
if !strings.Contains(reply, c.want) {
|
||||
t.Errorf("reply = %q, want it to contain %q", reply, c.want)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestResolveTaskStatusOpensACandidateFirst(t *testing.T) {
|
||||
// The store refuses candidate → done. Saying it is done is the confirmation
|
||||
// the candidate was waiting for, so the arm makes both legal moves.
|
||||
api := &taskAPI{tasks: []ipc.Task{{ID: 9, Text: "продлить домен", Status: "candidate"}}}
|
||||
h := taskHandler(api)
|
||||
h.resolveTaskStatus(context.Background(), router.Decision{
|
||||
Slots: router.Slots{Fn: router.TaskStatusFn, HasFn: true, Value: "done", Text: "продлить домен"},
|
||||
})
|
||||
if len(api.moved) != 2 {
|
||||
t.Fatalf("moved %+v, want open then done", api.moved)
|
||||
}
|
||||
if api.moved[0].status != "open" || api.moved[1].status != "done" {
|
||||
t.Errorf("moved %+v, want open then done", api.moved)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,186 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/delivery"
|
||||
"github.com/kami/maven/internal/loop"
|
||||
"github.com/kami/maven/internal/phraser"
|
||||
"github.com/kami/maven/internal/store"
|
||||
)
|
||||
|
||||
// The sev4 repeat path had no off switch (Vikunja #535): it re-sent every
|
||||
// pending telegram nudge every repeat_interval, and nothing in the tree could
|
||||
// ever mark one acked. None of what follows can be reproduced by hand without
|
||||
// sitting in front of the box for hours, so it is covered here or nowhere.
|
||||
|
||||
// seedDown writes one kuma monitor fact at ts. value is "down" or "up".
|
||||
func seedDown(t *testing.T, st *store.Store, ctx context.Context, value string, ts time.Time) {
|
||||
t.Helper()
|
||||
if _, err := st.SetValue(ctx, store.KindSelf, "service_down:db", "poll:uptimekuma", value, ts); err != nil {
|
||||
t.Fatalf("seed service_down:db=%s: %v", value, err)
|
||||
}
|
||||
}
|
||||
|
||||
// newAlarmTickLoop — like newTestTickLoop but with the ack tracker wired, which
|
||||
// the shared helper leaves nil. Without it RepeatUnacked returns early and the
|
||||
// repeat these tests are about never happens. The daemon wires it (main.go).
|
||||
func newAlarmTickLoop(t *testing.T, st *store.Store, sink delivery.Sink) *tickLoop {
|
||||
t.Helper()
|
||||
rules := loop.DefaultRules()
|
||||
g := loop.NewGatherer(st, rules)
|
||||
d := delivery.NewDispatcher(delivery.Config{
|
||||
Voice: sink, Ntfy: sink, Telegram: sink,
|
||||
Ack: st, Nudges: st, Reminders: st,
|
||||
})
|
||||
return newTickLoop(st, g, d, phraser.NewStub(), rules, time.Second, 5*time.Minute, 0, nil, nil, nil, nil)
|
||||
}
|
||||
|
||||
// telegramSends counts sends that went out on the telegram reach.
|
||||
func telegramSends(sink *fakeSink, rule string) int {
|
||||
n := 0
|
||||
for _, s := range sink.sends {
|
||||
if s.RuleName == rule {
|
||||
n++
|
||||
}
|
||||
}
|
||||
return n
|
||||
}
|
||||
|
||||
// outcomes returns the outcome of every nudge row for a rule, newest first.
|
||||
func outcomes(t *testing.T, st *store.Store, ctx context.Context, rule string) []string {
|
||||
t.Helper()
|
||||
rows, err := st.RecentNudges(ctx, 50)
|
||||
if err != nil {
|
||||
t.Fatalf("recent nudges: %v", err)
|
||||
}
|
||||
var out []string
|
||||
for _, n := range rows {
|
||||
if n.Rule == rule {
|
||||
out = append(out, n.Outcome)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func TestAlarmStopsWhenTheServiceComesBackUp(t *testing.T) {
|
||||
// The condition clearing is the ending that should happen. StillTrue reads
|
||||
// the same DownServices helper the phraser reads, so the repeat stops on
|
||||
// exactly the monitor he was told about.
|
||||
st := newTestStore(t)
|
||||
ctx := context.Background()
|
||||
now := refNow()
|
||||
seedDown(t, st, ctx, "down", now.Add(-5*time.Minute))
|
||||
|
||||
sink := &fakeSink{}
|
||||
tl := newAlarmTickLoop(t, st, sink)
|
||||
tl.tick(ctx, now)
|
||||
if telegramSends(sink, "service_down") == 0 {
|
||||
t.Fatal("the alarm never went out; the rest of this test proves nothing")
|
||||
}
|
||||
|
||||
seedDown(t, st, ctx, "up", now.Add(time.Minute))
|
||||
sink.sends = nil
|
||||
tl.tick(ctx, now.Add(6*time.Minute)) // past repeat_interval
|
||||
|
||||
if n := telegramSends(sink, "service_down"); n != 0 {
|
||||
t.Fatalf("repeated %d time(s) after the service came back up; want 0", n)
|
||||
}
|
||||
for _, o := range outcomes(t, st, ctx, "service_down") {
|
||||
if o != store.NudgeResolved {
|
||||
t.Fatalf("nudge outcome = %q, want %q", o, store.NudgeResolved)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestAlarmStopsAtTheAgeCapWhileStillDown(t *testing.T) {
|
||||
// Still down, still un-acked, and nobody has answered in two hours. That is
|
||||
// not one more repeat away from being answered.
|
||||
st := newTestStore(t)
|
||||
ctx := context.Background()
|
||||
now := refNow()
|
||||
seedDown(t, st, ctx, "down", now.Add(-5*time.Minute))
|
||||
|
||||
sink := &fakeSink{}
|
||||
tl := newAlarmTickLoop(t, st, sink)
|
||||
tl.tick(ctx, now)
|
||||
|
||||
sink.sends = nil
|
||||
tl.tick(ctx, now.Add(6*time.Minute))
|
||||
if n := telegramSends(sink, "service_down"); n == 0 {
|
||||
t.Fatal("no repeat inside the cap; the cap is not what stopped it later")
|
||||
}
|
||||
|
||||
sink.sends = nil
|
||||
tl.tick(ctx, now.Add(maxAlarmAge+time.Minute))
|
||||
if n := telegramSends(sink, "service_down"); n != 0 {
|
||||
t.Fatalf("repeated %d time(s) past the %s cap; want 0", n, maxAlarmAge)
|
||||
}
|
||||
// Ignored, not resolved: nothing says the service got better.
|
||||
for _, o := range outcomes(t, st, ctx, "service_down") {
|
||||
if o != store.NudgeIgnored {
|
||||
t.Fatalf("nudge outcome = %q, want %q", o, store.NudgeIgnored)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestAFlapRaisesAFreshAlarmRatherThanReviveTheClosedOne(t *testing.T) {
|
||||
// Down, up, down again. Closing the first run must not make the second run
|
||||
// unreportable, and must not silently reopen the closed rows either.
|
||||
st := newTestStore(t)
|
||||
ctx := context.Background()
|
||||
now := refNow()
|
||||
seedDown(t, st, ctx, "down", now.Add(-5*time.Minute))
|
||||
|
||||
sink := &fakeSink{}
|
||||
tl := newAlarmTickLoop(t, st, sink)
|
||||
tl.tick(ctx, now)
|
||||
first := len(outcomes(t, st, ctx, "service_down"))
|
||||
|
||||
seedDown(t, st, ctx, "up", now.Add(time.Minute))
|
||||
tl.tick(ctx, now.Add(2*time.Minute))
|
||||
if got := outcomes(t, st, ctx, "service_down"); len(got) != first {
|
||||
t.Fatalf("closing the run changed the row count: %d → %d", first, len(got))
|
||||
}
|
||||
|
||||
seedDown(t, st, ctx, "down", now.Add(25*time.Minute))
|
||||
sink.sends = nil
|
||||
tl.tick(ctx, now.Add(31*time.Minute))
|
||||
|
||||
if n := telegramSends(sink, "service_down"); n == 0 {
|
||||
t.Fatal("the second outage said nothing; the first alarm's ending swallowed it")
|
||||
}
|
||||
got := outcomes(t, st, ctx, "service_down")
|
||||
if len(got) <= first {
|
||||
t.Fatalf("no new nudge row for the second outage (%d rows, was %d)", len(got), first)
|
||||
}
|
||||
}
|
||||
|
||||
func TestARuleThatSaysNothingAboutItsConditionOnlyStopsOnAge(t *testing.T) {
|
||||
// StillTrue == nil means "I cannot tell you", never "it cleared". A rule
|
||||
// that says nothing must keep its alarm until the age cap, or a rule author
|
||||
// silences their own alarm by omission.
|
||||
st := newTestStore(t)
|
||||
ctx := context.Background()
|
||||
now := refNow()
|
||||
|
||||
sink := &fakeSink{}
|
||||
tl := newAlarmTickLoop(t, st, sink)
|
||||
tl.rules = []loop.Rule{{Name: "mute", Severity: loop.Sev4}} // no StillTrue
|
||||
|
||||
if _, err := st.RecordNudge(ctx, "mute", string(delivery.ChannelTelegram), "still bad", now); err != nil {
|
||||
t.Fatalf("record nudge: %v", err)
|
||||
}
|
||||
|
||||
live := tl.stopFinishedAlarms(ctx, []string{"mute"}, loop.State{}, now.Add(time.Minute))
|
||||
if len(live) != 1 {
|
||||
t.Fatalf("a nil StillTrue was read as resolved: live = %v", live)
|
||||
}
|
||||
|
||||
live = tl.stopFinishedAlarms(ctx, []string{"mute"}, loop.State{}, now.Add(maxAlarmAge+time.Minute))
|
||||
if len(live) != 0 {
|
||||
t.Fatalf("the age cap did not stop a rule with no StillTrue: live = %v", live)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,125 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// An empty attention list used to be answered "ничего не требует внимания"
|
||||
// unconditionally, which is an all-clear Maven had no way to know was true
|
||||
// (ECOSYSTEM-SPEC §2.6, Vikunja #540).
|
||||
func TestAttentionEmptyWithHealthySourcesIsAllClear(t *testing.T) {
|
||||
praxis := newFakePraxisWithSources(t, `[]`, `[{"source_id":"src_ntfy","health":"ok"}]`)
|
||||
h := newPraxisTestHandler(t, praxis)
|
||||
|
||||
reply := h.handlePraxisAct(context.Background(), praxisActDec("list_attention"))
|
||||
if !strings.Contains(reply, "ничего не требует внимания") {
|
||||
t.Fatalf("healthy and quiet should be an all-clear, got %q", reply)
|
||||
}
|
||||
}
|
||||
|
||||
func TestAttentionEmptyWithAFailedSourceHedges(t *testing.T) {
|
||||
praxis := newFakePraxisWithSources(t, `[]`, `[
|
||||
{"source_id":"src_ntfy","health":"ok"},
|
||||
{"source_id":"src_llamacpp","health":"failed"},
|
||||
{"source_id":"src_imap","health":"stale"}
|
||||
]`)
|
||||
h := newPraxisTestHandler(t, praxis)
|
||||
|
||||
reply := h.handlePraxisAct(context.Background(), praxisActDec("list_attention"))
|
||||
if strings.Contains(reply, "ничего не требует внимания") {
|
||||
t.Fatalf("a failed source must not read as all-clear, got %q", reply)
|
||||
}
|
||||
for _, want := range []string{"src_llamacpp", "src_imap"} {
|
||||
if !strings.Contains(reply, want) {
|
||||
t.Errorf("reply names no %s: %q", want, reply)
|
||||
}
|
||||
}
|
||||
if strings.Contains(reply, "src_ntfy") {
|
||||
t.Errorf("the healthy source is named as a problem: %q", reply)
|
||||
}
|
||||
}
|
||||
|
||||
// A Praxis that polls nothing knows nothing, which is the state the box is in.
|
||||
func TestAttentionEmptyWithNoSourcesHedges(t *testing.T) {
|
||||
praxis := newFakePraxisWithSources(t, `[]`, `[]`)
|
||||
h := newPraxisTestHandler(t, praxis)
|
||||
|
||||
reply := h.handlePraxisAct(context.Background(), praxisActDec("list_attention"))
|
||||
if strings.Contains(reply, "ничего не требует внимания") {
|
||||
t.Fatalf("a Praxis with no sources must not answer all-clear, got %q", reply)
|
||||
}
|
||||
if !strings.Contains(reply, "источник") {
|
||||
t.Errorf("reply does not say why she cannot tell: %q", reply)
|
||||
}
|
||||
}
|
||||
|
||||
// The spec's own mechanism, which the deployed Praxis does not send yet: the
|
||||
// envelope's degraded array is believed without a second call.
|
||||
func TestAttentionDegradedEnvelopeIsReadWithoutASourcesCall(t *testing.T) {
|
||||
praxis := newFakePraxisWithSources(t,
|
||||
`{"items":[],"degraded":["src_metrics"]}`,
|
||||
`[{"source_id":"src_ntfy","health":"ok"}]`)
|
||||
h := newPraxisTestHandler(t, praxis)
|
||||
|
||||
reply := h.handlePraxisAct(context.Background(), praxisActDec("list_attention"))
|
||||
if !strings.Contains(reply, "src_metrics") {
|
||||
t.Fatalf("the envelope's degraded source is not named: %q", reply)
|
||||
}
|
||||
for _, r := range praxis.Requests() {
|
||||
if r.Path == "/api/v1/sources" {
|
||||
t.Error("sources was read even though the response carried degraded")
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A sources endpoint that errors is not evidence of a fault: the attention call
|
||||
// itself succeeded, and hedging on it would make her permanently uncertain.
|
||||
func TestAttentionKeepsAllClearWhenSourcesCannotBeRead(t *testing.T) {
|
||||
praxis := newFakePraxisWithSources(t, `[]`, `[]`)
|
||||
praxis.SetRouteFault("/api/v1/sources", 500)
|
||||
h := newPraxisTestHandler(t, praxis)
|
||||
|
||||
reply := h.handlePraxisAct(context.Background(), praxisActDec("list_attention"))
|
||||
if !strings.Contains(reply, "ничего не требует внимания") {
|
||||
t.Fatalf("an unreadable sources list should leave the answer alone, got %q", reply)
|
||||
}
|
||||
}
|
||||
|
||||
// Both response shapes decode, because the spec says one and the box sends the
|
||||
// other.
|
||||
func TestPraxisAttentionDecodesBothShapes(t *testing.T) {
|
||||
var bare praxisAttention
|
||||
if err := json.Unmarshal([]byte(`[{"id":"item_1"}]`), &bare); err != nil {
|
||||
t.Fatalf("bare array: %v", err)
|
||||
}
|
||||
if len(bare.Items) != 1 || len(bare.Degraded) != 0 {
|
||||
t.Errorf("bare array decoded as %+v", bare)
|
||||
}
|
||||
|
||||
var env praxisAttention
|
||||
if err := json.Unmarshal([]byte(`{"items":[{"id":"item_2"}],"degraded":["src_a"]}`), &env); err != nil {
|
||||
t.Fatalf("envelope: %v", err)
|
||||
}
|
||||
if len(env.Items) != 1 || len(env.Degraded) != 1 || env.Degraded[0] != "src_a" {
|
||||
t.Errorf("envelope decoded as %+v", env)
|
||||
}
|
||||
}
|
||||
|
||||
// A source that reports no health at all counts as healthy. A Praxis that never
|
||||
// fills the field would otherwise make every quiet turn a hedge.
|
||||
func TestUnhealthySourcesTreatsAnAbsentHealthFieldAsHealthy(t *testing.T) {
|
||||
praxis := newFakePraxisWithSources(t, `[]`, `[{"source_id":"src_a"},{"id":"src_b","health":"stale"}]`)
|
||||
bad, total, err := newPraxisClient(praxis.URL).UnhealthySources(context.Background())
|
||||
if err != nil {
|
||||
t.Fatalf("UnhealthySources: %v", err)
|
||||
}
|
||||
if total != 2 {
|
||||
t.Errorf("total = %d, want 2", total)
|
||||
}
|
||||
if len(bad) != 1 || bad[0] != "src_b" {
|
||||
t.Errorf("bad = %v, want [src_b]", bad)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,61 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
|
||||
// TestCalendarStepsAsideForTheWorld — the defect (Vikunja #552). Weather was
|
||||
// one instance of a wider class, and V-474 fixed only that instance. Every one
|
||||
// of these answered "на 05.08.2026 ничего нет" on the deployed daemon, and
|
||||
// every one of them has an answer in search, which sits below the calendar.
|
||||
func TestCalendarStepsAsideForTheWorld(t *testing.T) {
|
||||
h, api := contQueryHandler()
|
||||
for _, u := range []string{
|
||||
"во сколько закат сегодня",
|
||||
"какой сегодня курс доллара",
|
||||
"какой сегодня праздник",
|
||||
"что интересного произошло сегодня в мире",
|
||||
} {
|
||||
if reply, ok := h.queryCalendar(context.Background(), &queryTurn{
|
||||
dec: router.Decision{Intent: router.IntentQuery, Utterance: u},
|
||||
}); ok {
|
||||
t.Errorf("the calendar claimed %q with %q", u, reply)
|
||||
}
|
||||
}
|
||||
if api.events != 0 {
|
||||
t.Errorf("CalendarEvents called %d times for world questions, want 0", api.events)
|
||||
}
|
||||
}
|
||||
|
||||
// The other half of the same narrowing: a question about his own day still
|
||||
// reaches the calendar, including the one that names no subject at all.
|
||||
func TestCalendarStillAnswersHisDay(t *testing.T) {
|
||||
for _, u := range []string{
|
||||
"что у меня сегодня",
|
||||
"во сколько у меня встреча сегодня",
|
||||
"какие встречи завтра",
|
||||
"что в календаре на завтра",
|
||||
"что сегодня?",
|
||||
} {
|
||||
h, _ := contQueryHandler()
|
||||
if _, ok := h.queryCalendar(context.Background(), &queryTurn{
|
||||
dec: router.Decision{Intent: router.IntentQuery, Utterance: u},
|
||||
}); !ok {
|
||||
t.Errorf("the calendar passed on %q", u)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A continuation carries its subject in the turn before it, and the calendar
|
||||
// is the only date-aware source, so the narrowing must not reach it.
|
||||
func TestCalendarStillAnswersAContinuation(t *testing.T) {
|
||||
h, _ := contQueryHandler()
|
||||
if _, ok := h.queryCalendar(context.Background(), &queryTurn{
|
||||
dec: router.Decision{Intent: router.IntentQuery, Utterance: "а завтра?", Continued: true},
|
||||
}); !ok {
|
||||
t.Error("the calendar passed on a continuation")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,50 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/llm"
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
|
||||
// The chat path must answer when llama-server is down (Vikunja #45 step 4).
|
||||
// Both halves of a turn call the model — the router and the replier — and each
|
||||
// has its own floor: the cascade falls to the classifier, the replier falls to
|
||||
// the stub. This wires a client at a closed port so both floors are exercised
|
||||
// by a dial error rather than by a stubbed error value.
|
||||
func TestChatAnswersWithNoLlamaServer(t *testing.T) {
|
||||
h, _, _ := newClarifyHandler(t)
|
||||
dead := llm.New("http://127.0.0.1:1", 500*time.Millisecond)
|
||||
emb := router.NewHashEmbedder(1024)
|
||||
h.recall.embedder = emb
|
||||
h.router = buildRouter(emb, h.matcher, 0.55, pickLLMRouter(true, dead))
|
||||
h.replier = newLLMReplier(dead, nil)
|
||||
|
||||
ctx := withDialogueID(context.Background(), dialogueIDFor(sourceText, "web"))
|
||||
for _, utt := range []string{
|
||||
"привет",
|
||||
"запиши что кофе закончился",
|
||||
"что у меня сегодня",
|
||||
} {
|
||||
reply := h.handleText(ctx, "web", utt)
|
||||
if reply == "" {
|
||||
t.Errorf("%q answered with nothing; a dead model must degrade to the stub", utt)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// daemonAPI.Chat reports an error only when the voice path was never wired.
|
||||
// A turn that reaches handleText always carries text, which is what keeps
|
||||
// mavweb's /api/chat off its error branch when the model is down.
|
||||
func TestChatAPIErrsOnlyWhenUnwired(t *testing.T) {
|
||||
d := &daemonAPI{}
|
||||
if _, err := d.Chat(context.Background(), "web", "привет"); err == nil {
|
||||
t.Fatal("an unwired daemon must say so")
|
||||
}
|
||||
d.chatFn = func(context.Context, string, string) string { return "" }
|
||||
if _, err := d.Chat(context.Background(), "web", "привет"); err != nil {
|
||||
t.Fatalf("a wired daemon must not error: %v", err)
|
||||
}
|
||||
}
|
||||
+69
-12
@@ -156,6 +156,11 @@ func (h *reactiveHandler) askClarify(ctx context.Context, dec router.Decision) (
|
||||
if !ok {
|
||||
return "", false
|
||||
}
|
||||
// An act she could not resolve is a refusal, not a question (Vikunja #556).
|
||||
if slot == dialogue.SlotFn {
|
||||
log.Printf("voice: clarify — act %q matched no capability; saying so instead of asking", dec.Utterance)
|
||||
return actNotRecognized, true
|
||||
}
|
||||
h.clarifyStore.Put(dialogueIDOf(ctx), &dialogue.PendingQuestion{
|
||||
Intent: dialogue.Intent(dec.Intent),
|
||||
Slots: toDialogueSlots(dec.Slots),
|
||||
@@ -170,14 +175,31 @@ func (h *reactiveHandler) askClarify(ctx context.Context, dec router.Decision) (
|
||||
return question, true
|
||||
}
|
||||
|
||||
// resolveClarifyAnswer reads an utterance as the answer to a parked question.
|
||||
// Returns ("", false) when no live question is parked (or it expired), so the
|
||||
// caller routes the utterance normally as a fresh request. Sibling of
|
||||
// resolveConfirm and checked in the same place.
|
||||
// clarifyCancelled — he called the half-built request off. Said out loud, like
|
||||
// every other way it can end: a silent drop reads as "done". Feminine
|
||||
// self-reference ("отменила"), as everywhere.
|
||||
const clarifyCancelled = "Хорошо, отменила."
|
||||
|
||||
// clarifyDropped — he asked for something else instead, so the parked request
|
||||
// is gone. Glued in front of the answer to what he actually asked, because
|
||||
// nothing may be dropped in silence. V-561 suspends and resumes it instead of
|
||||
// letting it go, and this line goes away with it.
|
||||
const clarifyDropped = "Прошлую просьбу отпускаю."
|
||||
|
||||
// resolveClarifyAnswer reads an utterance against the parked question and
|
||||
// decides what it IS before deciding what to do with it. Returns ("", false)
|
||||
// when the turn is not this resolver's — nothing parked, or the utterance turned
|
||||
// out to be a request of its own — so the caller dispatches it normally.
|
||||
//
|
||||
// The answer is parsed with the same extractor the router uses, for the intent
|
||||
// she parked — no second parser. If it still does not fill the gap she asks
|
||||
// again, up to MaxAttempts; after that she says out loud that she did not
|
||||
// The order is the point (Vikunja #560). The utterance is ROUTED first, and the
|
||||
// role is read off that decision: a routed decision that stands on its own is
|
||||
// not an answer, whatever the extractor found inside it. Before this the
|
||||
// extractor decided, so "какая сейчас погода в Риме?" became the time of a
|
||||
// reminder on the strength of the word "сейчас".
|
||||
//
|
||||
// The answer itself is parsed with the same extractor the router uses, for the
|
||||
// intent she parked — no second parser. If it still does not fill the gap she
|
||||
// asks again, up to MaxAttempts; after that she says out loud that she did not
|
||||
// understand. She never drops the request in silence.
|
||||
func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string) (string, bool) {
|
||||
if h.clarifyStore == nil {
|
||||
@@ -185,11 +207,36 @@ func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string)
|
||||
}
|
||||
q := h.clarifyStore.Get(dialogueIDOf(ctx), h.now())
|
||||
if q == nil {
|
||||
return "", false
|
||||
return "", false // not_applicable: nothing is pending
|
||||
}
|
||||
|
||||
intent := router.Intent(q.Intent)
|
||||
answer := h.extractor.Extract(ctx, intent, text, h.now())
|
||||
|
||||
var (
|
||||
routed router.Decision
|
||||
routedOK bool
|
||||
)
|
||||
if needsRoute(text) {
|
||||
routed, routedOK = h.routeForRole(ctx, text)
|
||||
}
|
||||
role := classifyTurnRole(q, text, toDialogueSlots(answer), routed, routedOK)
|
||||
log.Printf("voice: clarify — %q is a %s against %s (routed=%v)", text, role, dialogue.CapabilityFor(q.Intent), routedOK)
|
||||
|
||||
switch role {
|
||||
case roleCancel:
|
||||
h.clarifyStore.Delete(dialogueIDOf(ctx))
|
||||
return clarifyCancelled, true
|
||||
case roleSideQuery, roleNewRequest:
|
||||
// He moved on. A parked question used to swallow whatever came next, so
|
||||
// one act she could not fulfil ate the following three turns (Vikunja
|
||||
// #554) and a world question set a reminder for a time nobody asked for
|
||||
// (#558). Drop the question, say so, and let these words be themselves.
|
||||
h.clarifyStore.Delete(dialogueIDOf(ctx))
|
||||
h.noteDropped(ctx)
|
||||
return "", false
|
||||
}
|
||||
|
||||
merged := q.Answer(text, toDialogueSlots(answer))
|
||||
// Fold a newly answered subject into the raw utterance. Downstream actions
|
||||
// phrase from Utterance, not from the text slot — actionReminder stores it
|
||||
@@ -226,6 +273,16 @@ func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string)
|
||||
return h.finishClarified(ctx, dec), true
|
||||
}
|
||||
|
||||
// noteDropped records that the parked request was let go this turn, so runTurn
|
||||
// can say it in front of whatever these words are answered with. Nothing to
|
||||
// record outside runTurn — a unit test calling one resolver has no turn to glue
|
||||
// a notice onto.
|
||||
func (h *reactiveHandler) noteDropped(ctx context.Context) {
|
||||
if rt := turnRouteFrom(ctx); rt != nil {
|
||||
rt.dropped = clarifyDropped
|
||||
}
|
||||
}
|
||||
|
||||
// foldAnswerIntoUtterance appends an answered subject to the original words,
|
||||
// unless they already carry it. "напомни" + "позвонить маме" reads as the
|
||||
// request he would have made in one breath. Nothing is appended when the
|
||||
@@ -305,9 +362,9 @@ func (h *reactiveHandler) reaskOrGiveUp(ctx context.Context, q *dialogue.Pending
|
||||
func (h *reactiveHandler) finishClarified(ctx context.Context, dec router.Decision) string {
|
||||
if h.dialogueSessions != nil {
|
||||
now := h.now()
|
||||
prev := h.dialogueSessions.Get(voiceDialogueID, now)
|
||||
prev := h.dialogueSessions.Get(dialogueIDOf(ctx), now)
|
||||
dec = followUpMerge(prev, dec, now)
|
||||
h.rememberTurn(prev, dec, now)
|
||||
h.rememberTurn(ctx, prev, dec, now)
|
||||
}
|
||||
reply := h.applyAction(ctx, dec)
|
||||
if reply == "" {
|
||||
@@ -323,7 +380,7 @@ func (h *reactiveHandler) finishClarified(ctx context.Context, dec router.Decisi
|
||||
// rememberTurn stores this turn as the dialogue session the next follow-up
|
||||
// inherits from, carrying up to 4 prior turns of history for anaphora. Capped so
|
||||
// one long conversation can't grow the session unboundedly.
|
||||
func (h *reactiveHandler) rememberTurn(prev *dialogue.Session, dec router.Decision, now time.Time) {
|
||||
func (h *reactiveHandler) rememberTurn(ctx context.Context, prev *dialogue.Session, dec router.Decision, now time.Time) {
|
||||
var history []dialogue.Turn
|
||||
if prev != nil {
|
||||
history = append(history, dialogue.Turn{
|
||||
@@ -359,7 +416,7 @@ func (h *reactiveHandler) rememberTurn(prev *dialogue.Session, dec router.Decisi
|
||||
if !dec.Continued && (dec.Intent == router.IntentSystem || dec.Intent == router.IntentQuery) {
|
||||
slots.Text = dec.Utterance
|
||||
}
|
||||
h.dialogueSessions.Put(voiceDialogueID, &dialogue.Session{
|
||||
h.dialogueSessions.Put(dialogueIDOf(ctx), &dialogue.Session{
|
||||
Intent: dialogue.Intent(dec.Intent),
|
||||
Slots: slots,
|
||||
Timestamp: now,
|
||||
|
||||
+184
-25
@@ -10,6 +10,7 @@ import (
|
||||
|
||||
"github.com/kami/maven/internal/dialogue"
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/memory"
|
||||
"github.com/kami/maven/internal/phraser/eval"
|
||||
"github.com/kami/maven/internal/router"
|
||||
"github.com/kami/maven/internal/store"
|
||||
@@ -231,34 +232,35 @@ func TestClarifyRestatedAnswerWins(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestClarifiedActOffAllowlistIsStillRefused — clarification fills in an
|
||||
// argument, it never grants authority.
|
||||
func TestClarifiedActOffAllowlistIsStillRefused(t *testing.T) {
|
||||
// TestActOffAllowlistIsStillRefused — naming a capability is not being granted
|
||||
// one. Since Vikunja #556 an unresolved act no longer parks a question, so this
|
||||
// goes through applyAction, which is the only way an act runs.
|
||||
func TestActOffAllowlistIsStillRefused(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, st, _ := newClarifyHandler(t)
|
||||
marker := filepath.Join(t.TempDir(), "not-allowed-ran")
|
||||
if err := st.EnableTool(ctx, "uptime", []string{"true"}, false, "test", h.now()); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
if _, asked := h.askClarify(ctx, clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это")); !asked {
|
||||
t.Fatal("an act with no fn should be asked about")
|
||||
}
|
||||
reply, handled := h.resolveClarifyAnswer(ctx, "rm "+marker)
|
||||
if !handled {
|
||||
t.Fatal("the answer should be consumed")
|
||||
}
|
||||
reply := h.applyAction(ctx, router.Decision{
|
||||
Utterance: "rm " + marker,
|
||||
Intent: router.IntentAct,
|
||||
Slots: router.Slots{Fn: "rm " + marker, HasFn: true},
|
||||
})
|
||||
if strings.Contains(reply, "готово") {
|
||||
t.Fatalf("an act that is not on the allowlist must not report success: %q", reply)
|
||||
}
|
||||
if _, err := os.Stat(marker); !os.IsNotExist(err) {
|
||||
t.Fatalf("a clarified act off the allowlist ran anyway: %v", err)
|
||||
}
|
||||
if tools, err := st.ListTools(ctx, "enabled"); err != nil || len(tools) != 0 {
|
||||
t.Fatalf("clarify must not enable a tool: tools=%+v err=%v", tools, err)
|
||||
if tools, err := st.ListTools(ctx, "enabled"); err != nil || len(tools) != 1 {
|
||||
t.Fatalf("an act must not enable a tool: tools=%+v err=%v", tools, err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestClarifiedDestructiveActStillNeedsConfirm — the confirm gate survives the
|
||||
// clarify path.
|
||||
func TestClarifiedDestructiveActStillNeedsConfirm(t *testing.T) {
|
||||
// TestDestructiveActStillNeedsConfirm — the confirm gate stands on the act path.
|
||||
func TestDestructiveActStillNeedsConfirm(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, st, _ := newClarifyHandler(t)
|
||||
marker := filepath.Join(t.TempDir(), "destructive-ran")
|
||||
@@ -266,18 +268,16 @@ func TestClarifiedDestructiveActStillNeedsConfirm(t *testing.T) {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
if _, asked := h.askClarify(ctx, clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это")); !asked {
|
||||
t.Fatal("expected a question")
|
||||
}
|
||||
reply, handled := h.resolveClarifyAnswer(ctx, "delete_backups")
|
||||
if !handled {
|
||||
t.Fatal("the answer should be consumed")
|
||||
}
|
||||
reply := h.applyAction(ctx, router.Decision{
|
||||
Utterance: "delete_backups",
|
||||
Intent: router.IntentAct,
|
||||
Slots: router.Slots{Fn: "delete_backups", HasFn: true},
|
||||
})
|
||||
if !strings.Contains(reply, "да") || h.pending == nil {
|
||||
t.Fatalf("a clarified destructive act must still park a confirm: reply=%q pending=%+v", reply, h.pending)
|
||||
t.Fatalf("a destructive act must park a confirm: reply=%q pending=%+v", reply, h.pending)
|
||||
}
|
||||
if _, err := os.Stat(marker); !os.IsNotExist(err) {
|
||||
t.Fatalf("a clarified destructive act ran before confirmation: %v", err)
|
||||
t.Fatalf("a destructive act ran before confirmation: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -419,7 +419,7 @@ func TestClarifyProseHoldsThePersona(t *testing.T) {
|
||||
eval.CheckAddress: true,
|
||||
eval.CheckCringe: true,
|
||||
}
|
||||
lines := append([]string{clarifyGaveUp}, clarifyExpiredVariants...)
|
||||
lines := append([]string{clarifyGaveUp, clarifyCancelled, clarifyDropped}, clarifyExpiredVariants...)
|
||||
lines = append(lines, clarifyMissedVariants...)
|
||||
for _, variants := range clarifyQuestionVariants {
|
||||
lines = append(lines, variants...)
|
||||
@@ -564,3 +564,162 @@ func TestARestartExpiresTheParkedQuestion(t *testing.T) {
|
||||
t.Fatalf("notice = %q, want silence: nothing survived to expire", notice)
|
||||
}
|
||||
}
|
||||
|
||||
// TestClarifyStepsAsideForItsOwnRequest — Vikunja #554. An act she could not
|
||||
// fulfil parked "Что сделать?", and the three turns after it were scored as
|
||||
// answers to that question: a world question, then "как дела", then the give-up
|
||||
// line. None of them was ever an answer.
|
||||
func TestClarifyStepsAsideForItsOwnRequest(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, _, _ := newClarifyHandler(t)
|
||||
|
||||
// A reminder, not the act this bug was found on: since Vikunja #556 an act
|
||||
// no longer parks anything, so it can no longer eat the turn after it.
|
||||
if _, asked := h.askClarify(ctx, clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме")); !asked {
|
||||
t.Fatal("a reminder with no time should be asked about")
|
||||
}
|
||||
if reply, handled := h.resolveClarifyAnswer(ctx, "кто изобрёл телефон"); handled {
|
||||
t.Fatalf("a world question must route as itself, got %q", reply)
|
||||
}
|
||||
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
|
||||
t.Error("the parked question must be dropped, not left to eat the turn after this one")
|
||||
}
|
||||
}
|
||||
|
||||
// TestClarifyStillRetriesOnAnAnswerThatMissed — the other half of #554, and the
|
||||
// reason the test above is narrow. A bare noun answers nothing either, but it
|
||||
// carries no request of its own, so she asks again as before.
|
||||
func TestClarifyStillRetriesOnAnAnswerThatMissed(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, _, _ := newClarifyHandler(t)
|
||||
|
||||
if _, asked := h.askClarify(ctx, clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме")); !asked {
|
||||
t.Fatal("expected the time question")
|
||||
}
|
||||
reply, handled := h.resolveClarifyAnswer(ctx, "ага")
|
||||
if !handled || reply == "" {
|
||||
t.Fatalf("a missed answer must still be re-asked, handled=%v reply=%q", handled, reply)
|
||||
}
|
||||
if h.clarifyStore.Get(voiceDialogueID, h.now()) == nil {
|
||||
t.Error("the question must survive a missed answer")
|
||||
}
|
||||
}
|
||||
|
||||
// TestClarifyQuestionShapedAnswerThatFillsTheGapStillLands — the guard runs only
|
||||
// where nothing was filled. "во сколько?" is question-shaped and is also how a
|
||||
// time gets said back, so an answer that closes the gap wins whatever its shape.
|
||||
func TestClarifyQuestionShapedAnswerThatFillsTheGapStillLands(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, st, _ := newClarifyHandler(t)
|
||||
|
||||
if _, asked := h.askClarify(ctx, clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме")); !asked {
|
||||
t.Fatal("expected the time question")
|
||||
}
|
||||
if reply, handled := h.resolveClarifyAnswer(ctx, "а что если в 11:00"); !handled || reply == clarifyGaveUp {
|
||||
t.Fatalf("an answer that fills the gap must land, handled=%v reply=%q", handled, reply)
|
||||
}
|
||||
if reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour)); err != nil || len(reminders) != 1 {
|
||||
t.Fatalf("reminder was not created: reminders=%v err=%v", reminders, err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestUnresolvedActSaysItDoesNotKnowTheCommand — Vikunja #556. "Что сделать?"
|
||||
// has no answer he can give, so an act that matched no capability is refused in
|
||||
// one line and nothing is parked. It does not recite what she can do instead.
|
||||
func TestUnresolvedActSaysItDoesNotKnowTheCommand(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, st, _ := newClarifyHandler(t)
|
||||
// Enabled tools change nothing here: this act matched none of them.
|
||||
if err := st.EnableTool(ctx, "uptime", []string{"true"}, false, "test", h.now()); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
reply, spoken := h.askClarify(ctx, clarifyDec(router.IntentAct, router.Slots{Text: "выключи свет"}, "выключи свет"))
|
||||
if !spoken || reply != actNotRecognized {
|
||||
t.Fatalf("reply = %q spoken=%v, want %q", reply, spoken, actNotRecognized)
|
||||
}
|
||||
if strings.Contains(reply, "uptime") {
|
||||
t.Errorf("reply = %q, want no list of capabilities he did not ask about", reply)
|
||||
}
|
||||
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
|
||||
t.Error("nothing to ask about, so nothing may be parked")
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
|
||||
// newRoutingClarifyHandler wires the real cascade (hash embedder, no model) onto
|
||||
// the clarify handler, so a test can drive handleText end to end and see which
|
||||
// gate claimed the turn.
|
||||
func newRoutingClarifyHandler(t *testing.T) (*reactiveHandler, *store.Store) {
|
||||
t.Helper()
|
||||
h, st, _ := newClarifyHandler(t)
|
||||
h.router = buildRouter(router.NewHashEmbedder(1024), h.matcher, 0.55, nil)
|
||||
h.recall = recallWiring{embedder: router.NewHashEmbedder(1024), memStore: memory.NewInMemoryStore()}
|
||||
return h, st
|
||||
}
|
||||
|
||||
// TestIncompleteReminderAsksInsteadOfFailing — Vikunja #557. "напомни позвонить"
|
||||
// is routed confidently and is still half a request. It used to reach applyAction,
|
||||
// fail on the missing time and park nothing, so the "в семь вечера" that followed
|
||||
// was routed as a world question and web-searched.
|
||||
func TestIncompleteReminderAsksInsteadOfFailing(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, st := newRoutingClarifyHandler(t)
|
||||
|
||||
reply := h.handleText(ctx, "web", "напомни позвонить маме")
|
||||
want, _ := clarifyQuestionFor(dialogue.SlotTime, 1)
|
||||
if reply != want {
|
||||
t.Fatalf("reply = %q, want the time question %q", reply, want)
|
||||
}
|
||||
if h.clarifyStore.Get(dialogueIDFor(sourceText, "web"), h.now()) == nil {
|
||||
t.Fatal("the request must be parked, or the answer has nowhere to land")
|
||||
}
|
||||
if reply := h.handleText(ctx, "web", "в семь вечера"); strings.Contains(reply, "нашла") {
|
||||
t.Fatalf("the answer to her own question must not be looked up: %q", reply)
|
||||
}
|
||||
if reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour)); err != nil || len(reminders) != 1 {
|
||||
t.Fatalf("the answer did not complete the reminder: reminders=%v err=%v", reminders, err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestBareCaptureVerbAsksWhatToRecord — the other half of #557. A bare "запиши"
|
||||
// went to the resident model as chat, which agreed to a wording change nobody
|
||||
// asked for. It is a fact with no key, and that gap has a question.
|
||||
func TestBareCaptureVerbAsksWhatToRecord(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, _ := newRoutingClarifyHandler(t)
|
||||
|
||||
reply := h.handleText(ctx, "web", "запиши")
|
||||
want, _ := clarifyQuestionFor(dialogue.SlotKey, 1)
|
||||
if reply != want {
|
||||
t.Fatalf("reply = %q, want %q", reply, want)
|
||||
}
|
||||
if h.clarifyStore.Get(dialogueIDFor(sourceText, "web"), h.now()) == nil {
|
||||
t.Fatal("the request must be parked so the next utterance completes it")
|
||||
}
|
||||
}
|
||||
|
||||
// TestACompleteTurnStillDoesNotAsk — the gate reads a missing slot, not any
|
||||
// slot, so a request she can act on must never turn into a question. Checked on
|
||||
// the decision rather than through the cascade: what is at stake is the gate's
|
||||
// condition, and driving it through the hash embedder would measure routing.
|
||||
func TestACompleteTurnStillDoesNotAsk(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, _, _ := newClarifyHandler(t)
|
||||
|
||||
complete := []router.Decision{
|
||||
{Intent: router.IntentReminder, Slots: router.Slots{Text: "позвонить маме", HasTime: true}, Utterance: "напомни в 11 позвонить маме"},
|
||||
{Intent: router.IntentFact, Slots: router.Slots{Key: "water", Value: "выпил", HasKey: true}, Utterance: "я выпил воды"},
|
||||
{Intent: router.IntentNote, Slots: router.Slots{Text: "купить хлеб"}, Utterance: "запиши купить хлеб"},
|
||||
{Intent: router.IntentQuery, Slots: router.Slots{Text: "что у меня сегодня"}, Utterance: "что у меня сегодня"},
|
||||
}
|
||||
for _, dec := range complete {
|
||||
if gaps := missingFor(dec); len(gaps) > 0 {
|
||||
t.Errorf("%q reads as incomplete: %v", dec.Utterance, gaps)
|
||||
}
|
||||
if reply, asked := h.askClarify(ctx, dec); asked {
|
||||
t.Errorf("%q was answered with a question: %q", dec.Utterance, reply)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -40,6 +40,9 @@ var clarifyQuestionVariants = map[dialogue.Slot][]string{
|
||||
"Что именно отметить?",
|
||||
"Назови, что записать — например, «выпил воды».",
|
||||
},
|
||||
// Not spoken since Vikunja #556: askClarify answers actNotRecognized for a
|
||||
// missing capability rather than asking. Kept because clarifyQuestion still
|
||||
// reports the gap, and a re-ask deck with a hole in it is harder to read.
|
||||
dialogue.SlotFn: {
|
||||
"Что сделать?",
|
||||
"Какое действие выполнить?",
|
||||
@@ -47,6 +50,17 @@ var clarifyQuestionVariants = map[dialogue.Slot][]string{
|
||||
},
|
||||
}
|
||||
|
||||
// actNotRecognized is what an act she cannot run gets (Vikunja #556).
|
||||
//
|
||||
// The deck used to ask "Что сделать?" instead. That question has no answer he
|
||||
// can give: he already said what he wanted, and nothing he repeats will match a
|
||||
// capability that is not there. So she asked, failed, asked again and gave up —
|
||||
// three turns spent on one refusal. She says it once now, and parks nothing.
|
||||
//
|
||||
// It does not recite the allowlist. A list of names he did not ask about is not
|
||||
// an answer to the thing he did ask about.
|
||||
const actNotRecognized = "Такую команду я не знаю."
|
||||
|
||||
// clarifyQuestionFor picks the wording for this attempt. attempt is 1-based, as
|
||||
// PendingQuestion.Attempts counts it; anything past the list uses the last and
|
||||
// most explicit phrasing rather than wrapping round to the short one, because
|
||||
|
||||
+112
-33
@@ -3,9 +3,14 @@ package main
|
||||
import (
|
||||
"context"
|
||||
"log"
|
||||
"slices"
|
||||
"sort"
|
||||
"strings"
|
||||
"time"
|
||||
"unicode"
|
||||
|
||||
"github.com/kami/maven/internal/lexicon"
|
||||
"github.com/kami/maven/internal/phraser"
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
|
||||
@@ -56,10 +61,18 @@ func (h *reactiveHandler) park(fn string, args []string, phrase string) {
|
||||
// resolveConfirm interprets an utterance as the answer to a parked destructive
|
||||
// act OR a parked routine proposal. Returns (reply, true) when it consumed the
|
||||
// utterance as a y/n answer; ("", false) when there's nothing pending (or the
|
||||
// parked act expired), so the caller routes the utterance normally. An
|
||||
// unrecognised answer cancels the pending and routes normally — a confirm that
|
||||
// can't be answered clearly is safer abandoned than left armed.
|
||||
// parked act expired), so the caller routes the utterance normally.
|
||||
//
|
||||
// An utterance that is not clearly yes or no is not an answer at all, so it is
|
||||
// handed straight back and the pending stays parked until it expires (V-567).
|
||||
// This resolver runs before routing and holds the most dangerous trigger on the
|
||||
// box; it may only claim a turn it is certain about.
|
||||
func (h *reactiveHandler) resolveConfirm(ctx context.Context, text string) (string, bool) {
|
||||
verdict := classifyConfirm(text)
|
||||
if verdict == confirmUnknown {
|
||||
return "", false
|
||||
}
|
||||
|
||||
h.mu.Lock()
|
||||
defer h.mu.Unlock()
|
||||
|
||||
@@ -67,16 +80,12 @@ func (h *reactiveHandler) resolveConfirm(ctx context.Context, text string) (stri
|
||||
if !r.claim() {
|
||||
continue
|
||||
}
|
||||
// The slot is already cleared by claim(): every branch below drops the
|
||||
// pending, including the unclear one — a confirm that can't be
|
||||
// answered clearly is safer abandoned than left armed.
|
||||
switch classifyConfirm(text) {
|
||||
// The slot is already cleared by claim().
|
||||
switch verdict {
|
||||
case confirmYes:
|
||||
return r.yes(), true
|
||||
case confirmNo:
|
||||
return r.no(), true
|
||||
default:
|
||||
return "", false
|
||||
return r.no(), true
|
||||
}
|
||||
}
|
||||
return "", false
|
||||
@@ -118,13 +127,13 @@ func (h *reactiveHandler) confirmResolvers(ctx context.Context) []confirmResolve
|
||||
//
|
||||
// Acceptance itself is recorded by /routines, and the tick
|
||||
// loop nudges on the interval from there (Vikunja #366).
|
||||
return "поняла — подтверди на странице рутин, и начну напоминать."
|
||||
return phraser.C(phraser.ConfirmRoutineAuthed, nil)
|
||||
},
|
||||
no: func() string {
|
||||
if err := h.dataStore.DismissProposedRoutine(ctx, pr.routineID); err != nil {
|
||||
log.Printf("voice: dismiss proposed routine: %v", err)
|
||||
}
|
||||
return "хорошо, не буду."
|
||||
return phraser.C(phraser.ConfirmRoutineNo, nil)
|
||||
},
|
||||
},
|
||||
// Hexis execution confirm. Bound to the exact capability + target that
|
||||
@@ -137,7 +146,7 @@ func (h *reactiveHandler) confirmResolvers(ctx context.Context) []confirmResolve
|
||||
yes: func() string {
|
||||
return h.execHexis(ctx, hx.capabilityID, hx.capName, hx.entityID, hx.displayName)
|
||||
},
|
||||
no: func() string { return "отменила." },
|
||||
no: func() string { return phraser.C(phraser.ConfirmCancelled, nil) },
|
||||
},
|
||||
// Tool confirm.
|
||||
{
|
||||
@@ -150,16 +159,16 @@ func (h *reactiveHandler) confirmResolvers(ctx context.Context) []confirmResolve
|
||||
if err != nil {
|
||||
log.Printf("voice: tool %s (confirmed): %v", p.fn, err)
|
||||
if out != "" {
|
||||
return "не получилось выполнить команду: " + firstLine(out)
|
||||
return phraser.A(phraser.ActFailOut, map[string]string{"out": firstLine(out)})
|
||||
}
|
||||
return "не получилось выполнить команду."
|
||||
return phraser.A(phraser.ActFail, nil)
|
||||
}
|
||||
if out != "" {
|
||||
return "готово: " + firstLine(out)
|
||||
return phraser.A(phraser.ActDoneOut, map[string]string{"out": firstLine(out)})
|
||||
}
|
||||
return "готово."
|
||||
return phraser.A(phraser.ActDone, nil)
|
||||
},
|
||||
no: func() string { return "отменила." },
|
||||
no: func() string { return phraser.C(phraser.ConfirmCancelled, nil) },
|
||||
},
|
||||
}
|
||||
}
|
||||
@@ -170,17 +179,18 @@ func (h *reactiveHandler) confirmResolvers(ctx context.Context) []confirmResolve
|
||||
func (h *reactiveHandler) proposeGap(ctx context.Context, dec router.Decision) string {
|
||||
name := firstWord(stripWake(dec.Utterance))
|
||||
if name == "" {
|
||||
return "не разобрала команду — попробуй иначе."
|
||||
return phraser.C(phraser.ProposeNoVerb, nil)
|
||||
}
|
||||
vars := map[string]string{"name": name}
|
||||
newly, err := h.api.ProposeTool(ctx, name, dec.Utterance, "", h.now())
|
||||
if err != nil {
|
||||
log.Printf("voice: propose tool %q: %v", name, err)
|
||||
return "команды «" + name + "» нет в списке разрешённых."
|
||||
return phraser.C(phraser.ProposeFailed, vars)
|
||||
}
|
||||
if newly {
|
||||
return "команды «" + name + "» нет в списке. Предложила её добавить — включи через клиент."
|
||||
return phraser.C(phraser.ProposeNew, vars)
|
||||
}
|
||||
return "команды «" + name + "» пока нет в списке — она уже предложена, включи через клиент."
|
||||
return phraser.C(phraser.ProposeAlready, vars)
|
||||
}
|
||||
|
||||
// confirmVerdict — the parse of a y/n confirm answer.
|
||||
@@ -192,23 +202,92 @@ const (
|
||||
confirmNo
|
||||
)
|
||||
|
||||
// classifyConfirm reads a short ru/en yes-or-no answer. Substring match on the
|
||||
// stems so inflections/fillers ("да, давай", "нет, отмени") still land.
|
||||
// confirmWords are the two closed sets, tokenized once and ordered
|
||||
// longest-first so "не надо" is read before "нет" could claim any of it.
|
||||
var (
|
||||
confirmYesPhrases = confirmPhrases(lexicon.ConfirmYes())
|
||||
confirmNoPhrases = confirmPhrases(lexicon.ConfirmNo())
|
||||
)
|
||||
|
||||
// confirmPhrases splits each lexicon member into tokens and sorts the result
|
||||
// longest-first, so a walk that tries them in order matches the longest member
|
||||
// that fits.
|
||||
func confirmPhrases(words []string) [][]string {
|
||||
out := make([][]string, 0, len(words))
|
||||
for _, w := range words {
|
||||
if toks := confirmTokens(w); len(toks) > 0 {
|
||||
out = append(out, toks)
|
||||
}
|
||||
}
|
||||
sort.SliceStable(out, func(i, j int) bool { return len(out[i]) > len(out[j]) })
|
||||
return out
|
||||
}
|
||||
|
||||
// confirmTokens splits an utterance into lowercase word tokens. Punctuation and
|
||||
// spacing are separators; an apostrophe is not, because "don't" is one word.
|
||||
func confirmTokens(text string) []string {
|
||||
return strings.FieldsFunc(strings.ToLower(text), func(r rune) bool {
|
||||
if r == '\'' || r == '’' {
|
||||
return false
|
||||
}
|
||||
return !unicode.IsLetter(r) && !unicode.IsDigit(r)
|
||||
})
|
||||
}
|
||||
|
||||
// classifyConfirm reads a short ru/en yes-or-no answer to a parked confirm.
|
||||
//
|
||||
// The whole utterance must consist of confirmation words and filler, matched as
|
||||
// whole tokens against the closed lexicon sets. Anything else is
|
||||
// confirmUnknown, which leaves the confirm parked and routes the turn — see
|
||||
// resolveConfirm. Both halves of that are the fix for V-567: this used to be a
|
||||
// substring test over bare stems, so "погода", "дальше", "надо" and "давление"
|
||||
// all read as "да", and "покажи" and "около" read as "ок". A parked destructive
|
||||
// act fired on a question about the weather.
|
||||
//
|
||||
// Requiring the WHOLE utterance is the second half. A leading confirm word does
|
||||
// not make a sentence an answer: "давай посмотрим погоду" opens a request, and
|
||||
// the only safe reading of a sentence that carries its own subject is that he
|
||||
// moved on. Guessing wrong here executes something; guessing wrong the other way
|
||||
// asks again.
|
||||
func classifyConfirm(text string) confirmVerdict {
|
||||
t := strings.ToLower(strings.TrimSpace(text))
|
||||
// negatives first — "не надо" contains no "да", but check no-stems before
|
||||
// yes so a leading "нет" isn't shadowed.
|
||||
for _, no := range []string{"нет", "не надо", "отмен", "стоп", "no", "cancel", "stop", "don't"} {
|
||||
if strings.Contains(t, no) {
|
||||
tokens := confirmTokens(text)
|
||||
if len(tokens) == 0 {
|
||||
return confirmUnknown
|
||||
}
|
||||
verdict := confirmUnknown
|
||||
for i := 0; i < len(tokens); {
|
||||
// Negatives first: "не надо" and "не хочу" open with a token that is
|
||||
// not itself an answer, and a yes hit must never shadow them.
|
||||
if n := matchConfirm(confirmNoPhrases, tokens[i:]); n > 0 {
|
||||
return confirmNo
|
||||
}
|
||||
if n := matchConfirm(confirmYesPhrases, tokens[i:]); n > 0 {
|
||||
verdict, i = confirmYes, i+n
|
||||
continue
|
||||
}
|
||||
if lexicon.IsFillerParticle(tokens[i]) {
|
||||
i++
|
||||
continue
|
||||
}
|
||||
// A word that is neither an answer nor filler carries a subject of its
|
||||
// own, so this utterance is not an answer to her question.
|
||||
return confirmUnknown
|
||||
}
|
||||
for _, yes := range []string{"да", "ага", "давай", "подтвер", "конечно", "yes", "yeah", "yep", "confirm", "ок", "okay", "ok"} {
|
||||
if strings.Contains(t, yes) {
|
||||
return confirmYes
|
||||
return verdict
|
||||
}
|
||||
|
||||
// matchConfirm reports the length of the longest phrase matching at the head of
|
||||
// tokens, or 0.
|
||||
func matchConfirm(phrases [][]string, tokens []string) int {
|
||||
for _, p := range phrases {
|
||||
if len(p) > len(tokens) {
|
||||
continue
|
||||
}
|
||||
if slices.Equal(p, tokens[:len(p)]) {
|
||||
return len(p)
|
||||
}
|
||||
}
|
||||
return confirmUnknown
|
||||
return 0
|
||||
}
|
||||
|
||||
// actPhrase renders "fn arg1 arg2" for the confirm prompt.
|
||||
|
||||
@@ -0,0 +1,103 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/tool"
|
||||
)
|
||||
|
||||
// TestClassifyConfirmRejectsSubstrings — V-567. The old matcher tested bare
|
||||
// stems with strings.Contains, so every word below answered a question she had
|
||||
// asked about something else: "погода", "дальше", "надо" and "давление" carry
|
||||
// "да"; "покажи", "около" and "окно" carry "ок". A parked destructive act fired
|
||||
// on a question about the weather.
|
||||
func TestClassifyConfirmRejectsSubstrings(t *testing.T) {
|
||||
for _, text := range []string{
|
||||
"погода",
|
||||
"какая погода",
|
||||
"что дальше",
|
||||
"надо ещё",
|
||||
"покажи заметки",
|
||||
"около окна",
|
||||
"давление",
|
||||
"давай посмотрим погоду",
|
||||
"не забудь купить хлеб",
|
||||
"окно открыто",
|
||||
"стоит ли брать зонт",
|
||||
"",
|
||||
} {
|
||||
if got := classifyConfirm(text); got != confirmUnknown {
|
||||
t.Errorf("classifyConfirm(%q) = %v, want confirmUnknown", text, got)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestClassifyConfirmAcceptsAnswers keeps every genuine answer the substring
|
||||
// matcher accepted, and pins the pair the fix could most easily get wrong:
|
||||
// "надо" is not an answer and "не надо" is the opposite of one.
|
||||
func TestClassifyConfirmAcceptsAnswers(t *testing.T) {
|
||||
yes := []string{"да", "Да!", "ага", "давай", "да, давай", "конечно", "подтверждаю", "ну да", "yes", "yeah", "ok", "okay", "confirm"}
|
||||
no := []string{"нет", "Нет.", "не надо", "не нужно", "не сейчас", "отмена", "отмени", "стоп", "нет, отмени", "no", "nope", "cancel", "stop", "don't"}
|
||||
|
||||
for _, text := range yes {
|
||||
if got := classifyConfirm(text); got != confirmYes {
|
||||
t.Errorf("classifyConfirm(%q) = %v, want confirmYes", text, got)
|
||||
}
|
||||
}
|
||||
for _, text := range no {
|
||||
if got := classifyConfirm(text); got != confirmNo {
|
||||
t.Errorf("classifyConfirm(%q) = %v, want confirmNo", text, got)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestUnrelatedTurnLeavesConfirmParked — the whole point of V-567. An utterance
|
||||
// that is not an answer must not execute the parked act, must not consume the
|
||||
// turn, and must not disarm the confirm either: the answer he has not given yet
|
||||
// is still answerable until it expires.
|
||||
func TestUnrelatedTurnLeavesConfirmParked(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
st := newTestStore(t)
|
||||
api := ipc.NewStoreAPI(st)
|
||||
now := time.Date(2026, 8, 6, 9, 0, 0, 0, time.UTC)
|
||||
h := &reactiveHandler{
|
||||
api: api,
|
||||
dataStore: st,
|
||||
now: func() time.Time { return now },
|
||||
tools: tool.NewExecutor(api, time.Second),
|
||||
}
|
||||
|
||||
marker := filepath.Join(t.TempDir(), "destructive-tool-ran")
|
||||
if err := st.EnableTool(ctx, "delete_backups", []string{"touch", marker}, true, "test", h.now()); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
h.park("delete_backups", nil, "delete_backups")
|
||||
|
||||
if reply, handled := h.resolveConfirm(ctx, "какая погода"); handled {
|
||||
t.Fatalf("the weather question was consumed as a confirm: %q", reply)
|
||||
}
|
||||
if _, err := os.Stat(marker); !os.IsNotExist(err) {
|
||||
t.Fatalf("the parked destructive command ran on an unrelated turn: %v", err)
|
||||
}
|
||||
if h.pending == nil {
|
||||
t.Fatal("the confirm was disarmed by a turn that did not answer it")
|
||||
}
|
||||
|
||||
// It is still answerable, and answering it still runs the act.
|
||||
reply, handled := h.resolveConfirm(ctx, "да")
|
||||
if !handled || !strings.Contains(reply, "готово") {
|
||||
t.Fatalf("the still-parked confirm did not resolve: handled=%v reply=%q", handled, reply)
|
||||
}
|
||||
if _, err := os.Stat(marker); err != nil {
|
||||
t.Fatalf("confirmed destructive command did not run: %v", err)
|
||||
}
|
||||
if h.pending != nil {
|
||||
t.Fatal("the confirm stayed parked after being answered")
|
||||
}
|
||||
}
|
||||
@@ -1,6 +1,7 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
@@ -149,12 +150,13 @@ func TestRememberTurnRefreshesTheTopic(t *testing.T) {
|
||||
now: func() time.Time { return contNow },
|
||||
dialogueSessions: dialogue.NewSessionStore(2 * time.Minute),
|
||||
}
|
||||
h.rememberTurn(nil, router.Decision{
|
||||
ctx := context.Background()
|
||||
h.rememberTurn(ctx, nil, router.Decision{
|
||||
Intent: router.IntentQuery, Utterance: "во сколько у меня встреча",
|
||||
}, contNow)
|
||||
// The second turn arrives with the first turn's Text already merged in.
|
||||
prev := h.dialogueSessions.Get(voiceDialogueID, contNow)
|
||||
h.rememberTurn(prev, router.Decision{
|
||||
h.rememberTurn(ctx, prev, router.Decision{
|
||||
Intent: router.IntentQuery,
|
||||
Utterance: "какие у меня планы",
|
||||
Slots: router.Slots{Text: "во сколько у меня встреча"},
|
||||
|
||||
@@ -0,0 +1,132 @@
|
||||
// mavend/decisiontrace.go — the daemon's half of the per-turn decision record.
|
||||
//
|
||||
// V-564. The router says what the cascade did (internal/router/decisiontrace.go);
|
||||
// this file covers the two claimant sets that live in the daemon: the stateful
|
||||
// resolvers that run BEFORE routing and pre-empt it unconditionally, and the
|
||||
// query source chain that runs after. Those two are where the arbitration is
|
||||
// least visible, because both are a hardcoded order of functions that each
|
||||
// answer "is this mine?" alone and none of which answers "is this more mine
|
||||
// than yours?" (V-558).
|
||||
//
|
||||
// Recording changes no route. Every helper here is a no-op on a context with no
|
||||
// record, which is what every test that does not ask for one gets.
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
|
||||
"github.com/kami/maven/internal/decision"
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
|
||||
// preRouteLadder — the resolvers runTurn offers the utterance to before the
|
||||
// router sees it, in the order they get their say. Kept here as a roster rather
|
||||
// than derived from the code, so a resolver that returns early and skips the
|
||||
// rest still leaves the rest NAMED in the record: a claimant that never looked
|
||||
// and one that looked and passed are the distinction the ordering hides, and
|
||||
// they are the difference between a bug in the ladder and a bug in a resolver.
|
||||
//
|
||||
// Adding a step to runTurn means adding its name here. Nothing enforces that,
|
||||
// and nothing should: a missing name costs one line of the record, while a
|
||||
// check that walks the ladder would have to run the ladder.
|
||||
var preRouteLadder = []string{
|
||||
"confirm", "clarify-answer", "quiet-toggle", "snooze", "ack", "repair", "ordinal",
|
||||
}
|
||||
|
||||
// notePreRoute records one rung of that ladder and passes its verdict through
|
||||
// unchanged, so the call site stays the single `if handled` it already was.
|
||||
func notePreRoute(ctx context.Context, name string, handled bool) bool {
|
||||
rec := decision.From(ctx)
|
||||
if rec == nil {
|
||||
return handled
|
||||
}
|
||||
if handled {
|
||||
rec.Note(decision.Claim{
|
||||
Stage: decision.StagePreRoute, Claimant: name, Outcome: decision.Won,
|
||||
Reason: "it pre-empted routing, so the router never saw this turn",
|
||||
})
|
||||
return handled
|
||||
}
|
||||
rec.Note(decision.Claim{
|
||||
Stage: decision.StagePreRoute, Claimant: name, Outcome: decision.Declined,
|
||||
Reason: "nothing of its own was pending",
|
||||
})
|
||||
return handled
|
||||
}
|
||||
|
||||
// noteTerminal records whoever actually produced the reply, but only if the
|
||||
// turn is still unclaimed. A route decides the intent; it does not answer, and
|
||||
// on a thinned route or a plain act nothing downstream keeps a scoreboard. So
|
||||
// the record would otherwise close with an empty winner, which reads as a lost
|
||||
// turn instead of an asked question.
|
||||
func noteTerminal(ctx context.Context, claimant string, intent router.Intent, reason string) {
|
||||
decision.From(ctx).NoteIfUnclaimed(decision.Claim{
|
||||
Stage: decision.StageAction, Claimant: claimant,
|
||||
Intent: string(intent), Reason: reason,
|
||||
})
|
||||
}
|
||||
|
||||
// noteMerge records the follow-up merge, which is the one claimant that edits
|
||||
// the winning decision instead of taking the turn from it. It is compared on
|
||||
// the four slots the merge can fill, because a Decision holds a slice and is
|
||||
// not comparable.
|
||||
func noteMerge(ctx context.Context, before, after router.Decision) {
|
||||
rec := decision.From(ctx)
|
||||
if rec == nil {
|
||||
return
|
||||
}
|
||||
changed := before.Slots.HasTime != after.Slots.HasTime ||
|
||||
before.Slots.HasKey != after.Slots.HasKey ||
|
||||
before.Slots.HasFn != after.Slots.HasFn ||
|
||||
before.Slots.Text != after.Slots.Text ||
|
||||
before.Intent != after.Intent
|
||||
if !changed {
|
||||
rec.Note(decision.Claim{
|
||||
Stage: decision.StageMerge, Claimant: "follow-up-merge", Outcome: decision.Declined,
|
||||
Reason: "no slot of this turn was left for a previous one to fill",
|
||||
})
|
||||
return
|
||||
}
|
||||
rec.Note(decision.Claim{
|
||||
Stage: decision.StageMerge, Claimant: "follow-up-merge", Intent: string(after.Intent),
|
||||
Outcome: decision.Merged, Reason: "filled this turn's gaps from the previous turn",
|
||||
})
|
||||
}
|
||||
|
||||
// turnDecisionsFn — the reader mavweb gets, or nil when voice was never wired.
|
||||
// Same shape as intakeEventsFn: the daemon holds the ring, the IPC layer only
|
||||
// converts it.
|
||||
func turnDecisionsFn(w *voiceWiring) func(int) []ipc.TurnDecision {
|
||||
if w == nil || w.handler == nil || w.handler.decisions == nil {
|
||||
return nil
|
||||
}
|
||||
ring := w.handler.decisions
|
||||
return func(n int) []ipc.TurnDecision {
|
||||
recs := ring.Recent(n)
|
||||
out := make([]ipc.TurnDecision, 0, len(recs))
|
||||
for _, rec := range recs {
|
||||
claims := make([]ipc.TurnClaim, 0, len(rec.Claims))
|
||||
for _, c := range rec.Claims {
|
||||
claims = append(claims, ipc.TurnClaim{
|
||||
Stage: c.Stage, Claimant: c.Claimant, Intent: c.Intent,
|
||||
Score: c.Score, HasScore: c.HasScore,
|
||||
Outcome: c.Outcome, Reason: c.Reason,
|
||||
})
|
||||
}
|
||||
out = append(out, ipc.TurnDecision{
|
||||
Ts: rec.Ts, Utterance: rec.Utterance, Winner: rec.Winner, Claims: claims,
|
||||
})
|
||||
}
|
||||
return out
|
||||
}
|
||||
}
|
||||
|
||||
// querySourceNames — the query chain's roster, in chain order.
|
||||
func querySourceNames() []string {
|
||||
names := make([]string, len(querySources))
|
||||
for i, src := range querySources {
|
||||
names[i] = src.name
|
||||
}
|
||||
return names
|
||||
}
|
||||
@@ -0,0 +1,163 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/decision"
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/memory"
|
||||
"github.com/kami/maven/internal/router"
|
||||
"github.com/kami/maven/internal/tool"
|
||||
"github.com/kami/maven/internal/voice"
|
||||
)
|
||||
|
||||
// traceHandler — a handler with the decision ring wired, the same shape the
|
||||
// daemon builds in wireVoice.
|
||||
func traceHandler(t *testing.T, ring *decision.Ring) *reactiveHandler {
|
||||
t.Helper()
|
||||
st := newTestStore(t)
|
||||
api := ipc.NewStoreAPI(st)
|
||||
now := time.Now()
|
||||
emb := router.NewHashEmbedder(1024)
|
||||
return &reactiveHandler{
|
||||
api: api,
|
||||
recall: recallWiring{embedder: emb, memStore: memory.NewInMemoryStore()},
|
||||
router: buildRouter(emb, tool.NewMatcher(api), 0.55, nil),
|
||||
replier: voice.NewStubReplier(),
|
||||
now: func() time.Time { return now },
|
||||
dataStore: st,
|
||||
decisions: ring,
|
||||
}
|
||||
}
|
||||
|
||||
// findClaim — the first claim for a claimant, or nil.
|
||||
func findClaim(rec *decision.Record, claimant string) *decision.Claim {
|
||||
for i := range rec.Claims {
|
||||
if rec.Claims[i].Claimant == claimant {
|
||||
return &rec.Claims[i]
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// TestTurnRecordNamesWinnerAndLosers — the point of V-564. A turn a stage-0
|
||||
// grammar claims must leave a record naming that grammar as the winner, naming
|
||||
// a pre-route resolver that declined, and naming the routing engines that were
|
||||
// never reached at all. The last of those is the fact the hardcoded ordering
|
||||
// hides: "classifier" absent from the record and "classifier" never asked read
|
||||
// the same to a human, and only one of them is the truth.
|
||||
func TestTurnRecordNamesWinnerAndLosers(t *testing.T) {
|
||||
ring := decision.NewRing()
|
||||
h := traceHandler(t, ring)
|
||||
|
||||
reply := h.handleText(context.Background(), "web", "сколько сейчас времени")
|
||||
if reply == "" {
|
||||
t.Fatal("turn produced no reply")
|
||||
}
|
||||
|
||||
recs := ring.Recent(5)
|
||||
if len(recs) != 1 {
|
||||
t.Fatalf("want 1 record, got %d", len(recs))
|
||||
}
|
||||
rec := recs[0]
|
||||
if rec.Utterance != "сколько сейчас времени" {
|
||||
t.Errorf("utterance = %q", rec.Utterance)
|
||||
}
|
||||
if !strings.HasPrefix(rec.Winner, "stage0:") {
|
||||
t.Errorf("want a stage-0 grammar as the winner, got %q", rec.Winner)
|
||||
}
|
||||
|
||||
// A loser that examined the turn: the confirm resolver ran first and had
|
||||
// nothing pending.
|
||||
confirm := findClaim(rec, "confirm")
|
||||
if confirm == nil || confirm.Outcome != decision.Declined {
|
||||
t.Errorf("confirm claim = %+v, want a decline", confirm)
|
||||
}
|
||||
|
||||
// A loser that never looked: stage 0 answered, so neither routing engine
|
||||
// was reached.
|
||||
for _, name := range []string{"llm-router", "classifier"} {
|
||||
c := findClaim(rec, name)
|
||||
if c != nil && c.Outcome == decision.Won {
|
||||
t.Errorf("%s cannot have won a stage-0 turn: %+v", name, c)
|
||||
}
|
||||
}
|
||||
|
||||
// And every rung of the ladder below the winner is named, not omitted.
|
||||
for _, name := range preRouteLadder {
|
||||
if findClaim(rec, name) == nil {
|
||||
t.Errorf("ladder rung %q is missing from the record", name)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestRecordingDoesNotChangeTheReply — instrumentation, so a turn with the ring
|
||||
// wired and the same turn without it must answer identically. If this ever
|
||||
// fails, a claim site is doing more than noting.
|
||||
func TestRecordingDoesNotChangeTheReply(t *testing.T) {
|
||||
for _, utt := range []string{
|
||||
"сколько сейчас времени",
|
||||
"запиши что я пил воду",
|
||||
"что у меня сегодня",
|
||||
} {
|
||||
withRing := traceHandler(t, decision.NewRing()).handleText(context.Background(), "web", utt)
|
||||
without := traceHandler(t, nil).handleText(context.Background(), "web", utt)
|
||||
if withRing != without {
|
||||
t.Errorf("%q: recorded reply %q != unrecorded %q", utt, withRing, without)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestQueryChainRecordsWhoWasNeverAsked — a query source below the claimant is
|
||||
// never consulted, and the record must say so rather than leave it out. This is
|
||||
// the arm that would have explained the Rome misroute in one read.
|
||||
func TestQueryChainRecordsWhoWasNeverAsked(t *testing.T) {
|
||||
ring := decision.NewRing()
|
||||
h := traceHandler(t, ring)
|
||||
h.handleText(context.Background(), "web", "что у меня сегодня")
|
||||
|
||||
rec := ring.Recent(1)[0]
|
||||
var asked, never int
|
||||
for _, c := range rec.Claims {
|
||||
if c.Stage != decision.StageQuery {
|
||||
continue
|
||||
}
|
||||
if c.Outcome == decision.NeverAsked {
|
||||
never++
|
||||
} else {
|
||||
asked++
|
||||
}
|
||||
}
|
||||
if asked == 0 {
|
||||
t.Fatal("no query source reported at all")
|
||||
}
|
||||
if never == 0 {
|
||||
t.Fatal("no query source was recorded as never asked; the chain cannot have run to the end")
|
||||
}
|
||||
if got := len(querySourceNames()); asked+never != got {
|
||||
t.Errorf("record covers %d of %d query sources", asked+never, got)
|
||||
}
|
||||
}
|
||||
|
||||
// TestTurnDecisionsFnConvertsTheRing — the IPC read path. Nil when voice was
|
||||
// never wired, because a box with no turns is an empty page and not an error.
|
||||
func TestTurnDecisionsFnConvertsTheRing(t *testing.T) {
|
||||
if fn := turnDecisionsFn(nil); fn != nil {
|
||||
t.Error("no wiring should mean no reader")
|
||||
}
|
||||
ring := decision.NewRing()
|
||||
h := traceHandler(t, ring)
|
||||
h.handleText(context.Background(), "web", "сколько сейчас времени")
|
||||
|
||||
fn := turnDecisionsFn(&voiceWiring{handler: h})
|
||||
if fn == nil {
|
||||
t.Fatal("wired handler produced no reader")
|
||||
}
|
||||
out := fn(10)
|
||||
if len(out) != 1 || out[0].Winner == "" || len(out[0].Claims) == 0 {
|
||||
t.Fatalf("conversion lost the record: %+v", out)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,556 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"log"
|
||||
"os"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/dialogue"
|
||||
"github.com/kami/maven/internal/memory"
|
||||
"github.com/kami/maven/internal/router"
|
||||
"github.com/kami/maven/internal/store"
|
||||
)
|
||||
|
||||
// Dialogue contract tests (V-563, child of V-558).
|
||||
//
|
||||
// Every other clarify test is single-shot: one ask, one answer, one assertion.
|
||||
// Three bugs of the same family shipped in two days that way — V-554 (a parked
|
||||
// question ate the three turns after it), V-557 (a confidently routed but
|
||||
// incomplete reminder parked nothing, so the answer was web-searched) and the
|
||||
// Rome case in V-558 (a side question was eaten as the time answer). None of
|
||||
// them is visible in one turn. The dialogue path is a state machine, so it can
|
||||
// be enumerated instead: whole traces, each with a per-turn expectation and an
|
||||
// expected END state — what was written to the store, and what is still parked.
|
||||
//
|
||||
// Two rules for the rows below.
|
||||
//
|
||||
// Where today's behaviour is correct, it is asserted. Where it is WRONG, the row
|
||||
// carries the CORRECT expectation and is skipped with the Vikunja id that will
|
||||
// unskip it. A weakened expectation would be worse than no row: it would pin the
|
||||
// bug as the contract.
|
||||
//
|
||||
// Everything runs on the offline floor — hash embedder, no llama-server, no
|
||||
// ONNX, StubDateTimeParser. That has one consequence worth knowing before
|
||||
// reading a fire time here: the stub reads "в 11:00" and "через час" and does
|
||||
// not read "на 9" or "на завтра", so a trace that needs those is noted where it
|
||||
// sits.
|
||||
|
||||
// claim — which claimant consumed an utterance. Not asserted: it is derived from
|
||||
// the log lines the daemon already emits and printed on every failure, because
|
||||
// "the reply differed" does not distinguish a wrong claimant from wrong copy,
|
||||
// and that distinction is the whole point of V-558.
|
||||
type claim struct {
|
||||
utterance string
|
||||
steps []string
|
||||
}
|
||||
|
||||
func (c claim) String() string { return c.utterance + " ⇒ " + strings.Join(c.steps, " → ") }
|
||||
|
||||
// claimMarkers — log fragment to claimant name, in the order runTurn checks
|
||||
// them. The fragments are the daemon's own words (clarify.go, repair.go,
|
||||
// voice.go); a rename there shows up here as an "unclaimed" step rather than a
|
||||
// silent mislabel.
|
||||
var claimMarkers = []struct{ fragment, name string }{
|
||||
{"parked question expired", "clarify:expired"},
|
||||
{"is its own request", "clarify:stepped-aside"},
|
||||
{"gave up on", "clarify:gave-up"},
|
||||
{"did not fill", "clarify:re-ask"},
|
||||
{"one gap filled", "clarify:ask-second-gap"},
|
||||
{"asked about", "clarify:ask"},
|
||||
{"repair —", "repair"},
|
||||
{"route result: intent=", "route"},
|
||||
}
|
||||
|
||||
// claimsOf reads the turn's log output and names the claimants that touched it.
|
||||
func claimsOf(utterance, logged string) claim {
|
||||
c := claim{utterance: utterance}
|
||||
for _, line := range strings.Split(logged, "\n") {
|
||||
for _, m := range claimMarkers {
|
||||
if strings.Contains(line, m.fragment) {
|
||||
name := m.name
|
||||
if m.name == "route" {
|
||||
name = "route:" + intentInLine(line)
|
||||
}
|
||||
c.steps = append(c.steps, name)
|
||||
break
|
||||
}
|
||||
}
|
||||
}
|
||||
if len(c.steps) == 0 {
|
||||
c.steps = []string{"unclaimed"}
|
||||
}
|
||||
return c
|
||||
}
|
||||
|
||||
func intentInLine(line string) string {
|
||||
_, rest, ok := strings.Cut(line, "intent=")
|
||||
if !ok {
|
||||
return "?"
|
||||
}
|
||||
intent, _, _ := strings.Cut(rest, " ")
|
||||
return intent
|
||||
}
|
||||
|
||||
// parkedWant — the question that must be armed after a turn. Attempt matters:
|
||||
// a claimant that spends a retry on an utterance that was never an answer is
|
||||
// exactly the V-554 shape, and the count is the only place it shows.
|
||||
type parkedWant struct {
|
||||
slot dialogue.Slot
|
||||
attempt int
|
||||
// carries — a substring the parked utterance must still hold, so a re-park
|
||||
// that lost the answered subject fails here rather than three turns later.
|
||||
carries string
|
||||
}
|
||||
|
||||
// turn — one utterance and everything that must be true right after it.
|
||||
type turn struct {
|
||||
say string
|
||||
// wait — the clock moves this far BEFORE the utterance. The only way to
|
||||
// reach the TTL without sleeping.
|
||||
wait time.Duration
|
||||
// question — the reply must be exactly this clarify question, worded for
|
||||
// this attempt. Zero slot ⇒ not checked.
|
||||
question dialogue.Slot
|
||||
attempt int
|
||||
contains []string
|
||||
notContain []string
|
||||
// noQuestion — the reply must not be any clarify question. Used where the
|
||||
// correct behaviour is known but her wording for it is not written yet: a
|
||||
// cancel must not be answered with another question, whatever it does say.
|
||||
noQuestion bool
|
||||
expired bool // the reply must open with the TTL notice
|
||||
// parked — what is armed after the turn. nil ⇒ nothing may be armed.
|
||||
parked *parkedWant
|
||||
}
|
||||
|
||||
// endState — what the store holds once the trace is over. Counts and
|
||||
// substrings, not rows: a trace is about who claimed what, and a payload
|
||||
// substring is enough to catch a request landing under the wrong words.
|
||||
type endState struct {
|
||||
reminders []reminderWant
|
||||
factKeys []string
|
||||
notes int
|
||||
tasks []string
|
||||
}
|
||||
|
||||
type reminderWant struct {
|
||||
payload string // substring of the stored payload
|
||||
fireAt string // "2006-01-02 15:04" in UTC, "" ⇒ not checked
|
||||
}
|
||||
|
||||
// trace — a named conversation, its turns, and the end state.
|
||||
type trace struct {
|
||||
name string
|
||||
skip string // non-empty ⇒ t.Skip: today's behaviour is wrong, this names the fix
|
||||
turns []turn
|
||||
end endState
|
||||
}
|
||||
|
||||
// newDialogueHandler — the offline floor with the real cascade and a movable
|
||||
// clock: newClarifyHandler's wiring (stub date parser, real fact parser, tool
|
||||
// matcher) plus the router newRoutingClarifyHandler builds, and the `now`
|
||||
// pointer so a turn can carry a wait.
|
||||
func newDialogueHandler(t *testing.T) (*reactiveHandler, *store.Store, *time.Time) {
|
||||
t.Helper()
|
||||
h, st, now := newClarifyHandler(t)
|
||||
h.router = buildRouter(router.NewHashEmbedder(1024), h.matcher, 0.55, nil)
|
||||
h.recall = recallWiring{embedder: router.NewHashEmbedder(1024), memStore: memory.NewInMemoryStore()}
|
||||
return h, st, now
|
||||
}
|
||||
|
||||
// runTrace drives one trace through handleText and checks every turn, then the
|
||||
// end state. Every failure carries the decision trace so far, so a wrong
|
||||
// claimant reads differently from wrong copy.
|
||||
func runTrace(t *testing.T, tr trace) {
|
||||
t.Helper()
|
||||
// MAVEN_DIALOGUE_NO_SKIP=1 runs the rows that fail today. That is how
|
||||
// whoever lands V-560, V-561 or V-562 sees their row go green before
|
||||
// deleting its skip, and it is also the check that a skip is still earned:
|
||||
// a row that passes with the skip in place is a fix nobody noticed.
|
||||
if tr.skip != "" && os.Getenv("MAVEN_DIALOGUE_NO_SKIP") == "" {
|
||||
t.Skip(tr.skip)
|
||||
}
|
||||
ctx := context.Background()
|
||||
h, st, now := newDialogueHandler(t)
|
||||
const conversation = "web"
|
||||
id := dialogueIDFor(sourceText, conversation)
|
||||
|
||||
var claims []claim
|
||||
fail := func(turnIdx int, format string, args ...any) {
|
||||
t.Helper()
|
||||
lines := make([]string, 0, len(claims))
|
||||
for _, c := range claims {
|
||||
lines = append(lines, " "+c.String())
|
||||
}
|
||||
t.Fatalf("turn %d: "+format+"\n who claimed what:\n%s",
|
||||
append([]any{turnIdx}, append(args, strings.Join(lines, "\n"))...)...)
|
||||
}
|
||||
|
||||
for i, tn := range tr.turns {
|
||||
if tn.wait > 0 {
|
||||
*now = now.Add(tn.wait)
|
||||
}
|
||||
var logged bytes.Buffer
|
||||
prev := log.Writer()
|
||||
log.SetOutput(&logged)
|
||||
reply := h.handleText(ctx, conversation, tn.say)
|
||||
log.SetOutput(prev)
|
||||
claims = append(claims, claimsOf(tn.say, logged.String()))
|
||||
|
||||
body := reply
|
||||
if tn.expired {
|
||||
if !isClarifyExpired(reply) {
|
||||
fail(i, "reply %q must open with the expiry notice", reply)
|
||||
}
|
||||
body = trimClarifyExpired(reply)
|
||||
// The notice is glued in front of this turn's reply, and both halves
|
||||
// have to survive: the words he just said are routed fresh, and
|
||||
// answering only "I let the old one go" drops them.
|
||||
if body == "" {
|
||||
fail(i, "the notice was the whole reply; the fresh words were never answered")
|
||||
}
|
||||
} else if isClarifyExpired(reply) {
|
||||
fail(i, "reply %q announced an expiry nothing asked for", reply)
|
||||
}
|
||||
if tn.question != "" {
|
||||
want, ok := clarifyQuestionFor(tn.question, tn.attempt)
|
||||
if !ok {
|
||||
fail(i, "no question exists for slot %s attempt %d", tn.question, tn.attempt)
|
||||
}
|
||||
if body != want {
|
||||
fail(i, "reply %q, want the %s question worded for attempt %d, %q", body, tn.question, tn.attempt, want)
|
||||
}
|
||||
}
|
||||
if tn.noQuestion && isAnyClarifyQuestion(body) {
|
||||
fail(i, "reply %q is another question; this turn is not something to ask about", body)
|
||||
}
|
||||
for _, want := range tn.contains {
|
||||
if !strings.Contains(body, want) {
|
||||
fail(i, "reply %q does not carry %q", body, want)
|
||||
}
|
||||
}
|
||||
for _, unwanted := range tn.notContain {
|
||||
if strings.Contains(body, unwanted) {
|
||||
fail(i, "reply %q carries %q and must not", body, unwanted)
|
||||
}
|
||||
}
|
||||
checkParked(t, fail, i, h.clarifyStore.Get(id, h.now()), tn.parked)
|
||||
}
|
||||
checkEnd(t, ctx, st, h, tr.end, claims)
|
||||
}
|
||||
|
||||
// isAnyClarifyQuestion — is this reply one of her clarify questions, at any
|
||||
// attempt wording? Reads the templates rather than a list of its own.
|
||||
func isAnyClarifyQuestion(reply string) bool {
|
||||
for _, variants := range clarifyQuestionVariants {
|
||||
for _, v := range variants {
|
||||
if reply == v {
|
||||
return true
|
||||
}
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
func checkParked(t *testing.T, fail func(int, string, ...any), i int, got *dialogue.PendingQuestion, want *parkedWant) {
|
||||
t.Helper()
|
||||
if want == nil {
|
||||
if got != nil {
|
||||
fail(i, "a question about %v is still armed and nothing should be: %+v", got.Missing, got.Slots)
|
||||
}
|
||||
return
|
||||
}
|
||||
if got == nil {
|
||||
fail(i, "nothing is armed, want a question about %s (attempt %d)", want.slot, want.attempt)
|
||||
return
|
||||
}
|
||||
if len(got.Missing) != 1 || got.Missing[0] != want.slot {
|
||||
fail(i, "armed question is about %v, want %s", got.Missing, want.slot)
|
||||
}
|
||||
if got.Attempts != want.attempt {
|
||||
fail(i, "armed question is on attempt %d, want %d — a retry spent on something that was never an answer is the V-554 shape", got.Attempts, want.attempt)
|
||||
}
|
||||
if want.carries != "" && !strings.Contains(got.Utterance, want.carries) {
|
||||
fail(i, "the parked request no longer carries %q: %q", want.carries, got.Utterance)
|
||||
}
|
||||
}
|
||||
|
||||
func checkEnd(t *testing.T, ctx context.Context, st *store.Store, h *reactiveHandler, want endState, claims []claim) {
|
||||
t.Helper()
|
||||
lines := make([]string, 0, len(claims))
|
||||
for _, c := range claims {
|
||||
lines = append(lines, " "+c.String())
|
||||
}
|
||||
trace := "\n who claimed what:\n" + strings.Join(lines, "\n")
|
||||
|
||||
reminders, err := st.DueReminders(ctx, h.now().Add(14*24*time.Hour))
|
||||
if err != nil {
|
||||
t.Fatalf("DueReminders: %v", err)
|
||||
}
|
||||
if len(reminders) != len(want.reminders) {
|
||||
t.Fatalf("end state: %d reminder(s), want %d: %+v%s", len(reminders), len(want.reminders), reminders, trace)
|
||||
}
|
||||
for i, w := range want.reminders {
|
||||
if !strings.Contains(reminders[i].Payload, w.payload) {
|
||||
t.Fatalf("end state: reminder %d payload %q does not carry %q%s", i, reminders[i].Payload, w.payload, trace)
|
||||
}
|
||||
if w.fireAt != "" {
|
||||
if got := reminders[i].FireTs.UTC().Format("2006-01-02 15:04"); got != w.fireAt {
|
||||
t.Fatalf("end state: reminder %d fires at %s, want %s%s", i, got, w.fireAt, trace)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
facts, err := st.RecentFacts(ctx, 20)
|
||||
if err != nil {
|
||||
t.Fatalf("RecentFacts: %v", err)
|
||||
}
|
||||
if len(facts) != len(want.factKeys) {
|
||||
t.Fatalf("end state: %d fact(s), want %d: %+v%s", len(facts), len(want.factKeys), facts, trace)
|
||||
}
|
||||
for i, key := range want.factKeys {
|
||||
if facts[i].Key != key {
|
||||
t.Fatalf("end state: fact %d is %q, want %q%s", i, facts[i].Key, key, trace)
|
||||
}
|
||||
}
|
||||
|
||||
notes, err := st.RecentNotes(ctx, 20)
|
||||
if err != nil {
|
||||
t.Fatalf("RecentNotes: %v", err)
|
||||
}
|
||||
if len(notes) != want.notes {
|
||||
t.Fatalf("end state: %d note(s), want %d%s", len(notes), want.notes, trace)
|
||||
}
|
||||
|
||||
tasks, err := st.ListTasks(ctx, store.TaskOpen)
|
||||
if err != nil {
|
||||
t.Fatalf("ListTasks: %v", err)
|
||||
}
|
||||
if len(tasks) != len(want.tasks) {
|
||||
t.Fatalf("end state: %d open task(s), want %d: %+v%s", len(tasks), len(want.tasks), tasks, trace)
|
||||
}
|
||||
for i, text := range want.tasks {
|
||||
if !strings.Contains(tasks[i].Text, text) {
|
||||
t.Fatalf("end state: task %d is %q, want it to carry %q%s", i, tasks[i].Text, text, trace)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestDialogueTraces(t *testing.T) {
|
||||
for _, tr := range dialogueTraces() {
|
||||
tr := tr
|
||||
t.Run(tr.name, func(t *testing.T) { runTrace(t, tr) })
|
||||
}
|
||||
}
|
||||
|
||||
// dialogueTraces — the fixture. Order is the order the shapes were found, not a
|
||||
// dependency: each trace builds its own handler and store.
|
||||
func dialogueTraces() []trace {
|
||||
return []trace{
|
||||
// The plain two-turn shape, and the one every other row is a deviation
|
||||
// from: she asks for the time, he gives it, the reminder lands with the
|
||||
// subject he said in the FIRST turn.
|
||||
{
|
||||
name: "reminder completed over two turns",
|
||||
turns: []turn{
|
||||
{say: "напомни позвонить маме", question: dialogue.SlotTime, attempt: 1,
|
||||
parked: &parkedWant{slot: dialogue.SlotTime, attempt: 1, carries: "маме"}},
|
||||
{say: "в 11:00", contains: []string{"11:00"}, notContain: []string{"?"}},
|
||||
},
|
||||
end: endState{reminders: []reminderWant{{payload: "позвонить маме", fireAt: "2026-07-31 11:00"}}},
|
||||
},
|
||||
// The same shape on the fact path, where the answer carries both halves
|
||||
// of what was missing — the key and the value — in one breath.
|
||||
{
|
||||
name: "fact completed over two turns",
|
||||
turns: []turn{
|
||||
{say: "запиши", question: dialogue.SlotKey, attempt: 1,
|
||||
parked: &parkedWant{slot: dialogue.SlotKey, attempt: 1}},
|
||||
{say: "пил воду", contains: []string{"water"}},
|
||||
},
|
||||
end: endState{factKeys: []string{"water"}},
|
||||
},
|
||||
// An answer past the TTL is a new request, not an answer (V-385). She
|
||||
// says the old one is gone and routes the words fresh. A bare time on
|
||||
// its own carries no request, so the fresh routing lands on the canned
|
||||
// reply — the point of the row is that NOTHING is created: a reminder
|
||||
// here would fire with the subject of a request she had already let go.
|
||||
{
|
||||
name: "answer arrives after the TTL",
|
||||
turns: []turn{
|
||||
{say: "напомни позвонить маме", question: dialogue.SlotTime, attempt: 1,
|
||||
parked: &parkedWant{slot: dialogue.SlotTime, attempt: 1}},
|
||||
{say: "в 11:00", wait: clarifyTTL + time.Second, expired: true},
|
||||
},
|
||||
end: endState{},
|
||||
},
|
||||
// Three questions is the budget, and running out is SPOKEN: a mute
|
||||
// give-up reads as "done" and he would wait for a reminder that was
|
||||
// never set. The wording changes with the attempt (V-457).
|
||||
{
|
||||
name: "three unclear answers then the give-up line",
|
||||
turns: []turn{
|
||||
{say: "напомни позвонить маме", question: dialogue.SlotTime, attempt: 1,
|
||||
parked: &parkedWant{slot: dialogue.SlotTime, attempt: 1}},
|
||||
{say: "ну не знаю", question: dialogue.SlotTime, attempt: 2,
|
||||
parked: &parkedWant{slot: dialogue.SlotTime, attempt: 2}},
|
||||
{say: "ну не знаю", question: dialogue.SlotTime, attempt: 3,
|
||||
parked: &parkedWant{slot: dialogue.SlotTime, attempt: 3}},
|
||||
{say: "ну не знаю", contains: []string{clarifyGaveUp}, noQuestion: true},
|
||||
},
|
||||
end: endState{},
|
||||
},
|
||||
// A correction points at the previous ACTED turn (repair.go): she redoes
|
||||
// it under the intent he names and says so out loud, because a
|
||||
// correction he cannot see is indistinguishable from one that was
|
||||
// dropped. The task she filed first stays filed — repair redoes, it does
|
||||
// not retract, and V-455 decided that deliberately.
|
||||
//
|
||||
// The corrected-to intent has to differ from the one she used, or repair
|
||||
// declines: teaching the classifier the label it already produced is
|
||||
// worse than doing nothing.
|
||||
{
|
||||
name: "correction of the previous turn",
|
||||
turns: []turn{
|
||||
{say: "добавь в задачи купить молоко", contains: []string{"купить молоко"}},
|
||||
{say: "нет, это был вопрос", contains: []string{"поняла, это вопрос"}},
|
||||
},
|
||||
end: endState{tasks: []string{"купить молоко"}},
|
||||
},
|
||||
// He walks away from his own request: a question is parked, the next
|
||||
// utterance is an unrelated request of its own, and nothing follows.
|
||||
// V-554's fix is what makes this row pass — the question steps aside
|
||||
// rather than scoring "добавь в задачи" as the time. The reminder is
|
||||
// dropped in silence and that is the decision: if he meant it he says it
|
||||
// again, and a question left armed eats the turn after next.
|
||||
{
|
||||
name: "abandoned flow: parked, then an unrelated request",
|
||||
turns: []turn{
|
||||
{say: "напомни позвонить маме", question: dialogue.SlotTime, attempt: 1,
|
||||
parked: &parkedWant{slot: dialogue.SlotTime, attempt: 1}},
|
||||
{say: "добавь в задачи купить молоко", contains: []string{"купить молоко"}},
|
||||
{say: "спасибо"},
|
||||
},
|
||||
end: endState{tasks: []string{"купить молоко"}},
|
||||
},
|
||||
|
||||
// ---- rows below carry the CORRECT expectation and fail today ----
|
||||
|
||||
// The owner's target transcript, V-561. He asks for a reminder, she asks
|
||||
// when, he asks something else entirely, and then comes back to her
|
||||
// question. On the box this created a reminder at 00:12 and never
|
||||
// answered Rome; on the offline floor the side question is recognised as
|
||||
// its own request and the flow is dropped instead, so the wrong reminder
|
||||
// is not made and the right one is not either.
|
||||
//
|
||||
// Both are the same defect: there is no suspend and resume. The correct
|
||||
// shape is the middle turn answered on its own and the parked question
|
||||
// still standing, on the same attempt — a side query is not a failed
|
||||
// answer and must not spend a retry.
|
||||
//
|
||||
// Unskipping this needs more than V-561. "на 9" and "на завтра" are not
|
||||
// read by StubDateTimeParser, which is what the offline floor runs, so
|
||||
// the row below it is the same shape in words the floor can parse and is
|
||||
// the one to watch first.
|
||||
{
|
||||
name: "the owner's transcript from V-561",
|
||||
skip: "V-561: a parked question is not suspended for a side query and never resumes",
|
||||
turns: []turn{
|
||||
{say: "напомни позвонить маме", question: dialogue.SlotTime, attempt: 1,
|
||||
parked: &parkedWant{slot: dialogue.SlotTime, attempt: 1}},
|
||||
{say: "какая сейчас погода в Риме?",
|
||||
parked: &parkedWant{slot: dialogue.SlotTime, attempt: 1, carries: "маме"}},
|
||||
{say: "а, да, прости - на 9.",
|
||||
parked: &parkedWant{slot: dialogue.SlotTime, attempt: 1, carries: "маме"}},
|
||||
{say: "на завтра."},
|
||||
},
|
||||
end: endState{reminders: []reminderWant{{payload: "позвонить маме", fireAt: "2026-08-01 09:00"}}},
|
||||
},
|
||||
// The same shape said in words StubDateTimeParser reads, so this row
|
||||
// turns green on V-561 alone. Same three claims: Rome is answered, the
|
||||
// question survives the side query on the same attempt, and the answer
|
||||
// after it completes the reminder he actually asked for.
|
||||
{
|
||||
name: "nested question: a parked question, then one of his own",
|
||||
skip: "V-561: a side query drops the parked question instead of suspending it",
|
||||
turns: []turn{
|
||||
{say: "напомни позвонить маме", question: dialogue.SlotTime, attempt: 1,
|
||||
parked: &parkedWant{slot: dialogue.SlotTime, attempt: 1}},
|
||||
{say: "какая сейчас погода в Риме?",
|
||||
parked: &parkedWant{slot: dialogue.SlotTime, attempt: 1, carries: "маме"}},
|
||||
{say: "в 11:00", contains: []string{"11:00"}},
|
||||
},
|
||||
end: endState{reminders: []reminderWant{{payload: "позвонить маме", fireAt: "2026-07-31 11:00"}}},
|
||||
},
|
||||
// A cancel is one of the five turn roles V-560 names, and today it is
|
||||
// none of them: "неважно" fills no slot and carries no request of its
|
||||
// own, so it reads as a failed answer and spends a retry. Two turns
|
||||
// later she is still asking about a reminder he called off.
|
||||
//
|
||||
// The row asserts what is knowable — nothing armed, nothing written, and
|
||||
// not another question — rather than her wording for it, which is not
|
||||
// written yet and is not this task's to invent.
|
||||
{
|
||||
name: "cancel: a parked question, then never mind",
|
||||
skip: "V-560: a cancel is scored as a failed answer, not as a cancel",
|
||||
turns: []turn{
|
||||
{say: "напомни позвонить маме", question: dialogue.SlotTime, attempt: 1,
|
||||
parked: &parkedWant{slot: dialogue.SlotTime, attempt: 1}},
|
||||
{say: "неважно", noQuestion: true},
|
||||
},
|
||||
end: endState{},
|
||||
},
|
||||
// Order in runTurn is the whole arbitration (V-558), and this is what it
|
||||
// costs: the clarify answer is checked at step 3 and the repair marker at
|
||||
// step 4d, so while a question is parked no correction can be made. She
|
||||
// scores "нет, это была заметка" as a bad time answer and asks again.
|
||||
{
|
||||
name: "correction while a question is parked",
|
||||
skip: "V-560: clarify pre-empts the repair marker, so a correction cannot be spoken mid-flow",
|
||||
turns: []turn{
|
||||
{say: "добавь в задачи купить молоко", contains: []string{"купить молоко"}},
|
||||
{say: "напомни позвонить маме", question: dialogue.SlotTime, attempt: 1,
|
||||
parked: &parkedWant{slot: dialogue.SlotTime, attempt: 1}},
|
||||
{say: "нет, это был вопрос", contains: []string{"поняла, это вопрос"},
|
||||
parked: &parkedWant{slot: dialogue.SlotTime, attempt: 1, carries: "маме"}},
|
||||
},
|
||||
end: endState{tasks: []string{"купить молоко"}},
|
||||
},
|
||||
// A reminder said whole, in one breath, with the hour in it — and she
|
||||
// asks when. ReminderGrammar (stage0.go) builds its slots by hand and
|
||||
// never runs the extractor, so a stage-0 reminder carries no time
|
||||
// whatever the sentence says, and the clarify gate reads the gap as
|
||||
// real. It costs a turn on the commonest reminder shape there is.
|
||||
//
|
||||
// Hermetic despite the date parser: stage 0 calls no parser at all, so
|
||||
// this fails the same way with or without python dateparser installed.
|
||||
{
|
||||
name: "a reminder said whole is not asked about",
|
||||
skip: "V-562: a stage-0 decision never meets the extractor, so its slots are never validated",
|
||||
turns: []turn{
|
||||
{say: "напомни в 11:00 позвонить маме", contains: []string{"11:00"}, noQuestion: true},
|
||||
},
|
||||
end: endState{reminders: []reminderWant{{payload: "позвонить маме", fireAt: "2026-07-31 11:00"}}},
|
||||
},
|
||||
// The same gap on the repair path. A correction redoes the request
|
||||
// through finishClarified, which goes straight to applyAction — it never
|
||||
// passes the clarify gate — so a redo that lands short answers with the
|
||||
// parse error V-557 removed from the routing path: "не поняла, на когда
|
||||
// напомнить." She should ask, exactly as she does for a fresh reminder
|
||||
// with no time.
|
||||
{
|
||||
name: "a correction that lands short asks rather than failing",
|
||||
skip: "V-562: finishClarified skips the clarify gate, so a repaired decision is never checked for gaps",
|
||||
turns: []turn{
|
||||
{say: "добавь в задачи купить молоко", contains: []string{"купить молоко"}},
|
||||
{say: "нет, это было напоминание", contains: []string{"поняла, это напоминание"},
|
||||
parked: &parkedWant{slot: dialogue.SlotTime, attempt: 1}},
|
||||
},
|
||||
end: endState{tasks: []string{"купить молоко"}},
|
||||
},
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,82 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
|
||||
// The list she read at the mic is not the list a browser is looking at
|
||||
// (Vikunja #45 step 3). The clarify store was keyed per reach in #466; the
|
||||
// dialogue session was still one slot for the box, so "второй" typed on the web
|
||||
// closed the second task she had recited out loud.
|
||||
func TestCandidatesDoNotCrossReaches(t *testing.T) {
|
||||
h, st, _ := newClarifyHandler(t)
|
||||
voiceCtx := withDialogueID(context.Background(), dialogueIDFor(sourceVoice, ""))
|
||||
webCtx := withDialogueID(context.Background(), dialogueIDFor(sourceText, "web"))
|
||||
|
||||
ids := seedTasks(t, st, "купить хлеб", "позвонить маме")
|
||||
putCandidates(h, voiceCtx, ids, "купить хлеб", "позвонить маме")
|
||||
|
||||
if reply, handled := h.resolveCandidate(webCtx, "первую сделал", sourceText); handled {
|
||||
t.Fatalf("a web turn picked from the list she read aloud: %q", reply)
|
||||
}
|
||||
live, err := st.ListTasks(context.Background(), "live")
|
||||
if err != nil {
|
||||
t.Fatalf("list tasks: %v", err)
|
||||
}
|
||||
if len(live) != 2 {
|
||||
t.Fatalf("%d tasks live, want 2 — the web turn moved one", len(live))
|
||||
}
|
||||
// The reach that was offered the list still owns it.
|
||||
if _, handled := h.resolveCandidate(voiceCtx, "первую сделал", sourceVoice); !handled {
|
||||
t.Fatal("the mic lost its own list")
|
||||
}
|
||||
}
|
||||
|
||||
// A selection writes a fact, so the fact must name the reach the words arrived
|
||||
// on. It said "tap:voice" for a typed turn.
|
||||
func TestCandidateProvenanceFollowsTheReach(t *testing.T) {
|
||||
h, st, _ := newClarifyHandler(t)
|
||||
ctx := withDialogueID(context.Background(), dialogueIDFor(sourceText, "web"))
|
||||
ids := seedTasks(t, st, "купить хлеб")
|
||||
putCandidates(h, ctx, ids, "купить хлеб")
|
||||
|
||||
if _, handled := h.resolveCandidate(ctx, "первую сделал", sourceText); !handled {
|
||||
t.Fatal("the pick was not acted on")
|
||||
}
|
||||
done, err := st.ListTasks(context.Background(), "done")
|
||||
if err != nil {
|
||||
t.Fatalf("list tasks: %v", err)
|
||||
}
|
||||
if len(done) != 1 {
|
||||
t.Fatalf("%d tasks done, want 1", len(done))
|
||||
}
|
||||
if by := done[0].ResolvedBy; by != string(sourceText) {
|
||||
t.Errorf("resolved_by = %q, want %q", by, sourceText)
|
||||
}
|
||||
}
|
||||
|
||||
// Anaphora is per reach too: an ellipsis typed on the web must not continue the
|
||||
// question he asked at the mic. Both surfaces stay usable at once, which is the
|
||||
// case a single-owner box actually hits — a phone open while he talks.
|
||||
func TestAnaphoraDoesNotCrossReaches(t *testing.T) {
|
||||
h, _, _ := newClarifyHandler(t)
|
||||
voiceCtx := withDialogueID(context.Background(), dialogueIDFor(sourceVoice, ""))
|
||||
webCtx := withDialogueID(context.Background(), dialogueIDFor(sourceText, "web"))
|
||||
now := h.now()
|
||||
|
||||
h.rememberTurn(voiceCtx, nil, router.Decision{
|
||||
Intent: router.IntentQuery, Utterance: "во сколько у меня встреча",
|
||||
}, now)
|
||||
|
||||
if sess := h.dialogueSessions.Get(dialogueIDOf(webCtx), now); sess != nil {
|
||||
t.Fatalf("the web reach inherited the mic's turn: %+v", sess)
|
||||
}
|
||||
sess := h.dialogueSessions.Get(dialogueIDOf(voiceCtx), now)
|
||||
if sess == nil || !strings.Contains(sess.Slots.Text, "встреча") {
|
||||
t.Fatalf("the mic lost its own turn: %+v", sess)
|
||||
}
|
||||
}
|
||||
+71
-4
@@ -263,18 +263,85 @@ func (c *praxisClient) getJSON(ctx context.Context, op, path string, out any) er
|
||||
return nil
|
||||
}
|
||||
|
||||
func (c *praxisClient) ListAttention(ctx context.Context, limit int) ([]map[string]any, error) {
|
||||
var out []map[string]any
|
||||
// praxisAttention — an attention response in either of the two shapes Praxis
|
||||
// may send (Vikunja #540).
|
||||
//
|
||||
// ECOSYSTEM-SPEC §2.6 says the response carries `degraded: [source_ids]` when a
|
||||
// source is failed or stale, and that Maven is required to say so rather than
|
||||
// report all-clear. The deployed Praxis answers with a bare JSON array and no
|
||||
// envelope at all, so both are decoded here: an array is the items, an object is
|
||||
// the spec envelope. This lands the Maven half without waiting on the server,
|
||||
// and the sources read below is what makes the hedge work meanwhile.
|
||||
type praxisAttention struct {
|
||||
Items []map[string]any
|
||||
Degraded []string
|
||||
}
|
||||
|
||||
func (a *praxisAttention) UnmarshalJSON(data []byte) error {
|
||||
trimmed := bytes.TrimSpace(data)
|
||||
if len(trimmed) > 0 && trimmed[0] == '[' {
|
||||
return json.Unmarshal(trimmed, &a.Items)
|
||||
}
|
||||
var env struct {
|
||||
Items []map[string]any `json:"items"`
|
||||
Degraded []string `json:"degraded"`
|
||||
}
|
||||
if err := json.Unmarshal(trimmed, &env); err != nil {
|
||||
return err
|
||||
}
|
||||
a.Items, a.Degraded = env.Items, env.Degraded
|
||||
return nil
|
||||
}
|
||||
|
||||
func (c *praxisClient) ListAttention(ctx context.Context, limit int) (praxisAttention, error) {
|
||||
var out praxisAttention
|
||||
err := c.getJSON(ctx, "attention", fmt.Sprintf("/api/v1/tools/attention?limit=%d", limit), &out)
|
||||
return out, err
|
||||
}
|
||||
|
||||
// praxisSource — one polled source, as much of it as the hedge needs. The tools
|
||||
// API does not expose sources, so this decodes the plain `/api/v1/sources` rows.
|
||||
type praxisSource struct {
|
||||
ID string `json:"id"`
|
||||
SourceID string `json:"source_id"`
|
||||
Health string `json:"health"`
|
||||
}
|
||||
|
||||
func (s praxisSource) name() string {
|
||||
if s.SourceID != "" {
|
||||
return s.SourceID
|
||||
}
|
||||
return s.ID
|
||||
}
|
||||
|
||||
// UnhealthySources reports which sources cannot be trusted to have reported,
|
||||
// and how many sources Praxis has at all (Vikunja #540).
|
||||
//
|
||||
// Only read when the attention list came back empty, which is the one turn where
|
||||
// an all-clear is at stake. A source whose health field is absent counts as
|
||||
// healthy: a Praxis that never reports health would otherwise make every quiet
|
||||
// turn a hedge, and an unreported field is not evidence of a fault. Everything it
|
||||
// does report other than "ok" — failed, stale, degraded, unknown — counts as
|
||||
// cannot-tell, because none of them mean the source has spoken.
|
||||
func (c *praxisClient) UnhealthySources(ctx context.Context) (bad []string, total int, err error) {
|
||||
var out []praxisSource
|
||||
if err := c.getJSON(ctx, "sources", "/api/v1/sources", &out); err != nil {
|
||||
return nil, 0, err
|
||||
}
|
||||
for _, s := range out {
|
||||
if s.Health != "" && s.Health != "ok" {
|
||||
bad = append(bad, s.name())
|
||||
}
|
||||
}
|
||||
return bad, len(out), nil
|
||||
}
|
||||
|
||||
// ListAttentionForEntity is ListAttention scoped to a single canonical Nexus
|
||||
// entity, so callers already holding a resolved entity_id (e.g. after
|
||||
// resolveEntityReference) can ask "what needs attention for this entity"
|
||||
// instead of filtering the unscoped list client-side.
|
||||
func (c *praxisClient) ListAttentionForEntity(ctx context.Context, entityID string, limit int) ([]map[string]any, error) {
|
||||
var out []map[string]any
|
||||
func (c *praxisClient) ListAttentionForEntity(ctx context.Context, entityID string, limit int) (praxisAttention, error) {
|
||||
var out praxisAttention
|
||||
err := c.getJSON(ctx, "attention_for_entity",
|
||||
fmt.Sprintf("/api/v1/tools/attention?limit=%d&entity_id=%s", limit, url.QueryEscape(entityID)), &out)
|
||||
return out, err
|
||||
|
||||
@@ -5,6 +5,7 @@ import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"log"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
@@ -105,6 +106,13 @@ func (h *reactiveHandler) handlePraxisAct(ctx context.Context, dec router.Decisi
|
||||
ctx = withCorrelationID(ctx, newCorrelationID())
|
||||
}
|
||||
px := h.ecosystem.praxis
|
||||
dec, ok := h.resolveSurfacedPosition(dec)
|
||||
if !ok {
|
||||
// A demonstrative with no digest behind it. "я это сделал" is a sentence
|
||||
// about his day, so the rest of the cascade gets it back rather than
|
||||
// hearing "какой пункт?" for something that was never about a пункт.
|
||||
return ""
|
||||
}
|
||||
for _, capability := range praxisCapabilities {
|
||||
for _, alias := range capability.aliases() {
|
||||
if alias == dec.Slots.Fn {
|
||||
@@ -154,18 +162,23 @@ func (listAttentionCapability) aliases() []string {
|
||||
|
||||
func (listAttentionCapability) handle(ctx context.Context, h *reactiveHandler, px *praxisClient, _ router.Decision) string {
|
||||
started := h.now()
|
||||
items, err := px.ListAttention(ctx, 20)
|
||||
att, err := px.ListAttention(ctx, 20)
|
||||
if err != nil {
|
||||
log.Printf("ecosystem: praxis attention: %v", err)
|
||||
h.recordEcosystemTrace(ctx, "praxis", "list_attention", traceStatusForError(err),
|
||||
started, traceErrorFields(err))
|
||||
return phraser.A(phraser.AttentionFail, nil)
|
||||
}
|
||||
items := att.Items
|
||||
if len(items) == 0 {
|
||||
if hedge := h.attentionCannotTell(ctx, px, att.Degraded, started); hedge != "" {
|
||||
return hedge
|
||||
}
|
||||
return phraser.A(phraser.AttentionNone, nil)
|
||||
}
|
||||
h.recordPraxisTrace(ctx, "list_attention", started, map[string]any{"count": len(items)})
|
||||
var parts []string
|
||||
var spoken []string
|
||||
for _, item := range items {
|
||||
title, _ := item["title"].(string)
|
||||
// importance arrives as JSON number ⇒ float64 over the HTTP contract.
|
||||
@@ -190,11 +203,15 @@ func (listAttentionCapability) handle(ctx context.Context, h *reactiveHandler, p
|
||||
// (ECOSYSTEM-SPEC.md §2.3: surfaced != acknowledged). Best-effort:
|
||||
// a failed surface call must not block delivering the digest.
|
||||
if id, ok := item["id"].(string); ok && id != "" {
|
||||
// Recorded in the order she says them, and only for items she could
|
||||
// say: an item skipped above has no position in what he heard (#516).
|
||||
spoken = append(spoken, id)
|
||||
if _, err := px.Surface(ctx, id); err != nil {
|
||||
log.Printf("ecosystem: praxis surface %s: %v", id, err)
|
||||
}
|
||||
}
|
||||
}
|
||||
h.rememberSurfaced(spoken)
|
||||
if len(parts) == 0 {
|
||||
// Praxis returned items and not one of them could be said. "ничего не
|
||||
// требует внимания" is the honest answer; the list line would render as
|
||||
@@ -299,14 +316,14 @@ func (entityAttentionCapability) handle(ctx context.Context, h *reactiveHandler,
|
||||
}
|
||||
|
||||
queried := h.now()
|
||||
items, err := px.ListAttentionForEntity(ctx, entityID, 20)
|
||||
att, err := px.ListAttentionForEntity(ctx, entityID, 20)
|
||||
if err != nil {
|
||||
log.Printf("ecosystem: praxis attention for %s: %v", entityID, err)
|
||||
h.recordEcosystemTrace(ctx, "praxis", "entity_attention", traceStatusForError(err),
|
||||
queried, mergeFields(traceErrorFields(err), map[string]any{"entity_id": entityID}))
|
||||
return phraser.A(phraser.AttentionFailEntity, map[string]string{"name": displayName})
|
||||
}
|
||||
items, scoped := scopedToEntity(items, entityID)
|
||||
items, scoped := scopedToEntity(att.Items, entityID)
|
||||
if !scoped {
|
||||
// A Praxis old enough to ignore an unknown query parameter answers the
|
||||
// scoped question with the unscoped list. Reading that back as "по
|
||||
@@ -339,6 +356,11 @@ func (entityAttentionCapability) handle(ctx context.Context, h *reactiveHandler,
|
||||
parts = append(parts, known)
|
||||
}
|
||||
if len(parts) == 0 {
|
||||
// The scoped list is as exposed to a silent source as the unscoped one,
|
||||
// and a per-entity all-clear is the more convincing of the two (#540).
|
||||
if hedge := h.attentionCannotTell(ctx, px, att.Degraded, queried); hedge != "" {
|
||||
return hedge
|
||||
}
|
||||
return phraser.A(phraser.AttentionNoneEntity, map[string]string{"name": displayName})
|
||||
}
|
||||
return phraser.A(phraser.AttentionListEntity, map[string]string{"name": displayName, "items": strings.Join(parts, "; ")})
|
||||
@@ -783,3 +805,105 @@ func (h *reactiveHandler) hexisBeforeClarify(ctx context.Context, dec router.Dec
|
||||
}
|
||||
return h.handleHexisAct(ctx, dec)
|
||||
}
|
||||
|
||||
// attentionCannotTell returns the hedge to say instead of an all-clear, or ""
|
||||
// when an empty attention list really does mean nothing needs looking at
|
||||
// (ECOSYSTEM-SPEC §2.6, Vikunja #540).
|
||||
//
|
||||
// "Nothing needs attention" and "I cannot currently tell" are different answers
|
||||
// and only one of them was ever said. The spec's mechanism is a `degraded` array
|
||||
// on the attention response, which the deployed Praxis does not send, so the
|
||||
// source health read is the half that works today. It costs one HTTP call and
|
||||
// only on the empty-list turn, which is the only turn where an all-clear is at
|
||||
// stake.
|
||||
//
|
||||
// A failed sources read is deliberately NOT a hedge. The attention call itself
|
||||
// succeeded, and not being able to ask about health is not evidence of a fault —
|
||||
// hedging on it would turn one flaky endpoint into a permanently uncertain
|
||||
// assistant.
|
||||
func (h *reactiveHandler) attentionCannotTell(ctx context.Context, px *praxisClient, degraded []string, started time.Time) string {
|
||||
if len(degraded) > 0 {
|
||||
h.recordPraxisTrace(ctx, "attention_degraded", started, map[string]any{
|
||||
"degraded": strings.Join(degraded, ","), "source": "response",
|
||||
})
|
||||
return phraser.A(phraser.AttentionDegraded, map[string]string{"items": strings.Join(degraded, ", ")})
|
||||
}
|
||||
bad, total, err := px.UnhealthySources(ctx)
|
||||
if err != nil {
|
||||
log.Printf("ecosystem: praxis sources: %v", err)
|
||||
return ""
|
||||
}
|
||||
if total == 0 {
|
||||
// A Praxis that polls nothing knows nothing, so its silence is not an
|
||||
// all-clear either. This is the state the box is in as of 2026-08-05:
|
||||
// /api/v1/sources answers with an empty array.
|
||||
h.recordPraxisTrace(ctx, "attention_no_sources", started, map[string]any{"sources": 0})
|
||||
return phraser.A(phraser.AttentionNoSources, nil)
|
||||
}
|
||||
if len(bad) > 0 {
|
||||
h.recordPraxisTrace(ctx, "attention_degraded", started, map[string]any{
|
||||
"degraded": strings.Join(bad, ","), "sources": total, "source": "health",
|
||||
})
|
||||
return phraser.A(phraser.AttentionDegraded, map[string]string{"items": strings.Join(bad, ", ")})
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// rememberSurfaced records the item ids she just read out, replacing whatever the
|
||||
// previous digest left. Called with the ids in speaking order (Vikunja #516).
|
||||
func (h *reactiveHandler) rememberSurfaced(ids []string) {
|
||||
h.mu.Lock()
|
||||
defer h.mu.Unlock()
|
||||
h.surfacedItems = ids
|
||||
}
|
||||
|
||||
// resolveSurfacedPosition turns a positional item reference into a Praxis item
|
||||
// id, using the list she last read out.
|
||||
//
|
||||
// The router names a position and not an id, because only the daemon has the
|
||||
// list: PraxisGrammars fills the value slot with "2", "last" or "this". An id is
|
||||
// left alone, since "item_ab12" is already one.
|
||||
//
|
||||
// The second return says whether the turn is still Praxis's. A position that
|
||||
// names nothing keeps the turn and clears the slot, so the capability answers its
|
||||
// own "какой пункт?" — he said "второй пункт" and deserves to hear that there is
|
||||
// no second one. A demonstrative that resolves to nothing gives the turn BACK,
|
||||
// because "я это сделал" was probably never about a пункт at all. "это" also
|
||||
// needs the list to hold exactly one item: pointing at one of five is a guess,
|
||||
// and a wrong guess here transitions the wrong item.
|
||||
func (h *reactiveHandler) resolveSurfacedPosition(dec router.Decision) (router.Decision, bool) {
|
||||
ref := dec.Slots.Value
|
||||
if ref == "" || strings.HasPrefix(ref, "item") {
|
||||
return dec, true
|
||||
}
|
||||
h.mu.Lock()
|
||||
ids := h.surfacedItems
|
||||
h.mu.Unlock()
|
||||
|
||||
idx := -1
|
||||
switch {
|
||||
case ref == "this":
|
||||
if len(ids) != 1 {
|
||||
log.Printf("ecosystem: praxis \"это\" has no single item (%d surfaced)", len(ids))
|
||||
return dec, false
|
||||
}
|
||||
idx = 0
|
||||
case ref == "last":
|
||||
idx = len(ids) - 1
|
||||
default:
|
||||
n, err := strconv.Atoi(ref)
|
||||
if err != nil || n < 1 {
|
||||
// Neither a position nor an id: leave it for the capability to
|
||||
// reject rather than silently rewriting what he said.
|
||||
return dec, true
|
||||
}
|
||||
idx = n - 1
|
||||
}
|
||||
if idx < 0 || idx >= len(ids) {
|
||||
log.Printf("ecosystem: praxis position %q has no item (%d surfaced)", ref, len(ids))
|
||||
dec.Slots.Value = ""
|
||||
return dec, true
|
||||
}
|
||||
dec.Slots.Value = ids[idx]
|
||||
return dec, true
|
||||
}
|
||||
|
||||
@@ -167,3 +167,69 @@ func TestActionFact_AskedToRememberAComplaintStillWrites(t *testing.T) {
|
||||
t.Fatalf("an explicit capture was refused: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// A second tap of the same key supersedes the first, so recall must hold one
|
||||
// vector and it must be the new value (Vikunja #493). Before this the id
|
||||
// carried a timestamp, both rows stayed, and the superseded value went on
|
||||
// competing for the turn.
|
||||
func TestActionFact_ARetapSupersedesTheOldVector(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, _ := newFactGateHandler(t, time.Now())
|
||||
|
||||
h.actionFact(ctx, router.Decision{
|
||||
Intent: router.IntentFact,
|
||||
Utterance: "запиши что я пил воду",
|
||||
Slots: router.Slots{Key: "water", HasKey: true, Value: `"250мл"`},
|
||||
})
|
||||
// A later tap of the same key. The clock moves, so the old id and the new
|
||||
// one differ — which is exactly what used to leave two rows behind.
|
||||
h.now = func() time.Time { return time.Now().Add(time.Hour) }
|
||||
h.actionFact(ctx, router.Decision{
|
||||
Intent: router.IntentFact,
|
||||
Utterance: "запиши что я пил воду",
|
||||
Slots: router.Slots{Key: "water", HasKey: true, Value: `"500мл"`},
|
||||
})
|
||||
|
||||
hits, err := h.recall.memStore.Search(ctx, mustEmbedPassage(t, h, "вода"), 5)
|
||||
if err != nil {
|
||||
t.Fatalf("memory search: %v", err)
|
||||
}
|
||||
if len(hits) != 1 {
|
||||
t.Fatalf("want one vector for the key, got %d: %+v", len(hits), hits)
|
||||
}
|
||||
if got := hits[0].Meta["text"]; got != "water — 500мл" {
|
||||
t.Errorf("indexed text = %q, want the current value", got)
|
||||
}
|
||||
}
|
||||
|
||||
// Another key is not this key. A prefix delete that widened would take the
|
||||
// whole index with it.
|
||||
func TestActionFact_ARetapLeavesOtherKeysAlone(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, _ := newFactGateHandler(t, time.Now())
|
||||
|
||||
h.actionFact(ctx, router.Decision{
|
||||
Intent: router.IntentFact,
|
||||
Utterance: "запиши что я обедал",
|
||||
Slots: router.Slots{Key: "meal", HasKey: true, Value: `"суп"`},
|
||||
})
|
||||
h.actionFact(ctx, router.Decision{
|
||||
Intent: router.IntentFact,
|
||||
Utterance: "запиши что я пил воду",
|
||||
Slots: router.Slots{Key: "water", HasKey: true, Value: `"250мл"`},
|
||||
})
|
||||
|
||||
hits, err := h.recall.memStore.Search(ctx, mustEmbedPassage(t, h, "обед"), 5)
|
||||
if err != nil {
|
||||
t.Fatalf("memory search: %v", err)
|
||||
}
|
||||
var found bool
|
||||
for _, hit := range hits {
|
||||
if hit.Meta["text"] == "meal — суп" {
|
||||
found = true
|
||||
}
|
||||
}
|
||||
if !found {
|
||||
t.Fatalf("writing water dropped the meal vector: %+v", hits)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -178,6 +178,14 @@ func (fs *fakeServer) Requests() []capturedRequest {
|
||||
return out
|
||||
}
|
||||
|
||||
// ResetRequests drops the captured requests, so a test can assert about one
|
||||
// turn without subtracting the setup turn's calls.
|
||||
func (fs *fakeServer) ResetRequests() {
|
||||
fs.mu.Lock()
|
||||
defer fs.mu.Unlock()
|
||||
fs.requests = nil
|
||||
}
|
||||
|
||||
func jsonHandler(status int, body string) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
w.Header().Set("Content-Type", "application/json")
|
||||
@@ -330,7 +338,17 @@ func newFakeNexus(t *testing.T, resolveBody string) *fakeServer {
|
||||
// Maven's praxisClient calls. Every route returns its fixed body until a
|
||||
// fault is injected via SetFault.
|
||||
func newFakePraxis(t *testing.T, attentionBody string) *fakeServer {
|
||||
// One healthy source by default: an empty attention list only means
|
||||
// all-clear when something is actually polling (Vikunja #540), and the
|
||||
// other tests here are about attention rather than about source health.
|
||||
return newFakePraxisWithSources(t, attentionBody, `[{"source_id":"src_ntfy","health":"ok"}]`)
|
||||
}
|
||||
|
||||
// newFakePraxisWithSources is newFakePraxis with the /api/v1/sources body
|
||||
// under the test's control, for the degraded and no-sources hedges.
|
||||
func newFakePraxisWithSources(t *testing.T, attentionBody, sourcesBody string) *fakeServer {
|
||||
return newFakeServer(t, map[string]http.HandlerFunc{
|
||||
"GET /api/v1/sources": jsonHandler(http.StatusOK, sourcesBody),
|
||||
"GET /api/v1/tools/attention": jsonHandler(http.StatusOK, attentionBody),
|
||||
"GET /api/v1/tools/changes": jsonHandler(http.StatusOK, `[]`),
|
||||
"POST /api/v1/tools/surface": jsonHandler(http.StatusOK, `{}`),
|
||||
|
||||
+20
-5
@@ -8,10 +8,10 @@ import (
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
|
||||
// voiceDialogueID — the dialogue-session key for the microphone, and the
|
||||
// clarify key for it too. This is a single-user box (ponytail), so one slot
|
||||
// suffices; a second speaker would need per-speaker ids, which waits on
|
||||
// voice-print attribution (see PROGRESS multi-user deferral).
|
||||
// voiceDialogueID — the dialogue-session and clarify key for the microphone.
|
||||
// This is a single-user box (ponytail), so one slot per reach suffices; a
|
||||
// second speaker would need per-speaker ids, which waits on voice-print
|
||||
// attribution (see PROGRESS multi-user deferral).
|
||||
const voiceDialogueID = "voice"
|
||||
|
||||
// textDialogueID — the clarify key for a text turn that named no conversation.
|
||||
@@ -97,7 +97,8 @@ var anaphoraResolver router.AnaphoraResolver
|
||||
// followUpMerge fills the current turn's missing slots from a prior
|
||||
// non-expired session — the multi-turn seam. It handles three cases:
|
||||
//
|
||||
// 1. Same-intent: inherit missing slots via InheritSlots (existing behavior).
|
||||
// 1. Same-intent: inherit missing slots via InheritSlots (existing behavior),
|
||||
// except a reminder time the current sentence named and the parser missed.
|
||||
// 2. Cross-intent anaphora: if the current utterance contains a pronoun
|
||||
// ("это" / "он" / "она" etc.) AND the prior session has a key, inherit
|
||||
// the key for fact-lookup queries and reminder creation.
|
||||
@@ -113,8 +114,22 @@ func followUpMerge(prev *dialogue.Session, dec router.Decision, now time.Time) r
|
||||
|
||||
// Case 1: same-intent inheritance (existing).
|
||||
if prev.Intent == dialogue.Intent(dec.Intent) {
|
||||
// A reminder that named an hour nobody could read must not borrow the
|
||||
// last one's. Two reminders in a row and the second landed at the
|
||||
// first's time, confirmed as if it had been read from the sentence:
|
||||
// "напомни без четверти восемь выходить" fired at 07:30 (V-543). The
|
||||
// hour is also what fills before the action's own fallback parse can
|
||||
// run, so inheriting it hid a time that did parse.
|
||||
//
|
||||
// Inheriting is still right when the sentence names no time at all,
|
||||
// which is the follow-up this seam exists for.
|
||||
blockTime := dec.Intent == router.IntentReminder &&
|
||||
!dec.Slots.HasTime && router.MentionsTime(dec.Utterance)
|
||||
merged := dialogue.InheritSlots(prev.Slots, toDialogueSlots(dec.Slots))
|
||||
dec.Slots = applyDialogueSlots(dec.Slots, merged)
|
||||
if blockTime {
|
||||
dec.Slots.Time, dec.Slots.HasTime = time.Time{}, false
|
||||
}
|
||||
return dec
|
||||
}
|
||||
|
||||
|
||||
@@ -35,6 +35,42 @@ func TestFollowUpMerge(t *testing.T) {
|
||||
}
|
||||
})
|
||||
|
||||
// V-543, measured on the box: four reminders in a row all landed at the
|
||||
// first one's hour, each confirmed as if it had been read from the sentence.
|
||||
// A sentence that names a time and fails to parse must ask, not borrow.
|
||||
t.Run("a named time that did not parse is not inherited", func(t *testing.T) {
|
||||
for _, utt := range []string{
|
||||
"напомни без четверти восемь выходить",
|
||||
"напомни в половине первого пообедать",
|
||||
"напомни завтра принять лекарство",
|
||||
"remind me at noon to stretch",
|
||||
} {
|
||||
cur := router.Decision{
|
||||
Intent: router.IntentReminder,
|
||||
Utterance: utt,
|
||||
Slots: router.Slots{Text: utt},
|
||||
}
|
||||
got := followUpMerge(prev, cur, base.Add(30*time.Second))
|
||||
if got.Slots.HasTime {
|
||||
t.Errorf("%q borrowed the previous hour %v", utt, got.Slots.Time)
|
||||
}
|
||||
}
|
||||
})
|
||||
|
||||
// The follow-up this seam exists for still works: the sentence names no
|
||||
// time, so the previous one is the only one it could mean.
|
||||
t.Run("a follow-up naming no time still inherits", func(t *testing.T) {
|
||||
cur := router.Decision{
|
||||
Intent: router.IntentReminder,
|
||||
Utterance: "и ещё полить цветы",
|
||||
Slots: router.Slots{Text: "полить цветы"},
|
||||
}
|
||||
got := followUpMerge(prev, cur, base.Add(30*time.Second))
|
||||
if !got.Slots.HasTime || !got.Slots.Time.Equal(fireAt) {
|
||||
t.Errorf("time not inherited: HasTime=%v Time=%v", got.Slots.HasTime, got.Slots.Time)
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("current slot wins over prior (gaps only)", func(t *testing.T) {
|
||||
own := base.Add(48 * time.Hour)
|
||||
cur := router.Decision{
|
||||
|
||||
+40
-7
@@ -54,30 +54,56 @@ var historyMarkersEn = [][2]string{
|
||||
// answers a topic far better than a list of the last five facts does.
|
||||
var historyRecall = []string{" про ", " об ", " о ", " about "}
|
||||
|
||||
// historySide — whose turn the question asks about. The rows read are the same
|
||||
// either way, because a tapped fact is one act seen from two sides, but the
|
||||
// sentence is not: answering "что ты записала сегодня?" with "ты говорил…"
|
||||
// hands the question back instead of answering it (Vikunja #456).
|
||||
type historySide int
|
||||
|
||||
const (
|
||||
historyAskedHim historySide = iota // "что я тебе говорил"
|
||||
historyAskedHer // "что ты записала сегодня"
|
||||
)
|
||||
|
||||
// isHistoryQuery reports whether he is asking what he told her.
|
||||
func isHistoryQuery(u string) bool {
|
||||
_, ok := historyAsks(u)
|
||||
return ok
|
||||
}
|
||||
|
||||
// historyAsks reports whether this is a history question, and whose turn it is
|
||||
// about.
|
||||
func historyAsks(u string) (historySide, bool) {
|
||||
s := " " + strings.ToLower(strings.TrimSpace(u)) + " "
|
||||
if s == " " {
|
||||
return false
|
||||
return historyAskedHim, false
|
||||
}
|
||||
for _, r := range historyRecall {
|
||||
if strings.Contains(s, r) {
|
||||
return false
|
||||
return historyAskedHim, false
|
||||
}
|
||||
}
|
||||
for _, pair := range historyMarkersEn {
|
||||
if strings.Contains(s, pair[0]) && strings.Contains(s, pair[1]) {
|
||||
return true
|
||||
if strings.Contains(pair[0], "you") {
|
||||
return historyAskedHer, true
|
||||
}
|
||||
return historyAskedHim, true
|
||||
}
|
||||
}
|
||||
toks := historyTokens(s)
|
||||
if !hasAny(toks, "что", "чего") {
|
||||
return false
|
||||
return historyAskedHim, false
|
||||
}
|
||||
// His side is tested first: "отмечать" is on both verb lists, so "что я
|
||||
// отметил" must not read as a question about her.
|
||||
if hasAny(toks, firstPersonSubjects...) && hasVerbForm(toks, historySpokenVerbs) {
|
||||
return true
|
||||
return historyAskedHim, true
|
||||
}
|
||||
return hasAny(toks, secondPersonSubjects...) && hasVerbForm(toks, historyRecordedVerbs)
|
||||
if hasAny(toks, secondPersonSubjects...) && hasVerbForm(toks, historyRecordedVerbs) {
|
||||
return historyAskedHer, true
|
||||
}
|
||||
return historyAskedHim, false
|
||||
}
|
||||
|
||||
// historyTokens splits an utterance into bare words. The punctuation goes
|
||||
@@ -142,7 +168,8 @@ const historyWindow = 24 * time.Hour
|
||||
// pass: the notes pass would otherwise answer this from whatever note happens
|
||||
// to be nearest, which reads as an answer and is not one.
|
||||
func (h *reactiveHandler) queryHistory(ctx context.Context, t *queryTurn) (string, bool) {
|
||||
if !isHistoryQuery(t.dec.Utterance) {
|
||||
side, ok := historyAsks(t.dec.Utterance)
|
||||
if !ok {
|
||||
return "", false
|
||||
}
|
||||
facts, err := h.api.RecentFacts(ctx, historyScan)
|
||||
@@ -164,8 +191,14 @@ func (h *reactiveHandler) queryHistory(ctx context.Context, t *queryTurn) (strin
|
||||
if len(said) == 0 {
|
||||
// Claim the turn rather than fall through. "ничего не говорил" is the
|
||||
// true answer, and recall would answer it with an old note instead.
|
||||
if side == historyAskedHer {
|
||||
return "за последние сутки я ничего с твоих слов не записывала.", true
|
||||
}
|
||||
return "за последние сутки ты мне ничего такого не говорил.", true
|
||||
}
|
||||
if side == historyAskedHer {
|
||||
return "я записала: " + strings.Join(said, "; "), true
|
||||
}
|
||||
return "ты говорил: " + strings.Join(said, "; "), true
|
||||
}
|
||||
|
||||
|
||||
@@ -89,6 +89,33 @@ func TestHistoryReadsOnlyWhatHeSaid(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// The rows are the same either way, because a tapped fact is one act seen from
|
||||
// two sides. The sentence is not: "что ты записала" answered with "ты говорил"
|
||||
// hands the question back (Vikunja #456).
|
||||
func TestHistoryAnswersTheSideItWasAsked(t *testing.T) {
|
||||
now := time.Date(2026, 8, 4, 20, 0, 0, 0, time.UTC)
|
||||
h, _ := historyHandler(now, ipc.Fact{Key: "water", Value: "выпил", Source: "tap:voice", Ts: now.Add(-time.Hour)})
|
||||
|
||||
his, ok := askHistory(h, "что я тебе говорил?")
|
||||
if !ok || !strings.HasPrefix(his, "ты говорил") {
|
||||
t.Errorf("reply = %q, ok = %v, want his side", his, ok)
|
||||
}
|
||||
hers, ok := askHistory(h, "что ты записала сегодня?")
|
||||
if !ok || !strings.HasPrefix(hers, "я записала") {
|
||||
t.Errorf("reply = %q, ok = %v, want her side", hers, ok)
|
||||
}
|
||||
// "отмечать" is on both verb lists, so his subject has to win.
|
||||
if side, ok := historyAsks("что я отметил?"); !ok || side != historyAskedHim {
|
||||
t.Errorf("historyAsks(что я отметил) = %v, %v", side, ok)
|
||||
}
|
||||
|
||||
empty, _ := historyHandler(now)
|
||||
none, ok := askHistory(empty, "что ты записала сегодня?")
|
||||
if !ok || !strings.Contains(none, "не записывала") {
|
||||
t.Errorf("empty reply = %q, ok = %v, want her side", none, ok)
|
||||
}
|
||||
}
|
||||
|
||||
// Nothing said is an answer of its own. Falling through would hand the question
|
||||
// to recall, which answers it with an old note.
|
||||
func TestHistorySaysWhenThereIsNothing(t *testing.T) {
|
||||
|
||||
@@ -0,0 +1,44 @@
|
||||
package main
|
||||
|
||||
import "testing"
|
||||
|
||||
func TestHasCyrillic(t *testing.T) {
|
||||
for _, s := range []string{"что такое фотосинтез", "кто такой Elon Musk", "фотосинтез"} {
|
||||
if !hasCyrillic(s) {
|
||||
t.Errorf("hasCyrillic(%q) = false; it is a Russian question", s)
|
||||
}
|
||||
}
|
||||
for _, s := range []string{"what is photosynthesis", "", "3:2"} {
|
||||
if hasCyrillic(s) {
|
||||
t.Errorf("hasCyrillic(%q) = true; there is no Cyrillic in it", s)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The book choice and the rewrite decision are the same decision: a Russian
|
||||
// book reads his question as he asked it, an English one needs it translated
|
||||
// into keywords first (V-508).
|
||||
func TestKiwixBookChoice(t *testing.T) {
|
||||
for _, tc := range []struct {
|
||||
name string
|
||||
wiring kiwixWiring
|
||||
utterance string
|
||||
wantBook string
|
||||
wantVerb bool
|
||||
}{
|
||||
{"a russian question reads the russian book verbatim",
|
||||
kiwixWiring{book: "en", bookRU: "ru"}, "что такое фотосинтез", "ru", true},
|
||||
{"an english question reads the english book",
|
||||
kiwixWiring{book: "en", bookRU: "ru"}, "what is photosynthesis", "en", false},
|
||||
{"no russian book configured leaves every question on the english one",
|
||||
kiwixWiring{book: "en"}, "что такое фотосинтез", "en", false},
|
||||
} {
|
||||
book, verbatim := tc.wiring.book, false
|
||||
if tc.wiring.bookRU != "" && hasCyrillic(tc.utterance) {
|
||||
book, verbatim = tc.wiring.bookRU, true
|
||||
}
|
||||
if book != tc.wantBook || verbatim != tc.wantVerb {
|
||||
t.Errorf("%s: book=%q verbatim=%v, want %q/%v", tc.name, book, verbatim, tc.wantBook, tc.wantVerb)
|
||||
}
|
||||
}
|
||||
}
|
||||
+10
-2
@@ -19,8 +19,12 @@ type kiwixWiring struct {
|
||||
client *kiwix.Client
|
||||
rewriter *kiwix.Rewriter // nil ⇒ the question is searched verbatim
|
||||
book string
|
||||
max int
|
||||
runes int
|
||||
// bookRU — searched instead of book when the question is Cyrillic, and
|
||||
// searched verbatim because it is in his language already (V-508). Empty ⇒
|
||||
// every question goes to book.
|
||||
bookRU string
|
||||
max int
|
||||
runes int
|
||||
}
|
||||
|
||||
// wireKiwix builds the ZIM reader from the `kiwix` block, or returns nil when
|
||||
@@ -38,9 +42,13 @@ func wireKiwix(cfg *config.Config, c *llm.Client) *kiwixWiring {
|
||||
w := &kiwixWiring{
|
||||
client: kiwix.New(kc.URL),
|
||||
book: kc.Book,
|
||||
bookRU: kc.BookRU,
|
||||
max: kc.MaxResults,
|
||||
runes: kc.SnippetRunes,
|
||||
}
|
||||
if kc.BookRU != "" {
|
||||
log.Printf("voice: kiwix: russian questions read %q verbatim", kc.BookRU)
|
||||
}
|
||||
switch {
|
||||
case !kc.RewriteEnabled():
|
||||
log.Printf("voice: kiwix at %s (book %q, query rewriting off by config)", kc.URL, kc.Book)
|
||||
|
||||
@@ -127,8 +127,12 @@ func run(args []string) error {
|
||||
cfgPath := flag.String("config", defaultConfigPath(), "path to mavend JSON config")
|
||||
wrappedKeyPath := flag.String("wrapped-key-file", "", "path to wrapped encryption key blob (enables cold-start unlock)")
|
||||
reembed := flag.Bool("reembed", false, "re-embed every stored note and fact with the configured embedder, then serve normally (run once after an embedder swap; the daemon does not answer until it finishes)")
|
||||
allowSeed := flag.Bool("allow-seed", false, "enable the backdated seed_event write path (QA only: it lets a caller place a fact in the past and mint a routine the tick loop will then act on; off means the method has nothing to write with)")
|
||||
wipe := flag.Bool("wipe", false, "print every table and its row count, then exit without serving; add -confirm-wipe to delete all of it")
|
||||
confirmWipe := flag.Bool("confirm-wipe", false, "with -wipe, actually remove every piece of personal data (facts, notes, vectors, events, tasks, sessions, traces, voiceprints). config, models, passkeys and the encryption key are files and survive")
|
||||
flag.CommandLine.Parse(args)
|
||||
reembedOnStart = *reembed
|
||||
allowSeedOnStart = *allowSeed
|
||||
cfg, err := config.Load(*cfgPath)
|
||||
if err != nil {
|
||||
return err
|
||||
@@ -208,6 +212,14 @@ func run(args []string) error {
|
||||
}()
|
||||
}
|
||||
|
||||
// ----- wipe: never serves, exits when it is done (Vikunja #494) -----
|
||||
if *wipe {
|
||||
if locked {
|
||||
return fmt.Errorf("wipe: the store is locked and there is no key to open it with")
|
||||
}
|
||||
return runWipe(ctx, st, os.Stdout, *confirmWipe)
|
||||
}
|
||||
|
||||
// ----- daemon components (only wired when unlocked) -----
|
||||
// Pre-declare so the unlock path can wire them later.
|
||||
var (
|
||||
@@ -338,6 +350,9 @@ func run(args []string) error {
|
||||
getMorningStatus: func(ctx context.Context) []ipc.MorningRoutineStatus { return tl.morningStatus(ctx, time.Now()) },
|
||||
getDayPlan: func(ctx context.Context) ipc.DayPlan { return tl.dayPlan(ctx, time.Now()) },
|
||||
getEvents: intakeEventsFn(evBus),
|
||||
getDecisions: turnDecisionsFn(voiceW),
|
||||
seedStore: seedStoreIfAllowed(st),
|
||||
nexus: nexusOf(voiceW),
|
||||
}
|
||||
if voiceW != nil && voiceW.handler != nil {
|
||||
api := coreAPI.(*daemonAPI)
|
||||
@@ -605,6 +620,8 @@ func run(args []string) error {
|
||||
getMorningStatus: func(ctx context.Context) []ipc.MorningRoutineStatus { return tl.morningStatus(ctx, time.Now()) },
|
||||
getDayPlan: func(ctx context.Context) ipc.DayPlan { return tl.dayPlan(ctx, time.Now()) },
|
||||
getEvents: intakeEventsFn(evBus),
|
||||
getDecisions: turnDecisionsFn(voiceW),
|
||||
seedStore: seedStoreIfAllowed(st),
|
||||
}
|
||||
if voiceW != nil && voiceW.handler != nil {
|
||||
newAPI.chatFn = voiceW.handler.handleText
|
||||
|
||||
+29
-31
@@ -8,6 +8,7 @@ import (
|
||||
"unicode"
|
||||
|
||||
"github.com/kami/maven/internal/dialogue"
|
||||
"github.com/kami/maven/internal/lexicon"
|
||||
"github.com/kami/maven/internal/store"
|
||||
)
|
||||
|
||||
@@ -26,44 +27,41 @@ import (
|
||||
// does not say what to do with it. Acting on the bare word would guess, and a
|
||||
// wrong guess here closes work he never finished.
|
||||
|
||||
// candidateOrdinals — the words that pick a position, by index. Prefix match,
|
||||
// because Russian declines them: "первый", "первую", "первое".
|
||||
var candidateOrdinals = []struct {
|
||||
word string
|
||||
nth int
|
||||
}{
|
||||
{"перв", 1}, {"втор", 2}, {"трет", 3}, {"четв", 4}, {"пят", 5},
|
||||
{"first", 1}, {"second", 2}, {"third", 3},
|
||||
}
|
||||
// The position words come from the lexicon, which lists every form with its
|
||||
// position and "последний" as -1 (V-522). They used to be stem prefixes here —
|
||||
// {"перв", 1}, {"втор", 2} — which is the shape that sweep removed: a stem
|
||||
// decides meaning by guessing where a word ends, and "трет" also opens
|
||||
// "third-party". The lexicon runs to twelve rather than five, so he can pick
|
||||
// past the fifth of a longer list; resolveCandidate already answers a position
|
||||
// she did not read.
|
||||
|
||||
// candidateDigits — "второй" said as a number. Matched whole, never by prefix:
|
||||
// "15" starts with "1" and is a time, not a position.
|
||||
// "15" starts with "1" and is a time, not a position. Digits are not a Russian
|
||||
// word list, so they stay here rather than in the lexicon.
|
||||
var candidateDigits = map[string]int{"1": 1, "2": 2, "3": 3, "4": 4, "5": 5}
|
||||
|
||||
// candidateLast — "последний" picks the end of the list whatever its length.
|
||||
var candidateLast = []string{"последн", "last"}
|
||||
|
||||
// parseOrdinal reads which position he named. 0 and false when he named none.
|
||||
// A negative result means the last one.
|
||||
func parseOrdinal(text string) (int, bool) {
|
||||
// Token by token, not substring: " 1" would otherwise match inside
|
||||
// "напомни в 15:00" and turn a reminder into a selection.
|
||||
for _, tok := range strings.FieldsFunc(strings.ToLower(text), func(r rune) bool {
|
||||
toks := strings.FieldsFunc(strings.ToLower(text), func(r rune) bool {
|
||||
return !unicode.IsLetter(r) && !unicode.IsDigit(r)
|
||||
}) {
|
||||
for _, w := range candidateLast {
|
||||
if strings.HasPrefix(tok, w) {
|
||||
return -1, true
|
||||
}
|
||||
}
|
||||
})
|
||||
for i, tok := range toks {
|
||||
if n, ok := candidateDigits[tok]; ok {
|
||||
return n, true
|
||||
}
|
||||
for _, o := range candidateOrdinals {
|
||||
// Prefix, because Russian declines them: "первый", "первую".
|
||||
if strings.HasPrefix(tok, o.word) {
|
||||
return o.nth, true
|
||||
}
|
||||
// A spoken half hour names the hour it is entering with the same
|
||||
// genitive ordinal: "в половине восьмого" is 07:30, not the eighth
|
||||
// thing she read out. She reads a list and he answers with a time
|
||||
// often enough that this has to be declined here, or the reminder
|
||||
// becomes a selection.
|
||||
if i > 0 && lexicon.IsHalfHour(toks[i-1]) {
|
||||
continue
|
||||
}
|
||||
if n, ok := lexicon.Ordinal(tok); ok {
|
||||
return n, true
|
||||
}
|
||||
}
|
||||
return 0, false
|
||||
@@ -97,20 +95,20 @@ func parseCandidateVerb(text string) (status, say string, ok bool) {
|
||||
// offerCandidates records the list she just read, so his next words can pick
|
||||
// from it. Best effort: no session store, or a session that expired between the
|
||||
// question and the answer, means the words route normally.
|
||||
func (h *reactiveHandler) offerCandidates(cands []dialogue.Candidate) {
|
||||
func (h *reactiveHandler) offerCandidates(ctx context.Context, cands []dialogue.Candidate) {
|
||||
if h.dialogueSessions == nil || len(cands) == 0 {
|
||||
return
|
||||
}
|
||||
h.dialogueSessions.SetCandidates(voiceDialogueID, h.now(), cands)
|
||||
h.dialogueSessions.SetCandidates(dialogueIDOf(ctx), h.now(), cands)
|
||||
}
|
||||
|
||||
// resolveCandidate handles "второй", "первую сделал", "последнюю убери" against
|
||||
// the list she just read.
|
||||
func (h *reactiveHandler) resolveCandidate(ctx context.Context, text string) (string, bool) {
|
||||
func (h *reactiveHandler) resolveCandidate(ctx context.Context, text string, src turnSource) (string, bool) {
|
||||
if h.dialogueSessions == nil {
|
||||
return "", false
|
||||
}
|
||||
sess := h.dialogueSessions.Get(voiceDialogueID, h.now())
|
||||
sess := h.dialogueSessions.Get(dialogueIDOf(ctx), h.now())
|
||||
if sess == nil || len(sess.Candidates) == 0 {
|
||||
return "", false
|
||||
}
|
||||
@@ -133,13 +131,13 @@ func (h *reactiveHandler) resolveCandidate(ctx context.Context, text string) (st
|
||||
// a sentence, and the second half is the next turn.
|
||||
return pick.Label, true
|
||||
}
|
||||
if err := h.api.SetTaskStatus(ctx, pick.Ref, status, h.now(), "tap:voice"); err != nil {
|
||||
if err := h.api.SetTaskStatus(ctx, pick.Ref, status, h.now(), string(src)); err != nil {
|
||||
log.Printf("voice: candidate %d → %s: %v", pick.Ref, status, err)
|
||||
return "не получилось изменить задачу.", true
|
||||
}
|
||||
// Spent: the list she read is no longer the list, and a second ordinal
|
||||
// against it would close the wrong task.
|
||||
h.dialogueSessions.SetCandidates(voiceDialogueID, h.now(), nil)
|
||||
h.dialogueSessions.SetCandidates(dialogueIDOf(ctx), h.now(), nil)
|
||||
log.Printf("voice: candidate %d (%q) → %s", pick.Ref, pick.Label, status)
|
||||
return say + ": " + pick.Label, true
|
||||
}
|
||||
|
||||
+24
-12
@@ -27,6 +27,16 @@ func TestParseOrdinalReadsThePosition(t *testing.T) {
|
||||
{"", 0, false},
|
||||
// A digit inside a time is not a position.
|
||||
{"напомни в 15:00", 0, false},
|
||||
// Forms the stem list used to miss, and positions past its fifth.
|
||||
{"вторым", 2, true},
|
||||
{"седьмую", 7, true},
|
||||
{"одиннадцатый", 11, true},
|
||||
// A spoken half hour names its hour with the same genitive ordinal, so
|
||||
// this is 07:30 and not the eighth thing she read out (V-522).
|
||||
{"напомни в половине восьмого", 0, false},
|
||||
{"полвосьмого", 0, false},
|
||||
// The ordinal still wins when the half word is not in front of it.
|
||||
{"восьмую сделал", 8, true},
|
||||
}
|
||||
for _, c := range cases {
|
||||
got, ok := parseOrdinal(c.text)
|
||||
@@ -38,7 +48,7 @@ func TestParseOrdinalReadsThePosition(t *testing.T) {
|
||||
|
||||
func TestOrdinalPassesWithNothingOffered(t *testing.T) {
|
||||
h, _, _ := newClarifyHandler(t)
|
||||
if _, handled := h.resolveCandidate(context.Background(), "второй"); handled {
|
||||
if _, handled := h.resolveCandidate(context.Background(), "второй", sourceVoice); handled {
|
||||
t.Error("an ordinal with no list behind it was claimed")
|
||||
}
|
||||
}
|
||||
@@ -47,14 +57,14 @@ func TestOrdinalReadsBackWithoutAVerb(t *testing.T) {
|
||||
h, st, _ := newClarifyHandler(t)
|
||||
ctx := context.Background()
|
||||
ids := seedTasks(t, st, "купить хлеб", "позвонить маме")
|
||||
putCandidates(h, ids, "купить хлеб", "позвонить маме")
|
||||
putCandidates(h, context.Background(), ids, "купить хлеб", "позвонить маме")
|
||||
|
||||
reply, handled := h.resolveCandidate(ctx, "второй")
|
||||
reply, handled := h.resolveCandidate(ctx, "второй", sourceVoice)
|
||||
if !handled || !strings.Contains(reply, "позвонить маме") {
|
||||
t.Fatalf("a bare ordinal did not read the task back: %q handled=%v", reply, handled)
|
||||
}
|
||||
// Still live: naming one is often the first half of a sentence.
|
||||
if _, handled := h.resolveCandidate(ctx, "первый"); !handled {
|
||||
if _, handled := h.resolveCandidate(ctx, "первый", sourceVoice); !handled {
|
||||
t.Error("the list was spent by a read-back")
|
||||
}
|
||||
}
|
||||
@@ -63,9 +73,9 @@ func TestOrdinalWithAVerbMovesTheTask(t *testing.T) {
|
||||
h, st, _ := newClarifyHandler(t)
|
||||
ctx := context.Background()
|
||||
ids := seedTasks(t, st, "купить хлеб", "позвонить маме")
|
||||
putCandidates(h, ids, "купить хлеб", "позвонить маме")
|
||||
putCandidates(h, context.Background(), ids, "купить хлеб", "позвонить маме")
|
||||
|
||||
reply, handled := h.resolveCandidate(ctx, "первую сделал")
|
||||
reply, handled := h.resolveCandidate(ctx, "первую сделал", sourceVoice)
|
||||
if !handled || !strings.Contains(reply, "купить хлеб") {
|
||||
t.Fatalf("the pick was not acted on: %q handled=%v", reply, handled)
|
||||
}
|
||||
@@ -80,7 +90,7 @@ func TestOrdinalWithAVerbMovesTheTask(t *testing.T) {
|
||||
}
|
||||
// Spent: a second ordinal against a list that no longer holds would close
|
||||
// the wrong task.
|
||||
if _, handled := h.resolveCandidate(ctx, "второй"); handled {
|
||||
if _, handled := h.resolveCandidate(ctx, "второй", sourceVoice); handled {
|
||||
t.Error("the list survived the pick it was spent on")
|
||||
}
|
||||
}
|
||||
@@ -88,9 +98,9 @@ func TestOrdinalWithAVerbMovesTheTask(t *testing.T) {
|
||||
func TestOrdinalPastTheEndSaysHowMany(t *testing.T) {
|
||||
h, st, _ := newClarifyHandler(t)
|
||||
ids := seedTasks(t, st, "купить хлеб")
|
||||
putCandidates(h, ids, "купить хлеб")
|
||||
putCandidates(h, context.Background(), ids, "купить хлеб")
|
||||
|
||||
reply, handled := h.resolveCandidate(context.Background(), "третий")
|
||||
reply, handled := h.resolveCandidate(context.Background(), "третий", sourceVoice)
|
||||
if !handled || !strings.Contains(reply, "1") {
|
||||
t.Fatalf("a position she never read was not answered: %q handled=%v", reply, handled)
|
||||
}
|
||||
@@ -118,11 +128,13 @@ func seedTasks(t *testing.T, st *store.Store, texts ...string) []int64 {
|
||||
return ids
|
||||
}
|
||||
|
||||
func putCandidates(h *reactiveHandler, ids []int64, labels ...string) {
|
||||
// putCandidates binds a list to the reach the ctx names, the way queryTasks
|
||||
// does when she recites one.
|
||||
func putCandidates(h *reactiveHandler, ctx context.Context, ids []int64, labels ...string) {
|
||||
cands := make([]dialogue.Candidate, 0, len(ids))
|
||||
for i, id := range ids {
|
||||
cands = append(cands, dialogue.Candidate{Kind: "task", Ref: id, Label: labels[i]})
|
||||
}
|
||||
h.dialogueSessions.Put(voiceDialogueID, &dialogue.Session{Timestamp: h.now()})
|
||||
h.offerCandidates(cands)
|
||||
h.dialogueSessions.Put(dialogueIDOf(ctx), &dialogue.Session{Timestamp: h.now()})
|
||||
h.offerCandidates(ctx, cands)
|
||||
}
|
||||
|
||||
@@ -77,6 +77,32 @@ var worldSeeds = []string{
|
||||
"что мне почитать про историю",
|
||||
"что я должен знать про питон",
|
||||
"what can i watch tonight",
|
||||
// A third shape that looks personal and is not: asking when something
|
||||
// happens (Vikunja #553). "во сколько закат сегодня" scored personal,
|
||||
// because "что у меня сегодня" and "когда моя встреча" put that frame on
|
||||
// the personal side and nothing here answered it. The sunset is the one
|
||||
// thing on his list that is the same for everybody standing outside.
|
||||
// "сегодня" is carried on purpose. Without it these caught nothing: the
|
||||
// day word is most of what pulls the frame personal, because "что у меня
|
||||
// сегодня" is a personal seed and the day word is the half it shares.
|
||||
"во сколько сегодня открывается магазин",
|
||||
"когда сегодня начинается матч",
|
||||
"во сколько сегодня восход солнца",
|
||||
// The other frame a day word carries, and the same story: "что у меня
|
||||
// сегодня" is a personal seed, so "какой сегодня праздник" and "что
|
||||
// интересного произошло сегодня в мире" were refused as his after the
|
||||
// topic seeds had already let them past the weather source.
|
||||
"какой сегодня курс валют",
|
||||
"что сегодня происходит в мире",
|
||||
// The narrative shape (Vikunja #554). "расскажи про Байкал" was refused as
|
||||
// his by 0.0052, and nothing here was phrased as an order rather than a
|
||||
// question: every world seed above opens with an interrogative. So a world
|
||||
// question that names its subject and asks for prose landed nearer "я тебе
|
||||
// рассказывал об этом?", which is the same verb about his own words.
|
||||
"расскажи про байкал",
|
||||
"расскажи про древний рим",
|
||||
"объясни как работает двигатель",
|
||||
"tell me about the roman empire",
|
||||
}
|
||||
|
||||
// personalBoundary holds the embedded seeds. Zero value is usable and means
|
||||
|
||||
@@ -68,6 +68,27 @@ func TestONNXPersonalBoundary(t *testing.T) {
|
||||
{"я хочу узнать про рим", false},
|
||||
{"кто такой гагарин", false},
|
||||
{"how do i boil an egg", false},
|
||||
// Asking when a public thing happens (Vikunja #553). "во сколько закат
|
||||
// сегодня" was answered "не знаю — не нашла у тебя такой записи",
|
||||
// because the frame lived only on the personal side. The pair above it
|
||||
// is the control: "во сколько у меня встреча" is the same frame about
|
||||
// something that IS his, and it has to stay personal.
|
||||
{"во сколько закат сегодня", false},
|
||||
{"когда сегодня заканчивается концерт", false},
|
||||
{"во сколько завтра открывается аптека", false},
|
||||
// The "какой сегодня X" frame. These clear the weather topic after the
|
||||
// V-553 seeds and were then refused here, which is the same defect one
|
||||
// source further down the chain.
|
||||
{"какой сегодня праздник", false},
|
||||
{"что интересного произошло сегодня в мире", false},
|
||||
{"кто выиграл вчера матч", false},
|
||||
// The narrative shape, held out from the seeds above (Vikunja #554).
|
||||
// The control is the row after them: the same verb about his own words
|
||||
// is still his.
|
||||
{"расскажи про эверест", false},
|
||||
{"расскажи про войну 1812 года", false},
|
||||
{"объясни что такое инфляция", false},
|
||||
{"я рассказывал тебе про байкал?", true},
|
||||
}
|
||||
|
||||
h := &reactiveHandler{recall: recallWiring{embedder: emb}}
|
||||
|
||||
@@ -0,0 +1,160 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// "отметь второй пункт" names a position, and only the daemon knows which item
|
||||
// that is. The router fills the value slot with "2"; this is where it becomes an
|
||||
// item id (Vikunja #516).
|
||||
func TestPositionResolvesAgainstTheLastSpokenList(t *testing.T) {
|
||||
praxis := newFakePraxis(t, `[
|
||||
{"id":"item_a","title":"диск заканчивается"},
|
||||
{"id":"item_b","title":"бэкап не прошёл"},
|
||||
{"id":"item_c","title":"сертификат истекает"}
|
||||
]`)
|
||||
h := newPraxisTestHandler(t, praxis)
|
||||
|
||||
if reply := h.handlePraxisAct(context.Background(), praxisActDec("list_attention")); reply == "" {
|
||||
t.Fatal("attention returned nothing")
|
||||
}
|
||||
|
||||
cases := []struct{ ref, wantItem string }{
|
||||
{"2", "item_b"},
|
||||
{"1", "item_a"},
|
||||
{"last", "item_c"},
|
||||
}
|
||||
for _, c := range cases {
|
||||
praxis.ResetRequests()
|
||||
reply := h.handlePraxisAct(context.Background(), praxisItemDec("acknowledge_item", c.ref))
|
||||
if !strings.Contains(reply, "принято") {
|
||||
t.Errorf("ref %q: reply %q", c.ref, reply)
|
||||
}
|
||||
if !requestedPathContaining(praxis, c.wantItem) {
|
||||
t.Errorf("ref %q did not acknowledge %s; paths %v", c.ref, c.wantItem, paths(praxis))
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A position past the end must not acknowledge the wrong item. It asks.
|
||||
func TestPositionPastTheEndAsksInsteadOfGuessing(t *testing.T) {
|
||||
praxis := newFakePraxis(t, `[{"id":"item_a","title":"диск заканчивается"}]`)
|
||||
h := newPraxisTestHandler(t, praxis)
|
||||
h.handlePraxisAct(context.Background(), praxisActDec("list_attention"))
|
||||
|
||||
praxis.ResetRequests()
|
||||
reply := h.handlePraxisAct(context.Background(), praxisItemDec("resolve_item", "4"))
|
||||
if !strings.Contains(reply, "какой пункт") {
|
||||
t.Errorf("a position with no item should ask, got %q", reply)
|
||||
}
|
||||
if requestedPathContaining(praxis, "item_a") {
|
||||
t.Error("the only surfaced item was resolved for a position that did not name it")
|
||||
}
|
||||
}
|
||||
|
||||
// No digest yet means no positions. Nothing is mutated.
|
||||
func TestPositionWithNoSpokenListAsks(t *testing.T) {
|
||||
praxis := newFakePraxis(t, `[]`)
|
||||
h := newPraxisTestHandler(t, praxis)
|
||||
|
||||
reply := h.handlePraxisAct(context.Background(), praxisItemDec("acknowledge_item", "1"))
|
||||
if !strings.Contains(reply, "какой пункт") {
|
||||
t.Errorf("want the ask, got %q", reply)
|
||||
}
|
||||
}
|
||||
|
||||
// An explicit id is not a position and passes through untouched.
|
||||
func TestExplicitItemIDIsNotRewritten(t *testing.T) {
|
||||
praxis := newFakePraxis(t, `[{"id":"item_a","title":"диск"}]`)
|
||||
h := newPraxisTestHandler(t, praxis)
|
||||
h.handlePraxisAct(context.Background(), praxisActDec("list_attention"))
|
||||
|
||||
praxis.ResetRequests()
|
||||
h.handlePraxisAct(context.Background(), praxisItemDec("pin_item", "item_zz"))
|
||||
if !requestedPathContaining(praxis, "item_zz") {
|
||||
t.Errorf("the id he gave was not the one called; paths %v", paths(praxis))
|
||||
}
|
||||
}
|
||||
|
||||
// An item Praxis sent without a title is never spoken, so it holds no position.
|
||||
func TestUnspokenItemsHoldNoPosition(t *testing.T) {
|
||||
praxis := newFakePraxis(t, `[
|
||||
{"id":"item_silent"},
|
||||
{"id":"item_said","title":"бэкап не прошёл"}
|
||||
]`)
|
||||
h := newPraxisTestHandler(t, praxis)
|
||||
h.handlePraxisAct(context.Background(), praxisActDec("list_attention"))
|
||||
|
||||
praxis.ResetRequests()
|
||||
h.handlePraxisAct(context.Background(), praxisItemDec("acknowledge_item", "1"))
|
||||
if !requestedPathContaining(praxis, "item_said") {
|
||||
t.Errorf("position 1 is the first item she SAID; paths %v", paths(praxis))
|
||||
}
|
||||
}
|
||||
|
||||
// The item id travels in the POST body, so that is what these read.
|
||||
func paths(f *fakeServer) []string {
|
||||
var out []string
|
||||
for _, r := range f.Requests() {
|
||||
out = append(out, r.Path+" "+string(r.Body))
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func requestedPathContaining(f *fakeServer, want string) bool {
|
||||
for _, r := range f.Requests() {
|
||||
if strings.Contains(string(r.Body), want) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// "отметь это как сделанное" after a one-item digest points at that item.
|
||||
func TestDemonstrativeResolvesWhenOneItemWasSpoken(t *testing.T) {
|
||||
praxis := newFakePraxis(t, `[{"id":"item_only","title":"бэкап не прошёл"}]`)
|
||||
h := newPraxisTestHandler(t, praxis)
|
||||
h.handlePraxisAct(context.Background(), praxisActDec("list_attention"))
|
||||
|
||||
praxis.ResetRequests()
|
||||
reply := h.handlePraxisAct(context.Background(), praxisItemDec("acknowledge_item", "this"))
|
||||
if !strings.Contains(reply, "принято") {
|
||||
t.Errorf("reply %q", reply)
|
||||
}
|
||||
if !requestedPathContaining(praxis, "item_only") {
|
||||
t.Errorf("the one surfaced item was not acknowledged; paths %v", paths(praxis))
|
||||
}
|
||||
}
|
||||
|
||||
// Pointing at one of several is a guess, and a wrong guess transitions the wrong
|
||||
// item. The turn goes back to the cascade instead.
|
||||
func TestDemonstrativeWithSeveralItemsGivesTheTurnBack(t *testing.T) {
|
||||
praxis := newFakePraxis(t, `[
|
||||
{"id":"item_a","title":"диск"},
|
||||
{"id":"item_b","title":"бэкап"}
|
||||
]`)
|
||||
h := newPraxisTestHandler(t, praxis)
|
||||
h.handlePraxisAct(context.Background(), praxisActDec("list_attention"))
|
||||
|
||||
praxis.ResetRequests()
|
||||
if reply := h.handlePraxisAct(context.Background(), praxisItemDec("resolve_item", "this")); reply != "" {
|
||||
t.Errorf("want a fall-through, got %q", reply)
|
||||
}
|
||||
for _, p := range paths(praxis) {
|
||||
if strings.Contains(p, "resolve") {
|
||||
t.Error("an ambiguous demonstrative resolved an item anyway")
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// "я это сделал" with no digest behind it is a sentence about his day.
|
||||
func TestDemonstrativeWithNoDigestGivesTheTurnBack(t *testing.T) {
|
||||
praxis := newFakePraxis(t, `[]`)
|
||||
h := newPraxisTestHandler(t, praxis)
|
||||
|
||||
if reply := h.handlePraxisAct(context.Background(), praxisItemDec("resolve_item", "this")); reply != "" {
|
||||
t.Errorf("want a fall-through, got %q", reply)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,95 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/loop"
|
||||
"github.com/kami/maven/internal/store"
|
||||
)
|
||||
|
||||
// The row is the whole mechanism, and nothing wrote it (Vikunja #532).
|
||||
//
|
||||
// The existing hysteresis test in internal/store scores the pure function and
|
||||
// passed throughout, which is exactly why this went unnoticed: Resolve was
|
||||
// always correct and was always handed the cold-start Away. So this test asserts
|
||||
// the round trip — the tick writes what gather resolved, and the next load
|
||||
// reads it back — rather than re-testing the function.
|
||||
func TestTickPersistsTheResolvedBucket(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
st := newTestStore(t)
|
||||
now := time.Now()
|
||||
|
||||
// Cold start: no row, so a load must say Away and the zero time.
|
||||
b, score, updated, err := st.LoadPresenceState(ctx)
|
||||
if err != nil {
|
||||
t.Fatalf("load before: %v", err)
|
||||
}
|
||||
if b != store.Away || score != 0 || !updated.IsZero() {
|
||||
t.Fatalf("cold start = %s/%v/%v, want away/0/zero", b, score, updated)
|
||||
}
|
||||
|
||||
tl := &tickLoop{store: st}
|
||||
tl.savePresence(ctx, loop.State{Presence: store.Present, PresenceScore: 0.9}, now)
|
||||
|
||||
b, score, updated, err = st.LoadPresenceState(ctx)
|
||||
if err != nil {
|
||||
t.Fatalf("load after: %v", err)
|
||||
}
|
||||
if b != store.Present {
|
||||
t.Errorf("bucket = %s, want present", b)
|
||||
}
|
||||
if score != 0.9 {
|
||||
t.Errorf("score = %v, want 0.9", score)
|
||||
}
|
||||
if updated.IsZero() {
|
||||
t.Error("updated_ts was not written, so /dash still reads (never)")
|
||||
}
|
||||
}
|
||||
|
||||
// The singleton stays a singleton, and a later tick overwrites rather than
|
||||
// accumulating. A row per tick would make LoadPresenceState's single-row query
|
||||
// return whichever one SQLite felt like.
|
||||
func TestPresenceStateIsOverwrittenNotAppended(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
st := newTestStore(t)
|
||||
tl := &tickLoop{store: st}
|
||||
now := time.Now()
|
||||
|
||||
tl.savePresence(ctx, loop.State{Presence: store.Present, PresenceScore: 0.9}, now)
|
||||
tl.savePresence(ctx, loop.State{Presence: store.Away, PresenceScore: 0.1}, now.Add(time.Minute))
|
||||
|
||||
b, score, _, err := st.LoadPresenceState(ctx)
|
||||
if err != nil {
|
||||
t.Fatalf("load: %v", err)
|
||||
}
|
||||
if b != store.Away || score != 0.1 {
|
||||
t.Fatalf("got %s/%v, want the second write (away/0.1)", b, score)
|
||||
}
|
||||
}
|
||||
|
||||
// What the persisted row buys: the hold band. A score sitting between Exit and
|
||||
// Enter holds Present when the last bucket was Present, and stays Away when it
|
||||
// was Away. Before the write existed the second arm was the only one that could
|
||||
// ever run, so presence dropped at roughly four minutes of idle instead of
|
||||
// holding to the exit threshold at about nine.
|
||||
func TestPersistedBucketIsWhatFeedsHysteresis(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
st := newTestStore(t)
|
||||
tl := &tickLoop{store: st}
|
||||
|
||||
mid := (store.PresenceExit + store.PresenceEnter) / 2
|
||||
if mid <= store.PresenceExit || mid >= store.PresenceEnter {
|
||||
t.Fatalf("%v is not inside the hold band", mid)
|
||||
}
|
||||
|
||||
tl.savePresence(ctx, loop.State{Presence: store.Present, PresenceScore: 0.9}, time.Now())
|
||||
last, _, _, err := st.LoadPresenceState(ctx)
|
||||
if err != nil {
|
||||
t.Fatalf("load: %v", err)
|
||||
}
|
||||
if got := store.Resolve(mid, last); got != store.Present {
|
||||
t.Errorf("Resolve(%v, %s) = %s, want present — the hold band did not apply", mid, last, got)
|
||||
}
|
||||
}
|
||||
@@ -2,8 +2,11 @@ package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"math"
|
||||
"net/http"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
@@ -31,6 +34,57 @@ func (f *fixedEmbedder) Embed(_ context.Context, text string) ([]float32, error)
|
||||
return v, nil
|
||||
}
|
||||
|
||||
// brokenEmbedder fails every call, which is what an ONNX session error looks
|
||||
// like from the query chain's side.
|
||||
type brokenEmbedder struct{}
|
||||
|
||||
func (brokenEmbedder) Dim() int { return 4 }
|
||||
func (brokenEmbedder) Close() error { return nil }
|
||||
func (brokenEmbedder) Embed(context.Context, string) ([]float32, error) {
|
||||
return nil, errors.New("onnx: session failed")
|
||||
}
|
||||
|
||||
// TestQueryEmbedFailureDoesNotStopTheChain — V-568. The embed source used to
|
||||
// claim the turn on an embedder error, so one failing EmbedQuery answered every
|
||||
// question below it with "не смогла ответить", including the ones the search
|
||||
// answers without an embedder at all. A source that could not look must pass.
|
||||
func TestQueryEmbedFailureDoesNotStopTheChain(t *testing.T) {
|
||||
const q = "почему небо голубое"
|
||||
|
||||
h, _ := searchHandler(t, searchBody, http.StatusOK)
|
||||
h.api = ipc.NewStoreAPI(newTestStore(t))
|
||||
h.recall = recallWiring{embedder: brokenEmbedder{}, minScore: 0.55, minMargin: 0.008}
|
||||
h.now = time.Now
|
||||
|
||||
// The embed source itself passes rather than claiming.
|
||||
turn := &queryTurn{dec: router.Decision{Intent: router.IntentQuery, Utterance: q}}
|
||||
if reply, ok := h.queryEmbed(context.Background(), turn); ok {
|
||||
t.Fatalf("queryEmbed claimed the turn on an embedder error: %q", reply)
|
||||
}
|
||||
|
||||
// And the whole chain still reaches the search below it.
|
||||
reply := askQuery(t, h, q)
|
||||
if !strings.Contains(reply, "рэлеевского рассеяния") {
|
||||
t.Fatalf("reply = %q, want the search answer", reply)
|
||||
}
|
||||
}
|
||||
|
||||
// The recall sources read the empty vector the failed embed left behind, and
|
||||
// neither of them may turn that into an answer: no vector means they could not
|
||||
// look, which is not the same as looking and finding nothing.
|
||||
func TestQueryRecallPassesWithoutAVector(t *testing.T) {
|
||||
h, _ := buildRecallHandler(t, "где молоко", []recallCase{
|
||||
{text: "молоко стоит в холодильнике", score: 0.90, kind: "note"},
|
||||
})
|
||||
turn := &queryTurn{dec: router.Decision{Intent: router.IntentQuery, Utterance: "где молоко"}}
|
||||
if reply, ok := h.queryMemory(context.Background(), turn); ok {
|
||||
t.Errorf("queryMemory claimed with no vector: %q", reply)
|
||||
}
|
||||
if reply, ok := h.queryNotes(context.Background(), turn); ok {
|
||||
t.Errorf("queryNotes claimed with no vector: %q", reply)
|
||||
}
|
||||
}
|
||||
|
||||
// scoreVec builds a unit vector whose cosine against the query vector
|
||||
// (1,0,0,0) is exactly score.
|
||||
func scoreVec(score float64) []float32 {
|
||||
|
||||
@@ -0,0 +1,56 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"sync"
|
||||
)
|
||||
|
||||
// The query source that claimed a turn was visible in the daemon log and
|
||||
// nowhere else (V-539). A QA step reading /chat could see a wrong answer but
|
||||
// not tell a wrong answer from a wrongly ordered chain: "почему небо голубое"
|
||||
// answered badly reads the same whether search claimed it, the ZIM did, or the
|
||||
// resident model answered from memory.
|
||||
//
|
||||
// It rides the context rather than a return value because handleText answers
|
||||
// every reach through one string, and threading a second value through the
|
||||
// whole action dispatch would change a signature the mic, telegram and the web
|
||||
// all share. The sink is per turn, created by the caller that wants to read it;
|
||||
// a turn with no sink notes nothing, which is what the mic path does.
|
||||
type querySourceKey struct{}
|
||||
|
||||
// querySourceSink holds the name of the source that claimed one turn. The mutex
|
||||
// is there because a query source may fan out to goroutines of its own, not
|
||||
// because two turns share a sink.
|
||||
type querySourceSink struct {
|
||||
mu sync.Mutex
|
||||
name string
|
||||
}
|
||||
|
||||
func (s *querySourceSink) note(name string) {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
s.name = name
|
||||
}
|
||||
|
||||
// Name is the source that claimed, or empty when nothing did or the turn was
|
||||
// not a query at all.
|
||||
func (s *querySourceSink) Name() string {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
return s.name
|
||||
}
|
||||
|
||||
// withQuerySourceSink returns a context that collects the claiming source, and
|
||||
// the sink to read after the turn has answered.
|
||||
func withQuerySourceSink(ctx context.Context) (context.Context, *querySourceSink) {
|
||||
sink := &querySourceSink{}
|
||||
return context.WithValue(ctx, querySourceKey{}, sink), sink
|
||||
}
|
||||
|
||||
// noteQuerySource records which source claimed the turn. It is a no-op when the
|
||||
// caller did not ask for one.
|
||||
func noteQuerySource(ctx context.Context, name string) {
|
||||
if sink, ok := ctx.Value(querySourceKey{}).(*querySourceSink); ok {
|
||||
sink.note(name)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,35 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestQuerySourceSinkCollectsTheClaimingName(t *testing.T) {
|
||||
ctx, sink := withQuerySourceSink(context.Background())
|
||||
if sink.Name() != "" {
|
||||
t.Fatalf("a fresh sink names a source: %q", sink.Name())
|
||||
}
|
||||
noteQuerySource(ctx, "kiwix")
|
||||
if got := sink.Name(); got != "kiwix" {
|
||||
t.Errorf("sink.Name() = %q, want kiwix", got)
|
||||
}
|
||||
}
|
||||
|
||||
// A turn with no sink must not panic. The mic path asks for no source, and a
|
||||
// query source calls noteQuerySource unconditionally.
|
||||
func TestNoteQuerySourceWithoutASinkIsSilent(t *testing.T) {
|
||||
noteQuerySource(context.Background(), "search")
|
||||
}
|
||||
|
||||
// The last source to claim wins, because only one does: actionQuery returns on
|
||||
// the first claim. This pins that the sink overwrites rather than appends, so a
|
||||
// second turn on the same context could not read a stale name.
|
||||
func TestQuerySourceSinkKeepsTheLastNote(t *testing.T) {
|
||||
ctx, sink := withQuerySourceSink(context.Background())
|
||||
noteQuerySource(ctx, "search")
|
||||
noteQuerySource(ctx, "kiwix")
|
||||
if got := sink.Name(); got != "kiwix" {
|
||||
t.Errorf("sink.Name() = %q, want kiwix", got)
|
||||
}
|
||||
}
|
||||
@@ -44,9 +44,12 @@ func TestReactiveNotesReminders(t *testing.T) {
|
||||
HasTime: true,
|
||||
},
|
||||
}
|
||||
// The confirmation is phrased from the row now (Vikunja #507), so it
|
||||
// names the stored hour rather than leaving the replier to read one
|
||||
// out of the sentence.
|
||||
reply := h.applyAction(ctx, dec)
|
||||
if reply != "" {
|
||||
t.Errorf("expected empty reply from applyAction, got %q", reply)
|
||||
if want := "хорошо, напомню завтра в " + fireAt.Format("15:04") + "."; reply != want {
|
||||
t.Errorf("reply = %q, want %q", reply, want)
|
||||
}
|
||||
reminders, err := st.ListReminders(ctx, 10)
|
||||
if err != nil {
|
||||
|
||||
@@ -0,0 +1,51 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// A reminder confirmation is the one sentence that must match a database row.
|
||||
// It used to be phrased by the replier from Slots.Text, which meant it named
|
||||
// whatever hour the sentence contained — including an hour the parser had
|
||||
// rejected or read differently (Vikunja #507).
|
||||
|
||||
func TestReminderConfirmNamesTheStoredHour(t *testing.T) {
|
||||
now := time.Date(2026, 8, 5, 9, 0, 0, 0, time.UTC)
|
||||
got := reminderConfirm(now.Add(10*time.Hour), now) // 19:00 today
|
||||
if !strings.Contains(got, "19:00") {
|
||||
t.Fatalf("confirmation = %q, want the stored 19:00 in it", got)
|
||||
}
|
||||
if !strings.Contains(got, "сегодня") {
|
||||
t.Fatalf("confirmation = %q, want it to say сегодня", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestReminderConfirmUsesADateBeyondTheDayWords(t *testing.T) {
|
||||
// dayPrefix answers "это" past послезавтра, and "напомню это в 09:00" is
|
||||
// not a sentence. A date is.
|
||||
now := time.Date(2026, 8, 5, 9, 0, 0, 0, time.UTC)
|
||||
got := reminderConfirm(now.Add(10*24*time.Hour), now)
|
||||
if strings.Contains(got, "это") {
|
||||
t.Fatalf("confirmation = %q, want a date rather than the fallback day word", got)
|
||||
}
|
||||
if !strings.Contains(got, "15 августа") {
|
||||
t.Fatalf("confirmation = %q, want the date in it", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestReminderConfirmIsFeminineAndInformal(t *testing.T) {
|
||||
// The persona checks the phrasing eval enforces apply here too, and this
|
||||
// sentence never passes through a phraser.
|
||||
now := time.Date(2026, 8, 5, 9, 0, 0, 0, time.UTC)
|
||||
got := reminderConfirm(now.Add(time.Hour), now)
|
||||
for _, bad := range []string{"вы", "ваш", "напомнил ", "рад "} {
|
||||
if strings.Contains(strings.ToLower(got), bad) {
|
||||
t.Fatalf("confirmation = %q contains %q", got, bad)
|
||||
}
|
||||
}
|
||||
if !strings.HasPrefix(got, "хорошо, напомню") {
|
||||
t.Fatalf("confirmation = %q, want it to open with the promise", got)
|
||||
}
|
||||
}
|
||||
@@ -165,6 +165,8 @@ func formatTime(t time.Time) string {
|
||||
n := int(diff.Hours())
|
||||
return fmt.Sprintf("%d %s назад", n, say.CountWord(n, "час", "часа", "часов"))
|
||||
default:
|
||||
return t.Format("2 января 15:04")
|
||||
// Not t.Format("2 января …"): Go reads that as a literal, so every
|
||||
// fact older than a day used to read as January (Vikunja #507).
|
||||
return fmt.Sprintf("%d %s %s", t.Day(), lexicon.MonthGenitive(int(t.Month())), t.Format("15:04"))
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,107 @@
|
||||
// mavend/seed.go — the backdated-fact seam (Vikunja #518).
|
||||
//
|
||||
// The pattern detector needs four events for one action+object, spread by at
|
||||
// least pattern.MinIntervalDays, before it proposes a routine. Nothing could
|
||||
// produce that against a running daemon in one sitting: the only writer is a
|
||||
// fact write at time.Now(), so V-43, V-46, V-247 and V-254 all stopped at the
|
||||
// same missing step and had been stopped there since they were filed.
|
||||
//
|
||||
// This is the write path that unblocks them, and it is deliberately the narrow
|
||||
// one. It takes a fact, not an event, so pattern.Extract runs for real and a
|
||||
// key the extractor ignores seeds nothing. It runs detectAndPropose, so what a
|
||||
// seed proves is the daemon's own wiring rather than the detector in isolation
|
||||
// — which is what an eval-lab fixture would have proved, and is not what those
|
||||
// four tasks doubt.
|
||||
//
|
||||
// It is off unless mavend was started with -allow-seed, and AuthStepUp in the
|
||||
// authority table besides. See ipc.SeedEventReq and auth.Requirement.
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"errors"
|
||||
"fmt"
|
||||
"log"
|
||||
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/pattern"
|
||||
"github.com/kami/maven/internal/store"
|
||||
)
|
||||
|
||||
// errSeedDisabled — what a caller gets on an ordinary box. Named rather than
|
||||
// inline so the mavweb route can tell "not allowed here" apart from "the seed
|
||||
// ran and the extractor declined", which look the same to a reader otherwise.
|
||||
var errSeedDisabled = errors.New("mavend: seeding is off (start with -allow-seed)")
|
||||
|
||||
// seedSource — every seeded fact carries this, and no other writer uses it.
|
||||
// The point is that seeded data stays identifiable forever: a fact that came
|
||||
// from a QA sitting must never be mistaken for something he said, either by a
|
||||
// person reading /history or by the wipe in V-494 when it lands.
|
||||
const seedSource = "seed:qa"
|
||||
|
||||
// seedStoreIfAllowed returns st only when -allow-seed was passed, and logs the
|
||||
// fact loudly when it does. A box that can rewrite its own past should say so
|
||||
// in its boot log, so nobody reads a seeded routine months later as evidence of
|
||||
// something he actually did.
|
||||
func seedStoreIfAllowed(st *store.Store) *store.Store {
|
||||
if !allowSeedOnStart {
|
||||
return nil
|
||||
}
|
||||
log.Printf("seed: -allow-seed is ON — backdated fact writes are permitted under source %q (Vikunja #518)", seedSource)
|
||||
return st
|
||||
}
|
||||
|
||||
// SeedEvent writes the fact at the caller's timestamp, extracts an event from
|
||||
// it, and runs the same detect-and-propose step the voice path runs.
|
||||
//
|
||||
// Best-effort is NOT the shape here, unlike detectPattern: a seed that half
|
||||
// worked is a QA result nobody can trust, so every step reports its own
|
||||
// failure. Extraction declining is not a failure — it is the extractor's
|
||||
// documented answer for a value outside its lexicon, and Extracted says so.
|
||||
func (d *daemonAPI) SeedEvent(ctx context.Context, req ipc.SeedEventReq) (ipc.SeedEventResp, error) {
|
||||
if d.seedStore == nil {
|
||||
return ipc.SeedEventResp{}, errSeedDisabled
|
||||
}
|
||||
if req.Key == "" || req.Value == "" {
|
||||
return ipc.SeedEventResp{}, errors.New("mavend: seed needs a key and a value")
|
||||
}
|
||||
if req.Ts.IsZero() {
|
||||
return ipc.SeedEventResp{}, errors.New("mavend: seed needs an explicit timestamp")
|
||||
}
|
||||
|
||||
// No Subject, unlike the voice path: a seeded key must not queue a Nexus
|
||||
// resolution. QA data has no business reaching the ecosystem.
|
||||
factID, err := d.seedStore.WriteFact(ctx, req.Ts, store.KindSelf, req.Key, req.Value, seedSource, 1.0, sql.NullInt64{})
|
||||
if err != nil {
|
||||
return ipc.SeedEventResp{}, fmt.Errorf("seed write fact: %w", err)
|
||||
}
|
||||
resp := ipc.SeedEventResp{FactID: factID}
|
||||
|
||||
ev := pattern.Extract(factID, req.Key, req.Value, req.Ts)
|
||||
if ev == nil {
|
||||
// The fact is written and stays written. Saying so matters: a caller
|
||||
// that assumed a seed always produces an event would otherwise read
|
||||
// four silent successes and conclude the detector is broken.
|
||||
log.Printf("seed: %s=%s wrote fact %d, no event (value outside the action lexicon)", req.Key, req.Value, factID)
|
||||
return resp, nil
|
||||
}
|
||||
resp.Extracted, resp.Action, resp.Object = true, ev.Action, ev.Object
|
||||
|
||||
eventID, err := d.seedStore.CreateEvent(ctx, factID, ev.Action, ev.Object, req.Ts)
|
||||
if err != nil {
|
||||
return resp, fmt.Errorf("seed create event: %w", err)
|
||||
}
|
||||
resp.EventID = eventID
|
||||
|
||||
r, routineID, err := detectAndPropose(ctx, d.seedStore, ev.Action, ev.Object, req.Ts)
|
||||
if err != nil {
|
||||
return resp, fmt.Errorf("seed detect: %w", err)
|
||||
}
|
||||
if r == nil {
|
||||
return resp, nil // too few events yet, too irregular, or already decided
|
||||
}
|
||||
resp.Proposed, resp.RoutineID, resp.IntervalDays = true, routineID, r.IntervalDays
|
||||
log.Printf("seed: proposed routine %d — %s/%s every %.1f days", routineID, r.Action, r.Object, r.IntervalDays)
|
||||
return resp, nil
|
||||
}
|
||||
@@ -0,0 +1,105 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
)
|
||||
|
||||
// Off is the default and it must mean "nothing to write with", not "permission
|
||||
// to refuse later". A daemonAPI with no seedStore writes no fact at all.
|
||||
func TestSeedRefusedWithoutTheFlag(t *testing.T) {
|
||||
d := &daemonAPI{CoreAPI: ipc.UnimplementedCoreAPI{}}
|
||||
_, err := d.SeedEvent(context.Background(), ipc.SeedEventReq{
|
||||
Key: "cat_water_fountain", Value: "заправил", Ts: time.Now(),
|
||||
})
|
||||
if err == nil {
|
||||
t.Fatal("seed succeeded with no seedStore")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "-allow-seed") {
|
||||
t.Errorf("error does not name the flag: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// The whole point of the task: four seeds spread past the detector's floor
|
||||
// produce a proposal against the real daemon path, which is what nobody could
|
||||
// do before (Vikunja #518). Three seeds must NOT propose — MinEvents is four,
|
||||
// and a test that only checked the happy end would pass on an off-by-one.
|
||||
func TestSeedFourEventsProposesARoutine(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
d := &daemonAPI{CoreAPI: ipc.UnimplementedCoreAPI{}, seedStore: newTestStore(t)}
|
||||
now := time.Now()
|
||||
|
||||
var last ipc.SeedEventResp
|
||||
// Oldest first, three hours apart — past MinIntervalDays (two hours).
|
||||
for i := 3; i >= 0; i-- {
|
||||
var err error
|
||||
last, err = d.SeedEvent(ctx, ipc.SeedEventReq{
|
||||
Key: "cat_water_fountain",
|
||||
Value: "заправил",
|
||||
Ts: now.Add(-time.Duration(i) * 3 * time.Hour),
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatalf("seed %d: %v", i, err)
|
||||
}
|
||||
if !last.Extracted {
|
||||
t.Fatalf("seed %d: no event extracted from a lexicon verb", i)
|
||||
}
|
||||
if i > 0 && last.Proposed {
|
||||
t.Fatalf("proposed after only %d events, MinEvents is 4", 4-i)
|
||||
}
|
||||
}
|
||||
if !last.Proposed {
|
||||
t.Fatal("four spaced events did not propose a routine")
|
||||
}
|
||||
if last.Action != "refill" || last.Object != "cat_water_fountain" {
|
||||
t.Errorf("wrong pair: %s/%s", last.Action, last.Object)
|
||||
}
|
||||
if last.IntervalDays < 0.1 {
|
||||
t.Errorf("interval %v — the detector saw a burst, not a rhythm", last.IntervalDays)
|
||||
}
|
||||
|
||||
// The proposal is readable through the same list the /routines page uses,
|
||||
// which is the wiring an eval-lab fixture would not have proved.
|
||||
proposed, err := d.seedStore.ListProposedRoutines(ctx)
|
||||
if err != nil {
|
||||
t.Fatalf("list: %v", err)
|
||||
}
|
||||
if len(proposed) != 1 {
|
||||
t.Fatalf("expected 1 proposed routine, got %d", len(proposed))
|
||||
}
|
||||
}
|
||||
|
||||
// A value outside the action lexicon writes the fact and says it seeded
|
||||
// nothing. Silence here would read as four working seeds and a broken
|
||||
// detector.
|
||||
func TestSeedReportsWhenExtractionDeclines(t *testing.T) {
|
||||
d := &daemonAPI{CoreAPI: ipc.UnimplementedCoreAPI{}, seedStore: newTestStore(t)}
|
||||
resp, err := d.SeedEvent(context.Background(), ipc.SeedEventReq{
|
||||
Key: "mood", Value: "ok", Ts: time.Now(),
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatalf("seed: %v", err)
|
||||
}
|
||||
if resp.FactID == 0 {
|
||||
t.Error("fact was not written")
|
||||
}
|
||||
if resp.Extracted || resp.EventID != 0 || resp.Proposed {
|
||||
t.Errorf("claimed an event for a non-action value: %+v", resp)
|
||||
}
|
||||
}
|
||||
|
||||
// A seed with no timestamp is refused rather than defaulting to now: the only
|
||||
// reason this seam exists is the caller choosing when, so a zero Ts is a bug in
|
||||
// the caller and must not silently write a fact at the wrong time.
|
||||
func TestSeedRequiresAnExplicitTimestamp(t *testing.T) {
|
||||
d := &daemonAPI{CoreAPI: ipc.UnimplementedCoreAPI{}, seedStore: newTestStore(t)}
|
||||
if _, err := d.SeedEvent(context.Background(), ipc.SeedEventReq{
|
||||
Key: "cat_water_fountain", Value: "заправил",
|
||||
}); err == nil {
|
||||
t.Fatal("seed accepted a zero timestamp")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,116 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"log"
|
||||
"regexp"
|
||||
|
||||
"github.com/kami/maven/internal/phraser"
|
||||
)
|
||||
|
||||
// A question about her — "что ты умеешь", "кто ты" — used to have no answer at
|
||||
// all (Vikunja #555). It reached the personal boundary, which claimed it as his
|
||||
// and said "не знаю — не нашла у тебя такой записи", because the boundary knows
|
||||
// two sides and this is neither: her own description is not his data and it is
|
||||
// not the world's either. Letting it past the boundary is no better, because
|
||||
// then SearXNG answers about somebody else's assistant.
|
||||
//
|
||||
// The description does NOT live in the note store. Notes are his. A note about
|
||||
// her sitting in his index would come back for "что я записал", would be fed to
|
||||
// the digestion worker as something he said, and would be recalled by vector
|
||||
// proximity for questions that are not about her at all. It is her own text, so
|
||||
// it lives here, in one place, and it is the only copy.
|
||||
//
|
||||
// This source sits ABOVE the boundary, because a question about her never had
|
||||
// an answer below it.
|
||||
|
||||
// selfDescription — what she is and what this box actually does. Frozen text,
|
||||
// and the one rule for editing it: name only what is really wired. Anything
|
||||
// that depends on config — the house, the LAN, the feeds, telegram, search — is
|
||||
// named as depending on what he allowed, never claimed outright. Inventing a
|
||||
// capability here is the same defect as inventing a fact, and it is worse than
|
||||
// silence because he would plan around it.
|
||||
//
|
||||
// Written in her own voice, feminine, addressing him informally, because it is
|
||||
// handed to the phraser as the evidence for the answer and the phraser will
|
||||
// keep the words it is given.
|
||||
const selfDescription = `Я Мэйвен, твоя помощница. Я живу на твоём сервере, ` +
|
||||
`и наружу уходит только поисковый запрос — больше ничего.
|
||||
|
||||
Что я делаю сама: запоминаю, что ты мне говоришь, и потом отвечаю на вопросы ` +
|
||||
`об этом; веду заметки; ставлю напоминания; читаю твой календарь и задачи; ` +
|
||||
`отвечаю на вопросы о мире — сначала поиском, а если сети нет, то по ` +
|
||||
`офлайновой энциклопедии.
|
||||
|
||||
Что зависит от того, что ты мне разрешил: дом, локальная сеть, ленты, ` +
|
||||
`список покупок, погода, телеграм. Если что-то из этого не настроено, я ` +
|
||||
`скажу об этом прямо, а не буду выдумывать ответ.
|
||||
|
||||
Говорю по-русски и по-английски.`
|
||||
|
||||
// selfSeeds — the questions this source claims. Scoring data like every other
|
||||
// topic set: editing one moves the recogniser and has to be re-measured against
|
||||
// TestONNXTopics.
|
||||
//
|
||||
// All of them are about HER — what she is, what she can do, who made her. The
|
||||
// neighbouring set is topicAttend, "что требует внимания", which asks about the
|
||||
// state of his things; the two share almost nothing but the second person.
|
||||
var selfSeeds = []string{
|
||||
"что ты умеешь",
|
||||
"что ты можешь делать",
|
||||
"кто ты такая",
|
||||
"расскажи о себе",
|
||||
"какие у тебя возможности",
|
||||
// Added after measuring: it won self by 0.0002, under the margin, and the
|
||||
// floor does not carry it — "способна" names no verb the floor matches.
|
||||
"на что ты способна",
|
||||
"чем ты можешь помочь",
|
||||
"what can you do",
|
||||
"who are you",
|
||||
}
|
||||
|
||||
// selfFloor — the offline floor, for a handler with no embedder or a turn whose
|
||||
// vector never got computed. Narrow on purpose, like every other floor here: it
|
||||
// answers only when the seeds cannot, and a broad guess made blind is worse
|
||||
// than a narrow one.
|
||||
//
|
||||
// Go's \b is ASCII-only and never fires next to a Cyrillic letter, so the
|
||||
// Russian patterns spell the boundary out.
|
||||
var selfPatterns = []*regexp.Regexp{
|
||||
regexp.MustCompile(`(?i)(^|[^\p{L}\p{N}])ты\s+(умеешь|можешь)([^\p{L}\p{N}]|$)`),
|
||||
regexp.MustCompile(`(?i)(^|[^\p{L}\p{N}])кто\s+ты([^\p{L}\p{N}]|$)`),
|
||||
regexp.MustCompile(`(?i)(^|[^\p{L}\p{N}])(расскажи|поведай)\s+о\s+себе([^\p{L}\p{N}]|$)`),
|
||||
regexp.MustCompile(`(?i)\bwhat\s+can\s+you\s+do\b`),
|
||||
regexp.MustCompile(`(?i)\bwho\s+are\s+you\b`),
|
||||
}
|
||||
|
||||
func selfFloor(utterance string) bool {
|
||||
for _, re := range selfPatterns {
|
||||
if re.MatchString(utterance) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// querySelf answers a question about her from selfDescription. The description
|
||||
// goes through the phraser as evidence so the answer is shaped to what he
|
||||
// asked — "что ты умеешь" and "кто ты" want different halves of it — and falls
|
||||
// back to the text itself, which is already readable, if the model is down.
|
||||
func (h *reactiveHandler) querySelf(ctx context.Context, t *queryTurn) (string, bool) {
|
||||
if !h.turnIsAbout(ctx, t, topicSelf, selfFloor) {
|
||||
return "", false
|
||||
}
|
||||
var reply string
|
||||
if h.phraser != nil {
|
||||
var err error
|
||||
reply, err = h.phraser.PhraseSelf(ctx, t.dec.Utterance, selfDescription)
|
||||
if err != nil {
|
||||
log.Printf("voice: phrase self: %v", err)
|
||||
}
|
||||
}
|
||||
if reply == "" {
|
||||
reply = phraser.Q(phraser.QueryFound, map[string]string{"text": selfDescription})
|
||||
}
|
||||
return reply, true
|
||||
}
|
||||
@@ -0,0 +1,96 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
|
||||
// TestSelfFloorClaimsAQuestionAboutHerAndNothingElse — the offline floor, which
|
||||
// is what answers with no embedder. Narrow on purpose, so the rows that must
|
||||
// NOT match are the point.
|
||||
func TestSelfFloorClaimsAQuestionAboutHerAndNothingElse(t *testing.T) {
|
||||
claimed := []string{
|
||||
"что ты умеешь",
|
||||
"что ты можешь",
|
||||
"а что ты умеешь?",
|
||||
"кто ты",
|
||||
"кто ты такая?",
|
||||
"расскажи о себе",
|
||||
"what can you do",
|
||||
"who are you",
|
||||
}
|
||||
for _, u := range claimed {
|
||||
if !selfFloor(u) {
|
||||
t.Errorf("%q is a question about her and the floor missed it", u)
|
||||
}
|
||||
}
|
||||
declined := []string{
|
||||
"что у меня сегодня",
|
||||
"расскажи про байкал",
|
||||
"кто изобрёл телефон",
|
||||
"что требует внимания",
|
||||
"запиши что я пил воду",
|
||||
// The floor spells its own word boundaries out, because Go's \b never
|
||||
// fires next to a Cyrillic letter. Without that these would match.
|
||||
"кто тыкал в розетку",
|
||||
"расскажи о себестоимости",
|
||||
}
|
||||
for _, u := range declined {
|
||||
if selfFloor(u) {
|
||||
t.Errorf("%q is not about her and the floor claimed it", u)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestSelfSourceAnswersFromTheDescription — with no embedder the source falls
|
||||
// to the floor, and the answer has to be the description rather than silence.
|
||||
func TestSelfSourceAnswersFromTheDescription(t *testing.T) {
|
||||
h := personalHandler()
|
||||
reply, claimed := h.querySelf(context.Background(), &queryTurn{
|
||||
dec: router.Decision{Utterance: "что ты умеешь"},
|
||||
})
|
||||
if !claimed {
|
||||
t.Fatal("a question about her must be claimed above the boundary")
|
||||
}
|
||||
if !strings.Contains(reply, "напоминания") {
|
||||
t.Errorf("the answer must come from the description: %q", reply)
|
||||
}
|
||||
if _, claimed := h.querySelf(context.Background(), &queryTurn{
|
||||
dec: router.Decision{Utterance: "почему небо синее"},
|
||||
}); claimed {
|
||||
t.Error("a world question must pass this source")
|
||||
}
|
||||
}
|
||||
|
||||
// TestSelfDescriptionHoldsThePersona — it is her own text and she reads it out,
|
||||
// so the same rules the phrasing eval enforces apply to it. Feminine
|
||||
// self-reference, informal address, no pet names.
|
||||
func TestSelfDescriptionHoldsThePersona(t *testing.T) {
|
||||
lower := strings.ToLower(selfDescription)
|
||||
for _, bad := range []string{"я рад ", "я готов ", "вы ", "ваш", "милый", "дорогой"} {
|
||||
if strings.Contains(lower, bad) {
|
||||
t.Errorf("the description breaks the persona on %q", bad)
|
||||
}
|
||||
}
|
||||
for _, want := range []string{"тво", "ты"} {
|
||||
if !strings.Contains(lower, want) {
|
||||
t.Errorf("the description must address him directly, missing %q", want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestSelfDescriptionClaimsNothingUnconditionally — the constraint that makes
|
||||
// this text safe to read out. Every capability that depends on config has to be
|
||||
// named as depending on it, and inventing one here is the same defect as
|
||||
// inventing a fact.
|
||||
func TestSelfDescriptionClaimsNothingUnconditionally(t *testing.T) {
|
||||
conditional := selfDescription[strings.Index(selfDescription, "Что зависит"):]
|
||||
for _, cap := range []string{"дом", "локальная сеть", "ленты", "список покупок", "погода", "телеграм"} {
|
||||
if !strings.Contains(conditional, cap) {
|
||||
t.Errorf("%q is configured, not wired — it must sit under the conditional half", cap)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -347,6 +347,55 @@ func (s *scriptedLLM) Complete(_ context.Context, r llm.Req) (string, error) {
|
||||
map[bool]string{true: "route", false: "reply"}[routing], truncateRunes(r.User, 60))
|
||||
}
|
||||
|
||||
// scriptedPhraser answers the chat path from the same script the router reads.
|
||||
//
|
||||
// It exists because actionChat calls h.phraser.PhraseChat, and the production
|
||||
// implementation posts raw HTTP to /v1/chat/completions rather than going
|
||||
// through the llm client scriptedLLM stands in for. So until this, no scenario
|
||||
// could script what she SAYS on a chat turn: the simulator wired phraser.NewStub()
|
||||
// and every chat reply came back as a pick from fallbacks_ru_v1.json, four
|
||||
// variants deep, which varied between two runs of one scenario (V-542 item 4).
|
||||
//
|
||||
// Everything except PhraseChat is the Stub's, by embedding. A nudge and a
|
||||
// reminder are phrased by the tick loop, which has its own phraser and its own
|
||||
// assertions; this seam is only about the conversation.
|
||||
type scriptedPhraser struct {
|
||||
*phraser.Stub
|
||||
entries []scriptEntry
|
||||
}
|
||||
|
||||
// PhraseChat returns the scripted reply for the utterance, or an error when the
|
||||
// scenario scripted none. The error rather than a fallback is deliberate and
|
||||
// matches scriptedLLM: actionChat logs it and falls back to ChatFallback(), so a
|
||||
// scenario that never meant to assert on a chat reply behaves exactly as it did
|
||||
// before, and one that DID means to is told its script has a hole.
|
||||
func (p *scriptedPhraser) PhraseChat(_ context.Context, utterance string, _ []dialogue.Turn) (string, error) {
|
||||
for _, e := range p.entries {
|
||||
if e.Reply == "" {
|
||||
continue
|
||||
}
|
||||
if e.Match != "" && !strings.Contains(strings.ToLower(utterance), strings.ToLower(e.Match)) {
|
||||
continue
|
||||
}
|
||||
return chatReplyText(e.Reply), nil
|
||||
}
|
||||
return "", fmt.Errorf("simulator: no scripted chat reply for %q", truncateRunes(utterance, 60))
|
||||
}
|
||||
|
||||
// chatReplyText reads a scripted reply in either shape the phrasing contract
|
||||
// allows: the {"response","mood"} object the model emits, or plain text.
|
||||
// LLMPhraser does this parse itself, so a scenario writes one thing and both
|
||||
// paths understand it.
|
||||
func chatReplyText(reply string) string {
|
||||
var out struct {
|
||||
Response string `json:"response"`
|
||||
}
|
||||
if err := json.Unmarshal([]byte(reply), &out); err == nil && out.Response != "" {
|
||||
return out.Response
|
||||
}
|
||||
return reply
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Building the world
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -440,7 +489,7 @@ func newSimWorld(t *testing.T, sc scenario) *simWorld {
|
||||
api: api,
|
||||
matcher: matcher,
|
||||
tools: tool.NewExecutor(api, 5*time.Second),
|
||||
phraser: phraser.NewStub(),
|
||||
phraser: &scriptedPhraser{Stub: phraser.NewStub(), entries: sc.Script},
|
||||
replier: newLLMReplier(scripted, nil),
|
||||
now: clock.Now,
|
||||
dataStore: st,
|
||||
|
||||
@@ -0,0 +1,79 @@
|
||||
{
|
||||
"schema_version": 1,
|
||||
"name": "conversation_anaphora",
|
||||
"description": "Five consecutive Russian turns about one object, replayed from the run that found V-542 on the box on 05-08-2026. He names a monitor, then asks four questions that all say \"он\" and never name it again.\n\nThis scenario exists because the shape had nowhere to fail. The routing fixture scores one utterance at a time, so a conversation that breaks on its second turn cannot lose a point there, and V-44 step 2 could only be verified by hand. That is item 3 of V-542.\n\nFour of the five replies below are WRONG, and the assertions pin them anyway. Read them as the recorded defect rather than the contract: she has the last four turns in front of her and never once names the thing he is asking about. Every wrong assertion is marked in its step note with what it must become. When V-542 lands, those flip and the ones marked correct do not move.\n\nWhat the four assert is that the reply LACKS \"монитор\". Absence is the defect itself: she is answering a question about a thing she wrote down two minutes ago and cannot name it. It also survives the fallback picker, which matters on the three query turns — they refuse from internal/phraser/fallbacks_ru_v1.json, four variants deep, and the same scenario returned \"тут я пас.\" one run and \"не знаю, честно.\" the next, so a string assertion there would pin the picker rather than the daemon.\n\nTurn 4 asserts its text as well, because that turn goes through the chat path and the chat path is now scriptable. scriptedPhraser in simulator_test.go answers PhraseChat from the same script entries the router reads (V-542 item 4); before it, the simulator wired phraser.NewStub() and no scenario could say what she SAYS on a chat turn at all.\n\nThe routes are scripted exactly as the box produced them, because the failure is not the model's. Turn 1 went to fact despite \"давай поболтаем\", every question after it went to query, and turn 4 went to chat. A scripted route is what lets this scenario pin the daemon's half without a llama-server in the loop.",
|
||||
"start": "2026-08-05T14:00:00+03:00",
|
||||
"script": [
|
||||
{
|
||||
"match": "купил новый монитор",
|
||||
"route": "[{\"intent\":\"fact\",\"key\":\"purchase\",\"value\":\"новый монитор\"}]",
|
||||
"reply": "{\"response\":\"записала: новый монитор.\",\"mood\":\"neutral\"}"
|
||||
},
|
||||
{
|
||||
"match": "он большой",
|
||||
"route": "[{\"intent\":\"query\",\"text\":\"а он большой?\"}]"
|
||||
},
|
||||
{
|
||||
"match": "сколько он примерно стоит",
|
||||
"route": "[{\"intent\":\"query\",\"text\":\"сколько он примерно стоит по-твоему?\"}]"
|
||||
},
|
||||
{
|
||||
"match": "переплатил",
|
||||
"route": "[{\"intent\":\"chat\",\"text\":\"мне кажется я переплатил\"}]",
|
||||
"reply": "{\"response\":\"я не знаю, о каком именно устройстве ты говоришь.\",\"mood\":\"neutral\"}"
|
||||
},
|
||||
{
|
||||
"match": "стоит его вернуть",
|
||||
"route": "[{\"intent\":\"query\",\"text\":\"стоит его вернуть?\"}]"
|
||||
},
|
||||
{
|
||||
"match": "",
|
||||
"route": "[{\"intent\":\"chat\",\"text\":\"\"}]",
|
||||
"reply": "{\"response\":\"я рада тебя слышать.\",\"mood\":\"happy\"}"
|
||||
}
|
||||
],
|
||||
"steps": [
|
||||
{
|
||||
"at": "14:00",
|
||||
"note": "CORRECT, and it is the first half of the defect. \"давай поболтаем\" is an explicit request to converse and the turn is filed as a fact anyway. Storing what he said is not wrong on its own — he did buy a monitor — but the object then lives in the fact store and never enters the transcript PhraseChat reads. That is V-542 decision 2: either the marker claims the turn at stage 0, or it means nothing and comes out of the fixture.",
|
||||
"say": "давай поболтаем: я вчера купил новый монитор",
|
||||
"expect_events": ["purchase"],
|
||||
"expect_no_send": true
|
||||
},
|
||||
{
|
||||
"at": "14:01",
|
||||
"note": "WRONG. \"он\" is the monitor from one turn ago, and she says she has no record of it. followUpMerge inherits prev.Slots.Key, and a query turn asking about a pronoun has no key to merge, so the question reaches the query sources naked and the notes source answers the only way it can. Must become: an answer about the monitor, or a route to chat where the transcript is.",
|
||||
"say": "а он большой?",
|
||||
"expect_reply_lacks": ["монитор"],
|
||||
"expect_no_send": true
|
||||
},
|
||||
{
|
||||
"at": "14:02",
|
||||
"note": "WRONG, and it rules out one explanation. This is not the previous turn failing to stick — it is the same wall a second time, two turns from where the monitor was named. Nothing accumulates across query turns.",
|
||||
"say": "сколько он примерно стоит по-твоему?",
|
||||
"expect_reply_lacks": ["монитор"],
|
||||
"expect_no_send": true
|
||||
},
|
||||
{
|
||||
"at": "14:03",
|
||||
"note": "WRONG, and it is the same wall from the other side. This turn routed chat, so it HAD the history that Session.History holds, and it asks which device he means anyway — because turn 1's object went to the fact store rather than the transcript. So a source reading the conversation is not sufficient on its own; decision 1 has to say which store the referent comes from. This is the one step whose text is pinned: the reply is scripted and reaches PhraseChat, so it is the box's own words rather than a fallback pick. Must become: a reply that names the monitor.",
|
||||
"say": "мне кажется я переплатил",
|
||||
"expect_reply_contains": ["о каком именно устройстве"],
|
||||
"expect_reply_lacks": ["монитор"],
|
||||
"expect_no_send": true
|
||||
},
|
||||
{
|
||||
"at": "14:04",
|
||||
"note": "WRONG. The fifth turn is the one that shows the cost. A returns question about a purchase two minutes old is answered with \"не нашла у тебя такой записи\", which is wrong in kind rather than merely unhelpful: the record exists, she wrote it herself at 14:00 under the key purchase.",
|
||||
"say": "стоит его вернуть?",
|
||||
"expect_reply_lacks": ["монитор"],
|
||||
"expect_no_send": true
|
||||
},
|
||||
{
|
||||
"at": "14:05",
|
||||
"note": "CORRECT, and it is the control. Nothing in five conversational turns was sent at him unprompted, and a tick with him mid-conversation stays silent. Whatever V-542 changes must not change this.",
|
||||
"tick": true,
|
||||
"expect_no_send": true
|
||||
}
|
||||
]
|
||||
}
|
||||
+2
-2
@@ -74,9 +74,9 @@
|
||||
},
|
||||
{
|
||||
"at": "08:50",
|
||||
"note": "he asks. The query path answers from local recall only: nothing stored clears the score gate, so she refuses rather than inventing a morning summary, and the replier is never reached. That refusal is the no-hallucination floor and this step pins it. Note what the persona check here is and is not: the reply is a constant in the Go source, so expect_reply_lacks pins that constant, not anything the model wrote. The step below is the one that reads model output.",
|
||||
"note": "he asks what he missed, and Praxis holds one unresolved item — the morning medicine — so she reads that back. This step pinned \"не знаю\" until 05-08-2026, and that was the keyword floor's blind spot rather than a rule: isAttentionQuery does not match \"что я пропустил\", while the topicAttend seeds carry \"что важное я пропустил\" almost verbatim. The seeds only started deciding when turnVector fixed the empty query vector every topic source was reading (V-547). Reading a surfaced item aloud is not inventing a morning summary, so the no-hallucination floor still holds; what moved is which source answers. Note what the persona check here is and is not: the reply is a constant in the Go source, so expect_reply_lacks pins that constant, not anything the model wrote. The step below is the one that reads model output.",
|
||||
"say": "что я пропустил?",
|
||||
"expect_reply_contains": ["не знаю"],
|
||||
"expect_reply_contains": ["требует внимания", "morning_medicine"],
|
||||
"expect_reply_lacks": ["рад ", "милый", "ваш"]
|
||||
},
|
||||
{
|
||||
|
||||
@@ -10,6 +10,7 @@ package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"log"
|
||||
"os"
|
||||
@@ -160,6 +161,7 @@ func (t *tickLoop) tick(ctx context.Context, now time.Time) {
|
||||
log.Printf("tick: gather: %v", err)
|
||||
return
|
||||
}
|
||||
t.savePresence(ctx, state, now)
|
||||
|
||||
// proactive: at most one candidate, max severity.
|
||||
cand, trace := loop.ExplainTick(state, t.rules)
|
||||
@@ -249,6 +251,7 @@ func (t *tickLoop) tick(ctx context.Context, now time.Time) {
|
||||
return
|
||||
}
|
||||
keys = t.repeatableRules(keys)
|
||||
keys = t.stopFinishedAlarms(ctx, keys, state, now)
|
||||
if len(keys) == 0 {
|
||||
return
|
||||
}
|
||||
@@ -260,6 +263,108 @@ func (t *tickLoop) tick(ctx context.Context, now time.Time) {
|
||||
}
|
||||
}
|
||||
|
||||
// savePresence writes back the bucket GatherState just resolved.
|
||||
//
|
||||
// It lives here and not in GatherState because that method holds a read-only
|
||||
// transaction on purpose — one consistent snapshot per tick — and a write
|
||||
// inside it would either break that guarantee or quietly upgrade the
|
||||
// transaction. The tick is the layer that already owns writes.
|
||||
//
|
||||
// Nothing wrote this row before (Vikunja #532), and the row is the whole
|
||||
// mechanism, so two things were broken at once. Hysteresis was dead: lastBucket
|
||||
// read the cold-start Away on every tick, so store.Resolve only ever took the
|
||||
// `last == Away` arm and demanded a full PresenceEnter score to say he is
|
||||
// there. The 0.30-0.55 hold band the function exists to provide never applied
|
||||
// once. And every readout lied: /dash and ipc.Presence read this row, so they
|
||||
// showed "away — score 0.00 (never)" while desk_active facts were arriving
|
||||
// every sixty seconds.
|
||||
//
|
||||
// A failure logs and the tick continues. The gate reads the in-memory bucket,
|
||||
// which is why nudge routing kept working through all of this — losing the
|
||||
// write costs the next tick's hysteresis, not this tick's decisions.
|
||||
func (t *tickLoop) savePresence(ctx context.Context, state loop.State, now time.Time) {
|
||||
if err := t.store.SavePresenceState(ctx, state.Presence, state.PresenceScore, now); err != nil {
|
||||
log.Printf("tick: save presence state: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// maxAlarmAge — how long one un-acked telegram alarm may keep repeating.
|
||||
//
|
||||
// This is the floor brake and it applies to every rule, including one that
|
||||
// says nothing about its own condition (Vikunja #535). Nothing in the tree can
|
||||
// ack a telegram nudge: MarkAcked has no caller outside internal/store, and the
|
||||
// only ack that exists is a voice "готово" on a box that runs no voice loop. So
|
||||
// "repeat until acked" meant "repeat forever", and it did — every five minutes
|
||||
// for over two hours.
|
||||
//
|
||||
// Two hours at the five-minute default is about 24 messages, which is already
|
||||
// past the point of being read. An alarm nobody answered in two hours is not
|
||||
// one more repeat away from being answered, and the right move is to stop
|
||||
// talking, not to talk louder.
|
||||
const maxAlarmAge = 2 * time.Hour
|
||||
|
||||
// stopFinishedAlarms returns the keys that may still repeat, and closes the
|
||||
// rest.
|
||||
//
|
||||
// Two ways an alarm ends without him. The condition cleared, which the rule
|
||||
// answers through StillTrue — deliberately NOT Predicate, which is
|
||||
// edge-triggered and reads false one tick after the alarm is raised, so using
|
||||
// it would cancel every alarm immediately. Or the alarm simply got old, which
|
||||
// is the bound that does not need the rule's cooperation.
|
||||
//
|
||||
// A rule with no StillTrue is not treated as resolved. Silence about the
|
||||
// condition is not evidence the condition cleared, so those keys only ever stop
|
||||
// on age.
|
||||
func (t *tickLoop) stopFinishedAlarms(ctx context.Context, keys []string, state loop.State, now time.Time) []string {
|
||||
if len(keys) == 0 {
|
||||
return nil
|
||||
}
|
||||
byName := make(map[string]loop.Rule, len(t.rules))
|
||||
for _, r := range t.rules {
|
||||
byName[r.Name] = r
|
||||
}
|
||||
live := keys[:0:0]
|
||||
for _, key := range keys {
|
||||
outcome := ""
|
||||
switch r := byName[key]; {
|
||||
case r.StillTrue != nil && !r.StillTrue(state):
|
||||
outcome = store.NudgeResolved
|
||||
case t.alarmIsOlderThan(ctx, key, maxAlarmAge, now):
|
||||
// Not "resolved": nothing says the thing got better. This is her
|
||||
// giving up on being answered, and /notifications should say so.
|
||||
outcome = store.NudgeIgnored
|
||||
}
|
||||
if outcome == "" {
|
||||
live = append(live, key)
|
||||
continue
|
||||
}
|
||||
n, err := t.store.ResolvePendingTelegram(ctx, key, outcome, now)
|
||||
if err != nil {
|
||||
// Could not close it, so do not drop it either: repeating is the
|
||||
// lesser fault against losing the alarm entirely.
|
||||
log.Printf("tick: stop alarm %s: %v", key, err)
|
||||
live = append(live, key)
|
||||
continue
|
||||
}
|
||||
log.Printf("tick: alarm %s ended (%s), %d pending nudge(s) closed", key, outcome, n)
|
||||
}
|
||||
return live
|
||||
}
|
||||
|
||||
// alarmIsOlderThan reports whether the oldest un-acked send for this rule is
|
||||
// past the cap. A read failure answers false: an alarm that repeats one more
|
||||
// time is better than one silenced by a transient store error.
|
||||
func (t *tickLoop) alarmIsOlderThan(ctx context.Context, rule string, age time.Duration, now time.Time) bool {
|
||||
oldest, err := t.store.OldestPendingTelegram(ctx, rule)
|
||||
if err != nil {
|
||||
if !errors.Is(err, store.ErrNudgeNotFound) {
|
||||
log.Printf("tick: oldest pending %s: %v", rule, err)
|
||||
}
|
||||
return false
|
||||
}
|
||||
return now.Sub(oldest) >= age
|
||||
}
|
||||
|
||||
// repeatableRules drops keys whose rule is not wired any more.
|
||||
//
|
||||
// The repeat path reads the nudges table, not the rule set: any sev4 telegram
|
||||
|
||||
+74
-4
@@ -10,6 +10,7 @@ import (
|
||||
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/loop"
|
||||
"github.com/kami/maven/internal/store"
|
||||
)
|
||||
|
||||
// daemonAPI wraps a store-backed CoreAPI and overrides TickTrace with the
|
||||
@@ -22,6 +23,17 @@ type daemonAPI struct {
|
||||
chatFn func(ctx context.Context, conversation, text string) string
|
||||
getMCPServers func() []ipc.MCPServerStatus
|
||||
getEvents func(n int) []ipc.IntakeEvent
|
||||
getDecisions func(n int) []ipc.TurnDecision
|
||||
// nexus — the identity client, nil when no nexus block is configured. It
|
||||
// is what makes ResolveEntity answerable at all; without it the store
|
||||
// adapter's refusal stands, and a surface that wanted an entity id says so
|
||||
// instead of storing a name.
|
||||
nexus *nexusClient
|
||||
// seedStore — non-nil ONLY when mavend was started with -allow-seed. It is
|
||||
// the whole off-switch for the backdated write path (Vikunja #518), and it
|
||||
// is a store rather than a bool so that leaving the flag off means the
|
||||
// method has nothing to write with, not merely permission to refuse.
|
||||
seedStore *store.Store
|
||||
}
|
||||
|
||||
// RecentEvents — the unified intake journal (Vikunja #283). Empty, not an
|
||||
@@ -35,11 +47,58 @@ func (d *daemonAPI) RecentEvents(ctx context.Context, n int) ([]ipc.IntakeEvent,
|
||||
return d.getEvents(n), nil
|
||||
}
|
||||
|
||||
func (d *daemonAPI) Chat(ctx context.Context, conversation, text string) (string, error) {
|
||||
if d.chatFn == nil {
|
||||
return "", errors.New("mavend: chat not available")
|
||||
// nexusOf — the identity client the voice wiring built, or nil. Same shape as
|
||||
// embedderOf: a wiring that is absent and a wiring with no nexus block are one
|
||||
// answer here.
|
||||
func nexusOf(w *voiceWiring) *nexusClient {
|
||||
if w == nil || w.handler == nil || w.handler.ecosystem == nil {
|
||||
return nil
|
||||
}
|
||||
return d.chatFn(ctx, conversation, text), nil
|
||||
return w.handler.ecosystem.nexus
|
||||
}
|
||||
|
||||
// ResolveEntity asks Nexus for the canonical id behind a name (Vikunja #511).
|
||||
//
|
||||
// Three outcomes, kept apart on purpose. No nexus block is ErrNotImplemented,
|
||||
// so a surface can say "identity is not configured here" rather than invent an
|
||||
// id. A miss is ipc.ErrNoEntity. A match against several entities comes back
|
||||
// Ambiguous with the names, because picking one is how a task ends up blocked
|
||||
// on the wrong person and nobody can see it happened.
|
||||
func (d *daemonAPI) ResolveEntity(ctx context.Context, query string, types []string) (ipc.EntityRef, error) {
|
||||
if d.nexus == nil {
|
||||
return ipc.EntityRef{}, ipc.ErrNotImplemented
|
||||
}
|
||||
res, err := d.nexus.Resolve(ctx, query, types)
|
||||
if err != nil {
|
||||
return ipc.EntityRef{}, err
|
||||
}
|
||||
if len(res.Candidates) > 1 {
|
||||
names := make([]string, 0, len(res.Candidates))
|
||||
for _, c := range res.Candidates {
|
||||
names = append(names, c.DisplayName)
|
||||
}
|
||||
return ipc.EntityRef{Ambiguous: true, Candidates: names}, nil
|
||||
}
|
||||
if res.Entity == nil || res.Entity.ID == "" {
|
||||
return ipc.EntityRef{}, ipc.ErrNoEntity
|
||||
}
|
||||
return ipc.EntityRef{
|
||||
ID: res.Entity.ID,
|
||||
Type: res.Entity.Type,
|
||||
DisplayName: res.Entity.DisplayName,
|
||||
}, nil
|
||||
}
|
||||
|
||||
// Chat runs one text turn and reports which query source claimed it. The sink
|
||||
// rides the context so handleText keeps the one string signature the mic,
|
||||
// telegram and the web all call it through (V-539).
|
||||
func (d *daemonAPI) Chat(ctx context.Context, conversation, text string) (ipc.ChatReply, error) {
|
||||
if d.chatFn == nil {
|
||||
return ipc.ChatReply{}, errors.New("mavend: chat not available")
|
||||
}
|
||||
ctx, sink := withQuerySourceSink(ctx)
|
||||
reply := d.chatFn(ctx, conversation, text)
|
||||
return ipc.ChatReply{Reply: reply, Source: sink.Name()}, nil
|
||||
}
|
||||
|
||||
// MCPServers — the configured MCP servers and their health (Vikunja #251).
|
||||
@@ -60,6 +119,17 @@ func (d *daemonAPI) TickTrace(ctx context.Context) (ipc.TickTrace, error) {
|
||||
return toIPCTickTrace(*trace), nil
|
||||
}
|
||||
|
||||
// TurnDecisions — the arbitration records of the last few turns (V-564). Nil
|
||||
// getter means voice was never wired, and that is an empty list rather than an
|
||||
// error: a box with no voice path has had no turns to arbitrate, which is not a
|
||||
// fault and renders as an empty table.
|
||||
func (d *daemonAPI) TurnDecisions(ctx context.Context, n int) ([]ipc.TurnDecision, error) {
|
||||
if d.getDecisions == nil {
|
||||
return nil, nil
|
||||
}
|
||||
return d.getDecisions(n), nil
|
||||
}
|
||||
|
||||
func (d *daemonAPI) MorningStatus(ctx context.Context) ([]ipc.MorningRoutineStatus, error) {
|
||||
if d.getMorningStatus == nil {
|
||||
return nil, errors.New("mavend: morning status not available")
|
||||
|
||||
@@ -584,7 +584,9 @@ func TestDigestSev4BypassesQueue(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
now := refNow()
|
||||
markPresent(t, st, ctx, now)
|
||||
if _, err := st.SetValue(ctx, store.KindSelf, "service_down:db", "poll:uptimekuma", "down", now); err != nil {
|
||||
// Older than loop.MinDownAge, so this tests the digest bypass and not the
|
||||
// flap debounce (Vikunja #536).
|
||||
if _, err := st.SetValue(ctx, store.KindSelf, "service_down:db", "poll:uptimekuma", "down", now.Add(-5*time.Minute)); err != nil {
|
||||
t.Fatalf("seed service_down: %v", err)
|
||||
}
|
||||
sink := &fakeSink{}
|
||||
|
||||
+121
-4
@@ -9,7 +9,7 @@ import (
|
||||
)
|
||||
|
||||
// Which subject is this question about — the weather, the house, the LAN, what
|
||||
// needs looking at, or none of them. Third of the three mechanisms replacing hand-written Russian
|
||||
// needs looking at, his feeds, or none of them. Third of the three mechanisms replacing hand-written Russian
|
||||
// patterns (Vikunja #522, owner's call 2026-08-04). internal/lexicon holds the
|
||||
// sets that can be finished and internal/morph answers the grammar questions;
|
||||
// this is for the sets that can never be finished, because "is this about the
|
||||
@@ -37,6 +37,11 @@ import (
|
||||
// внимания", which is Praxis's operational state and reached the web search
|
||||
// before the source existed (Vikunja #475).
|
||||
//
|
||||
// A fifth joined on 05-08-2026: the feeds, "что нового в лентах". Its word lists
|
||||
// were the last pair of hand-written Russian stem lists in the router (V-522),
|
||||
// and they carried the same admission in their own comments — vagueNouns exists
|
||||
// because "что нового?" is a greeting that matched a feed noun.
|
||||
//
|
||||
// The regexes stay as the offline floor, unchanged, for a handler with no
|
||||
// embedder or a turn whose vector never got computed. They are allowed to remain
|
||||
// narrow now precisely because they are no longer the only answer.
|
||||
@@ -52,6 +57,9 @@ const (
|
||||
topicHome topicLabel = "home"
|
||||
topicNetwork topicLabel = "network"
|
||||
topicAttend topicLabel = "attention"
|
||||
topicFeed topicLabel = "feeds"
|
||||
topicList topicLabel = "list"
|
||||
topicSelf topicLabel = "self"
|
||||
topicOther topicLabel = "other"
|
||||
)
|
||||
|
||||
@@ -117,7 +125,56 @@ var topicSeedSets = map[topicLabel][]string{
|
||||
"what needs attention",
|
||||
"what needs looking at right now",
|
||||
},
|
||||
topicFeed: {
|
||||
"что нового в лентах",
|
||||
"какие новости",
|
||||
"что нового по технологиям",
|
||||
"почитай заголовки",
|
||||
// Two seeds carrying a day word beside the headlines. Without them
|
||||
// "какие сегодня заголовки" read as weather, because "какая сегодня
|
||||
// погода" is the nearest thing in the whole set with "сегодня" in it.
|
||||
"заголовки за сегодня",
|
||||
"какие главные новости за день",
|
||||
"покажи новости за сегодня",
|
||||
"что пишут в новостях",
|
||||
"что нового про политику",
|
||||
"расскажи что нового в ленте",
|
||||
"what is new in the feeds",
|
||||
"any news headlines today",
|
||||
},
|
||||
// Reading a standing list back, and only that. Adding to one and clearing
|
||||
// one stay on the phrase tables in internal/router/list.go — see its header
|
||||
// for why a span and a delete are not seed-shaped work.
|
||||
topicList: {
|
||||
"что в списке покупок",
|
||||
"что мне нужно купить",
|
||||
"прочитай список покупок",
|
||||
"покажи что в списке",
|
||||
"что осталось купить в магазине",
|
||||
"что мне нужно в аптеке",
|
||||
"какой у меня список покупок",
|
||||
"what is on my shopping list",
|
||||
"read me the grocery list",
|
||||
},
|
||||
// Questions about her (Vikunja #555). The set lives in self.go beside the
|
||||
// description it unlocks, so the two are edited together — a seed claiming
|
||||
// a question the description does not answer is the failure mode.
|
||||
//
|
||||
// "что ты умеешь" was a topicOther seed until this existed, put there so an
|
||||
// attention question had something to lose to. It is a self seed now, and
|
||||
// it cannot be both: a phrasing on two sides never clears the margin.
|
||||
topicSelf: selfSeeds,
|
||||
topicOther: {
|
||||
// A task question is not a list read-back. They collide on "что у меня",
|
||||
// and the list has its own table to lose to as well.
|
||||
"какие у меня задачи",
|
||||
"что у меня в делах",
|
||||
// The bare newness opener, which is a greeting and not a request for
|
||||
// headlines. It sits here on purpose: it is close enough to the feed
|
||||
// seeds that it will not clear topicMargin, and a thin call goes to
|
||||
// ParseFeedQuery, which declines a vague noun with no topic beside it.
|
||||
"что нового",
|
||||
"как дела",
|
||||
// Complaints, which are not requests to scan or to read the house.
|
||||
// isNetworkQuery's comment names this one: a scan she runs unasked is
|
||||
// the noisy behaviour the bounds exist to prevent.
|
||||
@@ -137,9 +194,40 @@ var topicSeedSets = map[topicLabel][]string{
|
||||
"что я говорил про бэкапы",
|
||||
"что у меня сегодня по календарю",
|
||||
"напомни мне позвонить маме",
|
||||
// An attention question is about the state of his things; this is not.
|
||||
"что ты умеешь",
|
||||
"what did i say about backups",
|
||||
// World questions that name a day (Vikunja #553). Weather was the only
|
||||
// topic whose seeds carry a day word — four of its eight do — so every
|
||||
// "какой сегодня X" landed nearest it and cleared the margin: the
|
||||
// dollar rate by 0.0220 and a public holiday by 0.0398, against 0.0883
|
||||
// for a real weather question. The gate then asked "для какого города?"
|
||||
// about the dollar.
|
||||
//
|
||||
// The margin was not the knob. 0.0398 is not a coin flip, and raising
|
||||
// the bar far enough to catch it would take real weather questions with
|
||||
// it. What was missing is the negative class: a day word means the
|
||||
// question is about a day, and says nothing about whether it is about
|
||||
// the sky.
|
||||
"сколько стоит биткоин сегодня",
|
||||
"какой завтра праздник в стране",
|
||||
"во сколько сегодня восход солнца",
|
||||
"кто вчера победил в чемпионате",
|
||||
// The frame itself, twice. "какая сегодня погода" is a weather seed,
|
||||
// and the four above did not move "какой сегодня курс доллара" or
|
||||
// "что интересного произошло сегодня в мире" off weather, because what
|
||||
// pulls them is the frame and not the noun. A frame that both topics
|
||||
// use has to sit on both sides, or the side that owns it wins every
|
||||
// noun it has never seen.
|
||||
"какой сегодня курс валют",
|
||||
"что сегодня происходит в мире",
|
||||
// The same story one topic over, found while verifying V-554 on the
|
||||
// box: "кто изобрёл телефон" ran a LAN scan and answered "нашла 3
|
||||
// устройства". The network set opens with "кто в сети сейчас" and
|
||||
// names devices throughout, so a "кто ..." question about any device
|
||||
// noun landed there. A device has a history, and asking about it is
|
||||
// not asking what is plugged in.
|
||||
"кто изобрёл телефон",
|
||||
"когда появился первый компьютер",
|
||||
"как работает роутер",
|
||||
},
|
||||
}
|
||||
|
||||
@@ -201,6 +289,35 @@ func (x *topicIndex) best(vec []float32) (label topicLabel, margin float64, ok b
|
||||
return label, first - second, true
|
||||
}
|
||||
|
||||
// turnVector returns the turn's query vector, computing it on first ask and
|
||||
// caching it on the turn.
|
||||
//
|
||||
// It exists because every topic source sits ABOVE the "embed" source in
|
||||
// querySources, and that source was the only thing that ever set t.vec. So
|
||||
// turnIsAbout was reading an empty vector on every deployed turn, best returned
|
||||
// ok=false, and all six recognisers ran on their keyword floors — the seeds
|
||||
// decided nothing outside the tests, which embed the utterance themselves and
|
||||
// call best directly. Found on the box on 05-08-2026: "что мне нужно купить" was
|
||||
// answered from an old note, and the seeds place it as the list by 0.0841.
|
||||
//
|
||||
// Computing here rather than moving the embed source up: the cost is paid by the
|
||||
// turns that ask, the cache means queryEmbed below reuses this one, and the
|
||||
// order of querySources stays what its comments argue for.
|
||||
func (h *reactiveHandler) turnVector(ctx context.Context, t *queryTurn) []float32 {
|
||||
if len(t.vec) > 0 || h.recall.embedder == nil {
|
||||
return t.vec
|
||||
}
|
||||
vec, err := router.EmbedQuery(ctx, h.recall.embedder, t.dec.Utterance)
|
||||
if err != nil {
|
||||
// The floor answers. A topic source is not the place to fail a turn:
|
||||
// the recall sources below hit the same embedder and report it there.
|
||||
log.Printf("voice: topic vector for %q: %v", t.dec.Utterance, err)
|
||||
return nil
|
||||
}
|
||||
t.vec = vec
|
||||
return vec
|
||||
}
|
||||
|
||||
// turnIsAbout — the recogniser every topic source calls. The seeds decide when
|
||||
// the embedder is there, which is every deployed box; floor is the source's own
|
||||
// keyword test, which answers when they are not.
|
||||
@@ -213,7 +330,7 @@ func (x *topicIndex) best(vec []float32) (label topicLabel, margin float64, ok b
|
||||
// network by 0.0055; isNetworkQuery says no, so it stays the complaint it is.
|
||||
func (h *reactiveHandler) turnIsAbout(ctx context.Context, t *queryTurn, want topicLabel, floor func(string) bool) bool {
|
||||
h.recall.topics.load(ctx, h.recall.embedder)
|
||||
label, margin, ok := h.recall.topics.best(t.vec)
|
||||
label, margin, ok := h.recall.topics.best(h.turnVector(ctx, t))
|
||||
if !ok {
|
||||
return floor(t.dec.Utterance)
|
||||
}
|
||||
|
||||
@@ -25,6 +25,8 @@ func TestTopicFloorAnswersWithoutSeeds(t *testing.T) {
|
||||
{"что включено в доме?", topicHome, isHomeQuery, true},
|
||||
{"какие устройства в сети?", topicNetwork, isNetworkQuery, true},
|
||||
{"что требует внимания?", topicAttend, isAttentionQuery, true},
|
||||
{"что нового в лентах?", topicFeed, feedFloor, true},
|
||||
{"что в списке покупок?", topicList, listFloor, true},
|
||||
{"почему небо синее", topicWeather, isWeatherQuery, false},
|
||||
{"я дома", topicHome, isHomeQuery, false},
|
||||
{"интернет не работает", topicNetwork, isNetworkQuery, false},
|
||||
@@ -86,6 +88,50 @@ func TestONNXTopics(t *testing.T) {
|
||||
{"что требует моего внимания сейчас", topicAttend, isAttentionQuery},
|
||||
{"что не так с базой данных", topicAttend, isAttentionQuery},
|
||||
{"есть что-то срочное на сегодня", topicAttend, isAttentionQuery},
|
||||
{"что нового в ленте за сегодня", topicFeed, feedFloor},
|
||||
{"какие сегодня заголовки", topicFeed, feedFloor},
|
||||
{"что нового про искусственный интеллект", topicFeed, feedFloor},
|
||||
// The greeting. It has to lose to topicOther, or fall thin enough that
|
||||
// ParseFeedQuery — which declines a vague noun with no topic — answers.
|
||||
{"что нового?", topicOther, feedFloor},
|
||||
{"что мне надо купить в магазине", topicList, listFloor},
|
||||
{"прочитай мне список", topicList, listFloor},
|
||||
{"что там в аптеке нужно взять", topicList, listFloor},
|
||||
// A task read-back is not a list read-back, and the two collide on
|
||||
// "что у меня".
|
||||
{"какие у меня сейчас задачи", topicOther, listFloor},
|
||||
// World questions that name a day (Vikunja #553). Weather was the only
|
||||
// topic carrying day words, so all of these read as weather and two of
|
||||
// them cleared the margin: the gate asked "для какого города?" about
|
||||
// the dollar. The last two are far from any seed on purpose — the
|
||||
// first three are close enough to the new topicOther seeds that they
|
||||
// would pass on similarity alone.
|
||||
{"какой сегодня курс доллара", topicOther, isWeatherQuery},
|
||||
{"какой сегодня праздник", topicOther, isWeatherQuery},
|
||||
{"что интересного произошло сегодня в мире", topicOther, isWeatherQuery},
|
||||
{"во сколько завтра открывается музей", topicOther, isWeatherQuery},
|
||||
{"кто сегодня играет в лиге чемпионов", topicOther, isWeatherQuery},
|
||||
// The control the seeds above must not cost: real weather still reads
|
||||
// as weather, including the two that lean on the keyword floor.
|
||||
{"будет ли завтра дождь в москве", topicWeather, isWeatherQuery},
|
||||
{"какая температура завтра утром", topicWeather, isWeatherQuery},
|
||||
// The same shape one topic over, seen on the box (Vikunja #554): a
|
||||
// device has a history, and asking about it is not asking what is
|
||||
// plugged in. "кто изобрёл телефон" answered "нашла 3 устройства".
|
||||
// Held out from the seeds, which name the telephone and the computer.
|
||||
{"кто придумал радио", topicOther, isNetworkQuery},
|
||||
{"когда изобрели телевизор", topicOther, isNetworkQuery},
|
||||
{"как устроен телефон внутри", topicOther, isNetworkQuery},
|
||||
// The control: a real scan is still a scan.
|
||||
{"какие устройства подключены к вайфаю", topicNetwork, isNetworkQuery},
|
||||
// Questions about her (Vikunja #555), held out from selfSeeds.
|
||||
{"а что ты вообще умеешь делать", topicSelf, selfFloor},
|
||||
{"какие у тебя навыки", topicSelf, selfFloor},
|
||||
{"расскажи мне о себе", topicSelf, selfFloor},
|
||||
{"what are you able to do", topicSelf, selfFloor},
|
||||
// The control: an attention question is about the state of his things,
|
||||
// and it is the neighbour these seeds could have taken.
|
||||
{"что требует внимания у меня в сервисах", topicAttend, isAttentionQuery},
|
||||
}
|
||||
|
||||
h := &reactiveHandler{recall: recallWiring{embedder: emb}}
|
||||
|
||||
@@ -0,0 +1,251 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"unicode"
|
||||
|
||||
"github.com/kami/maven/internal/dialogue"
|
||||
"github.com/kami/maven/internal/lexicon"
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
|
||||
// turnRole — what this utterance IS, relative to the action Maven is in the
|
||||
// middle of assembling. The five roles are the owner's vocabulary (Vikunja
|
||||
// #558), plus the sixth answer a resolver is allowed to give: not_applicable,
|
||||
// which hands the turn back to generic dispatch.
|
||||
//
|
||||
// It exists because the arbitration used to be ordering. The clarify resolver,
|
||||
// the confirm gate, the follow-up merge and the repair marker all ran BEFORE
|
||||
// the router, so the claimant holding conversational state decided what an
|
||||
// utterance was without asking the one component whose job that is — and on
|
||||
// 2026-08-05 "какая сейчас погода в Риме?" became the time of a reminder,
|
||||
// because the extractor found "сейчас" in it and nothing looked at the rest.
|
||||
//
|
||||
// The rule that fixes that: a routed decision which stands on its own — its own
|
||||
// intent, its own slots filled from its own words — is not an answer, whatever
|
||||
// the extractor found inside it.
|
||||
type turnRole string
|
||||
|
||||
const (
|
||||
roleAnswer turnRole = "answer" // it fills the slot she asked about
|
||||
roleCorrection turnRole = "correction" // it replaces a value she already had
|
||||
roleSideQuery turnRole = "side_query" // a question of its own, asked mid-flow
|
||||
roleNewRequest turnRole = "new_request" // a different request entirely
|
||||
roleCancel turnRole = "cancel" // call the pending action off
|
||||
roleNotApplicable turnRole = "not_applicable" // nothing is pending; not our turn
|
||||
)
|
||||
|
||||
// frameWords — the words that can stand around a bare slot value without adding
|
||||
// a request. Every member is a closed class from internal/lexicon: the frame
|
||||
// itself, the interrogatives, the parts of a spoken clock, the days and the
|
||||
// months. Assembled once; the sets are copies, so this cannot edit them.
|
||||
var frameWords = buildFrameWords()
|
||||
|
||||
func buildFrameWords() map[string]bool {
|
||||
out := make(map[string]bool)
|
||||
add := func(list []string) {
|
||||
for _, w := range list {
|
||||
out[strings.ToLower(w)] = true
|
||||
}
|
||||
}
|
||||
add(lexicon.SlotValueFrame())
|
||||
add(lexicon.Interrogatives())
|
||||
add(lexicon.PartsOfDay())
|
||||
add(lexicon.HalfHourWords())
|
||||
add(lexicon.DayOffsetWords())
|
||||
for i := 0; i < 7; i++ {
|
||||
out[lexicon.Weekday(i)] = true
|
||||
}
|
||||
for m := 1; m <= 12; m++ {
|
||||
out[lexicon.MonthGenitive(m)] = true
|
||||
}
|
||||
for hh := 0; hh <= 23; hh++ {
|
||||
add(strings.Fields(lexicon.HourSpoken(hh)))
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// cancelWords — the same, for the words that call the pending action off.
|
||||
var cancelWords = buildCancelWords()
|
||||
|
||||
func buildCancelWords() map[string]bool {
|
||||
out := make(map[string]bool)
|
||||
for _, w := range lexicon.DialogueCancel() {
|
||||
out[strings.ToLower(w)] = true
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// turnTokens splits an utterance the way the router's own predicates do: over
|
||||
// letters and digits, lowercased, so punctuation and a clock's colon fall out.
|
||||
func turnTokens(text string) []string {
|
||||
return strings.FieldsFunc(strings.ToLower(text), func(r rune) bool {
|
||||
return !unicode.IsLetter(r) && !unicode.IsDigit(r)
|
||||
})
|
||||
}
|
||||
|
||||
// ownContent lists the tokens of an utterance that are neither frame nor value:
|
||||
// what it is about, over and above the thing she asked for. Numbers go out
|
||||
// because a number is the commonest slot value there is, and the closed number
|
||||
// and day lexicons go out with them.
|
||||
//
|
||||
// Empty ⇒ the utterance is a slot value and nothing else, however it is dressed
|
||||
// up. That is the whole test, and it is what separates "а что если в 11:00" —
|
||||
// which is a time, hedged — from "какая сейчас погода в Риме?", which leaves
|
||||
// "погода" and "риме" behind and is therefore about something.
|
||||
func ownContent(text string) []string {
|
||||
var out []string
|
||||
for _, tok := range turnTokens(text) {
|
||||
if frameWords[tok] || lexicon.IsFillerParticle(tok) {
|
||||
continue
|
||||
}
|
||||
if _, ok := lexicon.Cardinal(tok); ok {
|
||||
continue
|
||||
}
|
||||
if _, ok := lexicon.Ordinal(tok); ok {
|
||||
continue
|
||||
}
|
||||
if isNumeric(tok) {
|
||||
continue
|
||||
}
|
||||
out = append(out, tok)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func isNumeric(tok string) bool {
|
||||
for _, r := range tok {
|
||||
if !unicode.IsDigit(r) {
|
||||
return false
|
||||
}
|
||||
}
|
||||
return tok != ""
|
||||
}
|
||||
|
||||
// isCancel reports whether the utterance is nothing but a call-off. Every
|
||||
// content token has to be a cancel word, so "забудь" ends the exchange and
|
||||
// "забудь купить молоко" does not.
|
||||
func isCancel(text string) bool {
|
||||
content := ownContent(text)
|
||||
if len(content) == 0 {
|
||||
return false
|
||||
}
|
||||
for _, tok := range content {
|
||||
if !cancelWords[tok] {
|
||||
return false
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// carriesOwnRequest reads the ROUTED decision for the thing that decides this:
|
||||
// does the utterance ask for something in its own right? Each intent is asked
|
||||
// the question in its own terms, because "its own slots filled from its own
|
||||
// words" means a different field for each of them.
|
||||
//
|
||||
// A query or a system question needs no further evidence — the router already
|
||||
// read a question in these words. The write intents need the verb or the slot
|
||||
// that names the request, so a bare value the router guessed a home for does
|
||||
// not count as one.
|
||||
func carriesOwnRequest(dec router.Decision, text string) bool {
|
||||
if dec.Clarify {
|
||||
// The router itself was unsure. An utterance she could not route is
|
||||
// not an utterance that outranks the question in front of it.
|
||||
return false
|
||||
}
|
||||
switch dec.Intent {
|
||||
case router.IntentQuery, router.IntentSystem:
|
||||
return true
|
||||
case router.IntentReminder:
|
||||
return carriesReminderVerb(text)
|
||||
case router.IntentFact, router.IntentNote:
|
||||
return router.CarriesCaptureVerb(text)
|
||||
case router.IntentAct:
|
||||
// An act that resolved to a capability is a command. One that did not
|
||||
// is words she cannot execute anyway, so it stays an answer and gets
|
||||
// re-asked — the same thing that happens to it today.
|
||||
return dec.Slots.HasFn
|
||||
default: // chat
|
||||
return false
|
||||
}
|
||||
}
|
||||
|
||||
// carriesReminderVerb — "напомни" and its forms, matched over tokens. The
|
||||
// reminder verbs are a closed lexicon and are not capture verbs, so
|
||||
// CarriesCaptureVerb never sees them.
|
||||
func carriesReminderVerb(text string) bool {
|
||||
toks := turnTokens(text)
|
||||
for _, v := range lexicon.ReminderVerbs() {
|
||||
for _, t := range toks {
|
||||
if t == strings.ToLower(v) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// offlineOwnRequest is the shape half of the evidence: the offline token tests,
|
||||
// which cost nothing and never depend on the model that produced the routing.
|
||||
// It is also the whole answer when there is no route to read — the classifier
|
||||
// is the failure floor and a turn must never break on the model.
|
||||
func offlineOwnRequest(text string) bool {
|
||||
return router.IsQuestionShaped(text) || router.CarriesCaptureVerb(text) || carriesReminderVerb(text)
|
||||
}
|
||||
|
||||
// classifyTurnRole decides what this utterance is against the pending action.
|
||||
//
|
||||
// The fast path is the first line and it is a fast path to the SAME answer, not
|
||||
// a second decision procedure: an utterance with no content of its own can
|
||||
// never be a request of its own, so it can never be anything but an answer, and
|
||||
// the route below would spend a second on the resident model to say so. Every
|
||||
// other utterance is routed first, and the role is read off the decision.
|
||||
//
|
||||
// `routed` is the turn's routing, already computed; ok is false when there was
|
||||
// none to compute (no router wired, or the route failed). A failed route falls
|
||||
// to the offline shape tests rather than breaking the turn.
|
||||
func classifyTurnRole(q *dialogue.PendingQuestion, text string, answer dialogue.Slots, routed router.Decision, ok bool) turnRole {
|
||||
if isCancel(text) {
|
||||
return roleCancel
|
||||
}
|
||||
// Two pieces of evidence, and the content gate in front of both. The shape
|
||||
// tests are the floor and answer for free; the route is what sees a request
|
||||
// with no shape to it — "погода в риме" asks a question and carries neither
|
||||
// a question mark nor an interrogative, and only the router knows that.
|
||||
own := false
|
||||
if len(ownContent(text)) > 0 {
|
||||
own = offlineOwnRequest(text) || (ok && carriesOwnRequest(routed, text))
|
||||
}
|
||||
if !own {
|
||||
if replacesFilledSlot(q, answer) {
|
||||
return roleCorrection
|
||||
}
|
||||
return roleAnswer
|
||||
}
|
||||
if router.IsQuestionShaped(text) || (ok && (routed.Intent == router.IntentQuery || routed.Intent == router.IntentSystem)) {
|
||||
return roleSideQuery
|
||||
}
|
||||
return roleNewRequest
|
||||
}
|
||||
|
||||
// replacesFilledSlot reports whether the utterance overwrites something the
|
||||
// pending action already had, rather than filling the gap she asked about —
|
||||
// "нет, на девять" while she is waiting for the subject. Both are handled the
|
||||
// same way (dialogue.Answer already prefers the newer value), so this only
|
||||
// names the turn honestly for the log and for the decision trace V-564 adds.
|
||||
func replacesFilledSlot(q *dialogue.PendingQuestion, answer dialogue.Slots) bool {
|
||||
if q == nil {
|
||||
return false
|
||||
}
|
||||
asked := make(map[dialogue.Slot]bool, len(q.Missing))
|
||||
for _, s := range q.Missing {
|
||||
asked[s] = true
|
||||
}
|
||||
if answer.HasTime && q.Slots.HasTime && !asked[dialogue.SlotTime] && !answer.Time.Equal(q.Slots.Time) {
|
||||
return true
|
||||
}
|
||||
if answer.HasKey && q.Slots.HasKey && !asked[dialogue.SlotKey] && answer.Key != q.Slots.Key {
|
||||
return true
|
||||
}
|
||||
return false
|
||||
}
|
||||
@@ -0,0 +1,233 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/dialogue"
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
|
||||
// TestOwnContentSeparatesAValueFromAQuestion pins the test the whole role
|
||||
// classifier rests on: after the frame, the numbers and the closed time sets
|
||||
// come out, does anything of his own remain? A hedged time leaves nothing. A
|
||||
// question about the weather leaves the weather.
|
||||
func TestOwnContentSeparatesAValueFromAQuestion(t *testing.T) {
|
||||
cases := []struct {
|
||||
text string
|
||||
own bool
|
||||
}{
|
||||
{"в 11:00", false},
|
||||
{"в семь вечера", false},
|
||||
{"нет, в 15:00", false},
|
||||
{"а что если в 11:00", false},
|
||||
{"на 9", false},
|
||||
{"а, да, прости — на 9", false},
|
||||
{"завтра", false},
|
||||
{"в половине восьмого", false},
|
||||
{"какая сейчас погода в Риме?", true},
|
||||
{"кто изобрёл телефон", true},
|
||||
{"напомни в 11:00", true},
|
||||
{"позвонить маме", true},
|
||||
{"запиши что я пил воду", true},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
if got := len(ownContent(tc.text)) > 0; got != tc.own {
|
||||
t.Errorf("ownContent(%q) = %v, want own content = %v", tc.text, ownContent(tc.text), tc.own)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestCancelIsTheWholeUtterance — a call-off calls the request off, and a
|
||||
// sentence that merely contains the word does not.
|
||||
func TestCancelIsTheWholeUtterance(t *testing.T) {
|
||||
for _, yes := range []string{"отмена", "забудь", "неважно", "проехали", "cancel", "ой, отмена"} {
|
||||
if !isCancel(yes) {
|
||||
t.Errorf("isCancel(%q) = false, want true", yes)
|
||||
}
|
||||
}
|
||||
for _, no := range []string{"забудь купить молоко", "в 11:00", "позвонить маме", ""} {
|
||||
if isCancel(no) {
|
||||
t.Errorf("isCancel(%q) = true, want false", no)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestTurnRoleReadsTheRoutedDecision — the inversion itself. The same utterance
|
||||
// gets a different role depending on what the router made of it, which is the
|
||||
// evidence the old guard never had.
|
||||
func TestTurnRoleReadsTheRoutedDecision(t *testing.T) {
|
||||
q := &dialogue.PendingQuestion{
|
||||
Intent: dialogue.Intent(router.IntentReminder),
|
||||
Missing: []dialogue.Slot{dialogue.SlotTime},
|
||||
}
|
||||
dec := func(in router.Intent, s router.Slots) router.Decision {
|
||||
return router.Decision{Intent: in, Slots: s}
|
||||
}
|
||||
cases := []struct {
|
||||
name string
|
||||
text string
|
||||
routed router.Decision
|
||||
ok bool
|
||||
answer dialogue.Slots
|
||||
want turnRole
|
||||
}{
|
||||
{
|
||||
// The measured defect. The extractor finds "сейчас" and would have
|
||||
// closed the gap with it; the route says this is a question of its
|
||||
// own, and the question wins.
|
||||
name: "a world question mid-flow is a side query",
|
||||
text: "какая сейчас погода в Риме?",
|
||||
routed: dec(router.IntentQuery, router.Slots{Text: "какая сейчас погода в Риме?"}),
|
||||
ok: true,
|
||||
answer: dialogue.Slots{HasTime: true, Time: time.Now()},
|
||||
want: roleSideQuery,
|
||||
},
|
||||
{
|
||||
name: "a hedged time is an answer even routed as a query",
|
||||
text: "а что если в 11:00",
|
||||
routed: dec(router.IntentQuery, router.Slots{Text: "а что если в 11:00"}),
|
||||
ok: true,
|
||||
answer: dialogue.Slots{HasTime: true, Time: time.Now()},
|
||||
want: roleAnswer,
|
||||
},
|
||||
{
|
||||
name: "a fresh reminder is a new request",
|
||||
text: "напомни завтра позвонить маме",
|
||||
routed: dec(router.IntentReminder, router.Slots{Text: "позвонить маме", HasTime: true}),
|
||||
ok: true,
|
||||
want: roleNewRequest,
|
||||
},
|
||||
{
|
||||
name: "a capture is a new request",
|
||||
text: "запиши что я пил воду",
|
||||
routed: dec(router.IntentFact, router.Slots{Key: "water", HasKey: true}),
|
||||
ok: true,
|
||||
want: roleNewRequest,
|
||||
},
|
||||
{
|
||||
name: "an act that resolved to a capability is a new request",
|
||||
text: "выключи свет в спальне",
|
||||
routed: dec(router.IntentAct, router.Slots{Fn: "light_off", HasFn: true}),
|
||||
ok: true,
|
||||
want: roleNewRequest,
|
||||
},
|
||||
{
|
||||
// She could not route it. An utterance she did not understand does
|
||||
// not outrank the question in front of it.
|
||||
name: "a clarify decision is not a request of its own",
|
||||
text: "выключи свет",
|
||||
routed: router.Decision{Intent: router.IntentAct, Slots: router.Slots{Fn: "light_off", HasFn: true}, Clarify: true},
|
||||
ok: true,
|
||||
want: roleAnswer,
|
||||
},
|
||||
{
|
||||
name: "a bare noun that answers nothing is still an answer",
|
||||
text: "ага",
|
||||
ok: false,
|
||||
want: roleAnswer,
|
||||
},
|
||||
{
|
||||
name: "no route to read falls back to the shape",
|
||||
text: "кто изобрёл телефон",
|
||||
ok: false,
|
||||
want: roleSideQuery,
|
||||
},
|
||||
{
|
||||
name: "a call-off needs no route at all",
|
||||
text: "отмена",
|
||||
ok: false,
|
||||
want: roleCancel,
|
||||
},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
if got := classifyTurnRole(q, tc.text, tc.answer, tc.routed, tc.ok); got != tc.want {
|
||||
t.Errorf("%s: classifyTurnRole(%q) = %s, want %s", tc.name, tc.text, got, tc.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestTurnRoleNamesACorrection — the answer overwrites a slot she was not
|
||||
// asking about. Handled like an answer, named as what it is.
|
||||
func TestTurnRoleNamesACorrection(t *testing.T) {
|
||||
nine := time.Date(2026, 8, 6, 9, 0, 0, 0, time.UTC)
|
||||
q := &dialogue.PendingQuestion{
|
||||
Intent: dialogue.Intent(router.IntentReminder),
|
||||
Missing: []dialogue.Slot{dialogue.SlotText},
|
||||
Slots: dialogue.Slots{HasTime: true, Time: nine.Add(2 * time.Hour)},
|
||||
}
|
||||
got := classifyTurnRole(q, "нет, на 9", dialogue.Slots{HasTime: true, Time: nine}, router.Decision{}, false)
|
||||
if got != roleCorrection {
|
||||
t.Fatalf("role = %s, want %s", got, roleCorrection)
|
||||
}
|
||||
}
|
||||
|
||||
// TestRomeIsAnsweredAndTheReminderIsNotInvented — the measured failure of
|
||||
// 2026-08-05, end to end through the real cascade. "напомни позвонить маме"
|
||||
// parks the time question; the weather question that follows must not become
|
||||
// its answer, must not create a reminder for a time nobody asked for, and must
|
||||
// not be dropped in silence.
|
||||
func TestRomeIsAnsweredAndTheReminderIsNotInvented(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, st := newRoutingClarifyHandler(t)
|
||||
|
||||
if reply := h.handleText(ctx, "web", "напомни позвонить маме"); !strings.Contains(reply, "?") {
|
||||
t.Fatalf("expected the time question, got %q", reply)
|
||||
}
|
||||
reply := h.handleText(ctx, "web", "какая сейчас погода в Риме?")
|
||||
if strings.Contains(reply, "напомню") {
|
||||
t.Fatalf("the question was eaten as the reminder's time again: %q", reply)
|
||||
}
|
||||
if !strings.HasPrefix(reply, clarifyDropped) {
|
||||
t.Fatalf("the parked request died without a word: %q", reply)
|
||||
}
|
||||
if reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour)); err != nil || len(reminders) != 0 {
|
||||
t.Fatalf("a reminder was invented for a time nobody asked for: %v err=%v", reminders, err)
|
||||
}
|
||||
if h.clarifyStore.Get(dialogueIDFor(sourceText, "web"), h.now()) != nil {
|
||||
t.Fatal("the parked question must be gone, not left to eat the next turn")
|
||||
}
|
||||
}
|
||||
|
||||
// TestClarifyCancelEndsTheExchange — "отмена" while she is waiting calls the
|
||||
// half-built request off, out loud, and creates nothing.
|
||||
func TestClarifyCancelEndsTheExchange(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, st := newRoutingClarifyHandler(t)
|
||||
|
||||
if reply := h.handleText(ctx, "web", "напомни позвонить маме"); !strings.Contains(reply, "?") {
|
||||
t.Fatalf("expected the time question, got %q", reply)
|
||||
}
|
||||
if reply := h.handleText(ctx, "web", "отмена"); reply != clarifyCancelled {
|
||||
t.Fatalf("reply = %q, want %q", reply, clarifyCancelled)
|
||||
}
|
||||
if reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour)); err != nil || len(reminders) != 0 {
|
||||
t.Fatalf("a cancelled request still landed: %v err=%v", reminders, err)
|
||||
}
|
||||
if h.clarifyStore.Get(dialogueIDFor(sourceText, "web"), h.now()) != nil {
|
||||
t.Fatal("a cancelled exchange must leave nothing parked")
|
||||
}
|
||||
}
|
||||
|
||||
// TestTheTurnIsRoutedOnce — the cost bound. A turn with a question parked pays
|
||||
// for one extra route and not two: the clarify resolver and the pipeline read
|
||||
// the same memo.
|
||||
func TestTheTurnIsRoutedOnce(t *testing.T) {
|
||||
h, _ := newRoutingClarifyHandler(t)
|
||||
rt := h.newTurnRoute("какая сейчас погода в Риме?", h.now())
|
||||
ctx := withTurnRoute(withDialogueID(context.Background(), voiceDialogueID), rt)
|
||||
|
||||
first, ok := h.routeForRole(ctx, rt.text)
|
||||
if !ok {
|
||||
t.Fatal("the cascade must produce a decision to classify against")
|
||||
}
|
||||
second, _, _, err := rt.resolve(ctx)
|
||||
if err != nil {
|
||||
t.Fatalf("resolve: %v", err)
|
||||
}
|
||||
if second.Intent != first.Intent || second.Utterance != first.Utterance {
|
||||
t.Fatalf("the pipeline routed again and got something else: %+v vs %+v", second, first)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,104 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"log"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/dialogue"
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
|
||||
// turnRoute is this turn's routing, computed at most once.
|
||||
//
|
||||
// It exists because the arbitration was inverted (Vikunja #560): the clarify
|
||||
// resolver now reads the routed decision before deciding what the utterance is,
|
||||
// and the pipeline then acts on that same decision. Routing twice would cost a
|
||||
// second on the resident model and — worse — could disagree with itself, which
|
||||
// is exactly the class of bug this task is about.
|
||||
type turnRoute struct {
|
||||
h *reactiveHandler
|
||||
text string
|
||||
now time.Time
|
||||
|
||||
once sync.Once
|
||||
dec router.Decision
|
||||
cont bool
|
||||
prev *dialogue.Session
|
||||
err error
|
||||
|
||||
// dropped — what she let go of this turn and must say out loud. A parked
|
||||
// request that dies without a word leaves him thinking it landed.
|
||||
dropped string
|
||||
}
|
||||
|
||||
type turnRouteKey struct{}
|
||||
|
||||
func (h *reactiveHandler) newTurnRoute(text string, now time.Time) *turnRoute {
|
||||
return &turnRoute{h: h, text: text, now: now}
|
||||
}
|
||||
|
||||
func withTurnRoute(ctx context.Context, rt *turnRoute) context.Context {
|
||||
return context.WithValue(ctx, turnRouteKey{}, rt)
|
||||
}
|
||||
|
||||
// turnRouteFrom returns the turn's memo, or nil when the caller is not inside
|
||||
// runTurn — a unit test calling one resolver directly, most often.
|
||||
func turnRouteFrom(ctx context.Context) *turnRoute {
|
||||
rt, _ := ctx.Value(turnRouteKey{}).(*turnRoute)
|
||||
return rt
|
||||
}
|
||||
|
||||
// resolve does the routing exactly as step 5 of runTurn does it: an elliptical
|
||||
// follow-up is answered from the previous turn, everything else goes to the
|
||||
// router. One copy of that, so the pre-route the clarify resolver reads and the
|
||||
// decision the pipeline acts on cannot drift apart.
|
||||
func (r *turnRoute) resolve(ctx context.Context) (router.Decision, bool, *dialogue.Session, error) {
|
||||
r.once.Do(func() {
|
||||
if r.h.dialogueSessions != nil {
|
||||
r.prev = r.h.dialogueSessions.Get(dialogueIDOf(ctx), r.now)
|
||||
}
|
||||
if dec, cont := continuationDecision(r.prev, r.text, r.now); cont {
|
||||
log.Printf("voice: continuation of %s from the previous turn", dec.Intent)
|
||||
r.dec, r.cont = dec, true
|
||||
return
|
||||
}
|
||||
if r.h.router == nil {
|
||||
r.err = router.ErrNoIntents
|
||||
return
|
||||
}
|
||||
r.dec, r.err = r.h.router.Route(ctx, r.text, r.now)
|
||||
})
|
||||
return r.dec, r.cont, r.prev, r.err
|
||||
}
|
||||
|
||||
// routeForRole gives the role classifier the turn's routed decision. The second
|
||||
// return is false when there is no usable decision — no router wired, or the
|
||||
// route failed — and the classifier falls back to its offline tests then. A
|
||||
// turn must never break on the model, so the error is logged and swallowed
|
||||
// here; step 5 reads the same memo and reports it the way it always has.
|
||||
func (h *reactiveHandler) routeForRole(ctx context.Context, text string) (router.Decision, bool) {
|
||||
rt := turnRouteFrom(ctx)
|
||||
if rt == nil {
|
||||
rt = h.newTurnRoute(text, h.now())
|
||||
}
|
||||
dec, _, _, err := rt.resolve(ctx)
|
||||
if err != nil {
|
||||
log.Printf("voice: role — no route to classify against (%v), falling back to the offline tests", err)
|
||||
return router.Decision{}, false
|
||||
}
|
||||
return dec, true
|
||||
}
|
||||
|
||||
// needsRoute reports whether classifying this utterance's role is worth a
|
||||
// route. It is not: an utterance with no content of its own carries no request
|
||||
// of its own, so the classifier reaches the same answer without the model. A
|
||||
// call-off is the same — it is read off a closed lexicon and nothing else.
|
||||
//
|
||||
// This is a fast path to the SAME answer and must stay one. If it ever needs a
|
||||
// rule the classifier does not have, it has become a second decision procedure
|
||||
// and it is the thing V-560 deleted.
|
||||
func needsRoute(text string) bool {
|
||||
return !isCancel(text) && len(ownContent(text)) > 0
|
||||
}
|
||||
+73
-33
@@ -53,6 +53,7 @@ import (
|
||||
|
||||
"github.com/kami/maven/internal/audio"
|
||||
"github.com/kami/maven/internal/crawl"
|
||||
"github.com/kami/maven/internal/decision"
|
||||
"github.com/kami/maven/internal/dialogue"
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/lexicon"
|
||||
@@ -132,10 +133,18 @@ type reactiveHandler struct {
|
||||
// The production dateparser will replace StubDateTimeParser here too.
|
||||
timeParser router.DateTimeParser
|
||||
|
||||
// dialogueSessions carries slots across turns for follow-ups (single-user
|
||||
// box → one session slot, keyed voiceDialogueID). nil ⇒ no carry-over.
|
||||
// dialogueSessions carries slots across turns for follow-ups. Keyed by the
|
||||
// reach the turn arrived on (dialogueIDOf), like the clarify store: one
|
||||
// slot per reach, not one for the box. nil ⇒ no carry-over.
|
||||
dialogueSessions *dialogue.SessionStore
|
||||
|
||||
// decisions holds the last few turns' arbitration records (V-564): who
|
||||
// claimed the turn, who lost it and who was never asked. In memory and
|
||||
// bounded, because a turn record is read minutes later or never, and none
|
||||
// of his words belong in a table that outlives the diagnosis. nil ⇒ nothing
|
||||
// is recorded, which is what a test that did not ask for one gets.
|
||||
decisions *decision.Ring
|
||||
|
||||
// clarifyStore parks the request behind an open question she asked (see
|
||||
// clarify.go). nil ⇒ she falls back to the canned "не поняла" reply.
|
||||
clarifyStore *dialogue.ClarifyStore
|
||||
@@ -158,6 +167,15 @@ type reactiveHandler struct {
|
||||
pendingRoutine *pendingRoutineConfirm // routine proposal awaiting y/n
|
||||
pendingHexis *pendingHexisExec // mutating Hexis capability awaiting y/n
|
||||
|
||||
// surfacedItems — the Praxis item ids she last read out, in the order she
|
||||
// read them, so "отметь второй пункт" has a second pункт to mean (Vikunja
|
||||
// #516). Same single-slot posture as pending above: the next attention digest
|
||||
// replaces the list, because a position only refers to the last one spoken.
|
||||
// No TTL — a stale position resolves to an item that Praxis will report as
|
||||
// already acknowledged, which is a harmless answer, unlike a stale
|
||||
// confirmation that would execute something.
|
||||
surfacedItems []string
|
||||
|
||||
ecosystem *ecosystemWiring // nexus + hexis + praxis clients
|
||||
}
|
||||
|
||||
@@ -238,6 +256,26 @@ const (
|
||||
//
|
||||
// The ordering is load-bearing — see the step comments.
|
||||
func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSource) string {
|
||||
// 0. the decision record (V-564). Installed here rather than in the IPC
|
||||
// entry point, so the mic, telegram and the web all leave the same trail —
|
||||
// a record only the web produced would be missing exactly the turns that
|
||||
// are hardest to reproduce. It rides the context, costs a few dozen structs
|
||||
// on a human-rate path, and no claim site can change a route with it.
|
||||
if h.decisions != nil {
|
||||
var rec *decision.Record
|
||||
ctx, rec = decision.With(ctx, text)
|
||||
decision.Expect(ctx, decision.StagePreRoute, preRouteLadder)
|
||||
defer func() { h.decisions.Push(rec.Finish(h.now())) }()
|
||||
}
|
||||
|
||||
// 0b. the turn's routing, computed at most once and shared (Vikunja #560).
|
||||
// The clarify resolver reads it to decide what this utterance IS before
|
||||
// claiming it, and step 5 acts on the same decision — routing twice would
|
||||
// cost a second on the resident model and could disagree with itself.
|
||||
now := h.now()
|
||||
rt := h.newTurnRoute(text, now)
|
||||
ctx = withTurnRoute(ctx, rt)
|
||||
|
||||
// 1. expired clarify — a question was parked but its TTL ran out, so the
|
||||
// request behind it is gone. Say that out loud (see clarify.go) and carry
|
||||
// on: these words are still routed as a fresh utterance below, with the
|
||||
@@ -253,7 +291,7 @@ func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSour
|
||||
// 2. confirm turn — if a destructive act is parked, this utterance is its
|
||||
// y/n answer, not a fresh command. Handled before routing so "да" doesn't
|
||||
// get classified as some other intent.
|
||||
if reply, handled := h.resolveConfirm(ctx, text); handled {
|
||||
if reply, handled := h.resolveConfirm(ctx, text); notePreRoute(ctx, "confirm", handled) {
|
||||
return withNotice(expiredNotice, reply)
|
||||
}
|
||||
|
||||
@@ -265,15 +303,19 @@ func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSour
|
||||
// so the notice is empty here in practice. withNotice anyway: every exit
|
||||
// from runTurn carries it, and that is what stops the next one from
|
||||
// forgetting.
|
||||
if reply, handled := h.resolveClarifyAnswer(ctx, text); handled {
|
||||
if reply, handled := h.resolveClarifyAnswer(ctx, text); notePreRoute(ctx, "clarify-answer", handled) {
|
||||
return withNotice(expiredNotice, reply)
|
||||
}
|
||||
// It did not claim the turn. If it let a parked request go to get out of the
|
||||
// way, that has to be said in front of whatever these words are answered
|
||||
// with — carried on the same notice, so every exit below keeps it.
|
||||
expiredNotice = withNotice(expiredNotice, rt.dropped)
|
||||
|
||||
// 4. quiet-hours toggle — keyword match, not classifier-dependent.
|
||||
// "тихий режим" / "quiet on" would route through the classifier
|
||||
// unreliably (it's a command, not a free-form query), so we match it
|
||||
// before routing. Same pattern as the confirm turn above.
|
||||
if reply, handled := h.resolveQuietToggle(ctx, text, src); handled {
|
||||
if reply, handled := h.resolveQuietToggle(ctx, text, src); notePreRoute(ctx, "quiet-toggle", handled) {
|
||||
return withNotice(expiredNotice, reply)
|
||||
}
|
||||
|
||||
@@ -281,14 +323,14 @@ func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSour
|
||||
// sent. Only handled when a pending nudge is actually inside the window
|
||||
// (snooze.go); otherwise the words route normally, because "потом" is an
|
||||
// ordinary word and eating every one of them would break real sentences.
|
||||
if reply, handled := h.resolveSnooze(ctx, text, src); handled {
|
||||
if reply, handled := h.resolveSnooze(ctx, text, src); notePreRoute(ctx, "snooze", handled) {
|
||||
return withNotice(expiredNotice, reply)
|
||||
}
|
||||
|
||||
// 4c. spoken ack — "готово" closes that same nudge as `acted`. Only the
|
||||
// contentless form is intercepted here; "выпил воды" keeps routing and
|
||||
// closes the nudge after its fact lands (ackFromFact, step 8b).
|
||||
if reply, handled := h.resolveAck(ctx, text, src); handled {
|
||||
if reply, handled := h.resolveAck(ctx, text, src); notePreRoute(ctx, "ack", handled) {
|
||||
return withNotice(expiredNotice, reply)
|
||||
}
|
||||
|
||||
@@ -296,7 +338,7 @@ func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSour
|
||||
// turn and names what it should have been (repair.go). Before routing,
|
||||
// like the confirm and clarify turns: routing the correction as a fresh
|
||||
// utterance files the correction itself instead of fixing anything.
|
||||
if reply, handled := h.resolveRepair(ctx, text); handled {
|
||||
if reply, handled := h.resolveRepair(ctx, text); notePreRoute(ctx, "repair", handled) {
|
||||
return withNotice(expiredNotice, reply)
|
||||
}
|
||||
|
||||
@@ -304,7 +346,7 @@ func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSour
|
||||
// just read (ordinal.go). Before routing, and only when a list is actually
|
||||
// bound to the session: with nothing offered, "второй" is an ordinary word
|
||||
// and keeps routing.
|
||||
if reply, handled := h.resolveCandidate(ctx, text); handled {
|
||||
if reply, handled := h.resolveCandidate(ctx, text, src); notePreRoute(ctx, "ordinal", handled) {
|
||||
return withNotice(expiredNotice, reply)
|
||||
}
|
||||
|
||||
@@ -313,22 +355,7 @@ func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSour
|
||||
// missing, so no amount of routing recovers it, and the model's guess
|
||||
// costs seconds to obtain and is close to a coin flip. Everything else
|
||||
// goes to the router.
|
||||
var (
|
||||
dec router.Decision
|
||||
err error
|
||||
prev *dialogue.Session
|
||||
)
|
||||
now := h.now()
|
||||
if h.dialogueSessions != nil {
|
||||
prev = h.dialogueSessions.Get(voiceDialogueID, now)
|
||||
}
|
||||
cont := false
|
||||
if dec, cont = continuationDecision(prev, text, now); cont {
|
||||
log.Printf("voice: continuation of %s from the previous turn", dec.Intent)
|
||||
}
|
||||
if !cont {
|
||||
dec, err = h.router.Route(ctx, text, now)
|
||||
}
|
||||
dec, cont, prev, err := rt.resolve(ctx)
|
||||
if err != nil {
|
||||
// ErrNoIntents ⇒ classifier unseeded (cold boot). reply with a
|
||||
// "still warming up" rather than a wire error.
|
||||
@@ -349,21 +376,31 @@ func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSour
|
||||
// ("а завтра?" … "а послезавтра?") keeps working.
|
||||
if h.dialogueSessions != nil {
|
||||
if !cont {
|
||||
dec = followUpMerge(prev, dec, now)
|
||||
merged := followUpMerge(prev, dec, now)
|
||||
noteMerge(ctx, dec, merged)
|
||||
dec = merged
|
||||
}
|
||||
if !dec.Clarify {
|
||||
h.rememberTurn(prev, dec, now)
|
||||
h.rememberTurn(ctx, prev, dec, now)
|
||||
}
|
||||
}
|
||||
|
||||
// 7. clarify — she is not sure. If one named thing is missing, ask about it
|
||||
// and park the request (clarify.go); otherwise the replier's canned reply
|
||||
// stands.
|
||||
if dec.Clarify {
|
||||
// 7. clarify — something she needs is missing. If one named thing is missing,
|
||||
// ask about it and park the request (clarify.go); otherwise the replier's
|
||||
// canned reply stands.
|
||||
//
|
||||
// Not gated on dec.Clarify alone (Vikunja #557). A turn the cascade routed
|
||||
// confidently but incompletely skipped this entirely: "напомни позвонить"
|
||||
// reached applyAction, failed on the missing time, parked nothing, and the
|
||||
// "в семь вечера" that followed was web-searched as a world question. A
|
||||
// required slot that missingFor names is a gap whatever the confidence.
|
||||
if dec.Clarify || len(missingFor(dec)) > 0 {
|
||||
if reply := h.hexisBeforeClarify(ctx, dec); reply != "" {
|
||||
return withNotice(expiredNotice, reply)
|
||||
}
|
||||
if question, asked := h.askClarify(ctx, dec); asked {
|
||||
noteTerminal(ctx, "clarify-ask", dec.Intent,
|
||||
"the route was below the threshold, so she asked instead of acting")
|
||||
return withNotice(expiredNotice, question)
|
||||
}
|
||||
}
|
||||
@@ -380,6 +417,9 @@ func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSour
|
||||
// the round-trip stays alive.
|
||||
replyText := h.applyAction(ctx, dec)
|
||||
log.Printf("voice: applyAction returned: %q", replyText)
|
||||
// A query turn was already claimed by a source inside the chain; every other
|
||||
// intent has no chain and no scoreboard, so the handler is the winner.
|
||||
noteTerminal(ctx, "action-handler", dec.Intent, "")
|
||||
|
||||
// 8b. a fact that answers a live nudge closes it as `acted` (ack.go).
|
||||
// Silent: the fact reply stands, she does not congratulate him for it.
|
||||
@@ -483,12 +523,12 @@ func (h *reactiveHandler) replySystem(ctx context.Context, dec router.Decision)
|
||||
// chatHistory collects dialogue turns from the session store for the current
|
||||
// conversation. Returns prior user utterances (newest last) up to a depth of
|
||||
// 4 turns. Returns nil when there's no session or no history.
|
||||
func (h *reactiveHandler) chatHistory() []dialogue.Turn {
|
||||
func (h *reactiveHandler) chatHistory(ctx context.Context) []dialogue.Turn {
|
||||
if h.dialogueSessions == nil {
|
||||
return nil
|
||||
}
|
||||
now := h.now()
|
||||
prev := h.dialogueSessions.Get(voiceDialogueID, now)
|
||||
prev := h.dialogueSessions.Get(dialogueIDOf(ctx), now)
|
||||
if prev == nil {
|
||||
return nil
|
||||
}
|
||||
|
||||
+23
-1
@@ -11,6 +11,7 @@ import (
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/config"
|
||||
"github.com/kami/maven/internal/decision"
|
||||
"github.com/kami/maven/internal/delivery"
|
||||
"github.com/kami/maven/internal/delivery/voicesink"
|
||||
"github.com/kami/maven/internal/dialogue"
|
||||
@@ -293,7 +294,11 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
|
||||
},
|
||||
dataStore: dataStore,
|
||||
dialogueSessions: dialogueSessions,
|
||||
clarifyStore: clarifyStore,
|
||||
// Always on (V-564). The record is the instrument the rest of V-558 is
|
||||
// measured with, and one that only runs when a flag is set is not there
|
||||
// on the night the misroute happens.
|
||||
decisions: decision.NewRing(),
|
||||
clarifyStore: clarifyStore,
|
||||
// 0 here (unset config) ⇒ the dialogue default.
|
||||
clarifyMaxAttempts: cfg.Voice.ClarifyMaxAttempts,
|
||||
extractor: router.Extractor{Time: timeParser, Acts: matcher, Facts: router.DefaultFactParser{}},
|
||||
@@ -392,10 +397,22 @@ func buildRouter(emb router.Embedder, acts router.ActMatcher, threshold float64,
|
||||
grammars = append(grammars, router.TaskListGrammar())
|
||||
grammars = append(grammars, router.ListGrammars()...)
|
||||
grammars = append(grammars, router.ReminderGrammar())
|
||||
// Before the capture marker, because "отметь" is a capture verb and "отметь
|
||||
// второй пункт" is not a note. The Praxis rules are the narrower claim — a
|
||||
// lifecycle verb AND an item named — so they get first refusal (Vikunja #516).
|
||||
grammars = append(grammars, router.PraxisGrammars()...)
|
||||
// Last, and it matches any utterance shape — its Build is the filter. An
|
||||
// explicit capture marker beats the model, which called it an act and
|
||||
// rewrote the task text (Vikunja #467). After the rules above because a
|
||||
// marker never collides with a clock or agenda question.
|
||||
// After Praxis, whose bare "закрой" claim this rule cannot reach (it needs the
|
||||
// board noun), and before the capture marker, which would otherwise read
|
||||
// "убери из задач купить молоко" as a new task (Vikunja #512).
|
||||
grammars = append(grammars, router.TaskStatusGrammar())
|
||||
// Before the capture markers, which all need an object. A capture verb
|
||||
// alone is a fact with no key, and the clarify path asks for it rather than
|
||||
// letting the model invent an answer (Vikunja #557).
|
||||
grammars = append(grammars, router.BareCaptureGrammar()...)
|
||||
grammars = append(grammars, router.TaskCaptureGrammar())
|
||||
// After the capture marker, so "запиши" still wins over "расскажи", and
|
||||
// last overall because it matches on the first word alone: "расскажи про
|
||||
@@ -552,6 +569,11 @@ func repairFactVectors(dataStore *store.Store, emb router.Embedder) {
|
||||
// runReembed.
|
||||
var reembedOnStart bool
|
||||
|
||||
// allowSeedOnStart is the -allow-seed flag (set in run()). Opt-in, and the
|
||||
// default is the one that matters: a box nobody is testing has no live path to
|
||||
// write a fact into the past. See seed.go and Vikunja #518.
|
||||
var allowSeedOnStart bool
|
||||
|
||||
// checkStoredEmbedder compares the embedder we just loaded with the one that
|
||||
// wrote the vectors already in the DB (Vikunja #378).
|
||||
//
|
||||
|
||||
@@ -0,0 +1,45 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"io"
|
||||
|
||||
"github.com/kami/maven/internal/store"
|
||||
)
|
||||
|
||||
// runWipe implements the -wipe flag: it prints what the database holds, and
|
||||
// removes it only when the operator also passed -confirm-wipe (Vikunja #494).
|
||||
//
|
||||
// Two flags rather than one, because the destructive reading of a single flag
|
||||
// is the reading a mistyped command gets. Without the confirmation this is a
|
||||
// dry run that costs nothing and answers the question a QA session actually
|
||||
// has — what is on this box right now.
|
||||
//
|
||||
// It runs before any daemon component is wired, so nothing is writing while
|
||||
// the tables go. The daemon exits afterwards rather than serving a store it
|
||||
// just emptied, because every component that read the old rows at boot would
|
||||
// still be holding them.
|
||||
func runWipe(ctx context.Context, st *store.Store, out io.Writer, confirmed bool) error {
|
||||
counts, err := st.WipeCounts(ctx)
|
||||
if err != nil {
|
||||
return fmt.Errorf("wipe: read counts: %w", err)
|
||||
}
|
||||
total := 0
|
||||
for _, c := range counts {
|
||||
total += c.Rows
|
||||
fmt.Fprintf(out, " %-24s %d\n", c.Table, c.Rows)
|
||||
}
|
||||
fmt.Fprintf(out, " %-24s %d rows in %d tables\n", "TOTAL", total, len(counts))
|
||||
|
||||
if !confirmed {
|
||||
fmt.Fprintln(out, "\nnothing was deleted. pass -confirm-wipe to delete all of it.")
|
||||
fmt.Fprintln(out, "config, models, passkeys and the encryption key are files and are never touched.")
|
||||
return nil
|
||||
}
|
||||
if err := st.Wipe(ctx); err != nil {
|
||||
return err
|
||||
}
|
||||
fmt.Fprintf(out, "\nwiped. %d rows gone, the schema is intact, mavend knows nobody.\n", total)
|
||||
return nil
|
||||
}
|
||||
@@ -31,12 +31,11 @@ import (
|
||||
// not a guesser-of-truth, and a mailbox of noise rendered as invented meetings
|
||||
// is worse than a gap.
|
||||
//
|
||||
// KNOWN GAP: this writes calendar_event_* and nothing else, so an ambient
|
||||
// meeting is good enough to recite and not good enough to stop a nudge —
|
||||
// calendar_busy is still written only by the CalDAV poller. That is backwards,
|
||||
// since suppressing a nudge is the lower-risk use of a low-confidence signal.
|
||||
// calendar_busy is a level rather than an event, so an ambient writer needs an
|
||||
// expiry, which is its own task and not a change here.
|
||||
// This writes calendar_event_* and nothing else, and since Vikunja #513 that is
|
||||
// enough to stop a nudge as well as to recite: the loop gatherer reads the event
|
||||
// family and asks whether any span covers the instant. So there is no ambient
|
||||
// calendar_busy and no expiry to pick — a level needs one and an event carries
|
||||
// its own. calendar_busy stays the CalDAV poller's key.
|
||||
|
||||
// ambientMaxBody bounds the request. A notification is two short lines.
|
||||
const ambientMaxBody = 8 << 10
|
||||
|
||||
@@ -4,7 +4,7 @@
|
||||
{{if .Error}}<div class="msg msg-err">{{.Error}}</div>{{end}}
|
||||
<div class="scroll chat-scroll" id=chatHistory>
|
||||
{{range .Messages}}
|
||||
<div class="chat-msg {{.Role}}"><strong>{{if eq .Role "user"}}you{{else}}maven{{end}}:</strong> {{.Text}}</div>
|
||||
<div class="chat-msg {{.Role}}"><strong>{{if eq .Role "user"}}you{{else}}maven{{end}}:</strong> {{.Text}}{{if .Source}} <span class="badge badge-accent" title="the query source that claimed this turn">{{.Source}}</span>{{end}}</div>
|
||||
{{else}}
|
||||
<div class=empty>
|
||||
<svg class=icon width="20" height="20"><use href="/ethos-icons.svg#i-message"/></svg>
|
||||
|
||||
@@ -65,10 +65,12 @@ type fakeCore struct {
|
||||
// for handleTrace tests
|
||||
tickTrace ipc.TickTrace
|
||||
traceErr error
|
||||
turns []ipc.TurnDecision
|
||||
|
||||
// for handleChatAPI tests
|
||||
chatText string
|
||||
chatErr error
|
||||
chatText string
|
||||
chatSource string
|
||||
chatErr error
|
||||
|
||||
// for the MCP section of /tools
|
||||
mcpServers []ipc.MCPServerStatus
|
||||
@@ -79,12 +81,12 @@ func (f *fakeCore) MCPServers(context.Context) ([]ipc.MCPServerStatus, error) {
|
||||
return f.mcpServers, f.mcpErr
|
||||
}
|
||||
|
||||
func (f *fakeCore) Chat(_ context.Context, _, text string) (string, error) {
|
||||
func (f *fakeCore) Chat(_ context.Context, _, text string) (ipc.ChatReply, error) {
|
||||
f.chatText = text
|
||||
if f.chatErr != nil {
|
||||
return "", f.chatErr
|
||||
return ipc.ChatReply{}, f.chatErr
|
||||
}
|
||||
return "поняла", nil
|
||||
return ipc.ChatReply{Reply: "поняла", Source: f.chatSource}, nil
|
||||
}
|
||||
|
||||
func (f *fakeCore) EnableTool(_ context.Context, name string, cmd []string, destructive bool, scope string, _ time.Time) error {
|
||||
@@ -172,6 +174,10 @@ func (f *fakeCore) RevertFact(_ context.Context, key string) (int64, error) {
|
||||
return f.revertNewID, nil
|
||||
}
|
||||
|
||||
func (f *fakeCore) TurnDecisions(_ context.Context, _ int) ([]ipc.TurnDecision, error) {
|
||||
return f.turns, nil
|
||||
}
|
||||
|
||||
func (f *fakeCore) TickTrace(_ context.Context) (ipc.TickTrace, error) {
|
||||
if f.traceErr != nil {
|
||||
return ipc.TickTrace{}, f.traceErr
|
||||
@@ -824,6 +830,41 @@ func TestHandleTrace(t *testing.T) {
|
||||
t.Error("rendered 'nothing fired' but a winner was set")
|
||||
}
|
||||
})
|
||||
|
||||
// The turn arbitration shares this page (V-564). A reader must see the
|
||||
// winner, a loser and the claimants that were never asked, because the last
|
||||
// of those is what the hardcoded ordering hides.
|
||||
t.Run("renders the turn decision record", func(t *testing.T) {
|
||||
core := &fakeCore{turns: []ipc.TurnDecision{{
|
||||
Ts: time.Date(2025, 6, 1, 12, 0, 0, 0, time.UTC),
|
||||
Utterance: "какая погода в риме",
|
||||
Winner: "query:weather",
|
||||
Claims: []ipc.TurnClaim{
|
||||
{Stage: "query", Claimant: "weather", Intent: "query", Outcome: "won"},
|
||||
{Stage: "query", Claimant: "calendar", Outcome: "declined", Reason: "no answer"},
|
||||
{Stage: "query", Claimant: "kiwix", Outcome: "never_asked"},
|
||||
},
|
||||
}}}
|
||||
rr := httptest.NewRecorder()
|
||||
handleTrace(rr, httptest.NewRequest(http.MethodGet, "/trace", nil), core)
|
||||
body := rr.Body.String()
|
||||
for _, want := range []string{"какая погода в риме", "query:weather", "calendar", "kiwix", "never_asked"} {
|
||||
if !strings.Contains(body, want) {
|
||||
t.Errorf("rendered page is missing %q", want)
|
||||
}
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("no turns renders the empty note, not an error", func(t *testing.T) {
|
||||
rr := httptest.NewRecorder()
|
||||
handleTrace(rr, httptest.NewRequest(http.MethodGet, "/trace", nil), &fakeCore{})
|
||||
if rr.Code != http.StatusOK {
|
||||
t.Fatalf("status = %d, want 200", rr.Code)
|
||||
}
|
||||
if !strings.Contains(rr.Body.String(), "no turn has run") {
|
||||
t.Error("empty ring did not render its note")
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
// --- handleRevert ---
|
||||
@@ -1285,3 +1326,37 @@ func TestHandleNotifications_ShowsTheOutbox(t *testing.T) {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// --- the query source badge (V-539) ---
|
||||
//
|
||||
// Which query source claimed a turn was readable in the daemon log and nowhere
|
||||
// else, so a QA step could not tell a wrong answer from a wrongly ordered
|
||||
// chain. It now rides the redirect and renders beside the reply.
|
||||
|
||||
func TestHandleChatAPI_CarriesTheClaimingSource(t *testing.T) {
|
||||
core := &fakeCore{chatSource: "kiwix"}
|
||||
rr := httptest.NewRecorder()
|
||||
handleChatAPI(rr, postChat("почему небо голубое"), core, nil, false)
|
||||
loc := rr.Header().Get("Location")
|
||||
if !strings.Contains(loc, "s=kiwix") {
|
||||
t.Errorf("redirect = %q; want the claiming source in it", loc)
|
||||
}
|
||||
}
|
||||
|
||||
func TestHandleChatAPI_OmitsTheSourceWhenNothingClaimed(t *testing.T) {
|
||||
core := &fakeCore{}
|
||||
rr := httptest.NewRecorder()
|
||||
handleChatAPI(rr, postChat("запиши что я пил воду"), core, nil, false)
|
||||
if loc := rr.Header().Get("Location"); strings.Contains(loc, "s=") {
|
||||
t.Errorf("redirect = %q; a turn no source claimed carries no badge", loc)
|
||||
}
|
||||
}
|
||||
|
||||
func TestChatPageRendersTheSourceBadge(t *testing.T) {
|
||||
req := httptest.NewRequest(http.MethodGet, "/chat?q=%D1%82%D0%B5%D1%81%D1%82&r=%D0%BE%D1%82%D0%B2%D0%B5%D1%82&s=search", nil)
|
||||
rr := httptest.NewRecorder()
|
||||
handleChatPage(rr, req, &fakeCore{})
|
||||
if body := rr.Body.String(); !strings.Contains(body, ">search</span>") {
|
||||
t.Errorf("chat page does not render the source badge; body=%s", body)
|
||||
}
|
||||
}
|
||||
|
||||
+301
-47
@@ -883,14 +883,19 @@ type taskRow struct {
|
||||
Created string
|
||||
Resolved string
|
||||
ResolvedBy string
|
||||
// DueValue and Weight are the raw values the edit form posts back
|
||||
// (Vikunja #509). Due above is for reading and says "—" for no date; a
|
||||
// date input needs "2026-08-07" or the empty string.
|
||||
DueValue string
|
||||
Weight int
|
||||
// Why — the ranker's reason for this row's position (Vikunja #129), in
|
||||
// Russian, empty when nothing distinguished the task. Blank is the honest
|
||||
// rendering: he never said this one mattered more.
|
||||
Why string
|
||||
}
|
||||
|
||||
// handleTasks serves the task review surface (GET) and the four writes it
|
||||
// offers (POST): add, confirm, done, drop.
|
||||
// handleTasks serves the task review surface (GET) and the five writes it
|
||||
// offers (POST): add, edit, confirm, done, drop.
|
||||
//
|
||||
// Not step-up gated, unlike /tools and /routines, and the difference is the
|
||||
// point: enabling a tool defines argv Maven will execute, and accepting a
|
||||
@@ -900,6 +905,13 @@ type taskRow struct {
|
||||
// still sits behind whatever transport auth fronts mavweb, like every other
|
||||
// page.
|
||||
//
|
||||
// "edit" was re-argued on the same terms rather than inheriting the exemption
|
||||
// (Vikunja #509), and it stays ungated. It rewrites a line on a list he reads
|
||||
// himself, the same blast radius "drop" already has on this page, and the store
|
||||
// refuses the two edits that would cost something: a resolved task keeps the
|
||||
// text it was finished under, and a text collision with another live row is
|
||||
// named instead of merged.
|
||||
//
|
||||
// "confirm" is the only interesting move: it promotes a candidate Maven derived
|
||||
// from something she read into work he owns. That review step is why derived
|
||||
// tasks are captured as candidates in the first place.
|
||||
@@ -965,6 +977,7 @@ func handleTasks(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
|
||||
ID: t.ID, Text: t.Text, Source: t.Source, Evidence: t.Evidence,
|
||||
Status: t.Status, Created: fmtTaskTime(&t.CreatedTs),
|
||||
Due: fmtTaskDate(t.Due), Resolved: fmtTaskTime(t.Resolved),
|
||||
DueValue: fmtTaskDateValue(t.Due), Weight: t.Weight,
|
||||
Why: r.Reason,
|
||||
}
|
||||
if t.Status == "candidate" {
|
||||
@@ -979,11 +992,12 @@ func handleTasks(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
|
||||
w.Header().Set("Content-Type", "text/html; charset=utf-8")
|
||||
if err := tasksTmpl.Execute(w, struct {
|
||||
Msg, Err string
|
||||
Stalls []tasks.Stall
|
||||
Candidates []taskRow
|
||||
Open []taskRow
|
||||
Resolved []taskRow
|
||||
ResolvedMore bool
|
||||
}{msg, errMsg, cands, open, resolved, resolvedTotal > len(resolved)}); err != nil {
|
||||
}{msg, errMsg, tasks.Stalls(live, now()), cands, open, resolved, resolvedTotal > len(resolved)}); err != nil {
|
||||
log.Printf("tasks render: %v", err)
|
||||
}
|
||||
}
|
||||
@@ -999,27 +1013,16 @@ func applyTaskPost(ctx context.Context, core ipc.CoreAPI, r *http.Request) (stri
|
||||
return "", errors.New("empty task text")
|
||||
}
|
||||
req := ipc.CaptureTaskReq{Text: text, Source: "tap:web", Status: "open", Ts: now()}
|
||||
// Importance is his, stated on the form. Out-of-range values are
|
||||
// clamped rather than rejected — a bad select is not worth a 400.
|
||||
if v := r.FormValue("weight"); v != "" {
|
||||
// strconv, not Sscanf: Sscanf("3junk", "%d") succeeds with 3, and a
|
||||
// form value is not a place to accept trailing garbage.
|
||||
wgt, err := strconv.Atoi(v)
|
||||
if err != nil || wgt < 0 {
|
||||
return "", fmt.Errorf("bad weight %q", v)
|
||||
}
|
||||
if wgt > tasks.MaxWeight {
|
||||
wgt = tasks.MaxWeight
|
||||
}
|
||||
req.Weight = wgt
|
||||
wgt, err := formWeight(r)
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
if d := r.FormValue("due"); d != "" {
|
||||
due, err := time.ParseInLocation("2006-01-02", d, now().Location())
|
||||
if err != nil {
|
||||
return "", fmt.Errorf("bad due date %q", d)
|
||||
}
|
||||
req.Due = &due
|
||||
req.Weight = wgt
|
||||
due, err := formDue(r, now())
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
req.Due = due
|
||||
resp, err := core.CaptureTask(ctx, req)
|
||||
if err != nil {
|
||||
return "", err
|
||||
@@ -1037,6 +1040,44 @@ func applyTaskPost(ctx context.Context, core ipc.CoreAPI, r *http.Request) (stri
|
||||
if err != nil {
|
||||
return "", errors.New("invalid id")
|
||||
}
|
||||
|
||||
if action == "promote" {
|
||||
msg, err := promoteCandidate(ctx, core, r, id)
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
return msg, nil
|
||||
}
|
||||
|
||||
if action == "edit" {
|
||||
// The three fields capture set, and only those (Vikunja #509). Status
|
||||
// is not editable here: that ladder is one-way and has its own buttons.
|
||||
text := strings.TrimSpace(r.FormValue("text"))
|
||||
if text == "" {
|
||||
return "", errors.New("empty task text")
|
||||
}
|
||||
wgt, err := formWeight(r)
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
due, err := formDue(r, now())
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
switch err := core.EditTask(ctx, id, text, due, wgt); {
|
||||
case err == nil:
|
||||
return "saved task", nil
|
||||
case errors.Is(err, ipc.ErrTaskDuplicate):
|
||||
// Naming the collision instead of merging: two live rows carry two
|
||||
// provenances, and picking one is not the page's call.
|
||||
return "", errors.New("another open task already says this — drop one of the two")
|
||||
case errors.Is(err, ipc.ErrTaskResolved):
|
||||
return "", errors.New("a resolved task keeps the text it was finished under")
|
||||
default:
|
||||
return "", err
|
||||
}
|
||||
}
|
||||
|
||||
var status, msg string
|
||||
switch action {
|
||||
case "confirm":
|
||||
@@ -1049,11 +1090,134 @@ func applyTaskPost(ctx context.Context, core ipc.CoreAPI, r *http.Request) (stri
|
||||
return "", fmt.Errorf("unknown action %q", action)
|
||||
}
|
||||
if err := core.SetTaskStatus(ctx, id, status, now(), "tap:web"); err != nil {
|
||||
if errors.Is(err, ipc.ErrTaskNoDoneWhen) {
|
||||
// The refusal has to name what is missing, or the button looks
|
||||
// broken. The field it asks for arrives with the intake form
|
||||
// (Vikunja #511).
|
||||
return "", errors.New("write a definition of done before confirming this candidate")
|
||||
}
|
||||
return "", err
|
||||
}
|
||||
return msg, nil
|
||||
}
|
||||
|
||||
// fmtTaskDateValue renders a due date the way <input type=date> requires, or
|
||||
// "" for no date. Separate from fmtTaskDate, which renders it for reading.
|
||||
// promoteCandidate turns a candidate into open work with the three things the
|
||||
// board needs (Vikunja #511): a definition of done, an optional blocker, and an
|
||||
// optional date.
|
||||
//
|
||||
// The definition of done is required, and the refusal is the store's — this
|
||||
// only reaches it in a readable order. The blocker is a NAME here and an entity
|
||||
// id in the row: identity lives in Nexus, so the name is resolved first and a
|
||||
// name Nexus cannot resolve stops the promotion instead of being stored.
|
||||
//
|
||||
// A date set here writes a reminder, which is the one unprompted delivery the
|
||||
// persona allows: he asked to be told, on a day he named.
|
||||
func promoteCandidate(ctx context.Context, core ipc.CoreAPI, r *http.Request, id int64) (string, error) {
|
||||
doneWhen := strings.TrimSpace(r.FormValue("done_when"))
|
||||
if doneWhen == "" {
|
||||
return "", errors.New("write a definition of done — what has to be true for this to be finished")
|
||||
}
|
||||
text := strings.TrimSpace(r.FormValue("text"))
|
||||
if text == "" {
|
||||
return "", errors.New("empty task text")
|
||||
}
|
||||
due, err := formDue(r, now())
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
|
||||
blockedOn := ""
|
||||
if name := strings.TrimSpace(r.FormValue("blocked_on")); name != "" {
|
||||
ref, err := core.ResolveEntity(ctx, name, []string{"person"})
|
||||
switch {
|
||||
case errors.Is(err, ipc.ErrNotImplemented):
|
||||
return "", errors.New("no identity service here, so blocked-on cannot be stored — leave it empty")
|
||||
case errors.Is(err, ipc.ErrNoEntity):
|
||||
return "", fmt.Errorf("nexus does not know %q", name)
|
||||
case err != nil:
|
||||
return "", fmt.Errorf("resolving %q: %w", name, err)
|
||||
case ref.Ambiguous:
|
||||
// Asking, not picking: a task blocked on the wrong person is a
|
||||
// mistake nobody can see afterwards.
|
||||
return "", fmt.Errorf("%q matches %s — say which", name, strings.Join(ref.Candidates, ", "))
|
||||
}
|
||||
blockedOn = ref.ID
|
||||
}
|
||||
|
||||
if err := core.SetTaskFields(ctx, id, doneWhen, blockedOn); err != nil {
|
||||
return "", err
|
||||
}
|
||||
if due != nil {
|
||||
wgt, err := formWeight(r)
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
if err := core.EditTask(ctx, id, text, due, wgt); err != nil {
|
||||
return "", err
|
||||
}
|
||||
}
|
||||
if err := core.SetTaskStatus(ctx, id, "open", now(), "tap:web"); err != nil {
|
||||
if errors.Is(err, ipc.ErrTaskNoDoneWhen) {
|
||||
return "", errors.New("write a definition of done before confirming this candidate")
|
||||
}
|
||||
return "", err
|
||||
}
|
||||
if due == nil {
|
||||
return "confirmed", nil
|
||||
}
|
||||
// A date-only field has no hour. Nine in the morning, because the reminder
|
||||
// is about a day's work and being told at midnight is being told the night
|
||||
// before.
|
||||
fire := time.Date(due.Year(), due.Month(), due.Day(), 9, 0, 0, 0, due.Location())
|
||||
if _, err := core.CreateReminder(ctx, fire, text, ""); err != nil {
|
||||
// The task IS promoted; only the reminder failed. Saying "confirmed"
|
||||
// and nothing else would leave him expecting a nudge that will not come.
|
||||
return "", fmt.Errorf("confirmed, but the reminder did not save: %w", err)
|
||||
}
|
||||
return "confirmed, and maven will remind you that morning", nil
|
||||
}
|
||||
|
||||
// formWeight reads the importance select. Out-of-range clamps rather than
|
||||
// rejects — a bad select is not worth a 400 — but trailing garbage is refused,
|
||||
// because strconv is not Sscanf and "3junk" is not a 3.
|
||||
func formWeight(r *http.Request) (int, error) {
|
||||
v := r.FormValue("weight")
|
||||
if v == "" {
|
||||
return 0, nil
|
||||
}
|
||||
wgt, err := strconv.Atoi(v)
|
||||
if err != nil || wgt < 0 {
|
||||
return 0, fmt.Errorf("bad weight %q", v)
|
||||
}
|
||||
if wgt > tasks.MaxWeight {
|
||||
wgt = tasks.MaxWeight
|
||||
}
|
||||
return wgt, nil
|
||||
}
|
||||
|
||||
// formDue reads the date input. An empty field is nil, which on an edit means
|
||||
// "clear the date" — the form has no other way to say it.
|
||||
func formDue(r *http.Request, now time.Time) (*time.Time, error) {
|
||||
d := r.FormValue("due")
|
||||
if d == "" {
|
||||
return nil, nil
|
||||
}
|
||||
due, err := time.ParseInLocation("2006-01-02", d, now.Location())
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("bad due date %q", d)
|
||||
}
|
||||
return &due, nil
|
||||
}
|
||||
|
||||
func fmtTaskDateValue(t *time.Time) string {
|
||||
if t == nil || t.IsZero() {
|
||||
return ""
|
||||
}
|
||||
return t.Local().Format("2006-01-02")
|
||||
}
|
||||
|
||||
func fmtTaskTime(t *time.Time) string {
|
||||
if t == nil || t.IsZero() {
|
||||
return "—"
|
||||
@@ -1094,34 +1258,53 @@ func handleRoutines(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI, se
|
||||
var msg string
|
||||
if r.Method == http.MethodPost {
|
||||
action := r.FormValue("action")
|
||||
idStr := r.FormValue("id")
|
||||
var rid int64
|
||||
if n, _ := fmt.Sscanf(idStr, "%d", &rid); n != 1 {
|
||||
http.Error(w, "invalid id", http.StatusBadRequest)
|
||||
return
|
||||
}
|
||||
switch action {
|
||||
case "accept":
|
||||
// "seed" is the one action with no routine to act on — it is what
|
||||
// MAKES a routine (Vikunja #518), so it runs before the id parse. It
|
||||
// lives on this route rather than a page of its own because it is
|
||||
// already the step-up-gated surface for this table, and a second gated
|
||||
// surface is a second thing to get wrong.
|
||||
if action == "seed" {
|
||||
if !stepUpOK(session, requireStepUp) {
|
||||
http.Error(w, "step-up required: assert a passkey first", http.StatusForbidden)
|
||||
return
|
||||
}
|
||||
if err := acceptRoutine(ctx, core, rid); err != nil {
|
||||
log.Printf("routines: accept %d: %v", rid, err)
|
||||
http.Error(w, "accept failed: "+err.Error(), http.StatusBadGateway)
|
||||
out, err := seedRoutineEvent(ctx, core, r)
|
||||
if err != nil {
|
||||
log.Printf("routines: seed: %v", err)
|
||||
http.Error(w, "seed failed: "+err.Error(), http.StatusBadGateway)
|
||||
return
|
||||
}
|
||||
msg = "accepted routine — maven will remind you"
|
||||
case "dismiss":
|
||||
if err := core.DismissProposedRoutine(ctx, rid); err != nil {
|
||||
log.Printf("routines: dismiss %d: %v", rid, err)
|
||||
http.Error(w, "dismiss failed: "+err.Error(), http.StatusBadGateway)
|
||||
msg = out
|
||||
} else {
|
||||
idStr := r.FormValue("id")
|
||||
var rid int64
|
||||
if n, _ := fmt.Sscanf(idStr, "%d", &rid); n != 1 {
|
||||
http.Error(w, "invalid id", http.StatusBadRequest)
|
||||
return
|
||||
}
|
||||
switch action {
|
||||
case "accept":
|
||||
if !stepUpOK(session, requireStepUp) {
|
||||
http.Error(w, "step-up required: assert a passkey first", http.StatusForbidden)
|
||||
return
|
||||
}
|
||||
if err := acceptRoutine(ctx, core, rid); err != nil {
|
||||
log.Printf("routines: accept %d: %v", rid, err)
|
||||
http.Error(w, "accept failed: "+err.Error(), http.StatusBadGateway)
|
||||
return
|
||||
}
|
||||
msg = "accepted routine — maven will remind you"
|
||||
case "dismiss":
|
||||
if err := core.DismissProposedRoutine(ctx, rid); err != nil {
|
||||
log.Printf("routines: dismiss %d: %v", rid, err)
|
||||
http.Error(w, "dismiss failed: "+err.Error(), http.StatusBadGateway)
|
||||
return
|
||||
}
|
||||
msg = "dismissed routine"
|
||||
default:
|
||||
http.Error(w, "unknown action", http.StatusBadRequest)
|
||||
return
|
||||
}
|
||||
msg = "dismissed routine"
|
||||
default:
|
||||
http.Error(w, "unknown action", http.StatusBadRequest)
|
||||
return
|
||||
}
|
||||
}
|
||||
proposed, err := core.ListProposedRoutines(ctx)
|
||||
@@ -1158,6 +1341,45 @@ func toRoutineViews(rs []ipc.ProposedRoutine) []routineView {
|
||||
// may do it (Vikunja #367): accepting gives the tick loop a standing new
|
||||
// reason to speak, which DESIGN.md puts at layer 3, and the button here is
|
||||
// behind step-up. Voice can park the question and dismiss, never accept.
|
||||
// seedRoutineEvent drives one backdated fact write through core (Vikunja #518),
|
||||
// so the pattern detector can be exercised against a running daemon instead of
|
||||
// over real days. Refused unless mavend was started with -allow-seed; on an
|
||||
// ordinary box the error says so and nothing is written.
|
||||
//
|
||||
// Takes "ago" rather than an absolute timestamp — hours before now, as a float
|
||||
// so a QA sitting can space four seeds three hours apart without doing clock
|
||||
// arithmetic. The detector's floor is two hours, and "0" is a legal answer
|
||||
// meaning now.
|
||||
func seedRoutineEvent(ctx context.Context, core ipc.CoreAPI, r *http.Request) (string, error) {
|
||||
key := strings.TrimSpace(r.FormValue("key"))
|
||||
value := strings.TrimSpace(r.FormValue("value"))
|
||||
if key == "" || value == "" {
|
||||
return "", errors.New("seed needs a key and a value")
|
||||
}
|
||||
agoHours, err := strconv.ParseFloat(strings.TrimSpace(r.FormValue("ago")), 64)
|
||||
if err != nil {
|
||||
return "", fmt.Errorf("seed: bad ago (hours before now): %w", err)
|
||||
}
|
||||
if agoHours < 0 {
|
||||
return "", errors.New("seed: ago is hours BEFORE now, so it cannot be negative")
|
||||
}
|
||||
resp, err := core.SeedEvent(ctx, ipc.SeedEventReq{
|
||||
Key: key,
|
||||
Value: value,
|
||||
Ts: time.Now().Add(-time.Duration(agoHours * float64(time.Hour))),
|
||||
})
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
if !resp.Extracted {
|
||||
return fmt.Sprintf("wrote fact %d, but %q is not in the action lexicon — no event, no pattern", resp.FactID, value), nil
|
||||
}
|
||||
if !resp.Proposed {
|
||||
return fmt.Sprintf("seeded %s/%s (fact %d, event %d) — not enough yet to propose", resp.Action, resp.Object, resp.FactID, resp.EventID), nil
|
||||
}
|
||||
return fmt.Sprintf("seeded %s/%s and PROPOSED routine %d, every %.1f days", resp.Action, resp.Object, resp.RoutineID, resp.IntervalDays), nil
|
||||
}
|
||||
|
||||
func acceptRoutine(ctx context.Context, core ipc.CoreAPI, id int64) error {
|
||||
proposed, err := core.ListProposedRoutines(ctx)
|
||||
if err != nil {
|
||||
@@ -1193,12 +1415,29 @@ func handleTrace(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
|
||||
http.Error(w, "core read failed", http.StatusBadGateway)
|
||||
return
|
||||
}
|
||||
// The turn records share this page rather than getting one of their own
|
||||
// (V-564): both answer the same question — who won, who lost and why — and
|
||||
// one is about nudges while the other is about utterances. A read failure
|
||||
// here is not fatal to the page: the rule trace above it still renders, and
|
||||
// a daemon too old to know the method is the ordinary case during a rolling
|
||||
// deploy.
|
||||
turns, err := core.TurnDecisions(ctx, 25)
|
||||
if err != nil {
|
||||
log.Printf("trace: turn decisions: %v", err)
|
||||
}
|
||||
w.Header().Set("Content-Type", "text/html; charset=utf-8")
|
||||
if err := traceTmpl.Execute(w, trace); err != nil {
|
||||
if err := traceTmpl.Execute(w, traceData{Tick: trace, Turns: turns}); err != nil {
|
||||
log.Printf("trace render: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// traceData — what trace.html renders: the last tick's rule arbitration and the
|
||||
// last turns' claim arbitration.
|
||||
type traceData struct {
|
||||
Tick ipc.TickTrace
|
||||
Turns []ipc.TurnDecision
|
||||
}
|
||||
|
||||
func handleMorning(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
|
||||
if core == nil {
|
||||
http.Error(w, "morning disabled (no -core)", http.StatusServiceUnavailable)
|
||||
@@ -1540,7 +1779,13 @@ func handlePTT(w http.ResponseWriter, r *http.Request, voiceAddr string, session
|
||||
return
|
||||
}
|
||||
w.Header().Set("Content-Type", "audio/l16;rate=16000;channels=1")
|
||||
w.Header().Set("X-Reply-Text", url.QueryEscape(pttResp.ReplyText))
|
||||
// PathEscape, not QueryEscape (Vikunja #533). QueryEscape writes a space
|
||||
// as "+", which is form encoding, and the client decodes this header
|
||||
// with decodeURIComponent, which only knows "%20" — so every space in a
|
||||
// spoken reply reached the on-page log as a plus sign. PathEscape is the
|
||||
// flavour decodeURIComponent actually reverses, which keeps the encoding
|
||||
// a property of the header rather than something the client has to know.
|
||||
w.Header().Set("X-Reply-Text", url.PathEscape(pttResp.ReplyText))
|
||||
w.Write(pttResp.ReplyAudio.Bytes)
|
||||
return
|
||||
}
|
||||
@@ -1563,8 +1808,8 @@ func handleChatPage(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
|
||||
if q := r.URL.Query().Get("q"); q != "" {
|
||||
msgs = append(msgs, chatMsg{Role: "user", Text: q})
|
||||
}
|
||||
if r := r.URL.Query().Get("r"); r != "" {
|
||||
msgs = append(msgs, chatMsg{Role: "assistant", Text: r})
|
||||
if reply := r.URL.Query().Get("r"); reply != "" {
|
||||
msgs = append(msgs, chatMsg{Role: "assistant", Text: reply, Source: r.URL.Query().Get("s")})
|
||||
}
|
||||
w.Header().Set("Content-Type", "text/html; charset=utf-8")
|
||||
if err := chatTmpl.Execute(w, struct {
|
||||
@@ -1613,13 +1858,22 @@ func handleChatAPI(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI, ses
|
||||
http.Redirect(w, r, "/chat", http.StatusSeeOther)
|
||||
return
|
||||
}
|
||||
http.Redirect(w, r, "/chat?q="+url.QueryEscape(text)+"&r="+url.QueryEscape(reply), http.StatusSeeOther)
|
||||
// The claiming query source rides back on the redirect so the page can show
|
||||
// it. Empty for a turn no source claimed, which is most of them.
|
||||
dest := "/chat?q=" + url.QueryEscape(text) + "&r=" + url.QueryEscape(reply.Reply)
|
||||
if reply.Source != "" {
|
||||
dest += "&s=" + url.QueryEscape(reply.Source)
|
||||
}
|
||||
http.Redirect(w, r, dest, http.StatusSeeOther)
|
||||
}
|
||||
|
||||
// chatMsg — one message in the conversation history.
|
||||
type chatMsg struct {
|
||||
Role string // "user" | "assistant"
|
||||
Text string
|
||||
// Source — the query source that claimed the turn, shown as a badge beside
|
||||
// the reply. Empty for a turn no source claimed (V-539).
|
||||
Source string
|
||||
}
|
||||
|
||||
func mustMarshal(v any) json.RawMessage {
|
||||
|
||||
@@ -0,0 +1,59 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"net/url"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// decodeURIComponent is what static/app.js calls on X-Reply-Text. PathUnescape
|
||||
// is its Go equivalent for this purpose: both turn %XX into bytes and both
|
||||
// leave a literal "+" alone. That last part is the whole defect — QueryEscape
|
||||
// wrote spaces as "+" and the client had no way to tell those from a plus the
|
||||
// speaker actually said.
|
||||
func decodeURIComponent(t *testing.T, s string) string {
|
||||
t.Helper()
|
||||
out, err := url.PathUnescape(s)
|
||||
if err != nil {
|
||||
t.Fatalf("decodeURIComponent(%q): %v", s, err)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// The reply the QA session actually saw was "на+04.08.2026+ничего+нет."
|
||||
// (Vikunja #533). Round-tripping through the client's decoder is the assertion
|
||||
// that matters — checking the encoder in isolation would have passed with
|
||||
// QueryEscape too.
|
||||
func TestReplyTextSurvivesTheClientDecoder(t *testing.T) {
|
||||
cases := []string{
|
||||
"на 04.08.2026 ничего нет.",
|
||||
"Я поставила тебе напоминание позвонить маме через час.",
|
||||
// A literal plus must stay a plus, which is the case that makes
|
||||
// "just replace + with space on the JS side" the wrong fix.
|
||||
"два плюс два = 2+2",
|
||||
// Headers cannot carry a raw newline. PathEscape writes %0A.
|
||||
"первая строка\nвторая строка",
|
||||
"", // no reply text at all
|
||||
}
|
||||
for _, want := range cases {
|
||||
encoded := url.PathEscape(want)
|
||||
if strings.ContainsAny(encoded, "\r\n") {
|
||||
t.Errorf("encoded %q contains a raw newline, which is not a legal header value", want)
|
||||
}
|
||||
if got := decodeURIComponent(t, encoded); got != want {
|
||||
t.Errorf("round trip: got %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The specific regression, named. QueryEscape is form encoding and this header
|
||||
// is not a form.
|
||||
func TestReplyTextDoesNotUseFormEncoding(t *testing.T) {
|
||||
const spoken = "на 04.08.2026 ничего нет."
|
||||
if got := decodeURIComponent(t, url.QueryEscape(spoken)); got == spoken {
|
||||
t.Skip("QueryEscape round-trips here, so this test proves nothing — check the decoder stand-in")
|
||||
}
|
||||
if strings.Contains(url.PathEscape(spoken), "+") {
|
||||
t.Errorf("PathEscape(%q) still writes a plus", spoken)
|
||||
}
|
||||
}
|
||||
+44
-6
@@ -21,20 +21,43 @@
|
||||
</form>
|
||||
</section>
|
||||
|
||||
{{if .Stalls}}
|
||||
<section class=card>
|
||||
<h2 class=card-title>shapes</h2>
|
||||
<!-- Counts, and nothing about what they mean (V-512). Whether a task should be
|
||||
dropped is his call and Maven does not have an opinion to show here. -->
|
||||
<div class=scroll><table>
|
||||
<tr><th>count</th><th>shape</th></tr>
|
||||
{{range .Stalls}}<tr><td>{{.N}}</td><td>{{.Line}}</td></tr>{{end}}
|
||||
</table></div>
|
||||
</section>
|
||||
{{end}}
|
||||
|
||||
{{if .Candidates}}
|
||||
<section class=card>
|
||||
<h2 class=card-title>found, not confirmed <span class=badge>{{len .Candidates}}</span></h2>
|
||||
<div class=hint>maven derived these from something she read. nothing counts as your work until you confirm it.</div>
|
||||
<div class=hint>confirming asks for a definition of done: what has to be true for this to be finished. a task without one can never leave the board. a date here also books a reminder that morning.</div>
|
||||
<div class=scroll><table>
|
||||
<tr><th>task</th><th>where from</th><th>due</th><th>captured</th><th></th><th></th></tr>
|
||||
<tr><th>task</th><th>where from</th><th>captured</th><th>confirm</th><th></th></tr>
|
||||
{{range .Candidates}}<tr>
|
||||
<td class=text-max>{{.Text}}</td>
|
||||
<td class=hint>{{.Source}}{{if .Evidence}} — {{.Evidence}}{{end}}</td>
|
||||
<td>{{.Due}}</td>
|
||||
<td class=muted>{{.Created}}</td>
|
||||
<td><form method=post action=/tasks class=inline-form>
|
||||
<input type=hidden name=id value="{{.ID}}">
|
||||
<input type=hidden name=action value=confirm>
|
||||
<input type=hidden name=action value=promote>
|
||||
<input type=hidden name=text value="{{.Text}}">
|
||||
<input type=text name=done_when placeholder="готово, когда…" size=26 required>
|
||||
<!-- A name, not an id. It is resolved against nexus before anything is
|
||||
stored, and a name nexus cannot place stops the confirmation. -->
|
||||
<input type=text name=blocked_on placeholder="ждёт кого-то" size=14>
|
||||
<input type=date name=due value="{{.DueValue}}" title="due date">
|
||||
<select name=weight title=importance>
|
||||
<option value=0 {{if eq .Weight 0}}selected{{end}}>normal</option>
|
||||
<option value=2 {{if eq .Weight 2}}selected{{end}}>важно</option>
|
||||
<option value=3 {{if eq .Weight 3}}selected{{end}}>срочно</option>
|
||||
</select>
|
||||
<button class=btn>confirm</button></form></td>
|
||||
<td><form method=post action=/tasks class=inline-form>
|
||||
<input type=hidden name=id value="{{.ID}}">
|
||||
@@ -48,12 +71,27 @@
|
||||
<h2 class=card-title>open <span class=badge>{{len .Open}}</span></h2>
|
||||
<div class=hint>most pressing first — by the deadlines and the urgency you gave. nothing about a task is guessed; the only signal that is not yours is age, which lifts anything sitting here for weeks.</div>
|
||||
{{if .Open}}<div class=scroll><table>
|
||||
<tr><th>task</th><th>why</th><th>from</th><th>due</th><th>captured</th><th></th><th></th></tr>
|
||||
<tr><th>task</th><th>why</th><th>from</th><th>captured</th><th></th><th></th></tr>
|
||||
{{range .Open}}<tr>
|
||||
<td class=text-max>{{.Text}}</td>
|
||||
<!-- The text, the date and the importance are editable in place (V-509): a
|
||||
dictated task can carry a typo, and a deadline moves. The status is not
|
||||
here — that ladder is one-way and has its own two buttons. -->
|
||||
<td class=text-max><form method=post action=/tasks class=inline-form>
|
||||
<input type=hidden name=id value="{{.ID}}">
|
||||
<input type=hidden name=action value=edit>
|
||||
<input type=text name=text value="{{.Text}}" size=30 required>
|
||||
<input type=date name=due value="{{.DueValue}}" title="due date">
|
||||
<select name=weight title=importance>
|
||||
<!-- Any weight that is not one of the three rungs keeps its own option, or
|
||||
saving an unrelated edit would silently reset it to normal. -->
|
||||
{{if and (ne .Weight 0) (ne .Weight 2) (ne .Weight 3)}}<option value={{.Weight}} selected>{{.Weight}}</option>{{end}}
|
||||
<option value=0 {{if eq .Weight 0}}selected{{end}}>normal</option>
|
||||
<option value=2 {{if eq .Weight 2}}selected{{end}}>важно</option>
|
||||
<option value=3 {{if eq .Weight 3}}selected{{end}}>срочно</option>
|
||||
</select>
|
||||
<button class="btn btn-muted">save</button></form></td>
|
||||
<td class=hint>{{.Why}}</td>
|
||||
<td class=hint>{{.Source}}</td>
|
||||
<td>{{.Due}}</td>
|
||||
<td class=muted>{{.Created}}</td>
|
||||
<td><form method=post action=/tasks class=inline-form>
|
||||
<input type=hidden name=id value="{{.ID}}">
|
||||
|
||||
@@ -77,10 +77,14 @@ func TestHandleTasksSplitsCandidatesFromOpen(t *testing.T) {
|
||||
t.Errorf("body missing %q", want)
|
||||
}
|
||||
}
|
||||
// The candidate must offer confirm, and the open task must not.
|
||||
if !strings.Contains(body, "value=confirm") {
|
||||
// The candidate must offer the intake form, and it asks for a definition of
|
||||
// done before it will confirm anything (V-511).
|
||||
if !strings.Contains(body, "value=promote") {
|
||||
t.Error("candidate row has no confirm action")
|
||||
}
|
||||
if !strings.Contains(body, "name=done_when") {
|
||||
t.Error("the confirm form does not ask for a definition of done")
|
||||
}
|
||||
}
|
||||
|
||||
func TestHandleTasksAddCaptures(t *testing.T) {
|
||||
|
||||
+22
-2
@@ -1,9 +1,9 @@
|
||||
{{template "shellTop" "trace"}}
|
||||
<h1>Rule Trace</h1>
|
||||
<div class="hint mb-4">{{.Now | ago}} — winner: <strong>{{if .Winner}}{{.Winner}}{{else}}nothing fired{{end}}</strong></div>
|
||||
<div class="hint mb-4">{{.Tick.Now | ago}} — winner: <strong>{{if .Tick.Winner}}{{.Tick.Winner}}{{else}}nothing fired{{end}}</strong></div>
|
||||
<div class=scroll><table class=mono>
|
||||
<tr><th>rule<th>sev<th>predicate<th>gate<th>blocked by<th>detail<th>selected<th>lost to</tr>
|
||||
{{range .Rules}}<tr>
|
||||
{{range .Tick.Rules}}<tr>
|
||||
<td>{{.RuleName}}</td>
|
||||
<td>{{.Severity}}</td>
|
||||
<td class={{if .PredicateResult}}green{{else}}gray{{end}}>{{.PredicateResult}}</td>
|
||||
@@ -21,5 +21,25 @@
|
||||
<td>{{.LostTo}}</td>
|
||||
</tr>{{end}}
|
||||
</table></div>
|
||||
|
||||
<h1 class=mt-4>Turn Decisions</h1>
|
||||
<div class="hint mb-4">Who claimed each utterance, who lost it, and who was never asked. In memory, newest first, cleared on restart.</div>
|
||||
{{if not .Turns}}<div class=hint>no turn has run since the daemon started</div>{{end}}
|
||||
{{range .Turns}}
|
||||
<details class=mb-4>
|
||||
<summary><span class=mono>{{.Utterance}}</span> — <strong>{{if .Winner}}{{.Winner}}{{else}}nobody{{end}}</strong> <span class=hint>{{.Ts | ago}}</span></summary>
|
||||
<div class=scroll><table class=mono>
|
||||
<tr><th>stage<th>claimant<th>would have been<th>score<th>outcome<th>why</tr>
|
||||
{{range .Claims}}<tr>
|
||||
<td>{{.Stage}}</td>
|
||||
<td>{{.Claimant}}</td>
|
||||
<td>{{if .Intent}}{{.Intent}}{{else}}—{{end}}</td>
|
||||
<td>{{if .HasScore}}{{printf "%.3f" .Score}}{{else}}—{{end}}</td>
|
||||
<td class={{if eq .Outcome "won"}}green{{else if eq .Outcome "never_asked"}}red{{else}}gray{{end}}>{{.Outcome}}</td>
|
||||
<td>{{.Reason}}</td>
|
||||
</tr>{{end}}
|
||||
</table></div>
|
||||
</details>
|
||||
{{end}}
|
||||
{{template "shellBottom"}}
|
||||
</html>
|
||||
|
||||
@@ -90,6 +90,7 @@
|
||||
"kiwix": {
|
||||
"url": "http://kiwix-server:8080",
|
||||
"book": "wikipedia_en_all_maxi_2026-02",
|
||||
"book_ru": "wikipedia_ru_all_maxi_2026-02",
|
||||
"max_results": 5,
|
||||
"snippet_runes": 1500
|
||||
},
|
||||
|
||||
+6
-1
@@ -12,7 +12,12 @@ x-image: &image
|
||||
# build on EVERY service (same image name ⇒ built once) so `docker compose
|
||||
# build <anyservice>` actually rebuilds. With build on only one service, the
|
||||
# others silently no-op and you deploy a stale binary.
|
||||
build: .
|
||||
build:
|
||||
context: .
|
||||
# the zone is declared once, here. The image points /etc/localtime at it so
|
||||
# a caller reading the system zone agrees with one reading TZ (V-545).
|
||||
args:
|
||||
TZ: Europe/Samara
|
||||
pull_policy: never # only ever the locally-built image
|
||||
restart: unless-stopped
|
||||
# local time for clock/date replies AND quiet-hours evaluation. Change to
|
||||
|
||||
@@ -0,0 +1,93 @@
|
||||
# Routing from audio: four paths, one fixture
|
||||
|
||||
**05-08-2026. Vikunja #486.** Workstation `gemma-4-12B-it-qat-UD-Q4_K_XL` with
|
||||
`mmproj-F16.gguf`, homesrv whisper `ggml-small`, piper `ru_RU-irina-medium`.
|
||||
|
||||
**Verdict: transcribe, then route.** One call from audio straight to a route loses 36
|
||||
points, so it is not a candidate. Moving speech-to-text to the workstation buys 375ms and
|
||||
better transcripts at no measurable accuracy cost. So #486 proceeds on the two-call shape.
|
||||
|
||||
## The numbers
|
||||
|
||||
72 Russian cases from `internal/router/eval/ru_routing_v1.json`, rendered by piper at
|
||||
16kHz mono, 153.9s of audio, mean 2.14s per clip. Every path used the daemon's own
|
||||
`routeSystem` prompt and `routeGrammar`, read out of `internal/router/llmrouter.go` at run
|
||||
time, at `temperature 0` and `enable_thinking:false`.
|
||||
|
||||
| Path | Intent-only | Verbatim transcripts | p50 | p95 |
|
||||
|---|---|---|---|---|
|
||||
| text in, the ceiling | **90.3%** (65/72) | — | 361ms | 495ms |
|
||||
| whisper on homesrv, then route | **84.7%** (61/72) | 29/72 | 1372ms | 1546ms |
|
||||
| workstation transcribes, then routes | **83.3%** (60/72) | 48/72 | 997ms | 1177ms |
|
||||
| workstation, one call from audio | **54.2%** (39/72) | — | 425ms | 756ms |
|
||||
|
||||
The 90.3% ceiling is the same model on the same 72 cases with the utterance as text. It is
|
||||
not the 93.5% in `docs/evals/2026-08-02-workstation-gemma4-12b.md`, which scored all 87
|
||||
cases including the English ones.
|
||||
|
||||
The two speech-to-text paths differ by one case, which is noise on 72. So the choice
|
||||
between them is latency and transcript quality, and the workstation wins both.
|
||||
|
||||
## One call from audio is not a transcription failure
|
||||
|
||||
The obvious reading of 54.2% is that the audio encoder cannot hear Russian. It can. Eight
|
||||
of the failing clips were sent back with a transcribe instruction instead of the router
|
||||
prompt:
|
||||
|
||||
| Clip | Said | Heard, transcribing | Routing from audio |
|
||||
|---|---|---|---|
|
||||
| ru-sys-002 | какое число завтра | Какое число завтра? | `unknown` |
|
||||
| ru-sys-003 | переходи в тихий режим | Переходи в тихий режим. | `unknown` |
|
||||
| ru-query-001 | сколько воды я выпил с утра | Сколько воды я выпил с утра? | `fact`, value "выпил с утра" |
|
||||
| ru-act-002 | выключи свет в спальне | Выключи свет в спальне. | `fact`, value "включен" |
|
||||
|
||||
Four clips it transcribes word for word, and routes wrong or refuses. The `ru-query-001`
|
||||
row shows the mechanism: the emitted slot holds the tail of the sentence and the
|
||||
interrogative head is gone. The model is not deaf, it stops attending to the audio once it
|
||||
is also holding a 3.5k-character classification prompt.
|
||||
|
||||
That pattern decides the whole task. A long system prompt and an audio part compete, so the
|
||||
transcription has to be its own call with a short instruction. It also means the number
|
||||
would not be rescued by a better prompt, a longer clip, or a bigger `mmproj`.
|
||||
|
||||
The failures cluster where the head of the sentence carries the intent: `ru-act` 1/6,
|
||||
`ru-sys` 2/5, `ru-query` 12/25. Reminders scored 10/10, because "напомни" is the first word
|
||||
and nothing after it changes the answer.
|
||||
|
||||
## Transcript quality and routing accuracy come apart
|
||||
|
||||
The workstation transcribes 48 of 72 verbatim against whisper's 29, and routes one case
|
||||
worse. Both directions of that appear in the same run:
|
||||
|
||||
- `ru-chat-002`: whisper heard "Кто думаешь про переезд", the workstation heard "Что ты
|
||||
думаешь про переезд". The correct transcript routed to `chat`, the broken one to `query`.
|
||||
- `ru-query-020`: whisper heard "Кто дальше?", the workstation heard the correct "Что
|
||||
дальше?". The **broken** transcript routed correctly and the correct one missed.
|
||||
|
||||
A word error rate is not a proxy for routing accuracy here. Judge a speech-to-text change
|
||||
on the routing fixture, not on transcripts.
|
||||
|
||||
Three cases only the text path gets right. No speech-to-text path recovers them, so they
|
||||
are lost in the rendering rather than in the model.
|
||||
|
||||
## Latency
|
||||
|
||||
Whisper `ggml-small` on homesrv CPU costs p50 998ms for a 2.14s clip, which is nearly
|
||||
all of that path's 1372ms. The workstation does the same job inside its 997ms end-to-end
|
||||
total for two calls. So the transfer is worth about 375ms per turn at p50, and more at p95.
|
||||
|
||||
Both are above the one-call 425ms, and that is the trade the table settles: 29 points of
|
||||
accuracy for 572ms.
|
||||
|
||||
## Notes for the next run
|
||||
|
||||
- `--mmproj /mnt/D/AI/gemma4/mmproj-F16.gguf` has to be in `llama_args` in
|
||||
`~/.config/mavgpud.json`, or `/props` reports `modalities.audio: false` and every audio
|
||||
part is dropped silently. It was added for this measurement and removed afterwards, so
|
||||
the box is back to the text-only config.
|
||||
- `enable_thinking:false` is mandatory. It was set for all 224 calls here.
|
||||
- The degenerate `<|channel>thought` output recorded against #486 did not reproduce, in 80
|
||||
transcribe calls or in 144 routing calls.
|
||||
- Piper renders at 22050Hz mono. Every clip was resampled with
|
||||
`ffmpeg -ar 16000 -ac 1 -c:a pcm_s16le`, because 16kHz is what `audio.PCM16kMono`
|
||||
declares and what the earlier measurement used.
|
||||
@@ -0,0 +1,43 @@
|
||||
# Half-past and quarter-to hours, 2026-08-05
|
||||
|
||||
Vikunja V-538. `rewriteHalfPast` in `internal/router/halfpast.go`, run in front of
|
||||
the token pass inside `SpellOutDigits`, so both date parsers see digits.
|
||||
|
||||
## What the shapes are
|
||||
|
||||
Russian names a half hour by the hour being ENTERED, in the genitive. "половина
|
||||
восьмого" is 07:30. "без четверти восемь" counts the other way, from a cardinal,
|
||||
and is 07:45. Both are minus one from the word in the sentence, and the arithmetic
|
||||
lives in one function, `clockHourBefore`.
|
||||
|
||||
## Result
|
||||
|
||||
| | before | after |
|
||||
|---|---|---|
|
||||
| classifier + onnx over the routing fixture | 58/82 (70.7%) | 62/87 (71.3%) |
|
||||
| new fixture cases passing | — | 2 of 3 |
|
||||
| stub parser reads a half hour | no | yes |
|
||||
|
||||
The three new cases are ru-rem-008, ru-rem-009 and ru-rem-010. No existing case
|
||||
regressed and no new clarify appeared.
|
||||
|
||||
ru-rem-009, "разбуди меня полвосьмого", still misses the intent. Its two siblings
|
||||
without a half hour miss it the same way. ru-rem-005 "разбуди меня в 6:30" routes
|
||||
to `fact`, and en-rem-002 "wake me at 6:15" does too. So the miss is the "разбуди"
|
||||
phrasing against the classifier, not the half hour. The time slot now fills.
|
||||
|
||||
## Not measured here
|
||||
|
||||
Python dateparser. It is not installed on this host, so only the stub was run.
|
||||
The rewrite emits "в 7:30 вечера". The script's own qualifier rewrite turns that
|
||||
trailing "вечера" into "pm", which is the shape it already reads for a whole hour.
|
||||
Judge it on the box.
|
||||
|
||||
The LLM arm. No llama-server in this run, so the cascade number is the classifier
|
||||
floor.
|
||||
|
||||
## Left out on purpose
|
||||
|
||||
Minutes a spoken clock does not use. "без семи восемь" is not rewritten, because
|
||||
nobody says it and a guess in this shape is a missed dose. The parsers fail on it
|
||||
as they did before.
|
||||
@@ -0,0 +1,65 @@
|
||||
# Does the ZIM answer when the line is down? (V-508)
|
||||
|
||||
Measured 2026-08-05 on the deploy, through `POST /api/chat`. The question was
|
||||
`что такое фотосинтез` in every run. Which source claimed is read off
|
||||
`voice: query claimed by source` and off the badge V-539 added.
|
||||
|
||||
## It fires, and it is fast when the host is gone
|
||||
|
||||
`docker stop searxng`, then one question:
|
||||
|
||||
| | Claimed by | Turn |
|
||||
|---|---|---|
|
||||
| Search reachable | search | 3.5 s |
|
||||
| Container stopped | kiwix | 3.5 s |
|
||||
| Host blackholed | kiwix | 15.4 s |
|
||||
|
||||
With the container stopped, DNS failed and the ZIM answered inside the same
|
||||
second:
|
||||
|
||||
```
|
||||
15:40:51 voice: search "что такое фотосинтез": ... lookup searxng: no such host
|
||||
15:40:51 voice: kiwix: "photosynthesis" → 5 hits, top "Photosynthesis"
|
||||
15:40:53 voice: query claimed by source "kiwix"
|
||||
```
|
||||
|
||||
The rewrite, the search and the reply all fit in the same turn budget as a live
|
||||
search. The fallback works.
|
||||
|
||||
## The blackhole is the case that hurts
|
||||
|
||||
192.0.2.1 is reserved and routed nowhere. Pointing `search.url` at it is the
|
||||
shape of a real outage: the router drops the packet instead of refusing it. The
|
||||
search sat for its full 8-second budget before the ZIM was asked. The turn took
|
||||
15.4 seconds against 3.5. He waits through all of it with nothing
|
||||
being said.
|
||||
|
||||
Fixed by capping the connect phase alone at 1.5 s (`dialTimeout` in
|
||||
`internal/websearch/searxng.go`). The instance is on the LAN, so a connection it
|
||||
will ever accept is accepted in milliseconds. A reachable instance that is
|
||||
merely slow still gets the whole 8 seconds. It is fanning out to real engines,
|
||||
which is worth waiting for.
|
||||
|
||||
## The Russian ZIM is now on the box and is read directly
|
||||
|
||||
`wikipedia_ru_all_maxi_2026-02` (41 GB) was copied to the kiwix zims directory
|
||||
and kiwix-serve picked it up. Note that the catalog name is derived from the
|
||||
filename. `books.name=wikipedia_ru_all_maxi_2026-02` returns Фотосинтез,
|
||||
С4-фотосинтез and Википедия. The `<name>` field in the catalog says
|
||||
`wikipedia_ru_all`, which returns nothing.
|
||||
|
||||
A Cyrillic question now searches that book verbatim (`book_ru` in the `kiwix`
|
||||
block). The rewriter was never a feature. An English ZIM cannot match a Russian
|
||||
sentence, so the resident model translated the question into English keywords
|
||||
first. That costs a model call. It also drops whatever the keywords do not carry.
|
||||
Against a Russian book it is a translation of his own words back at him.
|
||||
|
||||
## Not measured here
|
||||
|
||||
- The Russian book answering a driven turn. The `book_ru` field is a binary
|
||||
change, so it needs a rebuild the owner runs. The book itself was verified by
|
||||
querying kiwix-serve directly.
|
||||
- Recall against the Russian book compared with the rewrite path. Reading his
|
||||
own language directly should win, and it was not scored.
|
||||
- `ru.stackoverflow.com_mul_all_2026-02.zim` is still in the staging directory
|
||||
and is wired to nothing.
|
||||
@@ -0,0 +1,63 @@
|
||||
# Praxis reach at stage 0, 2026-08-05
|
||||
|
||||
Vikunja #516. Measured with `make eval-reach` on the held-out ecosystem fixture
|
||||
(`internal/router/eval/ru_ecosystem_v1.json`, 30 cases), classifier + ONNX embedder,
|
||||
no llama-server in the run. The LLM arm was not measured, so judge a cascade
|
||||
number again before quoting one.
|
||||
|
||||
## Result
|
||||
|
||||
| | before | after |
|
||||
|---|---|---|
|
||||
| overall | 16/30 (53.3%) | 27/30 (90.0%) |
|
||||
| by want: praxis | 0/12 | 11/12 |
|
||||
| by want: hexis | 9/10 | 9/10 |
|
||||
| by want: none | 7/8 | 7/8 |
|
||||
| by tag: lifecycle | 0/5 | 5/5 |
|
||||
| by tag: attention | 0/7 | 6/7 |
|
||||
| by tag: reading | 0/7 | 6/7 |
|
||||
| wrong praxis arm | 0 | 0 |
|
||||
| p50 latency | 20.6ms | 16.5ms |
|
||||
|
||||
## Why it was zero
|
||||
|
||||
Not a tuning gap. `handlePraxisAct` dispatches on exact equality between
|
||||
`Slots.Fn` and a capability alias, and the fn slot is filled by `DefaultActMatcher`
|
||||
from the deployment's enabled tool names. No Praxis alias is on that list, so no
|
||||
utterance could put one in the slot. The Russian aliases in `praxisCapabilities`
|
||||
read as if they matched speech. They are compared against a fn slot and never
|
||||
against an utterance.
|
||||
|
||||
`PraxisGrammars()` (`internal/router/praxis.go`) fills the slot at stage 0, wired in
|
||||
`buildRouter` before the capture marker because "отметь" is a capture verb.
|
||||
|
||||
## The three misses that remain
|
||||
|
||||
- `eco-ru-006` "запусти бэкап на нексусе", a Hexis case, routed note. Pre-existing.
|
||||
- `eco-ru-028` "выключи", reached Hexis, should have asked. Pre-existing.
|
||||
- `eco-ru-021` "что там с нексусом" wants scoped attention. Deliberately not
|
||||
claimed. "что там с X" also opens "что там с погодой". Routing a weather
|
||||
question to Nexus is worse than one missed fixture case.
|
||||
|
||||
## Two judgement calls worth re-arguing
|
||||
|
||||
**A lifecycle word alone does not transition an item.** "готово" is what he says
|
||||
about the thing he just finished. So the rules split lifecycle words by mood. An
|
||||
imperative he says to her ("закрывай") claims the turn bare, and the capability
|
||||
asks which пункт. A stative ("готово", "принято") needs an item named beside it.
|
||||
|
||||
The bare-imperative arm also requires that nothing else in the sentence is being
|
||||
acted on. "закрой шторы в комнате" is an imperative too. Without that guard it took
|
||||
a house command to Praxis, measured at hexis 8/10 mid-change.
|
||||
|
||||
**A demonstrative resolves only against a one-item digest.** "отметь это как
|
||||
сделанное" points at what she just read. `resolveSurfacedPosition` maps it to an id
|
||||
only when exactly one item was spoken. With two or more it gives the turn back to
|
||||
the cascade rather than transitioning one of them at random. With no digest at all
|
||||
it gives the turn back too, because "я это сделал" was never about a пункт.
|
||||
|
||||
## Routing fixture
|
||||
|
||||
`make eval-router`, same run: classifier + ONNX 60/84 (71.4% full and intent-only),
|
||||
0 false clarifies, 6 missed clarifies (the known `amb-*` set). No failure in that
|
||||
list comes from a stage-0 decision. Every one carries a classifier confidence score.
|
||||
@@ -0,0 +1,70 @@
|
||||
# Ecosystem reach with the resident model as router
|
||||
|
||||
Date: 2026-08-05. Vikunja #517, split out of #405.
|
||||
Fixture: `internal/router/eval/ru_ecosystem_v1.json`, 30 held-out Russian cases.
|
||||
Model: Qwen3-1.7B-UD-Q4_K_XL, llama-server on the host at 127.0.0.1:8899.
|
||||
Harness: `TestReachWithLLMRouter` in `internal/router/eval/llmrouter_test.go`.
|
||||
|
||||
## The numbers
|
||||
|
||||
| configuration | reached the right place | praxis | hexis | none |
|
||||
|---|---|---|---|---|
|
||||
| classifier + hash (V-405 floor) | 16/30 | 0/12 | — | — |
|
||||
| classifier + ONNX, after the V-516 grammars | 27/30 | 11/12 | — | — |
|
||||
| **llm-only** (resident model alone) | **17/30 (56.7%)** | **0/12** | 10/10 | 7/8 |
|
||||
| **cascade + llm** + hash fallback | **28/30 (93.3%)** | **11/12** | 10/10 | 7/8 |
|
||||
|
||||
Latency: llm-only p50 1.29s, p95 1.65s. Cascade p50 1.11s, p95 1.64s.
|
||||
No case errored in either configuration.
|
||||
|
||||
## The open question is answered: the model never reaches Praxis
|
||||
|
||||
The route grammar lets the model write any string into the `fn` slot. So it
|
||||
could in principle emit a literal Praxis capability name, and reach a service
|
||||
the classifier structurally cannot. It does not. **Praxis is 0/12 with the
|
||||
model alone.** That is exactly what the classifier alone scores. Every one of
|
||||
the twelve fails the same way: the utterance stays local with an empty `fn`.
|
||||
|
||||
So the stage-0 Praxis grammars from V-516 are not a determinism argument. They
|
||||
are the only path to Praxis that exists. Deleting them takes reach from 11/12
|
||||
back to 0/12 whichever engine is answering.
|
||||
|
||||
The failure is not that the model routes these badly in its own terms. It
|
||||
spreads them across `query`, `fact`, `system` and `chat`. Those are reasonable
|
||||
readings of "что требует внимания" and "готово, закрывай" for a model that has
|
||||
never been told Praxis exists. Nothing in the prompt names a Praxis capability,
|
||||
so there is no string for it to write.
|
||||
|
||||
## What the model does buy
|
||||
|
||||
Hexis is 10/10 with the model alone, and the mutating tag is 10/15 llm-only
|
||||
against 15/15 through the cascade. The model reaches everything Hexis owns
|
||||
without help, which is the half the act allowlist already names in the prompt.
|
||||
|
||||
Cascade + llm scores one point above the classifier baseline: 28/30 against
|
||||
27/30, the difference being one attention case. That is the same shape as the
|
||||
routing fixture, where the router buys about 4 points rather than a doubling.
|
||||
|
||||
## The two that still miss
|
||||
|
||||
- `eco-ru-021 "что там с нексусом"`. Routes `query`, stays local, wants Praxis.
|
||||
Asking after a named service reads as a question about a thing, and no
|
||||
grammar claims a service name.
|
||||
- `eco-ru-029 "сделай это"`. Routes `act` and reaches Hexis. The fixture wants
|
||||
nothing reached, because "это" names no target. This is the overreach case
|
||||
and it is the one direction worth failing on. The confirmation binding
|
||||
downstream still resolves a canonical entity id before anything executes.
|
||||
The fixture is right that the turn should have asked.
|
||||
|
||||
Overreach is 1 in both configurations, under the 4 the harness asserts.
|
||||
|
||||
## How to re-run
|
||||
|
||||
```sh
|
||||
MAVEN_LLM_URL=http://127.0.0.1:8899 \
|
||||
deps/go/go/bin/go test -v -count=1 -timeout 40m \
|
||||
-run TestReachWithLLMRouter ./internal/router/eval/
|
||||
```
|
||||
|
||||
The host `http_proxy` answers 503 for 127.0.0.1. `noProxyLoopback` in the test
|
||||
excludes it. A run that scores every case as a route error measured the proxy.
|
||||
@@ -0,0 +1,69 @@
|
||||
# Routing with the resident model, re-measured
|
||||
|
||||
Date: 2026-08-05. Vikunja #320 items 2 and 3.
|
||||
Fixture: `internal/router/eval/ru_routing_v1.json`, now **91 cases** (76 ru, 15 en).
|
||||
Model: Qwen3-1.7B-UD-Q4_K_XL, llama-server on the host at 127.0.0.1:8899.
|
||||
Harness: `TestLLMRouterBaseline`, `make eval-models`.
|
||||
|
||||
## How the block was cleared
|
||||
|
||||
Item 2 was blocked because the resident llama-server binds `--host 127.0.0.1
|
||||
--port 0` inside `maven-mavend-1`. The port is kernel-assigned, scraped from
|
||||
stderr and never published, so no `go test` on the host can reach it. The task
|
||||
listed three ways out. This run took the first: a **second** llama-server on
|
||||
the same gguf, on a fixed host port. The Vega takes the second copy of a 1.7B
|
||||
without complaint.
|
||||
|
||||
## The numbers
|
||||
|
||||
| configuration | full | intent-only | p50 | p95 |
|
||||
|---|---|---|---|---|
|
||||
| llm-only | 34/91 (37.4%) | 61.5% | 1.24s | 1.65s |
|
||||
| cascade + llm + hash fallback | 69/91 (75.8%) | 80.2% | 1.19s | 1.65s |
|
||||
|
||||
By language, through the cascade: ru 57/76, en 12/15.
|
||||
Clarify: 3 false, 1 missed. No errors. Six slots deferred to the daemon.
|
||||
|
||||
For comparison, the figures that stood in CLAUDE.md were 72.7% full and 77.9%
|
||||
intent-only, measured on 77 cases. The fixture has grown by 14 cases since, so
|
||||
this is a new baseline rather than a movement.
|
||||
|
||||
## llm-only is low for a reason that is not routing
|
||||
|
||||
37.4% full against 61.5% intent-only is the gap, and it is almost entirely
|
||||
slots. Every reminder case fails with "no time slot, want one". The model
|
||||
routes `reminder` correctly and leaves the time to the daemon, which is what
|
||||
the contract asks of it. The cascade fills those slots. That is why the same
|
||||
model scores 38 points higher inside it.
|
||||
|
||||
Three cases errored in the llm-only arm and none in the cascade, which is the
|
||||
fallback working as designed.
|
||||
|
||||
## Item 3: latency
|
||||
|
||||
Router p50 1.19s, p95 1.65s, max 1.79s through the cascade. The one earlier
|
||||
data point in the task, roughly 6s wall clock for `привет` through
|
||||
`POST /api/chat`, was the whole path and not the router. It is not comparable
|
||||
and should not be quoted as a routing number.
|
||||
|
||||
These numbers are the homesrv floor. With the workstation up, routing completes
|
||||
against gemma-4-12b at p50 329ms, measured separately in
|
||||
`docs/evals/2026-08-02-workstation-gemma4-12b.md`.
|
||||
|
||||
## What still misses
|
||||
|
||||
The confusion is concentrated in one direction: `query→fact ×4`,
|
||||
`query→note ×3`, `query→system ×3`. A question about his own rows that carries
|
||||
no interrogative reads as a statement to the model. Ten of the twenty-two
|
||||
failures are that shape, including "я сегодня вообще пил воду" and "чем я
|
||||
занимался в среду". This is the case V-546's three-head classifier is aimed at.
|
||||
|
||||
The two `разбуди меня` cases clarify at 0.300 instead of routing `reminder`.
|
||||
|
||||
## Item 4 is still not run
|
||||
|
||||
Killing the resident llama-server to confirm the classifier floor needs a
|
||||
permission this session does not have. The test is otherwise ready. It now has
|
||||
a second half. With the workstation up, killing the resident server should
|
||||
still complete a turn through `modelSeam`. Only killing both proves the
|
||||
classifier answers.
|
||||
@@ -0,0 +1,78 @@
|
||||
# Does SearXNG claim a question it cannot answer? (V-539)
|
||||
|
||||
Measured 2026-08-05 against the configured instance, `http://127.0.0.1:9563`,
|
||||
`max_results: 4`, `language: auto`. Sixteen Russian questions: eight real, eight
|
||||
invented from non-words. The probe read SearXNG's JSON directly, so this measures
|
||||
the search, not the cascade around it.
|
||||
|
||||
## The premise no longer reproduces
|
||||
|
||||
V-539 was filed on the 2026-08-02 measurement, where SearXNG returned four
|
||||
results for every query including `зыркабулентный флогистон Мшанского`, and no
|
||||
`voice: kiwix:` line ever appeared. Today the same shape of query returns
|
||||
nothing:
|
||||
|
||||
| Query set | Zero results | Four results claimed |
|
||||
|---|---|---|
|
||||
| Eight real questions | 0 | 8 |
|
||||
| Eight invented questions | 7 | 1 |
|
||||
|
||||
`Response.Empty()` is already the gate. Seven of eight invented questions now
|
||||
pass the turn to the ZIM with no code change at all. What changed is upstream.
|
||||
Every real answer today comes from `google cse`. It answers a non-word with an
|
||||
empty result set, where the engine set of three days ago answered with
|
||||
something.
|
||||
|
||||
## The one that still claims
|
||||
|
||||
`трюмбальная нидроскопия` returned four results, all about a lumbar puncture:
|
||||
|
||||
```
|
||||
Люмбальная пункция - адреса и стоимость в больницах в СПб
|
||||
Пункция спинного мозга - Больница «Шиба
|
||||
Педиатрический фантом люмбальной пункции новорожденного
|
||||
```
|
||||
|
||||
The engine read the invented word as a misspelling of a real one and answered
|
||||
the real one. That is the whole remaining failure, and it is a near-miss
|
||||
spelling rather than a catch-all.
|
||||
|
||||
## The three candidate signals do not separate the sets
|
||||
|
||||
V-539 named three signals a quality gate could read. Each was recorded per
|
||||
query:
|
||||
|
||||
- **No result title shares a token with the query.** Useless. It is true of the
|
||||
one bad claim, and also true of `столица Франции`, whose four titles are
|
||||
`Париж`, `Франция`, `Париж — Путеводитель`, `Париж - Море Трэвел`. The right
|
||||
answer to a capital-city question is the city, which is not a word in the
|
||||
question. Two more real questions score 3 of 4 rather than 4.
|
||||
- **Every snippet is empty.** Never fired. Zero empty snippets across all
|
||||
sixteen queries, real or invented. `ParseResponse` already drops a hit with no
|
||||
text, so this signal cannot fire by construction.
|
||||
- **A spelling-suggestion or catch-all engine answered.** Never fired. SearXNG
|
||||
returned no `corrections` and no `suggestions` for any query, including the one
|
||||
that silently corrected the spelling itself.
|
||||
|
||||
## Decision: do not build the threshold
|
||||
|
||||
A gate on token overlap would cost `столица Франции` a correct answer to save
|
||||
one invented word, and the other two signals cannot fire. The task said a wrong
|
||||
threshold costs a real answer and needs measuring first. It was measured and it
|
||||
loses.
|
||||
|
||||
What ships instead is the second half of V-539. The claiming query source now
|
||||
crosses the IPC seam on `ipc.ChatReply.Source`. It renders as a badge beside the
|
||||
reply on `/chat`. The only evidence before it was a `voice:` log line, which is
|
||||
why this was hard to judge. The next occurrence is readable off the UI rather
|
||||
than off the box.
|
||||
|
||||
## Not measured here
|
||||
|
||||
- The cascade. This probe read SearXNG directly. It says nothing about how
|
||||
`querySearch` phrases what it gets, or whether the resident model turns four
|
||||
weak snippets into a confident wrong sentence.
|
||||
- Kiwix. It was healthy on 2026-08-02 and was not re-probed today.
|
||||
- English questions. The premise was about Russian, where the invented words are.
|
||||
- Whether the engine set is stable. The whole finding is that it moved in three
|
||||
days, so this table is a reading of one day.
|
||||
@@ -0,0 +1,60 @@
|
||||
# Talk fixture against the resident model, 2026-08-05
|
||||
|
||||
Vikunja #44 step 1. `MAVEN_LLM_URL=http://127.0.0.1:8899 make eval-phrasing`,
|
||||
Qwen3-1.7B-UD-Q4_K_XL on the host, no workstation in the run. The fixture holds
|
||||
36 cases now, against 27 when the bakeoff measured it. So the old score is not
|
||||
a column in this table.
|
||||
|
||||
## Result
|
||||
|
||||
| | before the escape fix | after |
|
||||
|---|---|---|
|
||||
| talk, passes every check | 2/36 (5.6%) | 25/36 (69.4%) |
|
||||
| failed generations | 31 | 0 |
|
||||
| by path: chat | 0/9 | 4/9 |
|
||||
| by path: knowledge | 1/9 | 6/9 |
|
||||
| by path: query | 0/9 | 9/9 |
|
||||
| by path: reply | 1/9 | 6/9 |
|
||||
| feminine | 5/36 | 36/36 |
|
||||
| address | 5/36 | 33/36 |
|
||||
| ontopic | 2/36 | 28/36 |
|
||||
| p50 latency | 3.05s | 2.97s |
|
||||
| nudges (15 cases) | 15/15 | 15/15 |
|
||||
|
||||
## What the 31 errors were
|
||||
|
||||
Not the model. `escapeRawControls` in `internal/phraser/llmphraser.go`, added
|
||||
for #537 to repair a raw newline written inside a string, escaped the whole
|
||||
object. Qwen3-1.7B pretty-prints: it opens `{` and writes three newlines before
|
||||
the first key. Those newlines became a literal backslash-n, which is legal
|
||||
nowhere outside a string, so the object stopped parsing and `parseResponseMood`
|
||||
reported `errBrokenJSON`.
|
||||
|
||||
The comment said escaping unconditionally could not turn valid JSON into
|
||||
anything else, because JSON permits no control character outside a string. It
|
||||
permits three. Newline, tab and return are whitespace between tokens, and that
|
||||
is what pretty-printing is made of.
|
||||
|
||||
Every chat reply and every knowledge answer the resident model wrote was being
|
||||
discarded for a stub line. The nudge path never showed it, because the nudge
|
||||
prompt gets compact JSON back.
|
||||
|
||||
## The 11 that still fail
|
||||
|
||||
Eight are `ontopic`, three are `address`.
|
||||
|
||||
The address failures are all plural imperatives written to a formal listener:
|
||||
`держите`, `уточните`, `попробуйте`. Feminine self-reference held in all 36,
|
||||
which is the half #122 is training for. So the persona gap the CPT is aimed at
|
||||
is now the address half, not the gender half.
|
||||
|
||||
The ontopic failures are the resident model answering next to the question
|
||||
rather than in it. `chat-joke` describes crying dolls instead of telling one,
|
||||
`know-hiccups` calls hiccups an icon, `know-boil-egg` answers about an omelette.
|
||||
`chat-about-me` answers "Я - записка", which is the same confabulation the
|
||||
bakeoff recorded.
|
||||
|
||||
## Not measured here
|
||||
|
||||
The workstation. Every number above is the homesrv floor. `make eval-phrasing`
|
||||
points at one URL, so a gemma-4-12b column needs its own run.
|
||||
@@ -0,0 +1,58 @@
|
||||
# Talk temperature sweep: Qwen3-1.7B, 4 temperatures × 3 runs
|
||||
|
||||
Date: 05-08-2026. Model: Qwen3-1.7B-UD-Q4_K_XL, the resident model, on homesrv.
|
||||
Harness: `TestTalkTemperatureSweep` (`internal/phraser/eval/temperature_test.go`),
|
||||
gated on `MAVEN_LLM_URL` + `MAVEN_TEMP_SWEEP`. Fixture: the 36-case talk set.
|
||||
Wall clock: 3394s for all twelve runs. Vikunja #402.
|
||||
|
||||
## What was asked
|
||||
|
||||
Whether 0.7 is the right sampling temperature for phrasing, and whether a lower
|
||||
one buys persona compliance.
|
||||
|
||||
## Numbers
|
||||
|
||||
| temp | run 1 | run 2 | run 3 | mean | errors |
|
||||
|---|---|---|---|---|---|
|
||||
| 0.70 | 24/36 | 24/36 | 22/36 | 23.3 (64.8%) | 7, 4, 6 |
|
||||
| 0.40 | 25/36 | 25/36 | 26/36 | 25.3 (70.4%) | 3, 4, 5 |
|
||||
| 0.20 | 22/36 | 23/36 | 26/36 | 23.7 (65.7%) | 6, 5, 4 |
|
||||
| 0.05 | 23/36 | 25/36 | 23/36 | 23.7 (65.7%) | 5, 5, 6 |
|
||||
|
||||
## What it says
|
||||
|
||||
**The sweep does not separate the temperatures.** 0.40 leads by 5.6 points on
|
||||
the mean. The spread inside a single temperature is 11 points: 0.20 ranges 22 to
|
||||
26 across three runs of the same setting. Three runs cannot tell a 5.6-point
|
||||
effect from that noise. Lowering the temperature to 0.05 does not help either.
|
||||
That is the result that would have been most useful if it had.
|
||||
|
||||
**So the default stays 0.7.** `Config.Temperature` is now a config field, so
|
||||
setting it is a one-line change. No measurement here justifies moving it. Anyone
|
||||
re-running this needs more runs per setting, not more settings.
|
||||
|
||||
## The finding that is not about temperature
|
||||
|
||||
Sixty of the failures across twelve runs are one error:
|
||||
`phraser: model output starts as JSON but does not parse`. The case fails with
|
||||
an empty string, so it costs a whole case rather than one check.
|
||||
|
||||
They are not spread evenly. Every one lands in the `reply` family, and the
|
||||
distribution is:
|
||||
|
||||
| case | runs failed (of 12) |
|
||||
|---|---|
|
||||
| reply-reminder-tomorrow | 12 |
|
||||
| reply-reminder-evening | 12 |
|
||||
| reply-question-bait | 11 |
|
||||
| reply-formality-bait | 9 |
|
||||
| reply-note-router | 8 |
|
||||
| reply-fact-weight | 6 |
|
||||
|
||||
Two cases fail in every single run at every temperature. That is not sampling
|
||||
noise, and no temperature will fix it. It is a defect in the reply phrasing
|
||||
path. It caps the talk fixture at 30/36 before persona is scored at all. Filed
|
||||
as Vikunja #537.
|
||||
|
||||
The talk score of 27/36 recorded on 2026-08-04 went through a different call
|
||||
path. It is not comparable to the numbers above.
|
||||
+8
-1
@@ -1,6 +1,6 @@
|
||||
# Offloading model work to the workstation
|
||||
|
||||
*Last verified: 2026-08-03 @ 12530c8. Living doc: correct it in place, do not append.*
|
||||
*Last verified: 2026-08-05 @ b789676. Living doc: correct it in place, do not append.*
|
||||
|
||||
Owner's call, 2026-08-02. Vikunja #483 is the umbrella. Tasks #484 to #487 are the
|
||||
work, and this file holds the shape and the rules all four must obey.
|
||||
@@ -149,6 +149,13 @@ flips. It is wired anyway: `PhraseReminder` is on the same transport and is on.
|
||||
Then the embedder above, **whisper.cpp** in `mavsttd`, and **piper** in `mavttsd`.
|
||||
`mavwaked` uses no model at all: an energy-threshold VAD over 30ms frames.
|
||||
|
||||
Speech-to-text stays two stages when it moves. One call carrying both a clip and the router
|
||||
prompt was measured on 05-08-2026. It scores 54.2% intent-only against 84.7% for whisper on
|
||||
homesrv, on the same 72 cases. The model transcribes clips it then routes wrong, so a long
|
||||
classification prompt and an audio part compete for attention. Transcribing on the
|
||||
workstation and routing the text scores 83.3% at p50 997ms. So the transfer buys 375ms and
|
||||
cleaner transcripts, not accuracy. See `docs/evals/2026-08-05-audio-in-routing.md`.
|
||||
|
||||
## Order
|
||||
|
||||
1. **Transport** (#484). Nothing else is possible until a seam can cross a host.
|
||||
|
||||
@@ -47,7 +47,13 @@ for.
|
||||
|
||||
QA session 1 step 2 should say what it actually covers, which is push-to-talk through
|
||||
`/dash`. It should not read as though it covers the voice loop. The wake path is checked
|
||||
on the client machine or it is not checked, and today there is no client machine.
|
||||
on the client machine or it is not checked.
|
||||
|
||||
**Correction, 2026-08-05.** This plan said there was no client machine. There is: workpc,
|
||||
where he sits most of the day and where the microphone is. The sentence was written when
|
||||
the workstation was only a model host. The verdict above is unchanged, and so is
|
||||
everything about the seam. What changes is the size of the remaining work: deploying two
|
||||
daemons and asking mavend to listen on TCP, not acquiring hardware.
|
||||
|
||||
That is the honest state, and it is worse than the task suggests: this is not a
|
||||
configuration gap that a compose entry closes. Until a machine with a microphone runs
|
||||
|
||||
@@ -0,0 +1,148 @@
|
||||
# Plan: route with heads on e5-small, not with a generative model
|
||||
|
||||
**Owner's call, 05-08-2026. Vikunja #546.**
|
||||
|
||||
**Verdict: the routing model is the 118M multilingual-e5-small already resident on
|
||||
homesrv.** It gets one classification head per output. No LoRA on a decoder, no 100M model
|
||||
trained from scratch. Routing has a bounded output space, so it is classification. A model
|
||||
that generates is being asked to do the wrong job.
|
||||
|
||||
Last verified: 05-08-2026 @ 52fd218
|
||||
|
||||
## The two options this rules out
|
||||
|
||||
**A LoRA on Qwen3-0.6B or 1.7B.** About 1 to 4 GPU hours on the workstation, for 20k
|
||||
examples over three epochs. It works, and it still generates. So the output still needs a
|
||||
GBNF grammar in front of it. The confidence still has to be rebuilt from structure, the way
|
||||
`gateLLMDecision` does today.
|
||||
|
||||
**A 100M decoder from scratch.** It needs roughly 2B tokens to be a usable language model.
|
||||
That is 6 × 1e8 × 2e9, about 1.2e18 FLOPs. A 16GB card
|
||||
does that in 10 to 20 GPU hours at its effective throughput. It also needs a Russian tokenizer built and a corpus assembled.
|
||||
What it buys is a small model generating Russian, and `CLAUDE.md` already records that as
|
||||
the thing that does not work. LFM2.5-350M routes at 5.2% and answered "столица Франции?"
|
||||
with the invented non-word "Сторзит".
|
||||
|
||||
That finding is about generation, not about size. A 118M encoder classifying Russian is a
|
||||
different job with a bounded output space. The measured recall of e5-small on this box is
|
||||
the evidence it reads the language well enough.
|
||||
|
||||
## What the model becomes
|
||||
|
||||
The encoder body stays as it is. Three heads sit on top of one forward pass:
|
||||
|
||||
| Output | Head | Reads |
|
||||
|---|---|---|
|
||||
| intent | `Linear(384, 7)` | the mean-pooled vector |
|
||||
| mood | `Linear(384, 5)` | the mean-pooled vector |
|
||||
| slots | `Linear(384, 9)` per token | `last_hidden_state` |
|
||||
|
||||
Seven intents are the existing enum: fact, reminder, note, query, act, chat, system. Five
|
||||
moods are the existing enum: neutral, happy, thinking, tired, confused. Nine slot tags are
|
||||
BIO over `key`, `value`, `when` and `fn`, plus outside.
|
||||
|
||||
Total head size is about 12k parameters. That number decides how this is served. See the
|
||||
serving section below.
|
||||
|
||||
## Two things this buys that the current router cannot
|
||||
|
||||
**Constrained output stops being a grammar problem.** There is no free generation, so the
|
||||
heads can only emit values that exist. Four mechanisms exist today because a decoder can
|
||||
write anything. The GBNF grammar, the JSON parse, the fallback to plain text, the legacy
|
||||
`{"body","summary"}` path. A softmax cannot write anything.
|
||||
|
||||
**Confidence becomes a real number.** `Confidence: 1.0` was hardcoded in `llmrouter.go`,
|
||||
so the router could never ask for clarification. V-359 had to rebuild a signal out of
|
||||
structure: single-token utterance, keyless fact, act with no allowlisted fn. Max softmax
|
||||
over the intent head is calibratable against the fixture. `r.threshold` and the stage-3
|
||||
gate would read a probability instead of a proxy. Two false clarifies survived the V-359
|
||||
fix, both in the act-with-no-allowlisted-fn arm. That is the arm a calibrated score
|
||||
replaces.
|
||||
|
||||
## Cost
|
||||
|
||||
About 20k examples at 64 tokens over five epochs is 6.4M tokens. So 6 × 1.2e8 × 6.4e6,
|
||||
roughly 5e15 FLOPs. **10 to 30 minutes on the workstation. Under 2GB of VRAM.** It also
|
||||
finishes overnight on homesrv's CPU when the card is busy. That matters, because the
|
||||
workstation is never assumed up.
|
||||
|
||||
Freeze the embedding table. The XLM-R vocabulary is about 96M of the 118M parameters, and
|
||||
it is the part that overfits 20k examples. Train the twelve layers and the heads, at 2e-5
|
||||
on the body and 1e-3 on the heads, batch 32, sequence 64. Loss is cross-entropy on intent
|
||||
plus mood plus per-token tags, with the tag term down-weighted.
|
||||
|
||||
## The trap: fine-tune a copy
|
||||
|
||||
The resident embedder backs memory recall. `docs/evals/2026-08-04-recall-e5-small.md`
|
||||
measured ten points of recall@1 above MiniLM, at 2.5× the speed. `CLAUDE.md` says it stays
|
||||
on homesrv permanently, because it backs the floor.
|
||||
|
||||
Training it in place couples routing accuracy to recall. Say a run gains four points of
|
||||
intent accuracy and quietly loses six of recall@1. It would look like a win, and nothing
|
||||
in the test suite would name the trade. So the routing weights are a second file. About
|
||||
45MB extra, quantized. `modelIDFromPath` already derives the DB marker from the filename.
|
||||
So two files means two ids, and no ambiguity about which model wrote an embedding.
|
||||
|
||||
## Serving: the heads do not need a runtime
|
||||
|
||||
`internal/router/onnxembedder.go` already asks the ONNX session for `last_hidden_state` at
|
||||
`[1, 128, 384]` and mean-pools in Go. Both tensors the heads need already cross into Go on
|
||||
every call.
|
||||
|
||||
At 12k parameters the heads are three dot products. They can be plain Go over a weights
|
||||
file rather than a second ONNX graph. Then the export step covers the encoder only, and
|
||||
the head weights are data. That keeps the whole thing inside the existing session, the
|
||||
existing vendored tokenizer and the existing `TestONNX*` tests.
|
||||
|
||||
Wire it where the router sits. `pickLLMRouter` becomes a three-way choice. The classifier
|
||||
stays underneath as the floor. The rule that keeps it there now still holds: a turn must
|
||||
never break on a model. The resident model keeps chat, world answers and phrasing. None of that
|
||||
is classification, and this cannot do it.
|
||||
|
||||
## The labeled set is the whole project
|
||||
|
||||
There are 77 RU routing cases and 30 Praxis cases today. That is a test set, not a
|
||||
training set.
|
||||
|
||||
The stage 0 grammars are high-precision label functions. `AgendaQueryGrammars`,
|
||||
`NarrativeQueryGrammar` and `PraxisGrammars` each decide a shape deterministically, so
|
||||
running them over the turn history self-labels it. Training on their output distils the
|
||||
rules into the model. That is the point rather than a compromise: the model generalizes
|
||||
past a regex, where Go's `\b` never fires after a Cyrillic letter.
|
||||
|
||||
Two rules for the data:
|
||||
|
||||
- **The fixtures stay out of training.** Otherwise the measurement reads the rules and
|
||||
reports them as the model.
|
||||
- **Keep cases the grammars do not cover.** A set labeled only by the rules teaches only
|
||||
the rules. The hard cases carry no interrogative and no question mark, which is why
|
||||
V-498 existed.
|
||||
|
||||
Expanding to 20k runs through gemma-4-12b on the workstation. `docs/evals/2026-08-02-workstation-gemma4-12b.md`
|
||||
measured 329ms per call. So that is 2 to 4 hours of the card, plus a day of the owner
|
||||
reading it. That is the real cost of this plan.
|
||||
|
||||
## How it gets judged
|
||||
|
||||
The 77-case RU routing fixture, against the two numbers that always answer:
|
||||
|
||||
| Path | Full | Intent-only | p50 |
|
||||
|---|---|---|---|
|
||||
| classifier | 68.8% | — | 16.6µs |
|
||||
| resident Qwen3-1.7B, through the cascade | 72.7% | 77.9% | 0.80 to 1.04s |
|
||||
| heads on e5-small | to measure | to measure | to measure |
|
||||
|
||||
Latency should land near the classifier rather than near the router. This is one encoder
|
||||
pass and three dot products, against the classifier's one encoder pass and a
|
||||
nearest-neighbour scan. If it does, V-464 answers itself, and the latency trade the LLM
|
||||
router asks for stops being a trade.
|
||||
|
||||
## Not decided here
|
||||
|
||||
- Whether mood belongs on this model at all. Mood is a phrasing property, and the router
|
||||
emitting it is a convenience. A fifth head is cheap, so this is a question about where
|
||||
the value is read, not about cost.
|
||||
- Whether the slot head replaces the stage 0 grammars or sits behind them. Cascade order
|
||||
is a separate measurement, and the rules are currently faster and exact.
|
||||
- The confidence calibration method. Temperature scaling on a held-out split is the
|
||||
obvious first try, and it has not been measured.
|
||||
@@ -0,0 +1,135 @@
|
||||
# Plan: dialogue arbitration, one channel and many claimants
|
||||
|
||||
Umbrella V-558. This file collects the design for its children.
|
||||
|
||||
Last verified: 06-08-2026 @ b6305f1
|
||||
|
||||
## A common unit for claims on an utterance (V-565)
|
||||
|
||||
**Verdict: four ordinal bands, and the band is the tie-break rather than the decision.
|
||||
Coverage decides first.** The measurement below says no claimant Maven has today can produce
|
||||
a graded confidence. A float would be an invention either way. What is available is the KIND
|
||||
of evidence a claimant holds, and there are exactly four kinds.
|
||||
|
||||
### What the claimants report today
|
||||
|
||||
Measured 06-08-2026 on the 91-case RU fixture (`internal/router/eval`), through the deployed
|
||||
cascade with the quantized multilingual-e5-small embedder. The harness is
|
||||
`TestONNXClaimConfidenceDistribution` and `TestStage0Contention` in
|
||||
`internal/router/eval/claims_test.go`. Correct means the right intent, or a refusal where the
|
||||
fixture wants one. Slots are excluded, because a slot miss is a parser question and would
|
||||
blur what the number is being asked to predict.
|
||||
|
||||
| Claimant | Values it can emit | Distribution on the fixture | Correct |
|
||||
|---|---|---|---|
|
||||
| Stage 0 grammars, 21 of them | `1.0`, always | claimed 20 of 91 cases | 20/20 (100%) |
|
||||
| Classifier, cosine | continuous in principle | observed range 0.859 to 0.942 over 71 cases | 44/71 (62%) |
|
||||
| LLM router | `1.0` or `0.3`, nothing between | not run here, no llama-server | see below |
|
||||
| Query sources, 22 of them | a bool | not routed by the fixture | n/a |
|
||||
| Stateful four | nothing at all | n/a | n/a |
|
||||
|
||||
Four findings, and each one constrains the band set.
|
||||
|
||||
**The classifier's cosine carries no signal about correctness.** It scores 62% below the
|
||||
median and 62% above it. That is 13/21 in 0.8 to 0.9, and 31/50 in 0.9 to 1.0. The spread is
|
||||
0.083 wide. Every case sits above the 0.55 threshold, so the gate never fires here. A number
|
||||
flat against correctness, which never crosses its own gate, is not a confidence.
|
||||
|
||||
**Nor does the margin between its top two intents.** Top1 minus top2 is min 0.000, p50
|
||||
0.009, max 0.025. Sixty-eight of the 71 classified cases sit under 0.02 and score 60%. Three
|
||||
clear 0.02 and score 3/3, which is a sample of three. So the ledger's question is answered:
|
||||
a calibrated float is NOT cheaply available from the classifier alone. Nearest-centroid over
|
||||
frozen seeds ranks intents, and the ranking is decided in the third decimal place. It can say
|
||||
which intent is nearest. It cannot say how near.
|
||||
|
||||
**Stage 0 asserts 1.0 by fiat, and on this fixture the fiat is right.** Twenty of twenty.
|
||||
That is not evidence that a hand-written anchored pattern is always right. It is evidence
|
||||
that anchored and nearest are different kinds of claim, and must not share a scale. The gap
|
||||
is 100% against 62% on the same 91 utterances.
|
||||
|
||||
**Stage 0 contention is rarer than the list order suggests.** Exactly one case of 91 draws
|
||||
two grammars. That is `ru-query-019`, where `calendar-query` and `agenda-query` both match,
|
||||
and `calendar-query` wins because it is earlier in `buildRouter`. Both would route
|
||||
`IntentQuery`, so the ordering costs nothing there. The finding is not that ordering is
|
||||
harmless. It is that the fixture barely exercises what V-558 is about. Part of what a claim
|
||||
object buys is making the contention countable.
|
||||
|
||||
**The LLM router emits two values, and one of them is not a confidence.** `llmFullConfidence`
|
||||
is 1.0 and `llmThinConfidence` is 0.3. `gateLLMDecision` moves a decision to 0.3 through
|
||||
three named arms. A fact with no key, an act with no allowlisted fn, a reminder with no
|
||||
subject. Each is a self-veto with a reason, flattened into a number that then loses the
|
||||
reason. Both values are meaningful only against `config.DefaultRouterThreshold`. 0.3 is below
|
||||
0.55 and 1.0 is above it, and nothing anywhere reads any other property of either.
|
||||
|
||||
### The band set
|
||||
|
||||
Four bands, ordinal, highest first. They name the kind of evidence, because that is the one
|
||||
thing every claimant can report without inventing it.
|
||||
|
||||
**`BandAnchored`.** A literal pattern anchored in the utterance matched, and the matched span
|
||||
is what decides the intent. Stage 0 grammars and query-source matchers. The claimant is
|
||||
certain about the shape of the sentence. That is not the same as being certain about the
|
||||
answer. Measured 20/20.
|
||||
|
||||
**`BandStructural`.** A claimant read the whole sentence and produced a complete route. Every
|
||||
slot the intent requires is filled. The LLM router at `llmFullConfidence` sits here, and so
|
||||
does a stateful claimant holding a pending question. Not anchored, because nothing in the
|
||||
utterance is pointed at.
|
||||
|
||||
**`BandNearest`.** The claim rests only on resemblance to something else. No anchor in the
|
||||
utterance, no structural check behind it. The classifier. One band rather than a graded
|
||||
scale, and the measurement is the argument. 62% at both ends of the cosine range, and a
|
||||
top-two margin that never reaches 0.03.
|
||||
|
||||
**`BandVetoed`.** The claimant will take the turn only if nobody else will, and says why it
|
||||
should not. The three arms of `gateLLMDecision` land here with their reason preserved. A
|
||||
vetoed claim is still a claim. Maven asking "о чём напомнить?" beats silence.
|
||||
|
||||
There is no fifth band, and that is a measurement result rather than a preference. No
|
||||
claimant in the cascade today can report what a fifth band would carry. V-546 lands a softmax
|
||||
head whose max probability is a calibrated number. That one gets read as a number, not
|
||||
squeezed into these four.
|
||||
|
||||
### Coverage decides before the band does
|
||||
|
||||
The band is the tie-break. The first question is how much of the utterance a claim explains,
|
||||
and that is `Consumed` against `Unexplained` on the claim object. Two reasons.
|
||||
|
||||
It is the fix for the failure that opened V-558. "какая сейчас погода в Риме?" arrived while
|
||||
a reminder was pending. The pending claimant ate the whole utterance as a time answer while
|
||||
explaining none of it. Not "погода", not "Риме", not the question mark. A weather claim
|
||||
explains all of it. Coverage-first arbitration prefers the weather claim without knowing that
|
||||
a pending reminder is less trustworthy than a grammar. The pending question then survives to
|
||||
be asked again.
|
||||
|
||||
It also keeps the stateful four out of the top slot without special-casing them. They sit at
|
||||
`BandStructural`, below any anchored claim. That is the whole V-558 complaint about the
|
||||
highest-priority claimants being the least informed, expressed as one rule.
|
||||
|
||||
### The claim object
|
||||
|
||||
```go
|
||||
type Claim struct {
|
||||
Claimant string // who wants the turn
|
||||
Intent string // plain string: internal/dialogue must not import internal/router
|
||||
Filled []string // the slots this claim would fill
|
||||
Consumed []string // utterance tokens this claim explains
|
||||
Unexplained []string // the rest, in order
|
||||
Band Band
|
||||
Veto string // why this claim should NOT win, empty when there is none
|
||||
}
|
||||
```
|
||||
|
||||
`Intent` is a plain `string` rather than `router.Intent` on purpose. `internal/dialogue` must
|
||||
not import `internal/router`, so the claim package must not either, and a shared string costs
|
||||
one conversion at each edge.
|
||||
|
||||
`Unexplained` is carried rather than derived at read time. A claimant can then decline to
|
||||
explain a span it did match.
|
||||
|
||||
### What this task does not do
|
||||
|
||||
`router.Decision.Confidence` stays and keeps its float. `r.threshold` and `gateLLMDecision`
|
||||
read it, and the classifier is the failure floor. A rewire that broke either would trade a
|
||||
measured floor for an unmeasured design. V-565 lands the type and the builder beside the
|
||||
existing path. The arbiter that reads claims is V-560.
|
||||
+17
-8
@@ -1,6 +1,6 @@
|
||||
# QA plan: checking Maven properly
|
||||
|
||||
*Last verified: 2026-08-04 @ 8d816f4. Living doc: correct it in place, do not append.*
|
||||
*Last verified: 2026-08-05 @ 12667fd. Living doc: correct it in place, do not append.*
|
||||
|
||||
Written 2026-08-01, after the 35-PR stack landed and the box came back up.
|
||||
Refreshed 2026-08-02 against the live list, after PRs #85-#90.
|
||||
@@ -255,8 +255,13 @@ room — **463**, written up in `docs/plans/17-where-the-voice-loop-runs.md`.
|
||||
So the wake word and the VAD gate are covered by their unit tests and by
|
||||
nothing else, and no session at this box changes that. Checking them needs a
|
||||
machine with a microphone running both binaries against a TCP-listening mavend.
|
||||
`ipc.Dial` already speaks `tcp://host:port?token=...`, so the work is a machine
|
||||
and a config line, not protocol work. Until then, **287** can only be
|
||||
`ipc.Dial` already speaks `tcp://host:port?token=...`, so the work is not
|
||||
protocol work.
|
||||
|
||||
That machine is workpc (owner's correction, 05-08-2026). This section used to
|
||||
call it a machine Maven does not have. That was written when the workstation
|
||||
was only a model host. So the remaining work is deploying two daemons and
|
||||
asking mavend to listen on TCP. Until that is done, **287** can only be
|
||||
half-answered, and step 2 above is push-to-talk, not the voice loop.
|
||||
|
||||
**319's single-token bug is fixed** (01-08-2026). Single-word Russian utterances no longer come
|
||||
@@ -418,11 +423,15 @@ something only a ZIM answers, with the search block on. Live search leads and th
|
||||
ZIMs are the fallback since 02-08-2026. **286**'s remaining half is doc and
|
||||
git ingestion, which is build work, not a check.
|
||||
|
||||
**Do not read `/trace` for this.** `/trace` is the nudge-rule trace: rule,
|
||||
severity, predicate, gate, selected. No query-source field exists anywhere in the
|
||||
codebase. The only evidence of which query source claimed a turn is the
|
||||
`voice: search:` and `voice: kiwix:` lines in `docker compose logs mavend`
|
||||
(`actions_query.go:589` and `:660`).
|
||||
**Read `/trace` for this.** It carries two tables since 06-08-2026 (V-564). The
|
||||
nudge-rule trace it always had, and below it the **turn decisions**: one
|
||||
collapsible record per utterance. Each names every claimant, what it would have
|
||||
made the turn, the score it reported, and whether it won, declined, lost or was
|
||||
**never asked**. That last one answers "did Kiwix pass, or was it never
|
||||
reached". The log lines cannot tell you that. The ring holds the last 25 turns
|
||||
in daemon memory and is empty after a restart, so read it in the same sitting.
|
||||
`/chat` still shows the claiming source as a badge, and the `voice: query
|
||||
claimed by source` line is still in `docker compose logs mavend`.
|
||||
|
||||
Run 02-08-2026, 20 turns. **Search leads and the personal boundary holds.** Every
|
||||
world question that reached the boundary was claimed by search. All three
|
||||
|
||||
@@ -348,8 +348,8 @@ func TestGate_IpcServer_ChatAllowedForEnrolledCaller(t *testing.T) {
|
||||
if err != nil {
|
||||
t.Fatalf("Chat: %v", err)
|
||||
}
|
||||
if reply != "echo: привет" {
|
||||
t.Fatalf("Chat reply = %q; want %q", reply, "echo: привет")
|
||||
if reply.Reply != "echo: привет" {
|
||||
t.Fatalf("Chat reply = %q; want %q", reply.Reply, "echo: привет")
|
||||
}
|
||||
if fake.chats != 1 {
|
||||
t.Fatalf("CoreAPI.Chat calls = %d; want 1", fake.chats)
|
||||
@@ -373,9 +373,9 @@ func (r *recordingAPI) WriteFact(_ context.Context, _ ipc.WriteFactReq) (int64,
|
||||
return int64(r.writes), nil
|
||||
}
|
||||
|
||||
func (r *recordingAPI) Chat(_ context.Context, _, text string) (string, error) {
|
||||
func (r *recordingAPI) Chat(_ context.Context, _, text string) (ipc.ChatReply, error) {
|
||||
r.chats++
|
||||
return "echo: " + text, nil
|
||||
return ipc.ChatReply{Reply: "echo: " + text}, nil
|
||||
}
|
||||
|
||||
// mustWriteFactParams — minimal WriteFactReq JSON with only the source field,
|
||||
|
||||
@@ -91,6 +91,19 @@ func Requirement(m ipc.Method) Authority {
|
||||
// can do is make Maven stop recognising someone, which is the state the
|
||||
// box ships in anyway.
|
||||
return AuthWrite
|
||||
case ipc.MethodSeedEvent:
|
||||
// The one backdating write path in the tree (Vikunja #518). AuthStepUp,
|
||||
// the same rung as mutating the tool allowlist, and for a reason that is
|
||||
// not about privilege: every other write records when something actually
|
||||
// happened, and this one asserts it. A caller who can place a fact in the
|
||||
// past can manufacture a routine Maven will then act on forever, which is
|
||||
// the tick loop obeying evidence nobody produced.
|
||||
//
|
||||
// Step-up is not the real gate and is not meant to be. mavend refuses the
|
||||
// method entirely unless started with -allow-seed, so the ordinary state
|
||||
// of the box is that no gesture reaches it. This rung is what stops a
|
||||
// module from calling it on a box where QA left the flag on.
|
||||
return AuthStepUp
|
||||
case ipc.MethodWriteFact:
|
||||
return AuthWrite
|
||||
case ipc.MethodIngestMail:
|
||||
|
||||
@@ -16,6 +16,7 @@ package calendar
|
||||
import (
|
||||
"fmt"
|
||||
"sort"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
"unicode"
|
||||
@@ -125,6 +126,69 @@ func FactSummary(value string) string {
|
||||
return value[:i]
|
||||
}
|
||||
|
||||
// EventKeyPrefix — every calendar event fact starts with this. The loop scans
|
||||
// the family to work out whether a meeting covers right now (Vikunja #513).
|
||||
const EventKeyPrefix = "calendar_event_"
|
||||
|
||||
// FactSpan reads an event fact back into the instants it covers, against loc.
|
||||
// The day comes from the key and the two clock readings from the value's
|
||||
// "@ HH:MM-HH:MM" tail, which is everything FactValue wrote.
|
||||
//
|
||||
// ok is false for anything that does not parse. A fact whose span cannot be
|
||||
// read tells you nothing about now, and guessing a span is how a signal that
|
||||
// was meant to suppress one nudge starts suppressing all of them.
|
||||
//
|
||||
// An end at or before the start is read as crossing midnight, so a 23:30-00:15
|
||||
// meeting covers the quarter hour it actually covers.
|
||||
func FactSpan(key, value string, loc *time.Location) (start, end time.Time, ok bool) {
|
||||
if !strings.HasPrefix(key, EventKeyPrefix) {
|
||||
return time.Time{}, time.Time{}, false
|
||||
}
|
||||
rest := key[len(EventKeyPrefix):]
|
||||
if len(rest) < 8 {
|
||||
return time.Time{}, time.Time{}, false
|
||||
}
|
||||
day, err := time.ParseInLocation("20060102", rest[:8], loc)
|
||||
if err != nil {
|
||||
return time.Time{}, time.Time{}, false
|
||||
}
|
||||
// SetValue stores a string fact JSON-encoded, so the value comes back
|
||||
// quoted. Reading the tail off the quote is how this returned false for
|
||||
// every real event the first time it ran.
|
||||
if unq, err := strconv.Unquote(value); err == nil {
|
||||
value = unq
|
||||
}
|
||||
i := strings.LastIndex(value, " @ ")
|
||||
if i < 0 {
|
||||
return time.Time{}, time.Time{}, false
|
||||
}
|
||||
tail := value[i+len(" @ "):]
|
||||
from, to, found := strings.Cut(tail, "-")
|
||||
if !found {
|
||||
return time.Time{}, time.Time{}, false
|
||||
}
|
||||
sh, sm, ok1 := parseHM(strings.TrimSpace(from))
|
||||
eh, em, ok2 := parseHM(strings.TrimSpace(to))
|
||||
if !ok1 || !ok2 {
|
||||
return time.Time{}, time.Time{}, false
|
||||
}
|
||||
start = day.Add(time.Duration(sh)*time.Hour + time.Duration(sm)*time.Minute)
|
||||
end = day.Add(time.Duration(eh)*time.Hour + time.Duration(em)*time.Minute)
|
||||
if !end.After(start) {
|
||||
end = end.Add(24 * time.Hour)
|
||||
}
|
||||
return start, end, true
|
||||
}
|
||||
|
||||
// parseHM reads "15:04" and nothing else.
|
||||
func parseHM(s string) (h, m int, ok bool) {
|
||||
t, err := time.Parse("15:04", s)
|
||||
if err != nil {
|
||||
return 0, 0, false
|
||||
}
|
||||
return t.Hour(), t.Minute(), true
|
||||
}
|
||||
|
||||
// KeyPrefixForDay is the fact-key prefix covering one calendar day. The store
|
||||
// range-scans between two of these.
|
||||
func KeyPrefixForDay(day time.Time) string {
|
||||
|
||||
@@ -0,0 +1,192 @@
|
||||
// Package claim is the common unit for the many claimants that compete for one
|
||||
// utterance (V-565, umbrella V-558, design in
|
||||
// docs/plans/19-dialogue-arbitration.md).
|
||||
//
|
||||
// Maven's cascade has roughly ten stage-0 grammars, seven router intents,
|
||||
// twenty-two query sources and four stateful pre-emptors, and every one of them
|
||||
// answers "is this mine?" alone. None can answer "is this more mine than
|
||||
// yours?", because their scores are not comparable: stage 0 asserts 1.0 by
|
||||
// fiat, the classifier reports a cosine, the LLM router derives one from
|
||||
// structure. So list order is the whole arbitration.
|
||||
//
|
||||
// A Claim carries evidence rather than a verdict. Two things read that evidence
|
||||
// and neither needs a float:
|
||||
//
|
||||
// - specificity — a claim explaining more of the utterance is preferred, and
|
||||
// that is Consumed against Unexplained;
|
||||
// - negative constraint — a claimant may veto itself and say why, and that is
|
||||
// Veto.
|
||||
//
|
||||
// Where a number is unavoidable, it is an ordinal Band and not a probability.
|
||||
// The band set is argued from measurement in the plan doc: the classifier's
|
||||
// cosine is flat against correctness (62% at both ends of a spread 0.083 wide)
|
||||
// and its top-two margin is p50 0.009, so no claimant Maven has today can
|
||||
// produce a graded confidence.
|
||||
//
|
||||
// This package deliberately imports nothing from the rest of Maven.
|
||||
// internal/dialogue must not import internal/router, so a shared unit that
|
||||
// pulled in router.Intent would smuggle that edge back in. Intent is a plain
|
||||
// string and the conversion happens at each edge.
|
||||
package claim
|
||||
|
||||
import "strings"
|
||||
|
||||
// Band — the kind of evidence behind a claim, ordinal and comparable. Higher
|
||||
// wins a tie. Four values, because four is what the claimants can report.
|
||||
type Band int
|
||||
|
||||
const (
|
||||
// BandUnknown — the zero value. A claim that never set a band is a bug in
|
||||
// its builder, not a weak claim, so it must not silently rank as one.
|
||||
BandUnknown Band = iota
|
||||
|
||||
// BandVetoed — the claimant will take the turn only if nobody else will,
|
||||
// and Veto says why it should not. The three arms of gateLLMDecision (a
|
||||
// fact with no key, an act with no allowlisted fn, a reminder with no
|
||||
// subject) land here. Still a claim: asking "о чём напомнить?" beats
|
||||
// silence.
|
||||
BandVetoed
|
||||
|
||||
// BandNearest — the claim rests only on resemblance to something else,
|
||||
// with no anchor in the utterance and no structural check behind it. The
|
||||
// nearest-centroid classifier. One band and not a graded scale, because
|
||||
// the cosine measured flat against correctness.
|
||||
BandNearest
|
||||
|
||||
// BandStructural — the claimant read the whole sentence and produced a
|
||||
// complete route, every slot its intent requires filled. The LLM router at
|
||||
// full confidence, and a stateful claimant holding a pending question.
|
||||
// Below BandAnchored on purpose: the four stateful claimants pre-empt
|
||||
// unconditionally today, and that is the V-558 defect.
|
||||
BandStructural
|
||||
|
||||
// BandAnchored — a literal pattern anchored in the utterance matched, and
|
||||
// the matched span is what decides the intent. Stage 0 grammars and
|
||||
// query-source matchers. Certainty about the shape of the sentence, which
|
||||
// is not certainty about the answer.
|
||||
BandAnchored
|
||||
)
|
||||
|
||||
// String — the band's name, for a trace line and for a test failure that has to
|
||||
// say which band it got.
|
||||
func (b Band) String() string {
|
||||
switch b {
|
||||
case BandVetoed:
|
||||
return "vetoed"
|
||||
case BandNearest:
|
||||
return "nearest"
|
||||
case BandStructural:
|
||||
return "structural"
|
||||
case BandAnchored:
|
||||
return "anchored"
|
||||
default:
|
||||
return "unknown"
|
||||
}
|
||||
}
|
||||
|
||||
// Claim — one claimant's bid for one utterance.
|
||||
type Claim struct {
|
||||
// Claimant — who wants the turn. A grammar name, a query source name, a
|
||||
// stage label. Read by the trace and by a test naming a loser.
|
||||
Claimant string
|
||||
|
||||
// Intent — the route this claim would take. A plain string and not
|
||||
// router.Intent: see the package comment.
|
||||
Intent string
|
||||
|
||||
// Filled — the slot names this claim would fill ("time", "fn", "key",
|
||||
// "text"). Names and not values, because arbitration compares shape.
|
||||
Filled []string
|
||||
|
||||
// Consumed — the utterance tokens this claim explains, in the order they
|
||||
// appear. The numerator of specificity.
|
||||
Consumed []string
|
||||
|
||||
// Unexplained — the tokens this claim does not explain, in order. Carried
|
||||
// rather than derived, so a claimant may decline a span it did match.
|
||||
Unexplained []string
|
||||
|
||||
// Band — the kind of evidence. The tie-break, after coverage.
|
||||
Band Band
|
||||
|
||||
// Veto — why this claim should NOT win, empty when there is none. A
|
||||
// non-empty Veto and a Band above BandVetoed is legal: a claim can be
|
||||
// well-evidenced and still name a reason to prefer somebody else.
|
||||
Veto string
|
||||
}
|
||||
|
||||
// Coverage — the fraction of the utterance this claim explains, in [0,1]. A
|
||||
// claim with no tokens either way covers nothing; it is not division by zero
|
||||
// and it is not a full claim.
|
||||
func (c Claim) Coverage() float64 {
|
||||
total := len(c.Consumed) + len(c.Unexplained)
|
||||
if total == 0 {
|
||||
return 0
|
||||
}
|
||||
return float64(len(c.Consumed)) / float64(total)
|
||||
}
|
||||
|
||||
// Vetoed reports whether the claimant named a reason against itself.
|
||||
func (c Claim) Vetoed() bool { return c.Veto != "" }
|
||||
|
||||
// MoreSpecificThan — the ordering V-560's arbiter will read. Coverage first,
|
||||
// because that is what fixes the failure this program opened with: a pending
|
||||
// reminder ate "какая сейчас погода в Риме?" as a time answer while explaining
|
||||
// none of it. Band only breaks a coverage tie.
|
||||
//
|
||||
// Deliberately NOT wired into the cascade by V-565. It is here so the ordering
|
||||
// is one function with tests on it, rather than a rule restated at each of the
|
||||
// sites that will eventually call it.
|
||||
func (c Claim) MoreSpecificThan(other Claim) bool {
|
||||
cc, oc := c.Coverage(), other.Coverage()
|
||||
if cc != oc {
|
||||
return cc > oc
|
||||
}
|
||||
return c.Band > other.Band
|
||||
}
|
||||
|
||||
// Tokens — the utterance split for coverage accounting. Whitespace, then
|
||||
// trailing and leading punctuation, then lowercased.
|
||||
//
|
||||
// This is tokenization over the raw string and not a Russian pattern: it
|
||||
// contains no word list, and its output is a count rather than a fact or a
|
||||
// route (CLAUDE.md § Russian patterns). Lowercasing is Unicode-aware, so
|
||||
// Cyrillic folds the same way Latin does.
|
||||
func Tokens(utterance string) []string {
|
||||
fields := strings.FieldsFunc(utterance, func(r rune) bool {
|
||||
return r == ' ' || r == '\t' || r == '\n' || r == '\r'
|
||||
})
|
||||
out := make([]string, 0, len(fields))
|
||||
for _, f := range fields {
|
||||
t := strings.Trim(strings.ToLower(f), ".,!?;:()\"'«»…-–—")
|
||||
if t == "" {
|
||||
continue
|
||||
}
|
||||
out = append(out, t)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// Split partitions the utterance's tokens into the ones a claim explains and
|
||||
// the rest, preserving order in both. A token is explained when it appears in
|
||||
// one of the spans the claimant filled (a slot value, a matched substring).
|
||||
//
|
||||
// Duplicates are handled by membership and not by count: "напомни напомни
|
||||
// позвонить" with span "напомни" explains both copies. The alternative is a
|
||||
// multiset, and no claimant Maven has can say which copy it meant.
|
||||
func Split(utterance string, spans ...string) (consumed, unexplained []string) {
|
||||
explained := map[string]bool{}
|
||||
for _, s := range spans {
|
||||
for _, t := range Tokens(s) {
|
||||
explained[t] = true
|
||||
}
|
||||
}
|
||||
for _, t := range Tokens(utterance) {
|
||||
if explained[t] {
|
||||
consumed = append(consumed, t)
|
||||
} else {
|
||||
unexplained = append(unexplained, t)
|
||||
}
|
||||
}
|
||||
return consumed, unexplained
|
||||
}
|
||||
@@ -0,0 +1,120 @@
|
||||
package claim
|
||||
|
||||
import (
|
||||
"reflect"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestTokensStripsPunctuationAndCase(t *testing.T) {
|
||||
got := Tokens("Какая сейчас погода в Риме?")
|
||||
want := []string{"какая", "сейчас", "погода", "в", "риме"}
|
||||
if !reflect.DeepEqual(got, want) {
|
||||
t.Errorf("Tokens = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
// The dash forms matter: an STT transcript routinely carries "а, да, прости -
|
||||
// на 9", and a stray dash counted as a token would dilute every coverage score
|
||||
// in the sentence.
|
||||
func TestTokensDropsBareDashes(t *testing.T) {
|
||||
got := Tokens("а, да, прости — на 9")
|
||||
want := []string{"а", "да", "прости", "на", "9"}
|
||||
if !reflect.DeepEqual(got, want) {
|
||||
t.Errorf("Tokens = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSplitPartitionsInOrder(t *testing.T) {
|
||||
consumed, unexplained := Split("напомни позвонить маме", "позвонить маме")
|
||||
if want := []string{"позвонить", "маме"}; !reflect.DeepEqual(consumed, want) {
|
||||
t.Errorf("consumed = %q, want %q", consumed, want)
|
||||
}
|
||||
if want := []string{"напомни"}; !reflect.DeepEqual(unexplained, want) {
|
||||
t.Errorf("unexplained = %q, want %q", unexplained, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestCoverageIsZeroWithoutTokens(t *testing.T) {
|
||||
if got := (Claim{}).Coverage(); got != 0 {
|
||||
t.Errorf("Coverage of an empty claim = %v, want 0", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestCoverageFraction(t *testing.T) {
|
||||
c := Claim{Consumed: []string{"a", "b", "c"}, Unexplained: []string{"d"}}
|
||||
if got := c.Coverage(); got != 0.75 {
|
||||
t.Errorf("Coverage = %v, want 0.75", got)
|
||||
}
|
||||
}
|
||||
|
||||
// The band order is load-bearing, so it is asserted rather than assumed from
|
||||
// the iota. BandUnknown must sit at the bottom: a claim whose builder forgot to
|
||||
// set a band is a bug and must not outrank a measured one.
|
||||
func TestBandOrder(t *testing.T) {
|
||||
ordered := []Band{BandUnknown, BandVetoed, BandNearest, BandStructural, BandAnchored}
|
||||
for i := 1; i < len(ordered); i++ {
|
||||
if !(ordered[i-1] < ordered[i]) {
|
||||
t.Errorf("%v is not below %v", ordered[i-1], ordered[i])
|
||||
}
|
||||
}
|
||||
for _, b := range ordered {
|
||||
if b.String() == "" {
|
||||
t.Errorf("band %d has no name", b)
|
||||
}
|
||||
}
|
||||
if BandUnknown.String() != "unknown" {
|
||||
t.Errorf("BandUnknown.String() = %q", BandUnknown.String())
|
||||
}
|
||||
}
|
||||
|
||||
// The failure V-558 opened with, as an ordering test. A pending reminder eats
|
||||
// "какая сейчас погода в Риме?" as a time answer and explains none of it; the
|
||||
// weather source explains all of it. Coverage decides, and the band never gets
|
||||
// consulted, which is the point: the pending claimant is structural and the
|
||||
// weather claim is anchored, but even if the bands were equal the weather claim
|
||||
// wins.
|
||||
func TestCoverageBeatsBand(t *testing.T) {
|
||||
utterance := "какая сейчас погода в Риме?"
|
||||
pendingConsumed, pendingRest := Split(utterance, "сейчас")
|
||||
pending := Claim{
|
||||
Claimant: "reminder-followup", Intent: "reminder",
|
||||
Consumed: pendingConsumed, Unexplained: pendingRest,
|
||||
Band: BandStructural,
|
||||
}
|
||||
weatherConsumed, weatherRest := Split(utterance, utterance)
|
||||
weather := Claim{
|
||||
Claimant: "weather", Intent: "query",
|
||||
Consumed: weatherConsumed, Unexplained: weatherRest,
|
||||
Band: BandNearest,
|
||||
}
|
||||
if !weather.MoreSpecificThan(pending) {
|
||||
t.Errorf("weather (%v cover) did not beat pending (%v cover)",
|
||||
weather.Coverage(), pending.Coverage())
|
||||
}
|
||||
if pending.MoreSpecificThan(weather) {
|
||||
t.Error("pending beat weather, so the ordering is not asymmetric")
|
||||
}
|
||||
}
|
||||
|
||||
// Where coverage ties, the band decides. Two grammars claiming the same
|
||||
// utterance is the stage-0 contention case (ru-query-019 on the fixture), and
|
||||
// today list order settles it with nothing recorded.
|
||||
func TestBandBreaksACoverageTie(t *testing.T) {
|
||||
anchored := Claim{Claimant: "calendar-query", Consumed: []string{"a"}, Band: BandAnchored}
|
||||
nearest := Claim{Claimant: "classifier", Consumed: []string{"a"}, Band: BandNearest}
|
||||
if !anchored.MoreSpecificThan(nearest) {
|
||||
t.Error("anchored did not beat nearest on equal coverage")
|
||||
}
|
||||
if nearest.MoreSpecificThan(anchored) {
|
||||
t.Error("nearest beat anchored on equal coverage")
|
||||
}
|
||||
}
|
||||
|
||||
func TestVetoed(t *testing.T) {
|
||||
if (Claim{}).Vetoed() {
|
||||
t.Error("a claim with no veto reports itself vetoed")
|
||||
}
|
||||
if !(Claim{Veto: "fact with no key"}).Vetoed() {
|
||||
t.Error("a claim with a veto reason does not report itself vetoed")
|
||||
}
|
||||
}
|
||||
@@ -1085,6 +1085,17 @@ type KiwixConfig struct {
|
||||
// query at a time.
|
||||
Book string `json:"book,omitempty"`
|
||||
|
||||
// BookRU — the ZIM to search when the question is in Russian, by the same
|
||||
// catalog name. Empty ⇒ every question goes to Book.
|
||||
//
|
||||
// It exists because the rewriter is a workaround, not a feature (V-508). An
|
||||
// English ZIM cannot match a Russian sentence, so the resident model turns
|
||||
// the question into English keywords first, and that costs a model call and
|
||||
// loses whatever the keywords drop. A Russian ZIM matches the question as he
|
||||
// asked it. So a Cyrillic question searches this book verbatim and skips the
|
||||
// rewrite, and the English book keeps answering English ones.
|
||||
BookRU string `json:"book_ru,omitempty"`
|
||||
|
||||
// MaxResults — how many hits are asked for. 0 ⇒ DefaultKiwixResults.
|
||||
// Only the top few reach the phraser regardless; the rest are context the
|
||||
// snippet ranking throws away.
|
||||
|
||||
@@ -0,0 +1,185 @@
|
||||
// Package decision records who claimed one turn and who lost it (V-564).
|
||||
//
|
||||
// Arbitration between the claimants on the utterance stream is order, hardcoded
|
||||
// in the resolver ladder, in buildRouter and in querySources (V-558). Order is
|
||||
// invisible in a log: the daemon says which intent won and which query source
|
||||
// answered, never who else wanted the turn, with what score, or why it did not
|
||||
// get it. The Rome misroute took a probe, a log read and a code read to explain,
|
||||
// which is one diagnosis too many for a defect family already three deep.
|
||||
//
|
||||
// The record rides the context, the same seam querysource.go uses and for the
|
||||
// same reason: a turn answers through one string that the mic, telegram and the
|
||||
// web all share, so a second return value is not threadable. A context with no
|
||||
// recorder notes nothing, so every Note here is free in a test or a tool that
|
||||
// did not ask for one.
|
||||
//
|
||||
// The most important thing it holds is not a loss but a silence. A claimant
|
||||
// that was NEVER ASKED — because something earlier in the ladder returned
|
||||
// first — looks identical to one that examined the turn and declined, and it is
|
||||
// that confusion the hardcoded ordering hides. So a stage declares its roster
|
||||
// up front and Finish names everyone who never reported.
|
||||
package decision
|
||||
|
||||
import (
|
||||
"context"
|
||||
"sync"
|
||||
"time"
|
||||
)
|
||||
|
||||
// Stages, in the order a turn passes through them.
|
||||
const (
|
||||
StagePreRoute = "pre-route" // clarify, confirm, repair and their siblings
|
||||
StageZero = "stage0" // grammar rules
|
||||
StageRoute = "route" // LLM router, classifier
|
||||
StageMerge = "merge" // the follow-up merge, which edits rather than claims
|
||||
StageQuery = "query" // the query source chain
|
||||
StageAction = "action" // whoever actually produced the reply
|
||||
)
|
||||
|
||||
// Outcomes. Coarse on purpose: the record answers "who wanted this turn and
|
||||
// what happened to their claim", not "re-derive the branch".
|
||||
const (
|
||||
Won = "won" // this claimant produced the turn
|
||||
Declined = "declined" // it looked at the turn and said not mine
|
||||
LostOnOrder = "lost_on_order" // it wanted the turn, something earlier had it
|
||||
LostOnScore = "lost_on_score" // it was scored against a rival and scored lower
|
||||
Thinned = "thinned" // it claimed, and a gate cut its confidence
|
||||
Merged = "merged" // it changed the winning claim without owning it
|
||||
NeverAsked = "never_asked" // it never got to look at all
|
||||
)
|
||||
|
||||
// Claim is one claimant's say on one turn.
|
||||
type Claim struct {
|
||||
Stage string `json:"stage"`
|
||||
Claimant string `json:"claimant"`
|
||||
Intent string `json:"intent,omitempty"` // what it would have made the turn
|
||||
Score float64 `json:"score,omitempty"` // only meaningful with HasScore
|
||||
HasScore bool `json:"has_score,omitempty"`
|
||||
Outcome string `json:"outcome"`
|
||||
Reason string `json:"reason,omitempty"` // why it lost, in its own terms
|
||||
}
|
||||
|
||||
// Record is one turn's arbitration. Utterance is held because a record with no
|
||||
// utterance is unreadable, and this store is diagnostics with a short life —
|
||||
// unlike the facts table, which is the audit trail.
|
||||
type Record struct {
|
||||
Ts time.Time `json:"ts"`
|
||||
Utterance string `json:"utterance"`
|
||||
Winner string `json:"winner"`
|
||||
Claims []Claim `json:"claims"`
|
||||
|
||||
mu sync.Mutex
|
||||
rosters []roster
|
||||
}
|
||||
|
||||
type roster struct {
|
||||
stage string
|
||||
names []string
|
||||
}
|
||||
|
||||
// Expect declares the claimants a stage could have asked, so Finish can tell a
|
||||
// decline from a silence. The slice is held, not copied: every caller passes a
|
||||
// package-level table.
|
||||
func (r *Record) Expect(stage string, names []string) {
|
||||
if r == nil {
|
||||
return
|
||||
}
|
||||
r.mu.Lock()
|
||||
defer r.mu.Unlock()
|
||||
r.rosters = append(r.rosters, roster{stage: stage, names: names})
|
||||
}
|
||||
|
||||
// Note appends one claim. The winner is whoever noted Won last, which is the
|
||||
// claimant that actually returned the reply.
|
||||
func (r *Record) Note(c Claim) {
|
||||
if r == nil {
|
||||
return
|
||||
}
|
||||
r.mu.Lock()
|
||||
defer r.mu.Unlock()
|
||||
r.Claims = append(r.Claims, c)
|
||||
if c.Outcome == Won {
|
||||
r.Winner = c.Stage + ":" + c.Claimant
|
||||
}
|
||||
}
|
||||
|
||||
// NoteIfUnclaimed records a win only when nobody has claimed the turn yet. It
|
||||
// is what closes a record whose route was decided but whose reply came from
|
||||
// somewhere with no scoreboard — a clarify question, or an action handler with
|
||||
// no chain in front of it. Without it a thinned route leaves the record with no
|
||||
// winner at all, which reads as a lost turn rather than an asked question.
|
||||
func (r *Record) NoteIfUnclaimed(c Claim) {
|
||||
if r == nil {
|
||||
return
|
||||
}
|
||||
r.mu.Lock()
|
||||
claimed := r.Winner != ""
|
||||
r.mu.Unlock()
|
||||
if claimed {
|
||||
return
|
||||
}
|
||||
c.Outcome = Won
|
||||
r.Note(c)
|
||||
}
|
||||
|
||||
// Finish fills in the never-asked claimants and returns the record. Called once
|
||||
// by whoever installed the recorder, after the turn has answered.
|
||||
func (r *Record) Finish(now time.Time) *Record {
|
||||
if r == nil {
|
||||
return nil
|
||||
}
|
||||
r.mu.Lock()
|
||||
defer r.mu.Unlock()
|
||||
r.Ts = now
|
||||
reported := map[string]bool{}
|
||||
for _, c := range r.Claims {
|
||||
reported[c.Stage+":"+c.Claimant] = true
|
||||
}
|
||||
for _, ros := range r.rosters {
|
||||
for _, name := range ros.names {
|
||||
if !reported[ros.stage+":"+name] {
|
||||
r.Claims = append(r.Claims, Claim{
|
||||
Stage: ros.stage, Claimant: name, Outcome: NeverAsked,
|
||||
})
|
||||
}
|
||||
}
|
||||
}
|
||||
return r
|
||||
}
|
||||
|
||||
// --- the context seam ---
|
||||
|
||||
type recorderKey struct{}
|
||||
|
||||
// With returns a context carrying a fresh record, and the record to read after
|
||||
// the turn has answered.
|
||||
func With(ctx context.Context, utterance string) (context.Context, *Record) {
|
||||
rec := &Record{Utterance: utterance}
|
||||
return context.WithValue(ctx, recorderKey{}, rec), rec
|
||||
}
|
||||
|
||||
// From returns the record on the context, or nil. Every method on *Record is
|
||||
// nil-safe, so a caller does not have to check.
|
||||
func From(ctx context.Context) *Record {
|
||||
rec, _ := ctx.Value(recorderKey{}).(*Record)
|
||||
return rec
|
||||
}
|
||||
|
||||
// Note is the shorthand every claim site uses: a no-op when nobody is recording.
|
||||
func Note(ctx context.Context, c Claim) {
|
||||
From(ctx).Note(c)
|
||||
}
|
||||
|
||||
// Expect is the roster shorthand, likewise a no-op with no recorder.
|
||||
func Expect(ctx context.Context, stage string, names []string) {
|
||||
From(ctx).Expect(stage, names)
|
||||
}
|
||||
|
||||
// Scored is a claim carrying a confidence, kept as a constructor so a caller
|
||||
// cannot forget HasScore and have a real 0.0 read as "no score".
|
||||
func Scored(stage, claimant, intent string, score float64, outcome, reason string) Claim {
|
||||
return Claim{
|
||||
Stage: stage, Claimant: claimant, Intent: intent,
|
||||
Score: score, HasScore: true, Outcome: outcome, Reason: reason,
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,98 @@
|
||||
package decision
|
||||
|
||||
import (
|
||||
"context"
|
||||
"sync"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
func TestFinishNamesTheNeverAsked(t *testing.T) {
|
||||
ctx, rec := With(context.Background(), "какая погода в риме?")
|
||||
Expect(ctx, StageQuery, []string{"calendar", "weather", "search", "kiwix"})
|
||||
Note(ctx, Claim{Stage: StageQuery, Claimant: "calendar", Outcome: Declined})
|
||||
Note(ctx, Claim{Stage: StageQuery, Claimant: "weather", Outcome: Won})
|
||||
|
||||
rec.Finish(time.Now())
|
||||
|
||||
outcomes := map[string]string{}
|
||||
for _, c := range rec.Claims {
|
||||
outcomes[c.Claimant] = c.Outcome
|
||||
}
|
||||
if outcomes["weather"] != Won || rec.Winner != StageQuery+":weather" {
|
||||
t.Errorf("winner = %q, weather = %q", rec.Winner, outcomes["weather"])
|
||||
}
|
||||
if outcomes["calendar"] != Declined {
|
||||
t.Errorf("calendar = %q, want a decline", outcomes["calendar"])
|
||||
}
|
||||
// The two below the winner never looked, and saying so is the whole point.
|
||||
for _, name := range []string{"search", "kiwix"} {
|
||||
if outcomes[name] != NeverAsked {
|
||||
t.Errorf("%s = %q, want %q", name, outcomes[name], NeverAsked)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A real 0.0 confidence must not read as "this claimant has no score".
|
||||
func TestScoredKeepsAZeroScore(t *testing.T) {
|
||||
c := Scored(StageRoute, "classifier", "chat", 0, LostOnScore, "")
|
||||
if !c.HasScore || c.Score != 0 {
|
||||
t.Errorf("claim = %+v", c)
|
||||
}
|
||||
}
|
||||
|
||||
func TestNoteIfUnclaimedYieldsToARealWinner(t *testing.T) {
|
||||
ctx, rec := With(context.Background(), "x")
|
||||
Note(ctx, Claim{Stage: StageQuery, Claimant: "weather", Outcome: Won})
|
||||
rec.NoteIfUnclaimed(Claim{Stage: StageAction, Claimant: "action-handler"})
|
||||
if rec.Winner != StageQuery+":weather" {
|
||||
t.Errorf("winner = %q, want the query source", rec.Winner)
|
||||
}
|
||||
}
|
||||
|
||||
// A context with no record must cost nothing and crash nothing: that is what
|
||||
// makes the claim sites safe to leave in every test and every fixture run.
|
||||
func TestNoRecorderIsANoOp(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
Note(ctx, Claim{Claimant: "x", Outcome: Won})
|
||||
Expect(ctx, StageQuery, []string{"y"})
|
||||
if From(ctx) != nil {
|
||||
t.Error("bare context reported a record")
|
||||
}
|
||||
From(ctx).NoteIfUnclaimed(Claim{Claimant: "z"})
|
||||
if rec := From(ctx).Finish(time.Now()); rec != nil {
|
||||
t.Error("finishing a nil record produced one")
|
||||
}
|
||||
}
|
||||
|
||||
func TestRingIsBoundedAndNewestFirst(t *testing.T) {
|
||||
r := NewRing()
|
||||
for i := 0; i < ringSize+5; i++ {
|
||||
r.Push(&Record{Utterance: string(rune('a' + i))})
|
||||
}
|
||||
got := r.Recent(ringSize + 10)
|
||||
if len(got) != ringSize {
|
||||
t.Fatalf("kept %d records, want %d", len(got), ringSize)
|
||||
}
|
||||
if got[0].Utterance != string(rune('a'+ringSize+4)) {
|
||||
t.Errorf("newest = %q", got[0].Utterance)
|
||||
}
|
||||
}
|
||||
|
||||
// A query source may fan out to goroutines of its own, so two of them noting at
|
||||
// once must not race. Run under -race, which is where this earns its keep.
|
||||
func TestConcurrentNotes(t *testing.T) {
|
||||
ctx, rec := With(context.Background(), "x")
|
||||
var wg sync.WaitGroup
|
||||
for i := 0; i < 8; i++ {
|
||||
wg.Add(1)
|
||||
go func() {
|
||||
defer wg.Done()
|
||||
Note(ctx, Claim{Stage: StageQuery, Claimant: "fanout", Outcome: Declined})
|
||||
}()
|
||||
}
|
||||
wg.Wait()
|
||||
if len(rec.Claims) != 8 {
|
||||
t.Errorf("recorded %d claims, want 8", len(rec.Claims))
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,49 @@
|
||||
package decision
|
||||
|
||||
import "sync"
|
||||
|
||||
// ringSize is how many turns are kept. Turns arrive at human rate, not machine
|
||||
// rate, so the whole store is memory: no migration, no insert on the answer
|
||||
// path, and nothing of his words survives a restart. That is what makes this
|
||||
// cheap enough to leave on always (V-564). Ecosystem traces went to SQLite
|
||||
// because one act writes several hops and they must outlive the turn; an
|
||||
// arbitration record is read minutes later or never.
|
||||
const ringSize = 25
|
||||
|
||||
// Ring holds the newest records, newest first on read.
|
||||
type Ring struct {
|
||||
mu sync.Mutex
|
||||
recs []*Record
|
||||
}
|
||||
|
||||
func NewRing() *Ring { return &Ring{} }
|
||||
|
||||
// Push adds one finished record and drops the oldest past the bound.
|
||||
func (r *Ring) Push(rec *Record) {
|
||||
if r == nil || rec == nil {
|
||||
return
|
||||
}
|
||||
r.mu.Lock()
|
||||
defer r.mu.Unlock()
|
||||
r.recs = append(r.recs, rec)
|
||||
if len(r.recs) > ringSize {
|
||||
r.recs = r.recs[len(r.recs)-ringSize:]
|
||||
}
|
||||
}
|
||||
|
||||
// Recent returns up to n records, newest first.
|
||||
func (r *Ring) Recent(n int) []*Record {
|
||||
if r == nil || n <= 0 {
|
||||
return nil
|
||||
}
|
||||
r.mu.Lock()
|
||||
defer r.mu.Unlock()
|
||||
if n > len(r.recs) {
|
||||
n = len(r.recs)
|
||||
}
|
||||
out := make([]*Record, 0, n)
|
||||
for i := 0; i < n; i++ {
|
||||
out = append(out, r.recs[len(r.recs)-1-i])
|
||||
}
|
||||
return out
|
||||
}
|
||||
+137
-24
@@ -38,21 +38,34 @@ type PendingQuestion struct {
|
||||
MaxAttempts int
|
||||
}
|
||||
|
||||
// maxAttempts is MaxAttempts with the default filled in.
|
||||
func (q *PendingQuestion) maxAttempts() int {
|
||||
if q.MaxAttempts <= 0 {
|
||||
return DefaultMaxAttempts
|
||||
// Action reads the parked question as the typed action it is assembling
|
||||
// (pending.go). Derived rather than stored: the question's fields stay the one
|
||||
// copy of the truth, so a caller that fills them the old way cannot end up with
|
||||
// a capability that disagrees with the intent.
|
||||
func (q *PendingQuestion) Action() PendingAction {
|
||||
return PendingAction{
|
||||
Capability: CapabilityFor(q.Intent),
|
||||
Slots: q.Slots,
|
||||
Missing: q.Missing,
|
||||
Utterance: q.Utterance,
|
||||
Asked: q.Asked,
|
||||
TTL: q.TTL,
|
||||
Attempts: q.Attempts,
|
||||
MaxAttempts: q.MaxAttempts,
|
||||
}
|
||||
return q.MaxAttempts
|
||||
}
|
||||
|
||||
// IsExpired and CanAsk answer through the action, so there is exactly one copy
|
||||
// of the TTL and attempt-cap rules and the widening cannot drift from them.
|
||||
func (q *PendingQuestion) IsExpired(now time.Time) bool {
|
||||
return now.After(q.Asked.Add(q.TTL))
|
||||
a := q.Action()
|
||||
return a.IsExpired(now)
|
||||
}
|
||||
|
||||
// CanAsk reports whether Maven may ask another question about this request.
|
||||
func (q *PendingQuestion) CanAsk() bool {
|
||||
return q.Attempts < q.maxAttempts()
|
||||
a := q.Action()
|
||||
return a.CanAsk()
|
||||
}
|
||||
|
||||
// ClarifyStore holds the parked questions. Same shape and locking as
|
||||
@@ -66,11 +79,26 @@ func (q *PendingQuestion) CanAsk() bool {
|
||||
// next words route fresh, which is the right answer with or without a notice.
|
||||
// Do not give this store a persister without re-arguing that.
|
||||
type ClarifyStore struct {
|
||||
mu sync.RWMutex
|
||||
questions map[string]*PendingQuestion
|
||||
mu sync.RWMutex
|
||||
// stacks — one stack of parked questions per dialogue id, newest last. It
|
||||
// was a single question per id until V-559; a side query has to be able to
|
||||
// suspend the active flow and find it still there afterwards (V-561 does
|
||||
// the suspending, this only holds the room for it).
|
||||
stacks map[string][]*PendingQuestion
|
||||
defaultTTL time.Duration
|
||||
}
|
||||
|
||||
// MaxStackDepth — how many parked questions one dialogue id may hold.
|
||||
//
|
||||
// Two, not three. One is the flow he is in, one is the thing he interrupted it
|
||||
// with, and out loud he does not nest deeper than that: a side query inside a
|
||||
// side query is a shape typed conversation has and spoken conversation does
|
||||
// not. The bound is also a promise — every level she keeps is a level she must
|
||||
// be able to SPEAK when it dies (clarifyGaveUp, clarifyExpiredVariants), and
|
||||
// two lines of "and the other thing I dropped" is already the limit of what a
|
||||
// reply can carry.
|
||||
const MaxStackDepth = 2
|
||||
|
||||
func NewClarifyStore(defaultTTL time.Duration) *ClarifyStore {
|
||||
if defaultTTL <= 0 {
|
||||
// Short, like confirmTTL in voice.go: a clarifying question is a
|
||||
@@ -78,27 +106,58 @@ func NewClarifyStore(defaultTTL time.Duration) *ClarifyStore {
|
||||
defaultTTL = 90 * time.Second
|
||||
}
|
||||
return &ClarifyStore{
|
||||
questions: make(map[string]*PendingQuestion),
|
||||
stacks: make(map[string][]*PendingQuestion),
|
||||
defaultTTL: defaultTTL,
|
||||
}
|
||||
}
|
||||
|
||||
// Put parks a question. Called on a clarify decision (cmd/mavend/clarify.go).
|
||||
// Put parks a question, replacing the one on top. Called on a clarify decision
|
||||
// (cmd/mavend/clarify.go), and it is still what the daemon uses: re-asking the
|
||||
// same request is a new question about the SAME action, so it overwrites rather
|
||||
// than growing the stack. Push is the deeper one, and nothing calls it yet.
|
||||
func (s *ClarifyStore) Put(id string, q *PendingQuestion) {
|
||||
if q.TTL <= 0 {
|
||||
q.TTL = s.defaultTTL
|
||||
}
|
||||
s.fillTTL(q)
|
||||
s.mu.Lock()
|
||||
s.questions[id] = q
|
||||
s.mu.Unlock()
|
||||
defer s.mu.Unlock()
|
||||
stack := s.stacks[id]
|
||||
if len(stack) == 0 {
|
||||
s.stacks[id] = []*PendingQuestion{q}
|
||||
return
|
||||
}
|
||||
stack[len(stack)-1] = q
|
||||
}
|
||||
|
||||
// Get returns the live parked question, or nil when there is none.
|
||||
func (s *ClarifyStore) Get(id string, now time.Time) *PendingQuestion {
|
||||
// Push suspends whatever is parked and puts q on top. The returned question is
|
||||
// one the depth bound forced out of the bottom of the stack, and the caller MUST
|
||||
// tell him about it — a parked request that dies without a word leaves him
|
||||
// thinking it landed, which is the whole reason clarifyGaveUp exists. nil is the
|
||||
// ordinary case.
|
||||
func (s *ClarifyStore) Push(id string, q *PendingQuestion) *PendingQuestion {
|
||||
s.fillTTL(q)
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
stack := append(s.stacks[id], q)
|
||||
var dropped *PendingQuestion
|
||||
if len(stack) > MaxStackDepth {
|
||||
dropped = stack[0]
|
||||
stack = stack[1:]
|
||||
}
|
||||
s.stacks[id] = stack
|
||||
return dropped
|
||||
}
|
||||
|
||||
// Peek returns the live question on top, or nil when there is none. Expired
|
||||
// entries below it are left alone: TakeExpired is what reports those, and
|
||||
// dropping one here would be the silent death this store is careful about.
|
||||
func (s *ClarifyStore) Peek(id string, now time.Time) *PendingQuestion {
|
||||
s.mu.RLock()
|
||||
q, ok := s.questions[id]
|
||||
stack := s.stacks[id]
|
||||
var q *PendingQuestion
|
||||
if len(stack) > 0 {
|
||||
q = stack[len(stack)-1]
|
||||
}
|
||||
s.mu.RUnlock()
|
||||
if !ok {
|
||||
if q == nil {
|
||||
return nil
|
||||
}
|
||||
if q.IsExpired(now) {
|
||||
@@ -108,27 +167,81 @@ func (s *ClarifyStore) Get(id string, now time.Time) *PendingQuestion {
|
||||
return q
|
||||
}
|
||||
|
||||
// Get is Peek under the name every caller already uses. Kept because a clarify
|
||||
// answer is always about the top of the stack, so the two are the same call.
|
||||
func (s *ClarifyStore) Get(id string, now time.Time) *PendingQuestion {
|
||||
return s.Peek(id, now)
|
||||
}
|
||||
|
||||
// Pop takes the live question off the top and returns it, so the flow beneath
|
||||
// becomes current again. nil when the top is empty or expired — an expired top
|
||||
// is dropped along with the rest of the stack, exactly as Peek does, because the
|
||||
// clock that killed it has been running for everything underneath too.
|
||||
func (s *ClarifyStore) Pop(id string, now time.Time) *PendingQuestion {
|
||||
s.mu.Lock()
|
||||
stack := s.stacks[id]
|
||||
if len(stack) == 0 {
|
||||
s.mu.Unlock()
|
||||
return nil
|
||||
}
|
||||
q := stack[len(stack)-1]
|
||||
if q.IsExpired(now) {
|
||||
delete(s.stacks, id)
|
||||
s.mu.Unlock()
|
||||
return nil
|
||||
}
|
||||
if len(stack) == 1 {
|
||||
delete(s.stacks, id)
|
||||
} else {
|
||||
s.stacks[id] = stack[:len(stack)-1]
|
||||
}
|
||||
s.mu.Unlock()
|
||||
return q
|
||||
}
|
||||
|
||||
// Depth — how many questions are parked for this id, expired ones included.
|
||||
// Diagnostic; the arbiter in V-560 reads it to know it is inside a flow.
|
||||
func (s *ClarifyStore) Depth(id string) int {
|
||||
s.mu.RLock()
|
||||
defer s.mu.RUnlock()
|
||||
return len(s.stacks[id])
|
||||
}
|
||||
|
||||
// TakeExpired reports whether a question was parked here but its TTL ran out,
|
||||
// and drops it. Get drops such a question silently, which leaves the user
|
||||
// thinking his request is still alive — the caller uses this to tell him it is
|
||||
// gone before treating his words as a fresh utterance.
|
||||
//
|
||||
// It looks at the top only, and drops the whole stack when that one is dead: one
|
||||
// notice is what a reply can carry, and anything parked under a question that
|
||||
// timed out has been waiting at least as long.
|
||||
func (s *ClarifyStore) TakeExpired(id string, now time.Time) bool {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
q, ok := s.questions[id]
|
||||
if !ok || !q.IsExpired(now) {
|
||||
stack := s.stacks[id]
|
||||
if len(stack) == 0 || !stack[len(stack)-1].IsExpired(now) {
|
||||
return false
|
||||
}
|
||||
delete(s.questions, id)
|
||||
delete(s.stacks, id)
|
||||
return true
|
||||
}
|
||||
|
||||
// Delete drops every question parked for this id. The old single-slot Delete
|
||||
// under the old name: at depth one the two are the same, and a caller that means
|
||||
// "this exchange is over" means all of it.
|
||||
func (s *ClarifyStore) Delete(id string) {
|
||||
s.mu.Lock()
|
||||
delete(s.questions, id)
|
||||
delete(s.stacks, id)
|
||||
s.mu.Unlock()
|
||||
}
|
||||
|
||||
// fillTTL applies the store default to a question parked without one.
|
||||
func (s *ClarifyStore) fillTTL(q *PendingQuestion) {
|
||||
if q.TTL <= 0 {
|
||||
q.TTL = s.defaultTTL
|
||||
}
|
||||
}
|
||||
|
||||
// Answer merges the slots parsed from the user's answer into the parked ones.
|
||||
// Only the slots listed in Missing are touched. Within those, a value the answer
|
||||
// carries WINS over what was parked: she asked about this slot, so «нет, в пять»
|
||||
|
||||
@@ -0,0 +1,101 @@
|
||||
package dialogue
|
||||
|
||||
import "time"
|
||||
|
||||
// Capability names the thing being assembled across a clarify exchange —
|
||||
// "reminder.create", not "reminder". A router intent says what she heard; a
|
||||
// capability says what she is about to do, and those are not the same word:
|
||||
// three intents currently reach exactly one capability each, but a fact key
|
||||
// that turns out to be a Hexis target does not. Named in the ecosystem's
|
||||
// dotted form because that is what a confirmation binds (cmd/mavend/confirm.go)
|
||||
// and what Hexis registers.
|
||||
//
|
||||
// This package must stay free of internal/router (the cycle rule that makes
|
||||
// Slots a hand-kept copy), so the mapping from an intent lives here and reads
|
||||
// off dialogue.Intent only.
|
||||
type Capability string
|
||||
|
||||
const (
|
||||
CapReminderCreate Capability = "reminder.create"
|
||||
CapFactWrite Capability = "fact.write"
|
||||
CapNoteWrite Capability = "note.write"
|
||||
CapActRun Capability = "act.run"
|
||||
CapQueryAnswer Capability = "query.answer"
|
||||
CapChatReply Capability = "chat.reply"
|
||||
CapSystemControl Capability = "system.control"
|
||||
)
|
||||
|
||||
// intentCapability — the one place an intent becomes a capability. Every intent
|
||||
// is listed, including the four that are never worth a clarifying question, so a
|
||||
// parked action always knows what it is even when nothing asks it.
|
||||
var intentCapability = map[Intent]Capability{
|
||||
IntentReminder: CapReminderCreate,
|
||||
IntentFact: CapFactWrite,
|
||||
IntentNote: CapNoteWrite,
|
||||
IntentAct: CapActRun,
|
||||
IntentQuery: CapQueryAnswer,
|
||||
IntentChat: CapChatReply,
|
||||
IntentSystem: CapSystemControl,
|
||||
}
|
||||
|
||||
// CapabilityFor maps a router intent (already narrowed to dialogue.Intent by
|
||||
// the caller) to the capability being assembled. "" for an intent she does not
|
||||
// recognise — an unknown intent must not silently become a real capability.
|
||||
func CapabilityFor(in Intent) Capability {
|
||||
return intentCapability[in]
|
||||
}
|
||||
|
||||
// PendingAction is the action Maven is assembling, as an object rather than as
|
||||
// conversational history: which capability, the slots it already has, the slots
|
||||
// it is still missing, when she asked, how many questions that has cost and how
|
||||
// long the answer stays welcome.
|
||||
//
|
||||
// It exists because the resolver used to have to infer all of that from a
|
||||
// parked question plus the previous turn (Vikunja #558): "is this his answer or
|
||||
// a new request" is answerable against an object and guessy against a
|
||||
// transcript. PendingQuestion carries one of these and keeps its own flat
|
||||
// fields, so this is a widening — nothing reads the capability yet.
|
||||
type PendingAction struct {
|
||||
Capability Capability
|
||||
Slots Slots // what is filled so far
|
||||
Missing []Slot // what she is waiting for, in the order to ask about
|
||||
Utterance string // his original raw words, as the action's provenance
|
||||
Asked time.Time
|
||||
TTL time.Duration
|
||||
Attempts int // questions already asked about this action
|
||||
// MaxAttempts caps Attempts. 0 ⇒ DefaultMaxAttempts.
|
||||
MaxAttempts int
|
||||
}
|
||||
|
||||
// maxAttempts is MaxAttempts with the default filled in.
|
||||
func (a *PendingAction) maxAttempts() int {
|
||||
if a.MaxAttempts <= 0 {
|
||||
return DefaultMaxAttempts
|
||||
}
|
||||
return a.MaxAttempts
|
||||
}
|
||||
|
||||
// IsExpired — the answer came too late for this action to still be his answer.
|
||||
func (a *PendingAction) IsExpired(now time.Time) bool {
|
||||
return now.After(a.Asked.Add(a.TTL))
|
||||
}
|
||||
|
||||
// CanAsk reports whether she may ask another question about this action.
|
||||
func (a *PendingAction) CanAsk() bool {
|
||||
return a.Attempts < a.maxAttempts()
|
||||
}
|
||||
|
||||
// Gaps lists the slots this action asked for and still does not have. Computed
|
||||
// from the slots rather than trusted from Missing, because Missing is what she
|
||||
// asked about and the slots are what she got — an answer can fill a gap she
|
||||
// never asked about, and a re-park must not ask again for something now filled.
|
||||
func (a *PendingAction) Gaps() []Slot {
|
||||
return StillMissing(a.Missing, a.Slots)
|
||||
}
|
||||
|
||||
// Complete reports whether every slot this action was waiting for is filled, so
|
||||
// it can run. Note that this is completeness against what she ASKED, not
|
||||
// against the capability's whole schema — validating that is V-562.
|
||||
func (a *PendingAction) Complete() bool {
|
||||
return len(a.Gaps()) == 0
|
||||
}
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user