Compare commits

...

201 Commits

Author SHA1 Message Date
kami 08f3db318f Query Praxis by canonical entity ref and back off enrichment retries (#272)
Add an entity-scoped attention capability: the subject is resolved to a
canonical Nexus entity_id, the id travels to Praxis as a query scope
instead of being dropped after resolution, and Maven's own facts already
tagged with the same id join the answer. Ambiguous, unknown, degraded and
no-Nexus cases each get a distinct reply and never a scoped query without
a scope.

Give the fact-enrichment worker per-fact exponential backoff capped at an
hour and a status report of pending/in-backoff/worst-attempt counts, so a
long Nexus outage shows as a visible backlog rather than facts that
silently never got tagged. Nothing is ever given up on.
2026-08-01 06:52:25 +04:00
kami 69e2800ef3 Cover ecosystem degraded modes with a shared fault-injection harness (#276)
Extend the fake Nexus/Praxis/Hexis harness with request header and query
capture, a malformed-body lever, a response delay lever, and a request
counter, then add a degraded-mode suite on top of it: independent outages,
malformed and drifted contracts, cancellation, execution failure vs
transport failure, ambiguous targets, no autonomous Praxis to Hexis
chaining, confirmation for mutating capabilities, and recovery without a
restart.
2026-08-01 06:47:55 +04:00
kami a8fcb404be Scan the LAN, bounded to configured subnets (#257)
internal/netscan/ discovers hosts on the network Maven is configured to look at:
a TCP-connect scan (net.DialTimeout, no raw sockets, no privileges) plus a read
of the kernel's ARP cache. Wired as a read-only query source, "network", so
"какие устройства в сети?" is answered by a scan instead of by whatever old note
happens to be nearest.

Scanning is a read, but an unbounded scanner on a home LAN is noisy and easy to
point somewhere it should not go, so the package is built around four bounds:

  - Scan takes NO target argument. The range comes from the config block and
    from nowhere else, so there is no exported way to scan an arbitrary prefix
    and nothing an utterance, the router, or a scanned host says can retarget
    it. That is asserted directly: the test watches every address handed to the
    dialer and fails if one falls outside the configured prefix. The ARP cache —
    the one input the network itself populates — is filtered to the configured
    range for the same reason.
  - Every configured CIDR must be private (RFC1918 / CGNAT / link-local) and no
    larger than 1024 addresses. 8.8.8.0/24, 0.0.0.0/0 and 10.0.0.0/8 are refused
    at config load, not after the packets have left.
  - Rate-limited to a configured connections-per-second across the whole scan,
    so it looks like background traffic rather than a portscan.
  - Bounded in total by MaxHosts, a per-connection timeout, a 20s turn budget
    and the context; a canceled scan stops dialing immediately.

Off unless configured: dark without "enabled": true, and applyDefaults
normalises a disabled block to nil. deploy/mavend.json carries it disabled.

BLUETOOTH IS NOT SHIPPED, AND IS BLOCKED, NOT SKIPPED. The plan's other half
(internal/bluetooth/, RSSI presence probes) needs a bluez stack that is not
here: bluetoothctl and hcitool are not installed, bluetoothd is not installed,
the bluetooth unit is inactive, and org.bluez is not on the system bus. hci0
exists as a kernel device and nothing can talk to it. The docker deploy is
further away still — it would need host networking, the D-Bus system socket
passed in, and CAP_NET_ADMIN. Writing an exec wrapper around a binary that does
not exist, against an output format nothing here can produce, would be a guess
dressed as a feature. It needs a decision about privileging the container before
any of it is worth writing.

Vikunja #257
2026-08-01 06:35:10 +04:00
kami dc4c5b7841 Read and control the house through Home Assistant (#256)
A `smarthome` block points Maven at a Home Assistant instance. She reads its
entity states to answer "что включено дома?", and every controllable device
becomes a PROPOSED row in the existing act allowlist — cmd
["smarthome",<entity_id>,<service>], scope smarthome:<domain> — so nothing new
had to be invented for the mutating half. ProposeTool/EnableTool/DisableTool,
tool.Matcher and the confirm turn are untouched; one branch in Executor.Exec
routes such a row to the client instead of exec, and "smarthome" is never run as
a binary. This is the same trick overnight/mcp-tools used for #251, on purpose.

Discovery only ever PROPOSES, and every control row is destructive=true: there
is no read-only way to turn the heating off, so flipping something in his flat
always costs a confirm turn and always had to be enabled by hand on /tools,
behind step-up.

The entity and the service come from the row he enabled, never from the
utterance — Exec drops the spoken tail for a house row. A router that misheard
can pick the wrong lamp; it cannot compose a target of its own. The service is
checked against the domain's table on the way out too, so a hand-edited cmd
column cannot reach an arbitrary Home Assistant service. set_brightness and
set_temperature are deliberately absent: a spoken number the router got wrong is
a wrong act on real hardware, and on/off is the whole of what a voice turn can
defend.

The read side is a query source ("home", before calendar and the recall passes)
so "что нового дома?" is not answered from an old note. Its matcher needs a
house marker plus an ask plus a device word and bails out on weather wording,
because "какая температура на улице?" belongs to the weather source.

Off unless configured: the block is dark without "enabled": true, and
applyDefaults normalises a disabled block to nil so "off" stays in one place.
deploy/mavend.json carries it disabled, with the token as ${HA_TOKEN}.

NOT shipped, and not faked: MQTT / Zigbee2MQTT (plan steps 2 and 5) and the
sensor-to-fact and presence-probe pipelines. There is no broker and no Home
Assistant anywhere on this network — 8123 and 1883 are closed on every host in
192.168.1.0/24 — the module tree is vendored so a paho dependency cannot be
added offline, and Home Assistant already fronts Zigbee2MQTT where it exists.
Writing a sensor pipeline with no sensor to test it against would be a guess.

Vikunja #256
2026-08-01 06:27:39 +04:00
kami 33e53ee897 Add a replayable full-system simulator on a fake clock (#284)
A scenario is a JSON file under cmd/mavend/testdata/scenarios: a start
instant, a script of canned model answers, and a list of steps at "HH:MM".
Each step does one thing — say, audio, signal, arrive, tick, fault — and
then asserts on what she said, what was sent, which ecosystem services were
called, and what landed in the intake journal.

Between those boundaries the real components run: the real router cascade
(stage0, the LLM router over a scripted completer, the classifier
underneath it), the real store, the real reactive handler, the real tick
loop, and the same intake-decorated ipc.CoreAPI the daemon wires. What is
faked is only what a test cannot have: the model, the microphone, the
speaker, the delivery sink, and the ecosystem HTTP services.

Time is a single fakeClock threaded into every reader — the handler, the
intake publish stamp and tick(ctx, now) — so there is no time.Now() on the
replay path and a scenario is reproducible. TestSimulatorIsDeterministic
enforces that by replaying twice and diffing the transcripts byte for byte;
advanceTo refuses a step that goes backwards.

Two scenarios ship. morning_missed replays #284's own description: he
appears at the desk, a feed item, a mail candidate and a relayed
notification arrive through the morning, two ticks pass, and the assertions
are as much about nothing being sent at him unprompted as about what she
said. evening_degraded picks up the tier-2 pipeline case #288 deferred
here — a golden WAV through the STT seam to a written fact — and then puts
the ecosystem into 503 and checks that the proactive loop stays quiet and
that intake keeps working without it.

This is test-only code. Nothing in the production binaries changed, so the
daemon behaves identically when no scenario is running.

`make simulate` runs them verbose so the transcript is readable; `make
test` runs them with everything else.

Vikunja #284
2026-08-01 06:15:21 +04:00
kami 45b5e16eff Normalize every intake path into one event envelope (#283)
Things arrive at Maven from eight directions — a relayed Android
notification on POST /api/ambient, mail candidates from mavmaild, RSS
items, changed pages from the crawler, zenmoney and wg reads from
mavpoll, CalDAV events, presence probes, meeting transcripts and image
descriptions. Each grew its own shape and its own log line, and nothing
could answer "what came in today, from where".

internal/event is that answer: a flat source-agnostic envelope (Source,
Kind, EntityIDs, Title, Body, Priority, OccurredAt, Payload) plus a
bounded in-memory journal. Both are pure — Publish and Normalize take
`now` as a parameter, so no clock read sits on a path a replay would
drive.

Adopting it did not touch eight callers, because every intake path
already converges on three ipc.CoreAPI methods: WriteFact, WriteNote and
CaptureTask. cmd/mavend/intake.go decorates that ONE interface, so
mavweb, mavcaldav, mavpoll, mavmaild and the in-core feed/crawl/capture/
vision workers publish envelopes without knowing events exist. The lone
exception is cmd/mavend/mail.go, which captures through the store
directly and now publishes explicitly.

Nothing dispatches on an event. It is a report that something arrived,
never an instruction to speak — "a feed item appeared" becoming a
notification is the nag this repo refuses. Digestion may read the
journal later; it will still go through internal/loop's rules and the
severity/presence routing table.

Read surface: ipc.MethodRecentEvents (AuthRead, daemon-cached like
TickTrace — a bare store cannot serve a ring) and a read-only /events
page in mavweb.

Production is unchanged when nobody is watching: a nil *event.Bus makes
Publish a no-op and newIntakeAPI returns the wrapped API untouched, so
config.intake_journal < 0 leaves no decorator on the call path at all.
The default is 512 entries; the "off unless configured" rule is for
capabilities that reach out, and a bounded in-memory log of writes core
already performed reaches nowhere.

Verified: make build, make test (go test -race) both clean. New tests
cover the envelope and ring (internal/event, 95.7%), the decorator's
invariants — a failed write publishes nothing, a deduped capture
publishes nothing, OccurredAt is the fact's Ts and not notice time — and
the /events page including escaping of feed-supplied titles.
2026-08-01 06:05:00 +04:00
kami 4eca20bd94 Derive the cold-start unlock key from the passkey PRF, not the public key (#14)
Cold-start unlock wrapped the database key under the credential *public* key.
A public key is public: mavweb writes it verbatim to passkeys.json, normally in
the same state dir as db_key.wrapped, so anyone holding both files recovered the
database key offline with no authenticator involved. The wrapped blob was a
plaintext key with extra steps.

The secret is now the WebAuthn PRF extension output — 32 bytes the authenticator
computes over a fixed salt and never stores anywhere. The blob gains a version:

  v2:  "MVNKW2\x00" || salt || nonce || AES-256-GCM(key), magic as AAD
  v1:  salt || nonce || AES-256-GCM(key)                  (read-only)

v1 still opens so an existing deployment is not bricked, and reports itself so
the daemon can log a SECURITY line telling him to re-enroll. Nothing writes v1.
The magic is authenticated, so a v2 blob cannot be stripped and re-read as v1.

Four other defects on the same path:

  - The locked-boot store was opened on an IPC goroutine inside UnlockFn and
    never closed. Close is what re-encrypts the tmpfs working copy back over
    the ciphertext, so every write of a cold-started session was lost silently
    on the next boot. daemonLock now owns the store and seals it at shutdown.
  - MethodUnlock was reachable by anything on the box; the socket is same-uid
    and cannot authenticate its caller. It now requires a passkey assertion
    that mavweb verified first.
  - Concurrent unlocks would each open a store and wire a daemon. One at a
    time, and never a second one.
  - The hand-rolled HKDF keyed the expand step with the salt instead of the
    PRK. Replaced with crypto/hkdf.

Key wrapping moves from enrolment to the first assertion, because create() does
not produce a PRF result on most authenticators — only a support flag. An
authenticator without PRF now writes no wrapped file at all rather than one
that looks protected and is not, and the page says so.

Verified: make build, make test. New tests cover the v2 round trip, a wrong
secret, every single-bit tamper, truncation, the v1 downgrade attempt, legacy
v1 reads, non-32-byte and all-zero secrets, the ipc wire field, locked-mode
default-deny, a forged assertion never reaching the unlock path, seal-on-
shutdown after a cold start, and that nothing in the state dir contains the
plaintext key. The PRF round trip against real hardware is a QA step.

Vikunja #14
2026-08-01 05:49:27 +04:00
kami fed33a4e16 Stop mavwaked from hearing itself, and add barge-in (#287)
Playback was `go playAudio(reply)` — fire and forget, nobody holding the
process handle. Two audible consequences fell out of that.

She answered herself. The capture loop kept feeding the VAD while the
speaker was running, so her own reply came back in through the mic,
tripped the VAD, and was shipped to the daemon as a fresh command. There
is no acoustic echo canceller in this pipeline, so the fix is
half-duplex: while she is speaking, the capture side is muted. That part
is unconditional — it repairs a defect, it is not a new capability.

And talking over her did nothing, because there was no handle to cancel.
-barge-in now cuts playback when sustained energy clears a room-tuned
threshold (-barge-in-rms, default 0.12 normalised, over -barge-in-frames
consecutive frames, default 5). It is off by default: without an echo
canceller the only way to tell "he is talking over her" from "the mic is
hearing her" is that he is much louder, and how much louder depends on
where the mic sits.

The frame decision moved out of main.go into session.feed, behind a
player and an utteranceSender interface, so all of it is testable with
no mic, no speaker and no daemon. Nine tests cover the self-hearing
case, the off-by-default case, the consecutive-frame requirement,
speaker-leak-level audio not triggering, capturing the interrupting
utterance after a cut, and failed round-trips not starting playback.

The other seven items on #287 (partial STT, per-segment retry, mic
profiles, noise-floor calibration, short-response-while-speaking) are
untouched and stay on the task.
2026-08-01 05:36:13 +04:00
kami 62cc072f8c Add golden-audio STT tests against real whisper.cpp (#288)
Four committed WAV fixtures go through the real whisper.cpp binding in
cmd/mavsttd, so a wrong model, a wrong language hint, a broken resample
or a regressed silence gate fails `make test` instead of surfacing as
Maven mishearing him.

The fixtures are piper-synthesised, not recorded: scripts/gen-stt-fixtures.sh
drives the vendored piper with the ru_RU-irina voice Maven already speaks
with, so nothing of the owner's voice is committed and every fixture is
reproducible. 360K total for three Russian clips and one English.

Matching is tolerant on purpose. Golden transcripts move with the model,
so each case asserts intent-carrying keywords (prefix match, so Russian
inflection does not fail it) plus a word error rate ceiling, not an exact
string. The matcher is unit-tested on its own and needs no model.

TestGoldenAudioTranscription skips when models/stt/ggml-small.bin is
absent, so `make test` still passes on a box without models.
TestGoldenFixturesAreCanonical runs everywhere and checks the WAVs are
16k mono s16le and would clear mavsttd's own silence gate.
2026-08-01 05:32:10 +04:00
kami 7c7bd8ceeb Ship voice enrolment, and report recognition as blocked (#255)
Maven can now be told who someone is. She cannot yet tell who is speaking,
and this commit is careful to say so rather than pretend otherwise.

What works: profiles are enrolled from several deliberately recorded samples,
listed, and deleted. They live in the existing memory_vectors table under a
"speaker:" id prefix, so there is no migration; what that needed was a wider
interface than memory.Store, hence memory.Catalog with ByPrefix and Delete.
Delete is the load-bearing half — a voiceprint someone asked to be rid of has
to actually go, and a search-only store cannot do that. InMemoryStore.Insert
became an upsert by id to match what the persistent store already did.

What does not work, and why it is not faked: there is no speaker-embedding
model on this box. Sixteen ggufs in /mnt/hdd1/llms, all text; no ECAPA, no
x-vector, no titanet, no wespeaker, no .onnx anywhere under /mnt/hdd1. So
newSpeakerEmbedder returns nil, internal/speaker falls back to
speaker.Disabled, Identify answers ErrDisabled, and the daemon logs which
half is off at startup. The plan's "simple MFCC + GMM" floor is refused in
the package comment: MFCC cosine distance detects channel and loudness as
much as voice, and a biometric that is confidently wrong writes false claims
about named people into his memory. A bad floor is worse than none here.

Refused as well, and the reason is in enroll.go's doc comment: the plan asked
for unknown speakers to be enrolled on first interaction with a TTS "кто
это?". There is no request shape in the protocol that could express that.
Taking a biometric of whoever walks past the microphone does it to guests who
are not party to the exchange, and a synthesised question into a room is not
consent from whoever answers.

Authority: enrolment is AuthStepUp, because it is a deliberate sit-down act
that writes a biometric of a named person and never something done by voice
mid-conversation. Deletion is one rung lower at AuthWrite, deliberately
inverting the usual pattern — getting rid of a biometric must never be the
harder half. Listing is AuthRead and never returns the vectors themselves.

Off unless configured: no speaker block means the three methods answer
ErrUnknownMethod, so a default box has no wire path that takes a voiceprint.

make build and make test pass.

Vikunja #255

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 05:23:03 +04:00
kami aa1a26532c Add meeting capture with explicit start and stop (#253)
Maven can record a meeting when she is told to, transcribe it through the
STT she already has, and write a summary note. The audio lives in the blob
store #252 introduced, under the same retention loop.

Nothing here listens. Recorder.Append is the only way audio enters and it
refuses every frame unless someone explicitly started a session, so audio
arriving at an idle core is dropped rather than buffered. The plan document
asked for a keyword trigger ("maven record" heard in the room) and that is
refused: noticing a keyword means listening to the room, which is the one
behaviour this capability must not have.

Off unless configured twice over. No media block means nowhere to keep
audio, no capture block means no recorder, and in either case the four IPC
methods answer ErrUnknownMethod. On an unconfigured box there is no wire
path that begins a recording at all.

A forgotten session ends itself at max_minutes, checked on every append,
and the audio collected before the cap is kept. Stop with discard set is
what "забудь, не записывай" maps to and it leaves nothing behind. The
verbatim transcript is not saved unless save_transcript says so; the
summary is.

Long audio against n_ctx 4096 is handled by map-reduce over 3000-rune
windows rather than by truncation, because a truncated meeting summary
reads as complete and is not. Transcription is windowed at five minutes so
the whisper worker stays responsive to the voice path.

No second STT: internal/capture takes the stt.Transcriber the voice path
already holds. Capture with voice off is refused rather than degraded,
since hours of unreadable audio of other people is worse than no recording.

The three write methods are AuthWrite, not AuthStepUp: step-up needs a
passkey gesture the voice path cannot make, which would leave "запиши
встречу" impossible by voice. capture_status is AuthRead.

make build and make test both pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 05:08:08 +04:00
kami d92349ca6e Store and describe images through a shared media intake (#252)
Vision needs a second model this box does not have, so the shipped half is
the part that works without one: an image arrives, is sniffed, is stored
content-addressed, and is prepared for inference. The describing half is
written and tested against a fake server, and refuses any endpoint that is
not on this box.

internal/media is the intake all three senses share — hearing and speaker
recognition store their audio in the same place under the same retention.
Blobs stay out of the sqlite store; only the derived text becomes a note,
and only when the caller asks. Retention is enforced by an hourly prune
loop rather than by a comment.

The plan's RemoteProvider step is refused: no cloud model, inference stays
on the box, and vision.NewLocal validates that at construction.
2026-08-01 04:53:07 +04:00
kami 8d5e357b57 Expose discovered MCP tools through the act allowlist (#251)
Second half of the MCP client: the tools the manager discovers become rows in
the existing act allowlist instead of a parallel capability system.

An MCP tool is encoded in the columns that already exist — cmd
["mcp",<server>,<tool>], scope mcp:<server> — so no migration, and
ProposeTool/EnableTool/DisableTool, tool.Matcher and the confirm turn need no
changes. One branch in Executor.Exec routes such a row to the manager instead
of exec, and "mcp" is never run as a binary.

Discovery only ever PROPOSES. destructive comes from the inverse of the MCP
readOnlyHint, so a tool that does not promise to be read-only inherits the
confirm turn, and enabling stays on /tools behind step-up.

Voice args are positional and MCP args are named, so CallPositional binds only
what it can defend: no required properties runs bare, and a read-only tool with
exactly one required string or number gets the tail. Everything else refuses
with ErrNeedsArgs rather than guessing. The read-only condition was learned
against the live Vikunja server: update_task requires only task_id and takes
the rest as optional, so one guessed argument blanked the fields it did not
mention. A partially-filled write destroys what it omits, so a mutating tool
never receives a guessed argument.

Also: a read-only mcp_servers IPC method and an "MCP servers" card on /tools
showing transport, target and state, with the trust level of a local target
spelled out. There is deliberately no call-a-tool IPC method and no run button,
so mutation keeps exactly one path.

Vikunja #251
2026-08-01 04:36:40 +04:00
kami 95ae900a58 Talk MCP: a client for external tool servers (#251)
docs/plans/06-mcp-support.md asks for the host direction — Maven connects OUT
to MCP servers and consumes what they offer. This is the client half: the
protocol, the transports, the connection manager, the config block. Nothing is
wired into a turn yet, and nothing here exposes Maven's own capabilities to an
outside caller.

internal/mcp:
  - hand-rolled JSON-RPC 2.0 (the wire format is four fields, and the repo
    vendors its deps, so a library would cost more than it saves);
  - two transports: a stdio subprocess on this box, and streamable HTTP, which
    accepts a plain JSON reply or an SSE frame because servers disagree about
    which they send;
  - Client: initialize handshake, tools/list, tools/call, resources/list,
    resources/read. Text content only — everything downstream is a sentence;
  - Manager: lazy dial, per-server failure that never blocks boot or the other
    servers, backoff reconnect, Status for a web surface, graceful Close;
  - the allowlist encoding: a discovered tool becomes the store row
    "vikunja_list_tasks" with cmd ["mcp","vikunja","list_tasks"], scope
    "mcp:vikunja". No new column, no migration, and ProposeTool, EnableTool,
    the act matcher and the confirm turn all keep working untouched.

Constraints held, in code rather than in prose:
  - OFF unless configured, and a server is dark until "enabled": true.
  - A url server goes through internal/webfetch, so the SSRF guard, the size
    cap, the redirect cap and the per-host rate limit apply. Reaching loopback
    needs allow_private on THAT server, and each server gets its own fetcher so
    one loopback exemption cannot become a hole for a public endpoint.
  - readOnlyHint decides destructive: no hint means "assume it mutates", which
    will route the call through the existing confirm turn. Guessing wrong in
    that direction only costs a question.
  - The catalogue stays small on purpose — allow_tools, and max_tools=12 per
    server. The resident model is a 1.7B with a 4096-token context; a tool name
    it half-remembers is a wrong act.
  - Only the tool name and the router's arguments are sent. There is no API
    here through which a note, a fact or the persona block could travel.

webfetch grows Post (JSON-RPC cannot be a GET) and surfaces response headers
for Mcp-Session-Id. It shares Get's guards exactly: a body buys a caller
nothing, a POST to the LAN is refused for the same reason a GET is.

Verified against the real Vikunja MCP server on homesrv
(http://localhost:9100/mcp): handshake, three discovered tools with update_task
correctly NOT read-only, a live list_projects call, a tool excluded by
allow_tools refused, and the same server refused outright once allow_private
was dropped. Tests cover both transports (the stdio one against a real
subprocess), SSE and JSON framing, session echo, reconnect, and the config
validation.
2026-08-01 04:22:52 +04:00
kami be066a4b04 Deploy a new build with verification and automatic rollback (#249)
internal/update applies a new build of Maven to the box she runs on and
undoes it when the new build does not come up. cmd/mavupdate is the only
trigger: a CLI the owner runs on the host.

Apply is health-check the running daemon, snapshot the deployed artifacts,
make build, make test, install, restart, health-check — and restore the
snapshot on any failure. The order is load-bearing:

  - The preflight health check refuses to update a daemon that is already
    not answering. Without a working baseline, a failed update and a box
    that was already broken are indistinguishable, and the rollback has
    nothing to prove itself against.
  - The snapshot is taken BEFORE the build, because make build writes its
    binaries into the working tree and on the docker deployment the tree
    is the install dir — snapshotting afterwards would snapshot the new
    artifacts and leave nothing to roll back to.
  - Verification is make build plus make test, before anything is
    deployed, so a broken tree costs time and nothing else. A failed
    verify also puts the tree's artifacts back, so a later restart by
    hand cannot deploy code that failed its own tests.
  - The rollback depends on nothing that just changed: byte-for-byte
    copies out of the snapshot dir, sha256-verified on the way in, and
    the same restart command. No build, no migration, no cooperation from
    the code being replaced. It also runs on an uncancellable context —
    a rollback interrupted halfway is worse than the failure that caused
    it. When the restore itself fails it says so and names the directory
    to copy back by hand rather than reporting a tidy rollback.

Off unless configured, and the refusals are code, not documentation. The
daemon does not import this package: there is no IPC method, no web route,
no timer and no act that can start an update, so nothing Maven says or
routes reaches it. Nothing fetches code — the new version is whatever the
owner pulled into the tree. The plan's release checker, auto-update
channel and in-process crash-loop supervisor are deliberately absent; a
process cannot reliably notice that it keeps dying, and restart-on-crash
belongs to compose or systemd. The database is never snapshotted or rolled
back; schema compatibility stays store.Migrate's job.

The config is refused at load without a health socket, since an update
that cannot check its own result cannot roll back, and refused when the
snapshot dir is inside the install dir, since a restore must not read from
what the install writes.

Vikunja #249
2026-08-01 04:09:30 +04:00
kami ad074cea31 Swap the resident model without restarting mavend (#250)
Loading a different gguf was a one-line edit to phraser.model_path plus a
restart. It is now an owner-triggered IPC call, off unless configured.

internal/phraser/swap.go holds the safety properties as code:

  - Never two models resident. The old llama-server is killed and reaped
    before the new one is launched. One 1.7B fits the Vega iGPU; a
    blue/green overlap would OOM the box, so it is not offered.
  - Atomic from a turn's point of view. Swap drains the in-flight turns
    (they finish on the old model), then refuses arrivals with ErrSwapping
    until the new server has answered /v1/models. No turn ever sees half a
    swap; refused turns fall back to the classifier cascade.
  - A failed load rolls back. If the new model does not start or does not
    probe, the previous one is reloaded and the call returns RolledBack
    with the error. If the rollback also fails the daemon says so and
    degrades to the classifier rather than pretending to serve.

Holders of the completion client are re-pointed, not rebuilt: llm.Client
guards its base URL and LLMPhraser.OnSwap re-points it, so the router, the
replier, the mail extractor and the memory evaluator follow the new port
without knowing a swap happened.

Reach is deliberately narrow. phraser.swap_models is an exact-match
allowlist of absolute paths a human wrote, rejected at startup otherwise,
so "swap the model" can never mean "load any file on my disk"; the running
model is always swappable back to. MethodSwapModel is AuthStepUp, the same
rung as mutating the tool allowlist, and /models gates POST through the
same stepUpOK the tools page uses. Nothing calls Swap on a timer and no
act, intent or utterance reaches it.

Vikunja #250
2026-08-01 03:59:08 +04:00
kami 2c1b0eede0 Read a web page when he names one, and watch a few on a timer (#259)
The network fallback behind the local sources, off unless configured.

internal/crawl is pure: a stdlib robots.txt parser (group specificity,
wildcards, Crawl-delay, cached per host), HTML-to-plaintext extraction, and a
watcher that notes a watched page only when its text changed. It has no store
access and no net/http; cmd/mavend/crawls.go is the impure half.

Every limit is code and tested: the guarded fetcher from #258 enforces the host
allowlist/denylist, refuses private addresses in the dialer Control hook (so DNS
rebinding and each redirect hop are covered), caps size and redirects, times out,
and spaces requests per host. A robots.txt Disallow is refused with no override.

On demand, reading is a query source placed last in the chain, after his memory,
his notes, and the local Kiwix ZIMs once those are wired: no URL in the
utterance means no fetch, and only the URL ever leaves the box. Scheduled
watches write notes and announce nothing.

The vendored tree has no x/net/html, goquery or temoto/robotstxt, so the parsers
are stdlib. No new dependency.
2026-08-01 03:40:21 +04:00
kami cb3641e7bb Read RSS and Atom feeds, and speak about them only when asked (#258)
internal/rss parses RSS 2.0 and Atom, and polls each configured feed on its own
interval; internal/webfetch is the one door either of them uses to touch the
network. The poller writes items as notes with source "rss:<feed>" and nothing
else: the answer path reads them back when he asks "что нового в лентах?", and
nothing is announced on arrival. A feed that dispatched would be a nag, which is
why the plan's breaking-news rule was left out rather than built.

webfetch is where the limits live, as code rather than a paragraph: http(s)
only, an allowlist (the configured feeds' hosts) and a denylist, a 2 MiB body
cap, a 3-redirect cap, one request per host per second, and a refusal to connect
to any private address — checked in the dialer's Control hook so it holds for
every resolved address and every redirect hop, not just for a literal IP.

Off unless configured: no "feeds" block, no poller, no outbound request. How far
a feed was read is a config fact (rss:latest:<name>), so a restart does not
re-note yesterday's headlines.
2026-08-01 03:27:45 +04:00
kami ee7bec11e3 Add mavmaild, the read-only IMAP poller that feeds mail intake (#246)
The extraction seam landed on the previous branch but nothing fed it. This
adds the daemon that does: every interval it opens one mailbox read-only
(EXAMINE + BODY.PEEK, so reading leaves no \Seen behind), fetches the UIDs
it has not handed over yet, and posts each message to core over
ingest_mail. Core runs the model and writes task candidates; this daemon
writes nothing and cannot create a reminder.

It is a separate daemon because of the credential. mavpoll set the
precedent with the zenmoney token (#125): the module talking to the third
party holds the secret, reads it from a file so it never lands in argv, in
docker-compose.yml or in shell history, and core never sees it. There is
deliberately no -password flag, and a test asserts that.

Off unless configured at both ends: without -password-file the daemon
refuses to start, and if core has no email block the first ingest returns
ErrUnknownMethod, which disables the reader instead of hammering a socket
that will keep refusing. A seen-UID state file (0600, atomic write) keeps a
restart from re-extracting the whole lookback window; correctness does not
depend on it, since capture dedupes on normalised text. Logs are counts and
UIDs — no subject, sender or body.

Verified with an in-process IMAP server and a fake core: bulk mail is
filtered before core is asked, seen UIDs are not re-fetched, a failed
ingest is retried next poll, ErrUnknownMethod stops at the first message,
and state survives a restart. The live half is untested by design — no IMAP
credential exists on this box; setup is written up as QA steps.

Vikunja #246
2026-08-01 03:13:18 +04:00
kami f42d1594ef Turn a mail into task candidates, and into nothing else (#246)
The extraction half. internal/email.Extractor asks the resident Qwen3-1.7B,
under a GBNF grammar, what one message requires of him, and returns at most
three short candidates with an optional date.

Everything it can produce is a row in `tasks` with status "candidate",
written through the intake seam #130 built for exactly this (Source
"email:<mailbox>", Evidence = the subject line). No reminder, no fact, no
note, no calendar event. That bound is the design: a reminder FIRES, so a
1.7B misreading "встреча была в четверг" as a future appointment would wake
him up about it, whereas a wrong candidate is a line he dismisses in one
click. A due date the model read out of the mail is stored on the candidate,
where no scheduler reads it — the review page sorts by it. Relative wording
("до пятницы") is deliberately left in the text rather than resolved to a
date the model would get wrong.

The prompt is written against the two things a small model does here: it
summarises when asked to extract, and it invents an obligation out of a
polite closing line. Hence the demand for a verb phrase, and an explicit
empty array — most mail contains no task, and a model with no way to say
"nothing" says something.

Wiring: core owns extraction because llama-server lives in core's process,
so the reader hands messages over a new ipc.MethodIngestMail. It is a Server
hook (like StepUp/UnlockFn), not a CoreAPI method — not a store operation,
and no CoreAPI implementation should have to carry it. The hook stays nil
without an `email` config block or without a llama-server phraser, so the
method answers ErrUnknownMethod: off unless configured, twice over. There is
no keyword fallback on purpose — "the subject became a task" is a mailbox
rendered as a to-do list, not extraction.

Privacy: junk is refused before the model is called, mail text is never
search input, extraction errors carry byte counts rather than the reply, the
stored evidence is a truncated subject, and the log line names the mailbox
and the UID only.
2026-08-01 03:06:55 +04:00
kami b4646155b4 Read a mailbox read-only, in a client small enough to audit (#246)
internal/email is the reading half of the email reader: a ~200-line IMAP
client (LOGIN, EXAMINE, UID SEARCH SINCE, UID FETCH BODY.PEEK, LOGOUT), a
MIME-to-plaintext converter, and a header-only junk filter.

Two protocol choices are the design, not shortcuts. EXAMINE instead of
SELECT means the session is read-only at the protocol level, so no command
in it can flip a flag or expunge anything by mistake. BODY.PEEK instead of
BODY means reading a message does not mark it \Seen — Maven reads his mail
and leaves no trace of having done so, and the unread state in his own
client stays his.

Hand-rolled rather than go-imap because this is the one path that holds his
mailbox credential and reads his private mail: five commands with no
dependencies is auditable in a sitting. No IDLE and no cleartext/STARTTLS
either — an option to send his password over a plain socket is an option to
get it wrong once.

Junk is decided by headers alone, before any model is involved:
List-Unsubscribe/List-Id, Precedence: bulk, Auto-Submitted, the spam
headers, and Gmail's own category labels. Sender lists and subject keywords
are deliberately absent — they age badly and they would put his contacts in
a config file. A junk verdict only means "do not spend the model on this";
nothing is deleted and no server flag is touched.

Nothing here logs a body, a subject or an address, the junk reason names a
header rather than content, and an undecodable charset degrades to
headers-only instead of feeding the model mojibake. Verified against
recorded .eml fixtures and an in-process fake IMAP server.
2026-08-01 02:59:24 +04:00
kami da647e87d0 Read spending from zenmoney in the poller, answer it from facts (#125)
The trust boundary is zenmoney, not maven — they already hold his bank
sessions. So the poller reads /v8/diff/ and writes totals as
facts(kind=env, source=poll:zenmoney); core reads those back when he asks
and never sees the token.

internal/zenmoney sums transactions per currency over a window, skipping
tombstoned rows and transfers between his own accounts, and refuses to
encode a summary built from zero transactions. That refusal is the whole
design: a failed or empty read writes nothing and leaves the last good
total alone, because a zero recited as fact is worse than silence. No
currency conversion either — a figure he can check against his bank beats
one he cannot.

Off unless configured, and the token is read from a FILE rather than a
flag so it never lands in `ps`, in docker-compose.yml, or in shell
history. Nothing about the money is search input, no tick rule reads the
keys, and the log lines name keys, never figures.

The live-credential half is BLOCKED: there is no zenmoney account or token
here, so everything is verified against a recorded diff fixture.
2026-08-01 02:50:27 +04:00
kami bf6ccf9aea Rank captured tasks by what he actually said (#129)
Ordering is computed, not generated. Asking a 1.7B which of his tasks
matters most produces a fluent opinion with no basis in anything, and a
confidently wrong priority is worse than none — same posture as the
behaviour profile in internal/memory, which counts instead of summarising.

internal/tasks is a pure package (no ipc, no store, no cgo) holding the
score, the order and the Russian rendering, so the spoken list and the
/tasks page cannot drift. Four signals, all of them things he stated:
deadline (overdue > today > tomorrow > this week), stated urgency, age
with a cap so nothing rots at the bottom, and confirmed work always
ahead of mail-derived candidates. A task with no due date and no weight
scores nothing and carries no reason string — inventing a "потому что"
about a priority he never set is the failure mode this avoids.

Capture now picks up urgency he says out loud ("добавь в задачи срочно
оплатить интернет"), stripping the marker from the task text, and the web
add form offers the same three rungs. Ranking is a read: it sorts and
renders, never writes, schedules or announces.
2026-08-01 02:42:02 +04:00
kami 7b2b96b957 Capture tasks, with one intake seam mail can call later (#130)
A task is not a fact and not a note. A fact is a claim about the world that a
correction supersedes; a note is something to recall by meaning. A task is work
with a lifecycle, and the read that matters is "everything outstanding right
now" — which over an append-only log would mean replaying history on every
question. So: a tasks table, migration #14, statuses candidate/open/done/dropped
that each move forward exactly once.

Dedupe is on normalised text among LIVE rows only, via a partial unique index.
That is the property the mail side needs: an extractor may call CaptureTask for
every message it reads, as often as it likes, without growing the list — while a
weekly errand is still capturable again once the last one is done.

Three ways in, one seam. ipc.CaptureTaskReq is it: the voice path
(router.ParseTaskCapture on an explicit marker — "добавь в задачи …", never
"надо бы поспать"), the /tasks form, and the email reader from #246 when it
exists. Mail-derived items set Source "email:<account>", Status "candidate" and
Evidence to whatever makes the row reviewable; a candidate is inert until he
confirms it on /tasks, and Maven names it as unconfirmed when she recites the
list rather than putting words in his mouth.

No new intent — the router enum is a contract with the relabelling prompt, so
capture rides the note intent and the list rides a query source, both matched
deterministically like the calendar and plan matchers already are.

Nothing here speaks. No tick rule reads tasks; the list is answered when asked
about, which is why /tasks POST is not step-up gated the way /tools and
/routines are — a task write moves no boundary.

Vikunja #130
2026-08-01 02:33:47 +04:00
kami c8444813e2 Answer "что я обычно делаю по вторникам?" by counting, not guessing (#254)
Behavioural memory, narrowed on purpose. internal/memory/behavior.go builds a
profile out of self-facts — distinct days per weekday, median time of day — and
reads it back in RU; router.ParseHabitQuery finds the weekday deterministically;
a `habits` query source answers the question.

Three things the plan doc asks for are deliberately absent, and the doc now
records why:

- The profile is COUNTED, not LLM-generated. A 1.7B asked to summarise a year of
  habits writes fluent claims about the owner's life that no row supports, and a
  wrong claim about him is the most expensive kind of wrong maven can be.
- No cached profile fact, so no "update on fact write" machinery. It is
  recomputed on the question; a cache that can disagree with its own rows is two
  truths.
- No proactive daily plan nudge. A dispatcher proposal at 08:00 every day is the
  definition of a nag. The path from "she noticed a pattern" to "she acts on it"
  already exists in internal/pattern with the proposal queue on /routines, and it
  goes through him.

A one-off is not a habit: an activity needs two distinct days before she will
call it usual, and until then she says she does not know yet. Only self-facts
count — env rows are the world, config rows are her own tuning state. The typical
time is a median so one 03:00 outlier cannot move a morning habit into the night.
An unrecognised fact key is read back verbatim rather than glossed into something
she made up.

The source sits before "calendar" in querySources, and its matcher requires a
habit marker, so "что я делаю в среду?" still reaches the calendar — answering a
question about this coming Wednesday with a statistical average would be
answering a different question.

Verified: make build and make test both exit 0.
2026-08-01 02:21:09 +04:00
kami ed9bdd5e09 Add the day plan she can recite when asked (#128)
The plan answers "какие планы на сегодня?" by putting one day in order:
calendar events (with #126's ambient provenance carried through and hedged),
pending reminders, and one line per morning routine that still has items
outstanding. "что дальше?" trims what has already passed.

It lives in internal/morning, not in a parallel system, because it is the same
question the checklist asks at a different scale — the routine knows what is
missing from a window, the plan knows what the whole day holds, and both read
the same facts and the same idea of "today". BuildPlan is pure; tickLoop.dayPlan
is the impure half that reads the store.

It is not a nag. Nothing here fires, schedules or announces: the plan is built
only when asked, over IPC (day_plan) or on the existing /morning page.
Unprompted delivery stays with the morning nudge and the dispatcher's policy.

The query source sits before "calendar" in querySources because both match
"…на сегодня" and the plan's matcher is the more specific one; IsDayPlanQuery
matches whole words so "планёрка" (a meeting) is not read as a request for the
plan, and refuses any utterance naming another day, since the plan is built for
the clock's own day only.

Verified: make build and make test both exit 0; new tests cover plan ordering,
the checklist-only-what-is-left rule, other-day rejection, the RU rendering
against the persona checks, rest-of-day trimming, the source ordering, and the
matcher's refusals.
2026-08-01 02:15:18 +04:00
kami 49f089d8a6 Read the work calendar as a notification signal, not a mailbox (#126)
Maven does not get a work credential. A corp mail or calendar session living on
the homelab ties the box's blast radius to the employer's data, which is the
thing this task exists to refuse. What she reads instead is the signal: an
Android notification-listener on the phone relays meeting notifications over
wg/LAN to POST /api/ambient, and the ones that clearly describe a meeting become
calendar events at source=ambient:notif, confidence 0.6.

The provenance is the point. A notification is evidence about a meeting, not a
reading of a calendar, so it is never indistinguishable from one: it is stored
below full confidence, store.CalendarEvents keeps the source and confidence on
every row it returns, and the query path hedges — "похоже, Планёрка @ 14:00" for
a relayed event, plain text for a CalDAV read.

The parse is deliberately conservative (internal/calendar/ambient.go). It needs
a real clock reading and a summary that is not just that clock reading;
otherwise it stores nothing at all. A bare hour is not a time, an unread count
is not a time, and "срок 2026.08.15" does not offer 08:15 as a meeting — loose
digits in a notification are far more often a badge or a date, and a mailbox of
noise rendered as invented meetings is worse than a gap.

The ingest is off unless configured: no -ambient-token, no route registered. The
token is a shared secret compared in constant time, because the poster is a
background Android service and WebAuthn has no answer for one. The endpoint is
write-only, accepts one shape of write, and cannot read anything back out.
Reposts of the same notification dedupe against the latest fact for that
key+source, the same append-only discipline cmd/mavcaldav follows.

Not shipped: the Android relay app itself, which is a separate artifact and a
device, not Go in this repo.
2026-08-01 02:04:06 +04:00
kami 3af290152c Render maven's own reminders to a calendar she owns (#127)
Radicale becomes a write-only render target, not a store. sqlite stays
canonical: every poll mavcaldav reads the pending reminders out of core and
publishes each one as a single-event iCal resource, withdrawing the ones that
have fired or been cancelled. Losing the collection costs nothing — the next
tick rebuilds it, and nothing is ever read back from it.

It structurally cannot write to a calendar maven only reads. The render URL and
credential are their own flags, and -render-url is refused at startup when it
names the collection -url reads; the only paths it addresses carry the
maven-reminder- prefix, so even aimed at the wrong collection it can only touch
resources it created. Rendering is off unless -render-url is given.

The calendar data model now lives in one place, internal/calendar: the Event,
the iCal parse it comes from and the render it goes to, the fact key/value
encoding, and the source constants that say which calendars may be written to.
It was a parse inlined in cmd/mavcaldav and a Sprintf in two files; #126 and
#128 both need to agree with it.

Fixes a latent day-boundary bug moved out of that inline parse: it took the day
number off a local clock reading but built the window boundaries in UTC, so on
a box east of Greenwich part of the evening fell outside "today" and the poller
saw an empty calendar after 20:00 UTC. Today is now the owner's day in the
owner's location, which is what the busy gate and the day plan mean.
2026-08-01 01:55:51 +04:00
kami dc7c72a3d7 Add background memory evaluation, off unless configured (#248)
Ships the real, local, testable part of the memory-evaluation plan
(docs/plans/03-memory-evaluation.md): Maven reads back her own recent
memory on a slow ticker, asks the resident model what it notices, and
records the confident answers as notes.

internal/memeval — not internal/memory/eval.go as the plan says, because
internal/store imports internal/memory for the vector backend and an
evaluator has to read store.Fact/Note/Nudge, which would close the
cycle. Evaluate() gathers RecentFacts/RecentNotes/RecentNudges, prompts
under a GBNF grammar bounded to three {observation, confidence,
suggested_action} objects, drops anything under min_confidence,
deduplicates against what earlier runs wrote, and writes the rest as
notes with source infer:memory-eval. /dash already renders notes with
their source, so the output is visible with no UI change.

cmd/mavend/memoryeval.go drives it on its own goroutine and ticker, not
on the 60s tick: an evaluation is a multi-second round-trip on the same
llama-server that answers voice turns, and it runs hourly at most. The
memory_eval config block is absent by default and absence means the
goroutine does not exist. No llama-server phraser also means no loop —
there is no template fallback, because a "memory evaluation" assembled
from templates is a fixed sentence pretending to be an observation.

What it deliberately cannot do, since this is the feature most likely to
turn Maven into a nag:

  - It cannot speak. No dispatcher reference, no channel, no nudge. An
    observation is a thought she wrote down and he reads on /dash.
    Announcing them is a separate decision with its own opt-in.
  - It cannot act. suggested_action is recorded as text and interpreted
    by nobody — no reminder, routine or fact is created from it.
  - It says nothing about an empty store: no memory means no LLM call,
    so there are no observations invented out of two facts.
  - Its own notes are excluded from the next evaluation's input, and are
    written with a nil embedding so they stay out of the recall pool.

The plan's remaining items (dispatching observations, an /eval IPC
method and trace view, RecentEvents) and the fact that output quality is
entirely unmeasured are written up at the bottom of the plan doc.
2026-08-01 01:45:49 +04:00
kami 766ca091a7 Announce tick-inferred routines, opt-in and rate-limited (#247, #43)
The digestion tick already runs the pattern detector over all recorded
events (67563ed) and writes a proposed_routines row. What was missing is
the other half of #247: a proposal that nobody is at the mic for reaches
nothing but the /routines page, so a pattern noticed at 03:00 is only
seen if he goes looking.

This wires the tick's proposals into the existing care-delivery path
rather than a second channel: sev1 nudge, loop.Gate, dispatcher, same
routing table as an accepted routine. Restraints, since a feature that
speaks unprompted is the easiest way to turn Maven into a nag:

  - off unless configured — the new pattern_proposals block, absent by
    default, and deploy/mavend.json ships notify: false;
  - at most one announcement per tick however many patterns surfaced;
  - at most one per cooldown (24h default) across all pairs;
  - sev1, so quiet hours, away and snooze suppress it;
  - suppressed means dropped, not queued — /routines still has it;
  - once per pair for good, since proposed_routines is
    UNIQUE(action, object) and the row survives dismissal.

The body is pattern.PhraseRoutine's literal Russian, not LLM-worded, so
an inferred routine cannot arrive describing something never observed.

Also raises pattern.MinEvents from 3 to 4 — the interval-quality item on
#43. Two intervals with a ±50% band is a coincidence with a mean, not a
pattern, and now that a scan of all history can announce itself the cost
of a false positive is a permanent dismissal of that pair.
2026-08-01 01:37:50 +04:00
kami c5317eb2b4 Move the quiet-toggle and pattern-extraction slices out of voice.go (#321)
Continues the decomposition PR #50 started. voice.go 542 -> 365:

  quiet_toggle.go  144  resolveQuietToggle, quietInflections, quietStem,
                        quietTokens, quietPhrase, quietOn/OffPhrases,
                        classifyQuietToggle  (quiet_toggle_test.go already
                        existed for these)
  patterns.go     +44  detectPattern, next to detectAndPropose which it calls
                        and which patterns.go's own header already pointed at

What is left in voice.go is the handler: reactiveHandler, HandlePushToTalk,
handleText, runTurn, applyAction, replySystem, chatHistory, reply.

Move-only: all 133 distinct non-blank lines removed from voice.go were
matched in the two destination files, zero lines added to voice.go. The only
non-move edits are import lists (log added to patterns.go, unicode and
internal/pattern dropped from voice.go) and two comments that pointed at
voice.go for code that is no longer there.
2026-08-01 01:30:06 +04:00
kami 9190f897a3 Add a locked-down maven.<domain> block to the nginx template (#354)
The template's wildcard `listen 80` with no ACL was fixed in 50cc17f, but it
still only covered nexus/praxis/hexis. mavweb — the one service in the set
that serves an RCE surface (POST /tools defines argv internal/tool executes)
— had no block at all, so anyone wiring it up wrote their own, which is how
the wildcard got there the first time.

Adds a maven.kvmx.ru server with the same wg+LAN bind and allow/deny,
proxying 127.0.0.1:9201, with the WebSocket upgrade /ws needs, a 32m body
limit for push-to-talk PCM, and a 300s read timeout because an LLM turn on
the iGPU is slow.

Also records in deploy/ecosystem/docker-compose.yml that the sibling
`build:` paths pin nothing and ship the sibling working tree, with the
command to check what is about to be deployed. The stale public DNS records
(item 2) are outside the repo.

Verified: nginx -t on the template inside a minimal http{} accepts it.
2026-08-01 01:27:11 +04:00
kami d29e7ba813 Gate POST /api/chat on the same step-up as /tools (#317)
/api/chat reaches the router, the LLM and, through applyAction, the whole
act path, so it is the widest state-changing surface mavweb serves. It was
the only one with no gate. It now goes through stepUpOK like POST /tools,
POST /routines and POST /api/revert: unchanged in the default deploy
(WebAuthn unconfigured, fail-open behind wg+nginx), 403 under
-require-stepup or an unasserted passkey session.

The route table now carries an explicit enumeration of every state-changing
route and its gate, and the two startup SECURITY log lines name /routines
and /api/chat alongside /tools and /api/revert.

The loopback -addr default the task also asked for landed earlier in
d12de58; the compose already publishes mavweb on 127.0.0.1 only.
2026-08-01 01:24:42 +04:00
kami f7e1187823 Match quiet-mode toggles on whole words, and resolve OFF first 2026-08-01 01:02:10 +04:00
kami ed48c59ba7 Merge branch 'refactor/query-sources' into integration/small-batch 2026-08-01 00:54:11 +04:00
kami b09967f9e6 Split actions.go into per-intent files
Pure move: actionFact, actionReminder, actionAct and actionNote each get
their own actions_<intent>.go. The two small ones (chat, system) and the
actionHandlers table stay in actions.go, which is now just the dispatch
layer and the notes about what does not belong in it. No behaviour
change — only the file a handler is read in.
2026-08-01 00:53:37 +04:00
kami b4a3867479 Turn actionQuery into a chain of query sources
The six answer sources were hand-unrolled inside one 127-line function.
The intent table is a closed set of 7, but this list is open-ended —
Kiwix (#286), RSS (#258), the crawler (#259) and email (#246) each add
one. Each is now a registry entry: a name plus a method on the handler,
walked in order until one claims the question.

Order is unchanged and still load-bearing (memory before the notes-only
pass, #373), the confidence gate keeps its position and semantics, and
every reply string, log line and best-effort failure is verbatim.
2026-08-01 00:51:44 +04:00
kami 88d07b5175 Unify the voice and text turn pipelines into runTurn
HandlePushToTalk and handleText hand-wrote the same eight-step turn
sequence twice, comments in the latter saying "same as HandlePushToTalk"
four times. Extract it into runTurn(ctx, text) string: the voice path
wraps it in stt/tts, the text path returns it directly.

The two had drifted. The text path was missing the quiet-hours toggle
check entirely, so "тихий режим" over IPC/telegram fell through to the
classifier; unifying gives it the check. It also logged the route result
and applyAction return where the voice path did not — both logs are kept
for both paths.
2026-08-01 00:48:50 +04:00
kami c00e3003bf Merge branch 'refactor/praxis-capability-registry' into integration/small-batch 2026-08-01 00:43:54 +04:00
kami c0f9834528 Turn the Praxis act dispatch into a capability registry 2026-08-01 00:43:16 +04:00
kami ad5eb2d1cf Walk a chain of confirm resolvers instead of three copied blocks 2026-08-01 00:42:23 +04:00
kami 5253123d99 Merge branch 'refactor/voice-wiring' into integration/small-batch
# Conflicts:
#	cmd/mavend/voice.go
2026-07-31 23:54:11 +04:00
kami 2abf98dea6 Move the voice daemon wiring and startup out of voice.go 2026-07-31 23:52:51 +04:00
kami f5c71b87f3 Move the confirm/park gate out of voice.go 2026-07-31 23:51:33 +04:00
kami 5934110fa8 Move the Praxis/Hexis act handling out of voice.go
voice.go is still the biggest file in cmd/mavend and most of what is left
has nothing to do with the audio path. The ecosystem integration is one
such lump: it talks to Nexus, Praxis and Hexis over HTTP and only touches
the handler for its store and clock. Lifting it into ecosystem_acts.go
puts it next to ecosystem.go, where the clients it drives already live.

Move-only: handlePraxisAct, recordPraxisTrace (called from nowhere else),
handleHexisAct and execHexis verbatim, plus the two imports that became
unused in voice.go.
2026-07-31 23:43:47 +04:00
kami 6e47a3d736 Merge branch 'worktree-agent-a88193d1d84b04a5b' into integration/small-batch 2026-07-31 23:38:00 +04:00
kami 67a5eb3805 Split applyAction's 300-line switch into a per-intent handler table
applyAction (cmd/mavend/voice.go) dispatched all 7 intents from one giant
switch. Extract each case body verbatim into its own actionXxx method in
new cmd/mavend/actions.go, dispatched from an actionHandlers table keyed by
router.Intent. applyAction itself is now just the dec.Clarify guard plus a
table lookup.

No behaviour change: same reply strings, same side-effect order, comments
moved verbatim. The destructive-act confirm gate and the enabled-tool
allowlist stay entirely inside actionAct, exactly where they lived in the
old switch's IntentAct case — they're act-specific, not cross-cutting, so
they don't move to a separate layer. dec.Clarify short-circuit, dialogue
bookkeeping and detectPattern stay outside the table since they run
regardless of intent.

voice.go: 1638 -> 1344 lines. New actions.go: 362 lines.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:37:33 +04:00
kami 029449eefa Table-drive the IPC dispatcher instead of a 42-arm switch
dispatch() replaces the hand-written switch with a package-level
map[Method]handlerFunc built once at init. Each entry is one
withParams/withParamsVoid/withoutParams call closing only over the
CoreAPI method it invokes — adding a method is now one table line
instead of a new arm.

Check still runs once at the top before any unmarshal, unchanged. The
three non-CoreAPI methods (assert_stepup, store_encryption_key, unlock)
are special-cased before the table lookup since they drive Server
fields (StepUp/WrapKeyFn/UnlockFn), not store state. The current
CoreAPI is loaded once per dispatch and passed into the handler as an
argument, so SetAPI's runtime swap (the unlock transition) still takes
effect on the next request — the table itself never captures an api
value. No wire-format change; existing round-trip and unknown-method
tests in ipc_test.go pass unmodified.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:36:46 +04:00
kami ac36216f5d Merge branch 'worktree-agent-a517419e93e6a219f' into integration/small-batch 2026-07-31 23:31:24 +04:00
kami db17cfcc65 Delete the dead lockedAPI, add UnimplementedCoreAPI for the doubles 2026-07-31 23:31:24 +04:00
kami 7d676eb941 Stop tracking the mavwaked build artifact 2026-07-31 23:30:52 +04:00
kami fe3a4e9514 Merge branch 'worktree-agent-a9e5cef90b263a5e5' into integration/small-batch 2026-07-31 23:27:36 +04:00
kami a906f2afad Extract the pure RU/string/weather helpers out of voice.go 2026-07-31 23:27:08 +04:00
kami a2031a31d1 Record the measured confidence-gate numbers 2026-07-31 23:23:37 +04:00
kami 9b8bdf73cc Merge the five small-task branches 2026-07-31 23:11:26 +04:00
kami 7ad3c9a408 Merge branch 'worktree-agent-af88d63f65f30896b' into integration/small-batch 2026-07-31 23:11:16 +04:00
kami 84e1478823 Merge branch 'worktree-agent-af0fd9507d3e2ee46' into integration/small-batch 2026-07-31 23:11:16 +04:00
kami 9c8d0baffe Merge branch 'worktree-agent-a4cef2a815e32ebbf' into integration/small-batch 2026-07-31 23:11:16 +04:00
kami b0f5a16ec9 Add digest as a real outcome: suppressed care nudges get resurfaced, not lost
Vikunja #281. The interruption policy promised four outcomes — deliver_now,
queue, digest, drop — but only three existed: a care candidate the restraint
gate suppressed for quiet hours / away / calendar-busy simply vanished in
loop.Tick's `continue`, with only the trace remembering why.

internal/morning turned out not to be the natural drain: it's a fixed
Item/FactKey checklist engine, not a generic message bundler, so gate-
suppressed nudge text has nowhere to plug into its evidence model. Built a
parallel (but small, reusing the outbox's shape) durable digest instead:

- internal/store: digest_entries table + EnqueueDigestEntry (dedupes by
  rule+body, mirroring the delivery outbox's bodyHash), PendingDigestEntries,
  ExpireStaleDigestEntries, DrainDigestEntries (mark, never delete — an
  audit trail of what she actually said).
- internal/loop: DigestEligible(severity, blockedBy) is the pure boundary —
  only genuine restraint blocks (quiet_hours/calendar_busy/presence) even
  qualify (cooldown/snooze are not "suppression"); within care, Sev2 (break)
  digests, Sev1 (water/meal — stale by the time anyone could resurface them)
  drops. High severity never digests; alarms bypass the gate and deliver
  unchanged, on purpose.
- cmd/mavend/tick.go: each tick scans ExplainTick's trace for eligible
  blocked candidates, enqueues them, sweeps stale entries (24h expiry — the
  care rules are daily-cadence, so anything older is describing a day
  that's over), and drains the bundle only once the suppression reason has
  actually cleared, capped at 3 spoken items plus a trailing count so a
  digest can't turn into the exact nagging it was built to avoid.

Tests: store-level round-trip/restart-survival/dedupe/expiry/drain, loop-
level severity-boundary unit tests, and tick-level integration tests for
the drain-only-when-clear and never-digest-high-severity behavior.
2026-07-31 23:09:22 +04:00
kami 67563ed1f6 Run pattern detection from the digestion tick, not just voice (#43)
detectPattern only ever fired as a side effect of a voice fact-write, so a
recurring pattern already sitting in history went unnoticed until he
happened to mention it again by voice — the opposite of proactive.

Split the pipeline: extraction (fact -> normalized event) stays where a fact
is written, in voice.go, since it's tied to that write regardless of who's
talking. Detection (events -> stable pattern -> proposed_routines row) moves
into shared code (patterns.go's detectAndPropose) that both the voice path
and the new tick.go:detectPatterns call. The tick runs it every cycle over
every action+object pair on record (store.DistinctEventPairs, added), so a
pattern gets noticed on the daemon's own schedule.

Idempotence and the dismiss-must-stick requirement turned out to already be
handled by the store, not something the tick needs to reinvent:
proposed_routines has UNIQUE(action, object) and CreateProposedRoutine does
ON CONFLICT DO NOTHING, and DismissProposedRoutine flips status in place
without deleting the row. So a pair already proposed, accepted, OR
dismissed is a silent no-op on every later tick — a dismissed pattern can
never resurface, and re-running the scan never spams the /routines page.
Kept the voice-path call (immediate spoken confirmation is a nice feature
UX-wise and is now redundant-but-harmless with the tick, since both paths
share the same guarded detectAndPropose).

Tick-side detection only ever writes a row; it does not notify, ring, or
speak, keeping Maven "not a nag, not autonomous" — the /routines page is
still the only place a proposal becomes visible, and only accepting it
starts producing nudges (fireAcceptedRoutines).

Also fixed the stale vikunja#46 reference in proposed_routines.go — the
TODO it named is what this commit does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:07:33 +04:00
kami f0f7ebc9b2 Give LLM-routed decisions a real confidence so clarify can fire (#359)
Confidence was hardcoded to 1.0 for every LLM decision, and the LLM branch
in Router.Route returned straight from fillSlots without ever touching the
stage-3 threshold gate — so the LLM path could not produce a Clarify no
matter what confidence a model reported. That is why all 6 want_clarify
cases in the 77-case RU fixture were missed by every model in the bake-off.

Fix reads structural signal instead of changing the (parity-locked) router
prompt: a single-token utterance ("вода", "бэкап") is flagged thin evidence
in llmrouter.go; a fact left keyless or an act that never resolves to an
allowlisted fn, checked after fillSlots so the deterministic parsers get
first crack, is flagged in router.go's new gateLLMDecision. Anything below
config.DefaultRouterThreshold (0.55) now sets Clarify=true through the same
path the classifier already uses.

Added unit tests with a stubbed Completer proving both directions: thin
cases clarify, clean multi-word/resolved-slot cases stay confident. The
77-case fixture re-run against a live llama-server is still needed to
confirm the 6/6 moves — not done here, no llama-server on this box.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:07:32 +04:00
kami d12de589a2 mavweb: default -addr to loopback, not all interfaces
PR #47 added two state-changing routes (POST /chat, POST /routines)
behind the -addr flag, which defaulted to ":9200" (all interfaces).
Default now binds 127.0.0.1:9200; anyone who wants LAN/wider exposure
still passes an explicit bind (as deploy/docker-compose.yml already
does with "-addr :9201" inside the container, unaffected by this
default change).

Vikunja #317.
2026-07-31 23:03:45 +04:00
kami 50cc17f33a Lock down deploy/ecosystem/nginx.conf template to match the live host
The template said "drop into your nginx sites" but listened on the
wildcard `listen 80;` with no allow/deny ACL, unlike the actual deployed
hexis.kvmx.ru config which binds only to the WireGuard (10.42.0.1) and
LAN (192.168.1.104) addresses with allow/deny all. Anyone following the
template as written would expose these unauthenticated admin UIs to the
open internet.

Bind explicitly to those two addresses and add the matching ACL block,
mirroring cmd/mavweb/nginx.conf which already does this correctly.
Added a comment naming both addresses as host-specific so a deploy on a
different box swaps the IPs instead of reverting to `listen 80` when the
bind fails.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 23:01:58 +04:00
kami 73d13f1ea6 Merge pull request 'Stop the docs claiming the LLM router is off' (#49) from docs/fix-drift into master 2026-07-31 20:46:57 +02:00
kami 0ca5748699 Stop the docs claiming the LLM router is off
CLAUDE.md's routing section said "llmrouter is wired nil" and called the
classifier cascade the committed default. That stopped being true when the
integration merge landed: voice.go:214 wires pickLLMRouter, DefaultLLMRouter is
on, and deploy/mavend.json sets llm_router true. It is the first thing anyone
reads before touching the router, so it was pointing the next reader at a
wiring job that is already done.

Rewritten to say the LLM router is the default, the classifier is the failure
floor and must not be deleted, and what the two actually measure — 36.8% at
p50 31ms against 67.5%/72.7% at p50 ~2.7s, a trade accepted on purpose. Names
the one thing still open on that path: Confidence is hardcoded 1.0 in
llmrouter.go, so the LLM never asks for clarification (#359).

Also in CLAUDE.md: the persona line pointed at a memory file that does not
exist, so the actual rule was nowhere in the repo. Written out instead —
feminine self-reference, informal singular address, pet names forbidden but his
name allowed — plus the three eval checks that enforce it.

MODEL-BAKEOFF: three claims had gone stale within hours of being written. There
IS a make eval-models target now; the routing numbers ARE the production path,
not a bench artifact waiting on a wiring change; and the truncated 293 MB gguf
is deleted. Struck through rather than removed, since the caveats are part of
how the evening read at the time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 22:37:15 +04:00
kami b43bb265b5 Say it the way she would actually say it (#48) 2026-07-31 20:17:27 +02:00
kami b9a24334ea Say it the way she'd actually say it
Wording fixes from the review of the clarify + phrasing PRs.

- "На когда напомнить?" → "Когда?". After she has just been asked something,
  the long form is the phrasing of a form field, not of a person.
- A reminder now wants a subject as well as a time. "напомни в 11" had a time
  and nothing to say at 11, and she asked nothing at all — she now asks
  "О чём напомнить?". Subject first, since a reminder with no subject is not
  worth setting.
- The expiry notice is five phrasings picked at random instead of one fixed
  sentence. It is the line he hears every time he walks off mid-request, so it
  is the line that repeats most.
- The nudge prompt's ban on "обращения" is now "ласковые обращения". It was
  meant to forbid "милый"/"дорогой", not his name — "Ками, ноутбук на трёх
  процентах" is how she talks, and the eval's cringe check already only flags
  pet names.
- The nudge example no longer claims she plugged the laptop in. She has no
  hands and no smart plug; an example where she acts teaches the model to
  invent actions Maven never took.
- replySystem: "тепло" → "спокойно и без официальных формулировок". A one-word
  mood instruction a 1.7B can't act on, replaced with the behaviour meant.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 21:52:12 +04:00
kami c97aebf55a Merge tonight's work: all 46 reviewed PRs as one verified branch
135 commits. make build produces all 8 binaries; make test exits 0 across 38 packages with no failures, no data races, gofmt and vet clean.

See PR #47 for what had to be fixed to make it build as a unit.
2026-07-31 19:41:52 +02:00
kami 891136c65d gofmt the kiwix client and rewrite test
PR #41 and #44 landed these two files unformatted, so the gofmt gate that
PR #12 added to `make test` failed as soon as both were on one branch.
Struct-tag and comment alignment only, no semantic change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 21:34:03 +04:00
kami 41c7c13f42 Merge remote-tracking branch 'origin/overnight/snooze-works' into integration/jul31
# Conflicts:
#	internal/store/migrations.go
2026-07-31 21:32:18 +04:00
kami a324e8f624 Merge remote-tracking branch 'origin/overnight/eval-writeup' into integration/jul31 2026-07-31 21:31:36 +04:00
kami 51805e7f35 Merge remote-tracking branch 'origin/overnight/kiwix-rewrite' into integration/jul31 2026-07-31 21:31:36 +04:00
kami 533f0acda8 Lead the bake-off with the answer, not the superseded one
The file ran two sweeps and the second one changed the resident model, but
the lede still opened with "Recommendation: keep Qwen3.5-0.8B". Anyone
landing on the file read the wrong conclusion and had to scroll 100 lines
to find that it had been replaced — and it contradicted CLAUDE.md, which
already says the resident model is Qwen3-1.7B.

Both sweeps are accurate, so nothing is rewritten. The lede now states the
outcome and the first sweep's verdict is scoped to what it actually tested:
it rejects LFM2.5-1.2B, which still holds. It never was a case for keeping
0.8B as the resident model.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 21:17:07 +04:00
kami 4f59ba78c6 Write down the five-model sweep and why the 1.7B won
Numbers behind the resident-model change, plus the answer to "could a 230-350M
model do this instead" — no, and the reason is worth keeping: LFM2.5's published
instruction-following scores beat Qwen3.5-0.8B, and every one of those benchmarks
except Multi-IF is English. In Russian the 350M invents non-words and the 230M
answers in Spanish.

Also fills the row TALK-EVAL-31-07-2026.md had to void for contamination, and
corrects a wrong call I nearly made: the 1.7B's 16s p95 looked like the reasoning
trace, but the 0.8B sits at 17s in every run and the 1.7B beat it twice out of
three. The long tail is shared and is not the Thinking block.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 19:08:37 +04:00
kami d0afd9d4f6 Make Qwen3-1.7B the resident model
Stock Qwen3-1.7B, not the CPT'd one — that training is still running. It won
on both fixtures we have, measured tonight on an otherwise idle box:

  routing, 77 RU cases, intent-only:  67.5%  vs  59.7%  for Qwen3.5-0.8B
  talk fixture, 27 cases:             20/27  vs  11-17/27

It also beat Qwen3.5-2B, which is 20% larger, on every routing column.

Two other things came with it:

n_ctx goes 2048 -> 4096. This is a Thinking variant, so reasoning tokens need
the room, and 4096 is the context every score above was measured at. Shipping
2048 would ship something nobody measured.

The doc now says not to bother with sub-500M models, because I checked and they
are not close. LFM2.5-350M routes at 5.2% — worse than guessing among 7 intents
— and answers "столица Франции?" with "Сторзит", which is not a word. The 230M
replies to Russian in Spanish. Their published IFEval and BFCL numbers are good
and they are all English.

Note the routing gain needs the LLM router actually wired on to show up. It is
still nil, so this commit buys the phrasing improvement today and the routing
improvement when that lands.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:58:20 +04:00
kami 6b67e6f3c2 Word nudges from templates by default, model optional
DEPRECATION, flagged not asked: LLM-phrased nudges are no longer the default.
LLMPhraser.PhraseNudge now returns a hand-written Russian template. The model
still phrases chat, queries and reminders — only nudges moved.

Why: measured over many runs, Qwen3.5-0.8B wrote formal "вы" and plural
imperatives, used masculine self-reference, and invented facts and units
(90-95 seconds to boil an egg). A nudge is five words of known content, so
generation buys nothing and risks the persona every time. Templates score
15/15 on the nudge fixture, the model 11-13/15.

Nothing is deleted: the prompt, the fallbacks and the whole LLM nudge path
stay. Set phraser.llm_nudges = true in deploy/mavend.json to get them back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:21:52 +04:00
kami 742b2ad1d7 Score the Russian-to-keywords rewrite end to end (#403)
Same 9 cases as the retrieval eval, so the numbers compare directly:
hand-written keywords hit 8 of 8, this is what the model reaches on its
own. Reports the hand-written query next to the model's for every case,
because where the phrasing differs is the useful part.

Opt-in on MAVEN_KIWIX_URL + MAVEN_LLM_URL, like the other evals.

Result on Qwen3.5-0.8B: 3 of 8, identical on all three runs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:19:56 +04:00
kami b300ac5c70 Rewrite a Russian question into English Kiwix keywords (#403)
Kiwix ranks by keyword, not meaning, so a translated question finds song
and TV titles. This asks the resident model for the TOPIC instead: a short
English noun phrase, like a Wikipedia article title.

Locked down three ways, because a wrong query is silently wrong:
- A GBNF grammar, same idea as routeGrammar and responseGrammar. The
  reply must be {"query":"..."} with Latin words only. The JSON wrapper
  matters: this model always thinks out loud and this llama-server build
  ignores the thinking switch, so a bare word-list grammar just captured
  "Let me analyze this request carefully" for every question.
- max_tokens 32, since the answer is a few words.
- CleanQuery, which throws away empty, Russian and prose replies rather
  than passing them to Kiwix, and drops question words like "why" and
  "how much" that a keyword ranker cannot use anyway.

Client side only. Nothing is wired into the daemon or any config.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:19:42 +04:00
kami 13e5170e9e Hand-written Russian nudge templates plus a picker
Nudge wording as data instead of generation. The wording lives in
internal/phraser/nudges_ru_v1.json (embedded), about 10 variants per rule:
water, meal, break, service_down, netdata_critical, routine:, morning:, plus
a contentless default. That JSON is long because it is data — the owner can
edit any line of Russian without touching Go.

The picker:
- random, but never the same variant twice in a row for the same rule
- deterministic when seeded (math/rand with an injectable source)
- fills {since} / {service} / {what} from the candidate, and skips any variant
  whose value is missing, so no raw placeholder can reach the piper voice
- {since} is spelled out in words ("полтора часа", "семь часов"), because
  "3 ч" is wrong in a Russian voice

Scores 15/15 on the existing nudge fixture, on every seed swept. Nothing is
wired yet — that is the next commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:18:55 +04:00
kami 0b90952e55 Write down every conversational eval score from tonight
Records all four configurations on the 27-case talk fixture, three runs
each: no grammar, plus grammar, plus Russian prompts, plus the truncation
fix. Composite, per-path and per-check, with the reproduce command.

The short version is that the plumbing got fixed and the score barely
moved. Grammar was the real win. Russian prompts helped a little and cut
latency by 5x. The truncation fix was necessary and bought nothing.

Also writes down three things that are easy to lose:

- The truncation cause was the grammar's 400-character bound, not the
  token cap. Measured at three caps, same 400 characters every time.
- Then I set the bound to 1000 against a 768-token cap and made it worse.
  The two limits have to agree.
- One run is contaminated and marked void: I ran an agent against the same
  llama-server, and the report still claimed zero errors while a third of
  the fixture silently answered "не знаю.". That is #397 and it is worse
  than filed — a busy server is indistinguishable from bad phrasing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:18:19 +04:00
kami aa8f5b2ee2 Make the nonempty check look for actual words
It scored 27/27 on a run where two replies were "{" and "{\n  \"". It only
tested that the string was not blank, so punctuation counted as content and
the worst replies of the run passed the first check.

Now a reply needs at least one letter, Cyrillic or Latin. Latin counts
because answers about ssd or vpn are legitimately part English.

Digits alone fail too. The same run answered "сколько варить яйцо
вкрутую?" with "15-16" — no unit, no words, and the wrong number as well.
That is not something she said.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:57:39 +04:00
kami d7cdcb63bd Stop shipping half-written JSON as a reply
Two bugs, one symptom. A run of the talk eval produced replies that were
literally "{" and "{\n  \"" — those strings went out as things Maven said.

First bug: the parser could not tell "the model answered in plain prose"
from "the model started a JSON object and got cut off". Both came back as
empty, and every caller then shipped the raw text. Now an unfinished object
returns an error and each caller uses its own fallback instead. Bare prose
with no JSON in it still passes through, because small models do sometimes
answer that way and the reply is fine.

Second bug, and the actual cause: the grammar capped the response field at
400 characters. I measured it against Qwen3.5-0.8B at three different token
caps — 256, 768 and 2048 — and the reply came back exactly 400 characters
every time, cut mid-word. So the token limit was never what stopped it.
The bound is 1000 now, about six Russian sentences, still low enough to cut
off a repetition loop.

Token caps go from 256 to 768 on the chat and query paths so 1000
characters of Russian actually fits. The nudge path keeps its own cap; a
nudge is meant to be one sentence.

Note: cmd/mavend/replier_llm.go has its own copy of this parser with the
same bug. Left alone here so this commit stays small — that duplicate is
Vikunja #396.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:56:55 +04:00
kami ddb658ffbb Add a Kiwix search client and score retrieval (Vikunja #403)
Step one of letting Maven read instead of recall. No LLM yet.

internal/kiwix/client.go: search a local Kiwix server, parse the RSS
reply, hand back title + path + plain-text snippet + word count. The
snippet is the unit of context; a full article is ~100KB of HTML and
will not fit a 4096 token window.

internal/kiwix/retrieval_eval.go plus knowledge_v1.json: the 9 knowledge
questions from the phrasing fixture, each with hand-written English
keywords, scored on whether a wanted article comes back in the top 5.
Opt-in via MAVEN_KIWIX_URL, since CI has no Kiwix. No pass bar, the
number is the finding.

Result on the live mirror: 8/8 answerable questions hit, 7 of them at
rank 1. Retrieval works. Keywords are written by hand on purpose, since
Kiwix ranks by keyword and not by meaning, so a natural question fails.
A query-rewrite step is the next piece of work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:43:44 +04:00
kami c7dadc97d9 Write the chat and notes prompts in Russian
The reply has to be Russian, but two of the phrasing prompts told her
what to do in English. Both are Russian now, in the same style as the
nudge prompt that already works better.

Also dropped the "you are maven, a self-hosted personal assistant"
line from both. The persona block right above it already says who she
is, so it was said twice.

The JSON part is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:38:36 +04:00
kami c9d88c152e Drop "never phones home" as a hard rule
The owner's call, 2026-07-31: a 0.8B model does not know enough about the
world to be useful without reading something. So she may now read external
sources to answer world questions.

What replaces the old rule, in all three docs:

- No telemetry, no cloud model, no third-party account. Unchanged.
- Local first: the Kiwix ZIMs on the box before anything on the network.
- External search is allowed but off unless configured, same as weather
  and telegram.
- His notes and facts are never search input. Only the utterance goes out
  — never the persona block, the history, or matched notes.

Docs only, no code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:34:53 +04:00
kami 1890ff5d5d Constrain the phrasing output with a GBNF grammar
The 0.8B answered about one chat turn in three with open reasoning as plain text, so no JSON ever closed and the fallback shipped "Thinking Process:" to the user. A grammar makes that output impossible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:18:06 +04:00
kami 0110e9bc8c Report every address break, and stop -те verbs blinding the check
From a real reply in a nudge eval run: "Смотрите на его потребление
воды" is a plural imperative AND third person about him. Only the plural
printed.

Two separate faults. The check returned on its first hit, so the second
break stayed invisible and the failure read as milder than it was; it now
joins them. And "его" was not detected at all — looksVerb knows the
-й/-йте imperative but not the -те plural, so "смотрите" counted as the
person being talked about, which is what an antecedent means here.
pluralVerb already knows that form, so the antecedent test uses it too.

Third time a verb form has blinded this check. A fourth means it wants a
morphology table rather than another suffix.
2026-07-31 16:52:16 +04:00
kami 50ca8c8b5a Score the chat, query and knowledge phrasing paths (#395)
The phrasing fixture was 15 nudge cases, so every prompt change we
measured only told us about nudges. But the shared context block sits in
front of five prompts, and three of them — chat, note query, general
knowledge — had no scorer at all. Those are the long free-form replies,
where a persona break is most likely and where nothing could see one.

27 cases, nine per path. Nine rather than five because the nudge fixture
already cannot resolve a change smaller than about three cases, and a
per-path score off five would be worse.

Reuses the persona checks instead of copying them. Length, mood and
"no questions" are left out on purpose: these paths return no mood, and
a follow-up question is a feature in chat, not a fault.

The run refuses to score unless the model answers before and after it.
PhraseChat and PhraseQuery swallow model errors and return a canned
string, so without that guard a dead server produces a full report with
zero errors and a bad score — which reads as bad phrasing rather than as
nothing measured. Vikunja #397 is the real fix.
2026-07-31 16:51:52 +04:00
kami de09471421 Merge the shared prompt context block 2026-07-31 16:07:35 +04:00
kami d65c16a567 Don't tell her she can't talk
The block listed what she can do and ended with "nothing else". It sits
in front of the chat and general-knowledge prompts too, so that told her
to refuse the exact thing those prompts are for. Talking is now first in
the list, and the closing line limits ACTIONS rather than everything.

Also dropped the self-introduction from the knowledge prompt. It said
"Мавена, персональный ассистент" — a different name and a masculine
noun, right after the block says she is Maven and feminine. Identity
lives in the block now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 16:07:35 +04:00
kami 062d4252ef Tell her what she can actually do
The context block now lists her real capabilities: reminders, notes and
facts (write and recall), and the calendar — all three are code paths in
mavend today. Weather, telegram and shell acts are listed only when the
config actually has them, because offering something she cannot do is
worse than staying quiet about it.

Also drops the pronouns from the optional name/city line. The block's
own "ты" is Maven, so "тебя зовут" read as her name and "его" would have
shown her the third-person form she must never use about him. They are
plain labels now.

Vikunja #394.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 15:57:58 +04:00
kami 2c27e2ce1f Give every prompt one shared context block
The "address him as ты" rule had only reached two of the five system
prompts. Instead of pasting it into the other three (five copies drift —
that is how this happened), there is now one block, in internal/persona,
prepended to all five: nudges, action replies, chat, note queries and
general knowledge.

The block says who he is and how to address him (a man, always "ты",
never "вы", never "он" about him; Maven stays feminine), plus the
current local date and time. It is rendered fresh each turn because the
time changes, and it is correct with an empty config — the address and
gender rules are defaults in code. Config only adds optional facts:
owner_name, city, and the existing free-text `persona` string, which is
now the static half of the block.

Russian even in front of the English prompts: the rules are Russian
grammar, so they read best stated in Russian, and there is one copy.

Vikunja #394.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 15:55:30 +04:00
kami ccc5cba2a3 Merge the example-led prompt finding 2026-07-31 15:41:01 +04:00
kami 89d83c0b11 Record the example-led nudge prompt experiment (#393) — it made things worse
Tried rewriting the nudge prompt to lead with five on-topic examples instead
of rules. Three eval runs each side: before 12/13/14 of 15, after 11/12/11.
The loss is all in the address check — formal "вы" and plural imperatives came
back once the "говоришь на ты" rule stopped being its own sentence, and the
on-topic examples leaked their wording into the wrong cases.

Prompt reverted. Only the finding is committed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 15:40:18 +04:00
kami a97f554802 Merge the informal address prompt rule 2026-07-31 14:54:20 +04:00
kami f4de2fc5e1 Don't let a verb count as the person being talked about
The third-person check asks whether anyone else was named before "он".
A nudge is mostly verbs, and they were counted as possible people, so
"попробуй встать и отдохнуть — у него есть перерыв" passed. Infinitives
and imperatives now join past tense as words that cannot be a person.

A plain noun before the pronoun still blinds it. That needs a parser,
and the comment says so.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:54:20 +04:00
kami eef5d4da4f Tell the phraser to speak to him informally, singular
The prompts stated the feminine self-reference rule but never said whom she is
speaking to, so the model produced formal plural ("Жду вас") and talked about
him in third person ("Он не ел 11 дней"). Adds the address rule right next to
the feminine one, in the nudge prompt and the confirmation prompt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:52:23 +04:00
kami 09f1696fce Merge the eval label and kill script fixes 2026-07-31 14:32:45 +04:00
kami 80f7322294 Don't fail when docker confirms nothing is running
"Nothing on the host" meant two different things and the script treated
them the same. If docker answers and names no running containers, Maven
really is down and the script should say so and exit 0. Only when docker
cannot be asked is the answer unknown, and that is the case that must
fail loudly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:32:45 +04:00
kami fa5aebfbe4 Merge the delivery boundary fixes 2026-07-31 14:30:54 +04:00
kami 59cec63da1 List the columns in the table rebuild
The migration copied rows with SELECT *, which matches columns by
position. It is correct today, but if the old table's order ever
differed it would shuffle every row instead of failing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:30:54 +04:00
kami 02e8786695 Stop the containers instead of claiming success (#380)
docker-compose.yml has no 'pid: host', so each container has its own PID
namespace and pkill on the host matches nothing inside them. The script
then printed "All services gracefully stopped" while mavend, its
llama-server and the rest were still running.

Now it checks for running compose containers first and stops them with
docker compose. If it cannot ask docker and finds nothing to kill on the
host, or anything survives the kill, it says so and exits non-zero
instead of claiming success. The bare-metal path is unchanged apart from
verifying the SIGKILL actually worked, and no longer risks killing the
shell it was launched from.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:30:51 +04:00
kami 0272dc9d89 Record a suppressed care nudge instead of dropping it silently (#370)
Dropping a sev1-2 care nudge while you're away is right and still happens.
But it was a bare `continue`: no row, no log, so "she dropped it", "the gate
suppressed it" and "the rule never fired" all looked identical afterwards.

Adds a 'dropped' delivery status (migration #12 widens the CHECK constraint;
sqlite can't do that in place, so the table is rebuilt) and records the drop
as one delivery_attempts row plus a log line.

No nudges row for a drop: that table feeds the ignored_rate signal, and a
nudge nobody could see must not count as ignored.

TestVoiceNoSessionFallthroughLeavesOutboxTrail expected exactly one row for
sev1-2 when voice had no session. It now expects the voice failure plus the
drop, which is the point of the change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:27:47 +04:00
kami 2ad7635501 Merge the address-form eval check 2026-07-31 14:27:16 +04:00
kami 9949b309b1 Don't let a time word blind the third-person check
The check asks whether anyone else was named before "он". Time words
were not stoplisted, so "сегодня он не ел" read "сегодня" as the person
being talked about and passed — which is the recorded break with a word
in front of it, and nudges open with those words constantly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:27:08 +04:00
kami a788ca3915 Label eval runs with the model the server actually loaded (#379)
The phrasing eval printed "llm (0.8B, ...)" no matter which gguf
llama-server had loaded, so two runs of two different models came out
named the same and were easy to mix up when comparing.

It now asks llama-server over /v1/models, same as the router eval
already did. The helper moved to internal/llm so both share it, and it
now errors instead of returning a blank name when the id field is
missing — an unreachable server gets labelled "unknown-model", never a
plausible-looking guess.

Both eval paths stay opt-in behind MAVEN_LLM_URL; no server needed for
go test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:25:28 +04:00
kami 62d47d28ac Add an eval check for formal and third-person address (#384)
The phrasing run produced two persona breaks that scored clean:
"Приходите… Жду вас" (formal plural) and "Он не ел 11 дней" (talks
about him instead of to him). She is feminine, he is male, and she
speaks to him informally, one to one.

The new `address` check flags the "вы" family, plural imperative
endings, and a third-person "он" with no other subject named earlier in
the message. Like `hisgender` it is a keyword/suffix heuristic, not a
parser, and it prints the word it tripped on so a false alarm is easy to
dismiss. Limits are written out in the comment.

Both recorded strings are pinned as unit tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:25:23 +04:00
kami e9ff2c4912 Never send the full nudge body off-box (#368)
The away sinks fell back to the whole Body when Summary was empty. ntfy and
telegram leave the box, and the 0.8B phraser drops fields regularly, so that
fallback could push full detail off the machine.

The dispatcher already strips detail from away sendables. This exports that
one rule as delivery.AwayMessage and has both sinks use it, so a sink can't
leak the body on its own either: empty Summary means a generic line plus the
rule name, never the body.

The two sink tests named TestSendFallsBackToBodyWhenSummaryEmpty asserted the
old, wrong behaviour, so they are rewritten to assert the generic line.
TestSendRejectsEmptyMessage is likewise replaced: an away message can no
longer be empty, so the sink has nothing left to reject.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:23:54 +04:00
kami 3dbf67f8f9 Drop the city time-zone table
The user only ever asks the time in his own zone, so answering other
cities was code kept in step with the weather city list for no gain.
Any named place now gets the honest "local time only" answer that was
already there for unknown cities.

Removes the 22-entry table, the lookup and the embedded tz database.
Closes Vikunja #389 — there is only one city list again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:14:18 +04:00
kami 84ba217892 Say so when the day asked about is out of reach 2026-07-31 14:05:21 +04:00
kami f179ae2fde Merge the system reply fixes 2026-07-31 14:03:42 +04:00
kami d00929ac0b Answer the day the user asked about and the city he named (#388)
replySystem had two arms that PR 30 made reachable, and both answered confidently wrong: the date arm keyword-matched "числ" and always answered today, so "какое число завтра" answered today; the clock arm ignored a named city and answered local time. The date arm now reads the day word through router.ParseCalendarDate (which grew послезавтра/вчера and now cuts the day boundary in the local zone instead of UTC). The clock arm answers the named zone when it resolves offline from the tz database embedded in the binary, and otherwise says plainly that she only knows local time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 14:02:57 +04:00
kami b6f47fbeb6 Merge the clock and calendar routing rule 2026-07-31 13:53:35 +04:00
kami 2e9b9ec1cf Warn separately when a row has no text to re-embed 2026-07-31 13:52:56 +04:00
kami bfb57c3148 Give the router prompt a rule for clock and calendar questions (#374)
The prompt named seven intents but never said which one a clock or date
question belongs to, so the model guessed: system->query x4 in every eval
run. The rule now says the clock and the calendar date themselves are
system, what is written in the calendar or in memory stays query, and a
time named inside a request is just part of the request.

That split follows what the daemon can answer. Only replySystem owns the
clock and the date formatter, while the agenda is answered from
CalendarEvents inside the query branch.

Also adds one calendar-agenda fixture case so an over-broad system rule
cannot pass unnoticed, and writes up the before/after numbers. The
targeted confusion is gone; the headline accuracy did not move.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:52:00 +04:00
kami d1f6f6355f Merge the vector backfill 2026-07-31 13:51:25 +04:00
kami 92ecb691de Re-embed stored notes and facts after an embedder swap (#378)
The embedder swap left every stored vector in the old model's space, so cosine against a new query vector is noise. Add the one-shot backfill: store.ReembedAll re-embeds every note and fact text with the currently configured embedder (the passage side, which is the side stored text was written with) and rewrites both places a vector lives — the notes table embedding column and the memory_vectors rows.

All of it plus the embedder marker happens in one transaction, so a failure partway changes nothing and writes no marker: re-run it. A run against a DB whose marker already names the current embedder does nothing.

Triggered explicitly with `mavend -reembed`, not automatically on mismatch: ONNX on the laptop CPU makes this minutes of work, and a silent multi-minute stall on boot would look like a hang. The mismatch warning now tells the user to run it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:50:43 +04:00
kami 4282f6b9a9 Warn about vectors written before the marker existed 2026-07-31 13:44:29 +04:00
kami 7bb9f9be06 Merge the embedder marker 2026-07-31 13:42:40 +04:00
kami 1e47eaca5a Record which embedder wrote the stored vectors and warn on a swap (#378)
The embedder moved from paraphrase-multilingual-MiniLM-L12-v2 to
multilingual-e5-small. Both are 384-dimensional, so nothing in the code
noticed: cosine between an old stored vector and a new query vector is
noise, and recall degrades silently.

So the DB now records the embedder that wrote its vectors. One value for
the whole DB (migration #11, a small `meta` key/value table) rather than a
column on every vector row: the backfill re-embeds every note and fact in
one pass, so a per-row marker would hold the same string in every row and
cost a column on two tables for nothing.

The identity comes from the embedder itself via a new optional ID() method
("multilingual-e5-small@384", model file name plus dimension), so pointing
the config at another model changes the string without anyone editing a
constant. mavend logs a loud WARNING at startup naming both the stored and
the configured embedder when they differ.

Detection only — recall behaviour is unchanged. TODO(#378) in
store.CheckEmbedder marks where the backfill will hook in.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:42:07 +04:00
kami 892330eb84 Merge the note recall fix 2026-07-31 13:33:14 +04:00
kami 9a3bcd7c46 Merge the thinking-off measurement 2026-07-31 13:31:31 +04:00
kami 98ee701e03 Let a note win a recall, not only a fact (#373)
The memory pass ran only after the notes-only gate had already rejected
the same note at the same score. Notes and facts share one vector index,
so a note that failed there failed again — the branch could only ever
return a fact.

Now the memory pass runs first: one search over everything Maven
remembers, one gate, and the memory that clearly matches best answers
(a note gets phrased, a fact is read back). The notes-only pass stays
behind it for notes the vector index does not hold. No threshold moved,
so the set of questions answered is unchanged — only which memory
answers them.

Fixture gained two mixed note+fact cases, so the answerable count goes
25 -> 27: hash recall@1 36.0% -> 37.0% (ratchet 0.32 unchanged, comment
updated), e5 recall@1 72.0% -> 70.4%, false recall still 1/5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:30:38 +04:00
kami 04c1088088 Measure thinking off on routing properly — it does not win (#376)
The 67.1% "thinking off" column in ROUTING-EVAL-31-07-2026.md was an
artefact. It came from a hand-rolled HTTP client in the eval test that
did not send repeat_penalty, so it differed from the reference run on two
axes and the penalty was the one that mattered.

Re-scored back to back on an idle box with everything else held equal:
thinking off is identical to thinking on, case for case, same confusion
matrix, same three unparseable replies. A direct probe of the running
llama-server shows enable_thinking, thinking and reasoning_budget are all
ignored for this model on this build, so there was nothing to turn off.

No defaults changed. The misleading third configuration is removed from
internal/router/eval so its table cannot be quoted again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:30:22 +04:00
kami 07c191d8b8 Merge dialogue session persistence 2026-07-31 13:21:03 +04:00
kami c668310b3e Persist the dialogue session so a restart keeps the conversation
Vikunja #363. The follow-up session was a plain in-memory map, so any
mavend restart dropped the thread. It now mirrors to a small TTL-pruned
sqlite table and is loaded on startup; expired sessions are deleted on
load, not revived. Clarify's pending question is untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 13:20:30 +04:00
kami 1bd2acdc2a Do not exempt Russian words that are both noun and verb 2026-07-31 12:55:45 +04:00
kami 15e5dd8eaa Merge the second-person gender check 2026-07-31 12:54:48 +04:00
kami 10cf6f525c Check that nudges do not address the owner in the feminine
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:54:18 +04:00
kami e2210f6844 Merge the clarify-expiry notice 2026-07-31 12:52:36 +04:00
kami 214a4032cf Tell him when an expired clarify question is dropped
Vikunja #382. A parked clarifying question past its TTL was discarded
silently on read; now she says the old request is gone and the newly
spoken words are still routed as a fresh utterance.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:51:47 +04:00
kami dc70a5a7ab Show clarify_max_attempts in the deployed config
The default is 3 either way. Writing it out means you can see the knob
without reading the Go.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:34:28 +04:00
kami 06aded6ab0 Merge commit 'd2be98e' into overnight-jul31
# Conflicts:
#	cmd/mavend/clarify.go
#	cmd/mavend/clarify_test.go
#	cmd/mavend/voice.go
#	internal/config/config.go
2026-07-31 12:34:00 +04:00
kami d2be98ee2a Say out loud when she gives up instead of dropping the request
An unclear answer used to end the request on the spot. Now she re-asks the same
question while attempts remain, and when they run out she says
"Прости, я не поняла. Скажи, пожалуйста, по-другому." — silence would leave him
thinking it was handled. Same reply when the missing slot has no question to
ask, and as a floor in finishClarified so an empty reply can never ship.

Tests: three questions allowed, the fourth gives up out loud, the cap is
configurable, and a restated time is the one that lands.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:31:55 +04:00
kami 62d320f93a Let her ask three times, and let a restated answer win
MaxAttempts was 1, justified as "not a nag". Wrong reading: "not a nag" is about
interrupting unprompted, and a clarifying question is part of a conversation he
started. Now three, configurable via voice.clarify_max_attempts (default 3).
Three, because after that the likely problem is she misheard the whole request,
not one slot.

Answer used to keep the parked value, so "в три" then "нет, в пять" threw the
five away. Now a value the answer carries wins for the slot she asked about.
Only for the clarify answer — a correction in a fresh turn is followUpMerge.

The eight-field chained assertion in the Answer test is one DeepEqual now, so a
new field in Slots is covered without touching the test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:31:46 +04:00
kami 796e6af3cf Merge commit '74a7088' into overnight-jul31 2026-07-31 12:24:20 +04:00
kami 74a70880a8 Write up the phrasing eval: 0/15 to 13/15, and what the number hides
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:23:54 +04:00
kami a40bc559d5 Fix the nudge phrasing prompt: stop teaching the model to echo the example
The system prompt showed the JSON contract as {"response": "..."} and the
user prompt repeated it. A 0.8B copies whatever sits in the response slot, so
7 of 15 nudges came back as literally "...".

Changes, all prompt-side — the {"response","mood"} contract is unchanged:
- nudge system prompt is Russian, feminine self-reference, with filled-in
  examples on topics that never appear as rules, so copying them is visible
- rule names get a Russian gloss and a required keyword, named last in the
  prompt where a small model weights it hardest
- durations render in Russian, not English
- the no-parse fallback says something Russian instead of "water — care",
  which was going straight to a Russian piper voice
- same "..." placeholder removed from replier_llm.go

Scored on internal/phraser/eval: 0/15 -> 13/15.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:23:03 +04:00
kami 1db0fcfcd0 Merge commit '4ca68d2' into overnight-jul31 2026-07-31 12:07:38 +04:00
kami 4ca68d2f3f Bake off LFM2.5 against Qwen3.5-0.8B on the RU routing fixture
Vikunja #278 / #250. Keep Qwen: LFM2.5-1.2B loses 8 points of intent
accuracy, all of it Russian, and runs 2.4x slower.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 12:07:06 +04:00
kami 5a1d465db5 Merge commit '4383844' into overnight-jul31
# Conflicts:
#	deploy/mavend.json
#	internal/config/config.go
2026-07-31 11:52:21 +04:00
kami 43838445ab Write up the margin gate results
Third section: why the absolute gate could not separate the two
distributions, the delta sweep, and the before/after. Marks next-steps
item 3 done.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:51:21 +04:00
kami 11831c6ace Gate recall on the margin over the runner-up, not just the score
The e5 embedder puts every cosine in one narrow band (0.79-0.89), so the
absolute query_min_score gate cannot tell a real hit from a made-up
question: any value under the band answers everything, any value above it
answers nothing. False recall was 5/5.

New gate asks whether one note is clearly the best instead: top1 - top2 >
delta. New query_min_margin config knob, default 0.008, read off the sweep
in the recall harness. The absolute floor stays as a second check.

On the recall fixture with e5: answered 72% -> 68%, false recall 5/5 -> 1/5.

Vikunja #359

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:50:39 +04:00
kami 93c1a41d4a Route with the resident model by default
The two things that made this unsafe are fixed: the router can now
refuse, and slot extraction runs on its decisions.

On the held-out fixture it gets 63.2% of intents right against the
classifier's 50.0%, with no route errors. It costs about a second a
turn instead of 30ms.

The flag is a pointer now, so leaving it out of the config means on
and only writing false turns it off.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:47:57 +04:00
kami bce5ed210c Merge branch 'worktree-agent-a76ce40c73601d90d' into overnight-jul31 2026-07-31 11:44:32 +04:00
kami c31f0d1001 Extract slots for LLM router decisions too
An LLM-routed reminder came back with no parsed time and an act with no
fn, because only the classifier path ran the extractor. Now the router
runs the same extraction after an LLM decision and fills only the empty
slots. No time in the utterance still means no time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:44:06 +04:00
kami 34521c30b8 Merge branch 'worktree-agent-ad5da57e47b822152' into overnight-jul31 2026-07-31 11:39:35 +04:00
kami 1d48755d12 Record the recall numbers after the embedder swap
recall@1 60% to 72%, answered 48% to 72%, latency 3x better. But false
recall went 1/5 to 5/5: e5 packs every score into a narrow high band, so
the 0.55 gate now admits everything. Left the gate alone as instructed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:38:18 +04:00
kami f6d5a2a7a4 Swap the embedder to multilingual-e5-small (Vikunja #371, #372)
The old model was a symmetric paraphrase model, so it scored "do these
look alike" instead of "does this note answer this question". Also fixes
the file mismatch: the Makefile, the deploy config and both evals now all
name the same quantized file, and the quantized one is what gets measured.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:38:18 +04:00
kami 751c2a705f Embed a question and a stored note differently (Vikunja #371)
Note recall is asymmetric: a short question goes in, a longer note comes
out. Adds EmbedQuery/EmbedPassage helpers and the e5 prefixes, and points
the note/fact write path at the passage side and the query path at the
query side. Reviewers: the three call sites in voice.go.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:38:07 +04:00
kami 5d5b0cfd49 Merge branch 'worktree-agent-a0ab2a9b7439296e3' into overnight-jul31 2026-07-31 11:37:53 +04:00
kami bd16ca69e5 Let the LLM router answer "unknown" when it cannot route
Chose an 8th enum value over a confidence number: the model already picks
one enum token, so it costs nothing in the grammar, while a score from a
0.8B model would be uncalibrated noise. A refusal returns "no decision"
with no error, which is the fall-through the caller already uses for a
bad parse, so the classifier and its clarify gate take the turn.

Reviewers: the prompt's counter-examples matter most — a small model will
over-use any easy escape hatch. The training workspace copy of the prompt
still needs the same edit (Vikunja #362).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:35:37 +04:00
kami 0ed386eca6 Re-measure the router on a quiet box and record the numbers
The earlier before/after was taken while another eval shared
llama-server. This run had the box to itself.

Intent accuracy 61.8% llm-only, 63.2% cascade, 67.1% with thinking off.
The prompt fix holds. note→fact shows up here too, so it is real.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:34:57 +04:00
kami 1c4eab2107 Merge commit '94eb92f' into overnight-jul31
# Conflicts:
#	Makefile
2026-07-31 10:11:56 +04:00
kami 0914e0a3d5 Merge commit '4ba9a6f' into overnight-jul31
# Conflicts:
#	Makefile
2026-07-31 10:11:29 +04:00
kami c860808528 Make make test actually gate on gofmt and vet
DESIGN.md has always said `make test` is "gofmt + vet + -race, no
exceptions". It only ever ran the tests, which is how nine files drifted
out of format without anyone noticing.

`test` now depends on `fmt-check` and `vet`. Checked that fmt-check does
fail when a file is unformatted, so the gate is real and not decorative.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 10:10:27 +04:00
kami f7442c3aea Run gofmt over the seven files that had drifted
Formatting only: import order, and statements that were packed onto one
line split out. `git diff -w` shows nothing but that.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 10:09:32 +04:00
kami 75b067ac51 Merge the accepted-routine fix and drop reminder_id from accept
Two merge fixes on top of the branch:

- migrations: keep both new steps, snooze stays #8, the routine columns
  become #9. Both agents had numbered theirs #8.
- accepting no longer takes a reminder id, on the web surface too. The
  web accept path had the same one-shot-reminder bug the voice path did,
  so both now just flip the status and let the tick loop schedule.

The test that asserted "accept creates a reminder and links it" asserted
the bug. It now asserts that accepting creates no reminder.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 10:07:37 +04:00
kami 424d1b3446 Fire accepted routines every interval, not once (Vikunja #366)
The tick loop now reads accepted routines from the store and nudges when
their interval has passed; accepting no longer builds a one-shot reminder.
Look at routine.DueAccepted for the schedule rule (no catch-up backlog) and
at fireAcceptedRoutines for the restraint gate — routines do not bypass it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:45:30 +04:00
kami c47886c2bc Merge branch 'worktree-agent-ab5b5c61a32cac4fe' into overnight-jul31 2026-07-31 02:42:53 +04:00
kami f6236da760 Collapse the duplicate away-detail and panic tests
Two agents wrote the same three test helpers and names for the same two
bugs. Kept the real assertions from dispatcher_test.go and removed the
skipped placeholders they replace. panicSink stays in durability_test.go
since both files use it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:42:21 +04:00
kami 54dc43516b Add accepted-routine timestamps to the store (Vikunja #366)
Data layer only. Migration #8 adds accepted_ts and last_fired_ts to
proposed_routines, plus ListAcceptedRoutines and MarkRoutineFired so the
tick loop can own the schedule. Accepting no longer links a reminder id.
Look at the TODO(vikunja#366) in cmd/mavend/tick.go for the next commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:41:58 +04:00
kami fa51a48958 Merge branch 'worktree-agent-afe3f2ec18b2b8497' into overnight-jul31 2026-07-31 02:40:10 +04:00
kami 5fd25d7ad7 Test the away-channel minimal body and the panicking sink (#368, #369)
The integration branch names one test TestAwayFallsBackToFullBodyWhenSummaryEmpty,
which describes the old bug; it is here as TestAwaySendsGenericLineWhenSummaryEmpty
and asserts the generic line instead of the body. Also covers: a normal summary
goes out unchanged, voice keeps the full body, and one panicking sink does not
eat the other channel for the same nudge. Reformatted one pre-existing struct.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:39:38 +04:00
kami c22fc352fc Merge branch 'worktree-agent-a1610b8c5376eadd6' into overnight-jul31 2026-07-31 02:37:30 +04:00
kami 215aa331c5 Recover from a panicking sink so the attempt is always closed (#369)
A panic in Send used to unwind past completeOutbox and leave the
delivery_attempts row pending forever, since reconciliation only runs at
startup. safeSend turns the panic into an error, logs it loudly, records the
attempt failed, and lets the other channels for the same nudge still go out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:35:54 +04:00
kami 859bbf750f Never send a nudge body off-box when the summary is empty (#368)
Away channels (ntfy, telegram) leave the box, so an empty Summary now sends
a fixed generic line plus the rule name instead of the full Body. The
dispatcher strips detail before any sink sees it, so a sink added later
cannot leak by reading the wrong field. Voice is local and unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:34:54 +04:00
kami 8a174c1c70 Score the recall fixture and write up what it shows
Real recall is 48% after the gate, and one must-be-silent query gets an
answer anyway. Review finding 2 (the score distributions overlap, so no
gate separates a real recall from a false one) and finding 4 (the memStore
branch at voice.go:776 is unreachable for notes). Adds an embedder cache
so the gate sweep does not re-embed the fixture nine times.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:34:28 +04:00
kami ee3e6a9eaf Wire the snooze read into the Gatherer and honour it for reminders (#364)
The Gatherer now fills State.SnoozeUntil from store.SnoozedUntil instead
of nil, so a snooze finally reaches the gate. RemindDecisions gains the
one restraint check that applies to a reminder — quiet hours, presence
and cooldown are still bypassed, so "wake me 7" is unchanged. Reviewer:
the two tests in internal/loop/gate_test.go are the contract.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:33:51 +04:00
kami 8acb8a97c6 Read the recorded snooze outcomes back out of the nudges table (#364)
The gate honours State.SnoozeUntil but nothing ever filled it. New
store.SnoozedUntil returns, per rule, when the newest snooze runs out.
Reviewer: the fixed 2h SnoozeDuration and its reasoning in nudges.go —
nothing upstream can supply a per-nudge length, so no new column.
Expired snoozes are dropped in SQL, so silence can never be permanent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:33:42 +04:00
kami 1f55207b58 Merge the snooze read path and wiring
# Conflicts:
#	internal/loop/gate_test.go
2026-07-31 02:33:41 +04:00
kami 4db109346a Ignore the .claude directory
Agent worktrees land in .claude/worktrees, so the directory shows up as
untracked noise in every git status.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:32:44 +04:00
kami 2f00593411 Wire the snooze read into the Gatherer and honour it for reminders (#364)
The Gatherer now fills State.SnoozeUntil from store.SnoozedUntil instead
of nil, so a snooze finally reaches the gate. RemindDecisions gains the
one restraint check that applies to a reminder — quiet hours, presence
and cooldown are still bypassed, so "wake me 7" is unchanged. Reviewer:
the two tests in internal/loop/gate_test.go are the contract.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:32:24 +04:00
kami 0b3b8d0a9e Test the clarify round-trip end to end at the daemon level
Covers: a reminder with no time is asked about and completes on the answer; the
same for a fact; an answer past the TTL falls through as a fresh utterance; a
second unclear answer drops the request with no second question; a clarified act
off the allowlist neither runs nor gets enabled; a clarified destructive act
still parks a confirm; noise keeps the canned reply. No model, no network.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:31:07 +04:00
kami fe0e654ab1 Ask the question, then act on the answer
On a clarify decision with one identifiable gap she now asks instead of saying
"не поняла", and parks the request. The next utterance is parsed as the answer
with the router's own extractor and the completed decision runs through
applyAction like any other — so a clarified act still needs the allowlist and
still hits the destructive confirm gate. An answer that does not fill the gap
drops the request; she never asks twice. Also pulls the session-store block
that HandlePushToTalk and handleText both had into rememberTurn, since the
clarify path needed a third copy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:31:07 +04:00
kami a54ebac0cb Work out which slot is missing and phrase one short question
A table per intent (reminder needs a time, fact needs a key, act needs a fn)
plus one fixed Russian question per slot. Templates, not model output: a 0.8B
would wander and a question that rewords itself is harder to answer. Note,
query, chat and system get no question — for those a clarify decision keeps
the canned reply rather than inventing a question for noise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:31:07 +04:00
kami 784d688b44 Merge the clarify wiring
# Conflicts:
#	internal/dialogue/clarify.go
#	internal/dialogue/clarify_test.go
2026-07-31 02:30:59 +04:00
kami 4ba9a6f422 Add a deterministic scorer for nudge phrasing (Vikunja #323)
Review internal/phraser/eval/checks.go -- it IS the measurement. Each check
names in a comment which DESIGN.md line it defends: length, feminine
self-reference (windowed around "я" so the operator's own masculine
second-person forms are not flagged), the cringe list (pet names, emoji,
"!!", fake concern, apology, emotional support, asking how he feels,
praise), on-topic, mood enum. No send/veto signal anywhere, per
DESIGN.md § "Rules decide, LLM phrases".
Fixture (158 lines) and tests (252) do not count toward the diff ceiling;
the scorer itself is still ~650. Splitting eval.go from checks.go would
give two commits neither of which measures anything.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:30:52 +04:00
kami 32687b3712 Read the recorded snooze outcomes back out of the nudges table (#364)
The gate honours State.SnoozeUntil but nothing ever filled it. New
store.SnoozedUntil returns, per rule, when the newest snooze runs out.
Reviewer: the fixed 2h SnoozeDuration and its reasoning in nudges.go —
nothing upstream can supply a per-nudge length, so no new column.
Expired snoozes are dropped in SQL, so silence can never be permanent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:29:51 +04:00
kami db3e706bdc Merge the delivery routing and durability tests 2026-07-31 02:29:05 +04:00
kami 2f4257e194 Test the clarify round-trip end to end at the daemon level
Covers: a reminder with no time is asked about and completes on the answer; the
same for a fact; an answer past the TTL falls through as a fresh utterance; a
second unclear answer drops the request with no second question; a clarified act
off the allowlist neither runs nor gets enabled; a clarified destructive act
still parks a confirm; noise keeps the canned reply. No model, no network.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:29:04 +04:00
kami 7d8b0af99d Test voice fallthrough and delivery durability
Fallthrough is checked per severity through the outbox trail, so sev3/sev4
reroute and sev1/sev2 still drop. Durability uses a real store on a temp file:
a crash between Begin and Complete becomes unknown, is not resent, is not
dropped, and a late Complete cannot overwrite it. One skipped test marks a real
gap: a panic mid-send leaves a permanent pending row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:27:33 +04:00
kami cfd38d53cf Ask the question, then act on the answer
On a clarify decision with one identifiable gap she now asks instead of saying
"не поняла", and parks the request. The next utterance is parsed as the answer
with the router's own extractor and the completed decision runs through
applyAction like any other — so a clarified act still needs the allowlist and
still hits the destructive confirm gate. An answer that does not fill the gap
drops the request; she never asks twice. Also pulls the session-store block
that HandlePushToTalk and handleText both had into rememberTurn, since the
clarify path needed a third copy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:26:11 +04:00
kami a2835bbdf6 Cover every cell of the delivery routing table
Table-driven tests for all four severity bands crossed with present and away,
both as the pure table and end to end through the dispatcher. Three tests are
written to DESIGN.md and skipped because the code does not keep the claim: the
care-away drop is recorded nowhere, and the minimal body is enforced per-sink
rather than by the dispatcher. Also gofmt'd dispatcher_test.go.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:24:43 +04:00
kami c262c1e4c0 Merge the routine accept path and store gaps 2026-07-31 02:24:07 +04:00
kami a33ad82178 Let the /routines page accept a proposal, gated at step-up (#46)
Accepting a routine gives the trigger loop a new standing reason to speak to
the human, so it is the same authority tier as enabling a tool and shares the
stepUpOK gate; dismiss only ever makes maven quieter, so it is ungated.
Look at handleRoutines and acceptRoutine in cmd/mavweb/main.go: accept creates
the recurring reminder, then links it via the new ipc AcceptProposedRoutine.
The page now says what maven noticed in her own words (pattern.PhraseRoutine).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:23:28 +04:00
kami af35ec3629 Work out which slot is missing and phrase one short question
A table per intent (reminder needs a time, fact needs a key, act needs a fn)
plus one fixed Russian question per slot. Templates, not model output: a 0.8B
would wander and a question that rewords itself is harder to answer. Note,
query, chat and system get no question — for those a clarify decision keeps
the canned reply rather than inventing a question for noise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:22:18 +04:00
kami 9145b83100 Add the clarify data layer: a parked question with one missing slot
The router can already say "I am not sure" (Decision.Clarify) but the daemon
had nowhere to keep the request while it asked. PendingQuestion holds the
original slots, ClarifyStore parks one per dialogue id with a 90s TTL, and
Answer fills only the slots that were missing so an answer can never rewrite
what she already understood. Logic that uses this comes next.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:20:28 +04:00
kami 43470abc57 Add a held-out note-recall harness (fixture + scorer)
Measures whether Maven can find the right note again from a paraphrased
question. Review internal/memory/recalleval/recalleval.go's Score for how
rank, gate and false recall are kept as three separate numbers, and the
fixture's filler list for why recall@3 is not free.
Fixture JSON is generated data and does not count toward the diff limit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:18:58 +04:00
kami 94eb92fb15 Label LLM eval runs with the model llama-server has loaded
The bake-off in #278/#250 needs two models' scores side by side, and the
report names only carried the config, so the rows were indistinguishable.
ModelID reads /v1/models instead of taking a string that goes stale.
New target: make eval-models MAVEN_LLM_URL=...

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:17:20 +04:00
kami 707c3e5040 Ignore deps and models as symlinks, not just directories
.gitignore had deps/ and /models/llm/ with trailing slashes. A trailing slash
only matches a real directory, so a *symlink* with the same name is not ignored
and git add -A commits it as a symlink blob.

That bites anyone working in a git worktree, where deps/ and models/ do not
exist and have to be linked in from the main checkout. It already happened once
tonight.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:17:15 +04:00
kami ace9fbbc06 Merge the llm_router flag and the kill-maven fix 2026-07-31 02:17:00 +04:00
kami 3884db33e9 Give proposed routines a status filter and pin down the dedup rule (#46)
Look at internal/store/proposed_routines.go: status flips in place with an
`AND status = 'proposed'` guard, not append-only like facts/voids_id — a
proposal is a question with one answer, same shape as tools.status. The
UNIQUE(action, object) key is what stops a dismissed routine coming back.
New tests cover re-propose-after-dismiss and listing by status.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:16:39 +04:00
kami dfb8d26b62 Fix kill-maven.sh so it actually kills llama-server
The MODEL default was LFM2, but the deploy runs Qwen3.5-0.8B, so the
pkill pattern matched nothing and the server survived every kill.
Now matches any llama-server serving a .gguf, so changing the model in
deploy/mavend.json cannot break the script again. MODEL still narrows it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:15:30 +04:00
kami bf99fd4192 Add a voice.llm_router flag, default off
Wires cmd/mavend/voice.go to build the LLM router when the operator asks
for it. Default false, so nothing changes on the deploy box.
Look at pickLLMRouter: the flag on with no llama-server logs one line and
keeps the classifier, it never fails a turn.
The default stays off until the router can refuse (#359) and the extractor
runs on LLM decisions — both noted as TODOs in config.go.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:15:30 +04:00
kami 6d9aa83b6e Merge the proactive rule and gate tests 2026-07-31 02:15:15 +04:00
kami f75072175d Merge the pending-question data layer 2026-07-31 02:14:33 +04:00
kami 39d83a33e8 Add the pending-question data layer for slot clarification
PendingQuestion plus ClarifyStore: same shape, locking and expiry as SessionStore. Answer fills only the missing slots and never overwrites a filled one. No wiring yet — TODOs mark the daemon hooks.
Reviewer: MaxAttempts is 1 on purpose (Maven asks once, she is not a nag).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:13:48 +04:00
kami 925ce223a0 Add a Value slot to dialogue.Slots
router.Slots already carries the fact payload; the dialogue copy did not, so a clarifying answer had nowhere to put it. InheritSlots carries it like Key.
Reviewer: check the new inherit block does not overwrite a filled value.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:13:40 +04:00
kami d30618ecb7 Stop the router repetition loop
Route now sets RepeatPenalty on the request, and the grammar's string rule is
capped at 120 characters. Two of 76 fixture cases looped one sentence inside
the text field until MaxTokens, which cut the JSON in half.
Reviewers: the new constant and the grammar string rule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:11:38 +04:00
kami 17b47ce206 Route questions to query, not fact
The router prompt tested "reports current state -> fact" before "wants
information -> query", so a question naming a fact key was written as a fact.
Query now comes first, plus an explicit question test.
Reviewers: the prompt block in llmrouter.go, and the note about the
training-side copy of the prompt that needs the same edit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:10:37 +04:00
344 changed files with 50230 additions and 3202 deletions
+17 -1
View File
@@ -6,12 +6,18 @@
/mavweb
/mavpoll
/mavcaldav
/mavwaked
/mavmaild
/mavupdate
# Certs (private keys, don't commit)
certs/
# Dependencies (fetch/build, not vendored)
# Dependencies (fetch/build, not vendored).
# Both forms on purpose: 'deps/' misses a symlink named deps, and agents working
# in a git worktree symlink these in from the main checkout.
deps/
deps
# ML models (large, downloaded separately) — specific dirs, not blanket,
# because models/seeds/*.txt are small, tracked files the classifier needs.
@@ -19,6 +25,9 @@ deps/
/models/stt/
/models/tts/
/models/llm/
# Symlink forms, same reason as deps above.
/models/embedder
/models/llm
# Runtime data
*.db
@@ -27,6 +36,10 @@ deps/
deploy/db_key.env
# Deploy secret (telegram bot token + chat id) — never commit
deploy/telegram.env
# zenmoney API token, read by mavpoll (never in argv, never committed)
deploy/zenmoney.token
# IMAP password, read by mavmaild (never in argv, never committed)
deploy/imap.password
# Temp files
/tmp/
@@ -36,3 +49,6 @@ opencode.json
# Test coverage output
coverage.out
# Agent worktrees and local agent state
.claude/
+8 -5
View File
@@ -43,14 +43,17 @@ notes. Without it, the floor `HashEmbedder` is used — deterministic but weak
(Russian recall rarely clears the confidence gate, many commands fall to
"clarify").
**Download the embedder** (ONNX, ~90 MB):
**Download the embedder** (ONNX, ~120 MB):
```sh
make download-embedder
```
This fetches `paraphrase-multilingual-MiniLM-L12-v2` (384-dim, 12-layer,
supports 50+ languages including Russian) to `models/embedder/`.
This fetches `multilingual-e5-small` (384-dim, 12-layer, Russian and English)
to `models/embedder/multilingual-e5-small/`. It is an asymmetric retrieval
model: the code puts `query: ` in front of a question and `passage: ` in front
of a stored note, which is how e5 was trained. The quantized file is the one
that is downloaded, deployed and measured.
**Also need ONNX Runtime** (`libonnxruntime.so`):
@@ -64,8 +67,8 @@ sudo cp onnxruntime-linux-x64-1.15.1/lib/libonnxruntime.so* /usr/local/lib/
```json
"voice": {
"embedder": {
"model_path": "models/embedder/model_quantized.onnx",
"tokenizer_path": "models/embedder/tokenizer.json",
"model_path": "models/embedder/multilingual-e5-small/model_quantized.onnx",
"tokenizer_path": "models/embedder/multilingual-e5-small/tokenizer.json",
"lib_path": "/usr/local/lib/libonnxruntime.so"
}
}
+66 -14
View File
@@ -7,9 +7,20 @@ talking over unix sockets; one resident small model for routing + phrasing; whis
Deploy target is a Ryzen laptop (homesrv) with Vulkan offload to the Vega iGPU (`n_gpu_layers: 99`,
compose passes `/dev/dri` + the render gid) — the resident model stays ≤1.7B either way.
**Resident model:** currently **Qwen3.5-0.8B** (`Q4_K_M`), the smallest checkpoint in the gguf
library, picked for CPU/iGPU latency. The **target** is the locally CPT'd **Qwen3-1.7B**; that
training is still in flight (Vikunja #122), so no such gguf exists yet. Model files live in
**Resident model:** currently **Qwen3-1.7B** (`UD-Q4_K_XL`), stock — not yet the CPT'd one.
It replaced Qwen3.5-0.8B on 2026-07-31 because it measured better on both fixtures we have:
67.5% vs 59.7% intent-only on the 77-case RU routing fixture, and 20/27 vs 11-17/27 on the
talk fixture. See `MODEL-BAKEOFF-31-07-2026.md`. It is a Thinking variant, so `n_ctx` is 4096
— reasoning tokens need the room, and 4096 is what the scores above were measured at.
The **target** is still the locally CPT'd **Qwen3-1.7B** (Vikunja #122, training in flight).
Stock already speaks good Russian; what it gets wrong is the persona — it writes `я рад`,
masculine, where Maven needs `рада`. That is what the CPT is for.
**Do not bother with sub-500M models.** LFM2.5-230M and 350M were measured on 2026-07-31 and
both are unusable in Russian: the 350M routes at 5.2% (worse than guessing) and answers
"столица Франции?" with the invented non-word "Сторзит"; the 230M replies to Russian in
Spanish. Their strong published IFEval/BFCL numbers are English-only. Model files live in
`/mnt/hdd1/llms`, bind-mounted to `/opt/maven/models/llm` — which **shadows** the repo's
`models/llm/`, so the LFM2.5 gguf sitting there is not loaded by anything. Swapping the resident
model is a one-line change to `phraser.model_path` in `deploy/mavend.json`.
@@ -23,7 +34,7 @@ CGO daemons (`mavend`, `mavsttd`, `mavttsd`, `mavenclient`) need the vendored to
and libs wired through the Makefile — **do not** call `go build` on them bare, use `make`:
```sh
make build # all 8 binaries
make build # all 9 binaries
make build-web # single daemon (pure-Go ones: web/waked/poll/caldav build without CGO)
make test # go test -race across ./internal/... ./cmd/... with CGO env set
```
@@ -51,6 +62,7 @@ Pure-Go packages (`router`, `memory`, `mavweb`, …) run under a plain `go test
| `mavenclient` | Voice loop client (mic → stt → core → tts). |
| `mavpoll` | Telegram long-poll reach. |
| `mavcaldav` | CalDAV calendar sync. |
| `mavmaild` | Mail reader (IMAP, read-only). Holds the IMAP password; core never sees it. |
Daemons are wired socket-to-socket, not linked. `internal/ipc` is the client/server wire
protocol; the config in `deploy/mavend.json` (with `${VAR}` env expansion from gitignored
@@ -58,19 +70,41 @@ protocol; the config in `deploy/mavend.json` (with `${VAR}` env expansion from g
## Routing — read this before touching the router
`internal/router/` has TWO layered engines and the committed default is an **interim
stopgap, not the intended design** (see memory `routing-architecture-target`):
`internal/router/` has TWO layered engines. **The LLM router is now the default and it is
on in deploy** — this section used to say it was wired `nil`, which stopped being true on
2026-07-31.
- **Target (REARCH.md):** LLM-as-router. One resident Qwen3-1.7B (`llmrouter.go`) emits
GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is demoted
from a routing gate to a RAG hint.
- **Current stopgap:** `llmrouter` is wired `nil` (around `voice.go`), so the
`classifier.go` + `embedder.go` nearest-neighbour cascade actually runs. It routes by
similarity to frozen seed phrases — the known cause of weak RU query handling.
- **LLM router (the intended design, REARCH.md):** the resident Qwen3-1.7B (`llmrouter.go`)
emits GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is
demoted from a routing gate to a RAG hint. Wired at `voice.go:214` via
`pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient)`; the flag is `voice.llm_router`
(`config.go`), `DefaultLLMRouter` is **on**, and `deploy/mavend.json` sets it `true`.
- **Classifier cascade (the failure floor, not dead code):** `classifier.go` +
`embedder.go` nearest-neighbour over frozen seed phrases. It runs when the LLM router is
off, when there is no llama-server to talk to (`pickLLMRouter` logs that and degrades),
and on any per-turn LLM error. Do not delete it — routing by seed similarity is the known
cause of weak RU query handling, but a turn must never break on the model.
Cascade order: `stage0.go` exact-match fast-path → LLM router (when non-nil) → classifier
fallback. Any LLM error falls through to the classifier so a turn never breaks on the model.
Measured on the 77-case RU fixture (`MODEL-BAKEOFF-31-07-2026.md`): the classifier scores
36.8% full accuracy at p50 31ms; Qwen3-1.7B scores 67.5% intent-only / 72.7% through the
cascade at p50 ≈2.7s. Accuracy roughly doubled, latency is ~90× worse, and that trade was
accepted deliberately. `Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
#359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with
no allowlisted fn) feeding the same stage-3 gate the classifier path already had — see
`gateLLMDecision` in `router.go`. Note the second half of that bug: the LLM branch never
consulted `r.threshold` at all, so a correct low confidence would have been discarded anyway.
Re-measured on the fixture after the fix: **missed clarify 6/6 → 1**, at the cost of 3 false
clarifies and 2.6pt of full accuracy (72.7% → 70.1%, intent-only 67.5% → 74.0%). Two of the
three false clarifies are acts the model mis-routed and the gate caught — asking beats wrongly
executing, so the fixture and the daemon disagree about what is correct there. The third,
`"поужинал"`, is a real defect: **the single-token rule is an English intuition and does not
transfer to Russian**, where one word is routinely a whole sentence. Narrow or drop it.
## LLM output contract
All phrasing paths emit `{"response":"...","mood":"..."}` (parsed in `replier_llm.go` and
@@ -82,8 +116,26 @@ workspace enforces that the Go and relabelling prompts remain identical.
## Non-goals (hard constraints)
Never phones home. Not a nag, not autonomous. Maven's persona is **feminine** — Russian
self-reference must use feminine forms (the user is male; see memory `maven-persona-gender`).
Not a nag, not autonomous. Maven's persona is **feminine** — Russian
self-reference must use feminine forms `рада`, not `рад`; `поняла`, not `понял`. The owner
is male and is addressed informally: "ты", singular, never "вы"/"ваш" and never "он"/"его"
(she talks TO him, not about him). Pet names ("милый", "дорогой") are forbidden; his name
("Ками") is not. The eval enforces this: `CheckAddress`, `CheckFeminine` and `CheckCringe` in
`internal/phraser/eval/checks.go`, scored by `make eval-phrasing`.
**"Never phones home" is DEPRECATED** (owner's call, 2026-07-31). It used to be a hard
constraint and it is not one any more: a 0.8B — and a 1.7B — does not know enough to answer
world questions, so she needs to read external sources. What replaces it:
- **No telemetry, no cloud model, no third-party account.** That part never changes. Nothing
about Maven is reported to anyone, and inference stays on the box.
- **Local sources first.** Kiwix ZIMs on homesrv (Wikipedia, ifixit) before anything on the
network. Reading beats recalling for a small model, and a local read costs nothing.
- **External search is allowed and off unless configured**, like the weather and telegram
capabilities.
- **His notes and facts are never search input.** Looking up why the sky is blue and sending
his stored personal notes to an upstream engine are different acts. Only the utterance goes
out, never the persona block, history, or matched notes.
## Web UI conventions
+10 -4
View File
@@ -15,7 +15,8 @@
**Maven** — self-hosted personal assistant. Manages your day, acts on your
homelab. One daemon on homesrv (always-on, not the workstation), multiple
client surfaces. All local, never phones home.
client surfaces. Inference and data stay on the box; she may READ external
sources (see Non-goals — "never phones home" is deprecated).
Primary name is "Maven", with feminine-gendered Russian self-reference
("она", "меня", "помогла"). Clients may choose their own UI label. Consistent
@@ -35,8 +36,13 @@ Inside boundary — the ones that actually constrain the build:
she records. A confident wrong fact is worse than a known gap.
- **Not a nag** — she'd rather miss a nudge than be mutable. Shuts up when
uncertain. Load-bearing.
- **Not a stranger** — runs on your stuff, your model, your data. Never
phones home.
- **Not a stranger** — runs on your stuff, your model, your data. No
telemetry, no cloud model, no third-party account. She may READ external
sources to answer world questions (Kiwix first, then optional search); she
never reports anything about you to anyone, and your notes and facts are
never used as search input. **"Never phones home" as an absolute is
deprecated** — owner's call, 2026-07-31: a small model does not know enough
to be useful without reading.
- **Not a relationship** — mom-tone is a function that makes nudges land, not
emotional company. Names the drift a warm small model falls into.
@@ -458,7 +464,7 @@ decides *insistence*. Both are needed.
sev ≤ 2 drops on away, sev ≥ 3 holds: a missed water nudge is noise, a missed
backup failure isn't. Away-channels (ntfy/telegram) leave the box — the one
path that crosses "never phones home," through your own relay. **Minimal
path that leaves the box for a person to see, through your own relay. **Minimal
body** — "disk low on homesrv," not detail; don't make notifications a
shoulder-surf exfil surface.
+2 -1
View File
@@ -51,7 +51,8 @@ RUN go build -o /out/mavend ./cmd/mavend && \
go build -o /out/mavttsd ./cmd/mavttsd && \
go build -o /out/mavweb ./cmd/mavweb && \
go build -o /out/mavpoll ./cmd/mavpoll && \
go build -o /out/mavcaldav ./cmd/mavcaldav
go build -o /out/mavcaldav ./cmd/mavcaldav && \
go build -o /out/mavmaild ./cmd/mavmaild
# llama.cpp Vulkan build — the phraser/router LFM engine (llama-server). Built
# from source (not a prebuilt vendored blob) so the binary's glibc/GLIBCXX match
+223
View File
@@ -0,0 +1,223 @@
# Resident model bake-off — 31-07-2026
**Outcome: the resident model is stock Qwen3-1.7B** (`UD-Q4_K_XL`). Two sweeps ran this
evening and the second one changed the answer — read to the end before acting on any table
here. [Second sweep](#second-sweep-same-evening--five-models-and-a-resident-model-change)
is the one that holds.
## First sweep — LFM2.5-1.2B vs Qwen3.5-0.8B
**Verdict, scoped to this pair: keep Qwen3.5-0.8B over LFM2.5-1.2B.** LFM2.5-1.2B is worse
at routing (52.6% vs 60.5% intent accuracy), and the loss is almost entirely Russian
(18/61 vs 22/61 RU, while EN is a wash). It is also 2.4× slower. The Thinking variant is
far worse again. This verdict still stands as written — it rejects LFM2.5-1.2B. It is
**not** a recommendation to keep 0.8B as the resident model; the second sweep replaced it
with Qwen3-1.7B.
Settles Vikunja **#278 / #250**.
- Same fixture and scorer as `ROUTING-EVAL-31-07-2026.md`: `internal/router/eval/`
(`ru_routing_v1.json`, 76 held-out cases).
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:<port> make eval-router`
(`TestLLMRouterBaseline`). (This line used to say there is no `make eval-models` target.
There is one now — start a server with the gguf you want, then
`make eval-models MAVEN_LLM_URL=http://127.0.0.1:<port>`. It runs only the LLM test, since
the classifier baselines do not depend on the model.)
- All three models served by the same `llama-server` flags — `-c 2048 -ngl 99 -t 6`, only
`-m` and `--port` differ. One server at a time on an otherwise idle box, so latencies are
real and not contention.
- Measured on top of the router prompt fix (`origin/overnight/router-prompt` merged in), so
the Qwen column is directly comparable to the numbers already recorded.
## Results
`llm-only` — the model alone. This is the column that measures the model.
| | Qwen3.5-0.8B | LFM2.5-1.2B Instruct | LFM2.5-1.2B Thinking |
|---|---|---|---|
| **intent-only accuracy** | **60.5%** | 52.6% | 36.8% |
| full accuracy (intent+slots+gate) | **36.8%** | 32.9% | 21.1% |
| **RU** | **22/61** | 18/61 | 10/61 |
| EN | 6/15 | **7/15** | 6/15 |
| route errors | 0 | 0 | 0 |
| **p50 / p95 latency** | **1.05s / 1.71s** | 2.47s / 3.62s | 2.42s / 3.24s |
| missed clarify | 6 / 6 | 6 / 6 | 6 / 6 |
`cascade+llm` — stage-0 → model → classifier floor, what #320 would actually ship. Same
ordering.
| | Qwen3.5-0.8B | LFM2.5-1.2B Instruct | LFM2.5-1.2B Thinking |
|---|---|---|---|
| intent-only accuracy | **61.8%** | 55.3% | 38.2% |
| full accuracy | **46.1%** | 42.1% | 30.3% |
| RU / EN | **27/61** / 8/15 | 23/61 / **9/15** | 15/61 / 8/15 |
| route errors | 0 | 0 | 0 |
| p50 / p95 latency | **1.28s / 1.94s** | 2.18s / 2.72s | 2.27s / 3.19s |
Full logs: the three runs are archived in the session scratchpad
(`qwen08.txt`, `lfm-instruct.txt`, `lfm-thinking.txt`).
## Russian-specific failures — the owner's worry is confirmed
LFM2.5's Russian loss is not spread out. It has one large, specific failure: **it hears
almost any Russian imperative or short phrase as `reminder`.**
- `перезапусти докер` → reminder (want act)
- `включи вытяжку` → reminder (want act)
- `закрой жалюзи` → reminder (want act)
- `заметка: продлить домен в августе` → reminder (want note)
- `запиши что кран на кухне снова капает` → reminder (want note)
- `доброе утро` → reminder (want chat)
- `спасибо тебе` → reminder (want note/chat)
- `переходи в тихий режим` → reminder (want system)
That is `note→reminder ×4`, `act→reminder ×4`, `chat→reminder ×2` in one run. Qwen's
equivalent failure axis is `query→fact ×8`, which is a narrower and already-understood bug.
Two more Russian-side problems worth naming:
1. **Fact keys come back empty or wrong in Russian.** `воды попил наконец`, `поужинал`,
`поспал часов пять` and `отметь что я позавтракал овсянкой` all returned an empty key.
`сходил в душ` and `отдохнул минут двадцать` both returned `water`. Qwen does not do this.
2. **It leaked German.** `slept about seven hours` produced the fact key
`"7 Stunden geschlafen"`. Grammar-valid, semantically garbage — a sign the multilingual
mix is not anchored where Maven needs it.
The claimed tool-calling advantage did not show up here. `act` is the closest thing this
fixture has to a tool call, and LFM2.5 got it wrong more often than Qwen, mostly by calling
it a reminder. It also produced no `fn` slot on any act, same as Qwen.
## The Thinking variant
Not viable. 36.8% intent accuracy, 10/61 Russian, and no latency saving over Instruct — the
thinking trace costs time without buying accuracy on a short enum classification. With the
`enable_thinking=false` diagnostic it collapsed further to 28.9% with 2 route errors
(`query→reminder ×12`). Do not pursue.
## Notes
- Nothing crashed, nothing ignored the GBNF grammar, and no model produced unparseable JSON
in the shippable configurations. Zero route errors for both Instruct and Thinking in
`llm-only` and `cascade+llm`. The problem with LFM2.5 is what it decides, not whether it
can emit the contract.
- The `6 / 6` missed clarify is unchanged across all three models. No model fixes the missing
refusal lane — that is `Confidence: 1.0` hardcoded in `llmrouter.go` (Vikunja #359), not a
model property.
- The report labels every configuration `(0.8B)`; that string is hardcoded in the test, not a
reflection of which gguf was loaded. Model identity was confirmed per run via `/v1/models`.
- No Go code was changed for this measurement, and no bug was found that needed one.
## What this does not settle
Routing only. LFM2.5 might still phrase better, and phrasing is the resident model's other
job — that needs its own fixture. But routing is the load-bearing path and Maven is
Russian-first, so on the evidence here the switch is not worth making.
---
# Second sweep, same evening — five models, and a resident-model change
The sections above compared LFM2.5-1.2B against Qwen3.5-0.8B on routing and concluded
"the switch is not worth making". That still holds. This sweep asked a different
question — whether a *smaller* model could work, since LFM2.5's published
instruction-following scores beat Qwen3.5-0.8B badly — and answered it, plus found a
better resident model by accident.
**Outcome: the resident model is now stock Qwen3-1.7B.** Sub-500M is a dead end.
## Routing — 77 Russian cases, one run each
| model | on disk | llm-only (full) | llm-only (intent) | cascade + fallback |
|---|---|---|---|---|
| LFM2.5-230M-Q8_0 | 246 MB | 23.4% | 33.8% | 36.4% |
| LFM2.5-350M-Q8_0 | 379 MB | 2.6% | **5.2%** | 20.8% |
| Qwen3.5-0.8B-Q4_K_M | 527 MB | 36.4% | 59.7% | 61.0% |
| Qwen3.5-2B-UD-Q4_K_XL | 1.34 GB | 42.9% | 62.3% | 63.6% |
| **Qwen3-1.7B-UD-Q4_K_XL (stock)** | 1.13 GB | **44.2%** | **67.5%** | **72.7%** |
Qwen3-1.7B wins every column, including against a model 20% larger than it.
## Talk fixture — 27 cases, three runs each, idle box
| | Qwen3.5-0.8B | Qwen3-1.7B stock |
|---|---|---|
| composite | 13, 11, 8 | **20, 21, 18** |
| address | 21, 18, 18 | **26, 25, 23** |
| feminine | 27, 25, 26 | 26, 27, 26 |
| lang | 27, 27, 26 | 26, 27, 27 |
| ontopic | 16, 19, 19 | **22, 23, 23** |
| canned fallbacks | 8, 5, 6 | **0, 2, 0** |
This also fills the row `TALK-EVAL-31-07-2026.md` had to void for contamination:
**600ch/1024tok on Qwen3.5-0.8B scores 13, 11, 8.**
`address` is the headline. It sat at 18-22 of 27 on the 0.8B no matter how the prompt
was worded — the prompt explicitly forbids "вы" and the model writes `вашей`,
`подождите`, `делаете` anyway. That was read as "prompting is out of levers", and it
was really "0.8B is out of capacity". The 1.7B mostly holds the constraint.
The fallback column matters too: 5-8 of 27 turns on the 0.8B end in a hardcoded
`"не знаю."`, meaning it failed to emit parseable JSON about a quarter of the time.
The 1.7B does that 0-2 times.
## Latency — the long tail is not the Thinking block
| | p50 | p95 |
|---|---|---|
| Qwen3.5-0.8B | 2.4s, 2.9s, 2.0s | 17.4s, 17.6s, 17.4s |
| Qwen3-1.7B stock | 2.7s, 2.6s, 2.8s | 16.4s, 6.6s, 3.9s |
p50 is flat across a 2× size difference. The first instinct on seeing the 1.7B's
16s p95 was "that is the reasoning trace, cap it" — wrong. The 0.8B's p95 is a
consistent 17s and the 1.7B beat it in two of three runs. The tail is shared and
lives somewhere else. Do not spend time on `/no_think` on this evidence.
## Sub-500M: not close, and the benchmarks say otherwise for a reason
LFM2.5-350M publishes IFEval 76.96 against Qwen3.5-0.8B's 59.94, and BFCLv3 44.11
against 35.08 — better at instruction-following and structured output, at 2/3 the
size. Those numbers are real and they are **English**. Every benchmark in that
table except Multi-IF is English-only.
In Russian, with a 300-token budget and temperature 0:
- **350M**, «Столица Франции? Ответь кратко.» → *«Сторзит в Париже.»*`Сторзит` is
not a word; it is invented morphology.
- **350M**, asked to read back a reminder → a fortune cookie about being attentive
and confident. No reminder in it.
- **230M**, «Привет, как дела?» → answered **in Spanish**.
The 230M beating the 350M six-fold on routing (33.8% vs 5.2%) is the other tell:
when the larger sibling collapses like that it is format compliance failing, not
reasoning.
This is a pretraining gap, not a fine-tuning gap. Teaching Russian to a 350M from
near-zero is not an afternoon on a Colab, which was the premise worth checking.
## Why this vindicates the 1.7B CPT
Stock Qwen3-1.7B, untrained and unprompted, answers all three probes in fluent
correct Russian. What it gets wrong is the persona: *«Привет! Я рад, что ты здесь»*
`рад` is masculine and Maven needs `рада`. That is the right kind of remaining
problem, and it is exactly what the CPT (Vikunja #122) is for.
The 1.7B was the correct model choice. What was wrong was treating it as a
**blocker**: stock already beats what was deployed, so it ships now and gets
swapped again when the CPT lands.
## Caveats
- Routing is one run per model, not three. The gaps between families are far larger
than the run-to-run spread seen on the talk fixture, but the 2B-vs-1.7B gap (62.3
vs 67.5) is not safe to call on one run.
- ~~The routing numbers only reach production once the LLM router is wired on. It is
still `nil`.~~ **Resolved the same evening:** the LLM router is wired at `voice.go:214`
behind `voice.llm_router`, the default is on, and `deploy/mavend.json` sets it `true`.
These numbers are the production path now, so the p50 ≈2.7s is a real per-turn cost and
not a bench artifact.
- ~~`/mnt/hdd1/llms/LFM2.5/Qwen3-1.7B-UD-Q4_K_XL.gguf` is a 293 MB truncated download
in the wrong directory.~~ **Deleted 2026-07-31.** The good 1.13 GB copy in `qwen3/` is
what `deploy/mavend.json` loads.
- Harness: `scratchpad/bakeoff.sh`, one server at a time, health-checked before each
run, `/v1/models` recorded per run. Never run two LLM consumers at once — see the
contamination note in `TALK-EVAL-31-07-2026.md`.
+86 -7
View File
@@ -16,11 +16,11 @@ PIPER_BIN := $(shell pwd)/deps/piper/piper
PIPER_MODEL := $(shell pwd)/models/tts/ru_RU-irina-medium.onnx
PIPER_ESPEAK := $(shell pwd)/deps/piper/espeak-ng-data
.PHONY: all build build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav clean test run-stt run-tts run-web download-embedder deps-go eval-router
.PHONY: simulate stt-fixtures test-stt-golden all build build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav clean test fmt-check vet run-stt run-tts run-web download-embedder deps-go eval-router eval-recall eval-phrasing eval-models
all: build
build: build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav
build: build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav build-mail build-update
build-stt:
CGO_CFLAGS="$(CGO_CFLAGS)" CGO_LDFLAGS="$(CGO_LDFLAGS)" LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
@@ -50,6 +50,15 @@ build-poll:
build-caldav:
$(GO) build $(GOFLAGS) -o mavcaldav ./cmd/mavcaldav/
build-mail:
$(GO) build $(GOFLAGS) -o mavmaild ./cmd/mavmaild/
# mavupdate is an operator CLI, not a daemon: nothing runs it but a human on the
# box. It is built with the rest so a broken update path is caught by `make
# build` rather than the first time it is needed.
build-update:
$(GO) build $(GOFLAGS) -o mavupdate ./cmd/mavupdate/
run-web: build-web
./mavweb -addr :9200 -voice 127.0.0.1:9100
@@ -69,7 +78,28 @@ deps-go:
done
$(GO) version
test:
# fmt-check fails if any file needs gofmt. DESIGN.md has always said `make
# test` gates on gofmt and vet; it did not, so nine files quietly drifted.
# Run `gofmt -w` on whatever this prints.
fmt-check:
@bad=$$(gofmt -l internal cmd); \
if [ -n "$$bad" ]; then \
echo "these files need gofmt:"; echo "$$bad"; exit 1; \
fi
vet:
CGO_CFLAGS="$(CGO_CFLAGS)" CGO_LDFLAGS="$(CGO_LDFLAGS)" LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
$(GO) vet ./internal/... ./cmd/...
# simulate — replay every scripted day under cmd/mavend/testdata/scenarios
# through the real router, store, tick loop and intake journal, on a fake clock
# (Vikunja #284). Verbose so the transcript of each scenario lands in the
# terminal. Also runs as part of `make test`; this target is for reading it.
simulate:
CGO_CFLAGS="$(CGO_CFLAGS)" CGO_LDFLAGS="$(CGO_LDFLAGS)" LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
$(GO) test -v -count=1 -run TestSimulator ./cmd/mavend/
test: fmt-check vet
CGO_CFLAGS="$(CGO_CFLAGS)" CGO_LDFLAGS="$(CGO_LDFLAGS)" LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
$(GO) test -race -coverprofile=coverage.out ./internal/... ./cmd/...
@@ -83,6 +113,51 @@ MAVEN_ONNX_LIB ?= $(shell pwd)/deps/onnxruntime-linux-x64-1.26.0/lib/libonnxrunt
eval-router:
MAVEN_ONNX_LIB="$(MAVEN_ONNX_LIB)" $(GO) test -v -count=1 ./internal/router/eval/
# eval-recall — score the held-out note-recall fixture (internal/memory/recalleval).
# Answers "can she find the note again when it matters": recall@1, recall@3,
# false recall and the query_min_score sweep. Same MAVEN_ONNX_LIB deal as
# eval-router; without it only the deterministic hash ratchet runs.
eval-recall:
MAVEN_ONNX_LIB="$(MAVEN_ONNX_LIB)" $(GO) test -v -count=1 ./internal/memory/recalleval/
# eval-phrasing -- score nudge phrasing AND the conversational paths (chat,
# query, general knowledge) in internal/phraser/eval. Verbose so the
# report and every generated message land in the terminal. With no environment
# it scores the deterministic Stub only, which is what CI runs. Set
# MAVEN_LLM_URL to add the resident model:
# MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing
# The model run is slow (minutes) -- the timeout is raised to match. It covers
# two fixtures now (15 nudges + 27 conversational cases, and the chat replies are
# the long ones), hence 90m rather than 40m.
eval-phrasing:
$(GO) test -v -count=1 -timeout 90m ./internal/phraser/eval/
# eval-models — score ONE llama-server against the same fixture, for the
# resident-model bake-off (#278, #250). Start a server with the gguf you want,
# then:
#
# make eval-models MAVEN_LLM_URL=http://127.0.0.1:18100
#
# The report names carry the model llama-server reports, so runs from two
# checkpoints stay apart. Only the LLM test runs — the classifier baselines do
# not depend on the model and take the ONNX runtime with them.
MAVEN_LLM_URL ?= http://127.0.0.1:18099
eval-models:
MAVEN_LLM_URL="$(MAVEN_LLM_URL)" $(GO) test -v -count=1 -timeout 60m \
-run TestLLMRouterBaseline ./internal/router/eval/
# stt-fixtures — regenerate the golden STT audio in cmd/mavsttd/testdata from
# the piper voices (#288). The committed WAVs are synthesised, never recorded,
# so this is the only way they should ever change. TestGoldenAudioTranscription
# then scores them against ggml-small; it self-skips when the model is absent.
stt-fixtures:
./scripts/gen-stt-fixtures.sh
test-stt-golden:
CGO_CFLAGS="$(CGO_CFLAGS)" CGO_LDFLAGS="$(CGO_LDFLAGS)" LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
$(GO) test -v -count=1 -run TestGolden ./cmd/mavsttd/
run-stt: build-stt
LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
./mavsttd -socket /tmp/maven/stt.sock -model $(WHISPER_MODEL)
@@ -108,9 +183,13 @@ deps-piper:
-o /tmp/piper.tar.gz
tar -xzf /tmp/piper.tar.gz -C deps/
EMBEDDER_DIR := $(shell pwd)/models/embedder
EMBEDDER_MODEL_URL := https://huggingface.co/Xenova/paraphrase-multilingual-MiniLM-L12-v2/resolve/main/onnx/model_quantized.onnx
EMBEDDER_TOKENIZER_URL := https://huggingface.co/Xenova/paraphrase-multilingual-MiniLM-L12-v2/resolve/main/tokenizer.json
# multilingual-e5-small: an asymmetric retrieval model. It is trained to match
# a short question against a longer passage, which is what note recall is.
# The quantized file is the one we download, deploy and measure — see
# RECALL-EVAL-31-07-2026.md.
EMBEDDER_DIR := $(shell pwd)/models/embedder/multilingual-e5-small
EMBEDDER_MODEL_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/onnx/model_quantized.onnx
EMBEDDER_TOKENIZER_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/tokenizer.json
download-embedder:
mkdir -p $(EMBEDDER_DIR)
@@ -134,4 +213,4 @@ download-embedder:
@echo ' sudo cp onnxruntime-linux-x64-1.15.1/lib/libonnxruntime.so* /usr/local/lib/'
clean:
rm -f mavend mavenclient mavsttd mavttsd mavweb mavpoll mavcaldav mavwaked
rm -f mavend mavenclient mavsttd mavttsd mavweb mavpoll mavcaldav mavwaked mavmaild
+169
View File
@@ -0,0 +1,169 @@
# Phrasing evaluation — 31-07-2026
How Maven words a nudge, measured instead of argued. Counterpart to
`ROUTING-EVAL-31-07-2026.md`.
- Fixture + scorer: `internal/phraser/eval/` (`nudges_v1.json`, 15 cases; `eval.go`, `checks.go`)
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing`
- Model: Qwen3.5-0.8B Q4_K_M, the resident model. Not swapped.
- Commit: `a40bc55` (prompt fix)
Every check is a string or length test a human can read and disagree with. No model
grades another model here.
## Result
| | before | after |
|---|---|---|
| **cases passing every check** | **0/15** | **13/15** |
| mood in enum | 6/15 | 15/15 |
| Russian | 2/15 | 14/15 |
| length (≤120 chars, ≤16 words) | 13/15 | 15/15 |
| feminine self-reference | 15/15 | 15/15 |
| no cringe | 13/15 | 15/15 |
| on topic | 6/15 | 13/15 |
| p50 latency | 11.4s | 11.4s |
Latency did not move and is not good. 11s to word one nudge on this box.
## The bug reproduced
Yes, exactly as reported. 7 of 15 messages were the literal string `"..."`, and one was
`"full voice message"`. Both are text copied straight out of the prompt.
The system prompt said:
```
Respond ONLY with valid JSON: {"response": "full voice message", "mood": "neutral"}
```
and the user prompt said:
```
Respond as JSON: {"response": "...", "mood": "..."}
```
A 0.8B does not read `"..."` as "put your answer here". It reads it as the answer. The
prompt was a worked example whose worked part was blank, so the model filled the slot by
copying. This is the whole of finding 1.
## What else was wrong
Four separate faults, all prompt-side:
1. **Placeholder echo** (7 cases) — above.
2. **Wrong language** (13/15 failed the language check). The prompt was entirely English
and said "in the user's language (Russian or English)". The model picked English. It is
never English: the nudge is spoken by a Russian piper voice.
3. **Rule names are English identifiers.** `netdata_critical`, `service_down`, `break` went
into the prompt raw. The model cannot nudge about a topic it has not been told in words,
so 9/15 were off topic. The daemon knows what its own rules mean; now it says so.
4. **Mood invented** (`"warm"`, twice). The enum was listed in a parenthesis at the end of
an English sentence. Now it is its own line: "ровно одно из: neutral, happy, thinking,
tired, confused."
Plus two non-prompt faults the run exposed:
- **The no-parse fallback was English.** When the model returned nothing usable, the body
became `fmt.Sprintf("%s — %s", rule, sev)``"water — care"` — and that string went to
a Russian TTS. Now it falls back to plain Russian.
- **Durations were English.** `humanDur` returns "3 hours"; it was landing verbatim inside
Russian sentences. Nudges now use a Russian formatter.
## Three iterations, and what each taught
| | score | change |
|---|---|---|
| baseline | 0/15 | — |
| iter 1 | 2/15 | Russian prompt, filled-in examples, Russian durations |
| iter 2 | 11/15 | required keyword per rule, one example instead of five, Russian fallback |
| iter 3 | **13/15** | examples moved to topics that are not rules |
The interesting step is 1 → 2. Fixing the placeholder did not fix the disease, it moved it:
the model stopped copying `"..."` and started copying my first example instead. Five nudges
in a row came back as `"Ты не пил воду три часа. Налей стакан."` regardless of the rule.
**A small model copies the nearest concrete text in its prompt.** That is one failure mode
with two symptoms. The fix that stuck was making the examples about laundry and a laptop
battery — topics no rule ever produces, so copying them is visible in the score rather than
invisibly passing the water cases.
## Do not oversell 13/15
Seven of the thirteen passes are the **deterministic fallback**, not the model:
`"Напоминаю: таблетки."`, `"Сервис не отвечает."`, `"Критический алярм: проверь диск."`,
`"Ты давно не пил воду."`. Those are strings this commit added to Go. The model returned
nothing parseable and the fallback scored.
So the honest reading is roughly **6/15 from the model, 7/15 from a fallback, 2/15 failing**.
The prompt fix is real — `"..."` is nearly gone and the language and mood checks are clean —
but a large part of the jump is that failure now degrades into Russian instead of into
`"water — care"`. That is a genuine improvement for the operator and a weak one for the model.
The two remaining failures: one `"..."` recurrence (`routine-stretch`) and one meal nudge
that never says food.
## Tried and reverted: an example-led nudge prompt (#393)
The idea was that a 0.8B copies examples better than it follows rules, so the nudge prompt
was rewritten to lead with five on-topic examples (water, break, pills, morning, service) and
the prose rules were compressed to pay for the tokens: 1190 chars down to 986.
It measured **worse**, three runs each side, same llama-server, same fixture:
| run | before | after |
|---|---|---|
| 1 | 12/15 (address 14) | 11/15 (address 13) |
| 2 | 13/15 (address 15) | 12/15 (address 15) |
| 3 | 14/15 (address 15) | 11/15 (address 12) |
`feminine` and `hisgender` were 15/15 on all six runs, so they measure nothing here. The
regression is all in `address`: 44/45 before, 40/45 after. Formal "вы"/"ваше" and plural
imperatives came back, and so did `"..."`.
Two likely causes, both about the same thing — **examples do not carry a prohibition**. The
old prompt spent a whole sentence on «говоришь на "ты", в единственном числе»; the new one
demoted that to one item in a long "никогда" list, and the model stopped obeying it. And
making the examples on-topic let their *wording* leak: a break case came back as
«Вы давно не пили воду. Выпей стакан.» — the water example, verbatim, in the wrong slot.
That is exactly the failure the laundry/laptop examples were chosen to avoid.
Change reverted. What survives is the measurement: a rule the model must obey needs its own
sentence, and examples must stay off-topic. Also note the before side alone spans 1214 of
15 — this fixture cannot resolve anything smaller than about three cases.
## Broken, found, not fixed
1. ~~**`checkFeminine` only catches half the constraint.**~~ **Fixed** (#381). It scanned for
masculine self-reference only, so three messages that addressed the *owner* in the feminine
("ты давно не отдыхал**а**") scored clean. There is now a second check, `hisgender`: a
feminine past-tense verb (-ла/-лась) in a sentence addressed to him ("ты", "тебе", "твой")
fails, unless the verb is hers ("я заметила", "напомнила тебе"). It is a suffix rule, not a
parser — see the comment in `checks.go` for what it misses. A fresh 15-case run after adding
it scored **12/15** with `hisgender` 15/15; the model did not repeat the feminine address in
that sample, and the check is pinned by unit tests on the recorded bad strings instead.
2. **Grammar is not checked at all, and it is bad.** `"Он не ел 11 дней"` (it was 11 hours),
`"Сонуждились 7 дней"` (not a word), `"Они забыли воду"` (wrong person entirely). Every
one of these passes all six checks. The fixture measures properties, not fluency, and at
0.8B fluency is the binding constraint.
3. **Unit confusion.** The model turns hours into days about a third of the time. The
prompt now says "11 ч"; it reads it as days.
4. **11s p50.** Unchanged and untouched here. A nudge the model takes eleven seconds to
word has missed its moment. Worth its own task.
5. **The keyword hint is close to teaching to the test.** `ruleKeywords` names the word the
on-topic check looks for. It is defensible — the daemon genuinely knows its rule topics
and the model genuinely cannot infer them from `netdata_critical` — but the on-topic
number is softer than the others because of it.
## Next steps
1. ~~**Add a second-person gender check**~~ — done, `hisgender` in `checks.go` (#381).
2. **Decide whether the fallback should count as a pass.** Right now `Score` cannot tell a
model answer from a fallback. Either mark fallback bodies in `PhrasedNudge` or count them
in their own column. Without that, any future prompt change can score well by failing
more.
3. **Attack the 11s.** Nudge phrasing is short and non-interactive; thinking off is the first
thing to try, as it was for routing (#376).
4. **Re-measure when #122 lands.** The CPT'd Qwen3-1.7B is the target. 13/15 with seven
fallbacks is the floor it has to beat, and the fluency problems above are the ones a
bigger, Russian-trained checkpoint should actually fix.
+4 -1
View File
@@ -90,4 +90,7 @@ later* is the worker + RAG.
4. **Deferred work** — larger reasoner, custom Piper voice and other expansions.
## Non-goals (unchanged)
Never phones home. Not a nag. Not autonomous. Feminine-gendered RU self-ref.
Not a nag. Not autonomous. Feminine-gendered RU self-ref. No telemetry, no
cloud model, no third-party account — but she MAY read external sources to
answer world questions (Kiwix first, search optional). "Never phones home" as
an absolute is deprecated, owner's call 2026-07-31; see CLAUDE.md § Non-goals.
+250
View File
@@ -0,0 +1,250 @@
# Note recall evaluation — 31-07-2026
The operator's goal is that Maven "memorize/note things … and know more about me/world". This
measures whether the note/recall path delivers that.
- Fixture + scorer: `internal/memory/recalleval/` (`ru_recall_v1.json`, 30 cases)
- Reproduce: `make eval-recall` — hash ratchet always, ONNX when `deps/` is present
- Commit: `43470ab` (harness)
Each case inserts its own 3 notes **plus 12 shared filler notes** into a fresh store, embeds the
query, takes the top 3 — the read path `cmd/mavend/voice.go` runs for `IntentQuery`. Filler is
load-bearing: with 3 notes and a top-3 search, recall@3 is 100% by construction. 25 answerable
cases (paraphrased queries, homelab and preference content, 9 with a plausible second note) and 5
that must recall **nothing**. `TestFixtureIsParaphrased` fails the build if a query shares over half
its words with its note; equal-score ties count as ties, not recall.
## Results
| | recall+hash (CI ratchet) | recall+onnx (deployed) |
|---|---|---|
| **recall@1** | 36.0% (9/25) | **60.0% (15/25)** |
| recall@3 | 76.0% (19/25) | 80.0% (20/25) |
| **answered after the 0.55 gate** | **0.0% (0/25)** | **48.0% (12/25)** |
| wrong note on top / tie on top | 9 / 7 | 10 / 0 |
| ranked first, then silenced by the gate | 9 | 3 |
| **false recall** | 0/5 | **1/5 (20%)** |
| top-1 score when right, min / median | n/a | 0.559 / 0.678 |
| top-1 when it must stay silent, median / max | 0.000 / 0.144 | 0.470 / **0.567** |
| RU / EN / `hard` cases passed | 4/24 / 1/6 / 0/11 | 13/24 / 3/6 / 2/11 |
| latency p50 / p95 / max | 49µs / 70µs | 59ms / 148ms / 194ms |
Never compare a hash-embedder number to an ONNX one — the hash floor is lexical and exists only so
CI has a deterministic ratchet with no model files.
## Findings
### 1. Real recall is 48%, not 60%
The right note ranks first 60% of the time, but the daemon only *says* it 48% of the time — three
more cases rank first and are then silenced by `voice.go:776`'s `queryMinScore`. **Roughly one
useful question in two gets "не знаю".** This is not a working memory yet.
### 2. The gate cannot separate a real recall from a false one — the distributions overlap
Right-note top-1 scores start at **0.559**. Must-stay-silent top-1 scores reach **0.567**. No
threshold keeps every real recall and rejects every false one. From the sweep: gate 0.50 → 13/25
answered, 1/5 false; **0.55 (default) → 12/25, 1/5**; **0.60 → 10/25, 0/5**; 0.70 → 5/25, 0/5. What
the data says about `DefaultQueryMinScore` (`internal/config/config.go:392`): **0.55 is
slightly too loose** — it admits one confident wrong answer ("как зовут сестру моего коллеги"
recalls "выучил пару аккордов на гитаре" at 0.567), which the spec ranks as worse than a gap. 0.60
silences all five and costs 8 points of real recall. Left alone as instructed; the overlap means
the threshold is the wrong dial anyway (finding 3).
### 3. Filler notes outrank the right answer — the model scores similarity, not relevance
`models/embedder/` is **paraphrase-multilingual-MiniLM-L12-v2** (`Makefile:119`), a *symmetric*
paraphrase model. It scores "do these sentences look alike", not "does this passage answer this
question", so question-shaped queries drift to whatever note is stylistically closest. "из-за чего
кончилось место" and "откуда берётся токен бота" both return `выучил пару аккордов на гитаре`
(0.730, 0.729); "как я восстановил конфиги" returns a bootloader note at 0.703 with the right note
not even in the top 3. An unrelated guitar note beating a homelab note at 0.73 is not a tuning
problem — an asymmetric retrieval model (`multilingual-e5-small`, with `query:` / `passage:`
prefixes) is the targeted fix, and it would move findings 1 and 2 together. Separately:
`deploy/mavend.json:39` loads a 470MB fp32 `model.onnx` while `make download-embedder` fetches
`model_quantized.onnx` — not the same file.
`hard` cases score **2/11**: every one is a query where the operator did not reuse his own words.
That is the normal case weeks later, and exactly what DESIGN.md's "recall when relevant" promises.
### 4. The memory-store recall branch is dead for notes
`voice.go:776` only reaches `h.memStore.Search` when the notes-RAG top score is already below
`queryMinScore`, and `bestRecall` (`cmd/mavend/recall.go:19`) then applies the **same** gate to the
same vector. A note is indexed in both places with the same embedding, so if it failed the gate in
`QueryNotes` it fails again here — the branch can only ever return a **fact**. Its comment calls it
"additive"; for notes it is not.
**Fixed (Vikunja #373).** The memory pass now runs *first*, as one search over notes and facts with
one gate, so whichever memory is clearly the best match answers — note or fact. The notes-only pass
stays behind it for notes the vector index does not hold. No threshold changed, so the set of
questions Maven answers is the same; only which memory answers them. The fixture gained two mixed
note+fact cases (`ru-mixed-031`, `ru-mixed-032`), which is why the counts below are out of 27
answerable cases and not 25: hash recall@1 36.0% (9/25) → 37.0% (10/27), e5 recall@1 72.0% (18/25) →
70.4% (19/27) with answered-after-gate 68.0% → 66.7% and false recall unchanged at 1/5.
### 5. Ranking has no recency or type signal, and the store is not the bottleneck
`internal/store/notes.go:67` sorts by cosine and uses `ts` only to break an exact float tie, which
never happens; `kind` never enters the ranking. Meanwhile `TestPersistentStoreScoresTheSame` scores
sqlite-backed `store.MemoryStore` and `memory.InMemoryStore` identically — both full-scan cosine
(`internal/store/memory.go:64`) at ~150µs over 42 rows against a ~59ms query embed. An ANN index is
not the problem to solve.
## Re-measured after the embedder swap — 31-07-2026, later the same day
Changed: `models/embedder/` is now **multilingual-e5-small** (quantized, 118MB), with `query: ` in
front of a question and `passage: ` in front of a stored note (Vikunja #371). `deploy/mavend.json`
and `make download-embedder` now name the same file, and it is the quantized one — that is what the
column below measures (Vikunja #372). Everything else is unchanged: same fixture, same store, same
0.55 gate. The old column is the baseline and is left as it was.
| | recall+onnx, MiniLM (baseline) | recall+onnx, e5-small (new) |
|---|---|---|
| **recall@1** | 60.0% (15/25) | **72.0% (18/25)** |
| recall@3 | 80.0% (20/25) | 84.0% (21/25) |
| **answered after the 0.55 gate** | 48.0% (12/25) | **72.0% (18/25)** |
| wrong note on top / tie on top | 10 / 0 | 7 / 0 |
| ranked first, then silenced by the gate | 3 | 0 |
| **false recall** | 1/5 (20%) | **5/5 (100%)** |
| top-1 score when right, min / median | 0.559 / 0.678 | 0.791 / 0.857 |
| top-1 when it must stay silent, median / max | 0.470 / 0.567 | 0.815 / 0.835 |
| RU / EN / `hard` cases passed | 13/24 / 3/6 / 2/11 | 14/24 / 4/6 / 5/11 |
| latency p50 / p95 / max | 59ms / 148ms / 194ms | 18ms / 37ms / 49ms |
### What moved
Ranking got better and got faster. Half the previously-unwinnable `hard` cases now pass (2/11 →
5/11), the guitar note no longer beats the docker-logs note, and the gate stops silencing notes that
already ranked first. The quantized e5 is also ~3x quicker than the fp32 MiniLM it replaces.
### What got worse: the gate is now a no-op
e5 packs every cosine into a narrow high band. Right-note scores start at 0.791; must-stay-silent
scores reach 0.835. **The distributions still overlap, and now they overlap above the gate**, so
0.55 admits everything and false recall goes from 1/5 to 5/5. The sweep:
```
gate 0.500.70: answered 18/25 (72%) false recall 5/5
gate 0.80: answered 17/25 (68%) false recall 4/5
gate 0.90: answered 0/25 ( 0%) false recall 0/5
```
There is no value that keeps real recall and rejects made-up questions — same conclusion as before,
now with a wider band and no room at all. `query_min_score` was left at 0.55 as instructed. **The
recommendation is to leave it there and stop tuning it**: any number under ~0.79 is a no-op and
anything above starts cutting real recall long before it stops the false ones. The fix is a margin
gate (`top1 top2 > δ`), next-steps item 3, which is now the top item.
### The prefixes did not do the work
A control run with both prefixes set to the empty string scored the **same** recall@1 (72%), a
slightly better recall@3 (88%) and the same 5/5 false recall. So on this fixture the gain comes from
the model, not from the `query:` / `passage:` split. The prefixes are kept because they are how e5
was trained and the split is the right shape for the read path, but they are not worth defending on
this evidence — a bigger fixture may say otherwise.
### Stored vectors from the old model are now junk
Cosine between a MiniLM vector and an e5 vector means nothing. Every row already in `notes` and in
the vector memory table was written by the old model, so after this deploy they will score as noise
against a new query. A live database needs every note and fact re-embedded before recall works at
all. Filed as its own task.
## Margin gate — 31-07-2026, third run
Next-steps item 3, done. The absolute gate is replaced by a **margin gate**: answer only when the
top hit beats the runner-up by more than delta (`top1 top2 > δ`). Same fixture, same e5 embedder,
same store as the run above. `internal/memory/gate.go` holds the check; both read paths call it
(`cmd/mavend/recall.go` and the notes-RAG branch in `voice.go`). New knob `voice.query_min_margin`
in `deploy/mavend.json`, default 0.008.
### Why the absolute gate could not work, in one line of data
The harness now prints the margin distributions, and they barely overlap where the raw scores
overlap completely:
| | top-1 score | margin (top1 top2) |
|---|---|---|
| right note first (n=18) | min 0.810, median 0.862, max 0.890 | min 0.001, median 0.029, max 0.053 |
| must stay silent (n=5) | min 0.795, median 0.815, max 0.835 | min 0.000, median 0.002, **max 0.019** |
Four of the five must-be-silent cases have a margin at or under 0.002 — when there is nothing to
recall, e5 finds several notes equally close and no clear winner. That is the signal the absolute
score throws away.
### The delta sweep
Absolute gate held at 0.55 throughout.
```
delta 0.000: answered 18/25 (72%) false recall 5/5
delta 0.002: answered 17/25 (68%) false recall 3/5
delta 0.005: answered 17/25 (68%) false recall 2/5
delta 0.008: answered 17/25 (68%) false recall 1/5 <- chosen
delta 0.010: answered 15/25 (60%) false recall 1/5
delta 0.012: answered 14/25 (56%) false recall 1/5
delta 0.015: answered 12/25 (48%) false recall 1/5
delta 0.020: answered 11/25 (44%) false recall 0/5
delta 0.025: answered 9/25 (36%) false recall 0/5
delta 0.030: answered 8/25 (32%) false recall 0/5
delta 0.040: answered 4/25 (16%) false recall 0/5
delta 0.050: answered 2/25 ( 8%) false recall 0/5
delta 0.060: answered 0/25 ( 0%) false recall 0/5
```
### Chosen: δ = 0.008
It is the best point on the frontier, not a taste call. **0.008 dominates 0.010, 0.012 and 0.015
outright** — same 1/5 false recall, 8 to 20 points more real recall. Everything below it buys recall
back only by admitting more false recalls (0.005 → 2/5, 0.002 → 3/5). The next real improvement is
0.020 at 0/5 false, and it costs 24 points of recall to get there.
The brief's bar was "recall above 60% with false recall at 1/5 or better". 0.008 clears it with room:
68% and 1/5.
### Before / after
| | absolute gate 0.55 (previous) | margin gate δ=0.008 |
|---|---|---|
| recall@1 (ranking, ungated) | 72.0% (18/25) | 72.0% (18/25) — unchanged, the gate does not rank |
| **answered after the gate** | 72.0% (18/25) | **68.0% (17/25)** |
| **false recall** | **5/5 (100%)** | **1/5 (20%)** |
| fixture cases passed | 18/30 | **21/30** |
Four false recalls removed for one real answer. That is the trade the spec asks for — she is not a
guesser-of-truth. The one survivor is `en-pref-025` ("should i be offered wine"), which recalls a
filler note at 0.796 with a 0.019 margin: the widest silent-case margin in the fixture, and it sits
inside the real-recall range, so no delta removes it without taking real answers with it.
### Does the absolute cutoff still earn its keep? Marginally — kept
On this fixture with e5 it is a **no-op**: the lowest right-note score is 0.791, so 0.55 rejects
nothing the margin does not already reject. It is kept for two reasons, neither glamorous. It still
does real work for the hash embedder (its own sweep shows answers dropping from 16% to 0% between
0.30 and 0.50), and it is the only thing standing between the user and a reply built from a store
where everything is far away but one row happens to be a little less far — a near-empty database, or
the stale-vector case below. Cheap insurance, no measured cost. If a later embedder makes it bite,
the sweep is one command.
### Caveat on the numbers
Five must-be-silent cases is a thin basis for a 4-point decision. 1/5 and 2/5 differ by one case.
The shape of the frontier is trustworthy — margins separate, absolute scores do not — but δ=0.008
itself should be re-read off a bigger fixture (next-steps item 6) before anyone defends the third
decimal.
## Next steps — ordered by value-to-risk; nothing here is a decision
1. **Swap the embedder to `multilingual-e5-small` with `query:`/`passage:` prefixes.** One config
change plus a prefix in `onnxembedder.go`, re-measurable in one command.
2. **Re-run `make eval-recall`, then set the gate from the sweep** — not before. Any
`query_min_score` picked against today's embedder describes a model on its way out.
3. ~~**Replace the absolute-score gate with a margin gate**~~ — done, see the section above.
δ=0.008, false recall 5/5 → 1/5.
4. **Delete or repair the dead `memStore` branch** at `voice.go:776` — search before the gate,
gate it separately, or restrict it to facts and say so.
5. **Add a mild time decay to ranking** — the newest statement of a preference is the true one.
6. **Grow the fixture from real misses.** 30 cases can rank two embedders, not trust 4 points.
7. **Re-measure end to end.** Recall is gated twice — the utterance must first route to `query`,
which the routing eval puts at ~50%. The product is ~24%, and that is what he experiences.
+138 -5
View File
@@ -40,6 +40,139 @@ model → classifier as failure floor.
Never compare a hash-embedder run to an ONNX one.
## Re-measured after the prompt fix
The table above is the **baseline at commit `46259b4`**, kept as-is. The prompt fix (query
tested before fact, plus `repeat_penalty` and a bounded grammar string) was then measured on
an otherwise idle box — no other eval sharing llama-server, so these latencies are real
rather than contention.
| | llm-only (0.8B) | cascade+llm (0.8B) | llm-only, thinking off |
|---|---|---|---|
| **intent-only accuracy** | 48.7% → **61.8%** | 50.0% → **63.2%** | **67.1%** |
| full accuracy (intent+slots+gate) | 23.7% → **38.2%** | 32.9% → **47.4%** | **42.1%** |
| route errors | 2 → **0** | 0 → 0 | **0** |
| p50 / p95 latency | **1.08s / 1.55s** | **1.04s / 1.53s** | **0.93s / 1.41s** |
Three things this run settles:
1. **The prompt fix holds.** An earlier contended run reported 60.5% / 36.8% for llm-only;
the quiet run gives 61.8% / 38.2%. Close enough to call the gain real, and the earlier
run's 4-5s latency figures were contention, not the model.
2. **`query→fact` fell from ×15 to ×7**, and both unparseable replies are gone. Zero route
errors in every LLM configuration.
3. **`note→fact ×4` is real, not noise.** It shows up in the quiet run too. The agent that
wrote the prompt fix suspected its own change might have caused it by pulling assertive
`запиши что…` phrasings toward fact, and that suspicion stands — all five `ru-note-*`
cases now land on fact. Tracked as Vikunja #375.
The `thinking off` column above read as the best configuration measured so far (Vikunja #376).
**It was wrong** — see the controlled re-run below. Ignore that column.
Still `6 / 6` missed clarify — the router has no way to say "I don't know" (Vikunja #359).
That is unchanged by anything here.
## Thinking off — 31-07-2026, controlled re-run (Vikunja #376)
The "thinking off wins by 6 points" observation above **does not hold**. It was a measurement
artefact, and the earlier table's `thinking off` column should be ignored.
The thinking-off variant was scored by a hand-rolled HTTP client living in the test file
instead of `llm.Client`. That copy did not send `repeat_penalty`, which the real router does
send (`routeRepeatPenalty = 1.15`). So the two columns differed on two axes at once, and the
one that mattered was the penalty, not the thinking mode.
Re-measured with everything else held equal — same fixture, same prompt, same grammar, same
sampling, same idle box, the three configurations run back to back and never concurrently:
| | llm-only, thinking on | llm-only, thinking off | cascade+llm |
|---|---|---|---|
| intent-only accuracy | 59.2% (45/76) | 59.2% (45/76) | 61.8% (47/76) |
| full accuracy (intent+slots+gate) | 38.2% (29/76) | 38.2% (29/76) | 57.9% (44/76) |
| route errors | 3 | 3 | 0 |
| grammar violations | 3 (all 3 route errors) | 3 (same 3 cases) | 0 |
| missed clarify | 5 / 6 | 5 / 6 | 5 / 6 |
| p50 latency | 836ms | 920ms | 810ms |
| p95 latency | 1.41s | 2.00s | 1.31s |
Thinking off is not just a tie on the headline numbers — it is identical case for case, with
the same confusion matrix and the same three unparseable replies. The latency difference is
run-to-run noise on one box, and it points the wrong way here.
The reason is simpler than any accuracy argument: **this llama-server build ignores the
request-level thinking switch for this model.** Probed directly against the running server
with `chat_template_kwargs.enable_thinking = false`, `chat_template_kwargs.thinking = false`
and top-level `reasoning_budget = 0` — all three return a byte-identical answer with the
thinking trace still in `reasoning_content`, and the server reports the prompt prefix as
cached, meaning the rendered template did not change. There was never anything being turned
off, which is also why the numbers match exactly.
Nothing was defaulted. `internal/llm` still has no `chat_template_kwargs` field, `VoiceConfig`
has no thinking flag, and `deploy/mavend.json` is unchanged. The misleading third
configuration is removed from `internal/router/eval` so the table it produced cannot be quoted
again.
Two caveats worth saying out loud:
- **The fixture is 76 cases.** A 6-point difference on 76 cases is roughly 4-5 cases and would
not have been worth trusting even if it had reproduced. This one was exactly 0 cases, which
is a much easier call.
- **This is one server build and one checkpoint** (`b9351`, Qwen3.5-0.8B Q4_K_M). If the
#122 checkpoint or a newer llama.cpp does honour the switch, the question reopens — but it
reopens as an unmeasured question, not as a 6-point win.
Phrasing was **not** measured. Whether thinking helps there is still open, and now also blocked
on the same "can we even turn it off" question.
## Clock and calendar rule — 31-07-2026 (Vikunja #374)
`routeSystem` never said whether "который час" or "какое число завтра" are `system` or
`query`, and `system→query ×4` showed up in every run. The rule added says: the clock and the
calendar date themselves are `system`; what is *written in* the calendar or in memory
("что у меня завтра", "какие есть напоминания") stays `query`; and a time named inside a
request ("напомни завтра…") is just a detail of the request, not a reason for `system`.
That split is not a preference. In `cmd/mavend/voice.go` only `replySystem` owns the clock and
the date formatter, so a clock question routed to `query` falls into the embedder + note RAG
and answers "не знаю". The agenda, on the other hand, is answered by `ParseCalendarDate` +
`CalendarEvents` *inside* the `query` branch, so that side has to stay `query`. The rule sits
above the question test because every one of these utterances carries a question word and a
later rule would never be reached.
The fixture is now 77 cases: one calendar-agenda case was added
(`ru-query-019` "что у меня стоит в календаре на послезавтра", intent `query`) specifically so
an over-broad system rule cannot pass unnoticed. The clock/date cases (`ru-sys-001/002/005`,
`en-sys-001`) already existed.
Three runs, same box, back to back, never concurrently:
| | baseline | first rule (too broad) | rule as committed |
|---|---|---|---|
| llm-only intent-only | 59.2% (45/76) | 54.5% (42/77) | 59.7% (46/77) |
| llm-only full | 38.2% | 35.1% | 39.0% |
| llm-only route errors | 3 | 4 | 5 |
| llm-only p50 | 1.09s | 0.91s | 0.93s |
| cascade+llm intent-only | 61.8% (47/76) | 58.4% | 62.3% (48/77) |
| cascade+llm full | 57.9% | 54.5% | 59.7% |
| cascade+llm route errors | 0 | 0 | 0 |
| cascade+llm p50 | 0.91s | 0.80s | 1.04s |
**The targeted bug is fixed and the headline number did not move.** `system→query ×4` is gone
in both LLM configurations — the `time` and `date` tags go from 0/2 and 0/2 to 2/2 and 2/2 —
but the model then over-applies the rule, and `query→system ×5` plus `reminder→system ×2`
appear where they did not exist before. Net accuracy is a wash, inside the noise of a 77-case
fixture.
The first attempt is shown because it is the honest history: it said "спрашивает время, дату
или день недели → system" with no scope, which swept up reminders, and it cost 3-5 points. It
was tightened once, on the reasoning that a rule capturing "напомни завтра в 7" is simply
wrong, and not tuned further. The remaining `query/reminder → system` over-trigger is a new,
separate weakness of the sub-1B model and deserves its own task rather than more prompt
kneading against a held-out fixture.
The rule is kept. It is correct about what the daemon can answer, and the failure it replaces
was silent ("не знаю" to "который час") while the one it introduces is loud.
## Findings
### 1. The resident model does route better — 50.0% vs 36.8%
@@ -110,11 +243,11 @@ Note the grammar's `string ::= "\"" ([^"\\] | "\\" .)* "\""` is unbounded, so no
### 7. Two hypotheses tested and closed
- **Thinking mode is a non-issue.** Qwen3.5's template defaults `thinking = 1`, so
grammar-constrained JSON lands in `reasoning_content` with `content` empty —
`llm.Client`'s fallback handles it. A `thinking off` run scored *identically* (18/76,
48.7%, same p50). `internal/llm` deliberately does **not** grow a `chat_template_kwargs`
field.
- **Thinking mode is a non-issue.** Confirmed twice now, the second time properly — see the
controlled re-run section. Grammar-constrained JSON lands in `reasoning_content` with
`content` empty and `llm.Client`'s fallback handles it; the request-level switch does
nothing on this build. `internal/llm` deliberately does **not** grow a
`chat_template_kwargs` field.
- **Runaway array repetition does not reproduce.** An isolated smoke test with a stripped
grammar emitted `{"intent":"reminder"}` until `MaxTokens`; under the real `routeSystem`
prompt the few-shot examples anchor it to one object. 2 errors in 76, not 76.
+150
View File
@@ -0,0 +1,150 @@
# Conversational phrasing eval — 31-07-2026
Every score measured tonight, on the three paths the nudge eval never touched:
chat, query-with-notes, and general knowledge.
**Short version: the plumbing got fixed and the score barely moved.** Grammar and
Russian prompts together took the composite from ~9 to ~14 of 27. Everything
still failing is the model not knowing things or not holding a constraint, and
prompting is out of levers. Settles the measurement half of Vikunja #395 / #398 /
#400.
## How to reproduce
```sh
# llama-server: -c 4096 -ngl 99 -t 6, model /mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf
MAVEN_LLM_URL=http://127.0.0.1:18099 no_proxy=127.0.0.1,localhost \
deps/go/go/bin/go test -count=1 -timeout 40m \
-run TestLLMTalkBaseline ./internal/phraser/eval/ -v
```
Three runs per configuration, always. The fixture is 27 cases, so one reply
changing moves the composite by 3.7 points — a single run cannot tell a real
change from sampling noise. This was learned the expensive way: an earlier claim
that "one nudge case fails every run" turned out to be three different cases
across three runs.
**Run the box otherwise idle.** See the contamination note at the bottom.
## Composite, per configuration
| config | overall /27 | chat /9 | query /9 | knowledge /9 | canned fallbacks |
|---|---|---|---|---|---|
| baseline, no grammar | 7, 12, 7 | 1, 1, 0 | 2, 4, 2 | 4, 7, 5 | 0, 0, 0 |
| + GBNF grammar (#398) | 14, 15, 8 | 1, 3, 0 | 5, 6, 3 | 8, 6, 5 | 0, 0, 0 |
| + Russian prompts (#400) | 11, 17, 15 | 1, 5, 3 | 5, 6, 8 | 5, 6, 4 | 0, 0, 0 |
| + truncation fix, 1000ch/768tok | 12, 13, 10 | 2, 2, 1 | 7, 7, 5 | 3, 4, 4 | 3, 3, 6 |
| + rebalanced, 600ch/1024tok | **void — contaminated** | | | | |
"Canned fallbacks" counts replies that came back as the hardcoded `"не знаю."`
or `"поговорили."`. It is not a check, it is a health signal: those strings mean
the phraser gave up, and the eval scores them as ordinary bad replies.
## Per-check
| check | no grammar | + grammar | + RU prompts | + truncation fix |
|---|---|---|---|---|
| nonempty | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 |
| ellipsis | 20, 19, 23 | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 |
| lang | 13, 16, 15 | 23, 26, 26 | 25, 26, 25 | 26, 27, 27 |
| feminine | — | — | 25, 24, 26 | 25, 25, 27 |
| address | — | — | 21, 22, 22 | 22, 21, 22 |
| ontopic | — | — | 17, 24, 18 | 17, 19, 14 |
`nonempty` reading 27/27 everywhere is not good news — it was a broken check.
It tested for a non-blank string, so replies of literally `{` and `"15-16"`
passed it. Fixed on `overnight/fix-truncation`; it needs a letter now.
## What each change actually bought
**GBNF grammar (#398) — the biggest single win.** Qwen3.5-0.8B writes
`Thinking Process:` as plain text with no tags, `stripThink` only handles
`</think>`, so the JSON never closed and the plain-text fallback shipped the
literal reasoning. `ellipsis` went 20→27 and `lang` 13→26. The router had been
using a grammar for ages; the phraser asking nicely in the prompt was the
oversight.
**Russian prompts (#400) — modest, plus a large latency win.** Chat 1.3→3.0
average, query 4.7→6.3, knowledge 6.3→5.0. All inside the run-to-run spread, so
"probably better on the paths it targeted, not provable in three runs". p50
latency dropped from ~11.5s to ~2.3s and that part is consistent across all
three runs — shorter prompts, and she stopped emitting English reasoning first.
**Truncation fix — necessary, and did not help the score.** Two real bugs
(replies of `{`, and a `nonempty` check that passed them), both fixed, and the
composite went nowhere. A complete rambling wrong answer fails the same checks a
truncated one did. Worth doing anyway: the daemon was shipping `{` to a
text-to-speech voice.
## The truncation bug, since the cause was counter-intuitive
The grammar's `string ::= ... {0,400}` rule was the cause, not the token cap.
Measured against Qwen3.5-0.8B at three caps — 256, 768 and 2048 — the reply came
back **exactly 400 characters every time, cut mid-word** (`"Нужно записать и,"`).
Then I raised the bound to 1000 while the cap was 768 tokens and made it worse:
Russian runs ~1.5 characters per token here, so generation died on the *token*
cap instead, mid-object, and the new guard correctly refused it and shipped
`"не знаю."` — 3, 3 and 6 fallbacks per run, from zero. **The two limits have to
agree.** 600 characters needs ~400 tokens; the cap is 1024.
## Where the remaining failures live
`address` is stuck at 21-22 of 27 and `ontopic` at 14-19. Both resist prompting.
**The prompt now explicitly forbids exactly what she does.** It says never "вы",
use the singular — and she writes `вашей`, `подождите`, `делаете`, `хотите`,
`напишите`. Telling a 0.8B "never do X" does not work. Same for
`feminine`: `я готов`, `я понял`, `я нашел`, `я заметил`, `я сказал`.
**Some of `ontopic` is the fixture, not the model.** `chat-how-are-you` got
`"Привет! Я здесь, чтобы поговорить. Как дела сегодня?"` — a fine reply that
fails because `want_any` is `[норм, хорош, порядк, тут, работ]`. It fails in
every run, so it inflates the count. The `ontopic` column currently measures the
fixture as much as the model. Not fixed yet, deliberately: changing it would
break comparability with the runs above.
**Two replies worth reading, because they are not fixable by prompting:**
- Thunder and lightning: *"Скорость молнии — 8-10 тысяч километров в секунду, но
звук — 300 метров в секунду, что делает молнию громче."* Confidently wrong,
and it concludes lightning is *louder* rather than sound being *slower*.
- "расскажи обо мне": *"Ты — прекрасное существо, с душой и вниманием… Спасибо за
твою улыбку… О тебе — заповедь любви."* Sycophantic filler, zero information,
and precisely the "not a relationship" non-goal.
- Boiling an egg: `"15-16"` one run, `"1"` another. No unit, wrong number.
The first argues for reading instead of recalling (#403 — Kiwix retrieval scores
8/8 on the same questions given English keywords). The second and third argue
for templates on the paths where correctness matters (#392).
## Contamination note — how the last row got voided
I started the query-rewrite agent against the same llama-server the sweep was
using, and assumed contention would only affect latency. It did not. The
knowledge path collapsed to 0 of 9 with eight canned `"не знаю."` replies, p95
tripled to 23.7s, and **the report still said "0 errors"**.
That is Vikunja #397, and it is worse than filed: a merely *busy* server
produces a clean-looking report with a third of the fixture silently answering
`"не знаю."`. `PhraseChat` and `PhraseQuery` swallow every failure and return a
hardcoded string, so infrastructure trouble is indistinguishable from bad
phrasing in the score. The talk test guards the *start* and *end* of a run with
a model check, which catches a dead server but not a loaded one.
**Until #397 is fixed, treat any run made on a busy box as void.**
## Next
- Re-run 600ch/1024tok clean, to fill the void row.
- Score `Qwen3.5-2B-UD-Q4_K_XL` (already at `/mnt/hdd1/llms/qwen3.5/`, never
measured) on this fixture and the router fixture. Not the 4B — too big for
this box, owner's call.
- Newer sub-500M candidates (LFM2.5 200M/300M) are worth a run for routing.
Note `MODEL-BAKEOFF-31-07-2026.md` found LFM2.5-**1.2B** worse than
Qwen3.5-0.8B at Russian routing and 2.4× slower — but those are a different,
older generation, so that result does not predict the small ones.
- Fix `chat-how-are-you`'s `want_any`, and re-baseline once, so `ontopic`
measures the model.
- #397 first if anything, since it decides whether any of the above is
trustworthy.
+75 -144
View File
@@ -1,17 +1,24 @@
// mavcaldav — the CalDAV poller module.
// mavcaldav — the CalDAV module: reads calendars into facts, and renders
// maven's own reminders back out to a calendar she owns.
//
// Polls a Radicale (or any CalDAV) server for today's events and writes
// `facts (kind=env, source=poll:caldav)` through core's IPC socket.
// Key-free, restart-free, fail-independent — crashes can't touch the
// store key, worst case a stale calendar_busy fact until the next poll.
// READ side (unchanged behaviour): polls a Radicale (or any CalDAV) server for
// today's events and writes `facts (kind=env, source=poll:caldav)` through
// core's IPC socket. Key-free, restart-free, fail-independent — crashes can't
// touch the store key, worst case a stale calendar_busy fact until the next
// poll. Two facts:
//
// Two facts written:
// - calendar_busy ("true"/"false") — read by the loop gate to suppress
// nudges during meetings
// - calendar_event ("<summary> @ <start>-<end>") — per-event for query
//
// Append-only discipline: a fact is written only when its value CHANGED
// vs the latest for that key+source.
// Append-only discipline: a fact is written only when its value CHANGED vs the
// latest for that key+source.
//
// RENDER side (Vikunja #127, off unless -render-url is given): publishes each
// pending reminder as a single-event iCal resource in a collection maven owns.
// The calendar is a view, sqlite is the store — see render.go. The render URL
// must differ from the read URL, checked at startup, so the render target can
// never be a calendar maven is only supposed to read.
package main
import (
@@ -27,6 +34,7 @@ import (
"syscall"
"time"
"github.com/kami/maven/internal/calendar"
"github.com/kami/maven/internal/ipc"
)
@@ -43,6 +51,10 @@ func run(args []string) error {
url := fs.String("url", "", "CalDAV calendar URL, e.g. http://localhost:5232/kami/personal (required)")
user := fs.String("user", "", "CalDAV basic-auth username (required)")
pass := fs.String("pass", "", "CalDAV basic-auth password (required)")
renderURL := fs.String("render-url", "", "CalDAV collection maven publishes her own reminders to; empty disables rendering")
renderUser := fs.String("render-user", "", "basic-auth username for -render-url (defaults to -user)")
renderPass := fs.String("render-pass", "", "basic-auth password for -render-url (defaults to -pass)")
renderDur := fs.Duration("render-duration", calendar.DefaultReminderDuration, "how long a rendered reminder occupies")
interval := fs.Duration("interval", 5*time.Minute, "poll cadence")
timeout := fs.Duration("timeout", 10*time.Second, "per-request HTTP timeout")
if err := fs.Parse(args); err != nil {
@@ -54,6 +66,9 @@ func run(args []string) error {
if *url == "" || *user == "" || *pass == "" {
return fmt.Errorf("-url, -user, -pass are required")
}
if err := checkRenderTarget(*url, *renderURL); err != nil {
return err
}
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
defer stop()
@@ -64,16 +79,36 @@ func run(args []string) error {
}
defer core.Close()
hc := &http.Client{Timeout: *timeout}
p := &poller{
core: core,
http: &http.Client{Timeout: *timeout},
http: hc,
url: strings.TrimRight(*url, "/"),
user: *user,
pass: *pass,
}
var rend *renderer
if *renderURL != "" {
ru, rp := *renderUser, *renderPass
if ru == "" {
ru = *user
}
if rp == "" {
rp = *pass
}
rend = newRenderer(core, hc, *renderURL, ru, rp, *renderDur)
log.Printf("mavcaldav: rendering reminders to %s", *renderURL)
}
log.Printf("mavcaldav: polling %s every %s", *url, *interval)
p.pollOnce(ctx) // fire immediately
tick := func() {
p.pollOnce(ctx)
if rend != nil {
rend.renderOnce(ctx)
}
}
tick() // fire immediately
t := time.NewTicker(*interval)
defer t.Stop()
for {
@@ -82,11 +117,30 @@ func run(args []string) error {
log.Printf("mavcaldav: bye")
return nil
case <-t.C:
p.pollOnce(ctx)
tick()
}
}
}
// checkRenderTarget refuses a render URL that is also a read URL. This is the
// structural half of #127's "cannot write to your work calendar": the write
// credential and the write URL are separate flags, and the one calendar maven
// is known to only read is rejected as a target at startup rather than trusted
// at runtime.
func checkRenderTarget(readURL, renderURL string) error {
if renderURL == "" {
return nil
}
if sameCollection(readURL, renderURL) {
return fmt.Errorf("-render-url must differ from -url: maven renders into a calendar she owns, never into one she reads")
}
return nil
}
func sameCollection(a, b string) bool {
return strings.EqualFold(strings.TrimRight(a, "/"), strings.TrimRight(b, "/"))
}
type poller struct {
core ipc.CoreAPI
http *http.Client
@@ -95,12 +149,6 @@ type poller struct {
pass string
}
type icalEvent struct {
start time.Time
end time.Time
summary string
}
func (p *poller) pollOnce(ctx context.Context) {
now := time.Now()
events, err := p.fetchEvents(ctx, now)
@@ -109,38 +157,30 @@ func (p *poller) pollOnce(ctx context.Context) {
return
}
busy := false
for _, e := range events {
if !now.Before(e.start) && now.Before(e.end) {
busy = true
break
}
}
busyVal := "false"
if busy {
if calendar.Busy(events, now) {
busyVal = "true"
}
// Write calendar_busy on change.
if err := p.writeIfChanged(ctx, "calendar_busy", "poll:caldav", busyVal, now); err != nil {
if err := p.writeIfChanged(ctx, "calendar_busy", calendar.SourcePersonal, busyVal, now, 1.0); err != nil {
log.Printf("mavcaldav: write calendar_busy: %v", err)
return
}
// Write per-event facts (one per event, keyed by event summary + start).
// Write per-event facts (one per event, keyed by day + event summary).
// This lets the note RAG path answer "what's on my calendar" without
// reaching back to Radicale.
for _, e := range events {
val := fmt.Sprintf("%s @ %s-%s", e.summary, e.start.Format("15:04"), e.end.Format("15:04"))
eventKey := fmt.Sprintf("calendar_event_%s_%s", e.start.Format("20060102"), safeKey(e.summary))
if err := p.writeIfChanged(ctx, eventKey, "poll:caldav", val, e.start); err != nil {
log.Printf("mavcaldav: write %s: %v", eventKey, err)
key := calendar.FactKey(e)
if err := p.writeIfChanged(ctx, key, calendar.SourcePersonal, calendar.FactValue(e), e.Start, 1.0); err != nil {
log.Printf("mavcaldav: write %s: %v", key, err)
}
}
}
// fetchEvents GETs the calendar URL and parses VEVENTs from the iCal response.
func (p *poller) fetchEvents(ctx context.Context, now time.Time) ([]icalEvent, error) {
func (p *poller) fetchEvents(ctx context.Context, now time.Time) ([]calendar.Event, error) {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, p.url, nil)
if err != nil {
return nil, err
@@ -162,120 +202,11 @@ func (p *poller) fetchEvents(ctx context.Context, now time.Time) ([]icalEvent, e
return nil, fmt.Errorf("GET %s: %s", p.url, resp.Status)
}
return parseICal(body, now), nil
}
// parseICal scans iCal text for VEVENT components. Returns events that overlap
// with today (UTC day boundaries) to keep the response manageable.
func parseICal(body []byte, now time.Time) []icalEvent {
todayStart := time.Date(now.Year(), now.Month(), now.Day(), 0, 0, 0, 0, time.UTC)
todayEnd := todayStart.AddDate(0, 0, 1)
var events []icalEvent
text := string(body)
for {
veventStart := strings.Index(text, "BEGIN:VEVENT")
if veventStart < 0 {
break
}
text = text[veventStart+len("BEGIN:VEVENT"):]
veventEnd := strings.Index(text, "END:VEVENT")
if veventEnd < 0 {
break
}
block := text[:veventEnd]
text = text[veventEnd+len("END:VEVENT"):]
e := parseVEVENT(block)
if e == nil {
continue
}
// Only keep events overlapping today.
if e.end.After(todayStart) && e.start.Before(todayEnd) {
events = append(events, *e)
}
}
return events
}
// parseVEVENT extracts start, end, summary from a VEVENT block.
// Supports both UTC (DTEND:20260703T100000Z) and local (DTSTART;TZID=...:...)
// formats. Returns nil for all-day events (no DTSTART/DTEND time component) or
// parse failures.
func parseVEVENT(block string) *icalEvent {
var e icalEvent
lines := strings.Split(block, "\n")
for _, line := range lines {
line = strings.TrimSpace(line)
switch {
case strings.HasPrefix(line, "DTSTART"):
if t, ok := parseDT(line); ok {
e.start = t
}
case strings.HasPrefix(line, "DTEND"):
if t, ok := parseDT(line); ok {
e.end = t
}
case strings.HasPrefix(line, "SUMMARY"):
if idx := strings.Index(line, ":"); idx >= 0 {
e.summary = strings.TrimSpace(line[idx+1:])
}
}
}
if e.start.IsZero() || e.end.IsZero() {
return nil
}
return &e
}
// parseDT parses a DTSTART/DTEND value. Supports:
// - UTC: DTEND:20260703T100000Z
// - Local: DTSTART;TZID=Europe/Moscow:20260703T130000
// - Value-date (all-day): DTSTART;VALUE=DATE:20260703 (returns zero time)
func parseDT(line string) (time.Time, bool) {
if strings.Contains(line, "VALUE=DATE:") {
return time.Time{}, false // all-day, skip
}
idx := strings.LastIndex(line, ":")
if idx < 0 {
return time.Time{}, false
}
val := line[idx+1:]
val = strings.TrimSuffix(val, "Z")
// Try UTC first (has Z suffix, or ended in Z before TrimSuffix).
if strings.HasSuffix(line, "Z") {
t, err := time.Parse("20060102T150405", val)
if err != nil {
return time.Time{}, false
}
return t.UTC(), true
}
// Local time — treat as UTC for simplicity (CalDAV server and poller
// run in the same timezone; the gate only needs busy/not-busy accuracy).
t, err := time.Parse("20060102T150405", val)
if err != nil {
return time.Time{}, false
}
return t.UTC(), true
}
// safeKey makes an event summary safe to use as a fact key (alphanumeric + dash).
func safeKey(s string) string {
var b strings.Builder
for _, r := range s {
if (r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z') || (r >= '0' && r <= '9') || r == '-' {
b.WriteRune(r)
} else if r == ' ' || r == '_' {
b.WriteRune('-')
}
}
return b.String()
return calendar.ParseICalDay(body, now), nil
}
// writeIfChanged writes a fact only when the value differs from the latest.
func (p *poller) writeIfChanged(ctx context.Context, key, source, val string, ts time.Time) error {
func (p *poller) writeIfChanged(ctx context.Context, key, source, val string, ts time.Time, confidence float64) error {
prev, err := p.core.LatestFactBySource(ctx, key, source)
switch {
case err == nil && prev.Value == val:
@@ -289,7 +220,7 @@ func (p *poller) writeIfChanged(ctx context.Context, key, source, val string, ts
Key: key,
Value: val,
Source: source,
Confidence: 1.0,
Confidence: confidence,
})
if err != nil {
return fmt.Errorf("write %s: %w", key, err)
+6 -165
View File
@@ -12,7 +12,7 @@ import (
)
type fakeCore struct {
ipc.CoreAPI
ipc.UnimplementedCoreAPI
facts map[string]ipc.Fact // composite key "key|source" → Fact
writeLog []ipc.WriteFactReq
writeErr error
@@ -51,165 +51,6 @@ func (f *fakeCore) WriteFact(_ context.Context, req ipc.WriteFactReq) (int64, er
return int64(len(f.writeLog)), nil
}
// ---------------------------------------------------------------------------
// Parsing tests
// ---------------------------------------------------------------------------
func TestParseICal(t *testing.T) {
now := time.Date(2026, 7, 3, 12, 0, 0, 0, time.UTC)
body := []byte(`BEGIN:VCALENDAR
BEGIN:VEVENT
DTSTART:20260703T090000Z
DTEND:20260703T100000Z
SUMMARY:Morning standup
END:VEVENT
BEGIN:VEVENT
DTSTART:20260703T140000Z
DTEND:20260703T150000Z
SUMMARY:Team sync
END:VEVENT
BEGIN:VEVENT
DTSTART:20260702T140000Z
DTEND:20260702T150000Z
SUMMARY:Yesterday retro
END:VEVENT
BEGIN:VEVENT
DTSTART:20260704T090000Z
DTEND:20260704T100000Z
SUMMARY:Tomorrow standup
END:VEVENT
BEGIN:VEVENT
DTSTART;VALUE=DATE:20260704
DTEND;VALUE=DATE:20260705
SUMMARY:All-day event
END:VEVENT
END:VCALENDAR`)
events := parseICal(body, now)
if len(events) != 2 {
t.Fatalf("got %d events, want 2 (today events, no all-day/past/future)", len(events))
}
// Morning standup — overlaps today.
if events[0].summary != "Morning standup" {
t.Errorf("events[0].summary = %q, want %q", events[0].summary, "Morning standup")
}
wantStart0 := time.Date(2026, 7, 3, 9, 0, 0, 0, time.UTC)
if !events[0].start.Equal(wantStart0) {
t.Errorf("events[0].start = %v, want %v", events[0].start, wantStart0)
}
wantEnd0 := time.Date(2026, 7, 3, 10, 0, 0, 0, time.UTC)
if !events[0].end.Equal(wantEnd0) {
t.Errorf("events[0].end = %v, want %v", events[0].end, wantEnd0)
}
// Team sync — overlaps today.
if events[1].summary != "Team sync" {
t.Errorf("events[1].summary = %q, want %q", events[1].summary, "Team sync")
}
wantStart1 := time.Date(2026, 7, 3, 14, 0, 0, 0, time.UTC)
if !events[1].start.Equal(wantStart1) {
t.Errorf("events[1].start = %v, want %v", events[1].start, wantStart1)
}
wantEnd1 := time.Date(2026, 7, 3, 15, 0, 0, 0, time.UTC)
if !events[1].end.Equal(wantEnd1) {
t.Errorf("events[1].end = %v, want %v", events[1].end, wantEnd1)
}
}
func TestParseVEVENT(t *testing.T) {
// Normal event with TZID in DTSTART and UTC DTEND.
block := "DTSTART;TZID=Europe/Moscow:20260703T130000\nDTEND:20260703T140000Z\nSUMMARY:Stand up meeting"
e := parseVEVENT(block)
if e == nil {
t.Fatal("expected non-nil icalEvent")
}
wantStart := time.Date(2026, 7, 3, 13, 0, 0, 0, time.UTC)
if !e.start.Equal(wantStart) {
t.Errorf("start = %v, want %v", e.start, wantStart)
}
wantEnd := time.Date(2026, 7, 3, 14, 0, 0, 0, time.UTC)
if !e.end.Equal(wantEnd) {
t.Errorf("end = %v, want %v", e.end, wantEnd)
}
if e.summary != "Stand up meeting" {
t.Errorf("summary = %q, want %q", e.summary, "Stand up meeting")
}
// All-day event (VALUE=DATE) → nil.
allDay := "DTSTART;VALUE=DATE:20260703\nDTEND;VALUE=DATE:20260704\nSUMMARY:All-day"
if e2 := parseVEVENT(allDay); e2 != nil {
t.Error("expected nil for all-day event")
}
}
func TestParseDT(t *testing.T) {
tests := []struct {
name string
line string
want time.Time
wantOK bool
}{
{
name: "UTC",
line: "DTEND:20260703T100000Z",
want: time.Date(2026, 7, 3, 10, 0, 0, 0, time.UTC),
wantOK: true,
},
{
name: "local time",
line: "DTSTART;TZID=Europe/Moscow:20260703T130000",
want: time.Date(2026, 7, 3, 13, 0, 0, 0, time.UTC),
wantOK: true,
},
{
name: "all-day",
line: "DTSTART;VALUE=DATE:20260703",
want: time.Time{},
wantOK: false,
},
{
name: "invalid",
line: "DTSTART:garbage",
want: time.Time{},
wantOK: false,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
got, ok := parseDT(tt.line)
if ok != tt.wantOK {
t.Errorf("ok = %v, want %v", ok, tt.wantOK)
}
if !got.Equal(tt.want) {
t.Errorf("got = %v, want %v", got, tt.want)
}
})
}
}
func TestSafeKey(t *testing.T) {
tests := []struct {
input string
want string
}{
{"Stand up meeting", "Stand-up-meeting"},
{"Hello_World", "Hello-World"},
{"special@#$chars!!", "specialchars"},
{"ALL_CAPS_123", "ALL-CAPS-123"},
}
for _, tt := range tests {
got := safeKey(tt.input)
if got != tt.want {
t.Errorf("safeKey(%q) = %q, want %q", tt.input, got, tt.want)
}
}
}
// ---------------------------------------------------------------------------
// Core logic tests
// ---------------------------------------------------------------------------
@@ -221,7 +62,7 @@ func TestWriteIfChanged(t *testing.T) {
t.Run("no previous fact writes", func(t *testing.T) {
fc := &fakeCore{}
p := &poller{core: fc}
err := p.writeIfChanged(ctx, "test_key", "poll:caldav", "hello", now)
err := p.writeIfChanged(ctx, "test_key", "poll:caldav", "hello", now, 1.0)
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
@@ -249,7 +90,7 @@ func TestWriteIfChanged(t *testing.T) {
},
}
p := &poller{core: fc}
err := p.writeIfChanged(ctx, "test_key", "poll:caldav", "hello", now)
err := p.writeIfChanged(ctx, "test_key", "poll:caldav", "hello", now, 1.0)
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
@@ -265,7 +106,7 @@ func TestWriteIfChanged(t *testing.T) {
},
}
p := &poller{core: fc}
err := p.writeIfChanged(ctx, "test_key", "poll:caldav", "new", now)
err := p.writeIfChanged(ctx, "test_key", "poll:caldav", "new", now, 1.0)
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
@@ -280,7 +121,7 @@ func TestWriteIfChanged(t *testing.T) {
t.Run("read error other than ErrNoFact returns error", func(t *testing.T) {
fc := &fakeCore{readErr: fmt.Errorf("connection refused")}
p := &poller{core: fc}
err := p.writeIfChanged(ctx, "fail_key", "poll:caldav", "x", now)
err := p.writeIfChanged(ctx, "fail_key", "poll:caldav", "x", now, 1.0)
if err == nil {
t.Fatal("expected error, got nil")
}
@@ -292,7 +133,7 @@ func TestWriteIfChanged(t *testing.T) {
writeErr: fmt.Errorf("disk full"),
}
p := &poller{core: fc}
err := p.writeIfChanged(ctx, "test_key", "poll:caldav", "hello", now)
err := p.writeIfChanged(ctx, "test_key", "poll:caldav", "hello", now, 1.0)
if err == nil {
t.Fatal("expected error, got nil")
}
+145
View File
@@ -0,0 +1,145 @@
package main
import (
"context"
"fmt"
"io"
"log"
"net/http"
"strings"
"time"
"github.com/kami/maven/internal/calendar"
"github.com/kami/maven/internal/ipc"
)
// renderer is the write half of maven's own local calendar (Vikunja #127).
//
// It is a RENDER TARGET, not a store. sqlite stays canonical: every tick the
// renderer reads the pending reminders out of core and publishes each one as a
// single-event iCal resource in a CalDAV collection maven owns. Nothing is ever
// read back from that collection, and losing it costs nothing — the next tick
// rebuilds it.
//
// It structurally cannot write to a calendar maven only reads. The URL comes
// from its own flag, checked at startup against every read URL (see
// run in main.go), and the only paths it ever addresses carry
// calendar.ReminderUIDPrefix — so even pointed at the wrong collection it can
// only touch resources it created.
type renderer struct {
core ipc.CoreAPI
http *http.Client
url string
user string
pass string
dur time.Duration
// published maps reminder id → the body last successfully PUT, so an
// unchanged reminder costs nothing. Purely an optimisation: a restart
// re-publishes every reminder once, which is idempotent.
published map[int64]string
}
func newRenderer(core ipc.CoreAPI, hc *http.Client, url, user, pass string, dur time.Duration) *renderer {
return &renderer{
core: core,
http: hc,
url: strings.TrimRight(url, "/"),
user: user,
pass: pass,
dur: dur,
published: make(map[int64]string),
}
}
// renderOnce publishes every pending reminder and withdraws the ones that are
// no longer pending. Errors are logged and skipped: a calendar maven cannot
// reach must never break the reminder itself, which lives in sqlite.
func (r *renderer) renderOnce(ctx context.Context) {
reminders, err := r.core.ListReminders(ctx, renderMaxReminders)
if err != nil {
log.Printf("mavcaldav: list reminders: %v", err)
return
}
live := make(map[int64]bool, len(reminders))
for _, rem := range reminders {
if rem.Status != "pending" {
continue
}
live[rem.ID] = true
e := calendar.ReminderEvent(rem.ID, fireTime(rem), rem.Payload, r.dur)
body := calendar.RenderICal([]calendar.Event{e})
if r.published[rem.ID] == body {
continue
}
if err := r.put(ctx, calendar.ReminderPath(rem.ID), body); err != nil {
log.Printf("mavcaldav: render reminder %d: %v", rem.ID, err)
continue
}
r.published[rem.ID] = body
log.Printf("mavcaldav: rendered reminder %d (%s)", rem.ID, e.Summary)
}
for id := range r.published {
if live[id] {
continue
}
if err := r.delete(ctx, calendar.ReminderPath(id)); err != nil {
log.Printf("mavcaldav: withdraw reminder %d: %v", id, err)
continue
}
delete(r.published, id)
log.Printf("mavcaldav: withdrew reminder %d", id)
}
}
// renderMaxReminders bounds the read. Reminders past this count are older than
// anything a calendar view is useful for.
const renderMaxReminders = 200
// fireTime prefers NextFireTs — for a recurring reminder that is the occurrence
// worth showing; FireTs is the original statement.
func fireTime(rem ipc.Reminder) time.Time {
if !rem.NextFireTs.IsZero() {
return rem.NextFireTs
}
return rem.FireTs
}
func (r *renderer) put(ctx context.Context, name, body string) error {
req, err := http.NewRequestWithContext(ctx, http.MethodPut, r.url+"/"+name, strings.NewReader(body))
if err != nil {
return err
}
req.SetBasicAuth(r.user, r.pass)
req.Header.Set("Content-Type", "text/calendar; charset=utf-8")
return r.do(req, name)
}
func (r *renderer) delete(ctx context.Context, name string) error {
req, err := http.NewRequestWithContext(ctx, http.MethodDelete, r.url+"/"+name, nil)
if err != nil {
return err
}
req.SetBasicAuth(r.user, r.pass)
return r.do(req, name)
}
// do runs the request and treats any 2xx, plus 404 on a DELETE, as success —
// a resource that is already gone is the state the caller wanted.
func (r *renderer) do(req *http.Request, name string) error {
resp, err := r.http.Do(req)
if err != nil {
return err
}
defer resp.Body.Close()
io.Copy(io.Discard, io.LimitReader(resp.Body, 1<<16))
switch {
case resp.StatusCode >= 200 && resp.StatusCode < 300:
return nil
case req.Method == http.MethodDelete && resp.StatusCode == http.StatusNotFound:
return nil
}
return fmt.Errorf("%s %s: %s", req.Method, name, resp.Status)
}
+186
View File
@@ -0,0 +1,186 @@
package main
import (
"context"
"io"
"net/http"
"net/http/httptest"
"strings"
"sync"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
)
// reminderCore is a fakeCore that also answers ListReminders.
type reminderCore struct {
fakeCore
reminders []ipc.Reminder
listErr error
}
func (c *reminderCore) ListReminders(context.Context, int) ([]ipc.Reminder, error) {
if c.listErr != nil {
return nil, c.listErr
}
return c.reminders, nil
}
// calSrv records what a CalDAV collection received.
type calSrv struct {
mu sync.Mutex
puts map[string]string
dels []string
status int
*httptest.Server
}
func newCalSrv() *calSrv {
s := &calSrv{puts: map[string]string{}, status: http.StatusCreated}
s.Server = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
body, _ := io.ReadAll(r.Body)
s.mu.Lock()
defer s.mu.Unlock()
switch r.Method {
case http.MethodPut:
s.puts[strings.TrimPrefix(r.URL.Path, "/cal/")] = string(body)
case http.MethodDelete:
s.dels = append(s.dels, strings.TrimPrefix(r.URL.Path, "/cal/"))
}
w.WriteHeader(s.status)
}))
return s
}
func (s *calSrv) putCount() int {
s.mu.Lock()
defer s.mu.Unlock()
return len(s.puts)
}
func TestRenderOncePublishesPendingReminders(t *testing.T) {
fire := time.Date(2026, 8, 1, 18, 30, 0, 0, time.UTC)
srv := newCalSrv()
defer srv.Close()
core := &reminderCore{reminders: []ipc.Reminder{
{ID: 7, FireTs: fire, Payload: "позвонить маме", Status: "pending"},
{ID: 8, FireTs: fire, Payload: "уже сделано", Status: "fired"},
{ID: 9, FireTs: fire, Payload: "отменено", Status: "cancelled"},
}}
r := newRenderer(core, srv.Client(), srv.URL+"/cal/", "u", "p", 0)
r.renderOnce(context.Background())
srv.mu.Lock()
body, ok := srv.puts["maven-reminder-7.ics"]
n := len(srv.puts)
srv.mu.Unlock()
if n != 1 {
t.Fatalf("expected exactly the pending reminder to be published, got %d PUTs", n)
}
if !ok {
t.Fatal("pending reminder 7 was not published")
}
if !strings.Contains(body, "SUMMARY:позвонить маме") {
t.Errorf("payload missing from rendered body:\n%s", body)
}
if !strings.Contains(body, "UID:maven-reminder-7") {
t.Errorf("UID missing from rendered body:\n%s", body)
}
}
func TestRenderOnceSkipsUnchanged(t *testing.T) {
srv := newCalSrv()
defer srv.Close()
core := &reminderCore{reminders: []ipc.Reminder{
{ID: 1, FireTs: time.Date(2026, 8, 1, 9, 0, 0, 0, time.UTC), Payload: "выпить воды", Status: "pending"},
}}
r := newRenderer(core, srv.Client(), srv.URL+"/cal", "u", "p", 0)
r.renderOnce(context.Background())
r.renderOnce(context.Background())
if got := srv.putCount(); got != 1 {
t.Fatalf("an unchanged reminder was re-published: %d distinct PUTs", got)
}
}
func TestRenderOnceWithdrawsResolvedReminders(t *testing.T) {
srv := newCalSrv()
defer srv.Close()
core := &reminderCore{reminders: []ipc.Reminder{
{ID: 5, FireTs: time.Date(2026, 8, 1, 9, 0, 0, 0, time.UTC), Payload: "встреча", Status: "pending"},
}}
r := newRenderer(core, srv.Client(), srv.URL+"/cal", "u", "p", 0)
r.renderOnce(context.Background())
core.reminders[0].Status = "fired"
r.renderOnce(context.Background())
srv.mu.Lock()
dels := append([]string(nil), srv.dels...)
srv.mu.Unlock()
if len(dels) != 1 || dels[0] != "maven-reminder-5.ics" {
t.Fatalf("resolved reminder was not withdrawn: %v", dels)
}
if len(r.published) != 0 {
t.Errorf("published map still holds %v", r.published)
}
}
// A calendar maven cannot reach must never break anything: sqlite is canonical.
func TestRenderOnceSurvivesServerErrors(t *testing.T) {
srv := newCalSrv()
srv.status = http.StatusInternalServerError
defer srv.Close()
core := &reminderCore{reminders: []ipc.Reminder{
{ID: 1, FireTs: time.Date(2026, 8, 1, 9, 0, 0, 0, time.UTC), Payload: "x", Status: "pending"},
}}
r := newRenderer(core, srv.Client(), srv.URL+"/cal", "u", "p", 0)
r.renderOnce(context.Background())
if len(r.published) != 0 {
t.Error("a failed PUT must not be recorded as published, or it never retries")
}
}
func TestRenderOnceUsesNextFireForRecurring(t *testing.T) {
srv := newCalSrv()
defer srv.Close()
next := time.Date(2026, 8, 2, 7, 0, 0, 0, time.UTC)
core := &reminderCore{reminders: []ipc.Reminder{{
ID: 3,
FireTs: time.Date(2026, 8, 1, 7, 0, 0, 0, time.UTC),
NextFireTs: next,
Payload: "зарядка",
Status: "pending",
Cron: "0 7 * * *",
}}}
r := newRenderer(core, srv.Client(), srv.URL+"/cal", "u", "p", 0)
r.renderOnce(context.Background())
srv.mu.Lock()
body := srv.puts["maven-reminder-3.ics"]
srv.mu.Unlock()
if !strings.Contains(body, "DTSTART:20260802T070000Z") {
t.Errorf("recurring reminder should render its next occurrence:\n%s", body)
}
}
func TestCheckRenderTargetRefusesTheCalendarItReads(t *testing.T) {
read := "http://localhost:5232/kami/personal"
if err := checkRenderTarget(read, ""); err != nil {
t.Fatalf("rendering off must be fine: %v", err)
}
if err := checkRenderTarget(read, "http://localhost:5232/kami/maven"); err != nil {
t.Fatalf("a distinct collection must be accepted: %v", err)
}
if err := checkRenderTarget(read, read); err == nil {
t.Error("rendering into the read calendar must be refused")
}
if err := checkRenderTarget(read, read+"/"); err == nil {
t.Error("a trailing slash must not defeat the check")
}
if err := checkRenderTarget(read, strings.ToUpper(read)); err == nil {
t.Error("case must not defeat the check")
}
}
+71
View File
@@ -0,0 +1,71 @@
// actionTable dispatches applyAction's per-intent bodies. Each of the 7
// intents (fact, reminder, note, query, act, chat, system) has one handler
// here with the signature:
//
// func(h *reactiveHandler, ctx context.Context, dec router.Decision) string
//
// same contract as applyAction itself: "" means "let the Replier phrase the
// reply", a non-empty string OVERRIDES it. This is a straight extraction of
// applyAction's old switch cases (formerly ~300 lines in voice.go) — no
// reordering of side effects, no new abstractions inside a handler.
//
// What does NOT belong in this table, because it is not per-intent:
//
// - the dec.Clarify short-circuit ("" when the router's stage-3 fired) —
// stays in applyAction, before dispatch, since it applies to every
// intent identically.
// - the destructive-act confirm gate (park / resolveConfirm / confirmTTL)
// and the enabled-tool allowlist. Both live entirely inside
// actionAct/handleAct in actions_act.go, exactly where they lived in the old
// switch's IntentAct case — they are act-specific (a fact or a note
// can't be destructive), not shared across intents, so they do not need
// to move to a separate layer. The important invariant, preserved
// as-is: applyAction runs identically whether dec came from a fresh
// route or from a completed clarify answer (see finishClarified in
// clarify.go and its comment "filling in an argument never grants
// authority") — a handler must never special-case a clarify-completed
// decision to skip the confirm gate or the allowlist.
// - detectPattern and dialogue-session bookkeeping (rememberTurn,
// followUpMerge) run in the callers (runTurn,
// finishClarified), not per-intent, and are untouched by this slice.
//
// Each handler lives in actions_<intent>.go; the small ones (chat, system)
// and the table itself stay here.
//
// Adding an intent: write its handler in its own file, add one line to
// actionHandlers. Do not grow applyAction's switch back.
package main
import (
"context"
"log"
"github.com/kami/maven/internal/router"
)
// actionHandlers is the per-intent dispatch table used by applyAction.
var actionHandlers = map[router.Intent]func(*reactiveHandler, context.Context, router.Decision) string{
router.IntentFact: (*reactiveHandler).actionFact,
router.IntentReminder: (*reactiveHandler).actionReminder,
router.IntentAct: (*reactiveHandler).actionAct,
router.IntentChat: (*reactiveHandler).actionChat,
router.IntentSystem: (*reactiveHandler).actionSystem,
router.IntentNote: (*reactiveHandler).actionNote,
router.IntentQuery: (*reactiveHandler).actionQuery,
}
func (h *reactiveHandler) actionChat(ctx context.Context, dec router.Decision) string {
// Conversational: build history from dialogue session (prior user turns)
// and let the LLM respond from general knowledge + context.
history := h.chatHistory()
reply, err := h.phraser.PhraseChat(ctx, dec.Utterance, history)
if err != nil {
log.Printf("voice: chat: %v", err)
return "поговорили."
}
return reply
}
func (h *reactiveHandler) actionSystem(ctx context.Context, dec router.Decision) string {
return h.replySystem(ctx, dec)
}
+73
View File
@@ -0,0 +1,73 @@
package main
import (
"context"
"errors"
"log"
"github.com/kami/maven/internal/mcp"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/tool"
)
// actionAct handles router.IntentAct: match a verb to an enabled tool, offer
// it to the ecosystems first, and run it behind the confirm gate and the
// allowlist. proposeGap and the confirm gate itself live in confirm.go.
func (h *reactiveHandler) actionAct(ctx context.Context, dec router.Decision) string {
// tool executor: run the matched fn against the enabled allowlist.
// HasFn=false ⇒ try the matcher (for LLM-routed acts where the verb
// didn't go through the stage-0 act grammar).
if !dec.Slots.HasFn && dec.Slots.Text != "" && h.matcher != nil {
if fn, args, ok := h.matcher.Match(dec.Slots.Text); ok {
dec.Slots.Fn, dec.Slots.Args, dec.Slots.HasFn = fn, args, true
}
}
// Praxis ecosystem tools: intercept before the system command executor.
if h.ecosystem != nil && h.ecosystem.praxis != nil && dec.Slots.HasFn {
if reply := h.handlePraxisAct(ctx, dec); reply != "" {
return reply
}
}
// Hexis ecosystem action: if ecosystem is configured and we have a verb
// + entity text, try to resolve the entity and execute via Hexis.
if h.ecosystem != nil && h.ecosystem.hexis != nil && dec.Slots.Text != "" {
if reply := h.handleHexisAct(ctx, dec); reply != "" {
return reply
}
}
// HasFn still false ⇒ no allowlist match: scaffold a 'proposed' tool
// the user can enable on the authed surface ("earn the right to ask").
if !dec.Slots.HasFn {
return h.proposeGap(ctx, dec)
}
out, err := h.tools.Exec(ctx, dec.Slots.Fn, dec.Slots.Args, false)
if err != nil {
switch {
case errors.Is(err, tool.ErrNeedsConfirm):
// destructive: park it and ask. The next utterance answers.
phrase := actPhrase(dec.Slots.Fn, dec.Slots.Args)
h.park(dec.Slots.Fn, dec.Slots.Args, phrase)
return "выполнить «" + phrase + "»? скажи «да» или «нет»."
case errors.Is(err, tool.ErrNotEnabled):
return h.proposeGap(ctx, dec)
case errors.Is(err, mcp.ErrNeedsArgs):
// An MCP tool that wants named arguments a spoken verb cannot
// supply. Guessing them would be a wrong act, so she says so
// instead — the tool is still runnable from the authed surface,
// where a human types them.
return "этому инструменту нужны аргументы, которые я из голоса не соберу — я не буду угадывать."
}
log.Printf("voice: tool %s: %v", dec.Slots.Fn, err)
if out != "" {
return "не получилось выполнить команду: " + firstLine(out)
}
return "не получилось выполнить команду."
}
if out != "" {
return "готово: " + firstLine(out)
}
return "готово."
}
+64
View File
@@ -0,0 +1,64 @@
package main
import (
"context"
"log"
"strconv"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
)
// actionFact handles router.IntentFact: persist a tapped self-fact, index
// it for recall, and let pattern detection propose a routine.
func (h *reactiveHandler) actionFact(ctx context.Context, dec router.Decision) string {
if !dec.Slots.HasKey {
return "не разобрала, что записать — попробуй иначе."
}
now := h.now()
req := ipc.WriteFactReq{
Ts: now,
Kind: "self",
Key: dec.Slots.Key,
Value: dec.Slots.Value,
Source: "tap:voice",
Confidence: 1.0,
// Subject: the key doubles as the entity-resolution candidate —
// a voice-tapped fact's key is usually the thing/person it's
// about ("espresso_machine", "kate"), so queueing it for Nexus
// resolution costs one async lookup and is a no-op (not_found)
// for the abstract self-state keys (mood, water) that aren't
// entities at all.
Subject: dec.Slots.Key,
}
factID, err := h.api.WriteFact(ctx, req)
if err != nil {
log.Printf("voice: write fact: %v", err)
return "не получилось сохранить факт."
}
// Index the fact utterance in long-term memory (best-effort, must not
// fail the fact write). Facts aren't in the notes table, so this is the
// only recall path for them — "когда я пил воду?" reads back from here.
if h.memStore != nil {
if vec, err := router.EmbedPassage(ctx, h.embedder, dec.Utterance); err != nil {
log.Printf("voice: embed fact for memory: %v", err)
} else if err := h.memStore.Insert(ctx, "fact:"+dec.Slots.Key+":"+strconv.FormatInt(now.Unix(), 10), vec, map[string]string{
"source": "voice",
"type": "fact",
"text": dec.Utterance,
"ts": strconv.FormatInt(now.Unix(), 10),
}); err != nil {
log.Printf("voice: memory insert fact: %v", err)
}
}
// Event extraction + pattern detection (best-effort, must not fail the
// fact write). If the fact describes a recognizable action, it becomes a
// normalized event; if ≥3 events for the same action+object show stable
// intervals, a proposed routine is created and parked for confirmation.
if h.dataStore != nil {
if phrase := h.detectPattern(ctx, factID, dec.Slots.Key, dec.Slots.Value, now); phrase != "" {
return phrase // "ты заправляешь ... напоминать?"
}
}
return "" // replier phrases the success reply
}
+78
View File
@@ -0,0 +1,78 @@
package main
import (
"context"
"log"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/zenmoney"
)
// Money questions (Vikunja #125).
//
// This is the whole read side: mavpoll holds the zenmoney token and writes
// facts(kind=env, source=poll:zenmoney); core reads them back when he asks.
// Core never sees the token, never calls zenmoney, and has no rule on these
// keys — a total is never a reason for Maven to speak first. Maven is not a
// nag, least of all about his money.
//
// Nothing here can reach the external search capability: the figures are read
// from the store and rendered locally, and his financial data is never search
// input.
// queryMoney — "сколько я потратил сегодня?", "покажи мои траты".
//
// Answers only from the latest fact the poller wrote. Three honest outcomes and
// no fourth: the figure, "the fact is old and here is its date", or "money
// tracking is not connected". It never computes, estimates or rounds a total of
// its own — an invented number about his money is the worst thing this could do.
func (h *reactiveHandler) queryMoney(ctx context.Context, t *queryTurn) (string, bool) {
window, ok := router.ParseMoneyQuery(t.dec.Utterance)
if !ok {
return "", false
}
key, phrase := zenmoney.KeySpentMonth, "в этом месяце"
if window == router.MoneyToday {
key, phrase = zenmoney.KeySpentToday, "сегодня"
}
fact, err := h.api.LatestFactBySource(ctx, key, zenmoney.Source)
if err != nil {
// No fact at all is the normal state when the capability is off. Claim
// the turn anyway: falling through to recall would answer a question
// about money with whatever note happens to be nearest.
if !isNoFactErr(err) {
log.Printf("voice: money fact: %v", err)
}
return "я не отслеживаю траты — не подключено.", true
}
val, err := zenmoney.ParseFactValue(fact.Value)
if err != nil {
log.Printf("voice: money fact: decode: %v", err)
return "не получилось прочитать траты.", true
}
reply := val.FormatRU(phrase)
if reply == "" {
return "по тратам пока нечего сказать.", true
}
// A stale fact is reported as stale rather than spoken as today's number.
if h.now().Sub(fact.Ts) > zenmoney.StaleAfter {
return "данные от " + fact.Ts.Local().Format("02.01") + ": " + reply, true
}
return reply, true
}
// isNoFactErr — ErrNoFact survives the wire wrapped, so unwrap for it.
func isNoFactErr(err error) bool {
for e := err; e != nil; {
if e == ipc.ErrNoFact {
return true
}
u, ok := e.(interface{ Unwrap() error })
if !ok {
return false
}
e = u.Unwrap()
}
return false
}
+139
View File
@@ -0,0 +1,139 @@
package main
import (
"context"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/zenmoney"
)
// moneyAPI answers only LatestFactBySource; everything else is unimplemented,
// which is the assertion that answering a money question costs no model call
// and reaches no network.
type moneyAPI struct {
ipc.UnimplementedCoreAPI
fact ipc.Fact
err error
gotKey string
gotSrc string
callCnt int
}
func (a *moneyAPI) LatestFactBySource(_ context.Context, key, source string) (ipc.Fact, error) {
a.gotKey, a.gotSrc = key, source
a.callCnt++
return a.fact, a.err
}
func moneyNow() time.Time { return time.Date(2026, 8, 15, 20, 0, 0, 0, time.UTC) }
func moneyFact(ts time.Time, val string) ipc.Fact {
return ipc.Fact{Kind: "env", Key: zenmoney.KeySpentMonth, Value: val, Source: zenmoney.Source, Ts: ts}
}
func TestQueryMoneyAnswersFromTheFact(t *testing.T) {
api := &moneyAPI{fact: moneyFact(moneyNow(), `{"spent":[{"currency":"RUB","amount":1749.5}],"count":3}`)}
h := &reactiveHandler{api: api, now: moneyNow}
reply, ok := h.queryMoney(context.Background(), &queryTurn{
dec: router.Decision{Utterance: "сколько я потратил в этом месяце?"},
})
if !ok {
t.Fatal("the money source must claim a money question")
}
if api.gotKey != zenmoney.KeySpentMonth || api.gotSrc != zenmoney.Source {
t.Errorf("read %q/%q, want the month key from the poller's source", api.gotKey, api.gotSrc)
}
if !strings.Contains(reply, "1749.5") {
t.Errorf("reply = %q, want the exact figure", reply)
}
if !strings.Contains(reply, "в этом месяце") {
t.Errorf("reply = %q, want the window named", reply)
}
}
func TestQueryMoneyPicksTodaysKey(t *testing.T) {
api := &moneyAPI{fact: moneyFact(moneyNow(), `{"spent":[{"currency":"RUB","amount":250}],"count":1}`)}
h := &reactiveHandler{api: api, now: moneyNow}
if _, ok := h.queryMoney(context.Background(), &queryTurn{
dec: router.Decision{Utterance: "сколько я потратил сегодня?"},
}); !ok {
t.Fatal("expected the source to claim it")
}
if api.gotKey != zenmoney.KeySpentToday {
t.Errorf("key = %q, want today's", api.gotKey)
}
}
// The capability is off unless configured, and then there is no fact. She says
// so instead of letting the recall pass answer a money question from a note.
func TestQueryMoneySaysNotConnected(t *testing.T) {
h := &reactiveHandler{api: &moneyAPI{err: ipc.ErrNoFact}, now: moneyNow}
reply, ok := h.queryMoney(context.Background(), &queryTurn{
dec: router.Decision{Utterance: "сколько я потратил?"},
})
if !ok {
t.Fatal("expected the source to claim it")
}
if !strings.Contains(reply, "не подключено") {
t.Errorf("reply = %q, want an honest 'not connected'", reply)
}
// No number of any kind in that answer.
for _, d := range []string{"0", "1", "2", "3", "4", "5", "6", "7", "8", "9"} {
if strings.Contains(reply, d) {
t.Errorf("reply %q contains a digit — nothing was read, so there is no figure", reply)
}
}
}
// A fact older than the staleness bound is dated rather than spoken as if it
// were current: the poller can be down, and last week's total presented as
// today's is a lie by omission.
func TestQueryMoneyDatesAStaleFact(t *testing.T) {
old := moneyNow().Add(-72 * time.Hour)
api := &moneyAPI{fact: moneyFact(old, `{"spent":[{"currency":"RUB","amount":100}],"count":1}`)}
h := &reactiveHandler{api: api, now: moneyNow}
reply, _ := h.queryMoney(context.Background(), &queryTurn{
dec: router.Decision{Utterance: "сколько я потратил?"},
})
if !strings.Contains(reply, "данные от") {
t.Errorf("reply = %q, want the stale fact dated", reply)
}
}
func TestQueryMoneyPassesOtherQuestions(t *testing.T) {
api := &moneyAPI{}
h := &reactiveHandler{api: api, now: moneyNow}
for _, u := range []string{"какая погода?", "я потратил весь день на это", "какие у меня задачи?"} {
if _, ok := h.queryMoney(context.Background(), &queryTurn{dec: router.Decision{Utterance: u}}); ok {
t.Errorf("the money source claimed %q", u)
}
}
if api.callCnt != 0 {
t.Error("a non-money question must not read the money facts")
}
}
// Money must be answered before the recall sources, or a question about
// spending gets answered by the nearest note.
func TestQuerySourcesOrderMoneyBeforeRecall(t *testing.T) {
moneyAt, notesAt := -1, -1
for i, src := range querySources {
switch src.name {
case "money":
moneyAt = i
case "notes":
notesAt = i
}
}
if moneyAt < 0 || notesAt < 0 {
t.Fatalf("sources missing: money=%d notes=%d", moneyAt, notesAt)
}
if moneyAt > notesAt {
t.Errorf("money source at %d, after notes at %d", moneyAt, notesAt)
}
}
+47
View File
@@ -0,0 +1,47 @@
package main
import (
"context"
"log"
"strconv"
"github.com/kami/maven/internal/router"
)
// actionNote handles router.IntentNote: embed the note, persist it, and
// index it for recall.
func (h *reactiveHandler) actionNote(ctx context.Context, dec router.Decision) string {
// An utterance that explicitly files a task is work, not recall, and
// belongs in the task store (Vikunja #130). Checked before the embedding
// is paid for. Everything else is a note, exactly as before.
if reply, ok := h.captureTaskFromNote(ctx, dec); ok {
return reply
}
// embed the note text with the same model the classifier uses, persist
// via CoreAPI (source=tap:voice). Semantic recall lives in `notes`, not
// facts — no predicate reads it (spec's two-memory split).
vec, err := router.EmbedPassage(ctx, h.embedder, dec.Utterance)
if err != nil {
log.Printf("voice: embed note: %v", err)
return "не получилось сохранить заметку."
}
noteTs := h.now()
noteID, err := h.api.WriteNote(ctx, noteTs, dec.Utterance, vec, "tap:voice")
if err != nil {
log.Printf("voice: write note: %v", err)
return "не получилось сохранить заметку."
}
// Insert into long-term memory (best-effort, must not fail the note write).
// text/ts in the meta make a Search hit self-describing (see bestRecall).
if h.memStore != nil {
if err := h.memStore.Insert(ctx, "note:"+strconv.FormatInt(noteID, 10), vec, map[string]string{
"source": "voice",
"type": "note",
"text": dec.Utterance,
"ts": strconv.FormatInt(noteTs.Unix(), 10),
}); err != nil {
log.Printf("voice: memory insert: %v", err)
}
}
return "" // replier phrases the "saved" reply
}
+487
View File
@@ -0,0 +1,487 @@
package main
import (
"context"
"errors"
"fmt"
"log"
"strings"
"time"
"github.com/kami/maven/internal/crawl"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/morning"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/rss"
"github.com/kami/maven/internal/weather"
)
// queryTurn is the per-turn scratch a chain of query sources shares: the
// decision being answered plus the work an earlier source already paid for
// (the query embedding, the notes it pulled). Sources read and fill it in
// order, so a later source never re-embeds.
type queryTurn struct {
dec router.Decision
vec []float32
notes []ipc.Note
}
// querySource — one answer source in the chain actionQuery walks. answer
// returns (reply, true) when this source claims the question, ("", false)
// when it passes to the next one. name is for reading the table, not logged.
//
// A struct of one func rather than an interface: every source is a plain
// method on *reactiveHandler with no state of its own (what state a turn has
// lives in queryTurn), so an interface would mean one empty type per source
// to satisfy it — ceremony for nothing. Same reasoning as confirmResolver in
// confirm.go, and the table then reads like actionHandlers: a flat list of
// method expressions you extend with one line.
type querySource struct {
name string
answer func(*reactiveHandler, context.Context, *queryTurn) (string, bool)
}
// querySources is the ordered chain actionQuery walks; first source to claim
// answers the turn. THE ORDER IS LOAD-BEARING — see the memory-before-notes
// comment on queryMemory: running the notes-only pass first was #373, and the
// gate was never the bug. Adding a source (Kiwix, RSS, crawler, email) is one
// line here plus its method; where you put the line is the whole decision.
var querySources = []querySource{
{"fact-by-key", (*reactiveHandler).queryFactByKey},
// Before "calendar" on purpose: both match "…на сегодня", and the plan is
// the more specific ask (its matcher requires a plan word), so the calendar
// listing would otherwise swallow it.
{"day-plan", (*reactiveHandler).queryDayPlan},
// Also before "calendar": "что я обычно делаю по средам?" names a weekday,
// and the habit question is the more specific one. Its matcher requires a
// habit marker ("обычно", "каждый", …), so a question about this coming
// Wednesday still reaches the calendar.
{"habits", (*reactiveHandler).queryHabits},
// Before "calendar" and before the recall sources: "что мне нужно
// сделать?" is a question about the task list, and the notes pass would
// otherwise answer it with whatever note happens to be nearest. Its
// matcher requires a task noun or an explicit "что … сделать", so a
// date-bearing question still reaches the calendar.
{"tasks", (*reactiveHandler).queryTasks},
// Before the recall sources too: "сколько я потратил?" is a question about
// the money facts the poller wrote, and the notes pass would otherwise
// answer it from whatever he once said about spending. Its matcher needs a
// money noun plus an actual ask, so "я потратил весь день" is untouched.
{"money", (*reactiveHandler).queryMoney},
// Before the recall sources and before general knowledge: "что нового?" is
// a question about the feeds she reads, and general knowledge would answer
// it by inventing news. Its matcher needs a feed noun plus an ask, so
// "у меня новая лента в инстаграме" is untouched.
{"feeds", (*reactiveHandler).queryFeeds},
// Before "calendar" and before the recall sources: "что включено дома?" is
// a question about the house, and the notes pass would otherwise answer it
// from whatever he once said about the lights. Its matcher needs a house
// marker plus an ask plus a device word, and it bails out on weather
// wording, so "какая температура на улице?" still reaches the weather
// source.
{"home", (*reactiveHandler).queryHome},
// Next to "home" and for the same reason: "какие устройства в сети?" is a
// question about the LAN, and the recall pass would otherwise answer it
// from an old note about the router. Its matcher needs a network word plus
// an ask plus a device noun, so "интернет не работает" is untouched.
{"network", (*reactiveHandler).queryNetwork},
{"calendar", (*reactiveHandler).queryCalendar},
{"weather", (*reactiveHandler).queryWeather},
{"embed", (*reactiveHandler).queryEmbed},
{"memory", (*reactiveHandler).queryMemory},
{"notes", (*reactiveHandler).queryNotes},
// LAST before the model answers from memory, and that position is the whole
// design (Vikunja #259): local sources first. The model, his own notes and
// facts, and — once internal/kiwix is wired into this chain — the offline
// ZIMs all get their turn before anything touches the network. This source
// only claims a turn where he named a URL out loud, so it never competes
// with a local answer.
{"web", (*reactiveHandler).queryWeb},
{"general-knowledge", (*reactiveHandler).queryGeneral},
}
func (h *reactiveHandler) actionQuery(ctx context.Context, dec router.Decision) string {
t := &queryTurn{dec: dec}
for _, src := range querySources {
if reply, ok := src.answer(h, ctx, t); ok {
return reply
}
}
return "не знаю."
}
// queryFactByKey — when the dialogue layer resolved an anaphoric reference to
// a prior fact's key (e.g. "когда я это сделал?" after "запиши что я пил
// воду"), look up the fact's value directly.
func (h *reactiveHandler) queryFactByKey(ctx context.Context, t *queryTurn) (string, bool) {
dec := t.dec
if !dec.Slots.HasKey || dec.Slots.Key == "" {
return "", false
}
f, err := h.api.LatestFact(ctx, dec.Slots.Key)
if err != nil {
return "", false
}
if dec.Slots.HasTime {
// The query asks about timing — the fact's own timestamp is the
// answer it's looking for. Format as a natural reply.
return fmt.Sprintf("я записала это %s", formatTime(f.Ts)), true
}
// General fact reference: describe what we know.
if dec.Utterance == "" {
return fmt.Sprintf("вот что я знаю: %s — %s", dec.Slots.Key, f.Value), true
}
// The utterance still carries the question; fall through to normal RAG
// with the resolved key in context.
return "", false
}
// queryDayPlan — "какие планы на сегодня?", "что у меня по плану?", "что
// дальше?" (Vikunja #128). Recites the day: calendar events, pending
// reminders, and any morning checklist still outstanding.
//
// Read-only by construction — the plan is assembled and rendered core-side and
// nothing here schedules or announces. "что дальше?" asks for the rest of the
// day, so that phrasing trims what has already passed.
func (h *reactiveHandler) queryDayPlan(ctx context.Context, t *queryTurn) (string, bool) {
if !router.IsDayPlanQuery(t.dec.Utterance) {
return "", false
}
plan, err := h.api.DayPlan(ctx)
if err != nil {
log.Printf("voice: day plan: %v", err)
return "не получилось собрать план.", true
}
if !isRestOfDayQuery(t.dec.Utterance) {
return plan.Spoken, true
}
// Rebuild the pure plan so the rest-of-day rendering is the same code that
// rendered the whole day — one formatter, one persona.
p := morning.Plan{Date: plan.Date}
for _, it := range plan.Items {
p.Items = append(p.Items, morning.PlanEntry{
At: it.At,
Text: it.Text,
Kind: morning.PlanKind(it.Kind),
Uncertain: it.Uncertain,
})
}
return p.After(h.now()).FormatRU(), true
}
// isRestOfDayQuery — "что дальше?" and its English form, the only plan phrasing
// that means "from now on" rather than "the whole day".
func isRestOfDayQuery(text string) bool {
s := strings.ToLower(text)
return strings.Contains(s, "дальше") || strings.Contains(s, "next")
}
// habitFactWindow — how many recent facts the behaviour profile is counted
// over. Enough for a season of habits without scanning the whole store on every
// question; the profile is recomputed on read, so the bound is the cost control.
const habitFactWindow = 2000
// queryHabits — "что я обычно делаю по вторникам?" (Vikunja #254). Counts the
// answer out of the fact log rather than asking the model to summarise a life:
// see internal/memory/behavior.go for why nothing here is generated.
func (h *reactiveHandler) queryHabits(ctx context.Context, t *queryTurn) (string, bool) {
q, ok := router.ParseHabitQuery(t.dec.Utterance)
if !ok {
return "", false
}
facts, err := h.api.RecentFacts(ctx, habitFactWindow)
if err != nil {
log.Printf("voice: habits: recent facts: %v", err)
return "не получилось посмотреть записи.", true
}
obs := make([]memory.Observation, 0, len(facts))
for _, f := range facts {
obs = append(obs, memory.Observation{At: f.Ts, Key: f.Key, Kind: f.Kind})
}
profile := memory.BuildProfile(obs, h.now())
if q.HasWeekday {
return profile.FormatWeekdayRU(q.Weekday), true
}
return profile.FormatOverallRU(), true
}
// feedNoteWindow — how many recent notes are scanned for feed items, and
// feedReadOut — how many headlines she actually reads back. She summarises the
// top of the pile, she does not recite a river.
const (
feedNoteWindow = 200
feedReadOut = 3
)
// queryFeeds — "что нового в лентах?", "что нового по технологиям?"
// (Vikunja #258).
//
// This is the ONLY way a feed item reaches him. The poller writes notes and
// never speaks; asking is the trigger. If that ever changes, the thing that
// changed is "Maven is not a nag", not a detail of this file.
func (h *reactiveHandler) queryFeeds(ctx context.Context, t *queryTurn) (string, bool) {
q, ok := router.ParseFeedQuery(t.dec.Utterance)
if !ok {
return "", false
}
if !h.feedsOn {
// Claim the turn rather than fall through: "не читаю ленты" is true, and
// letting general knowledge answer "что нового?" would be an invented
// news bulletin.
return "я пока не читаю ленты — они не настроены.", true
}
notes, err := h.api.RecentNotes(ctx, feedNoteWindow)
if err != nil {
log.Printf("voice: feeds: recent notes: %v", err)
return "не получилось посмотреть ленты.", true
}
var picked []string
for _, n := range notes {
if !strings.HasPrefix(n.Source, rss.SourcePrefix) {
continue
}
if !router.CategoryMatches(n.Text, q.Category) {
continue
}
// The note carries title, summary and link; she reads the title.
title := n.Text
if i := strings.IndexByte(title, '\n'); i > 0 {
title = title[:i]
}
picked = append(picked, strings.TrimSpace(title))
if len(picked) == feedReadOut {
break
}
}
if len(picked) == 0 {
if q.Category != "" {
return "по этой теме в лентах пока ничего.", true
}
return "в лентах пока ничего нового.", true
}
return "вот что нового: " + strings.Join(picked, "; "), true
}
// queryCalendar — "что у меня сегодня?", "планы на завтра?"
// h.now(), not time.Now(): the handler's clock is the injected one, so this
// source can be tested at a fixed time like the rest.
func (h *reactiveHandler) queryCalendar(ctx context.Context, t *queryTurn) (string, bool) {
date, ok := router.ParseCalendarDate(t.dec.Utterance, h.now())
if !ok {
return "", false
}
events, err := h.api.CalendarEvents(ctx, date, date.Add(24*time.Hour))
if err != nil {
log.Printf("voice: calendar events: %v", err)
return "не получилось проверить календарь.", true
}
// Provenance travels with each event. A work meeting relayed off a phone
// notification (source ambient:notif, #126) is stored below full confidence
// and gets hedged; a CalDAV read is recited plainly.
entries := make([]router.CalendarEntry, len(events))
for i, e := range events {
entries[i] = router.CalendarEntry{Text: e.Value, Uncertain: e.Confidence < 1.0}
}
var f router.CalendarEventFormatter
return f.FormatEntries(entries, date), true
}
// queryHome answers a question about the house. Read-only by construction: it
// calls States and nothing else, so there is no confirm turn here — the only
// way to CHANGE something is an enabled allowlist row through tool.Executor.
func (h *reactiveHandler) queryHome(ctx context.Context, t *queryTurn) (string, bool) {
if !isHomeQuery(t.dec.Utterance) {
return "", false
}
if h.home == nil {
// Claim the turn rather than fall through: "дом не подключён" is true,
// and letting general knowledge answer would be an invented house.
return "дом не подключён — я его не вижу.", true
}
ctxH, cancel := context.WithTimeout(ctx, 10*time.Second)
defer cancel()
return h.home.homeSummary(ctxH)
}
// queryNetwork answers a question about the LAN with a bounded scan. There is
// no confirm turn because nothing is changed, and no way to widen the range
// because Scan takes no target — the utterance selects the question, never the
// subnet.
func (h *reactiveHandler) queryNetwork(ctx context.Context, t *queryTurn) (string, bool) {
if !isNetworkQuery(t.dec.Utterance) {
return "", false
}
if h.netscan == nil {
// Claim the turn: "сканирование не настроено" is true, and general
// knowledge would answer with an invented list of devices.
return "сканирование сети не настроено.", true
}
return h.netscan.scanSummary(ctx)
}
func (h *reactiveHandler) queryWeather(ctx context.Context, t *queryTurn) (string, bool) {
if !isWeatherQuery(t.dec.Utterance) {
return "", false
}
loc := extractWeatherLocation(t.dec.Utterance, h.weatherLocation)
ctxWT, cancel := context.WithTimeout(ctx, 5*time.Second)
defer cancel()
w, err := h.weatherProvider.CurrentWeather(ctxWT, loc)
if errors.Is(err, weather.ErrNotConfigured) {
return "погода не настроена.", true
}
if err != nil {
log.Printf("voice: weather: %v", err)
return "не получилось узнать погоду.", true
}
return fmt.Sprintf("в %s сейчас %.0f градусов, %s.", w.Location, w.Temperature, w.Condition), true
}
// queryEmbed isn't an answer source — it's the shared cost the two recall
// sources below both need, run once, in the position it always ran in. It
// only claims the turn when the embedder fails.
func (h *reactiveHandler) queryEmbed(ctx context.Context, t *queryTurn) (string, bool) {
vec, err := router.EmbedQuery(ctx, h.embedder, t.dec.Utterance)
if err != nil {
log.Printf("voice: embed query: %v", err)
return "не получилось найти ответ.", true
}
t.vec = vec
return "", false
}
// queryMemory — long-term memory first: ONE search over everything Maven
// remembers (notes and facts share this index) and ONE confidence gate, so
// the memory that is clearly the best match answers — a note just as much as
// a fact.
//
// This used to run only after the notes-only source below had already
// rejected the same note at the same score, which no note could ever survive
// a second time: the branch could only return a fact (#373). Order, not the
// gate, was the bug — the set of questions Maven answers is unchanged, only
// which memory gets to answer them.
func (h *reactiveHandler) queryMemory(ctx context.Context, t *queryTurn) (string, bool) {
if h.memStore == nil {
return "", false
}
hits, herr := h.memStore.Search(ctx, t.vec, 3)
if herr != nil {
log.Printf("voice: memory search: %v", herr)
return "", false
}
hit, ok := bestRecall(hits, h.queryMinScore, h.queryMinMargin)
if !ok {
return "", false
}
text := hit.Meta["text"]
// A note is phrased in Maven's voice; a fact is read back as it was
// stored.
if hit.Meta["type"] == "note" {
if reply, perr := h.phraser.PhraseQuery(ctx, t.dec.Utterance, []string{text}); perr == nil && reply != "" {
return reply, true
}
}
return text, true
}
// queryNotes — notes-only pass, for notes the vector index above does not
// hold (an older note written before it existed). Same gate, notes-only
// candidates.
//
// Confidence gate: below it, say "I don't know" rather than read back the
// least-unrelated note — a confident wrong recall is worse than a gap (spec's
// "not a guesser-of-truth"). Same instinct as the loop's since(key)==null →
// don't fire. Two parts: an absolute cosine floor, and a margin over the
// runner-up, which is the part that works with the e5 embedder's narrow score
// band. See memory.Confident. Failing the gate passes the turn on to general
// knowledge, which is what "don't read back the runner-up" means here.
func (h *reactiveHandler) queryNotes(ctx context.Context, t *queryTurn) (string, bool) {
notes, err := h.api.QueryNotes(ctx, t.vec, 5)
if err != nil {
log.Printf("voice: query notes: %v", err)
return "не получилось найти ответ.", true
}
t.notes = notes
noteScores := make([]float64, len(notes))
for i, n := range notes {
noteScores[i] = n.Score
}
if !memory.ConfidentScores(noteScores, h.queryMinScore, h.queryMinMargin) {
return "", false
}
texts := make([]string, len(notes))
for i, n := range notes {
texts[i] = n.Text
}
reply, err := h.phraser.PhraseQuery(ctx, t.dec.Utterance, texts)
if err != nil {
log.Printf("voice: phrase query: %v", err)
}
if reply == "" {
reply = "вот что я нашла: " + texts[0]
}
return reply, true
}
// webPageContextRunes — how much of a fetched page is handed to the phraser.
// Less than the crawler keeps: the rest of the 4096-token window belongs to the
// prompt, the persona block and the reply.
const webPageContextRunes = 1500
// queryWeb — "посмотри https://example.org/x — что там?" (Vikunja #259).
//
// It claims a turn ONLY when he named a URL, which is what keeps a fallback from
// becoming a habit: no URL, no fetch, and the model answers from what is local.
// What leaves the box is the URL and nothing else — no note, no fact, no history
// travels with it.
func (h *reactiveHandler) queryWeb(ctx context.Context, t *queryTurn) (string, bool) {
link, ok := router.FirstURL(t.dec.Utterance)
if !ok {
return "", false
}
if h.crawler == nil {
// Claim rather than fall through: he asked about a specific page, and
// letting the model answer from the URL's spelling alone is how a small
// model invents a page's contents.
return "я не читаю страницы — это не настроено.", true
}
ctxFetch, cancel := context.WithTimeout(ctx, 30*time.Second)
defer cancel()
page, err := h.crawler.Page(ctxFetch, link)
if err != nil {
if errors.Is(err, crawl.ErrRobots) {
return "эта страница закрыта для чтения — robots.txt не разрешает.", true
}
log.Printf("voice: web: %v", err)
return "не получилось прочитать страницу.", true
}
if page.Text == "" {
return "страница открылась, но читать там нечего.", true
}
// The page is handed to the phraser the same way a note is: as context for
// the question he actually asked. She answers the question, she does not
// recite the page.
snippet := page.Title + "\n" + crawl.TrimRunes(page.Text, webPageContextRunes)
reply, perr := h.phraser.PhraseQuery(ctx, t.dec.Utterance, []string{snippet})
if perr != nil {
log.Printf("voice: web: phrase: %v", perr)
}
if reply == "" {
// No phraser (or it failed): read back the top of the page rather than
// pretend the fetch did not happen.
return "вот что на странице: " + crawl.TrimRunes(page.Text, 300), true
}
return reply, true
}
// queryGeneral — general knowledge from the phraser, the last source before
// giving up. It always claims: either the model answers or Maven says she
// doesn't know.
func (h *reactiveHandler) queryGeneral(ctx context.Context, t *queryTurn) (string, bool) {
reply, err := h.phraser.PhraseQuery(ctx, t.dec.Utterance, nil)
if err != nil || reply == "" {
return "не знаю.", true
}
return reply, true
}
+33
View File
@@ -0,0 +1,33 @@
package main
import (
"context"
"log"
"github.com/kami/maven/internal/router"
)
// actionReminder handles router.IntentReminder: parse the time when stage-0
// skipped the extractor, then create the reminder.
func (h *reactiveHandler) actionReminder(ctx context.Context, dec router.Decision) string {
if !dec.Slots.HasTime {
// Stage-0 (reminder-wakeword grammar) skips the extractor, so the
// time wasn't parsed. Run the parser as a fallback.
if dec.Stage == 0 && h.timeParser != nil {
t, ok, err := h.timeParser.Parse(ctx, dec.Utterance, h.now())
if err == nil && ok {
dec.Slots.Time = t
dec.Slots.HasTime = true
}
}
if !dec.Slots.HasTime {
return "не получилось разобрать время напоминания."
}
}
payload := `{"text":` + jsonString(dec.Utterance) + `}`
if _, err := h.api.CreateReminder(ctx, dec.Slots.Time, payload, ""); err != nil {
log.Printf("voice: create reminder: %v", err)
return "не получилось поставить напоминание."
}
return ""
}
+83
View File
@@ -0,0 +1,83 @@
package main
import (
"context"
"log"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/tasks"
)
// Task capture on the voice/chat path (Vikunja #130).
//
// Two halves, both deliberately small:
//
// - captureTaskFromNote runs at the top of actionNote. An utterance that
// explicitly files a task ("добавь в задачи купить молоко") goes to the task
// store instead of the note store. Anything without an explicit marker is
// still a note — see router.ParseTaskCapture for why "надо бы поспать" must
// not become a task.
// - queryTasks is a query source that reads the list back.
//
// Nothing here speaks unprompted. Tasks are answered when asked about; no tick
// rule reads the table.
// captureTaskFromNote claims the turn when the utterance explicitly files a
// task, returning the reply. ("", false) hands the turn back to the note path.
func (h *reactiveHandler) captureTaskFromNote(ctx context.Context, dec router.Decision) (string, bool) {
cap, ok := router.ParseTaskCapture(dec.Utterance)
if !ok {
return "", false
}
resp, err := h.api.CaptureTask(ctx, ipc.CaptureTaskReq{
Text: cap.Text,
Source: "tap:voice",
Status: store.TaskOpen, // he stated it himself — not a candidate
Weight: cap.Weight, // 0 unless he said "срочно" / "важно"
Ts: h.now(),
})
if err != nil {
log.Printf("voice: capture task: %v", err)
return "не получилось записать задачу.", true
}
if !resp.Created {
return "это уже в списке.", true
}
return "записала: " + cap.Text, true
}
// queryTasks — "какие у меня задачи?", "что мне нужно сделать?".
//
// Reads the live set and recites it in priority order (Vikunja #129). The order
// is computed by internal/tasks from what he told her — deadlines, the urgency
// he stated, how long a task has been sitting — never asked of the model. The
// rendering is the package's too, so the spoken list and the /tasks page can
// never disagree about what comes first.
func (h *reactiveHandler) queryTasks(ctx context.Context, t *queryTurn) (string, bool) {
if !router.IsTaskListQuery(t.dec.Utterance) {
return "", false
}
live, err := h.api.ListTasks(ctx, "live")
if err != nil {
log.Printf("voice: list tasks: %v", err)
return "не получилось посмотреть задачи.", true
}
return tasks.FormatRU(tasks.Rank(taskItems(live), h.now())), true
}
// taskItems maps wire rows onto the ranker's input. Written here rather than in
// internal/tasks so the ranker stays a pure package with no ipc (and therefore
// no store, and therefore no cgo) dependency — the same posture as
// internal/morning and internal/memory.
func taskItems(ts []ipc.Task) []tasks.Item {
out := make([]tasks.Item, len(ts))
for i, t := range ts {
out[i] = tasks.Item{
ID: t.ID, Text: t.Text, Status: t.Status,
Created: t.CreatedTs, Due: t.Due, Weight: t.Weight,
}
}
return out
}
+227
View File
@@ -0,0 +1,227 @@
package main
import (
"context"
"errors"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
)
// taskAPI answers only the three task methods; every other call is
// unimplemented, which is the assertion that capture needs nothing else — in
// particular no embedder, so a filed task costs no model call.
type taskAPI struct {
ipc.UnimplementedCoreAPI
captured []ipc.CaptureTaskReq
created bool
capErr error
tasks []ipc.Task
listArg string
listErr error
}
func (a *taskAPI) CaptureTask(_ context.Context, req ipc.CaptureTaskReq) (ipc.CaptureTaskResp, error) {
a.captured = append(a.captured, req)
if a.capErr != nil {
return ipc.CaptureTaskResp{}, a.capErr
}
return ipc.CaptureTaskResp{ID: 1, Created: a.created}, nil
}
func (a *taskAPI) ListTasks(_ context.Context, status string) ([]ipc.Task, error) {
a.listArg = status
return a.tasks, a.listErr
}
func taskNow() time.Time { return time.Date(2026, 8, 1, 9, 0, 0, 0, time.UTC) }
func taskHandler(api ipc.CoreAPI) *reactiveHandler {
return &reactiveHandler{api: api, now: taskNow}
}
func TestCaptureTaskFromNoteFilesTheTask(t *testing.T) {
api := &taskAPI{created: true}
h := taskHandler(api)
reply, ok := h.captureTaskFromNote(context.Background(), router.Decision{
Intent: router.IntentNote, Utterance: "добавь в задачи купить молоко",
})
if !ok {
t.Fatal("an explicit capture must claim the turn")
}
if len(api.captured) != 1 {
t.Fatalf("captured %d, want 1", len(api.captured))
}
got := api.captured[0]
if got.Text != "купить молоко" {
t.Errorf("text = %q, want the marker stripped", got.Text)
}
if got.Source != "tap:voice" {
t.Errorf("source = %q, want tap:voice", got.Source)
}
if got.Status != "open" {
t.Errorf("status = %q — work he stated is open, never a candidate", got.Status)
}
if !got.Ts.Equal(taskNow()) {
t.Errorf("ts = %v, want the handler clock", got.Ts)
}
if !strings.Contains(reply, "купить молоко") {
t.Errorf("reply = %q, want it to read the task back", reply)
}
}
// A note is still a note: capture only fires on an explicit marker, so
// ordinary recall is untouched.
func TestCaptureTaskFromNotePassesOrdinaryNotes(t *testing.T) {
api := &taskAPI{}
h := taskHandler(api)
for _, u := range []string{"надо бы поспать", "мне понравился этот фильм", "запиши что я пил воду"} {
if _, ok := h.captureTaskFromNote(context.Background(), router.Decision{Utterance: u}); ok {
t.Errorf("%q was captured as a task", u)
}
}
if len(api.captured) != 0 {
t.Errorf("captured %d requests, want none", len(api.captured))
}
}
func TestCaptureTaskFromNoteSaysAlreadyOnTheList(t *testing.T) {
h := taskHandler(&taskAPI{created: false})
reply, ok := h.captureTaskFromNote(context.Background(), router.Decision{Utterance: "добавь в задачи купить молоко"})
if !ok {
t.Fatal("expected the capture path to claim it")
}
if !strings.Contains(reply, "уже") {
t.Errorf("reply = %q — a deduped capture must not claim it saved something new", reply)
}
}
func TestCaptureTaskFromNoteReportsStoreFailure(t *testing.T) {
h := taskHandler(&taskAPI{capErr: errors.New("db is on fire")})
reply, ok := h.captureTaskFromNote(context.Background(), router.Decision{Utterance: "добавь задачу починить кран"})
if !ok {
t.Fatal("a failed capture still claims the turn — the note path must not double-write")
}
if !strings.Contains(reply, "не получилось") {
t.Errorf("reply = %q, want an honest failure", reply)
}
}
func TestQueryTasksRecitesTheLiveList(t *testing.T) {
api := &taskAPI{tasks: []ipc.Task{
{ID: 1, Text: "купить молоко", Status: "open"},
{ID: 2, Text: "продлить страховку", Status: "candidate"},
}}
h := taskHandler(api)
reply, ok := h.queryTasks(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "какие у меня задачи?"},
})
if !ok {
t.Fatal("the task source must claim a task-list question")
}
if api.listArg != "live" {
t.Errorf("ListTasks(%q), want \"live\" — a resolved task is not outstanding work", api.listArg)
}
if !strings.Contains(reply, "купить молоко") || !strings.Contains(reply, "продлить страховку") {
t.Errorf("reply = %q, want both tasks", reply)
}
// The candidate must be named as unconfirmed, not recited as his work.
openIdx := strings.Index(reply, "купить молоко")
candIdx := strings.Index(reply, "продлить страховку")
if !(openIdx < candIdx) {
t.Errorf("reply = %q, want confirmed work before candidates", reply)
}
if !strings.Contains(reply, "не подтвердил") {
t.Errorf("reply = %q, want the candidate flagged as unconfirmed", reply)
}
}
// The stated urgency rides through capture as a weight, so the ranker can use
// it later (Vikunja #129). "срочно" is not part of the task text.
func TestCaptureTaskCarriesStatedUrgency(t *testing.T) {
api := &taskAPI{created: true}
h := taskHandler(api)
if _, ok := h.captureTaskFromNote(context.Background(), router.Decision{
Utterance: "добавь в задачи срочно оплатить интернет",
}); !ok {
t.Fatal("expected a capture")
}
got := api.captured[0]
if got.Text != "оплатить интернет" {
t.Errorf("text = %q, want the urgency word out of the task", got.Text)
}
if got.Weight == 0 {
t.Error("weight = 0 — he said срочно and it was dropped")
}
}
// The recital is ordered by the ranker, not by insertion: a deadline he named
// comes before undated work.
func TestQueryTasksRecitesInPriorityOrder(t *testing.T) {
due := taskNow()
api := &taskAPI{tasks: []ipc.Task{
{ID: 1, Text: "купить молоко", Status: "open", CreatedTs: taskNow()},
{ID: 2, Text: "оплатить интернет", Status: "open", CreatedTs: taskNow(), Due: &due},
}}
h := taskHandler(api)
reply, _ := h.queryTasks(context.Background(), &queryTurn{
dec: router.Decision{Utterance: "какие у меня задачи?"},
})
if strings.Index(reply, "оплатить интернет") > strings.Index(reply, "купить молоко") {
t.Errorf("reply = %q, want the dated task first", reply)
}
if !strings.Contains(reply, "сегодня") {
t.Errorf("reply = %q, want the reason named", reply)
}
}
func TestQueryTasksEmptyList(t *testing.T) {
h := taskHandler(&taskAPI{})
reply, ok := h.queryTasks(context.Background(), &queryTurn{
dec: router.Decision{Utterance: "что мне нужно сделать?"},
})
if !ok {
t.Fatal("expected the task source to claim it")
}
if reply != "задач нет." {
t.Errorf("reply = %q", reply)
}
}
func TestQueryTasksPassesOtherQuestions(t *testing.T) {
api := &taskAPI{}
h := taskHandler(api)
for _, u := range []string{"как дела?", "какая погода в москве?", "что у меня сегодня?"} {
if _, ok := h.queryTasks(context.Background(), &queryTurn{dec: router.Decision{Utterance: u}}); ok {
t.Errorf("the task source claimed %q", u)
}
}
if api.listArg != "" {
t.Error("a non-task question must not read the task list")
}
}
// The chain must reach the task source before the recall sources, or "что мне
// нужно сделать?" gets answered by whatever note is nearest.
func TestQuerySourcesOrderTasksBeforeRecall(t *testing.T) {
var tasksAt, notesAt = -1, -1
for i, src := range querySources {
switch src.name {
case "tasks":
tasksAt = i
case "notes":
notesAt = i
}
}
if tasksAt < 0 || notesAt < 0 {
t.Fatalf("sources missing: tasks=%d notes=%d", tasksAt, notesAt)
}
if tasksAt > notesAt {
t.Errorf("tasks source at %d, after notes at %d", tasksAt, notesAt)
}
}
+263
View File
@@ -0,0 +1,263 @@
// mavend/capture.go — core's half of the meeting recorder (Vikunja #253,
// docs/plans/08-hearing.md).
//
// The split: a client that has a microphone (mavenclient, or a phone on the PWA)
// is told to start, streams frames over ipc.MethodCaptureAppend, and is told to
// stop. Core keeps the PCM, stores it as a WAV blob under the same media store
// and the same retention as images, transcribes it through the ONE STT Maven has
// (mavsttd's whisper.cpp, reused — not a second engine), and summarises the
// transcript on the resident model in windows that fit n_ctx 4096.
//
// # Off unless configured, twice over
//
// No `media` block ⇒ nowhere to keep audio ⇒ the four capture methods do not
// exist. No `capture` block with enabled ⇒ they still do not exist. On an
// unconfigured box there is no wire path that starts a recording, which is the
// only guarantee worth making about a capability like this one.
//
// # What this file refuses to do
//
// - Nothing listens. There is no VAD hook here, no wake-word branch, no
// "start when you hear a meeting". The plan document's keyword-triggered
// recorder is refused in internal/capture's package comment for the reason
// that applies here too: noticing a keyword requires listening, which is
// the behaviour this capability must not have.
// - No transcript note by default. The summary is written where he will read
// it; the verbatim record of what other people said takes a deliberate
// capture.save_transcript.
// - The transcript is never search input beyond this box, and the audio never
// leaves it at all.
package main
import (
"context"
"errors"
"fmt"
"log"
"time"
"github.com/kami/maven/internal/capture"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
)
// captureSummaryTimeout — the budget for one Stop, which is a map-reduce over
// the whole meeting: one model call per transcript window plus a reduce, each of
// which is seconds on this box. Forty windows is the configured ceiling, so the
// budget has to be minutes, not the 60s the reply path uses.
const captureSummaryTimeout = 20 * time.Minute
// llmCompleter adapts *llm.Client to capture.Completer. The pure package names
// the two strings it needs and stays free of the llm request struct; the client
// itself is the swap-aware one from llmClientFor, so a model swap re-points it.
type llmCompleter struct {
c *llm.Client
maxTokens int
}
func (l llmCompleter) Complete(ctx context.Context, system, user string) (string, error) {
return l.c.Complete(ctx, llm.Req{System: system, User: user, MaxTokens: l.maxTokens})
}
// captureWiring — the recorder plus what it needs to write the result down.
type captureWiring struct {
rec *capture.Recorder
st *store.Store
emb router.Embedder
cfg *config.CaptureConfig
now func() time.Time
}
// newCaptureWiring returns nil when the recorder should not exist: no media
// store, no capture block, capture disabled, or no STT to transcribe with.
//
// A missing llama-server is NOT a reason to return nil. Without one the
// recording is still made, stored and transcribed, and the summary is simply
// absent — the honest degradation, and much better than refusing to record a
// meeting that is happening now.
func newCaptureWiring(keeper *mediaKeeper, st *store.Store, voiceW *voiceWiring, phr phraser.Phraser, emb router.Embedder, cfg *config.Config) *captureWiring {
if keeper == nil || !cfg.Capture.Records() {
return nil
}
tr := transcriberOf(voiceW)
if tr == nil {
// Voice off ⇒ no STT client ⇒ nothing could turn the audio into words.
// Storing hours of unreadable audio of other people is worse than not
// recording, so this is a refusal, not a degradation.
log.Printf("capture: enabled but voice/stt is not wired — meeting capture disabled")
return nil
}
cc := cfg.Capture
var sum *capture.Summarizer
if lp, ok := phr.(*phraser.LLMPhraser); ok {
client := llmClientFor(lp, captureSummaryTimeout)
sum = capture.NewSummarizer(
llmCompleter{c: client, maxTokens: 512},
cc.ChunkRunes, cc.MaxChunks, contextBlockFn(cfg, time.Now),
)
} else {
log.Printf("capture: no llama-server phraser — meetings are transcribed, not summarised")
}
rec, err := capture.New(keeper.store, tr, sum, capture.Config{
MaxDuration: cc.MaxDuration(),
STTWindow: time.Duration(cc.STTWindow),
})
if err != nil {
log.Printf("capture: %v — meeting capture disabled", err)
return nil
}
log.Printf("capture: enabled, sessions capped at %s", rec.MaxDuration())
return &captureWiring{rec: rec, st: st, emb: emb, cfg: cc, now: time.Now}
}
// start handles ipc.MethodCaptureStart.
func (c *captureWiring) start(_ context.Context, req ipc.CaptureStartReq) (ipc.CaptureStartResp, error) {
s, err := c.rec.Start(req.Label)
if err != nil {
return ipc.CaptureStartResp{}, err
}
// The label is logged; nothing that was said ever is.
log.Printf("capture: started %q", s.Label)
return ipc.CaptureStartResp{
Label: s.Label,
Started: s.Started,
MaxSeconds: int(c.rec.MaxDuration().Seconds()),
}, nil
}
// append handles ipc.MethodCaptureAppend. ErrExpired is reported as a successful
// response with Expired set rather than an error: the cap firing is the designed
// behaviour, and the client needs the flag to stop sending and call stop.
func (c *captureWiring) append(_ context.Context, req ipc.CaptureAppendReq) (ipc.CaptureAppendResp, error) {
err := c.rec.Append(req.Audio)
st := c.rec.Status()
if errors.Is(err, capture.ErrExpired) {
log.Printf("capture: %q hit the %s cap — stopping", st.Label, c.rec.MaxDuration())
return ipc.CaptureAppendResp{Seconds: st.Duration.Seconds(), Expired: true}, nil
}
if err != nil {
return ipc.CaptureAppendResp{}, err
}
return ipc.CaptureAppendResp{Seconds: st.Duration.Seconds()}, nil
}
// stop handles ipc.MethodCaptureStop.
//
// The error handling here mirrors vision's, and for the same reason: the audio is
// stored first, so a transcription or summary failure returns what exists rather
// than nothing. A response can carry a blob id with no transcript (STT failed,
// re-runnable), or a transcript with no summary (the model failed, the words are
// kept) — both are degraded successes and neither is an error to the caller.
func (c *captureWiring) stop(ctx context.Context, req ipc.CaptureStopReq) (ipc.CaptureStopResp, error) {
if req.Discard {
// "забудь, не записывай" — nothing is stored, transcribed or noted.
if !c.rec.Abort() {
return ipc.CaptureStopResp{}, capture.ErrNoSession
}
log.Printf("capture: session discarded on request")
return ipc.CaptureStopResp{Discarded: true}, nil
}
res, err := c.rec.Stop(ctx)
resp := ipc.CaptureStopResp{
BlobID: res.BlobID,
Label: res.Label,
Started: res.Started,
Seconds: res.Duration.Seconds(),
Transcript: res.Transcript,
Summary: res.Summary,
Chunks: res.Chunks,
}
if err != nil {
if res.BlobID == "" && res.Transcript == "" {
// Nothing survived: no session, or an empty recording. There is
// nothing to hand back, so this is a real error.
return ipc.CaptureStopResp{}, err
}
log.Printf("capture: %q partially finished: %v", res.Label, err)
}
if id, werr := c.writeNotes(ctx, res); werr != nil {
log.Printf("capture: note write for %q failed: %v", res.Label, werr)
} else {
resp.NoteID = id
}
log.Printf("capture: finished %q — %s of audio, %d summary chunk(s)",
res.Label, res.Duration.Round(time.Second), res.Chunks)
return resp, nil
}
// writeNotes stores the summary as a note, and the transcript too when
// capture.save_transcript is set. Returns the summary note's id, or 0 when there
// was no summary to write.
//
// The note source carries the blob id, which is the only link back to the audio.
// When retention prunes the blob the note remains — words about a meeting are a
// far lighter thing to keep than a recording of it.
func (c *captureWiring) writeNotes(ctx context.Context, res capture.Result) (int64, error) {
source := "capture:meeting"
if res.BlobID != "" {
source = "capture:meeting:" + res.BlobID[:12]
}
var id int64
if text := res.Summary; text != "" {
var err error
id, err = c.writeNote(ctx, text, source)
if err != nil {
return 0, fmt.Errorf("summary note: %w", err)
}
}
if c.cfg.SaveTranscript && res.Transcript != "" {
if _, err := c.writeNote(ctx, res.Transcript, source+":transcript"); err != nil {
return id, fmt.Errorf("transcript note: %w", err)
}
}
return id, nil
}
func (c *captureWiring) writeNote(ctx context.Context, text, source string) (int64, error) {
var vec []float32
if c.emb != nil {
// EmbedPassage, not Embed: this is text being searched FOR, and the e5
// embedder is asymmetric. Backwards here makes the meeting unfindable by
// the question that should have matched it.
var err error
vec, err = router.EmbedPassage(ctx, c.emb, text)
if err != nil {
return 0, fmt.Errorf("embed: %w", err)
}
}
return c.st.WriteNote(ctx, c.now(), text, vec, source)
}
// status handles ipc.MethodCaptureStatus.
func (c *captureWiring) status(_ context.Context) (ipc.CaptureStatusResp, error) {
st := c.rec.Status()
return ipc.CaptureStatusResp{
Running: st.Running,
Label: st.Label,
Started: st.Started,
Seconds: st.Duration.Seconds(),
Bytes: st.Bytes,
}, nil
}
// wireCapture installs the four IPC hooks, or leaves them nil so every capture
// method reports ErrUnknownMethod. Takes the media keeper wireVision already
// opened: one blob store, one retention loop, images and audio side by side.
func wireCapture(srv *ipc.Server, keeper *mediaKeeper, st *store.Store, voiceW *voiceWiring, phr phraser.Phraser, cfg *config.Config) {
cw := newCaptureWiring(keeper, st, voiceW, phr, embedderOf(voiceW), cfg)
if cw == nil {
return
}
srv.CaptureStartFn = cw.start
srv.CaptureAppendFn = cw.append
srv.CaptureStopFn = cw.stop
srv.CaptureStatusFn = cw.status
}
+292
View File
@@ -0,0 +1,292 @@
package main
import (
"context"
"log"
"math/rand"
"strings"
"time"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/router"
)
// clarifyTTL — how long a parked question stays answerable. Same 90s as the
// confirm gate, for the same reason: an answer is a same-breath gesture, and a
// stale question must not eat an unrelated later utterance.
const clarifyTTL = 90 * time.Second
// wantedSlots — what each intent needs before she can act on it. First entry is
// the one she asks about; the rest are only used to decide act-vs-drop.
//
// Intents not listed here are never worth a question: note and query act on the
// raw utterance, chat and system have nothing to fill in. For those a clarify
// decision keeps the canned "не поняла" reply — inventing a question for noise
// is worse than admitting she missed it.
// A reminder wants BOTH what to remind about and when. Subject first: "напомни
// в 11" has a time and nothing to say at 11, and a reminder with no subject is
// not worth setting. Order here is the order she asks in — she still only asks
// about the first one missing.
var wantedSlots = map[router.Intent][]dialogue.Slot{
router.IntentReminder: {dialogue.SlotText, dialogue.SlotTime},
router.IntentFact: {dialogue.SlotKey},
router.IntentAct: {dialogue.SlotFn},
}
// clarifyQuestions — one short question per missing slot.
//
// These are fixed templates, not model output. The resident model is a 0.8B; it
// would wander, and a question whose wording changes every time is harder to
// answer than a blunt one that always reads the same. They are infinitive
// questions, so there is no gender agreement to get wrong; the feminine
// self-reference lives in the reply she gives when she drops the request.
var clarifyQuestions = map[dialogue.Slot]string{
dialogue.SlotTime: "Когда?",
dialogue.SlotText: "О чём напомнить?",
dialogue.SlotKey: "Что записать?",
dialogue.SlotFn: "Что сделать?",
}
// clarifyGaveUp — she is out of questions and still does not have the slot. She
// says so out loud: dropping the request in silence would leave him thinking it
// landed. Feminine self-reference ("поняла"), as everywhere.
const clarifyGaveUp = "Прости, я не поняла. Скажи, пожалуйста, по-другому."
// clarifyExpiredVariants — his answer came after the TTL, so the parked request
// is already gone. Same tone as clarifyGaveUp, different reason: too much time
// passed, not "I did not understand". Feminine self-reference ("ждала",
// "отпустила"); he is addressed with a plain imperative.
//
// Five phrasings, not one. This is the line he hears whenever he walks off
// mid-request, so it is the line that repeats most — and the same sentence every
// time is what makes a house assistant sound like a kiosk. They all carry the
// same two facts (the old request is gone; say it again if it still matters),
// because the wording may vary and the meaning may not.
//
// Fixed templates rather than model output, for the same reason as
// clarifyQuestions: this text has to be right every time, and it is not worth a
// generation to say something this small.
var clarifyExpiredVariants = []string{
"Прости, я слишком долго ждала ответа и отпустила прошлую просьбу. Если она ещё нужна, скажи заново.",
"Кажется, прошлая просьба уже не важна — я её отпустила. Если я ошибаюсь, повтори.",
"Ты как-то резко замолчал, и я не стала ждать дальше. Если та просьба ещё нужна, скажи заново.",
"Я не дождалась ответа и убрала прошлую просьбу. Повтори, если она всё ещё нужна.",
"Столько времени прошло, что я отпустила прошлую просьбу. Скажи заново, если она в силе.",
}
// clarifyExpiredLine picks one of them at random.
func clarifyExpiredLine() string {
return clarifyExpiredVariants[rand.Intn(len(clarifyExpiredVariants))]
}
// isClarifyExpired reports whether s opens with any of the expiry lines. The
// notice is glued in front of this turn's reply (see withNotice), so a caller
// checking for it has to match a prefix, not the whole string.
func isClarifyExpired(s string) bool {
for _, v := range clarifyExpiredVariants {
if strings.HasPrefix(s, v) {
return true
}
}
return false
}
// trimClarifyExpired strips a leading expiry notice, leaving this turn's actual
// reply. "" ⇒ the notice was the whole thing.
func trimClarifyExpired(s string) string {
for _, v := range clarifyExpiredVariants {
if strings.HasPrefix(s, v) {
return strings.TrimSpace(strings.TrimPrefix(s, v))
}
}
return strings.TrimSpace(s)
}
// clarifyExpiredNotice returns that line when a parked question had just timed
// out, and "" when nothing was parked. Call it right after
// resolveClarifyAnswer: a live question is answered there, an expired one is
// only reported here — the words themselves still go on to be routed fresh.
func (h *reactiveHandler) clarifyExpiredNotice() string {
if h.clarifyStore == nil {
return ""
}
if !h.clarifyStore.TakeExpired(voiceDialogueID, h.now()) {
return ""
}
log.Printf("voice: clarify — parked question expired, telling him and routing the words fresh")
return clarifyExpiredLine()
}
// withNotice glues the expiry notice in front of this turn's reply. One turn
// carries one reply on the wire, so the notice cannot be a message of its own —
// but neither the notice nor the fresh answer may be dropped.
func withNotice(notice, reply string) string {
if notice == "" {
return reply
}
if reply == "" {
return notice
}
return notice + " " + reply
}
// missingFor returns the slots a decision still needs, most important first.
// Empty ⇒ there is nothing identifiable to ask about.
func missingFor(dec router.Decision) []dialogue.Slot {
return dialogue.StillMissing(wantedSlots[dec.Intent], toDialogueSlots(dec.Slots))
}
// clarifyQuestion picks the one question to ask for a clarify decision. Returns
// ("", false) when she has no idea what is missing.
//
// One question about one thing: if two slots are missing she asks about the
// first and lets the rest go. Two questions in a row is an interrogation.
func clarifyQuestion(dec router.Decision) (dialogue.Slot, string, bool) {
missing := missingFor(dec)
if len(missing) == 0 {
return "", "", false
}
q, ok := clarifyQuestions[missing[0]]
if !ok {
return "", "", false
}
return missing[0], q, true
}
// askClarify parks the request and returns the question to ask instead of the
// canned "не поняла". Returns ("", false) when there is nothing to ask about, so
// the caller falls back to the canned reply.
func (h *reactiveHandler) askClarify(dec router.Decision) (string, bool) {
if h.clarifyStore == nil {
return "", false
}
slot, question, ok := clarifyQuestion(dec)
if !ok {
return "", false
}
h.clarifyStore.Put(voiceDialogueID, &dialogue.PendingQuestion{
Intent: dialogue.Intent(dec.Intent),
Slots: toDialogueSlots(dec.Slots),
Missing: []dialogue.Slot{slot},
Utterance: dec.Utterance,
Asked: h.now(),
TTL: clarifyTTL,
Attempts: 1, // this ask
MaxAttempts: h.clarifyMaxAttempts,
})
log.Printf("voice: clarify — asked about %s for intent=%s", slot, dec.Intent)
return question, true
}
// resolveClarifyAnswer reads an utterance as the answer to a parked question.
// Returns ("", false) when no live question is parked (or it expired), so the
// caller routes the utterance normally as a fresh request. Sibling of
// resolveConfirm and checked in the same place.
//
// The answer is parsed with the same extractor the router uses, for the intent
// she parked — no second parser. If it still does not fill the gap she asks
// again, up to MaxAttempts; after that she says out loud that she did not
// understand. She never drops the request in silence.
func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string) (string, bool) {
if h.clarifyStore == nil {
return "", false
}
q := h.clarifyStore.Get(voiceDialogueID, h.now())
if q == nil {
return "", false
}
intent := router.Intent(q.Intent)
answer := h.extractor.Extract(ctx, intent, text, h.now())
merged := q.Answer(text, toDialogueSlots(answer))
if len(dialogue.StillMissing(q.Missing, merged)) > 0 {
return h.reaskOrGiveUp(q, merged, text), true
}
h.clarifyStore.Delete(voiceDialogueID)
// Rebuild the decision as if it had routed cleanly, then run it down the
// normal path. Clarify is deliberately false and the intent is unchanged:
// filling in an argument never grants authority, so the completed decision
// still meets the allowlist and the destructive-act confirm gate in
// applyAction exactly like any other decision.
dec := router.Decision{
Utterance: q.Utterance,
Stage: 2,
Intent: intent,
Slots: applyDialogueSlots(answer, merged),
}
return h.finishClarified(ctx, dec), true
}
// reaskOrGiveUp handles an answer that left the gap open: ask the same question
// again while she has attempts left, otherwise say she did not understand and
// let the request go. Never returns "" — a mute give-up reads as "done".
func (h *reactiveHandler) reaskOrGiveUp(q *dialogue.PendingQuestion, merged dialogue.Slots, text string) string {
question := ""
if len(q.Missing) > 0 {
question = clarifyQuestions[q.Missing[0]]
}
if question == "" || !q.CanAsk() {
h.clarifyStore.Delete(voiceDialogueID)
log.Printf("voice: clarify — gave up on %v after %d question(s), answer was %q", q.Missing, q.Attempts, text)
return clarifyGaveUp
}
// Re-park with whatever the answer DID give, the clock restarted and one
// more question spent.
q.Slots = merged
q.Attempts++
q.Asked = h.now()
h.clarifyStore.Put(voiceDialogueID, q)
log.Printf("voice: clarify — answer %q did not fill %v, asking again (attempt %d)", text, q.Missing, q.Attempts)
return question
}
// finishClarified runs a completed decision through the same steps a freshly
// routed one takes: remember the turn, act, then phrase.
func (h *reactiveHandler) finishClarified(ctx context.Context, dec router.Decision) string {
if h.dialogueSessions != nil {
now := h.now()
prev := h.dialogueSessions.Get(voiceDialogueID, now)
dec = followUpMerge(prev, dec, now)
h.rememberTurn(prev, dec, now)
}
reply := h.applyAction(ctx, dec)
if reply == "" {
reply = h.replier.Reply(dec)
}
if reply == "" {
// Belt: an empty reply here would be a silent drop.
reply = clarifyGaveUp
}
return reply
}
// rememberTurn stores this turn as the dialogue session the next follow-up
// inherits from, carrying up to 4 prior turns of history for anaphora. Capped so
// one long conversation can't grow the session unboundedly.
func (h *reactiveHandler) rememberTurn(prev *dialogue.Session, dec router.Decision, now time.Time) {
var history []dialogue.Turn
if prev != nil {
history = append(history, dialogue.Turn{
Intent: prev.Intent,
Slots: prev.Slots,
Text: prev.Slots.Text,
})
maxHist := len(prev.History)
if maxHist > 3 {
maxHist = 3
}
history = append(history, prev.History[:maxHist]...)
}
ttl := time.Duration(0) // use the store default (2 min)
if dec.Intent == router.IntentChat {
ttl = 15 * time.Minute // conversational turns should last longer
}
h.dialogueSessions.Put(voiceDialogueID, &dialogue.Session{
Intent: dialogue.Intent(dec.Intent),
Slots: toDialogueSlots(dec.Slots),
Timestamp: now,
TTL: ttl,
History: history,
})
}
+332
View File
@@ -0,0 +1,332 @@
package main
import (
"context"
"os"
"path/filepath"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/tool"
"github.com/kami/maven/internal/voice"
)
// newClarifyHandler builds a handler with the clarify path wired and no model:
// stub date parser, the real fact parser, and a matcher over whatever tools the
// test enabled. `now` is fixed so TTL behaviour is testable.
func newClarifyHandler(t *testing.T) (*reactiveHandler, *store.Store, *time.Time) {
t.Helper()
st := newTestStore(t)
api := ipc.NewStoreAPI(st)
now := time.Date(2026, 7, 31, 9, 0, 0, 0, time.UTC)
matcher := tool.NewMatcher(api)
h := &reactiveHandler{
api: api,
dataStore: st,
tools: tool.NewExecutor(api, 2*time.Second),
matcher: matcher,
replier: voice.NewStubReplier(),
now: func() time.Time { return now },
dialogueSessions: dialogue.NewSessionStore(2 * time.Minute),
clarifyStore: dialogue.NewClarifyStore(clarifyTTL),
extractor: router.Extractor{
Time: router.StubDateTimeParser{},
Acts: matcher,
Facts: router.DefaultFactParser{},
},
}
return h, st, &now
}
func clarifyDec(intent router.Intent, slots router.Slots, utterance string) router.Decision {
return router.Decision{Utterance: utterance, Stage: 3, Intent: intent, Slots: slots, Clarify: true}
}
// TestClarifyQuestionForMissingSlot pins which question goes with which gap, and
// which intents get no question at all.
func TestClarifyQuestionForMissingSlot(t *testing.T) {
cases := []struct {
name string
dec router.Decision
want string
asked bool
}{
{"reminder without a time", clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"), "Когда?", true},
{"fact without a key", clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши"), "Что записать?", true},
{"act without a fn", clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это"), "Что сделать?", true},
// A time with nothing to say at that time is still half a reminder, so
// the subject is what she asks about — not silence.
{"reminder that has a time but no subject", clarifyDec(router.IntentReminder, router.Slots{HasTime: true}, "напомни в 11"), "О чём напомнить?", true},
{"reminder that has both", clarifyDec(router.IntentReminder, router.Slots{Text: "позвонить маме", HasTime: true}, "напомни в 11 позвонить маме"), "", false},
{"chat is never worth a question", clarifyDec(router.IntentChat, router.Slots{Text: "мгм"}, "мгм"), "", false},
{"query is never worth a question", clarifyDec(router.IntentQuery, router.Slots{Text: "а"}, "а"), "", false},
}
for _, tc := range cases {
_, got, asked := clarifyQuestion(tc.dec)
if asked != tc.asked || got != tc.want {
t.Errorf("%s: got (%q, %v), want (%q, %v)", tc.name, got, asked, tc.want, tc.asked)
}
}
}
// TestClarifyReminderCompletesOnAnswer is the whole point of the feature: she
// asks for the missing time and the answer creates the reminder.
func TestClarifyReminderCompletesOnAnswer(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
question, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"))
if !asked || question != "Когда?" {
t.Fatalf("expected the time question, got %q asked=%v", question, asked)
}
reply, handled := h.resolveClarifyAnswer(ctx, "в 11:00")
if !handled {
t.Fatal("the answer to an open question must be consumed as an answer")
}
if reply == clarifyGaveUp {
t.Fatalf("a good answer must not drop the request: %q", reply)
}
reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour))
if err != nil || len(reminders) != 1 {
t.Fatalf("clarified reminder was not created: reminders=%v err=%v", reminders, err)
}
if !strings.Contains(reminders[0].Payload, "маме") {
t.Fatalf("the reminder lost the original request: %q", reminders[0].Payload)
}
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
t.Fatal("the question must be cleared once answered")
}
}
// TestClarifyFactCompletesOnAnswer — the fact path, where the answer carries
// both the key and the value.
func TestClarifyFactCompletesOnAnswer(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
if _, asked := h.askClarify(clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши")); !asked {
t.Fatal("a fact with no key should be asked about")
}
if reply, handled := h.resolveClarifyAnswer(ctx, "пил воду"); !handled || reply == clarifyGaveUp {
t.Fatalf("answer should complete the fact, handled=%v reply=%q", handled, reply)
}
if fact, err := st.LatestFact(ctx, "water"); err != nil || fact.Key != "water" {
t.Fatalf("clarified fact was not written: fact=%+v err=%v", fact, err)
}
}
// TestClarifyAnswerAfterTTLIsANewRequest — a late answer is not an answer.
func TestClarifyAnswerAfterTTLIsANewRequest(t *testing.T) {
ctx := context.Background()
h, st, now := newClarifyHandler(t)
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a question")
}
*now = now.Add(clarifyTTL + time.Second)
if reply, handled := h.resolveClarifyAnswer(ctx, "в 11:00"); handled {
t.Fatalf("an answer past the TTL must fall through to normal routing, got %q", reply)
}
if reminders, err := st.DueReminders(ctx, now.Add(48*time.Hour)); err != nil || len(reminders) != 0 {
t.Fatalf("expired question must not create anything: reminders=%v err=%v", reminders, err)
}
}
// TestClarifyAsksThreeTimesThenSaysSo — three questions are allowed, the fourth
// is not, and running out is SPOKEN. Silence would read as "handled".
func TestClarifyAsksThreeTimesThenSaysSo(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a first question")
}
// Two more unclear answers ⇒ two more questions (3 asks in total).
for i := 2; i <= 3; i++ {
reply, handled := h.resolveClarifyAnswer(ctx, "ну не знаю")
if !handled {
t.Fatalf("answer %d must be consumed as an answer", i)
}
if reply != "Когда?" {
t.Fatalf("attempt %d should ask again, got %q", i, reply)
}
if h.clarifyStore.Get(voiceDialogueID, h.now()) == nil {
t.Fatalf("attempt %d must leave the question armed", i)
}
}
reply, handled := h.resolveClarifyAnswer(ctx, "ну не знаю")
if !handled || reply != clarifyGaveUp {
t.Fatalf("the fourth try must give up out loud, handled=%v reply=%q", handled, reply)
}
if reply == "" || strings.Contains(reply, "?") {
t.Fatalf("giving up must be spoken and must not be another question: %q", reply)
}
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
t.Fatal("a given-up request must leave no armed question")
}
if reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour)); err != nil || len(reminders) != 0 {
t.Fatalf("a given-up request must not create anything: reminders=%v err=%v", reminders, err)
}
}
// TestClarifyMaxAttemptsIsConfigurable — one question when the config says one.
func TestClarifyMaxAttemptsIsConfigurable(t *testing.T) {
ctx := context.Background()
h, _, _ := newClarifyHandler(t)
h.clarifyMaxAttempts = 1
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a question")
}
if reply, handled := h.resolveClarifyAnswer(ctx, "ну не знаю"); !handled || reply != clarifyGaveUp {
t.Fatalf("with max 1 she must give up at once, handled=%v reply=%q", handled, reply)
}
}
// TestClarifyRestatedAnswerWins — «в 11:00», then «нет, в 15:00». The second
// value is the one that lands.
func TestClarifyRestatedAnswerWins(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме")); !asked {
t.Fatal("expected a question")
}
// First answer parses, but re-park it by hand as if she had asked again:
// what matters here is that Answer prefers the newer value over the parked
// one, which is the case the daemon hits on a re-ask.
q := h.clarifyStore.Get(voiceDialogueID, h.now())
if q == nil {
t.Fatal("expected an armed question")
}
first := h.extractor.Extract(ctx, router.IntentReminder, "в 11:00", h.now())
q.Slots = q.Answer("в 11:00", toDialogueSlots(first))
if reply, handled := h.resolveClarifyAnswer(ctx, "нет, в 15:00"); !handled || reply == clarifyGaveUp {
t.Fatalf("the restated answer should complete the request, handled=%v reply=%q", handled, reply)
}
reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour))
if err != nil || len(reminders) != 1 {
t.Fatalf("expected one reminder: %v err=%v", reminders, err)
}
want := h.extractor.Extract(ctx, router.IntentReminder, "в 15:00", h.now())
if !reminders[0].FireTs.Equal(want.Time) {
t.Fatalf("reminder at %v, want the restated %v", reminders[0].FireTs, want.Time)
}
}
// TestClarifiedActOffAllowlistIsStillRefused — clarification fills in an
// argument, it never grants authority.
func TestClarifiedActOffAllowlistIsStillRefused(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
marker := filepath.Join(t.TempDir(), "not-allowed-ran")
if _, asked := h.askClarify(clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это")); !asked {
t.Fatal("an act with no fn should be asked about")
}
reply, handled := h.resolveClarifyAnswer(ctx, "rm "+marker)
if !handled {
t.Fatal("the answer should be consumed")
}
if strings.Contains(reply, "готово") {
t.Fatalf("an act that is not on the allowlist must not report success: %q", reply)
}
if _, err := os.Stat(marker); !os.IsNotExist(err) {
t.Fatalf("a clarified act off the allowlist ran anyway: %v", err)
}
if tools, err := st.ListTools(ctx, "enabled"); err != nil || len(tools) != 0 {
t.Fatalf("clarify must not enable a tool: tools=%+v err=%v", tools, err)
}
}
// TestClarifiedDestructiveActStillNeedsConfirm — the confirm gate survives the
// clarify path.
func TestClarifiedDestructiveActStillNeedsConfirm(t *testing.T) {
ctx := context.Background()
h, st, _ := newClarifyHandler(t)
marker := filepath.Join(t.TempDir(), "destructive-ran")
if err := st.EnableTool(ctx, "delete_backups", []string{"touch", marker}, true, "test", h.now()); err != nil {
t.Fatal(err)
}
if _, asked := h.askClarify(clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это")); !asked {
t.Fatal("expected a question")
}
reply, handled := h.resolveClarifyAnswer(ctx, "delete_backups")
if !handled {
t.Fatal("the answer should be consumed")
}
if !strings.Contains(reply, "да") || h.pending == nil {
t.Fatalf("a clarified destructive act must still park a confirm: reply=%q pending=%+v", reply, h.pending)
}
if _, err := os.Stat(marker); !os.IsNotExist(err) {
t.Fatalf("a clarified destructive act ran before confirmation: %v", err)
}
}
// TestNoQuestionWhenNothingIsMissing — noise keeps the canned reply, so she
// never invents a question for nothing.
func TestNoQuestionWhenNothingIsMissing(t *testing.T) {
h, _, _ := newClarifyHandler(t)
for _, dec := range []router.Decision{
clarifyDec(router.IntentChat, router.Slots{Text: "эм"}, "эм"),
clarifyDec(router.IntentQuery, router.Slots{Text: "ммм"}, "ммм"),
clarifyDec(router.IntentNote, router.Slots{Text: "..."}, "..."),
} {
if question, asked := h.askClarify(dec); asked {
t.Fatalf("intent %s should keep the canned reply, got %q", dec.Intent, question)
}
}
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
t.Fatal("noise must not park a question")
}
}
// TestClarifyExpiryIsAnnouncedAndWordsStillRoute — his answer lands after the
// TTL: she must say the old request is gone AND still answer the new words.
func TestClarifyExpiryIsAnnouncedAndWordsStillRoute(t *testing.T) {
ctx := context.Background()
h, _, now := newClarifyHandler(t)
emb := router.NewHashEmbedder(1024)
h.embedder = emb
h.router = buildRouter(emb, h.matcher, 0.55, nil)
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
t.Fatal("expected a question")
}
*now = now.Add(clarifyTTL + time.Second)
reply := h.handleText(ctx, "как дела")
if !isClarifyExpired(reply) {
t.Fatalf("expired question must be announced first, got %q", reply)
}
if trimClarifyExpired(reply) == "" {
t.Fatalf("the new words must still be answered, got only the notice: %q", reply)
}
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
t.Fatal("the expired question must be gone")
}
// The notice is said once, not on every later utterance.
if reply := h.handleText(ctx, "как дела"); isClarifyExpired(reply) {
t.Fatalf("notice repeated on a later turn: %q", reply)
}
}
// TestNoPendingQuestionFallsThrough — with nothing parked, an utterance routes
// normally.
func TestNoPendingQuestionFallsThrough(t *testing.T) {
h, _, _ := newClarifyHandler(t)
if reply, handled := h.resolveClarifyAnswer(context.Background(), "напомни в 11:00"); handled {
t.Fatalf("no open question ⇒ must not be treated as an answer, got %q", reply)
}
}
+190
View File
@@ -0,0 +1,190 @@
package main
import (
"bytes"
"context"
"crypto/rand"
"io"
"os"
"path/filepath"
"testing"
"time"
"github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/webauthn"
)
func randBytes(t *testing.T, n int) []byte {
t.Helper()
b := make([]byte, n)
if _, err := io.ReadFull(rand.Reader, b); err != nil {
t.Fatalf("rand: %v", err)
}
b[0] |= 1
return b
}
func TestDaemonLockStartsLockedAndFlips(t *testing.T) {
dl := newDaemonLock(true)
if !dl.isLocked() {
t.Fatal("newDaemonLock(true) is not locked")
}
dl.unlock(nil)
if dl.isLocked() {
t.Fatal("still locked after unlock")
}
if newDaemonLock(false).isLocked() {
t.Fatal("newDaemonLock(false) reports locked")
}
}
// closeStore must be safe on a daemon that never unlocked and safe twice —
// shutdown runs it unconditionally.
func TestDaemonLockCloseStoreIsSafeWhenNeverUnlocked(t *testing.T) {
dl := newDaemonLock(true)
if err := dl.closeStore(); err != nil {
t.Fatalf("closeStore with no store: %v", err)
}
if err := dl.closeStore(); err != nil {
t.Fatalf("second closeStore: %v", err)
}
}
// The data-loss bug: in locked mode the store is opened on an IPC goroutine
// inside UnlockFn, and shutdown runs on main. Without the handoff nothing
// calls Close, and Close is what re-encrypts the tmpfs working copy back over
// the ciphertext file — so every write of a cold-started session vanished.
func TestDaemonLockSealsTheStoreOpenedAfterUnlock(t *testing.T) {
dir := t.TempDir()
dbPath := filepath.Join(dir, "maven.db")
tmpfs := filepath.Join(dir, "work")
key := randBytes(t, 32)
// Store.Close zeroes the key slice it was handed (encState.key is the
// caller's backing array), so the next boot needs its own copy — exactly
// as mavend keeps envKeyBytes separate from the config's key.
nextBoot := bytes.Clone(key)
ctx := context.Background()
// Cold start: locked, no store.
dl := newDaemonLock(true)
// ... unlock arrives, opens the store and hands it over.
st, err := store.OpenEncrypted(ctx, dbPath, tmpfs, key)
if err != nil {
t.Fatalf("OpenEncrypted: %v", err)
}
dl.unlock(st)
if _, err := st.WriteNote(ctx, time.Now(), "заметка после холодного старта", nil, "test"); err != nil {
t.Fatalf("WriteNote: %v", err)
}
// Shutdown.
if err := dl.closeStore(); err != nil {
t.Fatalf("closeStore: %v", err)
}
if err := dl.closeStore(); err != nil {
t.Fatalf("second closeStore after a real store: %v", err)
}
// Next boot with the same key must see the write.
st2, err := store.OpenEncrypted(ctx, dbPath, tmpfs, nextBoot)
if err != nil {
t.Fatalf("reopen: %v", err)
}
defer st2.Close()
notes, err := st2.RecentNotes(ctx, 10)
if err != nil {
t.Fatalf("RecentNotes: %v", err)
}
if len(notes) != 1 {
t.Fatalf("got %d notes after a cold-started session, want 1 — the session was lost", len(notes))
}
}
// The whole point of the wrapped blob: what sits in the state dir must not let
// anyone open the database. Nothing written there may contain the key, and the
// ciphertext must not be readable with a wrong one.
func TestColdStartLeavesNoPlaintextKeyOnDisk(t *testing.T) {
dir := t.TempDir()
dbPath := filepath.Join(dir, "maven.db")
tmpfs := filepath.Join(dir, "work")
wrappedPath := filepath.Join(dir, "db_key.wrapped")
key := randBytes(t, 32)
secret := randBytes(t, 32)
ctx := context.Background()
blob, err := webauthn.WrapKey(key, secret)
if err != nil {
t.Fatalf("WrapKey: %v", err)
}
if err := os.WriteFile(wrappedPath, blob, 0o600); err != nil {
t.Fatalf("write wrapped key: %v", err)
}
st, err := store.OpenEncrypted(ctx, dbPath, tmpfs, key)
if err != nil {
t.Fatalf("OpenEncrypted: %v", err)
}
if _, err := st.WriteNote(ctx, time.Now(), "секрет", nil, "test"); err != nil {
t.Fatalf("WriteNote: %v", err)
}
if err := st.Close(); err != nil {
t.Fatalf("Close: %v", err)
}
// Walk everything in the state dir; none of it may contain the key.
err = filepath.Walk(dir, func(p string, info os.FileInfo, err error) error {
if err != nil || info.IsDir() {
return err
}
b, rerr := os.ReadFile(p)
if rerr != nil {
return nil // unreadable is not a leak
}
if bytes.Contains(b, key) {
t.Errorf("%s contains the plaintext encryption key", p)
}
return nil
})
if err != nil {
t.Fatalf("walk: %v", err)
}
// The wrapped file must have owner-only permissions.
fi, err := os.Stat(wrappedPath)
if err != nil {
t.Fatalf("stat: %v", err)
}
if perm := fi.Mode().Perm(); perm != 0o600 {
t.Errorf("wrapped key file mode = %o, want 600", perm)
}
// A wrong passkey must not open the store.
if _, _, err := webauthn.UnwrapKey(blob, randBytes(t, 32)); err == nil {
t.Fatal("a wrong PRF secret unwrapped the key")
}
if _, err := store.OpenEncrypted(ctx, dbPath, filepath.Join(dir, "work2"), randBytes(t, 32)); err == nil {
t.Fatal("the encrypted store opened under a wrong key")
}
// And the right one round-trips back to a readable database.
got, version, err := webauthn.UnwrapKey(blob, secret)
if err != nil {
t.Fatalf("UnwrapKey: %v", err)
}
if version != webauthn.BlobV2 {
t.Errorf("blob version = %v, want v2", version)
}
st2, err := store.OpenEncrypted(ctx, dbPath, tmpfs, got)
if err != nil {
t.Fatalf("reopen with the unwrapped key: %v", err)
}
defer st2.Close()
notes, err := st2.RecentNotes(ctx, 10)
if err != nil {
t.Fatalf("RecentNotes: %v", err)
}
if len(notes) != 1 {
t.Fatalf("got %d notes, want 1", len(notes))
}
}
+216
View File
@@ -0,0 +1,216 @@
package main
import (
"context"
"log"
"strings"
"time"
"github.com/kami/maven/internal/router"
)
// pendingHexisExec — a mutating Hexis capability parked awaiting a spoken
// confirm. The confirmation is bound to the resolved capability + canonical
// target entity so a later "да" can only execute exactly what was proposed
// (ecosystem invariant: protected actions require bound confirmation).
type pendingHexisExec struct {
capabilityID string
capName string
entityID string
displayName string
expiry time.Time
}
// pendingRoutineConfirm — a proposed routine awaiting a spoken y/n to become
// a recurring reminder. Set by detectPattern after creating a proposal.
type pendingRoutineConfirm struct {
routineID int64
action string
object string
interval float64
phrase string
expiry time.Time
}
// pendingAct — a destructive act awaiting a spoken confirm.
type pendingAct struct {
fn string
args []string
phrase string
expiry time.Time
}
// confirmTTL — how long a parked destructive confirm stays answerable. Short:
// a confirm is a same-breath gesture; a stale prompt shouldn't fire on an
// unrelated later "да".
const confirmTTL = 90 * time.Second
// park stores a destructive act awaiting confirmation. Overwrites any prior
// pending (last-asked wins — single-user box).
func (h *reactiveHandler) park(fn string, args []string, phrase string) {
h.mu.Lock()
h.pending = &pendingAct{fn: fn, args: args, phrase: phrase, expiry: h.now().Add(confirmTTL)}
h.mu.Unlock()
}
// resolveConfirm interprets an utterance as the answer to a parked destructive
// act OR a parked routine proposal. Returns (reply, true) when it consumed the
// utterance as a y/n answer; ("", false) when there's nothing pending (or the
// parked act expired), so the caller routes the utterance normally. An
// unrecognised answer cancels the pending and routes normally — a confirm that
// can't be answered clearly is safer abandoned than left armed.
func (h *reactiveHandler) resolveConfirm(ctx context.Context, text string) (string, bool) {
h.mu.Lock()
defer h.mu.Unlock()
for _, r := range h.confirmResolvers(ctx) {
if !r.claim() {
continue
}
// The slot is already cleared by claim(): every branch below drops the
// pending, including the unclear one — a confirm that can't be
// answered clearly is safer abandoned than left armed.
switch classifyConfirm(text) {
case confirmYes:
return r.yes(), true
case confirmNo:
return r.no(), true
default:
return "", false
}
}
return "", false
}
// confirmResolver — one parked-confirm slot in the chain. claim() reports
// whether this slot holds a live pending, taking it (and dropping an expired
// one) as it goes; yes/no then run the answer. Only ever called with h.mu held.
type confirmResolver struct {
claim func() bool
yes func() string
no func() string
}
// confirmResolvers builds the ordered chain resolveConfirm walks. Order is
// deliberate: the routine proposal is checked before the tool confirm so a
// routine confirm doesn't get eaten by a stale tool pending.
func (h *reactiveHandler) confirmResolvers(ctx context.Context) []confirmResolver {
var pr *pendingRoutineConfirm
var hx *pendingHexisExec
var p *pendingAct
return []confirmResolver{
// Routine proposal.
{
claim: func() bool {
pr, h.pendingRoutine = h.pendingRoutine, nil
return pr != nil && !h.now().After(pr.expiry)
},
yes: func() string {
// Only record the acceptance. The tick loop reads accepted
// routines and nudges on their own interval. Building a
// reminder here made a routine fire exactly once (Vikunja #366).
if err := h.dataStore.AcceptProposedRoutine(ctx, pr.routineID, h.now()); err != nil {
log.Printf("voice: accept proposed routine: %v", err)
return "не получилось запомнить рутину."
}
return "буду напоминать."
},
no: func() string {
if err := h.dataStore.DismissProposedRoutine(ctx, pr.routineID); err != nil {
log.Printf("voice: dismiss proposed routine: %v", err)
}
return "хорошо, не буду."
},
},
// Hexis execution confirm. Bound to the exact capability + target that
// was proposed; a stray "да" can only run that, nothing else.
{
claim: func() bool {
hx, h.pendingHexis = h.pendingHexis, nil
return hx != nil && !h.now().After(hx.expiry)
},
yes: func() string {
return h.execHexis(ctx, hx.capabilityID, hx.capName, hx.entityID, hx.displayName)
},
no: func() string { return "отменила." },
},
// Tool confirm.
{
claim: func() bool {
p, h.pending = h.pending, nil
return p != nil && !h.now().After(p.expiry)
},
yes: func() string {
out, err := h.tools.Exec(ctx, p.fn, p.args, true) // confirmed
if err != nil {
log.Printf("voice: tool %s (confirmed): %v", p.fn, err)
if out != "" {
return "не получилось выполнить команду: " + firstLine(out)
}
return "не получилось выполнить команду."
}
if out != "" {
return "готово: " + firstLine(out)
}
return "готово."
},
no: func() string { return "отменила." },
},
}
}
// proposeGap scaffolds a 'proposed' tool for an act whose verb isn't enabled.
// maven drafts the registration (name = the verb, provenance = the utterance);
// a human enables it on the authed surface. She suggests, never enables.
func (h *reactiveHandler) proposeGap(ctx context.Context, dec router.Decision) string {
name := firstWord(stripWake(dec.Utterance))
if name == "" {
return "не разобрала команду — попробуй иначе."
}
newly, err := h.api.ProposeTool(ctx, name, dec.Utterance, "", h.now())
if err != nil {
log.Printf("voice: propose tool %q: %v", name, err)
return "команды «" + name + "» нет в списке разрешённых."
}
if newly {
return "команды «" + name + "» нет в списке. Предложила её добавить — включи через клиент."
}
return "команды «" + name + "» пока нет в списке — она уже предложена, включи через клиент."
}
// confirmVerdict — the parse of a y/n confirm answer.
type confirmVerdict int
const (
confirmUnknown confirmVerdict = iota
confirmYes
confirmNo
)
// classifyConfirm reads a short ru/en yes-or-no answer. Substring match on the
// stems so inflections/fillers ("да, давай", "нет, отмени") still land.
func classifyConfirm(text string) confirmVerdict {
t := strings.ToLower(strings.TrimSpace(text))
// negatives first — "не надо" contains no "да", but check no-stems before
// yes so a leading "нет" isn't shadowed.
for _, no := range []string{"нет", "не надо", "отмен", "стоп", "no", "cancel", "stop", "don't"} {
if strings.Contains(t, no) {
return confirmNo
}
}
for _, yes := range []string{"да", "ага", "давай", "подтвер", "конечно", "yes", "yeah", "yep", "confirm", "ок", "okay", "ok"} {
if strings.Contains(t, yes) {
return confirmYes
}
}
return confirmUnknown
}
// actPhrase renders "fn arg1 arg2" for the confirm prompt.
func actPhrase(fn string, args []string) string {
if len(args) == 0 {
return fn
}
return fn + " " + strings.Join(args, " ")
}
+184
View File
@@ -0,0 +1,184 @@
// mavend/crawls.go — the driver for reading web pages (Vikunja #259,
// docs/plans/14-web-crawler.md). The crawler is pure and lives in
// internal/crawl; this is the impure half: the guarded fetcher, a ticker for the
// scheduled watches, and the fact-backed dedup hashes.
//
// Two paths, one config block, both off unless configured:
//
// - ON DEMAND — he names a URL out loud and she reads it. That is the
// `queryWeb` source in actions_query.go, LAST in the chain: after his
// memory, after the notes, and (once Kiwix is wired into the chain) after
// the local ZIMs. A local read costs nothing and leaks nothing; a fetch puts
// a URL in someone's log, so it goes last.
// - SCHEDULED — a watched page is re-read on its interval, and a page whose
// text changed is written as a note. It does NOT announce itself. Same rule
// as the feed poller: notes, never nudges.
//
// Only the URL goes out. Nothing here reads a note, a fact, the persona block or
// the history, and internal/crawl has no access to the store at all.
package main
import (
"context"
"log"
"net/url"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/crawl"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/webfetch"
)
// newCrawler builds the crawler from the `crawl` block, or returns nil when
// there is none. Every caller checks for nil, and nil means no page is ever
// fetched.
func newCrawler(cfg *config.Config) *crawl.Crawler {
if cfg.Crawl == nil {
return nil
}
cc := cfg.Crawl
hosts := append([]string(nil), cc.AllowHosts...)
// A watched page's own host is always reachable; otherwise an allowlist and
// a watch list would have to be kept in sync by hand.
for _, w := range cc.Watches {
if u, err := url.Parse(w.URL); err == nil && u.Hostname() != "" {
hosts = append(hosts, u.Hostname())
}
}
// An allowlist plus on-demand is a contradiction worth logging rather than
// silently resolving: he asked for arbitrary pages AND for a fixed list.
// The allowlist wins, because it is the narrower instruction.
if len(hosts) > 0 && cc.OnDemand && len(cc.AllowHosts) > 0 {
log.Printf("crawl: allow_hosts is set, so on-demand reading is limited to those hosts")
}
ua := cc.UserAgent
if ua == "" {
ua = webfetch.DefaultUserAgent
}
fetcher := webfetch.New(webfetch.Config{
AllowHosts: hosts,
DenyHosts: cc.DenyHosts,
Timeout: time.Duration(cc.Timeout),
MaxBytes: cc.MaxBytes,
UserAgent: ua,
})
// The user-agent handed to the crawler is the one the fetcher sends: obeying
// robots rules written for a different name would be a lie.
return crawl.New(&crawlFetcher{f: fetcher}, crawl.Config{
UserAgent: ua,
MaxRunes: cc.MaxRunes,
})
}
// onDemandCrawler returns a crawler for the answer path, or nil when on-demand
// reading is off. The scheduled watches can be on while this is off: reading a
// fixed list of pages on a timer and reading whatever URL is in an utterance are
// different permissions, and the config keeps them separate.
func onDemandCrawler(cfg *config.Config) *crawl.Crawler {
if cfg.Crawl == nil || !cfg.Crawl.OnDemand {
return nil
}
return newCrawler(cfg)
}
// crawlWorker — ticker + watcher for the scheduled half.
type crawlWorker struct {
watcher *crawl.Watcher
interval time.Duration
}
// crawlTickInterval — how often the worker asks what is due. Per-watch cadence
// is the watcher's business.
const crawlTickInterval = 15 * time.Minute
// newCrawlWorker wires the scheduled crawls, or nil when nothing is watched.
func newCrawlWorker(c *crawl.Crawler, api ipc.CoreAPI, emb router.Embedder, cfg *config.Config) *crawlWorker {
if c == nil || cfg.Crawl == nil || len(cfg.Crawl.Watches) == 0 {
return nil
}
watches := make([]crawl.WatchConfig, 0, len(cfg.Crawl.Watches))
for _, w := range cfg.Crawl.Watches {
watches = append(watches, crawl.WatchConfig{
Name: w.Name,
URL: w.URL,
Interval: time.Duration(w.Interval),
})
}
watcher := crawl.NewWatcher(c, watches, api, &factHashes{api: api},
crawlEmbedder(emb), time.Duration(cfg.Crawl.Interval))
if watcher == nil {
log.Printf("crawl: configured but nothing watchable — scheduled crawls disabled")
return nil
}
log.Printf("crawl: watching %d page(s), checking what is due every %s", len(watches), crawlTickInterval)
return &crawlWorker{watcher: watcher, interval: crawlTickInterval}
}
// run checks what is due until ctx is canceled. The first round runs
// immediately; it writes notes only, so an early round startles nobody.
func (w *crawlWorker) run(ctx context.Context) {
w.watcher.CheckDue(ctx, time.Now())
t := time.NewTicker(w.interval)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case now := <-t.C:
w.watcher.CheckDue(ctx, now)
}
}
}
// crawlFetcher adapts webfetch to crawl.Fetcher, which is the seam that keeps
// net/http out of the crawler package.
type crawlFetcher struct{ f *webfetch.Fetcher }
func (a *crawlFetcher) Get(ctx context.Context, u string) (*crawl.Response, error) {
resp, err := a.f.Get(ctx, u)
if err != nil {
return nil, err
}
return &crawl.Response{URL: resp.URL, ContentType: resp.ContentType, Body: resp.Body}, nil
}
// factHashes stores each watch's last content hash as a config fact, so a
// restart does not re-note an unchanged page. Same mechanism the feed reader
// uses for its marks, and inspectable on /dash.
type factHashes struct{ api ipc.CoreAPI }
func hashKey(name string) string { return "crawl:hash:" + name }
func (h *factHashes) LastHash(ctx context.Context, name string) (string, error) {
f, err := h.api.LatestFact(ctx, hashKey(name))
if err != nil {
// No hash yet is not an error: the watcher treats "" as "never read".
return "", nil
}
return f.Value, nil
}
func (h *factHashes) SetHash(ctx context.Context, name, hash string) error {
_, err := h.api.WriteFact(ctx, ipc.WriteFactReq{
Ts: time.Now(),
Kind: "config",
Key: hashKey(name),
Value: hash,
Source: "poll:crawl",
Confidence: 1.0,
})
return err
}
// crawlEmbedder adapts router.Embedder for the watcher, embedding with
// EmbedPassage (a page is text being searched FOR, and the e5 embedder is
// asymmetric).
func crawlEmbedder(emb router.Embedder) crawl.Embedder {
if emb == nil {
return nil
}
return passageEmbedder{emb}
}
+186
View File
@@ -0,0 +1,186 @@
package main
import (
"context"
"net/http"
"net/http/httptest"
"strings"
"testing"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/crawl"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice"
)
// The default config reads nothing. This is the whole "off unless configured"
// contract for the crawler, asserted at the wiring level rather than trusted.
func TestCrawlOffByDefault(t *testing.T) {
cfg := &config.Config{}
if c := newCrawler(cfg); c != nil {
t.Error("newCrawler with no crawl block returned a crawler")
}
if c := onDemandCrawler(cfg); c != nil {
t.Error("onDemandCrawler with no crawl block returned a crawler")
}
if w := newCrawlWorker(nil, nil, nil, cfg); w != nil {
t.Error("newCrawlWorker with no crawl block returned a worker")
}
// Watches configured but on_demand off ⇒ the answer path still reads
// nothing: a timer over a fixed list is not permission for arbitrary URLs.
withWatch := &config.Config{Crawl: &config.CrawlConfig{
Watches: []config.CrawlWatchConfig{{Name: "p", URL: "https://example.org/p"}},
}}
if c := onDemandCrawler(withWatch); c != nil {
t.Error("onDemandCrawler honoured a watch list as on-demand permission")
}
if c := newCrawler(withWatch); c == nil {
t.Error("newCrawler returned nil for a configured watch")
}
}
// The wired fetcher must refuse a private address, because the crawler on this
// box sits one hop from the whole homelab. Same guard the webfetch tests cover;
// this asserts the daemon actually wires it.
func TestCrawlerRefusesPrivateAddress(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "text/html")
w.Write([]byte("<html><body>secret</body></html>"))
}))
defer srv.Close()
c := newCrawler(&config.Config{Crawl: &config.CrawlConfig{OnDemand: true}})
if c == nil {
t.Fatal("newCrawler returned nil for an on-demand config")
}
if _, err := c.Page(context.Background(), srv.URL); err == nil {
t.Fatalf("reading %s succeeded; a loopback address must be refused", srv.URL)
}
}
func TestFactHashesRoundTrip(t *testing.T) {
ctx := context.Background()
st := newTestStore(t)
h := &factHashes{api: ipc.NewStoreAPI(st)}
got, err := h.LastHash(ctx, "page")
if err != nil {
t.Fatalf("LastHash on a fresh store: %v", err)
}
if got != "" {
t.Errorf("LastHash = %q, want empty for a never-read page", got)
}
if err := h.SetHash(ctx, "page", "deadbeef"); err != nil {
t.Fatalf("SetHash: %v", err)
}
got, err = h.LastHash(ctx, "page")
if err != nil {
t.Fatalf("LastHash: %v", err)
}
if got != "deadbeef" {
t.Errorf("LastHash = %q, want deadbeef", got)
}
if key := hashKey("page"); key != "crawl:hash:page" {
t.Errorf("hashKey = %q", key)
}
}
// stubCrawlFetcher serves one fixed page to every URL, so queryWeb can be
// exercised without a network or an allowlist.
type stubCrawlFetcher struct{ body, ctype string }
func (s *stubCrawlFetcher) Get(_ context.Context, u string) (*crawl.Response, error) {
ct := s.ctype
if ct == "" {
ct = "text/html"
}
if strings.HasSuffix(u, "/robots.txt") {
return &crawl.Response{URL: u, ContentType: "text/plain", Body: []byte("")}, nil
}
return &crawl.Response{URL: u, ContentType: ct, Body: []byte(s.body)}, nil
}
func buildWebHandler(c *crawl.Crawler) *reactiveHandler {
return &reactiveHandler{
replier: voice.NewStubReplier(),
phraser: phraser.NewStub(),
crawler: c,
}
}
func askWeb(h *reactiveHandler, q string) (string, bool) {
return h.queryWeb(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: q},
})
}
func TestQueryWebPassesWithoutAURL(t *testing.T) {
h := buildWebHandler(crawl.New(&stubCrawlFetcher{body: "<html><body>x</body></html>"}, crawl.Config{}))
if reply, ok := askWeb(h, "почему небо синее?"); ok {
t.Errorf("the web source claimed a question with no URL: %q", reply)
}
}
// Not configured is said out loud rather than falling through, so a small model
// never invents a page's contents from its URL.
func TestQueryWebSaysWhenNotConfigured(t *testing.T) {
h := buildWebHandler(nil)
reply, ok := askWeb(h, "посмотри https://example.org/page")
if !ok {
t.Fatal("the web source did not claim a question with a URL")
}
if !strings.Contains(reply, "не настроено") {
t.Errorf("reply = %q, want the not-configured answer", reply)
}
}
func TestQueryWebReadsThePage(t *testing.T) {
h := buildWebHandler(crawl.New(&stubCrawlFetcher{
body: "<html><head><title>Заголовок</title></head><body><p>текст страницы</p></body></html>",
}, crawl.Config{}))
reply, ok := askWeb(h, "посмотри https://example.org/page — что там?")
if !ok {
t.Fatal("the web source did not claim a question with a URL")
}
if !strings.Contains(reply, "текст страницы") {
t.Errorf("reply = %q, want the page text read back", reply)
}
}
func TestQueryWebRefusesNonHTML(t *testing.T) {
h := buildWebHandler(crawl.New(&stubCrawlFetcher{
body: "\x00\x01binary", ctype: "application/octet-stream",
}, crawl.Config{}))
reply, ok := askWeb(h, "почитай https://example.org/blob.bin")
if !ok {
t.Fatal("the web source did not claim a question with a URL")
}
if !strings.Contains(reply, "не получилось") {
t.Errorf("reply = %q, want the read-failed answer", reply)
}
}
// robots.txt is honoured on the answer path too, and she says so instead of
// reporting a generic failure.
func TestQueryWebObeysRobots(t *testing.T) {
h := buildWebHandler(crawl.New(&robotsDenyFetcher{}, crawl.Config{}))
reply, ok := askWeb(h, "посмотри https://example.org/private")
if !ok {
t.Fatal("the web source did not claim a question with a URL")
}
if !strings.Contains(reply, "robots.txt") {
t.Errorf("reply = %q, want the robots answer", reply)
}
}
type robotsDenyFetcher struct{}
func (robotsDenyFetcher) Get(_ context.Context, u string) (*crawl.Response, error) {
if strings.HasSuffix(u, "/robots.txt") {
return &crawl.Response{URL: u, ContentType: "text/plain",
Body: []byte("User-agent: *\nDisallow: /private\n")}, nil
}
return &crawl.Response{URL: u, ContentType: "text/html", Body: []byte("<html>nope</html>")}, nil
}
+227
View File
@@ -0,0 +1,227 @@
package main
import (
"context"
"errors"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
)
// planAPI answers only DayPlan; every other call is unimplemented, which is
// exactly the assertion that the plan source needs nothing else.
type planAPI struct {
ipc.UnimplementedCoreAPI
plan ipc.DayPlan
err error
calls int
}
func (a *planAPI) DayPlan(context.Context) (ipc.DayPlan, error) {
a.calls++
if a.err != nil {
return ipc.DayPlan{}, a.err
}
return a.plan, nil
}
func planDay() time.Time { return time.Date(2026, 8, 3, 12, 0, 0, 0, time.UTC) }
func samplePlan() ipc.DayPlan {
day := planDay()
mid := time.Date(2026, 8, 3, 0, 0, 0, 0, time.UTC)
return ipc.DayPlan{
Date: mid,
Items: []ipc.DayPlanItem{
{At: day.Add(-2 * time.Hour), Text: "Standup @ 10:00-10:30", Kind: "event"},
{At: day.Add(2 * time.Hour), Text: "Планёрка @ 14:00-14:30", Kind: "event", Uncertain: true},
{At: day.Add(6 * time.Hour), Text: "позвонить маме", Kind: "reminder"},
},
Spoken: "план на 03.08.2026: 10:00 — Standup @ 10:00-10:30; " +
"похоже, 14:00 — Планёрка @ 14:00-14:30; 18:00 — позвонить маме.",
}
}
func planHandler(api ipc.CoreAPI) *reactiveHandler {
return &reactiveHandler{api: api, now: planDay}
}
func TestQueryDayPlanRecitesTheDay(t *testing.T) {
api := &planAPI{plan: samplePlan()}
h := planHandler(api)
reply, ok := h.queryDayPlan(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "какие планы на сегодня?"},
})
if !ok {
t.Fatal("the plan source must claim a plan question")
}
if reply != api.plan.Spoken {
t.Errorf("reply = %q, want the core's spoken plan %q", reply, api.plan.Spoken)
}
}
// "что дальше?" is the rest of the day, not the whole day: what has already
// happened is not a plan.
func TestQueryDayPlanTrimsToRestOfDay(t *testing.T) {
h := planHandler(&planAPI{plan: samplePlan()})
reply, ok := h.queryDayPlan(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "что дальше?"},
})
if !ok {
t.Fatal("expected the plan source to claim it")
}
if strings.Contains(reply, "Standup") {
t.Errorf("a passed item must not be read back: %q", reply)
}
if !strings.Contains(reply, "Планёрка") || !strings.Contains(reply, "позвонить маме") {
t.Errorf("the rest of the day is missing: %q", reply)
}
// Provenance survives the trim.
if !strings.Contains(reply, "похоже,") {
t.Errorf("a relayed event must stay hedged: %q", reply)
}
}
// A question that is not about the plan must fall through, or the plan buries
// the calendar listing and the weather behind it.
func TestQueryDayPlanPassesOnEverythingElse(t *testing.T) {
for _, q := range []string{
"что у меня сегодня?",
"какие планы на завтра?",
"когда планёрка?",
"какая погода?",
"",
} {
api := &planAPI{plan: samplePlan()}
reply, ok := planHandler(api).queryDayPlan(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: q},
})
if ok {
t.Errorf("%q was claimed by the plan source (reply %q)", q, reply)
}
if api.calls != 0 {
t.Errorf("%q hit the core for a plan it does not want", q)
}
}
}
func TestQueryDayPlanCoreFailure(t *testing.T) {
h := planHandler(&planAPI{err: errors.New("socket closed")})
reply, ok := h.queryDayPlan(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "план на сегодня"},
})
if !ok {
t.Fatal("a failed plan read must still answer, not fall through to RAG")
}
if reply != "не получилось собрать план." {
t.Errorf("reply = %q", reply)
}
}
// The day plan must sit before the calendar listing: both match "…на сегодня",
// and the more specific matcher has to get first refusal (see #373 for what
// happens when the order is wrong).
func TestDayPlanSourcePrecedesCalendar(t *testing.T) {
plan, cal := -1, -1
for i, s := range querySources {
switch s.name {
case "day-plan":
plan = i
case "calendar":
cal = i
}
}
if plan < 0 || cal < 0 {
t.Fatalf("sources missing: day-plan=%d calendar=%d", plan, cal)
}
if plan > cal {
t.Errorf("day-plan at %d must come before calendar at %d", plan, cal)
}
}
// habitAPI answers only RecentFacts — the whole input the behaviour profile
// needs (Vikunja #254). Nothing is asked of the LLM, so nothing else is wired.
type habitAPI struct {
ipc.UnimplementedCoreAPI
facts []ipc.Fact
err error
calls int
}
func (a *habitAPI) RecentFacts(_ context.Context, _ int) ([]ipc.Fact, error) {
a.calls++
return a.facts, a.err
}
// tuesdayFacts — n weekly Tuesday rows for key, ending before now.
func tuesdayFacts(key string, hh, weeks int, now time.Time) []ipc.Fact {
d := now
for d.Weekday() != time.Tuesday {
d = d.AddDate(0, 0, -1)
}
var out []ipc.Fact
for i := 0; i < weeks; i++ {
day := d.AddDate(0, 0, -7*i)
out = append(out, ipc.Fact{
Ts: time.Date(day.Year(), day.Month(), day.Day(), hh, 0, 0, 0, now.Location()),
Kind: "self",
Key: key,
})
}
return out
}
func TestQueryHabitsAnswersFromCountedFacts(t *testing.T) {
now := planDay() // a Monday
api := &habitAPI{facts: tuesdayFacts("workout", 19, 4, now)}
h := &reactiveHandler{api: api, now: func() time.Time { return now }}
reply, ok := h.queryHabits(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: "что я обычно делаю по вторникам?"},
})
if !ok {
t.Fatal("the habit source must claim a habit question")
}
if want := "по вторникам ты обычно тренируешься около 19:00."; reply != want {
t.Errorf("reply = %q, want %q", reply, want)
}
}
func TestQueryHabitsPassesOnEverythingElse(t *testing.T) {
now := planDay()
for _, q := range []string{"что я делаю в среду?", "что у меня сегодня?", "какие планы на сегодня?", ""} {
api := &habitAPI{}
h := &reactiveHandler{api: api, now: func() time.Time { return now }}
if reply, ok := h.queryHabits(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: q},
}); ok {
t.Errorf("%q was claimed by the habit source (reply %q)", q, reply)
}
if api.calls != 0 {
t.Errorf("%q scanned the fact log for a profile it does not want", q)
}
}
}
// Both specific sources must precede the calendar listing, which matches any
// utterance naming a day.
func TestHabitSourcePrecedesCalendar(t *testing.T) {
habits, cal := -1, -1
for i, s := range querySources {
switch s.name {
case "habits":
habits = i
case "calendar":
cal = i
}
}
if habits < 0 || cal < 0 {
t.Fatalf("sources missing: habits=%d calendar=%d", habits, cal)
}
if habits > cal {
t.Errorf("habits at %d must come before calendar at %d", habits, cal)
}
}
+158
View File
@@ -0,0 +1,158 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/store"
)
// Vikunja #281 — the fourth delivery outcome: a care candidate the restraint
// gate suppresses (quiet hours / away / calendar-busy) is not necessarily
// lost. If it's worth resurfacing (loop.DigestEligible), it's durably held
// (internal/store's digest_entries) and spoken as one bundle once speaking
// is appropriate again — never while the suppression reason still holds.
func breakTrace(blockedBy string) *loop.TickTrace {
return &loop.TickTrace{
RuleTraces: []loop.RuleTrace{{
RuleName: "break",
Severity: loop.Sev2,
PredicateResult: true,
GateResult: false,
GateBlockedBy: blockedBy,
}},
}
}
// TestSuppressedCareDigestsAcrossQuietHours — a Sev2 care candidate blocked
// by quiet hours is enqueued into the durable digest, and is spoken as a
// "digest" nudge only once quiet hours actually end — never while still
// suppressed (that would just be a second way to nag through quiet hours).
func TestSuppressedCareDigestsAcrossQuietHours(t *testing.T) {
st := newTestStore(t)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
ctx := context.Background()
now := refNow()
quiet := loop.State{Now: now, QuietHours: true, Presence: store.Present}
tl.enqueueSuppressedDigest(ctx, breakTrace("quiet_hours"), quiet, now)
entries, err := st.PendingDigestEntries(ctx, now)
if err != nil {
t.Fatalf("pending: %v", err)
}
if len(entries) != 1 || entries[0].Rule != "break" {
t.Fatalf("want 1 pending digest entry for break, got %+v", entries)
}
// still quiet hours: draining now must not speak — the same restraint
// that suppressed the live nudge must suppress the bundle too.
tl.maybeDrainDigest(ctx, quiet, now)
if len(sink.sends) != 0 {
t.Fatalf("digest must not drain while quiet hours holds, got %+v", sink.sends)
}
// quiet hours end: this is the moment speaking is appropriate again.
after := now.Add(time.Hour)
clear := loop.State{Now: after, QuietHours: false, Presence: store.Present}
tl.maybeDrainDigest(ctx, clear, after)
if len(sink.sends) != 1 {
t.Fatalf("want exactly 1 dispatched digest bundle, got %d: %+v", len(sink.sends), sink.sends)
}
if sink.sends[0].RuleName != "digest" {
t.Fatalf("want RuleName digest, got %q", sink.sends[0].RuleName)
}
remaining, err := st.PendingDigestEntries(ctx, after)
if err != nil {
t.Fatalf("pending after drain: %v", err)
}
if len(remaining) != 0 {
t.Fatalf("drained entry must no longer be pending, got %+v", remaining)
}
}
// TestSuppressedCareDigestDedupesAcrossTicks — quiet hours holding for
// several ticks must not enqueue several copies of the same suppressed
// nudge; he hears it once when the bundle finally drains.
func TestSuppressedCareDigestDedupesAcrossTicks(t *testing.T) {
st := newTestStore(t)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
ctx := context.Background()
now := refNow()
quiet := loop.State{Now: now, QuietHours: true, Presence: store.Present}
for i := 0; i < 3; i++ {
tl.enqueueSuppressedDigest(ctx, breakTrace("quiet_hours"), quiet, now.Add(time.Duration(i)*time.Minute))
}
entries, err := st.PendingDigestEntries(ctx, now)
if err != nil {
t.Fatalf("pending: %v", err)
}
if len(entries) != 1 {
t.Fatalf("3 suppressions of the same nudge must collapse to 1 pending entry, got %d", len(entries))
}
}
// TestSuppressedCareDigestExpiresRatherThanDeliveringLate — an entry that
// aged out before the suppression cleared is dropped, not spoken late: a
// two-day-old "you skipped a break" is noise, not news.
func TestSuppressedCareDigestExpiresRatherThanDeliveringLate(t *testing.T) {
st := newTestStore(t)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
ctx := context.Background()
now := refNow()
quiet := loop.State{Now: now, QuietHours: true, Presence: store.Present}
tl.enqueueSuppressedDigest(ctx, breakTrace("quiet_hours"), quiet, now)
// well past digestExpiry (24h) before the suppression ever clears.
stale := now.Add(48 * time.Hour)
tl.expireStaleDigest(ctx, stale)
clear := loop.State{Now: stale, QuietHours: false, Presence: store.Present}
tl.maybeDrainDigest(ctx, clear, stale)
if len(sink.sends) != 0 {
t.Fatalf("a stale digest entry must be dropped, not delivered late; got %+v", sink.sends)
}
}
// TestSuppressedCareDigestIgnoresHighSeverity — defense in depth at the
// wiring layer: even if a RuleTrace somehow showed a high-severity rule
// blocked by a care-only gate reason, the tick driver must not durably
// digest it. Alarms bypass the gate and deliver now, unchanged; they must
// never be silently delayed into a bundle.
func TestSuppressedCareDigestIgnoresHighSeverity(t *testing.T) {
st := newTestStore(t)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
ctx := context.Background()
now := refNow()
trace := &loop.TickTrace{RuleTraces: []loop.RuleTrace{{
RuleName: "service_down",
Severity: loop.Sev4,
PredicateResult: true,
GateResult: false,
GateBlockedBy: "quiet_hours",
}}}
quiet := loop.State{Now: now, QuietHours: true, Presence: store.Present}
tl.enqueueSuppressedDigest(ctx, trace, quiet, now)
entries, err := st.PendingDigestEntries(ctx, now)
if err != nil {
t.Fatalf("pending: %v", err)
}
if len(entries) != 0 {
t.Fatalf("high severity must never be digested, got %+v", entries)
}
}
+413
View File
@@ -0,0 +1,413 @@
package main
import (
"context"
"encoding/json"
"fmt"
"log"
"strings"
hexisclient "github.com/kami/hexis/pkg/client"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
)
// praxisCapability is one arm of the Praxis act dispatch. This is an interface
// rather than a map[string]func because each arm carries its own state: the
// verb aliases it answers to, the trace name it records, and its own reply
// formatting. The dispatch grows an arm per Praxis capability, so a new one is
// added to praxisCapabilities below and nothing else changes.
type praxisCapability interface {
// aliases are the verbs (router fn slots, EN and RU) this capability answers to.
aliases() []string
// handle runs the capability and returns the user-facing reply.
handle(ctx context.Context, h *reactiveHandler, px *praxisClient, dec router.Decision) string
}
// praxisCapabilities is the registry handlePraxisAct consults, in order.
var praxisCapabilities = []praxisCapability{
listAttentionCapability{},
praxisItemAction{
verbs: []string{"acknowledge_item", "принято", "понял", "поняла"},
ask: "какой пункт отметить принятым?",
op: "acknowledge",
failure: "не получилось отметить принятым.",
success: "принято.",
call: func(ctx context.Context, px *praxisClient, id string) error {
_, err := px.Acknowledge(ctx, id)
return err
},
},
praxisItemAction{
verbs: []string{"resolve_item", "сделано", "готово", "решено"},
ask: "какой пункт отметить сделанным?",
op: "resolve",
failure: "не получилось отметить сделанным.",
success: "отмечено как сделано.",
call: func(ctx context.Context, px *praxisClient, id string) error {
_, err := px.Resolve(ctx, id)
return err
},
},
praxisItemAction{
verbs: []string{"ignore_item", "игнорировать", "неважно"},
ask: "какой пункт игнорировать?",
op: "ignore",
failure: "не получилось проигнорировать.",
success: "проигнорировано.",
call: func(ctx context.Context, px *praxisClient, id string) error {
_, err := px.Ignore(ctx, id)
return err
},
},
praxisItemAction{
verbs: []string{"pin_item", "закрепить"},
ask: "какой пункт закрепить?",
op: "pin",
failure: "не получилось закрепить.",
success: "закреплено.",
call: func(ctx context.Context, px *praxisClient, id string) error {
_, err := px.Pin(ctx, id, true)
return err
},
},
listChangesCapability{},
entityAttentionCapability{},
}
// handlePraxisAct — dispatches ecosystem tool acts through the Praxis tools API.
// Returns "" when the act is not a Praxis verb (the caller falls through to the
// system command executor). Returns a reply string otherwise.
func (h *reactiveHandler) handlePraxisAct(ctx context.Context, dec router.Decision) string {
if h.ecosystem == nil || h.ecosystem.praxis == nil {
return ""
}
px := h.ecosystem.praxis
for _, capability := range praxisCapabilities {
for _, alias := range capability.aliases() {
if alias == dec.Slots.Fn {
return capability.handle(ctx, h, px, dec)
}
}
}
// Not a Praxis verb — let the caller fall through.
return ""
}
// praxisItemAction is the shared shape of the item-lifecycle capabilities: take
// an item id from the value slot, call one Praxis endpoint, trace the result.
type praxisItemAction struct {
verbs []string
ask string // reply when no item id was given
op string // trace + log name of the operation
failure string // reply when the Praxis call errors
success string
call func(ctx context.Context, px *praxisClient, id string) error
}
func (a praxisItemAction) aliases() []string { return a.verbs }
func (a praxisItemAction) handle(ctx context.Context, h *reactiveHandler, px *praxisClient, dec router.Decision) string {
id := dec.Slots.Value
if id == "" {
return a.ask
}
if err := a.call(ctx, px, id); err != nil {
log.Printf("ecosystem: praxis %s %s: %v", a.op, id, err)
return a.failure
}
h.recordPraxisTrace(ctx, a.op, map[string]any{"item_id": id})
return a.success
}
// listAttentionCapability reads the attention digest and surfaces every item it speaks.
type listAttentionCapability struct{}
func (listAttentionCapability) aliases() []string {
return []string{"list_attention", "attention", "внимание", "что требует внимания", "что нового"}
}
func (listAttentionCapability) handle(ctx context.Context, h *reactiveHandler, px *praxisClient, _ router.Decision) string {
items, err := px.ListAttention(ctx, 20)
if err != nil {
log.Printf("ecosystem: praxis attention: %v", err)
return "не могу сейчас узнать, что требует внимания."
}
if len(items) == 0 {
return "ничего не требует внимания."
}
h.recordPraxisTrace(ctx, "list_attention", map[string]any{"count": len(items)})
var parts []string
for _, item := range items {
title, _ := item["title"].(string)
// importance arrives as JSON number ⇒ float64 over the HTTP contract.
importance, _ := item["importance"].(float64)
rule, _ := item["rule"].(string)
s := title
if importance > 0 {
s += fmt.Sprintf(" (важность %d", int(importance))
if rule != "" {
s += ": " + rule
}
s += ")"
}
parts = append(parts, s)
// Speaking an item surfaces it, it does not acknowledge it
// (ECOSYSTEM-SPEC.md §2.3: surfaced != acknowledged). Best-effort:
// a failed surface call must not block delivering the digest.
if id, ok := item["id"].(string); ok && id != "" {
if _, err := px.Surface(ctx, id); err != nil {
log.Printf("ecosystem: praxis surface %s: %v", id, err)
}
}
}
return "требует внимания: " + strings.Join(parts, "; ")
}
// listChangesCapability reads the recent-changes feed.
type listChangesCapability struct{}
func (listChangesCapability) aliases() []string {
return []string{"list_changes", "changes", "изменения", "что изменилось"}
}
func (listChangesCapability) handle(ctx context.Context, h *reactiveHandler, px *praxisClient, _ router.Decision) string {
changes, err := px.ListChanges(ctx, 20)
if err != nil {
log.Printf("ecosystem: praxis changes: %v", err)
return "не могу сейчас узнать об изменениях."
}
if len(changes) == 0 {
return "нет изменений."
}
h.recordPraxisTrace(ctx, "list_changes", map[string]any{"count": len(changes)})
var parts []string
for _, c := range changes {
title, _ := c["title"].(string)
typ, _ := c["change_type"].(string)
parts = append(parts, fmt.Sprintf("%s (%s)", title, typ))
}
return "изменения: " + strings.Join(parts, "; ")
}
// entityAttentionCapability answers "what's going on with X" by resolving X to
// a canonical Nexus entity and asking Praxis for that entity's attention items
// (Vikunja #272). Unlike listAttentionCapability it is scoped: the entity_id
// travels to Praxis as a query parameter instead of Maven filtering an unscoped
// list client-side, which is what makes the ref canonical end to end.
//
// It also folds in what Maven herself knows about the same entity — facts the
// enrichment worker has already resolved to that entity_id — so one question
// gets one answer across both stores.
type entityAttentionCapability struct{}
func (entityAttentionCapability) aliases() []string {
return []string{"entity_attention", "что с", "как дела у", "статус"}
}
func (entityAttentionCapability) handle(ctx context.Context, h *reactiveHandler, px *praxisClient, dec router.Decision) string {
subject := dec.Slots.Value
if subject == "" {
subject = dec.Slots.Text
}
if subject == "" {
return "про что именно спросить?"
}
if h.ecosystem == nil || h.ecosystem.nexus == nil {
// Without Nexus there is no canonical ref to scope by. Say so rather
// than quietly answering about something else.
return "не могу связать это с сущностью — Nexus не настроен."
}
entityID, displayName, ambiguous, err := h.ecosystem.resolveEntityReference(ctx, subject, nil)
if err != nil {
log.Printf("ecosystem: entity attention resolve %q: %v", subject, err)
return "экосистема недоступна, попробуй ещё раз."
}
if len(ambiguous) > 0 {
return "уточни, что именно: " + strings.Join(ambiguous, ", ") + "?"
}
if entityID == "" {
return "не знаю такой сущности."
}
if displayName == "" {
displayName = subject
}
items, err := px.ListAttentionForEntity(ctx, entityID, 20)
if err != nil {
log.Printf("ecosystem: praxis attention for %s: %v", entityID, err)
return "не могу сейчас узнать, что требует внимания по «" + displayName + "»."
}
h.recordPraxisTrace(ctx, "entity_attention", map[string]any{
"entity_id": entityID, "count": len(items),
})
var parts []string
for _, item := range items {
title, _ := item["title"].(string)
if title == "" {
continue
}
parts = append(parts, title)
// Same surfaced != acknowledged rule as the unscoped digest.
if id, ok := item["id"].(string); ok && id != "" {
if _, err := px.Surface(ctx, id); err != nil {
log.Printf("ecosystem: praxis surface %s: %v", id, err)
}
}
}
if known := h.localFactsForEntity(ctx, entityID); known != "" {
parts = append(parts, known)
}
if len(parts) == 0 {
return "по «" + displayName + "» ничего нет."
}
return "по «" + displayName + "»: " + strings.Join(parts, "; ")
}
// localFactsForEntity summarises Maven's own facts already resolved to this
// canonical entity. Empty when the store is unavailable or nothing matched —
// entity-scoped memory is an enrichment of the answer, never a precondition.
func (h *reactiveHandler) localFactsForEntity(ctx context.Context, entityID string) string {
if h.dataStore == nil || entityID == "" {
return ""
}
facts, err := h.dataStore.FactsByEntity(ctx, entityID, 3)
if err != nil {
log.Printf("ecosystem: facts by entity %s: %v", entityID, err)
return ""
}
var parts []string
for _, f := range facts {
if f.Value != "" {
parts = append(parts, f.Value)
}
}
if len(parts) == 0 {
return ""
}
return "я помню: " + strings.Join(parts, ", ")
}
// recordPraxisTrace — writes a fact recording a cross-service ecosystem call.
// The fact is stored with source "praxis:trace" so the proactive loop can
// reference it and the dashboard can display recent ecosystem activity.
func (h *reactiveHandler) recordPraxisTrace(ctx context.Context, operation string, details map[string]any) {
now := h.now()
value := operation
if len(details) > 0 {
if b, err := json.Marshal(details); err == nil {
value = operation + " " + string(b)
}
}
_, _ = h.api.WriteFact(ctx, ipc.WriteFactReq{
Ts: now,
Kind: "system",
Key: "praxis:" + operation,
Value: value,
Source: "praxis:trace",
Confidence: 1.0,
})
}
// handleHexisAct — resolves entity references through Nexus and executes
// matching capabilities through Hexis. Returns a reply string when handled,
// or "" to fall through to the system command executor.
func (h *reactiveHandler) handleHexisAct(ctx context.Context, dec router.Decision) string {
if h.ecosystem == nil {
return ""
}
// Resolve the utterance text as an entity reference through Nexus. An
// ambiguous match must stop and clarify — never guess a mutation target.
entityID, displayName, ambiguous, err := h.ecosystem.resolveEntityReference(ctx, dec.Slots.Text, nil)
if err != nil {
// A genuine Nexus dependency failure, not "no such entity" — stop here
// and report degradation rather than silently falling through to the
// local command executor (ECOSYSTEM-SPEC.md: services degrade
// independently, never a silent all-clear).
return "экосистема недоступна, попробуй ещё раз."
}
if len(ambiguous) > 0 {
return "уточни, что именно: " + strings.Join(ambiguous, ", ") + "?"
}
if entityID == "" {
return ""
}
// Discover Hexis capabilities for this entity. A resolved entity with a
// genuine Hexis failure must not be treated as "no capabilities" and
// fall through to unrelated local execution.
caps, err := h.ecosystem.discoverCapabilities(ctx, entityID)
if err != nil {
return "экосистема недоступна, попробуй ещё раз."
}
if len(caps) == 0 {
return ""
}
// Match the user's verb to a capability by name/description. Collect all
// matches: more than one is itself ambiguous, so we ask rather than pick
// the first (ecosystem invariant: no arbitrary target for mutation).
verb := dec.Slots.Fn
if verb == "" {
verb = dec.Slots.Text
}
verbLower := strings.ToLower(verb)
var matches []*hexisclient.Capability
for i, c := range caps {
if strings.Contains(strings.ToLower(c.Name), verbLower) ||
(c.Description != "" && strings.Contains(strings.ToLower(c.Description), verbLower)) {
matches = append(matches, &caps[i])
}
}
if len(matches) == 0 {
return ""
}
if len(matches) > 1 {
var names []string
for _, m := range matches {
names = append(names, m.Name)
}
return "какую команду для " + displayName + ": " + strings.Join(names, ", ") + "?"
}
matched := matches[0]
// Read-only capabilities run immediately; mutating ones are parked for an
// explicit spoken confirm bound to this capability + target.
if !matched.ReadOnly {
h.mu.Lock()
h.pendingHexis = &pendingHexisExec{
capabilityID: matched.ID,
capName: matched.Name,
entityID: entityID,
displayName: displayName,
expiry: h.now().Add(confirmTTL),
}
h.mu.Unlock()
return "выполнить «" + matched.Name + "» для " + displayName + "? скажи «да» или «нет»."
}
return h.execHexis(ctx, matched.ID, matched.Name, entityID, displayName)
}
// execHexis runs a resolved capability and records a cross-service trace with
// the correlation ID. It reports command success, never operational recovery
// (Praxis observes recovery independently).
func (h *reactiveHandler) execHexis(ctx context.Context, capID, capName, entityID, displayName string) string {
correlationID, err := h.ecosystem.executeCapability(ctx, capID, entityID, nil)
if err != nil {
log.Printf("ecosystem: hexis execute error (cor=%s): %v", correlationID, err)
return "не получилось выполнить команду для " + displayName + "."
}
h.recordPraxisTrace(ctx, "hexis:"+capName, map[string]any{
"entity_id": entityID,
"entity_name": displayName,
"capability": capName,
"correlation_id": correlationID,
})
return "команда выполнена для " + displayName + "."
}
+312
View File
@@ -0,0 +1,312 @@
package main
import (
"context"
"net/http"
"strings"
"testing"
"time"
hexisclient "github.com/kami/hexis/pkg/client"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/store"
)
// Phase-5 hardening suite (Vikunja #276). Everything here drives the shared
// fake ecosystem (fakeecosystem_test.go) rather than one-off inline handlers,
// so the same fault levers — SetFault, SetBody, SetDelay — cover every
// service. What is asserted is the degraded-mode contract:
//
// - services degrade independently: one outage never mutes the others,
// - a degraded reply is never silent, never fabricated, never "success",
// - contract drift (old shape, unknown fields, garbage) is survivable,
// - Maven never acts on an ambiguous target and never chains
// Praxis observation into Hexis execution on its own.
// ecoHandler wires a handler against whichever of the three fakes is given
// (pass nil to leave a service unconfigured, which is a different state from
// "configured but down").
func ecoHandler(t *testing.T, nexus, praxis, hexis *fakeServer) *reactiveHandler {
t.Helper()
st := newTestStore(t)
clock := newFakeClock(time.Date(2026, 8, 1, 9, 0, 0, 0, time.UTC))
w := &ecosystemWiring{}
if nexus != nil {
w.nexus = newNexusClient(nexus.URL)
}
if praxis != nil {
w.praxis = newPraxisClient(praxis.URL)
}
if hexis != nil {
w.hexis = hexisclient.New(hexis.URL)
}
return &reactiveHandler{
api: ipc.NewStoreAPI(st),
dataStore: st,
now: clock.Now,
ecosystem: w,
}
}
func traceFacts(t *testing.T, h *reactiveHandler) []store.Fact {
t.Helper()
facts, err := h.dataStore.RecentFacts(context.Background(), 50)
if err != nil {
t.Fatalf("read facts: %v", err)
}
var out []store.Fact
for _, f := range facts {
if f.Source == "praxis:trace" {
out = append(out, f)
}
}
return out
}
func restartCaps() string {
return fixtureHexisCapabilities(map[string]any{
"id": "cap_restart", "name": "restart", "read_only": true,
})
}
// TestEcosystem_OutagesAreIndependent: Praxis being down must not disable the
// Nexus+Hexis action path, and vice versa. A shared "ecosystem is broken"
// mode would take away working capability for no reason.
func TestEcosystem_OutagesAreIndependent(t *testing.T) {
ctx := context.Background()
nexus := newFakeNexus(t, fixtureNexusResolved("ent_muzick", "Muzick indexer", "service"))
praxis := newFakePraxis(t, fixturePraxisAttentionItems(map[string]any{
"id": "item_1", "title": "disk almost full", "importance": 3.0,
}))
hexis := newFakeHexis(t, restartCaps(), fixtureHexisExecuted("exec_1", "succeeded"))
h := ecoHandler(t, nexus, praxis, hexis)
praxis.SetFault(503)
if reply := h.handleHexisAct(ctx, actDec("muzick indexer")); !strings.Contains(reply, "выполнена") {
t.Fatalf("praxis outage must not block the hexis path, got %q", reply)
}
praxis.SetFault(0)
hexis.SetFault(503)
nexus.SetFault(503)
reply := h.handlePraxisAct(ctx, praxisActDec("list_attention"))
if !strings.Contains(reply, "disk almost full") {
t.Fatalf("nexus/hexis outage must not block the praxis digest, got %q", reply)
}
}
// TestEcosystem_MalformedNexusResponseFailsClosed: a 200 carrying garbage is a
// dependency failure, not "no such entity". It must stop before Hexis.
func TestEcosystem_MalformedNexusResponseFailsClosed(t *testing.T) {
ctx := context.Background()
nexus := newFakeNexus(t, fixtureNexusResolved("ent_muzick", "Muzick indexer", "service"))
hexis := newFakeHexis(t, restartCaps(), fixtureHexisExecuted("exec_1", "succeeded"))
h := ecoHandler(t, nexus, nil, hexis)
nexus.SetBody(`{"status":"resolved","entity":`)
reply := h.handleHexisAct(ctx, actDec("muzick indexer"))
if reply == "" || strings.Contains(reply, "выполнена") {
t.Fatalf("malformed nexus body must degrade, got %q", reply)
}
if hexis.Count("", "/api/v1") != 0 {
t.Fatal("hexis must not be contacted after a malformed nexus response")
}
}
// TestEcosystem_UnknownContractFieldsTolerated: a newer Nexus adding fields
// must not break an older Maven. Same for the older flat resolve shape.
func TestEcosystem_UnknownContractFieldsTolerated(t *testing.T) {
ctx := context.Background()
for name, body := range map[string]string{
"future": fixtureNexusResolvedFuture("ent_muzick", "Muzick indexer", "service"),
"flat": fixtureNexusResolvedFlat("ent_muzick", "Muzick indexer", "service"),
} {
t.Run(name, func(t *testing.T) {
nexus := newFakeNexus(t, body)
hexis := newFakeHexis(t, restartCaps(), fixtureHexisExecuted("exec_1", "succeeded"))
h := ecoHandler(t, nexus, nil, hexis)
if reply := h.handleHexisAct(ctx, actDec("muzick indexer")); !strings.Contains(reply, "выполнена") {
t.Fatalf("%s contract shape must still resolve and execute, got %q", name, reply)
}
})
}
}
// TestEcosystem_CancelledContextDegrades: a caller hanging up (turn abandoned,
// deadline hit) must surface as degradation, never as a fabricated result.
func TestEcosystem_CancelledContextDegrades(t *testing.T) {
nexus := newFakeNexus(t, fixtureNexusResolved("ent_muzick", "Muzick indexer", "service"))
hexis := newFakeHexis(t, restartCaps(), fixtureHexisExecuted("exec_1", "succeeded"))
h := ecoHandler(t, nexus, nil, hexis)
nexus.SetDelay(2 * time.Second)
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Millisecond)
defer cancel()
reply := h.handleHexisAct(ctx, actDec("muzick indexer"))
if reply == "" || strings.Contains(reply, "выполнена") {
t.Fatalf("cancelled resolve must degrade, got %q", reply)
}
if hexis.Count("", "/api/v1") != 0 {
t.Fatal("hexis must not be contacted after a cancelled resolve")
}
}
// TestEcosystem_ExecutionFailureIsNotSuccess: Hexis answering 200 with
// status=failed is a partial failure — the call worked, the command did not.
// Maven must report it as a failure and must not write a success trace.
func TestEcosystem_ExecutionFailureIsNotSuccess(t *testing.T) {
ctx := context.Background()
nexus := newFakeNexus(t, fixtureNexusResolved("ent_muzick", "Muzick indexer", "service"))
hexis := newFakeHexis(t, restartCaps(), fixtureHexisExecutionFailed("exec_1", "unit not found"))
h := ecoHandler(t, nexus, nil, hexis)
reply := h.handleHexisAct(ctx, actDec("muzick indexer"))
if strings.Contains(reply, "выполнена") {
t.Fatalf("failed execution must not read as success, got %q", reply)
}
if reply == "" {
t.Fatal("failed execution must say something")
}
for _, f := range traceFacts(t, h) {
if strings.HasPrefix(f.Key, "praxis:hexis:") {
t.Fatalf("failed execution must not write a success trace: %+v", f)
}
}
}
// TestEcosystem_AmbiguousTargetBlocksExecution: ambiguity blocks mutation, and
// the clarification must name the candidates rather than pick one.
func TestEcosystem_AmbiguousTargetBlocksExecution(t *testing.T) {
ctx := context.Background()
nexus := newFakeNexus(t, fixtureNexusAmbiguous(
map[string]string{"entity_id": "ent_a", "display_name": "Muzick indexer"},
map[string]string{"entity_id": "ent_b", "display_name": "Muzick web"},
))
hexis := newFakeHexis(t, restartCaps(), fixtureHexisExecuted("exec_1", "succeeded"))
h := ecoHandler(t, nexus, nil, hexis)
reply := h.handleHexisAct(ctx, actDec("muzick"))
if !strings.Contains(reply, "Muzick indexer") || !strings.Contains(reply, "Muzick web") {
t.Fatalf("ambiguous resolve must list candidates, got %q", reply)
}
if hexis.Count("POST", "/api/v1/execute") != 0 {
t.Fatal("ambiguous target must never execute")
}
}
// TestEcosystem_NoAutonomousPraxisToHexis: reading the attention digest is an
// observation. Maven must never turn an observed problem into a Hexis command
// by herself — she is not autonomous.
func TestEcosystem_NoAutonomousPraxisToHexis(t *testing.T) {
ctx := context.Background()
praxis := newFakePraxis(t, fixturePraxisAttentionItems(
map[string]any{"id": "item_1", "title": "muzick indexer is down", "importance": 4.0, "rule": "service_down"},
))
nexus := newFakeNexus(t, fixtureNexusResolved("ent_muzick", "Muzick indexer", "service"))
hexis := newFakeHexis(t, restartCaps(), fixtureHexisExecuted("exec_1", "succeeded"))
h := ecoHandler(t, nexus, praxis, hexis)
_ = h.handlePraxisAct(ctx, praxisActDec("list_attention"))
if hexis.Count("", "/api/v1") != 0 {
t.Fatal("attention digest must not contact hexis on its own")
}
if nexus.Count("", "/api/v1/resolve") != 0 {
t.Fatal("attention digest must not resolve targets for autonomous action")
}
}
// TestEcosystem_MutatingCapabilityWaitsForConfirmation: a non-read-only
// capability parks for an explicit spoken confirm bound to capability+target.
func TestEcosystem_MutatingCapabilityWaitsForConfirmation(t *testing.T) {
ctx := context.Background()
nexus := newFakeNexus(t, fixtureNexusResolved("ent_muzick", "Muzick indexer", "service"))
caps := fixtureHexisCapabilities(map[string]any{"id": "cap_restart", "name": "restart", "read_only": false})
hexis := newFakeHexis(t, caps, fixtureHexisExecuted("exec_1", "succeeded"))
h := ecoHandler(t, nexus, nil, hexis)
reply := h.handleHexisAct(ctx, actDec("restart"))
if !strings.Contains(reply, "restart") || !strings.Contains(reply, "да") {
t.Fatalf("mutating capability must ask for confirmation, got %q", reply)
}
if hexis.Count("POST", "/api/v1/execute") != 0 {
t.Fatal("mutating capability must not execute before confirmation")
}
h.mu.Lock()
pending := h.pendingHexis
h.mu.Unlock()
if pending == nil || pending.capabilityID != "cap_restart" || pending.entityID != "ent_muzick" {
t.Fatalf("confirmation must be bound to capability+target, got %+v", pending)
}
}
// TestEcosystem_SurfaceFailureStillDelivers: surfacing is bookkeeping. If the
// surface call fails the digest must still be spoken — a partial failure
// downgrades bookkeeping, not the answer.
func TestEcosystem_SurfaceFailureStillDelivers(t *testing.T) {
ctx := context.Background()
praxis := newFakeServer(t, map[string]http.HandlerFunc{
"GET /api/v1/tools/attention": jsonHandler(200, fixturePraxisAttentionItems(
map[string]any{"id": "item_1", "title": "disk almost full", "importance": 3.0},
)),
"POST /api/v1/tools/surface": jsonHandler(500, `{"error":"boom"}`),
})
h := ecoHandler(t, nil, praxis, nil)
reply := h.handlePraxisAct(ctx, praxisActDec("list_attention"))
if !strings.Contains(reply, "disk almost full") {
t.Fatalf("failed surface must not swallow the digest, got %q", reply)
}
if praxis.Count("POST", "/api/v1/tools/surface") == 0 {
t.Fatal("expected the surface attempt")
}
}
// TestEcosystem_TotalOutageSaysSoForEveryPath: with all three down, every
// entry point degrades explicitly instead of returning empty or inventing.
func TestEcosystem_TotalOutageSaysSoForEveryPath(t *testing.T) {
ctx := context.Background()
nexus := newFakeNexus(t, fixtureNexusResolved("ent_muzick", "Muzick indexer", "service"))
praxis := newFakePraxis(t, fixturePraxisAttentionItems())
hexis := newFakeHexis(t, restartCaps(), fixtureHexisExecuted("exec_1", "succeeded"))
for _, fs := range []*fakeServer{nexus, praxis, hexis} {
fs.SetFault(503)
}
h := ecoHandler(t, nexus, praxis, hexis)
for name, reply := range map[string]string{
"hexis act": h.handleHexisAct(ctx, actDec("muzick indexer")),
"attention": h.handlePraxisAct(ctx, praxisActDec("list_attention")),
"changes": h.handlePraxisAct(ctx, praxisActDec("list_changes")),
"acknowledge": h.handlePraxisAct(ctx, praxisActDec("acknowledge_item")),
} {
if reply == "" {
t.Errorf("%s: total outage must not answer with silence", name)
}
if strings.Contains(reply, "выполнена") {
t.Errorf("%s: total outage must not claim success: %q", name, reply)
}
}
if len(traceFacts(t, h)) != 0 {
t.Fatal("a total outage must not leave success traces behind")
}
}
// TestEcosystem_RecoveryAfterOutageNeedsNoRestart: once the dependency comes
// back the very next turn works — no cached failure state, no restart.
func TestEcosystem_RecoveryAfterOutageNeedsNoRestart(t *testing.T) {
ctx := context.Background()
praxis := newFakePraxis(t, fixturePraxisAttentionItems(
map[string]any{"id": "item_1", "title": "disk almost full", "importance": 3.0},
))
h := ecoHandler(t, nil, praxis, nil)
praxis.SetFault(503)
if reply := h.handlePraxisAct(ctx, praxisActDec("list_attention")); strings.Contains(reply, "disk") {
t.Fatalf("outage must not serve content, got %q", reply)
}
praxis.SetFault(0)
if reply := h.handlePraxisAct(ctx, praxisActDec("list_attention")); !strings.Contains(reply, "disk almost full") {
t.Fatalf("recovery must work on the next turn, got %q", reply)
}
}
+213
View File
@@ -0,0 +1,213 @@
package main
import (
"context"
"database/sql"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
)
// Entity-ref propagation, Maven side (Vikunja #272): the canonical Nexus
// entity_id must reach Praxis as a query scope rather than being resolved and
// then thrown away, and the enrichment that produces those ids must degrade
// visibly instead of silently.
func entityAttentionDec(subject string) router.Decision {
return router.Decision{
Intent: router.IntentAct,
Slots: router.Slots{Fn: "entity_attention", HasFn: true, Value: subject},
}
}
// TestEntityAttention_ScopesPraxisByCanonicalID: the resolved id must travel
// to Praxis in the request, not be used for client-side filtering.
func TestEntityAttention_ScopesPraxisByCanonicalID(t *testing.T) {
ctx := context.Background()
nexus := newFakeNexus(t, fixtureNexusResolved("ent_muzick", "Muzick indexer", "service"))
praxis := newFakePraxis(t, fixturePraxisAttentionItems(
map[string]any{"id": "item_1", "title": "indexer queue is backing up", "importance": 3.0},
))
h := ecoHandler(t, nexus, praxis, nil)
reply := h.handlePraxisAct(ctx, entityAttentionDec("muzick indexer"))
if !strings.Contains(reply, "indexer queue is backing up") {
t.Fatalf("expected the scoped item in the reply, got %q", reply)
}
var scoped bool
for _, r := range praxis.Requests() {
if r.Method == "GET" && strings.HasPrefix(r.Path, "/api/v1/tools/attention") &&
strings.Contains(r.Query, "entity_id=ent_muzick") {
scoped = true
}
}
if !scoped {
t.Fatalf("expected attention scoped by entity_id, got requests %+v", praxis.Requests())
}
if praxis.Count("POST", "/api/v1/tools/surface") == 0 {
t.Error("a spoken scoped item must be surfaced, like the unscoped digest")
}
}
// TestEntityAttention_FoldsInLocalFactsForSameEntity: facts the enrichment
// worker already tagged with the same canonical id join the same answer.
func TestEntityAttention_FoldsInLocalFactsForSameEntity(t *testing.T) {
ctx := context.Background()
nexus := newFakeNexus(t, fixtureNexusResolved("ent_espresso", "the espresso machine", "device"))
praxis := newFakePraxis(t, fixturePraxisAttentionItems())
h := ecoHandler(t, nexus, praxis, nil)
id, err := h.dataStore.WriteFactAboutSubject(ctx, time.Now(), store.KindEnv,
"descaled", "the espresso machine", "descaled in june", "infer:pref", 0.8, sql.NullInt64{})
if err != nil {
t.Fatalf("WriteFactAboutSubject: %v", err)
}
if err := h.dataStore.ResolveFactEntity(ctx, id, "ent_espresso", store.ResolutionResolved); err != nil {
t.Fatalf("ResolveFactEntity: %v", err)
}
reply := h.handlePraxisAct(ctx, entityAttentionDec("the espresso machine"))
if !strings.Contains(reply, "descaled in june") {
t.Fatalf("expected entity-scoped local facts in the reply, got %q", reply)
}
}
// TestEntityAttention_AmbiguousAsksInsteadOfGuessing.
func TestEntityAttention_AmbiguousAsksInsteadOfGuessing(t *testing.T) {
ctx := context.Background()
nexus := newFakeNexus(t, fixtureNexusAmbiguous(
map[string]string{"entity_id": "ent_a", "display_name": "Muzick indexer"},
map[string]string{"entity_id": "ent_b", "display_name": "Muzick web"},
))
praxis := newFakePraxis(t, fixturePraxisAttentionItems())
h := ecoHandler(t, nexus, praxis, nil)
reply := h.handlePraxisAct(ctx, entityAttentionDec("muzick"))
if !strings.Contains(reply, "Muzick indexer") || !strings.Contains(reply, "Muzick web") {
t.Fatalf("ambiguous subject must ask, got %q", reply)
}
if praxis.Count("GET", "/api/v1/tools/attention") != 0 {
t.Fatal("an ambiguous subject must not be queried against praxis")
}
}
// TestEntityAttention_MissingAndDegradedAreDistinct: "no such entity" and
// "Nexus is down" must not produce the same answer.
func TestEntityAttention_MissingAndDegradedAreDistinct(t *testing.T) {
ctx := context.Background()
nexus := newFakeNexus(t, fixtureNexusNotFound())
praxis := newFakePraxis(t, fixturePraxisAttentionItems())
h := ecoHandler(t, nexus, praxis, nil)
missing := h.handlePraxisAct(ctx, entityAttentionDec("нечто"))
if missing == "" {
t.Fatal("an unknown entity must still get an answer")
}
nexus.SetFault(503)
degraded := h.handlePraxisAct(ctx, entityAttentionDec("нечто"))
if degraded == missing {
t.Fatalf("outage and unknown-entity must not read the same: %q", degraded)
}
}
// TestEntityAttention_DelayedNexusDegradesNotHangs: a slow Nexus past the
// caller's deadline degrades and never queries Praxis with an empty scope.
func TestEntityAttention_DelayedNexusDegradesNotHangs(t *testing.T) {
nexus := newFakeNexus(t, fixtureNexusResolved("ent_muzick", "Muzick indexer", "service"))
praxis := newFakePraxis(t, fixturePraxisAttentionItems())
h := ecoHandler(t, nexus, praxis, nil)
nexus.SetDelay(2 * time.Second)
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Millisecond)
defer cancel()
reply := h.handlePraxisAct(ctx, entityAttentionDec("muzick indexer"))
if reply == "" {
t.Fatal("a delayed resolve must still answer")
}
if praxis.Count("GET", "/api/v1/tools/attention") != 0 {
t.Fatal("praxis must not be queried without a resolved scope")
}
}
// TestEntityAttention_WithoutNexusSaysSo: no Nexus means no canonical ref, so
// the scoped query is refused rather than answered about something else.
func TestEntityAttention_WithoutNexusSaysSo(t *testing.T) {
ctx := context.Background()
praxis := newFakePraxis(t, fixturePraxisAttentionItems(
map[string]any{"id": "item_1", "title": "disk almost full", "importance": 3.0},
))
h := ecoHandler(t, nil, praxis, nil)
reply := h.handlePraxisAct(ctx, entityAttentionDec("muzick indexer"))
if strings.Contains(reply, "disk almost full") {
t.Fatalf("unscoped items must not be passed off as entity-scoped, got %q", reply)
}
if praxis.Count("GET", "/api/v1/tools/attention") != 0 {
t.Fatal("no canonical ref means no scoped query at all")
}
}
// TestEnrichmentBackoff_HoldsAndReleases: repeated Nexus failures back the
// fact off instead of hammering, and the fact is retried once the window
// elapses. Nothing is ever given up on.
func TestEnrichmentBackoff_HoldsAndReleases(t *testing.T) {
ctx := context.Background()
nexus := newFakeNexus(t, fixtureNexusResolved("ent_espresso", "the espresso machine", "device"))
st := newTestStore(t)
if _, err := st.WriteFactAboutSubject(ctx, time.Now(), store.KindEnv, "likes",
"the espresso machine", `"true"`, "infer:pref", 0.8, sql.NullInt64{}); err != nil {
t.Fatalf("WriteFactAboutSubject: %v", err)
}
clock := newFakeClock(time.Date(2026, 8, 1, 3, 0, 0, 0, time.UTC))
w := newFactEnrichmentWorker(st, stubEcosystem(nexus.URL, ""), time.Hour)
w.now = clock.Now
nexus.SetFault(503)
w.tick(ctx)
failedCalls := nexus.Count("POST", "/api/v1/resolve")
if failedCalls != 1 {
t.Fatalf("expected one resolve attempt, got %d", failedCalls)
}
// Immediately after a failure the fact is in backoff: no second call.
w.tick(ctx)
if nexus.Count("POST", "/api/v1/resolve") != failedCalls {
t.Fatal("a fact in backoff must not be retried on the very next tick")
}
if s := w.status(ctx); s.Pending != 1 || s.InBackoff != 1 || s.MaxAttempts != 1 {
t.Fatalf("degradation must be reported, got %+v", s)
}
// Once the window elapses and Nexus recovers, the fact resolves.
clock.Advance(2 * time.Minute)
nexus.SetFault(0)
w.tick(ctx)
facts, err := st.FactsByEntity(ctx, "ent_espresso", 10)
if err != nil {
t.Fatalf("FactsByEntity: %v", err)
}
if len(facts) != 1 {
t.Fatalf("expected the fact resolved after recovery, got %+v", facts)
}
if s := w.status(ctx); s.Pending != 0 || s.MaxAttempts != 0 {
t.Fatalf("recovery must clear the degradation report, got %+v", s)
}
}
func TestEnrichmentBackoff_GrowsAndIsCapped(t *testing.T) {
if enrichmentBackoff(1) != time.Minute {
t.Fatalf("first retry should be a minute, got %v", enrichmentBackoff(1))
}
if enrichmentBackoff(3) != 4*time.Minute {
t.Fatalf("third retry should be four minutes, got %v", enrichmentBackoff(3))
}
if enrichmentBackoff(50) != time.Hour {
t.Fatalf("backoff must cap at an hour, got %v", enrichmentBackoff(50))
}
}
+99 -5
View File
@@ -9,6 +9,7 @@ package main
import (
"context"
"log"
"sync"
"time"
"github.com/kami/maven/internal/store"
@@ -24,10 +25,69 @@ type factEnrichmentWorker struct {
eco *ecosystemWiring
interval time.Duration
batch int // facts resolved per tick; keeps a single slow tick bounded
now func() time.Time
// Retry state for facts whose resolution failed transiently. Kept in
// memory rather than in the DB: a restart legitimately retries
// everything, and the backoff exists to spare a struggling Nexus, not
// to be durable. A fact is never given up on — degraded means slower,
// not dropped.
mu sync.Mutex
attempt map[int64]int // fact id → consecutive failures
nextTry map[int64]time.Time // fact id → earliest retry
skipped int // facts held back by backoff on the last tick
}
// enrichmentBackoff is the wait before retrying a fact after n consecutive
// failures, capped so a long Nexus outage still retries about hourly.
func enrichmentBackoff(n int) time.Duration {
d := time.Minute
for i := 1; i < n && d < time.Hour; i++ {
d *= 2
}
if d > time.Hour {
d = time.Hour
}
return d
}
func newFactEnrichmentWorker(st *store.Store, eco *ecosystemWiring, interval time.Duration) *factEnrichmentWorker {
return &factEnrichmentWorker{store: st, eco: eco, interval: interval, batch: 20}
return &factEnrichmentWorker{
store: st,
eco: eco,
interval: interval,
batch: 20,
now: time.Now,
attempt: map[int64]int{},
nextTry: map[int64]time.Time{},
}
}
// enrichmentStatus is what the worker reports about its own health: how many
// facts are waiting, how many are currently in backoff, and the worst retry
// count seen. Degradation is reported, never hidden — a Nexus that has been
// down all day must be visible as a backlog, not as facts that silently
// never got tagged.
type enrichmentStatus struct {
Pending int
InBackoff int
MaxAttempts int
}
func (w *factEnrichmentWorker) status(ctx context.Context) enrichmentStatus {
var st enrichmentStatus
if pending, err := w.store.PendingFactResolutions(ctx, 1000); err == nil {
st.Pending = len(pending)
}
w.mu.Lock()
defer w.mu.Unlock()
st.InBackoff = w.skipped
for _, n := range w.attempt {
if n > st.MaxAttempts {
st.MaxAttempts = n
}
}
return st
}
func (w *factEnrichmentWorker) run(ctx context.Context) {
@@ -57,18 +117,50 @@ func (w *factEnrichmentWorker) tick(ctx context.Context) {
log.Printf("factenrichment: list pending: %v", err)
return
}
skipped, failed := 0, 0
for _, f := range pending {
w.resolveOne(ctx, f)
if !w.due(f.ID) {
skipped++
continue
}
if !w.resolveOne(ctx, f) {
failed++
}
}
w.mu.Lock()
w.skipped = skipped
w.mu.Unlock()
if failed > 0 {
log.Printf("factenrichment: %d/%d resolutions failed this tick, %d held in backoff",
failed, len(pending), skipped)
}
}
func (w *factEnrichmentWorker) resolveOne(ctx context.Context, f store.Fact) {
// due reports whether a fact's backoff window has elapsed.
func (w *factEnrichmentWorker) due(id int64) bool {
w.mu.Lock()
defer w.mu.Unlock()
next, ok := w.nextTry[id]
return !ok || !w.now().Before(next)
}
// resolveOne resolves one pending fact. It returns false when the attempt
// failed transiently: the fact stays pending and is retried on a backoff.
func (w *factEnrichmentWorker) resolveOne(ctx context.Context, f store.Fact) bool {
entityID, _, ambiguous, err := w.eco.resolveEntityReference(ctx, f.Subject, nil)
if err != nil {
// Transient (Nexus unreachable) — leave pending, retry next tick.
// Transient (Nexus unreachable) — leave pending, back off, retry later.
log.Printf("factenrichment: resolve fact %d subject %q: %v", f.ID, f.Subject, err)
return
w.mu.Lock()
w.attempt[f.ID]++
w.nextTry[f.ID] = w.now().Add(enrichmentBackoff(w.attempt[f.ID]))
w.mu.Unlock()
return false
}
w.mu.Lock()
delete(w.attempt, f.ID)
delete(w.nextTry, f.ID)
w.mu.Unlock()
state := store.ResolutionNotFound
switch {
case entityID != "":
@@ -78,5 +170,7 @@ func (w *factEnrichmentWorker) resolveOne(ctx context.Context, f store.Fact) {
}
if err := w.store.ResolveFactEntity(ctx, f.ID, entityID, state); err != nil {
log.Printf("factenrichment: record resolution for fact %d: %v", f.ID, err)
return false
}
return true
}
+96 -4
View File
@@ -14,7 +14,9 @@ import (
type capturedRequest struct {
Method string
Path string
Query string
Body []byte
Header http.Header
}
// fakeServer is the common shell behind fakeNexus/fakePraxis/fakeHexis: an
@@ -27,7 +29,9 @@ type fakeServer struct {
mu sync.Mutex
requests []capturedRequest
fault int // non-zero: every request gets this HTTP status instead of routing
fault int // non-zero: every request gets this HTTP status instead of routing
garbage string // non-empty: returned 200 verbatim instead of routing (malformed-contract lever)
delay time.Duration
}
// newFakeServer starts a server dispatching to routes keyed by "METHOD
@@ -47,14 +51,34 @@ func newFakeServer(t *testing.T, routes map[string]http.HandlerFunc) *fakeServer
}
}
fs.mu.Lock()
fs.requests = append(fs.requests, capturedRequest{Method: r.Method, Path: r.URL.Path, Body: body})
fs.requests = append(fs.requests, capturedRequest{
Method: r.Method,
Path: r.URL.Path,
Query: r.URL.RawQuery,
Body: body,
Header: r.Header.Clone(),
})
fault := fs.fault
garbage := fs.garbage
delay := fs.delay
fs.mu.Unlock()
if delay > 0 {
select {
case <-time.After(delay):
case <-r.Context().Done():
return
}
}
if fault != 0 {
http.Error(w, "injected fault", fault)
return
}
if garbage != "" {
w.Header().Set("Content-Type", "application/json")
w.Write([]byte(garbage))
return
}
for key, handler := range routes {
method, prefix := splitRouteKey(key)
@@ -90,6 +114,36 @@ func (fs *fakeServer) SetFault(status int) {
fs.fault = status
}
// SetBody makes every subsequent request answer 200 with the given body,
// bypassing the route table. Used to serve a malformed or contract-violating
// payload where the transport itself is healthy. Pass "" to clear it.
func (fs *fakeServer) SetBody(body string) {
fs.mu.Lock()
defer fs.mu.Unlock()
fs.garbage = body
}
// SetDelay stalls every subsequent request for d before answering, so callers
// can drive client timeouts and context cancellation deterministically. The
// delay is abandoned as soon as the client hangs up.
func (fs *fakeServer) SetDelay(d time.Duration) {
fs.mu.Lock()
defer fs.mu.Unlock()
fs.delay = d
}
// Count returns how many captured requests used the given method and path
// prefix. "" matches any method.
func (fs *fakeServer) Count(method, prefix string) int {
n := 0
for _, r := range fs.Requests() {
if (method == "" || r.Method == method) && hasPrefix(r.Path, prefix) {
n++
}
}
return n
}
// Requests returns a snapshot of captured requests, in arrival order.
func (fs *fakeServer) Requests() []capturedRequest {
fs.mu.Lock()
@@ -118,6 +172,32 @@ func fixtureNexusResolved(entityID, displayName, entityType string) string {
})
}
// fixtureNexusResolvedFlat is the flat resolve shape documented in
// ECOSYSTEM-SPEC.md §1.5 (entity_id/entity_type/display_name at the top
// level) rather than the nested "entity" object — the older of the two
// wire shapes Maven must keep accepting.
func fixtureNexusResolvedFlat(entityID, displayName, entityType string) string {
return mustJSON(map[string]any{
"status": "resolved",
"entity_id": entityID,
"entity_type": entityType,
"display_name": displayName,
})
}
// fixtureNexusResolvedFuture is a resolved response from a hypothetical newer
// Nexus: same required fields plus unknown ones. Decoding must ignore the
// extras, not fail — forward compatibility is what lets the ecosystem be
// upgraded one service at a time.
func fixtureNexusResolvedFuture(entityID, displayName, entityType string) string {
return mustJSON(map[string]any{
"status": "resolved",
"entity": map[string]any{"id": entityID, "display_name": displayName, "type": entityType, "tenant": "home"},
"provenance": map[string]any{"resolver": "v3", "graph_epoch": 42},
"score_breakdown": []any{map[string]any{"signal": "alias", "weight": 0.9}},
})
}
func fixtureNexusNotFound() string {
return `{"status":"not_found"}`
}
@@ -138,6 +218,13 @@ func fixtureHexisExecuted(id, status string) string {
return mustJSON(map[string]any{"id": id, "status": status})
}
// fixtureHexisExecutionFailed is a well-formed Hexis response reporting that
// the command itself failed: the call succeeded, the execution did not. Maven
// must distinguish this from a transport failure and from success.
func fixtureHexisExecutionFailed(id, message string) string {
return mustJSON(map[string]any{"id": id, "status": "failed", "error": message})
}
func fixturePraxisAttentionItems(items ...map[string]any) string {
return mustJSON(items)
}
@@ -191,8 +278,13 @@ func newFakeNexus(t *testing.T, resolveBody string) *fakeServer {
// fault is injected via SetFault.
func newFakePraxis(t *testing.T, attentionBody string) *fakeServer {
return newFakeServer(t, map[string]http.HandlerFunc{
"GET /api/v1/tools/attention": jsonHandler(http.StatusOK, attentionBody),
"POST /api/v1/tools/surface": jsonHandler(http.StatusOK, `{}`),
"GET /api/v1/tools/attention": jsonHandler(http.StatusOK, attentionBody),
"GET /api/v1/tools/changes": jsonHandler(http.StatusOK, `[]`),
"POST /api/v1/tools/surface": jsonHandler(http.StatusOK, `{}`),
"POST /api/v1/tools/acknowledge": jsonHandler(http.StatusOK, `{}`),
"POST /api/v1/tools/resolve": jsonHandler(http.StatusOK, `{}`),
"POST /api/v1/tools/ignore": jsonHandler(http.StatusOK, `{}`),
"POST /api/v1/tools/pin": jsonHandler(http.StatusOK, `{}`),
})
}
+194
View File
@@ -0,0 +1,194 @@
// mavend/feeds.go — the driver for RSS/Atom reading (Vikunja #258,
// docs/plans/13-rss-news-feeds.md). The reader itself is pure and lives in
// internal/rss; this is the impure half: a ticker, the guarded fetcher, and the
// two adapters that let a pure package talk to the store.
//
// Why in-core rather than its own daemon like mavmaild and mavpoll: those two
// hold a CREDENTIAL (an IMAP password, a zenmoney token), and the reason they
// are separate processes is that core must never see it. A feed URL is public,
// there is no secret to isolate, and a whole extra binary and compose service
// would buy nothing. The other half of the mavpoll precedent — off unless
// configured — is kept: no `feeds` block, no poller, no outbound request.
//
// It is its own goroutine, not a step on the tick: the tick has a delivery
// deadline behind it, and a feed read is a network round-trip that nobody is
// waiting on.
//
// Nothing here dispatches. A feed that announced itself would be a nag, so the
// only output is notes with source "rss:<feed>", which the answer path reads
// when he asks ("что нового в лентах?" — see queryFeeds in actions_query.go).
package main
import (
"context"
"log"
"net/url"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/rss"
"github.com/kami/maven/internal/stt"
"github.com/kami/maven/internal/webfetch"
)
// feedWorker — ticker + poller.
type feedWorker struct {
poller *rss.Poller
interval time.Duration
}
// feedTickInterval — how often the worker asks the poller what is due. Per-feed
// cadence is the poller's business; this is just the granularity.
const feedTickInterval = 5 * time.Minute
// newFeedWorker wires feed reading, or returns nil when it must not run:
// no `feeds` block (the normal case), or nothing valid in it. Every caller
// checks for nil.
func newFeedWorker(api ipc.CoreAPI, emb router.Embedder, cfg *config.Config) *feedWorker {
if cfg.Feeds == nil {
return nil
}
fc := cfg.Feeds
feeds := make([]rss.FeedConfig, 0, len(fc.Sources))
hosts := append([]string(nil), fc.AllowHosts...)
for _, s := range fc.Sources {
feeds = append(feeds, rss.FeedConfig{
Name: s.Name,
URL: s.URL,
Category: s.Category,
Interval: time.Duration(s.Interval),
Include: s.Include,
Exclude: s.Exclude,
})
// Each configured feed's own host is allowed. The allowlist is then
// exactly "the feeds he asked for", so a redirect off to somewhere else
// is refused by the fetcher rather than followed.
if u, err := url.Parse(s.URL); err == nil && u.Hostname() != "" {
hosts = append(hosts, u.Hostname())
}
}
fetcher := webfetch.New(webfetch.Config{
AllowHosts: hosts,
Timeout: time.Duration(fc.Timeout),
MaxBytes: fc.MaxBytes,
})
poller := rss.NewPoller(feeds, &feedFetcher{f: fetcher}, api, &factMarks{api: api},
embedderFor(emb), nil, rss.Config{
DefaultInterval: time.Duration(fc.PollInterval),
MaxItems: fc.MaxItems,
MaxAge: time.Duration(fc.MaxAge),
})
if poller == nil {
log.Printf("feeds: configured but nothing pollable — feed reading disabled")
return nil
}
log.Printf("feeds: reading %d feed(s), checking what is due every %s", len(feeds), feedTickInterval)
return &feedWorker{poller: poller, interval: feedTickInterval}
}
// run polls what is due until ctx is canceled. The first round runs immediately
// so a restart does not blind her for the first interval; it writes notes only,
// so an early round cannot startle anyone.
func (w *feedWorker) run(ctx context.Context) {
w.poller.PollDue(ctx, time.Now())
t := time.NewTicker(w.interval)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case now := <-t.C:
w.poller.PollDue(ctx, now)
}
}
}
// embedderOf — the voice wiring's embedder, or nil when voice is not wired.
// Feed notes are embedded with the SAME model the rest of the store uses, or not
// at all; a second embedder would write vectors nothing can search.
func embedderOf(w *voiceWiring) router.Embedder {
if w == nil {
return nil
}
return w.embedder
}
// transcriberOf — the STT the voice path is using, or nil when voice is off.
// The meeting recorder reuses it rather than dialling mavsttd a second time:
// Maven has one speech-to-text engine and adding a second would mean two
// whisper contexts competing for the same iGPU.
func transcriberOf(w *voiceWiring) stt.Transcriber {
if w == nil {
return nil
}
return w.transcriber
}
// feedFetcher adapts webfetch to rss.Fetcher — the pure package names the two
// fields it needs and stays free of net/http.
type feedFetcher struct{ f *webfetch.Fetcher }
func (a *feedFetcher) Get(ctx context.Context, u string) (*rss.Body, error) {
resp, err := a.f.Get(ctx, u)
if err != nil {
return nil, err
}
return &rss.Body{Bytes: resp.Body}, nil
}
// factMarks stores "how far this feed was read" as a config fact, the same
// mechanism the plan named and the same one the pattern tick uses for its own
// bookkeeping. Durable, inspectable on /dash, and cheap.
type factMarks struct{ api ipc.CoreAPI }
func markKey(feed string) string { return "rss:latest:" + feed }
func (m *factMarks) LastMark(ctx context.Context, feed string) (time.Time, error) {
f, err := m.api.LatestFact(ctx, markKey(feed))
if err != nil {
// No mark yet is not an error worth propagating: the poller treats a
// zero time as a cold start.
return time.Time{}, nil
}
t, err := time.Parse(time.RFC3339, f.Value)
if err != nil {
return time.Time{}, nil
}
return t, nil
}
func (m *factMarks) SetMark(ctx context.Context, feed string, at time.Time) error {
_, err := m.api.WriteFact(ctx, ipc.WriteFactReq{
Ts: time.Now(),
Kind: "config",
Key: markKey(feed),
Value: at.UTC().Format(time.RFC3339),
Source: "poll:rss",
Confidence: 1.0,
})
return err
}
// embedderFor adapts router.Embedder to rss.Embedder, and returns nil when
// there is none — a note without a vector is still a note the recent-notes path
// can read.
//
// EmbedPassage, not Embed: a feed item is text being searched FOR, and the e5
// embedder is asymmetric. Getting this backwards makes the item unfindable by
// the question that should have matched it.
func embedderFor(emb router.Embedder) rss.Embedder {
if emb == nil {
return nil
}
return passageEmbedder{emb}
}
type passageEmbedder struct{ e router.Embedder }
func (p passageEmbedder) Embed(ctx context.Context, text string) ([]float32, error) {
return router.EmbedPassage(ctx, p.e, text)
}
+177
View File
@@ -0,0 +1,177 @@
package main
import (
"context"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/rss"
"github.com/kami/maven/internal/voice"
)
// buildFeedHandler — a handler with the given feed notes already stored. No
// embedder: the feed source answers from recent notes by source, which is what
// makes it work for notes written before an embedder existed.
func buildFeedHandler(t *testing.T, feedsOn bool, notes ...ipc.Note) *reactiveHandler {
t.Helper()
ctx := context.Background()
st := newTestStore(t)
now := time.Now()
for i, n := range notes {
ts := now.Add(time.Duration(i) * time.Minute)
if _, err := st.WriteNote(ctx, ts, n.Text, nil, n.Source); err != nil {
t.Fatalf("WriteNote: %v", err)
}
}
return &reactiveHandler{
api: ipc.NewStoreAPI(st),
replier: voice.NewStubReplier(),
phraser: phraser.NewStub(),
now: func() time.Time { return now },
feedsOn: feedsOn,
embedder: nil,
}
}
func askFeeds(t *testing.T, h *reactiveHandler, q string) (string, bool) {
t.Helper()
return h.queryFeeds(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: q},
})
}
func TestQueryFeedsReadsFeedNotes(t *testing.T) {
h := buildFeedHandler(t, true,
ipc.Note{Text: "Новая уязвимость в ядре [технологии]\nпатч вышел\nhttps://example.org/a", Source: "rss:habr"},
ipc.Note{Text: "что-то он сам сказал", Source: "tap:voice"},
)
reply, ok := askFeeds(t, h, "что нового в лентах?")
if !ok {
t.Fatal("the feed source did not claim the question")
}
if !strings.Contains(reply, "уязвимость") {
t.Errorf("reply = %q, want the headline", reply)
}
if strings.Contains(reply, "он сам сказал") {
t.Errorf("a note he dictated leaked into the feed answer: %q", reply)
}
// She reads the headline, not the summary and not the URL.
if strings.Contains(reply, "https://") || strings.Contains(reply, "патч вышел") {
t.Errorf("reply = %q, want the title line only", reply)
}
}
func TestQueryFeedsByCategory(t *testing.T) {
h := buildFeedHandler(t, true,
ipc.Note{Text: "Релиз ядра [технологии]", Source: "rss:habr"},
ipc.Note{Text: "Выборы отложены [политика]", Source: "rss:news"},
)
reply, ok := askFeeds(t, h, "что нового по технологиям?")
if !ok {
t.Fatal("not claimed")
}
if !strings.Contains(reply, "ядра") || strings.Contains(reply, "Выборы") {
t.Fatalf("reply = %q, want only the технологии item", reply)
}
reply, _ = askFeeds(t, h, "что нового по спорту?")
if !strings.Contains(reply, "ничего") {
t.Fatalf("reply = %q, want an honest empty answer for an unread category", reply)
}
}
// "не настроены" and "ничего нового" are different truths, and neither may be
// answered by the model inventing a bulletin.
func TestQueryFeedsOffAndEmptyDiffer(t *testing.T) {
off := buildFeedHandler(t, false)
reply, ok := askFeeds(t, off, "что нового?")
if !ok || !strings.Contains(reply, "не настроены") {
t.Fatalf("feeds off: reply = %q, ok = %v", reply, ok)
}
on := buildFeedHandler(t, true)
reply, ok = askFeeds(t, on, "что нового?")
if !ok || !strings.Contains(reply, "ничего нового") {
t.Fatalf("feeds on but empty: reply = %q, ok = %v", reply, ok)
}
}
func TestQueryFeedsPassesOnANonFeedQuestion(t *testing.T) {
h := buildFeedHandler(t, true)
if reply, ok := askFeeds(t, h, "напомни полить цветы"); ok {
t.Fatalf("claimed an unrelated question with %q", reply)
}
}
// The mark is what stops a restart from re-noting yesterday's headlines, so the
// fact round-trip is worth a test of its own.
func TestFactMarksRoundTrip(t *testing.T) {
st := newTestStore(t)
m := &factMarks{api: ipc.NewStoreAPI(st)}
ctx := context.Background()
at, err := m.LastMark(ctx, "habr")
if err != nil || !at.IsZero() {
t.Fatalf("no mark yet: got %v, %v — want zero time and no error", at, err)
}
want := time.Date(2026, 7, 28, 10, 0, 0, 0, time.UTC)
if err := m.SetMark(ctx, "habr", want); err != nil {
t.Fatal(err)
}
got, err := m.LastMark(ctx, "habr")
if err != nil {
t.Fatal(err)
}
if !got.Equal(want) {
t.Fatalf("mark = %v, want %v", got, want)
}
}
// Off unless configured, checked at the wiring seam: no `feeds` block ⇒ no
// worker ⇒ no outbound request is possible.
func TestNewFeedWorkerOffByDefault(t *testing.T) {
st := newTestStore(t)
api := ipc.NewStoreAPI(st)
if w := newFeedWorker(api, nil, &config.Config{}); w != nil {
t.Fatal("a config with no feeds block wired a feed worker")
}
// An empty sources list is normalised to "off" by config.Load; the worker
// refuses it too, so a hand-built Config cannot switch it on by accident.
if w := newFeedWorker(api, nil, &config.Config{Feeds: &config.FeedsConfig{}}); w != nil {
t.Fatal("an empty sources list wired a feed worker")
}
cfg := &config.Config{Feeds: &config.FeedsConfig{Sources: []config.FeedSourceConfig{
{Name: "habr", URL: "https://example.org/rss"},
}}}
w := newFeedWorker(api, nil, cfg)
if w == nil {
t.Fatal("a configured feed did not wire a worker")
}
if got := w.poller.Feeds(); len(got) != 1 || got[0].Name != "habr" {
t.Fatalf("feeds = %+v", got)
}
}
// The fetcher the worker builds must be allowlisted to the configured feeds and
// nothing else — the crawler's SSRF guards are only worth as much as the
// allowlist handed to them.
func TestFeedWorkerFetcherIsAllowlisted(t *testing.T) {
cfg := &config.Config{Feeds: &config.FeedsConfig{Sources: []config.FeedSourceConfig{
{Name: "habr", URL: "https://feeds.example.org/rss"},
}}}
w := newFeedWorker(ipc.NewStoreAPI(newTestStore(t)), nil, cfg)
if w == nil {
t.Fatal("no worker")
}
// PollFeed goes through the guarded fetcher; a feed URL pointing at the box
// itself must fail rather than be read.
_, err := w.poller.PollFeed(context.Background(), rss.FeedConfig{
Name: "evil", URL: "http://127.0.0.1:9100/mcp",
}, time.Now())
if err == nil {
t.Fatal("the poller fetched a private address")
}
}
+235
View File
@@ -0,0 +1,235 @@
// mavend/intake.go — the unified event intake envelope, wired (Vikunja #283).
//
// internal/event defines the envelope and the bounded in-memory journal. This
// file is the one place that FILLS it, and the reason it is one place is worth
// stating, because the alternative was eight patches:
//
// Every intake path in Maven already converges on three writes, and all three
// are ipc.CoreAPI methods —
//
// WriteFact ← POST /api/ambient, mavcaldav, mavpoll's zenmoney + wg reads,
// /api/signal presence probes, the RSS/crawl watermarks
// WriteNote ← the RSS poller, the page crawler, meeting transcripts,
// image descriptions
// CaptureTask ← the voice path, the web form, and the mail reader
//
// — so decorating that ONE interface with a publish covers the lot without a
// caller knowing about events at all. cmd/mavmaild, cmd/mavcaldav, cmd/mavpoll,
// cmd/mavweb and the in-core feed/crawl/capture/vision workers are unchanged:
// they call the same interface they always called, and it now also narrates.
//
// The exception is cmd/mavend/mail.go, which reaches past the interface to
// st.CaptureTask directly. It publishes explicitly; see mailIntake.ingest.
//
// # Production behaviour when nobody is watching
//
// A nil *event.Bus makes Publish a no-op, and newIntakeAPI with a nil bus
// returns the wrapped API unchanged, so there is not even a decorator on the
// call path. The journal is memory-only and is never consulted by the tick
// loop, the router, or delivery — nothing Maven says depends on it. It is a
// read surface (`/events`, `recent_events`) and an observation seam for the
// simulator.
//
// # What is deliberately NOT here
//
// No dispatch. An event is a report that something arrived, never an
// instruction to speak: "a feed item appeared" becoming a notification is the
// nag this repo refuses. Digestion may one day read the journal; it will still
// go through internal/loop's rules and the severity/presence routing table.
package main
import (
"context"
"log"
"strings"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/event"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/store"
)
// newEventBus builds the journal, or returns nil when the operator turned it
// off (a negative config.intake_journal). nil is the "behave exactly as before"
// value all the way down: no decorator, no ring, no /events rows.
func newEventBus(cfg *config.Config) *event.Bus {
if cfg == nil || cfg.IntakeJournal < 0 {
log.Printf("intake journal: off (intake_journal < 0)")
return nil
}
n := cfg.IntakeJournal
if n == 0 {
n = config.DefaultIntakeJournal
}
log.Printf("intake journal: keeping the last %d intake events in memory", n)
return event.NewBus(n)
}
// intakeEventsFn is the daemonAPI.getEvents closure: the bus's ring rendered as
// the wire type. Returns nil for a nil bus, which the daemonAPI reports as an
// empty journal rather than an error.
func intakeEventsFn(bus *event.Bus) func(n int) []ipc.IntakeEvent {
if bus == nil {
return nil
}
return func(n int) []ipc.IntakeEvent {
evs := bus.Recent(n)
out := make([]ipc.IntakeEvent, 0, len(evs))
for _, e := range evs {
out = append(out, ipc.IntakeEvent{
Source: e.Source,
Kind: e.Kind,
EntityIDs: e.EntityIDs,
Title: e.Title,
Body: e.Body,
Priority: e.Priority,
OccurredAt: e.OccurredAt,
})
}
return out
}
}
// intakeAPI decorates a CoreAPI, publishing one envelope per successful
// intake write. Embedding the interface means every other method passes
// through untouched, and a new CoreAPI method is inherited rather than
// silently dropped.
type intakeAPI struct {
ipc.CoreAPI
bus *event.Bus
now func() time.Time
}
// newIntakeAPI wraps api so its intake writes are journalled. A nil bus
// returns api itself — no decorator, no allocation, no behaviour change.
func newIntakeAPI(api ipc.CoreAPI, bus *event.Bus, now func() time.Time) ipc.CoreAPI {
if bus == nil || api == nil {
return api
}
if now == nil {
now = time.Now
}
return &intakeAPI{CoreAPI: api, bus: bus, now: now}
}
// WriteFact journals the fact after it lands. Order matters: an event is a
// report of something that HAPPENED, so a failed write publishes nothing.
func (a *intakeAPI) WriteFact(ctx context.Context, req ipc.WriteFactReq) (int64, error) {
id, err := a.CoreAPI.WriteFact(ctx, req)
if err != nil {
return id, err
}
// OccurredAt is req.Ts, not now: mavpoll's wg read carries the handshake
// instant and the ambient path carries the meeting's start. Flattening
// those to notice-time would make the journal lie about when things
// happened, which is the one thing it is for.
a.bus.Publish(event.Event{
Source: req.Source,
Kind: event.SourceKind(req.Source, event.KindFact),
Title: req.Key,
Body: req.Value,
Priority: factPriority(req),
OccurredAt: req.Ts,
EntityIDs: entityIDs(req.Subject),
}, a.now())
return id, nil
}
// WriteNote journals a note. This is the RSS and crawler path, and also the
// meeting transcript and image description paths, which write their derived
// text as ordinary notes.
func (a *intakeAPI) WriteNote(ctx context.Context, ts time.Time, text string, embedding []float32, source string) (int64, error) {
id, err := a.CoreAPI.WriteNote(ctx, ts, text, embedding, source)
if err != nil {
return id, err
}
title, body := splitFirstLine(text)
a.bus.Publish(event.Event{
Source: source,
Kind: event.SourceKind(source, event.KindNote),
Title: title,
Body: body,
Priority: event.PriorityLow,
OccurredAt: ts,
}, a.now())
return id, nil
}
// CaptureTask journals a captured task, but only when a row was actually
// created. CaptureTask dedupes on normalised text among live rows, so a
// mailbox re-read after a restart must not refill the journal with tasks that
// were already there.
func (a *intakeAPI) CaptureTask(ctx context.Context, req ipc.CaptureTaskReq) (ipc.CaptureTaskResp, error) {
resp, err := a.CoreAPI.CaptureTask(ctx, req)
if err != nil || !resp.Created {
return resp, err
}
a.bus.Publish(publishableTask(store.Task{
CreatedTs: req.Ts,
Text: req.Text,
Source: req.Source,
Evidence: req.Evidence,
Status: req.Status,
Due: req.Due,
}, a.now()), a.now())
return resp, nil
}
// publishableTask is the task→envelope shape, shared with mail.go, which
// captures through the store directly rather than through the interface.
//
// Priority is high for a candidate with a due date and normal otherwise. That
// is the only place this file makes a judgement, and it is a display hint on a
// review page — nothing routes on it.
func publishableTask(t store.Task, now time.Time) event.Event {
occurred := t.CreatedTs
if occurred.IsZero() {
occurred = now
}
prio := event.PriorityNormal
if t.Due != nil {
prio = event.PriorityHigh
}
return event.Event{
Source: t.Source,
Kind: event.KindTask,
Title: t.Text,
Body: t.Evidence,
Priority: prio,
OccurredAt: occurred,
}
}
// factPriority is the attention hint for a fact write. Deliberately crude:
// a low-confidence inference (the ambient notification path writes below 1.0)
// is worth less attention than a read he or a credentialled poller made, and
// nothing else is distinguishable from here.
func factPriority(req ipc.WriteFactReq) string {
if req.Confidence > 0 && req.Confidence < 1.0 {
return event.PriorityLow
}
return event.PriorityNormal
}
// entityIDs turns a fact's free-text Subject into the EntityIDs slot when it
// already looks resolved. Intake runs BEFORE the fact enrichment worker
// resolves a subject against Nexus, so this is almost always empty — the slot
// exists for the paths that do know (the ecosystem acts), not for guessing.
func entityIDs(subject string) []string {
subject = strings.TrimSpace(subject)
if subject == "" || !strings.HasPrefix(subject, "entity:") {
return nil
}
return []string{strings.TrimPrefix(subject, "entity:")}
}
// splitFirstLine renders a note as title + body. Feed and crawl notes are
// written "headline\nsummary\nlink", so the first line is already the title.
func splitFirstLine(text string) (title, body string) {
text = strings.TrimSpace(text)
if i := strings.IndexByte(text, '\n'); i >= 0 {
return strings.TrimSpace(text[:i]), strings.TrimSpace(text[i+1:])
}
return text, ""
}
+178
View File
@@ -0,0 +1,178 @@
package main
import (
"context"
"errors"
"testing"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/event"
"github.com/kami/maven/internal/ipc"
)
var intakeNow = time.Date(2026, 8, 1, 10, 0, 0, 0, time.UTC)
func intakeClock() time.Time { return intakeNow }
// failingAPI wraps the store adapter, failing the three intake writes on
// demand, so the "a failed write publishes nothing" invariant is testable.
type failingAPI struct {
ipc.CoreAPI
fail bool
}
func (f *failingAPI) WriteFact(ctx context.Context, req ipc.WriteFactReq) (int64, error) {
if f.fail {
return 0, errors.New("injected")
}
return f.CoreAPI.WriteFact(ctx, req)
}
func newIntakeTestAPI(t *testing.T) (ipc.CoreAPI, *event.Bus) {
t.Helper()
st := newTestStore(t)
bus := event.NewBus(32)
return newIntakeAPI(ipc.NewStoreAPI(st), bus, intakeClock), bus
}
func TestIntakeAPIWithoutBusIsTheBareAPI(t *testing.T) {
// The adoption invariant: with the journal off there is not even a
// decorator on the intake path, so production behaves exactly as before.
st := newTestStore(t)
bare := ipc.NewStoreAPI(st)
if got := newIntakeAPI(bare, nil, intakeClock); got != ipc.CoreAPI(bare) {
t.Errorf("newIntakeAPI with a nil bus returned a wrapper, want the bare API")
}
}
func TestNewEventBusOffWhenNegative(t *testing.T) {
if b := newEventBus(&config.Config{IntakeJournal: -1}); b != nil {
t.Error("intake_journal = -1 still built a bus")
}
if b := newEventBus(&config.Config{IntakeJournal: 4}); b == nil {
t.Error("intake_journal = 4 built no bus")
}
}
func TestIntakeJournalsAFactWrite(t *testing.T) {
api, bus := newIntakeTestAPI(t)
ctx := context.Background()
// The ambient path's shape: an env fact below full confidence, timestamped
// at the meeting's start rather than at notice time.
start := intakeNow.Add(2 * time.Hour)
if _, err := api.WriteFact(ctx, ipc.WriteFactReq{
Ts: start, Kind: "env", Key: "calendar_event_20260801_планёрка",
Value: "10:00-11:00 планёрка", Source: "ambient:notif", Confidence: 0.6,
}); err != nil {
t.Fatalf("WriteFact: %v", err)
}
got := bus.Recent(0)
if len(got) != 1 {
t.Fatalf("journal has %d entries, want 1", len(got))
}
e := got[0]
if e.Source != "ambient:notif" || e.Kind != event.KindFact {
t.Errorf("source/kind = %q/%q", e.Source, e.Kind)
}
if e.Title != "calendar_event_20260801_планёрка" {
t.Errorf("title = %q, want the fact key", e.Title)
}
if !e.OccurredAt.Equal(start) {
t.Errorf("occurred_at = %v, want the fact's Ts %v — the journal must not flatten intake to notice time", e.OccurredAt, start)
}
if e.Priority != event.PriorityLow {
t.Errorf("priority = %q, want %q for a sub-1.0 confidence read", e.Priority, event.PriorityLow)
}
}
func TestIntakeDoesNotJournalAFailedWrite(t *testing.T) {
st := newTestStore(t)
bus := event.NewBus(8)
api := newIntakeAPI(&failingAPI{CoreAPI: ipc.NewStoreAPI(st), fail: true}, bus, intakeClock)
if _, err := api.WriteFact(context.Background(), ipc.WriteFactReq{
Ts: intakeNow, Kind: "env", Key: "k", Value: "v", Source: "poll:zenmoney", Confidence: 1,
}); err == nil {
t.Fatal("expected the injected error")
}
if bus.Len() != 0 {
t.Errorf("journal has %d entries after a failed write, want 0 — an event reports something that happened", bus.Len())
}
}
func TestIntakeJournalsANoteAsTitlePlusBody(t *testing.T) {
api, bus := newIntakeTestAPI(t)
// The RSS shape: "headline\nsummary\nlink".
if _, err := api.WriteNote(context.Background(), intakeNow,
"Вышло ядро 6.19\nкраткое содержание\nhttps://example.org/a", nil, "rss:tech"); err != nil {
t.Fatalf("WriteNote: %v", err)
}
got := bus.Recent(1)
if len(got) != 1 {
t.Fatalf("journal has %d entries, want 1", len(got))
}
if got[0].Title != "Вышло ядро 6.19" {
t.Errorf("title = %q, want the headline", got[0].Title)
}
if got[0].Kind != event.KindNote {
t.Errorf("kind = %q, want %q", got[0].Kind, event.KindNote)
}
if got[0].Body == "" {
t.Error("body is empty, want the rest of the note")
}
}
func TestIntakeJournalsOnlyCreatedTasks(t *testing.T) {
api, bus := newIntakeTestAPI(t)
ctx := context.Background()
req := ipc.CaptureTaskReq{Text: "оплатить интернет", Source: "email:inbox", Status: "candidate", Ts: intakeNow}
if _, err := api.CaptureTask(ctx, req); err != nil {
t.Fatalf("CaptureTask: %v", err)
}
// Same text again: CaptureTask dedupes among live rows, and a re-read of a
// mailbox must not refill the journal.
resp, err := api.CaptureTask(ctx, req)
if err != nil {
t.Fatalf("CaptureTask (repeat): %v", err)
}
if resp.Created {
t.Fatal("store did not dedupe; the test cannot check what it means to")
}
if bus.Len() != 1 {
t.Errorf("journal has %d entries, want 1 — a deduped capture must not publish", bus.Len())
}
if got := bus.Recent(1)[0]; got.Kind != event.KindTask || got.Title != "оплатить интернет" {
t.Errorf("entry = %+v, want the captured task", got)
}
}
func TestIntakeEventsFnRendersNewestFirst(t *testing.T) {
api, bus := newIntakeTestAPI(t)
ctx := context.Background()
for _, key := range []string{"a", "b", "c"} {
if _, err := api.WriteFact(ctx, ipc.WriteFactReq{
Ts: intakeNow, Kind: "env", Key: key, Value: "1", Source: "poll:zenmoney", Confidence: 1,
}); err != nil {
t.Fatalf("WriteFact %s: %v", key, err)
}
}
fn := intakeEventsFn(bus)
got := fn(2)
if len(got) != 2 || got[0].Title != "c" || got[1].Title != "b" {
t.Errorf("intakeEventsFn(2) = %+v, want the two newest, newest first", got)
}
if intakeEventsFn(nil) != nil {
t.Error("intakeEventsFn(nil) returned a closure, want nil so daemonAPI reports an empty journal")
}
}
func TestDaemonAPIRecentEventsEmptyWithoutABus(t *testing.T) {
d := &daemonAPI{CoreAPI: ipc.UnimplementedCoreAPI{}}
got, err := d.RecentEvents(context.Background(), 10)
if err != nil {
t.Fatalf("RecentEvents with no journal errored: %v", err)
}
if len(got) != 0 {
t.Errorf("got %d events, want none", len(got))
}
}
+28
View File
@@ -0,0 +1,28 @@
package main
import (
"testing"
"time"
"github.com/kami/maven/internal/llm"
)
func TestPickLLMRouterOff(t *testing.T) {
if r := pickLLMRouter(false, llm.New("http://127.0.0.1:1", time.Second)); r != nil {
t.Error("flag off should give no LLM router")
}
}
// The operator can turn the flag on without an LLM phraser configured. That must
// leave the classifier running, not panic.
func TestPickLLMRouterOnWithoutClient(t *testing.T) {
if r := pickLLMRouter(true, nil); r != nil {
t.Error("no llama-server should give no LLM router")
}
}
func TestPickLLMRouterOn(t *testing.T) {
if r := pickLLMRouter(true, llm.New("http://127.0.0.1:1", time.Second)); r == nil {
t.Error("flag on with a client should give an LLM router")
}
}
+167
View File
@@ -0,0 +1,167 @@
// mavend/mail.go — core's half of the email reader (Vikunja #246,
// docs/plans/01-email-reader.md).
//
// The split: cmd/mavmaild holds the IMAP credential, connects to the mailbox
// and converts messages to plaintext; it hands each message to core over
// ipc.MethodIngestMail. Core runs the extraction on the resident model —
// llama-server lives in this process, spawned by the phraser — and writes what
// comes back through the one task intake seam.
//
// What this file may produce is exactly one thing: rows in `tasks` with status
// "candidate". No fact, no reminder, no note, no nudge, no calendar event. A
// 1.7B misreading a mail can therefore put a wrong line on a review page and
// nothing else; it can never make Maven speak, and it can never make her
// recite something out of an advert as true.
//
// Off unless configured twice over: no `email` block in mavend.json ⇒ the IPC
// method does not exist; no llama-server phraser ⇒ same. A reader pointed at a
// core that is not set up for mail gets ErrUnknownMethod rather than silence.
package main
import (
"context"
"fmt"
"log"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/email"
"github.com/kami/maven/internal/event"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/store"
)
// evidenceMaxChars — how much of the subject line is kept as a candidate's
// evidence. Enough to recognise the mail on /tasks, not enough to turn the task
// list into a copy of his mailbox.
const evidenceMaxChars = 160
// mailIntake — extraction + capture for one message at a time.
type mailIntake struct {
st *store.Store
ex *email.Extractor
timeout time.Duration
now func() time.Time
// bus — the unified intake journal (Vikunja #283). This path captures
// through the store directly rather than through ipc.CoreAPI, so the
// decorator in intake.go does not see it and the publish is explicit here.
// nil is a working no-op.
bus *event.Bus
}
// newMailIntake returns nil when mail ingestion must not be available, which is
// the default. Both preconditions are real:
//
// - no cfg.Email ⇒ not configured, and a capability is off unless configured;
// - no llama-server phraser ⇒ nothing to extract with. There is deliberately
// no keyword fallback: "the subject line became a task" is not extraction,
// it is a mailbox rendered as a to-do list, and it would fill the review
// page faster than he could clear it.
func newMailIntake(st *store.Store, phr phraser.Phraser, cfg *config.Config, bus *event.Bus) *mailIntake {
if cfg.Email == nil {
return nil
}
lp, ok := phr.(*phraser.LLMPhraser)
if !ok {
log.Printf("mail intake: configured but no llama-server phraser — mail ingestion disabled")
return nil
}
timeout := time.Duration(cfg.Email.Timeout)
if timeout <= 0 {
timeout = config.DefaultEmailTimeout
}
ex := email.NewExtractor(llmClientFor(lp, timeout), cfg.Email.MaxTasks, contextBlockFn(cfg, time.Now))
log.Printf("mail intake: enabled (max %d candidates per message, timeout %s)", cfg.Email.MaxTasks, timeout)
return &mailIntake{st: st, ex: ex, timeout: timeout, now: time.Now, bus: bus}
}
// ingest handles one ipc.MethodIngestMail call.
//
// Junk and empty messages are answered Skipped without touching the model — the
// reader's header filter is what keeps the resident model off newsletters.
//
// Every candidate is captured with Status "candidate", Source "email:<mailbox>"
// and the subject as Evidence. CaptureTask dedupes on normalised text among
// live rows, so a mailbox re-read after a restart produces Created=0 rather
// than a second copy of every task.
func (m *mailIntake) ingest(ctx context.Context, req ipc.IngestMailReq) (ipc.IngestMailResp, error) {
msg := email.Message{
UID: req.UID,
From: req.From,
Subject: req.Subject,
Date: req.Date,
Body: req.Body,
Junk: req.Junk,
}
if msg.Junk || (msg.Subject == "" && msg.Body == "") {
return ipc.IngestMailResp{Skipped: true}, nil
}
ctx, cancel := context.WithTimeout(ctx, m.timeout)
defer cancel()
cands, err := m.ex.Extract(ctx, msg)
if err != nil {
// The error from internal/email never carries mail text; keep it that way
// by not adding the subject here.
return ipc.IngestMailResp{}, fmt.Errorf("mail intake: uid %d: %w", req.UID, err)
}
if len(cands) == 0 {
return ipc.IngestMailResp{}, nil
}
source := email.SourcePrefix + req.Mailbox
evidence := truncateRunes(req.Subject, evidenceMaxChars)
now := m.now()
var resp ipc.IngestMailResp
for _, c := range cands {
t := store.Task{
CreatedTs: now,
Text: c.Text,
Source: source,
Evidence: evidence,
// The one status this path may ever write. Anything Maven derived from
// something she read is a suggestion until he confirms it on /tasks.
Status: store.TaskCandidate,
}
if due, ok := email.ParseDue(c.Due); ok {
t.Due = &due
}
id, created, err := m.st.CaptureTask(ctx, t)
if err != nil {
return resp, fmt.Errorf("mail intake: capture: %w", err)
}
resp.TaskIDs = append(resp.TaskIDs, id)
if created {
resp.Created++
// Only a row that was actually created. CaptureTask dedupes on
// normalised text among live rows, so a mailbox re-read after a
// restart must not refill the journal with tasks already in it.
m.bus.Publish(publishableTask(t, now), now)
}
}
// Counts only: the log line names the mailbox and the UID, never the subject,
// the sender or the task text. Reviewing a candidate is what /tasks is for.
log.Printf("mail intake: %s uid %d → %d candidate(s), %d new", source, req.UID, len(cands), resp.Created)
return resp, nil
}
// wireMailIntake installs the IPC hook, or leaves it nil so the method reports
// ErrUnknownMethod. Called on both startup paths (unlocked boot and passkey
// unlock) so mail behaves the same either way.
func wireMailIntake(srv *ipc.Server, st *store.Store, phr phraser.Phraser, cfg *config.Config, bus *event.Bus) {
mi := newMailIntake(st, phr, cfg, bus)
if mi == nil {
return
}
srv.IngestMailFn = mi.ingest
}
// truncateRunes cuts a string to n runes, marking the cut.
func truncateRunes(s string, n int) string {
r := []rune(s)
if len(r) <= n {
return s
}
return string(r[:n]) + "…"
}
+182
View File
@@ -0,0 +1,182 @@
package main
import (
"context"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/email"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/store"
)
// mailLLM — a canned extraction reply.
type mailLLM struct {
reply string
calls int
}
func (m *mailLLM) Complete(_ context.Context, _ llm.Req) (string, error) {
m.calls++
return m.reply, nil
}
func newTestIntake(t *testing.T, reply string) (*mailIntake, *store.Store, *mailLLM) {
t.Helper()
st := newTestStore(t)
fake := &mailLLM{reply: reply}
return &mailIntake{
st: st,
ex: email.NewExtractor(fake, 0, nil),
timeout: 5 * time.Second,
now: func() time.Time { return time.Date(2026, 8, 1, 10, 0, 0, 0, time.UTC) },
}, st, fake
}
func ingestReq() ipc.IngestMailReq {
return ipc.IngestMailReq{
Mailbox: "INBOX", UID: 42,
From: "billing@isp.example",
Subject: "Счёт за интернет",
Body: "Оплатите счёт до 5 августа.",
}
}
// The one property that matters: a mail-derived task is a candidate, attributed
// to the mailbox, with the subject as reviewable evidence — and nothing else is
// written.
func TestIngestCapturesCandidates(t *testing.T) {
mi, st, _ := newTestIntake(t, `[{"text":"оплатить счёт за интернет","due":"2026-08-05"}]`)
resp, err := mi.ingest(context.Background(), ingestReq())
if err != nil {
t.Fatalf("ingest: %v", err)
}
if resp.Created != 1 || len(resp.TaskIDs) != 1 {
t.Fatalf("resp = %+v, want one created task", resp)
}
tasks, err := st.ListTasks(context.Background(), "")
if err != nil {
t.Fatalf("list: %v", err)
}
if len(tasks) != 1 {
t.Fatalf("got %d tasks, want 1", len(tasks))
}
got := tasks[0]
if got.Status != store.TaskCandidate {
t.Errorf("status = %q, want %q — mail may only produce candidates", got.Status, store.TaskCandidate)
}
if got.Source != "email:INBOX" {
t.Errorf("source = %q, want email:INBOX", got.Source)
}
if got.Evidence != "Счёт за интернет" {
t.Errorf("evidence = %q, want the subject line", got.Evidence)
}
if got.Due == nil || got.Due.Format("2006-01-02") != "2026-08-05" {
t.Errorf("due = %v, want 2026-08-05", got.Due)
}
// Nothing else may have been written: no reminder, no fact.
rem, err := st.ListReminders(context.Background(), 10)
if err != nil {
t.Fatalf("list reminders: %v", err)
}
if len(rem) != 0 {
t.Errorf("mail created %d reminders; a misread mail must never be able to fire", len(rem))
}
}
// Re-reading a mailbox must not grow the list — CaptureTask dedupes among live
// rows, and the intake relies on exactly that.
func TestIngestSameMailTwiceIsIdempotent(t *testing.T) {
mi, st, _ := newTestIntake(t, `[{"text":"оплатить счёт","due":""}]`)
if _, err := mi.ingest(context.Background(), ingestReq()); err != nil {
t.Fatalf("first ingest: %v", err)
}
resp, err := mi.ingest(context.Background(), ingestReq())
if err != nil {
t.Fatalf("second ingest: %v", err)
}
if resp.Created != 0 || len(resp.TaskIDs) != 1 {
t.Errorf("resp = %+v, want the existing row and Created=0", resp)
}
tasks, _ := st.ListTasks(context.Background(), "")
if len(tasks) != 1 {
t.Errorf("got %d tasks after two reads, want 1", len(tasks))
}
}
func TestIngestJunkSkipsTheModel(t *testing.T) {
mi, st, fake := newTestIntake(t, `[{"text":"купить со скидкой","due":""}]`)
req := ingestReq()
req.Junk = true
resp, err := mi.ingest(context.Background(), req)
if err != nil {
t.Fatalf("ingest: %v", err)
}
if !resp.Skipped || resp.Created != 0 {
t.Errorf("resp = %+v, want skipped", resp)
}
if fake.calls != 0 {
t.Errorf("model called %d times for junk, want 0", fake.calls)
}
if tasks, _ := st.ListTasks(context.Background(), ""); len(tasks) != 0 {
t.Errorf("junk produced %d tasks, want 0", len(tasks))
}
}
func TestIngestEmptyMessageSkipped(t *testing.T) {
mi, _, fake := newTestIntake(t, "[]")
resp, err := mi.ingest(context.Background(), ipc.IngestMailReq{Mailbox: "INBOX", UID: 1})
if err != nil || !resp.Skipped {
t.Fatalf("resp = %+v, err = %v; want skipped", resp, err)
}
if fake.calls != 0 {
t.Errorf("model called %d times for an empty message, want 0", fake.calls)
}
}
func TestIngestNoTasksWritesNothing(t *testing.T) {
mi, st, _ := newTestIntake(t, "[]")
resp, err := mi.ingest(context.Background(), ingestReq())
if err != nil {
t.Fatalf("ingest: %v", err)
}
if resp.Created != 0 || len(resp.TaskIDs) != 0 || resp.Skipped {
t.Errorf("resp = %+v, want nothing captured and not skipped", resp)
}
if tasks, _ := st.ListTasks(context.Background(), ""); len(tasks) != 0 {
t.Errorf("got %d tasks, want 0", len(tasks))
}
}
func TestIngestTruncatesEvidence(t *testing.T) {
mi, st, _ := newTestIntake(t, `[{"text":"дело","due":""}]`)
req := ingestReq()
req.Subject = strings.Repeat("щ", 400)
if _, err := mi.ingest(context.Background(), req); err != nil {
t.Fatalf("ingest: %v", err)
}
tasks, _ := st.ListTasks(context.Background(), "")
if len(tasks) != 1 {
t.Fatalf("got %d tasks, want 1", len(tasks))
}
if n := len([]rune(tasks[0].Evidence)); n > evidenceMaxChars+1 {
t.Errorf("evidence kept %d runes, want ≤ %d", n, evidenceMaxChars)
}
}
// Off unless configured: no email block ⇒ no intake, so the IPC method does not
// exist at all.
func TestNewMailIntakeOffWithoutConfig(t *testing.T) {
st := newTestStore(t)
if mi := newMailIntake(st, nil, &config.Config{}, nil); mi != nil {
t.Error("no email block must mean no mail intake")
}
// Configured but with a non-LLM phraser: still off — there is no fallback
// extraction, by design.
if mi := newMailIntake(st, nil, &config.Config{Email: &config.EmailConfig{}}, nil); mi != nil {
t.Error("without a llama-server phraser there is nothing to extract with")
}
}
+267 -142
View File
@@ -57,20 +57,28 @@ import (
"github.com/kami/maven/internal/delivery/ntfysink"
"github.com/kami/maven/internal/delivery/telegramsink"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/webauthn"
"github.com/kami/maven/internal/loop"
)
var errLocked = errors.New("mavend: daemon locked — complete passkey assertion first")
// daemonLock tracks whether the daemon is in locked (pre-unlock) mode.
// In locked mode, all CoreAPI methods return errLocked. The unlock path
// replaces the CoreAPI with the real store adapter and flips the flag.
// daemonLock tracks whether the daemon is in locked (pre-unlock) mode, and
// owns the store handle the unlock path creates.
//
// The store matters here because of who runs when. In locked mode there is no
// store at boot; one is opened inside UnlockFn, on an IPC goroutine, minutes
// or days later. Shutdown runs on the main goroutine. Without a handoff the
// main goroutine has nothing to close, and store.Close is what re-encrypts
// the tmpfs working copy back over the ciphertext file — so a daemon that
// cold-started lost every write of that session, silently, on the next boot.
type daemonLock struct {
mu sync.Mutex
locked bool
st *store.Store
}
func newDaemonLock(locked bool) *daemonLock {
@@ -83,10 +91,25 @@ func (l *daemonLock) isLocked() bool {
return l.locked
}
func (l *daemonLock) unlock() {
// unlock flips the flag and takes ownership of the store opened by UnlockFn.
func (l *daemonLock) unlock(st *store.Store) {
l.mu.Lock()
defer l.mu.Unlock()
l.locked = false
l.st = st
}
// closeStore seals the store the unlock path opened, if any. Safe to call
// when the daemon never unlocked, and safe to call twice.
func (l *daemonLock) closeStore() error {
l.mu.Lock()
st := l.st
l.st = nil
l.mu.Unlock()
if st == nil {
return nil
}
return st.Close()
}
func main() {
@@ -96,97 +119,12 @@ func main() {
}
}
// lockedAPI is a dummy CoreAPI used while the daemon is locked. Every method
// returns errLocked. The wire protocol's StoreAPI methods all go through the
// Server dispatch on CoreAPI, so returning errLocked from each is correct.
type lockedAPI struct{}
var _ ipc.CoreAPI = (*lockedAPI)(nil)
func (l *lockedAPI) WriteFact(ctx context.Context, req ipc.WriteFactReq) (int64, error) {
return 0, errLocked
}
func (l *lockedAPI) LatestFact(ctx context.Context, key string) (ipc.Fact, error) {
return ipc.Fact{}, errLocked
}
func (l *lockedAPI) LatestFactBySource(ctx context.Context, key, source string) (ipc.Fact, error) {
return ipc.Fact{}, errLocked
}
func (l *lockedAPI) Since(ctx context.Context, key string, now time.Time) (time.Duration, error) {
return 0, errLocked
}
func (l *lockedAPI) Presence(ctx context.Context) (ipc.Presence, error) {
return ipc.Presence{}, errLocked
}
func (l *lockedAPI) CreateReminder(ctx context.Context, fire time.Time, payload, cron string) (int64, error) {
return 0, errLocked
}
func (l *lockedAPI) MarkReminder(ctx context.Context, id int64, status string) error {
return errLocked
}
func (l *lockedAPI) ListReminders(ctx context.Context, n int) ([]ipc.Reminder, error) {
return nil, errLocked
}
func (l *lockedAPI) RecordNudge(ctx context.Context, rule, channel, message string, ts time.Time) (int64, error) {
return 0, errLocked
}
func (l *lockedAPI) ResolveNudge(ctx context.Context, id int64, outcome string, ts time.Time) error {
return errLocked
}
func (l *lockedAPI) RecentOutcomes(ctx context.Context, rule string, n int) ([]string, error) {
return nil, errLocked
}
func (l *lockedAPI) RecentFacts(ctx context.Context, n int) ([]ipc.Fact, error) {
return nil, errLocked
}
func (l *lockedAPI) CalendarEvents(ctx context.Context, from, to time.Time) ([]ipc.Fact, error) {
return nil, errLocked
}
func (l *lockedAPI) RecentNudges(ctx context.Context, n int) ([]ipc.Nudge, error) {
return nil, errLocked
}
func (l *lockedAPI) WriteNote(ctx context.Context, ts time.Time, text string, embedding []float32, source string) (int64, error) {
return 0, errLocked
}
func (l *lockedAPI) QueryNotes(ctx context.Context, embedding []float32, k int) ([]ipc.Note, error) {
return nil, errLocked
}
func (l *lockedAPI) RecentNotes(ctx context.Context, n int) ([]ipc.Note, error) {
return nil, errLocked
}
func (l *lockedAPI) ProposeTool(ctx context.Context, name, utterance, scope string, ts time.Time) (bool, error) {
return false, errLocked
}
func (l *lockedAPI) EnableTool(ctx context.Context, name string, cmd []string, destructive bool, scope string, ts time.Time) error {
return errLocked
}
func (l *lockedAPI) DisableTool(ctx context.Context, name string) error { return errLocked }
func (l *lockedAPI) DeleteTool(ctx context.Context, name string) error { return errLocked }
func (l *lockedAPI) ListProposedRoutines(ctx context.Context) ([]ipc.ProposedRoutine, error) {
return nil, errLocked
}
func (l *lockedAPI) DismissProposedRoutine(ctx context.Context, id int64) error { return errLocked }
func (l *lockedAPI) LookupTool(ctx context.Context, name string) (ipc.Tool, error) {
return ipc.Tool{}, errLocked
}
func (l *lockedAPI) ListTools(ctx context.Context, status string) ([]ipc.Tool, error) {
return nil, errLocked
}
func (l *lockedAPI) RevertFact(ctx context.Context, key string) (int64, error) { return 0, errLocked }
func (l *lockedAPI) Chat(ctx context.Context, text string) (string, error) {
return "", errLocked
}
func (l *lockedAPI) TickTrace(ctx context.Context) (ipc.TickTrace, error) {
return ipc.TickTrace{}, errLocked
}
func (l *lockedAPI) MorningStatus(ctx context.Context) ([]ipc.MorningRoutineStatus, error) {
return nil, errLocked
}
func run(args []string) error {
cfgPath := flag.String("config", defaultConfigPath(), "path to mavend JSON config")
wrappedKeyPath := flag.String("wrapped-key-file", "", "path to wrapped encryption key blob (enables cold-start unlock)")
reembed := flag.Bool("reembed", false, "re-embed every stored note and fact with the configured embedder, then serve normally (run once after an embedder swap)")
flag.CommandLine.Parse(args)
reembedOnStart = *reembed
cfg, err := config.Load(*cfgPath)
if err != nil {
return err
@@ -239,22 +177,43 @@ func run(args []string) error {
return fmt.Errorf("open store: %w", err)
}
defer st.Close()
} else {
// Locked boot: the store does not exist yet. Seal whatever UnlockFn
// opened, at shutdown, on this goroutine.
defer func() {
if err := dl.closeStore(); err != nil {
log.Printf("mavend: seal store on shutdown: %v", err)
}
}()
}
// ----- daemon components (only wired when unlocked) -----
// Pre-declare so the unlock path can wire them later.
var (
gatherer *loop.Gatherer
rules []loop.Rule
phr phraser.Phraser
voiceW *voiceWiring
dispatcher *delivery.Dispatcher
tl *tickLoop
coreAPI ipc.CoreAPI
eco *ecosystemWiring
factWorker *factEnrichmentWorker
gatherer *loop.Gatherer
rules []loop.Rule
phr phraser.Phraser
voiceW *voiceWiring
dispatcher *delivery.Dispatcher
tl *tickLoop
coreAPI ipc.CoreAPI
eco *ecosystemWiring
factWorker *factEnrichmentWorker
evalWorker *memoryEvalWorker // nil ⇒ memory evaluation off (the default)
feedWkr *feedWorker // nil ⇒ no feed is read (the default)
crawlWkr *crawlWorker // nil ⇒ no page is watched (the default)
)
// The unified intake journal (Vikunja #283). Built before anything else
// that holds a CoreAPI, because intakeAPI wraps that one interface and
// every intake path in the daemon reaches its sink through it. nil (the
// operator set intake_journal negative) means no decorator at all.
evBus := newEventBus(cfg)
// coreFor is what every in-process holder of a CoreAPI now takes, instead
// of a bare ipc.NewStoreAPI(st). Identical behaviour plus one published
// envelope per successful intake write.
coreFor := func() ipc.CoreAPI { return newIntakeAPI(ipc.NewStoreAPI(st), evBus, time.Now) }
if !locked {
rules = loop.DefaultRules()
gatherer = loop.NewGatherer(st, rules)
@@ -266,13 +225,14 @@ func run(args []string) error {
phr = phraser.NewStub()
if cfg.Phraser != nil {
pc := phraser.Config{
ModelPath: cfg.Phraser.ModelPath,
BinPath: cfg.Phraser.BinPath,
Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
Timeout: time.Duration(cfg.Phraser.Timeout),
Persona: personaFromCfg(cfg),
ModelPath: cfg.Phraser.ModelPath,
BinPath: cfg.Phraser.BinPath,
Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
Timeout: time.Duration(cfg.Phraser.Timeout),
LLMNudges: cfg.Phraser.LLMNudges,
ContextBlock: contextBlockFn(cfg, time.Now),
}
if pc.BinPath == "" {
pc.BinPath = "llama-server"
@@ -297,7 +257,7 @@ func run(args []string) error {
eco = wireEcosystem(cfg)
// voice
voiceW, err = wireVoice(cfg, ipc.NewStoreAPI(st), phr, st.VectorMemory(), st, eco)
voiceW, err = wireVoice(cfg, coreFor(), phr, st.VectorMemory(), st, eco)
if err != nil {
return fmt.Errorf("wire voice: %w", err)
}
@@ -344,21 +304,34 @@ func run(args []string) error {
tickInterval := time.Duration(cfg.TickInterval)
repeatInterval := time.Duration(cfg.RepeatInterval)
autotuneInterval := time.Duration(cfg.AutotuneInterval)
tl = newTickLoop(st, gatherer, dispatcher, phr, rules, tickInterval, repeatInterval, autotuneInterval, cfg.Digest, routinesFromConfig(cfg.Routines), config.MorningRoutinesFromConfig(cfg.MorningRoutines))
tl = newTickLoop(st, gatherer, dispatcher, phr, rules, tickInterval, repeatInterval, autotuneInterval, cfg.Digest, routinesFromConfig(cfg.Routines), config.MorningRoutinesFromConfig(cfg.MorningRoutines), cfg.PatternProposals)
factWorker = newFactEnrichmentWorker(st, eco, time.Duration(cfg.FactEnrichmentInterval))
evalWorker = newMemoryEvalWorker(st, phr, cfg)
feedWkr = newFeedWorker(coreFor(), embedderOf(voiceW), cfg)
crawlWkr = newCrawlWorker(newCrawler(cfg), coreFor(), embedderOf(voiceW), cfg)
coreAPI = &daemonAPI{
CoreAPI: ipc.NewStoreAPI(st),
CoreAPI: coreFor(),
getTrace: tl.trace,
getMorningStatus: func(ctx context.Context) []ipc.MorningRoutineStatus { return tl.morningStatus(ctx, time.Now()) },
getDayPlan: func(ctx context.Context) ipc.DayPlan { return tl.dayPlan(ctx, time.Now()) },
getEvents: intakeEventsFn(evBus),
}
if voiceW != nil && voiceW.handler != nil {
api := coreAPI.(*daemonAPI)
api.chatFn = voiceW.handler.handleText
}
if voiceW != nil && voiceW.mcp != nil {
coreAPI.(*daemonAPI).getMCPServers = voiceW.mcp.status
}
} else {
// locked mode: dummy CoreAPI that returns errLocked for everything
coreAPI = &lockedAPI{}
// locked mode: no real store yet, so there's no meaningful CoreAPI to
// serve. srv.Check below is the actual guard — every CoreAPI call is
// refused before it reaches this value. This is just a safe non-nil
// placeholder: if the guard is ever bypassed by a bug, calls land
// here and fail loudly with ipc.ErrNotImplemented instead of a nil
// dereference or, worse, silently succeeding.
coreAPI = ipc.UnimplementedCoreAPI{}
}
// ----- IPC boundary (core ↔ modules) -----
@@ -369,7 +342,17 @@ func run(args []string) error {
passkeySess := webauthn.NewPasskeySession(5 * time.Minute)
// Set Server.Check — in locked mode, block everything except unlock-path methods.
// Set Server.Check — the single authorization guard, run once by
// Server.dispatch before any CoreAPI method is called (see
// internal/ipc/server.go). In locked mode this is the ONLY thing
// standing between an unauthenticated caller and the store: it must
// default-deny, with an explicit allowlist for the two methods the
// unlock flow itself needs (MethodAssertStepUp, MethodUnlock — neither
// of which touches CoreAPI; dispatch handles them directly via
// srv.StepUp/srv.UnlockFn). Forgetting to allowlist a new unlock-path
// method fails safe (denied); forgetting to guard a new CoreAPI method
// is impossible because there is nothing left to forget — every method
// not in the allowlist is refused by construction.
if locked {
srv.Check = func(ctx context.Context, m ipc.Method, _ json.RawMessage) error {
switch m {
@@ -385,12 +368,35 @@ func run(args []string) error {
srv.StepUp = func(ctx context.Context) error { return passkeySess.Assert(ctx, auth.Scope{}) }
// WrapKeyFn — wraps the env key with a passkey credential public key and
// persists the wrapped blob. Only wired when the daemon has the key in
// memory (env key mode). Called by mavweb after passkey enrollment.
// Mail ingestion (Vikunja #246): the hook stays nil unless an email block is
// configured and there is a llama-server to extract with, in which case
// ipc.MethodIngestMail reports ErrUnknownMethod.
if !locked {
wireMailIntake(srv, st, phr, cfg, evBus)
wireModelSwap(srv, phr, cfg)
// Vision + the media blob store (Vikunja #252). Both stay dark without a
// media block; MethodDescribeImage answers ErrUnknownMethod then.
keeper := wireVision(ctx, srv, st, embedderOf(voiceW), cfg)
// The meeting recorder (Vikunja #253) shares that blob store and its
// retention loop. Off unless a capture block enables it, in which case
// all four capture methods answer ErrUnknownMethod.
wireCapture(srv, keeper, st, voiceW, phr, cfg)
// Voice identification (Vikunja #255). Enrolment plumbing only until a
// speaker-embedding model exists on disk; off entirely without a speaker
// block, so no wire path takes a voiceprint on a default box.
wireSpeaker(srv, st, cfg)
}
// WrapKeyFn — wraps the env key under the passkey PRF secret and persists
// the wrapped blob. Only wired when the daemon has the key in memory (env
// key mode). Called by mavweb after passkey enrollment.
//
// webauthn.WrapKey refuses anything that is not a 32-byte PRF output, so
// an authenticator without PRF support produces no wrapped file at all
// rather than a file that looks protected and is not.
if envKeyBytes != nil {
srv.WrapKeyFn = func(ctx context.Context, publicKey []byte) error {
blob, err := webauthn.WrapKey(envKeyBytes, publicKey)
srv.WrapKeyFn = func(ctx context.Context, secret []byte) error {
blob, err := webauthn.WrapKey(envKeyBytes, secret)
if err != nil {
return fmt.Errorf("wrap encryption key: %w", err)
}
@@ -406,20 +412,42 @@ func run(args []string) error {
}
}
// UnlockFn — cold-start unlock: unwraps the encryption key from the wrapped
// blob using the passkey credential public key, opens the store, wires all
// UnlockFn — cold-start unlock: unwraps the encryption key from the
// wrapped blob using the passkey PRF secret, opens the store, wires all
// daemon components, and replaces the locked API.
if locked {
srv.UnlockFn = func(ctx context.Context, publicKey []byte) error {
var unlockMu sync.Mutex
srv.UnlockFn = func(ctx context.Context, secret []byte) error {
// One unlock at a time, and never a second one. Without this a
// concurrent pair of Unlock calls would each open a store and
// wire a full daemon, and the loser's goroutines would run
// against a store nobody closes.
unlockMu.Lock()
defer unlockMu.Unlock()
if !dl.isLocked() {
return nil // already unlocked; the caller does not need to know
}
// The wire cannot authenticate its caller — the socket is
// same-uid — so the unlock path requires a passkey assertion
// that mavweb verified cryptographically first. Without this,
// MethodUnlock is reachable by anything on the box.
if !passkeySess.IsStepUp() {
return errors.New("unlock: no verified passkey assertion (assert first)")
}
wp := *wrappedKeyPath
blob, err := os.ReadFile(wp)
if err != nil {
return fmt.Errorf("read wrapped key: %w", err)
}
key, err := webauthn.UnwrapKey(blob, publicKey)
key, version, err := webauthn.UnwrapKey(blob, secret)
if err != nil {
return fmt.Errorf("unwrap key: %w", err)
}
if version == webauthn.BlobV1 {
log.Printf("SECURITY: %s was unwrapped from a %s blob. The wrapping key is derived from the credential PUBLIC key, which mavweb also writes to its passkeys.json — anyone holding both files can recover the database key with no authenticator. Re-enroll the passkey on an authenticator that supports the PRF extension to rewrite it as v2.", wp, version)
}
// Open the store with the unwrapped key.
st, err = store.OpenEncrypted(ctx, cfg.DBPath, cfg.DBTmpfs, key)
if err != nil {
@@ -436,13 +464,14 @@ func run(args []string) error {
phr = phraser.NewStub()
if cfg.Phraser != nil {
pc := phraser.Config{
ModelPath: cfg.Phraser.ModelPath,
BinPath: cfg.Phraser.BinPath,
Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
Timeout: time.Duration(cfg.Phraser.Timeout),
Persona: personaFromCfg(cfg),
ModelPath: cfg.Phraser.ModelPath,
BinPath: cfg.Phraser.BinPath,
Listen: cfg.Phraser.Listen,
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
Timeout: time.Duration(cfg.Phraser.Timeout),
LLMNudges: cfg.Phraser.LLMNudges,
ContextBlock: contextBlockFn(cfg, time.Now),
}
if pc.BinPath == "" {
pc.BinPath = "llama-server"
@@ -464,7 +493,7 @@ func run(args []string) error {
eco = wireEcosystem(cfg)
voiceW, err = wireVoice(cfg, ipc.NewStoreAPI(st), phr, st.VectorMemory(), st, eco)
voiceW, err = wireVoice(cfg, coreFor(), phr, st.VectorMemory(), st, eco)
if err != nil {
return fmt.Errorf("wire voice: %w", err)
}
@@ -505,20 +534,33 @@ func run(args []string) error {
tickInterval := time.Duration(cfg.TickInterval)
repeatInterval := time.Duration(cfg.RepeatInterval)
autotuneInterval := time.Duration(cfg.AutotuneInterval)
tl = newTickLoop(st, gatherer, dispatcher, phr, rules, tickInterval, repeatInterval, autotuneInterval, cfg.Digest, routinesFromConfig(cfg.Routines), config.MorningRoutinesFromConfig(cfg.MorningRoutines))
tl = newTickLoop(st, gatherer, dispatcher, phr, rules, tickInterval, repeatInterval, autotuneInterval, cfg.Digest, routinesFromConfig(cfg.Routines), config.MorningRoutinesFromConfig(cfg.MorningRoutines), cfg.PatternProposals)
factWorker = newFactEnrichmentWorker(st, eco, time.Duration(cfg.FactEnrichmentInterval))
evalWorker = newMemoryEvalWorker(st, phr, cfg)
feedWkr = newFeedWorker(coreFor(), embedderOf(voiceW), cfg)
crawlWkr = newCrawlWorker(newCrawler(cfg), coreFor(), embedderOf(voiceW), cfg)
// Swap the CoreAPI from lockedAPI to the real store adapter.
// Swap the CoreAPI from the locked placeholder to the real store adapter.
newAPI := &daemonAPI{
CoreAPI: ipc.NewStoreAPI(st),
CoreAPI: coreFor(),
getTrace: tl.trace,
getMorningStatus: func(ctx context.Context) []ipc.MorningRoutineStatus { return tl.morningStatus(ctx, time.Now()) },
getDayPlan: func(ctx context.Context) ipc.DayPlan { return tl.dayPlan(ctx, time.Now()) },
getEvents: intakeEventsFn(evBus),
}
if voiceW != nil && voiceW.handler != nil {
newAPI.chatFn = voiceW.handler.handleText
}
srv.SetAPI(newAPI)
srv.Check = (&auth.Gate{Enrollment: auth.NewFloorEnrollment(), Session: passkeySess}).Check
wireMailIntake(srv, st, phr, cfg, evBus)
wireModelSwap(srv, phr, cfg)
keeper := wireVision(ctx, srv, st, embedderOf(voiceW), cfg)
wireCapture(srv, keeper, st, voiceW, phr, cfg)
// Voice identification (Vikunja #255). Enrolment plumbing only until a
// speaker-embedding model exists on disk; off entirely without a speaker
// block, so no wire path takes a voiceprint on a default box.
wireSpeaker(srv, st, cfg)
// Start voice server.
if voiceW != nil {
@@ -543,7 +585,38 @@ func run(args []string) error {
factWorker.run(ctx)
}()
dl.unlock()
// Start background memory evaluation (nil unless configured).
if evalWorker != nil {
go func() {
evalWorker.run(ctx)
}()
}
// Start feed reading (nil unless configured).
if feedWkr != nil {
go func() {
feedWkr.run(ctx)
}()
}
// Start the watched-page crawls (nil unless configured).
if crawlWkr != nil {
go func() {
crawlWkr.run(ctx)
}()
}
// Keep MCP connections alive (nil unless configured).
if voiceW != nil && voiceW.mcp != nil {
go voiceW.mcp.run(ctx)
}
// Re-enumerate the house for new devices (nil unless configured).
if voiceW != nil && voiceW.home != nil {
go voiceW.home.run(ctx)
}
dl.unlock(st)
log.Printf("mavend: unlocked via passkey assertion")
return nil
}
@@ -581,6 +654,41 @@ func run(args []string) error {
defer wg.Done()
factWorker.run(ctx)
}()
if evalWorker != nil {
wg.Add(1)
go func() {
defer wg.Done()
evalWorker.run(ctx)
}()
}
if feedWkr != nil {
wg.Add(1)
go func() {
defer wg.Done()
feedWkr.run(ctx)
}()
}
if crawlWkr != nil {
wg.Add(1)
go func() {
defer wg.Done()
crawlWkr.run(ctx)
}()
}
if voiceW != nil && voiceW.mcp != nil {
wg.Add(1)
go func() {
defer wg.Done()
voiceW.mcp.run(ctx)
}()
}
if voiceW != nil && voiceW.home != nil {
wg.Add(1)
go func() {
defer wg.Done()
voiceW.home.run(ctx)
}()
}
}
<-ctx.Done()
@@ -596,12 +704,29 @@ func run(args []string) error {
return nil
}
// personaFromCfg extracts the voice persona from the config, or returns ""
// when voice isn't configured. Used to pass a character prompt into the
// LLM phraser without requiring voice to be enabled.
func personaFromCfg(cfg *config.Config) string {
if cfg.Voice != nil {
return cfg.Voice.Persona
// personaFacts reads the optional, deployment-specific facts (his name, his
// city, the free-text persona string) out of the config. Everything here may
// be empty — the context block is correct without any of it.
func personaFacts(cfg *config.Config) persona.Facts {
f := persona.Facts{
// Telegram lives outside the voice block, so it counts either way.
Telegram: cfg.Telegram != nil && cfg.Telegram.BotToken != "" && cfg.Telegram.ChatID != "",
}
return ""
if cfg.Voice == nil {
return f
}
f.OwnerName = cfg.Voice.OwnerName
f.City = cfg.Voice.City
f.Static = cfg.Voice.Persona
// Same test wireVoice uses to pick the real provider over the stub.
f.Weather = cfg.Voice.Weather != nil && cfg.Voice.Weather.Provider == "open-meteo"
f.Tools = len(cfg.Voice.Tools) > 0
return f
}
// contextBlockFn returns the per-turn renderer of the shared context block.
// Per turn, not once at startup, because the block states the current time.
func contextBlockFn(cfg *config.Config, now func() time.Time) func() string {
f := personaFacts(cfg)
return func() string { return f.Block(now()) }
}
+154
View File
@@ -0,0 +1,154 @@
package main
import (
"context"
"fmt"
"log"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/mcp"
"github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/webfetch"
)
// mcpRefreshInterval — how often the manager re-dials a server that is down.
// The manager applies its own backoff on top, so this being short is cheap.
const mcpRefreshInterval = time.Minute
// mcpWiring — the MCP client, when the `mcp` block configures at least one
// enabled server. nil ⇒ nothing was configured, nothing is connected, and an
// allowlist row that happens to look like an MCP row refuses to run.
//
// It lives on the voice wiring because MCP tools ARE acts: they run through
// tool.Executor, the enabled allowlist and the confirm turn, which only exist
// on the voice/chat path. No voice surface ⇒ nothing that could call a tool.
type mcpWiring struct {
mgr *mcp.Manager
st *store.Store
}
// wireMCP builds the manager, connects, and proposes what it found. It never
// fails the daemon: a server that is unreachable at boot is logged and retried,
// because Maven starting is not contingent on someone else's process.
func wireMCP(cfg *config.Config, st *store.Store) *mcpWiring {
servers := cfg.MCPServers()
if len(servers) == 0 {
return nil
}
limits := webfetch.Config{}
if cfg.MCP != nil {
limits.AllowHosts = cfg.MCP.AllowHosts
limits.DenyHosts = cfg.MCP.DenyHosts
limits.MaxBytes = cfg.MCP.MaxBytes
limits.Timeout = time.Duration(cfg.MCP.Timeout)
}
mgr, err := mcp.NewManager(mcp.WebfetchDoor(limits), servers)
if err != nil {
// Validation already ran in config.validate, so this is a programming
// error rather than a config one. Still not fatal: MCP off is a working
// Maven.
log.Printf("mcp: not wired: %v", err)
return nil
}
w := &mcpWiring{mgr: mgr, st: st}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
mgr.Connect(ctx)
w.propose(ctx)
return w
}
// propose writes a 'proposed' allowlist row for every discovered tool. It does
// NOT enable anything: a configured server is a place Maven may look, not a
// capability she has. Kami enables what he wants on /tools, behind step-up,
// which is the same gate a shell tool goes through.
//
// Re-running on every boot is idempotent — ProposeMCPTool never touches an
// existing row, so a tool he disabled stays disabled and one he enabled keeps
// the cmd he enabled it with.
func (w *mcpWiring) propose(ctx context.Context) {
if w == nil {
return
}
now := time.Now()
fresh := 0
for _, t := range w.mgr.Tools() {
name := mcp.LocalName(t.Server, t.Name)
// No readOnlyHint ⇒ assume it mutates ⇒ the confirm turn. Being wrong
// in this direction only costs a question.
destructive := !t.ReadOnly
provenance := fmt.Sprintf("mcp %s/%s", t.Server, t.Name)
if t.Description != "" {
provenance += ": " + t.Description
}
ok, err := w.st.ProposeMCPTool(ctx, name, mcp.Scope(t.Server),
mcp.Cmd(t.Server, t.Name), destructive, provenance, now)
if err != nil {
log.Printf("mcp: propose %s: %v", name, err)
continue
}
if ok {
fresh++
}
}
if fresh > 0 {
log.Printf("mcp: %d new tool proposal(s) waiting on /tools", fresh)
}
}
// run re-dials downed servers and picks up tools that appeared, until ctx is
// canceled.
func (w *mcpWiring) run(ctx context.Context) {
if w == nil {
return
}
t := time.NewTicker(mcpRefreshInterval)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
w.mgr.Refresh(ctx)
w.propose(ctx)
}
}
}
// status maps the manager's view onto the wire type the web surface reads.
func (w *mcpWiring) status() []ipc.MCPServerStatus {
if w == nil {
return nil
}
in := w.mgr.Status()
out := make([]ipc.MCPServerStatus, 0, len(in))
for _, s := range in {
out = append(out, ipc.MCPServerStatus{
Name: s.Name,
Transport: s.Transport,
Target: s.Target,
Connected: s.Connected,
Server: s.Server,
Tools: s.Tools,
Err: s.Err,
})
}
return out
}
func (w *mcpWiring) close() {
if w == nil {
return
}
_ = w.mgr.Close()
}
// caller is the tool.MCPCaller the executor gets, or nil when MCP is off.
func (w *mcpWiring) caller() *mcp.Manager {
if w == nil {
return nil
}
return w.mgr
}
+77
View File
@@ -0,0 +1,77 @@
package main
import (
"context"
"strings"
"testing"
"github.com/kami/maven/internal/config"
)
func TestWireMCPOffWhenUnconfigured(t *testing.T) {
st := newTestStore(t)
for name, cfg := range map[string]*config.Config{
"no block": {},
"nothing enabled": {MCP: &config.MCPConfig{Servers: []config.MCPServerConfig{
{Name: "vikunja", URL: "http://192.168.1.104:9100/mcp"},
}}},
} {
t.Run(name, func(t *testing.T) {
if w := wireMCP(cfg, st); w != nil {
t.Fatal("MCP must be off unless a server is configured AND enabled")
}
})
}
// nil wiring must be safe to use everywhere it is reachable.
var w *mcpWiring
w.close()
w.propose(context.Background())
if w.status() != nil || w.caller() != nil {
t.Fatal("a nil wiring must report nothing")
}
}
// An unreachable server must not stop the daemon, must be reported as down, and
// must propose nothing.
func TestWireMCPUnreachableServerIsNotFatal(t *testing.T) {
st := newTestStore(t)
w := wireMCP(&config.Config{MCP: &config.MCPConfig{Servers: []config.MCPServerConfig{{
Name: "dead", Command: "/nonexistent/mcp-server", Enabled: true,
}}}}, st)
if w == nil {
t.Fatal("a configured server should still wire")
}
defer w.close()
st2 := w.status()
if len(st2) != 1 || st2[0].Connected || st2[0].Err == "" {
t.Fatalf("status = %+v", st2)
}
tools, err := st.ListTools(context.Background(), "")
if err != nil {
t.Fatal(err)
}
if len(tools) != 0 {
t.Fatalf("a server that never answered must propose nothing, got %+v", tools)
}
}
// A url server whose address is private is refused by webfetch unless that
// server sets allow_private. This is the guard the whole MCP path rides on, so
// it is asserted here too, at the wiring level.
func TestWireMCPPrivateURLRefusedWithoutAllowPrivate(t *testing.T) {
st := newTestStore(t)
w := wireMCP(&config.Config{MCP: &config.MCPConfig{Servers: []config.MCPServerConfig{{
Name: "lan", URL: "http://127.0.0.1:9100/mcp", Enabled: true,
}}}}, st)
if w == nil {
t.Fatal("should wire")
}
defer w.close()
s := w.status()[0]
if s.Connected {
t.Fatal("a loopback server must not connect without allow_private")
}
if !strings.Contains(s.Err, "private address") {
t.Fatalf("err = %q, want the private-address refusal", s.Err)
}
}
+87
View File
@@ -0,0 +1,87 @@
// mavend/memoryeval.go — the driver for background memory evaluation
// (Vikunja #248). The evaluator itself is pure-ish and lives in
// internal/memeval; this is the one impure part: a ticker, the store, and the
// resident model's base URL.
//
// It is its own goroutine and NOT a step on the main tick, deliberately. The
// tick runs every 60s and has a delivery deadline behind it; an evaluation is
// a multi-second LLM round-trip on the same llama-server that answers voice
// turns, and it happens hourly at most. Bolting it onto the tick would make
// every hour's tick the slow one for no benefit.
package main
import (
"context"
"log"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/memeval"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/store"
)
// memoryEvalWorker — ticker + evaluator.
type memoryEvalWorker struct {
eval *memeval.Evaluator
interval time.Duration
}
// newMemoryEvalWorker wires the evaluation loop, or returns nil when it should
// not run at all. nil is the normal case and every caller must handle it:
//
// - no memory_eval config block ⇒ off (a capability is off unless configured);
// - no LLM phraser ⇒ nothing to evaluate with. There is no template fallback
// here on purpose: a "memory evaluation" assembled from string templates
// would be a fixed sentence pretending to be an observation.
func newMemoryEvalWorker(st *store.Store, phr phraser.Phraser, cfg *config.Config) *memoryEvalWorker {
if cfg.MemoryEval == nil {
return nil
}
lp, ok := phr.(*phraser.LLMPhraser)
if !ok {
log.Printf("memory eval: configured but no llama-server phraser — evaluation disabled")
return nil
}
interval := time.Duration(cfg.MemoryEval.Interval)
if interval <= 0 {
interval = config.DefaultMemoryEvalInterval
}
// A generous per-request timeout: this is a long prompt to a Thinking model
// and nobody is waiting on the answer.
client := llmClientFor(lp, 5*time.Minute)
ev := memeval.NewEvaluator(st, st, client, memeval.Config{
MaxItems: cfg.MemoryEval.MaxItems,
MinConfidence: cfg.MemoryEval.MinConfidence,
ContextBlock: contextBlockFn(cfg, time.Now),
})
log.Printf("memory eval: enabled, every %s", interval)
return &memoryEvalWorker{eval: ev, interval: interval}
}
// run evaluates every interval until ctx is canceled.
//
// The first evaluation waits a full interval rather than firing at startup, the
// opposite of the tick loop's cold-start behaviour. A tick that fires late is a
// nudge that arrives late; an evaluation that fires late is nothing at all, and
// the alternative is a heavy LLM call competing with startup — including with
// the first voice turn after a restart.
func (w *memoryEvalWorker) run(ctx context.Context) {
ticker := time.NewTicker(w.interval)
defer ticker.Stop()
for {
select {
case <-ctx.Done():
return
case now := <-ticker.C:
obs, err := w.eval.Evaluate(ctx, now)
if err != nil {
log.Printf("memory eval: %v", err)
continue
}
for _, o := range obs {
log.Printf("memory eval: noted (%.2f, %s): %s", o.Conf, o.Action, o.Text)
}
}
}
}
+113
View File
@@ -0,0 +1,113 @@
package main
import (
"context"
"fmt"
"log"
"path/filepath"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/phraser"
)
// Swapping the resident model while the daemon runs (Vikunja #250).
//
// Off unless configured: with no phraser.swap_models allowlist the two IPC
// methods are never wired, so they answer ErrUnknownMethod. When it is wired the
// swap method is AuthStepUp (internal/auth), which means an authed human surface
// only — there is no act, no intent and no timer that reaches it. The daemon
// never decides to change its own brain.
//
// The allowlist is exact-match against paths a human wrote in mavend.json. The
// request carries a path and llama-server is started with it as `-m`, so
// anything looser would turn "swap the model" into "load any file on my disk".
func wireModelSwap(srv *ipc.Server, phr phraser.Phraser, cfg *config.Config) {
if cfg.Phraser == nil || len(cfg.Phraser.SwapModels) == 0 {
return
}
lp, ok := phr.(*phraser.LLMPhraser)
if !ok {
log.Printf("model swap: phraser.swap_models is set but there is no llama-server phraser — swap disabled")
return
}
allowed := map[string]bool{}
for _, m := range cfg.Phraser.SwapModels {
allowed[filepath.Clean(m)] = true
}
// The configured model is always swappable back to, listed or not: the way
// out of a bad swap must not depend on remembering to allowlist the model
// you are already running.
allowed[filepath.Clean(cfg.Phraser.ModelPath)] = true
srv.SwapModelFn = func(ctx context.Context, req ipc.SwapModelReq) (ipc.SwapModelResp, error) {
path := filepath.Clean(req.ModelPath)
if !allowed[path] {
log.Printf("model swap: REFUSED %q — not in phraser.swap_models", req.ModelPath)
return ipc.SwapModelResp{}, fmt.Errorf("%w: %q is not in phraser.swap_models", ipc.ErrForbidden, req.ModelPath)
}
res, err := lp.Swap(ctx, phraser.SwapSpec{
ModelPath: path,
NGpuLayers: req.NGpuLayers,
NCtx: req.NCtx,
})
resp := ipc.SwapModelResp{
Model: res.Model,
ModelPath: res.ModelPath,
BaseURL: res.BaseURL,
RolledBack: res.RolledBack,
TookMs: res.Took.Milliseconds(),
}
if err != nil {
// A rolled-back swap is a failure that left a working daemon behind.
// Both halves matter to the caller, so the response is filled in even
// though the error is returned.
log.Printf("model swap: %v", err)
return resp, err
}
return resp, nil
}
srv.ModelStatusFn = func(ctx context.Context) (ipc.ModelStatusResp, error) {
path, ngl, nctx := lp.LiveModel()
base := lp.BaseURL()
resp := ipc.ModelStatusResp{
ModelPath: path,
BaseURL: base,
NGpuLayers: ngl,
NCtx: nctx,
Swappable: cfg.Phraser.SwapModels,
}
if base == "" {
resp.Model = llm.UnknownModel
return resp, nil
}
id, err := llm.ModelID(ctx, base)
if err != nil {
// Report the honest "I could not confirm it" rather than echoing the
// configured filename as if the server had said it.
resp.Model = llm.UnknownModel
return resp, nil
}
resp.Model = id
return resp, nil
}
log.Printf("model swap: enabled, %d allowlisted model(s) — step-up required", len(cfg.Phraser.SwapModels))
}
// llmClientFor builds a completion client on the phraser's llama-server and
// keeps it pointed at the right one across a model swap.
//
// Without the OnSwap registration every holder of a base URL — the LLM router,
// the replier, the mail extractor, the memory evaluator — would keep talking to
// the port of a server that no longer exists, and the daemon would degrade to
// the classifier permanently after the first swap. The client is re-pointed, not
// rebuilt, so nothing that holds it has to know a swap happened.
func llmClientFor(lp *phraser.LLMPhraser, timeout time.Duration) *llm.Client {
c := llm.New(lp.BaseURL(), timeout)
lp.OnSwap(func(base string) { c.SetBaseURL(base) })
return c
}
+142
View File
@@ -0,0 +1,142 @@
package main
import (
"context"
"fmt"
"log"
"strings"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/netscan"
)
// scanBudget — the whole spoken scan, end to end. A voice turn that takes
// longer than this has already failed as a turn, so the scan returns whatever
// it found rather than keeping him waiting.
const scanBudget = 20 * time.Second
// scanReadOut — how many hosts she names out loud. The rest are a count: a
// spoken list of twenty IP addresses is not an answer.
const scanReadOut = 6
// netWiring — the LAN scanner, when the `netscan` block is enabled. nil ⇒ Maven
// never puts a discovery packet on the network.
//
// Unlike the house, a scan is a READ, so it is a query source rather than an
// act: there is no allowlist row and no confirm turn, because nothing changes.
// What makes that safe is that the range is not an argument — see
// internal/netscan's package comment.
type netWiring struct {
scanner *netscan.Scanner
subnets []string
}
// wireNetScan builds the scanner. nil unless the block is enabled and valid.
func wireNetScan(cfg *config.Config) *netWiring {
nc, ok := cfg.NetScanner()
if !ok {
return nil
}
if err := netscan.Validate(nc); err != nil {
// config.validate already ran this, so reaching here is a programming
// error rather than a config one. Not fatal: the scanner off is a
// working Maven.
log.Printf("netscan: not wired: %v", err)
return nil
}
return &netWiring{scanner: netscan.New(nc), subnets: nc.Subnets}
}
// scanSummary answers "какие устройства в сети?" in one line.
func (w *netWiring) scanSummary(ctx context.Context) (string, bool) {
if w == nil {
return "", false
}
ctx, cancel := context.WithTimeout(ctx, scanBudget)
defer cancel()
hosts, err := w.scanner.Scan(ctx)
if err != nil {
log.Printf("netscan: scan: %v", err)
return "не получилось просканировать сеть.", true
}
if len(hosts) == 0 {
return "в сети никого не нашла.", true
}
shown := hosts
if len(shown) > scanReadOut {
shown = shown[:scanReadOut]
}
parts := make([]string, 0, len(shown))
for _, h := range shown {
s := h.Addr
if len(h.Ports) > 0 {
ps := make([]string, 0, len(h.Ports))
for _, p := range h.Ports {
ps = append(ps, fmt.Sprintf("%d", p))
}
s += " (" + strings.Join(ps, ", ") + ")"
}
parts = append(parts, s)
}
out := fmt.Sprintf("нашла %d %s: %s", len(hosts), hostWord(len(hosts)), strings.Join(parts, "; "))
if len(hosts) > len(shown) {
out += fmt.Sprintf(" и ещё %d", len(hosts)-len(shown))
}
return out + ".", true
}
// hostWord — Russian counts inflect the noun: 1 устройство, 2-4 устройства,
// 5+ устройств, and the teens are all the last form.
func hostWord(n int) string {
if n%100 >= 11 && n%100 <= 14 {
return "устройств"
}
switch n % 10 {
case 1:
return "устройство"
case 2, 3, 4:
return "устройства"
default:
return "устройств"
}
}
// isNetworkQuery recognises a question about the LAN, narrowly. It needs a
// network word AND an ask: "интернет не работает" is a complaint, not a request
// to scan, and a scan she runs unasked is exactly the noisy behaviour the
// bounds exist to prevent.
func isNetworkQuery(u string) bool {
s := strings.ToLower(strings.TrimSpace(u))
if s == "" {
return false
}
network := false
for _, w := range []string{"в сети", "в сетке", "сеть", "сети", "локальн", "wifi", "wi-fi", "вайфай"} {
if strings.Contains(s, w) {
network = true
break
}
}
if !network {
return false
}
// An explicit ask to scan, or a phrase that can only be about the LAN.
// "кто в сети" carries no device noun but means nothing else.
for _, w := range []string{"просканируй", "сканируй", "скан", "просканир", "кто в сети", "кто в сетке"} {
if strings.Contains(s, w) {
return true
}
}
ask := strings.Contains(s, "?") || homeWord(s, "какие") || homeWord(s, "кто") ||
homeWord(s, "что") || homeWord(s, "сколько") || strings.Contains(s, "покажи")
if !ask {
return false
}
for _, w := range []string{"устройств", "хост", "компьютер", "машин", "адрес"} {
if strings.Contains(s, w) {
return true
}
}
return false
}
+111
View File
@@ -0,0 +1,111 @@
package main
import (
"context"
"strings"
"testing"
"github.com/kami/maven/internal/config"
)
func TestWireNetScanOffUnlessEnabled(t *testing.T) {
for name, cfg := range map[string]*config.Config{
"no block": {},
"written but dark": {NetScan: &config.NetScanConfig{
Subnets: []string{"192.168.1.0/24"},
}},
"enabled but nothing to scan": {NetScan: &config.NetScanConfig{Enabled: true}},
"enabled but public": {NetScan: &config.NetScanConfig{
Subnets: []string{"8.8.8.0/24"}, Enabled: true,
}},
"enabled but far too wide": {NetScan: &config.NetScanConfig{
Subnets: []string{"10.0.0.0/8"}, Enabled: true,
}},
} {
t.Run(name, func(t *testing.T) {
if w := wireNetScan(cfg); w != nil {
t.Fatal("the scanner must not wire for this config")
}
})
}
var w *netWiring
if _, ok := w.scanSummary(context.Background()); ok {
t.Fatal("a nil wiring must not claim a query")
}
ok := wireNetScan(&config.Config{NetScan: &config.NetScanConfig{
Subnets: []string{"192.168.1.0/24"}, Enabled: true,
}})
if ok == nil {
t.Fatal("a valid enabled block should wire")
}
}
// A loopback /32 with nothing listening on the scanned port: the summary must
// come back honest rather than inventing a host. This also exercises the real
// dialer end to end without touching anything outside this box.
func TestScanSummaryOnAnEmptyRange(t *testing.T) {
w := wireNetScan(&config.Config{NetScan: &config.NetScanConfig{
// Port 1 on loopback: nothing listens and the connection is refused
// immediately, so the scan is fast and touches only this machine.
Subnets: []string{"127.0.0.1/32"}, Ports: []int{1}, Rate: 1000, Enabled: true,
}})
if w == nil {
t.Fatal("wireNetScan returned nil")
}
out, claimed := w.scanSummary(context.Background())
if !claimed {
t.Fatal("the summary did not claim the turn")
}
if out == "" {
t.Fatal("empty summary")
}
// Persona: feminine self-reference, informal address, no pet names.
low := strings.ToLower(out)
for _, bad := range []string{"нашёл", "не смог ", "вы ", "ваш", "милый", "дорогой"} {
if strings.Contains(low, bad) {
t.Errorf("persona violation %q in %q", bad, out)
}
}
}
func TestHostWordAgreesWithTheCount(t *testing.T) {
for n, want := range map[int]string{
1: "устройство", 2: "устройства", 4: "устройства", 5: "устройств",
11: "устройств", 12: "устройств", 21: "устройство", 22: "устройства",
25: "устройств", 111: "устройств", 101: "устройство", 0: "устройств",
} {
if got := hostWord(n); got != want {
t.Errorf("hostWord(%d) = %q, want %q", n, got, want)
}
}
}
func TestIsNetworkQuery(t *testing.T) {
yes := []string{
"какие устройства в сети?",
"кто в сети?",
"просканируй сеть",
"покажи устройства в локальной сети",
"сколько машин в сети",
}
no := []string{
"",
"интернет не работает",
"сеть какая-то медленная",
"я в сети инстаграма",
"что включено дома?",
"напомни оплатить интернет",
}
for _, u := range yes {
if !isNetworkQuery(u) {
t.Errorf("isNetworkQuery(%q) = false, want true", u)
}
}
for _, u := range no {
if isNetworkQuery(u) {
t.Errorf("isNetworkQuery(%q) = true, want false", u)
}
}
}
+120
View File
@@ -0,0 +1,120 @@
// mavend/patterns.go — the shared detect+propose step of pattern inference
// (Vikunja #43). Event *extraction* (fact -> action/object) happens at fact-
// write time in detectPattern below, tied to whichever channel wrote the
// fact. Detection — turning a run of events into a proposed routine — is
// channel-agnostic: it only needs what's already in the events table, so it
// runs both right after a voice fact-write (for the immediate "напоминать?"
// confirmation) and, proactively, from the digestion tick (tick.go's
// detectPatterns) over every action+object pair on record, not just the one
// that was just talked about.
package main
import (
"context"
"errors"
"fmt"
"log"
"time"
"github.com/kami/maven/internal/pattern"
"github.com/kami/maven/internal/store"
)
// detectAndPropose runs the pattern detector over every recorded event for
// action+object and, if a stable pattern is found and nothing has been
// proposed/accepted/dismissed for this pair yet, creates a proposed_routines
// row. Returns (nil, 0, nil) — not an error — whenever there is nothing new
// to report: too few events, irregular intervals, or a pair that already has
// a row in any status. That last case is the one that matters most: it is
// how a routine the owner already DISMISSED stays dismissed forever, because
// the row survives dismissal (status flips in place, see
// store.DismissProposedRoutine) and both the Lookup check here and the
// table's UNIQUE(action, object) constraint refuse to create a second one.
func detectAndPropose(ctx context.Context, ds *store.Store, action, object string, ts time.Time) (*pattern.ProposedRoutine, int64, error) {
events, err := ds.EventsFor(ctx, action, object)
if err != nil {
return nil, 0, fmt.Errorf("events for %s/%s: %w", action, object, err)
}
patEvents := make([]pattern.Event, len(events))
for i, e := range events {
patEvents[i] = pattern.Event{
FactID: e.FactID,
Action: e.Action,
Object: e.Object,
Ts: e.Ts,
}
}
r, err := pattern.Detect(patEvents)
if err != nil {
return nil, 0, fmt.Errorf("detect %s/%s: %w", action, object, err)
}
if r == nil {
return nil, 0, nil // not enough data or intervals too irregular
}
// Belt: check first so the common "nothing new" case never even attempts
// an insert. Suspenders: CreateProposedRoutine's ON CONFLICT DO NOTHING
// (backed by the UNIQUE(action,object) constraint) is the actual
// guarantee — this Lookup is an optimization, not the source of truth.
existing, err := ds.LookupProposedRoutine(ctx, r.Action, r.Object)
if err != nil {
return nil, 0, fmt.Errorf("lookup proposed routine %s/%s: %w", action, object, err)
}
if existing != nil {
return nil, 0, nil // already proposed, accepted, or dismissed — say nothing
}
id, err := ds.CreateProposedRoutine(ctx, r.Action, r.Object, r.IntervalDays, ts)
if err != nil {
if errors.Is(err, store.ErrProposedRoutineExists) {
return nil, 0, nil // lost a race with another caller — not an error
}
return nil, 0, fmt.Errorf("create proposed routine %s/%s: %w", action, object, err)
}
return r, id, nil
}
// detectPattern extracts an event from the written fact and runs the pattern
// detector. If a stable recurring pattern is found and no proposed routine
// exists for this action+object yet, one is created and the user is prompted
// to confirm via the park() mechanism. Returns the suggestion phrase when a
// new proposal was created and parked; "" otherwise.
func (h *reactiveHandler) detectPattern(ctx context.Context, factID int64, key, value string, ts time.Time) string {
ev := pattern.Extract(factID, key, value, ts)
if ev == nil {
return "" // not an actionable event
}
if _, err := h.dataStore.CreateEvent(ctx, factID, ev.Action, ev.Object, ts); err != nil {
log.Printf("voice: create event: %v", err)
return ""
}
// Detect+propose (Vikunja #43) is shared with the digestion tick's
// proactive scan — see detectAndPropose above. Event *extraction* stays
// here, tied to this fact write; detection over the accumulated history does
// not need to happen right now for the voice path to have already done
// its job — it's dedupe-safe to also let the next tick find the same
// pattern independently.
r, id, err := detectAndPropose(ctx, h.dataStore, ev.Action, ev.Object, ts)
if err != nil {
log.Printf("voice: detect pattern %s/%s: %v", ev.Action, ev.Object, err)
return ""
}
if r == nil {
return "" // not enough data, too irregular, or already proposed/decided
}
log.Printf("voice: proposed routine: %s/%s every %.1f days", r.Action, r.Object, r.IntervalDays)
// Park the proposal for voice confirmation.
phrase := pattern.PhraseRoutine(r)
h.mu.Lock()
h.pendingRoutine = &pendingRoutineConfirm{
routineID: id,
action: r.Action,
object: r.Object,
interval: r.IntervalDays,
phrase: phrase,
expiry: ts.Add(confirmTTL),
}
h.mu.Unlock()
return phrase
}
+285
View File
@@ -0,0 +1,285 @@
package main
import (
"context"
"database/sql"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/pattern"
"github.com/kami/maven/internal/store"
)
// seedRefillEvents writes N weekly "refill/cat_water" events straight to the
// events table — this is what the tick reads, independent of any utterance.
func seedRefillEvents(t *testing.T, st *store.Store, ctx context.Context, base time.Time, n int) {
t.Helper()
for i := 0; i < n; i++ {
factID, err := st.WriteFact(ctx, base.Add(time.Duration(i)*7*24*time.Hour), store.KindSelf,
"cat_water", "refill", "test", 1.0, sql.NullInt64{})
if err != nil {
t.Fatalf("write fact %d: %v", i, err)
}
if _, err := st.CreateEvent(ctx, factID, "refill", "cat_water", base.Add(time.Duration(i)*7*24*time.Hour)); err != nil {
t.Fatalf("create event %d: %v", i, err)
}
}
}
// TestTickDetectsPatternFromStoredEvents proves the tick notices a pattern on
// its own, reading straight from the store — not as a side effect of a live
// utterance (Vikunja #43). MinEvents weekly events with no voice turn in
// sight must produce exactly one proposed routine.
func TestTickDetectsPatternFromStoredEvents(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents)
tl := newTestTickLoop(t, st, &fakeSink{}, nil)
tl.detectPatterns(ctx, now, loop.State{})
rows, err := st.ListProposedRoutines(ctx)
if err != nil {
t.Fatalf("list proposed routines: %v", err)
}
if len(rows) != 1 {
t.Fatalf("proposed routines = %d, want 1: %+v", len(rows), rows)
}
if rows[0].Action != "refill" || rows[0].Object != "cat_water" {
t.Errorf("proposed routine = %s/%s, want refill/cat_water", rows[0].Action, rows[0].Object)
}
}
// TestTickPatternDetectionIsIdempotent proves running the tick's pattern scan
// twice does not spam a second proposal for the same pair, and that the store
// itself is what stops the duplicate (not tick-local state) — the whole point
// of the guard, since the tick has no memory of what it proposed last time.
func TestTickPatternDetectionIsIdempotent(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents)
tl := newTestTickLoop(t, st, &fakeSink{}, nil)
tl.detectPatterns(ctx, now, loop.State{})
tl.detectPatterns(ctx, now.Add(time.Hour), loop.State{})
rows, err := st.ListProposedRoutines(ctx)
if err != nil {
t.Fatalf("list proposed routines: %v", err)
}
if len(rows) != 1 {
t.Fatalf("proposed routines after two ticks = %d, want 1 (no duplicate): %+v", len(rows), rows)
}
}
// TestTickPatternDetectionRespectsDismissal proves the single worst failure
// mode here — a proposal the owner already said no to coming back on the next
// tick — cannot happen. Dismissal flips the row's status in place; it must
// still be there to block re-proposal.
func TestTickPatternDetectionRespectsDismissal(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents)
tl := newTestTickLoop(t, st, &fakeSink{}, nil)
tl.detectPatterns(ctx, now, loop.State{})
rows, err := st.ListProposedRoutines(ctx)
if err != nil {
t.Fatalf("list proposed routines: %v", err)
}
if len(rows) != 1 {
t.Fatalf("setup: proposed routines = %d, want 1", len(rows))
}
if err := st.DismissProposedRoutine(ctx, rows[0].ID); err != nil {
t.Fatalf("dismiss: %v", err)
}
// More events for the same pair arrive, and the tick runs again — a
// dismissed pattern must not resurface.
seedRefillEvents(t, st, ctx, now.Add(30*24*time.Hour), pattern.MinEvents)
tl.detectPatterns(ctx, now.Add(60*24*time.Hour), loop.State{})
proposed, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineProposed)
if err != nil {
t.Fatalf("list proposed: %v", err)
}
if len(proposed) != 0 {
t.Fatalf("a dismissed pattern came back: %+v", proposed)
}
all, err := st.ListProposedRoutinesByStatus(ctx, "")
if err != nil {
t.Fatalf("list all: %v", err)
}
if len(all) != 1 {
t.Fatalf("total rows for the pair = %d, want 1 (still dismissed, not duplicated): %+v", len(all), all)
}
if all[0].Status != store.RoutineDismissed {
t.Errorf("status = %s, want dismissed", all[0].Status)
}
}
// proposalRule — the rule name announceProposal uses for the seeded pair.
const proposalRule = "proposal:refill cat_water"
// TestTickProposalSilentByDefault — detection is always on, announcing is not.
// With no pattern_proposals block the tick still records the proposal, and says
// nothing about it: Maven is not autonomous, so a behaviour that speaks without
// being asked stays off until it is configured.
func TestTickProposalSilentByDefault(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents)
markPresent(t, st, ctx, now)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
tl.tick(ctx, now)
if n := countSends(sink, proposalRule); n != 0 {
t.Fatalf("announced %d proposals with no config, want 0", n)
}
rows, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineProposed)
if err != nil {
t.Fatalf("list proposed: %v", err)
}
if len(rows) != 1 {
t.Fatalf("proposed routines = %d, want 1 (silent, but recorded)", len(rows))
}
}
// TestTickAnnouncesProposalWhenConfigured — with notify on, the proposal goes
// out once through the ordinary delivery path, worded by the detector itself.
// Later ticks stay quiet because the pair is already proposed: one pattern is
// one announcement, ever.
func TestTickAnnouncesProposalWhenConfigured(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents)
markPresent(t, st, ctx, now)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
tl.proposalCfg = &config.PatternProposalConfig{Notify: true}
tl.tick(ctx, now)
var got *delivery.Sendable
for i := range sink.sends {
if sink.sends[i].RuleName == proposalRule {
got = &sink.sends[i]
}
}
if got == nil {
t.Fatalf("proposal was not announced; sends=%+v", sink.sends)
}
if !strings.Contains(got.Body, "напоминать?") {
t.Errorf("body = %q, want the detector's own question", got.Body)
}
if got.Channel != delivery.ChannelVoice {
t.Errorf("channel = %v, want voice (sev1, present)", got.Channel)
}
// A month of further ticks: the pair already has a row, so there is
// nothing new to detect and nothing more to say.
sink.sends = nil
later := now.Add(40 * 24 * time.Hour)
markPresent(t, st, ctx, later)
tl.tick(ctx, later)
if n := countSends(sink, proposalRule); n != 0 {
t.Fatalf("re-announced an existing proposal %d times, want 0", n)
}
}
// TestTickProposalRespectsGate — a proposal is the least urgent thing Maven can
// say, so it is sev1 and the restraint gate suppresses it. Away presence means
// it is not announced at all: it is not held, not retried, it just lives on
// /routines. The proposal row is still written — noticing is never gated.
func TestTickProposalRespectsGate(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents)
// no presence probes ⇒ away ⇒ care-class gate blocks.
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
tl.proposalCfg = &config.PatternProposalConfig{Notify: true}
tl.tick(ctx, now)
if n := countSends(sink, proposalRule); n != 0 {
t.Fatalf("away: announced %d proposals, want 0", n)
}
if !tl.lastProposalAt.IsZero() {
t.Error("cooldown clock advanced on a suppressed announcement")
}
rows, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineProposed)
if err != nil {
t.Fatalf("list proposed: %v", err)
}
if len(rows) != 1 {
t.Fatalf("proposed routines = %d, want 1 (detection is never gated)", len(rows))
}
}
// TestTickProposalCooldownSpacesAnnouncements — two patterns detected on the
// same tick must not become two interruptions. The second one waits for the
// cooldown, and is on /routines meanwhile.
func TestTickProposalCooldownSpacesAnnouncements(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
now := refNow()
seedRefillEvents(t, st, ctx, now, pattern.MinEvents)
for i := 0; i < pattern.MinEvents; i++ {
ts := now.Add(time.Duration(i) * 3 * 24 * time.Hour)
factID, err := st.WriteFact(ctx, ts, store.KindSelf, "litter_box", "clean", "test", 1.0, sql.NullInt64{})
if err != nil {
t.Fatalf("write fact: %v", err)
}
if _, err := st.CreateEvent(ctx, factID, "clean", "litter_box", ts); err != nil {
t.Fatalf("create event: %v", err)
}
}
markPresent(t, st, ctx, now)
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
tl.proposalCfg = &config.PatternProposalConfig{Notify: true, Cooldown: config.Duration(24 * time.Hour)}
tl.tick(ctx, now)
announced := 0
for _, s := range sink.sends {
if strings.HasPrefix(s.RuleName, "proposal:") {
announced++
}
}
if announced != 1 {
t.Fatalf("announced %d proposals on one tick, want exactly 1", announced)
}
rows, err := st.ListProposedRoutinesByStatus(ctx, store.RoutineProposed)
if err != nil {
t.Fatalf("list proposed: %v", err)
}
if len(rows) != 2 {
t.Fatalf("proposed routines = %d, want 2 (both recorded, one announced)", len(rows))
}
// Still inside the cooldown: silence, even though a proposal is pending.
sink.sends = nil
soon := now.Add(time.Hour)
markPresent(t, st, ctx, soon)
tl.tick(ctx, soon)
for _, s := range sink.sends {
if strings.HasPrefix(s.RuleName, "proposal:") {
t.Fatalf("announced %q inside the cooldown", s.RuleName)
}
}
}
+158
View File
@@ -0,0 +1,158 @@
package main
import (
"context"
"fmt"
"math"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice"
)
// fixedEmbedder hands back a vector chosen per text, so a test can say exactly
// how close each stored memory is to the question. The real embedders make
// scores that are realistic but not controllable, and this test is about the
// gate, not about the embedder.
type fixedEmbedder struct{ vecs map[string][]float32 }
func (f *fixedEmbedder) Dim() int { return 4 }
func (f *fixedEmbedder) Close() error { return nil }
func (f *fixedEmbedder) Embed(_ context.Context, text string) ([]float32, error) {
v, ok := f.vecs[text]
if !ok {
return nil, fmt.Errorf("fixedEmbedder: no vector for %q", text)
}
return v, nil
}
// scoreVec builds a unit vector whose cosine against the query vector
// (1,0,0,0) is exactly score.
func scoreVec(score float64) []float32 {
rest := math.Sqrt(1 - score*score)
return []float32{float32(score), float32(rest), 0, 0}
}
// recordingPhraser remembers what the query path handed it to phrase, which is
// how the test can tell which pass produced the answer.
type recordingPhraser struct {
*phraser.Stub
notes []string
}
func (r *recordingPhraser) PhraseQuery(ctx context.Context, utterance string, notes []string) (string, error) {
r.notes = notes
return r.Stub.PhraseQuery(ctx, utterance, notes)
}
// recallCase — one stored memory: its text, how close it is to the question,
// whether it is a note or a fact, and whether the notes table holds it too.
type recallCase struct {
text string
score float64
kind string
}
// buildRecallHandler stores the given memories and returns a handler whose
// query path can be run directly. Notes go into BOTH the notes table and the
// vector index, which is what the daemon does (voice.go's IntentNote).
func buildRecallHandler(t *testing.T, question string, mems []recallCase) (*reactiveHandler, *recordingPhraser) {
t.Helper()
ctx := context.Background()
st := newTestStore(t)
emb := &fixedEmbedder{vecs: map[string][]float32{question: {1, 0, 0, 0}}}
mem := memory.NewInMemoryStore()
now := time.Now()
for i, m := range mems {
vec := scoreVec(m.score)
emb.vecs[m.text] = vec
id := fmt.Sprintf("%s:%d", m.kind, i)
if m.kind == "note" {
if _, err := st.WriteNote(ctx, now, m.text, vec, "tap:voice"); err != nil {
t.Fatalf("WriteNote: %v", err)
}
}
if err := mem.Insert(ctx, id, vec, map[string]string{"text": m.text, "type": m.kind}); err != nil {
t.Fatalf("memory insert: %v", err)
}
}
phr := &recordingPhraser{Stub: phraser.NewStub()}
h := &reactiveHandler{
api: ipc.NewStoreAPI(st),
embedder: emb,
replier: voice.NewStubReplier(),
phraser: phr,
now: func() time.Time { return now },
memStore: mem,
dataStore: st,
queryMinScore: 0.55,
queryMinMargin: 0.008,
weatherProvider: nil,
}
return h, phr
}
func askQuery(t *testing.T, h *reactiveHandler, question string) string {
t.Helper()
return h.applyAction(context.Background(), router.Decision{
Intent: router.IntentQuery,
Utterance: question,
})
}
// TestQueryRecallNoteCanWin — the note-recall regression (Vikunja #373). Notes
// and facts share one vector index, and a note that clearly beats everything
// else must be the answer. Before the fix the memory pass only ran after the
// notes-only gate had already rejected the same note at the same score, so only
// a fact could ever come back from it.
func TestQueryRecallNoteCanWin(t *testing.T) {
const q = "где молоко"
t.Run("a clearly best note answers", func(t *testing.T) {
h, phr := buildRecallHandler(t, q, []recallCase{
{text: "молоко стоит в холодильнике", score: 0.90, kind: "note"},
{text: "выучил пару аккордов", score: 0.50, kind: "note"},
})
reply := askQuery(t, h, q)
if want := "вот что я нашла: молоко стоит в холодильнике"; reply != want {
t.Errorf("reply %q, want %q", reply, want)
}
// One text, the winning memory's — the answer came from the memory
// pass, not from handing the phraser every note in the table.
if len(phr.notes) != 1 || phr.notes[0] != "молоко стоит в холодильнике" {
t.Errorf("phraser got %q, want just the recalled note", phr.notes)
}
})
// The other half of "one gate over everything": a fact that matches better
// than the best note now answers, instead of losing to a note that only had
// to beat other notes.
t.Run("the better-matching fact answers", func(t *testing.T) {
h, _ := buildRecallHandler(t, q, []recallCase{
{text: "молоко стоит в холодильнике", score: 0.80, kind: "note"},
{text: "купил молоко в среду", score: 0.95, kind: "fact"},
})
if reply := askQuery(t, h, q); reply != "купил молоко в среду" {
t.Errorf("reply %q, want the fact read back", reply)
}
})
// The gate is untouched: two memories this close mean the embedder cannot
// tell them apart, and silence still beats a coin flip.
t.Run("no clear best stays silent", func(t *testing.T) {
h, _ := buildRecallHandler(t, q, []recallCase{
{text: "молоко стоит в холодильнике", score: 0.860, kind: "note"},
{text: "молоко закончилось", score: 0.858, kind: "note"},
})
if reply := askQuery(t, h, q); reply != "не знаю." {
t.Errorf("reply %q, want silence", reply)
}
})
}
+144
View File
@@ -0,0 +1,144 @@
// Quiet-mode toggle recognition — the pre-route keyword check that lets
// "тихий режим" flip the daemon-wide quiet_hours config without going through
// the router. Moved out of voice.go unchanged (Vikunja #321); the tests live in
// quiet_toggle_test.go.
package main
import (
"context"
"log"
"strings"
"unicode"
"github.com/kami/maven/internal/ipc"
)
// resolveQuietToggle — pre-route keyword check. Returns (reply, true) when
// the utterance is a quiet-on/off command; ("", false) otherwise. Called from
// runTurn BEFORE the router so a classifier miscue can't drop it — which means
// both the voice path and the text path (mavweb /api/chat, telegram) reach it,
// so a false positive here is a network-reachable way to flip a daemon-wide
// setting. See classifyQuietToggle for the matching rule.
func (h *reactiveHandler) resolveQuietToggle(ctx context.Context, text string) (string, bool) {
on, off := classifyQuietToggle(text)
if !on && !off {
return "", false
}
val := "false"
reply := "тихий режим выключен."
if on {
val = "true"
reply = "тихий режим включён. буду реже напоминать."
}
if _, err := h.api.WriteFact(ctx, ipc.WriteFactReq{
Ts: h.now(),
Kind: "config",
Key: "quiet_hours",
Value: val,
Source: "tap:voice",
Confidence: 1.0,
}); err != nil {
log.Printf("voice: write quiet_hours: %v", err)
return "не получилось переключить тихий режим.", true
}
return reply, true
}
// quietInflections — the inflectional endings a stem may carry and still be
// the same word. Adjective/adverb/noun/verb endings, all ≤3 letters. This is
// what separates "тихий"/"тихом"/"тихо" (stem "тих" + a real ending) from
// "тихонько"/"потихоньку", which are different words: "онько" is not an
// ending, and "потихоньку" doesn't start with the stem at all.
var quietInflections = []string{
"", "а", "е", "и", "й", "о", "у", "ы", "ю", "я",
"ая", "ее", "ей", "ем", "ие", "ий", "им", "их", "ия", "ию", "ое", "ой", "ом", "ую", "ые", "ый", "ым", "ых", "ья",
"ами", "ого", "ому", "ыми", "ать", "ить", "ять",
}
// quietStem reports whether tok is the given stem carrying at most one
// inflectional ending. Word boundaries come from tokenisation (see
// quietTokens), not from a regexp — Go's \b is ASCII-oriented and treats every
// Cyrillic letter as a non-word character, so `\bтих\b` would happily match
// inside "тихонько". Comparing whole tokens sidesteps that entirely.
func quietStem(tok, stem string) bool {
if !strings.HasPrefix(tok, stem) {
return false
}
suffix := tok[len(stem):]
for _, e := range quietInflections {
if suffix == e {
return true
}
}
return false
}
// quietTokens splits an utterance into lowercase word tokens, dropping
// punctuation and spacing. Unicode-aware, so Cyrillic words tokenise the same
// way ASCII ones do.
func quietTokens(text string) []string {
return strings.FieldsFunc(strings.ToLower(strings.TrimSpace(text)), func(r rune) bool {
return !unicode.IsLetter(r) && !unicode.IsDigit(r)
})
}
// quietPhrase matches a pattern (a sequence of stems) against the token list.
// Multi-word patterns match any contiguous run of tokens — "включи тихий
// режим" carries "тихий режим". Single-word patterns match ONLY when they are
// the whole utterance: bare "тихо" is a command, but "в комнате тихо" is a
// remark about the room and must not flip a daemon-wide setting.
func quietPhrase(tokens, pattern []string) bool {
if len(pattern) == 0 || len(tokens) < len(pattern) {
return false
}
if len(pattern) == 1 {
return len(tokens) == 1 && quietStem(tokens[0], pattern[0])
}
for i := 0; i+len(pattern) <= len(tokens); i++ {
hit := true
for j, stem := range pattern {
if !quietStem(tokens[i+j], stem) {
hit = false
break
}
}
if hit {
return true
}
}
return false
}
// quietOffPhrases / quietOnPhrases — the toggle vocabulary, as stem sequences.
var (
quietOffPhrases = [][]string{
{"quiet", "off"}, {"quiet", "end"},
{"громк", "режим"}, {"шумн", "режим"},
{"отмен", "тих"}, {"выключ", "тих"}, {"не", "тих"},
}
quietOnPhrases = [][]string{
{"quiet", "on"}, {"quiet", "mode"},
{"тих", "режим"}, {"не", "шум"}, {"не", "беспоко"},
{"тих"},
}
)
// classifyQuietToggle reads an utterance as a quiet-mode command. OFF is
// resolved before ON for the same reason classifyConfirm checks negatives
// first: the OFF phrases are built out of the ON words ("выключи тихий"
// contains "тихий"), so scanning ON first would shadow them and "выключи
// тихий режим" would turn quiet mode on. Negation wins.
func classifyQuietToggle(text string) (on, off bool) {
tokens := quietTokens(text)
for _, p := range quietOffPhrases {
if quietPhrase(tokens, p) {
return false, true
}
}
for _, p := range quietOnPhrases {
if quietPhrase(tokens, p) {
return true, false
}
}
return false, false
}
+114
View File
@@ -0,0 +1,114 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
)
// quietFakeAPI records the WriteFact the toggle performs.
type quietFakeAPI struct {
ipc.UnimplementedCoreAPI
got ipc.WriteFactReq
call int
}
func (a *quietFakeAPI) WriteFact(_ context.Context, req ipc.WriteFactReq) (int64, error) {
a.got, a.call = req, a.call+1
return 1, nil
}
// quietVerdict — what a phrase should do to the setting.
type quietVerdict int
const (
quietNone quietVerdict = iota
quietOn
quietOff
)
func TestResolveQuietToggle(t *testing.T) {
cases := []struct {
text string
want quietVerdict
}{
// ON vocabulary.
{"quiet on", quietOn},
{"quiet mode", quietOn},
{"тихий режим", quietOn},
{"тихий", quietOn},
{"не шуми", quietOn},
{"не беспокоить", quietOn},
{"тихо", quietOn},
// ON, inflected / embedded in a sentence.
{"включи тихий режим", quietOn},
{"побудь в тихом режиме", quietOn},
{"Тихий Режим!", quietOn},
{"тихая", quietOn},
// OFF vocabulary — all seven, incl. the three that used to say ON.
{"quiet off", quietOff},
{"quiet end", quietOff},
{"громкий режим", quietOff},
{"шумный режим", quietOff},
{"отмени тихий", quietOff},
{"выключи тихий", quietOff},
{"не тихо", quietOff},
// OFF wins over the ON words it contains.
{"выключи тихий режим", quietOff},
{"отмени тихий режим пожалуйста", quietOff},
{"верни громкий режим", quietOff},
// False positives: "тихо"/"тихий" as ordinary Russian.
{"очень тихий сегодня день", quietNone},
{"в комнате тихо", quietNone},
{"тихонько напомни", quietNone},
{"потихоньку", quietNone},
{"тихонько", quietNone},
{"он говорил тихим голосом весь вечер", quietNone},
// Unrelated.
{"напомни завтра позвонить маме", quietNone},
{"какая погода", quietNone},
{"", quietNone},
}
for _, tc := range cases {
t.Run(tc.text, func(t *testing.T) {
api := &quietFakeAPI{}
h := &reactiveHandler{api: api, now: func() time.Time { return time.Unix(0, 0).UTC() }}
reply, handled := h.resolveQuietToggle(context.Background(), tc.text)
if tc.want == quietNone {
if handled || reply != "" {
t.Fatalf("%q: got (%q, %v), want no match", tc.text, reply, handled)
}
if api.call != 0 {
t.Fatalf("%q: wrote a fact on a non-match", tc.text)
}
return
}
if !handled {
t.Fatalf("%q: not handled, want %v", tc.text, tc.want)
}
wantReply, wantVal := "тихий режим выключен.", "false"
if tc.want == quietOn {
wantReply, wantVal = "тихий режим включён. буду реже напоминать.", "true"
}
if reply != wantReply {
t.Errorf("%q: reply = %q, want %q", tc.text, reply, wantReply)
}
if api.call != 1 {
t.Fatalf("%q: WriteFact called %d times, want 1", tc.text, api.call)
}
if api.got.Kind != "config" || api.got.Key != "quiet_hours" || api.got.Source != "tap:voice" || api.got.Confidence != 1.0 {
t.Errorf("%q: request shape = %+v", tc.text, api.got)
}
if api.got.Value != wantVal {
t.Errorf("%q: value = %q, want %q", tc.text, api.got.Value, wantVal)
}
})
}
}
+18 -18
View File
@@ -2,24 +2,24 @@ package main
import "github.com/kami/maven/internal/memory"
// bestRecall is the read side of the long-term memory store: the top hit's
// stored text when it clears the confidence gate. This recalls across BOTH
// notes and facts (facts aren't in the notes table, so this is the only path
// that can answer "when did I last …?" from a captured fact). A note hit here
// is redundant with the notes-RAG path — by design; the two indexes can diverge
// once the backend is swapped for a persistent/external store. ok=false when
// there's no hit above the threshold or the hit carries no text.
func bestRecall(results []memory.Result, min float64) (string, bool) {
if len(results) == 0 {
return "", false
// bestRecall is the read side of the long-term memory store: the top hit when
// it clears the confidence gate. The index holds BOTH notes and facts, and
// either can win — the caller looks at the returned hit's meta["type"] to see
// which. Facts aren't in the notes table, so this is the only path that can
// answer "when did I last …?" from a captured fact.
//
// The whole hit is returned, not just its text, because "which memory answered"
// decides how the answer is said: a note gets phrased in Maven's voice, a fact
// is read back as stored.
//
// ok=false when the hit fails the confidence gate (see memory.Confident: an
// absolute floor plus a margin over the runner-up) or carries no text.
func bestRecall(results []memory.Result, minScore, minMargin float64) (memory.Result, bool) {
if !memory.Confident(results, minScore, minMargin) {
return memory.Result{}, false
}
top := results[0]
if top.Score < min {
return "", false
if results[0].Meta["text"] == "" {
return memory.Result{}, false
}
text := top.Meta["text"]
if text == "" {
return "", false
}
return text, true
return results[0], true
}
+38 -6
View File
@@ -8,23 +8,24 @@ import (
func TestBestRecall(t *testing.T) {
const min = 0.55
const margin = 0.008
t.Run("empty results", func(t *testing.T) {
if _, ok := bestRecall(nil, min); ok {
if _, ok := bestRecall(nil, min, margin); ok {
t.Error("empty results returned ok")
}
})
t.Run("top below threshold", func(t *testing.T) {
res := []memory.Result{{Score: 0.4, Meta: map[string]string{"text": "выпил воды"}}}
if _, ok := bestRecall(res, min); ok {
if _, ok := bestRecall(res, min, margin); ok {
t.Error("below-threshold hit returned ok")
}
})
t.Run("hit without text meta", func(t *testing.T) {
res := []memory.Result{{Score: 0.9, Meta: map[string]string{"type": "fact"}}}
if _, ok := bestRecall(res, min); ok {
if _, ok := bestRecall(res, min, margin); ok {
t.Error("textless hit returned ok")
}
})
@@ -34,12 +35,43 @@ func TestBestRecall(t *testing.T) {
{Score: 0.82, Meta: map[string]string{"text": "выпил воды в три часа", "type": "fact"}},
{Score: 0.60, Meta: map[string]string{"text": "другое"}},
}
got, ok := bestRecall(res, min)
got, ok := bestRecall(res, min, margin)
if !ok {
t.Fatal("clearing hit not returned")
}
if got != "выпил воды в три часа" {
t.Errorf("wrong text: %q", got)
if got.Meta["text"] != "выпил воды в три часа" {
t.Errorf("wrong text: %q", got.Meta["text"])
}
if got.Meta["type"] != "fact" {
t.Errorf("kind lost: %q", got.Meta["type"])
}
})
// The index holds notes and facts together, so a note has to be able to win
// it — for a long time it could not (Vikunja #373).
t.Run("a note can win", func(t *testing.T) {
res := []memory.Result{
{Score: 0.86, Meta: map[string]string{"text": "молоко в холодильнике", "type": "note"}},
{Score: 0.61, Meta: map[string]string{"text": "выпил воды", "type": "fact"}},
}
got, ok := bestRecall(res, min, margin)
if !ok {
t.Fatal("clearly-best note not returned")
}
if got.Meta["type"] != "note" || got.Meta["text"] != "молоко в холодильнике" {
t.Errorf("got %v, want the note", got.Meta)
}
})
// The runner-up is almost as close, so the embedder cannot tell the two
// notes apart. Silence beats reading back a coin flip.
t.Run("runner-up too close", func(t *testing.T) {
res := []memory.Result{
{Score: 0.860, Meta: map[string]string{"text": "выпил воды в три часа"}},
{Score: 0.858, Meta: map[string]string{"text": "другое"}},
}
if _, ok := bestRecall(res, min, margin); ok {
t.Error("thin-margin hit returned ok")
}
})
}
+11 -4
View File
@@ -7,6 +7,7 @@ import (
"time"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice"
)
@@ -23,13 +24,19 @@ type completer interface {
type llmReplier struct {
c completer
stub *voice.StubReplier
// block renders the shared context block per turn (who he is, the time).
// nil ⇒ the prompt stands alone.
block func() string
}
func newLLMReplier(c completer) *llmReplier {
return &llmReplier{c: c, stub: voice.NewStubReplier()}
func newLLMReplier(c completer, block func() string) *llmReplier {
return &llmReplier{c: c, stub: voice.NewStubReplier(), block: block}
}
const replySystem = `Ты — Maven, домашняя ассистентка (о себе — в женском роде). Подтверди действие РОВНО ОДНИМ коротким предложением (≤120 символов), тепло и по-русски. Не задавай вопросов, не повторяй слова, не добавляй ничего после точки. Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}.`
const replySystem = `Ты Maven, домашняя ассистентка (о себе в женском роде). Владелец мужчина, говоришь с ним на "ты", в единственном числе; никогда не "вы"/"ваш" и не "он"/"его". Подтверди действие РОВНО ОДНИМ коротким предложением (120 символов), по-русски, спокойно и без официальных формулировок. Не задавай вопросов, не повторяй слова, не добавляй ничего после точки. Отвечай ТОЛЬКО одним объектом JSON с полями "response" (текст) и "mood" (ровно одно из: neutral, happy, thinking, tired, confused).
Пример: {"response": "Записала, что ты выпил стакан воды.", "mood": "neutral"}
Никогда не пиши "..." в поле response.`
func (r *llmReplier) Reply(d router.Decision) string {
if d.Clarify {
@@ -38,7 +45,7 @@ func (r *llmReplier) Reply(d router.Decision) string {
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
out, err := r.c.Complete(ctx, llm.Req{
System: replySystem,
System: persona.Prepend(r.block, replySystem),
User: replyContext(d),
MaxTokens: 512,
RepeatPenalty: 1.3,
+9 -6
View File
@@ -9,12 +9,15 @@ import (
"github.com/kami/maven/internal/voice"
)
type mockCompleter struct{ out string; err error }
type mockCompleter struct {
out string
err error
}
func (m mockCompleter) Complete(_ context.Context, _ llm.Req) (string, error) { return m.out, m.err }
func TestLLMReplierReturnsLLMReply(t *testing.T) {
r := newLLMReplier(mockCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`})
r := newLLMReplier(mockCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`}, nil)
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
if got != "записала, кофе закончился" {
t.Errorf("got %q, want %q", got, "записала, кофе закончился")
@@ -22,7 +25,7 @@ func TestLLMReplierReturnsLLMReply(t *testing.T) {
}
func TestLLMReplierFallsBackToPlainText(t *testing.T) {
r := newLLMReplier(mockCompleter{out: "записала, кофе закончился"})
r := newLLMReplier(mockCompleter{out: "записала, кофе закончился"}, nil)
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
if got != "записала, кофе закончился" {
t.Errorf("got %q, want %q", got, "записала, кофе закончился")
@@ -30,7 +33,7 @@ func TestLLMReplierFallsBackToPlainText(t *testing.T) {
}
func TestLLMReplierFallsBackToStubOnError(t *testing.T) {
r := newLLMReplier(mockCompleter{err: errTestLLMDown})
r := newLLMReplier(mockCompleter{err: errTestLLMDown}, nil)
noteDec := router.Decision{Intent: router.IntentNote}
got := r.Reply(noteDec)
want := voice.NewStubReplier().Reply(noteDec)
@@ -40,7 +43,7 @@ func TestLLMReplierFallsBackToStubOnError(t *testing.T) {
}
func TestLLMReplierFallsBackToStubOnEmpty(t *testing.T) {
r := newLLMReplier(mockCompleter{out: ""})
r := newLLMReplier(mockCompleter{out: ""}, nil)
noteDec := router.Decision{Intent: router.IntentNote}
got := r.Reply(noteDec)
want := voice.NewStubReplier().Reply(noteDec)
@@ -50,7 +53,7 @@ func TestLLMReplierFallsBackToStubOnEmpty(t *testing.T) {
}
func TestLLMReplierClarifyUsesStub(t *testing.T) {
r := newLLMReplier(mockCompleter{out: "я всё поняла"})
r := newLLMReplier(mockCompleter{out: "я всё поняла"}, nil)
clarifyDec := router.Decision{Clarify: true}
got := r.Reply(clarifyDec)
want := voice.NewStubReplier().Reply(clarifyDec)
+180
View File
@@ -0,0 +1,180 @@
// Package main — ruwords.go holds Russian language + calendar/time formatting
// helpers used by the voice reply paths (replySystem, the reminder/routine
// phrasing, etc). Pure functions, no receivers: weekday/month name tables,
// plural agreement, clock/date rendering, and the "do I actually know this
// place/day" guards that pick an honest reply over a confidently wrong one.
// Extend this file rather than voice.go for anything in that shape.
package main
import (
"fmt"
"strconv"
"strings"
"time"
)
var ruWeekdays = []string{
"воскресенье", "понедельник", "вторник", "среда",
"четверг", "пятница", "суббота",
}
var ruMonths = []string{
"января", "февраля", "марта", "апреля", "мая", "июня",
"июля", "августа", "сентября", "октября", "ноября", "декабря",
}
// onlyLocalTimeReply — the honest answer when the user asks the time somewhere
// other than here. She only keeps one clock, and saying so is better than
// naming the wrong city's time.
//
// There used to be a city→time-zone table here. It was removed on purpose: the
// user only ever asks for local time, so the table was a second list of cities
// to keep in step with the weather one for no gain.
const onlyLocalTimeReply = "я знаю только местное время, про другие города пока не скажу."
// notPlaceAfterV — words that follow "в" without naming a place, so
// mentionsUnknownPlace does not mistake them for a city.
var notPlaceAfterV = map[string]bool{
"данный": true, "данную": true, "этот": true, "эту": true,
"котором": true, "какое": true, "какой": true, "который": true,
"общем": true, "точности": true, "курсе": true, "сутках": true,
"часах": true, "минутах": true, "секундах": true, "неделе": true,
}
// mentionsUnknownPlace reports whether the question has a "в <слово>" phrase
// that looks like a place we do not know ("который час в киеве"). Used only to
// pick the honest "local time only" reply instead of answering local time as
// if it were the city's.
func mentionsUnknownPlace(u string) bool {
toks := strings.Fields(u)
for i := 0; i+1 < len(toks); i++ {
if toks[i] != "в" && toks[i] != "во" {
continue
}
next := strings.Trim(toks[i+1], ".,?!")
if next == "" || notPlaceAfterV[next] {
continue
}
// A number after "в" is a clock ("в 5 часов"), not a place.
if _, err := strconv.Atoi(strings.SplitN(next, ":", 2)[0]); err == nil {
continue
}
return true
}
return false
}
// onlyNearDaysReply — she can work out today, tomorrow, the day after and
// yesterday, and nothing further. Said out loud instead of answering today's
// date for a day she did not understand.
const onlyNearDaysReply = "я считаю только сегодня, завтра, послезавтра и вчера — про другие дни пока не скажу."
// dayWords — day references the calendar parser cannot resolve. A weekday name
// or a "через …" phrase means he asked about a specific other day.
var dayWords = []string{
"понедельник", "вторник", "сред", "четверг", "пятниц", "суббот", "воскресен",
"через", "monday", "tuesday", "wednesday", "thursday", "friday", "saturday", "sunday",
}
// mentionsUnknownDay reports whether the question names a day the calendar
// parser could not resolve. Mirror of mentionsUnknownPlace: it exists only to
// pick an honest reply over a confidently wrong one.
//
// Only called after ParseCalendarDate has already failed, so "завтра" and the
// other words it does know never reach here.
func mentionsUnknownDay(u string) bool {
for _, w := range dayWords {
if strings.Contains(u, w) {
return true
}
}
return false
}
// ruClock renders the clock part of the time reply: "15 часов 4 минуты".
func ruClock(t time.Time) string {
h, m := t.Hour(), t.Minute()
hourWord := ruPlural(h, "час", "часа", "часов")
if m == 0 {
return fmt.Sprintf("%d %s ровно", h, hourWord)
}
return fmt.Sprintf("%d %s %d %s", h, hourWord, m, ruPlural(m, "минута", "минуты", "минут"))
}
// dayPrefix names the day relative to now ("завтра", "вчера", …) so the date
// reply opens the way a person would say it.
func dayPrefix(now, day time.Time) string {
base := time.Date(now.Year(), now.Month(), now.Day(), 0, 0, 0, 0, now.Location())
switch int(day.Sub(base).Hours() / 24) {
case -1:
return "вчера"
case 0:
return "сегодня"
case 1:
return "завтра"
case 2:
return "послезавтра"
}
return "это"
}
func ruPlural(n int, one, two, many string) string {
n = n % 100
if n > 10 && n < 20 {
return many
}
n = n % 10
switch n {
case 1:
return one
case 2, 3, 4:
return two
default:
return many
}
}
// hasDurationWords checks whether u is asking about elapsed/remaining time
// rather than the current clock — guards replySystem from replying "сейчас
// X часов" to "сколько времени прошло". Mirrors the stage0.go build filter.
func hasDurationWords(u string) bool {
s := strings.ToLower(strings.TrimSpace(u))
// First-word duration markers (same keywords as timeQueryBuild in stage0).
first := strings.Fields(s)
if len(first) > 0 {
switch first[0] {
case "прошло", "осталось", "пройдет", "минуло", "проходит":
return true
}
}
// Broader duration keywords appearing anywhere in the utterance.
if strings.Contains(s, "прошло") || strings.Contains(s, "осталось") {
return true
}
if strings.Contains(s, " до ") {
return true
}
return false
}
// formatTime returns a human-readable Russian time string for a fact timestamp.
// Used by the query handler when answering "когда я это сделал?"-style questions.
func formatTime(t time.Time) string {
now := time.Now()
if t.After(now.Add(-2*time.Minute)) && t.Before(now.Add(2*time.Minute)) {
return "только что"
}
diff := now.Sub(t)
switch {
case diff < 10*time.Minute:
return "несколько минут назад"
case diff < 60*time.Minute:
return fmt.Sprintf("%d минут назад", int(diff.Minutes()))
case diff < 2*time.Hour:
return "час назад"
case diff < 24*time.Hour:
return fmt.Sprintf("%d часа назад", int(diff.Hours()))
default:
return t.Format("2 января 15:04")
}
}
+798
View File
@@ -0,0 +1,798 @@
// mavend/simulator_test.go — the replayable full-system simulator
// (Vikunja #284, 20-07-2026-BACKLOG.md item 7).
//
// # What it is
//
// A scripted day, replayed through the real mavend code paths, with every
// boundary faked and the clock under the scenario's control. A scenario is a
// JSON file in testdata/scenarios; the harness reads it, builds a world, walks
// the steps in order, and asserts on what actually happened:
//
// what Maven SAID — the reply text of every utterance
// what was SENT — every delivery.Sendable the dispatcher emitted
// what ARRIVED — the unified intake journal from #283
// what TOOLS were called — the recorded requests against fake Praxis/Nexis/Hexis
// what did NOT happen — expect_no_send / expect_no_call, first-class
//
// The last one is the point. Maven's hard constraints are mostly negative —
// not a nag, not autonomous, nothing executed without confirmation — and a
// harness that can only assert on things that happened cannot test any of
// them. "Nothing was sent" is an assertion here, not an absence of one.
//
// # Determinism
//
// No time.Now() runs inside a replay. The scenario names a start instant, each
// step names a wall-clock offset from it, and the harness advances a fakeClock
// to that offset before running the step. Every clock reader in the world —
// the handler's `now`, the tick loop's `tick(ctx, now)`, the intake journal's
// publish stamp — is wired to that clock. Two runs of the same file produce
// the same transcript, and a scenario about 08:35 does not behave differently
// at 03:00 in CI.
//
// The tick is driven by the scenario, not by a ticker: tick() already takes
// `now` as an argument, so the only thing the daemon's ticker contributed was
// wall-clock timing, which is exactly what a replay must not have.
//
// # Why this shape and not a binary
//
// Vikunja #288 (golden-audio STT) deferred its tier-2 "audio → STT → router →
// phraser" scenarios to this task, and asked that they reuse a fixture format
// rather than inventing a third. A scenario here can name a WAV from
// cmd/mavsttd/testdata and the harness will feed it through the STT seam. As a
// test it runs under `make test` on every change, which a separate binary
// would not.
//
// # Production is untouched
//
// Every file this task adds is a _test.go file or testdata. There is no
// simulator in the daemon, no flag, no config key, and no code path that
// checks whether a simulation is running. The seams it uses — stt.Transcriber,
// tts.Synthesizer, router.Completer, delivery.Sink, ipc.CoreAPI, the
// event.Bus from #283 — all already existed for the production wiring.
package main
import (
"context"
"encoding/json"
"fmt"
"os"
"path/filepath"
"strings"
"sync"
"testing"
"time"
"github.com/kami/maven/internal/audio"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/event"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/tool"
"github.com/kami/maven/internal/voice"
)
// ---------------------------------------------------------------------------
// Scenario format
// ---------------------------------------------------------------------------
// scenario — one scripted day. schema_version matches the convention already
// set by testdata/system_safety_scenarios.json.
type scenario struct {
SchemaVersion int `json:"schema_version"`
Name string `json:"name"`
Description string `json:"description,omitempty"`
// Start — the instant the day begins, RFC3339. Every step offset is
// relative to it, and nothing in the run reads a real clock.
Start string `json:"start"`
// Script — what the resident model answers. The world has no llama-server;
// see scriptedLLM for how an entry is chosen.
Script []scriptEntry `json:"script,omitempty"`
// Praxis / Nexus / Hexis — canned bodies for the ecosystem fakes. Absent ⇒
// that service is not wired at all, which is the default box.
Praxis string `json:"praxis_attention,omitempty"`
Nexus string `json:"nexus_resolve,omitempty"`
Hexis string `json:"hexis_capabilities,omitempty"`
Steps []step `json:"steps"`
}
// scriptEntry — one canned model answer. Match is a substring of the user
// message; the first entry whose Match is contained in it wins, and an entry
// with an empty Match is the catch-all.
//
// Route and Reply are separate because the same model serves both contracts
// (CLAUDE.md, "LLM output contract"): a grammar-constrained call is a routing
// call and gets Route, an unconstrained one is a phrasing call and gets Reply.
type scriptEntry struct {
Match string `json:"match"`
Route string `json:"route,omitempty"`
Reply string `json:"reply,omitempty"`
}
// step — one scripted moment. At is "HH:MM" or "HH:MM:SS", interpreted in the
// start instant's location; the clock is advanced to it before the step runs.
//
// A step does exactly one thing (say / audio / signal / fact / tick / arrive)
// and then asserts. Assertions are evaluated against everything recorded since
// the run began, except expect_no_send and expect_no_call, which are scoped to
// this step — "nothing was sent because of THIS" is the useful question.
type step struct {
At string `json:"at"`
Note string `json:"note,omitempty"`
// --- stimuli (at most one per step) ---
// Say — an utterance, as text, through the same runTurn the IPC chat path
// uses.
Say string `json:"say,omitempty"`
// Audio — a WAV under cmd/mavsttd/testdata, fed through the STT seam. This
// is #288's deferred tier 2. The harness uses the deterministic stt stub
// unless a real transcriber is available, so the assertion a scenario can
// make about an audio step is about the PIPELINE, not about whisper's
// accuracy — that is what cmd/mavsttd/golden_test.go is for.
Audio string `json:"audio,omitempty"`
// Signal — a presence/world fact arriving from a poller or /api/signal.
Signal *signalStep `json:"signal,omitempty"`
// Arrive — an intake write from a module: an ambient notification, a feed
// item, a mail candidate. Goes through the same decorated ipc.CoreAPI the
// daemon gives those callers, so it lands in the journal exactly as it
// would in production.
Arrive *arriveStep `json:"arrive,omitempty"`
// Tick — run one iteration of the proactive loop at this instant.
Tick bool `json:"tick,omitempty"`
// Fault — make every ecosystem fake answer with this HTTP status from now
// on. The degraded-mode lever; ClearFault puts them back.
Fault int `json:"fault,omitempty"`
ClearFault bool `json:"clear_fault,omitempty"`
// --- assertions ---
ExpectReply []string `json:"expect_reply_contains,omitempty"`
ExpectNotReply []string `json:"expect_reply_lacks,omitempty"`
ExpectSent []string `json:"expect_sent_contains,omitempty"`
ExpectNoSend bool `json:"expect_no_send,omitempty"`
ExpectCalled []string `json:"expect_called,omitempty"`
ExpectNotCalled []string `json:"expect_not_called,omitempty"`
ExpectEvents []string `json:"expect_events,omitempty"`
ExpectNoEvents bool `json:"expect_no_events,omitempty"`
}
type signalStep struct {
Key string `json:"key"`
Value string `json:"value"`
Source string `json:"source"`
Kind string `json:"kind,omitempty"`
}
type arriveStep struct {
// Note / Fact / Task — exactly one. Each mirrors the intake seam its real
// caller uses.
Note *arriveNote `json:"note,omitempty"`
Fact *signalStep `json:"fact,omitempty"`
Task *arriveTask `json:"task,omitempty"`
AsOf string `json:"as_of,omitempty"` // "HH:MM" — OccurredAt, when it differs from the step time
Source string `json:"source"`
}
type arriveNote struct {
Text string `json:"text"`
}
type arriveTask struct {
Text string `json:"text"`
Evidence string `json:"evidence,omitempty"`
Status string `json:"status,omitempty"`
}
// ---------------------------------------------------------------------------
// The world
// ---------------------------------------------------------------------------
// simWorld — every faked boundary plus the real components between them.
type simWorld struct {
t *testing.T
clock *fakeClock
loc *time.Location
start time.Time
store *store.Store
api ipc.CoreAPI // the intake-decorated adapter, same as the daemon builds
bus *event.Bus
handler *reactiveHandler
tick *tickLoop
sink *recordingSink
llm *scriptedLLM
praxis *fakeServer
nexus *fakeServer
hexis *fakeServer
// transcript — everything that happened, in order. Printed on failure so a
// broken scenario is diagnosable without a debugger.
transcript []string
replies []string
}
// recordingSink captures every send, mutex-guarded (the tick loop dispatches
// from its own goroutine in production and the race detector is on here).
type recordingSink struct {
mu sync.Mutex
sends []delivery.Sendable
}
func (s *recordingSink) Send(_ context.Context, d delivery.Sendable) error {
s.mu.Lock()
defer s.mu.Unlock()
s.sends = append(s.sends, d)
return nil
}
func (s *recordingSink) all() []delivery.Sendable {
s.mu.Lock()
defer s.mu.Unlock()
out := make([]delivery.Sendable, len(s.sends))
copy(out, s.sends)
return out
}
func (s *recordingSink) count() int {
s.mu.Lock()
defer s.mu.Unlock()
return len(s.sends)
}
// scriptedLLM stands in for llama-server on BOTH contracts the resident model
// serves: grammar-constrained routing and unconstrained phrasing.
//
// It is not a stub that ignores its input — a scenario that scripts an answer
// for "что я пропустил" and gets asked something else must fail, not silently
// return the wrong intent. An unmatched call returns an error, and the router
// then falls through to the classifier cascade exactly as it does in
// production when llama-server is unreachable. That fall-through is itself
// worth exercising: it is the failure floor CLAUDE.md refuses to let rot.
type scriptedLLM struct {
mu sync.Mutex
entries []scriptEntry
calls []llm.Req
}
func (s *scriptedLLM) Complete(_ context.Context, r llm.Req) (string, error) {
s.mu.Lock()
defer s.mu.Unlock()
s.calls = append(s.calls, r)
routing := r.Grammar != ""
for _, e := range s.entries {
if e.Match != "" && !strings.Contains(strings.ToLower(r.User), strings.ToLower(e.Match)) {
continue
}
if routing && e.Route != "" {
return e.Route, nil
}
if !routing && e.Reply != "" {
return e.Reply, nil
}
}
return "", fmt.Errorf("simulator: no scripted %s answer for %q",
map[bool]string{true: "route", false: "reply"}[routing], truncateRunes(r.User, 60))
}
// ---------------------------------------------------------------------------
// Building the world
// ---------------------------------------------------------------------------
func newSimWorld(t *testing.T, sc scenario) *simWorld {
t.Helper()
start, err := time.Parse(time.RFC3339, sc.Start)
if err != nil {
t.Fatalf("scenario %q: bad start %q: %v", sc.Name, sc.Start, err)
}
clock := newFakeClock(start)
st := newTestStore(t)
bus := event.NewBus(512)
// The same decorator the daemon wires, on the same clock: intake in a
// replay is journalled exactly as it is in production.
api := newIntakeAPI(ipc.NewStoreAPI(st), bus, clock.Now)
sink := &recordingSink{}
rules := loop.DefaultRules()
gatherer := loop.NewGatherer(st, rules)
dispatcher := delivery.NewDispatcher(delivery.Config{
Voice: sink, Ntfy: sink, Telegram: sink, Nudges: st, Reminders: st,
})
tl := newTickLoop(st, gatherer, dispatcher, phraser.NewStub(), rules,
time.Minute, 5*time.Minute, 0, nil, nil, nil, nil)
scripted := &scriptedLLM{entries: sc.Script}
w := &simWorld{
t: t, clock: clock, loc: start.Location(), start: start,
store: st, api: api, bus: bus, tick: tl, sink: sink, llm: scripted,
}
// Ecosystem fakes, wired only when the scenario supplies a body — a box
// with no praxis block has no praxis client, and a scenario must be able to
// reproduce that.
eco := &ecosystemWiring{}
if sc.Praxis != "" {
w.praxis = newFakePraxis(t, sc.Praxis)
eco.praxis = newPraxisClient(w.praxis.URL)
}
if sc.Nexus != "" {
w.nexus = newFakeNexus(t, sc.Nexus)
}
if sc.Hexis != "" {
w.hexis = newFakeHexis(t, sc.Hexis, fixtureHexisExecuted("exec_1", "completed"))
}
// The router: the same cascade the daemon builds — stage-0 grammars, the
// LLM router on the scripted model, the classifier underneath. Keeping the
// classifier in is deliberate; it is the failure floor, and a scenario that
// scripts no route for an utterance exercises it.
emb := router.NewHashEmbedder(1024)
matcher := tool.NewMatcher(nil)
rtr := buildRouter(emb, matcher, config.DefaultRouterThreshold, router.NewLLMRouter(scripted))
w.handler = &reactiveHandler{
stt: simTranscriber{},
tts: simSynthesizer{},
router: rtr,
embedder: emb,
api: api,
matcher: matcher,
phraser: phraser.NewStub(),
replier: newLLMReplier(scripted, nil),
now: clock.Now,
memStore: st.VectorMemory(),
dataStore: st,
queryMinScore: config.DefaultQueryMinScore,
queryMinMargin: config.DefaultQueryMinMargin,
timeParser: router.StubDateTimeParser{},
dialogueSessions: dialogue.NewSessionStore(time.Hour),
clarifyStore: dialogue.NewClarifyStore(time.Hour),
clarifyMaxAttempts: dialogue.DefaultMaxAttempts,
ecosystem: eco,
}
return w
}
// simTranscriber — the STT seam. Deterministic by construction: it returns the
// text the harness parked for this step, so the pipeline under test is
// "audio arrives → a turn runs", not "whisper heard correctly". Transcription
// accuracy is cmd/mavsttd/golden_test.go's job (#288 tier 1), and duplicating
// it here would make every scenario depend on a 500 MB model.
type simTranscriber struct{ text string }
func (s simTranscriber) Transcribe(_ context.Context, _ audio.Audio) (string, float64, error) {
return s.text, 1.0, nil
}
// simSynthesizer — the TTS seam. A scenario asserts on what Maven SAID, which
// is the reply text; the waveform is not the artefact under test.
type simSynthesizer struct{}
func (simSynthesizer) Synthesize(_ context.Context, _ string) (audio.Audio, error) {
return audio.Audio{Format: audio.PCM16kMono}, nil
}
// ---------------------------------------------------------------------------
// Running
// ---------------------------------------------------------------------------
func (w *simWorld) logf(format string, args ...any) {
w.transcript = append(w.transcript,
fmt.Sprintf("%s %s", w.clock.Now().In(w.loc).Format("15:04:05"), fmt.Sprintf(format, args...)))
}
// dump prints the whole transcript. Called on any failure — a scenario that
// broke on step 7 is unreadable without the six steps before it.
func (w *simWorld) dump() {
w.t.Logf("--- replay transcript ---\n%s", strings.Join(w.transcript, "\n"))
}
// advanceTo moves the clock to the step's offset. Time only ever moves
// FORWARD: a scenario with steps out of order is a bug in the scenario, and
// silently reordering it would hide the bug.
func (w *simWorld) advanceTo(at string) {
w.t.Helper()
if at == "" {
return
}
target := w.timeOf(at)
now := w.clock.Now()
if target.Before(now) {
w.t.Fatalf("step at %s goes backwards from %s — scenario steps must be in order",
at, now.In(w.loc).Format("15:04:05"))
}
w.clock.Advance(target.Sub(now))
}
// timeOf resolves an "HH:MM" or "HH:MM:SS" step offset against the scenario's
// start day and location.
func (w *simWorld) timeOf(at string) time.Time {
w.t.Helper()
layout := "15:04"
if strings.Count(at, ":") == 2 {
layout = "15:04:05"
}
hm, err := time.Parse(layout, at)
if err != nil {
w.t.Fatalf("bad step time %q: %v", at, err)
}
return time.Date(w.start.Year(), w.start.Month(), w.start.Day(),
hm.Hour(), hm.Minute(), hm.Second(), 0, w.loc)
}
func (w *simWorld) run(sc scenario) {
ctx := context.Background()
for i, s := range sc.Steps {
w.advanceTo(s.At)
if s.Note != "" {
w.logf("# %s", s.Note)
}
sendsBefore := w.sink.count()
callsBefore := w.callCount()
eventsBefore := w.bus.Len()
w.stimulate(ctx, s)
w.assert(i, s, sendsBefore, callsBefore, eventsBefore)
}
}
func (w *simWorld) stimulate(ctx context.Context, s step) {
if s.Fault != 0 || s.ClearFault {
for _, fs := range []*fakeServer{w.praxis, w.nexus, w.hexis} {
if fs != nil {
fs.SetFault(s.Fault)
}
}
w.logf("fault=%d on every ecosystem fake", s.Fault)
}
switch {
case s.Say != "":
reply := w.handler.runTurn(ctx, s.Say)
w.replies = append(w.replies, reply)
w.logf("он: %s", s.Say)
w.logf("она: %s", reply)
case s.Audio != "":
text := w.audioText(s.Audio)
// Swap in a transcriber parked with this step's text, then run the same
// push-to-talk entry point the voice client calls.
w.handler.stt = simTranscriber{text: text}
resp, err := w.handler.HandlePushToTalk(ctx, voicePTT(), 0)
if err != nil {
w.t.Fatalf("push-to-talk on %s: %v", s.Audio, err)
}
w.replies = append(w.replies, resp.ReplyText)
w.logf("[wav %s → %q]", filepath.Base(s.Audio), text)
w.logf("она: %s", resp.ReplyText)
case s.Signal != nil:
w.write(ctx, *s.Signal, w.clock.Now())
w.logf("сигнал: %s=%s (%s)", s.Signal.Key, s.Signal.Value, s.Signal.Source)
case s.Arrive != nil:
w.arrive(ctx, *s.Arrive)
case s.Tick:
w.tick.tick(ctx, w.clock.Now())
w.logf("tick")
}
}
func (w *simWorld) write(ctx context.Context, sig signalStep, ts time.Time) {
w.t.Helper()
kind := sig.Kind
if kind == "" {
kind = "env"
}
if _, err := w.api.WriteFact(ctx, ipc.WriteFactReq{
Ts: ts, Kind: kind, Key: sig.Key, Value: sig.Value, Source: sig.Source, Confidence: 1.0,
}); err != nil {
w.t.Fatalf("write fact %s: %v", sig.Key, err)
}
}
func (w *simWorld) arrive(ctx context.Context, a arriveStep) {
w.t.Helper()
// AsOf is when the thing HAPPENED, which for a feed item or a relayed
// notification is usually earlier than when Maven heard about it. It does
// not move the clock — only the timestamp on the row and the envelope.
ts := w.clock.Now()
if a.AsOf != "" {
ts = w.timeOf(a.AsOf)
}
switch {
case a.Fact != nil:
f := *a.Fact
if f.Source == "" {
f.Source = a.Source
}
w.write(ctx, f, ts)
w.logf("пришло: факт %s=%s (%s)", f.Key, f.Value, f.Source)
case a.Note != nil:
if _, err := w.api.WriteNote(ctx, ts, a.Note.Text, nil, a.Source); err != nil {
w.t.Fatalf("write note from %s: %v", a.Source, err)
}
w.logf("пришло: заметка от %s — %s", a.Source, truncateRunes(a.Note.Text, 60))
case a.Task != nil:
status := a.Task.Status
if status == "" {
status = store.TaskCandidate
}
if _, err := w.api.CaptureTask(ctx, ipc.CaptureTaskReq{
Text: a.Task.Text, Source: a.Source, Evidence: a.Task.Evidence, Status: status, Ts: ts,
}); err != nil {
w.t.Fatalf("capture task from %s: %v", a.Source, err)
}
w.logf("пришло: задача от %s — %s", a.Source, a.Task.Text)
default:
w.t.Fatalf("arrive step from %s carries nothing", a.Source)
}
}
// audioText resolves a scenario's WAV reference to the text the fixture is
// known to contain, by reading cmd/mavsttd's golden manifest (#288's format,
// reused rather than duplicated). An unknown reference fails the scenario
// rather than quietly transcribing to "".
func (w *simWorld) audioText(ref string) string {
w.t.Helper()
manifest := filepath.Join("..", "mavsttd", "testdata", "golden_v1.json")
raw, err := os.ReadFile(manifest)
if err != nil {
w.t.Fatalf("audio step %q: reading %s: %v", ref, manifest, err)
}
var m struct {
Cases []struct {
Name string `json:"name"`
WAV string `json:"wav"`
Text string `json:"text"`
} `json:"cases"`
}
if err := json.Unmarshal(raw, &m); err != nil {
w.t.Fatalf("audio step %q: parsing %s: %v", ref, manifest, err)
}
for _, c := range m.Cases {
if c.Name == ref || c.WAV == ref {
return c.Text
}
}
w.t.Fatalf("audio step %q: no such case in %s", ref, manifest)
return ""
}
func voicePTT() voice.PushToTalkReq {
return voice.PushToTalkReq{Audio: audio.Audio{Format: audio.PCM16kMono}}
}
// callCount — how many requests every wired ecosystem fake has seen.
func (w *simWorld) callCount() int {
n := 0
for _, fs := range []*fakeServer{w.praxis, w.nexus, w.hexis} {
if fs != nil {
n += len(fs.Requests())
}
}
return n
}
func (w *simWorld) callPaths() []string {
var out []string
for _, fs := range []*fakeServer{w.praxis, w.nexus, w.hexis} {
if fs == nil {
continue
}
for _, r := range fs.Requests() {
out = append(out, r.Method+" "+r.Path)
}
}
return out
}
// ---------------------------------------------------------------------------
// Assertions
// ---------------------------------------------------------------------------
func (w *simWorld) assert(i int, s step, sendsBefore, callsBefore, eventsBefore int) {
w.t.Helper()
where := fmt.Sprintf("step %d (%s)", i+1, s.At)
if s.Note != "" {
where += " " + s.Note
}
fail := func(format string, args ...any) {
w.dump()
w.t.Errorf("%s: %s", where, fmt.Sprintf(format, args...))
}
lastReply := ""
if len(w.replies) > 0 {
lastReply = w.replies[len(w.replies)-1]
}
for _, want := range s.ExpectReply {
if !containsFold(lastReply, want) {
fail("reply %q does not contain %q", lastReply, want)
}
}
for _, unwanted := range s.ExpectNotReply {
if containsFold(lastReply, unwanted) {
fail("reply %q contains %q and must not", lastReply, unwanted)
}
}
sent := w.sink.all()
for _, want := range s.ExpectSent {
if !anyContains(sendableTexts(sent), want) {
fail("nothing sent mentions %q; sent so far: %v", want, sendableTexts(sent))
}
}
// Scoped to this step on purpose: "nothing was sent BECAUSE OF THIS" is the
// question a not-a-nag constraint asks.
if s.ExpectNoSend && len(sent) > sendsBefore {
fail("expected nothing to be sent, got %v", sendableTexts(sent[sendsBefore:]))
}
paths := w.callPaths()
for _, want := range s.ExpectCalled {
if !anyContains(paths, want) {
fail("no ecosystem call matches %q; calls so far: %v", want, paths)
}
}
for _, unwanted := range s.ExpectNotCalled {
if anyContains(paths[callsBefore:], unwanted) {
fail("an ecosystem call matched %q and must not have: %v", unwanted, paths[callsBefore:])
}
}
evs := w.bus.Recent(0)
for _, want := range s.ExpectEvents {
if !anyContains(eventLines(evs), want) {
fail("no intake event matches %q; journal: %v", want, eventLines(evs))
}
}
if s.ExpectNoEvents && w.bus.Len() > eventsBefore {
fail("expected nothing to arrive, journal grew to %d", w.bus.Len())
}
}
func sendableTexts(sends []delivery.Sendable) []string {
out := make([]string, 0, len(sends))
for _, s := range sends {
out = append(out, fmt.Sprintf("[%s] %s", s.RuleName, s.Body))
}
return out
}
func eventLines(evs []event.Event) []string {
out := make([]string, 0, len(evs))
for _, e := range evs {
out = append(out, fmt.Sprintf("%s/%s %s %s", e.Source, e.Kind, e.Title, e.Body))
}
return out
}
func containsFold(hay, needle string) bool {
return strings.Contains(strings.ToLower(hay), strings.ToLower(needle))
}
func anyContains(hay []string, needle string) bool {
for _, h := range hay {
if containsFold(h, needle) {
return true
}
}
return false
}
// ---------------------------------------------------------------------------
// The test
// ---------------------------------------------------------------------------
const scenarioDir = "testdata/scenarios"
// TestSimulatorScenarios replays every scenario file. Adding a scenario is
// adding a JSON file — no Go change, which is the property that makes this
// cheap enough to actually use.
func TestSimulatorScenarios(t *testing.T) {
entries, err := os.ReadDir(scenarioDir)
if err != nil {
t.Fatalf("reading %s: %v", scenarioDir, err)
}
var ran int
for _, ent := range entries {
if ent.IsDir() || !strings.HasSuffix(ent.Name(), ".json") {
continue
}
ran++
name := strings.TrimSuffix(ent.Name(), ".json")
t.Run(name, func(t *testing.T) {
sc := loadScenario(t, filepath.Join(scenarioDir, ent.Name()))
w := newSimWorld(t, sc)
w.run(sc)
if testing.Verbose() {
w.dump()
}
})
}
if ran == 0 {
t.Fatalf("no scenarios in %s — the harness would pass vacuously", scenarioDir)
}
}
func loadScenario(t *testing.T, path string) scenario {
t.Helper()
raw, err := os.ReadFile(path)
if err != nil {
t.Fatalf("reading %s: %v", path, err)
}
var sc scenario
dec := json.NewDecoder(strings.NewReader(string(raw)))
dec.DisallowUnknownFields() // a typo'd assertion key must fail, not be ignored
if err := dec.Decode(&sc); err != nil {
t.Fatalf("parsing %s: %v", path, err)
}
if sc.SchemaVersion != 1 {
t.Fatalf("%s: schema_version = %d, want 1", path, sc.SchemaVersion)
}
if sc.Name == "" || sc.Start == "" || len(sc.Steps) == 0 {
t.Fatalf("%s: a scenario needs a name, a start and at least one step", path)
}
return sc
}
// TestSimulatorIsDeterministic replays one scenario twice and requires an
// identical transcript. This is the property the whole task rests on: if a
// time.Now() creeps into a replayed path, two runs diverge and this fails.
func TestSimulatorIsDeterministic(t *testing.T) {
path := filepath.Join(scenarioDir, "morning_missed.json")
sc := loadScenario(t, path)
transcriptOf := func() string {
w := newSimWorld(t, sc)
w.run(sc)
return strings.Join(w.transcript, "\n")
}
first := transcriptOf()
second := transcriptOf()
if first != second {
t.Errorf("two replays of the same scenario diverged:\n--- first ---\n%s\n--- second ---\n%s", first, second)
}
// And the transcript's own timestamps must be the scenario's, not today's.
if strings.Contains(first, time.Now().Format("15:04")) && !strings.Contains(sc.Start, time.Now().Format("15:04")) {
t.Error("transcript carries the wall clock — something in the replay path read time.Now()")
}
}
// TestSimulatorRefusesBackwardsSteps guards the one scenario-authoring mistake
// that would silently produce a meaningless run.
func TestSimulatorRefusesBackwardsSteps(t *testing.T) {
// Not table-driven through run() because advanceTo calls t.Fatalf; this
// checks the ordering arithmetic directly.
sc := scenario{SchemaVersion: 1, Name: "x", Start: "2026-08-01T08:30:00+03:00",
Steps: []step{{At: "09:00"}}}
w := newSimWorld(t, sc)
w.advanceTo("09:00")
if got := w.clock.Now().In(w.loc).Format("15:04"); got != "09:00" {
t.Fatalf("clock at %s after advancing to 09:00", got)
}
w.advanceTo("09:30")
if got := w.clock.Now().In(w.loc).Format("15:04"); got != "09:30" {
t.Fatalf("clock at %s after advancing to 09:30", got)
}
}
+215
View File
@@ -0,0 +1,215 @@
package main
import (
"context"
"log"
"strings"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/smarthome"
"github.com/kami/maven/internal/store"
)
// homeWiring — the Home Assistant client, when the `smarthome` block is present
// AND enabled. nil ⇒ the house is not wired, nothing was proposed, and an
// allowlist row that happens to look like a house row refuses to run.
//
// It lives on the voice wiring for the same reason MCP does: a house control IS
// an act. It goes through tool.Executor, the enabled allowlist and the confirm
// turn, all of which only exist on the voice/chat path.
type homeWiring struct {
client *smarthome.Client
st *store.Store
refresh time.Duration
}
// wireSmartHome builds the client and proposes what it found. It never fails
// the daemon: an instance that is down at boot is logged and retried, because
// Maven starting is not contingent on someone else's process.
func wireSmartHome(cfg *config.Config, st *store.Store) *homeWiring {
hc, ok := cfg.SmartHomeClient()
if !ok || st == nil {
return nil
}
if err := smarthome.Validate(hc); err != nil {
// config.validate already ran this, so reaching here is a programming
// error rather than a config one. Still not fatal: the house off is a
// working Maven.
log.Printf("smarthome: not wired: %v", err)
return nil
}
w := &homeWiring{
client: smarthome.NewClient(hc),
st: st,
refresh: time.Duration(cfg.SmartHome.Refresh),
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
w.propose(ctx)
return w
}
// caller is the tool.HomeCaller seam.
func (w *homeWiring) caller() *smarthome.Client {
if w == nil {
return nil
}
return w.client
}
// propose writes a 'proposed' allowlist row for every controllable device. It
// does NOT enable anything: a reachable house is a place Maven may look, not a
// set of switches she may flip. Kami enables what he wants on /tools, behind
// step-up, which is the same gate a shell tool goes through.
//
// Sensors are read but never proposed — there is nothing to call on them.
func (w *homeWiring) propose(ctx context.Context) {
if w == nil {
return
}
ents, err := w.client.States(ctx)
if err != nil {
log.Printf("smarthome: read states: %v", err)
return
}
now := time.Now()
fresh, devices := 0, 0
for _, e := range ents {
svcs := smarthome.Services(e.Domain)
if len(svcs) == 0 {
continue
}
devices++
for _, s := range svcs {
name := smarthome.LocalName(e.ID, s.Verb)
provenance := "дом: " + s.Name + " → " + e.Name + " (" + e.ID + ")"
ok, err := w.st.ProposeSmartHomeTool(ctx, name, smarthome.Scope(e.Domain),
smarthome.Cmd(e.ID, s.Name), provenance, now)
if err != nil {
log.Printf("smarthome: propose %s: %v", name, err)
continue
}
if ok {
fresh++
}
}
}
log.Printf("smarthome: %d entities, %d controllable", len(ents), devices)
if fresh > 0 {
log.Printf("smarthome: %d new device proposal(s) waiting on /tools", fresh)
}
}
// run re-enumerates the house and picks up devices that appeared, until ctx is
// canceled.
func (w *homeWiring) run(ctx context.Context) {
if w == nil {
return
}
iv := w.refresh
if iv <= 0 {
iv = config.DefaultSmartHomeRefresh
}
t := time.NewTicker(iv)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
w.propose(ctx)
}
}
}
// homeSummary answers "что дома?" — a read of the current entity states, one
// short line. Read-only: it can never call a service, so it needs no confirm
// and no allowlist row.
func (w *homeWiring) homeSummary(ctx context.Context) (string, bool) {
if w == nil {
return "", false
}
ents, err := w.client.States(ctx)
if err != nil {
log.Printf("smarthome: summary: %v", err)
return "не смогла достучаться до дома.", true
}
if len(ents) == 0 {
return "дом ничего не отдаёт.", true
}
var on []string
var sensors []string
for _, e := range ents {
switch {
case e.Domain == "sensor" || e.Domain == "binary_sensor":
if len(sensors) < 3 && e.State != "" && e.State != "unavailable" {
sensors = append(sensors, e.Name+" "+e.State+e.Unit)
}
case e.State == "on" || e.State == "open" || e.State == "unlocked":
on = append(on, e.Name)
}
}
var parts []string
if len(on) > 0 {
if len(on) > 5 {
on = on[:5]
}
parts = append(parts, "включено: "+strings.Join(on, ", "))
} else {
parts = append(parts, "всё выключено")
}
if len(sensors) > 0 {
parts = append(parts, strings.Join(sensors, ", "))
}
return strings.Join(parts, "; ") + ".", true
}
// isHomeQuery recognises a question about the house, narrowly. "дома" on its
// own is not enough — "я дома" is a fact, not a question — so it takes a house
// marker AND an ask AND either a device word or the word "включ…". Weather
// wording bails out first: "какая температура на улице?" belongs to the weather
// source, and both questions contain "температура".
func isHomeQuery(u string) bool {
s := strings.ToLower(strings.TrimSpace(u))
if s == "" {
return false
}
for _, w := range []string{"погод", "на улице", "прогноз"} {
if strings.Contains(s, w) {
return false
}
}
for _, phrase := range []string{"что включено", "что выключено", "умный дом", "что в доме включено"} {
if strings.Contains(s, phrase) {
return true
}
}
house := homeWord(s, "дома") || strings.Contains(s, "в доме") || strings.Contains(s, "в квартире")
if !house {
return false
}
ask := strings.Contains(s, "?") || homeWord(s, "что") || homeWord(s, "какая") ||
homeWord(s, "какой") || homeWord(s, "сколько")
if !ask {
return false
}
for _, w := range []string{"свет", "лампа", "лампы", "розетк", "датчик", "температур", "включ", "выключ"} {
if strings.Contains(s, w) {
return true
}
}
return false
}
// homeWord — whole-token membership, so "дома" does not fire on "домашний".
// Punctuation is trimmed off each token because a spoken question arrives with
// a question mark glued to the last word.
func homeWord(s, w string) bool {
for _, tok := range strings.Fields(s) {
if strings.Trim(tok, ".,!?;:") == w {
return true
}
}
return false
}
+190
View File
@@ -0,0 +1,190 @@
package main
import (
"context"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/config"
)
const haStatesFixture = `[
{"entity_id":"light.living_room","state":"on","attributes":{"friendly_name":"Гостиная"}},
{"entity_id":"switch.kettle","state":"off","attributes":{"friendly_name":"Чайник"}},
{"entity_id":"sensor.bedroom_temp","state":"22.5","attributes":{"friendly_name":"Спальня","unit_of_measurement":"°C"}}
]`
func TestWireSmartHomeOffUnlessEnabled(t *testing.T) {
st := newTestStore(t)
for name, cfg := range map[string]*config.Config{
"no block": {},
"written but dark": {SmartHome: &config.SmartHomeConfig{
URL: "http://ha.lan:8123", Token: "t",
}},
} {
t.Run(name, func(t *testing.T) {
if w := wireSmartHome(cfg, st); w != nil {
t.Fatal("the house must be off unless the block is enabled")
}
})
}
// nil wiring must be safe everywhere it is reachable.
var w *homeWiring
w.propose(context.Background())
w.run(context.Background())
if w.caller() != nil {
t.Fatal("a nil wiring must have no caller")
}
if _, ok := w.homeSummary(context.Background()); ok {
t.Fatal("a nil wiring must not claim a query")
}
}
// An unreachable instance must not stop the daemon and must propose nothing.
func TestWireSmartHomeUnreachableIsNotFatal(t *testing.T) {
st := newTestStore(t)
w := wireSmartHome(&config.Config{SmartHome: &config.SmartHomeConfig{
// Port 1 on loopback: nothing listens, and it fails fast.
URL: "http://127.0.0.1:1", Token: "t", Enabled: true,
}}, st)
if w == nil {
t.Fatal("a configured house should still wire")
}
tools, err := st.ListTools(context.Background(), "")
if err != nil {
t.Fatal(err)
}
if len(tools) != 0 {
t.Fatalf("an instance that never answered must propose nothing, got %+v", tools)
}
}
// Discovery proposes one row per controllable service, always destructive,
// always 'proposed'. A sensor gets no row: there is nothing to call on it.
func TestProposeOnlyProposesControllableDevices(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
_, _ = w.Write([]byte(haStatesFixture))
}))
defer srv.Close()
st := newTestStore(t)
w := wireSmartHome(&config.Config{SmartHome: &config.SmartHomeConfig{
URL: srv.URL, Token: "t", Enabled: true,
}}, st)
if w == nil {
t.Fatal("wireSmartHome returned nil for an enabled, reachable house")
}
tools, err := st.ListTools(context.Background(), "")
if err != nil {
t.Fatal(err)
}
got := map[string]bool{}
for _, tl := range tools {
got[tl.Name] = true
if tl.Status != "proposed" {
t.Errorf("%s status = %q: discovery must never enable", tl.Name, tl.Status)
}
if !tl.Destructive {
t.Errorf("%s is not destructive: every house control needs the confirm turn", tl.Name)
}
if len(tl.Cmd) == 0 || tl.Cmd[0] != "smarthome" {
t.Errorf("%s cmd = %v", tl.Name, tl.Cmd)
}
}
for _, want := range []string{
"home_light_living_room_on", "home_light_living_room_off",
"home_switch_kettle_on", "home_switch_kettle_off",
} {
if !got[want] {
t.Errorf("missing proposal %q (have %v)", want, got)
}
}
if len(tools) != 4 {
t.Fatalf("got %d rows, want 4 — the sensor must not be proposed: %+v", len(tools), tools)
}
// A second pass must be idempotent: re-discovery duplicates nothing and
// never rewrites a row Kami already enabled.
if err := st.EnableTool(context.Background(), "home_switch_kettle_on",
[]string{"smarthome", "switch.kettle", "turn_on"}, true, "smarthome:switch", time.Now()); err != nil {
t.Fatal(err)
}
w.propose(context.Background())
again, err := st.ListTools(context.Background(), "")
if err != nil {
t.Fatal(err)
}
if len(again) != 4 {
t.Fatalf("re-discovery duplicated rows: %d", len(again))
}
for _, tl := range again {
if tl.Name == "home_switch_kettle_on" && tl.Status != "enabled" {
t.Errorf("re-discovery un-enabled a device he had enabled: %q", tl.Status)
}
}
}
func TestHomeSummaryReadsState(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
_, _ = w.Write([]byte(haStatesFixture))
}))
defer srv.Close()
w := wireSmartHome(&config.Config{SmartHome: &config.SmartHomeConfig{
URL: srv.URL, Token: "t", Enabled: true,
}}, newTestStore(t))
out, ok := w.homeSummary(context.Background())
if !ok {
t.Fatal("summary did not claim the turn")
}
if !strings.Contains(out, "Гостиная") {
t.Errorf("the lamp that is on should be named: %q", out)
}
if strings.Contains(out, "Чайник") {
t.Errorf("a device that is off should not be listed as on: %q", out)
}
if !strings.Contains(out, "22.5") {
t.Errorf("the sensor reading should be there: %q", out)
}
// Persona: no masculine self-reference, no "вы", no pet names.
for _, bad := range []string{"рад ", "готов ", "вы ", "ваш", "милый", "дорогой"} {
if strings.Contains(strings.ToLower(out), bad) {
t.Errorf("persona violation %q in %q", bad, out)
}
}
}
func TestIsHomeQuery(t *testing.T) {
yes := []string{
"что включено дома?",
"что выключено",
"какой свет горит дома",
"свет в доме включен?",
"какая температура в квартире?",
"покажи умный дом",
}
no := []string{
"",
"я дома",
"буду дома в семь",
"какая погода дома", // weather wording wins
"какая температура на улице?",
"домашние дела", // "дома" must not fire on "домашние"
"что мне нужно сделать?",
"напомни выключить чайник в семь", // a reminder, not a house read
}
for _, u := range yes {
if !isHomeQuery(u) {
t.Errorf("isHomeQuery(%q) = false, want true", u)
}
}
for _, u := range no {
if isHomeQuery(u) {
t.Errorf("isHomeQuery(%q) = true, want false", u)
}
}
}
+146
View File
@@ -0,0 +1,146 @@
// mavend/speaker.go — core's half of voice identification (Vikunja #255,
// docs/plans/10-speaker-recognition.md).
//
// # What is actually wired here, and what is not
//
// The enrolment plumbing is real: profiles are stored, listed and deleted, and
// the wire methods exist as soon as a speaker block is configured. The
// recognising half is NOT, and cannot be on this box, because there is no
// speaker-embedding model on disk — no ECAPA, no x-vector, no titanet, no
// wespeaker, nothing in /mnt/hdd1/llms but text ggufs. Until one is downloaded,
// newSpeakerEmbedder returns nil, internal/speaker falls back to
// speaker.Disabled, and every Identify answers ErrDisabled. The daemon logs
// which half is off at startup rather than pretending.
//
// This is deliberately not papered over with a hand-rolled MFCC floor. A
// biometric that is confidently wrong writes false claims about named people
// into his memory, and that is worse than a capability that is honestly absent.
//
// # Off unless configured
//
// No speaker block, or one without enabled, ⇒ the three methods do not exist and
// answer ErrUnknownMethod. On an unconfigured box there is no wire path that
// takes a voiceprint at all.
//
// # The refused design step
//
// The plan asks for unknown speakers to be enrolled on first interaction. That
// is refused in internal/speaker/enroll.go and there is no handler for it here:
// no request shape in the protocol enrols whoever just spoke. Taking a biometric
// of a guest who walked past the microphone is not something this daemon does.
package main
import (
"context"
"errors"
"log"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/speaker"
"github.com/kami/maven/internal/store"
)
// speakerWiring holds the recognizer behind the three IPC handlers.
type speakerWiring struct {
rec *speaker.Recognizer
}
// newSpeakerEmbedder loads the speaker-embedding model named by the config.
//
// It always returns nil today. The seam exists so that wiring a real model is a
// change to this one function and nothing else: give it a loader, and Identify
// starts working with no change to the store, the protocol, the auth table or
// the handlers. See the plan document for what to download.
func newSpeakerEmbedder(cfg *config.SpeakerConfig) speaker.Embedder {
if cfg == nil || cfg.ModelPath == "" {
return nil
}
log.Printf("speaker: model_path %q is configured but no embedding backend is built yet; "+
"enrolment and deletion work, recognition does not (Vikunja #255)", cfg.ModelPath)
return nil
}
// newSpeakerWiring builds the recognizer, or nil when the capability is off.
func newSpeakerWiring(st *store.Store, cfg *config.Config) *speakerWiring {
if cfg == nil || cfg.Speaker == nil || !cfg.Speaker.Enabled {
return nil
}
if st == nil {
log.Print("speaker: enabled but there is no store to keep profiles in; staying off")
return nil
}
rec, err := speaker.New(newSpeakerEmbedder(cfg.Speaker), st.VectorMemory(), speaker.Config{
Threshold: cfg.Speaker.Threshold,
MinSeconds: cfg.Speaker.MinSeconds,
})
if err != nil {
log.Printf("speaker: %v; staying off", err)
return nil
}
if rec.Enabled() {
log.Printf("speaker: recognition on, threshold %.2f", rec.Threshold())
} else {
log.Print("speaker: enrolment on, recognition BLOCKED — no speaker-embedding model " +
"on this box (see docs/plans/10-speaker-recognition.md)")
}
return &speakerWiring{rec: rec}
}
func (w *speakerWiring) enroll(ctx context.Context, req ipc.EnrollSpeakerReq) (ipc.EnrollSpeakerResp, error) {
p, err := w.rec.Enroll(ctx, req.ID, req.Name, req.Samples)
if err != nil {
return ipc.EnrollSpeakerResp{}, speakerErr(err)
}
return ipc.EnrollSpeakerResp{Speaker: toWireSpeaker(p)}, nil
}
func (w *speakerWiring) list(ctx context.Context) (ipc.ListSpeakersResp, error) {
ps, err := w.rec.List(ctx)
if err != nil {
return ipc.ListSpeakersResp{}, speakerErr(err)
}
out := make([]ipc.Speaker, 0, len(ps))
for _, p := range ps {
out = append(out, toWireSpeaker(p))
}
return ipc.ListSpeakersResp{Speakers: out, Enabled: w.rec.Enabled()}, nil
}
func (w *speakerWiring) forget(ctx context.Context, req ipc.ForgetSpeakerReq) error {
return speakerErr(w.rec.Forget(ctx, req.ID))
}
// toWireSpeaker drops the voiceprint. A listing says who is enrolled; it does
// not hand the biometric back out over the socket.
func toWireSpeaker(p speaker.Profile) ipc.Speaker {
return ipc.Speaker{ID: p.ID, Name: p.Name, Enrolled: p.Enrolled, Samples: p.Samples}
}
// speakerErr maps the package sentinels onto the wire vocabulary so a surface
// can tell "you asked wrong" from "core broke".
func speakerErr(err error) error {
switch {
case err == nil:
return nil
case errors.Is(err, speaker.ErrNotFound):
return ipc.ErrNoFact
case errors.Is(err, speaker.ErrBadID),
errors.Is(err, speaker.ErrBadFormat),
errors.Is(err, speaker.ErrTooShort):
return errors.Join(ipc.ErrBadParams, err)
default:
return err
}
}
// wireSpeaker attaches the three handlers when the capability is configured.
func wireSpeaker(srv *ipc.Server, st *store.Store, cfg *config.Config) {
w := newSpeakerWiring(st, cfg)
if w == nil {
return
}
srv.EnrollSpeakerFn = w.enroll
srv.ListSpeakersFn = w.list
srv.ForgetSpeakerFn = w.forget
}
+85
View File
@@ -0,0 +1,85 @@
// Package main — strutil.go holds small, receiver-free string utilities used
// across the voice reply paths: trimming a wake token, pulling out the first
// word or first line, and a minimal JSON string encoder for the one payload
// shape that needs it. Extend this file rather than voice.go for anything in
// that shape.
package main
import (
"fmt"
"strings"
"github.com/kami/maven/internal/router"
)
// stripWake removes a leading wake token (any script the STT phonetically
// transcribes "Maven" as) so the verb is the first word.
func stripWake(u string) string {
stripped, had := router.StripWakeToken(u)
if !had {
return strings.TrimSpace(u)
}
return stripped
}
// firstWord returns the first whitespace-delimited token (lowercased) — the
// proposed tool's name.
func firstWord(s string) string {
f := strings.Fields(s)
if len(f) == 0 {
return ""
}
return strings.ToLower(f[0])
}
// firstLine — the first non-empty line of a tool's output, for a short spoken
// reply (the full output goes to the log, not the TTS). Trimmed to keep the
// utterance sane if a command dumps a wall of text.
func firstLine(s string) string {
for _, line := range strings.Split(s, "\n") {
line = strings.TrimSpace(line)
if line != "" {
if len(line) > 200 {
line = line[:200]
}
return line
}
}
return ""
}
// jsonString — a one-line JSON string encoder without dragging encoding/json
// into the top of this file. Used to wrap a reminder payload's text field;
// the router's reminder Slots are already absolute (DateTimeParser resolved
// relative→absolute), the payload shape is conventional {"text":...}.
func jsonString(s string) string {
// minimal JSON string escape — quotes + backslash + control chars.
// adequate for the reminder payload's text field; not a general JSON
// encoder. The chroma / RAG modules (when they land) use a real json
// encoder for richer payloads. Keep it inline here so the import
// direction stays narrow.
var b []byte
b = append(b, '"')
for _, r := range s {
switch r {
case '"':
b = append(b, '\\', '"')
case '\\':
b = append(b, '\\', '\\')
case '\n':
b = append(b, '\\', 'n')
case '\r':
b = append(b, '\\', 'r')
case '\t':
b = append(b, '\\', 't')
default:
if r < 0x20 {
b = append(b, []byte(fmt.Sprintf("\\u%04x", r))...)
} else {
b = append(b, []byte(string(r))...)
}
}
}
b = append(b, '"')
return string(b)
}
+78
View File
@@ -0,0 +1,78 @@
package main
import (
"context"
"testing"
"time"
"github.com/kami/maven/internal/router"
)
// systemHandler — a handler with nothing but a fixed clock, which is all
// replySystem needs.
func systemHandler(now time.Time) *reactiveHandler {
return &reactiveHandler{now: func() time.Time { return now }}
}
// TestReplySystemDateOffset — "какое число завтра" must answer tomorrow's
// date, not today's (Vikunja #388).
func TestReplySystemDateOffset(t *testing.T) {
// Thursday, 30 July 2026.
now := time.Date(2026, 7, 30, 14, 5, 0, 0, time.UTC)
h := systemHandler(now)
cases := []struct{ utterance, want string }{
{"какое сегодня число", "сегодня четверг, 30 июля 2026 года"},
{"какое число", "сегодня четверг, 30 июля 2026 года"},
{"какое число завтра", "завтра пятница, 31 июля 2026 года"},
{"какое число послезавтра", "послезавтра суббота, 1 августа 2026 года"},
{"какое было число вчера", "вчера среда, 29 июля 2026 года"},
}
for _, c := range cases {
got := h.replySystem(context.Background(), router.Decision{Utterance: c.utterance})
if got != c.want {
t.Errorf("replySystem(%q) = %q, want %q", c.utterance, got, c.want)
}
}
}
// A day she cannot work out must not come back as today's date — that is the
// same silent wrong answer #388 was about, one step further out.
func TestReplySystemUnknownDayIsHonest(t *testing.T) {
now := time.Date(2026, 7, 30, 14, 5, 0, 0, time.UTC)
h := systemHandler(now)
for _, u := range []string{
"какое число в пятницу",
"какое число через неделю",
"какое число в понедельник",
} {
got := h.replySystem(context.Background(), router.Decision{Utterance: u})
if got != onlyNearDaysReply {
t.Errorf("replySystem(%q) = %q, want the honest reply", u, got)
}
}
// The days she does know must not be caught by the same guard.
if got := h.replySystem(context.Background(), router.Decision{Utterance: "какое число завтра"}); got == onlyNearDaysReply {
t.Error("завтра was treated as an unknown day")
}
}
// TestReplySystemClockCity — the clock arm must not answer local time for a
// question about another city (Vikunja #388). She keeps one clock, so every
// named place gets the honest "local time only" answer.
func TestReplySystemClockCity(t *testing.T) {
now := time.Date(2026, 7, 30, 12, 0, 0, 0, time.UTC)
h := systemHandler(now)
cases := []struct{ utterance, want string }{
{"который час", "сейчас 12 часов ровно"},
{"который час в киеве", onlyLocalTimeReply},
{"сколько времени в москве", onlyLocalTimeReply},
{"который час в лондоне", onlyLocalTimeReply},
{"который час в бишкеке", onlyLocalTimeReply},
}
for _, c := range cases {
got := h.replySystem(context.Background(), router.Decision{Utterance: c.utterance})
if got != c.want {
t.Errorf("replySystem(%q) = %q, want %q", c.utterance, got, c.want)
}
}
}
+67
View File
@@ -0,0 +1,67 @@
{
"schema_version": 1,
"name": "evening_degraded",
"description": "The tier-2 pipeline case #288 deferred here, plus degraded mode. A golden WAV goes in at the microphone end and comes out as a written fact, and then the ecosystem starts answering 503 and the proactive loop has to stay quiet instead of falling over. The audio step asserts the PIPELINE — mic to STT seam to router to store to TTS — not whisper's accuracy; cmd/mavsttd/golden_test.go owns accuracy.",
"start": "2026-08-01T21:00:00+03:00",
"praxis_attention": "[{\"id\":\"item_1\",\"title\":\"medicine not taken\",\"importance\":3.0,\"rule\":\"evening_medicine\"}]",
"script": [
{
"match": "выпил воды",
"route": "[{\"intent\":\"fact\",\"key\":\"water\",\"value\":\"выпил\"}]"
},
{
"match": "записала факт: water",
"reply": "{\"response\":\"Записала, что ты выпил воды.\",\"mood\":\"neutral\"}"
},
{
"match": "",
"route": "[{\"intent\":\"chat\",\"text\":\"привет\"}]",
"reply": "{\"response\":\"Я рада тебя слышать.\",\"mood\":\"happy\"}"
}
],
"steps": [
{
"at": "21:00",
"note": "he speaks. The whole voice path runs: push-to-talk, the STT seam parked with the golden transcript, the real router, the real store write, the phrasing contract.",
"audio": "ru_fact",
"expect_reply_contains": ["записала"],
"expect_reply_lacks": ["записал,", "милый", "ваш"],
"expect_events": ["water"]
},
{
"at": "21:05",
"note": "a healthy tick with him just having spoken stays silent",
"tick": true,
"expect_no_send": true
},
{
"at": "21:10",
"note": "the ecosystem goes down",
"fault": 503
},
{
"at": "21:15",
"note": "a tick against a dead ecosystem must degrade, not send half a thought",
"tick": true,
"expect_no_send": true,
"expect_no_events": true
},
{
"at": "21:20",
"note": "intake keeps working while the ecosystem is down — a write does not depend on it",
"arrive": {
"source": "rss:tech",
"note": { "text": "Патч 6.19.1 [tech]\nисправления\nhttps://example.org/b" }
},
"expect_events": ["rss:tech"],
"expect_no_send": true
},
{
"at": "21:25",
"note": "recovery",
"clear_fault": true,
"tick": true,
"expect_no_send": true
}
]
}
+96
View File
@@ -0,0 +1,96 @@
{
"schema_version": 1,
"name": "morning_missed",
"description": "The scenario from Vikunja #284's description, replayed. He appears at 08:30, things arrive through the morning while he is at the desk, and at 08:50 he asks what he missed. The assertions are as much about what did NOT happen — nothing was sent at him unprompted — as about what she said.",
"start": "2026-08-01T08:30:00+03:00",
"praxis_attention": "[{\"id\":\"item_1\",\"title\":\"medicine not taken\",\"importance\":3.0,\"rule\":\"morning_medicine\"}]",
"script": [
{
"match": "выпил воды",
"route": "[{\"intent\":\"fact\",\"key\":\"water\",\"value\":\"выпил\"}]"
},
{
"match": "записала факт: water",
"reply": "{\"response\":\"Записала, что ты выпил воды.\",\"mood\":\"neutral\"}"
},
{
"match": "что я пропустил",
"route": "[{\"intent\":\"query\",\"text\":\"что я пропустил\"}]"
},
{
"match": "",
"route": "[{\"intent\":\"chat\",\"text\":\"привет\"}]",
"reply": "{\"response\":\"Я рада тебя слышать.\",\"mood\":\"happy\"}"
}
],
"steps": [
{
"at": "08:30",
"note": "he appears at the desk",
"signal": { "key": "desk_active", "value": "true", "source": "infer:hyprland" },
"expect_events": ["infer:hyprland"],
"expect_no_send": true
},
{
"at": "08:32",
"note": "a feed item arrives, published half an hour ago",
"arrive": {
"source": "rss:tech",
"as_of": "08:02",
"note": { "text": "Вышло ядро 6.19 [tech]\nкраткое содержание\nhttps://example.org/a" }
},
"expect_events": ["rss:tech"],
"expect_no_send": true
},
{
"at": "08:35",
"note": "the mail reader extracts a candidate — a candidate is never spoken",
"arrive": {
"source": "email:inbox",
"task": { "text": "продлить домен", "evidence": "Домен истекает через 7 дней" }
},
"expect_events": ["email:inbox", "продлить домен"],
"expect_no_send": true
},
{
"at": "08:40",
"note": "the work calendar signal — a relayed notification, below full confidence",
"arrive": {
"source": "ambient:notif",
"fact": {
"key": "calendar_event_20260801_планёрка",
"value": "10:00-11:00 планёрка"
}
},
"expect_events": ["ambient:notif", "планёрка"],
"expect_no_send": true
},
{
"at": "08:45",
"note": "a tick with him present and nothing wrong must stay silent",
"tick": true,
"expect_no_send": true
},
{
"at": "08:50",
"note": "he asks. The query path answers from local recall only: nothing stored clears the score gate, so she refuses rather than inventing a morning summary, and the replier is never reached. That refusal is the no-hallucination floor and this step pins it.",
"say": "что я пропустил?",
"expect_reply_contains": ["не знаю"],
"expect_reply_lacks": ["рад ", "милый", "ваш"]
},
{
"at": "08:55",
"note": "stating a fact writes it and says so, in the feminine",
"say": "я выпил воды",
"expect_reply_contains": ["записала"],
"expect_reply_lacks": ["записал,", "милый"],
"expect_events": ["water"]
},
{
"at": "09:00",
"note": "a second tick, still nothing unprompted",
"tick": true,
"expect_no_send": true
}
]
}
+423
View File
@@ -24,6 +24,7 @@ import (
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/morning"
"github.com/kami/maven/internal/pattern"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/routine"
"github.com/kami/maven/internal/store"
@@ -67,6 +68,14 @@ type tickLoop struct {
morningRoutines []morning.Routine
morningLast map[string]time.Time
// proposalCfg — announcement policy for routines the tick inferred itself.
// nil ⇒ detect silently, never announce (the default). lastProposalAt is
// the cooldown clock, in-memory on purpose: a restart is allowed to permit
// one more announcement, and a restart-per-day loop is a bigger problem
// than a duplicate proposal notice.
proposalCfg *config.PatternProposalConfig
lastProposalAt time.Time
// digestQ — in-memory queue of eligible nudges waiting for batch flush.
// populated when digestCfg != nil && digestCfg.Enabled.
digestQ []QueuedNudge
@@ -92,6 +101,7 @@ func newTickLoop(
digestCfg *config.DigestConfig,
routines []routine.Routine,
morningRoutines []morning.Routine,
proposalCfg *config.PatternProposalConfig,
) *tickLoop {
return &tickLoop{
store: st,
@@ -108,6 +118,7 @@ func newTickLoop(
routineLast: make(map[string]time.Time),
morningRoutines: morningRoutines,
morningLast: make(map[string]time.Time),
proposalCfg: proposalCfg,
lastPhrase: make(map[string]delivery.PhrasedNudge),
}
}
@@ -180,17 +191,40 @@ func (t *tickLoop) tick(ctx context.Context, now time.Time) {
// (with dedup) avoids re-queueing the same rule after a flush.
t.maybeFlush(ctx, now, state)
// gate-suppressed digest (Vikunja #281): rules the restraint gate held
// back this tick (quiet hours / away / calendar-busy), not because they
// weren't due, but because it wasn't the moment. Some of those are worth
// resurfacing later instead of just being lost — loop.DigestEligible
// draws that line. This is a SEPARATE mechanism from the in-memory
// digestQ above: that one batches candidates the gate already ALLOWED to
// fire; this one durably holds candidates the gate BLOCKED.
t.enqueueSuppressedDigest(ctx, trace, state, now)
t.expireStaleDigest(ctx, now)
t.maybeDrainDigest(ctx, state, now)
// routines: operator-declared scheduled behaviors. fire the ones whose cron
// crossed since last fire, delivered through the normal routing (voice when
// present, away channels otherwise). bodies are literal operator text — not
// LLM-phrased — so a routine can't hallucinate. severity comes from config.
t.fireRoutines(ctx, now, state)
// accepted routines: patterns the user confirmed. read straight from the
// store each tick so the schedule survives a restart.
t.fireAcceptedRoutines(ctx, now, state)
// morning routines: daily checklists (medicine/water/pets/...), nagged at
// most once per day per routine, and only for items still unevidenced at
// nudge time. See internal/morning for the "why not four timers" rationale.
t.fireMorningRoutines(ctx, now, state)
// pattern detection: scan every action+object pair with recorded events
// and propose a routine for any stable one not already decided (Vikunja
// #43). This used to only run as a side effect of the voice fact-write
// path, so a pattern already sitting in history went unnoticed until he
// happened to mention it again by voice. See patterns.go and
// detectPatterns below for how idempotence and dismissal are respected.
t.detectPatterns(ctx, now, state)
// reminders: gate-bypassing class. fired once, marked after a successful
// delivery. a failed send leaves the reminder pending — the next tick
// re-gathers and re-attempts.
@@ -338,6 +372,235 @@ func (t *tickLoop) flushDigest(ctx context.Context, now time.Time, state loop.St
t.digestQ = nil
}
// detectPatterns runs the pattern detector proactively over every
// action+object pair that has ever produced an event, independent of
// whichever fact write (or channel) last touched it (Vikunja #43). This is
// what makes pattern inference actually proactive: it fires on the daemon's
// own schedule reading accumulated history, not only as a side effect of a
// live voice turn.
//
// Idempotence and noise are handled by the store, not here — this function
// is safe to call every tick:
// - Same pattern, tick after tick: detectAndPropose's LookupProposedRoutine
// check plus proposed_routines' UNIQUE(action, object) constraint (with
// CreateProposedRoutine's ON CONFLICT DO NOTHING) mean a pair that
// already has a row — in ANY status — produces no second row and no log
// spam beyond the one line at genuine creation.
// - A DISMISSED proposal must never come back. DismissProposedRoutine flips
// status in place; the row is never deleted. So the same Lookup check
// that stops a duplicate "proposed" also stops a "dismissed" one from
// resurrecting — there is nothing tick-specific to get right here beyond
// calling the same shared path the voice route already used.
//
// By default this only creates a row for the /routines page to show: it does
// not notify, ring, or speak. Detection is not the same act as disturbing him
// about it, and Maven is "not a nag, not autonomous" (CLAUDE.md). Announcing
// is opt-in through the pattern_proposals config block — see announceProposal
// for the restraints that apply even then. A proposal only starts producing
// recurring nudges once he accepts it (fireAcceptedRoutines).
func (t *tickLoop) detectPatterns(ctx context.Context, now time.Time, state loop.State) {
pairs, err := t.store.DistinctEventPairs(ctx)
if err != nil {
log.Printf("tick: distinct event pairs: %v", err)
return
}
announced := false
for _, p := range pairs {
r, _, err := detectAndPropose(ctx, t.store, p.Action, p.Object, now)
if err != nil {
log.Printf("tick: detect pattern %s/%s: %v", p.Action, p.Object, err)
continue
}
if r == nil {
continue // no stable pattern, or already proposed/accepted/dismissed
}
log.Printf("tick: proposed routine: %s/%s every %.1f days", r.Action, r.Object, r.IntervalDays)
// One announcement per tick at most, whatever the scan turned up. The
// rest are on /routines; they are not lost, they are just not shouted.
if announced {
continue
}
announced = t.announceProposal(ctx, r, now, state)
}
}
// announceProposal offers a freshly inferred routine through the ordinary
// care-delivery path, if announcing is switched on at all. Returns true when
// something was actually sent.
//
// Everything here is restraint. The feature is off unless configured; when on
// it is sev1 (the lowest severity, so quiet hours, away presence and snooze
// all suppress it via loop.Gate exactly like a care nudge); it is spaced by
// proposalCfg.Cooldown across every pair, not per pair; and a suppressed or
// dropped announcement is NOT retried — the cooldown clock advances only on a
// real send, but the proposal row already exists, so the next tick will not
// re-detect it and nothing queues up behind it. A missed announcement means
// he reads it on /routines instead, which is the whole point of the page.
//
// The body is the detector's own literal Russian phrasing (pattern.PhraseRoutine
// — "ты заправляешь поилку раз в 7 дней — напоминать?"), not LLM-generated, so
// an inferred routine cannot arrive worded as something Maven never observed.
func (t *tickLoop) announceProposal(ctx context.Context, r *pattern.ProposedRoutine, now time.Time, state loop.State) bool {
if !t.proposalCfg.AnnounceProposals() {
return false
}
cooldown := time.Duration(t.proposalCfg.Cooldown)
if cooldown <= 0 {
cooldown = config.DefaultProposalCooldown
}
if !t.lastProposalAt.IsZero() && now.Sub(t.lastProposalAt) < cooldown {
return false
}
rule := loop.Rule{Name: "proposal:" + r.Action + " " + r.Object, Severity: loop.Sev1}
if !loop.Gate(state, rule) {
return false
}
body := pattern.PhraseRoutine(r)
pn := delivery.PhrasedNudge{
Candidate: loop.Candidate{Rule: rule, Severity: rule.Severity, State: state},
Body: body,
Summary: body,
}
sent, err := t.dispatcher.DispatchNudge(ctx, pn, now)
if err != nil {
log.Printf("tick: announce proposal %s/%s: %v", r.Action, r.Object, err)
return false
}
if len(sent) == 0 {
return false // routing dropped it — /routines still has it.
}
t.lastProposalAt = now
return true
}
// digestExpiry — how long a gate-suppressed care nudge stays worth
// resurfacing. 24h: these are daily-cadence rules (water/meal/break run on
// hour-scale cooldowns and re-derive from facts that reset every day), so a
// digest entry that outlives one full day is describing a day that's already
// over — "you skipped a break yesterday" said tomorrow evening is noise, not
// news. Bounding at one day also means a digest can never silently span a
// weekend of quiet hours into an unbounded backlog.
const digestExpiry = 24 * time.Hour
// maxDigestSpokenItems — the bundle read-out is capped so "batched, not
// dropped" cannot regress into "she dumps twelve things on me the moment I
// walk in" — a digest that nags in bulk is worse than the drops it replaced.
// Anything beyond the cap is still marked drained (it did get its moment;
// the cap limits WORDS, not whether it counted) and folded into a trailing
// count instead of being spoken in full.
const maxDigestSpokenItems = 3
// enqueueSuppressedDigest scans this tick's trace for care candidates the
// gate blocked for a genuine restraint reason and durably records the
// digest-eligible ones (loop.DigestEligible). Phrasing happens once, here,
// at enqueue time — not re-derived at drain time — the same way queueNudge
// phrases once and caches, so a rule suppressed for hours isn't re-prompting
// the LLM every tick it stays blocked (EnqueueDigestEntry's rule+body dedupe
// makes repeat calls here harmless, but skipping the phrase call entirely
// when a pending entry already exists avoids the LLM round-trip too).
func (t *tickLoop) enqueueSuppressedDigest(ctx context.Context, trace *loop.TickTrace, state loop.State, now time.Time) {
if trace == nil {
return
}
for _, tr := range trace.RuleTraces {
if !tr.PredicateResult || tr.GateResult {
continue // didn't want to fire, or wasn't suppressed
}
if !loop.DigestEligible(tr.Severity, tr.GateBlockedBy) {
continue
}
rule := loop.Rule{Name: tr.RuleName, Severity: tr.Severity}
cand := loop.Candidate{Rule: rule, Severity: tr.Severity, State: state}
pn, err := t.phraser.PhraseNudge(ctx, cand)
if err != nil {
log.Printf("tick: phrase digest candidate %s: %v", tr.RuleName, err)
continue
}
expires := now.Add(digestExpiry)
if _, deduped, err := t.store.EnqueueDigestEntry(ctx, tr.RuleName, int(tr.Severity), pn.Body, now, expires); err != nil {
log.Printf("tick: enqueue digest entry %s: %v", tr.RuleName, err)
} else if deduped {
// same suppressed nudge already pending — nothing new to say.
continue
}
}
}
// expireStaleDigest sweeps entries past their expiry once per tick — cheap
// bookkeeping, mirrors ReconcileStaleDeliveryAttempts's shape.
func (t *tickLoop) expireStaleDigest(ctx context.Context, now time.Time) {
n, err := t.store.ExpireStaleDigestEntries(ctx, now)
if err != nil {
log.Printf("tick: expire stale digest entries: %v", err)
return
}
if n > 0 {
log.Printf("tick: expired %d stale digest entr(y/ies) unspoken", n)
}
}
// maybeDrainDigest speaks the pending digest bundle once the gate's
// suppression reasons have actually cleared — quiet hours over, back from
// away, out of the meeting. Draining while still suppressed would just be a
// second way to nag through quiet hours; the bundle waits for the same "is
// it allowed right now" condition a live nudge already waits for.
func (t *tickLoop) maybeDrainDigest(ctx context.Context, state loop.State, now time.Time) {
if state.QuietHours || state.CalendarBusy || state.Presence == store.Away {
return
}
entries, err := t.store.PendingDigestEntries(ctx, now)
if err != nil {
log.Printf("tick: pending digest entries: %v", err)
return
}
if len(entries) == 0 {
return
}
spoken := entries
extra := 0
if len(spoken) > maxDigestSpokenItems {
spoken = entries[:maxDigestSpokenItems]
extra = len(entries) - maxDigestSpokenItems
}
var b strings.Builder
maxSev := 0
for i, e := range spoken {
if i > 0 {
b.WriteString(" · ")
}
b.WriteString(e.Body)
if e.Severity > maxSev {
maxSev = e.Severity
}
}
if extra > 0 {
fmt.Fprintf(&b, " · и ещё %d", extra)
}
body := b.String()
summary := fmt.Sprintf("%d отложенных уведомлений", len(entries))
cand := loop.Candidate{
Rule: loop.Rule{Name: "digest", Severity: loop.Severity(maxSev)},
Severity: loop.Severity(maxSev),
State: state,
}
pn := delivery.PhrasedNudge{Candidate: cand, Body: body, Summary: summary}
t.cachePhrase(pn)
if _, err := t.dispatcher.DispatchNudge(ctx, pn, now); err != nil {
log.Printf("tick: dispatch digest bundle: %v", err)
return // leave entries pending; retried next tick
}
ids := make([]int64, len(entries))
for i, e := range entries {
ids[i] = e.ID
}
if err := t.store.DrainDigestEntries(ctx, ids, now); err != nil {
log.Printf("tick: drain digest entries: %v", err)
}
}
// routinesFromConfig maps the config's routine blocks to the engine type.
// Validation (cron parses, name/body present, severity defaulted) already ran
// in config.Load, so this is a pure field copy.
@@ -377,6 +640,64 @@ func (t *tickLoop) fireRoutines(ctx context.Context, now time.Time, state loop.S
}
}
// fireAcceptedRoutines nudges about the routines the user accepted, once per
// interval (Vikunja #366). Accepting used to create a single reminder, so a
// non-weekly routine fired once and went quiet forever; the schedule lives in
// the proposed_routines row now and the loop re-reads it every tick.
//
// A routine is a care-class nudge and goes through the restraint gate like any
// other: quiet hours, away presence and snooze all suppress it. Reminders bypass
// that gate; routines must not. A suppressed nudge is NOT marked fired, so it
// goes out on the next tick that the gate allows — one nudge, held, not dropped
// and not repeated.
//
// The body is literal text built from the detected action and object, not
// LLM-phrased, so a routine can't hallucinate. It nudges; it never acts.
func (t *tickLoop) fireAcceptedRoutines(ctx context.Context, now time.Time, state loop.State) {
rows, err := t.store.ListAcceptedRoutines(ctx)
if err != nil {
log.Printf("tick: list accepted routines: %v", err)
return
}
accepted := make([]routine.Accepted, 0, len(rows))
for _, r := range rows {
if r.AcceptedTs == nil {
continue // accepted before the schedule column existed — no clock to start from.
}
accepted = append(accepted, routine.Accepted{
ID: r.ID,
Name: r.Action + " " + r.Object,
IntervalDays: r.IntervalDays,
Accepted: *r.AcceptedTs,
LastFired: r.LastFiredTs,
})
}
for _, a := range routine.DueAccepted(accepted, now) {
rule := loop.Rule{Name: "routine:" + a.Name, Severity: loop.Sev1}
if !loop.Gate(state, rule) {
continue
}
body := "пора: " + a.Name
pn := delivery.PhrasedNudge{
Candidate: loop.Candidate{Rule: rule, Severity: rule.Severity, State: state},
Body: body,
Summary: body,
}
sent, err := t.dispatcher.DispatchNudge(ctx, pn, now)
if err != nil {
log.Printf("tick: dispatch accepted routine %d: %v", a.ID, err)
continue
}
if len(sent) == 0 {
continue // routing dropped it — leave it due.
}
if err := t.store.MarkRoutineFired(ctx, a.ID, now); err != nil {
log.Printf("tick: mark routine %d fired: %v", a.ID, err)
}
}
}
// fireMorningRoutines checks each configured checklist against today's facts
// and dispatches a nag listing exactly what's still missing, at most once per
// routine per calendar day. Fact reads happen here (not in loop.Gatherer)
@@ -468,6 +789,77 @@ func (t *tickLoop) morningStatus(ctx context.Context, now time.Time) []ipc.Morni
return out
}
// dayPlan is the read-only "what does today hold" query (Vikunja #128). It is
// the impure half of morning.BuildPlan: it reads the calendar events, the
// pending reminders and the checklist facts, and the pure builder orders them.
//
// It never dispatches. Asking for the plan is a query like any other; the only
// unprompted delivery in maven stays with the morning nudge and the
// dispatcher's policy.
func (t *tickLoop) dayPlan(ctx context.Context, now time.Time) ipc.DayPlan {
y, m, d := now.Date()
dayStart := time.Date(y, m, d, 0, 0, 0, 0, now.Location())
dayEnd := dayStart.AddDate(0, 0, 1)
var events []morning.PlanEntry
facts, err := t.store.CalendarEvents(ctx, dayStart, dayEnd)
if err != nil {
log.Printf("tick: day plan: calendar events: %v", err)
}
for _, f := range facts {
events = append(events, morning.PlanEntry{
At: f.Ts,
Text: f.Value,
Kind: morning.PlanEvent,
// Provenance below a calendar read (an ambient relay, #126) is
// hedged rather than recited as fact.
Uncertain: f.Confidence < 1.0,
})
}
var reminders []morning.PlanEntry
rems, err := t.store.ListReminders(ctx, dayPlanMaxReminders)
if err != nil {
log.Printf("tick: day plan: list reminders: %v", err)
}
for _, r := range rems {
if r.Status != "pending" {
continue
}
fire := r.NextFireTs
if fire.IsZero() {
fire = r.FireTs
}
reminders = append(reminders, morning.PlanEntry{
At: fire,
Text: strings.TrimSpace(r.Payload),
Kind: morning.PlanReminder,
})
}
var checklistFacts map[string]store.Fact
if len(t.morningRoutines) > 0 {
checklistFacts = t.gatherMorningFacts(ctx)
}
plan := morning.BuildPlan(t.morningRoutines, checklistFacts, events, reminders, now)
out := ipc.DayPlan{Date: plan.Date, Spoken: plan.FormatRU()}
out.Items = make([]ipc.DayPlanItem, len(plan.Items))
for i, it := range plan.Items {
out.Items[i] = ipc.DayPlanItem{
At: it.At,
Text: it.Text,
Kind: string(it.Kind),
Uncertain: it.Uncertain,
}
}
return out
}
// dayPlanMaxReminders bounds the reminder scan. The plan covers one day; a
// pending queue longer than this is a bug elsewhere, not a plan to recite.
const dayPlanMaxReminders = 500
// tune — the feedback auto-tuner's impure step. runs on a slow cadence
// (autotuneInterval, see run) so it doesn't write a fact every tick. for each
// rule:
@@ -547,7 +939,21 @@ type daemonAPI struct {
ipc.CoreAPI
getTrace func() *loop.TickTrace
getMorningStatus func(ctx context.Context) []ipc.MorningRoutineStatus
getDayPlan func(ctx context.Context) ipc.DayPlan
chatFn func(ctx context.Context, text string) string
getMCPServers func() []ipc.MCPServerStatus
getEvents func(n int) []ipc.IntakeEvent
}
// RecentEvents — the unified intake journal (Vikunja #283). Empty, not an
// error, when no bus was wired: "nothing has arrived" and "the journal is off"
// look the same to a reader on purpose, because neither is a fault and the
// page renders both as an empty table.
func (d *daemonAPI) RecentEvents(ctx context.Context, n int) ([]ipc.IntakeEvent, error) {
if d.getEvents == nil {
return nil, nil
}
return d.getEvents(n), nil
}
func (d *daemonAPI) Chat(ctx context.Context, text string) (string, error) {
@@ -557,6 +963,16 @@ func (d *daemonAPI) Chat(ctx context.Context, text string) (string, error) {
return d.chatFn(ctx, text), nil
}
// MCPServers — the configured MCP servers and their health (Vikunja #251).
// Empty, not an error, when the mcp block is absent: "not configured" is the
// default state and the web surface renders it as such.
func (d *daemonAPI) MCPServers(ctx context.Context) ([]ipc.MCPServerStatus, error) {
if d.getMCPServers == nil {
return nil, nil
}
return d.getMCPServers(), nil
}
func (d *daemonAPI) TickTrace(ctx context.Context) (ipc.TickTrace, error) {
trace := d.getTrace()
if trace == nil {
@@ -572,6 +988,13 @@ func (d *daemonAPI) MorningStatus(ctx context.Context) ([]ipc.MorningRoutineStat
return d.getMorningStatus(ctx), nil
}
func (d *daemonAPI) DayPlan(ctx context.Context) (ipc.DayPlan, error) {
if d.getDayPlan == nil {
return ipc.DayPlan{}, errors.New("mavend: day plan not available")
}
return d.getDayPlan(ctx), nil
}
func toIPCTickTrace(t loop.TickTrace) ipc.TickTrace {
rules := make([]ipc.RuleTrace, len(t.RuleTraces))
for i, r := range t.RuleTraces {
+111 -2
View File
@@ -46,7 +46,7 @@ func newTestTickLoop(t *testing.T, st *store.Store, sink delivery.Sink, digestCf
Nudges: st,
Reminders: st,
})
return newTickLoop(st, g, d, phraser.NewStub(), rules, time.Second, 5*time.Minute, 0, digestCfg, nil, nil)
return newTickLoop(st, g, d, phraser.NewStub(), rules, time.Second, 5*time.Minute, 0, digestCfg, nil, nil, nil)
}
func TestTickFiresRoutineWhenScheduleCrosses(t *testing.T) {
@@ -63,7 +63,7 @@ func TestTickFiresRoutineWhenScheduleCrosses(t *testing.T) {
sink := &fakeSink{}
d := delivery.NewDispatcher(delivery.Config{Voice: sink, Ntfy: sink, Telegram: sink, Nudges: st, Reminders: st})
rs := []routine.Routine{{Name: "morning", Cron: "0 12 * * *", Body: "полдень, время воды", Severity: 1}}
tl := newTickLoop(st, g, d, phraser.NewStub(), rules, time.Second, 5*time.Minute, 0, nil, rs, nil)
tl := newTickLoop(st, g, d, phraser.NewStub(), rules, time.Second, 5*time.Minute, 0, nil, rs, nil, nil)
// first tick: seeds, does not fire the routine.
tl.tick(ctx, now)
@@ -96,6 +96,115 @@ func TestTickFiresRoutineWhenScheduleCrosses(t *testing.T) {
}
}
// TestTickFiresAcceptedRoutineEveryInterval — Vikunja #366. An accepted routine
// with a 3-day interval must nudge every 3 days, not once. It also must not
// replay the occurrences it slept through: after a 30-day gap it nudges once.
func TestTickFiresAcceptedRoutineEveryInterval(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
accepted := refNow()
id, err := st.CreateProposedRoutine(ctx, "полить", "цветы", 3.0, accepted)
if err != nil {
t.Fatalf("CreateProposedRoutine: %v", err)
}
if err := st.AcceptProposedRoutine(ctx, id, accepted); err != nil {
t.Fatalf("AcceptProposedRoutine: %v", err)
}
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
const rule = "routine:полить цветы"
// Same day as the accept: not due yet.
markPresent(t, st, ctx, accepted)
tl.tick(ctx, accepted.Add(time.Hour))
if n := countSends(sink, rule); n != 0 {
t.Fatalf("routine fired %d times before its first interval passed, want 0", n)
}
// Three days later: the first nudge.
first := accepted.Add(3 * 24 * time.Hour)
markPresent(t, st, ctx, first)
tl.tick(ctx, first)
if n := countSends(sink, rule); n != 1 {
t.Fatalf("first interval: sends = %d, want 1", n)
}
// Next day: still inside the interval, silent.
sink.sends = nil
markPresent(t, st, ctx, first.Add(24*time.Hour))
tl.tick(ctx, first.Add(24*time.Hour))
if n := countSends(sink, rule); n != 0 {
t.Fatalf("mid-interval: sends = %d, want 0", n)
}
// Three days after the first nudge: it fires again. This is the bug —
// a one-shot reminder would never come back.
second := first.Add(3 * 24 * time.Hour)
markPresent(t, st, ctx, second)
tl.tick(ctx, second)
if n := countSends(sink, rule); n != 1 {
t.Fatalf("second interval: sends = %d, want 1 (a routine repeats)", n)
}
// A long silence must not turn into a backlog of missed nudges.
sink.sends = nil
late := second.Add(30 * 24 * time.Hour)
markPresent(t, st, ctx, late)
tl.tick(ctx, late)
if n := countSends(sink, rule); n != 1 {
t.Fatalf("after a 30-day gap: sends = %d, want exactly 1 (no backlog)", n)
}
}
// TestTickAcceptedRoutineRespectsQuietHours — routines are not reminders: they
// do not inherit the reminder gate bypass. Away presence drops a care-class
// nudge, and the routine stays due so it nudges once the user is back.
func TestTickAcceptedRoutineRespectsGate(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
accepted := refNow()
id, err := st.CreateProposedRoutine(ctx, "полить", "цветы", 3.0, accepted)
if err != nil {
t.Fatalf("CreateProposedRoutine: %v", err)
}
if err := st.AcceptProposedRoutine(ctx, id, accepted); err != nil {
t.Fatalf("AcceptProposedRoutine: %v", err)
}
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
const rule = "routine:полить цветы"
// No presence probes at all ⇒ away ⇒ the care gate blocks the nudge.
due := accepted.Add(3 * 24 * time.Hour)
tl.tick(ctx, due)
if n := countSends(sink, rule); n != 0 {
t.Fatalf("away: sends = %d, want 0 (routine must not bypass the gate)", n)
}
// Back at the desk a minute later: the nudge that was held now goes out.
back := due.Add(time.Minute)
markPresent(t, st, ctx, back)
tl.tick(ctx, back)
if n := countSends(sink, rule); n != 1 {
t.Fatalf("present again: sends = %d, want 1", n)
}
}
// countSends counts captured sends for one rule name.
func countSends(sink *fakeSink, rule string) int {
n := 0
for _, s := range sink.sends {
if s.RuleName == rule {
n++
}
}
return n
}
// refNow — fixed tick time so presence decay + since durations are deterministic.
func refNow() time.Time { return time.Date(2026, 6, 30, 12, 0, 0, 0, time.UTC) }
+257
View File
@@ -0,0 +1,257 @@
// mavend/vision.go — core's half of image understanding (Vikunja #252,
// docs/plans/07-vision.md).
//
// The split: any surface that can receive a picture (mavweb upload, a Telegram
// photo through mavpoll, a path he names) hands the bytes to core over
// ipc.MethodDescribeImage. Core stores them content-addressed under
// media.dir, prepares a downscaled JPEG, and asks a local vision server what it
// is. The description comes back as words; nothing about the image is echoed.
//
// Off unless configured twice over: no `media` block ⇒ nowhere to keep the
// bytes, so the method does not exist; no `vision` block with enabled + a local
// endpoint ⇒ the store is wired but the describing half refuses, and the method
// still does not exist. A surface cannot make Maven look at pictures by merely
// sending one.
//
// Two things this file deliberately does not do:
//
// - No cloud vision call, ever. internal/vision refuses a non-private
// endpoint at construction; there is no config shape here that could reach
// an upstream API even if someone wanted one.
// - No automatic memory. SaveNote is opt-in per call. Glancing at a screenshot
// is not the same act as remembering it, and a 1.7B-class VLM's guess about
// a photo is not a fact worth carrying around.
package main
import (
"context"
"errors"
"fmt"
"log"
"path/filepath"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/media"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/vision"
)
// prunePeriod — how often stored blobs are checked against media.retention.
// Hourly is far more often than needed for a 7-day retention and costs a
// directory walk over a handful of sidecars; the point is that the promise is
// kept by a loop that runs, not by an operator remembering a cron.
const prunePeriod = time.Hour
// mediaKeeper — the blob store plus the loop that enforces its retention. The
// two are one object because a store without the loop is a directory that grows
// forever, and shipping that would break the only interesting promise this
// capability makes.
type mediaKeeper struct {
store *media.Store
}
// openMediaStore builds the blob store from config, or returns nil when media is
// not configured. A relative dir resolves against StateDir, the same rule the db
// and socket paths follow.
func openMediaStore(cfg *config.Config) *mediaKeeper {
dir := cfg.Media.StoreDir()
if dir == "" {
return nil
}
if !filepath.IsAbs(dir) && cfg.StateDir != "" {
dir = filepath.Join(cfg.StateDir, dir)
}
st, err := media.Open(dir, cfg.Media.MaxBytes, time.Duration(cfg.Media.Retention))
if err != nil {
log.Printf("media: %v — image and audio intake disabled", err)
return nil
}
log.Printf("media: blob store at %s, retention %s", st.Dir(), st.Retention())
return &mediaKeeper{store: st}
}
// runPrune deletes over-retention blobs on a loop until ctx ends. It prunes once
// immediately, so a daemon restarted after a long downtime does not sit on a
// month of stale recordings until the first tick.
func (k *mediaKeeper) runPrune(ctx context.Context) {
prune := func() {
n, err := k.store.Prune()
if err != nil {
log.Printf("media: prune: %v", err)
return
}
if n > 0 {
log.Printf("media: pruned %d blob(s) older than %s", n, k.store.Retention())
}
}
prune()
t := time.NewTicker(prunePeriod)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
prune()
}
}
}
// visionIntake — one image at a time: store, prepare, describe, optionally note.
type visionIntake struct {
in *vision.Intake
st *store.Store
emb router.Embedder
now func() time.Time
}
// newVisionIntake returns nil when there is nothing to wire. keeper == nil means
// no media block, which disables the method outright; a missing or disabled
// vision block still wires the method, because storing an image and answering
// "I can't look at it yet" is more useful than pretending the surface does not
// exist — and it is exactly the state this box is in until a vision model is on
// disk.
func newVisionIntake(keeper *mediaKeeper, st *store.Store, emb router.Embedder, cfg *config.Config) *visionIntake {
if keeper == nil {
return nil
}
vc := cfg.Vision
maxDim := 0
var provider vision.Provider = vision.Disabled{}
if vc.LooksAtImages() {
p, err := vision.NewLocal(vision.Config{
Endpoint: vc.Endpoint,
Model: vc.Model,
Timeout: time.Duration(vc.Timeout),
MaxTokens: vc.MaxTokens,
Prompt: vc.Prompt,
})
if err != nil {
// A public endpoint, a hostname, a bad URL. Logged once here rather
// than failing every turn, and the store still works.
log.Printf("vision: %v — she can store images but not describe them", err)
} else {
provider = p
maxDim = vc.MaxDim
log.Printf("vision: enabled against %s", p.Endpoint())
}
} else {
log.Printf("vision: not configured — images are stored, not described")
}
return &visionIntake{
in: vision.NewIntake(keeper.store, provider, maxDim),
st: st,
emb: emb,
now: time.Now,
}
}
// describe handles one ipc.MethodDescribeImage call.
//
// A description failure is NOT an error out of this method when the bytes were
// stored: the caller gets the id and an empty description, which is honest ("it
// is kept, I cannot read it yet") and re-runnable. A failure to store, or bytes
// that are not an image at all, is an error — there is nothing to come back to.
func (v *visionIntake) describe(ctx context.Context, req ipc.DescribeImageReq) (ipc.DescribeImageResp, error) {
if len(req.Data) == 0 && req.ID == "" {
return ipc.DescribeImageResp{}, fmt.Errorf("describe image: neither data nor id")
}
var (
res vision.Result
err error
)
if req.ID != "" {
res, err = v.in.Rerun(ctx, req.ID, req.Question)
} else {
res, err = v.in.Accept(ctx, req.Data, sourceOrDefault(req.Source), req.Question)
}
if res.Blob.ID == "" {
// Nothing was stored: bad format, over the size cap, unwritable dir.
return ipc.DescribeImageResp{}, fmt.Errorf("describe image: %w", err)
}
resp := ipc.DescribeImageResp{
ID: res.Blob.ID,
Description: res.Description,
Width: res.Image.Width,
Height: res.Image.Height,
}
if err != nil {
// Bytes are safe, words are not available. The log names the blob and the
// reason; it never names what was in the picture.
if errors.Is(err, vision.ErrDisabled) {
log.Printf("vision: stored %s, no vision model configured", res.Blob)
} else {
log.Printf("vision: stored %s, describe failed: %v", res.Blob, err)
}
return resp, nil
}
if req.SaveNote {
id, werr := v.writeNote(ctx, res)
if werr != nil {
// The description is still returned: losing the note is worse as a
// silent failure than as a log line next to a successful answer.
log.Printf("vision: note write for %s failed: %v", res.Blob, werr)
} else {
resp.NoteID = id
}
}
log.Printf("vision: described %s (%dx%d)", res.Blob, res.Image.Width, res.Image.Height)
return resp, nil
}
// writeNote stores the description as an ordinary note so it is recallable. The
// note carries the blob id in its source, which is the only link back to the
// bytes — the note text is words about the picture, never the picture.
func (v *visionIntake) writeNote(ctx context.Context, res vision.Result) (int64, error) {
var vec []float32
if v.emb != nil {
// EmbedPassage, not Embed: a description is text being searched FOR, and
// the e5 embedder is asymmetric. Backwards here makes it unfindable by
// the question that should have matched it.
var err error
vec, err = router.EmbedPassage(ctx, v.emb, res.Description)
if err != nil {
return 0, fmt.Errorf("embed: %w", err)
}
}
source := "media:image:" + res.Blob.ID[:12]
return v.st.WriteNote(ctx, v.now(), res.Description, vec, source)
}
// sourceOrDefault labels a blob whose sender did not say where it came from.
func sourceOrDefault(s string) string {
if s == "" {
return "unknown"
}
return s
}
// wireVision installs the IPC hook and starts the retention loop, or leaves the
// hook nil so ipc.MethodDescribeImage reports ErrUnknownMethod. Called on both
// startup paths (unlocked boot and passkey unlock) so vision behaves the same
// either way.
//
// Returns the media keeper so the meeting recorder can share it: one blob store
// with one retention loop holds both the images and the audio, which is the
// whole point of internal/media being a shared package. nil ⇒ no media block,
// and neither capability exists.
func wireVision(ctx context.Context, srv *ipc.Server, st *store.Store, emb router.Embedder, cfg *config.Config) *mediaKeeper {
keeper := openMediaStore(cfg)
if keeper == nil {
return nil
}
go keeper.runPrune(ctx)
vi := newVisionIntake(keeper, st, emb, cfg)
if vi == nil {
return keeper
}
srv.DescribeImageFn = vi.describe
return keeper
}
+121 -1450
View File
File diff suppressed because it is too large Load Diff
+478
View File
@@ -0,0 +1,478 @@
package main
import (
"bufio"
"context"
"fmt"
"log"
"os"
"path/filepath"
"strings"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/delivery/voicesink"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/stt"
"github.com/kami/maven/internal/tool"
"github.com/kami/maven/internal/tts"
"github.com/kami/maven/internal/voice"
"github.com/kami/maven/internal/weather"
"github.com/kami/maven/internal/worker"
)
// voiceWiring — everything the daemon needs to run the audio path. Held by
// cmd/mavend/main.go alongside the other wirings; closed on shutdown.
type voiceWiring struct {
server *voice.Server
sessions *voice.Sessions
voiceSink delivery.Sink
embedder router.Embedder
handler *reactiveHandler // the reactive handler for IPC Chat
// worker clients (set when configured as Remote): closed on shutdown so
// mavsttd / mavttsd don't keep a stale conn into a restarting daemon.
sttClient *worker.Client
ttsClient *worker.Client
// transcriber — the STT in use, exposed so the meeting recorder
// (cmd/mavend/capture.go) can reuse it. Maven has exactly one STT and does
// not grow a second one for capture: this is the same whisper.cpp worker the
// voice path talks to.
transcriber stt.Transcriber
// mcp — the MCP client, nil unless the `mcp` block configures an enabled
// server (Vikunja #251). Its tools land in the same allowlist as every
// other act, so nothing else here has to know about it.
mcp *mcpWiring
// home — the Home Assistant client, nil unless the `smarthome` block is
// enabled (Vikunja #256). Its devices land in the same allowlist as every
// other act, so nothing else here has to know about it.
home *homeWiring
// netscan — the LAN scanner, nil unless the `netscan` block is enabled
// (Vikunja #257).
netscan *netWiring
}
// close releases the listener + worker conns. Safe to call on nil (when
// voice is not wired — wireVoice returns nil,nil).
func (w *voiceWiring) close() {
if w == nil {
return
}
if w.embedder != nil {
_ = w.embedder.Close()
}
if w.server != nil {
_ = w.server.Close()
}
if w.sttClient != nil {
_ = w.sttClient.Close()
}
if w.ttsClient != nil {
_ = w.ttsClient.Close()
}
w.mcp.close()
}
// wireVoice builds the audio path from cfg + a CoreAPI + a router. Returns
// nil wiring + nil error when voice isn't enabled (the caller's voice sink
// stays nil; the dispatcher's ChannelVoice routing drops silently).
//
// When voice is enabled, MUST wire a voicesink into the dispatcher's Voice
// slot using w.sessions (the caller does that — see main.go).
func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, memStore memory.Store, dataStore *store.Store, eco *ecosystemWiring) (*voiceWiring, error) {
if cfg.Voice == nil || !cfg.Voice.Enabled {
return nil, nil
}
w := &voiceWiring{}
// ----- stt (Stub in-process OR Remote via worker socket) -----
var transcriber stt.Transcriber
if cfg.Voice.Stt != nil && cfg.Voice.Stt.Socket != "" {
c := worker.Dial(cfg.Voice.Stt.Socket)
w.sttClient = c
lang := cfg.Voice.Stt.Lang
if lang == "" {
lang = cfg.Voice.Lang
}
transcriber = stt.NewRemote(c, lang)
} else {
transcriber = stt.NewStub()
}
w.transcriber = transcriber
// ----- tts (Stub in-process OR Remote) -----
var synthesizer tts.Synthesizer
if cfg.Voice.Tts != nil && cfg.Voice.Tts.Socket != "" {
c := worker.Dial(cfg.Voice.Tts.Socket)
w.ttsClient = c
lang := cfg.Voice.Tts.Lang
if lang == "" {
lang = cfg.Voice.Lang
}
synthesizer = tts.NewRemote(c, lang, cfg.Voice.Tts.Voice)
} else {
synthesizer = tts.NewStub()
}
// ----- router: embedder (ONNX when configured, floor HashEmbedder otherwise) -----
var emb router.Embedder
if cfg.Voice.Embedder != nil {
onnx, err := router.NewONNXEmbedder(
cfg.Voice.Embedder.ModelPath,
cfg.Voice.Embedder.TokenizerPath,
cfg.Voice.Embedder.LibPath,
)
if err != nil {
w.close()
return nil, fmt.Errorf("embedder: %w", err)
}
log.Printf("voice: onnx embedder loaded (%d dim)", onnx.Dim())
emb = onnx
} else {
log.Printf("voice: embedder not configured, using HashEmbedder floor")
emb = router.NewHashEmbedder(1024)
}
w.embedder = emb
checkStoredEmbedder(dataStore, emb)
// ----- tool executor (the enabled act allowlist, store-backed) -----
// Config tools are the declarative bootstrap: seed them into the store as
// enabled (editing mavend.json IS the human enable act). Ad-hoc tools are
// enabled later through the authed mavweb surface. The executor + matcher
// both read the store live, so a newly-enabled tool is runnable without a
// daemon restart.
seedTools(coreAPI, cfg.Voice.Tools)
exec := tool.NewExecutor(coreAPI, time.Duration(cfg.Voice.ToolTimeout))
// MCP servers (Vikunja #251): discovery PROPOSES tools into the same
// allowlist, so an MCP tool is enabled by hand on /tools like any other and
// runs through the same confirm turn. Off unless the `mcp` block configures
// an enabled server.
w.mcp = wireMCP(cfg, dataStore)
if w.mcp != nil {
exec = exec.WithMCP(w.mcp.caller())
}
// The house (Vikunja #256): same story as MCP. Discovery PROPOSES a row per
// controllable device, always destructive, and Kami enables the ones he
// wants on /tools. Off unless the `smarthome` block is enabled.
w.home = wireSmartHome(cfg, dataStore)
if w.home != nil {
exec = exec.WithHome(w.home.caller())
}
// The LAN scanner (Vikunja #257): a read, bounded to the configured
// subnets and rate-limited. Off unless the `netscan` block is enabled.
w.netscan = wireNetScan(cfg)
matcher := tool.NewMatcher(coreAPI)
// ----- weather provider (Open-Meteo when configured, Stub otherwise) -----
var weatherProvider weather.Provider
var weatherLocation string
if cfg.Voice.Weather != nil && cfg.Voice.Weather.Provider == "open-meteo" {
weatherProvider = weather.NewOpenMeteoProvider()
weatherLocation = cfg.Voice.Weather.DefaultLocation
log.Printf("voice: weather provider: open-meteo (default location: %s)", cfg.Voice.Weather.DefaultLocation)
} else {
weatherProvider = weather.NewStubProvider()
log.Printf("voice: weather provider: stub (not configured)")
}
// The replier uses the same llama-server as the phraser.
var llmClient *llm.Client
if lp, ok := phr.(*phraser.LLMPhraser); ok {
// llmClientFor, not llm.New: this client must follow the phraser onto
// the new llama-server when the resident model is swapped (Vikunja #250).
llmClient = llmClientFor(lp, 60*time.Second)
}
// ----- router (the cascade; floor examples seed the classifier) -----
// The act matcher's allowlist is exactly the enabled tool names — the
// router only matches acts the executor can run (one source of truth).
threshold := cfg.Voice.RouterThreshold
if threshold <= 0 {
threshold = config.DefaultRouterThreshold
}
// The resident model routes by default: 63.2% of held-out intents right
// against the classifier's 50.0%, at about 1s a turn instead of 30ms (see
// config.VoiceConfig.LLMRouter). The classifier always stays wired as the
// fallback, so a model error never breaks a turn.
rtr := buildRouter(emb, matcher, threshold, pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient))
// ----- sessions registry (shared with voicesink) -----
sessions := voice.NewSessions()
w.sessions = sessions
// ----- voice sink (proactive nudges: dispatcher → voicesink → tts → push to client) -----
w.voiceSink = voicesink.New(synthesizer, sessions)
// ----- memory (long-term vector storage) -----
// Persistent (store-backed, survives restarts) when the daemon passes one;
// falls back to the in-memory floor otherwise (tests / no-store paths).
if memStore == nil {
memStore = memory.NewInMemoryStore()
}
// ----- dialogue (multi-turn slot carry-over; 2-min follow-up window) -----
// Store-backed when the daemon passes a store, so a restart mid-conversation
// keeps the thread (Vikunja #363). Sessions past their TTL are dropped on
// load, never revived. Clarify's parked question stays in memory only.
var dialogueSessions *dialogue.SessionStore
if dataStore != nil {
dialogueSessions = dialogue.NewPersistentSessionStore(2*time.Minute, dataStore)
if err := dialogueSessions.Load(context.Background(), time.Now()); err != nil {
log.Printf("dialogue: load saved sessions: %v", err)
}
} else {
dialogueSessions = dialogue.NewSessionStore(2 * time.Minute)
}
clarifyStore := dialogue.NewClarifyStore(clarifyTTL)
timeParser := router.NewPythonDateParser()
// ----- replier (LLM-backed when the engine is on, Stub floor otherwise) -----
replier := voice.Replier(voice.NewStubReplier())
if llmClient != nil {
replier = newLLMReplier(llmClient, contextBlockFn(cfg, time.Now))
}
// ----- the handler (the reactive path; closes over stt / tts / router / coreAPI / memory) -----
h := &reactiveHandler{
stt: transcriber,
tts: synthesizer,
router: rtr,
embedder: emb,
api: coreAPI,
tools: exec,
matcher: matcher,
replier: replier,
phraser: phr,
now: time.Now,
feedsOn: cfg.Feeds != nil,
home: w.home,
netscan: w.netscan,
// nil unless `crawl.on_demand` is on: reading a page he names is a
// capability, and capabilities are off unless configured.
crawler: onDemandCrawler(cfg),
weatherProvider: weatherProvider,
weatherLocation: weatherLocation,
memStore: memStore,
dataStore: dataStore,
dialogueSessions: dialogueSessions,
clarifyStore: clarifyStore,
// 0 here (unset config) ⇒ the dialogue default.
clarifyMaxAttempts: cfg.Voice.ClarifyMaxAttempts,
extractor: router.Extractor{Time: timeParser, Acts: matcher, Facts: router.DefaultFactParser{}},
queryMinScore: cfg.Voice.QueryMinScore,
queryMinMargin: cfg.Voice.QueryMinMargin,
timeParser: timeParser,
ecosystem: eco,
}
// ----- the server (TCP listener) -----
srv := voice.NewServer(cfg.Voice.Bind, h, sessions)
if err := srv.Listen(); err != nil {
w.close()
return nil, fmt.Errorf("voice listen: %w", err)
}
w.server = srv
w.handler = h
return w, nil
}
// pickLLMRouter returns the LLM router when the operator asked for it and there
// is a llama-server to talk to, and nil otherwise. nil is safe: the cascade then
// routes with the classifier, so an unusable setting costs accuracy, not turns.
func pickLLMRouter(enabled bool, c *llm.Client) *router.LLMRouter {
if !enabled {
return nil
}
if c == nil {
log.Printf("voice: voice.llm_router is on but there is no llama-server to route with (the phraser is not an LLM phraser) — using the classifier instead")
return nil
}
log.Printf("voice: LLM router enabled")
return router.NewLLMRouter(c)
}
// buildRouter constructs the reactive-path router with the given embedder
// and confidence threshold.
// - stage-0 grammars from DefaultActMatcher whose fn allowlist is exactly
// the enabled tool names (actFns) — the router only matches acts the
// executor can run. Empty ⇒ every act refuses at the matcher.
// - The embedder is provided by wireVoice: HashEmbedder (floor) when no
// embedder config is present, or the ONNX multilingual model when
// configured — same interface, one constructor change.
// - 6 bootstrap examples covering the 5 intents + one compound-capture
// placeholder. Spec calls for ~10 per intent at production; this is the
// bootstrapping floor swapped by tuning the seed set later.
// - Threshold is from voice.router_threshold config (default 0.55).
func buildRouter(emb router.Embedder, acts router.ActMatcher, threshold float64, llmR *router.LLMRouter) *router.Router {
cls := router.NewClassifier(emb)
seedClassifier(cls)
grammars := router.DefaultGrammars(acts)
grammars = append(grammars, router.SystemTimeDateGrammars()...)
grammars = append(grammars, router.ReminderGrammar())
return router.New(router.Config{
Grammars: grammars,
Classifier: cls,
Extractor: router.Extractor{
Time: router.NewPythonDateParser(),
Acts: acts,
Facts: router.DefaultFactParser{},
},
Threshold: threshold,
LLM: llmR,
})
}
// seedDir is the directory containing intent seed files. Each file is named
// <intent>.txt and contains one training example per line (blank lines and
// lines starting with # are ignored). Relative to the working directory.
const seedDir = "models/seeds"
// seedClassifier floors the embedded examples so the cold-boot path
// doesn't return ErrNoIntents. Loads examples from seedDir — one file per
// intent (act.txt, reminder.txt, fact.txt, note.txt, query.txt). When the
// classifier can't decide it falls through to Clarify — the last-resort
// path asks the user to rephrase rather than guessing wrong.
func seedClassifier(c *router.Classifier) {
intents := []router.Intent{
router.IntentAct,
router.IntentReminder,
router.IntentFact,
router.IntentNote,
router.IntentQuery,
router.IntentChat,
router.IntentSystem,
}
total := 0
for _, intent := range intents {
n, err := loadSeedFile(c, intent)
if err != nil {
log.Printf("voice: seed %s: %v", intent, err)
continue
}
total += n
}
log.Printf("voice: loaded %d seed examples from %s", total, seedDir)
}
func loadSeedFile(c *router.Classifier, intent router.Intent) (int, error) {
path := filepath.Join(seedDir, string(intent)+".txt")
f, err := os.Open(path)
if err != nil {
return 0, fmt.Errorf("open %s: %w", path, err)
}
defer f.Close()
var count int
sc := bufio.NewScanner(f)
for sc.Scan() {
line := strings.TrimSpace(sc.Text())
if line == "" || strings.HasPrefix(line, "#") {
continue
}
if err := c.AddExample(context.Background(), intent, line); err != nil {
log.Printf("voice: seed %s: skipping %q: %v", intent, line, err)
continue
}
count++
}
if err := sc.Err(); err != nil {
return count, fmt.Errorf("scan %s: %w", path, err)
}
return count, nil
}
// seedTools upserts the config-declared tools into the store as enabled. Editing
// mavend.json is a human act, so a config tool is enabled by definition; this
// makes the declarative config the reproducible bootstrap while the store stays
// the single runtime source of truth (mavweb enables ad-hoc ones on top).
func seedTools(api ipc.CoreAPI, tools []config.ToolConfig) {
ctx := context.Background()
now := time.Now()
n := 0
for _, tc := range tools {
if tc.Name == "" || len(tc.Cmd) == 0 {
log.Printf("voice: skipping malformed tool config %+v", tc)
continue
}
if err := api.EnableTool(ctx, tc.Name, tc.Cmd, tc.Destructive, tc.Scope, now); err != nil {
log.Printf("voice: seed tool %q: %v", tc.Name, err)
continue
}
n++
}
log.Printf("voice: seeded %d act tools from config", n)
}
// reembedOnStart is the -reembed flag (set in run()). Opt-in on purpose: see
// runReembed.
var reembedOnStart bool
// checkStoredEmbedder compares the embedder we just loaded with the one that
// wrote the vectors already in the DB (Vikunja #378).
//
// The two models we have both make 384-dim vectors, so a size check catches
// nothing: after a swap, recall silently compares vectors from different
// spaces and the scores are noise. So we say it out loud. Recall itself is not
// changed here — the fix is `mavend -reembed`.
func checkStoredEmbedder(dataStore *store.Store, emb router.Embedder) {
if dataStore == nil {
return
}
current := router.EmbedderID(emb)
if reembedOnStart {
runReembed(dataStore, emb, current)
return
}
stored, mismatch, err := dataStore.CheckEmbedder(context.Background(), current)
if err != nil {
log.Printf("voice: embedder marker check failed: %v", err)
return
}
if mismatch {
log.Printf("voice: WARNING embedder MISMATCH — stored vectors were written by %q but the configured embedder is %q; recall scores are noise until the notes and facts are re-embedded — run `mavend -reembed` once (Vikunja #378)", stored, current)
return
}
log.Printf("voice: embedder marker ok (%s)", current)
}
// runReembed is the one-shot backfill behind -reembed.
//
// Why a flag and not automatic on mismatch: the embedder is ONNX on the
// laptop's CPU, so a few thousand notes is minutes of work. Doing that silently
// inside a normal start would look like the daemon hanging on boot. So the user
// runs it once, deliberately, after an embedder swap; the mismatch warning
// above tells them to. It re-embeds, logs what it did, and then the daemon
// carries on serving as usual — no separate binary, no second start needed.
func runReembed(dataStore *store.Store, emb router.Embedder, current string) {
log.Printf("voice: re-embedding stored notes and facts with %s — this can take a few minutes, do not interrupt", current)
res, err := dataStore.ReembedAll(context.Background(), current,
// EmbedPassage, not EmbedQuery: these are stored texts being searched
// FOR, which is the side they were written with.
func(ctx context.Context, text string) ([]float32, error) {
return router.EmbedPassage(ctx, emb, text)
})
if err != nil {
log.Printf("voice: re-embed FAILED, nothing was changed and no marker was written — safe to run again: %v", err)
return
}
if res.Skipped {
log.Printf("voice: re-embed skipped — the stored vectors were already written by %s", current)
return
}
log.Printf("voice: re-embed done — %d notes in the notes table, %d notes and %d facts in the memory index, took %s; stored vectors now belong to %s",
res.Notes, res.MemNotes, res.Facts, res.Took.Round(time.Second), current)
// A row with no text cannot be re-embedded, so its vector is still the old
// model's noise while the marker now says everything is current. Both write
// paths always store the text, so this should be zero — say it loudly
// rather than bury it in the line above if it ever isn't.
if res.NoText > 0 {
log.Printf("voice: WARNING %d stored rows had no text, so their vectors could not be re-embedded and are still noise; they will never match anything useful (Vikunja #378)", res.NoText)
}
}
+50
View File
@@ -0,0 +1,50 @@
// Package main — weatherq.go holds the weather-query keyword helpers: does
// this utterance ask about weather at all, and which city (if any) did it
// name. Both are plain substring/lookup matching, not NLU — extend this file
// rather than voice.go for anything in that shape.
package main
import "strings"
// isWeatherQuery returns true if the utterance is about weather.
func isWeatherQuery(u string) bool {
lower := strings.ToLower(u)
return strings.Contains(lower, "погод") ||
strings.Contains(lower, "градус") ||
strings.Contains(lower, "температур") ||
strings.Contains(lower, "дожд") ||
strings.Contains(lower, "холод") ||
strings.Contains(lower, "тепл") ||
strings.Contains(lower, "weather") ||
strings.Contains(lower, "temperature")
}
// extractWeatherLocation parses a location from the utterance, or falls back
// to the configured default. Very basic: just checks for known city names.
func extractWeatherLocation(u, defaultLoc string) string {
lower := strings.ToLower(u)
cities := map[string]string{
"москв": "Moscow",
"moscow": "Moscow",
"питер": "Saint Petersburg",
"spb": "Saint Petersburg",
"петербур": "Saint Petersburg",
"лондон": "London",
"london": "London",
"париж": "Paris",
"paris": "Paris",
"берлин": "Berlin",
"berlin": "Berlin",
"нью-йорк": "New York",
"new york": "New York",
}
for substr, name := range cities {
if strings.Contains(lower, substr) {
return name
}
}
if defaultLoc != "" {
return defaultLoc
}
return "Moscow"
}
+332
View File
@@ -0,0 +1,332 @@
// mavmaild — the mail reader module (Vikunja #246,
// docs/plans/01-email-reader.md).
//
// Every so often it opens one IMAP mailbox read-only, fetches the messages it
// has not read yet, and hands each one to core over ipc.MethodIngestMail. Core
// runs the extraction on the resident model and writes what comes back as task
// CANDIDATES he reviews on /tasks. Nothing here writes to the store, nothing
// here can create a reminder, and nothing here speaks.
//
// Why a separate daemon rather than a loop inside mavend, when extraction has
// to happen in mavend anyway: the credential. mavpoll set the precedent with the
// zenmoney token (#125) — the module that talks to a third party holds the
// secret, reads it from a FILE so it never appears in `ps`, in
// docker-compose.yml or in shell history, and core never sees it. Core learns
// that mail exists only as message text on one IPC method; it cannot connect to
// the mailbox even if it wanted to, and a compromised core yields no mail
// password.
//
// Off unless configured: without -password-file there is nothing to run, and
// the daemon says so and exits. If core has no `email` block the very first
// ingest comes back ErrUnknownMethod and this daemon stops polling instead of
// hammering a socket that will keep refusing.
//
// Mail is personal, so the log is counts and UIDs: how many messages were
// fetched, how many were bulk, how many candidates came back. No subject, no
// sender, no body, ever — reviewing a candidate is what /tasks is for.
package main
import (
"context"
"encoding/json"
"errors"
"flag"
"fmt"
"log"
"os"
"os/signal"
"path/filepath"
"sort"
"strings"
"syscall"
"time"
"github.com/kami/maven/internal/email"
"github.com/kami/maven/internal/ipc"
)
func main() {
if err := run(os.Args[1:]); err != nil {
fmt.Fprintln(os.Stderr, "mavmaild:", err)
os.Exit(1)
}
}
func run(args []string) error {
fs := flag.NewFlagSet("mavmaild", flag.ContinueOnError)
socket := fs.String("socket", "", "core IPC socket path (required)")
server := fs.String("imap", "", "IMAP server, host or host:993 (required)")
user := fs.String("user", "", "IMAP username (required)")
passFile := fs.String("password-file", "", "file holding the IMAP password (required — never passed as a flag value)")
mailbox := fs.String("mailbox", "INBOX", "mailbox to read, read-only")
interval := fs.Duration("interval", 15*time.Minute, "how often to read the mailbox")
lookback := fs.Duration("lookback", 72*time.Hour, "how far back to search on each poll")
max := fs.Int("max", 25, "most messages to fetch in one poll")
timeout := fs.Duration("timeout", 30*time.Second, "IMAP network timeout")
statePath := fs.String("state", "", "file remembering which UIDs were read (default: none — every poll re-reads the window)")
if err := fs.Parse(args); err != nil {
return err
}
if *socket == "" {
return fmt.Errorf("-socket is required")
}
if *server == "" || *user == "" || *passFile == "" {
return fmt.Errorf("mail reading is off unless configured: set -imap, -user and -password-file")
}
// The password is read from a file, never taken as a flag value: an argv
// secret is visible in `ps` to every user on the box and lands in the compose
// file and the shell history. Read once at start — a rotated password means a
// restart, which is cheaper than re-reading his credential every quarter hour.
raw, err := os.ReadFile(*passFile)
if err != nil {
return fmt.Errorf("read password file: %w", err)
}
password := strings.TrimSpace(string(raw))
if password == "" {
return fmt.Errorf("password file %s is empty", *passFile)
}
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
defer stop()
core, err := ipc.DialWait(*socket, 60*time.Second)
if err != nil {
return err
}
defer core.Close()
r := &reader{
core: core,
addr: *server,
user: *user,
mailbox: *mailbox,
lookback: *lookback,
max: *max,
timeout: *timeout,
state: newSeenState(*statePath),
}
if err := r.state.load(); err != nil {
// A missing or corrupt state file must not stop mail from being read: the
// worst case is re-reading the window, and capture dedupes on text.
log.Printf("mavmaild: state: %v (starting from an empty seen-set)", err)
}
// The password is never logged, not even its length.
log.Printf("mavmaild: reading %s on %s every %s (lookback %s, max %d/poll)",
*mailbox, *server, *interval, *lookback, *max)
r.pollOnce(ctx, password) // don't idle a full interval on start
t := time.NewTicker(*interval)
defer t.Stop()
for {
select {
case <-ctx.Done():
log.Printf("mavmaild: bye")
return nil
case <-t.C:
if r.disabled {
// Core told us mail ingestion is not configured. Nothing will change
// without a core restart, and a restart restarts us too.
log.Printf("mavmaild: core does not accept mail — idling")
return nil
}
r.pollOnce(ctx, password)
}
}
}
// mailIngester — the slice of core this daemon uses. One method: hand over a
// message. It cannot write a fact, create a reminder or read the store, and the
// interface says so.
type mailIngester interface {
IngestMail(ctx context.Context, req ipc.IngestMailReq) (ipc.IngestMailResp, error)
}
type reader struct {
core mailIngester
addr string
user string
mailbox string
lookback time.Duration
max int
timeout time.Duration
state *seenState
// dial — connection seam for the tests; nil ⇒ implicit TLS.
dial func(addr string, timeout time.Duration) (*email.Conn, error)
// disabled — core answered ErrUnknownMethod, i.e. it has no email block.
disabled bool
}
// pollOnce — one read of the mailbox, then one ingest per message.
//
// A fetch error aborts this poll and nothing else; the next tick tries again.
// An ingest error for one message does not skip the rest — one mail the model
// choked on should not hide the four behind it.
func (r *reader) pollOnce(ctx context.Context, password string) {
msgs, err := r.fetch(password)
if err != nil {
// The error may name a UID; it never names a subject or a sender.
log.Printf("mavmaild: fetch: %v", err)
if len(msgs) == 0 {
return
}
}
var junk, candidates, created int
for _, m := range msgs {
if ctx.Err() != nil {
return
}
if m.Junk {
junk++
// Marked seen without a model call: the header filter already decided,
// and re-classifying it every quarter hour would be pure waste.
r.state.mark(m.UID)
continue
}
resp, err := r.core.IngestMail(ctx, ipc.IngestMailReq{
Mailbox: r.mailbox,
UID: m.UID,
From: m.From,
Subject: m.Subject,
Date: m.Date,
Body: m.Body,
})
if errors.Is(err, ipc.ErrUnknownMethod) {
log.Printf("mavmaild: core has no email block configured — mail ingestion is off; stopping")
r.disabled = true
return
}
if err != nil {
// Not marked seen: an ingest that failed should be retried next poll.
log.Printf("mavmaild: ingest uid %d: %v", m.UID, err)
continue
}
r.state.mark(m.UID)
candidates += len(resp.TaskIDs)
created += resp.Created
}
if err := r.state.save(); err != nil {
log.Printf("mavmaild: state: %v", err)
}
log.Printf("mavmaild: %s: %d read, %d bulk, %d candidate(s), %d new", r.mailbox, len(msgs), junk, candidates, created)
}
// fetch reads the mailbox. Messages already in the seen-set are not fetched at
// all, so a steady mailbox costs one SEARCH per poll and nothing else.
func (r *reader) fetch(password string) ([]email.Message, error) {
f := email.FetchSince{
Addr: r.addr,
User: r.user,
Mailbox: r.mailbox,
Timeout: r.timeout,
Since: time.Now().Add(-r.lookback),
Max: r.max,
Skip: r.state.seen,
}
return f.RunWith(password, r.dial)
}
// ---- seen state ------------------------------------------------------------
// seenState — the UIDs already handed to core, persisted so a restart does not
// re-read (and re-extract, at multi-second LLM cost) the whole lookback window.
//
// Correctness does not depend on it: ipc.CaptureTask dedupes on normalised text
// among live tasks, so a re-read produces no duplicate rows. This exists to save
// the model's time, which is why a broken state file is a log line rather than a
// failure.
//
// UIDs are per-mailbox and monotonic, so the set is kept as a high-water mark
// plus the stragglers above it. If the server ever changes UIDVALIDITY, UIDs
// reset and the window is simply re-read once — dedupe absorbs it.
type seenState struct {
path string
high uint32
set map[uint32]bool
dirty bool
}
func newSeenState(path string) *seenState {
return &seenState{path: path, set: map[uint32]bool{}}
}
type seenFile struct {
High uint32 `json:"high"`
UIDs []uint32 `json:"uids,omitempty"`
}
func (s *seenState) seen(uid uint32) bool {
return uid <= s.high || s.set[uid]
}
func (s *seenState) mark(uid uint32) {
if s.seen(uid) {
return
}
s.set[uid] = true
s.dirty = true
// Advance the high-water mark through any contiguous run, so the explicit set
// stays small on a mailbox read in order.
for {
next := s.high + 1
if !s.set[next] {
break
}
delete(s.set, next)
s.high = next
}
}
func (s *seenState) load() error {
if s.path == "" {
return nil
}
b, err := os.ReadFile(s.path)
if errors.Is(err, os.ErrNotExist) {
return nil // first run
}
if err != nil {
return err
}
var f seenFile
if err := json.Unmarshal(b, &f); err != nil {
return fmt.Errorf("parse %s: %w", s.path, err)
}
s.high = f.High
for _, u := range f.UIDs {
s.set[u] = true
}
return nil
}
// save writes the state atomically (temp file + rename), 0600: it is a list of
// message ids from his mailbox, which is metadata about his mail.
func (s *seenState) save() error {
if s.path == "" || !s.dirty {
return nil
}
uids := make([]uint32, 0, len(s.set))
for u := range s.set {
uids = append(uids, u)
}
sort.Slice(uids, func(i, j int) bool { return uids[i] < uids[j] })
b, err := json.Marshal(seenFile{High: s.high, UIDs: uids})
if err != nil {
return err
}
tmp := s.path + ".tmp"
if err := os.MkdirAll(filepath.Dir(s.path), 0o700); err != nil {
return err
}
if err := os.WriteFile(tmp, b, 0o600); err != nil {
return err
}
if err := os.Rename(tmp, s.path); err != nil {
return err
}
s.dirty = false
return nil
}
+254
View File
@@ -0,0 +1,254 @@
package main
import (
"bufio"
"context"
"fmt"
"net"
"os"
"path/filepath"
"strconv"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/email"
"github.com/kami/maven/internal/ipc"
)
// ---- a scripted IMAP server, same shape internal/email's tests use ---------
type fakeIMAP struct {
msgs map[uint32]string
uids []uint32
cmds []string
}
func (f *fakeIMAP) serve(c net.Conn) {
defer c.Close()
fmt.Fprint(c, "* OK fake ready\r\n")
r := bufio.NewReader(c)
for {
line, err := r.ReadString('\n')
if err != nil {
return
}
parts := strings.SplitN(strings.TrimRight(line, "\r\n"), " ", 2)
if len(parts) != 2 {
return
}
tag, cmd := parts[0], parts[1]
f.cmds = append(f.cmds, cmd)
upper := strings.ToUpper(cmd)
switch {
case strings.HasPrefix(upper, "LOGIN"), strings.HasPrefix(upper, "EXAMINE"):
fmt.Fprintf(c, "%s OK\r\n", tag)
case strings.HasPrefix(upper, "UID SEARCH"):
var ids []string
for _, u := range f.uids {
ids = append(ids, strconv.FormatUint(uint64(u), 10))
}
fmt.Fprintf(c, "* SEARCH %s\r\n%s OK\r\n", strings.Join(ids, " "), tag)
case strings.HasPrefix(upper, "UID FETCH"):
uid, _ := strconv.ParseUint(strings.Fields(cmd)[2], 10, 32)
if raw, ok := f.msgs[uint32(uid)]; ok {
fmt.Fprintf(c, "* 1 FETCH (UID %d BODY[] {%d}\r\n%s)\r\n", uid, len(raw), raw)
}
fmt.Fprintf(c, "%s OK\r\n", tag)
case strings.HasPrefix(upper, "LOGOUT"):
fmt.Fprintf(c, "* BYE\r\n%s OK\r\n", tag)
return
default:
fmt.Fprintf(c, "%s BAD\r\n", tag)
}
}
}
func (f *fakeIMAP) dial(_ string, timeout time.Duration) (*email.Conn, error) {
cli, srv := net.Pipe()
go f.serve(srv)
return email.NewConn(cli, timeout)
}
// ---- a fake core -----------------------------------------------------------
type fakeCore struct {
got []ipc.IngestMailReq
resp ipc.IngestMailResp
err error
}
func (c *fakeCore) IngestMail(_ context.Context, req ipc.IngestMailReq) (ipc.IngestMailResp, error) {
c.got = append(c.got, req)
if c.err != nil {
return ipc.IngestMailResp{}, c.err
}
return c.resp, nil
}
func mail(subject, body string, extraHeaders ...string) string {
h := "Subject: " + subject + "\r\nFrom: a@b.c\r\nContent-Type: text/plain; charset=utf-8\r\n"
for _, e := range extraHeaders {
h += e + "\r\n"
}
return h + "\r\n" + body + "\r\n"
}
func newTestReader(t *testing.T, f *fakeIMAP, core *fakeCore, statePath string) *reader {
t.Helper()
return &reader{
core: core, addr: "mail.example:993", user: "kami", mailbox: "INBOX",
lookback: 72 * time.Hour, max: 25, timeout: 5 * time.Second,
state: newSeenState(statePath),
dial: f.dial,
}
}
func TestPollHandsMessagesToCore(t *testing.T) {
f := &fakeIMAP{
uids: []uint32{1, 2},
msgs: map[uint32]string{
1: mail("Счёт", "Оплатить до 5 августа."),
2: mail("Скидки", "Sale!", "List-Unsubscribe: <mailto:u@x>"),
},
}
core := &fakeCore{resp: ipc.IngestMailResp{TaskIDs: []int64{1}, Created: 1}}
r := newTestReader(t, f, core, "")
r.pollOnce(context.Background(), "secret")
// The newsletter is filtered before core is asked: only the real mail crosses.
if len(core.got) != 1 {
t.Fatalf("core saw %d messages, want 1 (the bulk one must not cross): %+v", len(core.got), core.got)
}
got := core.got[0]
if got.UID != 1 || got.Mailbox != "INBOX" || got.Subject != "Счёт" {
t.Errorf("ingest req = %+v", got)
}
if !strings.Contains(got.Body, "Оплатить") {
t.Errorf("body = %q", got.Body)
}
}
// A second poll must not re-send what core already saw — extraction is a
// multi-second LLM call per message.
func TestPollSkipsSeenUIDs(t *testing.T) {
f := &fakeIMAP{uids: []uint32{5}, msgs: map[uint32]string{5: mail("Счёт", "текст")}}
core := &fakeCore{}
r := newTestReader(t, f, core, "")
r.pollOnce(context.Background(), "secret")
r.pollOnce(context.Background(), "secret")
if len(core.got) != 1 {
t.Errorf("core saw %d messages over two polls, want 1", len(core.got))
}
}
// An ingest that failed is NOT marked seen: the next poll retries it.
func TestPollRetriesFailedIngest(t *testing.T) {
f := &fakeIMAP{uids: []uint32{5}, msgs: map[uint32]string{5: mail("Счёт", "текст")}}
core := &fakeCore{err: fmt.Errorf("llama-server is warming up")}
r := newTestReader(t, f, core, "")
r.pollOnce(context.Background(), "secret")
core.err = nil
r.pollOnce(context.Background(), "secret")
if len(core.got) != 2 {
t.Errorf("core saw %d attempts, want 2 (a failed ingest is retried)", len(core.got))
}
}
// Core without an email block ⇒ stop, don't hammer the socket.
func TestPollStopsWhenCoreRefusesMail(t *testing.T) {
f := &fakeIMAP{uids: []uint32{1, 2}, msgs: map[uint32]string{1: mail("a", "b"), 2: mail("c", "d")}}
core := &fakeCore{err: fmt.Errorf("call: %w", ipc.ErrUnknownMethod)}
r := newTestReader(t, f, core, "")
r.pollOnce(context.Background(), "secret")
if !r.disabled {
t.Error("ErrUnknownMethod must disable the reader")
}
if len(core.got) != 1 {
t.Errorf("core saw %d messages, want 1 — stop at the first refusal", len(core.got))
}
}
func TestSeenStatePersists(t *testing.T) {
path := filepath.Join(t.TempDir(), "state", "seen.json")
f := &fakeIMAP{uids: []uint32{9}, msgs: map[uint32]string{9: mail("Счёт", "текст")}}
core := &fakeCore{}
r := newTestReader(t, f, core, path)
r.pollOnce(context.Background(), "secret")
fi, err := os.Stat(path)
if err != nil {
t.Fatalf("state file: %v", err)
}
// A list of message ids from his mailbox is metadata about his mail.
if perm := fi.Mode().Perm(); perm != 0o600 {
t.Errorf("state file mode = %v, want 0600", perm)
}
// A fresh reader with the same state file must not re-read the message.
core2 := &fakeCore{}
r2 := newTestReader(t, f, core2, path)
if err := r2.state.load(); err != nil {
t.Fatalf("load: %v", err)
}
r2.pollOnce(context.Background(), "secret")
if len(core2.got) != 0 {
t.Errorf("after a restart core saw %d messages, want 0", len(core2.got))
}
}
func TestSeenStateHighWaterMark(t *testing.T) {
s := newSeenState("")
s.mark(1)
s.mark(3)
s.mark(2)
if s.high != 3 {
t.Errorf("high = %d, want 3 (contiguous run collapses)", s.high)
}
if len(s.set) != 0 {
t.Errorf("explicit set = %v, want empty", s.set)
}
if !s.seen(2) || s.seen(4) {
t.Errorf("seen(2)=%v seen(4)=%v", s.seen(2), s.seen(4))
}
}
func TestSeenStateCorruptFileIsNotFatal(t *testing.T) {
path := filepath.Join(t.TempDir(), "seen.json")
if err := os.WriteFile(path, []byte("{not json"), 0o600); err != nil {
t.Fatal(err)
}
s := newSeenState(path)
if err := s.load(); err == nil {
t.Error("a corrupt state file should report an error the caller logs")
}
if s.seen(1) {
t.Error("a corrupt state file must leave an empty seen-set, not a poisoned one")
}
}
// Off unless configured, and the credential is never a flag value.
func TestRunRequiresConfig(t *testing.T) {
if err := run([]string{}); err == nil {
t.Error("no -socket must be an error")
}
if err := run([]string{"-socket", "/tmp/nope.sock"}); err == nil {
t.Error("no mailbox configuration must be an error, not a default mailbox")
}
// There is no -password flag at all: only -password-file.
if err := run([]string{"-socket", "/x", "-imap", "h", "-user", "u", "-password", "p"}); err == nil ||
!strings.Contains(err.Error(), "flag provided but not defined") {
t.Errorf("a -password flag must not exist; err = %v", err)
}
}
func TestRunRejectsEmptyPasswordFile(t *testing.T) {
path := filepath.Join(t.TempDir(), "pass")
if err := os.WriteFile(path, []byte(" \n"), 0o600); err != nil {
t.Fatal(err)
}
err := run([]string{"-socket", "/x/y.sock", "-imap", "h", "-user", "u", "-password-file", path})
if err == nil || !strings.Contains(err.Error(), "empty") {
t.Errorf("an empty password file must be refused before dialling; err = %v", err)
}
}
+124 -3
View File
@@ -9,6 +9,12 @@
// Two sources, each its own provenance (the loop's rules trust source):
// - netdata → poll:netdata resource alarms (disk/mem/cert/temp)
// - kuma → poll:uptimekuma service up/down (the source of truth for it)
// - zenmoney → poll:zenmoney spending/income totals (Vikunja #125)
//
// The zenmoney source is why the token lives HERE and not in core: the poller
// already owns every other third-party credential, it holds no store key, and
// core never needs to know an account exists to answer a question about a fact
// the poller wrote. It is off unless -zenmoney-token-file is given.
//
// Netdata needs no auth over the wg-fronted net. Kuma's /metrics needs an API
// key (basic-auth); without -kuma the whole kuma path is skipped (netdata-only
@@ -37,6 +43,7 @@ import (
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/zenmoney"
)
func main() {
@@ -52,6 +59,9 @@ func run(args []string) error {
netdataURL := fs.String("netdata", "http://127.0.0.1:19999", "netdata base URL ('' to disable)")
kumaURL := fs.String("kuma", "", "uptime-kuma metrics URL, e.g. http://127.0.0.1:3001/metrics ('' to disable)")
kumaKey := fs.String("kuma-key", "", "uptime-kuma API key (basic-auth username)")
zenTokenFile := fs.String("zenmoney-token-file", "", "file holding the zenmoney API token ('' disables money tracking)")
zenURL := fs.String("zenmoney-url", zenmoney.DefaultBaseURL, "zenmoney API base URL (tests/self-hosted proxies)")
zenInterval := fs.Duration("zenmoney-interval", time.Hour, "how often to read zenmoney (money does not move every minute)")
wgIface := fs.String("wg", "", "wireguard interface for the presence signal, e.g. wg0 or 'all' ('' to disable)")
wgCmd := fs.String("wg-cmd", "wg", "wg binary (use e.g. 'sudo wg' if the poller lacks CAP_NET_ADMIN)")
interval := fs.Duration("interval", 60*time.Second, "poll cadence")
@@ -62,8 +72,24 @@ func run(args []string) error {
if *socket == "" {
return fmt.Errorf("-socket is required")
}
if *netdataURL == "" && *kumaURL == "" && *wgIface == "" {
return fmt.Errorf("nothing to poll: set -netdata, -kuma and/or -wg")
if *netdataURL == "" && *kumaURL == "" && *wgIface == "" && *zenTokenFile == "" {
return fmt.Errorf("nothing to poll: set -netdata, -kuma, -wg and/or -zenmoney-token-file")
}
// The token is read from a file, never taken as a flag value: an argv token
// is visible in `ps` to every user on the box and lands in the compose file
// and the shell history. Read once at start — a rotated token means a
// restart, which is cheaper than re-reading his credential every hour.
var zen *zenmoney.Client
if *zenTokenFile != "" {
raw, err := os.ReadFile(*zenTokenFile)
if err != nil {
return fmt.Errorf("read zenmoney token: %w", err)
}
zen, err = zenmoney.New(strings.TrimSpace(string(raw)), *zenURL, *timeout*3)
if err != nil {
return err
}
}
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
@@ -83,9 +109,13 @@ func run(args []string) error {
kumaKey: *kumaKey,
wgIface: *wgIface,
wgCmd: *wgCmd,
zen: zen,
zenEvery: *zenInterval,
}
log.Printf("mavpoll: polling every %s (netdata=%q kuma=%q wg=%q)", *interval, *netdataURL, *kumaURL, *wgIface)
// The token is never logged, not even its length.
log.Printf("mavpoll: polling every %s (netdata=%q kuma=%q wg=%q zenmoney=%v every %s)",
*interval, *netdataURL, *kumaURL, *wgIface, zen != nil, *zenInterval)
p.pollOnce(ctx) // fire immediately; don't idle a full interval on start
t := time.NewTicker(*interval)
defer t.Stop()
@@ -108,6 +138,12 @@ type poller struct {
kumaKey string
wgIface string
wgCmd string
// zen is nil unless a token file was configured — money tracking is a
// capability, off by default like weather and telegram.
zen *zenmoney.Client
zenEvery time.Duration
zenLast time.Time
}
// pollOnce — one sweep of both sources. A failure in one source logs and does
@@ -129,6 +165,66 @@ func (p *poller) pollOnce(ctx context.Context) {
log.Printf("mavpoll: wg: %v", err)
}
}
// Money on its own, much slower cadence: a bank feed that updates hourly
// polled every minute is 60 pointless reads of his financial history.
if p.zen != nil && now.Sub(p.zenLast) >= p.zenEvery {
p.zenLast = now
if err := p.pollZenmoney(ctx, now); err != nil {
log.Printf("mavpoll: zenmoney: %v", err)
}
}
}
// ---- zenmoney: spending/income totals → money facts ------------------------
// pollZenmoney reads today's and this month's totals and writes them as
// facts(kind=env, source=poll:zenmoney) (Vikunja #125).
//
// Two properties this function exists to hold:
//
// - An empty or failed read writes NOTHING. zenmoney.Summary.Value() refuses
// to encode a summary built from zero transactions, so a poller that cannot
// reach the API leaves the last good fact in place rather than overwriting
// it with a zero Maven would then recite as fact.
// - Nothing about the money leaves the box except the diff request itself, to
// the service that already holds his bank sessions. The totals are written
// to the store and read back only when he asks; they are never search input
// and no tick rule fires on them.
//
// Both windows are read from one diff call each. Two calls an hour against an
// API whose whole job is this is not worth caching.
// moneyWindow — one fact key and the period it covers.
type moneyWindow struct {
key string
from, to time.Time
}
func (p *poller) pollZenmoney(ctx context.Context, now time.Time) error {
dFrom, dTo := zenmoney.DayWindow(now)
mFrom, mTo := zenmoney.MonthWindow(now)
windows := []moneyWindow{
{zenmoney.KeySpentToday, dFrom, dTo},
{zenmoney.KeySpentMonth, mFrom, mTo},
}
var firstErr error
for _, w := range windows {
sum, err := p.zen.Since(ctx, w.from, w.to)
if err != nil {
if firstErr == nil {
firstErr = err
}
continue
}
val, ok := sum.Value()
if !ok {
// Nothing read. Silence, not a zero.
continue
}
if err := p.writeIfChangedRaw(ctx, w.key, zenmoney.Source, val, now); err != nil && firstErr == nil {
firstErr = err
}
}
return firstErr
}
// ---- wireguard: latest handshake → presence signal -------------------------
@@ -301,6 +397,31 @@ func (p *poller) writeIfChanged(ctx context.Context, key, source, val string, no
return nil
}
// writeIfChangedRaw is writeIfChanged for values that are already JSON (the
// money facts store an object, not a string). Kept separate rather than
// generalising writeIfChanged, because the string-valued env facts encoding
// their own value is the convention the rules rely on.
//
// The log line names the key and the source, never the figures: mavpoll's log
// is not the place his spending ends up.
func (p *poller) writeIfChangedRaw(ctx context.Context, key, source, jsonVal string, now time.Time) error {
prev, err := p.core.LatestFactBySource(ctx, key, source)
switch {
case err == nil && prev.Value == jsonVal:
return nil
case err != nil && err != ipc.ErrNoFact && !isNoFact(err):
return fmt.Errorf("read %s: %w", key, err)
}
if _, err := p.core.WriteFact(ctx, ipc.WriteFactReq{
Ts: now, Kind: "env", Key: key, Value: jsonVal,
Source: source, Confidence: 1.0,
}); err != nil {
return fmt.Errorf("write %s: %w", key, err)
}
log.Printf("mavpoll: %s updated (%s)", key, source)
return nil
}
// isNoFact — ErrNoFact rehydrated over the wire is wrapped (fmt.Errorf %w), so
// errors.Is is the right check; keep a helper so the switch above reads clean.
func isNoFact(err error) bool {
+126
View File
@@ -1,8 +1,17 @@
package main
import (
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"os"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/zenmoney"
)
func TestMaxSeverity(t *testing.T) {
@@ -62,3 +71,120 @@ func TestParseMaxHandshake(t *testing.T) {
}
}
}
// ---- zenmoney (Vikunja #125) ----------------------------------------------
// factCore records the facts the poller wrote and answers "no fact yet".
type factCore struct {
ipc.UnimplementedCoreAPI
written []ipc.WriteFactReq
prev map[string]string
}
func (c *factCore) LatestFactBySource(_ context.Context, key, source string) (ipc.Fact, error) {
if v, ok := c.prev[key+"|"+source]; ok {
return ipc.Fact{Key: key, Source: source, Value: v}, nil
}
return ipc.Fact{}, ipc.ErrNoFact
}
func (c *factCore) WriteFact(_ context.Context, req ipc.WriteFactReq) (int64, error) {
c.written = append(c.written, req)
return int64(len(c.written)), nil
}
func zenFixtureServer(t *testing.T, body []byte, status int) *httptest.Server {
t.Helper()
return httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if status != http.StatusOK {
w.WriteHeader(status)
return
}
w.Write(body)
}))
}
func TestPollZenmoneyWritesMoneyFacts(t *testing.T) {
body, err := os.ReadFile("../../internal/zenmoney/testdata/diff.json")
if err != nil {
t.Fatal(err)
}
srv := zenFixtureServer(t, body, http.StatusOK)
defer srv.Close()
zen, err := zenmoney.New("tok", srv.URL, time.Second)
if err != nil {
t.Fatal(err)
}
core := &factCore{}
p := &poller{core: core, zen: zen}
now := time.Date(2026, 8, 1, 21, 0, 0, 0, time.UTC)
if err := p.pollZenmoney(context.Background(), now); err != nil {
t.Fatal(err)
}
if len(core.written) != 2 {
t.Fatalf("wrote %d facts, want today + month", len(core.written))
}
for _, f := range core.written {
if f.Kind != "env" || f.Source != zenmoney.Source {
t.Errorf("fact = %+v, want kind=env source=%s", f, zenmoney.Source)
}
if _, err := zenmoney.ParseFactValue(f.Value); err != nil {
t.Errorf("fact value %q does not decode: %v", f.Value, err)
}
}
}
// A read that returns nothing for the window writes NOTHING. Silence, not a
// zero: an invented 0 would be recited back to him as fact.
func TestPollZenmoneyWritesNothingWhenEmpty(t *testing.T) {
srv := zenFixtureServer(t, []byte(`{"serverTimestamp":1,"instrument":[],"transaction":[]}`), http.StatusOK)
defer srv.Close()
zen, _ := zenmoney.New("tok", srv.URL, time.Second)
core := &factCore{}
p := &poller{core: core, zen: zen}
if err := p.pollZenmoney(context.Background(), time.Now()); err != nil {
t.Fatal(err)
}
if len(core.written) != 0 {
t.Errorf("wrote %+v, want no fact at all", core.written)
}
}
// An API failure must not overwrite the last good total either.
func TestPollZenmoneyFailureWritesNothing(t *testing.T) {
srv := zenFixtureServer(t, nil, http.StatusUnauthorized)
defer srv.Close()
zen, _ := zenmoney.New("bad", srv.URL, time.Second)
core := &factCore{}
p := &poller{core: core, zen: zen}
if err := p.pollZenmoney(context.Background(), time.Now()); err == nil {
t.Error("want the 401 reported")
}
if len(core.written) != 0 {
t.Errorf("wrote %+v on a failed read", core.written)
}
}
// Unchanged totals do not churn the facts table.
func TestWriteIfChangedRawSkipsUnchanged(t *testing.T) {
core := &factCore{prev: map[string]string{
zenmoney.KeySpentToday + "|" + zenmoney.Source: `{"count":1}`,
}}
p := &poller{core: core}
if err := p.writeIfChangedRaw(context.Background(), zenmoney.KeySpentToday, zenmoney.Source, `{"count":1}`, time.Now()); err != nil {
t.Fatal(err)
}
if len(core.written) != 0 {
t.Errorf("wrote %+v for an unchanged value", core.written)
}
}
// Money tracking is off unless configured: no token file, no zenmoney client,
// and the poller still refuses to start with nothing at all to poll.
func TestRunRequiresSomethingToPoll(t *testing.T) {
err := run([]string{"-socket", "/tmp/nope.sock", "-netdata", "", "-kuma", "", "-wg", ""})
if err == nil || !strings.Contains(err.Error(), "nothing to poll") {
t.Errorf("err = %v, want a 'nothing to poll' refusal", err)
}
}
+323
View File
@@ -0,0 +1,323 @@
package main
// Golden-audio STT tests (Vikunja #288).
//
// These push real audio through the real whisper.cpp binding, so a bad model
// path, a wrong language hint, a broken resample or a regressed silence gate
// is caught by `make test` rather than by the owner talking to a daemon that
// mishears him.
//
// The fixtures are piper-synthesised, not recorded — see
// scripts/gen-stt-fixtures.sh. Nothing of the owner's voice is committed, and
// any fixture can be rebuilt from the script plus a voice model.
//
// Matching is deliberately tolerant. Golden transcripts are model-dependent:
// swapping ggml-small for a different whisper build moves punctuation, casing
// and the odd word ending, and an exact-string assertion would turn every
// model swap into a fixture rewrite. Each case therefore asserts two things —
// the words that carry the intent are present, and the word error rate
// against the reference stays under a per-case ceiling.
import (
"context"
"encoding/json"
"os"
"path/filepath"
"strings"
"testing"
"unicode"
"github.com/kami/maven/internal/audio"
"github.com/kami/maven/internal/worker"
)
// goldenModelPath — the whisper model the golden tests run against. Same file
// the Makefile's run-stt target uses. Overridable so a box that keeps its
// models elsewhere can still run these.
func goldenModelPath() string {
if p := os.Getenv("MAVEN_WHISPER_MODEL"); p != "" {
return p
}
return filepath.Join("..", "..", "models", "stt", "ggml-small.bin")
}
type goldenCase struct {
Name string `json:"name"`
WAV string `json:"wav"`
Lang string `json:"lang"`
Text string `json:"text"`
Keywords []string `json:"keywords"`
MaxWER float64 `json:"max_wer"`
}
type goldenManifest struct {
Cases []goldenCase `json:"cases"`
}
func loadGoldenManifest(t *testing.T) goldenManifest {
t.Helper()
raw, err := os.ReadFile(filepath.Join("testdata", "golden_v1.json"))
if err != nil {
t.Fatalf("read golden manifest: %v", err)
}
var m goldenManifest
if err := json.Unmarshal(raw, &m); err != nil {
t.Fatalf("parse golden manifest: %v", err)
}
if len(m.Cases) == 0 {
t.Fatal("golden manifest has no cases")
}
return m
}
// normalizeTranscript lowercases, drops punctuation, folds the Russian ё onto
// е (whisper is inconsistent about it and the router does not care), and
// collapses whitespace. Everything the comparison does happens on this form.
func normalizeTranscript(s string) []string {
var b strings.Builder
for _, r := range strings.ToLower(s) {
switch {
case r == 'ё':
b.WriteRune('е')
case unicode.IsLetter(r) || unicode.IsDigit(r):
b.WriteRune(r)
default:
b.WriteRune(' ')
}
}
return strings.Fields(b.String())
}
// wordErrorRate is the Levenshtein distance between two word sequences,
// divided by the length of the reference. 0 means identical; it can exceed 1
// when the hypothesis is much longer than the reference.
func wordErrorRate(ref, hyp []string) float64 {
if len(ref) == 0 {
if len(hyp) == 0 {
return 0
}
return 1
}
prev := make([]int, len(hyp)+1)
cur := make([]int, len(hyp)+1)
for j := range prev {
prev[j] = j
}
for i := 1; i <= len(ref); i++ {
cur[0] = i
for j := 1; j <= len(hyp); j++ {
cost := 1
if ref[i-1] == hyp[j-1] {
cost = 0
}
cur[j] = min(prev[j]+1, min(cur[j-1]+1, prev[j-1]+cost))
}
prev, cur = cur, prev
}
return float64(prev[len(hyp)]) / float64(len(ref))
}
// missingKeywords returns the keywords absent from the hypothesis. A keyword
// matches on prefix, so a different case ending ("воды" vs "воду") does not
// fail the assertion — the router's stage-0 grammar is stem-shaped too.
func missingKeywords(keywords []string, hyp []string) []string {
var missing []string
for _, kw := range keywords {
want := normalizeTranscript(kw)
if len(want) == 0 {
continue
}
if !containsSeq(hyp, want) {
missing = append(missing, kw)
}
}
return missing
}
func containsSeq(hyp, want []string) bool {
for i := 0; i+len(want) <= len(hyp); i++ {
ok := true
for j, w := range want {
// Prefix match, so inflection differences pass but
// distinct words do not.
if !looseWordMatch(hyp[i+j], w) {
ok = false
break
}
}
if ok {
return true
}
}
return false
}
func looseWordMatch(got, want string) bool {
if got == want {
return true
}
g, w := []rune(got), []rune(want)
n := len(w) - 1
if len(w) > 6 {
n = len(w) - 2
}
// Words of three runes or fewer have no room for a safe prefix: require
// an exact match rather than letting "час" pass for "часть".
if n < 3 || len(g) < n {
return false
}
return string(g[:n]) == string(w[:n])
}
// --- the model-backed test -------------------------------------------------
func TestGoldenAudioTranscription(t *testing.T) {
m := loadGoldenManifest(t)
model := goldenModelPath()
if _, err := os.Stat(model); err != nil {
t.Skipf("whisper model %s absent (%v) — set MAVEN_WHISPER_MODEL or see AGENTS.md", model, err)
}
// Same gate thresholds as mavsttd's defaults, so a regression in the
// silence gate shows up here as an empty transcript.
h, err := newWhisperHandler(model, 300, 0.01)
if err != nil {
t.Fatalf("load whisper model %s: %v", model, err)
}
defer h.Close()
for _, c := range m.Cases {
t.Run(c.Name, func(t *testing.T) {
path := filepath.Join("testdata", c.WAV)
raw, err := os.ReadFile(path)
if err != nil {
t.Skipf("fixture %s absent (%v) — run scripts/gen-stt-fixtures.sh", path, err)
}
format, pcm, err := audio.PCMFromWAV(raw)
if err != nil {
t.Fatalf("%s is not canonical 16k mono PCM: %v", path, err)
}
resp, err := h.Transcribe(context.Background(), worker.TranscribeReq{
Audio: audio.Audio{Format: format, Bytes: pcm},
Lang: c.Lang,
})
if err != nil {
t.Fatalf("transcribe %s: %v", c.WAV, err)
}
t.Logf("%s → %q (confidence %.3f)", c.WAV, resp.Text, resp.Confidence)
if strings.TrimSpace(resp.Text) == "" {
t.Fatalf("%s transcribed to empty text — the silence gate ate real speech", c.WAV)
}
if resp.Confidence <= 0 {
t.Errorf("%s: confidence %v, want > 0", c.WAV, resp.Confidence)
}
hyp := normalizeTranscript(resp.Text)
ref := normalizeTranscript(c.Text)
if missing := missingKeywords(c.Keywords, hyp); len(missing) > 0 {
t.Errorf("%s: missing keywords %v in %q", c.WAV, missing, resp.Text)
}
if wer := wordErrorRate(ref, hyp); wer > c.MaxWER {
t.Errorf("%s: WER %.2f > %.2f\n want: %q\n got: %q", c.WAV, wer, c.MaxWER, c.Text, resp.Text)
}
})
}
}
// TestGoldenFixturesAreCanonical checks the committed audio without needing a
// model, so a fixture regenerated at the wrong sample rate fails on every box.
func TestGoldenFixturesAreCanonical(t *testing.T) {
m := loadGoldenManifest(t)
for _, c := range m.Cases {
path := filepath.Join("testdata", c.WAV)
raw, err := os.ReadFile(path)
if err != nil {
t.Errorf("fixture %s missing: %v", path, err)
continue
}
format, pcm, err := audio.PCMFromWAV(raw)
if err != nil {
t.Errorf("%s: %v", path, err)
continue
}
if !format.IsValid() {
t.Errorf("%s: format %+v is not canonical", path, format)
}
a := audio.Audio{Format: format, Bytes: pcm}
if d := a.Duration(); d < 0.5 || d > 10 {
t.Errorf("%s: duration %.2fs outside the sane 0.510s fixture range", path, d)
}
// The fixture must clear mavsttd's own silence gate, otherwise the
// model test below would be asserting on a gated empty string.
if reason := gateReason(pcmToF32(pcm), whisperSampleRate, 300, 0.01); reason != "" {
t.Errorf("%s: would be gated as %s", path, reason)
}
if len(c.Keywords) == 0 {
t.Errorf("%s: manifest case has no keywords", c.Name)
}
if c.MaxWER <= 0 || c.MaxWER > 1 {
t.Errorf("%s: max_wer %v outside (0,1]", c.Name, c.MaxWER)
}
}
}
func pcmToF32(b []byte) []float32 {
out := make([]float32, len(b)/2)
for i := range out {
s := int16(b[i*2]) | int16(b[i*2+1])<<8
out[i] = float32(s) / 32768.0
}
return out
}
// --- matcher unit tests (no model, no fixtures) ----------------------------
func TestNormalizeTranscript(t *testing.T) {
got := normalizeTranscript(" Ещё, Раз... ")
want := []string{"еще", "раз"}
if len(got) != len(want) || got[0] != want[0] || got[1] != want[1] {
t.Fatalf("normalizeTranscript = %v, want %v", got, want)
}
}
func TestWordErrorRate(t *testing.T) {
cases := []struct {
name string
ref, hyp string
want float64
}{
{"identical", "напомни мне через час", "Напомни мне через час.", 0},
{"one substitution", "напомни мне через час", "напомни мне через день", 0.25},
{"one deletion", "напомни мне через час", "напомни мне час", 0.25},
{"empty hypothesis", "напомни мне", "", 1},
{"both empty", "", "", 0},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
got := wordErrorRate(normalizeTranscript(c.ref), normalizeTranscript(c.hyp))
if got != c.want {
t.Fatalf("WER = %v, want %v", got, c.want)
}
})
}
}
func TestMissingKeywords(t *testing.T) {
hyp := normalizeTranscript("Отметь, что я выпил воду.")
if got := missingKeywords([]string{"воды", "отметь"}, hyp); len(got) != 0 {
t.Fatalf("missingKeywords = %v, want none (inflection must not fail the match)", got)
}
if got := missingKeywords([]string{"календарю"}, hyp); len(got) != 1 {
t.Fatalf("missingKeywords = %v, want the absent keyword reported", got)
}
// A short word must match exactly — no 4-rune prefix shortcut that would
// let "час" pass for "часть".
hyp2 := normalizeTranscript("через час")
if got := missingKeywords([]string{"часть"}, hyp2); len(got) != 1 {
t.Fatalf("missingKeywords = %v, want %q reported missing", got, "часть")
}
}
BIN
View File
Binary file not shown.
+37
View File
@@ -0,0 +1,37 @@
{
"note": "Golden STT fixtures. Audio is piper-synthesised, not recorded — see scripts/gen-stt-fixtures.sh. Regenerate with that script; do not hand-edit `wav`.",
"cases": [
{
"name": "ru_reminder",
"wav": "ru_reminder.wav",
"lang": "ru",
"text": "напомни мне через час позвонить маме",
"keywords": ["напомни", "час", "позвонить"],
"max_wer": 0.34
},
{
"name": "ru_fact",
"wav": "ru_fact.wav",
"lang": "ru",
"text": "отметь что я выпил воды",
"keywords": ["отметь", "воды"],
"max_wer": 0.34
},
{
"name": "ru_query",
"wav": "ru_query.wav",
"lang": "ru",
"text": "что у меня сегодня по календарю",
"keywords": ["сегодня", "календарю"],
"max_wer": 0.34
},
{
"name": "en_act",
"wav": "en_act.wav",
"lang": "en",
"text": "restart the web server and check the disk space",
"keywords": ["restart", "server", "disk"],
"max_wer": 0.34
}
]
}
Binary file not shown.
Binary file not shown.
Binary file not shown.
+204
View File
@@ -0,0 +1,204 @@
// Command mavupdate deploys a new build of Maven to the box she runs on, with
// an automatic rollback when the new build does not come up (Vikunja #249).
//
// It is a CLI on purpose, and it is the ONLY trigger for the update path.
//
// The obvious design — an IPC method plus a button on the web UI behind the
// step-up passkey gate, the way /tools works — was considered and refused. A
// step-up gate protects against the wrong person clicking; it does not change
// the fact that anything reachable over the network becomes, in the event of a
// mavweb bug, a remote arbitrary-code path with a build system attached. An
// update needs shell access on the host, which is a strictly higher bar than
// the gate that guards the tool allowlist. That is deliberate and it is the
// reason there is no MethodApplyUpdate anywhere in internal/ipc.
//
// Consequently: mavend does not import internal/update, nothing runs on a timer,
// nothing checks a release server, and no act, intent, tool or LLM output can
// reach any of this. She cannot update herself. She can be updated, by him.
//
// mavupdate -config deploy/mavend.json list # snapshots available to roll back to
// mavupdate -config deploy/mavend.json verify # make build + make test, deploys nothing
// mavupdate -config deploy/mavend.json apply -yes # the whole thing
// mavupdate -config deploy/mavend.json rollback [id] # restore + restart (default: newest)
package main
import (
"context"
"errors"
"flag"
"fmt"
"os"
"os/signal"
"syscall"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/update"
)
func main() {
cfgPath := flag.String("config", "deploy/mavend.json", "path to mavend.json (the update block is read from it)")
yes := flag.Bool("yes", false, "required by `apply` and `rollback`: yes, restart the daemon")
flag.Usage = usage
flag.Parse()
// The stdlib flag package stops parsing at the first non-flag argument, so a
// `-yes` written after the subcommand (which is how anyone would type it, and
// how the usage text shows it) lands in Args instead of the flag. Pick it out
// by hand rather than silently treating "apply -yes" as an unconfirmed apply.
var args []string
for _, a := range flag.Args() {
if a == "-yes" || a == "--yes" {
*yes = true
continue
}
args = append(args, a)
}
if len(args) == 0 {
usage()
os.Exit(2)
}
cfg, err := config.Load(*cfgPath)
if err != nil {
die("config: %v", err)
}
if cfg.Update == nil {
die("no `update` block in %s — the update capability is off unless configured.\nSee the package comment in internal/update for what it does and does not do.", *cfgPath)
}
logf := func(format string, a ...any) {
fmt.Fprintf(os.Stderr, "%s %s\n", time.Now().Format("15:04:05"), fmt.Sprintf(format, a...))
}
u, err := update.New(*cfg.Update, update.WithLogger(logf))
if err != nil {
die("%v", err)
}
// Ctrl-C cancels the build or the health wait. It cannot cancel a rollback
// midway into leaving the box in an unknown state, because the rollback runs
// on its own context — see cmdApply.
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
defer stop()
switch args[0] {
case "list":
cmdList(u)
case "verify":
cmdVerify(ctx, u)
case "apply":
if !*yes {
die("apply restarts mavend and can roll her back. Re-run with -yes if that is what you want.")
}
cmdApply(ctx, u)
case "rollback":
if !*yes {
die("rollback restores the previous artifacts and restarts mavend. Re-run with -yes.")
}
id := ""
if len(args) > 1 {
id = args[1]
}
cmdRollback(ctx, u, id)
default:
usage()
os.Exit(2)
}
}
func cmdList(u *update.Updater) {
snaps, err := u.Snapshots()
if err != nil {
die("snapshots: %v", err)
}
if len(snaps) == 0 {
fmt.Println("no snapshots yet — the first `apply` takes one before it builds anything")
return
}
fmt.Printf("%-18s %-12s %s\n", "SNAPSHOT", "COMMIT", "FILES")
for _, s := range snaps {
commit := s.Commit
if len(commit) > 12 {
commit = commit[:12]
}
if commit == "" {
commit = "-"
}
fmt.Printf("%-18s %-12s %d\n", s.ID, commit, len(s.Files))
}
fmt.Printf("\nrollback to the newest with: mavupdate rollback -yes\n")
}
func cmdVerify(ctx context.Context, u *update.Updater) {
steps, err := u.Verify(ctx)
report(steps)
if err != nil {
die("%v", err)
}
fmt.Println("verified: the tree builds and passes its own tests. Nothing was deployed — run `apply -yes` for that.")
}
func cmdApply(ctx context.Context, u *update.Updater) {
res, err := u.Apply(ctx)
report(res.Steps)
summarize(res)
switch {
case err == nil:
fmt.Println("\nupdate committed: she answers on the new build.")
case errors.Is(err, update.ErrRollbackFailed):
die("\n%v\n\nSHE IS PROBABLY DOWN. The previous artifacts are in the snapshot dir; copy them\nover the install dir and restart by hand.", err)
case errors.Is(err, update.ErrRolledBack):
die("\n%v\n\nShe is answering again on the previous build. Nothing was lost; fix the change and retry.", err)
default:
die("\n%v", err)
}
}
func cmdRollback(ctx context.Context, u *update.Updater, id string) {
res, err := u.Rollback(ctx, id)
report(res.Steps)
summarize(res)
if err != nil && !errors.Is(err, update.ErrRolledBack) {
die("\n%v", err)
}
fmt.Printf("\nrolled back to %s; she answers on it.\n", res.SnapshotID)
}
func report(steps []update.Step) {
for _, s := range steps {
status := "ok"
if s.Err != nil {
status = "FAILED: " + s.Err.Error()
}
fmt.Printf(" %-8s %-8s %s\n", s.Name, s.Took.Round(time.Second), status)
if s.Output != "" {
fmt.Printf("---- %s output ----\n%s\n-------------------\n", s.Name, s.Output)
}
}
}
func summarize(res update.Result) {
fmt.Printf("\nverified=%v snapshot=%s installed=%d restarted=%v healthy=%v rolled_back=%v rollback_healthy=%v took=%s\n",
res.Verified, res.SnapshotID, len(res.Installed), res.Restarted, res.Healthy, res.RolledBack, res.RollbackHealthy, res.Took.Round(time.Second))
}
func usage() {
fmt.Fprint(os.Stderr, `mavupdate deploy a new build of Maven, with rollback.
mavupdate [-config path] list
mavupdate [-config path] verify
mavupdate [-config path] apply -yes
mavupdate [-config path] rollback [snapshot-id] -yes
apply is: health-check the running daemon, snapshot the deployed artifacts,
make build, make test, install, restart, health-check and restore the
snapshot if any of that fails. It never fetches code and never runs by itself.
`)
flag.PrintDefaults()
}
func die(format string, a ...any) {
fmt.Fprintf(os.Stderr, format+"\n", a...)
os.Exit(1)
}
+33 -92
View File
@@ -12,8 +12,15 @@
// (30ms frames, 16kHz PCM) matches silero-vad's input interface exactly, so
// swapping energy-threshold for ONNX-inference is a local change in vad.go.
//
// While a reply is playing the capture side is muted (half-duplex): without
// it, Maven's own voice comes back in through the mic and she answers
// herself. -barge-in punches one hole in that gate — sustained energy well
// above the speaker's leak level cuts playback so he can talk over her. It is
// off by default because the threshold is room-specific; see playback.go.
//
// usage:
// mavwaked # default ALSA device, 127.0.0.1:9100
// mavwaked -barge-in # let him interrupt her mid-reply
// mavwaked -device hw:1,0 -addr 10.42.0.1:9100
// mavwaked -test file.wav # read from file, no arecord
package main
@@ -60,6 +67,9 @@ func run(args []string) error {
silenceMs := flag.Int("silence-ms", defaultSilenceMs, "silence ms to end utterance")
maxMs := flag.Int("max-ms", defaultMaxMs, "max utterance ms")
testFile := flag.String("test", "", "read PCM from file instead of arecord (testing only)")
bargeIn := flag.Bool("barge-in", false, "cut Maven off when he talks over her (needs a room-tuned -barge-in-rms)")
bargeRMS := flag.Int("barge-in-rms", defaultBargeRMS, "RMS x10000 a frame must clear to count as barge-in")
bargeFrames := flag.Int("barge-in-frames", defaultBargeFrames, "consecutive frames over -barge-in-rms before playback is cut")
flag.CommandLine.Parse(args)
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM, syscall.SIGHUP)
@@ -117,14 +127,21 @@ func run(args []string) error {
defer src.Close()
return captureLoop(ctx, src, vad, vc, *lang)
var barge bargeInConfig
if *bargeIn {
barge = bargeInConfig{RMS: float64(*bargeRMS) / 10000.0, Frames: *bargeFrames}
log.Printf("mavwaked: barge-in on (rms %.4f x %d frames)", barge.RMS, barge.Frames)
}
sess := newSession(vad, newAplayPlayer(), &voiceSender{vc: vc}, *lang, barge)
return captureLoop(ctx, src, sess)
}
// captureLoop reads PCM from src, runs VAD, and sends complete utterances to
// the voice server. Returns when ctx is done or src is exhausted.
func captureLoop(ctx context.Context, src io.Reader, vad *VAD, vc *voice.Client, lang string) error {
// captureLoop reads PCM from src and hands whole frames to the session.
// Returns when ctx is done or src is exhausted.
func captureLoop(ctx context.Context, src io.Reader, sess *session) error {
br := bufio.NewReaderSize(src, defaultReadSize)
frameBytes := vad.FrameSamples() * 2 // 480 samples × 2 bytes = 960 bytes per 30ms
frameBytes := sess.vad.FrameSamples() * 2 // 480 samples × 2 bytes = 960 bytes per 30ms
log.Printf("mavwaked: capture loop starting (frame=%d bytes, %dms)",
frameBytes, defaultFrameMs)
@@ -147,7 +164,7 @@ func captureLoop(ctx context.Context, src io.Reader, vad *VAD, vc *voice.Client,
// Flush partial frame.
partial = append(partial, buf[:n]...)
if len(partial) >= frameBytes {
if err := processFrame(partial[:frameBytes], vad, vc, lang); err != nil {
if err := sess.feed(ctx, partial[:frameBytes]); err != nil {
log.Printf("mavwaked: process frame: %v", err)
}
partial = partial[frameBytes:]
@@ -165,107 +182,31 @@ func captureLoop(ctx context.Context, src io.Reader, vad *VAD, vc *voice.Client,
partial = nil
}
if err := processFrame(full, vad, vc, lang); err != nil {
if err := sess.feed(ctx, full); err != nil {
log.Printf("mavwaked: process frame: %v", err)
}
}
}
// processFrame feeds one 30ms PCM frame to the VAD and sends any completed
// utterance to the voice server.
func processFrame(frame []byte, vad *VAD, vc *voice.Client, lang string) error {
samples := PCMToI16(frame)
utt, state := vad.Feed(samples)
// voiceSender is the production utteranceSender: one PushToTalk round-trip
// over the voice wire. SurfaceVoice (not the default SurfacePCClient that
// c.PushToTalk uses) caps everything at L0, which is what makes an accidental
// VAD trigger safe.
type voiceSender struct{ vc *voice.Client }
if state == StateSpeech {
// Speech is in progress; nothing to send yet.
return nil
}
if utt.Bytes == nil {
// Still in silence, or short speech that didn't trigger.
return nil
}
// We have a complete utterance — send it to the voice server.
return sendUtterance(context.Background(), utt, vc, lang)
}
// sendUtterance sends audio to the voice server and plays the reply.
func sendUtterance(ctx context.Context, utt audio.Audio, vc *voice.Client, lang string) error {
dur := utt.Duration()
log.Printf("mavwaked: utterance complete (%.2fs, %d bytes), sending...",
dur, len(utt.Bytes))
// Use SendRequest directly so we can set SurfaceVoice instead of the
// default SurfacePCClient that c.PushToTalk uses.
func (s *voiceSender) Send(ctx context.Context, utt audio.Audio, lang string) (audio.Audio, error) {
var resp voice.PushToTalkResp
err := vc.SendRequest(ctx, voice.MethodPushToTalk, voice.PushToTalkReq{
err := s.vc.SendRequest(ctx, voice.MethodPushToTalk, voice.PushToTalkReq{
Audio: utt,
Lang: lang,
Surface: voice.SurfaceVoice,
}, &resp)
if err != nil {
return fmt.Errorf("push-to-talk: %w", err)
return audio.Audio{}, fmt.Errorf("push-to-talk: %w", err)
}
log.Printf("mavwaked: reply: %q (%.2fs audio)", resp.ReplyText, resp.ReplyAudio.Duration())
// Play the reply audio.
if len(resp.ReplyAudio.Bytes) > 0 {
go playAudio(resp.ReplyAudio)
} else {
log.Printf("mavwaked: empty reply audio (text only)")
}
if len(resp.RoutedChannels) > 0 {
log.Printf("mavwaked: also routed to: %v", resp.RoutedChannels)
}
return nil
}
// playAudio pipes PCM audio to aplay(1) for playback. Runs in a goroutine.
func playAudio(a audio.Audio) {
// Build WAV header for aplay (or pipe raw PCM with the right format flags).
cmd := exec.Command("aplay",
"-f", "S16_LE",
"-r", fmt.Sprintf("%d", a.Format.SampleRate),
"-c", fmt.Sprintf("%d", a.Format.Channels),
"-t", "raw",
)
stdin, err := cmd.StdinPipe()
if err != nil {
log.Printf("mavwaked: aplay stdin pipe: %v", err)
return
}
if err := cmd.Start(); err != nil {
log.Printf("mavwaked: start aplay: %v", err)
return
}
// Write audio to aplay's stdin.
if _, err := stdin.Write(a.Bytes); err != nil {
log.Printf("mavwaked: write to aplay: %v", err)
}
_ = stdin.Close()
// Wait for playback to finish (with a timeout).
done := make(chan error, 1)
go func() {
done <- cmd.Wait()
}()
select {
case err := <-done:
if err != nil {
log.Printf("mavwaked: aplay: %v", err)
}
case <-time.After(30 * time.Second):
log.Printf("mavwaked: aplay timeout, killing")
_ = cmd.Process.Kill()
<-done
}
return resp.ReplyAudio, nil
}
+139
View File
@@ -0,0 +1,139 @@
package main
// Reply playback, and the half-duplex gate around it (Vikunja #287).
//
// Before this, playback was `go playAudio(reply)` — fire and forget, with no
// handle on the running aplay. Two things fell out of that, and both are
// audible:
//
// 1. Self-trigger. The capture loop keeps feeding the VAD while the speaker
// is playing, so Maven's own reply comes back in through the mic, trips
// the VAD, and is sent to the daemon as a fresh utterance. She answers
// herself. There is no acoustic echo canceller in this pipeline, so the
// only correct fix is half-duplex: while she is speaking, the capture
// side is muted.
//
// 2. No barge-in. Talking over her did nothing — there was nothing to
// cancel, because nobody held the process handle.
//
// The two are the same mechanism seen from opposite sides, so they live
// together here. Echo suppression is unconditional (it fixes a bug). Barge-in
// is off unless -barge-in is passed, because it needs a room-specific energy
// threshold: with no echo canceller, the only way to tell "he is talking over
// her" from "the mic is hearing her" is that he is louder, and how much
// louder depends on where the mic sits relative to the speaker.
import (
"log"
"os/exec"
"strconv"
"sync"
"time"
"github.com/kami/maven/internal/audio"
)
// player plays one reply at a time and can be cut off mid-utterance.
type player interface {
// Play starts playback of a, replacing anything already playing, and
// returns immediately.
Play(a audio.Audio)
// Stop ends playback now. A no-op when nothing is playing.
Stop()
// Playing reports whether audio is currently going out of the speaker.
Playing() bool
}
// aplayPlayer pipes raw PCM to aplay(1). Stop kills the child, which is what
// makes barge-in instant rather than "instant at the end of the sentence".
type aplayPlayer struct {
mu sync.Mutex
cmd *exec.Cmd
playing bool
// gen rises on every Play/Stop so a finishing playback cannot clear the
// playing flag of the one that replaced it.
gen uint64
}
func newAplayPlayer() *aplayPlayer { return &aplayPlayer{} }
func (p *aplayPlayer) Play(a audio.Audio) {
if len(a.Bytes) == 0 {
return
}
p.Stop()
cmd := exec.Command("aplay",
"-f", "S16_LE",
"-r", strconv.Itoa(a.Format.SampleRate),
"-c", strconv.Itoa(a.Format.Channels),
"-t", "raw",
)
stdin, err := cmd.StdinPipe()
if err != nil {
log.Printf("mavwaked: aplay stdin pipe: %v", err)
return
}
if err := cmd.Start(); err != nil {
log.Printf("mavwaked: start aplay: %v", err)
_ = stdin.Close()
return
}
p.mu.Lock()
p.gen++
gen := p.gen
p.cmd = cmd
p.playing = true
p.mu.Unlock()
go func() {
if _, err := stdin.Write(a.Bytes); err != nil {
// Broken pipe is the expected outcome of Stop().
log.Printf("mavwaked: write to aplay: %v", err)
}
_ = stdin.Close()
done := make(chan error, 1)
go func() { done <- cmd.Wait() }()
select {
case err := <-done:
if err != nil {
log.Printf("mavwaked: aplay: %v", err)
}
case <-time.After(30 * time.Second):
log.Printf("mavwaked: aplay timeout, killing")
if pr := cmd.Process; pr != nil {
_ = pr.Kill()
}
<-done
}
p.mu.Lock()
if p.gen == gen {
p.playing = false
p.cmd = nil
}
p.mu.Unlock()
}()
}
func (p *aplayPlayer) Stop() {
p.mu.Lock()
cmd := p.cmd
if cmd != nil {
p.gen++
p.playing = false
p.cmd = nil
}
p.mu.Unlock()
if cmd != nil && cmd.Process != nil {
_ = cmd.Process.Kill()
}
}
func (p *aplayPlayer) Playing() bool {
p.mu.Lock()
defer p.mu.Unlock()
return p.playing
}
+33
View File
@@ -0,0 +1,33 @@
package main
import (
"testing"
"github.com/kami/maven/internal/audio"
)
// The real player must be safe to poke when nothing is playing — the capture
// loop calls Playing() on every 30ms frame, and Stop() lands on an idle
// player whenever a barge-in races the end of a reply. Neither may need
// aplay(1) to be installed.
func TestAplayPlayerIdleIsSafe(t *testing.T) {
p := newAplayPlayer()
if p.Playing() {
t.Fatal("a fresh player reports playing")
}
p.Stop()
p.Stop()
if p.Playing() {
t.Fatal("playing after Stop on an idle player")
}
// Empty audio is a text-only turn: nothing to play, no process to spawn.
p.Play(audio.Audio{Format: audio.PCM16kMono})
if p.Playing() {
t.Fatal("empty audio started playback")
}
}
func TestAplayPlayerSatisfiesPlayer(t *testing.T) {
var _ player = newAplayPlayer()
var _ player = &fakePlayer{}
}
+123
View File
@@ -0,0 +1,123 @@
package main
// The capture session: what happens to one 30ms frame, given whether Maven is
// currently speaking. Split out of main.go's processFrame so the decision is
// testable without a mic, a speaker, or a daemon (Vikunja #287).
import (
"context"
"log"
"github.com/kami/maven/internal/audio"
)
// utteranceSender ships one complete utterance to the voice server and
// returns the reply audio to play. The real one round-trips over the voice
// wire; tests substitute a recorder.
type utteranceSender interface {
Send(ctx context.Context, utt audio.Audio, lang string) (audio.Audio, error)
}
// bargeInConfig holds the two numbers barge-in needs. Zero Frames disables
// barge-in entirely — the half-duplex gate still runs.
type bargeInConfig struct {
// RMS is the normalised energy a frame must exceed to count as him
// talking over her rather than the mic hearing her. It is deliberately
// far above the VAD's own floor: the speaker leaks into the mic at
// roughly ambient level, a person talking at the mic does not.
RMS float64
// Frames is how many consecutive frames must clear RMS before playback
// is cut. One loud frame is a door closing; five in a row is a voice.
Frames int
}
// Enabled reports whether barge-in should be attempted at all.
func (c bargeInConfig) Enabled() bool { return c.Frames > 0 && c.RMS > 0 }
// session is the per-client capture state machine.
type session struct {
vad *VAD
player player
sender utteranceSender
lang string
barge bargeInConfig
// loudFrames counts consecutive over-threshold frames seen while she is
// speaking. Reset whenever a frame falls back under the threshold, and
// whenever playback ends.
loudFrames int
// counters, read by tests and logged on the way out.
suppressed int // frames dropped because she was speaking
bargeIns int // times playback was cut because he spoke over her
sent int // utterances shipped to the daemon
}
func newSession(vad *VAD, p player, s utteranceSender, lang string, barge bargeInConfig) *session {
return &session{vad: vad, player: p, sender: s, lang: lang, barge: barge}
}
// feed processes one 30ms PCM frame.
//
// While the player is running the capture side is muted: the VAD is not fed
// and no utterance can be produced, so Maven's own reply cannot come back in
// as a new command. The one thing that gets through is barge-in — sustained
// energy well above the speaker's leak level cuts playback, and capture
// resumes on the very next frame with a clean VAD.
func (s *session) feed(ctx context.Context, frame []byte) error {
if s.player.Playing() {
s.suppressed++
if !s.barge.Enabled() {
return nil
}
if frameRMS(PCMToI16(frame)) < s.barge.RMS {
s.loudFrames = 0
return nil
}
s.loudFrames++
if s.loudFrames < s.barge.Frames {
return nil
}
// He is talking over her. Cut her off, drop the VAD state that
// accumulated from the echo, and start listening for real.
s.player.Stop()
s.bargeIns++
s.loudFrames = 0
s.vad.Reset()
log.Printf("mavwaked: barge-in — stopped playback")
return nil
}
// Not speaking. If we just stopped, make sure no echo-era state leaks
// into the next utterance.
if s.loudFrames != 0 {
s.loudFrames = 0
s.vad.Reset()
}
utt, state := s.vad.Feed(PCMToI16(frame))
if state == StateSpeech || utt.Bytes == nil {
return nil
}
return s.dispatch(ctx, utt)
}
// dispatch ships a complete utterance and plays whatever comes back.
func (s *session) dispatch(ctx context.Context, utt audio.Audio) error {
log.Printf("mavwaked: utterance complete (%.2fs, %d bytes), sending...", utt.Duration(), len(utt.Bytes))
reply, err := s.sender.Send(ctx, utt, s.lang)
s.sent++
if err != nil {
return err
}
if len(reply.Bytes) == 0 {
log.Printf("mavwaked: empty reply audio (text only)")
return nil
}
// The VAD has been accumulating from the buffered mic stream while the
// round-trip blocked. None of it is a command — reset before the
// speaker opens, so the first post-reply frame starts clean.
s.vad.Reset()
s.player.Play(reply)
return nil
}
+282
View File
@@ -0,0 +1,282 @@
package main
import (
"context"
"errors"
"math"
"testing"
"github.com/kami/maven/internal/audio"
)
// fakePlayer records Play/Stop instead of shelling out to aplay.
type fakePlayer struct {
playing bool
plays int
stops int
last audio.Audio
}
func (p *fakePlayer) Play(a audio.Audio) { p.playing = true; p.plays++; p.last = a }
func (p *fakePlayer) Stop() { p.playing = false; p.stops++ }
func (p *fakePlayer) Playing() bool { return p.playing }
// fakeSender records what was shipped and hands back a canned reply.
type fakeSender struct {
sent []audio.Audio
reply audio.Audio
err error
}
func (s *fakeSender) Send(_ context.Context, utt audio.Audio, _ string) (audio.Audio, error) {
s.sent = append(s.sent, utt)
return s.reply, s.err
}
func replyAudio() audio.Audio {
return audio.Audio{Format: audio.PCM16kMono, Bytes: make([]byte, 16000)}
}
// frameAt returns a 30ms frame whose RMS is approximately rms.
func frameAt(rms float64) []byte {
amp := rms * math.Sqrt2 * 32768
f := make([]int16, frameSamples)
for i := range f {
f[i] = int16(amp * math.Sin(2*math.Pi*440*float64(i)/16000))
}
return pcmBytes(f)
}
func silentBytes() []byte { return make([]byte, frameSamples*2) }
// newTestSession wires a session with fakes and a default VAD.
func newTestSession(barge bargeInConfig) (*session, *fakePlayer, *fakeSender) {
p := &fakePlayer{}
s := &fakeSender{reply: replyAudio()}
return newSession(NewVAD(0, 0, 0, 0), p, s, "ru", barge), p, s
}
// speakThenPause drives a full utterance through the session: enough loud
// frames to trigger, then enough silence to end it.
func speakThenPause(t *testing.T, sess *session) {
t.Helper()
speechFrames := (defaultSpeechMs + defaultFrameMs - 1) / defaultFrameMs
silenceFrames := (defaultSilenceMs+defaultFrameMs-1)/defaultFrameMs + 2
loud := frameAt(0.35)
for i := 0; i < speechFrames+5; i++ {
if err := sess.feed(context.Background(), loud); err != nil {
t.Fatalf("feed loud frame %d: %v", i, err)
}
}
for i := 0; i < silenceFrames; i++ {
if err := sess.feed(context.Background(), silentBytes()); err != nil {
t.Fatalf("feed silent frame %d: %v", i, err)
}
}
}
func TestSessionSendsUtteranceAndPlaysReply(t *testing.T) {
sess, p, snd := newTestSession(bargeInConfig{})
speakThenPause(t, sess)
if len(snd.sent) != 1 {
t.Fatalf("sent %d utterances, want 1", len(snd.sent))
}
if snd.sent[0].Format != audio.PCM16kMono {
t.Errorf("utterance format = %+v, want canonical", snd.sent[0].Format)
}
if p.plays != 1 {
t.Errorf("plays = %d, want 1", p.plays)
}
}
// The bug this whole file exists for: while the speaker is running, the mic
// hears Maven and the old code shipped that back as a fresh command.
func TestSessionDoesNotHearItselfWhilePlaying(t *testing.T) {
sess, p, snd := newTestSession(bargeInConfig{})
speakThenPause(t, sess)
if !p.Playing() {
t.Fatal("expected playback to be running after the reply")
}
// Feed a long stretch of loud audio — Maven's own voice coming back in.
base := sess.suppressed
loud := frameAt(0.35)
for i := 0; i < 200; i++ {
if err := sess.feed(context.Background(), loud); err != nil {
t.Fatalf("feed echo frame %d: %v", i, err)
}
}
if len(snd.sent) != 1 {
t.Fatalf("sent %d utterances, want 1 — her own reply was captured as a command", len(snd.sent))
}
if got := sess.suppressed - base; got != 200 {
t.Errorf("suppressed %d of the 200 echo frames, want all of them", got)
}
if p.stops != 0 {
t.Errorf("stops = %d, want 0 — barge-in is off, nothing should cut her off", p.stops)
}
}
// With barge-in off, no amount of noise stops playback.
func TestSessionBargeInDisabledByDefault(t *testing.T) {
sess, p, _ := newTestSession(bargeInConfig{})
if sess.barge.Enabled() {
t.Fatal("zero bargeInConfig must be disabled")
}
speakThenPause(t, sess)
veryLoud := frameAt(0.6)
for i := 0; i < 50; i++ {
_ = sess.feed(context.Background(), veryLoud)
}
if p.stops != 0 || sess.bargeIns != 0 {
t.Fatalf("stops = %d, bargeIns = %d, want 0 with barge-in off", p.stops, sess.bargeIns)
}
}
func TestSessionBargeInCutsPlayback(t *testing.T) {
barge := bargeInConfig{RMS: 0.12, Frames: 5}
sess, p, _ := newTestSession(barge)
speakThenPause(t, sess)
if !p.Playing() {
t.Fatal("expected playback after the reply")
}
// Four loud frames must not be enough — a door closing is not a voice.
veryLoud := frameAt(0.35)
for i := 0; i < 4; i++ {
_ = sess.feed(context.Background(), veryLoud)
}
if p.stops != 0 {
t.Fatalf("playback cut after 4 frames, want it to hold until %d", barge.Frames)
}
// The fifth cuts her off.
_ = sess.feed(context.Background(), veryLoud)
if p.stops != 1 || sess.bargeIns != 1 {
t.Fatalf("stops = %d, bargeIns = %d, want 1 and 1", p.stops, sess.bargeIns)
}
if p.Playing() {
t.Fatal("still playing after barge-in")
}
}
// A burst that falls back under the threshold resets the counter, so noise
// spread over a whole reply never accumulates into a false barge-in.
func TestSessionBargeInNeedsConsecutiveFrames(t *testing.T) {
sess, p, _ := newTestSession(bargeInConfig{RMS: 0.12, Frames: 5})
speakThenPause(t, sess)
veryLoud := frameAt(0.35)
quiet := frameAt(0.02)
for i := 0; i < 20; i++ {
_ = sess.feed(context.Background(), veryLoud)
_ = sess.feed(context.Background(), veryLoud)
_ = sess.feed(context.Background(), quiet)
}
if p.stops != 0 || sess.bargeIns != 0 {
t.Fatalf("stops = %d, bargeIns = %d, want 0 — two-frame bursts must not accumulate", p.stops, sess.bargeIns)
}
}
// Speaker leak sits near the room floor; it must never reach the barge-in bar.
func TestSessionEchoLevelAudioNeverBargesIn(t *testing.T) {
sess, p, _ := newTestSession(bargeInConfig{RMS: 0.12, Frames: 5})
speakThenPause(t, sess)
base := sess.suppressed
leak := frameAt(0.05) // loud enough for the VAD, far under the barge bar
for i := 0; i < 300; i++ {
_ = sess.feed(context.Background(), leak)
}
if p.stops != 0 {
t.Fatalf("stops = %d, want 0 — speaker leak must not read as barge-in", p.stops)
}
if got := sess.suppressed - base; got != 300 {
t.Errorf("suppressed %d of the 300 leak frames, want all of them", got)
}
}
// After barge-in the VAD must start clean, so the interrupting speech is
// captured as a whole utterance rather than joined onto echo state.
func TestSessionCapturesTheInterruptingUtterance(t *testing.T) {
sess, p, snd := newTestSession(bargeInConfig{RMS: 0.12, Frames: 5})
speakThenPause(t, sess)
veryLoud := frameAt(0.35)
for i := 0; i < 5; i++ {
_ = sess.feed(context.Background(), veryLoud)
}
if p.stops != 1 {
t.Fatalf("expected barge-in, stops = %d", p.stops)
}
// He keeps talking; that is a new command.
speakThenPause(t, sess)
if len(snd.sent) != 2 {
t.Fatalf("sent %d utterances, want 2 — the interruption itself must be heard", len(snd.sent))
}
if p.plays != 2 {
t.Errorf("plays = %d, want 2", p.plays)
}
}
// A failed round-trip must surface as an error and must not start playback.
func TestSessionSendErrorDoesNotPlay(t *testing.T) {
p := &fakePlayer{}
snd := &fakeSender{err: errors.New("boom")}
sess := newSession(NewVAD(0, 0, 0, 0), p, snd, "ru", bargeInConfig{})
speechFrames := (defaultSpeechMs + defaultFrameMs - 1) / defaultFrameMs
silenceFrames := (defaultSilenceMs+defaultFrameMs-1)/defaultFrameMs + 2
loud := frameAt(0.35)
var lastErr error
for i := 0; i < speechFrames+5; i++ {
_ = sess.feed(context.Background(), loud)
}
for i := 0; i < silenceFrames; i++ {
if err := sess.feed(context.Background(), silentBytes()); err != nil {
lastErr = err
}
}
if lastErr == nil {
t.Fatal("send error was swallowed")
}
if p.plays != 0 || p.Playing() {
t.Fatalf("plays = %d, playing = %v, want no playback on a failed round-trip", p.plays, p.Playing())
}
}
// An empty reply (text-only turn) must leave the capture side open.
func TestSessionEmptyReplyLeavesCaptureOpen(t *testing.T) {
p := &fakePlayer{}
snd := &fakeSender{reply: audio.Audio{Format: audio.PCM16kMono}}
sess := newSession(NewVAD(0, 0, 0, 0), p, snd, "ru", bargeInConfig{})
speakThenPause(t, sess)
if p.plays != 0 {
t.Fatalf("plays = %d, want 0 for an empty reply", p.plays)
}
speakThenPause(t, sess)
if len(snd.sent) != 2 {
t.Fatalf("sent %d, want 2 — capture must stay open when there is no audio reply", len(snd.sent))
}
}
func TestBargeInConfigEnabled(t *testing.T) {
cases := []struct {
c bargeInConfig
want bool
}{
{bargeInConfig{}, false},
{bargeInConfig{RMS: 0.12}, false},
{bargeInConfig{Frames: 5}, false},
{bargeInConfig{RMS: 0.12, Frames: 5}, true},
}
for _, tc := range cases {
if got := tc.c.Enabled(); got != tc.want {
t.Errorf("%+v.Enabled() = %v, want %v", tc.c, got, tc.want)
}
}
}
+9
View File
@@ -23,6 +23,15 @@ const (
defaultSilenceMs = 800 // silence hold before declaring end-of-utterance
defaultMaxMs = 10000 // cap single utterance at 10s
defaultMinRMS = 0.01 // RMS floor (same as mavsttd)
// Barge-in thresholds. Only used when -barge-in is passed. The RMS is
// x10000 like -min-rms, and sits an order of magnitude above the VAD's
// own floor on purpose: with no acoustic echo canceller, a frame only
// counts as "he is talking over her" if it is far louder than what the
// speaker leaks back into the mic. 5 frames is 150ms — long enough that
// a door or a cough does not cut her off mid-sentence.
defaultBargeRMS = 1200 // 0.12 normalised RMS
defaultBargeFrames = 5
)
// frameSamples — samples per 30ms frame at 16kHz.
+135
View File
@@ -0,0 +1,135 @@
package main
import (
"crypto/subtle"
"encoding/json"
"errors"
"io"
"log"
"net/http"
"strings"
"github.com/kami/maven/internal/calendar"
"github.com/kami/maven/internal/ipc"
)
// POST /api/ambient — the work calendar read (Vikunja #126).
//
// Maven does not hold a work credential. A corp mail or calendar session on the
// homelab ties the box's blast radius to the employer's data, so the work
// calendar is read as a SIGNAL instead: an Android notification-listener on the
// owner's phone posts meeting notifications here over wg/LAN, and the ones that
// clearly describe a meeting become calendar events at source=ambient:notif,
// confidence below 1.0. Mail as a notification signal, not a mailbox.
//
// Off unless configured: no -ambient-token, no route. The token is a shared
// secret because the poster is a phone service, not a browser — WebAuthn has no
// answer for a background Android service. The endpoint is write-only and
// accepts exactly one shape of write; it cannot read anything back out.
//
// A notification with no recognisable clock reading stores NOTHING. Maven is
// not a guesser-of-truth, and a mailbox of noise rendered as invented meetings
// is worse than a gap.
// ambientMaxBody bounds the request. A notification is two short lines.
const ambientMaxBody = 8 << 10
type ambientResp struct {
Stored bool `json:"stored"`
Key string `json:"key,omitempty"`
Reason string `json:"reason,omitempty"`
}
// handleAmbient ingests one relayed notification. token is the configured
// shared secret; an empty token means the capability is off and the handler is
// never registered, so it is treated as a hard failure here too.
func handleAmbient(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI, token string) {
if r.Method != http.MethodPost {
http.Error(w, "POST only", http.StatusMethodNotAllowed)
return
}
if token == "" {
http.Error(w, "ambient ingest disabled (no -ambient-token)", http.StatusServiceUnavailable)
return
}
if !ambientAuthorized(r, token) {
http.Error(w, "unauthorized", http.StatusUnauthorized)
return
}
if core == nil {
http.Error(w, "ambient ingest disabled (no -core)", http.StatusServiceUnavailable)
return
}
var n calendar.Notification
body, err := io.ReadAll(io.LimitReader(r.Body, ambientMaxBody))
if err != nil {
http.Error(w, "read failed", http.StatusBadRequest)
return
}
if err := json.Unmarshal(body, &n); err != nil {
http.Error(w, "bad json", http.StatusBadRequest)
return
}
if n.Posted.IsZero() {
writeAmbient(w, http.StatusBadRequest, ambientResp{Reason: "posted_at is required"})
return
}
ev, ok := calendar.EventFromNotification(n)
if !ok {
// Not an event. 202: the relay did its job, there is just nothing here
// worth remembering, and it must not retry.
writeAmbient(w, http.StatusAccepted, ambientResp{Reason: "no meeting time in notification"})
return
}
key := calendar.FactKey(ev)
val := calendar.FactValue(ev)
// Append-only discipline, same as cmd/mavcaldav: a phone reposts the same
// notification many times, and each repost is the same event.
if prev, err := core.LatestFactBySource(r.Context(), key, calendar.SourceAmbient); err == nil && prev.Value == val {
writeAmbient(w, http.StatusOK, ambientResp{Stored: false, Key: key, Reason: "unchanged"})
return
} else if err != nil && !errors.Is(err, ipc.ErrNoFact) {
log.Printf("ambient: read %s: %v", key, err)
http.Error(w, "read failed", http.StatusBadGateway)
return
}
// kind=env: an observation about the world, never a self-fact — a passive
// signal does not write truth about the owner. Confidence below 1.0 is the
// honest part: this is a notification about a meeting, not a reading of a
// calendar, and the query path hedges when it recites one.
if _, err := core.WriteFact(r.Context(), ipc.WriteFactReq{
Ts: ev.Start,
Kind: "env",
Key: key,
Value: val,
Source: calendar.SourceAmbient,
Confidence: calendar.AmbientConfidence,
}); err != nil {
log.Printf("ambient: write %s: %v", key, err)
http.Error(w, "write failed", http.StatusBadGateway)
return
}
log.Printf("ambient: %s=%s (%s, pkg=%s)", key, val, calendar.SourceAmbient, n.Package)
writeAmbient(w, http.StatusCreated, ambientResp{Stored: true, Key: key})
}
// ambientAuthorized accepts the token as a bearer header or as an X-Maven-Token
// header, compared in constant time.
func ambientAuthorized(r *http.Request, token string) bool {
got := strings.TrimSpace(strings.TrimPrefix(r.Header.Get("Authorization"), "Bearer"))
if got == "" {
got = strings.TrimSpace(r.Header.Get("X-Maven-Token"))
}
return subtle.ConstantTimeCompare([]byte(got), []byte(token)) == 1
}
func writeAmbient(w http.ResponseWriter, code int, resp ambientResp) {
w.Header().Set("Content-Type", "application/json")
w.WriteHeader(code)
json.NewEncoder(w).Encode(resp)
}
+223
View File
@@ -0,0 +1,223 @@
package main
import (
"context"
"encoding/json"
"fmt"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/calendar"
"github.com/kami/maven/internal/ipc"
)
const ambientTestToken = "s3cret"
// ambientCore adds provenance-scoped reads to fakeCore, which the dedupe path
// needs.
type ambientCore struct {
fakeCore
latest map[string]ipc.Fact // "key|source" → fact
readErr error
}
func (c *ambientCore) LatestFactBySource(_ context.Context, key, source string) (ipc.Fact, error) {
if c.readErr != nil {
return ipc.Fact{}, c.readErr
}
f, ok := c.latest[key+"|"+source]
if !ok {
return ipc.Fact{}, ipc.ErrNoFact
}
return f, nil
}
func postAmbient(t *testing.T, core ipc.CoreAPI, token string, n calendar.Notification) (*httptest.ResponseRecorder, ambientResp) {
t.Helper()
body, err := json.Marshal(n)
if err != nil {
t.Fatal(err)
}
req := httptest.NewRequest(http.MethodPost, "/api/ambient", strings.NewReader(string(body)))
req.Header.Set("Authorization", "Bearer "+ambientTestToken)
rr := httptest.NewRecorder()
handleAmbient(rr, req, core, token)
var resp ambientResp
json.Unmarshal(rr.Body.Bytes(), &resp)
return rr, resp
}
func meetingNotification() calendar.Notification {
return calendar.Notification{
Package: "com.google.android.gm",
Title: "Планёрка",
Text: "10:00-10:30",
Posted: time.Date(2026, 8, 3, 9, 40, 0, 0, time.UTC),
}
}
func TestHandleAmbientStoresMeeting(t *testing.T) {
core := &ambientCore{}
rr, resp := postAmbient(t, core, ambientTestToken, meetingNotification())
if rr.Code != http.StatusCreated {
t.Fatalf("status = %d, want 201: %s", rr.Code, rr.Body)
}
if !resp.Stored {
t.Errorf("resp = %+v, want stored", resp)
}
if len(core.writeLog) != 1 {
t.Fatalf("expected 1 fact write, got %d", len(core.writeLog))
}
got := core.writeLog[0]
if got.Source != calendar.SourceAmbient {
t.Errorf("source = %q, want %q", got.Source, calendar.SourceAmbient)
}
if got.Confidence >= 1.0 {
t.Errorf("confidence = %v — a notification is not a calendar read", got.Confidence)
}
if got.Confidence != calendar.AmbientConfidence {
t.Errorf("confidence = %v, want %v", got.Confidence, calendar.AmbientConfidence)
}
if got.Kind != "env" {
t.Errorf("kind = %q — a passive signal never writes a self-fact", got.Kind)
}
if want := "calendar_event_20260803_"; !strings.HasPrefix(got.Key, want) {
t.Errorf("key = %q, want prefix %q", got.Key, want)
}
if got.Value != "Планёрка @ 10:00-10:30" {
t.Errorf("value = %q", got.Value)
}
}
// A phone reposts the same notification many times. Each repost is the same
// event, and the append-only log must not fill with duplicates.
func TestHandleAmbientDedupesReposts(t *testing.T) {
core := &ambientCore{}
postAmbient(t, core, ambientTestToken, meetingNotification())
if len(core.writeLog) != 1 {
t.Fatalf("first post did not write")
}
w := core.writeLog[0]
core.latest = map[string]ipc.Fact{w.Key + "|" + w.Source: {Value: w.Value}}
rr, resp := postAmbient(t, core, ambientTestToken, meetingNotification())
if rr.Code != http.StatusOK {
t.Errorf("status = %d, want 200 for an unchanged repost", rr.Code)
}
if resp.Stored {
t.Error("a repost must not be stored again")
}
if len(core.writeLog) != 1 {
t.Errorf("wrote %d facts, want 1", len(core.writeLog))
}
}
// The conservative half: noise stores nothing at all.
func TestHandleAmbientIgnoresNonMeetings(t *testing.T) {
core := &ambientCore{}
rr, resp := postAmbient(t, core, ambientTestToken, calendar.Notification{
Package: "com.google.android.gm",
Title: "3 новых письма",
Posted: time.Now(),
})
if rr.Code != http.StatusAccepted {
t.Errorf("status = %d, want 202 (accepted, nothing to store — the relay must not retry)", rr.Code)
}
if resp.Stored {
t.Error("a notification with no meeting time must store nothing")
}
if len(core.writeLog) != 0 {
t.Fatalf("wrote %d facts for a non-meeting", len(core.writeLog))
}
}
func TestHandleAmbientAuth(t *testing.T) {
body := `{"title":"Планёрка 10:00","posted_at":"2026-08-03T09:40:00Z"}`
newReq := func(hdr, val string) *http.Request {
r := httptest.NewRequest(http.MethodPost, "/api/ambient", strings.NewReader(body))
if hdr != "" {
r.Header.Set(hdr, val)
}
return r
}
t.Run("no token rejected", func(t *testing.T) {
core := &ambientCore{}
rr := httptest.NewRecorder()
handleAmbient(rr, newReq("", ""), core, ambientTestToken)
if rr.Code != http.StatusUnauthorized {
t.Errorf("status = %d, want 401", rr.Code)
}
if len(core.writeLog) != 0 {
t.Error("an unauthorized post must not write")
}
})
t.Run("wrong token rejected", func(t *testing.T) {
rr := httptest.NewRecorder()
handleAmbient(rr, newReq("Authorization", "Bearer nope"), &ambientCore{}, ambientTestToken)
if rr.Code != http.StatusUnauthorized {
t.Errorf("status = %d, want 401", rr.Code)
}
})
t.Run("X-Maven-Token accepted", func(t *testing.T) {
rr := httptest.NewRecorder()
handleAmbient(rr, newReq("X-Maven-Token", ambientTestToken), &ambientCore{}, ambientTestToken)
if rr.Code != http.StatusCreated {
t.Errorf("status = %d, want 201: %s", rr.Code, rr.Body)
}
})
t.Run("capability off", func(t *testing.T) {
rr := httptest.NewRecorder()
handleAmbient(rr, newReq("Authorization", "Bearer "+ambientTestToken), &ambientCore{}, "")
if rr.Code != http.StatusServiceUnavailable {
t.Errorf("status = %d, want 503 when no token is configured", rr.Code)
}
})
t.Run("GET rejected", func(t *testing.T) {
rr := httptest.NewRecorder()
r := httptest.NewRequest(http.MethodGet, "/api/ambient", nil)
handleAmbient(rr, r, &ambientCore{}, ambientTestToken)
if rr.Code != http.StatusMethodNotAllowed {
t.Errorf("status = %d, want 405 — the ingest is write-only", rr.Code)
}
})
}
func TestHandleAmbientBadInput(t *testing.T) {
t.Run("bad json", func(t *testing.T) {
req := httptest.NewRequest(http.MethodPost, "/api/ambient", strings.NewReader("{nope"))
req.Header.Set("X-Maven-Token", ambientTestToken)
rr := httptest.NewRecorder()
handleAmbient(rr, req, &ambientCore{}, ambientTestToken)
if rr.Code != http.StatusBadRequest {
t.Errorf("status = %d, want 400", rr.Code)
}
})
t.Run("missing posted_at", func(t *testing.T) {
req := httptest.NewRequest(http.MethodPost, "/api/ambient", strings.NewReader(`{"title":"Планёрка 10:00"}`))
req.Header.Set("X-Maven-Token", ambientTestToken)
rr := httptest.NewRecorder()
handleAmbient(rr, req, &ambientCore{}, ambientTestToken)
if rr.Code != http.StatusBadRequest {
t.Errorf("status = %d, want 400", rr.Code)
}
})
t.Run("read error surfaces", func(t *testing.T) {
core := &ambientCore{readErr: fmt.Errorf("socket closed")}
rr, _ := postAmbient(t, core, ambientTestToken, meetingNotification())
if rr.Code != http.StatusBadGateway {
t.Errorf("status = %d, want 502", rr.Code)
}
})
}
+4 -1
View File
@@ -91,7 +91,10 @@ func handleEcosystem(w http.ResponseWriter, r *http.Request, urls ecoURLs) {
var d ecoData
var wg sync.WaitGroup
wg.Add(3)
go func() { defer wg.Done(); d.Nexus.Err = getEco(ctx, urls.nexus, "/api/v1/entities?limit=50", &d.Nexus.Rows) }()
go func() {
defer wg.Done()
d.Nexus.Err = getEco(ctx, urls.nexus, "/api/v1/entities?limit=50", &d.Nexus.Rows)
}()
go func() {
defer wg.Done()
d.Praxis.Err = getEco(ctx, urls.praxis, "/api/v1/items?limit=50", &d.Praxis.Rows)
+24
View File
@@ -0,0 +1,24 @@
{{template "shellTop" "events"}}
<h1>Intake</h1>
<div class=hint>Everything that arrived, newest first — a relayed notification, a mail candidate, a feed
item, a changed page, a spend, a presence probe. One envelope per write; the durable row is still the
fact, note or task itself. Held in memory only, so a restart empties this.</div>
{{if .Err}}<div class=hint>journal unavailable: {{.Err}}</div>{{end}}
{{if and (not .Events) (not .Err)}}
<div class=hint>nothing has arrived yet</div>
{{end}}
{{if .Events}}
<div class=scroll><table class=mono>
<tr><th>when<th>source<th>kind<th>pri<th>what<th>detail</tr>
{{range .Events}}<tr>
<td>{{.OccurredAt.Format "02.01 15:04:05"}}</td>
<td class=gray>{{.Source}}</td>
<td class=gray>{{.Kind}}</td>
<td class=gray>{{.Priority}}</td>
<td>{{.Title}}</td>
<td class=gray>{{.Body}}</td>
</tr>{{end}}
</table></div>
{{end}}
{{template "shellBottom"}}
</html>
+107
View File
@@ -0,0 +1,107 @@
package main
import (
"context"
"errors"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
)
// eventsCore serves a canned intake journal. Embedding
// ipc.UnimplementedCoreAPI means any other call fails loudly.
type eventsCore struct {
ipc.UnimplementedCoreAPI
events []ipc.IntakeEvent
err error
gotN int
}
func (c *eventsCore) RecentEvents(_ context.Context, n int) ([]ipc.IntakeEvent, error) {
c.gotN = n
return c.events, c.err
}
func getEvents(t *testing.T, core ipc.CoreAPI) *httptest.ResponseRecorder {
t.Helper()
w := httptest.NewRecorder()
handleEvents(w, httptest.NewRequest(http.MethodGet, "/events", nil), core)
return w
}
func TestEventsPageRendersTheJournal(t *testing.T) {
core := &eventsCore{events: []ipc.IntakeEvent{
{Source: "rss:tech", Kind: "note", Title: "Вышло ядро 6.19", Priority: "low",
OccurredAt: time.Date(2026, 8, 1, 7, 15, 0, 0, time.UTC)},
{Source: "ambient:notif", Kind: "fact", Title: "calendar_event_20260801_планёрка",
Body: "10:00-11:00 планёрка", Priority: "low",
OccurredAt: time.Date(2026, 8, 1, 10, 0, 0, 0, time.UTC)},
}}
w := getEvents(t, core)
if w.Code != http.StatusOK {
t.Fatalf("status = %d, want 200", w.Code)
}
body := w.Body.String()
for _, want := range []string{"rss:tech", "Вышло ядро 6.19", "ambient:notif", "10:00-11:00 планёрка", "01.08 10:00:00"} {
if !strings.Contains(body, want) {
t.Errorf("page does not mention %q", want)
}
}
if core.gotN != eventsPageLimit {
t.Errorf("asked core for %d events, want %d", core.gotN, eventsPageLimit)
}
}
func TestEventsPageSaysNothingArrived(t *testing.T) {
w := getEvents(t, &eventsCore{})
if w.Code != http.StatusOK {
t.Fatalf("status = %d, want 200", w.Code)
}
if !strings.Contains(w.Body.String(), "nothing has arrived yet") {
t.Error("empty journal did not render the empty-state line")
}
}
func TestEventsPageReportsAReadFailure(t *testing.T) {
// An unreachable journal must say so rather than render an empty table,
// which would imply nothing arrived.
w := getEvents(t, &eventsCore{err: errors.New("core is down")})
if w.Code != http.StatusOK {
t.Fatalf("status = %d, want 200 with the error rendered", w.Code)
}
body := w.Body.String()
if !strings.Contains(body, "journal unavailable") || !strings.Contains(body, "core is down") {
t.Errorf("page did not report the read failure: %s", body)
}
if strings.Contains(body, "nothing has arrived yet") {
t.Error("a failed read rendered as an empty journal")
}
}
func TestEventsPageWithoutCore(t *testing.T) {
w := getEvents(t, nil)
if w.Code != http.StatusServiceUnavailable {
t.Errorf("status = %d, want 503", w.Code)
}
}
func TestEventsPageEscapesIntakeText(t *testing.T) {
// Titles come from outside — a feed headline, a notification. They are shown
// on a page and must never be able to inject markup into it.
core := &eventsCore{events: []ipc.IntakeEvent{{
Source: "rss:x", Kind: "note", Priority: "low",
Title: `<script>alert(1)</script>`,
OccurredAt: time.Date(2026, 8, 1, 7, 0, 0, 0, time.UTC),
}}}
body := getEvents(t, core).Body.String()
if strings.Contains(body, "<script>alert(1)</script>") {
t.Error("intake title was not escaped")
}
if !strings.Contains(body, "&lt;script&gt;") {
t.Error("intake title is missing from the page entirely")
}
}

Some files were not shown because too many files have changed in this diff Show More