Compare commits

...

7 Commits

Author SHA1 Message Date
kami 2c1b0eede0 Read a web page when he names one, and watch a few on a timer (#259)
The network fallback behind the local sources, off unless configured.

internal/crawl is pure: a stdlib robots.txt parser (group specificity,
wildcards, Crawl-delay, cached per host), HTML-to-plaintext extraction, and a
watcher that notes a watched page only when its text changed. It has no store
access and no net/http; cmd/mavend/crawls.go is the impure half.

Every limit is code and tested: the guarded fetcher from #258 enforces the host
allowlist/denylist, refuses private addresses in the dialer Control hook (so DNS
rebinding and each redirect hop are covered), caps size and redirects, times out,
and spaces requests per host. A robots.txt Disallow is refused with no override.

On demand, reading is a query source placed last in the chain, after his memory,
his notes, and the local Kiwix ZIMs once those are wired: no URL in the
utterance means no fetch, and only the URL ever leaves the box. Scheduled
watches write notes and announce nothing.

The vendored tree has no x/net/html, goquery or temoto/robotstxt, so the parsers
are stdlib. No new dependency.
2026-08-01 03:40:21 +04:00
kami cb3641e7bb Read RSS and Atom feeds, and speak about them only when asked (#258)
internal/rss parses RSS 2.0 and Atom, and polls each configured feed on its own
interval; internal/webfetch is the one door either of them uses to touch the
network. The poller writes items as notes with source "rss:<feed>" and nothing
else: the answer path reads them back when he asks "что нового в лентах?", and
nothing is announced on arrival. A feed that dispatched would be a nag, which is
why the plan's breaking-news rule was left out rather than built.

webfetch is where the limits live, as code rather than a paragraph: http(s)
only, an allowlist (the configured feeds' hosts) and a denylist, a 2 MiB body
cap, a 3-redirect cap, one request per host per second, and a refusal to connect
to any private address — checked in the dialer's Control hook so it holds for
every resolved address and every redirect hop, not just for a literal IP.

Off unless configured: no "feeds" block, no poller, no outbound request. How far
a feed was read is a config fact (rss:latest:<name>), so a restart does not
re-note yesterday's headlines.
2026-08-01 03:27:45 +04:00
kami ee7bec11e3 Add mavmaild, the read-only IMAP poller that feeds mail intake (#246)
The extraction seam landed on the previous branch but nothing fed it. This
adds the daemon that does: every interval it opens one mailbox read-only
(EXAMINE + BODY.PEEK, so reading leaves no \Seen behind), fetches the UIDs
it has not handed over yet, and posts each message to core over
ingest_mail. Core runs the model and writes task candidates; this daemon
writes nothing and cannot create a reminder.

It is a separate daemon because of the credential. mavpoll set the
precedent with the zenmoney token (#125): the module talking to the third
party holds the secret, reads it from a file so it never lands in argv, in
docker-compose.yml or in shell history, and core never sees it. There is
deliberately no -password flag, and a test asserts that.

Off unless configured at both ends: without -password-file the daemon
refuses to start, and if core has no email block the first ingest returns
ErrUnknownMethod, which disables the reader instead of hammering a socket
that will keep refusing. A seen-UID state file (0600, atomic write) keeps a
restart from re-extracting the whole lookback window; correctness does not
depend on it, since capture dedupes on normalised text. Logs are counts and
UIDs — no subject, sender or body.

Verified with an in-process IMAP server and a fake core: bulk mail is
filtered before core is asked, seen UIDs are not re-fetched, a failed
ingest is retried next poll, ErrUnknownMethod stops at the first message,
and state survives a restart. The live half is untested by design — no IMAP
credential exists on this box; setup is written up as QA steps.

Vikunja #246
2026-08-01 03:13:18 +04:00
kami f42d1594ef Turn a mail into task candidates, and into nothing else (#246)
The extraction half. internal/email.Extractor asks the resident Qwen3-1.7B,
under a GBNF grammar, what one message requires of him, and returns at most
three short candidates with an optional date.

Everything it can produce is a row in `tasks` with status "candidate",
written through the intake seam #130 built for exactly this (Source
"email:<mailbox>", Evidence = the subject line). No reminder, no fact, no
note, no calendar event. That bound is the design: a reminder FIRES, so a
1.7B misreading "встреча была в четверг" as a future appointment would wake
him up about it, whereas a wrong candidate is a line he dismisses in one
click. A due date the model read out of the mail is stored on the candidate,
where no scheduler reads it — the review page sorts by it. Relative wording
("до пятницы") is deliberately left in the text rather than resolved to a
date the model would get wrong.

The prompt is written against the two things a small model does here: it
summarises when asked to extract, and it invents an obligation out of a
polite closing line. Hence the demand for a verb phrase, and an explicit
empty array — most mail contains no task, and a model with no way to say
"nothing" says something.

Wiring: core owns extraction because llama-server lives in core's process,
so the reader hands messages over a new ipc.MethodIngestMail. It is a Server
hook (like StepUp/UnlockFn), not a CoreAPI method — not a store operation,
and no CoreAPI implementation should have to carry it. The hook stays nil
without an `email` config block or without a llama-server phraser, so the
method answers ErrUnknownMethod: off unless configured, twice over. There is
no keyword fallback on purpose — "the subject became a task" is a mailbox
rendered as a to-do list, not extraction.

Privacy: junk is refused before the model is called, mail text is never
search input, extraction errors carry byte counts rather than the reply, the
stored evidence is a truncated subject, and the log line names the mailbox
and the UID only.
2026-08-01 03:06:55 +04:00
kami b4646155b4 Read a mailbox read-only, in a client small enough to audit (#246)
internal/email is the reading half of the email reader: a ~200-line IMAP
client (LOGIN, EXAMINE, UID SEARCH SINCE, UID FETCH BODY.PEEK, LOGOUT), a
MIME-to-plaintext converter, and a header-only junk filter.

Two protocol choices are the design, not shortcuts. EXAMINE instead of
SELECT means the session is read-only at the protocol level, so no command
in it can flip a flag or expunge anything by mistake. BODY.PEEK instead of
BODY means reading a message does not mark it \Seen — Maven reads his mail
and leaves no trace of having done so, and the unread state in his own
client stays his.

Hand-rolled rather than go-imap because this is the one path that holds his
mailbox credential and reads his private mail: five commands with no
dependencies is auditable in a sitting. No IDLE and no cleartext/STARTTLS
either — an option to send his password over a plain socket is an option to
get it wrong once.

Junk is decided by headers alone, before any model is involved:
List-Unsubscribe/List-Id, Precedence: bulk, Auto-Submitted, the spam
headers, and Gmail's own category labels. Sender lists and subject keywords
are deliberately absent — they age badly and they would put his contacts in
a config file. A junk verdict only means "do not spend the model on this";
nothing is deleted and no server flag is touched.

Nothing here logs a body, a subject or an address, the junk reason names a
header rather than content, and an undecodable charset degrades to
headers-only instead of feeding the model mojibake. Verified against
recorded .eml fixtures and an in-process fake IMAP server.
2026-08-01 02:59:24 +04:00
kami da647e87d0 Read spending from zenmoney in the poller, answer it from facts (#125)
The trust boundary is zenmoney, not maven — they already hold his bank
sessions. So the poller reads /v8/diff/ and writes totals as
facts(kind=env, source=poll:zenmoney); core reads those back when he asks
and never sees the token.

internal/zenmoney sums transactions per currency over a window, skipping
tombstoned rows and transfers between his own accounts, and refuses to
encode a summary built from zero transactions. That refusal is the whole
design: a failed or empty read writes nothing and leaves the last good
total alone, because a zero recited as fact is worse than silence. No
currency conversion either — a figure he can check against his bank beats
one he cannot.

Off unless configured, and the token is read from a FILE rather than a
flag so it never lands in `ps`, in docker-compose.yml, or in shell
history. Nothing about the money is search input, no tick rule reads the
keys, and the log lines name keys, never figures.

The live-credential half is BLOCKED: there is no zenmoney account or token
here, so everything is verified against a recorded diff fixture.
2026-08-01 02:50:27 +04:00
kami bf6ccf9aea Rank captured tasks by what he actually said (#129)
Ordering is computed, not generated. Asking a 1.7B which of his tasks
matters most produces a fluent opinion with no basis in anything, and a
confidently wrong priority is worse than none — same posture as the
behaviour profile in internal/memory, which counts instead of summarising.

internal/tasks is a pure package (no ipc, no store, no cgo) holding the
score, the order and the Russian rendering, so the spoken list and the
/tasks page cannot drift. Four signals, all of them things he stated:
deadline (overdue > today > tomorrow > this week), stated urgency, age
with a cap so nothing rots at the bottom, and confirmed work always
ahead of mail-derived candidates. A task with no due date and no weight
scores nothing and carries no reason string — inventing a "потому что"
about a priority he never set is the failure mode this avoids.

Capture now picks up urgency he says out loud ("добавь в задачи срочно
оплатить интернет"), stripping the marker from the task text, and the web
add form offers the same three rungs. Ranking is a read: it sorts and
renders, never writes, schedules or announces.
2026-08-01 02:42:02 +04:00
78 changed files with 8256 additions and 89 deletions
+5
View File
@@ -7,6 +7,7 @@
/mavpoll
/mavcaldav
/mavwaked
/mavmaild
# Certs (private keys, don't commit)
certs/
@@ -34,6 +35,10 @@ deps
deploy/db_key.env
# Deploy secret (telegram bot token + chat id) — never commit
deploy/telegram.env
# zenmoney API token, read by mavpoll (never in argv, never committed)
deploy/zenmoney.token
# IMAP password, read by mavmaild (never in argv, never committed)
deploy/imap.password
# Temp files
/tmp/
+2 -1
View File
@@ -34,7 +34,7 @@ CGO daemons (`mavend`, `mavsttd`, `mavttsd`, `mavenclient`) need the vendored to
and libs wired through the Makefile — **do not** call `go build` on them bare, use `make`:
```sh
make build # all 8 binaries
make build # all 9 binaries
make build-web # single daemon (pure-Go ones: web/waked/poll/caldav build without CGO)
make test # go test -race across ./internal/... ./cmd/... with CGO env set
```
@@ -62,6 +62,7 @@ Pure-Go packages (`router`, `memory`, `mavweb`, …) run under a plain `go test
| `mavenclient` | Voice loop client (mic → stt → core → tts). |
| `mavpoll` | Telegram long-poll reach. |
| `mavcaldav` | CalDAV calendar sync. |
| `mavmaild` | Mail reader (IMAP, read-only). Holds the IMAP password; core never sees it. |
Daemons are wired socket-to-socket, not linked. `internal/ipc` is the client/server wire
protocol; the config in `deploy/mavend.json` (with `${VAR}` env expansion from gitignored
+2 -1
View File
@@ -51,7 +51,8 @@ RUN go build -o /out/mavend ./cmd/mavend && \
go build -o /out/mavttsd ./cmd/mavttsd && \
go build -o /out/mavweb ./cmd/mavweb && \
go build -o /out/mavpoll ./cmd/mavpoll && \
go build -o /out/mavcaldav ./cmd/mavcaldav
go build -o /out/mavcaldav ./cmd/mavcaldav && \
go build -o /out/mavmaild ./cmd/mavmaild
# llama.cpp Vulkan build — the phraser/router LFM engine (llama-server). Built
# from source (not a prebuilt vendored blob) so the binary's glibc/GLIBCXX match
+5 -2
View File
@@ -20,7 +20,7 @@ PIPER_ESPEAK := $(shell pwd)/deps/piper/espeak-ng-data
all: build
build: build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav
build: build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav build-mail
build-stt:
CGO_CFLAGS="$(CGO_CFLAGS)" CGO_LDFLAGS="$(CGO_LDFLAGS)" LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
@@ -50,6 +50,9 @@ build-poll:
build-caldav:
$(GO) build $(GOFLAGS) -o mavcaldav ./cmd/mavcaldav/
build-mail:
$(GO) build $(GOFLAGS) -o mavmaild ./cmd/mavmaild/
run-web: build-web
./mavweb -addr :9200 -voice 127.0.0.1:9100
@@ -185,4 +188,4 @@ download-embedder:
@echo ' sudo cp onnxruntime-linux-x64-1.15.1/lib/libonnxruntime.so* /usr/local/lib/'
clean:
rm -f mavend mavenclient mavsttd mavttsd mavweb mavpoll mavcaldav mavwaked
rm -f mavend mavenclient mavsttd mavttsd mavweb mavpoll mavcaldav mavwaked mavmaild
+78
View File
@@ -0,0 +1,78 @@
package main
import (
"context"
"log"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/zenmoney"
)
// Money questions (Vikunja #125).
//
// This is the whole read side: mavpoll holds the zenmoney token and writes
// facts(kind=env, source=poll:zenmoney); core reads them back when he asks.
// Core never sees the token, never calls zenmoney, and has no rule on these
// keys — a total is never a reason for Maven to speak first. Maven is not a
// nag, least of all about his money.
//
// Nothing here can reach the external search capability: the figures are read
// from the store and rendered locally, and his financial data is never search
// input.
// queryMoney — "сколько я потратил сегодня?", "покажи мои траты".
//
// Answers only from the latest fact the poller wrote. Three honest outcomes and
// no fourth: the figure, "the fact is old and here is its date", or "money
// tracking is not connected". It never computes, estimates or rounds a total of
// its own — an invented number about his money is the worst thing this could do.
func (h *reactiveHandler) queryMoney(ctx context.Context, t *queryTurn) (string, bool) {
window, ok := router.ParseMoneyQuery(t.dec.Utterance)
if !ok {
return "", false
}
key, phrase := zenmoney.KeySpentMonth, "в этом месяце"
if window == router.MoneyToday {
key, phrase = zenmoney.KeySpentToday, "сегодня"
}
fact, err := h.api.LatestFactBySource(ctx, key, zenmoney.Source)
if err != nil {
// No fact at all is the normal state when the capability is off. Claim
// the turn anyway: falling through to recall would answer a question
// about money with whatever note happens to be nearest.
if !isNoFactErr(err) {
log.Printf("voice: money fact: %v", err)
}
return "я не отслеживаю траты — не подключено.", true
}
val, err := zenmoney.ParseFactValue(fact.Value)
if err != nil {
log.Printf("voice: money fact: decode: %v", err)
return "не получилось прочитать траты.", true
}
reply := val.FormatRU(phrase)
if reply == "" {
return "по тратам пока нечего сказать.", true
}
// A stale fact is reported as stale rather than spoken as today's number.
if h.now().Sub(fact.Ts) > zenmoney.StaleAfter {
return "данные от " + fact.Ts.Local().Format("02.01") + ": " + reply, true
}
return reply, true
}
// isNoFactErr — ErrNoFact survives the wire wrapped, so unwrap for it.
func isNoFactErr(err error) bool {
for e := err; e != nil; {
if e == ipc.ErrNoFact {
return true
}
u, ok := e.(interface{ Unwrap() error })
if !ok {
return false
}
e = u.Unwrap()
}
return false
}
+139
View File
@@ -0,0 +1,139 @@
package main
import (
"context"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/zenmoney"
)
// moneyAPI answers only LatestFactBySource; everything else is unimplemented,
// which is the assertion that answering a money question costs no model call
// and reaches no network.
type moneyAPI struct {
ipc.UnimplementedCoreAPI
fact ipc.Fact
err error
gotKey string
gotSrc string
callCnt int
}
func (a *moneyAPI) LatestFactBySource(_ context.Context, key, source string) (ipc.Fact, error) {
a.gotKey, a.gotSrc = key, source
a.callCnt++
return a.fact, a.err
}
func moneyNow() time.Time { return time.Date(2026, 8, 15, 20, 0, 0, 0, time.UTC) }
func moneyFact(ts time.Time, val string) ipc.Fact {
return ipc.Fact{Kind: "env", Key: zenmoney.KeySpentMonth, Value: val, Source: zenmoney.Source, Ts: ts}
}
func TestQueryMoneyAnswersFromTheFact(t *testing.T) {
api := &moneyAPI{fact: moneyFact(moneyNow(), `{"spent":[{"currency":"RUB","amount":1749.5}],"count":3}`)}
h := &reactiveHandler{api: api, now: moneyNow}
reply, ok := h.queryMoney(context.Background(), &queryTurn{
dec: router.Decision{Utterance: "сколько я потратил в этом месяце?"},
})
if !ok {
t.Fatal("the money source must claim a money question")
}
if api.gotKey != zenmoney.KeySpentMonth || api.gotSrc != zenmoney.Source {
t.Errorf("read %q/%q, want the month key from the poller's source", api.gotKey, api.gotSrc)
}
if !strings.Contains(reply, "1749.5") {
t.Errorf("reply = %q, want the exact figure", reply)
}
if !strings.Contains(reply, "в этом месяце") {
t.Errorf("reply = %q, want the window named", reply)
}
}
func TestQueryMoneyPicksTodaysKey(t *testing.T) {
api := &moneyAPI{fact: moneyFact(moneyNow(), `{"spent":[{"currency":"RUB","amount":250}],"count":1}`)}
h := &reactiveHandler{api: api, now: moneyNow}
if _, ok := h.queryMoney(context.Background(), &queryTurn{
dec: router.Decision{Utterance: "сколько я потратил сегодня?"},
}); !ok {
t.Fatal("expected the source to claim it")
}
if api.gotKey != zenmoney.KeySpentToday {
t.Errorf("key = %q, want today's", api.gotKey)
}
}
// The capability is off unless configured, and then there is no fact. She says
// so instead of letting the recall pass answer a money question from a note.
func TestQueryMoneySaysNotConnected(t *testing.T) {
h := &reactiveHandler{api: &moneyAPI{err: ipc.ErrNoFact}, now: moneyNow}
reply, ok := h.queryMoney(context.Background(), &queryTurn{
dec: router.Decision{Utterance: "сколько я потратил?"},
})
if !ok {
t.Fatal("expected the source to claim it")
}
if !strings.Contains(reply, "не подключено") {
t.Errorf("reply = %q, want an honest 'not connected'", reply)
}
// No number of any kind in that answer.
for _, d := range []string{"0", "1", "2", "3", "4", "5", "6", "7", "8", "9"} {
if strings.Contains(reply, d) {
t.Errorf("reply %q contains a digit — nothing was read, so there is no figure", reply)
}
}
}
// A fact older than the staleness bound is dated rather than spoken as if it
// were current: the poller can be down, and last week's total presented as
// today's is a lie by omission.
func TestQueryMoneyDatesAStaleFact(t *testing.T) {
old := moneyNow().Add(-72 * time.Hour)
api := &moneyAPI{fact: moneyFact(old, `{"spent":[{"currency":"RUB","amount":100}],"count":1}`)}
h := &reactiveHandler{api: api, now: moneyNow}
reply, _ := h.queryMoney(context.Background(), &queryTurn{
dec: router.Decision{Utterance: "сколько я потратил?"},
})
if !strings.Contains(reply, "данные от") {
t.Errorf("reply = %q, want the stale fact dated", reply)
}
}
func TestQueryMoneyPassesOtherQuestions(t *testing.T) {
api := &moneyAPI{}
h := &reactiveHandler{api: api, now: moneyNow}
for _, u := range []string{"какая погода?", "я потратил весь день на это", "какие у меня задачи?"} {
if _, ok := h.queryMoney(context.Background(), &queryTurn{dec: router.Decision{Utterance: u}}); ok {
t.Errorf("the money source claimed %q", u)
}
}
if api.callCnt != 0 {
t.Error("a non-money question must not read the money facts")
}
}
// Money must be answered before the recall sources, or a question about
// spending gets answered by the nearest note.
func TestQuerySourcesOrderMoneyBeforeRecall(t *testing.T) {
moneyAt, notesAt := -1, -1
for i, src := range querySources {
switch src.name {
case "money":
moneyAt = i
case "notes":
notesAt = i
}
}
if moneyAt < 0 || notesAt < 0 {
t.Fatalf("sources missing: money=%d notes=%d", moneyAt, notesAt)
}
if moneyAt > notesAt {
t.Errorf("money source at %d, after notes at %d", moneyAt, notesAt)
}
}
+127
View File
@@ -8,10 +8,12 @@ import (
"strings"
"time"
"github.com/kami/maven/internal/crawl"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/morning"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/rss"
"github.com/kami/maven/internal/weather"
)
@@ -62,11 +64,28 @@ var querySources = []querySource{
// matcher requires a task noun or an explicit "что … сделать", so a
// date-bearing question still reaches the calendar.
{"tasks", (*reactiveHandler).queryTasks},
// Before the recall sources too: "сколько я потратил?" is a question about
// the money facts the poller wrote, and the notes pass would otherwise
// answer it from whatever he once said about spending. Its matcher needs a
// money noun plus an actual ask, so "я потратил весь день" is untouched.
{"money", (*reactiveHandler).queryMoney},
// Before the recall sources and before general knowledge: "что нового?" is
// a question about the feeds she reads, and general knowledge would answer
// it by inventing news. Its matcher needs a feed noun plus an ask, so
// "у меня новая лента в инстаграме" is untouched.
{"feeds", (*reactiveHandler).queryFeeds},
{"calendar", (*reactiveHandler).queryCalendar},
{"weather", (*reactiveHandler).queryWeather},
{"embed", (*reactiveHandler).queryEmbed},
{"memory", (*reactiveHandler).queryMemory},
{"notes", (*reactiveHandler).queryNotes},
// LAST before the model answers from memory, and that position is the whole
// design (Vikunja #259): local sources first. The model, his own notes and
// facts, and — once internal/kiwix is wired into this chain — the offline
// ZIMs all get their turn before anything touches the network. This source
// only claims a turn where he named a URL out loud, so it never competes
// with a local answer.
{"web", (*reactiveHandler).queryWeb},
{"general-knowledge", (*reactiveHandler).queryGeneral},
}
@@ -175,6 +194,63 @@ func (h *reactiveHandler) queryHabits(ctx context.Context, t *queryTurn) (string
return profile.FormatOverallRU(), true
}
// feedNoteWindow — how many recent notes are scanned for feed items, and
// feedReadOut — how many headlines she actually reads back. She summarises the
// top of the pile, she does not recite a river.
const (
feedNoteWindow = 200
feedReadOut = 3
)
// queryFeeds — "что нового в лентах?", "что нового по технологиям?"
// (Vikunja #258).
//
// This is the ONLY way a feed item reaches him. The poller writes notes and
// never speaks; asking is the trigger. If that ever changes, the thing that
// changed is "Maven is not a nag", not a detail of this file.
func (h *reactiveHandler) queryFeeds(ctx context.Context, t *queryTurn) (string, bool) {
q, ok := router.ParseFeedQuery(t.dec.Utterance)
if !ok {
return "", false
}
if !h.feedsOn {
// Claim the turn rather than fall through: "не читаю ленты" is true, and
// letting general knowledge answer "что нового?" would be an invented
// news bulletin.
return "я пока не читаю ленты — они не настроены.", true
}
notes, err := h.api.RecentNotes(ctx, feedNoteWindow)
if err != nil {
log.Printf("voice: feeds: recent notes: %v", err)
return "не получилось посмотреть ленты.", true
}
var picked []string
for _, n := range notes {
if !strings.HasPrefix(n.Source, rss.SourcePrefix) {
continue
}
if !router.CategoryMatches(n.Text, q.Category) {
continue
}
// The note carries title, summary and link; she reads the title.
title := n.Text
if i := strings.IndexByte(title, '\n'); i > 0 {
title = title[:i]
}
picked = append(picked, strings.TrimSpace(title))
if len(picked) == feedReadOut {
break
}
}
if len(picked) == 0 {
if q.Category != "" {
return "по этой теме в лентах пока ничего.", true
}
return "в лентах пока ничего нового.", true
}
return "вот что нового: " + strings.Join(picked, "; "), true
}
// queryCalendar — "что у меня сегодня?", "планы на завтра?"
// h.now(), not time.Now(): the handler's clock is the injected one, so this
// source can be tested at a fixed time like the rest.
@@ -303,6 +379,57 @@ func (h *reactiveHandler) queryNotes(ctx context.Context, t *queryTurn) (string,
return reply, true
}
// webPageContextRunes — how much of a fetched page is handed to the phraser.
// Less than the crawler keeps: the rest of the 4096-token window belongs to the
// prompt, the persona block and the reply.
const webPageContextRunes = 1500
// queryWeb — "посмотри https://example.org/x — что там?" (Vikunja #259).
//
// It claims a turn ONLY when he named a URL, which is what keeps a fallback from
// becoming a habit: no URL, no fetch, and the model answers from what is local.
// What leaves the box is the URL and nothing else — no note, no fact, no history
// travels with it.
func (h *reactiveHandler) queryWeb(ctx context.Context, t *queryTurn) (string, bool) {
link, ok := router.FirstURL(t.dec.Utterance)
if !ok {
return "", false
}
if h.crawler == nil {
// Claim rather than fall through: he asked about a specific page, and
// letting the model answer from the URL's spelling alone is how a small
// model invents a page's contents.
return "я не читаю страницы — это не настроено.", true
}
ctxFetch, cancel := context.WithTimeout(ctx, 30*time.Second)
defer cancel()
page, err := h.crawler.Page(ctxFetch, link)
if err != nil {
if errors.Is(err, crawl.ErrRobots) {
return "эта страница закрыта для чтения — robots.txt не разрешает.", true
}
log.Printf("voice: web: %v", err)
return "не получилось прочитать страницу.", true
}
if page.Text == "" {
return "страница открылась, но читать там нечего.", true
}
// The page is handed to the phraser the same way a note is: as context for
// the question he actually asked. She answers the question, she does not
// recite the page.
snippet := page.Title + "\n" + crawl.TrimRunes(page.Text, webPageContextRunes)
reply, perr := h.phraser.PhraseQuery(ctx, t.dec.Utterance, []string{snippet})
if perr != nil {
log.Printf("voice: web: phrase: %v", perr)
}
if reply == "" {
// No phraser (or it failed): read back the top of the page rather than
// pretend the fetch did not happen.
return "вот что на странице: " + crawl.TrimRunes(page.Text, 300), true
}
return reply, true
}
// queryGeneral — general knowledge from the phraser, the last source before
// giving up. It always claims: either the model answers or Maven says she
// doesn't know.
+23 -39
View File
@@ -3,11 +3,11 @@ package main
import (
"context"
"log"
"strings"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/tasks"
)
// Task capture on the voice/chat path (Vikunja #130).
@@ -27,14 +27,15 @@ import (
// captureTaskFromNote claims the turn when the utterance explicitly files a
// task, returning the reply. ("", false) hands the turn back to the note path.
func (h *reactiveHandler) captureTaskFromNote(ctx context.Context, dec router.Decision) (string, bool) {
text, ok := router.ParseTaskCapture(dec.Utterance)
cap, ok := router.ParseTaskCapture(dec.Utterance)
if !ok {
return "", false
}
resp, err := h.api.CaptureTask(ctx, ipc.CaptureTaskReq{
Text: text,
Text: cap.Text,
Source: "tap:voice",
Status: store.TaskOpen, // he stated it himself — not a candidate
Weight: cap.Weight, // 0 unless he said "срочно" / "важно"
Ts: h.now(),
})
if err != nil {
@@ -44,56 +45,39 @@ func (h *reactiveHandler) captureTaskFromNote(ctx context.Context, dec router.De
if !resp.Created {
return "это уже в списке.", true
}
return "записала: " + text, true
return "записала: " + cap.Text, true
}
// queryTasks — "какие у меня задачи?", "что мне нужно сделать?".
//
// Reads the live set and recites it. Newest first, which is the order the store
// returns: this source has no opinion about which task matters more, and
// pretending otherwise would be a guess. Ranking is Vikunja #129.
// Reads the live set and recites it in priority order (Vikunja #129). The order
// is computed by internal/tasks from what he told her — deadlines, the urgency
// he stated, how long a task has been sitting — never asked of the model. The
// rendering is the package's too, so the spoken list and the /tasks page can
// never disagree about what comes first.
func (h *reactiveHandler) queryTasks(ctx context.Context, t *queryTurn) (string, bool) {
if !router.IsTaskListQuery(t.dec.Utterance) {
return "", false
}
tasks, err := h.api.ListTasks(ctx, "live")
live, err := h.api.ListTasks(ctx, "live")
if err != nil {
log.Printf("voice: list tasks: %v", err)
return "не получилось посмотреть задачи.", true
}
return formatTaskListRU(tasks), true
return tasks.FormatRU(tasks.Rank(taskItems(live), h.now())), true
}
// formatTaskListRU renders the live task list the way Maven says it. Candidates
// are named as candidates — a task she pulled out of his mail is something she
// suggests, and saying it in the same breath as work he actually stated would
// put words in his mouth.
func formatTaskListRU(tasks []ipc.Task) string {
var open, cands []string
for _, t := range tasks {
switch t.Status {
case store.TaskCandidate:
cands = append(cands, t.Text)
default:
open = append(open, t.Text)
// taskItems maps wire rows onto the ranker's input. Written here rather than in
// internal/tasks so the ranker stays a pure package with no ipc (and therefore
// no store, and therefore no cgo) dependency — the same posture as
// internal/morning and internal/memory.
func taskItems(ts []ipc.Task) []tasks.Item {
out := make([]tasks.Item, len(ts))
for i, t := range ts {
out[i] = tasks.Item{
ID: t.ID, Text: t.Text, Status: t.Status,
Created: t.CreatedTs, Due: t.Due, Weight: t.Weight,
}
}
if len(open) == 0 && len(cands) == 0 {
return "задач нет."
}
var b strings.Builder
if len(open) > 0 {
b.WriteString("в списке: ")
b.WriteString(strings.Join(open, "; "))
b.WriteString(".")
}
if len(cands) > 0 {
if b.Len() > 0 {
b.WriteString(" ")
}
b.WriteString("ещё я нашла, но ты не подтвердил: ")
b.WriteString(strings.Join(cands, "; "))
b.WriteString(".")
}
return b.String()
return out
}
+39
View File
@@ -141,6 +141,45 @@ func TestQueryTasksRecitesTheLiveList(t *testing.T) {
}
}
// The stated urgency rides through capture as a weight, so the ranker can use
// it later (Vikunja #129). "срочно" is not part of the task text.
func TestCaptureTaskCarriesStatedUrgency(t *testing.T) {
api := &taskAPI{created: true}
h := taskHandler(api)
if _, ok := h.captureTaskFromNote(context.Background(), router.Decision{
Utterance: "добавь в задачи срочно оплатить интернет",
}); !ok {
t.Fatal("expected a capture")
}
got := api.captured[0]
if got.Text != "оплатить интернет" {
t.Errorf("text = %q, want the urgency word out of the task", got.Text)
}
if got.Weight == 0 {
t.Error("weight = 0 — he said срочно and it was dropped")
}
}
// The recital is ordered by the ranker, not by insertion: a deadline he named
// comes before undated work.
func TestQueryTasksRecitesInPriorityOrder(t *testing.T) {
due := taskNow()
api := &taskAPI{tasks: []ipc.Task{
{ID: 1, Text: "купить молоко", Status: "open", CreatedTs: taskNow()},
{ID: 2, Text: "оплатить интернет", Status: "open", CreatedTs: taskNow(), Due: &due},
}}
h := taskHandler(api)
reply, _ := h.queryTasks(context.Background(), &queryTurn{
dec: router.Decision{Utterance: "какие у меня задачи?"},
})
if strings.Index(reply, "оплатить интернет") > strings.Index(reply, "купить молоко") {
t.Errorf("reply = %q, want the dated task first", reply)
}
if !strings.Contains(reply, "сегодня") {
t.Errorf("reply = %q, want the reason named", reply)
}
}
func TestQueryTasksEmptyList(t *testing.T) {
h := taskHandler(&taskAPI{})
reply, ok := h.queryTasks(context.Background(), &queryTurn{
+184
View File
@@ -0,0 +1,184 @@
// mavend/crawls.go — the driver for reading web pages (Vikunja #259,
// docs/plans/14-web-crawler.md). The crawler is pure and lives in
// internal/crawl; this is the impure half: the guarded fetcher, a ticker for the
// scheduled watches, and the fact-backed dedup hashes.
//
// Two paths, one config block, both off unless configured:
//
// - ON DEMAND — he names a URL out loud and she reads it. That is the
// `queryWeb` source in actions_query.go, LAST in the chain: after his
// memory, after the notes, and (once Kiwix is wired into the chain) after
// the local ZIMs. A local read costs nothing and leaks nothing; a fetch puts
// a URL in someone's log, so it goes last.
// - SCHEDULED — a watched page is re-read on its interval, and a page whose
// text changed is written as a note. It does NOT announce itself. Same rule
// as the feed poller: notes, never nudges.
//
// Only the URL goes out. Nothing here reads a note, a fact, the persona block or
// the history, and internal/crawl has no access to the store at all.
package main
import (
"context"
"log"
"net/url"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/crawl"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/webfetch"
)
// newCrawler builds the crawler from the `crawl` block, or returns nil when
// there is none. Every caller checks for nil, and nil means no page is ever
// fetched.
func newCrawler(cfg *config.Config) *crawl.Crawler {
if cfg.Crawl == nil {
return nil
}
cc := cfg.Crawl
hosts := append([]string(nil), cc.AllowHosts...)
// A watched page's own host is always reachable; otherwise an allowlist and
// a watch list would have to be kept in sync by hand.
for _, w := range cc.Watches {
if u, err := url.Parse(w.URL); err == nil && u.Hostname() != "" {
hosts = append(hosts, u.Hostname())
}
}
// An allowlist plus on-demand is a contradiction worth logging rather than
// silently resolving: he asked for arbitrary pages AND for a fixed list.
// The allowlist wins, because it is the narrower instruction.
if len(hosts) > 0 && cc.OnDemand && len(cc.AllowHosts) > 0 {
log.Printf("crawl: allow_hosts is set, so on-demand reading is limited to those hosts")
}
ua := cc.UserAgent
if ua == "" {
ua = webfetch.DefaultUserAgent
}
fetcher := webfetch.New(webfetch.Config{
AllowHosts: hosts,
DenyHosts: cc.DenyHosts,
Timeout: time.Duration(cc.Timeout),
MaxBytes: cc.MaxBytes,
UserAgent: ua,
})
// The user-agent handed to the crawler is the one the fetcher sends: obeying
// robots rules written for a different name would be a lie.
return crawl.New(&crawlFetcher{f: fetcher}, crawl.Config{
UserAgent: ua,
MaxRunes: cc.MaxRunes,
})
}
// onDemandCrawler returns a crawler for the answer path, or nil when on-demand
// reading is off. The scheduled watches can be on while this is off: reading a
// fixed list of pages on a timer and reading whatever URL is in an utterance are
// different permissions, and the config keeps them separate.
func onDemandCrawler(cfg *config.Config) *crawl.Crawler {
if cfg.Crawl == nil || !cfg.Crawl.OnDemand {
return nil
}
return newCrawler(cfg)
}
// crawlWorker — ticker + watcher for the scheduled half.
type crawlWorker struct {
watcher *crawl.Watcher
interval time.Duration
}
// crawlTickInterval — how often the worker asks what is due. Per-watch cadence
// is the watcher's business.
const crawlTickInterval = 15 * time.Minute
// newCrawlWorker wires the scheduled crawls, or nil when nothing is watched.
func newCrawlWorker(c *crawl.Crawler, api ipc.CoreAPI, emb router.Embedder, cfg *config.Config) *crawlWorker {
if c == nil || cfg.Crawl == nil || len(cfg.Crawl.Watches) == 0 {
return nil
}
watches := make([]crawl.WatchConfig, 0, len(cfg.Crawl.Watches))
for _, w := range cfg.Crawl.Watches {
watches = append(watches, crawl.WatchConfig{
Name: w.Name,
URL: w.URL,
Interval: time.Duration(w.Interval),
})
}
watcher := crawl.NewWatcher(c, watches, api, &factHashes{api: api},
crawlEmbedder(emb), time.Duration(cfg.Crawl.Interval))
if watcher == nil {
log.Printf("crawl: configured but nothing watchable — scheduled crawls disabled")
return nil
}
log.Printf("crawl: watching %d page(s), checking what is due every %s", len(watches), crawlTickInterval)
return &crawlWorker{watcher: watcher, interval: crawlTickInterval}
}
// run checks what is due until ctx is canceled. The first round runs
// immediately; it writes notes only, so an early round startles nobody.
func (w *crawlWorker) run(ctx context.Context) {
w.watcher.CheckDue(ctx, time.Now())
t := time.NewTicker(w.interval)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case now := <-t.C:
w.watcher.CheckDue(ctx, now)
}
}
}
// crawlFetcher adapts webfetch to crawl.Fetcher, which is the seam that keeps
// net/http out of the crawler package.
type crawlFetcher struct{ f *webfetch.Fetcher }
func (a *crawlFetcher) Get(ctx context.Context, u string) (*crawl.Response, error) {
resp, err := a.f.Get(ctx, u)
if err != nil {
return nil, err
}
return &crawl.Response{URL: resp.URL, ContentType: resp.ContentType, Body: resp.Body}, nil
}
// factHashes stores each watch's last content hash as a config fact, so a
// restart does not re-note an unchanged page. Same mechanism the feed reader
// uses for its marks, and inspectable on /dash.
type factHashes struct{ api ipc.CoreAPI }
func hashKey(name string) string { return "crawl:hash:" + name }
func (h *factHashes) LastHash(ctx context.Context, name string) (string, error) {
f, err := h.api.LatestFact(ctx, hashKey(name))
if err != nil {
// No hash yet is not an error: the watcher treats "" as "never read".
return "", nil
}
return f.Value, nil
}
func (h *factHashes) SetHash(ctx context.Context, name, hash string) error {
_, err := h.api.WriteFact(ctx, ipc.WriteFactReq{
Ts: time.Now(),
Kind: "config",
Key: hashKey(name),
Value: hash,
Source: "poll:crawl",
Confidence: 1.0,
})
return err
}
// crawlEmbedder adapts router.Embedder for the watcher, embedding with
// EmbedPassage (a page is text being searched FOR, and the e5 embedder is
// asymmetric).
func crawlEmbedder(emb router.Embedder) crawl.Embedder {
if emb == nil {
return nil
}
return passageEmbedder{emb}
}
+186
View File
@@ -0,0 +1,186 @@
package main
import (
"context"
"net/http"
"net/http/httptest"
"strings"
"testing"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/crawl"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/voice"
)
// The default config reads nothing. This is the whole "off unless configured"
// contract for the crawler, asserted at the wiring level rather than trusted.
func TestCrawlOffByDefault(t *testing.T) {
cfg := &config.Config{}
if c := newCrawler(cfg); c != nil {
t.Error("newCrawler with no crawl block returned a crawler")
}
if c := onDemandCrawler(cfg); c != nil {
t.Error("onDemandCrawler with no crawl block returned a crawler")
}
if w := newCrawlWorker(nil, nil, nil, cfg); w != nil {
t.Error("newCrawlWorker with no crawl block returned a worker")
}
// Watches configured but on_demand off ⇒ the answer path still reads
// nothing: a timer over a fixed list is not permission for arbitrary URLs.
withWatch := &config.Config{Crawl: &config.CrawlConfig{
Watches: []config.CrawlWatchConfig{{Name: "p", URL: "https://example.org/p"}},
}}
if c := onDemandCrawler(withWatch); c != nil {
t.Error("onDemandCrawler honoured a watch list as on-demand permission")
}
if c := newCrawler(withWatch); c == nil {
t.Error("newCrawler returned nil for a configured watch")
}
}
// The wired fetcher must refuse a private address, because the crawler on this
// box sits one hop from the whole homelab. Same guard the webfetch tests cover;
// this asserts the daemon actually wires it.
func TestCrawlerRefusesPrivateAddress(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "text/html")
w.Write([]byte("<html><body>secret</body></html>"))
}))
defer srv.Close()
c := newCrawler(&config.Config{Crawl: &config.CrawlConfig{OnDemand: true}})
if c == nil {
t.Fatal("newCrawler returned nil for an on-demand config")
}
if _, err := c.Page(context.Background(), srv.URL); err == nil {
t.Fatalf("reading %s succeeded; a loopback address must be refused", srv.URL)
}
}
func TestFactHashesRoundTrip(t *testing.T) {
ctx := context.Background()
st := newTestStore(t)
h := &factHashes{api: ipc.NewStoreAPI(st)}
got, err := h.LastHash(ctx, "page")
if err != nil {
t.Fatalf("LastHash on a fresh store: %v", err)
}
if got != "" {
t.Errorf("LastHash = %q, want empty for a never-read page", got)
}
if err := h.SetHash(ctx, "page", "deadbeef"); err != nil {
t.Fatalf("SetHash: %v", err)
}
got, err = h.LastHash(ctx, "page")
if err != nil {
t.Fatalf("LastHash: %v", err)
}
if got != "deadbeef" {
t.Errorf("LastHash = %q, want deadbeef", got)
}
if key := hashKey("page"); key != "crawl:hash:page" {
t.Errorf("hashKey = %q", key)
}
}
// stubCrawlFetcher serves one fixed page to every URL, so queryWeb can be
// exercised without a network or an allowlist.
type stubCrawlFetcher struct{ body, ctype string }
func (s *stubCrawlFetcher) Get(_ context.Context, u string) (*crawl.Response, error) {
ct := s.ctype
if ct == "" {
ct = "text/html"
}
if strings.HasSuffix(u, "/robots.txt") {
return &crawl.Response{URL: u, ContentType: "text/plain", Body: []byte("")}, nil
}
return &crawl.Response{URL: u, ContentType: ct, Body: []byte(s.body)}, nil
}
func buildWebHandler(c *crawl.Crawler) *reactiveHandler {
return &reactiveHandler{
replier: voice.NewStubReplier(),
phraser: phraser.NewStub(),
crawler: c,
}
}
func askWeb(h *reactiveHandler, q string) (string, bool) {
return h.queryWeb(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: q},
})
}
func TestQueryWebPassesWithoutAURL(t *testing.T) {
h := buildWebHandler(crawl.New(&stubCrawlFetcher{body: "<html><body>x</body></html>"}, crawl.Config{}))
if reply, ok := askWeb(h, "почему небо синее?"); ok {
t.Errorf("the web source claimed a question with no URL: %q", reply)
}
}
// Not configured is said out loud rather than falling through, so a small model
// never invents a page's contents from its URL.
func TestQueryWebSaysWhenNotConfigured(t *testing.T) {
h := buildWebHandler(nil)
reply, ok := askWeb(h, "посмотри https://example.org/page")
if !ok {
t.Fatal("the web source did not claim a question with a URL")
}
if !strings.Contains(reply, "не настроено") {
t.Errorf("reply = %q, want the not-configured answer", reply)
}
}
func TestQueryWebReadsThePage(t *testing.T) {
h := buildWebHandler(crawl.New(&stubCrawlFetcher{
body: "<html><head><title>Заголовок</title></head><body><p>текст страницы</p></body></html>",
}, crawl.Config{}))
reply, ok := askWeb(h, "посмотри https://example.org/page — что там?")
if !ok {
t.Fatal("the web source did not claim a question with a URL")
}
if !strings.Contains(reply, "текст страницы") {
t.Errorf("reply = %q, want the page text read back", reply)
}
}
func TestQueryWebRefusesNonHTML(t *testing.T) {
h := buildWebHandler(crawl.New(&stubCrawlFetcher{
body: "\x00\x01binary", ctype: "application/octet-stream",
}, crawl.Config{}))
reply, ok := askWeb(h, "почитай https://example.org/blob.bin")
if !ok {
t.Fatal("the web source did not claim a question with a URL")
}
if !strings.Contains(reply, "не получилось") {
t.Errorf("reply = %q, want the read-failed answer", reply)
}
}
// robots.txt is honoured on the answer path too, and she says so instead of
// reporting a generic failure.
func TestQueryWebObeysRobots(t *testing.T) {
h := buildWebHandler(crawl.New(&robotsDenyFetcher{}, crawl.Config{}))
reply, ok := askWeb(h, "посмотри https://example.org/private")
if !ok {
t.Fatal("the web source did not claim a question with a URL")
}
if !strings.Contains(reply, "robots.txt") {
t.Errorf("reply = %q, want the robots answer", reply)
}
}
type robotsDenyFetcher struct{}
func (robotsDenyFetcher) Get(_ context.Context, u string) (*crawl.Response, error) {
if strings.HasSuffix(u, "/robots.txt") {
return &crawl.Response{URL: u, ContentType: "text/plain",
Body: []byte("User-agent: *\nDisallow: /private\n")}, nil
}
return &crawl.Response{URL: u, ContentType: "text/html", Body: []byte("<html>nope</html>")}, nil
}
+182
View File
@@ -0,0 +1,182 @@
// mavend/feeds.go — the driver for RSS/Atom reading (Vikunja #258,
// docs/plans/13-rss-news-feeds.md). The reader itself is pure and lives in
// internal/rss; this is the impure half: a ticker, the guarded fetcher, and the
// two adapters that let a pure package talk to the store.
//
// Why in-core rather than its own daemon like mavmaild and mavpoll: those two
// hold a CREDENTIAL (an IMAP password, a zenmoney token), and the reason they
// are separate processes is that core must never see it. A feed URL is public,
// there is no secret to isolate, and a whole extra binary and compose service
// would buy nothing. The other half of the mavpoll precedent — off unless
// configured — is kept: no `feeds` block, no poller, no outbound request.
//
// It is its own goroutine, not a step on the tick: the tick has a delivery
// deadline behind it, and a feed read is a network round-trip that nobody is
// waiting on.
//
// Nothing here dispatches. A feed that announced itself would be a nag, so the
// only output is notes with source "rss:<feed>", which the answer path reads
// when he asks ("что нового в лентах?" — see queryFeeds in actions_query.go).
package main
import (
"context"
"log"
"net/url"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/rss"
"github.com/kami/maven/internal/webfetch"
)
// feedWorker — ticker + poller.
type feedWorker struct {
poller *rss.Poller
interval time.Duration
}
// feedTickInterval — how often the worker asks the poller what is due. Per-feed
// cadence is the poller's business; this is just the granularity.
const feedTickInterval = 5 * time.Minute
// newFeedWorker wires feed reading, or returns nil when it must not run:
// no `feeds` block (the normal case), or nothing valid in it. Every caller
// checks for nil.
func newFeedWorker(api ipc.CoreAPI, emb router.Embedder, cfg *config.Config) *feedWorker {
if cfg.Feeds == nil {
return nil
}
fc := cfg.Feeds
feeds := make([]rss.FeedConfig, 0, len(fc.Sources))
hosts := append([]string(nil), fc.AllowHosts...)
for _, s := range fc.Sources {
feeds = append(feeds, rss.FeedConfig{
Name: s.Name,
URL: s.URL,
Category: s.Category,
Interval: time.Duration(s.Interval),
Include: s.Include,
Exclude: s.Exclude,
})
// Each configured feed's own host is allowed. The allowlist is then
// exactly "the feeds he asked for", so a redirect off to somewhere else
// is refused by the fetcher rather than followed.
if u, err := url.Parse(s.URL); err == nil && u.Hostname() != "" {
hosts = append(hosts, u.Hostname())
}
}
fetcher := webfetch.New(webfetch.Config{
AllowHosts: hosts,
Timeout: time.Duration(fc.Timeout),
MaxBytes: fc.MaxBytes,
})
poller := rss.NewPoller(feeds, &feedFetcher{f: fetcher}, api, &factMarks{api: api},
embedderFor(emb), nil, rss.Config{
DefaultInterval: time.Duration(fc.PollInterval),
MaxItems: fc.MaxItems,
MaxAge: time.Duration(fc.MaxAge),
})
if poller == nil {
log.Printf("feeds: configured but nothing pollable — feed reading disabled")
return nil
}
log.Printf("feeds: reading %d feed(s), checking what is due every %s", len(feeds), feedTickInterval)
return &feedWorker{poller: poller, interval: feedTickInterval}
}
// run polls what is due until ctx is canceled. The first round runs immediately
// so a restart does not blind her for the first interval; it writes notes only,
// so an early round cannot startle anyone.
func (w *feedWorker) run(ctx context.Context) {
w.poller.PollDue(ctx, time.Now())
t := time.NewTicker(w.interval)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case now := <-t.C:
w.poller.PollDue(ctx, now)
}
}
}
// embedderOf — the voice wiring's embedder, or nil when voice is not wired.
// Feed notes are embedded with the SAME model the rest of the store uses, or not
// at all; a second embedder would write vectors nothing can search.
func embedderOf(w *voiceWiring) router.Embedder {
if w == nil {
return nil
}
return w.embedder
}
// feedFetcher adapts webfetch to rss.Fetcher — the pure package names the two
// fields it needs and stays free of net/http.
type feedFetcher struct{ f *webfetch.Fetcher }
func (a *feedFetcher) Get(ctx context.Context, u string) (*rss.Body, error) {
resp, err := a.f.Get(ctx, u)
if err != nil {
return nil, err
}
return &rss.Body{Bytes: resp.Body}, nil
}
// factMarks stores "how far this feed was read" as a config fact, the same
// mechanism the plan named and the same one the pattern tick uses for its own
// bookkeeping. Durable, inspectable on /dash, and cheap.
type factMarks struct{ api ipc.CoreAPI }
func markKey(feed string) string { return "rss:latest:" + feed }
func (m *factMarks) LastMark(ctx context.Context, feed string) (time.Time, error) {
f, err := m.api.LatestFact(ctx, markKey(feed))
if err != nil {
// No mark yet is not an error worth propagating: the poller treats a
// zero time as a cold start.
return time.Time{}, nil
}
t, err := time.Parse(time.RFC3339, f.Value)
if err != nil {
return time.Time{}, nil
}
return t, nil
}
func (m *factMarks) SetMark(ctx context.Context, feed string, at time.Time) error {
_, err := m.api.WriteFact(ctx, ipc.WriteFactReq{
Ts: time.Now(),
Kind: "config",
Key: markKey(feed),
Value: at.UTC().Format(time.RFC3339),
Source: "poll:rss",
Confidence: 1.0,
})
return err
}
// embedderFor adapts router.Embedder to rss.Embedder, and returns nil when
// there is none — a note without a vector is still a note the recent-notes path
// can read.
//
// EmbedPassage, not Embed: a feed item is text being searched FOR, and the e5
// embedder is asymmetric. Getting this backwards makes the item unfindable by
// the question that should have matched it.
func embedderFor(emb router.Embedder) rss.Embedder {
if emb == nil {
return nil
}
return passageEmbedder{emb}
}
type passageEmbedder struct{ e router.Embedder }
func (p passageEmbedder) Embed(ctx context.Context, text string) ([]float32, error) {
return router.EmbedPassage(ctx, p.e, text)
}
+177
View File
@@ -0,0 +1,177 @@
package main
import (
"context"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/router"
"github.com/kami/maven/internal/rss"
"github.com/kami/maven/internal/voice"
)
// buildFeedHandler — a handler with the given feed notes already stored. No
// embedder: the feed source answers from recent notes by source, which is what
// makes it work for notes written before an embedder existed.
func buildFeedHandler(t *testing.T, feedsOn bool, notes ...ipc.Note) *reactiveHandler {
t.Helper()
ctx := context.Background()
st := newTestStore(t)
now := time.Now()
for i, n := range notes {
ts := now.Add(time.Duration(i) * time.Minute)
if _, err := st.WriteNote(ctx, ts, n.Text, nil, n.Source); err != nil {
t.Fatalf("WriteNote: %v", err)
}
}
return &reactiveHandler{
api: ipc.NewStoreAPI(st),
replier: voice.NewStubReplier(),
phraser: phraser.NewStub(),
now: func() time.Time { return now },
feedsOn: feedsOn,
embedder: nil,
}
}
func askFeeds(t *testing.T, h *reactiveHandler, q string) (string, bool) {
t.Helper()
return h.queryFeeds(context.Background(), &queryTurn{
dec: router.Decision{Intent: router.IntentQuery, Utterance: q},
})
}
func TestQueryFeedsReadsFeedNotes(t *testing.T) {
h := buildFeedHandler(t, true,
ipc.Note{Text: "Новая уязвимость в ядре [технологии]\nпатч вышел\nhttps://example.org/a", Source: "rss:habr"},
ipc.Note{Text: "что-то он сам сказал", Source: "tap:voice"},
)
reply, ok := askFeeds(t, h, "что нового в лентах?")
if !ok {
t.Fatal("the feed source did not claim the question")
}
if !strings.Contains(reply, "уязвимость") {
t.Errorf("reply = %q, want the headline", reply)
}
if strings.Contains(reply, "он сам сказал") {
t.Errorf("a note he dictated leaked into the feed answer: %q", reply)
}
// She reads the headline, not the summary and not the URL.
if strings.Contains(reply, "https://") || strings.Contains(reply, "патч вышел") {
t.Errorf("reply = %q, want the title line only", reply)
}
}
func TestQueryFeedsByCategory(t *testing.T) {
h := buildFeedHandler(t, true,
ipc.Note{Text: "Релиз ядра [технологии]", Source: "rss:habr"},
ipc.Note{Text: "Выборы отложены [политика]", Source: "rss:news"},
)
reply, ok := askFeeds(t, h, "что нового по технологиям?")
if !ok {
t.Fatal("not claimed")
}
if !strings.Contains(reply, "ядра") || strings.Contains(reply, "Выборы") {
t.Fatalf("reply = %q, want only the технологии item", reply)
}
reply, _ = askFeeds(t, h, "что нового по спорту?")
if !strings.Contains(reply, "ничего") {
t.Fatalf("reply = %q, want an honest empty answer for an unread category", reply)
}
}
// "не настроены" and "ничего нового" are different truths, and neither may be
// answered by the model inventing a bulletin.
func TestQueryFeedsOffAndEmptyDiffer(t *testing.T) {
off := buildFeedHandler(t, false)
reply, ok := askFeeds(t, off, "что нового?")
if !ok || !strings.Contains(reply, "не настроены") {
t.Fatalf("feeds off: reply = %q, ok = %v", reply, ok)
}
on := buildFeedHandler(t, true)
reply, ok = askFeeds(t, on, "что нового?")
if !ok || !strings.Contains(reply, "ничего нового") {
t.Fatalf("feeds on but empty: reply = %q, ok = %v", reply, ok)
}
}
func TestQueryFeedsPassesOnANonFeedQuestion(t *testing.T) {
h := buildFeedHandler(t, true)
if reply, ok := askFeeds(t, h, "напомни полить цветы"); ok {
t.Fatalf("claimed an unrelated question with %q", reply)
}
}
// The mark is what stops a restart from re-noting yesterday's headlines, so the
// fact round-trip is worth a test of its own.
func TestFactMarksRoundTrip(t *testing.T) {
st := newTestStore(t)
m := &factMarks{api: ipc.NewStoreAPI(st)}
ctx := context.Background()
at, err := m.LastMark(ctx, "habr")
if err != nil || !at.IsZero() {
t.Fatalf("no mark yet: got %v, %v — want zero time and no error", at, err)
}
want := time.Date(2026, 7, 28, 10, 0, 0, 0, time.UTC)
if err := m.SetMark(ctx, "habr", want); err != nil {
t.Fatal(err)
}
got, err := m.LastMark(ctx, "habr")
if err != nil {
t.Fatal(err)
}
if !got.Equal(want) {
t.Fatalf("mark = %v, want %v", got, want)
}
}
// Off unless configured, checked at the wiring seam: no `feeds` block ⇒ no
// worker ⇒ no outbound request is possible.
func TestNewFeedWorkerOffByDefault(t *testing.T) {
st := newTestStore(t)
api := ipc.NewStoreAPI(st)
if w := newFeedWorker(api, nil, &config.Config{}); w != nil {
t.Fatal("a config with no feeds block wired a feed worker")
}
// An empty sources list is normalised to "off" by config.Load; the worker
// refuses it too, so a hand-built Config cannot switch it on by accident.
if w := newFeedWorker(api, nil, &config.Config{Feeds: &config.FeedsConfig{}}); w != nil {
t.Fatal("an empty sources list wired a feed worker")
}
cfg := &config.Config{Feeds: &config.FeedsConfig{Sources: []config.FeedSourceConfig{
{Name: "habr", URL: "https://example.org/rss"},
}}}
w := newFeedWorker(api, nil, cfg)
if w == nil {
t.Fatal("a configured feed did not wire a worker")
}
if got := w.poller.Feeds(); len(got) != 1 || got[0].Name != "habr" {
t.Fatalf("feeds = %+v", got)
}
}
// The fetcher the worker builds must be allowlisted to the configured feeds and
// nothing else — the crawler's SSRF guards are only worth as much as the
// allowlist handed to them.
func TestFeedWorkerFetcherIsAllowlisted(t *testing.T) {
cfg := &config.Config{Feeds: &config.FeedsConfig{Sources: []config.FeedSourceConfig{
{Name: "habr", URL: "https://feeds.example.org/rss"},
}}}
w := newFeedWorker(ipc.NewStoreAPI(newTestStore(t)), nil, cfg)
if w == nil {
t.Fatal("no worker")
}
// PollFeed goes through the guarded fetcher; a feed URL pointing at the box
// itself must fail rather than be read.
_, err := w.poller.PollFeed(context.Background(), rss.FeedConfig{
Name: "evil", URL: "http://127.0.0.1:9100/mcp",
}, time.Now())
if err == nil {
t.Fatal("the poller fetched a private address")
}
}
+158
View File
@@ -0,0 +1,158 @@
// mavend/mail.go — core's half of the email reader (Vikunja #246,
// docs/plans/01-email-reader.md).
//
// The split: cmd/mavmaild holds the IMAP credential, connects to the mailbox
// and converts messages to plaintext; it hands each message to core over
// ipc.MethodIngestMail. Core runs the extraction on the resident model —
// llama-server lives in this process, spawned by the phraser — and writes what
// comes back through the one task intake seam.
//
// What this file may produce is exactly one thing: rows in `tasks` with status
// "candidate". No fact, no reminder, no note, no nudge, no calendar event. A
// 1.7B misreading a mail can therefore put a wrong line on a review page and
// nothing else; it can never make Maven speak, and it can never make her
// recite something out of an advert as true.
//
// Off unless configured twice over: no `email` block in mavend.json ⇒ the IPC
// method does not exist; no llama-server phraser ⇒ same. A reader pointed at a
// core that is not set up for mail gets ErrUnknownMethod rather than silence.
package main
import (
"context"
"fmt"
"log"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/email"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/store"
)
// evidenceMaxChars — how much of the subject line is kept as a candidate's
// evidence. Enough to recognise the mail on /tasks, not enough to turn the task
// list into a copy of his mailbox.
const evidenceMaxChars = 160
// mailIntake — extraction + capture for one message at a time.
type mailIntake struct {
st *store.Store
ex *email.Extractor
timeout time.Duration
now func() time.Time
}
// newMailIntake returns nil when mail ingestion must not be available, which is
// the default. Both preconditions are real:
//
// - no cfg.Email ⇒ not configured, and a capability is off unless configured;
// - no llama-server phraser ⇒ nothing to extract with. There is deliberately
// no keyword fallback: "the subject line became a task" is not extraction,
// it is a mailbox rendered as a to-do list, and it would fill the review
// page faster than he could clear it.
func newMailIntake(st *store.Store, phr phraser.Phraser, cfg *config.Config) *mailIntake {
if cfg.Email == nil {
return nil
}
lp, ok := phr.(*phraser.LLMPhraser)
if !ok {
log.Printf("mail intake: configured but no llama-server phraser — mail ingestion disabled")
return nil
}
timeout := time.Duration(cfg.Email.Timeout)
if timeout <= 0 {
timeout = config.DefaultEmailTimeout
}
ex := email.NewExtractor(llm.New(lp.BaseURL(), timeout), cfg.Email.MaxTasks, contextBlockFn(cfg, time.Now))
log.Printf("mail intake: enabled (max %d candidates per message, timeout %s)", cfg.Email.MaxTasks, timeout)
return &mailIntake{st: st, ex: ex, timeout: timeout, now: time.Now}
}
// ingest handles one ipc.MethodIngestMail call.
//
// Junk and empty messages are answered Skipped without touching the model — the
// reader's header filter is what keeps the resident model off newsletters.
//
// Every candidate is captured with Status "candidate", Source "email:<mailbox>"
// and the subject as Evidence. CaptureTask dedupes on normalised text among
// live rows, so a mailbox re-read after a restart produces Created=0 rather
// than a second copy of every task.
func (m *mailIntake) ingest(ctx context.Context, req ipc.IngestMailReq) (ipc.IngestMailResp, error) {
msg := email.Message{
UID: req.UID,
From: req.From,
Subject: req.Subject,
Date: req.Date,
Body: req.Body,
Junk: req.Junk,
}
if msg.Junk || (msg.Subject == "" && msg.Body == "") {
return ipc.IngestMailResp{Skipped: true}, nil
}
ctx, cancel := context.WithTimeout(ctx, m.timeout)
defer cancel()
cands, err := m.ex.Extract(ctx, msg)
if err != nil {
// The error from internal/email never carries mail text; keep it that way
// by not adding the subject here.
return ipc.IngestMailResp{}, fmt.Errorf("mail intake: uid %d: %w", req.UID, err)
}
if len(cands) == 0 {
return ipc.IngestMailResp{}, nil
}
source := email.SourcePrefix + req.Mailbox
evidence := truncateRunes(req.Subject, evidenceMaxChars)
now := m.now()
var resp ipc.IngestMailResp
for _, c := range cands {
t := store.Task{
CreatedTs: now,
Text: c.Text,
Source: source,
Evidence: evidence,
// The one status this path may ever write. Anything Maven derived from
// something she read is a suggestion until he confirms it on /tasks.
Status: store.TaskCandidate,
}
if due, ok := email.ParseDue(c.Due); ok {
t.Due = &due
}
id, created, err := m.st.CaptureTask(ctx, t)
if err != nil {
return resp, fmt.Errorf("mail intake: capture: %w", err)
}
resp.TaskIDs = append(resp.TaskIDs, id)
if created {
resp.Created++
}
}
// Counts only: the log line names the mailbox and the UID, never the subject,
// the sender or the task text. Reviewing a candidate is what /tasks is for.
log.Printf("mail intake: %s uid %d → %d candidate(s), %d new", source, req.UID, len(cands), resp.Created)
return resp, nil
}
// wireMailIntake installs the IPC hook, or leaves it nil so the method reports
// ErrUnknownMethod. Called on both startup paths (unlocked boot and passkey
// unlock) so mail behaves the same either way.
func wireMailIntake(srv *ipc.Server, st *store.Store, phr phraser.Phraser, cfg *config.Config) {
mi := newMailIntake(st, phr, cfg)
if mi == nil {
return
}
srv.IngestMailFn = mi.ingest
}
// truncateRunes cuts a string to n runes, marking the cut.
func truncateRunes(s string, n int) string {
r := []rune(s)
if len(r) <= n {
return s
}
return string(r[:n]) + "…"
}
+182
View File
@@ -0,0 +1,182 @@
package main
import (
"context"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/config"
"github.com/kami/maven/internal/email"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/store"
)
// mailLLM — a canned extraction reply.
type mailLLM struct {
reply string
calls int
}
func (m *mailLLM) Complete(_ context.Context, _ llm.Req) (string, error) {
m.calls++
return m.reply, nil
}
func newTestIntake(t *testing.T, reply string) (*mailIntake, *store.Store, *mailLLM) {
t.Helper()
st := newTestStore(t)
fake := &mailLLM{reply: reply}
return &mailIntake{
st: st,
ex: email.NewExtractor(fake, 0, nil),
timeout: 5 * time.Second,
now: func() time.Time { return time.Date(2026, 8, 1, 10, 0, 0, 0, time.UTC) },
}, st, fake
}
func ingestReq() ipc.IngestMailReq {
return ipc.IngestMailReq{
Mailbox: "INBOX", UID: 42,
From: "billing@isp.example",
Subject: "Счёт за интернет",
Body: "Оплатите счёт до 5 августа.",
}
}
// The one property that matters: a mail-derived task is a candidate, attributed
// to the mailbox, with the subject as reviewable evidence — and nothing else is
// written.
func TestIngestCapturesCandidates(t *testing.T) {
mi, st, _ := newTestIntake(t, `[{"text":"оплатить счёт за интернет","due":"2026-08-05"}]`)
resp, err := mi.ingest(context.Background(), ingestReq())
if err != nil {
t.Fatalf("ingest: %v", err)
}
if resp.Created != 1 || len(resp.TaskIDs) != 1 {
t.Fatalf("resp = %+v, want one created task", resp)
}
tasks, err := st.ListTasks(context.Background(), "")
if err != nil {
t.Fatalf("list: %v", err)
}
if len(tasks) != 1 {
t.Fatalf("got %d tasks, want 1", len(tasks))
}
got := tasks[0]
if got.Status != store.TaskCandidate {
t.Errorf("status = %q, want %q — mail may only produce candidates", got.Status, store.TaskCandidate)
}
if got.Source != "email:INBOX" {
t.Errorf("source = %q, want email:INBOX", got.Source)
}
if got.Evidence != "Счёт за интернет" {
t.Errorf("evidence = %q, want the subject line", got.Evidence)
}
if got.Due == nil || got.Due.Format("2006-01-02") != "2026-08-05" {
t.Errorf("due = %v, want 2026-08-05", got.Due)
}
// Nothing else may have been written: no reminder, no fact.
rem, err := st.ListReminders(context.Background(), 10)
if err != nil {
t.Fatalf("list reminders: %v", err)
}
if len(rem) != 0 {
t.Errorf("mail created %d reminders; a misread mail must never be able to fire", len(rem))
}
}
// Re-reading a mailbox must not grow the list — CaptureTask dedupes among live
// rows, and the intake relies on exactly that.
func TestIngestSameMailTwiceIsIdempotent(t *testing.T) {
mi, st, _ := newTestIntake(t, `[{"text":"оплатить счёт","due":""}]`)
if _, err := mi.ingest(context.Background(), ingestReq()); err != nil {
t.Fatalf("first ingest: %v", err)
}
resp, err := mi.ingest(context.Background(), ingestReq())
if err != nil {
t.Fatalf("second ingest: %v", err)
}
if resp.Created != 0 || len(resp.TaskIDs) != 1 {
t.Errorf("resp = %+v, want the existing row and Created=0", resp)
}
tasks, _ := st.ListTasks(context.Background(), "")
if len(tasks) != 1 {
t.Errorf("got %d tasks after two reads, want 1", len(tasks))
}
}
func TestIngestJunkSkipsTheModel(t *testing.T) {
mi, st, fake := newTestIntake(t, `[{"text":"купить со скидкой","due":""}]`)
req := ingestReq()
req.Junk = true
resp, err := mi.ingest(context.Background(), req)
if err != nil {
t.Fatalf("ingest: %v", err)
}
if !resp.Skipped || resp.Created != 0 {
t.Errorf("resp = %+v, want skipped", resp)
}
if fake.calls != 0 {
t.Errorf("model called %d times for junk, want 0", fake.calls)
}
if tasks, _ := st.ListTasks(context.Background(), ""); len(tasks) != 0 {
t.Errorf("junk produced %d tasks, want 0", len(tasks))
}
}
func TestIngestEmptyMessageSkipped(t *testing.T) {
mi, _, fake := newTestIntake(t, "[]")
resp, err := mi.ingest(context.Background(), ipc.IngestMailReq{Mailbox: "INBOX", UID: 1})
if err != nil || !resp.Skipped {
t.Fatalf("resp = %+v, err = %v; want skipped", resp, err)
}
if fake.calls != 0 {
t.Errorf("model called %d times for an empty message, want 0", fake.calls)
}
}
func TestIngestNoTasksWritesNothing(t *testing.T) {
mi, st, _ := newTestIntake(t, "[]")
resp, err := mi.ingest(context.Background(), ingestReq())
if err != nil {
t.Fatalf("ingest: %v", err)
}
if resp.Created != 0 || len(resp.TaskIDs) != 0 || resp.Skipped {
t.Errorf("resp = %+v, want nothing captured and not skipped", resp)
}
if tasks, _ := st.ListTasks(context.Background(), ""); len(tasks) != 0 {
t.Errorf("got %d tasks, want 0", len(tasks))
}
}
func TestIngestTruncatesEvidence(t *testing.T) {
mi, st, _ := newTestIntake(t, `[{"text":"дело","due":""}]`)
req := ingestReq()
req.Subject = strings.Repeat("щ", 400)
if _, err := mi.ingest(context.Background(), req); err != nil {
t.Fatalf("ingest: %v", err)
}
tasks, _ := st.ListTasks(context.Background(), "")
if len(tasks) != 1 {
t.Fatalf("got %d tasks, want 1", len(tasks))
}
if n := len([]rune(tasks[0].Evidence)); n > evidenceMaxChars+1 {
t.Errorf("evidence kept %d runes, want ≤ %d", n, evidenceMaxChars)
}
}
// Off unless configured: no email block ⇒ no intake, so the IPC method does not
// exist at all.
func TestNewMailIntakeOffWithoutConfig(t *testing.T) {
st := newTestStore(t)
if mi := newMailIntake(st, nil, &config.Config{}); mi != nil {
t.Error("no email block must mean no mail intake")
}
// Configured but with a non-LLM phraser: still off — there is no fallback
// extraction, by design.
if mi := newMailIntake(st, nil, &config.Config{Email: &config.EmailConfig{}}); mi != nil {
t.Error("without a llama-server phraser there is nothing to extract with")
}
}
+42
View File
@@ -170,6 +170,8 @@ func run(args []string) error {
eco *ecosystemWiring
factWorker *factEnrichmentWorker
evalWorker *memoryEvalWorker // nil ⇒ memory evaluation off (the default)
feedWkr *feedWorker // nil ⇒ no feed is read (the default)
crawlWkr *crawlWorker // nil ⇒ no page is watched (the default)
)
if !locked {
@@ -265,6 +267,8 @@ func run(args []string) error {
tl = newTickLoop(st, gatherer, dispatcher, phr, rules, tickInterval, repeatInterval, autotuneInterval, cfg.Digest, routinesFromConfig(cfg.Routines), config.MorningRoutinesFromConfig(cfg.MorningRoutines), cfg.PatternProposals)
factWorker = newFactEnrichmentWorker(st, eco, time.Duration(cfg.FactEnrichmentInterval))
evalWorker = newMemoryEvalWorker(st, phr, cfg)
feedWkr = newFeedWorker(ipc.NewStoreAPI(st), embedderOf(voiceW), cfg)
crawlWkr = newCrawlWorker(newCrawler(cfg), ipc.NewStoreAPI(st), embedderOf(voiceW), cfg)
coreAPI = &daemonAPI{
CoreAPI: ipc.NewStoreAPI(st),
@@ -320,6 +324,13 @@ func run(args []string) error {
srv.StepUp = func(ctx context.Context) error { return passkeySess.Assert(ctx, auth.Scope{}) }
// Mail ingestion (Vikunja #246): the hook stays nil unless an email block is
// configured and there is a llama-server to extract with, in which case
// ipc.MethodIngestMail reports ErrUnknownMethod.
if !locked {
wireMailIntake(srv, st, phr, cfg)
}
// WrapKeyFn — wraps the env key with a passkey credential public key and
// persists the wrapped blob. Only wired when the daemon has the key in
// memory (env key mode). Called by mavweb after passkey enrollment.
@@ -444,6 +455,8 @@ func run(args []string) error {
tl = newTickLoop(st, gatherer, dispatcher, phr, rules, tickInterval, repeatInterval, autotuneInterval, cfg.Digest, routinesFromConfig(cfg.Routines), config.MorningRoutinesFromConfig(cfg.MorningRoutines), cfg.PatternProposals)
factWorker = newFactEnrichmentWorker(st, eco, time.Duration(cfg.FactEnrichmentInterval))
evalWorker = newMemoryEvalWorker(st, phr, cfg)
feedWkr = newFeedWorker(ipc.NewStoreAPI(st), embedderOf(voiceW), cfg)
crawlWkr = newCrawlWorker(newCrawler(cfg), ipc.NewStoreAPI(st), embedderOf(voiceW), cfg)
// Swap the CoreAPI from the locked placeholder to the real store adapter.
newAPI := &daemonAPI{
@@ -457,6 +470,7 @@ func run(args []string) error {
}
srv.SetAPI(newAPI)
srv.Check = (&auth.Gate{Enrollment: auth.NewFloorEnrollment(), Session: passkeySess}).Check
wireMailIntake(srv, st, phr, cfg)
// Start voice server.
if voiceW != nil {
@@ -488,6 +502,20 @@ func run(args []string) error {
}()
}
// Start feed reading (nil unless configured).
if feedWkr != nil {
go func() {
feedWkr.run(ctx)
}()
}
// Start the watched-page crawls (nil unless configured).
if crawlWkr != nil {
go func() {
crawlWkr.run(ctx)
}()
}
dl.unlock()
log.Printf("mavend: unlocked via passkey assertion")
return nil
@@ -533,6 +561,20 @@ func run(args []string) error {
evalWorker.run(ctx)
}()
}
if feedWkr != nil {
wg.Add(1)
go func() {
defer wg.Done()
feedWkr.run(ctx)
}()
}
if crawlWkr != nil {
wg.Add(1)
go func() {
defer wg.Done()
crawlWkr.run(ctx)
}()
}
}
<-ctx.Done()
+10
View File
@@ -52,6 +52,7 @@ import (
"time"
"github.com/kami/maven/internal/audio"
"github.com/kami/maven/internal/crawl"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/memory"
@@ -82,6 +83,15 @@ type reactiveHandler struct {
replier voice.Replier
now func() time.Time
// crawler reads a web page he names out loud (queryWeb). nil ⇒ on-demand
// page reading is off, which is the default: no `crawl` block, no fetch.
crawler *crawl.Crawler
// feedsOn — whether any RSS feed is configured (config.Feeds). It changes
// only what she SAYS when asked and nothing is there: "ленты не настроены"
// instead of "ничего нового", which are different truths.
feedsOn bool
weatherProvider weather.Provider
weatherLocation string // default location for weather queries
+14 -10
View File
@@ -201,16 +201,20 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
// ----- the handler (the reactive path; closes over stt / tts / router / coreAPI / memory) -----
h := &reactiveHandler{
stt: transcriber,
tts: synthesizer,
router: rtr,
embedder: emb,
api: coreAPI,
tools: exec,
matcher: matcher,
replier: replier,
phraser: phr,
now: time.Now,
stt: transcriber,
tts: synthesizer,
router: rtr,
embedder: emb,
api: coreAPI,
tools: exec,
matcher: matcher,
replier: replier,
phraser: phr,
now: time.Now,
feedsOn: cfg.Feeds != nil,
// nil unless `crawl.on_demand` is on: reading a page he names is a
// capability, and capabilities are off unless configured.
crawler: onDemandCrawler(cfg),
weatherProvider: weatherProvider,
weatherLocation: weatherLocation,
memStore: memStore,
+332
View File
@@ -0,0 +1,332 @@
// mavmaild — the mail reader module (Vikunja #246,
// docs/plans/01-email-reader.md).
//
// Every so often it opens one IMAP mailbox read-only, fetches the messages it
// has not read yet, and hands each one to core over ipc.MethodIngestMail. Core
// runs the extraction on the resident model and writes what comes back as task
// CANDIDATES he reviews on /tasks. Nothing here writes to the store, nothing
// here can create a reminder, and nothing here speaks.
//
// Why a separate daemon rather than a loop inside mavend, when extraction has
// to happen in mavend anyway: the credential. mavpoll set the precedent with the
// zenmoney token (#125) — the module that talks to a third party holds the
// secret, reads it from a FILE so it never appears in `ps`, in
// docker-compose.yml or in shell history, and core never sees it. Core learns
// that mail exists only as message text on one IPC method; it cannot connect to
// the mailbox even if it wanted to, and a compromised core yields no mail
// password.
//
// Off unless configured: without -password-file there is nothing to run, and
// the daemon says so and exits. If core has no `email` block the very first
// ingest comes back ErrUnknownMethod and this daemon stops polling instead of
// hammering a socket that will keep refusing.
//
// Mail is personal, so the log is counts and UIDs: how many messages were
// fetched, how many were bulk, how many candidates came back. No subject, no
// sender, no body, ever — reviewing a candidate is what /tasks is for.
package main
import (
"context"
"encoding/json"
"errors"
"flag"
"fmt"
"log"
"os"
"os/signal"
"path/filepath"
"sort"
"strings"
"syscall"
"time"
"github.com/kami/maven/internal/email"
"github.com/kami/maven/internal/ipc"
)
func main() {
if err := run(os.Args[1:]); err != nil {
fmt.Fprintln(os.Stderr, "mavmaild:", err)
os.Exit(1)
}
}
func run(args []string) error {
fs := flag.NewFlagSet("mavmaild", flag.ContinueOnError)
socket := fs.String("socket", "", "core IPC socket path (required)")
server := fs.String("imap", "", "IMAP server, host or host:993 (required)")
user := fs.String("user", "", "IMAP username (required)")
passFile := fs.String("password-file", "", "file holding the IMAP password (required — never passed as a flag value)")
mailbox := fs.String("mailbox", "INBOX", "mailbox to read, read-only")
interval := fs.Duration("interval", 15*time.Minute, "how often to read the mailbox")
lookback := fs.Duration("lookback", 72*time.Hour, "how far back to search on each poll")
max := fs.Int("max", 25, "most messages to fetch in one poll")
timeout := fs.Duration("timeout", 30*time.Second, "IMAP network timeout")
statePath := fs.String("state", "", "file remembering which UIDs were read (default: none — every poll re-reads the window)")
if err := fs.Parse(args); err != nil {
return err
}
if *socket == "" {
return fmt.Errorf("-socket is required")
}
if *server == "" || *user == "" || *passFile == "" {
return fmt.Errorf("mail reading is off unless configured: set -imap, -user and -password-file")
}
// The password is read from a file, never taken as a flag value: an argv
// secret is visible in `ps` to every user on the box and lands in the compose
// file and the shell history. Read once at start — a rotated password means a
// restart, which is cheaper than re-reading his credential every quarter hour.
raw, err := os.ReadFile(*passFile)
if err != nil {
return fmt.Errorf("read password file: %w", err)
}
password := strings.TrimSpace(string(raw))
if password == "" {
return fmt.Errorf("password file %s is empty", *passFile)
}
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
defer stop()
core, err := ipc.DialWait(*socket, 60*time.Second)
if err != nil {
return err
}
defer core.Close()
r := &reader{
core: core,
addr: *server,
user: *user,
mailbox: *mailbox,
lookback: *lookback,
max: *max,
timeout: *timeout,
state: newSeenState(*statePath),
}
if err := r.state.load(); err != nil {
// A missing or corrupt state file must not stop mail from being read: the
// worst case is re-reading the window, and capture dedupes on text.
log.Printf("mavmaild: state: %v (starting from an empty seen-set)", err)
}
// The password is never logged, not even its length.
log.Printf("mavmaild: reading %s on %s every %s (lookback %s, max %d/poll)",
*mailbox, *server, *interval, *lookback, *max)
r.pollOnce(ctx, password) // don't idle a full interval on start
t := time.NewTicker(*interval)
defer t.Stop()
for {
select {
case <-ctx.Done():
log.Printf("mavmaild: bye")
return nil
case <-t.C:
if r.disabled {
// Core told us mail ingestion is not configured. Nothing will change
// without a core restart, and a restart restarts us too.
log.Printf("mavmaild: core does not accept mail — idling")
return nil
}
r.pollOnce(ctx, password)
}
}
}
// mailIngester — the slice of core this daemon uses. One method: hand over a
// message. It cannot write a fact, create a reminder or read the store, and the
// interface says so.
type mailIngester interface {
IngestMail(ctx context.Context, req ipc.IngestMailReq) (ipc.IngestMailResp, error)
}
type reader struct {
core mailIngester
addr string
user string
mailbox string
lookback time.Duration
max int
timeout time.Duration
state *seenState
// dial — connection seam for the tests; nil ⇒ implicit TLS.
dial func(addr string, timeout time.Duration) (*email.Conn, error)
// disabled — core answered ErrUnknownMethod, i.e. it has no email block.
disabled bool
}
// pollOnce — one read of the mailbox, then one ingest per message.
//
// A fetch error aborts this poll and nothing else; the next tick tries again.
// An ingest error for one message does not skip the rest — one mail the model
// choked on should not hide the four behind it.
func (r *reader) pollOnce(ctx context.Context, password string) {
msgs, err := r.fetch(password)
if err != nil {
// The error may name a UID; it never names a subject or a sender.
log.Printf("mavmaild: fetch: %v", err)
if len(msgs) == 0 {
return
}
}
var junk, candidates, created int
for _, m := range msgs {
if ctx.Err() != nil {
return
}
if m.Junk {
junk++
// Marked seen without a model call: the header filter already decided,
// and re-classifying it every quarter hour would be pure waste.
r.state.mark(m.UID)
continue
}
resp, err := r.core.IngestMail(ctx, ipc.IngestMailReq{
Mailbox: r.mailbox,
UID: m.UID,
From: m.From,
Subject: m.Subject,
Date: m.Date,
Body: m.Body,
})
if errors.Is(err, ipc.ErrUnknownMethod) {
log.Printf("mavmaild: core has no email block configured — mail ingestion is off; stopping")
r.disabled = true
return
}
if err != nil {
// Not marked seen: an ingest that failed should be retried next poll.
log.Printf("mavmaild: ingest uid %d: %v", m.UID, err)
continue
}
r.state.mark(m.UID)
candidates += len(resp.TaskIDs)
created += resp.Created
}
if err := r.state.save(); err != nil {
log.Printf("mavmaild: state: %v", err)
}
log.Printf("mavmaild: %s: %d read, %d bulk, %d candidate(s), %d new", r.mailbox, len(msgs), junk, candidates, created)
}
// fetch reads the mailbox. Messages already in the seen-set are not fetched at
// all, so a steady mailbox costs one SEARCH per poll and nothing else.
func (r *reader) fetch(password string) ([]email.Message, error) {
f := email.FetchSince{
Addr: r.addr,
User: r.user,
Mailbox: r.mailbox,
Timeout: r.timeout,
Since: time.Now().Add(-r.lookback),
Max: r.max,
Skip: r.state.seen,
}
return f.RunWith(password, r.dial)
}
// ---- seen state ------------------------------------------------------------
// seenState — the UIDs already handed to core, persisted so a restart does not
// re-read (and re-extract, at multi-second LLM cost) the whole lookback window.
//
// Correctness does not depend on it: ipc.CaptureTask dedupes on normalised text
// among live tasks, so a re-read produces no duplicate rows. This exists to save
// the model's time, which is why a broken state file is a log line rather than a
// failure.
//
// UIDs are per-mailbox and monotonic, so the set is kept as a high-water mark
// plus the stragglers above it. If the server ever changes UIDVALIDITY, UIDs
// reset and the window is simply re-read once — dedupe absorbs it.
type seenState struct {
path string
high uint32
set map[uint32]bool
dirty bool
}
func newSeenState(path string) *seenState {
return &seenState{path: path, set: map[uint32]bool{}}
}
type seenFile struct {
High uint32 `json:"high"`
UIDs []uint32 `json:"uids,omitempty"`
}
func (s *seenState) seen(uid uint32) bool {
return uid <= s.high || s.set[uid]
}
func (s *seenState) mark(uid uint32) {
if s.seen(uid) {
return
}
s.set[uid] = true
s.dirty = true
// Advance the high-water mark through any contiguous run, so the explicit set
// stays small on a mailbox read in order.
for {
next := s.high + 1
if !s.set[next] {
break
}
delete(s.set, next)
s.high = next
}
}
func (s *seenState) load() error {
if s.path == "" {
return nil
}
b, err := os.ReadFile(s.path)
if errors.Is(err, os.ErrNotExist) {
return nil // first run
}
if err != nil {
return err
}
var f seenFile
if err := json.Unmarshal(b, &f); err != nil {
return fmt.Errorf("parse %s: %w", s.path, err)
}
s.high = f.High
for _, u := range f.UIDs {
s.set[u] = true
}
return nil
}
// save writes the state atomically (temp file + rename), 0600: it is a list of
// message ids from his mailbox, which is metadata about his mail.
func (s *seenState) save() error {
if s.path == "" || !s.dirty {
return nil
}
uids := make([]uint32, 0, len(s.set))
for u := range s.set {
uids = append(uids, u)
}
sort.Slice(uids, func(i, j int) bool { return uids[i] < uids[j] })
b, err := json.Marshal(seenFile{High: s.high, UIDs: uids})
if err != nil {
return err
}
tmp := s.path + ".tmp"
if err := os.MkdirAll(filepath.Dir(s.path), 0o700); err != nil {
return err
}
if err := os.WriteFile(tmp, b, 0o600); err != nil {
return err
}
if err := os.Rename(tmp, s.path); err != nil {
return err
}
s.dirty = false
return nil
}
+254
View File
@@ -0,0 +1,254 @@
package main
import (
"bufio"
"context"
"fmt"
"net"
"os"
"path/filepath"
"strconv"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/email"
"github.com/kami/maven/internal/ipc"
)
// ---- a scripted IMAP server, same shape internal/email's tests use ---------
type fakeIMAP struct {
msgs map[uint32]string
uids []uint32
cmds []string
}
func (f *fakeIMAP) serve(c net.Conn) {
defer c.Close()
fmt.Fprint(c, "* OK fake ready\r\n")
r := bufio.NewReader(c)
for {
line, err := r.ReadString('\n')
if err != nil {
return
}
parts := strings.SplitN(strings.TrimRight(line, "\r\n"), " ", 2)
if len(parts) != 2 {
return
}
tag, cmd := parts[0], parts[1]
f.cmds = append(f.cmds, cmd)
upper := strings.ToUpper(cmd)
switch {
case strings.HasPrefix(upper, "LOGIN"), strings.HasPrefix(upper, "EXAMINE"):
fmt.Fprintf(c, "%s OK\r\n", tag)
case strings.HasPrefix(upper, "UID SEARCH"):
var ids []string
for _, u := range f.uids {
ids = append(ids, strconv.FormatUint(uint64(u), 10))
}
fmt.Fprintf(c, "* SEARCH %s\r\n%s OK\r\n", strings.Join(ids, " "), tag)
case strings.HasPrefix(upper, "UID FETCH"):
uid, _ := strconv.ParseUint(strings.Fields(cmd)[2], 10, 32)
if raw, ok := f.msgs[uint32(uid)]; ok {
fmt.Fprintf(c, "* 1 FETCH (UID %d BODY[] {%d}\r\n%s)\r\n", uid, len(raw), raw)
}
fmt.Fprintf(c, "%s OK\r\n", tag)
case strings.HasPrefix(upper, "LOGOUT"):
fmt.Fprintf(c, "* BYE\r\n%s OK\r\n", tag)
return
default:
fmt.Fprintf(c, "%s BAD\r\n", tag)
}
}
}
func (f *fakeIMAP) dial(_ string, timeout time.Duration) (*email.Conn, error) {
cli, srv := net.Pipe()
go f.serve(srv)
return email.NewConn(cli, timeout)
}
// ---- a fake core -----------------------------------------------------------
type fakeCore struct {
got []ipc.IngestMailReq
resp ipc.IngestMailResp
err error
}
func (c *fakeCore) IngestMail(_ context.Context, req ipc.IngestMailReq) (ipc.IngestMailResp, error) {
c.got = append(c.got, req)
if c.err != nil {
return ipc.IngestMailResp{}, c.err
}
return c.resp, nil
}
func mail(subject, body string, extraHeaders ...string) string {
h := "Subject: " + subject + "\r\nFrom: a@b.c\r\nContent-Type: text/plain; charset=utf-8\r\n"
for _, e := range extraHeaders {
h += e + "\r\n"
}
return h + "\r\n" + body + "\r\n"
}
func newTestReader(t *testing.T, f *fakeIMAP, core *fakeCore, statePath string) *reader {
t.Helper()
return &reader{
core: core, addr: "mail.example:993", user: "kami", mailbox: "INBOX",
lookback: 72 * time.Hour, max: 25, timeout: 5 * time.Second,
state: newSeenState(statePath),
dial: f.dial,
}
}
func TestPollHandsMessagesToCore(t *testing.T) {
f := &fakeIMAP{
uids: []uint32{1, 2},
msgs: map[uint32]string{
1: mail("Счёт", "Оплатить до 5 августа."),
2: mail("Скидки", "Sale!", "List-Unsubscribe: <mailto:u@x>"),
},
}
core := &fakeCore{resp: ipc.IngestMailResp{TaskIDs: []int64{1}, Created: 1}}
r := newTestReader(t, f, core, "")
r.pollOnce(context.Background(), "secret")
// The newsletter is filtered before core is asked: only the real mail crosses.
if len(core.got) != 1 {
t.Fatalf("core saw %d messages, want 1 (the bulk one must not cross): %+v", len(core.got), core.got)
}
got := core.got[0]
if got.UID != 1 || got.Mailbox != "INBOX" || got.Subject != "Счёт" {
t.Errorf("ingest req = %+v", got)
}
if !strings.Contains(got.Body, "Оплатить") {
t.Errorf("body = %q", got.Body)
}
}
// A second poll must not re-send what core already saw — extraction is a
// multi-second LLM call per message.
func TestPollSkipsSeenUIDs(t *testing.T) {
f := &fakeIMAP{uids: []uint32{5}, msgs: map[uint32]string{5: mail("Счёт", "текст")}}
core := &fakeCore{}
r := newTestReader(t, f, core, "")
r.pollOnce(context.Background(), "secret")
r.pollOnce(context.Background(), "secret")
if len(core.got) != 1 {
t.Errorf("core saw %d messages over two polls, want 1", len(core.got))
}
}
// An ingest that failed is NOT marked seen: the next poll retries it.
func TestPollRetriesFailedIngest(t *testing.T) {
f := &fakeIMAP{uids: []uint32{5}, msgs: map[uint32]string{5: mail("Счёт", "текст")}}
core := &fakeCore{err: fmt.Errorf("llama-server is warming up")}
r := newTestReader(t, f, core, "")
r.pollOnce(context.Background(), "secret")
core.err = nil
r.pollOnce(context.Background(), "secret")
if len(core.got) != 2 {
t.Errorf("core saw %d attempts, want 2 (a failed ingest is retried)", len(core.got))
}
}
// Core without an email block ⇒ stop, don't hammer the socket.
func TestPollStopsWhenCoreRefusesMail(t *testing.T) {
f := &fakeIMAP{uids: []uint32{1, 2}, msgs: map[uint32]string{1: mail("a", "b"), 2: mail("c", "d")}}
core := &fakeCore{err: fmt.Errorf("call: %w", ipc.ErrUnknownMethod)}
r := newTestReader(t, f, core, "")
r.pollOnce(context.Background(), "secret")
if !r.disabled {
t.Error("ErrUnknownMethod must disable the reader")
}
if len(core.got) != 1 {
t.Errorf("core saw %d messages, want 1 — stop at the first refusal", len(core.got))
}
}
func TestSeenStatePersists(t *testing.T) {
path := filepath.Join(t.TempDir(), "state", "seen.json")
f := &fakeIMAP{uids: []uint32{9}, msgs: map[uint32]string{9: mail("Счёт", "текст")}}
core := &fakeCore{}
r := newTestReader(t, f, core, path)
r.pollOnce(context.Background(), "secret")
fi, err := os.Stat(path)
if err != nil {
t.Fatalf("state file: %v", err)
}
// A list of message ids from his mailbox is metadata about his mail.
if perm := fi.Mode().Perm(); perm != 0o600 {
t.Errorf("state file mode = %v, want 0600", perm)
}
// A fresh reader with the same state file must not re-read the message.
core2 := &fakeCore{}
r2 := newTestReader(t, f, core2, path)
if err := r2.state.load(); err != nil {
t.Fatalf("load: %v", err)
}
r2.pollOnce(context.Background(), "secret")
if len(core2.got) != 0 {
t.Errorf("after a restart core saw %d messages, want 0", len(core2.got))
}
}
func TestSeenStateHighWaterMark(t *testing.T) {
s := newSeenState("")
s.mark(1)
s.mark(3)
s.mark(2)
if s.high != 3 {
t.Errorf("high = %d, want 3 (contiguous run collapses)", s.high)
}
if len(s.set) != 0 {
t.Errorf("explicit set = %v, want empty", s.set)
}
if !s.seen(2) || s.seen(4) {
t.Errorf("seen(2)=%v seen(4)=%v", s.seen(2), s.seen(4))
}
}
func TestSeenStateCorruptFileIsNotFatal(t *testing.T) {
path := filepath.Join(t.TempDir(), "seen.json")
if err := os.WriteFile(path, []byte("{not json"), 0o600); err != nil {
t.Fatal(err)
}
s := newSeenState(path)
if err := s.load(); err == nil {
t.Error("a corrupt state file should report an error the caller logs")
}
if s.seen(1) {
t.Error("a corrupt state file must leave an empty seen-set, not a poisoned one")
}
}
// Off unless configured, and the credential is never a flag value.
func TestRunRequiresConfig(t *testing.T) {
if err := run([]string{}); err == nil {
t.Error("no -socket must be an error")
}
if err := run([]string{"-socket", "/tmp/nope.sock"}); err == nil {
t.Error("no mailbox configuration must be an error, not a default mailbox")
}
// There is no -password flag at all: only -password-file.
if err := run([]string{"-socket", "/x", "-imap", "h", "-user", "u", "-password", "p"}); err == nil ||
!strings.Contains(err.Error(), "flag provided but not defined") {
t.Errorf("a -password flag must not exist; err = %v", err)
}
}
func TestRunRejectsEmptyPasswordFile(t *testing.T) {
path := filepath.Join(t.TempDir(), "pass")
if err := os.WriteFile(path, []byte(" \n"), 0o600); err != nil {
t.Fatal(err)
}
err := run([]string{"-socket", "/x/y.sock", "-imap", "h", "-user", "u", "-password-file", path})
if err == nil || !strings.Contains(err.Error(), "empty") {
t.Errorf("an empty password file must be refused before dialling; err = %v", err)
}
}
+124 -3
View File
@@ -9,6 +9,12 @@
// Two sources, each its own provenance (the loop's rules trust source):
// - netdata → poll:netdata resource alarms (disk/mem/cert/temp)
// - kuma → poll:uptimekuma service up/down (the source of truth for it)
// - zenmoney → poll:zenmoney spending/income totals (Vikunja #125)
//
// The zenmoney source is why the token lives HERE and not in core: the poller
// already owns every other third-party credential, it holds no store key, and
// core never needs to know an account exists to answer a question about a fact
// the poller wrote. It is off unless -zenmoney-token-file is given.
//
// Netdata needs no auth over the wg-fronted net. Kuma's /metrics needs an API
// key (basic-auth); without -kuma the whole kuma path is skipped (netdata-only
@@ -37,6 +43,7 @@ import (
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/zenmoney"
)
func main() {
@@ -52,6 +59,9 @@ func run(args []string) error {
netdataURL := fs.String("netdata", "http://127.0.0.1:19999", "netdata base URL ('' to disable)")
kumaURL := fs.String("kuma", "", "uptime-kuma metrics URL, e.g. http://127.0.0.1:3001/metrics ('' to disable)")
kumaKey := fs.String("kuma-key", "", "uptime-kuma API key (basic-auth username)")
zenTokenFile := fs.String("zenmoney-token-file", "", "file holding the zenmoney API token ('' disables money tracking)")
zenURL := fs.String("zenmoney-url", zenmoney.DefaultBaseURL, "zenmoney API base URL (tests/self-hosted proxies)")
zenInterval := fs.Duration("zenmoney-interval", time.Hour, "how often to read zenmoney (money does not move every minute)")
wgIface := fs.String("wg", "", "wireguard interface for the presence signal, e.g. wg0 or 'all' ('' to disable)")
wgCmd := fs.String("wg-cmd", "wg", "wg binary (use e.g. 'sudo wg' if the poller lacks CAP_NET_ADMIN)")
interval := fs.Duration("interval", 60*time.Second, "poll cadence")
@@ -62,8 +72,24 @@ func run(args []string) error {
if *socket == "" {
return fmt.Errorf("-socket is required")
}
if *netdataURL == "" && *kumaURL == "" && *wgIface == "" {
return fmt.Errorf("nothing to poll: set -netdata, -kuma and/or -wg")
if *netdataURL == "" && *kumaURL == "" && *wgIface == "" && *zenTokenFile == "" {
return fmt.Errorf("nothing to poll: set -netdata, -kuma, -wg and/or -zenmoney-token-file")
}
// The token is read from a file, never taken as a flag value: an argv token
// is visible in `ps` to every user on the box and lands in the compose file
// and the shell history. Read once at start — a rotated token means a
// restart, which is cheaper than re-reading his credential every hour.
var zen *zenmoney.Client
if *zenTokenFile != "" {
raw, err := os.ReadFile(*zenTokenFile)
if err != nil {
return fmt.Errorf("read zenmoney token: %w", err)
}
zen, err = zenmoney.New(strings.TrimSpace(string(raw)), *zenURL, *timeout*3)
if err != nil {
return err
}
}
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
@@ -83,9 +109,13 @@ func run(args []string) error {
kumaKey: *kumaKey,
wgIface: *wgIface,
wgCmd: *wgCmd,
zen: zen,
zenEvery: *zenInterval,
}
log.Printf("mavpoll: polling every %s (netdata=%q kuma=%q wg=%q)", *interval, *netdataURL, *kumaURL, *wgIface)
// The token is never logged, not even its length.
log.Printf("mavpoll: polling every %s (netdata=%q kuma=%q wg=%q zenmoney=%v every %s)",
*interval, *netdataURL, *kumaURL, *wgIface, zen != nil, *zenInterval)
p.pollOnce(ctx) // fire immediately; don't idle a full interval on start
t := time.NewTicker(*interval)
defer t.Stop()
@@ -108,6 +138,12 @@ type poller struct {
kumaKey string
wgIface string
wgCmd string
// zen is nil unless a token file was configured — money tracking is a
// capability, off by default like weather and telegram.
zen *zenmoney.Client
zenEvery time.Duration
zenLast time.Time
}
// pollOnce — one sweep of both sources. A failure in one source logs and does
@@ -129,6 +165,66 @@ func (p *poller) pollOnce(ctx context.Context) {
log.Printf("mavpoll: wg: %v", err)
}
}
// Money on its own, much slower cadence: a bank feed that updates hourly
// polled every minute is 60 pointless reads of his financial history.
if p.zen != nil && now.Sub(p.zenLast) >= p.zenEvery {
p.zenLast = now
if err := p.pollZenmoney(ctx, now); err != nil {
log.Printf("mavpoll: zenmoney: %v", err)
}
}
}
// ---- zenmoney: spending/income totals → money facts ------------------------
// pollZenmoney reads today's and this month's totals and writes them as
// facts(kind=env, source=poll:zenmoney) (Vikunja #125).
//
// Two properties this function exists to hold:
//
// - An empty or failed read writes NOTHING. zenmoney.Summary.Value() refuses
// to encode a summary built from zero transactions, so a poller that cannot
// reach the API leaves the last good fact in place rather than overwriting
// it with a zero Maven would then recite as fact.
// - Nothing about the money leaves the box except the diff request itself, to
// the service that already holds his bank sessions. The totals are written
// to the store and read back only when he asks; they are never search input
// and no tick rule fires on them.
//
// Both windows are read from one diff call each. Two calls an hour against an
// API whose whole job is this is not worth caching.
// moneyWindow — one fact key and the period it covers.
type moneyWindow struct {
key string
from, to time.Time
}
func (p *poller) pollZenmoney(ctx context.Context, now time.Time) error {
dFrom, dTo := zenmoney.DayWindow(now)
mFrom, mTo := zenmoney.MonthWindow(now)
windows := []moneyWindow{
{zenmoney.KeySpentToday, dFrom, dTo},
{zenmoney.KeySpentMonth, mFrom, mTo},
}
var firstErr error
for _, w := range windows {
sum, err := p.zen.Since(ctx, w.from, w.to)
if err != nil {
if firstErr == nil {
firstErr = err
}
continue
}
val, ok := sum.Value()
if !ok {
// Nothing read. Silence, not a zero.
continue
}
if err := p.writeIfChangedRaw(ctx, w.key, zenmoney.Source, val, now); err != nil && firstErr == nil {
firstErr = err
}
}
return firstErr
}
// ---- wireguard: latest handshake → presence signal -------------------------
@@ -301,6 +397,31 @@ func (p *poller) writeIfChanged(ctx context.Context, key, source, val string, no
return nil
}
// writeIfChangedRaw is writeIfChanged for values that are already JSON (the
// money facts store an object, not a string). Kept separate rather than
// generalising writeIfChanged, because the string-valued env facts encoding
// their own value is the convention the rules rely on.
//
// The log line names the key and the source, never the figures: mavpoll's log
// is not the place his spending ends up.
func (p *poller) writeIfChangedRaw(ctx context.Context, key, source, jsonVal string, now time.Time) error {
prev, err := p.core.LatestFactBySource(ctx, key, source)
switch {
case err == nil && prev.Value == jsonVal:
return nil
case err != nil && err != ipc.ErrNoFact && !isNoFact(err):
return fmt.Errorf("read %s: %w", key, err)
}
if _, err := p.core.WriteFact(ctx, ipc.WriteFactReq{
Ts: now, Kind: "env", Key: key, Value: jsonVal,
Source: source, Confidence: 1.0,
}); err != nil {
return fmt.Errorf("write %s: %w", key, err)
}
log.Printf("mavpoll: %s updated (%s)", key, source)
return nil
}
// isNoFact — ErrNoFact rehydrated over the wire is wrapped (fmt.Errorf %w), so
// errors.Is is the right check; keep a helper so the switch above reads clean.
func isNoFact(err error) bool {
+126
View File
@@ -1,8 +1,17 @@
package main
import (
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"os"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/zenmoney"
)
func TestMaxSeverity(t *testing.T) {
@@ -62,3 +71,120 @@ func TestParseMaxHandshake(t *testing.T) {
}
}
}
// ---- zenmoney (Vikunja #125) ----------------------------------------------
// factCore records the facts the poller wrote and answers "no fact yet".
type factCore struct {
ipc.UnimplementedCoreAPI
written []ipc.WriteFactReq
prev map[string]string
}
func (c *factCore) LatestFactBySource(_ context.Context, key, source string) (ipc.Fact, error) {
if v, ok := c.prev[key+"|"+source]; ok {
return ipc.Fact{Key: key, Source: source, Value: v}, nil
}
return ipc.Fact{}, ipc.ErrNoFact
}
func (c *factCore) WriteFact(_ context.Context, req ipc.WriteFactReq) (int64, error) {
c.written = append(c.written, req)
return int64(len(c.written)), nil
}
func zenFixtureServer(t *testing.T, body []byte, status int) *httptest.Server {
t.Helper()
return httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if status != http.StatusOK {
w.WriteHeader(status)
return
}
w.Write(body)
}))
}
func TestPollZenmoneyWritesMoneyFacts(t *testing.T) {
body, err := os.ReadFile("../../internal/zenmoney/testdata/diff.json")
if err != nil {
t.Fatal(err)
}
srv := zenFixtureServer(t, body, http.StatusOK)
defer srv.Close()
zen, err := zenmoney.New("tok", srv.URL, time.Second)
if err != nil {
t.Fatal(err)
}
core := &factCore{}
p := &poller{core: core, zen: zen}
now := time.Date(2026, 8, 1, 21, 0, 0, 0, time.UTC)
if err := p.pollZenmoney(context.Background(), now); err != nil {
t.Fatal(err)
}
if len(core.written) != 2 {
t.Fatalf("wrote %d facts, want today + month", len(core.written))
}
for _, f := range core.written {
if f.Kind != "env" || f.Source != zenmoney.Source {
t.Errorf("fact = %+v, want kind=env source=%s", f, zenmoney.Source)
}
if _, err := zenmoney.ParseFactValue(f.Value); err != nil {
t.Errorf("fact value %q does not decode: %v", f.Value, err)
}
}
}
// A read that returns nothing for the window writes NOTHING. Silence, not a
// zero: an invented 0 would be recited back to him as fact.
func TestPollZenmoneyWritesNothingWhenEmpty(t *testing.T) {
srv := zenFixtureServer(t, []byte(`{"serverTimestamp":1,"instrument":[],"transaction":[]}`), http.StatusOK)
defer srv.Close()
zen, _ := zenmoney.New("tok", srv.URL, time.Second)
core := &factCore{}
p := &poller{core: core, zen: zen}
if err := p.pollZenmoney(context.Background(), time.Now()); err != nil {
t.Fatal(err)
}
if len(core.written) != 0 {
t.Errorf("wrote %+v, want no fact at all", core.written)
}
}
// An API failure must not overwrite the last good total either.
func TestPollZenmoneyFailureWritesNothing(t *testing.T) {
srv := zenFixtureServer(t, nil, http.StatusUnauthorized)
defer srv.Close()
zen, _ := zenmoney.New("bad", srv.URL, time.Second)
core := &factCore{}
p := &poller{core: core, zen: zen}
if err := p.pollZenmoney(context.Background(), time.Now()); err == nil {
t.Error("want the 401 reported")
}
if len(core.written) != 0 {
t.Errorf("wrote %+v on a failed read", core.written)
}
}
// Unchanged totals do not churn the facts table.
func TestWriteIfChangedRawSkipsUnchanged(t *testing.T) {
core := &factCore{prev: map[string]string{
zenmoney.KeySpentToday + "|" + zenmoney.Source: `{"count":1}`,
}}
p := &poller{core: core}
if err := p.writeIfChangedRaw(context.Background(), zenmoney.KeySpentToday, zenmoney.Source, `{"count":1}`, time.Now()); err != nil {
t.Fatal(err)
}
if len(core.written) != 0 {
t.Errorf("wrote %+v for an unchanged value", core.written)
}
}
// Money tracking is off unless configured: no token file, no zenmoney client,
// and the poller still refuses to start with nothing at all to poll.
func TestRunRequiresSomethingToPoll(t *testing.T) {
err := run([]string{"-socket", "/tmp/nope.sock", "-netdata", "", "-kuma", "", "-wg", ""})
if err == nil || !strings.Contains(err.Error(), "nothing to poll") {
t.Errorf("err = %v, want a 'nothing to poll' refusal", err)
}
}
+49 -6
View File
@@ -25,6 +25,7 @@ import (
"github.com/kami/maven/internal/audio"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/pattern"
"github.com/kami/maven/internal/tasks"
"github.com/kami/maven/internal/voice"
"github.com/kami/maven/internal/webauthn"
)
@@ -840,6 +841,10 @@ type taskRow struct {
Due string
Created string
Resolved string
// Why — the ranker's reason for this row's position (Vikunja #129), in
// Russian, empty when nothing distinguished the task. Blank is the honest
// rendering: he never said this one mattered more.
Why string
}
// handleTasks serves the task review surface (GET) and the four writes it
@@ -878,20 +883,46 @@ func handleTasks(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
http.Error(w, "tasks error: "+err.Error(), http.StatusBadGateway)
return
}
var cands, open, resolved []taskRow
// Live rows are ordered by the same ranker the spoken list uses, so the page
// and the voice reply can never disagree about what comes first. Resolved
// rows keep store order (newest first) — ranking finished work is pointless.
var live []tasks.Item
var resolved []taskRow
for _, t := range all {
switch t.Status {
case "candidate", "open":
live = append(live, tasks.Item{
ID: t.ID, Text: t.Text, Status: t.Status,
Created: t.CreatedTs, Due: t.Due, Weight: t.Weight,
})
default:
resolved = append(resolved, taskRow{
ID: t.ID, Text: t.Text, Source: t.Source, Evidence: t.Evidence,
Status: t.Status, Created: fmtTaskTime(&t.CreatedTs),
Due: fmtTaskDate(t.Due), Resolved: fmtTaskTime(t.Resolved),
})
}
}
byID := make(map[int64]ipc.Task, len(all))
for _, t := range all {
byID[t.ID] = t
}
var cands, open []taskRow
for _, r := range tasks.Rank(live, time.Now()) {
t := byID[r.ID]
row := taskRow{
ID: t.ID, Text: t.Text, Source: t.Source, Evidence: t.Evidence,
Status: t.Status, Created: fmtTaskTime(&t.CreatedTs),
Due: fmtTaskDate(t.Due), Resolved: fmtTaskTime(t.Resolved),
Why: r.Reason,
}
switch t.Status {
case "candidate":
if t.Status == "candidate" {
// A candidate's due date is Maven's reading of a mail, so its
// ranking reason is not shown as if he had set a priority.
row.Why = ""
cands = append(cands, row)
case "open":
} else {
open = append(open, row)
default:
resolved = append(resolved, row)
}
}
w.Header().Set("Content-Type", "text/html; charset=utf-8")
@@ -916,6 +947,18 @@ func applyTaskPost(ctx context.Context, core ipc.CoreAPI, r *http.Request) (stri
return "", errors.New("empty task text")
}
req := ipc.CaptureTaskReq{Text: text, Source: "tap:web", Status: "open", Ts: time.Now()}
// Importance is his, stated on the form. Out-of-range values are
// clamped rather than rejected — a bad select is not worth a 400.
if v := r.FormValue("weight"); v != "" {
var wgt int
if n, _ := fmt.Sscanf(v, "%d", &wgt); n != 1 || wgt < 0 {
return "", fmt.Errorf("bad weight %q", v)
}
if wgt > tasks.MaxWeight {
wgt = tasks.MaxWeight
}
req.Weight = wgt
}
if d := r.FormValue("due"); d != "" {
due, err := time.ParseInLocation("2006-01-02", d, time.Local)
if err != nil {
+8 -1
View File
@@ -9,6 +9,11 @@
<input type=hidden name=action value=add>
<input type=text name=text placeholder="что нужно сделать" size=44 required>
<input type=date name=due title="due date (optional)">
<select name=weight title="importance (optional)">
<option value=0>normal</option>
<option value=2>важно</option>
<option value=3>срочно</option>
</select>
<button class=btn>add</button>
</form>
</section>
@@ -38,10 +43,12 @@
<section class=card>
<h2 class=card-title>open <span class=badge>{{len .Open}}</span></h2>
<div class=hint>most pressing first — by the deadlines and the urgency you gave, nothing guessed.</div>
{{if .Open}}<div class=scroll><table>
<tr><th>task</th><th>from</th><th>due</th><th>captured</th><th></th><th></th></tr>
<tr><th>task</th><th>why</th><th>from</th><th>due</th><th>captured</th><th></th><th></th></tr>
{{range .Open}}<tr>
<td class=text-max>{{.Text}}</td>
<td class=hint>{{.Why}}</td>
<td class=hint>{{.Source}}</td>
<td>{{.Due}}</td>
<td class=muted>{{.Created}}</td>
+59
View File
@@ -10,6 +10,7 @@ import (
"time"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/tasks"
)
// fakeTaskCore serves the /tasks handler: a canned list plus a log of the
@@ -164,3 +165,61 @@ func TestHandleTasksNoCore(t *testing.T) {
t.Errorf("status = %d, want 503", rec.Code)
}
}
// The open list is ordered by the ranker, and the reason is shown so the page
// says why a task is first instead of asking him to trust the order.
func TestHandleTasksOrdersOpenByRank(t *testing.T) {
now := time.Now()
due := now
core := &fakeTaskCore{tasks: []ipc.Task{
{ID: 1, Text: "купить молоко", Status: "open", CreatedTs: now},
{ID: 2, Text: "оплатить интернет", Status: "open", CreatedTs: now, Due: &due},
}}
rec := httptest.NewRecorder()
handleTasks(rec, httptest.NewRequest(http.MethodGet, "/tasks", nil), core)
body := rec.Body.String()
if strings.Index(body, "оплатить интернет") > strings.Index(body, "купить молоко") {
t.Error("want the dated task rendered first")
}
if !strings.Contains(body, "сегодня") {
t.Error("want the ranker's reason shown in the why column")
}
}
// A candidate is ranked into place but never carries a priority reason: its due
// date is Maven's reading of a mail, not something he stated.
func TestHandleTasksHidesCandidateReason(t *testing.T) {
now := time.Now()
due := now
core := &fakeTaskCore{tasks: []ipc.Task{
{ID: 1, Text: "продлить страховку", Status: "candidate", CreatedTs: now, Due: &due},
}}
rec := httptest.NewRecorder()
handleTasks(rec, httptest.NewRequest(http.MethodGet, "/tasks", nil), core)
if strings.Contains(rec.Body.String(), "сегодня") {
t.Error("a candidate must not be shown with a priority reason")
}
}
func TestApplyTaskPostCarriesWeight(t *testing.T) {
core := &fakeTaskCore{created: true}
form := url.Values{"action": {"add"}, "text": {"оплатить интернет"}, "weight": {"3"}}
req := httptest.NewRequest(http.MethodPost, "/tasks", strings.NewReader(form.Encode()))
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
handleTasks(httptest.NewRecorder(), req, core)
if len(core.captured) != 1 || core.captured[0].Weight != 3 {
t.Fatalf("captured = %+v, want weight 3", core.captured)
}
}
// Out of range clamps rather than 400s; a non-number is a real client error.
func TestApplyTaskPostClampsWeight(t *testing.T) {
core := &fakeTaskCore{created: true}
form := url.Values{"action": {"add"}, "text": {"что-то"}, "weight": {"99"}}
req := httptest.NewRequest(http.MethodPost, "/tasks", strings.NewReader(form.Encode()))
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
handleTasks(httptest.NewRecorder(), req, core)
if core.captured[0].Weight != tasks.MaxWeight {
t.Errorf("weight = %d, want the cap", core.captured[0].Weight)
}
}
+61
View File
@@ -36,6 +36,67 @@ present (see `.dockerignore`).
| `/var/lib/maven` (volume) | encrypted db at rest |
| `/dev/shm` (tmpfs) | decrypted db working copy (RAM only) |
## Reading the outside world (off by default)
`mavend.json` ships without a `feeds` block, which means no RSS/Atom feed is
fetched and no outbound request is made. Switching it on is adding the block:
```json
"feeds": {
"poll_interval": "30m",
"max_items": 5,
"max_age": "24h",
"sources": [
{ "name": "habr", "url": "https://habr.com/ru/rss/best/daily/",
"category": "технологии", "exclude": ["реклама"] }
]
}
```
What it does and does not do:
- items are written as notes with source `rss:<name>`, visible on `/dash`;
- **nothing is announced.** She reads them back when asked — "что нового в
лентах?", "что нового по технологиям?" — and never on arrival. There is no
severity or channel knob here on purpose;
- the fetcher is allowlisted to the hosts of the configured feeds, plus any
`allow_hosts`. It refuses non-http(s) schemes and every private address
(loopback, the LAN, the `10.42.0.0/24` wg range, cloud metadata). It caps the
response at 2 MiB and redirects at 3, and makes at most one request per host
per second. See `internal/webfetch`;
- how far each feed was read is stored as a config fact `rss:latest:<name>`, so
a restart does not re-note yesterday's headlines.
### Reading a page (`crawl`, also off by default)
There is no `crawl` block either, so no page is fetched. Two halves, separately
switched:
```json
"crawl": {
"on_demand": true,
"interval": "6h",
"max_runes": 4000,
"watches": [
{ "name": "changelog", "url": "https://example.org/changelog", "interval": "12h" }
]
}
```
- `on_demand` lets her read a page he names in the utterance: "посмотри
https://example.org/x — что там?". The page becomes context for his question,
and only the URL leaves the box. Without a URL nothing is fetched, so this is
a fallback and not a habit;
- `watches` re-reads a fixed list on its interval and writes a note when the
text changed. Like the feeds, it announces nothing;
- the answer path sits **last** in the query chain, behind his memory, his notes
and (once wired) the local Kiwix ZIMs. A local read costs nothing;
- `robots.txt` is fetched first and obeyed with no override; a `Disallow` is a
refusal she says out loud. `Crawl-delay` is honoured;
- same guarded fetcher as the feeds: allowlist/denylist, no private addresses,
size cap, redirect cap, timeout, one request per host per second;
- dedup state is the config fact `crawl:hash:<name>`.
## Not yet verified / host-dependent
This stack is correct-by-construction but has **not been build-tested here**
+33
View File
@@ -101,9 +101,42 @@ services:
"-netdata", "http://127.0.0.1:19999",
"-kuma", "http://127.0.0.1:3001/metrics",
"-kuma-key", "uk5_mavpoll-key"]
# Money tracking (Vikunja #125) is OFF: it needs a zenmoney token,
# which mavpoll reads from a FILE so it never appears in `ps`, in
# this file, or in shell history. To enable, mount the token and
# append: "-zenmoney-token-file", "/run/secrets/zenmoney.token"
# (optionally "-zenmoney-interval", "1h"). Core never sees the
# token — the poller writes facts(kind=env, source=poll:zenmoney)
# and mavend only reads those back when he asks.
depends_on: [mavend]
volumes:
- sockets:/run/maven
# - ./deploy/zenmoney.token:/run/secrets/zenmoney.token:ro
# The mail reader (Vikunja #246) is OFF and commented out: it needs an IMAP
# account, and there is none on this box. mavmaild reads the password from a
# FILE so it never appears in `ps`, in this file, or in shell history — the
# same rule mavpoll follows for the zenmoney token. Core never sees the
# password: the reader hands core message text on one IPC method, and core
# writes what the model extracts as task CANDIDATES he reviews on /tasks.
# Nothing here can create a reminder, so a misread mail cannot fire.
#
# To enable: write the password to deploy/imap.password (0600, gitignored),
# add an "email": {} block to deploy/mavend.json, and uncomment this service.
# mavmaild:
# <<: *image
# command: ["mavmaild", "-socket", "/run/maven/mavend.sock",
# "-imap", "imap.example.org:993",
# "-user", "kami@example.org",
# "-password-file", "/run/secrets/imap.password",
# "-mailbox", "INBOX",
# "-interval", "15m",
# "-state", "/var/lib/maven/mail-seen.json"]
# depends_on: [mavend]
# volumes:
# - sockets:/run/maven
# - dbdata:/var/lib/maven
# - ./deploy/imap.password:/run/secrets/imap.password:ro
volumes:
dbdata:
+25
View File
@@ -27,3 +27,28 @@
7. Add voice query handler — `"что нового?"` queries `RecentNotes` filtered by source prefix `rss:` and phrases via `phraser.PhraseQuery`
8. Add `feeds` block to `config.Config` and `deploy/mavend.json`
9. Test with a live RSS feed (e.g., `https://news.ycombinator.com/rss`) — verify items appear in notes table
---
## Shipped 2026-08-01 (#258)
`internal/webfetch` (the guarded HTTP door: scheme, allow/deny hosts, private-address
refusal in the dialer, size cap, redirect cap, per-host rate limit), `internal/rss`
(RSS 2.0 + Atom parser, poller with durable marks and a keyword filter),
`cmd/mavend/feeds.go` (ticker, fetcher adapter, `rss:latest:<feed>` fact marks),
config block `feeds`, and the `feeds` query source with `router.ParseFeedQuery`.
Deviations from the plan above, both deliberate:
- **Step 5 (breaking-news nudges) was not built.** A feed that dispatches is a nag,
and the one thing Maven is not is a nag. Items are read when asked and nowhere else.
If breaking news is ever wanted, it belongs behind the existing delivery policy
(severity, quiet hours, digest), not in the poller.
- **Step 3 (embedder relevance) is a seam, not an implementation.** `rss.Ranker`
exists and is wired nil. Scoring items against an "interest profile" needs a
profile, and there is none yet; a threshold with nothing to compare against is a
random filter with a confident name. The filter that runs is the per-feed
include/exclude keyword list, which he can read and predict.
No new dependency: stdlib `encoding/xml`, no gofeed. Stock deploy config has no
`feeds` block, so the capability is off.
+38
View File
@@ -28,3 +28,41 @@
7. Add IPC methods `MethodTriggerCrawl(name)`, `MethodListCrawls`, `MethodGetCrawlResult(name)`
8. Add `crawls` block to `config.Config` and `deploy/mavend.json`
9. Test with a static HTML page — verify extraction matches expected values, verify scheduling fires correctly
## Shipped 2026-08-01 (#259)
Built as `internal/crawl` (pure: robots, extraction, watcher) plus
`cmd/mavend/crawls.go` (fetcher, ticker, dedup facts), on top of the guarded
`internal/webfetch` door added with the feed reader (#258). Off unless
configured, in two separately-switched halves: `crawl.on_demand` for a URL he
names, `crawl.watches` for a scheduled re-read.
**Limits are code, not documentation** (`internal/webfetch`, tested one test per
limit): host allowlist/denylist, no private addresses (loopback, RFC1918 —
hence the LAN and the `10.42.0.0/24` wg range —, link-local incl. cloud
metadata, CGNAT, v6 ULA) enforced in the dialer's `Control` hook so DNS
rebinding and every redirect hop are covered, response size cap, redirect cap,
timeout, one request per host per second. `robots.txt` is fetched first, cached
per host, and a `Disallow` is refused with no override.
Deliberate deviations from the plan above:
- **No CSS selectors and no LLM structured extraction** (steps 2). The output is
plaintext handed to the phraser as context for the question he asked. A 1.7B
extracting a JSON price table from 4000 runes is a worse bet than reading, and
`goquery` is not vendored.
- **No `crawl` act verb and no new IPC methods** (steps 5, 7). Reading a page is
a query source (`queryWeb` in `actions_query.go`, last in the chain, behind
Kiwix once that is wired), not an action he commands. Nothing needs a new wire
method to work.
- **Notes, not facts.** A page's text is not a fact about him. Only the dedup
hash is a fact (`crawl:hash:<name>`, kind `config`, source `poll:crawl`).
- **Nothing is dispatched.** A changed page writes a note; it does not nudge.
Not a nag.
- **No `/tools` crawl history page.** The notes and the hash facts are already
visible on `/dash`.
**No new dependency.** The vendored tree has no `x/net/html`, no `goquery` and
no `temoto/robotstxt`, so robots parsing and HTML-to-text are stdlib
(`regexp`, `html`) — RE2 has no backreferences, hence the `pairsRE` builder in
`extract.go`.
+6 -1
View File
@@ -74,7 +74,12 @@ func Requirement(m ipc.Method) Authority {
// existing analogue.
ipc.MethodCaptureTask,
ipc.MethodListTasks,
ipc.MethodSetTaskStatus:
ipc.MethodSetTaskStatus,
// Mail ingestion (Vikunja #246). AuthRead because of what the method can
// produce: candidate tasks and nothing else. It cannot write a fact, set a
// reminder, or touch the tool allowlist, so a compromised mail reader can
// at worst put junk on a review page he clears in one click.
ipc.MethodIngestMail:
return AuthRead
}
// Unknown method ⇒ AuthRead, but ipc.dispatch returns ErrUnknownMethod
+156
View File
@@ -150,6 +150,21 @@ type Config struct {
// absent ⇒ no evaluation loop at all. See MemoryEvalConfig.
MemoryEval *MemoryEvalConfig `json:"memory_eval,omitempty"`
// Email — mail ingestion (Vikunja #246). nil / absent ⇒ core refuses
// ipc.MethodIngestMail outright, so a mail reader cannot make Maven read a
// mailbox by merely existing. See EmailConfig; the IMAP host and credential
// live in the reader (cmd/mavmaild), never here.
Email *EmailConfig `json:"email,omitempty"`
// Feeds — RSS/Atom feed reading (Vikunja #258). nil / absent ⇒ no feed is
// ever fetched: reading the outside world is off unless configured, like
// the weather and telegram. See FeedsConfig.
Feeds *FeedsConfig `json:"feeds,omitempty"`
// Crawl — reading a web page (Vikunja #259). nil / absent ⇒ Maven never
// fetches a page: not on request, not on a schedule. See CrawlConfig.
Crawl *CrawlConfig `json:"crawl,omitempty"`
// Praxis — the ecosystem attention-state service. When configured, maven
// calls the Praxis HTTP tools API for attention listing and item lifecycle.
// Maven never touches Praxis's database directly (ecosystem invariant: no
@@ -396,6 +411,109 @@ func (p *PatternProposalConfig) AnnounceProposals() bool {
return p != nil && p.Notify
}
// FeedsConfig — the RSS/Atom reader (Vikunja #258, docs/plans/13-rss-news-feeds.md).
//
// Absent ⇒ off, and off means no outbound request at all. Present with an empty
// `sources` list is also off — a poller with nothing to poll is not wired.
//
// What a feed may NOT do here: speak. Items are written as notes with source
// "rss:<name>" and read back when he asks; nothing is dispatched, nudged or
// announced on arrival. That is the "not a nag" constraint, and it is why there
// is no severity or channel field in this block to reach for.
type FeedsConfig struct {
// Sources — the feeds to read. Empty ⇒ the reader stays down.
Sources []FeedSourceConfig `json:"sources,omitempty"`
// PollInterval — default per-feed cadence. 0 ⇒ rss.DefaultPollInterval (30m).
PollInterval Duration `json:"poll_interval,omitempty"`
// MaxItems — most items kept from one feed in one poll. 0 ⇒
// rss.DefaultMaxItems (5). This is the "не завали мне /dash" knob.
MaxItems int `json:"max_items,omitempty"`
// MaxAge — on a first poll (no saved mark), how far back to take items.
// 0 ⇒ rss.DefaultMaxAge (24h), so switching a feed on imports today, not
// the archive.
MaxAge Duration `json:"max_age,omitempty"`
// AllowHosts — when set, the reader may only connect to these hosts (and
// their subdomains). The feed URLs' own hosts are added automatically, so
// this is only needed to be stricter than that.
AllowHosts []string `json:"allow_hosts,omitempty"`
// Timeout — per-request budget. 0 ⇒ webfetch.DefaultTimeout.
Timeout Duration `json:"timeout,omitempty"`
// MaxBytes — response size cap. 0 ⇒ webfetch.DefaultMaxBytes (2 MiB).
MaxBytes int64 `json:"max_bytes,omitempty"`
}
// FeedSourceConfig — one feed.
type FeedSourceConfig struct {
Name string `json:"name"` // note source is "rss:<name>"
URL string `json:"url"` // http(s) only
Category string `json:"category,omitempty"` // "технологии" — what "что нового по X?" matches
Interval Duration `json:"interval,omitempty"` // 0 ⇒ FeedsConfig.PollInterval
Include []string `json:"include,omitempty"` // keep only items containing one of these
Exclude []string `json:"exclude,omitempty"` // drop items containing any of these
}
// CrawlConfig — the web crawler (Vikunja #259, docs/plans/14-web-crawler.md).
//
// Absent ⇒ off, and off means no page is ever fetched. Present with neither
// `on_demand` nor a `watches` entry is also off: there would be nothing to do.
//
// The crawler is the LAST place an answer is looked for, behind the model, his
// own memory and the local Kiwix ZIMs. That ordering lives in the query-source
// chain (cmd/mavend/actions_query.go), not here, but it is the reason this block
// is small: it is a fallback, not a search engine.
//
// Only the URL leaves the box. His notes, facts, persona block and history are
// never part of a request — the crawler package cannot even read the store.
type CrawlConfig struct {
// OnDemand — may he ask her to read a page he names out loud
// ("посмотри https://… — что там пишут?"). false ⇒ the on-demand answer
// source stays off and only the watches below run.
OnDemand bool `json:"on_demand,omitempty"`
// Watches — pages re-read on a schedule. A page whose text changed is
// written as a note (source "crawl:<name>"); nothing is announced.
Watches []CrawlWatchConfig `json:"watches,omitempty"`
// Interval — default watch cadence. 0 ⇒ crawl.DefaultWatchInterval (6h).
Interval Duration `json:"interval,omitempty"`
// AllowHosts — when set, the ONLY hosts the crawler may reach (subdomains
// included). Watched pages' own hosts are added automatically. Setting this
// is how "she may read the arch wiki and nothing else" is expressed.
AllowHosts []string `json:"allow_hosts,omitempty"`
// DenyHosts — never reachable, checked first. Private addresses do not need
// to be listed: they are refused unconditionally (see internal/webfetch).
DenyHosts []string `json:"deny_hosts,omitempty"`
// UserAgent — sent on every request AND matched against robots.txt groups.
// Empty ⇒ webfetch.DefaultUserAgent.
UserAgent string `json:"user_agent,omitempty"`
// Timeout — per-request budget. 0 ⇒ webfetch.DefaultTimeout.
Timeout Duration `json:"timeout,omitempty"`
// MaxBytes — response size cap. 0 ⇒ webfetch.DefaultMaxBytes (2 MiB).
MaxBytes int64 `json:"max_bytes,omitempty"`
// MaxRunes — how much extracted text is kept. 0 ⇒ crawl.DefaultMaxRunes
// (4000), which is what fits a 4096-token context alongside a prompt.
MaxRunes int `json:"max_runes,omitempty"`
}
// CrawlWatchConfig — one page kept an eye on.
type CrawlWatchConfig struct {
Name string `json:"name"` // note source is "crawl:<name>"
URL string `json:"url"`
Interval Duration `json:"interval,omitempty"` // 0 ⇒ CrawlConfig.Interval
}
// MemoryEvalConfig — the background memory-evaluation loop (Vikunja #248).
// Absent ⇒ off, like every other capability that costs something the owner did
// not ask for. Each evaluation is a full LLM round-trip on the one resident
@@ -418,6 +536,26 @@ type MemoryEvalConfig struct {
MinConfidence float64 `json:"min_confidence,omitempty"`
}
// EmailConfig — core's half of the email reader: how many task candidates one
// message may produce, and how long the extraction call may take.
//
// There is deliberately nothing about a mailbox here. Core does not connect to
// IMAP, does not know an account exists, and holds no mail credential — the
// reader daemon does, the same split mavpoll uses for the zenmoney token. This
// block only says "extraction is allowed, with these bounds".
type EmailConfig struct {
// MaxTasks — candidates per message. 0 ⇒ email.MaxCandidates (3).
MaxTasks int `json:"max_tasks,omitempty"`
// Timeout — per-message extraction budget. 0 ⇒ DefaultEmailTimeout. This is
// a Thinking model reading a mail; nobody is waiting on the answer, but a
// hung llama-server must not pin the reader's connection forever.
Timeout Duration `json:"timeout,omitempty"`
}
// DefaultEmailTimeout — extraction budget per message.
const DefaultEmailTimeout = 2 * time.Minute
// PhraserConfig — the LLM-backed phraser seam. The daemon spawns llama-server
// as a managed subprocess and sends chat-completion requests to phrase nudge
// and reminder messages. nil ⇒ the template-based Stub is used instead.
@@ -607,6 +745,24 @@ func (c *Config) applyDefaults() {
c.MemoryEval.Interval = Duration(DefaultMemoryEvalInterval)
}
// Same rule again: absent stays nil (⇒ mail ingestion refused), present gets
// the timeout default so `{}` is a valid "on with the defaults".
if c.Email != nil && c.Email.Timeout <= 0 {
c.Email.Timeout = Duration(DefaultEmailTimeout)
}
// A feeds block with no sources is the same as no block: nothing to poll,
// nothing wired. Normalising it to nil keeps that "off" in one place.
if c.Feeds != nil && len(c.Feeds.Sources) == 0 {
c.Feeds = nil
}
// Same rule for the crawler: a block that neither answers on demand nor
// watches anything has nothing to do, so it is normalised to "off".
if c.Crawl != nil && !c.Crawl.OnDemand && len(c.Crawl.Watches) == 0 {
c.Crawl = nil
}
if c.Voice != nil {
if c.Voice.RouterThreshold <= 0 {
c.Voice.RouterThreshold = DefaultRouterThreshold
+174
View File
@@ -0,0 +1,174 @@
// Package crawl reads a web page: fetch, robots check, HTML to text.
//
// It is the LAST place Maven looks for an answer, and that ordering is the whole
// design. "Never phones home" is deprecated, but what replaced it puts local
// sources first: the resident model, then his own memory, then the Kiwix ZIMs on
// the box (internal/kiwix), and only then the network. A local read costs
// nothing and leaks nothing; a fetch costs a round-trip and puts a URL in
// someone's access log. So this package exists to be the fallback, not the
// front door — see the querySources chain in cmd/mavend/actions_query.go for
// where it actually sits.
//
// What never leaves the box: his notes, his facts, the persona block, the
// conversation history. Only the URL is requested and, for the on-demand path,
// only because he said it out loud. Nothing here reads the store.
//
// The limits are not in this package — they are in internal/webfetch, which is
// the only way anything here touches a socket: http(s) only, host allow/deny,
// private-address refusal, size cap, redirect cap, per-host rate limit. What
// this package adds is politeness (robots.txt) and dedup.
package crawl
import (
"context"
"crypto/sha256"
"encoding/hex"
"errors"
"fmt"
"net/url"
"strings"
"time"
)
// Errors callers distinguish.
var (
ErrRobots = errors.New("crawl: robots.txt disallows this path")
ErrNotHTML = errors.New("crawl: response is not html or text")
)
// Fetcher is the guarded HTTP door (internal/webfetch adapted by the daemon). An
// interface so this package constructs no http.Client of its own and can be
// tested without a network.
type Fetcher interface {
Get(ctx context.Context, url string) (*Response, error)
}
// Response is the minimum a crawl needs from a fetch.
type Response struct {
URL string
ContentType string
Body []byte
}
// Config — crawler knobs.
type Config struct {
// UserAgent is the name matched against robots.txt groups. It must be the
// same string the fetcher sends, or Maven would be claiming one identity
// and obeying the rules for another.
UserAgent string
// MaxRunes caps extracted text. 0 ⇒ DefaultMaxRunes.
MaxRunes int
// RobotsTTL — how long a parsed robots.txt is trusted. 0 ⇒ 1h.
RobotsTTL time.Duration
// Now is injectable for tests. nil ⇒ time.Now.
Now func() time.Time
}
// Crawler fetches and extracts pages. Safe for concurrent use.
type Crawler struct {
fetch Fetcher
cfg Config
robots *robotsCache
}
// New builds a crawler. Returns nil when there is no fetcher, which is how the
// daemon expresses "crawling is off unless configured".
func New(fetch Fetcher, cfg Config) *Crawler {
if fetch == nil {
return nil
}
if cfg.UserAgent == "" {
cfg.UserAgent = "Maven"
}
if cfg.MaxRunes <= 0 {
cfg.MaxRunes = DefaultMaxRunes
}
if cfg.RobotsTTL <= 0 {
cfg.RobotsTTL = time.Hour
}
if cfg.Now == nil {
cfg.Now = time.Now
}
return &Crawler{fetch: fetch, cfg: cfg, robots: newRobotsCache(cfg.RobotsTTL)}
}
// Page fetches rawURL and returns its text. It checks robots.txt first and
// refuses a disallowed path with ErrRobots — there is no override.
func (c *Crawler) Page(ctx context.Context, rawURL string) (Page, error) {
u, err := url.Parse(strings.TrimSpace(rawURL))
if err != nil {
return Page{}, fmt.Errorf("crawl: bad url %q: %w", rawURL, err)
}
ok, err := c.allowed(ctx, u)
if err != nil {
return Page{}, err
}
if !ok {
return Page{}, fmt.Errorf("%w: %s", ErrRobots, u.Path)
}
resp, err := c.fetch.Get(ctx, u.String())
if err != nil {
return Page{}, err
}
// A PDF or an image is bytes Maven cannot read; saying so beats storing
// binary garbage as a "note".
ct := strings.ToLower(resp.ContentType)
if ct != "" && !strings.Contains(ct, "html") && !strings.Contains(ct, "text/") &&
!strings.Contains(ct, "xml") && !strings.Contains(ct, "json") {
return Page{}, fmt.Errorf("%w: %s", ErrNotHTML, resp.ContentType)
}
return Extract(resp.URL, resp.Body, c.cfg.MaxRunes), nil
}
// allowed consults robots.txt for u's host, reading it at most once per TTL.
//
// A robots.txt that cannot be fetched (404, a timeout, a blocked host) means
// allow, per the standard. The one thing that is NOT fail-open is an explicit
// Disallow.
func (c *Crawler) allowed(ctx context.Context, u *url.URL) (bool, error) {
host := u.Host
now := c.cfg.Now()
rules, ok := c.robots.get(host, now)
if !ok {
robotsURL := u.Scheme + "://" + host + "/robots.txt"
resp, err := c.fetch.Get(ctx, robotsURL)
switch {
case err != nil:
// Note what is NOT swallowed: a refusal from the guarded fetcher.
// If webfetch says this host is denied or private, the page fetch
// would fail the same way, and reporting the real reason beats
// reporting a robots verdict we never got.
if isFatalFetchError(err) {
return false, err
}
rules = Rules{}
default:
rules = ParseRobots(string(resp.Body), c.cfg.UserAgent)
}
c.robots.put(host, rules, now)
}
path := u.EscapedPath()
if u.RawQuery != "" {
path += "?" + u.RawQuery
}
return rules.Allowed(path), nil
}
// isFatalFetchError — a fetch failure that means "this host is off limits"
// rather than "there is no robots.txt here". The sentinel set is webfetch's, but
// this package must not import it (the interface exists precisely so it does
// not), so the check is on the message. Ugly and honest: the alternative is a
// dependency inversion for two strings.
func isFatalFetchError(err error) bool {
s := err.Error()
return strings.Contains(s, "not allowed") || strings.Contains(s, "private address") ||
strings.Contains(s, "only http and https")
}
// Hash is the dedup key for a crawl result: the sha256 of the extracted text,
// hex, first 16 chars. Text and not raw HTML, because a page whose only change
// is a rotating ad slot or a CSRF token has not changed.
func Hash(text string) string {
sum := sha256.Sum256([]byte(strings.TrimSpace(text)))
return hex.EncodeToString(sum[:])[:16]
}
+155
View File
@@ -0,0 +1,155 @@
package crawl
import (
"context"
"errors"
"strings"
"testing"
"time"
)
// fakeFetcher serves canned pages by URL and counts requests, so a test can
// assert that robots.txt was read once and that a refusal never reached the page.
type fakeFetcher struct {
pages map[string]Response
err error
calls []string
}
func (f *fakeFetcher) Get(_ context.Context, u string) (*Response, error) {
f.calls = append(f.calls, u)
if f.err != nil {
return nil, f.err
}
r, ok := f.pages[u]
if !ok {
return nil, errors.New("http 404")
}
if r.URL == "" {
r.URL = u
}
if r.ContentType == "" {
r.ContentType = "text/html; charset=utf-8"
}
return &r, nil
}
const htmlPage = `<html><head><title>Почему небо синее</title>
<style>body{color:red}</style><script>track()</script></head>
<body><nav>меню</nav><h1>Небо</h1>
<p>Свет рассеивается на молекулах воздуха.</p>
<p>Короткие волны рассеиваются сильнее.</p>
<footer>© 2026</footer></body></html>`
func newTestCrawler(f *fakeFetcher) *Crawler {
return New(f, Config{UserAgent: "Maven/1.0", Now: func() time.Time { return time.Unix(0, 0) }})
}
func TestPageExtractsText(t *testing.T) {
f := &fakeFetcher{pages: map[string]Response{
"https://example.org/sky": {Body: []byte(htmlPage)},
}}
page, err := newTestCrawler(f).Page(context.Background(), "https://example.org/sky")
if err != nil {
t.Fatal(err)
}
if page.Title != "Почему небо синее" {
t.Errorf("title = %q", page.Title)
}
if !strings.Contains(page.Text, "Свет рассеивается") {
t.Errorf("body text missing: %q", page.Text)
}
for _, junk := range []string{"track()", "color:red", "меню", "© 2026"} {
if strings.Contains(page.Text, junk) {
t.Errorf("%q survived extraction: %q", junk, page.Text)
}
}
}
func TestRobotsIsCheckedAndObeyed(t *testing.T) {
f := &fakeFetcher{pages: map[string]Response{
"https://example.org/robots.txt": {Body: []byte("User-agent: *\nDisallow: /secret\n"), ContentType: "text/plain"},
"https://example.org/secret/x": {Body: []byte(htmlPage)},
"https://example.org/open": {Body: []byte(htmlPage)},
}}
c := newTestCrawler(f)
if _, err := c.Page(context.Background(), "https://example.org/secret/x"); !errors.Is(err, ErrRobots) {
t.Fatalf("error = %v, want ErrRobots", err)
}
for _, u := range f.calls {
if strings.Contains(u, "/secret") {
t.Fatal("the disallowed page was fetched anyway")
}
}
if _, err := c.Page(context.Background(), "https://example.org/open"); err != nil {
t.Fatalf("allowed page: %v", err)
}
// robots.txt was read once for the host, not once per page.
robotsReads := 0
for _, u := range f.calls {
if strings.HasSuffix(u, "/robots.txt") {
robotsReads++
}
}
if robotsReads != 1 {
t.Fatalf("robots.txt read %d times, want 1", robotsReads)
}
}
// No robots.txt means allow — that is the standard, and the alternative makes
// most of the web unreadable.
func TestMissingRobotsAllows(t *testing.T) {
f := &fakeFetcher{pages: map[string]Response{
"https://example.org/page": {Body: []byte(htmlPage)},
}}
if _, err := newTestCrawler(f).Page(context.Background(), "https://example.org/page"); err != nil {
t.Fatalf("err = %v, want the page", err)
}
}
// A refusal from the guarded fetcher must surface as itself, not be laundered
// into "no robots.txt, go ahead".
func TestFetcherRefusalIsNotSwallowed(t *testing.T) {
f := &fakeFetcher{err: errors.New("webfetch: refusing to connect to a private address: 127.0.0.1")}
_, err := newTestCrawler(f).Page(context.Background(), "http://127.0.0.1:9100/mcp")
if err == nil || !strings.Contains(err.Error(), "private address") {
t.Fatalf("error = %v, want the fetcher's refusal", err)
}
}
func TestNonTextIsRefused(t *testing.T) {
f := &fakeFetcher{pages: map[string]Response{
"https://example.org/f.pdf": {Body: []byte("%PDF-1.7"), ContentType: "application/pdf"},
}}
if _, err := newTestCrawler(f).Page(context.Background(), "https://example.org/f.pdf"); !errors.Is(err, ErrNotHTML) {
t.Fatalf("error = %v, want ErrNotHTML", err)
}
}
func TestMaxRunesCapsText(t *testing.T) {
long := "<html><body><p>" + strings.Repeat("привет ", 2000) + "</p></body></html>"
f := &fakeFetcher{pages: map[string]Response{"https://example.org/l": {Body: []byte(long)}}}
c := New(f, Config{MaxRunes: 50})
page, err := c.Page(context.Background(), "https://example.org/l")
if err != nil {
t.Fatal(err)
}
if n := len([]rune(page.Text)); n > 51 {
t.Fatalf("text = %d runes, want the 50-rune cap", n)
}
}
func TestNewWithoutFetcherIsNil(t *testing.T) {
if New(nil, Config{}) != nil {
t.Fatal("a crawler with no fetcher must be nil — crawling is off unless configured")
}
}
func TestHashIgnoresNothingButText(t *testing.T) {
if Hash("a") == Hash("b") {
t.Fatal("different text hashed the same")
}
if Hash(" same \n") != Hash("same") {
t.Fatal("surrounding whitespace changed the hash")
}
}
+106
View File
@@ -0,0 +1,106 @@
package crawl
import (
"html"
"regexp"
"strings"
)
// HTML → text, with a regexp and no tokenizer.
//
// golang.org/x/net/html is not vendored and the network is not assumed, so this
// is stdlib. That is less of a compromise than it sounds: the unit of context
// here is a few hundred words for a 4096-token model to read, exactly like the
// Kiwix snippet, so what matters is dropping script/style/nav noise and keeping
// paragraph boundaries. A DOM would buy correctness on malformed markup that is
// then thrown away by truncation anyway.
//
// What this deliberately does NOT do: run JavaScript, follow links, or extract
// structured fields with CSS selectors or an LLM prompt. The plan's step 2 asked
// for the last of those; see docs/plans/14-web-crawler.md for why it was left
// out for now.
var (
// RE2 has no backreferences, so each tag pair is spelled out rather than
// captured and matched against itself.
dropRE = regexp.MustCompile(pairsRE("script", "style", "noscript", "svg", "head", "nav", "footer", "form"))
titleRE = regexp.MustCompile(`(?is)<title\b[^>]*>(.*?)</title>`)
h1RE = regexp.MustCompile(`(?is)<h1\b[^>]*>(.*?)</h1>`)
// Block-level tags become newlines so paragraphs survive as paragraphs.
blockRE = regexp.MustCompile(`(?is)</?(p|div|br|li|tr|h[1-6]|section|article|blockquote|pre)\b[^>]*>`)
tagRE = regexp.MustCompile(`(?s)<[^>]*>`)
commentRE = regexp.MustCompile(`(?s)<!--.*?-->`)
spaceRE = regexp.MustCompile(`[ \t\f\v]+`)
blankRE = regexp.MustCompile(`\n{2,}`)
)
// pairsRE builds `(?is)<tag …>…</tag>|…` for the given tags.
func pairsRE(tags ...string) string {
parts := make([]string, 0, len(tags))
for _, t := range tags {
parts = append(parts, `<`+t+`\b[^>]*>.*?</`+t+`>`)
}
return `(?is)` + strings.Join(parts, "|")
}
// Page is an extracted page.
type Page struct {
URL string
Title string
Text string // plain text, paragraphs separated by single newlines
}
// Extract turns a fetched HTML document into a Page. maxRunes caps the text (0 ⇒
// DefaultMaxRunes); the cap is on runes, not bytes, because a Russian page cut
// at a byte boundary ends in half a letter.
func Extract(url string, body []byte, maxRunes int) Page {
if maxRunes <= 0 {
maxRunes = DefaultMaxRunes
}
s := string(body)
s = commentRE.ReplaceAllString(s, " ")
title := firstGroup(titleRE, s)
if title == "" {
title = firstGroup(h1RE, s)
}
s = dropRE.ReplaceAllString(s, "\n")
s = blockRE.ReplaceAllString(s, "\n")
s = tagRE.ReplaceAllString(s, " ")
s = html.UnescapeString(s)
s = spaceRE.ReplaceAllString(s, " ")
var lines []string
for _, l := range strings.Split(s, "\n") {
if l = strings.TrimSpace(l); l != "" {
lines = append(lines, l)
}
}
text := blankRE.ReplaceAllString(strings.Join(lines, "\n"), "\n")
return Page{URL: url, Title: title, Text: TrimRunes(text, maxRunes)}
}
// DefaultMaxRunes — how much of a page is kept. ~4000 runes is a long answer's
// worth of context and still leaves room in a 4096-token window for the prompt
// and the reply.
const DefaultMaxRunes = 4000
func firstGroup(re *regexp.Regexp, s string) string {
m := re.FindStringSubmatch(s)
if len(m) < 2 {
return ""
}
t := tagRE.ReplaceAllString(m[1], " ")
return strings.TrimSpace(strings.Join(strings.Fields(html.UnescapeString(t)), " "))
}
// TrimRunes cuts s to at most max runes, on a rune boundary.
func TrimRunes(s string, max int) string {
r := []rune(s)
if len(r) <= max {
return s
}
return strings.TrimSpace(string(r[:max])) + "…"
}
+211
View File
@@ -0,0 +1,211 @@
package crawl
import (
"regexp"
"strings"
"sync"
"time"
)
// robots.txt, parsed the small way: no wildcards beyond the two the standard
// actually defines (`*` inside a path and `$` at the end), no sitemaps, no
// crawl-delay-per-agent gymnastics. A personal assistant reading a handful of
// pages does not need a spec-complete implementation; it needs to not be rude,
// and to be auditable in one sitting.
//
// Two rules worth stating because they are choices, not accidents:
//
// - a missing or unreadable robots.txt means ALLOW. That is what the standard
// says (404 ⇒ unrestricted), and the alternative would make a site that
// simply has no robots.txt unreadable;
// - an explicit Disallow means REFUSE, and Maven does not offer an override.
// There is no "but he asked me to" flag: the page is not read.
// Rules is a parsed robots.txt for one user-agent.
type Rules struct {
allow []string
disallow []string
// Delay is Crawl-delay in seconds when the group named one, 0 otherwise.
// The fetcher's own per-host rate limit is the floor; this can only make
// Maven slower, never faster.
Delay time.Duration
}
// ParseRobots reads robots.txt and returns the rules that apply to agent.
//
// Group selection follows the standard: the most specific matching group wins,
// which here means an exact user-agent match beats `*`. Lines that are neither
// are ignored rather than guessed at.
func ParseRobots(body string, agent string) Rules {
agent = strings.ToLower(agent)
type group struct {
agents []string
allow []string
disallow []string
delay time.Duration
}
var groups []group
var cur *group
// startNew tracks whether the next User-agent line opens a new group or
// joins the current one: consecutive User-agent lines share their rules.
startNew := true
for _, raw := range strings.Split(body, "\n") {
line := raw
if i := strings.IndexByte(line, '#'); i >= 0 {
line = line[:i]
}
line = strings.TrimSpace(line)
if line == "" {
continue
}
key, val, ok := strings.Cut(line, ":")
if !ok {
continue
}
key = strings.ToLower(strings.TrimSpace(key))
val = strings.TrimSpace(val)
switch key {
case "user-agent":
if startNew || cur == nil {
groups = append(groups, group{})
cur = &groups[len(groups)-1]
startNew = false
}
cur.agents = append(cur.agents, strings.ToLower(val))
case "disallow":
if cur == nil {
continue
}
startNew = true
// "Disallow:" with an empty value allows everything, and is not a
// path rule at all.
if val != "" {
cur.disallow = append(cur.disallow, val)
}
case "allow":
if cur == nil {
continue
}
startNew = true
if val != "" {
cur.allow = append(cur.allow, val)
}
case "crawl-delay":
if cur == nil {
continue
}
startNew = true
if d, err := time.ParseDuration(val + "s"); err == nil && d > 0 {
cur.delay = d
}
}
}
var star, exact *group
for i := range groups {
for _, a := range groups[i].agents {
if a == "*" && star == nil {
star = &groups[i]
}
// A robots.txt names "maven", we send "Maven/1.0 (…)": match on
// prefix, which is how every crawler reads this field.
if a != "*" && a != "" && strings.HasPrefix(agent, a) {
exact = &groups[i]
}
}
}
g := exact
if g == nil {
g = star
}
if g == nil {
return Rules{}
}
return Rules{allow: g.allow, disallow: g.disallow, Delay: g.delay}
}
// Allowed reports whether path may be fetched. Longest matching rule wins, and
// Allow beats Disallow at equal length — the standard's tie-break, and the one
// that makes "Disallow: /" plus "Allow: /public" mean what it looks like.
func (r Rules) Allowed(path string) bool {
if path == "" {
path = "/"
}
best, allowed := -1, true
for _, p := range r.disallow {
if n, ok := matchPath(p, path); ok && n > best {
best, allowed = n, false
}
}
for _, p := range r.allow {
if n, ok := matchPath(p, path); ok && n >= best {
best, allowed = n, true
}
}
return allowed
}
// matchPath applies a robots path pattern and returns the pattern's length as
// the specificity score. `*` matches any run of characters, `$` anchors the end.
// A pattern is a PREFIX match otherwise, which is what "Disallow: /admin" means.
func matchPath(pattern, path string) (int, bool) {
score := len(pattern)
re, err := robotsRegexp(pattern)
if err != nil {
return 0, false
}
return score, re.MatchString(path)
}
// robotsRegexp turns a robots path pattern into an anchored-at-the-start
// regexp. Everything but `*` and a trailing `$` is a literal, so the pattern is
// quoted first and the two metacharacters are put back afterwards.
func robotsRegexp(pattern string) (*regexp.Regexp, error) {
end := ""
if strings.HasSuffix(pattern, "$") {
pattern = strings.TrimSuffix(pattern, "$")
end = "$"
}
parts := strings.Split(pattern, "*")
for i, p := range parts {
parts[i] = regexp.QuoteMeta(p)
}
return regexp.Compile("^" + strings.Join(parts, ".*") + end)
}
// robotsCache holds parsed rules per host so a crawl of ten pages on one site
// reads robots.txt once. TTL because a site may change its mind, and a daemon
// that runs for weeks would otherwise never notice.
type robotsCache struct {
ttl time.Duration
mu sync.Mutex
m map[string]robotsEntry
}
type robotsEntry struct {
rules Rules
at time.Time
}
func newRobotsCache(ttl time.Duration) *robotsCache {
return &robotsCache{ttl: ttl, m: map[string]robotsEntry{}}
}
func (c *robotsCache) get(host string, now time.Time) (Rules, bool) {
c.mu.Lock()
defer c.mu.Unlock()
e, ok := c.m[host]
if !ok || now.Sub(e.at) > c.ttl {
return Rules{}, false
}
return e.rules, true
}
func (c *robotsCache) put(host string, r Rules, now time.Time) {
c.mu.Lock()
defer c.mu.Unlock()
c.m[host] = robotsEntry{rules: r, at: now}
}
+84
View File
@@ -0,0 +1,84 @@
package crawl
import (
"testing"
"time"
)
const robotsBody = `# a comment
User-agent: *
Disallow: /private
Disallow: /tmp/
Crawl-delay: 5
User-agent: Maven
Disallow: /
Allow: /public
`
func TestParseRobotsPicksTheMostSpecificGroup(t *testing.T) {
// The Maven group applies to us even though we send a longer UA string.
r := ParseRobots(robotsBody, "Maven/1.0 (self-hosted personal assistant)")
if r.Allowed("/anything") {
t.Error("Disallow: / in our own group was ignored")
}
if !r.Allowed("/public/page") {
t.Error("Allow: /public must beat the shorter Disallow: /")
}
// A different agent falls into the * group.
star := ParseRobots(robotsBody, "SomeoneElse/2")
if !star.Allowed("/anything") {
t.Error("the * group disallows nothing but /private and /tmp/")
}
if star.Allowed("/private/x") || star.Allowed("/tmp/") {
t.Error("the * group's disallows were not applied")
}
if star.Delay != 5*time.Second {
t.Errorf("crawl-delay = %v, want 5s", star.Delay)
}
}
func TestParseRobotsEmptyMeansAllowAll(t *testing.T) {
for _, body := range []string{"", "# nothing here\n", "User-agent: *\nDisallow:\n"} {
if !ParseRobots(body, "Maven").Allowed("/whatever") {
t.Errorf("body %q must allow everything", body)
}
}
}
func TestRobotsWildcards(t *testing.T) {
r := ParseRobots("User-agent: *\nDisallow: /*.pdf$\nDisallow: /a/*/secret\n", "Maven")
if r.Allowed("/docs/manual.pdf") {
t.Error("*.pdf$ did not match")
}
if !r.Allowed("/docs/manual.pdf.html") {
t.Error("$ must anchor at the end")
}
if r.Allowed("/a/b/secret") {
t.Error("/a/*/secret did not match")
}
if !r.Allowed("/a/b/public") {
t.Error("unrelated path was refused")
}
}
// Consecutive User-agent lines share one group, which is common in the wild.
func TestRobotsSharedGroup(t *testing.T) {
r := ParseRobots("User-agent: Googlebot\nUser-agent: Maven\nDisallow: /x\n", "Maven/1.0")
if r.Allowed("/x/y") {
t.Fatal("a shared group's rules were not applied to the second agent")
}
}
func TestRobotsCacheTTL(t *testing.T) {
c := newRobotsCache(time.Minute)
now := time.Now()
c.put("example.com", ParseRobots("User-agent: *\nDisallow: /\n", "Maven"), now)
if _, ok := c.get("example.com", now.Add(30*time.Second)); !ok {
t.Error("a fresh entry must be served from cache")
}
if _, ok := c.get("example.com", now.Add(2*time.Minute)); ok {
t.Error("an expired entry must be re-read")
}
}
+177
View File
@@ -0,0 +1,177 @@
package crawl
import (
"context"
"fmt"
"log"
"strings"
"time"
)
// Scheduled crawls: a page is re-read on an interval, and when its TEXT changed
// the new text is written as a note. Nothing is dispatched — same rule as the
// feed poller (Vikunja #258). A page that announced its own change would be a
// nag, and "the docs page changed" is not worth interrupting anyone for.
//
// Dedup is by content hash, so a page that re-renders identically writes nothing
// and a rotating ad slot does not count as news.
// WatchConfig — one page to keep an eye on.
type WatchConfig struct {
Name string // note source is "crawl:<Name>"
URL string // http(s), guarded by the fetcher
Interval time.Duration // 0 ⇒ Watcher's default
}
// Notes is core's note-writing half (same shape as ipc.CoreAPI's method).
type Notes interface {
WriteNote(ctx context.Context, ts time.Time, text string, embedding []float32, source string) (int64, error)
}
// Hashes remembers the last text hash per watch, durably, so a restart does not
// re-note an unchanged page. The daemon backs this with config facts
// ("crawl:hash:<name>").
type Hashes interface {
LastHash(ctx context.Context, name string) (string, error)
SetHash(ctx context.Context, name, hash string) error
}
// Embedder embeds a note on its way into the store. nil ⇒ no vector.
type Embedder interface {
Embed(ctx context.Context, text string) ([]float32, error)
}
// DefaultWatchInterval — pages change slowly, and every check is a request in
// someone's log.
const DefaultWatchInterval = 6 * time.Hour
// Watcher re-reads watched pages on their interval.
type Watcher struct {
c *Crawler
watches []WatchConfig
notes Notes
hashes Hashes
embed Embedder
interval time.Duration
nextDue map[string]time.Time
}
// NewWatcher wires the scheduled half, or returns nil when there is nothing to
// watch. Callers check for nil: no watches, no goroutine, no request.
func NewWatcher(c *Crawler, watches []WatchConfig, notes Notes, hashes Hashes, embed Embedder, defaultInterval time.Duration) *Watcher {
if c == nil || notes == nil {
return nil
}
var valid []WatchConfig
for _, w := range watches {
if strings.TrimSpace(w.Name) == "" || strings.TrimSpace(w.URL) == "" {
log.Printf("crawl: skipping a watch with no name or no url")
continue
}
valid = append(valid, w)
}
if len(valid) == 0 {
return nil
}
if defaultInterval <= 0 {
defaultInterval = DefaultWatchInterval
}
return &Watcher{
c: c, watches: valid, notes: notes, hashes: hashes, embed: embed,
interval: defaultInterval, nextDue: map[string]time.Time{},
}
}
// Watches returns the configured watches.
func (w *Watcher) Watches() []WatchConfig { return w.watches }
// CheckDue re-reads every watch whose interval elapsed and returns how many
// notes were written. Errors are logged per watch, never returned: one dead page
// must not stop the others.
func (w *Watcher) CheckDue(ctx context.Context, now time.Time) int {
written := 0
for _, watch := range w.watches {
if due, ok := w.nextDue[watch.Name]; ok && now.Before(due) {
continue
}
interval := watch.Interval
if interval <= 0 {
interval = w.interval
}
w.nextDue[watch.Name] = now.Add(interval)
changed, err := w.Check(ctx, watch, now)
if err != nil {
log.Printf("crawl: watch %s: %v", watch.Name, err)
continue
}
if changed {
log.Printf("crawl: watch %s: page changed, noted", watch.Name)
written++
}
}
return written
}
// Check re-reads one watch now and reports whether it wrote a note.
func (w *Watcher) Check(ctx context.Context, watch WatchConfig, now time.Time) (bool, error) {
page, err := w.c.Page(ctx, watch.URL)
if err != nil {
return false, err
}
// Title included: a page whose headline changed has changed.
h := Hash(page.Title + "\n" + page.Text)
if w.hashes != nil {
prev, err := w.hashes.LastHash(ctx, watch.Name)
if err != nil {
log.Printf("crawl: watch %s: read hash: %v", watch.Name, err)
}
if prev == h {
return false, nil
}
}
text := NoteText(watch, page)
var vec []float32
if w.embed != nil {
v, err := w.embed.Embed(ctx, text)
if err != nil {
log.Printf("crawl: watch %s: embed: %v", watch.Name, err)
} else {
vec = v
}
}
if _, err := w.notes.WriteNote(ctx, now, text, vec, SourceFor(watch.Name)); err != nil {
return false, fmt.Errorf("write note: %w", err)
}
if w.hashes != nil {
if err := w.hashes.SetHash(ctx, watch.Name, h); err != nil {
log.Printf("crawl: watch %s: save hash: %v", watch.Name, err)
}
}
return true, nil
}
// SourceFor is the note source for a watch, and SourcePrefix is what the answer
// path matches to recognise one.
func SourceFor(name string) string { return SourcePrefix + name }
// SourcePrefix — provenance for anything read off the network on a schedule.
const SourcePrefix = "crawl:"
// noteRunes — how much of a watched page goes into a note. Shorter than what the
// on-demand path reads: a note is a record of a change, not an archive.
const noteRunes = 800
// NoteText renders a watched page as a note body.
func NoteText(watch WatchConfig, page Page) string {
var b strings.Builder
if page.Title != "" {
b.WriteString(page.Title)
} else {
b.WriteString(watch.Name)
}
b.WriteString("\n")
b.WriteString(TrimRunes(page.Text, noteRunes))
b.WriteString("\n")
b.WriteString(watch.URL)
return b.String()
}
+117
View File
@@ -0,0 +1,117 @@
package crawl
import (
"context"
"strings"
"testing"
"time"
)
type note struct {
text string
source string
}
type fakeNotes struct{ notes []note }
func (n *fakeNotes) WriteNote(_ context.Context, _ time.Time, text string, _ []float32, source string) (int64, error) {
n.notes = append(n.notes, note{text, source})
return int64(len(n.notes)), nil
}
type fakeHashes struct{ m map[string]string }
func newHashes() *fakeHashes { return &fakeHashes{m: map[string]string{}} }
func (f *fakeHashes) LastHash(_ context.Context, name string) (string, error) {
return f.m[name], nil
}
func (f *fakeHashes) SetHash(_ context.Context, name, h string) error { f.m[name] = h; return nil }
var t0 = time.Date(2026, 8, 1, 9, 0, 0, 0, time.UTC)
func TestWatchNotesAChangedPage(t *testing.T) {
f := &fakeFetcher{pages: map[string]Response{
"https://example.org/docs": {Body: []byte(htmlPage)},
}}
notes := &fakeNotes{}
hashes := newHashes()
w := NewWatcher(newTestCrawler(f), []WatchConfig{{Name: "docs", URL: "https://example.org/docs"}},
notes, hashes, nil, time.Hour)
if w == nil {
t.Fatal("NewWatcher returned nil for a configured watch")
}
if n := w.CheckDue(context.Background(), t0); n != 1 {
t.Fatalf("first check wrote %d notes, want 1", n)
}
if notes.notes[0].source != "crawl:docs" {
t.Errorf("source = %q, want crawl:docs", notes.notes[0].source)
}
if !strings.Contains(notes.notes[0].text, "https://example.org/docs") {
t.Errorf("note does not carry the url: %q", notes.notes[0].text)
}
// Unchanged page, interval elapsed: nothing written.
if n := w.CheckDue(context.Background(), t0.Add(2*time.Hour)); n != 0 {
t.Fatalf("an unchanged page wrote %d notes", n)
}
// Changed page: one note.
f.pages["https://example.org/docs"] = Response{Body: []byte(strings.Replace(htmlPage, "синее", "серое", 1))}
if n := w.CheckDue(context.Background(), t0.Add(4*time.Hour)); n != 1 {
t.Fatalf("a changed page wrote %d notes, want 1", n)
}
}
func TestWatchIntervalIsRespected(t *testing.T) {
f := &fakeFetcher{pages: map[string]Response{"https://example.org/d": {Body: []byte(htmlPage)}}}
w := NewWatcher(newTestCrawler(f), []WatchConfig{{Name: "d", URL: "https://example.org/d", Interval: time.Hour}},
&fakeNotes{}, newHashes(), nil, 0)
w.CheckDue(context.Background(), t0)
before := len(f.calls)
w.CheckDue(context.Background(), t0.Add(time.Minute))
if len(f.calls) != before {
t.Fatal("the page was re-read inside its interval")
}
}
// The hash is durable so a restart does not re-note an unchanged page.
func TestWatchHashSurvivesRestart(t *testing.T) {
f := &fakeFetcher{pages: map[string]Response{"https://example.org/d": {Body: []byte(htmlPage)}}}
hashes := newHashes()
watches := []WatchConfig{{Name: "d", URL: "https://example.org/d"}}
NewWatcher(newTestCrawler(f), watches, &fakeNotes{}, hashes, nil, time.Hour).CheckDue(context.Background(), t0)
notes2 := &fakeNotes{}
NewWatcher(newTestCrawler(f), watches, notes2, hashes, nil, time.Hour).CheckDue(context.Background(), t0.Add(time.Hour))
if len(notes2.notes) != 0 {
t.Fatalf("a fresh watcher re-noted an unchanged page: %q", notes2.notes[0].text)
}
}
func TestWatchDeadPageDoesNotStopTheOthers(t *testing.T) {
f := &fakeFetcher{pages: map[string]Response{"https://example.org/live": {Body: []byte(htmlPage)}}}
notes := &fakeNotes{}
w := NewWatcher(newTestCrawler(f), []WatchConfig{
{Name: "dead", URL: "https://example.org/gone"},
{Name: "live", URL: "https://example.org/live"},
}, notes, newHashes(), nil, time.Hour)
if n := w.CheckDue(context.Background(), t0); n != 1 {
t.Fatalf("wrote %d notes, want 1 (the live page)", n)
}
if notes.notes[0].source != "crawl:live" {
t.Fatalf("source = %q", notes.notes[0].source)
}
}
func TestNoWatchesMeansNoWatcher(t *testing.T) {
c := newTestCrawler(&fakeFetcher{})
if NewWatcher(c, nil, &fakeNotes{}, nil, nil, 0) != nil {
t.Fatal("no watches must mean no watcher")
}
if NewWatcher(nil, []WatchConfig{{Name: "a", URL: "u"}}, &fakeNotes{}, nil, nil, 0) != nil {
t.Fatal("no crawler must mean no watcher")
}
if NewWatcher(c, []WatchConfig{{Name: "", URL: ""}}, &fakeNotes{}, nil, nil, 0) != nil {
t.Fatal("a watch with no name or url is not a configuration")
}
}
+226
View File
@@ -0,0 +1,226 @@
package email
import (
"context"
"encoding/json"
"fmt"
"strings"
"time"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/persona"
)
// Extraction — turning one mail into task CANDIDATES, and nothing else.
//
// The output of this file can only ever become rows in `tasks` with status
// "candidate" (store.TaskCandidate), written through the one intake seam
// (ipc.CaptureTaskReq, Vikunja #130). That bound is the whole design:
//
// - No reminder. A reminder FIRES; it speaks to him unprompted. A 1.7B that
// misreads "встреча была в четверг" as a future appointment would then wake
// him up about it. A candidate that is wrong is a line on a review page he
// dismisses in one click, which is the correct cost of a model being wrong
// about someone's mail.
// - No fact. A fact is a claim Maven will later recite as true. Nothing read
// out of a marketing mail deserves that standing.
// - No calendar event, no note, no action. Extraction writes candidates or
// writes nothing.
//
// The due date the model may return is stored on the candidate (tasks.due_ts),
// which no scheduler reads — it is there so the review page can sort by it.
//
// Privacy: the mail text goes to the resident model on this box and nowhere
// else. It is never search input (CLAUDE.md: "his notes and facts are never
// search input" — mail is the same class), and Evidence keeps only the subject
// line, so the review page shows him where a candidate came from without the
// store growing a copy of his mailbox.
// MaxCandidates — at most this many candidates per message, enforced by the
// grammar. A mail with four tasks in it is a mail he has to read himself; a
// model allowed ten will produce ten.
const MaxCandidates = 3
// SourcePrefix — provenance for everything this package captures. The mailbox
// name is appended: "email:INBOX". Same vocabulary as tap:voice / poll:netdata.
const SourcePrefix = "email:"
// Candidate — one piece of work the model thinks the mail is asking for.
type Candidate struct {
Text string `json:"text"`
// Due — "YYYY-MM-DD" or empty. A date the model read out of the text, not a
// date it computed: relative wording ("до пятницы") is left in Text, because
// a small model resolving "пятница" against today's date gets it wrong often
// enough that a stored wrong date is worse than no date.
Due string `json:"due"`
}
// Completer — the llama-server seam, same shape memeval and the router use, so
// the one resident model serves this caller too.
type Completer interface {
Complete(ctx context.Context, r llm.Req) (string, error)
}
// Extractor reads a message and returns candidates. It holds no store and no
// writer on purpose: this type cannot persist anything, so "extraction never
// acts" is a property of the code, not of a review.
type Extractor struct {
llm Completer
// MaxCandidates — 0 ⇒ MaxCandidates.
max int
// ContextBlock — the shared persona block, optional. Extraction output is
// not spoken, so the persona matters less here than in the phraser; it is
// wired anyway so a candidate reads in her voice on the review page.
contextBlock func() string
}
func NewExtractor(c Completer, max int, contextBlock func() string) *Extractor {
if max <= 0 || max > MaxCandidates {
max = MaxCandidates
}
return &Extractor{llm: c, max: max, contextBlock: contextBlock}
}
// extractGrammar — GBNF pinning the answer to a bounded array of fixed-shape
// candidates. Same reasoning as memeval's evalGrammar and the router's
// routeGrammar: the shape and the length bound are what keep a small model from
// drifting into prose or spending the token budget repeating one field.
//
// The empty array is reachable, deliberately: most mail contains no task, and a
// model with no way to say "nothing" invents something.
const extractGrammar = `
root ::= "[" ws (item ("," ws item){0,2})? ws "]"
item ::= "{" ws "\"text\"" ws ":" ws text "," ws "\"due\"" ws ":" ws due ws "}"
text ::= "\"" ([^"\\] | "\\" .){1,120} "\""
due ::= "\"\"" | "\"" [0-9]{4} "-" [0-9]{2} "-" [0-9]{2} "\""
ws ::= [ \t\n]*
`
// extractSystem — the extraction prompt.
//
// Written around the two failure modes a small model has on this task: it
// summarises when asked to extract (turning a mail into "письмо от Антона"),
// and it invents an obligation from any polite closing sentence. Hence the
// insistence on a verb phrase, and the explicit permission to return [].
const extractSystem = `Ты читаешь одно письмо из его почты и достаёшь из него дела, которые письмо от него требует.
Правила:
- Отвечай ТОЛЬКО массивом JSON. Каждый элемент: {"text": "...", "due": "ГГГГ-ММ-ДД" или ""}.
- text — короткая формулировка дела по-русски, с глаголом: "оплатить счёт за интернет", "отправить акт". Не пересказывай письмо и не описывай его.
- Дело — это то, что должен сделать ОН. Рассылка, реклама, уведомление, отчёт, письмо «просто к сведению» — дел не содержат.
- Если письмо ничего от него не требует, верни пустой массив []. Это нормальный ответ, так бывает чаще всего.
- Ничего не придумывай. Если срока в письме нет — "".
- due заполняй только когда в письме стоит конкретная дата. Слова вроде «до пятницы» оставь в text, дату не вычисляй.
- Максимум три дела. Лучше одно точное, чем три общих.`
// Extract returns the candidates in one message.
//
// Junk is refused without an LLM call — cheapest possible defence, and the
// reason the header filter exists. An empty message (no subject, no body) is
// likewise not worth a round trip.
//
// A parse failure is an error the caller logs and moves past. It is never
// silently turned into zero candidates, because "the model went off the rails"
// and "the mail contains no task" want different reactions from a human reading
// the log.
func (e *Extractor) Extract(ctx context.Context, msg Message) ([]Candidate, error) {
if msg.Junk {
return nil, nil
}
user := renderForModel(msg)
if user == "" {
return nil, nil
}
raw, err := e.llm.Complete(ctx, llm.Req{
System: persona.Prepend(e.contextBlock, extractSystem),
User: user,
Grammar: extractGrammar,
MaxTokens: 512,
RepeatPenalty: 1.1,
})
if err != nil {
return nil, fmt.Errorf("email: extract: %w", err)
}
items, err := parseCandidates(raw)
if err != nil {
// The raw reply is NOT in the error: it is a transformation of his mail,
// and this error reaches the daemon log.
return nil, fmt.Errorf("email: extract: unparsable reply (%d bytes)", len(raw))
}
out := make([]Candidate, 0, len(items))
seen := map[string]bool{}
for _, it := range items {
it.Text = strings.TrimSpace(it.Text)
if it.Text == "" {
continue
}
key := strings.ToLower(strings.Join(strings.Fields(it.Text), " "))
if seen[key] {
continue // the model repeating itself is not two tasks
}
seen[key] = true
if _, ok := ParseDue(it.Due); !ok {
it.Due = "" // a date the grammar allowed but the calendar does not
}
out = append(out, it)
if len(out) >= e.max {
break
}
}
return out, nil
}
// renderForModel is the user turn: subject, sender and body, labelled. Only
// these three fields — no headers, no recipient list, no message-id, nothing
// that would let the model start reasoning about routing metadata.
func renderForModel(msg Message) string {
var b strings.Builder
if msg.From != "" {
fmt.Fprintf(&b, "От: %s\n", msg.From)
}
if msg.Subject != "" {
fmt.Fprintf(&b, "Тема: %s\n", msg.Subject)
}
if msg.Body != "" {
fmt.Fprintf(&b, "\n%s\n", msg.Body)
}
if msg.Subject == "" && msg.Body == "" {
return ""
}
return b.String()
}
// parseCandidates decodes the grammar-constrained reply, tolerating the
// wrappers a Thinking model sometimes leaves around it (a fenced block, or
// leading reasoning before the array).
func parseCandidates(raw string) ([]Candidate, error) {
s := strings.TrimSpace(raw)
if i := strings.Index(s, "["); i > 0 {
s = s[i:]
}
if j := strings.LastIndex(s, "]"); j >= 0 {
s = s[:j+1]
}
var out []Candidate
if err := json.Unmarshal([]byte(s), &out); err != nil {
return nil, err
}
return out, nil
}
// ParseDue turns the model's "YYYY-MM-DD" into a time in UTC. Exported because
// the daemon-side intake stores it on the candidate.
//
// The zero-value/empty case returns ok=false rather than an error: no date is
// the common answer, not a failure.
func ParseDue(s string) (time.Time, bool) {
s = strings.TrimSpace(s)
if s == "" {
return time.Time{}, false
}
t, err := time.Parse("2006-01-02", s)
if err != nil {
return time.Time{}, false
}
return t, true
}
+143
View File
@@ -0,0 +1,143 @@
package email
import (
"context"
"strings"
"testing"
"github.com/kami/maven/internal/llm"
)
// fakeLLM returns a canned reply and records the request, so a test can assert
// on the grammar and on what of the mail was sent.
type fakeLLM struct {
reply string
err error
got llm.Req
calls int
}
func (f *fakeLLM) Complete(_ context.Context, r llm.Req) (string, error) {
f.calls++
f.got = r
return f.reply, f.err
}
func msgFor(subject, body string) Message {
return Message{UID: 1, From: "anton@example.org", Subject: subject, Body: body}
}
func TestExtractCandidates(t *testing.T) {
f := &fakeLLM{reply: `[{"text":"отправить акт","due":""},{"text":"оплатить счёт","due":"2026-08-05"}]`}
e := NewExtractor(f, 0, nil)
got, err := e.Extract(context.Background(), msgFor("Акт и счёт", "Надо отправить акт и оплатить счёт до 5 августа."))
if err != nil {
t.Fatalf("extract: %v", err)
}
if len(got) != 2 {
t.Fatalf("got %d candidates, want 2: %+v", len(got), got)
}
if got[0].Text != "отправить акт" || got[1].Due != "2026-08-05" {
t.Errorf("candidates = %+v", got)
}
if f.got.Grammar == "" {
t.Error("extraction must be grammar-constrained")
}
// The subject and body go to the model; nothing else about the message does.
if !strings.Contains(f.got.User, "Акт и счёт") || !strings.Contains(f.got.User, "оплатить счёт") {
t.Errorf("user turn = %q", f.got.User)
}
}
func TestExtractEmptyArrayIsNotAnError(t *testing.T) {
f := &fakeLLM{reply: "[]"}
got, err := NewExtractor(f, 0, nil).Extract(context.Background(), msgFor("FYI", "Просто к сведению."))
if err != nil || len(got) != 0 {
t.Fatalf("got (%v, %v), want (empty, nil) — no task is the normal answer", got, err)
}
}
// Junk must never reach the model: the header filter exists so the resident
// model is not spent on newsletters.
func TestExtractSkipsJunkWithoutCallingModel(t *testing.T) {
f := &fakeLLM{reply: `[{"text":"купить всё со скидкой","due":""}]`}
msg := msgFor("Скидки", "Sale!")
msg.Junk = true
got, err := NewExtractor(f, 0, nil).Extract(context.Background(), msg)
if err != nil || got != nil {
t.Fatalf("got (%v, %v), want (nil, nil)", got, err)
}
if f.calls != 0 {
t.Errorf("model called %d times for junk, want 0", f.calls)
}
}
func TestExtractEmptyMessageIsNotSent(t *testing.T) {
f := &fakeLLM{reply: "[]"}
if _, err := NewExtractor(f, 0, nil).Extract(context.Background(), Message{UID: 3}); err != nil {
t.Fatalf("extract: %v", err)
}
if f.calls != 0 {
t.Errorf("model called %d times for an empty message, want 0", f.calls)
}
}
func TestExtractCaps(t *testing.T) {
f := &fakeLLM{reply: `[{"text":"a","due":""},{"text":"b","due":""},{"text":"c","due":""}]`}
got, err := NewExtractor(f, 2, nil).Extract(context.Background(), msgFor("s", "b"))
if err != nil {
t.Fatalf("extract: %v", err)
}
if len(got) != 2 {
t.Errorf("got %d, want the configured cap of 2", len(got))
}
}
func TestExtractDropsRepeatsAndBadDates(t *testing.T) {
f := &fakeLLM{reply: `[{"text":"Отправить акт","due":"2026-02-31"},{"text":"отправить акт","due":""},{"text":" ","due":""}]`}
got, err := NewExtractor(f, 0, nil).Extract(context.Background(), msgFor("s", "b"))
if err != nil {
t.Fatalf("extract: %v", err)
}
if len(got) != 1 {
t.Fatalf("got %d candidates, want 1 (repeat and blank dropped): %+v", len(got), got)
}
if got[0].Due != "" {
t.Errorf("due = %q, want empty — 2026-02-31 is not a date", got[0].Due)
}
}
// A Thinking model sometimes wraps the array; and when it emits something
// unparsable the caller must hear about it rather than see "no tasks".
func TestParseCandidatesTolerance(t *testing.T) {
got, err := parseCandidates("думаю... [{\"text\":\"x\",\"due\":\"\"}] всё")
if err != nil || len(got) != 1 || got[0].Text != "x" {
t.Fatalf("got (%+v, %v)", got, err)
}
if _, err := parseCandidates("нет никакого JSON"); err == nil {
t.Error("unparsable output must be an error")
}
}
func TestExtractParseErrorHidesMailText(t *testing.T) {
f := &fakeLLM{reply: "он просил отправить акт, вот такой ответ"}
_, err := NewExtractor(f, 0, nil).Extract(context.Background(), msgFor("Акт", "секретный текст"))
if err == nil {
t.Fatal("want an error")
}
if strings.Contains(err.Error(), "акт") || strings.Contains(err.Error(), "секретный") {
t.Errorf("error text leaks mail content: %v", err)
}
}
func TestParseDue(t *testing.T) {
if _, ok := ParseDue(""); ok {
t.Error("empty due must be (zero, false)")
}
if got, ok := ParseDue("2026-08-05"); !ok || got.Year() != 2026 || got.Month() != 8 || got.Day() != 5 {
t.Errorf("ParseDue = (%v, %v)", got, ok)
}
if _, ok := ParseDue("05.08.2026"); ok {
t.Error("a non-ISO date must not parse")
}
}
+98
View File
@@ -0,0 +1,98 @@
package email
import (
"fmt"
"time"
)
// FetchSince is the whole read path in one call: connect, log in, examine the
// mailbox read-only, list what arrived since a date, fetch and parse the ones
// the caller has not seen, log out.
//
// It is a function rather than a long-lived object because a mail poller should
// not hold an authenticated session (and therefore his credential in a live TLS
// state) between polls. Connect, read, drop.
//
// skip decides which UIDs are already known — the poller's seen-set. max bounds
// one poll: a mailbox that received 400 messages overnight must not turn into
// 400 LLM calls, and the newest max are the ones a task could still be hiding
// in. Junk messages are returned too, flagged, so the caller can mark them seen
// without a second protocol round.
type FetchSince struct {
Addr string // host or host:993
User string
Mailbox string // e.g. "INBOX"
Timeout time.Duration
Since time.Time
Max int
Skip func(uid uint32) bool
}
// Run performs one read. password is passed here, not stored in the struct, so
// the configuration of a mailbox and the secret for it are never the same value
// sitting in the same place.
func (f FetchSince) Run(password string) ([]Message, error) {
return f.RunWith(password, nil)
}
// RunWith is Run with an explicit connection function, which is how the reader
// daemon and the tests substitute an in-process server. nil ⇒ Dial, i.e.
// implicit TLS with certificate verification; there is no configuration path
// that reaches this, so no deployment can end up talking cleartext IMAP.
func (f FetchSince) RunWith(password string, dial func(addr string, timeout time.Duration) (*Conn, error)) ([]Message, error) {
if f.Addr == "" || f.User == "" || f.Mailbox == "" {
return nil, fmt.Errorf("email: mailbox not configured (addr/user/mailbox)")
}
if dial == nil {
dial = Dial
}
c, err := dial(f.Addr, f.Timeout)
if err != nil {
return nil, err
}
defer c.Close()
if err := c.Login(f.User, password); err != nil {
return nil, err
}
defer c.Logout()
if err := c.Select(f.Mailbox); err != nil {
return nil, err
}
uids, err := c.SearchSince(f.Since)
if err != nil {
return nil, err
}
// Newest UIDs first — IMAP hands them back ascending, and when Max clips the
// list the recent mail is what matters.
wanted := make([]uint32, 0, len(uids))
for i := len(uids) - 1; i >= 0; i-- {
if f.Skip != nil && f.Skip(uids[i]) {
continue
}
wanted = append(wanted, uids[i])
if f.Max > 0 && len(wanted) >= f.Max {
break
}
}
out := make([]Message, 0, len(wanted))
for _, uid := range wanted {
raw, err := c.Fetch(uid)
if err != nil {
// One unreadable message does not abandon the poll; the rest of the
// mailbox is still worth reading. The error names the UID, not the
// message.
return out, fmt.Errorf("email: fetch uid %d: %w", uid, err)
}
if len(raw) == 0 {
continue // vanished between SEARCH and FETCH
}
msg, err := ParseMessage(uid, raw)
if err != nil {
continue // unparsable headers — nothing to review, skip silently
}
out = append(out, msg)
}
return out, nil
}
+49
View File
@@ -0,0 +1,49 @@
package email
import (
"net"
"strings"
"testing"
"time"
)
func TestFetchSinceRun(t *testing.T) {
mk := func(subject string) string {
return "Subject: " + subject + "\r\nContent-Type: text/plain; charset=utf-8\r\n\r\nbody\r\n"
}
f := &fakeIMAP{
uids: []uint32{1, 2, 3},
msgs: map[uint32]string{1: mk("one"), 2: mk("two"), 3: mk("three")},
}
fs := FetchSince{
Addr: "mail.example:993", User: "kami", Mailbox: "INBOX",
Timeout: 5 * time.Second,
Since: time.Date(2026, 7, 30, 0, 0, 0, 0, time.UTC),
Max: 2,
Skip: func(uid uint32) bool { return uid == 3 },
}
msgs, err := fs.RunWith("secret", func(addr string, timeout time.Duration) (*Conn, error) {
cli, srv := net.Pipe()
go f.serve(t, srv)
return NewConn(cli, timeout)
})
if err != nil {
t.Fatalf("run: %v", err)
}
// Newest first, the already-seen UID skipped, Max respected.
if len(msgs) != 2 {
t.Fatalf("got %d messages, want 2: %+v", len(msgs), msgs)
}
if msgs[0].Subject != "two" || msgs[1].Subject != "one" {
t.Errorf("subjects = %q,%q, want two,one (newest first)", msgs[0].Subject, msgs[1].Subject)
}
if strings.Contains(strings.Join(f.cmds, " "), "UID FETCH 3") {
t.Error("a skipped UID must not be fetched again")
}
}
func TestFetchSinceRequiresConfig(t *testing.T) {
if _, err := (FetchSince{}).Run("secret"); err == nil {
t.Fatal("an unconfigured mailbox must not be read")
}
}
+280
View File
@@ -0,0 +1,280 @@
package email
import (
"bufio"
"crypto/tls"
"fmt"
"io"
"net"
"regexp"
"strconv"
"strings"
"time"
)
// A minimal IMAP4rev1 client — LOGIN, SELECT, UID SEARCH, UID FETCH with
// BODY.PEEK, LOGOUT, and nothing else.
//
// Why hand-rolled instead of go-imap: the whole surface Maven needs is five
// commands, and this is the one code path that holds his mailbox credential and
// reads his private mail. A ~200-line client with no dependencies is auditable
// in one sitting; a general-purpose IMAP library is a much larger amount of
// code doing much more than we asked, in the most sensitive place in the tree.
// If IDLE, CONDSTORE or server-side threading ever become worth having, that
// trade should be re-made deliberately.
//
// BODY.PEEK[] rather than BODY[] is load-bearing: Maven reads his mail and must
// leave no trace of having done so. Reading a message here does not mark it
// \Seen, so the unread state in his own mail client stays his.
// DefaultIMAPPort — implicit-TLS IMAP. There is no cleartext and no STARTTLS
// path in this client: an option to send his password over a plain socket is an
// option to get it wrong once.
const DefaultIMAPPort = "993"
// Conn — one authenticated IMAP connection. Not safe for concurrent use; the
// poller drives one connection at a time.
type Conn struct {
rwc io.ReadWriteCloser
r *bufio.Reader
tag int
timeout time.Duration
}
// Dial opens an implicit-TLS connection and reads the server greeting.
func Dial(addr string, timeout time.Duration) (*Conn, error) {
host, _, err := net.SplitHostPort(addr)
if err != nil {
host, addr = addr, net.JoinHostPort(addr, DefaultIMAPPort)
}
d := &net.Dialer{Timeout: timeout}
// ServerName is set from the host we asked for: certificate verification is
// the only thing standing between his password and a MITM on the way out.
c, err := tls.DialWithDialer(d, "tcp", addr, &tls.Config{ServerName: host, MinVersion: tls.VersionTLS12})
if err != nil {
return nil, fmt.Errorf("email: dial %s: %w", addr, err)
}
return NewConn(c, timeout)
}
// NewConn wraps an already-open stream (the tests speak IMAP over a pipe) and
// consumes the greeting.
func NewConn(rwc io.ReadWriteCloser, timeout time.Duration) (*Conn, error) {
c := &Conn{rwc: rwc, r: bufio.NewReaderSize(rwc, 64<<10), timeout: timeout}
line, err := c.readLine()
if err != nil {
return nil, fmt.Errorf("email: greeting: %w", err)
}
if !strings.HasPrefix(line, "* OK") && !strings.HasPrefix(line, "* PREAUTH") {
c.rwc.Close()
return nil, fmt.Errorf("email: server refused connection: %s", line)
}
return c, nil
}
func (c *Conn) Close() error { return c.rwc.Close() }
// Login authenticates with LOGIN. The password is passed as an argument and
// never stored on the Conn: nothing in this package keeps a credential alive
// past the command that uses it, so no struct dump or panic trace can carry it.
func (c *Conn) Login(user, pass string) error {
// The command line itself is never logged (see exec) — a LOGIN line IS the
// credential.
if _, err := c.exec(fmt.Sprintf("LOGIN %s %s", quote(user), quote(pass))); err != nil {
return fmt.Errorf("email: login: %w", err)
}
return nil
}
// Select opens a mailbox read-only. EXAMINE, not SELECT: read-only at the
// protocol level means no command in this session can change a flag, expunge a
// message, or move anything, even by mistake.
func (c *Conn) Select(mailbox string) error {
if _, err := c.exec(fmt.Sprintf("EXAMINE %s", quote(mailbox))); err != nil {
return fmt.Errorf("email: examine %s: %w", mailbox, err)
}
return nil
}
// SearchSince returns the UIDs of messages received on or after since. An
// unlimited search is not offered: the first poll against a years-old mailbox
// would otherwise fetch everything and hand a decade of mail to the model.
//
// The IMAP SINCE key has date granularity (and compares the server's internal
// date), so the result can include messages slightly older than since. The
// caller dedupes by UID anyway, so a wider window costs one extra fetch.
func (c *Conn) SearchSince(since time.Time) ([]uint32, error) {
cmd := fmt.Sprintf("UID SEARCH SINCE %s", since.Format("2-Jan-2006"))
lines, err := c.exec(cmd)
if err != nil {
return nil, fmt.Errorf("email: search: %w", err)
}
var uids []uint32
for _, l := range lines {
rest, ok := untagged(l, "SEARCH")
if !ok {
continue
}
for _, f := range strings.Fields(rest) {
n, err := strconv.ParseUint(f, 10, 32)
if err == nil {
uids = append(uids, uint32(n))
}
}
}
return uids, nil
}
var literalSize = regexp.MustCompile(`\{(\d+)\}$`)
// Fetch returns the raw RFC 5322 bytes of one message, by UID.
//
// Returns (nil, nil) when the UID no longer exists — a message he deleted
// between SEARCH and FETCH is normal, not an error.
func (c *Conn) Fetch(uid uint32) ([]byte, error) {
tag := c.nextTag()
if err := c.send(fmt.Sprintf("%s UID FETCH %d (BODY.PEEK[])", tag, uid)); err != nil {
return nil, err
}
var raw []byte
for {
line, err := c.readLine()
if err != nil {
return nil, fmt.Errorf("email: fetch %d: %w", uid, err)
}
if done, err := c.tagged(tag, line); done {
if err != nil {
return nil, fmt.Errorf("email: fetch %d: %w", uid, err)
}
return raw, nil
}
m := literalSize.FindStringSubmatch(strings.TrimSpace(line))
if m == nil {
continue
}
n, err := strconv.Atoi(m[1])
if err != nil {
continue
}
buf := make([]byte, n)
if _, err := io.ReadFull(c.r, buf); err != nil {
return nil, fmt.Errorf("email: fetch %d: literal: %w", uid, err)
}
if raw == nil {
raw = buf
}
}
}
// Logout ends the session politely. A failure is not worth reporting — the
// connection is being closed either way.
func (c *Conn) Logout() {
_, _ = c.exec("LOGOUT")
}
// ---- protocol plumbing -----------------------------------------------------
func (c *Conn) nextTag() string {
c.tag++
return fmt.Sprintf("a%03d", c.tag)
}
// exec sends one command and returns the untagged response lines.
//
// Neither the command nor the response is ever logged here. LOGIN goes through
// this function, and a debug line "sent: a001 LOGIN ..." is how a credential
// ends up in a log file forever.
func (c *Conn) exec(cmd string) ([]string, error) {
tag := c.nextTag()
if err := c.send(tag + " " + cmd); err != nil {
return nil, err
}
var lines []string
for {
line, err := c.readLine()
if err != nil {
return nil, err
}
if done, err := c.tagged(tag, line); done {
return lines, err
}
lines = append(lines, line)
// A response line may carry a literal (e.g. a header FETCH). Nothing we
// send asks for one outside Fetch, but skip it if it appears so the
// stream stays aligned.
if m := literalSize.FindStringSubmatch(strings.TrimSpace(line)); m != nil {
if n, err := strconv.Atoi(m[1]); err == nil {
if _, err := io.CopyN(io.Discard, c.r, int64(n)); err != nil {
return nil, err
}
}
}
}
}
// tagged reports whether line completes the command with this tag, and turns a
// NO/BAD completion into an error. The error text is the server's, which never
// echoes a password.
func (c *Conn) tagged(tag, line string) (bool, error) {
if !strings.HasPrefix(line, tag+" ") {
return false, nil
}
rest := strings.TrimSpace(line[len(tag):])
switch {
case strings.HasPrefix(rest, "OK"):
return true, nil
case strings.HasPrefix(rest, "NO"), strings.HasPrefix(rest, "BAD"):
return true, fmt.Errorf("server said: %s", rest)
default:
return true, fmt.Errorf("unexpected completion: %s", rest)
}
}
func (c *Conn) send(line string) error {
c.setDeadline()
if _, err := io.WriteString(c.rwc, line+"\r\n"); err != nil {
return fmt.Errorf("email: write: %w", err)
}
return nil
}
func (c *Conn) readLine() (string, error) {
c.setDeadline()
line, err := c.r.ReadString('\n')
if err != nil {
return "", err
}
return strings.TrimRight(line, "\r\n"), nil
}
// setDeadline applies the per-connection timeout when the transport supports
// one. A hung IMAP server must not park the poller forever.
func (c *Conn) setDeadline() {
if c.timeout <= 0 {
return
}
if d, ok := c.rwc.(interface{ SetDeadline(time.Time) error }); ok {
_ = d.SetDeadline(time.Now().Add(c.timeout))
}
}
// untagged splits "* SEARCH 1 2 3" into its payload when the key matches.
func untagged(line, key string) (string, bool) {
if !strings.HasPrefix(line, "* ") {
return "", false
}
rest := strings.TrimSpace(line[2:])
if !strings.HasPrefix(rest, key) {
return "", false
}
return strings.TrimSpace(rest[len(key):]), true
}
// quote renders an IMAP quoted string. Passwords routinely contain characters
// that would otherwise end the argument early, and CR/LF are stripped rather
// than escaped because there is no legal way to send them — a credential file
// with a stray newline must not become a second command.
func quote(s string) string {
s = strings.NewReplacer("\r", "", "\n", "").Replace(s)
return `"` + strings.NewReplacer(`\`, `\\`, `"`, `\"`).Replace(s) + `"`
}
+162
View File
@@ -0,0 +1,162 @@
package email
import (
"bufio"
"fmt"
"net"
"strconv"
"strings"
"testing"
"time"
)
// fakeIMAP is a scripted server: enough of IMAP to exercise the client, and
// nothing more. It records the commands it received so a test can assert on the
// protocol (BODY.PEEK rather than BODY, EXAMINE rather than SELECT).
type fakeIMAP struct {
msgs map[uint32]string
uids []uint32
cmds []string
failOn string // substring of a command to answer NO
}
func (f *fakeIMAP) serve(t *testing.T, c net.Conn) {
t.Helper()
defer c.Close()
fmt.Fprint(c, "* OK fake IMAP ready\r\n")
r := bufio.NewReader(c)
for {
line, err := r.ReadString('\n')
if err != nil {
return
}
line = strings.TrimRight(line, "\r\n")
parts := strings.SplitN(line, " ", 2)
if len(parts) != 2 {
return
}
tag, cmd := parts[0], parts[1]
f.cmds = append(f.cmds, cmd)
if f.failOn != "" && strings.Contains(cmd, f.failOn) {
fmt.Fprintf(c, "%s NO computer says no\r\n", tag)
continue
}
upper := strings.ToUpper(cmd)
switch {
case strings.HasPrefix(upper, "LOGIN"), strings.HasPrefix(upper, "EXAMINE"):
fmt.Fprintf(c, "%s OK done\r\n", tag)
case strings.HasPrefix(upper, "UID SEARCH"):
var ids []string
for _, u := range f.uids {
ids = append(ids, strconv.FormatUint(uint64(u), 10))
}
fmt.Fprintf(c, "* SEARCH %s\r\n", strings.Join(ids, " "))
fmt.Fprintf(c, "%s OK search done\r\n", tag)
case strings.HasPrefix(upper, "UID FETCH"):
uid64, _ := strconv.ParseUint(strings.Fields(cmd)[2], 10, 32)
raw, ok := f.msgs[uint32(uid64)]
if ok {
fmt.Fprintf(c, "* 1 FETCH (UID %d BODY[] {%d}\r\n", uid64, len(raw))
fmt.Fprint(c, raw)
fmt.Fprint(c, ")\r\n")
}
fmt.Fprintf(c, "%s OK fetch done\r\n", tag)
case strings.HasPrefix(upper, "LOGOUT"):
fmt.Fprint(c, "* BYE\r\n")
fmt.Fprintf(c, "%s OK bye\r\n", tag)
return
default:
fmt.Fprintf(c, "%s BAD unknown\r\n", tag)
}
}
}
// dialFake wires a client Conn to an in-process server over net.Pipe.
func dialFake(t *testing.T, f *fakeIMAP) *Conn {
t.Helper()
cli, srv := net.Pipe()
go f.serve(t, srv)
c, err := NewConn(cli, 5*time.Second)
if err != nil {
t.Fatalf("greeting: %v", err)
}
t.Cleanup(func() { c.Close() })
return c
}
func TestIMAPRoundTrip(t *testing.T) {
body := "Subject: hello\r\nContent-Type: text/plain; charset=utf-8\r\n\r\nCall the bank.\r\n"
f := &fakeIMAP{uids: []uint32{4, 9}, msgs: map[uint32]string{4: body, 9: body}}
c := dialFake(t, f)
if err := c.Login("kami", `pa"ss\word`); err != nil {
t.Fatalf("login: %v", err)
}
if err := c.Select("INBOX"); err != nil {
t.Fatalf("select: %v", err)
}
uids, err := c.SearchSince(time.Date(2026, 8, 1, 0, 0, 0, 0, time.UTC))
if err != nil {
t.Fatalf("search: %v", err)
}
if len(uids) != 2 || uids[0] != 4 || uids[1] != 9 {
t.Fatalf("uids = %v, want [4 9]", uids)
}
raw, err := c.Fetch(9)
if err != nil {
t.Fatalf("fetch: %v", err)
}
if string(raw) != body {
t.Errorf("fetched %q, want the literal verbatim", raw)
}
c.Logout()
joined := strings.Join(f.cmds, "\n")
// Read-only at the protocol level, and peeking — Maven must leave no trace
// of having read his mail.
if !strings.Contains(joined, "EXAMINE") || strings.Contains(joined, "SELECT ") {
t.Errorf("want EXAMINE (read-only), got:\n%s", joined)
}
if !strings.Contains(joined, "BODY.PEEK[]") {
t.Errorf("want BODY.PEEK, got:\n%s", joined)
}
// The password must have been quoted and escaped, not truncated at the quote.
if !strings.Contains(joined, `"pa\"ss\\word"`) {
t.Errorf("password not quoted correctly:\n%s", joined)
}
// SINCE must carry the IMAP date form.
if !strings.Contains(joined, "SINCE 1-Aug-2026") {
t.Errorf("want a SINCE date, got:\n%s", joined)
}
}
func TestIMAPServerNoIsAnError(t *testing.T) {
f := &fakeIMAP{failOn: "LOGIN"}
c := dialFake(t, f)
err := c.Login("kami", "wrong")
if err == nil {
t.Fatal("a NO completion must be an error")
}
// The error is the server's text; it must not echo the credential.
if strings.Contains(err.Error(), "wrong") {
t.Errorf("error leaks the password: %v", err)
}
}
func TestIMAPFetchMissingUID(t *testing.T) {
f := &fakeIMAP{uids: []uint32{1}, msgs: map[uint32]string{}}
c := dialFake(t, f)
raw, err := c.Fetch(1)
if err != nil {
t.Fatalf("fetch: %v", err)
}
if raw != nil {
t.Errorf("a vanished UID should give nil, got %q", raw)
}
}
func TestQuoteStripsNewlines(t *testing.T) {
if got := quote("pass\r\nA1 LOGOUT"); strings.ContainsAny(got, "\r\n") {
t.Errorf("quote kept a line break: %q", got)
}
}
+80
View File
@@ -0,0 +1,80 @@
package email
import (
"net/mail"
"strings"
)
// The junk filter — the cheapest and most important half of reading mail.
//
// A mailbox is mostly machine-generated: newsletters, receipts nobody acts on,
// social notifications, marketing. Sending all of it to a 1.7B and asking "is
// there a task here" produces confident nonsense at a rate proportional to the
// volume, so junk is decided by HEADERS, before any model sees the message.
//
// The rules are all bulk-mail markers that senders set on themselves, never
// guesses about content:
//
// - List-Unsubscribe / List-Id — by definition a mailing list. If he can
// unsubscribe from it, it is not asking him to do anything.
// - Precedence: bulk|junk|list — the sender declaring itself bulk.
// - Auto-Submitted other than "no" (RFC 3834) — generated by a machine.
// - X-Spam-Flag: YES, X-Spam-Status: Yes — the spam filter upstream already
// decided; we do not second-guess it in the other direction.
// - X-GM-LABELS / X-Gmail-Labels containing a Gmail category — Gmail's own
// Promotions/Social/Forums/Spam classification, when the server sends it.
//
// Deliberately NOT here: sender allow/deny lists and subject keyword matching.
// Both are configuration that ages badly and both would be a place for his
// contacts to end up in a config file. If a real correspondent's mail is being
// dropped, the fix is a rule about a header, not a list of names.
//
// A junk verdict never deletes anything and never touches a flag on the server.
// It means "do not spend the model on this", nothing more.
// junkHeaders — headers whose mere presence marks bulk mail.
var junkPresence = []string{"List-Unsubscribe", "List-Id", "List-Post"}
// gmailCategories — Gmail's category labels, lowercased as they appear in
// X-GM-LABELS. "important" and "inbox" are labels too, and are NOT categories.
// Matching is by these exact tokens (substring is fine — they are namespaced
// and cannot appear in a hand-made label by accident), so a user label named
// "Social Club" is not mistaken for Gmail's Social category.
var gmailCategories = []string{
"category_promotions", "category_social", "category_forums", "category_updates",
`\spam`, `\junk`,
}
// classifyJunk returns whether the message is bulk/automated and why. The
// reason is a short header name, safe to log — it names the marker, never the
// sender or the subject.
func classifyJunk(h mail.Header) (bool, string) {
for _, name := range junkPresence {
if strings.TrimSpace(h.Get(name)) != "" {
return true, strings.ToLower(name)
}
}
switch strings.ToLower(strings.TrimSpace(h.Get("Precedence"))) {
case "bulk", "junk", "list":
return true, "precedence"
}
if v := strings.ToLower(strings.TrimSpace(h.Get("Auto-Submitted"))); v != "" && v != "no" {
return true, "auto-submitted"
}
if strings.EqualFold(strings.TrimSpace(h.Get("X-Spam-Flag")), "yes") {
return true, "x-spam-flag"
}
if v := strings.ToLower(strings.TrimSpace(h.Get("X-Spam-Status"))); strings.HasPrefix(v, "yes") {
return true, "x-spam-status"
}
labels := strings.ToLower(h.Get("X-GM-LABELS") + " " + h.Get("X-Gmail-Labels"))
for _, c := range gmailCategories {
if c == "" {
continue
}
if strings.Contains(labels, c) {
return true, "gmail-category"
}
}
return false, ""
}
+59
View File
@@ -0,0 +1,59 @@
package email
import (
"net/mail"
"strings"
"testing"
)
func headers(t *testing.T, raw string) mail.Header {
t.Helper()
m, err := mail.ReadMessage(strings.NewReader(strings.ReplaceAll(raw, "\n", "\r\n") + "\r\n\r\nbody\r\n"))
if err != nil {
t.Fatalf("read headers: %v", err)
}
return m.Header
}
func TestClassifyJunk(t *testing.T) {
cases := []struct {
name string
raw string
junk bool
reason string
}{
{"personal", "From: a@b.c\nSubject: привет", false, ""},
{"list-unsubscribe", "From: a@b.c\nList-Unsubscribe: <mailto:u@b.c>", true, "list-unsubscribe"},
{"list-id", "From: a@b.c\nList-Id: <golang-nuts.example>", true, "list-id"},
{"precedence bulk", "From: a@b.c\nPrecedence: bulk", true, "precedence"},
{"auto-submitted", "From: a@b.c\nAuto-Submitted: auto-generated", true, "auto-submitted"},
{"auto-submitted no", "From: a@b.c\nAuto-Submitted: no", false, ""},
{"spam flag", "From: a@b.c\nX-Spam-Flag: YES", true, "x-spam-flag"},
{"spam status", "From: a@b.c\nX-Spam-Status: Yes, score=9.1", true, "x-spam-status"},
{"spam status no", "From: a@b.c\nX-Spam-Status: No, score=0.1", false, ""},
{"gmail promo", "From: a@b.c\nX-Gmail-Labels: Inbox,CATEGORY_PROMOTIONS", true, "gmail-category"},
{"user label", "From: a@b.c\nX-Gmail-Labels: Social Club,Important", false, ""},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
junk, reason := classifyJunk(headers(t, c.raw))
if junk != c.junk || reason != c.reason {
t.Errorf("classifyJunk = (%v, %q), want (%v, %q)", junk, reason, c.junk, c.reason)
}
})
}
}
func TestNewsletterFixtureIsJunk(t *testing.T) {
msg, err := ParseMessage(9, fixture(t, "newsletter.eml"))
if err != nil {
t.Fatalf("parse: %v", err)
}
if !msg.Junk {
t.Fatal("a newsletter with List-Unsubscribe + Precedence: bulk must be junk")
}
// The reason is what gets logged, so it must never carry mail content.
if strings.Contains(msg.JunkReason, "@") || strings.Contains(msg.JunkReason, "Скидки") {
t.Errorf("junk reason leaks content: %q", msg.JunkReason)
}
}
+258
View File
@@ -0,0 +1,258 @@
// Package email is the reading half of the email reader (Vikunja #246,
// docs/plans/01-email-reader.md): a small IMAP client, a MIME-to-plaintext
// converter, and the junk filter that decides a message is not worth reading at
// all. Extraction lives in extract.go and writes nothing itself.
//
// Two constraints shape everything here, both from CLAUDE.md:
//
// - Mail is personal. Nothing in this package logs a body, a subject, or an
// address; callers get the text and decide. Mail text is never search input
// — no function here reaches the network except the IMAP connection itself.
// - Off unless configured. There is no default host, no default account, and
// no fallback that would make a mailbox get read because a field was empty.
//
// The IMAP subset is deliberately tiny (LOGIN, SELECT, UID SEARCH, UID FETCH
// with BODY.PEEK, LOGOUT). No IDLE: a poll every few minutes is what a task
// candidate needs, and IDLE would mean holding a connection and a credential
// open forever for latency nobody is waiting on.
package email
import (
"encoding/base64"
"fmt"
"io"
"mime"
"mime/multipart"
"mime/quotedprintable"
"net/mail"
"regexp"
"strings"
)
// MaxBodyBytes — how much of one message body is kept. A task hides in the
// first screenful; the rest is signature, quoted history and legal boilerplate,
// and it would only spend the resident model's 4096-token context.
const MaxBodyBytes = 4000
// Message — one mail, reduced to the fields extraction and review need.
//
// Raw is deliberately absent: once a message is parsed the original bytes are
// dropped, so no caller can accidentally log or forward the whole mail.
type Message struct {
UID uint32
From string
Subject string
Date string // as sent, unparsed — display only
Body string // plaintext, decoded, HTML-stripped, truncated
// Junk is set by the junk filter (see junk.go). A junk message is carried
// rather than dropped so the poller can count it and still mark it seen.
Junk bool
JunkReason string
}
// ParseMessage turns one RFC 5322 message into a Message.
//
// It never fails on a body it cannot understand: an unparsable or
// unsupported-charset body yields an empty Body and the headers still come
// through, because a subject line alone is often the whole task ("Счёт за
// интернет"). Only a message whose headers cannot be read at all is an error.
func ParseMessage(uid uint32, raw []byte) (Message, error) {
m, err := mail.ReadMessage(strings.NewReader(string(raw)))
if err != nil {
return Message{}, fmt.Errorf("email: parse message: %w", err)
}
msg := Message{
UID: uid,
From: decodeHeader(m.Header.Get("From")),
Subject: decodeHeader(m.Header.Get("Subject")),
Date: m.Header.Get("Date"),
}
msg.Junk, msg.JunkReason = classifyJunk(m.Header)
body, err := plaintextBody(m.Header.Get("Content-Type"), m.Header.Get("Content-Transfer-Encoding"), m.Body)
if err == nil {
msg.Body = truncate(collapse(body), MaxBodyBytes)
}
return msg, nil
}
// plaintextBody walks the MIME tree and returns the best plaintext it can.
//
// Preference order inside a multipart: text/plain first, text/html stripped
// only when there is no plain part. multipart/mixed attachments are skipped
// wholesale — an attachment is a file, not a sentence, and reading one would
// mean parsing arbitrary formats from the network.
func plaintextBody(contentType, encoding string, body io.Reader) (string, error) {
mediaType, params, err := mime.ParseMediaType(contentType)
if contentType == "" || err != nil {
// No Content-Type at all is legal and means text/plain; a broken one is
// treated the same rather than dropping the message.
mediaType, params = "text/plain", nil
}
switch {
case strings.HasPrefix(mediaType, "multipart/"):
boundary := params["boundary"]
if boundary == "" {
return "", fmt.Errorf("email: multipart without boundary")
}
return multipartText(multipart.NewReader(body, boundary))
case mediaType == "text/html":
raw, err := decodeBody(body, encoding, params["charset"])
if err != nil {
return "", err
}
return stripHTML(raw), nil
case mediaType == "text/plain":
return decodeBody(body, encoding, params["charset"])
default:
// A single-part non-text message (a bare PDF, say). No body, headers only.
return "", nil
}
}
// multipartText reads one multipart level, recursing into nested multiparts.
// Returns the plain part if any part yielded one, else the stripped HTML.
func multipartText(mr *multipart.Reader) (string, error) {
var plain, html string
for {
part, err := mr.NextPart()
if err == io.EOF {
break
}
if err != nil {
// A truncated multipart still gives up whatever came before it.
break
}
if part.FileName() != "" {
part.Close()
continue // attachment
}
ct := part.Header.Get("Content-Type")
mediaType, _, _ := mime.ParseMediaType(ct)
text, err := plaintextBody(ct, part.Header.Get("Content-Transfer-Encoding"), part)
part.Close()
if err != nil || strings.TrimSpace(text) == "" {
continue
}
if mediaType == "text/html" && !strings.HasPrefix(mediaType, "multipart/") {
if html == "" {
html = text
}
continue
}
if plain == "" {
plain = text
}
}
if strings.TrimSpace(plain) != "" {
return plain, nil
}
return html, nil
}
// decodeBody applies the transfer encoding, then the charset.
//
// Charset support is UTF-8 (and ASCII, its subset) only, on purpose: x/text's
// encoding tables are not vendored here, and guessing at windows-1251 bytes
// would feed the model mojibake it would happily extract a task from. An
// unsupported charset returns an error, which ParseMessage turns into an empty
// body — subject-only, which is honest.
func decodeBody(r io.Reader, encoding, charset string) (string, error) {
switch strings.ToLower(strings.TrimSpace(encoding)) {
case "quoted-printable":
r = quotedprintable.NewReader(r)
case "base64":
r = newBase64Reader(r)
}
b, err := io.ReadAll(io.LimitReader(r, 1<<20))
if err != nil && len(b) == 0 {
return "", fmt.Errorf("email: read body: %w", err)
}
switch cs := strings.ToLower(strings.TrimSpace(charset)); cs {
case "", "utf-8", "utf8", "us-ascii", "ascii":
return string(b), nil
default:
return "", fmt.Errorf("email: unsupported charset %q", cs)
}
}
// decodeHeader decodes RFC 2047 encoded words ("=?utf-8?B?...?="), which is how
// every Russian subject line arrives. Undecodable headers come back as-is
// rather than empty: a mangled subject is still a hint, and it is only ever
// shown to him as evidence.
func decodeHeader(v string) string {
dec := new(mime.WordDecoder)
out, err := dec.DecodeHeader(v)
if err != nil {
return collapse(v)
}
return collapse(out)
}
var (
scriptStyle = regexp.MustCompile(`(?is)<(script|style)\b[^>]*>.*?</\s*(script|style)\s*>`)
htmlBreak = regexp.MustCompile(`(?i)<\s*(br\s*/?|/p|/div|/tr|/li|/h[1-6])\s*>`)
htmlTag = regexp.MustCompile(`(?s)<[^>]*>`)
htmlComment = regexp.MustCompile(`(?s)<!--.*?-->`)
)
// stripHTML reduces an HTML part to text. A regex stripper, not a parser:
// x/net/html is not vendored, and the consumer is a model reading prose — a
// stray angle bracket costs nothing, whereas a new dependency for the privacy-
// sensitive path costs review.
func stripHTML(s string) string {
s = scriptStyle.ReplaceAllString(s, " ")
s = htmlComment.ReplaceAllString(s, " ")
s = htmlBreak.ReplaceAllString(s, "\n")
s = htmlTag.ReplaceAllString(s, " ")
return unescapeEntities(s)
}
var entities = strings.NewReplacer(
"&nbsp;", " ", "&amp;", "&", "&lt;", "<", "&gt;", ">",
"&quot;", `"`, "&#39;", "'", "&apos;", "'", "&mdash;", "—", "&ndash;", "",
)
func unescapeEntities(s string) string { return entities.Replace(s) }
// collapse squeezes runs of whitespace, keeping single newlines. Mail bodies
// arrive with hard-wrapped lines and blocks of blank space; the model does not
// need them and they are pure context budget.
func collapse(s string) string {
lines := strings.Split(strings.ReplaceAll(s, "\r\n", "\n"), "\n")
var out []string
blank := 0
for _, l := range lines {
l = strings.TrimSpace(strings.Join(strings.Fields(l), " "))
if l == "" {
blank++
if blank > 1 {
continue
}
out = append(out, "")
continue
}
blank = 0
out = append(out, l)
}
return strings.TrimSpace(strings.Join(out, "\n"))
}
// truncate cuts to n bytes on a rune boundary.
func truncate(s string, n int) string {
if len(s) <= n {
return s
}
cut := s[:n]
for len(cut) > 0 && !isRuneStart(cut[len(cut)-1]) {
cut = cut[:len(cut)-1]
}
return strings.TrimSpace(cut) + "…"
}
func isRuneStart(b byte) bool { return b&0xC0 != 0x80 }
// newBase64Reader — base64.NewDecoder already skips the CRLFs mail bodies wrap
// with, so this is only a named seam for decodeBody to read cleanly.
func newBase64Reader(r io.Reader) io.Reader {
return base64.NewDecoder(base64.StdEncoding, r)
}
+111
View File
@@ -0,0 +1,111 @@
package email
import (
"os"
"path/filepath"
"strings"
"testing"
)
func fixture(t *testing.T, name string) []byte {
t.Helper()
b, err := os.ReadFile(filepath.Join("testdata", name))
if err != nil {
t.Fatalf("read fixture %s: %v", name, err)
}
return b
}
func TestParsePlainRussian(t *testing.T) {
msg, err := ParseMessage(7, fixture(t, "plain_ru.eml"))
if err != nil {
t.Fatalf("parse: %v", err)
}
if msg.UID != 7 {
t.Errorf("uid = %d, want 7", msg.UID)
}
if want := "Нужно закрыть задачу"; msg.Subject != want {
t.Errorf("subject = %q, want %q", msg.Subject, want)
}
if !strings.Contains(msg.From, "Антон") {
t.Errorf("from = %q, want the decoded display name", msg.From)
}
if !strings.Contains(msg.Body, "Надо отправить акт до пятницы.") {
t.Errorf("body = %q, want the quoted-printable text decoded", msg.Body)
}
if msg.Junk {
t.Errorf("a personal mail must not be junk (%s)", msg.JunkReason)
}
}
func TestParseHTMLOnlyIsStripped(t *testing.T) {
msg, err := ParseMessage(1, fixture(t, "html_only.eml"))
if err != nil {
t.Fatalf("parse: %v", err)
}
if strings.Contains(msg.Body, "<") || strings.Contains(msg.Body, "color:red") || strings.Contains(msg.Body, "x()") {
t.Errorf("body still has markup/script/style: %q", msg.Body)
}
for _, want := range []string{"Счёт за интернет: 700", "Оплатить до 5 августа."} {
if !strings.Contains(msg.Body, want) {
t.Errorf("body = %q, want it to contain %q", msg.Body, want)
}
}
// &nbsp; must have become a real space, not vanished into the number.
if strings.Contains(msg.Body, "&nbsp;") {
t.Errorf("entity left unescaped: %q", msg.Body)
}
}
func TestParsePrefersPlainAndSkipsAttachments(t *testing.T) {
msg, err := ParseMessage(2, fixture(t, "mixed_attachment.eml"))
if err != nil {
t.Fatalf("parse: %v", err)
}
if got := strings.TrimSpace(msg.Body); got != "Sign the contract before Monday." {
t.Errorf("body = %q, want the text/plain alternative only", got)
}
if strings.Contains(msg.Body, "PDF") {
t.Errorf("attachment bytes leaked into the body: %q", msg.Body)
}
}
// An unsupported charset must degrade to headers-only rather than to mojibake
// the model would then extract a task from.
func TestParseUnsupportedCharsetKeepsHeaders(t *testing.T) {
msg, err := ParseMessage(3, fixture(t, "cp1251.eml"))
if err != nil {
t.Fatalf("parse: %v", err)
}
if msg.Subject != "Legacy" {
t.Errorf("subject = %q, want Legacy", msg.Subject)
}
if msg.Body != "" {
t.Errorf("body = %q, want empty for an undecodable charset", msg.Body)
}
}
func TestParseTruncatesLongBody(t *testing.T) {
var b strings.Builder
b.WriteString("Subject: long\r\nContent-Type: text/plain; charset=utf-8\r\n\r\n")
for i := 0; i < 2000; i++ {
b.WriteString("длинная строка ")
}
msg, err := ParseMessage(4, []byte(b.String()))
if err != nil {
t.Fatalf("parse: %v", err)
}
if len(msg.Body) > MaxBodyBytes+8 {
t.Errorf("body kept %d bytes, want ≤ %d", len(msg.Body), MaxBodyBytes)
}
if !strings.HasSuffix(msg.Body, "…") {
t.Errorf("truncated body should be marked: %q", msg.Body[len(msg.Body)-20:])
}
}
func TestCollapseSqueezesBlankLines(t *testing.T) {
got := collapse(" a b \r\n\r\n\r\n\r\n c \r\n")
if got != "a b\n\nc" {
t.Errorf("collapse = %q, want %q", got, "a b\n\nc")
}
}
+7
View File
@@ -0,0 +1,7 @@
From: legacy@example.org
To: kami@example.org
Subject: Legacy
Date: Fri, 01 Aug 2026 05:00:00 +0400
Content-Type: text/plain; charset="windows-1251"
Ï
+13
View File
@@ -0,0 +1,13 @@
From: billing@isp.example
To: kami@example.org
Subject: =?utf-8?B?0KHRh9GR0YIg0LfQsCDQuNC90YLQtdGA0L3QtdGC?=
Date: Fri, 01 Aug 2026 08:00:00 +0400
MIME-Version: 1.0
Content-Type: multipart/alternative; boundary="B1"
--B1
Content-Type: text/html; charset="utf-8"
Content-Transfer-Encoding: base64
PGh0bWw+PGhlYWQ+PHN0eWxlPnB7Y29sb3I6cmVkfTwvc3R5bGU+PC9oZWFkPjxib2R5PjxwPtCh0YfRkdGCINC30LAg0LjQvdGC0LXRgNC90LXRgjogNzAwJm5ic3A74oK9PC9wPjxwPtCe0L/Qu9Cw0YLQuNGC0Ywg0LTQviA1INCw0LLQs9GD0YHRgtCwLjwvcD48c2NyaXB0PngoKTwvc2NyaXB0PjwvYm9keT48L2h0bWw+
--B1--
+26
View File
@@ -0,0 +1,26 @@
From: hr@work.example
To: kami@example.org
Subject: Contract
Date: Fri, 01 Aug 2026 07:00:00 +0400
MIME-Version: 1.0
Content-Type: multipart/mixed; boundary="M1"
--M1
Content-Type: multipart/alternative; boundary="A1"
--A1
Content-Type: text/plain; charset="utf-8"
Sign the contract before Monday.
--A1
Content-Type: text/html; charset="utf-8"
<p>Sign the contract before Monday.</p>
--A1--
--M1
Content-Type: application/pdf; name="contract.pdf"
Content-Disposition: attachment; filename="contract.pdf"
Content-Transfer-Encoding: base64
JVBERi0xLjQgbm90IHJlYWxseSBhIHBkZg==
--M1--
+9
View File
@@ -0,0 +1,9 @@
From: news@shop.example
To: kami@example.org
Subject: =?utf-8?B?0KHQutC40LTQutC4INGC0L7Qu9GM0LrQviDRgdC10LPQvtC00L3Rjw==?=
Date: Fri, 01 Aug 2026 06:00:00 +0400
List-Unsubscribe: <mailto:unsub@shop.example>
Precedence: bulk
Content-Type: text/plain; charset="utf-8"
Sale!
+14
View File
@@ -0,0 +1,14 @@
From: =?utf-8?B?0JDQvdGC0L7QvQ==?= <anton@example.org>
To: kami@example.org
Subject: =?utf-8?B?0J3Rg9C20L3QviDQt9Cw0LrRgNGL0YLRjCDQt9Cw0LTQsNGH0YM=?=
Date: Fri, 01 Aug 2026 09:12:00 +0400
Content-Type: text/plain; charset="utf-8"
Content-Transfer-Encoding: quoted-printable
Message-ID: <plain-ru@example.org>
=D0=9F=D1=80=D0=B8=D0=B2=D0=B5=D1=82! =D0=9D=D0=B0=D0=B4=D0=BE =D0=BE=D1=82=
=D0=BF=D1=80=D0=B0=D0=B2=D0=B8=D1=82=D1=8C =D0=B0=D0=BA=D1=82 =D0=B4=D0=BE =
=D0=BF=D1=8F=D1=82=D0=BD=D0=B8=D1=86=D1=8B.
--
Anton
+38
View File
@@ -138,6 +138,44 @@ type CaptureTaskResp struct {
Created bool `json:"created"`
}
// IngestMailReq — one message a mail reader has fetched, handed to core for
// extraction (Vikunja #246).
//
// The mail reader (cmd/mavmaild) holds the IMAP credential and core never sees
// it, the same split mavpoll uses for the zenmoney token. What crosses this
// boundary is only the message text, because extraction runs on the resident
// model and llama-server lives inside core's process.
//
// Body is already plaintext and truncated by internal/email; core does not
// re-parse MIME and never stores the body. Junk means the reader's header
// filter already classified the message as bulk — core is told rather than
// asked, so a junk message can be counted without a model call.
//
// This method is available only when core has an email block configured AND a
// llama-server phraser; otherwise it answers ErrUnknownMethod, which is what
// "off unless configured" looks like at the wire.
type IngestMailReq struct {
Mailbox string `json:"mailbox"`
UID uint32 `json:"uid"`
From string `json:"from,omitempty"`
Subject string `json:"subject,omitempty"`
Date string `json:"date,omitempty"`
Body string `json:"body,omitempty"`
Junk bool `json:"junk,omitempty"`
}
// IngestMailResp — what core did with the message. TaskIDs are the rows
// CaptureTask returned; Created counts the ones that were new (a re-read
// mailbox dedupes to Created=0). Skipped is set when nothing was asked of the
// model at all — junk, or an empty message.
//
// Nothing here echoes the mail back. The reader logs counts.
type IngestMailResp struct {
TaskIDs []int64 `json:"task_ids,omitempty"`
Created int `json:"created"`
Skipped bool `json:"skipped,omitempty"`
}
type listTasksReq struct {
Status string `json:"status"` // "" all | "live" | candidate|open|done|dropped
}
+11
View File
@@ -448,6 +448,17 @@ func (c *Client) SetTaskStatus(ctx context.Context, id int64, status string, ts
return c.call(ctx, MethodSetTaskStatus, setTaskStatusReq{ID: id, Status: status, Ts: ts}, nil)
}
// IngestMail hands one fetched message to core for extraction. ErrUnknownMethod
// means core has no email block configured — the caller should stop asking, not
// retry.
func (c *Client) IngestMail(ctx context.Context, req IngestMailReq) (IngestMailResp, error) {
var r IngestMailResp
if err := c.call(ctx, MethodIngestMail, req, &r); err != nil {
return IngestMailResp{}, err
}
return r, nil
}
func (c *Client) DismissProposedRoutine(ctx context.Context, id int64) error {
return c.call(ctx, MethodDismissProposedRoutine, dismissProposedRoutineReq{ID: id}, nil)
}
+33
View File
@@ -565,3 +565,36 @@ func mustJSON(v any) []byte {
}
return b
}
// TestIngestMail_OffUnlessConfigured — with no IngestMailFn set (the default,
// and what an unconfigured core looks like) the method does not exist. A mail
// reader gets a refusal it can act on rather than a silent success.
func TestIngestMail_OffUnlessConfigured(t *testing.T) {
_, _, cli, _ := newServerWithStore(t)
if _, err := cli.IngestMail(context.Background(), IngestMailReq{Mailbox: "INBOX", UID: 1}); !errors.Is(err, ErrUnknownMethod) {
t.Fatalf("IngestMail error = %v, want ErrUnknownMethod", err)
}
}
// TestIngestMail_Hook — when the daemon wires the hook, the message crosses the
// boundary intact and the response comes back.
func TestIngestMail_Hook(t *testing.T) {
_, srv, cli, _ := newServerWithStore(t)
var got IngestMailReq
srv.IngestMailFn = func(_ context.Context, req IngestMailReq) (IngestMailResp, error) {
got = req
return IngestMailResp{TaskIDs: []int64{7}, Created: 1}, nil
}
resp, err := cli.IngestMail(context.Background(), IngestMailReq{
Mailbox: "INBOX", UID: 12, Subject: "Счёт", Body: "Оплатить.", Junk: false,
})
if err != nil {
t.Fatalf("IngestMail: %v", err)
}
if resp.Created != 1 || len(resp.TaskIDs) != 1 || resp.TaskIDs[0] != 7 {
t.Errorf("resp = %+v", resp)
}
if got.UID != 12 || got.Subject != "Счёт" || got.Body != "Оплатить." {
t.Errorf("req across the wire = %+v", got)
}
}
+33 -4
View File
@@ -421,6 +421,17 @@ type Server struct {
// Set by the daemon; nil ⇒ MethodStoreEncryptionKey returns ErrUnknownMethod.
WrapKeyFn WrapKeyFunc
// IngestMailFn — extracts task candidates from one fetched message. Set by
// the daemon only when an email block is configured AND there is a
// llama-server to extract with; nil ⇒ MethodIngestMail returns
// ErrUnknownMethod, so a mail reader pointed at a core that is not
// configured for mail is refused rather than silently ignored.
//
// Like StepUp/WrapKeyFn/UnlockFn this bypasses CoreAPI: it is not a store
// operation, it needs the resident model, and it must not become a method
// every CoreAPI implementation has to carry.
IngestMailFn IngestMailFunc
// UnlockFn — unwraps the store encryption key from the wrapped blob using
// the passkey credential public key, opens the encrypted store, and wires
// the rest of the daemon (voice, loop, delivery). Set by the daemon when
@@ -439,6 +450,9 @@ type WrapKeyFunc func(ctx context.Context, publicKey []byte) error
// public key and completes daemon initialization.
type UnlockFunc func(ctx context.Context, publicKey []byte) error
// IngestMailFunc — core-side mail extraction. Returns what was captured.
type IngestMailFunc func(ctx context.Context, req IngestMailReq) (IngestMailResp, error)
// CheckFunc — the auth hook signature. Wired by the daemon (auth.Gate.Check
// satisfies this); dispatch calls it once per request after param-unmarshal
// independence (it gets the raw params, may unmarshal what it needs — ipc
@@ -606,9 +620,10 @@ func withoutParams[R any](fn func(ctx context.Context, api CoreAPI) (R, error))
// existed) as an argument — so SetAPI's runtime swap (the unlock transition)
// is still honored on the very next request with no extra plumbing here.
//
// MethodAssertStepUp, MethodStoreEncryptionKey and MethodUnlock are NOT in
// this table: they bypass CoreAPI entirely (s.StepUp / s.WrapKeyFn /
// s.UnlockFn), so dispatch special-cases them before consulting the table.
// MethodAssertStepUp, MethodStoreEncryptionKey, MethodUnlock and
// MethodIngestMail are NOT in this table: they bypass CoreAPI entirely
// (s.StepUp / s.WrapKeyFn / s.UnlockFn / s.IngestMailFn), so dispatch
// special-cases them before consulting the table.
var methodTable = map[Method]handlerFunc{
MethodWriteFact: withParams(func(ctx context.Context, api CoreAPI, p WriteFactReq) (idResp, error) {
id, err := api.WriteFact(ctx, p)
@@ -816,7 +831,7 @@ func (s *Server) dispatch(ctx context.Context, req Request) (json.RawMessage, er
}
}
// These three bypass CoreAPI entirely — they drive Server fields set
// These bypass CoreAPI entirely — they drive Server fields set
// directly by the daemon (StepUp / WrapKeyFn / UnlockFn), not store
// state, so they can never be table entries keyed on a CoreAPI method.
switch req.Method {
@@ -845,6 +860,20 @@ func (s *Server) dispatch(ctx context.Context, req Request) (json.RawMessage, er
return marshalResult(nil), s.UnlockFn(ctx, p.PublicKey)
}
return nil, fmt.Errorf("%w: %s", ErrUnknownMethod, req.Method)
case MethodIngestMail:
if s.IngestMailFn != nil {
var p IngestMailReq
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
resp, err := s.IngestMailFn(ctx, p)
if err != nil {
return nil, err
}
return marshalResult(resp), nil
}
return nil, fmt.Errorf("%w: %s", ErrUnknownMethod, req.Method)
}
h, ok := methodTable[req.Method]
+1
View File
@@ -50,6 +50,7 @@ const (
MethodCaptureTask Method = "capture_task"
MethodListTasks Method = "list_tasks"
MethodSetTaskStatus Method = "set_task_status"
MethodIngestMail Method = "ingest_mail"
)
// Request — one frame from module to core. Params is the JSON-encoded argument
+109
View File
@@ -0,0 +1,109 @@
package router
import "strings"
// Feed queries — "что нового в лентах?", "что нового по технологиям?"
// (Vikunja #258).
//
// Deterministic matching, like the calendar, plan and habit matchers above it:
// the LLM router says this is a query, and this decides whether it is a question
// about the feeds. A model deciding that would occasionally answer "что нового?"
// out of world knowledge, which is the one thing a feed reader exists to avoid.
// FeedQuery — a parsed "what's new" question. Category is the topic he named
// ("технологии"), empty when he asked about the feeds in general.
type FeedQuery struct {
Category string
}
// feedNouns — the words that make a question about the feeds themselves.
var feedNouns = []string{
"лента", "ленте", "ленты", "лентах", "лентам",
"новости", "новостей", "новостях", "новостям",
"новое", "нового", "новенького",
"feed", "feeds", "news", "headlines",
}
// newnessMarkers — the "что нового" half. "нового" alone is in feedNouns
// because it carries the question on its own ("что нового?"); a bare "лента"
// needs the ask, which is what askMarkers below is for.
var askMarkers = []string{
"что", "какие", "какое", "расскажи", "почитай", "прочитай", "покажи",
"what", "any", "tell", "show", "read",
}
// ParseFeedQuery reports whether an utterance asks what is new in the feeds, and
// which topic if it names one after "по"/"о"/"про"/"about".
//
// Both a feed noun and an ask are required. "у меня новая лента в инстаграме" is
// a statement and must not be read as a request to recite headlines.
func ParseFeedQuery(text string) (FeedQuery, bool) {
toks := planTokens(text)
noun, ask := false, false
for _, t := range toks {
for _, n := range feedNouns {
if t == n {
noun = true
break
}
}
for _, a := range askMarkers {
if t == a {
ask = true
break
}
}
}
if !noun || !ask {
return FeedQuery{}, false
}
return FeedQuery{Category: feedCategory(toks)}, true
}
// categoryPreps — the prepositions a topic follows. Russian marks the topic with
// a preposition ("по технологиям", "про политику"), so the word after one is the
// category; there is no stemming here, and the match against the configured
// category is a prefix comparison for exactly that reason.
var categoryPreps = map[string]bool{"по": true, "о": true, "об": true, "про": true, "about": true, "on": true}
func feedCategory(toks []string) string {
for i, t := range toks {
if categoryPreps[t] && i+1 < len(toks) {
next := toks[i+1]
// "по новостям" names no topic, it repeats the noun.
for _, n := range feedNouns {
if next == n {
return ""
}
}
return next
}
}
return ""
}
// CategoryMatches reports whether a note's text plausibly belongs to the
// category he named. Russian inflects the topic ("технологиям" vs the configured
// "технологии"), and there is no stemmer in this repo, so the comparison is on a
// common prefix — long enough that "полит" and "погод" stay apart, short enough
// to survive a case ending.
func CategoryMatches(text, category string) bool {
if category == "" {
return true
}
stem := categoryStem(category)
if stem == "" {
return false
}
return strings.Contains(strings.ToLower(text), stem)
}
// categoryStem cuts a word down to the part inflection leaves alone. 5 runes is
// the compromise: shorter words are used whole.
func categoryStem(word string) string {
r := []rune(strings.ToLower(strings.TrimSpace(word)))
if len(r) > 5 {
r = r[:5]
}
return string(r)
}
+48
View File
@@ -0,0 +1,48 @@
package router
import "testing"
func TestParseFeedQuery(t *testing.T) {
cases := []struct {
text string
ok bool
category string
}{
{"что нового в лентах?", true, ""},
{"что нового?", true, ""},
{"какие новости?", true, ""},
{"что нового по технологиям?", true, "технологиям"},
{"расскажи новости про политику", true, "политику"},
{"что нового по новостям", true, ""},
{"what's new in the feeds?", true, ""},
{"any news about kubernetes", true, "kubernetes"},
// Statements, not requests.
{"у меня новая лента в инстаграме", false, ""},
{"новости меня утомили", false, ""},
{"напомни полить цветы", false, ""},
{"", false, ""},
}
for _, c := range cases {
q, ok := ParseFeedQuery(c.text)
if ok != c.ok {
t.Errorf("ParseFeedQuery(%q) ok = %v, want %v", c.text, ok, c.ok)
continue
}
if ok && q.Category != c.category {
t.Errorf("ParseFeedQuery(%q) category = %q, want %q", c.text, q.Category, c.category)
}
}
}
func TestCategoryMatches(t *testing.T) {
// The inflected form he says must match the form the config spells.
if !CategoryMatches("Новый релиз [технологии]", "технологиям") {
t.Error("inflected category did not match")
}
if CategoryMatches("Новый релиз [технологии]", "политику") {
t.Error("unrelated category matched")
}
if !CategoryMatches("anything", "") {
t.Error("an empty category must match everything")
}
}
+71
View File
@@ -0,0 +1,71 @@
package router
import "strings"
// Money questions, matched deterministically (Vikunja #125).
//
// No new intent, for the same reason as tasks: the intent enum is a contract
// with the relabelling prompt. "сколько я потратил?" is a query; which figure
// it asks for is a lookup, not something to ask a 1.7B — and a model asked to
// invent a spending total will happily do it.
// MoneyWindow — which period a money question asks about.
type MoneyWindow int
const (
MoneyNone MoneyWindow = iota
MoneyToday
MoneyMonth
)
// moneyNouns — the words that make a question be about his money.
var moneyNouns = []string{
"потратил", "потратила", "тратил", "траты", "трат", "расходы", "расходов",
"заработал", "потрачено", "денег", "spend", "spent", "expenses",
}
// ParseMoneyQuery reports whether an utterance asks about spending or income,
// and over which window. Defaults to the month: "сколько я потратил?" without a
// period is the month-to-date question, which is the one worth answering.
//
// Narrow on purpose. A money noun alone is not enough — "я потратил весь день
// на это" is him talking about his day, so an amount word or an explicit
// question word has to be there too.
func ParseMoneyQuery(text string) (MoneyWindow, bool) {
toks := planTokens(text)
if len(toks) == 0 {
return MoneyNone, false
}
hasNoun := false
for _, t := range toks {
for _, n := range moneyNouns {
if t == n {
hasNoun = true
}
}
}
if !hasNoun {
return MoneyNone, false
}
// "весь день", "время", "силы" — spending that is not money.
for _, t := range toks {
switch t {
case "день", "дня", "время", "времени", "силы", "сил", "нервы":
return MoneyNone, false
}
}
asking := hasTok(toks, "сколько") || hasTok(toks, "какие") || hasTok(toks, "покажи") ||
hasTok(toks, "how") || hasTok(toks, "much") || hasTok(toks, "my") ||
hasTok(toks, "мои") || hasTok(toks, "траты") || hasTok(toks, "расходы")
if !asking {
return MoneyNone, false
}
lower := strings.ToLower(text)
switch {
case hasTok(toks, "сегодня") || strings.Contains(lower, "today"):
return MoneyToday, true
case hasTok(toks, "месяц") || hasTok(toks, "месяце") || strings.Contains(lower, "month"):
return MoneyMonth, true
}
return MoneyMonth, true
}
+31
View File
@@ -0,0 +1,31 @@
package router
import "testing"
func TestParseMoneyQuery(t *testing.T) {
cases := []struct {
in string
window MoneyWindow
ok bool
}{
{"сколько я потратил сегодня?", MoneyToday, true},
{"сколько я потратил в этом месяце?", MoneyMonth, true},
{"сколько я потратил?", MoneyMonth, true}, // month-to-date by default
{"покажи мои траты", MoneyMonth, true},
{"какие у меня расходы за месяц", MoneyMonth, true},
{"how much did I spend today", MoneyToday, true},
{"сколько я заработал в этом месяце", MoneyMonth, true},
// Not about money.
{"я потратил весь день на это", MoneyNone, false},
{"потратил много сил", MoneyNone, false},
{"какая погода?", MoneyNone, false},
{"я купил молоко", MoneyNone, false},
{"", MoneyNone, false},
}
for _, c := range cases {
w, ok := ParseMoneyQuery(c.in)
if ok != c.ok || w != c.window {
t.Errorf("ParseMoneyQuery(%q) = (%v, %v), want (%v, %v)", c.in, w, ok, c.window, c.ok)
}
}
}
+51 -4
View File
@@ -38,11 +38,33 @@ var taskCapturePrefixes = []string{
"new task",
}
// TaskCapture — a parsed capture: the task itself, plus the importance he
// stated out loud if he stated one (Vikunja #129). Weight 0 means he said
// nothing about importance, which the ranker treats as exactly that — no
// urgency is inferred from the wording.
type TaskCapture struct {
Text string
Weight int
}
// urgencyMarkers — the words that set a weight, strongest first. Only these
// two rungs: "срочно" is a deadline he has not named, "важно" is a preference,
// and a third shade of urgent would be a distinction he never makes out loud.
var urgencyMarkers = []struct {
word string
weight int
}{
{"срочно", 3},
{"urgent", 3},
{"важно", 2},
{"important", 2},
}
// ParseTaskCapture reports whether an utterance explicitly files a task, and
// returns the task text with the marker stripped. A marker with nothing after it
// is not a capture (there is no task in "добавь в задачи") — the caller falls
// through to whatever it would otherwise have done with the turn.
func ParseTaskCapture(text string) (string, bool) {
func ParseTaskCapture(text string) (TaskCapture, bool) {
trimmed := strings.TrimSpace(text)
lower := strings.ToLower(trimmed)
best := ""
@@ -52,7 +74,7 @@ func ParseTaskCapture(text string) (string, bool) {
}
}
if best == "" {
return "", false
return TaskCapture{}, false
}
// Cut on the rune length of the matched prefix. ToLower does not change the
// byte length of Russian or English letters, so the index carries over.
@@ -60,10 +82,35 @@ func ParseTaskCapture(text string) (string, bool) {
rest = strings.TrimLeft(rest, ":—- ")
rest = strings.TrimSpace(rest)
rest = strings.TrimRight(rest, ".!")
rest, weight := stripUrgency(rest)
if rest == "" {
return "", false
return TaskCapture{}, false
}
return rest, true
return TaskCapture{Text: rest, Weight: weight}, true
}
// stripUrgency pulls a leading or trailing urgency word out of the task text
// and returns the weight it implies. Only at the edges: "срочно оплатить
// интернет" and "оплатить интернет срочно" are the same instruction, while
// "позвонить в срочную помощь" is a task whose text happens to contain the
// stem, and cutting a word out of the middle of it would mangle the task.
//
// The word is removed from the text, because the list should read "оплатить
// интернет (важно)" and not "важно оплатить интернет (важно)".
func stripUrgency(text string) (string, int) {
for _, m := range urgencyMarkers {
lower := strings.ToLower(text)
switch {
case strings.HasPrefix(lower, m.word+" "):
return strings.TrimSpace(text[len(m.word):]), m.weight
case strings.HasSuffix(lower, " "+m.word):
return strings.TrimSpace(text[:len(text)-len(m.word)]), m.weight
case lower == m.word:
// Nothing but the marker — no task in it.
return "", 0
}
}
return text, 0
}
// taskListWords — the nouns that make a question be about the task list.
+25 -17
View File
@@ -4,28 +4,36 @@ import "testing"
func TestParseTaskCapture(t *testing.T) {
cases := []struct {
in string
text string
ok bool
in string
text string
weight int
ok bool
}{
{"добавь в задачи купить молоко", "купить молоко", true},
{"Добавь в список дел: позвонить в банк", "позвонить в банк", true},
{"запиши задачу починить кран.", "починить кран", true},
{"новая задача — оплатить интернет", "оплатить интернет", true},
{"add a task buy milk", "buy milk", true},
{"добавь в задачи купить молоко", "купить молоко", 0, true},
{"Добавь в список дел: позвонить в банк", "позвонить в банк", 0, true},
{"запиши задачу починить кран.", "починить кран", 0, true},
{"новая задача — оплатить интернет", "оплатить интернет", 0, true},
{"add a task buy milk", "buy milk", 0, true},
// Urgency he stated out loud, leading or trailing, stripped from the text.
{"добавь в задачи срочно оплатить интернет", "оплатить интернет", 3, true},
{"добавь в задачи оплатить интернет срочно", "оплатить интернет", 3, true},
{"новая задача важно позвонить маме", "позвонить маме", 2, true},
// The stem inside the task text is part of the task, not a marker.
{"добавь в задачи позвонить в срочную помощь", "позвонить в срочную помощь", 0, true},
// A marker with nothing after it files nothing.
{"добавь в задачи", "", false},
{"новая задача", "", false},
{"добавь в задачи", "", 0, false},
{"новая задача", "", 0, false},
{"добавь в задачи срочно", "", 0, false},
// Not a capture: he is talking, not filing.
{"надо бы поспать", "", false},
{"я не добавил молоко в список", "", false},
{"какие у меня задачи?", "", false},
{"", "", false},
{"надо бы поспать", "", 0, false},
{"я не добавил молоко в список", "", 0, false},
{"какие у меня задачи?", "", 0, false},
{"", "", 0, false},
}
for _, c := range cases {
text, ok := ParseTaskCapture(c.in)
if ok != c.ok || text != c.text {
t.Errorf("ParseTaskCapture(%q) = (%q, %v), want (%q, %v)", c.in, text, ok, c.text, c.ok)
got, ok := ParseTaskCapture(c.in)
if ok != c.ok || got.Text != c.text || got.Weight != c.weight {
t.Errorf("ParseTaskCapture(%q) = (%+v, %v), want (%q, w=%d, %v)", c.in, got, ok, c.text, c.weight, c.ok)
}
}
}
+39
View File
@@ -0,0 +1,39 @@
package router
import (
"regexp"
"strings"
)
// Finding a URL in an utterance (Vikunja #259).
//
// This is deliberately strict: a scheme is required. "посмотри на example.org"
// is not treated as a fetch request, because a bare dotted word is also how
// people say file names, versions and Russian abbreviations, and the cost of a
// false positive here is an outbound request nobody asked for.
//
// Note where this runs: an utterance from STT. Whisper will mangle a spoken URL,
// which is fine — the URL that survives is one he pasted into the web chat, and
// a mangled one simply fails to match.
var urlRE = regexp.MustCompile(`(?i)\bhttps?://[^\s<>"']+`)
// FirstURL returns the first http(s) URL in text.
//
// Trailing punctuation is trimmed: he ends sentences, and "…/page." is not a
// path component. A closing bracket is only trimmed when it has no opener,
// because a wikipedia URL legitimately ends in one.
func FirstURL(text string) (string, bool) {
m := urlRE.FindString(text)
if m == "" {
return "", false
}
m = strings.TrimRight(m, ".,;:!?…")
if strings.HasSuffix(m, ")") && strings.Count(m, "(") == 0 {
m = strings.TrimSuffix(m, ")")
}
// A scheme with nothing after it is not a URL.
if rest := strings.SplitN(m, "//", 2); len(rest) < 2 || rest[1] == "" {
return "", false
}
return m, true
}
+34
View File
@@ -0,0 +1,34 @@
package router
import "testing"
func TestFirstURL(t *testing.T) {
cases := []struct {
text string
want string
}{
{"посмотри https://example.org/page — что там?", "https://example.org/page"},
{"почитай http://example.org/a/b?x=1 и скажи", "http://example.org/a/b?x=1"},
{"вот ссылка: https://example.org/page.", "https://example.org/page"},
{"https://ru.wikipedia.org/wiki/Небо_(значения)", "https://ru.wikipedia.org/wiki/Небо_(значения)"},
// No scheme ⇒ no fetch. A bare dotted word is not an instruction to
// reach out to the network.
{"посмотри на example.org", ""},
{"открой файл config.json", ""},
{"что нового?", ""},
{"https://", ""},
{"", ""},
}
for _, c := range cases {
got, ok := FirstURL(c.text)
if c.want == "" {
if ok {
t.Errorf("FirstURL(%q) = %q, want no match", c.text, got)
}
continue
}
if !ok || got != c.want {
t.Errorf("FirstURL(%q) = %q, %v; want %q", c.text, got, ok, c.want)
}
}
}
+179
View File
@@ -0,0 +1,179 @@
// Package rss reads RSS 2.0 and Atom feeds, and does nothing else with them.
//
// Parsing and polling are split from delivery on purpose: a feed is a source
// Maven can be ASKED about, not a thing that speaks. Nothing in this package
// dispatches, nudges or notifies — the poller writes notes, and the answer path
// reads them when he asks "что нового в лентах?". "Not a nag" is the oldest
// constraint in the spec, and a news feed is the single most tempting way to
// break it.
//
// Stdlib only (encoding/xml). Feeds are XML from strangers, so the parser takes
// what it recognises and ignores the rest rather than failing a whole feed over
// one malformed item.
package rss
import (
"encoding/xml"
"fmt"
"html"
"io"
"regexp"
"strings"
"time"
)
// Item is one feed entry, normalised across RSS and Atom.
type Item struct {
Title string
Link string
Summary string // plain text, tags stripped, entities decoded
Published time.Time // zero when the feed did not say
ID string // guid / atom id, falling back to the link
}
// Feed is a parsed document.
type Feed struct {
Title string
Items []Item
}
// feedDoc covers both dialects in one struct. RSS puts items under
// channel>item, Atom puts entries at the top level, and the field names barely
// overlap — so both sets are declared and whichever the document filled in wins.
type feedDoc struct {
ChannelTitle string `xml:"channel>title"`
AtomTitle string `xml:"title"`
Items []struct {
Title string `xml:"title"`
Link string `xml:"link"`
Description string `xml:"description"`
Encoded string `xml:"encoded"` // content:encoded
GUID string `xml:"guid"`
PubDate string `xml:"pubDate"`
Date string `xml:"date"` // dc:date
} `xml:"channel>item"`
Entries []struct {
Title string `xml:"title"`
Links []struct {
Href string `xml:"href,attr"`
Rel string `xml:"rel,attr"`
} `xml:"link"`
Summary string `xml:"summary"`
Content string `xml:"content"`
ID string `xml:"id"`
Updated string `xml:"updated"`
Published string `xml:"published"`
} `xml:"entry"`
}
// Parse reads a feed document.
func Parse(r io.Reader) (Feed, error) {
var doc feedDoc
dec := xml.NewDecoder(r)
// Feeds in the wild declare windows-1251 and worse. We only ever read
// UTF-8; a charset we cannot decode is a feed we do not read, which is
// better than mojibake in his notes.
dec.Strict = false
if err := dec.Decode(&doc); err != nil {
return Feed{}, fmt.Errorf("rss: bad xml: %w", err)
}
f := Feed{Title: strings.TrimSpace(doc.ChannelTitle)}
if f.Title == "" {
f.Title = strings.TrimSpace(doc.AtomTitle)
}
for _, it := range doc.Items {
item := Item{
Title: PlainText(it.Title),
Link: strings.TrimSpace(it.Link),
Summary: PlainText(firstNonEmpty(it.Description, it.Encoded)),
Published: parseTime(firstNonEmpty(it.PubDate, it.Date)),
ID: strings.TrimSpace(firstNonEmpty(it.GUID, it.Link)),
}
if item.Title != "" || item.Link != "" {
f.Items = append(f.Items, item)
}
}
for _, e := range doc.Entries {
link := ""
for _, l := range e.Links {
if l.Rel == "" || l.Rel == "alternate" {
link = strings.TrimSpace(l.Href)
break
}
}
if link == "" && len(e.Links) > 0 {
link = strings.TrimSpace(e.Links[0].Href)
}
item := Item{
Title: PlainText(e.Title),
Link: link,
Summary: PlainText(firstNonEmpty(e.Summary, e.Content)),
Published: parseTime(firstNonEmpty(e.Published, e.Updated)),
ID: strings.TrimSpace(firstNonEmpty(e.ID, link)),
}
if item.Title != "" || item.Link != "" {
f.Items = append(f.Items, item)
}
}
return f, nil
}
func firstNonEmpty(vals ...string) string {
for _, v := range vals {
if strings.TrimSpace(v) != "" {
return v
}
}
return ""
}
// timeLayouts — RFC1123/822 for RSS, RFC3339 for Atom, plus the near-misses
// real feeds ship (no seconds, numeric zone where a name is expected).
var timeLayouts = []string{
time.RFC1123Z,
time.RFC1123,
time.RFC822Z,
time.RFC822,
time.RFC3339,
"2006-01-02T15:04:05Z0700",
"2006-01-02 15:04:05",
"2006-01-02",
"Mon, 02 Jan 2006 15:04:05 -0700",
"Mon, 2 Jan 2006 15:04:05 -0700",
"Mon, 2 Jan 2006 15:04:05 MST",
}
// parseTime returns the zero time on anything it cannot read. An undated item
// is still an item; the poller dedupes by ID, so a missing date costs nothing.
func parseTime(s string) time.Time {
s = strings.TrimSpace(s)
if s == "" {
return time.Time{}
}
for _, l := range timeLayouts {
if t, err := time.Parse(l, s); err == nil {
return t.UTC()
}
}
return time.Time{}
}
var (
// RE2 has no backreferences, so the two tags are spelled out rather than
// captured and matched against themselves.
scriptRE = regexp.MustCompile(`(?is)<script\b[^>]*>.*?</script>|<style\b[^>]*>.*?</style>`)
tagRE = regexp.MustCompile(`(?s)<[^>]*>`)
)
// PlainText strips markup and decodes entities — feed summaries are HTML, and
// what reaches a note (and possibly the TTS) must be text. Exported because the
// crawler's extractor needs exactly this on a bigger input.
func PlainText(s string) string {
s = scriptRE.ReplaceAllString(s, " ")
s = tagRE.ReplaceAllString(s, " ")
s = html.UnescapeString(s)
return strings.TrimSpace(strings.Join(strings.Fields(s), " "))
}
+112
View File
@@ -0,0 +1,112 @@
package rss
import (
"strings"
"testing"
"time"
)
const rss2 = `<?xml version="1.0"?>
<rss version="2.0">
<channel>
<title>Хабр</title>
<item>
<title>Новая уязвимость в ядре</title>
<link>https://example.org/a</link>
<description>&lt;p&gt;Патч уже &lt;b&gt;вышел&lt;/b&gt;.&lt;/p&gt;</description>
<guid>tag:example.org,a</guid>
<pubDate>Mon, 28 Jul 2026 10:00:00 +0000</pubDate>
</item>
<item>
<title>Без даты</title>
<link>https://example.org/b</link>
</item>
</channel>
</rss>`
const atom = `<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
<title>Example Atom</title>
<entry>
<title>Release 2.0</title>
<link rel="alternate" href="https://example.com/rel"/>
<link rel="edit" href="https://example.com/edit"/>
<id>urn:uuid:1</id>
<updated>2026-07-30T12:30:00Z</updated>
<summary>Ships &amp; works</summary>
</entry>
</feed>`
func TestParseRSS2(t *testing.T) {
f, err := Parse(strings.NewReader(rss2))
if err != nil {
t.Fatal(err)
}
if f.Title != "Хабр" {
t.Fatalf("title = %q", f.Title)
}
if len(f.Items) != 2 {
t.Fatalf("items = %d, want 2", len(f.Items))
}
it := f.Items[0]
if it.Title != "Новая уязвимость в ядре" {
t.Errorf("title = %q", it.Title)
}
if it.Summary != "Патч уже вышел ." && it.Summary != "Патч уже вышел." {
t.Errorf("summary = %q — tags must be stripped and entities decoded", it.Summary)
}
if it.ID != "tag:example.org,a" {
t.Errorf("id = %q", it.ID)
}
if want := time.Date(2026, 7, 28, 10, 0, 0, 0, time.UTC); !it.Published.Equal(want) {
t.Errorf("published = %v, want %v", it.Published, want)
}
if !f.Items[1].Published.IsZero() {
t.Errorf("undated item got a date: %v", f.Items[1].Published)
}
if f.Items[1].ID != "https://example.org/b" {
t.Errorf("id falls back to the link, got %q", f.Items[1].ID)
}
}
func TestParseAtom(t *testing.T) {
f, err := Parse(strings.NewReader(atom))
if err != nil {
t.Fatal(err)
}
if f.Title != "Example Atom" || len(f.Items) != 1 {
t.Fatalf("feed = %+v", f)
}
it := f.Items[0]
if it.Link != "https://example.com/rel" {
t.Errorf("link = %q, want the alternate link", it.Link)
}
if it.Summary != "Ships & works" {
t.Errorf("summary = %q", it.Summary)
}
if want := time.Date(2026, 7, 30, 12, 30, 0, 0, time.UTC); !it.Published.Equal(want) {
t.Errorf("published = %v, want %v", it.Published, want)
}
}
func TestParseGarbage(t *testing.T) {
if _, err := Parse(strings.NewReader("<html><body>not a feed")); err == nil {
t.Fatal("want an error on a non-feed document")
}
// A feed with an item that has neither title nor link contributes nothing
// rather than an empty note.
f, err := Parse(strings.NewReader(`<rss><channel><item><description>x</description></item></channel></rss>`))
if err != nil {
t.Fatal(err)
}
if len(f.Items) != 0 {
t.Fatalf("items = %d, want 0", len(f.Items))
}
}
func TestPlainTextDropsScript(t *testing.T) {
got := PlainText(`<p>hi</p><script>alert("x")</script><style>b{}</style> there`)
if got != "hi there" {
t.Fatalf("got %q", got)
}
}
+324
View File
@@ -0,0 +1,324 @@
package rss
import (
"context"
"fmt"
"log"
"strings"
"time"
)
// FeedConfig — one feed to read. A feed with no Name or no URL is ignored.
type FeedConfig struct {
Name string // short id; the note source is "rss:<Name>"
URL string // http(s) only, enforced by the fetcher
Category string // free text ("технологии"), used to answer "что по X?"
Interval time.Duration // 0 ⇒ the poller's default
Include []string // when non-empty, keep only items matching one of these
Exclude []string // drop items matching any of these, even if included
}
// Fetcher is the guarded HTTP door (internal/webfetch). An interface so the
// poller is testable without a network and so it CANNOT fetch by any other
// means: no http.Client is constructed in this package.
type Fetcher interface {
Get(ctx context.Context, url string) (*Body, error)
}
// Body is the minimum the poller needs from a response.
type Body struct{ Bytes []byte }
// Notes is core's note-writing half. Same shape as ipc.CoreAPI's method, so the
// daemon passes its API straight in.
type Notes interface {
WriteNote(ctx context.Context, ts time.Time, text string, embedding []float32, source string) (int64, error)
}
// Marks remembers how far a feed was read. Durable, because the alternative is
// re-writing yesterday's headlines as fresh notes after every restart. The
// daemon backs this with config facts (key "rss:latest:<feed>").
type Marks interface {
LastMark(ctx context.Context, feed string) (time.Time, error)
SetMark(ctx context.Context, feed string, at time.Time) error
}
// Embedder embeds a note on its way into the store so recall can find it. nil ⇒
// notes are written without a vector (still readable by the recent-notes path).
type Embedder interface {
Embed(ctx context.Context, text string) ([]float32, error)
}
// Ranker is the relevance seam. The plan called for scoring each item against
// an interest profile built from his notes; that profile does not exist yet, and
// a threshold over an embedder with no profile to compare to is a random filter
// with a confident name. So the seam is here, nil in the daemon, and the filter
// that actually runs is the per-feed keyword one — a rule he can read and
// predict. When there IS a profile, implement this and pass it.
//
// Note what a Ranker must NOT be: anything that sends his notes outward. The
// scoring happens locally against a local embedder; the feed item is the input,
// his memory is never the payload.
type Ranker interface {
Relevant(ctx context.Context, text string) (bool, error)
}
// Config — poller-wide settings.
type Config struct {
DefaultInterval time.Duration // 0 ⇒ DefaultPollInterval
MaxItems int // most notes written per feed per poll; 0 ⇒ DefaultMaxItems
MaxAge time.Duration // ignore items older than this on a cold start; 0 ⇒ DefaultMaxAge
}
// Defaults chosen to be quiet: a feed read every half hour, at most a handful of
// items kept, and a cold start that does not import a month of history.
const (
DefaultPollInterval = 30 * time.Minute
DefaultMaxItems = 5
DefaultMaxAge = 24 * time.Hour
)
// Poller reads feeds on a schedule and writes what survives filtering as notes.
type Poller struct {
feeds []FeedConfig
fetch Fetcher
notes Notes
marks Marks
embed Embedder
ranker Ranker
cfg Config
nextDue map[string]time.Time
seen map[string]map[string]bool // feed → item ID, for items with no date
}
// NewPoller wires a poller. Returns nil when there is nothing to poll — a
// capability is off unless configured, and callers check for nil.
func NewPoller(feeds []FeedConfig, fetch Fetcher, notes Notes, marks Marks, embed Embedder, ranker Ranker, cfg Config) *Poller {
var valid []FeedConfig
for _, f := range feeds {
if strings.TrimSpace(f.Name) == "" || strings.TrimSpace(f.URL) == "" {
log.Printf("rss: skipping a feed with no name or no url")
continue
}
valid = append(valid, f)
}
if len(valid) == 0 || fetch == nil || notes == nil {
return nil
}
if cfg.DefaultInterval <= 0 {
cfg.DefaultInterval = DefaultPollInterval
}
if cfg.MaxItems <= 0 {
cfg.MaxItems = DefaultMaxItems
}
if cfg.MaxAge <= 0 {
cfg.MaxAge = DefaultMaxAge
}
return &Poller{
feeds: valid, fetch: fetch, notes: notes, marks: marks,
embed: embed, ranker: ranker, cfg: cfg,
nextDue: map[string]time.Time{},
seen: map[string]map[string]bool{},
}
}
// Feeds returns the configured feeds (the answer path lists categories).
func (p *Poller) Feeds() []FeedConfig { return p.feeds }
// PollDue reads every feed whose interval has elapsed and returns how many
// notes were written. Errors are logged per feed, never returned: one dead feed
// must not stop the others, and there is nobody waiting on this.
func (p *Poller) PollDue(ctx context.Context, now time.Time) int {
written := 0
for _, f := range p.feeds {
if due, ok := p.nextDue[f.Name]; ok && now.Before(due) {
continue
}
interval := f.Interval
if interval <= 0 {
interval = p.cfg.DefaultInterval
}
p.nextDue[f.Name] = now.Add(interval)
n, err := p.PollFeed(ctx, f, now)
if err != nil {
// The URL is configured by him and not a secret, so it is loggable;
// item titles are not logged, only counts.
log.Printf("rss: feed %s: %v", f.Name, err)
continue
}
if n > 0 {
log.Printf("rss: feed %s: %d new item(s) noted", f.Name, n)
}
written += n
}
return written
}
// PollFeed reads one feed now, regardless of its schedule.
func (p *Poller) PollFeed(ctx context.Context, f FeedConfig, now time.Time) (int, error) {
body, err := p.fetch.Get(ctx, f.URL)
if err != nil {
return 0, err
}
feed, err := Parse(strings.NewReader(string(body.Bytes)))
if err != nil {
return 0, err
}
mark := p.mark(ctx, f.Name, now)
newest := mark
written := 0
for _, it := range feed.Items {
if written >= p.cfg.MaxItems {
break
}
if !p.fresh(f, it, mark, now) {
continue
}
if !Matches(f, it) {
continue
}
if p.ranker != nil {
ok, err := p.ranker.Relevant(ctx, it.Title+" "+it.Summary)
if err != nil {
log.Printf("rss: feed %s: relevance: %v", f.Name, err)
} else if !ok {
continue
}
}
if err := p.write(ctx, f, it, now); err != nil {
return written, err
}
written++
if it.Published.After(newest) {
newest = it.Published
}
}
if p.marks != nil && newest.After(mark) {
if err := p.marks.SetMark(ctx, f.Name, newest); err != nil {
log.Printf("rss: feed %s: save mark: %v", f.Name, err)
}
}
return written, nil
}
// mark — how far this feed was read. A feed with no mark starts MaxAge ago, so
// a first poll takes today's headlines instead of the whole archive.
func (p *Poller) mark(ctx context.Context, feed string, now time.Time) time.Time {
cold := now.Add(-p.cfg.MaxAge)
if p.marks == nil {
return cold
}
at, err := p.marks.LastMark(ctx, feed)
if err != nil || at.IsZero() {
return cold
}
return at
}
// fresh — two dedup rules, because feeds are inconsistent about dates. A dated
// item must be newer than the mark; an undated one is kept once per process by
// ID. Both are needed: dates alone re-import undated feeds forever, IDs alone
// lose their memory on restart.
func (p *Poller) fresh(f FeedConfig, it Item, mark, now time.Time) bool {
if !it.Published.IsZero() {
if !it.Published.After(mark) {
return false
}
// A feed that dates its items in the future (or a clock skew) must not
// win the mark and mute everything after it.
return !it.Published.After(now.Add(time.Hour))
}
id := it.ID
if id == "" {
id = it.Title
}
if p.seen[f.Name] == nil {
p.seen[f.Name] = map[string]bool{}
}
if p.seen[f.Name][id] {
return false
}
p.seen[f.Name][id] = true
return true
}
// write stores one item as a note. Source "rss:<feed>" is what the answer path
// filters on, and what makes a feed note distinguishable from something he said.
func (p *Poller) write(ctx context.Context, f FeedConfig, it Item, now time.Time) error {
text := NoteText(f, it)
var vec []float32
if p.embed != nil {
v, err := p.embed.Embed(ctx, text)
if err != nil {
log.Printf("rss: feed %s: embed: %v", f.Name, err)
} else {
vec = v
}
}
ts := it.Published
if ts.IsZero() {
ts = now
}
if _, err := p.notes.WriteNote(ctx, ts, text, vec, SourceFor(f.Name)); err != nil {
return fmt.Errorf("write note: %w", err)
}
return nil
}
// SourceFor is the note source for a feed.
func SourceFor(feed string) string { return "rss:" + feed }
// SourcePrefix — what the answer path matches to find feed notes.
const SourcePrefix = "rss:"
// NoteText renders an item as the note body. The category is included because
// "что нового по технологиям?" is answered by reading notes, and a note has to
// carry enough to be recognised as belonging to that category.
func NoteText(f FeedConfig, it Item) string {
var b strings.Builder
b.WriteString(it.Title)
if f.Category != "" {
fmt.Fprintf(&b, " [%s]", f.Category)
}
if it.Summary != "" {
b.WriteString("\n")
b.WriteString(trimRunes(it.Summary, 500))
}
if it.Link != "" {
b.WriteString("\n")
b.WriteString(it.Link)
}
return b.String()
}
// trimRunes cuts on a rune boundary — a note is Russian as often as English and
// half a cyrillic letter is a broken note.
func trimRunes(s string, max int) string {
r := []rune(s)
if len(r) <= max {
return s
}
return strings.TrimSpace(string(r[:max])) + "…"
}
// Matches applies the per-feed keyword filter: keep when Include is empty or one
// include matches, drop when any exclude matches. Case-insensitive substring,
// which for Russian is the honest choice — no stemmer here, so "выборы" does not
// match "выборах", and a filter he writes is a filter he can predict.
func Matches(f FeedConfig, it Item) bool {
hay := strings.ToLower(it.Title + " " + it.Summary)
for _, x := range f.Exclude {
if x = strings.ToLower(strings.TrimSpace(x)); x != "" && strings.Contains(hay, x) {
return false
}
}
if len(f.Include) == 0 {
return true
}
for _, in := range f.Include {
if in = strings.ToLower(strings.TrimSpace(in)); in != "" && strings.Contains(hay, in) {
return true
}
}
return false
}
+210
View File
@@ -0,0 +1,210 @@
package rss
import (
"context"
"errors"
"strings"
"testing"
"time"
)
type fakeFetch struct {
body string
err error
calls int
urls []string
}
func (f *fakeFetch) Get(_ context.Context, url string) (*Body, error) {
f.calls++
f.urls = append(f.urls, url)
if f.err != nil {
return nil, f.err
}
return &Body{Bytes: []byte(f.body)}, nil
}
type writtenNote struct {
ts time.Time
text string
source string
vec []float32
}
type fakeNotes struct{ notes []writtenNote }
func (n *fakeNotes) WriteNote(_ context.Context, ts time.Time, text string, vec []float32, source string) (int64, error) {
n.notes = append(n.notes, writtenNote{ts, text, source, vec})
return int64(len(n.notes)), nil
}
type fakeMarks struct{ m map[string]time.Time }
func newMarks() *fakeMarks { return &fakeMarks{m: map[string]time.Time{}} }
func (f *fakeMarks) LastMark(_ context.Context, feed string) (time.Time, error) {
return f.m[feed], nil
}
func (f *fakeMarks) SetMark(_ context.Context, feed string, at time.Time) error {
f.m[feed] = at
return nil
}
var now = time.Date(2026, 7, 28, 12, 0, 0, 0, time.UTC)
func TestPollWritesNotesWithSource(t *testing.T) {
fetch := &fakeFetch{body: rss2}
notes := &fakeNotes{}
marks := newMarks()
p := NewPoller([]FeedConfig{{Name: "habr", URL: "https://example.org/rss", Category: "технологии"}},
fetch, notes, marks, nil, nil, Config{})
if p == nil {
t.Fatal("NewPoller returned nil for a configured feed")
}
n := p.PollDue(context.Background(), now)
if n != 2 || len(notes.notes) != 2 {
t.Fatalf("wrote %d notes (returned %d), want 2", len(notes.notes), n)
}
if notes.notes[0].source != "rss:habr" {
t.Errorf("source = %q, want rss:habr", notes.notes[0].source)
}
if !strings.Contains(notes.notes[0].text, "технологии") {
t.Errorf("note does not carry its category: %q", notes.notes[0].text)
}
if !strings.Contains(notes.notes[0].text, "https://example.org/a") {
t.Errorf("note does not carry its link: %q", notes.notes[0].text)
}
// The undated item is stamped with now, the dated one with its own date.
if !notes.notes[1].ts.Equal(now) {
t.Errorf("undated item ts = %v, want now", notes.notes[1].ts)
}
}
// The whole point of a mark: polling twice must not re-note the same headlines.
func TestSecondPollIsQuiet(t *testing.T) {
fetch := &fakeFetch{body: rss2}
notes := &fakeNotes{}
p := NewPoller([]FeedConfig{{Name: "habr", URL: "u", Interval: time.Minute}}, fetch, notes, newMarks(), nil, nil, Config{})
p.PollDue(context.Background(), now)
before := len(notes.notes)
p.PollDue(context.Background(), now.Add(2*time.Minute))
if len(notes.notes) != before {
t.Fatalf("second poll wrote %d extra notes", len(notes.notes)-before)
}
}
// A mark that survives a restart is the durable half; simulate one by building a
// fresh poller over the same marks.
func TestMarkSurvivesRestart(t *testing.T) {
marks := newMarks()
fetch := &fakeFetch{body: rss2}
notes := &fakeNotes{}
feeds := []FeedConfig{{Name: "habr", URL: "u"}}
NewPoller(feeds, fetch, notes, marks, nil, nil, Config{}).PollDue(context.Background(), now)
if len(notes.notes) != 2 {
t.Fatalf("first run wrote %d", len(notes.notes))
}
notes2 := &fakeNotes{}
NewPoller(feeds, fetch, notes2, marks, nil, nil, Config{}).PollDue(context.Background(), now.Add(time.Hour))
// The dated item is behind the mark. The undated one has no date to compare,
// so it comes back — accepted and documented in fresh(): an undated feed is
// deduped per process, not forever.
for _, n := range notes2.notes {
if strings.Contains(n.text, "уязвимость") {
t.Fatalf("dated item re-noted after restart: %q", n.text)
}
}
}
func TestIntervalIsRespected(t *testing.T) {
fetch := &fakeFetch{body: rss2}
p := NewPoller([]FeedConfig{{Name: "habr", URL: "u", Interval: time.Hour}}, fetch, &fakeNotes{}, newMarks(), nil, nil, Config{})
p.PollDue(context.Background(), now)
p.PollDue(context.Background(), now.Add(time.Minute))
if fetch.calls != 1 {
t.Fatalf("fetched %d times inside one interval, want 1", fetch.calls)
}
p.PollDue(context.Background(), now.Add(2*time.Hour))
if fetch.calls != 2 {
t.Fatalf("fetched %d times, want 2 after the interval elapsed", fetch.calls)
}
}
func TestColdStartIgnoresOldItems(t *testing.T) {
old := `<rss><channel><item><title>Старое</title><link>l</link>` +
`<pubDate>Mon, 01 Jun 2026 10:00:00 +0000</pubDate></item></channel></rss>`
notes := &fakeNotes{}
p := NewPoller([]FeedConfig{{Name: "f", URL: "u"}}, &fakeFetch{body: old}, notes, newMarks(), nil, nil, Config{MaxAge: 24 * time.Hour})
if n := p.PollDue(context.Background(), now); n != 0 {
t.Fatalf("cold start imported %d old items, want 0", n)
}
}
func TestMaxItemsCap(t *testing.T) {
var b strings.Builder
b.WriteString("<rss><channel>")
for i := 0; i < 10; i++ {
b.WriteString("<item><title>t")
b.WriteByte(byte('0' + i))
b.WriteString("</title><link>https://example.org/")
b.WriteByte(byte('0' + i))
b.WriteString("</link></item>")
}
b.WriteString("</channel></rss>")
notes := &fakeNotes{}
p := NewPoller([]FeedConfig{{Name: "f", URL: "u"}}, &fakeFetch{body: b.String()}, notes, newMarks(), nil, nil, Config{MaxItems: 3})
if n := p.PollDue(context.Background(), now); n != 3 {
t.Fatalf("wrote %d notes, want the cap of 3", n)
}
}
func TestKeywordFilter(t *testing.T) {
f := FeedConfig{Include: []string{"ядр"}, Exclude: []string{"реклама"}}
if !Matches(f, Item{Title: "Новое ядро"}) {
t.Error("include did not match")
}
if Matches(f, Item{Title: "Новое ядро", Summary: "Реклама внутри"}) {
t.Error("exclude must win over include")
}
if Matches(f, Item{Title: "Погода"}) {
t.Error("non-matching item passed the include filter")
}
if !Matches(FeedConfig{}, Item{Title: "что угодно"}) {
t.Error("an unfiltered feed must keep everything")
}
}
type fakeRanker struct{ keep bool }
func (r fakeRanker) Relevant(context.Context, string) (bool, error) { return r.keep, nil }
func TestRankerCanDropEverything(t *testing.T) {
notes := &fakeNotes{}
p := NewPoller([]FeedConfig{{Name: "f", URL: "u"}}, &fakeFetch{body: rss2}, notes, newMarks(), nil, fakeRanker{false}, Config{})
if n := p.PollDue(context.Background(), now); n != 0 {
t.Fatalf("ranker rejected everything but %d notes were written", n)
}
}
func TestFetchErrorIsSurvivable(t *testing.T) {
notes := &fakeNotes{}
p := NewPoller([]FeedConfig{
{Name: "dead", URL: "u1"},
{Name: "live", URL: "u2"},
}, &fakeFetch{err: errors.New("boom")}, notes, newMarks(), nil, nil, Config{})
if n := p.PollDue(context.Background(), now); n != 0 {
t.Fatalf("n = %d", n)
}
// Both feeds were attempted: one dead feed does not abort the round.
if p.nextDue["live"].IsZero() {
t.Fatal("the second feed was never attempted")
}
}
func TestNoFeedsMeansNoPoller(t *testing.T) {
if p := NewPoller(nil, &fakeFetch{}, &fakeNotes{}, nil, nil, nil, Config{}); p != nil {
t.Fatal("NewPoller must return nil when nothing is configured")
}
if p := NewPoller([]FeedConfig{{Name: "", URL: ""}}, &fakeFetch{}, &fakeNotes{}, nil, nil, nil, Config{}); p != nil {
t.Fatal("a feed with no name or url is not a configuration")
}
}
+238
View File
@@ -0,0 +1,238 @@
// Package tasks ranks captured work (Vikunja #129).
//
// The ordering is COMPUTED, not generated. Asking a 1.7B model which of his
// tasks matters most would produce a fluent opinion about his life with no
// basis in anything, and a confidently wrong priority is worse than no
// priority at all — the same reasoning as the behaviour profile in
// internal/memory, which counts instead of summarising.
//
// So: four signals, all of them things he told her, and a reason string naming
// the one that decided each row. Nothing here invents urgency. A task with no
// due date and no weight scores nothing and sits where its age puts it, which
// is the honest answer to "which of these matters?" when he never said.
//
// Ranking is a READ. It sorts and renders; it never writes, schedules or
// announces. Maven is not a nag: a task rising to the top of this list is not a
// reason to speak, only the order she recites in when asked.
package tasks
import (
"fmt"
"sort"
"strings"
"time"
)
// Status values, mirroring internal/store so a caller can rank ipc.Task rows
// without importing the store.
const (
StatusCandidate = "candidate"
StatusOpen = "open"
)
// Item — one task to rank. The subset of a task that ranking depends on;
// callers map their own row type onto it.
type Item struct {
ID int64
Text string
Status string
Created time.Time
Due *time.Time
Weight int
}
// Ranked — one task with its score and the reason that decided it.
type Ranked struct {
Item
Score float64
// Reason — the dominant signal, in Russian, for the page and the spoken
// list. Empty when nothing distinguished this task: no due date, no
// weight, not old. Saying "потому что" about a task he never prioritised
// would be making something up.
Reason string
}
// Scoring weights. Deliberately coarse round numbers: this is a knob, not
// math, and the only property that has to hold is the ordering between classes
// (overdue beats today beats this week beats undated).
const (
scoreOverdue = 100 // he already missed it
scoreOverduePer = 5 // per further day late, capped
scoreOverdueCap = 40
scoreDueToday = 60
scoreDueTomorrow = 40
scoreDueWeek = 20
scoreDueLater = 5
scorePerWeight = 15 // "срочно" / "важно" / the web form's select
scorePerWeekOld = 1 // so nothing rots at the bottom forever
scoreAgeCap = 10
// MaxWeight — the highest importance hint capture accepts. Three rungs is
// as many as anyone can rank by hand honestly.
MaxWeight = 3
)
// Rank scores every item and returns them ordered: confirmed work first, then
// candidates, each by score descending, oldest first on a tie.
//
// Candidates never outrank open work, whatever their due date. A task Maven
// derived from something she read is a suggestion until he confirms it, and
// putting her guess above his own stated work would be reading his priorities
// back to him wrong.
func Rank(items []Item, now time.Time) []Ranked {
out := make([]Ranked, 0, len(items))
for _, it := range items {
score, reason := score(it, now)
out = append(out, Ranked{Item: it, Score: score, Reason: reason})
}
sort.SliceStable(out, func(i, j int) bool {
ci, cj := out[i].Status == StatusCandidate, out[j].Status == StatusCandidate
if ci != cj {
return !ci // open before candidate
}
if out[i].Score != out[j].Score {
return out[i].Score > out[j].Score
}
return out[i].Created.Before(out[j].Created) // oldest first, FIFO
})
return out
}
// score — the per-item scoring function. Returns the score and the dominant
// reason. Deadline beats weight when both are present: a date is a fact about
// the world, a weight is how he felt when he filed it.
func score(it Item, now time.Time) (float64, string) {
var total float64
reason := ""
if it.Due != nil {
days := dayDelta(*it.Due, now)
switch {
case days < 0:
late := -days
bonus := float64(late * scoreOverduePer)
if bonus > scoreOverdueCap {
bonus = scoreOverdueCap
}
total += scoreOverdue + bonus
reason = "просрочено"
if late == 1 {
reason = "просрочено на день"
} else if late > 1 {
reason = fmt.Sprintf("просрочено на %d дн.", late)
}
case days == 0:
total += scoreDueToday
reason = "сегодня"
case days == 1:
total += scoreDueTomorrow
reason = "завтра"
case days <= 7:
total += scoreDueWeek
reason = fmt.Sprintf("через %d дн.", days)
default:
total += scoreDueLater
}
}
w := it.Weight
if w > MaxWeight {
w = MaxWeight
}
if w > 0 {
total += float64(w * scorePerWeight)
if reason == "" {
reason = "важно"
}
}
if !it.Created.IsZero() {
weeks := int(now.Sub(it.Created).Hours() / (24 * 7))
if weeks > 0 {
age := float64(weeks * scorePerWeekOld)
if age > scoreAgeCap {
age = scoreAgeCap
}
total += age
if reason == "" && weeks >= 2 {
reason = "давно в списке"
}
}
}
return total, reason
}
// dayDelta — calendar days from now to due, in due's own location. Whole days,
// not hours: a task due today is due today whether it is 09:00 or 23:00, and an
// hours-based comparison would call this evening's task "overdue" all afternoon.
func dayDelta(due, now time.Time) int {
loc := due.Location()
d := time.Date(due.Year(), due.Month(), due.Day(), 0, 0, 0, 0, loc)
n := now.In(loc)
n = time.Date(n.Year(), n.Month(), n.Day(), 0, 0, 0, 0, loc)
return int(d.Sub(n).Hours() / 24)
}
// SpokenLimit — how many tasks the spoken list names before it summarises the
// rest. A recital of twenty items is noise; five is a list he can hold.
const SpokenLimit = 5
// FormatRU renders a ranked list the way Maven says it. Confirmed work first,
// with the reason attached where there is one; candidates named as
// unconfirmed, never recited as his work.
//
// One renderer for the voice reply and the web page, for the same reason
// DayPlan.Spoken is built core-side: two formatters drift, and then she says
// one order and shows another.
func FormatRU(ranked []Ranked) string {
var open, cands []Ranked
for _, r := range ranked {
if r.Status == StatusCandidate {
cands = append(cands, r)
} else {
open = append(open, r)
}
}
if len(open) == 0 && len(cands) == 0 {
return "задач нет."
}
var b strings.Builder
if len(open) > 0 {
b.WriteString("сначала: ")
b.WriteString(joinRU(open, SpokenLimit, true))
b.WriteString(".")
}
if len(cands) > 0 {
if b.Len() > 0 {
b.WriteString(" ")
}
b.WriteString("ещё я нашла, но ты не подтвердил: ")
b.WriteString(joinRU(cands, SpokenLimit, false))
b.WriteString(".")
}
return b.String()
}
// joinRU lists up to limit tasks, then says how many are left. withReasons
// attaches the parenthesised reason — candidates are listed bare, since their
// due dates are Maven's reading of a mail and not something he stated.
func joinRU(rs []Ranked, limit int, withReasons bool) string {
shown := rs
rest := 0
if len(rs) > limit {
shown, rest = rs[:limit], len(rs)-limit
}
parts := make([]string, 0, len(shown))
for _, r := range shown {
if withReasons && r.Reason != "" {
parts = append(parts, r.Text+" ("+r.Reason+")")
} else {
parts = append(parts, r.Text)
}
}
s := strings.Join(parts, "; ")
if rest > 0 {
s += fmt.Sprintf("; и ещё %d", rest)
}
return s
}
+177
View File
@@ -0,0 +1,177 @@
package tasks
import (
"strings"
"testing"
"time"
)
func at(y int, m time.Month, d int) *time.Time {
t := time.Date(y, m, d, 0, 0, 0, 0, time.UTC)
return &t
}
func now() time.Time { return time.Date(2026, 8, 1, 14, 0, 0, 0, time.UTC) }
func texts(rs []Ranked) []string {
out := make([]string, len(rs))
for i, r := range rs {
out[i] = r.Text
}
return out
}
func TestRankOrdersByDeadline(t *testing.T) {
items := []Item{
{ID: 1, Text: "через неделю", Status: StatusOpen, Due: at(2026, 8, 7), Created: now()},
{ID: 2, Text: "просрочено", Status: StatusOpen, Due: at(2026, 7, 28), Created: now()},
{ID: 3, Text: "без срока", Status: StatusOpen, Created: now()},
{ID: 4, Text: "сегодня", Status: StatusOpen, Due: at(2026, 8, 1), Created: now()},
{ID: 5, Text: "завтра", Status: StatusOpen, Due: at(2026, 8, 2), Created: now()},
}
got := texts(Rank(items, now()))
want := []string{"просрочено", "сегодня", "завтра", "через неделю", "без срока"}
for i := range want {
if got[i] != want[i] {
t.Fatalf("order = %v, want %v", got, want)
}
}
}
func TestRankCandidatesNeverOutrankOpenWork(t *testing.T) {
items := []Item{
{ID: 1, Text: "его задача", Status: StatusOpen, Created: now()},
// Everything about this one screams urgent — and it is still a guess.
{ID: 2, Text: "из письма", Status: StatusCandidate, Due: at(2026, 7, 1), Weight: 3, Created: now()},
}
got := Rank(items, now())
if got[0].Text != "его задача" {
t.Errorf("order = %v, want his own work first", texts(got))
}
}
func TestRankWeightLiftsUndatedWork(t *testing.T) {
items := []Item{
{ID: 1, Text: "обычная", Status: StatusOpen, Created: now()},
{ID: 2, Text: "важная", Status: StatusOpen, Weight: 2, Created: now()},
}
got := Rank(items, now())
if got[0].Text != "важная" {
t.Errorf("order = %v, want the weighted task first", texts(got))
}
if got[0].Reason != "важно" {
t.Errorf("reason = %q, want важно", got[0].Reason)
}
// A deadline still beats a weight: a date is a fact, a weight is a feeling.
items = append(items, Item{ID: 3, Text: "сегодня", Status: StatusOpen, Due: at(2026, 8, 1), Created: now()})
got = Rank(items, now())
if got[0].Text != "сегодня" {
t.Errorf("order = %v, want the dated task first", texts(got))
}
}
func TestRankOldestFirstOnATie(t *testing.T) {
old := now().AddDate(0, 0, -3)
items := []Item{
{ID: 1, Text: "новая", Status: StatusOpen, Created: now()},
{ID: 2, Text: "старая", Status: StatusOpen, Created: old},
}
got := Rank(items, now())
if got[0].Text != "старая" {
t.Errorf("order = %v, want FIFO on equal urgency", texts(got))
}
}
func TestRankNoInventedReason(t *testing.T) {
got := Rank([]Item{{ID: 1, Text: "что-то", Status: StatusOpen, Created: now()}}, now())
if got[0].Reason != "" {
t.Errorf("reason = %q — nothing distinguished this task, so there is nothing to say", got[0].Reason)
}
if got[0].Score != 0 {
t.Errorf("score = %v, want 0", got[0].Score)
}
}
func TestRankAgeIsCappedAndNamed(t *testing.T) {
items := []Item{
{ID: 1, Text: "прошлогодняя", Status: StatusOpen, Created: now().AddDate(-1, 0, 0)},
{ID: 2, Text: "трёхнедельная", Status: StatusOpen, Created: now().AddDate(0, 0, -21)},
}
got := Rank(items, now())
if got[0].Score != scoreAgeCap {
t.Errorf("oldest score = %v, want the cap %v", got[0].Score, float64(scoreAgeCap))
}
if got[0].Reason != "давно в списке" {
t.Errorf("reason = %q", got[0].Reason)
}
}
// A task due at 23:00 today is due today, not overdue since this morning.
func TestRankDueTodayIsNotOverdue(t *testing.T) {
due := time.Date(2026, 8, 1, 23, 0, 0, 0, time.UTC)
got := Rank([]Item{{ID: 1, Text: "вечером", Status: StatusOpen, Due: &due, Created: now()}}, now())
if got[0].Reason != "сегодня" {
t.Errorf("reason = %q, want сегодня", got[0].Reason)
}
}
func TestRankOverdueDaysAreCounted(t *testing.T) {
got := Rank([]Item{
{ID: 1, Text: "вчера", Status: StatusOpen, Due: at(2026, 7, 31), Created: now()},
{ID: 2, Text: "давно", Status: StatusOpen, Due: at(2026, 7, 20), Created: now()},
}, now())
if got[0].Text != "давно" {
t.Errorf("order = %v, want the later-overdue task first", texts(got))
}
if got[0].Reason != "просрочено на 12 дн." {
t.Errorf("reason = %q", got[0].Reason)
}
if got[1].Reason != "просрочено на день" {
t.Errorf("reason = %q", got[1].Reason)
}
}
func TestFormatRUNamesReasonsAndSeparatesCandidates(t *testing.T) {
ranked := Rank([]Item{
{ID: 1, Text: "оплатить интернет", Status: StatusOpen, Due: at(2026, 8, 1), Created: now()},
{ID: 2, Text: "купить молоко", Status: StatusOpen, Created: now()},
{ID: 3, Text: "продлить страховку", Status: StatusCandidate, Due: at(2026, 7, 1), Created: now()},
}, now())
got := FormatRU(ranked)
if !strings.HasPrefix(got, "сначала: оплатить интернет (сегодня)") {
t.Errorf("reply = %q", got)
}
if !strings.Contains(got, "не подтвердил: продлить страховку") {
t.Errorf("candidate not named as unconfirmed: %q", got)
}
// A candidate's due date is Maven's reading of a mail, not his statement.
if strings.Contains(got, "продлить страховку (") {
t.Errorf("a candidate must be listed without a reason: %q", got)
}
// Persona: nothing masculine, no pet names, informal address only.
for _, bad := range []string{"рад ", "понял ", "милый", "дорогой", "вам", "ваши"} {
if strings.Contains(got, bad) {
t.Errorf("reply %q contains %q", got, bad)
}
}
}
func TestFormatRUCapsTheSpokenList(t *testing.T) {
var items []Item
for i := 0; i < SpokenLimit+3; i++ {
items = append(items, Item{ID: int64(i), Text: "задача", Status: StatusOpen, Created: now()})
}
got := FormatRU(Rank(items, now()))
if !strings.Contains(got, "и ещё 3") {
t.Errorf("reply = %q, want the tail summarised", got)
}
if strings.Count(got, "задача") != SpokenLimit {
t.Errorf("reply = %q, want exactly %d named", got, SpokenLimit)
}
}
func TestFormatRUEmpty(t *testing.T) {
if got := FormatRU(nil); got != "задач нет." {
t.Errorf("reply = %q", got)
}
}
+301
View File
@@ -0,0 +1,301 @@
// Package webfetch is the one door Maven uses to read something off the
// network, and it is a narrow one.
//
// "Never phones home" stopped being a hard constraint on 2026-07-31, but what
// replaced it is not "she may fetch anything": local sources come first (Kiwix
// on the box), external fetching is off unless configured, and only the
// utterance ever leaves — never his notes, facts or history. That policy is
// enforced by the callers. What THIS package enforces is the part that must be
// code rather than a paragraph in a plan, because it protects the homelab from
// its own assistant:
//
// - http/https only — no file://, no ftp://, no gopher;
// - no private address, ever: loopback, RFC1918 (which is what makes the
// 10.42.0.0/24 wireguard tunnel and the 192.168.1.0/24 LAN unreachable),
// link-local incl. the 169.254.169.254 cloud metadata address, CGNAT,
// unique-local v6. Checked in the dialer's Control hook, so it holds for
// every address the resolver returns AND for every hop of a redirect
// chain — a DNS name that resolves to 127.0.0.1 is refused at connect
// time, which a pre-flight lookup could not promise (rebinding);
// - an allowlist, when one is configured, and a denylist that always wins;
// - a response size cap, a total timeout, a redirect cap;
// - one request per host per interval, so a poll loop with a bug is slow
// rather than an outbound flood.
//
// Everything above is on by default with sane numbers: a zero Config is a
// usable, conservative fetcher. There is no cache and no retry — a feed poll
// or a page read that fails is simply not answered this round.
package webfetch
import (
"context"
"errors"
"fmt"
"io"
"net"
"net/http"
"net/url"
"strings"
"sync"
"syscall"
"time"
)
// Defaults. Small on purpose: this reads feeds and article pages, not ISOs.
const (
DefaultTimeout = 20 * time.Second
DefaultMaxBytes = 2 << 20 // 2 MiB
DefaultMaxRedirects = 3
DefaultHostInterval = time.Second
DefaultUserAgent = "Maven/1.0 (self-hosted personal assistant)"
)
// Errors callers distinguish. Everything else is wrapped transport error.
var (
ErrScheme = errors.New("webfetch: only http and https are allowed")
ErrBlocked = errors.New("webfetch: host is not allowed")
ErrPrivate = errors.New("webfetch: refusing to connect to a private address")
ErrTooLarge = errors.New("webfetch: response exceeds the size cap")
ErrRedirects = errors.New("webfetch: too many redirects")
ErrStatus = errors.New("webfetch: non-2xx status")
)
// Config are the limits. Every zero value means "the default above", so
// Config{} is safe; the only field that changes behaviour by being empty is
// AllowHosts (empty ⇒ any public host that is not denied).
type Config struct {
// AllowHosts — when non-empty, the ONLY hosts that may be fetched. An
// entry matches the host itself and its subdomains ("example.com" allows
// "news.example.com"). This is the knob to reach for when a capability
// should read two feeds and nothing else.
AllowHosts []string
// DenyHosts — same matching, checked first and always winning.
DenyHosts []string
Timeout time.Duration // whole request, including redirects and body read
MaxBytes int64 // response body cap
MaxRedirects int // 0 ⇒ default; negative ⇒ no redirects followed
HostInterval time.Duration // minimum spacing between requests to one host
UserAgent string
// AllowPrivate disables the private-address guard. It exists for tests
// (httptest listens on 127.0.0.1) and for an explicitly configured
// on-box mirror. Nothing in deploy/mavend.json sets it, and it should
// stay that way: with it on, any URL Maven is handed becomes an SSRF
// probe of the LAN and the wireguard range.
AllowPrivate bool
}
// Response is a fetched body, already bounded by MaxBytes.
type Response struct {
URL string // final URL after redirects
Status int
ContentType string
Body []byte
}
// Fetcher performs guarded GETs. Safe for concurrent use; the per-host rate
// limiter is shared, which is the point of sharing one Fetcher.
type Fetcher struct {
cfg Config
http *http.Client
mu sync.Mutex
last map[string]time.Time // host → when we last dialed it
}
// New builds a fetcher from cfg, filling in defaults.
func New(cfg Config) *Fetcher {
if cfg.Timeout <= 0 {
cfg.Timeout = DefaultTimeout
}
if cfg.MaxBytes <= 0 {
cfg.MaxBytes = DefaultMaxBytes
}
if cfg.MaxRedirects == 0 {
cfg.MaxRedirects = DefaultMaxRedirects
}
if cfg.HostInterval <= 0 {
cfg.HostInterval = DefaultHostInterval
}
if cfg.UserAgent == "" {
cfg.UserAgent = DefaultUserAgent
}
f := &Fetcher{cfg: cfg, last: map[string]time.Time{}}
dialer := &net.Dialer{Timeout: 10 * time.Second}
if !cfg.AllowPrivate {
// The guard lives here rather than in a pre-flight net.LookupHost so
// that it sees the address actually being connected to: every A/AAAA
// the resolver handed back, on every redirect hop, with no window in
// which the name could be re-pointed at the LAN.
dialer.Control = func(_, address string, _ syscall.RawConn) error {
host, _, err := net.SplitHostPort(address)
if err != nil {
return err
}
ip := net.ParseIP(host)
if ip == nil || IsPrivateIP(ip) {
return fmt.Errorf("%w: %s", ErrPrivate, host)
}
return nil
}
}
f.http = &http.Client{
Timeout: cfg.Timeout,
Transport: &http.Transport{DialContext: dialer.DialContext},
CheckRedirect: func(req *http.Request, via []*http.Request) error {
if len(via) > f.cfg.MaxRedirects {
return ErrRedirects
}
// A redirect is a fresh URL and gets the full check: an allowed
// host must not be able to bounce us onto a denied one.
return f.checkURL(req.URL)
},
}
return f
}
// Get fetches rawURL. The body is capped: a larger response is an error, not a
// truncation, because half an XML document is worse than none.
func (f *Fetcher) Get(ctx context.Context, rawURL string) (*Response, error) {
u, err := url.Parse(strings.TrimSpace(rawURL))
if err != nil {
return nil, fmt.Errorf("webfetch: bad url %q: %w", rawURL, err)
}
if err := f.checkURL(u); err != nil {
return nil, err
}
if err := f.waitTurn(ctx, u.Hostname()); err != nil {
return nil, err
}
req, err := http.NewRequestWithContext(ctx, http.MethodGet, u.String(), nil)
if err != nil {
return nil, err
}
req.Header.Set("User-Agent", f.cfg.UserAgent)
req.Header.Set("Accept-Encoding", "identity")
resp, err := f.http.Do(req)
if err != nil {
// http.Client wraps our sentinels in *url.Error; unwrap so callers can
// still tell "blocked" from "the network is down".
for _, sentinel := range []error{ErrPrivate, ErrBlocked, ErrRedirects, ErrScheme} {
if errors.Is(err, sentinel) {
return nil, err
}
}
return nil, err
}
defer resp.Body.Close()
body, err := io.ReadAll(io.LimitReader(resp.Body, f.cfg.MaxBytes+1))
if err != nil {
return nil, err
}
if int64(len(body)) > f.cfg.MaxBytes {
return nil, fmt.Errorf("%w (%d bytes)", ErrTooLarge, f.cfg.MaxBytes)
}
if resp.StatusCode < 200 || resp.StatusCode > 299 {
return nil, fmt.Errorf("%w: %d", ErrStatus, resp.StatusCode)
}
return &Response{
URL: resp.Request.URL.String(),
Status: resp.StatusCode,
ContentType: resp.Header.Get("Content-Type"),
Body: body,
}, nil
}
// checkURL applies the scheme rule and the host lists. The address rule is the
// dialer's job (see New).
func (f *Fetcher) checkURL(u *url.URL) error {
switch u.Scheme {
case "http", "https":
default:
return fmt.Errorf("%w: %q", ErrScheme, u.Scheme)
}
host := strings.ToLower(u.Hostname())
if host == "" {
return fmt.Errorf("%w: no host", ErrBlocked)
}
if HostMatches(host, f.cfg.DenyHosts) {
return fmt.Errorf("%w: %s is denied", ErrBlocked, host)
}
if len(f.cfg.AllowHosts) > 0 && !HostMatches(host, f.cfg.AllowHosts) {
return fmt.Errorf("%w: %s is not on the allowlist", ErrBlocked, host)
}
// A literal private address is refused here as well as in the dialer, so
// the error is the specific one even when no connection is attempted.
if !f.cfg.AllowPrivate {
if ip := net.ParseIP(host); ip != nil && IsPrivateIP(ip) {
return fmt.Errorf("%w: %s", ErrPrivate, host)
}
}
return nil
}
// waitTurn blocks until this host's rate-limit interval has elapsed. It holds
// no lock while sleeping, so two hosts never wait on each other.
func (f *Fetcher) waitTurn(ctx context.Context, host string) error {
for {
f.mu.Lock()
now := time.Now()
earliest := f.last[host].Add(f.cfg.HostInterval)
if !now.Before(earliest) {
f.last[host] = now
f.mu.Unlock()
return nil
}
f.mu.Unlock()
wait := time.NewTimer(earliest.Sub(now))
select {
case <-ctx.Done():
wait.Stop()
return ctx.Err()
case <-wait.C:
}
}
}
// HostMatches reports whether host equals one of pats or is a subdomain of one.
// Exported because the crawler applies the same rule to links it decides not to
// follow, before it ever builds a request.
func HostMatches(host string, pats []string) bool {
host = strings.ToLower(strings.TrimSuffix(host, "."))
for _, p := range pats {
p = strings.ToLower(strings.TrimSpace(strings.TrimPrefix(p, "*.")))
if p == "" {
continue
}
if host == p || strings.HasSuffix(host, "."+p) {
return true
}
}
return false
}
// cgnat is 100.64.0.0/10 — carrier NAT, not covered by net.IP's helpers and not
// somewhere a personal assistant has business connecting.
var cgnat = &net.IPNet{IP: net.IPv4(100, 64, 0, 0).To4(), Mask: net.CIDRMask(10, 32)}
// IsPrivateIP reports whether ip is somewhere Maven must never reach out to:
// the box itself, the LAN, the wireguard range (10.42.0.0/24 ⊂ 10/8), the cloud
// metadata address (169.254.169.254 ⊂ link-local), or anything unroutable.
func IsPrivateIP(ip net.IP) bool {
if ip.IsLoopback() || ip.IsPrivate() || ip.IsUnspecified() ||
ip.IsLinkLocalUnicast() || ip.IsLinkLocalMulticast() ||
ip.IsInterfaceLocalMulticast() || ip.IsMulticast() {
return true
}
if v4 := ip.To4(); v4 != nil && cgnat.Contains(v4) {
return true
}
// IPv4-mapped/compatible forms of the above are handled by To4() inside the
// stdlib helpers; what is left is v6 unique-local (fc00::/7).
if len(ip) == net.IPv6len && ip.To4() == nil && ip[0]&0xfe == 0xfc {
return true
}
return false
}
+227
View File
@@ -0,0 +1,227 @@
package webfetch
import (
"context"
"errors"
"net"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
)
// The limits in this package are the reason a crawler is allowed to exist on
// this box at all, so each one has a test that fails loudly if it is removed.
func TestPrivateAddressesAreRefused(t *testing.T) {
// The wireguard range (10.42.0.0/24), the LAN (192.168.1.0/24) and the
// cloud metadata address are the three that matter here; the rest come
// along for free.
for _, s := range []string{
"127.0.0.1", "127.1.2.3", "10.42.0.7", "10.0.0.5", "192.168.1.104",
"172.16.4.4", "169.254.169.254", "100.64.1.1", "0.0.0.0",
"::1", "fc00::1", "fd12:3456::1", "fe80::1",
} {
if !IsPrivateIP(net.ParseIP(s)) {
t.Errorf("IsPrivateIP(%s) = false, want true", s)
}
}
for _, s := range []string{"8.8.8.8", "1.1.1.1", "93.184.216.34", "2606:2800:220:1::1"} {
if IsPrivateIP(net.ParseIP(s)) {
t.Errorf("IsPrivateIP(%s) = true, want false", s)
}
}
}
func TestGetRefusesPrivateLiteral(t *testing.T) {
f := New(Config{})
for _, u := range []string{
"http://127.0.0.1:8034/search",
"http://10.42.0.1/",
"http://192.168.1.104/dash",
"http://[::1]:9100/mcp",
} {
if _, err := f.Get(context.Background(), u); !errors.Is(err, ErrPrivate) {
t.Errorf("Get(%s) error = %v, want ErrPrivate", u, err)
}
}
}
// A hostname that resolves into private space must fail too — that is the
// rebinding case, and it is why the check lives in the dialer.
func TestGetRefusesPrivateResolution(t *testing.T) {
f := New(Config{})
if _, err := f.Get(context.Background(), "http://localhost:8034/"); !errors.Is(err, ErrPrivate) {
t.Fatalf("Get(localhost) error = %v, want ErrPrivate", err)
}
}
func TestGetRefusesNonHTTPSchemes(t *testing.T) {
f := New(Config{})
for _, u := range []string{"file:///etc/passwd", "ftp://example.com/x", "gopher://example.com"} {
if _, err := f.Get(context.Background(), u); !errors.Is(err, ErrScheme) {
t.Errorf("Get(%s) error = %v, want ErrScheme", u, err)
}
}
}
// testFetcher — a fetcher pointed at an httptest server, which necessarily
// listens on loopback. AllowPrivate is the test-only escape hatch.
func testFetcher(t *testing.T, cfg Config) *Fetcher {
t.Helper()
cfg.AllowPrivate = true
if cfg.HostInterval == 0 {
cfg.HostInterval = time.Nanosecond
}
return New(cfg)
}
func TestAllowAndDenyLists(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Write([]byte("ok"))
}))
defer srv.Close()
f := testFetcher(t, Config{AllowHosts: []string{"example.com"}})
if _, err := f.Get(context.Background(), srv.URL); !errors.Is(err, ErrBlocked) {
t.Fatalf("off-allowlist host: error = %v, want ErrBlocked", err)
}
f = testFetcher(t, Config{DenyHosts: []string{"127.0.0.1"}})
if _, err := f.Get(context.Background(), srv.URL); !errors.Is(err, ErrBlocked) {
t.Fatalf("denied host: error = %v, want ErrBlocked", err)
}
f = testFetcher(t, Config{AllowHosts: []string{"127.0.0.1"}})
if _, err := f.Get(context.Background(), srv.URL); err != nil {
t.Fatalf("allowlisted host: %v", err)
}
}
func TestHostMatchesSubdomains(t *testing.T) {
pats := []string{"example.com", "*.news.org"}
for _, h := range []string{"example.com", "news.example.com", "a.b.example.com", "news.org", "feeds.news.org"} {
if !HostMatches(h, pats) {
t.Errorf("HostMatches(%q) = false, want true", h)
}
}
for _, h := range []string{"notexample.com", "example.com.evil.net", "org"} {
if HostMatches(h, pats) {
t.Errorf("HostMatches(%q) = true, want false", h)
}
}
}
func TestSizeCap(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Write([]byte(strings.Repeat("x", 5000)))
}))
defer srv.Close()
f := testFetcher(t, Config{MaxBytes: 100})
if _, err := f.Get(context.Background(), srv.URL); !errors.Is(err, ErrTooLarge) {
t.Fatalf("error = %v, want ErrTooLarge", err)
}
f = testFetcher(t, Config{MaxBytes: 6000})
resp, err := f.Get(context.Background(), srv.URL)
if err != nil {
t.Fatalf("under the cap: %v", err)
}
if len(resp.Body) != 5000 {
t.Fatalf("body = %d bytes, want 5000", len(resp.Body))
}
}
func TestRedirectCap(t *testing.T) {
var srv *httptest.Server
srv = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Redirect(w, r, srv.URL+"/again", http.StatusFound)
}))
defer srv.Close()
f := testFetcher(t, Config{MaxRedirects: 2})
if _, err := f.Get(context.Background(), srv.URL); !errors.Is(err, ErrRedirects) {
t.Fatalf("error = %v, want ErrRedirects", err)
}
}
// A redirect off the allowlist is the interesting redirect: the first hop is
// permitted, the second must not be.
func TestRedirectRecheckedAgainstDenylist(t *testing.T) {
target := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Write([]byte("secret"))
}))
defer target.Close()
hop := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Redirect(w, r, target.URL, http.StatusFound)
}))
defer hop.Close()
// Reach the hop under the name "localhost" and allow only that name; the
// redirect lands on the same box under its literal address, which the
// allowlist does not cover. Without the CheckRedirect hook this fetch
// succeeds and returns "secret".
f := testFetcher(t, Config{AllowHosts: []string{"localhost"}})
viaName := strings.Replace(hop.URL, "127.0.0.1", "localhost", 1)
if _, err := f.Get(context.Background(), viaName); !errors.Is(err, ErrBlocked) {
t.Fatalf("error = %v, want ErrBlocked", err)
}
}
func TestPerHostRateLimit(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Write([]byte("ok"))
}))
defer srv.Close()
f := testFetcher(t, Config{HostInterval: 60 * time.Millisecond})
start := time.Now()
for i := 0; i < 3; i++ {
if _, err := f.Get(context.Background(), srv.URL); err != nil {
t.Fatalf("request %d: %v", i, err)
}
}
if elapsed := time.Since(start); elapsed < 120*time.Millisecond {
t.Fatalf("three requests took %s, want at least 120ms of spacing", elapsed)
}
}
func TestRateLimitHonoursContext(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {}))
defer srv.Close()
f := testFetcher(t, Config{HostInterval: 10 * time.Second})
if _, err := f.Get(context.Background(), srv.URL); err != nil {
t.Fatalf("first request: %v", err)
}
ctx, cancel := context.WithTimeout(context.Background(), 20*time.Millisecond)
defer cancel()
if _, err := f.Get(ctx, srv.URL); !errors.Is(err, context.DeadlineExceeded) {
t.Fatalf("error = %v, want DeadlineExceeded", err)
}
}
func TestNon2xxIsAnError(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Error(w, "nope", http.StatusInternalServerError)
}))
defer srv.Close()
f := testFetcher(t, Config{})
if _, err := f.Get(context.Background(), srv.URL); !errors.Is(err, ErrStatus) {
t.Fatalf("error = %v, want ErrStatus", err)
}
}
func TestUserAgentIsSent(t *testing.T) {
got := make(chan string, 1)
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
got <- r.Header.Get("User-Agent")
}))
defer srv.Close()
f := testFetcher(t, Config{UserAgent: "Maven/test"})
if _, err := f.Get(context.Background(), srv.URL); err != nil {
t.Fatal(err)
}
if ua := <-got; ua != "Maven/test" {
t.Fatalf("user-agent = %q", ua)
}
}
+267
View File
@@ -0,0 +1,267 @@
// Package zenmoney reads spending and income from ZenMoney's /v8/diff/ API
// (Vikunja #125).
//
// Trust boundary: ZenMoney, not Maven. They already hold his bank sessions —
// this package only reads back what they have, over a token that lives in the
// poller module and is never handed to core. Nothing here writes to ZenMoney;
// diff is called read-only (an empty change set in, a change set out).
//
// Two rules the code exists to enforce:
//
// - NEVER invent a number. Every figure in a Summary is a sum of amounts the
// API returned. A request that fails, or returns nothing, produces no
// summary and therefore no fact — silence, not a zero. A confidently wrong
// "ты потратил 0" is worse than no answer.
// - His money is never search input. This package holds no notes, no
// utterances and no persona text, and it has no path to the external search
// capability. The only thing that leaves the box here is the diff request
// itself, to the service that already has the data.
package zenmoney
import (
"bytes"
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"sort"
"strings"
"time"
)
// DefaultBaseURL — ZenMoney's API root. Overridable so the tests can point at
// an httptest server replaying a recorded response.
const DefaultBaseURL = "https://api.zenmoney.ru"
// Client is a ZenMoney diff reader. The token is held here, in the poller's
// address space; core never receives it and never learns it exists.
type Client struct {
BaseURL string
Token string
HTTP *http.Client
}
// New returns a client with a bounded HTTP timeout. An empty token is a
// programming error the caller must catch — the capability is off unless
// configured, so a client is only ever built when a token was supplied.
func New(token, baseURL string, timeout time.Duration) (*Client, error) {
if strings.TrimSpace(token) == "" {
return nil, fmt.Errorf("zenmoney: empty token")
}
if baseURL == "" {
baseURL = DefaultBaseURL
}
if timeout <= 0 {
timeout = 20 * time.Second
}
return &Client{
BaseURL: strings.TrimRight(baseURL, "/"),
Token: token,
HTTP: &http.Client{Timeout: timeout},
}, nil
}
// diffRequest — the smallest body /v8/diff/ accepts. serverTimestamp is the
// incremental cursor: the server returns objects changed at or after it.
type diffRequest struct {
CurrentClientTimestamp int64 `json:"currentClientTimestamp"`
ServerTimestamp int64 `json:"serverTimestamp"`
}
// diffResponse — only the fields spending needs. ZenMoney returns a dozen more
// object types (tags, merchants, budgets, reminders); decoding them would mean
// holding more of his financial life in memory than the question needs.
type diffResponse struct {
ServerTimestamp int64 `json:"serverTimestamp"`
Instrument []instrument `json:"instrument"`
Transaction []transaction `json:"transaction"`
}
type instrument struct {
ID int64 `json:"id"`
ShortTitle string `json:"shortTitle"`
}
type transaction struct {
ID string `json:"id"`
Date string `json:"date"` // "2026-07-15"
Deleted bool `json:"deleted"`
Income float64 `json:"income"`
Outcome float64 `json:"outcome"`
IncomeInstrument int64 `json:"incomeInstrument"`
OutcomeInstrmnt int64 `json:"outcomeInstrument"`
IncomeAccount string `json:"incomeAccount"`
OutcomeAccount string `json:"outcomeAccount"`
}
// Money — an amount in one currency. Kept as the currency's own short title
// ("RUB", "EUR") rather than converted: ZenMoney's rates are a snapshot, and
// converting would turn a figure he can check against his bank into one he
// cannot.
type Money struct {
Currency string `json:"currency"`
Amount float64 `json:"amount"`
}
// Summary — what was spent and earned over a window, per currency, plus how
// many transactions it was computed from. Count is the honesty check: a
// summary built from zero transactions is not "you spent nothing", it is "there
// was nothing to read", and callers treat it as no answer.
type Summary struct {
From, To time.Time
Spent []Money `json:"spent"`
Earned []Money `json:"earned"`
Count int `json:"count"`
// ServerTimestamp — the cursor the API returned, for the caller to log or
// carry. Not used as an incremental cursor for summaries; see Since.
ServerTimestamp int64 `json:"-"`
}
// Since returns the summary of transactions dated in [from, to).
//
// The diff cursor is set to `from` so the server only sends objects changed
// since then, which for a "this month" window is everything filed this month.
// The caveat, deliberately accepted: a transaction he EDITED this month but
// dated last month arrives too, and is then excluded by date — so editing old
// records cannot inflate this month's total. The reverse case (a transaction
// dated this month, filed and last changed before `from`) cannot exist.
func (c *Client) Since(ctx context.Context, from, to time.Time) (Summary, error) {
resp, err := c.diff(ctx, from.Unix())
if err != nil {
return Summary{}, err
}
return summarize(resp, from, to), nil
}
func (c *Client) diff(ctx context.Context, serverTimestamp int64) (diffResponse, error) {
body, err := json.Marshal(diffRequest{
CurrentClientTimestamp: time.Now().Unix(),
ServerTimestamp: serverTimestamp,
})
if err != nil {
return diffResponse{}, err
}
req, err := http.NewRequestWithContext(ctx, http.MethodPost, c.BaseURL+"/v8/diff/", bytes.NewReader(body))
if err != nil {
return diffResponse{}, err
}
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Authorization", "Bearer "+c.Token)
hc := c.HTTP
if hc == nil {
hc = &http.Client{Timeout: 20 * time.Second}
}
res, err := hc.Do(req)
if err != nil {
return diffResponse{}, err
}
defer res.Body.Close()
raw, err := io.ReadAll(io.LimitReader(res.Body, 32<<20))
if err != nil {
return diffResponse{}, err
}
if res.StatusCode != http.StatusOK {
// The status only. The body of a failed diff can echo account data, and
// this string reaches the log.
return diffResponse{}, fmt.Errorf("zenmoney diff: %s", res.Status)
}
var out diffResponse
if err := json.Unmarshal(raw, &out); err != nil {
return diffResponse{}, fmt.Errorf("zenmoney diff: decode: %w", err)
}
return out, nil
}
// summarize sums the transactions dated inside the window.
//
// Excluded, in order: deleted rows (ZenMoney tombstones rather than removes),
// transfers and currency exchanges (income and outcome both non-zero — moving
// his own money between his own accounts is not spending), and anything dated
// outside the window.
func summarize(resp diffResponse, from, to time.Time) Summary {
cur := map[int64]string{}
for _, in := range resp.Instrument {
cur[in.ID] = in.ShortTitle
}
spent := map[string]float64{}
earned := map[string]float64{}
count := 0
for _, t := range resp.Transaction {
if t.Deleted {
continue
}
d, err := time.ParseInLocation("2006-01-02", t.Date, from.Location())
if err != nil {
continue // an undated row is not a number we can place
}
if d.Before(from) || !d.Before(to) {
continue
}
if t.Income > 0 && t.Outcome > 0 {
continue // transfer / exchange
}
switch {
case t.Outcome > 0:
spent[currency(cur, t.OutcomeInstrmnt)] += t.Outcome
count++
case t.Income > 0:
earned[currency(cur, t.IncomeInstrument)] += t.Income
count++
}
}
return Summary{
From: from, To: to,
Spent: sortMoney(spent), Earned: sortMoney(earned),
Count: count, ServerTimestamp: resp.ServerTimestamp,
}
}
// currency names the instrument, or says it does not know. An unknown id keeps
// the amount rather than dropping it: a sum without a currency label is still
// his money, and silently discarding it would understate the total.
func currency(names map[int64]string, id int64) string {
if s := names[id]; s != "" {
return s
}
return "?"
}
// sortMoney gives the amounts a stable order (largest first) so the rendered
// string and the written fact do not churn between polls.
func sortMoney(m map[string]float64) []Money {
out := make([]Money, 0, len(m))
for c, a := range m {
out = append(out, Money{Currency: c, Amount: a})
}
sort.Slice(out, func(i, j int) bool {
if out[i].Amount != out[j].Amount {
return out[i].Amount > out[j].Amount
}
return out[i].Currency < out[j].Currency
})
return out
}
// Empty reports whether the summary rests on no transactions at all. Callers
// must treat an empty summary as "nothing to say", never as a zero: the
// difference between "he spent nothing" and "the read returned nothing" is the
// difference between an answer and an invented one.
func (s Summary) Empty() bool { return s.Count == 0 }
// MonthWindow — the first instant of now's month, and now's own day-end
// exclusive bound, in now's location. The window a "сколько я потратил в этом
// месяце?" question means.
func MonthWindow(now time.Time) (from, to time.Time) {
loc := now.Location()
from = time.Date(now.Year(), now.Month(), 1, 0, 0, 0, 0, loc)
to = time.Date(now.Year(), now.Month(), now.Day(), 0, 0, 0, 0, loc).AddDate(0, 0, 1)
return from, to
}
// DayWindow — today, in now's location.
func DayWindow(now time.Time) (from, to time.Time) {
loc := now.Location()
from = time.Date(now.Year(), now.Month(), now.Day(), 0, 0, 0, 0, loc)
return from, from.AddDate(0, 0, 1)
}
+178
View File
@@ -0,0 +1,178 @@
package zenmoney
import (
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"os"
"strings"
"testing"
"time"
)
// fixtureServer replays testdata/diff.json and records the request, so the
// tests can assert the wire contract (Bearer token, POST, /v8/diff/) without a
// ZenMoney account.
func fixtureServer(t *testing.T, got *diffRequest, auth *string) *httptest.Server {
t.Helper()
body, err := os.ReadFile("testdata/diff.json")
if err != nil {
t.Fatal(err)
}
return httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.Method != http.MethodPost {
t.Errorf("method = %s, want POST", r.Method)
}
if r.URL.Path != "/v8/diff/" {
t.Errorf("path = %s, want /v8/diff/", r.URL.Path)
}
if auth != nil {
*auth = r.Header.Get("Authorization")
}
if got != nil {
if err := json.NewDecoder(r.Body).Decode(got); err != nil {
t.Errorf("decode request: %v", err)
}
}
w.Header().Set("Content-Type", "application/json")
w.Write(body)
}))
}
func aug(day int) time.Time { return time.Date(2026, 8, day, 0, 0, 0, 0, time.UTC) }
func TestSinceSumsSpendingPerCurrency(t *testing.T) {
var req diffRequest
var auth string
srv := fixtureServer(t, &req, &auth)
defer srv.Close()
c, err := New("tok", srv.URL, time.Second)
if err != nil {
t.Fatal(err)
}
s, err := c.Since(context.Background(), aug(1), aug(6))
if err != nil {
t.Fatal(err)
}
if auth != "Bearer tok" {
t.Errorf("Authorization = %q", auth)
}
if req.ServerTimestamp != aug(1).Unix() {
t.Errorf("serverTimestamp = %d, want the window start", req.ServerTimestamp)
}
// 1500 + 249.5 RUB spent, 12 EUR spent, 3000 RUB in. The transfer (t4), the
// deleted row (t6) and July's salary (t3) are all excluded.
want := map[string]float64{"RUB": 1749.5, "EUR": 12}
if len(s.Spent) != 2 {
t.Fatalf("spent = %+v, want two currencies", s.Spent)
}
for _, m := range s.Spent {
if want[m.Currency] != m.Amount {
t.Errorf("spent %s = %v, want %v", m.Currency, m.Amount, want[m.Currency])
}
}
if len(s.Earned) != 1 || s.Earned[0].Amount != 3000 || s.Earned[0].Currency != "RUB" {
t.Errorf("earned = %+v, want 3000 RUB (July's salary is outside the window)", s.Earned)
}
if s.Count != 4 {
t.Errorf("count = %d, want 4 counted transactions", s.Count)
}
// Largest first, so the fact value does not churn between polls.
if s.Spent[0].Currency != "RUB" {
t.Errorf("spent order = %+v, want the largest amount first", s.Spent)
}
}
// A window with nothing in it is NOT a zero. No transactions means no answer,
// and the caller must be able to tell the difference.
func TestSinceEmptyWindowIsNotAZero(t *testing.T) {
srv := fixtureServer(t, nil, nil)
defer srv.Close()
c, _ := New("tok", srv.URL, time.Second)
s, err := c.Since(context.Background(), time.Date(2026, 9, 1, 0, 0, 0, 0, time.UTC), time.Date(2026, 9, 30, 0, 0, 0, 0, time.UTC))
if err != nil {
t.Fatal(err)
}
if !s.Empty() {
t.Fatalf("summary = %+v, want empty", s)
}
if _, ok := s.Value(); ok {
t.Error("an empty summary must not produce a fact value")
}
}
func TestSinceReportsHTTPFailureWithoutTheBody(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Error(w, `{"account":"acc-card","secret":"leaky"}`, http.StatusUnauthorized)
}))
defer srv.Close()
c, _ := New("tok", srv.URL, time.Second)
_, err := c.Since(context.Background(), aug(1), aug(6))
if err == nil {
t.Fatal("want an error on 401")
}
if strings.Contains(err.Error(), "acc-card") || strings.Contains(err.Error(), "leaky") {
t.Errorf("error %q echoes the response body — it reaches the log", err)
}
}
func TestNewRequiresAToken(t *testing.T) {
if _, err := New(" ", "", 0); err == nil {
t.Error("want an error for an empty token — the capability is off unless configured")
}
}
func TestMonthAndDayWindows(t *testing.T) {
now := time.Date(2026, 8, 15, 21, 30, 0, 0, time.UTC)
from, to := MonthWindow(now)
if from != time.Date(2026, 8, 1, 0, 0, 0, 0, time.UTC) || to != time.Date(2026, 8, 16, 0, 0, 0, 0, time.UTC) {
t.Errorf("month window = %v..%v", from, to)
}
from, to = DayWindow(now)
if from != time.Date(2026, 8, 15, 0, 0, 0, 0, time.UTC) || to != time.Date(2026, 8, 16, 0, 0, 0, 0, time.UTC) {
t.Errorf("day window = %v..%v", from, to)
}
}
func TestFactValueRoundTripAndFormat(t *testing.T) {
s := Summary{Spent: []Money{{"RUB", 1749.5}}, Earned: []Money{{"RUB", 3000}}, Count: 3}
raw, ok := s.Value()
if !ok {
t.Fatal("want a fact value")
}
v, err := ParseFactValue(raw)
if err != nil {
t.Fatal(err)
}
got := v.FormatRU("в этом месяце")
if !strings.Contains(got, "1749.5 RUB") || !strings.Contains(got, "3000 RUB") {
t.Errorf("reply = %q, want the exact figures", got)
}
// Persona: informal, feminine, no commentary on his spending.
for _, bad := range []string{"вы", "ваш", "милый", "дорогой", "рад ", "слишком", "много"} {
if strings.Contains(got, bad) {
t.Errorf("reply %q contains %q", got, bad)
}
}
if strings.Contains(got, "он ") {
t.Errorf("reply %q talks about him in the third person", got)
}
}
// An empty fact value renders to nothing, so a caller cannot accidentally
// speak a zero.
func TestFormatRUEmptyRendersNothing(t *testing.T) {
if got := (FactValue{}).FormatRU("сегодня"); got != "" {
t.Errorf("reply = %q, want empty", got)
}
}
func TestFormatAmountKeepsTheTruth(t *testing.T) {
for in, want := range map[float64]string{1500: "1500", 249.5: "249.5", 0.99: "0.99", 1749.55: "1749.55"} {
if got := formatAmount(in); got != want {
t.Errorf("formatAmount(%v) = %q, want %q", in, got, want)
}
}
}
+99
View File
@@ -0,0 +1,99 @@
package zenmoney
import (
"encoding/json"
"fmt"
"strings"
"time"
)
// Fact keys the poller writes, all under source "poll:zenmoney". Two windows,
// because they are the two questions he actually asks; a per-category
// breakdown would mean storing what he bought, and the store is not a ledger.
const (
KeySpentToday = "money_today"
KeySpentMonth = "money_month"
)
// Source — the provenance every money fact carries. The loop's rules trust
// source, and nothing in Maven has a rule on these keys: they are read when he
// asks, never a reason to speak. Maven is not a nag, least of all about money.
const Source = "poll:zenmoney"
// FactValue — the JSON stored in a money fact. A wire shape of its own rather
// than the Summary struct so From/To (which carry a timezone and a clock) stay
// out of the store; the key already says which window it is.
type FactValue struct {
Spent []Money `json:"spent"`
Earned []Money `json:"earned"`
Count int `json:"count"`
}
// Value encodes the summary for the facts table. Returns ok=false for an empty
// summary: no transactions read means no fact written, so that a failed or
// empty poll can never be recited back to him as a zero.
func (s Summary) Value() (string, bool) {
if s.Empty() {
return "", false
}
b, err := json.Marshal(FactValue{Spent: s.Spent, Earned: s.Earned, Count: s.Count})
if err != nil {
return "", false
}
return string(b), true
}
// ParseFactValue decodes a stored money fact.
func ParseFactValue(raw string) (FactValue, error) {
var v FactValue
if err := json.Unmarshal([]byte(raw), &v); err != nil {
return FactValue{}, err
}
return v, nil
}
// FormatRU renders a money fact the way Maven says it — feminine, informal,
// and only about numbers that came from ZenMoney. window is the Russian phrase
// for the period ("сегодня", "в этом месяце").
//
// No commentary. She reports the figure and stops: an opinion about his
// spending is exactly the nagging Maven is not for.
func (v FactValue) FormatRU(window string) string {
if v.Count == 0 {
return ""
}
var parts []string
if len(v.Spent) > 0 {
parts = append(parts, "потратил "+joinMoney(v.Spent))
}
if len(v.Earned) > 0 {
parts = append(parts, "получил "+joinMoney(v.Earned))
}
if len(parts) == 0 {
return ""
}
return window + " ты " + strings.Join(parts, ", ") + "."
}
func joinMoney(ms []Money) string {
parts := make([]string, 0, len(ms))
for _, m := range ms {
parts = append(parts, fmt.Sprintf("%s %s", formatAmount(m.Amount), m.Currency))
}
return strings.Join(parts, " и ")
}
// formatAmount — whole units when the amount is whole, two decimals otherwise.
// Never rounded to something prettier than the truth.
func formatAmount(a float64) string {
if a == float64(int64(a)) {
return fmt.Sprintf("%d", int64(a))
}
return strings.TrimRight(strings.TrimRight(fmt.Sprintf("%.2f", a), "0"), ".")
}
// StaleAfter — how old a money fact may be and still be worth reciting. The
// poller is off unless configured and can be down; answering with last week's
// total as if it were today's would be a lie by omission, so a stale fact is
// reported as stale.
const StaleAfter = 26 * time.Hour
+35
View File
@@ -0,0 +1,35 @@
{
"serverTimestamp": 1785312000,
"instrument": [
{"id": 2, "title": "Российский рубль", "shortTitle": "RUB", "symbol": "₽", "rate": 1},
{"id": 3, "title": "Евро", "shortTitle": "EUR", "symbol": "€", "rate": 100}
],
"account": [
{"id": "acc-card", "title": "карта", "instrument": 2},
{"id": "acc-cash", "title": "наличные", "instrument": 2},
{"id": "acc-eur", "title": "евро", "instrument": 3}
],
"transaction": [
{"id": "t1", "date": "2026-08-01", "changed": 1785300000, "income": 0, "outcome": 1500,
"incomeInstrument": 2, "outcomeInstrument": 2, "incomeAccount": "acc-card", "outcomeAccount": "acc-card",
"payee": "пятёрочка", "deleted": false},
{"id": "t2", "date": "2026-08-01", "changed": 1785300001, "income": 0, "outcome": 249.5,
"incomeInstrument": 2, "outcomeInstrument": 2, "incomeAccount": "acc-card", "outcomeAccount": "acc-card",
"payee": "метро", "deleted": false},
{"id": "t3", "date": "2026-07-20", "changed": 1785300002, "income": 120000, "outcome": 0,
"incomeInstrument": 2, "outcomeInstrument": 2, "incomeAccount": "acc-card", "outcomeAccount": "acc-card",
"payee": "зарплата", "deleted": false},
{"id": "t4", "date": "2026-08-02", "changed": 1785300003, "income": 5000, "outcome": 5000,
"incomeInstrument": 2, "outcomeInstrument": 2, "incomeAccount": "acc-cash", "outcomeAccount": "acc-card",
"payee": "", "comment": "снял наличные", "deleted": false},
{"id": "t5", "date": "2026-08-03", "changed": 1785300004, "income": 0, "outcome": 12,
"incomeInstrument": 3, "outcomeInstrument": 3, "incomeAccount": "acc-eur", "outcomeAccount": "acc-eur",
"payee": "hosting", "deleted": false},
{"id": "t6", "date": "2026-08-04", "changed": 1785300005, "income": 0, "outcome": 999,
"incomeInstrument": 2, "outcomeInstrument": 2, "incomeAccount": "acc-card", "outcomeAccount": "acc-card",
"payee": "удалённая", "deleted": true},
{"id": "t7", "date": "2026-08-05", "changed": 1785300006, "income": 3000, "outcome": 0,
"incomeInstrument": 2, "outcomeInstrument": 2, "incomeAccount": "acc-card", "outcomeAccount": "acc-card",
"payee": "возврат", "deleted": false}
]
}