Read a web page when he names one, and watch a few on a timer (#259)

The network fallback behind the local sources, off unless configured.

internal/crawl is pure: a stdlib robots.txt parser (group specificity,
wildcards, Crawl-delay, cached per host), HTML-to-plaintext extraction, and a
watcher that notes a watched page only when its text changed. It has no store
access and no net/http; cmd/mavend/crawls.go is the impure half.

Every limit is code and tested: the guarded fetcher from #258 enforces the host
allowlist/denylist, refuses private addresses in the dialer Control hook (so DNS
rebinding and each redirect hop are covered), caps size and redirects, times out,
and spaces requests per host. A robots.txt Disallow is refused with no override.

On demand, reading is a query source placed last in the chain, after his memory,
his notes, and the local Kiwix ZIMs once those are wired: no URL in the
utterance means no fetch, and only the URL ever leaves the box. Scheduled
watches write notes and announce nothing.

The vendored tree has no x/net/html, goquery or temoto/robotstxt, so the parsers
are stdlib. No new dependency.
This commit is contained in:
kami
2026-08-01 03:40:21 +04:00
parent cb3641e7bb
commit 2c1b0eede0
18 changed files with 1696 additions and 11 deletions
+5
View File
@@ -52,6 +52,7 @@ import (
"time"
"github.com/kami/maven/internal/audio"
"github.com/kami/maven/internal/crawl"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/memory"
@@ -82,6 +83,10 @@ type reactiveHandler struct {
replier voice.Replier
now func() time.Time
// crawler reads a web page he names out loud (queryWeb). nil ⇒ on-demand
// page reading is off, which is the default: no `crawl` block, no fetch.
crawler *crawl.Crawler
// feedsOn — whether any RSS feed is configured (config.Feeds). It changes
// only what she SAYS when asked and nothing is there: "ленты не настроены"
// instead of "ничего нового", which are different truths.