Read a web page when he names one, and watch a few on a timer (#259)
The network fallback behind the local sources, off unless configured. internal/crawl is pure: a stdlib robots.txt parser (group specificity, wildcards, Crawl-delay, cached per host), HTML-to-plaintext extraction, and a watcher that notes a watched page only when its text changed. It has no store access and no net/http; cmd/mavend/crawls.go is the impure half. Every limit is code and tested: the guarded fetcher from #258 enforces the host allowlist/denylist, refuses private addresses in the dialer Control hook (so DNS rebinding and each redirect hop are covered), caps size and redirects, times out, and spaces requests per host. A robots.txt Disallow is refused with no override. On demand, reading is a query source placed last in the chain, after his memory, his notes, and the local Kiwix ZIMs once those are wired: no URL in the utterance means no fetch, and only the URL ever leaves the box. Scheduled watches write notes and announce nothing. The vendored tree has no x/net/html, goquery or temoto/robotstxt, so the parsers are stdlib. No new dependency.
This commit is contained in:
@@ -52,6 +52,7 @@ import (
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/audio"
|
||||
"github.com/kami/maven/internal/crawl"
|
||||
"github.com/kami/maven/internal/dialogue"
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/memory"
|
||||
@@ -82,6 +83,10 @@ type reactiveHandler struct {
|
||||
replier voice.Replier
|
||||
now func() time.Time
|
||||
|
||||
// crawler reads a web page he names out loud (queryWeb). nil ⇒ on-demand
|
||||
// page reading is off, which is the default: no `crawl` block, no fetch.
|
||||
crawler *crawl.Crawler
|
||||
|
||||
// feedsOn — whether any RSS feed is configured (config.Feeds). It changes
|
||||
// only what she SAYS when asked and nothing is there: "ленты не настроены"
|
||||
// instead of "ничего нового", which are different truths.
|
||||
|
||||
Reference in New Issue
Block a user