Read a web page when he names one, and watch a few on a timer (#259)

The network fallback behind the local sources, off unless configured.

internal/crawl is pure: a stdlib robots.txt parser (group specificity,
wildcards, Crawl-delay, cached per host), HTML-to-plaintext extraction, and a
watcher that notes a watched page only when its text changed. It has no store
access and no net/http; cmd/mavend/crawls.go is the impure half.

Every limit is code and tested: the guarded fetcher from #258 enforces the host
allowlist/denylist, refuses private addresses in the dialer Control hook (so DNS
rebinding and each redirect hop are covered), caps size and redirects, times out,
and spaces requests per host. A robots.txt Disallow is refused with no override.

On demand, reading is a query source placed last in the chain, after his memory,
his notes, and the local Kiwix ZIMs once those are wired: no URL in the
utterance means no fetch, and only the URL ever leaves the box. Scheduled
watches write notes and announce nothing.

The vendored tree has no x/net/html, goquery or temoto/robotstxt, so the parsers
are stdlib. No new dependency.
This commit is contained in:
kami
2026-08-01 03:40:21 +04:00
parent cb3641e7bb
commit 2c1b0eede0
18 changed files with 1696 additions and 11 deletions
+30
View File
@@ -67,6 +67,36 @@ What it does and does not do:
- how far each feed was read is stored as a config fact `rss:latest:<name>`, so
a restart does not re-note yesterday's headlines.
### Reading a page (`crawl`, also off by default)
There is no `crawl` block either, so no page is fetched. Two halves, separately
switched:
```json
"crawl": {
"on_demand": true,
"interval": "6h",
"max_runes": 4000,
"watches": [
{ "name": "changelog", "url": "https://example.org/changelog", "interval": "12h" }
]
}
```
- `on_demand` lets her read a page he names in the utterance: "посмотри
https://example.org/x — что там?". The page becomes context for his question,
and only the URL leaves the box. Without a URL nothing is fetched, so this is
a fallback and not a habit;
- `watches` re-reads a fixed list on its interval and writes a note when the
text changed. Like the feeds, it announces nothing;
- the answer path sits **last** in the query chain, behind his memory, his notes
and (once wired) the local Kiwix ZIMs. A local read costs nothing;
- `robots.txt` is fetched first and obeyed with no override; a `Disallow` is a
refusal she says out loud. `Crawl-delay` is honoured;
- same guarded fetcher as the feeds: allowlist/denylist, no private addresses,
size cap, redirect cap, timeout, one request per host per second;
- dedup state is the config fact `crawl:hash:<name>`.
## Not yet verified / host-dependent
This stack is correct-by-construction but has **not been build-tested here**