Read a web page when he names one, and watch a few on a timer (#259)

The network fallback behind the local sources, off unless configured.

internal/crawl is pure: a stdlib robots.txt parser (group specificity,
wildcards, Crawl-delay, cached per host), HTML-to-plaintext extraction, and a
watcher that notes a watched page only when its text changed. It has no store
access and no net/http; cmd/mavend/crawls.go is the impure half.

Every limit is code and tested: the guarded fetcher from #258 enforces the host
allowlist/denylist, refuses private addresses in the dialer Control hook (so DNS
rebinding and each redirect hop are covered), caps size and redirects, times out,
and spaces requests per host. A robots.txt Disallow is refused with no override.

On demand, reading is a query source placed last in the chain, after his memory,
his notes, and the local Kiwix ZIMs once those are wired: no URL in the
utterance means no fetch, and only the URL ever leaves the box. Scheduled
watches write notes and announce nothing.

The vendored tree has no x/net/html, goquery or temoto/robotstxt, so the parsers
are stdlib. No new dependency.
This commit is contained in:
kami
2026-08-01 03:40:21 +04:00
parent cb3641e7bb
commit 2c1b0eede0
18 changed files with 1696 additions and 11 deletions
+38
View File
@@ -28,3 +28,41 @@
7. Add IPC methods `MethodTriggerCrawl(name)`, `MethodListCrawls`, `MethodGetCrawlResult(name)`
8. Add `crawls` block to `config.Config` and `deploy/mavend.json`
9. Test with a static HTML page — verify extraction matches expected values, verify scheduling fires correctly
## Shipped 2026-08-01 (#259)
Built as `internal/crawl` (pure: robots, extraction, watcher) plus
`cmd/mavend/crawls.go` (fetcher, ticker, dedup facts), on top of the guarded
`internal/webfetch` door added with the feed reader (#258). Off unless
configured, in two separately-switched halves: `crawl.on_demand` for a URL he
names, `crawl.watches` for a scheduled re-read.
**Limits are code, not documentation** (`internal/webfetch`, tested one test per
limit): host allowlist/denylist, no private addresses (loopback, RFC1918 —
hence the LAN and the `10.42.0.0/24` wg range —, link-local incl. cloud
metadata, CGNAT, v6 ULA) enforced in the dialer's `Control` hook so DNS
rebinding and every redirect hop are covered, response size cap, redirect cap,
timeout, one request per host per second. `robots.txt` is fetched first, cached
per host, and a `Disallow` is refused with no override.
Deliberate deviations from the plan above:
- **No CSS selectors and no LLM structured extraction** (steps 2). The output is
plaintext handed to the phraser as context for the question he asked. A 1.7B
extracting a JSON price table from 4000 runes is a worse bet than reading, and
`goquery` is not vendored.
- **No `crawl` act verb and no new IPC methods** (steps 5, 7). Reading a page is
a query source (`queryWeb` in `actions_query.go`, last in the chain, behind
Kiwix once that is wired), not an action he commands. Nothing needs a new wire
method to work.
- **Notes, not facts.** A page's text is not a fact about him. Only the dedup
hash is a fact (`crawl:hash:<name>`, kind `config`, source `poll:crawl`).
- **Nothing is dispatched.** A changed page writes a note; it does not nudge.
Not a nag.
- **No `/tools` crawl history page.** The notes and the hash facts are already
visible on `/dash`.
**No new dependency.** The vendored tree has no `x/net/html`, no `goquery` and
no `temoto/robotstxt`, so robots parsing and HTML-to-text are stdlib
(`regexp`, `html`) — RE2 has no backreferences, hence the `pairsRE` builder in
`extract.go`.