Read a web page when he names one, and watch a few on a timer (#259)
The network fallback behind the local sources, off unless configured. internal/crawl is pure: a stdlib robots.txt parser (group specificity, wildcards, Crawl-delay, cached per host), HTML-to-plaintext extraction, and a watcher that notes a watched page only when its text changed. It has no store access and no net/http; cmd/mavend/crawls.go is the impure half. Every limit is code and tested: the guarded fetcher from #258 enforces the host allowlist/denylist, refuses private addresses in the dialer Control hook (so DNS rebinding and each redirect hop are covered), caps size and redirects, times out, and spaces requests per host. A robots.txt Disallow is refused with no override. On demand, reading is a query source placed last in the chain, after his memory, his notes, and the local Kiwix ZIMs once those are wired: no URL in the utterance means no fetch, and only the URL ever leaves the box. Scheduled watches write notes and announce nothing. The vendored tree has no x/net/html, goquery or temoto/robotstxt, so the parsers are stdlib. No new dependency.
This commit is contained in:
@@ -67,6 +67,36 @@ What it does and does not do:
|
||||
- how far each feed was read is stored as a config fact `rss:latest:<name>`, so
|
||||
a restart does not re-note yesterday's headlines.
|
||||
|
||||
### Reading a page (`crawl`, also off by default)
|
||||
|
||||
There is no `crawl` block either, so no page is fetched. Two halves, separately
|
||||
switched:
|
||||
|
||||
```json
|
||||
"crawl": {
|
||||
"on_demand": true,
|
||||
"interval": "6h",
|
||||
"max_runes": 4000,
|
||||
"watches": [
|
||||
{ "name": "changelog", "url": "https://example.org/changelog", "interval": "12h" }
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
- `on_demand` lets her read a page he names in the utterance: "посмотри
|
||||
https://example.org/x — что там?". The page becomes context for his question,
|
||||
and only the URL leaves the box. Without a URL nothing is fetched, so this is
|
||||
a fallback and not a habit;
|
||||
- `watches` re-reads a fixed list on its interval and writes a note when the
|
||||
text changed. Like the feeds, it announces nothing;
|
||||
- the answer path sits **last** in the query chain, behind his memory, his notes
|
||||
and (once wired) the local Kiwix ZIMs. A local read costs nothing;
|
||||
- `robots.txt` is fetched first and obeyed with no override; a `Disallow` is a
|
||||
refusal she says out loud. `Crawl-delay` is honoured;
|
||||
- same guarded fetcher as the feeds: allowlist/denylist, no private addresses,
|
||||
size cap, redirect cap, timeout, one request per host per second;
|
||||
- dedup state is the config fact `crawl:hash:<name>`.
|
||||
|
||||
## Not yet verified / host-dependent
|
||||
|
||||
This stack is correct-by-construction but has **not been build-tested here**
|
||||
|
||||
Reference in New Issue
Block a user