crawl: stop letting a watch widen on-demand reading, and honour Crawl-delay

The on-demand crawler was built over allow_hosts plus every watched host.
webfetch reads a non-empty allow list as these and nothing else, so a config
with one watch and no allow_hosts at all silently narrowed on-demand reading
to the watched site. Every other url he pasted came back as a flat refusal
with nothing in the log to explain it. The two crawlers now take two host
lists from one crawlHosts helper.

Crawl-delay was parsed into Rules and never read. The only pacing was the
fetcher's flat one request per host per second, which cannot express what a
site asked for, and deploy/README claimed the field was honoured. Page now
waits it out between the robots fetch and the page fetch, and a delay longer
than the turn fails the read instead of hanging it.

A robots.txt that failed was treated as no rules, so a site whose server was
having a bad minute became a site with no restrictions. A 5xx now refuses the
crawl. A 404 still means unrestricted, which is what the standard says.

The refusal check matched substrings of webfetch's message text from a package
that cannot import webfetch, so a reworded error would have silently turned
into a robots verdict. internal/crawl now exports ErrFetchRefused and
ErrFetchStatus and the adapter in cmd/mavend maps the webfetch sentinels onto
them. Robots group selection picks the longest matching agent prefix instead
of the first one in file order.

queryWeb passed a claim it could not serve when no crawler was configured, so
an unconfigured deployment answered a web question with an apology instead of
falling through to the model.

Found in review of #67.
This commit is contained in:
kami
2026-08-01 14:33:26 +04:00
parent 57161fb762
commit 327726a06a
8 changed files with 384 additions and 79 deletions
+12 -8
View File
@@ -92,10 +92,12 @@ var querySources = []querySource{
{"memory", (*reactiveHandler).queryMemory},
{"notes", (*reactiveHandler).queryNotes},
// LAST before the model answers from memory, and that position is the whole
// design (Vikunja #259): local sources first. The model, his own notes and
// facts, and — once internal/kiwix is wired into this chain — the offline
// ZIMs all get their turn before anything touches the network. This source
// only claims a turn where he named a URL out loud, so it never competes
// design (Vikunja #259): local sources first. His memory, his notes and
// once internal/kiwix is wired into this chain — the offline ZIMs all get
// their turn before anything touches the network. The model does NOT: it
// answers after this, because a URL he said out loud is an instruction and
// a 1.7B guessing at a page it cannot read is how contents get invented.
// This source only claims a turn where he named a URL, so it never competes
// with a local answer.
{"web", (*reactiveHandler).queryWeb},
{"general-knowledge", (*reactiveHandler).queryGeneral},
@@ -444,10 +446,12 @@ func (h *reactiveHandler) queryWeb(ctx context.Context, t *queryTurn) (string, b
return "", false
}
if h.crawler == nil {
// Claim rather than fall through: he asked about a specific page, and
// letting the model answer from the URL's spelling alone is how a small
// model invents a page's contents.
return "я не читаю страницы — это не настроено.", true
// Fall through. Reading pages is off unless configured, and on a daemon
// where it was never turned on the older behaviour is right: the model
// answers the question as if the URL had not been said. Announcing a
// configuration status is for a capability that exists and failed, not
// for one he never asked for.
return "", false
}
ctxFetch, cancel := context.WithTimeout(ctx, 30*time.Second)
defer cancel()