crawl: stop letting a watch widen on-demand reading, and honour Crawl-delay

The on-demand crawler was built over allow_hosts plus every watched host.
webfetch reads a non-empty allow list as these and nothing else, so a config
with one watch and no allow_hosts at all silently narrowed on-demand reading
to the watched site. Every other url he pasted came back as a flat refusal
with nothing in the log to explain it. The two crawlers now take two host
lists from one crawlHosts helper.

Crawl-delay was parsed into Rules and never read. The only pacing was the
fetcher's flat one request per host per second, which cannot express what a
site asked for, and deploy/README claimed the field was honoured. Page now
waits it out between the robots fetch and the page fetch, and a delay longer
than the turn fails the read instead of hanging it.

A robots.txt that failed was treated as no rules, so a site whose server was
having a bad minute became a site with no restrictions. A 5xx now refuses the
crawl. A 404 still means unrestricted, which is what the standard says.

The refusal check matched substrings of webfetch's message text from a package
that cannot import webfetch, so a reworded error would have silently turned
into a robots verdict. internal/crawl now exports ErrFetchRefused and
ErrFetchStatus and the adapter in cmd/mavend maps the webfetch sentinels onto
them. Robots group selection picks the longest matching agent prefix instead
of the first one in file order.

queryWeb passed a claim it could not serve when no crawler was configured, so
an unconfigured deployment answered a web question with an apology instead of
falling through to the model.

Found in review of #67.
This commit is contained in:
kami
2026-08-01 14:33:26 +04:00
parent 57161fb762
commit 327726a06a
8 changed files with 384 additions and 79 deletions
+44 -9
View File
@@ -20,6 +20,8 @@ package main
import (
"context"
"errors"
"fmt"
"log"
"net/url"
"time"
@@ -39,21 +41,36 @@ func newCrawler(cfg *config.Config) *crawl.Crawler {
return nil
}
cc := cfg.Crawl
// The WATCH crawler, and only it, reaches the watched hosts. webfetch reads
// a non-empty allow list as "these and nothing else", so folding the watch
// hosts in turned a single watch into an allowlist for everything: a config
// with one watch and on_demand true silently refused every other page he
// pasted, with "не получилось прочитать страницу." and no clue why.
return crawlerWithHosts(cc, crawlHosts(cc, true))
}
// crawlHosts — the allowlist for one of the two crawlers. forWatches adds the
// watched pages' own hosts, so a watch does not have to be allowlisted by hand.
//
// The on-demand crawler gets his allow_hosts and nothing else. webfetch reads a
// non-empty list as "these and nothing else", so adding the watch hosts there
// would silently narrow on-demand reading to the watched sites.
func crawlHosts(cc *config.CrawlConfig, forWatches bool) []string {
hosts := append([]string(nil), cc.AllowHosts...)
// A watched page's own host is always reachable; otherwise an allowlist and
// a watch list would have to be kept in sync by hand.
if !forWatches {
return hosts
}
for _, w := range cc.Watches {
if u, err := url.Parse(w.URL); err == nil && u.Hostname() != "" {
hosts = append(hosts, u.Hostname())
}
}
// An allowlist plus on-demand is a contradiction worth logging rather than
// silently resolving: he asked for arbitrary pages AND for a fixed list.
// The allowlist wins, because it is the narrower instruction.
if len(hosts) > 0 && cc.OnDemand && len(cc.AllowHosts) > 0 {
log.Printf("crawl: allow_hosts is set, so on-demand reading is limited to those hosts")
}
return hosts
}
// crawlerWithHosts builds a crawler over one allowlist. Two callers, two lists:
// see newCrawler and onDemandCrawler.
func crawlerWithHosts(cc *config.CrawlConfig, hosts []string) *crawl.Crawler {
ua := cc.UserAgent
if ua == "" {
ua = webfetch.DefaultUserAgent
@@ -81,7 +98,14 @@ func onDemandCrawler(cfg *config.Config) *crawl.Crawler {
if cfg.Crawl == nil || !cfg.Crawl.OnDemand {
return nil
}
return newCrawler(cfg)
cc := cfg.Crawl
// His own allow_hosts, and nothing added behind his back. Empty means "any
// host that is not denied and not private", which is what on-demand reading
// of a URL he just said out loud has to mean.
if len(cc.AllowHosts) > 0 {
log.Printf("crawl: allow_hosts is set, so on-demand reading is limited to those %d host(s)", len(cc.AllowHosts))
}
return crawlerWithHosts(cc, crawlHosts(cc, false))
}
// crawlWorker — ticker + watcher for the scheduled half.
@@ -137,9 +161,20 @@ func (w *crawlWorker) run(ctx context.Context) {
// net/http out of the crawler package.
type crawlFetcher struct{ f *webfetch.Fetcher }
// Get maps webfetch's sentinels onto crawl's. This adapter is the one place
// that imports both packages, so the mapping belongs here; the crawler used to
// match on three substrings of a message it could not see the definition of,
// and a reworded error would have quietly turned a blocked host into "there is
// no robots.txt here".
func (a *crawlFetcher) Get(ctx context.Context, u string) (*crawl.Response, error) {
resp, err := a.f.Get(ctx, u)
if err != nil {
switch {
case errors.Is(err, webfetch.ErrBlocked), errors.Is(err, webfetch.ErrPrivate), errors.Is(err, webfetch.ErrScheme):
return nil, fmt.Errorf("%w: %v", crawl.ErrFetchRefused, err)
case errors.Is(err, webfetch.ErrStatus):
return nil, fmt.Errorf("%w: %v", crawl.ErrFetchStatus, err)
}
return nil, err
}
return &crawl.Response{URL: resp.URL, ContentType: resp.ContentType, Body: resp.Body}, nil