feat(research): deterministic HTML→markdown extraction (research-workflow §5)

First slice of the research workflow: the pure, zero-inference extraction pipeline
that the WebFetchTool will consume. New :infrastructure:research module (jsoup).

HtmlMarkdownExtractor: content-type dispatch (HTML extracted; JSON/text pass
through; binary/media/pdf rejected at the header), then structural strip
(MainContentSelector: drop script/style/nav/footer/aside/form + class/id blocklist)
+ density extraction (text − link-text score, semantic <article>/<main> preferred)
+ DOM→markdown (MarkdownRenderer/Table/Inline: headings, lists incl. nested, tables,
fenced code, blockquotes, links, emphasis). Output below the min-length threshold is
flagged LOW_QUALITY so the workflow can route around dead sources (SPAs, paywalls).

extractorVersion ("html-md-1") is pinned on every result so replay reads recorded
artifacts and never re-extracts — the embedding-hash discipline (§5, ADR-0000 §10).

Next slices: WebFetchTool + egress policy + CAS wiring (emits the fetch +
LowQualityExtraction events), then WebSearchTool/SearXNG, workflow graph, synthesis.
This commit is contained in:
2026-06-13 22:44:57 +04:00
parent f8fd2601a8
commit fb1d97058a
7 changed files with 505 additions and 0 deletions
+1
View File
@@ -39,6 +39,7 @@ include ':infrastructure:tools'
include ':infrastructure:tools:filesystem'
include ':infrastructure:workflow'
include ':infrastructure:artifacts-cas'
include ':infrastructure:research'
include ':testing:replay'
include ':testing:contracts'