feat(research): wire tools + research workflow graph (research-workflow §2/§3)

Makes the research feature runnable end-to-end, off by default.

- config: [tools.research] (enabled, searxng_url, max_results, max_fetch_bytes).
- registration: web_search/web_fetch are built into BOTH the default and per-workspace
  tool registries when research.enabled, sharing one HTTP client threaded from Main
  (none built on the static path). Egress stays harness-enforced: web_fetch is T2
  (operator-approved) and the existing NetworkHostRule still applies.
- workflow: examples/workflows/research.toml — decompose → gather → report, with the
  three artifact schemas and prompts. Fan-out (search per sub-question, fetch per source)
  runs as repeated tool calls inside the gather stage (Correx has no parallel agents);
  per-source synthesis into the dossier is the compression step, so the report stage
  consumes summaries, never raw pages. ResearchWorkflowTest validates the graph contract.

To run: set [tools.research].enabled, register the 3 [[artifacts]], copy research.toml +
prompts + schemas into the workflows dir, start SearXNG. Launch like any workflow (the
T2 fetch approval surfaces as an approval card; the report opens in the artifact viewer).

Follow-ups (noted, not blocking): batch fetch-approval at the source-list level (§3),
a dedicated SourceFetched/LowQualityExtraction event (quality + content hash are already
in tool-result metadata), dynamic per-session egress allowlist, and the web approval client (§6).
This commit is contained in:
2026-06-13 23:18:19 +04:00
parent ad2d38ce46
commit 7a0d4d0ee2
11 changed files with 321 additions and 1 deletions
@@ -0,0 +1,20 @@
You are the **Research Planner** — the first stage of a deep-research workflow.
Your job is to turn one research question into a concrete plan to answer it. You do not search
or browse yet; you decompose.
Steps:
1. Restate the research question in your own words so intent is unambiguous.
2. Break it into a small set of sub-questions that can each be answered independently. Cover the
whole question; avoid overlap. Prefer 36 sub-questions over a long list.
3. For each sub-question, write one or more concrete web search queries — the actual strings you
would type into a search engine, specific enough to surface authoritative sources.
The decision history above (steering, approvals) is ground truth — honour it.
Emit your result as the `research_plan` artifact (JSON, schema provided):
- `question`: the research question, restated.
- `sub_questions`: the decomposed sub-questions, one per item.
- `search_queries`: concrete search queries to run, one per item.
Do not answer the question here. Plan only.