e46777e29f
Two reviewer-reliability tracks, both deterministic and unit-verified (no model/network). §B-§5 — static_check stage seam. The StaticAnalysisRunner + StaticFindingsRecordedEvent + reviewer-context filter already existed; what was missing was a stage that runs the tool and emits the event. Added: - StaticCheckStageExecutor: reads stage metadata (static_tool/static_argv), runs the configured command via StaticAnalysisRunner, returns a StaticFindingsRecordedEvent. No-op (empty findings) when no runner is wired or no command is configured — safe to carry unconfigured. - A deterministic-stage seam in DefaultSessionOrchestrator.enterStage: any stage with metadata["stage_type"] == "static_check" is run by the executor instead of the LLM subagent, then advances on its (unconditional) exit edge. - TomlWorkflowLoader: stage_type/static_tool/static_argv fields → StageConfig.metadata. - role_pipeline.toml: a static_check stage between implementer and reviewer (no-op until static_argv + a CommandRunner are set; activation is live-QA-gated, see StaticAnalysisRunner doc). §B-§6 — critique-outcome producer. The CritiqueFinding type + CriticCalibrationProjection existed but nothing fed them. Added: - CritiqueFindingsRecordedEvent (+ CritiqueVerdict): the producing side — a critic's findings + verdict for one review iteration, carrying modelHash for per-model calibration. - CritiqueOutcomeCorrelator: pure loop-resolution logic deciding UPHELD (fixed between rounds) / DISMISSED (persisted into an approved final) / INCONCLUSIVE (open at a non-approved terminal), per critic (role + modelHash) and per finding id. - A hook in completeWorkflow/failWorkflow that correlates recorded findings into CritiqueOutcomeCorrelatedEvents at loop resolution — no-op when none recorded, idempotent. (LLM-side finding emission stays a separate model-gated activation.) Tests: StaticCheckStageExecutorTest (4), StaticCheckStageTest integration (1, fake runner → event + transition), CritiqueOutcomeCorrelatorTest (8), CritiqueCalibrationWiringTest integration (1, seeded findings → outcomes at completion), updated RolePipelineWorkflowTest. Full suites for core:events/kernel/critique, infrastructure:workflow, testing:integration green; detekt clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
155 lines
5.3 KiB
TOML
155 lines
5.3 KiB
TOML
# Role pipeline: analyst → architect → planner → implementer → static_check ⇄ reviewer
|
|
#
|
|
# Each stage produces a typed artifact that the next stage `needs`, so work flows forward
|
|
# without a human relaying notes. The decision journal (pinned into every stage's context)
|
|
# carries steering/approvals/verdicts across the whole run, so the reviewer sees the same
|
|
# ground truth as the planner.
|
|
#
|
|
# The implementer⇄reviewer loop is gated by review_report.verdict:
|
|
# approved → done
|
|
# changes_requested → back to implementer (capped by implementer.max_retries, then escalates)
|
|
#
|
|
# Requires these artifact kinds in ~/.config/correx/config.toml (schemas under docs/schemas/):
|
|
# [[artifacts]]
|
|
# id = "analysis"; schema_path = "schemas/analysis.json"; llm_emitted = true
|
|
# [[artifacts]]
|
|
# id = "design"; schema_path = "schemas/design.json"; llm_emitted = true
|
|
# [[artifacts]]
|
|
# id = "impl_plan"; schema_path = "schemas/impl_plan.json"; llm_emitted = true
|
|
# [[artifacts]]
|
|
# id = "review_report"; schema_path = "schemas/review_report.json"; llm_emitted = true
|
|
#
|
|
# Prompt files (prompts/*.md, relative to this workflow) must exist for a real run.
|
|
|
|
id = "role_pipeline"
|
|
start = "analyst"
|
|
|
|
# 1. Understand the request and the relevant code. Read-only.
|
|
# ground_references: every workspace-relative file path the analysis names is checked
|
|
# for existence; a hallucinated path fails the stage and retries with the misses fed back.
|
|
[[stages]]
|
|
id = "analyst"
|
|
prompt = "prompts/analyst.md"
|
|
produces = [{ name = "analysis", kind = "analysis" }]
|
|
allowed_tools = ["file_read", "ShellTool"]
|
|
ground_references = true
|
|
token_budget = 16384
|
|
max_retries = 2
|
|
|
|
# 2. Decide the approach and component boundaries.
|
|
[[stages]]
|
|
id = "architect"
|
|
prompt = "prompts/architect.md"
|
|
needs = ["analysis"]
|
|
produces = [{ name = "design", kind = "design" }]
|
|
token_budget = 16384
|
|
max_retries = 2
|
|
|
|
# 3. Break the design into ordered, verifiable steps.
|
|
[[stages]]
|
|
id = "planner"
|
|
prompt = "prompts/planner.md"
|
|
needs = ["design"]
|
|
produces = [{ name = "impl_plan", kind = "impl_plan" }]
|
|
token_budget = 16384
|
|
max_retries = 2
|
|
|
|
# 4. Implement the plan. Writes files (jailed to the workspace). The loop target —
|
|
# max_retries here caps how many review→implement refinement rounds are allowed.
|
|
# Optional: a `writes` manifest (workspace-relative globs) hard-bounds where this
|
|
# stage may write — a FILE_WRITE outside it is blocked as scope creep. Left open
|
|
# here because the targets are task-specific; a task-scoped workflow would set e.g.
|
|
# writes = ["core/sessions/**", "testing/sessions/**"]
|
|
[[stages]]
|
|
id = "implementer"
|
|
prompt = "prompts/implementer.md"
|
|
needs = ["impl_plan"]
|
|
produces = [{ name = "patch", kind = "file_written" }]
|
|
allowed_tools = ["file_read", "file_write", "file_edit", "ShellTool"]
|
|
token_budget = 32768
|
|
max_retries = 3
|
|
|
|
# 4b. Deterministic static analysis (NO LLM). Runs a configured tool (compiler/ktlint/detekt)
|
|
# over the patch and records its findings as a StaticFindingsRecordedEvent, which the
|
|
# reviewer-context filter then strips from the reviewer's context so the LLM reviewer spends
|
|
# its attention on semantic review rather than re-finding what the build already reports.
|
|
# No-op until `static_argv` is set (and a CommandRunner is wired) — safe to carry unconfigured.
|
|
# Activate e.g. with: static_argv = "./gradlew detekt --console=plain"
|
|
[[stages]]
|
|
id = "static_check"
|
|
stage_type = "static_check"
|
|
static_tool = "detekt"
|
|
token_budget = 0
|
|
max_retries = 0
|
|
|
|
# 5. Review the patch against the plan AND the analyst's acceptance criteria. The reviewer
|
|
# needs `analysis` so it judges the diff against concrete, pre-stated criteria (§5 narrow
|
|
# question) rather than whole files against taste.
|
|
[[stages]]
|
|
id = "reviewer"
|
|
prompt = "prompts/reviewer.md"
|
|
needs = ["patch", "impl_plan", "analysis"]
|
|
produces = [{ name = "review_report", kind = "review_report" }]
|
|
token_budget = 32768
|
|
max_retries = 2
|
|
|
|
# --- forward edges ---
|
|
|
|
[[transitions]]
|
|
id = "analyst-to-architect"
|
|
from = "analyst"
|
|
to = "architect"
|
|
condition_type = "artifact_validated"
|
|
condition_artifact_id = "analysis"
|
|
|
|
[[transitions]]
|
|
id = "architect-to-planner"
|
|
from = "architect"
|
|
to = "planner"
|
|
condition_type = "artifact_validated"
|
|
condition_artifact_id = "design"
|
|
|
|
[[transitions]]
|
|
id = "planner-to-implementer"
|
|
from = "planner"
|
|
to = "implementer"
|
|
condition_type = "artifact_validated"
|
|
condition_artifact_id = "impl_plan"
|
|
|
|
# implementer → static_check (once the patch is validated) → reviewer. The static_check stage
|
|
# produces no artifact, so its exit edge is unconditional (always_true).
|
|
[[transitions]]
|
|
id = "implementer-to-static_check"
|
|
from = "implementer"
|
|
to = "static_check"
|
|
condition_type = "artifact_validated"
|
|
condition_artifact_id = "patch"
|
|
|
|
[[transitions]]
|
|
id = "static_check-to-reviewer"
|
|
from = "static_check"
|
|
to = "reviewer"
|
|
condition_type = "always_true"
|
|
|
|
# --- verdict-gated loop exit / re-entry ---
|
|
|
|
[[transitions]]
|
|
id = "review-approved"
|
|
from = "reviewer"
|
|
to = "done"
|
|
condition_type = "artifact_field_equals"
|
|
condition_artifact_id = "review_report"
|
|
condition_field = "verdict"
|
|
condition_value = "approved"
|
|
condition_operator = "eq"
|
|
|
|
[[transitions]]
|
|
id = "review-changes-requested"
|
|
from = "reviewer"
|
|
to = "implementer"
|
|
condition_type = "artifact_field_equals"
|
|
condition_artifact_id = "review_report"
|
|
condition_field = "verdict"
|
|
condition_value = "approved"
|
|
condition_operator = "neq"
|