Three generic harness fixes from the web-ui postmortem dataset. Nothing here keys
on a language, framework, build tool or task type.
1. Failure attribution. WorkflowFailedEvent carries one primary FailureAttribution
(AGENT | HARNESS | WORKFLOW | ENVIRONMENT | PROVIDER | OPERATOR | UNKNOWN),
defaulted to UNKNOWN so pre-field events replay unchanged. FailureAttributor is
the deterministic reason->layer mapping, used both at emission and when
classifying history, so the baseline and the live metric are one measurement.
Emission sites set it: failWorkflow derives from the reason unless the caller
knows the layer, cancellation is OPERATOR, the server catch-all falls back to
HARNESS, a grounding-rejected plan is AGENT. Multi-cause chains stay on
FailureTicketOpened — no second causal structure.
GET /metrics/failure-attribution (FailureAttributionInspectionService, mirroring
ToolReliabilityInspectionService) reports counts, share, UNKNOWN share, the
preserved reasons and the ticket categories from the same sessions. Read-only:
historical events are classified at READ time and reported as `inferred`, never
written back over an append-only log.
Baseline over the local log, 122 terminal failures: AGENT 51 (41.8%),
OPERATOR 28 (23.0%), WORKFLOW 19 (15.6%), PROVIDER 15 (12.3%), HARNESS 6 (4.9%),
ENVIRONMENT 3 (2.5%), UNKNOWN 0.
2. The `~` guard bug. ToolPath is now the ONE canonical normalization rule
(expand a leading `~`/`~/`, keep absolutes, anchor relatives on the session
working dir). Every filesystem tool, all six plane-2 path rules and the approval
preview resolve through it, so policy and existence checks inspect the path the
tool will operate on. `~/.gradle/init.d/offline.gradle` used to resolve to
`<workspace>/~/.gradle/...`: reported non-existent AND in-workspace, so the
reference gate called a real file a hallucination and the out-of-workspace prompt
never fired. Containment and external-read approval behaviour are unchanged —
the expanded path is simply outside the workspace, where it always belonged.
3. file_copy (#713). A first-class tool with the writer's jail, tier, receipt,
replay and CAS pre/post images; static and binary assets no longer move through
the model's token stream. Needed one generic split: ParamRole.SOURCE_PATH marks a
path a call reads FROM, so containment gates judge both params while write-target
gates (read-before-write, stale-write, write scope, write manifest) judge the
mutated one. ReadBeforeWriteRule exempts any call declaring a SOURCE_PATH: its
content comes from disk, not from memory, and requiring a read of a binary is
unsatisfiable. Existing tools declare no SOURCE_PATH, so their behaviour is
byte-identical.
Tests: ToolPathTest (9), FailureAttributionTest (10), PathNormalizationRuleTest (6),
FileCopyToolTest (10), plus a home-relative FileReadTool read. ./gradlew check green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This branch's uncommitted WIP, committed together (entangled at file level).
Distinct pieces of work:
Freestyle QA fixes (this session):
- FileEditTool: pre-validate replace anchor in validateRequest — reject a
missing/ambiguous target BEFORE the approval gate, mirroring read/write's
file-not-found / read-before-write pre-checks. Shared not-found/ambiguous
messages between validate and execute so they can't drift.
- PlanGrounder: add `scanned` flag; when no RepoMapComputedEvent was recorded,
repoMapPaths is "unknown" not "empty workspace" — skip scope grounding
(which proves a path ABSENT) so real paths (apps/server/**) aren't falsely
rejected. Build-manifest check still runs.
- FreestyleDriver: wire scanned=(repoMap!=null); on plan rejection emit a
session-terminal WorkflowFailedEvent so a rejected run reads FAILED, not the
COMPLETED-lie (last verdict was the planning-phase WorkflowCompleted).
- ServerModule: resolve project-memory workspace root from the session's bound
workspace (sessionWorkspaceRoot) instead of boot-static pm.repoRoot(), fixing
the workspace-binding divergence (correx vs empty scratch dir). Retire tracked
in Vikunja #266.
- LaunchRegistrationRaceTest: join registered jobs before asserting launchCount
— computeIfAbsent returns the Job immediately but the fire-and-forget launch
body lagged awaitAll (the 49-vs-50 flake).
ACR concept-compiler experiment (pre-existing WIP on this branch):
- ExecutionPlanCompiler/Model/PlanLinter, #264 needs-seam (sessionArtifacts),
LSP diagnostics subsystem (LspDiagnosticEvents/Runner/Lsp4j), BootWorkspace,
config surface, workflow prompts/schemas, orchestrator advance-don't-rerun.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Enable autonomous QA through a remote OpenAI-compatible provider (NVIDIA NIM)
and harden the tool/approval path so unattended multi-stage runs complete.
- inference: add openai_compat provider (Bearer chat-completions for NIM/OpenAI),
dispatched by provider type "nim"/"openai"; key via api_key/api_key_env.
- server: bind configured [server] host/port instead of a hardcoded 8080;
POST /sessions accepts an optional `intent` (WS parity) for intent-driven workflows.
- kernel: thread the bound operator profile's approval_mode into per-tool gating so
auto/yolo enable unattended approval (engine still consulted; policy/plane-2 BLOCK
stays terminal); on a recoverable tool failure feed the tool's arg-schema back into
context so the model self-corrects instead of repeating a malformed call.
- tools: split deletion out of file_write into a separate, explicitly-named file_delete
tool — a model can no longer delete a file by getting a write-mode parameter wrong.
- server: add GET /metrics/tool-reliability — per-model tool-call validity from the
event log (measurement groundwork for capability-aware routing).
- docs: update AGENTS.md across kernel, tools, server, inference.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Root Child DOX Index assembled, plus a per-module AGENTS.md across the tree
(core/*, infrastructure/*, apps/*, testing/*, and docs/examples/frontend/etc),
each following the DOX section shape: Purpose, Ownership, Local Contracts,
Work Guidance, Verification, Child DOX Index.