fix(kernel,server,workflow): per-stage retry budget, sane freestyle budget, tail /events
Three defects surfaced by a live headless freestyle_planning run (all made a long execution workflow fail or stall in ways invisible to a REST operator): 1. retryCount was session-global (DefaultOrchestrationReducer): TransitionExecuted updated currentStageId but never reset retryCount, while shouldRetry gates on it per-stage. An early stage that burned its 3 retries starved every later stage of its own — the next stage's first *retryable* failure went terminal. Reset retryCount on stage entry; back-edge loops stay bounded by the separate refinementIterations counter. Regression test added. 2. Freestyle stages ran at StageConfig's 4096-token default (ExecutionPlanCompiler never set a budget) vs 16384-32768 for static role_pipeline stages. A tool-heavy stage truncated its context every round, evicting the file_read results it just gathered, so it re-read forever and never accumulated enough to write. Set 16384. 3. GET /events returned the OLDEST `limit` events (readFrom(0) then take): for a session past `limit` events the tail — including a pending ApprovalRequested — was invisible to a headless/REST poller, the exact caller POST /approve serves. Default now tails; explicit fromSeq still paginates forward.
This commit is contained in:
+7
-1
@@ -53,7 +53,13 @@ class DefaultOrchestrationReducer : OrchestrationReducer {
|
||||
pauseReason = null,
|
||||
)
|
||||
|
||||
is TransitionExecutedEvent -> state.copy(currentStageId = p.to)
|
||||
// Entering a new stage resets the retry budget: retryCount is per-stage (RetryAttemptedEvent
|
||||
// is stage-scoped, maxAttempts is a per-stage cap), so without this reset the budget is shared
|
||||
// session-wide — an early stage that exhausts its retries starves every later stage of its own,
|
||||
// turning the next stage's first retryable failure into a terminal WorkflowFailed (live repro:
|
||||
// implement burned 3/3, then review's retryable JSON-decode error got 0 retries). Back-edge
|
||||
// loops stay bounded by the separate refinementIterations counter, so this can't loop forever.
|
||||
is TransitionExecutedEvent -> state.copy(currentStageId = p.to, retryCount = 0)
|
||||
|
||||
is RetryAttemptedEvent -> state.copy(
|
||||
status = OrchestrationStatus.RUNNING,
|
||||
|
||||
Reference in New Issue
Block a user