Files
manga-recap-pipeline/caveats/audit-open.md
T
kami 63f7918a3e Record all six defects and the 9% baseline
Adds decision entries for the unpaired set-of-mark label, the interjection
verifier false positive, and the vision-blob clearing bug, plus the per-run
speaker audit script used to measure the chapter.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 23:59:32 +04:00

7.7 KiB

Open limits from the 2026-08-11 audit

Everything here was read in source during the audit and deliberately left unfixed in Phase 1. The fixed findings live in decisions/audit-phase1.md. Line numbers are from the audit and may drift.

Reconcile deletes the losing character irreversibly

Reconciliation deletes the losing character row (db.py:506). Clearing the reconcile stage does not undo it, and name claims attached to the merged-away character are not repointed.

Costs: one bad merge is unrecoverable without rebuilding the identity stage for the whole manga. Revisit when: identity work resumes, or a reviewer reports a wrong merge on a real chapter. Workaround: none. Clear identity and rerun, which loses the good merges too.

Clearing a stage does not undo what it wrote

The dialogue and direct half of this is fixed and proven (decisions/storage-layout.md#clear-vision-blob). It cost a wasted rerun on 2026-08-11 first: the clear returned {"ok": true}, deleted nothing, and the stage skipped all 116 panels.

What remains: clearing identity preserves the per-manga registry by design, so a rerun inherits every character it ever minted (caveats/speaker-attribution.md#registry-pollution). Nothing verifies that a clear emptied what it claimed.

Costs: a rerun silently reuses stale output, which reads as a reproducible result. Revisit when: any stage is scheduled concurrently or resumed automatically. A stage must be idempotent before either is safe.

ComfyUI uses the GPU outside the session mutex

The layers stage calls ComfyUI directly and takes no lease. Another job can load gemma, siglip2, or dots while ComfyUI holds the same GPU.

Costs: out-of-memory failures that look random and land on an unrelated stage. Revisit when: layers is enabled on a real run, or a second concurrent job is allowed. Workaround: MAX_CONCURRENT_JOBS=1 keeps one pipeline at a time, which is the current default.

The JSON repair pass can fabricate dialogue

call_gemma4_json hands the model its own truncated text and asks for the JSON it should have been (worker_vision.py). The repair call carries no image. On a response truncated by max_tokens, the model completes dialogue it can no longer see. What it adds is indistinguishable downstream from transcribed text.

Costs: invented lines enter the script with normal provenance. Revisit when: schema-constrained generation lands, which removes most of this path. Workaround: resend the image on repair, or retry a truncated response instead of repairing it.

/review/preview pins a solo clip into the final video

review_preview renders one panel and calls save_clip(panel_id, ...). _render_one_beat returns early when a clip already exists for the leader. Previewing a beat leader therefore drops the rest of the beat's panels. review_retts gets this right by calling delete_clip first.

Costs: a reviewer silently corrupts the output by looking at it. Revisit when: the review UI is used on a real chapter. Workaround: never preview a beat leader, or delete the clip row afterwards.

Stored embeddings carry no model or pooling version

embed_crop uses pooler_output when present and falls back to mean-pooled patch tokens otherwise. The two paths produce different vector spaces, and the fixed 0.85 threshold is valid for one of them. Nothing beside a stored .npy records which model, revision, or pooling produced it.

Costs: a transformers upgrade mixes incompatible vectors into one gallery with no error. Revisit when: transformers or the siglip2 revision is upgraded. Before, not after. Workaround: none. Write the model id and pooling mode beside the vector and refuse cross-version comparison.

Worker endpoints block the event loop

Endpoints declared async def run blocking MinIO, OpenCV, torch, and ffmpeg calls directly on the event loop. A busy worker cannot answer /health or /unload.

Costs: /health/workers reports a working worker as unreachable, and the session manager's 30-second /unload can time out exactly when VRAM needs freeing. Revisit when: a stage stalls on /unload, or before any bounded parallelism lands. Workaround: def instead of async def moves each handler to the threadpool. Tracked as [#199] for the render worker.

SQLite has no busy timeout

get_conn opens a connection per call with no busy_timeout. WAL tolerates one writer. PIPELINE=1 already writes clips from concurrent tasks while TTS writes audio.

Costs: planned CPU parallelism will surface as database is locked before it surfaces as throughput. Revisit when: Phase 2 concurrency work starts. Set the timeout first. Workaround: keep PIPELINE off.

layers runs after tts

STAGES orders layers after tts, while run_stage_tts warns that eager rendering under PIPELINE=1 needs layers to run first.

Costs: with the flag on, solo beats always render without parallax. Revisit when: a real run enables PIPELINE=1. Workaround: keep PIPELINE off, or reorder STAGES.

completed means something different in each stage

A dropped vision panel fails the stage and halts the pipeline. A failed direction window counts its panels as done. Layers always finishes completed.

Costs: the acceptance metrics in ROADMAP.md cannot be read across stages. Revisit when: a baseline measurement is taken. The numbers are meaningless until then. Workaround: none.

Identity worker caches characters the orchestrator has deleted

worker_identity._known_cache is invalidated only by _persist_char. The reconcile stage deletes losing characters directly in the orchestrator database. A long-lived worker keeps shortlisting and assigning ids that no longer exist.

Costs: assignments point at rows that are gone. Revisit when: reconcile runs on a chapter without a worker restart between stages. Tracked as [#201], which proposes caching per manga_id. Workaround: restart the identity worker after reconcile.

A character seen once gets no assignment at all

worker_identity._pending holds full crop images in worker memory for a whole chapter, keyed by session. That is durable per-chapter state inside a worker documented as stateless, and it is lost on restart. A character seen exactly once receives no assignment, not even a chapter-local handle.

Costs: one-off characters vanish from the scene graph. Revisit when: chapter-local tracklet persistence lands (ROADMAP.md, Phase 3). Workaround: none.

MinIO credentials are hardcoded in committed source

Defaults live in transport.py and service.py.

Costs: the credentials are in git history for anyone who gets the repo. Revisit when: the repo leaves this machine, or MinIO is reachable outside the LAN. Workaround: the environment variables already override them. Set them and remove the defaults.

Assemble marks a job completed with no clips

run_stage_assemble does not check that clip_uris is non-empty before assembling, then marks the job completed.

Costs: a failed chapter reports success. Revisit when: any run reports completed without a video. One if not clip_uris guard fixes it. Workaround: none.

Reviewer timestamps drift against the crossfaded video

/review/panels sums per-panel audio durations. Assemble crossfades clips using the per-beat transitions, so every non-cut transition shortens the real video.

Costs: reviewer timestamps drift further out of sync the further into the chapter they scrub. Revisit when: the review UI is used for timing work. Workaround: subtract the transition overlaps by hand.