Files
manga-recap-pipeline/NEXT.md
T
kami a9d64fe80a Record panel 7 against the art, and the A/V gap it hid
Somebody watched chapter.mp4 for the first time. Two failures came out of it
that no stage counter could see.

chapter.mp4 is video 436.39s over audio 363.67s, so narration finishes 72.7s
before the picture. The 49 clips are clean: all 25fps, video and audio agree
to 0.03s, summing to 363.6s. A per-round probe puts the loss in the final
round of _assemble_batched, which turns 359s of video into 100s while the
audio survives. Round 0 is correct. Round 1 differs by holding a 7th input,
the leftover clip that skips encoding, so the tree mixes concat output, xfade
output and a raw clip. Not fixed.

worker_render.py gains an FPS constant, fps normalization in the xfade branch
to match concat, _stream_dur, and a self-check that compares video against
audio instead of asserting the file is non-empty. That old check is how a 20%
sync failure shipped. The fps inconsistency is real but not proven to be the
shipped cause. Pinning -r on the output was tried and reverted: it drops
frames to force CFR, which the concat branch comment already warned about.

Panel 7 checked against the art has zero correct identity bindings out of two,
and Seonho, the one character who matters, is unbound. bbox values are
consumed as absolute pixels; on a 900x1650 panel that puts all six boxes in
the top third, two inside a speech balloon. Identity therefore embeds crops of
balloon edges and window frames, which is how confidence 0.9 lands on the
wrong person. The colleague has no name in the story and was labelled Choi
Haeseon; that row holds 25 of 26 assignments, so it is the label the pipeline
stamps on any unnamed woman.

Four caveats added. Two earlier claims are withdrawn in place: rescaling bbox
by 1000 does not make the boxes correct, and the constraint is not 16 nameless
rows needing names. Cast profiles already exist, since all 53 rows populate
ref_image_uris and embedding_uri, but they are enrolled from the wrong crops.

worker_render.py self-check passes. No pipeline ran.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 01:07:17 +04:00

9.7 KiB

NEXT

Updated 2026-08-12. What this session did is in HANDOFF.md.

State

The chapter runs end to end and the output is not watchable. That is now measured, not guessed.

Job 778297bc-e7ce-439d-91b5-8a027060d17f, chapter 7c944dd4-e972-42c7-ba60-9f6939548e80, 116 panels, status=completed, finished 2026-08-11T20:08:16Z. s3://video/ holds 49 clips and a 50MiB chapter.mp4. The user watched it and read out 19 defects. They are grouped by cause in HANDOFF.md.

Two numbers set the agenda:

  • chapter.mp4 is video 436.39s over audio 363.67s. The narration finishes 72.7s before the picture.
  • Panel 7 checked against the art has zero correct identity bindings out of two, and the one character who matters is unbound. HANDOFF.md#panel-7-walked-against-the-art has the table.

Uncommitted work sits in the tree: worker_render.py has an FPS = 25 constant, fps normalization in the xfade branch, a new _stream_dur, and a self-check that compares video against audio. The self-check passes. It does not yet fix the chapter. Details and one dead end in HANDOFF.md.

Next

  1. Fix chapter assembly. _assemble_batched turns 359s of video into 100s while the audio survives. It reproduces offline in two minutes, no GPU. Round 0 is correct and round 1 collapses. Round 1 is the only round holding a raw clip that skipped encoding, so suspect the single-item passthrough first. The recommended shape is one path, not three: normalize every input, then xfade every boundary, treating cut as a 0.05s fade. acrossfade and xfade shorten audio and video equally, so the streams stay locked. HANDOFF.md holds the per-round table, the filtergraph, and the repro commands. Nothing downstream is worth judging until this lands.

  2. Fix identity, in this order. Panel 7 is the worked example and HANDOFF.md#panel-7-walked-against-the-art carries the evidence. Do not start at the registry.

    a. Settle the bbox coordinate space. Consumed as pixels, all six boxes on panel 7 land in the top third of the panel, two inside a speech balloon. Divided by 1000 they mostly land on their subjects. worker_vision.py:271 and worker_identity.py:91 both assert pixels, and the art says otherwise. Identity embeds _crop_bbox(img, ch["bbox"]) at worker_identity.py:200, so today it matches faces against crops of balloons and window frames. Nothing downstream can be judged until this is right. Rescaling alone is not the fix: after scaling, one box still sits on an empty window frame and another clips its subject. b. Let identity abstain and stay abstained. The colleague has no name in the story and was labelled Choi Haeseon at 0.9. An unnamed recurring person needs a stable anonymous identity so narration says "the colleague" every time. match() already returns None below threshold. Check whether the Tier-2 gemma resolver can answer "none of these"; that was not verified. c. Separate extra from cast. Four of the six detections on panel 7 are background extras or nothing at all, and all six reach identity as equal candidates. d. Only then merge Seonho into Lim Seonho and split character_afa7623b, which still needs the reversible-merge design (caveats/audit-open.md#destructive-reconcile), not a patch.

    Cast profiles already exist. Do not rebuild them. The user asked whether the main cast could get a profile built from reference frames and reused. characters already carries ref_image_uris and embedding_uri, and all 53 rows have both populated. The mechanism is not missing, it is enrolled from the wrong crops, so today it stores references to balloon edges and window frames. Step (a) is what makes it work. Three things are genuinely absent and are the smaller follow-on:

    • no quality gate on enrollment, so nothing checks that a reference crop holds a face at all
    • nothing re-enrolls a reference set once it is written, so the wrong crops persist

    The visual "is this them?" check is already built. Do not write it again. /vision/resolve at worker_vision.py:963 sends the query crop plus up to 3 labelled reference images per candidate. build_resolve_prompt tells the model to judge face shape first, to treat hair and outfit as secondary, that two people sharing a hair colour are not the same, and to answer 0 for NONE when unsure. choice: 0 becomes a new character, an out-of-range index becomes unresolved, and ref_image_uris is republished as reference_image_uris at worker_identity.py:152 and :161. The mechanism, the prompt and the abstain path are all correct. They are fed crops of the wrong region, which is step (a).

    • no human gate to name, merge or split the clusters. The user wants this as a minor adjustment on top, not as the mechanism. The gates table and the review gates from [#136] are the place to hang it

    The chibi at 1:35 will survive all of this. He genuinely is brown hair plus a yellow shirt, so a profile match is correct on appearance and wrong on reality. That needs item 4 below, plus requiring a real face before a crop can enroll.

  3. Stop the narration inventing facts. 0:43, 2:03, 2:05 and 2:15 assert things no panel shows. The correctness verifier passed 116/116 because it checks quotes and names, never invented claims.

  4. Teach vision that art inside a panel is not the scene. A chibi on a monitor became "a man holding a drink" at 1:35. A colleague pointing into the distance became "pointing towards the screen" at 1:59.

  5. layers writes nothing and reports completed 116/116, so no clip has parallax and a still holds for 28s from 2:24 (caveats/audit-open.md#layers-writes-nothing).

  6. Clear the stale job error. The completed job still carries error: "partial: 112/116 completed" (caveats/audit-open.md#stale-job-error).

  7. Balloon-to-speaker geometry via the unused det/seg heads (caveats/speaker-attribution.md#tail-is-not-geometry) is now behind item 2. With no name to attach, geometry buys nothing.

  8. Resolve a speaker answer across the whole dialogue window, not just the answering panel. The last 3 unresolved refs describe a neighbouring panel in the same 8-panel call.

  9. Start Phase 2 from ROADMAP.md. Set SQLite busy_timeout before any concurrency work (caveats/audit-open.md#sqlite-locking).

Lesson worth keeping

Every metric recorded before this session said the pipeline was fine or nearly fine. script 116/116, "9 named speech lines", layers 116/116, assemble 1/1. Watching two and a half minutes of output found a 20% sync failure, a cast that is 84% anonymous, invented narration, and a stage that writes nothing while reporting success. Stage counters measure whether code ran. They say nothing about whether the result is correct. Watch the output before trusting a number.

Running the pieces

./start_workers.sh              # session_manager + 9 workers, each a uvicorn in a tmux window
tmux attach -t manga-workers    # per-worker logs
.venv/bin/python worker_render.py    # self-check, runs real ffmpeg, about 4 minutes

Read the state, or clear a stage and resume:

/usr/bin/ssh kami@192.168.1.104 "curl -s 'http://127.0.0.1:9090/job/status?job_id=778297bc-e7ce-439d-91b5-8a027060d17f'"
/usr/bin/ssh kami@192.168.1.104 "curl -s -X POST http://127.0.0.1:9090/stage/clear -H 'Content-Type: application/json' -d '{\"job_id\":\"778297bc-e7ce-439d-91b5-8a027060d17f\",\"stage\":\"<stage>\"}'"
/usr/bin/ssh kami@192.168.1.104 "curl -s -X POST 'http://127.0.0.1:9090/job/resume?job_id=778297bc-e7ce-439d-91b5-8a027060d17f'"

Traps: plain ssh is the kitty ssh kitten and refuses non-interactive stdin, so use /usr/bin/ssh. mc aliases on homesrv are homesrv and mio. local returns Access Denied and rfs is the empty rustfs. cp is aliased to cp -i and hangs on overwrite, so use /usr/bin/cp -f.

Re-fixing assembly needs the real clips, which the session scratchpad no longer holds:

/usr/bin/ssh kami@192.168.1.104 'P=homesrv/video/ef105a86-4b7e-4ac4-b45c-b7d83b8f5b5e/7c944dd4-e972-42c7-ba60-9f6939548e80; mc cp -q -r $P/clips/ /tmp/rclips/; cd /tmp/rclips && tar cf - .' | tar xf - -C clips/

Storage and viewer, tasks #116/#117

[#117] is done. stowage serves the manga buckets. It was never a MinIO problem: the container had been dead since 2026-07-19 on an arm64 digest pin.

[#116] is closer but not cut over. Artifacts split one bucket per class (decisions/storage-layout.md#bucket-per-artifact), and both MinIO and rustfs hold all six buckets. rustfs on 127.0.0.1:9010/9011 is still empty and nothing is repointed, so MinIO serves every read and write. Remaining: mc mirror the live buckets, verify counts and sizes, then decide on cutover (decisions/storage-layout.md#rustfs-staged).

Two containers on homesrv had been dead for two weeks and now run. manga-fetch is the one /job/create needs. manga-web is what manga.kvmx.ru proxies to on 8083. Nothing watches them, and nothing watches the workers.

Open questions

Four Phase 1 items have no Vikunja task, because writing to the tracker was not asked for: the speaker contract fix, the verifier rules, the tracklet constraints, and the flag resolution path. Only [#203] existed and is closed by decisions/audit-phase1.md#unlocked-model-load.

Three audit items are deliberately not done and are recorded as caveats rather than silently dropped: honest stage clearing, ComfyUI under the session mutex, and reversible identity merges. Each needs a design decision, not a patch.

Carried over from the reconstruction: .venv needs the ROCm torch wheel reinstalled, and dots.tts/, legacy/, RESUME_SPEC.md, pipeline-design-notes.md, spec-v2.md are unrecoverable.