Files
manga-recap-pipeline/NEXT.md
T
kami 97cb4831f9 Wire the caption merge, and write the target architecture down
merge_faceless_captions had been written and never called; both crop
endpoints called the non-destructive context_fragment_links instead, with no
decision recording that choice. Wiring it changes panel count and every panel
index, so the chapter needs a re-crop with the panels prefix cleared first --
crop_webtoon skips an upload when the key already exists, which is right for a
resume and silently wrong after a slicing change. Noted at the line.

It does not cover the head-in-one-shot body-in-the-next split that prompted
the question. _merge_plan only folds a fragment that has text and no face.

ARCHITECTURE.md is the target shape from the user's design, with what exists
against each section today. Nothing in it is built.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 00:39:57 +04:00

16 KiB

NEXT

Updated 2026-08-12 (sixth session). What the fifth session did is in HANDOFF.md, the runs are in JOURNAL.md.

State

The chapter runs end to end. The A/V sync defect is fixed and chapter.mp4 is rebuilt: video 364.120s against audio 364.122s at 25/1.

Job 778297bc-e7ce-439d-91b5-8a027060d17f, chapter 7c944dd4-e972-42c7-ba60-9f6939548e80, 116 panels. Four GPU cycles ran on 2026-08-12, the last 19:44-19:52 UTC on a reset registry. The job is running and parked at dialogue waiting. Everything from dialogue down is cleared and stale. All 9 workers are current.

Eleven changes are proven on real panels. This session added bbox corner ordering. It then deployed and ran the tracklet span cap, the shared GENERIC tokenizer and the cast gallery together.

The registry is now roughly right and is worth reading. Checked by eye, crop by crop, confirmed by the user. The lead holds 16 assignments. 14 are him and 2 are art inside a panel, the photograph at order 17 and the chibi at order 20. character_2b1b12a1 holds 13, all of them her, and she is a main character the registry never named. character_f0d4e901 holds 9, of which 7 are her and 2 are 2b1b12a1. On the 18:07 run the lead alone held 36, covering six men, a chibi and a cat.

So every wrong crop on the lead is one defect, item 7, and not a linker failure. The registry's other weakness is that its biggest character has no name.

metric 18:07 run 19:44 run
detections 119 119
assignments 60 60
coverage 57% 50%
tracklets over crops 12 / 64 33 / 72
lead's assignments 36 16
characters after reconcile 18 14

Coverage fell because gemma now clears 12 crops instead of naming them from a shortlist that did not contain the right person. Coverage is still the has_face gate plus those refusals, and nothing else.

Next

  1. Re-crop the chapter and run the fifth cycle. Four changes are written and tested since the fourth cycle, and none has touched a GPU.

    • merge_faceless_captions is wired into both crop endpoints. It was written, never called, and context_fragment_links was called instead. A stranded caption fragment now vstacks into the face-bearing fragment it belongs to, so panel count and every panel index change.
    • three naming fixes in db.add_name_claim (decisions/identity-naming.md): alias grouping, a caption as strong evidence, and a name held by another character refusing to promote onto a second one.

    Clear s3://panels/<manga>/<chapter>/panels/ before re-cropping. crop_webtoon skips the upload when the key exists, so a re-crop after a slicing change silently keeps the previous run's images. Wiring the merge is a slicing change. Everything downstream is invalidated by it, so this is a full re-run and not a stage rerun.

    Expected: fewer than 116 panels, 2b1b12a1 named Choi Haeseon from the p040 caption, the lead's conflicting-name-claims flag gone, and the green-dress woman no longer named Seonho but carrying a name-already-taken flag instead.

    Not fixed by any of it. Bare hair colour still links different men. Do not add a crop-to-crop cosine to close that. Measured on 22 embeddings, different people reach 0.93 and the same person reaches 0.96, so no threshold exists (caveats/audit-open.md#cosine-not-identity).

    Then, separately, test embedding the FACE box rather than the person box. face_detect already finds the face and pairs it for has_face. That is the likely root cause of cosine measuring scene instead of person. The test is to re-embed these same 22 detections and recompute the matrix.

1b. The head/body split that started the crop question is NOT fixed. The wired merge only folds a fragment that has text and no face. A body fragment carries no dialogue, so _merge_plan leaves it solo and it becomes its own panel and its own shot. Finding it needs a different signal, most likely a face touching the bottom edge of one fragment with a textless fragment below. No evidence has been gathered yet on how often this chapter does it.

1c. Three items from ARCHITECTURE.md are cheap against the current code. A plane field per detection for art-in-art, same-panel co-presence demoted to a weak cannot-link once the plane exists, and offscreen as a fourth speaker_ref kind.

  1. Order the corners in _bbox_to_pixels. Done 2026-08-12, run and verified on a GPU. 0 degenerate boxes over 119 detections (decisions/identity-bbox.md#bbox-corners-ordered).

  2. Vision boxes animals as people and dresses them. p081 and p108 are cats, described short brown / yellow plaid shirt and short brown / white t-shirt. A detection prompt problem, not a linker one, and it feeds every stage below.

  3. Detection still names from in-panel text and can hit the wrong body. p011 and p026 emit name: "Seonho" on person_2 with the roster hint gone. That is the residue of the naming defect and the reason the glasses man is named at all (decisions/identity-bbox.md#roster-does-not-name, "not covered").

  4. Decide what a back-turned character costs. This is now the whole of the coverage number. has_face gates 40% of detections. On the 18:07 run panel 7 lost 5 of its 6 detections to the gate, and its one face-bearing detection went unassigned. Panel 7's lead is at his desk from behind. face_detect finds one face on the whole panel at conf 0.599, and nothing else above 0.056 even at a 0.04 threshold. The gate is right and the detector is right. The lead is still unidentifiable there. Options are a body or head detector alongside the face one, or letting the tracklet carry identity across a back-turned panel. Do not lower FACE_CONF, the sweep shows nothing to find.

    Smaller follow-on: nine _audio_dur calls in worker_render.py measure finished clips with format=duration. So the durations reported to the orchestrator are blind to per-clip drift. They position no filter, so invariant 9 does not cover them. Worth converting to _stream_dur.

  5. Fix identity, in this order. All of 2a-2d below are done and now proven on a GPU. Kept for the evidence trail.

    a. Settle the bbox coordinate space. Done 2026-08-12, proven over all 113 detections and checked by eye on panel 7, where five of six converted boxes land on their subject (decisions/identity-bbox.md#bbox-is-normalized). The rerun settles the last doubt: person_1 now frames the lead at his desk, not the window mullion the pre-fix box caught. b. Let identity abstain and stay abstained. Done 2026-08-12, run and verified on a GPU (decisions/identity-bbox.md#none-mints-an-anonymous-character). The resolver could always answer "none of these". The orchestrator was discarding the answer: it read only character_id, so a deliberate NONE and a hallucinated index both unassigned every crop of the tracklet. A NONE now mints an anonymous character from the crop, using the embedding /identity/resolve ships beside it as emb_uri. Deployed: image rebuilt, resolve_outcome verified inside the container. c. Separate extra from cast. Done 2026-08-12, run and verified on a GPU (decisions/identity-bbox.md#face-gates-enrollment, decisions/identity-bbox.md#extras-gate-consumers). has_face stops a faceless detection enrolling, and two more consumers now skip it: worker_scene's actions, which is the script prompt's content and the verifier's evidence, and service._beat, the director's "who". _present_characters stays ungated on purpose, reasoned out in the decision. The remaining gap is that vision still emits extras into the blob, which is deliberate so the audit can see what was gated. d. Merge and split. Safety net done and now exercised. The 18:07 run merged 9 rows and stamped 22 assignments merged_from:, with no chain deeper than one hop. The non-destructive path and the flattening fix are both proven on real data.

    A merge no longer deletes the losing row: it sets merged_into, and stamps every repointed assignment with method = merged_from:<loser_id>. A wrong merge now costs a hand-written SQL walk, not a full rebaseline (caveats/audit-open.md#destructive-reconcile).

    Deliberately not built: the unmerge path and the split. No wrong merge has been observed since the crops were fixed, so the consumer of those records waits for one. The forward case is partly covered by 2b, because a resolver NONE now mints rather than folding a stranger into the nearest match. Splitting a character that is ALREADY over-merged still needs a re-embed pass over its detections.

    Cast profiles already exist. Do not rebuild them. The user asked whether the main cast could get a profile built from reference frames and reused. characters already carries ref_image_uris and embedding_uri, and every row has both populated (16 rows after the rerun, 53 before it). The mechanism was never missing. It was enrolled from the wrong crops, so it stored balloon edges and window frames. Step (a) fixed that, and the audit now reports 0 characters missing a ref crop. Three things are genuinely absent and are the smaller follow-on:

    • no quality gate on enrollment, so nothing checks that a reference crop holds a face at all
    • nothing re-enrolls a reference set once it is written, so the wrong crops persist

    The visual "is this them?" check is already built. Do not write it again. /vision/resolve at worker_vision.py:963 sends the query crop plus up to 3 labelled reference images per candidate. build_resolve_prompt tells the model to judge face shape first, to treat hair and outfit as secondary, that two people sharing a hair colour are not the same, and to answer 0 for NONE when unsure. choice: 0 becomes a new character, an out-of-range index becomes unresolved, and ref_image_uris is republished as reference_image_uris at worker_identity.py:152 and :161. The mechanism, the prompt and the abstain path are all correct. They are fed crops of the wrong region, which is step (a).

    • no human gate to name, merge or split the clusters. The user wants this as a minor adjustment on top, not as the mechanism. The gates table and the review gates from [#136] are the place to hang it

    The chibi at 1:35 will survive all of this. He genuinely is brown hair plus a yellow shirt, so a profile match is correct on appearance and wrong on reality. That needs item 4 below, plus requiring a real face before a crop can enroll.

  6. Stop the narration inventing facts. 0:43, 2:03, 2:05 and 2:15 assert things no panel shows. The correctness verifier passed 116/116 because it checks quotes and names, never invented claims.

  7. Teach vision that art inside a panel is not the scene. A chibi on a monitor became "a man holding a drink" at 1:35. A colleague pointing into the distance became "pointing towards the screen" at 1:59.

  8. layers writes nothing and reports completed 116/116, so no clip has parallax and a still holds for 28s from 2:24 (caveats/audit-open.md#layers-writes-nothing).

  9. Clear the stale job error. The completed job still carries error: "partial: 112/116 completed" (caveats/audit-open.md#stale-job-error).

  10. Balloon-to-speaker geometry via the unused det/seg heads (caveats/speaker-attribution.md#tail-is-not-geometry) is now behind item 5. With no name to attach, geometry buys nothing.

  11. Resolve a speaker answer across the whole dialogue window, not just the answering panel. The last 3 unresolved refs describe a neighbouring panel in the same 8-panel call.

  12. Start Phase 2 from ROADMAP.md. Set SQLite busy_timeout before any concurrency work (caveats/audit-open.md#sqlite-locking).

Lesson worth keeping

Every metric recorded before this session said the pipeline was fine or nearly fine. script 116/116, "9 named speech lines", layers 116/116, assemble 1/1. Watching two and a half minutes of output found a 20% sync failure, a cast that is 84% anonymous, invented narration, and a stage that writes nothing while reporting success. Stage counters measure whether code ran. They say nothing about whether the result is correct. Watch the output before trusting a number.

Running the pieces

./start_workers.sh              # session_manager + 9 workers, each a uvicorn in a tmux window
tmux attach -t manga-workers    # per-worker logs
.venv/bin/python worker_render.py    # self-check, runs real ffmpeg, about 4 minutes

Read the state, or clear a stage and resume:

/usr/bin/ssh kami@192.168.1.104 "curl -s 'http://127.0.0.1:9090/job/status?job_id=778297bc-e7ce-439d-91b5-8a027060d17f'"
/usr/bin/ssh kami@192.168.1.104 "curl -s -X POST http://127.0.0.1:9090/stage/clear -H 'Content-Type: application/json' -d '{\"job_id\":\"778297bc-e7ce-439d-91b5-8a027060d17f\",\"stage\":\"<stage>\"}'"
/usr/bin/ssh kami@192.168.1.104 "curl -s -X POST 'http://127.0.0.1:9090/job/resume?job_id=778297bc-e7ce-439d-91b5-8a027060d17f'"

Traps: tmux respawn-window -k does NOT re-run the window's command, it drops to a bare shell and the worker stays down. Restart one worker by re-sending the send-keys line from start_workers.sh, and confirm with /health plus ./check_stale.sh. docker compose up --build orchestrator recreates the container, which drops any docker cp'd script such as audit_registry.py. Plain ssh is the kitty ssh kitten and refuses non-interactive stdin, so use /usr/bin/ssh. mc aliases on homesrv are homesrv and mio. local returns Access Denied and rfs is the empty rustfs. cp is aliased to cp -i and hangs on overwrite, so use /usr/bin/cp -f.

Re-fixing assembly needs the real clips, which the session scratchpad no longer holds:

/usr/bin/ssh kami@192.168.1.104 'P=homesrv/video/ef105a86-4b7e-4ac4-b45c-b7d83b8f5b5e/7c944dd4-e972-42c7-ba60-9f6939548e80; mc cp -q -r $P/clips/ /tmp/rclips/; cd /tmp/rclips && tar cf - .' | tar xf - -C clips/

Storage and viewer, tasks #116/#117

[#117] is done. stowage serves the manga buckets. It was never a MinIO problem: the container had been dead since 2026-07-19 on an arm64 digest pin.

[#116] is closer but not cut over. Artifacts split one bucket per class (decisions/storage-layout.md#bucket-per-artifact), and both MinIO and rustfs hold all six buckets. rustfs on 127.0.0.1:9010/9011 is still empty and nothing is repointed, so MinIO serves every read and write. Remaining: mc mirror the live buckets, verify counts and sizes, then decide on cutover (decisions/storage-layout.md#rustfs-staged).

Two containers on homesrv had been dead for two weeks and now run. manga-fetch is the one /job/create needs. manga-web is what manga.kvmx.ru proxies to on 8083. Nothing watches them, and nothing watches the workers.

Open questions

Four Phase 1 items have no Vikunja task, because writing to the tracker was not asked for: the speaker contract fix, the verifier rules, the tracklet constraints, and the flag resolution path. Only [#203] existed and is closed by decisions/audit-phase1.md#unlocked-model-load.

Three audit items are deliberately not done and are recorded as caveats rather than silently dropped: honest stage clearing, ComfyUI under the session mutex, and reversible identity merges. Each needs a design decision, not a patch.

Carried over from the reconstruction: .venv needs the ROCm torch wheel reinstalled, and dots.tts/, legacy/, RESUME_SPEC.md, pipeline-design-notes.md, spec-v2.md are unrecoverable.