Files
manga-recap-pipeline/NEXT.md
T
kami a9d64fe80a Record panel 7 against the art, and the A/V gap it hid
Somebody watched chapter.mp4 for the first time. Two failures came out of it
that no stage counter could see.

chapter.mp4 is video 436.39s over audio 363.67s, so narration finishes 72.7s
before the picture. The 49 clips are clean: all 25fps, video and audio agree
to 0.03s, summing to 363.6s. A per-round probe puts the loss in the final
round of _assemble_batched, which turns 359s of video into 100s while the
audio survives. Round 0 is correct. Round 1 differs by holding a 7th input,
the leftover clip that skips encoding, so the tree mixes concat output, xfade
output and a raw clip. Not fixed.

worker_render.py gains an FPS constant, fps normalization in the xfade branch
to match concat, _stream_dur, and a self-check that compares video against
audio instead of asserting the file is non-empty. That old check is how a 20%
sync failure shipped. The fps inconsistency is real but not proven to be the
shipped cause. Pinning -r on the output was tried and reverted: it drops
frames to force CFR, which the concat branch comment already warned about.

Panel 7 checked against the art has zero correct identity bindings out of two,
and Seonho, the one character who matters, is unbound. bbox values are
consumed as absolute pixels; on a 900x1650 panel that puts all six boxes in
the top third, two inside a speech balloon. Identity therefore embeds crops of
balloon edges and window frames, which is how confidence 0.9 lands on the
wrong person. The colleague has no name in the story and was labelled Choi
Haeseon; that row holds 25 of 26 assignments, so it is the label the pipeline
stamps on any unnamed woman.

Four caveats added. Two earlier claims are withdrawn in place: rescaling bbox
by 1000 does not make the boxes correct, and the constraint is not 16 nameless
rows needing names. Cast profiles already exist, since all 53 rows populate
ref_image_uris and embedding_uri, but they are enrolled from the wrong crops.

worker_render.py self-check passes. No pipeline ran.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 01:07:17 +04:00

153 lines
9.7 KiB
Markdown

# NEXT
Updated 2026-08-12. What this session did is in `HANDOFF.md`.
## State
The chapter runs end to end and the output is **not watchable**. That is now measured, not guessed.
Job `778297bc-e7ce-439d-91b5-8a027060d17f`, chapter `7c944dd4-e972-42c7-ba60-9f6939548e80`, 116 panels,
`status=completed`, finished 2026-08-11T20:08:16Z. `s3://video/` holds 49 clips and a 50MiB
`chapter.mp4`. The user watched it and read out 19 defects. They are grouped by cause in `HANDOFF.md`.
Two numbers set the agenda:
- `chapter.mp4` is video 436.39s over audio 363.67s. The narration finishes 72.7s before the picture.
- Panel 7 checked against the art has **zero correct identity bindings** out of two, and the one
character who matters is unbound. `HANDOFF.md#panel-7-walked-against-the-art` has the table.
Uncommitted work sits in the tree: `worker_render.py` has an `FPS = 25` constant, fps normalization in
the xfade branch, a new `_stream_dur`, and a self-check that compares video against audio. The
self-check passes. It does **not** yet fix the chapter. Details and one dead end in `HANDOFF.md`.
## Next
1. **Fix chapter assembly.** `_assemble_batched` turns 359s of video into 100s while the audio survives.
It reproduces offline in two minutes, no GPU. Round 0 is correct and round 1 collapses. Round 1 is
the only round holding a raw clip that skipped encoding, so suspect the single-item passthrough
first. The recommended shape is one path, not three: normalize every input, then xfade every
boundary, treating `cut` as a 0.05s fade. `acrossfade` and `xfade` shorten audio and video equally,
so the streams stay locked. `HANDOFF.md` holds the per-round table, the filtergraph, and the repro
commands. Nothing downstream is worth judging until this lands.
2. **Fix identity, in this order.** Panel 7 is the worked example and
`HANDOFF.md#panel-7-walked-against-the-art` carries the evidence. Do not start at the registry.
a. **Settle the `bbox` coordinate space.** Consumed as pixels, all six boxes on panel 7 land in the
top third of the panel, two inside a speech balloon. Divided by 1000 they mostly land on their
subjects. `worker_vision.py:271` and `worker_identity.py:91` both assert pixels, and the art says
otherwise. Identity embeds `_crop_bbox(img, ch["bbox"])` at `worker_identity.py:200`, so today it
matches faces against crops of balloons and window frames. Nothing downstream can be judged until
this is right. Rescaling alone is **not** the fix: after scaling, one box still sits on an empty
window frame and another clips its subject.
b. **Let identity abstain and stay abstained.** The colleague has no name in the story and was
labelled `Choi Haeseon` at 0.9. An unnamed recurring person needs a stable anonymous identity so
narration says "the colleague" every time. `match()` already returns `None` below threshold. Check
whether the Tier-2 gemma resolver can answer "none of these"; that was not verified.
c. **Separate extra from cast.** Four of the six detections on panel 7 are background extras or
nothing at all, and all six reach identity as equal candidates.
d. Only then merge `Seonho` into `Lim Seonho` and split `character_afa7623b`, which still needs the
reversible-merge design (`caveats/audit-open.md#destructive-reconcile`), not a patch.
**Cast profiles already exist. Do not rebuild them.** The user asked whether the main cast could get a
profile built from reference frames and reused. `characters` already carries `ref_image_uris` and
`embedding_uri`, and all 53 rows have both populated. The mechanism is not missing, it is enrolled
from the wrong crops, so today it stores references to balloon edges and window frames. Step (a) is
what makes it work. Three things are genuinely absent and are the smaller follow-on:
- no quality gate on enrollment, so nothing checks that a reference crop holds a face at all
- nothing re-enrolls a reference set once it is written, so the wrong crops persist
**The visual "is this them?" check is already built. Do not write it again.** `/vision/resolve` at
`worker_vision.py:963` sends the query crop plus up to 3 labelled reference images per candidate.
`build_resolve_prompt` tells the model to judge face shape first, to treat hair and outfit as
secondary, that two people sharing a hair colour are not the same, and to answer `0` for NONE when
unsure. `choice: 0` becomes a new character, an out-of-range index becomes `unresolved`, and
`ref_image_uris` is republished as `reference_image_uris` at `worker_identity.py:152` and `:161`. The
mechanism, the prompt and the abstain path are all correct. They are fed crops of the wrong region,
which is step (a).
- no human gate to name, merge or split the clusters. The user wants this as a minor adjustment on
top, not as the mechanism. The `gates` table and the review gates from [#136] are the place to hang
it
The chibi at 1:35 will survive all of this. He genuinely is brown hair plus a yellow shirt, so a
profile match is correct on appearance and wrong on reality. That needs item 4 below, plus requiring
a real face before a crop can enroll.
3. **Stop the narration inventing facts.** 0:43, 2:03, 2:05 and 2:15 assert things no panel shows. The
correctness verifier passed 116/116 because it checks quotes and names, never invented claims.
4. **Teach vision that art inside a panel is not the scene.** A chibi on a monitor became "a man holding
a drink" at 1:35. A colleague pointing into the distance became "pointing towards the screen" at
1:59.
5. **`layers` writes nothing** and reports `completed 116/116`, so no clip has parallax and a still
holds for 28s from 2:24 (`caveats/audit-open.md#layers-writes-nothing`).
6. **Clear the stale job error.** The completed job still carries `error: "partial: 112/116 completed"`
(`caveats/audit-open.md#stale-job-error`).
7. Balloon-to-speaker geometry via the unused `det`/`seg` heads
(`caveats/speaker-attribution.md#tail-is-not-geometry`) is now behind item 2. With no name to attach,
geometry buys nothing.
8. Resolve a speaker answer across the whole dialogue window, not just the answering panel. The last 3
unresolved refs describe a neighbouring panel in the same 8-panel call.
9. Start Phase 2 from `ROADMAP.md`. Set SQLite `busy_timeout` before any concurrency work
(`caveats/audit-open.md#sqlite-locking`).
## Lesson worth keeping
Every metric recorded before this session said the pipeline was fine or nearly fine. `script` 116/116,
"9 named speech lines", `layers` 116/116, `assemble` 1/1. Watching two and a half minutes of output
found a 20% sync failure, a cast that is 84% anonymous, invented narration, and a stage that writes
nothing while reporting success. Stage counters measure whether code ran. They say nothing about whether
the result is correct. Watch the output before trusting a number.
## Running the pieces
```bash
./start_workers.sh # session_manager + 9 workers, each a uvicorn in a tmux window
tmux attach -t manga-workers # per-worker logs
.venv/bin/python worker_render.py # self-check, runs real ffmpeg, about 4 minutes
```
Read the state, or clear a stage and resume:
```bash
/usr/bin/ssh kami@192.168.1.104 "curl -s 'http://127.0.0.1:9090/job/status?job_id=778297bc-e7ce-439d-91b5-8a027060d17f'"
/usr/bin/ssh kami@192.168.1.104 "curl -s -X POST http://127.0.0.1:9090/stage/clear -H 'Content-Type: application/json' -d '{\"job_id\":\"778297bc-e7ce-439d-91b5-8a027060d17f\",\"stage\":\"<stage>\"}'"
/usr/bin/ssh kami@192.168.1.104 "curl -s -X POST 'http://127.0.0.1:9090/job/resume?job_id=778297bc-e7ce-439d-91b5-8a027060d17f'"
```
Traps: plain `ssh` is the kitty ssh kitten and refuses non-interactive stdin, so use `/usr/bin/ssh`.
`mc` aliases on homesrv are `homesrv` and `mio`. `local` returns Access Denied and `rfs` is the empty
rustfs. `cp` is aliased to `cp -i` and hangs on overwrite, so use `/usr/bin/cp -f`.
Re-fixing assembly needs the real clips, which the session scratchpad no longer holds:
```bash
/usr/bin/ssh kami@192.168.1.104 'P=homesrv/video/ef105a86-4b7e-4ac4-b45c-b7d83b8f5b5e/7c944dd4-e972-42c7-ba60-9f6939548e80; mc cp -q -r $P/clips/ /tmp/rclips/; cd /tmp/rclips && tar cf - .' | tar xf - -C clips/
```
## Storage and viewer, tasks #116/#117
[#117] is done. `stowage` serves the manga buckets. It was never a MinIO problem: the container had been
dead since 2026-07-19 on an arm64 digest pin.
[#116] is closer but not cut over. Artifacts split one bucket per class
(`decisions/storage-layout.md#bucket-per-artifact`), and both MinIO and `rustfs` hold all six buckets.
`rustfs` on `127.0.0.1:9010/9011` is still empty and nothing is repointed, so MinIO serves every read
and write. Remaining: `mc mirror` the live buckets, verify counts and sizes, then decide on cutover
(`decisions/storage-layout.md#rustfs-staged`).
Two containers on homesrv had been dead for two weeks and now run. `manga-fetch` is the one
`/job/create` needs. `manga-web` is what `manga.kvmx.ru` proxies to on 8083. Nothing watches them, and
nothing watches the workers.
## Open questions
Four Phase 1 items have no Vikunja task, because writing to the tracker was not asked for: the speaker
contract fix, the verifier rules, the tracklet constraints, and the flag resolution path. Only [#203]
existed and is closed by `decisions/audit-phase1.md#unlocked-model-load`.
Three audit items are deliberately not done and are recorded as caveats rather than silently dropped:
honest stage clearing, ComfyUI under the session mutex, and reversible identity merges. Each needs a
design decision, not a patch.
Carried over from the reconstruction: `.venv` needs the ROCm torch wheel reinstalled, and `dots.tts/`,
`legacy/`, `RESUME_SPEC.md`, `pipeline-design-notes.md`, `spec-v2.md` are unrecoverable.