8113bdfc8b
The rebuild after 1457556 came out byte-identical to the broken file,
which proved the xfade fix never runs for this chapter. An all-cut chapter
goes down the concat demuxer with -c copy, which writes the output in the
FIRST input's time_base and reinterprets every later packet in it. 14 of
49 clips are 30/1 at 1/15360 against 35 at 25/1 at 1/12800, so those 14
play 15360/12800 = 1.2 too long with their audio untouched. collage_cmd
hardcoded -r 30 and yesterday's FPS sweep missed it.
collage_cmd now emits -r FPS, and assemble probes r_frame_rate across the
clips and routes mixed rates through the re-encoding tree. Rebuilt
chapter.mp4 is 364.120s video against 364.122s audio at 25/1, from
436.392 over 363.675.
Also settle the bbox coordinate space, measured over all 113 detections:
47 boxes have x2 past the 900px panel width, none has y2 past 1000 on
panels up to 2307px tall, and the range is exactly [0, 1000]. It is
gemma's normalized grid, not pixels, whatever the prompt asks for.
/vision converts before returning, which fixes identity's crop, the gated
face pairing that was comparing pixel face boxes against grid boxes, the
set-of-mark boxes and the review UI at once. Checked by eye on panel 7:
five of six boxes now land on their subject, including the foreground
character who had no identity.
The registry still holds boxes and embeddings enrolled from the wrong
space. vision and identity have to re-run, which is GPU work and was not
started.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
309 lines
20 KiB
Markdown
309 lines
20 KiB
Markdown
# JOURNAL
|
|
|
|
Append-only, newest last. One block per session or run. Not a changelog: this records what happened on
|
|
the day a number was produced, so a later postmortem can find it.
|
|
|
|
## 2026-08-11 Audit second pass [no task]
|
|
|
|
Command: none. Source reading only.
|
|
Outcome: finished. `AUDIT.md` grew from 565 to 771 lines with a `## Second-pass findings` section:
|
|
4 new P0, 6 new P1, 13 P2, 5 additions to the Phase 1 list.
|
|
Produced: commit `6d9df5b`, `AUDIT.md:566`.
|
|
|
|
## 2026-08-11 Audit Phase 1 implemented [#203]
|
|
|
|
Command: `python worker_scene.py worker_script.py worker_vision.py session_manager.py`,
|
|
`pytest -q --ignore=test_api.py` in the orchestrator.
|
|
Outcome: finished. All self-checks pass, 108 orchestrator tests pass. No GPU work, no pipeline run.
|
|
`test_api.py` was skipped because fastapi is not installed in the workpc venv.
|
|
Produced: `decisions/audit-phase1.md`, `caveats/audit-open.md`, `ROADMAP.md`, and this scaffold.
|
|
|
|
One existing test asserted the bug: `test_name_binding.test_conflict_flags_and_stays_unnamed` relied on
|
|
orphan flags leaking into every chapter, because it never created panel rows. It now creates them.
|
|
|
|
## 2026-08-11 Orchestrator half committed and deployed [no task]
|
|
|
|
Command: `pytest -q --ignore=test_api.py`, then `docker compose up -d --build orchestrator` on homesrv.
|
|
Outcome: finished. 108 tests pass. Commits `1c60710` (orchestrator half) and `94bd4d8` (minio pin) in
|
|
`/mnt/server/home/kami/docker-apps`. Orchestrator and minio both answer health on homesrv.
|
|
|
|
The rebuild recreated `minio` as a side effect and it crash-looped with `exec format error`: the
|
|
compose pin was the arm64 manifest digest of `minio/minio:latest` and homesrv is amd64. Repinned to the
|
|
amd64 digest. Nothing about Phase 1 caused this, but any compose action that recreates minio would have
|
|
hit it, so it was latent, not new.
|
|
|
|
Still unrun against a real chapter.
|
|
|
|
## 2026-08-11 S3 viewer and storage swap, tasks #116/#117 [#116 #117]
|
|
|
|
Command: docker compose on homesrv, `dig`, `openssl s_client`. No pipeline, no GPU.
|
|
Outcome: partial. Viewer works, storage swap staged and unfinished.
|
|
|
|
#117 needed no new software. `stowage` at `~/docker-apps/stowage` was already configured against the
|
|
manga MinIO and had been dead since 2026-07-19 with `exec /sbin/tini: exec format error`: its digest
|
|
pin was the arm64 manifest. Repinned to amd64 `sha256:91be7f13`, chowned `data/` to uid 65532 for the
|
|
new image, and it serves. MinIO had the identical bug, repinned to `sha256:a1a8bd4a`. A sweep of all
|
|
470 local images on homesrv found exactly those two arm64; nothing else in the homelab is affected.
|
|
|
|
#116 is staged, not done. `rustfs` runs alongside MinIO on `127.0.0.1:9010/9011`, pinned
|
|
`sha256:19b105cc`, data at `/mnt/hdd2/rustfs`. Buckets are empty: the `mc` mirror of
|
|
`audio layers manga panels raw video` (350M, all in `manga`) has NOT run. `/mnt/hdd2/minio/data` is
|
|
untouched and is the rollback. RustFS is `1.0.0-beta.12`, labeled `build-type=prerelease`. Cutover
|
|
would give rustfs 9000/9001 and repoint `MINIO_ENDPOINT=minio:9000` in the orchestrator plus
|
|
`stowage/config.yaml`; `transport.py:95` needs no change if rustfs takes `192.168.1.104:9000`.
|
|
|
|
Side quest, unrelated to the pipeline: the shared 41-domain cert stopped renewing. Root cause was DNS,
|
|
not nginx. Every `*.kvmx.ru` name pointed at a hard A record for `109.229.102.117` while the line had
|
|
moved to `109.229.127.149`; the Mercusys DDNS at `kvmx-home.mercusysddns.com` was correct the whole
|
|
time but nothing in the zone referenced it. Fixed with `CNAME * -> kvmx-home.mercusysddns.com` at
|
|
reg.ru. Certificate now issues.
|
|
|
|
Two measurement traps worth remembering. The ISP transparently intercepts ports 80 and 443 by
|
|
Host/SNI, so `curl` from workpc to ANY address returns kvmx.ru content and proves nothing about
|
|
external reachability; bare TCP connects also succeed against arbitrary addresses and then hang. Three
|
|
wrong root causes came out of trusting those probes before checking them.
|
|
|
|
Also patched `~/scripts/migrate-kvmx-https.sh:54` on homesrv. `need_stream_module` used
|
|
`sudo -n nginx -V` and `sudo -n nginx -T`; the NOPASSWD rule covers only `nginx -t`, so it reported
|
|
"stream module is not loaded" whenever it meant "could not ask for a password". Both checks now run
|
|
without sudo. `bash -n` passes and both conditions evaluate true.
|
|
|
|
## 2026-08-11 Per-artifact buckets, rustfs buckets, baseline chapter run [#116]
|
|
|
|
Command: `mc mb` on rustfs, `docker compose up -d --build orchestrator`, `pytest -q --ignore=test_api.py`,
|
|
`./start_workers.sh`, then `/job/create` + `/stage/clear` + `/job/resume` for chapter
|
|
`7c944dd4-e972-42c7-ba60-9f6939548e80` of "Teto X Egen" as job `778297bc-e7ce-439d-91b5-8a027060d17f`.
|
|
Outcome: partial. Storage split landed and is proven by the run. The run itself was still in `direct`
|
|
when the session ended.
|
|
Produced: `decisions/storage-layout.md`, `caveats/speaker-attribution.md`, 109 orchestrator tests pass.
|
|
|
|
Artifacts now split one bucket per class instead of everything under `manga`
|
|
(`decisions/storage-layout.md#bucket-per-artifact`). Both MinIO and rustfs hold all six buckets. The
|
|
run put 79 pages in `raw` and 116 panel crops in `panels`, so the split works end to end.
|
|
|
|
Two containers on homesrv had been dead for two weeks and blocked the work. `manga-fetch` was exited,
|
|
so `/job/create` failed with `httpx.ConnectError`; `manga-web` was exited, so `manga.kvmx.ru` had
|
|
nothing behind it on port 8083. Both started with `docker compose up -d`. Neither is related to the
|
|
storage change. Neither was caught by any check, because nothing watches these containers.
|
|
|
|
Stage timings, 116 panels: crop 85s, vision ~4min, identity ~1min, reconcile ~7min for 35 pairs,
|
|
dialogue ~8min. Faster than the 2026-07-17 run at 75 panels. The webtoon crop that 500'd in July
|
|
succeeded this time.
|
|
|
|
Quality cross-check against the panel images, the point of the run. Dialogue text extraction is
|
|
accurate. Character detection is accurate. Speaker attribution is not: three of three sampled
|
|
two-character panels attribute both speakers to the wrong person, always swapped
|
|
(`caveats/speaker-attribution.md#tail-is-not-geometry`). 24 of 81 speech lines resolve to a named
|
|
character, which is the Phase 1 headline metric at 30%, and the sample says that 30% is not
|
|
trustworthy. 26 of 113 detected people got an identity, and 25 of those 26 went to one character that
|
|
turns out to cover two different women.
|
|
|
|
The run then reached `scene` 116/116 and failed in `script` at 87/116, not on OOM: 28 beats were
|
|
rejected by the script verifier as `unsupported-proper-noun: ['Choi', 'Haeseon']`
|
|
(`caveats/speaker-attribution.md#multiword-name-verifier`). No two-word cast name can pass that check.
|
|
|
|
## 2026-08-11 Speaker provenance and the multi-word cast name
|
|
|
|
Command: `.venv/bin/python worker_vision.py`, `pytest -q --ignore=test_api.py` in the orchestrator.
|
|
Outcome: both pass, 110 orchestrator tests. Nothing deployed, no GPU work, no pipeline run.
|
|
Produced: `decisions/speaker-attribution.md`, two caveats rewritten.
|
|
|
|
`_annotate_speaker_methods` stopped stamping `tail` on a model guess. With two or more characters
|
|
present the guess is dropped to `unknown` at confidence 0.0. With one present it is kept as
|
|
`model_solo` at 0.7, the same claim the solo backstop already makes
|
|
(`decisions/speaker-attribution.md#no-fake-tail`). Grounded `som_face` and `solo_prior` rows are
|
|
untouched. Nothing outside `worker_vision.py` reads the literal `tail`, checked across both repos.
|
|
|
|
`verify_script` now tokenizes each cast name into `allowed`, so `Choi Haeseon` passes as two tokens
|
|
(`decisions/speaker-attribution.md#multiword-cast-names`). That is the 28 beats job `778297bc` lost.
|
|
|
|
The remaining `['Blur']` beat is a true positive that still halts the whole chapter, now recorded as
|
|
`caveats/speaker-attribution.md#one-word-halts-chapter`.
|
|
|
|
Neither fix is live. The orchestrator container is not rebuilt and the workers are not restarted.
|
|
|
|
## 2026-08-11 Rerun from dialogue: the honest speaker number is 9%
|
|
|
|
Command: `/job/cancel`, `/stage/clear dialogue`, `./start_workers.sh`, `/job/resume` on job
|
|
`778297bc-e7ce-439d-91b5-8a027060d17f`, twice. `docker compose up -d --build orchestrator` three times.
|
|
Outcome: `dialogue` 116/116. 112 orchestrator tests pass, `worker_vision.py` self-check passes.
|
|
|
|
The named-speaker share is 9%, 9 of 95 speech lines, down from a reported 30% that counted fake tails.
|
|
Multi-character panels contribute 0 of 40 lines by design. Single-character panels give 9 of 55. All 9
|
|
binds are `Choi Haeseon`, the row that covers two different women.
|
|
|
|
Five defects, four of them found by measuring the run rather than by reading code.
|
|
|
|
1. The fake `tail` label, fixed before the run (`decisions/speaker-attribution.md#no-fake-tail`).
|
|
2. The multi-word cast name in the script verifier
|
|
(`decisions/speaker-attribution.md#multiword-cast-names`).
|
|
3. `/stage/clear dialogue` deleted nothing and reported success. dialogue and direct write onto the
|
|
per-panel vision blob and had no `_STAGE_TABLES` entry, so `run_stage_dialogue` saw
|
|
`"dialogue" in vision` and would have skipped all 116 panels. The proof is the second clear:
|
|
116 dialogue blobs and 75 direct blobs stripped that the first had left. This is
|
|
`caveats/audit-open.md#dishonest-clearing` firing exactly where it was filed.
|
|
4. gemma answers the speaker field with whatever the prompt showed, most often the character
|
|
description, and every such answer became a free-form name that no registry entry matched. 28 of 51
|
|
sampled lines (`decisions/speaker-attribution.md#prompt-label-answers`). After the fix, 3 of 95.
|
|
5. All 7 `som_face` lines pointed at a mark whose face paired to no present character, so the
|
|
highest-trust provenance sat on a line with no speaker. Same defect class as the fake tail.
|
|
|
|
Identity is now the binding constraint, not attribution. 26 of 113 detected people carry an identity,
|
|
23%, and 25 of the 26 are the one over-merged row. Even perfect balloon binding caps this chapter near
|
|
23% named. The person who does hold an identity is stored as `Lim Seonho` while a separate row is named
|
|
`Seonho` with alias `Lim Seonho`, so either name matches two rows, raises `ambiguous-speaker` and binds
|
|
nothing.
|
|
|
|
Two more defects surfaced after `dialogue` finished. All 7 `som_face` lines pointed at a mark whose face
|
|
paired to no present character (`decisions/speaker-attribution.md#unpaired-mark`). Then `script` halted
|
|
at 112/116 because the narrator wrote `"...Hm?"` for the source line `"Uh... hum...?"`, and both verifier
|
|
rules fired on that two-letter interjection
|
|
(`decisions/speaker-attribution.md#interjection-false-positive`). After the fix, `script` passed 116/116,
|
|
the first time this chapter has cleared the verifier. `tts` then ran for the first time.
|
|
|
|
The run then completed end to end for the first time: `tts` 116/116, `layers` 116/116, `render` 116/116,
|
|
`assemble` 1/1, finished 2026-08-11T20:08:16Z. `s3://video/` holds 49 clips and a 50MiB `chapter.mp4`,
|
|
`s3://audio/` 49 objects at 32MiB. The per-artifact bucket split is now proven for every class except
|
|
layers (`decisions/storage-layout.md#bucket-per-artifact`).
|
|
|
|
Two honesty defects surfaced at the finish, both recorded rather than fixed. `layers` reported
|
|
`completed 116/116` with an empty bucket, and the completed job still carries
|
|
`error: "partial: 112/116 completed"` from the failure three resumes earlier.
|
|
|
|
## 2026-08-12 — the video got watched
|
|
|
|
No pipeline ran. The user watched `chapter.mp4` for the first time and read out 19 timestamped defects.
|
|
That found more than the previous four sessions of measuring, because the recorded metrics were all
|
|
measuring whether code ran rather than whether the result was right.
|
|
|
|
Two measurements came out of it. First, `chapter.mp4` is video 436.39s over audio 363.67s, so the
|
|
narration finishes 72.7s before the picture and the gap accumulates. The 49 clips are clean: every one is
|
|
25fps exactly, video and audio agree to 0.03s, and they sum to 363.6s. Assembly adds 72.7s of video and
|
|
no audio. Second, this manga holds 19 character rows of which 3 carry a name, and `Choi Haeseon` holds 25
|
|
of the chapter's 26 identity assignments. That is why the video calls the colleague Choi, never names the
|
|
MC, and flips gender.
|
|
|
|
The A/V bug was narrowed with a per-round probe over the 49 real clips. Round 0 of `_assemble_batched` is
|
|
correct, losing only the xfade overlap per group. Round 1 turns 359s of video into 100s while the audio
|
|
survives at 358.79s. The round-1 filtergraph is arithmetically correct, and re-running the same chain by
|
|
hand over only the 6 encoded intermediates gives a correct 348.24s with no warnings. Round 1 differs by
|
|
holding a 7th input: the leftover 49th clip, which `_assemble_batched` passes through un-encoded. That
|
|
passthrough is a third path beside `concat` and `xfade` and is the prime suspect.
|
|
|
|
`worker_render.py` gained an `FPS = 25` constant, fps normalization in the xfade branch to match the
|
|
concat branch, a `_stream_dur` helper, and a self-check that compares video against audio rather than
|
|
asserting the file is non-empty. The old check only asserted `getsize(out) > 0`, which is how a 20% sync
|
|
failure shipped. Pinning `-r FPS` on the output encodes was tried and reverted: it collapsed the chapter
|
|
to exactly 100.00s by dropping frames to force CFR, which the comment at the concat branch already
|
|
warned about. None of it is committed and none of it fixes the chapter yet.
|
|
|
|
The fps inconsistency between the two branches is real but not proven to be the shipped cause. The scene
|
|
graphs hold 356 `cut` against 6 `fade_black`, so the real run's final round most likely stayed on the
|
|
concat branch where no mixing happens.
|
|
|
|
### Panel 7, checked against the art
|
|
|
|
The same day, the user pulled up panel 7 and checked every detection by eye. It overturned the framing
|
|
this file carried an hour earlier, and it overturned two theories I proposed before being corrected.
|
|
|
|
Panel `7c944dd4-e972-42c7-ba60-9f6939548e80_p007`, a wide establishing shot of an office through a
|
|
window, crop 900x1650. Vision emitted 6 characters. Zero of the two identity bindings are correct and the
|
|
one character who matters is unbound. `person_5`, described as "yellow sweater", is Seonho in the
|
|
foreground and got no identity. `person_6` is the colleague, who has no name in the story, and was
|
|
assigned `Choi Haeseon` at 0.9. `person_2` is a background extra and was assigned `Lim Seonho` at 0.9.
|
|
`person_1` is a window frame with nobody in it. `person_3` and `person_4` are background extras.
|
|
|
|
Three defects stack, recorded as `caveats/speaker-attribution.md#bbox-wrong-space`,
|
|
`#no-anonymous-identity` and `#extras-as-cast`. The `bbox` values are consumed as absolute pixels, and on
|
|
this panel that puts all six boxes in the top third with two inside a speech balloon. Divided by 1000
|
|
four of the six fit tightly. Identity therefore embedded crops of balloon edges and window frames, which
|
|
is how a 0.9 confidence lands on the wrong person. Blank crops embed alike, a plausible mechanism for one
|
|
row absorbing 25 of 26 assignments.
|
|
|
|
Two claims I made and had to withdraw. First, that rescaling by 1000 makes the boxes correct: after
|
|
scaling, `person_1` still sits on an empty window frame and `person_6` clips its subject, and the
|
|
descriptions are unreliable anyway, since `person_6` reads "white shirt" for a green dress. Second, that
|
|
the constraint is 16 nameless rows needing names. The opposite is true. The pipeline mints names onto
|
|
people who have none, and at least one nameless row is a real recurring person who should stay nameless.
|
|
|
|
The "26 of 113 detected people carry an identity" figure that framed the roadmap counted mostly
|
|
background extras. It should not be quoted again.
|
|
|
|
## 2026-08-12, chapter assembly, root cause and fix
|
|
|
|
Reproduced the A/V collapse offline with 49 synthetic clips at `ASSEMBLE_BATCH=8` and six `fade_black`
|
|
boundaries. It came out worse than the shipped run: **two round-0 groups of 8 fresh clips collapsed on
|
|
their own**, so the single-item passthrough theory from yesterday is dead
|
|
(`decisions/chapter-assembly.md#passthrough-innocent`).
|
|
|
|
Bisected one collapsing group by truncating the chain stage by stage:
|
|
|
|
```
|
|
k=7 out= 52.52 correct
|
|
k=8 out= 52.52 the last xfade contributed nothing
|
|
[v6][n7]xfade=duration=0.050:offset=52.500 <- [v6] is 52.52s long, 0.02s of margin
|
|
```
|
|
|
|
`_xfade_chain` took its durations from `_audio_dur`, which is `format=duration`, which is
|
|
`max(video, audio)`. Each clip's audio outlasts its video by about a frame, so the offset accumulator
|
|
crept ahead of the picture. Once the creep passed the transition width, xfade emitted the transition and
|
|
threw away the second input and every clip after it, at `rc 0` with nothing on stderr.
|
|
|
|
Fix: offsets come from `min(_stream_dur(v), _stream_dur(a))`, every input is floored to a whole frame
|
|
count and `trim`/`atrim`ed on both streams, and `_check_assembled` now verifies each encode against the
|
|
predicted timeline instead of trusting the exit code
|
|
(`decisions/chapter-assembly.md#offsets-from-min-stream`, `#check-assembled`).
|
|
|
|
Verified on the 49 real clips of chapter `7c944dd4`, re-downloaded from MinIO:
|
|
|
|
```
|
|
before r1 n=7 XFADE in v=359.29 a=359.60 -> out v= 99.96 a=358.79
|
|
after r1 n=7 XFADE in v=359.61 a=359.62 -> out v=358.76 a=358.76
|
|
chapter v=358.76 a=358.76 gap=+0.00 (shipped: v=436.39 a=363.67 gap=+72.72)
|
|
```
|
|
|
|
`worker_render.py` `__main__` passes. Two checks were added there, because the existing 4-clip A/V assert
|
|
passed all the way through the broken build. One asserts the frame-exact `trim` on both streams, one
|
|
assembles three clips whose audio outlasts their video by 0.4s. Mutation-tested by putting `_audio_dur`
|
|
back: the new check fires with `video=1.80 audio=3.56 expected=3.56`.
|
|
|
|
Not done: `s3://video/.../chapter.mp4` is still the broken 436s file. Rebuilding it means clearing the
|
|
`assemble` stage and resuming, which is CPU-only and was not run.
|
|
|
|
## 2026-08-12, the chapter rebuilt, and the bbox space settled
|
|
|
|
**The rebuild came out byte-identical to the broken file.** Clearing `assemble` and resuming produced
|
|
video 436.392031s over audio 363.674667s and `nb_frames` 9902 again, which proved the xfade fix committed
|
|
earlier today never runs for this chapter. With all-`cut` transitions `assemble` takes the `else` branch,
|
|
a `concat` demuxer with `-c copy`.
|
|
|
|
Reproduced that path offline in seconds and got the shipped numbers exactly. The cause is mixed frame
|
|
rates: 14 of the 49 clips are `r_frame_rate=30/1` at `time_base=1/15360`, the other 35 are `25/1` at
|
|
`1/12800`. `-c copy` writes the output in the first input's timebase, so those 14 play `15360/12800 = 1.2`
|
|
too long with their audio untouched. `collage_cmd` hardcoded `-r 30`, which yesterday's `FPS` sweep
|
|
missed. `decisions/chapter-assembly.md#mixed-rate-stream-copy`.
|
|
|
|
Fixed `collage_cmd` to emit `-r FPS`, and made `assemble` probe `r_frame_rate` across the clips and route
|
|
mixed rates through the re-encoding tree. Rebuilt:
|
|
|
|
```
|
|
before v=436.392 a=363.675 nb_frames=9902 avg_frame_rate=22.69
|
|
after v=364.120 a=364.122 nb_frames=9101 r=25/1
|
|
```
|
|
|
|
The 14 clips in the bucket are still 30fps. Assembly normalizes them, so the chapter is correct without
|
|
re-rendering, but the fast stream-copy path stays disabled for this chapter until `render` re-runs.
|
|
|
|
**The `bbox` space is 0-1000, not pixels.** Pulled all 113 detections from `/review/identity` and
|
|
measured: 47 boxes have `x2` past the 900px panel width, none has `y2` past 1000 on panels 1257 to 2307px
|
|
tall, 21 clamp at exactly 1000 in x, and the whole range is `[0, 1000]`. `/vision` now converts to pixels
|
|
before returning, so identity crops, gated face pairing, the set-of-mark boxes and the review UI all read
|
|
pixels (`decisions/identity-bbox.md#bbox-is-normalized`).
|
|
|
|
Checked by eye the way the user did. Drew the converted boxes on panel 7: five of six land on their
|
|
subject, including `person_5`, who is Seonho in the foreground with headphones and carried no identity.
|
|
`person_1` still frames an empty window mullion, which is the extra-versus-cast caveat, not this one.
|
|
|
|
Not done: `vision` and `identity` have not re-run, so every box, embedding and `ref_image_uris` in the
|
|
registry is still from the wrong space. That rerun is GPU work and was not started.
|