Files
manga-recap-pipeline/manga-recap-pipeline-spec.md
T
kami ff6a512630 Reconstruct repo from Claude Code + codex transcripts
Working tree (including .git) was lost to an rm. Rebuilt by replaying Write/Edit/
Read/attachment events from 25 Claude sessions and 22 successful codex apply_patch
blocks into one timestamp-ordered timeline.

Verified against ground truth recorded in the transcripts: wc -l on 10 files and
ls -l on 5 files at 2026-07-18T13:13:44Z both match exactly; 18 files are
byte-identical to their newest ~/.claude/file-history blob.

See HANDOFF.md for sources, gaps, and how to rebuild .venv.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 02:42:41 +04:00

249 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Manga/Manhwa Recap Pipeline — Technical Spec
**Status:** draft v1
**Target:** fully local, single-GPU, resumable batch pipeline. No publishing (copyright out of scope).
**Core constraint:** no new PyTorch / ROCm wheels. Existing llama.cpp (gemma) stays. Everything else runs on `onnxruntime + opencv + numpy + scikit-learn`.
---
## 1. Goal
Given raw manga pages or manhwa strips for a series, automatically produce a narrated recap video with burned-in subtitles, with minimal human involvement (one-time character-bank setup per series).
This spec replaces a dependency on the `magi`/`magiv2` model by **decomposing its subtasks** into swappable, low-dependency components.
---
## 2. Design principles
- **Dependency isolation.** gemma runs on llama.cpp (unchanged). All CV/ML runs via ONNX Runtime. Detection models are small enough to run on CPU EP; ROCm EP optional. No `transformers`, no second torch install.
- **Specialists over generalist.** Do not ask one VLM to detect + identify + attribute. Each subtask is a dumb specialized component; a clean data bus (JSON files) connects them.
- **Resumable.** Every stage reads its input artifact from disk and writes its output artifact to disk. A stage is skipped if its output exists and is newer than its input (unless `--force`).
- **Supervised shortcut for identity.** Replace magi's unsupervised character clustering with a per-series character bank (exemplar crops + names). Simpler, more robust, and how magiv2 gets names anyway.
- **Fail loud, fail per-panel.** Low-confidence panels are flagged, not silently guessed. Confidence thresholds are config, not magic numbers in code.
---
## 3. Architecture overview
```
ingest → [1] panels → [2] detect → [3] identity → [4] text
[5] speaker-bind ←──────────┤
[6] filter essential ←──────┘
[7] transcript (ordered)
┌─────────────────────┴───────────────────┐
[8] scene-action (gemma) (transcript.json)
└─────────────────────┬───────────────────┘
[9] script gen (LLM, chapter + rolling summary)
[10] TTS (dots.tts) ──► [11] visual assembly (ffmpeg) ──► [12] subs (faster-whisper) ──► [13] mux
```
Stages 27 are the magi replacement. Stages 813 are the existing recap backend.
---
## 4. Filesystem layout
```
work/
<series>/
source/ # input pages/strips
ch<NN>/
panels/ # panel crops, ordered
panels.json # [1]
detections.json # [2]
identities.json # [3]
texts.json # [4]
transcript.json # [5][6][7] merged
script.json # [9]
audio/ # [10] wav per segment
video.mp4 # [13] final
bank/
bank.json # character bank (per series, hand-built once)
crops/ # exemplar images
```
Artifact-per-stage = resumability. Delete an artifact to re-run that stage forward.
---
## 5. Stage specs
### [1] Panel extraction
- **Manga (page-based):** `kumiko` → ordered panel polygons. Pure OpenCV.
- **Manhwa (vertical strip):** slice at horizontal whitespace bands. OpenCV: row-wise background uniformity → cut points. Produces pseudo-panels.
- **In:** `source/*` · **Out:** `panels/`, `panels.json`
- **Dep:** kumiko, opencv. No torch.
```jsonc
// panels.json
{
"type": "manga", // or "manhwa"
"reading_order": "rtl", // rtl | ltr | ttb
"panels": [
{ "id": "p001", "page": 1, "bbox": [x,y,w,h], "file": "panels/p001.png", "order": 0 }
]
}
```
### [2] Detection (text + balloons + characters)
- **Text + balloons:** `comic-text-detector` (ships ONNX). Returns text regions + balloon masks. This is what kills the OCR problem — you never OCR a full page again.
- **Character boxes:** YOLOv8 anime face/person model exported to ONNX. CPU-fine.
- **In:** `panels/` · **Out:** `detections.json`
- **Dep:** onnxruntime, opencv, numpy.
```jsonc
// detections.json (per panel)
{
"p001": {
"balloons": [
{ "id": "b0", "mask_poly": [[x,y],...], "centroid": [x,y], "bbox": [x,y,w,h] }
],
"text_regions": [
{ "id": "t0", "bbox": [x,y,w,h], "in_balloon": "b0" } // null if floating (sfx/sign)
],
"chars": [
{ "id": "c0", "bbox": [x,y,w,h], "crop": "..." }
]
}
}
```
### [3] Character identity (bank match)
- Embed every `chars[*]` crop with SigLIP or an anime ArcFace model (ONNX).
- Cosine-match against `bank.json` entries. Above threshold → assign name; else `"unknown"`.
- **Bank is built once per series by hand** (10 min): a few exemplar crops + a name each. This is the only required human touch.
- **In:** `detections.json`, `bank/` · **Out:** `identities.json`
- **Dep:** onnxruntime, numpy, sklearn (cosine / nearest-neighbour).
- **Note:** this is the accuracy-critical stage. Generic embeddings + same-face syndrome = the weakest link vs magi. The bank is what rescues it. Do NOT attempt unsupervised clustering as primary — it merges lookalikes.
```jsonc
// bank.json
{ "characters": [
{ "name": "Aria", "exemplars": ["bank/crops/aria_0.png","bank/crops/aria_1.png"] }
]}
// identities.json (per panel)
{ "p001": { "c0": { "name": "Aria", "score": 0.82 }, "c1": { "name": "unknown", "score": 0.41 } } }
```
### [4] Text extraction
- For each `in_balloon` text region: crop the balloon, feed the clean crop to gemma (llama.cpp) → read text. Isolated crops read far better than full pages.
- **In:** `detections.json`, `panels/` · **Out:** `texts.json`
- **Dep:** llama.cpp (existing). No torch.
```jsonc
// texts.json
{ "p001": { "t0": "We can't stay here." } }
```
### [5] Speaker binding
Two paths — start geometry, fall back to gemma on low confidence.
- **A — tail geometry (default, pure CV):** from the balloon mask, find the tail = sharpest protrusion off the centroid. Vector centroid→tip; nearest `chars[*]` box along that ray = speaker. Hand-rolled magiv2 tail logic.
- **B — set-of-mark + gemma (fallback):** draw numbered boxes on chars + balloons on the panel image, ask gemma "balloon 2 → face #3 or #5?". Grounding via drawn marks beats free-form spatial reasoning. No new deps.
- **Trigger fallback when:** >2 candidate chars, ambiguous/absent tail, or geometry confidence < threshold.
- **In:** `detections.json`, `identities.json`, `texts.json` · **Out:** merged into `transcript.json`.
### [6] Essential vs non-essential filter
- Pure geometry: `text_region.in_balloon != null` → dialogue. Floating on raw art → sfx/sign → **drop**. Kills "THUD" and street-sign garbage with no classifier.
### [7] Transcript assembly (reading order)
- Panel order from `panels.json`. Within panel: manga = sort (top → right-to-left); manhwa = top-down.
- Emit ordered speaker+line list.
```jsonc
// transcript.json
{ "chapter": 12, "lines": [
{ "panel": "p001", "speaker": "Aria", "line": "We can't stay here.", "conf": 0.79, "flagged": false }
]}
```
### [8] Scene-action description
- gemma describes physical action per panel ("she draws a sword") — the one thing magi does NOT do and gemma is good at. Runs in parallel with 27.
- **Out:** `scene` field per panel, merged for the script stage.
### [9] Script generation
- LLM input = **one chapter of transcript + scene-actions + a ~500-token rolling summary** of prior chapters. Do NOT token-max the context; attention degrades mid-window and coherence/attribution drop.
- Prompt: compress and narrate, not transcribe. Output narration segments each mapped to source panel IDs (needed for visual timing).
- **Out:** `script.json`
```jsonc
// script.json
{ "segments": [
{ "id": "s0", "text": "Cornered in the ruins, Aria makes her choice...", "panels": ["p001","p002"] }
]}
```
### [10] TTS
- `dots.tts`, wav per segment. Keep durations — they drive visual timing.
### [11] Visual assembly
- ffmpeg ken-burns (pan/zoom) per panel. Panel display time = proportional to its segment's audio length. moviepy to orchestrate or raw filtergraphs for lean/fast.
### [12] Subtitles
- `faster-whisper` on generated audio → timestamped SRT → burn in with ffmpeg. Easier than aligning from the script side.
### [13] Mux
- Combine video + audio + burned subs → `video.mp4`.
---
## 6. Dependency matrix
| Stage | Tooling | New torch? |
|---|---|---|
| 1 panels | kumiko, opencv | no |
| 2 detect | comic-text-detector (onnx), yolo-anime (onnx) | no |
| 3 identity | siglip/arcface (onnx), sklearn | no |
| 4 text | llama.cpp gemma | no |
| 5 bind | numpy geometry + gemma fallback | no |
| 6 filter | numpy geometry | no |
| 7 order | numpy | no |
| 8 scene | llama.cpp gemma | no |
| 9 script | local LLM | no |
| 10 tts | dots.tts | (its own env) |
| 11 video | ffmpeg, moviepy | no |
| 12 subs | faster-whisper | (ctranslate2, not torch) |
| 13 mux | ffmpeg | no |
Net: the magi replacement (27) adds **only ONNX Runtime + a couple of onnx model files**. No ROCm wheel churn.
---
## 7. Orchestration
- Bash driver calls python stage scripts; artifacts handed off as files.
- Each stage: `stage_N.py --series X --chapter NN [--force]`.
- Skip logic: if output exists and `mtime(output) > mtime(input)` and not `--force`, skip.
- Model loading: sequence stages so you never hold two large models in VRAM at once (gemma vs detectors vs faster-whisper). One card handles all, serially.
- Do NOT wrap in a Spring service. This is a batch job, not request/response — a service is pure added attack surface and state for zero benefit.
---
## 8. Hardware notes
- gemma (mmproj) is the VRAM heavyweight; detectors are CPU-viable.
- Sequence model loads; never co-resident. Target: single decent GPU, staged.
- Detection/embedding on CPU EP is fine and frees VRAM for gemma.
---
## 9. Known risks / caveats
- **Re-ID accuracy** is the weak link vs magi's end-to-end association. Mitigation: character bank (mandatory), tune cosine threshold per series, flag `unknown` rather than guess.
- **Manhwa layout:** comic-text-detector and the anime detectors are manga/anime-trained. Vertical webtoon art + non-japanese layout = degraded results. Expect tuning on the whitespace slicer and lower binding confidence.
- **Tail geometry** fails on off-panel speakers and thought bubbles. Fallback to gemma SoM; if still ambiguous, flag the line.
- **Rolling summary drift:** long series accumulate summary error. Periodically re-anchor the summary from a canonical synopsis if available.
---
## 10. Pre-build check (do this first)
Before building any of this: magi's HF weights may load through your **existing** ROCm torch directly, ignoring its pinned `requirements.txt` (the version pain is usually the wrapper deps, not torch). 5-minute test. If it instantiates and runs, you skip this entire rebuild. If it OOMs or the arch won't load, decompose per this spec.