Reconstruct repo from Claude Code + codex transcripts

Working tree (including .git) was lost to an rm. Rebuilt by replaying Write/Edit/
Read/attachment events from 25 Claude sessions and 22 successful codex apply_patch
blocks into one timestamp-ordered timeline.

Verified against ground truth recorded in the transcripts: wc -l on 10 files and
ls -l on 5 files at 2026-07-18T13:13:44Z both match exactly; 18 files are
byte-identical to their newest ~/.claude/file-history blob.

See HANDOFF.md for sources, gaps, and how to rebuild .venv.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-11 02:42:41 +04:00
commit ff6a512630
32 changed files with 6759 additions and 0 deletions
+248
View File
@@ -0,0 +1,248 @@
# Manga/Manhwa Recap Pipeline — Technical Spec
**Status:** draft v1
**Target:** fully local, single-GPU, resumable batch pipeline. No publishing (copyright out of scope).
**Core constraint:** no new PyTorch / ROCm wheels. Existing llama.cpp (gemma) stays. Everything else runs on `onnxruntime + opencv + numpy + scikit-learn`.
---
## 1. Goal
Given raw manga pages or manhwa strips for a series, automatically produce a narrated recap video with burned-in subtitles, with minimal human involvement (one-time character-bank setup per series).
This spec replaces a dependency on the `magi`/`magiv2` model by **decomposing its subtasks** into swappable, low-dependency components.
---
## 2. Design principles
- **Dependency isolation.** gemma runs on llama.cpp (unchanged). All CV/ML runs via ONNX Runtime. Detection models are small enough to run on CPU EP; ROCm EP optional. No `transformers`, no second torch install.
- **Specialists over generalist.** Do not ask one VLM to detect + identify + attribute. Each subtask is a dumb specialized component; a clean data bus (JSON files) connects them.
- **Resumable.** Every stage reads its input artifact from disk and writes its output artifact to disk. A stage is skipped if its output exists and is newer than its input (unless `--force`).
- **Supervised shortcut for identity.** Replace magi's unsupervised character clustering with a per-series character bank (exemplar crops + names). Simpler, more robust, and how magiv2 gets names anyway.
- **Fail loud, fail per-panel.** Low-confidence panels are flagged, not silently guessed. Confidence thresholds are config, not magic numbers in code.
---
## 3. Architecture overview
```
ingest → [1] panels → [2] detect → [3] identity → [4] text
[5] speaker-bind ←──────────┤
[6] filter essential ←──────┘
[7] transcript (ordered)
┌─────────────────────┴───────────────────┐
[8] scene-action (gemma) (transcript.json)
└─────────────────────┬───────────────────┘
[9] script gen (LLM, chapter + rolling summary)
[10] TTS (dots.tts) ──► [11] visual assembly (ffmpeg) ──► [12] subs (faster-whisper) ──► [13] mux
```
Stages 27 are the magi replacement. Stages 813 are the existing recap backend.
---
## 4. Filesystem layout
```
work/
<series>/
source/ # input pages/strips
ch<NN>/
panels/ # panel crops, ordered
panels.json # [1]
detections.json # [2]
identities.json # [3]
texts.json # [4]
transcript.json # [5][6][7] merged
script.json # [9]
audio/ # [10] wav per segment
video.mp4 # [13] final
bank/
bank.json # character bank (per series, hand-built once)
crops/ # exemplar images
```
Artifact-per-stage = resumability. Delete an artifact to re-run that stage forward.
---
## 5. Stage specs
### [1] Panel extraction
- **Manga (page-based):** `kumiko` → ordered panel polygons. Pure OpenCV.
- **Manhwa (vertical strip):** slice at horizontal whitespace bands. OpenCV: row-wise background uniformity → cut points. Produces pseudo-panels.
- **In:** `source/*` · **Out:** `panels/`, `panels.json`
- **Dep:** kumiko, opencv. No torch.
```jsonc
// panels.json
{
"type": "manga", // or "manhwa"
"reading_order": "rtl", // rtl | ltr | ttb
"panels": [
{ "id": "p001", "page": 1, "bbox": [x,y,w,h], "file": "panels/p001.png", "order": 0 }
]
}
```
### [2] Detection (text + balloons + characters)
- **Text + balloons:** `comic-text-detector` (ships ONNX). Returns text regions + balloon masks. This is what kills the OCR problem — you never OCR a full page again.
- **Character boxes:** YOLOv8 anime face/person model exported to ONNX. CPU-fine.
- **In:** `panels/` · **Out:** `detections.json`
- **Dep:** onnxruntime, opencv, numpy.
```jsonc
// detections.json (per panel)
{
"p001": {
"balloons": [
{ "id": "b0", "mask_poly": [[x,y],...], "centroid": [x,y], "bbox": [x,y,w,h] }
],
"text_regions": [
{ "id": "t0", "bbox": [x,y,w,h], "in_balloon": "b0" } // null if floating (sfx/sign)
],
"chars": [
{ "id": "c0", "bbox": [x,y,w,h], "crop": "..." }
]
}
}
```
### [3] Character identity (bank match)
- Embed every `chars[*]` crop with SigLIP or an anime ArcFace model (ONNX).
- Cosine-match against `bank.json` entries. Above threshold → assign name; else `"unknown"`.
- **Bank is built once per series by hand** (10 min): a few exemplar crops + a name each. This is the only required human touch.
- **In:** `detections.json`, `bank/` · **Out:** `identities.json`
- **Dep:** onnxruntime, numpy, sklearn (cosine / nearest-neighbour).
- **Note:** this is the accuracy-critical stage. Generic embeddings + same-face syndrome = the weakest link vs magi. The bank is what rescues it. Do NOT attempt unsupervised clustering as primary — it merges lookalikes.
```jsonc
// bank.json
{ "characters": [
{ "name": "Aria", "exemplars": ["bank/crops/aria_0.png","bank/crops/aria_1.png"] }
]}
// identities.json (per panel)
{ "p001": { "c0": { "name": "Aria", "score": 0.82 }, "c1": { "name": "unknown", "score": 0.41 } } }
```
### [4] Text extraction
- For each `in_balloon` text region: crop the balloon, feed the clean crop to gemma (llama.cpp) → read text. Isolated crops read far better than full pages.
- **In:** `detections.json`, `panels/` · **Out:** `texts.json`
- **Dep:** llama.cpp (existing). No torch.
```jsonc
// texts.json
{ "p001": { "t0": "We can't stay here." } }
```
### [5] Speaker binding
Two paths — start geometry, fall back to gemma on low confidence.
- **A — tail geometry (default, pure CV):** from the balloon mask, find the tail = sharpest protrusion off the centroid. Vector centroid→tip; nearest `chars[*]` box along that ray = speaker. Hand-rolled magiv2 tail logic.
- **B — set-of-mark + gemma (fallback):** draw numbered boxes on chars + balloons on the panel image, ask gemma "balloon 2 → face #3 or #5?". Grounding via drawn marks beats free-form spatial reasoning. No new deps.
- **Trigger fallback when:** >2 candidate chars, ambiguous/absent tail, or geometry confidence < threshold.
- **In:** `detections.json`, `identities.json`, `texts.json` · **Out:** merged into `transcript.json`.
### [6] Essential vs non-essential filter
- Pure geometry: `text_region.in_balloon != null` → dialogue. Floating on raw art → sfx/sign → **drop**. Kills "THUD" and street-sign garbage with no classifier.
### [7] Transcript assembly (reading order)
- Panel order from `panels.json`. Within panel: manga = sort (top → right-to-left); manhwa = top-down.
- Emit ordered speaker+line list.
```jsonc
// transcript.json
{ "chapter": 12, "lines": [
{ "panel": "p001", "speaker": "Aria", "line": "We can't stay here.", "conf": 0.79, "flagged": false }
]}
```
### [8] Scene-action description
- gemma describes physical action per panel ("she draws a sword") — the one thing magi does NOT do and gemma is good at. Runs in parallel with 27.
- **Out:** `scene` field per panel, merged for the script stage.
### [9] Script generation
- LLM input = **one chapter of transcript + scene-actions + a ~500-token rolling summary** of prior chapters. Do NOT token-max the context; attention degrades mid-window and coherence/attribution drop.
- Prompt: compress and narrate, not transcribe. Output narration segments each mapped to source panel IDs (needed for visual timing).
- **Out:** `script.json`
```jsonc
// script.json
{ "segments": [
{ "id": "s0", "text": "Cornered in the ruins, Aria makes her choice...", "panels": ["p001","p002"] }
]}
```
### [10] TTS
- `dots.tts`, wav per segment. Keep durations — they drive visual timing.
### [11] Visual assembly
- ffmpeg ken-burns (pan/zoom) per panel. Panel display time = proportional to its segment's audio length. moviepy to orchestrate or raw filtergraphs for lean/fast.
### [12] Subtitles
- `faster-whisper` on generated audio → timestamped SRT → burn in with ffmpeg. Easier than aligning from the script side.
### [13] Mux
- Combine video + audio + burned subs → `video.mp4`.
---
## 6. Dependency matrix
| Stage | Tooling | New torch? |
|---|---|---|
| 1 panels | kumiko, opencv | no |
| 2 detect | comic-text-detector (onnx), yolo-anime (onnx) | no |
| 3 identity | siglip/arcface (onnx), sklearn | no |
| 4 text | llama.cpp gemma | no |
| 5 bind | numpy geometry + gemma fallback | no |
| 6 filter | numpy geometry | no |
| 7 order | numpy | no |
| 8 scene | llama.cpp gemma | no |
| 9 script | local LLM | no |
| 10 tts | dots.tts | (its own env) |
| 11 video | ffmpeg, moviepy | no |
| 12 subs | faster-whisper | (ctranslate2, not torch) |
| 13 mux | ffmpeg | no |
Net: the magi replacement (27) adds **only ONNX Runtime + a couple of onnx model files**. No ROCm wheel churn.
---
## 7. Orchestration
- Bash driver calls python stage scripts; artifacts handed off as files.
- Each stage: `stage_N.py --series X --chapter NN [--force]`.
- Skip logic: if output exists and `mtime(output) > mtime(input)` and not `--force`, skip.
- Model loading: sequence stages so you never hold two large models in VRAM at once (gemma vs detectors vs faster-whisper). One card handles all, serially.
- Do NOT wrap in a Spring service. This is a batch job, not request/response — a service is pure added attack surface and state for zero benefit.
---
## 8. Hardware notes
- gemma (mmproj) is the VRAM heavyweight; detectors are CPU-viable.
- Sequence model loads; never co-resident. Target: single decent GPU, staged.
- Detection/embedding on CPU EP is fine and frees VRAM for gemma.
---
## 9. Known risks / caveats
- **Re-ID accuracy** is the weak link vs magi's end-to-end association. Mitigation: character bank (mandatory), tune cosine threshold per series, flag `unknown` rather than guess.
- **Manhwa layout:** comic-text-detector and the anime detectors are manga/anime-trained. Vertical webtoon art + non-japanese layout = degraded results. Expect tuning on the whitespace slicer and lower binding confidence.
- **Tail geometry** fails on off-panel speakers and thought bubbles. Fallback to gemma SoM; if still ambiguous, flag the line.
- **Rolling summary drift:** long series accumulate summary error. Periodically re-anchor the summary from a canonical synopsis if available.
---
## 10. Pre-build check (do this first)
Before building any of this: magi's HF weights may load through your **existing** ROCm torch directly, ignoring its pinned `requirements.txt` (the version pain is usually the wrapper deps, not torch). 5-minute test. If it instantiates and runs, you skip this entire rebuild. If it OOMs or the arch won't load, decompose per this spec.