Files
Maven/docs/plans/07-vision.md
T
kami d92349ca6e Store and describe images through a shared media intake (#252)
Vision needs a second model this box does not have, so the shipped half is
the part that works without one: an image arrives, is sniffed, is stored
content-addressed, and is prepared for inference. The describing half is
written and tested against a fake server, and refuses any endpoint that is
not on this box.

internal/media is the intake all three senses share — hearing and speaker
recognition store their audio in the same place under the same retention.
Blobs stay out of the sqlite store; only the derived text becomes a note,
and only when the caller asks. Retention is enforced by an hourly prune
loop rather than by a comment.

The plan's RemoteProvider step is refused: no cloud model, inference stays
on the box, and vision.NewLocal validates that at construction.
2026-08-01 04:53:07 +04:00

107 lines
5.3 KiB
Markdown

# Plan: Vision — Image Understanding Capability
**Goal:** Maven can "see" — accept images (from mavweb upload, Telegram, or filesystem paths), store them, run inference via a **local** multimodal model, and answer questions about the image content or extract text from it.
**Status (2026-08-01):** intake, storage, config seam and the provider are shipped. The
describing half is **BLOCKED on a model download** — see "What is blocked" below.
## What shipped
| Piece | Where |
|---|---|
| Blob store (content-addressed, retention-pruned) | `internal/media/store.go` |
| Image decode / flatten / downscale / JPEG | `internal/media/image.go` |
| `Provider` seam + `Disabled` floor + `LocalProvider` | `internal/vision/vision.go` |
| Store-then-describe orchestration, re-runnable | `internal/vision/intake.go` |
| Config blocks `media` and `vision` | `internal/config/config.go` |
| IPC method `describe_image` (`AuthRead`) | `internal/ipc/{wire,api,client,server}.go`, `internal/auth/policy.go` |
| Daemon wiring + hourly retention prune | `cmd/mavend/vision.go` |
`internal/media` is deliberately shared: hearing (#253) and speaker recognition (#255) have
the same intake problem — a blob arrives, gets stored, gets described — and they store their
audio in the same place under the same retention.
## Design decisions worth knowing
**Store before describe.** `Intake.Accept` writes the blob to disk *first*, then asks the
model. If the model is missing or broken — which is this box's actual state — the answer is
"it's kept, I can't read it yet" with a content-addressed id, and `Intake.Rerun(id, question)`
describes it later. Nothing is lost to a missing model.
**No `RemoteProvider`.** The original step 3 called for "an OpenAI-compatible vision API
endpoint". Refused. The surviving hard constraint in CLAUDE.md after "never phones home" was
deprecated is *no cloud model, inference stays on the box*, and a photo of his flat is the
worst possible exception. `vision.NewLocal` therefore validates the endpoint at construction:
loopback, a private IP, or `localhost`. A hostname is refused too — it could resolve anywhere,
and resolving it would mean trusting DNS with his pictures.
**Blobs are not in the database.** The sqlite store is small, encrypted and read every tick;
a 40 MB blob has no business there. What lands in the database is the *text* the blob produced,
as an ordinary note (`source: media:image:<id-prefix>`), and only when the caller asks for it
(`save_note`). Glancing at a screenshot is not the same act as remembering it.
**Images are never search input and never embedded.** Only the derived description
participates in recall, and only after he can see it as a note.
**Retention is enforced by a loop, not by a promise.** `media.retention` defaults to 7 days
and `cmd/mavend` prunes hourly, starting at boot. A store that grows forever would be the real
failure mode of this capability.
**No webp.** The stdlib has no webp decoder and this repo takes no new dependencies (the box
is offline). `media.SniffImage` recognises webp well enough to refuse it *by name*, so the log
says "webp is not supported" instead of "not an image". Telegram sends webp for stickers; that
is a known gap, not a mystery.
**Text extraction is not a second method.** "прочитай текст с картинки" is a prompt. A VLM has
no separate OCR mode to select, and a second interface method would only duplicate the first.
## What is blocked, and on what
There is **no vision-capable gguf and no mmproj file on this box**. Checked 2026-08-01:
```
/mnt/hdd1/llms/{Bonsai,LFM2.5,llama3.2,ministral,nemotron3-nano,qwen3,qwen3.5}
```
— sixteen ggufs, all text-only, no `*mmproj*` anywhere. The resident Qwen3-1.7B is text-only
by construction, so vision needs a *second* model. The ≤1.7B ceiling in CLAUDE.md is about the
resident router/phraser, not about a second model loaded on demand — but iGPU VRAM still is,
so keep it small.
To unblock, download one pair to `/mnt/hdd1/llms/vision/` (bind-mounted to
`/opt/maven/models/llm`), a gguf **and** its mmproj:
- `Qwen2.5-VL-3B-Instruct` (Q4_K_M + `mmproj-F16.gguf`) — the safe default; reads Russian, and
its OCR is the best of this size class.
- `SmolVLM2-2.2B-Instruct` — smaller and faster, weaker at Cyrillic text in images.
- `moondream2` — smallest, English-only in practice. Do not bother, per the sub-500M lesson.
Then run a second llama-server on 8081 with `--mmproj`, point `vision.endpoint` at it, and
walk the QA steps on Vikunja #252.
## Config
```json
"media": { "dir": "media", "retention": "168h", "max_bytes": 67108864 },
"vision": {
"enabled": true,
"endpoint": "http://127.0.0.1:8081",
"model": "qwen2.5-vl-3b",
"max_dim": 896,
"max_tokens": 300,
"timeout": "90s"
}
```
Both absent by default. No `media` block ⇒ `describe_image` does not exist at all; a `media`
block with no `vision` block ⇒ images are stored and honestly not described.
## Still open
- **Router intent.** "что на картинке?" does not route anywhere yet. Adding an intent is
premature while nothing can answer it; the IPC method is the surface a Telegram photo or a
mavweb upload calls today.
- **Telegram photo path** in `mavpoll` (download the file, call `DescribeImage`).
- **mavweb upload page** and a `/media` listing so stored blobs are visible and deletable from
the authed surface.