Store and describe images through a shared media intake (#252)

Vision needs a second model this box does not have, so the shipped half is
the part that works without one: an image arrives, is sniffed, is stored
content-addressed, and is prepared for inference. The describing half is
written and tested against a fake server, and refuses any endpoint that is
not on this box.

internal/media is the intake all three senses share — hearing and speaker
recognition store their audio in the same place under the same retention.
Blobs stay out of the sqlite store; only the derived text becomes a note,
and only when the caller asks. Retention is enforced by an hourly prune
loop rather than by a comment.

The plan's RemoteProvider step is refused: no cloud model, inference stays
on the box, and vision.NewLocal validates that at construction.
This commit is contained in:
kami
2026-08-01 04:53:07 +04:00
parent 8d5e357b57
commit d92349ca6e
19 changed files with 2444 additions and 26 deletions
+101 -22
View File
@@ -1,27 +1,106 @@
# Plan: Vision — Image Understanding Capability
**Goal:** Maven can "see" — accept images (from mavweb upload, Telegram, or filesystem paths), run vision inference via a local or remote multimodal model, and answer questions about the image content or extract structured information.
**Goal:** Maven can "see" — accept images (from mavweb upload, Telegram, or filesystem paths), store them, run inference via a **local** multimodal model, and answer questions about the image content or extract text from it.
**Done when:**
- Vision model backend is configurable: local multimodal LLM (e.g., LLaVA, Qwen-VL via `llama-server` mmproj) or remote API
- `internal/vision/` package handles image preprocessing, model inference, result parsing
- Voice/text commands like "что на картинке?" or "прочитай текст с экрана" route to the vision handler
- Extracted information can be written as facts/notes through `ipc.CoreAPI`
- Telegram image messages are processed through the same pipeline
**Status (2026-08-01):** intake, storage, config seam and the provider are shipped. The
describing half is **BLOCKED on a model download** — see "What is blocked" below.
**Scope:**
- New `internal/vision/` package — image loader (Go stdlib `image` + `golang.org/x/image`), inference client
- New config block: `voice.vision` in `config.Config``{enabled, provider, model_path, mmproj_path, remote_url}`
- Router intent extension: new `IntentVision` or reuse `IntentQuery` with a vision flag
- Reuses `internal/llm.Client` for API-compatible backends (OpenAI-compatible vision API)
- Reuses `internal/ipc.CoreAPI` for writing extracted data
## What shipped
**Steps:**
1. Create `internal/vision/provider.go``Provider` interface with `Describe(image []byte, prompt string) (string, error)` and `ExtractText(image []byte) (string, error)`
2. Implement `LocalProvider` — spawns `llama-server` with mmproj, sends multimodal chat completion requests
3. Implement `RemoteProvider` — calls an OpenAI-compatible vision API endpoint, reuses `internal/llm.Client`
4. Create `internal/vision/processor.go` — image preprocessing (resize, format conversion to JPEG/PNG, base64 encoding)
5. Wire vision into `cmd/mavend/voice.go:reactiveHandler` — detect vision intent from router (new `IntentVision` or a `Slots.HasImage` flag)
6. Add IPC method `MethodDescribeImage` for programmatic access (mavweb upload, telegram bot)
7. Add vision config block to `config.Config` and wire in `cmd/mavend/main.go`
8. Test with a local multimodal model: send an image via mavweb, verify description and text extraction
| Piece | Where |
|---|---|
| Blob store (content-addressed, retention-pruned) | `internal/media/store.go` |
| Image decode / flatten / downscale / JPEG | `internal/media/image.go` |
| `Provider` seam + `Disabled` floor + `LocalProvider` | `internal/vision/vision.go` |
| Store-then-describe orchestration, re-runnable | `internal/vision/intake.go` |
| Config blocks `media` and `vision` | `internal/config/config.go` |
| IPC method `describe_image` (`AuthRead`) | `internal/ipc/{wire,api,client,server}.go`, `internal/auth/policy.go` |
| Daemon wiring + hourly retention prune | `cmd/mavend/vision.go` |
`internal/media` is deliberately shared: hearing (#253) and speaker recognition (#255) have
the same intake problem — a blob arrives, gets stored, gets described — and they store their
audio in the same place under the same retention.
## Design decisions worth knowing
**Store before describe.** `Intake.Accept` writes the blob to disk *first*, then asks the
model. If the model is missing or broken — which is this box's actual state — the answer is
"it's kept, I can't read it yet" with a content-addressed id, and `Intake.Rerun(id, question)`
describes it later. Nothing is lost to a missing model.
**No `RemoteProvider`.** The original step 3 called for "an OpenAI-compatible vision API
endpoint". Refused. The surviving hard constraint in CLAUDE.md after "never phones home" was
deprecated is *no cloud model, inference stays on the box*, and a photo of his flat is the
worst possible exception. `vision.NewLocal` therefore validates the endpoint at construction:
loopback, a private IP, or `localhost`. A hostname is refused too — it could resolve anywhere,
and resolving it would mean trusting DNS with his pictures.
**Blobs are not in the database.** The sqlite store is small, encrypted and read every tick;
a 40 MB blob has no business there. What lands in the database is the *text* the blob produced,
as an ordinary note (`source: media:image:<id-prefix>`), and only when the caller asks for it
(`save_note`). Glancing at a screenshot is not the same act as remembering it.
**Images are never search input and never embedded.** Only the derived description
participates in recall, and only after he can see it as a note.
**Retention is enforced by a loop, not by a promise.** `media.retention` defaults to 7 days
and `cmd/mavend` prunes hourly, starting at boot. A store that grows forever would be the real
failure mode of this capability.
**No webp.** The stdlib has no webp decoder and this repo takes no new dependencies (the box
is offline). `media.SniffImage` recognises webp well enough to refuse it *by name*, so the log
says "webp is not supported" instead of "not an image". Telegram sends webp for stickers; that
is a known gap, not a mystery.
**Text extraction is not a second method.** "прочитай текст с картинки" is a prompt. A VLM has
no separate OCR mode to select, and a second interface method would only duplicate the first.
## What is blocked, and on what
There is **no vision-capable gguf and no mmproj file on this box**. Checked 2026-08-01:
```
/mnt/hdd1/llms/{Bonsai,LFM2.5,llama3.2,ministral,nemotron3-nano,qwen3,qwen3.5}
```
— sixteen ggufs, all text-only, no `*mmproj*` anywhere. The resident Qwen3-1.7B is text-only
by construction, so vision needs a *second* model. The ≤1.7B ceiling in CLAUDE.md is about the
resident router/phraser, not about a second model loaded on demand — but iGPU VRAM still is,
so keep it small.
To unblock, download one pair to `/mnt/hdd1/llms/vision/` (bind-mounted to
`/opt/maven/models/llm`), a gguf **and** its mmproj:
- `Qwen2.5-VL-3B-Instruct` (Q4_K_M + `mmproj-F16.gguf`) — the safe default; reads Russian, and
its OCR is the best of this size class.
- `SmolVLM2-2.2B-Instruct` — smaller and faster, weaker at Cyrillic text in images.
- `moondream2` — smallest, English-only in practice. Do not bother, per the sub-500M lesson.
Then run a second llama-server on 8081 with `--mmproj`, point `vision.endpoint` at it, and
walk the QA steps on Vikunja #252.
## Config
```json
"media": { "dir": "media", "retention": "168h", "max_bytes": 67108864 },
"vision": {
"enabled": true,
"endpoint": "http://127.0.0.1:8081",
"model": "qwen2.5-vl-3b",
"max_dim": 896,
"max_tokens": 300,
"timeout": "90s"
}
```
Both absent by default. No `media` block ⇒ `describe_image` does not exist at all; a `media`
block with no `vision` block ⇒ images are stored and honestly not described.
## Still open
- **Router intent.** "что на картинке?" does not route anywhere yet. Adding an intent is
premature while nothing can answer it; the IPC method is the surface a Telegram photo or a
mavweb upload calls today.
- **Telegram photo path** in `mavpoll` (download the file, call `DescribeImage`).
- **mavweb upload page** and a `/media` listing so stored blobs are visible and deletable from
the authed surface.