Store and describe images through a shared media intake (#252)
Vision needs a second model this box does not have, so the shipped half is the part that works without one: an image arrives, is sniffed, is stored content-addressed, and is prepared for inference. The describing half is written and tested against a fake server, and refuses any endpoint that is not on this box. internal/media is the intake all three senses share — hearing and speaker recognition store their audio in the same place under the same retention. Blobs stay out of the sqlite store; only the derived text becomes a note, and only when the caller asks. Retention is enforced by an hourly prune loop rather than by a comment. The plan's RemoteProvider step is refused: no cloud model, inference stays on the box, and vision.NewLocal validates that at construction.
This commit is contained in:
+101
-22
@@ -1,27 +1,106 @@
|
||||
# Plan: Vision — Image Understanding Capability
|
||||
|
||||
**Goal:** Maven can "see" — accept images (from mavweb upload, Telegram, or filesystem paths), run vision inference via a local or remote multimodal model, and answer questions about the image content or extract structured information.
|
||||
**Goal:** Maven can "see" — accept images (from mavweb upload, Telegram, or filesystem paths), store them, run inference via a **local** multimodal model, and answer questions about the image content or extract text from it.
|
||||
|
||||
**Done when:**
|
||||
- Vision model backend is configurable: local multimodal LLM (e.g., LLaVA, Qwen-VL via `llama-server` mmproj) or remote API
|
||||
- `internal/vision/` package handles image preprocessing, model inference, result parsing
|
||||
- Voice/text commands like "что на картинке?" or "прочитай текст с экрана" route to the vision handler
|
||||
- Extracted information can be written as facts/notes through `ipc.CoreAPI`
|
||||
- Telegram image messages are processed through the same pipeline
|
||||
**Status (2026-08-01):** intake, storage, config seam and the provider are shipped. The
|
||||
describing half is **BLOCKED on a model download** — see "What is blocked" below.
|
||||
|
||||
**Scope:**
|
||||
- New `internal/vision/` package — image loader (Go stdlib `image` + `golang.org/x/image`), inference client
|
||||
- New config block: `voice.vision` in `config.Config` — `{enabled, provider, model_path, mmproj_path, remote_url}`
|
||||
- Router intent extension: new `IntentVision` or reuse `IntentQuery` with a vision flag
|
||||
- Reuses `internal/llm.Client` for API-compatible backends (OpenAI-compatible vision API)
|
||||
- Reuses `internal/ipc.CoreAPI` for writing extracted data
|
||||
## What shipped
|
||||
|
||||
**Steps:**
|
||||
1. Create `internal/vision/provider.go` — `Provider` interface with `Describe(image []byte, prompt string) (string, error)` and `ExtractText(image []byte) (string, error)`
|
||||
2. Implement `LocalProvider` — spawns `llama-server` with mmproj, sends multimodal chat completion requests
|
||||
3. Implement `RemoteProvider` — calls an OpenAI-compatible vision API endpoint, reuses `internal/llm.Client`
|
||||
4. Create `internal/vision/processor.go` — image preprocessing (resize, format conversion to JPEG/PNG, base64 encoding)
|
||||
5. Wire vision into `cmd/mavend/voice.go:reactiveHandler` — detect vision intent from router (new `IntentVision` or a `Slots.HasImage` flag)
|
||||
6. Add IPC method `MethodDescribeImage` for programmatic access (mavweb upload, telegram bot)
|
||||
7. Add vision config block to `config.Config` and wire in `cmd/mavend/main.go`
|
||||
8. Test with a local multimodal model: send an image via mavweb, verify description and text extraction
|
||||
| Piece | Where |
|
||||
|---|---|
|
||||
| Blob store (content-addressed, retention-pruned) | `internal/media/store.go` |
|
||||
| Image decode / flatten / downscale / JPEG | `internal/media/image.go` |
|
||||
| `Provider` seam + `Disabled` floor + `LocalProvider` | `internal/vision/vision.go` |
|
||||
| Store-then-describe orchestration, re-runnable | `internal/vision/intake.go` |
|
||||
| Config blocks `media` and `vision` | `internal/config/config.go` |
|
||||
| IPC method `describe_image` (`AuthRead`) | `internal/ipc/{wire,api,client,server}.go`, `internal/auth/policy.go` |
|
||||
| Daemon wiring + hourly retention prune | `cmd/mavend/vision.go` |
|
||||
|
||||
`internal/media` is deliberately shared: hearing (#253) and speaker recognition (#255) have
|
||||
the same intake problem — a blob arrives, gets stored, gets described — and they store their
|
||||
audio in the same place under the same retention.
|
||||
|
||||
## Design decisions worth knowing
|
||||
|
||||
**Store before describe.** `Intake.Accept` writes the blob to disk *first*, then asks the
|
||||
model. If the model is missing or broken — which is this box's actual state — the answer is
|
||||
"it's kept, I can't read it yet" with a content-addressed id, and `Intake.Rerun(id, question)`
|
||||
describes it later. Nothing is lost to a missing model.
|
||||
|
||||
**No `RemoteProvider`.** The original step 3 called for "an OpenAI-compatible vision API
|
||||
endpoint". Refused. The surviving hard constraint in CLAUDE.md after "never phones home" was
|
||||
deprecated is *no cloud model, inference stays on the box*, and a photo of his flat is the
|
||||
worst possible exception. `vision.NewLocal` therefore validates the endpoint at construction:
|
||||
loopback, a private IP, or `localhost`. A hostname is refused too — it could resolve anywhere,
|
||||
and resolving it would mean trusting DNS with his pictures.
|
||||
|
||||
**Blobs are not in the database.** The sqlite store is small, encrypted and read every tick;
|
||||
a 40 MB blob has no business there. What lands in the database is the *text* the blob produced,
|
||||
as an ordinary note (`source: media:image:<id-prefix>`), and only when the caller asks for it
|
||||
(`save_note`). Glancing at a screenshot is not the same act as remembering it.
|
||||
|
||||
**Images are never search input and never embedded.** Only the derived description
|
||||
participates in recall, and only after he can see it as a note.
|
||||
|
||||
**Retention is enforced by a loop, not by a promise.** `media.retention` defaults to 7 days
|
||||
and `cmd/mavend` prunes hourly, starting at boot. A store that grows forever would be the real
|
||||
failure mode of this capability.
|
||||
|
||||
**No webp.** The stdlib has no webp decoder and this repo takes no new dependencies (the box
|
||||
is offline). `media.SniffImage` recognises webp well enough to refuse it *by name*, so the log
|
||||
says "webp is not supported" instead of "not an image". Telegram sends webp for stickers; that
|
||||
is a known gap, not a mystery.
|
||||
|
||||
**Text extraction is not a second method.** "прочитай текст с картинки" is a prompt. A VLM has
|
||||
no separate OCR mode to select, and a second interface method would only duplicate the first.
|
||||
|
||||
## What is blocked, and on what
|
||||
|
||||
There is **no vision-capable gguf and no mmproj file on this box**. Checked 2026-08-01:
|
||||
|
||||
```
|
||||
/mnt/hdd1/llms/{Bonsai,LFM2.5,llama3.2,ministral,nemotron3-nano,qwen3,qwen3.5}
|
||||
```
|
||||
|
||||
— sixteen ggufs, all text-only, no `*mmproj*` anywhere. The resident Qwen3-1.7B is text-only
|
||||
by construction, so vision needs a *second* model. The ≤1.7B ceiling in CLAUDE.md is about the
|
||||
resident router/phraser, not about a second model loaded on demand — but iGPU VRAM still is,
|
||||
so keep it small.
|
||||
|
||||
To unblock, download one pair to `/mnt/hdd1/llms/vision/` (bind-mounted to
|
||||
`/opt/maven/models/llm`), a gguf **and** its mmproj:
|
||||
|
||||
- `Qwen2.5-VL-3B-Instruct` (Q4_K_M + `mmproj-F16.gguf`) — the safe default; reads Russian, and
|
||||
its OCR is the best of this size class.
|
||||
- `SmolVLM2-2.2B-Instruct` — smaller and faster, weaker at Cyrillic text in images.
|
||||
- `moondream2` — smallest, English-only in practice. Do not bother, per the sub-500M lesson.
|
||||
|
||||
Then run a second llama-server on 8081 with `--mmproj`, point `vision.endpoint` at it, and
|
||||
walk the QA steps on Vikunja #252.
|
||||
|
||||
## Config
|
||||
|
||||
```json
|
||||
"media": { "dir": "media", "retention": "168h", "max_bytes": 67108864 },
|
||||
"vision": {
|
||||
"enabled": true,
|
||||
"endpoint": "http://127.0.0.1:8081",
|
||||
"model": "qwen2.5-vl-3b",
|
||||
"max_dim": 896,
|
||||
"max_tokens": 300,
|
||||
"timeout": "90s"
|
||||
}
|
||||
```
|
||||
|
||||
Both absent by default. No `media` block ⇒ `describe_image` does not exist at all; a `media`
|
||||
block with no `vision` block ⇒ images are stored and honestly not described.
|
||||
|
||||
## Still open
|
||||
|
||||
- **Router intent.** "что на картинке?" does not route anywhere yet. Adding an intent is
|
||||
premature while nothing can answer it; the IPC method is the surface a Telegram photo or a
|
||||
mavweb upload calls today.
|
||||
- **Telegram photo path** in `mavpoll` (download the file, call `DescribeImage`).
|
||||
- **mavweb upload page** and a `/media` listing so stored blobs are visible and deletable from
|
||||
the authed surface.
|
||||
|
||||
Reference in New Issue
Block a user