Store and describe images through a shared media intake (#252)

Vision needs a second model this box does not have, so the shipped half is
the part that works without one: an image arrives, is sniffed, is stored
content-addressed, and is prepared for inference. The describing half is
written and tested against a fake server, and refuses any endpoint that is
not on this box.

internal/media is the intake all three senses share — hearing and speaker
recognition store their audio in the same place under the same retention.
Blobs stay out of the sqlite store; only the derived text becomes a note,
and only when the caller asks. Retention is enforced by an hourly prune
loop rather than by a comment.

The plan's RemoteProvider step is refused: no cloud model, inference stays
on the box, and vision.NewLocal validates that at construction.
This commit is contained in:
kami
2026-08-01 04:53:07 +04:00
parent 8d5e357b57
commit d92349ca6e
19 changed files with 2444 additions and 26 deletions
+4
View File
@@ -333,6 +333,9 @@ func run(args []string) error {
if !locked {
wireMailIntake(srv, st, phr, cfg)
wireModelSwap(srv, phr, cfg)
// Vision + the media blob store (Vikunja #252). Both stay dark without a
// media block; MethodDescribeImage answers ErrUnknownMethod then.
wireVision(ctx, srv, st, embedderOf(voiceW), cfg)
}
// WrapKeyFn — wraps the env key with a passkey credential public key and
@@ -476,6 +479,7 @@ func run(args []string) error {
srv.Check = (&auth.Gate{Enrollment: auth.NewFloorEnrollment(), Session: passkeySess}).Check
wireMailIntake(srv, st, phr, cfg)
wireModelSwap(srv, phr, cfg)
wireVision(ctx, srv, st, embedderOf(voiceW), cfg)
// Start voice server.
if voiceW != nil {