Compare commits
20 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| af4eeceb6a | |||
| 7b507dec94 | |||
| 2c0334c4fe | |||
| 92cbdbfdd3 | |||
| e78b2d8992 | |||
| 9d58922462 | |||
| b5ac48c126 | |||
| 69d0f5ee78 | |||
| 661b5c1099 | |||
| ff70637a0d | |||
| 06c1cf247e | |||
| 400653810e | |||
| b3936348f5 | |||
| c61b0b3968 | |||
| 0a5211b038 | |||
| 38be702188 | |||
| d42372e996 | |||
| 45231ba69e | |||
| 42c7b8b927 | |||
| e5a1db995d |
@@ -53,7 +53,7 @@ CGO daemons (`mavend`, `mavsttd`, `mavttsd`, `mavenclient`) need the vendored to
|
||||
and libs wired through the Makefile — **do not** call `go build` on them bare, use `make`:
|
||||
|
||||
```sh
|
||||
make build # all 9 binaries
|
||||
make build # all 11 binaries
|
||||
make build-web # single daemon (pure-Go ones: web/waked/poll/caldav build without CGO)
|
||||
make test # go test -race across ./internal/... ./cmd/... with CGO env set
|
||||
```
|
||||
@@ -82,13 +82,31 @@ Pure-Go packages (`router`, `memory`, `mavweb`, …) run under a plain `go test
|
||||
| `mavpoll` | Environment poller: netdata alarms, uptime-kuma, zenmoney, wireguard presence. Writes facts, sends nothing. Telegram is `internal/delivery/telegramsink`, not this. |
|
||||
| `mavcaldav` | CalDAV calendar sync. |
|
||||
| `mavmaild` | Mail reader (IMAP, read-only). Holds the IMAP password; core never sees it. |
|
||||
| `mavgpud` | GPU supervisor. **Runs on workpc, not homesrv** — own unit, `deploy/mavgpud.service`. Keeps llama-server loaded while the card is free (V-488). Maven never asks it for anything, it reads `/health` through `llm.Pair`. |
|
||||
| `mavupdate` | Not a daemon. Operator CLI a human runs on the box to deploy a new build. |
|
||||
|
||||
Two more binaries have no Makefile target and are built with `go run` or `go build` when
|
||||
they are needed. Neither is deployed.
|
||||
|
||||
| Binary | Role |
|
||||
|---|---|
|
||||
| `mavseal` | Recovery tool. Encrypts a live tmpfs working copy back to the ciphertext file when mavend was killed before `defer st.Close()` sealed it. |
|
||||
| `labelgen` | Runs the stage 0 grammars over utterances and prints JSONL, the training data for the routing heads (V-546). |
|
||||
|
||||
Daemons are wired socket-to-socket, not linked. `internal/ipc` is the client/server wire
|
||||
protocol; the config in `deploy/mavend.json` (with `${VAR}` env expansion from gitignored
|
||||
`deploy/telegram.env`) sets socket paths, model paths, and the phraser/embedder blocks.
|
||||
|
||||
**Seven of the nine run on homesrv. `mavwaked` and `mavenclient` do not, and that is the
|
||||
decision, not an oversight** (Vikunja #463, `docs/plans/17-where-the-voice-loop-runs.md`).
|
||||
**`docker-compose.yml` runs five: `mavend`, `mavsttd`, `mavttsd`, `mavweb`, `mavpoll`.**
|
||||
Count against compose, not against the table. Four of the nine daemons are absent, and each
|
||||
absence has a different reason.
|
||||
|
||||
`mavmaild` is commented out in compose, with the reason written beside it: it needs a mail
|
||||
account and this box has none. `mavcaldav` appears nowhere at all, and unlike the other
|
||||
three that is an oversight rather than a decision (V-644).
|
||||
|
||||
**`mavwaked` and `mavenclient` are absent by decision, not oversight** (Vikunja #463,
|
||||
`docs/plans/17-where-the-voice-loop-runs.md`).
|
||||
homesrv has a microphone — it is a laptop — but it is in the wrong room, so a wake-word
|
||||
daemon there listens to nobody. They belong on a client machine where the owner is standing.
|
||||
|
||||
@@ -305,9 +323,20 @@ not a transcript. The transcript still expires. The gesture that writes one is
|
||||
two buttons beside the reply on `/chat`, reached over `ipc.CorrectTurn` and the
|
||||
trace id that now rides back on `ipc.ChatReply`. A turn marked wrong with no
|
||||
target is a usable negative, so naming the intent is never required. The target
|
||||
is one of the seven intents and never free text. Only `/chat` offers it: the wire
|
||||
op assumes no browser, but telegram and voice do not call it yet, and
|
||||
`docs/plans/22-correcting-a-turn.md` says why voice is the hard one. Adding a rung to the ladder
|
||||
is one of the seven intents and never free text. **All three reaches offer it as
|
||||
of 06-08-2026**, and this section used to say only `/chat` did. Voice is the
|
||||
`repair` rung, which has read spoken corrections since V-455 and now writes the
|
||||
durable label beside the classifier seed it always wrote; a spoken negative with
|
||||
no target is its own rung, `repair-negative` (V-636, `docs/plans/22-correcting-a-turn.md`).
|
||||
Telegram is an inline keyboard under the reply, and it needed the chat to become
|
||||
readable first — **telegram is no longer outbound only** (V-637,
|
||||
`docs/plans/23-inbound-telegram.md`). The poller is dark unless the `telegram`
|
||||
block says `intake`, it long-polls because the box takes no inbound connections,
|
||||
it accepts `chat_id` and no other sender, and it drops whatever queued while the
|
||||
daemon was down. It reaches the daemon through `ipc.CoreAPI` alone, so a chat
|
||||
turn takes the path `POST /api/chat` takes. Note that the turn source is still
|
||||
`tap:text` for both, so provenance cannot tell a chat turn from a typed one.
|
||||
Adding a rung to the ladder
|
||||
in `runTurn` means adding its name to `preRouteLadder` in
|
||||
`cmd/mavend/decisiontrace.go`, or that rung is silently missing from the record.
|
||||
|
||||
|
||||
@@ -104,4 +104,3 @@ func (h *reactiveHandler) actionAct(ctx context.Context, dec router.Decision) st
|
||||
}
|
||||
return phraser.A(phraser.ActDone, nil)
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,120 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"log"
|
||||
"net"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/event"
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/store"
|
||||
)
|
||||
|
||||
// The two boot paths meet here. run() wires the daemon twice: once at boot
|
||||
// when a key is in the environment, and once inside UnlockFn after a passkey
|
||||
// assertion, minutes or days later. Listing the same wiring in both places is
|
||||
// what let them drift — seven workers started untracked on the unlock path and
|
||||
// two daemonAPI fields were never set there, silently, for as long as anyone
|
||||
// had been cold-starting (V-639).
|
||||
//
|
||||
// So both paths call newDaemonAPI and startBackground and nothing else. A
|
||||
// field or a worker added later reaches both paths or neither.
|
||||
|
||||
// bootDeps is everything the two constructors below read. It is filled from
|
||||
// the same variables on both paths, by depsNow in run().
|
||||
type bootDeps struct {
|
||||
coreFor func() ipc.CoreAPI
|
||||
tl *tickLoop
|
||||
evBus *event.Bus
|
||||
voiceW *voiceWiring
|
||||
st *store.Store
|
||||
factWorker *factEnrichmentWorker
|
||||
evalWorker *memoryEvalWorker // nil ⇒ memory evaluation off (the default)
|
||||
feedWkr *feedWorker // nil ⇒ no feed is read (the default)
|
||||
crawlWkr *crawlWorker // nil ⇒ no page is watched (the default)
|
||||
}
|
||||
|
||||
// newDaemonAPI builds the real CoreAPI, with every field set. The unlock path
|
||||
// used to leave nexus and getMCPServers nil, so after a cold start
|
||||
// ResolveEntity refused with a nexus block configured and /tools rendered
|
||||
// "not configured" with an mcp block configured. Empty is a wrong answer
|
||||
// there, not a degraded one.
|
||||
func newDaemonAPI(d bootDeps) *daemonAPI {
|
||||
api := &daemonAPI{
|
||||
CoreAPI: d.coreFor(),
|
||||
getTrace: d.tl.trace,
|
||||
getMorningStatus: func(ctx context.Context) []ipc.MorningRoutineStatus { return d.tl.morningStatus(ctx, time.Now()) },
|
||||
getDayPlan: func(ctx context.Context) ipc.DayPlan { return d.tl.dayPlan(ctx, time.Now()) },
|
||||
getEvents: intakeEventsFn(d.evBus),
|
||||
getDecisions: turnDecisionsFn(d.voiceW),
|
||||
seedStore: seedStoreIfAllowed(d.st),
|
||||
nexus: nexusOf(d.voiceW),
|
||||
}
|
||||
if d.voiceW != nil && d.voiceW.handler != nil {
|
||||
api.chatFn = d.voiceW.handler.handleText
|
||||
// And the reverse: the handler was wired with the bare store adapter,
|
||||
// which cannot serve the day plan. See upgradeAPI.
|
||||
d.voiceW.handler.upgradeAPI(api)
|
||||
}
|
||||
if d.voiceW != nil && d.voiceW.mcp != nil {
|
||||
api.getMCPServers = d.voiceW.mcp.status
|
||||
}
|
||||
return api
|
||||
}
|
||||
|
||||
// namedWorker is one long-running goroutine. The name exists so the set is
|
||||
// assertable from a test and readable in a log; nothing dispatches on it.
|
||||
type namedWorker struct {
|
||||
name string
|
||||
run func(ctx context.Context)
|
||||
}
|
||||
|
||||
// backgroundWorkers lists what this deployment runs. It is pure — it starts
|
||||
// nothing — so a test can compare the set the two paths would start without
|
||||
// standing a daemon up.
|
||||
func backgroundWorkers(d bootDeps) []namedWorker {
|
||||
var ws []namedWorker
|
||||
if d.voiceW != nil && d.voiceW.server != nil {
|
||||
ws = append(ws, namedWorker{"voice", func(context.Context) {
|
||||
if err := d.voiceW.server.Serve(); err != nil && !errors.Is(err, net.ErrClosed) {
|
||||
log.Printf("voice serve: %v", err)
|
||||
}
|
||||
}})
|
||||
}
|
||||
ws = append(ws,
|
||||
namedWorker{"tick", d.tl.run},
|
||||
namedWorker{"fact-enrichment", d.factWorker.run},
|
||||
)
|
||||
if d.evalWorker != nil {
|
||||
ws = append(ws, namedWorker{"memory-eval", d.evalWorker.run})
|
||||
}
|
||||
if d.feedWkr != nil {
|
||||
ws = append(ws, namedWorker{"feed", d.feedWkr.run})
|
||||
}
|
||||
if d.crawlWkr != nil {
|
||||
ws = append(ws, namedWorker{"crawl", d.crawlWkr.run})
|
||||
}
|
||||
if d.voiceW != nil && d.voiceW.mcp != nil {
|
||||
ws = append(ws, namedWorker{"mcp", d.voiceW.mcp.run})
|
||||
}
|
||||
if d.voiceW != nil && d.voiceW.home != nil {
|
||||
ws = append(ws, namedWorker{"home", d.voiceW.home.run})
|
||||
}
|
||||
return ws
|
||||
}
|
||||
|
||||
// startBackground starts every worker through goWorker, so waitWorkers can
|
||||
// wait for it at shutdown. A worker started as a bare `go func()` is the
|
||||
// shutdown bug documented at the end of run(): run() never returns, the
|
||||
// deferred Close never seals the database, and the ciphertext goes stale.
|
||||
func startBackground(ctx context.Context, wg *sync.WaitGroup, d bootDeps) {
|
||||
for _, w := range backgroundWorkers(d) {
|
||||
goWorker(wg, func() { w.run(ctx) })
|
||||
}
|
||||
if d.voiceW != nil && d.voiceW.server != nil {
|
||||
log.Printf("mavend: voice listening on %s", d.voiceW.server.Addr())
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,96 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"reflect"
|
||||
"testing"
|
||||
|
||||
"github.com/kami/maven/internal/decision"
|
||||
"github.com/kami/maven/internal/event"
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/store"
|
||||
"github.com/kami/maven/internal/voice"
|
||||
)
|
||||
|
||||
// fullDeps — a deployment with every optional piece present. Nothing here is
|
||||
// run: newDaemonAPI takes method values and backgroundWorkers is pure, so
|
||||
// zero-value wirings are enough to say what WOULD be started.
|
||||
func fullDeps() bootDeps {
|
||||
h := &reactiveHandler{
|
||||
ecosystem: &ecosystemWiring{nexus: &nexusClient{}},
|
||||
decisions: decision.NewRing(),
|
||||
}
|
||||
return bootDeps{
|
||||
coreFor: func() ipc.CoreAPI { return ipc.UnimplementedCoreAPI{} },
|
||||
tl: &tickLoop{},
|
||||
evBus: event.NewBus(4),
|
||||
st: &store.Store{},
|
||||
factWorker: &factEnrichmentWorker{},
|
||||
evalWorker: &memoryEvalWorker{},
|
||||
feedWkr: &feedWorker{},
|
||||
crawlWkr: &crawlWorker{},
|
||||
voiceW: &voiceWiring{
|
||||
server: &voice.Server{},
|
||||
handler: h,
|
||||
mcp: &mcpWiring{},
|
||||
home: &homeWiring{},
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// The unlock path used to build its own daemonAPI literal and leave nexus and
|
||||
// getMCPServers nil (V-639). Both paths call newDaemonAPI now, so the drift
|
||||
// that can still happen is a field added to the struct and not to the
|
||||
// constructor. This catches that one, by name.
|
||||
func TestNewDaemonAPISetsEveryField(t *testing.T) {
|
||||
prev := allowSeedOnStart
|
||||
allowSeedOnStart = true
|
||||
defer func() { allowSeedOnStart = prev }()
|
||||
|
||||
api := newDaemonAPI(fullDeps())
|
||||
v := reflect.ValueOf(*api)
|
||||
for i := range v.NumField() {
|
||||
if v.Field(i).IsZero() {
|
||||
t.Errorf("newDaemonAPI left %s unset — a fully wired deployment must fill every field", v.Type().Field(i).Name)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The handler is wired with the bare store adapter and cannot serve the day
|
||||
// plan until upgradeAPI hands it the real one. The unlocked path did that and
|
||||
// the unlock path did it too; keep it a property of the constructor.
|
||||
func TestNewDaemonAPIUpgradesTheHandler(t *testing.T) {
|
||||
d := fullDeps()
|
||||
api := newDaemonAPI(d)
|
||||
if d.voiceW.handler.api != ipc.CoreAPI(api) {
|
||||
t.Fatal("newDaemonAPI did not hand the handler the API it built")
|
||||
}
|
||||
}
|
||||
|
||||
// Every worker the daemon runs goes through startBackground, so shutdown can
|
||||
// wait for it. The unlock path used to start seven of these as bare
|
||||
// `go func()` under a shadowed WaitGroup.
|
||||
func TestBackgroundWorkersFullSet(t *testing.T) {
|
||||
want := []string{"voice", "tick", "fact-enrichment", "memory-eval", "feed", "crawl", "mcp", "home"}
|
||||
var got []string
|
||||
for _, w := range backgroundWorkers(fullDeps()) {
|
||||
got = append(got, w.name)
|
||||
}
|
||||
if !reflect.DeepEqual(got, want) {
|
||||
t.Errorf("workers = %v, want %v", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
// A default box configures none of the optional blocks. Two workers always run
|
||||
// and the rest stay dark, rather than a nil run being scheduled.
|
||||
func TestBackgroundWorkersFloor(t *testing.T) {
|
||||
d := fullDeps()
|
||||
d.evalWorker, d.feedWkr, d.crawlWkr, d.voiceW = nil, nil, nil, nil
|
||||
want := []string{"tick", "fact-enrichment"}
|
||||
var got []string
|
||||
for _, w := range backgroundWorkers(d) {
|
||||
got = append(got, w.name)
|
||||
}
|
||||
if !reflect.DeepEqual(got, want) {
|
||||
t.Errorf("workers = %v, want %v", got, want)
|
||||
}
|
||||
}
|
||||
@@ -571,7 +571,7 @@ func (h *reactiveHandler) finishClarified(ctx context.Context, dec router.Decisi
|
||||
}
|
||||
reply := h.applyAction(ctx, dec)
|
||||
if reply == "" {
|
||||
reply = h.replier.Reply(dec)
|
||||
reply = h.replier.Reply(ctx, dec)
|
||||
}
|
||||
if reply == "" {
|
||||
// Belt: an empty reply here would be a silent drop.
|
||||
|
||||
@@ -37,8 +37,8 @@ type factEnrichmentWorker struct {
|
||||
nextTry map[int64]time.Time // fact id → earliest retry
|
||||
}
|
||||
|
||||
// enrichmentScanLimit bounds how deep a single tick (or status report) walks
|
||||
// the pending queue looking for facts whose backoff has elapsed. The queue is
|
||||
// enrichmentScanLimit bounds how deep a single tick walks the pending queue
|
||||
// looking for facts whose backoff has elapsed. The queue is
|
||||
// ordered by id, so without a scan the oldest facts hold every batch slot
|
||||
// whether or not they are eligible, and one permanently failing fact stalls
|
||||
// every younger one behind it.
|
||||
@@ -75,8 +75,8 @@ func newFactEnrichmentWorker(st *store.Store, eco *ecosystemWiring, interval tim
|
||||
// has been down all day must be visible as a backlog, not as facts that
|
||||
// silently never got tagged.
|
||||
//
|
||||
// All three numbers describe the same set of rows, the first
|
||||
// enrichmentScanLimit pending facts. Counting Pending over a thousand rows
|
||||
// All three numbers describe the same set of rows, whatever is still pending
|
||||
// out of the first enrichmentScanLimit facts. Counting Pending over a thousand rows
|
||||
// while counting InBackoff over the twenty that reached the head of a batch
|
||||
// described two different populations under one struct.
|
||||
type enrichmentStatus struct {
|
||||
@@ -86,13 +86,22 @@ type enrichmentStatus struct {
|
||||
Scanned int // rows the other three counts were taken over
|
||||
}
|
||||
|
||||
// status reads the queue and counts over it. For a caller with no batch in
|
||||
// hand — anything asking the worker how it is doing from outside the tick.
|
||||
func (w *factEnrichmentWorker) status(ctx context.Context) enrichmentStatus {
|
||||
var st enrichmentStatus
|
||||
pending, err := w.store.PendingFactResolutions(ctx, enrichmentScanLimit)
|
||||
if err != nil {
|
||||
log.Printf("factenrichment: status: %v", err)
|
||||
return st
|
||||
return enrichmentStatus{}
|
||||
}
|
||||
return w.statusOf(pending)
|
||||
}
|
||||
|
||||
// statusOf counts over a batch the caller already has. The batch is the query
|
||||
// the tick already ran, so reporting the backlog costs no second read of the
|
||||
// scan limit — up to a thousand rows, on a database that serialises them.
|
||||
func (w *factEnrichmentWorker) statusOf(pending []store.Fact) enrichmentStatus {
|
||||
var st enrichmentStatus
|
||||
st.Pending = len(pending)
|
||||
st.Scanned = len(pending)
|
||||
w.mu.Lock()
|
||||
@@ -144,17 +153,24 @@ func (w *factEnrichmentWorker) tick(ctx context.Context) {
|
||||
}
|
||||
w.forgetDeparted(pending)
|
||||
skipped, failed, attempted := 0, 0, 0
|
||||
// A resolved fact leaves the pending queue, so the batch in hand overstates
|
||||
// the backlog by however many succeeded. Drop them here rather than
|
||||
// re-reading the queue to find out.
|
||||
remaining := make([]store.Fact, 0, len(pending))
|
||||
for _, f := range pending {
|
||||
if attempted >= w.batch {
|
||||
break
|
||||
remaining = append(remaining, f)
|
||||
continue
|
||||
}
|
||||
if !w.due(f.ID) {
|
||||
skipped++
|
||||
remaining = append(remaining, f)
|
||||
continue
|
||||
}
|
||||
attempted++
|
||||
if !w.resolveOne(ctx, f) {
|
||||
failed++
|
||||
remaining = append(remaining, f)
|
||||
}
|
||||
}
|
||||
if failed > 0 {
|
||||
@@ -164,7 +180,7 @@ func (w *factEnrichmentWorker) tick(ctx context.Context) {
|
||||
// Report the backlog every tick, not only when something failed: the
|
||||
// stalled state worth seeing is the one where nothing failed because
|
||||
// nothing was attempted.
|
||||
if st := w.status(ctx); st.Pending > 0 {
|
||||
if st := w.statusOf(remaining); st.Pending > 0 {
|
||||
log.Printf("factenrichment: %d facts pending entity resolution, %d in backoff, worst attempt %d (scanned %d)",
|
||||
st.Pending, st.InBackoff, st.MaxAttempts, st.Scanned)
|
||||
}
|
||||
|
||||
+30
-112
@@ -252,6 +252,23 @@ func run(args []string) error {
|
||||
// envelope per successful intake write.
|
||||
coreFor := func() ipc.CoreAPI { return newIntakeAPI(ipc.NewStoreAPI(st), evBus, time.Now) }
|
||||
|
||||
// depsNow reads whatever the current path has wired. Both boot paths build
|
||||
// the CoreAPI and start the workers from this one value, so neither can
|
||||
// hold a field the other misses. See cmd/mavend/boot.go.
|
||||
depsNow := func() bootDeps {
|
||||
return bootDeps{
|
||||
coreFor: coreFor,
|
||||
tl: tl,
|
||||
evBus: evBus,
|
||||
voiceW: voiceW,
|
||||
st: st,
|
||||
factWorker: factWorker,
|
||||
evalWorker: evalWorker,
|
||||
feedWkr: feedWkr,
|
||||
crawlWkr: crawlWkr,
|
||||
}
|
||||
}
|
||||
|
||||
if !locked {
|
||||
rules = wireRules(cfg)
|
||||
gatherer = wireGatherer(st, cfg, rules)
|
||||
@@ -284,26 +301,7 @@ func run(args []string) error {
|
||||
feedWkr = newFeedWorker(coreFor(), embedderOf(voiceW), cfg)
|
||||
crawlWkr = newCrawlWorker(newCrawler(cfg), coreFor(), embedderOf(voiceW), cfg)
|
||||
|
||||
coreAPI = &daemonAPI{
|
||||
CoreAPI: coreFor(),
|
||||
getTrace: tl.trace,
|
||||
getMorningStatus: func(ctx context.Context) []ipc.MorningRoutineStatus { return tl.morningStatus(ctx, time.Now()) },
|
||||
getDayPlan: func(ctx context.Context) ipc.DayPlan { return tl.dayPlan(ctx, time.Now()) },
|
||||
getEvents: intakeEventsFn(evBus),
|
||||
getDecisions: turnDecisionsFn(voiceW),
|
||||
seedStore: seedStoreIfAllowed(st),
|
||||
nexus: nexusOf(voiceW),
|
||||
}
|
||||
if voiceW != nil && voiceW.handler != nil {
|
||||
api := coreAPI.(*daemonAPI)
|
||||
api.chatFn = voiceW.handler.handleText
|
||||
// And the reverse: the handler was wired with the bare store
|
||||
// adapter, which cannot serve the day plan. See upgradeAPI.
|
||||
voiceW.handler.upgradeAPI(api)
|
||||
}
|
||||
if voiceW != nil && voiceW.mcp != nil {
|
||||
coreAPI.(*daemonAPI).getMCPServers = voiceW.mcp.status
|
||||
}
|
||||
coreAPI = newDaemonAPI(depsNow())
|
||||
} else {
|
||||
// locked mode: no real store yet, so there's no meaningful CoreAPI to
|
||||
// serve. srv.Check below is the actual guard — every CoreAPI call is
|
||||
@@ -364,6 +362,9 @@ func run(args []string) error {
|
||||
if !locked {
|
||||
wireMailIntake(srv, st, phr, cfg, evBus)
|
||||
wireModelSwap(srv, phr, cfg)
|
||||
// Inbound telegram (V-637). Dark unless the telegram block says intake,
|
||||
// and it reads one chat.
|
||||
wireTelegramIntake(ctx, &wg, coreAPI, cfg)
|
||||
// Vision + the media blob store (Vikunja #252). Both stay dark without a
|
||||
// media block; MethodDescribeImage answers ErrUnknownMethod then.
|
||||
keeper := wireVision(ctx, &wg, srv, st, embedderOf(voiceW), cfg)
|
||||
@@ -494,23 +495,14 @@ func run(args []string) error {
|
||||
crawlWkr = newCrawlWorker(newCrawler(cfg), coreFor(), embedderOf(voiceW), cfg)
|
||||
|
||||
// Swap the CoreAPI from the locked placeholder to the real store adapter.
|
||||
newAPI := &daemonAPI{
|
||||
CoreAPI: coreFor(),
|
||||
getTrace: tl.trace,
|
||||
getMorningStatus: func(ctx context.Context) []ipc.MorningRoutineStatus { return tl.morningStatus(ctx, time.Now()) },
|
||||
getDayPlan: func(ctx context.Context) ipc.DayPlan { return tl.dayPlan(ctx, time.Now()) },
|
||||
getEvents: intakeEventsFn(evBus),
|
||||
getDecisions: turnDecisionsFn(voiceW),
|
||||
seedStore: seedStoreIfAllowed(st),
|
||||
}
|
||||
if voiceW != nil && voiceW.handler != nil {
|
||||
newAPI.chatFn = voiceW.handler.handleText
|
||||
voiceW.handler.upgradeAPI(newAPI)
|
||||
}
|
||||
newAPI := newDaemonAPI(depsNow())
|
||||
srv.SetAPI(newAPI)
|
||||
srv.Check = (&auth.Gate{Enrollment: auth.NewFloorEnrollment(), Session: passkeySess}).Check
|
||||
wireMailIntake(srv, st, phr, cfg, evBus)
|
||||
wireModelSwap(srv, phr, cfg)
|
||||
// Same on the unlock path, with the API that has just replaced the
|
||||
// locked placeholder (V-637).
|
||||
wireTelegramIntake(ctx, &wg, newAPI, cfg)
|
||||
keeper := wireVision(ctx, &wg, srv, st, embedderOf(voiceW), cfg)
|
||||
wireCapture(ctx, &wg, srv, keeper, st, voiceW, phr, cfg)
|
||||
// Voice identification (Vikunja #255). Enrolment plumbing only until a
|
||||
@@ -518,59 +510,10 @@ func run(args []string) error {
|
||||
// block, so no wire path takes a voiceprint on a default box.
|
||||
wireSpeaker(srv, st, cfg)
|
||||
|
||||
// Start voice server.
|
||||
if voiceW != nil {
|
||||
var wg sync.WaitGroup
|
||||
wg.Add(1)
|
||||
go func() {
|
||||
defer wg.Done()
|
||||
if err := voiceW.server.Serve(); err != nil && !errors.Is(err, net.ErrClosed) {
|
||||
log.Printf("voice serve: %v", err)
|
||||
}
|
||||
}()
|
||||
log.Printf("mavend: voice listening on %s", voiceW.server.Addr())
|
||||
}
|
||||
|
||||
// Start tick loop.
|
||||
go func() {
|
||||
tl.run(ctx)
|
||||
}()
|
||||
|
||||
// Start fact-entity enrichment worker.
|
||||
go func() {
|
||||
factWorker.run(ctx)
|
||||
}()
|
||||
|
||||
// Start background memory evaluation (nil unless configured).
|
||||
if evalWorker != nil {
|
||||
go func() {
|
||||
evalWorker.run(ctx)
|
||||
}()
|
||||
}
|
||||
|
||||
// Start feed reading (nil unless configured).
|
||||
if feedWkr != nil {
|
||||
go func() {
|
||||
feedWkr.run(ctx)
|
||||
}()
|
||||
}
|
||||
|
||||
// Start the watched-page crawls (nil unless configured).
|
||||
if crawlWkr != nil {
|
||||
go func() {
|
||||
crawlWkr.run(ctx)
|
||||
}()
|
||||
}
|
||||
|
||||
// Keep MCP connections alive (nil unless configured).
|
||||
if voiceW != nil && voiceW.mcp != nil {
|
||||
go voiceW.mcp.run(ctx)
|
||||
}
|
||||
|
||||
// Re-enumerate the house for new devices (nil unless configured).
|
||||
if voiceW != nil && voiceW.home != nil {
|
||||
go voiceW.home.run(ctx)
|
||||
}
|
||||
// The voice server and every background worker, on the outer wg
|
||||
// so shutdown waits for them. This used to be nine bare
|
||||
// `go func()` calls and a shadowed WaitGroup (V-639).
|
||||
startBackground(ctx, &wg, depsNow())
|
||||
|
||||
dl.unlock(st)
|
||||
log.Printf("mavend: unlocked via passkey assertion")
|
||||
@@ -585,33 +528,8 @@ func run(args []string) error {
|
||||
})
|
||||
log.Printf("mavend: ipc listening on %s", srv.Path())
|
||||
|
||||
if !locked && voiceW != nil {
|
||||
goWorker(&wg, func() {
|
||||
if err := voiceW.server.Serve(); err != nil && !errors.Is(err, net.ErrClosed) {
|
||||
log.Printf("voice serve: %v", err)
|
||||
}
|
||||
})
|
||||
log.Printf("mavend: voice listening on %s", voiceW.server.Addr())
|
||||
}
|
||||
|
||||
if !locked {
|
||||
goWorker(&wg, func() { tl.run(ctx) })
|
||||
goWorker(&wg, func() { factWorker.run(ctx) })
|
||||
if evalWorker != nil {
|
||||
goWorker(&wg, func() { evalWorker.run(ctx) })
|
||||
}
|
||||
if feedWkr != nil {
|
||||
goWorker(&wg, func() { feedWkr.run(ctx) })
|
||||
}
|
||||
if crawlWkr != nil {
|
||||
goWorker(&wg, func() { crawlWkr.run(ctx) })
|
||||
}
|
||||
if voiceW != nil && voiceW.mcp != nil {
|
||||
goWorker(&wg, func() { voiceW.mcp.run(ctx) })
|
||||
}
|
||||
if voiceW != nil && voiceW.home != nil {
|
||||
goWorker(&wg, func() { voiceW.home.run(ctx) })
|
||||
}
|
||||
startBackground(ctx, &wg, depsNow())
|
||||
}
|
||||
|
||||
<-ctx.Done()
|
||||
|
||||
@@ -22,7 +22,7 @@ func newLLMReplier(c phraser.Completer, block func() string) *llmReplier {
|
||||
|
||||
// Reply never fails: a clarify, a model error and an unusable generation all
|
||||
// answer from the stub, which is what keeps a turn from breaking on the model.
|
||||
func (r *llmReplier) Reply(d router.Decision) string {
|
||||
func (r *llmReplier) Reply(ctx context.Context, d router.Decision) string {
|
||||
if d.Clarify {
|
||||
// The deck, not the stub's single sentence: a clarify she cannot turn
|
||||
// into a question is the line he hears most often when she misses him,
|
||||
@@ -39,14 +39,14 @@ func (r *llmReplier) Reply(d router.Decision) string {
|
||||
// что ты выпел стакан воды" for "я выпил воды".
|
||||
return phraser.FactAck(d.Utterance)
|
||||
}
|
||||
out, err := r.p.PhraseReply(context.Background(), d)
|
||||
out, err := r.p.PhraseReply(ctx, d)
|
||||
if err != nil || out == "" {
|
||||
return r.stub.Reply(d)
|
||||
return r.stub.Reply(ctx, d)
|
||||
}
|
||||
// The persona checks, on the live path (personaguard.go). A reply that
|
||||
// leaks reasoning or calls him "вы" is worse than a flat one.
|
||||
if _, ok := guardSpoken("reply", out); !ok {
|
||||
return r.stub.Reply(d)
|
||||
return r.stub.Reply(ctx, d)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
@@ -22,7 +22,7 @@ func (s stubCompleter) Complete(_ context.Context, _ llm.Req) (string, error) {
|
||||
|
||||
func TestLLMReplierPassesTheModelReplyThrough(t *testing.T) {
|
||||
r := newLLMReplier(stubCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`}, nil)
|
||||
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
|
||||
got := r.Reply(context.Background(), router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
|
||||
if got != "записала, кофе закончился" {
|
||||
t.Errorf("got %q, want %q", got, "записала, кофе закончился")
|
||||
}
|
||||
@@ -42,7 +42,7 @@ func TestLLMReplierFallsBackToStubOnEmpty(t *testing.T) {
|
||||
// the clarify deck rather than the stub's single sentence.
|
||||
func TestLLMReplierClarifyReadsTheDeck(t *testing.T) {
|
||||
r := newLLMReplier(stubCompleter{out: "я всё поняла"}, nil)
|
||||
got := r.Reply(router.Decision{Clarify: true, Utterance: "мгм"})
|
||||
got := r.Reply(context.Background(), router.Decision{Clarify: true, Utterance: "мгм"})
|
||||
if got == "я всё поняла" {
|
||||
t.Fatal("a clarify must not be phrased by the model")
|
||||
}
|
||||
@@ -50,7 +50,7 @@ func TestLLMReplierClarifyReadsTheDeck(t *testing.T) {
|
||||
t.Errorf("on clarify: got %q, want %q", got, want)
|
||||
}
|
||||
// Two different misses do not sound identical.
|
||||
if same := r.Reply(router.Decision{Clarify: true, Utterance: "а"}); same == got {
|
||||
if same := r.Reply(context.Background(), router.Decision{Clarify: true, Utterance: "а"}); same == got {
|
||||
t.Log("two utterances hashed to the same line, which is allowed but should be rare")
|
||||
}
|
||||
}
|
||||
@@ -60,14 +60,14 @@ func TestLLMReplierClarifyReadsTheDeck(t *testing.T) {
|
||||
// produce, which is the same claim without pinning one wording.
|
||||
func assertAck(t *testing.T, r *llmReplier, d router.Decision, key, what string) {
|
||||
t.Helper()
|
||||
if got := r.Reply(d); !phraser.IsAck(key, nil, got) {
|
||||
if got := r.Reply(context.Background(), d); !phraser.IsAck(key, nil, got) {
|
||||
t.Errorf("on %s: got %q, want a %q line", what, got, key)
|
||||
}
|
||||
}
|
||||
|
||||
func assertStub(t *testing.T, r *llmReplier, d router.Decision, what string) {
|
||||
t.Helper()
|
||||
got, want := r.Reply(d), voice.NewStubReplier().Reply(d)
|
||||
got, want := r.Reply(context.Background(), d), voice.NewStubReplier().Reply(context.Background(), d)
|
||||
if got != want {
|
||||
t.Errorf("on %s: got %q, want stub %q", what, got, want)
|
||||
}
|
||||
|
||||
@@ -0,0 +1,59 @@
|
||||
// mavend/telegramintake.go — wiring the inbound telegram poller (V-637).
|
||||
//
|
||||
// The poller reaches the daemon through ipc.CoreAPI and nothing else, so a
|
||||
// telegram turn takes exactly the path the web's POST /api/chat takes: Chat
|
||||
// returns the reply and the persisted trace id, and CorrectTurn writes the
|
||||
// label. Nothing in internal/delivery knows what a handler is.
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"log"
|
||||
"sync"
|
||||
|
||||
"github.com/kami/maven/internal/config"
|
||||
"github.com/kami/maven/internal/delivery/telegramsink"
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
)
|
||||
|
||||
// wireTelegramIntake starts the poller, or returns having done nothing. It is
|
||||
// nil-safe in every argument, because it is called from both boot paths — the
|
||||
// unlocked start and the passkey unlock — and telegram must behave the same on
|
||||
// either.
|
||||
//
|
||||
// A sink that will not build is logged rather than fatal here. The push half
|
||||
// already failed the boot in wireDispatcher for the same config, so a second
|
||||
// hard failure would only lose that message.
|
||||
func wireTelegramIntake(ctx context.Context, wg *sync.WaitGroup, api ipc.CoreAPI, cfg *config.Config) {
|
||||
if cfg == nil || cfg.Telegram == nil || !cfg.Telegram.Intake || api == nil {
|
||||
return
|
||||
}
|
||||
sink, err := telegramsink.New(*cfg.Telegram)
|
||||
if err != nil {
|
||||
log.Printf("telegram intake: %v", err)
|
||||
return
|
||||
}
|
||||
poller, err := telegramsink.NewPoller(sink, chatTurnFn(api), api.CorrectTurn)
|
||||
if err != nil {
|
||||
log.Printf("telegram intake: %v", err)
|
||||
return
|
||||
}
|
||||
wg.Add(1)
|
||||
go func() {
|
||||
defer wg.Done()
|
||||
poller.Run(ctx)
|
||||
}()
|
||||
}
|
||||
|
||||
// chatTurnFn adapts ipc.Chat to the poller's Turn. The trace id comes back on
|
||||
// the reply because the daemon's Chat collects it off the context (V-630), so
|
||||
// the chat can offer the same correction the web does without a second op.
|
||||
func chatTurnFn(api ipc.CoreAPI) telegramsink.Turn {
|
||||
return func(ctx context.Context, conversation, text string) (string, int64, error) {
|
||||
reply, err := api.Chat(ctx, conversation, text)
|
||||
if err != nil {
|
||||
return "", 0, err
|
||||
}
|
||||
return reply.Reply, reply.TraceID, nil
|
||||
}
|
||||
}
|
||||
+1
-1
@@ -458,7 +458,7 @@ func (h *reactiveHandler) runTurn(ctx context.Context, text string, src turnSour
|
||||
|
||||
// 9. replier — phrase the reply across the router decision.
|
||||
if replyText == "" {
|
||||
replyText = h.replier.Reply(dec)
|
||||
replyText = h.replier.Reply(ctx, dec)
|
||||
}
|
||||
return withNotice(expiredNotice, replyText)
|
||||
}
|
||||
|
||||
+19
-1
@@ -58,6 +58,12 @@ func main() {
|
||||
// mutex, so sharing the connection would freeze every other page for the
|
||||
// length of the load. See handleModels.
|
||||
var swapConn modelController
|
||||
// turnConn — a third connection, for POST /api/chat and nothing else, for
|
||||
// the same reason /models has one (V-638). A chat turn routes, phrases and
|
||||
// may act, bounded only by phraser.timeout at 60s, and every other handler
|
||||
// on this server queues behind it on the shared client's one mutex. Nil ⇒
|
||||
// chat shares the main connection, which is how it behaved before.
|
||||
var turnConn ipc.CoreAPI
|
||||
if *coreSock != "" {
|
||||
c, err := ipc.DialWait(*coreSock, 60*time.Second)
|
||||
if err != nil {
|
||||
@@ -71,6 +77,12 @@ func main() {
|
||||
defer sc.Close()
|
||||
swapConn = sc
|
||||
}
|
||||
if tc, err := ipc.Dial(*coreSock); err != nil {
|
||||
log.Printf("chat: third core connection failed (%v) — /api/chat will share the main one and a turn will block the other pages", err)
|
||||
} else {
|
||||
defer tc.Close()
|
||||
turnConn = tc
|
||||
}
|
||||
}
|
||||
|
||||
// stepUpSession stays nil unless the passkey endpoints are wired below — it
|
||||
@@ -208,7 +220,13 @@ func main() {
|
||||
// decides how every utterance is routed and how every reply is worded.
|
||||
mux.HandleFunc("/tools", gatedPage(handleTools))
|
||||
mux.HandleFunc("/routines", gatedPage(handleRoutines))
|
||||
mux.HandleFunc("/api/chat", gatedPage(handleChatAPI))
|
||||
mux.HandleFunc("/api/chat", func(w http.ResponseWriter, r *http.Request) {
|
||||
c := turnConn
|
||||
if c == nil {
|
||||
c = core
|
||||
}
|
||||
handleChatAPI(w, r, c, stepUpSession, *requireStepUp)
|
||||
})
|
||||
mux.HandleFunc("/api/revert", gatedPage(handleRevert))
|
||||
mux.HandleFunc("/api/correct", gatedPage(handleCorrectAPI))
|
||||
mux.HandleFunc("/models", func(w http.ResponseWriter, r *http.Request) {
|
||||
|
||||
+9
-1
@@ -37,7 +37,15 @@
|
||||
"This needs a matching ufw rule or the container's SYN is dropped:",
|
||||
" ufw allow from 192.168.240.0/20 to any port 10808 proto tcp"
|
||||
],
|
||||
"proxy": "socks5://192.168.240.1:10808"
|
||||
"proxy": "socks5://192.168.240.1:10808",
|
||||
"//intake": [
|
||||
"Read the chat as well as write to it (V-637). The poller long-polls",
|
||||
"getUpdates through the same relay and accepts chat_id as the only",
|
||||
"sender. Deleting this key turns inbound off again.",
|
||||
"chat_id must be numeric here or the daemon refuses to start: an inbound",
|
||||
"update names its chat by number, so an @-name would match nothing."
|
||||
],
|
||||
"intake": true
|
||||
},
|
||||
|
||||
"//workstation": [
|
||||
|
||||
@@ -0,0 +1,71 @@
|
||||
# The routing trajectory, and the number that is missing
|
||||
|
||||
**06-08-2026. V-464.** Not a new measurement. This collates the figures already recorded
|
||||
in `docs/evals/` and CLAUDE.md, and names one measurement that has not been taken. Dated
|
||||
because the conclusion expires the moment the missing number is measured.
|
||||
|
||||
## The question
|
||||
|
||||
126 of the 1023 commits between 03-07-2026 and 06-08-2026 touch `internal/router`. Is the
|
||||
routing between the core functions and his speech getting better?
|
||||
|
||||
## The trajectory
|
||||
|
||||
RU routing fixture, classifier plus the ONNX embedder, no LLM arm in any of these runs.
|
||||
|
||||
| date | change | fixture | source |
|
||||
|---|---|---|---|
|
||||
| 02-08-2026 | classifier re-measured | 68.8% of 77 | CLAUDE.md |
|
||||
| 04-08-2026 | V-498, rest-of-day and narrative rules | 58/82, 70.7% | CLAUDE.md |
|
||||
| 06-08-2026 | V-626 baseline | 64/91, 70.3% | `2026-08-06-seeds-to-prompt-boundary.md` |
|
||||
| 06-08-2026 | V-626, seeds onto the prompt boundary | 66/91, 72.5% | same |
|
||||
| 06-08-2026 | V-627, alarm verbs reach stage 0 | 69/91, 75.8% | `2026-08-06-alarm-verbs-reach-stage-0.md` |
|
||||
| 06-08-2026 | V-633, Russian acts reach tools | 69/91, unchanged | `2026-08-06-russian-acts-reach-tools.md` |
|
||||
|
||||
The fixture grew from 77 to 82 to 91 cases across this window. So the percentages are
|
||||
comparable and the counts are not.
|
||||
|
||||
## Accuracy moved late
|
||||
|
||||
It sat near 70% for a month. V-626 and V-627 landed the same day and took the
|
||||
deterministic path from 64/91 to 69/91. That is the first real accuracy movement since the
|
||||
stage-0 rules went in.
|
||||
|
||||
## Most of the work was reach, not accuracy
|
||||
|
||||
Praxis went 0/12 to 11/12 and lifecycle 0/5 to 5/5 (V-516,
|
||||
`2026-08-05-praxis-reach.md`). No Russian utterance could reach a tool before V-633. That
|
||||
one landed at 69/91 unchanged, because the fixture holds no case for it. Alarm verbs,
|
||||
ordinal selection, spoken corrections and the claimant ladder share the shape.
|
||||
|
||||
So the fixture undercounts the month. Things that were structurally unreachable now reach,
|
||||
and a fixture that never asked about them cannot show it. Judge reach against
|
||||
`make eval-reach` and the ecosystem fixture, not against the routing one.
|
||||
|
||||
## The missing number
|
||||
|
||||
On 05-08-2026 the cascade with the resident model scored 69/91, 75.8% full, 80.2%
|
||||
intent-only, at p50 1.19s (`2026-08-05-routing-resident-model.md`).
|
||||
|
||||
On 06-08-2026 the classifier and stage 0 alone reached 69/91, 75.8% full, at p50 22.9ms.
|
||||
|
||||
Those are the same full-accuracy score. The cascade has not been re-measured since V-626
|
||||
and V-627 landed. Both are stage-0 changes, and stage 0 runs inside the cascade, so the
|
||||
cascade should have gained from them too.
|
||||
|
||||
One of two things is true, and nothing on the box says which:
|
||||
|
||||
- The cascade gained as well, the model still separates from the floor on intent-only, and
|
||||
it earns its place.
|
||||
- The deterministic floor has caught up on this fixture, and the resident model is costing
|
||||
1.17 seconds a turn for nothing measurable.
|
||||
|
||||
Take that measurement before planning more routing work. It needs a second llama-server on
|
||||
a fixed host port, because the resident one binds `--port 0` inside the container.
|
||||
|
||||
## What this does not settle
|
||||
|
||||
Intent-only is the more honest comparison for the model arm. The model routes `reminder`
|
||||
and leaves the time to the daemon, which is what the contract asks. The 05-08 run puts it
|
||||
at 80.2% through the cascade and 61.5% for the model alone. There is no 06-08 intent-only
|
||||
figure for the deterministic path to set beside those.
|
||||
@@ -0,0 +1,69 @@
|
||||
# Does one sqlite connection make reads queue? No (V-642)
|
||||
|
||||
Measured 07-08-2026 at `7b507de`, on homesrv. The harness is
|
||||
`internal/store/conncap_test.go`. It stays in the repo, because this claim gets
|
||||
re-argued and the numbers should be re-runnable rather than quoted.
|
||||
|
||||
`internal/store/store.go` opens the database with `SetMaxOpenConns(1)`, while
|
||||
`schema.sql` sets `journal_mode=WAL`. WAL exists to let readers run beside one
|
||||
writer, so the cap gives up the thing the journal mode was chosen for. The
|
||||
question was whether that costs anything.
|
||||
|
||||
## What was measured
|
||||
|
||||
A fixed two-second window. One writer calling `SetValue` paced at 2ms, and a
|
||||
reader loop calling `RecentFacts(50)` over 500 seeded rows as fast as it can.
|
||||
Same schema, same modernc driver, same machine, three runs per cap.
|
||||
|
||||
The window is wall-clock rather than a read count on purpose. A first version ran
|
||||
a fixed 300 reads. That finished sooner at the higher cap, so it received fewer
|
||||
writes, and two runs that did different work cannot be compared.
|
||||
|
||||
| cap | reads | writes | p50 | p95 | max |
|
||||
|---|---|---|---|---|---|
|
||||
| 1 | ~3050 | ~760 | 594µs | 900µs | 16-19ms |
|
||||
| 4 | ~3600 | ~340 | 525µs | 710µs | 1-2ms |
|
||||
|
||||
## What it says
|
||||
|
||||
**Reads do not queue behind writes.** Four connections buy about 70µs at p50. A
|
||||
turn spends 1.19s in the resident model. The tail does improve, from 19ms to 2ms,
|
||||
and 19ms is still not a figure anyone notices in a spoken reply.
|
||||
|
||||
**Write throughput more than halves at the higher cap**, 760 writes against 340.
|
||||
inference, not measured directly: at one connection the reader and the writer take
|
||||
turns with no lock contention. At four the writer contends for the WAL write lock
|
||||
with a live reader. Whatever the mechanism, the trade runs the opposite way from
|
||||
the one the task expected.
|
||||
|
||||
**The cap was not the source of the 2.7s router figure.** CLAUDE.md records that
|
||||
figure as contention rather than the model. This task was a candidate for where
|
||||
that contention came from. A 19ms worst case cannot produce it. That line of
|
||||
enquiry is closed.
|
||||
|
||||
**One transaction is what the cap cannot survive.** With a read-only transaction
|
||||
open, a second read at cap 1 never completes. The harness gave it two seconds and
|
||||
got `context deadline exceeded`. The same read at cap 4 took 1ms. The transaction
|
||||
holds the only connection, so this is not a slow read, it is a stalled database.
|
||||
|
||||
## What was done
|
||||
|
||||
The cap stays at 1. The reason is now written where the cap is set, rather than
|
||||
inferred from a four-word comment.
|
||||
|
||||
`Store.DB` was deleted. It handed out exactly the read-only transaction measured
|
||||
above. It had been there since the initial commit with no production caller, and
|
||||
its doc comment described a loop that never materialised. Its one user was a test
|
||||
helper reading `delivery_attempts` by raw SQL. `ListDeliveryAttempts` has covered
|
||||
that since V-390, and the helper now goes through the reader.
|
||||
|
||||
So the hazard is gone by construction, not by documentation.
|
||||
`TestConnCap_ReadBlocksBehindOpenSnapshot` is the standing measurement of what
|
||||
re-adding the seam would cost.
|
||||
|
||||
## Not answered
|
||||
|
||||
Whether reads queue on the deployed box under real load, as opposed to a
|
||||
synthetic loop. The harness writes and reads one table. Digestion reads four and
|
||||
embeds while it does. The finding that closes this task is the transaction stall,
|
||||
which is structural and does not depend on load.
|
||||
@@ -0,0 +1,70 @@
|
||||
# Inbound telegram
|
||||
|
||||
Last verified: 06-08-2026 @ c61b0b3
|
||||
|
||||
V-637, under V-628. Reads with `22-correcting-a-turn.md`.
|
||||
|
||||
## What was missing
|
||||
|
||||
Telegram was a reach and nothing else. `telegramsink` pushed an away message and the chat
|
||||
had no way to answer, so the correction gesture reached the web and voice only.
|
||||
|
||||
That skews the labels. V-546 fits routing heads on them, and a sample drawn from wherever
|
||||
the owner happens to be sitting is the wrong sample.
|
||||
|
||||
## Long-poll, not a webhook
|
||||
|
||||
The box takes no inbound connections and reaches api.telegram.org through a relay, so the
|
||||
connection has to open outward. `getUpdates` with a 25 second hold, one goroutine in the
|
||||
daemon's WaitGroup.
|
||||
|
||||
A failed poll waits 15 seconds and retries without escalating. The relay going down is the
|
||||
normal cause and it comes back on its own.
|
||||
|
||||
## The backlog is dropped on start
|
||||
|
||||
Telegram keeps undelivered updates for 24 hours. A daemon that was down overnight would
|
||||
otherwise wake and answer every queued message in order.
|
||||
|
||||
That is worse than missing them. A question asked eight hours ago has been answered
|
||||
already. A reminder set from it lands at the wrong time. So the first call moves the offset
|
||||
past whatever is queued and acts on none of it.
|
||||
|
||||
## One chat
|
||||
|
||||
`ChatID` is the only accepted sender, and it is the same chat the push half already sends
|
||||
to. A message from anywhere else is dropped with no reply, because a reply confirms the bot
|
||||
exists and whose it is.
|
||||
|
||||
Chat ids are not guessable. They are also not secret, since they travel in every forwarded
|
||||
message. So this is the whole authorisation and it is an allowlist of one.
|
||||
|
||||
## The gesture
|
||||
|
||||
Two taps at most. The reply carries one button, `не то`. Tapping it writes nothing and opens
|
||||
the seven intents plus `просто неверно`. The untargeted negative stays reachable, because he
|
||||
may have opened the row without meaning to name anything.
|
||||
|
||||
Callback data carries the trace id and the target, under telegram's 64 byte cap. It comes
|
||||
off the wire. So an id that will not parse is dropped, and so is a target that is not one of
|
||||
the seven. A label nothing can score is worse than no label.
|
||||
|
||||
A failed write says so on the button and leaves the keyboard up. A successful one takes the
|
||||
keyboard off, because a live keyboard on an answered turn invites correcting it twice.
|
||||
|
||||
## The seam
|
||||
|
||||
`NewPoller` takes two functions and no daemon type. `cmd/mavend/telegramintake.go` fills
|
||||
them from `ipc.CoreAPI`: `Chat` returns the reply and the trace id it collected off the
|
||||
context, and `CorrectTurn` writes the label. So a chat turn takes the path
|
||||
`POST /api/chat` already takes, and nothing in `internal/delivery` knows what a handler is.
|
||||
|
||||
## What is not done
|
||||
|
||||
The turn source is still `tap:text`, which telegram shares with the web. Provenance cannot
|
||||
tell a chat turn from a typed one, so a label's `source` column cannot either.
|
||||
That matters the first time someone asks whether corrections given in the chat differ from
|
||||
corrections given at the desk.
|
||||
|
||||
Voice messages are ignored. The poller reads `message.text` and nothing else, so a voice
|
||||
note in the chat does not reach `mavsttd`.
|
||||
@@ -0,0 +1,99 @@
|
||||
# No deadline on the turn path
|
||||
|
||||
Last verified: 06-08-2026 @ 60e64dd
|
||||
|
||||
**All four steps landed on 06-08-2026.** What follows describes the defect as it was and
|
||||
the work as it was planned. Two things came out differently. `Client.Close` read the conn
|
||||
field with no lock while `roundtrip` re-dialed and dropped it. `-race` caught that on the
|
||||
new cancellation test. So the conn field now has a mutex of its own, held only across a
|
||||
read or an assignment. And `/api/ptt` needed nothing: it proxies to the voice port and never
|
||||
touches the shared client, so only `/api/chat` got the extra connection. The pool inside
|
||||
`ipc.Client` is still unbuilt and still waiting on a second module measured queueing.
|
||||
|
||||
V-638. Sibling of V-607, which is the same class of bug in `internal/worker`.
|
||||
Reads with `docs/offload.md` and `docs/protocol.md`.
|
||||
|
||||
## What is missing
|
||||
|
||||
A chat turn starts in a mavweb HTTP handler and ends at llama-server. Nothing between those
|
||||
two points can be cancelled, and one hop has a timeout.
|
||||
|
||||
Four places, all on the same path.
|
||||
|
||||
`voice.Replier.Reply` takes no context (`internal/voice/replier.go:41`). So `llmReplier`
|
||||
calls `PhraseReply(context.Background(), d)` at `cmd/mavend/replier_llm.go:42`. The turn
|
||||
cannot deadline its own reply. The only bound is `phraser.timeout`, 60s in deploy.
|
||||
|
||||
`ipc.Client.roundtrip` sets no connection deadline (`internal/ipc/client.go:202`). A daemon
|
||||
that stops answering parks the caller for as long as the socket stays open.
|
||||
|
||||
`ipc.Client.call` checks the context once, before sending (`client.go:149`), then blocks in
|
||||
`roundtrip`. Cancelling mid-call does nothing.
|
||||
|
||||
`ipc.Server.serveConn` dispatches under `context.Background()` (`internal/ipc/server.go:253`).
|
||||
A client that hangs up does not cancel the turn, and neither does `Server.Close`.
|
||||
|
||||
## And every call queues behind the slowest one
|
||||
|
||||
`ipc.Client` serialises on one connection and one mutex. mavweb routes `/api/chat` and
|
||||
`/api/ptt` through the shared client, so one turn blocks all 28 handlers while it runs.
|
||||
Worst case is a 60s page load.
|
||||
|
||||
This is understood for exactly one route already. `cmd/mavweb/main.go:57` opens a second
|
||||
connection for `/models`, and the comment there says why. A model swap is a multi-minute
|
||||
call, and sharing the connection would freeze every other page.
|
||||
|
||||
## The pattern is already in the repo
|
||||
|
||||
`internal/voice/client.go:101` derives a connection deadline from the caller's context,
|
||||
falls back to 120s, and clears it with a defer. `internal/ipc/client.go` never learned it.
|
||||
Copy that rather than inventing a second convention.
|
||||
|
||||
## The work
|
||||
|
||||
One commit each.
|
||||
|
||||
**Context on the reply seam.** `phraser.Replier.PhraseReply` already takes a context and the
|
||||
interface has two implementations, so this is small. Change `Reply` to take a context, have
|
||||
`StubReplier` ignore it, and pass it through `llmReplier` to `PhraseReply`. Both call sites
|
||||
already hold one: `cmd/mavend/voice.go:461` and `cmd/mavend/clarify.go:574`.
|
||||
|
||||
**Deadlines and cancellation on the client.** Pass the context into `roundtrip` and set
|
||||
`SetDeadline` from it. For cancellation mid-call, a watchdog goroutine that calls `c.drop()`
|
||||
on `ctx.Done()` is enough. `drop` exists, and the retry split already separates a lost write
|
||||
from a lost read. So a cancelled call lands in `errReadLost` and is never retried for a
|
||||
mutation. Check that against `internal/ipc/maperr_test.go`.
|
||||
|
||||
**A request context on the server.** `serveConn` should derive from a server-scoped context
|
||||
so `Close` cancels a dispatch in flight. `Server` already carries `done` and a conn registry
|
||||
for this class of problem. The registry comment records what the last version of it cost:
|
||||
eleven days of stale ciphertext.
|
||||
|
||||
**Stop serialising mavweb.** Give `/api/chat` and `/api/ptt` their own connection, the way
|
||||
`/models` has one. Roughly ten lines, and it changes no shared code.
|
||||
|
||||
A connection pool inside `ipc.Client` is the general form and is deliberately not the first
|
||||
step. Each connection is already its own request and response stream. So a pool preserves
|
||||
frame pairing by construction. It still has to keep re-dial on drop, the
|
||||
`errWriteLost` and `errReadLost` split, and `Close`. Do the narrow fix, measure, and reach
|
||||
for the pool only if a second module turns out to queue.
|
||||
|
||||
## How it is judged
|
||||
|
||||
`make test` stays green. It is green at `06c1cf2`.
|
||||
|
||||
Nothing here changes routing or recall, so `make eval-router` and `make eval-recall` are
|
||||
unchanged rather than re-measured.
|
||||
|
||||
By hand: load `/dash` while a chat turn is in flight. Before the change it waits for the
|
||||
length of the turn.
|
||||
|
||||
There is no test today that a cancelled context aborts an in-flight `ipc.Client` call. That
|
||||
absence is why two of these four went unnoticed, so the test is part of the work.
|
||||
|
||||
## What is not done here
|
||||
|
||||
The store is still `SetMaxOpenConns(1)` (`internal/store/store.go:99`) under WAL. WAL is
|
||||
built for concurrent readers against one writer, and the cap makes every read queue.
|
||||
`Store.DB(ctx)` hands the digestion worker a read transaction on that same connection. This
|
||||
plan does not touch it. It is measurable first and should be measured before it is changed.
|
||||
@@ -0,0 +1,99 @@
|
||||
# The two boot paths have drifted
|
||||
|
||||
Last verified: 06-08-2026 @ 69d0f5e
|
||||
|
||||
V-639. Reads with `docs/operations.md`.
|
||||
|
||||
## What landed
|
||||
|
||||
`cmd/mavend/boot.go`. `newDaemonAPI(deps)` builds the CoreAPI with every field
|
||||
set, and `startBackground(ctx, &wg, deps)` starts the voice server and every
|
||||
worker through `goWorker`. `backgroundWorkers(deps)` is the pure list behind it,
|
||||
so a test can compare the set without standing a daemon up. Both paths in
|
||||
`run()` now read `coreAPI = newDaemonAPI(depsNow())` and one
|
||||
`startBackground(...)`, where `depsNow` reads whatever the current path wired.
|
||||
|
||||
The shadowed `wg` is gone. Four tests in `cmd/mavend/boot_test.go`. Every
|
||||
`daemonAPI` field is set on a fully wired deployment. The handler gets the API
|
||||
it was built with. The worker set is asserted by name, at the full set and at
|
||||
the floor.
|
||||
|
||||
Still by hand: unlock a locked box by passkey, ask something that needs Nexus,
|
||||
and check `/tools` lists the MCP servers.
|
||||
|
||||
## What is wrong
|
||||
|
||||
`run()` in `cmd/mavend/main.go` brings the daemon up two ways. A box with a key in the
|
||||
environment starts unlocked and wires everything at lines 280 to 621. A box without one
|
||||
starts locked. It wires the same things again inside the unlock closure, at lines 500 to
|
||||
579, after a passkey assertion.
|
||||
|
||||
The two lists have drifted apart. Three ways.
|
||||
|
||||
**Seven workers start untracked.** The unlocked path puts every one through
|
||||
`goWorker(&wg, ...)`, so `waitWorkers` at line 637 can wait for them. The unlock path
|
||||
starts `tl.run`, `factWorker`, `evalWorker`, `feedWkr`, `crawlWkr`, `mcp.run` and
|
||||
`home.run` as bare `go func()`. Nothing waits for any of them.
|
||||
|
||||
That is the shutdown bug the code already documents at lines 631 to 636, reintroduced on
|
||||
the other path. The comment there records what it cost the first time. `run()` never
|
||||
returned, so `defer st.Close()` never sealed the database. The deployed ciphertext was
|
||||
eleven days stale before anyone noticed.
|
||||
|
||||
**A shadowed WaitGroup hides it.** Line 529 declares `var wg sync.WaitGroup` inside the
|
||||
`if voiceW != nil` block, shadowing the one from line 359. It is `Add`ed and `Done`d and
|
||||
never waited. Reading the block, the voice server looks tracked. It is not.
|
||||
|
||||
**Two `daemonAPI` fields are never set.** The unlocked path fills `nexus` at line 295 and
|
||||
`getMCPServers` at line 305. The unlock path fills neither. So after a passkey unlock,
|
||||
`ResolveEntity` answers `ErrNotImplemented` with a `nexus` block configured, and
|
||||
`MCPServers` answers empty with an `mcp` block configured.
|
||||
|
||||
The second is the worse one. Empty is not a degraded answer, it is a wrong answer, and
|
||||
`/tools` renders it as "not configured".
|
||||
|
||||
## Why it drifted
|
||||
|
||||
`wireTelegramIntake` was added to both paths on 06-08-2026 (V-637) and it does use the
|
||||
outer `wg`, at line 519. So the newest line on that path is correct and the older ones
|
||||
around it are not. The path gets touched one line at a time and is never read whole.
|
||||
|
||||
The shape of `cmd/mavend` is what allows that. It is 155 files and 9,551 lines of code.
|
||||
Six things live in it with no seam between them:
|
||||
|
||||
- the handler
|
||||
- the action dispatch
|
||||
- the 19 query sources
|
||||
- the wiring functions
|
||||
- the six background workers
|
||||
- these two boot paths
|
||||
|
||||
Nothing in the package makes the divergence visible.
|
||||
|
||||
## The fix
|
||||
|
||||
Make the two paths call one function instead of listing the same wiring twice.
|
||||
|
||||
One `startBackground(ctx, &wg, deps)` that takes what it needs and starts every worker
|
||||
through `goWorker`. One `newDaemonAPI(deps)` that fills every field, including `nexus` and
|
||||
`getMCPServers`, so a field added later cannot reach one path and miss the other. Both
|
||||
call sites then read as one call each, and a future addition has one place to go.
|
||||
|
||||
Delete the shadowed `wg` at line 529 as part of it.
|
||||
|
||||
## How it is judged
|
||||
|
||||
`make test` stays green.
|
||||
|
||||
The regression that matters is a test asserting the two paths wire the same set. Compare
|
||||
the constructed `daemonAPI` field by field, and assert the worker count started under the
|
||||
outer `wg` matches. Without that, this drifts again the next time a wiring line is added.
|
||||
|
||||
Then confirm on a locked box: unlock by passkey, ask something that needs Nexus, and check
|
||||
`/tools` lists the MCP servers. Both answer wrongly today.
|
||||
|
||||
## Priority
|
||||
|
||||
Latent, not live. `deploy/mavend.json` sets `db_key_env`, so homesrv boots unlocked and
|
||||
takes the correct path. This bites the locked deployment that `docs/operations.md`
|
||||
describes, and it bites silently.
|
||||
@@ -456,9 +456,30 @@ func (c *Config) validate() error {
|
||||
if err := c.validateCapture(); err != nil {
|
||||
return err
|
||||
}
|
||||
if err := c.validateTelegram(); err != nil {
|
||||
return err
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// validateTelegram refuses an intake half that cannot read the chat it is
|
||||
// pointed at. The push half accepts an @channelusername and the intake half
|
||||
// does not, so a box configured with both boots clean, keeps pushing, and
|
||||
// answers nothing — the failure is invisible from the chat. Same shape as
|
||||
// validateNetScan: fail the config rather than the turn.
|
||||
func (c *Config) validateTelegram() error {
|
||||
if c.Telegram == nil || !c.Telegram.Intake {
|
||||
return nil
|
||||
}
|
||||
// An unset ${TELEGRAM_*} expands to empty, and the daemon already reads an
|
||||
// empty token or chat id as telegram not being wired at all. Validating a
|
||||
// block that wires nothing would fail a box that merely has no bot.
|
||||
if c.Telegram.BotToken == "" || c.Telegram.ChatID == "" {
|
||||
return nil
|
||||
}
|
||||
return telegramsink.ValidateIntakeChatID(c.Telegram.ChatID)
|
||||
}
|
||||
|
||||
// DBEncryptionKey resolves the at-rest encryption key: DBKeyEnv (if set) wins
|
||||
// over DBKeyB64. Returns (nil, nil) when neither is set — the caller then opens
|
||||
// a plaintext store. A configured-but-invalid key is an error (fail closed,
|
||||
|
||||
@@ -466,3 +466,29 @@ func TestNormaliseKeepsExplicitWorkstationHealth(t *testing.T) {
|
||||
t.Errorf("Health = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTelegramIntakeRefusesNamedChat(t *testing.T) {
|
||||
// The push half accepts an @channelusername and the intake half cannot use
|
||||
// one, so a box with both boots clean and answers nothing. Refuse the
|
||||
// config instead.
|
||||
p := writeConfig(t, `{"telegram":{"bot_token":"t","chat_id":"@maven","intake":true}}`)
|
||||
if _, err := Load(p); err == nil {
|
||||
t.Fatal("Load succeeded for intake with an @-name chat id; want error")
|
||||
}
|
||||
}
|
||||
|
||||
func TestTelegramNamedChatOKWithoutIntake(t *testing.T) {
|
||||
// Push-only is what the @-name is for, so nothing changes for a box that
|
||||
// never turned intake on.
|
||||
p := writeConfig(t, `{"telegram":{"bot_token":"t","chat_id":"@maven"}}`)
|
||||
if _, err := Load(p); err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTelegramIntakeAcceptsNumericChat(t *testing.T) {
|
||||
p := writeConfig(t, `{"telegram":{"bot_token":"t","chat_id":"-1001234567890","intake":true}}`)
|
||||
if _, err := Load(p); err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -94,20 +94,22 @@ func openTestStore(t *testing.T) *store.Store {
|
||||
|
||||
// attemptStatus reads one attempt row back. Returns ok=false when the row is
|
||||
// gone, which would itself be a broken promise (a dropped attempt).
|
||||
//
|
||||
// It goes through ListDeliveryAttempts rather than raw SQL. This helper used to
|
||||
// reach past the store into store.DB, which was the tell that the outbox was
|
||||
// write-only; the reader landed in V-390 and this caller was not moved over.
|
||||
func attemptStatus(t *testing.T, st *store.Store, id int64) (status string, completed bool, ok bool) {
|
||||
t.Helper()
|
||||
tx, err := st.DB(context.Background())
|
||||
attempts, err := st.ListDeliveryAttempts(context.Background(), "", 200)
|
||||
if err != nil {
|
||||
t.Fatalf("read tx: %v", err)
|
||||
t.Fatalf("ListDeliveryAttempts: %v", err)
|
||||
}
|
||||
defer func() { _ = tx.Rollback() }()
|
||||
var completedTS *int64
|
||||
err = tx.QueryRowContext(context.Background(),
|
||||
`SELECT status, completed_ts FROM delivery_attempts WHERE id = ?`, id).Scan(&status, &completedTS)
|
||||
if err != nil {
|
||||
return "", false, false
|
||||
for _, a := range attempts {
|
||||
if a.ID == id {
|
||||
return a.Status, a.HasComplete, true
|
||||
}
|
||||
}
|
||||
return status, completedTS != nil, true
|
||||
return "", false, false
|
||||
}
|
||||
|
||||
// TestCrashBetweenBeginAndCompleteBecomesUnknown — simulate the crash window:
|
||||
|
||||
@@ -0,0 +1,175 @@
|
||||
// botapi.go — the telegram bot API calls the intake half makes, and the inbound
|
||||
// shapes it reads (V-637). Split out of intake.go so the poller reads as the
|
||||
// policy it is, with the wire in one place under it.
|
||||
package telegramsink
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"io"
|
||||
"log"
|
||||
"net/http"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// getUpdates long-polls. The offset is telegram's own acknowledgement: asking
|
||||
// for lastSeen+1 is what drops everything before it from the queue, so an
|
||||
// update is handled once even across a restart.
|
||||
func (p *Poller) getUpdates(ctx context.Context, timeoutSec int) ([]update, error) {
|
||||
body, err := json.Marshal(map[string]any{
|
||||
"offset": p.offset,
|
||||
"timeout": timeoutSec,
|
||||
"allowed_updates": []string{"message", "callback_query"},
|
||||
})
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
var env struct {
|
||||
telegramResp
|
||||
Result []update `json:"result"`
|
||||
}
|
||||
if err := p.call(ctx, "getUpdates", body, &env); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
for _, u := range env.Result {
|
||||
if u.UpdateID >= p.offset {
|
||||
p.offset = u.UpdateID + 1
|
||||
}
|
||||
}
|
||||
return env.Result, nil
|
||||
}
|
||||
|
||||
func (p *Poller) send(ctx context.Context, text string, kb *inlineKeyboard) error {
|
||||
body, err := json.Marshal(sendMessageReq{
|
||||
ChatID: p.cfgChatID(),
|
||||
Text: text,
|
||||
// A reply to something he just typed is not an alarm, but it is still his
|
||||
// own data in a third party's chat, so it stays unforwardable like the
|
||||
// away messages the sink pushes.
|
||||
ProtectContent: true,
|
||||
ReplyMarkup: kb,
|
||||
})
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
return p.call(ctx, "sendMessage", body, nil)
|
||||
}
|
||||
|
||||
// answerCallback stops the clock on the tapped button. text empty is a silent
|
||||
// acknowledgement; anything else shows as a toast.
|
||||
func (p *Poller) answerCallback(ctx context.Context, id, text string) {
|
||||
body, err := json.Marshal(map[string]any{"callback_query_id": id, "text": text})
|
||||
if err != nil {
|
||||
return
|
||||
}
|
||||
if err := p.call(ctx, "answerCallbackQuery", body, nil); err != nil {
|
||||
log.Printf("telegram intake: answer callback: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// editKeyboard replaces the buttons under a message the bot sent. kb nil takes
|
||||
// them off.
|
||||
func (p *Poller) editKeyboard(ctx context.Context, chatID string, messageID int64, kb *inlineKeyboard) error {
|
||||
payload := map[string]any{"chat_id": chatID, "message_id": messageID}
|
||||
if kb != nil {
|
||||
payload["reply_markup"] = kb
|
||||
} else {
|
||||
payload["reply_markup"] = inlineKeyboard{Rows: [][]inlineButton{}}
|
||||
}
|
||||
body, err := json.Marshal(payload)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
return p.call(ctx, "editMessageReplyMarkup", body, nil)
|
||||
}
|
||||
|
||||
// call posts one bot API method and checks the envelope. out may be nil when
|
||||
// only the ok flag matters. Every error goes through the sink's redaction: the
|
||||
// token is in the URL path because telegram accepts it nowhere else, and
|
||||
// net/http prints that URL in transport errors.
|
||||
func (p *Poller) call(ctx context.Context, method string, body []byte, out any) error {
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodPost,
|
||||
p.sink.base+"/bot"+p.sink.cfg.BotToken+"/"+method, bytes.NewReader(body))
|
||||
if err != nil {
|
||||
return p.sink.redact(err)
|
||||
}
|
||||
req.Header.Set("Content-Type", "application/json")
|
||||
resp, err := p.hc.Do(req)
|
||||
if err != nil {
|
||||
return fmt.Errorf("telegramsink: %s: %w", method, p.sink.redact(err))
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
rb, _ := io.ReadAll(io.LimitReader(resp.Body, maxIntakeRespBytes))
|
||||
|
||||
var tr telegramResp
|
||||
if err := json.Unmarshal(rb, &tr); err != nil {
|
||||
return fmt.Errorf("telegramsink: %s: %d with a body that is not the bot API envelope: %s",
|
||||
method, resp.StatusCode, snippet(rb))
|
||||
}
|
||||
if !tr.Ok {
|
||||
return fmt.Errorf("telegramsink: %s: telegram returned error %d: %s",
|
||||
method, tr.ErrorCode, strings.TrimSpace(tr.Description))
|
||||
}
|
||||
if out == nil {
|
||||
return nil
|
||||
}
|
||||
if err := json.Unmarshal(rb, out); err != nil {
|
||||
return fmt.Errorf("telegramsink: %s: decode result: %w", method, err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// maxIntakeRespBytes — a getUpdates batch carries up to 100 messages, so the
|
||||
// send path's cap is too small here. Still bounded: the body is wire-controlled
|
||||
// and a relay sits in front of it.
|
||||
const maxIntakeRespBytes = 4 << 20
|
||||
|
||||
// The inbound shapes, cut to what the poller reads.
|
||||
type update struct {
|
||||
UpdateID int64 `json:"update_id"`
|
||||
Message *message `json:"message,omitempty"`
|
||||
CallbackQuery *callbackQuery `json:"callback_query,omitempty"`
|
||||
}
|
||||
|
||||
type message struct {
|
||||
MessageID int64 `json:"message_id"`
|
||||
Chat chat `json:"chat"`
|
||||
Text string `json:"text"`
|
||||
}
|
||||
|
||||
type callbackQuery struct {
|
||||
ID string `json:"id"`
|
||||
Data string `json:"data"`
|
||||
Message message `json:"message"`
|
||||
}
|
||||
|
||||
// chat — the id arrives as a JSON number for a user and a string for a channel,
|
||||
// and the config holds whichever was written. json.Number keeps both without
|
||||
// choosing.
|
||||
type chat struct {
|
||||
ID json.Number `json:"id"`
|
||||
Username string `json:"username,omitempty"`
|
||||
}
|
||||
|
||||
func (c chat) idString() string {
|
||||
if s := c.ID.String(); s != "" {
|
||||
return s
|
||||
}
|
||||
if c.Username != "" {
|
||||
return "@" + c.Username
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// inlineKeyboard — the reply_markup shape. Rows of buttons, each carrying
|
||||
// callback data.
|
||||
type inlineKeyboard struct {
|
||||
Rows [][]inlineButton `json:"inline_keyboard"`
|
||||
}
|
||||
|
||||
type inlineButton struct {
|
||||
Text string `json:"text"`
|
||||
Data string `json:"callback_data"`
|
||||
}
|
||||
@@ -0,0 +1,101 @@
|
||||
// correction.go — the correction gesture as it appears in the chat (V-637).
|
||||
// Two taps at most: "не то" opens the seven intents, and one of them writes the
|
||||
// label. The web's version of the same gesture is cmd/mavweb/chat.go.
|
||||
package telegramsink
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"strconv"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// CorrectionTargets — the intents a correction may name, in the order the
|
||||
// buttons are drawn. It mirrors the seven the web offers, and it is a closed
|
||||
// list for the same reason: V-632 fits prototypes from the label table, and a
|
||||
// label nothing can score is worse than no label.
|
||||
var CorrectionTargets = []string{"fact", "note", "reminder", "query", "act", "chat", "system"}
|
||||
|
||||
// correctionKeyboard — the one gesture beside the reply. Nothing when the turn
|
||||
// did not persist: a button that cannot name a row would report a failure the
|
||||
// owner cannot act on.
|
||||
func (p *Poller) correctionKeyboard(traceID int64) *inlineKeyboard {
|
||||
if traceID <= 0 || p.correct == nil {
|
||||
return nil
|
||||
}
|
||||
return &inlineKeyboard{Rows: [][]inlineButton{{
|
||||
{Text: "не то", Data: fmt.Sprintf("%s%d", prefixAsk, traceID)},
|
||||
}}}
|
||||
}
|
||||
|
||||
// targetKeyboard — the seven intents, plus the cheap half kept reachable. He
|
||||
// opened the row without knowing he had to name something, and closing it with
|
||||
// no way out would price the negative he was willing to give.
|
||||
func targetKeyboard(traceID int64) *inlineKeyboard {
|
||||
var rows [][]inlineButton
|
||||
row := []inlineButton{}
|
||||
for _, t := range CorrectionTargets {
|
||||
row = append(row, inlineButton{Text: t, Data: fmt.Sprintf("%s%d:%s", prefixTarget, traceID, t)})
|
||||
if len(row) == 4 {
|
||||
rows, row = append(rows, row), nil
|
||||
}
|
||||
}
|
||||
if len(row) > 0 {
|
||||
rows = append(rows, row)
|
||||
}
|
||||
return &inlineKeyboard{Rows: append(rows, []inlineButton{
|
||||
{Text: "просто неверно", Data: fmt.Sprintf("%s%d:", prefixTarget, traceID)},
|
||||
})}
|
||||
}
|
||||
|
||||
// Callback data is capped at 64 bytes by telegram, so it carries the trace id
|
||||
// and the target and nothing else.
|
||||
const (
|
||||
prefixAsk = "w:"
|
||||
prefixTarget = "t:"
|
||||
)
|
||||
|
||||
type callbackKind int
|
||||
|
||||
const (
|
||||
callbackUnknown callbackKind = iota
|
||||
callbackAskTarget
|
||||
callbackTarget
|
||||
)
|
||||
|
||||
// parseCallback reads button data. An unparseable id, or a target that is not
|
||||
// one of the seven, is callbackUnknown — the data came off the wire, and a
|
||||
// label the fitting code cannot score is worse than no label.
|
||||
func parseCallback(data string) (traceID int64, target string, kind callbackKind) {
|
||||
switch {
|
||||
case strings.HasPrefix(data, prefixAsk):
|
||||
id, err := strconv.ParseInt(strings.TrimPrefix(data, prefixAsk), 10, 64)
|
||||
if err != nil || id <= 0 {
|
||||
return 0, "", callbackUnknown
|
||||
}
|
||||
return id, "", callbackAskTarget
|
||||
case strings.HasPrefix(data, prefixTarget):
|
||||
rest := strings.TrimPrefix(data, prefixTarget)
|
||||
idPart, target, ok := strings.Cut(rest, ":")
|
||||
if !ok {
|
||||
return 0, "", callbackUnknown
|
||||
}
|
||||
id, err := strconv.ParseInt(idPart, 10, 64)
|
||||
if err != nil || id <= 0 {
|
||||
return 0, "", callbackUnknown
|
||||
}
|
||||
if target != "" && !isCorrectionTarget(target) {
|
||||
return 0, "", callbackUnknown
|
||||
}
|
||||
return id, target, callbackTarget
|
||||
}
|
||||
return 0, "", callbackUnknown
|
||||
}
|
||||
|
||||
func isCorrectionTarget(s string) bool {
|
||||
for _, t := range CorrectionTargets {
|
||||
if t == s {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
@@ -0,0 +1,30 @@
|
||||
package telegramsink
|
||||
|
||||
import "testing"
|
||||
|
||||
// Button data comes off the wire. An unparseable id or an intent that is not one
|
||||
// of the seven must not reach the label table V-632 fits prototypes from.
|
||||
func TestParseCallbackRejectsWhatCannotBeALabel(t *testing.T) {
|
||||
for _, data := range []string{
|
||||
"", "nonsense", "w:", "w:0", "w:-3", "w:abc",
|
||||
"t:77", "t:0:note", "t:abc:note", "t:77:погода", "t:77:fact:extra",
|
||||
} {
|
||||
if _, _, kind := parseCallback(data); kind != callbackUnknown {
|
||||
t.Errorf("%q was accepted, want callbackUnknown", data)
|
||||
}
|
||||
}
|
||||
if id, target, kind := parseCallback("t:77:reminder"); id != 77 || target != "reminder" || kind != callbackTarget {
|
||||
t.Errorf("got %d %q %v, want the reminder correction", id, target, kind)
|
||||
}
|
||||
}
|
||||
|
||||
// Every intent the web offers has a button here, so a new intent cannot exist
|
||||
// with no way to correct a chat turn into it.
|
||||
func TestIntakeTargetsAreTheSeven(t *testing.T) {
|
||||
if len(CorrectionTargets) != 7 {
|
||||
t.Fatalf("%d targets, want the seven public intents", len(CorrectionTargets))
|
||||
}
|
||||
if isCorrectionTarget("") {
|
||||
t.Error("empty is the absence of a target, not one of them")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,225 @@
|
||||
// intake.go — the inbound half of the telegram channel (V-637).
|
||||
//
|
||||
// Until this file, telegram was a reach and nothing else: the sink pushes an
|
||||
// away message and the chat has no way to answer. That made the correction
|
||||
// gesture (V-630) reachable from the web and from voice only, and the sample of
|
||||
// labels skews to wherever the owner happens to be standing.
|
||||
//
|
||||
// Long-poll getUpdates, not a webhook. The box takes no inbound connections and
|
||||
// it reaches api.telegram.org through a relay, so the direction of the
|
||||
// connection has to stay outbound. The poller is off unless the telegram block
|
||||
// says intake, and it accepts messages from exactly one chat.
|
||||
package telegramsink
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"log"
|
||||
"net/http"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
// longPollSeconds — how long telegram holds an empty getUpdates open. The HTTP
|
||||
// client's own timeout has to sit above it or every poll ends as a transport
|
||||
// error, which is why the poller does not reuse the sink's client.
|
||||
const longPollSeconds = 25
|
||||
|
||||
// pollBackoff — the wait after a failed poll. The relay going down is the
|
||||
// normal cause and it comes back on its own, so this is a quiet retry rather
|
||||
// than an escalation.
|
||||
const pollBackoff = 15 * time.Second
|
||||
|
||||
// Turn runs one utterance as a turn and reports the reply and the persisted
|
||||
// trace id. traceID 0 means nothing persisted, and then the reply carries no
|
||||
// correction buttons — there is no row for them to point at.
|
||||
type Turn func(ctx context.Context, conversation, text string) (reply string, traceID int64, err error)
|
||||
|
||||
// Correct records the owner's correction of one turn. shouldBe empty is the
|
||||
// cheap half of the gesture: wrong, target unstated.
|
||||
type Correct func(ctx context.Context, traceID int64, shouldBe string) error
|
||||
|
||||
// Poller reads the configured chat and answers in it. One per daemon.
|
||||
type Poller struct {
|
||||
sink *Sink
|
||||
turn Turn
|
||||
correct Correct
|
||||
hc *http.Client
|
||||
offset int64
|
||||
}
|
||||
|
||||
// ValidateIntakeChatID refuses a chat id the intake half cannot use. The push
|
||||
// half accepts @channelusername as a destination. The intake half cannot: an
|
||||
// inbound update names its chat by numeric id, so an @-name would match nothing
|
||||
// and the poller would read the chat and answer none of it. Config validation
|
||||
// calls this, so the box refuses to boot rather than running a dead reach —
|
||||
// NewPoller returning an error is too late, because the daemon is already up.
|
||||
func ValidateIntakeChatID(chatID string) error {
|
||||
id := strings.TrimSpace(chatID)
|
||||
if id == "" {
|
||||
return errors.New("telegramsink: intake needs a chat id")
|
||||
}
|
||||
digits := strings.TrimPrefix(id, "-")
|
||||
if digits == "" || strings.TrimLeft(digits, "0123456789") != "" {
|
||||
return fmt.Errorf("telegramsink: intake needs the numeric chat id, not %s", chatID)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// NewPoller builds the intake half around an already-validated sink, so the
|
||||
// token, the base URL and the relay are resolved in one place. turn is
|
||||
// required; correct may be nil, and then the reply carries no buttons.
|
||||
func NewPoller(s *Sink, turn Turn, correct Correct) (*Poller, error) {
|
||||
if s == nil {
|
||||
return nil, errors.New("telegramsink: intake needs a sink")
|
||||
}
|
||||
if turn == nil {
|
||||
return nil, errors.New("telegramsink: intake needs a turn handler")
|
||||
}
|
||||
if err := ValidateIntakeChatID(s.cfg.ChatID); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
// The sink's transport already carries the relay. Only the timeout differs,
|
||||
// and it has to clear the long poll.
|
||||
hc := &http.Client{
|
||||
Timeout: (longPollSeconds + 10) * time.Second,
|
||||
Transport: s.hc.Transport,
|
||||
}
|
||||
return &Poller{sink: s, turn: turn, correct: correct, hc: hc}, nil
|
||||
}
|
||||
|
||||
// Run polls until the context ends. It never returns an error: a chat that
|
||||
// cannot be read is a degraded reach, not a reason to stop the daemon.
|
||||
func (p *Poller) Run(ctx context.Context) {
|
||||
p.discardBacklog(ctx)
|
||||
log.Printf("telegram intake: reading chat %s", p.sink.cfg.ChatID)
|
||||
for ctx.Err() == nil {
|
||||
updates, err := p.getUpdates(ctx, longPollSeconds)
|
||||
if err != nil {
|
||||
if ctx.Err() != nil {
|
||||
return
|
||||
}
|
||||
log.Printf("telegram intake: poll: %v", err)
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return
|
||||
case <-time.After(pollBackoff):
|
||||
}
|
||||
continue
|
||||
}
|
||||
for _, u := range updates {
|
||||
p.handle(ctx, u)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// discardBacklog moves the offset past whatever is already queued, without
|
||||
// acting on any of it.
|
||||
//
|
||||
// Telegram holds undelivered updates for 24 hours, so a daemon that was down
|
||||
// overnight would otherwise wake up and answer every question in order. A
|
||||
// question asked eight hours ago has been answered by the owner himself or has
|
||||
// stopped mattering, and a reminder set from it would land at the wrong time.
|
||||
// Missing it is the safe direction.
|
||||
func (p *Poller) discardBacklog(ctx context.Context) {
|
||||
// getUpdates returns at most 100 per call, so one call is not the queue. The
|
||||
// loop is bounded rather than "until empty": the timeout is 0, so an instance
|
||||
// that keeps handing back a full batch would spin, and a thousand skipped
|
||||
// messages is already a box that was down for a long time.
|
||||
skipped := 0
|
||||
for range 10 {
|
||||
updates, err := p.getUpdates(ctx, 0)
|
||||
if err != nil {
|
||||
// Not fatal. The offset stays where it was, so the first real poll sees
|
||||
// what is left and answers it late. Say so rather than hide it.
|
||||
log.Printf("telegram intake: could not skip the backlog, old messages may be answered: %v", err)
|
||||
return
|
||||
}
|
||||
skipped += len(updates)
|
||||
if len(updates) == 0 {
|
||||
break
|
||||
}
|
||||
}
|
||||
if skipped > 0 {
|
||||
log.Printf("telegram intake: skipped %d message(s) queued while the daemon was down", skipped)
|
||||
}
|
||||
}
|
||||
|
||||
// handle dispatches one update. Anything that is neither a message from the
|
||||
// owner's chat nor a callback on one of Maven's own keyboards is dropped in
|
||||
// silence: a reply to a stranger confirms the bot exists and who it belongs to.
|
||||
func (p *Poller) handle(ctx context.Context, u update) {
|
||||
switch {
|
||||
case u.CallbackQuery != nil:
|
||||
p.onCallback(ctx, u.CallbackQuery)
|
||||
case u.Message != nil:
|
||||
p.onMessage(ctx, u.Message)
|
||||
}
|
||||
}
|
||||
|
||||
func (p *Poller) onMessage(ctx context.Context, m *message) {
|
||||
text := strings.TrimSpace(m.Text)
|
||||
if text == "" || !p.fromOwner(m.Chat.idString()) {
|
||||
return
|
||||
}
|
||||
// The conversation id keys the dialogue, so a clarify question asked in the
|
||||
// chat is not answered by an utterance typed on the web.
|
||||
reply, traceID, err := p.turn(ctx, "telegram:"+m.Chat.idString(), text)
|
||||
if err != nil {
|
||||
log.Printf("telegram intake: turn: %v", err)
|
||||
return
|
||||
}
|
||||
if strings.TrimSpace(reply) == "" {
|
||||
return
|
||||
}
|
||||
if err := p.send(ctx, reply, p.correctionKeyboard(traceID)); err != nil {
|
||||
log.Printf("telegram intake: reply: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// onCallback handles a tap on a correction button. Every path from the owner
|
||||
// answers the callback: telegram spins a clock on the button until it is
|
||||
// answered, and an unanswered tap reads as a gesture that was dropped. A tap
|
||||
// from anyone else gets silence, the same as a message from a stranger.
|
||||
func (p *Poller) onCallback(ctx context.Context, cb *callbackQuery) {
|
||||
if !p.fromOwner(cb.Message.Chat.idString()) {
|
||||
return
|
||||
}
|
||||
traceID, target, kind := parseCallback(cb.Data)
|
||||
if kind == callbackUnknown || p.correct == nil {
|
||||
p.answerCallback(ctx, cb.ID, "")
|
||||
return
|
||||
}
|
||||
// A tap on "не то" only opens the second row. Nothing is written yet: the
|
||||
// target is worth much more than the negative, so he gets the chance to name
|
||||
// it before the gesture is spent.
|
||||
if kind == callbackAskTarget {
|
||||
p.answerCallback(ctx, cb.ID, "")
|
||||
if err := p.editKeyboard(ctx, cb.Message.Chat.idString(), cb.Message.MessageID, targetKeyboard(traceID)); err != nil {
|
||||
log.Printf("telegram intake: open the target row: %v", err)
|
||||
}
|
||||
return
|
||||
}
|
||||
if err := p.correct(ctx, traceID, target); err != nil {
|
||||
log.Printf("telegram intake: correct turn %d: %v", traceID, err)
|
||||
p.answerCallback(ctx, cb.ID, "не записалось")
|
||||
return
|
||||
}
|
||||
p.answerCallback(ctx, cb.ID, "записала")
|
||||
// The buttons come off, because the correction is given and a live keyboard
|
||||
// on an answered turn invites correcting it twice.
|
||||
if err := p.editKeyboard(ctx, cb.Message.Chat.idString(), cb.Message.MessageID, nil); err != nil {
|
||||
log.Printf("telegram intake: clear the keyboard: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// fromOwner — one chat, and it is the one the sink already sends to. Telegram
|
||||
// chat ids are not guessable, but they are also not secret: they travel in
|
||||
// every forwarded message. So this is the whole authorisation and it is an
|
||||
// allowlist of one.
|
||||
func (p *Poller) fromOwner(chatID string) bool {
|
||||
return chatID != "" && chatID == p.cfgChatID()
|
||||
}
|
||||
|
||||
func (p *Poller) cfgChatID() string { return strings.TrimSpace(p.sink.cfg.ChatID) }
|
||||
@@ -0,0 +1,216 @@
|
||||
package telegramsink
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// The turn he types in the chat is the turn the web would run, and the reply
|
||||
// carries the one gesture beside it.
|
||||
func TestIntakeRunsTheTurnAndOffersTheCorrection(t *testing.T) {
|
||||
b := newFakeBot(t)
|
||||
rec := &recorder{reply: "поняла", traceID: 91}
|
||||
p := newTestPoller(t, b, rec)
|
||||
|
||||
p.handle(context.Background(), msg(ownerChat, " поужинал "))
|
||||
|
||||
if got := rec.took(); len(got) != 1 || got[0] != "поужинал" {
|
||||
t.Fatalf("turns %q, want the trimmed utterance once", got)
|
||||
}
|
||||
// The dialogue is keyed per chat, so a clarify asked here is not answered on
|
||||
// the web.
|
||||
if rec.conversation != "telegram:"+ownerChat {
|
||||
t.Errorf("conversation %q does not name the chat", rec.conversation)
|
||||
}
|
||||
sends := b.called("sendMessage")
|
||||
if len(sends) != 1 {
|
||||
t.Fatalf("%d sends, want 1", len(sends))
|
||||
}
|
||||
if sends[0].body["text"] != "поняла" {
|
||||
t.Errorf("sent %v, want the reply", sends[0].body["text"])
|
||||
}
|
||||
if sends[0].body["protect_content"] != true {
|
||||
t.Error("his own data went out forwardable")
|
||||
}
|
||||
kb, _ := json.Marshal(sends[0].body["reply_markup"])
|
||||
if !strings.Contains(string(kb), "w:91") {
|
||||
t.Errorf("keyboard %s does not point at the turn's trace", kb)
|
||||
}
|
||||
}
|
||||
|
||||
// A turn nothing persisted has no row to correct, and a button that would name
|
||||
// one reports a failure he cannot act on.
|
||||
func TestIntakeSkipsTheGestureWithNoTrace(t *testing.T) {
|
||||
b := newFakeBot(t)
|
||||
p := newTestPoller(t, b, &recorder{reply: "поняла", traceID: 0})
|
||||
|
||||
p.handle(context.Background(), msg(ownerChat, "привет"))
|
||||
|
||||
sends := b.called("sendMessage")
|
||||
if len(sends) != 1 {
|
||||
t.Fatalf("%d sends, want 1", len(sends))
|
||||
}
|
||||
if _, ok := sends[0].body["reply_markup"]; ok {
|
||||
t.Error("offered a correction on a turn with no trace")
|
||||
}
|
||||
}
|
||||
|
||||
// One chat, and a stranger is not answered at all: a reply confirms the bot
|
||||
// exists and whose it is.
|
||||
func TestIntakeIgnoresAnyOtherChat(t *testing.T) {
|
||||
b := newFakeBot(t)
|
||||
rec := &recorder{reply: "поняла", traceID: 5}
|
||||
p := newTestPoller(t, b, rec)
|
||||
|
||||
p.handle(context.Background(), msg("9999", "включи свет"))
|
||||
p.handle(context.Background(), update{UpdateID: 8, CallbackQuery: &callbackQuery{
|
||||
ID: "cb", Data: "t:5:note", Message: message{Chat: chat{ID: json.Number("9999")}},
|
||||
}})
|
||||
|
||||
if got := rec.took(); len(got) != 0 {
|
||||
t.Errorf("ran %q for a chat that is not the owner's", got)
|
||||
}
|
||||
if len(rec.corrections) != 0 {
|
||||
t.Errorf("wrote %v from a chat that is not the owner's", rec.corrections)
|
||||
}
|
||||
if len(b.calls) != 0 {
|
||||
t.Errorf("answered a stranger: %v", b.calls)
|
||||
}
|
||||
}
|
||||
|
||||
// Tapping "не то" opens the seven and writes nothing yet. The target is worth
|
||||
// much more than the negative, so it must not be spent before he can name it.
|
||||
func TestIntakeFirstTapOnlyOpensTheTargets(t *testing.T) {
|
||||
b := newFakeBot(t)
|
||||
rec := &recorder{}
|
||||
p := newTestPoller(t, b, rec)
|
||||
|
||||
p.handle(context.Background(), update{UpdateID: 9, CallbackQuery: &callbackQuery{
|
||||
ID: "cb", Data: "w:77", Message: message{MessageID: 11, Chat: chat{ID: json.Number(ownerChat)}},
|
||||
}})
|
||||
|
||||
if len(rec.corrections) != 0 {
|
||||
t.Fatalf("wrote %v before he named a target", rec.corrections)
|
||||
}
|
||||
if len(b.called("answerCallbackQuery")) != 1 {
|
||||
t.Error("left the clock spinning on the button")
|
||||
}
|
||||
edits := b.called("editMessageReplyMarkup")
|
||||
if len(edits) != 1 {
|
||||
t.Fatalf("%d edits, want the target row", len(edits))
|
||||
}
|
||||
kb, _ := json.Marshal(edits[0].body["reply_markup"])
|
||||
for _, want := range CorrectionTargets {
|
||||
if !strings.Contains(string(kb), `"`+want+`"`) {
|
||||
t.Errorf("target row %s is missing %s", kb, want)
|
||||
}
|
||||
}
|
||||
// And the way out, because he opened the row without knowing he had to name
|
||||
// anything.
|
||||
if !strings.Contains(string(kb), `"t:77:"`) {
|
||||
t.Errorf("target row %s prices out the untargeted negative", kb)
|
||||
}
|
||||
}
|
||||
|
||||
func TestIntakeWritesTheCorrection(t *testing.T) {
|
||||
for _, tc := range []struct {
|
||||
name, data, want string
|
||||
}{
|
||||
{"with a target", "t:77:note", "note"},
|
||||
{"untargeted", "t:77:", ""},
|
||||
} {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
b := newFakeBot(t)
|
||||
rec := &recorder{}
|
||||
p := newTestPoller(t, b, rec)
|
||||
|
||||
p.handle(context.Background(), update{UpdateID: 9, CallbackQuery: &callbackQuery{
|
||||
ID: "cb", Data: tc.data, Message: message{MessageID: 11, Chat: chat{ID: json.Number(ownerChat)}},
|
||||
}})
|
||||
|
||||
if len(rec.corrections) != 1 || rec.corrections[0] != (correction{77, tc.want}) {
|
||||
t.Fatalf("corrections %v, want trace 77 → %q", rec.corrections, tc.want)
|
||||
}
|
||||
// The buttons come off once the gesture is given.
|
||||
edits := b.called("editMessageReplyMarkup")
|
||||
if len(edits) != 1 {
|
||||
t.Fatalf("%d edits, want the keyboard cleared", len(edits))
|
||||
}
|
||||
kb, _ := json.Marshal(edits[0].body["reply_markup"])
|
||||
if strings.Contains(string(kb), "t:77") {
|
||||
t.Errorf("keyboard %s still invites a second correction", kb)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// A write that failed says so on the button. Silence would read as recorded.
|
||||
func TestIntakeSaysWhenTheLabelDidNotLand(t *testing.T) {
|
||||
b := newFakeBot(t)
|
||||
rec := &recorder{correctErr: errors.New("no such routing trace")}
|
||||
p := newTestPoller(t, b, rec)
|
||||
|
||||
p.handle(context.Background(), update{UpdateID: 9, CallbackQuery: &callbackQuery{
|
||||
ID: "cb", Data: "t:77:fact", Message: message{MessageID: 11, Chat: chat{ID: json.Number(ownerChat)}},
|
||||
}})
|
||||
|
||||
answers := b.called("answerCallbackQuery")
|
||||
if len(answers) != 1 || answers[0].body["text"] == "" {
|
||||
t.Fatalf("answers %v, want a toast saying it did not land", answers)
|
||||
}
|
||||
if len(b.called("editMessageReplyMarkup")) != 0 {
|
||||
t.Error("cleared the buttons after a failed write, so he cannot try again")
|
||||
}
|
||||
}
|
||||
|
||||
// A question asked while the daemon was down has been answered by him or has
|
||||
// stopped mattering, and a reminder set from it would land at the wrong time.
|
||||
func TestIntakeDiscardsTheBacklog(t *testing.T) {
|
||||
b := newFakeBot(t, []update{msg(ownerChat, "напомни в 7 позвонить маме")})
|
||||
rec := &recorder{reply: "поняла", traceID: 3}
|
||||
p := newTestPoller(t, b, rec)
|
||||
|
||||
p.discardBacklog(context.Background())
|
||||
|
||||
if got := rec.took(); len(got) != 0 {
|
||||
t.Errorf("answered %q from the overnight queue", got)
|
||||
}
|
||||
// And the offset moved past it, so the next poll does not see it again.
|
||||
if p.offset != 8 {
|
||||
t.Errorf("offset %d, want the skipped update acknowledged", p.offset)
|
||||
}
|
||||
}
|
||||
|
||||
// The poller does not start without somewhere to send the turn.
|
||||
func TestNewPollerNeedsATurn(t *testing.T) {
|
||||
sink, err := New(Config{BotToken: "t", ChatID: ownerChat})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if _, err := NewPoller(sink, nil, nil); err == nil {
|
||||
t.Error("built a poller that reads the chat and answers nothing")
|
||||
}
|
||||
if _, err := NewPoller(nil, func(context.Context, string, string) (string, int64, error) {
|
||||
return "", 0, nil
|
||||
}, nil); err == nil {
|
||||
t.Error("built a poller with no sink to answer through")
|
||||
}
|
||||
}
|
||||
|
||||
// A chat id the intake half cannot match is refused before anything reads the
|
||||
// chat. Config validation calls the same check, so this is the boot error.
|
||||
func TestValidateIntakeChatID(t *testing.T) {
|
||||
for _, ok := range []string{"123", "-1001234567890", " 42 "} {
|
||||
if err := ValidateIntakeChatID(ok); err != nil {
|
||||
t.Errorf("ValidateIntakeChatID(%q): %v", ok, err)
|
||||
}
|
||||
}
|
||||
for _, bad := range []string{"", "@maven", "-", "12a", "1 2"} {
|
||||
if err := ValidateIntakeChatID(bad); err == nil {
|
||||
t.Errorf("ValidateIntakeChatID(%q) accepted; want error", bad)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,123 @@
|
||||
// intakeharness_test.go — a fake bot API and a recorder for what the poller
|
||||
// asked the daemon to do. Shared by the intake tests beside it.
|
||||
package telegramsink
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"io"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"strings"
|
||||
"sync"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// fakeBot stands in for the bot API. It hands out queued updates once, records
|
||||
// every other call, and answers the ok=true envelope the poller checks.
|
||||
type fakeBot struct {
|
||||
mu sync.Mutex
|
||||
updates [][]update // one batch per getUpdates call, then empty
|
||||
calls []botCall
|
||||
srv *httptest.Server
|
||||
}
|
||||
|
||||
type botCall struct {
|
||||
method string
|
||||
body map[string]any
|
||||
}
|
||||
|
||||
func newFakeBot(t *testing.T, batches ...[]update) *fakeBot {
|
||||
t.Helper()
|
||||
b := &fakeBot{updates: batches}
|
||||
b.srv = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
method := r.URL.Path[strings.LastIndex(r.URL.Path, "/")+1:]
|
||||
raw, _ := io.ReadAll(r.Body)
|
||||
var body map[string]any
|
||||
_ = json.Unmarshal(raw, &body)
|
||||
|
||||
b.mu.Lock()
|
||||
b.calls = append(b.calls, botCall{method: method, body: body})
|
||||
var batch []update
|
||||
if method == "getUpdates" && len(b.updates) > 0 {
|
||||
batch, b.updates = b.updates[0], b.updates[1:]
|
||||
}
|
||||
b.mu.Unlock()
|
||||
|
||||
w.Header().Set("Content-Type", "application/json")
|
||||
_ = json.NewEncoder(w).Encode(map[string]any{"ok": true, "result": batch})
|
||||
}))
|
||||
t.Cleanup(b.srv.Close)
|
||||
return b
|
||||
}
|
||||
|
||||
func (b *fakeBot) called(method string) []botCall {
|
||||
b.mu.Lock()
|
||||
defer b.mu.Unlock()
|
||||
var out []botCall
|
||||
for _, c := range b.calls {
|
||||
if c.method == method {
|
||||
out = append(out, c)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// recorder collects what the poller asked the daemon to do.
|
||||
type recorder struct {
|
||||
mu sync.Mutex
|
||||
turns []string
|
||||
conversation string
|
||||
traceID int64
|
||||
corrections []correction
|
||||
reply string
|
||||
err error
|
||||
correctErr error
|
||||
}
|
||||
|
||||
type correction struct {
|
||||
traceID int64
|
||||
shouldBe string
|
||||
}
|
||||
|
||||
func (r *recorder) turn(_ context.Context, conversation, text string) (string, int64, error) {
|
||||
r.mu.Lock()
|
||||
defer r.mu.Unlock()
|
||||
r.turns = append(r.turns, text)
|
||||
r.conversation = conversation
|
||||
return r.reply, r.traceID, r.err
|
||||
}
|
||||
|
||||
func (r *recorder) correct(_ context.Context, traceID int64, shouldBe string) error {
|
||||
r.mu.Lock()
|
||||
defer r.mu.Unlock()
|
||||
r.corrections = append(r.corrections, correction{traceID, shouldBe})
|
||||
return r.correctErr
|
||||
}
|
||||
|
||||
func (r *recorder) took() []string {
|
||||
r.mu.Lock()
|
||||
defer r.mu.Unlock()
|
||||
return append([]string(nil), r.turns...)
|
||||
}
|
||||
|
||||
const ownerChat = "4242"
|
||||
|
||||
func newTestPoller(t *testing.T, b *fakeBot, rec *recorder) *Poller {
|
||||
t.Helper()
|
||||
sink, err := New(Config{BotToken: "secret-token", ChatID: ownerChat, BaseURL: b.srv.URL})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
p, err := NewPoller(sink, rec.turn, rec.correct)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return p
|
||||
}
|
||||
|
||||
func msg(chatID, text string) update {
|
||||
return update{UpdateID: 7, Message: &message{
|
||||
MessageID: 11, Text: text, Chat: chat{ID: json.Number(chatID)},
|
||||
}}
|
||||
}
|
||||
@@ -72,6 +72,13 @@ type Config struct {
|
||||
// Timeout — per-request; 0 = DefaultTimeout. a dead relay can't hang the
|
||||
// tick loop.
|
||||
Timeout time.Duration
|
||||
|
||||
// Intake — read the chat as well as write to it (V-637). Off by default,
|
||||
// like the search and weather blocks: a bot that only pushes cannot be
|
||||
// talked into anything, and turning that off has to stay a deletion. When
|
||||
// set, a message from ChatID becomes a turn and its reply carries the
|
||||
// correction gesture. ChatID is the only accepted sender.
|
||||
Intake bool `json:"intake,omitempty"`
|
||||
}
|
||||
|
||||
// Sink — implements delivery.Sink via the telegram bot sendMessage API. one
|
||||
@@ -130,6 +137,11 @@ type sendMessageReq struct {
|
||||
Text string `json:"text"`
|
||||
DisableNotification bool `json:"disable_notification"` // false = ring (always — these are alarms)
|
||||
ProtectContent bool `json:"protect_content"` // true = no forwarding out of chat
|
||||
|
||||
// ReplyMarkup — the inline keyboard, used only by the intake half (V-637):
|
||||
// a reply to a turn he typed carries the correction gesture. nil on every
|
||||
// push the sink sends, and omitted from the wire when nil.
|
||||
ReplyMarkup *inlineKeyboard `json:"reply_markup,omitempty"`
|
||||
}
|
||||
|
||||
// telegramResp — the shape telegram returns. ok=false on logical error with
|
||||
|
||||
@@ -0,0 +1,192 @@
|
||||
package ipc
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"net"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// A cancelled context has to abort a call that is already in flight. It did not
|
||||
// until V-638: call checked ctx once before sending and then blocked in
|
||||
// roundtrip with no connection deadline, so a daemon that read the frame and
|
||||
// never answered parked the caller for as long as the socket stayed open.
|
||||
//
|
||||
// The server here is that daemon: it accepts, reads nothing, replies nothing.
|
||||
|
||||
func deafServer(t *testing.T) string {
|
||||
t.Helper()
|
||||
sock := filepath.Join(t.TempDir(), "deaf.sock")
|
||||
ln, err := net.Listen("unix", sock)
|
||||
if err != nil {
|
||||
t.Fatalf("listen: %v", err)
|
||||
}
|
||||
t.Cleanup(func() { _ = ln.Close() })
|
||||
go func() {
|
||||
for {
|
||||
conn, err := ln.Accept()
|
||||
if err != nil {
|
||||
return
|
||||
}
|
||||
// Hold it open and say nothing. Closed by the listener cleanup.
|
||||
t.Cleanup(func() { _ = conn.Close() })
|
||||
}
|
||||
}()
|
||||
return sock
|
||||
}
|
||||
|
||||
func TestClientCancelAbortsAReadInFlight(t *testing.T) {
|
||||
c, err := Dial(deafServer(t))
|
||||
if err != nil {
|
||||
t.Fatalf("dial: %v", err)
|
||||
}
|
||||
defer c.Close()
|
||||
|
||||
ctx, cancel := context.WithCancel(context.Background())
|
||||
go func() {
|
||||
time.Sleep(50 * time.Millisecond)
|
||||
cancel()
|
||||
}()
|
||||
|
||||
done := make(chan error, 1)
|
||||
go func() {
|
||||
_, err := c.Ping(ctx)
|
||||
done <- err
|
||||
}()
|
||||
|
||||
select {
|
||||
case err := <-done:
|
||||
// Ping is read-only, so the cancellation is reported as itself rather
|
||||
// than as an ambiguous mutation.
|
||||
if !errors.Is(err, context.Canceled) {
|
||||
t.Errorf("got %v, want context.Canceled", err)
|
||||
}
|
||||
case <-time.After(5 * time.Second):
|
||||
t.Fatal("a cancelled Ping did not return")
|
||||
}
|
||||
}
|
||||
|
||||
// A mutation cancelled while awaiting the reply may already have committed, so
|
||||
// it is ErrAmbiguousOutcome and never a retry. That split is the invariant
|
||||
// internal/ipc/maperr_test.go's neighbours rest on.
|
||||
func TestClientCancelLeavesAMutationAmbiguous(t *testing.T) {
|
||||
c, err := Dial(deafServer(t))
|
||||
if err != nil {
|
||||
t.Fatalf("dial: %v", err)
|
||||
}
|
||||
defer c.Close()
|
||||
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 50*time.Millisecond)
|
||||
defer cancel()
|
||||
|
||||
done := make(chan error, 1)
|
||||
go func() {
|
||||
_, err := c.WriteFact(ctx, WriteFactReq{Key: "water", Value: "drank"})
|
||||
done <- err
|
||||
}()
|
||||
|
||||
select {
|
||||
case err := <-done:
|
||||
if !errors.Is(err, ErrAmbiguousOutcome) {
|
||||
t.Errorf("got %v, want ErrAmbiguousOutcome", err)
|
||||
}
|
||||
case <-time.After(5 * time.Second):
|
||||
t.Fatal("a cancelled WriteFact did not return")
|
||||
}
|
||||
}
|
||||
|
||||
// The deadline itself, with no cancellation: a call on a context with no
|
||||
// deadline used to have no bound at all. This one has one and must respect it.
|
||||
func TestClientDeadlineBoundsACall(t *testing.T) {
|
||||
c, err := Dial(deafServer(t))
|
||||
if err != nil {
|
||||
t.Fatalf("dial: %v", err)
|
||||
}
|
||||
defer c.Close()
|
||||
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 100*time.Millisecond)
|
||||
defer cancel()
|
||||
|
||||
start := time.Now()
|
||||
if _, err := c.Ping(ctx); err == nil {
|
||||
t.Fatal("a deaf server answered a Ping")
|
||||
}
|
||||
if elapsed := time.Since(start); elapsed > 3*time.Second {
|
||||
t.Errorf("Ping took %v, want the context deadline to bound it", elapsed)
|
||||
}
|
||||
}
|
||||
|
||||
// blockingAPI parks Presence until its context is cancelled and records what
|
||||
// cancelled it. Every other method is the unimplemented floor.
|
||||
type blockingAPI struct {
|
||||
UnimplementedCoreAPI
|
||||
entered chan struct{}
|
||||
err chan error
|
||||
}
|
||||
|
||||
func (b *blockingAPI) Presence(ctx context.Context) (Presence, error) {
|
||||
close(b.entered)
|
||||
<-ctx.Done()
|
||||
b.err <- ctx.Err()
|
||||
return Presence{}, ctx.Err()
|
||||
}
|
||||
|
||||
// serveConn dispatched under context.Background() until V-638, so Close could
|
||||
// only abandon a dispatch in flight and never tell it to stop.
|
||||
func TestServerCloseCancelsADispatchInFlight(t *testing.T) {
|
||||
api := &blockingAPI{entered: make(chan struct{}), err: make(chan error, 1)}
|
||||
srv, err := Listen(filepath.Join(t.TempDir(), "core.sock"), api)
|
||||
if err != nil {
|
||||
t.Fatalf("listen: %v", err)
|
||||
}
|
||||
served := make(chan struct{})
|
||||
go func() { _ = srv.Serve(); close(served) }()
|
||||
|
||||
cli, err := Dial(srv.Path())
|
||||
if err != nil {
|
||||
t.Fatalf("dial: %v", err)
|
||||
}
|
||||
defer cli.Close()
|
||||
go func() { _, _ = cli.Presence(context.Background()) }()
|
||||
|
||||
select {
|
||||
case <-api.entered:
|
||||
case <-time.After(5 * time.Second):
|
||||
t.Fatal("the handler was never dispatched")
|
||||
}
|
||||
|
||||
_ = srv.Close()
|
||||
<-served
|
||||
select {
|
||||
case got := <-api.err:
|
||||
if !errors.Is(got, context.Canceled) {
|
||||
t.Errorf("handler saw %v, want context.Canceled", got)
|
||||
}
|
||||
case <-time.After(5 * time.Second):
|
||||
t.Fatal("Close did not cancel the dispatch")
|
||||
}
|
||||
}
|
||||
|
||||
// The watchdog closes the conn, and it races the end of the call: a
|
||||
// cancellation landing as the reply arrives can close a conn the call was
|
||||
// already done with. That is survivable either way, because a write to a closed
|
||||
// socket is errWriteLost and errWriteLost re-dials and retries, so this test
|
||||
// passes with or without the drop in roundtrip's defer. What it pins is that
|
||||
// the recovery is real and costs one round trip at most, never an error the
|
||||
// caller sees.
|
||||
func TestClientSurvivesACancelledCall(t *testing.T) {
|
||||
_, _, cli, _ := newServerWithStore(t)
|
||||
|
||||
for i := 0; i < 20; i++ {
|
||||
ctx, cancel := context.WithCancel(context.Background())
|
||||
go cancel() // races the reply on purpose
|
||||
_, _ = cli.Ping(ctx)
|
||||
cancel()
|
||||
|
||||
if _, err := cli.Ping(context.Background()); err != nil {
|
||||
t.Fatalf("call %d after a cancelled one: %v", i, err)
|
||||
}
|
||||
}
|
||||
}
|
||||
+94
-12
@@ -26,9 +26,20 @@ type Client struct {
|
||||
conn net.Conn
|
||||
path string // the address as configured, kept for errors and logs
|
||||
addr netaddr.Addr // parsed, so a dropped conn can be re-dialed (core restart)
|
||||
mu sync.Mutex
|
||||
mu sync.Mutex // one request at a time, so a frame and its reply pair up
|
||||
|
||||
// connMu guards the conn field alone, and is held only across an assignment
|
||||
// or a read. It exists so Close and the cancellation watchdog can reach the
|
||||
// connection without waiting for the call that is holding c.mu (V-638).
|
||||
connMu sync.Mutex
|
||||
}
|
||||
|
||||
// defaultCallTimeout bounds a call whose context carries no deadline. It is
|
||||
// the same 120s internal/voice/client.go settles on: long enough for a model
|
||||
// call on a cold resident model, short enough that a daemon which stopped
|
||||
// answering does not park the caller forever.
|
||||
const defaultCallTimeout = 120 * time.Second
|
||||
|
||||
// errWriteLost marks a conn drop while sending the request frame: the request
|
||||
// never reached the server (or the server never saw a complete frame), so
|
||||
// retrying is always safe regardless of method — nothing was applied to
|
||||
@@ -103,11 +114,22 @@ func Dial(path string) (*Client, error) {
|
||||
return &Client{conn: c, path: path, addr: addr}, nil
|
||||
}
|
||||
|
||||
// Close closes the connection out from under a call in flight, on purpose: a
|
||||
// shutdown must not wait out a parked read. It takes connMu and never c.mu, so
|
||||
// it cannot block behind the call it is interrupting.
|
||||
//
|
||||
// The lock is taken and released by hand, around the two field accesses and
|
||||
// nothing else. The socket close happens outside it, because a close on a tcp
|
||||
// conn can block and connMu is on the path of every call.
|
||||
func (c *Client) Close() error {
|
||||
if c.conn == nil {
|
||||
c.connMu.Lock()
|
||||
conn := c.conn
|
||||
c.conn = nil
|
||||
c.connMu.Unlock()
|
||||
if conn == nil {
|
||||
return nil
|
||||
}
|
||||
return c.conn.Close()
|
||||
return conn.Close()
|
||||
}
|
||||
|
||||
// DialWait is Dial with patience: it retries with capped backoff until the
|
||||
@@ -163,16 +185,26 @@ func (c *Client) call(ctx context.Context, m Method, params, result any) error {
|
||||
}
|
||||
|
||||
var resp Response
|
||||
err := c.roundtrip(m, raw, &resp)
|
||||
err := c.roundtrip(ctx, m, raw, &resp)
|
||||
switch {
|
||||
case errors.Is(err, errWriteLost):
|
||||
// The request never left; a duplicate send can't double-apply.
|
||||
// Redial (roundtrip re-dials on a nil conn) and retry exactly once.
|
||||
err = c.roundtrip(m, raw, &resp)
|
||||
// Not when the caller has given up — a retry would only be a second
|
||||
// frame nobody is waiting for.
|
||||
if ctx.Err() == nil {
|
||||
err = c.roundtrip(ctx, m, raw, &resp)
|
||||
}
|
||||
case errors.Is(err, errReadLost):
|
||||
if readOnlyMethods[m] {
|
||||
if ctx.Err() != nil {
|
||||
// The caller cancelled the read it was waiting for. Nothing
|
||||
// was applied, so this is the cancellation and not an
|
||||
// ambiguity.
|
||||
return ctx.Err()
|
||||
}
|
||||
// A duplicate read can't double-apply either — safe to replay.
|
||||
err = c.roundtrip(m, raw, &resp)
|
||||
err = c.roundtrip(ctx, m, raw, &resp)
|
||||
} else {
|
||||
// The mutation may have already committed server-side. Do not
|
||||
// retry: report the ambiguity instead of guessing.
|
||||
@@ -199,19 +231,55 @@ func (c *Client) call(ctx context.Context, m Method, params, result any) error {
|
||||
// failure is wrapped in errReadLost (ambiguous — call() only retries it for
|
||||
// read-only methods). Either way a failed conn is dropped so the next call
|
||||
// re-dials clean. Caller holds c.mu.
|
||||
func (c *Client) roundtrip(m Method, raw json.RawMessage, resp *Response) error {
|
||||
if c.conn == nil {
|
||||
conn, err := netaddr.Dial(c.addr)
|
||||
//
|
||||
// The connection carries a deadline derived from ctx, falling back to
|
||||
// defaultCallTimeout, and a watchdog closes it if ctx is cancelled mid-call
|
||||
// (V-638). Before that a daemon which stopped answering parked the caller for
|
||||
// as long as the socket stayed open. The watchdog closes the conn rather than
|
||||
// calling drop, because drop wants c.mu and the caller is holding it — the
|
||||
// closed socket fails the read, and roundtrip drops it on the way out.
|
||||
func (c *Client) roundtrip(ctx context.Context, m Method, raw json.RawMessage, resp *Response) error {
|
||||
conn := c.currentConn()
|
||||
if conn == nil {
|
||||
dialed, err := netaddr.Dial(c.addr)
|
||||
if err != nil {
|
||||
return fmt.Errorf("%w: dial %s: %v", errWriteLost, c.addr, err)
|
||||
}
|
||||
c.conn = conn
|
||||
c.setConn(dialed)
|
||||
conn = dialed
|
||||
}
|
||||
if err := writeFrame(c.conn, Request{Method: m, Params: raw}); err != nil {
|
||||
if dl, ok := ctx.Deadline(); ok {
|
||||
_ = conn.SetDeadline(dl)
|
||||
} else {
|
||||
_ = conn.SetDeadline(time.Now().Add(defaultCallTimeout))
|
||||
}
|
||||
defer conn.SetDeadline(time.Time{})
|
||||
|
||||
// The watchdog and the end of the call race by construction: a cancellation
|
||||
// landing just as the reply arrives can close a conn this call is already
|
||||
// done with, and c.conn would still point at the closed socket. So a call
|
||||
// whose context ended does not leave the conn behind for the next one,
|
||||
// whichever of the two got there first.
|
||||
done := make(chan struct{})
|
||||
defer func() {
|
||||
close(done)
|
||||
if ctx.Err() != nil {
|
||||
c.drop()
|
||||
}
|
||||
}()
|
||||
go func() {
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
_ = conn.Close()
|
||||
case <-done:
|
||||
}
|
||||
}()
|
||||
|
||||
if err := writeFrame(conn, Request{Method: m, Params: raw}); err != nil {
|
||||
c.drop()
|
||||
return fmt.Errorf("%w: %v", errWriteLost, err)
|
||||
}
|
||||
if err := readFrame(c.conn, resp); err != nil {
|
||||
if err := readFrame(conn, resp); err != nil {
|
||||
c.drop()
|
||||
return fmt.Errorf("%w: %v", errReadLost, err)
|
||||
}
|
||||
@@ -220,12 +288,26 @@ func (c *Client) roundtrip(m Method, raw json.RawMessage, resp *Response) error
|
||||
|
||||
// drop closes and forgets the current conn so the next call re-dials.
|
||||
func (c *Client) drop() {
|
||||
c.connMu.Lock()
|
||||
defer c.connMu.Unlock()
|
||||
if c.conn != nil {
|
||||
_ = c.conn.Close()
|
||||
c.conn = nil
|
||||
}
|
||||
}
|
||||
|
||||
func (c *Client) currentConn() net.Conn {
|
||||
c.connMu.Lock()
|
||||
defer c.connMu.Unlock()
|
||||
return c.conn
|
||||
}
|
||||
|
||||
func (c *Client) setConn(conn net.Conn) {
|
||||
c.connMu.Lock()
|
||||
defer c.connMu.Unlock()
|
||||
c.conn = conn
|
||||
}
|
||||
|
||||
// hydrate rehydrates a wire RpcError into the matching package sentinel. The
|
||||
// code↔sentinel table is the only place the wire "knows" about errors; keep it
|
||||
// in sync with codeOf in wire.go.
|
||||
|
||||
+36
-8
@@ -17,9 +17,11 @@ import (
|
||||
// Server — the core side of the boundary. Listens on a unix domain socket,
|
||||
// accepts module connections, frames requests to a CoreAPI and responses back.
|
||||
// One Server per daemon process; concurrent connections are handled in their
|
||||
// own goroutine but share the single CoreAPI (and therefore the single store
|
||||
// writer — store is single-connection, SetMaxOpenConns(1), so serialization is
|
||||
// already guaranteed at the db; the Server adds no locking of its own).
|
||||
// own goroutine but share the single CoreAPI, and so the single store writer.
|
||||
// The store opens at SetMaxOpenConns(1), so serialisation is already guaranteed
|
||||
// at the database and the Server adds no locking of its own. That cap is an
|
||||
// invariant this comment depends on, measured and kept on 07-08-2026 (V-642,
|
||||
// docs/evals/2026-08-07-store-connection-cap.md).
|
||||
type Server struct {
|
||||
api atomic.Value // stores CoreAPI
|
||||
path string
|
||||
@@ -30,6 +32,14 @@ type Server struct {
|
||||
done chan struct{}
|
||||
accept sync.Mutex // guards wg.Add vs Close's wg.Wait sequence
|
||||
|
||||
// ctx — server-scoped, cancelled by Close, and the parent of every request
|
||||
// context. serveConn dispatched under context.Background() until V-638, so
|
||||
// a dispatch in flight during shutdown could not be told to stop and the
|
||||
// closeGrace below could only abandon it. Cancelling gives a handler that
|
||||
// respects its context the chance to return instead.
|
||||
ctx context.Context
|
||||
cancel context.CancelFunc
|
||||
|
||||
// conns — every accepted connection still being served. Close needs these
|
||||
// because closing the listener does nothing to a connection already
|
||||
// accepted: serveConn is parked in readFrame waiting for a peer that may
|
||||
@@ -208,11 +218,14 @@ func Listen(path string, api CoreAPI) (*Server, error) {
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
ctx, cancel := context.WithCancel(context.Background())
|
||||
s := &Server{
|
||||
path: path,
|
||||
addr: addr,
|
||||
ln: ln,
|
||||
done: make(chan struct{}),
|
||||
path: path,
|
||||
addr: addr,
|
||||
ln: ln,
|
||||
done: make(chan struct{}),
|
||||
ctx: ctx,
|
||||
cancel: cancel,
|
||||
}
|
||||
s.api.Store(api)
|
||||
return s, nil
|
||||
@@ -250,7 +263,10 @@ func (s *Server) Serve() error {
|
||||
|
||||
func (s *Server) serveConn(c net.Conn) {
|
||||
caller, callerOK := peerCaller(c)
|
||||
ctx := context.Background()
|
||||
// Derived from the server's, so Close cancels a dispatch in flight, and
|
||||
// cancelled when this conn ends so nothing a handler spawned outlives it.
|
||||
ctx, cancel := context.WithCancel(s.serverContext())
|
||||
defer cancel()
|
||||
if callerOK {
|
||||
ctx = WithCaller(ctx, caller)
|
||||
}
|
||||
@@ -274,6 +290,15 @@ func (s *Server) serveConn(c net.Conn) {
|
||||
}
|
||||
}
|
||||
|
||||
// serverContext is s.ctx, or Background for a Server built as a zero value
|
||||
// rather than by Listen (the wiring tests do that).
|
||||
func (s *Server) serverContext() context.Context {
|
||||
if s.ctx == nil {
|
||||
return context.Background()
|
||||
}
|
||||
return s.ctx
|
||||
}
|
||||
|
||||
func (s *Server) safeDispatch(ctx context.Context, req Request) (result json.RawMessage, err error) {
|
||||
defer func() {
|
||||
if r := recover(); r != nil {
|
||||
@@ -720,6 +745,9 @@ func (s *Server) Close() error {
|
||||
default:
|
||||
close(s.done)
|
||||
}
|
||||
if s.cancel != nil {
|
||||
s.cancel()
|
||||
}
|
||||
err := s.ln.Close()
|
||||
// Closing the listener stops new connections; it does nothing to the ones
|
||||
// already accepted. Close those too, or every serveConn parked in readFrame
|
||||
|
||||
@@ -0,0 +1,154 @@
|
||||
package store
|
||||
|
||||
import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"fmt"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"sync"
|
||||
"sync/atomic"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// conncap_test.go measures whether a read queues behind a write at
|
||||
// SetMaxOpenConns(1), which is what openAt sets (V-642). It is a measurement
|
||||
// harness, not an assertion: the numbers it prints are the evidence, and the
|
||||
// decision to move the cap or leave it belongs in docs/evals.
|
||||
//
|
||||
// Run it with -v, and note that it is skipped under -short because it spends
|
||||
// seconds on purpose.
|
||||
|
||||
// openCapped opens a plaintext store at the given connection cap. In-package,
|
||||
// so it can reach the handle openAt caps at 1.
|
||||
func openCapped(t *testing.T, cap int) *Store {
|
||||
t.Helper()
|
||||
path := filepath.Join(t.TempDir(), "cap.db")
|
||||
db, err := openAt(context.Background(), path)
|
||||
if err != nil {
|
||||
t.Fatalf("openAt: %v", err)
|
||||
}
|
||||
db.SetMaxOpenConns(cap)
|
||||
s := &Store{db: db}
|
||||
t.Cleanup(func() { _ = s.Close() })
|
||||
return s
|
||||
}
|
||||
|
||||
func percentile(d []time.Duration, p float64) time.Duration {
|
||||
if len(d) == 0 {
|
||||
return 0
|
||||
}
|
||||
i := int(float64(len(d)-1) * p)
|
||||
return d[i]
|
||||
}
|
||||
|
||||
// seedFacts writes n facts so a read has rows to decode.
|
||||
func seedFacts(t *testing.T, s *Store, n int) {
|
||||
t.Helper()
|
||||
ctx := context.Background()
|
||||
now := time.Now().UTC()
|
||||
for i := 0; i < n; i++ {
|
||||
key := fmt.Sprintf("seed_%d", i)
|
||||
if _, err := s.SetValue(ctx, KindSelf, key, "tap:test",
|
||||
map[string]int{"ml": i}, now.Add(time.Duration(i)*time.Millisecond)); err != nil {
|
||||
t.Fatalf("seed %d: %v", i, err)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// measureReadsUnderWrites reports read latency percentiles while a writer
|
||||
// writes at a fixed pace. The pace matters: an unpaced writer completes a
|
||||
// different number of writes at each cap, because at a higher cap it competes
|
||||
// with the readers for the write lock instead of taking turns on one
|
||||
// connection. Two runs that did different work cannot be compared.
|
||||
// It runs for a fixed wall-clock window rather than a fixed read count, so the
|
||||
// paced writer does the same work at every cap. Tying the window to a read
|
||||
// count made the faster configuration receive fewer writes.
|
||||
func measureReadsUnderWrites(t *testing.T, s *Store, window, pace time.Duration) []time.Duration {
|
||||
t.Helper()
|
||||
ctx := context.Background()
|
||||
|
||||
var stop atomic.Bool
|
||||
var writes atomic.Int64
|
||||
var wg sync.WaitGroup
|
||||
wg.Add(1)
|
||||
go func() {
|
||||
defer wg.Done()
|
||||
now := time.Now().UTC()
|
||||
for i := 0; !stop.Load(); i++ {
|
||||
key := fmt.Sprintf("hot_%d", i%16)
|
||||
if _, err := s.SetValue(ctx, KindSelf, key, "tap:test",
|
||||
map[string]int{"n": i}, now.Add(time.Duration(i)*time.Millisecond)); err != nil {
|
||||
t.Errorf("write: %v", err)
|
||||
return
|
||||
}
|
||||
writes.Add(1)
|
||||
time.Sleep(pace)
|
||||
}
|
||||
}()
|
||||
|
||||
var lat []time.Duration
|
||||
deadline := time.Now().Add(window)
|
||||
for time.Now().Before(deadline) {
|
||||
start := time.Now()
|
||||
if _, err := s.RecentFacts(ctx, 50); err != nil {
|
||||
t.Fatalf("RecentFacts: %v", err)
|
||||
}
|
||||
lat = append(lat, time.Since(start))
|
||||
}
|
||||
stop.Store(true)
|
||||
wg.Wait()
|
||||
t.Logf("in %v: %d reads, %d writes", window, len(lat), writes.Load())
|
||||
|
||||
sort.Slice(lat, func(i, j int) bool { return lat[i] < lat[j] })
|
||||
return lat
|
||||
}
|
||||
|
||||
// TestConnCap_ReadLatencyUnderWrites is the V-642 measurement: read latency at
|
||||
// cap 1 against cap 4, same workload, same schema, same driver.
|
||||
func TestConnCap_ReadLatencyUnderWrites(t *testing.T) {
|
||||
if testing.Short() {
|
||||
t.Skip("measurement harness; runs for seconds")
|
||||
}
|
||||
for _, cap := range []int{1, 4} {
|
||||
t.Run(fmt.Sprintf("cap=%d", cap), func(t *testing.T) {
|
||||
s := openCapped(t, cap)
|
||||
seedFacts(t, s, 500)
|
||||
lat := measureReadsUnderWrites(t, s, 2*time.Second, 2*time.Millisecond)
|
||||
t.Logf("cap=%d reads=%d p50=%v p95=%v max=%v",
|
||||
cap, len(lat), percentile(lat, 0.50), percentile(lat, 0.95), lat[len(lat)-1])
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestConnCap_ReadBlocksBehindOpenSnapshot is the sharper claim: at cap 1 an
|
||||
// open read-only transaction holds the only connection, so an unrelated read
|
||||
// cannot proceed until it commits. This is why the store exposes no way to
|
||||
// begin one — `Store.DB` used to, and was deleted in V-642 with no caller. The
|
||||
// test stays as the reason, so re-adding that seam fails a measurement rather
|
||||
// than shipping a stall.
|
||||
func TestConnCap_ReadBlocksBehindOpenSnapshot(t *testing.T) {
|
||||
if testing.Short() {
|
||||
t.Skip("measurement harness; waits on a timeout")
|
||||
}
|
||||
for _, cap := range []int{1, 4} {
|
||||
t.Run(fmt.Sprintf("cap=%d", cap), func(t *testing.T) {
|
||||
s := openCapped(t, cap)
|
||||
seedFacts(t, s, 50)
|
||||
|
||||
tx, err := s.db.BeginTx(context.Background(), &sql.TxOptions{ReadOnly: true})
|
||||
if err != nil {
|
||||
t.Fatalf("BeginTx: %v", err)
|
||||
}
|
||||
defer func() { _ = tx.Rollback() }()
|
||||
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Second)
|
||||
defer cancel()
|
||||
start := time.Now()
|
||||
_, err = s.RecentFacts(ctx, 10)
|
||||
t.Logf("cap=%d read alongside an open snapshot: waited %v, err=%v",
|
||||
cap, time.Since(start).Round(time.Millisecond), err)
|
||||
})
|
||||
}
|
||||
}
|
||||
+14
-8
@@ -95,7 +95,20 @@ func openAt(ctx context.Context, path string) (*sql.DB, error) {
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("open %s: %w", path, err)
|
||||
}
|
||||
// single writer expected; the daemon is the only process touching the db.
|
||||
// One connection, so every statement is serialised at the database and no
|
||||
// caller above needs a lock of its own. internal/ipc's Server relies on
|
||||
// exactly this, which is why the cap is an invariant rather than a tuning
|
||||
// knob: raising it moves the serialisation guarantee somewhere it is not
|
||||
// written down.
|
||||
//
|
||||
// Measured on 07-08-2026 (V-642, docs/evals/2026-08-07-store-connection-cap.md).
|
||||
// WAL exists to let readers run beside one writer, and the cap gives that
|
||||
// up, but reads do not queue: p50 594µs against 525µs at a cap of four,
|
||||
// while write throughput more than halves. The one thing the cap cannot
|
||||
// survive is a long-lived transaction, which holds the only connection and
|
||||
// stalls every read for its lifetime. So the store begins none, and
|
||||
// TestConnCap_ReadBlocksBehindOpenSnapshot is the standing measurement of
|
||||
// what re-adding one would cost.
|
||||
db.SetMaxOpenConns(1)
|
||||
if _, err := db.ExecContext(ctx, schemaSQL); err != nil {
|
||||
if closeErr := db.Close(); closeErr != nil {
|
||||
@@ -133,13 +146,6 @@ func (s *Store) Close() error {
|
||||
return s.enc.closeAndSeal(s.db)
|
||||
}
|
||||
|
||||
// DB exposes the underlying handle for internal read-only snapshots.
|
||||
// Used by the loop to take a consistent read under a single transaction.
|
||||
// Modules never receive this handle — core mediates.
|
||||
func (s *Store) DB(ctx context.Context) (*sql.Tx, error) {
|
||||
return s.db.BeginTx(ctx, &sql.TxOptions{ReadOnly: true})
|
||||
}
|
||||
|
||||
var (
|
||||
// ErrNoFact — no non-voided row exists for this key.
|
||||
ErrNoFact = errors.New("store: no fact for key")
|
||||
|
||||
@@ -328,7 +328,7 @@ func TestExecEmptyCmdRefuses(t *testing.T) {
|
||||
func TestExecRefusesATargetTheSystemCannotHave(t *testing.T) {
|
||||
api := fakeAPI{tools: map[string]ipc.Tool{
|
||||
"restart": {Name: "restart", Cmd: []string{"systemctl", "restart"}, Status: "enabled"},
|
||||
"drop": {Name: "drop", Cmd: []string{"dropdb"}, Destructive: true, Status: "enabled"},
|
||||
"drop": {Name: "drop", Cmd: []string{"dropdb"}, Destructive: true, Status: "enabled"},
|
||||
}}
|
||||
ran := false
|
||||
e := NewExecutor(api, 0)
|
||||
|
||||
@@ -26,6 +26,8 @@
|
||||
package voice
|
||||
|
||||
import (
|
||||
"context"
|
||||
|
||||
"github.com/kami/maven/internal/phraser"
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
@@ -38,8 +40,10 @@ import (
|
||||
// decision's Intent + Slots + Clarify. The Intent largely names the reply
|
||||
// shape (act/reminder/fact/note/query/clarify); the Slots carry the
|
||||
// specifics that personalise it ("got it: water at 14:00").
|
||||
// The context is the turn's, and it is the only bound an LLM-backed impl has
|
||||
// besides the phraser timeout (V-638). A floor impl ignores it.
|
||||
type Replier interface {
|
||||
Reply(d router.Decision) string
|
||||
Reply(ctx context.Context, d router.Decision) string
|
||||
}
|
||||
|
||||
// StubReplier — the deterministic, no-model floor. Canned per intent;
|
||||
@@ -54,7 +58,8 @@ func NewStubReplier() *StubReplier { return &StubReplier{} }
|
||||
|
||||
// Reply dispatches on Intent + Clarify. Each branch is short; the LLM impl
|
||||
// will replace this with prompted text and the same dispatch shape.
|
||||
func (s *StubReplier) Reply(d router.Decision) string {
|
||||
// It makes no model call, so the context is unused.
|
||||
func (s *StubReplier) Reply(_ context.Context, d router.Decision) string {
|
||||
if d.Clarify {
|
||||
return "не совсем поняла — можешь переформулировать?"
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user