9e2a9690bfce87f708d0710f6b5cab9873de7d7d
mavweb (and every ipc.Client) held one net.Conn from Dial and reused it for the life of the process. When mavend restarted, the socket got a new inode, the cached conn went dead, and every call failed forever with "broken pipe" — the dash and page-heartbeat 502'd until mavweb was manually restarted. Fix in the one place all 25 methods route through (call): on a lost connection — write failure OR read EOF, since a peer restart can surface on either phase depending on socket-buffer timing — drop the conn, re-dial the remembered path, and retry once. Safe for the case that happens (core restarted, request never processed); the rare committed-then-died window can double-apply a write, but the store is append-only so a duplicate is a superseding row, not corruption. ponytail: retry-once, not request-ids — revisit if double-apply ever bites. Test reproduces the exact incident: server restart on the same socket path, and asserts the next call transparently reconnects. Note (not fixed here): Server.Close waits on its handler goroutines, which park reading live client conns — so a graceful core shutdown with a client attached blocks until the client disconnects. Minor; surfaces as a slow SIGTERM. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Description
No description provided
Languages
Go
97.1%
HTML
0.9%
Shell
0.6%
CSS
0.5%
Makefile
0.3%
Other
0.6%