# QA Plan: EVENT_STORE health probe — 7a1362c Observability spec §4. Drafted from the diff (commit `7a1362c`) per the BACKLOG QA rule. **Status:** DRAFT **Run date / operator:** **BACKLOG item:** §A "§4 EVENT_STORE health probe — append + fsync latency heartbeat". --- ## Preconditions - [ ] **server build/branch:** master @ `7a1362c` (or later) — _rebuild only with the server stopped_ - [ ] **config:** `[health].enabled = true`; new field `[health].event_store_latency_warn_ms` (default 500). - [ ] **interval:** `[health].interval_ms` (default 30_000) — how long a status edge takes to surface. - [ ] **how to read state:** `correx health` / `GET /health/checks`; raw events via `correx events __system__`. - [ ] **(for the slow-path check)** a way to make a real SQLite read slow — e.g. point the event store at a path on a slow/contended/network filesystem, or temporarily lower `event_store_latency_warn_ms` to a value below normal read latency to force a DEGRADED. ## Acceptance gate (one sentence) > The probe is correct **iff** a responsive store reads HEALTHY (latency recorded), an unresponsive or > slow store emits exactly one `HealthDegradedEvent(subject=EVENT_STORE)` (visible in `correx health` and > the `__system__` log), recovery emits exactly one `HealthRestoredEvent`, the probe NEVER writes to the > log itself, and replay re-derives all of it from recorded events. ## Checks | # | Action | Expected observable evidence | Result | |---|--------|------------------------------|--------| | 1 | Server up, `health.enabled=true`, wait one interval | `correx health` shows EVENT_STORE HEALTHY; detail "event store responsive: lastGlobalSequence() in Nms"; `observedValue` = a small latency | | | 2 | **No log pollution:** note `correx events __system__` event count, let several health intervals pass | The `__system__` event count does NOT grow per tick — only on actual status *edges*. The probe must not append a heartbeat event every tick (Invariant #1) | | | 3 | Force slow reads (slow FS, or set `event_store_latency_warn_ms` below normal latency), wait one interval | Exactly one `HealthDegradedEvent` subject=EVENT_STORE, metric `read_latency_ms`, detail "lastGlobalSequence() took Nms (warn ≥ …ms)" | | | 4 | Restore normal latency (or raise the threshold back), wait one interval | Exactly one `HealthRestoredEvent(EVENT_STORE)`; back to HEALTHY. No repeated Degraded/Restored on later ticks (hysteresis = edge-only) | | | 5 | Simulate store failure (e.g. make the DB unreadable mid-run, if feasible) | `HealthDegradedEvent` with detail "lastGlobalSequence() threw: …"; the monitor tick survives (no crash, next tick still runs) | | | 6 | Restart the **correx** server while the store is HEALTHY | No spurious Degraded/Restored re-emission on boot (HealthMonitor seeds status from the projection) | | | 7 | `correx replay __system__` | Replay re-derives the same health timeline from recorded `HealthObservation` events; double-read digest stable; no live store-latency re-probe (invariants #8/#9) | | Evidence sources: `correx events __system__` (recorded truth + the no-growth check), `correx health` / `GET /health/checks`, server logs (probe tick + emitted edge events), `correx replay __system__`. ## Out of scope (NOT covered this pass) - LLAMA_SERVER probe (separate QA plan, `docs/qa/QA-llama-health-probe.md`). - Live health TUI pane (separate BACKLOG item, §A). - True fsync-latency instrumentation of the append path — this probe is a read-responsiveness heartbeat by design (writing probe events to time appends would pollute the log); deeper append/fsync metrics, if ever wanted, belong in MetricsProjection over real append events, not here. ## Disposition - **PASS** → MOVE the §A "EVENT_STORE health probe" entry into `RETRO.md` (cite `7a1362c`, run date, evidence). - **FAIL** → file each failure as a numbered finding back into `BACKLOG.md` (action + repro + wrong/missing signal), fix, re-run only the failed checks. Set Status: FAILED until green.