feat(health): gate routing event-driven on connection drop, not just periodic poll (#300)
HealthMonitor's periodic poll took ~18s to notice a dead provider — long after a retry had already re-selected it and died. Health needs to GATE routing, not just log reactively. - InferenceRouter.reportFailure(providerId, reason): new default-no-op method; DefaultInferenceRouter implements it by writing Unavailable straight into its health cache, bypassing healthCheck()/TTL entirely. - SessionOrchestrator.runInference: on a connection-level exception (ConnectException/SocketException/IOException or a message matching "prematurely closed"/"connection refused"/"connection reset"), call reportFailure immediately so the very next route() call — the retry driven by #299 — sees the provider as down right away instead of re-selecting it and dying again. Pairs with #299: routing now (1) reacts to a connection drop instantly and (2) waits/backs off for the capability to recover before declaring terminal. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GeyGFXczJb8RUWGBKmkm6G
This commit is contained in:
+20
@@ -409,6 +409,13 @@ abstract class SessionOrchestrator(
|
||||
} catch (e: CancellationException) {
|
||||
throw e // never swallow
|
||||
} catch (e: Exception) {
|
||||
if (isConnectionLevelFailure(e)) {
|
||||
// Mark it down NOW instead of waiting for the next periodic health poll (~18s lag,
|
||||
// see #300) — the retry's route() call must see this provider as unavailable
|
||||
// immediately so it gates/waits (#299) rather than instantly re-selecting the dead
|
||||
// provider again.
|
||||
inferenceRouter.reportFailure(provider.id, e.message ?: "connection failure")
|
||||
}
|
||||
emit(
|
||||
sessionId,
|
||||
InferenceFailedEvent(
|
||||
@@ -423,6 +430,19 @@ abstract class SessionOrchestrator(
|
||||
}
|
||||
}
|
||||
|
||||
// Connection-level failures (provider crashed/restarting mid-request) should gate routing
|
||||
// immediately; other failures (bad request, model error, HTTP 4xx) should not mark the
|
||||
// provider down since the provider itself is still reachable.
|
||||
private fun isConnectionLevelFailure(e: Exception): Boolean {
|
||||
val message = e.message.orEmpty()
|
||||
return e is java.net.ConnectException ||
|
||||
e is java.net.SocketException ||
|
||||
e is java.io.IOException || // covers ktor/CIO's IOException, which extends java.io.IOException on the JVM
|
||||
message.contains("prematurely closed", ignoreCase = true) ||
|
||||
message.contains("connection refused", ignoreCase = true) ||
|
||||
message.contains("connection reset", ignoreCase = true)
|
||||
}
|
||||
|
||||
// --- token estimation ---
|
||||
|
||||
internal open suspend fun estimateTokens(content: String): Int {
|
||||
|
||||
Reference in New Issue
Block a user