SecurityService pool exhaustion caused liveness kills

FixSecurityService
Ya
10 Agosti 2026, 17:14 UTC
Mwandishi
Kamo
Ahadi ya
2fb944b

Hikari had no explicit maximum-pool-size, so it ran on the default of 10 while fronting org/rights lookups for every service and the internal site. Observed total=10 active=10 waiting=33, which stalled the /actuator/health db indicator to 18.8s; the 3s probe timeout then killed the pod (exit 137), 29 readiness failures over 154 minutes. That was surfacing to users as 'Upstream returned 500' on the internal site. Pool raised to 40 (safe now that YSQL Connection Manager multiplexes clients rather than one PostgreSQL backend each), and the probe timeouts relaxed from 3s to 10s.

Mabadiliko yote

Je, unaona nini kuhusu usafiri?

Kila moja ya hizi updates ardhi katika nafasi yako ya kazi moja kwa moja. Kuanza bure na kuangalia kukua wiki baada ya wiki.

Kuwa Huru MileleMtazamo wa bei