SecurityService pool exhaustion caused liveness kills

FixSecurityService
Shipped
August 10, 2026 at 5:14 PM UTC
Author
Kamo
Commit
2fb944b

Hikari had no explicit maximum-pool-size, so it ran on the default of 10 while fronting org/rights lookups for every service and the internal site. Observed total=10 active=10 waiting=33, which stalled the /actuator/health db indicator to 18.8s; the 3s probe timeout then killed the pod (exit 137), 29 readiness failures over 154 minutes. That was surfacing to users as 'Upstream returned 500' on the internal site. Pool raised to 40 (safe now that YSQL Connection Manager multiplexes clients rather than one PostgreSQL backend each), and the probe timeouts relaxed from 3s to 10s.

All changes

Like what you see shipping?

Every one of these updates lands in your workspace automatically. Start free and watch it grow week after week.

Start Free ForeverView Pricing