- Shipped
- August 27, 2026 at 5:09 AM UTC
- Author
- Kamo
- Commit
- 795a0d4
Root cause of the crash-loop was not a leak. The 485Mi steady state that kept tripping the old 512Mi limit was two normal things stacked: ~350Mi Go heap at the top of a stock GOGC=100 sawtooth. Measured over an hour: peaks at 283Mi, hands back 40-80Mi per GC cycle, floor around 178Mi. Connection count does not track it (108 conns at 186Mi, 58 conns at 283Mi), so it is not connection retention. ~133Mi page cache from the 205 MiB traefik binary's text demand-paging in, which takes about a day to reach that level. Those two do not fit in 512Mi. The limit was only survivable while the binary was still cold, which is why it held for months and then started failing every ~25 minutes. GOMEMLIMIT=750MiB bounds the Go runtime alone, leaving the binary's cache room under the 1Gi limit (1024 - 205 - 70 slack). It is a soft limit, so a future live-heap increase makes Traefik collect harder instead of being SIGKILLed — which matters because killing the cluster's only ingress pod blanks TLS platform-wide for ~11s. The 256Mi request also badly under-declared a pod whose steady state is ~270Mi, leaving the only ingress pod a prime eviction candidate under node memory pressure. Raised to 512Mi.