KamoCRM

Cap llama-server's prompt cache below the memory limit, so it stops being OOMKilled

FixKlusterServices
Shipped
September 28, 2026 at 2:09 PM UTC
Author
Kamo
Commit
f4d979d

llama-server's host-RAM prompt cache defaults to 8192 MiB (--cache-ram), the container's whole limit, so it grew until the kernel killed ollama: on 2026-09-28 it held 14 prompts in 3 548 MiB on a ~4.4 GiB model and the 15th save was OOMKilled (8 restarts in 9.5 h, a 502 to every call in flight). LLAMA_ARG_CACHE_RAM=1024 keeps about four prompts (agent steps share ~340 tokens of prefix, so nothing is lost); LLAMA_ARG_CTX_CHECKPOINTS=8 bounds the per-slot checkpoints (2-3 per prompt at ~55 MiB) that the default 32 left at 1.7 GiB. Ollama starts llama-server with neither flag, and the env reaches it.

All changes

Like what you see shipping?

All of it arrives in your workspace on its own. Start on the free plan and read this page again in a month.

Start Free ForeverView Pricing