KamoCRM

Run llama-server with one thread per core of its CPU limit

FixKlusterServices
Shipped
September 28, 2026 at 2:35 AM UTC
Author
Kamo
Commit
6186ce6

Ollama starts llama-server without --threads, so it counted k3m1's 24 physical cores and ran 24 compute threads inside the container's 8-core quota. The container spent twice as long throttled as running (16 560 s throttled against 8 405 s used), prompts went at ~25 tokens/s and answers at ~3, and a 3 600-token agent step could not finish inside AIService's 120 s per attempt. LLAMA_ARG_THREADS is read by llama-server (and sets the batch threads too); the Downward API ties it to limits.cpu so the two cannot drift apart.

All changes

Like what you see shipping?

All of it arrives in your workspace on its own. Start on the free plan and read this page again in a month.

Start Free ForeverView Pricing