- Shipped
- September 28, 2026 at 2:35 AM UTC
- Author
- Kamo
- Commit
- 6186ce6
Ollama starts llama-server without --threads, so it counted k3m1's 24 physical cores and ran 24 compute threads inside the container's 8-core quota. The container spent twice as long throttled as running (16 560 s throttled against 8 405 s used), prompts went at ~25 tokens/s and answers at ~3, and a 3 600-token agent step could not finish inside AIService's 120 s per attempt. LLAMA_ARG_THREADS is read by llama-server (and sets the batch threads too); the Downward API ties it to limits.cpu so the two cannot drift apart.
