- Shipped
- September 28, 2026 at 4:36 AM UTC
- Author
- Kamo
- Commit
- 8f24b41
The owner asked for a local model that runs substantially better on k3m1's AVX2-only Xeons (2026-09-27, ruling SP01-LLM-granite4). granite4:tiny-h is a hybrid Mamba-2/transformer MoE (7B total, ~1B active): several times qwen3:4b's prompt and answer speed on CPU, native tool calling, and no thinking pass. - postStart pulls granite4:tiny-h; - CPU limit 8 -> 16 (LLAMA_ARG_THREADS follows it), request 2 -> 4; - OLLAMA_CONTEXT_LENGTH 8192 -> 16384 (agent steps run 3-5k tokens); - OLLAMA_FLASH_ATTENTION=1. Memory stays 8Gi.
