KamoCRM

An Ollama model's context is the window the server serves, not the trained one

FixAIService
Ya
28 Septemba 2026, 10:06 UTC
Mwandishi
Kamo
Ahadi ya
68050c5

OllamaDiscovery took model_info.*.context_length from /api/show: the window the model was TRAINED for. The registry listed granite4:tiny-h at 1 048 576 (and qwen3:4b at 262 144) while kamo-llm serves 16 384 (OLLAMA_CONTEXT_LENGTH, klusterservices), so the router never excluded a prompt past 16k and trimming aimed at a million tokens; Ollama cuts such a prompt. The window now comes from GET /api/ps, the loaded model's context_length. A model the server has not loaded is reported without a window, and a sync keeps the stored one. At 16 384 the automated LONG_CONTEXT claim goes away, so the row's 1 048 576 no longer mirrors back into max_context_length. SP99-T12.

Mabadiliko yote

Je, unaona nini kuhusu usafiri?

Kila kitu kinaingia kwenye tovuti yako mwenyewe. Anza kwenye mpango wa bure na usome ukurasa huu tena katika mwezi mmoja.

Kuwa Huru MileleMtazamo wa bei