KamoCRM

An Ollama model's context is the window the server serves, not the trained one

FixAIService
Name
lúc 10:06 28 tháng 9, 2026 UTC
Tác giả
Kamo
Cam kết
68050c5

OllamaDiscovery took model_info.*.context_length from /api/show: the window the model was TRAINED for. The registry listed granite4:tiny-h at 1 048 576 (and qwen3:4b at 262 144) while kamo-llm serves 16 384 (OLLAMA_CONTEXT_LENGTH, klusterservices), so the router never excluded a prompt past 16k and trimming aimed at a million tokens; Ollama cuts such a prompt. The window now comes from GET /api/ps, the loaded model's context_length. A model the server has not loaded is reported without a window, and a sync keeps the stored one. At 16 384 the automated LONG_CONTEXT claim goes away, so the row's 1 048 576 no longer mirrors back into max_context_length. SP99-T12.

Mọi thay đổi

Như những gì anh thấy vận chuyển?

Tất cả những thứ đó đều đến trong không gian làm việc của anh. Bắt đầu với kế hoạch miễn phí và đọc lại trang này trong một tháng.

Bắt đầu tự do mãi mãiXem truy cập