KamoCRM

An in-cluster model gets the time to read its prompt, inside the walk's budget

FixAIService
Verschifft
28. September 2026 um 10:06 UTC
Autor
Kamo
Ausschuss
a1d2743

kamo-llm reads a prompt at about 110 tokens a second on k3m1 before it writes a word, so an E2 prompt past ~6k tokens spent E2's whole 60 s per-attempt bound (and a stream its 60 s first-event bound) before the answer began, and the walk moved on. An attempt on an IN_CLUSTER model now also gets the prompt's reading time at **************** (50: half the measured rate, for the chars/4 estimate on JSON tool schemas and a short queue). No attempt runs past the walk's own budget, maxAttempts x the per-attempt bound, so the callers' rule (master section 11 'facade timeouts', SP04-F2: E10 3 x 120 s + 30 s) still holds; the ParameterFallback retry, which used to get a fresh full bound, is held to it too. SP99-T12.

Alle Änderungen

Wie, was Sie sehen Versand?

Alles kommt in Ihrem Arbeitsbereich für sich. Starten Sie mit dem kostenlosen Plan und lesen Sie diese Seite in einem Monat wieder.

Free Forever startenPreisgestaltung anzeigen