KamoCRM

An in-cluster model gets the time to read its prompt, inside the walk's budget

FixAIService
शिप
28 सितंबर 2026 को 10:06 am बजे UTC
लेखक
Kamo
Commit
a1d2743

kamo-llm reads a prompt at about 110 tokens a second on k3m1 before it writes a word, so an E2 prompt past ~6k tokens spent E2's whole 60 s per-attempt bound (and a stream its 60 s first-event bound) before the answer began, and the walk moved on. An attempt on an IN_CLUSTER model now also gets the prompt's reading time at **************** (50: half the measured rate, for the chars/4 estimate on JSON tool schemas and a short queue). No attempt runs past the walk's own budget, maxAttempts x the per-attempt bound, so the callers' rule (master section 11 'facade timeouts', SP04-F2: E10 3 x 120 s + 30 s) still holds; the ParameterFallback retry, which used to get a fresh full bound, is held to it too. SP99-T12.

सभी बदलाव

जैसा कि आप शिपिंग देखते हैं?

यह सब अपने कार्यक्षेत्र में आता है। मुफ्त योजना शुरू करें और इस पृष्ठ को एक महीने में फिर से पढ़ें।.

Foreverमूल्य निर्धारण देखें