KamoCRM

A Hugging Face text stream is metered as the router counted it

FixAIService
Shipped
September 28, 2026 at 10:06 AM UTC
Author
Kamo
Commit
c9be36f

HuggingFaceAdapter streamed through the interface's text-only default: every chunk was read for its text and nothing else, so no usage ever reached the executor or member chat and each streamed answer was metered as a chars/4 estimate. The stream now asks for the usage (stream_options.include_usage, in the router's chat-completion spec) and reads each chunk for its text, finish_reason and usage. A router that refuses the ask is asked again without it by the executor's ParameterFallback; a stream without a usage chunk stays an estimate, flagged as one. A cut-off answer now ends as max_tokens instead of reading as finished. The stream no longer prints the request body (the prompt) to stdout. SP99-T12.

All changes

Like what you see shipping?

All of it arrives in your workspace on its own. Start on the free plan and read this page again in a month.

Start Free ForeverView Pricing