- Shipped
- October 5, 2026 at 3:57 PM UTC
- Author
- Kamo
- Commit
- 2a848b8
KBService sends a help article's Lexical JSON to be translated: one long line with few sentence ends, ~0.46 model tokens a character. A 3,980-character article reached the model as a 1,363-token piece padded beside six short ones; two at once (both TranslateService pods) peaked at 1,960 MiB against the 2 GiB limit, and the engine was OOMKilled 21 times in 6.5 days (last 2026-10-05T11:40:31Z). - No piece reaches the model longer than the 256 tokens it was trained on: a sentence over 600 characters is cut at its spaces, and a piece still over 256 tokens is halved at the space nearest its middle (never inside a KMPH sentinel). - CTranslate2 batches by tokens (MAX_BATCH_TOKENS=512), not 16 examples. - MALLOC_ARENA_MAX=2. - Memory limit 3Gi (twice the worst measured peak, 1,164 MiB for two 20,000-character texts at once); the request stays 1Gi. Same load now: 974 MiB (two articles), 943 MiB (four), no slower; short strings unchanged. MemoryTest pins it on the real model (old engine 2,442 MiB: red). SP98 final review I-1.
