- Shipped
- September 29, 2026 at 2:32 AM UTC
- Author
- Kamo
- Commit
- 16db29f
The adapter hard-coded output_format=mp3_44100_128 on every call, so a caller asking E5 for wav (or kamoai-voice asking for pcm, on every phone call) still got mp3 back — one extra lossy decode plus a resample on the call path, and a silent contract violation for any other caller of /api/ai/v1/audio/speech. synthesize() now maps AudioController's response_format to ElevenLabs' own output_format: pcm/wav both fetch pcm_24000 (AIService's own PCM convention, matching kamoai-voice's hardcoded 24 kHz decode) with wav wrapped in a real RIFF/WAVE header here (ElevenLabs never sends one), mp3 keeps mp3_44100_128, opus maps to opus_48000_128, and an explicit ulaw maps to ulaw_8000. A format ElevenLabs can't speak (aac, flac) is refused up front so the router falls back to a model that can, instead of silently mismapping it to mp3. The declared mimeType now always matches the bytes returned, not whatever Content-Type the vendor happened to send. kamoai-voice already asks for pcm and reads AIService's declared Content-Type rather than trusting its own request, so no change was needed there: it gets real, matching-rate PCM now instead of a decoded-and-resampled mp3. Tests (SpeechVendorFamiliesTest, red then green against a fake ElevenLabs server): each response_format maps to the right output_format, the wav wrapper is a valid 44-byte RIFF/WAVE header around the untouched PCM bytes, an unsupported format is rejected before any HTTP call, and mp3 is unchanged.
