Models
Natural & NaturalHD
Pre-built voices. Browse available voices via the Voices API or the Voice Design.Audio Format
Default: MP3. NaturalHD supportsaudio_format query parameter to override:
pcm, wav.
Voice Settings
Qwen3TTS
Voice cloning model. Thevoice_id is the name of a clone created in the Voice Design. Cloned voice usage may require identity verification.
Requires the clone to belong to your organization.
Audio Format
Always raw PCM — 24kHz, signed 16-bit little-endian, mono. Forced by the backend regardless of anyoutput_format value sent.
Voice Settings
KokoroTTS
Lightweight, low-latency model. Suitable for high-throughput applications where quality tradeoffs are acceptable.Bayan
Arabic voice model. 113 speakers across 13 dialects (Modern Standard Arabic, Egyptian, Emirati, Saudi, Jordanian, Iraqi, Lebanese, Syrian, Palestinian, Kuwaiti, Bahraini, Qatari, Omani) plus English speakers. Native audio is 16kHz.Voice Settings
Only the native
16000 sample rate is supported.
Sukhan
Urdu voice model. 14 curated voices. No prosody controls or language selection — Urdu only. Native audio is 22050Hz.Voice Settings
None.voice_speed and other prosody settings are not supported and have no effect.
Only the native 22050 sample rate is supported. Supported response_format/audio_format: pcm, mp3 (no wav).