Skip to main content

Models

Ultra is not available over WebSocket. Use the REST API for Ultra.

Natural & NaturalHD

Pre-built voices. Browse available voices via the Voices API or the Voice Design.

Audio Format

Default: MP3. NaturalHD supports audio_format query parameter to override:
Accepted values: pcm, wav.

Voice Settings

Qwen3TTS

Voice cloning model. The voice_id is the name of a clone created in the Voice Design. Cloned voice usage may require identity verification. Requires the clone to belong to your organization.

Audio Format

Always raw PCM — 24kHz, signed 16-bit little-endian, mono. Forced by the backend regardless of any output_format value sent.

Voice Settings

KokoroTTS

Lightweight, low-latency model. Suitable for high-throughput applications where quality tradeoffs are acceptable.

Bayan

Arabic voice model. 113 speakers across 13 dialects (Modern Standard Arabic, Egyptian, Emirati, Saudi, Jordanian, Iraqi, Lebanese, Syrian, Palestinian, Kuwaiti, Bahraini, Qatari, Omani) plus English speakers. Native audio is 16kHz.

Voice Settings

Only the native 16000 sample rate is supported.

Sukhan

Urdu voice model. 14 curated voices. No prosody controls or language selection — Urdu only. Native audio is 22050Hz.

Voice Settings

None. voice_speed and other prosody settings are not supported and have no effect. Only the native 22050 sample rate is supported. Supported response_format/audio_format: pcm, mp3 (no wav).