Skip to main content
Voice format: Telnyx.Ultra.<voice> Sub-100ms latency. 36 languages.
REST only — Ultra is not available over public WebSocket.

Voice Samples


SSML Emotions

Ultra supports inline SSML emotion tags. Place the tag before the text:
Primary emotions: angry, excited, content, sad, scared. Additional: happy, enthusiastic, curious, calm, grateful, affectionate, sarcastic, surprised, confident, hesitant, apologetic, determined, frustrated, disappointed. Omitting the tag = neutral delivery. Use sparingly — Ultra interprets emotional subtext from the text itself.

Nonverbal Cues

Insert [laughter] inline for natural laughing:

Language Support

Set language_boost to improve pronunciation for the target language: Arabic, Bengali, Bulgarian, Chinese, Czech, Danish, Dutch, English, Finnish, French, German, Gujarati, Hebrew, Hindi, Indonesian, Italian, Japanese, Korean, Malay, Marathi, Māori, Norwegian, Polish, Portuguese, Punjabi, Romanian, Russian, Slovak, Spanish, Swedish, Tamil, Telugu, Thai, Turkish, Ukrainian, Vietnamese.

REST API

Fields

Response

Default (binary_output): chunked audio bytes with Content-Type: audio/mpeg. With output_type: "base64_output": JSON with base64-encoded audio. With output_type: "audio_id": JSON with an audio_url for deferred retrieval.

See also

xAI Grok is the second TTS provider supporting Expressive Mode. For Grok voice options, see Grok Voices. Note: Grok voices have higher latency than Ultra.