Stream text to speech over WebSocket
AsyncAPI specification for the Telnyx Text-to-Speech WebSocket endpoint. Real-time speech synthesis by streaming text and receiving audio chunks.
Supported Providers
-
telnyx- Telnyx native voices (KokoroTTS, Qwen3TTS) -
aws- Amazon Polly -
azure- Microsoft Azure TTS -
elevenlabs- ElevenLabs voices -
minimax- MiniMax voices -
resemble- Resemble AI voices -
xai- xAI voices (Eve, Ara, Rex, Sal, Leo) -
inworld- Inworld AI voices -
fishaudio- Fish Audio voices (s2.1-pro, s2-pro, s1 models) -
soniox- Soniox voices (tts-rt-v2 model, every voice speaks 60+ languages)
Connection Flow
- Open WebSocket connection to
wss://api.telnyx.com/v2/text-to-speech/speechwith query parameters. - Send an initial handshake message
{"text": " "}(single space) with optionalvoice_settings. - Send text messages as
{"text": "Hello world"}. - Receive audio chunks as JSON frames with base64-encoded audio.
- A final frame with
isFinal: trueindicates the end of audio for the current text.
Authentication
Requires authentication via a Bearer token (Telnyx API v2 key).
{
"text": " ",
"voice_settings": {
"voice_speed": 1.2
}
}{
"audio": "QmFzZTY0RW5jb2RlZEF1ZGlv",
"text": "Hello world",
"isFinal": false,
"cached": false,
"timeToFirstAudioFrameMs": 245
}{
"audio": null,
"text": "",
"isFinal": true
}{
"error": "Invalid voice_id specified"
}Telnyx API v2 Bearer token authentication.
Query parameters passed when opening the WebSocket connection.
Client-to-server frame containing text to synthesize. The initial handshake message should be {"text": " "} (single space) with optional voice_settings. Subsequent messages contain actual text. To interrupt synthesis mid-stream, send {"force": true}.
Server-to-client frame containing a base64-encoded audio chunk. For providers that stream audio in real-time (Telnyx, Minimax, Resemble, Inworld, Fish Audio), text will be null because audio is streamed before full text alignment is available, and cached will be false. For other providers, text contains the corresponding text segment.
Server-to-client frame indicating synthesis is complete for the current text. The connection remains open for additional text messages.
Server-to-client frame indicating an error during synthesis. The connection will be closed shortly after sending this frame.
Was this page helpful?
{
"text": " ",
"voice_settings": {
"voice_speed": 1.2
}
}{
"audio": "QmFzZTY0RW5jb2RlZEF1ZGlv",
"text": "Hello world",
"isFinal": false,
"cached": false,
"timeToFirstAudioFrameMs": 245
}{
"audio": null,
"text": "",
"isFinal": true
}{
"error": "Invalid voice_id specified"
}