Stream speech to text over WebSocket
AsyncAPI specification for the Telnyx Speech-to-Text WebSocket endpoint. Real-time speech transcription by streaming audio and receiving transcript frames.
Supported Engines
Azure- Microsoft Azure Speech ServicesDeepgram- Deepgram Nova and Flux modelsGoogle- Google Cloud Speech-to-TextTelnyx- Telnyx native transcription (OpenAI Whisper models)xAI- xAI Grok STTAssemblyAI- AssemblyAI Universal-Streaming (backed by Universal-3.5 Pro Realtime)Speechmatics- Speechmatics real-time transcriptionSoniox- Soniox real-time transcriptionParakeet- Self-hosted NVIDIA Parakeet multilingual transcriptionReson8- Reson8 turn-based multilingual transcriptionCohere- Self-hosted Cohere Arabic/English transcription (batch, no auto-detect)
Connection Flow
- Open WebSocket connection to
wss://api.telnyx.com/v2/speech-to-text/transcriptionwith query parameters. - Send binary audio frames (mp3, wav, linear16, or linear32 format, per input_format).
- Receive JSON transcript frames with
transcript,is_final, andconfidencefields. - Close connection when done.
Authentication
Requires authentication via a Bearer token (Telnyx API v2 key).
{}{
"type": "transcript",
"transcript": "Hello, this is",
"is_final": false,
"confidence": 0.85
}{
"type": "error",
"error": "Invalid transcription_engine specified"
}Telnyx API v2 Bearer token authentication.
Query parameters passed when opening the WebSocket connection.
Client-to-server binary frame containing audio data to transcribe. Audio should be in mp3, wav, linear16, or linear32 format as specified in the input_format query parameter.
Server-to-client frame containing a transcription result. When interim_results is enabled, you may receive multiple interim results (is_final=false) before the final result (is_final=true) for each utterance.
Reson8 transcribes per turn of speech: each interim frame carries the full transcript of the turn so far and supersedes the previous frame, the final frame arrives when the turn ends, and confidence is omitted from its frames.
Server-to-client frame indicating an error during transcription. The connection may be closed shortly after sending this frame.
Was this page helpful?
{}{
"type": "transcript",
"transcript": "Hello, this is",
"is_final": false,
"confidence": 0.85
}{
"type": "error",
"error": "Invalid transcription_engine specified"
}