Telnyx Voice — Documentation Index
Voice API, speech-to-text, text-to-speech, and voice design. One section of the Telnyx developer docs (https://developers.telnyx.com). Root index: https://developers.telnyx.com/llms.txt · Full content for this section: https://developers.telnyx.com/development/llms/voice-llms-full-txt.md
Subsections
Focused per-subsection files (index + full content):Overview
- Overview: Overview of Telnyx Voice — standalone components for building voice AI applications, including STT, TTS, programmable voice, WebRTC, and SIP trunking.
STT
- Overview: Real-time and batch audio transcription via WebSocket, REST, or in-call.
- Quickstart: Stream audio to Telnyx Speech-to-Text and see live transcripts in under 5 minutes with this end-to-end Python and JavaScript quickstart.
- Models: Compare Telnyx Speech-to-Text engines and models — Deepgram, Whisper, Google, Azure, xAI, AssemblyAI, Speechmatics, Soniox, Parakeet — by accuracy, latency, language coverage, and price.
- Migration: Migrate from Deepgram, AssemblyAI, Google, or other speech-to-text providers to Telnyx STT in minutes by changing only two to three lines of code.
- Lifecycle: How the Telnyx Speech-to-Text WebSocket endpoint works for real-time streaming, including connection, audio frames, transcript messages, and shutdown.
- Overview: Reference for query parameters on the Telnyx Speech-to-Text WebSocket endpoint, including engine, model, language, and redaction options.
- Audio Formats: Supported audio input formats, sample rates, and engine compatibility for the Telnyx Speech-to-Text WebSocket streaming endpoint and binary frames.
- Engines & Models: Available speech-to-text engines and models on Telnyx — Deepgram, Telnyx native, Google, Azure, xAI, AssemblyAI, Speechmatics, Soniox, Parakeet — selectable via query parameters.
- End-of-Turn Detection: Flux end-of-turn detection parameters for voice agent turn-taking.
- Language: Configure the language for Telnyx Speech-to-Text WebSocket streaming using BCP-47 codes, with per-engine differences in supported languages and behavior.
- Interim Results: Enable interim results to receive partial transcripts as audio streams to the Telnyx Speech-to-Text WebSocket endpoint (Deepgram engine only).
- Endpointing: Configure silence-based endpointing for utterance boundary detection on the Telnyx Speech-to-Text WebSocket endpoint (Deepgram, xAI, Google, Speechmatics, and Soniox).
- Keyword Boosting: Boost recognition of specific terms (keyterm and keywords parameters).
- Redaction: Automatically redact PII such as numbers, names, and SSNs from Telnyx streaming transcription results using Deepgram redaction parameters.
- Messages: Reference for binary audio frames and JSON message types exchanged over the Telnyx Speech-to-Text WebSocket connection in both directions.
- Errors: Error codes and troubleshooting tips for the Telnyx Speech-to-Text WebSocket endpoint, including invalid parameters, engine mismatches, and disconnects.
- Examples: Complete Python and JavaScript code examples for streaming audio to the Telnyx Speech-to-Text WebSocket API and printing live transcription results.
- Production Patterns: Reconnect, backoff, partial handling, buffering, and monitoring patterns for WebSocket STT.
- Pricing: Pricing for Telnyx Speech-to-Text WebSocket streaming. Billed per minute of audio streamed in real time, with model-tier rates and volume discounts.
- Overview: Transcribe audio files synchronously with the Telnyx Speech-to-Text REST API by uploading a file or providing a public URL and receiving text.
- Overview: Reference of all parameters for the Telnyx Speech-to-Text REST API — audio source, language, model, diarization, punctuation, and response options.
- Models: Available speech-to-text models on the Telnyx STT REST API. Compare general, telephony, and specialty models by language, accuracy, and latency.
- Audio Formats: Supported audio formats for the Telnyx Speech-to-Text REST API — WAV, MP3, FLAC, OGG, and more. Includes recommended sample rates and encodings.
- Model Config: Model configuration options for the Telnyx Speech-to-Text REST API — biasing, hotwords, language hints, and per-request tuning of transcription accuracy.
- Response Format: Response format reference for the Telnyx Speech-to-Text REST API — transcript text, words, timestamps, confidence, and per-channel diarization fields.
- Pricing: Pricing for the Telnyx Speech-to-Text REST API. Billed per minute of audio processed, with model-tier rates and volume discounts available on request.
- In-Call Transcription: Real-time speech-to-text during live Telnyx voice calls via Voice API or TeXML.
TTS
- Overview: Synthesize natural speech from text via WebSocket streaming, REST API, or in-call playback.
- Lifecycle: How the Telnyx Text-to-Speech WebSocket endpoint works for real-time streaming, including connection, text messages, audio frames, and shutdown signals.
- Configuration: Configuration surfaces for Telnyx Text-to-Speech WebSocket streaming, including connection-time query parameters and per-message voice settings.
- Messages: Reference for WebSocket frame types — client-to-server text messages and server-to-client audio frames — used in the Telnyx Text-to-Speech streaming API.
- Errors: WebSocket TTS error codes and troubleshooting tips for handshake failures, authentication issues, and runtime streaming errors on the Telnyx platform.
- Examples: Working Python and JavaScript code samples showing how to stream text-to-speech audio over the Telnyx WebSocket TTS endpoint with basic and advanced setups.
- Overview: Single-request text-to-speech with HTTP chunked streaming — start playing before synthesis finishes.
- Request: REST TTS request body fields — text, voice, output type, and provider-specific settings.
- Response: REST TTS response formats — streaming audio, base64, and async retrieval.
- Examples: Code examples for the Telnyx Text-to-Speech REST API showing OpenAI SDK compatibility, synchronous and streaming playback, and async retrieval.
- API Reference: OpenAPI reference for the Telnyx Text-to-Speech REST endpoints, including the generate speech endpoint and request/response schemas for all parameters.
- Overview: Overview of Telnyx native text-to-speech models, comparing latency, quality, language coverage, and expressive control for the TTS REST API.
- Natural: Telnyx Natural is a low-latency English text-to-speech model backed by Rime Mist, designed for real-time voice agents and conversational applications.
- NaturalHD: Telnyx NaturalHD is a high-fidelity multilingual text-to-speech model backed by Rime Arcana, offering studio-quality voices for premium voice applications.
- KokoroTTS: Telnyx KokoroTTS is a lightweight, lowest-latency text-to-speech model ideal for real-time voice agents and interactive applications on the Telnyx TTS API.
- Qwen3TTS: Telnyx Qwen3TTS provides high-quality voice cloning with native support for 11 languages, ideal for multilingual voice agents on the Telnyx TTS API.
- Ultra: Telnyx Ultra text-to-speech delivers sub-100ms latency across 44 languages, available exclusively through the TTS REST API for ultra-fast voice synthesis.
- Grok: xAI Grok voices for expressive, multilingual text-to-speech in Telnyx Voice AI Assistants.
- Bayan: Telnyx Bayan is an Arabic text-to-speech model covering 13 dialects plus English, ideal for Arabic-language voice agents on the Telnyx TTS API.
- Sukhan: Telnyx Sukhan is an Urdu text-to-speech model with a curated set of 14 voices, ideal for Urdu-language voice agents on the Telnyx TTS API.
- Rime: Configure Rime as a text-to-speech provider on Telnyx with Coda and ArcanaV3 models, voice format strings, speed control, and language coverage.
- Minimax: Configure Minimax as a text-to-speech provider on Telnyx with expressive voices and fine-grained speed, volume, and pitch controls per request.
- Resemble: Configure Resemble AI as a text-to-speech provider on Telnyx using your own Resemble API key, with voice cloning and custom voice support.
- Inworld: Configure Inworld as a text-to-speech provider on Telnyx with Mini low-latency, Max high-quality, and TTS-2 latest-generation models, voice format strings, and language support.
- xAI: xAI Grok TTS provider — expressive multilingual voices with speech tags and auto language detection.
- Fish Audio: Configure Fish Audio as a text-to-speech provider on Telnyx with curated voices, cross-lingual synthesis, and multiple audio formats and sample rates.
- AWS Polly: Configure AWS Polly as a text-to-speech provider on Telnyx, with neural, generative, and long-form synthesis engines and voice format strings.
- SSML Tags: Programmable voice with SSML tags from Telnyx - perfect for your
- Azure: Configure Microsoft Azure Speech as a text-to-speech provider on Telnyx, with multilingual neural voices, voice format strings, and SSML support.
- ElevenLabs: Configure ElevenLabs as a text-to-speech provider on Telnyx using your own ElevenLabs API key, with voice cloning and premium voice selection.
- Pronunciation Dictionaries: Control how specific words are spoken during TTS synthesis with custom pronunciation dictionaries.
- In-Call Playback: Play Telnyx text-to-speech audio during live voice calls using the Programmable Voice API or TeXML, with options for streaming and per-call voice selection.
- Pricing: Pricing for the Telnyx Text-to-Speech REST API, including per-character rates by engine, voice tier, and supported provider model.
Voice Design
- Overview: Create custom voices from text descriptions or audio recordings, then use them across all Telnyx voice products.
- Overview: Use AI to generate voices from natural language descriptions — understand the concepts, then create one via the portal or API.
- Quickstart: Create a custom synthetic voice step-by-step using the Telnyx Voice Design portal or API, including prompts, reference audio, and provider selection.
- Parameters: Provider differences and generation parameters for the Voice Design API.
- Prompting Guide: Write voice descriptions that produce consistent, high-quality results — format templates, dimension guides, and common pitfalls.
- Overview: Clone a voice from a short audio recording — capture a speaker’s identity from a sample.
- Quickstart: Clone a voice step-by-step — upload a file, record in the browser, or use the API.
- Parameters: Models, audio requirements, and async flows for the Voice Clone API.
- Responses: Reference for the Voice Clone API response, including the voice ID format, status fields, sample audio URLs, and timestamps returned for each clone.
- Errors: Reference for error codes returned by the Telnyx Voice Clone API, including general errors and provider-specific failures with troubleshooting tips.
- Using Custom Voices: Use your custom voice clones across AI Assistants, Call Control, and the TTS API.
API Reference (Voice)
Audio
- Transcribe speech to text: Transcribe speech to text. This endpoint is consistent with the OpenAI Transcription API and may be used with the OpenAI JS or Python SDK.
Text To Speech Commands
- List available voices: Retrieve a list of available voices from one or all TTS providers. When
provideris specified, returns voices for that provider only. Otherwise, returns voic…
Voice Designs
- List voice designs: Returns a paginated list of voice designs belonging to the authenticated account.
- Create or add a version to a voice design: Creates a new voice design (version 1) when
voice_design_idis omitted. Whenvoice_design_idis provided, adds a new version to the existing design instead… - Get a voice design: Returns the latest version of a voice design, or a specific version when
?version=Nis provided. Theidparameter accepts either a UUID or the design name. - Rename a voice design: Updates the name of a voice design. All versions retain their other properties.
- Delete a voice design: Permanently deletes a voice design and all of its versions. This action cannot be undone.
- Download voice design audio sample: Downloads the WAV audio sample for the voice design. Returns the latest version’s sample by default, or a specific version when
?version=Nis provided. The `… - Delete a specific version of a voice design: Permanently deletes a specific version of a voice design. The version number must be a positive integer.
Voice Clones
- List voice clones: Returns a paginated list of voice clones belonging to the authenticated account.
- Create a voice clone from a voice design: Creates a new voice clone by capturing the voice identity of an existing voice design. The clone can then be used for text-to-speech synthesis.
- Create a voice clone from an audio file upload: Creates a new voice clone by uploading an audio file directly. Supported formats: WAV, MP3, FLAC, OGG, M4A. For best results, provide 5–10 seconds of clear spe…
- Update a voice clone: Updates the name, language, or gender of a voice clone.
- Delete a voice clone: Permanently deletes a voice clone. This action cannot be undone.
- Download voice clone audio sample: Downloads the WAV audio sample that was used to create the voice clone.