The voice layer turns an Agent into a real-time voice agent: a call looks like one text turn per utterance, and the platform runs speech-to-text, synthesis, pacing, barge-in, and the per-turn clock. See the Build a Voice Agent guide for the end-to-end tutorial.
voice(Base)
The mixin. Wrap Agent (or an Agent subclass) with voice() to add the turn engine and the hooks below. The class stays a normal Agent — durable state, messages, RPC, and everything else work unchanged.
Hooks
Override these protected methods. Only reply() is required.
reply(transcript, call)
Required. One turn. Called the moment the recognizer guesses the caller has stopped — about 300 ms before it is sure. Return an AsyncGenerator<string> (or a string, or undefined to stay silent); the tokens are synthesized and streamed as they arrive.
Do not await call.confirmed at the top of reply(). Start the model on transcript now — that speculative head start is the largest latency win the layer offers. Commit history from call.confirmed inside the work the generator does; if the guess was wrong the turn is withdrawn and leaves no trace.
greeting(call)
Spoken as soon as the call is answered, before the caller says anything. Return a string (or undefined for no greeting).
webSocket(ws)
The media stream arrives here. Call this.answer(ws, options) to hand it to the voice layer. This override is required — without it the socket opens and the call is never picked up.
turnEnded(metrics, call)
Called once per turn with a VoiceTurnMetrics — the per-stage clock. Use it for latency logging.
bargedIn(call, cut)
Called when the caller interrupts the agent mid-sentence. cut is a VoiceBargeIn with what the agent was going to say and what the caller is confirmed to have heard.
keyPressed(digit, call)
A DTMF digit. Return a reply like reply(), or undefined.
streamError(error, call) · callEnded(call, reason)
Error and lifecycle hooks. callEnded fires once with the end reason ("closed", "error", …).
speechProviders()
Return the STT/TTS pair, e.g. telnyxSpeechProviders({ synthesizer: { voice: "..." } }). The default pair is Telnyx Flux STT + Telnyx TTS over the runtime’s bound connection.
answer(ws, options?)
Hands the media socket to the voice layer and returns the VoiceSession. Per-call knobs go in options (AnswerOptions).
AnswerOptions
VoiceCall
The call handle passed to every hook. Read-only fields plus the Call Control verbs.
VoiceTurnMetrics
The per-turn clock, passed to turnEnded. Headline figure: toFirstAudioMs — caller speech-stop → first assistant audio.
VoiceBargeIn
Passed to bargedIn.
Providers
telnyxSpeechProviders(options?)
Returns the { recognizer, synthesizer } pair. Override the voice via synthesizer.voice:
createCallControlClient()
A pre-credentialed Call Control client (for the webhook’s answer / streaming_start), using the runtime’s bound connection — no API key in your code.
Also exported
VoiceSession, attachVoiceSocket, sessionOf, TelnyxSpeechRecognizer, TelnyxSpeechSynthesizer, SentenceAssembler, PlayoutQueue, JitterBuffer, and the codec registry — for building a voice loop without the mixin, or replacing a stage. See the source TSDoc in @telnyx/edge-runtime/voice for the full surface.