How it works
A voice agent is an Agent that extends thevoice() mixin. One agent instance is created per call, so each conversation keeps its own history and state.
A call reaches the agent in two hops:
- Call Control answers the incoming call and starts a bidirectional media stream, pointing it at your function.
- The media stream — a WebSocket — is handed straight to the agent instance by
mountAgents. From there the voice layer runs the loop: recognize speech, call yourreply()hook for the answer, synthesize it, and stream audio back to the caller.
reply() — that returns the spoken answer as a stream of text. Everything else is default.
The speech providers (Telnyx Flux STT + Telnyx TTS) run over the runtime’s own bound connection — no API key appears anywhere in your code.
Prerequisites
- An Edge Compute account and the Telnyx CLI — see the Quick Start.
- A Telnyx phone number and a Call Control application whose webhook URL points at your deployed function’s
/route.
Create the project
Scaffold a StatefulActor project with the CLI. This registers the function and generatespackage.json, tsconfig.json, and a telnyx.toml that already carries the function’s identity:
telnyx client for the model call, and the ws types for the WebSocket parameter — so add them, then install:
new-func writes an [edge_compute] block (the func_id) into telnyx.toml. telnyx-edge ship requires it — keep that block when you edit the manifest below.The agent
The whole agent is one class. Two hooks are required:reply(), which returns the spoken answer, and webSocket(), which hands the media stream to the voice layer (the base Agent has no default, so the call is never picked up without it). greeting() is an optional override shown for completeness.
src/index.ts
The front door
Two more pieces: the Call Control webhook that hands a call to the agent, andmountAgents, which routes the media-stream WebSocket to the right instance.
src/index.ts
[telnyx] binding, which arrives already credentialed:
src/index.ts
Configure and deploy
Edit the generatedtelnyx.toml so it declares the Support actor and the [telnyx] binding — nothing about speech. Leave the [edge_compute] block exactly as new-func wrote it (ship needs the func_id):
telnyx.toml
1
Ship the function
2
Point your Call Control app's webhook at the deployed function's / route
The webhook answers the call and starts the media stream.
3
Dial your number
The agent greets you, and you can talk to it.
The head start
The single most important detail is inreply(): do not await call.confirmed.
The recognizer calls reply() the moment it guesses the caller has stopped — roughly 300 ms before it is certain. Handing that guess to the model immediately overlaps the model’s first-token latency with the tail of the caller’s speech. History is written only once call.confirmed resolves inside respond(); if the guess was wrong, the turn is withdrawn and leaves no trace — so a turn that ran ahead and one that did not write identical history.
Tuning
Per-call options are passed toanswer() inside webSocket():
bargeIn— let the caller interrupt the agent mid-sentence ({ minWords: 2 }requires two words before cutting, to ignore backchannels like “uh-huh”).keypressInterrupts— whether DTMF digits stop playback.eager— the speculative head start above; on by default.warmup— hold the TTS connection open on answer so no turn pays a cold synthesis handshake. Effective only for voices that accept a streaming connection.
telnyxSpeechProviders({ synthesizer: { voice: "..." } }).
Next steps
Mounting agents
How
mountAgents routes WebSocket, SSE, and RPC to one agent on one address.Message history
The durable conversation log the agent reads and writes each turn.