Skip to main content
Answer a phone call and hold a real-time spoken conversation, all from an Edge Compute function at the edge, next to the caller. Speech-to-text, a large language model, and text-to-speech are wired together for you — you implement one method that returns the answer, and the platform runs the audio pipeline with barge-in, pacing, and turn-taking handled. This guide builds the smallest complete voice agent: it greets the caller, listens, thinks, and replies, over a real PSTN call. Nothing about speech is configured — no STT model, no TTS voice, no codec, no API key. The defaults are the product.

How it works

A voice agent is an Agent that extends the voice() mixin. One agent instance is created per call, so each conversation keeps its own history and state. A call reaches the agent in two hops:
  1. Call Control answers the incoming call and starts a bidirectional media stream, pointing it at your function.
  2. The media stream — a WebSocket — is handed straight to the agent instance by mountAgents. From there the voice layer runs the loop: recognize speech, call your reply() hook for the answer, synthesize it, and stream audio back to the caller.
You implement one method — reply() — that returns the spoken answer as a stream of text. Everything else is default.
The speech providers (Telnyx Flux STT + Telnyx TTS) run over the runtime’s own bound connection — no API key appears anywhere in your code.

Prerequisites

  • An Edge Compute account and the Telnyx CLI — see the Quick Start.
  • A Telnyx phone number and a Call Control application whose webhook URL points at your deployed function’s / route.

Create the project

Scaffold a StatefulActor project with the CLI. This registers the function and generates package.json, tsconfig.json, and a telnyx.toml that already carries the function’s identity:
The agent pulls in two packages beyond the scaffold’s defaults — the telnyx client for the model call, and the ws types for the WebSocket parameter — so add them, then install:
new-func writes an [edge_compute] block (the func_id) into telnyx.toml. telnyx-edge ship requires it — keep that block when you edit the manifest below.

The agent

The whole agent is one class. Two hooks are required: reply(), which returns the spoken answer, and webSocket(), which hands the media stream to the voice layer (the base Agent has no default, so the call is never picked up without it). greeting() is an optional override shown for completeness.
src/index.ts

The front door

Two more pieces: the Call Control webhook that hands a call to the agent, and mountAgents, which routes the media-stream WebSocket to the right instance.
src/index.ts
The model is any streaming source that yields text. This one reads the account’s inference endpoint through the [telnyx] binding, which arrives already credentialed:
src/index.ts

Configure and deploy

Edit the generated telnyx.toml so it declares the Support actor and the [telnyx] binding — nothing about speech. Leave the [edge_compute] block exactly as new-func wrote it (ship needs the func_id):
telnyx.toml
1

Ship the function

2

Point your Call Control app's webhook at the deployed function's / route

The webhook answers the call and starts the media stream.
3

Dial your number

The agent greets you, and you can talk to it.

The head start

The single most important detail is in reply(): do not await call.confirmed. The recognizer calls reply() the moment it guesses the caller has stopped — roughly 300 ms before it is certain. Handing that guess to the model immediately overlaps the model’s first-token latency with the tail of the caller’s speech. History is written only once call.confirmed resolves inside respond(); if the guess was wrong, the turn is withdrawn and leaves no trace — so a turn that ran ahead and one that did not write identical history.
Awaiting call.confirmed at the top of reply() throws that head start away — it is the largest avoidable latency cost in the layer.

Tuning

Per-call options are passed to answer() inside webSocket():
  • bargeIn — let the caller interrupt the agent mid-sentence ({ minWords: 2 } requires two words before cutting, to ignore backchannels like “uh-huh”).
  • keypressInterrupts — whether DTMF digits stop playback.
  • eager — the speculative head start above; on by default.
  • warmup — hold the TTS connection open on answer so no turn pays a cold synthesis handshake. Effective only for voices that accept a streaming connection.
Speech providers and the voice are defaults you can override through telnyxSpeechProviders({ synthesizer: { voice: "..." } }).

Next steps

Mounting agents

How mountAgents routes WebSocket, SSE, and RPC to one agent on one address.

Message history

The durable conversation log the agent reads and writes each turn.