> ## Documentation Index
> Fetch the complete documentation index at: https://developers.telnyx.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Voice

> The voice() mixin, the turn hooks, answer() and its options, and the VoiceCall and VoiceTurnMetrics types.

The voice layer turns an [`Agent`](/docs/agent-sdk/api-reference/agent) into a real-time voice agent: a call looks like one text turn per utterance, and the platform runs speech-to-text, synthesis, pacing, barge-in, and the per-turn clock. See the [Build a Voice Agent](/docs/agent-sdk/voice-agent) guide for the end-to-end tutorial.

```ts theme={null}
import { voice } from "@telnyx/edge-runtime/voice";
import { Agent } from "@telnyx/edge-runtime";

export class Support extends voice(Agent<Env>) {
  protected override reply(transcript: string, call: VoiceCall) {
    return this.respond(transcript, call.confirmed, call.signal);
  }
  override webSocket(ws: WebSocket) {
    this.answer(ws);
  }
}
```

## `voice(Base)`

The mixin. Wrap `Agent` (or an `Agent` subclass) with `voice()` to add the turn engine and the hooks below. The class stays a normal `Agent` — durable state, `messages`, RPC, and everything else work unchanged.

## Hooks

Override these `protected` methods. Only `reply()` is required.

### `reply(transcript, call)`

**Required.** One turn. Called the moment the recognizer *guesses* the caller has stopped — about 300 ms before it is sure. Return an `AsyncGenerator<string>` (or a string, or `undefined` to stay silent); the tokens are synthesized and streamed as they arrive.

<Warning>
  Do **not** `await call.confirmed` at the top of `reply()`. Start the model on `transcript` now — that speculative head start is the largest latency win the layer offers. Commit history from `call.confirmed` inside the work the generator does; if the guess was wrong the turn is withdrawn and leaves no trace.
</Warning>

### `greeting(call)`

Spoken as soon as the call is answered, before the caller says anything. Return a string (or `undefined` for no greeting).

### `webSocket(ws)`

The media stream arrives here. Call `this.answer(ws, options)` to hand it to the voice layer. **This override is required** — without it the socket opens and the call is never picked up.

### `turnEnded(metrics, call)`

Called once per turn with a [`VoiceTurnMetrics`](#voiceturnmetrics) — the per-stage clock. Use it for latency logging.

### `bargedIn(call, cut)`

Called when the caller interrupts the agent mid-sentence. `cut` is a [`VoiceBargeIn`](#voicebargein) with what the agent *was* going to say and what the caller is confirmed to have heard.

### `keyPressed(digit, call)`

A DTMF digit. Return a reply like `reply()`, or `undefined`.

### `streamError(error, call)` · `callEnded(call, reason)`

Error and lifecycle hooks. `callEnded` fires once with the end reason (`"closed"`, `"error"`, …).

### `speechProviders()`

Return the STT/TTS pair, e.g. `telnyxSpeechProviders({ synthesizer: { voice: "..." } })`. The default pair is Telnyx Flux STT + Telnyx TTS over the runtime's bound connection.

## `answer(ws, options?)`

Hands the media socket to the voice layer and returns the `VoiceSession`. Per-call knobs go in `options` ([`AnswerOptions`](#answeroptions)).

```ts theme={null}
this.answer(ws, { bargeIn: { minWords: 2 }, warmup: true });
```

## `AnswerOptions`

| Option | Type | Meaning |
| - | - | - |
| `eager` | `boolean` | Start the model on the recognizer's guess, before end-of-turn is confirmed (\~300 ms head start). **Default `true`.** |
| `warmup` | `boolean` | Open the TTS connection on answer and hold it warm, so no turn pays a cold synthesis handshake. Effective only for voices that accept a streaming connection. |
| `bargeIn` | `{ minWords?: number }` | Let the caller interrupt the agent. `minWords` ignores backchannels ("uh-huh") until that many words are heard. |
| `keypressInterrupts` | `boolean` | Whether a DTMF digit stops playback. |
| `track`, `preRollMs`, `sentences`, `callControl`, `socket` | — | Lower-level media/session tuning. |
| `recognizer`, `synthesizer` | `SpeechProvider` | Override the STT/TTS pair for this call (same shape as `speechProviders()`). |

## `VoiceCall`

The call handle passed to every hook. Read-only fields plus the Call Control verbs.

| Member | Type | Meaning |
| - | - | - |
| `callControlId` | `string` | The call's Call Control id — the same id the verbs act through. |
| `from`, `to` | `string` | The caller and callee numbers. |
| `channel` | `VoiceChannel` | The transport the turn arrived on. |
| `history` | `MessageLog` | The agent's durable conversation log — one conversation across phone and chat. The layer never writes to it; your turn code does. |
| `confirmed` | `Promise<string \| undefined>` | Resolves to the confirmed transcript, or `undefined` if the turn was withdrawn. Commit history from this. |
| `signal` | `AbortSignal` | Aborts when the turn is superseded/interrupted/ended — pass it to the model call. |
| `turn` | `number` | Which turn of the call, from 1. |
| `speculative` | `boolean` | Whether this call fired on the recognizer's guess (eager) rather than a confirmed end-of-turn. |
| `aborted` | `boolean` | `true` once this turn was abandoned. Check it around anything slow. |
| `hangup()`, `say()`, … | — | Call Control verbs on the live call. |

## `VoiceTurnMetrics`

The per-turn clock, passed to `turnEnded`. Headline figure: **`toFirstAudioMs`** — caller speech-stop → first assistant audio.

| Field | Meaning |
| - | - |
| `turn` | Turn number (from 1). |
| `transcript` | What the caller said (confirmed, or the guess if withdrawn). |
| `started` | How the turn began. |
| `outcome` | How it ended (`"completed"`, `"barged-in"`, `"withdrawn"`, `"error"`, …). |
| `toFirstAudioMs` | Speech-stop → first audio (the headline). |
| `anchor` | Which signal `toFirstAudioMs` is measured from (`"speechStop"` / `"eagerEndOfTurn"` / `"endOfTurn"`). |
| `at` | Timeline of stage timestamps: `speechStop`, `endOfTurn`, `firstToken`, `synthesisRequested`, `synthesisFirstAudio`, `firstAudio`. |
| `synthesis` | `{ path, cached, timeToFirstAudioMs }` — how TTS was served (warm vs cold). |

## `VoiceBargeIn`

Passed to `bargedIn`.

| Field | Meaning |
| - | - |
| `text` | Everything the interrupted stretch was going to say. |
| `heard` | The part the caller is confirmed to have heard before the cut (sentences whose playback was confirmed). Empty if the cut landed in the first sentence. |

## Providers

### `telnyxSpeechProviders(options?)`

Returns the `{ recognizer, synthesizer }` pair. Override the voice via `synthesizer.voice`:

```ts theme={null}
telnyxSpeechProviders({ synthesizer: { voice: "Telnyx.NaturalHD.Alloy" } });
```

### `createCallControlClient()`

A pre-credentialed Call Control client (for the webhook's `answer` / `streaming_start`), using the runtime's bound connection — no API key in your code.

## Also exported

`VoiceSession`, `attachVoiceSocket`, `sessionOf`, `TelnyxSpeechRecognizer`, `TelnyxSpeechSynthesizer`, `SentenceAssembler`, `PlayoutQueue`, `JitterBuffer`, and the codec registry — for building a voice loop without the mixin, or replacing a stage. See the source TSDoc in `@telnyx/edge-runtime/voice` for the full surface.
