> ## Documentation Index
> Fetch the complete documentation index at: https://developers.telnyx.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Build a Voice Agent on Edge Compute

> Answer a phone call and hold a spoken conversation from an Edge Compute function — speech-to-text, an LLM, and text-to-speech are wired in. You write the agent's turn.

Answer a phone call and hold a real-time spoken conversation, all from an Edge Compute function at the edge, next to the caller. Speech-to-text, a large language model, and text-to-speech are wired together for you — you implement one method that returns the answer, and the platform runs the audio pipeline with barge-in, pacing, and turn-taking handled.

This guide builds the smallest complete voice agent: it greets the caller, listens, thinks, and replies, over a real PSTN call. Nothing about speech is configured — no STT model, no TTS voice, no codec, no API key. The defaults are the product.

***

## How it works

A voice agent is an [Agent](/docs/agent-sdk/concepts/how-agents-run) that extends the `voice()` mixin. One agent instance is created per call, so each conversation keeps its own history and state.

A call reaches the agent in two hops:

1. **Call Control** answers the incoming call and starts a bidirectional media stream, pointing it at your function.
2. **The media stream** — a WebSocket — is handed straight to the agent instance by `mountAgents`. From there the voice layer runs the loop: recognize speech, call your `reply()` hook for the answer, synthesize it, and stream audio back to the caller.

You implement one method — `reply()` — that returns the spoken answer as a stream of text. Everything else is default.

<Note>
  The speech providers (Telnyx Flux STT + Telnyx TTS) run over the runtime's own bound connection — no API key appears anywhere in your code.
</Note>

## Prerequisites

* An Edge Compute account and the Telnyx CLI — see the [Quick Start](/docs/edge-compute/quickstart).
* A Telnyx phone number and a [Call Control application](/docs/voice/programmable-voice) whose webhook URL points at your deployed function's `/` route.

## Create the project

Scaffold a StatefulActor project with the CLI. This registers the function and generates `package.json`, `tsconfig.json`, and a `telnyx.toml` that already carries the function's identity:

```bash theme={null}
telnyx-edge new-func --actor -l ts -n voice-minimal
cd voice-minimal
```

The agent pulls in two packages beyond the scaffold's defaults — the `telnyx` client for the model call, and the `ws` types for the `WebSocket` parameter — so add them, then install:

```bash theme={null}
npm install telnyx
npm install -D ws @types/ws
npm install
```

<Note>
  `new-func` writes an `[edge_compute]` block (the `func_id`) into `telnyx.toml`. `telnyx-edge ship` requires it — keep that block when you edit the manifest [below](#configure-and-deploy).
</Note>

## The agent

The whole agent is one class. Two hooks are required: `reply()`, which returns the spoken answer, and `webSocket()`, which hands the media stream to the voice layer (the base `Agent` has no default, so the call is never picked up without it). `greeting()` is an optional override shown for completeness.

```typescript src/index.ts theme={null}
import { Agent, rpc, type ActorNamespace } from "@telnyx/edge-runtime";
import { mountAgents } from "@telnyx/edge-runtime/mount";
import { voice, createCallControlClient, type VoiceCall } from "@telnyx/edge-runtime/voice";
import type { WebSocket } from "ws";
import type Telnyx from "telnyx";

const MODEL = "meta-llama/Llama-3.3-70B-Instruct";
const SYSTEM =
  "You are a concise, friendly voice assistant reached over a phone call. " +
  "Answer in one or two short sentences of plain spoken English. No markdown, no lists.";

type ChatMessage = Telnyx.AI.OpenAI.Chat.ChatCreateCompletionParams.Message;

export class Support extends voice(Agent<Env>) {
  /** One turn: start the model on `text` now, commit history once confirmed. */
  async *respond(text: string, commit: Promise<string | undefined>, signal?: AbortSignal): AsyncGenerator<string> {
    const history = (await this.messages.toOpenAI()).slice(-12) as ChatMessage[];
    const convo: ChatMessage[] = [{ role: "system", content: SYSTEM }, ...history, { role: "user", content: text }];

    let answer = "";
    for await (const token of completeStream(this.env.TELNYX, convo, signal)) {
      answer += token;
      yield token;
    }

    // Nothing durable is written until the transcript is confirmed. A turn that
    // ran ahead and one that did not write byte-identical history.
    const confirmed = await commit;
    if (confirmed === undefined) return;
    await this.messages.add("user", confirmed);
    await this.messages.add("assistant", answer);
  }

  /** Spoken as soon as the call is answered, before the caller says anything. */
  protected override greeting(): string {
    return "Hi, this is the Telnyx Edge voice agent. Ask me anything.";
  }

  /** The one required hook. Called the moment the recognizer *guesses* the
   *  caller stopped — do NOT await `call.confirmed` here (see the head start). */
  protected override reply(transcript: string, call: VoiceCall): AsyncGenerator<string> {
    return this.respond(transcript, call.confirmed, call.signal);
  }

  /** Required. The media stream arrives here; `answer` hands it to the voice layer.
   *  Per-call knobs live here, e.g. `this.answer(ws, { bargeIn: { minWords: 2 } })`. */
  override webSocket(ws: WebSocket): void {
    this.answer(ws);
  }
}
```

## The front door

Two more pieces: the Call Control webhook that hands a call to the agent, and `mountAgents`, which routes the media-stream WebSocket to the right instance.

```typescript src/index.ts theme={null}
const calls = createCallControlClient();

async function webhook(req: Request): Promise<Response> {
  const host = req.headers.get("x-forwarded-host") ?? req.headers.get("host") ?? "";

  let event: { data?: { event_type?: string; payload?: Record<string, unknown> } };
  try {
    event = (await req.json()) as typeof event;
  } catch {
    return new Response("bad json", { status: 400 });
  }

  const type = event.data?.event_type;
  const payload = event.data?.payload ?? {};
  const callControlId = String(payload.call_control_id ?? "");
  if (callControlId === "") return new Response("ok");

  if (type === "call.initiated" && payload.direction === "incoming") {
    await calls.command(callControlId, "answer", {});
  } else if (type === "call.answered") {
    await calls.command(callControlId, "streaming_start", {
      // The id goes in whole, percent-encoded; mountAgents decodes it and runs
      // it through the SDK's reversible encoder, so one call maps to exactly
      // one agent instance.
      stream_url: `wss://${host}/agents/support/${encodeURIComponent(callControlId)}`,
      stream_track: "inbound_track",
      stream_bidirectional_mode: "rtp",
      stream_bidirectional_codec: "PCMU",
      stream_bidirectional_target_legs: "self",
      stream_bidirectional_sampling_rate: 8000,
    });
  }

  return new Response("ok");
}

interface Env {
  SUPPORT: ActorNamespace;
  /** The `[telnyx]` binding: a Telnyx client the platform hands you already credentialed. */
  TELNYX: Telnyx;
}

export default {
  fetch: mountAgents<Env>((env) => ({ support: env.SUPPORT }), { fallback: webhook }),
};
```

The model is any streaming source that yields text. This one reads the account's inference endpoint through the `[telnyx]` binding, which arrives already credentialed:

```typescript src/index.ts theme={null}
async function* completeStream(telnyx: Telnyx, messages: ChatMessage[], signal?: AbortSignal): AsyncGenerator<string> {
  const res = await telnyx.ai.openai.chat
    .createCompletion({ model: MODEL, messages, stream: true, max_tokens: 120 }, { signal })
    .asResponse();
  if (!res.ok || res.body === null) throw new Error(`chat completion ${res.status}`);

  const reader = res.body.getReader();
  const decoder = new TextDecoder();
  let buffered = "";
  for (;;) {
    const { value, done } = await reader.read();
    if (done) break;
    buffered += decoder.decode(value, { stream: true });
    let nl: number;
    while ((nl = buffered.indexOf("\n")) >= 0) {
      const line = buffered.slice(0, nl).trim();
      buffered = buffered.slice(nl + 1);
      if (!line.startsWith("data:")) continue;
      const data = line.slice(5).trim();
      if (data === "[DONE]") return;
      try {
        const token = JSON.parse(data)?.choices?.[0]?.delta?.content;
        if (token) yield token;
      } catch {
        /* keep-alive or partial line */
      }
    }
  }
}
```

## Configure and deploy

Edit the generated `telnyx.toml` so it declares the `Support` actor and the `[telnyx]` binding — nothing about speech. Leave the `[edge_compute]` block exactly as `new-func` wrote it (`ship` needs the `func_id`):

```toml telnyx.toml theme={null}
name = "voice-minimal"
main = "src/index.ts"
compatibility_date = "2026-09-15"

[[actors]]
binding = "SUPPORT"
type    = "Support"

[telnyx]
binding = "TELNYX"

# Written by `new-func` — keep as generated (do not edit or remove).
[edge_compute]
func_id   = "00000000-0000-0000-0000-000000000000"
func_name = "voice-minimal"
```

<Steps>
  <Step title="Ship the function">
    ```bash theme={null}
    telnyx-edge ship
    ```
  </Step>

  <Step title="Point your Call Control app's webhook at the deployed function's `/` route">
    The webhook answers the call and starts the media stream.
  </Step>

  <Step title="Dial your number">
    The agent greets you, and you can talk to it.
  </Step>
</Steps>

## The head start

The single most important detail is in `reply()`: **do not `await call.confirmed`.**

The recognizer calls `reply()` the moment it *guesses* the caller has stopped — roughly 300 ms before it is certain. Handing that guess to the model immediately overlaps the model's first-token latency with the tail of the caller's speech. History is written only once `call.confirmed` resolves inside `respond()`; if the guess was wrong, the turn is withdrawn and leaves no trace — so a turn that ran ahead and one that did not write identical history.

<Warning>
  Awaiting `call.confirmed` at the top of `reply()` throws that head start away — it is the largest avoidable latency cost in the layer.
</Warning>

## Tuning

Per-call options are passed to `answer()` inside `webSocket()`:

* **`bargeIn`** — let the caller interrupt the agent mid-sentence (`{ minWords: 2 }` requires two words before cutting, to ignore backchannels like "uh-huh").
* **`keypressInterrupts`** — whether DTMF digits stop playback.
* **`eager`** — the speculative head start above; on by default.
* **`warmup`** — hold the TTS connection open on answer so no turn pays a cold synthesis handshake. Effective only for voices that accept a streaming connection.

Speech providers and the voice are defaults you can override through `telnyxSpeechProviders({ synthesizer: { voice: "..." } })`.

## Next steps

<CardGroup cols={2}>
  <Card title="Mounting agents" href="/docs/agent-sdk/mounting-agents">
    How `mountAgents` routes WebSocket, SSE, and RPC to one agent on one address.
  </Card>

  <Card title="Message history" href="/docs/agent-sdk/message-history">
    The durable conversation log the agent reads and writes each turn.
  </Card>
</CardGroup>
