> ## Documentation Index
> Fetch the complete documentation index at: https://developers.telnyx.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Delegation

> Split an AI assistant between a fast model that talks to the caller and a capable model that does the work — configure the backend, or answer delegations yourself over your own WebSocket.

Delegation splits a conversation between two models: a **frontend** model that talks to the caller, and a **backend** model that does the work. The frontend keeps the caller company — fast, cheap, low latency — while the backend looks things up and runs tools.

This matters most on [GPT-Live](/docs/voice/sip-trunking/gpt-live-configuration-guide) speech-to-speech conversations, where the frontend model cannot call tools at all. Without delegation such an assistant can hold a conversation but can never look anything up or act on the caller's behalf, which is why delegation is **enabled by default**.

<Note>
  **Beta.** Configure delegation with `delegation_settings` on the assistant.
</Note>

## How a delegation is raised

The two conversation paths differ only in how the frontend asks for help.

**GPT-Live (speech-to-speech).** The frontend model has no tools. When it needs work done it raises a delegation and *waits* — so the backend's answer streams back sentence by sentence, and the model speaks it as it arrives. The frontend prompt is deliberately small; the business rules belong on the backend.

**Chat completion (STT → LLM → TTS).** The frontend is stripped down to a single `delegate` tool. Calling it hands the work over and returns immediately, so the conversation carries on while the backend works:

> **Caller:** Can you check the status of my order?
> **Assistant:** Sure, let me take a look. *(calls `delegate`, keeps talking)*
> *(backend looks up the order)*
> **Assistant:** It shipped Tuesday and should arrive Thursday.

Either way the result is injected back as context, not spoken verbatim by the backend.

## Configure it

```bash theme={null}
curl -X POST https://api.telnyx.com/v2/ai/assistants/{assistant_id} \
  -H "Authorization: Bearer $TELNYX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "delegation_settings": {
      "enabled": true,
      "mode": "telnyx",
      "model": "openai/gpt-4o",
      "instructions": "You have access to the order system. Always confirm the order number before answering.",
      "speak_results": true
    }
  }'
```

| Field             | Default          | Description                                                             |
| ----------------- | ---------------- | ----------------------------------------------------------------------- |
| `enabled`         | `true`           | Whether the assistant delegates work to a backend model.                |
| `mode`            | `"telnyx"`       | Who answers a delegation — see [Choosing a mode](#choosing-a-mode).     |
| `model`           | platform default | The backend model. Must be a model available for AI Assistants.         |
| `llm_api_key_ref` | —                | Integration secret holding the API key for `model`.                     |
| `instructions`    | —                | Extra instructions for the backend, in addition to the assistant's own. |
| `speak_results`   | `true`           | Whether the backend's answer is spoken to the caller.                   |
| `external_llm`    | —                | Run the backend on your own OpenAI-compatible endpoint.                 |

`{"enabled": true}` on its own is a valid configuration: with no `model`, the backend runs on the platform default.

<Warning>
  `delegation_settings` does not accept a raw `api_key`, and neither does its nested `external_llm`. Passing one is **rejected**, not ignored — a plaintext credential in the assistant payload is a leak, and silently falling back to a different key would hide the misconfiguration. Reference an [integration secret](/docs/inference/ai-assistants/integrations) with `llm_api_key_ref` (or `external_llm.llm_api_key_ref`) instead.
</Warning>

### Spoken results versus silent context

`speak_results` decides what happens to the backend's answer:

* `true` (default) — the result is appended as **commentary** and paraphrased aloud. Use this when the caller is waiting on the answer.
* `false` — the result is kept as **silent context** that informs later answers without being read out. Use this when the backend is enriching what the assistant knows rather than answering a direct question.

### Writing backend instructions

`instructions` are added to the assistant's own instructions for the backend model only. Put the business rules, lookup procedures and tool guidance here — the frontend model does not need them, and on GPT-Live it has a small context window that is better spent on holding a natural conversation.

## Choosing a mode

### `telnyx` — Telnyx runs the backend

The default. Telnyx runs the backend model with the assistant's own tools, MCP servers and observability, so a delegation can do anything the assistant could do.

Validation happens when you save the assistant: an unavailable model, or an `llm_api_key_ref` that does not resolve, is rejected there rather than surfacing mid-call as an assistant that talks but can never look anything up.

<Note>
  On GPT-Live, an OpenAI backend model is handed to OpenAI's own delegation at session start and runs there, calling back into Telnyx for the assistant's tools. MCP servers are not available on that path.
</Note>

### `client` — you answer the delegation

Set `mode: "client"` and Telnyx relays each delegation to your own server over the WebSocket configured in [`websocket_settings`](/docs/inference/ai-assistants/conversation-event-stream). Use this when the answer has to come from a system that cannot be reached as a Telnyx tool — an internal service behind your own auth, a model you host, a human in the loop.

Telnyx sends:

```json theme={null}
{
  "type": "session.delegation.created",
  "delegation": {
    "id": "5b1d9e02-4c77-4f3a-8a2e-1c6f0b9d7a54",
    "request": "Check the status of order 40192.",
    "instructions": "You have access to the order system."
  }
}
```

You answer with the same `id`:

```json theme={null}
{
  "type": "session.delegation.completed",
  "delegation": {
    "id": "5b1d9e02-4c77-4f3a-8a2e-1c6f0b9d7a54",
    "output": "Order 40192 shipped on Tuesday and is due to arrive Thursday."
  }
}
```

Things to know about `client` mode:

* **It requires an active event-stream socket.** If none is connected when a delegation is raised, the delegation is refused: the assistant tells the caller it cannot look things up right now and carries on from what it already knows. It never leaves the caller waiting in silence.
* **The answer is text only.** The event stream offers your server no tool vocabulary, so `output` is plain text that gets injected as context.
* **An empty answer is a failure, not a quiet success.** The assistant has already told the caller it is checking, so an `output` that is blank after trimming is rejected rather than producing dead air.
* **`request` is `null` on GPT-Live.** The live model raises a delegation with no text of its own, so you work from the conversation events the socket is already streaming you.
* **Answer promptly.** A delegation that is not answered in time falls back the same way an unavailable socket does.

## Model requirements

GPT-Live conversations are selected by the assistant's `model` and `voice`:

| Setting                | Requirement                                                                      |
| ---------------------- | -------------------------------------------------------------------------------- |
| `model`                | The model family must start with `gpt-live` (for example `openai/gpt-live-...`). |
| `voice_settings.voice` | Must be prefixed `OpenAILive.`                                                   |

The two have to be paired. An assistant with a GPT-Live model and a non-live voice — or the reverse — is rejected when you save it, rather than failing when a call arrives. Conversation flows are not supported on the GPT-Live route.

The backend model in `delegation_settings.model` has no such restriction: it is an ordinary text model, and it is where tool use actually happens.

## Next steps

* [Conversation event stream](/docs/inference/ai-assistants/conversation-event-stream) — the WebSocket that `mode: "client"` delegations travel over
* [Tools library](/docs/inference/ai-assistants/tools-library) — the tools a `telnyx`-mode backend can call
* [Custom LLM](/docs/inference/ai-assistants/custom-llm) — running a model on your own endpoint with `external_llm`
