Skip to main content
Delegation splits a conversation between two models: a frontend model that talks to the caller, and a backend model that does the work. The frontend keeps the caller company — fast, cheap, low latency — while the backend looks things up and runs tools. This matters most on GPT-Live speech-to-speech conversations, where the frontend model cannot call tools at all. Without delegation such an assistant can hold a conversation but can never look anything up or act on the caller’s behalf, which is why delegation is enabled by default.
Beta. Configure delegation with delegation_settings on the assistant.

How a delegation is raised

The two conversation paths differ only in how the frontend asks for help. GPT-Live (speech-to-speech). The frontend model has no tools. When it needs work done it raises a delegation and waits — so the backend’s answer streams back sentence by sentence, and the model speaks it as it arrives. The frontend prompt is deliberately small; the business rules belong on the backend. Chat completion (STT → LLM → TTS). The frontend is stripped down to a single delegate tool. Calling it hands the work over and returns immediately, so the conversation carries on while the backend works:
Caller: Can you check the status of my order? Assistant: Sure, let me take a look. (calls delegate, keeps talking) (backend looks up the order) Assistant: It shipped Tuesday and should arrive Thursday.
Either way the result is injected back as context, not spoken verbatim by the backend.

Configure it

{"enabled": true} on its own is a valid configuration: with no model, the backend runs on the platform default.
delegation_settings does not accept a raw api_key, and neither does its nested external_llm. Passing one is rejected, not ignored — a plaintext credential in the assistant payload is a leak, and silently falling back to a different key would hide the misconfiguration. Reference an integration secret with llm_api_key_ref (or external_llm.llm_api_key_ref) instead.

Spoken results versus silent context

speak_results decides what happens to the backend’s answer:
  • true (default) — the result is appended as commentary and paraphrased aloud. Use this when the caller is waiting on the answer.
  • false — the result is kept as silent context that informs later answers without being read out. Use this when the backend is enriching what the assistant knows rather than answering a direct question.

Writing backend instructions

instructions are added to the assistant’s own instructions for the backend model only. Put the business rules, lookup procedures and tool guidance here — the frontend model does not need them, and on GPT-Live it has a small context window that is better spent on holding a natural conversation.

Choosing a mode

telnyx — Telnyx runs the backend

The default. Telnyx runs the backend model with the assistant’s own tools, MCP servers and observability, so a delegation can do anything the assistant could do. Validation happens when you save the assistant: an unavailable model, or an llm_api_key_ref that does not resolve, is rejected there rather than surfacing mid-call as an assistant that talks but can never look anything up.
On GPT-Live, an OpenAI backend model is handed to OpenAI’s own delegation at session start and runs there, calling back into Telnyx for the assistant’s tools. MCP servers are not available on that path.

client — you answer the delegation

Set mode: "client" and Telnyx relays each delegation to your own server over the WebSocket configured in websocket_settings. Use this when the answer has to come from a system that cannot be reached as a Telnyx tool — an internal service behind your own auth, a model you host, a human in the loop. Telnyx sends:
You answer with the same id:
Things to know about client mode:
  • It requires an active event-stream socket. If none is connected when a delegation is raised, the delegation is refused: the assistant tells the caller it cannot look things up right now and carries on from what it already knows. It never leaves the caller waiting in silence.
  • The answer is text only. The event stream offers your server no tool vocabulary, so output is plain text that gets injected as context.
  • An empty answer is a failure, not a quiet success. The assistant has already told the caller it is checking, so an output that is blank after trimming is rejected rather than producing dead air.
  • request is null on GPT-Live. The live model raises a delegation with no text of its own, so you work from the conversation events the socket is already streaming you.
  • Answer promptly. A delegation that is not answered in time falls back the same way an unavailable socket does.

Model requirements

GPT-Live conversations are selected by the assistant’s model and voice: The two have to be paired. An assistant with a GPT-Live model and a non-live voice — or the reverse — is rejected when you save it, rather than failing when a call arrives. Conversation flows are not supported on the GPT-Live route. The backend model in delegation_settings.model has no such restriction: it is an ordinary text model, and it is where tool use actually happens.

Next steps