Beta. Configure delegation with
delegation_settings on the assistant.How a delegation is raised
The two conversation paths differ only in how the frontend asks for help. GPT-Live (speech-to-speech). The frontend model has no tools. When it needs work done it raises a delegation and waits — so the backend’s answer streams back sentence by sentence, and the model speaks it as it arrives. The frontend prompt is deliberately small; the business rules belong on the backend. Chat completion (STT → LLM → TTS). The frontend is stripped down to a singledelegate tool. Calling it hands the work over and returns immediately, so the conversation carries on while the backend works:
Caller: Can you check the status of my order?
Assistant: Sure, let me take a look. (calls delegate, keeps talking)
(backend looks up the order)
Assistant: It shipped Tuesday and should arrive Thursday.
Either way the result is injected back as context, not spoken verbatim by the backend.
Configure it
{"enabled": true} on its own is a valid configuration: with no model, the backend runs on the platform default.
Spoken results versus silent context
speak_results decides what happens to the backend’s answer:
true(default) — the result is appended as commentary and paraphrased aloud. Use this when the caller is waiting on the answer.false— the result is kept as silent context that informs later answers without being read out. Use this when the backend is enriching what the assistant knows rather than answering a direct question.
Writing backend instructions
instructions are added to the assistant’s own instructions for the backend model only. Put the business rules, lookup procedures and tool guidance here — the frontend model does not need them, and on GPT-Live it has a small context window that is better spent on holding a natural conversation.
Choosing a mode
telnyx — Telnyx runs the backend
The default. Telnyx runs the backend model with the assistant’s own tools, MCP servers and observability, so a delegation can do anything the assistant could do.
Validation happens when you save the assistant: an unavailable model, or an llm_api_key_ref that does not resolve, is rejected there rather than surfacing mid-call as an assistant that talks but can never look anything up.
On GPT-Live, an OpenAI backend model is handed to OpenAI’s own delegation at session start and runs there, calling back into Telnyx for the assistant’s tools. MCP servers are not available on that path.
client — you answer the delegation
Set mode: "client" and Telnyx relays each delegation to your own server over the WebSocket configured in websocket_settings. Use this when the answer has to come from a system that cannot be reached as a Telnyx tool — an internal service behind your own auth, a model you host, a human in the loop.
Telnyx sends:
id:
client mode:
- It requires an active event-stream socket. If none is connected when a delegation is raised, the delegation is refused: the assistant tells the caller it cannot look things up right now and carries on from what it already knows. It never leaves the caller waiting in silence.
- The answer is text only. The event stream offers your server no tool vocabulary, so
outputis plain text that gets injected as context. - An empty answer is a failure, not a quiet success. The assistant has already told the caller it is checking, so an
outputthat is blank after trimming is rejected rather than producing dead air. requestisnullon GPT-Live. The live model raises a delegation with no text of its own, so you work from the conversation events the socket is already streaming you.- Answer promptly. A delegation that is not answered in time falls back the same way an unavailable socket does.
Model requirements
GPT-Live conversations are selected by the assistant’smodel and voice:
The two have to be paired. An assistant with a GPT-Live model and a non-live voice — or the reverse — is rejected when you save it, rather than failing when a call arrives. Conversation flows are not supported on the GPT-Live route.
The backend model in
delegation_settings.model has no such restriction: it is an ordinary text model, and it is where tool use actually happens.
Next steps
- Conversation event stream — the WebSocket that
mode: "client"delegations travel over - Tools library — the tools a
telnyx-mode backend can call - Custom LLM — running a model on your own endpoint with
external_llm