Skip to main content
Use GPT Live as the speaking model in a Telnyx AI Assistant. GPT Live receives speech and generates speech in one model, so its voice latency is a single speech-to-speech estimate rather than the sum of separate transcription, language model, and speech synthesis stages. For lookups and actions, keep delegation enabled. GPT Live handles the conversation; a backend model or a server you host performs the work and returns the result.
This guide configures a Telnyx AI Assistant through the Assistants API. For a SIP connection that routes calls directly to an OpenAI session controlled by your own sideband application, use Connect Telnyx to GPT-Live over SIP. These are separate integrations; the direct SIP setup does not configure an AI Assistant’s delegation_settings.

Configure the speaking model

Use a GPT Live model available in the model selector for AI Assistants in the Portal. Store an OpenAI API key with access to that model as an integration secret. Set these fields on the assistant: The model and voice must be paired: a GPT Live model with a non-live voice, or an OpenAILive. voice with a non-live model, is rejected when the assistant is saved. Configure GPT Live through the top-level model; external_llm selects an external text-model endpoint and takes precedence over it. Send this body to Create an assistant, POST /v2/ai/assistants. Replace openai-live-key with the integration secret identifier and confirm that the model is available to your account:
The example explicitly selects moonshotai/Kimi-K2.6 as the delegation backend. Confirm that this model is available to your account, or replace it with another supported backend model. It does not attach an order system: attach a shared lookup tool using the tools library before testing order lookups. The top-level llm_api_key_ref authenticates the speaking model. If the delegation backend needs its own provider key, configure delegation_settings.llm_api_key_ref or delegation_settings.external_llm.llm_api_key_ref separately. Use secret references; delegation does not accept raw API keys.

Configure delegation for work

GPT Live cannot call the assistant’s tools directly. When it needs a lookup or action, it raises a delegation and waits for the result. The result can stream back sentence by sentence for GPT Live to speak. Disabling delegation leaves a conversational assistant that cannot perform those lookups or actions. Choose the backend according to where the work runs: The delegation guide owns the full settings reference, backend selection, result handling, and event payloads. For GPT Live, account for these differences:
  • Backend instructions: put detailed business rules, lookup procedures, and tool guidance in delegation_settings.instructions. These supplement the assistant’s instructions for the backend; keep the speaking model’s conversational instructions concise.
  • OpenAI backend: with telnyx mode and an OpenAI backend model, delegation runs through OpenAI and calls back into Telnyx for the assistant’s tools. MCP servers are unavailable on that path.
  • Client requests: session.delegation.created has a null request on GPT Live. Build the request context from the conversation events already received on the socket. Return a non-empty text result with the matching delegation ID.
  • Spoken results: leave speak_results: true for an answer the caller is waiting for. Use false when the result should become silent context for later answers.

Connect and verify the assistant

Use the assistant ID with the existing voice assistant calling setup or realtime voice conversation WebSocket. A client delegation backend uses the separate conversation event stream: Telnyx connects to your server through websocket_settings. Verify both conversation and delegated work:
  1. Confirm the assistant speaks its configured greeting, then ask a conversational question and verify the response uses its selected GPT Live voice.
  2. Ask for information that requires a configured tool. Confirm the backend performs the lookup and the assistant speaks the returned result without inventing data.
  3. In client mode, verify a delegation with request: null is answered from the streamed conversation context. Test an unavailable backend as well as a successful result.

Next steps

  • Delegation — backend settings, spoken results, and the client event contract
  • Conversation event stream — connect a server that receives events and answers client delegations
  • Tools library — attach tools for delegated lookups and actions