Skip to main content
Configure websocket_settings on an assistant and Telnyx opens a WebSocket to a server you host, once per conversation, streaming transcript, response, telephony and delegation events to it for the life of that conversation. Your server can push messages back to inject a turn. The assistant keeps the brain. This socket observes and injects — it does not replace the model, and it is not in the call path. If you want your own server to replace the LLM entirely, use ConversationRelay instead.
Beta. The set of events carried on this stream is expected to grow. New event types can appear at any time, so ignore types you do not recognize rather than treating them as errors.

What you can build with it

Configure the assistant

Set websocket_settings when you create or update an assistant:
The url must be externally reachable. Localhost, private IP ranges and .local domains are rejected when you save the assistant, and an auth_ref that does not resolve to a secret in your organization is rejected too — so a typo fails at save time rather than silently producing a socket that never authenticates. Use wss:// in production. ws:// is accepted, but it sends your bearer token and your customers’ transcripts in the clear.
Telnyx resolves auth_ref on every connection attempt, so rotating the secret takes effect on the next reconnect without touching the assistant.

How a session flows

Frames are bare JSON objects discriminated by type, in the shape the OpenAI Realtime vocabulary uses — there is no envelope to unwrap. Telephony events are namespaced telnyx. so they can never collide with an event name a future OpenAI release introduces. session.created is always the first frame, and it is not sent until the conversation is ready. Anything you send before it is refused with an error frame, so wait for it before writing.

Events Telnyx sends

Only events that mean something to someone monitoring a conversation are forwarded — media-plumbing events like call.playback.started are deliberately left off the stream. For every payload, field by field, see the full reference under Assistants API → Assistant Event Stream in the sidebar.

Rebuilding the assistant’s reply

A turn arrives as a response.created, then one response.output_item.added and response.content_part.added opening the text, then a run of response.text.delta frames. Concatenate the delta values sharing the same item_id and content_index:
The conversation.item.created frame for the same turn carries the finished text, so you can either stream the deltas for responsiveness or just wait for the completed item.

Injecting messages

After session.created, send a conversation.item.create frame to put a message into the conversation:
A user item triggers an answer exactly as a spoken turn does. An assistant item is recorded in the history without provoking a reply — useful for injecting context the model should know about but not respond to. There is no response.create verb: the assistant owns turn-taking, so a user item is the way to ask for a response. Only text content parts are read (input_text, output_text, text), and an item whose text is empty after trimming is rejected.

Delivery guarantees

This is a side channel, and it is built so that a customer endpoint can never affect a live call. That shapes what you can rely on:
  • Events are dropped, not queued, while the socket is down. Nothing is replayed on reconnect. Queueing is what turns a dead endpoint into unbounded memory on a live call, and a replayed backlog is of little use to a monitoring client anyway. If your endpoint goes away mid-conversation, you have a permanent hole in your view of it.
  • No socket failure reaches the call. A refused connection, a wedged endpoint, a crash — all of it costs you the side channel for the rest of the conversation and nothing more. The caller notices nothing.
  • Telnyx reconnects with exponential backoff, from 1 second up to 30 seconds. A connection has to survive 10 seconds before it counts as stable and resets the backoff, so an endpoint that accepts the upgrade and then closes does not become a reconnect loop.
  • Some failures are terminal for the conversation. Telnyx stops reconnecting after three consecutive failed sends, ten consecutive invalid inbound frames, an oversized frame, or an endpoint that resolves to an address Telnyx refuses to talk to.
Because of the first point, treat this stream as a live view, not a system of record. If you need a guaranteed-complete transcript, read it from the Conversations API (under Assistants API → Conversations API in the sidebar) after the conversation ends.

Limits

Rate-limited and refused frames are answered and discarded without counting toward the invalid-frame limit, so a burst from a legitimate client does not cost it the connection.

Handoffs and private consults

The socket stays bound to the assistant that opened the conversation. It is not re-dialled against the new assistant’s websocket_settings on an agent handoff — re-dialling would drop events across the switch — so a telnyx.assistant.handoff event is how you learn the conversation changed hands. A handoff that cascades through several assistants is reported once, naming the assistant it landed on. During the private consult phase of a warm transfer the socket goes silent in both directions. The consult is between the assistant and the transfer target, and its transcript is not part of the conversation you are monitoring. Frames you send during a consult are refused with a deliberately non-specific error.

A minimal server

Next steps

  • Assistant Event Stream reference — every frame, field by field, under Assistants API → Assistant Event Stream in the sidebar
  • Delegation — answer the assistant’s lookups from your own backend over this socket
  • Agent handoff — what telnyx.assistant.handoff is reporting