> ## Documentation Index
> Fetch the complete documentation index at: https://developers.telnyx.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Stream conversation events to your WebSocket

> Point an AI assistant at a WebSocket server you host and receive live transcript, response, telephony and delegation events — and inject messages back into the conversation.

Configure `websocket_settings` on an assistant and Telnyx opens a WebSocket to a server you host, once per conversation, streaming transcript, response, telephony and delegation events to it for the life of that conversation. Your server can push messages back to inject a turn.

The assistant keeps the brain. This socket observes and injects — it does not replace the model, and it is not in the call path. If you want your own server to replace the LLM entirely, use [ConversationRelay](/docs/voice/programmable-voice/conversation-relay) instead.

<Note>
  **Beta.** The set of events carried on this stream is expected to grow. New event types can appear at any time, so ignore types you do not recognize rather than treating them as errors.
</Note>

## What you can build with it

| Use case                                                      | How                                                                               |
| ------------------------------------------------------------- | --------------------------------------------------------------------------------- |
| Live agent-assist screen showing the transcript as it happens | Render `conversation.item.created` and `response.text.delta`                      |
| Supervisor dashboard tracking calls in flight                 | Track `session.created` / `session.ended` and the `telnyx.call.*` events          |
| CRM logging a conversation as it unfolds, not after           | Persist conversation items as they arrive                                         |
| Nudge the assistant mid-call from your own system             | Send `conversation.item.create` with a `user` or `assistant` item                 |
| Answer the assistant's lookups from your own backend          | Pair with [delegation](/docs/inference/ai-assistants/delegation) in `client` mode |

## Configure the assistant

Set `websocket_settings` when you create or update an assistant:

```bash theme={null}
curl -X POST https://api.telnyx.com/v2/ai/assistants/{assistant_id} \
  -H "Authorization: Bearer $TELNYX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "websocket_settings": {
      "enabled": true,
      "url": "wss://events.example.com/telnyx",
      "auth_ref": "my_websocket_token"
    }
  }'
```

| Field      | Description                                                                                                                                                       |
| ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `enabled`  | Whether Telnyx opens a socket for each conversation. Defaults to `false`.                                                                                         |
| `url`      | The `ws://` or `wss://` endpoint Telnyx connects to. Required when `enabled` is `true`.                                                                           |
| `auth_ref` | Optional. An [integration secret](/docs/inference/ai-assistants/integrations) whose value Telnyx sends as `Authorization: Bearer <value>` on the upgrade request. |

The `url` must be externally reachable. Localhost, private IP ranges and `.local` domains are rejected when you save the assistant, and an `auth_ref` that does not resolve to a secret in your organization is rejected too — so a typo fails at save time rather than silently producing a socket that never authenticates.

Use `wss://` in production. `ws://` is accepted, but it sends your bearer token and your customers' transcripts in the clear.

<Note>
  Telnyx resolves `auth_ref` on every connection attempt, so rotating the secret takes effect on the next reconnect without touching the assistant.
</Note>

## How a session flows

```
Telnyx → You      (connect wss://events.example.com/telnyx, Authorization: Bearer ...)
Telnyx → You      {"type":"session.created","session":{...}}
Telnyx → You      {"type":"telnyx.call.answered","call_control_id":"..."}
Telnyx → You      {"type":"conversation.item.created","item":{"role":"user", ...}}
Telnyx → You      {"type":"response.created","response":{"id":"resp_9c1f4a2b", ...}}
Telnyx → You      {"type":"response.text.delta","delta":"We are open ", ...}
You    → Telnyx   {"type":"conversation.item.create","item":{"role":"user", ...}}
Telnyx → You      {"type":"telnyx.call.hangup","cause":"normal_clearing", ...}
Telnyx → You      {"type":"session.ended","reason":"normal","duration_sec":84, ...}
```

Frames are bare JSON objects discriminated by `type`, in the shape the OpenAI Realtime vocabulary uses — there is no envelope to unwrap. Telephony events are namespaced `telnyx.` so they can never collide with an event name a future OpenAI release introduces.

`session.created` is always the first frame, and it is not sent until the conversation is ready. Anything you send before it is refused with an `error` frame, so wait for it before writing.

## Events Telnyx sends

| Event                                         | Meaning                                                                                                           |
| --------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- |
| `session.created`                             | The conversation is ready. Carries `conversation_id`, `assistant_id` and the call identifiers.                    |
| `session.ended`                               | The conversation finished. Carries `reason`, `duration_sec` and `transfer_status`.                                |
| `conversation.item.created`                   | A message was added to history — a caller utterance, an assistant reply, a tool call, or a tool result.           |
| `conversation.item.deleted`                   | An item was removed, for example an assistant turn discarded after a barge-in.                                    |
| `conversation.participant.added` / `.removed` | A party joined or left a [multi-participant conversation](/docs/inference/ai-assistants/multi-participant-calls). |
| `response.created`                            | The assistant started a turn. `response.id` correlates the streaming frames that follow.                          |
| `response.output_item.added`                  | An output item was opened on that turn.                                                                           |
| `response.content_part.added`                 | A content part was opened on the item.                                                                            |
| `response.text.delta`                         | An incremental chunk of the assistant's reply.                                                                    |
| `telnyx.call.answered`                        | The call was answered.                                                                                            |
| `telnyx.call.hangup`                          | The call was hung up, with the SIP `cause`.                                                                       |
| `telnyx.call.dtmf.received`                   | The caller pressed a digit.                                                                                       |
| `telnyx.call.machine.detection.ended`         | Answering-machine detection finished, with the `result`.                                                          |
| `telnyx.call.transfer.completed` / `.failed`  | A transfer out of the conversation succeeded or failed.                                                           |
| `telnyx.assistant.handoff`                    | The conversation was handed to a different assistant.                                                             |
| `session.delegation.created`                  | The assistant is asking you to answer a [delegation](/docs/inference/ai-assistants/delegation).                   |
| `error`                                       | A frame you sent was invalid or cannot be accepted right now.                                                     |

Only events that mean something to someone monitoring a conversation are forwarded — media-plumbing events like `call.playback.started` are deliberately left off the stream.

For every payload, field by field, see the full reference under **Assistants API → Assistant Event Stream** in the sidebar.

### Rebuilding the assistant's reply

A turn arrives as a `response.created`, then one `response.output_item.added` and `response.content_part.added` opening the text, then a run of `response.text.delta` frames. Concatenate the `delta` values sharing the same `item_id` and `content_index`:

```javascript theme={null}
const buffers = new Map();

function onEvent(event) {
  if (event.type === "response.text.delta") {
    const key = `${event.item_id}:${event.content_index}`;
    buffers.set(key, (buffers.get(key) ?? "") + event.delta);
  }
  if (event.type === "conversation.item.created" && event.item.role === "assistant") {
    // The completed item carries the final text — use it as the source of truth.
    console.log(event.item.content[0]?.text);
  }
}
```

The `conversation.item.created` frame for the same turn carries the finished text, so you can either stream the deltas for responsiveness or just wait for the completed item.

## Injecting messages

After `session.created`, send a `conversation.item.create` frame to put a message into the conversation:

```json theme={null}
{
  "type": "conversation.item.create",
  "item": {
    "type": "message",
    "role": "user",
    "content": [{ "type": "input_text", "text": "Actually, make that a delivery instead." }]
  }
}
```

A `user` item triggers an answer exactly as a spoken turn does. An `assistant` item is recorded in the history without provoking a reply — useful for injecting context the model should know about but not respond to.

There is no `response.create` verb: the assistant owns turn-taking, so a user item is the way to ask for a response. Only text content parts are read (`input_text`, `output_text`, `text`), and an item whose text is empty after trimming is rejected.

## Delivery guarantees

This is a side channel, and it is built so that a customer endpoint can never affect a live call. That shapes what you can rely on:

* **Events are dropped, not queued, while the socket is down.** Nothing is replayed on reconnect. Queueing is what turns a dead endpoint into unbounded memory on a live call, and a replayed backlog is of little use to a monitoring client anyway. If your endpoint goes away mid-conversation, you have a permanent hole in your view of it.
* **No socket failure reaches the call.** A refused connection, a wedged endpoint, a crash — all of it costs you the side channel for the rest of the conversation and nothing more. The caller notices nothing.
* **Telnyx reconnects with exponential backoff**, from 1 second up to 30 seconds. A connection has to survive 10 seconds before it counts as stable and resets the backoff, so an endpoint that accepts the upgrade and then closes does not become a reconnect loop.
* **Some failures are terminal for the conversation.** Telnyx stops reconnecting after three consecutive failed sends, ten consecutive invalid inbound frames, an oversized frame, or an endpoint that resolves to an address Telnyx refuses to talk to.

Because of the first point, treat this stream as a live view, not a system of record. If you need a guaranteed-complete transcript, read it from the Conversations API (under **Assistants API → Conversations API** in the sidebar) after the conversation ends.

## Limits

| Limit                      | Value                | On breach                              |
| -------------------------- | -------------------- | -------------------------------------- |
| Inbound frame size         | 1 MiB                | `error`, then the connection is closed |
| Inbound frame rate         | 10 frames per second | `error`; the frame is discarded        |
| Consecutive invalid frames | 10                   | the connection is closed               |
| Binary frames              | not supported        | `error`; the frame is discarded        |

Rate-limited and refused frames are answered and discarded without counting toward the invalid-frame limit, so a burst from a legitimate client does not cost it the connection.

## Handoffs and private consults

The socket stays bound to the assistant that opened the conversation. It is **not** re-dialled against the new assistant's `websocket_settings` on an [agent handoff](/docs/inference/ai-assistants/agent-handoff) — re-dialling would drop events across the switch — so a `telnyx.assistant.handoff` event is how you learn the conversation changed hands. A handoff that cascades through several assistants is reported once, naming the assistant it landed on.

During the private consult phase of a [warm transfer](/docs/inference/ai-assistants/warm-transfer-acceptance) the socket goes silent in both directions. The consult is between the assistant and the transfer target, and its transcript is not part of the conversation you are monitoring. Frames you send during a consult are refused with a deliberately non-specific error.

## A minimal server

```javascript theme={null}
import { WebSocketServer } from "ws";

const wss = new WebSocketServer({ port: 8080 });

wss.on("connection", (socket, request) => {
  // Verify the credential Telnyx sends from your `auth_ref` secret.
  if (request.headers.authorization !== `Bearer ${process.env.TELNYX_WS_TOKEN}`) {
    socket.close(1008, "unauthorized");
    return;
  }

  let conversationId = null;

  socket.on("message", (data) => {
    const event = JSON.parse(data.toString());

    switch (event.type) {
      case "session.created":
        conversationId = event.session.conversation_id;
        console.log(`conversation ${conversationId} started`);
        break;

      case "conversation.item.created":
        console.log(`[${event.item.role}]`, event.item.content[0]?.text);
        break;

      case "session.ended":
        console.log(`ended: ${event.reason} after ${event.duration_sec}s`);
        break;

      default:
        // New event types can appear at any time — ignore what you don't know.
        break;
    }
  });
});
```

## Next steps

* **Assistant Event Stream reference** — every frame, field by field, under **Assistants API → Assistant Event Stream** in the sidebar
* [Delegation](/docs/inference/ai-assistants/delegation) — answer the assistant's lookups from your own backend over this socket
* [Agent handoff](/docs/inference/ai-assistants/agent-handoff) — what `telnyx.assistant.handoff` is reporting
