# Telnyx AI — Full Documentation > Inference, embedding, AI assistants, missions, and LiveKit integrations. Complete page content for the AI section of the Telnyx developer docs (https://developers.telnyx.com). > This file: https://developers.telnyx.com/docs/development/llms/ai-llms-full-txt · Root index: https://developers.telnyx.com/llms.txt ## Subsections Focused per-subsection full-content files: - [Inference](https://developers.telnyx.com/docs/development/llms/ai-inference-llms-full-txt) - [AI Gateway](https://developers.telnyx.com/docs/development/llms/ai-ai-gateway-llms-full-txt) - [Embedding](https://developers.telnyx.com/docs/development/llms/ai-embedding-llms-full-txt) - [Assistants](https://developers.telnyx.com/docs/development/llms/ai-assistants-llms-full-txt) - [LiveKit on Telnyx (Beta)](https://developers.telnyx.com/docs/development/llms/ai-livekit-llms-full-txt) ## Inference ### Models > Source: https://developers.telnyx.com/docs/inference/models.md Open-weight LLMs hosted on Telnyx GPU infrastructure. All chat models are accessible via the [Chat Completions API](/api-reference/openai-chat/create-a-chat-completion-openai-compatible) (OpenAI-compatible); embedding models are served on the [embeddings endpoints](/docs/inference/embeddings). For supported Flex, Default, and Priority options, see [Service tiers](/docs/inference/service-tiers). Not every model here is available to AI Assistants, and of the assistant models only a subset is verified for voice calls — see [Voice AI models](/docs/voice/conversational-ai/quickstart#voice-verified-models). ## Chat Models | Model ID | Parameters | Context Length | Best For | Voice AI | |----------|:----------:|:--------------:|----------|:--------:| | `zai-org/GLM-5.3-Flash` | 320B | 1M | Efficient multimodal coding and agentic workflows **(Recommended)** | Not available for assistants | | `zai-org/GLM-5.3` | 753.9B | 1M | Complex coding and long-horizon agent tasks | Not available for assistants | | `deepseek-ai/DeepSeek-V4.1-Flash` | 552B backbone | 1M | Low-latency multimodal inference | Not available for assistants | | `moonshotai/Kimi-K3` | 2.8T | 1M | State-of-the-art open-weight intelligence for coding, reasoning, and multimodal work | Not available for assistants | | `moonshotai/Kimi-K2.6` | 1.0T | 256K | Voice AI | Verified — default | | `zai-org/GLM-5.2` | 753.9B | 1M | Coding, reasoning, 1M context window | Verified | | `MiniMaxAI/MiniMax-M3-MXFP8` | 428B | 1M | Cheapest while maintaining high intelligence | Not available for assistants | | `deepseek-ai/DeepSeek-V4-Flash-0731` | 304B | 1M | Advanced coding, tool use, and long-horizon agentic workflows | Not available for assistants | | `Qwen/Qwen3.8-27B` | 27B | 256K | Multimodal coding and visual understanding across documents, diagrams, and video | Not supported | Reasoning (thinking) behavior differs by surface: reasoning models return their chain-of-thought in a separate `reasoning_content` field on chat completions (see [getting started](/docs/inference/getting-started)), but reasoning is always disabled on voice calls and cannot be enabled there — see [Reasoning on voice calls](/docs/voice/conversational-ai/quickstart#reasoning-on-voice-calls). ## Embedding Models | Model ID | Dimensions | Context Length | Best For | |----------|:----------:|:--------------:|----------| | `Qwen/Qwen3-Embedding-8B` | 4096 | 8,192 tokens | Multilingual and code retrieval; supports custom output dimensions via the `dimensions` parameter | | `thenlper/gte-large` | 1024 | 512 tokens | Text embeddings | | `intfloat/multilingual-e5-large` | 1024 | 512 tokens | Multilingual text embeddings | All three models are available on the OpenAI-compatible [Create embeddings](/api-reference/openai-embeddings/create-embeddings) endpoint (`POST /v2/ai/openai/embeddings`); [List embedding models](/api-reference/openai-embeddings/list-embedding-models) returns the current list programmatically. When you embed a Telnyx Storage bucket for AI use — via the [Embed documents](/api-reference/embeddings/embed-documents) endpoint or the portal's **Embed for AI Use** button (see [Embeddings](/docs/inference/embeddings)) — the `embedding_model` field accepts `thenlper/gte-large` and `intfloat/multilingual-e5-large` only, and defaults to `intfloat/multilingual-e5-large`. `Qwen/Qwen3-Embedding-8B` is available on the OpenAI-compatible embeddings endpoint only. Inputs longer than the model's context length are rejected with a `400` error on the OpenAI-compatible endpoint. --- ### Decision Models > Source: https://developers.telnyx.com/docs/inference/decision-models.md Decision Models is in **beta**. Choose between Flash and Pro using the `model` field. `POST https://api.telnyx.com/v2/ai/typesafe/v1/systemone` evaluates shared context against named questions and returns structured answers. Use `choice` to select a category, `noul` to evaluate a yes/no condition, and `score` to rate an ordered rubric. A request can combine all three question types. The endpoint supports a subset of the [TypeSafe System One API](https://docs.typesafe.ai/concepts/system-one) request format and preserves its typed answer shapes. It returns one complete JSON response. See the [API reference](/api-reference/decision-models/evaluate-decision-models-typesafe-compatible) for the full request and response schemas. ## Choose a model | Model | Use case | | --- | --- | | `telnyx/decision-flash` | Lowest cost and latency for high-volume decisions. | | `telnyx/decision-pro` | Decisions that need long context, including inputs beyond Jev’s 32k per-decision limit. | Set `model` once per request; every question uses that model. If omitted, it defaults to `telnyx/decision-flash`. Other model values are rejected. Telnyx manages the underlying models behind these public aliases. Compare both aliases using the same state and questions on representative examples. Measure answer quality and response time for the application; confidence scores alone do not establish which model is more accurate. ## Use Pro for decisions over long context Choose `telnyx/decision-pro` when the decision depends on more context than Jev can accept in one question. Examples include routing a case using its complete support history, checking an exception against a policy and its amendments, or rating an incident from a long transcript and operational records. [Jev 1.13 documents a 32k-token limit for state plus the longest question](https://docs.typesafe.ai/models), with a separate 64k limit for state plus all questions. Pro’s larger context capacity lets an application keep relevant evidence together when a decision would otherwise require splitting or summarizing the input to fit Jev. Include the question, criteria, and formatting when sizing requests. More context does not guarantee a more accurate answer: keep the evidence relevant and test against labeled examples from the application. Use Flash when cost and latency matter most and the decision fits its supported context. ## Classify a support incident Set `TELNYX_API_KEY` to a Telnyx API key. Send the incident as `state`, then use named questions to select the team, identify a production incident, and rate urgency in one request. ```bash curl --fail-with-body --max-time 100 \ 'https://api.telnyx.com/v2/ai/typesafe/v1/systemone' \ -H "Authorization: Bearer ${TELNYX_API_KEY}" \ -H 'Content-Type: application/json' \ --data '{ "model": "telnyx/decision-flash", "state": "Our production calls are failing. Every customer is affected.", "questions": { "team": { "type": "choice", "instructions": "Choose the team that should handle this incident.", "criteria": { "billing": "Payments and refunds", "technical_support": "Service faults and technical problems", "sales": "New purchases" } }, "production_incident": { "type": "noul", "instructions": "Does the message describe an active production incident?" }, "urgency": { "type": "score", "instructions": "Rate operational urgency.", "criteria": ["Low", "Normal", "High", "Critical"] } } }' ``` The response contains `model`, `answers`, and `usage` directly, with no `data` wrapper. The `model` value identifies the public alias used for the request. Each key in `answers` matches a key in `questions`. The following example rounds values for readability; results and token counts can vary. ```json { "model": "telnyx/decision-flash", "answers": { "team": { "type": "choice", "choice": "technical_support", "probabilities": { "billing": 0.002472, "technical_support": 0.997267, "sales": 0.000261 }, "confidence": 0.982052 }, "production_incident": { "type": "noul", "noul": 0.999196 }, "urgency": { "type": "score", "score": 2.997424, "legend": {"0": "Low", "1": "Normal", "2": "High", "3": "Critical"}, "probabilities": {"0": 0.000335, "1": 0.000035, "2": 0.001501, "3": 0.998129}, "confidence": 0.989420 } }, "usage": {"input_tokens": 267, "output_tokens": 4} } ``` Read `answers.team.choice` to choose a destination. Interpret `answers.production_incident.noul` as a numeric yes-score, and `answers.urgency.score` on the requested 0–3 rubric. Token usage includes shared-context preparation and question evaluation, so it can exceed the size of the unique input text. ## Choose question types Every question requires `type` and `instructions`. Instructions and `state` can be strings, JSON objects, or arrays. Text-only conversation histories are supported; image and audio inputs are not supported. | Type | Criteria | Answer | | --- | --- | --- | | choice | Object with 2–64 option keys. Values are strings or null; null uses the key as the option description. | choice contains the winning key, probabilities contains all option scores, and confidence describes concentration. | | noul | Optional object with string descriptions for true and false. Defaults to {"true":"Yes","false":"No"}. | noul is the positive outcome's score from 0 to 1. There is no separate confidence field. | | score | Ordered array of 2–64 description strings. | score is the expected zero-based index; legend and probabilities use stringified indices. | A `choice` question selects one option. Use separate `noul` questions when several independent conditions can be true at once. For example, a support message can both request a refund and report a service fault. A `score` answer can be fractional. For criteria `["Low", "Normal", "High", "Critical"]`, the range is 0–3. Compute the expected score as `sum(index * probability)`; do not treat it as a 0–1 probability or an array index without an application-specific decision rule. ## Use scores in application logic Option scores are normalized relative preferences across the supplied choices. Changing the choices or their wording can change the distribution. They are not calibrated probabilities that a decision is correct. For `choice` and `score`, `confidence` is normalized entropy: `1 - H(p) / ln(N)`, where `H(p) = -sum(p * ln(p))` and `N` is the number of options. It approaches 0 for a uniform distribution and 1 when the score is concentrated on one option. It is different from the winning option's probability. Choose review thresholds using representative examples from the application. The following Python example uses direct HTTP, selects a team, and falls back to manual review when the winning option has a low relative score. The `0.8` threshold is illustrative and must be evaluated for the application's data. ```python import json import os from urllib.error import HTTPError from urllib.request import Request, urlopen payload = { "model": "telnyx/decision-pro", "state": "The invoice looks right, but the usage page counts calls twice.", "questions": { "team": { "type": "choice", "instructions": "Choose the team for the underlying issue.", "criteria": { "billing": "Incorrect charges or refunds", "technical_support": "Service faults or software defects", "sales": "New purchases", }, } }, } request = Request( "https://api.telnyx.com/v2/ai/typesafe/v1/systemone", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": f"Bearer {os.environ['TELNYX_API_KEY']}", "Content-Type": "application/json", }, method="POST", ) try: with urlopen(request, timeout=100) as response: result = json.load(response) except HTTPError as error: raise RuntimeError(f"Decision model failed with HTTP {error.code}") from error answer = result["answers"]["team"] selected = answer["choice"] winning_score = answer["probabilities"][selected] destination = selected if winning_score >= 0.8 else "manual_review" print(destination) ``` ## Migrate a TypeSafe request Point HTTP requests to `https://api.telnyx.com/v2/ai/typesafe/v1/systemone` and authenticate with a Telnyx Bearer API key. Set `model` to a supported Telnyx alias and send `state` and `questions`, keeping the question IDs used by the application. Compatibility applies to the supported JSON request subset and typed answer shapes. It does not imply identical model predictions, confidence calibration, pricing, or token accounting. | Area | Telnyx behavior | | --- | --- | | Route | `POST /v2/ai/typesafe/v1/systemone`. | | Model selection | Set `model` to `telnyx/decision-flash` or `telnyx/decision-pro`. Omitted values default to Flash; unsupported values are rejected. The response returns the selected alias. | | Question instructions | Required for every question. TypeSafe also accepts omitted or null instructions; this subset does not. | | Criterion descriptions | Strings for all question types; `choice` also accepts `null`. TypeSafe's object and array descriptions are not supported here. | | Question and option limits | 1–64 questions; 2–64 options for each `choice` or `score`. TypeSafe's one-level score rubric is not supported. | | Answers | Named `choice`, `noul`, and `score` answers, with `model` and token `usage` at the top level. | | SDK base URL | Set `base_url="https://api.telnyx.com/v2/ai/typesafe"`. The SDK appends `/v1/systemone`; no route adapter is required. | The [official TypeSafe Python SDK](https://github.com/typesafe-ai/typesafe-sdk-python) appends `/v1/systemone` to the configured base URL. The route preserves that behavior. Its `client.models.list()` method uses a separate TypeSafe route and is outside this endpoint's compatibility scope. The SDK automatically sends a model value, so changing only the base URL is not sufficient: replace its default or existing model with a supported Telnyx alias. ## Use the TypeSafe Python SDK Install the [official SDK](https://github.com/typesafe-ai/typesafe-sdk-python): ```bash pip install typesafe-sdk ``` Set the Telnyx base URL, API key, and model alias. Existing `system_one()` calls using the supported question subset keep the same method and answer accessors. This example uses the SDK's `Choice`, `Noul`, and `Score` types. ```python import os from typesafe_sdk import Choice, Noul, Score, TypeSafeClient with TypeSafeClient( api_key=os.environ["TELNYX_API_KEY"], base_url="https://api.telnyx.com/v2/ai/typesafe", timeout=100, ) as client: result = client.system_one( model="telnyx/decision-flash", state="Our production calls are failing. Every customer is affected.", questions={ "team": Choice( instructions="Choose the team that should handle this incident.", criteria={ "billing": "Payments and refunds", "technical_support": "Service faults and technical problems", "sales": "New purchases", }, ), "production_incident": Noul( instructions="Does the message describe an active production incident?", ), "urgency": Score( instructions="Rate operational urgency.", criteria=["Low", "Normal", "High", "Critical"], ), }, ) print(result.choices["team"].choice) print(result.nouls["production_incident"].noul) print(result.scores["urgency"].score) ``` The SDK sends a Bearer authorization header from `api_key`. Use a Telnyx key for Telnyx requests. Keep `/v1/systemone` out of `base_url`; the SDK adds it. ## Limits and failures Use up to 64 named questions against one shared `state`. Split larger workloads into separate requests and bound client concurrency. Responses arrive after the complete evaluation; there is no streaming or per-question partial-success envelope. Request-body and token limits also apply, so question count alone does not guarantee that a request fits. The public request fields are `model`, `state`, and `questions`. `model` is optional and defaults to `telnyx/decision-flash`; `state` and `questions` are required. Unsupported model values and unknown fields are rejected. Do not send `input`, `labels`, `tier`, `stream`, `temperature`, or `max_tokens` to this endpoint. | Status | Action | | --- | --- | | `401` | Check the Telnyx API key. | | `413` | Reduce or fix the request body. | | `422` | Correct the schema or split context that exceeds token limits. | | `429`, `529` | Reduce concurrency and retry with bounded exponential backoff and jitter. | | `502`, `503` | Retry transient service failures with bounded backoff. | | `504` | Reduce request size or concurrency before retrying. | Check the HTTP status before parsing an answer. Honor `Retry-After` when present. Correct validation failures before retrying; retries repeat evaluation work. A decision model service error uses this shape, with a message describing the failure: ```json { "error": { "message": "Each question requires text/JSON instructions and 2–64 options" } } ``` --- ### Regions & Availability > Source: https://developers.telnyx.com/docs/inference/models/regions.md GPU infrastructure across five regions on four continents. Telnyx will endeavor to process requests in the region nearest the ingress domain you call, but this is not guaranteed. ## Routing Inference **processing is latency-based, influenced by the ingress domain you call**, not by your account's data locality setting. Telnyx will endeavor to process in the preferred region, but does not guarantee it: | Ingress domain | Preferred region | |----------------|--------| | `api.telnyx.com` | US | | `api.telnyx.eu` | EU | | `api.telnyx.com.au` | APAC | Calling a regional ingress domain (for example, `api.telnyx.eu`) directs requests to the nearest GPU region for that domain under normal conditions. Telnyx does **not guarantee** processing location: during failover or capacity events, requests are processed at the next-lowest-latency region rather than failing. ## Selecting a region per request Default routing is latency-based, as above. To ask for a specific region instead, send a `region` on the request, using the same vocabulary as your account's [Data Locality](/docs/account-setup/data-locality) setting: | `region` | Serving location | Matching ingress domain | |----------|------------------|-------------------------| | `USA` | United States | `api.telnyx.com` | | `EU` | Europe | `api.telnyx.eu` | | `AUS` | Australia | `api.telnyx.com.au` | | `UAE` | United Arab Emirates | `api.telnyx.me` | `mode` controls how strictly it is applied: ```bash Preferred (default) curl https://api.telnyx.com/v2/ai/chat/completions \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "moonshotai/Kimi-K2.6", "messages": [{"role": "user", "content": "Hello"}], "region": "EU" }' ``` ```bash Strict # Strict pins must be sent to the matching region's domain. curl https://api.telnyx.eu/v2/ai/chat/completions \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "moonshotai/Kimi-K2.6", "messages": [{"role": "user", "content": "Hello"}], "region": "EU", "mode": "strict" }' ``` - **`preferred`** — the default when `region` is set. Telnyx tries that region first and falls back to another when the model cannot be served there, so a request that would otherwise have succeeded still succeeds. - **`strict`** — the request is served from that region or it fails. It is never redirected to another region. A strict request that cannot be served in the region returns **422**, distinguishing the reasons: ```json { "errors": [{ "detail": "The requested model does not exist in region EU." }] } { "errors": [{ "detail": "The requested model is currently unavailable in region EU." }] } ``` A strict request sent to the wrong region's domain also returns 422 — see [`strict` requires the matching domain](#strict-requires-the-matching-domain) below. Other rules: - `mode` without `region` is a **400**, as are unknown values for either. - Supported for **Telnyx-hosted models only**. A request routed to an external provider (for example `openai/*` or `anthropic/*`) never passes through Telnyx model routing, so `region` cannot be enforced for it: `strict` returns a 400 rather than silently serving it from elsewhere, and `preferred` is ignored. - Not every model is deployed in every region. Availability changes as capacity moves; a `strict` request is the way to find out definitively, and `preferred` is the safe default when you want proximity without risking a failure. Both parameters are accepted on the Chat Completions, Responses and Anthropic Messages endpoints. ### `strict` requires the matching domain `region` controls where **inference** runs. The ingress domain you call is a separate hop — it is where the request is received and authenticated before it reaches a model. A strict pin must be **received** in the region it names, not merely served there. Send one to another region's domain and it returns a **422**: ```json { "errors": [{ "detail": "Requests pinned to region EU with mode strict must be received by a deployment in that region; this deployment serves USA." }] } ``` It is refused rather than forwarded because a request that could be forwarded has already entered the platform outside the region, so forwarding it would not keep the promise the pin makes. Call the domain for the same region from the table above: ```bash # Required for strict: domain and region agree curl https://api.telnyx.eu/v2/ai/chat/completions \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "moonshotai/Kimi-K2.6", "messages": [{"role": "user", "content": "Hello"}], "region": "EU", "mode": "strict" }' ``` Matching them also avoids a needless cross-region hop, so it is the lower-latency choice as well. `mode: "preferred"` and requests without a `region` are unaffected by which domain receives them. ## Data Residency Processing location and storage location are controlled separately: - **Processing in transit** is latency-based by default, influenced by the ingress domain you call (see [Routing](#routing) above); Telnyx will endeavor to process in the preferred region, but that alone is not a guaranteed processing location. For a request that must not leave a region, send `region` with `mode: "strict"` (see [Selecting a region per request](#selecting-a-region-per-request)), which fails the request rather than serving it elsewhere. - **Storage at rest** depends on the endpoint. The **chat completions** endpoint does not store request or response data. The **responses** endpoint stores conversations, and that storage is governed by your [Data Locality](/docs/account-setup/data-locality) setting (US, EU, APAC, or Middle East). For a full cross-product breakdown (including Voice AI Assistants), see the [Data Residency & Compliance FAQ](/docs/inference/data-residency). ## Roadmap - Region selection API parameter - Per-region model status and latency metrics - Edge inference for sub-50ms response times --- ### Service tiers > Source: https://developers.telnyx.com/docs/inference/service-tiers.md Service tiers select the serving capacity and pricing for a Telnyx-hosted model. Set `service_tier` on each request and keep the same public `model` ID. ## Choose a tier | Tier | Request value | When to use it | |------|---------------|----------------| | Flex | flex | Cost-sensitive work that can tolerate higher latency and variable availability, such as offline evaluation, document processing, and background summarization | | Default | default | General-purpose inference with standard pricing | | Priority | priority | Interactive applications where lower latency is more important than the lowest token price | For standard Inference API requests, omitting `service_tier` uses `default`. Flex and Priority are available only on select models. A tier does not change the model's capabilities or context window, and selecting Priority does not guarantee a fixed response time. Flex uses the same request APIs as the other tiers. A background workload sends individual requests; selecting Flex does not create an asynchronous batch job. ## Check model support Call [Get available models](/api-reference/openai-chat/get-available-models-openai-compatible) (`GET /v2/ai/openai/models`) and inspect `data[].service_tiers` for the desired model. Send one of those values as `service_tier`. ```bash curl https://api.telnyx.com/v2/ai/openai/models \ -H "Authorization: Bearer $TELNYX_API_KEY" ``` Use the returned public `id` unchanged. Do not append a tier suffix to the model name. Model support can change, so check the catalog before choosing Flex or Priority. The `service_tiers` field is omitted for externally hosted models; their provider determines supported tiers and behavior. ## Set the tier on a request ### Chat Completions After confirming that `deepseek-ai/DeepSeek-V4.1-Flash` lists `flex`, send this body to [Chat Completions](/api-reference/openai-chat/create-a-chat-completion-openai-compatible) (`POST /v2/ai/openai/chat/completions`): ```json { "model": "deepseek-ai/DeepSeek-V4.1-Flash", "service_tier": "flex", "messages": [ { "role": "user", "content": "Summarize this support note in one sentence: The customer could not sign in. Resetting the password restored access." } ] } ``` To use Default or Priority, set `service_tier` to `default` or `priority` and choose a model that advertises that tier. Streaming requests select the tier with the same field. ### Responses The [Responses API](/api-reference/openai-chat/create-an-openai-compatible-response) (`POST /v2/ai/openai/responses`) accepts the same top-level field: ```json { "model": "deepseek-ai/DeepSeek-V4.1-Flash", "service_tier": "flex", "input": "Summarize this support note in one sentence: The customer could not sign in. Resetting the password restored access." } ``` ## Handle unavailable capacity A supported tier can still be temporarily unavailable. For Flex workloads, allow a longer request timeout and use bounded retries with exponential backoff and jitter for transient capacity errors. Honor `Retry-After` when present. If the model does not support the requested tier, select a value advertised by the models endpoint instead of retrying the unchanged request. Do not rely on automatic switching to another tier. An application can explicitly retry on another supported tier if its latency requirements and budget permit; that new request uses the newly selected tier's rates. ## Pricing Input, cached input, and output token rates depend on both the model and service tier. Compare the applicable rates on the [Inference pricing page](https://telnyx.com/pricing/inference-api). See [Pricing and billing units](/docs/inference/models/pricing) for the billing basis of other Inference services. --- ### Spending limits > Source: https://developers.telnyx.com/docs/inference/spending-limits.md Use the Spend Limits API to set an inference budget in USD for the current UTC day, calendar month, or both. Limits apply to the authenticated user's organization, or their own account if they do not belong to an organization. Limits are shared across the organization. Reading or changing them requires the caller's access policy to permit the corresponding Spend Limits operation. An API key without the required permission receives HTTP 403, even when the organization has shared limits. This authorization error is separate from the HTTP 403 returned by inference requests when a spending limit blocks usage. ## Which usage counts The limit is shared across the account's billable inference usage, including LLM requests made by voice assistants. It is not a separate budget for each assistant, model, or API key. | LLM request | Counts toward the inference limit | | --- | --- | | Telnyx-hosted model, called directly or by a voice assistant | Yes | | OpenAI or another third-party model using Telnyx-managed provider credentials | Yes | | Third-party model using the customer's own provider key (BYOK) | No | | Customer-supplied model endpoint | No | BYOK refers to the third-party provider key, not the Telnyx API key used to authenticate a request. Supplying a provider key does not exempt usage of a Telnyx-hosted model. BYOK and customer-endpoint requests remain available when billable inference is blocked. If they fall back to a billable model, the fallback is checked against the account's block state. This limit covers costs recorded under the inference billing product, not every charge associated with a voice call. A voice assistant using BYOK can still incur other Telnyx charges. ## Understand enforcement The daily and monthly limits are independent. Exceeding either limit blocks new billable Chat Completions, Responses, Anthropic Messages, and classification requests. Requests already running finish normally. Voice-assistant LLM requests are subject to the same enforcement. An in-flight model request can finish, but a subsequent request or conversation turn can be refused after the block takes effect. - Spend must be **strictly greater** than the limit to trigger a block. A zero limit blocks at the first cent of spend; it does not disable inference before any spend occurs. - Daily periods reset at 00:00 UTC the next day. Monthly periods reset at 00:00 UTC on the first day of the next month. - Spend can lag actual usage by several minutes. After spend is recorded, a block appears within about two minutes for daily limits or ten minutes for monthly limits. Usage can exceed the configured amount before enforcement takes effect. - Creating or changing a limit checks existing spend immediately. Setting a limit below current spend can block inference at once. ## Monitor spend in the portal Use the [Inference dashboard](https://portal.telnyx.com/#/ai/reports/dashboard?product=inference) to investigate LLM token usage and reported cost by model and use case, including voice-assistant inference. BYOK can produce token activity without a corresponding Telnyx LLM charge. Use [Billing → Spending limits](https://portal.telnyx.com/#/billing/spending-limits) to view account-wide spend against the daily and monthly caps and edit limits across supported products. The Billing limit counters use the current UTC day and calendar month and include the whole inference billing product. The dashboard charts show LLM records and honor the selected dates, timezone, models, and use cases. Their totals can differ from the limit counters. Use `spend_usd` and `blocked` from `GET /v2/spend_limits`, also shown in the limit controls, to inspect the spend and block state used for enforcement. Reporting and enforcement can lag usage. ## Read current limits and spend Set `TELNYX_API_KEY` to an API key for the account being managed. Call [List spend limits](/api-reference/spend-limits/list-spend-limits): ```bash curl https://api.telnyx.com/v2/spend_limits \ -H "Authorization: Bearer $TELNYX_API_KEY" ``` The unpaginated `data` array includes each supported product and period, even when no limit is set. Select entries with `product: "inference"`; use the returned products rather than assuming future product support. | Field | Interpretation | | --- | --- | | `period` | `daily` or `monthly` | | `period_start`, `period_end` | UTC dates; `period_end` is exclusive | | `limit` | Configured limit, or `null` when none exists | | `effective_limit_usd` | Enforced amount as a decimal string; `null` means unlimited | | `spend_usd` | Current period spend as a decimal string; `null` means unavailable, not zero | | `spend_error` | Set when spend cannot be read | | `blocked`, `block` | Current period's block state and details, including `blocked_until` | A limit with `origin: "operator"` was set by Telnyx support. It can be updated or deleted through the same API. `limit: null` and a configured limit with `unlimited: true` both mean no cap for inference, but only the latter is an existing resource that can be updated. ## Create a limit When the selected entry has `limit: null`, call [Create a spend limit](/api-reference/spend-limits/create-a-spend-limit). This example sets a daily limit of USD 100: ```bash curl -X POST https://api.telnyx.com/v2/spend_limits \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -d '{"product":"inference","period":"daily","amount":100,"reason":"Team budget"}' ``` To also set a monthly limit, send a separate request with this body: ```json { "product": "inference", "period": "monthly", "amount": 2000, "reason": "Monthly inference budget" } ``` Send `amount` as a nonnegative JSON number. Response amounts are decimal strings. The optional audit `reason` accepts up to 500 characters. Send the period explicitly; the default is `daily`. Creation returns HTTP 409 if a limit already exists. Re-read the limits and update the existing resource instead. ## Change an existing limit Call [Update a spend limit](/api-reference/spend-limits/update-a-spend-limit). The product is a path parameter and the period is a query parameter, not a body field. This example raises the daily limit to USD 250: ```bash curl -X PATCH 'https://api.telnyx.com/v2/spend_limits/inference?period=daily' \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -d '{"amount":250,"reason":"Raised for launch traffic"}' ``` To keep an existing monthly resource but explicitly remove its cap, send `{"unlimited":true}` to `PATCH /v2/spend_limits/inference?period=monthly`. Send exactly one of `amount` and `unlimited: true`. Unknown body fields are rejected with HTTP 400. An update returns HTTP 404 if no limit exists; re-read the limits before creating one. ## Remove a limit Call [Delete a spend limit](/api-reference/spend-limits/delete-a-spend-limit) for the intended period: ```bash curl -X DELETE 'https://api.telnyx.com/v2/spend_limits/inference?period=daily' \ -H "Authorization: Bearer $TELNYX_API_KEY" ``` Inference has no default limit. Deletion returns `limit: null`, makes that period unlimited, and lifts its block. It does not remove the other period's limit or block. Deleting a missing limit returns HTTP 404. ## Verify a change and recover blocked requests Create, update, and delete responses include `data.evaluation`: | Field | Meaning | | --- | --- | | `blocked_now` | Existing spend exceeded the new limit and inference was blocked | | `released` | This period's block was lifted | | `still_over_limit` | This period remains blocked because spend is still above the new limit | | `still_blocked_other_period` | The other period still blocks inference | | `evaluation_deferred` | The change was saved, but spend could not be checked; evaluation applies within a few minutes | | `note` | Additional context when present, including exceptions after support lifts a block | Write responses always carry `blocked: false`. Use `evaluation` for the immediate result, then call `GET /v2/spend_limits` again to read the current block state for both periods. Do not treat a successful write or `evaluation.released` alone as confirmation that inference is unblocked. Blocked inference requests return HTTP 403 with error code `10039` and title `Inference spend limit reached`. Stop automatic retries for this error. Inspect both limits, then raise the blocking limits above current spend, remove them, or wait until their UTC periods end. Verify both periods are unblocked before resuming requests. Other HTTP 403 errors can have different causes. --- ### Pricing > Source: https://developers.telnyx.com/docs/inference/models/pricing.md Pay-per-token. No minimums, no commitments. For current per-model pricing, see [telnyx.com/pricing/inference-api](https://telnyx.com/pricing/inference-api). ## Service tiers For standard Inference API requests to Telnyx-hosted models, omitting `service_tier` uses `default`. Set `service_tier` to choose another tier supported by the model. | Service tier | Request value | Pricing and workload | |--------------|---------------|----------------------| | Flex | flex | Lower-cost inference for workloads that can tolerate higher latency and variable availability | | Default | default | Standard rates for general-purpose inference | | Priority | priority | Priority rates for latency-sensitive workloads | Flex and Priority are available for select models. Check the model's `service_tiers` field and use the corresponding tier's rates on the [pricing page](https://telnyx.com/pricing/inference-api). See [Service tiers](/docs/inference/service-tiers) for model availability, request examples, and guidance on choosing a tier. ## Billing units | Category | Basis | Notes | |----------|-------|-------| | Text generation | Per 1M tokens (input + output) | Input and output priced separately; cached input tokens at a discount | | Audio transcription | Per second of audio | Varies by model | | Text-to-speech | Per 1M characters | Varies by voice/model | | Embeddings | Per 1M tokens | Single rate | --- ### Data Residency & Compliance FAQ > Source: https://developers.telnyx.com/docs/inference/data-residency.md This page answers common customer questions about where data is processed and stored, which providers are used, retention, and model training. Telnyx AI spans products that handle **processing location** differently: - **Inference API** — chat completions, the responses endpoint, and related model APIs. - **Voice AI Assistants** — telephony-based conversational agents. - **AI Assistant web chat** — the assistant chat endpoint, which can opt in to a hard in-region processing guarantee. ## Processing vs. storage: the key distinction | | Processing in transit (not guaranteed) | Storage at rest (hard control) | | --- | --- | --- | | **Inference API** | Latency-based by default, influenced by the **ingress domain** you call (`api.telnyx.com`, `api.telnyx.eu`, `api.telnyx.com.au`) and not guaranteed. Opt in per request with **`region` + `mode: "strict"`** to pin inference to one region, or fail. | Chat completions: **not stored**. Responses endpoint (stores conversations): governed by your **data locality** setting (US, EU, APAC, or Middle East). | | **Voice AI Assistants** | Influenced by the **anchorsite** on the assistant's TeXML application. Telnyx will endeavor to honor it, but it is not guaranteed. | Governed by your **data locality** flag, plus the **data-retention** setting for conversation content. | | **AI Assistant web chat** | Latency-based by default. Opt in per assistant with **`privacy_settings.in_transit_data_locality`** to require every model call of a chat turn to be received and served in your region, or fail. | Governed by your **data locality** flag, plus the **data-retention** setting for conversation content. | [Data Locality](/docs/account-setup/data-locality) governs **storage at rest** for covered data types (available regions: US, EU, APAC, and Middle East). Neither the data locality flag nor the anchorsite is a hard guarantee of where live **processing** happens. --- ## Inference API ### Where is Inference processing performed? Inference **processing in transit is latency-based, influenced by the ingress domain you call**, not by your data locality setting. Telnyx will endeavor to process in the preferred region, but does not guarantee it: | Ingress domain | Preferred region | | --- | --- | | `api.telnyx.com` | US | | `api.telnyx.eu` | EU | | `api.telnyx.com.au` | APAC | Calling a regional ingress domain (for example, `api.telnyx.eu`) directs requests to the nearest GPU region for that domain. Telnyx will endeavor to route to that region, but does **not guarantee** the processing location: during failover or capacity events, requests are processed at the next-lowest-latency region rather than failing. See [Inference Regions & Availability](/docs/inference/models/regions) for the underlying GPU regions. To make the region a requirement rather than a preference, send `region` with `mode: "strict"` on the request — see [Can Inference traffic be pinned to a specific region?](#can-inference-traffic-be-pinned-to-a-specific-region) below. ### Does Inference store my data? It depends on the endpoint: - **Chat completions endpoint** — **does not store** request or response data. - **Responses endpoint** — **stores conversations**. For stored data, your [Data Locality](/docs/account-setup/data-locality) setting dictates the storage region. ### Can Inference traffic be pinned to a specific region? Yes, per request, as an opt-in. Send `region` with `mode: "strict"`: ```json { "model": "moonshotai/Kimi-K2.6", "messages": [{ "role": "user", "content": "Hello" }], "region": "EU", "mode": "strict" } ``` `region` uses the same vocabulary as your [Data Locality](/docs/account-setup/data-locality) setting — `USA`, `EU`, `AUS`, `UAE` — and `mode` decides how strictly it applies: - **`mode: "strict"`** — inference is performed in that region or the request **fails with a 422**. It is never redirected to another region, and no fallback (to another model, tier, or site) may take it out of the region. - **`mode: "preferred"`** — the default when `region` is set. That region is tried first, and another is used when the model cannot be served there. Useful for proximity; it is **not** a residency control. Omitting `region` leaves today's latency-based behavior unchanged. `strict` governs **where inference runs**. It is not a blanket "all processing stays in region X" guarantee: request ingress, authentication, logging, billing, and operational/security handling are separate systems with their own footprints, and the exclusion applies only to Telnyx-hosted models. Confirm written commitments with your account team and DPA before making representations to your own customers. Practical notes: - Supported for **Telnyx-hosted models only**. A request routed to an external provider (`openai/*`, `anthropic/*`, or a custom base URL) never passes through Telnyx model routing, so the region cannot be enforced: `strict` returns a **400** rather than silently serving it elsewhere. - Not every model is deployed in every region, and capacity moves. A strict request to a region where the model is not available returns a 422 that distinguishes *"does not exist in region X"* from *"currently unavailable in region X"*. - **`strict` requires the matching ingress domain.** A strict request must be *received* in the region it pins, not merely served there. Send it to another region's domain and it returns a **422** rather than being forwarded — by the time a request could be forwarded, it has already entered the platform outside the region, so forwarding would not keep the promise the pin makes. | `region` | Ingress domain | |----------|----------------| | `USA` | `api.telnyx.com` | | `EU` | `api.telnyx.eu` | | `AUS` | `api.telnyx.com.au` | | `UAE` | `api.telnyx.me` | `mode: "preferred"` and requests without a `region` are unaffected by where they are received. - Available on the Chat Completions, Responses and Anthropic Messages endpoints. See [Regions & Availability](/docs/inference/models/regions#selecting-a-region-per-request). --- ## Voice AI Assistants ### Where is Voice AI Assistant processing performed? For Voice AI Assistants, processing location is **influenced by the anchorsite** configured on the assistant's **TeXML application** — not by the data locality flag. Setting the anchorsite (for example, Frankfurt for the EU) directs media/processing to that region under normal conditions. The anchorsite is **not a hard control**. Telnyx will endeavor to honor it, but does not guarantee processing location: under failover or capacity events, processing can shift to another region rather than failing the call. ### Where is Voice AI Assistant data stored? **Storage location at rest is a hard control, governed by your [Data Locality](/docs/account-setup/data-locality) flag.** Retention of conversation content is further controlled by the **data-retention** setting (see [Data retention](#data-retention-and-model-training)). Recording storage can also be directed to your own storage destination, which Telnyx respects. ### Are call audio, transcripts, prompts, responses, summaries, or recordings ever handled outside the configured region? - **Processing** location is influenced by the assistant's anchorsite; Telnyx will endeavor to honor it, but it is **not guaranteed**. - **Storage at rest** is a hard control, following your **data locality** flag. Recordings can be directed to a customer-controlled storage destination, which Telnyx respects. Telnyx does **not** contractually guarantee blanket "EU-only processing." Telnyx will endeavor to honor processing controls but does not guarantee them. "Processing" is defined very broadly, and some components — for example, third-party STT/TTS providers, or operational/security/fraud handling — may involve activity outside a single region. The specifics depend on the providers and features you enable. Confirm written data commitments with your account team and DPA before making representations to your own customers. ### Example: EU-focused Voice AI setup A typical EU-oriented configuration combines: - **Data locality:** EU (Germany) — a hard control over storage at rest - **Anchorsite on the TeXML app:** an EU site (for example, Frankfurt) — Telnyx will endeavor to influence media/processing location, but it is not guaranteed - **Voice API endpoint:** `api.telnyx.eu` - **SIP endpoint:** `sip.telnyx.eu` This keeps storage in the EU (a hard control via data locality) and steers processing toward the EU (via the anchorsite, which Telnyx will endeavor to honor but does not guarantee). STT/TTS provider choice also matters — some providers are self-hosted by Telnyx and some are third parties (see below). --- ## AI Assistant web chat ### Can assistant chat be restricted to one region? Yes, per assistant, as an opt-in. Set `privacy_settings.in_transit_data_locality` to `true`: ```json { "name": "My assistant", "model": "Qwen/Qwen3-235B-A22B", "instructions": "You are a helpful assistant.", "privacy_settings": { "in_transit_data_locality": true } } ``` With it enabled, every model call made for a chat turn — the response itself, any tool-call follow-up turns, and the conversation-flow conditions that use a model — is required to be **received and served** inside your organization's data-locality region, rather than only stored there. A turn that cannot be handled in region **fails** instead of being served elsewhere. ### What it requires Enabling the setting is rejected with a `400` unless both hold: - **Your data-locality region has in-region inference.** Available for `USA`, `EU`, `AUS` and `UAE` (see [Inference regions](/docs/inference/models/regions)). Other regions have no in-region model capacity, so the guarantee cannot be offered there. - **Every model the assistant could use is Telnyx-hosted.** That means the assistant's `model`, its `fallback_config`, and any conversation-flow node that overrides the model. A third-party model (OpenAI, Anthropic, Google) or an `external_llm` anywhere in that set is rejected, because Telnyx cannot constrain where another vendor runs a model. The error names both the model and where it is configured. ### Send chat requests to your region's hostname Once enabled, a chat request must **enter the platform** in your region. A request that arrives elsewhere is rejected with a `400` naming the hostname to use, rather than being forwarded — forwarding it would already have carried the content across the border. | Data-locality region | API hostname | | --- | --- | | `USA` | `api.telnyx.com` | | `EU` | `api.telnyx.eu` | | `AUS` | `api.telnyx.com.au` | | `UAE` | `api.telnyx.me` | If a model cannot be served in your region at that moment, the turn returns a `422` rather than being served from another region. ### What it covers | | Covered | | --- | --- | | Web chat turns, including tool-call follow-ups and model-driven conversation-flow conditions | Yes | | PII redaction and conversation insights for those conversations | Yes | | Voice assistants | No — see [Voice AI Assistants](#voice-ai-assistants) | | SMS and WhatsApp | No | Messaging is deliberately out of scope rather than blocked: SMS and WhatsApp leave the platform through carriers, whose handling cannot be constrained to a region. An assistant with the setting enabled still serves those channels exactly as it would without it. This setting governs **where model calls for a chat turn are received and served**. It is not a blanket "all processing stays in region X" guarantee: request ingress, authentication, logging, billing, and operational/security handling are separate systems with their own footprints. Confirm written commitments with your account team and DPA before making representations to your own customers. --- ## STT, TTS, and LLM providers (Voice AI) For Voice AI Assistants, the STT, TTS, and LLM providers in use depend on the models and voices **you select**. Some are **self-hosted by Telnyx** (run on Telnyx-operated infrastructure); others are **third-party** services that Telnyx integrates with. This distinction matters for compliance: self-hosted models keep that processing step within Telnyx infrastructure, whereas third-party models route that step to the vendor. Hosting (self-hosted vs. third-party) is about *which infrastructure* performs the step, not a guarantee of *region*. Processing region is not guaranteed for any provider — see the [processing vs. storage](#processing-vs-storage-the-key-distinction) note above. ### Speech-to-text (STT) | Model | Provider | Hosting | | --- | --- | --- | | `deepgram/flux` | Deepgram | Self-hosted by Telnyx | | `deepgram/nova-3` | Deepgram | Self-hosted by Telnyx | | `deepgram/nova-2` | Deepgram | Self-hosted by Telnyx | | `assemblyai/universal-3-5-pro` | AssemblyAI | Self-hosted by Telnyx | | `assemblyai/universal-streaming` (legacy alias of `assemblyai/universal-3-5-pro`) | AssemblyAI | Self-hosted by Telnyx | | `speechmatics/standard` | Speechmatics | Self-hosted by Telnyx | | `distil-whisper/distil-large-v2` | Whisper (English-only) | Self-hosted by Telnyx | | `cohere/ar-stt` | Cohere | Self-hosted by Telnyx | | `azure/fast` | Azure | Third-party | | `soniox/stt-rt-v4` | Soniox | Third-party | | `xai/grok-stt` | xAI | Third-party | ### Text-to-speech (TTS) TTS is delivered through Telnyx's TTS gateway, which integrates multiple providers. The provider depends on the voice you select: | Provider | Hosting | | --- | --- | | Telnyx (in-house voices, including Telnyx Ultra) | Self-hosted by Telnyx | | Resemble | Self-hosted by Telnyx | | ElevenLabs | Third-party | | AWS | Third-party | | Azure | Third-party | | Minimax | Third-party | | Inworld | Third-party | | xAI | Third-party | See [Text to Speech voices](/docs/voice/tts/overview) for the current voice catalog. ### Large language model (LLM) The assistant's model is served through Telnyx's inference platform. The model in use is the one you configure on the assistant. **Self-hosted by Telnyx** (open models served on Telnyx infrastructure) include the **Qwen** and **Moonshot (Kimi)** model families — for example, `Qwen/Qwen3-235B-A22B`, `moonshotai/Kimi-K2.5`, and `moonshotai/Kimi-K2.6`. **Third-party** models — including those from **Anthropic** (Claude), **OpenAI** (GPT), and **Google** (Gemini) — are **not self-hosted**. When you select one of these, the prompt is sent to that external provider to generate the response. The available models evolve over time — for the current catalog and which models are recommended for assistants, see [Models](/docs/inference/models). If data residency or third-party data sharing is a concern, choose a self-hosted model (a Qwen or Moonshot/Kimi model) to keep prompt and response generation on Telnyx infrastructure. Region is not guaranteed even for self-hosted models. ### Can STT, TTS, or LLM processing be restricted to the EU? There is **no hard guarantee** of processing region for any provider — Telnyx will endeavor to honor the configured region, but does not guarantee it. In addition: - **Self-hosted** providers keep that processing step on Telnyx infrastructure, but region is not guaranteed. - **Third-party** providers route that step to the vendor, whose own region behavior applies. If you need STT, TTS, or LLM processing constrained to a specific region, [contact support](mailto:support@telnyx.com) so we can advise which self-hosted provider/model combinations best fit your requirement. This answer is about **Voice AI Assistants**, which do not currently expose a region control. Two other surfaces do: the **Inference API**, per request — see [Can Inference traffic be pinned to a specific region?](#can-inference-traffic-be-pinned-to-a-specific-region) — and **AI Assistant web chat**, per assistant, see [Can assistant chat be restricted to one region?](#can-assistant-chat-be-restricted-to-one-region). Bringing the same control to voice assistants is under consideration. --- ## Recordings ### Are call recordings disabled by default? No — for Voice AI Assistants, **call recordings are enabled by default**, and you can turn them off. When recordings are enabled, the recording is stored as Media Storage, which is subject to your [Data Locality](/docs/account-setup/data-locality) setting. Disable recording on the assistant (or per call) if you do not want recordings retained. --- ## Data retention and model training ### What does the data-retention setting control? Voice AI Assistants expose a **data-retention** privacy setting (`privacy_settings.data_retention`). It is **enabled by default**. When you disable it, the assistant stops persisting conversation **content** while continuing the minimum processing needed to run and bill the call. When `data_retention` is **disabled**, conversation content is **not retained**: | Item | Behavior when retention is off | | --- | --- | | Conversation messages / transcripts | Not persisted to the conversations store | | Insights | Not retained. An insight may be computed transiently in-memory to support live conversation behavior, but the conversation and its insights are not stored | | Transcript and assistant answer in observability logs | Not retained; replaced with placeholders (for example, `[transcript not available]` / `[answer not available]`) | | LLM request/response content logging | Disabled | | TTS cache | Disabled, so synthesized audio is not cached | A limited set of records is still **retained** even when conversation retention is off, because they are required to operate and bill the service: | Item | Behavior when retention is off | | --- | --- | | Latency / timing metrics | Retained (timing only, no conversation content) | | Billing, security, and fraud-prevention records | Retained as required for legitimate business and compliance purposes | The data-retention flag governs retention of **conversation content** for Voice AI Assistants. Disabling it stops persistence of conversation content and insights; it does not change where data that *is* retained lives — storage region is controlled by [Data Locality](/docs/account-setup/data-locality). Recordings are governed separately by the recording setting (see [Recordings](#recordings) above). For a guarantee tailored to your exact configuration (audio, tool inputs/outputs, memory, observability traces, and third-party provider logs), confirm in writing with your account team and DPA. ### Can a customer opt out of model improvement / training / evaluation? Customer data handling for model training is governed by Telnyx's applicable terms and DPA. If you require an opt-out from model improvement, training, or evaluation — for both input and output data, and covering Telnyx and any third-party AI providers in your configuration — [contact your account team](mailto:support@telnyx.com) to confirm the governing terms and document the opt-out. --- ## Usage reporting and billing ### Can usage be broken down by assistant, phone number, or metadata/tag? Usage and conversation data can be attributed using identifiers such as the assistant, the associated phone number, and metadata. For subscriber-level or per-tag billing breakdowns, [contact support](mailto:support@telnyx.com) to confirm which dimensions are available and how to structure metadata/tags for clean attribution. See [Agent Observability](/docs/inference/ai-assistants/agent-observability) and [Session Analysis](/docs/reporting/session-analysis). --- ## Related resources - [Data Locality](/docs/account-setup/data-locality) - [Inference Regions & Availability](/docs/inference/models/regions) - [Can assistant chat be restricted to one region?](#can-assistant-chat-be-restricted-to-one-region) - [Models](/docs/inference/models) - [Transcription Settings](/docs/inference/ai-assistants/transcription-settings) - [Text to Speech voices](/docs/voice/tts/overview) - [Agent Observability](/docs/inference/ai-assistants/agent-observability) --- ## Disclaimer This FAQ is provided for informational purposes only and describes the current design, functionality, and operation of the products, services, and features discussed herein. It is not intended to, and does not, create any contractual commitment, representation, warranty, guarantee, service level, or other binding obligation on the part of Telnyx. The information in this FAQ reflects Telnyx's current products, features, configurations, and operational practices as of the date of publication and may change from time to time. Any descriptions regarding processing locations, storage locations, data locality, service architecture, routing, infrastructure, providers, retention settings, product functionality, or the operation of the services should not be relied upon as contractual commitments or guarantees. References to particular regions, locations, providers, configurations, or operational outcomes do not constitute guarantees that data processing, storage, routing, or other service activities will occur exclusively in a particular location or manner. Customer rights and Telnyx obligations are governed exclusively by the applicable agreement(s) between the parties, including any Master Services Agreement, Data Processing Agreement, Order Form, and any other written contractual commitments expressly agreed by Telnyx. Nothing in this FAQ modifies or supplements such agreements. To the extent of any inconsistency between this FAQ and such agreements, the applicable agreement(s) shall control. --- ### Integrations > Source: https://developers.telnyx.com/docs/inference/integrations.md OpenAI-compatible API. Swap `base_url` and `api_key` in any framework that supports OpenAI. ## Quick Reference | Framework | Swap Method | Guide | |-----------|-------------|-------| | OpenAI SDK | `base_url` in client constructor | [OpenAI Migration](/docs/inference/openai) | | LangChain | `base_url` in `ChatOpenAI` | [LangChain](/docs/inference/langchain-integration) | | LlamaIndex | `api_base` in `OpenAILike` | [LlamaIndex](/docs/inference/llama-index) | | CrewAI | `OPENAI_BASE_URL` env var or `base_url` in LLM | [CrewAI](/docs/inference/crewai) | | LiveKit | Telnyx as LLM provider | [LiveKit](/docs/inference/livekit) | ## Environment Variables Route all OpenAI SDK calls through Telnyx with no code changes: ```shell export OPENAI_API_KEY=your_telnyx_api_key export OPENAI_BASE_URL=https://api.telnyx.com/v2/ai/openai ``` --- ### OpenAI Migration > Source: https://developers.telnyx.com/docs/inference/openai.md Swap two environment variables and change the model name. That's it. ```shell export OPENAI_BASE_URL='https://api.telnyx.com/v2/ai/openai' export OPENAI_API_KEY='KEY***' ``` ```python from openai import OpenAI client = OpenAI() # picks up env vars chat_completion = client.chat.completions.create( model="zai-org/GLM-5.3-Flash", messages=[{"role": "user", "content": "Tell me about Telnyx"}], reasoning_effort="high", temperature=0.0, stream=True, ) ``` Or pass explicitly: ```python import os from openai import OpenAI client = OpenAI( api_key=os.getenv("TELNYX_API_KEY"), base_url="https://api.telnyx.com/v2/ai/openai", ) chat_completion = client.chat.completions.create( model="zai-org/GLM-5.3-Flash", messages=[{"role": "user", "content": "Tell me about Telnyx"}], reasoning_effort="high", temperature=0.0, stream=True, ) ``` ## Reasoning models Reasoning models such as `zai-org/GLM-5.3-Flash` add a `reasoning_content` field alongside the usual `content`. It holds the model's chain-of-thought and appears on `message` (non-streaming) or `delta` (streaming). Read it the same way you read `content`: ```python chat_completion = client.chat.completions.create( model="zai-org/GLM-5.3-Flash", messages=[{"role": "user", "content": "Tell me about Telnyx"}], reasoning_effort="high", ) message = chat_completion.choices[0].message # Reasoning models populate reasoning_content; other models leave it None. if getattr(message, "reasoning_content", None): print("reasoning:", message.reasoning_content) print("answer:", message.content) ``` ## Reasoning effort Reasoning models accept a `reasoning_effort` parameter that controls how much compute the model spends on internal reasoning before generating its response. ```python chat_completion = client.chat.completions.create( model="zai-org/GLM-5.3-Flash", messages=[{"role": "user", "content": "Explain quantum computing"}], reasoning_effort="high", ) ``` Supported values: `none`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max`. Not all models support all values — unsupported combinations return a 400 error. When omitted, reasoning models use their default effort level. For vLLM/SGLang-hosted models, `enable_thinking=false` overrides `reasoning_effort` and disables reasoning entirely. ## Chat Completions Compatibility | Parameter | Telnyx | OpenAI | |-----------|:------:|:------:| | `messages` | ✅ | ✅ | | `model` | ✅ | ✅ | | `stream` | ✅ | ✅ | | `max_tokens` | ✅ | ✅ | | `temperature` | ✅ | ✅ | | `top_p` | ✅ | ✅ | | `frequency_penalty` | ✅ | ✅ | | `presence_penalty` | ✅ | ✅ | | `n` | ✅ | ✅ | | `stop` | ✅ | ✅ | | `logit_bias` | ✅ | ✅ | | `logprobs` | ✅ | ✅ | | `top_logprobs` | ✅ | ✅ | | `seed` | ✅ | ✅ | | `response_format` | ✅ | ✅ | | `tool_choice` | ✅ | ✅ | | `tools` | ✅ | ✅ | | `function` | ✅ | ✅ | | `reasoning_effort` | ✅ | ✅ (o-family/GPT-5 only) | | `retrieval` | ✅ | ❌ | | `guided_json` | ✅ | ❌ | | `guided_regex` | ✅ | ❌ | | `guided_choice` | ✅ | ❌ | | `min_p` | ✅ | ❌ | | `use_beam_search` | ✅ | ❌ | | `best_of` | ✅ | ❌ | | `length_penalty` | ✅ | ❌ | | `early_stopping` | ✅ | ❌ | | `user` | ❌ | ✅ | ## Transcriptions Compatibility | Parameter | Telnyx | OpenAI | |-----------|:------:|:------:| | `file` | ✅ | ✅ | | `model` | ✅ | ✅ | | `response_format` | ✅ | ✅ | | `timestamp_granularities[]` → `segment` | ✅ | ✅ | | `timestamp_granularities[]` → `word` | ❌ | ✅ | | `language` | ❌ | ✅ | | `prompt` | ❌ | ✅ | | `temperature` | ❌ | ✅ | --- ### Anthropic Migration > Source: https://developers.telnyx.com/docs/inference/anthropic.md Swap the base URL and pass your Telnyx API key as a Bearer token. That's it. The Telnyx Inference API now exposes an Anthropic-compatible Messages endpoint at `POST /v2/ai/anthropic/v1/messages`. It accepts the same request body as the [Anthropic Messages API](https://docs.anthropic.com/en/api/messages) and returns the same response shape — including streaming via Anthropic SSE event types (`message_start`, `content_block_start`, `content_block_delta`, `content_block_stop`, `message_delta`, `message_stop`). ## Authentication The Anthropic SDK sends requests with an `x-api-key` header by default. Telnyx uses `Authorization: Bearer ` instead. Pass the Telnyx key as a `default_headers` override and set the SDK's own `api_key` to any placeholder value — the gateway ignores it. ```shell export TELNYX_API_KEY='KEY***' ``` ```python import os from anthropic import Anthropic client = Anthropic( api_key="unused", # placeholder — the gateway uses the Bearer header below base_url="https://api.telnyx.com/v2/ai/anthropic", default_headers={ "Authorization": f"Bearer {os.environ['TELNYX_API_KEY']}", }, ) ``` ```javascript import Anthropic from "@anthropic-ai/sdk"; const client = new Anthropic({ apiKey: "unused", // placeholder — the gateway uses the Bearer header below baseURL: "https://api.telnyx.com/v2/ai/anthropic", defaultHeaders: { Authorization: `Bearer ${process.env.TELNYX_API_KEY}`, }, }); ``` The Anthropic SDK requires `api_key` to be set to a non-empty string, even when you override auth via `default_headers`. Use any placeholder — the Telnyx gateway only reads the `Authorization: Bearer` header. ## Quickstart ### Python ```python import os from anthropic import Anthropic client = Anthropic( api_key="unused", base_url="https://api.telnyx.com/v2/ai/anthropic", default_headers={ "Authorization": f"Bearer {os.environ['TELNYX_API_KEY']}", }, ) # Non-streaming message = client.messages.create( model="zai-org/GLM-5.3-Flash", max_tokens=1024, system="You are a friendly chatbot.", messages=[ {"role": "user", "content": "Tell me about Telnyx"} ], ) print(message.content[0].text) ``` ### JavaScript / TypeScript ```javascript import Anthropic from "@anthropic-ai/sdk"; const client = new Anthropic({ apiKey: "unused", baseURL: "https://api.telnyx.com/v2/ai/anthropic", defaultHeaders: { Authorization: `Bearer ${process.env.TELNYX_API_KEY}`, }, }); // Non-streaming const message = await client.messages.create({ model: "zai-org/GLM-5.3-Flash", max_tokens: 1024, system: "You are a friendly chatbot.", messages: [{ role: "user", content: "Tell me about Telnyx" }], }); console.log(message.content[0].text); ``` ### curl ```bash curl -sS -X POST "https://api.telnyx.com/v2/ai/anthropic/v1/messages" \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -H "anthropic-version: 2023-06-01" \ -d '{ "model": "zai-org/GLM-5.3-Flash", "max_tokens": 1024, "system": "You are a friendly chatbot.", "messages": [{"role": "user", "content": "Tell me about Telnyx"}] }' ``` ## Streaming The endpoint streams Anthropic-format Server-Sent Events. Use the SDK's built-in streaming just as you would with the native Anthropic API: ```python import os from anthropic import Anthropic client = Anthropic( api_key="unused", base_url="https://api.telnyx.com/v2/ai/anthropic", default_headers={ "Authorization": f"Bearer {os.environ['TELNYX_API_KEY']}", }, ) with client.messages.stream( model="zai-org/GLM-5.3-Flash", max_tokens=1024, system="You are a friendly chatbot.", messages=[{"role": "user", "content": "Tell me about Telnyx"}], ) as stream: for text in stream.text_stream: print(text, end="", flush=True) ``` ```javascript const stream = await client.messages.stream({ model: "zai-org/GLM-5.3-Flash", max_tokens: 1024, system: "You are a friendly chatbot.", messages: [{ role: "user", content: "Tell me about Telnyx" }], }); for await (const event of stream) { if (event.type === "content_block_delta" && event.delta.type === "text_delta") { process.stdout.write(event.delta.text); } } ``` ## Tool Calling Tool definitions and tool results follow the [Anthropic tool use format](https://docs.anthropic.com/en/docs/build-with-claude/tool-use): ```python import os import json from anthropic import Anthropic client = Anthropic( api_key="unused", base_url="https://api.telnyx.com/v2/ai/anthropic", default_headers={ "Authorization": f"Bearer {os.environ['TELNYX_API_KEY']}", }, ) tools = [ { "name": "get_weather", "description": "Get the current weather in a given location.", "input_schema": { "type": "object", "properties": { "location": {"type": "string", "description": "City and state, e.g. San Francisco, CA"}, }, "required": ["location"], }, } ] response = client.messages.create( model="zai-org/GLM-5.3-Flash", max_tokens=1024, tools=tools, messages=[{"role": "user", "content": "What's the weather in Lisbon?"}], ) # The model returns a tool_use content block for block in response.content: if block.type == "tool_use": print(f"Tool: {block.name}") print(f"Input: {block.input}") ``` ## Extended Thinking For models that support extended thinking (e.g. Claude reasoning models), pass the `thinking` parameter. On older SDK versions that reject unknown kwargs, use `extra_body` to forward it into the request JSON: ```python response = client.messages.create( model="zai-org/GLM-5.3-Flash", max_tokens=4096, messages=[{"role": "user", "content": "Exprove que 17 é primo."}], extra_body={ "thinking": {"type": "enabled", "budget_tokens": 2048}, }, ) ``` ## System Prompts The `system` parameter can be a plain string or an array of content blocks, matching the [Anthropic API format](https://docs.anthropic.com/en/api/messages): ```python # String system prompt response = client.messages.create( model="zai-org/GLM-5.3-Flash", max_tokens=1024, system="You are a helpful assistant.", messages=[{"role": "user", "content": "Hello!"}], ) # Array-style system prompt (cache_control, etc.) response = client.messages.create( model="zai-org/GLM-5.3-Flash", max_tokens=1024, system=[ {"type": "text", "text": "You are a helpful assistant.", "cache_control": {"type": "ephemeral"}}, ], messages=[{"role": "user", "content": "Hello!"}], ) ``` ## Models Anthropic models are available under the `anthropic/` prefix. See [Available Models](/docs/inference/models) for the full list. Open-source models hosted on Telnyx (e.g. `zai-org/GLM-5.3-Flash`, `moonshotai/Kimi-K2.6`) also work through this endpoint — the request is translated to the OpenAI-compatible format internally and the response is translated back to the Anthropic shape. ## Telnyx Extensions The endpoint accepts several Telnyx-specific fields alongside the standard Anthropic request body: | Field | Type | Description | |-------|------|-------------| | `api_key_ref` | string | Reference to an [integration secret](/api-reference/integration-secrets/create-a-secret) for external provider keys. | | `mcp_servers` | array | List of MCP (Model Context Protocol) server configs to expose to the model. | | `fallback_config` | object | Configuration for automatic model fallback when the primary model is unavailable. | | `billing_group_id` | uuid | Billing group to associate with this request. | | `timeout` | number | Request timeout in seconds (default: 300). | | `max_retries` | integer | Maximum retry attempts for the request. | | `service_tier` | string | Service tier for the request. | These fields pass through as extra body parameters in the SDK: ```python response = client.messages.create( model="zai-org/GLM-5.3-Flash", max_tokens=1024, messages=[{"role": "user", "content": "Hello!"}], extra_body={ "billing_group_id": "6a09cdc3-8948-47f0-aa62-74ac943d6c58", }, ) ``` ## Compatibility | Parameter | Telnyx | Anthropic | |-----------|:------:|:---------:| | `model` | ✅ | ✅ | | `messages` | ✅ | ✅ | | `max_tokens` | ✅ | ✅ | | `system` | ✅ | ✅ | | `stream` | ✅ | ✅ | | `temperature` | ✅ | ✅ | | `top_p` | ✅ | ✅ | | `top_k` | ✅ | ✅ | | `stop_sequences` | ✅ | ✅ | | `metadata` | ✅ | ✅ | | `tools` | ✅ | ✅ | | `tool_choice` | ✅ | ✅ | | `thinking` | ✅ | ✅ | | `api_key_ref` | ✅ | ❌ | | `mcp_servers` | ✅ | ❌ | | `fallback_config` | ✅ | ❌ | | `billing_group_id` | ✅ | ❌ | | `timeout` | ✅ | ❌ | | `service_tier` | ✅ | ❌ | --- ### LangChain > Source: https://developers.telnyx.com/docs/inference/langchain-integration.md OpenAI-compatible. Use `ChatOpenAI` with a `base_url` swap. ## Setup ```shell pip install langchain-openai ``` ## Usage ```python import os from langchain_openai import ChatOpenAI llm = ChatOpenAI( base_url="https://api.telnyx.com/v2/ai/openai", api_key=os.getenv("TELNYX_API_KEY"), model="zai-org/GLM-5.3-Flash", reasoning_effort="high", ) for chunk in llm.stream("Help me plan my vacation"): print(chunk.content, end="", flush=True) ``` ## Function Calling ```python import os from langchain_openai import ChatOpenAI from langchain_core.tools import tool @tool def get_weather(location: str) -> str: """Get the current weather for a location.""" return f"The weather in {location} is sunny and 72°F." llm_with_tools = ChatOpenAI( base_url="https://api.telnyx.com/v2/ai/openai", api_key=os.getenv("TELNYX_API_KEY"), model="zai-org/GLM-5.3-Flash", reasoning_effort="high", ).bind_tools([get_weather]) result = llm_with_tools.invoke("What's the weather in Chicago?") print(result.tool_calls) ``` ## Streaming ```python from langchain_core.messages import HumanMessage messages = [HumanMessage(content="Explain quantum computing in 3 sentences")] for chunk in llm.stream(messages): print(chunk.content, end="", flush=True) ``` --- ### LlamaIndex > Source: https://developers.telnyx.com/docs/inference/llama-index.md OpenAI-compatible. Use `OpenAILike` with `api_base` swap. ## Setup ```shell pip install llama-index-core llama-index-llms-openai-like ``` ## Usage ```python import os from llama_index.llms.openai_like import OpenAILike from llama_index.core.llms import ChatMessage llm = OpenAILike( api_base="https://api.telnyx.com/v2/ai/openai", api_key=os.getenv("TELNYX_API_KEY"), model="zai-org/GLM-5.3-Flash", additional_kwargs={"reasoning_effort": "high"}, is_chat_model=True, ) chat = llm.stream_chat([ChatMessage(role="user", content="Help me plan my vacation")]) for chunk in chat: print(chunk.delta, end="") ``` ## RAG with Embeddings Combine with [Telnyx Embeddings](/docs/inference/embeddings) for retrieval-augmented generation. See the [Embeddings guide](/docs/inference/embeddings) for document upload and indexing. --- ### CrewAI > Source: https://developers.telnyx.com/docs/inference/crewai.md OpenAI-compatible. Use as LLM backend for CrewAI agents. ## Setup ```shell pip install crewai ``` ## Usage Set environment variables for global routing: ```shell export TELNYX_API_KEY=your_telnyx_api_key export OPENAI_BASE_URL=https://api.telnyx.com/v2/ai/openai ``` Or configure per-agent: ```python import os from crewai import Agent, Task, Crew, LLM llm = LLM( model="zai-org/GLM-5.3-Flash", base_url="https://api.telnyx.com/v2/ai/openai", api_key=os.getenv("TELNYX_API_KEY"), reasoning_effort="high", ) researcher = Agent( role="Research Analyst", goal="Find and analyze information", backstory="You are an experienced research analyst.", llm=llm, ) writer = Agent( role="Technical Writer", goal="Write clear, accurate reports", backstory="You are a skilled technical writer.", llm=llm, ) research_task = Task( description="Research the latest trends in AI infrastructure", agent=researcher, ) write_task = Task( description="Write a summary report based on the research findings", agent=writer, ) crew = Crew(agents=[researcher, writer], tasks=[research_task, write_task]) result = crew.kickoff() print(result) ``` ## Tool Calling ```python from crewai.tools import tool @tool("Search the web") def search_web(query: str) -> str: """Search the web for information.""" return f"Results for: {query}" researcher = Agent( role="Research Analyst", goal="Find and analyze information", backstory="You are an experienced research analyst.", llm=llm, tools=[search_web], ) ``` --- ### LiveKit > Source: https://developers.telnyx.com/docs/inference/livekit.md LiveKit's [agent framework](https://docs.livekit.io/agents/overview/) lets you build real-time, programmable voice agents. Telnyx integrates with LiveKit through the OpenAI plugin for LLM inference and through `livekit-plugins-telnyx` for native STT and TTS. ## Voice assistant example This example is based on LiveKit's [agents examples repo](https://github.com/livekit/agents/tree/main/examples), modified to use Telnyx for LLM inference. ### Set up and activate a virtual env ```bash python -m venv venv source venv/bin/activate ``` ### Install requirements ```bash pip install -r requirements.txt pip install livekit-plugins-telnyx ``` ### Download files This downloads model weights for voice-activity detection: ```bash python agent.py download-files ``` ### Agent code The following code uses Telnyx for LLM inference via `openai.LLM.with_telnyx()`, with Telnyx STT and TTS. ```python from dotenv import load_dotenv from livekit.agents import AutoSubscribe, JobContext, WorkerOptions, cli, voice, llm from livekit.plugins import openai, silero, telnyx load_dotenv() async def entrypoint(ctx: JobContext): initial_ctx = llm.ChatContext().append( role="system", text=( "You are a helpful voice assistant powered by Telnyx. " "Keep responses short and conversational." ), ) await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY) session = voice.AgentSession( llm=openai.LLM.with_telnyx( model="zai-org/GLM-5.3-Flash", reasoning_effort="high", ), vad=silero.VAD.load(), stt=telnyx.STT(), tts=telnyx.TTS(voice="Telnyx.Ultra.Clara"), chat_ctx=initial_ctx, ) session.start(ctx.room) await session.say("Hey, how can I help you today?", allow_interruptions=True) if __name__ == "__main__": cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint)) ``` ### Set environment variables ```bash export TELNYX_API_KEY= export LIVEKIT_URL= export LIVEKIT_API_KEY= export LIVEKIT_API_SECRET= ``` ### Run the agent worker ```bash python agent.py dev ``` ### Test with a LiveKit frontend Use the [LiveKit Agents Playground](https://agents-playground.livekit.io) to test your agent without building a frontend. ## Telnyx STT & TTS plugin The `livekit-plugins-telnyx` package provides native Telnyx STT and TTS plugins for LiveKit agents. ```bash pip install livekit-plugins-telnyx ``` ### STT Use `telnyx.STT()` for real-time speech-to-text via Telnyx's WebSocket streaming API: ```python from livekit.plugins import telnyx session = voice.AgentSession( stt=telnyx.STT(), # ... other plugins ) ``` ### TTS Use `telnyx.TTS()` for real-time text-to-speech. Pass a `voice` parameter to select a specific voice: ```python from livekit.plugins import telnyx session = voice.AgentSession( tts=telnyx.TTS(voice="Telnyx.Ultra.Clara"), # ... other plugins ) ``` See [TTS available voices](/docs/voice/tts/overview) for the full list of voice options. ## Related resources - [LiveKit Telnyx LLM plugin docs](https://docs.livekit.io/agents/models/llm/plugins/telnyx/) - [SIP trunk configuration for LiveKit](/docs/voice/sip-trunking/livekit-configuration-guide) - [Inference getting started](/docs/inference/models) --- ### Inference API > Source: https://developers.telnyx.com/docs/inference/getting-started.md ## Prerequisites - [Telnyx account](https://telnyx.com/sign-up) - [API Key](https://portal.telnyx.com/#/app/auth/v2) - Python 3.8+ Install the OpenAI SDK: ```shell pip install openai ``` The Inference API is OpenAI-compatible. Any OpenAI SDK works with a `base_url` swap. ## Python ```python import os from openai import OpenAI client = OpenAI( api_key=os.getenv("TELNYX_API_KEY"), base_url="https://api.telnyx.com/v2/ai/openai", ) chat_completion = client.chat.completions.create( messages=[ {"role": "user", "content": "Tell me about Telnyx"} ], model="zai-org/GLM-5.3-Flash", reasoning_effort="high", stream=True ) # GLM-5.3 Flash is a reasoning model: it streams its thinking in `reasoning_content` # before the final answer in `content`. Print both so you can see the reasoning. reasoning_started = False content_started = False for chunk in chat_completion: delta = chunk.choices[0].delta if getattr(delta, "reasoning_content", None): if not reasoning_started: print("--- reasoning ---") reasoning_started = True print(delta.reasoning_content, end="", flush=True) if delta.content: if not content_started: print("\n--- answer ---") content_started = True print(delta.content, end="", flush=True) ``` Reasoning models such as `zai-org/GLM-5.3-Flash` return their chain-of-thought in a separate `reasoning_content` field (on `message` for non-streaming responses, or `delta` when streaming). Models without reasoning simply omit it, so the `getattr(..., "reasoning_content", None)` guard works for every model. ## Core Concepts ### Messages Chat history passed to the model. ### Roles Every message has a role: **system**, **user**, **assistant**, or **tool**. - **system** — model behavior instructions - **user** — end-user input - **assistant** — model output - **tool** — function call results. See [Function Calling](/docs/inference/functions). ### Models [Available Models](/docs/inference/models) lists all hosted LLMs with context lengths and capabilities. ### Streaming Server-sent events, same as OpenAI. ## What Next? | I want to... | Go to | |:-------------|:------| | Build a voice assistant | [No-Code Voice Assistant](/docs/inference/ai-assistants/no-code-voice-assistant) | | Call custom code from the model | [Function Calling](/docs/inference/functions) / [Streaming Functions](/docs/inference/streaming-functions) | | Ground responses in documents | [Embeddings](/docs/inference/embeddings) | | Identify themes in data | [Clusters](/docs/inference/clusters) | | Migrate from OpenAI | [OpenAI Migration](/docs/inference/openai) | | Browse all models | [Available Models](/docs/inference/models) | --- ### Function Calling > Source: https://developers.telnyx.com/docs/inference/functions.md In this tutorial, you'll learn how to connect large language models to external tools using our [chat completions API](/api-reference/openai-chat/create-a-chat-completion-openai-compatible). This includes: - Defining a function - Enabling the language model to choose the function - Executing the function - Sharing the results with the language model ## Introduction Using the `tools` field, you can enable a language model to choose functions to call. The [chat completions API](/api-reference/openai-chat/create-a-chat-completion-openai-compatible) does not call the function itself. It will return the arguments you need to execute the function yourself. Of the open-source language models hosted on Telnyx, `zai-org/GLM-5.3-Flash` is especially good at calling functions. While we recommend you start with this model, every model in our API supports the `tools` interface. ## Simple `get_current_weather` example A popular toy example for function calls is the `get_current_weather` example. The following code defines a function and passes it to the language model via the `tools` field. Make sure you have set the `TELNYX_API_KEY` environment variable. ```python import os from openai import OpenAI client = OpenAI( api_key=os.getenv("TELNYX_API_KEY"), base_url="https://api.telnyx.com/v2/ai/openai", ) tools = [ { "type": "function", "function": { "name": "get_current_weather", "description": "Get the current weather", "parameters": { "type": "object", "properties": { "location": { "type": "string", "description": "The city and state, e.g. San Francisco, CA", }, "unit": { "type": "string", "enum": ["celsius", "fahrenheit"], "description": "The temperature unit to use", }, }, "required": ["location", "unit"], }, } } ] messages = [ {"role": "user", "content": "How is the weather in Chicago?"} ] chat_completion = client.chat.completions.create( model="zai-org/GLM-5.3-Flash", messages=messages, reasoning_effort="high", tools=tools, tool_choice="auto" ) print(chat_completion.choices[0].message) ``` A `tool_choice` of `auto` lets the language model decide to call a function (or not). The options for `tool_choice` are: - `required`: this forces the language model to choose a tool - `none`: this forces the language model to NOT choose a tool - `auto`: this lets the language model decide If the language model chooses a function, the above will result in a response like this, with the `tool_calls` field populated. ``` ChatCompletionMessage(content=None, role='assistant', function_call=None, tool_calls=[ChatCompletionMessageToolCall(id='call_c31258d2-78a8-4566-b716-e3b2a774cbdb', function=Function(arguments='{"location": "Chicago", "unit": "fahrenheit"}', name='get_current_weather'), type='function')]) ``` ## Defining functions programmatically In the next example, we will implement and execute `get_current_weather`. To do this cleanly, we are first going to define a helper function `func_to_tool` that extends the `schema` function from Jeremy Howard's [A Hacker's Guide to Language Models](https://github.com/fastai/lm-hackers/blob/main/lm-hackers.ipynb). ```python import inspect import os import json from typing import Literal from openai import OpenAI from pydantic import create_model def func_to_tool(f): kw = { n: (o.annotation, ... if o.default==inspect.Parameter.empty else o.default) for n, o in inspect.signature(f).parameters.items() } s = create_model(f.__name__, **kw).model_json_schema() tool_json = { "type": "function", "function": { "name": s["title"], "description": inspect.getdoc(f), "parameters": s } } return tool_json TempUnit = Literal['celsius', 'fahrenheit'] def get_current_weather(location: str, unit: TempUnit = 'fahrenheit'): """Get the current weather in a given location""" if "tokyo" in location.lower(): return json.dumps({"location": "Tokyo", "temperature": "10", "unit": unit}) elif "san francisco" in location.lower(): return json.dumps({"location": "San Francisco", "temperature": "72", "unit": unit}) elif "chicago" in location.lower(): return json.dumps({"location": "Chicago", "temperature": "22", "unit": unit}) else: return json.dumps({"location": location, "temperature": "unknown"}) client = OpenAI( api_key=os.getenv("TELNYX_API_KEY"), base_url="https://api.telnyx.com/v2/ai/openai", ) tools = [func_to_tool(get_current_weather)] messages = [ {"role": "user", "content": "How is the weather in Chicago?"} ] chat_completion = client.chat.completions.create( model="zai-org/GLM-5.3-Flash", messages=messages, reasoning_effort="high", tools=tools, tool_choice="auto" ) print(chat_completion.choices[0].message) ``` This is now functionally equivalent to the first example, but instead of writing and maintaining the verbose JSON definition ourselves, we can generate it programmatically from the executable Python function. ## Executing functions Ok, it's nice that the language model wants to execute the `get_current_weather`, but how do we actually do that, and incorporate the results back into the interaction? Continuing from the `chat_completion` response in the previous example ```python assistant_message = chat_completion.choices[0].message tool_calls = assistant_message.tool_calls content = assistant_message.content if tool_calls: messages.append(assistant_message) available_functions = {"get_current_weather": get_current_weather} for tool_call in tool_calls: function_name = tool_call.function.name function_to_call = available_functions[function_name] function_args = json.loads(tool_call.function.arguments) function_response = function_to_call(**function_args) messages.append( { "tool_call_id": tool_call.id, "role": "tool", "name": function_name, "content": function_response, } ) second_chat_completion = client.chat.completions.create( model="zai-org/GLM-5.3-Flash", messages=messages, reasoning_effort="high", ) print(second_chat_completion.choices[0].message.content) else: print(content) ``` Now, we will get our answer from the language model, incorporating the output from the function call. ``` It's 22°F in Chicago. ``` --- ### Streaming and Parallel Calls > Source: https://developers.telnyx.com/docs/inference/streaming-functions.md In the [previous tutorial](/docs/inference/functions), we learned the basics for defining and executing functions using our [chat completions API](/api-reference/openai-chat/create-a-chat-completion-openai-compatible). In this tutorial, we will introduce more advanced use cases: - Streaming function calls - Passing multiple functions - Executing function calls in parallel For low-latency contexts, streaming and parallel calls are especially helpful. ## Defining our functions First, we will define two functions we want to execute in parallel: `sleep` and `dream`. Our goal is to use the `dream` function to make an API call to the Telnyx chat completions endpoint while we `sleep`. We will also re-use the `func_to_tool` helped function we defined in the [previous tutorial](/docs/inference/functions) to easily convert between our Python functions and the JSON we need to pass to the `tools` field for our chat completions API. Make sure you have set the `TELNYX_API_KEY` environment variable ```python import asyncio import inspect import json import os from openai import AsyncOpenAI from pydantic import create_model # Configuration API_KEY = os.getenv("TELNYX_API_KEY") BASE_URL = "https://api.telnyx.com/v2/ai/openai" MODEL = "zai-org/GLM-5.3-Flash" client = AsyncOpenAI(api_key=API_KEY, base_url=BASE_URL) async def sleep(seconds: int): """Sleep for a given number of seconds.""" await asyncio.sleep(seconds) return f"I slept for {seconds} seconds!" async def dream(subject: str): """Dream about a given subject.""" chat_completion = await client.chat.completions.create( model=MODEL, reasoning_effort="high", messages=[ { "role": "user", "content": f"BRIEFLY (one sentence max) describe a dream about {subject}" } ] ) return chat_completion.choices[0].message.content def func_to_tool(f): """Convert a function to a tool JSON schema.""" kw = { n: (o.annotation, ... if o.default == inspect.Parameter.empty else o.default) for n, o in inspect.signature(f).parameters.items() } schema = create_model(f.__name__, **kw).model_json_schema() tool_json = { "type": "function", "function": { "name": schema["title"], "description": inspect.getdoc(f), "parameters": schema } } return tool_json ``` ## Parsing Streaming Tools + Executing Tasks in Parallel Next we will define a few functions to help us parse and execute tasks in parallel. ### handle_tool_calls The `handle_tool_calls` function will iterate over streamed chunks from the chat completions endpoint. The language model may invoke multiple tool calls to be executed in parallel and will differentiate them using the `index` attribute on the chunk. As we progress through the stream, we will build our local copy of this list of function calls in the `tool_calls` list. The first chunk of a new tool call will contain the `name` of the function. This enables you to give early feedback to users that a function will be executed. In this example, we simply print the name of the function when it is detected. As we build the arguments from the streamed chunks, we attempt to parse what we have built as JSON. Once we have a valid JSON object, we create an async task to be scheduled for execution (if we have not already done so). **NB: Telnyx guarantees valid JSON is returned for tool calls, so you don't have to worry about lengthy retries or fuzzy matching.** ### execute_tasks This function executes the tasks from the previous function and returns the results as they are completed, enabling users to receive feedback as soon as possible. ### func_wrapper This is a trivial helper function that exposes the tool call ID and function name to `execute_tasks` ```python async def func_wrapper(func, tool_call_id, **kwargs): """Wrap a function to return its ID + name when executed.""" result = await func(**kwargs) return tool_call_id, func.__name__, result async def execute_tasks(tasks): """Execute asynchronous tasks and collect their results.""" results = [] for task in asyncio.as_completed(tasks): tool_call_id, func_name, result = await task print(f"Executed {func_name}, results: {result}") results.append( { "tool_call_id": tool_call_id, "role": "tool", "name": func_name, "content": result, } ) return results async def handle_tool_calls(chat_completion, function_map): """Handle streaming tool calls from chat completion.""" tool_calls = [] tasks = [] tasked_tool_ids = set() async for chunk in chat_completion: delta = chunk.choices[0].delta if delta and delta.tool_calls: # We have detected tool calls from the LLM tcchunklist = delta.tool_calls for tcchunk in tcchunklist: index = tcchunk.index or 0 if len(tool_calls) <= index: # Based on the index, we have a new tool call tool_calls.append( { "id": "", "type": "function", "function": { "name": "", "arguments": "" } } ) tc = tool_calls[index] if tcchunk.id: tc["id"] += tcchunk.id if tcchunk.function.name: tc["function"]["name"] += tcchunk.function.name print(f"Detected function: {tcchunk.function.name}") if tcchunk.function.arguments: tc["function"]["arguments"] += tcchunk.function.arguments try: kwargs = json.loads(tc["function"]["arguments"]) except json.JSONDecodeError: # We don't have the full arguments JSON yet continue else: if tc["id"] not in tasked_tool_ids: func_name = tc["function"]["name"] print(f"Executing {func_name} with {kwargs}") wrapped_func = func_wrapper(function_map[func_name], tc["id"], **kwargs) task = asyncio.create_task(wrapped_func) tasks.append(task) tasked_tool_ids.add(tc["id"]) return tool_calls, tasks ``` ## Putting it all together With our helper functions defined, we are ready to stream and execute multiple function calls in parallel. In this code, we: - Ask the language model to `sleep` and `dream` at the same time - Execute the returned tool calls in parallel - Provide the results back to the language model and get a final response ```python async def main(): prompt = "Take a quick 10 second power nap and dream about Telnyx. Then write a haiku about it!" messages = [{"role": "user", "content": prompt}] print(f"Prompt: {prompt}") functions = [sleep, dream] function_map = {f.__name__: f for f in functions} tools = [func_to_tool(func) for func in functions] chat_completion = await client.chat.completions.create( model=MODEL, messages=messages, reasoning_effort="high", tools=tools, tool_choice="required", stream=True ) tool_calls, tasks = await handle_tool_calls(chat_completion, function_map) messages.append( { "role": "assistant", "tool_calls": tool_calls, } ) task_results = await execute_tasks(tasks) messages.extend(task_results) print("Sending results back to LLM...") print() second_chat_completion = await client.chat.completions.create( model=MODEL, messages=messages, reasoning_effort="high", stream=True, ) # GLM-5.3 Flash is a reasoning model: stream reasoning_content (its thinking) and # content (the final answer) separately. Non-reasoning models omit the former. async for chunk in second_chat_completion: delta = chunk.choices[0].delta if getattr(delta, "reasoning_content", None): print(delta.reasoning_content, end="", flush=True) if delta.content: print(delta.content, end="", flush=True) print() if __name__ == "__main__": asyncio.run(main()) ``` The output of the print statements in this script will look something like this. Notice that `sleep` was detected and executed first, but `dream` still returned results first. ``` Prompt: Take a quick 10 second power nap and dream about Telnyx. Then write a haiku about it! Detected function: sleep Executing sleep with {'seconds': 10} Detected function: dream Executing dream with {'subject': 'Telnyx'} Executed dream, results: In my dream, I was walking through a futuristic cityscape where Telnyx's logo was emblazoned on skyscrapers, and I could hear the hum of millions of concurrent voice calls and messages being transmitted seamlessly through their network. Executed sleep, results: I slept for 10 seconds! Sending results back to LLM... Here is a haiku about Telnyx: Telnyx city glows Voices whisper through the air Connected we stand ``` --- ### JSON Mode and Beyond > Source: https://developers.telnyx.com/docs/inference/json-mode.md In this tutorial, you'll learn how to: * Guarantee structured output using our [chat completions API](/api-reference/openai-chat/create-a-chat-completion-openai-compatible) * This can be done using JSON Schema / Pydantic models, (schemaless) JSON Mode, and multiple choice. ## Sentiment Analysis using Multiple Choice One of the simplest forms of structured output is multiple choice. The schema below pins the classification to a single value — the response is always `{"sentiment": "positive"}` or `{"sentiment": "negative"}`. Make sure you have set the `TELNYX_API_KEY` environment variable `zai-org/GLM-5.3-Flash` is a reasoning model. Your structured output still arrives in `choices[0].message.content` — the model's chain-of-thought is returned separately in `choices[0].message.reasoning_content`, so it never pollutes the JSON you parse. To surface the reasoning, read it alongside `content`: ```python message = chat_completion.choices[0].message if getattr(message, "reasoning_content", None): print("reasoning:", message.reasoning_content) print(message.content) # the structured output ``` ```python import os from openai import OpenAI client = OpenAI( api_key=os.getenv("TELNYX_API_KEY"), base_url="https://api.telnyx.com/v2/ai/openai", ) chat_completion = client.chat.completions.create( model="zai-org/GLM-5.3-Flash", reasoning_effort="high", messages=[ {"role": "system", "content": "Classify the sentiment of this review as positive or negative."}, {"role": "user", "content": "The staff went above and beyond! I had a great stay."} ], temperature=0.0, response_format={ "type": "json_schema", "json_schema": { "name": "sentiment", "schema": { "type": "object", "properties": { "sentiment": { "type": "string", "enum": ["positive", "negative"] } }, "required": ["sentiment"], "additionalProperties": False } } } ) print(chat_completion.choices[0].message.content) ``` Since the schema's `sentiment` property is an enum with exactly two values, the model can only answer with one of `positive` or `negative`: ``` {"sentiment": "positive"} ``` ## Constraining the full response shape with `json_schema` Now let's say we want to capture an explanation for the classification as well. The `response_format` parameter with `"type": "json_schema"` guarantees the response parses as JSON matching your schema, and `"additionalProperties": false` on an object schema makes the response carry exactly the declared properties — without it, JSON Schema permits extra keys, so a schema listing `sentiment` and `explanation` doesn't by itself forbid a third key: ```python import os from openai import OpenAI client = OpenAI( api_key=os.getenv("TELNYX_API_KEY"), base_url="https://api.telnyx.com/v2/ai/openai", ) chat_completion = client.chat.completions.create( model="zai-org/GLM-5.3-Flash", reasoning_effort="high", messages=[ {"role": "system", "content": "First describe the sentiment of the following review and then use that to classify the sentiment as positive or negative."}, {"role": "user", "content": "The staff went above and beyond! I had a great stay."} ], temperature=0.0, response_format={ "type": "json_schema", "json_schema": { "name": "sentiment_analysis", "schema": { "type": "object", "properties": { "sentiment": { "type": "string", "enum": ["positive", "negative"] }, "explanation": {"type": "string"} }, "required": ["sentiment", "explanation"], "additionalProperties": False } } } ) print(chat_completion.choices[0].message.content) ``` This will ensure a JSON response with the same schema as the following ``` { "sentiment": "positive", "explanation": "The review mentions that the staff went above and beyond and the user had a great stay, which indicates a positive sentiment." } ``` The `guided_json`, `guided_choice`, and `guided_regex` request fields are accepted for compatibility with the vLLM/SGLang serving stack, but they are **not enforced** for Telnyx-hosted models today: a request carrying them is served as a plain chat completion, so the model may return free text that ignores the schema or choice list. Use `response_format` with `"type": "json_schema"` for guaranteed structured output on Telnyx-hosted models. ## Simplifying Schema Generation using Pydantic The above is helpful to see the raw JSON Schema specification being sent via API. However, for practical purposes, using Pydantic models can help simplify generating the specs. The following is functionally equivalent to the previous example. ```python import os from enum import Enum from openai import OpenAI from pydantic import BaseModel, ConfigDict client = OpenAI( api_key=os.getenv("TELNYX_API_KEY"), base_url="https://api.telnyx.com/v2/ai/openai", ) class Sentiment(str, Enum): positive = 'positive' negative = 'negative' class SentimentAnalysis(BaseModel): model_config = ConfigDict(extra='forbid') explanation: str sentiment: Sentiment chat_completion = client.chat.completions.create( model="zai-org/GLM-5.3-Flash", reasoning_effort="high", messages=[ {"role": "system", "content": "First describe the sentiment of the following review and then use that to classify the sentiment as positive or negative."}, {"role": "user", "content": "The staff went above and beyond! I had a great stay."} ], temperature=0.0, response_format={ "type": "json_schema", "json_schema": { "name": "SentimentAnalysis", "schema": SentimentAnalysis.model_json_schema() } } ) print(chat_completion.choices[0].message.content) ``` ## Schema-less JSON Mode We also support the schema-less JSON mode provided by OpenAI ```python import os from openai import OpenAI client = OpenAI( api_key=os.getenv("TELNYX_API_KEY"), base_url="https://api.telnyx.com/v2/ai/openai", ) chat_completion = client.chat.completions.create( model="zai-org/GLM-5.3-Flash", reasoning_effort="high", messages=[ {"role": "system", "content": "First describe the sentiment of the following review and then use that to classify the sentiment as positive or negative. Please respond using a JSON object."}, {"role": "user", "content": "The staff went above and beyond! I had a great stay."} ], temperature=0.0, response_format={"type": "json_object"} ) print(chat_completion.choices[0].message.content) ``` --- ### Audio Language Models > Source: https://developers.telnyx.com/docs/inference/audio-language-models.md In this tutorial, you'll learn how to: - Identify which Audio Language Models are available using our [models API](https://developers.telnyx.com/api-reference/openai-chat/get-available-models-openai-compatible) - Chat with open source Audio Language Models using our [chat completions API](/api-reference/openai-chat/create-a-chat-completion-openai-compatible) ## Getting started Audio Language Models are identified in our [models API](https://developers.telnyx.com/api-reference/openai-chat/get-available-models-openai-compatible) with a `task` type of `audio-text-to-text`. Audio is made available to the model in two main ways: - passing a link to the audio in a user message - passing the base64 encoded audio directly in a user message Make sure you have set the TELNYX_API_KEY environment variable ```python import base64 import os from openai import OpenAI import requests client = OpenAI( api_key=os.getenv("TELNYX_API_KEY"), base_url="https://api.telnyx.com/v2/ai/openai" ) def encode_audio_base64_from_url(audio_url: str) -> str: """Encode audio retrieved from a remote url to base64 format.""" with requests.get(audio_url) as response: response.raise_for_status() result = base64.b64encode(response.content).decode('utf-8') return f"data:audio/ogg;base64,{result}" def process_audio(audio_url, instructions): chat_completion = client.chat.completions.create( model="fixie-ai/ultravox-v0_4_1-llama-3_1-8b", messages=[ { "role": "system", "content": instructions }, { "role": "user", "content": [ { "type": "audio_url", "audio_url": { "url": audio_url, }, }, ], } ], ) return chat_completion.choices[0].message.content audio_url = "https://upload.wikimedia.org/wikipedia/commons/e/e0/Phrase_de_Neil_Armstrong.oga" audio_base64 = encode_audio_base64_from_url(audio_url) instructions = "Transcribe this verbatim. Do NOT respond with anything but the transcription." print(f"Transcribe link: {process_audio(audio_url, instructions)}") print(f"Transcribe base64: {process_audio(audio_base64, instructions)}") instructions = "Translate this to French. Be faithful to the original while sounding like a native French speaker. Do NOT respond with anything but the translation." print(f"Translate link: {process_audio(audio_url, instructions)}") ``` This will output something like the following for this audio Your browser does not support the audio element. ``` Transcribe link: That's one small step for man, one giant leap for mankind. Transcribe base64: That's one small step for man, one giant leap for mankind. Translate link: C'est un petit pas pour un homme, un grand pas pour l'humanité. ``` --- ### PR Reviewer - Github Action > Source: https://developers.telnyx.com/docs/inference/pr-reviewer.md ## Introduction Welcome to the PR Reviewer by Telnyx GitHub Action! This guide will teach you how to set up and use the PR Reviewer By Telnyx, which leverages open-source language models running on Telnyx GPUs to automatically review your pull requests. ## Prerequisites - [Sign up for a free Telnyx account](https://telnyx.com/sign-up) if you haven't already. ## Setup guide ### Step 1: Obtain Your Telnyx API Key 1. Log in to your [Telnyx account](https://portal.telnyx.com/). 2. Navigate to the **API Keys** section in the Telnyx portal. 3. Click on **Create API Key**. 4. Copy the generated API key and store it in a secure location. ### Step 2: Add Your Telnyx API Key as a Secret on GitHub 1. In your GitHub repository, go to **Settings** > **Secrets and variables** > **Actions**. 2. Click on **New repository secret**. 3. Name the secret `TELNYX_API_KEY`. 4. Paste your Telnyx API key in the **Value** field and click **Add secret**. ### Step 3: Create the GitHub workflow file To integrate the Telnyx PR Reviewer into your project, follow these steps: 1. In your repository, create a new file at `.github/workflows/review_pr.yml`. 2. Copy and paste the following configuration into the file: ```yaml name: PR Review on: pull_request: types: [opened, synchronize, reopened] permissions: pull-requests: write jobs: review: runs-on: ubuntu-latest steps: - name: PR Review uses: team-telnyx/reviewpr@main with: telnyx_api_key: ${{ secrets.TELNYX_API_KEY }} model_name: "zai-org/GLM-5.3-Flash" ``` 3. Commit the file to your repository. ### Step 4: Optional Configuration The `model_name` parameter in the workflow file is optional. If omitted, the action will use a default language model. If you wish to specify a different model, replace `'meta-llama/Meta-Llama-3.1-8B-Instruct'` with your desired model from the Telnyx [LLM Library](https://telnyx.com/products/llm-library). ## Core Concepts ### GitHub Actions GitHub Actions automate workflows directly in your GitHub repository. In this case, the PR Reviewer By Telnyx is triggered by pull request events, such as when a PR is opened or updated. ### Telnyx Inference API The PR Reviewer By Telnyx uses the Telnyx Inference API to analyze and review the content of pull requests. This API allows interaction with large language models (LLMs) hosted on Telnyx infrastructure. ### Model Selection Your choice of LLM will affect the quality and behavior of the reviews. You can experiment with different models from the Telnyx [LLM Library](https://telnyx.com/products/llm-library) to find the best fit for your project. ### Automatic PR Reviews Once configured, the PR Reviewer By Telnyx automatically generates a review for every pull request based on the content, providing suggestions or feedback powered by the chosen language model. ## Not sure how to get started? | I want to... | Relevant Tutorial | | :------------------------------ | :----------------------------------------------------------------- | | Learn more about GitHub Actions | [GitHub Actions Documentation](https://docs.github.com/en/actions) | | Explore more Telnyx models | [Telnyx LLM Library](https://telnyx.com/products/llm-library) | ## Additional references - Dive into our [Telnyx Inference API documentation](https://developers.telnyx.com/docs/inference) - Explore our full [API reference](/api-reference/openai-chat/create-a-chat-completion-openai-compatible) - Review our [OpenAI Compatibility Matrix](/docs/inference/openai) - Check out our [pricing page](https://telnyx.com/pricing/inference-api) --- ### AI SMS Outfit Recommender with OpenMeteo > Source: https://developers.telnyx.com/docs/inference/ai-outfit-recommender.md Today we will be making a fun script to text us a nice recommendation for outfits based on the weather every morning. It will look something like this at the end: ![Weather Recommendation Screenshot](/assets/images/telnyx-weather-sms-rec.jpg) This project has 3 main components: 1. Check the weather using the free OpenMeteo API 2. Pass the weather info to Telnyx Inference using the model of our choice 3. Send the recommendation to the user using Telnyx SMS Let's get started! ## Checking the weather with OpenMeteo [OpenMeteo](https://open-meteo.com/) is a great free API that allows you to retrieve the forecast for the current day. They have a ton of options for what you can retrieve, but for this demo, we will stick with just the temperature and weather code (although feel free to experiment with other things like humidity!). The following functions can be used to retrieve a weather_description that we can feed into our Telnyx Inference model: ```python def get_weather(latitude, longitude): url = f"https://api.open-meteo.com/v1/forecast?latitude={latitude}&longitude={longitude}¤t=temperature_2m,weathercode&temperature_unit=fahrenheit&timezone=auto" response = requests.get(url) data = response.json() if response.status_code == 200: current = data["current"] temperature = current["temperature_2m"] weathercode = current["weathercode"] weather_description = get_weather_description(weathercode) return f"Temperature: {temperature}°F, {weather_description}" else: return "Failed to fetch weather data" def get_weather_description(code): weather_codes = { 0: "Clear sky", 1: "Mainly clear", 2: "Partly cloudy", 3: "Overcast", 45: "Fog", 48: "Depositing rime fog", 51: "Light drizzle", 53: "Moderate drizzle", 55: "Dense drizzle", 61: "Slight rain", 63: "Moderate rain", 65: "Heavy rain", 71: "Slight snow fall", 73: "Moderate snow fall", 75: "Heavy snow fall", 77: "Snow grains", 80: "Slight rain showers", 81: "Moderate rain showers", 82: "Violent rain showers", 85: "Slight snow showers", 86: "Heavy snow showers", 95: "Thunderstorm", 96: "Thunderstorm with slight hail", 99: "Thunderstorm with heavy hail", } return weather_codes.get(code, "Unknown") ``` It looks like a lot of code, but the majority is just converting the weather codes to a nicely formatted string that the model will be able to understand. Notice also that the longitude and latitude are required to skip the need for an API key, so we will use Chicago's latitude and longitude, 41.9 and 87.6 for this example. ## Getting our recommendation text from Telnyx Inference Now that we have the current weather, let's make a call to Telnyx Inference to generate a text to send to our user. There are [many state-of-the-art open source models available through the Telnyx API](https://telnyx.com/products/llm-library), so for this one we will select [GLM-5.3 Flash from Zhipu AI](https://developers.telnyx.com/docs/inference/models). Let's write a function to retrieve a good weather recommendation from Telnyx Inference. We will just use the `requests` library to not add an additional pip requirement, but the Telnyx LLM API is also compatible with the OpenAI Python and JS SDKs, see the [OpenAI Migration Guide Here](https://developers.telnyx.com/docs/inference/openai). ```python def get_clothing_recommendation(weather_data): url = "https://api.telnyx.com/v2/ai/openai/chat/completions" payload = json.dumps( { "messages": [ { "role": "system", "content": "You are a helpful assistant that texts me every morning with a brief outfit recommendations based on the weather. Be friendly and brief.", }, { "role": "user", "content": f"The weather today is: {weather_data}. What should I wear?", }, ], "model": "zai-org/GLM-5.3-Flash", "reasoning_effort": "high", "max_tokens": 100, } ) headers = { "Content-Type": "application/json", "Accept": "application/json", "Authorization": f"Bearer {os.getenv('TELNYX_API_KEY')}", } response = requests.post(url, headers=headers, data=payload) return response.json()["choices"][0]["message"]["content"] ``` Make sure your `TELNYX_API_KEY` is set in your environment variables so that they can be loaded with `os.getenv('TELNYX_API_KEY')`. The API spec for chat completions can be found here if you would prefer to use HTTP requests instead of the OpenAI client or want to play around with some of the LLM parameters offered. The output of this function will be a string with our recommendation, for example: "Good morning. Perfect day ahead. Why not try a light, pastel-colored short-sleeved shirt, paired with some beige or light-gray shorts? Add some loafers or sneakers, and you're all set for a sunny day. Have a great one!" ## Sending our text to the user Great! Now that we have our weather and text recommendation, we can send the text to the user. Sending a message with Telnyx SMS is easy, follow the [tutorial here if you have not set up a Telnyx number yet](https://developers.telnyx.com/docs/messaging/messages/send-message). We can use the following snippet to send a text using Telnyx: ```python import telnyx telnyx.api_key = os.getenv("TELNYX_API_KEY") def send_sms(to_number, message): return telnyx.Message.create( from_=os.getenv("TELNYX_PHONE_NUMBER"), to=to_number, text=message, ) ``` ## Putting it all together Now that we have all the pieces in place, let's run the script! We can use the following sequence to chain everything together: ```python latitude = 40.7128 # Chicago latitude longitude = -74.0060 # Chicago longitude # Get weather data print("Getting weather data...") weather_description = get_weather(latitude, longitude) print(f"Description received: {weather_description}") to_number = "+1YOUR_DESTINATION_NUMBER" # Example phone number # Get clothing recommendation print("Getting clothing recommendation...") recommendation = get_clothing_recommendation(weather_description) print(f"Recommendation received: {recommendation}") full_text = f"{recommendation}\n\n{weather_description}" # Send SMS print("Sending SMS...") res = send_sms(to_number, full_text) print("SMS sent!") ``` And we see the following output: ``` Getting weather data... Description received: Temperature: 81.8°F, Clear sky Getting clothing recommendation... Recommendation received: Good morning. Perfect day ahead. Why not try a light, pastel-colored short-sleeved shirt, paired with some beige or light-gray shorts? Add some loafers or sneakers, and you're all set for a sunny day. Have a great one! Sending SMS... SMS sent! ``` Great work! We now have our script to send the user weather recommendations based on the weather. To improve this, test out different weather attributes such as humidity to influence the outfit recommendations, or adjust the prompt so the model knows what you generally like to wear or what you have in your closet. You could also run this script automatically every day at a certain time using a `cronjob` or the task scheduler of your choice. Thanks for following along! --- ## AI Gateway ### Overview > Source: https://developers.telnyx.com/docs/inference/ai-gateway.md AI Gateway issues **scoped inference credentials** for applications. Instead of sharing a Telnyx account API key with every service, create a token group, issue a token key inside it and hand that key to the application. The gateway enforces the model allowlist, budget and rate limits attached to the key, and records every request in a usage ledger attributed to the key, its user, its group and the end user it served. Applications call the gateway with the official OpenAI SDKs by changing the base URL and the API key. No request rewriting is required. Anthropic models used with your own Anthropic key can also be called with the official Anthropic SDKs. ## Capabilities | Capability | Behavior | | --- | --- | | Token groups | Define model access, group budgets and rate limits. | | Token users | Associate an application actor with one or more groups; aggregate its limits across all of its keys. | | Token keys | Issue scoped inference credentials for a user or a service; narrow model access, set limits and expiry, revoke access. | | End users | Apply account-scoped budget caps and blocks to a caller-asserted application user identifier. | | Provider keys (BYOK) | Store an OpenAI or Anthropic key once and attach it to groups; your provider bills requests on bring-your-own-key models. | | Inference | OpenAI Chat Completions, Anthropic Messages (Anthropic BYOK models), [model discovery](/docs/inference/ai-gateway/inference-api#models) and streaming. | | Guardrails | [Flag or block supported secrets and sensitive data](/docs/inference/ai-gateway/guardrails) with a policy on each token group. | | Usage | Durable request accounting with reservations, corrections and dimensional summaries. | ## Two planes, two credentials The gateway exposes a **management plane** for provisioning and reporting, and an **inference plane** that applications call. They use different hostnames and different credentials. | Plane | Base URL | Credential | | --- | --- | --- | | Management | `https://api.telnyx.com/v2/llm_token_gateway` | Telnyx account API key: `Authorization: Bearer $TELNYX_API_KEY` | | OpenAI-compatible inference | `https://llm.telnyx.com/v1` | AI Gateway token key: `Authorization: Bearer $AI_GATEWAY_TOKEN_KEY` | | Anthropic-compatible inference (Anthropic BYOK models) | `https://llm.telnyx.com` as the SDK base URL | Same token key via `x-api-key`; the SDK appends `/v1/messages` | Token keys start with `ltg_sk_`. The inference plane rejects Telnyx account API keys and provider secrets; the management plane rejects token keys. Keep every credential in a trusted backend or secret store, never in browser code, source control or logs. ## How it works 1. **Create a token group** with an explicit `allowed_models` list and optional budget and rate limits. To use your own OpenAI or Anthropic account, attach a [provider key](/docs/inference/ai-gateway/byok) to the group. 2. **Issue a token key** in that group, optionally bound to a token user. The secret is returned once, on the create response. 3. **Call models** from the application with an OpenAI SDK (or, for Anthropic BYOK models, an Anthropic SDK) pointed at the inference base URL and authenticated with the token key. 4. **Inspect usage** with the spend events and spend summary endpoints, filtered by group, user, key or end user. 5. **Revoke** the key when the application no longer needs it. New admissions stop immediately; spend history is retained. Telnyx owns authorization, admission, revocation and the authoritative usage ledger. Applications never hold provider credentials or configure model providers directly. Usage of Telnyx-hosted models is billed to your Telnyx account at standard Telnyx AI Inference pricing for each model. Requests on bring-your-own-key models are billed by your provider on your provider account. Budgets and reported `cost` values use a flat reference rate for enforcement and attribution; see [Budgets](/docs/inference/ai-gateway/controls#budgets). ## Next steps Create a group, issue a token key, make a request and revoke the key. Telnyx-hosted and BYOK models, OpenAI and Anthropic SDK configuration, streaming, request limits and supported fields. Groups, users, keys, end users and provider keys, with idempotency and ETag rules. How each control is enforced and what happens at the limit. Group policies, prompt and response inspection, streaming and findings. Use your own OpenAI or Anthropic key; your provider bills those requests. Spend events, dimensional summaries and snapshot pagination. Status codes, structured error codes and how to handle them. --- ### Quickstart > Source: https://developers.telnyx.com/docs/inference/ai-gateway/quickstart.md This quickstart provisions a group and a service token key through the management API, makes an inference request with that key, reads the resulting usage and revokes the key. Every step is a copy-paste request. ## Prerequisites - A Telnyx API key from the [portal](https://portal.telnyx.com/#/api-keys). - A model name from the [available models](/docs/inference/ai-gateway/inference-api#models). This guide uses `Kimi-K3`. The inference `GET /v1/models` endpoint is scoped to a token key, so it cannot be used to discover models before a key exists. - `curl`, `jq` and `uuidgen` (or another way to generate a UUID for the `Idempotency-Key` header). Export the credential, the model name and the two base URLs so the examples work as-is: ```bash export TELNYX_API_KEY="KEY..." export AI_GATEWAY_MODEL="Kimi-K3" export AI_GATEWAY_MANAGEMENT_BASE_URL="https://api.telnyx.com/v2/llm_token_gateway" export AI_GATEWAY_INFERENCE_BASE_URL="https://llm.telnyx.com/v1" ``` This walkthrough creates persistent account resources and can incur model charges. A token group defines which models its keys may call and the limits they share. `name` and `allowed_models` are required. An empty `allowed_models` list permits no inference. This example sets a USD 10 budget per anchored one-day period and 60 requests per minute. ```bash curl -X POST "$AI_GATEWAY_MANAGEMENT_BASE_URL/token_groups" \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: $(uuidgen)" \ -d "$(jq -n --arg model "$AI_GATEWAY_MODEL" '{ name: "support-assistant", allowed_models: [$model], max_budget: 10, budget_duration: "1d", rpm_limit: 60 }')" ``` ```json { "data": { "record_type": "token_group", "id": "5f1c9d2e-7b3a-4c8e-9f21-0a6d4e8b1c33", "name": "support-assistant", "allowed_models": ["Kimi-K3"], "max_budget": 10, "budget_duration": "1d", "rpm_limit": 60, "tpm_limit": null, "provider_key_ids": [], "blocked": false, "spend": 0, "reserved_spend": 0, "budget_started_at": "2026-09-22T10:00:00Z", "resets_at": "2026-09-23T10:00:00Z", "version": 1, "created_at": "2026-09-22T10:00:00Z", "updated_at": "2026-09-22T10:00:00Z" } } ``` Keep `data.id`; the next step needs it and the group is what you retire at the end. ```bash export GROUP_ID="5f1c9d2e-7b3a-4c8e-9f21-0a6d4e8b1c33" ``` Every mutation requires an `Idempotency-Key`. If a request times out, retry it with the **same** key and body; a new key starts a different operation and can create a second resource. A token key is the credential the application uses. `token_user_id: null` creates a **service key** owned by the group alone. Null `allowed_models` inherits the group's list. Key limits that are omitted are set to their maximums: USD 1,000 lifetime budget, 6,000 requests and 10,000,000 tokens per minute. See [Token key limits](/docs/inference/ai-gateway/controls#token-key-limits). ```bash curl -X POST "$AI_GATEWAY_MANAGEMENT_BASE_URL/token_keys" \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: $(uuidgen)" \ -d "$(jq -n --arg group "$GROUP_ID" '{ name: "support-backend", token_group_id: $group, token_user_id: null }')" ``` ```json { "data": { "record_type": "token_key", "id": "9b7e4a10-3c2d-4f5e-8a6b-1d2c3e4f5a60", "name": "support-backend", "token_group_id": "5f1c9d2e-7b3a-4c8e-9f21-0a6d4e8b1c33", "token_user_id": null, "allowed_models": null, "max_budget": 1000, "budget_duration": null, "rpm_limit": 6000, "tpm_limit": 10000000, "expires_at": null, "blocked": false, "required_end_user_id": false, "spend": 0, "reserved_spend": 0, "budget_started_at": null, "resets_at": null, "version": 1, "created_at": "2026-09-22T10:01:00Z", "updated_at": "2026-09-22T10:01:00Z", "token": "ltg_sk_..." } } ``` `data.token` is returned **only on the original create response**. GET, list and idempotent replays return metadata without the secret. If the response is lost, revoke the key and create a new one with a new idempotency key. Store the token in your secret store now, then export it and the key ID for the remaining steps: ```bash export TOKEN_KEY_ID="9b7e4a10-3c2d-4f5e-8a6b-1d2c3e4f5a60" export AI_GATEWAY_TOKEN_KEY="ltg_sk_..." ``` To issue a key for a specific application user instead, first `POST /token_users` with `name` and `token_group_ids: [GROUP_ID]`, then pass the returned ID as `token_user_id`. Key ownership cannot be changed later; issue a replacement key instead. The inference plane authenticates with the token key, not the account API key. `GET /v1/models` returns only the models this key may use. ```bash curl "$AI_GATEWAY_INFERENCE_BASE_URL/models" \ -H "Authorization: Bearer $AI_GATEWAY_TOKEN_KEY" ``` Send a Chat Completions request with one of those models: ```bash curl curl -X POST "$AI_GATEWAY_INFERENCE_BASE_URL/chat/completions" \ -H "Authorization: Bearer $AI_GATEWAY_TOKEN_KEY" \ -H "Content-Type: application/json" \ -d "$(jq -n --arg model "$AI_GATEWAY_MODEL" '{ model: $model, messages: [{role: "user", content: "Hello"}], max_tokens: 32 }')" ``` ```python Python import os from openai import OpenAI client = OpenAI( api_key=os.environ["AI_GATEWAY_TOKEN_KEY"], base_url=os.environ["AI_GATEWAY_INFERENCE_BASE_URL"], max_retries=0, ) completion = client.chat.completions.create( model=os.environ["AI_GATEWAY_MODEL"], messages=[{"role": "user", "content": "Hello"}], max_tokens=32, ) print(completion.choices[0].message.content) ``` ```javascript JavaScript import OpenAI from "openai"; const client = new OpenAI({ apiKey: process.env.AI_GATEWAY_TOKEN_KEY, baseURL: process.env.AI_GATEWAY_INFERENCE_BASE_URL, maxRetries: 0, }); const completion = await client.chat.completions.create({ model: process.env.AI_GATEWAY_MODEL, messages: [{ role: "user", content: "Hello" }], max_tokens: 32, }); console.log(completion.choices[0].message.content); ``` Inference requests are **not idempotent**. The examples disable SDK retries so that a timeout cannot silently trigger a second billed request. Set `max_tokens` on every request: a request without it reserves the model's full output allowance against budgets and `tpm_limit`. See [Inference API](/docs/inference/ai-gateway/inference-api) for streaming and request size limits. Spend events are queried over a half-open UTC date range `[start_date, end_date)` of at most 31 days, using ISO dates rather than timestamps. Filter by the key you just used: ```bash curl --globoff -G "$AI_GATEWAY_MANAGEMENT_BASE_URL/spend/events" \ -H "Authorization: Bearer $TELNYX_API_KEY" \ --data-urlencode "start_date=2026-09-22" \ --data-urlencode "end_date=2026-09-23" \ --data-urlencode "token_key_id=$TOKEN_KEY_ID" \ --data-urlencode "page[size]=100" ``` ```json { "data": [ { "id": "c0a1b2c3-d4e5-4f60-8a71-92b3c4d5e6f7", "request_id": "0b1c2d3e-4f50-4617-8a29-3b4c5d6e7f80", "created_at": "2026-09-22T10:02:14Z", "token_group_id": "5f1c9d2e-7b3a-4c8e-9f21-0a6d4e8b1c33", "token_user_id": null, "token_key_id": "9b7e4a10-3c2d-4f5e-8a6b-1d2c3e4f5a60", "end_user_id": null, "model": "Kimi-K3", "input_tokens": 9, "output_tokens": 12, "cost": 0.000225, "status": "succeeded", "usage_status": "known", "configuration_version": 1, "rate_version": "telnyx-reference-v1", "reservation_micro_usd": 0 } ], "meta": { "page_number": 1, "page_size": 100, "has_more": false, "snapshot": "..." } } ``` When `meta.has_more` is true, request the next `page[number]` with the same filters and `page[snapshot]=`. See [Usage reporting](/docs/inference/ai-gateway/usage) for summaries grouped by group, user, key or end user. `DELETE` requires the resource's current `ETag` in `If-Match`. Read the key first; the metadata GET never reveals the token. ```bash KEY_ETAG=$(curl -sS -I "$AI_GATEWAY_MANAGEMENT_BASE_URL/token_keys/$TOKEN_KEY_ID" \ -H "Authorization: Bearer $TELNYX_API_KEY" \ | awk 'tolower($1) == "etag:" { sub(/\r$/, "", $2); print $2 }') curl -X DELETE "$AI_GATEWAY_MANAGEMENT_BASE_URL/token_keys/$TOKEN_KEY_ID" \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Idempotency-Key: $(uuidgen)" \ -H "If-Match: $KEY_ETAG" ``` A `204` means the key is revoked: new requests with that token are rejected, while requests already admitted may finish. Spend history is retained. On `412`, the resource changed since the read; fetch it again rather than forcing the write. The group stays available for new keys. To retire it as well, `DELETE /token_groups/{id}` with a fresh ETag and idempotency key. Deleting a group cascades to its keys. ## Next: use your own provider key To run OpenAI or Anthropic models on your own provider account, store the provider key once, attach it to a group and allow a [bring-your-own-key model](/docs/inference/ai-gateway/inference-api#bring-your-own-key-models). Applications keep using the same token key. See [Bring your own key](/docs/inference/ai-gateway/byok). ## Next steps Available models, streaming, request limits and the supported request fields. Users, end users, PATCH semantics and pagination. What each limit does and how it is enforced. Status codes and structured error codes on both planes. Attach your own OpenAI or Anthropic key to a group. --- ### Inference API > Source: https://developers.telnyx.com/docs/inference/ai-gateway/inference-api.md The inference plane serves three endpoints. All of them authenticate with an AI Gateway token key (`ltg_sk_...`) issued through the [management API](/docs/inference/ai-gateway/management-api). Telnyx account API keys, provider secrets and any other credential are rejected on this plane. | Endpoint | Purpose | | --- | --- | | `GET https://llm.telnyx.com/v1/models` | Models available to the supplied token key. | | `POST https://llm.telnyx.com/v1/chat/completions` | OpenAI-compatible Chat Completions, including streaming. Works with every model. | | `POST https://llm.telnyx.com/v1/messages` | Anthropic-compatible Messages, including streaming. Only for [Anthropic BYOK models](#anthropic-sdk). | Compatibility is bounded by the AI Gateway contract, not by every option the OpenAI and Anthropic SDKs expose; see [Supported request fields](#supported-request-fields) and [Supported Messages fields](#supported-messages-fields). ## Models These Telnyx-hosted models are billed to your Telnyx account. Use these model names in the `model` field: - `Kimi-K3` - `Kimi-K2.6` - `Kimi-K2.5` - `GLM-5.3` - `GLM-5.3-Flash` - `GLM-5.2` - `GLM-5.1-FP8` - `MiniMax-M3-MXFP8` - `MiniMax-M2.7` - `Qwen3.8-27B` - `Qwen3-235B-A22B` - `DeepSeek-V4.1-Flash` - `DeepSeek-V4-Flash-0731` - `Llama-3.3-70B-Instruct` - `Meta-Llama-3.1-70B-Instruct` - `Meta-Llama-3.1-8B-Instruct` - `gemma-2b-it` A group lists the model names its keys may call in `allowed_models`; a key can narrow that list further. Per-request output and prompt limits are listed under [Request size limits](#request-size-limits). ### Bring-your-own-key models These models run on your own provider account with a [provider key](/docs/inference/ai-gateway/byok) and are billed by that provider. | Provider | Models | | --- | --- | | Anthropic | `claude-fable-5-1`, `claude-opus-5-5`, `claude-sonnet-5` | | OpenAI | `gpt-6-astra`, `gpt-6-sol`, `gpt-6-luna`, `gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna`, `gpt-5.5`, `gpt-5.4`, `gpt-5.4-mini`, `gpt-5.4-nano`, `gpt-4.1`, `gpt-4.1-mini`, `gpt-4o`, `gpt-4o-mini`, `o3` | Add them to `allowed_models` like any other model. A BYOK model works only when the calling key's group has an attached provider key for that model's provider. Without one, requests fail with `503` and code `enforcement_unavailable`; they never fall back to a Telnyx-hosted model. Every BYOK model works on `/v1/chat/completions`. Anthropic BYOK models also work on `/v1/messages`. ### Model discovery `GET /v1/models` returns the models the supplied key may use: the intersection of the key's `allowed_models` and its group's list. It is not a global catalog, and it does not charge against any budget. ```bash curl https://llm.telnyx.com/v1/models \ -H "Authorization: Bearer $AI_GATEWAY_TOKEN_KEY" ``` ```json { "object": "list", "data": [ { "id": "Kimi-K3", "object": "model", "created": 1758535200, "owned_by": "telnyx" } ] } ``` ## OpenAI SDK Point the OpenAI SDK at the inference base URL and use the token key as the API key. ```python Python import os from openai import OpenAI client = OpenAI( api_key=os.environ["AI_GATEWAY_TOKEN_KEY"], base_url="https://llm.telnyx.com/v1", max_retries=0, ) completion = client.chat.completions.create( model=os.environ["AI_GATEWAY_MODEL"], messages=[{"role": "user", "content": "Tell me about Telnyx"}], max_tokens=256, user="application-bound-user-id", ) print(completion.choices[0].message.content) ``` ```javascript JavaScript import OpenAI from "openai"; const client = new OpenAI({ apiKey: process.env.AI_GATEWAY_TOKEN_KEY, baseURL: "https://llm.telnyx.com/v1", maxRetries: 0, }); const completion = await client.chat.completions.create({ model: process.env.AI_GATEWAY_MODEL, messages: [{ role: "user", content: "Tell me about Telnyx" }], max_tokens: 256, user: "application-bound-user-id", }); console.log(completion.choices[0].message.content); ``` ```bash curl curl -X POST https://llm.telnyx.com/v1/chat/completions \ -H "Authorization: Bearer $AI_GATEWAY_TOKEN_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "Kimi-K3", "messages": [{"role": "user", "content": "Tell me about Telnyx"}], "max_tokens": 256, "user": "application-bound-user-id" }' ``` The optional `user` field attributes the request to an application end user for budgeting and reporting. See [End-user identity](/docs/inference/ai-gateway/controls#end-user-identity). ### Streaming Set `stream: true` and consume the stream to completion. A stream that starts with HTTP 200 can still fail later; handle SDK exceptions and close the stream. ```python Python import os from openai import OpenAI client = OpenAI( api_key=os.environ["AI_GATEWAY_TOKEN_KEY"], base_url="https://llm.telnyx.com/v1", max_retries=0, ) with client.chat.completions.create( model=os.environ["AI_GATEWAY_MODEL"], messages=[{"role": "user", "content": "Tell me about Telnyx"}], max_tokens=256, stream=True, ) as stream: for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="", flush=True) ``` ```javascript JavaScript const stream = await client.chat.completions.create({ model: process.env.AI_GATEWAY_MODEL, messages: [{ role: "user", content: "Tell me about Telnyx" }], max_tokens: 256, stream: true, }); for await (const chunk of stream) { process.stdout.write(chunk.choices[0]?.delta?.content ?? ""); } ``` ## Anthropic SDK `POST /v1/messages` is available **only for Anthropic models used with your own Anthropic key** ([Anthropic BYOK models](#bring-your-own-key-models)). A request to `/v1/messages` with a Telnyx-hosted model or an OpenAI BYOK model returns `400` with code `invalid_request`; call those models on `/v1/chat/completions`. Anthropic BYOK models work on both endpoints. The Anthropic SDK appends `/v1/messages` to its base URL, so set the base URL to the gateway root **without** the `/v1` suffix. The SDK sends the token key in `x-api-key`, which the gateway accepts. `Authorization: Bearer` is also accepted; if both headers are present they must carry the same token. ```python Python import os from anthropic import Anthropic client = Anthropic( api_key=os.environ["AI_GATEWAY_TOKEN_KEY"], base_url="https://llm.telnyx.com", max_retries=0, ) message = client.messages.create( model="claude-sonnet-5", max_tokens=256, messages=[{"role": "user", "content": "Tell me about Telnyx"}], metadata={"user_id": "application-bound-user-id"}, ) print(message.content[0].text) ``` ```javascript JavaScript import Anthropic from "@anthropic-ai/sdk"; const client = new Anthropic({ apiKey: process.env.AI_GATEWAY_TOKEN_KEY, baseURL: "https://llm.telnyx.com", maxRetries: 0, }); const message = await client.messages.create({ model: "claude-sonnet-5", max_tokens: 256, messages: [{ role: "user", content: "Tell me about Telnyx" }], metadata: { user_id: "application-bound-user-id" }, }); console.log(message.content[0].text); ``` ```bash curl curl -X POST https://llm.telnyx.com/v1/messages \ -H "x-api-key: $AI_GATEWAY_TOKEN_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-sonnet-5", "max_tokens": 256, "messages": [{"role": "user", "content": "Tell me about Telnyx"}], "metadata": {"user_id": "application-bound-user-id"} }' ``` Raw HTTP requests must send `anthropic-version: 2023-06-01`; the SDKs add it automatically. `metadata.user_id` and the OpenAI `user` field identify the same end-user namespace. The endpoint streams standard Anthropic server-sent events (`message_start`, `content_block_start`, `content_block_delta`, `content_block_stop`, `message_delta`, `message_stop`). An error after the stream has started arrives as an `event: error` frame. ```python Python with client.messages.stream( model="claude-sonnet-5", max_tokens=256, messages=[{"role": "user", "content": "Tell me about Telnyx"}], ) as stream: for text in stream.text_stream: print(text, end="", flush=True) ``` ```javascript JavaScript const stream = client.messages.stream({ model: "claude-sonnet-5", max_tokens: 256, messages: [{ role: "user", content: "Tell me about Telnyx" }], }); for await (const event of stream) { if (event.type === "content_block_delta" && event.delta.type === "text_delta") { process.stdout.write(event.delta.text); } } ``` ### Supported Messages fields | Field | Notes | | --- | --- | | `model` | Required. An Anthropic BYOK model available to this key. | | `messages` | Required. 1 to 256 messages with roles `user` and `assistant`. | | `max_tokens` | Required. At most 16,384. | | `system` | String or array of `text` blocks. | | `stream` | Anthropic server-sent events. | | `temperature` | 0 to 1. | | `top_p` | 0 to 1. | | `top_k` | 0 to 1000. | | `stop_sequences` | Up to 4. | | `metadata` | Only `user_id` is accepted; it maps to the end-user dimension. | | `tools`, `tool_choice` | Tool definitions with `name`, `description` and `input_schema`. `tool_choice` accepts `auto`, `any`, `tool`, `none` or an object. | Content blocks of type `text`, `tool_use` and `tool_result` are supported. `image` blocks are rejected. OpenAI-specific fields (`frequency_penalty`, `presence_penalty`, `n`, `seed`, `response_format`, `user`, `stream_options`, `parallel_tool_calls`) are rejected on the Messages endpoint. ## Retries and timeouts Inference is **not idempotent**. An SDK retry after a timeout or disconnect can dispatch a second billed request, and a timeout is not evidence that the first request did no provider work. The examples on this page set `max_retries=0`; if your application retries, do so deliberately and only for errors that occurred before dispatch (for example `401`, `403` or `429` with `Retry-After`). See [Errors](/docs/inference/ai-gateway/errors). Client-side timeouts are a client setting and do not extend the gateway's server-side request lifetime. ## Request size limits | Limit | Telnyx-hosted models | `gemma-2b-it` | BYOK models | | --- | --- | --- | --- | | Output tokens per request | 8,192 | 2,048 | 16,384 | | Prompt size | About 32k tokens | About 8k tokens | About 128k tokens (about 96k for `gpt-4o` and `gpt-4o-mini`) | Prompt size is measured conservatively from the request size, so the usable limit can be somewhat lower than the model's tokenizer would count. A request over either limit returns `400`. Set `max_tokens` (or `max_completion_tokens`) on every request, especially when the key, user or group has a low `tpm_limit`. A request without it reserves the model's full output allowance against budgets and `tpm_limit`, and can be rate-limited even when the actual response would be short. See [Reservations](/docs/inference/ai-gateway/controls#reservations). ## Supported request fields Unknown top-level fields are rejected with `400`. Options that a specific model does not support are also rejected with `400` rather than silently ignored. ### Chat Completions | Field | Notes | | --- | --- | | `model` | Required. Must be a model returned by `GET /v1/models` for this key. | | `messages` | Required. Roles: `system`, `developer`, `user`, `assistant`, `tool`. Content is a string or an array of `text` and `image_url` parts. | | `max_tokens` / `max_completion_tokens` | Send one or the other, never both. At most 8,192 for Telnyx-hosted models (2,048 for `gemma-2b-it`) and 16,384 for BYOK models. Strongly recommended; see [Request size limits](#request-size-limits). | | `stream`, `stream_options` | Server-sent events. | | `temperature` | 0 to 2. | | `top_p` | 0 to 1. | | `stop` | String or up to 4 strings. | | `seed`, `presence_penalty`, `frequency_penalty`, `n` | Passed through where the model supports them. | | `user` | End-user identifier for budgeting and reporting. | | `tools`, `tool_choice`, `parallel_tool_calls` | Function tools. `tool_choice` accepts `none`, `auto`, `required` or `{"type": "function", "function": {"name": ...}}`. | | `response_format` | Set `type` to `json_object` for a JSON object, or `json_schema` with `json_schema.name`, `json_schema.schema` and optional `json_schema.strict` for a supplied schema. Support depends on the model and `stream`; see below. | **Structured JSON on Telnyx-hosted models:** All 17 models support `json_schema` without streaming. All except `Meta-Llama-3.1-8B-Instruct` support `json_object`. Both formats support `stream: true` where enabled, except on `Llama-3.3-70B-Instruct`. Unsupported combinations return `400` before inference. `json_object` does not enforce requested fields or values. Use `json_schema` when a specific structure is required. Schemas are limited to 16 KiB, 16 levels and 512 schema nodes; remote references are rejected. Handle truncation, refusals and stream errors, and validate the completed output. Image parts must be inline `data:image/...;base64,...` URLs (PNG, JPEG, GIF or WebP). Remote image URLs are rejected. ### Not supported - Custom provider URLs, provider credentials, routing or fallback controls in the request body. - Remote image URLs. - Unlisted beta headers and options. - The Responses API, embeddings, image and audio generation, batches and realtime APIs. These are not part of the AI Gateway surface; see the [Telnyx Inference API](/docs/inference/getting-started) for those capabilities. --- ### Management API > Source: https://developers.telnyx.com/docs/inference/ai-gateway/management-api.md The management plane lives at `https://api.telnyx.com/v2/llm_token_gateway` and authenticates with a Telnyx account API key: ```text Authorization: Bearer $TELNYX_API_KEY ``` All paths on this page are relative to that base. Every resource is scoped to the authenticated account; a resource that belongs to another account returns `404`. ## Conventions ### Response envelope Success bodies wrap the resource in `data`. List responses add `meta` with pagination state. Every response carries an `X-Request-ID` correlation header; keep it when reporting a problem. Responses are `Cache-Control: no-store`. ### Idempotency Every `POST`, `PATCH`, `PUT` and `DELETE` requires an `Idempotency-Key` header. Keys are scoped to account, method and path and retained for 24 hours. - The same key with the same body returns the original outcome. - The same key with a different body returns `409` with code `idempotency_conflict`. - A key whose original request is still in progress returns `409` with `Retry-After`. - A replayed token-key create returns `200` with metadata only; the secret is never redisclosed. A provider-key secret is never returned, including on replay. When a mutation times out, retry it with the **same** key and body. Generating a fresh key starts a new operation and can create a duplicate resource. ### ETag preconditions Every resource carries an integer `version`, returned as a quoted `ETag` header on `GET`, create and update responses. `PATCH` and `DELETE` require that value in `If-Match`: | Condition | Response | | --- | --- | | `If-Match` matches the current version | The mutation is applied. | | `If-Match` is stale | `412` with code `precondition_failed`. Read the resource again. | | `If-Match` is missing | `428` with code `precondition_required`. | End-user caps use `PUT` as a full replacement: send `If-None-Match: *` to create and the current `If-Match` to replace. ### PATCH semantics - A field omitted from a `PATCH` body is preserved. - A nullable field set to `null` is cleared (for example `max_budget: null` removes a group or user cap). On a token key, a null limit is reset to its maximum instead; see [Token keys](#token-keys). - Policy changes apply to new admissions. They do not erase spend history or unresolved exposure, and a request already admitted under the old policy may finish. - A change can take a short time to propagate. During that window the management response is already committed, but inference may return `503` with code `enforcement_unavailable`. Do not treat a pending change as permission to rely on the old policy. ### Pagination Resource list endpoints accept `page[number]`, `page[size]` (1 to 100) and `page[snapshot]`. [Guardrail events](/docs/inference/ai-gateway/guardrails#query-findings) use offset pagination without a snapshot. ```json { "data": [ ... ], "meta": { "page_number": 1, "page_size": 100, "has_more": true, "snapshot": "..." } } ``` The first page returns a `meta.snapshot` bound to the account and filters and valid for 15 minutes. To fetch later pages, increment `page[number]` and pass `page[snapshot]=` with the **same filters**. A missing, expired or mismatched snapshot returns `409`. ## Token groups A group defines the model allowlist and shared limits for the keys inside it. | Operation | Path | | --- | --- | | Create | `POST /token_groups` | | List | `GET /token_groups` | | Read | `GET /token_groups/{id}` | | Update | `PATCH /token_groups/{id}` | | Delete | `DELETE /token_groups/{id}` | | Field | Type | Notes | | --- | --- | --- | | `name` | string | Required. 1 to 256 characters. | | `allowed_models` | string[] | Required. Up to 1000 unique [model names](/docs/inference/ai-gateway/inference-api#models). An empty list permits no inference. | | `max_budget` | number or null | USD, at most six decimal places. Null is uncapped; zero denies paid requests. | | `budget_duration` | `1d`, `7d`, `30d` or null | Anchored budget period. Null means a lifetime budget. | | `rpm_limit`, `tpm_limit` | integer or null | Requests and tokens per rolling 60-second window. Null removes the limit; zero denies. | | `provider_key_ids` | string[] | [Provider keys](#provider-keys) for [BYOK models](/docs/inference/ai-gateway/inference-api#bring-your-own-key-models). Defaults to `[]`. At most one key per provider; each key must belong to the same account (otherwise `404`) and its provider must match at least one BYOK model in `allowed_models` (otherwise `400`). Two keys for the same provider return `409`. | | `blocked` | boolean | Blocks every key in the group. | | `guardrails` | object or null | Group policy for secret and sensitive-data inspection. Defaults to null (disabled). See [Guardrails](/docs/inference/ai-gateway/guardrails) for actions, DLP profiles and streaming behavior. | Read-only fields on the response: `id`, `version`, `spend`, `reserved_spend`, `budget_started_at`, `resets_at`, `created_at`, `updated_at`. Deleting a group revokes its keys and removes user memberships; spend history is retained. ## Token users A user represents an application actor that may belong to more than one group. Limits set on the user aggregate across all of its keys in every group. | Operation | Path | | --- | --- | | Create | `POST /token_users` | | List | `GET /token_users` | | Read | `GET /token_users/{id}` | | Update | `PATCH /token_users/{id}` | | Delete | `DELETE /token_users/{id}` | | Field | Type | Notes | | --- | --- | --- | | `name` | string | Required. | | `token_group_ids` | string[] | Required. Groups this user may hold keys in. | | `external_id` | string or null | Your own identifier for the actor. | | `max_budget`, `budget_duration` | | Aggregate budget across the user's keys. | | `rpm_limit`, `tpm_limit` | | Aggregate rate limits across the user's keys. | Removing a group from `token_group_ids` while the user still holds active keys in that group returns `409`; revoke those keys first. Deleting a user revokes its keys and retains spend history. ## Token keys A key is the credential an application presents to the inference plane. | Operation | Path | | --- | --- | | Create | `POST /token_keys` | | List | `GET /token_keys` (filters: `token_group_id`, `token_user_id`) | | Read | `GET /token_keys/{id}` | | Update | `PATCH /token_keys/{id}` | | Revoke | `DELETE /token_keys/{id}` | | Field | Type | Notes | | --- | --- | --- | | `name` | string | Required. | | `token_group_id` | uuid | Required. Cannot be changed after creation. | | `token_user_id` | uuid or null | Null creates a service key. The user must be a member of the group. Cannot be changed after creation. | | `allowed_models` | string[] or null | Null inherits the group's list. An empty list denies every model. A non-empty list narrows the group's list. | | `max_budget` | number or null | USD budget scoped to this key, above 0 and at most 1000, with at most six decimal places. Omitted or null is stored as 1000. | | `budget_duration` | `1d`, `7d`, `30d` or null | Null means a lifetime budget. | | `rpm_limit` | integer or null | 1 to 6000. Omitted or null is stored as 6000. | | `tpm_limit` | integer or null | 1 to 10000000. Omitted or null is stored as 10000000. | | `expires_at` | date-time or null | After this instant the key is rejected. | | `blocked` | boolean | Rejects new requests without deleting the key. | | `required_end_user_id` | boolean | Requires a non-empty `user` / `metadata.user_id` on every request. Presence only, not authenticity. | Key limits differ from group and user limits: a key is never uncapped and cannot be denied with a zero limit. A value of 0 or above the maximum returns `400` with code `limit_out_of_range`; a negative, fractional or non-numeric value, or a budget with more than six decimal places, returns `400` with code `invalid_request`. On a blocked key, a null limit is kept until the key is unblocked. To deny a key, set `blocked: true` or revoke it. See [Token key limits](/docs/inference/ai-gateway/controls#token-key-limits). The create response is the **only** place `data.token` appears. It matches `^ltg_sk_[A-Za-z0-9_-]+$`. Store it immediately; a lost token cannot be recovered, only replaced. `DELETE` revokes the key. Acknowledged revocation blocks new admissions; requests already admitted may complete. ## End users An end-user cap applies an account-scoped budget or block to a caller-asserted identifier: the value an application sends as OpenAI `user` or Anthropic `metadata.user_id`. The identifier is the resource ID. | Operation | Path | | --- | --- | | List | `GET /end_users` | | Read | `GET /end_users/{id}` | | Create or replace | `PUT /end_users/{id}` | | Delete | `DELETE /end_users/{id}` | | Field | Type | Notes | | --- | --- | --- | | `max_budget` | number or null | Required in the body. Null is uncapped. | | `budget_duration` | `1d`, `7d`, `30d` or null | Required in the body. | | `blocked` | boolean | Required in the body. Denies every request carrying this identifier. | `PUT` is a full replacement and returns `200` for both create and replace. Send `If-None-Match: *` to create and the current `If-Match` to replace; a missing precondition returns `428`. End-user identifiers are assertions made by whoever holds the token key. Bind them to authenticated users in a trusted backend; see [End-user identity](/docs/inference/ai-gateway/controls#end-user-identity). ## Provider keys A provider key stores your own OpenAI or Anthropic secret for [bring your own key](/docs/inference/ai-gateway/byok). Attach it to groups through `provider_key_ids`. | Operation | Path | | --- | --- | | Create | `POST /provider_keys` | | List | `GET /provider_keys` | | Read | `GET /provider_keys/{id}` | | Delete | `DELETE /provider_keys/{id}` | | Field | Type | Notes | | --- | --- | --- | | `name` | string | Required. 1 to 256 characters. | | `provider` | `openai` or `anthropic` | Required. Any other value returns `400`. URLs are never accepted. | | `secret` | string | Required on create, write-only. Never returned by any response, list or replay. | Read-only fields on the response: `id`, `version`, `created_at`, `updated_at`. Provider keys cannot be edited; `PATCH` returns `405`. To change a secret, create a new provider key, attach it to the groups, then delete the old one. `DELETE` requires the current ETag in `If-Match`, like other deletes. Deleting a provider key detaches it from every group that references it; requests already in progress complete. ## Usage `GET /spend/events` and `GET /spend/summary` report the requests attributed to these resources. See [Usage reporting](/docs/inference/ai-gateway/usage). ## Guardrail findings `GET /guardrail_events` lists account-scoped flagged and blocked findings. Filter by date range, group, key, end user, stage or outcome. This endpoint uses offset pagination without `page[snapshot]`; see [Query findings](/docs/inference/ai-gateway/guardrails#query-findings). --- ### Budgets & rate limits > Source: https://developers.telnyx.com/docs/inference/ai-gateway/controls.md Every inference request passes an admission check before it is dispatched to a model. Admission evaluates the token key, its user, its group and the asserted end user together; the request is denied if any scope fails. This page describes each control and its observable behavior. ## Enforcement scopes | Scope | Set on | Applies to | | --- | --- | --- | | Key | `POST /token_keys` | Requests made with that key. | | User | `POST /token_users` | All keys owned by that user, across every group it belongs to. | | Group | `POST /token_groups` | All keys in the group. | | End user | `PUT /end_users/{id}` | Every request in the account that asserts that end-user identifier. | Model access is the intersection of the group's `allowed_models` and the key's `allowed_models` (when set). A blocked key, user, group or end user denies the request with `403`. ## Budgets Budgets are USD amounts with at most six decimal places, enforced independently at the key, user, group and end-user scopes. - On groups, users and end users, `max_budget: null` is uncapped and `max_budget: 0` denies every paid request. Token keys differ; see [Token key limits](#token-key-limits). - `budget_duration` of `1d`, `7d` or `30d` starts an anchored period at the moment the budget is committed. Periods roll from that anchor, not from midnight or the calendar month. A null duration is a lifetime budget. - Changing the amount keeps the current period. Changing the duration starts a new anchored period. Neither change erases history. - The response fields `spend`, `reserved_spend`, `budget_started_at` and `resets_at` on each resource show the current period. ### Reservations Before a request is dispatched, the gateway reserves a conservative upper bound on its cost from the input size and `max_tokens`. If any scope lacks that much headroom, the request is denied with `403` and code `budget_exceeded` (or `end_user_budget_exceeded`). The reservation is not shrunk to fit. After the response completes, the reservation is replaced by the actual cost. If the outcome is unknown, for example because the stream was interrupted before usage was reported, the reservation is retained as exposure and the spend event reports `cost: null`. Unknown usage is not zero usage, and a timeout is not a refund. Budgets are an application control, not an absolute guarantee of spend. ### Budgets and billing - Usage of Telnyx-hosted models is billed to your Telnyx account at standard [Telnyx AI Inference pricing](https://telnyx.com/pricing/inference-api) for each model. - Requests on [bring-your-own-key models](/docs/inference/ai-gateway/byok) are billed by your provider on your provider account, not by Telnyx. Budgets and rate limits still apply and act as a guard on that provider spend. - Budgets and the `cost` values in usage reporting are measured at a flat reference rate of USD 5 per million input tokens and USD 15 per million output tokens, for enforcing limits and attribution. They are not your invoice. ## Rate limits `rpm_limit` and `tpm_limit` are evaluated over rolling 60-second windows, not calendar minutes, at the key, user and group scopes. - On groups and users, null removes the limit and zero denies every request. Token keys differ; see [Token key limits](#token-key-limits). - A rate-limited request returns `429` with a `Retry-After` header. - `tpm_limit` counts the request's reserved tokens, including the full output allowance when `max_tokens` is not set. Set `max_tokens` to keep requests under a low `tpm_limit`. ## Token key limits Token keys are always capped. Group and user limits are unchanged: null is uncapped and zero denies. | Field | Range | Omitted or null | | --- | --- | --- | | `max_budget` | Above 0, at most USD 1,000, up to six decimal places | USD 1,000 | | `rpm_limit` | 1 to 6,000 | 6,000 | | `tpm_limit` | 1 to 10,000,000 | 10,000,000 | - An omitted or null key limit is stored as its maximum. On a blocked key, a null limit is kept until the key is unblocked. - The default USD 1,000 budget is a lifetime budget unless `budget_duration` is set. - A value of 0 or above the maximum returns `400` with code `limit_out_of_range`. - A negative, fractional or non-numeric value, or a budget with more than six decimal places, returns `400` with code `invalid_request`. - To deny a key, block it (`blocked: true`) or revoke it. ## End-user identity The OpenAI `user` field and the Anthropic Messages `metadata.user_id` field identify an account-scoped end user in the same namespace. The value is used for end-user budgets and blocks and appears as `end_user_id` in usage reporting. The identifier is an assertion by whoever holds the token key, not an authenticated identity. A key embedded in a client can assert any value. Bind end-user identifiers to authenticated sessions in a trusted backend, and set `required_end_user_id: true` on a key when every request must carry one. That flag checks presence only. ## Mutation safety Management mutations are protected by idempotency keys and ETag preconditions; see [Conventions](/docs/inference/ai-gateway/management-api#conventions). A committed policy change applies to new admissions once it has propagated. During propagation, inference can return `503` with code `enforcement_unavailable`. Fail closed in that case rather than retrying under the assumption that the previous policy still applies. ## Content guardrails Set a group-level `guardrails` policy to flag or block supported secrets and sensitive data before provider dispatch or before delivering a response. Prompt blocks reserve no budget; response blocks still account for provider work. See [Guardrails](/docs/inference/ai-gateway/guardrails) for configuration and streaming behavior. --- ### Guardrails > Source: https://developers.telnyx.com/docs/inference/ai-gateway/guardrails.md Guardrails inspect supported text fields in prompts and model responses. Configure a policy on a [token group](/docs/inference/ai-gateway/management-api#token-groups); every token key in that group uses it. Policies apply to new admissions. Existing groups have `guardrails: null`, which disables inspection. Guardrails detect supported credential and sensitive-data patterns. They do not redact text or provide model-based content moderation, prompt-injection detection, or a guarantee that all sensitive data will be detected. ## Configure a policy Include `guardrails` when creating a token group, or send it in `PATCH /v2/llm_token_gateway/token_groups/{id}`. Authenticate with a Telnyx account API key and supply the [idempotency and ETag headers](/docs/inference/ai-gateway/management-api#etag-preconditions) required for the mutation. This PATCH body blocks supported secrets in both directions and flags financial data: ```json { "guardrails": { "secrets": { "prompt": "block", "response": "block" }, "dlp": { "profiles": ["financial"], "prompt": "flag", "response": "flag" }, "streaming": "buffered" } } ``` Each detector has independent `prompt` and `response` actions: | Action | Behavior | | --- | --- | | `ignore` | Do not run the detector at that stage. This is the default. | | `flag` | Allow the request or response and record findings. | | `block` | Reject a matching prompt or withhold a matching response. | If several rules match, `block` takes precedence over `flag`. DLP needs at least one profile whenever either DLP action is `flag` or `block`. Unknown fields, duplicate profiles and unsupported values return `400`. Omitting `guardrails` from a PATCH preserves the policy. Supplying a `guardrails` object replaces the entire nested policy: omitted actions become `ignore`, and omitted DLP profiles become an empty list. Send the complete intended policy when changing one action. To disable inspection, send: ```json { "guardrails": null } ``` A policy update can briefly cause inference to return `503` while the change propagates. Read the group back to verify the saved policy. ## Detectors | Detector or profile | Supported patterns | | --- | --- | | `secrets` | AWS access key IDs; supported GitHub, Slack, Stripe, OpenAI, Anthropic, Google and Telnyx key patterns; AI Gateway token keys; JWT-shaped tokens; private-key block markers. | | DLP `financial` | Luhn-valid card-number patterns and checksum-validated IBAN patterns. | | DLP `government_id` | US Social Security number patterns and UK National Insurance number patterns. | | DLP `contact` | Email address and E.164 phone-number patterns. | These are pattern checks, not verification that a credential is active or an identifier belongs to a real person. Formats outside the supported patterns can be missed; text that resembles a supported format can match. Names, addresses and other free-text personal information are not detected. The gateway inspects supported message text, system text, tool definitions and tool inputs/arguments. Response inspection includes supported text, refusal, thinking and tool-use fields. Opaque provider signatures, encrypted thinking data, provider metadata and unrecognized content-block types are outside this inspection coverage. Guardrails do not inspect images or audio. ## Blocking and usage | Stage | Observable result | Usage behavior | | --- | --- | --- | | Prompt | HTTP `400`, code `prompt_blocked`. | No provider request or budget reservation; no spend event is created. | | Complete response | HTTP `400`, code `response_blocked`. | The provider already generated the response. Its usage is accounted for even though the response is withheld. | | Buffered streamed response | An SSE error carrying `response_blocked`, after the HTTP `200` headers. No provider content is released. | Provider usage is accounted for. | Guardrail findings are separate from the [usage ledger](/docs/inference/ai-gateway/usage). Do not interpret a blocked response as a refund or automatically retry it. When a cached response is available, it is checked against the current response policy before replay; a blocked cache hit makes no provider call and creates no spend event. ## Streaming `streaming` accepts `buffered` or `passthrough`: - **Buffered:** the gateway holds the stream until inspection completes, then replays its SSE events. This delays the first content event. Any response `block` action requires this mode; it is selected automatically if `streaming` is omitted. Explicit `passthrough` with a response `block` action is rejected with `400`. - **Passthrough:** events are forwarded immediately and findings are evaluated after the stream completes. This mode cannot block responses. It is the default when no response action is `block`. For Chat Completions, a streamed block is a `data:` event containing the error document. For Anthropic Messages, it is an `event: error` frame. Handle stream errors even when the initial HTTP status was `200`. ## Flagged-response header Successful responses with flagged findings can include `x-ltg-policy`. Its value is JSON containing detector codes, counts and actions: ```json { "outcome": "flagged", "findings": [ { "detector": "dlp", "code": "email", "count": 1, "action": "flag" } ] } ``` Prompt findings are available on successful complete and streamed responses. Response findings are available in the header only for non-streamed responses; streamed response findings are recorded after the headers have been sent. Blocked findings are available through the events endpoint, not this header. Finding counts represent pattern matches during inspection, not unique sensitive values. Structured data can be inspected in both its JSON spelling and unescaped form, so a value can contribute more than one match. ## Query findings Use `GET https://api.telnyx.com/v2/llm_token_gateway/guardrail_events` with the Telnyx account API key. Only that account's events are returned, newest first. Events identify the request, token key, group, optional user/end user, model, stage, outcome and findings. Findings contain codes and counts, never matched text or prompt/response excerpts. | Query parameter | Values or behavior | | --- | --- | | `start_date`, `end_date` | Supply both as `YYYY-MM-DD`. Start is inclusive and end exclusive, in UTC; the range must be positive and at most 31 days. With neither supplied, the range covers today and the previous six UTC dates. | | `token_group_id`, `token_key_id` | Filter by UUID. An unknown or foreign-account ID returns no matching events. | | `end_user_id` | Filter by the asserted end-user identifier. | | `stage` | `prompt` or `response`. | | `outcome` | `flagged`, `blocked`, `evaluated` or `unevaluated`. The current deterministic detectors emit `flagged` and `blocked`; clean passes create no event. | | `page[number]` | Positive page number; defaults to `1`. | | `page[size]` | `1` to `100`; defaults to `20`. | The response contains `data` and `meta.page_number`, `meta.page_size`, `meta.has_more`. Unlike resource and spend listings, guardrail events use offset pagination and do not accept `page[snapshot]`. Newly inserted events can shift page boundaries; deduplicate by event `id` when collecting several pages. A findings-storage failure does not make a blocked request pass, but it can leave no event for that decision. Use the inference error and `X-Request-ID` as well when investigating a blocked request. --- ### Bring your own key > Source: https://developers.telnyx.com/docs/inference/ai-gateway/byok.md By default, a token group uses Telnyx-hosted models and usage is billed to your Telnyx account. With bring your own key (BYOK), you store an OpenAI or Anthropic key through the management API and attach it to one or more groups. Requests on [bring-your-own-key models](/docs/inference/ai-gateway/inference-api#bring-your-own-key-models) from those groups run on your provider account, and your provider bills them. BYOK changes **who is billed**. It does not change attribution, budget enforcement or rate limiting: key, user, group and end-user controls apply exactly as they do for Telnyx-hosted models. ## Set up BYOK `POST /provider_keys` with a `name`, the `provider` (`openai` or `anthropic`) and the `secret`. The secret is accepted only on create and is never returned by any response, list or idempotent replay. ```bash curl -X POST https://api.telnyx.com/v2/llm_token_gateway/provider_keys \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: $(uuidgen)" \ -d '{ "name": "anthropic-production", "provider": "anthropic", "secret": "" }' ``` ```json { "data": { "record_type": "provider_key", "id": "2d4f6a8b-1c3e-4a5b-9d7f-0e1a2b3c4d5e", "name": "anthropic-production", "provider": "anthropic", "version": 1, "created_at": "2026-09-22T10:00:00Z", "updated_at": "2026-09-22T10:00:00Z" } } ``` Never keep a logged copy of the request that contains the secret. `GET /provider_keys` lists provider keys and `GET /provider_keys/{id}` reads one; neither shows the secret. Reference the provider key in the group's `provider_key_ids`, and include the BYOK models you want in `allowed_models`. A group holds at most one key per provider, the key's provider must match at least one BYOK model the group allows, and the key and group must be in the same account. Set both when creating a group, or update an existing group with an ETag-protected `PATCH`. `allowed_models` is replaced as a whole, so include every model the group should keep: ```bash curl -X PATCH "https://api.telnyx.com/v2/llm_token_gateway/token_groups/$GROUP_ID" \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: $(uuidgen)" \ -H "If-Match: $GROUP_ETAG" \ -d '{ "allowed_models": ["Kimi-K3", "claude-sonnet-5"], "provider_key_ids": ["2d4f6a8b-1c3e-4a5b-9d7f-0e1a2b3c4d5e"] }' ``` Applications keep using their `ltg_sk_` token key exactly as before. Never place the provider secret in SDK configuration or application code. Call BYOK models on `/v1/chat/completions`; Anthropic BYOK models also work on `/v1/messages` with the [Anthropic SDK](/docs/inference/ai-gateway/inference-api#anthropic-sdk). A BYOK model works only when the calling key's group has an attached provider key for that model's provider. Without one, requests fail with `503` and code `enforcement_unavailable`. They never fall back to a Telnyx-hosted model. ## Billing and reporting - Requests on BYOK models are billed by your provider on your provider account, not by Telnyx. - AI Gateway budgets and rate limits still apply. - Usage reporting records every BYOK request. Its `cost` is the [budget reference valuation](/docs/inference/ai-gateway/controls#budgets-and-billing), not a Telnyx charge. ## Rotate a provider secret Provider keys cannot be edited; `PATCH` returns `405`. To change a secret, create a new provider key, attach it to each group in place of the old one, then delete the old provider key. ## Delete a provider key `DELETE /provider_keys/{id}` requires the provider key's current `ETag` in `If-Match`. Read the key first to get it: ```bash PROVIDER_KEY_ETAG=$(curl -sS -I "https://api.telnyx.com/v2/llm_token_gateway/provider_keys/$PROVIDER_KEY_ID" \ -H "Authorization: Bearer $TELNYX_API_KEY" \ | awk 'tolower($1) == "etag:" { sub(/\r$/, "", $2); print $2 }') curl -X DELETE "https://api.telnyx.com/v2/llm_token_gateway/provider_keys/$PROVIDER_KEY_ID" \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Idempotency-Key: $(uuidgen)" \ -H "If-Match: $PROVIDER_KEY_ETAG" ``` Deleting a provider key detaches it from every group that references it. Requests already in progress complete; new requests on that provider's models from those groups fail with `503` until another key for the provider is attached. --- ### Usage reporting > Source: https://developers.telnyx.com/docs/inference/ai-gateway/usage.md Every inference request produces a durable spend event attributed to its token key, token user, token group and asserted end user. Two management endpoints expose that ledger: | Endpoint | Returns | | --- | --- | | `GET /spend/events` | One row per request. | | `GET /spend/summary` | Totals grouped by one dimension. | Both live under `https://api.telnyx.com/v2/llm_token_gateway`, authenticate with the Telnyx account API key and use the same date range, filters and pagination. ## Date range and filters | Parameter | Notes | | --- | --- | | `start_date` | Inclusive UTC date, `YYYY-MM-DD`. Required for `/spend/summary`. | | `end_date` | Exclusive UTC date, `YYYY-MM-DD`. Required for `/spend/summary`. | | `token_group_id`, `token_user_id`, `token_key_id` | Optional UUID filters. | | `end_user_id` | Optional end-user identifier filter. | | `group_by` | `/spend/summary` only, required: `token_group`, `token_user`, `token_key` or `end_user`. | | `page[number]`, `page[size]`, `page[snapshot]` | Snapshot pagination; see below. | The range is half-open, `[start_date, end_date)`, and spans at most 31 days. Send ISO dates, not timestamps. If you omit both dates on `/spend/events`, it returns the previous full UTC day. To include today, set `end_date` to tomorrow's date. ## Spend events ```bash curl --globoff -G https://api.telnyx.com/v2/llm_token_gateway/spend/events \ -H "Authorization: Bearer $TELNYX_API_KEY" \ --data-urlencode "start_date=2026-09-22" \ --data-urlencode "end_date=2026-09-23" \ --data-urlencode "token_group_id=$GROUP_ID" \ --data-urlencode "page[size]=100" ``` Each event carries: | Field | Meaning | | --- | --- | | `id`, `request_id`, `created_at` | Event identity and the `X-Request-ID` of the inference request. | | `token_group_id`, `token_user_id`, `token_key_id`, `end_user_id` | Attribution. `token_user_id` is null for service keys; `end_user_id` is null when the request asserted none. | | `model` | The model requested. | | `input_tokens`, `output_tokens` | Null when usage is unknown. | | `cost` | USD at the flat reference rate used for budgets; see [Budgets and billing](/docs/inference/ai-gateway/controls#budgets-and-billing). Null when usage is unknown. | | `status` | `succeeded`, `failed`, `partial` or `unknown`. | | `usage_status` | `known`, `unknown` or `reconciled` after a later correction. | | `configuration_version`, `rate_version` | The policy and rate snapshot the request was valued under. | | `reservation_micro_usd` | The reservation held for the request in micro-USD. It stays outstanding while usage is unknown. | ## Spend summary `group_by` selects the dimension. Each row totals the requests attributed to one value of that dimension within the range and filters. ```bash curl --globoff -G https://api.telnyx.com/v2/llm_token_gateway/spend/summary \ -H "Authorization: Bearer $TELNYX_API_KEY" \ --data-urlencode "start_date=2026-09-01" \ --data-urlencode "end_date=2026-10-01" \ --data-urlencode "group_by=token_key" ``` ```json { "data": [ { "dimension_id": "9b7e4a10-3c2d-4f5e-8a6b-1d2c3e4f5a60", "name": "support-backend", "token_group_id": "5f1c9d2e-7b3a-4c8e-9f21-0a6d4e8b1c33", "token_user_id": null, "token_key_id": "9b7e4a10-3c2d-4f5e-8a6b-1d2c3e4f5a60", "end_user_id": null, "spend": 1.204311, "requests": 4180, "input_tokens": 912340, "output_tokens": 401277, "unknown_requests": 2, "reserved_spend": 0.0125 } ], "meta": { "page_number": 1, "page_size": 100, "has_more": false, "snapshot": "...", "start_date": "2026-09-01", "end_date": "2026-10-01", "group_by": "token_key" } } ``` The four `group_by` dimensions are alternative views of the **same** requests. Do not add totals from different dimensions together. ## Pagination The first page returns `meta.snapshot`, a stable view bound to the account, range and filters for 15 minutes. While `meta.has_more` is true, request the next `page[number]` with the same parameters plus `page[snapshot]=`. Changing filters or mixing snapshots returns `409`. Treat an export as complete only once `has_more` is false. ## Interpreting the numbers - **Unknown usage is not zero.** A request whose usage never arrived, for example an interrupted stream, keeps its reservation as `reserved_spend` and reports `cost: null` with `usage_status: "unknown"`. Later reconciliation updates the same event to `reconciled`. - **Usage is billed at standard pricing.** Usage of Telnyx-hosted models is billed to your Telnyx account at standard [Telnyx AI Inference pricing](https://telnyx.com/pricing/inference-api) for each model. - **BYOK requests are reported like any other.** Requests on [bring-your-own-key models](/docs/inference/ai-gateway/byok) appear in spend events and summaries like any other request. Their `cost` is the budget reference valuation, not a Telnyx charge; your provider bills them. - **Spend is not an invoice.** `cost` and `spend` are measured at a flat reference rate (USD 5 per million input tokens, USD 15 per million output tokens) for enforcing budgets and attribution. They are not your invoice. - **Revoked and deleted resources keep their history.** Filters by ID continue to work after a key, user or group is deleted. --- ### Errors > Source: https://developers.telnyx.com/docs/inference/ai-gateway/errors.md Both planes return a structured `errors` array. The inference plane additionally wraps it in the envelope the calling SDK expects, so OpenAI and Anthropic SDK exceptions work unchanged while the Telnyx detail remains available. ## Error envelopes ```json Management { "errors": [ { "code": "precondition_failed", "title": "Stale If-Match", "detail": "The resource version has changed; read it again.", "meta": { "current_version": 4 } } ] } ``` ```json OpenAI-compatible { "error": { "message": "Send either max_tokens or max_completion_tokens, not both.", "type": "invalid_request_error", "param": "max_completion_tokens", "code": "invalid_request" }, "errors": [ { "code": "invalid_request", "title": "Invalid request", "detail": "Send either max_tokens or max_completion_tokens, not both.", "meta": {} } ] } ``` ```json Anthropic-compatible { "type": "error", "error": { "type": "permission_error", "message": "Model is not available to this token key." }, "request_id": "0b1c2d3e-4f50-4617-8a29-3b4c5d6e7f80", "errors": [ { "code": "model_not_in_catalog", "title": "Model not allowed", "detail": "Model is not available to this token key.", "meta": { "scope": "token_key" } } ] } ``` Every response carries an `X-Request-ID` header. Keep it, together with the `code`, when reporting a problem. Do not log authorization headers, token keys or full request objects. An error that occurs after a streaming response has started cannot change the HTTP status. On the OpenAI surface it arrives as an error chunk; on the Anthropic surface, as an `event: error` frame. Consume every stream to completion and handle the SDK's stream exceptions. ## Status codes | Status | Meaning | Action | | --- | --- | --- | | `400` | A guardrail policy blocked the prompt or response; invalid body, unsupported option or model-specific option, unknown field, request over the size limits, token key limit out of range, invalid date or page parameter. | Correct the request. | | `401` | Missing or wrong credential for this plane. | Use a Telnyx API key on the management plane and an `ltg_sk_` token key on the inference plane. | | `403` | Blocked, revoked or expired key; blocked resource; model not allowed; budget or end-user policy denied. | Read the `code`. Do not treat `403` as retryable. | | `404` | Resource does not exist or belongs to another account. | Check the ID. | | `405` | Method not allowed, for example `PATCH` on a provider key. | Provider keys cannot be edited; create a new one instead. | | `409` | Idempotency conflict, membership conflict or pagination snapshot conflict. | Read the `code`. | | `412` | Stale `If-Match`. | Read the resource and retry with the current ETag. | | `428` | Missing `If-Match` (or `If-None-Match: *` on end-user create). | Add the precondition header. | | `429` | Rate limit exceeded. | Wait for `Retry-After`. Remember that inference retries are not idempotent. | | `502` | The model provider failed. | Fail closed. The request may or may not have consumed provider work. | | `503` | Enforcement, catalog, policy propagation or a required dependency is unavailable, or a BYOK model has no attached provider key for its provider. | Fail closed. Do not switch credentials or bypass the gateway. | ## Error codes | Code | Status | Meaning | | --- | --- | --- | | `invalid_request` | 400 | Malformed or unsupported request, including a request over the [size limits](/docs/inference/ai-gateway/inference-api#request-size-limits) and a request to `/v1/messages` with a model that is not an Anthropic BYOK model (use `/v1/chat/completions` for those). | | `prompt_blocked` | 400 | The group guardrail policy blocked the prompt before provider dispatch or budget reservation. | | `response_blocked` | 400, or SSE error after HTTP 200 | The group guardrail policy withheld the generated response. Provider usage still applies. See [Guardrails](/docs/inference/ai-gateway/guardrails#blocking-and-usage). | | `limit_out_of_range` | 400 | A token key `max_budget`, `rpm_limit` or `tpm_limit` is 0 or above its maximum. See [Token key limits](/docs/inference/ai-gateway/controls#token-key-limits). | | `invalid_token_key` | 401 | The inference credential is missing, malformed or unknown. | | `unauthorized` | 401 / 403 | The credential is not valid for this plane or resource. | | `token_key_blocked` | 403 | The key is blocked, revoked or expired. | | `resource_blocked` | 403 | The key's user, group or asserted end user is blocked. | | `budget_exceeded` | 403 | A key, user or group budget lacks headroom for the request's reservation. | | `end_user_budget_exceeded` | 403 | The asserted end user's budget lacks headroom. | | `model_not_in_catalog` | 403 | The model is not in the key's or group's allowlist. | | `rate_limit_exceeded` | 429 | An RPM or TPM limit was hit. | | `not_found` | 404 | No such resource in this account. | | `conflict` | 409 | Membership or pagination snapshot conflict. | | `idempotency_conflict` | 409 | The `Idempotency-Key` was reused with a different body, or the original request is still in progress. | | `precondition_failed` | 412 | `If-Match` does not match the current version. | | `precondition_required` | 428 | A required precondition header is missing. | | `enforcement_unavailable` | 503 | Policy, catalog or a dependency needed to admit the request is unavailable, or the request uses a [BYOK model](/docs/inference/ai-gateway/byok) and the key's group has no attached provider key for that model's provider. | | `upstream_error` | 502 | The model provider returned an error. | ## Handling guidance - **Never resolve an error by escalating credentials.** A Telnyx API key, provider secret or any other credential is rejected on the inference plane by design. - **Retry management mutations with the same idempotency key.** A new key on retry can create a duplicate resource. - **Do not automatically retry inference.** A timeout or `502` is not proof that no provider work happened. Retry only errors that occurred before dispatch, and honor `Retry-After` on `429`. - **Re-read before re-writing.** On `412`, fetch the resource, review the change that landed, and apply your update to the current version. - **Fail closed on `503`.** The old policy may no longer apply; wait and retry rather than assuming the request is authorized. For a BYOK model, check that the group has a provider key attached for that model's provider. --- ## Embedding ### Overview > Source: https://developers.telnyx.com/docs/inference/embedding-rag.md Embedding and retrieval-augmented generation (RAG) let your applications search prior knowledge before asking a model to respond. Instead of relying solely on a model's training data, RAG retrieves relevant information from your own content and provides it as context at inference time -- producing more accurate, grounded, and up-to-date responses. ## What Is Embedding & RAG? **Embedding** converts text into numeric vectors that capture meaning. Two pieces of text with similar meaning produce similar vectors, even if the exact words differ. This makes it possible to search by concept rather than by keyword. **Retrieval-augmented generation (RAG)** uses those embeddings at query time. When a user asks a question, the system: 1. Embeds the query into a vector 2. Searches your indexed content for the most relevant chunks 3. Passes those chunks as context to a language model 4. The model generates a response grounded in your data The result is answers that reference your actual content -- call transcripts, documents, messages -- rather than guessing from training data. ## What's in This Section Telnyx's managed RAG product. Create searchable collections over your Telnyx communications data and query them with one retrieval API. Search persisted conversation records directly -- the same indexed history that backs AI Search's conversation sources. Lower-level primitives: embed documents in a Telnyx Storage bucket and run similarity search or clustering over them yourself. Rates for embedding, storage, and search events. Start with [AI Search](/docs/ai-search) -- it manages chunking, embedding, indexing, and ranking for you. Reach for the [embeddings APIs](/docs/inference/embeddings) when you need the primitives directly; [How AI Search Works](/docs/ai-search/how-it-works) compares the two. These primitives can be used with AI Assistants, custom agent runtimes, or your own application code. --- ### Pricing > Source: https://developers.telnyx.com/docs/inference/embedding-rag/pricing.md Pricing for the Embedding & RAG section is broken down by feature. Search covers collection management, embedding, storage, and search events. Embedding covers document embedding and storage for use with RAG workflows. Search pricing has three parts: embedding and persisting content when it is indexed for search, storage retention beyond 30 days, and search events when you query a collection. ## Rates | Usage | Price | Notes | | --- | --- | --- | | Embed and persist | `$0.0015 / 1K characters` | Billed once on input content characters when the content is indexed. Adding an already-indexed source to another collection is free. Includes 30 days of retention. | | Storage after 30 days | `$0.60 / GiB-month` | Applies only when extended retention is enabled. | | Search event | `$0.003 / search` | Charged per ranked search (a request with a `query` parameter). First 10,000 searches per month are free of charge. | | Catalog listing | Free | Browsing a collection's documents without a `query` parameter is not billed. | ## How Billing Works ### Embed and Persist Content is billed once when it is indexed. A source is indexed when its content is first ingested for search; once indexed, the same source can be added to any number of collections at no additional cost. You are not billed for a source simply because it is part of a collection. Billing covers the input content characters and includes 30 days of retention. Storage after 30 days is billed at the extended retention rate only if extended retention is enabled. ### Search Events A **search event** is a single `GET /v2/ai/knowledge/collections/{slug}/documents` request that includes a `query` parameter. Each ranked search counts as one billable event, regardless of the `top_k` value or the number of results returned. The first 10,000 searches per month are free. Requests without a `query` parameter (catalog listings) are free -- they return documents sorted by date without ranking, and are not counted as search events. ### Multi-Region Fan-Out When a collection's sources are stored across multiple regions, a search query fans out to all relevant regions in parallel. Each region retrieves its matching results independently, and the results are merged into a single ranked list by score. The user does not see or control which regions are searched -- the fan-out and merge happen automatically. Despite querying multiple regions, the entire fan-out counts as **one** billable search event, not one per region. The multi-region behavior does not increase the cost of a search. Region selection is automatic -- there is no region parameter on search requests. ## Example A collection with 10,000 characters of indexed content and one search within the free tier would be priced as: | Line Item | Calculation | Price | | --- | --- | --- | | Embed and persist | `10K characters * $0.0015` | `$0.015` | | Storage within 30 days | Included | `$0.00` | | One search | Included in the first 10,000 searches per month | `$0.00` | | Total | | `~$0.015` | Embedding pricing covers document embedding and storage for use with RAG workflows. Documents uploaded to a Telnyx Storage bucket are embedded and indexed for similarity search and retrieval. ## Rates | Usage | Price | Notes | | --- | --- | --- | | Embed and persist | `$0.0015 / 1K characters` | Billed once on input content characters when documents are embedded. Includes 30 days of retention. | | Storage after 30 days | `$0.60 / GiB-month` | Applies only when extended retention is enabled. | | Similarity search | `$0.003 / search` | Charged per similarity search request. First 10,000 searches per month are free of charge. | ## How Billing Works ### Embed and Persist When you embed documents in a Telnyx Storage bucket, the content is processed into sections and each section is embedded into a vector. You are billed once on the input content characters. This includes 30 days of retention. Storage after 30 days is billed at the extended retention rate only if extended retention is enabled. ### Similarity Search A similarity search request retrieves the most relevant document sections for a given query. Each search counts as one billable event. The first 10,000 searches per month are free. ## Example A bucket with 50,000 characters of document content and 100 searches in a month would be priced as: | Line Item | Calculation | Price | | --- | --- | --- | | Embed and persist | `50K characters * $0.0015` | `$0.075` | | Storage within 30 days | Included | `$0.00` | | 100 searches | Included in the first 10,000 searches per month | `$0.00` | | Total | | `~$0.075` | --- ### Embeddings > Source: https://developers.telnyx.com/docs/inference/embeddings.md In this tutorial, you'll learn how to: - Upload documents to [Telnyx Storage](https://telnyx.com/products/cloud-storage) - Transform these documents into embeddings, enabling a language model to retrieve relevant sections of your documents - Provide the storage bucket as context for the language model ## Upload your documents You can upload objects to Telnyx's S3-Compatible storage API using our [quickstart](https://developers.telnyx.com/docs/cloud-storage/quick-start) or with our [drag-and-drop interface in the portal](https://portal.telnyx.com/#/storage/buckets). ## Embed your documents Once you've uploaded your documents, you can [embed them via API](https://developers.telnyx.com/api-reference/embeddings/embed-url-content#embed-url-content) or by clicking the "Embed for AI Use" button in the portal while viewing your storage bucket's contents. Behind the scenes, your documents will be processed into sections and each section will be "embedded" based on its contents. Later, when a user asks a language model a question, it will automatically be provided with the most relevant sections of documents from the bucket to help answer the question. ## Chat over your documents Once your documents are embedded, you can try it out in the [AI Playground in the portal](https://portal.telnyx.com/#/ai/playground) by selecting your embedded bucket from the storage dropdown. You can also use embeddings via our [chat completions API](https://developers.telnyx.com/api-reference/openai-chat/create-a-chat-completion-openai-compatible). Here is a Python example. Make sure you have set the `TELNYX_API_KEY` environment variable. Also, update the `question` and `bucket` variables in the sample code. ```python import os from openai import OpenAI client = OpenAI( api_key=os.getenv("TELNYX_API_KEY"), base_url="https://api.telnyx.com/v2/ai/openai", ) question = "" bucket = "" chat_completion = client.chat.completions.create( messages=[ { "role": "user", "content": question } ], model="zai-org/GLM-5.3-Flash", reasoning_effort="high", stream=True, tools=[ { "type": "retrieval", "retrieval": { "bucket_ids": [bucket] } } ] ) for chunk in chat_completion: if chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="", flush=True) ``` --- ### Clusters > Source: https://developers.telnyx.com/docs/inference/clusters.md In this tutorial, you'll learn: - How [Embeddings](https://developers.telnyx.com/api-reference/embeddings/embed-url-content#embed-url-content) and [Clusters](https://developers.telnyx.com/api-reference/clusters/compute-new-clusters#compute-new-clusters) work - How to leverage them to identify common themes in your data # Embeddings and Clusters Embeddings are numerical representations of concepts within text, image, or audio data. [The representation is a real-valued vector that encodes the meaning of the word in such a way that the words that are closer in the vector space are expected to be similar in meaning](https://en.wikipedia.org/wiki/Word_embedding) Quantifying the semantic similarity of your data opens up several possibilities. For instance, by embedding a Telnyx storage bucket, you can [search for similar content](https://developers.telnyx.com/api-reference/embeddings/embed-url-content#embed-url-content) within your bucket. This tutorial is focused on another application of embeddings: analyzing how your semantic data is [clustered](https://en.wikipedia.org/wiki/Cluster_analysis), which provides insight into common themes and niche subtopics. For example, pictured below are clusters of embeddings computed for the novel The Great Gatsby. ![Gatsby clusters](/assets/images/gatsby-cluster.png) # Clustering content with Telnyx ## Embed your documents Embedding your content in a Telnyx storage bucket is a prerequisite for computing these clusters. For more information, check out our [Embeddings](https://developers.telnyx.com/docs/inference/embeddings) tutorial. ## Identify clusters Once your documents are embedded, you can [compute clusters](https://developers.telnyx.com/api-reference/clusters/compute-new-clusters#compute-new-clusters) via API. The optional `prefix` and `files` parameters allow you to specfiy a subset of your bucket you would like to cluster. The `min_cluster_size` and `min_subcluster_size` parameters control how clusters are identified. Top-level clusters should be thought of as identifying broad themes in your data. Choose `min_cluster_size` based on the minimum data points you would like to constitute a broader theme. Sub-clusters should be thought of as identifying more specific topics within a broader theme. Choose `min_subcluster_size` based on the minimum data points you would like to constitute a more niche subtopic. ## Identifying themes in The Great Gatsby To demonstrate embedding and clustering a Telnyx storage bucket, we will be using the text from The Great Gatsby. ### Upload to Telnyx Storage You can upload objects to Telnyx's S3-Compatible storage API using our [quickstart](https://developers.telnyx.com/docs/cloud-storage/quick-start) or with our [drag-and-drop interface in the portal](https://portal.telnyx.com/#/storage/buckets). ### Embed your documents Once you've uploaded your documents, you can [embed them via API](https://developers.telnyx.com/api-reference/embeddings/embed-url-content#embed-url-content) or by clicking the "Embed for AI Use" button in the portal while viewing your storage bucket's contents. Behind the scenes, your documents will be processed into chunks and each chunk will be "embedded" based on its contents. Each chunk will be a single data point used in the clustering step. ### Compute clusters You can compute multiple clusterings on the same data. This is helpful to tweak the parameters to find the best clusters for your data. Below is an example API request ``` $ curl --request POST \ --url https://api.telnyx.com/v2/ai/clusters \ --header "Authorization: Bearer $TELNYX_API_KEY" \ --header 'Content-Type: application/json' \ --data '{ "bucket": "cluster-gatsby", "min_cluster_size": 50, "min_subcluster_size": 10 }' ``` And the response ``` {"data":{"task_id":"04dd624f-c9b3-4fc8-8cec-492c8696e9ea"}} ``` ### Inspect clusters You can then take that `task_id` and view the clusters structured as JSON via ``` $ curl --request GET \ --url "https://api.telnyx.com/v2/ai/clusters/04dd624f-c9b3-4fc8-8cec-492c8696e9ea?show_subclusters=true" \ --header "Authorization: Bearer $TELNYX_API_KEY" ``` If you want to see example data from each cluster, you can also pass the `top_n_nodes` query parameter which will include the top N most central data points for each cluster. You can also view a simple graph of the clusters via ``` $ curl --request GET \ --url "https://api.telnyx.com/v2/ai/clusters/04dd624f-c9b3-4fc8-8cec-492c8696e9ea/graph" \ --header "Authorization: Bearer $TELNYX_API_KEY" --output clusters.png ``` If you want to look at a cluster's subclusters, you can pass the `cluster_id` query parameter. Here is a closer look at the sub-clusters related to the cluster for "Daisy's Past" using this endpoint ![Gatsby clusters](/assets/images/gatsby-daisy-subcluster.png) The initial parameters can have a large effect on the computed clusters, and the "right" clusters depend heavily on your data set and your goals, so you may have to play around a bit to find what works best. The general idea is that raising `min_cluster_size` will result in broader, more generic clusters. You can also compute as many configurations over your data as you like so you have multiple ways of clustering your data if you'd like. --- ## Assistants ### Voice Assistant > Source: https://developers.telnyx.com/docs/inference/ai-assistants/no-code-voice-assistant.md In this tutorial, you'll learn how to configure a voice assistant with Telnyx. You won't have to write a single line of code or create an account with anyone besides Telnyx. You'll be able to talk to your assistant over the phone in under five minutes. After this tutorial covers the basics, we will explore some optional enhancements, such as changing your voice or language model providers and empowering your assistant with built-in tools. Try out the [public demos](https://telnyx.com/) for a real example of the finished product. ## Video Tutorial Watch this step-by-step demonstration of creating a voice assistant: ## Requirements There are 2 required steps for this tutorial 1. Configure your AI Assistant 2. Configure the voice settings ### Configure your AI Assistant First, navigate to the [AI Assistants tab](https://portal.telnyx.com/#/ai/assistants) in the portal. You will create a new assistant to configure what context your assistant has and how it behaves. For this tutorial, we will use a blank template. ![AI Assistant Portal Templates](/assets/images/ai-assistant-templates.png) You can use the following instructions and greeting, or use your own. Instructions: ``` You are an intelligent and concise voice assistant. This is a {{telnyx_conversation_channel}} happening on {{telnyx_current_time}}. The agent is at {{telnyx_agent_target}} and the user is at {{telnyx_end_user_target}}. ``` Greeting: ``` Hi {{first_name}}, this is Nyx, your friendly Telnyx Assistant! How can I help you today? ``` ![AI Assistant Portal Config](/assets/images/create_assistant.png) Telnyx provides these system variables to inject details about a specific call into the instructions and greeting: | Variable | Description | Example | |----------|-------------|---------| | `{{telnyx_current_time}}` | The current date and time in UTC | `Monday, February 24 2025 04:04:15 PM UTC` | | `{{telnyx_conversation_channel}}` | This can be `phone_call` , `web_call`, or `sms_chat` | `phone_call` | | `{{telnyx_agent_target}}` | The phone number, SIP URI, or other identifier associated with the agent. | `+13128675309` | | `{{telnyx_end_user_target}}` | The phone number, SIP URI, or other identifier associated with the end user | `+13128675309` | | `{{telnyx_sip_header_user_to_user}}` | The User to User SIP header for the call, if applicable | `cmlkPTM0Nzg1O3A9dQ==;encoding=base64;purpose=call` | | `{{telnyx_sip_header_diversion}}` | The Diversion SIP header for the call, if applicable | `;reason=user-busy` | | `{{call_control_id}}` | The call control ID for the call, if applicable | `v3:u5OAKGEPT3Dx8SZSSDRWEMdNH2OripQhO` | Telnyx also supports timezone-aware date/time variants (e.g. `{{telnyx_current_time_America/New_York}}`), shorthands (`{{telnyx_current_date}}`, `{{telnyx_current_weekday}}`), and a custom `date` format filter. See [dynamic variables](https://developers.telnyx.com/docs/inference/ai-assistants/dynamic-variables) for the full list. You can also define your own [custom dynamic variables](https://developers.telnyx.com/docs/inference/ai-assistants/dynamic-variables) and set them via webhook, custom SIP headers, or an outbound API call. In this example, you've given the assistant access to the most commonly used system variables and a custom `{{first_name}}` variable in the greeting. Notice that by default, the Hangup tool is configured. This enables your assistant to end the call at an appropriate time. ### Configure the voice settings In this step, you can use the default voice settings, click `Create`, and enable the agent for calls when prompted. Or feel free to explore the wide range of voices in the playground. - TTS (Text-to-Speech): Telnyx, AWS, Azure, ElevenLabs, Inworld - STT (Speech-to-Text): Telnyx (whisper), Deepgram, Azure — see [Transcription Settings](/docs/inference/ai-assistants/transcription-settings) for model details and configuration Browse the list of available voices and configuration options on the [Text-to-Speech Available Voices](/docs/voice/tts/available-voices) page. [Ultra](/docs/voice/tts/providers/telnyx/ultra) and [xAI Grok](/docs/voice/tts/providers/telnyx/grok) voices support Expressive Mode with inline SSML emotion tags and nonverbal cues like `[laughter]`. **Background Audio** You can also configure background audio to play during the call. This provides a more realistic noise environment for voice calls, making longer pauses from tool calls feel more natural. Customers can select from our list of predefined options or share their own custom public URL. **Speaking Plan** Customers are in full control of when an agent starts talking and can distinguish between 4 types of pauses: 1. **Wait seconds** sets your baseline. A customer service agent might use 0.3 seconds for snappy responses. An agent calling into an IVR system needs 1.5 seconds to account for slower robotic speech. 2. **On punctuation seconds** handles high-confidence endpoints. When the transcription ends with a period or question mark, the user likely finished their thought. Set this to 0.1 seconds for minimal delay. 3. **On no punctuation seconds** handles uncertainty. The user said "my order number is" and paused. They're probably looking at their screen. Set this to 1.5 seconds so the agent doesn't interrupt with "I didn't catch that" while they're reading the digits. 4. **On number seconds** handles digit sequences specifically. People read numbers slowly: "4... 7... 2... 9." Each pause could trigger a response. Extending this to 1.0 seconds prevents the agent from cutting them off at "4... 7..." **Noise Suppression** Enable noise suppression to reduce background noise during calls. This is especially valuable for AI assistants — cleaner audio improves STT accuracy and leads to better AI responses. Three engines are available through the assistant's `telephony_settings`: | Engine | Description | |--------|-------------| | **AiCoustics** | STT-optimized noise suppression, tuned to maximize speech recognition accuracy in real-world conditions (recommended for AI assistants) | | **Krisp** | Industry-leading noise suppression, effective across diverse environments (home offices, contact centers, outdoor) | | **DeepFilterNet** | Configurable attenuation for fine-grained control | **AiCoustics** is configured per-assistant: set `noise_suppression` to `"aicoustics"` and tune it through `noise_suppression_config`: ```bash curl -X POST https://api.telnyx.com/v2/ai/assistants/{assistant_id} \ -H "Authorization: Bearer ***" \ -H "Content-Type: application/json" \ -d '{ "telephony_settings": { "noise_suppression": "aicoustics", "noise_suppression_config": { "family": "quail", "size": "vf", "enhancement_level": 0.8 } } }' ``` | Parameter | Values | Description | |-----------|--------|-------------| | `family` | `quail` | AiCoustics model family optimized for Voice AI and STT | | `size` | `vf` \| `vf_2_0_l` | `vf` tracks the latest model release; `vf_2_0_l` is pinned to version 2.0 | | `enhancement_level` | 0–1 (default `0.8`) | Enhancement intensity | `voice_gain` is not part of the per-assistant `noise_suppression_config`; it is only available per call via the [Call Control API `suppression_start` action](/docs/voice/programmable-voice/noise-suppression) with `noise_suppression_engine: "AiCoustics"`. See the [Noise Suppression guide](/docs/voice/programmable-voice/noise-suppression) for the full per-call engine list and configuration. To enable via the [Mission Control Portal](https://portal.telnyx.com/#/ai/assistants), select the noise suppression engine under the voice settings when creating or editing your assistant. ![AI Assistant Noise Suppression](/assets/images/ai-assistant-noise-suppression.png) To enable via the API, set `noise_suppression` in `telephony_settings`: ```bash curl -X POST https://api.telnyx.com/v2/ai/assistants/{assistant_id} \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "telephony_settings": { "noise_suppression": "krisp" } }' ``` Set to `"disabled"` to turn off noise suppression. See the [Noise Suppression guide](/docs/voice/programmable-voice/noise-suppression) for more details on engines and direction options. **Receiving DTMF tones** AI Assistants can receive DTMF (Dual-Tone Multi-Frequency) tones from callers by default. This allows users to interact with your assistant using their phone keypad — for example, pressing digits in response to a menu prompt or entering a PIN. No configuration is needed to enable inbound DTMF. If you need to disable it for a specific use case (for example, if a `pay` tool is configured on the assistant), set `disable_dtmf` to `true` in the assistant's `telephony_settings`: ```bash curl -X POST https://api.telnyx.com/v2/ai/assistants/{assistant_id} \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "telephony_settings": { "disable_dtmf": true } }' ``` When DTMF is received, the tones are transcribed and passed to the assistant as part of the conversation, so the assistant can respond to keypad input naturally. ![AI Assistant Voice Config](/assets/images/full-voice-config.png) ### (Optional) Assign a phone number If you have already purchased a Telnyx number with voice features, you can immediately assign it to your assistant. You can also click `Next`, as completing this step is not needed for testing your assistant. ![AI Assistant Number Config](/assets/images/assistant_assign_numbers.png) ### (Optional) Enable messaging Telnyx assistants work with messaging as well. This is not covered in this tutorial. ### Test it out! You should now be able to interact with your voice assistant. When you are ready, you can ask the assistant to hang up to test out the Hangup tool. ![Call Agent](/assets/images/test_assistant.png) ![Talking to Agent](/assets/images/talking_to_assistant.png) If you assigned a number to your assistant, you can also have your assistant call you: - from the portal - via API/CLI - in an automated workflow like Zapier using our [TeXML Outbound Call](https://zapier.com/apps/telnyx/integrations#triggers-and-actions) action. ![Outbound Selection](/assets/images/assistant_test_outbound.png) ![Outbound Calling Agent](/assets/images/outbound_agent_test.png) *If you are using the curl provided, make sure to set the `TELNYX_API_KEY` environment variable with your API Key* ### Outbound calls via API To initiate an outbound call with your AI assistant via API, use the `/v2/texml/ai_calls/` endpoint: ```bash curl --request POST \ --url https://api.telnyx.com/v2/texml/ai_calls/ \ --header "Authorization: Bearer $TELNYX_API_KEY" \ --header 'Content-Type: application/json' \ --data '{ "From": "+13128675309", "To": "+15551234567", "AIAssistantId": "assistant-6207ab25-b185-478f-b2ef-85159e226727" }' ``` If your assistant has [voicemail detection options](https://developers.telnyx.com/api-reference/assistants/create-an-assistant#body-telephony-settings-voicemail-detection) configured in the telephony settings, include these AMD parameters: ```bash curl --request POST \ --url https://api.telnyx.com/v2/texml/ai_calls/ \ --header "Authorization: Bearer $TELNYX_API_KEY" \ --header 'Content-Type: application/json' \ --data '{ "From": "+13128675309", "To": "+15551234567", "AIAssistantId": "assistant-6207ab25-b185-478f-b2ef-85159e226727", "MachineDetection": "Enable", "AsyncAmd": true, "DetectionMode": "Premium" }' ``` ### Review the conversation You can see all historical conversations in the **Conversation History** tab ![AI Assistant Transfer](/assets/images/handoff_conversation.png) This example shows a conversation where the Handoff tool was used. ## MMS integration during voice calls Your AI Assistant can now receive and process MMS messages during live voice calls, enabling visual context and real-time image analysis. This powerful feature allows your assistant to handle complex scenarios where visual information is crucial. ### How it works When a user sends an MMS message during an ongoing voice call with your AI Assistant, the agent can: - Automatically detect incoming MMS messages. - Access and analyze attached images using vision-capable Language Models (VLMs). - Provide real-time responses based on visual content. - Continue the voice conversation with enhanced context. ### Use cases | Category | Use Case | Description | |----------|----------|-------------| | **Visual Support** | Customer Service | Analyze product photos sent by customers | | | Technical Support | Review error screenshots or equipment photos | | | Healthcare | Examine medical documents or symptoms | | **Document Verification** | Insurance Claims | Process claims with photo evidence | | | Identity Verification | Verify identity with document images | | | Compliance | Conduct compliance checks with visual documentation | | **Real-time Analysis** | Quality Control | Perform inspections with photo submissions | | | Inventory Management | Manage inventory with visual confirmation | | | Damage Assessment | Assess damage with real-time photo analysis | ### Configuration MMS integration requires the following setup for your AI Assistant: 1. **Messaging must be enabled** - Ensure messaging is enabled for your Voice AI Agent to receive MMS during calls. ![AI Assistant Messaging Configuration](/assets/images/ai-assistant-messaging-enabled.png) 2. **Vision-capable models required** - Use one of the two supported vision models: `Groq/llama-4-maverick-17b-128e-instruct` or `OpenAI/gpt-4o` for image processing. 3. **Image processing** - The assistant can handle common image formats (JPEG, PNG, etc.). ### Best practices - **Model selection**: Choose one of the two supported vision models when configuring your assistant. ![AI Assistant Model Selection](/assets/images/ai-assistant-agent-model-selection.png) - **Response timing**: The assistant will process images and respond within the voice call flow. - **Image quality**: Higher resolution images provide better analysis results. - **Context integration**: The assistant seamlessly combines visual and conversational context. For detailed information about vision Language Models and how to use them with our API, see our [Models page](/docs/inference/models). ## Supported language models Telnyx AI Assistants support multiple language models. You can select a model in the **Agent** tab when creating or editing an assistant. If no model is set, the assistant defaults to `moonshotai/Kimi-K2.6`. Not every model selectable for an assistant is verified for voice calls. For the current voice model list — what is verified for voice, what is chat-only, and how reasoning behaves on calls — see [Models & Supported Languages](/docs/voice/conversational-ai/quickstart#voice-verified-models). | Model | Provider | Description | | --- | --- | --- | | `moonshotai/Kimi-K2.6` | Telnyx (native) | Default model — recommended balance of intelligence and cost, no API key required | | `moonshotai/Kimi-K2.5` | Telnyx (native) | Previous default | | `zai-org/GLM-5.2` | Telnyx (native) | No API key required | | `anthropic/claude-haiku-4-5` | Anthropic (native) | Fast, lightweight model — no API key required | | `openai/gpt-5.4-mini` | OpenAI | Compact high-efficiency model for production voice workflows — requires OpenAI API key | | `openai/gpt-4o` | OpenAI | Requires OpenAI API key (see [OpenAI integration](#openai-integration)) | For a complete list of available models, see [Available Models](/docs/inference/models). Native models run on Telnyx infrastructure with no external API key required. For models from external providers, see [Third-party integrations](#third-party-integrations) or [Custom LLMs for Assistants](/docs/inference/ai-assistants/custom-llm). Reasoning (thinking) is always disabled on voice calls, on every model, and cannot be enabled — see [Reasoning on voice calls](/docs/voice/conversational-ai/quickstart#reasoning-on-voice-calls). ## Optional enhancements ### Integrations You can integrate your assistant with enterprise platforms to access customer data, create tickets, update records, and automate workflows directly during conversations. Telnyx AI assistants support integrations with: - **Salesforce** - CRM and customer service - **ServiceNow** - IT service management - **Jira** - Project and issue tracking - **HubSpot** - Marketing and sales CRM - **Zendesk** - Customer support - **Intercom** - Customer messaging - **Github** - Code hosting and version control - **Greenhouse** - Applicant tracking Select the integration from the dropdown ![AI Integration Dropdown](/assets/images/ai-integration-dropdown.png) Enter your credentials ![AI Integration Creds](/assets/images/ai-integration-creds.png) Connect the integration to the assistant and choose the enabled tools ![AI Integration Dropdown](/assets/images/ai-integration-tools.png) **Learn more:** For comprehensive setup guides, required credentials, available tools, and best practices for each integration, see the [AI Assistant Integrations](/docs/inference/ai-assistants/integrations) documentation. ### Additional built-in tools Besides `Hangup`, we offer several additional built-in tools to empower your agent to take real-world actions. **Webhook** With the webhook tool, your agent can make API requests. You can configure headers (with integration secrets), along with path, query, and body parameters. You can also reference dynamic variables in the webhook path or in the parameter descriptions. ![AI Assistant Webhook](/assets/images/agent_webhook_tool.png) After you've saved a webhook tool on your assistant, you can test it out with sample data by clicking the play button icon on the tool. ![AI Assistant Webhook Test](/assets/images/agent_webhook_test.png) **Handoff** With the handoff tool, you can enable multiple assistants to support a user in a single conversation. By default, the handoff is transparent to the user: assistants share the same context and voice, allowing for a unified experience where a variety of tasks can be handled by a team of specialists working behind the scenes. You can also toggle to the distinct voice mode, where all assistants retain their voice configuration, providing the experience of a conference call with a team of assistants ![AI Assistant Transfer](/assets/images/handoff_config.png) Learn more about agent handoff, including best practices, industry-specific templates, and advanced configuration in the [Agent Handoff guide](/docs/inference/ai-assistants/agent-handoff). **Transfer** and **SIP Refer** With the transfer and SIP Refer tools, your agent can transfer or refer a call to a list of named targets. ![AI Assistant Transfer](/assets/images/agent_transfer.png) **Send DTMF** With the Send DTMF, your agent can interact with legacy IVR systems. **Client-Side Tools** Client-side tools let the assistant call functions that run directly in the browser during a WebRTC voice or chat conversation. This is useful for reading data the page already has, triggering UI actions, or calling APIs authenticated with the user's browser session. See the [Client-Side Tools guide](/docs/inference/ai-assistants/client-side-tools) for setup and implementation details. ### Model Context Protocol (MCP) Servers You can [configure an MCP Server with Telnyx](https://portal.telnyx.com/#/ai/mcp-servers) and then add it to an assistant. If the URL for the server must be kept secret (because the server is not otherwise authenticated), you may store it securely as an integration secret with Telnyx. ![AI Assistant MCP](/assets/images/mcp_server.png) ![AI Assistant MCP Config](/assets/images/mcp_assistant.png) When setting up your MCP servers with Telnyx AI Assistants, Telnyx automatically includes a `telnyx_conversation_id` with each MCP tool call. If you are managing your own MCP Server, the `telnyx_conversation_id` can be used for tracking and controlling the flow of conversations within your applications. This is sent on _meta field of MCP. To receive the conversation ID at the start of a voice conversation, you have two options: - For call control applications, the conversation ID is returned by the [Start AI Assistant command](/api-reference/call-commands/start-ai-assistant#start-ai-assistant) - If you have configured a [dynamic variables webhook URL](https://developers.telnyx.com/docs/inference/ai-assistants/dynamic-variables), the conversation ID will be sent in this request payload at the start of a voice conversation. *The `telnyx_conversation_id` is set by the Telnyx platform, not by the AI agent, and as such is not susceptible to prompt injection attacks.* ```json { "jsonrpc": "2.0", "id": 1, "method": "some_mcp_method", "params": { "_meta": { "progressToken": null, "telnyx_conversation_id": "123" } } } ``` ### Knowledge Bases You can use the **Knowledge Bases** tab to enable your assistant to retrieve your custom context. First, provide a name ![AI Assistant KB](/assets/images/create-knowledge-base.png) Then upload files or provide a URL ![AI Assistant KB URL](/assets/images/upload_for_knowledge_base.png) ### Insights You can automatically run structured and unstructured analysis on every assistant conversation using the **Insights** tab. ![AI Assistant Insights](/assets/images/insights-config.png) You can assign an **Insight Group** to your assistant, which can have one or more Insights. Insights can be reused across Groups, and Groups can be reused across Assistants. You can also configure a webhook URL to receive conversation insights after they are generated. ![AI Assistant Insights Webhook](/assets/images/insights-webhook.png) **Learn more:** For comprehensive guides on creating insights, using structured data schemas, organizing insight groups, configuring webhooks, and industry-specific use cases, see the [AI Insights documentation](https://developers.telnyx.com/docs/inference/ai-insights). ### Embeddable Widget You can easily embed a customizable voice and chat widget on your frontend in the **Widget** tab. ## Programmatic Voice You can also start your assistant as part of a programmatic voice application using the [Start Assistant](/api-reference/call-commands/start-ai-assistant#start-ai-assistant) command. ## Third-party integrations By default, every component of a Telnyx AI Assistant runs on Telnyx infrastructure. You can, however, BYO LLM or TTS using third-party providers. ### Vapi integration If you have voice assistants configured in Vapi, you can [import them](https://developers.telnyx.com/docs/inference/ai-assistants/importing) as Telnyx AI Assistants in a single click. If you want to use a voice from Vapi in your existing Telnyx assistant, you will 1. Create a Vapi API Key. 2. Reference the key in your Assistant voice configuration. ##### Create a Vapi API Key First, check out their guide on [creating an API Key](https://docs.vapi.ai/chat/web-widget#1-get-your-public-api-key) ##### Reference the key in your Assistant voice configuration In the voice tab for your assistant, you can select Vapi as the provider. A new dropdown will then appear to reference your API key. You can give the secret a friendly identifier and securely store your API key as the secret value. To enable a multilingual agent, set the transcription model to `deepgram/nova-3`. *You will not be able to access the value of a secret after it is stored.* You can also manage all your secrets in the [Integration Secrets](https://portal.telnyx.com/#/integration-secrets) tab in the portal. Choose a memorable identifier to refer to the API key and store it in the `Secret Value` field. ![Integration Secret Portal Config](/assets/images/integration_secrets.png) ### ElevenLabs integration If you have Conversational AI agents configured in ElevenLabs, you can [import them](https://developers.telnyx.com/docs/inference/ai-assistants/importing) as Telnyx AI Assistants in a single click. If you want to use a voice from ElevenLabs in your existing Telnyx assistant, you will 1. Create an ElevenLabs API Key 2. Reference the key in your Assistant voice configuration *Requests from a free plan are rejected. You will likely have to use a paid plan to set up this integration successfully.* ##### Create an ElevenLabs API Key First, check out their guide on [creating an API Key](https://help.elevenlabs.io/hc/en-us/articles/14599447207697-How-to-authorize-yourself-using-your-xi-api-key) ##### Reference the key in your Assistant voice configuration In the voice tab for your assistant, you can select ElevenLabs as the provider. A new dropdown will then appear to reference your API key. You can give the secret a friendly identifier and securely store your API key as the secret value. To enable a multilingual agent, set the transcription model to `deepgram/nova-3`. *You will not be able to access the value of a secret after it is stored.* ![ElevenLabs Voice Config](/assets/images/eleven_labs_voice_secret.png) You can also manage all your secrets in the [Integration Secrets](https://portal.telnyx.com/#/integration-secrets) tab in the portal. Choose a memorable identifier to refer to the API key and store it in the `Secret Value` field. ![Integration Secret Portal Config](/assets/images/integration_secrets.png) ### OpenAI integration To use an LLM from OpenAI in your assistant, you will 1. Create an OpenAI API Key 2. Configure the language model in your AI Assistant *Requests from a free plan are rejected. You will likely have to use a paid plan to set up this integration successfully.* ##### Create an OpenAI API Key First, check out their guide on [creating an API Key](https://help.openai.com/en/articles/4936850-where-do-i-find-my-openai-api-key) ##### Configure the language model in your AI Assistant Back at the [AI Assistants tab](https://portal.telnyx.com/#/ai/assistants) in the portal, edit your assistant. First, change the model to an OpenAI model like `openai/gpt-4o`. Then follow the same API Key steps as described in the ElevenLabs section above. --- ### Realtime Conversations > Source: https://developers.telnyx.com/docs/inference/ai-assistants/realtime-conversations.md Open a real-time voice conversation with a [Telnyx AI Assistant](https://portal.telnyx.com/#/ai/assistants) over a single WebSocket. Your application streams microphone audio to Telnyx as base64-encoded PCM16 frames, and Telnyx streams back lifecycle events, speech-detection events, user and assistant transcripts, and the assistant's synthesized speech. Everything about the assistant's behavior — instructions, model, voice, language, tools, transcription and interruption settings — is configured on the assistant itself, in the [Portal](https://portal.telnyx.com/#/ai/assistants) or through the Assistants API. The WebSocket carries only the conversation: audio in, events and audio out. The one thing you can pass on the socket is [dynamic variables](#pass-dynamic-variables) for the conversation, and there is no frame to request a response — the assistant owns turn-taking and answers automatically when the user finishes speaking. ## How a conversation flows ``` Client → Telnyx (connect wss://.../v2/ai/assistants/{assistant_id}/conversation) Client → Telnyx {"type":"session.update","session":{...}} (optional) Client ← Telnyx {"type":"session.created", ...} Client → Telnyx {"type":"input_audio_buffer.append","audio":""} Client → Telnyx {"type":"input_audio_buffer.append","audio":""} Client ← Telnyx {"type":"input_audio_buffer.speech_started"} Client ← Telnyx {"type":"input_audio_buffer.speech_stopped"} Client ← Telnyx {"type":"conversation.item.input_audio_transcription.completed","transcript":"..."} Client ← Telnyx {"type":"response.created","response":{"id":"resp_9c1f4a2b"}} Client ← Telnyx {"type":"response.output_audio.delta","delta":"", ...} Client ← Telnyx {"type":"response.output_audio_transcript.delta","delta":"We are open ", ...} Client ← Telnyx {"type":"response.output_audio.done","response_id":"resp_9c1f4a2b"} Client ← Telnyx {"type":"response.done","response":{"id":"resp_9c1f4a2b","status":"completed"}} ``` 1. Open a WebSocket connection, optionally setting the audio format through query parameters. 2. Optionally send a `session.update` frame as your first message to pass [dynamic variables](#pass-dynamic-variables) for the conversation. 3. Telnyx authenticates the connection, starts a conversation with the requested assistant, and sends a `session.created` frame confirming the negotiated input and output audio formats. 4. Stream microphone audio as `input_audio_buffer.append` frames. 5. Telnyx runs server-side voice-activity detection (VAD) and sends `input_audio_buffer.speech_started` and `input_audio_buffer.speech_stopped` on speech edges, followed by the user's transcript. 6. The assistant responds with a `response.created` frame, a stream of audio and transcript deltas, then `response.output_audio.done` and `response.done`. 7. Speaking while the assistant is talking triggers barge-in: Telnyx stops the response and starts a new turn. ## Connect and authenticate ``` wss://api.telnyx.com/v2/ai/assistants/{assistant_id}/conversation ``` Authenticate the WebSocket handshake with a Telnyx API v2 key in the `Authorization` header: ``` Authorization: Bearer YOUR_TELNYX_API_KEY ``` Browsers cannot set headers on WebSocket connections, so connect from your backend. For browser-based voice agents, use the WebRTC-based [`@telnyx/ai-agent-lib`](https://www.npmjs.com/package/@telnyx/ai-agent-lib) library instead — it handles media capture, playback, and [client-side tools](/docs/inference/ai-assistants/client-side-tools) for you. Audio negotiation happens at connect time through query parameters: | Parameter | Description | |-----------|-------------| | `input_sample_rate` | Sample rate, in Hz, of the PCM16 audio you stream to Telnyx. One of `8000`, `16000`, `24000`, `44100`, `48000`. Default: `16000`. | | `input_format` | Encoding of the audio you stream to Telnyx. Only `pcm16` is supported (the default). | | `output_format` | Encoding of the assistant audio Telnyx streams back. Only `pcm16` is supported (the default). | | `output_sample_rate` | Preferred sample rate, in Hz, for the assistant audio. Advisory only — the effective output rate is determined by the assistant's voice and reported in `session.created`. | ```javascript Node.js import WebSocket from "ws"; const ASSISTANT_ID = "assistant-0f4e8b2a"; const url = `wss://api.telnyx.com/v2/ai/assistants/${ASSISTANT_ID}/conversation?input_sample_rate=16000`; const ws = new WebSocket(url, { headers: { Authorization: `Bearer ${process.env.TELNYX_API_KEY}` }, }); ws.on("open", () => console.log("Connected")); ``` ```python Python import asyncio import os import websockets ASSISTANT_ID = "assistant-0f4e8b2a" URL = f"wss://api.telnyx.com/v2/ai/assistants/{ASSISTANT_ID}/conversation?input_sample_rate=16000" async def main(): headers = {"Authorization": f"Bearer {os.environ['TELNYX_API_KEY']}"} async with websockets.connect(URL, additional_headers=headers) as ws: print("Connected") asyncio.run(main()) ``` All frames in both directions are JSON text messages with a `type` field. Audio travels inside JSON frames as base64 — never as binary WebSocket frames. ## Session lifecycle The first frame Telnyx sends is `session.created`. It confirms the conversation has started and reports the negotiated audio formats — read the output rate from it before playing any assistant audio: ```json { "type": "session.created", "session": { "conversation_id": "3E6F995F-85F7-4705-9741-53B116D28237", "assistant_id": "assistant-0f4e8b2a", "audio": { "input": { "format": { "type": "audio/pcm", "rate": 16000 }, "turn_detection": { "type": "server_vad" } }, "output": { "format": { "type": "audio/pcm", "rate": 24000 } } } } } ``` Each WebSocket connection starts a new conversation with the assistant. Turn detection is always `server_vad` — there is no manual turn mode to configure. Only assistants whose voice streams raw PCM are supported over this WebSocket. Assistants whose voice outputs a compressed format (for example mp3 or opus) are rejected with an `unsupported_voice_output_format` error. ## Pass dynamic variables If the assistant's instructions, greeting, or tools use [dynamic variables](/docs/inference/ai-assistants/dynamic-variables), supply the values for this conversation with a `session.update` frame. Send it as your **first** frame, immediately on `open` — before any audio, and without waiting for `session.created`: ```json { "type": "session.update", "session": { "assistant": { "dynamic_variables": { "customer_name": "Ada", "account_tier": "pro" } } } } ``` Telnyx starts the conversation as soon as the frame arrives, so the variables are resolved before the assistant's first word. If you never send it, Telnyx starts the conversation with no variables after a short timeout. Only `type` is required. A bare `{ "type": "session.update" }` — or one whose `session` object carries no `dynamic_variables` — is valid: it starts the conversation immediately with no client-supplied variables, skipping that timeout. ```javascript Node.js ws.on("open", () => { ws.send( JSON.stringify({ type: "session.update", session: { assistant: { dynamic_variables: { customer_name: "Ada", account_tier: "pro" }, }, }, }) ); }); ``` ```python Python async with websockets.connect(URL, additional_headers=headers) as ws: await ws.send(json.dumps({ "type": "session.update", "session": { "assistant": { "dynamic_variables": {"customer_name": "Ada", "account_tier": "pro"}, }, }, })) ``` Values are string key/value pairs, and they reach the assistant exactly like dynamic variables passed on the REST conversation API. `session.update` configures nothing but dynamic variables, and only before the conversation starts. Sending it after `session.created` is rejected with a `session_update_after_start` error, and a `dynamic_variables` value that is not an object is rejected with `invalid_session_update`. The `telnyx_` namespace is reserved for [system variables](/docs/inference/ai-assistants/dynamic-variables#telnyx-system-variables), so `telnyx_conversation_channel` is set by Telnyx and cannot be overridden. ## Handling audio ### Stream microphone audio Send audio as `input_audio_buffer.append` frames. The `audio` field is base64-encoded, raw little-endian PCM16 (16-bit signed, mono) at the sample rate you chose with `input_sample_rate`: ```json { "type": "input_audio_buffer.append", "audio": "" } ``` There is no commit step. Keep appending audio continuously — server-side VAD detects when the user starts and stops speaking and turns the buffered audio into conversation turns automatically. A few practical rules: - **Stream at a real-time pace**, in small chunks (for example 20–100 ms of audio per frame). The server enforces an ingress budget; sending far faster than real time fails with an `ingress_budget_exceeded` error. - **Keep frames under 1 MiB.** Larger frames are rejected with a `frame_too_large` error. - **Wait for `session.created`** before sending audio. `session.update` is the only frame that goes before it. ```javascript Node.js import fs from "node:fs"; const SAMPLE_RATE = 16000; const CHUNK_MS = 100; const CHUNK_BYTES = (SAMPLE_RATE * 2 * CHUNK_MS) / 1000; // 16-bit mono // speech.raw is raw little-endian PCM16 at 16 kHz const pcm = fs.readFileSync("speech.raw"); let offset = 0; const timer = setInterval(() => { if (offset >= pcm.length) { clearInterval(timer); return; } const chunk = pcm.subarray(offset, offset + CHUNK_BYTES); offset += CHUNK_BYTES; ws.send( JSON.stringify({ type: "input_audio_buffer.append", audio: chunk.toString("base64"), }) ); }, CHUNK_MS); ``` ```python Python import base64 import json SAMPLE_RATE = 16000 CHUNK_MS = 100 CHUNK_BYTES = SAMPLE_RATE * 2 * CHUNK_MS // 1000 # 16-bit mono async def stream_audio(ws): # speech.raw is raw little-endian PCM16 at 16 kHz with open("speech.raw", "rb") as f: while chunk := f.read(CHUNK_BYTES): await ws.send(json.dumps({ "type": "input_audio_buffer.append", "audio": base64.b64encode(chunk).decode(), })) await asyncio.sleep(CHUNK_MS / 1000) # real-time pacing ``` If your capture pipeline produces float samples (for example the Web Audio API), convert them to PCM16 before encoding: ```javascript function floatTo16BitPCM(float32Array) { const buffer = new ArrayBuffer(float32Array.length * 2); const view = new DataView(buffer); for (let i = 0; i < float32Array.length; i++) { const s = Math.max(-1, Math.min(1, float32Array[i])); view.setInt16(i * 2, s < 0 ? s * 0x8000 : s * 0x7fff, true); } return Buffer.from(buffer); } ``` ### Speech detection and user transcripts Telnyx runs voice-activity detection server-side and reports speech edges. Both events are edge-triggered — each is sent only when the speaking state actually changes, and `speech_stopped` is never sent without a preceding `speech_started`: ```json { "type": "input_audio_buffer.speech_started" } { "type": "input_audio_buffer.speech_stopped" } ``` After the user's turn ends, Telnyx sends the transcript of what they said: ```json { "type": "conversation.item.input_audio_transcription.completed", "transcript": "What are your business hours?" } ``` ### Receive assistant audio Each assistant turn arrives as an ordered sequence of frames, correlated by `response_id`: | Frame | Meaning | |-------|---------| | `response.created` | The assistant turn started. `response.id` correlates everything that follows. | | `response.output_audio.delta` | A chunk of synthesized speech — base64 PCM16 at the output rate from `session.created`. | | `response.output_audio_transcript.delta` | A chunk of the text transcript of the spoken response, streamed alongside the audio. | | `response.output_audio.done` | The assistant finished streaming audio for the turn. | | `response.done` | The turn is over, with a final `status` of `completed` or `cancelled`. | ```javascript Node.js let outputRate = 24000; const playback = []; ws.on("message", async (raw) => { const event = JSON.parse(raw.toString()); switch (event.type) { case "session.created": outputRate = event.session.audio.output.format.rate; break; case "conversation.item.input_audio_transcription.completed": console.log(`You said: ${event.transcript}`); break; case "response.output_audio.delta": // Feed to your audio output as it arrives for lowest latency playback.push(Buffer.from(event.delta, "base64")); break; case "response.output_audio_transcript.delta": process.stdout.write(event.delta); break; case "response.done": console.log(`\n[turn ${event.response.id}: ${event.response.status}]`); break; case "error": console.error(`${event.error.code}: ${event.error.message}`); break; } }); ``` ```python Python import base64 import json async def handle_events(ws): output_rate = 24000 playback = bytearray() async for raw in ws: event = json.loads(raw) match event["type"]: case "session.created": output_rate = event["session"]["audio"]["output"]["format"]["rate"] case "conversation.item.input_audio_transcription.completed": print(f"You said: {event['transcript']}") case "response.output_audio.delta": # Feed to your audio output as it arrives for lowest latency playback.extend(base64.b64decode(event["delta"])) case "response.output_audio_transcript.delta": print(event["delta"], end="", flush=True) case "response.done": print(f"\n[turn {event['response']['id']}: {event['response']['status']}]") case "error": print(f"{event['error']['code']}: {event['error']['message']}") ``` For a natural conversation, play deltas as they arrive rather than waiting for `response.done`. Buffer the decoded PCM16 in a queue that your audio output drains at the output sample rate. ## Interruption and barge-in Barge-in is built in. If the user speaks while the assistant is talking, Telnyx detects it, stops generating the response, and starts a new turn — you don't send anything to make that happen. The interrupted turn finishes with `response.done` and `status: "cancelled"`. One thing remains your responsibility: audio you have already received and queued locally. When you see `input_audio_buffer.speech_started` during playback, flush your local playback queue so the assistant doesn't keep talking out of your speakers over the user: ```javascript case "input_audio_buffer.speech_started": playback.length = 0; // drop queued assistant audio immediately break; ``` You can also interrupt programmatically — for example when the user taps a stop button — with `response.cancel`: ```json { "type": "response.cancel" } ``` Include `response_id` to target a specific response, or omit it to cancel the current one. Cancelling flushes playback on the Telnyx side and is followed by a `response.done` frame with status `cancelled`. Cancelling a stale or already-finished response is a no-op. Interruption sensitivity is tuned on the assistant, not the socket — see [Interruption Settings](/docs/inference/ai-assistants/interruption-settings). ## Send text instead of audio Inject a completed user turn as text with `conversation.item.create`. The assistant answers it exactly as it would a spoken turn — including responding with audio: ```json { "type": "conversation.item.create", "item": { "type": "message", "role": "user", "content": [{ "type": "input_text", "text": "What are your business hours?" }] } } ``` There is no `response.create` frame — the assistant owns turn-taking and answers automatically. Items of any other shape are rejected with an `invalid_item` error. ## Tool calls Assistants can use two kinds of tools during a realtime conversation. Both are configured on the assistant — see the [Tools Library](/docs/inference/ai-assistants/tools-library). ### Server-side tools: observe Webhook and MCP tools execute on Telnyx. The socket surfaces them as informational frames so you can show tool activity in your UI, but you don't execute or respond to them: ```json { "type": "response.tool_call.started", "tool_call": { "id": "call_7a3f21b8", "name": "get_business_hours", "arguments": "{\"location\":\"downtown\"}" } } ``` ```json { "type": "response.tool_call.completed", "tool_call": { "id": "call_7a3f21b8", "name": "get_business_hours" }, "status": "success" } ``` ### Client-side tools: execute and respond When the assistant invokes a [client-side tool](/docs/inference/ai-assistants/client-side-tools), Telnyx sends a `conversation.item.created` frame containing a `function_call` item. Run the tool locally, then return the result with a `conversation.item.create` frame carrying a `function_call_output` item that references the same `call_id`. The assistant continues once the result arrives, or after the tool times out. ```javascript Node.js case "conversation.item.created": { const item = event.item; if (item.type === "function_call") { const args = JSON.parse(item.arguments); const result = await runTool(item.name, args); // your implementation ws.send( JSON.stringify({ type: "conversation.item.create", item: { type: "function_call_output", call_id: item.call_id, output: JSON.stringify(result), }, }) ); } break; } ``` ```python Python case "conversation.item.created": item = event["item"] if item["type"] == "function_call": args = json.loads(item["arguments"]) result = await run_tool(item["name"], args) # your implementation await ws.send(json.dumps({ "type": "conversation.item.create", "item": { "type": "function_call_output", "call_id": item["call_id"], "output": json.dumps(result), }, })) ``` ## Error handling Errors arrive as `error` frames. Some errors are non-fatal and leave the session open; others close the connection after the frame is sent. ```json { "type": "error", "error": { "code": "invalid_audio", "message": "Invalid base64-encoded audio payload" } } ``` | Code | Meaning | |------|---------| | `session_not_ready` | A frame arrived before the session was ready. Wait for `session.created`. | | `frame_too_large` | A frame exceeded the 1 MiB limit. Send smaller audio chunks. | | `unsupported_event` | The `type` field is not a supported client frame. | | `invalid_json` | The frame was not valid JSON. | | `invalid_audio` | The `audio` field was not valid base64. | | `invalid_item` | A `conversation.item.create` item had an unsupported shape. | | `invalid_session_update` | A `session.update` frame carried a `dynamic_variables` value that was not an object. | | `session_update_after_start` | A `session.update` frame arrived after the conversation had already started. Send it as your first frame. | | `ingress_budget_exceeded` | Audio arrived faster than the session's ingress budget allows. Stream at a real-time pace. | | `unsupported_voice_output_format` | The assistant's voice streams a compressed format (for example mp3 or opus), which this WebSocket does not support. | | `conversation_start_failed` | Telnyx could not start a conversation with the requested assistant. | | `conversation_ended` | The conversation has ended. | | `session_idle_timeout` | The session was closed after a period of inactivity. | | `session_max_duration_exceeded` | The session reached its maximum duration. | If the connection closes, reconnecting starts a new conversation — a fresh `conversation_id` is issued in `session.created`. ## Learn more - **Conversation WebSocket reference** — The full frame-by-frame reference, under **Assistants API → Conversation WebSocket** in the sidebar - **[Client-Side Tools](/docs/inference/ai-assistants/client-side-tools)** — Configure tools that execute in your application - **[Transcription Settings](/docs/inference/ai-assistants/transcription-settings)** and **[Interruption Settings](/docs/inference/ai-assistants/interruption-settings)** — Tune how the assistant hears and yields - **[Voice Assistant Quickstart](/docs/inference/ai-assistants/no-code-voice-assistant)** — Create and configure an assistant in the Portal --- ### Importing Assistants > Source: https://developers.telnyx.com/docs/inference/ai-assistants/importing.md If you have voice assistants with another provider, you can import them to Telnyx in the portal or [via API](https://developers.telnyx.com/api-reference/assistants/import-assistants-from-external-provider#import-assistants-from-external-provider). ## Supported Providers Telnyx currently supports importing assistants from: - **Vapi** - Import your Vapi voice assistants with all configurations. - **ElevenLabs** - Import your ElevenLabs conversational AI agents. - **Retell** - Import your single- and multi-prompt Retell agents. --- ## Video Tutorial Watch this quick demonstration of importing a Vapi assistant to Telnyx: --- ## Import Walkthrough ### Importing from Vapi #### Viewing your Assistants From the [assistants page](https://portal.telnyx.com/#/ai/assistants), you can select **Import Assistants**. ![AI Assistant List Import](/assets/images/list-assistants-import-vapi.png) #### Selecting your Provider On the next page, select **Vapi** from the provider list and securely store your Vapi API Key with Telnyx to initiate the import. *Any assistant that has already been imported will be overwritten with its latest version from Vapi.* ![Import Assistants Page](/assets/images/import-assistants-vapi.png) #### Selecting your Assistants You can then select which Vapi assistants to import. *[Trial accounts](https://developers.telnyx.com/docs/account-setup/account-upgrade) have a limit of 1 AI Assistant.* ![Import Assistants Page](/assets/images/select-vapi-ai-import.png) #### Testing your imported Assistant You should now see your new assistants listed and can click the pencil icon to edit or the telephone icon to test them in the portal. Check out the [quickstart tutorial](https://developers.telnyx.com/docs/inference/ai-assistants/no-code-voice-assistant) for more information on editing or creating our Assistants. ### Importing from ElevenLabs #### Viewing your Assistants From the [assistants page](https://portal.telnyx.com/#/ai/assistants), you can select **Import Assistants**. ![AI Assistant List Import](/assets/images/list-assistants-import.png) #### Selecting your Provider On the next page, select **ElevenLabs** from the provider list and securely store your ElevenLabs API Key with Telnyx to initiate the import. *Any assistant that has already been imported will be overwritten with its latest version from ElevenLabs.* ![Import Assistants Page](/assets/images/import-ai.png) #### Selecting your Assistants You can then select which ElevenLabs agents to import. *[Trial accounts](https://developers.telnyx.com/docs/account-setup/account-upgrade) have a limit of 1 AI Assistant.* ![Import Assistants Page](/assets/images/select-ai-import.png) #### Testing your imported Assistant You should now see your new assistants listed and can click the pencil icon to edit or the telephone icon to test them in the portal. Check out the [quickstart tutorial](https://developers.telnyx.com/docs/inference/ai-assistants/no-code-voice-assistant) for more information on editing or creating our Assistants. --- # Supported Import Functionality ### Instructions Your assistant's instructions will be imported as is. ### Greeting (a.k.a. first message) Your assistant's greeting will be imported as is. ### LLM By default, the LLM assistant will be set to our flagship on-prem LLM, which provides optimal latency, cost, and intelligence. You can BYO LLM with third-party providers by storing an API Key. ### Voice When importing from Vapi or ElevenLabs, the voice will be imported as is. Otherwise, the LLM assistant will be set to our flagship on-prem TTS, which provides optimal latency, cost, and realism. You can BYO TTS with other third-party providers by storing an API Key. ### Dynamic Variables References to dynamic variables and their defaults will be imported as is. See our [guide](https://developers.telnyx.com/docs/inference/ai-assistants/dynamic-variables) on how to set these values per call. ### Tools Hangup, transfer, and webhook tools will be imported as is. ### MCP Servers (a.k.a. Integrations) Integrations with MCP Servers will be imported as is. ### Insights (a.k.a. analysis, success criteria, structured data...) Structured and unstructured call analysis configuration is imported as is. ### Data Retention If you have disabled data retention with another provider, this setting will be imported as is. ### Knowledge Bases Telnyx will not import your knowledge base by default. You can easily drag and drop files or import website content for your assistant in the assistant builder. ![Knowledge Base Page](/assets/images/knowledge-base-upload.png) ### Secrets Telnyx will create placeholder integration secrets for you with the same name as they are configured in the importing provider. You will need to resupply the value of the secret on the Telnyx [integration secrets page](https://portal.telnyx.com/#/integration-secrets). ![Knowledge Base Page](/assets/images/edit-integration-secret.png) --- ### GPT Live > Source: https://developers.telnyx.com/docs/inference/ai-assistants/gpt-live.md Use GPT Live as the speaking model in a Telnyx AI Assistant. GPT Live receives speech and generates speech in one model, so its voice latency is a single speech-to-speech estimate rather than the sum of separate transcription, language model, and speech synthesis stages. For lookups and actions, keep [delegation](/docs/inference/ai-assistants/delegation) enabled. GPT Live handles the conversation; a backend model or a server you host performs the work and returns the result. This guide configures a **Telnyx AI Assistant** through the Assistants API. For a SIP connection that routes calls directly to an OpenAI session controlled by your own sideband application, use [Connect Telnyx to GPT-Live over SIP](/docs/voice/sip-trunking/gpt-live-configuration-guide). These are separate integrations; the direct SIP setup does not configure an AI Assistant's `delegation_settings`. ## Configure the speaking model Use a GPT Live model available in the model selector for [AI Assistants in the Portal](https://portal.telnyx.com/#/ai/assistants). Store an OpenAI API key with access to that model as an [integration secret](/api-reference/integration-secrets/create-a-secret). Set these fields on the assistant: | Field | Configuration | |-------|---------------| | `model` | The managed GPT Live model ID, such as `openai/gpt-live-1`. | | `llm_api_key_ref` | The integration secret identifier for the OpenAI key used by the speaking model. | | `voice_settings.voice` | A matching `OpenAILive.` voice, such as `OpenAILive.marin`. | | `enabled_features` | Include `"telephony"` to enable calls. | | `greeting` | The opening message the assistant speaks when the conversation starts. | | `delegation_settings` | The backend configuration for lookups and actions. Set `enabled: true`, `mode: "telnyx"`, and an available backend `model`. | The model and voice must be paired: a GPT Live model with a non-live voice, or an `OpenAILive.` voice with a non-live model, is rejected when the assistant is saved. Configure GPT Live through the top-level `model`; `external_llm` selects an external text-model endpoint and takes precedence over it. Send this body to [Create an assistant](/api-reference/assistants/create-an-assistant), `POST /v2/ai/assistants`. Replace `openai-live-key` with the integration secret identifier and confirm that the model is available to your account: ```json { "name": "GPT Live support assistant", "enabled_features": ["telephony"], "model": "openai/gpt-live-1", "llm_api_key_ref": "openai-live-key", "instructions": "Help callers with order questions. Ask for the order number before a lookup. Never invent an order status or claim an action succeeded without a result.", "greeting": "Hi, I'm your order support assistant. How can I help you today?", "voice_settings": { "voice": "OpenAILive.marin" }, "delegation_settings": { "enabled": true, "mode": "telnyx", "model": "moonshotai/Kimi-K2.6", "instructions": "Use the configured order lookup tool to retrieve order status. If no lookup tool is available or the lookup fails, explain that the status cannot be verified.", "speak_results": true } } ``` The example explicitly selects `moonshotai/Kimi-K2.6` as the delegation backend. Confirm that this model is available to your account, or replace it with another supported backend model. It does not attach an order system: attach a shared lookup tool using the [tools library](/docs/inference/ai-assistants/tools-library) before testing order lookups. The top-level `llm_api_key_ref` authenticates the speaking model. If the delegation backend needs its own provider key, configure `delegation_settings.llm_api_key_ref` or `delegation_settings.external_llm.llm_api_key_ref` separately. Use secret references; delegation does not accept raw API keys. ## Configure delegation for work GPT Live cannot call the assistant's tools directly. When it needs a lookup or action, it raises a delegation and waits for the result. The result can stream back sentence by sentence for GPT Live to speak. Disabling delegation leaves a conversational assistant that cannot perform those lookups or actions. Choose the backend according to where the work runs: | Mode | Who performs the work | Configuration details | |------|-----------------------|-----------------------| | `telnyx` | A backend model uses the assistant's configured tools. | [Telnyx-managed delegation](/docs/inference/ai-assistants/delegation#choosing-a-mode) | | `client` | A server you host answers over the assistant's conversation event-stream WebSocket. | [Client delegation](/docs/inference/ai-assistants/delegation#choosing-a-mode) | The [delegation guide](/docs/inference/ai-assistants/delegation) owns the full settings reference, backend selection, result handling, and event payloads. For GPT Live, account for these differences: - **Backend instructions:** put detailed business rules, lookup procedures, and tool guidance in `delegation_settings.instructions`. These supplement the assistant's instructions for the backend; keep the speaking model's conversational instructions concise. - **OpenAI backend:** with `telnyx` mode and an OpenAI backend model, delegation runs through OpenAI and calls back into Telnyx for the assistant's tools. MCP servers are unavailable on that path. - **Client requests:** `session.delegation.created` has a `null` request on GPT Live. Build the request context from the conversation events already received on the socket. Return a non-empty text result with the matching delegation ID. - **Spoken results:** leave `speak_results: true` for an answer the caller is waiting for. Use `false` when the result should become silent context for later answers. ## Connect and verify the assistant Use the assistant ID with the existing [voice assistant calling setup](/docs/inference/ai-assistants/no-code-voice-assistant) or [realtime voice conversation WebSocket](/docs/inference/ai-assistants/realtime-conversations). A `client` delegation backend uses the separate [conversation event stream](/docs/inference/ai-assistants/conversation-event-stream): Telnyx connects to your server through `websocket_settings`. Verify both conversation and delegated work: 1. Confirm the assistant speaks its configured greeting, then ask a conversational question and verify the response uses its selected GPT Live voice. 2. Ask for information that requires a configured tool. Confirm the backend performs the lookup and the assistant speaks the returned result without inventing data. 3. In `client` mode, verify a delegation with `request: null` is answered from the streamed conversation context. Test an unavailable backend as well as a successful result. ## Next steps - [Delegation](/docs/inference/ai-assistants/delegation) — backend settings, spoken results, and the client event contract - [Conversation event stream](/docs/inference/ai-assistants/conversation-event-stream) — connect a server that receives events and answers client delegations - [Tools library](/docs/inference/ai-assistants/tools-library) — attach tools for delegated lookups and actions --- ### Custom LLMs for Assistants > Source: https://developers.telnyx.com/docs/inference/ai-assistants/custom-llm.md In addition to standard third-party LLM providers like OpenAI, Gemini, and Groq, you can also power your AI Assistant with any public OpenAI-compatible chat completions endpoint. This includes models hosted using AWS Bedrock, Azure OpenAI, and Baseten or open source inference engines like vLLM and SGLang. You must have a publicly accessible OpenAI-compatible chat completions endpoint before proceeding with deployment. --- ## When to use custom LLM providers Custom LLM providers are ideal for scenarios where you need: - **Specific model requirements** - Access to proprietary models, fine-tuned models, or the latest releases not yet available through standard providers. - **Data residency and compliance** - Ensure your data stays within specific geographic regions or private cloud environments. - **Cost optimization** - Leverage enterprise agreements, reserved capacity, or self-hosted infrastructure for better economics at scale. - **Advanced model control** - Fine-tune parameters, adjust inference settings, or use specialized configurations for your use case. --- ## FastRouter [FastRouter](https://fastrouter.ai/) is an OpenAI-compatible LLM gateway that can route requests to a fixed model or a [Virtual Model](https://docs.fastrouter.ai/explore-features/virtual-model-aliases). A Virtual Model is an alias that represents a model pool and routing strategy configured in FastRouter. ### Configure FastRouter with the API First, store the FastRouter API key with the [Create Integration Secret API](/api-reference/integration-secrets/create-a-secret). The secret's `identifier` is referenced by the assistant; the key itself is not included in the assistant configuration. ```bash curl -L -X POST 'https://api.telnyx.com/v2/integration_secrets' \ -H 'Authorization: Bearer YOUR_TELNYX_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ "identifier": "fastrouter-api-key", "type": "bearer", "token": "YOUR_FASTROUTER_API_KEY" }' ``` Then pass an `external_llm` object to the [Create Assistant API](/api-reference/assistants/create-an-assistant). For a fixed model, set `external_llm.model` to its `provider/model` ID: ```bash curl -L -X POST 'https://api.telnyx.com/v2/ai/assistants' \ -H 'Authorization: Bearer YOUR_TELNYX_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ "name": "FastRouter assistant", "instructions": "You are a helpful assistant.", "external_llm": { "model": "openai/gpt-4.1-mini", "base_url": "https://api.fastrouter.ai/api/v1", "llm_api_key_ref": "fastrouter-api-key", "authentication_method": "token" } }' ``` For a Virtual Model, use its alias as the same `external_llm.model` value—for example, `"model": "OMNIMODEL"`. To change an existing assistant, send the same `external_llm` object to `POST /v2/ai/assistants/{assistant_id}` with the [Update Assistant API](/api-reference/assistants/update-an-assistant). Do not set the top-level `model` or `llm_api_key_ref` fields for this configuration. Those fields select Telnyx-managed third-party models; FastRouter uses `external_llm.model` and `external_llm.llm_api_key_ref`. ### Configure FastRouter in the Portal To configure FastRouter for an AI Assistant: 1. Create a FastRouter API key. 2. Create or edit a Telnyx AI Assistant and open the `Agent` tab. 3. Set `LLM Provider` to `FastRouter`. 4. Under `FastRouter API Key`, select or create an Integration Secret containing your FastRouter API key. 5. Choose a `Routing Mode`: - `Fixed model` is selected by default. Choose a model from the searchable catalog. Fixed model IDs use `provider/model` format. - Select `Virtual model` to enter the exact `Virtual Model Alias` from FastRouter, such as `OMNIMODEL`. 6. Select `Validate LLM connection`. 7. Save the assistant. FastRouter does not include Virtual Model aliases in its [Models API](https://docs.fastrouter.ai/api-reference/models). Enter Virtual Model aliases manually. The fixed-model dropdown loads from the Models API after you select an Integration Secret. Telnyx automatically uses FastRouter's OpenAI-compatible API base URL. You do not need to configure a custom base URL. --- ## Azure For this guide, we will deploy [gpt-4o on Azure AI Foundry](https://ai.azure.com/). First, create or select a resource. ![Azure AI Foundry homepage showing resource selection interface](/assets/images/foundry-homepage.png) Then, drill into the resource to find the API key and Azure OpenAI endpoint. ![Azure deployment details page displaying API key and endpoint URL](/assets/images/foundry-deployment.png) Create or edit a Telnyx AI Assistant and in the `Agent` tab, check `Use Custom LLM`. Input the endpoint URL as the `Base URL` and append `/openai/v1` and create a new Integration Secret with your API Key. You will see a dropdown of all possible Azure models but only ones that you have deployed will validate an LLM connection. ![Telnyx portal Agent tab with Azure custom LLM configuration and model dropdown](/assets/images/azure-in-portal.png) Once you save your assistant you will be able to immediately use your assistant in the Telnyx portal with the `Test Assistant` dropdown. ## Forward metadata to your custom LLM By default, Telnyx does not include your assistant's dynamic variables in requests to a custom LLM endpoint. If your model gateway or application needs those values, enable `forward_metadata` on the assistant's `external_llm` configuration. ```json { "external_llm": { "llm_api_key_ref": "integration_secret_id", "base_url": "https://your-llm-gateway.example.com/openai/v1", "model": "your-model-name", "forward_metadata": true } } ``` When `forward_metadata` is `true`, Telnyx adds a top-level `extra_metadata` object to the OpenAI-compatible chat completions request body sent to your custom LLM endpoint when dynamic variables are available. The field defaults to `false` when omitted. ```json { "model": "your-model-name", "messages": [ { "role": "system", "content": "..." }, { "role": "user", "content": "..." } ], "extra_metadata": { "customer_name": "Jane", "account_id": "acct_789", "telnyx_agent_target": "+13125550100", "telnyx_end_user_target": "+13125550123" } } ``` Use this when your external LLM service needs request context for routing, logging, personalization, or retrieval. `extra_metadata` is separate from OpenAI's native `metadata` field, so your endpoint must explicitly read `extra_metadata` from the request body. ## Baseten For this guide, we will deploy [Llama 3.3 70B on Baseten](https://www.baseten.co/library/llama-3-3-70b-instruct/). First, click `Deploy Now`. ![Baseten model library page showing Llama 3.3 70B with Deploy Now button](/assets/images/baseten-deploy-now.png) Navigate to the deployment. ![Baseten deployment overview page for Llama 3.3 70B model](/assets/images/baseten-deployment.png) Click the API Endpoint button to see the endpoint and generate an API Key. Save these details. ![Baseten deployment page showing API endpoint URL and key generation button](/assets/images/baseten-api-key.png) After about 15 minutes, the deployment should be live. When it is complete, create or edit a Telnyx AI Assistant and in the `Agent` tab, check `Use Custom LLM`. Input the Baseten Endpoint URL for your deployment as the `Base URL` and create a new Integration Secret with your Baseten API Key. ![Telnyx portal custom LLM configuration showing Base URL input field](/assets/images/custom-url.png) If your base URL supports an OpenAI-compatible /models endpoint the Model Name dropdown will populate automatically. Baseten deployments do not support this endpoint, so you can enter any name for your model here. You can also validate the connection is live before saving your assistant. ![Telnyx portal custom LLM validation interface with connection test button](/assets/images/validate-custom-llm.png) Once you save your assistant you will be able to immediately use your assistant in the Telnyx portal with the `Test Assistant` dropdown as well as review metrics in your Baseten deployment. ![Baseten deployment dashboard displaying usage metrics and performance data](/assets/images/baseten-metrics.png) --- ### Transcription Settings > Source: https://developers.telnyx.com/docs/inference/ai-assistants/transcription-settings.md Telnyx AI Assistants support multiple speech-to-text (STT) models for transcribing caller audio. The model you choose affects transcription accuracy, supported languages, and response latency. You can also tune provider-specific transcription behavior, such as end-of-turn detection, formatting, keyterm boosting, and Azure region selection. --- ## Available models | Model | Engine | Best for | | --- | --- | --- | | `deepgram/flux` | Deepgram | Conversational AI, optimized for turn-taking with multilingual support | | `deepgram/nova-3` | Deepgram | Fast multilingual transcription, recommended for multilingual assistants | | `deepgram/nova-2` | Deepgram | Fast multilingual transcription on Deepgram's previous-generation model | | `azure/fast` | Azure | Fast multilingual transcription with optional Azure region and API key configuration | | `assemblyai/universal-3-5-pro` | AssemblyAI | Conversational, multilingual streaming transcription with configurable turn detection, backed by Universal-3.5 Pro Realtime. The legacy alias `assemblyai/universal-streaming` resolves to the same model | | `xai/grok-stt` | xAI | Multilingual transcription using Grok STT | | `soniox/stt-rt-v4`, `soniox/stt-rt-v5` | Soniox | Multilingual streaming transcription with automatic language detection, term biasing, and endpoint detection | | `nvidia/parakeet-v3` | Parakeet | Multilingual transcription with automatic language detection | | `omi-health/omi-med-stt-v1` | Parakeet | English-only medical transcription (Omi Health, Parakeet-based) | | `cohere/ar-stt` | Cohere | Self-hosted Arabic/English transcription. Requires an explicit `ar` or `en` language — no auto-detect | | `reson8/turns` | Reson8 | Turn-based transcription of 10 European languages with automatic language detection | `deepgram/flux` supports English, Spanish, French, German, Hindi, Russian, Portuguese, Japanese, Italian, and Dutch. For broader language coverage, use `deepgram/nova-3`, `deepgram/nova-2`, `azure/fast`, `assemblyai/universal-3-5-pro`, `xai/grok-stt`, or `nvidia/parakeet-v3`. --- ## Selecting a model ### Portal In the [AI Assistants tab](https://portal.telnyx.com/#/ai/assistants), edit your assistant and navigate to the **Voice** tab. Select your preferred STT model from the **Transcription Model** dropdown. ![AI Assistant Transcription Model Selection](/assets/images/ai-assistant-transcription-model.png) When you change models in the Portal, related settings are reset to the defaults for that provider. For example: - `deepgram/flux` supports explicit languages, `auto`, and `multi` for its supported languages, and applies Flux end-of-turn defaults. - Other Deepgram models enable `smart_format` and `numerals` by default. - `assemblyai/universal-3-5-pro` applies AssemblyAI turn detection defaults. - `soniox/stt-rt-v4` and `soniox/stt-rt-v5` disable endpoint detection by default; term biasing and language hints are left unset. - `azure/fast` defaults the Azure region to `latency`, which auto-selects the closest supported Telnyx-managed region. - `nvidia/parakeet-v3` uses automatic multilingual transcription. - `omi-health/omi-med-stt-v1` pins the language to `en`. - `cohere/ar-stt` defaults the language to `ar`. Set it to `en` explicitly if needed — unlike other models, it does not auto-detect and rejects `auto`. - `reson8/turns` defaults the language to `auto` for automatic detection. Setting an explicit language reduces latency. ### API Set the `transcription.model` field when creating or updating an assistant: ```bash curl -X POST https://api.telnyx.com/v2/ai/assistants \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "name": "My Assistant", "model": "anthropic/claude-haiku-4-5", "instructions": "You are a helpful voice assistant.", "transcription": { "model": "deepgram/flux" } }' ``` You can also set the transcription language explicitly. If omitted or set to `auto`, supported models auto-detect the language: ```json "transcription": { "model": "deepgram/nova-3", "language": "es" } ``` --- ## Languages Supported language options depend on the selected model. | Model | Language behavior | | --- | --- | | `deepgram/flux` | English, Spanish, French, German, Hindi, Russian, Portuguese, Japanese, Italian, Dutch, plus `auto` and `multi` modes | | `deepgram/nova-3` | Auto-detect plus supported Deepgram Nova 3 language codes | | `deepgram/nova-2` | Auto-detect plus supported Deepgram Nova 2 language codes | | `azure/fast` | Explicit Azure locale codes, such as `en-US`, `es-MX`, or `fr-FR` | | `assemblyai/universal-3-5-pro` | `auto` for multilingual detection, or one of 18 supported languages: `en`, `es`, `de`, `fr`, `pt`, `it`, `tr`, `nl`, `sv`, `no`, `da`, `fi`, `hi`, `vi`, `ar`, `he`, `ja`, `zh` | | `xai/grok-stt` | `auto` plus supported Grok STT language codes | | `soniox/stt-rt-v4`, `soniox/stt-rt-v5` | `auto` (or unset) omits the language hint and lets Soniox auto-detect; ISO 639-1 codes (`en`, `es`, `nl`, `fr`, and others) bias detection toward that language. Use `language_hints` in `settings` to pin multiple languages at once. | | `nvidia/parakeet-v3` | Automatic multilingual detection | | `cohere/ar-stt` | `ar` (default) or `en` only — no auto-detect. `auto` is rejected. | | `reson8/turns` | `auto` (default) for automatic detection, or one of `nl`, `en`, `fr`, `fy`, `de`, `it`, `pl`, `pt`, `es`, `sv` | If your assistant has a language filter set elsewhere in the Voice tab, the Portal only shows transcription models and language choices that are compatible with that language. --- ## Deepgram settings ### Deepgram Flux end-of-turn detection `deepgram/flux` is optimized for live voice agents. It provides end-of-turn detection so the assistant can start responding as soon as the caller finishes speaking. It also supports eager end-of-turn, which starts large language model (LLM) processing before the caller fully stops speaking to reduce perceived response latency. When you select `deepgram/flux` in the Portal, these settings are applied by default: | Field | Type | Range | Portal default | Description | | --- | --- | --- | --- | --- | | `eot_threshold` | number | 0.5-0.9 | 0.8 | Confidence required to trigger a final end of turn. Higher values require more confidence and may add latency. | | `eot_timeout_ms` | integer | 500-10000 | 5000 | Maximum silence duration, in milliseconds, before forcing an end of turn. | | `eager_eot_threshold` | number | 0.3-0.9 | 0.4 | Confidence required to start speculative LLM processing before final end-of-turn confirmation. Lower values trigger earlier. | `eager_eot_threshold` must be less than or equal to `eot_threshold`. Setting both thresholds to the same value effectively disables eager end-of-turn behavior because the system waits for final end-of-turn confirmation before starting LLM processing. The `eager_eot_threshold` field is controlled by the `FE-eager-eot-threshold` Portal feature flag. When that flag is disabled, the Portal hides the field, but API payloads can still include it if your account supports the setting. When using Flux, the Portal may also prompt you to lower the assistant's start speaking plan timings. Flux works best with low start speaking delays, such as `0.1` seconds for wait time and endpointing plan thresholds. ### Keyterm Boost `deepgram/flux` and `deepgram/nova-3` support `keyterm`, a comma-separated list of terms to boost during recognition. Use it for product names, customer names, acronyms, or domain-specific vocabulary. Keyterm Boost also supports [dynamic variables](/docs/inference/ai-assistants/dynamic-variables). Use variables when boosted terms are caller-specific, such as a customer name, participant names, account name, or product names passed into the assistant at conversation start. ```json "transcription": { "model": "deepgram/nova-3", "settings": { "keyterm": "Telnyx,VoIP,SIP,{{customer_name}},{{product_name}}" } } ``` ### Smart Format and Numerals For Deepgram models other than Flux, the Portal exposes these settings and enables both by default when you select the model: | Field | Type | Default | Description | | --- | --- | --- | --- | | `smart_format` | boolean | `true` | Automatically formats transcripts for readability, including punctuation and casing. | | `numerals` | boolean | `true` | Converts spoken numbers to digits, for example "five hundred" to "500". | --- ## AssemblyAI settings `assemblyai/universal-3-5-pro` supports configurable turn detection. When you select it in the Portal, these defaults are applied: | Field | Type | Range | Portal default | Description | | --- | --- | --- | --- | --- | | `end_of_turn_confidence_threshold` | number | 0-1 | 0.4 | Confidence required to trigger an end of turn. Higher values require more certainty before ending a turn. | | `min_turn_silence` | integer | 100-5000 | 400 | Minimum silence duration, in milliseconds, before a turn can end. | | `max_turn_silence` | integer | 100-5000 | 1280 | Maximum silence duration, in milliseconds, before forcing an end of turn. | `min_turn_silence` must be less than or equal to `max_turn_silence`. ```json "transcription": { "model": "assemblyai/universal-3-5-pro", "language": "auto", "settings": { "end_of_turn_confidence_threshold": 0.4, "min_turn_silence": 400, "max_turn_silence": 1280 } } ``` --- ## Soniox settings Soniox offers two model versions: `soniox/stt-rt-v4` and `soniox/stt-rt-v5`. Both support multilingual streaming transcription with automatic language detection. ### Term Biasing `soniox/stt-rt-v4` and `soniox/stt-rt-v5` support `context`, a comma-separated list of terms to boost during recognition. Use it for staff names, building names, product names, or other domain-specific vocabulary. Term Biasing also supports [dynamic variables](/docs/inference/ai-assistants/dynamic-variables). Use variables when boosted terms are caller-specific, such as a customer name, participant names, account name, or product names passed into the assistant at conversation start. ```json "transcription": { "model": "soniox/stt-rt-v5", "settings": { "context": "Telnyx,VoIP,SIP,{{customer_name}},{{product_name}}" } } ``` ### Language Hints `language_hints` pins recognition to one or more languages at once, as an array of ISO 639-1 codes. Use it when an assistant regularly serves callers in a known set of languages, such as Dutch and French, instead of relying on full auto-detection. ```json "transcription": { "model": "soniox/stt-rt-v5", "settings": { "language_hints": ["nl", "fr"] } } ``` ### Endpoint detection | Field | Type | Range | Default | Description | | --- | --- | --- | --- | --- | | `interim_results` | boolean | — | `false` | Stream interim (non-final) transcripts in addition to final ones. Useful for live captions or low-latency UI feedback. | | `enable_endpoint_detection` | boolean | — | `false` | When enabled, Soniox emits end-of-utterance events at the cadence configured by `max_endpoint_delay_ms`. | | `max_endpoint_delay_ms` | integer | 500-3000 | — | Maximum silence, in milliseconds, before Soniox emits an end-of-utterance event. Only honored when `enable_endpoint_detection` is `true`. | ```json "transcription": { "model": "soniox/stt-rt-v5", "settings": { "enable_endpoint_detection": true, "max_endpoint_delay_ms": 1200 } } ``` --- ## Parakeet settings `nvidia/parakeet-v3` supports multilingual transcription with automatic language detection. It does not require provider-specific transcription settings. ```json "transcription": { "model": "nvidia/parakeet-v3", "language": "auto" } ``` --- ## Reson8 settings `reson8/turns` delivers transcripts per turn of speech: the assistant receives the full transcript of a turn when the caller finishes speaking. Language defaults to `auto` for automatic detection and can be fixed to one of the 10 supported languages. It does not require provider-specific transcription settings. ```json "transcription": { "model": "reson8/turns", "language": "auto" } ``` --- ## Cohere settings `cohere/ar-stt` is a self-hosted model for Arabic and English transcription. Unlike other models, it does not auto-detect the language and requires an explicit `ar` or `en` — `auto` (and any other value) is rejected. Omitting `language` defaults to `ar`. It does not require provider-specific transcription settings. ```json "transcription": { "model": "cohere/ar-stt", "language": "ar" } ``` --- ## Azure settings `azure/fast` supports region selection and an optional Azure API key reference. | Field | Type | Description | | --- | --- | --- | | `region` | string | Azure transcription region. The Portal defaults to `latency`, which auto-selects the closest supported Telnyx-managed region. | | `api_key_ref` | string | Optional integration secret reference for your Azure API key. When provided, the Portal only shows regions that support custom API keys. | Common Telnyx-managed regions include `latency`, `australiaeast`, `centralindia`, `eastus`, `northcentralus`, `westeurope`, and `westus2`. Additional Azure regions are available when using your own Azure API key. ```json "transcription": { "model": "azure/fast", "language": "en-US", "region": "latency" } ``` --- ## Configure advanced settings via API Create an assistant with tuned transcription settings: ```bash curl -X POST https://api.telnyx.com/v2/ai/assistants \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "name": "Low Latency Assistant", "model": "anthropic/claude-haiku-4-5", "instructions": "You are a helpful voice assistant.", "transcription": { "model": "deepgram/flux", "language": "en", "settings": { "eot_threshold": 0.8, "eot_timeout_ms": 5000, "eager_eot_threshold": 0.4, "keyterm": "Telnyx,VoIP,SIP,{{customer_name}},{{product_name}}" } } }' ``` --- ## Related resources - [Create an Assistant API Reference](/api-reference/assistants/create-an-assistant) - [Voice Assistant quickstart](/docs/inference/ai-assistants/no-code-voice-assistant) - [STT WebSocket Streaming](/docs/voice/stt/websocket-streaming) --- ### Interruption Settings > Source: https://developers.telnyx.com/docs/inference/ai-assistants/interruption-settings.md Interruption settings control whether and how callers can interrupt your AI assistant while it is speaking. By default, interruptions are enabled — callers can barge in at any point. You can disable interruptions entirely, protect only the greeting from being cut off, or tune the endpointing thresholds that determine when the assistant treats caller speech as a complete turn. In this guide, you will learn: - How interruption handling differs between turn-taking and non turn-taking transcription models. - How to configure interruption settings via the API. - How to protect greetings and tune endpointing thresholds. --- ## How it works Interruption behavior depends on which transcription model your assistant uses. ### Turn-taking models Turn-taking models like `deepgram/flux` have built-in end-of-turn detection. When you use one of these models, the assistant determines when the caller has finished speaking using the transcription-level settings `eot_threshold`, `eot_timeout_ms`, and `eager_eot_threshold` — configured under `transcription.settings`. See the [Transcription Settings](/docs/inference/ai-assistants/transcription-settings) guide for details. You can still use `interruption_settings.enable` and `interruption_settings.disable_greeting_interruption` with turn-taking models, but the `start_speaking_plan` and `transcription_endpointing_plan` thresholds are not relevant because the model handles turn detection natively. ### Non turn-taking models For models without built-in turn detection (such as `deepgram/nova-3`, `deepgram/nova-2`, `azure/fast`, or `nvidia/parakeet-v3`), the assistant relies on the `interruption_settings.start_speaking_plan` to decide when to start speaking after the caller stops. This plan uses silence-based endpointing thresholds to detect end of turn. --- ## Configuration ### API schema Set the `interruption_settings` object when creating or updating an assistant: ```bash curl -X POST https://api.telnyx.com/v2/ai/assistants \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "name": "My Assistant", "model": "anthropic/claude-haiku-4-5", "instructions": "You are a helpful voice assistant.", "transcription": { "model": "deepgram/nova-3" }, "interruption_settings": { "enable": true, "disable_greeting_interruption": true, "start_speaking_plan": { "wait_seconds": 0.4, "transcription_endpointing_plan": { "on_punctuation_seconds": 0.1, "on_no_punctuation_seconds": 1.5, "on_number_seconds": 0.5 } } } }' ``` ### Key fields | Field | Type | Default | Description | |-------|------|---------|-------------| | `enable` | boolean | `true` | Whether callers can interrupt the assistant while it is speaking. | | `disable_greeting_interruption` | boolean | — | When `true`, prevents callers from interrupting the assistant's greeting. | | `start_speaking_plan.wait_seconds` | float | `0.4` | Minimum seconds to wait after the caller stops speaking before the assistant responds. | | `start_speaking_plan.transcription_endpointing_plan.on_punctuation_seconds` | float | `0.1` | Seconds to wait after the transcript ends with punctuation before treating the turn as complete. | | `start_speaking_plan.transcription_endpointing_plan.on_no_punctuation_seconds` | float | `1.5` | Seconds to wait after the transcript ends without punctuation before treating the turn as complete. | | `start_speaking_plan.transcription_endpointing_plan.on_number_seconds` | float | `0.5` | Seconds to wait after the transcript ends with a number before treating the turn as complete. | ### Mission Control Portal You can also configure interruption settings through the Portal: 1. Navigate to **AI > Assistants**. 2. Select your assistant and click **Edit**. 3. Open the **Voice** tab. 4. Scroll to the **Interruption Settings** section. 5. Toggle interruptions on or off, enable greeting protection, and adjust the start speaking plan thresholds. --- ## Example configurations ### Default — allow interruptions Allow callers to interrupt at any time, including during the greeting. This is the default behavior when `interruption_settings` is omitted. ```json { "interruption_settings": { "enable": true } } ``` ### Protected greeting Allow interruptions during normal conversation but prevent callers from cutting off the assistant's initial greeting. Useful for flows where the greeting contains required disclosures or context. ```json { "interruption_settings": { "enable": true, "disable_greeting_interruption": true } } ``` ### Conservative endpointing for non turn-taking models Use longer silence thresholds to reduce false interruptions in noisy environments or when callers tend to pause mid-sentence. This configuration waits longer before treating silence as end of turn. ```json { "interruption_settings": { "enable": true, "start_speaking_plan": { "wait_seconds": 0.8, "transcription_endpointing_plan": { "on_punctuation_seconds": 0.3, "on_no_punctuation_seconds": 2.0, "on_number_seconds": 1.0 } } } } ``` --- ## Use cases - **Call centers with background noise**: Increase `wait_seconds` and endpointing thresholds to prevent ambient noise from triggering false interruptions. - **IVR-style interactions**: Keep interruptions enabled so callers can barge in to skip prompts they have already heard. - **Greeting-heavy flows**: Enable `disable_greeting_interruption` to ensure callers hear required disclosures, legal notices, or welcome messages before interacting. - **Dictation or number entry**: Increase `on_number_seconds` to give callers time to finish reading out long numbers like account IDs or phone numbers without the assistant cutting in. --- ## Related resources - **[Transcription Settings](/docs/inference/ai-assistants/transcription-settings)** — Configure STT models and end-of-turn detection for turn-taking models like `deepgram/flux`. - **[Create an Assistant API Reference](/api-reference/assistants/create-an-assistant)** — Full assistant configuration options including `interruption_settings`. - **[End-of-Turn Detection](/docs/voice/stt/websocket-streaming/parameters/end-of-turn)** — Lower-level end-of-turn parameters for WebSocket STT streaming. --- ### Dynamic Variables > Source: https://developers.telnyx.com/docs/inference/ai-assistants/dynamic-variables.md Dynamic variables let you configure a template for your agent's behavior. You can re-use the same general instructions while dynamically personalizing every conversation your agent has. In this tutorial, you will learn how to: - Template your AI assistant configuration with dynamic variables. - Supply the values via an outbound API request or the dynamic variable webhook. --- ## Overview Dynamic variables enable you to create a single AI assistant configuration that can handle personalized conversations for different users and contexts. Instead of creating multiple assistants for different scenarios, you can use placeholders that get filled with specific values at runtime. ### How dynamic variables work The dynamic variable lifecycle follows these steps: 1. **Define**: Create placeholders in your assistant's instructions, greeting, or tools using `{{variable_name}}` syntax. 2. **Inject**: Provide values through API calls, webhooks, SIP headers, or default configurations. 3. **Resolve**: Telnyx replaces placeholders with actual values when the conversation starts. 4. **Use**: Your assistant uses the personalized content throughout the conversation. ### Key benefits - **Scalability**: One assistant configuration serves multiple use cases. - **Personalization**: Each conversation can be tailored to specific users or contexts. - **Efficiency**: Reduce configuration overhead and maintenance complexity. - **Flexibility**: Update values dynamically without modifying the assistant. ### Common use cases - **Customer Service**: Personalize greetings with customer names and account details. - **Appointment Scheduling**: Include facility names, departments, and contact information. - **Account Management**: Reference specific account numbers, balances, or service details. - **Healthcare**: Customize interactions with patient names, appointment types, and provider information. --- ## Best practices Following these best practices will help you implement dynamic variables effectively and avoid common pitfalls. ### Variable naming and organization - **Use descriptive names**: Choose clear, meaningful variable names like `customer_name` instead of generic names like `name`. - **Follow consistent conventions**: Use snake_case for variable names (`facility_name`, `account_number`). - **Group related variables**: Organize logically related variables together (`facility_name`, `facility_department`, `facility_contact`). - **Avoid reserved namespaces**: Don't use the `telnyx_` prefix, which is reserved for system variables. ```json // Good variable naming { "customer_name": "Sarah Johnson", "account_number": "ACC-12345", "facility_name": "Memorial Hospital", "facility_department": "Cardiology" } // Poor variable naming { "name": "Sarah", "num": "12345", "place": "Hospital", "dept": "Card" } ``` ### Implementation strategy - **Start with defaults**: Set default values in the Assistant builder for testing and as fallback values. - **Use API injection for dynamic data**: Pass call-specific values through the `AIAssistantDynamicVariables` parameter in API calls. - **Leverage webhooks for complex lookups**: Use the dynamic variables webhook for real-time data retrieval or database lookups. - **Implement graceful degradation**: Ensure your assistant can handle scenarios where variables fail to resolve. ### Security and data handling - **Protect sensitive information**: Be cautious about including sensitive data in variables that may appear in logs. - **Validate webhook data**: Implement proper validation for data returned from webhook endpoints. - **Consider data retention**: Review your data retention and privacy requirements for variable content. - **Handle timeouts gracefully**: Webhook responses must return within the configured timeout (default 1.5 seconds, up to 10 seconds via `dynamic_variables_webhook_timeout_ms`) or the call proceeds with fallback values. ### Performance and reliability - **Optimize webhook response time**: Keep webhook responses fast (well under your configured timeout; the default is 1.5 seconds). - **Implement error handling**: Plan for scenarios where webhooks fail or return invalid data. - **Use caching strategies**: Cache frequently accessed data to improve response times. - **Test across channels**: Verify variable resolution works consistently across phone calls, web calls, and SMS. ### Testing and debugging - **Test with Portal defaults**: Use the Assistant builder's default values to test variable behavior during development. - **Monitor webhook logs**: Use the conversation transcript and webhook logs in the Portal to debug variable resolution issues. - **Validate fallback behavior**: Ensure unresolved variables (displayed as `{{variable_name}}`) don't break conversation flow. - **Test transfer scenarios**: Verify that variables pass correctly when using transfer tools with SIP headers. ### Cross-channel considerations - **Ensure universal compatibility**: Test that your variables work across phone calls, web calls, and SMS channels. - **Handle channel-specific data**: Consider how variables might behave differently across conversation channels. - **Plan for transfer scenarios**: When transferring calls, use SIP headers to pass variable data to the receiving assistant. --- ## Dynamic variables syntax Dynamic variables are placeholders surrounded by double curly braces. For example: ```json Hello, this is ABC Ambulance. Am I speaking with {{full_name}} with {{facility_name}} - {{facility_department}}? ``` You can template them in the following fields: - Instructions - Greeting - Tools ## Telnyx system variables Telnyx also provides these system variables: | Variable | Description | Example | |----------|-------------|---------| | `{{telnyx_current_time}}` | The current date and time in UTC | `Monday, February 24 2025 04:04:15 PM UTC` | | `{{telnyx_conversation_channel}}` | This can be `phone_call` , `web_call`, or `sms_chat` | `phone_call` | | `{{telnyx_agent_target}}` | The phone number, SIP URI, or other identifier associated with the agent. | `+13128675309` | | `{{telnyx_end_user_target}}` | The phone number, SIP URI, or other identifier associated with the end user | `+15551234567` | | `{{telnyx_shaken_stir_attestation}}` | The SHAKEN/STIR attestation level for inbound phone calls, if available. This can be `a`, `b`, or `c`. | `a` | | `{{call_control_id}}` | The call control ID for the call, if applicable | `v3:u5OAKGEPT3Dx8SZSSDRWEMdNH2OripQhO` | ### Date and time variables In addition to the default `{{telnyx_current_time}}`, Telnyx provides a family of reserved date/time variables. Use these to give the assistant timezone-aware context, convenient shorthands, or a fully custom format. The timezone variants and shorthands use the same default formatting as `{{telnyx_current_time}}` — a different format is only used when you explicitly specify one with the `date` filter. | Variable | Description | Example output | |----------|-------------|----------------| | `{{telnyx_current_time_}}` | Current date and time, in the given timezone | `Monday, February 24 2025 11:04:15 AM EST` | | `{{telnyx_current_date}}` | Current date (UTC) | `February 24 2025` | | `{{telnyx_current_time_of_day}}` | Current time of day (UTC) | `04:04:15 PM UTC` | | `{{telnyx_current_time_of_day_}}` | Current time of day, in the given timezone | `11:04:15 AM EST` | | `{{telnyx_current_weekday}}` | Current day of the week (UTC) | `Monday` | | `{{telnyx_current_month}}` | Current month name (UTC) | `February` | | `{{telnyx_current_day}}` | Current day of the month (UTC) | `24` | | `{{telnyx_current_year}}` | Current year (UTC) | `2025` | The shorthand variables (`telnyx_current_date`, `telnyx_current_time_of_day`, `telnyx_current_weekday`, `telnyx_current_month`, `telnyx_current_day`, `telnyx_current_year`) are evaluated in UTC. To render them for a specific timezone, use the `_` variants or the `date` filter described below. #### Timezone variants Append an [IANA timezone name](https://en.wikipedia.org/wiki/List_of_tz_database_time_zones) to `telnyx_current_time_` or `telnyx_current_time_of_day_` to render the value in that zone. These use the same default formatting as `{{telnyx_current_time}}`: - `{{telnyx_current_time_America/New_York}}` → `Monday, February 24 2025 11:04:15 AM EST` - `{{telnyx_current_time_Australia/Sydney}}` → `Tuesday, February 25 2025 03:04:15 AM AEDT` - `{{telnyx_current_time_of_day_Europe/London}}` → `04:04:15 PM GMT` #### Custom formatting with the `date` filter The default format is used everywhere unless you explicitly request a different one. For full control over formatting, pipe `telnyx_current_time` through the `date` filter. The filter takes a standard [`strftime`](https://docs.python.org/3/library/datetime.html#strftime-and-strptime-format-codes) format string and an optional IANA timezone. These are the same `%`-style codes used by `strftime` across most languages (C, Python, Ruby, and others): | Template | Example output | |----------|----------------| | `{{ telnyx_current_time \| date: "%A, %B %d, %Y" }}` | `Monday, February 24, 2025` (UTC) | | `{{ telnyx_current_time \| date: "%I:%M %p", "America/New_York" }}` | `11:04 AM` (in the given zone) | | `{{ telnyx_current_time \| date: "%A, %B %d, %Y at %I:%M %p %Z", "America/Los_Angeles" }}` | `Monday, February 24, 2025 at 08:04 AM PST` | When the timezone argument is omitted, the value is formatted in UTC. ##### Common format codes These are the most useful `strftime` codes for the `date` filter. For the complete list, see any [standard `strftime` reference](https://docs.python.org/3/library/datetime.html#strftime-and-strptime-format-codes): | Code | Meaning | Example | |------|---------|---------| | `%A` | Weekday name | `Monday` | | `%a` | Abbreviated weekday | `Mon` | | `%B` | Month name | `February` | | `%b` | Abbreviated month | `Feb` | | `%d` | Day of month (zero-padded) | `24` | | `%Y` | Four-digit year | `2025` | | `%I:%M %p` | 12-hour time | `11:04 AM` | | `%H:%M` | 24-hour time | `16:04` | | `%Z` | Timezone abbreviation | `PST` | | `%z` | UTC offset | `-0800` | Combine codes with any literal text, e.g. `%A, %B %d at %I:%M %p %Z`. If a timezone is unknown, a format string is invalid, or a variable is unrecognized, the template placeholder is left untouched rather than rendered to an empty or error value. ## Customer-defined dynamic variables You can also define your own variables. There are several ways to resolve a customer-defined dynamic variable before a conversation. They follow this order of precedence. On a [realtime WebSocket conversation](/docs/inference/ai-assistants/realtime-conversations#pass-dynamic-variables), pass the values in a `session.update` frame as your first message instead of on an outbound API call. The webhook and assistant-level defaults below still apply. **1. Pass the variables via the outbound API call** ```bash curl --request POST \ --url https://api.telnyx.com/v2/texml/ai_calls/ \ --header "Authorization: Bearer $TELNYX_API_KEY" \ --header 'Content-Type: application/json' \ --data '{ "From": "+13128675309", "To": "+15551234567", "AIAssistantId": "assistant-6207ab25-b185-478f-b2ef-85159e226727", "AIAssistantDynamicVariables": { "full_name": "James Smith", "facility_name": "Cleveland Clinic Main Campus", "facility_department": "Emergency Department" } }' ``` ![AI Assistant Outbound Call Variables](/assets/images/call-dynamic-vars.png) **2. Custom SIP Headers** Custom SIP headers using the `X-` prefix will be mapped to dynamic variables. Header names are case-insensitive, and `-` in the header names will be replaced by `_` in the variable names. For instance, `X-Full-Name` will be resolved as `{{full_name}}`. *Telnyx reserves the `{{telnyx_}}` namespace for system variables, meaning SIP headers beginning `X-Telnyx` will not be resolved to dynamic variables.* Here you can see an example of passing a dynamic variable to another AI Assistant using custom SIP headers in the Transfer tool. ![AI Assistant SIP Header Config](/assets/images/ai-assistant-custom-sip-variable.png) **3. Configure Dynamic Variables Webhook** If the `dynamic_variables_webhook_url` is set for the assistant, we will `POST` the following payload at the start of the conversation. Telnyx will [sign this webhook](https://developers.telnyx.com/docs/development/api-fundamentals/webhooks/receiving-webhooks#webhook-signing) so that the authenticity of the request can be verified. ![AI Assistant Variable Config](/assets/images/configure-dynamic.png) ``` { "data": { "record_type": "event", "id": "event_12345678-90ab-cdef-1234-567890abcdef", "event_type": "assistant.initialization", "occurred_at": "2025-04-07T10:00:00Z", "payload": { "telnyx_conversation_channel": "phone_call", "telnyx_agent_target": "+13128675309", "telnyx_end_user_target": "+15551234567", "telnyx_end_user_target_verified": false, "call_control_id": "v3:u5OAKGEPT3Dx8SZSSDRWEMdNH2OripQhO", "assistant_id": "assistant_12345678-90ab-cdef-1234-567890abcdef" } } } ``` *For inbound phone calls to an assistant, the `telnyx_end_user_target_verified` field will be set to `true` if the call has Full (A) [STIR/SHAKEN attestation](https://support.telnyx.com/en/articles/5402969-stir-shaken-with-telnyx) and Telnyx was able to verify the authenticity of the PASSporT token.* We expect a JSON response with the following structure. If we do not receive this response within the configured timeout (default 1.5 seconds, up to 10 seconds via the `dynamic_variables_webhook_timeout_ms` field on the assistant), the call will proceed "best effort" with the alternatives listed below. All fields (`dynamic_variables`, `memory`, `conversation`, and `encrypted_dynamic_variables`) are optional. The `dynamic_variables` field sets the values for the specified dynamic variables. You can read more about the `memory` and `conversation` fields in our [tutorial on Memory](https://developers.telnyx.com/docs/inference/ai-assistants/memory). To pass a credential that differs per caller — an MCP bearer token or a tool webhook credential belonging to the end user — return it encrypted in the `encrypted_dynamic_variables` section instead. See [Per-Caller Credentials](/docs/inference/ai-assistants/per-caller-credentials). Values in `dynamic_variables` are substituted into prompts and are not suitable for secrets. ```json { "dynamic_variables": { "full_name": "Rachel Thomas", "facility_name": "UCHealth", "facility_department": "Cardiology", "account_number": "ACC-12345", "appointment_time": "2:30 PM" }, "memory": { "conversation_query": "metadata->telnyx_end_user_target=eq.+13128675309&limit=5&order=last_message_at.desc" }, "conversation": { "metadata": { "customer_tier": "premium", "preferred_language": "en", "timezone": "America/Denver" } } } ``` **4. Configure default dynamic variables in the Assistant builder** In the above view, you can also set default values for dynamic variables at the agent level. These will serve as a fallback if any of the above mechanisms do not properly resolve a variable. **5. Unset variables** Variables not set by any of the above methods will remain in their raw form (`{{full_name}}`). ## Troubleshooting Dynamic variables can sometimes fail to resolve or behave unexpectedly. This section provides comprehensive guidance for identifying and resolving common issues. ### Common issues #### Variables not resolving When variables appear as raw text (`{{variable_name}}`) in conversations, check these potential causes: - **Webhook timeout**: Your webhook endpoint took longer than the configured timeout (default 1.5 seconds) to respond, causing fallback to defaults. - **Variable name mismatch**: Typos or case sensitivity differences between variable definition and usage. - **Incorrect precedence**: Understanding the resolution order (API > SIP headers > webhook > defaults). - **Missing configuration**: Variable not defined in any resolution method. #### Incorrect formatting Variables may not work correctly due to formatting issues: - **Invalid JSON**: Webhook responses contain malformed JSON syntax. - **Wrong data types**: Sending numbers instead of strings or vice versa. - **Missing required fields**: Webhook response missing `dynamic_variables` object. - **Template syntax errors**: Incorrect mustache template usage (`{variable}` instead of `{{variable}}`). #### Timeout issues Webhook timeouts are a common source of variable resolution failures: - **Slow database queries**: Optimize database lookups in webhook endpoints. - **Network connectivity**: Check firewall rules and network connectivity to webhook URLs. - **External service delays**: Implement proper timeout handling for third-party API calls. - **Processing overhead**: Minimize computation time in webhook response generation. ### Debugging checklist Follow this step-by-step process to diagnose variable issues: #### 1. Verify variable definition - Check variable name spelling and case sensitivity. - Confirm variables are used in supported fields (`instructions`, `greeting`, `tools`). - Validate mustache syntax: `{{variable_name}}`. - Ensure variable names don't use reserved `telnyx_` prefix. #### 2. Check resolution priority Test variables using the resolution hierarchy: ```bash # Test API injection (highest priority) curl --request POST \ --url https://api.telnyx.com/v2/texml/ai_calls/ \ --header "Authorization: Bearer $TELNYX_API_KEY" \ --header 'Content-Type: application/json' \ --data '{ "From": "+13128675309", "To": "+15551234567", "AIAssistantId": "assistant-6207ab25-b185-478f-b2ef-85159e226727", "AIAssistantDynamicVariables": { "test_variable": "API_VALUE" } }' ``` - Verify SIP header format for transfer scenarios: `X-Variable-Name` becomes `{{variable_name}}`. - Confirm webhook endpoint is reachable and responding correctly. - Set default values in Assistant builder as fallback. #### 3. Validate webhook implementation - Test webhook endpoint independently using tools like Postman. - Verify JSON response format matches expected structure. - Check response time (must be under the configured timeout; default is 1.5 seconds). - Validate all required payload fields are present. #### 4. Test in Portal - Use Assistant builder default values for initial testing. - Review conversation logs for variable resolution behavior. - Check webhook logs for errors, timeouts, or unexpected responses. ### Webhook debugging #### Using webhook logs effectively To facilitate troubleshooting webhook issues, logs are viewable per conversation in the portal alongside the conversation transcript and insights. ![AI Assistant Webhook Logs](/assets/images/dynamic-webhook-log.png) **Understanding webhook logs:** - **Request logs**: Show the exact payload Telnyx sent to your webhook - **Response logs**: Display your webhook's response and any errors - **Timing information**: Reveals if responses exceeded the configured timeout - **Error details**: Provide specific error messages for failed requests #### Common webhook issues **Authentication failures:** ```json // Webhook response indicating auth failure { "error": "Unauthorized: Invalid API key", "dynamic_variables": {} } ``` **Malformed JSON responses:** ```json // Incorrect - missing quotes around keys { dynamic_variables: { customer_name: "John Doe" } } // Correct JSON format { "dynamic_variables": { "customer_name": "John Doe" } } ``` **Timeout handling:** ```json // Webhook response for timeout scenarios { "dynamic_variables": { "customer_name": "Default Name" }, "message": "External service timeout, using cached data" } ``` #### Testing webhook endpoints **Independent webhook testing:** ```bash # Test your webhook with the exact payload Telnyx sends curl --request POST \ --url https://your-webhook-url.com/dynamic-variables \ --header 'Content-Type: application/json' \ --data '{ "data": { "record_type": "event", "id": "event_12345678-90ab-cdef-1234-567890abcdef", "event_type": "assistant.initialization", "occurred_at": "2025-04-07T10:00:00Z", "payload": { "telnyx_conversation_channel": "phone_call", "telnyx_agent_target": "+13128675309", "telnyx_end_user_target": "+15551234567", "call_control_id": "v3:test-call-control-id" } } }' ``` ### Variable validation #### Input validation **Sanitize variable values:** ```json // Good - Clean, safe variable values { "dynamic_variables": { "customer_name": "John Smith", "account_number": "ACC-12345", "appointment_time": "2:30 PM" } } // Avoid - Potentially problematic values { "dynamic_variables": { "customer_name": "", "account_number": "'; DROP TABLE users; --", "appointment_time": null } } ``` **Handle special characters properly:** - Escape quotes and backslashes in variable content. - Validate data types match expected formats. - Provide meaningful defaults for empty or null values. #### Response format validation **Ensure proper JSON structure:** ```json // Complete webhook response format { "dynamic_variables": { "customer_name": "Sarah Johnson", "facility_name": "Memorial Hospital" }, "memory": { "conversation_query": "metadata->customer_id=eq.12345" }, "conversation": { "metadata": { "customer_tier": "premium" } } } ``` #### Testing scenarios **Cross-channel testing:** - Verify variables work consistently across phone calls, web calls, and SMS. - Test variable behavior in transfer scenarios using SIP headers. - Validate fallback behavior when webhook endpoints are unavailable. **Load testing:** - Test webhook endpoint performance under expected call volumes. - Verify response times remain under your configured timeout (default 1.5 seconds) during peak usage. - Monitor for memory leaks or performance degradation over time. --- ### Conversation Keying > Source: https://developers.telnyx.com/docs/inference/ai-assistants/conversation-keying.md Every interaction between an AI Assistant and an end user is stored as a conversation in [AI Conversations](/api-reference/conversations/list-conversations). Conversations are keyed by channel and by the two endpoints of the conversation — the assistant's target (a phone number, SIP URI, or other identifier) and the end user's target — and the same assistant can have several independent conversations with the same end user, one per channel. ## Keying model per channel | Channel | Conversation granularity | Created by | | ------- | ----------------------- | ---------- | | `phone_call` | One conversation per phone call, inbound or outbound. | The platform, at call setup. | | `sms_chat` | One long-lived conversation per (assistant number, end-user number) pair. Texts between the same pair continue in the same conversation. | The platform, at the first text. | | `whatsapp_chat` | WhatsApp texts, marked with this channel by the messaging system. | The platform. | | `web_chat` | One conversation per `conversation_id` you pass to the [chat endpoint](/api-reference/assistants/create-assistant-chat-completion). | Your application, via the Assistants API. | | `websocket_call` | One conversation per realtime [WebSocket connection](/docs/inference/ai-assistants/realtime-conversations). | The platform, at socket start. | | `web_call` | Assistant test runs executed from the portal test simulator. | A [test run](/docs/inference/ai-assistants/scheduled-events) or the portal. | ## Conversation metadata Each conversation carries identifying metadata you can read on every [list](/api-reference/conversations/list-conversations) and [get](/api-reference/conversations/get-conversation) response: | Field | Present on | Description | | ----- | --------- | ----------- | | `assistant_id` | All channels | The assistant the conversation belongs to. | | `telnyx_conversation_channel` | All channels | The channel the conversation runs on — see the table above. | | `telnyx_agent_target` | All channels | The assistant's target: the phone number the assistant answers on, or another identifier. | | `telnyx_end_user_target` | All channels | The end user's target: the caller's phone number, or another identifier. | | `call_control_id`, `call_session_id`, `call_leg_id` | `phone_call` | The Call Control identifiers of the call the conversation belongs to. | The same fields are available to the assistant's instructions and tools as [dynamic variables](/docs/inference/ai-assistants/dynamic-variables): `{{telnyx_conversation_channel}}`, `{{telnyx_agent_target}}`, and `{{telnyx_end_user_target}}`. ## One pair, one conversation per channel A single number pair — one assistant number and one end-user number — can have several concurrent or sequential conversations, one per channel. A phone call to the assistant's number creates a `phone_call` conversation for that call, while texts with the same number pair continue in their own long-lived `sms_chat` conversation. The two records are independent: closing the call does not close the text conversation, and the assistant's [memory](/docs/inference/ai-assistants/memory) can span both by querying on `telnyx_end_user_target` without a channel filter. ## Querying conversations by channel or participant The [List Conversations endpoint](/api-reference/conversations/list-conversations) filters on any metadata field using PostgREST-style parameters. All SMS conversations with one end user: ``` GET /v2/ai/conversations?metadata->telnyx_conversation_channel=eq.sms_chat&metadata->telnyx_end_user_target=eq.+13125550123 ``` Every conversation — calls and texts — between one number pair: ``` GET /v2/ai/conversations?metadata->telnyx_agent_target=eq.+13125550100&metadata->telnyx_end_user_target=eq.+13125550123 ``` Phone-call conversations created in a window: ``` GET /v2/ai/conversations?metadata->telnyx_conversation_channel=eq.phone_call&created_at=gte.2026-09-01&created_at=lt.2026-10-01 ``` Message rows carry channel metadata too: a turn recorded with `telnyx_conversation_channel: sms_chat` inside a conversation whose own channel is `phone_call` did not come from the call. When inspecting a transcript, read the channel on both the conversation and each message. --- ### Memory > Source: https://developers.telnyx.com/docs/inference/ai-assistants/memory.md Memory enables your AI assistant to recall essential details from past conversations. Instead of starting each phone call or text exchange from scratch, your AI assistant naturally continues previous discussions. In this tutorial, you will learn how to: - Specify which conversations your AI Assistant has memory access to - Configure this dynamically at the start of every conversation *Telnyx Assistants natively support our Voice and Messaging APIs, meaning the same assistant can seamlessly remember conversations across channels.* --- ## Identifying the conversations to include There is no one-size-fits-all answer for which previous conversations an AI Assistant should remember during a specific conversation. You may want an AI Assistant to have memory access to: - Every conversation it had with any user - Every conversation it had with **this specific user** - Every conversation it had with **users in a specific group** in **the past 10 days** - Or something else entirely... To support this, we have exposed a flexible query language to give customers full control over their assistant's memory. Any query you can build with our [List Conversations endpoint](/api-reference/conversations/list-conversations), you can use to configure memory access. ## Configuring memory with the Dynamic Variables Webhook If the `dynamic_variables_webhook_url` is set for the assistant, we will send the following payload at the start of the conversation. ``` { "data" :{ "record_type": "event", "id": "event_id", "event_type": "assistant.initialization", "occurred_at": "2025-04-07T10:00:00Z", "payload": { "telnyx_conversation_channel": "phone_call", "telnyx_agent_target": "+1234567890", "telnyx_end_user_target": "+1234567890", "telnyx_end_user_target_verified": false } } } ``` *For inbound phone calls to an assistant, the `telnyx_end_user_target_verified` field will be set to `true` if the call has Full (A) [STIR/SHAKEN attestation](https://support.telnyx.com/en/articles/5402969-stir-shaken-with-telnyx) and Telnyx was able to verify the authenticity of the PASSporT token.* We expect a JSON response with the following structure. If we do not receive this response within a 1-second timeout, the call will proceed "best effort". ``` { "dynamic_variables": { "full_name": "Rachel Thomas", "facility_name": "UCHealth", "facility_department": "Cardiology" }, "memory": { "conversation_query": "metadata->telnyx_end_user_target=eq.+13128675309&limit=5&order=last_message_at.desc" } } ``` In this example, the optional `memory` field provides your AI assistant with memory access to the last 5 conversations with the current user's phone number. You can read more about the optional `dynamic_variables` field in our [tutorial on Dynamic Variables](https://developers.telnyx.com/docs/inference/ai-assistants/dynamic-variables). ![AI Assistant Variable Config](/assets/images/configure-dynamic.png) In addition to controlling which conversations are remembered, you can customize which information from a conversation is remembered. To do this, specify a comma-delimited list of insight IDs in the memory field. Insight IDs can be retrieved in the Insights tab for your assistant, as shown below. Only the results from the insights you specify will be included in the assistant's memory. ``` { "dynamic_variables": { "full_name": "Rachel Thomas", "facility_name": "UCHealth", "facility_department": "Cardiology" }, "memory": { "conversation_query": "metadata->telnyx_end_user_target=eq.+13128675309&limit=5&order=last_message_at.desc", "insight_query": "insight_ids=123,456"" } } ``` ![AI Assistant Insight Copy](/assets/images/copy-insight-id.png) ## Custom Metadata You may want to create your own memory access system based on custom metadata for conversations. To do this, you can add metadata to conversations in the dynamic variable webhook response: ``` { "dynamic_variables": { "full_name": "Rachel Thomas", "facility_name": "UCHealth", "facility_department": "Cardiology" }, "memory": { "conversation_query": "metadata->telnyx_end_user_target=eq.+13128675309&limit=5&order=last_message_at.desc" }, "conversation": { "metadata": { "your_custom_metadata": "your_custom_value" } } } ``` In future conversations, you can filter that metadata in the `memory` field using the following syntax `metadata->your_custom_metadata=eq.your_custom_value`. --- ### Tools Library > Source: https://developers.telnyx.com/docs/inference/ai-assistants/tools-library.md The Tools Library lets you create tools in a shared library and assign them to any assistant. Previously, tools were tied to individual assistants — if multiple assistants needed the same tool, you had to recreate it each time. --- ## Overview With the Tools Library, you define a tool once and reuse it across all your assistants. Updating a shared tool keeps behavior consistent across every assistant that uses it. --- ## Capabilities - **Shared tools** — Create tools in a central library and assign them to any assistant. - **Reduced duplication** — Eliminates recreating identical tools across assistants with similar workflows. - **Centralized configuration** — Update a tool in one place and maintain consistency across assistants. - **Faster deployment** — Speed up assistant creation using prebuilt tools from the library. All tool types are supported in the library, including [webhook tools](/docs/inference/ai-assistants/async-tools), [client-side tools](/docs/inference/ai-assistants/client-side-tools), handoff, transfer, and hangup tools. --- ## Getting started 1. Log in to the [Mission Control Portal](https://portal.telnyx.com). 2. Navigate to **AI, Storage and Compute** > **[AI Tools](https://portal.telnyx.com/#/ai/tools)**. 3. Click **Create New Tool** to add a new tool to the library. 4. When building or editing an assistant, assign tools from the library instead of creating them inline. ![Tools Library overview](/assets/images/tools-library-overview.png) You can also assign library tools directly from the assistant builder. --- ## Updating assistants with library tools When you update an assistant with `POST /v2/ai/assistants/{assistant_id}` (the [Update Assistant API](/api-reference/assistants/update-an-assistant) also accepts `PATCH`), `tools` (inline definitions) and `tool_ids` (library tools) are two independent channels: - A `tools` array you send **fully replaces** the assistant's inline tools; omit the field to leave them unchanged. The same applies to `tool_ids` for attached library tools. - Every tool type except `function`, `webhook`, and `client_side_tool` allows **at most one instance per assistant**, counted across inline `tools` and library `tool_ids` combined. Sending a duplicate of such a type returns HTTP 400 with error code `10015`. - GET responses merge library tools into the `tools` array, each flagged `"shared": true` (the flag is read-only — the server sets it; it is not accepted in requests). On update, **omit `shared: true` tools from the `tools` array** and manage them through `tool_ids` instead — re-sending their definitions creates an inline duplicate, rejected with `10015` when the type allows only one instance. --- ## Migrating existing tools Existing assistants with legacy (inline) tools continue to work. You can migrate legacy tools to the library at your own pace. Migration is optional. Legacy tools remain functional. You can migrate tools one at a time as needed. --- ### Integrations > Source: https://developers.telnyx.com/docs/inference/ai-assistants/integrations.md Telnyx AI assistants can integrate with leading enterprise platforms to access customer data, create tickets, update records, and automate workflows directly during conversations. ## Available integrations The Integrations tab in the assistant builder offers a growing catalog of enterprise connectors, organized by category. Use the category filter chips at the top of the **Add Integration** section to narrow the list, or browse everything below. ### Sales & CRM | Integration | Description | |-------------|-------------| | **Salesforce** | Manage leads, contacts, accounts, opportunities, tasks, and cases. | | **HubSpot** | Manage contacts, companies, deals, tickets, and custom objects. | | **Pipedrive** | Manage deals, contacts, organizations, activities, and pipelines. | | **Zoho CRM** | Manage leads, contacts, accounts, deals, tasks, and notes. | | **Gong** | Revenue intelligence: call analysis, deal insights, and user management. | ### Customer Support | Integration | Description | |-------------|-------------| | **Zendesk** | Manage tickets, users, organizations, and support operations. | | **Intercom** | Manage contacts, companies, conversations, and help center articles. | | **ServiceNow** | Enterprise ITSM: incidents, change requests, problems, and service catalog. | | **Jira** | Issue tracking, comments, transitions, and project management. | | **Jira Service Management** | Service desks, customer management, and request lifecycle. | ### Engineering & Product | Integration | Description | |-------------|-------------| | **GitHub** | Repository, issue, and pull request management. | | **Jira** | Issue tracking, comments, transitions, and project management. | | **Linear** | Project management: issues, projects, and team workflows. | ### IT Operations | Integration | Description | |-------------|-------------| | **ServiceNow** | Enterprise ITSM: incidents, change requests, problems, and service catalog. | | **Jira Service Management** | Service desks, customer management, and request lifecycle. | | **Microsoft Teams** | Manage teams, channels, messages, members, and file operations. | ### Work Management | Integration | Description | |-------------|-------------| | **Asana** | Manage tasks, projects, teams, workspaces, and custom fields. | | **Airtable** | Manage bases, tables, records, fields, comments, and attachments. | | **Notion** | Manage databases, pages, blocks, users, and content. | ### Knowledge & Documentation | Integration | Description | |-------------|-------------| | **Confluence** | Manage pages, spaces, content, and attachments. | | **Notion** | Manage databases, pages, blocks, users, and content. | | **SharePoint** | Manage sites, document libraries, lists, and content. | | **GitHub** | Repository, issue, and pull request management. | ### Communication & Collaboration | Integration | Description | |-------------|-------------| | **Microsoft Teams** | Manage teams, channels, messages, members, and file operations. | | **Outlook** | Manage email, calendar events, contacts, and mailbox organization. | ### File Storage & Productivity | Integration | Description | |-------------|-------------| | **OneDrive** | Manage files, folders, sharing, and metadata. | | **Outlook** | Manage email, calendar events, contacts, and mailbox organization. | | **SharePoint** | Manage sites, document libraries, lists, and content. | ### HR & Recruiting | Integration | Description | |-------------|-------------| | **Greenhouse** | ATS: candidates, jobs, applications, scorecards, and recruiting workflows. | | **SAP SuccessFactors** | Manage employees, time off, performance goals, positions, and recruiting. | ### Scheduling | Integration | Description | |-------------|-------------| | **Calendly** | Manage scheduling, events, invitees, event types, and organization operations. | ### Design & UX | Integration | Description | |-------------|-------------| | **Figma** | Manage files, projects, teams, comments, components, styles, and dev resources. | ### Accounting & Finance | Integration | Description | |-------------|-------------| | **QuickBooks Online** | Accounting: customers, invoices, payments, bills, and financial reporting. | ### E-commerce & Payments | Integration | Description | |-------------|-------------| | **Shopify** | Manage products, orders, customers, inventory, and store operations. | | **Stripe** | Payment processing: customers, charges, refunds, subscriptions, and invoices. | ### Testing & Evaluation | Integration | Description | |-------------|-------------| | **Coval** | Simulation and evaluation for voice and chat agents — automated testing, regression detection, and production monitoring. | The integration catalog is expanding regularly. The **Add Integration** section in the assistant builder always reflects the current, complete list — if you see a platform there that isn't documented below, it works the same way: connect it, enter the required credentials, and enable the tools you need. ## Getting started ### Prerequisites Before connecting an integration, make sure you have: Active access with the right permissions on the target platform. Platform-specific keys or tokens required by the integration. A configured assistant ready to connect the integration. ### Connection workflow Navigate to [AI Assistants](https://portal.telnyx.com/#/ai/assistants) and open an existing assistant or create a new one. Select the **Integrations** tab for that assistant. In the **Add Integration** section, optionally filter by category using the chips (for example *Sales & CRM* or *Customer Support*), pick a provider, and enter the required credentials. Enable the tools you need, adjust defaults, and save the assistant. ![Mission Control Portal showing the Add Integration section with category filter chips and a grid of available integration platforms](/assets/images/ai-integration-dropdown-full.jpeg) ## See integrations in action Watch how Telnyx Voice AI Agents connect to enterprise platforms like ServiceNow to create, update, and resolve tickets through natural voice conversation: This demonstration shows the integration workflow in the Mission Control Portal and real-time ticket management capabilities that work across all supported platforms. ## Platform-specific setup Select your integration platform below. If you don't see your platform, scroll horizontally to view all available options. ### GitHub Connect your AI assistant to GitHub for code hosting, version control, and development workflows. #### Prerequisites - GitHub account with repository access. - Permissions to create personal access tokens. - Appropriate repository scopes for integration needs. #### Required credentials When connecting GitHub, you'll need to provide: Generate a GitHub Personal Access Token at https://github.com/settings/tokens by clicking **Generate new token → Generate new token (classic)** and selecting scopes such as `repo`, `read:user`, and `read:org`, then copy the token (it only appears once). ![GitHub integration credentials dialog showing Personal Access Token input field](/assets/images/github-integration-credentials.png) #### Available tools After connecting GitHub, your assistant can: - **Create issues** - Log bugs, feature requests, or tasks in repositories. - **Search repositories** - Find repos by name, description, or topics. - **Manage pull requests** - Create, review, or update PRs. - **Access code** - Read file contents and repository structure. - **Manage labels** - Add or remove issue and PR labels. #### Use cases **Bug Reporting** ``` Developer: "We have a login bug affecting Safari users" Assistant: [Creates issue in GitHub repository] Assistant: [Adds labels: bug, priority:high, browser:safari] Assistant: "I've created issue #234 in the main repo and notified the team." ``` **Repository Search** ``` Developer: "Find our authentication libraries" Assistant: [Searches GitHub repos by keyword and topic] Assistant: "I found 3 authentication repos: auth-service, jwt-utils, and oauth-client." ``` ### Greenhouse Connect your AI assistant to Greenhouse for applicant tracking and recruiting workflows. #### Prerequisites - Greenhouse account with API access. - Access to Greenhouse Dev Center. - Harvest API permissions. #### Required credentials When connecting Greenhouse, you'll need to provide: Log in to Greenhouse, go to **Dev Center → API Credential Management**, create a Harvest API key with the permissions you need, and copy the generated token. ![Greenhouse integration credentials dialog showing API Key input field](/assets/images/greenhouse-integration-credentials.png) #### Available tools After connecting Greenhouse, your assistant can: - **Candidate lookup** - Search for candidates by name, email, or application ID. - **Interview scheduling** - View and coordinate interview schedules. - **Application management** - Track application status and progress. - **Job posting access** - View open positions and job details. - **Scorecard review** - Access candidate evaluations and feedback. #### Use cases **Candidate Status Check** ``` Recruiter: "What's the status of the candidate who interviewed yesterday?" Assistant: [Searches Greenhouse by interview date] Assistant: "Sarah Johnson completed her technical interview yesterday. She's in the Reference Check stage." ``` **Interview Coordination** ``` Recruiter: "Schedule the next round for top candidates" Assistant: [Retrieves candidate list and interview availability] Assistant: [Checks interviewer calendars] Assistant: "I can schedule 3 final interviews for next Tuesday and Wednesday." ``` ### HubSpot Connect your AI assistant to HubSpot for marketing, sales, and customer service workflows. #### Prerequisites - HubSpot account with API access. - Private app access token or OAuth credentials. #### Required credentials When connecting HubSpot, you'll need to provide: Create a private app in HubSpot (Settings → Integrations → Private Apps) and copy the access token from the **Auth** tab. ![HubSpot integration credentials dialog showing Private app token input field](/assets/images/hubspot-integration-credentials.png) #### Available tools After connecting HubSpot, your assistant can: - **Manage contacts** - Create, update, or search contacts. - **Deal tracking** - Create deals, update deal stages. - **Ticket management** - Create support tickets, update status. - **Company records** - Access and update company information. - **Engagement tracking** - Log calls, emails, and notes. #### Use cases **Lead Capture** ``` Prospect: "I'd like a demo of your product" Assistant: [Creates contact in HubSpot] Assistant: [Creates deal in pipeline] Assistant: [Schedules demo meeting] ``` **Support Ticketing** ``` Customer: "I have a billing question" Assistant: [Searches HubSpot for customer record] Assistant: [Creates ticket in support pipeline] Assistant: [Associates ticket with contact and deal] ``` ### Intercom Connect your AI assistant to Intercom for customer messaging and support workflows. #### Prerequisites - Intercom account with API access. - Permissions to create private apps. - Access token with appropriate scopes. #### Required credentials When connecting Intercom, you'll need to provide: Create a private app in Intercom (Settings → Developers → Developer Hub → Your Apps → New App) and copy the access token from the authentication section. ![Intercom integration credentials dialog showing Access Token input field](/assets/images/intercom-integration-credentials.png) #### Available tools After connecting Intercom, your assistant can: - **Access conversation history** - Retrieve past customer interactions and messages. - **Create notes** - Add internal notes to customer conversations. - **Update customer attributes** - Modify user data and custom attributes. - **Search users** - Find customers by email, user ID, or other identifiers. - **Manage tags** - Add or remove conversation tags for organization. #### Use cases **Customer Support Context** ``` Customer: "I need help with my subscription" Assistant: [Searches Intercom for customer by phone/email] Assistant: [Reviews conversation history] Assistant: "I can see you upgraded to Pro last month. How can I help with your subscription?" ``` **User Data Management** ``` Customer: "Please update my company name" Assistant: [Updates customer attributes in Intercom] Assistant: [Adds note documenting the change] Assistant: "I've updated your company name in our system." ``` ### Jira Connect your AI assistant to Jira for project management, issue tracking, and software development workflows. #### Prerequisites - Jira account (Cloud or Server). - API token or password. - Project access and permissions. #### Required credentials When connecting Jira, you'll need to provide: The Jira username or email that has access to the project. Generate a token at https://id.atlassian.com/manage-profile/security/api-tokens by selecting **Create API token**. Provide the base URL of your Jira instance (for example `yourcompany.atlassian.net`) without the `https://` prefix. ![Jira integration credentials dialog showing Email, API token, and Site URL input fields](/assets/images/jira-integration-credentials.png) #### Available tools After connecting Jira, your assistant can: - **Create issues** - Create bugs, tasks, stories, or epics. - **Update issues** - Change status, assignee, or priority. - **Search issues** - Find issues by project, assignee, or status. - **Add comments** - Comment on existing issues. - **Transition issues** - Move issues through workflow states. #### Use cases **Bug Reporting** ``` Developer: "Users are reporting a login error" Assistant: [Creates bug in Jira] Assistant: [Sets priority to High, component to Authentication] Assistant: "Created PROJ-1234. I've assigned it to the on-call engineer." ``` **Task Management** ``` Manager: "Create a task to update the documentation" Assistant: [Creates task in Jira project] Assistant: [Sets due date based on conversation] ``` ### Salesforce Connect your AI assistant to Salesforce to access customer records, create cases, update opportunities, and more. #### Prerequisites - Salesforce account with API access. - Username and password. - Security token (reset in Personal Settings → Reset My Security Token). - Organization ID (found in Setup → Company Settings → Company Information). #### Required credentials When connecting Salesforce, you'll need to provide: Your Salesforce hostname such as `acme.my.salesforce.com` (production) or `acme.sandbox.my.salesforce.com` (sandbox) without `https://` or trailing `/`. The Salesforce username or email with integration access. The password for that Salesforce user. The security token from Personal Settings → Reset My Security Token (it arrives via email). Your org ID from Setup → Company Settings → Company Information. ![Salesforce integration credentials dialog showing Instance domain, Username, Password, Security token, and Organization ID input fields](/assets/images/salesforce-integration-credentials.png) #### Available tools After connecting Salesforce, your assistant can use tools to: - **Search records** - Find accounts, contacts, leads, opportunities. - **Create records** - Create new cases, leads, tasks, or opportunities. - **Update records** - Modify existing records with new information. - **Query data** - Run SOQL queries for custom data retrieval. #### Use cases **Customer Service** ``` Customer: "I need help with my recent order" Assistant: [Searches Salesforce for customer by phone number] Assistant: "I found your account, let me check your recent orders..." ``` **Lead Qualification** ``` Prospect: "I'm interested in your enterprise plan" Assistant: [Creates lead in Salesforce with details from conversation] Assistant: [Updates lead score based on budget and timeline discussed] ``` **Case Management** ``` Customer: "My service is down" Assistant: [Creates high-priority case in Salesforce] Assistant: "I've created case #12345 for you. Our team will reach out within 2 hours." ``` ### ServiceNow Connect your AI assistant to ServiceNow for IT service management, incident tracking, and workflow automation. #### Prerequisites - ServiceNow instance with API access. - User account with appropriate roles (e.g., itil, admin). - Instance URL and credentials. #### Required credentials Provide the ServiceNow hostname, for example `acme.service-now.com` or `acme-dev.service-now.com`. Supply the ServiceNow user with the necessary roles. Enter the password for that user. ![ServiceNow integration credentials dialog showing Instance URL, Username, and Password input fields](/assets/images/servicenow-integration-credentials.png) #### Available tools After connecting ServiceNow, your assistant can: - **Create incidents** - Log IT incidents with priority and categorization. - **Update tickets** - Modify incident status, assignment, or details. - **Search knowledge base** - Find KB articles for issue resolution. - **Query records** - Access CMDB, user records, or service catalogs. #### Use cases **IT Support** ``` Employee: "My laptop won't connect to WiFi" Assistant: [Creates incident in ServiceNow] Assistant: [Categorizes as Network → WiFi] Assistant: "I've logged incident INC0012345. IT will assist you shortly." ``` **Service Requests** ``` Employee: "I need access to the marketing drive" Assistant: [Creates service request in ServiceNow] Assistant: [Routes to appropriate approval group] ``` ### Zendesk Connect your AI assistant to Zendesk for customer service and support workflows. #### Prerequisites - Zendesk account with API access. - Admin access to generate API tokens. - Subdomain and email credentials. #### Required credentials When connecting Zendesk, you'll need to provide: Enter only the subdomain portion (e.g., `company` if your portal is `company.zendesk.com`); the rest is added automatically. The Zendesk account email that owns the API token. Generate a token in Admin Center → Apps and integrations → APIs → Zendesk API. ![Zendesk integration credentials dialog showing Subdomain, Email, and API token input fields](/assets/images/zendesk-integration-credentials.png) #### Available tools After connecting Zendesk, your assistant can: - **Create tickets** - Log customer support requests with priority and categorization. - **Search customer history** - Find previous tickets and interactions by customer. - **Update ticket status** - Modify ticket status, assignment, or priority. - **Access knowledge base** - Search KB articles for issue resolution. #### Use cases **Support Ticketing** ``` Customer: "I'm having an issue with my account login" Assistant: [Creates ticket in Zendesk] Assistant: [Categorizes as Account → Login Issues] Assistant: "I've created ticket #12345. Our support team will reach out within 2 hours." ``` **Customer History Lookup** ``` Customer: "What's the status of my previous request?" Assistant: [Searches Zendesk by phone number or email] Assistant: "I found your ticket #12340 from last week. It was resolved on Monday." ``` ### Coval Connect your AI assistant to [Coval](https://www.coval.dev/) for automated simulation, evaluation, and production monitoring of voice and chat agents. Coval is a **testing and evaluation** tool — unlike the other integrations on this page, it does not add tools your assistant uses during conversations. Instead, it tests and monitors the assistant itself, helping you catch regressions and validate behavior at scale. #### Prerequisites - Coval account at [coval.dev](https://www.coval.dev/). - A configured Telnyx AI assistant to evaluate. #### Required credentials When connecting Coval, you'll need to provide: Log in to your Coval dashboard, navigate to your workspace settings, and copy your API key. #### Capabilities After connecting Coval, you can: - **Automated scenario simulation** — Generate thousands of test scenarios from a few seed cases, covering edge cases and unexpected user behaviors. - **CI/CD regression testing** — Automatically detect performance regressions on every code change before deploying to production. - **Production monitoring** — Log live calls, receive instant alerts for performance drops, and replay transcripts and audio. - **Built-in evaluation metrics** — Measure latency, accuracy, tool-call effectiveness, and instruction compliance across all interactions. #### Use cases **Pre-deployment validation** ``` Run hundreds of simulated conversations against your assistant to verify it handles greetings, edge cases, and tool calls correctly before going live. ``` **Regression testing in CI/CD** ``` Add Coval evaluation steps to your deployment pipeline so every assistant update is automatically tested against your scenario library. ``` **Production quality monitoring** ``` Monitor live assistant conversations for performance drops, then replay specific calls to diagnose issues. ``` ## Managing integrations ### Viewing connected integrations Navigate to the assistant you want to review in the Mission Control Portal. Open the **Integrations** tab. View everything listed under **Connected Integrations**. ![Integrations section displaying Jira under Connected Integrations with description and unassign button](/assets/images/connected-integration-view.png) ### Disconnecting an integration To disconnect an integration from your assistant: Navigate to your assistant and open the **Integrations** tab. Find it within the **Connected Integrations** list. Click the chain-link **unassign** button. Approve the confirmation dialog to finish disconnecting. ![Jira integration card in Connected Integrations showing unassign button (chain link icon)](/assets/images/jira-connected-integration-unassign.png) After disconnecting: - The integration is removed from this assistant. - All associated tools are disabled for this assistant. - The integration moves to **Available Integrations** and can be reconnected later. Disconnecting an integration only removes it from the current assistant. The integration remains in your account and can be connected to other assistants or reconnected to this one. ### Deleting an integration To permanently delete an integration from your account: Navigate to your AI assistant and open the **Integrations** tab. Look under **Available Integrations**. Click the trash icon next to the integration. Approve the deletion to remove stored credentials permanently. ![Jira integration card in Available Integrations showing connect button and delete button (trash icon)](/assets/images/jira-available-integrations-actions.png) Deleting an integration permanently removes it from your account, including all stored credentials. You will need to set it up again from scratch if you want to use it in the future. ## Best practices ### Security Create integration-specific profiles with only the permissions the workflow needs. Update API tokens and passwords on a routine cadence. Review integration activity through platform audit logs. Grant only the scopes required for each integration use case. Test and validate integrations in non-production environments first. ### Configuration Enable search/read capabilities first, then gradually introduce write actions. Document when and how each tool should be used so assistants respond correctly. Validate integration behavior across multiple conversation scenarios. Configure sensible defaults (e.g., priority, project) to reduce user input errors. Define fallback behavior when integration calls fail or return unexpected results. ### Performance Avoid duplicate searches or redundant requests where possible. Store reusable values (for example, via dynamic variables) for the duration of a session. Configure timeout thresholds that balance responsiveness with reliability. Track provider limits and design workflows to stay within allocation. ## Troubleshooting ### Connection failures **Symptom**: Unable to connect integration, credentials rejected. **Solutions**: - Verify credentials are correct and have not expired. - Check that the user account has API access enabled. - Ensure security tokens or API keys are current. - For Salesforce: Confirm security token is included. - For cloud platforms: Verify instance URL format (no `https://` or trailing `/`). ### Tools not appearing **Symptom**: Integration connected but no tools available. **Solutions**: - Refresh the page and check again. - Verify the integration account has required permissions. - Check that the platform subscription includes API access. - Disconnect and reconnect the integration. ### Authentication errors during calls **Symptom**: Tools fail with authentication errors during conversations. **Solutions**: - Regenerate API tokens or security tokens. - Update stored credentials in the integration. - Verify account has not been locked or suspended. - Check IP allowlists (if applicable). ### Missing data or records **Symptom**: Assistant cannot find expected records. **Solutions**: - Verify the integration account can access the records. - Check record permissions and sharing settings. - Confirm records exist in the platform. - Verify search parameters and filters. ### Rate limiting **Symptom**: Integration calls fail with rate limit errors. **Solutions**: - Reduce frequency of integration calls. - Implement caching for frequently accessed data. - Contact platform support to increase limits. - Distribute calls across multiple service accounts. ## Next steps - **[Voice Assistant Quickstart](https://developers.telnyx.com/docs/inference/ai-assistants/no-code-voice-assistant)** - Learn how to create and configure AI assistants. - **[Workflow](/docs/inference/ai-assistants/workflows)** - Visualize how your integrations and tools connect in your assistant's conversation flow. - **[Agent Handoff](https://developers.telnyx.com/docs/inference/ai-assistants/agent-handoff)** - Enable multiple specialized assistants with integrations. - **[Dynamic Variables](https://developers.telnyx.com/docs/inference/ai-assistants/dynamic-variables)** - Pass integration-specific context to your assistant. - **[API Reference](/api-reference/assistants/list-assistants)** - Programmatic assistant management. ## Related resources - [Integration Secrets](https://portal.telnyx.com/#/integration-secrets) - Securely store API keys and tokens. - [AI Assistants Portal](https://portal.telnyx.com/#/ai/assistants) - Configure assistants and integrations. --- ### Per-Caller Credentials > Source: https://developers.telnyx.com/docs/inference/ai-assistants/per-caller-credentials.md An MCP server or webhook tool normally authenticates with one static credential, stored as an [integration secret](/docs/inference/ai-assistants/integrations) and referenced from the assistant configuration. Every conversation uses the same one. That does not work when the credential belongs to the *end user*. If each caller has their own bearer token for an MCP server, a single static credential either over-shares — one token that can reach everyone's data — or cannot be used at all. Encrypted dynamic variables solve this. The [dynamic variables webhook](/docs/inference/ai-assistants/dynamic-variables) returns a credential encrypted with a key only the account holds, and Telnyx decrypts it at the moment it authenticates that one conversation's request. --- ## How it works 1. **Store a key.** Generate a 256-bit key and save it as an integration secret. 2. **Return a ciphertext.** The dynamic variables webhook returns the caller's credential in a new `encrypted_dynamic_variables` section, encrypted with that key. 3. **Reference it.** The assistant configuration points at the variable and the key with `{{variable | encryption_secret_ref}}`. 4. **Telnyx decrypts it** only where a credential is actually needed, for that conversation only. Decrypted values exist in memory at the moment of use. They are never stored, never written to a log, and never visible to the model. --- ## 1. Store the encryption key Generate 32 random bytes and store them base64url-encoded as an integration secret. The secret's identifier is what the configuration references. ```python import base64, os print(base64.urlsafe_b64encode(os.urandom(32)).decode()) ``` Store it with `POST /integration_secrets`. The `identifier` is what the configuration references: ```json { "identifier": "mcp_enc_key", "type": "bearer", "token": "" } ``` Padding is optional — a value with or without trailing `=` is accepted. Keep the raw key: the webhook needs it to encrypt. --- ## 2. Return encrypted variables from the webhook The dynamic variables webhook response gains an optional `encrypted_dynamic_variables` section, a sibling of `dynamic_variables`: ```json { "dynamic_variables": { "customer_name": "Ada" }, "encrypted_dynamic_variables": { "mcp_token": "q1zqGJeXi0…base64url…" } } ``` | Property | Rule | | --- | --- | | Type | Object; string keys to string values. | | Keys | Letters, digits and underscores, up to 128 characters. The reserved `telnyx_` prefix is rejected. | | Values | Ciphertext per the [encryption scheme](#encryption-scheme); base64url; decoded size between 29 and 8,220 bytes. | | Maximum entries | 32. | | Invalid entries | Dropped. The rest of the response is processed normally. | The two sections are separate namespaces. An encrypted variable is never substituted into instructions, greetings, messages, or tool descriptions — only into the credential positions in [step 3](#3-reference-the-credential). A plain dynamic variable is never usable as a credential. Using the same name in both sections is not an error, but it is almost always a mistake, and Telnyx flags it as one. ### Encryption scheme AES-256-GCM, nonce-prefixed, base64url-encoded: ``` base64url( nonce(12 bytes) || AES-256-GCM ciphertext+tag ) ``` - **Nonce**: 12 random bytes, unique per encryption, prefixed to the ciphertext. - **Tag**: the standard 16-byte GCM tag, appended by the cipher. - **Plaintext**: an opaque UTF-8 string, at most 8 KB. No associated data. Because GCM is authenticated, a wrong key or a modified ciphertext is detected and treated as a failed credential rather than partial plaintext. ```python Python import os, base64 from cryptography.hazmat.primitives.ciphers.aead import AESGCM key = base64.urlsafe_b64decode(ENCRYPTION_KEY) nonce = os.urandom(12) ciphertext = AESGCM(key).encrypt(nonce, b"user-bearer-token", None) value = base64.urlsafe_b64encode(nonce + ciphertext).decode() ``` ```javascript Node const crypto = require("crypto"); const key = Buffer.from(ENCRYPTION_KEY, "base64url"); const nonce = crypto.randomBytes(12); const cipher = crypto.createCipheriv("aes-256-gcm", key, nonce); const ciphertext = Buffer.concat([ cipher.update("user-bearer-token", "utf8"), cipher.final(), cipher.getAuthTag(), ]); const value = Buffer.concat([nonce, ciphertext]).toString("base64url"); ``` There is no `openssl enc` equivalent: that command does not support GCM. ### Key rotation Update the integration secret's value. Ciphertexts produced with the old key fail authentication — and the affected requests fail safely, per [failure behavior](#failure-behavior) — until the webhook encrypts with the new one. Rotate the key and the webhook together. --- ## 3. Reference the credential Reference an encrypted variable with pipe syntax, naming the variable and the secret holding its decryption key. Whitespace around the parts and the pipe is optional. ``` {{variable_name | encryption_secret_ref}} ``` It is accepted in exactly two places. ### MCP server credential The whole `api_key_ref` value is one reference: ```json { "name": "propertyradar", "type": "http", "url": "https://mcp.example.com", "api_key_ref": "{{mcp_token | mcp_enc_key}}" } ``` Each conversation's decrypted `mcp_token` becomes that conversation's bearer token for the server. A plain identifier in `api_key_ref` keeps its existing meaning — one static integration secret for every conversation. The two forms are mutually exclusive per server. ### Tool webhook header values As a token inside a configured header value, alongside the existing `{{#integration_secret}}` and `{{plain_variable}}` syntax: ```json { "headers": [ { "name": "Authorization", "value": "Bearer {{user_token | tool_enc_key}}" } ] } ``` ### Everywhere else Anywhere else — instructions, greetings, messages, tool descriptions, webhook URLs, preset body fields, preset query parameters — the reference is not resolved. Where Telnyx can tell at save time that a reference sits in a position that cannot resolve one, the configuration is rejected with a `422`. Reading a configuration back always returns the reference verbatim. A decrypted value is never echoed on a configuration endpoint. --- ## When credentials refresh | Channel | Resolved at | Refresh | | --- | --- | --- | | Voice | Call setup, from the webhook call that starts the conversation | Fixed for the call; refreshed on assistant handoff | | SMS | Each turn, from the conversation's stored ciphertexts | Refreshed by each platform-processed SMS webhook invocation | | Chat | Each turn, from the conversation's stored ciphertexts | Reuses the most recently stored set | Each webhook response **replaces** the stored set for that conversation in full, so a token refreshed on one turn is the token used on the next. A response that omits the section leaves the previous set in place; a response with an empty section clears it. --- ## Validation Rejected at save time with a `422`: - A value that looks like a reference but does not parse. A typo is never quietly downgraded to a plain variable, because that would send the literal `{{…}}` as a credential. - An `encryption_secret_ref` that names no integration secret on the account. - A reference in a position that cannot resolve one. ## Failure behavior A credential that cannot be resolved at conversation time — the variable is missing from that conversation's set, the key secret is unavailable, or decryption fails — is never worked around: | Position | Behavior | | --- | --- | | MCP server | The server is **excluded from the conversation**. No connection is attempted, so none is ever made unauthenticated, and its tools are unavailable for that conversation. | | Tool webhook header | The tool call **fails**. The request is not sent without the header that authenticates it. | There is no fallback to another credential, and the literal `{{…}}` is never sent. The failure is recorded with the variable name, the secret reference, and what went wrong — never with key material, ciphertext, or any part of the plaintext. An assistant whose MCP server drops out of a conversation loses that server's tools for the whole conversation. If a credential is optional for some callers, configure a second assistant or a [workflow](/docs/inference/ai-assistants/workflows) branch rather than relying on partial resolution. --- ## Confidentiality - Values are stored and logged only as ciphertext. - A decrypted value exists in memory only, at the moment it authenticates a request. It is never written to a log, a conversation record, or a webhook log. - Encrypted variables cannot reach model-visible text: the templating that renders instructions, greetings, and tool descriptions does not resolve the pipe form at all. - Decrypted values are never returned by a configuration endpoint. --- ## End-to-end example 1. Store `base64url(32 random bytes)` as integration secret `mcp_enc_key`. 2. Configure the MCP server with `"api_key_ref": "{{mcp_token | mcp_enc_key}}"`. 3. Point the assistant's `dynamic_variables_webhook_url` at the endpoint. 4. On each webhook call, identify the caller from the payload, fetch that user's token, encrypt it, and return it: ```json { "encrypted_dynamic_variables": { "mcp_token": "" } } ``` Every MCP request in that conversation now carries `Authorization: Bearer `. Two concurrent conversations for two different callers use two different tokens, and neither can reach the other's data. --- ## Related resources - [Dynamic Variables](/docs/inference/ai-assistants/dynamic-variables) - The webhook this section extends, and the plain-variable syntax. - [Integrations](/docs/inference/ai-assistants/integrations) - Store the encryption key as an integration secret. - [Preset Webhook Parameters](/docs/inference/ai-assistants/preset-webhook-parameters) - Fixed tool values the model never sees. - [Voice AI Assistant API Reference](/api-reference/assistants/create-an-assistant) - Assistant, MCP server, and tool configuration. --- ### Preset Webhook Parameters > Source: https://developers.telnyx.com/docs/inference/ai-assistants/preset-webhook-parameters.md Webhook tool parameters declared under `body_parameters`, `path_parameters`, and `query_parameters` are advertised to the model: the assistant decides their values at call time, and can get them wrong or omit them. Preset parameters are the opposite. `preset_body_fields` and `preset_query_params` are supplied by the assistant configuration, never appear in the tool definition the model sees, and are attached to every call the tool makes. Use them for values the model has no business choosing — an account identifier, an API key, a tenant or channel tag, a feature flag. --- ## How preset parameters behave - **Invisible to the model.** They are not part of the tool schema, so the LLM can neither read them nor be prompted into changing them. - **Always applied.** Every invocation of the tool carries them. - **They win on conflicts.** If a preset key has the same name as a model-supplied parameter, the preset value replaces it. - **They are templated.** Values go through the same mustache templating as the rest of the tool config, so they can carry [dynamic variables](/docs/inference/ai-assistants/dynamic-variables) and [integration secrets](/docs/inference/ai-assistants/integrations). - **Query values are encoded for you.** `preset_query_params` values are percent-encoded before they are appended to the URL, so a value like `+15551234567` arrives intact. Values templated directly into the `url` are not. - **`preset_body_fields` needs a body.** `GET` requests are sent without a body, so preset body fields are dropped on `GET` tools. Use `preset_query_params` instead. --- ## Configuration | Field | Type | Description | | --- | --- | --- | | `preset_body_fields` | object | Key/value pairs merged into the request body. | | `preset_query_params` | object | Key/value pairs merged into the query string. | Both accept arbitrary keys. String values are mustache-templated before the request is sent. --- ## Examples ### Tag every request with a fixed value Here the model supplies `order_id`; `source` and `account_id` are attached by the configuration, and the assistant never sees them: ```json { "type": "webhook", "webhook": { "name": "lookup_order", "description": "Look up the status of a customer order.", "url": "https://your-backend.com/orders/status", "method": "POST", "body_parameters": { "type": "object", "properties": { "order_id": { "type": "string", "description": "The order number the customer gave." } }, "required": ["order_id"] }, "preset_body_fields": { "source": "telnyx-assistant", "account_id": "{{customer_id}}" } } } ``` Your backend receives: ```json { "order_id": "A-4471", "source": "telnyx-assistant", "account_id": "cus_8123" } ``` ### Pass the caller's number as a query parameter `preset_query_params` is the safest way to forward a phone number, because the value is percent-encoded rather than pasted into the URL: ```json { "type": "webhook", "webhook": { "name": "recent_tickets", "description": "List the caller's recent support tickets.", "url": "https://your-backend.com/tickets", "method": "GET", "preset_query_params": { "caller": "{{telnyx_end_user_target}}", "channel": "voice" } } } ``` The request becomes `GET https://your-backend.com/tickets?caller=%2B15551234567&channel=voice`. ### Send a secret the model can never read Combine preset parameters with [integration secrets](/docs/inference/ai-assistants/integrations) when an API expects credentials in the query string or body rather than in a header: ```json { "type": "webhook", "webhook": { "name": "check_inventory", "description": "Check whether an item is in stock.", "url": "https://your-backend.com/inventory", "method": "GET", "query_parameters": { "type": "object", "properties": { "sku": { "type": "string", "description": "The item SKU to check." } }, "required": ["sku"] }, "preset_query_params": { "api_key": "{{#integration_secret}}inventory-api-key{{/integration_secret}}" } } } ``` Prefer a header for credentials when the API supports one — query strings are more likely to be logged by proxies and servers along the way. Webhook tool headers support the same integration secret templating. --- ## Related resources - [Voice AI Assistant API Reference](/api-reference/assistants/create-an-assistant#webhooktool) - Complete webhook tool API documentation, including `preset_body_fields` and `preset_query_params`. - [Dynamic Variables](/docs/inference/ai-assistants/dynamic-variables) - Values resolved per conversation that you can template into preset parameters. - [Integrations](/docs/inference/ai-assistants/integrations) - Store credentials as integration secrets and reference them from tool configuration. - [Async Tools & Deferred Context](/docs/inference/ai-assistants/async-tools) - Keep the conversation moving while a webhook tool runs. --- ### Async Tools > Source: https://developers.telnyx.com/docs/inference/ai-assistants/async-tools.md Async tools allow your AI assistant to trigger long-running operations without blocking the conversation. Combined with the [Add Messages API](https://developers.telnyx.com/api-reference/call-commands/add-messages-to-ai-assistant), you can inject results back into the conversation whenever they're ready—whether that's 5 seconds or 5 minutes later. In this guide, you will learn: - How to configure webhook tools to run asynchronously - How to use the Add Messages API to inject context mid-conversation - How to combine both features for powerful async workflows --- ## Overview Traditional webhook tools block the conversation until they complete. This works fine for fast operations, but creates awkward pauses for slow backend queries. Async tools solve this by letting the assistant continue the conversation while operations run in the background. If your backend responds within a few seconds and you'd prefer to keep using sync webhooks, [filler messages](/docs/inference/ai-assistants/filler-messages) offer a simpler alternative — scripted phrases that play at timed intervals to fill silence while the webhook executes. ### The two building blocks These features are **orthogonal**—each is useful on its own, but they become especially powerful when combined. | Feature | What it does | Use alone | |---------|--------------|-----------| | **Async webhook flag** | Lets the assistant continue talking while the webhook executes | Fire-and-forget operations (logging, notifications) | | **[Add Messages API](https://developers.telnyx.com/api-reference/call-commands/add-messages-to-ai-assistant)** | Injects new context into an active conversation | External triggers, scheduled reminders, supervisor interventions | ### Combined workflow When used together, these features enable a new pattern: 1. Assistant triggers an async webhook (e.g., order lookup) 2. Assistant continues chatting with the customer 3. Backend processes the request (5, 10, 30 seconds later) 4. Backend calls Add Messages API to inject the results 5. Assistant naturally incorporates the new information This creates a seamless experience where the assistant stays engaged while slow operations complete in the background. --- ## Async webhooks The `async` flag on webhook tools tells the assistant not to wait for the response. The webhook fires, and the assistant immediately continues the conversation. ### Configuring an async webhook Set `async: true` in your webhook tool configuration: ```json { "type": "webhook", "webhook": { "name": "lookup_order_status", "description": "Triggers an async order status lookup. Results will be delivered automatically when ready.", "url": "https://your-backend.com/order-lookup", "method": "POST", "async": true, "headers": [ {"name": "Content-Type", "value": "application/json"} ], "body_parameters": { "type": "object", "properties": { "order_id": { "type": "string", "description": "The customer's order ID" } }, "required": ["order_id"] } } } ``` ![Async webhook configuration in Portal](/assets/images/async-webhook-config.png) ### Key configuration options | Field | Description | |-------|-------------| | `async` | When `true`, the assistant continues without waiting for a response | | `url` | Your backend endpoint that will process the request | | `method` | HTTP method (typically `POST`) | | `body_parameters` | JSON schema defining the parameters the assistant should provide | For the complete webhook tool schema, see the [Create Assistant API reference](https://developers.telnyx.com/api-reference/assistants/create-an-assistant). ### What your backend receives When the assistant triggers an async webhook, your endpoint receives: - The configured body parameters (e.g., `order_id`) - The `x-telnyx-call-control-id` header identifying the active call ``` POST /order-lookup HTTP/1.1 Content-Type: application/json x-telnyx-call-control-id: v3:abc123def456... { "order_id": "ORD-12345" } ``` The `x-telnyx-call-control-id` header is critical—you'll need it to inject results back into the conversation using the [Add Messages API](https://developers.telnyx.com/api-reference/call-commands/add-messages-to-ai-assistant). --- ## Add Messages API The [Add Messages API](https://developers.telnyx.com/api-reference/call-commands/add-messages-to-ai-assistant) lets you inject new messages into an active conversation from outside the call flow. This is useful for delivering async results, supervisor interventions, or external triggers. ### API endpoint ``` POST /v2/calls/{call_control_id}/actions/ai_assistant_add_messages ``` ### Request format ```bash curl -X POST "https://api.telnyx.com/v2/calls/{call_control_id}/actions/ai_assistant_add_messages" \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "messages": [ { "role": "system", "content": "Order ORD-12345 status: SHIPPED. Tracking: 1Z999AA10123456784. Estimated delivery: Tomorrow. Share this with the customer now." } ] }' ``` For the complete API specification, see the [Add Messages API reference](https://developers.telnyx.com/api-reference/call-commands/add-messages-to-ai-assistant). ### Message roles | Role | Use case | |------|----------| | `system` | Instructions or context for the assistant (recommended for async results) | | `user` | Simulate user input | | `assistant` | Inject assistant responses | ### Standalone use cases The Add Messages API is valuable even without async webhooks: - **Supervisor intervention**: A human supervisor injects guidance during a difficult call - **Scheduled reminders**: External system reminds the assistant about time-sensitive information - **Cross-system triggers**: CRM or ticketing system pushes updates to an active call - **Escalation prompts**: Monitoring system detects frustration and injects de-escalation guidance --- ## Combining async webhooks with Add Messages The real power comes from combining these features. Here's a complete example of an async order lookup system. ### Architecture ``` ┌─────────────┐ 1. Trigger async webhook ┌─────────────────┐ │ │ ──────────────────────────────────▶│ │ │ Assistant │ │ Your Backend │ │ │◀────────────────────────────────── │ │ └─────────────┘ 4. Add Messages API └─────────────────┘ │ ▲ │ │ │ │ ▼ │ ▼ 2. Continue 3. Process request Query databases, conversation (5-30 seconds) external APIs, etc. ``` ### Step 1: Configure the assistant Create an assistant with async webhook tools. Notice how the instructions tell the assistant to continue engaging while waiting: ```json { "name": "Customer Service Agent", "instructions": "You are a helpful customer service agent for Acme Electronics.\n\nWhen a customer asks about an order, trigger the lookup_order_status tool. This runs asynchronously—results will arrive automatically in 10-20 seconds.\n\nAfter triggering the lookup, keep the customer engaged:\n- Mention current promotions\n- Ask about their experience\n- Offer to help with anything else\n\nWhen results arrive, naturally incorporate them: \"Great news, I have your order info now!\"", "tools": [ { "type": "webhook", "webhook": { "name": "lookup_order_status", "description": "Async order lookup. Results delivered automatically when ready.", "url": "https://your-backend.com/order-lookup", "method": "POST", "async": true, "body_parameters": { "type": "object", "properties": { "order_id": {"type": "string", "description": "Order ID to look up"} }, "required": ["order_id"] } } } ] } ``` ### Step 2: Build the backend service Your backend receives the webhook, processes the request, and calls the [Add Messages API](https://developers.telnyx.com/api-reference/call-commands/add-messages-to-ai-assistant) when done: ```python import os import time import requests from flask import Flask, request, jsonify app = Flask(__name__) TELNYX_API_KEY = os.environ.get("TELNYX_API_KEY") @app.route("/order-lookup", methods=["POST"]) def order_lookup(): data = request.get_json(silent=True) or {} # Get the call control ID from headers call_control_id = request.headers.get("x-telnyx-call-control-id") order_id = data.get("order_id") if not call_control_id: return jsonify({"error": "Missing call control ID"}), 400 # Simulate slow backend query (replace with real logic) time.sleep(15) # Build the result message result = { "status": "SHIPPED", "tracking": "1Z999AA10123456784", "delivery": "Tomorrow" } system_message = f"""[ORDER LOOKUP COMPLETE] Order {order_id}: {result['status']} Tracking: {result['tracking']} Estimated delivery: {result['delivery']} Share these details with the customer now.""" # Inject results back into the conversation inject_message(call_control_id, system_message) return jsonify({"status": "sent"}) def inject_message(call_control_id: str, message: str): """Send a message to an active conversation via Add Messages API.""" url = f"https://api.telnyx.com/v2/calls/{call_control_id}/actions/ai_assistant_add_messages" response = requests.post( url, headers={ "Authorization": f"Bearer {TELNYX_API_KEY}", "Content-Type": "application/json" }, json={ "messages": [{"role": "system", "content": message}] } ) return response.json() if __name__ == "__main__": app.run(host="0.0.0.0", port=8000) ``` ### Step 3: Test the flow 1. Call your assistant and ask about an order 2. The assistant triggers the async lookup and continues chatting 3. After 15 seconds, your backend injects the results 4. The assistant seamlessly shares the order details ![Conversation transcript showing async message injection](/assets/images/async-transcript-injection.png) --- ## Multiple parallel lookups You can trigger multiple async webhooks simultaneously. Each completes independently and injects results as they become available. ### Example: Staggered results Configure multiple tools with different backend processing times: | Tool | Processing time | Information returned | |------|----------------|---------------------| | `check_loyalty_points` | ~10 seconds | Points balance, membership tier | | `lookup_order_status` | ~20 seconds | Order status, tracking, delivery estimate | The assistant triggers both at once. Results drip into the conversation naturally: ``` Customer: "Where's my order 12345?" Agent: [Triggers both lookups] "Let me pull that up for you! By the way, we're running 20% off all accessories this week." [10 seconds pass - loyalty results arrive] Agent: "Oh nice, I see you have 2,500 reward points - that's Gold status! You've got $25 to use on your next purchase." [20 seconds pass - order results arrive] Agent: "And here's your order info - order 12345 is out for delivery! Tracking number is 1Z999AA10123456784, should arrive tomorrow." ``` ### Instructing the assistant For parallel lookups to work well, your assistant instructions should emphasize calling tools together: ``` When a customer asks about an order, trigger BOTH lookup tools at the same time: 1. check_loyalty_points 2. lookup_order_status Do not wait for one to complete before calling another. Call both immediately. Results will arrive automatically as each lookup completes. ``` --- ## Best practices ### Crafting system messages When injecting results via the [Add Messages API](https://developers.telnyx.com/api-reference/call-commands/add-messages-to-ai-assistant), format them clearly: ```python # Good - Clear, actionable system_message = """[ORDER LOOKUP COMPLETE] Order ORD-12345: SHIPPED Tracking: 1Z999AA10123456784 Estimated delivery: Tomorrow Share these details with the customer now.""" # Avoid - Ambiguous system_message = "The order was found in the system." ``` ### Handling edge cases **Call ended before results arrive:** ```python response = inject_message(call_control_id, message) if response.get("status_code") == 404: # Call has ended, log and move on print(f"Call {call_control_id} already ended") ``` **Multiple results for same lookup:** - Include identifiers in messages so the assistant knows which query the results belong to - Use timestamps or request IDs if needed ### Backend considerations - Your backend should return a 200 response quickly to acknowledge receipt - Process the actual work asynchronously (use background workers, Celery, AWS Lambda, etc.) - There's no timeout constraint on async webhooks—your backend can take as long as needed before calling the Add Messages API ### Testing tips - Use tools like ngrok to expose local backends during development - Log all headers to verify `x-telnyx-call-control-id` is received - Test with various delay lengths to ensure natural conversation flow - Monitor the conversation transcript in the Portal to see messages being injected --- ## Use cases ### Customer service - **Order lookups**: Query multiple systems (warehouse, shipping, payments) in parallel - **Account reviews**: Pull account history, loyalty status, and recent tickets simultaneously - **Product availability**: Check inventory across multiple warehouses ### Healthcare - **Patient record retrieval**: Fetch records from multiple systems while confirming appointment details - **Insurance verification**: Run eligibility checks while gathering patient information - **Lab results**: Query lab systems and deliver results when ready ### Financial services - **Loan pre-qualification**: Run credit checks and affordability calculations in background - **Account aggregation**: Pull balances from multiple accounts simultaneously - **Fraud alerts**: Inject real-time fraud warnings from monitoring systems ### Scheduling - **Multi-calendar availability**: Check availability across multiple calendars/resources - **Booking confirmations**: Process reservations and inject confirmation details - **Waitlist updates**: Notify assistant when spots become available --- ## Related resources - **[Add Messages API Reference](https://developers.telnyx.com/api-reference/call-commands/add-messages-to-ai-assistant)** - Complete API specification for injecting messages - **[Create Assistant API Reference](https://developers.telnyx.com/api-reference/assistants/create-an-assistant)** - Full webhook tool configuration options - **[Filler Messages](/docs/inference/ai-assistants/filler-messages)** - Scripted phrases that fill silence during sync webhook calls - **[Webhooks & Workflows](/docs/inference/ai-assistants/workflows)** - Learn more about configuring webhook tools - **[Dynamic Variables](/docs/inference/ai-assistants/dynamic-variables)** - Pass context into conversations at start time - **[Memory](/docs/inference/ai-assistants/memory)** - Persist information across conversations --- ### Filler Messages > Source: https://developers.telnyx.com/docs/inference/ai-assistants/filler-messages.md Filler messages let your AI assistant speak scripted phrases while synchronous webhook tools execute. Instead of silence during a database lookup or API call, callers hear natural phrases like "Let me look that up for you" at configurable intervals. In this guide, you will learn: - How filler messages work and when to use them - How to configure message types and timing - How to set them up via the API and Mission Control Portal ## Video tutorial --- ## Overview When an AI assistant calls a synchronous webhook tool, it waits for the response before speaking again. If your backend takes several seconds to respond, the caller hears silence — creating an awkward experience. Filler messages solve this by playing scripted phrases at the start of the request and at timed intervals while the assistant waits. Messages are **scripted per tool, not generated by the LLM** — you control the exact wording, timing, and brand voice. Filler messages work with **synchronous webhook tools** and **MCP tool calls** only. Async webhooks return immediately and don't produce dead air. If your backend consistently takes more than a few seconds, consider [async tools](/docs/inference/ai-assistants/async-tools) as an alternative. --- ## Filler message types There are two types of filler messages, each triggered at a different point during the webhook call. | Type | When it plays | `timing_ms` required | |------|---------------|----------------------| | `request_start` | Immediately when the webhook call begins | No | | `request_response_delayed` | After a specified delay if no response has arrived | Yes (100–120,000 ms) | You can configure multiple `request_response_delayed` messages at different delay thresholds to keep the caller engaged during longer operations. --- ## Configuration ### API schema Add a `messages` array to your webhook tool configuration: ```json { "type": "webhook", "webhook": { "name": "check_billing", "description": "Look up the customer's current billing information.", "url": "https://your-backend.com/billing", "method": "POST", "messages": [ { "type": "request_start", "content": "Let me look that up for you." }, { "type": "request_response_delayed", "content": "Still working on this.", "timing_ms": 5000 }, { "type": "request_response_delayed", "content": "Almost there, just a moment.", "timing_ms": 15000 } ], "body_parameters": { "type": "object", "properties": { "account_id": { "type": "string", "description": "The customer's account ID" } }, "required": ["account_id"] } } } ``` ### Key fields | Field | Description | |-------|-------------| | `messages` | Array of filler message objects on a webhook tool | | `type` | `request_start` or `request_response_delayed` | | `content` | The text the assistant speaks to the caller | | `timing_ms` | Delay in milliseconds before speaking (100–120,000). Required for `request_response_delayed` messages only | ### Mission Control Portal You can also configure filler messages through the Portal: 1. Navigate to **AI > Assistants** 2. Select your assistant and edit a webhook tool 3. Set the tool to **Sync** mode 4. Open the **Filler Messages** tab 5. Add your messages and timing thresholds ![Filler Messages configuration in Mission Control Portal](/assets/images/ai-assistant-filler-messages.png) --- ## Example: tiered filler messages This example shows a billing lookup tool with three filler messages that escalate in reassurance as the wait grows: ```json { "messages": [ { "type": "request_start", "content": "Let me pull up your billing details." }, { "type": "request_response_delayed", "content": "I'm still looking into that for you.", "timing_ms": 5000 }, { "type": "request_response_delayed", "content": "Just a bit longer, I want to make sure I have the latest information.", "timing_ms": 10000 } ] } ``` **What the caller hears:** ``` Caller: "What's my current balance?" Agent: "Let me pull up your billing details." [5 seconds pass] "I'm still looking into that for you." [5 more seconds pass] "Just a bit longer, I want to make sure I have the latest information." [Backend responds] "Your current balance is $142.50, due on August 1st." ``` --- ## Use cases - **Customer support**: Query billing or CRM systems mid-call without silence - **AI receptionists**: Book appointments via webhook while keeping the caller engaged - **Voice agents calling external LLMs**: Fill silence while a third-party LLM or RAG pipeline processes a request - **Appointment scheduling**: Check availability across calendar systems with natural wait phrases --- ## Filler messages vs. async tools Both features address the same problem — slow backend responses — but take different approaches. | | Filler messages | [Async tools](/docs/inference/ai-assistants/async-tools) | |---|---|---| | **Approach** | Keep the sync webhook, fill silence with scripted phrases | Make the webhook async, continue the conversation freely | | **Best for** | Backends that respond in under ~30 seconds | Backends that take 30+ seconds or have unpredictable latency | | **Caller experience** | Hears reassuring phrases at timed intervals | Assistant continues a natural conversation | | **Response handling** | Assistant receives the webhook response directly | Backend injects results via the Add Messages API | | **Configuration** | Add `messages` to the webhook configuration | Set `async: true` and build an Add Messages callback | Use filler messages when your backend is reasonably fast and you want simple, predictable caller reassurance. Use async tools when backends are slow or you want the assistant to have a free-flowing conversation while waiting. --- ## Related resources - **[Create Assistant API Reference](https://developers.telnyx.com/api-reference/assistants/create-an-assistant)** - Full webhook tool configuration options including `messages` - **[Async Tools & Deferred Context](/docs/inference/ai-assistants/async-tools)** - Alternative approach for long-running backend operations - **[Webhooks & Workflows](/docs/inference/ai-assistants/workflows)** - General webhook tool configuration --- ### Client-Side Tools > Source: https://developers.telnyx.com/docs/inference/ai-assistants/client-side-tools.md Client-side tools let your AI assistant invoke functions that run directly in the browser or client application during a voice or chat conversation. Unlike [webhook tools](/docs/inference/ai-assistants/async-tools) — which make HTTP requests to your backend — client-side tools execute JavaScript code on the client, with direct access to browser state, local data, and any API the page is authenticated to call. ## How they work 1. **You define a tool** in the [Portal](https://portal.telnyx.com/#/ai/assistants) with a name, description, and JSON Schema for parameters — the same way you configure other tool types. 2. **The assistant decides when to call it** based on the conversation context and the tool's description. 3. **Your client-side handler runs** in the browser, receiving the parsed arguments from the assistant. 4. **The return value is sent back** to the assistant so it can continue the conversation naturally. This flow uses the Voice SDK Proxy WebSocket to relay tool invocations and results between the assistant and your client code — no HTTP webhook required. Client-side tools are available for WebRTC-based conversations (voice and chat) using the `@telnyx/ai-agent-lib` JavaScript/React library. They are not available for SIP/phone-call-based conversations. ## When to use client-side tools | Scenario | Client-side tool | Webhook tool | |----------|:---:|:---:| | Read data the page already has (shopping cart, user session) | ✅ | | | Trigger a UI action (open a modal, navigate, play media) | ✅ | | | Call an API authenticated with the user's browser session | ✅ | | | Query a backend database or internal service | | ✅ | | Send a notification to an external system | | ✅ | | Needs to work for phone/SIP calls | | ✅ | ## Configure a client-side tool in the Portal 1. Open **AI Assistants** in the [Telnyx Portal](https://portal.telnyx.com/#/ai/assistants). 2. Select an assistant and scroll to the **Tools** section. 3. Click **Add Tool** and select **Client-Side Tool**. 4. Fill in the fields: | Field | Description | |-------|-------------| | **Name** | The tool name the assistant uses to call it. Must match the name you register in your client code. Only letters, numbers, underscores, and hyphens. | | **Description** | Tell the assistant what the tool does and when to use it. The clearer the description, the better the assistant is at deciding when to invoke it. | | **Parameters** | A JSON Schema object defining the tool's input parameters. Use the visual editor for simple schemas (strings, numbers, booleans, enums, arrays) or switch to **Advanced mode** for complex schemas with nested objects or custom keywords. | | **Timeout (ms)** | How long the client has to execute the tool before a timeout error is returned to the assistant. Default: 5000 ms. This per-tool value is sent to the client and applied automatically — you do not need to configure it in code. | 5. Click **Create & Add to Assistant**. You can also create client-side tools in the [Tools Library](/docs/inference/ai-assistants/tools-library) and share them across multiple assistants. The **Parameters** schema uses the same JSON Schema format as webhook tool body parameters. The visual editor supports `string`, `number`, `integer`, `boolean`, `enum`, and array types. If your schema uses keywords the visual editor doesn't support (like `minimum`, `pattern`, or `default`), it will be displayed in **Advanced mode** as raw JSON to preserve those keywords. ## Implement client-side tool handlers Install the Telnyx AI Agent library: ```bash npm install @telnyx/ai-agent-lib ``` Client-side tools require `@telnyx/ai-agent-lib` version **0.5.0** or later. ### Register tools at construction time ```typescript import { TelnyxAIAgent } from '@telnyx/ai-agent-lib'; const agent = new TelnyxAIAgent({ agentId: 'your-agent-id', clientTools: { // The key must match the tool name configured in the Portal lookup_order: async (args) => { // args is typed as `unknown` — cast to your expected shape const { orderId } = args as { orderId: string }; const response = await fetch(`/api/orders/${orderId}`); const order = await response.json(); return { status: 'found', orderId: order.id, total: order.total }; }, get_cart: () => { // No parameters needed — return data the page already has return { itemCount: cart.items.length, items: cart.items }; }, open_chat_panel: () => { // Trigger a UI action setChatOpen(true); return { success: true }; }, }, // Optional client-wide fallback timeout, used only for tools that have no // Timeout configured in the Portal. Each tool's Portal Timeout takes // precedence and is applied automatically. Default: 30000 ms. clientToolTimeoutMs: 15000, }); ``` ### Register tools at runtime ```typescript // Add a tool after construction agent.registerClientTool('get_user_profile', () => { return { name: currentUser.name, email: currentUser.email }; }); // Remove a tool agent.unregisterClientTool('get_user_profile'); // List registered tools agent.getClientTools(); // ['lookup_order', 'get_cart'] ``` ### Use with React ```tsx import { useClient } from '@telnyx/ai-agent-lib'; function ToolRegistration() { const client = useClient(); useEffect(() => { client.registerClientTool('lookup_order', async (args) => { const { orderId } = args as { orderId: string }; const response = await fetch(`/api/orders/${orderId}`); const order = await response.json(); return { status: 'found', orderId: order.id, total: order.total }; }); return () => client.unregisterClientTool('lookup_order'); }, [client]); return null; } ``` ### Handler contract ```typescript type ClientSideToolHandler = ( args: unknown, // parsed JSON arguments (or undefined if empty) context: ClientSideToolContext, // { callId, toolName, rawArguments } ) => unknown | Promise; ``` - The return value is serialized and sent back as the tool output. Strings are sent verbatim; anything else is JSON-stringified. - Handlers may be async. Each runs with its tool's **Timeout** configured in the Portal, falling back to `clientToolTimeoutMs` for tools that have none. If the handler exceeds it, a `{ "error": "timeout" }` output is returned to the assistant. - The library handles the `call_id` round-trip — you don't need to manage that yourself. ## Error handling The library always returns a result to the assistant so the conversation never hangs: | Scenario | What happens | |----------|-------------| | Unknown tool name | Safe error output `{ "error": "unknown_tool" }` | | Invalid JSON arguments | Safe error output `{ "error": "invalid_arguments" }` | | Handler throws or rejects | Safe error output `{ "error": "handler_error" }` | | Handler exceeds timeout | Safe error output `{ "error": "timeout" }` | | Disconnected before output sent | Output dropped, error event emitted | | Duplicate `call_id` (re-delivery) | Ignored — tool is not executed twice | Your handler should handle its own errors gracefully. If you want the assistant to know about a failure (e.g., an API returned 404), return an error object rather than throwing: ```typescript lookup_order: async (args) => { const { orderId } = args as { orderId: string }; const response = await fetch(`/api/orders/${orderId}`); if (!response.ok) { return { error: 'order_not_found', orderId }; } const order = await response.json(); return { status: 'found', orderId: order.id, total: order.total }; }, ``` ## Observe tool lifecycle events ```typescript agent.on('client.tool.invoked', ({ callId, toolName }) => { console.log(`Tool ${toolName} invoked (${callId})`); }); agent.on('client.tool.completed', ({ callId, toolName, isError }) => { console.log(`Tool ${toolName} completed, error=${isError}`); }); agent.on('client.tool.error', ({ callId, toolName, reason }) => { console.warn(`Tool ${toolName} failed: ${reason}`); }); ``` Tool arguments and outputs are **never logged** — they may contain customer data. Only safe correlation fields (`callId`, tool name) appear in debug logs. ## Client-side tools vs. webhook tools | Feature | Client-side tools | Webhook tools | |---------|:-:|:-:| | Execution location | Browser/client | Your backend server | | Requires public endpoint | No | Yes | | Works with phone/SIP calls | No | Yes | | Works with WebRTC (voice/chat) | Yes | Yes | | Can access browser state | Yes | No | | Latency | Minimal (no HTTP round-trip) | Network-dependent | | Can be tested in the Portal | No | Yes (test button) | ## Use cases - **E-commerce**: Look up cart contents, check order status from the browser session, apply a promo code - **Scheduling**: Read the user's calendar from the page, open a booking modal - **Customer support**: Fetch user details from the client session, trigger a screen share - **IoT dashboards**: Read device state from the dashboard, toggle a device on/off ## Learn more - **[Tools Library](/docs/inference/ai-assistants/tools-library)** — Create shared tools and reuse them across assistants - **[Async Tools & Deferred Context](/docs/inference/ai-assistants/async-tools)** — Webhook tools that don't block the conversation - **[Voice Assistant Quickstart](/docs/inference/ai-assistants/no-code-voice-assistant)** — Get started with AI assistants in the Portal - **[@telnyx/ai-agent-lib on npm](https://www.npmjs.com/package/@telnyx/ai-agent-lib)** — Full API reference for the JavaScript/React library --- ### AI Assistants and Edge Compute > Source: https://developers.telnyx.com/docs/edge-compute/guides/ai-assistant-backend.md Telnyx AI Assistants can call out to your own backend in several scenarios — resolving dynamic variables at the start of a conversation, executing webhook tool calls mid-conversation, and more. Whenever you need a backend for these callbacks, Telnyx Edge Compute is a natural fit: no server to manage, secrets injected at runtime, and deployment via a single CLI command. This guide walks through building a single Go function that handles both dynamic variables and webhook tool calls, using the demo app `telnyx-ai-edge` as the reference implementation. --- ## What you'll build A support assistant for "Telnyx Logistics" that: - Greets callers by name (dynamic variables resolved from the caller's phone number) - Has a `lookup-order` tool the assistant can call to retrieve order status, carrier, and estimated delivery Both the dynamic variable lookup and the tool call hit one Edge Compute function at a single URL. --- ## Prerequisites - A Telnyx account with [Edge Compute](/docs/edge-compute/quickstart) enabled. - The `telnyx-edge` CLI [installed and authenticated](/docs/edge-compute/quickstart). - An existing [AI Assistant](https://portal.telnyx.com/#/ai/assistants) (or you can create one via API as shown below). - Go 1.24+ installed locally (if following along with the Go sample). --- ## Key concepts ### Single function, two callbacks Edge Compute routes all HTTP methods and paths under your function URL to your handler — path handling is up to your code (see [Routes & Domains](/docs/edge-compute/configuration/routing)). The platform handles `/health/liveness` and `/health/readiness` probes automatically. In this guide, both the dynamic variables webhook and the webhook tool call point to the same function URL, so the handler dispatches on the **request body shape** rather than the URL path: - **Dynamic variables webhook** — Telnyx wraps the payload under `data.event_type`. - **Webhook tool call** — the body is the flat arguments object from the tool's `body_parameters` schema (e.g. `{"order_id": "ORD-10042"}`). You could also use separate paths (e.g. `/dynamic-variables` and `/tool/lookup-order`) if you prefer path-based routing — both approaches work. This guide uses body-shape dispatch to keep everything at a single URL. ### Webhook signature verification Telnyx signs every dynamic-variables webhook and webhook tool call with an Ed25519 key. The signature is in the `telnyx-signature-ed25519` header, and the timestamp is in `telnyx-timestamp`. The signed message is `"{timestamp}|{raw_body}"`. You must verify this signature to confirm the request is genuinely from Telnyx. Your org's public key is available at: ``` GET https://api.telnyx.com/v2/public_key Authorization: Bearer ``` The response contains `data.public` (not `data.public_key`) — the base64-encoded Ed25519 public key. ### Dynamic variables response format The response **must** nest variables under a `dynamic_variables` key. A flat object (e.g. `{"customer_name": "James"}`) is silently ignored — variables will remain unresolved. ```json { "dynamic_variables": { "customer_name": "James Smith", "account_tier": "premium" } } ``` ### Timeout The default dynamic variables webhook timeout is 1,500 ms. Edge Compute functions may occasionally need more time on a cold start, so consider setting `dynamic_variables_webhook_timeout_ms` on the assistant to a higher value (up to 10,000 ms). A value of 8,000 ms is a reasonable choice for edge backends. --- ## Step 1: Scaffold the function ```bash telnyx-edge new-func -l go -n telnyx-ai-edge cd telnyx-ai-edge ``` This creates a `func.toml` with the registered function ID and a Go handler scaffold. The Go module **must** be named `function` (package `function`, entrypoint `Handle(w, r)`). Other module names fail to build: "malformed module path: missing dot in first path element." Use `go 1.24` in `go.mod`. --- ## Step 2: Store the public key as a secret Fetch your org's public key and store it as an encrypted secret. The public key endpoint requires authentication — use your Telnyx API key: ```bash # Get the public key (requires authentication) PUBLIC_KEY=$(curl -s -H "Authorization: Bearer $TELNYX_API_KEY" \ https://api.telnyx.com/v2/public_key | jq -r '.data.public') # Store it as a secret (encrypted, org-scoped, injected as env var at runtime) telnyx-edge secrets add TELNYX_PUBLIC_KEY "$PUBLIC_KEY" ``` The function reads this secret from `os.Getenv("TELNYX_PUBLIC_KEY")` at startup. Secrets are never visible in `secrets list` — only the name is shown. --- ## Step 3: Write the handler The handler does three things: 1. Verifies the Telnyx Ed25519 signature on every request 2. Detects whether the request is a dynamic-variables webhook or a tool call 3. Returns the appropriate response ```go handler.go package function import ( "crypto/ed25519" "encoding/base64" "encoding/json" "io" "log" "net/http" "os" "strconv" "time" ) const maxSkew = 5 * time.Minute var publicKey ed25519.PublicKey func init() { raw := os.Getenv("TELNYX_PUBLIC_KEY") if raw == "" { log.Println("warning: TELNYX_PUBLIC_KEY is not set; all requests will be rejected") return } key, err := base64.StdEncoding.DecodeString(raw) if err != nil || len(key) != ed25519.PublicKeySize { log.Printf("warning: TELNYX_PUBLIC_KEY is invalid (len=%d, err=%v)", len(key), err) return } publicKey = ed25519.PublicKey(key) } func Handle(w http.ResponseWriter, r *http.Request) { // Health probes are handled by the platform — don't add custom health routes if r.Method != http.MethodPost { http.Error(w, "method not allowed", http.StatusMethodNotAllowed) return } body, err := io.ReadAll(r.Body) if err != nil { http.Error(w, "cannot read body", http.StatusBadRequest) return } if !verifyTelnyxSignature(r.Header, body) { http.Error(w, "invalid signature", http.StatusForbidden) return } // Dispatch on body shape: DV webhook has "data.event_type", // tool call is a flat args object if isDynamicVariablesRequest(body) { handleDynamicVariables(w, body) return } handleLookupOrder(w, body) } func isDynamicVariablesRequest(body []byte) bool { var probe struct { Data *struct { EventType string `json:"event_type"` } `json:"data"` } if err := json.Unmarshal(body, &probe); err != nil { return false } return probe.Data != nil } func verifyTelnyxSignature(h http.Header, body []byte) bool { if publicKey == nil { return false } sig := h.Get("telnyx-signature-ed25519") ts := h.Get("telnyx-timestamp") if sig == "" || ts == "" { return false } t, err := strconv.ParseInt(ts, 10, 64) if err != nil { return false } age := time.Since(time.Unix(t, 0)) if age < -maxSkew || age > maxSkew { return false } s, err := base64.StdEncoding.DecodeString(sig) if err != nil { return false } signed := append([]byte(ts+"|"), body...) return ed25519.Verify(publicKey, signed, s) } // --- Dynamic Variables --- type dvRequest struct { Data struct { EventType string `json:"event_type"` Payload struct { Channel string `json:"telnyx_conversation_channel"` AgentTarget string `json:"telnyx_agent_target"` EndUserTarget string `json:"telnyx_end_user_target"` CallControlID string `json:"call_control_id"` AssistantID string `json:"assistant_id"` } `json:"payload"` } `json:"data"` } type dvResponse struct { DynamicVariables map[string]any `json:"dynamic_variables"` } func handleDynamicVariables(w http.ResponseWriter, body []byte) { var req dvRequest if err := json.Unmarshal(body, &req); err != nil { http.Error(w, "bad json", http.StatusBadRequest) return } caller := req.Data.Payload.EndUserTarget resp := dvResponse{ DynamicVariables: map[string]any{ "customer_name": lookupCustomerName(caller), "account_tier": "premium", "open_order_id": "ORD-10042", "support_region": "US", }, } writeJSON(w, resp) } func lookupCustomerName(caller string) string { known := map[string]string{ "+13128675309": "James Smith", "+15551234567": "Rachel Thomas", } if name, ok := known[caller]; ok { return name } return "there" } // --- Webhook Tool: lookup-order --- type toolRequest struct { OrderID string `json:"order_id"` } type toolResponse struct { OrderID string `json:"order_id"` Status string `json:"status"` EstimatedDeliv string `json:"estimated_delivery"` Carrier string `json:"carrier"` } func handleLookupOrder(w http.ResponseWriter, body []byte) { var req toolRequest if err := json.Unmarshal(body, &req); err != nil { http.Error(w, "bad json", http.StatusBadRequest) return } resp := toolResponse{ OrderID: req.OrderID, Status: "shipped", EstimatedDeliv: "2025-04-10", Carrier: "Telnyx Logistics", } writeJSON(w, resp) } func writeJSON(w http.ResponseWriter, v any) { w.Header().Set("Content-Type", "application/json") json.NewEncoder(w).Encode(v) } ``` --- ## Step 4: Ship the function ```bash telnyx-edge ship ``` The ship process takes 2–3 minutes. After `uploaded successfully`, poll `telnyx-edge list` until the status shows `deploy_ok`. Don't trust a CLI timeout as a failure — the function may still be building server-side. ```bash telnyx-edge list # FUNC ID FUNCTION NAME STATUS INVOKE URL # 917210e7-... telnyx-ai-edge deploy_ok https://telnyx-ai-edge-.telnyxcompute.com ``` Save the invoke URL — you'll point the assistant at it next. --- ## Step 5: Configure the AI Assistant Set the function URL as both the dynamic variables webhook URL and the webhook tool URL on the assistant. ### Dynamic variables webhook In the [Portal](https://portal.telnyx.com/#/ai/assistants) or via the API: | Field | Value | |-------|-------| | `dynamic_variables_webhook_url` | `https://telnyx-ai-edge-.telnyxcompute.com/` | | `dynamic_variables_webhook_timeout_ms` | `8000` | Consider setting the timeout to 8,000 ms to give the function room on cold starts. The default 1,500 ms may be tight for a cold function. ### Template variables in the assistant Use `{{variable_name}}` in the assistant's instructions and greeting to reference the variables your function returns: ``` instructions: "You are a support agent for Telnyx Logistics. The caller is {{customer_name}} (tier: {{account_tier}}). They may have open order {{open_order_id}}." greeting: "Hi {{customer_name}}, thanks for calling Telnyx Logistics. How can I help you today?" ``` ### Webhook tool Add a webhook tool that points to the same function URL: ```json { "type": "webhook", "webhook": { "name": "lookup-order", "description": "Look up the current status of a customer order by its order id.", "url": "https://telnyx-ai-edge-.telnyxcompute.com/", "method": "POST", "body_parameters": { "type": "object", "properties": { "order_id": { "type": "string", "description": "The order id to look up, e.g. ORD-10042." } }, "required": ["order_id"] } } } ``` When the LLM decides to call `lookup-order`, Telnyx sends a POST with the tool arguments as the flat body (`{"order_id": "ORD-10042"}`), signed with the same Ed25519 key. Your function detects the body shape, handles it as a tool call, and returns the result. --- ## Step 6: Test end-to-end 1. **Call the function directly** (without a signature — it'll return 403, confirming it's live): ```bash curl -X POST https://telnyx-ai-edge-.telnyxcompute.com/ \ -H "Content-Type: application/json" \ -d '{"order_id":"ORD-10042"}' # → 403 invalid signature ← expected, signature verification is working ``` 2. **Make a test call** to the assistant from the Portal or via the API: ```bash curl --request POST \ --url https://api.telnyx.com/v2/texml/ai_calls/ \ --header "Authorization: Bearer $TELNYX_API_KEY" \ --header 'Content-Type: application/json' \ --data '{ "From": "+13128675309", "To": "+15551234567", "AIAssistantId": "assistant-" }' ``` 3. **Verify in the conversation transcript** that: - The greeting includes the resolved `customer_name` - The assistant can call `lookup-order` and read back real order data --- ## Tips and gotchas ### Choosing body-shape vs path-based dispatch Since Edge Compute routes all paths to your handler, you can use path-based routing (e.g. `r.URL.Path == "/tool/lookup-order"`) or body-shape dispatch as shown in this guide. Both work. If you configure separate URLs for the DV webhook and the tool on the assistant, path-based routing is natural. If you point both at the same URL, body-shape dispatch is the way to go. For a path-based routing example, see the [RESTful API example](https://github.com/team-telnyx/edge-compute-cli/tree/main/docs/examples/python/restful-api) in the Edge Compute CLI repo. ### Consider a higher webhook timeout The default dynamic variables webhook timeout is 1,500 ms. Edge Compute functions may occasionally need a bit more time on a cold start, so consider setting `dynamic_variables_webhook_timeout_ms` to 8,000 ms to give the function room. The maximum is 10,000 ms. ### Always verify signatures Without signature verification, anyone who knows your function URL can inject fake dynamic variables or tool responses. The `telnyx-signature-ed25519` and `telnyx-timestamp` headers are present on every request from Telnyx. ### Ship takes a few minutes A normal ship takes 2–3 minutes. The CLI's build monitor has a 5-minute timeout, but the build continues server-side regardless. If the CLI reports a timeout, check `telnyx-edge list` for the actual status before retrying — the function may have deployed successfully. ### Secrets require re-shipping Adding or changing a secret (`telnyx-edge secrets add`) does not affect an already-deployed function. Run `telnyx-edge ship` again to pick up the new secret. ### The `dynamic_variables` wrapper is mandatory Returning a flat JSON object like `{"customer_name": "James"}` will be silently ignored. Variables must be nested under `dynamic_variables`: ```json { "dynamic_variables": { "customer_name": "James" } } ``` --- ## Next steps - [Dynamic Variables](/docs/inference/ai-assistants/dynamic-variables) — full reference for the DV webhook payload and resolution precedence. - [Webhook signing](/docs/development/api-fundamentals/webhooks/receiving-webhooks#webhook-signing) — how Telnyx signs webhooks and how to verify signatures. - [Edge Compute quickstart](/docs/edge-compute/quickstart) — getting started with your first function. - [Secrets](/docs/edge-compute/configuration/secrets) — encrypted, org-scoped environment variables. - [Bindings](/docs/edge-compute/runtime/bindings) — pre-authenticated Telnyx API client for your function. --- ### Conversation Workflows > Source: https://developers.telnyx.com/docs/inference/ai-assistants/workflows.md Conversation workflows let you turn a single AI Assistant into a guided, multi-step experience. Instead of asking one prompt to handle every part of a call or chat, you can design a graph of conversation steps, define when the assistant should move between them, and tune the assistant's behavior at each step. Use workflows when your assistant needs to guide people through a process, such as intake, qualification, booking, verification, escalation, or support triage. ## Why use workflows? A single prompt works well for open-ended conversations. Workflows are better when the customer journey has structure. With workflows, you can: - **Break complex conversations into focused steps**: Give each stage its own name and instructions, such as `Intake`, `Billing`, `Schedule appointment`, or `Escalate to specialist`. - **Route with natural language or deterministic rules**: Move between steps when the LLM decides a condition is met, or when a dynamic variable matches a configured value. - **Tune behavior per step**: Append or replace the assistant's main instructions, and optionally override the model or voice for a specific node. - **Route to another assistant**: Route from one workflow to another assistant when the conversation should be handled by a different specialist. - **Debug what happened later**: Conversation transcripts can show the workflow node associated with assistant messages, so you can connect real conversations back to the workflow design. ## How workflows work A workflow is a directed graph stored on the AI Assistant as `conversation_flow`. The graph has two main building blocks: - **Nodes**: Conversation steps. A node is a **prompt node** (an LLM-driven step with its own label, instructions, instruction mode, model override, and voice override), a **speak node** (a deterministic step that plays a fixed scripted message with no LLM turn), or a **tool node** (a deterministic step that runs a single shared tool with no LLM turn). - **Edges**: Transitions between nodes, or from a node to another assistant. Each edge has a condition that decides when that path should be taken. Conditions can be natural-language (**LLM**), deterministic (**variable comparison**), or a **default** fallback. When a conversation starts, Telnyx begins at the workflow's start node. How routing is decided depends on the edge's condition type. Variable-comparison edges are evaluated by Telnyx against your dynamic variables. LLM-condition edges are decided by the assistant's model. On calls, this happens in the same turn that produces the reply — the edge's prompt is offered to the model as a tool choice, and the transition fires when the model selects it. On chat channels, the edge prompts are evaluated in a separate model call after the reply (see [LLM conditions](#llm-conditions)). If no edge fires, the conversation stays on the current node — except for speak and tool nodes, which advance automatically once their step completes. Workflows are optional. Assistants without a workflow continue to use their standard assistant-level instructions, model, voice, tools, and settings. ## What you can build with workflows Conversation workflows support the building blocks you need for guided, adaptive customer journeys: - **Start from a defined entry point**: Choose which node begins the workflow. - **Create focused conversation stages**: Add prompt nodes for each step and give each one its own instructions. - **Play scripted lines verbatim**: Add speak nodes for greetings, disclosures, or compliance statements that must be delivered word-for-word, with no model turn. - **Run tools as standalone steps**: Add tool nodes that execute a shared tool deterministically — no model turn — and route on the outcome. - **Control how prompts combine**: Append node instructions to the assistant's base instructions, or replace the base instructions for a specific step. - **Tune individual steps**: Override the model, voice, tools, or transcription behavior for a node when that step needs different capabilities. - **Connect steps with conditional routing**: Use LLM conditions for natural-language decisions and variable comparisons for deterministic decisions. - **React to tool outcomes**: Route based on whether a configured tool succeeds or fails. - **Move across assistants**: Send a conversation from one workflow into another assistant when a specialist configuration should take over. - **Review workflow context in transcripts**: Use conversation history to understand which workflow step produced assistant messages. ## Create a workflow in the Portal 1. Open **AI Assistants** in the Telnyx Portal. 2. Select an assistant. 3. Open the **Workflow** tab. 4. Enable workflows if the assistant does not already have one. 5. Add nodes for each major stage of the conversation. 6. Connect nodes with edges and configure the condition for each edge. 7. Save the assistant. The workflow canvas lets you pan, zoom, drag nodes, and inspect nodes or edges from the side panel. ![Workflow editor for a Front Desk Receptionist assistant showing a start node, transfer node, FAQ node, conditional edges, and a node inspector panel](/assets/images/ai-assistant-workflow-front-desk-editor.png) For example, a front desk assistant can start at **Greeting & Identify Intent**, route office-hours questions to **Answer FAQ**, and route callers who need a person or department to **Transfer Call**. If an assistant is still using the legacy handoff tool flow, the Portal shows the legacy handoff graph instead of the new workflow editor. Remove or migrate the legacy handoff setup before building a new conversation workflow for that assistant. ## Configure workflow nodes A workflow node represents one active stage of the conversation. ### Node types Workflows support three kinds of nodes: - **Prompt node**: An LLM-driven step. The assistant generates its response from the node's instructions (combined with the assistant's base instructions), using the node's model, voice, and tool settings. Most workflow steps are prompt nodes. This is the default node type. - **Speak node**: A deterministic step that plays a fixed, scripted message and then advances. The assistant does **not** call the model on a speak node, so the wording is exactly what you configure. Use speak nodes for greetings, disclosures, hold messages, compliance statements, or any line that must be delivered verbatim. - **Tool node**: A deterministic step that runs a single shared (org-level) tool and then advances. The assistant does **not** call the model on a tool node — reaching the node executes the tool directly. Use tool nodes for actions that must happen at a known point in the flow, such as charging a card, posting a booking, or ending the call, rather than leaving the timing to the model's judgment. Add a node from the **Add Node** menu and choose **Prompt node**, **Speak node**, or **Tool node**. You can also drag an edge from an existing node to an empty area of the canvas and pick the node type from the menu. #### Speak node message A speak node centers on a single **Message** field that defines the line to deliver. - The message is delivered verbatim — there is no model turn, so the caller hears exactly what you type. - The message supports `{{variable}}` placeholders. Both system variables and custom dynamic variables defined on the assistant are interpolated at runtime, with inline highlighting and autocomplete in the editor. - A speak node must always have **exactly one outgoing [default edge](#default-conditions)**. Because a speak node does not make a routing decision itself, it advances along this default edge after delivering the message. ```text Thanks for calling Acme, {{first_name}}. This call may be recorded for quality and training. ``` The message uses the assistant's (or flow's) configured voice to deliver the scripted line. #### Tool node execution A tool node runs one shared (org-level) tool as a deliberate step in the flow. When the conversation reaches a tool node, the tool executes immediately — there is no model turn — and the flow continues along the node's outgoing edges. - **Tool**: The node references a shared tool by its ID (`shared_tool_id`). In the Portal, pick the tool from the dropdown in the node editor. - **Arguments**: Arguments are resolved from the conversation's context — a dynamic variable whose name matches one of the tool's parameters supplies that argument's value. - **Optional announcement**: A tool node can carry an optional **message** delivered just before the tool runs — for example, *One moment while I confirm that for you.* The message is delivered verbatim, supports `{{variable}}` placeholders, and is spoken before the tool executes, so the tool's result is not available in it. - **Routing on the outcome**: After the tool runs, its HTTP status code is written to the reserved system variable `telnyx_last_tool_status_code`, so a [variable-comparison edge](#variable-comparison-conditions) can route on the result — compare against `200` for success (see the API example below for the exact condition shape), or route to a fallback node otherwise. A tool node with outgoing edges must carry exactly one [default edge](#default-conditions) to take when no other condition matches. - **Terminal tool nodes**: A tool node with **no** outgoing edges is a valid end step — the tool runs and the flow stops there. Use this for tool nodes whose action ends the conversation, such as a hangup tool. Some tool types have fixed behavior regardless of edges: - A **hangup** tool node ends the call; it accepts no outgoing edges. - A **transfer** tool node hands the call off to another destination; it accepts at most one outgoing default edge, used only if the transfer fails. The instructions, tool availability, model, and voice settings described below apply to **prompt nodes**, which generate their response with the model. ### Node name Use short, descriptive names. Node names are visible in the workflow canvas and can appear in conversation transcript context. Good node names: - `Greeting & Identify Intent` - `Answer FAQ` - `Transfer Call` - `Collect appointment details` ### Node instructions Each node has instructions that control the assistant while that node is active. You can choose how the node instructions combine with the assistant's main instructions: - **Append to assistant instructions**: Keep the assistant's base behavior and add step-specific guidance. - **Replace assistant instructions**: Use only the node's instructions for this step. Append mode is usually the safest default because it preserves global policies, tone, and business rules. Replace mode is useful for tightly-scoped steps where the assistant should follow a very different prompt. ### Tool availability Each node controls which of the assistant's tools the model can call while that node is active. This lets you configure every tool once on the assistant, then expose only the relevant subset at each step. A node's **Tools** tab shows two groups: - **Inherited tools**: Every tool configured on the assistant. Each tool has a toggle so you can enable or disable it for this specific node. Tools are enabled by default, so a new node starts with the full assistant toolset until you turn tools off. - **Added for this node**: Tools attached to this node only. Use the **Select a new tool** dropdown to add a node-specific tool that should not be available elsewhere in the workflow. For example, a `Take Message` node might disable the **Transfer** tool while keeping a `take-message` webhook and a **Hang Up** tool enabled, so the assistant can only capture and close out the message during that step. ![Workflow node inspector open to the Tools tab for a Take Message node, showing inherited tools with per-node toggles (Transfer off, take-message webhook and Hang Up on) and a Select a new tool dropdown](/assets/images/ai-assistant-workflow-node-tools.png) Scoping tools per node is one of the most effective ways to make each step reliable. When a node only exposes the tools that fit its job, the model has fewer choices to weigh, calls the right tool more consistently, and is far less likely to fire an unrelated action (such as transferring a caller during an FAQ answer). ### Model override A node can use the assistant's default LLM, or override it with another Telnyx-supported model. Use model overrides when one step needs different reasoning or latency characteristics. For example: - Use a faster model for intake. - Use a stronger model for complex qualification or policy-heavy decisions. - Keep the assistant default everywhere except one high-value decision point. ### Voice override A node can inherit the assistant's default voice, or use a node-specific voice configuration. Voice overrides are useful when different conversation stages should feel different. For example: - Use the standard brand voice during greeting and intent detection. - Switch to a calmer voice for sensitive support flows. - Use a different voice when routing to a specialist assistant persona. ## Configure workflow edges An edge defines where the conversation can go next. Each edge has: - **Source node**: The node the conversation is leaving. - **Target type**: Another workflow node, or another assistant. - **Condition type**: The logic that decides whether the edge should be followed — an **LLM** condition, a **variable comparison**, or a **default** fallback. ### Evaluation order When a node has several outgoing edges, how order matters depends on the condition types involved. Variable-comparison edges are evaluated by Telnyx in declaration order, and the first one that is true wins — the remaining edges are not considered, even if their conditions would also be true. The default edge is the exception: it is considered last, regardless of where it sits in the list, so it only fires when no conditioned edge has matched. Among LLM-condition edges, the channel decides: on calls they are offered to the model as transition tools in the same declaration order, but the model selects which one to call, so a later-declared edge can still fire even when an earlier one also matches; on chat channels the post-reply evaluation returns a verdict for every LLM edge, and the first true verdict in declaration order wins. A variable-comparison edge and an LLM edge do not compete in a single ordered pass. On calls, variable-comparison edges are evaluated before the model turn begins, and a matching edge moves the conversation before the model is offered any LLM-condition transition tools — so a variable-comparison edge takes precedence over an LLM edge on calls regardless of where each sits in the `edges` array. On chat channels, a variable-comparison edge that is true when the turn begins routes the conversation before the reply is generated, the same pre-emption; after the reply, all conditioned edges that remain are considered together in declaration order, so an LLM edge declared before a comparison edge that only became true during the turn wins there. The declaration order is the `edges` array order in the assistant's `conversation_flow` object. When you create or update an assistant through the [Assistants API](/api-reference/assistants/create-an-assistant), you control this order directly: for variable-comparison edges, the first edge in the array is the one evaluated first. Because order decides which variable-comparison edge wins, declare those edges in priority order. For example, when a tool returns availability for several days and each day has its own variable-comparison edge, list the edge for the first day you want to book first; the first matching day in the list is the one the conversation takes. ```text Check availability (tool node) ├── Book Tuesday when tuesday_available == true ← evaluated first ├── Book Friday when friday_available == true └── Offer callback when default ← evaluated last ``` ### LLM conditions Use an LLM condition when the routing decision depends on conversation meaning. How an LLM-condition edge is decided depends on the channel the conversation runs on: - **On calls**, LLM-condition edges are decided by the same model turn that produces the assistant's reply. Telnyx does not run a separate evaluation of the edge prompt after the response. Instead, each LLM-condition edge leaving the active node is offered to the model alongside the assistant's tools as a tool named `transition__`, with the edge's prompt as the tool description. The transition happens only when the model selects that tool; if the model replies without selecting a transition tool, the conversation stays on the current node. - **On chat channels** (web chat and the chat API), the reply is produced first, and then every LLM-condition edge leaving the active node is evaluated in a single separate model call: the edge prompts are listed as statements with the recent conversation history, and the model returns a true/false verdict for each edge. The first edge whose verdict is true, in declaration order, wins. Because this evaluation call does not use the assistant's instructions, forbidding tool calls in instructions does not stop chat LLM edges from firing. On calls, instructions are load-bearing for LLM-edge routing. Because the transition fires through a tool call, instructions that tell the assistant never to call tools (for example, "do not use the transfer tool; the call moves on by itself") can also stop workflow edges from firing: the model obeys the instruction, acknowledges the request, and the workflow silently stays on the node. If a call workflow routes transfers, hangups, or other actions through LLM edges, the instructions must permit the model to call the routing tool — or direct it explicitly, for example "when the caller asks to be transferred, call the transition tool." Whether an instruction actually blocks routing depends on the model, so test every routing path with the assistant's production model. Example conditions: ```text The caller has a general office-hours, location, or FAQ question. ``` ```text The caller wants to speak with a specific person or department. ``` ```text The user has asked to speak with a human agent. ``` LLM conditions are best for intent, sentiment, completeness, and other natural language judgments. ### Variable comparison conditions Use a variable comparison when the routing decision should be deterministic. Variable comparison conditions can use system variables and custom dynamic variables defined on the assistant. Examples: - `telnyx_conversation_channel == "phone_call"` - `customer_tier == "enterprise"` - `telnyx_conversation_duration_secs >= 30` - `telnyx_shaken_stir_attestation != "a"` - `telnyx_last_tool_status_code == "200"` (the tool node's last execution succeeded, voice) - `telnyx_last_tool_status_code == 200` (the same check on chat channels, where the status code is a number) Variable comparisons are best for account state, channel-specific behavior, elapsed conversation time, authentication flags, or data returned by a dynamic variables webhook. ### Default conditions A **default** condition is a fallback edge that is followed whenever no other outgoing edge's condition matches. It has no prompt or expression to configure — it simply defines where the conversation goes by default. Default conditions are required for nodes that do not make their own routing decision: - A **speak node** delivers a scripted message and then advances. It must have **exactly one** outgoing default edge so the conversation always has a defined next step. - A **tool node** runs its tool and then advances. If it has any outgoing edges, it must have **exactly one** default edge to take when no conditioned edge matches. A tool node with no outgoing edges at all is also valid — the tool runs and the flow ends there. - Default conditions are only valid on edges that **leave a speak or tool node**. They are not used on edges leaving a prompt node, which route based on LLM or variable comparison conditions. When you draw the first edge out of a speak node in the Portal, it is automatically created as a default edge. If a speak node already has its default edge, any additional edges you draw fall back to an LLM condition that you can configure. The Portal gives tool nodes the same treatment, oriented around the tool's result: the first edge you draw from a tool node is created as a variable comparison against `telnyx_last_tool_status_code` (success when the tool's HTTP status code is `200`), and a second edge is created as the default fallback. This matches what voice assistants need; on chat channels use a number literal, as shown in the API example below. A speak node with zero default edges, or more than one, is invalid and cannot be saved. Make sure every speak node has exactly one outgoing default edge. A tool node with outgoing edges follows the same rule — exactly one default edge among them. ## Configure workflows with the API You can also configure workflows programmatically through the [Assistants API](/api-reference/assistants/create-an-assistant). Workflows are stored on the assistant as `conversation_flow`. The API accepts the full workflow graph when you create or update an assistant. To change one node or edge, send the updated `conversation_flow` object with the assistant update request. Assistant updates treat `conversation_flow` atomically. If you omit the field, the existing workflow is unchanged. If you send `conversation_flow: null`, the workflow is cleared. A simplified workflow payload looks like this: ```json { "conversation_flow": { "start_node_id": "n_greeting", "nodes": [ { "id": "n_greeting", "name": "Greeting & Identify Intent", "instructions": "Quickly determine whether the caller needs a person, a department, or a general FAQ answer.", "instructions_mode": "append" }, { "id": "n_faq", "name": "Answer FAQ", "instructions": "Answer hours, location, and general information questions. If they need more detail, offer to connect them to the right person.", "instructions_mode": "append" }, { "id": "n_transfer", "name": "Transfer Call", "instructions": "Confirm the destination department or person, then verbally confirm the transfer before connecting the caller.", "instructions_mode": "append" } ], "edges": [ { "id": "e_greeting_to_faq", "start_node_id": "n_greeting", "target": { "type": "node", "node_id": "n_faq" }, "condition": { "type": "llm", "prompt": "The caller has a general office-hours, location, or FAQ question." } }, { "id": "e_greeting_to_transfer", "start_node_id": "n_greeting", "target": { "type": "node", "node_id": "n_transfer" }, "condition": { "type": "llm", "prompt": "The caller wants to speak with a specific person or department." } } ] } } ``` ### Speak nodes and default edges in the API A speak node uses `"type": "speak"` and carries its scripted text in the `message` field instead of `instructions`. Its single outgoing edge uses a `default` condition. ```json { "conversation_flow": { "start_node_id": "n_disclosure", "nodes": [ { "type": "speak", "id": "n_disclosure", "name": "Recording Disclosure", "message": "Thanks for calling Acme, {{first_name}}. This call may be recorded for quality and training." }, { "type": "prompt", "id": "n_intake", "name": "Identify Intent", "instructions": "Find out what the caller needs and route them accordingly.", "instructions_mode": "append" } ], "edges": [ { "id": "e_disclosure_to_intake", "start_node_id": "n_disclosure", "target": { "type": "node", "node_id": "n_intake" }, "condition": { "type": "default" } } ] } } ``` The prompt node's `"type"` field is optional and defaults to `"prompt"`. A `default` condition takes no `prompt` or `expression`, and is only valid on an edge that leaves a speak or tool node. ### Tool nodes in the API A tool node uses `"type": "tool"` and references a shared tool by `shared_tool_id`. The node below runs a webhook tool that books an appointment, announces it first with the optional `message` field, and routes on the outcome: a variable-comparison edge checks the `telnyx_last_tool_status_code` system variable for success, and a `default` edge covers the failure path. ```json { "conversation_flow": { "start_node_id": "n_confirm", "nodes": [ { "type": "prompt", "id": "n_confirm", "name": "Confirm details", "instructions": "Confirm the appointment details, then let the caller know the booking is being made.", "instructions_mode": "append" }, { "type": "tool", "id": "n_book", "name": "Book appointment", "message": "One moment while I book that for you.", "shared_tool_id": "e5d1a3aa-6e8d-4e23-9b7c-9d5b6e4c1a90" }, { "type": "prompt", "id": "n_done", "name": "Confirm booking", "instructions": "The booking is complete. Confirm the date and time to the caller and ask if they need anything else.", "instructions_mode": "append" }, { "type": "prompt", "id": "n_retry", "name": "Booking failed", "instructions": "Apologize and offer to take the caller's details and complete the booking manually.", "instructions_mode": "append" } ], "edges": [ { "id": "e_confirm_to_book", "start_node_id": "n_confirm", "target": { "type": "node", "node_id": "n_book" }, "condition": { "type": "llm", "prompt": "The caller has confirmed the details and the booking can be made." } }, { "id": "e_book_success", "start_node_id": "n_book", "target": { "type": "node", "node_id": "n_done" }, "condition": { "type": "expression", "expression": { "type": "comparison", "op": "==", "left": { "type": "variable", "name": "telnyx_last_tool_status_code" }, "right": { "type": "string_literal", "value": "200" } } } }, { "id": "e_book_default", "start_node_id": "n_book", "target": { "type": "node", "node_id": "n_retry" }, "condition": { "type": "default" } } ] } } ``` The `message` field is optional; omit it to run the tool silently. The success comparison above is the condition shape the Portal creates for voice assistants. On chat channels the status code is recorded as a number instead of a string, so compare against a number literal (`200`) there. ## Route to another assistant An edge can target another assistant instead of another node in the same workflow. Use assistant routing when a conversation should move to a different specialist configuration, such as: - A sales assistant routing qualified technical questions to a solutions assistant. - A front-desk assistant routing billing questions to a billing assistant. - A general support assistant routing high-risk cases to a stricter compliance assistant. This keeps each assistant focused while still letting customers move through a connected experience. ## Example workflow patterns ### Front desk receptionist Use one assistant to greet callers, identify intent, and route to the right next step. ```text Greeting & Identify Intent ├── Answer FAQ when the caller asks about hours, location, or general information └── Transfer Call when the caller wants a person or department ``` ### Appointment booking Guide the customer through a sequence of required information before confirmation. ```text Collect request → Collect availability → Confirm details → Final confirmation ``` Use variable comparisons for deterministic gates, such as whether a required dynamic variable exists, and LLM conditions for softer gates, such as whether the customer has verbally confirmed the details. ### Escalation after timeout Use the conversation duration system variable to escalate when the assistant has spent too long in a step. ```text Troubleshooting ├── Continue troubleshooting when the issue is not resolved └── Route to specialist when telnyx_conversation_duration_secs >= 300 ``` ### Multi-assistant specialization Use a workflow edge to route from a general assistant into another assistant. ```text Main support assistant ├── Billing assistant ├── Technical support assistant └── Sales assistant ``` This pattern is useful when each destination needs its own model, voice, instructions, tools, or operational ownership. ## Test and debug workflows After saving a workflow, test the assistant with realistic conversations that exercise each route. When reviewing conversation history: - Check whether the assistant followed the expected path. - Look for assistant messages labeled with workflow node context. - Open the workflow from transcript context when you need to inspect the node that produced a response. - Review dynamic variable values and webhook behavior if a variable comparison did not route as expected. - If an LLM-condition edge did not fire on a call, check the assistant's instructions for text that forbids or discourages tool calls — on calls the edge fires through a tool call, so such instructions stop routing (see [LLM conditions](#llm-conditions)). On chat channels, check the edge's prompt and the recent conversation history instead: the evaluation call judges only the prompt against that history. In this example, the assistant starts in **Greeting & Identify Intent**, the caller asks about hours, and the workflow routes the response through **Answer FAQ**. ![Conversation transcript showing a Front Desk Receptionist workflow moving from Greeting & Identify Intent to Answer FAQ after the user asks about hours](/assets/images/ai-assistant-workflow-front-desk-transcript.png) ## Best practices ### Keep nodes focused Each node should represent one clear job. If a node's instructions cover multiple unrelated tasks, split it into separate nodes. ### Limit each node to the tools it needs Configure all of your tools on the assistant, then disable the ones that do not apply on each node. A focused, node-specific toolset gives the model a clear purpose, reduces the chance it calls the wrong tool, and keeps each step predictable. As a rule of thumb, only leave a tool enabled on a node if that step is supposed to be able to use it. ### Prefer append mode for global rules Use append mode when the node should keep the assistant's normal safety rules, brand voice, and business constraints. Use replace mode only when the node truly needs a standalone prompt. ### Write edge conditions as clear decisions For LLM conditions, describe the moment when the edge should fire. Avoid vague conditions like `billing`. Prefer explicit criteria, such as `The caller has a general office-hours, location, or FAQ question.` ### Don't forbid the routing tool in instructions On calls, LLM-condition edges stop firing when instructions tell the assistant never to call tools — the transition depends on the model selecting the routing tool that the workflow offers it (see [LLM conditions](#llm-conditions)). A prohibition scoped to one named tool is not safe either: wording like "do not use the transfer tool" can suppress the workflow's routing tool as well. Keep such instructions off call assistants that run workflows, or replace them with wording that directs the routing action, such as "when the caller asks to be transferred, call the transition tool." ### Use variable comparisons for hard rules If the condition depends on structured data, use a variable comparison instead of an LLM condition. This makes routing predictable and easier to debug. ### Avoid too many paths from one node A node with many outgoing edges is harder to reason about and test. If the routing logic gets complex, add an intermediate triage node. ### Test every path before production Run through the happy path, fallback path, escalation path, and at least one negative case for every important node. ## Related resources - [Dynamic Variables](/docs/inference/ai-assistants/dynamic-variables): Personalize prompts and route with runtime data. - [Version Testing and Traffic Distribution](/docs/inference/ai-assistants/version-testing-traffic-distribution): Test assistant changes before sending all traffic to a new version. - [Integrations](/docs/inference/ai-assistants/integrations): Connect assistants to external systems and data sources. --- ### Delegation > Source: https://developers.telnyx.com/docs/inference/ai-assistants/delegation.md Delegation splits a conversation between two models: a **frontend** model that talks to the caller, and a **backend** model that does the work. The frontend keeps the caller company — fast, cheap, low latency — while the backend looks things up and runs tools. This matters most for [GPT Live voice assistants](/docs/inference/ai-assistants/gpt-live), where the frontend model cannot call tools at all. Without delegation such an assistant can hold a conversation but can never look anything up or act on the caller's behalf, which is why delegation is **enabled by default**. **Beta.** Configure delegation with `delegation_settings` on the assistant. ## How a delegation is raised The two conversation paths differ only in how the frontend asks for help. **GPT-Live (speech-to-speech).** The frontend model has no tools. When it needs work done it raises a delegation and *waits* — so the backend's answer streams back sentence by sentence, and the model speaks it as it arrives. The frontend prompt is deliberately small; the business rules belong on the backend. **Chat completion (STT → LLM → TTS).** The frontend is stripped down to a single `delegate` tool. Calling it hands the work over and returns immediately, so the conversation carries on while the backend works: > **Caller:** Can you check the status of my order? > **Assistant:** Sure, let me take a look. *(calls `delegate`, keeps talking)* > *(backend looks up the order)* > **Assistant:** It shipped Tuesday and should arrive Thursday. Either way the result is injected back as context, not spoken verbatim by the backend. ## Configure it ```bash curl -X POST https://api.telnyx.com/v2/ai/assistants/{assistant_id} \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "delegation_settings": { "enabled": true, "mode": "telnyx", "model": "openai/gpt-4o", "instructions": "You have access to the order system. Always confirm the order number before answering.", "speak_results": true } }' ``` | Field | Default | Description | |-------|---------|-------------| | `enabled` | `true` | Whether the assistant delegates work to a backend model. | | `mode` | `"telnyx"` | Who answers a delegation — see [Choosing a mode](#choosing-a-mode). | | `model` | — | The backend model. For enabled `telnyx` delegation, set this or `external_llm.model` to a supported model. | | `llm_api_key_ref` | — | Integration secret holding the API key for `model`. | | `instructions` | — | Extra instructions for the backend, in addition to the assistant's own. | | `speak_results` | `true` | Whether the backend's answer is spoken to the caller. | | `external_llm` | — | Run the backend on your own OpenAI-compatible endpoint. | When explicitly enabling `telnyx` delegation, supply `model` or `external_llm.model`. The production API rejects `{"enabled": true}` without a backend model. `delegation_settings` does not accept a raw `api_key`, and neither does its nested `external_llm`. Passing one is **rejected**, not ignored — a plaintext credential in the assistant payload is a leak, and silently falling back to a different key would hide the misconfiguration. Reference an [integration secret](/docs/inference/ai-assistants/integrations) with `llm_api_key_ref` (or `external_llm.llm_api_key_ref`) instead. ### Spoken results versus silent context `speak_results` decides what happens to the backend's answer: - `true` (default) — the result is appended as **commentary** and paraphrased aloud. Use this when the caller is waiting on the answer. - `false` — the result is kept as **silent context** that informs later answers without being read out. Use this when the backend is enriching what the assistant knows rather than answering a direct question. ### Writing backend instructions `instructions` are added to the assistant's own instructions for the backend model only. Put the business rules, lookup procedures and tool guidance here — the frontend model does not need them, and on GPT-Live it has a small context window that is better spent on holding a natural conversation. ## Choosing a mode ### `telnyx` — Telnyx runs the backend The default. Telnyx runs the backend model with the assistant's own tools, MCP servers and observability, so a delegation can do anything the assistant could do. Validation happens when you save the assistant: an unavailable model, or an `llm_api_key_ref` that does not resolve, is rejected there rather than surfacing mid-call as an assistant that talks but can never look anything up. On GPT-Live, an OpenAI backend model is handed to OpenAI's own delegation at session start and runs there, calling back into Telnyx for the assistant's tools. MCP servers are not available on that path. ### `client` — you answer the delegation Set `mode: "client"` and Telnyx relays each delegation to your own server over the WebSocket configured in [`websocket_settings`](/docs/inference/ai-assistants/conversation-event-stream). Use this when the answer has to come from a system that cannot be reached as a Telnyx tool — an internal service behind your own auth, a model you host, a human in the loop. Telnyx sends: ```json { "type": "session.delegation.created", "delegation": { "id": "5b1d9e02-4c77-4f3a-8a2e-1c6f0b9d7a54", "request": "Check the status of order 40192.", "instructions": "You have access to the order system." } } ``` You answer with the same `id`: ```json { "type": "session.delegation.completed", "delegation": { "id": "5b1d9e02-4c77-4f3a-8a2e-1c6f0b9d7a54", "output": "Order 40192 shipped on Tuesday and is due to arrive Thursday." } } ``` Things to know about `client` mode: - **It requires an active event-stream socket.** If none is connected when a delegation is raised, the delegation is refused: the assistant tells the caller it cannot look things up right now and carries on from what it already knows. It never leaves the caller waiting in silence. - **The answer is text only.** The event stream offers your server no tool vocabulary, so `output` is plain text that gets injected as context. - **An empty answer is a failure, not a quiet success.** The assistant has already told the caller it is checking, so an `output` that is blank after trimming is rejected rather than producing dead air. - **`request` is `null` on GPT-Live.** The live model raises a delegation with no text of its own, so you work from the conversation events the socket is already streaming you. - **Answer promptly.** A delegation that is not answered in time falls back the same way an unavailable socket does. ## Model requirements For a complete assistant configuration, including the speaking model's API key and a matching voice, see [GPT Live voice assistants](/docs/inference/ai-assistants/gpt-live#configure-the-speaking-model). The pairing requirements are: | Setting | Requirement | |---------|-------------| | `model` | The model family must start with `gpt-live` (for example `openai/gpt-live-...`). | | `voice_settings.voice` | Must be prefixed `OpenAILive.` | The two have to be paired. An assistant with a GPT-Live model and a non-live voice — or the reverse — is rejected when you save it, rather than failing when a call arrives. Conversation flows are not supported on the GPT-Live route. The backend model in `delegation_settings.model` has no such restriction: it is an ordinary text model, and it is where tool use actually happens. ## Next steps - [GPT Live voice assistants](/docs/inference/ai-assistants/gpt-live) — configure the speech-to-speech model and voice, then connect a call - [Conversation event stream](/docs/inference/ai-assistants/conversation-event-stream) — the WebSocket that `mode: "client"` delegations travel over - [Tools library](/docs/inference/ai-assistants/tools-library) — the tools a `telnyx`-mode backend can call - [Custom LLM](/docs/inference/ai-assistants/custom-llm) — running a model on your own endpoint with `external_llm` --- ### A2A Agents > Source: https://developers.telnyx.com/docs/inference/ai-assistants/a2a-agents.md The `a2a_agents` field lets an assistant hand a question to a remote agent that speaks the [A2A (Agent2Agent) protocol](https://a2a-protocol.org/). You register the agent's URL and credentials once; the skills it advertises on its agent card become tools the assistant can call. This is the point of configuring *agents* rather than tools. When the remote agent gains, renames, or drops a skill, the assistant follows without a change to your assistant configuration — on the first conversation that starts after the five-minute [card cache](#how-it-works) expires. It is the same bargain MCP makes with `list_tools`. **Beta:** A2A agent support is in beta. The `a2a_agents` configuration described here may change before general availability. In this guide, you will learn: - How an agent card becomes tools - How to configure an agent, authenticate to it, and pass secrets - When to make a delegation asynchronous - Which limits apply and what happens when an agent is unreachable --- ## How it works 1. **You configure an agent** — a name, a URL, and any headers it needs. 2. **The card is fetched at conversation start.** `/.well-known/agent-card.json` is appended to the URL's path, unless that path already ends in `.json`, in which case the URL is used as-is. 3. **Each advertised skill becomes one tool**, named `a2a__`, described by the skill's own description from the card. 4. **The model calls a tool** with a single free-text `message`. A2A agents route on the message itself, so there is no per-skill parameter schema to mirror. 5. **The agent's reply comes back as the tool result**, and the assistant speaks it. Cards are cached for five minutes, so only the first conversation to use an agent pays the round trip. --- ## Configure an agent Add `a2a_agents` when you create or update an assistant: ```json { "name": "Support assistant", "model": "meta-llama/Meta-Llama-3.1-70B-Instruct", "instructions": "You are a support assistant. Use the billing agent for anything about invoices or payments.", "a2a_agents": [ { "name": "billing_agent", "url": "https://agents.example.com/billing", "headers": [ { "name": "Authorization", "value": "Bearer {{#integration_secret}}billing_agent_token{{/integration_secret}}" } ], "timeout_ms": 30000 } ] } ``` With that configuration, the card at `https://agents.example.com/billing/.well-known/agent-card.json` is fetched at conversation start. A card advertising skills `refund_status` and `invoice_lookup` gives the assistant two tools: `a2a_billing_agent_refund_status` and `a2a_billing_agent_invoice_lookup`. ### Fields | Field | Required | Description | | --- | --- | --- | | `name` | Yes | Identifies the agent and seeds its derived tool names. At most 43 characters. | | `url` | Yes | The agent's base URL, or its agent-card URL directly. At most 2,048 bytes. | | `headers` | No | Sent both when fetching the card and on every call to the agent. | | `async` | No | Defaults to `false`. See [Synchronous and asynchronous delegation](#synchronous-and-asynchronous-delegation). | | `timeout_ms` | No | Total budget for one call, including polling. Defaults to the assistant's tool timeout. | | `poll_interval_ms` | No | How often to poll a task that has not settled. Defaults to `500`. | | `messages` | No | Filler messages spoken while the call is in progress. | ### URL rules The URL must use `http://` or `https://` and resolve to a host on the public internet. Internal destinations — `localhost`, private and reserved IP ranges, `.local` domains — are rejected when the assistant is saved, as are hostnames written in abbreviated, hexadecimal, or octal numeric form. The hostname may not contain a `{{...}}` placeholder: the host a placeholder resolves to cannot be checked at save time. Placeholders in the path are fine, because the host stays fixed. --- ## Authenticate to the agent Header values are stored exactly as written and resolved per conversation. A value can be: - a literal, such as `Bearer abc123` - a `{{dynamic_variable}}`, resolved from the conversation's [dynamic variables](/docs/inference/ai-assistants/dynamic-variables) - an `{{#integration_secret}}identifier{{/integration_secret}}` section, resolved from a stored [integration secret](/docs/inference/ai-assistants/integrations) Prefer an integration secret for anything long-lived: the secret itself never enters the assistant configuration, so it is never returned when you read the assistant back. A2A headers do **not** accept the encrypted `{{variable | encryption_secret_ref}}` form used for [per-caller credentials](/docs/inference/ai-assistants/per-caller-credentials). That reference is resolved only in an MCP server's `api_key_ref` and in tool webhook header values; using it here is rejected with a `422` when you save the assistant. For a value that has to vary by caller, return it as a plain `{{dynamic_variable}}` from the dynamic variables webhook — bearing in mind that plain dynamic variables are not encrypted and are also interpolated into instructions and messages. A header whose name or value resolves to an empty string is dropped for that conversation. If an agent starts returning `401`, check that every placeholder in its headers actually resolves. --- ## Synchronous and asynchronous delegation By default a delegation is synchronous: the turn waits for the remote agent, up to `timeout_ms`. If the agent answers with a task that is still running, it is polled every `poll_interval_ms` until it settles, and cancelled if the budget runs out. Set `"async": true` when the agent is slow enough that waiting would leave dead air. The model is told the request is on its way and the turn completes immediately; when the answer lands it is added to the conversation without interrupting the caller, and the assistant picks it up on its next turn. ```json { "name": "research_agent", "url": "https://agents.example.com/research", "async": true } ``` Use synchronous delegation when the answer is the caller's next sentence, and asynchronous delegation when the caller can keep talking without it. For a synchronous agent that takes a few seconds, `messages` fills the silence the same way [webhook filler messages](/docs/inference/ai-assistants/filler-messages) do: ```json { "name": "billing_agent", "url": "https://agents.example.com/billing", "messages": [ { "type": "request_start", "content": "Let me check that with our billing team." }, { "type": "request_response_delayed", "content": "Still looking, one moment.", "timing_ms": 4000 } ] } ``` Filler messages are not used when `async` is `true` — there is no silence to fill. --- ## Tool names Derived names are `a2a__`, truncated at 64 characters. Characters outside `[A-Za-z0-9_]` in the agent name are replaced with `_` before the name is built, so `billing agent` and `billing_agent` produce the same tool names and cannot both be configured on one assistant — saving the second is rejected. The 43-character limit on `name` exists for the same reason: a longer name would consume the whole 64-character budget, leaving every skill on that agent with the same truncated tool name. If a derived name still collides with another tool the assistant already has, a `_2`, `_3`, … suffix is added so that dispatch stays unambiguous. --- ## Limits These are applied when the conversation starts, not when the assistant is saved. Anything past them is dropped silently, so keep configurations inside them: | Limit | Value | | --- | --- | | Agents per assistant | 64 | | Skills read per card | 64 | | Tools derived per assistant | 128 | | Budget for all card fetches at conversation start | 6 seconds | | Agent card size | 1 MB | | Skill description, truncated past this | 2 KB | | Skill ID, skill skipped past this | 512 bytes | | Agent reply text kept | 8 KB | Agents are read in the order you configure them and the limits are applied in that order, so put the agents that matter most first. --- ## When an agent is unavailable A card that cannot be fetched, times out, or advertises no skills yields no tools. **The assistant loses that capability for the conversation; it does not lose the call.** If the model is asked to do something only that agent could do, it will say it cannot rather than fail. Once a tool exists and a call to it fails, the failure is reported to the model as ordinary text — "The agent did not answer in time", "The agent could not be reached" — so the assistant can acknowledge it and move on. That failure mode is quiet by design, which is why a malformed configuration is rejected at save time instead: a typo you can see in a `422` is better than a tool that silently never appears. ### Troubleshooting | Symptom | Check | | --- | --- | | The agent's tools never appear | Fetch the resolved card URL yourself with the same headers — your `url` unchanged if its path ends in `.json`, otherwise `/.well-known/agent-card.json`. It must return `200` with a card that lists skills. | | Only some skills appear | The card may advertise more than 64 skills, or the assistant may already be at its 128-tool ceiling. | | Tools appear but every call fails | Check the agent's JSON-RPC endpoint — the one advertised on the card, which is not necessarily the card's own URL. | | Changes to the card take a few minutes | Cards are cached for five minutes. | --- ## API reference - [Create an assistant](/api-reference/assistants/create-an-assistant) - [Update an assistant](/api-reference/assistants/update-an-assistant) --- ### Agent Handoff > Source: https://developers.telnyx.com/docs/inference/ai-assistants/agent-handoff.md Agent handoff enables your AI assistant to seamlessly transfer conversations to other specialized AI assistants while preserving full context. This allows you to build a team of expert agents, each focused on specific domains or tasks, working together to provide comprehensive support in a single conversation. In this guide, you will learn how to: - Configure agent handoff for your AI assistants. - Choose between Unified and Distinct handoff modes. - Design effective multi-agent architectures. - Implement agent handoff via API and Portal. --- ## Overview Agent handoff represents a powerful approach to building sophisticated AI systems. Instead of creating one complex agent that handles everything, you can build a team of specialized agents that collaborate seamlessly, each bringing domain expertise to the conversation. Agent handoff is model-agnostic and works with any AI model supported by Telnyx, including OpenAI GPT models, Meta Llama models, Anthropic Claude, Qwen, and others. Each agent in the handoff chain can use a different model based on its specialized needs. ### How agent handoff works The agent handoff lifecycle follows these steps: 1. **Detection**: The current agent identifies that another agent would be better suited to handle the user's request based on intent, domain, or task requirements. 2. **Agent selection**: The system determines which specialist agent should receive the handoff. 3. **Context transfer**: Full conversation history, user data, and relevant context is transferred to the target agent. 4. **Transition**: The handoff occurs, either seamlessly (Unified mode) or explicitly (Distinct mode). 5. **Continuation**: The target agent continues the conversation with full context awareness. ### Unified vs Distinct modes Agent handoff supports two modes that determine how the transition appears to the user: **Unified Mode (Default)** In Unified mode, assistants share the same context and voice, creating a seamless experience where specialists work behind the scenes. The user experiences one consistent agent, even though multiple specialized assistants are handling different parts of the conversation. - **Same voice**: All agents use the same voice configuration. - **Transparent transition**: User doesn't notice the handoff. - **Shared context**: Full conversation history available to all agents. - **Use case**: When you want a unified brand experience. **Distinct Mode** In Distinct mode, each assistant retains its voice configuration, creating a conference call experience with multiple distinct voices. The user hears different voices as they're transferred between specialists. - **Individual voices**: Each agent uses its own voice configuration. - **Explicit transition**: User hears "I'm transferring you to [specialist name]". - **Shared context**: Full conversation history available to all agents. - **Use case**: When you want to highlight specialist expertise. ### Key benefits - **Specialist routing**: Direct queries to domain experts who provide accurate, focused responses. - **Context preservation**: Full conversation history travels with the handoff, eliminating repetition. - **Reduced complexity**: Build focused agents instead of one mega-agent trying to do everything. - **Better user experience**: Users get the right expert for their specific needs. - **Easier maintenance**: Update individual specialist agents without affecting the entire system. - **Scalable architecture**: Add new specialists as your product or service expands. ### Common use cases - **Multi-domain support**: Transfer between technical support, billing, and sales departments with full context. - **Workflow automation**: Information gathering agent → processing agent → confirmation agent. - **Triage and routing**: Assessment agent evaluates needs, then routes to appropriate specialist. - **Language switching**: Detection agent identifies language, hands off to language-specific agent. - **Escalation tiers**: Level 1 support → Level 2 support → Level 3 specialist. - **Task segmentation**: Browse/Research agent → Purchase agent → Post-sale support agent. --- ## Best practices Following these best practices will help you design effective multi-agent systems and maximize the value of agent handoff. ### Agent architecture design **When to split agents vs one agent** Split into multiple agents when: - Domains require significantly different knowledge bases. - Agents need different tools or integrations. - Response patterns differ substantially. - You want to independently update different capabilities. Keep as one agent when: - Tasks are closely related with significant overlap. - Context switching would reduce quality. - User experience benefits from continuity. - The domain is narrow and well-defined. **Specialization patterns** - **By domain**: Technical support, billing, sales, product information. - **By task**: Information gathering, transaction processing, confirmation. - **By customer segment**: VIP customers, trial users, enterprise accounts. - **By complexity**: Simple queries, complex troubleshooting, escalations. - **By channel**: Phone, chat, email (if behavior differs significantly). **Context requirements** Each specialist agent should receive: - Full conversation history with previous agents. - User identification and account information. - Previous interactions and preferences. - Current intent and goals. - Any decisions or commitments made. ### Handoff trigger design **Clear criteria for handoff** Define explicit triggers: - **Intent-based**: "I need to talk to billing" → Billing agent. - **Keyword-based**: User mentions "refund" → Billing agent. - **Task completion**: Information gathered → Processing agent. - **Capability-based**: Technical question beyond triage scope → Technical specialist. - **Sentiment-based**: Frustrated customer → Senior support agent. **Avoiding handoff loops** Prevent agents from endlessly transferring to each other: - Define clear responsibility boundaries for each agent. - Implement handoff history tracking. - Set maximum handoff limits per conversation. - Design failsafe: after N handoffs, route to human agent. - Test handoff decision logic thoroughly. **Graceful transitions** Configure agents to announce handoffs clearly: **Unified mode**: ``` "Let me look into your billing question..." [Handoff to billing agent occurs seamlessly] [Billing agent continues without announcing the transition] ``` **Distinct mode**: ``` "I'm transferring you to Sarah, our billing specialist, who can help with your refund request." [Handoff occurs] "Hi, this is Sarah from billing. I can see you're asking about a refund..." ``` Use the [Transfer Call API](/api-reference/call-commands/transfer-call) to manage transfers between assistants programmatically. ### Context preservation strategies **What context to pass** Essential context to transfer: - **User information**: Name, account ID, customer tier. - **Conversation history**: What was discussed with previous agents. - **User intent**: What the user is trying to accomplish. - **Collected data**: Information gathered (order numbers, issue details, etc.). - **Agent actions**: What previous agents already tried. - **Sentiment**: User's emotional state (frustrated, satisfied, confused). **Using dynamic variables** Leverage dynamic variables to pass structured context: - `{{customer_name}}`, `{{account_id}}`, `{{customer_tier}}`. - `{{issue_type}}`, `{{priority_level}}`. - `{{previous_agent}}`, `{{handoff_reason}}`. - Custom variables specific to your domain. Learn more about [dynamic variables](https://developers.telnyx.com/docs/inference/ai-assistants/dynamic-variables). **Memory configuration** Configure agent memory settings to ensure context persistence: - Enable conversation memory for all agents in the handoff chain. - Use consistent memory keys across agents. - Set appropriate memory retention periods. Learn more about [memory configuration](https://developers.telnyx.com/docs/inference/ai-assistants/memory). ### Industry-specific templates **Healthcare: Triage to Specialist** ``` Scenario: Patient calls about symptoms Triage Agent: - Collects symptoms, medical history, urgency level. - Assesses severity and appropriate specialty. - Gathers insurance information. Handoff Trigger: - Specific condition identified (cardiology, dermatology, etc.). - Severity level requires specialist consultation. Specialist Agent (e.g., Cardiology): - Reviews triage notes and symptoms. - Provides condition-specific guidance. - Schedules appointment with cardiologist. - Explains what to expect. Benefits: - Triage agent handles initial screening efficiently. - Specialist agent provides expert medical guidance. - No need to repeat symptoms or medical history. ``` **E-commerce: Browse → Purchase → Support** ``` Scenario: Customer shopping for products Browse Agent: - Answers product questions. - Provides recommendations. - Helps compare options. - Adds items to cart. Handoff Trigger: User ready to purchase Purchase Agent: - Guides through checkout. - Handles payment processing. - Applies discounts/promo codes. - Confirms order. Handoff Trigger: Post-purchase question Support Agent: - Tracks order status. - Handles returns/exchanges. - Resolves delivery issues. Benefits: - Each agent optimized for specific task. - Smooth shopping experience. - Consistent context across journey. ``` **Financial Services: Authentication → Account → Fraud** ``` Scenario: Customer calling about suspicious activity Authentication Agent: - Verifies customer identity. - Security questions. - Multi-factor authentication. Handoff Trigger: Authentication successful Account Agent: - Accesses account information. - Reviews recent transactions. - Answers general account questions. Handoff Trigger: Suspicious activity detected Fraud Agent: - Investigates flagged transactions. - Freezes card if necessary. - Issues replacement card. - Sets up fraud monitoring. Benefits: - Security handled by specialized agent. - Fraud expertise when needed. - Reduces risk of errors. ``` **Customer Support: Tier 1 → Tier 2 → Tier 3** ``` Scenario: Technical support escalation Tier 1 Agent: - Handles common issues. - Tries basic troubleshooting. - Resolves 70% of queries. Handoff Trigger: Issue not resolved after standard troubleshooting Tier 2 Agent: - Advanced troubleshooting. - System configuration expertise. - Resolves complex technical issues. Handoff Trigger: System-level issue or bug identified Tier 3 Agent (Engineering Specialist): - Deep system knowledge. - Can create bug reports. - Provides workarounds for known issues. - Escalates to development team if needed. Benefits: - Efficient resource utilization. - Expertise matched to complexity. - Full troubleshooting history preserved. ``` --- ## Configuration Agent handoff can be configured through the Portal's AI Assistant builder or via API when creating or updating assistants. ### Portal configuration 1. Navigate to your AI Assistants in the [Telnyx Portal](https://portal.telnyx.com/#/ai/assistants). ![AI Assistants List](/assets/images/ai-assistants-list.png) 2. In the Tools section, add a Handoff tool. ![Add Tool Handoff Menu](/assets/images/add-tool-handoff-menu.png) 3. Choose your voice mode: **Unified** (seamless) or **Distinct** (conference call style). 4. Enter a display name for the target assistant. 5. Select the target assistant from the dropdown menu. 6. Click the plus (+) button to add additional target assistants if needed. 7. Save your configuration. ![AI Assistant Handoff Configuration](/assets/images/handoff-config-page.png) ### How handoff appears to users **Unified Mode Experience:** The user experiences a single, consistent agent voice throughout the conversation. When a handoff occurs, the transition is completely seamless - the agent simply continues helping without announcing any change. ``` User: "I need help with my billing" Triage Agent (Voice A): "I can help you with that. What's your question about billing?" User: "I was charged twice for my last order" [Handoff to Billing Agent occurs silently] Billing Agent (Voice A): "I see you were charged twice. Let me look into that for you..." ``` **Distinct Mode Experience:** The user hears different voices as they're transferred between specialists, creating a conference call experience where each expert introduces themselves. ``` User: "I need help with my billing" Triage Agent (Voice A): "I'll transfer you to Sarah from our billing team who can help." [Handoff occurs] Billing Agent (Voice B): "Hi! This is Sarah from billing. I can see you have a question about being charged twice..." ``` --- ## API implementation ### Creating an assistant with agent handoff When creating an AI assistant via API, include the Handoff tool with target assistant IDs in the tools array: ```bash curl -L 'https://api.telnyx.com/v2/ai/assistants' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "name": "Customer Support Triage Agent", "instructions": "You are a triage agent for customer support. Listen to customer needs and determine if they need technical support, billing help, or sales assistance. Handoff to the appropriate specialist when needed.", "model": "moonshotai/Kimi-K2.5", "tools": [ { "type": "handoff", "handoff": { "voice_mode": "unified", "ai_assistants": [ { "name": "Technical Support", "id": "asst_tech_abc123" }, { "name": "Billing Support", "id": "asst_billing_def456" }, { "name": "Sales", "id": "asst_sales_ghi789" } ] } } ] }' ``` ### Distinct mode configuration To enable distinct voice mode where each agent retains its voice configuration: ```bash curl -L 'https://api.telnyx.com/v2/ai/assistants' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "name": "Triage Agent", "instructions": "You are a triage agent. When handing off, announce the specialist name.", "model": "moonshotai/Kimi-K2.5", "tools": [ { "type": "handoff", "handoff": { "voice_mode": "distinct", "ai_assistants": [ { "name": "Technical Specialist", "id": "asst_tech_xyz789" } ] } } ] }' ``` ### Updating handoff configuration You can update handoff settings for an existing assistant: ```bash curl -L 'https://api.telnyx.com/v2/ai/assistants/{assistant_id}' \ -X PATCH \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "tools": [ { "type": "handoff", "handoff": { "voice_mode": "unified", "ai_assistants": [ { "name": "Priority Support", "id": "asst_priority_abc123" }, { "name": "Standard Support", "id": "asst_standard_def456" } ] } } ] }' ``` --- ## Troubleshooting ### Common issues #### Handoff loops **Symptoms**: Agents keep handing off to each other repeatedly, creating an endless loop. **Possible Causes**: - Overlapping agent responsibilities. - Unclear handoff criteria. - Agents don't recognize when they should handle the request. **Solutions**: - Define clear, mutually exclusive responsibility boundaries for each agent. - Document what each agent handles in their instructions. - Add specific instructions: "Only handoff if the request is truly outside your domain". #### Context loss after handoff **Symptoms**: Target agent doesn't have access to previous conversation history or user information. **Possible Causes**: - Dynamic variables webhook not configured for target agent. - Target agent's memory query doesn't include relevant past conversations. - Conversation metadata not being passed consistently. **Solutions**: - Configure dynamic variables webhook for all agents in the handoff chain. - Ensure all agents use the same conversation query logic to access relevant memory. - Use consistent conversation metadata across agents. - Test handoff flows in Portal conversation history. - Verify that dynamic variables are properly passed to target agents. #### Incorrect agent selection **Symptoms**: User gets handed off to wrong specialist, requiring additional transfers. **Possible Causes**: - Triage agent misunderstands user intent. - Agent instructions too vague about when to handoff. - Similar keywords trigger wrong agent selection. **Solutions**: - Refine triage agent instructions with specific examples. - Improve intent detection with more training examples. - Add confirmation step: "It sounds like you need billing help, is that correct?". - Review conversation logs to identify misrouting patterns. - Update agent instructions based on common mistakes. #### Handoff doesn't trigger **Symptoms**: Agent tries to handle request outside its expertise instead of handing off. **Possible Causes**: - Handoff tool not properly configured. - Agent instructions don't specify when to handoff. - Target assistant IDs incorrect or missing. **Solutions**: - Verify handoff tool is added to agent. - Check target assistant IDs are valid. - Add explicit handoff triggers in agent instructions. - Test with direct requests: "I need to talk to billing". - Review agent instructions for conflicting guidance. ### Testing strategies **Test handoff triggers**: - Provide various user requests to verify correct agent selection. - Try edge cases (ambiguous requests, multi-domain questions). - Test with different phrasings of the same intent. **Verify context preservation**: - Check that target agent has full conversation history. - Confirm user doesn't need to repeat information. - Validate that collected data (account numbers, etc.) transfers. **Monitor conversation flows**: - Review conversation transcripts in Portal. - Track average number of handoffs per conversation. - Identify common handoff patterns. - Measure resolution time with vs without handoffs. **Load testing**: - Test with multiple concurrent conversations. - Verify handoffs work reliably under load. - Check that context remains accurate with multiple handoffs. --- ## Related resources - [Workflow](/docs/inference/ai-assistants/workflows) - Visualize agent handoff flows and transitions in your workflow flowchart. - [Dynamic Variables](https://developers.telnyx.com/docs/inference/ai-assistants/dynamic-variables) - Customize context and personalize handoffs. - [Memory](https://developers.telnyx.com/docs/inference/ai-assistants/memory) - Configure persistent context across conversations. - [AI Assistants API Reference](https://developers.telnyx.com/api-reference/assistants/create-an-assistant#create-an-assistant) - Complete API documentation for assistants. - [Voice Assistant Quickstart](https://developers.telnyx.com/docs/inference/ai-assistants/no-code-voice-assistant#handoff) - See handoff tool in Portal. --- ### Multi-Participant Calls > Source: https://developers.telnyx.com/docs/inference/ai-assistants/multi-participant-calls.md Multi-participant Voice AI calls let an assistant bring another person into an active call, follow who is speaking, and continue using the same tools and instructions it would use in a one-to-one voice conversation. Use this pattern when your assistant needs to coordinate between people in real time, such as scheduling a meeting, connecting a customer with a specialist, or letting multiple callers complete a task together. Listen to this example call to hear how an assistant can invite a participant, stay silent while people talk to each other, and resume when asked to take action. Audio sample of a multi-participant Voice AI call (web only): https://developers.telnyx.com/assets/audio/ai-assistants-multi-participant/example-call.mp3 In this guide, you will learn how to: - Add an **Invite** tool so your assistant can invite another participant to the current call. - Design assistant instructions for multi-participant conversations. - Use the **Skip Turn** tool so the assistant can stay silent while people talk to each other. - Review a multi-participant call in Conversation History. - Configure the same behavior through the Portal or the [Assistants API](/api-reference/assistants/create-an-assistant). --- ## How multi-participant calls work A multi-participant Voice AI call starts like any other voice assistant call. The assistant speaks with the main caller, then uses an Invite tool when it needs to bring another participant into the conversation. Once the new participant joins, the assistant can: - Tell whether the main user or an invited participant is speaking. - Respond to either participant when appropriate. - Use its configured tools, memory, dynamic variables, and integrations. - Stay silent when the participants are talking to each other. - Resume when someone addresses the assistant again. By default, an assistant will respond to every turn it receives. For natural multi-participant calls, add a Skip Turn tool and describe when the assistant should remain silent. --- ## Requirements Before you start, create a voice assistant by following the [Voice Assistant Quickstart](/docs/inference/ai-assistants/no-code-voice-assistant). You can configure assistants and tools in the Portal or through the [Assistants API](/api-reference/assistants/create-an-assistant). You also need: - A phone number or SIP URI for the assistant. - A phone number or SIP URI for each participant the assistant may invite. - Any tools the assistant should use after the participant joins, such as a calendar, CRM, or booking integration. --- ## Step 1: Add an Invite tool The Invite tool lets your assistant invite another participant into the current call. 1. In the [Mission Control Portal](https://portal.telnyx.com), open your assistant. 2. Go to the assistant's **Tools** section. 3. Add an **Invite** tool. 4. Configure the invite target with the participant's phone number or SIP URI. 5. Save the assistant. You can also configure the assistant, its model, voice settings, and tools through the [Assistants API](/api-reference/assistants/create-an-assistant). When the assistant decides that another person should join, it calls the Invite tool. After the participant joins, the assistant receives the updated conversation context and can continue the call with both participants. In the example below, the user asks the assistant to invite Enzo. The assistant calls the Invite tool, receives a `Participant joined` response, then confirms that Enzo has joined. ![Invite tool call shown in Conversation History](/assets/images/ai-assistants-multi-participant/conversation-invite-tool-call.png) --- ## Step 2: Give the assistant multi-participant instructions Your assistant should know that the conversation may include multiple people. Add instructions that tell the assistant how to behave when participants are speaking to each other. Example instructions: ```text You are Amber, a voice assistant helping coordinate a multi-participant call. The call may include the main user and one or more invited participants. Pay close attention to who is speaking and who they are addressing. When the main user asks you to invite someone, use the Invite tool. After the participant joins, briefly confirm that they joined, then let the participants talk. If the participants are talking to each other and are not addressing you, use the Skip Turn tool and remain silent. If someone addresses you directly, respond normally. You can use any of your configured tools to help complete the task. ``` You can adapt this prompt for your use case. For example, a scheduling assistant might be instructed to listen while participants compare availability, then resume when asked to book a time. --- ## Step 3: Add a Skip Turn tool In a one-to-one call, an assistant usually responds after every user turn. In a multi-participant call, that can feel unnatural because the assistant may interrupt people who are talking to each other. The Skip Turn tool lets the assistant intentionally remain silent for a turn. Add a Skip Turn tool when you want the assistant to: - Stay silent while invited participants talk to the main user. - Avoid responding to side conversations. - Wait until someone directly addresses the assistant. - Let humans confirm details with each other before the assistant takes action. In this example, James asks Enzo how he is doing. The assistant recognizes that James is speaking to Enzo, not to the assistant, and uses `skip_turn` with the reason `James and Enzo are talking to each other, not addressing me`. ![Skip Turn tool call shown in Conversation History](/assets/images/ai-assistants-multi-participant/skip-turn-tool-call.png) Skip Turn does not end the call or disable the assistant. It only tells the assistant not to speak for that turn. The assistant can respond again when a participant addresses it. If your assistant should only respond when a specific name is spoken, or you know the participant names ahead of time, add those names to **Keyterm Boost** on the assistant's **Voice** tab. Boosting the names improves transcription accuracy on those exact words, which makes name-based Skip Turn rules far more reliable. Keyterm Boost accepts a comma-separated list, for example `Telnyx,Amber,Enzo`. It also supports [dynamic variables](/docs/inference/ai-assistants/dynamic-variables), so you can pass participant names or other caller-specific terms at conversation start, for example `Telnyx,{{participant_names}}`. It is supported by `deepgram/flux` and `deepgram/nova-3`. See the [Transcription Settings guide](/docs/inference/ai-assistants/transcription-settings) for details. ![Keyterm Boost field on the Voice tab of the AI Assistant settings](/assets/images/ai-assistants-multi-participant/keyterm-boost-voice-tab.png) --- ## Step 4: Use other assistant tools during the call After a participant joins, the assistant can do everything it can do in a one-to-one voice call. It can call APIs, look up information, schedule meetings, update records, send messages, and more. In this example, James asks the assistant to schedule time with Enzo. The assistant checks calendar availability, proposes a time, waits while James confirms with Enzo, then books the meeting. ![Calendar availability tool call in a multi-participant call](/assets/images/ai-assistants-multi-participant/calendar-tool-call.png) The assistant can continue to combine tools with turn-taking logic. For example: 1. The assistant checks availability. 2. The assistant proposes a time to both participants. 3. The main user asks the invited participant whether the time works. 4. The assistant stays silent while the invited participant answers. 5. The assistant books the meeting after the participants confirm. ![Meeting booking completed in a multi-participant call](/assets/images/ai-assistants-multi-participant/booking-complete.png) --- ## Best practices ### Be explicit about when to speak Tell the assistant when it should respond and when it should stay silent. Multi-participant calls work best when the assistant has clear turn-taking rules. Good instruction: ```text If James and Enzo are speaking directly to each other, use Skip Turn and stay silent. Only respond when someone addresses you by name or asks you to take an action. ``` ### Review calls in Conversation History Use Conversation History to inspect the transcript, audio, tool calls, and tool responses. This is especially useful when tuning Skip Turn instructions because you can see why the assistant decided to speak or remain silent. --- ## Next steps - Build your first assistant with the [Voice Assistant Quickstart](/docs/inference/ai-assistants/no-code-voice-assistant). - Add reusable tools from the [Tools Library](/docs/inference/ai-assistants/tools-library). - Use [dynamic variables](/docs/inference/ai-assistants/dynamic-variables) to personalize calls. - Review production calls with [Agent Observability](/docs/inference/ai-assistants/agent-observability). --- ### Warm Transfer Acceptance > Source: https://developers.telnyx.com/docs/inference/ai-assistants/warm-transfer-acceptance.md A warm transfer normally plays a recorded handoff message to the destination and then bridges the calls, whether or not the destination is ready to take them. Warm transfer acceptance adds a live consult step: once the destination answers, the assistant speaks with them privately, delivers the handoff context, and asks whether they accept the call. Only an explicit acceptance bridges the caller. While the consult happens, the caller stays with the assistant and keeps hearing ringback — they never hear the exchange. --- ## How it works 1. The assistant calls the transfer tool and Telnyx dials the selected target. 2. The destination answers. Instead of playing recorded audio, the assistant joins the destination on a private leg. 3. The assistant delivers the warm transfer message and asks whether they take the call. 4. The assistant finalizes the transfer with the built-in `complete_transfer` tool: - **Accept** — the caller and the destination are bridged and the assistant leaves the call. - **Decline** — the destination leg is hung up and the assistant returns to the caller with the reason the destination gave. 5. If the destination hangs up, or nobody decides within two minutes, the destination leg is dropped and the assistant returns to the caller. The `complete_transfer` tool is added to the assistant automatically when acceptance is enabled — you do not configure it. If your assistant already has a tool named `complete_transfer`, acceptance is disabled on that assistant to avoid a name collision, so pick a different name for your own tool. --- ## Configuration options Acceptance is configured inside the transfer tool, under `warm_transfer_acceptance`: | Field | Type | Default | Description | | --- | --- | --- | --- | | `enabled` | boolean | `false` | Whether the destination must accept before the calls are bridged. | | `end_user_target_context_mode` | `private` \| `shared` | `private` | Whether the consult with the destination is kept out of the conversation. | ### Context modes The consult is a real conversation between your assistant and the destination, so you choose whether it becomes part of the conversation record: | Mode | Behavior | | --- | --- | | `private` | The exchange never reaches the conversation history, AI Conversations, webhooks or insights. The caller-facing turn only sees the transfer tool result, rewritten with the outcome — including the reason on a decline. Use this when the handoff may contain notes the caller should not see. | | `shared` | The exchange stays in the conversation like any other messages, and is available in history, webhooks and insights. | ### Requirements - **A warm message must always be available.** Set `warm_transfer_instructions` on the transfer tool, or a `message` on every target. Saving an assistant with acceptance enabled and neither of these returns a validation error. When both are present, the target's `message` wins. - **Only for `ai_assistant_start` conversations.** Acceptance does not apply to `gather_using_ai`, Conversation Relay, or web calls. - **Single-caller conversations only.** If the conversation is in a conference or has more than one invited user participant, the transfer falls back to a regular warm transfer with recorded playback. If the consult cannot be set up for any reason, the transfer degrades gracefully: the warm transfer message is played to the destination as audio and the calls are bridged, exactly as a warm transfer without acceptance. --- ## Setting up via API Configure `warm_transfer_acceptance` within the transfer tool when creating or updating an assistant. ### Ask the destination to accept ```bash curl -L 'https://api.telnyx.com/v2/ai/assistants' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "name": "Support Agent", "instructions": "You are a support agent. Transfer calls to the support team when the caller asks for a human.", "model": "moonshotai/Kimi-K2.5", "tools": [ { "type": "transfer", "transfer": { "targets": [ { "name": "Support Team", "to": "+15551234567" } ], "from": "+15553456789", "warm_transfer_instructions": "Briefly summarize why the caller is being transferred and what they already tried. Then ask whether they can take the call.", "warm_transfer_acceptance": { "enabled": true } } } ] }' ``` ### Share the consult with the conversation record Set `end_user_target_context_mode` to `shared` when you want the exchange with the destination stored alongside the rest of the conversation: ```bash curl -L 'https://api.telnyx.com/v2/ai/assistants' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "name": "Support Agent", "instructions": "You are a support agent. Transfer calls to the support team when the caller asks for a human.", "model": "moonshotai/Kimi-K2.5", "tools": [ { "type": "transfer", "transfer": { "targets": [ { "name": "Support Team", "to": "+15551234567" } ], "from": "+15553456789", "warm_transfer_instructions": "Briefly summarize why the caller is being transferred and what they already tried. Then ask whether they can take the call.", "warm_transfer_acceptance": { "enabled": true, "end_user_target_context_mode": "shared" } } } ] }' ``` ### Give each target its own handoff message A `message` on a target replaces the message the assistant would compose from `warm_transfer_instructions`, and satisfies the warm-message requirement on its own: ```bash curl -L 'https://api.telnyx.com/v2/ai/assistants' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "name": "Support Agent", "instructions": "You are a support agent. Route callers to the right team.", "model": "moonshotai/Kimi-K2.5", "tools": [ { "type": "transfer", "transfer": { "targets": [ { "name": "Billing", "to": "+15551234567", "message": "I have a caller with a billing question. Can you take it?" }, { "name": "Technical Support", "to": "+15557654321", "message": "I have a caller with a device issue that I could not resolve. Can you take it?" } ], "from": "+15553456789", "warm_transfer_acceptance": { "enabled": true } } } ] }' ``` --- ## Combining with voicemail detection Acceptance works alongside [voicemail detection on transfer](/docs/inference/ai-assistants/voicemail-detection-on-transfer). When AMD is enabled on the transfer, the consult starts once the destination is identified as a human. If voicemail is detected, the configured voicemail action runs instead — both `stop_transfer` and `leave_message_and_stop_transfer` end the transfer, so no consult takes place and the assistant returns to the caller. --- ## Related resources - [Voice AI Assistant API Reference](/api-reference/assistants/create-an-assistant#transfertool) - Complete transfer tool API documentation, including the `warm_transfer_acceptance` parameters. - [Voicemail Detection on Transfer](/docs/inference/ai-assistants/voicemail-detection-on-transfer) - Detect voicemail on the transfer destination and respond automatically. - [Agent Handoff](/docs/inference/ai-assistants/agent-handoff) - Hand a conversation from one assistant to another instead of to a human. - [Attach an AI Assistant to a Call](/docs/voice/programmable-voice/ai-assistant-start) - Start an assistant on a live call with `ai_assistant_start`. --- ### Voicemail Detection on Transfer > Source: https://developers.telnyx.com/docs/inference/ai-assistants/voicemail-detection-on-transfer.md When a Voice AI Assistant transfers a call, the destination may go to voicemail. Without detection, the caller experiences dead air while the voicemail greeting plays. Voicemail detection on transfer solves this by identifying when a transferred call reaches voicemail and triggering a configured action automatically. The assistant stays on the line with the caller throughout, so there is never a gap in the conversation. --- ## How it works 1. The AI Assistant initiates a transfer using the transfer tool. 2. AMD (Answering Machine Detection) monitors the transfer destination. 3. If voicemail is detected, the configured action triggers immediately. 4. The assistant returns to the caller and continues the conversation. The caller remains connected to the assistant during the entire process. If a human answers the transfer destination, the call proceeds as a normal transfer. --- ## Configuration options ### Detection mode The `detection_mode` field controls whether voicemail detection is active on the transfer: | Value | Description | | --- | --- | | `disabled` | No voicemail detection (default). | | `premium` | ML-based detection with high accuracy. Recommended for production use. | ### Actions when voicemail is detected The `on_voicemail_detected.action` field determines what happens when voicemail is detected: | Action | Behavior | | --- | --- | | `stop_transfer` | Cancels the transfer immediately and returns the assistant to the caller. | | `leave_message_and_stop_transfer` | Delivers a TTS message to the voicemail, then cancels the transfer and returns to the caller. | ### Voicemail message options When using `leave_message_and_stop_transfer`, the `on_voicemail_detected.voicemail_message` object configures what message is left: | Field | Value | Description | | --- | --- | --- | | `type` | `message` | Plays a custom TTS text as the voicemail message. | | `type` | `warm_transfer_instructions` | Uses the warm transfer audio instructions as the voicemail message. | | `message` | string | The TTS text to deliver when `type` is `message`. | --- ## Setting up via the Portal 1. Navigate to **AI, Storage and Compute** in the [Telnyx Portal](https://portal.telnyx.com). 2. Select an existing AI Assistant or create a new one. 3. In the **Tools** section, open the transfer tool configuration. 4. Under **Voicemail Detection**, set the detection mode to **Premium**. 5. Choose your preferred action: **Stop the transfer** or **Leave a message and stop transfer**. 6. If leaving a message, enter the TTS text you want delivered to voicemail. 7. Save and test with a call to a number that goes to voicemail. --- ## Setting up via API Configure voicemail detection within the transfer tool when creating or updating an assistant. ### Stop the transfer on voicemail This configuration cancels the transfer and returns the assistant to the caller when voicemail is detected: ```bash curl -L 'https://api.telnyx.com/v2/ai/assistants' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "name": "Support Agent", "instructions": "You are a support agent. Transfer calls to the support team when needed.", "model": "moonshotai/Kimi-K2.5", "tools": [ { "type": "transfer", "transfer": { "targets": [ { "name": "Support Team", "to": "+15551234567" } ], "from": "+15553456789", "voicemail_detection": { "detection_mode": "premium", "on_voicemail_detected": { "action": "stop_transfer" } } } } ] }' ``` ### Leave a message on voicemail This configuration delivers a TTS message to voicemail before returning to the caller: ```bash curl -L 'https://api.telnyx.com/v2/ai/assistants' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "name": "Support Agent", "instructions": "You are a support agent. Transfer calls to the support team when needed.", "model": "moonshotai/Kimi-K2.5", "tools": [ { "type": "transfer", "transfer": { "targets": [ { "name": "Support Team", "to": "+15551234567" } ], "from": "+15553456789", "voicemail_detection": { "detection_mode": "premium", "on_voicemail_detected": { "action": "leave_message_and_stop_transfer", "voicemail_message": { "type": "message", "message": "Hello, this is a message from Telnyx support. We attempted to reach you regarding your open ticket. We will try again shortly. Thank you." } } } } } ] }' ``` ### Leave warm transfer instructions on voicemail This configuration uses the warm transfer audio instructions as the voicemail message instead of custom TTS text: ```bash curl -L 'https://api.telnyx.com/v2/ai/assistants' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "name": "Support Agent", "instructions": "You are a support agent. Transfer calls to the support team when needed.", "model": "moonshotai/Kimi-K2.5", "tools": [ { "type": "transfer", "transfer": { "targets": [ { "name": "Support Team", "to": "+15551234567" } ], "from": "+15553456789", "voicemail_detection": { "detection_mode": "premium", "on_voicemail_detected": { "action": "leave_message_and_stop_transfer", "voicemail_message": { "type": "warm_transfer_instructions" } } } } } ] }' ``` --- ## Related resources - [Voice AI Assistant API Reference](/api-reference/assistants/create-an-assistant#transfertool) - Complete transfer tool API documentation including voicemail detection parameters. - [Agent Handoff](/docs/inference/ai-assistants/agent-handoff) - Configure AI-to-AI handoffs between specialized assistants. - [Answering Machine Detection](/docs/voice/programmable-voice/answering-machine-detection) - AMD for outbound calls using programmable voice. - [No-Code Voice Assistant](/docs/inference/ai-assistants/no-code-voice-assistant) - Get started with Voice AI Assistants in the Portal. --- ### Scheduled Events > Source: https://developers.telnyx.com/docs/inference/ai-assistants/scheduled-events.md Scheduled events let an AI Assistant kick off an outbound interaction at a fixed point in the future — either a phone call or an SMS message to a specific recipient. The platform stores the event, dispatches it at the scheduled time, and reports the outcome through callbacks. For phone calls, you can also configure automatic retries when the recipient does not engage (busy, no-answer, failed, canceled), so the assistant keeps trying without you having to re-create the event. Typical uses: - Outbound appointment reminders, follow-ups, and renewal calls. - After-hours callbacks scheduled by an inbound assistant. - Automated SMS confirmations queued from a workflow. - Re-attempting a missed call on a configurable interval until it connects. --- ## How it works Each scheduled event flows through a small state machine: 1. **`pending`** — the event has been created and is waiting for its scheduled time. 2. **`in_progress`** — the executor has dispatched the call or SMS to Telnyx. 3. **`completed`** — the recipient was reached (call answered and ended normally, or SMS delivered). 4. **`failed`** — the event could not be completed and no further attempts will be made. For phone calls, the terminal `CallStatus` returned by Telnyx decides whether the event is treated as a success or a retryable failure: | `CallStatus` | Treated as | | --- | --- | | `completed` | Success | | `busy` | Retryable client-side outcome | | `no-answer` | Retryable client-side outcome | | `failed` | Retryable client-side outcome | | `canceled` | Retryable client-side outcome | If retries are configured and budget remains, the event is re-queued with an updated `scheduled_at_fixed_datetime` and returned to `pending`. If the budget is exhausted, the event is marked `failed`. --- ## Configuring retries for external failures Two fields on the create-event request control the retry policy for recipient-side outcomes (busy, no-answer, failed, canceled). | Field | Type | Description | | --- | --- | --- | | `max_retries_client_errors` | integer, `0`–`10` | Number of additional dispatches allowed after the initial attempt. `0` (default) disables retries. | | `retry_interval_secs` | integer, `60`–`86400` | Seconds to wait between attempts. Required whenever `max_retries_client_errors > 0`. | Retries are **phone-call only**. Setting either field on an SMS event is rejected with a 400. ### Validation rules - `retry_interval_secs` must be in the range `60`–`86400` (1 minute to 24 hours). - `max_retries_client_errors` must be in the range `0`–`10`. - If `max_retries_client_errors > 0`, `retry_interval_secs` must also be set. - `retry_interval_secs` is only meaningful when `max_retries_client_errors > 0` — supplying one without the other returns a 400. - Both fields must be omitted (or `0` / `null`) for `sms_chat` events. ### How retries are scheduled When a phone-call event finishes with a retryable `CallStatus` and budget remains: - A new `CallAttempt` record is appended to the event's `call_attempts` array. - `scheduled_at_fixed_datetime` is advanced to **now + `retry_interval_secs`** (not the original time + interval — the clock starts when the previous attempt's terminal status is received). - `status` returns to `pending` and the executor will pick the event up again on its next sweep. When the budget is exhausted, the final attempt is recorded, `status` becomes `failed`, and an `errors` entry explains that the client-error retry budget was exhausted. ### Total attempts `max_retries_client_errors` counts retries **on top of** the initial dispatch. So: - `max_retries_client_errors: 0` → 1 total attempt (no retries). - `max_retries_client_errors: 3` → up to 4 total attempts. - `max_retries_client_errors: 10` → up to 11 total attempts. --- ## Inspecting attempt history Each phone-call event exposes a `call_attempts` array containing one entry per terminal dispatch: ```json { "scheduled_event_id": "8f3c…", "telnyx_conversation_channel": "phone_call", "status": "pending", "max_retries_client_errors": 3, "retry_interval_secs": 300, "call_attempts": [ { "attempt_number": 1, "attempted_at": "2026-05-07T14:30:02.812Z", "call_status": "no-answer", "call_duration": null, "telnyx_call_control_id": "v3:abc…" }, { "attempt_number": 2, "attempted_at": "2026-05-07T14:35:08.114Z", "call_status": "busy", "call_duration": null, "telnyx_call_control_id": "v3:def…" } ] } ``` The audit trail stays on a single row across the entire retry sequence — you do not need to query multiple events to reconstruct what happened. --- ## API reference All endpoints are scoped to a specific assistant. | Method | Path | Purpose | | --- | --- | --- | | `POST` | `/v2/ai/assistants/{assistant_id}/scheduled_events` | Create a scheduled event | | `GET` | `/v2/ai/assistants/{assistant_id}/scheduled_events` | List scheduled events (paginated, filterable) | | `GET` | `/v2/ai/assistants/{assistant_id}/scheduled_events/{event_id}` | Fetch a single event | | `DELETE` | `/v2/ai/assistants/{assistant_id}/scheduled_events/{event_id}` | Cancel a scheduled event | ### Create-event fields | Field | Required | Channels | Description | | --- | --- | --- | --- | | `telnyx_conversation_channel` | Yes | both | `phone_call` or `sms_chat`. | | `telnyx_end_user_target` | Yes | both | The recipient (phone number, E.164 or Sip URI). | | `telnyx_agent_target` | Yes | both | The number the assistant calls or sends from. | | `scheduled_at_fixed_datetime` | Yes | both | ISO-8601 timestamp; must be in the future. | | `text` | Yes for SMS | `sms_chat` | The message body. | | `conversation_metadata` | No | both | Free-form key/value metadata attached to the resulting conversation. | | `dynamic_variables` | No | both | Variables passed to the assistant for prompt interpolation. | | `max_retries_client_errors` | No | `phone_call` | Retry budget for recipient-side outcomes. | | `retry_interval_secs` | No | `phone_call` | Seconds between retries. | --- ## Examples ### Schedule a phone call with retries on busy / no-answer This event will be tried up to four times total (initial + 3 retries), with a five-minute gap between attempts. ```bash curl -L 'https://api.telnyx.com/v2/ai/assistants/{assistant_id}/scheduled_events' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "telnyx_conversation_channel": "phone_call", "telnyx_end_user_target": "+15551234567", "telnyx_agent_target": "+15553456789", "scheduled_at_fixed_datetime": "2026-05-08T15:00:00Z", "max_retries_client_errors": 3, "retry_interval_secs": 300, "conversation_metadata": { "campaign": "renewal-q2", "customer_id": "cust_42" } }' ``` ### Schedule a phone call without retries Omit the retry fields (or set `max_retries_client_errors` to `0`) for a single-attempt call. ```bash curl -L 'https://api.telnyx.com/v2/ai/assistants/{assistant_id}/scheduled_events' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "telnyx_conversation_channel": "phone_call", "telnyx_end_user_target": "+15551234567", "telnyx_agent_target": "+15553456789", "scheduled_at_fixed_datetime": "2026-05-08T15:00:00Z" }' ``` ### Schedule an SMS Retry fields do not apply to SMS events. ```bash curl -L 'https://api.telnyx.com/v2/ai/assistants/{assistant_id}/scheduled_events' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "telnyx_conversation_channel": "sms_chat", "telnyx_end_user_target": "+15551234567", "telnyx_agent_target": "+15553456789", "scheduled_at_fixed_datetime": "2026-05-08T15:00:00Z", "text": "Hi! This is a reminder about your appointment tomorrow at 10 AM." }' ``` ### Inspect retry history ```bash curl -L 'https://api.telnyx.com/v2/ai/assistants/{assistant_id}/scheduled_events/{event_id}' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' ``` The response includes `status`, `call_attempts`, and (once retries finish) the final `call_status` and `errors` array. ### Cancel a pending event Deleting a `pending` event prevents any further dispatches. Deletion has no effect on attempts that have already completed. ```bash curl -X DELETE 'https://api.telnyx.com/v2/ai/assistants/{assistant_id}/scheduled_events/{event_id}' \ -H 'Authorization: Bearer ' ``` --- ## Choosing a retry policy A few rules of thumb when picking values: - **Short intervals (60–300s)** suit time-sensitive flows where the recipient is likely to be near their phone — verification calls, urgent reminders. They burn the retry budget quickly, which is sometimes what you want. - **Medium intervals (5–30 minutes)** are a good default for outreach campaigns. Long enough that a busy recipient may finish their previous call; short enough that the event still completes within an hour. - **Long intervals (1–24 hours)** match daily-cadence flows: missed bill reminders, callback queues, abandoned-cart nudges. - **Total reach time** = `max_retries_client_errors × retry_interval_secs`. Make sure that span fits inside the time window when calling is acceptable for your use case (and your jurisdiction's calling-hour rules). - **Higher `max_retries_client_errors` is not always better** — at some point the recipient is unlikely to answer regardless of how many times you try. Three to five retries is usually plenty. --- ## Related resources - [Voice AI Assistant API Reference](/api-reference/assistants/create-an-assistant) — Full assistant configuration. - [No-Code Voice Assistant](/docs/inference/ai-assistants/no-code-voice-assistant) — Get started with Voice AI Assistants in the Portal. - [Dynamic Variables](/docs/inference/ai-assistants/dynamic-variables) — Pass per-event context into the assistant prompt. - [Voicemail Detection on Transfer](/docs/inference/ai-assistants/voicemail-detection-on-transfer) — Handle voicemail on transferred calls (a separate but related feature for live calls). --- ### Testing, Versions & Traffic Distribution > Source: https://developers.telnyx.com/docs/inference/ai-assistants/version-testing-traffic-distribution.md This guide walks you through testing your AI assistant before production deployment and managing live traffic distribution between different versions. You'll learn how to create tests, iterate on your assistant, and safely roll out changes using A/B testing. --- ## Creating Your First Assistant Start by creating a new assistant using a template to establish a baseline for testing. 1. Navigate to the [AI Assistants page](https://portal.telnyx.com/#/ai/assistants) 2. Click "Create Assistant" and select the "Weather Assistant" template 3. This template provides a good foundation with a standard greeting and weather functionality ![Create Weather Assistant](/assets/images/create-weather-assistant-template.png) Take note of the default greeting message - we'll be testing and modifying this later. --- ## Setting Up Your First Test Testing your assistant ensures it behaves correctly before going live with users. ### Creating a Test 1. Navigate to the [AI Tests page](https://portal.telnyx.com/#/ai/tests) 2. Click "Create Test" to set up your first test scenario ![Create AI Test](/assets/images/create-ai-test-setup.png) ### Configuring Test Criteria 3. Configure your test with the following: - **Test Name**: "Weather Assistant Greeting Test" - **Assistant**: Select your weather assistant - **Success Criteria**: Add criteria to validate the greeting message content and that temperature is described by the assistant. ![Test Criteria Configuration](/assets/images/test-criteria-setup.png) ### Running Your Test 1. Click "Run Test" to execute your test scenario 2. Monitor the test progress in real-time 3. Review the detailed results once the test completes ![Test Execution Run](/assets/images/ai-test-run.png) ![Test Execution Results](/assets/images/test-execution-results.png) ![Test Execution Results](/assets/images/test-conversation-history.png) The results will show whether your assistant met all the defined criteria, helping you identify any issues before deployment. You can also review the conversation itself. --- ## Creating Assistant Versions Now you'll create a new version of your assistant with modified behavior to demonstrate A/B testing. To make it obvious that the A/B test is working, we make two visibly distinct versions of the AI Assistant using the frontend widget feature. Versions are not limited to the frontend, though. You can make versions from any configuration on the assistant including updated tools, instructions, and more. ### Modifying the Assistant 1. Return to your weather assistant in the [AI Assistants page](https://portal.telnyx.com/#/ai/assistants) 2. Click the edit icon (pencil) next to your assistant 3. Make the following changes to create a visually distinct version: - **Enable the frontend widget**: Navigate to the Widget tab and click enable - **Widget Appearance**: Navigate back to the Widget tab and change the widget theme from dark mode to light mode in the appearance settings ![Enable Assistant Widget](/assets/images/enable-assistant-widget.png) ### Creating a New Version 1. After making your changes, click "Save as New Version" 2. Give your version a descriptive name: "Light Theme with New Greeting" 3. Add version notes describing the changes made ![Edit Assistant Widget](/assets/images/edit-assistant-widget.png) You now have two versions of your assistant: - **Version 1**: Original greeting with dark theme widget - **Version 2**: New greeting with light theme widget ![View New Assistant Version](/assets/images/view-assistant-versions.png) --- ## Production Traffic Distribution Once you've validated a new version, use version routing to control which live calls receive it. Traffic routing now uses ordered rules, similar to feature-flag targeting. Each rule has: - **If** conditions that match the end user target for the conversation - **Serve** behavior that sends matching calls to one version, or splits them across several versions Rules are evaluated from top to bottom. The first matching rule wins. If no target rule matches, the assistant serves the main version unless you configure a default rule. ### Open the traffic routing editor 1. Open your assistant from the [AI Assistants page](https://portal.telnyx.com/#/ai/assistants). 2. Go to the assistant's version or deployment controls. 3. Open **Traffic distribution** to configure routing rules. ### Add a target rule Use target rules when you want a specific end user, customer, or test endpoint to receive a version. The target value is the same value exposed to assistants as [`{{telnyx_end_user_target}}`](/docs/inference/ai-assistants/dynamic-variables#telnyx-system-variables): the phone number, SIP URI, or other identifier associated with the end user. This works for both call directions: - **Inbound calls**: the end user target is the caller's phone number, SIP URI, or identifier. - **Outbound calls**: the end user target is the destination the assistant calls. For example, you can route calls to your own number to a test version. 1. Click **Add rule**. 2. In the **If** section, choose **End user target**. 3. Select an operator: - **is one of** for exact targets, such as `+13125550123` or `sip:customer@example.com` - **is not one of** to exclude specific targets - **starts with** for prefixes, such as `+1312` or `sip:qa-` ![AI Assistant traffic routing new end user target rule](/assets/images/ai-assistant-routing-new-target-rule.png) If your Traffic distribution editor still shows **Origination number**, treat it as the end user target. The field is being renamed because the same routing behavior applies to inbound callers and outbound call destinations. 4. Enter one or more target values. You can separate values with commas or new lines. 5. In the **Serve** section, choose **Send all matched calls to one version** and select the version that matching calls should receive. ![AI Assistant single-version end user target rule](/assets/images/ai-assistant-routing-single-version-rule.png) Target rules can contain multiple conditions. Conditions in the same rule are AND-joined, so every condition must match for the rule to apply. If multiple rules could match, only the first matching rule is used. ![AI Assistant end user target rule with multiple conditions](/assets/images/ai-assistant-routing-multiple-conditions.png) ### Split matching traffic by percentage For gradual rollouts, set a rule's **Serve** behavior to **Split by percentage**. Add version slots and assign each one a percentage. The allocation bar shows how matching traffic is split. Percentages must add up to less than 100. Any remaining percentage routes to the main version, which gives you a built-in safety fallback during canary releases. ![AI Assistant weighted rollout configuration](/assets/images/ai-assistant-routing-weighted-rollout.png) For example, if a target rule sends 25% of matching calls to Version 2 and 25% to Version 3, the remaining 50% of matching calls continue to use the main version. ### Configure the default rule The default rule handles calls that do not match any target rule. By default, unmatched calls serve the main version. Use **Configure default** when you want unmatched calls to receive another version or a percentage split. Use **Reset to main** to remove the custom default and return all unmatched traffic to the main version. ### Save, reorder, or rollback - Drag target rules to change their priority. Rule order matters because the first match wins. - Click **Save** to apply the routing configuration. - Click **Rollback** to clear all routing rules and send traffic back to the main version. This setup allows you to: - Test new versions with internal phone numbers or SIP URIs before a broad rollout - Run percentage-based canaries for matching call segments - Keep the main version as the fallback for unmatched calls and remaining rollout percentage - Quickly rollback if issues arise - Promote a validated version to main when you're ready ### Testing Live Traffic Distribution To verify your routing is working correctly, make repeated test calls with end user targets that should match each rule: 1. For inbound testing, call the assistant from a phone number or SIP URI listed in a target rule and confirm the routed version answers. 2. For outbound testing, have the assistant call a target listed in a rule, such as your own phone number, and confirm you receive the test version. 3. Call from or to a target that should not match the target rules and confirm the default behavior applies. 4. For weighted rollouts, make enough calls to confirm that matching traffic is distributed according to your configured percentages. --- ## Automated Evaluation with Coval The manual testing and A/B traffic distribution described above work well for targeted checks and gradual rollouts. For automated evaluation at scale, Telnyx integrates with [Coval](https://www.coval.dev/) — a simulation and evaluation platform purpose-built for voice and chat agents. ### What Coval adds | Capability | How it complements built-in testing | |---|---| | **Scenario simulation** | Generate thousands of test conversations from a few seed cases, covering edge cases that are difficult to script manually. | | **CI/CD evaluations** | Automatically run your scenario library on every assistant change and block deployments that introduce regressions. | | **Production monitoring** | Log live calls, surface performance drops in real time, and replay transcripts or audio for debugging. | | **Built-in metrics** | Measure latency, accuracy, tool-call effectiveness, and instruction compliance without custom instrumentation. | ### Getting started with Coval 1. Set up the integration on the [Integrations tab](/docs/inference/ai-assistants/integrations#coval) of your assistant. 2. Create seed scenarios in Coval that reflect your most important conversation paths. 3. Run simulations to validate assistant behavior before deploying new versions. 4. Add Coval evaluation steps to your CI/CD pipeline to catch regressions automatically. For setup details and required credentials, see the [Coval integration guide](/docs/inference/ai-assistants/integrations#coval). --- ### Conversation Event Stream > Source: https://developers.telnyx.com/docs/inference/ai-assistants/conversation-event-stream.md Configure `websocket_settings` on an assistant and Telnyx opens a WebSocket to a server you host, once per conversation, streaming transcript, response, telephony and delegation events to it for the life of that conversation. Your server can push messages back to inject a turn. The assistant keeps the brain. This socket observes and injects — it does not replace the model, and it is not in the call path. If you want your own server to replace the LLM entirely, use [ConversationRelay](/docs/voice/programmable-voice/conversation-relay) instead. **Beta.** The set of events carried on this stream is expected to grow. New event types can appear at any time, so ignore types you do not recognize rather than treating them as errors. ## What you can build with it | Use case | How | |----------|-----| | Live agent-assist screen showing the transcript as it happens | Render `conversation.item.created` and `response.text.delta` | | Supervisor dashboard tracking calls in flight | Track `session.created` / `session.ended` and the `telnyx.call.*` events | | CRM logging a conversation as it unfolds, not after | Persist conversation items as they arrive | | Nudge the assistant mid-call from your own system | Send `conversation.item.create` with a `user` or `assistant` item | | Answer the assistant's lookups from your own backend | Pair with [delegation](/docs/inference/ai-assistants/delegation) in `client` mode | ## Configure the assistant Set `websocket_settings` when you create or update an assistant: ```bash curl -X POST https://api.telnyx.com/v2/ai/assistants/{assistant_id} \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "websocket_settings": { "enabled": true, "url": "wss://events.example.com/telnyx", "auth_ref": "my_websocket_token" } }' ``` | Field | Description | |-------|-------------| | `enabled` | Whether Telnyx opens a socket for each conversation. Defaults to `false`. | | `url` | The `ws://` or `wss://` endpoint Telnyx connects to. Required when `enabled` is `true`. | | `auth_ref` | Optional. An [integration secret](/docs/inference/ai-assistants/integrations) whose value Telnyx sends as `Authorization: Bearer ` on the upgrade request. | The `url` must be externally reachable. Localhost, private IP ranges and `.local` domains are rejected when you save the assistant, and an `auth_ref` that does not resolve to a secret in your organization is rejected too — so a typo fails at save time rather than silently producing a socket that never authenticates. Use `wss://` in production. `ws://` is accepted, but it sends your bearer token and your customers' transcripts in the clear. Telnyx resolves `auth_ref` on every connection attempt, so rotating the secret takes effect on the next reconnect without touching the assistant. ## How a session flows ``` Telnyx → You (connect wss://events.example.com/telnyx, Authorization: Bearer ...) Telnyx → You {"type":"session.created","session":{...}} Telnyx → You {"type":"telnyx.call.answered","call_control_id":"..."} Telnyx → You {"type":"conversation.item.created","item":{"role":"user", ...}} Telnyx → You {"type":"response.created","response":{"id":"resp_9c1f4a2b", ...}} Telnyx → You {"type":"response.text.delta","delta":"We are open ", ...} You → Telnyx {"type":"conversation.item.create","item":{"role":"user", ...}} Telnyx → You {"type":"telnyx.call.hangup","cause":"normal_clearing", ...} Telnyx → You {"type":"session.ended","reason":"normal","duration_sec":84, ...} ``` Frames are bare JSON objects discriminated by `type`, in the shape the OpenAI Realtime vocabulary uses — there is no envelope to unwrap. Telephony events are namespaced `telnyx.` so they can never collide with an event name a future OpenAI release introduces. `session.created` is always the first frame, and it is not sent until the conversation is ready. Anything you send before it is refused with an `error` frame, so wait for it before writing. ## Events Telnyx sends | Event | Meaning | |-------|---------| | `session.created` | The conversation is ready. Carries `conversation_id`, `assistant_id` and the call identifiers. | | `session.ended` | The conversation finished. Carries `reason`, `duration_sec` and `transfer_status`. | | `conversation.item.created` | A message was added to history — a caller utterance, an assistant reply, a tool call, or a tool result. | | `conversation.item.deleted` | An item was removed, for example an assistant turn discarded after a barge-in. | | `conversation.participant.added` / `.removed` | A party joined or left a [multi-participant conversation](/docs/inference/ai-assistants/multi-participant-calls). | | `response.created` | The assistant started a turn. `response.id` correlates the streaming frames that follow. | | `response.output_item.added` | An output item was opened on that turn. | | `response.content_part.added` | A content part was opened on the item. | | `response.text.delta` | An incremental chunk of the assistant's reply. | | `telnyx.call.answered` | The call was answered. | | `telnyx.call.hangup` | The call was hung up, with the SIP `cause`. | | `telnyx.call.dtmf.received` | The caller pressed a digit. | | `telnyx.call.machine.detection.ended` | Answering-machine detection finished, with the `result`. | | `telnyx.call.transfer.completed` / `.failed` | A transfer out of the conversation succeeded or failed. | | `telnyx.assistant.handoff` | The conversation was handed to a different assistant. | | `session.delegation.created` | The assistant is asking you to answer a [delegation](/docs/inference/ai-assistants/delegation). | | `error` | A frame you sent was invalid or cannot be accepted right now. | Only events that mean something to someone monitoring a conversation are forwarded — media-plumbing events like `call.playback.started` are deliberately left off the stream. For every payload, field by field, see the full reference under **Assistants API → Assistant Event Stream** in the sidebar. ### Rebuilding the assistant's reply A turn arrives as a `response.created`, then one `response.output_item.added` and `response.content_part.added` opening the text, then a run of `response.text.delta` frames. Concatenate the `delta` values sharing the same `item_id` and `content_index`: ```javascript const buffers = new Map(); function onEvent(event) { if (event.type === "response.text.delta") { const key = `${event.item_id}:${event.content_index}`; buffers.set(key, (buffers.get(key) ?? "") + event.delta); } if (event.type === "conversation.item.created" && event.item.role === "assistant") { // The completed item carries the final text — use it as the source of truth. console.log(event.item.content[0]?.text); } } ``` The `conversation.item.created` frame for the same turn carries the finished text, so you can either stream the deltas for responsiveness or just wait for the completed item. ## Injecting messages After `session.created`, send a `conversation.item.create` frame to put a message into the conversation: ```json { "type": "conversation.item.create", "item": { "type": "message", "role": "user", "content": [{ "type": "input_text", "text": "Actually, make that a delivery instead." }] } } ``` A `user` item triggers an answer exactly as a spoken turn does. An `assistant` item is recorded in the history without provoking a reply — useful for injecting context the model should know about but not respond to. There is no `response.create` verb: the assistant owns turn-taking, so a user item is the way to ask for a response. Only text content parts are read (`input_text`, `output_text`, `text`), and an item whose text is empty after trimming is rejected. ## Delivery guarantees This is a side channel, and it is built so that a customer endpoint can never affect a live call. That shapes what you can rely on: - **Events are dropped, not queued, while the socket is down.** Nothing is replayed on reconnect. Queueing is what turns a dead endpoint into unbounded memory on a live call, and a replayed backlog is of little use to a monitoring client anyway. If your endpoint goes away mid-conversation, you have a permanent hole in your view of it. - **No socket failure reaches the call.** A refused connection, a wedged endpoint, a crash — all of it costs you the side channel for the rest of the conversation and nothing more. The caller notices nothing. - **Telnyx reconnects with exponential backoff**, from 1 second up to 30 seconds. A connection has to survive 10 seconds before it counts as stable and resets the backoff, so an endpoint that accepts the upgrade and then closes does not become a reconnect loop. - **Some failures are terminal for the conversation.** Telnyx stops reconnecting after three consecutive failed sends, ten consecutive invalid inbound frames, an oversized frame, or an endpoint that resolves to an address Telnyx refuses to talk to. Because of the first point, treat this stream as a live view, not a system of record. If you need a guaranteed-complete transcript, read it from the Conversations API (under **Assistants API → Conversations API** in the sidebar) after the conversation ends. ## Limits | Limit | Value | On breach | |-------|-------|-----------| | Inbound frame size | 1 MiB | `error`, then the connection is closed | | Inbound frame rate | 10 frames per second | `error`; the frame is discarded | | Consecutive invalid frames | 10 | the connection is closed | | Binary frames | not supported | `error`; the frame is discarded | Rate-limited and refused frames are answered and discarded without counting toward the invalid-frame limit, so a burst from a legitimate client does not cost it the connection. ## Handoffs and private consults The socket stays bound to the assistant that opened the conversation. It is **not** re-dialled against the new assistant's `websocket_settings` on an [agent handoff](/docs/inference/ai-assistants/agent-handoff) — re-dialling would drop events across the switch — so a `telnyx.assistant.handoff` event is how you learn the conversation changed hands. A handoff that cascades through several assistants is reported once, naming the assistant it landed on. During the private consult phase of a [warm transfer](/docs/inference/ai-assistants/warm-transfer-acceptance) the socket goes silent in both directions. The consult is between the assistant and the transfer target, and its transcript is not part of the conversation you are monitoring. Frames you send during a consult are refused with a deliberately non-specific error. ## A minimal server ```javascript import { WebSocketServer } from "ws"; const wss = new WebSocketServer({ port: 8080 }); wss.on("connection", (socket, request) => { // Verify the credential Telnyx sends from your `auth_ref` secret. if (request.headers.authorization !== `Bearer ${process.env.TELNYX_WS_TOKEN}`) { socket.close(1008, "unauthorized"); return; } let conversationId = null; socket.on("message", (data) => { const event = JSON.parse(data.toString()); switch (event.type) { case "session.created": conversationId = event.session.conversation_id; console.log(`conversation ${conversationId} started`); break; case "conversation.item.created": console.log(`[${event.item.role}]`, event.item.content[0]?.text); break; case "session.ended": console.log(`ended: ${event.reason} after ${event.duration_sec}s`); break; default: // New event types can appear at any time — ignore what you don't know. break; } }); }); ``` ## Next steps - **Assistant Event Stream reference** — every frame, field by field, under **Assistants API → Assistant Event Stream** in the sidebar - [Delegation](/docs/inference/ai-assistants/delegation) — answer the assistant's lookups from your own backend over this socket - [Agent handoff](/docs/inference/ai-assistants/agent-handoff) — what `telnyx.assistant.handoff` is reporting --- ### Post-Conversation Processing > Source: https://developers.telnyx.com/docs/inference/ai-assistants/post-conversation-processing.md Post-conversation processing gives an AI Assistant one extra LLM turn **after the conversation has ended**. Instead of finishing when the caller hangs up, the assistant is invoked one final time to perform wrap-up work — sending a summary, filing a ticket, or updating an internal record — using the tools available to it. Typical uses: - Writing a call summary to storage or a record after the call ends. - Filing an internal ticket or logging the outcome of the call. - Any final data capture that should not keep the caller waiting on the line. --- ## How it works When the conversation ends, the platform injects a system message into the conversation ("The conversation has ended. You may now perform any post-conversation tasks.") and runs one additional LLM turn. The model sees the full conversation history plus its configured instructions, and can call the tools that are available post-conversation. If a tool call returns results that warrant follow-up, the model can call further tools in sequence, up to a small fixed number of iterations, before the turn ends. The model's tool calls and their results are added to the conversation's message history. The final text the model produces in this turn, if any, is **not** stored as a conversation message and is not delivered to anyone — the turn is for executing actions, not for producing user-facing output. Enable it with the `post_conversation_settings` object on the assistant: | Field | Type | Description | | --- | --- | --- | | `post_conversation_settings.enabled` | boolean | Whether the assistant runs a post-conversation turn after each conversation ends. Defaults to `false`. | ``` POST /v2/ai/assistants { "name": "repro-20260917-receptionist", "instructions": "After the call, file a summary ticket.", "post_conversation_settings": { "enabled": true }, "enabled_features": ["telephony"] } ``` Post-conversation processing is a beta feature and is voice-only: the post-conversation turn runs for phone-call conversations. --- ## Tool availability post-conversation Not every tool the assistant has during a call is available once the conversation has ended. **Available post-conversation:** - Webhook tools, function tools, and other HTTP-backed tools attached to the assistant. - Built-in data tools such as knowledge-base retrieval. **Not available post-conversation:** - **Integrations and MCP server tools** (Salesforce, Outlook, HubSpot, and every connector in the [Integrations catalog](/docs/inference/ai-assistants/integrations)). Integration tools are not offered to the model in the post-conversation turn at all, and a post-conversation model attempt to call one is refused. Integration actions in post-conversation instructions — for example "after the call, send a summary email via Outlook" — will not execute and will not silently succeed. - **Call-control tools** (`hangup`, `transfer`, `refer`, and other telephony actions). There is no live call to act on. - **Client-side tools**, which require a connected Voice SDK client. - `update_dynamic_variables` — the conversation is over, so updated variables have nowhere to land. If your assistant's post-conversation instructions rely on an integration tool, that step will not happen. Configure a webhook tool that reaches your own backend (or a function tool) to perform the action instead, or trigger the action from the [`call.conversation.ended`](/api-reference/callbacks/call-conversation-ended) webhook in your application. ### Designing post-conversation instructions Because the available toolset is narrower than during a call, write post-conversation work into the assistant instructions only in terms of tools that survive the filter above. If the assistant's only configured tools are excluded (for example, an assistant with nothing but `transfer` and `hangup`), the post-conversation turn has nothing to call — the turn runs and ends without performing any action. --- ## Observability Post-conversation processing inherits the assistant's data-retention setting. With data retention enabled, the tool calls the model makes post-conversation and their results appear in the conversation's message history alongside the trigger message. The turn itself is billed as LLM inference like any other model turn. --- ### Voice Outreach > Source: https://developers.telnyx.com/docs/inference/missions.md Your AI agent can do more than answer questions — it can go out into the world and get things done. With **AI Missions** and the **Telnyx Missions skill**, you can give your agent a task like _"Find catering companies in Chicago and call them to negotiate quotes for a corporate event"_ and watch it execute the entire workflow: research, phone calls, follow-ups, and a final summary. This guide walks you through the full lifecycle — from setting up your agent with the Missions skill to reviewing call results in the Telnyx Portal. --- ## What you'll need Before you start, make sure you have: 1. **An OpenClaw agent** — This is the AI agent that will execute your mission. If you don't have one yet, follow the [OpenClaw quickstart](https://docs.openclaw.ai) to get set up. 2. **A Telnyx account with an API key** — Your agent needs API access to create assistants, schedule calls, and track missions. [Create a Telnyx account](https://telnyx.com/sign-up) if you don't have one, then [generate an API key](https://portal.telnyx.com/#/app/api-keys). 3. **The Telnyx Missions skill installed** — Install it from [ClawHub](https://clawhub.ai/teamtelnyx/skills/telnyx-toolkit) and configure your `TELNYX_API_KEY` as an environment variable for your agent. That's it. The skill handles phone number selection and assignment automatically — your agent will pick an available number from your account when it needs to make calls. --- ## How it works Your agent uses several Telnyx APIs to orchestrate the work: - **Missions API** — creates a mission, plans steps, logs events, and tracks status - **Assistants API** — creates a voice assistant with custom prompts and schedules calls - **Numbers API** — finds an available phone number and assigns it to the assistant The Missions API tracks every step — what your agent planned, what it did, what succeeded, and what failed. You get a complete audit trail without having to monitor the agent in real time. --- ## Give your agent a mission Once your agent has the Telnyx Missions skill installed, you can describe tasks in natural language. The agent handles the rest. ### Example: Catering quote negotiation Here's a real example — asking your agent to find caterers and negotiate pricing: > _"Find catering companies in Chicago for a corporate event with 50 people. Call Lakefront Catering, Chicago Grill Co, and Prairie Table Events. Get a baseline quote from the first one, then use that to negotiate better rates with the others. Compare all quotes and recommend the best option."_ Your agent will: 1. **Create a mission** in the Telnyx Missions API to track the work 2. **Build an execution plan** with steps (assistant setup → baseline call → negotiation calls → comparison) 3. **Create a voice assistant** — a Telnyx AI Assistant configured as a "Catering Quote Negotiator" with a professional script tailored to your request 4. **Assign a phone number** from your account to the assistant for outbound calling 5. **Call the first caterer** to establish a baseline price 6. **Call remaining caterers** armed with the best quote so far, negotiating for better rates 7. **Monitor call completions** and capture conversation insights 8. **Summarize results** — comparing quotes and recommending the best option ### What the agent plans When the mission starts, your agent creates a structured plan. In the Portal, you can see each step and its status: ![Mission plan view showing completed steps: create assistant, call Lakefront Catering (baseline), call Chicago Grill Co (with best quote), call Prairie Table Events (with best quote), compare and recommend](/assets/images/missions-plan-view.png) Each step has a status (`pending`, `in_progress`, `completed`, `failed`) so you can see exactly where things stand at any point. In this example, the agent's plan was: 1. Create catering negotiation assistant with dynamic variable placeholders 2. Call Lakefront Catering (baseline, no leverage) 3. Call Chicago Grill Co (with best quote from call 1) 4. Call Prairie Table Events (with best quote so far) 5. Compare quotes and recommend best option Notice how the agent strategically ordered the calls — getting a baseline first, then using that price as leverage in subsequent negotiations. --- ## Monitor progress in the Portal You don't need to watch your agent work. The Telnyx Portal gives you a dashboard view of every mission. ### Mission overview Navigate to the **AI Missions** section in the [Telnyx Portal](https://portal.telnyx.com) to see all missions, their status, and when they were last updated. The dashboard shows recent runs with result summaries at a glance: ![AI Missions overview showing recent runs including Corporate Catering Quotes (succeeded), Restaurant Reservation Screening, and Weather IVR Sweep missions](/assets/images/missions-overview.png) ### Run details Click **View Run** to see the full run detail — status, timing, the original input, and structured result payload. This is where you'll find the agent's final analysis: ![Mission Run Detail showing succeeded status, result summary with quotes from 3 caterers, and structured result payload with per-person pricing and recommendation](/assets/images/missions-run-detail.png) The result payload includes structured data — per-person quotes, negotiation notes, conversation IDs linking back to call recordings, and the agent's recommendation. This data is also accessible via the API. ### Linked agents and conversation history Scrolling down, the Portal shows which Telnyx AI Assistants your agent created and linked to the mission run. Below that, you can see every conversation — including which caterer was called, the call channel, and timestamps. ### Conversation playback Click any conversation to open the full transcript with audio playback. Here's the Chicago Grill Co negotiation — watch the agent use the Lakefront's $65 quote as leverage to negotiate down to $62: ![Conversation detail showing the Catering Quote Negotiator negotiating with Chicago Grill Co — starting at $58/person without dessert, then negotiating to $62 all-inclusive using the $65 competitor quote as leverage](/assets/images/missions-conversation-detail.png) The conversation panel shows the full back-and-forth, audio waveform with playback controls, and latency metrics (STT, LLM, TTS) for each assistant turn. --- ## Review call results After calls complete, your agent captures conversation insights — structured summaries of what was discussed. ### Conversation insights For each completed call, the Missions skill can extract structured information using the [AI Insights API](/docs/inference/ai-insights/creating-insights) — pulling out key data points like pricing, availability, and terms into a structured format. ### The final summary When all calls are done, your agent compiles everything into a summary with recommendations. Here's what that looks like for our catering example: > **Mission Complete: Catering Quote Negotiation** > > Called 3 caterers for a 50-person corporate event: > > | Caterer | Quote (per person) | Negotiated? | Notes | > |---------|-------------------|-------------|-------| > | Lakefront Catering | $65 | Baseline | First call, established reference price | > | Chicago Grill Co | $62 | ✅ Yes, down from $68 | Beat the Lakefront quote when presented with it | > | Prairie Table Events | $65 | ❌ No match | Couldn't beat $65, matched it | > > **Recommendation:** Chicago Grill Co offers the best value at $62/person — $150 savings over the baseline for 50 guests. This summary is also stored as the mission's `result_summary` and `result_payload`, so it's permanently accessible via the API and Portal. --- ## How the voice calls work Behind the scenes, your agent uses the Telnyx AI Assistants platform to make calls. Here's what it sets up: ### Voice assistant creation The agent creates a Telnyx AI Assistant with: - A **system prompt** tailored to your specific request (e.g., _"You are calling catering companies to negotiate quotes for a 50-person corporate event..."_) - A **greeting** that opens the conversation naturally (e.g., _"Hi, I'm calling to inquire about catering for a corporate event. Do you have a moment?"_) - **Dynamic variables** that update between calls (e.g., injecting the current best quote before each negotiation call) - **Telephony features** enabled for voice calls ### Phone number assignment The agent finds an available phone number from your Telnyx account and assigns it to the assistant's voice connection. This is the caller ID that businesses will see. You need at least one phone number in your Telnyx account. If none are available, the agent will let you know and link you to the [number purchase page](https://portal.telnyx.com/#/app/numbers/search-numbers). ### Call scheduling Calls are scheduled during business hours to maximize the chance of reaching someone. The agent handles timezone awareness and won't schedule calls at 3 AM. ### Call monitoring After scheduling, the agent polls for call completion and retrieves: - **Call status** — completed, failed, no answer, busy - **Conversation transcript** — full text of what was said - **Conversation insights** — structured extraction of key data points (quotes, availability, terms) --- ## Tips for effective missions ### Be specific in your request The more detail you give, the better your agent performs: ``` ❌ "Find me a caterer" ✅ "Find 3 catering companies in downtown Chicago for a 50-person corporate lunch. Call each one, get per-person pricing, ask about dietary accommodation options, and negotiate using the best quote you get as leverage." ``` ### Let the agent plan first Your agent will create a plan before executing. If you want to review the plan before calls go out, say so: > _"Find caterers and create a plan, but wait for my approval before making any calls."_ ### Check the Portal for real-time status You don't need to keep chatting with your agent. The Portal shows live progress — check back when you get a notification that the mission is complete. ### Start small Try a mission with 2-3 calls first to see how it works. You can always scale up once you're comfortable with the workflow. --- ## What's next - **[AI Missions API Reference](/docs/inference/missions)** — Full API documentation for creating and managing missions programmatically - **[AI Assistants](/docs/inference/ai-assistants/async-tools)** — Learn about async tools and advanced assistant configurations - **[AI Insights](/docs/inference/ai-insights/creating-insights)** — Extract structured data from conversations - **[Telnyx Missions Skill on ClawHub](https://clawhub.ai/teamtelnyx/skills/telnyx-toolkit)** — Install the skill for your OpenClaw agent --- ### Creating insights > Source: https://developers.telnyx.com/docs/inference/ai-insights/creating-insights.md This guide walks you through creating individual [AI Insights](https://portal.telnyx.com/#/ai/insights) in the Mission Control Portal. You'll learn how to define analysis instructions, use variables, and configure both normal and structured insights. ## Prerequisites - Access to the Telnyx Mission Control Portal. - At least one AI Assistant configured (recommended for testing). ## Accessing AI Insights 1. Log in to the [Mission Control Portal](https://portal.telnyx.com). 2. Navigate to **AI, Storage and Compute** > **AI Insights**. 3. You'll see a list of existing insights with their IDs, names, instructions, and creation dates. ![AI Insights Main Page](/assets/images/ai-insights-main-page.png) ## Creating an insight Insights return free-form text responses based on your instructions. ### Step 1: Open the create dialog Click the **Create Insight** button in the top-right corner of the AI Insights page. ### Step 2: Configure basic settings In the Create Insight modal, you'll see: ![Create Insight Modal](/assets/images/create-insight-modal.png) 1. **Name** (Required) - A descriptive identifier for your insight. - Example: "Conversation Summary", "Customer Sentiment", "Issue Classification". 2. **Instructions** - Detailed prompt describing what to analyze and extract. - Be specific about what information you want. - Include output format expectations. - Reference conversation elements (transcript, metadata, etc.). ### Step 3: Write effective instructions Good instructions are clear, specific, and actionable. Here are some examples: **Example 1: Conversation summary** ``` Summarize the conversation for use as future context. Include: - Key facts mentioned. - Decisions made. - User preferences expressed. - Action items or follow-ups needed. Keep the summary concise (2-3 sentences) and focus on information that would be useful in future conversations with this user. ``` **Example 2: Sentiment analysis** ``` Measure the positivity & negativity of the call and rate it from 1-5 in ascending order. Positivity: How positive, satisfied, or happy was the customer? (1=very negative, 5=very positive) Negativity: How frustrated, angry, or dissatisfied was the customer? (1=no negativity, 5=very negative) Provide your ratings and a brief explanation of why you assigned those scores. ``` **Example 3: Issue categorization** ``` Analyze the conversation and identify the primary issue or request. Categorize it into one of the following: - Technical Support. - Billing Question. - Feature Request. - General Inquiry. - Complaint. - Other. Also provide a brief description of the specific issue within that category. ``` ### Step 4: Use variables (optional) You can include dynamic variables in your instructions to provide context about the specific conversation. Click the **Add a variable** dropdown to see available options. **System variables:** - `{{telnyx_current_time}}` - Date and time of the conversation. - `{{telnyx_conversation_channel}}` - Channel type (phone_call, web_call, sms_chat). - `{{telnyx_agent_target}}` - Assistant's phone number or identifier. - `{{telnyx_end_user_target}}` - User's phone number or identifier. **Example with variables:** ``` Analyze this {{telnyx_conversation_channel}} conversation from {{telnyx_current_time}}. Identify if the user at {{telnyx_end_user_target}} expressed interest in any of our products or services. If so, list the products mentioned and their level of interest (high/medium/low). ``` You can also reference [custom dynamic variables](https://developers.telnyx.com/docs/inference/ai-assistants/dynamic-variables) that you've configured for your assistant. ### Step 5: Save the insight 1. Review your configuration. 2. Click **Save**. 3. The insight will appear in your insights list. ## Managing insights ### Editing an insight 1. Find the insight in the list. 2. Click the **edit icon** (pencil) on the right side of the row. 3. Modify the name or instructions. 4. Click **Save**. ### Copying an insight ID Each insight has a unique ID that you can use with the Memory API or for programmatic access: 1. Click the **copy icon** next to the insight ID. 2. The ID will be copied to your clipboard. 3. Use this ID in API calls or memory queries. Example ID format: `cfcc865c-d3d4-4823-8a4b-f0df57d9f56f` ## Configuring webhook delivery You can configure webhook URLs to automatically receive insight results when conversations complete. ### Via Insight Groups Set a webhook URL when creating or editing an Insight Group: 1. Navigate to [AI Insight Groups](https://portal.telnyx.com/#/ai/insights-groups). 2. Click **Create Insight Group** or edit an existing group. 3. Enter your webhook URL in the **Webhook URL** field. 4. Save the group. ![Create Insight Group Modal](/assets/images/create-insight-group-modal.png) **Example:** ``` Webhook URL: https://api.mycompany.com/webhooks/ai-insights ``` All assistants using this group will send insights to this URL. Editing a group's webhook URL changes delivery for every assistant using that group. ### Per-insight webhook There is no per-assistant webhook override: an assistant's insights are always delivered to the webhook URL configured on its Insight Group (or on a single insight when a run references that insight directly). If you need different destinations for different assistants, create a separate Insight Group per destination and assign each assistant to the matching group. ## Best practices ### Writing clear instructions 1. **Be Specific** - Clearly state what you want to extract. - ❌ "Analyze the call". - ✅ "Identify the customer's main complaint and rate the urgency from 1-5". 2. **Define Output Format** - Specify how you want the response structured. - ❌ "Tell me about sentiment". - ✅ "Rate sentiment from 1-10 and provide a one-sentence explanation". 3. **Provide Context** - Explain why the information matters. - ❌ "List products mentioned". - ✅ "List products the customer showed interest in purchasing, noting their budget concerns". 4. **Use Examples** - Show the format you expect ``` Categorize the issue as one of: - Billing (e.g., incorrect charges, payment questions). - Technical (e.g., service not working, setup help). - Account (e.g., upgrades, cancellations). ``` ### Testing your insights 1. **Start with Test Conversations** - Try your insight on a few sample conversations first. 2. **Review Results** - Check if the extracted information matches your expectations. 3. **Refine Instructions** - Adjust wording based on the results. 4. **Validate Accuracy** - Ensure consistent, reliable extraction across different conversation types. ### Naming conventions Use clear, descriptive names that indicate: - **What** is being analyzed: "Customer Sentiment", "Product Interest", "Issue Type". - **Why** it matters: "Escalation Needed", "Follow-up Required", "Compliance Check". - **Scope**: "Healthcare Compliance", "Sales Qualification", "Support Quality". ## Next steps - **[Use Structured Data](https://developers.telnyx.com/docs/inference/ai-insights/structured-insights)** - Create insights with consistent JSON schemas. - **[Create Insight Groups](https://developers.telnyx.com/docs/inference/ai-insights/insight-groups)** - Organize insights for your assistants. - **[Explore Use Cases](https://developers.telnyx.com/docs/inference/ai-insights/use-cases)** - Industry-specific examples. ## Related resources - [Dynamic Variables](https://developers.telnyx.com/docs/inference/ai-assistants/dynamic-variables) - Custom variables for your insights. - [Memory](https://developers.telnyx.com/docs/inference/ai-assistants/memory) - Using insight IDs in memory queries. - [Voice Assistant Configuration](https://developers.telnyx.com/docs/inference/ai-assistants/no-code-voice-assistant#insights) - Assigning insights to assistants. --- ### Use cases > Source: https://developers.telnyx.com/docs/inference/ai-insights/use-cases.md This guide provides complete, industry-specific examples of AI Insights implementations. Each use case includes insight configurations, group organization, webhook integration, and best practices. ## Healthcare ### Use case: Patient call quality & compliance **Business Need:** Ensure regulatory compliance (HIPAA) while monitoring patient interaction quality and identifying follow-up needs. ### Insight configurations **1. HIPAA compliance check** (structured) ``` Name: HIPAA Compliance Verification Instructions: Verify HIPAA compliance requirements were met during this call. Parameters: 1. disclosures_made - Type: boolean - Required: Yes - Description: Were required privacy disclosures made at the start of the call? 2. consent_obtained - Type: boolean - Required: Yes - Description: Was patient consent obtained before discussing PHI? 3. phi_handled_properly - Type: boolean - Required: Yes - Description: Was Protected Health Information handled according to HIPAA guidelines? 4. violations - Type: array - Required: No - Description: List any potential HIPAA violations or concerns 5. compliant - Type: boolean - Required: Yes - Description: Overall assessment: was this call HIPAA compliant? ``` **2. Patient care quality** (structured) ``` Name: Patient Care Assessment Instructions: Evaluate the quality of patient care and interaction. Parameters: 1. empathy_score - Type: number - Required: Yes - Description: Rate the assistant's empathy from 1-5 (1=robotic, 5=highly empathetic) 2. clarity_score - Type: number - Required: Yes - Description: Rate explanation clarity from 1-5 (1=confusing, 5=very clear) 3. questions_answered - Type: boolean - Required: Yes - Description: Were all patient questions adequately answered? 4. follow_up_needed - Type: boolean - Required: Yes - Description: Does this patient require follow-up contact? 5. urgency_level - Type: string - Required: Yes - Description: Urgency classification: "routine", "moderate", "urgent", "emergency" ``` **3. Appointment summary** (unstructured) ``` Name: Call Summary for Patient Record Instructions: Create a concise summary for the patient's medical record. Include: - Reason for call. - Symptoms or concerns discussed. - Instructions provided. - Any appointments scheduled. - Follow-up requirements. Keep summary professional and factual for medical record inclusion. ``` ### Insight group configuration ``` Name: Healthcare Patient Calls Webhook URL: https://ehr.healthcorp.com/api/webhooks/ai-insights Insights: - HIPAA Compliance Verification - Patient Care Assessment - Call Summary for Patient Record ``` ### Webhook integration example ```python @app.route('/api/webhooks/ai-insights', methods=['POST']) def handle_patient_call_insights(): data = request.json conversation_id = data['conversation_id'] insights = {i['insight_name']: i['result'] for i in data['insights']} # Check compliance compliance = insights.get('HIPAA Compliance Verification', {}) if not compliance.get('compliant'): # Alert compliance team alert_compliance_team( conversation_id=conversation_id, violations=compliance.get('violations', []) ) # Check if urgent follow-up needed care = insights.get('Patient Care Assessment', {}) if care.get('urgency_level') in ['urgent', 'emergency']: # Create urgent task create_urgent_followup( conversation_id=conversation_id, urgency=care['urgency_level'] ) # Store summary in EHR summary = insights.get('Call Summary for Patient Record') if summary: ehr.add_patient_note( conversation_id=conversation_id, note=summary, note_type='ai_assistant_call' ) return jsonify({'status': 'processed'}), 200 ``` ### Expected results ```json { "insights": [ { "insight_name": "HIPAA Compliance Verification", "result": { "disclosures_made": true, "consent_obtained": true, "phi_handled_properly": true, "violations": [], "compliant": true } }, { "insight_name": "Patient Care Assessment", "result": { "empathy_score": 5, "clarity_score": 4, "questions_answered": true, "follow_up_needed": true, "urgency_level": "routine" } }, { "insight_name": "Call Summary for Patient Record", "result": "Patient called regarding follow-up on recent lab results. Explained test findings indicating normal thyroid function. Patient had questions about medication dosing which were addressed. Scheduled 6-month follow-up appointment. No immediate concerns noted." } ] } ``` --- ## Customer support ### Use case: Support ticket automation & quality monitoring **Business Need:** Automatically categorize and route support calls, monitor agent performance, and ensure customer satisfaction. ### Insight configurations **1. Ticket classification** (structured) ``` Name: Support Ticket Classifier Instructions: Classify and triage the support request for ticket creation. Parameters: 1. issue_type - Type: string - Required: Yes - Description: Primary issue category: "technical", "billing", "account", "feature_request", "bug_report", "general" 2. product_area - Type: string - Required: No - Description: Which product or service is affected? 3. priority - Type: string - Required: Yes - Description: Priority level: "critical", "high", "medium", "low" 4. resolved - Type: boolean - Required: Yes - Description: Was the issue completely resolved during this call? 5. resolution_time_minutes - Type: number - Required: No - Description: If resolved, approximately how many minutes did resolution take? 6. tags - Type: array - Required: No - Description: Relevant tags (e.g., "password_reset", "refund_request", "api_error") ``` **2. Customer satisfaction** (structured) ``` Name: Customer Satisfaction Score Instructions: Assess customer satisfaction from the conversation. Parameters: 1. csat_score - Type: number - Required: Yes - Description: Customer satisfaction from 1-5 (1=very dissatisfied, 5=very satisfied) 2. sentiment - Type: string - Required: Yes - Description: Overall sentiment: "positive", "neutral", "negative" 3. frustration_level - Type: number - Required: Yes - Description: Customer frustration from 1-5 (1=not frustrated, 5=extremely frustrated) 4. likely_to_churn - Type: boolean - Required: Yes - Description: Based on the conversation, is the customer likely to cancel service? 5. positive_feedback - Type: array - Required: No - Description: Specific things the customer praised or appreciated ``` **3. Agent performance** (structured) ``` Name: Agent Quality Assessment Instructions: Evaluate the AI assistant's performance in handling this support request. Parameters: 1. professionalism_score - Type: number - Required: Yes - Description: Professional tone and communication from 1-5 2. problem_solving_score - Type: number - Required: Yes - Description: Effectiveness in solving the issue from 1-5 3. efficiency_score - Type: number - Required: Yes - Description: How efficiently was the issue handled from 1-5 4. escalation_needed - Type: boolean - Required: Yes - Description: Should this have been escalated to human agent? 5. improvement_areas - Type: array - Required: No - Description: Areas where the assistant could improve ``` ### Insight group configuration ``` Name: Support Call Analytics Webhook URL: https://support.mycompany.com/api/webhooks/ai-insights Insights: - Support Ticket Classifier - Customer Satisfaction Score - Agent Quality Assessment ``` ### Webhook integration example ```javascript app.post('/api/webhooks/ai-insights', async (req, res) => { res.status(200).send('OK'); const { conversation_id, insights } = req.body; const insightMap = Object.fromEntries( insights.map(i => [i.insight_name, i.result]) ); const classification = insightMap['Support Ticket Classifier']; const satisfaction = insightMap['Customer Satisfaction Score']; const performance = insightMap['Agent Quality Assessment']; // Create ticket if unresolved if (!classification.resolved) { await ticketing.create({ conversation_id, type: classification.issue_type, priority: classification.priority, product: classification.product_area, tags: classification.tags, customer_sentiment: satisfaction.sentiment }); } // Alert on potential churn if (satisfaction.likely_to_churn) { await alerts.send({ type: 'churn_risk', conversation_id, csat_score: satisfaction.csat_score, frustration: satisfaction.frustration_level }); } // Log agent performance metrics await analytics.track('agent_performance', { conversation_id, professionalism: performance.professionalism_score, problem_solving: performance.problem_solving_score, efficiency: performance.efficiency_score, escalation_needed: performance.escalation_needed }); }); ``` --- ## Sales ### Use case: Lead qualification & pipeline management **Business Need:** Automatically qualify inbound leads, score opportunities, and route to appropriate sales representatives. ### Insight configurations **1. Lead qualification** (structured) ``` Name: Sales Lead Qualifier Instructions: Assess the quality and readiness of this sales lead. Parameters: 1. budget_discussed - Type: boolean - Required: Yes - Description: Did the prospect discuss budget or pricing? 2. budget_range - Type: string - Required: No - Description: Budget range if mentioned: "under_10k", "10k_50k", "50k_100k", "over_100k", "not_disclosed" 3. decision_timeframe - Type: string - Required: Yes - Description: When will they decide: "immediate", "this_month", "this_quarter", "next_quarter", "exploring", "unknown" 4. authority_level - Type: string - Required: Yes - Description: Decision-making authority: "decision_maker", "influencer", "end_user", "researcher", "unknown" 5. pain_points - Type: array - Required: Yes - Description: Specific problems or needs mentioned by the prospect 6. competitor_mentions - Type: array - Required: No - Description: Names of competing solutions mentioned 7. lead_score - Type: number - Required: Yes - Description: Overall qualification score from 1-10 based on BANT criteria 8. recommended_action - Type: string - Required: Yes - Description: Next step: "immediate_followup", "schedule_demo", "send_proposal", "nurture", "disqualify" ``` **2. Product interest** (structured) ``` Name: Product Interest Tracking Instructions: Identify which products or features the prospect showed interest in. Parameters: 1. products_mentioned - Type: array - Required: Yes - Description: List of products discussed during the call 2. primary_interest - Type: string - Required: Yes - Description: The product they seemed most interested in 3. feature_priorities - Type: array - Required: No - Description: Specific features they asked about or emphasized 4. use_case - Type: string - Required: Yes - Description: Brief description of their intended use case 5. integration_requirements - Type: array - Required: No - Description: Systems they need to integrate with ``` **3. Objections & concerns** (unstructured) ``` Name: Sales Objections Analysis Instructions: Identify any objections, concerns, or hesitations the prospect expressed. Include: - Price/budget concerns. - Feature gaps or limitations. - Competitive comparisons. - Implementation concerns. - Trust or credibility questions. Provide actionable insights for sales team to address these objections. ``` ### Insight group configuration ``` Name: Sales Call Intelligence Webhook URL: https://crm.salesteam.com/api/webhooks/leads Insights: - Sales Lead Qualifier - Product Interest Tracking - Sales Objections Analysis ``` ### Webhook integration example ```javascript app.post('/api/webhooks/leads', async (req, res) => { res.status(200).send('OK'); const { conversation_id, insights } = req.body; const insightMap = Object.fromEntries( insights.map(i => [i.insight_name, i.result]) ); const qualification = insightMap['Sales Lead Qualifier']; const interest = insightMap['Product Interest Tracking']; const objections = insightMap['Sales Objections Analysis']; // Create or update lead in CRM const lead = await crm.leads.upsert({ source: 'ai_assistant', conversation_id, score: qualification.lead_score, budget_range: qualification.budget_range, timeframe: qualification.decision_timeframe, authority: qualification.authority_level, pain_points: qualification.pain_points, primary_interest: interest.primary_interest, products: interest.products_mentioned, use_case: interest.use_case, objections: objections }); // Route high-quality leads immediately if (qualification.lead_score >= 8) { const rep = await assignSalesRep(interest.primary_interest); await crm.tasks.create({ assigned_to: rep, lead_id: lead.id, type: 'immediate_followup', priority: 'high', notes: `High-quality lead (score: ${qualification.lead_score}). ${qualification.recommended_action}` }); // Send Slack notification await slack.notify({ channel: '#sales', message: `🔥 Hot lead! Score: ${qualification.lead_score}/10. Assigned to ${rep.name}.` }); } // Add to appropriate nurture campaign if (qualification.recommended_action === 'nurture') { await marketing.addToCampaign(lead.email, { campaign: `nurture_${qualification.decision_timeframe}`, interests: interest.products_mentioned }); } }); ``` --- ## E-commerce ### Use case: Customer service & order management **Business Need:** Handle order inquiries, identify dissatisfaction early, and capture product feedback. ### Insight configurations **1. Order inquiry classification** (structured) ``` Name: Order Support Classifier Instructions: Classify the type of order-related inquiry or issue. Parameters: 1. inquiry_type - Type: string - Required: Yes - Description: Type of inquiry: "order_status", "return_request", "product_question", "shipping_issue", "payment_problem", "cancel_order", "modify_order" 2. order_number - Type: string - Required: No - Description: Order number if mentioned 3. urgency - Type: string - Required: Yes - Description: Urgency level: "urgent", "moderate", "low" 4. resolution_provided - Type: boolean - Required: Yes - Description: Was a resolution or answer provided? 5. compensation_offered - Type: boolean - Required: No - Description: Was any compensation (refund, discount, credit) offered? 6. next_action - Type: string - Required: Yes - Description: Required next step: "none", "process_return", "issue_refund", "escalate", "ship_replacement", "cancel_order" ``` **2. Customer sentiment** (structured) ``` Name: E-commerce Customer Sentiment Instructions: Gauge customer satisfaction with their shopping experience. Parameters: 1. satisfaction_level - Type: string - Required: Yes - Description: Overall satisfaction: "very_satisfied", "satisfied", "neutral", "dissatisfied", "very_dissatisfied" 2. likely_to_repurchase - Type: boolean - Required: Yes - Description: Based on the conversation, is customer likely to purchase again? 3. likely_to_recommend - Type: number - Required: Yes - Description: NPS score from 0-10: How likely to recommend to a friend? 4. complaint_severity - Type: string - Required: No - Description: If complaining, severity: "minor", "moderate", "major" 5. praise_areas - Type: array - Required: No - Description: What aspects did they appreciate or praise? ``` **3. Product feedback** (unstructured) ``` Name: Product Feedback Collection Instructions: Capture any product feedback, suggestions, or quality issues mentioned. Include: - Specific products mentioned. - Positive feedback or features they loved. - Negative feedback or problems encountered. - Feature requests or suggestions. - Quality concerns. Focus on actionable insights for product and marketing teams. ``` ### Insight group configuration ``` Name: E-commerce Customer Insights Webhook URL: https://api.shop.com/webhooks/customer-insights Insights: - Order Support Classifier - E-commerce Customer Sentiment - Product Feedback Collection ``` ### Webhook integration example ```javascript app.post('/webhooks/customer-insights', async (req, res) => { res.status(200).send('OK'); const { conversation_id, insights } = req.body; const insightMap = Object.fromEntries( insights.map(i => [i.insight_name, i.result]) ); const orderClassification = insightMap['Order Support Classifier']; const sentiment = insightMap['E-commerce Customer Sentiment']; const feedback = insightMap['Product Feedback Collection']; // Process order actions switch (orderClassification.next_action) { case 'process_return': await orders.initiateReturn(orderClassification.order_number); break; case 'issue_refund': await orders.processRefund(orderClassification.order_number); break; case 'cancel_order': await orders.cancel(orderClassification.order_number); break; case 'ship_replacement': await orders.shipReplacement(orderClassification.order_number); break; } // Alert on poor experiences if (sentiment.satisfaction_level === 'very_dissatisfied') { await alerts.send({ type: 'customer_dissatisfaction', conversation_id, order_number: orderClassification.order_number, severity: orderClassification.complaint_severity, nps_score: sentiment.likely_to_recommend }); } // Collect product feedback if (feedback) { await productFeedback.create({ conversation_id, feedback: feedback, source: 'ai_assistant', sentiment: sentiment.satisfaction_level }); } // Segment for marketing if (!sentiment.likely_to_repurchase) { await marketing.addToWinBackCampaign({ conversation_id, reason: orderClassification.inquiry_type }); } }); ``` --- ## Financial services ### Use case: Fraud detection & compliance **Business Need:** Identify potential fraud, ensure regulatory compliance, and categorize financial inquiries. ### Insight configurations **1. Fraud risk assessment** (structured) ``` Name: Fraud Risk Detector Instructions: Assess potential fraud risk indicators in this conversation. Parameters: 1. risk_level - Type: string - Required: Yes - Description: Overall risk: "no_risk", "low", "medium", "high", "critical" 2. risk_indicators - Type: array - Required: No - Description: Specific fraud indicators detected (e.g., "urgency_pressure", "unusual_request", "inconsistent_info", "social_engineering_attempt") 3. verification_requested - Type: boolean - Required: Yes - Description: Did the caller request account changes or sensitive actions? 4. authentication_status - Type: string - Required: Yes - Description: Authentication status: "verified", "partially_verified", "not_verified", "failed_verification" 5. recommended_action - Type: string - Required: Yes - Description: Recommended action: "approve", "additional_verification", "escalate_to_fraud_team", "block_immediately" 6. confidence_score - Type: number - Required: Yes - Description: Confidence in risk assessment from 0-100% ``` **2. Compliance check** (structured) ``` Name: Financial Compliance Verification Instructions: Verify regulatory compliance requirements for financial services. Parameters: 1. required_disclosures_made - Type: boolean - Required: Yes - Description: Were all required financial disclosures made? 2. customer_consent_obtained - Type: boolean - Required: No - Description: Was consent obtained for account changes or data sharing? 3. pii_handled_properly - Type: boolean - Required: Yes - Description: Was Personally Identifiable Information handled securely? 4. recording_disclosure - Type: boolean - Required: Yes - Description: Was call recording disclosure made? 5. compliance_violations - Type: array - Required: No - Description: Any potential compliance violations or concerns 6. compliant - Type: boolean - Required: Yes - Description: Overall compliance status ``` **3. Inquiry categorization** (structured) ``` Name: Financial Inquiry Classifier Instructions: Categorize the type of financial inquiry or request. Parameters: 1. inquiry_category - Type: string - Required: Yes - Description: Primary category: "account_balance", "transaction_dispute", "card_activation", "fraud_report", "account_opening", "loan_inquiry", "investment_advice", "general_question" 2. account_type - Type: string - Required: No - Description: Account type involved: "checking", "savings", "credit_card", "loan", "investment", "multiple" 3. transaction_amount - Type: number - Required: No - Description: Dollar amount if transaction-related inquiry 4. resolution_status - Type: string - Required: Yes - Description: Resolution status: "resolved", "pending", "escalated", "requires_callback" 5. callback_required - Type: boolean - Required: Yes - Description: Does customer need a callback from specialist? ``` ### Insight group configuration ``` Name: Financial Services Security & Compliance Webhook URL: https://api.bank.com/webhooks/insights Insights: - Fraud Risk Detector - Financial Compliance Verification - Financial Inquiry Classifier ``` ### Webhook integration example ```python @app.route('/webhooks/insights', methods=['POST']) def handle_financial_insights(): data = request.json conversation_id = data['conversation_id'] insights = {i['insight_name']: i['result'] for i in data['insights']} fraud_risk = insights.get('Fraud Risk Detector', {}) compliance = insights.get('Financial Compliance Verification', {}) inquiry = insights.get('Financial Inquiry Classifier', {}) # Handle high fraud risk immediately if fraud_risk.get('risk_level') in ['high', 'critical']: fraud_team.alert({ 'conversation_id': conversation_id, 'risk_level': fraud_risk['risk_level'], 'indicators': fraud_risk['risk_indicators'], 'recommended_action': fraud_risk['recommended_action'], 'confidence': fraud_risk['confidence_score'] }) # Block if critical if fraud_risk['risk_level'] == 'critical': security.block_account_temporarily(conversation_id) # Handle compliance violations if not compliance.get('compliant'): compliance_team.alert({ 'conversation_id': conversation_id, 'violations': compliance['compliance_violations'], 'severity': 'high' }) # Log for audit audit_log.create({ 'event': 'compliance_violation', 'conversation_id': conversation_id, 'details': compliance }) # Route inquiry to appropriate team if inquiry.get('callback_required'): routing.create_callback({ 'conversation_id': conversation_id, 'category': inquiry['inquiry_category'], 'account_type': inquiry.get('account_type'), 'priority': 'high' if fraud_risk['risk_level'] != 'no_risk' else 'normal' }) return jsonify({'status': 'processed'}), 200 ``` --- ## Best practices across use cases ### 1. Start with core insights Begin with 2-3 essential insights: - ✅ Sentiment/satisfaction. - ✅ Primary classification. - ✅ Action required. Add more specialized insights once core metrics are validated. ### 2. Balance structured and unstructured Use structured insights for: - Metrics and scores. - Categories and classifications. - Boolean flags. Use unstructured insights for: - Summaries and context. - Open-ended feedback. - Nuanced analysis. ### 3. Configure appropriate webhooks - **Real-time action required** → Webhook to operational system. - **Analytics only** → Webhook to analytics platform. - **Manual review** → No webhook, use Portal. ### 4. Test with real conversations Before production: 1. Test insights on 10-20 real conversations. 2. Review accuracy of classifications. 3. Validate webhook integration. 4. Adjust instructions based on results. ### 5. Monitor and iterate Track these metrics: - Insight accuracy. - Webhook delivery success rate. - False positive/negative rates for classifications. - Time to process insights. Refine instructions monthly based on performance. ### 6. Secure sensitive data For regulated industries: - Use HTTPS for all webhooks. - Implement proper authentication. - Encrypt data at rest. - Maintain audit logs. - Follow industry-specific compliance requirements. ## Related resources - [Voice Assistant Quickstart](https://developers.telnyx.com/docs/inference/ai-assistants/no-code-voice-assistant) - Set up your AI assistant. - [Memory API](https://developers.telnyx.com/docs/inference/ai-assistants/memory) - Use insights in conversation memory. - [API Reference](/api-reference/assistants/list-assistants) - Programmatic insights access. --- ### Structured insights > Source: https://developers.telnyx.com/docs/inference/ai-insights/structured-insights.md Structured insights extract data in a predefined JSON schema format, providing consistent, machine-readable results perfect for analytics, dashboards, and downstream processing. ## When to use structured insights Choose structured insights when you need: - **Quantitative Metrics** - Scores, ratings, counts, percentages. - **Categorical Data** - Status values, types, priorities, classifications. - **Boolean Flags** - Yes/no decisions, presence checks, compliance indicators. - **Consistent Format** - Data that feeds into databases, dashboards, or analytics systems. - **Multiple Related Fields** - Complex data with multiple attributes. | Use Case | Unstructured | Structured | |----------|--------------|------------| | Open-ended summaries | ✅ | ❌ | | Sentiment scoring (1-5) | ❌ | ✅ | | Issue categorization | ❌ | ✅ | | Descriptive analysis | ✅ | ❌ | | Compliance flags | ❌ | ✅ | | Dashboard metrics | ❌ | ✅ | ## Creating a structured insight ### Step 1: Start creating an insight 1. Navigate to [AI Insights](https://portal.telnyx.com/#/ai/insights) in the Portal. 2. Click **Create Insight**. 3. Enter a name and basic instructions. ### Step 2: Enable structured data mode Click the **Collect as structured data** button to reveal the schema configuration interface. ![Structured Data Mode](/assets/images/create-insight-structured-data.png) ### Step 3: Define parameters For each piece of data you want to extract, add a parameter with: 1. **Name** - The field name in the JSON output (e.g., `sentiment_score`, `issue_type`). 2. **Type** - The data type (string, number, boolean, etc.). 3. **Required** - Whether this field must always be present. 4. **Description** - Instructions for extracting this specific field. Click **Add parameter** to add additional fields to your schema. ## Parameter types ### String Text values - use for categories, descriptions, identifiers. **Example:** ``` Name: issue_category Type: string Description: The primary category of the customer's issue. Must be one of: "billing", "technical", "account", "general" ``` **Output:** ```json { "issue_category": "technical" } ``` ### Enum Predefined categories - use when you have a fixed set of possible values. The AI agent will select one value from your provided list. **Example:** ``` Name: sentiment Type: enum Enum Values (comma separated): positive, negative, neutral, mixed Description: Overall sentiment of the customer during the conversation ``` **Output:** ```json { "sentiment": "positive" } ``` Enums provide better accuracy than asking the AI agent to choose from options in a string description. They enforce strict value validation and make the AI agent's task clearer. ### Number Numeric values - use for scores, ratings, counts, percentages. **Example:** ``` Name: satisfaction_score Type: number Description: Customer satisfaction rating from 1-10 based on their tone and statements ``` **Output:** ```json { "satisfaction_score": 8 } ``` ### Integer Whole number values - use when you need integers without decimals (counts, quantities, IDs). **Example:** ``` Name: message_count Type: integer Description: Total number of messages sent by the customer in this conversation ``` **Output:** ```json { "message_count": 5 } ``` Use `integer` instead of `number` when you specifically need whole numbers. This provides clearer intent and can help the AI agent avoid returning decimal values. ### Boolean True/false values - use for flags, presence checks, yes/no decisions. **Example:** ``` Name: escalation_needed Type: boolean Description: True if the conversation requires escalation to a supervisor or specialist ``` **Output:** ```json { "escalation_needed": false } ``` ### Array Lists of values - use for multiple items of the same type. **Example:** ``` Name: products_mentioned Type: array Description: List of product names or SKUs mentioned during the conversation ``` **Output:** ```json { "products_mentioned": ["Widget Pro", "Widget Basic", "Extended Warranty"] } ``` ### Array (string) Lists of text values - use for multiple string items with enforced type safety. **Example:** ``` Name: issue_keywords Type: array (string) Description: Key terms or phrases mentioned by the customer that describe their issue ``` **Output:** ```json { "issue_keywords": ["billing", "refund", "overcharge"] } ``` ### Array (number) Lists of numeric values - use for multiple numbers with enforced type safety. **Example:** ``` Name: mentioned_prices Type: array (number) Description: All price points or dollar amounts mentioned during the conversation ``` **Output:** ```json { "mentioned_prices": [29.99, 49.99, 99.99] } ``` ### Array (boolean) Lists of true/false values - use for multiple boolean flags with enforced type safety. **Example:** ``` Name: feature_preferences Type: array (boolean) Description: Customer preferences for features in order: email notifications, SMS alerts, push notifications ``` **Output:** ```json { "feature_preferences": [true, false, true] } ``` Typed arrays (string, number, boolean) provide better type safety than the generic `array` type. Use them when you know all elements will be of the same specific type. ### Object Nested structures - use for complex related data. **Example:** ``` Name: customer_info Type: object Description: Extracted customer information including name, account number, and contact preference ``` **Output:** ```json { "customer_info": { "name": "Jane Smith", "account_number": "ACC-12345", "preferred_contact": "email" } } ``` ## Complete examples ### Example 1: Sentiment analysis **Configuration:** ``` Name: Customer Sentiment Analysis Instructions: Analyze the customer's emotional state throughout the conversation. Parameters: 1. positivity_score - Type: number - Required: Yes - Description: Rate positive sentiment from 1-5 (1=very negative, 5=very positive) 2. frustration_level - Type: number - Required: Yes - Description: Rate customer frustration from 1-5 (1=not frustrated, 5=very frustrated) 3. overall_sentiment - Type: string - Required: Yes - Description: Overall sentiment classification: "positive", "neutral", or "negative" 4. key_emotions - Type: array - Required: No - Description: List of specific emotions detected (e.g., "happy", "confused", "angry", "satisfied") ``` **Sample Output:** ```json { "positivity_score": 4, "frustration_level": 2, "overall_sentiment": "positive", "key_emotions": ["satisfied", "relieved", "appreciative"] } ``` ### Example 2: Sales qualification **Configuration:** ``` Name: Lead Qualification Instructions: Assess the sales opportunity from this conversation. Parameters: 1. budget_mentioned - Type: boolean - Required: Yes - Description: Did the prospect mention or discuss budget? 2. budget_range - Type: string - Required: No - Description: If mentioned, what budget range? (e.g., "under $1000", "$1000-$5000", "over $5000") 3. decision_timeframe - Type: string - Required: Yes - Description: When do they need to make a decision? ("immediate", "this_month", "this_quarter", "exploring", "unknown") 4. pain_points - Type: array - Required: Yes - Description: List of specific problems or needs mentioned 5. competitor_mentions - Type: array - Required: No - Description: Names of competing solutions mentioned 6. lead_score - Type: number - Required: Yes - Description: Qualification score from 1-10 based on buying signals ``` **Sample Output:** ```json { "budget_mentioned": true, "budget_range": "$1000-$5000", "decision_timeframe": "this_month", "pain_points": [ "Manual data entry taking too long", "Errors in current system", "No mobile access" ], "competitor_mentions": ["CompetitorX", "LegacyTool"], "lead_score": 8 } ``` ### Example 3: Support ticket categorization **Configuration:** ``` Name: Support Ticket Classification Instructions: Categorize and triage the support request. Parameters: 1. issue_type - Type: string - Required: Yes - Description: Primary issue type: "technical", "billing", "account", "feature_request", "bug_report" 2. priority - Type: string - Required: Yes - Description: Urgency level: "critical", "high", "medium", "low" 3. affected_service - Type: string - Required: No - Description: Which product/service is affected? 4. resolved - Type: boolean - Required: Yes - Description: Was the issue resolved during this conversation? 5. resolution_time_minutes - Type: number - Required: No - Description: If resolved, how many minutes did it take? 6. follow_up_needed - Type: boolean - Required: Yes - Description: Does this require follow-up action? 7. tags - Type: array - Required: No - Description: Relevant tags for categorization (e.g., "password_reset", "billing_dispute", "api_error") ``` **Sample Output:** ```json { "issue_type": "technical", "priority": "high", "affected_service": "API Integration", "resolved": true, "resolution_time_minutes": 12, "follow_up_needed": false, "tags": ["api_error", "authentication", "resolved"] } ``` ## Advanced mode Enable **Advanced mode** (checkbox at the top of the structured data section) to access additional schema configuration options: - Custom validation rules. - Enum constraints for string values. - Min/max constraints for numbers. - Pattern matching for strings. - Nested object definitions. ## Best practices ### 1. Keep schemas focused Don't try to extract everything in one insight. Create multiple focused insights instead: - ✅ Separate insights for "Sentiment" and "Issue Classification" - ❌ One massive insight trying to capture sentiment, classification, entities, summary, etc. ### 2. Make instructions clear Each parameter's description should be crystal clear: ``` ❌ Bad: "sentiment score" ✅ Good: "Rate overall sentiment from 1-10, where 1 is very negative, 5 is neutral, and 10 is very positive" ``` ### 3. Use enums for categories When you have a fixed set of categories, use the **enum** type instead of listing values in a string description: **❌ Less effective (using string with description):** ``` Type: string Description: Classification must be one of: "billing_question", "technical_support", "feature_request", "complaint", "general_inquiry" ``` **✅ Better (using enum type):** ``` Type: enum Enum Values (comma separated): billing_question, technical_support, feature_request, complaint, general_inquiry Description: The primary classification of the customer inquiry ``` Using the enum type provides better accuracy and enforces strict value validation. See the [Enum parameter type](#enum) section for more details. ### 4. Mark optional appropriately Only mark fields as required if they should always be extractable: - **Required**: `overall_sentiment` - should always be detectable - **Optional**: `competitor_mentioned` - may not come up in every conversation ### 5. Provide value ranges For numeric fields, specify the range: ``` Description: Urgency score from 1-5, where 1 is low priority and 5 is critical/urgent ``` ### 6. Test with edge cases Test your structured insights with: - Very short conversations. - Conversations where some information is missing. - Ambiguous or unclear discussions. - Multiple topics in one conversation. ## Next steps - **[Create Insight Groups](https://developers.telnyx.com/docs/inference/ai-insights/insight-groups)** - Organize your structured insights. - **[Explore Use Cases](https://developers.telnyx.com/docs/inference/ai-insights/use-cases)** - Industry-specific structured insight examples. ## Related resources - [Creating Insights](https://developers.telnyx.com/docs/inference/ai-insights/creating-insights) - Basic insight creation. - [Insight Groups](https://developers.telnyx.com/docs/inference/ai-insights/insight-groups) - Organizing insights. - [API Reference](/api-reference/assistants/list-assistants) - Programmatic access to insights. --- ### Insight groups > Source: https://developers.telnyx.com/docs/inference/ai-insights/insight-groups.md Insight Groups allow you to organize related insights into collections that can be assigned to AI Assistants and configured with webhook delivery. Groups provide a modular, reusable approach to conversation analysis. ## Overview An **Insight Group** is a named collection of insights with optional webhook configuration. Key features: - **Reusable** - Assign the same group to multiple assistants. - **Modular** - Mix and match insights across different groups. - **Flexible Delivery** - Configure unique webhook URLs per group. - **Organized** - Group insights by use case, department, or business function. ## Accessing insight groups 1. Log in to the [Mission Control Portal](https://portal.telnyx.com). 2. Navigate to **AI, Storage and Compute** > **AI Insights**. 3. Click the **AI Insight Groups** tab. ![AI Insight Groups Page](/assets/images/ai-insight-groups-page.png) The page displays: - **ID** - Unique identifier for the group (copyable). - **Name** - Group name. - **Webhook URL** - Configured webhook endpoint (or "-" if none). - **Insights Count** - Number of insights in the group. - **Created At** - When the group was created. ## Creating an insight group ### Step 1: Open the create dialog Click the **Create Insight Group** button in the top-right corner. ### Step 2: Configure the group ![Create Insight Group Modal](/assets/images/create-insight-group-modal.png) Fill in the following fields: #### Name (required) A descriptive name that indicates the group's purpose. **Examples:** - "Customer Service Analytics". - "Sales Qualification Metrics". - "Healthcare Compliance Checks". - "E-commerce Customer Insights". - "Support Ticket Classification". **Best practices:** - Use clear, business-oriented names. - Include the use case or department. - Avoid generic names like "Group 1" or "Test Group". #### Webhook URL (optional) The HTTPS endpoint where insight results will be sent after each conversation. **Format:** `https://your-domain.com/webhooks/insights` **When to use:** - You want real-time insight delivery to your application. - You're building dashboards or analytics systems. - You need to trigger actions based on insights. - You're integrating with external systems. **When to skip:** - You only need to view insights in the Portal. - You're using the API to fetch insights on-demand. - You're still testing and refining your insights. #### Insights (multi-select) Select which insights to include in this group: 1. Click the **Select an insight** dropdown. 2. Search or browse available insights. 3. Click to add an insight to the group. 4. Repeat to add multiple insights. **You can:** - Add multiple insights to one group. - Use the same insight in multiple groups. - Create groups with a single insight. - Modify group membership after creation. ### Step 3: Save the group 1. Review your configuration. 2. Click **Save**. 3. The group will appear in the Insight Groups list. ## Example configurations ### Example 1: Customer support group ``` Name: Customer Support Analytics Webhook URL: https://api.mycompany.com/webhooks/support-insights Insights: - Customer Sentiment Analysis - Issue Classification - Resolution Status - Follow-up Required ``` **Use Case:** Automatically analyze support calls and send results to your ticketing system. ### Example 2: Sales qualification group ``` Name: Sales Lead Qualification Webhook URL: https://crm.mycompany.com/webhooks/lead-insights Insights: - Budget Discussion - Decision Timeframe - Pain Points Identified - Competitor Mentions - Lead Score ``` **Use Case:** Score sales calls and update CRM with qualification data. ### Example 3: Compliance monitoring group ``` Name: Healthcare Compliance Webhook URL: https://compliance.healthcorp.com/insights Insights: - HIPAA Compliance Check - Required Disclosures Made - Consent Verification - Patient Information Handled ``` **Use Case:** Ensure regulatory compliance and maintain audit trail. ### Example 4: Quality assurance group ``` Name: Call Quality Metrics Webhook URL: - Insights: - Agent Performance - Script Adherence - Professional Tone - Issue Resolution Quality ``` **Use Case:** Manual review of call quality without webhook integration. ## Managing insight groups ### Editing a group 1. Find the group in the list. 2. Click the **edit icon** (pencil) on the right. 3. Modify name, webhook URL, or insights. 4. Click **Save**. **Note:** Changes to the group will apply to all assistants using this group. ### Copying a group ID Each group has a unique ID for API access: 1. Click the **copy icon** next to the group ID. 2. The ID will be copied to your clipboard. Example ID format: `a2708926-c060-480a-8631-041cb7304117` ## Assigning groups to assistants ### Via portal (during assistant creation) When creating or editing an AI Assistant: 1. Navigate to the **Insights** tab under **Analysis**. 2. Select an Insight Group from the dropdown. 3. Save the assistant configuration. See the [Voice Assistant Quickstart](https://developers.telnyx.com/docs/inference/ai-assistants/no-code-voice-assistant#insights) for detailed steps. ### Via API Use the `insight_settings` field when creating or updating an assistant: ```bash curl -X POST https://api.telnyx.com/v2/ai/assistants \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "name": "Customer Support Assistant", "model": "moonshotai/Kimi-K2.5", "instructions": "You are a helpful support assistant...", "insight_settings": { "insight_group_id": "a2708926-c060-480a-8631-041cb7304117" } }' ``` ## Using insights from groups ### In conversation history View insight results in the conversation history: 1. Navigate to **AI Assistants** > select an assistant. 2. Go to the **Conversation History** tab under **Analysis**. 3. Click on a conversation. 4. Scroll to the **Insights** section to see results. ### With memory API Filter which insights are included when querying conversation memory: ```json { "memory": { "conversation_query": "metadata->user_id=eq.123&limit=5", "insight_query": "insight_ids=insight_abc,insight_def,insight_xyz" } } ``` **Get insight IDs from your group:** 1. View the group details in the Portal. 2. Copy the ID of each insight in the group. 3. Include those IDs in the `insight_query` parameter. Learn more in the [Memory documentation](https://developers.telnyx.com/docs/inference/ai-assistants/memory). ### Via webhooks If you configured a webhook URL (on the Insight Group or on a single insight), results are automatically delivered as an HTTPS POST with a JSON body: ```json { "record_type": "event", "event_type": "conversation_insight_result", "payload": { "request_id": "0b6a2b0e-...", "conversation_id": "7f3c9a52-...", "user_id": "0a1b2c3d-...", "status": "completed", "insights_instructions": "Summarize the conversation", "insight_id": null, "insight_group_id": "a2708926-c060-480a-8631-041cb7304117", "results": [ { "insight_id": "insight_xyz789", "result": "{\"score\": 8, \"sentiment\": \"positive\"}" }, { "insight_id": "insight_abc456", "result": "{\"category\": \"technical\", \"priority\": \"high\"}" } ], "metadata": { "your-conversation-metadata": "is echoed here" } } } ``` Field notes: - `event_type` is **`conversation_insight_result`** — one webhook per insight request, sent when the request completes (or fails after retries, with `status: "failed"`). - `payload.results` — array of one entry per insight that ran. Each `result` is the model output for that insight (a JSON string when the insight defines a `json_schema`, otherwise free text). `insight_id` is `null` when the run used free-form `insights_instructions` instead of a saved insight. - `payload.insight_group_id` — the group the run was generated for, `null` when the run targeted a single insight or free-form instructions. - `payload.metadata` — the conversation's metadata, echoed back to you. - Delivery is signed with the standard Telnyx Ed25519 webhook signature (`telnyx-signature-ed25519` + `telnyx-timestamp` headers), like other Telnyx webhooks. - Note for early adopters: an earlier version of this page documented an event named `conversation.insights.completed` with a flat body. That event was never shipped — all deliveries use `conversation_insight_result` as described above. ## Organization strategies ### By use case Group insights by business function: - "Sales Qualification". - "Customer Support". - "Compliance Monitoring". - "Quality Assurance". **Benefits:** - Clear ownership by department. - Focused analytics per use case. - Easy to assign to specialized assistants. ### By delivery destination Group insights by webhook endpoint: - "CRM Integration Group" → `https://crm.company.com/insights`. - "Analytics Dashboard Group" → `https://analytics.company.com/insights`. - "Ticket System Group" → `https://tickets.company.com/insights`. **Benefits:** - Simplified webhook management. - Targeted data delivery. - System-specific insight collections. ### By analysis type Group insights by the kind of analysis: - "Quantitative Metrics" (scores, ratings, counts). - "Categorical Classification" (types, statuses, priorities). - "Qualitative Analysis" (summaries, descriptions). **Benefits:** - Consistent data structures. - Easier downstream processing. - Clear analytical purpose. ### Hybrid approach Combine strategies for complex scenarios: - "Sales - Quantitative" (lead scoring, budget ranges). - "Sales - Qualitative" (pain points, objections). - "Support - Urgent" (critical issues only). - "Support - Complete" (all support metrics). ## Best practices ### 1. Start small Begin with a focused group: - ✅ 2-4 related insights. - ❌ 15+ insights covering everything. You can always add more insights later. ### 2. Test without webhooks first When creating a new group: 1. Leave webhook URL empty initially. 2. Assign to a test assistant. 3. Review results in the Portal. 4. Add webhook once validated. ### 3. Use descriptive names Make group purpose immediately clear: - ✅ "Healthcare Compliance Checks". - ✅ "E-commerce Order Analysis". - ❌ "Group 1". - ❌ "Test". ### 4. Document webhook endpoints Maintain documentation of: - What each webhook URL expects. - Who owns the endpoint. - What system processes the insights. - How to troubleshoot delivery issues. ### 5. Version your groups When making significant changes: 1. Create a new group (e.g., "Support Analytics v2"). 2. Test with a subset of assistants. 3. Migrate fully once validated. 4. Retire the old group. ### 6. Monitor insights count Keep groups manageable: - **1-5 insights**: Focused, fast analysis. - **5-10 insights**: Comprehensive, still efficient. - **10+ insights**: May be slower, consider splitting. ## Troubleshooting ### Insights not appearing **Check:** - Is the group assigned to the assistant? - Did a conversation complete after assignment? - Are the individual insights configured correctly? ### Webhook not receiving data **Check:** - Is the webhook URL correct and accessible? - Is the endpoint returning 200 OK status? - Check webhook logs in your application. ### Wrong insights in group **Solution:** 1. Edit the group. 2. Remove incorrect insights. 3. Add correct insights. 4. Save changes. Changes apply to future conversations immediately. ## Next steps - **[Explore Use Cases](https://developers.telnyx.com/docs/inference/ai-insights/use-cases)** - Industry-specific group examples. - **[Assign to Assistants](https://developers.telnyx.com/docs/inference/ai-assistants/no-code-voice-assistant#insights)** - Connect groups to your AI assistants. ## Related resources - [Creating Insights](https://developers.telnyx.com/docs/inference/ai-insights/creating-insights) - Build insights for your groups. - [Structured Insights](https://developers.telnyx.com/docs/inference/ai-insights/structured-insights) - Define consistent data schemas. - [Memory](https://developers.telnyx.com/docs/inference/ai-assistants/memory) - Use insights with conversation memory. - [API Reference](/api-reference/assistants/list-assistants) - Programmatic group management. --- ### Telnyx-managed insights > Source: https://developers.telnyx.com/docs/inference/ai-insights/telnyx-managed-insights.md ## Overview Telnyx provides a set of **built-in insights** that measure assistant quality out of the box — no prompt engineering or schema design required. These are called **Telnyx-managed insights** (also referred to as *default* insights in the API and Portal). Unlike [custom insights](https://developers.telnyx.com/docs/inference/ai-insights/creating-insights) that you define yourself, Telnyx-managed insights are maintained by Telnyx, have consistent scoring rubrics, and are available to every account. You just assign them to an assistant's Insight Group and start getting results. ## Available insights ### Agent Instruction Following Measures how well your assistant followed its system prompt and tool-use instructions during the conversation. | Score | Meaning | |-------|---------| | Excellent | Followed all instructions precisely | | Good | Followed instructions with minor deviations | | Fair | Missed one or more instructions but stayed on task | | Poor | Significantly deviated from instructions | | N/A | Could not be evaluated for this conversation | **When it matters:** Any time you care about prompt adherence — complex assistants with many tool instructions, compliance-sensitive flows, or when debugging unexpected assistant behavior. ### User Satisfaction Estimates how satisfied the caller was with the conversation based on their responses, tone, and engagement signals. | Score | Meaning | |-------|---------| | Excellent | User was clearly satisfied and engaged | | Good | User was generally satisfied | | Fair | User was neutral or mixed | | Poor | User was frustrated or dissatisfied | | N/A | Could not be evaluated for this conversation | **When it matters:** Customer support, sales calls, any voice flow where caller experience directly impacts business outcomes. The set of Telnyx-managed insights may grow over time. Check the Insight Group configuration in the Portal for the current list. ## Enabling Telnyx-managed insights ### Step 1: Add to an Insight Group 1. Navigate to [AI Insights](https://portal.telnyx.com/#/ai/insights) in the Portal. 2. Go to the **AI Insight Groups** tab. 3. Create a new group or edit an existing one. 4. In the insights dropdown, search for **Agent Instruction Following** and **User Satisfaction**. 5. Add both (or whichever you need) to the group. 6. Save the group. See [Insight Groups](https://developers.telnyx.com/docs/inference/ai-insights/insight-groups) for full group configuration details. ### Step 2: Assign the group to your assistant 1. Open your assistant in the [Portal](https://portal.telnyx.com/#/ai/assistants). 2. Go to the **Analysis** tab → **Insights** sub-tab. 3. Select the Insight Group you created above. 4. Save the assistant. ### Step 3: Have conversations Insights run automatically after each conversation completes. Results appear in two places: - **Per-conversation** — in the Conversation History tab for each individual call. - **Over time** — in the Insights Over Time tab as aggregated daily charts. ## Viewing results ### Per conversation 1. Open your assistant → **Analysis** tab → **Conversation History** sub-tab. 2. Click any conversation. 3. Scroll to the **Insights** section to see the scores for each insight in your group. Each result includes the score (e.g. "Good") and any additional detail the insight provides. ### Over time (7-day trend) The **Insights Over Time** sub-tab (third tab under Analysis) shows a **stacked-bar chart** of daily score counts for the last 7 days, scoped to the current assistant. **What the chart shows:** - One bar per day (UTC) for the last 7 days. - Each bar is segmented by score value, color-coded: - **Poor** — red - **Fair** — amber - **Good** — green - **Excellent** — blue - **N/A** — gray - Bars are stacked bottom-to-top with N/A pinned to the bottom and Poor through Excellent ascending, so the best results sit on top. **How to use it:** 1. Use the **Insight** dropdown to select which Telnyx-managed insight to view (only Telnyx-managed insights appear here — custom insights are excluded). 2. Read the chart left to right to spot trends — is the "Excellent" segment growing or shrinking over the week? 3. Check for sudden shifts on specific days, which often correlate with config changes or new assistant versions. ### Comparing assistant versions Toggle **Compare by assistant version** above the chart to split the data into separate charts — one per assistant version. This creates a small-multiples layout where each version gets its own labeled chart. Version comparison uses `metadata.assistant_version_id` to attribute conversations to versions. Conversations without a version ID are grouped as "Unknown version". This is useful when: - You've published a new version and want to see if quality metrics improved or regressed. - You're A/B testing different prompt configurations across versions. - You want to confirm a fix didn't introduce new issues before promoting a version. ## Tips for getting value from Telnyx-managed insights ### Spotting regressions after a change After updating your assistant's prompt, tools, or voice settings, watch the over-time chart for the next few days. If the "Poor" or "Fair" segments grow while "Good" or "Excellent" shrink, the change may have hurt quality. ### Correlating with volume Taller bars mean more conversations. A spike in volume can amplify small score shifts — check the raw counts in the chart tooltip before drawing conclusions. ### Understanding N/A rates A large gray (N/A) segment means the insight frequently couldn't be evaluated. For Agent Instruction Following, this may indicate conversations that were too short or didn't trigger tool use. For User Satisfaction, it may mean the conversation lacked enough user signal (e.g., one-sided calls). If N/A rates are high, check whether your assistant is handling the conversation types you expect. ### Using scores alongside conversation history The over-time chart tells you *when* quality shifted. The conversation history tells you *why*. When you see a dip, switch to Conversation History, filter to that day, and review individual conversations to find the root cause. ## Troubleshooting ### Chart shows "No insight data recorded" - Confirm the Insight Group includes Agent Instruction Following or User Satisfaction. - Confirm the group is assigned to the assistant in the Analysis tab. - Confirm the assistant has had conversations in the last 7 days. - Confirm conversations completed successfully — insights run after the conversation ends. ### Scores seem inconsistent Telnyx-managed insights use AI evaluation, which considers the full conversation context. Scores can vary based on conversation length, topic complexity, and user behavior. Look at trends over multiple days rather than individual conversations. ### Version comparison shows "Unknown version" Conversations that occurred before version tracking was enabled (or that lack version metadata) are grouped as "Unknown version." Ensure your assistant has published versions and that conversations are attributed correctly. ## Related resources - [Creating Insights](https://developers.telnyx.com/docs/inference/ai-insights/creating-insights) — Create your own custom insights. - [Structured Insights](https://developers.telnyx.com/docs/inference/ai-insights/structured-insights) — Define JSON schemas for consistent data extraction. - [Insight Groups](https://developers.telnyx.com/docs/inference/ai-insights/insight-groups) — Organize insights and configure webhooks. - [Use Cases](https://developers.telnyx.com/docs/inference/ai-insights/use-cases) — Industry-specific insight examples. - [Voice Assistant Configuration](https://developers.telnyx.com/docs/inference/ai-assistants/no-code-voice-assistant#insights) — Assigning insight groups to assistants. --- ## LiveKit on Telnyx (Beta) ### Overview > Source: https://developers.telnyx.com/docs/livekit.md # LiveKit on Telnyx Build, deploy, and scale voice AI agents on Telnyx's global network — the same LiveKit you know, hosted on Telnyx infrastructure with integrated telephony and AI models. ## Why LiveKit on Telnyx? LiveKit on Telnyx collapses your voice AI stack into a single platform: - **Built-in Telephony, No Third-Party SIP Fees** — Telnyx is the carrier. Buy numbers, configure SIP, connect to agents — all in one portal. On other platforms, you pay third-party SIP fees on top of your usage. On Telnyx, those are gone entirely. - **Ultra-Low Latency AI on Telnyx GPUs** — STT, TTS, and LLM run on Telnyx-owned GPUs, colocated in every region ~2ms from your agents. No round-trips to external APIs — faster responses, better conversations. - **Same LiveKit, Zero Migration Risk** — Same SDKs, same CLI, same agent framework. Swap three environment variables and redeploy — your code doesn't change. - **One Platform, One Bill** — Compute, models, telephony, and phone numbers — all on a single Telnyx invoice. No vendor sprawl. ## Not ready to migrate? Install the [Telnyx plugin](/docs/livekit/build#1-install-the-plugin) on your existing LiveKit Cloud or self-hosted setup. Same agent code, same infrastructure — just faster, colocated inference on Telnyx GPUs. See the [Plugin Reference](/docs/livekit/models) for available STT, TTS, and LLM options. ## Who is this for? - **LiveKit Cloud users** looking to consolidate vendors or reduce latency via Telnyx's network - **Self-hosted LiveKit operators** tired of managing infrastructure and wanting managed hosting - **New voice AI developers** who want a single platform for media, telephony, and AI ## How it works 1. **Connect** — Configure your LiveKit client to point to Telnyx's LiveKit cluster using your Telnyx credentials. 2. **Build** — Write agents using the LiveKit Agent framework with Telnyx plugins for STT, TTS, and LLM. 3. **Deploy** — Deploy agents to Telnyx's managed infrastructure. Telephony connects via built-in SIP. 4. **Scale** — Telnyx handles the infrastructure. You manage your agents and configuration. ## What's different from LiveKit Cloud? | Feature | LiveKit on Telnyx | LiveKit Cloud | |---------|-------------------|---------------| | SIP trunking | Built-in — no third-party SIP fees, HD voice (G.722 + Opus) | Third-party SIP fees on top of usage | | AI models | Colocated on Telnyx GPUs (~2ms from agents) | LiveKit Inference or third-party APIs | | Billing | Combined with Telnyx services | Separate | ## Next steps - [Quickstart](/docs/livekit/quickstart) — Get deployed in 5 minutes - [Connect](/docs/livekit/connect) — Configure your LiveKit client - [Telephony](/docs/livekit/telephony) — SIP and PSTN integration - [Plugin Reference](/docs/livekit/models) — STT, TTS, and LLM options --- ### Compatibility > Source: https://developers.telnyx.com/docs/livekit/compatibility.md ## What's the Same The LiveKit agent framework is 100% portable — your agent code does not change. Everything in LiveKit works identically on Telnyx. Notable sections from LiveKit docs: - **Agent Framework** — [docs.livekit.io/agents](https://docs.livekit.io/agents) - **SIP / Phone** — [docs.livekit.io/sip](https://docs.livekit.io/sip) - **API Reference** — [docs.livekit.io/api](https://docs.livekit.io/api) - **Cloud Deployment** — [docs.livekit.io/cloud](https://docs.livekit.io/cloud) - **Egress** — [docs.livekit.io/egress](https://docs.livekit.io/egress) ## What's Different These features work differently on Telnyx compared to LiveKit Cloud: - **STT / TTS / LLM** — LiveKit Cloud requires third-party AI providers. On Telnyx, hosted models are available out of the box via [`livekit-plugins-telnyx`](https://github.com/team-telnyx/telnyx-livekit-plugin). Prefer your own provider? Bring any API key. - **SIP trunking** — On other platforms, you pay third-party SIP fees on top of your usage. On Telnyx, those are gone entirely — SIP is built-in because you're already on the carrier. BYOT also supported if you prefer your existing provider. - **HD Voice** — G.722 (16 kHz) and Opus (48 kHz) codecs for wideband audio on SIP calls. G.722 is enabled by default. Opus requires SRTP — see [Telephony](/docs/livekit/telephony#hd-voice) for setup. - **Phone numbers** — Telnyx numbers in 140+ countries instead of LiveKit Phone (US only) - **Agent deployment** — Same `lk agent deploy` command. Just swap the URL: `lk agent deploy . --url .livekit.telnyx.com` - **Egress** — Telnyx call recording can capture SIP call audio. Enable it on the [phone number](/api-reference/phone-number-configurations/update-a-phone-number-with-voice-settings) for inbound calls or the SIP connection's [Outbound Voice Profile](/docs/voice/sip-trunking/configuration/outbound-voice-profiles#call-recording) for outbound calls. _Note: Available regions are nyc1, sfo3, atl1, and syd1. See the [Regions](/docs/livekit/regions) page for details._ ## What's Not Supported - **Ingress (RTMP/WHIP)** — importing external streams into a room - **Sandbox** — quick-start dev environments --- ### Quick Start > Source: https://developers.telnyx.com/docs/livekit/quickstart.md Go from zero to a working voice agent you can call from your phone. By the end, you'll have a LiveKit agent deployed on Telnyx infrastructure, connected to a real phone number. ## Prerequisites Before you start, make sure you have: - A [Telnyx account](https://portal.telnyx.com) - A Telnyx API key (generate one in the [portal](https://portal.telnyx.com)) - A secret key you create and keep safe — this is your `LIVEKIT_API_SECRET` - Python ≥ 3.10 - [LiveKit CLI](https://docs.livekit.io/home/cli/) (`lk`) version 2.16.0 or later Verify your CLI version: ```bash lk --version ``` ## Step 1: Configure the CLI Set your environment variables to point at a Telnyx LiveKit region: ```bash export TELNYX_API_KEY= export LIVEKIT_URL=https://.livekit-telnyx.com export LIVEKIT_API_KEY=$TELNYX_API_KEY export LIVEKIT_API_SECRET= ``` `LIVEKIT_API_KEY` is your Telnyx API key — the same one you use everywhere else on the platform. Set `TELNYX_API_KEY` once and reference it throughout. Register your credentials with the platform: ```bash curl -s -X POST "https://.livekit-telnyx.com/provision" \ -H "Content-Type: application/json" \ -d '{ "name": "my-project", "api_key": "'$TELNYX_API_KEY'", "api_secret": "'$LIVEKIT_API_SECRET'" }' ``` This creates your tenant record on the platform. You only need to do this once per region. ### Available regions | Region | Name (slug) | URL | |--------|------|-----| | New York | `nyc1` | `nyc1.livekit-telnyx.com` | | San Francisco | `sfo3` | `sfo3.livekit-telnyx.com` | | Atlanta | `atl1` | `atl1.livekit-telnyx.com` | | Sydney | `syd1` | `syd1.livekit-telnyx.com` | Choose the region closest to your users for the lowest latency. ## Step 2: Set up telephony To receive phone calls, you need a phone number (DID) pointed at the Telnyx LiveKit SIP servers. No third-party SIP fees — Telnyx is the carrier. ### Buy a phone number Search for an available number in your area code and order it: **1. Browse available numbers** (replace `512` with your area code): ```bash curl -sg "https://api.telnyx.com/v2/available_phone_numbers?filter[country_code]=US&filter[national_destination_code]=512" \ -H "Authorization: Bearer $TELNYX_API_KEY" | \ jq -r '["NUMBER","LOCATION","MONTHLY"], (.data[] | [.phone_number, (.region_information | map(.region_name) | join(", ")), .cost_information.monthly_cost]) | @tsv' | \ column -t | less ``` Note the number you want before exiting — you'll need it below. **2. Order the number you want:** ```bash curl -s -X POST "https://api.telnyx.com/v2/number_orders" \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -d '{"phone_numbers": [{"phone_number": "+1XXXXXXXXXX"}]}' | \ jq '{id: .data.id, status: .data.status, phone_number: .data.phone_numbers[0].phone_number}' ``` **3. Confirm the order completed:** ```bash curl -s "https://api.telnyx.com/v2/number_orders/ORDER_ID" \ -H "Authorization: Bearer $TELNYX_API_KEY" | \ jq '{id: .data.id, status: .data.status, phone_number: .data.phone_numbers[0].phone_number}' ``` Wait for `"success"` before moving on. If your order didn't complete successfully, [chat with us](https://telnyx.com/contact-us) — we're happy to help. 1. Log in to the [Telnyx portal](https://portal.telnyx.com) 2. Go to [**Real Time Communcations** → **Numbers** → **Buy Numbers**](https://portal.telnyx.com/#/numbers/buy-numbers) 3. Purchase a number by checking out ### Create a SIP connection A SIP connection tells Telnyx how to route inbound calls. You'll point it at one of our regional LiveKit SIP servers so calls land on the right infrastructure. | Region | Name (slug) | SIP FQDN | |--------|------|----------| | New York | `nyc1` | `nyc1.sip.livekit-telnyx.com` | | San Francisco | `sfo3` | `sfo3.sip.livekit-telnyx.com` | | Atlanta | `atl1` | `atl1.sip.livekit-telnyx.com` | | Sydney | `syd1` | `syd1.sip.livekit-telnyx.com` | **1. Create the FQDN connection** (replace `{slug}` with your region slug): ```bash curl -s -X POST "https://api.telnyx.com/v2/fqdn_connections" \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -d '{"connection_name": "{slug}-livekit-sip-connection"}' | \ jq '{id: .data.id, name: .data.connection_name}' ``` **2. Attach the regional SIP FQDN** (replace `CONNECTION_ID` and `{slug}` with your region slug): ```bash curl -s -X POST "https://api.telnyx.com/v2/fqdns" \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -d '{"connection_id": "CONNECTION_ID", "fqdn": "{slug}.sip.livekit-telnyx.com", "dns_record_type": "a", "port": 5060}' | \ jq '{fqdn: .data.fqdn, connection_id: .data.connection_id}' ``` **3. Assign your phone number to the connection** (replace `CONNECTION_ID` and phone number): ```bash curl -s -X PATCH "https://api.telnyx.com/v2/phone_numbers/+1XXXXXXXXXX" \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -d '{"connection_id": "CONNECTION_ID"}' | \ jq '{phone_number: .data.phone_number, connection_id: .data.connection_id}' ``` 1. In the portal, go to [**Real Time Communcations** → **Voice** → **SIP Connections** → **Create**](https://portal.telnyx.com/#/voice/connections) 2. Set the connection type to **FQDN** 3. Point it at your region's SIP endpoint from the table above 4. Assign your phone number to this SIP connection ### Register the number with the platform Set up an [inbound trunk](https://docs.livekit.io/telephony/accepting-calls/inbound-trunk/) for your number and a [dispatch rule](https://docs.livekit.io/telephony/accepting-calls/dispatch-rule/) to route calls to your agent. **1. Create an inbound SIP trunk:** ```bash echo '{ "trunk": { "name": "telnyx-inbound", "numbers": ["+1XXXXXXXXXX"], "allowed_addresses": ["192.76.120.0/22"] } }' | lk sip inbound create - ``` Note the trunk `sid` in the response — you'll need it in the next step. `allowed_addresses` restricts which source IPs can send SIP traffic to this trunk. `192.76.120.0/22` is Telnyx's SIP network — since Telnyx is your carrier, all inbound calls will originate from this range. **2. Create a dispatch rule:** ```bash echo '{ "dispatchRule": { "name": "route-to-agent", "trunkIds": [""], "rule": { "dispatchRuleIndividual": {} }, "roomConfig": { "agents": [{ "agentName": "agent" }] } } }' | lk sip dispatch create - ``` ## Step 3: Clone an example agent Clone the example agents repo and navigate to the restaurant agent: ```bash git clone https://github.com/team-telnyx/telnyx-livekit-agent-examples.git cd telnyx-livekit-agent-examples/restaurant ``` This is a fully working voice agent — an Italian restaurant ordering assistant built with [`livekit-plugins-telnyx`](https://github.com/team-telnyx/telnyx-livekit-plugin) for STT, TTS, and LLM. ## Step 4: Deploy Register the agent with the platform, then deploy it to your Telnyx LiveKit region: ```bash lk agent create . lk agent deploy . --secrets TELNYX_API_KEY=$TELNYX_API_KEY ``` The CLI uploads your agent code, builds a container image, and deploys it. You'll see build logs streaming in real-time. Check that your agent is running: ```bash lk agent status ``` `lk agent status` reads the agent ID from your `livekit.toml` file — you don't need to specify it manually. If you skip `lk agent create`, there's no ID in the file and the command will fail. ## Step 5: Call your agent Pick up your phone and dial the number you purchased. You should hear: > "Thanks for calling Bella's Kitchen!" Try ordering some pasta. The agent handles the full conversation — browsing the menu, answering questions, and taking your order. ### Troubleshooting - **Call doesn't connect** — Verify your SIP connection is pointed at the correct regional FQDN and your DID is assigned to it. - **Agent doesn't pick up** — Run `lk agent status` to confirm the agent is running. Check logs with `lk agent logs`. - **Audio quality issues** — Make sure you're using the region closest to you. ## Next steps You've deployed your first agent. Here's where to go from here: - [Build](/docs/livekit/build) — Write your own agent from scratch - [Deploy](/docs/livekit/deploy) — Regions, scaling, secrets, and production deployment - [Models](/docs/livekit/models) — STT, TTS, and LLM options available on Telnyx - [Telephony](/docs/livekit/telephony) — Inbound/outbound calls, dispatch rules, multiple numbers - [Compatibility](/docs/livekit/compatibility) — What's the same and different from LiveKit Cloud --- ### Connect > Source: https://developers.telnyx.com/docs/livekit/connect.md Connecting to the Telnyx LiveKit platform is a URL and credential swap. No code changes. ## Get your credentials Complete [Quickstart — Step 1](/docs/livekit/quickstart#step-1-configure-the-cli) to set your environment variables and register with the platform. That one-time setup covers everything you need to connect. Already set up? Just make sure your `LIVEKIT_URL`, `LIVEKIT_API_KEY`, and `LIVEKIT_API_SECRET` are exported in your shell. Verify the connection: ```bash lk room list ``` ## Available regions See [Regions](/docs/livekit/regions) for all available platform and SIP endpoints. ## SDK configuration Server SDKs (Python, Go, Node.js) just need the URL and credentials: ```python from livekit.api import LiveKitAPI api = LiveKitAPI( url="https://nyc1.livekit-telnyx.com", api_key="", api_secret="", ) ``` If you're migrating from LiveKit Cloud or self-hosted, change the three environment variables and everything else stays the same. --- ### From LiveKit Cloud > Source: https://developers.telnyx.com/docs/livekit/migration/from-livekit-cloud.md Zero code changes. Swap your credentials, redeploy, done. ## What changes | | LiveKit Cloud | Telnyx | |---|---|---| | `LIVEKIT_URL` | `https://your-project.livekit.cloud` | `https://nyc1.livekit-telnyx.com` | | `LIVEKIT_API_KEY` | LiveKit Cloud key | Telnyx API key | | `LIVEKIT_API_SECRET` | LiveKit Cloud secret | Your LiveKit API secret | ## What doesn't change - Your agent code - SDK usage - `lk` CLI commands - Dockerfiles - Room management, dispatch rules ## Steps 1. **Create a Telnyx account** at [portal.telnyx.com](https://portal.telnyx.com) 2. **Generate a Telnyx API key** in the portal 3. **Register with the platform** (one-time): ```bash curl -s -X POST "https://.livekit-telnyx.com/provision" \ -H "Content-Type: application/json" \ -d '{ "name": "my-project", "api_key": "'$TELNYX_API_KEY'", "api_secret": "'$LIVEKIT_API_SECRET'" }' ``` 4. **Update your environment variables** with the values above 5. **Deploy to Telnyx:** ```bash lk agent deploy . --secrets TELNYX_API_KEY= ``` 5. **Verify:** `lk room list` to confirm connectivity 6. **(Optional)** Install the [Telnyx plugin](/docs/livekit/build) for on-prem STT, TTS, and LLM 7. **(Optional)** [Buy a Telnyx number](/docs/livekit/telephony) for built-in telephony ## What you gain - **Built-in telephony, no third-party SIP fees** — On other platforms, you pay third-party SIP fees on top of your usage. On Telnyx, those are gone entirely. - **Ultra-low latency AI on Telnyx GPUs** — STT, TTS, and LLM colocated in every region ~2ms from your agents. No round-trips to external APIs. - **One platform, one bill** — Compute, models, telephony, and phone numbers on a single Telnyx invoice. No vendor sprawl. ## Rolling back It's just environment variables. Point them back at LiveKit Cloud and redeploy. --- ### From Self-Hosted > Source: https://developers.telnyx.com/docs/livekit/migration/from-self-hosted.md Keep your agents, ditch the infrastructure. No more managing Kubernetes clusters, TURN servers, SIP services, or autoscaling. ## What Telnyx replaces - LiveKit SFU server - SIP service / Otel SIP - TURN/STUN servers - Container orchestration / autoscaling - TLS termination and load balancing - Monitoring infrastructure ## What stays the same - Your agent code - SDK calls - Room management logic - Dispatch rules ## Steps 1. **Create a Telnyx account** at [portal.telnyx.com](https://portal.telnyx.com) 2. **Generate a Telnyx API key** in the portal 3. **Register with the platform** (one-time): ```bash curl -s -X POST "https://.livekit-telnyx.com/provision" \ -H "Content-Type: application/json" \ -d '{ "name": "my-project", "api_key": "'$TELNYX_API_KEY'", "api_secret": "'$LIVEKIT_API_SECRET'" }' ``` 4. **Update your agent** to point at `https://nyc1.livekit-telnyx.com` (or your [preferred region](/docs/livekit/connect#available-regions)) 4. **Move secrets** to the platform: ```bash lk agent deploy . --secrets TELNYX_API_KEY=,OTHER_SECRET= ``` 5. **Re-create dispatch rules** on the Telnyx platform 6. **Update SIP configuration** to point at Telnyx [SIP endpoints](/docs/livekit/telephony) 7. **Verify** your agent is running: `lk room list` 8. **Decommission** your self-hosted infrastructure ## What you gain - **No infra management** — No Kubernetes, no TURN, no Redis for multi-region. - **Built-in SIP, no third-party SIP fees** — On other platforms, you pay third-party SIP fees on top of your usage. On Telnyx, those are gone entirely. - **Managed autoscaling** — Platform handles container lifecycle. - **Ultra-low latency AI on Telnyx GPUs** — STT, TTS, and LLM colocated in every region ~2ms from your agents. No round-trips to external APIs. --- ### Build > Source: https://developers.telnyx.com/docs/livekit/build.md **First time building?** Complete [Quickstart — Step 1](/docs/livekit/quickstart#step-1-configure-the-cli) to configure your CLI and register with the platform. ## 1. Install the plugin ```bash pip install "telnyx-livekit-plugin @ git+https://github.com/team-telnyx/telnyx-livekit-plugin.git#subdirectory=telnyx-livekit-plugin" ``` This is the Telnyx-maintained plugin, which includes the latest features and fixes ahead of the upstream `livekit-plugins-telnyx` PyPI package. We host our own release to iterate faster than the upstream cadence allows. Use the git install above for the current recommended version. *[Plugin source on GitHub](https://github.com/team-telnyx/telnyx-livekit-plugin)* ## 2. Write your agent Create `agent.py`: ```python import asyncio from livekit.agents import Agent, AgentSession, JobContext, RoomInputOptions from livekit.plugins import openai, silero, telnyx class MyAgent(Agent): def __init__(self): super().__init__(instructions="You are a helpful voice assistant.") async def on_enter(self): self.session.generate_reply( instructions="Greet the caller and ask how you can help." ) async def entrypoint(ctx: JobContext): session = AgentSession( stt=telnyx.STT( transcription_engine="Deepgram", base_url="wss://api.telnyx.com/v2/speech-to-text/transcription", ), llm=openai.LLM.with_telnyx(model="zai-org/GLM-5.2"), tts=telnyx.TTS( voice="Telnyx.Ultra.Clara", sample_rate=24000, ), vad=silero.VAD.load(), ) await ctx.connect() await session.start( agent=MyAgent(), room=ctx.room, room_input_options=RoomInputOptions(), ) disconnect_event = asyncio.Event() @ctx.room.on("disconnected") def on_disconnect(*args): disconnect_event.set() await disconnect_event.wait() ``` The only Telnyx-specific parts are the `stt`, `tts`, and `llm` — everything else is standard LiveKit. ### Room name prefixes Telnyx prefixes LiveKit room names for routing. If your application logic reads `ctx.room.name`, you must strip the Telnyx prefix before comparing it with room names from your own system: ```python room_name = ctx.room.name if ":" in room_name: _, room_name = room_name.split(":", 1) ``` ## 3. Add a requirements file Create `requirements.txt`: ``` livekit-agents livekit-plugins-openai livekit-plugins-silero telnyx-livekit-plugin @ git+https://github.com/team-telnyx/telnyx-livekit-plugin.git#subdirectory=telnyx-livekit-plugin ``` ## 4. Deploy ```bash lk agent deploy . --secrets TELNYX_API_KEY=$TELNYX_API_KEY ``` The platform builds your image in-cluster, deploys it, and starts the worker. Check the status: ```bash lk agent list ``` Tail logs once it's running: ```bash lk agent logs --id ``` ## Swap models The example above uses Telnyx-hosted defaults. Swap any component independently: - **STT** — Deepgram, Nova-3, Nova-2, Flux → [STT plugin](/docs/livekit/models/stt) - **TTS** — Telnyx Ultra, MiniMax, ElevenLabs (BYOK), Azure → [TTS plugin](/docs/livekit/models/tts) - **LLM** — GPT-4o-mini, Kimi-K2.6, Claude (BYOK) → [LLM plugin](/docs/livekit/models/llm) ## Next steps - [Telephony](/docs/livekit/telephony) — Connect your agent to a phone number - [Deploy reference](/docs/livekit/deploy) — Secrets, rollbacks, multi-region --- ### Deploy > Source: https://developers.telnyx.com/docs/livekit/deploy.md Same commands as LiveKit Cloud. `lk agent deploy` uploads your code, builds a container image on Telnyx's build service, and deploys the worker. **First time deploying?** Complete [Quickstart — Step 1](/docs/livekit/quickstart#step-1-configure-the-cli) to configure your CLI and register with the platform. ## 1. Create your agent ```bash lk agent create . ``` Registers the agent and writes the agent ID to `livekit.toml`. Skip this if you've already created the agent. ## 2. Deploy ```bash lk agent deploy . --secrets TELNYX_API_KEY=$TELNYX_API_KEY ``` Subsequent deploys roll out a new version. ## 3. Check status ```bash lk agent list ``` ## 4. Tail logs ```bash lk agent logs --id ``` For build logs: ```bash lk agent logs --id --log-type build ``` ## 5. Update secrets Add or update secrets on a running agent: ```bash lk agent update-secrets --secrets "OPENAI_API_KEY=" ``` ## 6. Rollback ```bash lk agent rollback ``` ## What gets auto-injected The platform injects these into your agent's environment at runtime — you don't need to set them manually: - `LIVEKIT_URL` - `LIVEKIT_API_KEY` — your Telnyx API key - `LIVEKIT_API_SECRET` ## More - [Secrets](/docs/livekit/deploy/secrets) — Managing API keys and sensitive config - [Configuration](/docs/livekit/deploy/configuration) — Regions, multi-agent projects, env vars - [Management](/docs/livekit/deploy/management) — What Telnyx manages vs. what you control --- ### Configuration > Source: https://developers.telnyx.com/docs/livekit/deploy/configuration.md ## Regions Deploy to the region closest to your users. See [Regions](/docs/livekit/regions) for all available endpoints. **Telnyx-specific regional deployment:** `nyc1` is the default agent deployment region. To deploy to another region, set `LK_AGENTS_URL` to that region's agent deployment endpoint and set `LIVEKIT_URL` to the matching regional LiveKit endpoint. ```bash export LIVEKIT_URL=https://.livekit-telnyx.com export LK_AGENTS_URL=https://.agents.livekit-telnyx.com lk agent deploy . --secrets TELNYX_API_KEY=$TELNYX_API_KEY ``` Use the same region slug for both variables. For example, if your rooms and SIP traffic use `sfo3.livekit-telnyx.com`, set `LK_AGENTS_URL` to `https://sfo3.agents.livekit-telnyx.com` before deploying. Telnyx also prefixes room names for routing; see [Room name prefixes](/docs/livekit/build#room-name-prefixes). ## Environment variables and livekit.toml Set `LIVEKIT_URL`, `LIVEKIT_API_KEY`, and `LIVEKIT_API_SECRET` as environment variables. You also need a `livekit.toml` in your project directory that defines the subdomain and agent ID: ```toml [project] subdomain = "nyc1" [agent] id = "" ``` `lk agent create` writes the agent ID to `livekit.toml`. Don't edit the `[agent]` section manually. If you clone an example repo, remove the `[agent]` section before running `lk agent create` so a fresh ID is generated. For multi-agent projects, use the `--config` flag to point at different TOML files: ```bash lk agent deploy . --config agent-a.toml lk agent deploy . --config agent-b.toml ``` ## Multiple environments Use separate API keys for dev, staging, and production environments. Each key operates independently. ## Log access Stream logs from your running agent: ```bash lk agent logs ``` Log forwarding to third-party services (Datadog, Sentry, CloudWatch) is coming soon. --- ### Secrets > Source: https://developers.telnyx.com/docs/livekit/deploy/secrets.md ## Setting secrets on deploy Pass secrets when you deploy: ```bash lk agent deploy . --secrets TELNYX_API_KEY= ``` To set additional secrets after deploy: ```bash lk agent update-secrets --secrets "OPENAI_API_KEY=" ``` ## Updating secrets Update secrets on a running agent: ```bash lk agent update-secrets --secrets "KEY=value" ``` ## Auto-injected variables The platform automatically injects these into your agent's environment at runtime: - `LIVEKIT_URL` - `LIVEKIT_API_KEY` (your Telnyx API key — used for both platform auth and model access) - `LIVEKIT_API_SECRET` You don't need to set these manually. Note: `LIVEKIT_API_KEY` is your Telnyx API key. Telnyx can revoke it if needed, giving you a single credential to manage. ## Local development For local development, use a `.env.local` file: ```bash TELNYX_API_KEY=your-key-here LIVEKIT_URL=https://nyc1.livekit-telnyx.com LIVEKIT_API_KEY=your-key LIVEKIT_API_SECRET=your-secret ``` Secrets set via the CLI are encrypted at rest and injected at runtime. They cannot be retrieved after being set. --- ### Management & Access > Source: https://developers.telnyx.com/docs/livekit/deploy/management.md ## Access model All platform interaction is through the `lk` CLI and the [Telnyx Portal](https://portal.telnyx.com). There is no direct access to underlying infrastructure (no kubectl, no SSH). ## What Telnyx manages - LiveKit SFU (media server) - SIP service (built-in telephony) - Agent container runtime and orchestration - Autoscaling - TLS/SSL termination and load balancing - Health checks and rolling deploys ## What you control | Capability | How | |-----------|-----| | Agent code and Dockerfile | `lk agent deploy .` | | Secrets | `lk agent deploy --secrets` / `lk agent update-secrets` | | Rollback | `lk agent rollback` | | Logs | `lk agent logs` | | Phone numbers | Telnyx Portal | | SIP connections | Telnyx Portal | | Model selection | Telnyx plugin or BYOK | ## Current limitations - No volume mounts or persistent storage - No custom networking or network policies - No privileged containers - No custom domains - No direct database access — use external services with secrets See [Compatibility](/docs/livekit/compatibility) for the full list of supported and unsupported features. --- ### Overview > Source: https://developers.telnyx.com/docs/livekit/models.md All models are accessed through the [Telnyx plugin](/docs/livekit/build). Models run on Telnyx GPUs, so whether you're using the plugin from LiveKit Cloud or deployed on the Telnyx platform, you get on-prem inference without managing infrastructure. Deepgram Nova-3, Nova-2, and Flux — hosted on Telnyx. Extensive voice library across multiple providers and models. Hosted open-source and proprietary models via BYOK. --- ### STT > Source: https://developers.telnyx.com/docs/livekit/models/stt.md Telnyx hosts Deepgram models on dedicated GPUs. Access them through the Telnyx plugin: ```python from livekit.plugins import telnyx stt = telnyx.deepgram.STT(model="nova-3", language="en") ``` The plugin provides two STT classes: - **`telnyx.deepgram.STT`** — Recommended. Connects to Deepgram models hosted on Telnyx GPUs. Takes a `model` parameter (`nova-3`, `nova-2`, `flux`). - **`telnyx.STT`** — Connects to Telnyx's own transcription engine (default) or Deepgram via `transcription_engine="Deepgram"`. Takes a `transcription_engine` parameter instead of `model`. For most use cases, `telnyx.deepgram.STT` is the simpler interface. *[Plugin source on GitHub](https://github.com/team-telnyx/telnyx-livekit-plugin)* ## Available models ### Nova-3 (recommended) Latest generation, best accuracy. ```python stt = telnyx.deepgram.STT( model="nova-3", language="en", interim_results=True, keyterm=["YourBrand", "custom-term"], # keyword boosting ) ``` ### Nova-2 Previous generation, stable and reliable. Uses weighted keyword boosting. ```python stt = telnyx.deepgram.STT( model="nova-2", language="en", interim_results=True, keywords=["YourBrand:2.0", "custom-term:1.5"], ) ``` ### Flux Experimental, with built-in end-of-turn detection. Designed for real-time voice agents. ```python stt = telnyx.deepgram.STT( model="flux", language="en", interim_results=True, keyterm=["YourBrand", "custom-term"], eot_threshold=0.5, eot_timeout_ms=3000, eager_eot_threshold=0.3, ) ``` ## Parameters | Parameter | Default | Description | |-----------|---------|-------------| | `model` | `nova-3` | Model to use (`nova-3`, `nova-2`, `flux`) | | `language` | `en` | Language code | | `interim_results` | `True` | Stream partial transcriptions | | `keyterm` | — | Keyword boosting (Nova-3, Flux) | | `keywords` | — | Weighted keyword boosting (Nova-2) | | `eot_threshold` | — | End-of-turn confidence threshold (Flux only) | | `eot_timeout_ms` | — | End-of-turn timeout in ms (Flux only) | | `eager_eot_threshold` | — | Eager end-of-turn threshold (Flux only) | --- ### TTS > Source: https://developers.telnyx.com/docs/livekit/models/tts.md Telnyx offers an extensive library of voices across multiple providers and models, with broad language and accent support. Access them through the Telnyx plugin: ```python from livekit.plugins import telnyx tts = telnyx.TTS(voice="Telnyx.Ultra.Clara") ``` ## Voice ID format Voice IDs follow the pattern `Provider.Model.voice_name`. To find a voice: 1. Browse the [voice library](https://developers.telnyx.com/docs/voice/tts/overview/index) 2. Copy the voice ID (e.g. `Telnyx.Ultra.Clara`) 3. Pass it to `telnyx.TTS(voice="...")` ## Examples ```python # Telnyx Ultra tts = telnyx.TTS(voice="Telnyx.Ultra.Clara") # MiniMax Speech 2.8 Turbo tts = telnyx.TTS(voice="MiniMax.speech-2.8-turbo.Narrator") ``` ## Available providers and models | Provider | Models | |----------|--------| | **Telnyx** | Ultra, KokoroTTS | | **MiniMax** | speech-02-turbo, speech-2.6-turbo, speech-2.8-turbo | | **AWS** | Polly (Neural voices) | | **Azure** | Neural voices | | **Inworld** | Coming soon | | **ResembleAI** | Coming soon | Browse all voices and models in the [voice library →](https://developers.telnyx.com/docs/voice/tts/overview/index) ## Parameters | Parameter | Default | Description | |-----------|---------|-------------| | `voice` | — | Voice ID (e.g. `Telnyx.Ultra.Clara`) | | `sample_rate` | `16000` | Audio sample rate in Hz | --- ### LLM > Source: https://developers.telnyx.com/docs/livekit/models/llm.md Telnyx hosts models with an OpenAI-compatible API. No concurrency limits. Use the standard OpenAI plugin with the `.with_telnyx()` helper: ```python from livekit.plugins import openai llm = openai.LLM.with_telnyx( model="zai-org/GLM-5.3-Flash", reasoning_effort="high", ) ``` ## About `.with_telnyx()` This is a built-in static method on `openai.LLM` in the official [`livekit-plugins-openai`](https://pypi.org/project/livekit-plugins-openai/) package, maintained by LiveKit — not a Telnyx package or fork. It works the same way as the other OpenAI-compatible helpers in that package (`.with_azure()`, `.with_fireworks()`, etc.): it sets `base_url` to Telnyx's OpenAI-compatible inference endpoint (`https://api.telnyx.com/v2/ai/openai`) and reads your `TELNYX_API_KEY` from the environment. You don't need any additional packages beyond `livekit-plugins-openai`. ## Hosted models These run on Telnyx infrastructure — no external API key needed, just your `TELNYX_API_KEY`: | Model | Description | |-------|-------------| | `moonshotai/Kimi-K3` | Moonshot AI — state-of-the-art open-weight intelligence, native vision, 1M context | | `moonshotai/Kimi-K2.6` | Moonshot AI — voice AI, with thinking disabled **(Recommended)** | | `zai-org/GLM-5.3-Flash` | Zhipu AI — most efficient reasoning, function calling | | `MiniMaxAI/MiniMax-M3-MXFP8` | MiniMax — cheapest, high intelligence | ## Proprietary models (BYOK) For models like GPT-4o or Claude, Telnyx proxies the request using your own API key. Add your provider key in the [Telnyx Portal](https://portal.telnyx.com) under Inference settings. | Model | Provider | Description | |-------|----------|-------------| | `openai/gpt-5.4-mini` | OpenAI | Compact high-efficiency model for production voice workflows | | `openai/gpt-4o` | OpenAI | Multimodal flagship model | ```python # Proprietary model via BYOK (bring your own key) llm = openai.LLM.with_telnyx(model="openai/gpt-5.4-mini") ``` [Full models list →](/docs/inference/models) --- ### Observability > Source: https://developers.telnyx.com/docs/livekit/observability.md ## What's available today ### Agent logs Stream logs from your running agent via the CLI: ```bash lk agent logs ``` This gives you stdout/stderr from your agent in real time — same as LiveKit Cloud's log access. --- ### Telephony > Source: https://developers.telnyx.com/docs/livekit/telephony.md Telnyx is the carrier. Buy a number, connect it to your agent — no third-party SIP trunk setup, no FQDN auth dance. Calls route on-net from Telnyx SIP directly to your agent. For setup steps, see the [Quick Start](/docs/livekit/quickstart). ## Supported ### Inbound calls Inbound calls are routed to your agent via SIP dispatch rules. When someone calls your DID, Telnyx forwards it to the LiveKit SIP service, which dispatches it to your agent based on the rules you configure. ### Outbound calls Use the `lk` CLI to place an outbound call into a room: ```bash lk sip participant create \ --room "my-room" \ --trunk "" \ --call "+15551234567" \ --identity "outbound-caller" ``` ### DTMF DTMF tones are supported via RFC 2833/4733. Tones are forwarded to your agent as events and can be handled in code. ### Call transfers #### Cold transfer Use the stable [`TransferSIPParticipant`](https://docs.livekit.io/reference/telephony/sip-api/#transfersipparticipant) API to transfer the caller to another number or SIP endpoint using SIP REFER: ```python from livekit import api await ctx.api.sip.transfer_sip_participant( api.TransferSIPParticipantRequest( room_name=ctx.room.name, participant_identity="", transfer_to="tel:+14155550100", ) ) ``` See LiveKit's [cold transfer guide](https://docs.livekit.io/telephony/features/transfers/cold/) for transfer behavior and timeout handling. #### Warm transfer For a stable manual warm transfer, use [`CreateSIPParticipant`](https://docs.livekit.io/reference/telephony/sip-api/#createsipparticipant) to dial the transferee into a consultation room: ```python from livekit import api await ctx.api.sip.create_sip_participant( api.CreateSIPParticipantRequest( sip_trunk_id="", sip_call_to="+14155550100", room_name="", participant_identity="manager", wait_until_answered=True, ) ) ``` After the consultation, use LiveKit's stable [`MoveParticipant`](https://docs.livekit.io/intro/basics/rooms-participants-tracks/participants/#moveparticipant) API to move the transferee into the caller's room. ### SIP headers Pass call metadata through to your agent using `headers_to_attributes`. Header values are mapped to LiveKit participant attributes, available in your agent at runtime. ### HD voice Telnyx supports two HD voice codecs for higher-quality audio on SIP calls: - **G.722** — Wideband audio at 16 kHz sample rate. Enabled by default on all Telnyx SIP connections. No additional configuration needed. - **Opus** — Wideband audio at 48 kHz sample rate. Requires SRTP encryption on both sides — enable OPUS in your Telnyx SIP connection's inbound codec list and set `media_encryption: ALLOW` on your LiveKit inbound trunk. Delivers the highest audio quality for voice AI agents. Both codecs are negotiated automatically during SIP call setup. If the remote side supports Opus, it will be preferred over G.722. --- ### Architecture > Source: https://developers.telnyx.com/docs/livekit/architecture.md Telnyx LiveKit is a managed platform for deploying voice AI agents at scale. You ship agent code. The platform handles containers, scaling, SIP, and AI inference — colocated in each region. ## Agents Each agent you deploy is an isolated worker. Deploy as many as you need — a restaurant bot, a support agent, a scheduling assistant — each running independently under your account. ``` lk agent deploy . → isolated worker, autoscaled, managed containers ``` Agents are deployed per-account with full namespace isolation. Your rooms, SIP trunks, and dispatch rules are scoped to your API key and never visible to other tenants. ## Autoscaling The platform watches active room load and scales your agent workers up and down automatically. You don't configure `load_fnc` or `load_threshold` — the platform manages this. - **Scale up** — new workers spin up as concurrent calls increase - **Scale down** — idle workers are drained and removed - **Health checks** — unhealthy containers are replaced automatically with rolling deploys ## Inbound call flow ``` Caller → PSTN → Telnyx SIP → LiveKit SIP Bridge → Room → Agent Worker ``` Because Telnyx is both the carrier and the platform, the SIP leg is on-net — no external SIP trunk hop, no PSTN egress to a third party. The call lands directly in a LiveKit room where your agent is waiting. ## AI inference STT, TTS, and LLM inference runs on Telnyx GPUs colocated with the agent runtime — approximately 2ms from your agent. No round-trips to external APIs on the hot path. ``` Agent ──2ms──▶ Telnyx STT / TTS / LLM ``` You can also bring your own provider keys (OpenAI, Anthropic, etc.) — those route externally. ## Regions Each region is a full stack: SFU, SIP, agent runtime, and AI inference. See [Regions](/docs/livekit/regions) for available endpoints. --- ### Regions > Source: https://developers.telnyx.com/docs/livekit/regions.md ## Platform endpoints | Region | Endpoint | |--------|----------| | New York | `https://nyc1.livekit-telnyx.com` | | San Francisco | `https://sfo3.livekit-telnyx.com` | | Atlanta | `https://atl1.livekit-telnyx.com` | | Sydney | `https://syd1.livekit-telnyx.com` | ## Agent deployment endpoints Set `LK_AGENTS_URL` to the matching regional agent endpoint when deploying agents outside your default region. | Region | Agent Endpoint | |--------|----------------| | New York | `https://nyc1.agents.livekit-telnyx.com` | | San Francisco | `https://sfo3.agents.livekit-telnyx.com` | | Atlanta | `https://atl1.agents.livekit-telnyx.com` | | Sydney | `https://syd1.agents.livekit-telnyx.com` | ## SIP endpoints | Region | SIP Endpoint | |--------|-------------| | New York | `nyc1.sip.livekit-telnyx.com` | | San Francisco | `sfo3.sip.livekit-telnyx.com` | | Atlanta | `atl1.sip.livekit-telnyx.com` | | Sydney | `syd1.sip.livekit-telnyx.com` | Choose the region closest to your users for the lowest latency. --- ### Limits > Source: https://developers.telnyx.com/docs/livekit/limits.md ## Build limits | Limit | Value | |-------|-------| | Build timeout | 10 minutes | | Build context size | 1 GB | ## Agent limits | Limit | Value | |-------|-------| | Max concurrent agents | Contact us | | Max concurrent sessions per agent | Contact us | ## Telephony limits | Limit | Value | |-------|-------| | Max concurrent calls | Contact us | For specific limit increases, contact your Telnyx account team. --- ### Pricing > Source: https://developers.telnyx.com/docs/livekit/pricing.md All services — compute, models, and telephony — are billed on a single Telnyx invoice. On other platforms, you pay third-party SIP fees on top of your usage. On Telnyx, those are gone entirely. ## Telnyx Ultra — Platform Exclusive Our highest-fidelity voice. Available exclusively for agents deployed on LiveKit on Telnyx, not through the plugin alone. Head over to the [TTS voice library](https://developers.telnyx.com/docs/voice/tts/overview/index) to check out the variety of voices, languages, and accents we offer. ## STT See [telnyx.com/pricing/speech-to-text](https://telnyx.com/pricing/speech-to-text) for current rates. ## TTS See [telnyx.com/pricing/text-to-speech](https://telnyx.com/pricing/text-to-speech) for current rates. ## LLM See [telnyx.com/pricing/conversational-ai](https://telnyx.com/pricing/conversational-ai) for current rates. --- ## API Reference (AI) ### OpenAI Chat - [Create a chat completion (OpenAI-compatible)](https://developers.telnyx.com/api-reference/openai-chat/create-a-chat-completion-openai-compatible.md): Chat with a language model. This endpoint is consistent with the OpenAI Chat Completions API and may be used with the OpenAI JS or Python SDK by setting the ba… - [Get available models (OpenAI-compatible)](https://developers.telnyx.com/api-reference/openai-chat/get-available-models-openai-compatible.md): Lists every model currently available to your account on Telnyx Inference, including SOTA open-source LLMs hosted on Telnyx GPUs (for example `moonshotai/Kimi-… - [Create an OpenAI-compatible response](https://developers.telnyx.com/api-reference/openai-chat/create-an-openai-compatible-response.md): Create a response using Telnyx's OpenAI-compatible Responses API. This endpoint is compatible with the OpenAI Responses API and may be used with the OpenAI JS… ### Fine Tuning - [List fine tuning jobs](https://developers.telnyx.com/api-reference/fine-tuning/list-fine-tuning-jobs.md): Retrieve a list of all fine tuning jobs created by the user. - [Create a fine tuning job](https://developers.telnyx.com/api-reference/fine-tuning/create-a-fine-tuning-job.md): Creates a new fine-tuning job that trains a model on the provided dataset, and returns the created job. - [Get a fine tuning job](https://developers.telnyx.com/api-reference/fine-tuning/get-a-fine-tuning-job.md): Returns the details of a single fine-tuning job by its job_id, including its current status. - [Cancel a fine tuning job](https://developers.telnyx.com/api-reference/fine-tuning/cancel-a-fine-tuning-job.md): Cancels the specified in-progress fine-tuning job and returns the updated job. ### Anthropic Messages - [Create a message (Anthropic-compatible)](https://developers.telnyx.com/api-reference/anthropic-messages/create-a-message-anthropic-compatible.md): Send a message to a language model using the Anthropic Messages API format. This endpoint is compatible with the Anthropic Messages API and may be used with th… ### Chat - [Summarize file content](https://developers.telnyx.com/api-reference/chat/summarize-file-content.md): Generate a summary of a file's contents. ### Decision Models - [Evaluate decision models (TypeSafe-compatible)](https://developers.telnyx.com/api-reference/decision-models/evaluate-decision-models-typesafe-compatible.md): **Beta API.** Choose telnyx/decision-flash for the lowest cost and latency, or telnyx/decision-pro for decisions that require long context, including inputs be… ### OpenAI Embeddings - [Create embeddings](https://developers.telnyx.com/api-reference/openai-embeddings/create-embeddings.md): Creates an embedding vector representing the input text. This endpoint is compatible with the OpenAI Embeddings API and may be used with the OpenAI JS or Pytho… - [List embedding models](https://developers.telnyx.com/api-reference/openai-embeddings/list-embedding-models.md): Returns a list of available embedding models. This endpoint is compatible with the OpenAI Models API format. ### Embeddings - [Get Tasks by Status](https://developers.telnyx.com/api-reference/embeddings/get-tasks-by-status.md): Retrieve tasks for the user that are either `queued`, `processing`, `failed`, `success` or `partial_success` based on the query string. Defaults to `queued` an… - [Embed documents](https://developers.telnyx.com/api-reference/embeddings/embed-documents.md): Perform embedding on a Telnyx Storage Bucket using the a embedding model. - [List embedded buckets](https://developers.telnyx.com/api-reference/embeddings/list-embedded-buckets.md): Returns the list of storage buckets that have been embedded for your account, for use with similarity search. - [Disable AI for an Embedded Bucket](https://developers.telnyx.com/api-reference/embeddings/disable-ai-for-an-embedded-bucket.md): Deletes an entire bucket's embeddings and disables the bucket for AI-use, returning it to normal storage pricing. - [Get file-level embedding statuses for a bucket](https://developers.telnyx.com/api-reference/embeddings/get-file-level-embedding-statuses-for-a-bucket.md): Get all embedded files for a given user bucket, including their processing status. - [Search for documents](https://developers.telnyx.com/api-reference/embeddings/search-for-documents.md): Perform a similarity search on a Telnyx Storage Bucket, returning the most similar `num_docs` document chunks to the query. - [Embed URL content](https://developers.telnyx.com/api-reference/embeddings/embed-url-content.md): Embed website content from a specified URL, including child pages up to 5 levels deep within the same domain. The process crawls and loads content from the mai… - [Get an embedding task's status](https://developers.telnyx.com/api-reference/embeddings/get-an-embedding-tasks-status.md): Check the status of a current embedding task. Will be one of the following: ### Clusters - [List all clusters](https://developers.telnyx.com/api-reference/clusters/list-all-clusters.md): Retrieve a paginated list of clustering tasks and their statuses. - [Compute new clusters](https://developers.telnyx.com/api-reference/clusters/compute-new-clusters.md): Starts a background task to compute how the data in an embedded storage bucket is clustered. This helps identify common themes and patterns in the data. - [Delete a cluster](https://developers.telnyx.com/api-reference/clusters/delete-a-cluster.md): Delete a clustering task and its computed results. - [Fetch a cluster](https://developers.telnyx.com/api-reference/clusters/fetch-a-cluster.md): Fetch the results of a clustering task, including the discovered clusters. - [Fetch a cluster visualization](https://developers.telnyx.com/api-reference/clusters/fetch-a-cluster-visualization.md): Fetch a visualization image of the clusters computed by a clustering task. ### Conversation Histories - [Search conversation histories](https://developers.telnyx.com/api-reference/conversation-histories/search-conversation-histories.md): Performs semantic vector search across conversation history records. ### Assistants - [List assistants](https://developers.telnyx.com/api-reference/assistants/list-assistants.md): Retrieve a list of all AI Assistants configured by the user. - [Create an assistant](https://developers.telnyx.com/api-reference/assistants/create-an-assistant.md): Creates a new AI assistant from the provided configuration, including its model, instructions, and attached tools, and returns the created assistant. - [Delete an assistant](https://developers.telnyx.com/api-reference/assistants/delete-an-assistant.md): Delete an AI Assistant by `assistant_id`. - [Get an assistant](https://developers.telnyx.com/api-reference/assistants/get-an-assistant.md): Retrieve an AI Assistant configuration by `assistant_id`. - [Update an assistant](https://developers.telnyx.com/api-reference/assistants/update-an-assistant.md): Updates the specified AI assistant's attributes and returns the updated assistant. The request can also control how the change is promoted across assistant ver… - [Assistant Chat](https://developers.telnyx.com/api-reference/assistants/assistant-chat.md): This endpoint allows a client to send a chat message to a specific AI Assistant. The assistant processes the message and returns a relevant reply based on the… - [Assistant Sms Chat](https://developers.telnyx.com/api-reference/assistants/assistant-sms-chat.md): Send an SMS message for an assistant. This endpoint: - [Clone Assistant](https://developers.telnyx.com/api-reference/assistants/clone-assistant.md): Clone an existing assistant, excluding telephony and messaging settings. - [Enhance Assistant Instructions](https://developers.telnyx.com/api-reference/assistants/enhance-assistant-instructions.md): Enhance an assistant's instructions using an LLM. The endpoint reads the assistant's current instructions and tools, then streams back improved instructions as… - [Import assistants from external provider](https://developers.telnyx.com/api-reference/assistants/import-assistants-from-external-provider.md): Import assistants from external providers. Any assistant that has already been imported will be overwritten with its latest version from the importing provider. - [List scheduled events](https://developers.telnyx.com/api-reference/assistants/list-scheduled-events.md): Get scheduled events for an assistant with pagination and filtering - [Create a scheduled event](https://developers.telnyx.com/api-reference/assistants/create-a-scheduled-event.md): Create a scheduled event for an assistant - [Delete a scheduled event](https://developers.telnyx.com/api-reference/assistants/delete-a-scheduled-event.md): If the event is pending, this will cancel the event. Otherwise, this will simply remove the record of the event. - [Get a scheduled event](https://developers.telnyx.com/api-reference/assistants/get-a-scheduled-event.md): Returns the details of a single scheduled event configured for the specified assistant. - [List assistant tests with pagination](https://developers.telnyx.com/api-reference/assistants/list-assistant-tests-with-pagination.md): Retrieves a paginated list of assistant tests with optional filtering capabilities - [Create a new assistant test](https://developers.telnyx.com/api-reference/assistants/create-a-new-assistant-test.md): Creates a comprehensive test configuration for evaluating AI assistant performance - [Get all test suite names](https://developers.telnyx.com/api-reference/assistants/get-all-test-suite-names.md): Retrieves a list of all distinct test suite names available to the current user - [Get test suite run history](https://developers.telnyx.com/api-reference/assistants/get-test-suite-run-history.md): Retrieves paginated history of test runs for a specific test suite with filtering options - [Trigger test suite execution](https://developers.telnyx.com/api-reference/assistants/trigger-test-suite-execution.md): Executes all tests within a specific test suite as a batch operation - [Delete an assistant test](https://developers.telnyx.com/api-reference/assistants/delete-an-assistant-test.md): Permanently removes an assistant test and all associated data - [Get assistant test by ID](https://developers.telnyx.com/api-reference/assistants/get-assistant-test-by-id.md): Retrieves detailed information about a specific assistant test - [Update an assistant test](https://developers.telnyx.com/api-reference/assistants/update-an-assistant-test.md): Updates an existing assistant test configuration with new settings - [Get test run history for a specific test](https://developers.telnyx.com/api-reference/assistants/get-test-run-history-for-a-specific-test.md): Retrieves paginated execution history for a specific assistant test with filtering options - [Trigger a manual test run](https://developers.telnyx.com/api-reference/assistants/trigger-a-manual-test-run.md): Initiates immediate execution of a specific assistant test - [Get specific test run details](https://developers.telnyx.com/api-reference/assistants/get-specific-test-run-details.md): Retrieves detailed information about a specific test run execution - [Get all versions of an assistant](https://developers.telnyx.com/api-reference/assistants/get-all-versions-of-an-assistant.md): Retrieves all versions of a specific assistant with complete configuration and metadata - [Delete a specific assistant version](https://developers.telnyx.com/api-reference/assistants/delete-a-specific-assistant-version.md): Permanently removes a specific version of an assistant. Can not delete main version - [Get a specific assistant version](https://developers.telnyx.com/api-reference/assistants/get-a-specific-assistant-version.md): Retrieves a specific version of an assistant by assistant_id and version_id - [Update a specific assistant version](https://developers.telnyx.com/api-reference/assistants/update-a-specific-assistant-version.md): Updates the configuration of a specific assistant version. Can not update main version - [Promote an assistant version to main](https://developers.telnyx.com/api-reference/assistants/promote-an-assistant-version-to-main.md): Promotes a specific version to be the main/current version of the assistant. This will delete any existing canary deploy configuration and send all live produc… - [Delete Canary Deploy](https://developers.telnyx.com/api-reference/assistants/delete-canary-deploy.md): Endpoint to delete a canary deploy configuration for an assistant. - [Get Canary Deploy](https://developers.telnyx.com/api-reference/assistants/get-canary-deploy.md): Endpoint to get a canary deploy configuration for an assistant. - [Create Canary Deploy](https://developers.telnyx.com/api-reference/assistants/create-canary-deploy.md): Endpoint to create a canary deploy configuration for an assistant. - [Update Canary Deploy](https://developers.telnyx.com/api-reference/assistants/update-canary-deploy.md): Endpoint to update a canary deploy configuration for an assistant. - [Get assistant texml](https://developers.telnyx.com/api-reference/assistants/get-assistant-texml.md): Get an assistant texml by `assistant_id`. - [Test Assistant Tool](https://developers.telnyx.com/api-reference/assistants/test-assistant-tool.md): Executes a test invocation of the specified webhook tool for the assistant and returns the outcome, so you can verify the webhook's behavior before relying on… - [Get All Tags](https://developers.telnyx.com/api-reference/assistants/get-all-tags.md): Retrieve all tags that have been applied to your AI assistants. - [Add Assistant Tag](https://developers.telnyx.com/api-reference/assistants/add-assistant-tag.md): Add a tag to an AI assistant. Tags help you organize and filter your assistants. - [Remove Assistant Tag](https://developers.telnyx.com/api-reference/assistants/remove-assistant-tag.md): Removes the specified tag from the AI assistant and returns the assistant's updated tag list. - [Remove Assistant Tool](https://developers.telnyx.com/api-reference/assistants/remove-assistant-tool.md): Detaches the specified tool from the AI assistant so the assistant can no longer invoke it. - [Add Assistant Tool](https://developers.telnyx.com/api-reference/assistants/add-assistant-tool.md): Attach an existing tool to an AI assistant. ### Integrations - [List Integrations](https://developers.telnyx.com/api-reference/integrations/list-integrations.md): Returns the list of third-party integrations available to connect to your AI assistants and workflows. - [List User Integrations](https://developers.telnyx.com/api-reference/integrations/list-user-integrations.md): Returns the list of integration connections you have set up, linking your account to third-party services. - [Delete Integration Connection](https://developers.telnyx.com/api-reference/integrations/delete-integration-connection.md): Delete a specific integration connection. - [Get User Integration connection By Id](https://developers.telnyx.com/api-reference/integrations/get-user-integration-connection-by-id.md): Returns the details of a single integration connection by its ID. - [List Integration By Id](https://developers.telnyx.com/api-reference/integrations/list-integration-by-id.md): Returns the details of a single available integration, including its configuration details. ### MCP Servers - [List MCP Servers](https://developers.telnyx.com/api-reference/mcp-servers/list-mcp-servers.md): Returns a paginated list of the MCP servers configured on your account, with optional filtering by type or URL. - [Create MCP Server](https://developers.telnyx.com/api-reference/mcp-servers/create-mcp-server.md): Creates a new MCP server configuration on your account and returns the created server. - [Delete MCP Server](https://developers.telnyx.com/api-reference/mcp-servers/delete-mcp-server.md): Permanently deletes the specified MCP server configuration from your account. - [Get MCP Server](https://developers.telnyx.com/api-reference/mcp-servers/get-mcp-server.md): Retrieve details for a specific MCP server. - [Update MCP Server](https://developers.telnyx.com/api-reference/mcp-servers/update-mcp-server.md): Updates the specified MCP server's configuration and returns the updated server. ### Conversations - [List conversations](https://developers.telnyx.com/api-reference/conversations/list-conversations.md): Retrieve a list of all AI conversations configured by the user. Supports PostgREST-style query parameters for filtering. Examples are included for the standard… - [Create a conversation](https://developers.telnyx.com/api-reference/conversations/create-a-conversation.md): Creates a new AI conversation, the container for messages exchanged with an assistant, and returns the created conversation. - [Aggregate Conversation Insights](https://developers.telnyx.com/api-reference/conversations/aggregate-conversation-insights.md): Aggregate conversation insights by specified fields - [Get Insight Template Groups](https://developers.telnyx.com/api-reference/conversations/get-insight-template-groups.md): Returns a paginated list of your insight template groups. Groups organize related insight templates that are applied together when analyzing conversations. - [Create Insight Template Group](https://developers.telnyx.com/api-reference/conversations/create-insight-template-group.md): Creates a new insight template group for organizing related insight templates, and returns the created group. - [Delete Insight Template Group](https://developers.telnyx.com/api-reference/conversations/delete-insight-template-group.md): Permanently deletes the specified insight template group by its ID. - [Get Insight Template Group](https://developers.telnyx.com/api-reference/conversations/get-insight-template-group.md): Returns the details of a single insight template group, including the insight templates assigned to it. - [Update Insight Template Group](https://developers.telnyx.com/api-reference/conversations/update-insight-template-group.md): Updates the specified insight template group and returns the updated group. - [Assign Insight Template To Group](https://developers.telnyx.com/api-reference/conversations/assign-insight-template-to-group.md): Assigns the specified insight template to the specified insight template group. - [Unassign Insight Template From Group](https://developers.telnyx.com/api-reference/conversations/unassign-insight-template-from-group.md): Removes the specified insight template from the specified group. The insight template itself is not deleted. - [Get Insight Templates](https://developers.telnyx.com/api-reference/conversations/get-insight-templates.md): Returns a paginated list of your insight templates. Insight templates define analyses that run over AI conversations to extract structured findings. - [Create Insight Template](https://developers.telnyx.com/api-reference/conversations/create-insight-template.md): Creates a new insight template defining an analysis to run over conversations, and returns the created template. - [Delete Insight Template](https://developers.telnyx.com/api-reference/conversations/delete-insight-template.md): Permanently deletes the specified insight template by its ID. - [Get Insight Template](https://developers.telnyx.com/api-reference/conversations/get-insight-template.md): Returns the details of a single insight template by its ID, including its configuration. - [Update Insight Template](https://developers.telnyx.com/api-reference/conversations/update-insight-template.md): Updates the specified insight template and returns the updated template. - [Delete a conversation](https://developers.telnyx.com/api-reference/conversations/delete-a-conversation.md): Delete a specific conversation by its ID. - [Get a conversation](https://developers.telnyx.com/api-reference/conversations/get-a-conversation.md): Retrieve a specific AI conversation by its ID. - [Update conversation metadata](https://developers.telnyx.com/api-reference/conversations/update-conversation-metadata.md): Update metadata for a specific conversation. - [Get insights for a conversation](https://developers.telnyx.com/api-reference/conversations/get-insights-for-a-conversation.md): Retrieve insights for a specific conversation - [Create Message](https://developers.telnyx.com/api-reference/conversations/create-message.md): Add a new message to the conversation. Used to insert a new messages to a conversation manually ( without using chat endpoint ) - [Get conversation messages](https://developers.telnyx.com/api-reference/conversations/get-conversation-messages.md): Retrieve messages for a specific conversation, including tool calls made by the assistant. ### Missions - [List missions](https://developers.telnyx.com/api-reference/missions/list-missions.md): Returns a paginated list of all mission definitions in your organization. Missions describe a goal and the tools, knowledge bases, and MCP servers agents may u… - [Create mission](https://developers.telnyx.com/api-reference/missions/create-mission.md): Creates a new mission definition from the provided configuration and returns the created mission. Execute the mission by starting runs against it. - [List recent events](https://developers.telnyx.com/api-reference/missions/list-recent-events.md): Returns a paginated list of recent events across every mission in your organization, optionally filtered by event type. Useful for building activity feeds or m… - [List recent runs](https://developers.telnyx.com/api-reference/missions/list-recent-runs.md): Returns a paginated list of recent runs across every mission in your organization, optionally filtered by run status. Useful for monitoring overall mission act… - [Delete mission](https://developers.telnyx.com/api-reference/missions/delete-mission.md): Permanently deletes the specified mission definition and returns no content on success. - [Get mission](https://developers.telnyx.com/api-reference/missions/get-mission.md): Get a mission by ID (includes tools, knowledge_bases, mcp_servers) - [Update mission](https://developers.telnyx.com/api-reference/missions/update-mission.md): Replaces the specified mission's definition with the provided configuration and returns the updated mission. - [Clone mission](https://developers.telnyx.com/api-reference/missions/clone-mission.md): Creates a copy of the specified mission as a new mission definition, so you can iterate on its configuration without modifying the original. - [List runs for mission](https://developers.telnyx.com/api-reference/missions/list-runs-for-mission.md): Returns a paginated list of runs for the specified mission, optionally filtered by run status, so you can track the mission's execution history over time. - [Start a run](https://developers.telnyx.com/api-reference/missions/start-a-run.md): Starts a new run of the specified mission and returns the created run object. Track its progress through the run detail, plan, and events endpoints. - [Get run details](https://developers.telnyx.com/api-reference/missions/get-run-details.md): Returns the full details of a single run, including its current status. Use this to poll an in-flight run or inspect the outcome of a completed one. - [Update run](https://developers.telnyx.com/api-reference/missions/update-run.md): Updates a run's status and/or result and returns the updated run object. Typically used by executing agents to report progress or record the final outcome. - [Cancel run](https://developers.telnyx.com/api-reference/missions/cancel-run.md): Cancels a running or paused run and returns the updated run object. A cancelled run stops executing; start a new run to execute the mission again. - [List events](https://developers.telnyx.com/api-reference/missions/list-events.md): Returns a paginated list of events logged for the specified run, filterable by event type, plan step, and agent, so you can reconstruct exactly what happened d… - [Log event](https://developers.telnyx.com/api-reference/missions/log-event.md): Logs a new event against the specified run and returns the created event. Events form the run's audit trail and can reference a plan step or agent. - [Get event details](https://developers.telnyx.com/api-reference/missions/get-event-details.md): Returns the details of a single event logged for the specified run, including its type and payload. - [Pause run](https://developers.telnyx.com/api-reference/missions/pause-run.md): Pauses a currently running run and returns the updated run object. Execution halts until the run is resumed. - [Get plan](https://developers.telnyx.com/api-reference/missions/get-plan.md): Returns the plan for the specified run, including all plan steps and their statuses, so you can see how the mission was decomposed and how far execution has pr… - [Create initial plan](https://developers.telnyx.com/api-reference/missions/create-initial-plan.md): Creates the initial plan for the specified run from the provided steps and returns the created plan steps. Progress is subsequently reported by updating indivi… - [Add step(s) to plan](https://developers.telnyx.com/api-reference/missions/add-steps-to-plan.md): Add one or more steps to an existing plan - [Get step details](https://developers.telnyx.com/api-reference/missions/get-step-details.md): Returns the details of a single plan step within a run's plan, including its status. - [Update step status](https://developers.telnyx.com/api-reference/missions/update-step-status.md): Updates the status of a single plan step and returns the updated step. Typically called by the executing agent as it works through the plan. - [Resume run](https://developers.telnyx.com/api-reference/missions/resume-run.md): Resumes a previously paused run and returns the updated run object, letting execution continue from where it was paused. - [List linked Telnyx agents](https://developers.telnyx.com/api-reference/missions/list-linked-telnyx-agents.md): Returns the Telnyx agents currently linked to the specified run. Linked agents participate in executing the run's plan. - [Link Telnyx agent to run](https://developers.telnyx.com/api-reference/missions/link-telnyx-agent-to-run.md): Link a Telnyx AI agent (voice/messaging) to a run - [Unlink Telnyx agent](https://developers.telnyx.com/api-reference/missions/unlink-telnyx-agent.md): Unlinks the specified Telnyx agent from the run so it no longer participates in execution. The run itself and its history are unaffected. - [List knowledge bases](https://developers.telnyx.com/api-reference/missions/list-knowledge-bases.md): Returns the knowledge bases attached to the specified mission. Knowledge bases provide reference content agents can draw on during runs. - [Create knowledge base](https://developers.telnyx.com/api-reference/missions/create-knowledge-base.md): Create a new knowledge base for a mission - [Delete knowledge base](https://developers.telnyx.com/api-reference/missions/delete-knowledge-base.md): Detaches the specified knowledge base from the mission so its content is no longer available to agents in subsequent runs. - [Get knowledge base](https://developers.telnyx.com/api-reference/missions/get-knowledge-base.md): Returns the details of a single knowledge base attached to the specified mission. - [Update knowledge base](https://developers.telnyx.com/api-reference/missions/update-knowledge-base.md): Replaces the definition of the specified knowledge base on this mission. - [List MCP servers](https://developers.telnyx.com/api-reference/missions/list-mcp-servers.md): Returns the MCP servers configured on the specified mission. MCP servers expose external tools and data sources agents can use during runs. - [Create MCP server](https://developers.telnyx.com/api-reference/missions/create-mcp-server.md): Adds an MCP server to the specified mission, making the server's tools available to agents during runs of this mission. - [Delete MCP server](https://developers.telnyx.com/api-reference/missions/delete-mcp-server.md): Removes the specified MCP server from the mission, revoking agent access to its tools in subsequent runs. - [Get MCP server](https://developers.telnyx.com/api-reference/missions/get-mcp-server.md): Returns the configuration of a single MCP server attached to the specified mission. - [Update MCP server](https://developers.telnyx.com/api-reference/missions/update-mcp-server.md): Replaces the configuration of the specified MCP server on this mission. - [List tools](https://developers.telnyx.com/api-reference/missions/list-tools.md): Returns the tools configured on the specified mission. Tools define the actions agents may invoke while executing the mission's runs. - [Create tool](https://developers.telnyx.com/api-reference/missions/create-tool.md): Adds a new tool to the specified mission, defining an action agents can invoke during runs of this mission. - [Delete tool](https://developers.telnyx.com/api-reference/missions/delete-tool.md): Removes the specified tool from the mission so agents can no longer invoke it in subsequent runs. - [Get tool](https://developers.telnyx.com/api-reference/missions/get-tool.md): Returns the definition of a single tool configured on the specified mission. - [Update tool](https://developers.telnyx.com/api-reference/missions/update-tool.md): Replaces the definition of the specified tool on this mission.