Create an OpenAI-compatible response
Create a response using Telnyx’s OpenAI-compatible Responses API. This endpoint is compatible with the OpenAI Responses API and may be used with the OpenAI JS or Python SDK by setting the base URL to https://api.telnyx.com/v2/ai/openai.
The conversation parameter refers to a Telnyx Conversation rather than an OpenAI-hosted conversation object. To persist a thread across turns, first create a conversation with POST /ai/conversations, then pass that conversation’s id in the Responses request as conversation. The endpoint appends the new input, assistant output, reasoning, and tool-call messages to that conversation. Reuse the same conversation id on subsequent Responses requests, including tool-result followups, so the model receives the prior context.
If conversation is omitted, the request is processed without persisting messages to a Telnyx conversation. Use the Conversations API to manage history: list conversations (optionally filtered by metadata), fetch messages for a conversation, and optionally add messages outside the Responses flow.
You can attach arbitrary metadata when creating a conversation (for example to tag the conversation’s source, channel, or user) and later filter by it when listing conversations.
Authorizations
Bearer authentication header of the form Bearer <token>, where <token> is your auth token.
Body
Model identifier to use for the response, for example zai-org/GLM-5.1-FP8 or another model available from the Telnyx OpenAI-compatible models endpoint.
The service tier to use for this request. Supported values vary by model; use GET /v2/ai/openai/models and inspect the model's service_tiers field. If omitted, Telnyx-hosted models use default.
Optional data-residency region the request should be served from, using the same vocabulary as your account's Data Locality setting. Behavior depends on mode. Supported for Telnyx-hosted models only: a request routed to an external provider never passes through Telnyx model routing, so a region cannot be enforced for it. Omit for today's latency-based routing.
USA, EU, AUS, UAE How strictly region is applied. preferred (the default when region is set) tries that region first and falls back to another when the model cannot be served there, so a request that would have succeeded still succeeds. strict pins the request: it is served from that region or it fails with a 422, never redirected to another region. Requires region.
preferred, strict The input items for this turn, using the OpenAI Responses API input format.
Optional Telnyx Conversation ID from POST /ai/conversations. When provided, Telnyx stores this turn on that conversation and uses the conversation's prior messages as context. Reuse the same ID for subsequent turns and tool-result followups. Omit it for a non-persisted, stateless response.
Optional system/developer instructions for the model. When used with a persisted conversation, send these on the first request that creates the thread; subsequent turns can rely on the stored history.
Set to true to stream Server-Sent Events, matching OpenAI's Responses streaming format.
Response
Successful Response
The response is of type object.