Skip to main content
Use the Spend Limits API to set an inference budget in USD for the current UTC day, calendar month, or both. Limits apply to the authenticated user’s organization, or their own account if they do not belong to an organization. Limits are shared across the organization. Reading or changing them requires the caller’s access policy to permit the corresponding Spend Limits operation. An API key without the required permission receives HTTP 403, even when the organization has shared limits. This authorization error is separate from the HTTP 403 returned by inference requests when a spending limit blocks usage.

Which usage counts

The limit is shared across the account’s billable inference usage, including LLM requests made by voice assistants. It is not a separate budget for each assistant, model, or API key. BYOK refers to the third-party provider key, not the Telnyx API key used to authenticate a request. Supplying a provider key does not exempt usage of a Telnyx-hosted model. BYOK and customer-endpoint requests remain available when billable inference is blocked. If they fall back to a billable model, the fallback is checked against the account’s block state. This limit covers costs recorded under the inference billing product, not every charge associated with a voice call. A voice assistant using BYOK can still incur other Telnyx charges.

Understand enforcement

The daily and monthly limits are independent. Exceeding either limit blocks new billable Chat Completions, Responses, Anthropic Messages, and classification requests. Requests already running finish normally. Voice-assistant LLM requests are subject to the same enforcement. An in-flight model request can finish, but a subsequent request or conversation turn can be refused after the block takes effect.
  • Spend must be strictly greater than the limit to trigger a block. A zero limit blocks at the first cent of spend; it does not disable inference before any spend occurs.
  • Daily periods reset at 00:00 UTC the next day. Monthly periods reset at 00:00 UTC on the first day of the next month.
  • Spend can lag actual usage by several minutes. After spend is recorded, a block appears within about two minutes for daily limits or ten minutes for monthly limits. Usage can exceed the configured amount before enforcement takes effect.
  • Creating or changing a limit checks existing spend immediately. Setting a limit below current spend can block inference at once.

Monitor spend in the portal

Use the Inference dashboard to investigate LLM token usage and reported cost by model and use case, including voice-assistant inference. BYOK can produce token activity without a corresponding Telnyx LLM charge. Use Billing → Spending limits to view account-wide spend against the daily and monthly caps and edit limits across supported products.
The Billing limit counters use the current UTC day and calendar month and include the whole inference billing product. The dashboard charts show LLM records and honor the selected dates, timezone, models, and use cases. Their totals can differ from the limit counters. Use spend_usd and blocked from GET /v2/spend_limits, also shown in the limit controls, to inspect the spend and block state used for enforcement. Reporting and enforcement can lag usage.

Read current limits and spend

Set TELNYX_API_KEY to an API key for the account being managed. Call List spend limits:
The unpaginated data array includes each supported product and period, even when no limit is set. Select entries with product: "inference"; use the returned products rather than assuming future product support. A limit with origin: "operator" was set by Telnyx support. It can be updated or deleted through the same API. limit: null and a configured limit with unlimited: true both mean no cap for inference, but only the latter is an existing resource that can be updated.

Create a limit

When the selected entry has limit: null, call Create a spend limit. This example sets a daily limit of USD 100:
To also set a monthly limit, send a separate request with this body:
Send amount as a nonnegative JSON number. Response amounts are decimal strings. The optional audit reason accepts up to 500 characters. Send the period explicitly; the default is daily. Creation returns HTTP 409 if a limit already exists. Re-read the limits and update the existing resource instead.

Change an existing limit

Call Update a spend limit. The product is a path parameter and the period is a query parameter, not a body field. This example raises the daily limit to USD 250:
To keep an existing monthly resource but explicitly remove its cap, send {"unlimited":true} to PATCH /v2/spend_limits/inference?period=monthly. Send exactly one of amount and unlimited: true. Unknown body fields are rejected with HTTP 400. An update returns HTTP 404 if no limit exists; re-read the limits before creating one.

Remove a limit

Call Delete a spend limit for the intended period:
Inference has no default limit. Deletion returns limit: null, makes that period unlimited, and lifts its block. It does not remove the other period’s limit or block. Deleting a missing limit returns HTTP 404.

Verify a change and recover blocked requests

Create, update, and delete responses include data.evaluation:
Write responses always carry blocked: false. Use evaluation for the immediate result, then call GET /v2/spend_limits again to read the current block state for both periods. Do not treat a successful write or evaluation.released alone as confirmation that inference is unblocked.
Blocked inference requests return HTTP 403 with error code 10039 and title Inference spend limit reached. Stop automatic retries for this error. Inspect both limits, then raise the blocking limits above current spend, remove them, or wait until their UTC periods end. Verify both periods are unblocked before resuming requests. Other HTTP 403 errors can have different causes.