Which usage counts
The limit is shared across the account’s billable inference usage, including LLM requests made by voice assistants. It is not a separate budget for each assistant, model, or API key.
BYOK refers to the third-party provider key, not the Telnyx API key used to
authenticate a request. Supplying a provider key does not exempt usage of a
Telnyx-hosted model. BYOK and customer-endpoint requests remain available when
billable inference is blocked. If they fall back to a billable model, the
fallback is checked against the account’s block state.
This limit covers costs recorded under the inference billing product, not
every charge associated with a voice call. A voice assistant using BYOK can
still incur other Telnyx charges.
Understand enforcement
The daily and monthly limits are independent. Exceeding either limit blocks new billable Chat Completions, Responses, Anthropic Messages, and classification requests. Requests already running finish normally. Voice-assistant LLM requests are subject to the same enforcement. An in-flight model request can finish, but a subsequent request or conversation turn can be refused after the block takes effect.- Spend must be strictly greater than the limit to trigger a block. A zero limit blocks at the first cent of spend; it does not disable inference before any spend occurs.
- Daily periods reset at 00:00 UTC the next day. Monthly periods reset at 00:00 UTC on the first day of the next month.
- Spend can lag actual usage by several minutes. After spend is recorded, a block appears within about two minutes for daily limits or ten minutes for monthly limits. Usage can exceed the configured amount before enforcement takes effect.
- Creating or changing a limit checks existing spend immediately. Setting a limit below current spend can block inference at once.
Monitor spend in the portal
Use the Inference dashboard to investigate LLM token usage and reported cost by model and use case, including voice-assistant inference. BYOK can produce token activity without a corresponding Telnyx LLM charge. Use Billing → Spending limits to view account-wide spend against the daily and monthly caps and edit limits across supported products.The Billing limit counters use the current UTC day and calendar month and include
the whole inference billing product. The dashboard charts show LLM records
and honor the selected dates, timezone, models, and use cases. Their totals
can differ from the limit counters. Use
spend_usd and blocked from
GET /v2/spend_limits, also shown in the limit controls, to inspect the spend
and block state used for enforcement. Reporting and enforcement can lag usage.Read current limits and spend
SetTELNYX_API_KEY to an API key for the account being managed. Call
List spend limits:
data array includes each supported product and period, even
when no limit is set. Select entries with product: "inference"; use the returned
products rather than assuming future product support.
A limit with
origin: "operator" was set by Telnyx support. It can be updated or
deleted through the same API. limit: null and a configured limit with
unlimited: true both mean no cap for inference, but only the latter is an
existing resource that can be updated.
Create a limit
When the selected entry haslimit: null, call
Create a spend limit.
This example sets a daily limit of USD 100:
amount as a nonnegative JSON number. Response amounts are decimal
strings. The optional audit reason accepts up to 500 characters. Send the
period explicitly; the default is daily.
Creation returns HTTP 409 if a limit already exists. Re-read the limits and
update the existing resource instead.
Change an existing limit
Call Update a spend limit. The product is a path parameter and the period is a query parameter, not a body field. This example raises the daily limit to USD 250:{"unlimited":true} to PATCH /v2/spend_limits/inference?period=monthly.
Send exactly one of amount and unlimited: true. Unknown body fields are
rejected with HTTP 400. An update returns HTTP 404 if no limit exists; re-read
the limits before creating one.
Remove a limit
Call Delete a spend limit for the intended period:limit: null, makes that
period unlimited, and lifts its block. It does not remove the other period’s
limit or block. Deleting a missing limit returns HTTP 404.
Verify a change and recover blocked requests
Create, update, and delete responses includedata.evaluation:
Blocked inference requests return HTTP 403 with error code
10039 and title
Inference spend limit reached. Stop automatic retries for this error. Inspect
both limits, then raise the blocking limits above current spend, remove them,
or wait until their UTC periods end. Verify both periods are unblocked before
resuming requests. Other HTTP 403 errors can have different causes.