> ## Documentation Index
> Fetch the complete documentation index at: https://developers.telnyx.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Set inference spending limits

> Create daily and monthly inference spending limits, inspect usage and blocks, and change or remove limits through the Telnyx API.

Use the Spend Limits API to set an inference budget in USD for the current UTC
day, calendar month, or both. Limits apply to the authenticated user's
organization, or their own account if they do not belong to an organization.
Limits are shared across the organization. Reading or changing them requires
the caller's access policy to permit the corresponding Spend Limits operation.
An API key without the required permission receives HTTP 403, even when the
organization has shared limits. This authorization error is separate from the
HTTP 403 returned by inference requests when a spending limit blocks usage.

## Which usage counts

The limit is shared across the account's billable inference usage, including
LLM requests made by voice assistants. It is not a separate budget for each
assistant, model, or API key.

| LLM request                                                                   | Counts toward the inference limit |
| ----------------------------------------------------------------------------- | --------------------------------- |
| Telnyx-hosted model, called directly or by a voice assistant                  | Yes                               |
| OpenAI or another third-party model using Telnyx-managed provider credentials | Yes                               |
| Third-party model using the customer's own provider key (BYOK)                | No                                |
| Customer-supplied model endpoint                                              | No                                |

BYOK refers to the third-party provider key, not the Telnyx API key used to
authenticate a request. Supplying a provider key does not exempt usage of a
Telnyx-hosted model. BYOK and customer-endpoint requests remain available when
billable inference is blocked. If they fall back to a billable model, the
fallback is checked against the account's block state.

This limit covers costs recorded under the inference billing product, not
every charge associated with a voice call. A voice assistant using BYOK can
still incur other Telnyx charges.

## Understand enforcement

The daily and monthly limits are independent. Exceeding either limit blocks new
billable Chat Completions, Responses, Anthropic Messages, and classification
requests. Requests already running finish normally.

Voice-assistant LLM requests are subject to the same enforcement. An in-flight
model request can finish, but a subsequent request or conversation turn can
be refused after the block takes effect.

* Spend must be **strictly greater** than the limit to trigger a block. A zero
  limit blocks at the first cent of spend; it does not disable inference before
  any spend occurs.
* Daily periods reset at 00:00 UTC the next day. Monthly periods reset at
  00:00 UTC on the first day of the next month.
* Spend can lag actual usage by several minutes. After spend is recorded, a block
  appears within about two minutes for daily limits or ten minutes for monthly
  limits. Usage can exceed the configured amount before enforcement takes effect.
* Creating or changing a limit checks existing spend immediately. Setting a
  limit below current spend can block inference at once.

## Monitor spend in the portal

Use the [Inference dashboard](https://portal.telnyx.com/#/ai/reports/dashboard?product=inference)
to investigate LLM token usage and reported cost by model and use case,
including voice-assistant inference. BYOK can produce token activity without
a corresponding Telnyx LLM charge.

Use [Billing → Spending limits](https://portal.telnyx.com/#/billing/spending-limits)
to view account-wide spend against the daily and monthly caps and edit limits
across supported products.

<Note>
  The Billing limit counters use the current UTC day and calendar month and include
  the whole inference billing product. The dashboard charts show LLM records
  and honor the selected dates, timezone, models, and use cases. Their totals
  can differ from the limit counters. Use `spend_usd` and `blocked` from
  `GET /v2/spend_limits`, also shown in the limit controls, to inspect the spend
  and block state used for enforcement. Reporting and enforcement can lag usage.
</Note>

## Read current limits and spend

Set `TELNYX_API_KEY` to an API key for the account being managed. Call
[List spend limits](/api-reference/spend-limits/list-spend-limits):

```bash theme={null}
curl https://api.telnyx.com/v2/spend_limits \
  -H "Authorization: Bearer $TELNYX_API_KEY"
```

The unpaginated `data` array includes each supported product and period, even
when no limit is set. Select entries with `product: "inference"`; use the returned
products rather than assuming future product support.

| Field                        | Interpretation                                                               |
| ---------------------------- | ---------------------------------------------------------------------------- |
| `period`                     | `daily` or `monthly`                                                         |
| `period_start`, `period_end` | UTC dates; `period_end` is exclusive                                         |
| `limit`                      | Configured limit, or `null` when none exists                                 |
| `effective_limit_usd`        | Enforced amount as a decimal string; `null` means unlimited                  |
| `spend_usd`                  | Current period spend as a decimal string; `null` means unavailable, not zero |
| `spend_error`                | Set when spend cannot be read                                                |
| `blocked`, `block`           | Current period's block state and details, including `blocked_until`          |

A limit with `origin: "operator"` was set by Telnyx support. It can be updated or
deleted through the same API. `limit: null` and a configured limit with
`unlimited: true` both mean no cap for inference, but only the latter is an
existing resource that can be updated.

## Create a limit

When the selected entry has `limit: null`, call
[Create a spend limit](/api-reference/spend-limits/create-a-spend-limit).
This example sets a daily limit of USD 100:

```bash theme={null}
curl -X POST https://api.telnyx.com/v2/spend_limits \
  -H "Authorization: Bearer $TELNYX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"product":"inference","period":"daily","amount":100,"reason":"Team budget"}'
```

To also set a monthly limit, send a separate request with this body:

```json theme={null}
{
  "product": "inference",
  "period": "monthly",
  "amount": 2000,
  "reason": "Monthly inference budget"
}
```

Send `amount` as a nonnegative JSON number. Response amounts are decimal
strings. The optional audit `reason` accepts up to 500 characters. Send the
period explicitly; the default is `daily`.

Creation returns HTTP 409 if a limit already exists. Re-read the limits and
update the existing resource instead.

## Change an existing limit

Call [Update a spend limit](/api-reference/spend-limits/update-a-spend-limit).
The product is a path parameter and the period is a query parameter, not a body
field. This example raises the daily limit to USD 250:

```bash theme={null}
curl -X PATCH 'https://api.telnyx.com/v2/spend_limits/inference?period=daily' \
  -H "Authorization: Bearer $TELNYX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"amount":250,"reason":"Raised for launch traffic"}'
```

To keep an existing monthly resource but explicitly remove its cap, send
`{"unlimited":true}` to `PATCH /v2/spend_limits/inference?period=monthly`.
Send exactly one of `amount` and `unlimited: true`. Unknown body fields are
rejected with HTTP 400. An update returns HTTP 404 if no limit exists; re-read
the limits before creating one.

## Remove a limit

Call [Delete a spend limit](/api-reference/spend-limits/delete-a-spend-limit)
for the intended period:

```bash theme={null}
curl -X DELETE 'https://api.telnyx.com/v2/spend_limits/inference?period=daily' \
  -H "Authorization: Bearer $TELNYX_API_KEY"
```

Inference has no default limit. Deletion returns `limit: null`, makes that
period unlimited, and lifts its block. It does not remove the other period's
limit or block. Deleting a missing limit returns HTTP 404.

## Verify a change and recover blocked requests

Create, update, and delete responses include `data.evaluation`:

| Field                        | Meaning                                                                                       |
| ---------------------------- | --------------------------------------------------------------------------------------------- |
| `blocked_now`                | Existing spend exceeded the new limit and inference was blocked                               |
| `released`                   | This period's block was lifted                                                                |
| `still_over_limit`           | This period remains blocked because spend is still above the new limit                        |
| `still_blocked_other_period` | The other period still blocks inference                                                       |
| `evaluation_deferred`        | The change was saved, but spend could not be checked; evaluation applies within a few minutes |
| `note`                       | Additional context when present, including exceptions after support lifts a block             |

<Warning>
  Write responses always carry `blocked: false`. Use `evaluation` for the
  immediate result, then call `GET /v2/spend_limits` again to read the current
  block state for both periods. Do not treat a successful write or
  `evaluation.released` alone as confirmation that inference is unblocked.
</Warning>

Blocked inference requests return HTTP 403 with error code `10039` and title
`Inference spend limit reached`. Stop automatic retries for this error. Inspect
both limits, then raise the blocking limits above current spend, remove them,
or wait until their UTC periods end. Verify both periods are unblocked before
resuming requests. Other HTTP 403 errors can have different causes.
