> ## Documentation Index
> Fetch the complete documentation index at: https://developers.reflection.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Rate limits

> How rate limits apply and how to handle them

Rate limits cap how much you can use the API in a period of time. They keep the service reliable for everyone.

## How limits apply

* Limits apply to your **organization**, shared across all of its projects and API keys. Adding keys or projects doesn't raise them.
* Limits are measured in **requests** and in **tokens**, per minute and per UTC calendar day. Token usage includes input tokens, output tokens, and [reasoning tokens](/reasoning).
* The number of **concurrent requests**, those in progress at once, is also limited, for your organization and for each API key.
* Your organization's plan can also set a **daily token allowance**.

## When you exceed a limit

The API returns `429 Too Many Requests` with [error code](/errors#429-rate-limit) `rate_limit_exceeded`, whichever limit was exceeded. The `message` says which one:

```json theme={null}
{
  "error": {
    "message": "Your organization has reached its request limit. Retry after the indicated delay.",
    "type": "rate_limit_error",
    "param": null,
    "code": "rate_limit_exceeded"
  }
}
```

When the response includes a `Retry-After` header, it gives the minimum number of seconds to wait before retrying.

Daily limits and the daily allowance renew at 00:00 UTC. When one is spent, the response also includes `x-should-retry: false`, because retrying can't succeed until then. The OpenAI SDKs honor this header and don't retry.

The service can also return `503` with code `infrastructure_rate_limit_exceeded` or `inference_capacity_unavailable` when inference capacity is temporarily exhausted. These responses also include `Retry-After`; retry them the same way.

## Rate limit headers

Once a request is authenticated and accepted for rate limiting, its response reports your organization's limits and remaining capacity in these headers. Error responses carry them too, including a `429`.

| Header | Description |
| - | - |
| `x-ratelimit-limit-requests` | The number of requests allowed per minute. |
| `x-ratelimit-remaining-requests` | The number of requests remaining this minute. |
| `x-ratelimit-reset-requests` | How long until request capacity fully replenishes, as a duration such as `15s` or `0.75s`. |
| `x-ratelimit-limit-tokens` | The number of tokens allowed per minute. |
| `x-ratelimit-remaining-tokens` | The number of tokens remaining this minute. |
| `x-ratelimit-reset-tokens` | How long until token capacity fully replenishes. |

The same headers with `-day` appended, such as `x-ratelimit-remaining-tokens-day`, report your daily limits. Daily capacity resets at 00:00 UTC, and the daily reset headers give the time until then.

* Request headers appear on every endpoint, and token headers only on chat completions. Token values reflect an estimate of the request's input tokens, made when the request is accepted, not the request's final usage.
* A remaining value counts the current request as of when it was accepted. A rejected request uses no capacity.
* A reset value is the time until capacity fully replenishes if you send no more requests. It can be longer than a minute. It isn't when you can retry: after a `429`, wait for `Retry-After`.
* The headers are omitted when capacity can't be reported, so don't depend on them being present.

## Stay within your limits

* **Retry with backoff.** Wait at least `Retry-After` seconds when it is present, and back off exponentially with jitter when it isn't. The OpenAI SDKs do this automatically; see [Retrying](/errors#retrying).
* **Limit concurrency.** Cap the number of requests in flight rather than sending bursts.
* **Watch the [rate limit headers](#rate-limit-headers)** to slow down before you reach a limit.
* **Set `max_completion_tokens`** to what you need, and lower [`reasoning_effort`](/reasoning) for simple tasks.
* **Trim conversation history** you don't need, to use fewer input tokens.

Details on how to request higher limits from Reflection are coming soon.
