How limits apply
- Limits apply to your organization, shared across all of its projects and API keys. Adding keys or projects doesn’t raise them.
- Limits are measured in requests and in tokens, per minute and per UTC calendar day. Token usage includes input tokens, output tokens, and reasoning tokens.
- The number of concurrent requests, those in progress at once, is also limited, for your organization and for each API key.
- Your organization’s plan can also set a daily token allowance.
When you exceed a limit
The API returns429 Too Many Requests with error code rate_limit_exceeded, whichever limit was exceeded. The message says which one:
Retry-After header, it gives the minimum number of seconds to wait before retrying.
Daily limits and the daily allowance renew at 00:00 UTC. When one is spent, the response also includes x-should-retry: false, because retrying can’t succeed until then. The OpenAI SDKs honor this header and don’t retry.
The service can also return 503 with code infrastructure_rate_limit_exceeded or inference_capacity_unavailable when inference capacity is temporarily exhausted. These responses also include Retry-After; retry them the same way.
Rate limit headers
Once a request is authenticated and accepted for rate limiting, its response reports your organization’s limits and remaining capacity in these headers. Error responses carry them too, including a429.
The same headers with
-day appended, such as x-ratelimit-remaining-tokens-day, report your daily limits. Daily capacity resets at 00:00 UTC, and the daily reset headers give the time until then.
- Request headers appear on every endpoint, and token headers only on chat completions. Token values reflect an estimate of the request’s input tokens, made when the request is accepted, not the request’s final usage.
- A remaining value counts the current request as of when it was accepted. A rejected request uses no capacity.
- A reset value is the time until capacity fully replenishes if you send no more requests. It can be longer than a minute. It isn’t when you can retry: after a
429, wait forRetry-After. - The headers are omitted when capacity can’t be reported, so don’t depend on them being present.
Stay within your limits
- Retry with backoff. Wait at least
Retry-Afterseconds when it is present, and back off exponentially with jitter when it isn’t. The OpenAI SDKs do this automatically; see Retrying. - Limit concurrency. Cap the number of requests in flight rather than sending bursts.
- Watch the rate limit headers to slow down before you reach a limit.
- Set
max_completion_tokensto what you need, and lowerreasoning_effortfor simple tasks. - Trim conversation history you don’t need, to use fewer input tokens.