Docs · Production
Rate limits
To keep the platform stable for everyone, request rates and concurrency are limited. Default limits suit most production workloads.
Over the limit you receive 429. Limits apply per account and API key and may differ by model.
Handling 429
- Retry with exponential backoff, honoring
Retry-Afterwhen present. - Lower concurrency and smooth traffic spikes with a queue.
- Schedule batch jobs outside your peak hours.
- Need more throughput long-term? Contact us through the console to raise your limits.
Upstream limits
Even below your own limits, occasional capacity pressure at a provider can cause 429 or 503. The platform first tries other channels and only returns the error if they fail too; retry as described above.
Long generations with reasoning models can take minutes. Set your client timeout to at least 300 seconds, or use streaming.