ドキュメント · 本番運用
レート制限
To keep the platform stable for everyone, request rates and concurrency are limited. Default limits suit most production workloads.
本文は現在英語版のみ提供しています。
Over the limit you receive 429. Limits apply per account and API key and may differ by model.
Handling 429
- Retry with exponential backoff, honoring
Retry-Afterwhen present. - Lower concurrency and smooth traffic spikes with a queue.
- Schedule batch jobs outside your peak hours.
- Need more throughput long-term? Contact us through the console to raise your limits.
Upstream limits
Even below your own limits, occasional capacity pressure at a provider can cause 429 or 503. The platform first tries other channels and only returns the error if they fail too; retry as described above.
Long generations with reasoning models can take minutes. Set your client timeout to at least 300 seconds, or use streaming.