To restore API access, request a higher
approved usage limit
or
contact support
.
503 - Model temporarily overloaded
A
503
response with the
service_unavailable_error
type and
server_is_overloaded
code indicates that the requested model does not have enough capacity to process your request at the moment.
If a
Retry-After
header is present, wait at least as long as it specifies before retrying. If it’s missing, increase the delay between retries. If the error continues, check the
status page
for an active incident.
Python library error types
Python raises
RateLimitError
for
429
responses and
InternalServerError
for
503
responses. If your handler previously caught only one of these classes for throttling and overload, handle both and inspect
error.code
.

Reach out to your organization owner to increase the rate limits on your project
429 - Slow down
A
429
response with the
rate_limit_error
type and
slow_down
code indicates that your request rate increased faster than the service can safely handle. It can occur even when your traffic is within its requests-per-minute and tokens-per-minute limits.
As a rule of thumb, once your traffic reaches 1 million input tokens per minute (TPM), increase it by no more than 50% every 15 minutes. The exact point at which the ramp-rate limit applies can vary by model and traffic conditions.
To resolve this error:
If a
Retry-After
header is present, wait at least as long as it specifies before retrying. If it’s missing, increase the delay between retries and add a small random delay.
