API rate limits
Coming from another cloud?
▸AWS·EC2 Request Limits
This Quake AI feature maps to AWS’s EC2 Request Limits.
▸DigitalOcean·API Rate Limits
This Quake AI feature maps to DigitalOcean’s API Rate Limits.
▸Hetzner·Rate Limits
This Quake AI feature maps to Hetzner’s Rate Limits.
API rate limits
Quake AI distinguishes between two types of limits that affect API usage: resource quotas (how many resources you can create) and API rate limits (how many requests you can send per time window). These are separate mechanisms with different HTTP status codes, different causes, and different resolutions.
Quotas vs. rate limits#
| Dimension | Resource quotas | API rate limits |
|---|---|---|
| What it limits | Number of resources (instances, cores, IPs, volumes) | Number of API requests per time window |
| Enforced by | Compute (Nova), Block Storage (Cinder), Network (Neutron) | Gateway layer (openresty) |
| HTTP status code | 403 Forbidden or 413 Over Limit | 429 Too Many Requests |
| Error message | "Quota exceeded for resources: ['instances']" | Depends on gateway configuration |
| Response headers | None | Retry-After (when present) |
| Resolution | Delete resources or request a quota increase | Slow down and implement backoff |
| Persists until | Resources freed or quota raised | Time window resets |
Visible in /limits | Yes (in the absolute object) | No (rate array is always empty) |
If your API call fails with 403 or 413, you have a quota problem. See Quota and limits troubleshooting.
If your API call fails with 429, you have a rate limit problem. See Handling HTTP 429 responses below.
Current platform behavior#
The OpenStack services on Quake AI (running Antelope 2023.1) do not enforce application-layer API rate limits. Application-level rate limiting was removed from Nova in the Rocky release (2018), and the rate array returned by the /limits endpoint is always empty.
Any rate limiting that exists on Quake AI operates at the gateway layer (openresty), which sits in front of all API endpoints. Gateway rate limit thresholds are not currently published, and API responses do not include RateLimit-* headers.
The /limits endpoint#
The Compute API exposes a /limits endpoint that returns both quota usage and rate limit configuration. On Quake AI, the rate array is always empty. The absolute object contains your project's resource quotas.
Request#
curl -s "$OS_COMPUTE_URL/v2.1/limits" \
-H "X-Auth-Token: $OS_TOKEN" | python3 -m json.toolResponse#
{
"limits": {
"rate": [],
"absolute": {
"maxTotalInstances": 256,
"maxTotalCores": -1,
"maxTotalRAMSize": 40960,
"maxServerMeta": 128,
"maxImageMeta": 128,
"maxPersonality": 5,
"maxPersonalitySize": 10240,
"maxTotalKeypairs": 100,
"maxServerGroups": 100,
"maxServerGroupMembers": 100,
"maxTotalFloatingIps": -1,
"maxSecurityGroups": -1,
"maxSecurityGroupRules": -1,
"totalRAMUsed": 1024,
"totalCoresUsed": 1,
"totalInstancesUsed": 1,
"totalFloatingIpsUsed": 0,
"totalSecurityGroupsUsed": 1,
"totalServerGroupsUsed": 0
}
}
}Field reference#
| Field | Type | Description |
|---|---|---|
rate | array | Rate limit rules. Always empty on Quake AI (Antelope). |
absolute.maxTotal* | integer | Maximum allowed count for the resource. -1 means no application-layer limit. |
absolute.total*Used | integer | Current usage count for the resource. |
A value of -1 for floating IPs, security groups, and security group rules means the application layer does not enforce a cap. Neutron manages these quotas separately.
The Block Storage API (Cinder) exposes a similar endpoint at /v3/{project_id}/limits with volume, snapshot, and gigabyte quotas.
Handling HTTP 429 responses#
A 429 Too Many Requests response originates from the gateway layer, not from the OpenStack service itself. Handle it as follows:
- Check for a
Retry-Afterheader. If present, wait at least that many seconds before retrying. - If no
Retry-Afterheader is present, use exponential backoff with jitter (see Retry and resilience patterns). - Do not retry immediately. Use a minimum initial delay of 500 ms.
- Cap retries at 5 attempts with a maximum delay of 60 seconds per attempt.
- Do not retry in parallel. Sending concurrent retries increases the likelihood of further 429 responses.
Industry comparison#
Rate limit transparency varies across cloud providers. This context helps set expectations for API client design.
| Provider | Published limits | Status code | RateLimit-* headers | Retry-After header |
|---|---|---|---|---|
| DigitalOcean | 5,000 req/hr, 250 req/min burst | 429 | Yes, on every response | Yes, on 429 |
| Hetzner Cloud | 3,600 req/hr (token bucket) | 429 | Yes, on every response | Yes, on 429 |
| AWS EC2 | Per-action token buckets (undocumented) | 400 (RequestLimitExceeded) | No | No |
| Quake AI | Not currently published | 429 | No | When present |
Providers with RateLimit-Remaining and RateLimit-Reset headers on every response let clients self-throttle proactively. Design your Quake AI clients to handle 429 responses reactively.
Client guidance#
Implement these patterns in all API clients, regardless of whether specific rate limits are published.
Exponential backoff with jitter. Increase the wait time exponentially on each retry and add random jitter to prevent thundering herd problems. See Retry and resilience patterns for implementation examples in Bash and Python.
Respect Retry-After. Always check for a Retry-After header before applying your own backoff delay. Use whichever value is larger: the header's value or your calculated backoff.
Retry only retryable status codes. Configure retry adapters to act on 429 and 5xx responses only. Do not retry 4xx errors (except 429).
Cap total retries. Five attempts with a maximum 60-second delay per attempt is a reasonable default. If the error persists after five retries, log the failure and escalate.
Separate rate limits from quotas. A 429 response resolves on its own after the time window resets. A 403 or 413 response (quota exhaustion) never resolves on its own. See Quota and limits troubleshooting for quota recovery steps.
See also#
- Retry and resilience patterns: backoff algorithms, idempotency, and retry strategy by status code
- Quota and limits troubleshooting: diagnosing and recovering from quota exhaustion
- Compute API error reference: status codes for instance operations
- Network API error reference: status codes for network operations
- Block storage API error reference: status codes for volume operations
- Object storage API error reference: status codes for object storage operations
- API endpoints: service endpoint URLs