Skip to content

API rate limits

Reference · Updated May 2026

Coming from another cloud?

▸AWS·EC2 Request Limits

This Quake AI feature maps to AWS’s EC2 Request Limits.

▸DigitalOcean·API Rate Limits

This Quake AI feature maps to DigitalOcean’s API Rate Limits.

▸Hetzner·Rate Limits

This Quake AI feature maps to Hetzner’s Rate Limits.

API rate limits

Quake AI distinguishes between two types of limits that affect API usage: resource quotas (how many resources you can create) and API rate limits (how many requests you can send per time window). These are separate mechanisms with different HTTP status codes, different causes, and different resolutions.

Quotas vs. rate limits#

DimensionResource quotasAPI rate limits
What it limitsNumber of resources (instances, cores, IPs, volumes)Number of API requests per time window
Enforced byCompute (Nova), Block Storage (Cinder), Network (Neutron)Gateway layer (openresty)
HTTP status code403 Forbidden or 413 Over Limit429 Too Many Requests
Error message"Quota exceeded for resources: ['instances']"Depends on gateway configuration
Response headersNoneRetry-After (when present)
ResolutionDelete resources or request a quota increaseSlow down and implement backoff
Persists untilResources freed or quota raisedTime window resets
Visible in /limitsYes (in the absolute object)No (rate array is always empty)

If your API call fails with 403 or 413, you have a quota problem. See Quota and limits troubleshooting.

If your API call fails with 429, you have a rate limit problem. See Handling HTTP 429 responses below.

Current platform behavior#

The OpenStack services on Quake AI (running Antelope 2023.1) do not enforce application-layer API rate limits. Application-level rate limiting was removed from Nova in the Rocky release (2018), and the rate array returned by the /limits endpoint is always empty.

Any rate limiting that exists on Quake AI operates at the gateway layer (openresty), which sits in front of all API endpoints. Gateway rate limit thresholds are not currently published, and API responses do not include RateLimit-* headers.

The /limits endpoint#

The Compute API exposes a /limits endpoint that returns both quota usage and rate limit configuration. On Quake AI, the rate array is always empty. The absolute object contains your project's resource quotas.

Request#

bash
curl -s "$OS_COMPUTE_URL/v2.1/limits" \
  -H "X-Auth-Token: $OS_TOKEN" | python3 -m json.tool

Response#

JSON
{
  "limits": {
    "rate": [],
    "absolute": {
      "maxTotalInstances": 256,
      "maxTotalCores": -1,
      "maxTotalRAMSize": 40960,
      "maxServerMeta": 128,
      "maxImageMeta": 128,
      "maxPersonality": 5,
      "maxPersonalitySize": 10240,
      "maxTotalKeypairs": 100,
      "maxServerGroups": 100,
      "maxServerGroupMembers": 100,
      "maxTotalFloatingIps": -1,
      "maxSecurityGroups": -1,
      "maxSecurityGroupRules": -1,
      "totalRAMUsed": 1024,
      "totalCoresUsed": 1,
      "totalInstancesUsed": 1,
      "totalFloatingIpsUsed": 0,
      "totalSecurityGroupsUsed": 1,
      "totalServerGroupsUsed": 0
    }
  }
}

Field reference#

FieldTypeDescription
ratearrayRate limit rules. Always empty on Quake AI (Antelope).
absolute.maxTotal*integerMaximum allowed count for the resource. -1 means no application-layer limit.
absolute.total*UsedintegerCurrent usage count for the resource.

A value of -1 for floating IPs, security groups, and security group rules means the application layer does not enforce a cap. Neutron manages these quotas separately.

The Block Storage API (Cinder) exposes a similar endpoint at /v3/{project_id}/limits with volume, snapshot, and gigabyte quotas.

Handling HTTP 429 responses#

A 429 Too Many Requests response originates from the gateway layer, not from the OpenStack service itself. Handle it as follows:

  1. Check for a Retry-After header. If present, wait at least that many seconds before retrying.
  2. If no Retry-After header is present, use exponential backoff with jitter (see Retry and resilience patterns).
  3. Do not retry immediately. Use a minimum initial delay of 500 ms.
  4. Cap retries at 5 attempts with a maximum delay of 60 seconds per attempt.
  5. Do not retry in parallel. Sending concurrent retries increases the likelihood of further 429 responses.

Industry comparison#

Rate limit transparency varies across cloud providers. This context helps set expectations for API client design.

ProviderPublished limitsStatus codeRateLimit-* headersRetry-After header
DigitalOcean5,000 req/hr, 250 req/min burst429Yes, on every responseYes, on 429
Hetzner Cloud3,600 req/hr (token bucket)429Yes, on every responseYes, on 429
AWS EC2Per-action token buckets (undocumented)400 (RequestLimitExceeded)NoNo
Quake AINot currently published429NoWhen present

Providers with RateLimit-Remaining and RateLimit-Reset headers on every response let clients self-throttle proactively. Design your Quake AI clients to handle 429 responses reactively.

Client guidance#

Implement these patterns in all API clients, regardless of whether specific rate limits are published.

Exponential backoff with jitter. Increase the wait time exponentially on each retry and add random jitter to prevent thundering herd problems. See Retry and resilience patterns for implementation examples in Bash and Python.

Respect Retry-After. Always check for a Retry-After header before applying your own backoff delay. Use whichever value is larger: the header's value or your calculated backoff.

Retry only retryable status codes. Configure retry adapters to act on 429 and 5xx responses only. Do not retry 4xx errors (except 429).

Cap total retries. Five attempts with a maximum 60-second delay per attempt is a reasonable default. If the error persists after five retries, log the failure and escalate.

Separate rate limits from quotas. A 429 response resolves on its own after the time window resets. A 403 or 413 response (quota exhaustion) never resolves on its own. See Quota and limits troubleshooting for quota recovery steps.

See also#

Was this page helpful?