# API rate limits

Source: https://docs.quake.ai/docs/tools/api-rate-limits
Markdown: https://docs.quake.ai/docs/tools/api-rate-limits.md

---

# API rate limits

Quake AI distinguishes between two types of limits that affect API usage: **resource quotas** (how many resources you can create) and **API rate limits** (how many requests you can send per time window). These are separate mechanisms with different HTTP status codes, different causes, and different resolutions.

## Quotas vs. rate limits

| Dimension | Resource quotas | API rate limits |
|---|---|---|
| What it limits | Number of resources (instances, cores, IPs, volumes) | Number of API requests per time window |
| Enforced by | Compute (Nova), Block Storage (Cinder), Network (Neutron) | Gateway layer (openresty) |
| HTTP status code | 403 Forbidden or 413 Over Limit | 429 Too Many Requests |
| Error message | `"Quota exceeded for resources: ['instances']"` | Depends on gateway configuration |
| Response headers | None | `Retry-After` (when present) |
| Resolution | Delete resources or request a quota increase | Slow down and implement backoff |
| Persists until | Resources freed or quota raised | Time window resets |
| Visible in `/limits` | Yes (in the `absolute` object) | No (`rate` array is always empty) |

If your API call fails with 403 or 413, you have a quota problem. See [Quota and limits troubleshooting](/docs/operate/troubleshooting/quota-and-limits).

If your API call fails with 429, you have a rate limit problem. See [Handling HTTP 429 responses](#handling-http-429-responses) below.

## Current platform behavior

The OpenStack services on Quake AI (running [Antelope 2023.1](/resources/migration/openstack)) do not enforce application-layer API rate limits. Application-level rate limiting was removed from Nova in the Rocky release (2018), and the `rate` array returned by the `/limits` endpoint is always empty.

Any rate limiting that exists on Quake AI operates at the gateway layer (openresty), which sits in front of all API endpoints. Gateway rate limit thresholds are not currently published, and API responses do not include `RateLimit-*` headers.


Implement backoff and retry logic in all API clients. See [Client guidance](#client-guidance) for recommended patterns.


## The `/limits` endpoint

The Compute API exposes a `/limits` endpoint that returns both quota usage and rate limit configuration. On Quake AI, the `rate` array is always empty. The `absolute` object contains your project's resource quotas.

### Request

```bash
curl -s "$OS_COMPUTE_URL/v2.1/limits" \
  -H "X-Auth-Token: $OS_TOKEN" | python3 -m json.tool
```

### Response

```json
{
  "limits": {
    "rate": [],
    "absolute": {
      "maxTotalInstances": 256,
      "maxTotalCores": -1,
      "maxTotalRAMSize": 40960,
      "maxServerMeta": 128,
      "maxImageMeta": 128,
      "maxPersonality": 5,
      "maxPersonalitySize": 10240,
      "maxTotalKeypairs": 100,
      "maxServerGroups": 100,
      "maxServerGroupMembers": 100,
      "maxTotalFloatingIps": -1,
      "maxSecurityGroups": -1,
      "maxSecurityGroupRules": -1,
      "totalRAMUsed": 1024,
      "totalCoresUsed": 1,
      "totalInstancesUsed": 1,
      "totalFloatingIpsUsed": 0,
      "totalSecurityGroupsUsed": 1,
      "totalServerGroupsUsed": 0
    }
  }
}
```

### Field reference

| Field | Type | Description |
|---|---|---|
| `rate` | array | Rate limit rules. Always empty on Quake AI (Antelope). |
| `absolute.maxTotal*` | integer | Maximum allowed count for the resource. `-1` means no application-layer limit. |
| `absolute.total*Used` | integer | Current usage count for the resource. |

A value of `-1` for floating IPs, security groups, and security group rules means the application layer does not enforce a cap. Neutron manages these quotas separately.

The Block Storage API (Cinder) exposes a similar endpoint at `/v3/{project_id}/limits` with volume, snapshot, and gigabyte quotas.

## Handling HTTP 429 responses

A 429 Too Many Requests response originates from the gateway layer, not from the OpenStack service itself. Handle it as follows:

1. Check for a `Retry-After` header. If present, wait at least that many seconds before retrying.
2. If no `Retry-After` header is present, use exponential backoff with jitter (see [Retry and resilience patterns](/reference/api-conventions/retry-and-resilience)).
3. Do not retry immediately. Use a minimum initial delay of 500 ms.
4. Cap retries at 5 attempts with a maximum delay of 60 seconds per attempt.
5. Do not retry in parallel. Sending concurrent retries increases the likelihood of further 429 responses.


A 429 response means the gateway is actively throttling your requests. Retrying aggressively makes the situation worse and may extend the throttle window.


## Industry comparison

Rate limit transparency varies across cloud providers. This context helps set expectations for API client design.

| Provider | Published limits | Status code | `RateLimit-*` headers | `Retry-After` header |
|---|---|---|---|---|
| DigitalOcean | 5,000 req/hr, 250 req/min burst | 429 | Yes, on every response | Yes, on 429 |
| Hetzner Cloud | 3,600 req/hr (token bucket) | 429 | Yes, on every response | Yes, on 429 |
| AWS EC2 | Per-action token buckets (undocumented) | 400 (`RequestLimitExceeded`) | No | No |
| Quake AI | Not currently published | 429 | No | When present |

Providers with `RateLimit-Remaining` and `RateLimit-Reset` headers on every response let clients self-throttle proactively. Design your Quake AI clients to handle 429 responses reactively.

## Client guidance

Implement these patterns in all API clients, regardless of whether specific rate limits are published.

**Exponential backoff with jitter.** Increase the wait time exponentially on each retry and add random jitter to prevent thundering herd problems. See [Retry and resilience patterns](/reference/api-conventions/retry-and-resilience) for implementation examples in Bash and Python.

**Respect `Retry-After`.** Always check for a `Retry-After` header before applying your own backoff delay. Use whichever value is larger: the header's value or your calculated backoff.

**Retry only retryable status codes.** Configure retry adapters to act on 429 and 5xx responses only. Do not retry 4xx errors (except 429).

**Cap total retries.** Five attempts with a maximum 60-second delay per attempt is a reasonable default. If the error persists after five retries, log the failure and escalate.

**Separate rate limits from quotas.** A 429 response resolves on its own after the time window resets. A 403 or 413 response (quota exhaustion) never resolves on its own. See [Quota and limits troubleshooting](/docs/operate/troubleshooting/quota-and-limits) for quota recovery steps.

## See also

- [Retry and resilience patterns](/reference/api-conventions/retry-and-resilience): backoff algorithms, idempotency, and retry strategy by status code
- [Quota and limits troubleshooting](/docs/operate/troubleshooting/quota-and-limits): diagnosing and recovering from quota exhaustion
- [Compute API error reference](/reference/compute/api-errors): status codes for instance operations
- [Network API error reference](/reference/network/api-errors): status codes for network operations
- [Block storage API error reference](/reference/block-storage/api-errors): status codes for volume operations
- [Object storage API error reference](/reference/object-storage/api-errors): status codes for object storage operations
- [API endpoints](/docs/tools/api-endpoints): service endpoint URLs
