Retry and resilience patterns
When an API call fails, the correct response depends on why it failed. Some errors are transient and resolve on retry. Others indicate a permanent problem that retrying will never fix. This page explains how to classify errors, implement safe retries, and avoid common pitfalls.
Error classification#
API errors fall into three categories. Each category requires a different response.
Content errors (4xx: fix the request)#
The request itself is wrong. The server understood what you asked but cannot fulfill it because the input is invalid, the resource does not exist, or you lack permission.
Examples: 400 Bad Request, 401 Unauthorized, 403 Forbidden, 404 Not Found
Response: Do not retry. Fix the request: correct the parameters, re-authenticate, adjust permissions, or resolve quota exhaustion.
Server errors (5xx: retry with backoff)#
The server encountered an internal problem. The request may be valid, but the server cannot process it right now.
Examples: 500 Internal Server Error, 503 Service Unavailable
Response: Retry with exponential backoff. These are typically transient.
State-conflict errors (409: wait, then retry)#
The request is valid but conflicts with the resource's current state. The resource is in a transitional state (building, attaching, detaching) and cannot accept the action yet.
Examples: 409 Conflict on instance actions, volume operations, or network resource modifications
Response: Wait for the state transition to complete, then retry. Check the resource's status and task_state fields before retrying.
Status code retry map#
Use this table as the authoritative reference for retry behavior across all Quake AI APIs.
| Status | Name | Retry? | Strategy |
|---|---|---|---|
| 400 | Bad Request | No | Fix the request body or parameters |
| 401 | Unauthorized | No | Re-authenticate (openstack token issue) |
| 403 | Forbidden | No | Fix permissions or free quota |
| 404 | Not Found | No | Verify the resource exists |
| 409 | Conflict | Conditional | Wait for state transition, then retry |
| 412 | Precondition Failed | Yes | Re-read the resource, retry with current revision |
| 413 | Over Limit | No | Quota exhaustion: free resources or request increase |
| 429 | Too Many Requests | Yes | Gateway rate limit: backoff and retry. Check Retry-After header |
| 500 | Internal Server Error | Yes | Retry with exponential backoff |
| 503 | Service Unavailable | Yes | Retry with exponential backoff |
Exponential backoff with jitter#
When retrying, increase the wait time exponentially and add random jitter. This prevents thundering herd problems where many clients retry simultaneously after a service disruption.
Algorithm#
wait = min(base * 2^attempt + random(0, base), max_wait)- base: initial wait time (recommended: 1 second)
- attempt: retry counter starting at 0
- max_wait: ceiling to prevent unbounded waits (recommended: 60 seconds)
- jitter: random value between 0 and
baseto spread retries across time
Concrete timing example#
| Attempt | Base wait | Jitter (random) | Total wait |
|---|---|---|---|
| 0 | 1s | +0.4s | ~1.4s |
| 1 | 2s | +0.7s | ~2.7s |
| 2 | 4s | +0.2s | ~4.2s |
| 3 | 8s | +0.9s | ~8.9s |
| 4 | 16s | +0.5s | ~16.5s |
| 5 | 32s | +0.3s | ~32.3s |
After 5 retries with this schedule, you have waited approximately 66 seconds total. If the error persists, escalate to support.
Bash implementation#
retry_with_backoff() {
local max_attempts=5
local base=1
local max_wait=60
local attempt=0
until "$@"; do
attempt=$((attempt + 1))
if [ "$attempt" -ge "$max_attempts" ]; then
echo "Failed after $max_attempts attempts" >&2
return 1
fi
wait=$(echo "$base * (2 ^ ($attempt - 1))" | bc)
jitter=$(echo "scale=1; $RANDOM % 10 / 10" | bc)
sleep_time=$(echo "if ($wait + $jitter > $max_wait) $max_wait else $wait + $jitter" | bc)
echo "Attempt $attempt failed. Retrying in ${sleep_time}s..." >&2
sleep "$sleep_time"
done
}
retry_with_backoff openstack server create --flavor m2a.large --image ubuntu-24.04 --network YOUR_PRIVATE_NETWORK my-instancePython implementation#
import time
import random
def retry_with_backoff(fn, max_attempts=5, base=1, max_wait=60):
for attempt in range(max_attempts):
try:
return fn()
except Exception as e:
if attempt == max_attempts - 1:
raise
wait = min(base * (2 ** attempt) + random.uniform(0, base), max_wait)
print(f"Attempt {attempt + 1} failed: {e}. Retrying in {wait:.1f}s...")
time.sleep(wait)Idempotency#
An operation is idempotent if performing it multiple times produces the same result as performing it once. Idempotent operations are safe to retry without side effects.
Idempotent operations#
| Operation | Why it is safe |
|---|---|
| GET (any resource) | Read-only, no side effects |
| DELETE (already-deleted resource) | Returns 404, no additional effect |
| PUT (update with full resource body) | Replaces the resource with the same state |
| Setting a volume to a specific state | Final state is the same regardless of repetitions |
Non-idempotent operations (retry with caution)#
| Operation | Risk on retry |
|---|---|
| POST (create resource) | May create duplicate resources |
openstack server create | May launch multiple instances |
openstack floating ip create | May allocate multiple IPs |
Making non-idempotent operations safe#
- Check before retry. Before retrying a create operation, list existing resources to see if the first attempt succeeded:
openstack server list --name my-instance-
Use deterministic names. If your workflow creates resources with specific names, a duplicate-name error (409) confirms the first attempt succeeded.
-
Track request state externally. In automation scripts, record each operation's result before retrying. Use a state file or database to track which operations completed.
When NOT to retry#
- Authentication failures (401): Re-authenticate; the token is expired. Retrying the same request will not succeed.
- Permission errors (403): Fix the role assignment or project scope.
- Quota exhaustion (403 with quota message): Free resources first; retrying consumes API capacity without releasing the underlying quota.
- Validation errors (400): Fix the parameters; the request is malformed.
- Resource not found (404): The resource was deleted or never existed.
Client-side patterns#
OpenStack SDK#
The openstacksdk (Python) includes built-in retry configuration. Set max_retries in your connection profile:
import openstack
conn = openstack.connect(
cloud='rumble',
# SDK retries 5xx errors automatically
)The SDK retries on connection errors and 503 responses by default. Override per-call if needed.
HTTP clients (general)#
Most HTTP client libraries support retry adapters:
- Python
requests: Useurllib3.util.retry.Retrywith arequests.adapters.HTTPAdapter - Node.js: Libraries like
axios-retryorgothave built-in retry support - Go: Use
hashicorp/go-retryablehttp
Configure all retry adapters to:
- Retry only on 5xx and 429 status codes
- Use exponential backoff with jitter
- Respect
Retry-Afterheaders when present - Cap total retry attempts (5 is a reasonable default)
See also#
- API rate limits: quotas vs. rate limits,
/limitsendpoint, 429 handling - Troubleshooting overview: symptom-first diagnostic guide
- Compute API error reference: instance faults and state conflicts
- Network API error reference: port, router, and security group errors
- Block storage API error reference: volume state machine and delete conflicts
- Object storage API error reference: S3 authentication and upload errors
- Auth token diagnostics: 401/403 recovery
- Quota and limits troubleshooting: quota exhaustion recovery