Skip to content

Retry and resilience patterns

Reference · Updated Jun 2026

Coming from another cloud?

▸AWS·Error Retries

This Quake AI feature maps to AWS’s Error Retries.

Retry and resilience patterns

When an API call fails, the correct response depends on why it failed. Some errors are transient and resolve on retry. Others indicate a permanent problem that retrying will never fix. This page explains how to classify errors, implement safe retries, and avoid common pitfalls.

Error classification#

API errors fall into three categories. Each category requires a different response.

Content errors (4xx: fix the request)#

The request itself is wrong. The server understood what you asked but cannot fulfill it because the input is invalid, the resource does not exist, or you lack permission.

Examples: 400 Bad Request, 401 Unauthorized, 403 Forbidden, 404 Not Found

Response: Do not retry. Fix the request: correct the parameters, re-authenticate, adjust permissions, or resolve quota exhaustion.

Server errors (5xx: retry with backoff)#

The server encountered an internal problem. The request may be valid, but the server cannot process it right now.

Examples: 500 Internal Server Error, 503 Service Unavailable

Response: Retry with exponential backoff. These are typically transient.

State-conflict errors (409: wait, then retry)#

The request is valid but conflicts with the resource's current state. The resource is in a transitional state (building, attaching, detaching) and cannot accept the action yet.

Examples: 409 Conflict on instance actions, volume operations, or network resource modifications

Response: Wait for the state transition to complete, then retry. Check the resource's status and task_state fields before retrying.

Status code retry map#

Use this table as the authoritative reference for retry behavior across all Quake AI APIs.

StatusNameRetry?Strategy
400Bad RequestNoFix the request body or parameters
401UnauthorizedNoRe-authenticate (openstack token issue)
403ForbiddenNoFix permissions or free quota
404Not FoundNoVerify the resource exists
409ConflictConditionalWait for state transition, then retry
412Precondition FailedYesRe-read the resource, retry with current revision
413Over LimitNoQuota exhaustion: free resources or request increase
429Too Many RequestsYesGateway rate limit: backoff and retry. Check Retry-After header
500Internal Server ErrorYesRetry with exponential backoff
503Service UnavailableYesRetry with exponential backoff

Exponential backoff with jitter#

When retrying, increase the wait time exponentially and add random jitter. This prevents thundering herd problems where many clients retry simultaneously after a service disruption.

Algorithm#

wait = min(base * 2^attempt + random(0, base), max_wait)
  • base: initial wait time (recommended: 1 second)
  • attempt: retry counter starting at 0
  • max_wait: ceiling to prevent unbounded waits (recommended: 60 seconds)
  • jitter: random value between 0 and base to spread retries across time

Concrete timing example#

AttemptBase waitJitter (random)Total wait
01s+0.4s~1.4s
12s+0.7s~2.7s
24s+0.2s~4.2s
38s+0.9s~8.9s
416s+0.5s~16.5s
532s+0.3s~32.3s

After 5 retries with this schedule, you have waited approximately 66 seconds total. If the error persists, escalate to support.

Bash implementation#

bash
retry_with_backoff() {
  local max_attempts=5
  local base=1
  local max_wait=60
  local attempt=0

  until "$@"; do
    attempt=$((attempt + 1))
    if [ "$attempt" -ge "$max_attempts" ]; then
      echo "Failed after $max_attempts attempts" >&2
      return 1
    fi
    wait=$(echo "$base * (2 ^ ($attempt - 1))" | bc)
    jitter=$(echo "scale=1; $RANDOM % 10 / 10" | bc)
    sleep_time=$(echo "if ($wait + $jitter > $max_wait) $max_wait else $wait + $jitter" | bc)
    echo "Attempt $attempt failed. Retrying in ${sleep_time}s..." >&2
    sleep "$sleep_time"
  done
}

retry_with_backoff openstack server create --flavor m2a.large --image ubuntu-24.04 --network YOUR_PRIVATE_NETWORK my-instance

Python implementation#

Python
import time
import random

def retry_with_backoff(fn, max_attempts=5, base=1, max_wait=60):
    for attempt in range(max_attempts):
        try:
            return fn()
        except Exception as e:
            if attempt == max_attempts - 1:
                raise
            wait = min(base * (2 ** attempt) + random.uniform(0, base), max_wait)
            print(f"Attempt {attempt + 1} failed: {e}. Retrying in {wait:.1f}s...")
            time.sleep(wait)

Idempotency#

An operation is idempotent if performing it multiple times produces the same result as performing it once. Idempotent operations are safe to retry without side effects.

Idempotent operations#

OperationWhy it is safe
GET (any resource)Read-only, no side effects
DELETE (already-deleted resource)Returns 404, no additional effect
PUT (update with full resource body)Replaces the resource with the same state
Setting a volume to a specific stateFinal state is the same regardless of repetitions

Non-idempotent operations (retry with caution)#

OperationRisk on retry
POST (create resource)May create duplicate resources
openstack server createMay launch multiple instances
openstack floating ip createMay allocate multiple IPs

Making non-idempotent operations safe#

  1. Check before retry. Before retrying a create operation, list existing resources to see if the first attempt succeeded:
bash
openstack server list --name my-instance
  1. Use deterministic names. If your workflow creates resources with specific names, a duplicate-name error (409) confirms the first attempt succeeded.

  2. Track request state externally. In automation scripts, record each operation's result before retrying. Use a state file or database to track which operations completed.

When NOT to retry#

  • Authentication failures (401): Re-authenticate; the token is expired. Retrying the same request will not succeed.
  • Permission errors (403): Fix the role assignment or project scope.
  • Quota exhaustion (403 with quota message): Free resources first; retrying consumes API capacity without releasing the underlying quota.
  • Validation errors (400): Fix the parameters; the request is malformed.
  • Resource not found (404): The resource was deleted or never existed.

Client-side patterns#

OpenStack SDK#

The openstacksdk (Python) includes built-in retry configuration. Set max_retries in your connection profile:

Python
import openstack

conn = openstack.connect(
    cloud='rumble',
    # SDK retries 5xx errors automatically
)

The SDK retries on connection errors and 503 responses by default. Override per-call if needed.

HTTP clients (general)#

Most HTTP client libraries support retry adapters:

  • Python requests: Use urllib3.util.retry.Retry with a requests.adapters.HTTPAdapter
  • Node.js: Libraries like axios-retry or got have built-in retry support
  • Go: Use hashicorp/go-retryablehttp

Configure all retry adapters to:

  1. Retry only on 5xx and 429 status codes
  2. Use exponential backoff with jitter
  3. Respect Retry-After headers when present
  4. Cap total retry attempts (5 is a reasonable default)

See also#

Was this page helpful?