# Instance lifecycle troubleshooting

Source: https://docs.quake.ai/docs/operate/runbooks/instance-lifecycle
Markdown: https://docs.quake.ai/docs/operate/runbooks/instance-lifecycle.md
> Diagnose instances stuck in BUILD, ERROR, or VERIFY_RESIZE states and recover to a working state.

---

# Instance lifecycle troubleshooting

This runbook covers provisioning failures on Quake AI: instances stuck in BUILD, instances that end in ERROR, and create operations blocked by quota limits on the [Compute service (OpenStack Nova)](/resources/migration/openstack).

<Figure size="md" caption="Where to start: pick the branch that matches your symptom, then jump to that section below.">

```d2
direction: right

start: "Instance launch\nor lifecycle issue"

q_state: "What state\nis the instance in?" {shape: diamond}

leaf_build: "Stuck in BUILD"
leaf_error: "ERROR after launch"
leaf_quota: "Quota exceeded"

start -> q_state
q_state -> leaf_build: BUILD (no progress)
q_state -> leaf_error: ERROR (failure)
q_state -> leaf_quota: quota error message
```

</Figure>

## Quick triage

| Symptom | Most likely cause | Jump to |
| --- | --- | --- |
| Status stays **BUILD** and never becomes **ACTIVE** or **ERROR** | Scheduling or messaging failure, volume or network setup timeout during build | [Instance stuck in BUILD state](#instance-stuck-in-build-state) |
| Status moves **BUILD** → **ERROR**; `fault` is populated | Scheduler placement failure, Block Storage (Cinder) timeout, port binding failure, or repeated build retries | [Instance in ERROR state after launch](#instance-in-error-state-after-launch) |
| HTTP **403** or message `Quota exceeded for compute_units, ram: …` | Project quota exhausted (instances, vCPUs, RAM, volumes, floating IPs, etc.) | [Quota exceeded errors](#quota-exceeded-errors) |

---

## Instance stuck in BUILD state

### Symptoms

- Instance status shows **BUILD** indefinitely and never reaches **ACTIVE** (running) or **ERROR**.

### Diagnosis

Root causes include scheduling failure (no host with sufficient resources), messaging failure between services, volume creation timeout when you boot from a volume, or networking setup failure during the build.

Wait at least five minutes; some builds are slow under load. Inspect the fault field:

```bash
openstack server show YOUR_INSTANCE_NAME -c fault
```

If `fault` stays empty and status remains **BUILD** after about 10 minutes, treat the build as stuck. Check project quota headroom:

```bash
openstack quota show
```

If you boot from a volume, confirm the volume is **available** before launch and still healthy:

```bash
openstack volume list
```

### Resolution

1. Wait five minutes before you change anything.
2. Run `openstack server show YOUR_INSTANCE_NAME -c fault` and note any message.
3. If there is still no fault after 10 minutes in **BUILD**, delete the instance and retry creation.
4. Retry with a different or smaller flavor (`YOUR_FLAVOR`) to rule out scheduler resource constraints.
5. If quota is tight, free capacity or request an increase before you retry.

### Verification

```bash
openstack server show YOUR_INSTANCE_NAME -c status -c addresses -c fault
```

After a successful retry, `status` should be **ACTIVE** and `fault` empty.

### Prevention

- Check quota before batch or large launches.
- Use smaller flavors for iterative tests.
- When booting from a volume, wait until the volume is **available** before you create the instance.

### When to escalate

If the instance remains in **BUILD** for more than 15 minutes with an empty `fault`, the problem is likely platform-side on Quake AI. Open a support ticket with the instance ID and the UTC timestamp when creation started.

---

## Instance in ERROR state after launch

### Symptoms (instance in ERROR state after launch)

- Instance transitions from **BUILD** to **ERROR**.
- A fault message appears when you inspect the server record.

### Diagnosis (instance in ERROR state after launch)

Read the fault:

```bash
openstack server show YOUR_INSTANCE_NAME -c fault
```

Common messages and what they mean:

| Fault message (excerpt) | Layer |
| --- | --- |
| `No valid host was found` | Nova’s scheduler cannot place the instance on any hypervisor. |
| `Build of instance aborted: Volume did not finish being created` | Block Storage (Cinder) did not finish in time (typical with boot-from-volume). |
| `Exceeded maximum number of retries` | The build pipeline gave up after repeated failures. |
| `PortBindingFailed` | Network (Neutron) cannot bind the port to the dataplane. |

### Resolution (instance in ERROR state after launch)

1. Start from the full `fault` text from `openstack server show YOUR_INSTANCE_NAME -c fault`.
2. **No valid host:** Try a smaller flavor, run `openstack quota show` (and `openstack quota show --usage`), wait and retry if the region is busy, then delete the **ERROR** instance and recreate.
3. **Volume did not finish / Cinder timeout:** Check volume status with `openstack volume list`, delete orphaned or failed volumes where safe, and retry; for a faster path, boot from an image instead of a volume.
4. **PortBindingFailed:** Verify the network and subnet exist, are attached to your project, and match what you passed at create time; confirm the instance can reach the intended network path (see [Instance connectivity troubleshooting](/docs/operate/runbooks/instance-connectivity)).
5. Delete the **ERROR** instance after you address the cause, then create a new instance.

### Verification (instance in ERROR state after launch)

```bash
openstack server show YOUR_INSTANCE_NAME -c status -c fault
```

After recreation, `status` should be **ACTIVE** and `fault` empty.

### Prevention (instance in ERROR state after launch)

- Run `openstack quota show` or `openstack quota show --usage` before launch.
- Confirm your chosen flavor is appropriate with `openstack flavor list`.
- Prefer image-based boot unless you require boot-from-volume.

### When to escalate (instance in ERROR state after launch)

If **No valid host** appears for every reasonable flavor you try, the region may be capacity-constrained. Open a support ticket with the full fault text, flavor names or IDs, and UTC timestamps.

---

## Quota exceeded errors

### Symptoms (quota exceeded errors)

- The API returns HTTP **403** or a message such as `Quota exceeded for compute_units, ram: Requested ..., but already used ... of ... (HTTP 403)` (your resource list may differ). On Quake AI, `compute_units` is a Quake AI-specific quota dimension that can bind before `cores` or `ram`.
- Create operations fail for instances, cores, RAM, volumes, floating IPs, or other project-scoped resources.

### Diagnosis (quota exceeded errors)

Project quotas cap instances, vCPUs, RAM, volumes, storage, floating IPs, security groups, and related resources. Compare limits to current use:

```bash
openstack quota show --usage
```

Identify which resource shows usage at or above its limit.

### Resolution (quota exceeded errors)

1. Run `openstack quota show --usage`.
2. Note which resource is exhausted (instances, cores, RAM, volumes, gigabytes, floating IPs, etc.).
3. Delete unused instances.
4. Release unused floating IPs.
5. Delete unused volumes and snapshots where safe.
6. If cleanup does not restore enough headroom, request a quota increase through support.

### Verification (quota exceeded errors)

```bash
openstack quota show --usage
```

Retry the operation that failed.

### Prevention (quota exceeded errors)

- Review quota periodically, especially after test campaigns.
- Tear down dev and test resources after use.
- Run `openstack quota show --usage` as a preflight check before large changes.

### When to escalate (Quota exceeded errors)

Default quotas follow your plan tier. If you need more capacity than cleanup provides, open a structured quota increase request with your target limits and use case.

---

## Collecting evidence for support tickets

Gather this before you open a ticket:

| Evidence | Command or note |
| --- | --- |
| Instance ID, status, fault | `openstack server show YOUR_INSTANCE_NAME -c id -c status -c fault` |
| Console log (last lines) | `openstack console log show YOUR_INSTANCE_NAME --lines 50` |
| Quota usage | `openstack quota show --usage` |
| Volumes (if boot-from-volume) | `openstack volume list` |
| UTC timestamp | When creation started or when the failure occurred |

## See also

- [Compute API error reference](/reference/compute/api-errors): HTTP status codes, fault messages, and state-conflict handling
- [Instance connectivity troubleshooting](/docs/operate/runbooks/instance-connectivity)
- [Quota and limits](/docs/operate/troubleshooting/quota-and-limits)
- [Support ticket evidence](/docs/operate/troubleshooting/support-ticket-evidence)
