# Volume and storage troubleshooting

Source: https://docs.quake.ai/docs/operate/runbooks/volume-troubleshooting
Markdown: https://docs.quake.ai/docs/operate/runbooks/volume-troubleshooting.md
> Diagnose volumes stuck in attaching, detaching, or error states and resolve delete conflicts.

---

# Volume and storage troubleshooting

This runbook covers block volume and snapshot failures on Quake AI, managed by the [Block Storage service (Cinder)](/resources/migration/openstack) with Compute when you attach volumes, and full-disk conditions inside instances.

<Figure size="md" caption="Where to start: pick the branch that matches your symptom, then jump to that section below.">

```d2
direction: right

start: "Volume or storage\nissue"

q_kind: "What kind\nof failure?" {shape: diamond}

leaf_attach: "Volume stuck in\nattaching or detaching"
leaf_delete: "Unable to delete\nvolume"
leaf_snap: "Volume snapshot\ncreation failure"
leaf_disk: "Disk space exhaustion\ninside instance"

start -> q_kind
q_kind -> leaf_attach: stuck attaching/detaching
q_kind -> leaf_delete: delete blocked
q_kind -> leaf_snap: snapshot stuck
q_kind -> leaf_disk: filesystem full
```

</Figure>

## Quick triage

| Symptom | Likely cause | Section |
| --- | --- | --- |
| Volume status stays **attaching** or **detaching** indefinitely | API timeout (for example HTTP 504), storage backend fault, Nova–Cinder communication failure, volume resize through Heat | [Volume stuck in attaching or detaching](#volume-stuck-in-attaching-or-detaching) |
| `openstack volume delete` reports **in-use** or dependency errors | Active attachment, orphaned attachment, snapshots, or consistency group membership | [Unable to delete volume](#unable-to-delete-volume) |
| Snapshot fails or remains **creating** | Source volume in a transitional state, backend capacity, concurrent snapshots, quota | [Volume snapshot creation failure](#volume-snapshot-creation-failure) |
| Writes fail, apps crash, root filesystem at **100%** in `df` | Undersized root disk, log or temp growth, non-resizable ephemeral root | [Disk space exhaustion inside instance](#disk-space-exhaustion-inside-instance) |

## Volume stuck in attaching or detaching

### Symptoms

- Volume status remains **attaching** or **detaching** for an extended time
- You cannot attach, detach, or delete the volume until the state clears
- A recent operation may have hit an API timeout (for example HTTP **504**), or you resized the volume through Automation (Heat)

### Diagnosis

Inspect the volume and its attachments:

```bash
openstack volume show VOLUME_ID
openstack volume show VOLUME_ID -c attachments -c status
```

Note whether `attachments` is empty while `status` is still transitional. On Quake AI, Cinder advertises API microversions from 3.0 up to **3.70** (OpenStack Antelope). Some subcommands, such as `volume attachment list`, need a microversion above the client default; see the note in the resolution steps.

Common causes include client or gateway timeouts during attach or detach, storage backend errors, broken coordination between Compute and Cinder, or stack operations that change volume state while other operations are in flight.

### Resolution

1. Run `openstack volume show VOLUME_ID` and record `status` and `attachments`.



`openstack volume set --state` calls the Cinder admin policy `volume_extension:volume_admin_actions:reset_status`. A standard project or application credential does not hold that policy, so on Quake AI the command returns `HTTP 403` for most users (`Policy doesn't allow volume_extension:volume_admin_actions:reset_status to be performed`). The self-serve resolution is a clean detach with `openstack server remove volume`, or deleting a stale attachment when the instance is gone. If neither returns the volume to **available**, open a support ticket: only an operator can reset volume state. This is the same 403-under-standard-credential pattern documented in the [auth and token diagnostics runbook](/docs/operate/troubleshooting/auth-token-diagnostics). The command stays documented below for operators who hold the policy. It overrides the recorded state without verifying the storage backend, so run it only after the backend operation has completed or failed.



2. If an attachment still points to a live instance, detach it cleanly. This is the primary self-serve fix and needs no admin policy:

   ```bash
   openstack server remove volume SERVER_NAME_OR_ID VOLUME_ID
   ```

3. If an attachment row remains but the instance no longer exists, delete the stale attachment. The `volume attachment` subcommands need Cinder microversion 3.27, so prepend `--os-volume-api-version 3.27` (or export `OS_VOLUME_API_VERSION=3.27`); without it the client reports that 3.27 or greater is required:

   ```bash
   openstack --os-volume-api-version 3.27 volume attachment list --volume VOLUME_ID
   openstack --os-volume-api-version 3.27 volume attachment delete VOLUME_ID ATTACHMENT_ID
   ```

   To read the attachment id without raising the microversion, run `openstack volume show VOLUME_ID -c attachments`.

4. If `attachments` is empty but `status` is still **attaching**, **detaching**, or **error**, the volume needs a state reset to **available**. A standard credential cannot do this: `openstack volume set --state` returns `HTTP 403` (see the callout above). Open a support ticket with the volume ID, timestamps (UTC), and the commands you ran. An operator, or a credential that holds the `reset_status` policy, clears it with:

   ```bash
   openstack volume set --state available VOLUME_ID
   ```

5. Verify by attaching the volume to a small test instance, or repeat your original attach path once state is **available** or **in-use** as expected.

### Verification

```bash
openstack volume show VOLUME_ID -c status -c attachments
```

`status` should reflect the intended operational state (**available** when detached, **in-use** when attached), and `attachments` should match the instance you expect.

### Prevention

- Avoid rapid attach and detach loops; wait for each operation to finish before you start the next.
- Give Automation (Heat) and CI pipelines timeouts that allow Cinder and Nova to complete volume operations under load.

### When to escalate

A standard credential cannot run `openstack volume set --state available` (it returns `HTTP 403`), so a stuck transitional state that a clean detach does not clear needs operator action. Open a support ticket with the volume ID, timestamps (UTC), and the last commands you ran. Do the same if attachments reappear incorrectly after you delete them.

## Unable to delete volume

### Symptoms (Unable to delete volume)

- `openstack volume delete VOLUME_ID` fails with a message that the volume is **in-use** or lists dependency constraints
- The dashboard delete action fails with a similar error

### Diagnosis (Unable to delete volume)

List attachments and metadata:

```bash
openstack volume show VOLUME_ID -c attachments -c status
openstack volume snapshot list --volume VOLUME_ID
```

A volume stays non-deletable while it is attached, while Cinder still records an attachment after the instance is gone, while snapshots exist, or while it belongs to a consistency group (if you use that feature).

### Resolution (Unable to delete volume)

1. If the volume shows attachments, detach from the instance when that instance still exists:

   ```bash
   openstack server remove volume SERVER_NAME_OR_ID VOLUME_ID
   ```

2. If the instance is gone but an attachment remains, delete the attachment. The `volume attachment` subcommands need Cinder microversion 3.27, so prepend `--os-volume-api-version 3.27` (or export `OS_VOLUME_API_VERSION=3.27`):

   ```bash
   openstack --os-volume-api-version 3.27 volume attachment list --volume VOLUME_ID
   openstack --os-volume-api-version 3.27 volume attachment delete VOLUME_ID ATTACHMENT_ID
   ```



`openstack volume set --state` calls the Cinder admin policy `volume_extension:volume_admin_actions:reset_status`, which a standard project or application credential does not hold. On Quake AI the command returns `HTTP 403` for most users, so it is not a self-serve step. Clear the attachment first; if the volume still does not return to **available**, open a support ticket so an operator can reset the state.



3. If the volume does not return to **available** after the attachment is gone, the state needs a reset. A standard credential cannot do this because the command returns `HTTP 403`, so open a support ticket. An operator runs:

   ```bash
   openstack volume set --state available VOLUME_ID
   ```

4. If snapshots exist, delete them before you delete the source volume:

   ```bash
   openstack volume snapshot list --volume VOLUME_ID
   openstack volume snapshot delete SNAPSHOT_ID
   ```

5. Retry deletion:

   ```bash
   openstack volume delete VOLUME_ID
   ```

### Verification (Unable to delete volume)

Confirm the volume no longer appears:

```bash
openstack volume list | grep VOLUME_ID
```

The command should return no rows for that ID.

### Prevention (Unable to delete volume)

- Detach volumes before you delete instances when you plan to retire the volume separately.
- Delete snapshots before you delete their source volumes when policy allows.
- Use `openstack volume show VOLUME_ID` as a preflight check before destructive changes in scripts.

### When to escalate (Unable to delete volume)

If you cannot remove an attachment, snapshots are stuck in **deleting**, or error text references platform-managed resources you do not control, include `openstack volume show VOLUME_ID` output in your support ticket.

## Volume snapshot creation failure

### Symptoms (Volume snapshot creation failure)

- Snapshot create returns an error, or the snapshot stays **creating** for a long time
- Automated backup jobs fail intermittently

### Diagnosis (Volume snapshot creation failure)

Check the source volume state and project quota:

```bash
openstack volume show VOLUME_ID -c status
openstack quota show --usage
```

Snapshots require the source volume to be in **available** or **in-use**, not in other transitional states. Backend capacity, too many concurrent snapshot jobs, or exhausted volume or gigabyte quota can also block creation.

### Resolution (Volume snapshot creation failure)

1. If the source volume is **creating**, **attaching**, **detaching**, **extending**, or similar, wait until it settles, then retry the snapshot.
2. If the snapshot remains **creating** for more than about 10 minutes, re-check quota with `openstack quota show --usage` and free gigabytes or snapshot count if limits are tight.
3. Retry snapshot creation during a quieter window if you suspect concurrent load on the storage backend.
4. If failures persist after the source volume is stable and quota is healthy, open a support ticket; you may need operator action on the backend.

### Verification (Volume snapshot creation failure)

```bash
openstack volume snapshot list --volume VOLUME_ID
```

The new snapshot should reach **available** (or your platform’s equivalent success state) within a reasonable time.

### Prevention (Volume snapshot creation failure)

- Do not start snapshots during resize, attach, or detach operations.
- Monitor volume and snapshot quota the same way you monitor instance quota.

### When to escalate (Volume snapshot creation failure)

Escalate when stable volumes with free quota still cannot produce snapshots, or when snapshots stay **creating** well beyond 10 minutes after you have waited for other operations to finish.

## Disk space exhaustion inside instance

### Symptoms (Disk space exhaustion inside instance)

- Applications log write errors or exit without a clean shutdown
- SSH becomes slow or unresponsive on the instance
- `df -h` shows the root filesystem (or another mount) at **100%**

### Diagnosis (Disk space exhaustion inside instance)

On the instance (SSH or console), identify which filesystem is full and what consumes space:

```bash
df -h
sudo du -sh /var/log/* 2>/dev/null | sort -h
```

Typical causes are a root disk smaller than your workload needs, uncompressed or unrotated logs, large caches or temporary files, or using the ephemeral root disk for data that grows without bound. Root disk size follows the **flavor**; you do not resize the root disk the way you extend an attached Cinder volume.

### Resolution (Disk space exhaustion inside instance)

1. Confirm the full mount with `df -h`.
2. Inspect logs:

   ```bash
   sudo du -sh /var/log/* | sort -h
   ```

3. Trim systemd journal size if journals dominate disk use:

   ```bash
   sudo journalctl --vacuum-size=100M
   ```

4. Remove rotated logs only when you accept the data loss (for example old compressed archives):

   ```bash
   sudo rm /var/log/*.gz
   ```

5. Clear safe temporary directories (adjust paths to match your workload).
6. For durable extra capacity, attach a Cinder volume for application data and move large directories there; see [Create a block volume](/docs/block/how-to/create-volume) and [Extend a block volume](/docs/block/how-to/extend-volume).
7. If SSH is not usable, open the instance console:

   ```bash
   openstack console url show SERVER_NAME_OR_ID
   ```

   Complete cleanup from the console, or repair disk usage enough to restore SSH.

### Verification (Disk space exhaustion inside instance)

Re-run:

```bash
df -h
```

Free space should be a comfortable margin above zero on critical mounts, and services should write normally again.

### Prevention (Disk space exhaustion inside instance)

- Store growing data on Cinder volumes, not only on the root disk.
- Configure **logrotate** (or equivalent) for application logs.
- Monitor disk usage with your usual observability stack.

### When to escalate (Disk space exhaustion inside instance)

Escalate when you suspect filesystem corruption, read-only remounts you cannot clear, or platform issues unrelated to ordinary cleanup (for example missing volumes after reboot). Include `df -h` output and, if relevant, instance and volume IDs.

## Collecting evidence for support tickets

| Item | Command or value |
| --- | --- |
| Volume ID, status, attachments | `openstack volume show VOLUME_ID -c id -c status -c attachments` |
| Attachment list | `openstack --os-volume-api-version 3.27 volume attachment list --volume VOLUME_ID` |
| Snapshots for volume | `openstack volume snapshot list --volume VOLUME_ID` |
| Related instance | `openstack server show SERVER_NAME_OR_ID -c id -c status` |
| Quota usage | `openstack quota show --usage` |
| Timestamps | Operation start and failure times in **UTC** |
| In-guest disk use (if applicable) | `df -h` and, if you can collect it, `sudo du -sh /var/log/*` highlights |

For a broader checklist, see [Support ticket evidence](/docs/operate/troubleshooting/support-ticket-evidence).

## See also

- [Block storage API error reference](/reference/block-storage/api-errors): HTTP status codes, volume state machine, and delete conflicts
- [Create a block volume](/docs/block/how-to/create-volume)
- [Extend a block volume](/docs/block/how-to/extend-volume)
- [Create a volume snapshot](/docs/block/how-to/create-snapshot)
- [Support ticket evidence](/docs/operate/troubleshooting/support-ticket-evidence)
- [Instance lifecycle troubleshooting](/docs/operate/runbooks/instance-lifecycle)
