Volume and storage troubleshooting
Volume and storage troubleshooting
This runbook covers block volume and snapshot failures on Quake AI, managed by the Block Storage service (Cinder) with Compute when you attach volumes, and full-disk conditions inside instances.
Quick triage#
| Symptom | Likely cause | Section |
|---|---|---|
| Volume status stays attaching or detaching indefinitely | API timeout (for example HTTP 504), storage backend fault, Nova–Cinder communication failure, volume resize through Heat | Volume stuck in attaching or detaching |
openstack volume delete reports in-use or dependency errors | Active attachment, orphaned attachment, snapshots, or consistency group membership | Unable to delete volume |
| Snapshot fails or remains creating | Source volume in a transitional state, backend capacity, concurrent snapshots, quota | Volume snapshot creation failure |
Writes fail, apps crash, root filesystem at 100% in df | Undersized root disk, log or temp growth, non-resizable ephemeral root | Disk space exhaustion inside instance |
Volume stuck in attaching or detaching#
Symptoms#
- Volume status remains attaching or detaching for an extended time
- You cannot attach, detach, or delete the volume until the state clears
- A recent operation may have hit an API timeout (for example HTTP 504), or you resized the volume through Automation (Heat)
Diagnosis#
Inspect the volume and its attachments:
openstack volume show VOLUME_ID
openstack volume show VOLUME_ID -c attachments -c statusNote whether attachments is empty while status is still transitional. On Quake AI, Cinder advertises API microversions from 3.0 up to 3.70 (OpenStack Antelope). Some subcommands, such as volume attachment list, need a microversion above the client default; see the note in the resolution steps.
Common causes include client or gateway timeouts during attach or detach, storage backend errors, broken coordination between Compute and Cinder, or stack operations that change volume state while other operations are in flight.
Resolution#
- Run
openstack volume show VOLUME_IDand recordstatusandattachments.
-
If an attachment still points to a live instance, detach it cleanly. This is the primary self-serve fix and needs no admin policy:
bashopenstack server remove volume SERVER_NAME_OR_ID VOLUME_ID -
If an attachment row remains but the instance no longer exists, delete the stale attachment. The
volume attachmentsubcommands need Cinder microversion 3.27, so prepend--os-volume-api-version 3.27(or exportOS_VOLUME_API_VERSION=3.27); without it the client reports that 3.27 or greater is required:bashopenstack --os-volume-api-version 3.27 volume attachment list --volume VOLUME_ID openstack --os-volume-api-version 3.27 volume attachment delete VOLUME_ID ATTACHMENT_IDTo read the attachment id without raising the microversion, run
openstack volume show VOLUME_ID -c attachments. -
If
attachmentsis empty butstatusis still attaching, detaching, or error, the volume needs a state reset to available. A standard credential cannot do this:openstack volume set --statereturnsHTTP 403(see the callout above). Open a support ticket with the volume ID, timestamps (UTC), and the commands you ran. An operator, or a credential that holds thereset_statuspolicy, clears it with:bashopenstack volume set --state available VOLUME_ID -
Verify by attaching the volume to a small test instance, or repeat your original attach path once state is available or in-use as expected.
Verification#
openstack volume show VOLUME_ID -c status -c attachmentsstatus should reflect the intended operational state (available when detached, in-use when attached), and attachments should match the instance you expect.
Prevention#
- Avoid rapid attach and detach loops; wait for each operation to finish before you start the next.
- Give Automation (Heat) and CI pipelines timeouts that allow Cinder and Nova to complete volume operations under load.
When to escalate#
A standard credential cannot run openstack volume set --state available (it returns HTTP 403), so a stuck transitional state that a clean detach does not clear needs operator action. Open a support ticket with the volume ID, timestamps (UTC), and the last commands you ran. Do the same if attachments reappear incorrectly after you delete them.
Unable to delete volume#
Symptoms (Unable to delete volume)#
openstack volume delete VOLUME_IDfails with a message that the volume is in-use or lists dependency constraints- The dashboard delete action fails with a similar error
Diagnosis (Unable to delete volume)#
List attachments and metadata:
openstack volume show VOLUME_ID -c attachments -c status
openstack volume snapshot list --volume VOLUME_IDA volume stays non-deletable while it is attached, while Cinder still records an attachment after the instance is gone, while snapshots exist, or while it belongs to a consistency group (if you use that feature).
Resolution (Unable to delete volume)#
-
If the volume shows attachments, detach from the instance when that instance still exists:
bashopenstack server remove volume SERVER_NAME_OR_ID VOLUME_ID -
If the instance is gone but an attachment remains, delete the attachment. The
volume attachmentsubcommands need Cinder microversion 3.27, so prepend--os-volume-api-version 3.27(or exportOS_VOLUME_API_VERSION=3.27):bashopenstack --os-volume-api-version 3.27 volume attachment list --volume VOLUME_ID openstack --os-volume-api-version 3.27 volume attachment delete VOLUME_ID ATTACHMENT_ID
-
If the volume does not return to available after the attachment is gone, the state needs a reset. A standard credential cannot do this because the command returns
HTTP 403, so open a support ticket. An operator runs:bashopenstack volume set --state available VOLUME_ID -
If snapshots exist, delete them before you delete the source volume:
bashopenstack volume snapshot list --volume VOLUME_ID openstack volume snapshot delete SNAPSHOT_ID -
Retry deletion:
bashopenstack volume delete VOLUME_ID
Verification (Unable to delete volume)#
Confirm the volume no longer appears:
openstack volume list | grep VOLUME_IDThe command should return no rows for that ID.
Prevention (Unable to delete volume)#
- Detach volumes before you delete instances when you plan to retire the volume separately.
- Delete snapshots before you delete their source volumes when policy allows.
- Use
openstack volume show VOLUME_IDas a preflight check before destructive changes in scripts.
When to escalate (Unable to delete volume)#
If you cannot remove an attachment, snapshots are stuck in deleting, or error text references platform-managed resources you do not control, include openstack volume show VOLUME_ID output in your support ticket.
Volume snapshot creation failure#
Symptoms (Volume snapshot creation failure)#
- Snapshot create returns an error, or the snapshot stays creating for a long time
- Automated backup jobs fail intermittently
Diagnosis (Volume snapshot creation failure)#
Check the source volume state and project quota:
openstack volume show VOLUME_ID -c status
openstack quota show --usageSnapshots require the source volume to be in available or in-use, not in other transitional states. Backend capacity, too many concurrent snapshot jobs, or exhausted volume or gigabyte quota can also block creation.
Resolution (Volume snapshot creation failure)#
- If the source volume is creating, attaching, detaching, extending, or similar, wait until it settles, then retry the snapshot.
- If the snapshot remains creating for more than about 10 minutes, re-check quota with
openstack quota show --usageand free gigabytes or snapshot count if limits are tight. - Retry snapshot creation during a quieter window if you suspect concurrent load on the storage backend.
- If failures persist after the source volume is stable and quota is healthy, open a support ticket; you may need operator action on the backend.
Verification (Volume snapshot creation failure)#
openstack volume snapshot list --volume VOLUME_IDThe new snapshot should reach available (or your platform’s equivalent success state) within a reasonable time.
Prevention (Volume snapshot creation failure)#
- Do not start snapshots during resize, attach, or detach operations.
- Monitor volume and snapshot quota the same way you monitor instance quota.
When to escalate (Volume snapshot creation failure)#
Escalate when stable volumes with free quota still cannot produce snapshots, or when snapshots stay creating well beyond 10 minutes after you have waited for other operations to finish.
Disk space exhaustion inside instance#
Symptoms (Disk space exhaustion inside instance)#
- Applications log write errors or exit without a clean shutdown
- SSH becomes slow or unresponsive on the instance
df -hshows the root filesystem (or another mount) at 100%
Diagnosis (Disk space exhaustion inside instance)#
On the instance (SSH or console), identify which filesystem is full and what consumes space:
df -h
sudo du -sh /var/log/* 2>/dev/null | sort -hTypical causes are a root disk smaller than your workload needs, uncompressed or unrotated logs, large caches or temporary files, or using the ephemeral root disk for data that grows without bound. Root disk size follows the flavor; you do not resize the root disk the way you extend an attached Cinder volume.
Resolution (Disk space exhaustion inside instance)#
-
Confirm the full mount with
df -h. -
Inspect logs:
bashsudo du -sh /var/log/* | sort -h -
Trim systemd journal size if journals dominate disk use:
bashsudo journalctl --vacuum-size=100M -
Remove rotated logs only when you accept the data loss (for example old compressed archives):
bashsudo rm /var/log/*.gz -
Clear safe temporary directories (adjust paths to match your workload).
-
For durable extra capacity, attach a Cinder volume for application data and move large directories there; see Create a block volume and Extend a block volume.
-
If SSH is not usable, open the instance console:
bashopenstack console url show SERVER_NAME_OR_IDComplete cleanup from the console, or repair disk usage enough to restore SSH.
Verification (Disk space exhaustion inside instance)#
Re-run:
df -hFree space should be a comfortable margin above zero on critical mounts, and services should write normally again.
Prevention (Disk space exhaustion inside instance)#
- Store growing data on Cinder volumes, not only on the root disk.
- Configure logrotate (or equivalent) for application logs.
- Monitor disk usage with your usual observability stack.
When to escalate (Disk space exhaustion inside instance)#
Escalate when you suspect filesystem corruption, read-only remounts you cannot clear, or platform issues unrelated to ordinary cleanup (for example missing volumes after reboot). Include df -h output and, if relevant, instance and volume IDs.
Collecting evidence for support tickets#
| Item | Command or value |
|---|---|
| Volume ID, status, attachments | openstack volume show VOLUME_ID -c id -c status -c attachments |
| Attachment list | openstack --os-volume-api-version 3.27 volume attachment list --volume VOLUME_ID |
| Snapshots for volume | openstack volume snapshot list --volume VOLUME_ID |
| Related instance | openstack server show SERVER_NAME_OR_ID -c id -c status |
| Quota usage | openstack quota show --usage |
| Timestamps | Operation start and failure times in UTC |
| In-guest disk use (if applicable) | df -h and, if you can collect it, sudo du -sh /var/log/* highlights |
For a broader checklist, see Support ticket evidence.
See also#
- Block storage API error reference: HTTP status codes, volume state machine, and delete conflicts
- Create a block volume
- Extend a block volume
- Create a volume snapshot
- Support ticket evidence
- Instance lifecycle troubleshooting
Usage Guidelines
The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.
For the full policy, see Usage Guidelines.
Last validated: 19.06.2026