Skip to content

How to Set Up Multi-Cloud Object Storage Sync

How-to · Updated Sep 2026
Before this

How to set up multi-cloud object storage sync

Keep data synchronized between Quake AI and another S3-compatible provider. This guide covers active-passive replication, selective prefix sync, and monitoring for continuous sync workflows.

Prerequisites#

  • A Quake AI account with S3 credentials
  • Credentials for your other S3-compatible provider
  • rclone installed (v1.65+) on a server or VM that runs continuously
  • cron or systemd for scheduling

Configure rclone remotes#

Set up both providers in ~/.config/rclone/rclone.conf. This example uses AWS S3 as the primary, but any S3-compatible provider works identically.

ini
[primary]
type = s3
provider = AWS
access_key_id = YOUR_PRIMARY_ACCESS_KEY
secret_access_key = YOUR_PRIMARY_SECRET_KEY
region = us-east-1

[quakeai]
type = s3
provider = Ceph
access_key_id = YOUR_RUMBLE_ACCESS_KEY
secret_access_key = YOUR_RUMBLE_SECRET_KEY
endpoint = object.us-east-2.rumble.cloud
acl = private

See tools comparison for provider-specific rclone configuration.

Active-passive replication#

The most common pattern: your application writes to the primary provider, and rclone syncs changes to Quake AI on a schedule. Quake AI is a read-only replica and backup target.

Scheduled sync#

bash
rclone sync primary:my-bucket quakeai:my-bucket-replica \
  --transfers 16 \
  --checkers 32 \
  --progress \
  --log-file /var/log/rclone-sync.log \
  --log-level INFO

Add to cron for hourly replication:

0 * * * * /usr/bin/rclone sync primary:my-bucket quakeai:my-bucket-replica --transfers 16 --checkers 32 --log-file /var/log/rclone-sync.log --log-level INFO

Failover#

If the primary becomes unavailable, point your application at the Quake AI replica. See update application endpoint for SDK configuration changes. After the primary recovers, reverse the sync direction to bring it up to date:

bash
rclone sync quakeai:my-bucket-replica primary:my-bucket \
  --transfers 16 \
  --checkers 32 \
  --progress

Then revert your application to the primary endpoint.

Selective prefix sync#

Replicate only specific prefixes (e.g., backups and media) while leaving other data on the primary only:

bash
rclone sync primary:my-bucket quakeai:my-bucket-replica \
  --include "/backups/**" \
  --include "/media/**" \
  --transfers 16 \
  --progress

Everything outside /backups/ and /media/ is ignored.

Exclude patterns#

Alternatively, sync everything except specific prefixes:

bash
rclone sync primary:my-bucket quakeai:my-bucket-replica \
  --exclude "/tmp/**" \
  --exclude "/cache/**" \
  --transfers 16 \
  --progress

Bandwidth management#

Throttle during business hours#

bash
rclone sync primary:my-bucket quakeai:my-bucket-replica \
  --bwlimit "08:00,10M 18:00,off" \
  --transfers 16

Limits bandwidth to 10 MB/s between 08:00 and 18:00, unlimited outside those hours.

Fixed bandwidth cap#

bash
rclone sync primary:my-bucket quakeai:my-bucket-replica \
  --bwlimit 50M \
  --transfers 8

Monitoring#

Log-based monitoring#

rclone writes structured logs when --log-level INFO or higher is set. Monitor the log file for errors:

bash
#!/bin/bash
LOG="/var/log/rclone-sync.log"
ERRORS=$(grep -c "ERROR" "$LOG")
if [ "$ERRORS" -gt 0 ]; then
  echo "rclone sync had $ERRORS errors" | mail -s "Sync Alert" [email protected]
fi

Exit code checking#

rclone returns exit code 0 on success. Wrap cron jobs in a script that alerts on failure:

bash
#!/bin/bash
rclone sync primary:my-bucket quakeai:my-bucket-replica \
  --transfers 16 \
  --log-file /var/log/rclone-sync.log \
  --log-level INFO

if [ $? -ne 0 ]; then
  echo "rclone sync failed at $(date)" >> /var/log/rclone-alerts.log
  # Send alert via webhook, email, or monitoring system
fi

Drift detection#

Run a periodic check to detect objects that are out of sync without transferring data:

bash
rclone check primary:my-bucket quakeai:my-bucket-replica \
  --one-way \
  --log-file /var/log/rclone-check.log

Schedule weekly via cron. Review the log for mismatches.

Architecture considerations#

Where to run the sync process#

OptionProsCons
Quake AI VMClose to destination, low-latency writes to Quake AIEgress from primary provider costs money
Primary provider VMFree egress to internet from most providersHigher latency writing to Quake AI
Dedicated sync serverFull control, can throttle independentlyAdditional infrastructure to manage

For most setups, run the sync process on a small Quake AI VM. Data writes to Quake AI over the local network, and reads from the primary incur standard egress.

Conflict handling#

rclone sync is one-directional: it makes the destination match the source. If both sides are written to independently, the last sync overwrites the destination.

For workloads where both sides receive writes:

  • Use append-only patterns (timestamped filenames) so objects never conflict
  • Partition by prefix: one provider owns /region-a/, the other owns /region-b/
  • Use application-level conflict resolution rather than relying on storage-layer sync

See also#

Usage Guidelines

The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.

For the full policy, see Usage Guidelines.

Last validated: 08.09.2026

Quick answers

Was this page helpful?