Scientific computing and research
Scientific computing and research
Run simulation, analysis, and research workloads on Quake AI. You operate job schedulers, simulation binaries, and result stores; Quake AI provides CPU-only Compute, Block Storage, Object Storage, private networking, and optional self-managed Kubernetes. Quake AI flavors are AMD EPYC with no GPU option, which fits CPU-bound simulation and post-processing; see the compute FAQ for flavor details.
What this is for#
Research teams run batch simulation, analysis, and post-processing on infrastructure with predictable cost and no per-gigabyte data-transfer charges. Quake AI provides CPU compute for workers, Block Storage for scratch, and object storage for inputs and results; you operate the scheduler, simulation binaries, and result stores. The outcome is a self-operated CPU batch environment, deployable from validated OpenTofu templates and their companion tutorials.
Reference architecture#
Download diagram: SVG, PNG, and PDF.
The base is the Render farm worker template: a scheduler or login node and CPU batch workers behind a floating IP. The S3 Storage with ACLs template adds the bucket for input decks and result archives. Attach Block Storage scratch volumes to the workers as needed, and a Kubernetes cluster is the alternative when you schedule batch work on a self-managed cluster.
-
Researchers. Researchers and automation submit jobs to a scheduler or login node.
-
Scheduler or login node. A Compute instance runs the job scheduler (Slurm, PBS, or a workflow engine) or a login node. The Render farm worker template provisions the scheduler node and workers together; a Simple VM is the alternative for a standalone runner or login host.
-
Batch compute workers. Simulation and analysis jobs run on the Compute workers, optionally orchestrated on a self-managed Kubernetes cluster from the Kubernetes cluster template.
-
Scratch storage. Working data and intermediate files use Block Storage volumes attached to the workers.
-
Inputs and results. Input decks, mesh files, reference datasets, and result archives live in S3-compatible Object Storage.
Services involved#
| Service | Role in this architecture | Docs |
|---|---|---|
| Compute | Scheduler, login node, and batch workers (CPU) | Compute |
| Block Storage | Scratch and intermediate-file volumes | Block Storage |
| Object Storage | Input datasets and result archives | Object Storage |
| Network | Private networks and security groups for the cluster | Network |
| Kubernetes (Magnum) | Optional batch orchestration | Kubernetes |
Get started#
- Simple VM template and its deploy tutorial: single-instance OpenTofu pattern for standalone simulation runners or login nodes.
- Kubernetes cluster template and its deploy tutorial: self-managed cluster for batch jobs you deploy with manifests, operators, or workflow engines.
- Upload objects to Object Storage: stage input decks, checkpoints, and result archives in the bucket your template provisions.
- OpenTofu template library: browse validated IaC starting points for VM and cluster layouts.
Estimate the cost#
Monthly cost estimate
Pricing calculator ↗Sized as a custom package on dedicated vCPU.
Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.
What each resource is for
Dispatcher
c2a.large · 2 dedicated vCPU, 4 GiB RAM, 0.5 Gbps
2× Worker node
c2a.large · 2 dedicated vCPU, 4 GiB RAM, 0.5 Gbps
Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.
Included in baseline
c2a.large
2 dedicated vCPU, 4 GiB RAM, 0.5 Gbps
c2a.large
2 dedicated vCPU, 4 GiB RAM, 0.5 Gbps
c2a.large
2 dedicated vCPU, 4 GiB RAM, 0.5 Gbps
Compute + RAM rate basis
6 vCPU + 12 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.
Block storage (200 GiB)
200 GiB at $0.08/GiB/mo
Package promotional discount
Flat −$5.00/mo on the custom package (same promotion as named plans).
Object storage (usage-based)
Object storage
2 buckets. The first 1 TB is included, then $10.00 per TB each month. You pay for what you store, so this line depends on usage.
Assumes: 2 TB stored is $10/mo; 5 TB stored is $40/mo. Within the included allotment it stays $0.
Included at no charge
These line items are zero on Quake AI. Many other providers meter them separately.
Data transfer (inbound and outbound)
Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.
AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.
Learn moreObject storage upload and download
No separate charges for uploading or downloading object storage data.
Most object storage providers meter egress and API requests separately from stored capacity.
Learn morePrivate networking
Private networks, subnets, Neutron routers, and security groups are included with the plan.
VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.
Control-plane API requests
OpenStack API calls for provisioning and management are included.
Some managed services on other clouds meter API calls or charge for premium control-plane features.
Configure your estimate
Check the add-ons you plan to deploy to build a monthly total. Nothing is selected to start, so the total below begins at the baseline.
Starting template
The required baseline, always included.
Pick how much you expect to store to fold it into the total.
Dev/test vs production
Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.
Dev/test on shared CPU
Burstable s1a flavors; suited to prototyping and low or bursty load.
Production on dedicated CPU
The headline estimate above; predictable steady-load performance.
Saves $136.50/mo while you build on shared CPU.
Shared flavors carry less RAM (c2a.large (4 GiB RAM) -> s1a.small (2 GiB RAM); c2a.large (4 GiB RAM) -> s1a.small (2 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.
Pricing data last validated: . For current rates, check quake.ai/pricing.
Migrating an existing scientific computing / HPC-style workload?#
Move a simulation, analysis, or research pipeline that already runs on AWS EC2 and S3, Azure VMs and Blob, Google Compute Engine and Cloud Storage, Slurm or Kubernetes batch queues on Amazon EKS, or another hyperscaler fleet. The outcome is the same workloads on Quake AI CPU compute with input datasets, job definitions, and result stores cut over in a controlled order.
Follow this cutover path. Each step links an existing migration page; this section composes those pages into a workload-shaped sequence rather than duplicating their steps.
-
Map your source provider. Start with the concept-translation page for your current cloud: Coming from AWS, Coming from Azure, Coming from GCP, Coming from DigitalOcean, or Coming from Hetzner.
-
Stand up the target shape on Quake AI. Pick the greenfield template that matches how you run jobs after migration: Simple VM for dedicated runners or a small login plus worker layout, or Kubernetes cluster when you schedule batch work on a self-managed cluster.
-
Move compute workloads. Rebuild or migrate simulation nodes and batch workers with Migrate from EC2 (or the matching compute migration page for your source provider). For Kubernetes-hosted batch services, follow Migrate from EKS.
-
Move object and dataset data. Sync input decks, mesh files, reference genomes, and checkpoint objects with Migrate from S3 (or the matching object migration page for your source provider).
-
Cut over job submission and result access. Stage the Quake AI scheduler or API endpoint on a private hostname or an API gateway with a floating IP, run validation jobs against migrated inputs, then switch researchers and automation to the new submission URL and result-store base path.
Workload-specific cutover callouts#
- Large dataset transfer. Plan bandwidth and resume behavior for multi-terabyte inputs and archives. Use the object migration page's sync tooling, validate checksums on a sample of files, and keep read access to the source bucket until downstream jobs finish on Quake AI.
- Job scheduler and runner rebuild. Slurm, PBS, Airflow, Argo Workflows, or custom runners rarely lift verbatim across clouds. Recreate queues, resource limits, and container images on the target template, then replay a representative job mix before you retire the source cluster.
- Result store endpoint move. Update job scripts, notebooks, and CI pipelines to write and read from the Quake AI Object Storage endpoint and paths you provisioned. Confirm permissions, lifecycle rules, and any signed-URL behavior match what downstream analysis expects.
- CPU-only workloads. CPU-bound simulation, preprocessing, and post-processing fit Quake AI (compute FAQ). GPU-accelerated simulation or training that requires CUDA at production scale needs accelerators outside this flavor catalog.
Considerations and limits#
- You operate the scheduler and binaries. Quake AI provides compute, network, and storage; the scheduler, simulation software, and result stores are yours under the shared responsibility model.
- CPU-only compute. Compute is AMD EPYC with no GPU option (compute FAQ). CPU-bound simulation and post-processing fit; CUDA workloads need a different hosting path.
- Flat egress. Quake AI applies a no-egress-fee policy for outbound transfer, which suits large result-set downloads.
- Three US regions. All current regions are in the United States.
- No managed HPC scheduler. Slurm, PBS, or workflow engines run on Compute you operate; Quake AI has no managed batch-queue service.
- Compliance posture. Quake AI holds SOC 2 Type I and Type II attestations and SOC 3. See Compliance and certifications for the platform scope.