Skip to content
Solutions

AI/ML platforms

AI/ML platforms

Serve and fine-tune smaller machine learning models on Quake AI Compute. Quake AI flavors are CPU-only AMD EPYC with no GPU option, which fits inference, model serving, and smaller-model fine-tuning. See the compute FAQ for flavor details.

You operate training and serving pipelines, feature stores, and model registries; Quake AI provides compute, networking, block and object storage, and optional self-managed Kubernetes.

What this is for#

ML platform teams run experiment tracking, model registries, batch pipelines, and serving on infrastructure they control. Quake AI provides a self-managed Kubernetes cluster, CPU compute, a Postgres tier for metadata, and object storage for datasets and models; you operate the platform components (Kubeflow, MLflow, or custom controllers). The outcome is a self-hosted ML platform for CPU inference and smaller-model work, deployable from validated OpenTofu templates and their companion tutorials.

Reference architecture#

ML engineersQuake AIk8s-cluster templateSelf-managed Postgres (registry, tracking)Object Storage (datasets, models)Magnum control planeWorker nodes (serving, pipelines) schedule workloadsmetadata, registrydatasets, modelssubmit jobs, serve
Click to zoom
ML platform on Quake AI: the solid box is the k8s-cluster base template (Magnum control plane and worker nodes for serving and pipelines); a self-managed Postgres holds registry and tracking metadata, and Object Storage holds datasets and models.

Download diagram: SVG, PNG, and PDF.

The platform is four architecture decisions. Each maps to a tier in the diagram and to the template that builds it.

  1. ML engineers. Engineers submit training and batch jobs and deploy serving endpoints to the cluster.

  2. Magnum cluster. Serving controllers, pipeline runners, and batch jobs run on a self-managed Kubernetes cluster (Magnum). The Kubernetes cluster template provisions it.

  3. Metadata store. Experiment tracking, the model registry, and pipeline state live in self-managed Postgres on Block Storage. The self-managed PostgreSQL template scaffolds it.

  4. Object Storage. Training sets, processed features, serialized models, and container images live in S3-compatible Object Storage, which the cluster reads and writes across pipeline stages.

Services involved#

ServiceRole in this architectureDocs
Kubernetes (Magnum)Cluster for serving, pipelines, and batch jobsKubernetes
ComputeWorker nodes for training and serving (CPU)Compute
Block StorageMetadata store and dataset volumesBlock Storage
Object StorageDatasets, models, and container imagesObject Storage
Self-managed PostgresExperiment tracking and model registrySelf-managed PostgreSQL template

Get started#

Estimate the cost#

Monthly cost estimate

Pricing calculator ↗

Sized as a custom package on a mix of shared and dedicated vCPU.

Starting template$238.80/mo

Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.

What each resource is for

Control plane

m2a.xlarge · 4 dedicated vCPU, 16 GiB RAM, 1 Gbps

$132.00/mo

3× Worker node

s1a.medium · 4 shared vCPU, 4 GiB RAM, 0.5 Gbps

$99.00/mo

Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.

Included in baseline

m2a.xlarge

4 dedicated vCPU, 16 GiB RAM, 1 Gbps

$132.00

s1a.medium

4 shared vCPU, 4 GiB RAM, 0.5 Gbps

$33.00

s1a.medium

4 shared vCPU, 4 GiB RAM, 0.5 Gbps

$33.00

s1a.medium

4 shared vCPU, 4 GiB RAM, 0.5 Gbps

$33.00

Compute + RAM rate basis

16 vCPU + 28 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.

—

Block storage (160 GiB)

160 GiB at $0.08/GiB/mo

$12.80

Public IP (included)

1 included with the custom package

$0.00

Package promotional discount

Flat −$5.00/mo on the custom package (same promotion as named plans).

$-5.00

Included at no charge

These line items are zero on Quake AI. Many other providers meter them separately.

Data transfer (inbound and outbound)

Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.

AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.

Learn more
$0.00

Private networking

Private networks, subnets, Neutron routers, and security groups are included with the plan.

VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.

$0.00

Control-plane API requests

OpenStack API calls for provisioning and management are included.

Some managed services on other clouds meter API calls or charge for premium control-plane features.

$0.00

Kubernetes control plane

Magnum clusters run on Nova instances; there is no separate K8s platform fee in Quake AI pricing.

Managed Kubernetes on AWS, GCP, and Azure charges a control-plane fee on top of worker nodes.

Learn more
$0.00

Configure your estimate

Check the add-ons you plan to deploy to build a monthly total. Nothing is selected to start, so the total below begins at the baseline.

Starting template

The required baseline, always included.

$238.80/mo
Your configured estimate$238.80/mo

Dev/test vs production

Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.

Dev/test on shared CPU

Burstable s1a flavors; suited to prototyping and low or bursty load.

$139.80/mo

Production on the configured CPU

The headline estimate above; predictable steady-load performance.

$238.80/mo

Saves $99.00/mo while you build on shared CPU.

Shared flavors carry less RAM (m2a.xlarge (16 GiB RAM) -> s1a.medium (4 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.

Pricing data last validated: . For current rates, check quake.ai/pricing.

Migrating an existing AI/ML platform?#

Move an ML platform team that already runs on Amazon SageMaker, Google Vertex AI, Azure Machine Learning, or a self-managed stack on EKS, GKE, or VM fleets at another cloud. The outcome is the same models, datasets, and pipeline definitions on Quake AI compute or a self-managed Kubernetes cluster, with object storage, metadata stores, and inference endpoints cut over in a controlled order.

Follow this cutover path. Each step links an existing migration page; this section composes those pages into a workload-shaped sequence rather than duplicating their steps.

  1. Map your source provider. Start with the concept-translation page for your current cloud: Coming from AWS, Coming from Azure, Coming from GCP, Coming from DigitalOcean, or Coming from Hetzner.

  2. Stand up the target shape on Quake AI. Provision the migration target with the Kubernetes cluster template when serving and orchestration run on Kubernetes, and add Self-managed PostgreSQL when experiment metadata or a model registry lives in Postgres.

  3. Move compute workloads. For notebook servers, batch workers, or single-node serving on VMs, follow Migrate from EC2. For Kubernetes-hosted training jobs, operators, or inference deployments, follow Migrate from EKS.

  4. Move datasets, artifacts, and registry objects. Sync training sets, exported models, container images, and checkpoint prefixes with Migrate from S3 (or the matching object migration page for your source provider).

  5. Cut over cluster traffic and inference endpoints. Publish the serving layer through a Kubernetes LoadBalancer service or an API gateway on a floating IP. Verify application health checks, then switch DNS to the public address when error rates stay within budget.

Workload-specific cutover callouts#

  • Dataset and model-registry move. Copy or sync object-storage prefixes for raw data, processed features, and serialized models before you drain the source environment. Export registry metadata (model versions, stage labels, lineage fields) into Postgres or your registry's target schema, then validate row counts and artifact URIs on Quake AI.
  • Pipeline portability. Reconcile pipeline definitions (Kubeflow, Airflow, Argo, or vendor export formats) against CPU-only workers and your target storage classes. Run dry-run jobs on Quake AI before you schedule production batches on the new cluster.
  • Cluster cutover. Mirror container images to a registry the Quake AI cluster can pull from, reschedule inference Deployments or custom resources, and drain source nodes only after the Quake AI serving layer passes load tests. Keep the source cluster available for rollback until traffic stabilizes.
  • CPU-only workloads. Plan Quake AI for CPU inference, smaller-model fine-tuning, batch scoring, and control-plane services (compute FAQ). Large-scale GPU training, multi-billion-parameter fine-tuning, and latency-sensitive GPU inference need accelerators outside this flavor catalog.

Considerations and limits#

  • You operate the platform. Quake AI provides the cluster, compute, and storage; the ML platform components, pipelines, and registries are yours under the shared responsibility model.
  • CPU-only compute. Compute is AMD EPYC with no GPU option (compute FAQ). Plan for CPU inference, smaller-model fine-tuning, and batch scoring.
  • No managed database. Experiment tracking and the model registry run on self-managed Postgres on Compute and Block Storage.
  • Flat egress. Quake AI applies a no-egress-fee policy for outbound transfer, which suits dataset and model pulls.
  • Three US regions. All current regions are in the United States.
  • Compliance posture. Quake AI holds SOC 2 Type I and Type II attestations and SOC 3. See Compliance and certifications for the platform scope.
Was this page helpful?