Skip to content
Solutions

Data pipelines and analytics

Data pipelines and analytics

Run ETL and ELT pipelines and an analytics warehouse on Quake AI by composing validated OpenTofu templates. Apache Airflow schedules ingestion and transform jobs, Object Storage holds the data lake, a self-managed Postgres or ClickHouse warehouse stores query-ready tables, and Metabase or Apache Superset serves dashboards. You operate the orchestration, transform code, and BI layer; Quake AI provides the compute, network, and storage those templates provision.

What this is for#

Data engineering teams that want cost-stable pipelines outside hyperscaler-managed services compose a stack from first-party templates: Airflow for orchestration, dbt for SQL transforms, Object Storage for the lake, a self-managed warehouse, and a BI tool for analysts. Quake AI provides those templates with predictable cost and no per-query warehouse pricing; you operate Airflow, dbt project code, warehouse tuning, and dashboard access. The outcome is a self-hosted ETL or ELT pipeline and analytics warehouse you deploy from the template library and its companion tutorials.

Reference architecture#

Source systemsQuake AIairflow templates3-storage-acl templateself-managed-postgres templatesuperset templateoptional add-onsAirflow (orchestration + dbt jobs)Object Storage data lakePostgres warehouseSuperset BIAirbyte (ingestion)Redpanda (streaming)ClickHouse warehouse land rawread rawwrite curatedload tablesqueryextractreplicatestream events
Click to zoom
Data pipeline on Quake AI: the solid boxes are validated templates (Airflow orchestration, s3-storage-acl Object Storage lake, self-managed-postgres warehouse, Superset BI). Airflow schedules dbt transform jobs that read and write the lake and load the warehouse. The dashed box lists optional add-ons (Airbyte ingestion, Redpanda streaming, ClickHouse as an alternate warehouse).

Download diagram: SVG, PNG, and PDF.

The composed pipeline has six stages. Each maps to a tier in the diagram and to the template or deployment pattern that builds it.

  1. Source systems. Application databases, event streams, SaaS exports, and files feed the pipeline. Optional Airbyte connectors replicate SaaS and database sources on a schedule; optional Redpanda accepts Kafka-compatible event streams for real-time ingestion paths.

  2. Orchestration. The Airflow template provisions a self-hosted orchestration host. You define DAGs that schedule extract tasks, invoke dbt runs, and track dependencies and retries.

  3. Raw layer. Extracted data lands as-is in S3-compatible Object Storage via the S3 Storage with ACLs template, the immutable record of what arrived.

  4. Transform jobs. The dbt transforms deployment pattern runs dbt Core in a container that Airflow starts on a schedule. dbt reads the raw layer, applies business logic in SQL, and writes staging and mart models to the warehouse.

  5. Curated layer and warehouse. Transforms write a curated layer back to Object Storage and load query-ready tables into a self-managed warehouse. The self-managed PostgreSQL template scaffolds Postgres on Block Storage. For column-oriented analytics at scale, swap in the ClickHouse template as the warehouse tier.

  6. BI tools. Metabase or Apache Superset connects to the warehouse and serves dashboards and ad hoc queries to analysts.

Services involved#

ServiceRole in this architectureDocs
ComputeAirflow host, warehouse host, BI host, and optional ingestion workers (CPU)Compute
Object StorageRaw and curated data-lake layersObject Storage
Block StoragePostgres or ClickHouse data volumesBlock Storage
NetworkPrivate networks and security groups for pipeline tiersNetwork

Get started#

Each template below pairs with a step-by-step deploy tutorial. Stand up the lake and warehouse first, then Airflow, then wire dbt and BI.

Estimate the cost#

Monthly cost estimate

Pricing calculator ↗

Sized as a custom package on a mix of shared and dedicated vCPU.

Starting template$171.20/mo

Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.

What each resource is for

Airflow host

s1a.medium · 4 shared vCPU, 4 GiB RAM, 0.5 Gbps

Runs Apache Airflow in Docker with LocalExecutor (webserver, scheduler, and bundled metadata PostgreSQL), with DAG data and logs on an attached volume.

Airflow LocalExecutor with bundled metadata PostgreSQL runs on 4 vCPU and 4 GiB RAM. Size up for heavier schedules or higher task concurrency.

$33.00/mo

PostgreSQL database

m2a.xlarge · 4 dedicated vCPU, 16 GiB RAM, 1 Gbps

$132.00/mo

Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.

Included in baseline

s1a.medium

4 shared vCPU, 4 GiB RAM, 0.5 Gbps

$33.00

m2a.xlarge

4 dedicated vCPU, 16 GiB RAM, 1 Gbps

$132.00

Compute + RAM rate basis

8 vCPU + 20 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.

—

Block storage (140 GiB)

140 GiB at $0.08/GiB/mo

$11.20

Public IP (included)

1 included with the custom package

$0.00

Package promotional discount

Flat −$5.00/mo on the custom package (same promotion as named plans).

$-5.00

Object storage (usage-based)

Object storage

1 bucket. The first 1 TB is included, then $10.00 per TB each month. You pay for what you store, so this line depends on usage.

$0–$40/mo

Assumes: 2 TB stored is $10/mo; 5 TB stored is $40/mo. Within the included allotment it stays $0.

Included at no charge

These line items are zero on Quake AI. Many other providers meter them separately.

Data transfer (inbound and outbound)

Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.

AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.

Learn more
$0.00

Object storage upload and download

No separate charges for uploading or downloading object storage data.

Most object storage providers meter egress and API requests separately from stored capacity.

Learn more
$0.00

Private networking

Private networks, subnets, Neutron routers, and security groups are included with the plan.

VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.

$0.00

Control-plane API requests

OpenStack API calls for provisioning and management are included.

Some managed services on other clouds meter API calls or charge for premium control-plane features.

$0.00

Configure your estimate

Check the add-ons you plan to deploy to build a monthly total. Nothing is selected to start, so the total below begins at the baseline.

Starting template

The required baseline, always included.

$171.20/mo

Pick how much you expect to store to fold it into the total.

$0.00/mo
Your configured estimate$171.20/mo

Dev/test vs production

Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.

Dev/test on shared CPU

Burstable s1a flavors; suited to prototyping and low or bursty load.

$72.20/mo

Production on the configured CPU

The headline estimate above; predictable steady-load performance.

$171.20/mo

Saves $99.00/mo while you build on shared CPU.

Shared flavors carry less RAM (m2a.xlarge (16 GiB RAM) -> s1a.medium (4 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.

Pricing data last validated: . For current rates, check quake.ai/pricing.

Considerations and limits#

  • Self-managed stack. Quake AI provisions compute, network, and storage through the templates above. Your team operates Airflow, dbt project code, the warehouse, and BI access under the shared responsibility model.
  • CPU-only compute. Compute is AMD EPYC with no GPU option (compute FAQ). Transforms and queries run on CPU; GPU-accelerated query engines and large-scale ML feature engineering need a different hosting path.
  • dbt runs as a job. dbt Core executes in a short-lived container that Airflow or cron starts. There is no long-running dbt service; see the dbt transforms deployment pattern for the wiring.
  • Flat egress. Quake AI applies a no-egress-fee policy for outbound transfer, which helps pipelines that move large data sets to and from the lake.
  • Three US regions. All current regions are in the United States.
  • Compliance posture. Quake AI holds SOC 2 Type I and Type II attestations and SOC 3. See Compliance and certifications for the platform scope.
Was this page helpful?