Skip to content

Run dbt transforms as a scheduled job

Deployment · Updated Jul 2026

Run dbt transforms as a scheduled job

dbt Core runs SQL transforms inside your warehouse. Each run builds staging and mart models, runs tests, and writes documentation artifacts. dbt is a batch job: a container starts, executes dbt run, and exits. There is no long-running dbt service to keep online.

This deployment pattern composes two validated OpenTofu templates: Airflow schedules the work, and self-managed PostgreSQL holds the warehouse dbt writes to. You bring the dbt project, the container image, and the schedule. Quake AI supplies the compute and storage those templates provision.

ScheduleQuake AIAirflow scheduler(or cron)DAG run / cronfires on intervalAirflow instance(airflow template)dbt container(dbt run, then exit)PostgreSQL warehouse(self-managed-postgres template) start jobdocker runSQL transforms
Click to zoom
dbt runs as a short-lived container job. Airflow (or cron on a host with warehouse access) starts the container; dbt connects to Postgres, runs models, and exits.

Monthly cost estimate

Pricing calculator ↗

Sized as a custom package on a mix of shared and dedicated vCPU.

Starting template$171.20/mo

Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.

What each resource is for

Airflow host

s1a.medium · 4 shared vCPU, 4 GiB RAM, 0.5 Gbps

Runs Apache Airflow in Docker with LocalExecutor (webserver, scheduler, and bundled metadata PostgreSQL), with DAG data and logs on an attached volume.

Airflow LocalExecutor with bundled metadata PostgreSQL runs on 4 vCPU and 4 GiB RAM. Size up for heavier schedules or higher task concurrency.

$33.00/mo

PostgreSQL database

m2a.xlarge · 4 dedicated vCPU, 16 GiB RAM, 1 Gbps

$132.00/mo

Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.

Included in baseline

s1a.medium

4 shared vCPU, 4 GiB RAM, 0.5 Gbps

$33.00

m2a.xlarge

4 dedicated vCPU, 16 GiB RAM, 1 Gbps

$132.00

Compute + RAM rate basis

8 vCPU + 20 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.

—

Block storage (140 GiB)

140 GiB at $0.08/GiB/mo

$11.20

Public IP (included)

1 included with the custom package

$0.00

Package promotional discount

Flat −$5.00/mo on the custom package (same promotion as named plans).

$-5.00

Included at no charge

These line items are zero on Quake AI. Many other providers meter them separately.

Data transfer (inbound and outbound)

Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.

AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.

Learn more
$0.00

Private networking

Private networks, subnets, Neutron routers, and security groups are included with the plan.

VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.

$0.00

Control-plane API requests

OpenStack API calls for provisioning and management are included.

Some managed services on other clouds meter API calls or charge for premium control-plane features.

$0.00

Dev/test vs production

Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.

Dev/test on shared CPU

Burstable s1a flavors; suited to prototyping and low or bursty load.

$72.20/mo

Production on the configured CPU

The headline estimate above; predictable steady-load performance.

$171.20/mo

Saves $99.00/mo while you build on shared CPU.

Shared flavors carry less RAM (m2a.xlarge (16 GiB RAM) -> s1a.medium (4 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.

Pricing data last validated: . For current rates, check quake.ai/pricing.

Prerequisites#

Before you wire dbt, confirm you have:

  • Airflow running from the Deploy Airflow with the airflow template walkthrough. The scheduler and web UI must be reachable from the instance where dbt runs (typically the same Airflow host).
  • A Postgres warehouse from Deploy the self-managed PostgreSQL template with OpenTofu. Note the private IP, database name, and application role. dbt connects over the network; it does not co-locate with Postgres by default.
  • Network path between the Airflow instance and the Postgres private IP. Both templates can share a private subnet, or you can place a bastion or VPN between them. See How to set up SSH bastion access into a private subnet if your laptop needs to reach private addresses during setup.
  • A dbt project checked into git (models, dbt_project.yml, and a profiles.yml that reads connection settings from environment variables).
  • Docker available on the Airflow host (the airflow template installs Docker Compose for the Airflow stack; use the same engine for dbt runs).

If you prefer ClickHouse as the warehouse, apply the ClickHouse template instead of Postgres and point dbt at ClickHouse with the dbt-clickhouse adapter. The scheduling shapes below stay the same; only the profile target and adapter package change.

Step 1: Stand up the composed templates#

Apply each template in its own working directory if you have not already:

  1. Follow Deploy the self-managed PostgreSQL template with OpenTofu and record the warehouse private IP, database name, and role name from your tfvars file.
  2. Follow Deploy Airflow with the airflow template on a instance that can reach the warehouse security group. Add the Airflow host private IP to the Postgres template allowed_cidrs before apply if the database security group does not already permit it.

Store database credentials in Airflow Connections or in a restricted env file on the host. Do not commit passwords to your dbt repo or to DAG source in git.

Step 2: Add a dbt profile that reads from the environment#

On your workstation, create profiles.yml beside your dbt project. The profile below expects DBT_HOST, DBT_USER, DBT_PASSWORD, DBT_DATABASE, and DBT_SCHEMA at run time:

YAML
warehouse:
  target: prod
  outputs:
    prod:
      type: postgres
      host: "{{ env_var('DBT_HOST') }}"
      user: "{{ env_var('DBT_USER') }}"
      password: "{{ env_var('DBT_PASSWORD') }}"
      port: 5432
      dbname: "{{ env_var('DBT_DATABASE') }}"
      schema: "{{ env_var('DBT_SCHEMA') }}"
      threads: 4

Commit profiles.yml when every secret comes from env_var(). Keep production values in Airflow Variables, Airflow Connections, or a root-owned env file on the scheduler host.

Step 3: Invoke dbt from Airflow#

Mount your dbt project into the official dbt Postgres image and pass the env vars from an Airflow Connection. Example DAG using DockerOperator:

Python
from datetime import datetime

from airflow import DAG
from airflow.providers.docker.operators.docker import DockerOperator

with DAG(
    dag_id="dbt_daily_models",
    start_date=datetime(2026, 1, 1),
    schedule="@daily",
    catchup=False,
) as dag:
    DockerOperator(
        task_id="dbt_run",
        image="ghcr.io/dbt-labs/dbt-postgres:1.9.latest",
        api_version="auto",
        auto_remove="force",
        command="dbt run --profiles-dir /usr/app/profiles",
        mounts=[
            {"source": "/opt/dbt/my_project", "target": "/usr/app", "type": "bind"},
            {"source": "/opt/dbt/profiles", "target": "/usr/app/profiles", "type": "bind"},
        ],
        environment={
            "DBT_HOST": "{{ conn.warehouse_postgres.host }}",
            "DBT_USER": "{{ conn.warehouse_postgres.login }}",
            "DBT_PASSWORD": "{{ conn.warehouse_postgres.password }}",
            "DBT_DATABASE": "{{ conn.warehouse_postgres.schema }}",
            "DBT_SCHEMA": "analytics",
        },
    )

Copy your dbt project to /opt/dbt/my_project on the Airflow host (git pull, rsync, or a CI deploy step). Create an Airflow Connection named warehouse_postgres in the Airflow UI under Admin > Connections with the warehouse host, login, password, and database.

Unpause the DAG in the Airflow UI and trigger a run. Check task logs for dbt run output and confirm new relations appear in Postgres.

Step 4: Alternative: cron on a host with warehouse access#

When you do not need Airflow's dependency graph or UI, run the same container from cron on any host that can reach the warehouse private IP (the Airflow instance, a bastion, or a dedicated runner VM).

Create /etc/dbt/dbt.env on the host with root-only permissions:

bash
DBT_HOST=192.168.40.10
DBT_USER=app_role
DBT_PASSWORD=REPLACE_AT_DEPLOY_TIME
DBT_DATABASE=analytics
DBT_SCHEMA=analytics

Replace REPLACE_AT_DEPLOY_TIME with the value from your secret store before the first run. Do not check this file into git.

Add a cron entry:

15 2 * * * root docker run --rm --env-file /etc/dbt/dbt.env \
  -v /opt/dbt/my_project:/usr/app \
  -v /opt/dbt/profiles:/usr/app/profiles \
  ghcr.io/dbt-labs/dbt-postgres:1.9.latest \
  dbt run --profiles-dir /usr/app/profiles >> /var/log/dbt-run.log 2>&1

The container starts at 02:15, runs models, writes logs to /var/log/dbt-run.log, and exits.

Step 5: Verify transforms in the warehouse#

SSH to the Postgres host or connect through your bastion and list relations dbt created:

bash
psql -h localhost -U app_role -d analytics -c "\dt analytics.*"

Run a spot check on a mart model:

SQL
SELECT COUNT(*) FROM analytics.fct_orders;

If counts look wrong, re-run with dbt run --select model_name and inspect dbt logs on the Airflow task or in /var/log/dbt-run.log.

Clean up#

Tear down in reverse order when you finish testing:

  1. Pause or delete the Airflow DAG (or remove the cron entry).
  2. Run tofu destroy in the Airflow template working directory.
  3. Run tofu destroy in the self-managed-postgres template working directory.

Next steps#

Was this page helpful?