# Deploy Airflow with the airflow template

Source: https://docs.quake.ai/resources/deployments/deploy-airflow-template
Markdown: https://docs.quake.ai/resources/deployments/deploy-airflow-template.md

---

# Deploy Airflow with the airflow template

Stand up [Apache Airflow](https://airflow.apache.org), an open-source workflow orchestration platform, on a single Quake AI instance using the [validated OpenTofu template](/docs/platform/validation#how-infrastructure-templates-are-checked) `airflow`. You apply the template, read the generated admin credentials, sign in to the web UI, trigger a sample DAG, and optionally serve the UI over HTTPS.

Airflow schedules and monitors data pipelines on infrastructure you own. You run it yourself; this is a self-hosted tool you operate, not a managed orchestration cloud.

<Figure size="md" caption="What you'll build: an Airflow host on a single instance with LocalExecutor, bundled metadata PostgreSQL, and a sample DAG">

```d2
direction: right

dev: You {shape: person}
domain: Your domain\n(DNS A record)
fip: Floating IP
instance: Ubuntu instance {
  caddy: Caddy\nreverse proxy
  web: Airflow\nwebserver
  sched: Airflow\nscheduler
  pg: Metadata\nPostgreSQL
  dags: DAG files
  sched -> pg: metadata
  web -> pg: metadata
  sched -> dags: runs tasks
  caddy -> web: proxies 443 to 8080
}

dev -> domain: HTTPS UI
domain -> fip
fip -> instance.caddy
```

</Figure>

<PricingCompanion
  components={[
    { kind: "template", slug: "airflow", required: true },
  ]}
/>

## Prerequisites

You need:

- OpenTofu 1.6.0 or later (or Terraform 1.6.0 or later) installed locally.
- Your OpenStack credentials sourced into the shell (`source openrc.sh`). See [the OpenStack CLI guide](/docs/tools/openstack-cli).
- An SSH keypair that already exists in your project. Record its name for the `key_name` variable.
- A copy of the `airflow` template directory from [the template reference page](/resources/iac-templates/airflow).
- Your workstation's public IP address, so you can open the web UI port to it for first-boot setup. Find it with `curl -sS https://api.ipify.org`.

A domain is optional for first boot. You add it in step 4 to serve the UI over HTTPS.

## Step 1: Set the variables and apply the template

The web UI listens on port 8080 over plain HTTP. The template's security group restricts port 8080 to `ui_allowed_cidr`, which defaults to the private network only, so the raw UI stays off the public internet. To reach the UI from your workstation for first-boot setup, set `ui_allowed_cidr` to your own address.

Copy the template's example variables file and open it:

```bash
cp terraform.tfvars.example terraform.tfvars
```

Set `key_name` to the SSH keypair already in your project, and `ui_allowed_cidr` to your workstation's public IP with a `/32` suffix:

```hcl
key_name         = "YOUR_KEY_NAME"
ui_allowed_cidr  = "YOUR_IP/32"
```



If you would rather not expose port 8080 at all, leave `ui_allowed_cidr` at its default and reach the UI over an SSH tunnel instead: `ssh -L 8080:localhost:8080 ubuntu@YOUR_FLOATING_IP`, then open `http://localhost:8080`. Once you add a domain in step 4, Caddy serves the UI over HTTPS on port 443 and you no longer need port 8080 open.



Initialize the working directory, preview the plan, and apply:

```bash
tofu init
tofu plan
tofu apply
```

OpenTofu provisions a private network, a router, a security group, a block volume mounted at `/var/lib/docker`, an instance, and a floating IP. On first boot, cloud-init mounts the data volume, installs Docker Engine, generates secrets, and starts Airflow (webserver, scheduler, and metadata PostgreSQL) from a compose file on port 8080.

When the apply finishes, read the outputs:

```bash
tofu output
```

Record `floating_ip` and `web_ui_url`.

## Step 2: Sign in to the web UI

Airflow does not ship a default password. cloud-init generates the admin password on first boot and writes it to `/opt/airflow/credentials.txt` on the instance.

cloud-init takes several minutes after the instance reaches `ACTIVE`. Watch the containers start over SSH:

```bash
ssh ubuntu@YOUR_FLOATING_IP "sudo docker compose -f /opt/airflow/docker-compose.yml ps"
```

When the webserver and scheduler containers report healthy, read the login details:

```bash
ssh ubuntu@YOUR_FLOATING_IP "sudo cat /opt/airflow/credentials.txt"
```

Open `web_ui_url` (for example `http://YOUR_FLOATING_IP:8080`) in your browser. Sign in with username `admin` and the password from the credentials file. Store the password in your password manager and remove or restrict access to `credentials.txt` on the instance when you are done.



Until you attach a domain in step 4, the UI is served over unencrypted HTTP on port 8080, reachable only from `ui_allowed_cidr`. Avoid sending production credentials over it from a shared or public network. Adding a domain (step 4) moves the UI to HTTPS on port 443.



## Step 3: Trigger the sample DAG

The template ships a sample DAG at `/opt/airflow/dags/hello_quake.py`. New DAGs start paused.

1. In the Airflow UI, open **DAGs** and find `hello_quake`.
2. Toggle the pause switch to unpause the DAG.
3. Open the DAG and select **Trigger DAG** (play icon).
4. Open the **Graph** or **Grid** view and confirm the `hello` task reaches **success**.

The task runs `echo "Hello from Quake AI Airflow"` on the scheduler container. Check logs from the task instance detail page or over SSH:

```bash
ssh ubuntu@YOUR_FLOATING_IP "sudo docker compose -f /opt/airflow/docker-compose.yml logs airflow-scheduler --tail 20"
```

To add your own pipelines, copy Python DAG files into `/opt/airflow/dags` on the instance (or sync them from Git in a later iteration). Airflow picks up new files within its scan interval.

## Step 4: Serve the UI over HTTPS with Caddy

The template leaves ports 80 and 443 open for a reverse proxy. [Caddy](https://caddyserver.com) obtains and renews a TLS certificate automatically once a domain resolves to the instance.

1. Create a DNS **A record** for your domain (for example `airflow.example.com`) pointing at `YOUR_FLOATING_IP`. Follow [How to point a domain at a Quake AI resource](/docs/network/how-to/point-domain-to-quake-ai). Wait until the record resolves:

```bash
dig +short airflow.example.com
```

The command returns your floating IP once the record propagates.

2. SSH to the instance and add a Caddy service that proxies HTTPS to Airflow on port 8080. Create `/opt/airflow/Caddyfile`:

```text
airflow.example.com {
  reverse_proxy 127.0.0.1:8080
}
```

3. Add Caddy to the compose file at `/opt/airflow/docker-compose.yml` so it runs alongside Airflow:

```yaml
services:
  caddy:
    image: caddy:2
    restart: unless-stopped
    network_mode: host
    volumes:
      - /opt/airflow/Caddyfile:/etc/caddy/Caddyfile
      - caddy_data:/data
volumes:
  caddy_data:
```

4. Tell Airflow its public base URL. Append to `/opt/airflow/.env`:

```text
AIRFLOW__WEBSERVER__BASE_URL=https://airflow.example.com
```

5. Apply the changes and confirm Caddy and the Airflow containers run:

```bash
cd /opt/airflow
sudo docker compose up -d
sudo docker compose ps
```

Open `https://airflow.example.com` and confirm the padlock. For background on certificate issuance and renewal, see [How to issue and auto-renew a TLS certificate with Let's Encrypt](/docs/network/how-to/lets-encrypt-certificate). Once HTTPS works, close direct access to port 8080 by setting `ui_allowed_cidr` back to the private network in `terraform.tfvars` and running `tofu apply`.

## Kubernetes executor scale path

This deployment runs LocalExecutor on one VM. When schedules or task concurrency outgrow that shape, move to the Kubernetes executor:

1. Provision a cluster with the [Kubernetes cluster template](/resources/iac-templates/k8s-cluster).
2. Deploy Airflow worker pods on the cluster while keeping the scheduler and webserver on a control node, or migrate the full stack into the cluster with the official Helm chart or a compose overlay you maintain.
3. Point Airflow at the cluster API and configure worker pod templates for your task images (for example a dbt container invoked from a `KubernetesPodOperator`).

The single-instance template in this walkthrough stays the entry pattern. The cluster template is the reuse point for scale-out execution without changing how you author DAGs.

## What you built

- **Applied the `airflow` template** to provision a network, security group, data volume, instance, and floating IP, and let cloud-init install Docker and start Airflow with LocalExecutor
- **Signed in to the web UI** with the admin credentials cloud-init generated on first boot
- **Triggered the sample DAG** and confirmed the task succeeded
- **Served the UI over HTTPS** by pointing a domain at the floating IP and routing it through a Caddy reverse proxy

## Scope of this deployment

This template runs a single-VM Airflow host with LocalExecutor, not a managed orchestration cloud. The instance is CPU-only and runs in one region. You operate the instance, Docker, Airflow, and the data volume yourself: back them up, patch them, and watch resource use as pipeline volume grows. Snapshot the data volume before you resize or rebuild the host. For higher concurrency, follow the Kubernetes executor path above and reuse the `k8s-cluster` template.

## Next steps

- [Apache Airflow template](/resources/iac-templates/airflow): the template reference, parameters, and resource map
- [self-managed PostgreSQL template](/resources/iac-templates/self-managed-postgres): a warehouse or metadata store to query from DAG tasks
- [Kubernetes cluster template](/resources/iac-templates/k8s-cluster): the cluster footprint for Kubernetes executor workers
- [How to store application secrets and inject them at runtime](/docs/security/how-to/inject-app-secrets): move connection strings and API keys out of plain environment variables
- [Security hardening checklist](/docs/security/hardening-checklist): tighten SSH access and exposure before you serve real traffic

## Clean up

When you no longer need the deployment, destroy everything the template created:

```bash
tofu destroy
```

Then remove the DNS A record you created in step 4. Because Airflow, its metadata database, and DAG logs all live on the instance and its attached volume, `tofu destroy` removes them along with the infrastructure. Export any DAG files you want to keep before you destroy.
