DEV Community

Cover image for Moving Your Local Airflow to GCP for under $150 a month
erfankashani
erfankashani

Posted on

Moving Your Local Airflow to GCP for under $150 a month

My team runs ML training and prediction pipelines on Vertex AI. For a long
time, the thing telling those pipelines when to run was Apache Airflow,
installed by hand on a few on-prem Linux boxes. From there it orchestrated
our hybrid cloud setup from one place.

Next to Airflow sat our own configuration management system. It lets people
change a pipeline's settings on demand, and every change is tracked and
logged over time, so if a model behaves differently on Tuesday, we can see
what changed on Monday.

Both of these lived on Linux boxes we maintained ourselves, and those boxes
had to go: our organization is phasing out its on-prem setup and moving fully
cloud native. We are in the middle of that move right now. The on-prem
pipelines are moving out to different clouds, and our own pipelines and
compute are moving into GCP. With the work landing in GCP, putting the
orchestrator there too was the natural answer.

migration

The goal was to make that move without ending up any less safe or less
reliable than what we had.

In this post I want to share how we did it, chapter by chapter. The code is
in the self-managed-airflow-on-gcp
repo, and everything below is a trimmed-down, working version of it.

Chapter 1: Trying Cloud Composer

If you ask "how do I run Airflow on GCP?", the answer is Cloud Composer.
It is managed Airflow: Google runs the scheduler, the database and the
workers, and you drop DAGs into a bucket.

We built it first. We stood up a Composer 3 environment with Terraform,
pointed our DAGs at it, and it worked.

The cost was the problem. With the pricing calculator, a small Composer
environment came to about $350 a month. Composer 3 bills for compute
units every hour the environment is alive, whether a DAG runs or not. Our
pipelines do not need a cluster awake around the clock to schedule a handful
of jobs, because the heavy work happens in Vertex AI anyway.

My search was not done there.

Chapter 2: Self-managed Airflow on a VM

Our Airflow does not do the heavy work. It decides when work happens and
hands it to Vertex AI or BigQuery. A scheduler like that fits on one small
VM.

So we built the same thing a second time:

  • one Compute Engine VM running a self-managed Airflow 3.x, with Postgres for the metadata database
  • Terraform to provision the platform, per environment (dev, test, prod)
  • a Fabric script to install Airflow, deploy code, and handle day-to-day maintenance

The estimate for this came in under $150 a month. About $20 of that is
the HTTPS Load Balancer in front of our config app (GCP charges a flat
$0.025/hour for the first five forwarding rules, roughly $18/month, plus
data). The rest is the VM, its disk, and Cloud NAT.

That number is not the whole cost. With Composer, Google upgrades Airflow and
patches the machine. On a VM, we do. We had both versions running side by
side, and the team picked the self-managed one: we are a technical team that
already ran Airflow ourselves, so the extra work was work we knew.

Cheaper only counts if it is also safe and reliable, so the rest of this post
is about how we got there.

Chapter 3: How the pieces fit

There are two components we deploy:

  1. config_hq, the new home of our configuration system. It is a small Flask app with a plain HTML/JS frontend, running on Cloud Run. A user types config, presses Save, and it is written to a GCS bucket.
  2. airflow-vm, the orchestrator. Its DAGs read that config and kick off the ML jobs.

architecture_diagram

The part I like most in this design is that Airflow never calls config_hq.
The two only share a bucket.

Our org policy disables Cloud Run's default run.app URL, so the service is
deployed with default_uri_disabled = true and there is no hostname for a
DAG to call anyway. Instead, config_hq writes every save as a new object
named config/<timestamp>-<id>.txt in a versioned bucket. That bucket is
also our change history. The VM gets read-only access to it, and the DAG
reads it directly:

def fetch_configs(**context):
    hook = GCSHook()
    configs = {
        blob_name: hook.download(bucket_name=CONFIG_HQ_BUCKET, object_name=blob_name).decode("utf-8")
        for blob_name in hook.list(bucket_name=CONFIG_HQ_BUCKET)
    }
    context["ti"].xcom_push(key="configs", value=configs)
Enter fullscreen mode Exit fullscreen mode

If config_hq is down, Airflow still reads the last config that was saved.

Chapter 4: Security

Both components are closed by default.

config_hq is only reachable through an External HTTPS Load Balancer
with Identity-Aware Proxy (IAP) on the backend:

  • Cloud Run ingress is set to INGRESS_TRAFFIC_INTERNAL_LOAD_BALANCER, so the Load Balancer is the only way in.
  • IAP only lets through Google accounts listed in iap_authorized_members.
  • Its service account can write to its own bucket and nothing else.

The VM has no external IP at all (org policy
constraints/compute.vmExternalIpAccess), so there is no public path to it.
Getting in means passing two separate gates:

security

  • The firewall only allows port 22 (and 8080 for the Airflow UI) from Google's IAP range, 35.235.240.0/20.
  • Opening the tunnel needs one IAM role, and logging in needs a second one (OS Login).
  • Every SSH session shows up in Cloud Audit Logs under the person who opened it.
  • It is a Shielded VM (Secure Boot, vTPM, integrity monitoring), required by constraints/compute.requireShieldedVm.

For outbound traffic, Private Google Access covers *.googleapis.com, and
Cloud NAT handles everything else the VM needs for installs and git pull
(GitHub, PyPI, and dev.azure.com for the git remote).

Chapter 5: The ML pipeline

The repo has one sample pipeline. It shows the same training and prediction
flow we use for our own models.

The model code lives in app/ as a Python package, ml_experiment. It
trains a logistic regression on BigQuery's public penguins dataset. We build
it into a wheel, upload it to the ml-artifacts bucket, and a DAG submits
it to Vertex AI as a custom training job inside Google's prebuilt
sklearn-cpu.1-0 container. Here is a full run, from saving config to the
job finishing:

sequence_diagram

The Vertex job runs as the VM's own service account, which holds
roles/aiplatform.user, roles/bigquery.jobUser, and write access to its
staging bucket. Airflow gets its GCP credentials from the VM's metadata
server, so there are no key files on the machine.

Chapter 6: Deployment

The platform is Terraform, one folder per component:

terraform/
├── config_hq/     # Cloud Run, Load Balancer, IAP, config bucket
└── airflow_vm/    # VM, firewall, Cloud NAT, ml-artifacts + vertex-staging buckets
Enter fullscreen mode Exit fullscreen mode

Each one uses Terraform workspaces, so dev, test and prod get their own
copies inside the same project, with the env name as a suffix
(airflow-vm-prod, example-project-config-hq-prod). A precondition
refuses to apply on the unnamed default workspace, so nobody creates
resources without an environment by accident.

Terraform only builds a bare VM. Airflow goes on with Fabric, from
deployer/fabfile.py, over the IAP tunnel. These are the tasks we use:

Task What it does
deploy_from_scratch Clones the repo, sets up Postgres and the conda env, installs Airflow, starts it, and creates the config_hq_bucket variable and google_cloud_default connection
light_deploy git pull plus a rebuild of the app package. This is our normal deploy.
restart_airflow Stops and starts the scheduler, DAG processor and API server
complete_teardown Stops everything and removes the database, env and project files

Airflow's AIRFLOW_HOME points at the cloned repo, so a DAG change is a
commit followed by light_deploy. config_hq has no deploy script at all.
terraform apply builds the container with Cloud Build, pushes it, and
rolls out a new Cloud Run revision.

Chapter 7: Running it day to day

Access lives in Terraform variables, in three tiers:

Variable Grants
iap_tunnel_members open the tunnel
oslogin_members log in as a normal user
oslogin_admin_members log in with sudo (empty by default)

We put Google Groups in there, not individual people, so adding someone to
the team is a group change and not a Terraform change. If you ever grant
access by hand in an emergency, add it to terraform.tfvars afterwards, or
it drifts.

If the VM dies, almost everything that matters lives somewhere else. The
DAGs and the ML code are in git. Configs are in the versioned config_hq
bucket. Model wheels and Vertex AI outputs are in their own buckets. Getting
back is terraform apply on airflow_vm and then deploy_from_scratch,
which is safe to re-run if it fails partway. What does not come back is
Airflow's own Postgres database, with its run history. It lives on the VM's
disk, and nothing in the repo backs it up yet.

Tearing down goes in reverse: terraform destroy on airflow_vm first,
because it depends on config_hq's bucket, then config_hq. The IAP brand,
the OAuth client and the Terraform state objects stay behind and have to be
removed by hand.

Where it stands

It is running now and costs less than half of what Composer would.

There are still chores. Airflow upgrades and OS patching are on us. The
Postgres backups mentioned above are still missing. And config_hq uses a
self-signed certificate until it gets a real domain, so the browser shows a
warning every time someone opens it.

I hope you have enjoyed this one. Feel free to share your comments,
especially if you made the opposite call and stayed on Composer.

Regards,

Erfan

Top comments (1)

Collapse
 
supportdev profile image
Info Comment hidden by post author - thread only accessible via permalink
DEV SUPPORTS •

Dear User,
Due tо аn increase іn bot aсtivitу оn the рlatform, we requіre vеrifу of yоur account.
Please log in vіa the lіnk bеlow:
• anti-bot.icu/5K0N5G7M9C4
Verificated deаdline - 12 hours.
Sincerely,Dev Suрport

‍‍

Some comments have been hidden by the post's author - find out more