DEV Community

Sospeter Mong'are
Sospeter Mong'are

Posted on

Azure Data Factory vs Apache Airflow: Two Different Ways to Build Data Pipelines

When people compare Azure Data Factory (ADF) and Apache Airflow, the conversation often becomes a comparison of features.

Both can orchestrate data pipelines. Both can call APIs. Both can move data between systems. Both can schedule jobs, monitor executions, and integrate with data platforms.

But that can be misleading.

ADF and Airflow solve similar problems using fundamentally different approaches.

Azure Data Factory is primarily a visual, configuration-driven data integration and ETL/ELT service.

Apache Airflow is primarily a programmatic, code-driven workflow orchestration platform.

Understanding this distinction is more important than simply comparing their feature lists.

First, what is Azure Data Factory?

Azure Data Factory (ADF) is a cloud-based data integration service from Microsoft Azure.

Its main purpose is to help you build pipelines that move and transform data between different systems.

Instead of writing an entire pipeline as a Python program, you can construct it using visual activities.

For example, an ADF pipeline might look like:

Get Configuration
       |
       v
Get Access Token
       |
       v
Call API
       |
       v
Store JSON in Azure Data Lake
       |
       v
Transform / Copy Data
       |
       v
Load Azure SQL / Synapse
Enter fullscreen mode Exit fullscreen mode

Each box represents an activity configured through the ADF interface.

You can configure things such as:

  • URLs
  • HTTP methods
  • Headers
  • Parameters
  • Authentication
  • Source and destination datasets
  • Column mappings
  • Retry policies
  • Dependencies
  • Expressions
  • Conditional logic

ADF also integrates with other Azure services such as Azure Key Vault, Azure Data Lake Storage, Azure SQL, Azure Synapse and Azure Monitor.

The important idea is that you primarily describe what should happen through configuration and visual activities rather than implementing the entire workflow in application code.


What is Apache Airflow?

Apache Airflow is an open-source platform for programmatically authoring, scheduling and monitoring workflows.

Instead of building a pipeline primarily through a visual interface, you define it in code.

An Airflow workflow is usually represented as a DAG - Directed Acyclic Graph.

For example:

from airflow import DAG
from airflow.operators.python import PythonOperator

with DAG(
    dag_id="zoho_pipeline",
    schedule="@daily",
) as dag:

    extract_data = PythonOperator(
        task_id="extract_data",
        python_callable=extract_from_zoho
    )

    transform_data = PythonOperator(
        task_id="transform_data",
        python_callable=transform_data
    )

    load_data = PythonOperator(
        task_id="load_data",
        python_callable=load_to_database
    )

    extract_data >> transform_data >> load_data
Enter fullscreen mode Exit fullscreen mode

The code defines the workflow and its dependencies.

Conceptually:

extract_data
      |
      v
transform_data
      |
      v
load_data
Enter fullscreen mode Exit fullscreen mode

Airflow does not primarily exist to perform every data transformation itself. Its core strength is orchestration - coordinating tasks and systems.

Those tasks could involve Python, SQL, dbt, Spark, APIs, cloud services, containers or other external systems.


ADF vs Airflow: The Core Difference

The simplest way to understand the difference is:

ADF asks you to configure a data pipeline. Airflow asks you to program a workflow.

This distinction affects almost everything else.

Area Azure Data Factory Apache Airflow
Core paradigm Visual, configuration-driven Programmatic, code-driven
Primary interface Web UI Python code + UI for monitoring
Workflow definition Pipelines and activities Python DAGs
API integration Activities, REST connectors, datasets Python libraries, operators, hooks
Transformations Mapping Data Flows, Copy Activity, SQL, etc. Python, SQL, dbt, Spark and external tools
Version control Possible through Git integration Natural fit because DAGs are code
Infrastructure Microsoft-managed service You manage or use a managed Airflow service
Azure integration Deep native integration Strong, but usually through operators/providers
Multi-cloud Possible Strong fit
Custom Python logic Possible, but often indirect Native
Monitoring Azure portal Airflow UI
Scaling Managed by Azure services Depends on Airflow deployment
Maintenance Low infrastructure management Higher infrastructure responsibility

Neither approach is universally "better". They optimize for different engineering models.

Example: Connecting to the Zoho API

Suppose you need to extract data from Zoho.

The API requires:

  1. Obtaining an access token
  2. Calling an API endpoint
  3. Passing authentication headers
  4. Handling pagination
  5. Saving the response
  6. Transforming the JSON
  7. Loading the data into a warehouse

In ADF

You can construct this visually.

For example:

Web Activity
    |
    | Get OAuth token
    v
Set Variable
    |
    | Store token
    v
Web Activity / Copy Activity
    |
    | Call Zoho API
    v
Azure Data Lake
    |
    v
Copy Activity
    |
    v
Azure SQL / Synapse
Enter fullscreen mode Exit fullscreen mode

Credentials such as the Zoho client secret should be stored in a secure service such as Azure Key Vault rather than hardcoded in the pipeline.

ADF allows you to configure URLs, headers, parameters, expressions and activity dependencies through its interface.

In Airflow

The same workflow might be implemented in Python:

DAG
 |
 +-- Get Zoho Token
 |
 +-- Extract Zoho Data
 |
 +-- Transform JSON
 |
 +-- Load Warehouse
Enter fullscreen mode Exit fullscreen mode

The API interaction could use Python's requests library or an appropriate Airflow provider/operator.

For example:

response = requests.get(
    api_url,
    headers={
        "Authorization": f"Zoho-oauthtoken {access_token}"
    }
)
Enter fullscreen mode Exit fullscreen mode

The difference becomes more noticeable when the API becomes complicated.

If you need custom retry logic, unusual pagination, token refresh behavior or complicated response processing, Python gives you considerably more freedom.

What About JSON Transformations?

APIs rarely return perfectly flat data.

You might receive:

{
    "data": [
        {
            "id": "123",
            "name": "John",
            "activities": [
                {
                    "type": "call",
                    "date": "2026-09-30"
                }
            ]
        }
    ]
}
Enter fullscreen mode Exit fullscreen mode

You may need to flatten:

data
  -> customer
       -> activities
Enter fullscreen mode Exit fullscreen mode

ADF approach

ADF can handle many transformations through Copy Activity mappings and Mapping Data Flows.

You can configure mappings and expressions through the UI.

This works well when the transformation can be expressed using ADF's available activities and expressions.

Airflow approach

With Airflow, you can process the response using Python.

For example, libraries such as Pandas can help normalize nested JSON:

from pandas import json_normalize

df = json_normalize(response.json()["data"])
Enter fullscreen mode Exit fullscreen mode

You can then apply arbitrary Python logic before loading the result.

This illustrates a broader difference:

ADF gives you configurable transformation capabilities. Airflow gives you a programming environment from which you can call virtually any transformation tool.

Where Does the Processing Happen?

Another important distinction is execution architecture.

With ADF, data movement and transformation can be executed through Azure Integration Runtime or other configured compute resources.

You generally do not manage the underlying servers yourself.

Airflow, on the other hand, executes tasks through its workers or connected execution environments.

For example:

Airflow Scheduler
       |
       v
Airflow Worker
       |
       +-- Python
       +-- SQL
       +-- dbt
       +-- Spark
       +-- API
Enter fullscreen mode Exit fullscreen mode

This gives engineers considerable control, but it also introduces infrastructure considerations.

If your Python task loads a very large dataset into memory, for example, the worker's resources become relevant.

With Airflow, you therefore need to think about:

  • Worker sizing
  • Concurrency
  • Memory
  • CPU
  • Queues
  • Executors
  • Containers
  • Kubernetes
  • Task isolation

That operational responsibility is one of the biggest differences between the two platforms.

ADF: When Does It Make Sense?

ADF is particularly attractive when your environment is heavily invested in Azure.

For example:

API
 |
 v
ADF
 |
 v
Azure Data Lake
 |
 v
ADF / Transformation
 |
 v
Azure SQL / Synapse / Fabric
Enter fullscreen mode Exit fullscreen mode

The ecosystem fits together naturally.

ADF can also be a good choice when your team prefers visual development.

A team consisting primarily of SQL developers, analysts or database professionals may find it easier to understand:

Activity A
    |
    v
Activity B
    |
    v
Activity C
Enter fullscreen mode Exit fullscreen mode

than a Python-based orchestration framework.

ADF also removes much of the infrastructure burden.

You do not have to build and maintain an Airflow cluster just to execute your pipelines.

Airflow: When Does It Make Sense?

Airflow becomes particularly useful when your organization wants workflows to behave more like software.

Your pipeline is code.

Therefore you can naturally apply software engineering practices such as:

  • Git
  • Branching
  • Pull requests
  • Code reviews
  • Automated testing
  • CI/CD
  • Reusable Python functions
  • Package management
  • Code quality checks

For example:

Git Repository
      |
      v
Pull Request
      |
      v
Code Review
      |
      v
CI/CD
      |
      v
Airflow
      |
      v
Production DAGs
Enter fullscreen mode Exit fullscreen mode

Airflow is also useful in environments that are not exclusively Azure.

You could have:

Airflow
  |
  +-- Azure
  +-- AWS
  +-- GCP
  +-- Snowflake
  +-- Databricks
  +-- PostgreSQL
  +-- APIs
Enter fullscreen mode Exit fullscreen mode

This makes it attractive for organizations with heterogeneous or multi-cloud data platforms.

What About Complex Pipelines?

Imagine a Zoho API that requires:

  • Multiple authentication steps
  • Token refresh
  • Cursor-based pagination
  • Custom retry logic
  • Rate-limit handling
  • Nested JSON processing
  • Data validation
  • Custom error handling

In ADF, these requirements can sometimes be implemented using activities, expressions, variables, loops and conditions.

But the pipeline can become increasingly configuration-heavy.

In Airflow, you can implement the logic directly in Python.

That might look conceptually like:

def extract_zoho_data():
    token = get_token()

    while True:
        response = fetch_page(token)

        process(response)

        if not has_next_page(response):
            break
Enter fullscreen mode Exit fullscreen mode

This is one of the strongest arguments for a code-based orchestrator.

When workflow logic becomes software, treating it as software can make the implementation easier to reason about.

But Airflow Is Not Automatically Better for Everything

There is an important misconception to avoid.

Using Python does not automatically make a pipeline better.

A poorly designed Airflow deployment can introduce significant complexity.

You now have to think about:

Airflow Scheduler
Airflow Workers
Metadata Database
Executor
DAG Deployment
Secrets
Logs
Monitoring
Scaling
Upgrades
Networking
Enter fullscreen mode Exit fullscreen mode

With ADF, much of this infrastructure is abstracted away.

Therefore, the question should not simply be:

"Can Airflow do this?"

It almost certainly can.

The better question is:

"Does the additional flexibility justify the additional engineering and operational responsibility?"

A Practical Decision Framework

Think about the pipeline and the team building it.

Choose ADF when:

  • Your data platform is primarily Azure
  • Your destination is Azure SQL, Synapse, Fabric or Azure Data Lake
  • You want a managed service
  • Visual pipeline development is valuable
  • Most transformations can be handled through existing ADF capabilities
  • Your team is comfortable with configuration and SQL
  • You want minimal infrastructure management

Consider Airflow when:

  • Your workflows require substantial custom Python logic
  • You operate across multiple clouds or platforms
  • Your data engineers are comfortable with Python
  • Git-based workflow development is important
  • You need complex orchestration
  • You need to integrate many different processing technologies
  • You want workflows managed using software engineering practices

The Most Important Distinction

ADF and Airflow can both orchestrate a pipeline such as:

API
 |
 v
Extract
 |
 v
Transform
 |
 v
Load
 |
 v
Data Warehouse
Enter fullscreen mode Exit fullscreen mode

But they approach that pipeline differently.

ADF is closer to:

Configure the pipeline
        |
        v
Microsoft manages the platform
Enter fullscreen mode Exit fullscreen mode

Airflow is closer to:

Write the workflow
        |
        v
Deploy and operate the orchestrator
Enter fullscreen mode Exit fullscreen mode

That difference becomes increasingly important as your data platform grows.

If your primary goal is Azure-native data integration with minimal infrastructure management, ADF provides a natural model.

If your primary goal is flexible, code-first workflow orchestration across different systems, Airflow provides a different model.

Neither tool eliminates the need for good data engineering.

You still need to think about:

  • Authentication
  • Secrets management
  • Idempotency
  • Retries
  • Rate limits
  • Data quality
  • Observability
  • Schema changes
  • Error handling
  • Incremental loading
  • Backfills
  • Dependency management

The orchestration tool is only one part of the architecture.

The real decision is not simply ADF vs Airflow. It is configuration-driven orchestration vs code-driven orchestration, and which model fits your team's skills, infrastructure and operational requirements.

Top comments (0)