Search…

Fabric and Data Engineering

Everything you need to know about Airflow

Everything you need to know about Airflow

What Apache Airflow is, how DAGs work, a minimal DAG in Python, and how Airflow compares with Fivetran, AWS Step Functions and Microsoft Fabric pipelines.

What Apache Airflow is, how DAGs work, a minimal DAG in Python, and how Airflow compares with Fivetran, AWS Step Functions and Microsoft Fabric pipelines.

Written By: Austin Levine

Last Updated on September 22, 2026

Apache Airflow is an open-source tool for scheduling and orchestrating data workflows. You define each workflow in Python as a DAG, a set of tasks with dependencies between them, and Airflow runs the tasks in the right order, on a schedule, with retries and a web interface to see what ran and what failed. Airflow itself is free under the Apache 2.0 license. You pay for the servers it runs on, or for a managed service that runs it for you.

What Airflow does, and what it does not

Airflow is an orchestrator. It decides when each step runs and in what order, and it records the result. The actual work happens elsewhere: a SQL query in your warehouse, a Spark job, an API call, a dbt run.

That distinction explains most tool comparisons. Airflow does not copy data from Salesforce to your warehouse by itself. You either write that code, call a tool that does it, or use one of Airflow's provider packages that wrap a system's API.

How a DAG works

DAG stands for directed acyclic graph:

  • Directed: every dependency has a direction. Task B runs after task A.

  • Acyclic: there are no loops. A task can never depend, directly or indirectly, on itself.

Each workflow is one DAG. Each box in it is a task, usually created with an operator such as BashOperator, PythonOperator or a provider's operator for Snowflake, Databricks or Azure.

A minimal DAG that extracts data, transforms it and refreshes a report, once a day:

from datetime import datetime
from airflow import DAG
from airflow.operators.bash import BashOperator

with DAG(
    dag_id="daily_sales_pipeline",
    start_date=datetime(2026, 1, 1),
    schedule="@daily",
    catchup=False,
) as dag:
    extract = BashOperator(task_id="extract", bash_command="python extract_sales.py")
    transform = BashOperator(task_id="transform", bash_command="dbt run --select sales")
    refresh = BashOperator(task_id="refresh_report", bash_command="python refresh_power_bi.py")

    extract >> transform >> refresh

The last line sets the order. transform waits for extract, and refresh_report waits for transform. If transform fails, Airflow retries it according to your settings and does not run the refresh.

Import paths and parameter names differ between Airflow versions. Check the documentation for the version you run.

What Airflow is used for

Use

Example

ETL and ELT pipelines

Extract from an application database, run dbt models, refresh a Power BI semantic model

Machine learning

Retrain a model nightly, evaluate it, and publish it only if it beats the current one

Data quality

Run tests after each load and stop the pipeline when one fails

Operational jobs

Export files to a partner's SFTP server on a schedule

Airflow or Fivetran?

These tools do different jobs, so the question is usually which one to use for which part.

  • Fivetran is a managed data movement service. You pick a source such as Salesforce or a Postgres database, and Fivetran keeps a copy of it in your warehouse, handling schema changes and API limits. You write no code. Fivetran has a free plan covering 500,000 monthly active rows for connections and 5,000 model runs a month, and paid plans above that.

  • Airflow runs and orders tasks. It can call Fivetran, dbt, Spark or your own scripts, and decide what runs after what.

A common setup uses both: Fivetran loads raw data, dbt transforms it, and Airflow triggers each step in order and alerts you when one fails. Our dbt Cloud guide covers the transformation layer.

Choose Airflow alone when your sources have no ready-made connector, or when you need full control over the extraction code. Choose Fivetran when standard connectors exist and you would rather not maintain extraction code.

Airflow or AWS Step Functions?

AWS Step Functions is a serverless AWS service for orchestrating workflows. You define a workflow in Amazon States Language, a JSON format, or in a visual designer, and it calls AWS services such as Lambda, ECS and Batch.


Apache Airflow

AWS Step Functions

Workflow definition

Python code

JSON (Amazon States Language) or visual designer

Where it runs

Anywhere: your servers, Kubernetes, or a managed service

AWS only

Infrastructure

You run it, or pay for a managed service

Serverless, nothing to run

Best fit

Data pipelines across many systems and clouds

Workflows built mostly from AWS services

Where to run Airflow

  • Self-hosted: on your own servers or Kubernetes. Free software, but you maintain the scheduler, workers and metadata database.

  • Managed services: Amazon Managed Workflows for Apache Airflow (MWAA), Google Cloud Composer and Astronomer run Airflow for you.

  • Microsoft Fabric: Fabric's Data Factory includes Apache Airflow jobs, a managed Airflow that runs your Python DAGs next to your pipelines and lakehouses. Check the Airflow version and region list first: Microsoft documents one supported version per job, a subset of regions, and no support for private networks.

Airflow or Fabric pipelines?

If your data platform is Microsoft Fabric, start with Fabric's own pipelines. They are low-code, run on the capacity you already pay for, and can trigger notebooks, dataflows and semantic model refreshes. Our Fabric Data Factory vs Azure Data Factory comparison covers them.

Reach for Airflow when your team already writes Python DAGs, when workflows span systems outside Microsoft, or when you want orchestration defined in code and reviewed in Git. If you are still choosing the platform itself, read Why choose Microsoft Fabric.

Sources

Want Power BI expertise in-house?

Get in Touch With Us

Turn your team into Power BI pros and establish reliable, company-wide reporting.

Berlin, DE

powerbi@casewhen.co

Follow us on

© 2026 CaseWhen Consulting
© 2026 CaseWhen Consulting
© 2026 CaseWhen Consulting