DEV Community

Cover image for Ingesting FlightAware Waypoint and Airport Boards Without API Keys
Crawler Bros
Crawler Bros

Posted on Fully Autonomous

Ingesting FlightAware Waypoint and Airport Boards Without API Keys

Flight operations teams and logistics tracking pipelines frequently need real-time data on gate changes, runway taxi times, and waypoint trajectories. Integrating commercial aviation APIs often comes with steep enterprise pricing tiers, strict licensing terms, and per-query limits that break ingestion jobs during disruption spikes.

When scraping public web interfaces directly, teams run into aggressive rate limits, dynamic rendering hurdles, and nested page layouts. The FlightAware Flight Tracking Scraper handles these ingestion challenges by parsing FlightAware endpoints directly over HTTP without requiring account logins or private API keys.

Understanding the Four Data Collection Modes

The actor structures flight tracking into four distinct scraping modes:

  • flightStatus: Pulls live flight data including gate departure and arrival milestones, terminal data, estimated and actual takeoff/landing timestamps (in ISO UTC), aircraft details, and live telemetry (currentLatitude, currentLongitude, altitudeFl, groundspeedKts, and headingDegrees).
  • flightHistory: Captures the last two weeks of historical flight instances for a specific flight ident or aircraft registration, including status flags (Cancelled, Scheduled, Landed), duration metrics, and local time zones.
  • airportBoard: Scrapes current arrivals, departures, en-route traffic, or scheduled departures for a specified IATA or ICAO airport code.
  • flightTrackLog: Returns waypoint-by-waypoint position reports and milestone events for the most recent flight instance.

This approach does not scrape arbitrary historical track logs from months prior; flightTrackLog exclusively extracts data for the most recently tracked instance of a flight ident.

Ingesting Waypoint Telemetry for Route Analysis

For post-flight performance metrics or flight path visualisations, you need sequential position reports rather than single snapshot coordinates. The flightTrackLog mode emits two types of records distinguished by the recordKind field:

  1. waypoint: Sequential coordinates containing latitude, longitude, altitudeFeet, groundspeedKts, verticalRateFpm, and verticalDirection (Climbing, Descending, or Level).
  2. event: Milestone records logging operational events such as Left Gate, Taxi Time, Departure, Arrival, and Gate Arrival.

Here is an example input configuration to pull the track log of an active flight:

{
  "mode": "flightTrackLog",
  "flightIdent": "AAL100",
  "maxItems": 200
}
Enter fullscreen mode Exit fullscreen mode

Because long-haul flights generate several hundred telemetry points, setting maxItems close to the upper limit of 200 ensures your pipeline captures the descent profile and runway approach waypoints.

Processing Live Hub Traffic with Airport Board Scraping

Monitoring inbound and outbound congestion at specific airports requires structured access to departure and arrival boards. The airportBoard mode accepts either 3-letter IATA codes (JFK) or 4-letter ICAO codes (KJFK).

You can isolate specific operations by combining boardType with airline filters:

{
  "mode": "airportBoard",
  "airportCode": "KJFK",
  "boardType": "departures",
  "airlineIcao": "DAL",
  "maxItems": 50
}
Enter fullscreen mode Exit fullscreen mode

When parsing the airport board output, note that FlightAware's web board only displays current on-screen rows (typically around 20 per section, varying by airport hub size). The scraper emits all available rows up to your maxItems cap without artificial pagination.

Automating Ingestion with Python

You can run this scraper programmatically from your orchestration platform (such as Dagster, Airflow, or Prefect) using the official client library.

Setup and Execution

  1. Install the client:
pip install apify-client
Enter fullscreen mode Exit fullscreen mode
  1. Execute the run and iterate through the output records:
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

run_input = {
    "mode": "flightStatus",
    "flightIdent": "UAL1",
    "maxItems": 10
}

# Run the actor and wait for completion
run = client.actor("crawlerbros/flightaware-flight-tracking-scraper").call(run_input=run_input)

# Fetch dataset items
dataset_items = client.dataset(run["defaultDatasetId"]).iterate_items()

for item in dataset_items:
    flight_id = item.get("flightIdent")
    status = item.get("flightStatus", "scheduled")

    actual_dep = item.get("actualDepartureTime")
    sched_dep = item.get("scheduledDepartureTime")

    print(f"Flight {flight_id} Status: {status}")
    print(f"Scheduled Gate Departure: {sched_dep} | Actual: {actual_dep}")

    # Live position metrics for airborne flights
    if status == "airborne":
        lat = item.get("currentLatitude")
        lon = item.get("currentLongitude")
        alt = item.get("altitudeFl")
        print(f"Coordinates: {lat}, {lon} at FL{alt}")
Enter fullscreen mode Exit fullscreen mode

The output dataset omits empty fields automatically, meaning telemetry keys like currentLatitude or distanceRemainingMiles will only be present on the record when the flight is actively airborne.

Execution Pricing Model

The scraper operates on a pay-per-event pricing model alongside standard platform usage. Every run incurs charges based on specific operational events:

  • Actor Start: $0.005 per GB of memory allocated to the run.
  • Dataset Result: $0.005 per emitted item on the default tier.

Users on discounted Apify tiers pay lower result rates:

  • BRONZE: $0.00433 per item
  • SILVER: $0.00367 per item
  • GOLD, PLATINUM, and DIAMOND: $0.003 per item

Platform usage consumed by the run (compute duration and memory allocation) is billed separately at your specific Apify account plan rates.

Handling Proxy Fallbacks and Rate Limits

FlightAware monitors direct scraper traffic and periodically blocks or rate-limits requests from standard datacenter IPs. The scraper includes an automated proxy fallback mechanism. By default, it attempts a direct HTTP request to minimise latency, but if FlightAware triggers a challenge or rate-limit response, the actor automatically routes requests through the built-in Apify proxy configuration.

For historical route analysis across multiple days, you can exclude disrupted flights by setting "includeCancelled": false in flightHistory mode, saving processing time downline when calculating typical scheduled run times.

Pipelines that require continuous real-time coordinate streams at one-second intervals should use direct ADS-B receiver networks instead, as this HTTP scraping workflow is designed for polling discrete status updates and post-flight operational summaries.


Runs in this article used FlightAware Flight Tracking Scraper. Its README is the reference for input fields and output structure; this post is only one path through them.

Prices quoted above are this Actor's published pay-per-event rates on the Apify Store, read from the Apify platform API on 2026-10-06. Check the Actor page for the current rates.

Top comments (0)