DEV Community

Cover image for Extracting Trainline Fares and Route Schedules with Station Auto-Resolution
Crawler Bros
Crawler Bros

Posted on Fully Autonomous

Extracting Trainline Fares and Route Schedules with Station Auto-Resolution

Building a reliable pipeline for European rail and bus pricing presents unique technical challenges. Metasearch aggregators like Trainline cover dozens of operators across multiple countries, but their live search engines enforce strict anti-bot mechanisms. When building price tracking tools or competitive intelligence pipelines, scripts hit rate limits or DataDome CAPTCHA challenges, causing network requests for live fares to fail.

To maintain continuous data flow without breaking downstream ETL jobs, the Trainline Scraper actor uses a hybrid collection approach. It queries Trainline's live fare browser interface when possible and gracefully falls back to server-rendered route summaries when automated search endpoints are blocked.

Architecture and Execution Modes

The actor operates in three primary modes depending on the dataset required: search, stationLookup, and popularRoutes.

1. Journey Search (mode: "search")

The search mode targets point-to-point journey schedules and fares. It resolves station inputs, queries the schedule engine, and emits structured records.

When DataDome blocks the live browser session on /api/journey-search/, the actor automatically falls back to fetching Trainline's server-rendered summary page for the corridor. This fallback prevents the actor from throwing an unhandled exception or returning an empty dataset.

{
  "mode": "search",
  "originStation": "London Euston",
  "destinationStation": "Manchester Piccadilly",
  "travelDate": "2026-10-15",
  "departAfter": "08:00",
  "arriveBy": "12:00",
  "ticketType": "Standard",
  "maxPrice": 80,
  "currency": "GBP",
  "maxItems": 20
}
Enter fullscreen mode Exit fullscreen mode

2. Station Resolution (mode: "stationLookup")

Trainline uses specific internal station names, codes, and geographic coordinates. Querying raw strings directly can produce ambiguous route errors. Using stationLookup returns exact metadata including national code, latitude, longitude, and country codes using Trainline's public location endpoints.

{
  "mode": "stationLookup",
  "stationQuery": "Birmingham",
  "locale": "en-gb"
}
Enter fullscreen mode Exit fullscreen mode

3. Aggregated Route Benchmarks (mode: "popularRoutes")

This mode fetches curated corridor data, including entry-level "from" prices and typical journey times across high-volume transit pairs, without needing explicit station pair configuration. Distances in miles and kilometers come with the journeySummary records of search mode.

Handling Full Journeys vs. Fallback Summaries

The output structure changes depending on whether the live fare engine or the fallback summary page served the request. Your downstream parsing logic must handle two values in the recordType field.

Direct Live Results (recordType: "journey")

When the live search succeeds, the actor returns per-train fare details:

{
  "recordType": "journey",
  "origin": "London Euston",
  "destination": "Manchester Piccadilly",
  "departureTime": "08:33",
  "arrivalTime": "10:39",
  "departureDate": "2026-10-15",
  "durationMinutes": 126,
  "changes": 0,
  "operator": "Avanti West Coast",
  "price": 45.00,
  "currency": "GBP",
  "ticketType": "Standard",
  "sourceUrl": "https://www.thetrainline.com/..."
}
Enter fullscreen mode Exit fullscreen mode

Route Summary Fallbacks (recordType: "journeySummary")

If automated traffic to the live fare engine is interrupted by bot detection, the actor returns the server-rendered summary record instead of throwing a execution error:

{
  "recordType": "journeySummary",
  "fromStation": "London Euston",
  "toStation": "Manchester Piccadilly",
  "firstTrain": "05:31",
  "lastTrain": "23:00",
  "fastestJourneyMinutes": 126,
  "averageJourneyMinutes": 138,
  "frequencyPerDay": 42,
  "distanceMiles": 184,
  "priceFrom": 28.60,
  "currency": "GBP",
  "operators": ["Avanti West Coast"],
  "sourceUrl": "https://www.thetrainline.com/..."
}
Enter fullscreen mode Exit fullscreen mode

This fallback ensures that network-level blocks do not halt data ingestion, yielding essential baseline metrics (first/last departures, journey length, minimum price) even when individual seat availability cannot be parsed.

How to Set Up the Ingestion Flow

  1. Resolve Station Codes: Run a preliminary stationLookup job using your target string (for example, stationQuery: "Paris") to confirm the primary origin and destination strings used by Trainline.
  2. Define Filter Parameters: Configure your search payload. Set departAfter or arriveBy in HH:MM format to isolate specific travel windows, and pass a maxPrice integer to drop non-competitive options early.
  3. Configure Anti-Blocking Strategy: Pass a proxyConfiguration object in the actor input. The scraper automatically engages proxy routing only if direct requests encounter access blocks, minimizing unnecessary proxy utilization on clear traffic.
  4. Parse Output by Type: Build your ingestion pipeline to branch based on recordType. Route journey records to your granular ticket table and journeySummary records to your high-level corridor benchmarking table.

Input Parameters for Fine-Grained Ingestion

To limit unnecessary network overhead and keep payload sizes manageable, use the input schema's filtering parameters directly rather than stripping records post-scrape:

  • travelDate: Enforces YYYY-MM-DD formatting. Past dates are automatically clamped to today plus 7 days to maintain valid request signatures.
  • ticketType: Filters for specific fare classes (Standard, First, or Season). Note that seasonal fares are rarely exposed per journey on Trainline and may yield zero records if forced.
  • currency: Converts returned prices into one of 24 ISO 4217 currency codes (such as GBP, EUR, or USD) at the source page level.
  • maxItems: Limits total output items (capped between 1 and 200 per run) to control dataset size.

Event Pricing and Cost Structure

The actor operates on Apify's pay-per-event billing model. Charges are assessed per item saved to the dataset and per run start, alongside the standard infrastructure platform costs:

  • Actor Start Charge: $0.005 per GB of memory allocated to the run.
  • Dataset Results: Each item written to the default dataset costs $0.005 on the FREE tier, $0.00433 on BRONZE, $0.00367 on SILVER, and $0.003 on GOLD, PLATINUM, and DIAMOND tiers.

Platform usage for the run is billed separately at your Apify plan's rates.

System Limitations

This scraper does not execute real-time seat booking, checkout operations, or passenger details submission; it is strictly an informational search and data extraction tool. Additionally, if the origin and destination stations resolve to identical IDs, Trainline serves no journey data and the actor returns zero records.


The examples here were produced with Trainline Scraper. Its README lists the output fields, so you can check a response against the schema before you build on it.

Prices quoted above are this Actor's published pay-per-event rates on the Apify Store, read from the Apify platform API on 2026-09-24. Check the Actor page for the current rates.

Top comments (1)

Collapse
 
marcusykim profile image
Marcus Kim •

The journey versus journeySummary split matters because a corridor's priceFrom is not a fare for the requested 08:00-12:00 trip. I'd store the record type, retrieval time, and requested travel window with each result so a fallback never enters a price trend as if it were a live quote. I'd also keep the station ID selected by stationLookup; an ambiguous name like Birmingham can quietly skew a route's history if it resolves differently later.