DEV Community

Beck_Moulton
Beck_Moulton

Posted on

From XML Chaos to Quantified Self: Building a Lightning-Fast Apple Health Pipeline with Polars

Have you ever tried to open your Apple Health export.xml file only to have your text editor freeze and your laptop fans kick into overdrive? You aren't alone. As someone obsessed with Data Engineering and the "Quantified Self" movement, I quickly realized that Apple's export format—a massive, deeply nested XML file—is a nightmare for standard processing tools.

If you are tired of waiting minutes for Pandas to load your data, it's time to upgrade. In this tutorial, we will build a high-performance data pipeline using Polars, the Rust-backed DataFrame library, to transform gigabytes of health records into sleek, queryable Parquet files in seconds. By focusing on performance optimization and memory efficiency, we’ll turn that "dirty" data into a goldmine of personal insights. 🚀


The Architecture: From Raw XML to Insights

Processing large XML files requires a "Stream-then-Transform" approach. Instead of loading the entire 1GB+ file into memory (which will crash most 8GB RAM machines), we stream the XML elements and then leverage Polars' LazyFrame for heavy lifting.

graph TD
    A[export.xml] -->|ElementTree Iterparse| B(Streaming Parser)
    B -->|Dictionary List| C[Polars DataFrame]
    C -->|Lazy API| D{Data Cleaning}
    D -->|Filter Types| E[Heart Rate]
    D -->|Filter Types| F[Step Count]
    D -->|Filter Types| G[Sleep Analysis]
    E & F & G -->|Column Compression| H[Compressed Parquet Files]
    H -->|Analysis| I[Dashboard / Quantified Self]
Enter fullscreen mode Exit fullscreen mode

Prerequisites

To follow along, make sure you have the following tech stack installed:

  • Python 3.9+
  • Polars: The star of the show.
  • ElementTree: For memory-efficient XML streaming (built-in).
  • PyArrow: Required by Polars for Parquet export.
pip install polars pyarrow
Enter fullscreen mode Exit fullscreen mode

Step 1: Memory-Efficient XML Parsing

The secret to handling GB-scale XML is xml.etree.ElementTree.iterparse. This allows us to iterate through the file one "Record" at a time without loading the whole tree.

import xml.etree.ElementTree as ET
import polars as pl
from datetime import datetime

def parse_apple_health_xml(file_path):
    data = []
    # We use 'end' events to ensure the element is fully parsed
    context = ET.iterparse(file_path, events=("end",))

    for event, elem in context:
        if elem.tag == "Record":
            # Extract attributes from the XML tag
            record_type = elem.get("type")
            # Minimal extraction to keep memory low
            record = {
                "type": record_type.replace("HKQuantityTypeIdentifier", ""),
                "value": elem.get("value"),
                "unit": elem.get("unit"),
                "startDate": elem.get("startDate"),
            }
            data.append(record)

            # CRITICAL: Clear the element to free memory
            elem.clear()

    return pl.from_dicts(data)

# Usage
# df = parse_apple_health_xml("export.xml")
Enter fullscreen mode Exit fullscreen mode

Step 2: The Polars "Lazy" Magic

Now that we have a DataFrame, we need to clean it. Apple Health data is notoriously messy: timestamps are strings, values are mixed types, and "types" have long, redundant prefixes.

Instead of executing operations one by one, we use Polars LazyFrames. This allows Polars to optimize the query plan before executing.

def clean_health_data(df: pl.DataFrame):
    return (
        df.lazy()
        # 1. Convert timestamps to proper datetime objects
        .with_columns([
            pl.col("startDate").str.to_datetime("%Y-%m-%d %H:%M:%S %z"),
            pl.col("value").cast(pl.Float64, strict=False)
        ])
        # 2. Filter out nulls or irrelevant data
        .filter(pl.col("value").is_not_null())
        # 3. Rename or simplify column names
        .select([
            pl.col("type"),
            pl.col("startDate").alias("timestamp"),
            pl.col("value"),
            pl.col("unit")
        ])
        .collect() # Final execution happens here!
    )
Enter fullscreen mode Exit fullscreen mode

🥑 Advanced Patterns & Production Ready Pipelines

While the above logic works great for personal scripts, building a production-grade data engineering pipeline involves handling schema evolution and incremental loads.

Pro Tip: If you are looking for more production-ready examples, including how to deploy these pipelines to the cloud or set up automated "Quantified Self" dashboards, I highly recommend checking out the deep-dive articles at WellAlly Tech Blog. They cover advanced architectural patterns that take these concepts to the next level.


Step 3: Saving to Parquet for 100x Faster Reads

CSV is great for humans, but Parquet is built for machines. It uses columnar storage and heavy compression. A 1GB XML file can often be shrunk to a ~50MB Parquet file while remaining instantly searchable.

def export_to_warehouse(df: pl.DataFrame):
    # Partitioning data by type makes future queries instant
    unique_types = df["type"].unique().to_list()

    for t in unique_types:
        subset = df.filter(pl.col("type") == t)
        # Save as a compressed parquet file
        subset.write_parquet(f"health_data_{t}.parquet", compression="snappy")
        print(f"✅ Exported {t} with {len(subset)} rows.")

# Putting it all together
raw_data = parse_apple_health_xml("export.xml")
clean_data = clean_health_data(raw_data)
export_to_warehouse(clean_data)
Enter fullscreen mode Exit fullscreen mode

Conclusion

By ditching the standard "Load Everything" approach and embracing Polars and XML Streaming, we’ve turned a frustratingly slow task into a high-performance data pipeline. ⚡️

What we achieved:

  1. Memory Safety: Processed GBs of data with minimal RAM usage.
  2. Blazing Speed: Polars' Rust engine outperformed Pandas by a mile.
  3. Future-Proof Storage: Used Parquet to ensure our data is ready for BI tools or ML models.

Are you tracking your health data? What’s the weirdest insight you’ve found in your Apple Health export? Let me know in the comments! 👇


Happy coding! If you enjoyed this, don't forget to follow for more Data Engineering deep dives! 🥑💻

Top comments (0)