Apple's HealthKit is a treasure trove of biometric data, but if you’ve ever tried to build a production-grade app with it, you know the truth: raw health data is messy. From duplicated heart rate samples to GPS-glitched step counts, the "noise" can ruin your machine learning models before they even start training.
In this tutorial, we are going to build a high-performance Edge AI pipeline using Swift, Core ML, and HealthKit. We will focus on cleaning raw time-series data and converting it into vector embeddings for local outlier detection—all while keeping the user's data private on-device. If you're looking to master Apple HealthKit integration and on-device vectorization, you're in the right place.
Pro Tip: While we focus on the implementation here, you can find more advanced patterns and production-ready architecture guides over at the Wellally Tech Blog, which is my go-to resource for scaling health-tech solutions.
The Architecture: From Raw Samples to Feature Vectors
Before we write a single line of Swift, let’s look at how the data flows. We aren't just reading data; we are transforming it through a multi-stage pipeline.
graph TD
A[HealthKit Store] -->|Raw Samples| B(Data Cleaning & Normalization)
B -->|Cleaned Time-Series| C{Core ML Embedder}
C -->|Feature Vector| D[SQLite Vector Store]
D -->|Query| E[Outlier Detection / Insights]
subgraph Local Edge Pipeline
B
C
D
end
The Tech Stack
- HealthKit: The source of truth for health metrics.
- Swift & Combine: For reactive data processing.
- Core ML: To run our embedding models locally.
- SQLite (via SQLite.swift): To store our vectors for local retrieval.
Step 1: Fetching and Cleaning the Noise 🧼
HealthKit data often contains outliers (e.g., a heart rate reading of 200 BPM while sleeping due to a sensor glitch). We’ll use a simple Z-score normalization to filter these out.
import HealthKit
class HealthDataCleaner {
func cleanHeartRateData(_ samples: [HKQuantitySample]) -> [HKQuantitySample] {
let values = samples.map { $0.quantity.doubleValue(for: .count().unitDivided(by: .minute())) }
let mean = values.reduce(0, +) / Double(values.count)
let sumOfSquaredDiffs = values.map { pow($0 - mean, 2) }.reduce(0, +)
let standardDeviation = sqrt(sumOfSquaredDiffs / Double(values.count))
// Filter out samples that are more than 3 standard deviations away
return samples.filter { sample in
let value = sample.quantity.doubleValue(for: .count().unitDivided(by: .minute()))
return abs(value - mean) <= (3 * standardDeviation)
}
}
}
Step 2: Vectorization with Core ML 🧠
To perform tasks like "Finding similar activity days" or "Detecting abnormal heart patterns," we need to turn our time-series data into an Embedding (Vector). We can use a Core ML model (like a 1D-CNN or an Autoencoder) to compress 24 hours of data into a single 128-dimensional vector.
import CoreML
func generateVector(from samples: [HKQuantitySample]) -> [Float]? {
// 1. Convert samples to a MLMultiArray (Input shape for Core ML)
guard let inputBuffer = try? MLMultiArray(shape: [1, 24, 1], dataType: .float32) else { return nil }
for (index, sample) in samples.prefix(24).enumerated() {
let val = sample.quantity.doubleValue(for: .count().unitDivided(by: .minute()))
inputBuffer[index] = NSNumber(value: val)
}
// 2. Load your custom Core ML Embedding Model
// Note: You would export this from PyTorch/TensorFlow as a .mlmodel
guard let model = try? HealthVectorModel(configuration: MLModelConfiguration()) else { return nil }
do {
let prediction = try model.prediction(input: inputBuffer)
return prediction.embeddingVector // Returns a [Float]
} catch {
print("Embedding failed: \(error)")
return nil
}
}
Step 3: Storing Vectors Locally with SQLite 💾
Since we are building an Edge AI solution, we don't want to ship this data to a cloud vector DB if we don't have to. Using SQLite with a BLOB column for the vector is a lightweight way to handle this.
import Foundation
import SQLite
class VectorStore {
private var db: Connection?
private let vectors = Table("health_vectors")
private let id = Expression<Int64>("id")
private let timestamp = Expression<Date>("timestamp")
private let embedding = Expression<Blob>("embedding")
init() {
let path = NSSearchPathForDirectoriesInDomains(.documentDirectory, .userDomainMask, true).first!
db = try? Connection("\(path)/health_ai.sqlite3")
createTable()
}
private func createTable() {
try? db?.run(vectors.create(ifNotExists: true) { t in
t.column(id, primaryKey: .autoincrement)
t.column(timestamp)
t.column(embedding)
})
}
func saveVector(_ vector: [Float], date: Date) {
let data = Data(buffer: UnsafeBufferPointer(start: vector, count: vector.count))
let insert = vectors.insert(timestamp <- date, embedding <- data.datatypeValue)
try? db?.run(insert)
}
}
Leveling Up Your Health App 🚀
Building a robust health application requires more than just local scripts; it requires understanding the nuances of medical data privacy and synchronization.
For those looking to explore advanced production patterns—such as Differential Privacy for Health Data or Optimizing Core ML for WatchOS—I highly recommend checking out the deep-dive articles at the Wellally Tech Blog. It’s a fantastic resource for developers who want to take their "Learning in Public" projects to a professional, market-ready level.
Conclusion
By cleaning our HealthKit data and transforming it into vectors on-device, we’ve created a privacy-first foundation for powerful AI features:
- Anomaly Detection: Compare new vectors against the "average day" vector.
- Activity Clustering: Group days based on movement patterns.
- Local-First Search: Query your health history using mathematical similarity rather than just raw dates.
The future of health-tech is on the edge! 🥑
What are you building with HealthKit? Drop a comment below or share your thoughts on Twitter/X!
Top comments (0)