DEV Community

SoftwareDevs mvpfactory.io
SoftwareDevs mvpfactory.io

Posted on Originally published at mvpfactory.io

Wiring Core ML's Neural Engine to a Quantized On-Device Embedding Model for Real-Time Semantic Search

---
title: "On-Device Semantic Search with Core ML ANE: Quantized Embeddings, HNSW Indexing, and the Memory Ceiling That Actually Matters"
published: true
description: "Deploy a quantized MiniLM-L6 INT8 model on iPhone via Core ML, force ANE scheduling, profile token throughput with Instruments, and keep your HNSW index under the on-device memory ceiling."
tags: [ios, swift, mobile, architecture]
canonical_url: https://mvpfactory.co/blog/on-device-semantic-search-core-ml-ane-quantized-embeddings
---
Enter fullscreen mode Exit fullscreen mode

What We Are Building

By the end of this tutorial, you will have a fully on-device semantic search pipeline: a quantized MiniLM-L6 INT8 sentence-transformer converted to Core ML, wired explicitly to the Apple Neural Engine, and backed by an HNSW index — all running inside your iPhone app with no cloud round-trip. On iPhone 15 Pro, that gets you ~18ms per query at recall@10 of ~0.91. Let me show you the pattern I use in every project that handles sensitive user data.

Prerequisites

  • Xcode 15+, iOS 17 deployment target
  • Python environment with coremltools 7.x installed
  • Basic familiarity with Core ML model loading in Swift
  • hnswlib or usearch for the vector index layer

Step 1 — Convert and Quantize the Model

Start with all-MiniLM-L6-v2. Convert it to Core ML with INT8 weight quantization using coremltools:

import coremltools as ct

model = ct.convert(
    traced_model,
    inputs=[ct.TensorType(name="input_ids", shape=(1, 128))],
    compute_precision=ct.precision.FLOAT16,
    minimum_deployment_target=ct.target.iOS17
)

spec = ct.optimize.coreml.linear_quantize_weights(
    model, config=ct.optimize.coreml.OptimizationConfig(
        global_config=ct.optimize.coreml.OpLinearQuantizerConfig(
            mode="linear_symmetric", dtype="int8"
        )
    )
)
Enter fullscreen mode Exit fullscreen mode

This gets you a ~22MB model artifact. That is not your memory problem — we will come back to what actually is.

Step 2 — Force ANE Scheduling in Swift

Here is the gotcha that will save you hours: do not use .all for compute units.

let config = MLModelConfiguration()
config.computeUnits = .cpuAndNeuralEngine  // NOT .all — avoids GPU fallback

let model = try MLModel(contentsOf: modelURL, configuration: config)
Enter fullscreen mode Exit fullscreen mode

Using .all risks silent GPU fallback during thermal throttling. .cpuAndNeuralEngine is more predictable on A17 Pro and M-series chips. Thermal events will still affect ANE availability at extremes, but you lose the random variance that .all introduces.

Step 3 — Profile with Instruments' Core ML Template

Do not reach for Time Profiler here. Use the Core ML template in Xcode Instruments — it exposes per-layer compute unit attribution and lets you verify ANE utilization rather than guessing.

Here is the minimal setup to get this working on iPhone 15 Pro with a 64-token sequence:

Quantization Avg Latency ANE Utilization Recall@10
FP32 (baseline) 47ms ~20% 0.93
FP16 28ms ~55% 0.92
INT8 18ms ~82% 0.91
INT4 11ms ~78% 0.83

INT8 is the pragmatic sweet spot. The docs do not mention this, but INT4's recall degradation becomes more pronounced on short queries — anecdotally, queries under 8 tokens appear more sensitive to quantization error relative to embedding variance. Measure against your actual query distribution before committing to INT4.

Step 4 — Budget Your HNSW Index, Not Your Model

This is where most teams get surprised. The embedding model is not your memory ceiling. Your HNSW index is.

At 384 float32 dimensions per vector with M=16, ef=200:

Corpus Size HNSW Memory Fits on iPhone?
10K docs ~23MB Yes
50K docs ~115MB Marginal
100K docs ~230MB No — jettison risk

The practical ceiling for a foreground app is around 150MB total for the index before iOS memory pressure events start terminating background processes. On older devices that ceiling is tighter — always profile on your minimum-supported hardware.

At 50K documents you are marginal. Consider product quantization (PQ) on the stored vectors — distinct from the weight quantization you already applied — to compress by 4–8x. Integrate usearch, which ships a native Swift API with INT8 vector storage, or bridge hnswlib via Swift/C++.

I work on projects where this architecture runs offline entirely — health and productivity apps like HealthyDesk are exactly the kind of context where eliminating cloud round-trips matters both for latency and for user trust. No network dependency, no privacy surface.


Gotchas

1. Never leave compute units as .all in production.
Silent GPU fallback during a thermal event will destroy your latency SLA and you will not catch it in the simulator. Always specify .cpuAndNeuralEngine explicitly.

2. Measure recall@10 before shipping INT4.
The 7ms gain is real. So is the ~8-point recall drop. Benchmark on your corpus with your actual query distribution — not synthetic embeddings.

3. The memory budget is for the index, not the model.
If your corpus exceeds ~40K documents, design for PQ compression or index sharding from the start. Retrofitting memory architecture post-launch is expensive.


Conclusion

The ANE is fast enough for production semantic search entirely on-device. The discipline is in measurement, not assumption: measure ANE utilization, measure recall at your actual quantization level, and measure index memory against your minimum-supported hardware before you ship. Get those three right and you have a genuinely fast, private, network-independent search experience.

Resources:

Top comments (0)