DEV Community

Cover image for Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer
AI Explore
AI Explore

Posted on

Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer

Every GPU data library benchmarks itself at 50 million rows. Your dataframe has 80,000.

I build one of these libraries, and I published the 50-million-row numbers too. So this time I measured the other thing: for 111 operations, the row count at which the GPU on an Apple M4 Max first stays ahead of the fastest CPU code I could find, and below which a single CPU core beats it. The answer is later than the benchmarks imply, sixteen operations never get there at all, and the fix I shipped is to stop using the GPU below the line.

Everything here is from one Apple M4 Max (16 CPU cores, 64 GB) and every number names its row count. The CPU side is the fastest of Polars (eager and lazy), pyarrow (eager and 16-batch Acero), pandas and numpy at each size, or, where I say so, a single-core loop. The generated page behind this post is docs/CROSSOVER.md in the ArrowMetal repository, an Apache-2.0 Arrow compute library for the Apple silicon GPU. Nothing on that page is typed by hand.

Why the GPU is slower than the CPU on small data: the 100-microsecond tax

Ask the GPU to sum 1,000 integers and it takes 112 microseconds. Polars does it in less than one. The kernel is done in nanoseconds; what you pay for is the command-buffer round trip: encode, commit, wait, read back. On Apple silicon that floor is roughly 100 µs per call and it does not care how many rows you sent.

operation, 1,000 rows ArrowMetal GPU fastest CPU which GPU behind by (at least)
sum(int64) 0.112 ms < 0.001 ms Polars > 100×
filter(int64 > 0) 0.154 ms 0.003 ms Polars > 50×
group-by sum, 1,000 keys 0.401 ms 0.080 ms pandas 5×

Unified memory does not help here. It removes the transfer across a bus, which is what makes medium-sized work pointless on a discrete card. It does not remove the dispatch floor. The way to remove it would be a persistent GPU kernel polling shared memory for work; on Metal that cannot currently be made to work, for coherence reasons the project measured and filed with Apple (FB24858235). Batching many kernels into one command buffer shares the floor across them; a single small call cannot escape it.

sum(int64): GPU kernel vs one CPU core, Apple M4 Max, log-log; they cross at 220,881 rows

Read the amber line. From 1,000 to 1,000,000 rows the GPU goes from 115 µs to 197. A thousand times more data, 70% more time. One CPU core goes from 1 µs to 993. They cross at 220,881 rows. Below that, using the GPU for a sum turns a 50-microsecond job into a 120-microsecond one.

GPU vs CPU: where the crossover actually is, for 111 operations

The harder question: from what row count does the GPU stay ahead of the best CPU library, using all the cores it wants?

Crossover by family: bar from earliest to latest operation, ring at the median

family operations ever win earliest median latest
reductions 18 18 1,000 10,000,000 50,000,000
group-by 28 27 100,000 1,000,000 10,000,000
strings 19 16 100,000 1,000,000 10,000,000
sort 12 10 1,000,000 1,000,000 10,000,000
element-wise 25 20 1,000,000 10,000,000 50,000,000
compare + select 9 4 1,000,000 30,000,000 50,000,000

Two caveats on the number 111. It is the set of operations the size sweep covers, not a selection: chains, decimal, join, temporal and window are measured only at 10M and 50M and have no crossover to state yet. And the median of the 111 operations is 10,000,000 rows against the fastest multi-core library: 40 operations cross by 1,000,000, 55 need 10,000,000 or more. Against a single CPU core the line is nearer 1,000,000.

Three things that surprised me:

  • Group-by is the early winner, not arithmetic. 27 of 28 cross over, half of them by 100,000 rows. A sum by 100,000 keys is 1.4x ahead at 100,000 rows and 8.4x at 10,000,000. A GPU is a very good hash table.
  • Element-wise arithmetic is the late one. Adding two int64 columns does not stay ahead of Polars until 10,000,000 rows. One instruction per element, and the CPU library runs it on 12 cores at memory bandwidth.
  • Compare and select mostly never wins. Only 4 of 9. A filter keeping 90% of rows is 0.87x at 50,000,000. Single passes at bandwidth on both sides; the dispatch floor never gets paid back.

The sixteen operations where the GPU never beats the CPU

operation GPU ÷ fastest CPU at the largest size why
ln, sin, sqrt, power (float64) 0.10x, 0.16x, 0.87x, 0.95x Metal has no double; float64 maths is software binary64
slice (zero-copy view) 0.33x the CPU library returns a view and moves nothing
unique, value_counts (1,000 distinct) 0.83x, 0.84x answered by a full GPU radix sort; the CPU hashes
real regex, parse, to_strings 0.08x, 0.17x, 0.60x host fallbacks, not GPU kernels yet
drop_null, filter 90% kept, replace_with_mask, case_when, compare scalar 0.91x, 0.87x, 0.36x, 0.94x, 0.79x single-pass, memory-bound; Polars lazy already at bandwidth on 10 to 14 cores
variance by key (1,000 groups) 0.96x two-pass exact variance vs a fused one-pass on the CPU

Every one has its cause in the project's to-improve list. I would not use the GPU for logarithms either.

Where the GPU wins by 10x to 100x

operation, 50,000,000 rows crossover at 1M at 10M at 50M
any / all over booleans 100,000 3.1x 22.5x 106x
quantile (median), float64 1,000,000 2.4x 13.2x 38.7x
partition_nth_indices 1,000,000 2.3x 17.5x 24.5x
take, 25M random indices 1,000,000 1.4x 11.3x 24.2x
lexsort, two int32 keys 1,000,000 1.8x 16.2x 24.0x
argsort float64 1,000,000 4.7x 10.0x 10.5x

Gathers, sorts, order statistics and anything that is secretly a sort. If your job is 50,000,000 rows and a median, the GPU is 38.7x ahead of Polars. If your job is 80,000 rows and a sum, the GPU is the wrong tool and the benchmark page was never going to tell you.

What I did about it: ArrowMetal 0.2.0 refuses the GPU below the line

ArrowMetal 0.2.0, released 25 September 2026, ships a CPU/GPU router. For seven operations there is a second implementation, a tight single-threaded CPU loop over the Arrow layout, and the router runs it below a measured crossover and the GPU kernel at or above it. The table is generated by a script from a timing run of the shipped loops against the shipped kernels; a test fails if the committed table stops matching the results file.

routed operation CPU loop wins below fitted between (GPU µs / CPU µs)
sum 220,881 rows 100K: 125.6 / 46 · 300K: 161.5 / 213.6
max 227,546 100K: 126.3 / 66.8 · 300K: 178.4 / 212.2
min 234,408 100K: 124.2 / 65.8 · 300K: 180.3 / 208.8
filter by mask 591,365 300K: 375 / 198.1 · 1M: 406.8 / 654.9
multiply 678,287 own row: a 64-bit multiply costs the CPU more than an add
group-by sum, ≤ 1,024 keys 716,029 300K: 269.9 / 186.1 · 1M: 558 / 615.2
add, subtract 1,629,041 1M: 181.8 / 120.5 · 3M: 230.1 / 363.7
compare 2,668,823 1M: 141.5 / 77 · 3M: 211.2 / 224

Three rules I held myself to:

  • Byte-identical output. Same values, same validity bitmap, same null count, same buffer sizes. Float sums do not add in row order: the CPU loop reproduces the GPU's summation order exactly, thread by thread, so the sum is the GPU's to the bit, NaN payloads included.
  • No CPU path, no choice. Float32 arithmetic computes in hardware float on the GPU, where subnormals are flushed, so a CPU loop could not match its bits; it stays on the GPU and the decision says why.
  • Measured, then checked. Re-running the check against the fitted table, auto took the faster path in 71 of 72 cases. The one miss is compare(int64 > int64) at 3,000,000 rows, by 1.10.

Strings are not routed: the string rows are behind a 12-core Acero and a single-threaded loop would not change that, so a small string operation still pays the floor today.

You can watch it decide. This ran on the M4 Max against the 0.2.0 wheel from PyPI; the comments are the actual output:

pip install -U arrowmetal # 0.2.0 · macOS 14+ on Apple silicon

import numpy as np, pyarrow as pa, arrowmetal as am
am.router_crossovers()
# {'sum': 220881, 'min': 234408, 'max': 227546, 'compare': 2668823,
# 'arithmetic': 1629041, 'filter': 591365, 'group_by_sum': 716029, 'multiply': 678287}

for n in (1_000, 100_000, 1_000_000, 10_000_000):
 col = am.array(pa.array(np.arange(n, dtype=np.int64)))
 col.sum()
 r = am.last_route()
 print(f"{n:>11,} rows -> {r.path} ({r.reason})")
# 1,000 rows -> cpu (below the 220881-row crossover)
# 100,000 rows -> cpu (below the 220881-row crossover)
# 1,000,000 rows -> gpu (at or above the 220881-row crossover)
# 10,000,000 rows -> gpu (at or above the 220881-row crossover)

with am.router("gpu"): # or "cpu"; ARROWMETAL_ROUTER=gpu|cpu|auto for the process
 col.sum()
Enter fullscreen mode Exit fullscreen mode

The rest of 0.2.0 is the bigger half

0.1.0 was a compute library: 307 Arrow kernels that were very fast once your columns were already on the GPU. 0.2.0 is the version where the data arrives.

0.1.0, 8 September 0.2.0, 25 September
307 GPU kernels, always the GPU CPU/GPU router with a measured crossover table, byte-identical
Parquet: flat columns Nested Parquet on the GPU: structs, maps, lists to any depth; stored Arrow schema applied; page-index and bloom-filter skipping
No text readers GPU readers for CSV and newline-delimited JSON, checked against pyarrow
Files only Delta Lake and Apache Iceberg read through the GPU Parquet reader
Polars: convert a Series lf.collect(engine=am.MetalEngine()): subtrees of the optimised Polars plan run on the GPU, the rest stays with Polars
DuckDB: a Python bridge A DuckDB optimizer extension moving eligible aggregates of unchanged SQL onto the GPU
IPC: the common types Every Arrow type nested, LZ4/ZSTD, big-endian, view types, fixed-shape tensor
String sort: nulls partitioned on the CPU Nulls in the GPU sort key: 10M strings with 10% nulls, 122.7 ms → 42.0 ms

Parquet. A real Parquet file is structs of structs of strings, maps with nested values, lists inside lists. 0.2.0 reassembles all of it on the GPU: three levels per field, one flag kernel, one prefix sum, one scatter, checked value for value against pyarrow.parquet.read_table on files written by pyarrow, DuckDB and Polars. With a column index in the file, pages whose min/max cannot match are never read, decompressed or decoded; split-block bloom filters drop row groups for an equality filter before any page is touched; last_read_stats tells you how many pages were skipped. One bug I am glad was found first: a list column with more than 4,096 level entries in one read used to write past a 4-byte scratch buffer into host memory.

CSV, JSON, Delta Lake, Iceberg. CSV and NDJSON parse on the GPU. Delta and Iceberg tables read through the GPU Parquet reader, which I believe is the first time either format has been opened on an Apple GPU. Plainly: on day one both lakehouse readers are behind the fastest CPU reader, as are the IPC view layouts and some nested Parquet reads. The measurements are in Benchmarks/results/ and the lakehouse page says which.

Polars. lf.collect(engine=am.MetalEngine()) walks the optimised plan, translates the subtrees it can into ArrowMetal plans, replaces each with a GPU function returning a Polars DataFrame, and leaves everything else to Polars. engine.last_report says which nodes ran on Metal and why the rest did not. Building it is where four wrong-answer bugs came from, because its differential suite runs Polars queries against the GPU plans. The best one: a string filter or take meeting a null slot that still held bytes (valid Arrow; Polars exports them) copied those bytes over the next kept row and turned "banana" into "xanana". Fixed, tested, and the findings log lists the other three.

DuckDB. An optimizer extension moves eligible aggregates of ordinary SQL onto the GPU with the query unchanged; SET arrowmetal_rewrite = 'off' or 'force' overrides it, and it now refuses an unrecognised value instead of silently treating it as 'auto'.

Strings. The utf8 sort used to partition its null rows on the CPU, waiting for the GPU batch that produced the keys. 0.2.0 carries the null placement in bit 63 of every prefix key, so the radix passes leave nulls as one block and the wait is gone: 10,000,000 strings with 10% nulls, 122.7 ms to 42.0 ms. It also closed a real bug where the partition could read the index array before the GPU had written it.

The rule I would actually give you

Do not ask whether the GPU is faster. Ask at how many rows, for this operation, on this machine, and whether your data is already in Arrow memory. On an Apple M4 Max: group-by and strings from about 100,000 rows; sorts, gathers and order statistics from about 1,000,000; element-wise arithmetic from about 10,000,000; never for float64 transcendentals, views, or anything a multi-core CPU already runs at bandwidth. Below those lines one core is the right tool, and as of 0.2.0 the library picks it for you.

ArrowMetal is an independent Apache-2.0 project that implements Apache Arrow; it is version 0.2.0 and every number above is from one machine. pip install -U arrowmetal on macOS 14+. If you run Benchmarks/router_check.py on a different Apple silicon Mac, the crossovers you get are new information and an issue on GitHub would be welcome.

Top comments (0)