DEV Community

Anup Karanjkar
Anup Karanjkar

Posted on Originally published at wowhow.cloud

Polars 2.0 Upgrade Guide: 8 Breaking Changes That Hit Silently

Polars 2.0 changes the default engine behind LazyFrame.collect() from in-memory to streaming, and eight changes in the official upgrade guide can alter your pipeline output without raising a single error. The guide lives at docs.pola.rs/releases/upgrade/2 and states no release date. As of 8 October 2026 sources disagree on whether 2.0.0 final has shipped or only a release candidate, so check the GitHub releases page before you pin a version.

Short answer: pin the in-memory engine first, run your tests, then remove the pin one pipeline at a time. The loud breaks, such as removed methods and renamed arguments, fail fast and are easy to fix. The dangerous ones are quiet. Row order that used to look stable stops being stable, a horizontal concat that used to pad now raises, and a selector expression that used to mean a set intersection now means something else. This guide orders the changes by how likely they are to corrupt a result unnoticed.

Change 1: streaming is the default engine

With engine="auto", LazyFrame.collect() and collect_async() now pick the streaming engine. Eager DataFrame operations are unchanged, so code written in the eager style does not move. Lazy pipelines move to streaming, which uses less memory on large data and produces different execution behaviour.

Three switches restore the old behaviour. Use the one that fits your deployment:

# Polars 1.x default: in-memory
result = lf.collect()

# Polars 2.0: ask for the old engine per call
result = lf.collect(engine="in-memory")

# or globally in code
pl.Config.set_engine_affinity("in-memory")

# or from the environment, no code change
POLARS_ENGINE_AFFINITY=in-memory python run_pipeline.py
Enter fullscreen mode Exit fullscreen mode

The environment variable is the safest first step. It changes no source file, so you can run the whole test suite against 2.0 before touching code. One related rename: the variable POLARS_STREAMING_CHUNK_SIZE is now POLARS_IDEAL_MORSEL_SIZE, so update any deployment config that sets it.

Change 2: row order is no longer guaranteed

This is the change most likely to ship a wrong report. The streaming engine does not guarantee row order for unpivot, group_by and joins. If a downstream step took the first row of a group, wrote a CSV that a human diffed, or zipped two frames together by position, the output can differ run to run.

# Before: order looked stable
out = lf.group_by("customer").agg(pl.col("amount").sum()).collect()

# After: make the order explicit
out = (
    lf.group_by("customer", maintain_order=True)
    .agg(pl.col("amount").sum())
    .collect()
)
# or sort at the end
out = lf.group_by("customer").agg(pl.col("amount").sum()).sort("customer").collect()

# joins: preserve the left frame's order
joined = left.join(right, on="id", how="left", maintain_order="left")
Enter fullscreen mode Exit fullscreen mode

Sorting at the end is cheaper to reason about than preserving order inside every operation. Add one sort at the output boundary and be done. If your tests compare frames with assert_frame_equal, keep check_row_order on and let the failures show you where order mattered.

Changes 3 to 5: file readers, concat heights and casts

Change 3: read_csv is now scan_csv().collect(). The eager reader goes through the lazy one. It gains with_column_names, infer_schema_files (default 10), credential_provider, include_file_paths, extra_columns and missing_columns. It loses n_threads, batch_size, sample_size and rechunk. The dtypes argument is now schema_overrides, and a list passed to schema_overrides must cover every column. read_ipc follows the same pattern, as scan_ipc().collect(), and drops memory_map and rechunk.

# Before
df = pl.read_csv("orders.csv", dtypes={"id": pl.Int64}, n_threads=4)

# After
df = pl.read_csv("orders.csv", schema_overrides={"id": pl.Int64})
Enter fullscreen mode Exit fullscreen mode

If your CSV step is an ad hoc conversion from JSON exports, the free JSON to CSV Converter produces a clean file to test the new reader against before you point it at production data.

Change 4: horizontal concat requires equal heights. pl.concat(frames, how="horizontal") used to pad the shorter frames with nulls. In 2.0 it raises when heights differ. The padding behaviour still exists under a new name, how="horizontal_extend". This is a loud break, but only if you test with uneven frames. In a pipeline that always had matching heights it passes silently, and then fails in production the day an upstream filter returns fewer rows.

Change 5: casts and supertypes tightened. Casting integers to Categorical and Categorical to integers is disallowed. Casting String to Date, Datetime or Time is disallowed too, and the replacement is the explicit parser:

# Before
df.with_columns(pl.col("d").cast(pl.Date))

# After
df.with_columns(pl.col("d").str.to_date("%Y-%m-%d"))
Enter fullscreen mode Exit fullscreen mode

Mixing a signed integer with UInt64 now resolves to Int128 instead of Float64. That is better for correctness, but it changes the output dtype of any expression that mixes the two, and a bitwise AND between a boolean and an integer now raises. Parquet ENUM columns are read as String instead of Binary, and a new pl.Map dtype represents maps in Arrow, Parquet and Iceberg. Unknown Arrow extension types load as pl.Extension.

Changes 6 and 7: selectors and SQL

Change 6: selector operators changed meaning. The operators &, | and ^ applied between a selector and pl.col are now element-wise expressions, not set operations. A line like cs.numeric() & pl.col("flag") used to intersect two column sets. In 2.0 it combines values row by row. Nothing raises, because both readings are valid expressions. To keep set semantics, make both operands selectors, for example by using selector functions on each side. Search your code for every use of those three operators next to a selector.

Change 7: SQL resolves late. SQLContext.execute(eager=True) and pl.sql(..., eager=True) now run through collect(), which means streaming, so the row-order caveat applies to SQL queries. Resolution is also deferred. pl.sql(), execute() and LazyFrame.sql() resolve table and column names when you call collect() or collect_schema(), not when you build the query. A typo in a column name that used to fail at the call site now fails later, at collect time. The traceback points somewhere other than the line you wrote.

Change 8: the removals and renames table

These are loud. Python raises on the first run, so a smoke test finds them.

Removed or changed Replacement

| LazyFrame.profile() | Query Profiler in Polars Cloud |

| LazyFrame.fetch() | collect() then head() |

| melt() | unpivot(index, on) |

| with_row_count() | with_row_index() |

| group_by().count() | group_by().len() |

| join(join_nulls=) | join(nulls_equal=) |

| join(how="outer") | join(how="full") |

| pivot(columns=) | pivot(on=) |

| read_csv(dtypes=) | read_csv(schema_overrides=) |

| str.concat() | str.join(), default delimiter now empty, not "-" |

| Expr.where() | Expr.filter() |

| cut() and qcut() | Deprecated for bin_intervals(), bin_quantiles(), bin_ranks() |

| with_context() | concat(how="horizontal") |

| Dataframe Interchange Protocol | Removed, __dataframe__ raises |

The str.concat row hides a silent break inside a loud one. If you called str.concat() with no argument and relied on the "-" separator, the rename to str.join forces you to touch the line, but the new default delimiter is an empty string. Code that migrated by renaming alone produces joined strings with no separator. Pass the delimiter explicitly.

Smaller behaviour changes in the same release: explode() has empty_as_null set to False by default, shift(None) raises, show_graph defaults plan_stage to "physical", hash values and the single seed changed, and the ordering parameter on Categorical is gone. Anything that persists hashes, such as a partition key written to disk, needs a rebuild.

Loud breaks, quiet breaks and a migration order

Which changes are loud and which are quiet

Sort the work by failure mode before you estimate effort. A loud change stops the program. A quiet change returns a different answer. Your risk lives entirely in the quiet column.

Change Failure mode How you find it

| Streaming default | Quiet: different execution, same code | Run tests pinned to in-memory, then unpinned |

| Row order | Quiet: output order varies | Order-sensitive assertions, a final sort |

| read_csv and read_ipc | Loud for removed arguments, quiet for inference | Grep for dtypes, n_threads, memory_map |

| Horizontal concat | Loud, but only on uneven input | Test with a short frame |

| Casts and supertypes | Loud for casts, quiet for Int128 | Schema snapshot tests |

| Selector operators | Quiet: valid expression, new meaning | Grep for selector next to pl.col |

| SQL deferral | Loud, late and far from the cause | Call collect_schema() early |

| Removals | Loud, first run | Smoke test |

Three rows are quiet, and those are the three that deserve review time. The schema snapshot test is the cheapest defence for the dtype changes. Save df.schema for each pipeline output to a file, commit it, and fail the test when it changes. A signed integer column that quietly becomes Int128 shows up as a one-line diff instead of a downstream surprise.

Persisted hashes need their own line in the plan. The guide states that the hash seed handling and hash values changed. If any pipeline writes hash output to storage and later joins against it, for example a bucket key or a deduplication fingerprint, old and new values will not match. Recompute the stored values in the same release as the upgrade, or keep the old version running until you finish.

Finally, call collect_schema() early on SQL-heavy code. Because pl.sql(), execute() and LazyFrame.sql() resolve on collect() or collect_schema(), a single early schema call turns a late, confusing failure into an immediate one with the query still in view.

A migration order that works

Treat the upgrade as a branch with its own review, not a dependency bump in a Friday afternoon pull request. Upgrade one pipeline at a time, starting with the ones that have the best tests and the least downstream consumers, and record the engine setting you used for each so a rollback is a one-line change. Run the sequence below on that branch. It finds the loud breaks first and isolates the quiet ones.

grep -rnE "melt|with_row_count|join_nulls|how=.outer|dtypes=|str.concat|fetch|profile" src/
POLARS_ENGINE_AFFINITY=in-memory pytest
pytest   # now on streaming: failures here are row-order or SQL-timing issues
Enter fullscreen mode Exit fullscreen mode

Fix the grep hits, run the suite pinned to in-memory, then unpin and read the failures. If the second run fails and the first passes, the cause is streaming, and change 2 or 7 explains it nearly every time. Map which modules call collect() with the Codebase Graph Visualizer, so you migrate leaf pipelines before the shared ones.

Write the engine and ordering rules into your repository's agent instructions so coding assistants stop generating melt and with_row_count. The CLAUDE.md Production Rules pack has the format, and the Obsidian Developer Productivity Vault gives you a migration-log template to track which pipelines moved. Rules as files, not prompts, is the same approach described in agent skills beat agent crews.

Quick answers

Does Polars 2.0 break eager code?

Mostly no. The engine default applies to LazyFrame.collect() and collect_async(), and the guide states eager DataFrame operations are unchanged. Removed methods and renamed arguments still apply to eager code.

How do I keep the old in-memory engine?

Call collect(engine="in-memory"), set pl.Config.set_engine_affinity("in-memory"), or export POLARS_ENGINE_AFFINITY=in-memory. The environment variable needs no code change.

Why did my group_by output change order?

The streaming engine does not guarantee row order for group_by, joins and unpivot. Sort explicitly, or pass maintain_order to the operation.

Is Polars 2.0 released?

The upgrade guide exists and sources disagree on whether the final release shipped as of 8 October 2026. Check the GitHub releases page and pin an exact version.

Every product mentioned is available at wowhow.cloud — pay once, ship forever.

Originally published at wowhow.cloud

Top comments (0)