Polars 2.0 Is Out. Its Upgrade Guide Repeats One Warning 12 Times: Your Pipeline Results May Change Silently

On Tuesday, October 6, Ritchie Vink published "Release of Polars 2.0" on the Polars blog, and version 2.0.0 landed on PyPI the same day. When I opened the Hacker News thread on Wednesday, it had 410 points and 95 comments.
On September 2, announcing the first release candidate, Vink wrote that the team hoped 2.0 would be "a boring experience for you." The major version was there to change defaults, not to ship features.
Then I read the migration guide. It puts 16 changes in a "Danger" box, and 12 of those boxes carry the same sentence: "This change may silently impact the results of your pipelines."
I write Python pipelines that pull records, score them and push the results somewhere: content queues, lead lists, relevance filters. The dataframe library is rarely what slows them down. Its defaults are what bite. So here is 2.0 read from that seat.
The change that touches every lazy query
Calling collect() on a LazyFrame now runs on the streaming engine instead of the in-memory one. Eager DataFrame operations are unaffected, and sink_* calls were already streaming.
The reason it needed a major version is row order. The streaming engine does not guarantee it for operations that do not require it: joins, group_by, unpivot. The guide says joins are "easy to miss since nothing about the query looks order-sensitive." A left join that used to return rows in the order of the left table may not anymore. If anything downstream takes head(10), writes a CSV a human reads top to bottom, or diffs against yesterday's file, it just changed.
The fixes, copied from the guide:
left.join(right, on="k", how="left", maintain_order="left").collect()
pl.Config.set_engine_affinity("in-memory") # process-wide
lf.collect(engine="in-memory") # per query
The environment variable POLARS_ENGINE_AFFINITY=in-memory does the same. SQL follows the rule too: pl.sql(..., eager=True) now runs on the streaming engine.
The quiet ones
Most 2.0 changes raise an error, which is the easy case. These do not, and I would look for them first in scoring code.
explode() on an empty list now produces zero rows instead of one null row. Row counts change wherever a list column contains empty lists.
Adding a signed integer column to a UInt64 column used to give a lossy Float64. It now gives an exact Int128. More correct, but a different dtype flows into whatever comes next.
Combining a selector with pl.col(...) through &, | or ^ no longer means a column selection. For two integer columns, it silently becomes a bitwise operation.
pl.datetime(...) used to name its output "datetime". It now takes the name of its leftmost argument, which can overwrite an existing column in a with_columns call.
A schema passed to read_csv or scan_csv is now matched by column name instead of position. If your schema listed columns in a different order than the file, 1.x mislabelled them, as the guide's example shows, and 2.0 fixes it. Your output changes either way. Headerless files now get column_0 instead of column_1.
Hash values for the default seed changed, and the guide reminds you that Polars "doesn't guarantee hash stability across versions." If you bucket users or sample rows with hash() and stored the result, the buckets move.
Last, cut() and qcut() are deprecated in favor of bin_intervals(), bin_quantiles() and bin_ranks(). The first two are left-closed by default, so pass right_closed=True to keep the same bins (the API docs mark bin_intervals() experimental). For score bands, that one argument decides whether a 0.5 lands in "low" or "medium".
The loud ones are the good news
The rest of 2.0 is Polars getting stricter, and I like it. Casting a string to a date now raises; use str.to_date(). is_in() no longer coerces lossily: the pre-release post shows an Int64 user ID above 2^53 matching the wrong float ID under 1.x. Horizontal concat with unequal heights now raises instead of padding with nulls; how="horizontal_extend" pads on purpose. Removed methods raise typed errors (AttributeRemovedError, ArgumentRemovedError) that name the replacement, which also lets a coding agent fix the call on its own.
LazyFrame.profile() is gone, and so is the DataFrame Interchange Protocol: for a library like Seaborn, the guide's fix is df.to_pandas(). Lazy SQL queries are no longer validated when you build them, only at collect(); call collect_schema() if you relied on early errors.
Out-of-core is on, with a 64 GB default
The streaming engine now spills to disk by default. The release post says it starts at about 80 percent of RAM, that this "may need tuning", and that the default disk budget is 64 GB. Sort, window functions and many expressions can spill; joins and group-bys come later.
On a small VPS with 4 GB of RAM and a 40 GB disk, the kind I run side projects on, that default disk budget is bigger than the disk. In the 2.0.0 source, the defaults come from POLARS_OOC_MEMORY_BUDGET_FRACTION (0.8) and POLARS_OOC_DISK_BUDGET_MB (64,000), and on Linux the spill directory is /var/tmp/polars-$USER/spill. I did not find them in the user guide, so treat them as internal knobs that may change. Either way, check that directory's free space before a big job.
The benchmark, as Polars reports it
Polars ran queries derived from TPC-H and TPC-DS against DuckDB 1.5.6, a DuckDB 2.0 alpha and DataFusion 54.0.0, on AWS c7a.4xlarge (16 vCPUs, 32 GB) and c7a.metal (192 vCPUs, 384 GB). By its own summary, default Polars was fastest on all but one benchmark. The same post admits a constant overhead at 192 threads that hurts small queries: at scale factor 10, going from 16 to 192 vCPUs made default Polars 1.8x slower on TPC-DS, and Polars limited to 32 cores was competitive or winning everywhere.
On Hacker News, a Polars developer posting as orlp explained why: the join creates T partitions for each of T threads, so T squared on a 192-core machine. My reading: more cores is not automatically faster on small data, and vendor benchmarks, which Polars itself says are not comparable to published TPC results, do not replace a run on your own machine.
What I would do on Monday
First, pin. The latest polars on PyPI is now 2.0.0, so an unpinned requirement installs 2.0 on the next fresh build or CI run. Write polars<2 until you have tested. The last 1.x release is 1.44.2, from September 9.
Second, move to 1.44.2 and run your tests with deprecation warnings turned into errors. Polars says most removed functionality was deprecated long ago. A clean run does not prove a clean 2.0, though: the streaming default changes results, not calls.
Third, the golden test. Run the pipeline on a fixed input under 1.44.2, save the output as Parquet, then compare under 2.0 (checked against the assert_frame_equal docs):
from polars.testing import assert_frame_equal
assert_frame_equal(old, new, check_row_order=False) # same rows?
assert_frame_equal(old, new) # same order?
If only the second fails, it is ordering: add maintain_order or an explicit sort. If the first fails, check dtypes and row counts against the list above. Then grep for explode(, how="horizontal", pl.datetime(, .hash(, cut(, has_header=False and any schema= passed to a CSV reader.
Worth it, or not yet
For a new project, I would start on 2.0 rather than build on defaults that are already gone. For lazy pipelines that fight memory, 2.0 is the release to take: in September, Polars said it expected the streaming engine to be "easily 5x faster" in aggregate. I would keep pipelines whose consumers depend on row order or on stored hashes on a pinned 1.x until the golden test passes.
Pandas is a different question, and 2.0 does not settle it. If your work is a notebook that feeds plotting and stats libraries, you lose nothing by staying. If it is a job that runs unattended at 3 a.m., I want what Polars wrote in its release post: "Errors should ideally raise up-front, not 20 minutes into a pipeline."
The release is boring in features. In results, it is only boring once you have tested it.
Sources
- Polars, "Release of Polars 2.0" (October 6, 2026)
- Polars, "Pre-release of Polars 2.0" (September 2, 2026)
- Polars documentation, "Version 2.0" (upgrade guide, read October 7, 2026)
- Polars documentation, "polars.Expr.cut" (read October 7, 2026)
- Polars documentation, "polars.Expr.bin_intervals" (read October 7, 2026)
- Polars documentation, "polars.testing.assert_frame_equal" (read October 7, 2026)
- GitHub, pola-rs/polars, "Python Polars 2.0.0" (October 6, 2026)
- GitHub, pola-rs/polars, "crates/polars-config/src/lib.rs" at tag py-2.0.0 (October 6, 2026)
- GitHub, pola-rs/polars, "crates/polars-config/src/spill_path.rs" at tag py-2.0.0 (October 6, 2026)
- PyPI, "polars" release history (read October 7, 2026)
- Hacker News, "Polars 2.0" (October 6, 2026; 410 points and 95 comments when read on October 7)
- Hacker News, comment by orlp on "Polars 2.0" (October 6, 2026)
