Polars 2.0
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on networking and server gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Polars says version 2.0 switches LazyFrame collection to its streaming engine by default and enables initial out-of-core processing that can spill supported operations to disk. The release also strengthens SQL support and introduces a Map dtype; performance comparisons cited by the project are based on its own benchmark setup.

The Polars project has released Polars 2.0, making its streaming engine the default when users collect a LazyFrame and enabling initial spill-to-disk support for certain operations. The major-version release also expands SQL capabilities and adds a Map data type, changes aimed at broadening workloads Polars can handle and reducing memory pressure on supported queries.

Under the new default, calling collect on a LazyFrame uses the streaming engine. Polars says this can bring substantial memory and performance improvements on many queries. The change can affect observable row order for operations including joins, group-by and unpivot. Users who need to preserve order for those operations can set maintain_order=True.

Polars 2.0 also enables out-of-core processing, in which supported operations can spill data to disk when memory use rises. The release report says spilling starts at about 80% of RAM and the default disk budget is 64 GB; the project notes that the trigger may need tuning. Current support includes sorts, window functions and many expressions. Joins and group-by operations are not yet supported for out-of-core execution, though Polars says it plans to add them.

The release adds a native Map dtype corresponding to Arrow’s MapType, representing key-value data in a dictionary-like form. It also includes SQL and engine changes such as join reordering, common-subplan elimination and dynamic predicates or bloom filters. Polars says the changes improve SQL query execution, but the published performance results are the project’s own benchmark findings, not independent verification.

At a glance
announcementWhen: Announced in the supplied release repor…
The developmentThe Polars project has released version 2.0, changing default query execution and adding initial disk-spilling support, SQL improvements and a Map dtype.

How Disk Spilling Changes Query Limits

The new defaults matter to analysts and developers working with datasets that can exceed available memory. Streaming execution and disk spilling may let supported queries finish under tighter memory constraints, rather than requiring the entire workload to fit in RAM. This could make Polars more practical for larger or less predictable workloads, although the benefit depends on the operations a query uses and on available disk capacity.

There is also a behavior change users need to account for: row order is not guaranteed by the streaming engine for certain operations unless they request it explicitly. Applications that depend on a particular order may need code changes or checks during migration. Meanwhile, first-class SQL support gives teams another way to use the engine, but the benchmark claims should be read with the test conditions in mind rather than treated as a universal ranking.

Amazon

high performance SSD for data processing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Benchmark Behind the SQL Claims

To compare SQL performance, Polars says it ran queries on data derived from TPC-H and TPC-DS against DuckDB 1.5.6, a DuckDB 2.0 development build and DataFusion 54.0.0. Tests used two machines: one with 16 virtual CPUs and 32 GB of memory, and another with 192 virtual CPUs and 384 GB. The project used generated Parquet data stored on EBS, ran each query five times in a hot setting and reported the best run, comparing both total and geometric-mean query times.

Polars reports that its default configuration was fastest in all but one of the tested benchmarks. It also says the tested engines completed all queries except for DataFusion, which timed out on one TPC-DS query, timed out once on another, and ran out of memory on a TPC-H query on the smaller machine; those queries were excluded from results for all engines. Polars acknowledged that its default configuration has overhead on the 192-thread machine that hurts smaller queries, and said it hopes to address that in a later release. It shared a benchmark repository for others to reproduce the tests.

“Calling collect on a LazyFrame will now default to the streaming engine.”

— Polars, in its Polars 2.0 release report

Amazon

large capacity external hard drive 64GB

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of Current Disk Spilling

The release report does not provide independent validation of its performance claims, and results may differ with other hardware, datasets, storage and query mixes. Polars encourages users to replicate its benchmark, but the comparisons describe the project’s specified setup and methodology.

Out-of-core support is still partial: joins and group-by operations are not yet covered, and the roughly 80% RAM spill threshold may require tuning. The report does not specify when those additional operations will be supported. It also does not quantify the performance impact of the changed row-order behavior across different workloads.

Amazon

RAM disk for out-of-core processing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

More Operations and Benchmark Replication

Polars says it plans to extend spill-to-disk execution to joins and group-by operations, which would broaden the workloads able to continue when memory is limited. The release report does not give a schedule. The project also says it hopes to fix the thread-scaling overhead seen in its tests in a subsequent release.

For users evaluating the upgrade, the immediate next step is to test representative queries, check whether any operation depends on output order, and review memory and disk limits. Other teams can compare the SQL results using the benchmark repository Polars published, while accounting for differences in hardware and test conditions.

Amazon

SQL database management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main change in Polars 2.0?

Calling collect on a LazyFrame now uses the streaming engine by default. Polars 2.0 also enables initial spill-to-disk support and adds SQL improvements and a Map dtype.

Does Polars 2.0 preserve row order?

Not by default for some streaming operations, including joins, group-by and unpivot. The release report says users who need observable ordering can set maintain_order=True.

Which operations can currently spill data to disk?

Polars lists sorts, window functions and many expressions as currently supported. Joins and group-by operations are not yet supported for out-of-core execution.

Are Polars’ SQL benchmark results independently verified?

The cited comparisons were run and reported by Polars using its stated TPC-H and TPC-DS setup. The project shared a repository for replication, but the supplied report does not describe independent verification.

Source: hn

COLUMBUS DAY / I

Columbus Day / Indigenous Peoples' Day Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

MIT Qubit Design Could Speed Quantum Operations While Preserving Data

MIT researchers propose a new qubit architecture that could enhance quantum operation speed while maintaining data integrity, marking a potential breakthrough in quantum tech.

Cold Archive Retrieval: The Trade-Off Everyone Forgets

What you don’t realize about cold archive retrieval is how the slow access times and costs impact your data management strategy—discover the trade-offs everyone forgets.

China International Big Data Industry Expo Surges In Global Coverage

The China International Big Data Industry Expo has seen a surge in international media coverage, highlighting its growing influence in the global data industry.

Data Replication Explained: Synchronous Vs Asynchronous

Just understanding the differences between synchronous and asynchronous data replication can transform your data strategy—discover which method best suits your needs.