I've used Polars before and can only recommend it.
It's like you get a really good 'query planner' like a DB would give you, but for your notebooks/scripts/etc. Much better than pandas imo.
show comments
gozzoo
I'm not following the trends closely, but has Polars become a full replacement for Pandas? Are there use cases where one is better suited than the other?
show comments
tomrod
Well done, Polars team!
Everything that I build greenfield moving forward I plan to use DuckDB, Polars, or PyArrow. Pandas was a great grandfather of a project (I actually cut my OSS contrib teeth on it, how the time flies)! I'll always appreciate the improvement pandas brought over SAS.
show comments
popularonion
I worked on benchmarking in the past, including TPC benchmarks.
When you see a blog post like this, never interpret it as “database A is X% faster than database B”, there are just too many factors.
It’s more like “we put focused work into performance improvements and we expect certain workloads to perform better than the previous release”.
This seems like a good project, and benchmarking is a good way for a development team to iterate on performance. Just want to get my take out there.
show comments
therno
I use Polars 2.0(rc) to (pre)calculate billions of weather scores on https://therno.com and it has been a lifesaver
Happy that I can upgrade to 2.0 final tonight.
niltecedu
A bit surprised about the datafusion results from the post, I have tried it time and time again, but datafusion has always been the leading/trading blowers with polars for our workfloads with duckdb being vastly slower.
show comments
sgarland
TIL that Polars supports SQL. Amazing.
show comments
PLenz
Does polars have a good equivalent of geopandas yet?
raoulj
Looks like SQL does not support `asof_join()`, not sure what else it doesn't support.
show comments
jdefting
I would be interested in seeing memory usage differences in these benchmarks. I’ve had issues with excessive memory usage in polars compared to DuckDB.
I’m guessing most of the this disparity should be solved by the steaming engine.
show comments
dkgs
A coincidence with the fact duckdb is supposed to release 2.0.0 very soon? :)
show comments
hnd9q09qk4
Thing I care about most is whether the old eager-vs-lazy footguns got cleaned up. Half my bugs were a stray collect() in a loop killing the query plan.
dartharva
Been using polars for over a year now, it is fantastic.
Kinrany
How does Polars relate to DataFusion these days? There's no reason for them not to converge into a single ecosystem, is there?
show comments
breezybottom
Looks like the Claude skill hasn't been updated
Centigonal
out of core sounds awesome! biggest thing that forced me to switch from pandas/polars to other solutions back in the day.
Vaslo
My team is all moving over to polars and DuckDB
show comments
patrick_rtk
amazing project !
bjourne
Pandas is amazing, Polaris is amazinger!
yshvrdhn
what about daft project ?
jt-s
Although I find pandas a bit aggravating in many ways, for myself and my equally idiotic laboratory scientist pals, seems that it is the default way you might interface with other libraries like SciPy (i.e. they expect things as NumPy arrays or pandas dataframes). Is this a real issue or will most things happily accept a polars dataframe? We don’t work with such large datasets that speed is likely a huge concern tbh.
show comments
tonyhart7
finally long time coming
cant wait to upgrade my Quant trading bot
show comments
nzgrover
Spelling error in first paragraph. enthousiastic
mrtimo
I can use pandas to clean a dataset, but each cleaning task is usually one line of code. OTOH, With DuckDB with one SQL statement I can replace 40+ lines of polars/pandas. You may reply, SQL isn't as easy to understand! Fair point, it's a declarative language... which is why I use Malloy. Malloy is to TypeScript as Javascript is to SQL. Malloy is much easier to read and write (just as TypeScript is) because it has a built in semantic model -- all the joins, measures, and dimensions are done in one place.
Here is an example [1] of visualizing college football games. Here are all the queries, and semantic model that power all the visualizations [2] Here is the AI generated typescript/react that does the visualizations [3]. The Malloy ecosystem has Malloyyo and Publisher which are replacements for PowerBI and Tableau and Looker. Here is another example for visualizing global trade [4].
I've used Polars before and can only recommend it.
It's like you get a really good 'query planner' like a DB would give you, but for your notebooks/scripts/etc. Much better than pandas imo.
I'm not following the trends closely, but has Polars become a full replacement for Pandas? Are there use cases where one is better suited than the other?
Well done, Polars team!
Everything that I build greenfield moving forward I plan to use DuckDB, Polars, or PyArrow. Pandas was a great grandfather of a project (I actually cut my OSS contrib teeth on it, how the time flies)! I'll always appreciate the improvement pandas brought over SAS.
I worked on benchmarking in the past, including TPC benchmarks.
When you see a blog post like this, never interpret it as “database A is X% faster than database B”, there are just too many factors.
It’s more like “we put focused work into performance improvements and we expect certain workloads to perform better than the previous release”.
This seems like a good project, and benchmarking is a good way for a development team to iterate on performance. Just want to get my take out there.
I use Polars 2.0(rc) to (pre)calculate billions of weather scores on https://therno.com and it has been a lifesaver
Happy that I can upgrade to 2.0 final tonight.
A bit surprised about the datafusion results from the post, I have tried it time and time again, but datafusion has always been the leading/trading blowers with polars for our workfloads with duckdb being vastly slower.
TIL that Polars supports SQL. Amazing.
Does polars have a good equivalent of geopandas yet?
Looks like SQL does not support `asof_join()`, not sure what else it doesn't support.
I would be interested in seeing memory usage differences in these benchmarks. I’ve had issues with excessive memory usage in polars compared to DuckDB.
I’m guessing most of the this disparity should be solved by the steaming engine.
A coincidence with the fact duckdb is supposed to release 2.0.0 very soon? :)
Thing I care about most is whether the old eager-vs-lazy footguns got cleaned up. Half my bugs were a stray collect() in a loop killing the query plan.
Been using polars for over a year now, it is fantastic.
How does Polars relate to DataFusion these days? There's no reason for them not to converge into a single ecosystem, is there?
Looks like the Claude skill hasn't been updated
out of core sounds awesome! biggest thing that forced me to switch from pandas/polars to other solutions back in the day.
My team is all moving over to polars and DuckDB
amazing project !
Pandas is amazing, Polaris is amazinger!
what about daft project ?
Although I find pandas a bit aggravating in many ways, for myself and my equally idiotic laboratory scientist pals, seems that it is the default way you might interface with other libraries like SciPy (i.e. they expect things as NumPy arrays or pandas dataframes). Is this a real issue or will most things happily accept a polars dataframe? We don’t work with such large datasets that speed is likely a huge concern tbh.
finally long time coming
cant wait to upgrade my Quant trading bot
Spelling error in first paragraph. enthousiastic
I can use pandas to clean a dataset, but each cleaning task is usually one line of code. OTOH, With DuckDB with one SQL statement I can replace 40+ lines of polars/pandas. You may reply, SQL isn't as easy to understand! Fair point, it's a declarative language... which is why I use Malloy. Malloy is to TypeScript as Javascript is to SQL. Malloy is much easier to read and write (just as TypeScript is) because it has a built in semantic model -- all the joins, measures, and dimensions are done in one place.
Here is an example [1] of visualizing college football games. Here are all the queries, and semantic model that power all the visualizations [2] Here is the AI generated typescript/react that does the visualizations [3]. The Malloy ecosystem has Malloyyo and Publisher which are replacements for PowerBI and Tableau and Looker. Here is another example for visualizing global trade [4].
[1] - https://mrtimo.github.io/cfb-games/games-2026.html?week=Week... [2] - https://github.com/mrtimo/cfb-games/blob/main/drives.malloy [3] - https://github.com/mrtimo/cfb-games/blob/main/dashboards/gam... [4] - https://tradeexplorer.org/