← Back to context

Comment by 392

10 hours ago

my understanding is Polars is faster, scales better without using external solutions, better API, +Rust. Pandas wins if you want to use what the vast majority of folks are using and have used in the past. Probably has a more complete set of helpers / recipes for the little things you bump into when using it thoroughly, but in the age of LLMs, I think that's minor.

>Pandas wins if you want to use what the vast majority of folks are using

Vast majority of skilled developers are now using Polars, unless they are constrained by lack of Narwhals support in their third-party library of choice (e.g. Great Expectations, SHAP). That's the more important trend to follow.

It’s the API that gave me the push to leave Pandas. 10 or more years of occasional Pandas use and I still had to google for any non-trivial queries.

In that regard, I’m still waiting for a credible jq replacement…

  • Learn SQL and interface with Duck. You will be 100x faster than Pandas/Polaris duo at fraction of memory. Also SQL is supported literally everywhere with a much more capable than Pandas API. Duck outputs to a Pandas Dataframe, but just treat that like a dictionary. Do all your processing, filtering and aggregation in Duck.

    Also, try fx.wtf as a replacement for jq. it comes with a in-built tui viewer that supports vi-keybindings. Ecmascript is built into fx.wtf so you can query the JSON with JS notation (where JSON was born). You can use any JS functions including map/reduce/filter or perform any kind of transformation instead of learning jq dsl that you will forget tomorrow.