Comment by tomrod

9 hours ago

Well done, Polars team!

Everything that I build greenfield moving forward I plan to use DuckDB, Polars, or PyArrow. Pandas was a great grandfather of a project (I actually cut my OSS contrib teeth on it, how the time flies)! I'll always appreciate the improvement pandas brought over SAS.

If Pandas was the great grandfather, R data.frame is the great-great grandfather. R data.frames directly inspired Pandas.

  • Agreed. For its time, R was a lot of fun to play with data pipelines, visuals, Quarto, and frontier stats. It's a shame it's so hard to make it work for a large swath of production use cases.

Do you have a take on when each of these three choices is the best one? I totally agree that these are the good choices, but I still find myself hesitating about which thing to reach for when!

  • Depends on need. We started using PyArrow on a reporting microservice when we realized we needed no additional functionality that pandas provided since it has better data type ergonomics. DuckDb is a great go-to for SQL based transformations when working with parquet files outside a managed system like Databricks. I want to actually test duckdb versus polars with a few lower level places like iceberg on S3.

    • Yeah makes sense. The SQL thing has also been my differentiator (and yeah, straight up arrow if you aren't doing much transformation of the data), but now I'm curious whether polars sql might be just as good. I kind of like that duckdb allows me to work with a database file, like a sqlite db. But maybe persisting parquet (or arrow directly?) is just as good?

      This is why I asked someone else who is also figuring this out!

      1 reply →

    • You can just use polars by leveraging write_ipc if you need arrow functionality. Pyarrow is great but it is very heavy.