Comment by sanderjd
5 hours ago
Do you have a take on when each of these three choices is the best one? I totally agree that these are the good choices, but I still find myself hesitating about which thing to reach for when!
5 hours ago
Do you have a take on when each of these three choices is the best one? I totally agree that these are the good choices, but I still find myself hesitating about which thing to reach for when!
Depends on need. We started using PyArrow on a reporting microservice when we realized we needed no additional functionality that pandas provided since it has better data type ergonomics. DuckDb is a great go-to for SQL based transformations when working with parquet files outside a managed system like Databricks. I want to actually test duckdb versus polars with a few lower level places like iceberg on S3.
Yeah makes sense. The SQL thing has also been my differentiator (and yeah, straight up arrow if you aren't doing much transformation of the data), but now I'm curious whether polars sql might be just as good. I kind of like that duckdb allows me to work with a database file, like a sqlite db. But maybe persisting parquet (or arrow directly?) is just as good?
This is why I asked someone else who is also figuring this out!
Really comes down to query speed and management cognitive cost. I look forward to trying the new version for polars.