Comment by SukadarBukadar
2 hours ago
Pandas is better for slight in-place or per-row modifications, for loading from less conventional datasets, for transposition/more nuanced row-based aggregation/multi-axis manipulation, for performance when multiprocessing can be used, for interoperability with other libraries (e.g. plotting and statistics)... When you need to operate on huge datasets, use DuckDB, because its performance is even now on-par with Polars regarding speed, while handling huge or more complex joins is a huge win for DuckDB because it better offloads intermediate results to disk, while Polars just dies on me. I haven't tried such joins in Polars 2 though
No comments yet
Contribute on Hacker News ↗