← Back to context

Comment by desipenguin

9 hours ago

From recent Python Bytes podcast (https://pythonbytes.fm/episodes/show/496/a-lake-house-in-sea...)

> 1 Billion Row Challenge benchmark: Pandas took 4m28s vs. Polars 5.04s and DuckDB 5.19s — DuckDB also used 19x less memory

Python Vs Rust : In terms for speed - No comparison

(The above episode transcript has a link to blog post titled "Pandas should go extinct" )

I've nearly entirely switched to DuckDB for anything more than like 500 or 1,000 rows or if there are a tonne of columns.

Polars is great, but I'm just too used to the Pandas API to use it as a replacement for the cases where DuckDB is overkill.

  • That's a shame, because Pandas has a really quirky/legacy-burdened API and Polars is super clean by comparison. As someone who had Spark and Pandas experience before switching, Polars felt like Pyspark without the added mental overhead of needing you to think about multi-worker-node parallelism

    • It's just muscle memor, I've been using Pandas daily for over a decade, but I'll likely eventually switch over to Polars.

      Part of the problem, as I said, is that I'm just spoiled by DuckDB when performance matters.