Comment by fphhotchips
2 hours ago
It is absolutely without a doubt 100% not diverse enough.
I've spent years of my life running TPC benchmarks. They're useful guidelines, but they're simply not diverse enough to show the whole picture unless your picture is very simple. Certainly no single TPC benchmark in isolation. Maybe you can get a better idea by running all of them.
The largest gaps are around non-relational style data, JSON querying and the like. Nearly every organisation has something in that space now, and TPC-H/TPC-DS don't touch it at all.
> It is absolutely without a doubt 100% not diverse enough.
sure, nothing is 100%. But 90% may be good enough.
> JSON querying and the like. Nearly every organisation has something in that space now, and TPC-H/TPC-DS don't touch it at all.
unnesting json into relational data is some trivial op, and then you come back into tpc-h/ds realm.
> unnesting json into relational data is some trivial op
I mean, everything is representative if you just phrase the bits you don't want to do as "some trivial op". By this logic, Clickbench is a great and totally representative benchmark because joins are also trivial ops.
Billions of dollars are spent processing semi-structured data. It's a critical part of most analytical pipelines. You can't just ignore it and pretend like your benchmark is representative.
No. Clickbench:
1. operates on very small dataset (100M rows) with very low cardinality
2. It has very limited ops, trivial group by and aggregation
TPC benchs are way more diverse.
> It's a critical part of most analytical pipelines. You can't just ignore it and pretend like your benchmark is representative.
It is not ignored. All major DB engines provide utilities to work with json integrated into relational paradigm. Not sure what is your point.