Comment by saejox
6 hours ago
This is a project i wanted to implement for a long time. It regularly benchmarks cloud hosted models with private benchmarks. Not just openai & anthropic, popular openrouter models too.
Tests their intelligence, not their diligence.
Sadly i cant think of a way to monetize the service. Also if it ever gets famous enough labs would try to game the system, it would be cat&mouse game that i am not willing to waste time on without any monetary gain.
The only revenue model I for this is ads (like AI Stupid Level[0]). Or as a loss leader to get eyeballs to your service (like Margin Lab[1]).
EDIT: I forgot (and am shocked) that HN still doesn't seem to support Markdown-style links.
[0] https://aistupidlevel.info/
[1] https://marginlab.ai/trackers/claude-code/
I mean I think if this is done well, lots of companies would pay for access to that data. Think like Enterprise subscriptions.
Its similar to other data services I see around my F500 company.
I built GitHub.com/adrianco/retort to do this. It’s runs lots of experiments and you can contribute results if you have some spare tokens. You can add your own tests, and it runs Claude, Codex, Gemini, Hermes for local models.