← Back to context

Comment by Blikkentrekker

5 hours ago

That they do it is just concerning to me in that it says that home-ran models just aren't good enough. Surely researchers like this have the processing power to run them at home, they just don't have the processing power to train models of comparable level.

This is something I feared would happen and where open source would be left behind. Maybe they can do something with crowd-sourcing computational power from volunteirs. They were after all able to get Leela Chess Zero to be comparable to AlphaZero by training from volunteer processing power but it seems to me we live in a world now where the best models keep their stuff closed.

In imagine generation too. I'm not sure how well Stable Diffusion can compete in following instructions with all those advanced models that are kept secret.

> Surely researchers like this have the processing power to run them at home,

Nobody has that power. Certainly not mathematicians.

  • Yes people have to service the use case of a single person. All sorts of models exist that are designed to be ran at home. Yes, you need a relatively beefy graphics card for it but if your machine can handle the latest video games at good graphics it can handle these local models. How well they compare against these things remains to be seen but apparently OpenAI released gpt-oss-20b which is comparable to o3-mini apparently in terms of reasoning power. There's also oss-120b which does require at least a company server to run but this should well be within the budget of a university.