Comment by batshit_beaver
8 hours ago
The challenge with comparing these things to humans, is that humans learn. A newbie might not respect your organization’s set of policies on day one, but what about 3 months in? Or 3 years? Meanwhile there’s still no reasonable mechanism for automatically fine tuning LLMs or adjusting their harnesses to make them better at completing your organization’s objectives more successfully. They’re still overwhelmingly governed by the shared weights and harness policies found to be successful for the average case.
Models learn. It just costs $10B and 1 year to do what a human does every night.
LoRAs or even full fine-tunes would be much cheaper than that, and with some investment in the right infra could be updated regularly. And at least LoRAs can be swapped in and out cheaply, making them usable in large-scale inference providers. But there seems to be limited appetite in offering this. Both Anthropic and OpenAI no longer offer fine tuning for current models
does lora do a good job at teaching the model new things that werent in the training data?
without trillions of examples of following instructions at a million context length, im not convinced the behaviour is in the weights to begin with
1 reply →
At least these costs are currently preventing the planets surface from being covered in paperclip maximizers for the moment.