Comment by literalAardvark
16 hours ago
It's quite hard to measure until a wet lab gets behind the filter access to benchmark it.
But if you extrapolate from the ability it has in fields that aren't too strictly filtered, it looks pretty scary.
There are arguments against doing that but at first glance it seems like we just don't really know, and we likely won't: if governments decide they're interested in AI gain of function capabilities they won't be broadcasting that or allowing public benchmarks.
> But if you extrapolate from the ability it has in fields that aren't too strictly filtered, it looks pretty scary.
The closest unfiltered analogy to something as complex as chemistry or biology is most likely the softer fields like philosophy, the humanities and the softer end of the social sciences. Most practitioners and scholars in these fields would agree that AI is not nearly as compelling there as it might be in e.g. math, and that's putting it mildly and charitably.
Even coding shows the divide pretty well: AI writes code that manages to work (i.e. achieve its self-assessed functional goals) but the stuff is so unmaintainable that it ultimately poisons the AI's own context leading to mode collapse. This makes complete sense because maintainability is a soft objective that's especially hard to automatically optimize for in the short term, as part of a RL training run. The math folks themselves, too, now faced with a very real threat to their field from purportedly "hostile misaligned AIs", immediately zeroed in on education and exposition as something that LLMs are terrible at; with their abilities in systemizing and theory-building also being very much in question.