Comment by wokwokwok

3 years ago

If you accept that AGI is possible at all

How can a something that generates such a massive surge of interest, investment and research into AI not be a step toward it?

Saying it’s not a step towards AGI is basically saying AGI isn’t possible at all, because it means that all our efforts are making zero progress on AGI. That’s not a falsifiable position to take.

If you’re serious, the parent post literally said “AGI isnt going to look like this”.

…but realistically, how would a LLM that could easily refine itself from experiences, and had a very large context, let’s say, a billion tokens, be meaningfully different from AGI?

It could learn. It could remember things. It could generate human like output from a complex context.

Sure, it’s just a stochastic parrot… but if it can refine the model from real world inputs (learn new tricks, learn games, etc) and generate large scale (entire books worth) of coherent conversation and interactions… where do you draw the line between that and actual AGI?

Large contexts (35k tokens) are here right now. Refining models is here right now. They’re just expensive and slow (inference and training).

Maybe the current architecture doesn’t scale up beyond that and it’s a dead end, but my gosh.

If you don’t think what we have is a step towards AGI you really have to work hard to make your definition of AGI very very difficult to attain.

An AGI needs to be able to take an abstract concept and apply it to create a solution to a problem it has not encountered before at all - not sure LLMs can do that really. The lack of mathematics might be quite limiting there.

  • Can you give a concrete example of this problem that you expect an LLM to not be able to solve? It's fine saying "abstract concept" and "problem it has not encountered before at all" but these seem to me quite fuzzy concepts.

    • Sure. Ask it how to replicate the payoff of a financial derivative. It can explain the concept but it cannot use it on a specific payoff to arrive at a correct replication (beyond the odd widely published stuff). Taking ChatGPT, it will, however, talk about generic stuff, some incorrect stuff and some unrelated things when probed.

      Maybe also what I wrote a bit above: describe some greater than 3 dimensional objects and get it to stack them for some purpose could be another thing to try (I think, I will actually).

  • I think this whole “a problem never seen before” is something we need to rethink. Do people really work like that? I mean, I can’t expect a liberal arts major to solve a differential equation.