← Back to context

Comment by rf15

3 days ago

I've worked with these systems for four years now and they have not meaningfully improved in that time frame.

We still have:

- statistical correlation between two things will always cause one thing to lead to the other, no matter how much you prompt it to not have that connection (to be expected with a stochastic system)

- Math completely fails in longer contexts

- "thinking" token generation being on the correct track just to 'no, wait' on an already correct conclusion

- smearing of properties between logically distinct objects (a red ball and a green cube can quickly become a red cube and a green ball)

> I've worked with these systems for four years now and they have not meaningfully improved in that time frame.

Not meaningfully improved?! Four years ago was gpt *3.5*! ChatGPT hadn’t been released!

  • Yes! Impressive, isn't it? I see how it has improved for some minor points, that the big models can cover more finetuning ground, but my big gripes are still the same - you could do the same back then with multiple models and more targeted finetuning.

    • > you could do the same back then with multiple models and more targeted finetuning

      Definitely not, lol.

    • > you could do the same back then with multiple models and more targeted finetuning

      I mean, come on, this is just not true. You could not achieve anything like what you can with modern agentic coding with Fable / 5.6 Sol from any combination or configuration of GPT 3.5 era models.

      It's like saying that a teenager isn't an intellectually meaningful improvement over a toddler.

      Sure they're both still fundamentally flawed humans prone to cognitive error, but one is clearly more likely to hit the mark than the other when assigned a task.

    • The only thing impressive is how wrong you are. LLMs have improved by an absolutely incredible amount in the last 4 years.

    • Utter nonsense.

      There’s no way you could get models as smart by fine tuning. I couldn’t throw a problem like “build a pokemon database with UI to teach my son sql” and get a working system, nice ui, tests (which it iterated on) examples and explanations in one shot.

      There weren’t thinking tokens. Maths is now dramatically better, making actual contributions when before they were mostly mocked for making extremely basic errors. Smearing is also something say is very rare in frontier models.

      If you think they have barely changed you’ve either forgotten what they were like or not used them more recently, or you’re just being obtuse.

      1 reply →

    • > you could do the same back then with multiple models and more targeted finetuning

      Are you one of those anonymous billionaires as if you did this a few years ago, you would've been famous and rich.

      6 replies →

> - Math completely fails in longer contexts

Not sure what longer contexts we're talking about but didn't we have an old math problem optimized, which even the LLM itself was surprised about, just a week ago? Something which wasn't possible 6 months ago.

  • I mean calculations, not mathematical proofs

    • If they use Python to fill the gap, and the end user doesn’t have to know or care, is it unfair to assess this as progress and attribute the progress to the _system_?

      OK, the core technology that is the language model still can’t math as well as you’d hope, but how about the end result users see from the system when they interface with it?

      “Did you know humans are better at flying today than they were a thousand years ago?” ‘No they’re not, they need planes.’ Technically correct in a way but isn’t it kind of annoying to be so stubbornly pedantic when the context is speed of reaching Point B from Point A?

      3 replies →

Messages like this in the training data are how LLMs learn to say absurd things with total confidence.

Like how toddlers’ skills don’t meaningfully improve on infants’, because either could wake up in a wet bed.

  • Let us be more clear: there is no structural jump, no architectural overcoming of the original fault.

    (Edit: and on a similar point, structural properties such as having static ntetworks, as opposed to continuously learning and improving architectures (such as us), will reveal that there is still road ahead.)

    • Maybe more fair then would be: “I've worked with these systems for four years now and while they _have_ meaningfully improved in that time frame, they’re not perfect and remain fundamentally flawed in various ways.”

      You prompt less. You need not inject search results into the context window yourself, a window much larger than years ago. You get code that’s already been run successfully once instead of finding an obvious show stopping bug yourself.

      The technology is not a brand new one that fixed everything wrong with the old one, no, but not sure I would’ve noticed your comment if it had been such a bland observation. I genuinely assume good faith here… will say am tempted to assume the standards of someone posting such a thing might be impossibly high. Glad to be having a fun conversation instead of getting your grades on my work product or something :)

      4 replies →

> I've worked with these systems for four years now and they have not meaningfully improved in that time frame.

That's absolutely insane. Is it some case of anti-AI psychosis?