← Back to context

Comment by onlyrealcuzzo

2 days ago

If the benchmarks don't lie, this is getting very close to Opus 4.6 capability - which was the turning point for me for when AI was "good enough" that it became very hard to justify not using it.

I'm sure there's some benchmaxxing going on, and some things you get only with a a larger model.

But I'm feeling pretty confident if not by Gemma 5 than by mid 2028 we'll have local models that are almost always as good as Opus 4.6 was and in many cases far better.

What kind of things you only get with a larger model?

  • Asking it factual information[1]. You just can't compress the entire human knowledge into a 30GB file.

    [1] Without searching the internet. And even if you allow it, you'll get much worse results because search means browsing and parsing the top results, and search results are horrible, whereas internal knowledge from training encompasses the entire internet plus all books including very niche stuff.

    • > You just can't compress the entire human knowledge into a 30GB file

      Fortunately, that isn’t necessary! What LLMs need is a level of fluency with key concepts so that they can (1) make effective use of retrieval tools and (2) understand the material in the context window. 30GB-sized models can absolutely store enough knowledge to do this.

      Here is an example from the field of law. Most lawyers who have litigated contract disputes in California know about Civil Code § 1717, which makes any contract providing for attorney fees to a prevailing party mutual, so that even if the contract was written to be one-sided, it won’t be enforced that way. It’s a simple enough concept, but there are many more particulars to it, such as what happens when the fee provision is only written to apply to part of the contract. (Answer: it depends on other facts.)

      When a lawyer recognizes that they’re in a situation where § 1717 is relevant, the first thing they will do is pull the statute and read it, because nobody has it memorized. And they don’t need to.

      2 replies →

    • > You just can't compress the entire human knowledge into a 30GB file.

      ...I'm just asking questions here... how sure are we of this?

      If you'd asked me six years ago whether we could compress all of human knowledge into a 1 TB file, I would have said no, and yet here we are.

      If you can do 1 TB, why not 30 GB? It's, like, within one order of magnitude.

      10 replies →

  • IME using 5.6 Luna and DS V4 Flash, I notice that although they are excellent at programming, even Opus-like in the way they try to debug, the thing they are worst at is inferring user intent and making good decisions with little information. They are absolutely terrible at that, will misinterpret small wording ambiguities. I suspect that's an ability you can't add with RL training, that it requires the depth of understanding from vast pre-training.

  • Similar to the way they asked Sol to solve Erdos problems, that's what I want my model to do for programming.

    I don't want to try to take my best educated guess at what the best design is BEFORE implementation - especially if you're designing a feature for a codebase you're not an expert in, you don't know like the back of your hand (i.e. one that is mostly or entirely LLM generated).

    What sounds good on paper - often times becomes unideal in practice when you get to the reality of implementation.

    It may not be worth re-architecting your entire system to get to a "pure" design that would be the best - all things considered.

    Instead, I'd like the model to independently design many plausible and coherent good solutions, then implement each of them, then intelligently pick the few winners (after its fixed any bugs that could be causing promising solutions to look artificially bad) - unless there's an obvious one - and then give me the data I need to make an informed decision on which one to go with, all before I even look at the design or implementation.

    You're not getting this from a one shot prompt from a 30B model today. You can't even really get it from Sol or Fable - IME. But you can get somewhat close.

    • You're going to be waiting for a while.

      Even Fable is bad at this, I would constantly have to fix it going down architectural dead ends or just making obvious mistakes.

      Which sucks for people that want LLMs to do everything like a genie, but does mean senior engineers have a few more years before they become redundant.

      3 replies →

    • I think someone ought to encode this into a harness. It is really insightful into how we should be spending time if it is going to be spent reviewing AI code.

    • This is a harness problem not a model problem, try prime agent it can do that and it will do it well even :P but you need to prompt it in according to its tools and processes.

  • Smaller models are overconfident and have a hard time to self-correct.

    If it’s stuck, usually that’s it.

    Bigger models “understand” better, both the prompt and the contents. If you will try to read a paper together with a smaller model, the difference is immediately obvious.

    Bigger models will “forget” and drift much less.