Comment by lxe

8 hours ago

Why isn't this type of expert caching in the native llama.cpp yet? Why do we need a separate codebase?

One reason is that the llama.cpp team (GGML) has strict requirements that a human must understand the code they are contributing. If a project is fully vibe coded they can’t contribute. So a lot of projects where an AI went and coded a bunch of custom kernels to increase speed are left to their own devices.

I think this is a fine behavior. We can have upstream purists that are strict gatekeepers but don’t get in the way of downstream forks. Debian has some this in the Linux landscape for a long time, and it has enabled Ubuntu, Mint, etc. to flourish without compromising themselves.

  • >If a project is fully vibe coded they can’t contribute

    The irony (however mild) is apparently lost on the rest of the field.

  • "A human pretends to understand it" signifies what exactly?

    What you really mean is, the core team there doesn't want to lose control.

    Which isn't really predicated on contributions not being "vibe coded" or whatever.

    When quality is the problem, you need to be able to make your standards explicit, or you're just gatekeeping irrationally.

    • > What you really mean is, the core team there doesn't want to lose control.

      It's 100% this. They basically produce vague guidelines such that only the core maintainers are allowed to use LLMs, under the guise of "well of course we understand the code" and no one else is. It's also completely unenforceable, how are they going to prove whether someone understands the code or not? Even if they show sufficient evidence/understanding the maintainers can simply sabotage them and accuse them of using an LLM to explain the code. No one wins here.

      1 reply →

The "mainstream" inference engines are notoriously slow to integrate this stuff, to an extent understandably given the complexity of ensuring numerical accuracy alongside supporting a wide array of systems and models. Part of it is that not everyone is willing to bring what they develop into a pull request because they vibe coded it and don't care to deal with whatever quality requirements the more well known inference engines have.