Comment by Aurornis

7 hours ago

> Want it to go away, almost like magic? Local inference. When its under your control, and no longer being forced to hold it wrong, all of the common LLM defects will go away.

This is just not true. Any local LLM you can host on consumer-accessible hardware has all of these defects, too. Adjusting the knobs doesn’t solve everything.

The closest you can get to frontier performance is Kimi K3, but you’re not hosting that unless your budget is on the order of a nice house in a good metro area.

I like my local LLMs as much as the next person and my office is currently uncomfortably warm from the amount of compute happening, but I would never agree that local LLMs solve all of the common LLM defects. This is peak wishful thinking.

In my experience, the local models and even the larger ones that we can’t run at home suffer more from long context degradation than the frontier models. You are exactly right that you need to manage context length, but even at fp16/bf16 the local models have a lower ceiling for usable context length in my experience.

We're on the cusp of Kimi K3 becoming usable on sub-10k hardware.

https://github.com/gavamedia/deltafin

  • 14 seconds per token? Not tokens per second. Seconds per token? That’s nowhere near the cusp!

    • It's need it to be an order of magnitude cheaper, but for some queries ("What to discuss at tomorrow's meeting") I can wait 12+ hours.

    • At that rate, it would take me a mere 15 days to generate the number of tokens I typically use in a day.

      "Yes, I'm using AI to speed up development. I'll submit that PR in two weeks time!"

  • For values of "usable" that include "14.6 seconds/token". It's a cool accomplishment! And newer hardware would speed it up some. But I think I'd want something a bit faster before declaring it usable in practice.

This is a bit of a strawman. The parent comment wasn't suggesting that the inherent defects magically disappear nor was it suggesting that local inference today is sufficient for all tasks.

At the heart of it, self-hosting liberates your use cases from all the horribly opaque configuration, shadow prompting, etc. And local models are only getting better and more diverse every month.

  • >The parent comment wasn't suggesting that the inherent defects magically disappear

    that is almost exactly what they said though..?

    "Want it to go away, almost like magic? Local inference. [...] all of the common LLM defects will go away.

    • "[...]"

      -----

      edit: I suppose I have to spell this out.

      "[...]" = "When its under your control, and [you're] no longer being forced to hold it wrong," ≈ "liberates your use cases from all the horribly opaque configuration, shadow prompting, etc."

      and

      "Almost like magic" is cliché hyperbole. It also means "not magic" in the same way that "almost like a dog" means "not a dog." It was immediately followed by the explicit clarification that was surgically omitted above, in absolute and conscious bad faith.

      Hence, "[...]"

  • > This is a bit of a strawman. The parent comment wasn't suggesting that the inherent defects magically disappear

    The parent comment literally said that the common LLM defects would go away.

    Direct quote:

    > all of the common LLM defects will go away

  • I get your reading. My interpretation is that they’re claiming enshittification and that the local models don’t have that.

    It’s an understandable view, but I’d be astonished by any local model processing very long context better than any frontier model (and now many racks are we talking).

  • > Want it to go away, almost like magic?

    True, almost like magic and magic are not the same thing, but I'd hesitate to call this a strawman.