← Back to context

Comment by dmix

12 hours ago

Finding the bugs with LLMs is easy. Reviewing the output, cleaning it up, and making sure it doesn't break something else is the hard part.

This is where I believe strong typing (like, Haskell-strong or stronger) and functional programming in general will be a win. The confidence I have that my fixes are localised when fixing Haskell code is infinitely stronger than fixing even Java, not speak about C, code.

  • Haskell's type system would not easily prevent this bug. It's not good at numeric/logic issues like that. When people say "Haskell makes it impossible to write bugs" they mean "Haskell has enums" (ADTs).

    • Liquid Haskell might require you to prove that the divisor is nonzero, but even in standard Haskell there's common idioms for ensuring that a list is non-empty (data NonEmpty a = a :| [a]) or that text is non-empty (newtype NonEmptyText = NonEmptyText Text, with non-exported constructor, helpers like make :: Text -> NonEmptyText, or more advanced tricks like https://exploring-better-ways.bellroy.com/haskell-koan-type-... ).

      The big problem preventing this approach from working for numbers is that it's just so cumbersome there. Most of this is because all the arithmetic operators are bundled into a single Num typeclass, and `fromInteger :: Num a => Integer -> a` has a type that's impossible for a "non-zero number" wrapper to satisfy.

      6 replies →

    • I am not claiming you cant write buggy code in Haskell! But following good functional style, your bug will more likely be compartmentalised, and fixing it will not break some other part of your program.

      2 replies →

  • Imo, formal methods like more expressive/stricter type systems are key to making LLM generated code successful. Of course models will get better, but trusting the output will become much easier with a type system that proves more properties.

  • What's stronger than Haskell?

    • Dependent types is one possible direction. Not sure when a language with dependent types will arise which will be useful for making real programs.

      Agda is the most mature dependently typed programming languae (having been around since the 90s – it is basically Haskell on steroids), but has a more proof-assistant flavor than an actual programming language flavor. Opus & Fable write Agda quite well, so LLMs can understand dependent types.

If finding the bugs with LLMs is easy. Then making sure it doesn't break something else is just LLMs finding no bugs. Easy.

That hasn’t been that bad. My real issue has been the time sink involved in following along with the maintainer and jumper through their hoops. Even after I demonstrate a flaw and a potential fix. My schedule is just so busy I need to pencil in time to deal with them.

The missing part of this is that verifying the bug with LLMs is also easy, and so is adversarially reviewing the proposed fix with LLMs.

The only thing left for you to do should be directional decisions. The LLMs should pause and rope you in if the fix involves directional/invariant changes.

No one can keep up with the volume of code AI produces.

We wont stop using AI.

We will use AI to check AI.

Of course this is crazy, but it will also unlock pretty insane scaling and productivity and ultimately we will manage it on either end via requirements and tests.

  • > it will also unlock pretty insane scaling and productivity

    Insane scaling of bloat, bugs, and technical debt I'd say.

    > We will manage it on either end via requirements and tests

    It is so crazy that this is being touted as a sane strategy. When I was a much worse programmer, I tried to write a big complicated string manipulation function to take two types of scripts in a language and add diacritics. I had the requirements very clear. I had the tests very clearly with all the edge cases. But I didn't have a good and clear picture of how to attack the problem which was quite novel for me. As I got closer to passing all the tests it got exponentially more unruly and confusing. And nearing the end I was frantically changing little bits here and there wincing and praying and hoping the tests would pass. "Please work! Come on!" Then when I got close enough, I could never ever think about touching that mess again.

    I was a below average programmer then throwing myself at some novel problem I didn't understand. Throwing LLMs that produce below average code at novel problems and relying on tests and requirements is not where we want to go to make real progress.

    (Years later after much learning and coding myself I was able to redo the function in a totally different way. This time I actually understood how to attack the strange problem and made something clean, clear, and robust that just worked. The tests then become a secondary guardrail, not the main force of correction.)

    We are seeing such a massive regression from what we've learned over the years of CS.

    • >Insane scaling of bloat, bugs, and technical debt I'd say.

      You just described every legacy codebase. Many of which are widely used and do a lot of sales. You dont need a clean codebase to have a valuable product.

      >It is so crazy that this is being touted as a sane strategy.

      Re-read what I said. I literally called it crazy.

      It is the same dynamic that gave us customer service from some call center in India. Why would companies do this? Customer service got worse. Are they stupid? No, it's just worth it. The quality goes down but the business can scale more so it doesnt matter.

      AI will absolutely be good enough at doing things that we'll happily accept some jankiness at times so that we can devote an extra 3000 hours per year per person to other things.

      Im not even suggesting its a good thing. I just think the incentive structure dictates it. You're not going to have time to maintain a small slice of some service by hand.

    • I think all code is technical debt in a way. Good code is a necessary evil, bad code is more evil than necessary.

      Generating code automatically when you're not even quite sure what it is or even should be doing is insanity.

    • You shared a story of a novice incompetent human programmer and this should tell us that AI is bad at coding.

  • It's mostly (not entirely, but mostly) finding security issues in old human-written code. It'll eventually start running out of those.

    From that standpoint, it's not a crazy setup security-wise. Maybe still crazy for development.

    • You can point AI at any AI produced code and ask it to review it, get back 10 bullet points and a few pages of prose. And the fun part is, you can do that over and over and over again!

      1 reply →

  • You're suggesting that LLMs get better at fixing bugs/vulnerabilities, but at the same time stop getting better at finding them? What if this difference is inherent and essential?

    • > You're suggesting that LLMs get better at fixing bugs/vulnerabilities, but at the same time stop getting better at finding them?

      Are you implying that all code writing by LLMs atm is bug-free?

      2 replies →

  • Volume..... <sigh>

    It used to be considered a quality of good code that there would be less code, not more.

    Some people always tryin to get the highscore on golf.

    • You can have both less code per problem and more code overall when you make problem solving cheap enough.

  • In fairness at root this has been going on for awhile. No one can keep up with the volume of machine code that modern more abstracted codebases produce.

    We didn't stop using syntactic programming languages we used code to check code.

    Not sure it's really crazy at all. It's been an abstraction for programmers probably since we stopped soldering transistors to each other.