Comment by crazybonkersai

15 hours ago

This does not reflect my experience. AI code review has improved tremendously in the last years to the point that nobody in my team does manual code review any more. It used to be ridiculously bad two years ago, but not any more.

Do you use some "framework" specifically or (custom) skills, or is it just plain prompt to review the code?

  • At a workplace we have a review skill which includes things like skipping nitpicks and trivial issues. Myself I just use a simple prompt, as well as Copilot review on Github. There is a school of thought to use a different model for a code review and code generation, but I have no opinion about that.

As recently as yesterday, CodeRabbit was absolutely ridiculously bad.

I mean, depends on your reference point of course. Sometime around 2015 I participated in ICFP contest, where the task was in the code synthesis domain. At the time writing code that can generate basic arithmetical, well, forget it, even logical operations to implement some high-level description of a program was far out of hand. So, compared to that, CodeRabbit is light years ahead and is awesome beyond belief. But, compared to a trained human it still sucks.

Hearing conflicting reports on performance of AI aids, my attempt at explanation is that some problem domains have much better coverage. Essentially, the further away you are from "fullstack" the worse the performance is. So, maybe it does well on your end, it's because the project you work on is a well-researched problem that has many similar projects that help AI to distinguish the patterns it can then readily find and implement?