Comment by rfgplk
6 hours ago
> A proper code review usually takes something like one hour per 200-400 LOC and you should be spending at least that much time on code review alone.
Not only is this not enforceable (how do you enforce how long someone spent working on a codebase on their own local machine?) the metric is severely off which instantly makes me question the competence of the llama.cpp dev team. You can easily review 10-100x that in an hour, even if you're being super pedantic about it.
I also just ran _one_ of their files (with include deps) through Astra and it detected >100 vulnerabilities/correctness errors (with over 10 outright UB/memory corruption issues). It's actually outright shocking.
> question the competence of the llama.cpp dev team
Following them for years, they seem extremely well put-together, and have excellent judgment. They are using the same policy as Linux and Debian (in my words, the speed of light is human understanding and judgment). Whether it is reasonable is a different question from enforcement, which typically comes down to "this seems fishy, explain your reasoning".
As for code review, the rule of thumb I've used for decades is: it takes about as long to review and understand as it does to write. Your 100x metric is completely outside of anything I've seen in any hobby or professional project, ever.
I'd like to see specific files you scanned and specific vulnerabilities cited.
If you can thoroughly and accurately review 100 * 400 = 40,000 lines of CUDA kernel code per hour, I know roughly 1,000 people who would love to hire you right now.
> You can easily review 10-100x that in an hour
40k LoC per hour of pedantic review? That's eleven lines per second, every second, for an hour.
What you have quoted is a sentence elaborating on the requirements listed in the document. This is provided to help you better understand why the requirements exist and the goal they are trying to accomplish.
> should
https://www.rfc-editor.org/info/rfc2119/
The reason you SHOULD take that time to read the output is because you must read it to understand it.
And the way this is enforced is explicitly called out in the document (and again in more detail in the linked AGENTS.md): the maintainers may ask you to explain it.
> You can easily review 10-100x that in an hour,
Frankly, I don't believe you. I'm half decent at writing CUDA directly (a holdover from a project a few years ago and it is a nice skill to have), the degree to which these are optimized is unlike 99.9% of all other code out there and even a tiny slip-up is either going to kill your results, your performance or both and if you're lucky only in some edge case. Understanding this code is hard work. I made a couple of minor edits to some .cu files in llama.cpp yesterday because I have a pretty weird setup which they obviously did not anticipate and it took a couple of hours to get it 'just so'.
> You can easily review 10-100x that in an hour
200-400 LOC, 10-100x = 2.000-40.000 LOC/hour for human review?
reviewer: LGTM
Just merge in main, what are you even pretending to review?
AI review should happen before human review, not instead of it.
I see frontier AI giving up and finding only nitpicking things on huge PRs, then finding logic bugs that were always there after cleanup.
Split your PR in smaller ones, both humans and AI will work better.