Comment by jillesvangurp

14 hours ago

People forget that pull requests and doing code reviews in the context of those is still a fairly recent thing. People did some code reviews before that of course but nowhere near as strictly. Same with testing practices, static code analysis, etc. Most of that wasn't all that common until beginning of this century. I remember using findbugs with Java around 2004. It actually found bugs in my code the first time I used it. No review had caught those. And we got lucky not finding them in production. But they were definitely bugs. Our system didn't have unit tests; it was all manual. Junit was a fairly recent system that hadn't been around for that long yet. Our build was done with Ant. There was no test phase. Our tech lead would of course check my work and correct & educate me (I learned a lot). But a lot of bugs slipped through as well.

Git did not exist either, I migrated out cvs to a beta release of Subversion. We only used branches for releases. We'd cut a branch just before a release. Test it (manually) and then ship. That was a process I helped put in place actually. After release, master would diverge quickly so back porting fixes was not really a thing. We'd support releases for as long as our customers used them. Often that involved just upgrading them to the recent version. We shipped when things were good enough.

I think the notion of people reviewing any meaningful amount of generated code is simply delusional. As you say, we do need alternative means to replace those checks. And a lot of that is going to be AI driven as well. AI driven testing, code reviews, and all the rest. Essentially all the stuff we used to do manually (poorly).

And we do have an important new tool as well: clean room code replacement. That used to be prohibitively expensive but now it's not. If you have something that is well specified through documentation, APIs, specifications, tests, etc. replacing it is fairly straightforward now. There are some early examples of people using LLMs to generate functioning replacements for things like Postgresql, browsers, compilers and similarly large and complex systems. While not perfect, these things seem to work, pass their tests, and generally not be completely horrible. It's only going to get better from here.

The notion that people are going to ever manually review code that was generated for such systems in mere hours/days is beyond imagination. How? When? Who? Why? It simply does not scale. It's only going to be more and more code. The amount of code no person will have ever looked at will soon dwarf the amount of code that is still manually inspected/created pretty rapidly.