Comment by hatthew
8 hours ago
I feel like it's only within the past few months that opus got to the point where guiding the model is faster than doing things myself. I tried out sonnet recently and it was not a net positive to my work. I feel like anything that I'd trust haiku to handle isn't worth doing in the first place.
For context, I'm doing a range of tasks, everything from one-shotting adhoc scripts to having 4 hour 10M+ token conversations debugging things.
That is always amazing to me.
There is no way I can beat even local models at generating complex Python scripts fast.
hn is filled with uber geniuses.
I’m a different person and definitely not a genius but my experience today goes even beyond theirs.
I had Opus trying to simplify a query for me which was slow - it ran for maybe 30 minutes, including writing and running tests, and came up with a refactor across 9 files with a couple hundred lines changed. I was looking through the output before moving onto the next step, and noticed something a little fishy- I said “why does it do x, isn’t that a more complex y?”
Opus thought for another 20-30 seconds then output “Actually that would make the majority of the diff irrelevant, if we do that change it is just these 4 lines in this single file instead.
So then I had it do that. 5-10 minutes of writing and testing and that was done.
So my company spent $25 in tokens and I spent probably an hour in total for a 4 line change that, in the days before Claude, I probably could have found the correct file and thought through the problem, understood the solution, and written the 4 lines of code myself. Probably in the same amount of time.
So basically there was no benefit at all for my time, an extra cost to the company of $25, and now I understand our codebase a little bit less instead of more if I had done all the work.
As good as Claude is at building greenfield projects it still struggles a lot at complex ones
This is a recurring issue for us. Very small team. We move fast and pretty loose.
Dev + AI spend 3-4 hours on a project plan, there's a "wait a minute" moment, and finally they spend another hour dialing it back to a solution that could have been built, tested, and deployed in 2 hours.
Example: Someone was setting up a dev environment with multiple DB migrations from different branches - AI planned this wild 8 phase solution with a pretty fancy cutover event.
In review I essentially said... "Wait, isn't this a dev environment? It doesn't need 0 downtime, why not just destroy and recreate the DB" and it turned into a <1000LOC script.
Technically the original plan would have worked, it would have been more robust, but it would have taken a good deal more time to implement.
Some of this falls on the devs to know what fits our team well, what's realistic, what's obviously overengineered, etc... But some of it feels like AI just defaults to the most complex version of a thing. I catch it SUPER frequently. (And unfortunately some devs think that more complexity means it's a better solution)
Matches my experience to a tee.
And don’t forget the company also spent a bunch of money in tokens for the initial author to implement the thing poorly.
It's not so much that I can code fast, it's that it takes a significant amount of time to tell the model what exactly I want the script to accomplish, and at that point I might as well write the script myself. And often, figuring out what I want the script to do is that hardest part, so it doesn't really matter whether writing the script takes 10% or 20% of the total time.