← Back to context

Comment by simonw

2 days ago

A challenge with this kind of study is that coding agents (Claude Code, OpenAI Codex) only started working really well in late November, which for most people meant early January due to the December break.

General agents (OpenClaw, Anthropic Copilot, ChatGPT "Work") started working even later than that.

This category of software may have a much more meaningful impact on work than the mostly-chat systems we were using from 2022-2025.

Studies that mainly focus on 2022 to end of 2025 might be missing out on a material uptick in capabilities.

Anecdotally, I (along with the rest of my team) got laid off from a big tech company at the end of January. The stated reason was "AI", although I think as we all know the real reason was "we overhired in 2022". I managed to get a job by April without putting in a lot of work, though (mostly just listened to recruiters until something hit my fancy). My work experience is good, but I don't know if it's so good that I would be immune to market effects. The company that hired me is still hiring other software developers, and I was told that it took them a long time to find someone with my qualifications. (We use Claude, although so far I haven't seen a lot of AI psychosis like I have in other places.. that was also a big filter in my job search). Long story short, I think narrative of job loss is way overblown. It's more doom trolling from anthropic and openAI and their mouth pieces as far as I can tell.

  • >Long story short, I think narrative of job loss is way overblown.

    Long story short, I ate dinner so I think the narrative of people over the world being hungry is way overblown.

    • How many studies do we need showing that AI isn't causing massive job loss before you'll believe it? This isn't the first one showing these results.

      2 replies →

  • > I managed to get a job by April without putting in a lot of work

    Congrats! Probably, you have the blessing of luck as well. One of my ex-colleagues got laid off and managed to find an uplevel job within a month, but another has been looking for almost 2 years now. He showed me that most of the jobs he applied to over the past two-ish years are still open, while they rejected him by saying the usual - we are moving forward with the candidates more aligned with the position.

  • It's so early game, is the real problem. My experience so far, has been that it's barely a junior dev. I've met so many in my career that think reading stackoverflow, or watching a youtube video mysteriously makes them an expert. AI reminds me of this sort.

    Well anyone can use (prior to AI) a simple linter and learning to code isn't that big a deal. It's learning the pitfalls, the traps, that's the issue. And so far Opus just seems to fall into them again and again. I guess the best way to put it, is that it's not an architect. I sees no big picture, and that's not really a surprise with (compared to a human) an incredibly small context window. When I'm on a project, or working with a codebase, I often have years of "context window". And I have a career of "don't do this" context window.

    So what I wonder is, will this be resolved? Will that awareness of larger scope be solved? If that happens, we'll be in another ballpark of competency.

    Some companies have massive codebases. Are these companies slowly gaining rot in those codebases, a swiss cheese effect, which eventually will result in collapse? Because I've worked where a bad hire had just this effect over time. And what I worry about isn't using Claude to speed one up, it's the DEV that uses Claude and just "meh" and submits because it passes regression + other tests, and then a manager or code reviewer uses Claude and "meh" because it's a pass too.

    • > My experience so far, has been that it's barely a junior dev.

      At least on greenfield projects, Fable is more like having an endless succession of fly-by senior devs who lock themselves in an office for a week, and who come back with very reasonable code that I then need to maintain somehow.

      For longer-term maintenance, I am actually slowly warming to Sonnet 4.5-era models (so October 2025, right before the Opus revolution). They need to be given clear instructions and watched carefully. But since they force a human to stay in the loop, you don't have the institutional knowledge loss I see in some Opus projects, or the code that was one-shot with no human interaction at all that's a constant temptation in Fable projects.

      And yeah, the bit rot is painfully real once the humans step back too far. I've been dealing with a compelling prototype that someone built, and that stakeholders love (for good reasons). But it had to be put on a tech debt repayment plan for a couple of months.

    • > Are these companies slowly gaining rot in those codebases?

      Yes!

      And that "yes" holds regardless of whether they use AI or not.

      Let's not pretend tech debt accumulation is somehow an AI problem. Some of the world's biggest companies routinely ship code that reeks of years of rot and decay.

      8 replies →

    • I have seen a bad hire having instant impact on team and after person was let go, team was still getting back on track to pre-hire productivity for another 3 months.

      We know all managers that say "stack doesn't matter, we can swap devs". We had C/C++ dev hired to do web dev with C# and Typescript.

      Guy was utterly incompatible even if he was smart but management wanted someone who is local and can come to the office and that was this guy selling point.

> A challenge with this kind of study is that coding agents (Claude Code, OpenAI Codex) only started working really well in

2026? 5? 4? 3?

Heard this one way too many times.

  • I get the point having read much the same from Tesla (and fans) regarding self driving cars that still haven't done half the things that Musk said was just around the corner pending regulators a decade ago and repeatedly since then.

    And myself I keep making comparisons between AI and the progress in 90s video games where every minor improvement got called "photo realistic" and then forgotten with the next game engine: https://archive.org/details/nextgen-issue-26

    So I'm not gonna say "this is it" when the software quality really matters, and I absolutely won't speak to progress (or lack of it) outside of software.

    But I will say "you can look around and easily see small businesses using AI to generate posters, quite a lot of small business software and websites are in the same category: the mistakes are real but increasingly don't matter".

    • >the mistakes are real but increasingly don't matter...

      I think it would start to matter once again. People will get fed up of AI posters and art. I think they already are...and once some threshold is crossed, the business won't dare to use AI generated assets/designs.

      Turns out humans are much better at recognizing patterns in stuff that is generated ONLY using patterns from human generated content.

      8 replies →

    • I digress with your last quote, I feel like some people who are aware of AI are starting to develop a quasi allergic reaction to slop, and while the mistakes might not matter for most of the population, for others it does and will take notice.

    • I’m probably illustrating your point but as a FSD fan it really got ”good enough” recently with version 14. The tipping point was suddenly, much more often than not, it can drive end to end from start (my garage) to finish (parked at destination) with no interventions. I can text and watch videos on my phone and as long as I glance up once a minute, it doesn’t complain.

      Handling highway driving with lane changes was great when it got there years ago, but just in the last year or so it has gone from a nice to have to “from now on I will never buy a car that can’t do this”.

      AI has hit some milestones for replacing work as well. There’s still many more to go and maybe some of them will never get hit (much like I don’t think a coast to coast drive with zero interventions during winter conditions is ever going to happen) but there are points at which it forever meaningfully changes some field of work. I think it’s there for writing code.

      15 replies →

  • I guess it's important who one hears this from.

    I just spoke to a fried who is a headhunter and who's been trying to automate his processes for a while (he likes to fiddle and certainly has skills, but he's not an engineer). He kept trying, but it just wasn't good enough.

    Now he said with GPT Work and Sol, it worked, but the key point is: all of it suddenly worked.

    The problem was one of reliability, of handling edge cases. All previous attempts / model-harness-combinations were too brittle and needed too much observation and fiddling - cheaper to do it yourself.

    Now he says "I don't know why I would ever hire a recruiter [the folks doing the cold outreach] again. I can focus on the candidate screening and acquiring projects, everything else is fully automated".

    This doesn't come from an engineer or an AI lab, but a technically inclined power user, and I think this is where things get interesting.

    • You heard it from someone with no experience developing software. A lot of AI hype comes from that, even (or especially) from people in the actual business of developing software - a surprising amount of managers and adjacent or supporting roles in the field actually have very little clue about software development.

      It's cool that 'regular' people can now create solutions to many small problems, and automate stuff - genuinely a step forward. Like Excel, only vastly better. But for bigger projects, real software engineers know that what LLMs do today is only a tiny part of development. And it solves it in a way that might well make the rest of the lifecycle a lot harder. It's like that saying about tools that make easy things easier and hard things impossible.

      3 replies →

    • > I can focus on the candidate screening and acquiring projects, everything else is fully automated

      would a great candidate get excited about an AI agent reaching out to them? or would it be the desperate or clueless one?

      1 reply →

    • Then it just becomes a new baseline (everyone have access to the same LLMs), and recruiting moves up the philosophical ladder where human can add more value. What will it be? I don't know, I'm not a recruiter.

    • Again I've heard this since 2022 when gpt3.5 came out.

      This is like microprocessors in the 80s. Sure they double in capability every 18 months but the start is so pathetic it will be 30 years before they are good enough for everyday tasks.

      5 replies →

  • It seems to be true this time though; I have observed it myself and heard it from several experienced developers I personally know and respect. It feels like some threshold was crossed with Opus 4.5 and Gpt 5.3, where the models are now able to reliably solve certain classes of problems that were previously unreliable.

    Time will tell of course, and it’s early, but inflection points do exist with progress.

    • Thing is. You can find an extremely similar paragraph written about Claude 4.x or some equivalent gpt. And simultaneously, many people expressing their frustration and the shortcomings of <insert any model>

      “But it’s different this time” - several people, several times over the last couple of years.

      This is not at all a dig at you, I’m very sorry if it reads that way. My point is these things only get truly better in anecdotes. The ways in which they fail is yet to change. Just yesterday I had gpt 5.3 generate completely awful code for the Cinema 4D Python API. Also an anecdote. But for all of the people saying they are truly intelligent and truly reason, they still make obvious mistakes, write around problems, fail entirely at architectural decisions, fail at random, generate FAR too much code.

      And no amount of harnesses, methodologies, loops make much of a difference. If you listen to people on the internet they say it’s all working. You listen to people on the job and they mostly say it’s creating tech debt and a review bottleneck. Also burnout, so much burnout.

      I think LLMs are mediocre. I think it’s fine they’re mediocre. You can work with low expectations. But the hype cycles are so tiresome.

      6 replies →

    • I wonder how much of it is real and how much of it is people just being worn down by the hype to the point they can't fight it anymore

      Very smart people aren't immune to being worn down over time

      28 replies →

  • Nobody was saying coding agents started working in 2023 or 2024, because the category was defined by Claude Code which was first released in February 2025.

    • I would say that Aider is what defined coding agents. That was at least multiple months before Claude code. I remember seeing a coworker use aider for a hackathon project Adeline November-December 2024 , and it was already decent and pretty close to the DX we consider coding agents to have

      2 replies →

  • Since so many people are doubting you here, here is a post from ~a year ago that's pulling the same "LLMs 6 months ago were crap, now they're awesome" shtick: https://fly.io/blog/youre-all-nuts/. There's more posts along this vein being put out from 2024 on or so.

    • I find this attitude baffling.

      Things are allowed to get better more than once!

      The idea that "yeah, you said technology had improved in the past, and now you're saying it has improved again" is a gotcha just seems incoherent to me.

      10 replies →

    • A year ago it was capable but required much more hand-holding. I thought it was as good as it would ever get and was still happy and productive, but only because I didn't know how much better it would get.

  • Yep, the goalposts just keep shifting. In reality: they still don't work well, unless you're content with producing low quality work.

    • Your perspective is one that assumes a single entity is using AI, rather than disparate groups of people working on disparate problem spaces/topics, each with different "intelligence" thresholds for them to say "good enough". The goal post isn't changing, it's that there are multiple problem spaces with completely different, fixed, goalposts.

      And even then, within a single group, you'll have multiple thresholds, of "this really helps make my coding more productive" to "I no longer type code, just review" to eventually "I'm no longer employed".

    • "unless you're content with producing low quality work." - With the right guiding hand, it is a productivity multiplier without compromising quality. As a fully autonomous developer, it is a disaster.

      17 replies →

    • Which means good enough for all those companies that outsource their IT, especially for offshoring.

      What they care about isn't software delivery, is physical goods or services that aren't related to software, for them software is a cost center.

    • Here's the rub. They produce low quality work according to your rubric. If that was the universal rubric, they would already have RL'd against it, and you'd like the work they produce.

      3 replies →

    • What kind of work are you doing and what do you consider to be quality or not?

      Of course don’t let me assume, maybe you have a higher quality disproof for the Jacobian conjecture you could share with the class.

  • The timescale is well established: Late '25 was the start of agentic ai when capabilities of model + api + scaffold reached autonomous state. Any study comapring events before that timeframe is comparing apples with oranges.

  • Claude 4.5 was it (nov 2025?), without a doubt. It went from frequent hallucinations to highly usable with much less garbage output. If you were making demos of AI tools around this time your demo/pitch/product was saved and you probably looked like a genius.

  • I haven't. Around the start of 2026 is pretty widely mentioned as when they went from "this is broken slop" to "huh this is actually 90% what I would have written", which matches my experience.

  • I’ve been feeling gaslit about this too. Getting major “we’re still early!” crypto bro vibes from this constant goalpost moving.

    • The „you will be left behind if you don’t fully embrace the whole thing right now“ is a 1:1 match with cryptocurrency hype

    • The capex is still "early" for sure (i.e. data centers are still being planned and built out, we're hardware/energy constrained).

      If model scaling holds out, we're "early-ish" in terms of the reliability and performance of these systems, just based on utilization of the compute from the planned capex. If we hit hard diminishing returns and we don't find architectural/data workarounds, that would put a wrinkle in things, but I suspect that the AI we have now is capable of helping us find those workarounds and keep things moving.

      2 replies →

Anecdotally I'm seeing a lot more recruiter activity/interest now than this time last year.

But it seems more correlated with hype-cycle-stage than anything else. Right now a lot of founders seem to be convincing a lot of VCs that they can make $LOTS by replacing/changing $BIG_INDUSTRY/$BIG_PRODUCT with an agent-first blah blah replacement, and then using that money to hire more people to manage/execute/coordinate the coding agents...

Last year, by comparison, there seemed to be a mood of "software will stay the same but will require less people" while right now there's a lot of hype around "we can build different types of software or build it in different ways" and those early-stage things are in growth-mode. That guarantees nothing about how many people they'd need in the future, or their success at all, ofc.

The news that I'm getting from contacts in non-startup-land is a bit different - still layoff threats. Still pressure to use AI tools more. Mixed confidence on whether or not longer-running "agent" modes are that much more effective-without-breaking-things in legacy code if not used with care.

  • Is the recruiter activity you see very… Human? I’m getting a lot of emails but when I do my research, it all appears to be bots. I don’t trust any of it and I don’t respond.

> General agents (OpenClaw, ...) started working even later than that.

There is nothing special about the OpenClaw thing other than the enormous astroturfing campaign that benefited various "crypto" influencers and other scammers. Anyone who endorses it is either manipulated or trying to manipulate you.

  • Indeed. That's why I pointed out that OpenAI and Anthropic have entire own provides that serve the same purpose now.

However, companies have been using AI as an excuse for layoffs since well before January 2026, which corroborates the study's conclusion. (Source: https://layoffs.fyi/ai-layoffs/) There is certainly an uptick starting 2026, but that could be explained either by AI actually causing more layoffs, or by AI becoming an even better excuse for layoffs.

This isn’t the first study showing this though. It’s pretty simple, programmers and IT were severely overhired during the pandemic, there are massive job losses now, and it’s easy to blame AI when in reality there are a lot of economic factors and AI isn’t increasing productivity as much as anyone would think.

Maybe the future will change that for very specific things, but I think people should be learning and preparing for that, which isn’t any different than what everyone has been told in every job market since the start of the Industrial Revolution.

  • I personally hope that AI continues not to result in a noticeable negative impact on employment and that pandemic over-hiring turns out to be the major factor for all of the layoffs.

    I'm nervous that the studies which show that so far don't seem to be taking the 2026 improvements in coding and general agents into account.

  • >programmers and IT were severely overhired during the pandemic

    1) I hesitate to believe that losses were disproportionately technical roles as opposed to administrative.

    2) Over-hired by what metric? It's well known that hiring never fully recovered after the GFC; was the recruitment post-pandemic just bringing us to parity with where we had been 20 years earlier?

    Not to say that I disagree with your following point. The AI overspending and the layoff cost-cutting are not in a direct causal relationship; both are rather symptoms of a common corporate pathology.

    • They were hired at a higher rate than anytime since the 90s. In the US they ballooned by almost 50%. So when you take layoffs in consideration, the US is still has 25% more programmers than it did prior to the pandemic. And looking at the hiring rates and who is getting hired in programming, it looks like companies aren’t hiring entry level basically at all at this point, so shows the market is still over saturated. No wonder coding camps and CS majors are not being pushed any longer because it’s not a great market for them.

And this is why the labs cannot just "stop training and become profitable". I can't imagine they would like it if studies like this will actually become credible.

"Move fast like a blur so people can't see that you have no clothes"

Well, in addition to that you also have to consider that companies are now paying (increased) API pricing. It would be an understatement to say that my clients in the F100 range are skeptical at best regarding the gains they’ve seen compared to the costs.

This has led to many of them instilling dollar limits or demanding proof of increased productivity (not just output) with the implication being if you don’t provide value with it it’s getting taken away.

So that is to say, if they aren’t happy with the price now, how will they feel when it goes up again compared to just keeping a certain headcount?

  • >This has led to many of them instilling dollar limits or demanding proof of increased productivity (not just output)

    They should have done that from the beginning - demanding proof of increased productivity - if that was their goal. otherwise they were not using their brains well enough.

    And you doubly don't want to work with them, first because they confused output with productivity at first. and second, because they're parroting the productivity metric.

    You only need one guess for whose pockets the productivity benefits go into.

    10 . 9 . 8 . 7 . 6 ...

  •   > So that is to say, if they aren’t happy with the price now, how will they feel when it goes up again compared to just keeping a certain headcount?
    

    that got me thinking: how are companies expensing ai costs? as personnel expenses or r&d etc?

    • Oh wow. You just hit something there. There are government programs for R&D where I live. Grants. Tax credits. This sort of thing. I wonder if people are trying to expense to those programs too.

  • When I was a teenager, a friend had an old gas guzzler from the 70s. I live in a rural area. One time, my car broken, he drove to pick me up to go to College.

    This cost him an extra $40, in today's dollars. No, I'm not joking. That thing ate gas like a dry camel drinks water.

    This is what Fabel5 feels like. Crazy expensive. 10 minutes work pulled almost $80 is usage credits yesterday. I'd be exceptionally skeptical too, on costs, if I still had the DEV I had last week, but they were also eating that kind of cash on a very-improved, but still used as a linter.

    For $200+/hr, or ~$400k/year, I'd want to see a tripling of output at least. In a lot of US markets, you can hire 3 junior devs for that.

    Yes, there are cheaper options. Opus, etc. But it's really over-priced, and frankly I think the real gold now is making open models fully functional. Anyone predicating their business upon tie-in with the big boys is just going to fail, hard.

> This category of software may have a much more meaningful impact on work than the mostly-chat systems we were using from 2022-2025.

Mostly an impact on software development - I'm not seeing broad automation and layoffs in industries like law, finance etc. It will gradually happen but due to issues with memory, reliability and long term planning of LLMs there are real barriers. Even in software development - while it has completely transformed the field I don't think many people still believe we won't need devs in 2027 or that their amount will shrink by 50%.

  • I'm optimistic that coding agents will not obsolete software developers - my own experience is that coding agents have made my work harder, because the scope of projects I can take on has increased.

    My ideal version of all of this is that nobody loses their job and everyone gets to take on more ambitious projects.

    I'm not quite naïve enough to assume I'm right about that though!

    • > my own experience is that coding agents have made my work harder, because the scope of projects I can take on has increased

      Likewise. I find the types of problems I need to tackle and the challenges they represent are actually quite exhausting, too. And the agents unblock me relentlessly so I’m constantly pressed to do relatively difficult things. Either that or code review. I’m hoping it’ll only be an adjustment phase but this is the most challenging my career has been in over a decade.

  • If by 50% you mean all juniors/mid-level engineers will not find work, then I agree.

  • I think other industries are late, but also will see less productivity gains. In software engineering agents have the benefit of very short feedback loops that provide extremely actionable insights which allows the agent to very quickly iterate towards the final solution. This is much less the case in other industries such as law and finance I think. The first pass of the agentic loop is usually something that barely works (you don't even get to see it with the latest models). If that first pass is the primary output in other fields with no tight feedback loops, then the productivity boost will be much lower than in software.

It does feel like we are in a transition period, and its not clear what conclusions we can draw about any sort of "steady state" yet.

For whatever anecdotal evidence it's worth - that has been my experience as well. For the first time last December I noticed the harnesses performing like actual workers.

Everything impressive has happened in the last six months.

Any research on the impact of AI would have been lagging indicators. It's not the researchers' fault[0], but the nature of a field moving at neck breaking speed. Remember there was a paper saying programmers were 20% slower with AI?

[0]: well...

  • That early 2025 METR study was particularly interesting because participants self-evaluated themselves as 20% faster, but the measurements showed they were actually 19% slower.

    All the reports of productivity since then are self-reported, or using questionable measures such as SLOC and PRs, so it’s reasonable to say that productivity improvements are still unknown.

    Unfortunately, METR hasn’t been able to replicate the study because they couldn’t find enough willing participants.

  • > Remember there was a paper saying programmers were 20% slower with AI?

    Key thing was not that they were X% slower, but that they were slower while being convinced they are faster. Of course, any analogies with the current hype cycle are completely unfounded.

This. Also I'm finding out that, after being blown away by agent mode lately, non-agent mode still kind of sucks across frontier models. Using GPT and Gemini in non-agent mode is asking for inaccurate information confidently presented as the truth. Turning on agent mode fixed a lot of that for me.

  • I think it's because of the lack of feedback. Humans also can't do much without feedback. E.g. I doubt most people could write 100 lines of code that works first time without even compiling it once.

    • Disagree. When tool limitations meant that this was the way people had to work, many people could do this.

      It was much more inefficient, because it's easier to find bugs after compiling or running the code. But it is perfectly possible.

      2 replies →

    • Indeed. I wonder how much AI will do in industries with longer feedback loops that are much noisier. I think the field of software is a unique field in that sense.

> coding agents (Claude Code, OpenAI Codex) only started working really well in late November,

sigh

reset the clock everyone!

And yet I have felt a general code quality decrease and overall enshittification of software products since 2023 when people were already using copilot. Or I am just biased to use that to justify any overengineered piece of shit code with that because I refuse to think any sane person would come up with such contrived code and I am looking at the wrong places.

Ah, Simon with another load of bullshit! Your own blog articles contradict your comment. But sure, keep moving the goalpost for your overlords Anthropic and OpenAI.