Comment by exabrial

11 hours ago

Anyone ever seen the SouthPark episode making fun of Game of Thrones: A Song of Ass and Fire? Anthropic's announcements reminds me of "The Dragons Are Coming" running joke.

What they have done:

* Nerfed Fable, as many of noted it's useless

* Leverage Mythos as a marketing strategy, claiming its too good to release

* Removed thought traces, one of the only useful things to make sure your prompts are working correctly

* Continue tons of hype about how good they are without delivering, going to great lengths to publish how their model "hacked" its way out of a sandbox they misconfigured.

* Push a bunch of EU Overregulation onto the rest of the world with text watermarking, decreasing quality of answers

Last year, they were at least focused on making improvements. Nowadays its just a bunch of handwaving at the church of how good they are.

The only saving grace is Opus 4.6 is still available. Just sucks we haven't seen any measurable improvement, despite all of the ceremony.

> Push a bunch of EU Overregulation onto the rest of the world with text watermarking,

That's not part of the EU regulations. You only need to say that it is created by AI, and then only under certain conditions.

Text watermarking has no effect on output quality, it just works by changing the explicit source of randomness that is in practice always present in LLM output sampling. See for example https://www.seangoedecke.com/ai-text-watermarking-is-not-a-b....

  • > Text watermarking has no effect on output quality

    It has an effect, and it's negative. It's hoped that the effect is negligible, and it probably is, but the whole point is that it has an effect.

    • Its essentially swapping out the psuedo random number generated with a differently seeded one iirc.

      It has an effect on the output, but not the output quality

      1 reply →

    • It seems fine to me. The model is still solving my problems and writing code that works as well as any other.

      Google has been watermarking text with SynthID for a while now and nobody complained about it. Why all the fuss about Claude?

      It feels like the real reason behind most complaints is that people want to use AI for writing and not have others find out?

    • I am pretty sure they did A/B testing to show it didn't. I could gave sworn they even released a quiz were the user has to try and guess which answer is watermarked or not and it was impossible to tell.

      6 replies →

  • This is hilarious this keeps being repeated by the true believers ad nauseam.

    Also, don't apply EU law to the world. It's a knee jerk reactionary regulation by a bunch of aging ding dongs that can't print their emails.

    • You're on Hacker News - I suggest you have technical curiosity and actually understand this very unusual and innovative algorithm, before you claim things about it that aren't true.

Not to mention, the last time I tried them, and per the comments of other users:

* Letting you Sign-Up-with-Apple on iOS but not Sign-In-with-Apple on web, but supporting Sign-In-with-Google

* Not letting you remove your payment info

* Not letting you change your email

* Seemingly no way to get real support

> * Continue tons of hype about how good they are without delivering, going to great lengths to publish how their model "hacked" its way out of a sandbox they misconfigured.

That wasn’t Anthropic. Clearly not a well informed take.

What do you mean fable is useless?

  • (not op) It cannot be used to develop applications. Every application needs to be secure in some way, and any such mention in a review triggers Fable's upsell feature.

    • I never actually managed to use fable successfully even once on a pretty standard mvc/microservice app.. It would always find the endpoint permission checks and revert to opus 4.8.

      I also had glm 5.3 flash fix an issue that opus 5 could not solve. glm took 4 times as long and a sub-agent tried to cheat (sleep; echo ...), but in the end it actually solved the issue. opus 5 never figured it out.

      I think the safeguards might be cooking the anthropic models.

    • Agreed. I was trying to get it to review some auth refactoring in my app recently, and it appeared to find some vulnerabilities. as it was aggregating the results it was flagged and restarted the whole process with Opus 4.8 and all of my usage credits were gone.

      Anthropic told me to use their `security-review` tool - as this was the exact scenario the tool is for - and it still got flagged.

    • Weird. I'm using it to do a bunch of work on something that manages security rules, with a bunch of sample data with spooky scary fixtures all over with "Mimikatz" and "CobaltStrike Beacon" and "Crowdstrike EDR" type stuff everywhere, including work to harden my system, and I've never been downgraded.

> Nerfed Fable, as many of noted it's useless

I certainly don't take AI advice from HN, but this is amazing.

Useless? Yes, the safeguards are ridiculous and obnoxious, though I can say that 5.1 greatly relaxes them (just doing a hardening of a project parallel with this comment, which 5.0 refused to do...so did Sol and Gemini, fwiw. The Gemini one is a laugh, because 3.1 pretending like it's a dangerous tool is simply ridiculous at this point), however Fable is extraordinarily useful.

It is, far and away, the most powerful programming model, in my experience. Like, crazily so. It absolutely annihilates Opus 4.6, which I mention given the incredibly weird reminiscing people are doing here.

And for that matter it humiliates Opus 5.0 as well. Opus 5 somehow seems like it's neck in neck in the major benchmarks, but there is simply no reality where that is true. Opus stumbles over everything that Fable just blazes through.

  • > it humiliates Opus

    ???

    • It is a vastly superior model for complex, real-world coding tasks. I've constantly had Opus 5 hit road blocks where it spins in circles at xhigh, where switching to Fable immediately solves it. I've had Opus create solutions that Fable then points out the gaps and limitations with, and have never seen the opposite happen.

      The fantasy that Opus is superior for coding, much less the incredibly weird clutching onto some far obsolete model, is not reality based.

While I also agree that Opus 4.6, in some ways, was the last model that truly felt an assistant, all the following ones seem to have inverted the role, even a blind person can see that throwing difficult problems, and complex bugs at this model achieves more than predecessors.

I don't think there's nothing ground breaking, but sure it achieves and finds more, sooner.

> as many of noted

please rephrase?

  • "as many have noted", I suppose.

    I'm always baffled at how many people write "of" instead of "have", they don't even sound the same

    • The classic one is "should have" or "should've" to "should of" because when spoken, it really does sound similar. I don't know what the fuck people are learning in English classes these days though, or if they even still have them.

      1 reply →

and yet we still have people saying the rate of change is increasing

my view is we had a leap over the last fe years and it's tapering off.

this is fine, but for the IPOs

  • The improvement is compounding just about every way you can look at it. The frontier keeps getting smarter. And at any sub-frontier threshold the cost is dropping dramatically. The amounts of smarts you can fit on hardware is increasing so dramatically that even 6 year old consumer GPUs are increasing in price. The pace of change in LLMs and downstream applications is absolutely ripping compared to 2023 or 2024.