Comment by PedroBatista

13 hours ago

This post and comment makes me believe "science" is the new "code" for Anthropic now that the code advantage is mostly gone and lost for OpenAI, ie. they got much better and Claude become significantly worse over these months.

I write a lot of Rust and Lean, Fable 5 is in my experience better at both. Cost/performance is a different story.

  • Fable has become my go to in Agda as well. It just crunches hard technical tasks!

    I find Fable 5 still lacking in library design. But I guess there is no accounting for taste…

  • Yes, both of which are domains for which a verifier is readily available.

    You can generalize from them to "science".

  • This is really it imo. Fable 5 is better then Sol. But Fable is just of the table for anything even remotely long running. Unless you have very deep pockets. And the difference between Fable and Sol is not world shattering if you ask me. I also find codex a ton better than claude.

IMO, Codex is worse than Claude with Fable. At least at Rust.

That said, the open source models are not bad and I'm looking forward to more tools and products built on top of them. Code review, security review, etc.

Anthropic needs to change how it treats users though. I'm increasingly put off by Dario, the rug pulling, the lies, and the attempts to regulate open weights. I'm going to bail if this doesn't change. There's plenty enough that's good enough, and those things are hackable and extensible.

If Fable isn't available at subscription price via third party harnesses soon, I'm also going to bail.

  • The big issue I have with Fable is this. From the Anthropic email announcing Fable 5.1. So basically they're giving us a Ferrari, which will point blank refuse to do certain stuff - forcing us to go out in our Mustang. Their choice, not ours

    "Safeguards and automatic fallbacks (beta): Fable 5.1’s biology and cybersecurity classifiers block fewer benign requests and now permit vulnerability finding in source code. Blocked requests return an error and are not charged to you. On the Messages API, opt in to fall back to another model so users get a response instead of an error. We recommend Opus 5 for biology and Opus 4.8 for cybersecurity. In Managed Agents, fallback is built in."

  • Its "pure capabilities" are definitely worse than Fable, but I find codex has a much more pleasant style and is in comparison much more generous with its limits.

  • > IMO, Codex is worse than Claude with Fable.

    Fable easily trips its safe guards. You can be 95% complete with the plan for it to trip and then lose it all. Anything is better than nothing.

    • >> Fable easily trips its safe guards.

      Maybe it depends on the type of work you do, because for me it almost never happens.

      >> You can be 95% complete with the plan for it to trip and then lose it all.

      That's... not what happens though. The session will either seamlessly downgrade to another model mid-session, or it will stop with an alert and you can just re-prompt it. It will still have access to the context.

      5 replies →

  • Small reminder that the US government rug-pulled Fable, not Dario. Lots of the safety guards that users find annoying/objectionable were the results of negotiations to get the model back online after the US government forced them to take it down.

    Maybe Dario should have just "donated" $1M to Trump's inauguration fund like Altman, Meta, Amazon, Microsoft, Tim Cook, Elon, and Google. There's a reason they are the odd man out with this current Administration.

    • Maybe Dario shouldn’t have tried for regulatory capture. He was constantly on the news talking about how these models are so dangerous and that we need regulation to keep China from releasing open-source models without guardrails.

    • The US government didn't make the choices to release the worst version of Opus and label it 5.0, and then isolate portions of their subscribers to limited usage of Fable.

      They may have been unfairly targeted by the US government, but they are doing more damage to themselves without government help as well.

      3 replies →