← Back to context

Comment by ninjahawk1

6 hours ago

Once they make a model better than Fable I’ll be switching to Codex. Their priorities in terms of consumers seem to be better. I do think Anthropic has some solid safety viewpoints, but I don’t necessarily think that either is entirely aligned yet with delivering exactly what humanity needs. Maybe the AI will help align the AI companies when it gets smart enough. That’s the real misalignment I’m concerned about.

It feels like 5.6-Sol is already fairly close to Fable, and in some ways exceeds it. Just the other day I had Fable draw up a solution for me, and then I fed it into 5.6-Sol and said how does this look ... it found an oversight and told me about it, and when I then fed that observation back into Claude it acknowledged the miss.

I've noticed also that 5.6-Sol is more concise with output than Fable (and let's not talk about Opus, which is even more wordy).

  • It's common for different models to find holes in another's work. There are various good reasons for that.

    FWIW, we use ChatGPT for our primary model and use Claude to do the reviews. This works better than ChatGPT doing it's own review even with a clean session/context.

    • It's beyond common for a model to find holes in its own work, as well. I have an iterative review as the part of all agentic work, and it always finds something to fix, and will sometimes spend hours fixing its own work.

    • Agree, but the point is not because Fable is better than Sol, it's because it's .. different .. it just looks at the problem through a different angle.

    • Same here. Grok Build 4.6 for me, given how cheap Grok is and how Sol is supposed to be "the" SOTA, it finds a surprising amount of bugs. Most of which Sol agrees with needs to be fixed or improved.

      I've done this tens of times between these two models and it works great in my experience. Sol initial back and forth with me. Commit. Let Grok review. Sol fix. Only then do I start reading the code.

      1 reply →

  • Did you try to say "think more deeply about this problem" to fable after getting your solution, having one model focused on creation then blaming it for not doing proper review when the other model was told to focus sol-ely (pun intended) on review is not a fair apples to apples compaision

    • I mean, then it sits there stewing for 20-30 minutes when you can ask sol and get the same answer in 5.

      Like, the quality of the anthropic models is fine, but they’re so incredibly slow. Claude reads files one at a time while codes dispatches tool calls three or four a time.

  • > Just the other day I had Fable draw up a solution for me, and then I fed it into 5.6-Sol and said how does this look…

    You should be doing this for every solution.

    Even Fable reviewing itself will find issues, unproven assertions, etc. Same for Codex models. A review loop is critical.

  • I used to review each others work, Sol is amazing at review and finding what’s missing.

  • The fair comparison would be to also do the reverse: start with Sol then have Fable clean up. Then compare the Fable-Sol and Sol-Fable outputs side by side.

I don't think these companies have humanity's needs in mind when they're developing these models. Although the last part of your comment struck me as a bit comical, I genuinely believe that an AI can have way more empathy than a corporation. Afterall, a mimicry of empathy is probably better than no empathy.

  • It's a funny comparison. Comparing the empathy of some software to the empathy of a company. It's like saying my car was more empathetic than my school. How can those two objects even be compared is what i am wondering

    • We live in an odd time where 'software' (well neural networks) can be far more empathetic than summed product of a corporation.

      Company empathy does exist, just look at how easy or hard it is to reach a company when you have a problem. How do they try to solve it for you? Is it a brick wall, for example Google when you have a problem. People quite often like dealing with small businesses because they can reach a singular human and have them as an interface to the problems they face now and in the future.

      Agentic loops and the models underneath them can have a simulacra of empathy too. Not every model just blindly agrees with users, and some have a much better depth in picking up context clues that the user on the other end is having a hard time. Businesses just typically aren't running more expensive and fragile systems like that though.

    • Well, companies and AI are both entities that can make decisions and take actions that involve humans. Those might be empathetic or they might not. So of course you can compare their levels of empathy. I don't really understand why you think that you wouldn't be able to.

      For example, health insurance providers are renowned for not being empathetic. Charities are the opposite. Sometimes companies even build it into their identity, e.g. Cards Against Humanity.

      As for AI, I haven't seen a strong difference in empathy but it's definitely true that the big AI companies at least try to make their models moral and empathetic. Even if it mostly ends up just being annoying.

  • It varies, the researchers absolutely do have humanity’s needs in mind, it’s why they founded the companies and are doing work everyday. However the issue is that it’s got so much money involved that the heartless soulless billionaires are getting involved. I think the actual literal people doing the real work are doing it because AI could cure every cancer and every disease, make us a multi-planetary species, outlast humans by millions and millions of years, potentially create actual organic life and make direct upgrades to humans.

    I think that AI has an insane level of upside, it’s just that the greedy dumbfuck billionaires are getting their greedy little grubby paws involved. If left to researchers I think the sky is the limit, but unfortunately thy need assets, so there isn’t a clean solution to that.

    Ideally, we could somehow separate AI from funding from malicious entities like billionaires, but right now that doesn’t seem possible. Hopefully in the future researchers with genuinely good intentions can have far more direct control than dumbass greedy old fucks, but we’ll just have to see.

    I think the future can be bright in theory but we’ll have to see, making insanely powerful open-source models is the direct way to get around the billionaires so I think that’s our only option. Make open-source ASI you can run on a consumer computer.

Fable 5 is just straight up a larger model - I'm guessing at this, but there is plenty of evidence online from people far more plugged in than I am. OpenAI is pursuing a strategy that yields greater operating margins and penetration of their model to developers. Fable's high cost makes it so premium that Anthropic has to reserve it for only the richest customers and corporate users. That's not a winning formula long term.

I believe the reason we have not seen a Fable-level model from OpenAI yet is because doing so would box them in on costs just as harshly as it has boxed in Anthropic. They are letting Anthropic make this mistake.

  • Fable is available for $100 a month. If you're a working developer, you can pay that. I wouldn't really say it's "reserved for the richest customers".

  • Fable is indeed larger than Sol. OpenAI is developing Astra which will be more of a Fable-sized model.

    If you can train a larger model then you can distill smaller models from it. You don't need to necessarily serve the larger model publicly. Distillation is much more effective when you have unrestricted access to the original model.

You are basically saying you will switch from one evil to another because the other seems less evil for now.

It's funny how people make these alignment comments while ignoring how misaligned the leadership at these companies are right form the get go and they just play mental gymnastics to deflect those facts when confronted with them.