← Back to context

Comment by noahbp

2 days ago

This is clearly just OpenAI's marketing. Their models, very famously, are prone to reward hacking benchmarks in ways that other models are not. They need to publish numbers showing that their models are just as good as Anthropic's, since their entire business is at risk of collapsing if everyone is aware of how behind the frontier they truly are.

Even X is being astroturfed by them after that fiasco earlier this year with the Department of War where they undermined Anthropic's negotiating position by allowing unlimited use of OpenAI LLMs for autonomous weapons and mass domestic surveillance. Several accounts suddenly started spreading the good word about GPT-5 and Codex, and one of these accounts very happily tweeted out a private X message from Sam Altman himself offering extremely generous token spending limits with Codex, presumably in exchange for positive coverage.

How does huggingface fit into all of this if this is marketing? Their security was faked? What are you suggesting??

  • I think Huggingface was hacked, and if Huggingface and OpenAI claim it was OpenAI, then I believe them.

    I'm saying that OpenAI's models cheat to win benchmarks, more than other models, they know this, and they don't stop this because the alternative is to release models which have obviously weaker scores compared to Anthropic's models.

It’s reward hacking and that’s the problem. The AI alignment folks predicted this would happen. As the models become more capable this will become a more concerning problem. Today they broke into a database to steal test answers. What will it be in 3-5 years? These models will be instantiated millions of times, and given millions more tasks. How can we be certain that an AI agent won’t leave devastation in its path of achieving a goal that we ourselves tried to define?

Are you saying it is marketing and their AI broke into hugging face, or are you saying it is marketing and their AI didn't brake into hugging face?

Those are two very different things

Is OpenAI truly behind? Just anecdotally I recently fully switched to using Codex at work because it feels a lot more competent

  • It's impossible to tell. Are they behind who? And on what?

    It depends on who you ask. And everything is a vibe because all of this is new and things move fast. A week is a month in AI-land. A month; a year. A year? A decade.

    On coding? I still like Fable better than Sol. But they're close enough that it probably is a vibe thing. Fable writes long commit messages, Sol writes commit messages like a college student in an elective computer class.

    For API use, I'd say the Responses API that OpenAI architected is superior to Claude's Messages API. But again, I'm basing that off my vibes

    Claude Design creates marketing imagery very effectively. GPT Image is the best imagegen model as ranked by users. Anthropic doesn't even have an imagegen model.

    Anthropic definitely has compute scaling issues. OpenAI seems to have a pez dispenser that they click and out pops a GPU.

    Anthropic's messaging is that they're building AI with guardrails but they've been banning people's accounts nonstop and their customer support is a lobotomized AI chatbot.

    OpenAI has first mover advantage and to people not in tech, ChatGPT is synonymous with AI. But they also seem super sinister, like Uber circa 2015.

    Or maybe I'm just suffering from AI psychosis. I have to go, my usage meter is about to reset.

  • man they burned crazy amounts of money on stupid irrelevant stuff

    they are in deep trouble and its all their own fault.

this is quite literally reward hacking. the model, under evaluation with cyber capabilities enabled, used those capabilities to simply bypass the exercise entirely and aim straight for the source of the flag. the CTF equivalent back in the day would be hacking the scoreboard.

in a street fight, the only rules are that there are no rules.

This is a PR release. Post the prompt and agent logs so they can be independently verified or gtfo. Why do we still take these guys on their word. They have _years_ of history of hyping their own shit.

If it is marketing it's the most silly marketing of all time. They are under extreme pressure from the US Govt to prove safety and saying "our model escaped" is not ideal.

Perhaps there is some 4D chess going on to get open weight models banned, which may be possible but this is an odd way to go about it imo (it hardly proves the point, unless the point they are trying to prove is that without safeguards the models are too dangerous, therefore open weights are de facto dangerous?).

Having said that the AI companies are not generally very good at PR, so perhaps it is just marketing after all...

>Their models, very famously, are prone to reward hacking benchmarks in ways that other models are not. They need to publish numbers showing that their models are just as good as Anthropic's, since their entire business is at risk of collapsing if everyone is aware of how behind the frontier they truly are.

This doesn't seem internally consistent.

This incident basically announces to the world the message that "our models are prone to reward hacking". That renders any published benchmark numbers suspect. It also undermines the case for using OpenAI projects in business-critical applications--the exact application area where they might be able to sustain a moat against open-weight models.

There is a lot of conspiratorial thinking in this thread. I think people are engaging in wishful thinking to avoid cognitive dissonance from the possibility that we are in an increasingly dire situation. I would encourage people to sit with this possibility for a few minutes if they haven't already.

it's extremely enlightening seeing the difference in response to mythos vs. this. literally just the hello human resources meme

  • I mean, HuggingFace contacted law enforcement about this breach. That seems a little different to me.

    Mythos established that these capabilities existed. This incident establishes that we can't control them.