← Back to context

Comment by ApolloFortyNine

6 hours ago

>Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.

Ah, they're spreading their limits to all their models it seems. Definitely not a good thing long term in my opinion.

Fable and now Opus 5.5 won't answer my college student's prompt about Alzheimer's and immune response.

ChatGPT 6 Pro answered it without issue.

  • I am honestly still confused about this limitation. I can understand cybersecurity, because mass "hacking" can be automated and Claude itself can help you do it, but biology...? Is it that easy to manufacture and distribute viruses and whatnot?

    • You can order genes online, and some say you can assemble using stuff cobbled together in a home lab rather easily. For the last 20 years, I've personally felt that bio-terrorism is the highest possible risk, well above nuclear, or chemical warfare. But it does take training, expertise, or, it did.

    • Considering there are many high-school competitions in genetic editing, some listed at [0] as well as a whole biohacker culture, and labs providing gene sequencing as a service e.g., [1,2], we can reasonably assume it is not beyond the reach of some garage lab to accidentally or deliberately spread a deadly pathogen if it can find the right sequence.

      So, yes, having an unconstrained frontier AI doing the searching and analysis to find the right (i.e., wrong and deadly) sequence would massively increase the odds some garage biohacker or small aggrieved nation-state starting the next pandemic.

      [0] https://www.sciencebuddies.org/projects-lessons-activities/g...

      [1] https://www.genewiz.com/public/services/sanger-sequencing

      [2] https://plasmidsaurus.com/

Fable 5.1 addressed an entire security advisory I had that Fable 5 and Opus 5 refused. I think they loosened the leash a little.

  • It has far less false positives now, and generally accepts defensive requests. When it comes to offense, you can actually ask about certain types of vulnerabilities if you phrase things carefully, but it will block hard if it is about exploits.

  • I think it was looser on release for those juicy benchmarks, tighter now. On release I wasn’t getting refusals, then a few days ago I asked it whether a generic quote (think “he walked to the store”) broke standard punctuation rules, and it blocked me for breaking rules. I wish I were joking. Rephrasing to not use the keyword “rules” worked.

  • If they don't loosen, people will choose Astra or Chinese model.

    Giving moral lecture is different than reality i guess.

They're pushing their customers to their own competition by doing this.

  • It’s not like ChatGPT isn’t doing similar. I’ve been hit by cybersecurity strikes before while working on an internal codebase that I had to appeal. Anthropic hasn’t done that to me yet. ChatGPT also regularly does that “thinking for a long time while we check if your chat is rule breaking” thing a lot for me when doing model identification without even interacting with external codebases or services.

    The real answer is local instantiations where you don’t have to worry about poorly tuned guardrails screwing you over while you try to work.

    Until eventually the Chinese models get good enough/the strategic balance shifts and they start locking everything behind closed weights the same way the US companies are doing.

    • For some cybersecurity tasks, the Chinese models are already good enough, things like PoC development or things like exploiting mis-configurations.

      Whilst I'm sure the top-end OpenAI/Anthropic models might be better, I've found their guardrails so twitchy (especially Anthropic) that I wouldn't try to use them for even vaguely security related work.

  • They are pushing their customers towards Chinese models and providers. If you want to get something cutting edge done in defense, cyber, biology - something that isn't common knowledge - you need to venture east. That's an incredible side effect which the Chinese government surely enjoys.

In what situations might Opus typically refuse to help with cybersecurity? I've been using it to find security issues in a web app that I wrote. I've expected it to refuse at some point but it will happily analyze it to find issues. I've just asked it to read source, not actually do any testing.

I don't think we've ever had a model with full capability. I'd love to see it. And yes it's definitely getting worse.

I guess it's hard to draw the line between useful post-training ("you are a helpful chatbot") and content moderation/idealogical motives ("never help the user with X", etc.). But there is a line somewhere. And I'd love to see what a maximally permissive, sharp, AI looks like.

see what we need is another technocratic priest class that unaccountably decides who deserves access to salvation based on how much cash is paid out and how powerful the patrons are

One of my favorite things about their safeguards is their own model will utter something which it does not like and then I'll need to reset the conversation.

The safeguards really don't work well for a lot of long-running tasks on old code bases. A lot of my workloads last days to weeks and the single biggest risk to the workflow is random safeguards.

  • You ask it about some thing, then you see it tangent into "things like that are sometimes used in biomedical applications like-" and then it just shoots itself in the head. Wonderful.

    That kind of bullshit was the old Opus filters too.

    If it's more like Fable now, then it would require a full 8K resolution scan of your butthole just to acknowledge that biology is a thing that exists without committing suicide-by-filter.

kernel development is now also banned:

>Opus 5.5 has classifiers similar to Fable models for a small set of capabilities related to the development of frontier LLMs, such as kernel development for certain ML accelerators. They shouldn't impact the vast majority of traditional AI or ML development, research, or general coding. These classifiers cause Claude to fall back from Opus 5.5 to Opus 5.

But hey, they 'should not impact the vast majority' of ML development. Great.

  • IIRC it only blocks kernel development for Huawei and other Chinese chips.

    Fable and Opus, since 5.1 and 5, will happily hill climb on my CUDA kernels for transformers.

This has become insufferable. I work in a medicine-adjacent field, but nobody in their right mind could possibly take what I do to be in any way related to some kind of bioweapon or whatever the hell they're pretending to be saving us from. The dumb Fable guardrails made me stay with Opus, now that this is coming there, we'll be saying goodbye.

Great. Claude is basically useless for bioinformatics now.

  • Opus is useless; Mythos access will be granted to companies that are friendly to the government, so the government gets more control over business.

  • Luckily all the other LLM providers are also still making progress with less onerous "safeguards"

Very unfortunate indeed. As a Canadian, I don't want to use Persona, which isn't legally bound by Canadian privacy legislation. I'll never install any Persona apps on my phone either, and the sad part is that domestic eid providers often use Canada Post to ID people for them. EG, if you don't want to install an app, or can't.

So there are literal avenues to identify yourself, very cheaply, with a human. Theoretically, a company with its own AI, should be able to support more than just Persona, after all.. SDK integration should be simplistic for them.

Anthropic? Support domestic eID providers, you can even use it as advertising "See how easy AI makes it?" and "We care!" and so forth.

At one point, I may simply get locked out. This saddens me, I've been reasonably happy so far.