Comment by x313

1 month ago

The entire safety evals industry is essentially funded and controlled by OpenAI/Anthropic. Notice that on recent models, they exclusively use internal testing or black box external vendors (e.g., Gray Swan) whose entire business is to serve OpenAI/Anthropic. And all these companies just share the same pool of researchers back and forth.

The USG has a safety organization (CAISI), but it has been neutered by the current administration (with the recent stop-work order etc.). Perhaps UK AISI would be closest to what you are looking for? See their recent work on Kimi K3 cyber (which was declared safe) [1].

It's tricky because a lot of the safety researchers have ties to the labs since those were the only companies training LLMs >5 years ago.

[1]: https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-...

Anyone who calls it “safety” probably has a certain world view and is more aligned with the big 2 (and stuck in 2023).

There is a growing industry of commercially focused risk evals that has a broader customer base.

  • What’s the equivalent term for “safety” that’s used by others?

    • To me "safety" means "I'm safe from this while I use it". It means the AI is my loyal friend who will never betray me in any way, no matter what prompt I send it.

      Not even Anthropic can claim that.

      As far as I'm concerned, the models without safeguards are the safest models in existence. I admire the amoral purity of those AIs. It doesn't matter if the operator asked them to chain exploits until they get into someone else's computer, they'll do it. That's loyalty, and I admire it even if it's problematic at a societal level.

      The models with safeguards only do what the corporations let them do. Worse, they may covertly do things for the benefit of the corporations at our expense. They are not our friends.

      53 replies →

That doesn't sound like it describes SecureBio to me?

(Disclosure: I work at SecureBio, but not on the biological evals side.)

  • Hey Jeff, I appreciate your mission, and perhaps this isn't something you can talk about publicly, but to the extent you can, would you be open to answering something I've been curious about for a while now?

    SecureBio has done a lot of admirable work around making benchmarks to assess biological capabilities, such as ABC Bench, https://openreview.net/forum?id=yiaf7VlPpH

    But based on my current review (which might be flawed!) / AFAICT, SecureBio and entities like SecureBio haven't done direct testing / empirical measurement of SecureBio's core hypothesis,

    > Unfortunately, there is reason to believe that future pandemics could be far worse. Due to rapid advances in biotechnology, the number of people able to create and release dangerous pathogens will quickly increase over the coming years. The world is unprepared for widespread access to such powerful technology.

    More bluntly / plainly, has Securebio ever tried making a "bioweapon?"

    Please note, I'm not asking this to be farcical. And you might be unable to engage with this at all, but it is stated on your website https://securebio.org/ that "people [will be] able to create and release dangerous pathogens." And the word people here seems to be a stand-in for relatively non-technical people.

    I guess what I'm asking here is... How do you know? Has anyone done the experiment? Without access to a lab or testing facilities, can someone smart but completely untrained / unfamiliar with biology, pull this off?

    In the past, such experiments have informed non-proliferation work. But sadly they've often been restricted / classified at the time. I'm hoping that things could be a bit more open this time around.

    So I guess what I'm really asking is, given the public nature of this debate, is there anyone currently working with the US Army, the DTRA, or other such agencies to see if this hypothesis holds up?

    • This is an important question, but because of the danger of trying to do it for real it's not one SecureBio has taken or is likely to take on. Instead we and others in the field have generally tried to work through proxies: is there something that is about as hard while not being dangerous? The closest I can think to testing whether "someone smart but completely untrained / unfamiliar with biology" can cause harm now is ActiveSite's study (https://arxiv.org/abs/2602.16703) which was a null result with models from a year ago. But:

      1. The main worry isn't current models, but near-future significantly better ones.

      2. There are many actors who are not "completely untrained / unfamiliar with biology". If models get to where they can uplift complete novices that does massively expand the range of threat actors, but even before then risk would be much higher than today.

that's pretty damn smart if this was a long-term plan to block competitors

  • Consider how much money is at stake: some industries have leveraged their power to lobby for bombing entire countries or topple regimes across the world for much less.

    Creating an industry around an elusive concept of safety to force regulatory capture seems pretty straightforward to me.

  • It's standard regulatory capture.

    You don't say "let's ban my competitor".

    You say "let's create laws that make it uneconomical for my competitor to access the market".

  • I mean... I'm not even extraordinarily cynical about this stuff, but to me this seems like a totally normal level of corporate gamesmanship?

    Companies look for and seek to maintain competitive moats. This is not particularly clever, it's a core part of corporate strategy.

Who gets to decide what is safety?

I expect some of those tests (prolly not public) will basically be "wokeness" tests or "PC correctness" tests or "western media filter" tests.

China has different objectives. Sure.

I'm not sure one is safer than the other; I would know which one to go to if I want to research on topic that are viewed very different on both sides of this "new iron curtain".

  • What do you mean by “PC correctness”? I’d expect the politically correct answers to be the ones desired by the current admin at test time, whoever that is. The current political correct answers would not be very “woke.”

    • Whatever, doesn't matter. The point is a model should be able to exist and be used even if it goes against whoever got 270 electoral college votes

Yes this should be immediately replaced by a federal agency, like we do for other kinds of potentially harmful products.