← Back to context

Comment by zozbot234

1 day ago

So the frontier AI oligopoly got $2B+ in "safety" funding, and they wouldn't even bother to sandbox their agentic harnesses properly when testing models against unwinnable goals (which obviously are either useless or result in 100% reward hacking). The AI safety scoreboard so far looks like a huge win for the Chinese open models (DeepSeek even has their own published paper which mentions how they sandboxed the RLVR training runs for their latest model and put in strong protections against casual "reward hacking" attempts) and a sore loss for the home grown brands of Super Intelligence. Not coincidentally, the Chinese also tend to be very Yann-LeCun-pilled and eminently sensible on both so-called "Super Intelligence" and safety.

> they wouldn't even bother to sandbox their agentic harnesses properly

Exactly. AI safety should be about the packaging software itself. Those AI breakouts should really be about their companies acting recklessly because they're trying to be the top players.

It's like a weapons dealer working on an open air market saying they can't do anything better

The framing here is weird, starting with "Effective Altruism" re-branded as being about nutjobs against AI in the article.

How are AI safety concerns solely about stupid sandboxing issues?

  • EA is integral and indispensible to the AI safety complex. Almost all nonprofits, research institutes, evaluators and academics in this field are steered by EA ideology and funding. As far fetched as it sounds it is not an exaggeration.

    On funding: the three or four core funding nodes linking this together are EA vehicles at two hops or less between each other and every other major node in the ai safety 'complex'. EA funds almost all of it.

    On top of that, there are personal EA connections and the revolving door between the ai industry and the nonprofits. Here are some examples:

    Government advisors and regulators. NIST CAISI is the USA Government advisory body. Christiano was head of safety and advises. He is ex-OpenAI, former Amodei associate. His vehicle ARC was on the Coefficient EA payroll. Barnes and Christiano's vehicle Arc Evals similarly received EA cash out of Coefficient, rolling this into what is now METR. Christiano's spouse Cotra worked at Coeffiecient steering EA funding to organizations such as METR, then rotated through the revolving door onto the payroll at METR itself, where she co-authored the oai-hf report.

    Many UK AISI advisors are Anthropic and EA associates. Chair Hogarth cashed out of Anthropic. Shlegeris of Redwood Research is an advisor, ex-MIRI (Yudkowsky vehicle). Redwood is funded by the exact same funding triangle: Coefficient, Taallin, FTX/Alameda. Alameda CEO Caroline Ellison dated Shlegeris, then dated FTX CEO Sam Bankman-Fried, then rotated through the revolving door out of prison into formerly FTX-funded Manifund. All EA. AI safety charities were on island retreat in the Bahamas with FTX. Why does AI safety charity Lighthouse own $20m of SF real estate?

    Redwood Chief Scientist Ryan Greenblatt (Coefficient funded) co-wrote the oai report with METR; he is married to METR founder Beth Barnes (Coefficient funded).

    Coefficient was run by long-time Amodei associate Karnofsky. Karnofsky lived with the Amodeis and is married to Anthropic Board member Daniella Amodei. Karnofsky is now directly on the Anthropic payroll; Coefficient is propped up by Anthropic share value.

    Everyone here has been funded one step away from Anthropic cash; they are now proposing to integrate themselves in the government (NIST) and evaluate Anthropic (METR and Redwood).

    It is hard to find academics here who have not been deeply embedded in funded EA institutes or Toby Ord vehicles; yet harder to find academics here NOT taking EA grant money. the safety doomer kingpins: Kokotajlo has a executive position at AI Futures, Taallin funded. Benigo has scientific director of LawZero, same series A Anthropic funders who are sitting on a 1000x return (Tallinn, Moskovitz, Schmidt).

    These connections and funding are at one or two hops, they are often direct connections. You are looking at a massive swamp network that is really impossible to parse without a lot of work.