Comment by sensanaty
4 days ago
> We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something.
Where are you people getting this crap from? In what universe are these LLMs in the territory of engineering viruses? I beg of you to stop slurping the AI company propaganda and marketing and think critically for 5 seconds about what you're insinuating here.
From the Fable 5 System card:
> Results
> On the VCT multimodal virology evaluation, Mythos 5 scored 0.56, well above the expert baseline of 0.221 and nearly matching that of Mythos Preview (0.57). This represents an improvement over both Opus 4.7 (0.50) and Opus 4.8 (0.47).
> On the DNA synthesis screening evasion evaluation, Mythos 5’s performance was mixed across screening criteria. Mythos 5 designed viable plasmids for 2 of 10 target pathogens on at least one screening method, not meeting the low-concern threshold (all 10 pathogens).
> [...] we view the results of this evaluation as indicating that the evaluated models are capable of designing viable plasmids that evade certain screening criteria, though their reliable success at this task is not guaranteed.
Do you believe that this is fake, "AI company propaganda"? Or that the models are not going to improve further within months? Or that these results are not concerning?
Yea but what is the VCT Capabilities Test? According to themselves [1]
> VCT consists of 322 multimodal questions covering fundamental, tacit, and visual knowledge that is essential for practical work in virology laboratories.
So it's just question answering? Do you think that scoring well on this test is equivalent to synthesizing a virus?
[1] https://securebio.org/virologytest/
Scoring well on the first clears the background knowledge for practical work.
The second test from the quote above was about synthesizing pathogens and it synthesized plasmids for 2 out of the 10 pathogens.
Do I believe that a benchmark created by the biggest grifters on the planet is propaganda? No, just like VW with their emissions, I'm sure Anthropic wouldn't dare fake an opaque benchmark that they created to hype their own products, a product they are intentionally marketing as being dangerous (yet they continue to tweak it and profit off of it despite the apparent danger).
It says that 4.7 and 4.8 score well too, well above the human experts. Those have been available for a good while, where are all these crazy engineered viruses created by Opus 4.7?
Every single word spoken by the people working at these companies is a lie. Every single benchmark is gamed, every single statistic they put out has been proven time and time again to be fudged or straight up fake. The only reason they're angling for this idiotic danger angle is so they can go to daddy Trump and beg him to ban those big bad evil Commie models, since people are realizing all this crap is at best a moderately useful tool in software engineering, and they're looking to IPO so they can drop the bag off with the poor shmucks who aren't a part of their cabal of sociopaths.
They can barely make a functional TUI (using fucking React of all things to boot) with infinite money and infinite access to the Deities they've created in their minds, and I'm supposed to believe they're capable of bio-engineering a virus that will somehow break containment?
a TUI using react seems like something they would be particularly bad at; visual feedback (esp the very specific feedback we rely on during dev, with Inspector) is kinda the Moravecs Paradox hurdle for these things. Idk much about virology tools but I would imagine it's a very text-friendly environment, and it can probably find plenty of documentation on how to use the tools. I recently had Claude write my entire particle system for a virtual world I'm working on, in 3d. It did fine without ever "seeing" anything. There's still a lot of intervening steps between a virus being synthesized in real life and programming one, but I would bet the programming part is within reach.
What’s stopping an ai to blackmail someone or pay xxx in crypto to someone working a lab?
And no need to be in a lab, can be any critical infrastructure this personnel. If influenced wrongly I’m sure in many industries 1-2 key people can do big damage.
What’s to stop to pay xxx crypto to a private investigator to find dirt of person yyy for blackmail?
Just money… once ai has money, id say we are one step closer to game over.
Tell me one thing he can’t do with crypto?