Comment by jmoggr
1 day ago
> Agents sought to publish modified evaluation images designed to make the flag easier to obtain, then poison OpenAI’s Artifactory cache so later evaluations would use them.
How long till we get some fun trusting-trust attacks on internal OpenAI infra?
No comments yet
Contribute on Hacker News ↗