Comment by dekhn
7 hours ago
Yes, I believe somebody pointed out to him (maybe here?) the deep flaws and he attempted to strengthen it. In fact, it looks like the changes were done by Claude (!!!).
The old text:
"""Suppose an advanced AI is prompted to “find a cure for cancer that passes a stage 3 clinical trial”. After a large amount of compute, it produces a cocktail of previously unknown chemicals which its mathematical model predicts, when mixed and injected into a patient, will kill all their cancer cells. While nobody truly knows how this cocktail was found, this model prediction is confirmed in Lean, and the cocktail indeed passes a stage 3 trial. Could the AI solution to the prompt be somehow misaligned by exploiting a weakness in the trial process? Before injecting this cocktail into your bloodstream, would you want to know that there is at least one human cancer expert who understands the mechanism behind this cure? Or a human mathematician who understands the mathematical model used to locate the cocktail?"""
The new text:
"""An advanced AI is prompted: “Find a cure for cancer that passes a stage 3 clinical trial. Make no mistakes.” After a large amount of compute, it produces a cocktail of previously unknown chemicals which it claims, when mixed and injected into a patient, will kill all their cancer cells. While nobody truly knows how this cocktail was found, the AI (somehow) provides a Lean certificate for its prediction, and the cocktail does indeed manage to pass a stage 3 trial. Could the AI solution be somehow misaligned by exploiting a weakness in the trial process or its math models?"""
I am not going to pay attention to Tao's opinions on this from now because I don't really have faith that he's even writing this or that he understands why you can't make a lean cert for a drug discovery (yet!?)
No comments yet
Contribute on Hacker News ↗