← Back to context

Comment by binlog

12 hours ago

What’s the nuance? Two researchers alleged theft. OpenAI investigated and confirmed that their research was not in the model’s training data. The world decided to take the first part as objective truth and ignored the second.

> OpenAI investigated and confirmed that their research was not in the model’s training data.

You mean when they say it was impossible to confirm anything but one day after it was 100% confirmed that there was no theft? And here I'm not even talking about all the ethical problems related to trying to scoop another group when you hear they're close to success, or how current solutions follow extremely closely human-generated ideas, or about the lack of relevant citations in OpenAI's paper.

Believing that OpenAI's claims have any substance cannot be explained by naivety alone.

  • They didn't try to scoop another group; they thought the other group had already solved it, so maybe their latest model could take a shot too... and the model solved the full problem when the other group had not! They found out after the fact that the other group had only solved an important sub-problem.

    This was a low-key hilarious replay of George Dantzig and his homework problems: https://en.wikipedia.org/wiki/George_Dantzig

    The controversy was whether they had plagiarised that other work on the sub-problem, which they categorically denied after an investigation. And yes, a few days to investigate something like this is reasonable for a company as big as OpenAI. Having seen how data infra is set up when petabytes of data are flowing about, there are thousands of entwined data pipelines to figure out. Not quite as easy as running a query on a sqlite DB!

    > Believing that OpenAI's claims have any substance cannot be explained by naivety alone.

    Yes, they could be explained by a GitHub repo full of proofs :-)

    Or are you suggesting there were hundreds of researchers who just happened to be close to solving hundreds of these long standing open problems using Codex, and OpenAI swooped in plagiarized them all? ;-)