Comment by nl
10 hours ago
You are being downvoted for this because OpenAI subsequently checked and clarified than none of the relevant conversations were in the training data in anyway for the Navier-Stokes result.
10 hours ago
You are being downvoted for this because OpenAI subsequently checked and clarified than none of the relevant conversations were in the training data in anyway for the Navier-Stokes result.
> because OpenAI subsequently checked and clarified than none of the relevant conversations were in the training data in anyway for the Navier-Stokes result.
could you give link? Because I remember they said they couldn't verify:
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."
Sure.
> Following an investigation, we have confirmed that Buckmaster’s Codex prompts over the two months preceding this announcement and paper on September 8, 2026, could not have influenced the system in any way, including through training.
https://openai.com/index/navier-stokes-solution/
for 2 months prior. Not any of the relevant conversations. For a cutoff date a couple months before the announcement. They said they had been working on that problem for a year or more
Thanks, it's hard to stay up to date. However, we are just meant to believe that the mathematicians were going about the proof independently in the exact same way as the machines did it. It seems like a very odd coincidence to me.
True. They investigated themselves for one day.
Oh, well if notorious liar Sam Altman and his company notorious for lying says so…
Great that they're so transparent and honest, BTW, can you ask them if they trained on any copyrighted data that they pirated?
Courts have rejected the "training is piracy" interpretation.
I agree with the courts. I don't think learning from something is piracy in anyway.
Obviously though this is a very different issue to what the OP was claiming. In that case there is no legal argument at all that they could train on it and the argument is there about moral rights.
Buying and copying one training manual and distributing it to 1000s of human workers is considered illegal, but somehow scanning one book and sending it to 1000s of distributed training instances is not?
Also, you learning something is different than a model learning it, because a model is not a person. You can learn from a book and sell the skills you gained from it, but you can only be in one place at a time. The model can serve that knowledge to every person on the planet simultaneously. We obviously need new laws since this is a fundamentally different situation.
4 replies →
Training a model is not "learning something". Only people learn things. Whether training is a fair use is debatable, but it has nothing to do with the justification that people are allowed to learn from books.