Comment by enraged_camel
15 hours ago
>> I'm not sure how any of this provides evidence that OpenAI took any of their work.
Sorry, but the burden of proof lies in the other direction: OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published. Otherwise OpenAI simply stole the glory and the spotlight (and I'm being charitable here).
That's entirely unreasonable. Allegations of malfeasance always need to be backed up by evidence.
But there is evidence, the blog post says: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
In other words, yes, they had been using ChatGPT, and yes, ChatGPT could very well have trained on their data. Now that there is evidence, we need an investigation: yes or no, was it the case?
That is not an admission of malfeasance though? As I read it they don't know if anyone fed relevant private documents into the model under an account configured to permit training on user data.
If there's more to the story I'd be interested to hear it.
2 replies →
But the evidence is in the hand of the potential culprit. That's why allegations can be enough to force confiscation and intrusion to get evidence in safe hands before it is destroyed by the accused party.
Only in the event that there is some reason to suspect them of wrongdoing. Which would generally require evidence.
You don't just get to subpoena your neighbor's bank account because "I know he's stealing from me" you need to first present credible evidence that you were stolen from and that he is among the most likely culprits.
5 replies →
That is backwards. It is the responsibility of a researcher to do a thorough literature review and conscientiously avoid plagiarism or claiming false novelty.
Nobody except OpenAI knows whether or not OpenAI trained on their data. So the burden remains on OpenAI here.
Incorrect, Buckmaster and Alpöge can comment if they had the ChatGPT "Improve the model for everyone" setting enabled or disabled.
If it was enabled, then their work was included in the training dataset.
2 replies →
That is an absurd and entirely untenable position that breaks with approximately all western conventions.
Only the CIA knows whether or not they're actively covering up reptilian space aliens exerting control over the US government. Therefore the burden of proof remains on the CIA to prove that they are not actively participating in such a scheme.
5 replies →
How does one prove a negative ?
https://en.wikipedia.org/wiki/Burden_of_proof_(philosophy)#P...
> OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published.
I don’t think they’re too concerned about appeasing you, enraged_camel.
For most reasonable people, achievement in solving the other Millenium Prize problems at an unprecedented rate will be enough. At some point people will see models are capable of solving hard issues without whatever 0.00001% of the training data coming from irate individuals who believe their sample was the key component of the solution.