← Back to context

Comment by nulld3v

18 hours ago

Nobody except OpenAI knows whether or not OpenAI trained on their data. So the burden remains on OpenAI here.

Incorrect, Buckmaster and Alpöge can comment if they had the ChatGPT "Improve the model for everyone" setting enabled or disabled.

If it was enabled, then their work was included in the training dataset.

  • In order for this to be the strong evidence everyone also has to believe that the setting is absolutely true. That some logging from some piece of the system could not also leak the prompt information in such a way that it could have been included as training data. Perhaps the design of how data is collected for the training dataset is so rigorous as to make this a practical impossibility. But, it's asking a lot without sufficient detail to completely exclude from possibility that one setting is all that could possibly have been absolutely load bearing in deciding if the other researcher's active efforts meaningfully contaminated the internal model.

    At least, as an ignorant outsider, that's how it seems to me.

  • As I understand it, that setting does not prevent them training on user data, just which derivatives are used (i.e. just PII scrubbed vs certain types of synthetic summarization)

That is an absurd and entirely untenable position that breaks with approximately all western conventions.

Only the CIA knows whether or not they're actively covering up reptilian space aliens exerting control over the US government. Therefore the burden of proof remains on the CIA to prove that they are not actively participating in such a scheme.

  • I don't understand, OpenAI can just say: "yes/no we did/did not train on your data". It's not a hard question to answer, and it is a question that OpenAI should be able to answer for all data we feed into ChatGPT.

    • > It's not a hard question to answer

      I didn't realize you had insider knowledge about their systems. Do please explain for the class.

      As I understand it they will only have trained on his data if he consented to it. Do you have evidence that they do otherwise?

      2 replies →