← Back to context

Comment by Pulcinella

13 hours ago

Do you not log the training data? Seems like you should be able to just check what was in the training data. To not keep track is just sloppy work and certainly unprofessional science.

That training data does not preserve provenance seems a "smoking gun" in terms of intent to plagiarize.

The burden of proof is on Buckmaster and Alpöge to reveal if they had the "Improve the model for everyone" setting enabled or disabled. OpenAI shouldn't be expected to reveal private user configuration data. You're asking them to perform a user privacy violation.

  • If the reason that OpenAI is unable to state whether they trained on this data is because they (as policy) do not reveal whether a given member has turned on/off the "Improve the model for everyone" setting, they can at least say so.

    FWIW, publicly facing OAI docs are very unclear about whether this setting even applies to Codex conversations.

Does Anthropic, Grok, etc. log their training data? I had the impression it was rather a mess.