Comment by Pulcinella
14 hours ago
Do you not log the training data? Seems like you should be able to just check what was in the training data. To not keep track is just sloppy work and certainly unprofessional science.
14 hours ago
Do you not log the training data? Seems like you should be able to just check what was in the training data. To not keep track is just sloppy work and certainly unprofessional science.
That training data does not preserve provenance seems a "smoking gun" in terms of intent to plagiarize.
The burden of proof is on Buckmaster and Alpöge to reveal if they had the "Improve the model for everyone" setting enabled or disabled. OpenAI shouldn't be expected to reveal private user configuration data. You're asking them to perform a user privacy violation.
If the reason that OpenAI is unable to state whether they trained on this data is because they (as policy) do not reveal whether a given member has turned on/off the "Improve the model for everyone" setting, they can at least say so.
FWIW, publicly facing OAI docs are very unclear about whether this setting even applies to Codex conversations.
An OpenAI employee did say so: https://x.com/tszzl/status/2097393423808377173
it is exceptionally unlikely that anything they ever did made it into any part of training, and the chances are zero if they have opted out (likely). it would be a terrible precedent to break the the PII-scrubbing boundary to go and round it down to 0, and we won’t do it
3 replies →
Does Anthropic, Grok, etc. log their training data? I had the impression it was rather a mess.