← Back to context

Comment by mucha

16 hours ago

That's not what your Chief Research Officer, Mark Chen, says on X: "Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company."

https://x.com/markchen90/status/2097400166554993041

Can you explain what part of his post you believe is inconsistent with that quote?

  • "If they opted out of training, then we definitely did not train on them."

    Per OpenAI's privacy policy, they use de-identified data to improve their products. From Mark Chen's comment, improving products includes improving ChatGPT and Codex in a holistic way. Improving models in a holistic way sounds a lot like training to me.

    • > Per OpenAI's privacy policy, they use de-identified data to improve their products

      That’s not inconsistent with what you responded to. They use your data unless you opt out. If the user doesn’t opt out, their de-identified data is used to improve their products.

      2 replies →

    • > Improving models in a holistic way sounds a lot like training to me.

      I think that's quite a leap. Using de-indetified data to improve the products is what everyone has been doing since the dawn of web analytics.

Does OpenAI think de-identified data is no longer user data? Wild take for OpenAI and certainly not industry standard.