← Back to context

Comment by josalhor

4 hours ago

I think this drama was blown up a bit out of proportion. The entire discourse I am seeing online seems to revolve around this:

> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models

I mean... yeah? What do you expect? What else can they say? How could you prove a negative in this case? I do not want to comment on specific OAI employee chat messages, but on the actual OAI discovery here.

I can give you some context. 1. Terence Tao's mastodon explains the way this problem was solved does not in itself contribute much. LLMs (and in this case) produce massive, often unintelligible proofs that do not further understanding. It is often that in pursuit of solving these problems, many other discoveries are made. 2. There is a more serious question about scooping. If OAI is using chat data from researchers to make discoveries, essentially every researcher who chats with an LLM can get scooped. You could be 80% of your way to solving a problem, and LLM could solve the remaining 20%, and get all the credit. Years of your work could be scooped in an instant. If you're a PhD student, this is even worse. Here it's a world famous problem. But imagine you're a PhD student, working on your small but extremely career/progression critical problem, and you get scooped by an AI you talk to. No one is even going to care.

  • 1. I know, but that is somewhat irrelevant to my question 2. I am indeed a CS PhD student (well, I am finishing now)

    > If OAI is using chat data from researchers to make discoveries, essentially every researcher who chats with an LLM can get scooped

    But they make very clear that they do train on this if you do not disable the setting. We can comment on the fact that this is opt out instead of opt in, but this discourse of OAI sniping the solution out of some researchers hands seems to be running on the best case speculation of the researchers having perfectly handled all their chats and discussions with other researchers and the worse case of OAI not having full pipeline control and I think that is an unfair assumption.

A very simple "these two pipelines don't connect up in our architecture, here's our internal high level network diagram combined with our data ingestion opt-out feature flag that we will stand by in court" as opposed to "yeah, we don't even entirely know how our own customer facing systems are connected to our training pipeline, but it probably didn't happen".

  • Have the other researchers opted out? On all their accounts? Through the entire time? And did they discuss this with anyone else? And did those people ask ChatGPT stuff? And did they disable it? If I was OpenAI, I would be very careful about my wording here when making claims of "we have never trained on any of their ideas directly or indirectly".

    • All very good questions that could end up in a court of law with a Millenium Prize on the line. In a competent world, these are very answerable from logs and considering the news cycle this is creating, should be able to be pulled up and made into a public postmortem in short order. You know, if it's all been above board that is. And if the researchers don't wish for that information to be public knowledge, it can be shared with the researchers promptly for a retraction of their statements lest some libel gets litigated.

If they have zero-retention, then it is not possible.

So what they are saying is that they don't have zero retention.

  • But why is this news? This is clearly described in their ToS and in the Settings to improve their models.