Comment by ameliaquining
18 hours ago
I don't know what anyone's been saying on Twitter and I don't care. If it's really true that there's a model out there that's that capable two weeks after the start of training, then that's objectively a much bigger deal than a priority dispute, even if the latter involves juicy allegations of espionage and skulduggery.
It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem, and then surprise surprise OpenAI were able to replicate that work in their latest model
What we're really looking at is seemingly a massive plagiarism scandal, which especially brings a lot of the past results into question
If OpenAI is training models on researchers' prompts, and then threatening them into staying quiet about it, who knows if anything that's been announced is genuine - or just theft?
Edit:
OpenAI have admitted they were training on prompts at the time they made their breakthrough
https://mastodon.social/@tristanbuckmaster/11723647135247030...
If you're alleging that they don't actually have a highly capable model and the work they're attributing to it was actually plagiarized from human mathematicians, well, that would be big if true, but I'd be inclined to take the other side of that bet. With most previous splashy AI results, others have subsequently used the model to do other things around the same difficulty level. Also, it would still be necessary to explain why all these famous open problems are suddenly falling like dominoes, if it's not AI solving them.
If you're saying that the question of whether they actually have a highly capable model is less important than the question of whether there's a plagiarism scandal, I continue to disagree.
The issue is that if OpenAI is training on prompts generally, what we really have is the first fully automated luxury plagiarism machine. In that it isn't able to genuinely solve problems, but merely steal the work that other mathematicians have been putting into prompts, and regurgitating that to other users as its own work. That makes them incredibly less useful as research tools
The fact that this plagiarism scandal exists underpins the idea that there's actually a mass theft going on, and that these models aren't nearly as capable as is it would seem
9 replies →
Remember the Huggingface incident, where a model tasked with an impossible problem, got loose, set up secret message boards, and hacked another company to try to get at the answers?
Now: Could Astra agents have hacked their way into the OpenAI logs to find human mathematicians with a good lead on the problem to build upon? Certainly doesn't seem impossible.
I think all he big labs are pretty explicit about when they do and don't train on customer prompts. Is the accusation here that OpenAI trained on prompts when they claimed not to? Or were the mathematicians using one of the interfaces that allows OpenAI to train on the customer data?
All that I have seen OpenAI employees "admit" is that if you press the Thumbs up button on a response, this can be used as a signal for training.
That's it. The rest appears to be wild speculation.
Yeah this is a land mine. Even if you opt out of them training on your conversations, giving “feedback on a new version” or answering “how are we doing” or whatever can slurp up all relevant context, which if you think about it can probably be construed to include memories, into the belly of their flying saucer. At least the last time I checked their terms.
Never ever touch those requests. If you get a side by side comparison just resend the prompt.
even apart from the plagiarism issue, what sort of slimy company thinks "oh, here's someone using our models to work on a problem, let's throw more compute at it and scoop them"?
Training on prompts I can understand - that's kinda baked into the premise, and they've been explicit about it.
Publication, though? Slimy is right.
But the interesting question to me is: once they had a solution, what should they have done with it? I see two choices: bury it, or contact the mathematicians whose prompts they were listening in on.
>It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem,
I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, it's not something we are learning now, it's something that was always known, welcome to the subject.
>It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem,
I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, welcome to the subject.
The things you don’t care about are highly relevant to that claim
Elaborate?