Comment by ALLTaken
21 hours ago
Levent Alpöge 'additionally' proved OpenAI steals your findings & IP and plays dirty!
Ironically he proved two major findings in Navier Strokes and that unethical American companies violate laws, steal your breakthrough findings & IP and then threaten you if you dare to challenge them.
This is making the status-quo so bad for any of us working on serious capacity. My client's don't trust ChatGPT/Claude anymore and prefer on-premises and OSS models or even custom trained models.
There is nothing even close to a proof. A lot of accusations, a lot of people ready with pitchforks and torches (sadly, also here on HN), but not a lot of facts.
Did the researches opt out from data sharing on subsidised subs?
Did anyone prove that their methods enabled OpenAI models to produce the solution?
For a discussion about science, there is almost no scientifical method applied to proving anyone stole anything.
On one side, yes we don't have hard evidence that intentional plagiarism is exactly what happened.
On the other side, the lack of evidence is pretty damning. Only OpenAI can try to prove that they came by these results legitimately, and the case they're making is quite weak. They could make public metadata about what their model was trained on and whether it did train on the conversations in question; they have not. TBQH I read it as even they don't know.
And regardless of whether the result is legitimately obtained by their model, they've not at all conducted themselves well throughout this story. They set out to scoop researchers based on a rumor. They threatened to ruin a mathematicians career. They put up a paper that deliberately doesn't cite the most relevant research, despite building directly on it. No matter how you look at it, OpenAI has and should lose any standing they had in the research community.
That it was a dick move, I think there is no doubt about that. OpenAI wanted to scoop Anthropic, and the two guys working on the problem got caught in the crossfire.
Both OpenAI and the researchers know if the sessions in questions were subject to data sharing. Why neither the scientists nor OpenAI is clear about that is weird - it would seem at least one party has the incentive to report that. But even if their sessions were in training data sets its hard to tell whether it influenced the outcome. Those models are big, but are they big enough to preserve subtle, niche techniques enough to draw from them while solving a related problem? Probably nobody knows.
Actually no, the lack of evidence can be readily fixed by the researchers simply disclosing the pertinent parts of their notes and/or chats. The discovery has been scooped, so I don't see any value in keeping them private anymore. Then everybody can see how related the models' and the researchers' works are.
I take it you are referring to this[1], which has nothing to do with training.
1. https://thenextweb.com/news/bubeck-navier-stokes-account-apo...
Did they do so by looking at inference prompts against their explicit promises, or maybe just because somebody tipped them off about his unpublished work?
If it's not the former, while certainly concerning, I don't see how that's relevant here (other than maybe in a very vague general sense of "entities doing immoral/illegal thing X are likely to also do immoral/illegal thing Y").
I do trust Anthropic, i haven't observed them doing super shady stuff like hiding the training consent page and making you WAIT for it to activate.
You really shouldn't trust any company. Their goal is maximizing profit and that is (usually) not aligned with "doing what's right". Not in the long-term at least.
We really need a new Hanlon's razor for companies. Assume malice over ignorance because groups behave differently than individuals.
You mean the company which trained on others books, won't train on its own user generated data?
AFAIK, the court held that the problem was that they had acquired books by pirating them, not that they trained an AI on those books. They do train on user generated data by default, but you can opt out, that's the whole point.
Let's not down vote a factually correct comment that makes you emotionally disagree with me. I'm very concerned in regards to my privacy, I've studied ToS of: Openrouter, Anthropic, Anthropic business, opencode Go etc. I've double checked my understandings by GDPR exporting all my data. Anthropic sub listed a generic "logged in from Linux", "UA", "timestamp", "Country". That's it.
I want to discuss actual facts, observations and not your shallow social signaling replies of 0 value.
I mean weren't everyones privately shared chats Google searchable?
User doesn't understand how public links indexing works.
ah a proof by counterexample