Comment by jeswin

8 hours ago

All of these accusations could be true. But there's also no way for a company to casually claim "No, we did not train on your data", without verifying all the knobs the user might have turned to enable or disable data sharing.

I just don't understand getting the pitchforks out because a company did not give an answer immediately. And the effect such data entering training would have affected the output is even less clear.

The pitchforks are out because even without that part, it's still a scumbag move to try to frontrun the mathematicians who were working on this for years after OpenAI heard that they were close to releasing their results. Just identifying that one of these problems is solvable takes a lot of work. The only reason OpenAI got this result is because the mathematician shared with colleagues that he had made significant progress and was close to solving it, and OpenAI could not accept that so they decided to throw tens of millions to make sure it doesn't happen without them getting all the glory. Notice that their paper doesn't even have an author since they're probably all aware of how awful that would look, and no one wanted to take on the shame. They probably also knew that the paper was trash and no one involved could understand it, and didn't even cite many of the people who contributed to all of that knowledge. It's just a disgusting act any way you slice it, even without training on the prompts or the nasty communication by the OpenAI leaders.

  • > They probably also knew that the paper was trash

    Doesn't matter. This forum used to celebrate "because you can" with no riders. And solving a Millennium Prize problem is among the biggest stages for Because We Can.

    Now we're saying there are some qualifiers attached to it, such as (1) only if not done by companies with a lot of money, (2) only if it is inconsequential.

    I agree with some of what you're saying, but like everything else it isn't black and white. Maybe some day, someone will improve some particular treatment because we can.

At the very least, a company shrugging and saying it’s impossible to know whether academic plagiarism had occurred is a claim that needs to be justified, not taken at face value.

And even if so, it should be on the company to design systems to avoid academic plagiarism and offer the right transparency. It shouldn’t suffice to say “we don’t know what went into the model, when, or how” —- that’s a solvable problem that an accountable company can satisfy.