Comment by Arodex
16 hours ago
But I can subpoena my neighbours bank account when I see him driving a brand new 500'000$ car and I have a 490'000$ hole in my bank account and he works in the bank where my money is. And when questioned he evades some questions and threatens to destroy my career.
Any other argument, fc417fc802?
You're making a classic a burden-of-proof fallacy. The burden of proof lies on the person making the claim, not the person questioning it.
See Russell's teapot for an explanation https://en.wikipedia.org/wiki/Russell%27s_teapot
> You're making a classic a burden-of-proof fallacy
This is incorrect, and you invoke Russell's teapot incorrectly too.
It would only apply if the accusation rested solely on the fact that neither of us have evidence against the accusation.
But that's not the case. First, we know that there could be proof, it's just apparently burdensome and expensive to produce. At that point you're not in fallacy land anymore, you just need a way to balance the cost required of someone to prove the accusations against them false.
Second, we have an arguably plausible mechanism of action that OpenAI does not dispute is possible.
This isn't a legal dispute, so no one is going to force OpenAI to do anything here, but it's not unreasonable (and certainly not fallacious) to suggest that Buckmaster's suggestions are plausible enough it's up to OpenAI to stand behind their denial.
No, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set.
First, the conversations are anonymized, so there's no simple way to inspect the training dataset and identify which specific conversations belong to Buckmaster.
Second, OpenAI uses these anonymized chats to generate synthetic training data, i.e. they fabricate new conversations based on specific conversation patterns where the model performs poorly, and uses these synthetic conversations as training data for future models. The synthetic data could potentially contain some of selections of Buckmaster's original chats, but it is unknowable how his specific writing could have influenced these synthetic data sets or what portion belongs to him. This information is untraceable and effectively double anonymized.
Third, OpenAI explicitly uses user feedback (the thumbs up or thumbs down ratings), as RLHF to train models. However, this feedback is anonymized and stripped of user identifiers. It's not possible to trace a specific feedback to Buckmaster, nor do we know if Buckmaster ever used this feature. I doubt Buckmaster recalls or can provide a list of every time he used this feature over the past year. OpenAI doesn't have one.
Note that the first and second only happen if Buckmaster "Improve the model for everyone" setting enabled, which I find unlikely. But that doesn't exclude option three from this list.
You seem to think that it is some "gotcha" that OpenAI refuses to make a blanket denial, but they cannot do so in good faith, because they have a genuine understanding of their own system. They don't know where the data they have came from.
This situation meets the requirement of Russell's teapot, since neither party has enough evidence to prove nor disprove what information is actually in OpenAI's training set.
1 reply →
The accused party fails to answer half the questions and makes direct threats. I would say the accuser has already collected enough proof to trigger an investigation.