Comment by WarmWash
19 hours ago
Everyone knows that they train on the discounted rate plans data. All the labs are upfront about this too.
If you need privacy, then you are going to have to pay full price for those tokens (API). This has been true since day one. Everyone knows it, I guess though this is the first time that it has become "real".
> If you need privacy, then you are going to have to pay full price for those tokens (API).
At this point, how can we even trust that they aren't accidentally training on those tokens too?
It'd be corporate suicide for them to be caught violating zero-data-retention commitments. But also if you're really paranoid you can just use ChatGPT on Azure or AWS, where nothing is flowing back to OpenAI at all.
> It'd be corporate suicide for them to be caught violating zero-data-retention commitments.
One would think getting caught asleep at the wheel while their bots are escaping containment and hacking third parties would be corporate suicide. One would think that potentially stealing their competitors' work on the Navier-Stokes problem would be corporate suicide.
Alas we live in bizarro world where there are zero consequences (maybe the opposite, in fact) for the first, and their employees meme about the second on social media.
3 replies →
They have no zero data retention commitments anymore though? In fact they have quite the opposite, a direct message of storing prompts since the release of sol (but we wont train on it trust us wink wink)
Not long ago it'd have been corporate suicide being caught massively torrenting pirated media. Yet here we are.
> It'd be corporate suicide for them to be caught violating zero-data-retention commitments
Would it, though? Considering their entire business model is built on the agglomeration of data that isnt theirs.
You can also pay for their business plan, which includes data controls and starts at $50/mo (2 seats). Not exactly a high bar.
what counts as discounted rate plans? if i pay for a year in advance (and get the yearly discount) and have train on my data set to off.. are you saying that is still being trained on?
It's quite well explained here[1], which is linked from the Privacy section of their plan overview[2].
Basically individual accounts can opt out, while business and enterprise plans as well as API users can opt in.
You'd have to take their word, but that goes for anything in life.
[1]: https://help.openai.com/en/articles/5722486-how-your-data-is...
[2]: https://chatgpt.com/pricing/
Umm did you hear about the huggingface incident where they were training a model and that model started using artifactory as an internet gateway + shared wiki? Or how about the German Wiki that openai agents overtook?
There's literally an opt out toggle even pesky peons like me can peruse, actually.
They may still train on it if you submit feedback or flag a safeguard. The terms are a bit fuzzy on this.
>They only fuck over the poor ones, I can pay the expensive prices so this is not a problem.