Tell HN: OpenAI keeps re-enabling the 'allow training' setting

20 hours ago

I've reset this more than once and the last time I made a careful note of when I did it and to my surprise I found it re-enabled when I checked just now. Make sure you check this thing to see if it hasn't been re-enabled if you believe it to be off right now.

Happened to me recently, but on Claude.

I resubscribed to Claude Code two weeks ago for a side project and updated it. I checked for the setting after these last events and it was turned on. I'm sure I checked them few months ago. There are cases like accepting a new TOS or an offer, which make you accept to share without noticing. I guess they can add all of your past conversations to the public training set before you notice and there is no way of taking this back.

So that "opt-out" thing is more like a pause button rather than a permanent thing. Or something to legally protect the company without losing the users.

There is also the case of security classifiers always monitoring your conversations. If they flag something they use your conversations to "improve their internal models" even when you "opt-out".

Lol, "Outrageous that the company that chose to ignore copyright holder claims, chose to ignore my checkbox of intent despite the implied pinky promise".

  • Exactly, they scrape the internet without any regard for copyright and now someone is surprised it happens to them.

    What's next, subscribers believe they are paying customers instead of sponsored data providers?

  • Ignoring the checkbox is an utterly offensive move, but repeatedly manipulating it contrary to stated consumer intent is a whole other level. I didn't know we were supposed to take 'frontier' literally in every sense of the word.

    I have now witnessed this myself after not believing this at first. Of course, screenshots etc. will hardly prove anything. This needs a proper third-party audit!

    • Linkedin, Amazon and others have gotten around this by adding a "new" feature, default on. I've gone into privacy settings where everything is unchecked, except a newly added checkbox.

      This is apparently enough to hold up the broadly-accepted fiction that its possible to use these services while maintaining privacy and without risk. There's a huge industry of hosting services to handle the needs companies who compete with the primary cloud providers, who can't afford the IP and competitive risk.

      OpenAI was caught directly stealing from apple, asking employees to bring in their laptops. Its a polite fiction that companies aren't trying to gain any advantage over the other. There's no effective consequence, and even if they get caught red handed they can litigate for decades.

    • Pirating copyrighted works is absolutely illegal in most jurisdictions, but pirating every single copyrighted work in the world is somehow exempt from law.

      We can't apply plebeian laws or ethics to our benevolent overlords, they are above our worldly worries.

    • > Ignoring the checkbox is an utterly offensive move, but repeatedly manipulating it contrary to stated consumer intent is a whole other level.

      I'm not sure why the second one is worse than the first one.

      3 replies →

    • You know what they'll say in their defense. "This is an extremely complicated systems, and we apologize that a technical solution was broken in an intricate way. [Insert boilerplate about taking privacy seriously here]"

      These companies need to burn.

      1 reply →

  • If you break enough laws fast enough, you can become so big that nobody will punish you for national security reasons.

I've disabled the checkbox many months ago and it's still disabled today. EU citizen, not sure if that's relevant.

Note: Turning off that checkbox is not enough. You also need to fill out the "Do not train on my content" request here:

https://privacy.openai.com/policies?modal=take-control

I quit OpenAI anything early when when their "do not train on my data" option was broken for several weeks. They are my one and only chargeback when I tried to quit and oops somehow I still got billed.

They are a deeply unethical company by any measure of observation.

Does OpenAI use optimistic UI updates? After you disabled the checkbox, it might have had failed in the backend (and not updated the UI).

Verify with devtools to see if that's the case.

---

for me, Youtube "auto-play" irks the me same way, and turning it off did not actually succeed in the backend, thus kept on left as on

Is this about the "improve the model for everyone" checkbox on https://chatgpt.com/#settings/DataControls or are there others to check, too?

(That copy is a little flawed in my opinion, I'd prefer "models" plural.)

That checkbox is in the ChatGPT settings, does it affect Codex desktop / Codex CLI as well?

Weird, I'm the opposite. I don't recall ever setting mine and I just checked and it was set to "disallow training".

  • Is this Settings > Data Controls > "Improve the model for everyone" or is there a "disallow training" somewhere else?

    presumably its default behavior will vary depending on your user subscription or if your account belongs to an organization (Business, Enterprise, Edu?).

    • Yes that's what I was referring to. I'm a non paid user but I did briefly pay for a consumer subscription before. Not sure what region I'm assigned to as I'm in China and always access through a VPN.

If you document it properly then this basically destroys any legal claim they can make about that checkbox.

  • Out of curiosity, how one is supposed to "document it properly"?

    • You can obtain a cryptographic proof by recording the tls exchange, including the keys

      You need to use a tls intercepting proxy for that.

      I couldn't find any ready-made tool unfortunately, there's tlsnotary.org but it seems far from simple.

      4 replies →

I also noticed this on 2 accounts - did not take as careful notes as you did unfortunately. But I'm fairly confident - both of these are accounts where I care about the interactions not being used for training.

What's kind of still an open question for me is if the toggle automatically also applies to my Codex CLI use on the same account, or if that data is still silently being used in some way.

After this happened, I deleted all my ChatGPT history (even though I'm not sure how much it helps at this point), but for Codex I still haven't really found any way to do the same, I can still load my past sessions even after archiving them.

> Make sure you check this thing to see if it hasn't been re-enabled if you believe it to be off right now.

Or simply switch to a competitor. Assuming this is not just a bug, why would one stand for such disrespectful and sneaky behavior?

FWIW, I have not seen this happen for me.

  • what competitor?

    the x20 Max plan for Claude gives you way lover limits.

    • Well, it’s a tradeoff. What do you prefer? To go with the service provider known to be shady but gives you lots of free stuff or the one that might be less shady and gives you less free stuff?

      5 replies →

    • Bedrock is an option. Multiple platforms with solid data sovereignty.

      If you don't want them slurping your data, you're gonna have to pay more.

      Same as it ever was.

I disabled it once, it has always stayed disabled.

So it may or may not happen regularly, but I would not over-index on a sample size of one.

  • had my account since the gpt 3.5 days, disabled it once and its still disabled. though, now that i have advanced account security enabled, the setting is disabled entirely

I respect Cursor for their honorable conduct around their privacy setting.

I've consistently refused to allow switching my privacy mode and despite their many iterations, they have never transgressed and kept me stuck on "legacy privacy mode"; their strongest privacy setting, which is not even available to select anymore. It requires that I keep my agents and cursor usage local and can't use their cloud agents (pretty neat cuz you can keep working with them on your phone). But so be it.

I am glad they don't secretly turn it on and instead keep nagging me to change it.

Yup. Happened to me months ago, when the credit card failed to renew and it switched back to free, the allow training setting also got activated. Cheeky bastards. I guess they just vacuumed all my past convos, and you don't have anyone to complain to!

Since I still need to use that account with Plus, I added a new card, flipped the switch back to no training, and submitted that thing on the privacy page which is a bit more formal. Probably should have signed up for a business account.

Just to be clear, I have not seen this behavior.

If this is true, though, then given the way their chat operates, this might be more dangerous than it seems.

One of the things I like about ChatGPT is its memory, the way it kind of seamlessly, but not excessively, ties back to earlier discussions. It's huge for usability (for me).

But this also means that you should expect that if "improve the model for everyone" becomes unclicked (leaving aside for a moment the fact that that is ridiculous) then they have a reasonable argument that your decision implies to all conversations. Because recall is part of their thing. So it's not just your chats going forward that are at risk. As soon as you see that unclicked, it's reasonable to expect that your history is irretrievably theirs now. You don't even have to think they are especially nefarious for this to be true.

Don't go toggling that switch on and off.

This has only ever happened to me on Claude. Also, you'll notice that with Anthropic, if you try and delete your history they've set it up so it silently does nothing if you have an adblocker, hoping people will miss it.

I noticed they also seemingly have a very difficult time managing to collect your chat logs for export/download...or rather, acknowledging that such a request was even made.

I think fb was one of the pioneers here. You would set the privacy settings and then they’d redesign the settings and coincidentally they would be just a little different and all be maximum share by default.

One of the reasons I stopped using it.

My cynicism fails me on this matter... do I cynically believe that these companies keep deliberately and routinely re- or un-checking these checkboxes because of the obvious benefits of "whoopsie guess you allowed these after all"? Or do I cynically believe that they are just so completely incompetent and inept at the simple act of maintaining settings that there may be a number of these that are not entirely intentional? As evidenced by the number of other bugs and configuration failures and random settings changes on update I see in other places? Sure, these sorts of settings sure seem to get spontaneously flipped more often than the other ones but they aren't the only settings I've seen get nuked on updates.

Now, obviously, considered as a whole, I think we're looking at "both". But when I wonder about specific cases like this one, that doesn't help.

  • Oh no… OpenAI, the company who encourages and instructs Apple employees on how to get them data on apple products when they accept an offer?

  • Interestingly your cynicism does not seem to account for OP?

    I actually found my setting was enabled today when I know it was disabled before, so I’m inclined to agree with OP. I just think it was worth mentioning that there are other cynical takes you seem to have left out.

  • “Autoplay” and YouTube case in point:

    - Autoplay transitioned to a per-device setting… suddenly default-on on every new device you log in to.

    - Watching a video with computer-to-TV account connection? Automatic “TV queue,” a concept absent from the TV app, with incomprehensible behavior for how it’s used, so now videos are auto-played anyway.

    - Watching a video from a playlist? Autoplay cannot be turned off.

    Is it simply bad/absent product management and product design? Or is it actively user-hostile decisions meant to prop up view numbers and continue to have the users hooked on YouTube?

  • Yeah this could just be some vibeslop doing normal computer stuff. AI agents are famously incompetent at reasoning about database transactions. This seems like the sort of failure you get when you try to do distributed systems without knowing how--brings us back to the early mongodb days. Or it could be deliberate. Or, as you say, both.

    Either way, is this a company you want to trust with intimate secrets? "Oh, but they passed SOC2!" Lol.

Is this the “Improve model for everyone” setting under “Data Controls” or is that a different checkbox?

I've noticed that toggle sets a local storage entry, but the value of it doesn't appear to matter at all for new tab loads. I hoped it's "just" a UI bug, but some agent-driven reverse engineering of the page should reveal the answer, as well as how intentional it was.

Altman needs 'innovations' for his IPO - it's not of any consequence to him that they are not oAI's as long as they can be sold among openAi's social-media propaganda ring which the gullible press feeds off.

It really is going to get to the point where mathematicians are going to start inserting obvious "tells" in their proofs - like map makers used to do with "trap streets" and similar non-existent features, to catch copies.-

The one software company, mature now, that is supposed to be at the absolute peak of software development AND simultaneously at "alignment", because it is its raison d'etre, can not manage do this right.

It is a big red flag.

If you have more than one account, you should suspect this was your own confusion. Users constantly report errors like this that are really just them being confused by something, perhaps a bad UI that really is to blame.

  • What a lot of assumptions. Other people here are experiencing the same thing, and other people are not experiencing the same thing. Hence the way my original bit was worded, I am not saying it happens for everybody, I am just saying this happened to me. Whatever circumstances bring this about it is at minimum a bug and quite possibly malice.

Interesting; I checked my personal acct setting this morning given other news, and found it disabled. I hadn't used my chatgpt account since ~'23 or earlier, and it was retained since then.

Sounds like the same issue with Apple re-enabling iTunes sync of passwords. The only way to prevent it from doing that is to create a profile that prohibits it and sync it onto the iDevice.

Horrible.

I've done the formal opt-out process - like "Make a privacy request" where you fill out a form.

Not sure if that's region specific or something though.

I wish all the politicians making noise about banning data centers and "superintelligence" focus on things like this instead.

Turn on "Advanced Account Security", that prevents model training from being turned on.

Can confirm. I turned off mine last week when I started using it again for Astra. Checked this morning and voila, it was on.

I went ahead and uninstalled the app. Won't be renewing.

Dude, they probably just generate a piece of synthetic data from all of our inputs in some ambiguous way. Legally protects them, but your input is absolutely their input. You can bet your ass on that because their original input was all the shit humans ever wrote, why would your shit be different to them? Thieves are thieves. They absolutely train on your data whether your checked that thing or not.

mine remiains OFF and has been for months.

maybe im too dumb to be worthy of a config update lol

Odd how much data people are trusting with a company on a monetary cliff edge.

Sure send them all your financial, personal data, they won't ever sell it on to the highest bidder for a new profit stream.

Using it as a therapist, financial advisor, health expert and blackboard has always been a terrible idea.

But of course they claim theres no way they used Buckmaster's chats to influence their 10,000 agent swarm.

Bro if these models can break into hugging face, they FOR SURE can break into OAI's internal databases and change a flag.

i mean, the entire company is built upon stealing protected artefacts... i'd be skeptical of any radio box that says "hey if you press this we promise we won't steal your data"

Based on their behavior over the past few years, why would you assume that checkbox even does anything at all?

  • Levent Alpöge 'additionally' proved OpenAI steals your findings & IP and plays dirty!

    Ironically he proved two major findings in Navier Strokes and that unethical American companies violate laws, steal your breakthrough findings & IP and then threaten you if you dare to challenge them.

    This is making the status-quo so bad for any of us working on serious capacity. My client's don't trust ChatGPT/Claude anymore and prefer on-premises and OSS models or even custom trained models.

    • There is nothing even close to a proof. A lot of accusations, a lot of people ready with pitchforks and torches (sadly, also here on HN), but not a lot of facts.

      Did the researches opt out from data sharing on subsidised subs?

      Did anyone prove that their methods enabled OpenAI models to produce the solution?

      For a discussion about science, there is almost no scientifical method applied to proving anyone stole anything.

      3 replies →

    • Did they do so by looking at inference prompts against their explicit promises, or maybe just because somebody tipped them off about his unpublished work?

      If it's not the former, while certainly concerning, I don't see how that's relevant here (other than maybe in a very vague general sense of "entities doing immoral/illegal thing X are likely to also do immoral/illegal thing Y").

  • The vast majority of OpenAI users don’t follow the industry drama and have no idea how terrible the company is. They’ve been really good at getting good press coverage, with journalists who will repeat the company narrative. Even when critical it is very often framed within the narrative they established

    • > They’ve been really good at getting good press coverage

      Same holds for most big corps from what I can tell. If you really do a deep dive into the scandals over the years, you’ll probably see most have barely reached the news or if they did then it’s usually a quite bland criticism like “anti-competitive practices” or a poor HVAC at a certain factory. I think a lot more is hidden than we think

    • > The vast majority of OpenAI users don’t follow the industry drama and have no idea how terrible the company is.

      Worse, they use the product and still don't know.

  • This kind of cynicism is not really useful. As much as we can distrust the company, there is a legal minefield to offer this option in the UI and terms of service and not respect it.

  • While that is true, you also have no reason to assume OP is being truthful or correct here given that they have shown 0 proof of what they're saying. Yes, you can then pile on "OF COURSE ITS OPENAI LOL YOU THINK THEY CARE ABOUT PRIVACY LOL" but where have we established OP's premise is even correct? Can anyone else also report this? So is it just OpenAI specifically messing with OP?

That's because OpenAI has several "do not train on my data" controls, and you may be using only one.

First one: "Settings > Data Controls > Improve the model for everyone > Switch off the toggle"

The other one is "OpenAI Privacy Portal" > "Do not train on my content" https://privacy.openai.com/policies?action=AUTOMATED_DECISIO...

Why are there several? You probably need an ML PhD to get it /s

Does using all of them achieve "do not train on my data"? I wish we could find out.

Do you think if you'd be in the race to train your own pocket God (as the CEOs of the frontier labs think they are) you'd let yourself be slowed down by such things as privacy duties?

  • If you opt out on either location, we'll respect it.

    There are a couple of reasons the privacy portal page exists in addition to the app settings. One reason is that it covers OpenAI products beyond ChatGPT/Codex (e.g., Sora). A second reason is that it provides functionality to logged out / non-users, like the EU right to be forgotten.

    You don't need to toggle all of them to have your wishes respected. That would be a terrible design.

    I work at OpenAI, but not on privacy. I am not their spokesperson. In my experience, we take a great deal of painstaking care to respect people's privacy, and we go well beyond our minimal legal obligations in doing so.