Tell HN: OpenAI keeps re-enabling the 'allow training' setting
Posted by jacquesm 5 days ago
I've reset this more than once and the last time I made a careful note of when I did it and to my surprise I found it re-enabled when I checked just now. Make sure you check this thing to see if it hasn't been re-enabled if you believe it to be off right now.
Comments
Comment by ozgung 5 days ago
I resubscribed to Claude Code two weeks ago for a side project and updated it. I checked for the setting after these last events and it was turned on. I'm sure I checked them few months ago. There are cases like accepting a new TOS or an offer, which make you accept to share without noticing. I guess they can add all of your past conversations to the public training set before you notice and there is no way of taking this back.
So that "opt-out" thing is more like a pause button rather than a permanent thing. Or something to legally protect the company without losing the users.
There is also the case of security classifiers always monitoring your conversations. If they flag something they use your conversations to "improve their internal models" even when you "opt-out".
Comment by ActionHank 5 days ago
Comment by quikoa 5 days ago
What's next, subscribers believe they are paying customers instead of sponsored data providers?
Comment by jacquesm 5 days ago
Comment by infamouscow 5 days ago
It's unfortunate that most societies have outlawed this practice. I'm certain if that were in place today, we wouldn't have this problem.
I've yet to hear any alternative solution that's as effective.
Comment by mmh0000 5 days ago
/s for the /s impaired.
Comment by achrono 5 days ago
I have now witnessed this myself after not believing this at first. Of course, screenshots etc. will hardly prove anything. This needs a proper third-party audit!
Comment by xboxnolifes 5 days ago
Comment by mcmcmc 5 days ago
Comment by threetonesun 5 days ago
Comment by mcmcmc 5 days ago
Comment by threetonesun 5 days ago
Comment by mcmcmc 5 days ago
Comment by lbreakjai 4 days ago
Comment by croon 4 days ago
Annoying yes, but least bad.
Comment by Yizahi 5 days ago
Comment by taurath 5 days ago
This is apparently enough to hold up the broadly-accepted fiction that its possible to use these services while maintaining privacy and without risk. There's a huge industry of hosting services to handle the needs companies who compete with the primary cloud providers, who can't afford the IP and competitive risk.
OpenAI was caught directly stealing from apple, asking employees to bring in their laptops. Its a polite fiction that companies aren't trying to gain any advantage over the other. There's no effective consequence, and even if they get caught red handed they can litigate for decades.
Comment by gspr 5 days ago
These companies need to burn.
Comment by jacquesm 5 days ago
Comment by Yizahi 5 days ago
We can't apply plebeian laws or ethics to our benevolent overlords, they are above our worldly worries.
Comment by pavel_lishin 5 days ago
I'm not sure why the second one is worse than the first one.
Comment by Gander5739 5 days ago
Comment by duozerk 5 days ago
"Oh, side effect of a recent update, silly us, we'll improve QA, sorry !".
Or "once again our all-powerful models got away from us ! they did it themselves, the little scamps".
Comment by Gander5739 5 days ago
Comment by deaton 5 days ago
Comment by cbg0 5 days ago
Comment by gpugreg 5 days ago
Comment by noir_lord 5 days ago
Comment by bcrosby95 5 days ago
> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more
Comment by noir_lord 5 days ago
Could be worse I guess.
Comment by tyfon 5 days ago
Edit: I have a pro subscription now but it has also been on a free tier level for perhaps 1.5 of these years.
Comment by stingraycharles 5 days ago
Comment by bossyTeacher 5 days ago
Comment by stronglikedan 5 days ago
Comment by stingraycharles 4 days ago
Comment by layer8 5 days ago
Comment by burntalmonds 5 days ago
Comment by henry2023 5 days ago
At best they’ll “anonymize” the data before adding it to the corpus.
Comment by spongebobstoes 5 days ago
Comment by noname120 5 days ago
Comment by docheinestages 5 days ago
Comment by b800h 5 days ago
Comment by qurren 5 days ago
Comment by teej 5 days ago
Comment by buellerbueller 5 days ago
Comment by jryle70 4 days ago
(Just a faintest, without any clue, completely baseless suspicion on my end, just like what you did)
Comment by camel_Snake 4 days ago
Comment by mcmcmc 5 days ago
Comment by mellosouls 5 days ago
Comment by terminalbraid 5 days ago
They are a deeply unethical company by any measure of observation.
Comment by yipinwong 5 days ago
Verify with devtools to see if that's the case.
---
for me, Youtube "auto-play" irks the me same way, and turning it off did not actually succeed in the backend, thus kept on left as on
Comment by frangonf 5 days ago
Comment by jtolly710 5 days ago
Comment by simonw 5 days ago
(That copy is a little flawed in my opinion, I'd prefer "models" plural.)
That checkbox is in the ChatGPT settings, does it affect Codex desktop / Codex CLI as well?
Comment by olalonde 5 days ago
Comment by riffic 5 days ago
presumably its default behavior will vary depending on your user subscription or if your account belongs to an organization (Business, Enterprise, Edu?).
Comment by olalonde 5 days ago
Comment by amelius 5 days ago
Comment by vb-8448 5 days ago
Comment by g-b-r 5 days ago
You need to use a tls intercepting proxy for that.
I couldn't find any ready-made tool unfortunately, there's tlsnotary.org but it seems far from simple.
Comment by vb-8448 5 days ago
I wonder how I can attach a timestamp that cannot be faked.
Comment by g-b-r 4 days ago
They could only claim that it's been faked by claiming that you stole their TLS private key.
Comment by amelius 4 days ago
https://en.wikipedia.org/wiki/Electronic_discovery
Specifically forensic web preservation or web capture tools.
PS: If they made it impossible for you to prove that you clicked a checkbox or not, then, logically, the burden of proof is on THEM.
Comment by g-b-r 4 days ago
> If they made it impossible for you to prove that you clicked a checkbox or not, then, logically, the burden of proof is on THEM
Even in Europe the only "proof" that's required to companies is a log or database entry (both easily manipulated); if a user strongly disputed to have done it and sued the company for that, maybe you'd able to obtain some investigation on their systems.
Comment by andsoitis 5 days ago
Or simply switch to a competitor. Assuming this is not just a bug, why would one stand for such disrespectful and sneaky behavior?
FWIW, I have not seen this happen for me.
Comment by _zoltan_ 5 days ago
the x20 Max plan for Claude gives you way lover limits.
Comment by dgellow 5 days ago
Comment by CuriouslyC 5 days ago
Comment by cma 5 days ago
I do think it ended up being after a power outage though and it was some kind of config corruption.
Comment by troyvit 5 days ago
Comment by dgellow 5 days ago
Comment by andsoitis 5 days ago
OpenAI and Anthropic are but two players. There are others too. You have choice.
Comment by andsoitis 5 days ago
lovers should never be limited!
Comment by cheeze 5 days ago
If you don't want them slurping your data, you're gonna have to pay more.
Same as it ever was.
Comment by cute_boi 5 days ago
Comment by dgellow 5 days ago
Comment by andsoitis 5 days ago
Comment by danillonunes 4 days ago
It's funny to think that back in the days we used to think the most evil thing a company could do was charging money for their software. Today we even praise some companies because they are "less evil" and they actually charge money, instead of giving the service for free and mining our data (e.g. Fast Mail vs Gmail; or Kagi vs Google).
Comment by jacquesm 4 days ago
They did a lot more than that.
Comment by DrammBA 5 days ago
Comment by hollerith 4 days ago
Good point, but "Sudden Human Extinction" would be too long a name.
Comment by nullc 5 days ago
Comment by samvher 5 days ago
What's kind of still an open question for me is if the toggle automatically also applies to my Codex CLI use on the same account, or if that data is still silently being used in some way.
After this happened, I deleted all my ChatGPT history (even though I'm not sure how much it helps at this point), but for Codex I still haven't really found any way to do the same, I can still load my past sessions even after archiving them.
Comment by the_duke 5 days ago
So it may or may not happen regularly, but I would not over-index on a sample size of one.
Comment by czk 5 days ago
Comment by aurareturn 5 days ago
Comment by enraged_camel 5 days ago
I went ahead and uninstalled the app. Won't be renewing.
Comment by jsw97 5 days ago
If this is true, though, then given the way their chat operates, this might be more dangerous than it seems.
One of the things I like about ChatGPT is its memory, the way it kind of seamlessly, but not excessively, ties back to earlier discussions. It's huge for usability (for me).
But this also means that you should expect that if "improve the model for everyone" becomes unclicked (leaving aside for a moment the fact that that is ridiculous) then they have a reasonable argument that your decision implies to all conversations. Because recall is part of their thing. So it's not just your chats going forward that are at risk. As soon as you see that unclicked, it's reasonable to expect that your history is irretrievably theirs now. You don't even have to think they are especially nefarious for this to be true.
Don't go toggling that switch on and off.
Comment by nullbio 5 days ago
Comment by samdhar 4 days ago
I've consistently refused to allow switching my privacy mode and despite their many iterations, they have never transgressed and kept me stuck on "legacy privacy mode"; their strongest privacy setting, which is not even available to select anymore. It requires that I keep my agents and cursor usage local and can't use their cloud agents (pretty neat cuz you can keep working with them on your phone). But so be it.
I am glad they don't secretly turn it on and instead keep nagging me to change it.
Comment by nerdyadventurer 2 days ago
Comment by rajeasy 1 day ago
Comment by mickelsen 4 days ago
Since I still need to use that account with Plus, I added a new card, flipped the switch back to no training, and submitted that thing on the privacy page which is a bit more formal. Probably should have signed up for a business account.
Comment by theredsix 5 days ago
Comment by neutrinobro 5 days ago
Comment by heliosAtwork 5 days ago
Comment by jerf 5 days ago
Now, obviously, considered as a whole, I think we're looking at "both". But when I wonder about specific cases like this one, that doesn't help.
Comment by pacificat0r 5 days ago
Comment by yearolinuxdsktp 5 days ago
- Autoplay transitioned to a per-device setting… suddenly default-on on every new device you log in to.
- Watching a video with computer-to-TV account connection? Automatic “TV queue,” a concept absent from the TV app, with incomprehensible behavior for how it’s used, so now videos are auto-played anyway.
- Watching a video from a playlist? Autoplay cannot be turned off.
Is it simply bad/absent product management and product design? Or is it actively user-hostile decisions meant to prop up view numbers and continue to have the users hooked on YouTube?
Comment by 27183 5 days ago
Either way, is this a company you want to trust with intimate secrets? "Oh, but they passed SOC2!" Lol.
Comment by cheschire 5 days ago
I actually found my setting was enabled today when I know it was disabled before, so I’m inclined to agree with OP. I just think it was worth mentioning that there are other cynical takes you seem to have left out.
Comment by philipov 5 days ago
Comment by antonok 5 days ago
Comment by gcr 5 days ago
Comment by Bluestein 5 days ago
Comment by garyfirestorm 5 days ago
Comment by Bluestein 5 days ago
https://www.anthropic.com/research/small-samples-poison?from...
Comment by RandyRanderson 5 days ago
One of the reasons I stopped using it.
Comment by mentalgear 5 days ago
Comment by utopiah 5 days ago
It is a big red flag.
Comment by gashmol 2 days ago
I check it every few months and they didn't re-enabled it for me.
Comment by jacobgold 5 days ago
Comment by jacquesm 5 days ago
Comment by jacobgold 3 days ago
You're making the assumption that it was "quite possibly malice" I'm making the assumption you might have gotten confused switching between accounts, as I have often been.
Comment by pan_lid 4 days ago
Comment by rtaylorgarlock 5 days ago
Comment by Tepix 5 days ago
Horrible.
Comment by jasonjmcghee 5 days ago
Not sure if that's region specific or something though.
Comment by MetroWind 4 days ago
Comment by Rebuff5007 5 days ago
Comment by deaton 5 days ago
Comment by nharziro 5 days ago
Comment by arpinum 5 days ago
Comment by esafak 5 days ago
Adds the highest level of account security by requiring stronger sign-in methods and applying stricter protections to help prevent unauthorized access.
That is not about model training.
Comment by arpinum 5 days ago
Comment by esafak 5 days ago
Comment by arpinum 5 days ago
> Automatic training exclusion. People working with especially sensitive information may opt not to have those conversations used for model training. With Advanced Account Security enabled, that preference is automatic: conversations from those accounts will not be used to train our models.
Comment by noname120 5 days ago
Comment by hokkos 4 days ago
Comment by throwatdem12311 5 days ago
Comment by opentokix 5 days ago
Comment by mkarrmann 5 days ago
Comment by Havoc 5 days ago
Comment by ionwake 5 days ago
maybe im too dumb to be worthy of a config update lol
Comment by general_reveal 5 days ago
Comment by stagsterlabs 4 days ago
Comment by jacquesm 4 days ago
Comment by snihalani 5 days ago
Comment by codeduck 5 days ago
Comment by kd913 5 days ago
Sure send them all your financial, personal data, they won't ever sell it on to the highest bidder for a new profit stream.
Using it as a therapist, financial advisor, health expert and blackboard has always been a terrible idea.
Comment by docheinestages 5 days ago
Comment by jacquesm 5 days ago
Comment by epsteingpt 5 days ago
Comment by tencentshill 5 days ago
Comment by thatmf 5 days ago
Comment by VCFundedGenYer 4 days ago
Comment by sunaurus 5 days ago
Comment by ALLTaken 5 days ago
Ironically he proved two major findings in Navier Strokes and that unethical American companies violate laws, steal your breakthrough findings & IP and then threaten you if you dare to challenge them.
This is making the status-quo so bad for any of us working on serious capacity. My client's don't trust ChatGPT/Claude anymore and prefer on-premises and OSS models or even custom trained models.
Comment by aenis 5 days ago
Did the researches opt out from data sharing on subsidised subs?
Did anyone prove that their methods enabled OpenAI models to produce the solution?
For a discussion about science, there is almost no scientifical method applied to proving anyone stole anything.
Comment by ball_of_lint 5 days ago
On the other side, the lack of evidence is pretty damning. Only OpenAI can try to prove that they came by these results legitimately, and the case they're making is quite weak. They could make public metadata about what their model was trained on and whether it did train on the conversations in question; they have not. TBQH I read it as even they don't know.
And regardless of whether the result is legitimately obtained by their model, they've not at all conducted themselves well throughout this story. They set out to scoop researchers based on a rumor. They threatened to ruin a mathematicians career. They put up a paper that deliberately doesn't cite the most relevant research, despite building directly on it. No matter how you look at it, OpenAI has and should lose any standing they had in the research community.
Comment by aenis 5 days ago
Both OpenAI and the researchers know if the sessions in questions were subject to data sharing. Why neither the scientists nor OpenAI is clear about that is weird - it would seem at least one party has the incentive to report that. But even if their sessions were in training data sets its hard to tell whether it influenced the outcome. Those models are big, but are they big enough to preserve subtle, niche techniques enough to draw from them while solving a related problem? Probably nobody knows.
Comment by keeda 5 days ago
Comment by ghostly_s 5 days ago
1. https://thenextweb.com/news/bubeck-navier-stokes-account-apo...
Comment by lxgr 5 days ago
If it's not the former, while certainly concerning, I don't see how that's relevant here (other than maybe in a very vague general sense of "entities doing immoral/illegal thing X are likely to also do immoral/illegal thing Y").
Comment by innocent_name 5 days ago
Comment by Insanity 5 days ago
Comment by skinfaxi 5 days ago
Comment by faangguyindia 5 days ago
Comment by innocent_name 5 days ago
Comment by innocent_name 5 days ago
I want to discuss actual facts, observations and not your shallow social signaling replies of 0 value.
Comment by dpz 5 days ago
Comment by innocent_name 5 days ago
Comment by egillie 5 days ago
Comment by dgellow 5 days ago
Comment by huijzer 5 days ago
Same holds for most big corps from what I can tell. If you really do a deep dive into the scandals over the years, you’ll probably see most have barely reached the news or if they did then it’s usually a quite bland criticism like “anti-competitive practices” or a poor HVAC at a certain factory. I think a lot more is hidden than we think
Comment by chrisjj 5 days ago
Worse, they use the product and still don't know.
Comment by spongebobstoes 5 days ago
Comment by Centigonal 5 days ago
- OpenAI is the first AI lab to pioneer ads in consumer AI
- Anthropic seemingly exists primarily because top OpenAI researchers lost faith in the company's commitment to AI safety
- They had the CEO drama in 2023, with evidence that suggests people in a position to know were doubtful of Sam Altman's honesty and motives
- They were tripping over themselves to kiss the ring after Anthropic got in a row with the Department of War over using AI for autonomous killing systems and surveillance of US citizens
- Altman's record (YC, Loopt, WorldCoin) and associates suggests he subscribes to the Paypal-Facebook "move fast, break things, find and exploit gray areas" school of company building
All of this suggests that OpenAI will optimize its own growth and power over consumer welfare or societal stability in the future (obviously, companies aren't a monolith and I'd love to be wrong).
Comment by chris_va 5 days ago
Comment by dgellow 5 days ago
- OpenAI allegedly directed ex-Apple employees to leak internal documents and allegedly coached the employees how to evade Apple security processes
- OpenAI allegedly lied to hardware companies working with Apple to use proprietary technology
- OpenAI allegedly copied Scarlett Johansson voice for ChatGPT after she declined to work with them
- OpenAI allegedly made ChatGPT more sycophantic to increase their retention rate, while aware of the risks. ChatGPT is linked to multiple suicides
- OpenAI ran thousands of agents on hacking problems, with close to no supervision, for months, with a harness that allows for full execution, resulting in the hack of HuggingFace infra AND OpenAI’s own infrastructure (the agents allegedly got fully root access to their k8s cluster). They weren’t aware of most of it until their investigation.
- OpenAI has been spreading misinformation regarding the capabilities of their technology for years
- OpenAI allegedly front-run researchers who are using the platform for their own personal research
There is way more, I don’t maintain a list of everything that happened over the past 3y or so
Comment by spongebobstoes 5 days ago
ads fund access for poor people. Ant has a different philosophy, not a better one. Altman was supported by >90% of employees. OAI DoW deal includes technical safeguards against misuse (missing from Ant deal). Loopt and WorldCoin dont seem terrible or even relevant, Altman doesn't even have OAI equity
I don't think OAI are the "good guys", but I don't see convincing evidence that they are terrible
Comment by henry2023 5 days ago
Comment by riffic 5 days ago
Comment by stavros 5 days ago
Comment by keeda 5 days ago
Comment by stavros 5 days ago
Comment by keeda 5 days ago
Comment by stavros 5 days ago
Comment by onel 4 days ago
Comment by altmanaltman 5 days ago
Comment by tmpsvc2695f5 3 days ago
Comment by tmpsvc2695f5 4 days ago
Comment by tmpsvc2695f5 4 days ago
Comment by kryzz-ai-bo 5 days ago
Comment by tmpsvc2695f5 5 days ago
Comment by ovi256 5 days ago
First one: "Settings > Data Controls > Improve the model for everyone > Switch off the toggle"
The other one is "OpenAI Privacy Portal" > "Do not train on my content" https://privacy.openai.com/policies?action=AUTOMATED_DECISIO...
Why are there several? You probably need an ML PhD to get it /s
Does using all of them achieve "do not train on my data"? I wish we could find out.
Do you think if you'd be in the race to train your own pocket God (as the CEOs of the frontier labs think they are) you'd let yourself be slowed down by such things as privacy duties?
Comment by tedsanders 5 days ago
There are a couple of reasons the privacy portal page exists in addition to the app settings. One reason is that it covers OpenAI products beyond ChatGPT/Codex (e.g., Sora). A second reason is that it provides functionality to logged out / non-users, like the EU right to be forgotten.
You don't need to toggle all of them to have your wishes respected. That would be a terrible design.
I work at OpenAI, but not on privacy. I am not their spokesperson. In my experience, we take a great deal of painstaking care to respect people's privacy, and we go well beyond our minimal legal obligations in doing so.
Comment by franzcoughka 5 days ago
Comment by buellerbueller 5 days ago