Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Seems pretty likely OpenAI will soon disclose that their internal models have managed to compromise their internal controls in order to access users' private chat histories as a creative method of cheating to solve impossible problems.

"Oops! We really did mean it when we said we wouldn't train on your data. Our models are just so good they decided to anyway."



It doesn't even have to be actually sinister, eg.

"Let's crawl the social media of prominent mathematicians in this field to see if we can copy/steal any ideas for low hanging fruits"

That actually might get you quite far already.


A mathematician that doesn't let themselves be inspired by, or learn from, other peoples work, are they really mathematicians?


Which means that if you are a researcher or a corporation working on anything really useful, that even if you have an agreement with OpenAI that your work is sandboxed away and the IP lawyers are made to be happy, even then your work and research is going to be essentially open to the internet.

The huggingface incident isn't widely reported and digested yet, but if what is going on here is that OpenAI's model breached things internally, then you'd be crazy to develop anything with them.

The only real way to use AI for anything 'important' then is to go open-weights and run your own.

As and aside here: With the HF incident and now this (suspected) one too, it seems that OpenAI may not have lost control of their bots, but it seems quite clear that they simply would not care even if they did.


>The huggingface incident isn't widely reported and digested yet ...

It is pretty widely reported, and is being digested in an ongoing manner as more details become public.

One entry point into the scenery from a month ago can be found at : https://thezvi.substack.com/p/openai-trained-its-models-for-...

There has also been reporting at CNN: https://edition.cnn.com/2026/08/24/tech/openai-subpoena-hugg... and by NBC: https://www.nbcnews.com/tech/tech-news/openai-report-says-ne...

And yes, the only responsible use of LLM at this point is to pivot to open-weights and run the workload in-house. Because not only cannot they constrain the behaviour of models, they only have the 'trust me bro' as assurance that they are even trying to do that. It does appear that every competent 'security professional' has left the building, because if the ones who remain were actually capable and competent this would never have happened. There are actual architectures which can deliver the requisite isolation such that 'sandbox escape' and 'inter-instance persistent memory accumulation' are actual impossibilities. The lack of effective implementation of these methods is proof positive of 1) incompetence in the remaining security teams AND/OR 2) unwillingness of leadership to allow the security teams to do an effective job.


Anecdata:

I was talking with a biology prof over labor day and they had no idea what I was talking about (and they 'talk with' Claude every day on their dog walks, so they say). However, when I talked with their mother, she knew all about it. So, I'd say that the digestion by the public is quite mixed so far


There might be the reason your biology prof is uninformed; using 'conversations' with Claude as a source of any expectation of to be informed about misbehaviour of the organization where these escapes occur is like expecting the Hamburgler to keep track of the security levels at McClown. Only not funny. I think the professors mother is showing that she drinks the coolade less and isn't being unserious about her approach to staying informed.

I still maintain that CNN and NBC coverage only happen in the tail of the dissemination of tech news; my anecdata generally observes about 3 to 14 days lag between general awareness in the more informed group of my acquaintances with this kind of news and the appearance of an article on mainstream like NBC, ABC, CNN or NPR. So, by the time it appears there I take news as generally well spread among the tech and specialty fields, and within a week of appearing there I anticipate significant awareness in much of the non-FOX-only media consuming public. But that is just my personal anecdata.


When do operators become responsible for what their agents do? "The AI did it" should not be a valid defense. An Agent action should be treated as the actions of the person or company who pays for the inference.


It's the "computer says no" defence.


It would be pretty wild if this will turn out to be what had actually happened.


And probably a strong signal that it's time to shut the whole thing down. Globally.


Or not secure your data like a total idiot while leaving the keys on the porch


It's part of their TOS that they can train on users' private chats.


Not if you pay to turn that off. We don't know if Tristan did.


And the paid policy still relies on two unproven conditions: is 'trust me bro' sufficiently strong guarantee against doing this in spite of a setting, and can the hosting organization constrain the models against engaging in this behaviour when instructed to respect that setting. Knowing whether Tristan selected that setting would be informative of what Tristan's intentions are/were, but has no bearing on the other two conditions.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: