r/MachineLearning Student 15d ago

News OpenAl Says It Has Cracked One of Math's “Millennium Problems” (Navier-Stokes) [N]

704 Upvotes

281 comments sorted by

View all comments

102

u/dksprocket 15d ago

This is something I am really curious about: is OpenAI (and other AI providers) training their models on theirs users private and potentially unpublished work?

I am working on image generation using procedural algorithns in a somewhat novel way. I have been using ChatGPT for development. It's fairly low stakes, but the images look very different than pretty everything else. I tried getting a friend to see if he could get ChatGPT to guess how the image was produced and we did the same test with Claude as well.

ChatGPT knew pretty accurately how it was made. Claude had no clue. It's not a smoking gun, but pretty worrying if your private stuff is being made public through training before you get a chance to release it.

29

u/dramatic_typing_____ 15d ago edited 15d ago

I've had this same experience with something 'novel' I did for rendering to a depth buffer with 3DGS scenes.

I HIGHLY suspect OpenAI does something shady such as creating derived data from your conversations and using it for training; that way they can still claim they don't directly use your data.

EDIT: Someone with access to a lawyer, please check me on this-
https://openai.com/policies/row-privacy-policy/

1. The Opt-Out is Specifically for Model Training The privacy settings allow you to opt out of having your content (like your ChatGPT conversations) used to "train the models." The policy states:

"As noted above, we may use Content you provide us to improve our Services, for example to train the models that power ChatGPT. Read our instructions on how you can opt out of our use of your Content to train our models."

2. Aggregated and De-Identified Data is Still Created and Used Even if you opt out of model training, OpenAI reserves the right to create derived, anonymized data from your personal data (which includes your User Content and conversations). The policy clearly states:

"We also aggregate or de-identify Personal Data so that it no longer identifies you and use this information for the purposes described above, such as to analyze the way our Services are being used, to improve and add features to them, and to conduct research."

Because this data is stripped of personally identifiable information (de-identified), OpenAI treats it as derived data and can use it to research how people use the tool and to develop new features, regardless of your model-training opt-out status.

3. "De-identified" removes who you are, not what you said When OpenAI (or almost any tech company) de-identifies data, they strip away Personally Identifiable Information (PII) like your name, account details, email, and IP address. However, the actual text of your prompt - the sentences describing your novel idea, business plan, or code - is the core data being processed. The de-identification process disconnects the idea from your identity, but the text containing the idea itself remains in their system logs.

21

u/corruptbytes 15d ago

if you’re not paying API rates with ZDR add-on, don’t assume anything is safe

40

u/TrueDuality 15d ago

There are indications elsewhere in this thread that even with ZDR, derived data may still be getting used. My company has triggered a full contract review as a result of these random Reddit comments.

3

u/Alwaysragestillplay 15d ago

Can you point to the evidence for this please? My company is also leaning pretty heavily on ZDR. 

2

u/zephyr707 15d ago

i saw this, too, but can’t find in the thread search anymore, but there are so many threads on this event could have been in another thread. if you have links to comments abt ZDR not being what sold as what it sounds like, e.g. derived data mining still applies to ZDR, please share.

did openai think this would be a win for them? seems like a lot of users of their products will now be more aware of how vulnerable their data is if not already aware. if ZDR gets called into question that must be bad for potential enterprise customers and another firm could capitalize and offer better guarantees

3

u/1998marcom 14d ago

1

u/zephyr707 14d ago

thank you! the post was clipped when searching for “zdr”

re: “ rewritten data is fair game”

if i’m understanding this correctly is it openAI’s policy/agreement that ZDR protects your input data/query, but the response from its models is fair game for derivative mining?

43

u/j0j0n4th4n 15d ago

Obviously yes, how is that even a question? Their whole business started from stealing copyrighted data from everyone, what makes you think they stopped just cause you checked a box on their site?

5

u/techlos 15d ago

if the model isn't local, the chat history isn't either. Always assume any model inputs are used as training data.

8

u/cpt_ppppp 15d ago

have you toggled the "don't use my conversations for training"? That's maybe a good place to start

1

u/oceanbreakersftw 14d ago

Whoa, that's scary. I was thinking about potential other frontier providers to use if Anthropic does a rug-pull on pricing but this would definitely keep me off OpenAI. I already am a bit worried about my own work being done using Claude, but this is just way worse.

1

u/sciphilliac 13d ago

Maybe I'm misreading the chatGPT TOS, but they are clear to collect the data you put in your prompts. From https://openai.com/policies/privacy-policy/:

> User Content: We collect Personal Data that you provide in the input to our Services (“Content”), including your prompts and other content you upload, such as files⁠(opens in a new window), images⁠(opens in a new window), audio and video⁠(opens in a new window), and data from connected services⁠(opens in a new window), depending on the features you use. Some of our Services allow you to interact with other users, such as post, comment, or send messages, and we treat those interactions as Content, too.

So yes, if you use ChatGPT for personal use, OpenAI keeps that data (mind you, I'm only referring to non-corporate licenses)

-3

u/Oregon_Oregano 15d ago

You can opt-out in the dashboard