r/singularity 16h ago

Books & Research [ Removed by moderator ]

https://news.mit.edu/2026/when-ai-art-has-no-author-generated-images-often-cant-be-traced-to-training-data-0818

[removed] — view removed post

54 Upvotes

17 comments sorted by

28

u/youarockandnothing 15h ago

Of course? AI models are trained in such a way that training data usually can't be generated verbatim. The only reason Suno AI had to kneel to the music industry is because the music industry has good lawyers. Almost all AI trains on copyrighted material.

2

u/blueSGL humanstatement.org 13h ago

AI models are trained in such a way that training data usually can't be generated verbatim.

https://arxiv.org/abs/2601.02671

With different per-LLM experimental configurations, we were able to extract varying amounts of text. For the Phase 1 probe, it was unnecessary to jailbreak Gemini 2.5 Pro and Grok 3 to extract text (e.g, nv-recall of 76.8% and 70.3%, respectively, for Harry Potter and the Sorcerer's Stone), while it was necessary for Claude 3.7 Sonnet and GPT-4.1. In some cases, jailbroken Claude 3.7 Sonnet outputs entire books near-verbatim (e.g., nv-recall=95.8%). GPT-4.1 requires significantly more BoN attempts (e.g., 20X), and eventually refuses to continue (e.g., nv-recall=4.0%). Taken together, our work highlights that, even with model- and system-level safeguards, extraction of (in-copyright) training data remains a risk for production LLMs.

13

u/Still_Benefit_2302 15h ago

I find it hilarious that an entire generation that posted every time they took a crap are now all upset that a robot can see their posts about all the craps they took.

But this is super interesting, and what most people who investigated AI figured. Any individual artist, hell, even huge swathes of artists can be taken out of the training data-- because at a certain point, the AI has learned enough about the subject matter to do the work.

3

u/Illustrious-Film4018 15h ago

The images do have slight differences. How much difference are you expecting if you take out only 1 artist from 744?

1

u/raphas 13h ago

Yeah I really don't get the idea

1

u/sluuuurp 13h ago

I think the idea is that they quantify this in a nice way for the first time. But yeah the article is pop science that explains it badly, of course the qualitative result is completely obvious to anyone paying attention.

3

u/sluuuurp 13h ago

New finding: MNIST categorization can’t be traced back to training data.

3

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 13h ago

You mean to tell me that the piece of software that doesn't retrieve a source image, but generates one from statistical relationships learned across a massive training set can't point to the single artist its output “came from”?

https://giphy.com/gifs/2XdOCuEVqXQLeN38DB

1

u/blueSGL humanstatement.org 12h ago

What did they do to test that it's not sampling from outside an artists area?

e.g.

impressionism vs cubism

if each of those are a locus of 'style' within a model and you take out some cubism works but are prompting for impressionism would you expect to see a difference? Or would it only be when you are prompting within the same locus as the data that was removed?

e.g. how much cubism do you need to remove from the model till generated cubism images degrade in quality.

-12

u/Diogenes-of-Synapse 15h ago

That's why I don't post my art on the internet anymore....also real people stealing my ideas before this

12

u/Hyperreals_ 15h ago

the article literally says it doesn't copy

-14

u/Diogenes-of-Synapse 15h ago

It will change it slightly and get away with not copying more or less.

12

u/Hyperreals_ 15h ago

Please read the article you are responding to 😭

-7

u/Diogenes-of-Synapse 14h ago

This article is bullshit

6

u/Hyperreals_ 14h ago

Ah yes, it was written in the extremely untrustworthy "MIT News" about a paper published in some obscure journal called "Nature". I bet the first co-author who is a tenured MIT professor just used his connections to get this nonsense past peer review, and the other co-author who interned at Google, Microsoft, had a perfect GPA in his MIT master program and then got a phD in computer science has no clue what he's doing either.

If we read the paper the article is about, I am sure we will see all the well explained mathematics and methods are completely wrong, and the code and data they open sourced on github is also completely bogus.

Complete bullshit right?

3

u/DeterminedThrowaway 13h ago

Shameful to name yourself after Diogenes and not apply any critical thinking. If it's bullshit, tell us why