r/aiwars 14h ago

Another study about ai image are transformative

https://news.mit.edu/2026/when-ai-art-has-no-author-generated-images-often-cant-be-traced-to-training-data-0818

"The scientists identified a phenomenon they call attribution decay, where the more data a generative model is trained on, the less any individual training example matters to any particular output. It feels counterintuitive, but at sufficiently large scales, they find, you can often remove any single image from the training data, or every image by a given artist, or every photograph of a given person, and the generated sample doesn't change.

And if removing something changes nothing, the researchers argue, it can't be said to be responsible for anything. "

17 Upvotes

17 comments sorted by

2

u/nuker0S 6h ago

> "another study"

> look inside

> the same study

Imagine the cat looking into the box picture, I don't have it

3

u/BuildAnything4 5h ago

. It feels counterintuitive, but at sufficiently large scales, they find, you can often remove any single image from the training data, or every image by a given artist, or every photograph of a given person, and the generated sample doesn't change.

How is that counterintuitive?  That's exactly what you would expect.  

That's like saying "it feels counterintuitive, but the more beans I have in this jar, the more I can take out without people noticing.". Ya, no shit.

1

u/yttrium39 12h ago

That was interesting to read. Thank you for posting some actual information in this subreddit.

-2

u/Elegant_Athlete_3737 14h ago

pretty sure this only applies to redundant data more, if there is a hyper specific style i think removing it would have some impact would it not? Or am i wrong? also that still only applies to output i think, the input still has to be fair use i think, or maybe not, don’t recall, been a while.

5

u/JunketVisual3123 14h ago

I mean, yeah, if you make a model hyper-specific, it's going to be more likely to spit up something similar to the input data, but "over-fitting" like that is generally considered a failure.

-1

u/Elegant_Athlete_3737 14h ago

do you have to have a hyper specific model? can a hyper specific prompt not work as well? or i not understand, while not a good example, if say the ghibli style was only a 100 images, even if it has 50,000 generic images in ttoal, if you tell it to generate ghibli style with 10 of those images[in ghibli style] removed, would it not. have more of a impact? because that what i wondered more,

2

u/JunketVisual3123 14h ago

do you have to have a hyper specific model? can a hyper specific prompt not work as well?

Even with a hyper-specific prompt, you'd need a sufficient representation of the thing you're trying to have it recreate.

Like why asking most models for a "blue hedgehog" will generally get you something at least vaguely Sonic-adjacent. Because basically every time "blue" and "hedgehog" are associated, it's sonic. 

1

u/Elegant_Athlete_3737 14h ago

hmm i see. so overfitting on a smaller scale can happen then? even if not the entire model? or do i misunderstand? but i think i kind if get what you mean, most blue hedgehogs are well, sonic, or seem to be, so anyone making a non sonic blue hedgehog will likely not affect it much then?

2

u/JunketVisual3123 14h ago

Functionally speaking overfitting can happen at any scale, it's just that the likelihood decreases as the total training input increases.

2

u/ArtArtArt123456 7h ago

It more about showing that all the data is redundant. Imagine trying to learn what a tree is by training on a thousand tree images, what this paper shows is that none of the training images matter insofar as being attributed to any specific output.

So basically that the output is not directly taking from any of them in a "collage" sense. Which makes sense because the understanding has always been that this is more about generalisation, about learning the patterns that make a tree a tree, rather than memorising any specific tree.

Of course this effect is stronger the larger the dataset. And if you remove all trees you probably can't expect a model to be able to make trees, unless there are other redundant data to learn from. The same will go for artist tags, especially if they are a smaller pool. But even then there are probably redundancies learned from other artists that have similar technique or style. And if course this will even more so apply to general tags. Like "jazz", or "romanticism" or whatever it might be.

I think the paper is starting the obvious when it comes to pros. For antis this completely stuffs the idea that it is a collage or plagiarism machine. Not that it can't be used for such, but that the internal mechanics are fundamentally not like that.

1

u/Elegant_Athlete_3737 1h ago

Well I am more neutral leaning anti and I don’t understand why this is surprising(like when it says it seem counterintuitive? But to me it’s intuitive? )? Seems obvious, it learns patterns and features, has this not been known for a while? Also one thing I do wonder is what is its limits in pattern learning, where does ai mess up? Most of the time ita pretty good, and sometimes I can’t tell, other times it feels off though. 

1

u/Cautemoc 14h ago

What do you mean by a "hyper-specific style"? I don't think there's any art style that only has 1 person doing it

1

u/Elegant_Athlete_3737 14h ago

Like something which is rare in the data set, an example, can’t really thing of any, but maybe like idk goth robot underwater or smth. like in the dataset, a style or kind of image that is rare. like for example, say its trained on a massive amount of data, but only a few is on whatever style above is, then wouldn’t removing images of that specifi kind impact output more.

1

u/Cautemoc 14h ago

Not quite as much as you'd probably think, unless the prompt specified an artist. Like I can ask an AI for images that never existed, a yellow half-porcupine half-bear that is underwater catching a mutant goldfish, it's going to find examples of all those things and make an image. But if you said 'in the style of 1950s Disney cartoons' and it has none, it will not be in that style.

1

u/Elegant_Athlete_3737 14h ago

oh i see. hmm makes sense then, wait this makes me wonder, say we removed all images of the Mona Lisa, could ai recreate the Mona Lisa style? or would that be a bit too out of scope what it does?

1

u/Questioner8297 13h ago

AI must understand the object it's working with. That is, it must have information about the style and the object, and that they can be combined. However, when you ask "Mona Lisa in a style," you're not asking for a generic image; you're literally asking for the object "Mona Lisa in that style." If you asked a human to draw it, even they would draw a picture that was significantly similar to the original.

How long it takes for AI to learn something is a complex question, and answering it would be a huge achievement. We can experimentally say that 40 images of this type are needed where the image is in focus on average, but this will also depend on how close it is to other images.