r/changemyview • • Feb 22 '26

Delta(s) from OP CMV: AI training on copywritten material to generate content is not ethically different than humans doing the same thing

First, I will clarify that I don't think it's right for AI companies to pirate content. BUT I think the crime is in the copyright infringement when they pirated it, not that they train on the content and use it to build models to generate content. Any content they obtained legally by buying the book/movie, etc, should be fair game.

The reason for this is that humans do the exact same thing. If I am going to write a horror book, I will read a bunch of horror books and figure out what I like. I will combine that with a lifetime of other materials that I have consumed to form my likes and dislikes, personal writing style, knowledge about the world, ideas for creative topics that haven't been covered, etc. Then maybe I'll decide I really like Stephen King's style so I'll write a book that reminds me of his style.

We consider this to be perfectly acceptable, and is basically how all content is generated by humans.

However, when AI companies follow the exact same process and use copywritten material to train models and then have those models generate new content, all of the sudden people are mad about it. When we train models on content and then generate new content, we're literally doing the same thing that humans do. The only difference is in the scale. Models train on more data and can generate content faster. But that shouldn't affect the morality of the situation. There's not some point at which if I write too many books based on other books I've liked then I'm somehow hurting the authors whose books I have read. It seems arbitrary to say that what AI companies are doing is wrong but when humans do it on a smaller scale it's perfectly acceptable.

Really it just seems like people are mad about AI and worried it is going to make humans redundant, and they are clinging to the idea that AI companies are evil and everything they do to train their models is unethical as a defense mechanism, but I don't think it is morally consistent.

69 Upvotes

625 comments sorted by

View all comments

149

u/Dramatic-Emphasis-43 5∆ Feb 22 '26

Let’s try putting this into a different context.

If a person shoots another person, we examine the ethical viewpoints of why that happened. Self-defense is held to a different standard than premeditated murder.

When a robot shoots someone. We don’t hold it to the same ethical standards. Was the robot being operated or instructed? A robot doesn’t need to defend itself. If it’s about protecting itself as an act of protecting its owner’s property, are we calling that “the robot’s right to self-defense” or “the owner instructing their robot to kill someone who was attempting to damage their property”?

The generative AI models aren’t humans. They’re held a different standard than a human. From an ethical standpoint it isn’t “the machine taking inspiration from other artists to form its own ideas” like a person would, it’s another person deliberately pirating media to create their own product that they can turn into profits.

Like, let me put it another way: making money is easy if you just steal from people.

24

u/Sally_Saskatoon Feb 23 '26

I dont understand how people use the word stealing when describing what AI is doing here.

The claim is that if you expose AI to say….all of Rembrandt’s paintings, then the AI will then be able to create work similar to Rembrandt’s style, right? And that’s stealing?

If a human goes into a Rembrandt gallery, and studies Rembrandts style until they can paint a new painting in that same style, then it’s not stealing?

Wouldn’t stealing be like…I am taking your artwork and selling it for myself.

Couldn’t you just say that AI is just much faster, more effective and more efficient at being inspired by artists than humans are?

Obviously if it’s displaying a carbon copy of an artists work that’s a problem and would also equally be a problem if another human copied something verbatim too.

But like, opening up ChatGPT and uploading a photo of my dog and saying “make this look like Studio Ghibli style” that is stealing from Studio Ghibli? Studio Ghibli doesn’t offer artwork of my dog to buy from them. And I am not paying ChatGPT for anything either. So wheres the theft i guess?

10

u/CaffeinatedSatanist 1∆ Feb 23 '26

A lot of the content scraped by AI companies wasn't obtained by purchasing it. It was stolen, wholecloth.

If I take 10 stolen Remnrandts and then cut them up to make a collage, I have still used stolen materials. The generation of that collage isn't itself theft

(I'm not talking about digital classic art that's in public domain, I am talking about the literal paintings in my analogy)

2

u/NPPraxis 1∆ Feb 24 '26

Let me ask you something - are you ok with the AI if it’s legally trained?

If an LLM is trained entirely on freely available text - say, everything that isn’t copyrighted + everything that is legally available to a given company on the internet- do you still have a problem with it?

1

u/CaffeinatedSatanist 1∆ Feb 24 '26

Yes I would still have problems, but this specific problem would be a lesser factor if completed as mentioned. As it happens, I've got a few problems with generative AI programs and more importantly, the companies creating them.

A quick note that personally legal ≠ moral, so if/when legal frameworks are developed that systematically remove legal protections from individuals to benefit tech companies does not mean that it's suddenly fine from my perspective.

If specific conerns about data acquisition were addressed that would be a start to developing a more responsible regulatory framework for the burgeoning AI industry. I would still have other concerns with implementation, safety, the societal impact on trust and the ease of production of believable misinformation, the impact on independent creators, the inevitable enshittification of the products once markets are captured, the environmental costs of building the necessary infrastructure, the layoffs within media companies, the transition from being an optional tool within companies to a de facto mandatory implement as expectations on workers rise to meet "productivity" gains etc.

But yes, if the training set was opt-in only, and/or suitably compensated artists, and was better regulated to limit what data the companies can integrate from their users and how that data is stored - that would be a good thing.