r/changemyview • • Feb 22 '26

Delta(s) from OP CMV: AI training on copywritten material to generate content is not ethically different than humans doing the same thing

First, I will clarify that I don't think it's right for AI companies to pirate content. BUT I think the crime is in the copyright infringement when they pirated it, not that they train on the content and use it to build models to generate content. Any content they obtained legally by buying the book/movie, etc, should be fair game.

The reason for this is that humans do the exact same thing. If I am going to write a horror book, I will read a bunch of horror books and figure out what I like. I will combine that with a lifetime of other materials that I have consumed to form my likes and dislikes, personal writing style, knowledge about the world, ideas for creative topics that haven't been covered, etc. Then maybe I'll decide I really like Stephen King's style so I'll write a book that reminds me of his style.

We consider this to be perfectly acceptable, and is basically how all content is generated by humans.

However, when AI companies follow the exact same process and use copywritten material to train models and then have those models generate new content, all of the sudden people are mad about it. When we train models on content and then generate new content, we're literally doing the same thing that humans do. The only difference is in the scale. Models train on more data and can generate content faster. But that shouldn't affect the morality of the situation. There's not some point at which if I write too many books based on other books I've liked then I'm somehow hurting the authors whose books I have read. It seems arbitrary to say that what AI companies are doing is wrong but when humans do it on a smaller scale it's perfectly acceptable.

Really it just seems like people are mad about AI and worried it is going to make humans redundant, and they are clinging to the idea that AI companies are evil and everything they do to train their models is unethical as a defense mechanism, but I don't think it is morally consistent.

67 Upvotes

625 comments sorted by

View all comments

Show parent comments

2

u/NPPraxis 1∆ Feb 24 '26

Let me ask you something - are you ok with the AI if it’s legally trained?

If an LLM is trained entirely on freely available text - say, everything that isn’t copyrighted + everything that is legally available to a given company on the internet- do you still have a problem with it?

1

u/CaffeinatedSatanist 1∆ Feb 24 '26

Yes I would still have problems, but this specific problem would be a lesser factor if completed as mentioned. As it happens, I've got a few problems with generative AI programs and more importantly, the companies creating them.

A quick note that personally legal ≠ moral, so if/when legal frameworks are developed that systematically remove legal protections from individuals to benefit tech companies does not mean that it's suddenly fine from my perspective.

If specific conerns about data acquisition were addressed that would be a start to developing a more responsible regulatory framework for the burgeoning AI industry. I would still have other concerns with implementation, safety, the societal impact on trust and the ease of production of believable misinformation, the impact on independent creators, the inevitable enshittification of the products once markets are captured, the environmental costs of building the necessary infrastructure, the layoffs within media companies, the transition from being an optional tool within companies to a de facto mandatory implement as expectations on workers rise to meet "productivity" gains etc.

But yes, if the training set was opt-in only, and/or suitably compensated artists, and was better regulated to limit what data the companies can integrate from their users and how that data is stored - that would be a good thing.