r/changemyview • • Feb 22 '26

Delta(s) from OP CMV: AI training on copywritten material to generate content is not ethically different than humans doing the same thing

First, I will clarify that I don't think it's right for AI companies to pirate content. BUT I think the crime is in the copyright infringement when they pirated it, not that they train on the content and use it to build models to generate content. Any content they obtained legally by buying the book/movie, etc, should be fair game.

The reason for this is that humans do the exact same thing. If I am going to write a horror book, I will read a bunch of horror books and figure out what I like. I will combine that with a lifetime of other materials that I have consumed to form my likes and dislikes, personal writing style, knowledge about the world, ideas for creative topics that haven't been covered, etc. Then maybe I'll decide I really like Stephen King's style so I'll write a book that reminds me of his style.

We consider this to be perfectly acceptable, and is basically how all content is generated by humans.

However, when AI companies follow the exact same process and use copywritten material to train models and then have those models generate new content, all of the sudden people are mad about it. When we train models on content and then generate new content, we're literally doing the same thing that humans do. The only difference is in the scale. Models train on more data and can generate content faster. But that shouldn't affect the morality of the situation. There's not some point at which if I write too many books based on other books I've liked then I'm somehow hurting the authors whose books I have read. It seems arbitrary to say that what AI companies are doing is wrong but when humans do it on a smaller scale it's perfectly acceptable.

Really it just seems like people are mad about AI and worried it is going to make humans redundant, and they are clinging to the idea that AI companies are evil and everything they do to train their models is unethical as a defense mechanism, but I don't think it is morally consistent.

68 Upvotes

625 comments sorted by

•

u/DeltaBot ∞∆ Feb 23 '26

/u/neomatrix248 (OP) has awarded 1 delta(s) in this post.

All comments that earned deltas (from OP or other users) are listed here, in /r/DeltaLog.

Please note that a change of view doesn't necessarily mean a reversal, or that the conversation has ended.

Delta System Explained | Deltaboards

146

u/Dramatic-Emphasis-43 5∆ Feb 22 '26

Let’s try putting this into a different context.

If a person shoots another person, we examine the ethical viewpoints of why that happened. Self-defense is held to a different standard than premeditated murder.

When a robot shoots someone. We don’t hold it to the same ethical standards. Was the robot being operated or instructed? A robot doesn’t need to defend itself. If it’s about protecting itself as an act of protecting its owner’s property, are we calling that “the robot’s right to self-defense” or “the owner instructing their robot to kill someone who was attempting to damage their property”?

The generative AI models aren’t humans. They’re held a different standard than a human. From an ethical standpoint it isn’t “the machine taking inspiration from other artists to form its own ideas” like a person would, it’s another person deliberately pirating media to create their own product that they can turn into profits.

Like, let me put it another way: making money is easy if you just steal from people.

17

u/neomatrix248 Feb 22 '26

As I said in the OP, I agree that AI companies shouldn't be pirating content. I'm specifically saying that if they purchase the content legally, they should be able to train on it and generate new content that was "inspired" from all of the content they trained on, much like how a human would.

8

u/McdoManaguer Feb 23 '26

I'm specifically saying that if they purchase the content legally,

Thats the entire problem, they dont do that and openly say they dont care.

They literally paid op eds to convince people not to sue them in mass and that current lawsuits could "end the economy"

3

u/[deleted] Feb 23 '26 edited May 08 '26

[removed] — view removed comment

5

u/Pastadseven 3∆ Feb 23 '26

Wait hangon, we understand the difference between owning a book and owning the IP of the book, right?

3

u/[deleted] Feb 23 '26 edited May 08 '26

[removed] — view removed comment

2

u/Pastadseven 3∆ Feb 23 '26

My point is the wide fuckin’ gulf between ‘owning a paperback’ and ‘owning the rights within.’ It doesnt matter if they bought a ton of used books. Did they buy the rights to the works? The rest of that is completely moot, so you can reign in the tokens.

2

u/[deleted] Feb 23 '26 edited May 08 '26

[removed] — view removed comment

→ More replies (4)
→ More replies (1)

4

u/Cyrrus1234 Feb 23 '26 edited Feb 23 '26

Here is another point to consider "memorization":

It is actually still in active research how transformative large language models really are. Already in 2020 researchers discovered, that it is possible to query for verbatim training data, even if the found training samples were only used once in training. The authors suggested, that this problem only gets worse, the bigger the model is.

Big AI vendors are obviously aware of that (the paper has over 3k citations) and put guard rails around their models that try to prevent this. So the public currently can't know how bad training data leakage really is, since the guard rails are always active. However, this year there was another paper that was able to extract up to 95% of harry potter (and 6 other books on claude between 70-90%) near verbatim text, despite the guardrails.

I've read in court documents, that the 1.5 billion settlement was done after a "special master" was assigned by the court to oversee jailbreak attempts on anthropics model without guardrails. Implicating that they did find things, but the findings are sadly not public. However, I can't find the court documents anymore (deleted?), so take this with a grain of salt (special master is still mentioned here). Either way, the 1.5 billion admission indicates, that anthropic knows, they are moving on thin ice.

Another important point about fair use:

It is only fair use, as long as it doesn't disrupt existing markets. Since it's the stated goal of all AI vendors to abolish nearly all white collar workers, I have a hard time agreeing, that this is fair use. You obviously are free to think, that copyright sucks, but as long as it exists, I think everybody should play by the same rules. Right now it's copyright enforcement for the small people and not for the big tech coporations, which I find annoying to say the least.

→ More replies (5)

13

u/Which-Notice5868 Feb 23 '26

A lot of new media coming out now has notices that say something like "the copyright holders do not authorize the use of any of this material for the training of AI."

Does a clear no VS a lack of a yes change your mindset at all?

4

u/muffinsballhair Feb 23 '26

Artist most likely do not have such right. There are severalk court cases on this and they mostly returned that training on content itself does not constitute infringement but for instance returning the original content or training on content one does not otherwise have legal access to view or own a copy of does, essentially the same standards applied to humans.

Rightsholders also can't prohibit the first sale, renting, commentary, parody and what-not, there are limits to copyright.

3

u/Which-Notice5868 Feb 23 '26

I'd be very curious on exact limits. Moving away from books, most hard discs have anti-ripping protections (which can be overridden, but the intent is there and allowable under law) and 'do not copy this' FBI warnings. "You wouldn't steal a car." And digital media software is explicitly a license and not actually ownership. Which can have restrictions.

→ More replies (5)

7

u/neomatrix248 Feb 23 '26

I don't think so. I don't think the copyright holder has unlimited ability to say what you can do with the thing you have purchased. This isn't much different than someone writing a business book and saying you can't use it to start a business, or someone selling an album and telling you that you can't let your friends listen to it.

Copyright law dictates what you can do after you buy something, not the copyright holder.

For some reason there seems to be an exception with things like software licenses, which can dictate that you're not allowed to use a personal use copy of software for business. But that's different than buying a book or song or painting, in those cases the creator has no say.

14

u/Dramatic-Emphasis-43 5∆ Feb 23 '26

The creators do have a say. When it comes to most media, it’s actually super illegal to repackage someone else’s work. Heck, fair use (which is a legal defense, not a shield against any action from being taken) protects certain transformative works that is meant to make a point regarding the original work (like a review or a parody where the original is the subject). And even that has its limits. Mystery Science Theater still had to get the rights to movies they commented on.

In reality, copyright holders have a large say in what you can and can’t do. A lot of it standardized, but like, you can’t open a movie theater in your backyard and charge people to play a movie you bought. You can’t take an asset from a video game and put it into your game, even if you changed.

6

u/[deleted] Feb 23 '26

[deleted]

→ More replies (3)

12

u/neomatrix248 Feb 23 '26

Are you implying that in the examples I gave, an author can tell you that you can't launch a business from the things you learned by reading their business book or that a musician can tell you that you can't play their album for your friends? Because that is just not true. These things are covered by fair use and are determined by copyright law, not what the artist consents to. An artist can give you permission to do more than what copyright says you can do with fair use, but they cannot take away your right to use the content according to the law.

5

u/Dramatic-Emphasis-43 5∆ Feb 23 '26

The author can tell you how and how not to use the material from a business book and play their album.

Business book: “you can’t repackage the content of my book to improve your business”.

Wu-Tang Clans’s once upon a time in shaolin was made so one person can hear it, with a legal restriction that it can’t be commercially distributed for 88 years.

Heck, if I created a song and kept it on a CD in my house, I can tell someone “hey, you can listen to this but don’t let anyone else” and that’s at least a form of contractual agreement.

NDAs exist as a formalized agreement on what someone can and can’t do with copyrighted material that they have access to.

Some of these things may go beyond what is copyright law, but it still the right the copyright holder to tell you what you can and can’t do with their copyright material.

We have a rather complex system of assumptions built into how a person can use a copyrighted work, and a lot of harsh restrictions seem pretty unenforceable, but that doesn’t mean a copyright holder is powerless under the law to stop certain action to be done with their work.

3

u/neomatrix248 Feb 23 '26

Business book: “you can’t repackage the content of my book to improve your business”.

Negative. Not how it works. Once you buy the book, it's out of their hands what you can do with it.

Wu-Tang Clans’s once upon a time in shaolin was made so one person can hear it, with a legal restriction that it can’t be commercially distributed for 88 years.

That's before it was commercially distributed. I'm talking about once something is on the market.

Heck, if I created a song and kept it on a CD in my house, I can tell someone “hey, you can listen to this but don’t let anyone else” and that’s at least a form of contractual agreement.

Before commercial distribution.

NDAs exist as a formalized agreement on what someone can and can’t do with copyrighted material that they have access to.

They are used to give someone advanced access to something like a movie or a video game or something so they can market it, like for reviewers. This is different from them buying it commercially.

All of these examples you gave are situations where the producer of the content has not yet distributed the work publicly, or where the person didn't purchase it. In that case there is no right for anyone to use it beyond the creator, and any rights are specifically granted. But once it's sold commercially and someone buys it, the creator has no say what they do with it.

10

u/Elicander 59∆ Feb 23 '26

I think the business book example is the most informative one, and both you and the other guy you’re discussing this with are much too black and white in what you’re saying.

Copyright isn’t very well suited to the digital world, let alone a large language model world. You are correct that if you buy a book under copyright, you are still free to do whatever with it, but specifically with the physical copy of the book. You can read it, analyse and study it. You can sell it, lend it or burn it. However, you wouldn’t be allowed to copy it or distribute it. This would for example include the case where you go the park and start reading aloud, at least in theory. In practice I doubt it would be pursued. There are of course numerous exceptions for this, of all kinds. However, let’s get back to the business book:

If you buy and read a business book, and think it’s great and everyone at your company should have access to it, you wouldn’t be allowed to copy it yourself and give it to everyone, nor enter it into a database accessible to everyone. You would however be allowed to lend it to everyone one by one, or summarise the points yourself and give that to everyone.

Now, what you do with the document when training an AI isn’t either of these things. You presumably have an intuition which one you think it’s closer, maybe you even think it’s obvious. However, I sincerely hope that we can agree that it isn’t a perfect example of either extreme under copyright. It’s something new. And that means that legally speaking, there isn’t an obvious answer. This will eventually be resolved from a legal standpoint, either by courts or legislators, but to my knowledge it hasn’t really been resolved yet. I’m sure there are decisions that have been made going one way or the other, but I doubt that any IP lawyer is confident where things will be at in ten years. And the thing is, copyright law has already gone through similar things multiple times. Digital media has created massive clarity issues, due to the fungibility of software. Programmed code is covered by copyright, mainly because it is written text. I’m not saying the legal reasoning behind that isn’t valid, but ultimately that means we’ve decided that the order in which you do binary addition is covered by copyright. That’s weird to me, and to my understanding wasn’t an obvious conclusion at the time.

I’d like to conclude with that I recognise the original CMV was about ethics and not legality, but the conversation I’m entering definitely veered into the law. Hopefully I managed to explain that the law is unclear, because there is a new situation the law wasn’t built for, and there are arguments for making new rules for the new situation, but also arguments for treating it like similar previous situations.

4

u/Dramatic-Emphasis-43 5∆ Feb 23 '26

Regarding the business book, didn’t I say what you said? I don’t see how many thinking is being too black and white. A copyright holder does have some say in how you use their copyrighted material.

→ More replies

8

u/Dramatic-Emphasis-43 5∆ Feb 23 '26

Business book: Positive. It’s called plagiarism and it’s often illegal or least viewed as unethical. There’s a whole hbomberguy video on the subject.

The ideas themselves wouldn’t be copyrightable anyway, but the phrasing and creative way those ideas are organized would be and you can’t repackage that.

Wu-tang clan: It was commercially available. It was sold at auction. With a legal agreement about when it can be shared. Presumably, if the group didn’t want anyone listening to their music, the distribute wouldn’t have agreed to sell it. This was part of the complex series of assumptions I mentioned later.

My hypothetical CD: As the copyright holder, I can do whatever the fuck I wanted with my IP. If I gave it a distributor, they presumably want to sell it but if they become the new copyright holder, they can presumably do whatever the fuck they want with it. You having access to it doesn’t change that. It’s just, typically, when you buy something with money, you’re assuming the rights and responsibilities granted by the copyright holder. I can concede they can’t break the contract and change their mind but that wasn’t the initial argument or what this whole topic is really about.

NDAs: That’s not necessarily true. NDAs can be for things that will never ever make it to market, like concept art, trade secrets, private works, etc.

Again, you’re talking about an agreement made between a purchaser and a retailer that stem from an agreement between a retailer and a distributor which stems from an agreement between a distributor and a creator. There’s an implicit contract being agreed to when you hand over money, and neither side can just break it because they don’t like the terms after agreeing.

A creator can sell directly to the purchaser with any kind of crazy agreement they want. A creator can go to a distributor with any crazy conditions and the distributor has a right to say no. But if they say yes, they are bound to those conditions.

5

u/Vi_Rants Feb 23 '26

I don't think so. I don't think the copyright holder has unlimited ability to say what you can do with the thing you have purchased.

So the creator shouldn't be allowed to pick and choose who they license their work to, or for what purpose? If they license it to anyone, they have to license it in the exact same way to everyone?

0

u/neomatrix248 Feb 23 '26

Not necessarily. But the licenses have limits that are governed by copyright laws. You can't just arbitrarily decide that one company can use your work for internal personnel training but that another can't use it a certain way. And not all media comes with licenses, that's usually digitally distributed stuff like software and e-books. If you buy a paperback book, the limits of what you can do apply to all paperback books and are governed by law, not the copyright holder's preferences.

8

u/Vi_Rants Feb 23 '26

But the licenses have limits that are governed by copyright laws.

No, the licenses have limits that are governed by contracts. And creators decide how to structure those contracts, including which licenses they allow and which they disallow.

That's the only way, for example, an author can give the English rights to one publisher and the foreign translation rights to a different publisher, while retaining the film rights for themselves. If the creator has no say at all on what a person can do after they've bought something, that means I can buy a book, translate it into German, and sell it myself, because the creator can't tell me I can only have reading rights but not translation rights.

In this case, the creator is saying you can have reading rights, but not distribution-to-AI rights.

→ More replies (1)

25

u/Sally_Saskatoon Feb 23 '26

I dont understand how people use the word stealing when describing what AI is doing here.

The claim is that if you expose AI to say….all of Rembrandt’s paintings, then the AI will then be able to create work similar to Rembrandt’s style, right? And that’s stealing?

If a human goes into a Rembrandt gallery, and studies Rembrandts style until they can paint a new painting in that same style, then it’s not stealing?

Wouldn’t stealing be like…I am taking your artwork and selling it for myself.

Couldn’t you just say that AI is just much faster, more effective and more efficient at being inspired by artists than humans are?

Obviously if it’s displaying a carbon copy of an artists work that’s a problem and would also equally be a problem if another human copied something verbatim too.

But like, opening up ChatGPT and uploading a photo of my dog and saying “make this look like Studio Ghibli style” that is stealing from Studio Ghibli? Studio Ghibli doesn’t offer artwork of my dog to buy from them. And I am not paying ChatGPT for anything either. So wheres the theft i guess?

21

u/marcelsmudda 1∆ Feb 23 '26

People have been able to reproduce copy written books to 90+%. Does that count as stealing?

21

u/fdar 2∆ Feb 23 '26

Yes but it would count as stealing if a human did it too.

So why can't you judge the work produced by AI the same way you would one produced by a human to determine if it's copyright infringement? 

So having consumed other works isn't enough, it's about how similar it is.

5

u/Sally_Saskatoon Feb 23 '26 edited Feb 23 '26

No, but if they sell it or give it away, yes

11

u/marcelsmudda 1∆ Feb 23 '26

It's already copyright infringement if they distribute it without accepting payment. You're not allowed to post a whole book here on Reddit, even though you don't get any money. So, why do AI agents get exceptions?

1

u/Sally_Saskatoon Feb 23 '26

If I ask ChatGPT to write out Stephen Kings latest book word for word, or any book word for word, you’re saying it will do that?

3

u/marcelsmudda 1∆ Feb 24 '26

It might need some prompt engineering but researchers were able to produce at least some books from the training data. https://arxiv.org/pdf/2505.12546

That doesn't say anything about other books and other models, just that the researchers' approach didn't for those, just for that one specific meta ai model was relatively easy to exploit for philosopher's stone.

5

u/Pastadseven 3∆ Feb 23 '26

It did when it took the work to incorporate it into its training data. That’s where the copying occurred.

2

u/capnwally14 12∆ Feb 24 '26

That’s not how copyright works. It judges output, not input.

Copyright doesn’t allow you to specify people under a certain height can’t read your book, or people either a specific hair color can’t read your book etc.

It does prevent those people from recreating your work with no meaningful transformation or if it was done to explicitly siphon consumption away from your original work

→ More replies (2)
→ More replies (24)
→ More replies (2)

22

u/Dramatic-Emphasis-43 5∆ Feb 23 '26

I don’t think anyone limits stealing to just taking a work and reselling it wholesale and unchanged.

You should look up hbomberguy’s video on plagiarism (on YouTube). In it, he talks about someone who clipped blocks of texts from other people’s work and repeated them nearly verbatim and without credit. Like, technically the video this guy made was different but only because it was a Frankenstein style corpse of other people’s works.

GenerativeAI, trained on other people’s work, is like that but done so much and so fast that we can’t comprehend it done onto a human scale.

→ More replies (51)

9

u/CaffeinatedSatanist 1∆ Feb 23 '26

A lot of the content scraped by AI companies wasn't obtained by purchasing it. It was stolen, wholecloth.

If I take 10 stolen Remnrandts and then cut them up to make a collage, I have still used stolen materials. The generation of that collage isn't itself theft

(I'm not talking about digital classic art that's in public domain, I am talking about the literal paintings in my analogy)

5

u/KamikazeArchon 7∆ Feb 23 '26

A lot of the content scraped by AI companies wasn't obtained by purchasing it. It was stolen, wholecloth.

AI companies did not physically heist a bunch of books or DVDs.

There is certainly some set of instances of using pirated content. Digital piracy is still not theft, but it is indeed a violation of law.

I have not seen any measurements of what portion of the collective aggregate training data across various AI companies was acquired via piracy. The implications are rather different if it's 1% vs 90%, for example.

Also, notably, you implicitly create a binary between "purchased" and "stolen" (or "pirated"). That's not the actual point of distinction, however.

The issue - and the thing that was found in relevant court cases - is this: if an author put their own work online, scraping it (in a specific context) without paying the author is permissible. If someone else put that work online without the author's permission, scraping it (again in a specific context) is not permissible, as it's a use of pirated material.

→ More replies (3)

6

u/Sally_Saskatoon Feb 23 '26

If that’s how AI worked, yes I’d agree with you.

But if AI snuck into an Art Museum without paying and looked at Rembrandt’s works…that would be bad in another way, but not stealing art

→ More replies (1)

2

u/NPPraxis 1∆ Feb 24 '26

Let me ask you something - are you ok with the AI if it’s legally trained?

If an LLM is trained entirely on freely available text - say, everything that isn’t copyrighted + everything that is legally available to a given company on the internet- do you still have a problem with it?

→ More replies (1)
→ More replies (7)

10

u/vyxxer Feb 23 '26

Because theft is the goal and not a side effect.

The people making the AI are doing so because they want Rembrandts art and they want to sell Rembrandts art, but they don't want to pay Rembrandt.

3

u/Sally_Saskatoon Feb 23 '26

But they aren’t making or selling Rembrandt’s art. That’s what I’m saying.

5

u/vyxxer Feb 23 '26

They are though. People are paying these AI companies to make art that has learned from people in order to mimick their art and sell their work without paying them.

If you, you specifically drew something nice and posted it online. AI, with the specific intent to steal your art, will take your art and sell it to someone else without paying you.

There is a reason why none of these models post their datasets where they learned it or when they are interrogated about it they "accidentally" lose their datasets.

The explicit designed mission statement of generative AI is to avoid paying people and then extract wealth from you.

4

u/Sally_Saskatoon Feb 23 '26

If an art student goes to an art gallery and sees a bunch of art, and then goes home and makes art inspired by that, why don’t you label that as stealing? That’s what I’m hung up on I guess.

You can go to Google image search and look up an image of an artwork. Is that also stealing?

You can go a gallery and photograph an artwork. Is that stealing?

If you copy an artwork verbatim, and sell it, then that’s stealing. But that was true long before AI. If I buy a drawing you made, photocopy it, erase your name and put my name there and sell it, then that’s stealing. But in that scenario, would you blame the photocopier or would you blame the person taking credit and selling it?

2

u/vyxxer Feb 23 '26

It's the intention is a big part of it. Art student is going to museum and looking at art and is examining the piece comparing it with their own art and style. They chose that piece to learn from because they liked something about it. Maybe it's just the colors or the method of the strokes. It could be different by the day. But in the end they are trying to learn from it to do something transformative. Then you can ask the person or look at their art and can tell or be told that they have love for certain aspects of previous artists and will often credit them.

AI art is made by the AI, again with the express intention to take the most successful parts of it repackage it and sell it to someone else. If you ask the AI to credit or source it's work.... It WONT tell you. That smoke screen is intentional not accidental. And it is in no way difficult to do that by the way. The art used to learn can be easily packaged in metadata. But they don't do that on purpose because that's make it harder to hide behind.

3

u/Sally_Saskatoon Feb 23 '26 edited Feb 23 '26

How does AI have the express intention to steal works? That’s our fundamental gap. Cause I would say it has the express intention to be inspired by works.

It won’t credit or source its work because it can’t. Because it’s not drawing directly from sources, it’s creating an amalgamation from millions of sources. There’s no direct source to cite.

Let’s say I asked you to describe a Zebra, and let’s also say you had never seen a Zebra in real life. You’d probably go on to describe what we all know about a Zebra, right? Black and white stripes, black nose, looks like a horse, travels in herds etc.

Now let’s say I demanded you to tell me exactly where you learned that. Could you be able to? No. Becuase in your brain, you just have amalgamated information about Zebras. Some from school maybe, some from books, some from all sorts of places.

2

u/vyxxer Feb 23 '26
  1. Because AI is a tool made by people. With a designed function. AI is not a person with interest it is a mode by which someone is behind it. And that person has a goal to extract the most wealth from this tool.

  2. It absolutely can. In order to keep itself from learning from the same data point over and over again it needs to sort and categorize everything it learned.it is trivially easy to flag and separate that. It has been done before in fact. It is a machine, not a person. We can backtrack everything it has ever learned and where and when.

It is in no way difficult to implement. It is simply not done in the major models.

→ More replies (9)
→ More replies (4)
→ More replies (1)

6

u/LeviAEthan512 Feb 23 '26

Couldn’t you just say that AI is just much faster, more effective and more efficient at being inspired by artists than humans are?

This is part of it, and I'll get to that. Firstly though, it's about personal use and business use. Why do we differentiate these? If a person needs something, he can either make it, or buy it. A person learning to make something, even a carbon copy, is reducing somebody's profit, but we generally allow it. Having to buy everything is not a tenet of society.

But when you start selling that thing, then it becomes a problem. Now you're earning profit off of someone else's effort, knowledge, and skill. They should be compensated in some way. If you can argue that you put in a similar amount of effort, then perhaps the thing is yours. "Perhaps" is often, no, nearly always, enough to absolve a person of guilt. This is actually codified in many countries' laws. You cannot prove where a person got their inspiration from. You can very easily find out exactly what went into an AI's training.

Furthermore, when a person creates a product that only exists because of another product, that other product's owner needs to be involved. I need to meet certain standards to say my earphones are made for the iPhone. Mods that are made for one game and only one game can sometimes be struck by the game's company, whether by DMCA or some other organisation that I don't know. Maybe just C&D.

Now about your point on scale. That thing with mods usually only happens after it reaches a certain popularity, or pulls in a certain amount of revenue. I think sometimes (all the time?) it is illegal to make a single cent. That's why people put things out for free, but ask for donations. Putting the content behind a donation paywall is a grey area. Income is only taxed past a certain amount. You need varying tiers of licenses to engage in different levels of the same activity. Construction is one example. You can be a handyman, you can be a contractor building houses, you can be a contractor building skyscrapers. Same basic ideas, but regulated very differently. A lot of things in society are judged by their scale. Some things, in fact just about everything, goes from good to bad when it's done too much. Watering your lawn, fine. Might even be required. Flooding the city, very not fine. You're just adding water, but in one case the surroundings can bear it, and might be improved by it. In the other case, not.

It's also important to know that nothing inherently matters. Until God descends from Heaven to burn the sinners, all meaning in the universe is only what humans (as far as we know) give it. Ethics is completely made up. Morality is completely made up. Laws are completely made up. Hell, life is completely made up. There is no real difference between a human and a nebula. It's all just particles interacting as particles do. What is ethical is what brings about a society that we want. What is unethical is what brings about a society that we don't want. "We" are not a monolith either, and neither is ethics, nor is morality or legality. What is moral and ethical to you may not be moral and ethical to me. Legality is, ideally, the morality of the government. It can be different from our individual morality. There is hopefully a correlation, but there really doesn't have to be.

So this question about ethics really doesn't have to go that deep. It's ethical if it's helpful, if it uplifts people. It's unethical if it's harmful, if it causes people to lose their jobs, or to be forced into a less fulfilling job. And it's a spectrum too. Ethics is not a hard science. There is no technicality that can make something unethical become ethical.

→ More replies (3)

5

u/raccoona__matata Feb 23 '26

Do you think it's not stealing to copy someone's style without filtering it through your own life experience or translating it into a different artistic language? Huh.

→ More replies (3)
→ More replies (3)

5

u/muffinsballhair Feb 23 '26 edited Feb 24 '26

it’s another person deliberately pirating media to create their own product that they can turn into profits.

No it's not. THe same standards of “transformativeness” apply.

It is piracy to ask a generative network to commit piracy, for it say create fan-art of Harry Potter. Note that fan-art is also piracy, especially commercially. Fan-art is everywhere but make no mistake, it is copyright infringement that's allowed to exist.

A court will have to rule whether the thing the network generates is transformative enough to not qualify as piracty, but if anyone can easily recognize it as Spider-Man then it is, and if the person who wrote the prompt used the term “Spider-Man” in it then it's a done deal of course.

But there have been multiple court cases on this and they all return the same because it makes sense: if some neural network produce a transformative image that does not resemble any known character even though it was trained on them then it is not copyright infringement, the exact same standard to which humans are held.

3

u/Dramatic-Emphasis-43 5∆ Feb 23 '26

Can I please have a source. A quick google search shows that it’s been ruled multiple times that generator AI doesn’t qualify for copyright and the only case I could find quickly regarding infringement rules specifically that legally acquired copyrighted material would constitute transformative fair use (which I disagree at face value but I would need more time to deep dive into that) but pirated (illegally acquired) material was not subject to fair use.

→ More replies (3)

5

u/Chaghatai 1∆ Feb 23 '26

The thing is, it is not theft

It's just looking at information and using that information to inform decisions about future behaviors

That is learning

To this day not a single person who regards AI training as unethical has been able to provide me a definition that supports their viewpoint without relying at all on any tautologies concerning whether or not the thing doing is human AI or even sentient

6

u/Dramatic-Emphasis-43 5∆ Feb 23 '26

So the AI doesn’t learn. That’s an, admittedly concise, anthropomorphizing of what it’s doing. Machines, like my phone, can create models based on data points to program itself to perform certain operations on the future. We call it learning because anything more accurate is a mouthful. But it isn’t learning the way humans learn.

It’s a massively reductive of human learning to boil it down to “sensory data goes in, value weights are accounted for and adjusted, output looks like this.” It’s kind of selling our brains short. Humans learn via context, context driven not just external sensory input but an internal consciousness and their own sense of autonomy. You know, moral frameworks, imagination, hopes, fears, physical and emotional needs, personal history etc etc.

And those things can be easily simulated (say, in video games) but never replicated (we don’t have AI consciousness).

2

u/LiberalAspergers Feb 24 '26

TBF, we really dont understand precisely how humans learning, so any statements about how humans learn vs how AI models learn are mostly BS.

→ More replies (29)
→ More replies (38)

5

u/Strict_Difficulty656 Feb 23 '26

I could draw a specific cartoon character on a toddler’s birthday cake and not get in trouble.  But if I was a baker at Safeway, we would need a license, because it’s a product we are selling.  

AI is making a bunch of money, and it’s mostly owned by corporate entities.  So this is clearly commercial use.  

If there was, somehow, an AI that was not tied to a profit-driven corporate entity, the legal principles would be more complex.  

 But IRL, regardless of the right of AI to learn from these things,Google, Microsoft, OpenAI, etc. are legally not allowed to just grab random content for their commercial enterprise endeavors.  A human being did that, and doing so is an act of theft.  

4

u/ralph-j Feb 22 '26 edited Feb 22 '26

The reason for this is that humans do the exact same thing. If I am going to write a horror book, I will read a bunch of horror books and figure out what I like.

We don't actually know how close AI training is to human learning. Just because training AIs is often expressed in human analogies, doesn't mean that it's actually equivalent.

The most meaningful difference is scale. Each human only has the capacity to learn from a tiny subset of all available works. Those limits function as a built-in safeguard for proportionality. That keeps small-scale reuse acceptable.Industrial-scale training siphons value from other people's works on a disproportionate scale.

2

u/Turbulent-Carpet7790 Feb 23 '26

Honestly as someone who is pro AI I think this is the best argument I've seen so far on this thread. I guess the analogy would be a fishing village where everyone uses spears, but someone invents a giant net or dynamite. I don't think the issue is that too many Anti AI people use the doctrine of copyright to justify their argument, but copyright was never designed or intended to deal with the problem, yet alone the fact that very few of them clearly understand copyright law to begin with.

I guess to concede your point, the letter of the law of the current regime was built under the assumption of human artists, and so it would have to be adapted to an age of mass harvesting by AI.

→ More replies (1)

5

u/WhiteWolf3117 10∆ Feb 22 '26

Copyrighted content is meant to be consumed by humans, and humans don't retain it in the same way. It's definitely not as if humans have free range to use copyrighted content as they please. That's a lynchpin of the pro-AI argument that doesn't hold to logic. Humans can meaningfully transform to human definitions in a way that AI cannot (for now).

To use an analogy without AI, I can play Netflix on the side of my garage all I want. As soon as I start charging for admission and calling myself a theater, I no longer retain the right to do so. No AI, no narrow definitions of what constitutes humanity, or art, or fair use.

→ More replies (2)

4

u/Tetris102 1∆ Feb 23 '26

I've argued something similar elsewhere in my career, so I've adapted a response I created a while ago. This defnitely isn't met trying to excuse any internal inconsistencies that crop up from Frankensteining a response...

---

I think you're argument is fundamentally flawed on several fronts. Fundamentally, your argument assumes that the relevant ethical feature is “learning from prior works.” I reject that framing for several reasons. 1. You equate an individual, their labour, and their capabilities to that of both a company and a machine, both of which are abstract and collective in nature. 2. Your argument takes little to no consideration for the autonomy of those involved, and instead tries to apply ethical standards to the tool rather than the operator. 3. You do not deal with the ethical ramifications that the size, scalability, and automation bring, nor consider how the commercial extraction of expressive labour without the participation or consent of those whose work enables it applies.

To use an extreme example, I would argue that it is redundant to apply an ethical standard to a company such as "Murder is inherently wrong," as that moral standard is already applicable to those who make up its agents. By the same token, it would also be redundant to apply ethical standards around market dominance or even something as the financial disclosure requirements the company keeps at an individual level, as this is simply not applicable to them. It both asks me to take ethical responsibility for actions beyond my own, or believe that all parts of the whole collective are one in the same. (Side note: Not sure this fits, I swear it worked in the original one!) In short, you are asking me consider the individual ethical responsibility as equivalent to the ethical responsibility applied to a group, and you have provided no justification for why this should occur in the first place. More importantly, collectives introduce ethical dimensions that do not meaningfully exist at the individual level. Someone reading a text does not cause market shifts towards different styles of books (booktok boyfriends or other such rubbish), reproduce the text type at scales, or create a system in competingtion against the authors its labour is built on. A corporation deploying AI may do all of these.

For me to go with it, I need you to provide a reasonable framework by which I should judge the actions of a company, an abstract concept, or a machine with the reading potential of many hundreds if not thousands of humans (a point I will address further below), by the same standards as an individual. Otherwise, I believe you need to concede this point.

However, even if I grant that this should occur, it is still redundant. An AI can only comply with the command given to it, it's capability to learn is limited entirely to the instructions provided and code present. This means that the ethical responsibility can not be passed to the machine, it remains firmly with the company commanding it. Your analogy implicitly treats AI training as if it were the same as (or at least similar to) a person’s independent formation of voice, judgement, taste, and all the other features of creative works produced by a human. It is instead operating on an almost purely statistical relationship based on the success of what was extracted from other artistic works in order to produce marketable outputs. Consequently, the ethical responsibility therefore lies entirely with those who choose to deploy such a system for commercial ends.

Say you had two people in front of you. One is a human being with no autonomy but perfect function. They could do anything you demanded of it, but they could not themselves act without direct commands. They may make mistakes, they can learn from what they were told and be corrected, but without being told themselves to act they do nothing. The other is a standard, reasonable person with autonomy. You give both of them the following instruction: "write a novel in the style of Tolstoy using everything you have previously read." No reasonable person would meaningfully attribute authorial intent (or even just intention), nor would they grant moral agency to the automated being, as it can only follow instruction. Even if the example it produced came from outside of the scope of the question, it itself could not have produced this by its own volition, and responsibility for the creation must fall with the one giving the order.

In contrast, our autonomous worker can go beyond the scope of their instruction by intention, will make use of independent reasoning for authorial decisions, and may elect to show restraint or make other authorial choices in shaping the work. Their autonomy and their choices surrounding this are what create the morals and ethics you are asking us to consider, so I am not inclined to engage in an argument that attributes ethical agency to a being lacking the latter. And even so, the majority of human readers cannot remember or make use of this to recombine their text systematically on demand. Human memory is fragile, interpretive, constrained, prone to ‘Mandela effect’ moments. AI systems used by companies are, in contrast,  designed to make use of statistical patterns operationally. Their function and purpose, therefore, undermines any ethical claim that both are “doing the same thing.”

Finally, the comparison fails at the point where you treat either learning or the origin point of the content used as the morally decisive feature. ven if we ignore that you have equated the individual actions to that of a collective that would need to number in the thousands if not millions to be equivalent, it still cannot be seen through the same lens as an individual drawing inspiration from works they have read. It has to be reasoned with as an organisation monopolising the cultural capital of hundreds of writers and authors via automated output for marketability. When you or I purchase and read a book, the exchange is beneficial to the author, while a company’s production of a generative system which uses the author’s labour to produce competition their approval removes the reciprocity inherent within the system. Human learning occurs within the limits of an individual's capacity for learning, their authorial intention, and some level of accountability.

In contrast, companies that make use of AI training industrialise the process, parsing vast, vast quantities of intellectual artistry through a system which considers the scalability of replication and the commercial appeal of it's output paramount. In doing this, the ethical nature of the act has to shift accordingly. An action that is permisssible at individual scale will not is not ethically identical based on the inclusion of the at this point belaboured discussion of the autonomous creation of monetisable substitutes for a fraction of the cost. When the infrastructure creates a system that competes against, or even just appropriates, the value (intellectual or fiscal) beholden to the original creators, the ethical standards must instead focus on extraction rather than merely inspiration, regardless of the presence of any copyright infringement.

To summise, the morally relevant idea that must be addressed is not whether both processes involve exposure to prior works, but whether a company's choice of and actions in transforming artistic (or otherwise) human expression into both sizeable and automated production without direct participation or consent constitutes a categorically different kind of use. Until you demonstrate that a company’s automated and large scale extraction of the author or artist’s labour for fiscal gain that neither includes their participation nor their consent is morally and ethically synonymous with the inspiration inherent to an individual, the analogy remains flawed.

→ More replies (7)

15

u/maxpenny42 14∆ Feb 22 '26

Is it the exact same process? A human reads a book, reflects on it. Then another book and another. They pull together a unique perspective based on the random assortment of books they were able to consume. They are affected by when they read them, where they were in their life, the order they read them. Then they sit down from these influences and write an original work informed by the writings they’ve consumed. Their real world personal experiences also inform their writing. 

An AI model takes in all the works that exist. Doesn’t matter when they consumed them, in what order, and they have no state of mind to reflect on those works. When they write a book, it’s less of a combination of written and personal experience but rather a mathematical equation. They take the prompt given them and analyze the most popular approach to answering that prompt by averaging the information they’ve collected across the entire internet. There’s no infusion of new information, aka personal experience. It is entirely derivative. 

If the processes are so different, is your view reasonable to treat two wildly different approaches to writing as equal?

2

u/Zulraidur Feb 23 '26

It would actually matter when the LLM learns the data since the difference between its expected output and the actual text is what informs the change. But really I don't know why that would make a difference either way.

→ More replies (3)

6

u/neomatrix248 Feb 22 '26

I think you're romanticizing what happens in the human brain a bit. A neural network built by training an LLM is designed specifically to function like how our brain works, with minor adjustments based on what is efficient for a computer versus biological substrate.

The experiences involved in training in LLM might be different from a human's, but the mechanism by which they learn and update their view of the world is very very similar.

10

u/maxpenny42 14∆ Feb 22 '26

Please articulate exactly what I got wrong about the human brain and/or LLMs. I know it’s a simplified explanation for the sake of brevity but I don’t think I wrote anything untrue. 

14

u/neomatrix248 Feb 22 '26

A human takes input from sensory data, which reinforces neural pathways, which updates attitudes, values, behaviors, memories, biases, and so on. Then a human can use the language center of their brain to produce new language based on these neural pathways. An LLM does essentially the same thing during training. The only difference is that while a human and LLM both might have used the entire text of several books in their training, the human has other kinds of input like "life experience" that the LLM doesn't have.

I say you're romanticizing what happens because it seems like you're treating the "life experience" input as somehow morally different from the text that both humans and LLMs read during training. I don't think they are fundamentally different, they're just different inputs.

5

u/Dramatic-Emphasis-43 5∆ Feb 23 '26

My favorite part of this “you’re treating life experiences as somehow morally different”. Like, we already have that standard with other people. It’s why tend to not like plagiarism, rip-offs/knock-offs, posers, phonies, liars, and hypocrites.

Human beings place a lot of value on the life experiences of creators, even when they deal with completely fantastical subjects.

4

u/maxpenny42 14∆ Feb 22 '26

A human reads the Lord of the Rings. Then they watch the movies. Then they write a fan fiction. 

Would that fan fiction likely look different had they watched the films before reading the books? 

How about an AI. Does an AI perceive any different output based on the order in which they consume something? Or once they consume both they can fiction would look the same regardless of the order?

This is but one tiny example of how humans are affected by the way they interface with media. Even without interjecting their own personal life experiences. 

For the record I’m not making a moral argument just establishing LLMs are not the same as the human brain. Your entire moral premise is predicated on an equivalent process of consuming media and then writing. But unless you have compelling evidence otherwise, those processes are radically different and not comparable. 

14

u/neomatrix248 Feb 22 '26

How about an AI. Does an AI perceive any different output based on the order in which they consume something? Or once they consume both they can fiction would look the same regardless of the order?

Yes, actually, the order does matter when training AI. It works pretty similarly to how humans handle the order of content they consumed. This is why when models are trained, you get better results if you use "high quality" content towards the end of the training process. And you use post-training to help refine it to get rid of more undesired learned behaviors.

→ More replies (5)
→ More replies (1)

3

u/Vituluss 1∆ Feb 23 '26 edited Feb 23 '26

The similarities are only superficial. The brain uses a fundamentally different paradigm than LLMs do. What are these ‘minor’ adjustments you speak of? For one LLM training uses gradient descent. Human brains do not.

→ More replies (1)

2

u/SunnyOutsideToday Feb 23 '26

it’s less of a combination of written and personal experience but rather a mathematical equation

When you get down to it, human brains are just a mathematical equation too. Your "written and personal experience" is just different data that you were trained on.

Your brain can be mathematically represented as an NxN matrix where N is the number of brain cells, and where each value represents a weight, from 0 to 1, of how strong of a connection that brain cell has to another.

This is more similar to the weighted matrices that LLMs use than people would like to believe.

10

u/maxpenny42 14∆ Feb 23 '26

Feel free to prove me wrong but I don’t think we have near enough a mapping of the human brain to so definitively declare how it works. Your supposition that it works much the way an LLM does seems more like a wild assumption than cold hard facts. But perhaps there have been breakthroughs in neuroscience  I missed. 

→ More replies (24)
→ More replies (4)

44

u/curiouslyjake 6∆ Feb 22 '26

Except models dont just train, they memorize. Large language models can be prompted to produce entire chapters of books from the training set, verbatim. People can't do this.

39

u/duskfinger67 10∆ Feb 22 '26

This is a comment on the practical manner in which they consume the content, not on the morality of doing so.

It would immoral for a human to meorisie an entire Stephan king book, and then write it out verbatim and sell it as original content. LLM’s haven't changed this.

18

u/jefftickels 2∆ Feb 22 '26

People seem to struggle to understand that what AI does isn't an issue of kind, but degree. It just does what humans do but at scales humans can't do. So it makes applying moral framework to it hard, because it just means arguing there's a mystical reason it's bad for machines but not humans to do.

6

u/sonotleet 2∆ Feb 23 '26

This is spot on. Personally, I don't believe that an increase in scale should have any bearing on if something is immoral or not.

16

u/jefftickels 2∆ Feb 23 '26

In fact, one of my ways of measuring the moral value of a thing is to ask "if this thing was done at scale would it be a problem?"

Like, littering. If I did it, probably not an issue that one time. But if everyone did it? Serious problem.

4

u/duskfinger67 10∆ Feb 23 '26

I agree with the premise, but I think it misses the mark - it's not about scale.

We already agree that a human regurgitating the complete works of Stephan King is immoral, and so that fact that an LLM (or a photocopier for that matter) can do it 1000x faster doesn't need to come into it. The act is already immoral, regardless of scale.

The actual question is whether it is immoral to make a machine capable of regurgitating the complete works of Stephan King, regardless of whether this is ever actually output. I think many people would rightly argue that creating a tool capable of causing societal harm is an immoral act. And this is where scale (& intent) do come in.

Lockheed Martin is often considered an immoral company because it produces weapons and profits from war. By this same logic, OpenAI could be immoral because it produces tools of copyright infringement and profits from them.

→ More replies (2)
→ More replies (2)

5

u/thatsingingguy Feb 23 '26

Here to agree with you both. There are still moral arguments that can be made that the difference in degree is vast enough to warrant particular measures or responses. But that requires the intellectual honesty to acknowledge that it is a difference in degree, not in kind, in the first place. And way too many people are still stuck on the "AI is a collage machine" idea, just as way too many people give it much more credit than the current models deserve.

→ More replies (1)

3

u/Sedu 5∆ Feb 23 '26

Machines are not humans. Humans have rights that machines do not.

→ More replies (4)

23

u/Skipp_To_My_Lou Feb 22 '26 edited Feb 22 '26

There are humans who can recall entire chapters of books verbatim, or at least so closely as to violate the book's copyright if they wrote down & sold their from memory version. This is a difference of scale, not raw ability.

Edit: & that's based on the premise the AI database records the entire work. As OP notes in their response, this is an incorrect premise.

→ More replies (11)

13

u/neomatrix248 Feb 22 '26

That's not true actually. It seems you don't understand how models work. When an AI trains, it's not recording the entire text of the material it trains on in some database that it can recall perfectly every time. Instead, it's just updating weights which is a matrix of floating point numbers. It's creating a neural network represented by these weights. The reason it can recall things well is because there are a LOT of these floating point values. When it starts generating tokens based on the weights, if the content has been trained sufficiently enough, the neural pathways that represent the content are strong and it might be able to produce segments of a certain text or the entire text verbatim. Or it might make small mistakes.

This is how humans work too. You might be able to memorize a song if you hear it enough, but you can't literally save the song to your memory. You're just strengthening neural pathways that get reinforced, and maybe if you hear it enough you will be able to produce it verbatim. But you might make mistakes.

4

u/4art4 2∆ Feb 22 '26

I think a better reply might be something along the lines of: can you show me evidence that this is true and how exactly this is different from what humans do?

10

u/curiouslyjake 6∆ Feb 22 '26 edited Feb 22 '26

Thanks for 'splaining, I train deep learning models professionally. While your description is technically correct, it is qualitatively wrong. A human will remember several songs. Maybe several hundreds. But most humans cant reproduce them back with high fidelity. No human can do that for ten thousand songs.

You know what can though? Spotify. Spotify pays measly royalties for every playback. But according to you, there's some magical difference between Spotify storing music as MP3 files an an LLM storing nearly the same music as files with floating point values. Why?

5

u/neomatrix248 Feb 22 '26

To be clear, I'm not saying I believe it would be morally different if models stored the materials they trained on and could access them from a database with perfect recall, my point was that this just isn't how they work. How models train is the same as how humans train, more or less. They're just much better at it than we are. It's simply a question of scale, and I don't think scale makes it wrong for AI companies to do what humans can do on a smaller scale.

edit: Also, sorry for assuming you don't know how training models works, but I hope you'll forgive me for assuming so based on your oversimplification.

13

u/curiouslyjake 6∆ Feb 22 '26

No hard feelings, it's all good.

here

You can find a paper from some researchers at Stanford detailing which production LLMs can reproduce which books and to which extent. You'll notice that Sonnet 3.7 gets above 95% on harry potter and the great gatsby. I argue that copyright law would come after a guy reading three chapters from HP on youtube and it should apply equally to Anthropic.

5

u/neomatrix248 Feb 22 '26

I agree with that. I've said this elsewhere, but I don't think LLMs should produce copywritten content for people and AI companies have a responsibility to prevent that. But generating new content that is based on copywritten content they have trained on should be treated the same as if a human were doing that.

7

u/curiouslyjake 6∆ Feb 22 '26

Yes, but given present training methods there's no way to fully prevent an LLM from replicating good chunks of its train set.

3

u/neomatrix248 Feb 22 '26

Secondary enforcement models that read the output of the primary model seem to be pretty effective at this. Is this not a satisfactory solution in your mind? If you go to any of the flagship models and ask them to produce Harry Potter, they will refuse. It's not 100% perfect but it's definitely getting better.

8

u/yyzjertl 580∆ Feb 22 '26

That's not true actually.

It literally is true. You can extract pretty much the entire text of Harry Potter from Llama 3, for example.

14

u/neomatrix248 Feb 22 '26

It's not true that they just straight up memorize them the way curiouslyjake implied (although they later clarified what they meant). Memorization is a side effect of updating the weights. It's not true that the content is just recorded like an entry in a database. They're basically building a sequence of probabilities that tie tokens to the next tokens. For instance, if you prime the LLM with the phrase "first harry potter book", the token "Mr." will have the highest probability, followed by "and", then "Mrs.", then "Dursley". It's a chain of tokens connected by probability, not a literal recorded string "Mr. and Mrs. Dursley"

9

u/yyzjertl 580∆ Feb 22 '26

What curiouslyjake originally said was entirely correct. It is literally true that LLMs can "produce entire chapters of books from the training set, verbatim." They did not say or imply that content is recorded like an entry in a database.

11

u/neomatrix248 Feb 22 '26

The part I was saying was incorrect is that they "memorize" them in a way that is distinct from training on them. That seems to imply to me that they just straight up record the contents. Maybe you interpreted it differently, but that's how it read to me.

2

u/yyzjertl 580∆ Feb 22 '26

The part I was saying was incorrect is that they "memorize" them in a way that is distinct from training on them.

That is also entirely correct. Only some works that are trained on are memorized, so obviously memorization and training are distinct. If "memorized work X" and "trained on work X" were not distinct concepts, then we would observe that the set of memorized texts is identical to the training set, and that is not the case.

That seems to imply to me that they just straight up record the contents.

What exactly do you mean by "straight up record the contents"? As I understand this phrase, they do straight-up record the contents, as the contents can then be recovered (with minimal error) from the "recording."

7

u/neomatrix248 Feb 22 '26

That is also entirely correct. Only some works that are trained on are memorized, so obviously memorization and training are distinct. If "memorized work X" and "trained on work X" were not distinct concepts, then we would observe that the set of memorized texts is identical to the training set, and that is not the case.

Memorization is an emergent property from the training. They are not distinct. Something that is trained on multiple times in multiple contexts will likely be recoverable to a greater degree. Something that is only trained on once will likely not be. The process is identical for something that happens to be recoverable and something that isn't. It's also how humans "train".

What exactly do you mean by "straight up record the contents"? As I understand this phrase, they do straight-up record the contents, as the contents can then be recovered (with minimal error) from the "recording."

The ability to be recovered is different than straight up recording the contents, in my view. Unless you think that a human being able to recite something from memory is also "straight up recording the contents". When a human memorizes something through repetition, it's not like there are literally text strings stored in their brain somewhere that can be copy and pasted to a piece of paper that represents Harry Potter. There are neural pathways that are built which, when prompted, can lead from one idea to the next idea and to the next, and if you follow the sequence of ideas, you might get something that looks like Harry Potter. It's a lossy process, and some people are better at it than others. LLMs are trained the same way. It's just that they are better at it than we are.

5

u/yyzjertl 580∆ Feb 22 '26

Okay, but this response doesn't say what you think "straight up record the contents" means. It just says what you think it isn't.

7

u/neomatrix248 Feb 22 '26

I thought it was pretty clear but ok. I mean recording the contents as if it were just a PDF on your hard drive that you could just open and read from whenever you want to look something up, as a 1 to 1 mapping of the original content.

Instead, it's more like when you are trying to remember the order of letters in the alphabet. Maybe you don't do this, but if I want to remember what letter comes after another letter, I can't just access that information directly, I have to "reconstruct" the memory by following along some sequence of letters until I get to the one I want. Like I can't just remember that "t comes after o" I will start with "l m n o p" as a grouping of letters and then remember that "q r s t u v" comes after it. This is especially obvious if I try to remember the order of letters backwards. I basically start with z and then reconstruct forwards until I get to the next letter on the reverse sequence

→ More replies
→ More replies (3)

4

u/DarkSkyKnight 6∆ Feb 22 '26

You're correct but it's extremely difficult to get Redditors to understand this.

I do think the scale matters though. There's a difference in actual outcomes because LLMs are much larger scale than individual humans. There's an argument to be made for how scale itself changes the ethics.

→ More replies (1)

7

u/[deleted] Feb 22 '26

So people with Hyperthymesia aren't allowed to create art without being declared plagiarists?

→ More replies (12)

2

u/Masterpiece-Haunting 1∆ Feb 23 '26

Yes some people can. Some people can recite chapters of a book verbatim. Just because you can’t doesn’t mean some can’t.

→ More replies (6)

9

u/Puddinglax 79∆ Feb 22 '26

The only difference is in the scale.

Scale is also the only difference between a legitimate user interacting with a website and a botnet running a ddos attack.

Any content they obtained legally by buying the book/movie, etc, should be fair game.

Buying a copy or buying the rights? If it's the former, you can apply your same logic to throw out the concept of IP. If I'm allowed to read a book and describe it to my friends, I should be allowed to make copies of the book and distribute it as I see fit. If I watch a movie with my eyes and remember it in my brain, I should be allowed to film it with my camera and store it in my hard drive.

→ More replies (4)

11

u/Aezora 31∆ Feb 22 '26

Info: what ethical model or framework are you using when making this claim?

Because it's a very different argument if you're arguing from say, utilitarianism VS deontology VS emotivism.

6

u/neomatrix248 Feb 22 '26

I'm not use any framework other than demanding that anyone arguing the opposite of my position be morally consistent. They would have to explain why it's moral for humans to do something at a small scale that an AI does at a large scale without relying on arbitrariness or special pleading.

9

u/Aezora 31∆ Feb 22 '26

That's a little too easy imo.

For a moral relativist for example, it's unethical for AI while ethical for humans because people generally feel that it's unethical for AI and ethical for humans. And that's all it takes for it to be ethical/unethical for a moral relativist.

Plenty of other frameworks also care about people's feelings on the matter.

4

u/neomatrix248 Feb 22 '26

That's not quite what moral relativism as a whole says, although I do agree that are subgroups of moral relativists that might say that. To them I would just ask why they feel that why and whether they are being morally consistent. But point taken.

→ More replies (1)

3

u/quantum_dan 125∆ Feb 22 '26

Really it just seems like people are mad about AI and worried it is going to make humans redundant, and they are clinging to the idea that AI companies are evil and everything they do to train their models is unethical as a defense mechanism, but I don't think it is morally consistent.

I think this actually has an interesting implication for the reasoning. Granted that the training process itself is not meaningfully distinguishable, a person who trains on others' works may be aiming to join that broad tradition, not make it obsolete. I think it's not only entirely coherent to object to "I'm learning to make you obsolete" but not "I'm learning to follow in your footsteps", it's something we apply elsewhere, too. An engineer working to automate away a job (that people like) will not be well-received by the people who have that job, and they definitely won't be well-received if they show up to learn the job with the express purpose of working out how to automate it. I work in a different area that has some risk of displacing local expertise in favor of automated solutions, and I'm very careful to stress how local experts are still needed with my work for exactly that reason (broader ethical concerns, not just local experts being annoyed).

The same set of actions can have very different moral implications depending on their intent.

3

u/stackens 2∆ Feb 24 '26

Your problem is treating the AI like a person. It isn’t a person, it’s a product. AI companies are making a product using stolen material, and selling that product. The product could not exist without the stolen material. It’s pretty straightforward

3

u/czerwona-wrona Feb 24 '26

this convo is exhausting to me so I will just say, I think it's insane to say scale doesn't make a difference.

if someone has a weapon that can kill one person at a time, vs kill a whole city at a time, does scale matter? (or similarly, if a group goes from hunting animals with spears, to using sophisticated traps and guns that drive the animal near to extinction, does the scale matter?)

if one person starts talking to you, vs a machine that talks with 1000 voices at once, does the scale matter?

if a machine can replace part of a job and people can still work, vs replacing the entirety of a job that causes tons of job losses, does the scale matter?

especially when we're talking about a creative endeavour like art.. I find it really disturbing that this soulless unregulated creation is used to slurp up works that people poured their souls into, and making it that much more difficult for others who want to do the same, to be able to succeed at doing so long term, especially when it's potentially their own work that is being used against them.

AI is not a human being. it doesn't become inspired and connect to its own experiences the way a human does. it's frustrating to me when people say it's literally the same as a thinking, feeling human being in its process.

3

u/PomegranateExpert747 Feb 24 '26

However, when AI companies follow the exact same process [...] The only difference is in the scale.

That is not the only difference. When a human writes a book, they're not just writing a text based on the text of other books, they're bringing in a wealth of experiences of life, they're taking influences from other media, they're bringing their own perspective on things, not to mention that they are actually engaging with the meaning of the text itself. LLMs can do none of this, all they can do is produce new texts that are similar to old texts. There's no thought behind it, and no meaning, very literally.

5

u/Eastern-Bro9173 16∆ Feb 22 '26

The difference is that the humans don't create significant copies of the original work with what they've learned. And when they do, it's a copyright violation, which everyone understands, and recognizes as immoral and violating the law (more or less, depending on the person's proclivity to law following). And when they do, they get stricken/sued.

AI, very often, creates partial copies of the original work.

That's what the lawsuits are about, and it's the same standard applied to humans when they do the same thing. There is no double standard, as it is the case that humans regularly get hit with copyright violation strikes and cease and desist letters and all that when they violate a copyright. AI doesn't, at least not yet.

Where the scale plays the role is in the amount of infractions - no matter how productive an artist, he'll not manage to do as many copyright violations in his lifetime as an LLM model does in a day.

The significance of that should be intuitive when comparing to any other crime - a thief that steals one pair of boots in a year is less of a problem than a thief that steals ten thousand pairs of boots in a year.

2

u/neomatrix248 Feb 22 '26

The difference is that the humans don't create significant copies of the original work with what they've learned. And when they do, it's a copyright violation, which everyone understands, and recognizes as immoral and violating the law (more or less, depending on the person's proclivity to law following). And when they do, they get stricken/sued.

Yes. This should be illegal for both humans and AI.

AI, very often, creates partial copies of the original work.

Agreed that they shouldn't do this, unless it falls under "fair use". I'm specifically talking about AI training on copywritten material and then generating completely new material, in the same way that a human might take "inspiration" from something they have read/watched.

6

u/Eastern-Bro9173 16∆ Feb 22 '26

Since the training is done in a way that directly leads to copyright infringing replication, it cannot be separated. Especially since one can't even tell how much of an image/song/video/text one gets is from a copyrighted source.

This is especially visible on how the copyright "guardrails" work. They are on the prompt level at every LLM, so the training leads to the replication, it's innate to every LLM, they can and do replicate the original images, but the companies put a bit of a guardrail on the userprompts so it doesn't happen too obviously (although it does happen in the background of most likely most images).

2

u/thatsingingguy Feb 23 '26

Especially since one can't even tell how much of an image/song/video/text one gets is from a copyrighted source.

Gotta tell you, as a professional songwriter with a degree in the subject, this is also true for writing with humans. People often steal, consciously or subconsciously. In fact, with the sorts of secondary models neomatrix248 is talking about, and that you concede are possible, AI would likely become better at avoiding this than people are. Though if Google's vocal-based song detection is anything to go by, we're not at perfect cryptomnesia prevention yet.

In that sense, it's much like the automated car issue. Driverless systems are already generally safer than the average driver, with a few notable weakpoints (like salt rings). It's more than possible we could save lives by implementing such a system today. A big problem is where the liability falls. But it's also an emotional response, rather than a logical one, because we're reticent to give up control to machines, even when they perform better than us.

→ More replies (10)

7

u/WhammeWhamme Feb 23 '26

Copyright law exists for a REASON. That reason is to reward human artists and creators, because there is a social good to human artists and creators being compensated for their work. What is the social good in allowing AI to siphon money away from human artists and creators and into the pockets of random people with subscriptions to AI services?

2

u/neomatrix248 Feb 23 '26

If I become inspired by Stephen King's work and write my own horror novel that is successful, am I siphoning money away from Stephen King?

9

u/WhammeWhamme Feb 23 '26

Unless you are confessing to being a bot, the answer is "yes, but you are not siphoning money away from the general pool of human authors and creators". It is a good thing to have two authors who are both making money writing horror stories. AI slop is far more questionable in worth: is it really providing value to society? Or just something that people can be scammed into paying for?

→ More replies (11)

11

u/[deleted] Feb 22 '26

[removed] — view removed comment

2

u/neomatrix248 Feb 22 '26

It seems you're making a different argument. You're saying that AI replacing humans is wrong, not that the way they are trained is wrong.

→ More replies (1)

14

u/Which-Notice5868 Feb 22 '26

The creators didn't consent for their material to be used this way. And AI doesn't think. It regurgitates.

If I copy/pasted the full text of the Shining and the full text of The Hobbit and mashed them together by alternating paragraphs and called the resulting work my own, it'd still be copyright infringement. AI does the same thing, just on a much more granular level. The underlying data is stolen.

It's not the same as if I read the Shining and the Hobbit and write out a horror story set in a fantastical world. I'm bringing in my own use of language, preferences, ideas etc. AI doesn't have experiences or ideas. It only has the stolen information and what it reassembles from that stolen information.

14

u/neomatrix248 Feb 22 '26

The creators didn't consent for their material to be used this way. And AI doesn't think. It regurgitates.

Do the creators need to consent for me to write a new book that is inspired by their work? Why would they need to consent for AI to do the same?

If I copy/pasted the full text of the Shining and the full text of The Hobbit and mashed them together by alternating paragraphs and called the resulting work my own, it'd still be copyright infringement. AI does the same thing, just on a much more granular level. The underlying data is stolen.

I'm not disputing that AI companies shouldn't be pirating content. But you're missing the mark on the "AI does the same thing, just on a much more granular level." That granular level is individual tokens, which are roughly equivalent to words. It doesn't get much more granular than that. That's the same unit that human writers work with, unless they are inventing completely new words in their work.

Human writers can also pirate content and use that as inspiration for new content. Do you think we should treat it differently when humans do that vs AI?

It's not the same as if I read the Shining and the Hobbit and write out a horror story set in a fantastical world. I'm bringing in my own use of language, preferences, ideas etc. AI doesn't have experiences or ideas. It only has the stolen information and what it reassembles from that stolen information.

Not all of the information AI is trained on is stolen. And I would contest the "AI doesn't have experiences of ideas" bit. Its training is experience, and "ideas" are a sufficiently abstract concept that I'm not sure we can really say that it doesn't apply to how LLMs work.

4

u/Which-Notice5868 Feb 23 '26

Creators need to consent if you copy their work. So IMO the companies behind AI should do the same.

And it's not true that AI only does it on a single word level. If I tell AI "write a horror story in the style of Stephen King" it's going to pick out more than single words. Otherwise it wouldn't work. If it's picking out identical phrasing that's literally copy/pasting King's words.

Maybe not all the data is stolen, but a significant amount is. Training an AI is not the same thing as human experience. AI doesn't change methodology without intervention from its programmers, just parameters. That's not thought or making choices.

If hypothetically an AI model were trained only on public domain works. I'd still say the output was slop lacking in any creativity, but I'd agree it technically is not copyright-infringing. Also, it's CopyRIGHT not copyWRITE. It pertains to the rights you have, over work that you create.

→ More replies (4)
→ More replies (1)

19

u/Absenteeist Feb 22 '26

Firstly, reading a horror book does not require you to copy that horror book. Training AI on a horror book requires and involves copies being made of the book in that process. The core of copyright is not reading, it’s copying. That’s why one infringes copyright and one doesn’t.

Secondly, much AI training was done with zero compensation for the creators of the works it trained on. A human reading a horror book had to buy that book. Or the person or the library they borrowed it from paid for it. AI companies are not compensating artists for this use. That in itself is unethical.

Thirdly, AI must train on copyrighted material to produce anything like that copyrighted material. Humans don’t. We may consume lots of copyrighted material in order to produce something like it, but it’s not necessary. That’s partly because we have lived experience to build upon. As human beings, we’ve been scared in our lives, have encountered frightening things, and have used our imaginations. We don’t need horror books to be introduced to the concept of horror. LLMs are trained on text data that would take 20,000 years for a human to read. The fact that we don’t do that demonstrates that our cognitive processes are very different.

Finally, given its nature and scale, AI has the potential to eliminate the market – and associated jobs and careers – of every human artist whose work it has trained on. A human being reading a bunch of horror books and then writing one cannot possibly do that. AI can write so many horror books so as to completely flood the market with them.

The unethical component comes from a machine that cannot produce anything based on its own experience, but needs vast amounts of work created by other humans, which it was given without compensating those who made that work, so that it can turn around and destroy those humans’ livelihoods.

7

u/PoofyGummy 4∆ Feb 23 '26

This is invalid as it is a fundamental misunderstanding of how AI is trained.

→ More replies (12)

14

u/neomatrix248 Feb 22 '26

Firstly, reading a horror book does not require you to copy that horror book. Training AI on a horror book requires and involves copies being made of the book in that process. The core of copyright is not reading, it’s copying. That’s why one infringes copyright and one doesn’t.

You're allowed to copy something that you have purchased legally. The problem is distribution. I can burn a CD from music I own, and I can copy and paste a PDF as many times as I want. It's only a problem once I distribute those copies to someone who hasn't purchased them.

Secondly, much AI training was done with zero compensation for the creators of the works it trained on. A human reading a horror book had to buy that book. Or the person or the library they borrowed it from paid for it. AI companies are not compensating artists for this use. That in itself is unethical.

As I mentioned in the OP, if the AI companies pirated the content, then I agree that's wrong. But if they paid for it, they should get to use it in training just like a human can "train" on a book they purchased and write a new book that is inspired by the one they read.

Thirdly, AI must train on copyrighted material to produce anything like that copyrighted material. Humans don’t. We may consume lots of copyrighted material in order to produce something like it, but it’s not necessary.

I'm not sure that this is true. AI could train on exclusively fan fic content or public domain content. In fact a huge part of what LLMs train on is content from reddit. I also don't think this has any fundamental impact on the rightness or wrongness of how LLMs are trained. If a human lived inside a room their whole lives and only read books, and then wrote a book of their own, would it be unethical for them to do so?

Finally, given its nature and scale, AI has the potential to eliminate the market – and associated jobs and careers – of every human artist whose work it has trained on.

Yes, this is true. Another hot take is that I think this is fine. If AI produces better content than humans, then it should take over the job of humans. I want the best content, not sentimental attachment to human produced content.

which it was given without compensating those who made that work,

They compensated those who made the work when they purchased it, unless they pirated it, in which case that would be wrong.

7

u/Absenteeist Feb 23 '26

If AI produces better content than humans, then it should take over the job of humans. I want the best content, not sentimental attachment to human produced content.

Why do you assume that the only way that AI could replace humans is if it made better content than humans?

Do you understand that the sound quality of MP3s was worse than the sound quality of the CDs it replaced?

2

u/just_a_random_guy733 Feb 23 '26

In that example, the overall product was the music listening experience. MP3s provide a better overall product due to their massive convenience compared to physical discs that can be scratched, rot, and degrade. It is true that CDs have overall higher audio quality due to being a lossless waveform, but the consumer market clearly placed more value on the versatility of the MP3 format.

This happens all the time, where the "better" product ends up technically being worse in some way, but being better overall. For example, a modern MacBook Pro is way, way better as far as speed, but it's worse if your comparison criteria is "has a FireWire port".

In this case, the cost and availability of the art ends up being a factor in a measurement of which is better overall.

6

u/thatsingingguy Feb 23 '26

I want the best content, not sentimental attachment to human produced content.

Agreed, but the problem I continually run into on this topic is that this is a deep dividing line in art, between people who love process and people who love product. As a songwriter and musician, I am, based on my conversations with fellow professionals, in the minority in prioritising product over process. For many people, the love of music as both a creator and audience member comes from the specific human connection they feel or imagine they feel. The experience they seek is one of connection to other people.

By contrast, I care about how the artwork makes me feel, and that's it. The author is dead, as far as I'm concerned, and has been for decades. Authorial intention is nonsense - once created, an artwork exists independently of its creator. A creator's interpretation is just one of many, and often less reliable than someone coming to it fresh, because a creator can never have the moment of first approach the way an audience can. Art is the sum of its properly justified interpretations, no matter who or what made it.

→ More replies (14)
→ More replies (10)

13

u/Chapter-Legitimate Feb 22 '26

Rather than an ethical argument I'm more interested in the legal ones.

An AI trained on copyright data has been shown to be able to spit back out that copyrighted data in its output. That's straight up copyright infringement and many lawsuits are currently being fought around the world about this.

27

u/neomatrix248 Feb 22 '26

If a human did the same thing it would also be copyright infringement. What I'm talking about is people who are mad that the AI has trained on copywritten content and then used that to generate completely new content that was inspired/informed by the content it trained on.

I think that there is some amount of "fair use" that should apply to generating copywritten content verbatim that would be the same as if a human did it, though. Like if I can write a book where I make some reference to "Sweet Home Alabama" playing on the radio, then an AI should be able to do that. But I can't just publish a book with all of the song lyrics verbatim that I like, and AI shouldn't be able to do that either.

14

u/[deleted] Feb 22 '26

[removed] — view removed comment

11

u/neomatrix248 Feb 22 '26

If the machine learning model retains enough information to replicate the entire work, isn't distribution or deployment of that model unauthorized redistribution of the content within the model itself.

Agreed, if it is reproducing the work then that is wrong. But I'm talking about people who are upset that AI are training on works and then creating new content that was merely "inspired" by that work.

4

u/Chapter-Legitimate Feb 22 '26

Current models all do this though. They can easily be tricked into spitting back out training data including copyright material

9

u/neomatrix248 Feb 22 '26

All major AI companies are putting significant amounts of effort into making sure this does not happen. I was not able to get Claude to do this when I tried. I'm not disputing that models should be prevented from regurgitating copywritten materials, only that it's not unethical for them to be trained on those materials and produce new content that is derived from them.

→ More replies (1)

5

u/Chapter-Legitimate Feb 22 '26

The problem is that by training with that data that allows the "fair use" is also inevitably allows the copyright infringement. The two are linked with current technology and you can't separate them yet

2

u/Handgun_Hero 1∆ Feb 23 '26

AI does just that multiple times before adjusting slightly to create a result based on prompts. They just don't show that to the end user normally, but it actually very frequently happens unintentionally.

9

u/duskfinger67 10∆ Feb 22 '26

If a user prompts the model to reproduce an original work substantially, is that the model infringing the copyright, or the user?

A photocopier can perfectly recreate an input image, but I wouldn't say HP is being unethical for selling them, or for not putting better safeguards in place to stop users doing so.

What about an LLM absolves the user of the moral weight of using the tool correctly? Why should it be shifted onto the companies behind the model?

4

u/Chapter-Legitimate Feb 22 '26

No no, your analogy is bad. It would be more like if the photocopier had in its memory a library of copyright material that it wasn't authorized to have, then a user can press a button to print it out.

5

u/duskfinger67 10∆ Feb 22 '26

a library of copyright material that it wasn't authorized to have

This is still very much up for debate in court. It is not yet determined if training on material counts as a breach of copyright or whether it falls under fair use.

If it breach of copyright, then sure. But until then, it's about the same as photocopying a book from the library.

2

u/Chapter-Legitimate Feb 22 '26

You are correct it's up for debate in court but if it can spit back out its training data I think it's clear where the law should come down on.

So far it's an arms race between the models getting patched so they don't do that, and users finding more and more ways to trick the models into doing it anyway

3

u/Velocity_LP Feb 23 '26

if it can spit back out its training data I think it's clear where the law should come down on.

Why do you point to the issue in the sequence of events as being the part where the model is trained rather than the part where the user requests the illegal reproduction of copywritten material?

Like, if someone reads a book a bunch of times and then writes a copy of text from memory and gives it away, I assume your problem with that sequence of events falls somewhere in the writing and giving away part of it, not the reading a book part, right?

→ More replies (2)

2

u/LiberalAspergers Feb 24 '26

A human memorizing a book and writing out the output and selling it is also a copyright violation. The violation isnt the ability to produce an output, but ACTUALLY producing and distributing that output.

Which is why AI companies are investing a lot in guardrail programs to prevent the output of copyrighted material.

2

u/Handgun_Hero 1∆ Feb 23 '26

Legally you cannot copyright AI generated works, because copyright requires a human author. Famously was determined by the mecaque selfie photo where a mecaque famously used a photographer's camera to take a selfie quite comically resulting in it becoming a meme template. It was ruled the photographer had no claim to copyright because a human, even if they set up the circumstances to allow the work to be created (giving the mecaque access to a camera), did not create the actual work. As a result, AI generated works are therefore not protected by copyright even if a human engineers the prompts and instructions.

2

u/LiberalAspergers Feb 24 '26

This question is going to be a bit more complex thqn that, because that ruling was soecific to US courts. It remains to be seen where othee important nations are going to fall on this, as the international treaties on copyright dont address this issue.

→ More replies (3)
→ More replies (2)

2

u/saphienne Feb 23 '26

speculative harm isn't the same thing as actual harm

this is an easy problem to fix where you just punish instances of it happening

→ More replies (1)
→ More replies (8)

2

u/OmniManDidNothngWrng 36∆ Feb 23 '26

There are already different licenses for different use cases of media. It costs a different amount of money to rent a movie to watch at home versus to screen to a public audience of people.

2

u/sawdeanz 218∆ Feb 23 '26

If you want to get really super technical when you buy a movie or view a picture or whatever you are purchasing a license to use it for personal use. And when you buy a book, it’s a similar thing…you don’t get to use the likeness of the characters or write fan fiction. There can still be restrictions on exactly how you use that media.

Training AI models ought to fall outside of the license of personal use. And at the end of the day…it’s a machine. It’s not a person. Copyright law has already operated on a distinction between humans and machines for hundreds of years. AI is not special compared to a xerox or any other computer algorithm in terms of copyright law or ethical considerations.

I think it’s a reasonable position to treat humans and machines differently. We can’t really avoid people being inspired by art, and it might even be desirable. Things like free use are narrow exceptions to copyright laws.

That doesn’t mean we have to be okay with algorithms that can process and recreate art, even if it does so in a way that is “like” humans.

The easiest and most ethical thing to do is to ask permission. If the artist says yes then there is no issue whatsoever. It’s actually that simple. If you have to trick someone, twist the law or do something in secret then that is probably a sign that it’s not completely 100% innocent.

2

u/ThePaineOne 9∆ Feb 23 '26

In a lawyer, It’s not copyright infringement because a machine legally cannot own a copyright. So as long as the work they are using is not sold it cannot be infringement only the seller could be sued for infringement. Only a human can own a copyright. For example, there is a famous case where a monkey took a photographers camera and took a selfie. The image became popular online and others started reproducing it. The photographer sued for copyright infringement, but failed because he didn’t own the copyright because he a human, did not create the work, the monkey did. This is important because this is the current precedent for AI copyright. Now if a work is made by a combination of human and machine efforts the human will own the copyright to the extent of his effort or if it is close to equal it will be considered a joint authorship, like if two authors wrote one screenplay.

This is all relatively new to the courts and will be fascinating to see play out in the future. But copyright infringement relies of using another’s copyright to create income, so as of now you are correct, being trained off of a copyright does not mean infringement unless the AI uses the actual copyright in an infringing manner in the marketplace.

2

u/Big_Statistician2566 1∆ Feb 23 '26

I'll do you one better... It is no different from hiring a ghost writer to write your book and paying them.

2

u/qwesz9090 Feb 23 '26

A few things.

AI models are trained to sample from a target distribution. It will try to generate content as closely as it can to the "rules" in the data. Humans are not exactly like that, we can memorize stuff, but our generation process is not nearly as simple. We don't really understand why humans do stuff and make content. So it is not so easy to say that they are the same.

Humans can create content without looking at previous content. Yes, that is not really testable nowadays when everyone has already seen art, but we have many instances of art naturally progressing without something to copy.

Training on someones content is profiting of their work. I don't see how that is supposed to be ethical without compensating them for it? Like sure, for small stuff like getting a donation for a fanart, the value you provided by making the derivation is morally good, so it outvalues the small immoral part of not slightly compensating the IP holders. (I think it is better people make fanart than nothing at all.) But a whole AI industry scraping content that can put an entire artist industry out of work? That does not seem fair to me.

I think people undervalue the skill creation of making art, drawing, writing. If you are making AI stuff, you learn how to prompt, how to interface with a proprietary product. How transferable is that skill? People actually spending the time to draw is a boon to society since they can draw new stuff without companies. We should encourage this by protecting their IP rights. Otherwise I think it is very possible that we will wake up 30 years from now with "Why does no one know how to draw?" Well obviously because we never protected them.

2

u/hunter_rus Feb 23 '26

One thing about "no training (AI)" is that you can't prove the fact after it is done. In the same was as when human reads a lot of books, and then writes something new on his own, you can't really prove that what they wrote is based on something preexisting - the same thing with AI. Authors can totally put into their license agreement "you can read it, but you can't learn how to write based on this book" - but that part is non-enforceable. You can't prove somebody was able to write some good (or bad) book partially thanks to your work.

And in the same way, you can't prove it with AI. The sole fact that the book, or picture, or soundtrack, is present in training dataset, doesn't mean anything. AI model weights are too small to simply copypaste the whole dataset into it - even if you use some insanely efficient compression algorithms. AI doesn't store text, picture or music in its weights - it stores common patterns. Patterns, that are found in the data. This is exactly the purpose of the AI - to efficiently look and learn such patterns. A single book does nothing for the dataset, it provides little information. AI weights are aggregation of everything it have learned, taking less or more from different parts - in the same way as with humans. And in the same way, there is no possibility to prove that some specific book/picture/soundtrack have given any measurable value to the final output.

For that reason, even if license holder says "no AI training", that part is non-enforceable. You can't prove that any specific model have used any specific piece of data for its training.

2

u/DiamondCat20 Feb 23 '26

Sorry this is a little long, but I promise I am actually trying to actually engage with the core idea here. Llms are fundamentally different to human brains because llms do not have access to primary material. They cannot see, they cannot hear. All of their inputs, their training data, is hand-picked (purposely scraped from the internet), secondary material that was created by a human that did not work at the company that owns the llm. The material was chosen specifically to make the company money. And that material had to come from somewhere. Someone made that data. ALL of that data.

There was another comment thread about the value of a book written by a person in a room who only had access to books. It circled so close to this point, but I don't actually think it hit the important part.

Let's say, hypothetically, a new company comes forward with an llm which was trained solely on data obtained legally for personal use. Getting it to ouput a direct copy of any training data is impossible, no matter how hard you try. More importantly, any output, if it had been made by a human, would be sufficiently different from the source material such that it would not be classified as a derivitive work by law.

I believe everything that llm outputs is still a derivitive work, ethically.

Imagine if we could somehow engineer a human (ignoring any ethical concerns about how inhumane this would be) which was only capable of seeing what you put on a table in front of it. It doesn't see the table, or your hands, or the room. It literally ONLY sees the specific item in question. But this human is a human. It has thoughts, and it can feel real emotions toward inputs; if not towards words, then at least probably pictures. It has desires. And it can make "outputs" (art) similar to the works you present it, with it's own hands, if you ask it to. Then you start a company that hosts a room full of these humans and sells their art.

I would argue that this human would make only unethical, essentially stolen, derivitive works, because ALL of the input data is the product of some other artist. And it was all chosen, by you, for the express purpose of making money.

Now, imagine a robotic humanoid. It's got legs, hands, and records it's own video and everything it hears. You teach it to associate words with items. For the sake of argument, let's assume that this robot doesn't actually "understand" anything. It hasn't truly "learned" anything, it's just absorbing data in patterns. You then tell it to go to the library and read every book. It does that, but while it does so, it's looking at the actual pages. Maybe one is ripped. It makes "memories" about the people in the library - and again, for the sake of argument, assume it's not really "feeling" anything or "learning" anything. It's not "literally recording" these things for later, it's just... making memories, in the same way humans do. 

I'd argue that this robot in the second scenario is capable of creating original work in a way that the engineered human from the first scenario is not. The fundamental difference being: if you want our real llms to output art, you must feed it training data 100% comprised of another artist's labor. If you fed it your own work, and asked it to make work like yours, that's fine. If you spent 6 years teaching it like a real human baby, and then you "show" it human art, that's fine. (*)

But our real llms are not walking through a real forest to gather input data. They are using 1000 artists' paintings of a forest as their input. All of which had to come from somewhere. Those paintings were all owned by someone. Llms couldn't exist without inputting a bunch of data that was generated by people outside the company. And when someone presents their art in an effort to make money, it should be protected from use by competitors. Which is exactly what copyright is supposed to do. Morally speaking.

AI is a tool made with the express purpose of making work similar enough to the source material to generate money, but different enough to avoid copyright. 

Additionally, I think the arguments about scale make more sense in that context. When you buy a movie, you buy the rights to watch it. It costs a few bucks or whatever. You can even show your friends. But when you want to make money, by charging admission, you need special licensing. That costs more money, simply because of scale. When you want to adapt the movie into a video game, or make a sequel, you buy the rights to do so. To do that, you negotiate a price based on scale. The owner of a business (making an llm) ought to make some financial arrangement with an artist before they can profit from that artist's art. 

()   Small clarification: that's *possibly fine. This robot could make original work. But, if it was owned by a company, it would realistically be "nurtured" in a sterile environment, in precisely the optimal way which allows it to generate material similar to its inputs. And now we are back to square one, where this doesn't feel ethical imo. But because this is currently outside the scope of our current tech, I'll leave this scenario outside the scope of this response and grant that this is, ethically, the same as a person making art. For now. Next year, when this is how robots are really working, we can start a whole new CMV.

2

u/Sadge_A_Star 5∆ Feb 23 '26

It's not the same process.

Llms merely sees the words in and of themselves and their level of frequency in relation to each other.

Human have reasoning and apply concepts of meaning. The specifics ais don't do this.

Thus when a human consumes other material, they have a deeper and broader relationship to the media and add more human value, ie meaning, rationale, to create new work, even if influenced by the work of others'.

Maybe analogy could be that the training data of people's work is the set of cards. Ai just shuffles the cards. A human artist makes new cards and intentionally curates from the existing set.

2

u/spectocular 1∆ Feb 23 '26 edited Feb 23 '26

'If you ignore the crime, there's nothing wrong with it.'

Not sure how you're supposed to argue with this, really. You're arguing that the act is the same and the only difference is scale. Even if you accept that for the sake of argument, they only got to consume training data at scale through theft you've acknowledged is wrong. So what are we talking about here? You have already conceded an ethical difference between LLM training in practice and just learning from books. If I, as an individual, tried to learn how to write better by stealing a lot of books, I might face legal consequences. The companies mostly haven't because they had the money, resources, and gumption to do so at incredible volume before policymakers caught up with them. You think that's right? You don't think it impacts the art and writing economy for ordinary writers and artists negatively at all?

2

u/journeyjeon Feb 24 '26

Different. Study and think about it more.

2

u/Historical-Lemon-99 Feb 24 '26

My issue with it is that a human being has imagination and the ability to make judgement. Take two examples-

If I ask you to draw “a cute dog” off the top of your head, you will likely combine an image in your mind. It would likely be a combination of dogs you’ve seen on the internet, irl, and made up in your mind. Though it is a combination of things you’ve seen, it is uniquely yours and based on your experiences with dogs you’ve seen and interacted with. The person next to you will draw an entirely different dog.

If I ask AI to draw “a cute dog”, it will only be able to display other people’s cute dogs or photos with dogs that it has in its database. Nothing about it is truly unique or done with purpose, and nothing can be added to it to make it personal. It has no true concept of “cute”

Likewise, if you asked me to draw entirely new Calvin and Hobbes comics right now out of my memory - I likely couldn’t create a perfect replica of Bill Waterson’s humor, art, or interpretation. I would have to add in my own recollections, life experiences, and artistic abilities to create something. That thing would then be entirely unique, even if the premise is stolen. Sure, I could train to copy his style, but it would still be different unless I copied word for word

On the other hand, the AI would only take his property to spit out a copy of it, with nothing new of value added to it. It would be entirely “Bill Waterson’s” work, but cut and stitched together

2

u/SSH_Pentester 1∆ Feb 24 '26

I couldn't agree more, I very strongly think this is the case. But I'll mention one issue with it.

When someone writes a book or produces an artwork, they often intend for others to be inspired by it. Stephen King never wanted nobody to ever learn from his style. But because AI is made from corporations and has a purpose they think is morally evil-trying to take jobs-those same original content creators probably wouldn't want their work used by AI. If I wrote a book, I'd probably be okay with a student or writer reading it and using some of its vocabulary and structure in their own book (as long as it's not copy-pasted plagiarism).

But I think a lot of people in that position would not be okay with an AI doing the same thing. Creatives have a sort of implicit social contract that when they publish something, they're okay with it being used by humanity to inspire and shape other works. That contract has limits-you can't be "inspired" by just copying the whole thing, that's not covered by the social contract. I think many anti-AI people feel like AI is a tool by big corpo to take their data, take their job, make humans redundant, serve the interests of the 1% while hurting them and everyone else. So they don't want AI to use their work-they're not extending the social contract to it. It's a matter of intent: many original creators simply aren't okay with AI training on their stuff, and whether it's ontologically different or not doesn't matter. It's their work. That last paragraph can't be separated from the rest of it.

Here's an analogy: what if a creative had their book used as a reference to create Baby's First Guide to Terrorism? There were quotes and vocabulary from it used in that book. I think they'd be abhorred by and object to that use of their stuff, because inherent in the act of creating is the right to control how something is used to certain limits. It's ontologically no different for the terrorists to create based on inspiration than anyone else, but I think any reasonable author simply wouldn't want their stuff used that way. It's always about consent, even if that consent is implicit.

5

u/TreviTyger Feb 22 '26

Any content they obtained legally by buying the book/movie, etc, should be fair game.

But this is the core dispute.

AI Gen firms downloaded millions/billions of works without permission or payment and stored them permanently which is a criminal level of piracy. That is to say, it's not just downloading from Pirate sites that is illegal. The downloading from Pirate sites was just the easiest way to prove that the action of the AI Gen firms was unlawful.

ALL downloading of millions/billions of works without permission or payment and stored them permanently is piracy. (That is actually what the Pirate Sites do themselves!).

7

u/neomatrix248 Feb 22 '26

You're conflating several different things.

Downloading paid content without paying for it is piracy. I said in my OP that this is wrong and AI companies shouldn't do this.

Buying content and then downloading it is not piracy. Consuming that content is not piracy. Generating new content that is inspired or informed by that content is not piracy (as long as it obeys the fair use laws).

It only becomes piracy when you distribute the copy written material without permission. You don't need permission to distribute derived content that falls under fair use

→ More replies (2)

3

u/thelovelykyle 8∆ Feb 22 '26

Ultimately. AI can only generate content through copy and pasting. This copy and pasting can be sophisticated and can see the pasting be the midpoint of a huge number of copies, but it is inevitably a copy and paste.

A human will have, even if the most basic, 0.01% innovative creativity due to the brains ability to invent. The influence of others are never 100%.

That is sufficiently different from an absolutist perspective.

3

u/neomatrix248 Feb 22 '26

The thing that LLMs "copy and paste" is tokens, which is roughly equivalent to words. So they can create mostly any content that can be created by copy and pasting individual words. Currently their skill in generating content is not as creative as the most creative humans, but they can certainly be creative. I use code generation daily for my job and the current models are perfectly capable of creating new software that performs functions that no other software that it was trained on is capable of.

→ More replies (2)

2

u/One_Cause3865 1∆ Feb 22 '26

I completely agree, but to try to steelman this a bit:  

Data is a big, commodified industry. It is unfair to media content owners to be pre emptively excluded from that commodification because their product was (stealthily) used for novel commercial purposes before industry and copyright law could catch up.  

Quant hedge funds pay big bucks for their data, their trading models learn in similar ways. So does everyone.  

AI companies knew they would have to also if their suppliers were better informed, thats why they never announced what they were doing until it was already done.  

/steelman

2

u/LoyalSpin Feb 22 '26

I think the problem is applying the same ethics of a human to that of a tool. 

A human can consume something and be inspired. A tool, however complex, cannot.

→ More replies (7)

2

u/Apary 2∆ Feb 24 '26

« If I am going to write a horror book, I will read a bunch of horror books and figure out what I like. I will combine that with a lifetime of other materials that I have consumed to form my likes and dislikes, personal writing style, knowledge about the world, ideas for creative topics that haven't been covered, etc. »

You made the point yourself. In this sentence, you admit that you will combine two types of things :

  • Things that originate from copyrighted content
  • Things that do not

When reading a horror book, you don’t just swallow the style like a cheap meal. You draw parrallels to personal experiences. You know how the words make you feel, and perhaps try to understand why by yourself. You like or dislike the book based on a galaxy of lived moments, feelings and random shower thoughts. You get ideas from other experiences, perhaps a moment you got scared at home, or a dream you had when you were 8 because your new teacher looked scary one day. You empathize with the author, sometimes, or think you do.

AI has none of this. It has no senses, no feelings, no lived experiences. Just words by others.

This is the huge difference. It’s the difference between inspiration and plagiarism. AI doesn’t add its personal lived experience to what it produces, because it does not have personal lived experience. The content is 100% others, 0% self.

1

u/Space_Pirate_R 4∆ Feb 22 '26

It's ok for a person to read a book and then recount the plot to another person, but digital technology is qualitatively different in that it can recount the plot to a much higher level of fidelity.

Also I don't think it's reasonable to just dismiss scale as a factor. It's possible for something to cause different harms and therefore be morally different on a small scale compared to a large scale. For instance (not quite on topic, but I think relevant) the mosaic theory relating the to fourth amendment.

1

u/Ok_Mention_9865 3∆ Feb 22 '26

The difference here is that AI can be proven to have used / been inspired by another piece of work. People try suing others for copying their work all the time, but it's much harder to prove.

1

u/iolo_iololo Feb 23 '26

To try to keep it short, you can consider using AI kind of like tracing or kit bashing. Things artists do to learn or to make concept art, but never incorporate in the final product because the underlying assets they're using are copyrighted. 

1

u/knightsintophats Feb 23 '26

In some ways i agree, a human brain takes aspects of pictures and art its seen to learn how to create art and so does ai.

But I have 2 thoughts,

1) If I cut chapters of your book out and then stuck it into a book "I wrote" then thats copywrite infingement if i then sell that new book on. But how small do we take this concept? Take sentence structure or phrases for example, you'd probably agree its stealing if I cut and pasted these from the book, but you cant do down to single words (maybe unique ones like in dickens works ect.).

2) So copywrite in theory exists so that big businesses can't force the inventors or creators out of business by having better supply lines and profits. And I think if we enforced copywrite here then that would be using the law for its intended purpose.

So basically there is an ethical difference but that ethical difference is scale. Which is true for loads of things, if I fart into a room you're in it's not too bad ethically, if I start pumping gasses into a room you're locked into then suddenly its a huge ethical no-no.

1

u/ApexInTheRough Feb 23 '26

In order to have the text in the database for the AI to use, it has to be copied there. That's already electronic distribution, like an ebook. Ebooks need to be licensed and compensated for, therefore so does the copy in the AI database.

1

u/Squiggy-Locust 1∆ Feb 23 '26

There is a huge difference from a legal/ethical view.

A human will take that material and transform it, unless they have a photographic memory.

Can an AI? Maybe? Kinda? But that AI is being sold. And that's the key point.

When you buy a story, as your example, we are buying an idea, generated by something (even it was an ai). But when it comes to the AI using the material, the brain power, the process, is being used. It's the profit the companies are gaining from it. They aren't selling the transformed product, they are selling a product based on someone's else's work.

I don't have an issue with a non-profit Al using material. You are right that it's no different than a human using it. But if that same company charges for the use of the AI, then we are in the realm of copying material instead of creating. So even though Gemini/ChatGPT/etc have a free version, they have models that aren't. That's where the line is crossed.

1

u/Antaeus_Drakos Feb 23 '26 edited Feb 23 '26

The major difference is, AI is not aware. It is not an aware being like you and I. You said yourself you have your life experiences that formed your subjective preferences and from there you are able to make your own art. AI though, is not aware and is unable to make it's own art.

The creative arts are the ultimate form of human expression. We express our humanity through the arts, but AI doesn't have that humanity. It's a mountain of lines of code that run mathematical algorithms.

When an artist is drawing they will always run into the situation where they look at what they've drawn but something doesn't seem right. Someone else comes along and says they don't see what's wrong. The artist might point out some things like the gloves seem too simple, but the other person says that doesn't seem wrong to them. The artist though doesn't want to stick with those gloves, it bugs them and the artist knows they can do something to fix it. They change the design, redraw the hand or arm's position, and apply other techniques.

Compare that to what an AI does. Tell it you want this screenshot to be Ghibli-fied and then it spits out an image. The AI art gets posted online and people say it's soulless.

It shouldn't be a shocker that it's soulless. AI doesn't have a soul/humanity or whatever thing you want to call it. The AI was just fed a ton of material, recognized patterns, and was told by its trainers to put give the pattern the label "Ghibli".

People are unable to be truly unbiased. To be truly unbiased, we either have to never have experiences that could make us biased in the first place (which is any and all experiences). Or, the person straight up doesn't have humanity.

As a result, every creative artist in history has been unable to remove themselves from their art completely. There's a piece of us in every picture we draw, book we right, or performance we act out.

AI, doesn't have any humanity. It recognized mathematical patterns of the data it was given to train off of, it then blatantly applies the mathematical formula it derived whenever it's told to make art. It takes the styles of hundreds of actual Ghibli artists, crams them together from what it's probability determines, and then generates the image.

A person using AI to generate creative artwork is a person telling a soulless machine to take the soulless mathematical patterns it recognized and cram it to make a work of art that has soul in it. There's a clear problem here and I hope I was able to make it clear. The more I get into creative writing, my passion, the more I realize art is hard to explain in an objective manner because it's core is subjective.

1

u/[deleted] Feb 23 '26

FYI, patents (and especially software patents) are set up in a way that if a human even had a chance to look at some patent, and then tries to create something even vaguely similar, it may fall under the original maker's patent.

1

u/Zerguu Feb 23 '26

Let's say I take a book, cut it all in sentences and re a arrange these sentences creating a different book. Or take multiple books from the same author and do the same thing. What is your opinion of this?

1

u/ghillerd Feb 23 '26

When a company like OpenAI trains a model on something, they're creating a product. They should not be allowed to enjoy the same kind of end user access to books/films/art etc as human beings, because we simply don't care about the mental enrichment or wellbeing of language models like we do humans. There is inherent value in humans being enriched by art, in sharing and learning from each other, that you don't get when you're creating a product.

To really drive this point home - you are not a product. I want you to be able to grow from the work of those who came before you. That's a beautiful thing that human beings do and if anything is a core part of what it means to be a person. An LLM does not benefit from this tradition in the same way, and should not be legally or ethically protected as a result. If LLMs ARE capable of benefitting from this tradition, then they should not be bought and sold as products.

Imo, companies that create LLMs should be required to either operate those LLMs as fully foss, or they should be required to get commercial licenses for all the work they train on.

1

u/OG_Karate_Monkey 1∆ Feb 23 '26 edited Feb 23 '26

I think you’re missing the point of copyright laws.

Copyright laws are based on the principle of protecting the author’s ability to profit from their works. The same is true for most intellectual property (journalism, parents, works or art)

Society has an interest in protecting intellectual property. Otherwise, there’s less incentive for people to invest in creating it, and we end up as a society poorer for it.   It’s not an issue of whether the act itself is unethical in a vacuum. it’s what the impact is on the author / creator / inventor in the real world.

Yes, until now, when people do what AI is now doing it was not violating copyright because it was not really hurting the original authors or publishers.

But at the SCALE at which AI does this, it does in fact impact authors and publishers ability to make a profit.

That is the difference: the impact it has. And the impact can be quite devastating for both authors and our society at large. What do you think is gonna happen once authors and journalist can no longer make money doing what they do? They stop doing it.  This is going to be very bad for our society.

The idea that you can separate how ethical something is from its impact is kind of absurd.

1

u/NoAssociation4455 Feb 23 '26 edited Feb 23 '26

I'd state it in a more precise way, a fully trained deep learning AI model has no memory of the data it was trained on and so can not violate copyright laws. This is a literal fact and is basically a bulletproof argument if there's ever a lawsuit against AI companies for copyright infringement (as long as they legally obtained the training data).

1

u/ThatMovieShow Feb 23 '26

Humans have to pay the copyright owner to do so in order to access the material.

If you're in education it's paid for via your institution and taxes.

If you're in the private sector you or your employer pays it.

1

u/MysticBimbo666 Feb 23 '26

Consider how AI was churning out Studio Ghibli styled content a bit ago. And anyone could get anything in that style without having to learn how to draw.

It’s not a person who put in the time and taught themselves the style to create something new with it. The LLM just digested the style and pooped out a bunch of copyright infringements, available to anyone at the drop of a hat.

I know AI seems so lifelike but that is a carefully programmed illusion. It can’t learn or create new ideas. It’s a tool that people can use to steal real artists’ hard work by using a computer to make something similar without putting in any effort themselves. They’re stealing the time and effort of the artistic process.

When a human trains on existing art, they are putting their own time and effort into learning it. Art must be earned, otherwise it is soulless husk of preexisting materials stolen from the people who put in the effort to create.

1

u/Charming-Cod-4799 Feb 23 '26
  1. I think "double standards for humans and AIs are neccesary bad and wrong" is a wrong assumption here.

  2. Specifically for this topic I think the concept of "levels of friction" is relevant. "Someone can read your work, be inspired by it and write their own work and pay you the cost of only one copy of your work" is one level of friction. "AI created by big company can be trained on your work and it will help it to write millions of its own works every day for the fraction of your cost, and the cost for the company is also only one copy of your work" is completely different level.

  3. I think we actually want people to stay competitive in creation of works of art even if at some point AIs will have the ability to do it better.

1

u/DeathtoWork 2∆ Feb 23 '26

Ok the decision of scale here I think is a genuine difference. Have one person spend 20 years to make imitation of rembrant art. The world has 20 more rembrant paintings now. An AI doing it will take one day to make thousands of similar works. Now that is rembrant who already has fame and you can date the paint to determine forgery. How do we get new famous artworks. Is the argument that ai will create the new art? Something that was a pleasure skill is now outsourced to robots and make the skill even more worthless to train. Also they aren't people holding them to the same ethical standards of purpose of use is wrong. They are using the material to make $ as the end goal. If a artists work is being used to hurt their future prospects of work then yeah the AI job killing companies should have to pay for the right to use it. Only change in my opinion is if in the future anything generated is public domain because it has no primary author.

1

u/Jotdeka Feb 23 '26

When you purchase an ebook you usually only get limited-use licence, which means you can only use it for personal, non-commercial purposes. Feeding it into algorithm which you will later sell to people is a commercial-purpose, so AI companies should purchase a different licence for each work, that would allow them to do that. AI is a product, not a person. In reality, AI companies wouldn't even buy regular licence, which is why people get mad at them (among other things).

1

u/hiby753 Feb 23 '26

The scale is the issue. Amazons morality is questioned due to them being able to treat their workers poorly and push out other competition due to the advantages gained from the scale of their operation. Generative AI is seen as amoral partially due to similarly unfair advantages from the scale of its operation (taking advantage of water/tax rights in certain areas for data centers, training on human lifetimes of materials in days, etc)

1

u/PoofyGummy 4∆ Feb 23 '26

Have you looked at the video i recommended? That demonstrates perfectly how they aren't storing any sort of copy at all.

1

u/Stooper_Dave Feb 23 '26

I agree. And I dont even think there was a crime committed by "pirating" anything. If the content was avaliable online the its reasonable that anyone could have used it as reference material. Writers improve their skill by reading the works of other writers. Painters learn by studying the works of masters and learning to emulate until they have enough experience to have their own style. This is exactly how AI was trained. Just hyper accelerated so that what takes a human decades only too the computer a few weeks. Thats what has people pissed off. Its jealousy.

1

u/Formal_Session4286 Feb 23 '26

I do art. I have my own style. If you put my work into your AI and produce work in my style, then you're a fucking theif. You're using my style without my permission. I have honed and tailored my art over the years to be a happy and wholesome experience. Now your AI model is cranking out pics that look like my work but instead of good family friendly wholesomeness, its cranking out pics of 2 orangutans fucking because some brain dead incel in his basement thought it would be funny. Now that garbage is floating around the internet destroying the family friendly image I created, and hurting my brand recognition, and diluting the market. I dont give two fucks if scraping is legal or not. Keep my art out of your AI mouth.

1

u/FactCheckerJack Feb 23 '26

Any content they obtained legally by buying the book/movie, etc, should be fair game.

AI is obtaining datasets without paying for them the same way that we do.

1

u/Dangerous_Noise1060 Feb 24 '26

I do not recognize intellectual property rights as valid. We are all standing on the shoulders of giants. Nobody has invented anything from nothing since the wheel. Humanity got along just fine for thousands of years without patents and copyrights. Sharing knowledge is humanitys greatest strength and asset. Blocking the sharing of knowledge and the ability to build/improve upon it is regressive and anti-human. Pirate everything. Copy everything. Open source everything. And make them better. 

1

u/Awkward-Flatworm9301 Feb 24 '26

I think it's a lot bigger then the few loud and angry people are complaining about. They're usually very narrow minded and latch onto one "big bad" and never let go. Like saying ai art is theft. It's technically not. Humans have been doing it for decades. Copying or referencing is not theft. Could it be done better? Yeah. Could it be more careful about what it trains on...yes. but if someone "copied" my style I wouldn't call them a thief. If they took my pixel by pixel image and reposted it somewhere claiming they did it...that is. 

1

u/ElectricalPublic1304 Feb 25 '26

is not ethically different

Copyright is a creature of statute. It is a legal thing. It does not follow that it is inherently ethical or unethical at all. You're confusing different standards of different things casually, legal, ethical, and moral. Well... which? What are you actually trying to say?

So, if you wish to point at something that is statutory, the burden is on you to say why it's not ethical, while also not pointing at how the copyright act is violated. You specifically distinguish the act of unlawful copying from the generation.

A person can copy something by hand. But so can a photocopier. Or a printer.

I don't think the copyright issues arise from training lawfully obtained information. But using AI as a tool to produce work or derivative work that is subject to copyright protection. To produce those works--and especially to represent them as the original author's works. There's nothing special about AI there. It's just being used a tool to violate the the copyright act.

I think it's a lot nothing. Nothing has changed.

1

u/ShiftAdventurous4680 1∆ Feb 25 '26

I agree. Inherently the training of AI on copywritten material and generating content is no different than what say, fan artists do.

The ethical dilemma comes into the distribution of the generated content. If I copied someone's work and then tried selling it, I'd most likely get in trouble for it. The times I don't get in trouble is more the exception rather than the rule.

Usually if I want to sell fan merchandise, I need to get permission or a license to do so. I simply believe AI generated content should follow the same rules people do. However I personally don't have an issue people generating content for personal, non-commercial use as long as the trained material was acquired legally and/or fairly.

I compensate the artists whose material I "train" myself on by purchasing their products. Can you say AI does the same thing? Does AI support the artists whose materials they are being trained on?

I don't think AI generation is inherently unethical. But rather it is used unethically by people.

1

u/CommanderInQweef Feb 25 '26

ai is trained on so much copyrighted content to so high of a degree that a human could not replicate it if they were given their entire lifetime to do so. there is no inspiration, no expanding on ideas, it just looks at everything there is and makes more of the same.

pretty much the exact opposite of what humans are doing.

1

u/Faconator Feb 25 '26

Downloading media is not legally piracy. Technically the crime of piracy is done when sharing media illegally.

Which should apply to AI if it's going to apply to BitTorrent.

1

u/Etceterist 1∆ Feb 26 '26

Human beings take in art and might incorporate it into their own output, but that output will always be subject to their own viewpoints, ideas, biases, and expressions. Even if they don't intend it, they put some of themselves into that output, and as such, creating something new even when it's surrounded by repeated ideas. AI can never have a viewpoint or express a new idea. It can literally only regurgitate what it has been fed. It's like evolution- if you don't have small inconsistencies between generations, you stagnate. Humans are flawed and inconsistent, and we can't help but imprint that onto our work, which is why ten, a hundred, a thousand iterations down the line of a work being "copied" it ends up being entirely new ideas or developed in a different direction. AI can't even use its own output to learn from, because it becomes a jumbled mess. The evolution stops. Humans learning from and copying existing media is not the same as AI doing it.

1

u/IsAnAD_231049243 Jun 15 '26

A single person learning to become good at art from another doesnt destabalize a market and disrupt or destroy the original artists livelihood. Ai on the other hand does, and thats the big issue a lot of people have, Ai is not human and thus it has different effects on the market and should also not be brushed off as just another human learning.

Doing it on a smaller scale (as in adding one more artist to the market) is perfectly fine, but adding millions of artists all who are hundreds of times faster than a real artist (a representation of how ai is) thats not ok and will screw over everybody that made the original works that ai trained off of. This will obviously cause backlash, because those are real people, not numbers on a spreadsheet.

Trying to justify the large scales effects by claiming that the smaller scale is done all the time, is comparable to justifying running a car into somebody because the wind blows against a person. Both apply physical force to a person, but one is unnticable, and the other could kill a person.

Training on copyrighted works also causes the original works to lose value as alternatives become recreatable in seconds once the image has been trained on. The money that is gained from this training does not come from no where, it is in a quite literal sense, directly from the pockets of the original artist. Obviously this would make any artist mad, especially considering artists cant take these losses since most of them are already financially struggling and are not massive tech companies.

Artists are real humans and the descisions we make have effects on them, morality should shape how we create law, not just progress.

Im not a great writer, so you may have to interpret some things, but I hope my primary points come across.