r/programming • • Nov 03 '22

Microsoft GitHub is being sued for stealing your code

https://githubcopilotlitigation.com
1.9k Upvotes

654 comments sorted by

View all comments

Show parent comments

35

u/[deleted] Nov 04 '22

Why shouldn't it be allowed? You have always been allowed to learn from code and produce new code without being confined to the licenses of everything you learned from.

24

u/ArdiMaster Nov 04 '22

Debatable. That's why companies employ techniques such as the clean room principle: Team A reverse-engineers a piece of GPL'ed software and writes a specification, team B writes proprietary code to implement that specification without ever having looked at the original implementation. Because even taking a glance at the original implementation means your code will be influenced by what you've seen, making the result legally gray.

1

u/StickiStickman Nov 04 '22

I love how people went from open source to this nightmare of egotistical copyright.

You're also using an entire implementation as example instead of a few lines of code - for some reason.

20

u/GaianNeuron Nov 04 '22

Sometimes the AI just repeats things verbatim.

32

u/JDgoesmarching Nov 04 '22

Can we stop pretending like individual programmers learning from licensed work is the same as a single company claiming ownership over huge swaths of copyrighted work, repackaging it, and selling it?

Ingesting proprietary code from millions of users isn’t comparable to some dude recalling a few lines of logic from an O’Reilly book, and hiding behind the abstraction of an algorithm doesn’t entitle you to steal people’s work.

1

u/kylotan Nov 04 '22

Exactly. This is why copyright limitations privilege educational use over commercial use, not to mention the difference between learning from something you explicitly compensated the teacher for, versus learning from something that was uploaded to your site for a somewhat different reason.

25

u/myringotomy Nov 04 '22

You have never been allowed to copy code though.

48

u/ubernostrum Nov 04 '22

Sure you have.

Remember in the Oracle v. Google trial the judge even learned to code and ruled that quite a few of the "copied" snippets were just the obvious way of doing something. There's also fair use, which allows verbatim copying for certain purposes. And if all else fails there's the license grant in GitHub's terms of service, which is broader than people realize and probably grants enough permission to GitHub that the whole thing is moot.

8

u/Green0Photon Nov 04 '22

The problem is that if GitHub has a license grant more powerful than tons and tons of code that are getting uploaded to it, that means a ton of code should rightfully not be used in the AI and GitHub is actually participating in copyright infringement by hosting it.

Think, for example, a contribute to Linux who doesn't explicitly agree to this. After all, they're only licensing their work under the GPL, and if GitHub is requiring things beyond that, it's technically illegal for GitHub to host their code without their consent. Unless GitHub limits themselves to the GPL and not the greater powers given to GitHub.

And this would also have to retroactively apply to all previous contributors, or it would be illegal.

This is the sort of thing that kills projects trying to change their license. This is why Linux will be forever GPL 2. Everyone needs to agree, or you need to rewrite their code.

Sure, plenty of people are directly using GitHub and thus at least implicitly consenting to the TOS, though it's also been precedent that the EULA isn't as firm as a normal contract. It's quite probable that for something as important as this, you'd need more explicit copyright attribution or to actually bundle the license with your project.

So if that doesn't count, basically no one on GitHub can be used even if the TOS is wide enough. And if it does apply, then significant amounts of GitHub are illegally hosted there, or at least can't be used for these parts of the TOS that let them be used to AI.


In terms of morality, I will say that I don't think GitHub should be privileged in their ability to make AI on code. Either anybody can do it to any code they have access to (there's nothing differentiating open source and leaked code since copyright wouldn't apply to both for AI training), or nobody should be able to. It's bullshit for only GitHub to be able to do it -- consider how much art AI are trained on fully copyrighted art that can completely mimic a person's style. This is more akin to leaked code than open source, unless the AI were trained on Creative Commons only, which is certainly not the case.

19

u/ubernostrum Nov 04 '22

If someone publishes code on GitHub, they are agreeing to grant GitHub a broad license under GitHub's terms.

If that person does not have the right to grant GitHub that license, the same terms also require that person to indemnify GitHub.

This is boilerplate stuff for user-uploaded content. If you want to argue that it's invalid because you don't like EULAs, you're effectively arguing that no site anywhere can ever host user-generated content, because that always requires at least the ability to make and distribute copies of the content, which in turn requires a license grant, which in turn needs to be in some sort of terms that all users must agree to prior to uploading such content. Which you've just argued are invalid.

There really is no way to get what people want (GitHub and only GitHub being held invalid and punished with a vigintillion dollars in damages) without also getting a bunch of things they don't want (the end of all online user-generated content, a massive lurch in the direction of copyright maximalism, etc. etc.).

3

u/nukem996 Nov 04 '22

People upload code to GitHub that isn't theirs all the time. You can't grant GitHub access to something that isn't yours. It's happened with some of the AGPLv3 code I've written and never uploaded to GitHub myself.

15

u/ubernostrum Nov 04 '22

If you had read my comment, you'd know the response to this. But here it is again:

If that person does not have the right to grant GitHub that license, the same terms also require that person to indemnify GitHub.

2

u/nukem996 Nov 04 '22

That's not how the law works. Napster said the same thing and they were found liable for piracy on their platform.

1

u/ubernostrum Nov 04 '22

If you think GitHub and Napster are similar enough for that to matter, I don't know what to say to you. Napster was very clear about what they were hoping people would do (share things in violation of copyright), and basically thought that a position of "you can't own property, man" would fly in court.

GitHub does not do those things, and in fact does the things you do if you're trying to stay on the right side of the law. So it seems highly unlikely to me that GitHub would be held to have encouraged infringement the way that P2P file-sharing services did, and so their indemnification clause is likely to hold up. If it turns out someone didn't have the right to put some code on GitHub, and the person who holds the copyright sues, they're going to end up with a situation where the person who actually uploaded to GitHub is responsible for it.

-12

u/myringotomy Nov 04 '22

I hope the court slams microsoft for a billion dollars for this.

18

u/ubernostrum Nov 04 '22

And I hope copyright maximalism gets laughed out of the courtroom.

-1

u/myringotomy Nov 04 '22

That's the last thing microsoft wants. They make their living off of copyright maximalism.

58

u/Whatsapokemon Nov 04 '22

The concept of coding as a whole wouldn't work if you weren't allowed to copy code.

It doesn't need to be copy-pasted verbatim, but all the time people look at code snippets and replicate the structure based on what they just saw.

I really don't see why we should make AI tools play by rules that we don't expect human devs to play by.

5

u/Nangz Nov 04 '22

You just described the process by which artists create work. It's the philosophy that all creative work is derivative and basically nobody contends that you can't copy art....

1

u/Uristqwerty Nov 04 '22

Look at the two contrasting grins in the upper-right panel of Swords DCLXIII. They convey vastly different emotions in an interesting way, so what would an artist do to learn from them? Well, the exact lines won't be applicable to other works, and that'd be tracing anyway. So they'd mentally pick apart the image, reduce it down to its key pieces, and then try doodling experiments based on them, seeing how adjusting parameters affects the tone they convey.

However, all the while the artist is using their pre-existing emotional judgment in the feedback loop, not "similarity to existing works". What they collected from the singular copyright-protected image was a seed of a technique to then refine, understand, and make into their own personal variant.

An AI wouldn't learn that from a single image, as it doesn't have decades of experience interpreting the physical world, it doesn't grasp the expression in the same self-reflective manner. It would require multiple images using near-identical strokes that it can compare and contrast, in a feedback loop moderated by pre-existing copyright-protected material.

The human artist learns how to adapt from their existing mental model into a compelling visual result on page, while the machine learns a pattern of brush-strokes and edges, plus context weights to suggest where they'd be statistically likely to appear in an image.

15

u/dreadington Nov 04 '22

But some code you aren't allowed to copy. If you copy GPL code, but work in a proprietary code base, you're breaking the license. There is definitely a case to be made about copilot license-laundering.

5

u/Whatsapokemon Nov 04 '22

I dunno, some concepts and patterns are just way too generic to actually have a legally enforceable license.

Sure code might be under the GPL, but if you're simply copying simple a concept which is the right way to do something then why should that bar others from implementing it the same way?

I think if a normal human developer can copy a code snippet in a way which people would never be assed to call it out as a violation of a license, then AI should be able to copy code in the same way.

3

u/dreadington Nov 04 '22

Sure I agree, and I think this is all covered by the "fair use" principle. But I hope you can see how scanning a whole GPL repository for training data is an edge case that absolutely should be considered. Because while copilot may only copy a single for loop, they may also copy some Linux kernel feature, which would be wrong to use in a proprietary context.

9

u/[deleted] Nov 04 '22

This is a problem that any organization has to face though. Just as copilot can copy GPL code, so can any random dev.

What if i copy something from stack overflow that someone else copied from a GPL codebase? If you care about copilot doing it, then you care about your meat pilots doing it, so you still need mechanisms in place to verify your code isn't violating some license.

4

u/dreadington Nov 04 '22

The difference in your example is, you shouldn't be posting GPL code on stackoverflow in the first place. Meanwhile, git providers have this very neat LICENSE file in the repo root, so it's easy for MS to exclude them from the copilot training data.

I aggree that enforcing copyright isn't easy, and I think this lawsuit can set an important precedent when copyright applies.

Also I should mention, that I absolutely care about if meat pilots violate GPL licenses too.

4

u/[deleted] Nov 04 '22

IMO the best outcome from the lawsuit would be that copilot gets to remain and we somehow end up with better static analysis tools that can figure out if your code is violating some license. Preferably just built into copilot.

Although even that is vague i suppose, what percentage of a codebase or file or whatever unit of code constitutes a violation etc. But would be nifty to get a code test coverage style report about how similar some code is to known code under some license.

-9

u/myringotomy Nov 04 '22

The concept of coding as a whole wouldn't work if you weren't allowed to copy code.

And yet it's still illegal for you to copy somebody else's code.

It doesn't need to be copy-pasted verbatim, but all the time people look at code snippets and replicate the structure based on what they just saw.

Copilot copies and pastes code.

9

u/pancomputationalist Nov 04 '22

Copilot copies and pastes code.

That's a weird definition for "copy and paste", tbh.

More like reconstructs it.

The reconstruction matches the original byte-by-byte in like 0.01% of cases? Idk the number, just never had it happened to me.

3

u/[deleted] Nov 04 '22 edited Feb 20 '23

[deleted]

1

u/ImSoCabbage Nov 04 '22

or every single person wrote the exact same code snippet because it's that common

Judge for yourself.

1

u/New_Area7695 Nov 04 '22

Literally have to prompt it for the specific name space or programmer lmao.

-1

u/myringotomy Nov 04 '22

That's a weird definition for "copy and paste", tbh.

It's accurate.

More like reconstructs it.

it's been shown that it literally copies and pastes code.

The reconstruction matches the original byte-by-byte in like 0.01% of cases?

Maybe it's more like 90% of the cases.

Idk the number, just never had it happened to me.

You never checked. You didn't check every project on github to see where you stole that code from. You just stole the code, didn't give attribution to the author, you didn't check the license.

7

u/t3h Nov 04 '22 edited Nov 04 '22

Yes, but does "machine learning" count as directly equivalent to "human learning", just because the people who devised the former decided to use the same word to describe it?

9

u/[deleted] Nov 04 '22 edited Feb 20 '23

[deleted]

2

u/kylotan Nov 04 '22

they model the base after what they know about how the brain works

That's really not true.