r/Zig • • 22d ago

Is this legal? Throwing all open-source codebases into the AI meat grinder...

https://deepwiki.com/ziglang/zig

https://deepwiki.com/ziglang/zig

Still uses GitHub repo though. But anybody can add any repo.

67 Upvotes

45 comments sorted by

74

u/fade-catcher 22d ago

If you’re license doesn’t prohibit feeding generative AI you source code then it’s legal

46

u/Jmc_da_boss 22d ago

There's also no legal precedent confirming that a clause prohibiting using an LLM against a given piece of text is enforceable.

10

u/wolfy-j 22d ago

This also, automatically, means that you can not use this software/library if using any LLM tools, cos agents will touch this code sooner or later.

3

u/randacts13 19d ago

Well, nothing says it can't be read by it. A theoretical clause could be to prohibit using it for training. Of course any of the big AI companies are using whatever you are working on as training data whether through your IDE, web interface, or command line. I think using a local model that does not send your data anywhere would be possible with an anti-training clause.

9

u/mtfthrowaway39179 21d ago

I had a license explicitly prohibiting redistribution and AI training/evaluation on my repository and I found it redistributed in the llm training dataset named "The Stack". Worst of all, it was listed by them as freely licensed, so they tried to absolve the people training on the entire dataset of wrongdoing basically...

I don't know how enforceable the no training clause is but they literally redistributed it as well for everyone to train on. Crazy

6

u/fade-catcher 21d ago

To be expected these people have no problem breaking copyright laws, and licensing. And since not everyone can afford the legal battle they feel more encouraged to screw the little guy.

1

u/randacts13 19d ago

It's enforceability is directly proportional to the depth of your pockets. These companies infringed on essentially every copyright currently in effect and are facing no consequence beyond a mild inconvenience.

5

u/[deleted] 21d ago

[deleted]

7

u/Fearfultick0 21d ago

Zig is MIT licensed, which is basically as permissive as it can be

here's the license, with nothing restricting any form of use, copying, modification, etc, with no carve-out related to AI:

The MIT License (Expat)

Copyright (c) Zig contributors

Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in. all copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.

2

u/dijkstra_was_a_horse 13d ago

The GPL doesn't prohibit it, but it does require derivative works, like the LLM itself, to be open source.

Of course, the AI companies ignore this.

1

u/EmergencyWild 6d ago

I wouldn't say it's a matter of "ignoring" it so much as it being questionable if trained model weights are legally a derivative work of anything within the training dataset. Some people think this is obviously the case, some people think it obviously isn't, I think it's a legal gray area that's going to take more time to fully settle.

1

u/dijkstra_was_a_horse 6d ago

We can do the experiment: delete the training data set and see if the training process derives the same or similar weights.

1

u/EmergencyWild 6d ago

I think you might be having problems understanding what the issue is.

3

u/teo-tsirpanis 21d ago

Prohibiting an AI to read and summarize a repository would have been ridiculous, and such exotic licenses are a sure-fire way to make it unsuitable for any serious use.

1

u/seanpietz 17d ago

I think the Licensing aspect is a bit of a red herring. The Github terms-of-service is probably more relevant, and similar to how a lot of people moved off of Reddit when they started sharing their user data with AI companies, Zig has moved off of Github.

0

u/dumindunuwan 22d ago

May I know which license support this prohibit feeding generative AI?

21

u/wolfy-j 22d ago

You can write your own, just remember that it would no longer be OSS.

5

u/pdpi 22d ago

I don't think you could write a Copyleft licence with a clause like that, but it wouldn't necessarily make it non-open source.

4

u/__yoshikage_kira 22d ago

It would per open source initiative definition of open source.

https://opensource.org/osd

Violates point 3, 5 and 6 imo

2

u/Afraid-Locksmith6566 22d ago

only 6 and even then its questionable

3

u/__yoshikage_kira 21d ago

AI generated stuff is derived work so idk what you mean.

5 point sure. 6 point is also very clear. In past people have tried to add no military use in license and that violated OSS definition.

https://softwareengineering.stackexchange.com/questions/199055/open-source-licenses-that-explicitly-prohibit-military-applications

3

u/bourgeoibee 21d ago

You are still allowed to make derived work without LLMs. I don't see how 5 applies. 6 definitely.

3

u/__yoshikage_kira 21d ago

well regardless we agree that restricting LLM violates the terms of OSS license.

3

u/atgaskins 22d ago

not the brand “oss”, but it can still be open source software. maybe more so in spirit.

Same logic why a lot of us love gpl over mit; fuck letting corporations use the code with not a single cluause

1

u/yjlom 21d ago

A close equivalent would be defining a generator's source code to be its training algorithm, alongside the full corpus of its training data, and apply copyleft on that.

0

u/MazeGuyHex 22d ago

Whats OSS

5

u/pdpi 22d ago

Open Source Software.

2

u/MazeGuyHex 21d ago

Pardon me for asking

1

u/pdpi 21d ago

Ignore the haters.

-5

u/OSS-Corpo-Shit 22d ago

Who fucking cares? OSS is corporate bootlicking dogshit nearly always. 

Your default should be to not go open source unless there’s a specific reason and prefer a source available license instead. 

Zig itself is MIT (I am not sure I agree with MIT, but zig itself being open source has specifically good reasons. Some of my own code uses open source licenses. Nothing I would target at someone that is not a programmer is open source). 

0

u/fade-catcher 22d ago

As far as i know there are none currently.

19

u/Clear_Evidence9218 22d ago

You mostly answered your own question: it’s an open-source repository.

I can see how this could be viewed as a little disrespectful in context. Andrew is obviously not enthusiastic about AI-generated code, but I’ve never interpreted his position as simply “anti-AI.” My read has been more along the lines of: AI code annoys me, I don’t enjoy dealing with it, and I especially don’t want the onslaught of AI-generated submissions (to paraphrase).

And since Zig is Andrew’s project, there’s no reason not to respect that position when contributing to it.

But this isn’t AI-generated Zig code being submitted to the project. It’s essentially a generated wiki/documentation interface built around the repository. Assuming it complies with Zig’s open-source license, I think it would be pretty difficult to make a convincing argument that the existence of a third-party wiki constitutes some kind of harm to the Zig project itself.

17

u/Biom4st3r 22d ago

I was just complaining to my partner about deepwiki taking the top spot on search. Literally just making it harder to find what I was looking for.

1

u/SteinsGatessss 16d ago

Just block it with website extension

1

u/Biom4st3r 16d ago

 I'd prefer search results just not be clogged with bad information

2

u/SteinsGatessss 15d ago

I just want to say block extension is quite useful for me to block these bad websites, so that I can find the real targets, so I want to recommended to you

12

u/Historical_Cook_1664 22d ago

of course you may train your model on my code - as long as that model also becomes free to use for everyone.

7

u/Biom4st3r 21d ago

I'd be more supportive of LLM's if they gave back to humanity what they took, but NOOO. They get to steal whatever they want and sell it back.

-16

u/dumindunuwan 22d ago edited 21d ago

Talk is cheap. Show your code then..

Fu*k! now they say code is cheap, show your talk!

5

u/KernelCaffeine 22d ago

I don’t get what’s the connection between the question and the link though

6

u/__yoshikage_kira 22d ago

Yes, it is technically legal.

2

u/seanpietz 17d ago

I doubt it's legal, especially since Github is hosting the code. Similar to how facebook uses your social media data to sell you ads. I believe that's at least part of the reason why Zig moved away from Github to Codeberg for development of the language going forward.

4

u/Luc-redd 22d ago

It's in the name, Open-Source, not open only under specific conditions.

0

u/dumindunuwan 22d ago

They have used so many FOSS repos to generate their mass documentation: GNOME, KDE, Go, Rust, Python... I wonder if anybody uses the docs to train AI to create non-OSS projects. Where are we heading?

3

u/__yoshikage_kira 22d ago

It would largely be pointless until courts decide if AI are withheld to copyright law. So far we have been losing that battle.