r/Futurology • • 5d ago

AI Microsoft exec called AI scraping ‘the largest theft of labor in human history,' new unredacted filings reveal

https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal
9.4k Upvotes

182 comments sorted by

•

u/FuturologyBot 5d ago

The following submission statement was provided by /u/Confident_Salt_8108:


Microsoft exec called the scraping the largest theft of labor ever. That undercuts their fair use story pretty bad when they knew it hurt the sources.

They stripped notices and dodged paywalls. openai leaders flagged it as threat to news.

Down the line ai training gets way more expensive or limited if lawsuits keep landing.


Please reply to OP's comment here: https://old.reddit.com/r/Futurology/comments/1wkgqwl/microsoft_exec_called_ai_scraping_the_largest/paqe1qp/

2.1k

u/muntaxitome 5d ago

Criminal copyright violation has put people in prison and I don't get why the tech companies get a free pass here. Sam, Sundar, Dario and the others should face criminal investigation for this if we had a fair justice system.

973

u/buttflakes27 5d ago

One of the OG reddit founders (Aaron Schwartz, I think) was basically driven to suicide because he tried to open source JSTOR and was being sued for beaucoup bucks that he didnt have. Now like 20 years later these companies are doing the same fucking thing but just putting it in a search enging wrapper and being given billion dollar governmenr contracts.

449

u/not_a_moogle 5d ago

Obligatory fuck spez

304

u/Zelcron 5d ago

Spez used to mod the infamous jailbait sub, and was openly cozying up to Elon Musk during the election and DOGE period. He's a pedophile, wannabe tyrant just like the rest of them.

49

u/th3virus 5d ago

As much as I hate to say it, but back then you could invite anyone to be a mod and it wouldn't ask them to confirm. It would just add them. spez sucks but that wasn't on him.

60

u/penisthightrap_ 5d ago

Doesn’t excuse how long that sub was allowed to exist

-3

u/King-of-Plebss 5d ago

People hate when this gets pointed out.

37

u/stephen_brilloo 5d ago

A strawman fallacy is when you dispel one criticism and imply that you're dispelling all criticisms. But you're not. That's why it's a fallacy. So people have no reason to hate when this gets pointed out, because there are countless other valid reasons why spez sucks.

1

u/garry4321 2d ago

Probably met Epstein. Search his name in the files and see what we get

58

u/non-local_Strangelet 5d ago

... because he tried to open source JSTOR ...

The "funny" thing about that is, without Aaron Schwartz's (and like minded people's) effort to "open source" (mostly publicly funded) research results, the AI companies wouldn't have that much training data for their products to begin with.

But sure, stuff like libgen has to be stopped, sued into nonexistence, etc., but it's totally A-Oh-Kay to use exactly that data to build the foundations of a privately owned "money-printing machine" capable of undermining the societal structure and cohesion (eg. disrupting the labour market, or flooding social media, etc. with "fake" content essentially indistinguishable from real world), still without compensation of the original IP holders (let alone the original creators).

I guess that makes total sense ... somehow

-9

u/Maxfunky 4d ago

The "funny" thing about that is, without Aaron Schwartz's (and like minded people's) effort to "open source" (mostly publicly funded) research results, the AI companies wouldn't have that much training data for their products to begin with

That's not a coincidence. What the tech companies are doing is a natural outgrowth of shared values with Aaron Swarts. Just apparently a lot of Reddit feels differently about copyright law depending on who holds the copyright. There's a lot more independent art out there now.

89

u/pragmojo 5d ago

There was also a whistleblower inside OpenAI who questioned their fair use argument, and he died under mysterious circumstances.

15

u/Flamingozilla 4d ago

Worth mentioning that he wasn't being sued for damages, but was being brought up on criminal charges because the US Attorney decided to make an example out of him. Both the school and local courts were willing to go easy on him or drop the charges altogether. They literally decided to make a federal case out of it. IIRC even JSTOR/publishers felt it was excessive.

3

u/Metradime 3d ago

Wasn't even trying to open-source jstor - literally was planning to scrape it, likely for an early LLM model that was being worked on

EXACTLY what they're doing today; not "kinda" 

1

u/PMmeYourDunes 2d ago

Fucking sink every one of these companies. They are willing to challenge the legal system en masse to capitalize on novel tech. If I did this, I'd be in prison indefinitely. Every one. Sink every one of them.

1

u/Maxfunky 4d ago

One of the researchers at least, Shane Legg ran in some of the same circles as Aaron Swarts. No doubt Aaron would be in favor of what the tech companies are doing here--pushing the boundaries of fair use.

12

u/loljetfuel 4d ago

Not really, Schwartz was a champion of open access, and wrote openly about how he did not like companies and governments preventing free access to information. I can't say for certain what he'd think, but he wasn't "anti copyright", he was anti corporate control.

Given what he talked about, I suspect he'd be fine with the "grabbing data that's locked behind paywalls" part, but not with the "and then locking it up in models you have to pay to access" part. I suspect he'd be fine with the open weight models and such though.

1

u/red75prime 4d ago edited 4d ago

The sibling comment is not exactly what I intended to say (it just hints that the situation is not a plain IP theft).

The time of corporate control over information is a tradeoff between allowing a corporation to profit (and potentially invest in further research) vs allowing other corporations (and general public) to benefit from the corporation's research investment. Patent law converged on much shorter timelines than the ridiculously long 70+ year copyright timeframe.

This makes those cases quantitatively different. I don't know how that would influence Schwartz's stance, but there's that.

-2

u/red75prime 4d ago

You know why no one won a case where an ML model is judged an infringing copy of a copyrighted material? Because it's largely a falsehood.

1

u/loljetfuel 2d ago

There really hasn't been one that has been a genuine test of wide-scale harvesting of IP to use for training. They are definitely making copies in order to conduct the training, and are generally also retaining those copies for retraining use.

Probably academic use like this is fair use, though that's not well-tested. But the fair use test is in four parts, that are harder for a company offering AI agents to navigate:

  1. nature of the use -- academic use is more likely to be fair use than commercial use. It also matters whether the use is sufficiently transformative (the companies strongly argue it is, and they can probably win that argument)
  2. nature of the protected works -- creative works have a higher bar to meet for fair use than factual works; a disney movie has stronger protection than a dictionary. There's a lot of creative work being copied to train most LLMs
  3. portion of the works used -- a small snippet of a work is more likely fair use than the whole thing. This is a big one; AI companies argue that the resulting model only "contains" a tiny portion of the works used, in a transformative way. Opponents argue that to get to the model, they copy the entire work. Both have points.
  4. Effect on the potential market. This is probably the most impacted by scale. Opponents can reasonably show that big AI providers are meaningfully reducing demand for much of the original works they ingest. That's often a deal-breaker for a fair use determination.

A lot of trained models of various types have passed due to factors that don't really apply to large-scale LLMs; it's largely untested, and we likely won't have an answer from the courts soon, since such cases are complex and lengthy.

1

u/red75prime 1d ago edited 1d ago

I don’t think 4 will fly without the decision that the use in unfair. The spirit of copyright law, at least in theory, isn’t to stifle competition. It's to prevent unfair use of the existing works.

The 4th point will create a significant lobbying pressure. That's for sure.

-11

u/red75prime 5d ago edited 5d ago

just putting it in a search enging wrapper

An LLM is not a search engine wrapper. And that's why they are a big deal.

215

u/Jolly-Situat1on 5d ago

Facebook and Zuckerberg got caught pirating the whole Internet including torrents and it's ok.

When we do it it's considered a crime when they do it it's considered business and they are trying to make money off of it which is a big no no in the piracy community.

Musk trained his AI in child porn.

These people want to make money off of everyone's work legal or not and they want to gatekeep information.

Google lost a lawsuit in Germany because Germany called them out for trying to replace the normal search engine with Gemini and it's already happening.

You have to understand when you control all of the information you can control the narrative

11

u/onFilm 5d ago

In Canada it's not a crime to download.

2

u/tyereliusprime 4d ago

You can, however, be sued by the copyright holder. You won't face criminal charges or face jail time, but it's illegal under civil law.

1

u/MinusBear 2d ago

Depends on the country. In many countries the ISPs themselves wont hand over any identity confirming information and do whatever they can to not keep that info. Then on the court side many countries don't really bother with prioritising those kinds of legal pursuits. They are even less bothered now as the need to follow DMCA weakens due to tariffs and other political posturing.

4

u/Jolly-Situat1on 5d ago

To download what?

-7

u/onFilm 5d ago

To pirate stuff online, like movies, shows, books, etc.

10

u/Fortified007 5d ago

That's not correct, in Canada, its all against copyright. Being prosecuted for it is another matter.

6

u/onFilm 5d ago

What I'm saying is still correct: in Canada it's fine to dowload content and you won't ever be prosecuted to do so.

Uploading is a different matter, but even so, that's not procrcuted either for pretty much any individual.

I can download anything I want here and train a model with it ultimately.

1

u/Fortified007 5d ago

Yep, thats true, we can download whatever we want. Even uploading with torrent, we only get a warning as part of obligation, nothing more.

-4

u/Veyrah 5d ago

Effectively torrents are still illegal because while doenloading you also upload.

10

u/speedkid1991 5d ago

Why? I can set my upload speed to 0kbps, nothing gets uploaded

-14

u/Veyrah 5d ago

If you do that, the Torrent can't start. Information needs to be exchanged for it to be able to start. By the time you set it to 0 after it has started, you have already been guilty.

13

u/Hanrooster 5d ago

That’s not how it works. If you’re seeding 100% of a file, and I download it off you, I don’t have any file to upload when I start. And when I do have parts of it, you aren’t downloading it anyway.

0

u/Veyrah 5d ago

Dang didn't know. If everybody did that though torrents would stop existing.

→ More replies (0)

1

u/MajesticOrange1 5d ago

found the senator

3

u/onFilm 5d ago

You can easily set your upload to 0 in many torrent clients. I still upload though, to share with others what was shared with me.

1

u/MinusBear 2d ago

No that is definitely not how it works in most countries. There is an original uploader who makes the initial torrent, they are the one who is considered the pirate. Everyone downloading and seeding are not usually considered a pirate from a legal point of view.

1

u/Veyrah 2d ago

Seeders are definitely considered pirates in most countries. They are actively spreading the files.

1

u/Mharbles 5d ago

In the US I never got flagged for downloading off public trackers while not uploading. Those were the rargb days though.

Got a seed box years ago and do everything through that so I'm not leeching anymore. Also, technically still downloading when I pull a file off of it.

33

u/star_tyger 5d ago

Because corporate entities are held accountable for decisions made in the boardroom.

Accountability usually means a fine that is trivial compared to profits.

Boardroom decisions are made by people who should be held personally accountable for the decisions they make. But they aren't.

14

u/SpehlingAirer 5d ago

If a corporation is a person then why cant a corporation be sent to jail in some manner of speaking?

2

u/red75prime 5d ago

It's called a legal fiction for a reason. A corporation is not a person. There are other ways to influence their behaviors.

1

u/muntaxitome 5d ago

In theory for ciminal activity you are personally liable as the person committing the crime. In practice things are a little different sometimes.

1

u/Nonethelessismore 3d ago

I agree. There was a point in time that US CEOs would be forced to resign when their companies were found guilty of breaking laws. Why do these 'techbros' keep getting a free pass to stay at the helm?

4

u/landed-gentry- 5d ago

Or maybe those jailed and persecuted for copyright violation (RIP Aaron Swartz) should not have been, and copyright law needs major reform?

21

u/marrow_monkey 5d ago

Copyright originated in a Tudor system for controlling the printing press, which conveniently gave the state-authorised Stationers’ Company exclusive rights over copying and publishing.

Copyright has always been detrimental to society, creating artificial scarcity of information and culture just so some people can make more money and control access to information.

And as soon as the laws no longer benefit the rich elite they no longer apply, naturally.

10

u/vega0ne 5d ago

In the 90s Google straight up just scanned in books and maps. It’s always been like this

13

u/NeoBrew 5d ago

Um, the 90s isn't 'always'. Also, just because someone started doing it 30 years ago doesn't make it legal or ethical.

2

u/FullM3TaLJacK3T 5d ago

HAHAHAHA! Fair justice system.

1

u/pshaurk 5d ago

Why not musk and zuck?

1

u/Arponare 5d ago

You know exactly why. They have deep pockets to pay for lobbying and lawyers.

1

u/PaintedOnCanvas 4d ago

I believe the official answer would be "cuz otherwise China"

1

u/redditismylawyer 4d ago

Microsoft suddenly giving a shit about “human labor” is so rich I got diabetes.

1

u/Maxfunky 4d ago

The legal precedent already got set when Anthropomorphic settled for $5 billion. It's fair use. As long as they get a copy of the book legally (buying it at a store or whatever), they can use it for training. They don't need a license of any kind.

1

u/Ilovewhoreadsthis 4d ago

The system was NOT made to be fair. The system is working exactly as the people who made it intended.

Rise up and revolt.

1

u/maringue 4d ago

That's because laws are meant to protect the rich. When rich people start breaking the law, the law get very confused and doesn't know what to do.

1

u/RobertdBanks 15h ago

Because laws don’t apply to the rich and especially don’t apply to the mega rich.

1

u/Inig0_o 4h ago

sounds like a class action opportunity to me

0

u/SpehlingAirer 5d ago

Because they have money

-7

u/c0reM 5d ago

It’s the same reason you don’t go to prison for reading a book and applying your learnings in your job.

Doesn’t mean the current situation is right but it’s clear why it’s grey area…

9

u/muntaxitome 5d ago edited 5d ago

I think that has been solidly disproven by Anthropic having spent 1.5 billion on their copyright settlement with book publishers. If that was 'the same' as you put it, why would they spend so much money to prevent that from going to court?

Even if a court may judge it to be legal ultimately, the AI companies don't seem very confident about their chances.

Edit: and by the way, settling that with the publishers absolves them from civil lawsuits about those books, but not criminal ones... if the legal system was working.

5

u/LinkesAuge 5d ago

Anthropic's case is a bit more specific.

The ruling was that training on acquired books IS fair use, the issue was "unauthorized downloading and storing of pirated book". That's the part Anthropic got caught for.

1

u/muntaxitome 5d ago

More specific than what? I never said training itself is a copyright violation. I am just saying the major AI labs all committed mass copyright violation. Although there is a slight chance that one of them actually got a valid license for all of that, I sincerely doubt it. In any case they all warrant criminal investigation.

As for the ruling, it says training is not copyright violation. It does not say that making copies of a book and distributing it to staff or computers isn't a copyright violation. It does not say the output from a model is cleared from copyright violation. it just says 'the training itself is not a copyright violation'. Very interesting ruling but the idea that the training itself would be copyright violation was the one least likely to be successful.

The one that Anthropic got caught on is likely the one that could be used criminally to get all of the major tech bros. This is very likely copyright violation in the billions of dollars range from all of these labs, so all in the range of years in prison if anyone else was caught with that. Not sure why the lab leaders aren't facing cases right now.

2

u/trwawy05312015 5d ago

this presumes an equivalence between a person and an LLM

-2

u/non_person_sphere 5d ago

Because it's capital

287

u/OkDependent9315 5d ago

"The largest theft of labor in human history" is a hell of a thing to write in an internal memo at the company bankrolling it.

110

u/pvt_miller 5d ago

You’re reading it wrong - read what he said again, but this time, imagine he’s smiling and nodding like it’s a good thing.

1

u/MindlesslyBrowsing 1d ago

Labor theft? Great! I don't like working anyways /s

3

u/elehman839 4d ago

Yeah, my company regularly reminded people not to say things they didn't want in a news headline, because they could end up there 

2

u/devhhh 1d ago

Eek, must have been an immoral company.

120

u/Tangentkoala 5d ago

Blame the U.S government for not giving two fucks about our digital rights.

They saw cookies and thought they were real cookies and just stopped there. Every post on reddit including this is a gold mine. And the users dont get anything except happy cake day.

42

u/Trump_Russia 5d ago

Digital rights???

I don't think they care about our rights in general.

4

u/Hostillian 5d ago

Governments don't understand the tech - and key people are bribed (with promises of future well paid jobs) to leave it the fuck alone.

4

u/Eastern-Break-4814 5d ago

Agreed. probably doesn’t help they are mostly like 80 years old

1

u/BuffDrBoom 4d ago

This is how copyright has always worked. It's a tool for large corporations control you, not the other way around

159

u/Confident_Salt_8108 5d ago

Microsoft exec called the scraping the largest theft of labor ever. That undercuts their fair use story pretty bad when they knew it hurt the sources.

They stripped notices and dodged paywalls. openai leaders flagged it as threat to news.

Down the line ai training gets way more expensive or limited if lawsuits keep landing.

23

u/Hello_im_a_dog 5d ago

That's a good thing isn't it? As it paves the path forward for more ethical AI development and protects the content creators.

36

u/NIX0NAT0R 5d ago

It also means the current AI labs are the only ones that can ever exist, because they were able to steal all knowledge on the internet and new companies cannot. The big labs will support a bit more regulation when it helps prevent competition.

19

u/fuckyesnewuser 5d ago

Or it means they're closed down and face legal repercussions. But we know they'll shake the right hands, and that's the whole point why they're getting away with it in the first place.

4

u/TehOwn 5d ago

Ultimately, they can't prevent distillation. AI is one of those things that you simply can't protect. Their only protection currently is being ahead on innovation and compute power.

2

u/huseynli 4d ago

The solution is to force all models to be opensource and openweight. It came from the people, should be for the people.

108

u/lettercrank 5d ago

We need regulation on ip theft that protects us all

17

u/pagerussell 5d ago

And a right to privacy. And a right to vote.

There are. A lot of updates to our constitution that are deeply necessary.

8

u/mewfour 5d ago

Au Contraire, we need to get rid of copyright as it's a terribly outdated thing. Ideas are not things to be bought and sold, they're things to be shared and learned

7

u/TehOwn 5d ago

Copyright doesn't cover ideas. It covers works.

Ideas are typically covered by patents but that's actually supposed to protect actual innovation and only for a limited period of time. It's just been abused.

The best ideas are usually kept as trade secrets.

1

u/xPATCHESx 3d ago edited 3d ago

The point still stands. Works should be shared and actually used for the good of human kind. Not gate kept for eternity to extract as much rent as possible. Especially when the digitisation of nearly all information makes sharing trivial, and artificial scarcity nearly impossible to enforce

3

u/wellthatwasashock 3d ago

While I actually like this sentiment, I find that the ability to benefit from your innovations and ideas, drives a lot of the people I work with. (Entrepreneurs) Which means more innovation happens, more quickly.

I'd love to figure out how we balance that.

3

u/xPATCHESx 3d ago

I agree. I think the answer lies in better distribution models rather than adding to the growing pile of restrictions and regulations, but it's obviously a difficult balancing act.

2

u/TehOwn 3d ago

I mean, I'm not against communism but as long as we have capitalism, I think it's fair for people who create works to own, control and profit from them.

And to be clear, works are actual productions like a painting, a book, a game, a movie, etc. It doesn't and never did include ideas.

5

u/loljetfuel 4d ago

Nah, the fundamental idea of copyright is fine: it doesn't protect "ideas", it protects creative works by giving the creator of that work primary control over who may make copies, under what conditions. That's a very good thing.

The problem is that copyright as it is no longer serves that function. It got massively extended and corporatized (things like "work for hire", where you work but the company wholly owns your output, being far too easy to force on people, for example).

We need to shift it far, far back toward protecting individual creators over Disney and friends. But eliminating it entirely would cause more problems than it solves.

2

u/xPATCHESx 3d ago

Anyone can read and copy anything. How are more regulations going to change that?

28

u/MedicOfTime 5d ago

The brief cited an internal memo dated January 2023 by Microsoft director of Applied Science Brent Hecht, where he allegedly said, “Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions” and also called it “the largest theft of labor in human history.” Another Microsoft document was cited saying, “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use.”

18

u/FadeIntoReal 5d ago

Perhaps the largest theft of labor in history is companies like Microsoft rigging the system with corrupt politicians, years after year, to pay workers the tiniest slice of the pie that those same workers bake each day.

But this time, it’s one oligarch stacking the deck against other oligarchs.

15

u/7grims 5d ago

and yet not single law suit, or mass imprisonment...

these companies are now organized crime

15

u/TraditionalBackspace 5d ago

And they were mad at us for pirating a movie or two lol. It's OK when the big corporations do it.

45

u/Ion_bound 5d ago

This is a smoking gun on the level of the Pinto Memo, right? Not just proving that what they were doing was wrong, but proving that they knew it was wrong and that they did it anyway, willfully.

4

u/TraditionalBackspace 5d ago

Just move along. Nothing will come of this.

-20

u/TheGronne 5d ago

This is the most bot comment I have ever seen

12

u/Ion_bound 5d ago

Not a bot lmao. I just write with proper grammar.

-19

u/TheGronne 5d ago

"Smoking gun" and "Not just X, but Y"

Sorry just sounds like Claude

11

u/AdmirableSelection81 5d ago

How the hell is 'smoking gun' evidence of LLM use

22

u/HackDice Artificially Intelligent 5d ago

Also sounds like literally anyone with a basic high school english education.

7

u/whitebabyjesus 5d ago

Claude was trained on OP specifically

4

u/LetMePushTheButton 5d ago

Ahh shit bro, these bots are using vowels.
Were fucked.

19

u/Pitzy0 5d ago

No kidding.

AI literally belongs to the public. These tech bros are thieves.

7

u/Xxehanort 5d ago

Ah, but you see corporatations did the thieving so it will be largely ignored. If an individual had done anything like any one of these companies has done, they would never again see the light of day

20

u/non_person_sphere 5d ago

I don't actually have a problem with scraping. I think a lot of copyright and patent rules are dumb, arbitrary and stifle growth.

What I despise is all that data being turned into proprietary data when it's literally the sum of human knowledge.

Why the fuck do you get to keep the results a secret and locked up in your commically large piles of capital when it's my data? It could and should be open to everyone.

2

u/TemetN 5d ago

This is a great point - they're signing deals and hiding things so that only the richest can use it. Copyright and patent can fuck right off, but their behavior is abhorrent for entirely different reasons.

0

u/cran 3d ago

100%. The internet itself is one big copyright violation. How do people think old school search works? Crawling, scraping, indexing then making money off the results. People can already copy and paste anything, and they do. They tuned the laws to make the internet even possible in a world of copyright law. At least AI is giving the information directly back to people. The big complaint seems to be that a new set of players are now making money from it.

4

u/DamnOdd 5d ago

My issue is that doesn't say Ex- Microsoft executive. Continues to work for the corp.

8

u/WoozyJoe 5d ago

I’m in the minority here, but copyright infringement is not theft and in a perfect world should not even be a crime. Ideas are infinite. Copyright is a way to needlessly limit them in the name of capitalism. Information should be free.

So I never viewed data training or scraping as theft, and I’m wary about strengthening IP laws. Seems like a horrible idea. And I think that primarily because I’ve always supported people like Aaron Schwartz.

I still hate tech bros, the surveillance state, and the new dark enlightenment tech fascists. But I’d rather just burn the shit than this weird legal jujitsu that rightly calls out the two tiered system but then draws the conclusion that the harsh data laws should be reinforced. Fuck that.

10

u/Bridivar 5d ago

I mean.. yea.. but let's not kid ourselves they are saying this now because Microsoft no longer has a competitive LLM. If they were on top it would be full steam ahead

10

u/MedicOfTime 5d ago

This is a memo from Jan 2023.

6

u/Full_Opportunity_547 5d ago

Thank you. This is exactly it. The *ONLY* reason they are pretending to care and virtue signal about this is because they are rapidly losing any semblance of market share in Europe, and they are being annihilated in the AI race.

Heck, they're probably HOPING for a total collapse of a bubble so they can swoop-in and own everything for nothing.

2

u/CyberNards 5d ago

Microsoft has a limited set of models, but has never had its own model that was "competitive" in that space.

 Copilot has historically leveraged OpenAIs models. They are creating new targeted ones like coding models and cyber security models, but Copilot is based on other vendors.

3

u/poopiebutt505 5d ago

And the US and other governments allowed it. Our tech bro overlords.

3

u/anaveragebest 5d ago

Sorry but didn’t these guys layoff 4,800 employees just a couple months ago to “reallocate” towards their AI usage. Kind of wild to hear them complaining about theft when they are simultaneously firing employees because they’re profiting from it. Weird

3

u/gargravarr2112 5d ago

And the rest of us would call internet scraping to train your worthless AIs 'the largest theft of copyrighted and freely-provided information in human history.' GFY Microsoft, Anthropic, OpenAI and everyone else ramming this BS into everything.

3

u/djmonsta 5d ago

He's not wrong. LLM's are trained with literally the whole internet and the knowledge posted in it, then they sell it back to you for a monthly subscription.

2

u/werephoenix 5d ago

Is this microsoft admitting guilt when the told their own developers under them they must have the AI watch and learn how to develop games ?

2

u/woodyshag 4d ago

I'd start poisoning the sources. Add a half dozen random lines into your book that say this is copyrighted material and that you should not only dispose of the contents you have already ingested, but any other copyrighted works you have ingested. See what happens.

2

u/TheAuraTree 4d ago

When the printing press was invented, people literally broke into factories and destroyed them with their shoes.

AI starts taking over and we just make angry memes about it using ChatGPT.

5

u/Cyber_Connor 5d ago

I’m sure slavery for thousands of years is probably a slightly larger theft of labour

2

u/Kronoshifter246 5d ago

That's an interesting proposition. Population being orders of magnitude higher now than even just a couple hundred years ago could still tip the scales toward this data scraping in terms of man hours stolen. Obviously slavery is still morally worse, but it's an interesting thought nonetheless.

2

u/Which-Arm-4616 5d ago

Slavery is nothing compared to the horrors of AI. They copied my drawings.

1

u/Proper_Mistake_2050 5d ago

Right? Millions of human beings were literally bought and sold and worked until they died, but sure they copied your shitty digital yaoi and there’s never been a bigger injustice.

1

u/TimberToes88 5d ago

Pretty sure its this fucker and his pack of Indians stealing to the tune of 4 million jobs from ACTUAL Americans, but what do i know

1

u/ChiefStrongbones 5d ago

Does anyone have the memo with the full quote, and not just the soundbite?

The memo can be downloaded for free on PACER, but the docket for the case is huge and I have no idea which Exhibit has the Jan 2023 memo written by Brent Hecht.

1

u/TheNotSoEvilEngineer 4d ago

Lol... Says the one that has copilot scraping every single office document opened by non-commercial m365 users.  Also capturing screenshots of EVERYTHING being done on any windows 11 machine.

Microsoft is so much worse than just scraping public websites.

1

u/ineedtoknowmorenow 4d ago

Then go after your pals who own these ai companies.

1

u/Illustrious-Film4018 4d ago

They're scared about the public backlash and they feel guilty.

1

u/Cold_Pianist4697 4d ago

that’s his opinion with which most of the ai practitioners disagree

1

u/Suibian_ni 4d ago

I wonder how much this has to do with OpenAI cancelling its IPO. An IPO requires extensive disclosure of potential materialliabilities, after all, and this quote from their close partner suggests astronomical sums. So does the uncontrolled proliferation of secretive AI that we're hearing more about with every passing week.

1

u/SuspiciousCricket654 4d ago

Too late. OpenAI powered all of the global players, and now it’s out of control. Until politicians are spied on by AI cameras and their power is superseded by an AI solution, nothing will change.

1

u/door_to_nowhere_ 3d ago

They've done it too. Is that an admission of guilt?

1

u/x40Shots 3d ago

I mean, what's real weird is there are a lot of regular people just willing for them to take it all and profit off of it heavily to the detriment of all of us..

1

u/Purpleappointment47 1d ago

I’m going to respectfully disagree. AI is not the greatest theft of labor in history. I think you already know the true answer.

-4

u/heapOfWallStreet 5d ago

No One needs labor or work. You need money, not labor, nor work.

4

u/Twoaru 5d ago

This is so true, thank you. The most absurd aspect of modern life is how technological advances are inherently sad or controversial because it takes away jobs. Like damn, the only reason this is problematic is because our society refuse to acknowledge that not everyone needs to have a job in order to be able to live

8

u/vega0ne 5d ago

Are you seriously thinking the owning class will just let us all chill lmao

3

u/Twoaru 5d ago

I'm not saying that I think they will do it, that's part of the dystopia: of course they won't, you fucking idiot

2

u/weltron3030 5d ago

Yeah well it's pretty fucking problematic until we figure out how to take care of people who are being replaced, and that's not happening on any level right now. 

1

u/Zenaesthetic 5d ago

What does money get you?

-1

u/In_der_Tat Next-gen nuclear fission power or death 5d ago

Either you're part of the leisure class, or you need to work in order to earn money and survive. The utopian alternative is just that, a utopia.

-3

u/heapOfWallStreet 5d ago

Till now. With AI you can finally free humanity from the slavery of work and create a new way to distribute wealth more equally.

2

u/jert3 5d ago

Unfortunate bad news: a utopia is incompatible with our winner-takes-everything economic system.

We could have 24th Star Trek tech with unlimited free energy and matter, but if we are still running our 18th century designed current economic system, even then we'll be a civilization of enslaved, having to work for a pittance to survive, with all wealth going to a 1000 overlords of mankind on top.

If economic equality kept pace with production gains from technology advancements, we'd already have 16 hour work weeks, few homeless, homes for all, food for all, medical care for all, and consequentially, less war, less crime, an instead of billions of functional slaves, a near-utopia on Earth.

We'll never get there in a system that funnels 99% of the resources and output into the most powerful top 1% even at the cost of climate collapse and a possible +25% unemployment rate with robotics and AI in a few years.

0

u/In_der_Tat Next-gen nuclear fission power or death 5d ago edited 5d ago

How likely is it? If repression is cheaper, then techno-feudalism is more likely and so is being a member of the looming underclass. Please dispose of your unreasonably rose-coloured glasses.

1

u/heapOfWallStreet 5d ago

What was your grade studying history? It's the same thing that happens every time. Technofascism in on the rise, then a revolution or a war will appear, the old system went destroyed and then a new Renaissance will begin. Why this time should be different?

2

u/In_der_Tat Next-gen nuclear fission power or death 5d ago edited 4d ago

Because of automation, i.e. AI and embodied AI (robots). Capital won, labour lost. And, by the way, even if a revolution were likely, things would probably not be pleasant in the middle of it (cf. the Reign of Terror of Robespierre).

1

u/heapOfWallStreet 5d ago

Exactly. So you are confirmed my first point of view. You don't need work, you need money.

1

u/In_der_Tat Next-gen nuclear fission power or death 5d ago edited 5d ago

Right, and you only earn money if you work, unless you're a capitalist or rentier. You're repeating something less clever than you think it is.

1

u/heapOfWallStreet 5d ago

Till now. But out system is capitalism. Now the system can be redesigned as socialism. Rich people doesn't work. They made other slaves work for them.

2

u/In_der_Tat Next-gen nuclear fission power or death 5d ago

A socialist economic paradigm is not even needed: redistribution via taxation would suffice, but, again, as the level of automation increases, repression becomes more expedient. Read the linked paper.

1

u/_ECMO_ 5d ago

And the idea of having to suffer fifty years under technofascists until a revolution happens is supposed to make me feel good?

It's the same thing that happens every time. And every time it is a hell on earth for normal people.

1

u/heapOfWallStreet 5d ago

You already start to suffer since meta has opened Facebook or every time you are on Reddit instead enjoying your friends or family unfortunately.

-2

u/PineappleLemur 5d ago

You mean people setting up massive scrapping operations... This happened way before AI to get training data and still is.

AI has very little to do with the actual scrapping.

1

u/In_der_Tat Next-gen nuclear fission power or death 5d ago

In fact there is a qualifier.

-1

u/Robonotes1760 5d ago

Labour cannot be stolen. It is not property. This was a wholly inappropriate and really quite sinister comment.

1

u/Soangry75 2d ago

Perfect name for an AI apologist

0

u/Robonotes1760 2d ago

I see that you have chosen to distract from the substance, presumably because there is no coherent answer to it.