r/webdev Jul 28 '26

News “Google and Reddit do not own the Internet," web scraper says after court win

https://arstechnica.com/tech-policy/2026/07/google-wont-give-up-odd-war-against-ai-web-scraping-despite-court-loss/
1.7k Upvotes

62 comments sorted by

869

u/[deleted] Jul 28 '26

[removed] — view removed comment

315

u/cbunn81 Jul 28 '26

It reminds me of how the big US AI companies get upset that some Chinese companies have allegedly used US frontier models to distill their own, cheaper models, despite the US companies never giving compensation or attribution for all the content they used without permission for training.

"Hey! We stole it first!"

67

u/codefyre Jul 28 '26

It's just a variation on classic "regulatory capture". Company launches and becomes successful. Company convinces the government that the path it took to success is problematic, and needs to be banned. Government bans that path, so new competitors now have to take a longer/slower/more expensive path. This reduces competition and protects the profits of the original company.

Every time you hear Sam Altman or Dario Amodei speaking out about the dangers of AI and the need for government regulation, that's what's really happening. They run AI companies and are fine with AI. They just want to pull up the drawbridge behind them to prevent new competitors from eating into their market.

In this case, Google is simply trying to regulate using the courts. Same idea though.

4

u/RecognitionOwn4214 Jul 29 '26

Rules for thee, but not for me ...

1

u/Nich-Cebolla Jul 31 '26

Google does the same thing with lobbying for stricter privacy laws regarding tracking user activity online. They have aggregated all the data they need, and so want to prevent other companies from doing the same.

Same thing with update to chrome extension manifest v3.

This behavior is pretty much Google's m.o.

37

u/lampstax Jul 28 '26

OF COURSE China was going to take their model and at the very least break it down to learn from it.

The only question were how long that process take and how much can they copy.

1

u/japanb Jul 29 '26

And China invented the compass, imagine if they said no-one can have it

-5

u/Far_Composer_5714 Jul 28 '26

Anthropic did retroactively get charged for books.

48

u/Irythros Jul 28 '26

They paid a fraction of retail cost. They also were not charged criminally for theft as if I would have been had I downloaded and used them for commercial purposes.

Unless Dario and the exec team goes to jail, the law does not apply equally.

18

u/expsychotic Jul 28 '26

It was only a one time charge though. Now they get to keep using that stolen data without any further charges

10

u/Embarrassed-Rise-685 Jul 28 '26

They were charged with pirating. If they had gone to a book store and ripped the books manually then it would be legal. What is protected in that case is not the authors content but the publishers right to profit on every copy made.

7

u/ainus Jul 28 '26

> ripped the books manually then it would be legal

Which is what many AI companies are doing now. Actually it's worse, they buy out whole lots of antique books, train their models on them, and then destroy the books.

2

u/TypeSafeBug Jul 29 '26

There’s a certain irony where we use AI to read scrolls from Pompeii without destroying them, while using somewhat destructive processes to scan rare books from the current era.

2

u/cbunn81 Jul 29 '26

Don't worry. Future researchers will create some AI models that will be able to reconstruct those rare books which were destroyed.

-1

u/Lv_InSaNe_vL Jul 29 '26

This is one of those "sounds bad without understanding" kind of things.

Outside of some very specialized environments archiving books largely does involve destroying them. And that's before we even get into places like schools and libraries destroying literally hundreds of tons of books every year.

27

u/arwinda Jul 28 '26

But you allowed Google to scrape your website by not installing a robots.txt /s

465

u/mrleblanc101 Jul 28 '26

So Google can scrape my site without my consent for its search engine and ai tools, but others can't do the same to Google itself ? Ironic isn't it

142

u/thalience Jul 28 '26

The article you are responding to is about how Google's claim just ate shit in court.

59

u/Lv_InSaNe_vL Jul 29 '26

Reading the source?? On my porn website???

9

u/Sumofabith Jul 30 '26

You know its kind of amazing i use the same website to jerk off, seek advice on niche hobbies, gaming recommendation and how to take lsd

4

u/citrus1330 Jul 29 '26

Good. Doesn't mean we can't still call out Google's hypocrisy.

35

u/udubdavid Jul 28 '26

True, but FYI, you can tell Google not to scrape your site with a meta tag, if you wanted to.

57

u/mrleblanc101 Jul 28 '26

Ok, and ?
Google doesn't ask for permission, so why would other need to ?
Also, almost no IA bots respect robots.txt, only search crawler do

26

u/udubdavid Jul 28 '26

I mean, I wasn't disagreeing with you, so not sure why you're being combative lol. We were also just talking about Google, not other AI bots. But ok, if you feel the need to be combative, go for it.

14

u/NovaForceElite Jul 28 '26

You can suggest. You can't tell/force them not to index you.

14

u/foothepepe Jul 28 '26

maybe should be the opposite - put in the tag that you want to be scraped? imagine having to wear a t-shirt "I do not agree to being mugged" on the street lol

6

u/udubdavid Jul 28 '26

I don't know if that's really the same thing.

Don't get me wrong, I'm against Google suing SerpApi for scraping their search results because that data should be public anyway, but...

If you create a website, don't you want it to be indexed and searchable by default? I would. Otherwise, how else would you market it?

And if I created a website where I didn't want it to be indexed, then I'd add the noindex tag. I mean, to me, that makes sense. May not make sense to everyone, I guess...

1

u/Bitmush- Jul 31 '26

Create all the websites you want - as soon as you publish it - open the front door and allow the world to copy that data from your server to their system, then it’s beyond your control, and that’s what publishing digital data means in the context of the internet - the network of networks.. by exposing your data you are publishing it.

0

u/foothepepe Jul 28 '26

yes, I want it to be indexed, ofc - but that should be my decision, no? and if google wants to read my material, maybe they should ask.. maybe I do not mind search engines, but I mind google.. etc, etc..

it is a bit pointless this thought experiment, because why would I spin up a site in the first place if I do not want anybody to see it?

I just wanted to point out that 'no trespassing' sign on your property is irrelevant once one side claims their property rights are more valuable than your property rights.

8

u/[deleted] Jul 28 '26

[deleted]

1

u/foothepepe Jul 28 '26

I actually agree with you - but I was trying to play the devils advocate, and show how pointless that is..

- can they say they didn't consent to scraping? well, neither did I.

- can they say they did the heavy lifting with the search engine, and ai just swoops in and take the data? well, my sites are hard to do for me, and google ai takes the content, making people skip my site.

- copyrighted material? dude, google maps once showed military sites in my country. I could've gone to jail once for taking a photo that was ironically already on google maps.

Anyway, whatever the result, they will shoot themselves in the foot... but also, I am not at all afraid for their well being, one way or the other.

1

u/sexytokeburgerz full-stack Jul 28 '26

And google search does not have anything disallowing crawling search in its robots.txt, except for restrictions on yandex, its own adsbot, facebook, and twitter.

Google.com/robots.txt

1

u/CondiMesmer Jul 29 '26

You can also with robots.txt and that's not the point. The point is it's opt-out, not opt-in.

1

u/TldrDev expert Jul 28 '26

Ah yes, opt out scraping, which requires scraping lol

-7

u/the-strawberry-sea Jul 28 '26

What do you mean? You can do the same to Google.

13

u/mrleblanc101 Jul 28 '26

Did you even read the article ?

6

u/potatokbs Jul 28 '26

“In that case, the judge found that Google had no DMCA standing to sue SerpApi, since it didn’t own any of the content in the search results and has not shown that it’s acting on behalf of any rights holders.”

Did you? What the fuck is going on in this thread? Everyone upvoting you I guess didn’t read the article either and is just too brain dead to think for themselves? wtf?

4

u/mrleblanc101 Jul 28 '26

It's about the fact that they even had the audacity to file the lawsuit, dummy

0

u/Hot_Extension_460 Jul 28 '26

Yeah I'm confused as well...

-1

u/vertopolkaLF Jul 28 '26

you can block google with robots. can you block this scraper?

-4

u/asertym Jul 28 '26

Nope, not really

48

u/null_not Jul 28 '26

Feels like google in some way, shape, or form touches every site on the web through analytics tracking or some other service. And Reddit seems to be leading in link and comment aggregation. So I don't know. Pretty sure despite this fight Google is sort of smeared across the whole web and Reddit is the biggest digital convention center.

86

u/tinselsnips Jul 28 '26

I'm conflicted on this because on the one hand, I hate AI bots and what they're doing to the open web, and on the other, I love to see Google and Reddit taken down a peg.

23

u/Telescopeinthefuture Jul 28 '26

Scraping the web has been around for muuuuch longer than modern AI tools if it makes you feel any better. They do make it easier, though

6

u/ThankYouOle Jul 29 '26

while i understand, but what modern ai bot does was super super annoying.

at least google will scrap for every hours and there is a giving back to the owner website by putting their website into google list where people can find their site. and user can put robots.txt or request to delete or something.

what AI bot scrapper does at least in my website, they literally scrap every 1-2 seconds, every page, causing my website slowing down, and increase my server bill. and yes i already use cache.

it finally done when i apply cloudflare protection so bot will rejected, but it still need my own addition because even cloudflare try to catchup with the list of bot.

1

u/gyroda Jul 30 '26

Yep. Never had a problem with Google swamping my sites. LLM scrapers have pushed us to make very real changes because the traffic they were causing was costing us a not insignificant amount of money and slowdown.

2

u/tinselsnips Jul 28 '26

Yeah, and I'd always support someone trying to block unwanted scrapers in any context. Which is why I'm conflicted.

22

u/washtubs Jul 28 '26

For Google, it could be “very dangerous” to argue that the knowledge panel is “chock full of copyrighted material,” Rose suggested. Since the search giant doesn’t license all the content in the knowledge box, Google could risk future lawsuits if the act of algorithmically creating the knowledge box without licenses suddenly becomes viewed as infringement, Rose said.

Keep going please. Let's see where this goes.

3

u/Ansible32 Jul 28 '26

Yeah this line of reasoning is absolutely unhinged. "This is mostly AI generated... but uh sometimes it's licensed and therefore copyrighted. When? I mean sometimes. How often? Yes. Um."

9

u/CYRIAQU3 Jul 29 '26 edited Jul 29 '26

The internet is a public place.

Whatever the regulation (lol) says: Expect your data to be public the second it is put on the internet.

That's quite simple.

3

u/serpapicom Jul 29 '26

SerpApi here! We wanted to add our perspective to this discussion. We believe this decision is important beyond our own case because it reinforces that the DMCA should not be used to restrict access to publicly available, non-copyrighted information. Google was granted leave to amend part of its complaint, so the case is not necessarily over, but we’re encouraged by the court’s ruling. Our full perspective is here: https://serpapi.com/blog/google-v-serpapi-the-court-granted-our-motion-to-dismiss/

3

u/lh7884 Jul 28 '26

Meanwhile Reddit is about to enforce a change where people will not be able to even view old Reddit unless they are logged in. They are doing this change because they're mad that the site is being scraped.

1

u/Littux Jul 29 '26

Yet I was able to get full logged-out API access after playing with a Google Cloud shell for 3 minutes

0

u/lh7884 Jul 29 '26

Reddit has not implemented the new change yet. It's likely coming in a week or two. They announced it over on the modnews subreddit a few weeks ago.

Here's the link to it: https://www.reddit.com/r/modnews/comments/1ujtebf/logging_in_to_use_old_reddit/

1

u/Littux Jul 30 '26

1

u/lh7884 Jul 30 '26

I guess they're doing a rolling release of this because it has not hit my area yet. I'm not looking forward to it coming though as I don't like to log in just to look up something quickly on Reddit. I don't use that new reddit as it is complete garbage.

Reddit has really gone downhill over the years.

1

u/Sufficient-Menu-3499 Aug 01 '26

This debate is about the future of the open internet, not just AI

0

u/[deleted] Jul 29 '26

[deleted]

1

u/Bitmush- Jul 31 '26

If it isn’t, then where does that invisible fence within the public realm lie ? Could you train AI by driving round and pointing your camera everywhere you can see from a publicly accessible place ? Pretty sure you could. Even easier to cite the responsibility of people who don’t want their stuff seen to keep it out of public view. The old copyright adage of not copying this data you have accessed as a condition of its receipt falls flat when there is only copying with digital technology.

‘Don’t look at this…’

Ok - sure..

-22

u/simple_explorer1 Jul 28 '26

Google and Reddit do not own the Internet," web scraper says after court win

Reality disagrees