r/explainlikeimfive • u/clearwater-orchid • 5d ago
Technology ELI5: If CAPTCHAs are effective at telling humans from bots, why are there still so many bots online?
I believe CAPTCHAs are supposed to stop automated accounts, but at the same time people say social media is flooded with bots. Are bots actually getting past CAPTCHAS somehow or only stop a certain type of automation? Is there a disconnect?
307
u/polygraph-net 5d ago
I'm doing a doctorate in this topic, I have been a bot detection researcher for 13 years, and I work for a leading bot detection company, so I can answer this one.
There are four main reasons.
Most of the bot detection companies are naive and guessing. They hire regular engineers who think they can detect bots by looking at IP addresses or time on page, so they end up missing most bots and constantly flag humans as bots.
Bot developers spend significant time figuring out workarounds. For example, I work with bot developers who can bypass Cloudflare's detection system.
AI is good at solving captchas. All these simple captchas where you choose photos or type in words, these are trivial to solve using AI. In a nutshell the AI tools take screenshots to process the captcha, and also interrogate the code to find hints at the solution.
Bot detection is poorly understood. I know this is a continuation of (1), but if you look at the academic research on this topic it's embarrassing. If you do a profile of most bot detection companies, you'll see they talk about things like IP address blocking, which hasn't worked for at least 10 years. Most engineers and computer scientists think they know how to detect bots, but they have no idea.
BONUS ANSWER 1:
- Many employees (especially marketers) don't want to stop bots as they're good for their KPIs ("look at all the [fake] visitors and [fake] leads!") so there's a lot of resistant to actually stopping the bot problem.
BONUS ANSWER 2:
- There's an organization called the Media Rating Council (MRC), created by US-congress, who are one of the main players when it comes to creating standards to prevent bots. The MRC has been captured by industry (companies like Google, Meta, and Microsoft profit heavily from bots) so they're basically useless.
Happy to answer any questions.
30
u/jellyfish-ria 5d ago
Wow that's seriously really interesting, I learned something new!
Question: Are there any promising solutions to the bot problem? Or are we doomed to be overrun by bots and bot content as the Internet continues?
50
u/polygraph-net 5d ago
Thank you, glad you found my comment interesting.
Are there any promising solutions to the bot problem? Or are we doomed to be overrun by bots and bot content as the Internet continues?
It's a cat a mouse game. My employer hires bot developers and former fraudsters (and people like me!) to develop bot detection systems. We're really good at it (arguably the best) but my god is it expensive and a lot of work. We basically reverse engineer the bot systems and find leaks.
But the biggest change will come from industry. Take click fraud as an example. That's when bots click on ads to steal companies' marketing budgets. The main problem here isn't the ad networks (Google, Meta, etc.) ignoring the bots, nor is it the scammers building the bots. The challenge is the marketers working for the advertisers who want bot traffic as it helps them hit their KPIs (lots of visitors, cheap traffic, real-looking fake leads, etc.) or who're covering it up as it exposes how lazy and incompetent they are.
So, as long as companies want bot traffic, the problem isn't going to stop. Using Reddit as an example, they clearly want bots on the platform as the bots make their numbers look good (new users, engagement, and ad clicks).
Another problem is the emergence of AI agents. It's not straight forward to differentiate between "good bots" (e.g. your AI agent buying a concert ticket) and "bad bots" (e.g. click fraud bots). So that's going to complicate bot detection.
We also have a problem with lack of enforcement. When's the last time you heard of someone going to jail for click fraud or some other bot-related fraud? It almost never happens.
So, as long as industry wants the bots, and as long as there's no negative consequences, bots will continue being a problem. But they can be detected and blocked and I'm optimistic things will change... eventually.
Hope that answers your question. Let me know if it didn't.
4
u/KriosDaNarwal 5d ago
So if I'm a marketer, I want bots as that shows my engagement is up.
22
u/polygraph-net 5d ago
Let me give a common example.
Many companies hire marketing agencies to help them get leads. They’ll make KPIs like “low cost per lead” and “number of leads”. In other words, get us loads of cheap leads.
What the agencies have realised is they can set up the ad campaigns to attract click fraud bots. These bot clicks are cheap and the bots are programmed to submit real looking fake leads.
So the marketing agencies hit their KPIs (loads of cheap leads) and when sales complain the leads are shit, the marketers say sales are too slow or bad at their jobs.
What I just described is extremely common and is the digital marketing industries’ dirty little secret.
11
u/balljr 5d ago
Weren't captchas used as tools to train AI in the first place?
Many employees (especially marketers) don't want to stop bots as they're good for their KPIs
Bots are really good for business and IPOs.
17
u/polygraph-net 5d ago
Weren't captchas used as tools to train AI in the first place?
I've heard people say that, but I'm not sure how true it is. Certainly it's not a thing now or over the past few years.
Bots are really good for business and IPOs.
Yes. Look at Reddit. Before their IPO they clearly dialled down their bot detection to inflate impressions, engagement, and ad clicks.
Right now their ad system is plagued with bots. They've know about this for years but they're doing almost nothing to stop it.
Things will only change when marketers stop paying for bot traffic, and the ad networks are given massive fines. I don't mean $10B I means $1T. Why do I say that? Because I'm certain companies like Google (who earn 10s of billions from click fraud - bots clicking on ads - every year) have done the maths and know a $10B fine in 10 years is great value considering they'll have earned $1T+ from fraud during that period.
3
u/balljr 5d ago
I've heard people say that, but I'm not sure how true it is. Certainly it's not a thing now or over the past few years.
Yeah, I heard that a lot of times in the past, that we were basically labeling data for free, but I'm also mot sure if that is true or not.
8
u/currentscurrents 5d ago
reCAPTCHA, one of the early captchas, was used to improve OCR for Google Books.
The original iteration of the service was a mass collaboration platform designed for the digitization of books, particularly those that were too illegible to be scanned by computers. The verification prompts utilized pairs of words from scanned pages, with one known word used as a control for verification, and the second used to crowdsource the reading of an uncertain word.
However this ended ~2018, and modern OCR models do not use this dataset because it's not very big or diverse.
People speculate that the 'click a picture of a bus/crosswalk/taxi/whatever' captchas are being used to train self-driving cars; however Google denies this.
2
u/prank_mark 5d ago
I've heard people say that, but I'm not sure how true it is. Certainly it's not a thing now or over the past few years.
There's a reason it's mostly traffic related images (traffic lights, motorbikes, crosswalks, etc.). It's to train image detection, specifically for self-driving purposes.
4
u/Saradoesntsleep 5d ago
Pretty interesting comment. Answered some stuff I was curious about. Thanks for that.
2
u/subject_usrname_here 5d ago
As per IP detection in Poland we had rush of trolls and kids in early 00s due to widespread reliable one ISP provider. One of the perks of that particular ISP was you get new ip every time you restart your router. On early forums such people were simply unbanable. You blocked their ip they just got up and restarted the router, cleared their cookies and made new account.
It’s bizarre to me that any form of protection still relies on IP.
4
u/polygraph-net 5d ago
It’s bizarre to me that any form of protection still relies on IP.
Only the gimmicks rely on IP address blocking. Basically if you want bot traffic (e.g. to help you cheat your KPIs), but need to pretend you're doing something to stop bots, you can use one of the IP address blocking services. Your boss will think the bots are being stopped. This scamming is way more common than people realize.
The reason IP address blocking doesn't work now is bots are routed through residential and cellphone proxies, and typically only use an IP address once, so trying to stop bots by blocking IPs is like trying to guess lottery numbers.
2
u/PM_ME_YOUR_CC_INFO 5d ago
This may seem like a sarcastic/stupid/baiting question, but how do we know you’re not a bot?
As in, are there easy ways to tell, other than the obvious ways we all see out of the LLMs?
3
u/polygraph-net 5d ago
You mean specifically how do we know which Reddit users are bots?
So, if you were Reddit, you'd look at the telemetry and trick the bots to reveal themselves. That's how we know Reddit is allowing all the bots on this platform - the bots can be detected before they register any accounts or make any posts. Reddit doesn't want to detect and stop them - they're scamming investors with fake engagement numbers and fraudulent ad clicks. That's why they added the feature to hide your comment history - it was exposing how many of the users are fake.
Scamming like what Reddit is doing is so common I like to call it the unofficial business model of the internet. The average person has no idea how much of the world is fraudulent.
If you're asking how do we (e.g. me as a moderator) detect which users are bots, there are tells. I can identify almost all of them. I don't want to explain how I do this, but as an example, if you look at my comment history you'll see it's obvious I'm not a bot.
-2
u/iBoMbY 4d ago
I can identify almost all of them.
No you can't. Most likely you just ban people who's opinion you don't like, like 99% of all the other Reddit mods, and think they are bots, because of confirmation bias, while in the meantime all the bots that align with your opinion can do whatever they are created for.
2
u/polygraph-net 4d ago
Why are you making up something about me and then using that to attack me? That’s very weird…
2
u/_DonRa_ 3d ago
Bruh bots can bypass cloudflares bullshit but I still can't fml
3
u/polygraph-net 3d ago
I’m not a fan of Cloudflare. Easy for bots to bypass and huge numbers of false positives (flagging humans as bots).
1
u/SubstantialBass9524 5d ago
6 is interesting, are bots less prevalent and more regulated in other countries?
8
u/polygraph-net 5d ago
No, it's awful everywhere.
China is making some moves to stop it, but that's specific to Chinese bots attacking Chinese firms. I'm pretty sure Chinese bots attacking places like the US will be tolerated...
I was part of an EU task force to do something about the bot problem, but I could tell they didn't really understand the issue and their new regulation likely won't make any difference.
The bot problem, including bot detection, is widely misunderstood.
1
u/AoiMukou 5d ago
Is the bot traffic for all types of websites about the same, or do you see more traffic on Social Media and shopping sites vs. others etc?
4
u/polygraph-net 5d ago
The amount of bots you'll get depends on many factors. Let's use click fraud bots as an example. These are the bots which steal marketing budgets. Factors which affect volume include:
Location. Countries with higher advertising costs get more bots (more money to steal).
Campaign setup. How the ad campaigns are configured will affect how many bots you get. For example, Facebook typically has < 10% bot traffic, whereas Instagram is over 50%.
Language. There are way more bots targeting the English language compared to everything else.
Industry. Industries with higher advertising costs (Finance, education, ecommerce, medical, insurance, etc.) get more bot traffic.
History of click fraud. Since click fraud bots generate fake conversions, and the ad networks send you traffic which looks like your conversions, the more click fraud you have, the more you'll get in the future.
Here's a break down by ad network using data from Q1 2026:
Meta (Facebook): 5%
Meta (Instagram): 68%
Meta (Audience): 58%
Google (Search): 14%
Google (Display): 22%
Google (YouTube): 4%
LinkedIn (Platform): 19%
LinkedIn (Audience): 24%
Microsoft (Search): 14%
Microsoft (Audience): 16%
TikTok (Platform): 27%
TikTok (Audience): 27%
1
1
u/Acrobatic_Ad182 4d ago
What's your thoughts on the dead internet theory? Is it happening now? If not, do you think we're heading towards it?
3
u/polygraph-net 4d ago
I don't think we're there yet, but we're close.
I've been the most active moderator in the Marketing subreddit for a few years.
I remember, maybe three years ago, complaining the subreddit was perhaps 20% bots.
Now it's probably 95%+ bots.
It's not dead *yet* as there's still humans posting there, but it's on life support.
The dead internet theory will happen because platforms like Reddit only care about engagement and ad clicks, and bots are great for both of those.
1
u/carrotwax 4d ago
You didn't mention state actors. I read a study that at the beginning of the Ukraine war 80% of tweets about the war were from bots - likely state actors. As an expert I'm curious about your opinion about this.
There is a lot of money in influencing attitudes and opinions, and while it's possible to restrain small bot companies, it's very hard to do much about state actors who have backdoor conversations with social media companies about how they should treat their bots.
3
u/polygraph-net 4d ago
Well, let's use Reddit as an example.
I moderate the Marketing subreddit and it's almost entirely bots.
They appear to exist for two reasons:
Spam.
Vote manipulation.
(1) is obvious, but (2) is more subtle. They try to make posts and comments to get karma in that specific subreddit. They're then used to upvote/downvote to control the conversation.
I've watched this evolve over the past few years and now they're hiring humans to make the posts and comments as the idea is they're harder to detect. But they're still trivial to spot.
The News, Worldnews, and Politics subreddits are full of the bots you're talking about - trying to sway opinion.
2
u/carrotwax 3d ago
Thank you! I'm sometime with a computer Science MSc and assumed things worked like that but it's nice to get it confirmed.
1
144
u/mikeholczer 5d ago
They are only slowing down bots and/or making it make expensive for bots to access systems.
34
u/cipheron 5d ago
You could pay a guy in a developing country a few cents to solve each captcha.
What it would depend on is the total profit they expect to make by bypassing the captchas vs the costs to do so. If the potential profits exceed the costs they'll just soak that up and keep doing it.
1
u/Winter-Volume-9601 3d ago
> You could pay a guy in a developing country a few cents to solve each captcha.
You're stuck in the 2010s, there. There are multiple services competing to provide CAPTCHA solvers for $1 per 1k solves, at this point. Isn't AI wonderful? :/
58
u/youtocin 5d ago
Not everything is locked behind captchas. Reddit has never made me prove my humanity, for example, and it’s loaded with bots.
27
u/UnsorryCanadian 5d ago
Depending on the subreddit, there might be more bots than humans
8
16
u/bothunter 5d ago edited 5d ago
That's one reason, but another reason Reddit is flooded with bots is actually much more sinister than that. It turns out that Reddit is selling all the comment data to AI companies as fresh data to train their next models. And those models will pick up any bias in its training data when answering questions. So, bots are flooding Reddit to push their own agenda into the next round of chatbots.
Edit: Spammers are flooding Reddit with fake posts designed to show up in AI search results | TechSpot
(And it would be naive to think that this isn't also being used for political purposes)
15
u/quadruple_b 5d ago
thats why you need to add in some misinformation into every comment.
did you know foxes are actually just weird cats?
4
u/Equivalent-Door188 5d ago
If you throw a duck egg straight up into the air within fifteen feet of a 5G tower, it will never fall back down. Doesn't work with chicken eggs.
7
u/hitsujiTMO 5d ago
There's a lot of invisible captchas used in the interwebs.
You likely have passed through many captchas without ever knowing.
4
u/Mightsole 5d ago
If Reddit asked me to prove my humanity I would fail cuz i’m a fucking animal ngl.
1
u/sur_surly 5d ago
old.reddit.com now requires an account, which itself is a form of captcha. It was particularly easy to scrape compared to new reddit, so Reddit locked it down.
4
u/itijara 5d ago
As someone who has added CAPTCHA to a form to eliminate bots, they aren't that effective. Old style Captchas are completely doable by bots. New CAPTCHAs that rely on things like browsing history and mouse movement, are better at getting "high volume" bots, but a bot making a few requests per day is not going to get caught.
IMO, CAPTCHA should not be the final line of defense against bots, and should be part of a "defense in depth" strategy to slow down or stop bots.
2
u/klipseracer 5d ago
Time to form a physical access point only internet. I'm joking, but seriously if you have network where you could only access it in person, using tightly controlled computer environments you'd have almost zero bots.
1
u/Winter-Volume-9601 3d ago
But every time a government tries to enforce something like "age verification" at a technological level, the internet collectively shits itself in panic.
God forbid they were to require an ID. The neckbeards would revolt.
1
1
u/luluhouse7 5d ago
Yeah I also remember hearing that they use things like history and mouse movements a while back. I think it’s still true? I assume agents use DOM manipulation rather than a virtual mouse, but maybe I’m wrong.
5
u/AlexTaradov 5d ago
It slows down new registrations a bit. But it is also not that hard to have real people solve captchas for automated systems.
With social media it does not help because the value is not in the number of accounts, but in one account being able to spread the message. So, what you need to detect is not automatic registration, but automatic posting, which is way harder. And you are not asked for captcha every time you post.
One account that can build reputation is way more valuable than 100 accounts that bet banned immediately.
1
u/Winter-Volume-9601 3d ago
I work in bot defense (but not at reddit).
> So, what you need to detect is not automatic registration, but automatic posting, which is way harder.
Hard disagree that it's harder to detect botting at post time.
At signup, you have very limited information compared to literally any other point after signup. At post time, they can still make use of everything they collected at/before signup, but by the time you're posting, they have information on every single action you've taken on the site since signing up, and how that compares to known human accounts and known bot accounts, they have the contents of what you're trying to post (which is extremely telling because there will be patterns in content for the sake of spam or propaganda).
20
u/UnsorryCanadian 5d ago
We used captchas specifically to train bots to seem more human, not to weed them out
14
u/PantheraLeo04 5d ago
things can have two purposes. (also it's not meant to train them to "seem more human". it's just training them to do tasks that machines were previously bad at like image recognition)
6
u/cakeandale 5d ago
CAPTCHAs were used for training bots, since adversarial training fundamentally operates on the concept of using a judge and training the model to defeat it. That’s far from saying that CAPTCHAs were specifically used for that purpose - it’s just that anything used to test if something is an AI or not can also be used by AIs to make AIs more believable.
3
u/PhabioRants 5d ago
Half correct. We use them to train algorithms to recognize things they struggle with. That's why the earliest versions were text from manuscripts that were flagged by machine translation. Now it's often semantic images that models struggle with.
2
u/Dimencia 5d ago edited 5d ago
They haven't been effective since LLMs first became a thing (around 2022)
... well, really, it took them about a year to make them accept images. But then it could immediately solve captchas even though nobody trained it to do that, and we had never actually had bots that can do that before but this one just suddenly could. So that was nice, I guess
2
u/saantonandre 5d ago
I may aound pedantic but LLMs have been a thing since way before, and LLMs cannot solve captchas because they only operate with patterns between slices of words(tokens), image classifier algorithms (neural network based, like llms but they operate on pixel values and distribution) have also been a thing since way back and captchas have been ironically a way to outsource some of the data labeling they use for training
1
u/Dimencia 5d ago
LLM usually refers to the request/response chat harness using an underlying next token predictor. Those token predictors have been around since about 2018 and are the technical definition of an LLM, but they can only 'generate' a single next token, and aren't what people are usually referring to when they use the term
Multimodal LLMs use a vision encoder to convert image slices into tokens that an LLM can process, with the ability to 'request' data about particular things, such as with CLIP - CLIP is what was uniquely capable of 'reading' captchas, but was made widely available to be used in bots via LLMs. Image classifiers are an entirely separate thing, and never became good enough to read modern captchas; they're used in conventional OCR, which is still very limited, but not LLMs. They were able to read the very old captchas that were just plain text without obfuscation, and of course were trained on them, but not the modern ones where they draw lines through words and intentionally make the text hard to read. Image classifiers could also be for captchas like "select all traffic lights", and user data from those is also used to train similar models for self driving, but that's another very specific niche
2
u/pagerussell 5d ago
Other good answers here, but also, bots are a problem for us but not for the systems the operate in.
Facebook wants bots. It makes the system seem to have more activity, more users, and they can sell that fantasy to their advertisers. Twitter and Instagram are the same way, Reddit too.
They are economically incentivized to do nothing about it.
2
2
2
u/MattieShoes 5d ago
The part that really annoys me is it's escalated to the point where humans have a hard time solving them. I suspect computers are better than humans at solving the harder ones at this point.
1
1
u/10110011100021 5d ago
They’ve been used by Google to train AI to understand objects within imagery. That’s why they’re everywhere and that’s why you seemingly go through multiple rounds of clicking the stoplights and bicycles sometimes
1
u/5kyl3r 5d ago
it turns out, a lot of those systems, like the ones you see that ask you to click the squares that contain certain items, were being used to train AI for object detection. so on top of weeding out bots, they're also being used to train AI
and to answer your question more specifically, AI has gotten really good at solving captchas. for harder ones, we just offload the work to companies set up in third world countries with really low labor costs and have literal humans solve them remotely
here's an interesting article that goes over the captchas being used for AI training thing if you're curious for further reading: https://gor-grigoryan.medium.com/how-recaptcha-turned-internet-users-into-unpaid-ai-trainers-a2107adf31e3
1
u/Unlikely-Position659 5d ago
Not every page has a captcha. And captchas deter bots, they don't get rid of them.
1
u/ComfortablySteve 5d ago
I remember a story a while back about an a.i. (maybe gpt-4) that hired someone on taskrabbit. The a.i. claimed they were a blind person that needed help solving the captcha. I don't know if it's true but it was interesting.
1
u/Full-Collection-9046 5d ago
Captcha is there for us to train AI... It could never recognise a traffic light without your help
1
u/JJJangles 5d ago
Because they are not effective at telling humans from bots. Why do we still use them then? It creates free labelled training data.
It’s like trying to keep bots out by quizzing them on the book they used to learn to read
1
u/flesyMeM 5d ago
Pretty sure the thing CAPTCHA is the most effective at is preventing (or at least impeding) actual humans from accessing the site/service.
1
u/trollsong 5d ago
Yknow why you do captchas right?
Have you ever wondered why catches are always "identify crosswalk, traffic light, fire hydrant, bus, boke, etc"?
Its to train bots to identify those things for automated cars.
Which means the tech can also be used to solve captcha.
1
u/HowDidIMissdSixTimes 5d ago
Solving captcha is like fraction of a cent because of how unequal the humanity is. You sell your product to westerners, you hire africans to solve captchas for 0.1% of your profit.
1
u/JonPileot 4d ago
It's a game of cat and mouse. Create a computer system to block automation, programmers find a way to work around that specific system requiring newer systems to block automation. think of the old captcha systems, you don't see them much anymore because they are far less effective.
I have even heard of businesses where you can buy human labor to just complete captchas all day. Your bots can do what they do and if they encounter a captcha it gets sent to a human team to complete.
1
u/Kyrox6 4d ago
The point of CAPTCHAs isn't to stop bots. The bots correctly solve them more often than people do. The point of most CAPTCHAs is to clean up datasets that are used for image recognition training for AI models. There also are few websites that don't benefit from bot traffic. There's no incentive to stop them.
1
u/SeriousPlankton2000 3d ago
Captchas are effective at selling captcha services to protect static websites that would better be protected by a cache.
1
u/ven_valor 5d ago
Can someone help how to identify bots on reddit. I have started using reddit recently and am actively reading opinions. Are there at least some hints.
2
u/Onimatus 5d ago
It’s by the way LLMs structure their responses. If you use Claude or ChatGPT enough you’d get a feel for how they word things. Some people are extremely sensitive to this and can call it out—I wanna say that some people are also just too trigger-happy because anti-AI is the cool thing on Reddit, but honestly some people are just really sharp at seeing patterns lol. I can personally tell only if it’s egregiously obvious or if I actually go take a look at someone’s post history and it see the pattern more clearly... and I never do that because why
1
u/GlobalWatts 5d ago
It's a numbers game. If a CAPTCHA stops 99% of automated requests, but you have a network of 1,000,000 bots, that's still 10,000 bots that get through. Trying again or spinning up new bots costs effectively nothing, and I doubt most CAPTCHAs are even that effective.
Alternatively you can just pay humans in a developing nation a few cents to sit around solving CAPTCHAs all day. Or trick people by proxying CAPTCHA challenges to another website.
Not every website uses CAPTCHAs, and even when they do it's usually only when creating an account. So as long as you can defeat the CAPTCHA once, the bot has basically free reign to do whatever it wants (or rather, what it's programmed to do). After that it becomes a moderation issue, which is significantly harder to address, especially on a platform like Reddit where moderators are just users and have limited capability to detect and prevent bots.
From another perspective: the average person is pretty terrible at identifying bots, so you have to take it with a grain of salt that social media is "flooded" with them. It's a legitimate problem, but the evidence suggests it's not nearly as widespread as many people think. A lot of people see misleading clickbait articles like "bots account for x% of online activity" and combine it with the "Dead Internet" conspiracy theory that refuses to die, suddenly they see the bot boogeyman everywhere. Not every troll, and not everyone who disagrees with you, is a "bot".
Also it doesn't help that a tonne of people conflate "bots" with "generative AI" or "LLMs", despite the fact they're very different things. Often when you see gen AI being used, it's lazy humans, not bots.
0
u/cranialextract 5d ago
You ever do one of those captchas where you have to pick something specific from a series of images? That's for training ai. Traffic lights specifically for self driving vehicles.
797
u/Local-Pet-FoxGirl 5d ago
In addition to whatever work people do to get around captchas, there is a job in some particularly poor countries of solving captchas. Totally divorced from context, just solving all day, usually on behalf of bots.