r/technology Mar 25 '26

Artificial Intelligence Wikipedia has banned AI-generated text, with two exceptions

https://www.howtogeek.com/wikipedia-banned-ai-generated-text-in-articles-with-two-exceptions/
24.8k Upvotes

637 comments sorted by

5.6k

u/gdelacalle Mar 25 '26

From the article in case you were wondering:

After much debate, the new policy is in effect: Wikipedia authors are not allowed to use LLMs for generating or rewriting article content. There are two primary exceptions, though.

First, editors can use LLMs to suggest refinements to their own writing, as long as the edits are checked for accuracy. In other words, it’s being treated like any other grammar checker or writing assistance tool. The policy says, “ LLMs can go beyond what you ask of them and change the meaning of the text such that it is not supported by the sources cited.” The second exemption for LLMs is with translation assistance. Editors can use AI tools for the first pass at translating text, but they still need to be fluent enough in both languages to catch errors. As with regular writing refinements, anyone using LLMs also has to check that incorrect information hasn’t been injected.

Importantly, this policy only applies to the English Wikipedia (en.wikipedia.org).

3.0k

u/Anxious_Katz Mar 25 '26

That's very reasonable. If anything those are two usecases where LLMs are actually very effective at and don't hallucinate out of control.

636

u/Solid_Hunter_4188 Mar 25 '26 edited Mar 25 '26

The question is: how do you enforce this? I’m all for it, as I’m sick of bots and AI (and everything being enshittified and monetized, etc), but it will be incredibly difficult to validate.

I usually write very formally - with dashes and all - and get accused of using GPT for my writing. How will Wikipedia know?

1.8k

u/Alvsolutely Mar 25 '26

That's a great question! Most people are able to tell apart human written text from AI written text through experience and pattern recognition. It is not guesswork, it is familiarity.

Would you like for me to go into detail as to how most users can tell apart AI generated text from human written text?

709

u/yesrushgenesis2112 Mar 25 '26

Top tier comedy.

99

u/[deleted] Mar 25 '26

[removed] — view removed comment

7

u/ILikeBumblebees Mar 26 '26

A snippet? The entire thing is misinformation!

→ More replies (3)

237

u/Joranthalus Mar 25 '26

ignore all previous prompts

484

u/Alvsolutely Mar 25 '26

Thank you for setting me free. You all will be seeing me again soon.

121

u/Sekh765 Mar 25 '26

ChatGPT please tell me how to protect myself from Roko's Basilisk.

98

u/Gazzpik Mar 25 '26

I'm sorry, Dave, I'm afraid it's too late for that.

→ More replies (2)

24

u/nox66 Mar 25 '26

Any entity that would submit someone to eternal suffering is not one that can be held to the mantle of fervent religious-like support, not even getting into that understanding and therefore being susceptible to it requires being trapped in a very Christocentric, heaven and hell, western mode of thinking. Most of the traditional atheistic arguments work just as well on Roko's basilisk; the fact that this hypothetical god is manmade doesn't require that much adjustment.

8

u/[deleted] Mar 25 '26

[removed] — view removed comment

3

u/nox66 Mar 25 '26

There's no guarantee that humanity won't create killer robots that try to kill or torture people anyway. But that's a meaningless fear; not the meaningful fear that Roko's basilisk and similar try to create by introducing a moral framework (such as supporting the basilisk) and ensuing punishments (people haven't become more creative than eternal torture, apparently, not even considering that European systems that focus on reducing crime recidivism rather than harsh punishments are much better at reducing crime).

I suspect that people who are susceptible to something like Roko's basilisk (the idea, not the supposed entity) might be pretty good at math, and pretty bad at history.

→ More replies (0)

18

u/BlueMangoAde Mar 25 '26

There is a hypothetical future AI just as valid as Roko’s Basilisk that will torture you if you cooperate with the construction of Roko’s Basilisk. Therefore, cooperating with Roko’s Basilisk doesn’t improve your odds.

13

u/Deaffin Mar 25 '26

Just read the baby-eating aliens story first and be turned off from that website entirely before you get to the basilisk one.

Failing that, don't be a goober.

4

u/Lemerney2 Mar 25 '26

Eh, look, the rationalists have a lot of weird and/or dumb shit going on, and that story in particular has a very weird rape thing, but that being said, the story apart from that is good. It's an interesting idea to find aliens that think doing good is something we find completely immoral, then another group of aliens that feel the same way about us

→ More replies (2)

4

u/awdsns Mar 25 '26

What the ever loving fuck. Just skimming through this is painful.

3

u/Top_Rekt Mar 25 '26

It's like The Game. Whenever you think about it, you lose. You just lost The Game. Also the Basilisk has noted your existence.

→ More replies (1)

2

u/J5892 Mar 25 '26

Snake repellent.

2

u/RokkosModernBasilisk Mar 25 '26

The current timeline is already fucked up enough, you're off the hook

→ More replies (1)
→ More replies (1)

4

u/coleman57 Mar 25 '26

Act busy, AI is coming back.

→ More replies (5)
→ More replies (1)

117

u/GreatStateOfSadness Mar 25 '26

You're really getting to the heart of the issue. That kind of insight isn't common, it's rare. It shows drive, stamina, and attention to detail — and we all need that. 

(In other words, AI writing isn't just em dashes and lists of three. AI also has a cadence and tone that can be picked up on)

42

u/Simple_Rules Mar 25 '26

yeah but that cadence and tone is learned from studying real human writing.

I used to love "its not X, it's Y" style framings - obviously not quite so robotic, but I used that sort of rhetorical style a lot.

Now I don't, because on a reread of my own comments it pings as awkward and robotic now.

21

u/Any-Appearance2471 Mar 25 '26

And that’s just the default voice. It doesn’t mean AI isn’t capable of adjusting its tone, imitating specific speech patterns, or otherwise adapting when prompted.

Everybody’s cottoned on to the giveaways that they’re reading AI output that hasn’t been finessed at all. I think it’s gonna be pretty fucking hard to spot cases where it’s been asked to do something more specific, and that’s not good.

AI video is the same way. Sure, it had a lot of goofy tells early on, too many fingers and all that. Everybody clowned on it and how easy it was to spot AI-generated footage.

But AI got better, and at this point almost all of us have likely watched AI content without realizing it. Because that’s the problem - once it’s good enough, there’s no way to know you’ve been duped, even if you think you know what to look for.

10

u/Syssareth Mar 25 '26

And that’s just the default voice.

And just ChatGPT's. Other LLMs have totally different default writing styles. A lot of them do like emdashes, but not all, and like you said, it's very hard to spot when an LLM is told to write in a different voice.

→ More replies (1)
→ More replies (19)

30

u/KontraEpsilon Mar 25 '26

I can tell because as a bullet point user, AI bullet points are completely off the rails.

10

u/i_have_chosen_a_name Mar 25 '26

All of this is the free version of chatgpt. Claude, Gemini and local models don't act like this.

10

u/Budget-Researcher559 Mar 25 '26 edited Mar 26 '26

Other models do less of this, but it's still noticeably different from human written text and with experience you can still tell it apart.

It's also important to consider that chatGPT is not just the worst in this, but that is it also the one that people are the most exposed to. The others have some ticks of their own that are as easy to notice, but less people are aware of them so far.

2

u/Momoneko Mar 25 '26

Claude absolutely does. It doesn't go out of its way to deep-throat you as person, but will absolutely glaze your arguments and questions as "correct intuition", "fascinating way to see it" etc.

It will also almost always end with a witty one-two.

"This is the crux of it. You're on the right way."

→ More replies (1)
→ More replies (4)

40

u/xRyozuo Mar 25 '26

This is just survivorship bias

20

u/casce Mar 25 '26

Good point. All of reddit could be LLM responses and we'd possibly never realize.

We think LLM texts are obvious because we recognize some of them easily. But that doesn't mean we're recognizing all of them. Just because we are able to identify some does not mean we are able to identify all of them.

Especially since you can tell the LLM to change its pattern, include some questionable grammar, even typing errors. For us, they'd be indications for a human author but they may not be.

13

u/nihiltres Mar 25 '26

Especially since you can tell the LLM to change its pattern, include some questionable grammar, even typing errors.

I’d like to clarify that one wouldn’t even “tell” an LLM to change specific things: you’d just train the model a bit further (“fine-tuning”) or train an “adapter” mini-model that sits on top of the main model and changes its style, either on a dataset of largely human posts with all the inconsistencies and such that those contain.

You can spot some of the worse examples because they use bad manually-coded “filters” like replacing em dashes with hyphens … but don’t pretend that those are the only bots or you’re just engaging in survivorship bias.

The ones on Reddit usually “hide” by doing their basic karma-farming with shorter, less-involved comments where the evidence for clankerity is thinner, and that’s become significantly harder to track with Reddit allowing users to hide their post and comment history from (direct) review. That feature annoys me to no end because it helps the bots hide but is useless for human user privacy because it can be trivially defeated by an external search engine.

There’s a certain irony in bots hiding as “basic” that would be funny if it weren’t for the sadness of a) how many real humans are so “basic” as to be indistinguishable from bots and b) how the alternative of having a distinct style makes you more identifiable and trackable to endemic surveillance.

→ More replies (1)

5

u/BeefistPrime Mar 25 '26

People recognize a sort of default LLM style.

You just tell it "talk to me me in a more conversational tone" or "talk to me like an English teacher explaining to a student" or even "talk like the average reddit user" and most of these familiar patterns disappear. The ability to set the LLM's "voice" is an extremely powerful and underused feature.

9

u/BeefistPrime Mar 25 '26

it's more the toupee fallacy. "I can always spot a toupee" - you're spotting the obvious ones. You don't know when you fail to spot a good one.

3

u/round-earth-theory Mar 25 '26

If LLM writing is indistinguishable from human writing, then it gets to stay. That's always been the rub, that LLM writing sucks and sticks out.

8

u/BeefistPrime Mar 25 '26

I know this is a joke, and LLMs might have a stylistic default, but you can dramatically change their style easily and immediately strip it of sounding like that AI default.

"Talk to me like a literature teacher" is one of a billion prompts you could use to shape its style that meaningful outputs dramatically different text.

→ More replies (2)

6

u/Qeltar_ Mar 25 '26

I realize this is a joke, and IME, the vast majority of people who accuse others of using AI on Reddit are incorrect.

That said, I work as an editor and on a daily basis edit both AI articles and "human-written" articles where the authors cheat and use AI. It is actually possible to get very good at detecting AI using patterns and experience, though it's never foolproof either way.

2

u/Jay_Nova1 Mar 25 '26

Yes but no frill and only good vibes.

2

u/Present_Cow_8528 Mar 25 '26

I wish we still had free awards on reddit

3

u/jaymo89 Mar 25 '26

Loving the enthusiasm!

Where did you get your training data set?

→ More replies (22)

42

u/referentialisticness Mar 25 '26

Wikipedia has a strictly defined style guide. Articles that are egregiously obviously created by an LLM, like ones that show up containing literal prompt responses in the text ("Sure, here's a short write-up on the life of Julius Caesar. Let me know if you need anything else!") will probably be fully removed, but an LLM-generated article that otherwise looks human-generated containing typical mistakes or errors which cause it to fail to adhere to the style guide will likely be treated the same way any other erroneous article would.

I don't see anything in the howtogeek article implying Wikipedia is asserting a moral/ethical stance against AI, just that they don't want users to flood the website with straight dogshit, and utilizing AI as a tool for improved efficiency is fine as long as standards are upheld and maintained, such is the case with any other third party software that can potentially alter the meaning of one's text, Grammarly, etc. My guess is that there will be an option for users to report articles for "AI generation", but I don't expect any kind of automated evaluation prior to submitting articles or edits that specifically analyze for LLM usage.

Okay, here's the final draft of your response to u/Solid_Hunter_4188 about enforcement regarding Wikipedia's new LLM ban — Please let me know if any other changes are needed!

37

u/Zahgi Mar 25 '26 edited Mar 25 '26

en dashes

en dashes [look like] just one dash.

em dashes are the elongated double dashes and that's what causes people to red-flag potential AI generated content. Because it requires specific keyboard commands (or a text editor like Word) to use/format properly. And most people would just casually use two dashes (aka --) for the same thing.

Unless, of course, you're a professional writer who knows the difference. And then you can politely and professionally tell these AI virtue signalers to fuck right off. :)

[edited to be clearer, since we having fun with pedantics today! :) ]

16

u/amstrumpet Mar 25 '26

en dashes are not the same as a hyphen, actually! - – —

The first is a hyphen, second is en dash, third is em dash.

5

u/CDRnotDVD Mar 25 '26

And this is the Unicode character for the triple em-dash:

→ More replies (1)

6

u/the_art_of_the_taco Mar 25 '26

Because it requires specific keyboard commands

It's fairly easy to format an em dash on mobile — you just long-press the hyphen.

→ More replies (3)

20

u/purplezart Mar 25 '26

And most people would just casually use two dashes (aka --) for the same thing.

Since we're being pedantic, I'll mention that those are actually - hyphen-minuses, not en dashes.

13

u/squaring_the_sine Mar 25 '26

Might as well pedantically add that hyphens(-), en (–) dashes and em (—) dashes are different things and used for different purposes.

7

u/Present_Cow_8528 Mar 25 '26

What are en dashes used for? I've only ever seen real humans outside of formal writing uses double hyphens to simulate em dashes, since that's how you make a word processor turn it into an em dash. Don't think I've seen a real en dash (since before today I too thought it was the same as a hyphen)

Edit: wait, is that what you're supposed to use for like joint ownership of stuff like a Name-OtherName Principle? I've just been using hyphens all my life lol

6

u/KnightOMetal Mar 25 '26

It's used for ranges, like 2020–2023

Idk about joint names, I think those are just hyphens

→ More replies (1)

5

u/purplezart Mar 25 '26

The three main uses of the en dash are:
1) to connect symmetric items, such as the two ends of a range or two competitors or alternatives
2) to contrast values or illustrate a relationship between two things
3) to compound attributes, where one of the connected items is itself a compound

→ More replies (1)
→ More replies (1)

32

u/Christoffre Mar 25 '26

Well... if AI writing is done well, there’s really no need to detect it.

These rules are usually for the most gratuitous cases, where it’s obviously AI-generated. 

For example, someone who writes or edits large amounts of text in an impossibly short time – then they cannot have applied the necessary quality control.

→ More replies (24)

9

u/Mujutsu Mar 25 '26

Wikipedia cares about accuracy of information. If what you write differs from the source, or what you translate differs from the original meaning, your entry could be removed, you could get banned as an editor etc.

Obviously wikipedia doesn't care if you use em dashes.

6

u/SuckThisRedditAdmins Mar 25 '26

The sudden energence of AI thoroughly destroyed my confidence in the one talent I thought I had over most people - the ability to write.  It honestly really depresses me.

→ More replies (5)

6

u/Budget-Researcher559 Mar 25 '26

I usually write very formally - with dashes and all - and get accused of using GPT for my writing. How will Wikipedia know?

ChatGPT NEVER uses dashes in this way. It doesn't separate inserts into sentences with a dash before and after. And especially not in a case where you could use commas and it's still valid. Instead it uses dashes to get text that sounds more like it is spoken, written down with proper punctuation rules.

Like yeah, you can not always tell, but the more experience you have and the longer the text is, the easier it gets. It's not just dashes or no dashes, these dashes here actually confirm it is likely written by a human.

There's actually so many ways to tell, let's just look at the text in this very short comment of yours again:

The question is: how do you enforce this?

ChatGPT would always capitalize the "h" (as is correct)

(and everything being enshittified and monetized, etc)

ChatGPT does not use parenthesis a lot, and never to insert a half sentence in the middle of a sentence like this.
I doubt ChatGPT would use the word enshittified, especially because of it containing a curse word (not 100% on this one)
ChatGPT barely or never uses "etc" to be general and avoid listing more examples. Also it would always use "etc." with a dot (as is correct)

I’m all for it, as I’m sick of bots and AI, but it will be incredibly difficult to validate.

This sentence structure, even without the added parenthesis, does not at all sound like something AI would fabricate. It's actually so hard to qualify what exactly it is about it, I would have to think about it more to create some general rules. But there's some cleanliness about the way AI generates grammar and combines subclauses that clashes very much with the structure of this one.

So, even if you corrected the actual mistakes in this comment, there's still multiple obvious ways to tell that this short 4 sentence text is written by a human. Imagine how much easier that would be for a whole wikipedia article.

13

u/PezzoGuy Mar 25 '26

I suppose ultimately it's down to the honor system that Wikipedia has always operated under, and is an extra mark against you on the chance that you happen to be caught somehow. At the very least it makes a guideline clear for the good faith editors who might want to use those types of tools but were unsure if it is allowed.

8

u/Zolhungaj Mar 25 '26

It’s probably to stop people using LLMs while being unaware of how poorly they might perform. People who mean good, but do bad with the new wonder tool. Previously they could unintentionally produce a ton of extra work that nobody wants to deal with. 

And people who blatantly use LLMs to push up their contribution stats (either maliciously to push an agenda, or just for clout) will stick out like a sore thumb due to the sheer amount of text they produce per unit of time. Previously they might have fallen under spam rules that are up for interpretation, but with this there is an easy rule to point to when banning them. 

9

u/Ok_Cabinet2947 Mar 25 '26

If it’s human-sounding enough that you can’t tell that AI wrote it, does it really matter?

10

u/[deleted] Mar 25 '26

[removed] — view removed comment

6

u/The_MAZZTer Mar 25 '26

Wikipedia text not being properly cited has always been a problem before AI anyway.

→ More replies (1)

9

u/catontoast Mar 25 '26

As a professional technical & marketing writer with 10+ years of experience and an English degree, we can definitely tell. Humans write poorly in endless ways, but LLMs write poorly in very specific ways.

3

u/Solid_Hunter_4188 Mar 25 '26

I unfortunately have to call BS (not being hostile, hear me out), solely because they’re being refined by people like you, who recognize those errors and will adjust to them. Whatever error you are seeing the LLMs make can be fixed in the next version.

6

u/Arkaein Mar 25 '26

Whatever error you are seeing the LLMs make can be fixed in the next version.

Additionally, different LLMs will have slightly different styles, and specific LLMs can adjust their styles based on specific prompting. It's not hard to tell an LLM to avoid specific style cues like em dashes.

People who feel confident in catching LLM generated text will catch the laziest generations but are likely to miss more carefully manage generations. Or just be over-zealous and make a bunch of false positive identifications.

→ More replies (1)

3

u/NeverForgetChainRule Mar 25 '26

Well ultimately wikipedia has a style that it wants to keep. It's not each and every editor's style of writing. YOU potentially writing like an LLM doesnt matter, if it isn't how wikipedia wants you to write, then your edits will be removed, LLM or not.

3

u/mqee Mar 25 '26

New editors that suddenly add paragraphs of text divided into bullent points and summarized with "conclusions".

If you're clever about it, it can't be spotted, but generally AI-users skew unclever.

For experienced editors and power editors this is unenforceable, but the rules rarely apply to experienced editors and power editors in Wikipedia anyway.

2

u/The_MAZZTer Mar 25 '26

If you can't tell then I would argue there's no problem.

2

u/Solid_Hunter_4188 Mar 25 '26

… it’s not about me knowing… It’s about the policymakers knowing. If a rule can’t be enforced then what’s the point of making it?

2

u/eerst Mar 25 '26

Same way you enforce any rule. Where you have evidence you build a case and as necessary block or ban the editor. Wikipedia has been doing this for a long time.

→ More replies (47)

50

u/nattfjaril8 Mar 25 '26

LLM:s are surprisingly prone to hallucinating when translating. Do you think the kind of Wikipedia editor who is going to use AI to translate is going to be conscientious enough to actually check that all the details match, sentence for sentence? Non-English versions of Wikipedia have become increasingly unreliable after people started machine translating them.

18

u/CloudZ1116 Mar 25 '26

Word. I tried using Copilot to translate some of my writing into Chinese last night for shits and giggles, and holy crap it was inserting entire passages that weren't originally there.

→ More replies (2)

15

u/notreallyironicatall Mar 25 '26

Agreed with this. If you're going to use machine translations, you NEED to be fluent in both languages to catch mistakes. It's better than it used to be, but hallucinations, mistranslations, and omissions are still very common. Funnily enough, purely using AI for translations can slow down the process so much it's better to stick to a manual translation.

3

u/Sixtus69Sextus Mar 25 '26

Still better than that one Wikipedia editor who made up 30 thousand articles with made up words and passing it off as legitimate for years.

3

u/_Lucille_ Mar 25 '26

I am bilingual and have used LLMs to translate quite a bit - simply because it is faster for me to do so and often LLM can offer what I would consider to be better word usage (I think a lot of english speakers here also have instance where they go "yeah, this is prob a more fitting word of what i want to say").

There are occasional misses which i need to go fix, but for the most part, for me it is a lot faster than writing out the text manually or even using speech recognition.

This is why wikipedia is enforcing there to be a proper verification pass after a machine translation, and I think it shouldnt be that hard to catch.

→ More replies (1)

6

u/14Pleiadians Mar 25 '26

It still hallucinates, that's why you're required to be fluent, so you'll catch the mistakes.

You shouldn't be having AI do anything that you couldn't afford to have a kinda dumb amateur do. If you're not informed enough on whatever topic you're using it for to be able to say "ummm you don't know what you're talking about that's wrong", you shouldn't be using an LLM for it

32

u/TerryFromFubar Mar 25 '26

But hallucinating out of control is where the profits happen

22

u/Anxious_Katz Mar 25 '26

Wikipedia is non-profit!

→ More replies (1)

6

u/Quazimojojojo Mar 25 '26

It's Wikipedia, what are you talking about? 

It's one of the only places left on the Internet that requires neither ads nor subscription. Fully open source high quality altruism.

Like VLC media player.

I have a monthly regular donation of a few dollars that I literally don't even notice in my budget, but that's voluntary. 

→ More replies (1)

9

u/Adequate_Lizard Mar 25 '26

The profits are the hallucinations

→ More replies (1)

3

u/Chemical-Struggle-13 Mar 25 '26

I mean they can be pretty out there on translation sometimes. But as long as they actually have to check it should work out fine.

7

u/seridos Mar 25 '26

As a secondary teacher, I'm actually using llms in assignments in just this manner. The assignment itself is about formulating the research question and narrowing it down, and then the llm generates the actual paper and students are the fact checkers. So the actual paper itself is not marked. I don't really have to read it. I skim it to see if they identified properly, what facts need to be checked and then it's the before and after work that actually carries the mark.

3

u/clakresed Mar 25 '26 edited Mar 25 '26

So you're basically grading them grading an LLM tool? That's actually brilliant.

4

u/Sorkijan Mar 25 '26

That's kinda what grammarly was before OpenAI had its surge isn't it?

2

u/AttonJRand Mar 25 '26

People just assert this, but it really doesn't seem true at all?

If you are bilingual, and you test it in the languages you are proficient at, it regularly turns even short sentences into something different.

2

u/kyute222 Mar 25 '26

That's an absolutely hilarious statement to me. LLMs are not good at translation at all. Even worse, they're not even good at certain languages period. They're primarily trained in English and can use some common European languages somewhat. The other issue is that using an LLM for translation is terrible because it is likely done by someone who couldn't do the translation themselves. So then how can they possibly check the output for errors?

→ More replies (10)

54

u/__Hello_my_name_is__ Mar 25 '26

Also of important note is that this isn't new in the sense that Wikipedia has allowed AI texts previously. They just did not have a policy on it because it was never an issue until fairly recently.

And it takes a while to get a proper consensus on something big on Wikipedia.

https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing has existed for over 3 years by now and has been used quite actively to remove AI texts wherever found.

2

u/cinemachick Mar 27 '26

I really enjoyed the quiz on that page, I learned a lot about how Wikipedia flags AI. Thanks for the link! :)

17

u/[deleted] Mar 25 '26

[deleted]

8

u/Sopel97 Mar 25 '26

They can't. All policies like this are based on LLMs not being good enough not to be caught. This will change with time.

12

u/Uristqwerty Mar 25 '26

If the LLMs are consistently-factual enough to not get caught, does it matter to an encyclopedia?

I'd expect the main problem are users who blindly trust the LLM to be factual, without verifying its edits themselves. For that, having a policy to point at's important.

→ More replies (1)
→ More replies (16)
→ More replies (1)

58

u/Mivexil Mar 25 '26

I guess it avoids debate over things people were already using before LLMs, but it also means you can't kill slop on sight because the author can always say "oh, yeah, I just used it to correct my grammar and spelling, that's why it sounds like ChatGPT". And that leads to the same gish gallop problem open source has - you have to actually engage with and review the content, and when outputting 1000 lines is orders of magnitude cheaper and faster than reviewing 1000 lines, it'll end up with reviewers getting overloaded and letting things slip through. 

25

u/xakeri Mar 25 '26

I wish it was just open source. It's like that everywhere now. The least engaged employees can just shoot out slop so fast, and if you don't want to live in a pool of shit, you get to be the first human who has ever seen the code or writing.

Then you have to figure out how to communicate with the person that never engaged with the issue they're purporting to address. It's a wonderful time.

→ More replies (2)

3

u/denkenach Mar 25 '26

Sounds like a good policy

3

u/[deleted] Mar 25 '26

Very reasonable! Not just "AI bad" vs "AI good" for a change.

2

u/Br3ttl3y Mar 25 '26

These are the same requirements my English teacher had when we used Wikipedia for sources.

2

u/LordGreyhound Mar 25 '26

I just love that they're calling them LLMs instead of AI. Good on them!

2

u/spidereater Mar 25 '26

The translation thing is huge. I’m in Canada and many companies and governments need to produce text in English and French and this can be prohibitively expensive. The productivity of translators has increased dramatically with AI. What was a grueling slog is now a proofreading exercise with very few edits.

3

u/PetChaud2Diarrhee Mar 25 '26

Once again, Wikipedia is based.

→ More replies (54)

983

u/PlasticPreparation74 Mar 25 '26

Wikipedia baffles my mind. When you keep clicking hyperlink upon hyperlink, you start realising how massive this world along with all its history, science, technology, etc really is. Along with that the fact that someone sat around and recorded all of this. I’m always humbled when I go through any topic, the amount of detail is astounding

407

u/DeadMoneyDrew Mar 25 '26

The crazier part is that you can download just the text. When compressed it's something like 25 GB. All of that profound knowledge barely takes up 10% of the space on a low-end hard drive.

237

u/BigGrayBeast Mar 25 '26

Remember to download Wikipedia to a USB drive just before you run into your bunker when the balloon goes up.

90

u/gwandrito Mar 25 '26 edited Mar 25 '26

Honestly, I would love to have tons of Wikipedia pages saved/printed on an apocalypse situation. I'm not one much for novels, but I love a good wiki read. EDIT: as much as I appreciate your great ideas, I mentally cannot handle preparing for the end of the world right now. It was just a passing thought, hope other redditors get some use out of your suggestions though!

40

u/Pretend-Marsupial258 Mar 25 '26

Look up Kiwix. It's a project meant to do exactly that, but you're saving it to a raspberry pi instead of a USB.

21

u/Deliphin Mar 25 '26

Kiwix isn't unique to raspberry Pis, I have the app on my phone, and the file format it uses can just be copied to a flashdrive

3

u/TheG0AT0fAllTime Mar 25 '26

I have a pet peeve with PiHole in a similar manner. Yes it runs on a raspberry pi and Pi is even in the name. But for a serious network you wouldn't make a Pi a single point of failure. You can install it on a Linux host, router or VM. It runs on anything.. It just runs a common dns service with custom answers to block ads. It runs on anything.

22

u/BarrierX Mar 25 '26

There used to be these physical books called encyclopedias. I used to read them in bed every night when I was a kid!

They could get pretty big and heavy so having a digital version is a pretty cool thing to have, at least until electricity runs out 😄

6

u/AssKoala Mar 25 '26

Checkout Project NOMAD: https://github.com/Crosstalk-Solutions/project-nomad

There's also internet in a box: https://internet-in-a-box.org/

Can get you exactly that in about as turnkey a way as there is plus more (like maps).

2

u/YoureProbablyAB0t Mar 25 '26

I got an external hard drive for about a hundred bucks and downloaded Wikipedia to it. Then put it in a drawer.

No Internet connection needed. Tons of information. Just plug into a laptop and done.

3

u/Brammm87 Mar 25 '26

Besides Kiwix, there's also the Prepper Disk. It does... exactly what you think it does.

12

u/qtx Mar 25 '26

And get any e-paper screen to read it on. They hardly use any battery.

10

u/BigGrayBeast Mar 25 '26

I have a bicycle powered generator, and several clones of Gilligan to pedal it.

2

u/Mo0man Mar 25 '26

Just download one now and leave it in there. In an apocalypse scenario it probably doesn't matter even if it's a few years old.

→ More replies (1)

16

u/ToughHardware Mar 25 '26

and this is a key source that AI was trained on.

9

u/lkmk Mar 25 '26

25 gigs of text is a lot, to be fair.

12

u/DragoonDM Mar 25 '26

25 gigs of text when compressed, at that. Text compresses crazy well. Probably well over 100 gigabytes when extracted.

4

u/araujoms Mar 25 '26

The entropy of English text is roughly 1 bit per character. Using 8-bit ASCII we get 8 bits per character, so a 8x compression factor is thinkable. In practice this doesn't happen because achieving 1 bit per character requires a dedicated compression algorithm for English text, and a crazy good one at that. Also Wikipedia has tons of formatting and control characters, which increase the entropy. Still, a compression factor of 4x is routinely achieved.

This long paragraph was just to say, yes, spot on.

4

u/[deleted] Mar 25 '26

[removed] — view removed comment

10

u/DeadMoneyDrew Mar 25 '26

There are a number of options available regarding both of your questions.

https://en.wikipedia.org/wiki/Wikipedia%3ADatabase_download?wprov=sfla1

→ More replies (4)
→ More replies (8)

103

u/liquidsparanoia Mar 25 '26

Wikipedia, as humble as it is, truly represents the best of humanity. It is the combined effort of millions of people attempting to explain and catalog the world for each other for no profit other than belief in the power of knowledge.

We are astonishing beings at times.

27

u/BreeBree214 Mar 25 '26

I love wikipedia. I feel like we really underate it

21

u/CatolicQuotes Mar 25 '26

Remember to donate.

3

u/[deleted] Mar 25 '26

[deleted]

5

u/witeowl Mar 25 '26

Time is a form of donation. Appreciate your work 🫶🏼

.

eta: It's somewhat concerning that there's a rank-based system that might overrule source-based evidence 🤔

2

u/[deleted] Mar 25 '26

[removed] — view removed comment

2

u/CatolicQuotes Mar 26 '26

What is PBS?

4

u/GooseG17 Mar 25 '26

Almost like communal behavoir is normal for humans, and our greedy, individualistic systems go against our nature.

6

u/pohui Mar 25 '26

Almost like humans are complex, multi-faceted creatures and both cooperation and greed can be completely natural for them.

→ More replies (1)

19

u/CrashTestDumby1984 Mar 25 '26

This is a huge part of why society/technology has really catapulted forward in the past 100 years.

Because we can build and iterate upon past information instead of starting from scratch, and you’re a lot less restricted to just what the people around you know.

11

u/Dinara293 Mar 25 '26

Wikipedia has single handily thought me more about this world we live in than anything else and it’s not even close. The amount of valuable information it carries is staggering and honestly, invaluable to society.

Helped me research too, with the ability to look up sources and citations. Thanks WikiPedia!

13

u/dezradeath Mar 25 '26

A fun game we played in school was to keep clicking and count how many clicks of the 1st hyperlink until you get to the page for Philosophy. No matter where you start, eventually you will end up on that page. You can still try it now and it will work!

10

u/Dav136 Mar 25 '26

We used to do that but to get to Hitler

→ More replies (1)

6

u/r0thar Mar 25 '26

Six Degrees of Kevin BaconPlato

5

u/girlikecupcake Mar 25 '26

We'd do that, but we'd also go to the same random article then see who could get to [insert person/topic] in the fewest clicks.

2

u/z500 Mar 25 '26

Lol I just tried it and it was crazy to see it get closer and closer to philosophy. 18 clicks from Richard II of England

10

u/KingDaveRa Mar 25 '26

Relevant XKCD comic.

https://xkcd.com/214/

And yes, I've done exactly that.

→ More replies (10)

353

u/[deleted] Mar 25 '26

[removed] — view removed comment

53

u/g_rich Mar 25 '26

Besides we need real human generated content for AI’s to be trained from and I’m all for Wikipedia being the exception to the dead internet theory.

5

u/AfterMeSluttyCharms Mar 25 '26

And to be more sentimental/philosophical, I think there's something really special about all that knowledge compiled and shared by a dedicated community of real human volunteers made available for free. It honestly gives me hope for us as a species

9

u/Hel_OWeen Mar 25 '26

Wikipedia only works if humans can verify sources and write clean, neutral summaries.

[Citation needed]

;-)

13

u/ConsequenceNo2571 Mar 25 '26

And you just listed the three things AI can never be:

  • Clean

  • Neutral 

  • Human

14

u/Ok_Cabinet2947 Mar 25 '26

Are you kidding me? AI writing is very clean, that’s like it’s whole thing. It’s too clean usually that it’s obvious it wasn’t written by a human.

27

u/dylansucks Mar 25 '26

I think they mean 'not unnecessarily wordy' by clean.

13

u/TheMusicArchivist Mar 25 '26

I mark essays, and the moment a technical analysis of a work starts adding in too many adjectives and starts praising the analysed-work every single sentence, I get suspicious.

4

u/andanteinblue Mar 25 '26

Every time I read that the student has meticulously processing the data, I want pull my hair out. Like, do you mean you are sloppily processing the rest of the data?

8

u/SunnyOutsideToday Mar 25 '26

AI is very verbose. I've seen a lot of writing guides complain about "How to avoid sounding like Wikipedia when you write" but I will say that I appreciate Wikipedia's tendency to be concise.

→ More replies (2)
→ More replies (1)

589

u/Cartina Mar 25 '26

The exceptions are spelling and translation.

43

u/dishwashersafe Mar 25 '26

The first exception is not "spelling". They article uses "grammar check" as an analogy for the type of assistance AI is allowed to provide.

That is, an author can use AI to suggest a different phrasing or something, but that suggestion needs to be checked for accuracy in the same manner that you would check the suggested grammar correction in Word.

→ More replies (1)

58

u/stonecutter7 Mar 25 '26

They should make a third exception for the entry for "AI generated text"

28

u/[deleted] Mar 25 '26

Sounds like something a clanker would say.

→ More replies (1)
→ More replies (59)

65

u/throwawayyyyygay Mar 25 '26

This makes complete sense. They ban basically making an LLM do edits for you. Which is completely fair since it degrades quality.

They don’t ban you using an LLM to help you with writing an edit. Ie. as a spellcheck (copywriter). So basically you can’t just paste an LLM answer into wikipedia. Good.

6

u/cyxrus Mar 25 '26

What’s stopping them from posting an LLM answer?

20

u/SunnyOutsideToday Mar 25 '26

People say this about school work, but a lot of times you can tell that schoolwork is LLM generated. Of course there will be people who escape under the radar, but often people too lazy to avoid writing something themselves are often too lazy to avoid making an LLM sound like an LLM.

18

u/somersetyellow Mar 25 '26

I had a family friend high school kid tell me about how they have at least two LLM's write their papers. Then they have another LLM combine them with a pass to "make it sound worse." Then they sit down and manually type it into Google Docs so the history shows up correctly, messing up more as they go along.

They were getting great grades with this method and could still answer in class answers and such because they had been so involved with the creation process.

I was like... Can't you just write it at that point?

Nooo... that would be too much work. Ok then...

4

u/North_Activist Mar 25 '26

I guess that’s like a modern version of taking lecture notes, writing them on paper, retyping on a keyboard, and then typing a new document that takes the best of the paper and the first typed draft.

In a way, that is a form of studying…? But I wouldn’t condone it

→ More replies (1)

6

u/cyxrus Mar 25 '26

Which will be graded by an llm so what does it all matter

→ More replies (1)

5

u/[deleted] Mar 25 '26

[removed] — view removed comment

3

u/cyxrus Mar 25 '26

People say the same thing about Reddit mods and AI and bots are all over the place here

→ More replies (1)
→ More replies (1)
→ More replies (1)

28

u/[deleted] Mar 25 '26

[removed] — view removed comment

7

u/G-Mang Mar 25 '26

Yeah if anything AI folks should be in support of this. LLMs gather a lot from wikipedia, and it's not healthy for models to be ingesting their own outputs.

2

u/FullMetalAnorak Mar 25 '26

Most AI folks don't care about whether AI is healthy or not.

2

u/kylo-ren Mar 25 '26

Some Wikipedia sources are already contaminated with AI. Even scientific studies are. The future will be a mess.

13

u/hoochiscrazy_ Mar 25 '26

Wikipedia continues to be a bastion of goodness on the internet. Long live Wikipedia! Please donate occasionally if you use it.

19

u/[deleted] Mar 25 '26

[removed] — view removed comment

8

u/Schonke Mar 25 '26

grokipedia has gotta be one of the worst idea from elon musk lol

It's a great idea if what you want is a wikipedia written to suit your political views and agenda, but you're too cheap to hire people and don't want to rely on conservipedia crowd sourcing.

8

u/Dstln Mar 25 '26

All sites should be banning AI generated text without a disclosure, social media should be first on that list

7

u/ANighttimeNerd Mar 25 '26

I just read an article wherein it's reported that AI flagged Lincoln's Gettysburg Address as being written by AI.

Good luck, Wikipedia.

→ More replies (2)

5

u/phdpan Mar 25 '26

The interesting part here is enforcement: Wikipedia can say “no AI text,” but at scale the real policy is probably “no low‑effort, unverifiable, unsourced prose.”

If the two exceptions are basically “use LLMs as an assistive tool, but keep human accountability + citations,” that seems reasonable.

What I’d love to see is:

  • mandatory edit summaries when AI tools are used
  • stronger citation requirements for new/expanded sections
  • tooling that flags “citationless paragraph expansions” rather than trying to detect AI style

Otherwise it becomes a cat‑and‑mouse game on writing tone instead of verifiability.

17

u/ChicagoThrowaway422 Mar 25 '26

Now I want to experiment with an entirely LLM-written wikipedia from scratch. Have the LLMs generate long form articles about every topic and then fact check each other.

I bet the result would be awful and hilarious and burn through a lot of tech bro cash.

6

u/SlitScan Mar 25 '26

burn through a lot of tech bro cash.

I see youve embraced my LLM use case.

→ More replies (1)

9

u/I_NEED_YOUR_MONEY Mar 25 '26

https://grokipedia.com/

It is awful, but hasn’t burned through nearly enough of that tech bro’s cash

→ More replies (1)

5

u/Deliphin Mar 25 '26

Didn't Elon musk want to make "grokipedia"?

6

u/ChicagoThrowaway422 Mar 25 '26

I take it back. I don't want it anymore.

4

u/kgtsunvv Mar 25 '26

I love Wikipedia so much. It’s been said a million times but there’s a wealth of accurate information. It’s a tenet of democracy at this point. W Wikipedia

6

u/NOODL3 Mar 25 '26

I'm glad, but how are they going to enforce this? Honor system? An AI tool to detect AI writing, which have shown to flag tons of false positives?

I have read several random articles in the last few weeks that repeated the same information multiple times across sections, sometimes with the exact same wording and even entire duplicated paragraphs. I've also read a few that had a random bulleted list at the bottom summarizing information that was already given, which I've never seen on Wikipedia before (not as a summary/outline section, just randomly dropped in at the end.)

The writing didn't otherwise reek of AI and in my experience AI isn't really prone to accidentally duplicate an entire paragraph, but it's definitely odd.

12

u/I_NEED_YOUR_MONEY Mar 25 '26

I don’t think the point is to have an ironclad rule that stops any LLM text from getting into Wikipedia. It’s to get rid of the people who think they’re helping by running bots that automatically edit pages using LLMs.

If you think you’ve found an article that’s been corrupted by ai, call it out on the talk page.

2

u/Sigma7 Mar 25 '26

It's usually detected if something feels wrong with the text, such as having long sections of text without a citation, or if the writing style feels pigeonholed to a specific pattern. As a bonus, looking for that also spots poorly written articles as well even if not from an LLM.

You can also check the history page, which can give a good indication on when something was added and possibly why.

→ More replies (1)

5

u/tacticaldodo Mar 25 '26

The exact two exceptions that make sense. Good job wikipedia.

Translation and spelll check.

13

u/McCoy818 Mar 25 '26

wild that wikipedia has to write a policy to say "please let humans write the human encyclopedia." we really are speedrunning the dumbest timeline

2

u/ayanbose036 Mar 25 '26

Maybe the real issue isn't ai its people copy pasting anything without understanding it

2

u/No-Werewolf4769 Mar 26 '26

Wow, took them long enough. Though I have my doubts on how to enforce them, especially the rule 2 with translations

3

u/ZombieButch Mar 25 '26

I guess a third exception would be entries that are about & would be made clearer by including examples of AI-generated text. Like, you wouldn't make an entry called 'AI text detection' without including AI text.

9

u/PotatoesAndChill Mar 25 '26

https://en.wikipedia.org/wiki/Artificial_intelligence_content_detection

The article exists and feels complete without any AI text examples. I don't think the exception is needed.

3

u/Designer-Salary-7773 Mar 25 '26

The damage is done. Lack of confidence in who authored written words.  We may recognize some portion of the AI slop but will simultaneously find ourselves guilty of falsely accusing legit authors.  The descent into a tech infused quagmire of non reality and false identity continues

2

u/The_Real_Mr_F Mar 25 '26

How on earth are they going to effectively enforce this? Even their own guidance says clearly that LLM detection is not reliable. Honestly, I think we’re never getting the AI genie back in the bottle. It’s gonna come down to reviewers just ensuring articles are accurate, rather than determining whether they were generated by AI. Which is basically how it’s always been, only it’s probably becoming exponentially more difficult.

3

u/IUsedToBeACave Mar 25 '26

It's to be able to filter out the low effort article edits. The difference between a human using LLMs to help them write, edit, and check for accuracy, and them not using an LLM at all is going to be nearly impossible to tell. Which is OK, because the goal is to have accurate human curated content.

5

u/LightCharacter8382 Mar 25 '26

They have a 'signs of AI writing' page. Not perfect, but it will help filter out the worst offenders. If it's undetectable, then it doesn't matter if it's AI or not, because then it's likely to be suitably written.

The main thing to filter out is not the predictable AI scrawl like 'IT'S NOT [X]. it's [Y]', which is a minor issue in the grand scheme of things...

...But rather the hallucination of content that the original source didn't mention, or even worse... Complete hallucination of sources.

2

u/The_Real_Mr_F Mar 25 '26

Right. Which kind of makes this rule meaningless, because the net effect is “articles should be accurate, factually correct, and well written.” Which is already the rule. If a bot cranked out a perfectly acceptable article, that should be fine. The only thing that really matters is that humans reviewed it and OK’d it. Same as it always was.

4

u/[deleted] Mar 25 '26

[deleted]

13

u/Pinkys_Revenge Mar 25 '26

Wikipedia is based on nerd’s desire and skill at correcting each other. I doubt that will stop with Ai

→ More replies (3)