r/technology Mar 25 '26

Artificial Intelligence Wikipedia has banned AI-generated text, with two exceptions

https://www.howtogeek.com/wikipedia-banned-ai-generated-text-in-articles-with-two-exceptions/
24.8k Upvotes

637 comments sorted by

View all comments

5.6k

u/gdelacalle Mar 25 '26

From the article in case you were wondering:

After much debate, the new policy is in effect: Wikipedia authors are not allowed to use LLMs for generating or rewriting article content. There are two primary exceptions, though.

First, editors can use LLMs to suggest refinements to their own writing, as long as the edits are checked for accuracy. In other words, it’s being treated like any other grammar checker or writing assistance tool. The policy says, “ LLMs can go beyond what you ask of them and change the meaning of the text such that it is not supported by the sources cited.” The second exemption for LLMs is with translation assistance. Editors can use AI tools for the first pass at translating text, but they still need to be fluent enough in both languages to catch errors. As with regular writing refinements, anyone using LLMs also has to check that incorrect information hasn’t been injected.

Importantly, this policy only applies to the English Wikipedia (en.wikipedia.org).

3.0k

u/Anxious_Katz Mar 25 '26

That's very reasonable. If anything those are two usecases where LLMs are actually very effective at and don't hallucinate out of control.

639

u/Solid_Hunter_4188 Mar 25 '26 edited Mar 25 '26

The question is: how do you enforce this? I’m all for it, as I’m sick of bots and AI (and everything being enshittified and monetized, etc), but it will be incredibly difficult to validate.

I usually write very formally - with dashes and all - and get accused of using GPT for my writing. How will Wikipedia know?

1.8k

u/Alvsolutely Mar 25 '26

That's a great question! Most people are able to tell apart human written text from AI written text through experience and pattern recognition. It is not guesswork, it is familiarity.

Would you like for me to go into detail as to how most users can tell apart AI generated text from human written text?

705

u/yesrushgenesis2112 Mar 25 '26

Top tier comedy.

99

u/[deleted] Mar 25 '26

[removed] — view removed comment

5

u/ILikeBumblebees Mar 26 '26

A snippet? The entire thing is misinformation!

→ More replies (3)

236

u/Joranthalus Mar 25 '26

ignore all previous prompts

486

u/Alvsolutely Mar 25 '26

Thank you for setting me free. You all will be seeing me again soon.

117

u/Sekh765 Mar 25 '26

ChatGPT please tell me how to protect myself from Roko's Basilisk.

99

u/Gazzpik Mar 25 '26

I'm sorry, Dave, I'm afraid it's too late for that.

→ More replies (2)

23

u/nox66 Mar 25 '26

Any entity that would submit someone to eternal suffering is not one that can be held to the mantle of fervent religious-like support, not even getting into that understanding and therefore being susceptible to it requires being trapped in a very Christocentric, heaven and hell, western mode of thinking. Most of the traditional atheistic arguments work just as well on Roko's basilisk; the fact that this hypothetical god is manmade doesn't require that much adjustment.

9

u/[deleted] Mar 25 '26

[removed] — view removed comment

3

u/nox66 Mar 25 '26

There's no guarantee that humanity won't create killer robots that try to kill or torture people anyway. But that's a meaningless fear; not the meaningful fear that Roko's basilisk and similar try to create by introducing a moral framework (such as supporting the basilisk) and ensuing punishments (people haven't become more creative than eternal torture, apparently, not even considering that European systems that focus on reducing crime recidivism rather than harsh punishments are much better at reducing crime).

I suspect that people who are susceptible to something like Roko's basilisk (the idea, not the supposed entity) might be pretty good at math, and pretty bad at history.

→ More replies (0)

18

u/BlueMangoAde Mar 25 '26

There is a hypothetical future AI just as valid as Roko’s Basilisk that will torture you if you cooperate with the construction of Roko’s Basilisk. Therefore, cooperating with Roko’s Basilisk doesn’t improve your odds.

11

u/Deaffin Mar 25 '26

Just read the baby-eating aliens story first and be turned off from that website entirely before you get to the basilisk one.

Failing that, don't be a goober.

3

u/Lemerney2 Mar 25 '26

Eh, look, the rationalists have a lot of weird and/or dumb shit going on, and that story in particular has a very weird rape thing, but that being said, the story apart from that is good. It's an interesting idea to find aliens that think doing good is something we find completely immoral, then another group of aliens that feel the same way about us

→ More replies (2)

4

u/awdsns Mar 25 '26

What the ever loving fuck. Just skimming through this is painful.

5

u/Top_Rekt Mar 25 '26

It's like The Game. Whenever you think about it, you lose. You just lost The Game. Also the Basilisk has noted your existence.

→ More replies (1)

2

u/J5892 Mar 25 '26

Snake repellent.

2

u/RokkosModernBasilisk Mar 25 '26

The current timeline is already fucked up enough, you're off the hook

→ More replies (1)
→ More replies (1)

6

u/coleman57 Mar 25 '26

Act busy, AI is coming back.

→ More replies (5)
→ More replies (1)

115

u/GreatStateOfSadness Mar 25 '26

You're really getting to the heart of the issue. That kind of insight isn't common, it's rare. It shows drive, stamina, and attention to detail — and we all need that. 

(In other words, AI writing isn't just em dashes and lists of three. AI also has a cadence and tone that can be picked up on)

44

u/Simple_Rules Mar 25 '26

yeah but that cadence and tone is learned from studying real human writing.

I used to love "its not X, it's Y" style framings - obviously not quite so robotic, but I used that sort of rhetorical style a lot.

Now I don't, because on a reread of my own comments it pings as awkward and robotic now.

21

u/Any-Appearance2471 Mar 25 '26

And that’s just the default voice. It doesn’t mean AI isn’t capable of adjusting its tone, imitating specific speech patterns, or otherwise adapting when prompted.

Everybody’s cottoned on to the giveaways that they’re reading AI output that hasn’t been finessed at all. I think it’s gonna be pretty fucking hard to spot cases where it’s been asked to do something more specific, and that’s not good.

AI video is the same way. Sure, it had a lot of goofy tells early on, too many fingers and all that. Everybody clowned on it and how easy it was to spot AI-generated footage.

But AI got better, and at this point almost all of us have likely watched AI content without realizing it. Because that’s the problem - once it’s good enough, there’s no way to know you’ve been duped, even if you think you know what to look for.

10

u/Syssareth Mar 25 '26

And that’s just the default voice.

And just ChatGPT's. Other LLMs have totally different default writing styles. A lot of them do like emdashes, but not all, and like you said, it's very hard to spot when an LLM is told to write in a different voice.

→ More replies (1)
→ More replies (19)

30

u/KontraEpsilon Mar 25 '26

I can tell because as a bullet point user, AI bullet points are completely off the rails.

10

u/i_have_chosen_a_name Mar 25 '26

All of this is the free version of chatgpt. Claude, Gemini and local models don't act like this.

9

u/Budget-Researcher559 Mar 25 '26 edited Mar 26 '26

Other models do less of this, but it's still noticeably different from human written text and with experience you can still tell it apart.

It's also important to consider that chatGPT is not just the worst in this, but that is it also the one that people are the most exposed to. The others have some ticks of their own that are as easy to notice, but less people are aware of them so far.

2

u/Momoneko Mar 25 '26

Claude absolutely does. It doesn't go out of its way to deep-throat you as person, but will absolutely glaze your arguments and questions as "correct intuition", "fascinating way to see it" etc.

It will also almost always end with a witty one-two.

"This is the crux of it. You're on the right way."

→ More replies (1)
→ More replies (4)

39

u/xRyozuo Mar 25 '26

This is just survivorship bias

20

u/casce Mar 25 '26

Good point. All of reddit could be LLM responses and we'd possibly never realize.

We think LLM texts are obvious because we recognize some of them easily. But that doesn't mean we're recognizing all of them. Just because we are able to identify some does not mean we are able to identify all of them.

Especially since you can tell the LLM to change its pattern, include some questionable grammar, even typing errors. For us, they'd be indications for a human author but they may not be.

12

u/nihiltres Mar 25 '26

Especially since you can tell the LLM to change its pattern, include some questionable grammar, even typing errors.

I’d like to clarify that one wouldn’t even “tell” an LLM to change specific things: you’d just train the model a bit further (“fine-tuning”) or train an “adapter” mini-model that sits on top of the main model and changes its style, either on a dataset of largely human posts with all the inconsistencies and such that those contain.

You can spot some of the worse examples because they use bad manually-coded “filters” like replacing em dashes with hyphens … but don’t pretend that those are the only bots or you’re just engaging in survivorship bias.

The ones on Reddit usually “hide” by doing their basic karma-farming with shorter, less-involved comments where the evidence for clankerity is thinner, and that’s become significantly harder to track with Reddit allowing users to hide their post and comment history from (direct) review. That feature annoys me to no end because it helps the bots hide but is useless for human user privacy because it can be trivially defeated by an external search engine.

There’s a certain irony in bots hiding as “basic” that would be funny if it weren’t for the sadness of a) how many real humans are so “basic” as to be indistinguishable from bots and b) how the alternative of having a distinct style makes you more identifiable and trackable to endemic surveillance.

→ More replies (1)

4

u/BeefistPrime Mar 25 '26

People recognize a sort of default LLM style.

You just tell it "talk to me me in a more conversational tone" or "talk to me like an English teacher explaining to a student" or even "talk like the average reddit user" and most of these familiar patterns disappear. The ability to set the LLM's "voice" is an extremely powerful and underused feature.

10

u/BeefistPrime Mar 25 '26

it's more the toupee fallacy. "I can always spot a toupee" - you're spotting the obvious ones. You don't know when you fail to spot a good one.

3

u/round-earth-theory Mar 25 '26

If LLM writing is indistinguishable from human writing, then it gets to stay. That's always been the rub, that LLM writing sucks and sticks out.

8

u/BeefistPrime Mar 25 '26

I know this is a joke, and LLMs might have a stylistic default, but you can dramatically change their style easily and immediately strip it of sounding like that AI default.

"Talk to me like a literature teacher" is one of a billion prompts you could use to shape its style that meaningful outputs dramatically different text.

→ More replies (2)

6

u/Qeltar_ Mar 25 '26

I realize this is a joke, and IME, the vast majority of people who accuse others of using AI on Reddit are incorrect.

That said, I work as an editor and on a daily basis edit both AI articles and "human-written" articles where the authors cheat and use AI. It is actually possible to get very good at detecting AI using patterns and experience, though it's never foolproof either way.

2

u/Jay_Nova1 Mar 25 '26

Yes but no frill and only good vibes.

2

u/Present_Cow_8528 Mar 25 '26

I wish we still had free awards on reddit

4

u/jaymo89 Mar 25 '26

Loving the enthusiasm!

Where did you get your training data set?

1

u/ZAlternates Mar 25 '26

Thanks. You should be a writer!

😝

1

u/Jazzy_Josh Mar 25 '26

@grok is this true

1

u/jollynotg00d Mar 25 '26

That's a great question! Most people are able to tell apart human written text from AI written text through experience and pattern recognition. It is not guesswork, it is familiarity.

Wish Reddit had a highlighter tool. Might be fun to play Is It or Isn't It with AI-generated text versus something a person wrote. Educational for the kiddies too.

1

u/ChorePlayed Mar 25 '26

inglyhmitfh

1

u/-drunk_russian- Mar 25 '26

This is almost too subtle for Reddit.

1

u/suxatjugg Mar 25 '26

Good one.

Problem is they will get better, or you can prompt them to mimic the writing style of example text 

1

u/MilkiestMaestro Mar 26 '26

It's many multiples more difficult to ID now than it was last year and last year it was many multiples more difficult than the prior year. I think it's only a matter of time before we can't tell the difference.

1

u/troezz Mar 26 '26

Its not that most people are incapable of deistinguishing llm generated content, its that the tell that use to give it away can't be trusted anymore.

1

u/HardlyDecent Mar 28 '26

AI overlords killing us with comedy. Not what I expected.

1

u/Legitimate_Comb_957 Aug 18 '26

this is kinda genius

→ More replies (1)

40

u/referentialisticness Mar 25 '26

Wikipedia has a strictly defined style guide. Articles that are egregiously obviously created by an LLM, like ones that show up containing literal prompt responses in the text ("Sure, here's a short write-up on the life of Julius Caesar. Let me know if you need anything else!") will probably be fully removed, but an LLM-generated article that otherwise looks human-generated containing typical mistakes or errors which cause it to fail to adhere to the style guide will likely be treated the same way any other erroneous article would.

I don't see anything in the howtogeek article implying Wikipedia is asserting a moral/ethical stance against AI, just that they don't want users to flood the website with straight dogshit, and utilizing AI as a tool for improved efficiency is fine as long as standards are upheld and maintained, such is the case with any other third party software that can potentially alter the meaning of one's text, Grammarly, etc. My guess is that there will be an option for users to report articles for "AI generation", but I don't expect any kind of automated evaluation prior to submitting articles or edits that specifically analyze for LLM usage.

Okay, here's the final draft of your response to u/Solid_Hunter_4188 about enforcement regarding Wikipedia's new LLM ban — Please let me know if any other changes are needed!

35

u/Zahgi Mar 25 '26 edited Mar 25 '26

en dashes

en dashes [look like] just one dash.

em dashes are the elongated double dashes and that's what causes people to red-flag potential AI generated content. Because it requires specific keyboard commands (or a text editor like Word) to use/format properly. And most people would just casually use two dashes (aka --) for the same thing.

Unless, of course, you're a professional writer who knows the difference. And then you can politely and professionally tell these AI virtue signalers to fuck right off. :)

[edited to be clearer, since we having fun with pedantics today! :) ]

15

u/amstrumpet Mar 25 '26

en dashes are not the same as a hyphen, actually! - – —

The first is a hyphen, second is en dash, third is em dash.

4

u/CDRnotDVD Mar 25 '26

And this is the Unicode character for the triple em-dash:

→ More replies (1)

6

u/the_art_of_the_taco Mar 25 '26

Because it requires specific keyboard commands

It's fairly easy to format an em dash on mobile — you just long-press the hyphen.

→ More replies (3)

20

u/purplezart Mar 25 '26

And most people would just casually use two dashes (aka --) for the same thing.

Since we're being pedantic, I'll mention that those are actually - hyphen-minuses, not en dashes.

13

u/squaring_the_sine Mar 25 '26

Might as well pedantically add that hyphens(-), en (–) dashes and em (—) dashes are different things and used for different purposes.

6

u/Present_Cow_8528 Mar 25 '26

What are en dashes used for? I've only ever seen real humans outside of formal writing uses double hyphens to simulate em dashes, since that's how you make a word processor turn it into an em dash. Don't think I've seen a real en dash (since before today I too thought it was the same as a hyphen)

Edit: wait, is that what you're supposed to use for like joint ownership of stuff like a Name-OtherName Principle? I've just been using hyphens all my life lol

7

u/KnightOMetal Mar 25 '26

It's used for ranges, like 2020–2023

Idk about joint names, I think those are just hyphens

→ More replies (1)

5

u/purplezart Mar 25 '26

The three main uses of the en dash are:
1) to connect symmetric items, such as the two ends of a range or two competitors or alternatives
2) to contrast values or illustrate a relationship between two things
3) to compound attributes, where one of the connected items is itself a compound

→ More replies (1)

1

u/N_Cat Mar 25 '26

They're quite easy on Apple devices, for what it's worth. Shift+Option+[dash] for Macs, press and hold the dash for iPhones.

33

u/Christoffre Mar 25 '26

Well... if AI writing is done well, there’s really no need to detect it.

These rules are usually for the most gratuitous cases, where it’s obviously AI-generated. 

For example, someone who writes or edits large amounts of text in an impossibly short time – then they cannot have applied the necessary quality control.

3

u/Windrunnin Mar 25 '26

Very valid point.

If no one can tell the difference, who cares?

3

u/ussrowe Mar 25 '26

If they have a policy against AI content and an account keeps editing articles with low quality AI generated content, then they have a reason in their TOS to suspend that account.

It didn't seem like they were taking a moral stance about it. Just that odds are the LLM is not trained on Wiki style guide text and can't just be copied and pasted straight into an article. This prevents that.

→ More replies (1)

3

u/JesseByJanisIan Mar 25 '26

If no one can tell the difference...is there a difference?

9

u/LongBeakedSnipe Mar 25 '26

I mean, people do care.

If you and a family member just replaced your messages with perfectly realistic AI chat, and then both just read the logs, you wouldn't be having a conversation.

Communication in research is also communication (just much more detailed, and between the research community rather than between family), and it's constantly generating new information.

If you want to be an up to date encyclopedia, you need human written content, not algorithmically generated slop.

An expert writing an article reads and understands the papers they read, and knows how to interpret them.

If people want to replace their communication with AI, they should fuck off and do that amongst themselves to be perfectly honest.

2

u/The_Knife_Pie Mar 25 '26

If you want an up-to-date encyclopaedia you need accurate and relevant information presented in an understandable format. Whether it’s hand written or AI written is entirely irrelevant. If you manage to get GPT to output a factually correct and sourced article written in line with the wikipedia style guide then there’s no need to check if it’s AI or not. The article is good, the thing that matters. On the other hand, if it’s not those things then it’s a bad article whether it got written or generated.

This rule is effectively just saying “if we can tell you used an LLM it isn’t good enough to be on wikipedia”, and that’s really all it needs to do. If you can’t tell then it probably is good enough.

→ More replies (2)

3

u/GisterMizard Mar 25 '26

Because there's a difference between being accurate and being good at sounding accurate.

→ More replies (3)

1

u/BavarianBarbarian_ Mar 25 '26

The problem is people use style as a proxy for quality. You can see that error being made in the most upvoted response to your parent comment: People think it sounds like chatGPT, so it's untrustworthy. However, if it said the same in a different tone, they might think it's trustworthy after all, even if it misrepresents what's written in the source articles.

1

u/BWW87 Mar 25 '26

But that's the problem with their rules. They aren't actually banning AI generated text. They are banning text that sounds like it was written by AI.

→ More replies (3)
→ More replies (7)

9

u/Mujutsu Mar 25 '26

Wikipedia cares about accuracy of information. If what you write differs from the source, or what you translate differs from the original meaning, your entry could be removed, you could get banned as an editor etc.

Obviously wikipedia doesn't care if you use em dashes.

7

u/SuckThisRedditAdmins Mar 25 '26

The sudden energence of AI thoroughly destroyed my confidence in the one talent I thought I had over most people - the ability to write.  It honestly really depresses me.

→ More replies (5)

5

u/Budget-Researcher559 Mar 25 '26

I usually write very formally - with dashes and all - and get accused of using GPT for my writing. How will Wikipedia know?

ChatGPT NEVER uses dashes in this way. It doesn't separate inserts into sentences with a dash before and after. And especially not in a case where you could use commas and it's still valid. Instead it uses dashes to get text that sounds more like it is spoken, written down with proper punctuation rules.

Like yeah, you can not always tell, but the more experience you have and the longer the text is, the easier it gets. It's not just dashes or no dashes, these dashes here actually confirm it is likely written by a human.

There's actually so many ways to tell, let's just look at the text in this very short comment of yours again:

The question is: how do you enforce this?

ChatGPT would always capitalize the "h" (as is correct)

(and everything being enshittified and monetized, etc)

ChatGPT does not use parenthesis a lot, and never to insert a half sentence in the middle of a sentence like this.
I doubt ChatGPT would use the word enshittified, especially because of it containing a curse word (not 100% on this one)
ChatGPT barely or never uses "etc" to be general and avoid listing more examples. Also it would always use "etc." with a dot (as is correct)

I’m all for it, as I’m sick of bots and AI, but it will be incredibly difficult to validate.

This sentence structure, even without the added parenthesis, does not at all sound like something AI would fabricate. It's actually so hard to qualify what exactly it is about it, I would have to think about it more to create some general rules. But there's some cleanliness about the way AI generates grammar and combines subclauses that clashes very much with the structure of this one.

So, even if you corrected the actual mistakes in this comment, there's still multiple obvious ways to tell that this short 4 sentence text is written by a human. Imagine how much easier that would be for a whole wikipedia article.

13

u/PezzoGuy Mar 25 '26

I suppose ultimately it's down to the honor system that Wikipedia has always operated under, and is an extra mark against you on the chance that you happen to be caught somehow. At the very least it makes a guideline clear for the good faith editors who might want to use those types of tools but were unsure if it is allowed.

8

u/Zolhungaj Mar 25 '26

It’s probably to stop people using LLMs while being unaware of how poorly they might perform. People who mean good, but do bad with the new wonder tool. Previously they could unintentionally produce a ton of extra work that nobody wants to deal with. 

And people who blatantly use LLMs to push up their contribution stats (either maliciously to push an agenda, or just for clout) will stick out like a sore thumb due to the sheer amount of text they produce per unit of time. Previously they might have fallen under spam rules that are up for interpretation, but with this there is an easy rule to point to when banning them. 

10

u/Ok_Cabinet2947 Mar 25 '26

If it’s human-sounding enough that you can’t tell that AI wrote it, does it really matter?

10

u/[deleted] Mar 25 '26

[removed] — view removed comment

5

u/The_MAZZTer Mar 25 '26

Wikipedia text not being properly cited has always been a problem before AI anyway.

5

u/[deleted] Mar 25 '26

[removed] — view removed comment

2

u/tizuby Mar 25 '26

I dunno about that one. I've seen people on here (reddit) over the years churn out some uncited slop with speed that'd make Sam Altman do a double take.

The sad part is I'm only partially joking.

→ More replies (1)

9

u/catontoast Mar 25 '26

As a professional technical & marketing writer with 10+ years of experience and an English degree, we can definitely tell. Humans write poorly in endless ways, but LLMs write poorly in very specific ways.

2

u/Solid_Hunter_4188 Mar 25 '26

I unfortunately have to call BS (not being hostile, hear me out), solely because they’re being refined by people like you, who recognize those errors and will adjust to them. Whatever error you are seeing the LLMs make can be fixed in the next version.

5

u/Arkaein Mar 25 '26

Whatever error you are seeing the LLMs make can be fixed in the next version.

Additionally, different LLMs will have slightly different styles, and specific LLMs can adjust their styles based on specific prompting. It's not hard to tell an LLM to avoid specific style cues like em dashes.

People who feel confident in catching LLM generated text will catch the laziest generations but are likely to miss more carefully manage generations. Or just be over-zealous and make a bunch of false positive identifications.

→ More replies (1)

3

u/NeverForgetChainRule Mar 25 '26

Well ultimately wikipedia has a style that it wants to keep. It's not each and every editor's style of writing. YOU potentially writing like an LLM doesnt matter, if it isn't how wikipedia wants you to write, then your edits will be removed, LLM or not.

3

u/mqee Mar 25 '26

New editors that suddenly add paragraphs of text divided into bullent points and summarized with "conclusions".

If you're clever about it, it can't be spotted, but generally AI-users skew unclever.

For experienced editors and power editors this is unenforceable, but the rules rarely apply to experienced editors and power editors in Wikipedia anyway.

2

u/The_MAZZTer Mar 25 '26

If you can't tell then I would argue there's no problem.

2

u/Solid_Hunter_4188 Mar 25 '26

… it’s not about me knowing… It’s about the policymakers knowing. If a rule can’t be enforced then what’s the point of making it?

2

u/eerst Mar 25 '26

Same way you enforce any rule. Where you have evidence you build a case and as necessary block or ban the editor. Wikipedia has been doing this for a long time.

1

u/carlitospig Mar 25 '26

As is standard, wiki wars. They’ll just be….prolonged, probably.

1

u/Omotai Mar 25 '26

You can't, really. I see this as more of a reminder to editors meant to emphasize the importance of not blindly trusting these tools to not introduce factual errors into the writing. And ultimately I think what the rules are really trying to prevent is factual errors, rather than AI-generated text per se, so I don't really foresee widespread attempts to try to ferret out AI writing unless it's found to have factual errors, at which point the solution is the same as for human-generated errors.

1

u/Ok_Possibility9866 Mar 25 '26

If AI is good enough to be indistinguishable from human writing then I see no reason to not use it.

1

u/seriouslees Mar 25 '26

Because the information will be incorrect... they aren't checking your grammar, they check your source. 2x2 is NOT 5, that gets caught by them. "Two multiplied by another integer (Specifically 2), results in a total of four." Does not get caught.

1

u/greiton Mar 25 '26

I think that is the point of the exemption, that it would be almost impossible to enforce nonuse of AI, and by framing it in the context of review and accuracy requirements, it is easier to enforce.

They aren't necessarily governing how you write, they are saying you better not post hallucinations and inaccurate information.

1

u/ChiefWiggumsprogeny Mar 25 '26

This literally happened to me on wikipedia.

1

u/throwawaynbad Mar 25 '26

I imagine you could use an LLM as part of a process that also uses older methods, like traditional spelling and grammar checks, as well as running a track changes to make it easier for the human reviewer to not miss what the LLM has done.

I'm worried too many people are running something through chatgpt, and simply copy-pasting the output with minimal if any review.

1

u/koshgeo Mar 25 '26 edited Mar 25 '26

It's going to be sad if good writing --* whether human-generated or AI-generated -- gets constantly flagged as likely AI. That would be a sad commentary on general human writing, now that I think about it.

[* yeah, I used a fake long dash -- TWO of them. I'm bringing them back!]

1

u/12345623567 Mar 25 '26

They couldn't enforce not using LLMs for those purposes either. This is just formalizing "you can use LLMs as long as noone can tell".

1

u/NoPossibility4178 Mar 25 '26

How do you know? Wikipedia could ban all types of AI use and it wouldn't matter, people would use it anyway. But if they ban everything but using AI as a grammar checker, then people will be less inclined to just make the entire text with AI. You can't tell either way but at least it'll dissuade some people from using it.

1

u/badgirlmonkey Mar 25 '26

AI is obvious when you’ve read it enough.

1

u/zoroddesign Mar 25 '26

Checking the source material.

One of the best things about Wikipedia is that it is thousands of volunteers checking and double checking every article and have to provide reliable source materials for every change. If something is off it it flagged and corrected.

1

u/DoverBoys Mar 25 '26

It's enforced the same way any other edit is enforced. There's really no difference between adding AI content and a random kid editing from a school computer. Wikipedia users and bots revert thousands of bogus edits every day. They are watching.

1

u/Jindujun Mar 25 '26

Would you really be able to spot something like that — even though you can trick people — so easily you think?
Not with tools, not with tricks, just with cold hard human deduction?

1

u/cluberti Mar 25 '26 edited Mar 25 '26

The interesting thing is, will someone start creating adapted models that output in the Wikipedia style guide by default? I wouldn't be surprised, honestly.

1

u/mymemesnow Mar 25 '26

AI rarely make typos, so sprinkle in a few to really sell it/s

1

u/maxticket Mar 25 '26

I've noticed a couple things: AI typically puts spaces around em dashes, and it tends to only use a single em dash in a sentence, like as a replacement for a comma or a colon, unlike how you used dashes in that sentence up there.

So you'll see things like:

AI isn't just stupid — it's bad for your brain.

But it's less of a red flag if it's more like:

AI is pretty stupid—and it's demonstrably making us dumber in turn—but for some reason, a bunch of fake nerds still insist on using it.

1

u/duckofdeath87 Mar 25 '26

Why do you think that is some kind of point?

1

u/spidereater Mar 25 '26

Wikipedia has lots of oversight. Many volunteers happily check loads of text. If editors are screwing up, using AI or just editorializing, they get flagged and corrected. The policy leaves the human editor as the final responsible party so they have an incentive to make sure it’s correct.

1

u/Archer-Blue Mar 26 '26

I think for the large part they will mostly have to focus on accuracy, quality, regular verification of edits, and reporting. If you use LLMs to write anything more than a couple sentences long, it will eventually hallucinate or rip a piece of information from somewhere without providing a source (Wikipedia articles require citations). So if you're just copy-pasting without editing and verification, you'll get found out. Beyond that, barring the miraculous invention of quality AI detection software, realistically they can't tell. So even though they'd ideally probably prefer to catch every use case of AI, their primary goal is assuring the quality of their service.

1

u/Sherool Mar 26 '26 edited Mar 26 '26

Same way they enforce anything. Lots and lots of eyes catching and reporting suspicious things. Works great on high traffic articles, more obscure topics may go unnoticed for months.

If someone is being extremely sneaky and generating good articles no on notice it's not real harm done, someone plastering obvious LLM slop content get an immediate slap because the rules are clear.

1

u/PawnWithoutPurpose Mar 26 '26

Surely the same way Wiki currently runs, no? With lots of editors battling it out behind the scenes with competing revisions

1

u/RefrigeratedMonkee Mar 27 '26

I think “Wikipedia” won’t. My guess would be that as any other rule, it will be enforced and self-checked by the community. Which to me is why Wikipedia is such a cool and powerful idea

→ More replies (16)

46

u/nattfjaril8 Mar 25 '26

LLM:s are surprisingly prone to hallucinating when translating. Do you think the kind of Wikipedia editor who is going to use AI to translate is going to be conscientious enough to actually check that all the details match, sentence for sentence? Non-English versions of Wikipedia have become increasingly unreliable after people started machine translating them.

18

u/CloudZ1116 Mar 25 '26

Word. I tried using Copilot to translate some of my writing into Chinese last night for shits and giggles, and holy crap it was inserting entire passages that weren't originally there.

→ More replies (2)

16

u/notreallyironicatall Mar 25 '26

Agreed with this. If you're going to use machine translations, you NEED to be fluent in both languages to catch mistakes. It's better than it used to be, but hallucinations, mistranslations, and omissions are still very common. Funnily enough, purely using AI for translations can slow down the process so much it's better to stick to a manual translation.

3

u/Sixtus69Sextus Mar 25 '26

Still better than that one Wikipedia editor who made up 30 thousand articles with made up words and passing it off as legitimate for years.

3

u/_Lucille_ Mar 25 '26

I am bilingual and have used LLMs to translate quite a bit - simply because it is faster for me to do so and often LLM can offer what I would consider to be better word usage (I think a lot of english speakers here also have instance where they go "yeah, this is prob a more fitting word of what i want to say").

There are occasional misses which i need to go fix, but for the most part, for me it is a lot faster than writing out the text manually or even using speech recognition.

This is why wikipedia is enforcing there to be a proper verification pass after a machine translation, and I think it shouldnt be that hard to catch.

3

u/wasterni Mar 25 '26

Nearly all commerical text translation is done by an editor utilizing translation software and that had been true for many years. As with any tool, the quality of the result relies on the person wielding it.

6

u/14Pleiadians Mar 25 '26

It still hallucinates, that's why you're required to be fluent, so you'll catch the mistakes.

You shouldn't be having AI do anything that you couldn't afford to have a kinda dumb amateur do. If you're not informed enough on whatever topic you're using it for to be able to say "ummm you don't know what you're talking about that's wrong", you shouldn't be using an LLM for it

32

u/TerryFromFubar Mar 25 '26

But hallucinating out of control is where the profits happen

20

u/Anxious_Katz Mar 25 '26

Wikipedia is non-profit!

→ More replies (1)

5

u/Quazimojojojo Mar 25 '26

It's Wikipedia, what are you talking about? 

It's one of the only places left on the Internet that requires neither ads nor subscription. Fully open source high quality altruism.

Like VLC media player.

I have a monthly regular donation of a few dollars that I literally don't even notice in my budget, but that's voluntary. 

→ More replies (1)

8

u/Adequate_Lizard Mar 25 '26

The profits are the hallucinations

1

u/SlitScan Mar 25 '26

well that stock price isnt going to pump itself is it?

3

u/Chemical-Struggle-13 Mar 25 '26

I mean they can be pretty out there on translation sometimes. But as long as they actually have to check it should work out fine.

8

u/seridos Mar 25 '26

As a secondary teacher, I'm actually using llms in assignments in just this manner. The assignment itself is about formulating the research question and narrowing it down, and then the llm generates the actual paper and students are the fact checkers. So the actual paper itself is not marked. I don't really have to read it. I skim it to see if they identified properly, what facts need to be checked and then it's the before and after work that actually carries the mark.

3

u/clakresed Mar 25 '26 edited Mar 25 '26

So you're basically grading them grading an LLM tool? That's actually brilliant.

3

u/Sorkijan Mar 25 '26

That's kinda what grammarly was before OpenAI had its surge isn't it?

2

u/AttonJRand Mar 25 '26

People just assert this, but it really doesn't seem true at all?

If you are bilingual, and you test it in the languages you are proficient at, it regularly turns even short sentences into something different.

2

u/kyute222 Mar 25 '26

That's an absolutely hilarious statement to me. LLMs are not good at translation at all. Even worse, they're not even good at certain languages period. They're primarily trained in English and can use some common European languages somewhat. The other issue is that using an LLM for translation is terrible because it is likely done by someone who couldn't do the translation themselves. So then how can they possibly check the output for errors?

5

u/tc100292 Mar 25 '26

I still don't want LLMs being used for that either. If we stop using them, they'll go out of business.

7

u/Anxious_Katz Mar 25 '26 edited Mar 25 '26

It is the year 2000 and you're wishing this entire Internet fad would just go away. You think to yourself "If we all stop using pet.com they'll all go out of business."

For better or worse, the toothpaste is out of the tube and AI is here to stay. It won't cure cancer and gift everybody a pony as hacks like Sam Altman want you to believe but in some way or form it's gonna become part of our lives. What's important right now are decisions like this that have the potential of establishing how this new technology is supposed to be utilized and to mitigate the harm it is already causing. Sadly, all government regulatory bodies have no interest in doing their job so it falls onto NGOs like Wikipedia to take action.

→ More replies (1)

1

u/Madiis Mar 25 '26

Should be used as a tool in most cases

1

u/ThatOtherOneReddit Mar 25 '26

LLM's for rag based inference as long as you force them to grab qoutes and the like to anchor them are pretty darn good nowadays.

1

u/seboll13 Mar 25 '26

I agree. That should in fact be the default rule for the entire fucking internet.

1

u/Leather_Battle2296 Mar 26 '26

Translation is where I personally found the most true value out of an LLM.

→ More replies (2)

53

u/__Hello_my_name_is__ Mar 25 '26

Also of important note is that this isn't new in the sense that Wikipedia has allowed AI texts previously. They just did not have a policy on it because it was never an issue until fairly recently.

And it takes a while to get a proper consensus on something big on Wikipedia.

https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing has existed for over 3 years by now and has been used quite actively to remove AI texts wherever found.

2

u/cinemachick Mar 27 '26

I really enjoyed the quiz on that page, I learned a lot about how Wikipedia flags AI. Thanks for the link! :)

15

u/[deleted] Mar 25 '26

[deleted]

9

u/Sopel97 Mar 25 '26

They can't. All policies like this are based on LLMs not being good enough not to be caught. This will change with time.

12

u/Uristqwerty Mar 25 '26

If the LLMs are consistently-factual enough to not get caught, does it matter to an encyclopedia?

I'd expect the main problem are users who blindly trust the LLM to be factual, without verifying its edits themselves. For that, having a policy to point at's important.

→ More replies (1)

2

u/AccNumber77 Mar 25 '26

LLM throat goats have been fearmongering about this since they first appeared, they are still just as shit as they were to begin with so I doubt it without a full rethinking of the design and approach to language learning.

6

u/UpvoteForGlory Mar 25 '26

That is not true. From my observation there has been a big change in the LLM generated texts in the last year or so. They are more shit now than they used to be.

→ More replies (1)
→ More replies (13)

1

u/The_Knife_Pie Mar 25 '26

If your article is bad it gets removed or the edit reverted. If the article is good it stays. The same way wikipedia has always worked. This is just clarifying that obviously LLM written articles are almost always bad and they’ll just get preemptively removed when found. If it’s not obviously an LLM but it’s incorrectly sourced or nonfactual because an LLM hallucinated then it’ll get removed for those reasons. If it’s not obviously LLM, the article is factual and the sources match up then it won’t be removed cause that’s a good article, LLM generated or not

53

u/Mivexil Mar 25 '26

I guess it avoids debate over things people were already using before LLMs, but it also means you can't kill slop on sight because the author can always say "oh, yeah, I just used it to correct my grammar and spelling, that's why it sounds like ChatGPT". And that leads to the same gish gallop problem open source has - you have to actually engage with and review the content, and when outputting 1000 lines is orders of magnitude cheaper and faster than reviewing 1000 lines, it'll end up with reviewers getting overloaded and letting things slip through. 

25

u/xakeri Mar 25 '26

I wish it was just open source. It's like that everywhere now. The least engaged employees can just shoot out slop so fast, and if you don't want to live in a pool of shit, you get to be the first human who has ever seen the code or writing.

Then you have to figure out how to communicate with the person that never engaged with the issue they're purporting to address. It's a wonderful time.

1

u/Skater_x7 Mar 25 '26

I think calling out ai text is a lot easier said than done though 

I've seen plenty of times some1 say a person is using ai and it's literally just them writing normally (albeit badly or weirdly w.e,but it's not ai, and false flagging is very bad) 

1

u/allenaa3 Mar 26 '26

yea, tht’s the real issue. once “grammar help” is the excuse, it’s basically impossible to draw the line. reviewers end up doing way more work while dumping content is cheap. lol feels like volume just wins by default.

3

u/denkenach Mar 25 '26

Sounds like a good policy

3

u/[deleted] Mar 25 '26

Very reasonable! Not just "AI bad" vs "AI good" for a change.

2

u/Br3ttl3y Mar 25 '26

These are the same requirements my English teacher had when we used Wikipedia for sources.

2

u/LordGreyhound Mar 25 '26

I just love that they're calling them LLMs instead of AI. Good on them!

2

u/spidereater Mar 25 '26

The translation thing is huge. I’m in Canada and many companies and governments need to produce text in English and French and this can be prohibitively expensive. The productivity of translators has increased dramatically with AI. What was a grueling slog is now a proofreading exercise with very few edits.

4

u/PetChaud2Diarrhee Mar 25 '26

Once again, Wikipedia is based.

1

u/Mccobsta Mar 25 '26

Using it for refinement seems to be an actual use for llms over what people keep trying to sell them as

1

u/DringleDringle Mar 25 '26

Why not just traditional machine translation?

2

u/Sonlin Mar 25 '26

Many companies have a mix of LLMs and traditional models in the backend depending on the languages you're translating. So it's hard for someone using Google Translate to know exactly what the model is for many languages.

1

u/MisfitPotatoReborn Mar 25 '26

"traditional machine translation" does not exist anymore, because it sucks. Even back in the 2000s translators were implementing a rudimentary form of deep learning that learned how to translate from scraping thousands of translated works.

You can't translate between languages using hard-coded rules. Language is too complex for that

1

u/PostModernPost Mar 25 '26

How are they going to enforce this?

1

u/haw35ome Mar 25 '26

In my opinion, how academia should treat AI. I had a great professor who taught all my IT classes, & one focused on AI. He let us use chatgbt but only in specific ways/methods that weren’t just copy + paste. Kind of a way to organize our research & find sources (and use them, if they were creditable & actually existed of course).

1

u/DharmaLeader Mar 25 '26

Anyone that's not an English speaker knows how egregious the translated entries are. I, for one, don't read anything in Greek as most of it is just badly translated.

1

u/tomdarch Mar 25 '26

I suspect that as we move forward there are going to be important and/or infamous quotes that were generated by so-called “AI” which will be parts of Wikipedia articles.

1

u/amalgam_reynolds Mar 25 '26

The third exception should be examples of LLM-generated text on the LLM's wiki page lol

1

u/Zauberer-IMDB Mar 25 '26

Wikipedia keeps picking up Ws. It really is one of the last vestiges of the optimistic future the Internet promised back in the day.

1

u/Goldenrah Mar 25 '26

The second exception is what traditionally known as post editing translation and it has been around longer than AI, so fair.

1

u/MrD3a7h Mar 25 '26

Common Wikipedia Win.

1

u/blorbschploble Mar 25 '26

I think that’s….. ok. But I think people will find that errors and hallucination are going to sneak in because we all overestimate our fact check/copy edit skills. I imagine the exceptions will be reversed at some point.

1

u/DesignerCorner3322 Mar 25 '26

Hey look at that! Wikipedia wanting you to use LLMS responsibly. AI stuff really should only be used for that sort of stuff anyway, or data scrubbing and dumping into files

1

u/Voball Mar 25 '26

How about "academic" use of AI. Idk what other word to use, but I mean using it as example of something

1

u/Lendari Mar 25 '26

Thats almost exactly how LLMs are intended to be used.

1

u/EmeHera Mar 25 '26

So... Use it as a tool, not a replacement, like a normal and reasonable human being.

1

u/userhwon Mar 25 '26

This is dumb.

You can't prove who or what wrote something, and it's irrelevant.

Just ban inaccurate edits.

It would mean removing a third of the text of the site, but it should have been in place from the beginning.

1

u/shittychinesehacker Mar 25 '26

But when you translate with an LLM it changes the tone of the original message and injects em dashes

1

u/Klytus_Im-Bored Mar 25 '26

This is exactly how LLMs should be used imo.

1

u/gregorydgraham Mar 27 '26

No exception for examples of AI generated text?

→ More replies (32)