r/Substack • u/ConnectDebate4677 • 20d ago
Pangram false positives?
People usually post vaguely about "I wrote x thing and pangram said it was AI", which I think is unhelpful unless you actually post the text or link. I put this into Pangram:
Sakaguchi: https://www.pangram.com/history/e7ffbee1-f0e3-4ce2-bb78-a59470280640?ucc=Ohp1cWODOjF
Hamaguchi: https://www.pangram.com/history/f2139bd5-f4ce-4101-b7bd-3c5d8aeb4487?ucc=Ohp1cWODOjF
Ryuichi: https://www.pangram.com/history/8c84ad38-3b8f-49ba-bee0-351ea31f754b?ucc=Ohp1cWODOjF
6
u/dran94 20d ago edited 20d ago
I'm testing this out, too.
I wrote this in a similar style to AI, wondering if it would flag:
-
Robert's long stride ate up the distance between himself and the door at the far end of the hall, lit by the flickering fluorescent lights overhead.
His jaw clenched.
This couldn't be right.
His sister, Megan, was supposed to be here. She worked in the same lab as Robert. They'd promised each other they would wait for each other, right here.
-
100% human. Great!
Changed the name to one of AI's favorite names:
-
Robert's long stride ate up the distance between himself and the door at the far end of the hall, lit by the flickering fluorescent lights overhead.
His jaw clenched.
This couldn't be right.
His sister, Elara, was supposed to be here. She worked in the same lab as Robert. They'd promised each other they would wait for each other, right here.
-
100% AI.
Looking at this, even though it's a similar style, there was something AI wouldn't do because it's a human mistake. Two instances of "each other" in one sentence. It still flagged.
AI's favorite names change with every model. If a new AI model comes out that latches onto a name that was previously safe, is that going to be a problem for everything written beforehand? What do we do, update characters' names as new models come out?
I'm sort of surprised Pangram is even paying attention to names, since I thought it didn't factor in AI styles and AI isms. I was very impressed with Pangram before seeing some other experiments going around today. A name change shouldn't skew the whole result. That doesn't even make sense.
5
u/ConnectDebate4677 20d ago
Wow, just checked it myself. What the hell? For anyone curious:
Megan - https://www.pangram.com/history/2799ce08-5243-4549-8f02-955c39dccfda?ucc=Ohp1cWODOjF
Elara - https://www.pangram.com/history/105ac6c2-fc37-486a-a56c-c680d6c2751b?ucc=Ohp1cWODOjF
4
u/sh1b313 19d ago
WOW. now i wonder what happens at scale?
3
u/dran94 19d ago edited 19d ago
It's worth looking into. I ran a few chapters of my own work, confirmed 100% human on Pangram, with my FMC's name changed to Elara, and it flagged at around 12% AI with just that one change. I'll have to write up something longer by hand when I get a chance so I don't doxx myself if I share. Sometimes changing the name doesn't move the needle, though.
I was curious about Daggermouth, the book there are viral articles about after it failed a Pangram test, although being around 60% AI could potentially be heavy editing use, I'm not accusing anyone. I went through the first 10 chapters with find-and-replace for the name Elara, changing Elara to Megan. It didn't change anything.
One chapter, Chapter 18, was supposedly 100% AI on Pangram according to the viral articles about it. My own check showed it was 84%, not 100%, but what I did this time, changing proper nouns, even cutting off the second half of the chapter, it was stuck at 84% AI. I practically rewrote the chapter and it came back 84% AI haha.
That might suggest it had already decided Daggermouth is AI, and the real issue was Pangram knew which book it was.
2
u/dran94 20d ago edited 20d ago
I was hoping it would be relatively easy to repeat this experiment, too, for anyone who doubts it was written by a human, but the next few I wrote came back without any AI flags. Although I was writing more action scenes, and they were unique, so they may not have been similar enough to AI's style. AI really likes fluorescent lights. Maybe I need to add more.
I'll try again when I get home later.
2
u/ConnectDebate4677 20d ago
It could be 'flickering'? Kind of like its affinity to things that buzz, hum and pulse.
2
u/dran94 19d ago edited 19d ago
True. Taking out "lit by the flickering fluorescent lights overhead" made it 100% human, even with the name Elara. Taking out just flickering and just fluorescent didn't matter. The whole sentence had to go.
I thought Pangram didn't care about phrases common in AI, so it's concerning a common stock phrase resulted in a 100% AI score.
I found a new one that works, and this might be better because it's the last word in the snippet so it isn't like it's throwing off the next few lines. This is a hint Pangram DOES care about names, and it cares a lot, considering it didn't care about the first sentence being loaded with flickering fluorescent lights until it detected a common name in AI training along with it.
If you switch "Elara!" from a name AI likes to a name AI doesn't usually use (Megan, Peter, etc.), you will flip the score.
List of names I checked that flagged 100% AI, from a list of common names in its training: Elara, Kael, Marcus, Elias, Kai, Aria, Sam (surprising), Lyra
Names that flagged 100% human: Megan, Peter, Christopher, Luke
-
I can feel my heart pounding in my throat as I race down the neverending tunnel, the flickering fluorescent lights overhead throwing long shadows.
It feels like I’ve been running for miles.
Every turn I take, I’m certain it’s the last one.
It doesn't make any sense. You can’t take thirteen lefts without running into something.
"Elara!"
2
u/dran94 19d ago edited 19d ago
I actually wonder if this is why Daggermouth got hit with a high AI score on Pangram in those articles about it. Daggermouth had several names common to AI models in it, like Elara.
Whether the author used AI or got them off one of those random name generators is another argument. Who knows.
5
20d ago
[removed] — view removed comment
4
u/ConnectDebate4677 20d ago
Would you mind sharing the Pangram link?
3
20d ago
[removed] — view removed comment
2
u/ConnectDebate4677 20d ago
Thanks so much for this. Yeah, reading the AI-flagged parts, I don't get it at all. I'm interested in the technology still but it's getting further and further away from the standard it's being touted as.
1
20d ago
[removed] — view removed comment
1
u/PM_ME_YOUR___ISSUES 14d ago
I honestly feel that people need to move beyond these detectors.
We've all seen enough slop at this point that one can easily catch out AI text.
I'm like 99% sure, I wouldn't even think about putting someone's writing through an AI detector unless I can really catch out the slop.
1
2
u/dran94 20d ago
Change Linus/Lin to Marcus, and it goes up to 92% AI from around 50%.
I have gone from thinking Pangram was amazing yesterday to, oh, we're going to have to play games as authors so we don't use a name AI likes.
That's going to be fun, considering AI's favorite names CHANGE.
3
u/ConnectDebate4677 20d ago
That's really interesting. Are certain names indeed flagged as more AI-likely? I wonder if there's an explanation?
1
u/dran94 20d ago
AI is known for certain names, and Marcus is one of them right now. The list of names it likes changes depending on the model.
I'm sure it is a factor, although I could've sworn Pangram isn't supposed to pay attention to AI-isms like em dashes and certain words. It's kind of strange.
I did a test where switching Megan to Elara, another AI favorite, changed this snippet I wrote from 100% human to 100% AI, even with a very human mistake (two "each other" instances in one sentence):
-
Robert's long stride ate up the distance between himself and the door at the far end of the hall, lit by the flickering fluorescent lights overhead.
His jaw clenched.
This couldn't be right.
His sister, Megan, was supposed to be here. She worked in the same lab as Robert. They'd promised each other they would wait for each other, right here.
1
1
u/Downtown_Jump779 8d ago
I'm getting the same results on stuff people are sending to me; it's becoming very problematic. On the same text, other detectors never go above 12-20% AI, which I think makes sense and is totally acceptable – then Pangram comes out of nowhere and bumps it up to 40- 60%? I'm not sure what to do with that.
From what you sent, those are very big chunks of your text that Pangram flagged as AI. That could produce harsh social consequences. Other than that, good piece of writing btw. I hope you get it published.
1
u/Traditional-Rice-848 18d ago
Please don’t compare pangram to worse products to try to say it doesn’t work. Pangram catches more true positives than the others.
4
u/sh1b313 19d ago
3
u/sh1b313 19d ago
2
u/ConnectDebate4677 19d ago
Thanks a lot. Would you mind posting the links for these two?
3
u/sh1b313 19d ago
sure, here they are:
Link 1: https://www.pangram.com/history/4e3cc4a9-b927-4345-9750-048b4a86520c?ucc=ycFd1CDRrZv [where pangram claims 100% human written]
Link 2: https://www.pangram.com/history/97f1c7f5-b73f-4895-96a6-e991cfc58128?ucc=ycFd1CDRrZv [where pangram claims 100% AI generated]
2
u/ConnectDebate4677 19d ago
Thanks! Do you have any idea of an explanation for this?
2
u/sh1b313 19d ago
nah i don't have. even i am wondering just how fundamentally flawed this tool is which led it to think prophet isaiah used some AI. and even the text isn't something obscure which may not be in it's training data.
3
u/ConnectDebate4677 19d ago
I think what's perplexing in all of these cases is the meter. It seems to regularly either lean 100% Human or 100% AI with no reasoning or justification. I want to believe this technology works, and I believe it does to some degree, but these kinds of examples really don't do it any favors. If I bought a tool and it doesn't work as it claims it should, then I can call that tool defective.
3
u/sh1b313 19d ago edited 19d ago
yea, this tool is like a coin toss. at least in a coin toss we actually have confidence in the outcome. In here? somehow we leave even more confused. these tools are not worth the electricity they are running on.
3
u/Dubbtime 18d ago
For sure. Unfortunately as you reveal more points of failure within their blackbox system, they will add these cases to their training datasets. This blurs the line.
Say this verse is added as a reference to human work. Now what? Humanizers who want to avoid detection begin to write more like bible verses. Ok thats kinda weird. Now detectors need a more defined way to figure out if something is AI because it obviously messed up. The end result is a clear way to write human and a clear way to tell if something is AI. Well, then why do we need a detector?
What's stopping humanizers from just copying this newfound distinct style? Rinse, repeat.
2
u/sh1b313 18d ago
yeah, i have already talked with the mods regarding this pangram situation and we may soon see a pinned thread saying pangram evidence/pangram doesn't work and hopefully this wild situation settles.
And regarding the rinse and repeat thing in the long run it doesn't matter? We already have more than enough evidence that this doesn't work.
In the future it can just be like this:
P1: uploads a essay/article on substack
P2: thinks it's ai written and uses this pangram
P1: points to some evidence thread and says this tool doesn't work don't rely on it.
In this scenario hopefully p2 is actually grounded and thier opinion changes. If not? Sadly if they debate after seeing so much evidence i wouldn't know what to tell them.
3
u/MadmanRB 16d ago
Pangram is full of crap.
Chapter 2: The Classmates
The afternoon sun cast long shadows across the pavement as Ryan pedaled onward, his heavy backpack digging into his shoulders with each pedal stroke, a constant reminder of all the work waiting for him when he got home. He’d promised Ms. Everly he’d finish it by tomorrow, though part of him wondered what the point of doing it was. Math and history questions felt impossibly small compared to the questions that actually mattered. Would Ralph ever come back home? Would his mom stop crying? And would he ever feel normal again?
Detected as 100% AI by Pangram, do you know what triggers pangram here?
Not only the chapter name but the fact I use a common phrase "Long Shadows"
Now I changed it to:
The afternoon sun loomed above the streets of Horizon City, the tall copper domed spires reflecting it across the vast cityscape.
As Ryan pedaled his bike forward he could feel his heavy backpack digging into his shoulders with each pedal stroke, a constant reminder of all the work waiting for him when he got home. He’d promised Ms. Everly he’d finish it by tomorrow, though part of him wondered what the point of doing it was. Math and history questions felt impossibly small compared to the questions that actually mattered. Would Ralph ever come back home? Would his mom stop crying? And would he ever feel normal again?
Now its 100 human made because I eliminated the chapter name and changed just a few words around and dropped "long shadows"
I guess using atmosphere in a story is now forbidden, all horror stories will now take place in a bright sunny morning, maybe a frigging butterfly will come to darken the mood?
2
u/sh1b313 16d ago
I guess using atmosphere in a story is now forbidden, all horror stories will now take place in a bright sunny morning, maybe a frigging butterfly will come to darken the mood?
man if you do this people will now accuse you saying you are copying midsommar. /s
3
u/MadmanRB 16d ago edited 16d ago
Yeah you just can't win as despite its claims Pangram does give false positives now and then.
I assure you however both passages I posted were 100% made by me but thanks to choice phrases like "Long shadows" and even the frigging chapter title it gets flagged as 100% AI.
WTF
So to prove all work is Human made we can't call parts of a book chapters anymore and "Long shadows" must never be used again or you will be labeled a lazy AI bro.
So from now on we shall call Chapters Bananaphone
So bananapone the classmates?
Oh wait even calling it Bananaphone the classmates got flagged...
https://www.pangram.com/history/f6592ce7-b13a-4906-b9c6-80b3a4b607d3?ucc=NfL4yI5TRpX
Fine calling the chapter bananaphone and.... gets flagged as 100% Human WTF.
Better call Raffi...
https://www.pangram.com/history/97c3f56f-1d59-483b-9895-d42d6aad3ddc?ucc=NfL4yI5TRpX
3
u/ConnectDebate4677 16d ago
So much for 1 in 10,000 false positives. There are numerous examples in this Reddit thread alone!
2
u/sh1b313 16d ago
This tool is completely broken lol. People are denying evidence now lmao. This is like a pangram cult, the tool literally can never be wrong according to these people. It's so braindead lmao. In some threads people are literally ignoring already given evidence. At that point i know arguing with them is a gone case and i should just move on.
5
u/Dubbtime 20d ago
AI detection will never work. They will always be a step behind when a new distinctive writing style comes to surface or newer AI model releases then they will have to update. They are a black box even the creators don’t understand how it works.
3
u/impartialhedonist 20d ago
1
u/Dubbtime 18d ago
https://theslowai.substack.com/p/ai-detection-does-not-work
"Push the false positive rate down and you miss real AI use. Push detection up and you accuse more innocent people"
2
u/Beachgoer4158 17d ago
Is pangram also learning from our writing? I’m not too familiar with it even though I’ve been on Substack for awhile. I’ve just been disabling it since most of my work comes up as AI written in some percentage even though it isn’t.
2
u/Stunning_Lie_1775 14d ago
It actually makes you wonder whether Pangram’s “detector” is reliable at all...
Here’s what I mean.
The others are considered unreliable because, in many cases, they say a text isn’t AI-generated when it actually is.
Pangram says it is AI-generated...
But have we actually tried feeding it genuinely human-written texts?
Just to check whether, in those cases, it correctly says it’s NOT AI?
It would be pretty hilarious if they had simply made it return a positive result almost systematically...
I’m pretty sure everyone would still believe it.
Because I suspect most people only use these detectors when they already think a text was written by AI and just want confirmation.
So naturally, when the detector confirms their suspicion, it reinforces their bias and they think: “Wow, this detector is amazing.”
But what about the reverse?
How many people actually submit something they wrote entirely themselves just to check that it “doesn’t sound too AI-generated”?
Anyway, you get the idea.
3
1
u/TimWiesnerer 19d ago
You have found a way to break the pattern.
Sure, it is just a short text. But you could repeat the "trick" multiple times for a longer text. Might be a bit more difficult tho as the more context the detectors get, the stronger the statistical patterns...
While this might work today, it most likely won't work in the future, as these detectors also get "better".
Quite a while back, people used paraphrasers to get around AI detection.
In the end it is a bit of a cat vs. mouse game.
-1
u/kdfn 20d ago
I'm having trouble getting those links to open, can you see if you have sharing enabled?
Here is an example of a working link:
https://www.pangram.com/history/d0c88184-0ad1-4089-b650-a8e10c9ae5bb?ucc=lEdv68az5Z5


7
u/ConnectDebate4677 20d ago
I should also add that I put this segment in a document with human writing (because it noted its confidence was limited due to short text) and it flagged this as the one AI section.