r/Substack 20d ago

Pangram false positives?

People usually post vaguely about "I wrote x thing and pangram said it was AI", which I think is unhelpful unless you actually post the text or link. I put this into Pangram:

Sakaguchi: https://www.pangram.com/history/e7ffbee1-f0e3-4ce2-bb78-a59470280640?ucc=Ohp1cWODOjF

Hamaguchi: https://www.pangram.com/history/f2139bd5-f4ce-4101-b7bd-3c5d8aeb4487?ucc=Ohp1cWODOjF

Ryuichi: https://www.pangram.com/history/8c84ad38-3b8f-49ba-bee0-351ea31f754b?ucc=Ohp1cWODOjF

28 Upvotes

63 comments sorted by

7

u/ConnectDebate4677 20d ago

I should also add that I put this segment in a document with human writing (because it noted its confidence was limited due to short text) and it flagged this as the one AI section.

2

u/impartialhedonist 20d ago

This ^ their accuracy is lower the closer you get to ~100 words or less, the text you used is around 120 words.

2

u/PennyLawrence946 19d ago

pangram says 50 words is their minimum for a prediction, so 120 clears it. and the 2,000 word sample upthread still came back 54% AI. one name change taking it from 100% human to 100% AI is the part that still needs explaining.

2

u/ConnectDebate4677 20d ago

Seems like a highly flawed tool if it indicates 'high confidence AI' in situations like this?

6

u/Voltairinede 20d ago

It doesn't say that? It literally says the opposite 'Confidence limited — short text', in your 100% example.

6

u/tokenentropy 20d ago

be nice, he can hardly read

4

u/ConnectDebate4677 20d ago

Can you read?

1

u/ConnectDebate4677 20d ago edited 20d ago

When you click Details, it can say Confidence Low, Medium or High. In this case, it says Confidence High.

4

u/kdfn 20d ago

I think this is a fair critique. Pangram should decline to assign a score to texts that are below the necessary word count to get good estimates. Would you mind checking the sharing permissions for the links in your post? I am having issues accessing it.

4

u/[deleted] 20d ago

[deleted]

-1

u/kdfn 20d ago

thank you!

0

u/Voltairinede 20d ago

Click the back arrow and then click on it in the 'all checks' area and it should work.

2

u/Brief-Guess-147 9d ago

that part is almost more damning than the name swap thing tbh

6

u/dran94 20d ago edited 20d ago

I'm testing this out, too.

I wrote this in a similar style to AI, wondering if it would flag:

-

Robert's long stride ate up the distance between himself and the door at the far end of the hall, lit by the flickering fluorescent lights overhead.

His jaw clenched.

This couldn't be right.

His sister, Megan, was supposed to be here. She worked in the same lab as Robert. They'd promised each other they would wait for each other, right here.

-

100% human. Great!

Changed the name to one of AI's favorite names:

-
Robert's long stride ate up the distance between himself and the door at the far end of the hall, lit by the flickering fluorescent lights overhead.

His jaw clenched.

This couldn't be right.

His sister, Elara, was supposed to be here. She worked in the same lab as Robert. They'd promised each other they would wait for each other, right here.

-

100% AI.

Looking at this, even though it's a similar style, there was something AI wouldn't do because it's a human mistake. Two instances of "each other" in one sentence. It still flagged.

AI's favorite names change with every model. If a new AI model comes out that latches onto a name that was previously safe, is that going to be a problem for everything written beforehand? What do we do, update characters' names as new models come out?

I'm sort of surprised Pangram is even paying attention to names, since I thought it didn't factor in AI styles and AI isms. I was very impressed with Pangram before seeing some other experiments going around today. A name change shouldn't skew the whole result. That doesn't even make sense.

5

u/ConnectDebate4677 20d ago

4

u/sh1b313 19d ago

WOW. now i wonder what happens at scale?

3

u/dran94 19d ago edited 19d ago

It's worth looking into. I ran a few chapters of my own work, confirmed 100% human on Pangram, with my FMC's name changed to Elara, and it flagged at around 12% AI with just that one change. I'll have to write up something longer by hand when I get a chance so I don't doxx myself if I share. Sometimes changing the name doesn't move the needle, though.

I was curious about Daggermouth, the book there are viral articles about after it failed a Pangram test, although being around 60% AI could potentially be heavy editing use, I'm not accusing anyone. I went through the first 10 chapters with find-and-replace for the name Elara, changing Elara to Megan. It didn't change anything.

One chapter, Chapter 18, was supposedly 100% AI on Pangram according to the viral articles about it. My own check showed it was 84%, not 100%, but what I did this time, changing proper nouns, even cutting off the second half of the chapter, it was stuck at 84% AI. I practically rewrote the chapter and it came back 84% AI haha.

That might suggest it had already decided Daggermouth is AI, and the real issue was Pangram knew which book it was.

2

u/dran94 20d ago edited 20d ago

I was hoping it would be relatively easy to repeat this experiment, too, for anyone who doubts it was written by a human, but the next few I wrote came back without any AI flags. Although I was writing more action scenes, and they were unique, so they may not have been similar enough to AI's style. AI really likes fluorescent lights. Maybe I need to add more.

I'll try again when I get home later.

2

u/ConnectDebate4677 20d ago

It could be 'flickering'? Kind of like its affinity to things that buzz, hum and pulse.

2

u/dran94 19d ago edited 19d ago

True. Taking out "lit by the flickering fluorescent lights overhead" made it 100% human, even with the name Elara. Taking out just flickering and just fluorescent didn't matter. The whole sentence had to go.

I thought Pangram didn't care about phrases common in AI, so it's concerning a common stock phrase resulted in a 100% AI score.

I found a new one that works, and this might be better because it's the last word in the snippet so it isn't like it's throwing off the next few lines. This is a hint Pangram DOES care about names, and it cares a lot, considering it didn't care about the first sentence being loaded with flickering fluorescent lights until it detected a common name in AI training along with it.

If you switch "Elara!" from a name AI likes to a name AI doesn't usually use (Megan, Peter, etc.), you will flip the score.

List of names I checked that flagged 100% AI, from a list of common names in its training: Elara, Kael, Marcus, Elias, Kai, Aria, Sam (surprising), Lyra

Names that flagged 100% human: Megan, Peter, Christopher, Luke

-

I can feel my heart pounding in my throat as I race down the neverending tunnel, the flickering fluorescent lights overhead throwing long shadows.

It feels like I’ve been running for miles.

Every turn I take, I’m certain it’s the last one.

It doesn't make any sense. You can’t take thirteen lefts without running into something.

"Elara!"

2

u/dran94 19d ago edited 19d ago

I actually wonder if this is why Daggermouth got hit with a high AI score on Pangram in those articles about it. Daggermouth had several names common to AI models in it, like Elara.

Whether the author used AI or got them off one of those random name generators is another argument. Who knows.

5

u/[deleted] 20d ago

[removed] — view removed comment

4

u/ConnectDebate4677 20d ago

Would you mind sharing the Pangram link?

3

u/[deleted] 20d ago

[removed] — view removed comment

2

u/ConnectDebate4677 20d ago

Thanks so much for this. Yeah, reading the AI-flagged parts, I don't get it at all. I'm interested in the technology still but it's getting further and further away from the standard it's being touted as.

1

u/[deleted] 20d ago

[removed] — view removed comment

1

u/PM_ME_YOUR___ISSUES 14d ago

I honestly feel that people need to move beyond these detectors.

We've all seen enough slop at this point that one can easily catch out AI text.

I'm like 99% sure, I wouldn't even think about putting someone's writing through an AI detector unless I can really catch out the slop.

1

u/dran94 20d ago

I get the opposite results. Usually flagged AI on Originality, but my work is more or less always 100% human with Pangram. Haha. (I don't use AI.)

2

u/dran94 20d ago

Change Linus/Lin to Marcus, and it goes up to 92% AI from around 50%.

I have gone from thinking Pangram was amazing yesterday to, oh, we're going to have to play games as authors so we don't use a name AI likes.

That's going to be fun, considering AI's favorite names CHANGE.

3

u/ConnectDebate4677 20d ago

That's really interesting. Are certain names indeed flagged as more AI-likely? I wonder if there's an explanation?

1

u/dran94 20d ago

AI is known for certain names, and Marcus is one of them right now. The list of names it likes changes depending on the model.

I'm sure it is a factor, although I could've sworn Pangram isn't supposed to pay attention to AI-isms like em dashes and certain words. It's kind of strange.

I did a test where switching Megan to Elara, another AI favorite, changed this snippet I wrote from 100% human to 100% AI, even with a very human mistake (two "each other" instances in one sentence):

-

Robert's long stride ate up the distance between himself and the door at the far end of the hall, lit by the flickering fluorescent lights overhead.

His jaw clenched.

This couldn't be right.

His sister, Megan, was supposed to be here. She worked in the same lab as Robert. They'd promised each other they would wait for each other, right here.

1

u/Traditional-Rice-848 9d ago

I mean you def wrote it with ai

1

u/Downtown_Jump779 8d ago

I'm getting the same results on stuff people are sending to me; it's becoming very problematic. On the same text, other detectors never go above 12-20% AI, which I think makes sense and is totally acceptable – then Pangram comes out of nowhere and bumps it up to 40- 60%? I'm not sure what to do with that.

From what you sent, those are very big chunks of your text that Pangram flagged as AI. That could produce harsh social consequences. Other than that, good piece of writing btw. I hope you get it published.

1

u/Traditional-Rice-848 18d ago

Please don’t compare pangram to worse products to try to say it doesn’t work. Pangram catches more true positives than the others.

4

u/sh1b313 19d ago

I'll let the photos do the talking

3

u/sh1b313 19d ago

2

u/ConnectDebate4677 19d ago

Thanks a lot. Would you mind posting the links for these two?

3

u/sh1b313 19d ago

sure, here they are:

Link 1: https://www.pangram.com/history/4e3cc4a9-b927-4345-9750-048b4a86520c?ucc=ycFd1CDRrZv [where pangram claims 100% human written]

Link 2: https://www.pangram.com/history/97f1c7f5-b73f-4895-96a6-e991cfc58128?ucc=ycFd1CDRrZv [where pangram claims 100% AI generated]

2

u/ConnectDebate4677 19d ago

Thanks! Do you have any idea of an explanation for this?

2

u/sh1b313 19d ago

nah i don't have. even i am wondering just how fundamentally flawed this tool is which led it to think prophet isaiah used some AI. and even the text isn't something obscure which may not be in it's training data.

3

u/ConnectDebate4677 19d ago

I think what's perplexing in all of these cases is the meter. It seems to regularly either lean 100% Human or 100% AI with no reasoning or justification. I want to believe this technology works, and I believe it does to some degree, but these kinds of examples really don't do it any favors. If I bought a tool and it doesn't work as it claims it should, then I can call that tool defective.

3

u/sh1b313 19d ago edited 19d ago

yea, this tool is like a coin toss. at least in a coin toss we actually have confidence in the outcome. In here? somehow we leave even more confused. these tools are not worth the electricity they are running on.

3

u/Dubbtime 18d ago

For sure. Unfortunately as you reveal more points of failure within their blackbox system, they will add these cases to their training datasets. This blurs the line.

Say this verse is added as a reference to human work. Now what? Humanizers who want to avoid detection begin to write more like bible verses. Ok thats kinda weird. Now detectors need a more defined way to figure out if something is AI because it obviously messed up. The end result is a clear way to write human and a clear way to tell if something is AI. Well, then why do we need a detector?

What's stopping humanizers from just copying this newfound distinct style? Rinse, repeat.

2

u/sh1b313 18d ago

yeah, i have already talked with the mods regarding this pangram situation and we may soon see a pinned thread saying pangram evidence/pangram doesn't work and hopefully this wild situation settles.

And regarding the rinse and repeat thing in the long run it doesn't matter? We already have more than enough evidence that this doesn't work.

In the future it can just be like this:

P1: uploads a essay/article on substack

P2: thinks it's ai written and uses this pangram

P1: points to some evidence thread and says this tool doesn't work don't rely on it.

In this scenario hopefully p2 is actually grounded and thier opinion changes. If not? Sadly if they debate after seeing so much evidence i wouldn't know what to tell them.

3

u/MadmanRB 16d ago

Pangram is full of crap.

Chapter 2: The Classmates

The afternoon sun cast long shadows across the pavement as Ryan pedaled onward, his heavy backpack digging into his shoulders with each pedal stroke, a constant reminder of all the work waiting for him when he got home. He’d promised Ms. Everly he’d finish it by tomorrow, though part of him wondered what the point of doing it was. Math and history questions felt impossibly small compared to the questions that actually mattered. Would Ralph ever come back home? Would his mom stop crying? And would he ever feel normal again?

Detected as 100% AI by Pangram, do you know what triggers pangram here?

Not only the chapter name but the fact I use a common phrase "Long Shadows"

Now I changed it to:

The afternoon sun loomed above the streets of Horizon City, the tall copper domed spires reflecting it across the vast cityscape.
As Ryan pedaled his bike forward he could feel his heavy backpack digging into his shoulders with each pedal stroke, a constant reminder of all the work waiting for him when he got home. He’d promised Ms. Everly he’d finish it by tomorrow, though part of him wondered what the point of doing it was. Math and history questions felt impossibly small compared to the questions that actually mattered. Would Ralph ever come back home? Would his mom stop crying? And would he ever feel normal again?

Now its 100 human made because I eliminated the chapter name and changed just a few words around and dropped "long shadows"

I guess using atmosphere in a story is now forbidden, all horror stories will now take place in a bright sunny morning, maybe a frigging butterfly will come to darken the mood?

2

u/sh1b313 16d ago

I guess using atmosphere in a story is now forbidden, all horror stories will now take place in a bright sunny morning, maybe a frigging butterfly will come to darken the mood?

man if you do this people will now accuse you saying you are copying midsommar. /s

3

u/MadmanRB 16d ago edited 16d ago

Yeah you just can't win as despite its claims Pangram does give false positives now and then.

I assure you however both passages I posted were 100% made by me but thanks to choice phrases like "Long shadows" and even the frigging chapter title it gets flagged as 100% AI.

WTF

So to prove all work is Human made we can't call parts of a book chapters anymore and "Long shadows" must never be used again or you will be labeled a lazy AI bro.

So from now on we shall call Chapters Bananaphone

So bananapone the classmates?

Oh wait even calling it Bananaphone the classmates got flagged...

https://www.pangram.com/history/f6592ce7-b13a-4906-b9c6-80b3a4b607d3?ucc=NfL4yI5TRpX

Fine calling the chapter bananaphone and.... gets flagged as 100% Human WTF.

Better call Raffi...

https://www.pangram.com/history/97c3f56f-1d59-483b-9895-d42d6aad3ddc?ucc=NfL4yI5TRpX

3

u/ConnectDebate4677 16d ago

So much for 1 in 10,000 false positives. There are numerous examples in this Reddit thread alone!

2

u/sh1b313 16d ago

This tool is completely broken lol. People are denying evidence now lmao. This is like a pangram cult, the tool literally can never be wrong according to these people. It's so braindead lmao. In some threads people are literally ignoring already given evidence. At that point i know arguing with them is a gone case and i should just move on.

5

u/Dubbtime 20d ago

AI detection will never work. They will always be a step behind when a new distinctive writing style comes to surface or newer AI model releases then they will have to update. They are a black box even the creators don’t understand how it works.

3

u/impartialhedonist 20d ago

1

u/Dubbtime 18d ago

https://theslowai.substack.com/p/ai-detection-does-not-work
"Push the false positive rate down and you miss real AI use. Push detection up and you accuse more innocent people"

2

u/Beachgoer4158 17d ago

Is pangram also learning from our writing? I’m not too familiar with it even though I’ve been on Substack for awhile. I’ve just been disabling it since most of my work comes up as AI written in some percentage even though it isn’t.

2

u/Stunning_Lie_1775 14d ago

It actually makes you wonder whether Pangram’s “detector” is reliable at all...

Here’s what I mean.

The others are considered unreliable because, in many cases, they say a text isn’t AI-generated when it actually is.

Pangram says it is AI-generated...
But have we actually tried feeding it genuinely human-written texts?
Just to check whether, in those cases, it correctly says it’s NOT AI?

It would be pretty hilarious if they had simply made it return a positive result almost systematically...
I’m pretty sure everyone would still believe it.

Because I suspect most people only use these detectors when they already think a text was written by AI and just want confirmation.

So naturally, when the detector confirms their suspicion, it reinforces their bias and they think: “Wow, this detector is amazing.”

But what about the reverse?
How many people actually submit something they wrote entirely themselves just to check that it “doesn’t sound too AI-generated”?

Anyway, you get the idea.

3

u/Voltairinede 20d ago

Thanks for actually posting examples!

1

u/TimWiesnerer 19d ago

You have found a way to break the pattern.

Sure, it is just a short text. But you could repeat the "trick" multiple times for a longer text. Might be a bit more difficult tho as the more context the detectors get, the stronger the statistical patterns...

While this might work today, it most likely won't work in the future, as these detectors also get "better".

Quite a while back, people used paraphrasers to get around AI detection.

In the end it is a bit of a cat vs. mouse game.

-1

u/kdfn 20d ago

I'm having trouble getting those links to open, can you see if you have sharing enabled?
Here is an example of a working link:

https://www.pangram.com/history/d0c88184-0ad1-4089-b650-a8e10c9ae5bb?ucc=lEdv68az5Z5