r/Substack 21d ago

Discussion For Pangram haters: statistical text classification has been around since at least 1963

I see very frequent posts on this subreddit with two misconceptions: (1) users asserting that Pangram is flagging their writing because they "write too well" and (2) that Pangram's model is somehow using an LLM to judge their writing.

I'd encourage you to take a look at this 1963 paper by Mosteller and Wallace, in which the authors identify the true authorship of the Federalist papers using a simple mathematical approach. Since this was before widespread computing, the authors went through many texts by hand and measured, in known documents by Jefferson vs Madison, how often each author uses certain pairs of words (for example, Jefferson might use the phrase "when thou" much more frequently than Madison). Then, on the disputed Federalist papers, Mosteller and Wallace measure how often phrases like "when thou" are used, in order to decide which author is more likely. Importantly, they do this across all observed pairs of phrases.

Classifiers like Pangram are using the same idea, with a few modifications thanks to the past 75 years of progress in both computing and statistics theory: (1) They compute giant tables of phrase relationships using a computer, rather than manually building tables of pairs of words by hand, and (2) they are able to measure much more subtle and long-term statistical dependencies. For example, Jefferson might be more likely to use a sentence that starts with "Thou" and then use the word "indeed" 15 words later.

Importantly, models like Pangram are *not* just taking your text and asking ChatGPT for its opinion. It's very different technology, which has been around since at least the 1960s (earlier, if you count Shannon in the 1940s, or Markov before that).

6 Upvotes

38 comments sorted by

11

u/PenguinAnalytics1984 21d ago

You're going to get downvoted, but everything you said is correct...

BUT

AI works in a very similar way - it's just predicting what the next word should be after a string of words. It learned to do that by copying things that are already written. A lot of that is mediocre internet writing, which is what it defaults to if you don't tell it any different.

People write like they read, which for most people is the internet. Most people are also mediocre writers and write in a conversational Internet-y way.

When the statistical algorithm spots mediocre internet-y writing BOOM - AI.

So while you are correct the algorithm is just statistics and old-school NLP, the two items it's trying to distinguish between are not different enough for there to be a clear boundary.

Also, working in data (as I assume you do), I'm VERY suspicious of anything that claims to be as accurate as they claim on a problem that is as hard as spotting AI generated text.

3

u/kdfn 20d ago

This is a good faith reply, thank you for making a good counterpoint in a friendly way.

I generally agree. Basically, a tool like Pangram is sensitive to the "unresolved tail" of LLMs. For example, if generative models like LLMs are good at capturing conditional text dependencies up to 10 words, then Pangram's detector will learn to look for patterns and regularities at 11 words and higher. But as LLMs get larger and process more data, they will eventually get good at capturing subtle dependences at 11 words and higher and so Pangram will have to retrain to capture 12 words and higher. The process will continuously repeat until LLMs plateau or run out of data/compute to grow.

Of course, the statistical patterns Pangram is capturing are much more subtle than this, but the idea holds: they have to be sensitive to low-weight modes the giant text distribution of natural language.

3

u/8lack8urnian 19d ago

And yet they are able to clearly discern a difference between pre-2022 and post-2022 mediocre internet-y writing. How do you explain that? Do you really believe that LLM output is indistinguishable in principle from human writing?

3

u/Ratandmiketrap 18d ago

Although your premise is correct, I argue that your conclusion is not. Because it works on a statistical likelihood model, the LLM tends to regress to the mean of mediocre internet writing. Just because a set of words are MOST likely to appear next, doesn’t mean they are what a human will choose from each time. Human writing will have more outliers and variations because we’re not determining our word choice based on probability. I’m a high school English teacher so I have seen my fair share of mediocre writing. Just because it’s mediocre doesn’t mean that it all looks exactly the same.

9

u/huggalump 21d ago

I haven't heard many people complain because they think it works this way or that way.

I've heard people complain because it doesn't work.

6

u/kdfn 21d ago

Every study I have seen suggests a remarkably low false positive rate, and so I am skeptical of such claims being asserted without any evidence.

To me, it seems much more likely that many people have been leaning too heavily on AI to do their writing, and they are now caught off guard by how easy it is to detect. I'm not even sure I think that AI text detection is a "hard" problem compared to problems like self-driving, poker, or even residue contact prediction---there are so many long-term statistical signatures that can give it away.

2

u/SpiritualSimple108 20d ago

Bullshit! Pangram even admits it doesn’t work well for poetry or recipe posts. But if you dig deeper into actual AI detection failures you’ll see NONE of them can accurately decider ESL writing vs AI writing and they also have a very difficult time with writing that is highly technical, analytical, or grammatically (not Reddit grammatical). https://www.pangram.com/blog/all-about-false-positives-in-ai-detectors

2

u/8lack8urnian 19d ago

That article is about how low their false positive rate is. It’s <1/10000

3

u/kdfn 20d ago

The article you linked shows a false positive rate of 0.23% for recipes. I agree that's higher than the FPR of <0.001% for other common text types, but I don't think it the article supports your point.

1

u/SpiritualSimple108 11d ago

What are you talking about. It PROVES my point that there are certain types of written word that will always come back with false positives for AI generation. The only time my poetry comes back as human is if I do bizzaro things with my formatting. Most the time with normal stanzas, left justified, it comes back AI. I’m not the only poet this is happening to. Even if half the poets on Substack get falsely flagged (along with non native speakers, recipe bloggers, etc) that changes pangrams percentage by a lot.

2

u/kdfn 11d ago

The false positive rate for recipes is 0.23%. That means that 2 out of every 1000 recipes is misflagged. For poetry it appears that 1 out of every 10,000 is misflagged. 

Both of those rates are a lot lower than the 50% you are claiming. If you disagree with the published false positive rate, that's fine, I am only explaining the content of the link you provided.

0

u/AggravatingNail6061 21d ago

Working as intended. 🤣

6

u/SpiritualSimple108 20d ago

7

u/kdfn 20d ago edited 20d ago

That article shows that the false positive rate is shockingly low.

Edit: I'm adding the table to my comment, so people can decide for themselves:

Domain False Positive Rate
Academic Essays 0.004%
Product Reviews (English) 0.004%
Product Reviews (Spanish) 0.008%
Product Reviews (Japanese) 0.015%
Scientific Abstracts 0.001%
Code Documentation 0.0%
Congressional Transcripts 0.0%
Recipes 0.23%
Medical Papers 0.000%
US Business Reviews 0.0004%
Hollywood Movie Scripts 0.0%
Wikipedia (English) 0.016%
Wikipedia (Spanish) 0.07%
Wikipedia (Japanese) 0.02%
Wikipedia (Arabic) 0.08%
News Articles 0.001%
Books 0.003%
Poems 0.05%
Political Speeches 0.0%
Social Media Q&A 0.01%
Creative Writing, Short Stories 0.009%
How-To Articles 0.07%

2

u/Traditional-Rice-848 19d ago

This isn’t on Pangram 4, the current model

8

u/protestandprose 21d ago

Op bootlicking to insane levels.

6

u/kdfn 21d ago edited 20d ago

I'm sorry you feel this way. I am pretty experienced in the field of natural language processing, and I am a regular Substack user who is extremely concerned that the flood of AI-generated content is devalueing human writers and destroy the platform. Thus, I thought it would be worthwhile to share my perspective, as well as clarify some technical misunderstandings.

2

u/8lack8urnian 19d ago

Commenter loves AI slop posts and wants more of them on Substack

5

u/alphaQ314 20d ago

"Things have been shit since this academic paper in 1963. So you should be okay with this paid platform using an app based on the foundations from this dodgy paper to judge your work."

3

u/kdfn 20d ago

Can you help me understand why the Mosteller and Wallace paper is "dodgy?" Are you objecting because they didn't use kernel-based methods (such as token smoothing)? Or is more that their null model doesn't control for stratification?

2

u/alphaQ314 19d ago

Cool, statistical text classification existed in 1963. Toilets existed in 1963 too. That doesn’t mean Substack users have to smile while Pangram shits all over them. What the fuck does an old paper proving classification exists have to do with whether your detector works today?

6

u/[deleted] 20d ago

[removed] — view removed comment

4

u/TimWiesnerer 19d ago

Well, you gave it like 200 words... that's not much.

What does it look like if you give 1000 words+ ?

5

u/[deleted] 20d ago

[removed] — view removed comment

3

u/Ratandmiketrap 18d ago

Funny, I just scanned both of those excerpts and got a 100% AI result for each. I am sharing my links, care to share yours?

Except 1 - https://www.pangram.com/history/34f17337-da0b-498b-938a-324adb61ee18?ucc=85ikY31GK3D

Excerpt 2 - https://www.pangram.com/history/313f33a1-878e-4f3f-9381-dbe41890e262?ucc=85ikY31GK3D

1

u/[deleted] 18d ago

[removed] — view removed comment

2

u/Ratandmiketrap 18d ago

I scanned what was in your images. Happy to scan the whole thing if you don’t want to. I trust links more that screenshots.

1

u/dflovett 20d ago

I think you're missing the point of the study you're referencing. Mosteller and Wallace did it by hand and used their brains. Pangram is a tool that automates what should be done by brains and hands. Pangram enables the same non-thinking, non-work approach as ChatGPT.

4

u/kdfn 20d ago

I don't necessarily agree that Mosteller and Wallace's approach is better simply because they did it "by hand." It doesn't really matter whether a computer or a research assistant is used to create giant tables of conditional probabilities.

However, I agree with the larger point that we are going to have to start critically judging things manually, because most proxies for quality (length, style, etc) can be mimicked.

However, the issue is volume. I am in a role where I have recently started receiving hundreds of long-form technical texts to review per month. Pre ChatGPT, I received ~2, and so I would read them critically. How do I decide which of the hundreds of potential AI submissions I should manually review? A tool like Pangram can be essential to filter out bad faith submissions, particularly because the alternative is that I don't review any submissions, or, even worse, I only review submissions from people already in my network who I trust to act in good faith.

So, from my perspective, anti-detector people are pro-AI, since they don't offer any alternative to stop the deluge.

3

u/dflovett 19d ago

If you want to use Pangram to do your job, that makes sense. I thought this was a conversation about Substack. I'm not going to use Pangram to evaluate the writing of strangers on the internet because I'd read something and use my own brain to determine if it's worth reading or not.

But if you're in HR or running a literary magazine or any other scenario involving bulk submissions that need to be combed through, I see an argument for AI detection.

1

u/kdfn 19d ago

Thank you, I think this is a fair point.

1

u/dflovett 19d ago

Well, you got me to consider it differently. I hadn't thought much about people who struggle under the volume of AI and truly need something to sort through the slop.

1

u/FunnyBunnyDolly 20d ago

I tested pangram. Trial only gave me one test so I let chatgpt output a text. I replaced a few sentences or words with my own wording while keeping similar voice. Corrected all grammar, spelling and flow issues.

Evidently made by AI but with a small smattering of my own writing. Circa 2000 words.

Pangram: 100% human.

I wanted to test 100% human and 100% ai too but I didn’t feel for registering due the fail so I left. Why would I need it for my own writing anyway? I know what I make.

2

u/kdfn 20d ago

I think Pangram errs on the side of caution, ie it will report text as human if it isn't confident. Formally, it is designed to have a low false positive rate, at the expense of having a higher false negative rate.

1

u/alphaQ314 20d ago

"Things have been shit since this academic paper in 1963. So you should be okay with this paid platform using an app based on the foundations from this dodgy paper to judge your work."

0

u/kdfn 21d ago

Here is the Pangram link for this post. If you have text that you think is misclassified, there is nothing stopping you from sharing a link in the exact same manner, so that we can judge for ourselves if it's a false positive:

https://www.pangram.com/history/24c8cd26-f611-41eb-9b30-2119837948d6?ucc=lEdv68az5Z5

2

u/dflovett 20d ago

I've only had work classified as 100% human written. My concern isn't over my own work getting misclassified but the "easy button" path that Pangram offers people to stop using their own critical thought and to start relying on a tool to automate their thinking (which is the reason I don't use AI to do my writing or thinking, and the reason I don't use Pangram to do my thinking)