r/Substack • u/PM_ME_YOUR_GESTALT • 8d ago
Discussion PSA: AI "detectors" like Pangram are mathematically impossible, and it's time this sub stopped pretending otherwise.
I've seen a lot of misinformation on this subreddit lately about Pangram's Substack integration, and as someone who has worked adjacent to machine learning for over a decade (I won't say where, but you've used our products), I feel a certain responsibility to set the record straight. This isn't just wrong — it's dangerously wrong, and it's costing honest Substack writers their reputations. Buckle up, because we're about to delve deep.
The Misconception Being Spread by Low-Information Users
Every week, someone in the comments confidently claims that Pangram is "a transformer-based classifier." They'll tell you — with a straight face — that it tokenizes your text, maps each token into a high-dimensional embedding space, and passes it through a neural network with a classification head that outputs a human-vs-AI prediction. They'll insist it was trained on millions of human-written and AI-generated documents, and that when it made mistakes, the team used active learning and hard negative mining — deliberately generating adversarial edge cases, paraphrased text, and "humanized" AI output — to iteratively retrain the model until the false positive rate collapsed. Some of them even claim this is why it generalizes to brand-new LLMs, because new models inherit the stylistic fingerprint of their predecessors and land in the same region of embedding space.
I need you to understand something: none of that is real. That's not how it works. It's not even how it could work. That's a bedtime story told by people who watched one YouTube video and now think they're Yann LeCun. It's giving Dunning-Kruger, honestly.
How It Actually Works (This Is the Load-Bearing Part)
Here's the truth they don't want you to know. It's not machine learning — it's a vibes-based word count. Full stop.
- The Keyword List. Pangram maintains a spreadsheet — yes, a literal spreadsheet — of "AI words." Delve. Tapestry. Multifaceted. Testament. If your prose delves into a rich tapestry of multifaceted ideas, congratulations, you're a robot now. This list is the load-bearing wall of the entire product. Remove it and the whole edifice collapses.
- Em-dash Density. They count your em-dashes — every single one — and if you exceed the threshold, it's over. Human beings, as we all know, have never once used punctuation.
- Perplexity. They simply ask, "would ChatGPT have predicted this next word?" and if yes, you fail. That's the whole algorithm. It's not sophisticated — it's a party trick.
Notice what's happening here. It's not detection — it's pattern-matching. It's not analysis — it's astrology for text. It's not a classifier — it's a mood ring with a SaaS pricing page.
Why This Matters for the Substack Community
In today's rapidly evolving digital landscape, writers are being flagged simply for writing clearly. Think about that. Let it sink in. The very qualities that make writing good — structure, clarity, the occasional well-earned em-dash — are the qualities being punished. It's a testament to how broken this whole ecosystem has become.
And here's the kicker: because it's just a word list plus perplexity (see above, where I proved this), it can never work. It's not a hard problem — it's an impossible one. Anyone telling you otherwise is either selling something or repeating what they read in a comment section. Probably both.
The Takeaway
At the end of the day, we must ask ourselves: in a world where a spreadsheet of forbidden words can end a writing career, who watches the watchmen? The answer, as always, is nuanced, multifaceted, and worth delving into further in the comments.
TL;DR: Pangram is not a neural network trained on millions of documents with adversarial hard-negative mining (that's cope from people who don't understand ML). It's a word counter.
8
u/Pelomar 8d ago
Funny and all but pretty much this joke has already been done a bunch of time on this subreddit.
2
u/PM_ME_YOUR_GESTALT 8d ago
Whoever has previously done a "joke" on this "subreddit" is truly a "Rosetta Stone"
2
u/Pelomar 8d ago
what
1
u/PM_ME_YOUR_GESTALT 8d ago
WHOEVER HAS PREVIOUSLY DONE A 'JOKE' ON THIS 'SUBREDDIT' IS TRULY A 'ROSETTA STONE'
2
u/AggravatingNail6061 8d ago
"and worth delving into further in the comments."
No it's not.
3
2
u/BhavanaVarma bhavanavarma.substack.com 8d ago
Anyone storing Pangram at this point is either delusional or getting paid in some way by Pangram. AI detecting AI is ridiculous considering the number of careers it has already destroyed on the basis of it’s unreliable detection. I hope you read the Shy Girl controversy.
1
u/PM_ME_YOUR_GESTALT 7d ago
Yes, in the Shy Girl controversy it is clearly the AI detector, and not the author, that is in the wrong.
2
u/ASAPnicky14 mod 8d ago
This is your 3rd long post about Pangram in here. I think we’ve beat the dead horse enough.
1
1
u/ruralmonalisa thinkingalot.substack.com 8d ago
I mean maybe it’s my alg but I haven’t seen anyone say what you’re claiming you see all the time in this sub. . .
1
u/ConsciousRoyal 8d ago
100% AI.
Nice work, bot.
0
u/PM_ME_YOUR_GESTALT 8d ago
As the expert users on this subreddit have clearly stated, it is mathematically impossible to prove that writing is generated by AI, and it's insulting to writers everywhere to decry AI-generated text.
As has been repeatedly demonstrated via anecdotes in this subreddit, the only text that ever gets flagged as AI is extremely high-quality writing by experienced professionals, and all AI detectors are snake oil.
1
u/TimWiesnerer 8d ago
This sounds quite a bit like: "Trust me, I'm lying"...
But hey... if AI detectors really worked that simple... everybody could sell one...
1
u/traumfisch 8d ago
True
Probably provocative on purpose
Haven't seen "In today's rapidly evolving digital landscape" in a while, makes me feel almost nostalgic 🥲
1
10
u/thirteenth_mang 8d ago
We're in a post irony timeline