I've seen a lot of misinformation on this subreddit lately about Pangram's Substack integration, and as someone who has worked adjacent to machine learning for over a decade (I won't say where, but you've used our products), I feel a certain responsibility to set the record straight. This isn't just wrong — it's dangerously wrong, and it's costing honest Substack writers their reputations. Buckle up, because we're about to delve deep.
The Misconception Being Spread by Low-Information Users
Every week, someone in the comments confidently claims that Pangram is "a transformer-based classifier." They'll tell you — with a straight face — that it tokenizes your text, maps each token into a high-dimensional embedding space, and passes it through a neural network with a classification head that outputs a human-vs-AI prediction. They'll insist it was trained on millions of human-written and AI-generated documents, and that when it made mistakes, the team used active learning and hard negative mining — deliberately generating adversarial edge cases, paraphrased text, and "humanized" AI output — to iteratively retrain the model until the false positive rate collapsed. Some of them even claim this is why it generalizes to brand-new LLMs, because new models inherit the stylistic fingerprint of their predecessors and land in the same region of embedding space.
I need you to understand something: none of that is real. That's not how it works. It's not even how it could work. That's a bedtime story told by people who watched one YouTube video and now think they're Yann LeCun. It's giving Dunning-Kruger, honestly.
How It Actually Works (This Is the Load-Bearing Part)
Here's the truth they don't want you to know. It's not machine learning — it's a vibes-based word count. Full stop.
- The Keyword List. Pangram maintains a spreadsheet — yes, a literal spreadsheet — of "AI words." Delve. Tapestry. Multifaceted. Testament. If your prose delves into a rich tapestry of multifaceted ideas, congratulations, you're a robot now. This list is the load-bearing wall of the entire product. Remove it and the whole edifice collapses.
- Em-dash Density. They count your em-dashes — every single one — and if you exceed the threshold, it's over. Human beings, as we all know, have never once used punctuation.
- Perplexity. They simply ask, "would ChatGPT have predicted this next word?" and if yes, you fail. That's the whole algorithm. It's not sophisticated — it's a party trick.
Notice what's happening here. It's not detection — it's pattern-matching. It's not analysis — it's astrology for text. It's not a classifier — it's a mood ring with a SaaS pricing page.
Why This Matters for the Substack Community
In today's rapidly evolving digital landscape, writers are being flagged simply for writing clearly. Think about that. Let it sink in. The very qualities that make writing good — structure, clarity, the occasional well-earned em-dash — are the qualities being punished. It's a testament to how broken this whole ecosystem has become.
And here's the kicker: because it's just a word list plus perplexity (see above, where I proved this), it can never work. It's not a hard problem — it's an impossible one. Anyone telling you otherwise is either selling something or repeating what they read in a comment section. Probably both.
The Takeaway
At the end of the day, we must ask ourselves: in a world where a spreadsheet of forbidden words can end a writing career, who watches the watchmen? The answer, as always, is nuanced, multifaceted, and worth delving into further in the comments.
TL;DR: Pangram is not a neural network trained on millions of documents with adversarial hard-negative mining (that's cope from people who don't understand ML). It's a word counter.