Is Pangram's AI detector accurate? It flagged my own writing 100% AI
On October 2, 2026, I decided to test Pangram's AI detector on something pretty simple: short notes I had written myself. Pangram said two of them were 100% AI.
- Product
- Pangram
- Made by
- Pangram Labs
- What it is
- AI text detector
- What I used
- Pangram 4.0 on my own account, October 2, 2026
- Website
- pangram.com
And this is exactly where I have a problem with AI detectors. People see a number like 100% and treat it like proof. But when that number is wrong, an actual person is left having to prove they wrote their own words.
I tested my own writing
I ran all of these tests through Pangram 4.0, using my own account, on October 2, 2026.
At 12:09 PM, I scanned a short note I wrote myself. It starts, "Hey, just a heads up on why I shortened the email."
Pangram counted 53 words and labeled it "AI Generated." The result said "100% of this text is AI" and underneath that, "We believe that this entire text is AI."
Except I wrote it.
At 12:24 PM, I scanned the exact same note again. Same result. 100% AI.
So I tried another one.
At 12:27 PM, I scanned a second version I had written myself. This one started "Hey Lenny," and was 58 words.
Pangram called that one 100% AI too.
Now, to be fair, Pangram did not get everything wrong.
At 12:21 PM, I tested a different 54-word note I wrote that started "Hey Grooming Team," and that one came back "100% Human Written."
Each result also included a small note saying confidence was limited because the text was short. But the big gauge still said 100%.
I used Pangram's feedback box to tell them the writing was human and asked them to fix it.
Pangram does warn about short writing
I think this part matters because I'm not interested in making Pangram sound like it claims something it doesn't.
Pangram's model card from July 29, 2026 says Pangram 4 is "intended for long-form natural-language writing samples of at least 50 words" and lists short conversational replies as outside its primary scope.
Pangram's CEO, Max Spero, has also talked publicly about this problem.
On May 15, 2026, Max Spero wrote on X about a false flag involving an excerpt from Taylor Lorenz's story. The text was "just above our minimum of 50 words," Spero wrote, and "+1 sentence or -1 sentence both come back as human."
Then, in The Atlantic on May 30, 2026, Spero said Pangram should "never be the ending arbiter."
Pangram's own model card says "False accusations of AI usage can lead to serious consequences" and "our model has a non-zero error rate."
I agree.
Pangram's launch post for Pangram 4, also from July 29, 2026, claims a false positive rate of 0.0041%, roughly one in every 24,000 documents, on internal benchmarks. That post doesn't say how it performs on short texts.
So I'm not saying Pangram claims perfection. It doesn't.
My point is that the screen still said 100%.
Most people aren't going to stop and read a model card before deciding what that number means. They're going to see 100% and think the software knows.
It isn't just my test
The problem gets much bigger when a detector score leaves the screen and becomes an accusation against an actual person.
The Atlantic reported on May 30, 2026 that Taylor Lorenz had been accused on X of using AI to write a Vanity Fair story. Spero investigated and found Pangram had erred.
The same Atlantic story noted that a University of Chicago study that found almost no false positives tested samples of roughly 500 to 1,000 words. It also pointed out that with more than 10 million high schoolers and 20 million undergraduates in the US, even one in 10,000 means "plenty of false accusations."
The Atlantic also noted that Pangram can't point to much specific evidence for why it decides writing is AI or human.
In July 2026, Substack also added Pangram scans for posts, and 404 Media reported on July 28, 2026 that writers pushed back. "All it takes is one false accusation," writer Alice Lemee told 404 Media.
And that is exactly my concern. An AI detector false positive isn't just a technical mistake once somebody starts using it as evidence against a person.
How reliable are AI detectors?
The student cases are even more concerning, but I want to be really clear about something here. These examples are about AI detectors in general, and most involve Turnitin, not Pangram.
Spectrum News reported on May 15, 2025 that Kelsey Auman, a student at the University at Buffalo, had several assignments flagged by Turnitin's AI detector, putting Auman's graduation at risk. Auman was cleared after showing browser history and research, and later started a petition to turn the detector off.
Auman's conclusion was pretty direct: "the numbers don't mean anything."
Bloomberg Businessweek reported on October 18, 2024 that Moira Olmsted, an autistic student at Central Methodist University who writes in a formulaic style, received a zero on a reading summary after an AI detector flagged it. The grade was later changed, with a warning attached.
Bloomberg also ran its own test of AI detectors and found 1% to 2% of essays were falsely flagged, "in some cases claiming to have near 100% certainty." The same story reported that about two-thirds of teachers regularly use AI detectors.
ABC News Australia reported on October 9, 2025 that Australian Catholic University registered nearly 6,000 academic misconduct cases in 2024, about 90% of them AI-related, with many relying on Turnitin's detector. About a quarter were dismissed.
Madeleine, a 22-year-old nursing student, waited six months to be cleared. During that time, Madeleine applied for graduate jobs with a transcript that said "results withheld."
The university dropped the detector in March 2025.
Turnitin's own guidance says its score "should not be used as the sole basis for adverse actions" against a student.
To be fair, the university said the figures were "substantially overstated." The ABC also reported emails showing the university did rely on the detector report alone.
Vanderbilt University had already decided to turn off Turnitin's AI detector on August 16, 2023. It explained that a 1% false positive rate would have meant "around 750 student papers" incorrectly labeled out of the 75,000 papers it submitted to Turnitin in 2022.
And the debate is still going.
The Atlantic reported on September 21, 2026 that an MIT working group warned about an "atmosphere of distrust between instructors and students." Indiana University's Kelley School of Business barred professors from using AI detectors, and a few students have sued universities that accused them partly on the basis of AI detection.
That same Atlantic article noted that Pangram "occasionally errs on shorter snippets."
A 2023 Stanford study in Patterns also found that seven AI detectors flagged 61.22% of essays by non-native English writers as AI, on average. To be fair, that study came before Pangram launched in October 2023, and Pangram reports 0% false positives on that same set of essays.
So again, that study is not about Pangram. But it does show why being falsely accused of using AI isn't some abstract problem.
How to prove I didn't use AI
The cases above also tell us something practical.
Kelsey Auman was cleared after showing browser history and research. Taylor Lorenz credited edit history in an interview with The Atlantic.
So keep your drafts, version history and notes.
I don't think people should have to build a defense file for every paragraph they write. But when a system says your own work is AI, having evidence that shows how the work developed can matter.
On its own site, Pangram offers "LMS integrations for organizations looking to uphold academic integrity" among its products. LMS means the systems schools use for classes and grades.
And that makes the human part of this even more important.
A detector should be a signal, not proof
I'm not saying Pangram is useless. Tools like this may help find AI-manufactured writing.
As more schools, editors and platforms adopt this one, more people are going to put 100% trust in it.
There is also some irony here. Pangram's model card says it is built on "a causal, open-weight sparse mixture-of-experts language model."
So Pangram uses an AI language model to decide whether a person used AI.
AI flagging AI becomes a problem when no human checks the call.
And I think we may be asking the wrong question anyway.
Is the content useful?
Is it true?
Is it authentic?
Then there is the question nobody seems quite as comfortable asking. If a journalist takes their own notes, writes 85% of a piece, and uses AI to clean it up and catch typos, is that wrong?
Fortune ran a piece on September 29, 2026 arguing that instead of asking writers whether they used AI, we should ask "how much."
I think that is a much more useful conversation than treating a detector score like a verdict.
Because once a system puts 100% on the screen, it is very easy to forget the small print saying the system can be wrong.
I'm curious to see what other people think.
Sources
Product facts checked October 2, 2026. Every product fact above links to the maker's own page. The opinions are mine, from using it.
- Pangram: Pangram 4 Model Card
- Pangram: Introducing Pangram 4
- Pangram: Pangram enterprise page
- X: Max Spero on the Taylor Lorenz false flag
- The Atlantic: America Has a Pangram Problem
- The Atlantic: The Pangram Backlash Unfolding on College Campuses
- 404 Media: Substackers Say New AI Detection Tool Is a 'Witch Hunt'
- Spectrum News: UB student: False accusation over AI use inspired petition
- Bloomberg Businessweek: Do AI Detectors Work? Students Face False Cheating Accusations
- ABC News: University wrongly accuses students of using artificial intelligence to cheat
- Vanderbilt University: Guidance on AI Detection and Why We're Disabling Turnitin's AI Detector
- Patterns (Liang et al.): GPT detectors are biased against non-native English writers
- Fortune: Stop asking writers if they used AI. Just ask them how much
//Talk to our founders
Want to know which tools are worth it for your business?
Our founders look at how you sell today, then tell you which tools will earn their keep, which ones are still toys, and what to try first.
Talk to our foundersMore product reviews
- Gemini Notebook passed my pricing test.September 24, 2026
- Meta Muse is very fast. I still would not leave it alone.September 20, 2026

