Alex Kirshner ran his own draft through Pangram and watched the meter return “100% Human Written.” He still felt the machine in the room. The Slate writer had already cut em dashes by roughly half across 64,000 words of his recent work, then caught himself rewriting clean sentences so an editor would never wonder. The detector is supposed to protect human prose. Instead it is teaching writers to sand it down.
Kirshner’s August 18 account is personal and specific, yet it lands inside a wider shift. Novelists, magazine editors and LinkedIn posters are deliberately adding typos, first-person digressions and awkward phrasing so software will bless the page. The tools that score highest on academic benchmarks are also the ones reshaping the very habits they claim to defend.
What the Accuracy Numbers Show
Pangram markets itself as the most reliable commercial detector. Independent work from the University of Chicago Booth School of Business backs much of the claim. Researchers Brian Jabarian and Alex Imas tested commercial and open-source tools on nearly 2,000 human passages matched against AI versions from GPT-4.1, Claude and Gemini.
Pangram posted near-zero false positive rates in Booth tests on medium and long passages, and stayed under reasonable policy thresholds even on shorter text. Its accuracy metric (probability of ranking a random AI sample as more suspicious than a random human sample) hit 100 percent for most models and never dropped below 99.8 percent. Open-source RoBERTa, by contrast, misclassified up to 78 percent of human text.
| Detector | False positive rate (medium/long) | Policy-grade at 0.5% FPR cap | Notes from UChicago work |
|---|---|---|---|
| Pangram | Essentially zero | Yes | Only tool that held detection power under strict caps |
| GPTZero | Below 1% | Partial | Stronger on some false-negative trade-offs |
| Originality.ai | Below 1% | Partial | Higher false negatives on certain models |
| RoBERTa (open) | 30-78% | No | Unsuitable for high-stakes use |
The same team’s policy-cap framework for detector thresholds lets schools or employers set a hard ceiling on false accusations and then measure how much real AI still gets caught. Pangram alone kept working under a 0.5 percent false-positive ceiling. The company also points to third-party evaluations claiming industry lead from Chicago and Maryland. Kirshner notes the firm’s earlier public claim of a 1-in-10,000 false-positive rate; the Booth numbers are closer to zero on longer text yet still leave room for trouble at volume.
That gap between a near-zero rate and a zero rate matters once the same tool runs across thousands of pages. A detector can lead every benchmark table and still train writers to treat every flagged clause as a career risk. The accuracy story and the behavior story pull in opposite directions.
The Loop That Starts With a Clean Draft
Kirshner’s process now includes a final pass through the detector before an editor ever sees the file. Single-digit “AI-assisted” flags have appeared on passages he wrote by hand. He rewrote them anyway, then sent a flurry of texts to a longtime editor explaining himself. The editor shrugged. The anxiety stayed.
Claude Opus 5, asked to audit his Spate archive, found the em-dash drop of about 50 percent from early to late months. Antithetical frames (“not X but Y”) rose even as he tried to avoid them. He still used the construction twice in the very essay about the problem. Words such as “profound” and “tapestry” stayed mostly gone.
The pattern is no longer individual. Writers swap notes on which constructions trip the meter. Common targets include:
- Em dashes used for asides or emphasis
- The “not X but Y” or “It’s not A, it’s B” pivot
- Rule-of-three lists and stacked metaphors
- Words that flooded early LLM output: delve, tapestry, landscape (figurative), robust, pivotal
- Overly balanced sentence pairs and smooth transitions
Some of those markers are already outdated. A July Economist analysis found that only Claude still over-indexes on em dashes relative to human baselines; ChatGPT now uses fewer. Lack of punctuation overall has become a stronger statistical signal. The advice circulating among writers has not caught up. The fear remains.
Each clean draft therefore enters a second draft aimed at the software. The human editor arrives third. That order alone changes what reaches the page.
How Brittle the Meter Can Be
Freddie deBoer stress-tested Pangram after a reader flagged a year-old section of his work as 100 percent AI. The full 5,000-word essay cleared as 100 percent human. The 300-word excerpt that had been called out still scored 100 percent AI. He could reverse the trick: embed AI text inside human text and watch the whole pass, or drop human text into an AI frame and watch it fail.
He showed it is easy to provoke contradictory 100 percent scores simply by changing surrounding context. The percentage meter itself often snaps to polar 100 percent results rather than mixed shares that match actual word counts. In one controlled paste of three-quarters human writing plus a short ChatGPT tail, Pangram still returned 100 percent AI with high confidence.
DeBoer’s checks line up as a short set of contradictions:
- Full 5,000-word essay scored 100 percent human
- A 300-word excerpt from the same piece scored 100 percent AI
- AI text embedded in human text could pass as a whole
- Human text dropped into an AI frame could fail as a whole
- Three-quarters human writing plus a short ChatGPT tail still returned 100 percent AI
DeBoer is careful: the tool is not useless. It is brittle when treated as a one-shot verdict. At even a 0.19 percent false-positive rate, a writer who has published 1,200 posts since ChatGPT launched faces a non-zero expected number of spurious flags. Sentence-level highlights multiply the chances. Accusations built on a single screenshot travel farther than the caveats.
The Counterstyle Already Has Practitioners
While Kirshner rewrites to avoid flags, other writers lean into deliberate roughness. Wired reported in late July on a growing anti-AI literary counterculture. Novelist Laura Brooke Robson, drafting a book about an AI dating app, took “a more winding path to avoid an em dash or a tripartite list.” She invents phrases and leaves small mistakes as “winks to the reader.”
Australian novelist Ellena Savage said AI “has stolen several syntactical conventions I once held dear. I miss over-deploying the em dash. I miss the rule of three.” Editors are joining the push. Aeon’s Richard Fisher tells contributors to keep the rough edges. Pitchfork’s Mano Sundaresan encourages first-person voice and bans AI in review copy. Playboy’s Philip Picardi, flooded with AI-generated pitches, now prizes formally adventurous work that machines still struggle to fake.
On LinkedIn the effect is cruder: monthly internal memos that update banned AI tells, first-person mandates, and apps that intentionally insert typos and “sent from my iPhone” signatures because CEOs open those emails more often. One Clickhole piece titled “5 Reasons to Love Summer (NO AI)” left spelling errors on purpose. The joke lands because the fear is real.
I find myself writing less mechanically. I have this urge to make little mistakes, like winks to the reader, to show that there’s a human behind the words.
That is Robson, describing a change she believes improved her prose. Not every practitioner is so cheerful. The same pressures can produce self-conscious oddness that reads as performance rather than voice.
The split is practical. Some writers sand the surface so the meter stays quiet. Others roughen the surface so a human reader can still hear a person. Both responses treat the detector as an audience that arrives before the real one.
Students, Freelancers and the Scale Problem
High-stakes users feel the mismatch first. A college that runs every paper through a detector with a 0.5 percent false-positive floor will still generate dozens of incorrect flags per large class over a semester. Non-native English writers have long faced higher error rates on earlier detectors; commercial tools improved but the bias literature has not vanished. Freelancers who juggle multiple editors now treat the detector as a silent third party in the room.
Kirshner freelances for other outlets. One flagged draft produced a small “AI-assisted” percentage. He rewrote several sentences of pure human work, burned through free credits, and still hit send with an explanatory text blast. The business risk of looking dirty outweighed the aesthetic cost of sounding cleaner than he wanted.
Publishers have already pulled or delayed books after Pangram screens. Accusations against newspaper op-eds and prize-winning short stories circulate on the same screenshots. The tool’s own leadership has used it to call out named journalists. Once the accusation lands, the clearance score on the full document rarely travels as far.
Scale turns a tiny error rate into a steady drip of private rewrites and public doubt. The person who must answer the flag absorbs the cost long before any benchmark paper revisits the threshold.
Why Circulating Advice Falls Behind the Models
Writers still trade lists of banned tells built for earlier model habits. The July Economist finding undercuts several of those lists at once. Only Claude still over-indexes on em dashes relative to human baselines. ChatGPT now uses fewer. Lack of punctuation has become the stronger statistical signal, yet the shared advice has not shifted with it.
The lag is structural. Models change output distributions faster than group chats update their warning lists. A construction that once lit up a meter can go quiet while a new quiet pattern takes its place. Writers who optimize against last season’s tells may sand away voice without buying safety.
Kirshner’s own archive audit sits inside that lag. Em dashes fell by about half. Antithetical frames rose even as he tried to avoid them. The essay about the problem still used the pivot twice. The feedback loop rewards visible caution more than it rewards current calibration.
- July: Economist analysis shows em-dash habits diverging by model, with ChatGPT no longer over-indexing.
- Late July: Wired reports the anti-AI literary counterculture and deliberate roughness among novelists and editors.
- August 18: Kirshner publishes the personal account of cutting dashes, chasing clean scores, and still feeling the machine nearby.
By the time a tell becomes common knowledge, the underlying models may already have moved. The social list stays sticky. The prose keeps bending to it anyway.
What Travels After a Flag Appears
DeBoer’s excerpt test and the freelance experience point to the same distribution problem. A 300-word slice can score 100 percent AI while the 5,000-word parent scores 100 percent human. A small “AI-assisted” percentage on a handmade draft can still force a rewrite and a round of explanatory texts. Screenshots favor the alarming number. Context favors the full file. Screenshots move faster.
Publishers pulling or delaying books after screens, and leadership teams using the tool to call out named journalists, extend that pattern into institutional life. Clearance on the complete document does not automatically walk back the first charge. Sentence-level highlights multiply the surfaces where a polar score can attach.
For a writer with 1,200 posts since ChatGPT launched, even a 0.19 percent false-positive rate implies a non-zero expected count of spurious flags. Each one can demand labor: rechecking, rewriting, messaging editors, burning detector credits. The meter’s brittleness becomes a time tax whether or not the underlying page was machine-made.
Public Trust Collides With Private Habit
The same tension appears in other domains where technology arrives faster than norms. Public debates over tech ethics and trust often settle on paper long after people have already changed daily behavior. Writers are not waiting for the next benchmark paper. They are changing keyboard habits now.
Some of the change may prove useful. John Warner, author of a book on writing in the AI age, argues that the panic could finally kill the empty five-paragraph student essay that ChatGPT perfected. Spikier, more personal, more error-tolerant prose is harder to fake and, for some readers, more alive. The risk is that the opposite also spreads: a new median of carefully imperfect sentences engineered to clear a meter rather than to say something true.
Kirshner ends his piece with the only practical advice he found. The problem is not one more Slate essay. It is something to take to therapy. That line is both joke and diagnosis. When a writer’s own finished sentences feel unsafe until software blesses them, the detector has already done its deepest work. The page that finally scores 100 percent human may no longer sound like the person who wrote it.








