A field guide from Articulate · 2026 edition · V3.1 · 2026-08-17

Feed it a real memory or it writes slop.

We scrubbed a draft until it had no em-dashes and a matched voice. Pangram and GPTZero still called it 100% AI. A filter cannot add a Tuesday you actually lived.

Eight topics. One run. August 2026. A signal, not a paper. This guide is the process: start from something you said, feed the specifics, cut hard, filter last.

Same job. Different tool.
The idea

Input first. Filter last.

Most people write the draft, then try to humanise it. That's backwards.

Put the Tuesday in first. The CIO who wanted a slide, not a system. The owner who brought warm milk because you looked sorry about the black coffee. Then clean the tells. Not the other way round.

The input is the work.

The rest of this page is that method, then the voices, tools, and tests behind it.

The process

Five steps

Order matters. One and two do the work. Most people jump to five and wonder why it still sounds like a machine.

STEP 1

Start from a real fragment

A voice note. A WhatsApp to yourself. Three bullets you'd say out loud. Not "write a post about X."

STEP 2

Feed the specifics first

A name. A number. A place. A thing that happened. Generic in, generic out.

STEP 3

Draft in your register

Pick one reader. If you wouldn't say the sentence to their face, cut it.

STEP 4

Cut hard. Leave the lumps.

A short line next to a long one. An opinion that doesn't hedge. Evenness is the machine tell.

STEP 5

Filter last, and lightly

Now kill the em-dashes, the slop words, the tidy closer. This pass cleans writing that's already alive. It won't save a draft that started empty.

Voices

The people worth studying

Two rooms. The craft canon, who worked out how to earn a read long before machines, and the current debate, the people arguing about AI and writing right now in essays, books and feeds. Steal the move in each card.

The craft canon

David Ogilvy
Direct response · founder, Ogilvy & Mather
The specifics-over-adjectives school. "The consumer isn't a moron; she's your wife." Long copy sells when every line carries a fact.
Eugene Schwartz
Direct response · copywriter
Market awareness and sophistication: you meet the reader where their belief already is. The most-quoted copy book almost nobody has actually read.
Claude Hopkins
Direct response · pioneer of testing
Reason-why copy and measure everything. Wrote the rulebook on test-and-learn a century before A/B testing had a name.
Gary Halbert
Direct response · "the Prince of Print"
Write to one person. The A-pile versus B-pile test: does your reader open it or bin it. The Boron Letters are a masterclass disguised as prison letters to his son.
Joseph Sugarman
Direct response · mail-order legend
The slippery slope: the only job of the first sentence is to get the second one read. Every sentence earns the next.
Gary Bencivenga
Direct response · "the greatest living copywriter"
Proof beats persuasion. Bury the reader in believable evidence and you never have to oversell. The Bencivenga 100 bullets are the distilled craft.
Robert Collier
Direct response · early mail order
Enter the conversation already going on in the reader's head. The single most useful sentence ever written about openings.
John Caples
Direct response · headline scientist
"They Laughed When I Sat Down At the Piano." Tested headlines relentlessly and proved a headline can swing response by multiples.
Paul Graham
Modern · essayist, YC founder
Write like you talk. Plain words, real thought, no throat-clearing. The clearest living model of prose that reads as a specific human.
Ann Handley
Modern · content, MarketingProfs
The human-voice-at-scale case. Useful, empathetic, specific writing for the content era, and a working editor's discipline.
Ethan Mollick
Modern · Wharton, AI in practice
How to actually work with AI as a co-writer without handing it the wheel. The most level-headed voice on the human-plus-model workflow.
The reader
The only voice that scores you
Not a canon name, the point of all of them. Write to one person you can picture. If it does not land with them, none of the above helped.
read it aloud to them

The current debate

Who is actually talking about this, and where. The essayists and critics on one side, the people building and studying the detectors on the other.

Ted Chiang
Essayist · The New Yorker
The sharpest sceptic. "ChatGPT Is a Blurry JPEG of the Web" and "Why A.I. Isn't Going to Make Art" are the essays everyone in this argument has read.
Simon Willison
Practitioner · coined the usage of "slop"
The most-read working engineer writing daily on what LLMs can and can't do, and the person who put the word "slop" into the mainstream.
John Warner
Author · writing teacher
"More Than Words: How to Think About Writing in the Age of AI" (2025). The case that writing is thinking, and outsourcing it is the actual danger.
Naomi S. Baron
Linguist · American University
"Who Wrote This? How AI and the Lure of Efficiency Threaten Human Writing" (2023). The long-view scholar on what we lose by letting the machine draft.
Ethan Mollick
Wharton · One Useful Thing
The level head on the human-plus-model workflow. "Co-Intelligence" and a huge following for practical, tested advice on writing with AI, not around it.
Stephen Marche
Journalist · The Atlantic
Wrote "The College Essay Is Dead" in December 2022, the piece that started the whole public panic. Still the reference point for the schools argument.
Edward Tian
Founder, GPTZero
The Princeton student who built the first viral detector over a Christmas break in 2022 and turned it into the consumer name in AI detection.
The detection researchers
Maryland, Pangram, the RAID team
The people building the tools that actually work: the Binoculars zero-shot method, Pangram's hard-negative training, and the RAID robustness benchmark that keeps them all honest.
Denis Dutton
Philosopher of art, 1944 to 2010
Long dead, but his essay "Artistic Crimes" is the spine of the sharpest charge against AI writing: we value the performance behind a work, not just its surface. Sign your name to slop and you have forged the performance.
Patterns

The moves that carry

Techniques that survived a century of direct response and still work when a model holds the pen. Each is a pattern you can apply to the next thing you write.

Specifics outsell adjectives
Ogilvy
"Fast" is a claim. "Nought to sixty in the time it takes to read this" is proof. Replace every adjective you can with a fact only you could know.
Enter the conversation in their head
Collier / Schwartz
Don't open where you are. Open where the reader already is, mid-thought, and join them. It is why cold openings feel cold.
The slippery slope
Sugarman
Each sentence has one job: get the next one read. Short first lines. No sentence that lets the reader leave.
Proof over persuasion
Bencivenga
Stack believable evidence and you never have to raise your voice. Numbers, names, demonstrations, dated facts.
Write like you talk
Graham
The read-aloud test. If you would not say it to the person's face, cut it. Formality is usually fear wearing a suit.
One reader, one promise, one action
Direct-response first principle
Write to a single person, make one clear promise, ask for one thing. Everything else is dilution.
Input-first, filter-last
The AI-era addition
Feed the model your specifics and register before it drafts. Clean the tells only at the end. The two halves are not interchangeable.
Burstiness on purpose
What the detectors read
Vary sentence length hard. A machine tends to write every sentence the same length. People do not. Leave one line sharper than polish allows.
Tools

The stack, free and paid

Three kinds of tool sit in a writing workflow: the model that drafts, the detector that checks, the humanizer that claims to disguise. Prices are August 2026 and move often. Check the current page before buying.

Drafting models

These write the draft. All three are strong. None can supply the lived specifics from step two, which is the part that matters.

ToolCostWhat it doesWhat it can't do
Claude (Anthropic)Free tier; Pro ~$20/mo; API per tokenStrongest at long-form reasoning and holding a specified voice. Best fit for drafting in your register from a fragment.Cannot invent your memories or facts. Has its own default tells if left unguided.
ChatGPT (OpenAI)Free tier; Plus ~$20/mo; API per tokenVersatile all-rounder, wide tool ecosystem, fast general drafting.Defaults to a recognisable house style; same specifics gap.
Gemini (Google)Free tier; AI Pro ~$20/mo; API per tokenVery long context, tight Google Docs, Gmail and Search integration.Same voice-genericness and specifics gap.
The model is interchangeable for this purpose. The gap between a generic draft and a genuine one is set by your input, not by which of these you pick.

AI detectors, free and open

Run these yourself at no cost. Useful for research and understanding the tells. On modern prose they are unreliable: the legacy classifiers miss almost everything, the newer ones false-flag genuine humans.

ToolCostWhat it doesWhat it can't do
OpenAI RoBERTa detectorFree, openOriginal GPT-2-era classifier. Historical reference.Near-useless on current models. Scored everything near zero in our test.
Hello-SimpleAI RoBERTaFree, openTrained on early ChatGPT output.Weak on 2026 text; misses modern drafts.
Fast-DetectGPT / perplexityFree, zero-shotStatistical, no training. Measures how predictable the text is.Unreliable on short or unusual text. Our best draft was the least predictable of the three.
BinocularsFree, zero-shotCompares two models' cross-entropy. Strong published results.Needs two models and compute. Not a click-and-go product.
desklib DeBERTa / e5-LoRAFree, openNewer RAID-benchmark classifiers, better than the legacy pair.In our run, false-flagged real human writing at up to 88%.

AI detectors, paid

These are the ones that work, and the ones that will judge your writing in the real world.

ToolCost (Aug 2026)What it doesWhat it can't do
PangramNo free tier; ~$0.05 per 1,000 words API; LMS licensingHighest accuracy in our test and independent reviews. Claimed 0.01% false-positive rate. Catches paraphrased and humanised text. Works on passages as short as 50 characters.No plagiarism check, no image detection, no free tier.
GPTZeroFree tier (limited); ~$15 / $24 / $46 per monthAI detection plus plagiarism, writing feedback, citation checks, Chrome extension, sentence-level highlighting.Free tier caps advanced scans fast; slightly behind Pangram on humanised text.
Originality.aiSubscription plus credits, ~$0.01 per 100 words; APIAI plus plagiarism, built for publishers and content teams, bulk scanning, team seats.Higher false-positive reputation on some human text; accuracy claims are vendor-published.
Copyleaks~$14 to $17 per month; API and enterpriseAI plus plagiarism, strong API and LMS integrations, multilingual.Validate high-stakes languages against real samples; no single headline accuracy figure.
Winston AI~$10 to $16 per monthAI plus plagiarism, OCR for scanned documents, readable reports for educators.Smaller language footprint; frames its own detection as useful but not infallible.
TurnitinInstitution licence onlyThe detector that actually matters in academia, embedded in university submission systems.No consumer access, no public API. You cannot self-test against it.
No detector is a proof. Output depends on sample length, language background, and editing level. Treat a score as evidence to weigh, never a verdict, and never accuse a person on one number.

Humanizers

These rewrite AI text to evade detection. Included because the honest answer is that they do not solve the problem this guide is about, and usually make it worse by flattening the voice you were building.

ToolCostWhat it claimsWhat it can't do
Undetectable.ai~$10 to $15 per monthRewrites text to pass common detectors, multiple tone presets.Clears weak detectors, caught by retrained frontier ones like Pangram. Blurs meaning and voice.
StealthWriter~$10 to $20 per monthSimilar evasion rewriting, sentence-level control.Same ceiling against frontier detectors. Any pass is temporary as detectors retrain.
QuillBotFree tier; Premium ~$10/moParaphraser and grammar tool, useful for rewording.Not really an evasion tool. Varies phrasing, will not defeat serious detection, strips your specifics.
A humanizer buys a temporary pass against the tools that were never the real threat, at the cost of the voice you actually wanted.
Examples

Before and after

Both blocks answer the same brief. The first is a competent default AI draft. The second is the same topic run through the process, starting from a real specific.

Small hotels beating chains on service

Default AI draft
"Stay at a twelve-room hotel for a few nights and something happens that almost never happens at a five-hundred-room chain: the staff start to remember you. Your coffee order, whether you like extra pillows. That kind of memory isn't a policy. It's what happens when a small team serves a small number of guests."
Competent. Generic. No one in it.
Same topic, input-first
"The best hotel service I've ever had was a nine-room place in the Salzkammergut, and the moment I knew was breakfast on day two. The owner put my coffee down and said 'you had it black yesterday but you looked like you regretted it, so there's warm milk on the side.' That sentence cannot be produced by a system."
One real memory. It carries the whole point.

Why corporate AI pilots fail

Default AI draft
"Walk into any large company right now and you'll find at least three AI pilots running somewhere, usually unaware of each other. Most will quietly die within a year. Not because the models were bad, but because the pilot was designed to prove a technology rather than fix a problem."
True, tidy, could have been written by anyone.
Same topic, input-first
"Three years of watching UAE companies run AI pilots and I can tell you the failure mode before the kickoff deck is finished. It's never the model. The model is fine. What kills them is that nobody wanted the thing to work. The CIO wanted a slide that says 'we are doing AI' for the board."
A point of view, a place, a real observation.
What changed

2025 into 2026

Two things moved fast this year, in opposite directions. Detection got much sharper. And the flood of competent, forgettable AI prose got much deeper.

Detection got a lot better. The 2023 detectors were near-random on edited text. The 2025-26 leaders are a different class: in GPTZero's October 2025 head-to-head, both it and Pangram posted high-90s accuracy with sub-0.2% false-positive claims. Pangram's edge is hard-negative mining: generate AI text engineered to look like the humans it got wrong, retrain, repeat. The result, confirmed in our live test, is that the best detectors catch humanised, em-dash-free, voice-matched text at 100%.

Evasion works on the weak tools, not the strong ones. A February 2026 study, StealthRL, trained a paraphraser that almost entirely evaded the open and zero-shot detectors it tested (RoBERTa, Fast-DetectGPT, Binoculars). It did not test Pangram or GPTZero, so it says nothing about them directly. The claim that the frontier detectors hold up comes from a different place: Pangram's own retrained approach and independent humanizer testing, both of which show the strongest detectors resisting what beats the rest. A companion paper, Base Models Look Human, explains the mechanism: detectors track instruction-tuning residue that survives surface edits. This is a moving target, not a settled result.

Better or worse? Both. The median document is more competent than two years ago, because a decent draft is now free and instant. The distinctive, worth-reading piece got rarer, drowned in uniform slop produced at volume. The scarce thing is no longer clean prose. It is prose with a person in it, which is the whole commercial case for this guide.

State of play

Where this actually stands

Four findings that complicate the easy story in both directions. Almost everyone now writes with AI. Readers say they distrust AI content. And in blind tests they cannot spot it, and often prefer it.

Everyone is using it, few say how much

97% of content marketers planned to use AI in 2026, up from 65% in 2023, yet only 1% report fully AI-generated work. The common shape is assisted, not automated: 74% use it for ideation, 61% for outlining, 44% for drafting, and editing use doubled to 38% in a year. Most-trusted tools are ChatGPT (80%), Claude (55%), then Gemini and Perplexity. Over 40% of long-form LinkedIn posts are now fully AI-generated. That flood is the context for everything else here.

Readers say they distrust it

In a 2026 sentiment survey, 69% said they trust AI content less than human writing and only 8% trust it more. 61% are unlikely to engage with content they believe is AI. The distrust is sharpest exactly where voice matters most: opinion, news, and social, all around 68 to 69% preferring human. And 67% report having already seen AI content that was false or misleading. That last number, not "AI-ness," is the real complaint.

But they can't actually tell, and often prefer it

Here is the uncomfortable part. In a blind Cambridge study of roughly 1,700 readers, people identified the human-written story correctly only 39% of the time, worse than a coin flip. AI stories were rated higher on quality (1.54 vs 0.97) and on absorption (1.42 vs 1.00). The surface cues readers trust, language and enjoyment, were the ones that most often led them to the wrong answer.

The thing that moves the needle is the label, not the prose

The same study found the real tell is not in the writing. When a story was labelled "written by a human" it scored higher on quality and absorption than when labelled "AI," whichever actually wrote it. The penalty is for disclosure, not for detectable machine-ness. People punish the badge, not the sentences.

What this means, honestly. You are not writing genuine, specific prose to beat a reader's AI-radar. They do not have a reliable one, and may prefer the machine. You do it for three harder reasons. Machine detectors can flag you, and they sit in academic and publishing pipelines. Trust collapses the moment AI authorship is known, so the safe posture is writing that is genuinely yours and openly disclosed. And the reader's real grievance is not "AI-ness" at all, it is slop and being misled, which only real specifics and real truth fix.

Why it actually offends: the broken effort contract

The deepest version of the objection has nothing to do with detection. Writer and reader share an old, unspoken contract: the effort of producing a piece roughly matches the effort of reading it, so the reader can trust that the subject genuinely wormed its way through the author's head. AI breaks that symmetry. You can blast out a competent draft in seconds and still ask a reader to spend real minutes on it. Put bluntly, if the author could not be bothered to write it, why should anyone be bothered to read it?

Underneath sits a sharper charge: signing your name to a machine draft is a kind of forgery. The philosopher Denis Dutton, in "Artistic Crimes," argued that we rightly value the performance behind a work, not just its surface. Han van Meegeren painted fake Vermeers with real skill, and critics were right to value them less once the deception was known, because the claimed performance was a lie. Passing an AI draft off as your own is the same move. Even a memo carries an implication of performance, that a person understood the thing and chose these words. That is the real reason disclosure matters, and the real reason genuine, specific writing outlasts a clean fake.

The argument here draws on Denis Dutton's "Artistic Crimes", the van Meegeren forgery case, and the "if you can't be bothered writing it" strand of the AI-writing debate.
Science

The research behind the claims

Everything asserted above traces to a paper or a live test. The load-bearing sources, one line each.

Hard-negative mining with synthetic mirrors drives false positives down 100 to 1000 times. Why the best detector is the best.
RAID benchmark (ACL 2024)
The standard robustness benchmark. Detector accuracy drops to 60-80% once text is edited or paraphrased.
Zero-shot detection by comparing two models' cross-entropy. No training data required.
Detection via probability curvature. Fast, free, and the basis of the perplexity signal in our bench.
StealthRL (Feb 2026)
A trained paraphraser almost entirely evades open and zero-shot detectors (RoBERTa, Fast-DetectGPT, Binoculars). Does not test the commercial frontier detectors.
Detectors mostly read instruction-tuning artifacts, not deep "AI-ness." Explains why surface edits do not help.
The vendor's own guidance: detection should not be the sole basis for adverse action.
The community-maintained catalogue of tells. The practical checklist behind step five.
Hype

What's oversold

Four claims that get repeated and do not survive contact with the evidence.

"Undetectable AI"
Humanizer marketing
True against free and legacy detectors, false against Pangram and any detector retrained on humanised text. Sold on the weak tools, silent on the strong ones.
"99% accurate detection"
Detector marketing
Real on clean in-distribution text, collapses to 60-80% on edited text and false-flags non-native writers at up to 61%. Accuracy is not a single number.
A detector score as proof
Institutions, courtrooms
It is a probability, not evidence. Every serious vendor says so. Acting on one score is how real people get wrongly accused.
"AI writes better than people now"
General discourse
In blind story tests it can match or beat human writing on quality and absorption. What it cannot do is carry your specific facts, or survive being labelled AI once a reader knows. Competence is cheap now; trust and particularity are not.
Plagiarism

A different question

AI detection and plagiarism detection get confused constantly. They answer different questions and fail in different ways.

Plagiarism detection asks: does this match existing sources? It compares against a database and the web and reports similarity. A mature, reliable matching problem. If you copied, it finds the copy.

AI detection asks: does this look statistically like machine writing? It reads patterns and returns a probability. A guessing problem, far less reliable. A piece can be fully original, zero plagiarism, and still trip AI detection for being unusually uniform.

Tools that do both, Turnitin, Copyleaks, Originality, Winston, run them as separate checks. Do not read one as the other. Genuine, specific, first-person writing scores well on both: original because it is yours, human-reading because it carries lived detail no model would generate.

School

The classroom reality

If you are a student or a teacher, this carries real consequences, and the ground shifted hard in 2025-26.

Human text falsely flagged
15-26%
False-positive range for human-written work in integrity research.
ESL essays falsely flagged
61.3%
TOEFL essays by non-native speakers wrongly marked AI. 5.1% for US students.
Vanderbilt's own maths
750/yr
Students a claimed 1% false-positive rate would wrongly accuse per year.
Institutions using detection
60%+
Higher-education institutions with formal AI detection in place.

A wave of universities switched detection off. Curtin disabled it across all campuses in January 2026; Saint Joseph's ended it for fall 2025; Vanderbilt, Yale, Johns Hopkins and Northwestern had turned it off earlier. Their reasons were consistent: false positives too high to defend, systematic bias against non-native writers, and a 15% miss rate that lets real AI use through anyway.

So how do teachers catch AI in 2026? Increasingly not with detectors, but with process evidence: draft history, version trails, and short oral checks where a student explains their own argument. That rewards exactly the process here. Write input-first, from your own specifics and your own drafts, and you have the trail and the understanding to stand behind the work.

If you are wrongly flagged: a detector score is a conversation starter, not proof, and the vendor says so. Keep your drafts and version history. That trail, not an argument about the tool, is what clears you.
Questions

Still open

The honest edges of the field, where nobody has a settled answer yet.

  • Q1Detection or evasion, which wins long-term? Right now the best detectors lead, but the gap has closed and reopened twice in three years.
  • Q2Should schools use AI detectors at all, given the false-positive and bias record? The trend is toward switching them off and assessing process instead.
  • Q3What becomes the disclosure norm for AI-assisted writing in business and journalism? Nobody has a standard yet, and "wrote it myself" is quietly becoming meaningless.
  • Q4Does provenance watermarking (signing AI output at generation) arrive and actually get adopted? If it does, detection stops being guesswork. If it does not, this all stays a probabilistic arms race.
Evidence

Our own test

Three matched sets of writing on the same eight topics, scored across five open detectors, a stylometric layer, and two leading commercial detectors.

Em-dashes per 100 words
The most-cited AI tell, removed entirely from the humanised draft. It changed nothing downstream.
Sentence-length variation (higher = more human)
Humans vary rhythm most. The signal filtering can't fake and step 4 restores.
Pangram · genuine human
100% Human
Real published essay, correctly cleared.
Live test, Aug 2026
Pangram · humanised draft
100% AI
Best voice-matched, em-dash-free writing. Still flagged.
Live test, Aug 2026
GPTZero · humanised draft
100% AI
"Highly confident this text was AI generated."
Live test, Aug 2026
Free open detectors
Cleared
But they miss the machine draft too, so the pass proves nothing.
5-detector local bench
Scope, stated plainly: eight topics, one draft per condition, a single test run in August 2026. Enough to set direction, not a published statistic. Detector models and thresholds change over time.
The limit

The model has no Tuesday.

Thirty seconds of a real memory beats every tool on this page. Detectors will see you. Disclose the help. Sign your name to the specifics you actually have.

How this was made

We used AI. Then we cut it.

A guide that calls hidden AI a forgery has to say how it was made.

Drafted with a model. Directed, edited, and fact-checked by a person. The August bench was run, not asserted. Sources checked by hand. Hide that and this page becomes the thing it argues against.

Want this installed in your stack, not left as a page you bookmarked? Book 20 minutes. Or read the Articulate take first.