Input first. Filter last.
Most people write the draft, then try to humanise it. That's backwards.
Put the Tuesday in first. The CIO who wanted a slide, not a system. The owner who brought warm milk because you looked sorry about the black coffee. Then clean the tells. Not the other way round.
The input is the work.
The rest of this page is that method, then the voices, tools, and tests behind it.
Five steps
Order matters. One and two do the work. Most people jump to five and wonder why it still sounds like a machine.
Start from a real fragment
A voice note. A WhatsApp to yourself. Three bullets you'd say out loud. Not "write a post about X."
Feed the specifics first
A name. A number. A place. A thing that happened. Generic in, generic out.
Draft in your register
Pick one reader. If you wouldn't say the sentence to their face, cut it.
Cut hard. Leave the lumps.
A short line next to a long one. An opinion that doesn't hedge. Evenness is the machine tell.
Filter last, and lightly
Now kill the em-dashes, the slop words, the tidy closer. This pass cleans writing that's already alive. It won't save a draft that started empty.
The people worth studying
Two rooms. The craft canon, who worked out how to earn a read long before machines, and the current debate, the people arguing about AI and writing right now in essays, books and feeds. Steal the move in each card.
The craft canon
The current debate
Who is actually talking about this, and where. The essayists and critics on one side, the people building and studying the detectors on the other.
The moves that carry
Techniques that survived a century of direct response and still work when a model holds the pen. Each is a pattern you can apply to the next thing you write.
The stack, free and paid
Three kinds of tool sit in a writing workflow: the model that drafts, the detector that checks, the humanizer that claims to disguise. Prices are August 2026 and move often. Check the current page before buying.
Drafting models
These write the draft. All three are strong. None can supply the lived specifics from step two, which is the part that matters.
| Tool | Cost | What it does | What it can't do |
|---|---|---|---|
| Claude (Anthropic) | Free tier; Pro ~$20/mo; API per token | Strongest at long-form reasoning and holding a specified voice. Best fit for drafting in your register from a fragment. | Cannot invent your memories or facts. Has its own default tells if left unguided. |
| ChatGPT (OpenAI) | Free tier; Plus ~$20/mo; API per token | Versatile all-rounder, wide tool ecosystem, fast general drafting. | Defaults to a recognisable house style; same specifics gap. |
| Gemini (Google) | Free tier; AI Pro ~$20/mo; API per token | Very long context, tight Google Docs, Gmail and Search integration. | Same voice-genericness and specifics gap. |
AI detectors, free and open
Run these yourself at no cost. Useful for research and understanding the tells. On modern prose they are unreliable: the legacy classifiers miss almost everything, the newer ones false-flag genuine humans.
| Tool | Cost | What it does | What it can't do |
|---|---|---|---|
| OpenAI RoBERTa detector | Free, open | Original GPT-2-era classifier. Historical reference. | Near-useless on current models. Scored everything near zero in our test. |
| Hello-SimpleAI RoBERTa | Free, open | Trained on early ChatGPT output. | Weak on 2026 text; misses modern drafts. |
| Fast-DetectGPT / perplexity | Free, zero-shot | Statistical, no training. Measures how predictable the text is. | Unreliable on short or unusual text. Our best draft was the least predictable of the three. |
| Binoculars | Free, zero-shot | Compares two models' cross-entropy. Strong published results. | Needs two models and compute. Not a click-and-go product. |
| desklib DeBERTa / e5-LoRA | Free, open | Newer RAID-benchmark classifiers, better than the legacy pair. | In our run, false-flagged real human writing at up to 88%. |
AI detectors, paid
These are the ones that work, and the ones that will judge your writing in the real world.
| Tool | Cost (Aug 2026) | What it does | What it can't do |
|---|---|---|---|
| Pangram | No free tier; ~$0.05 per 1,000 words API; LMS licensing | Highest accuracy in our test and independent reviews. Claimed 0.01% false-positive rate. Catches paraphrased and humanised text. Works on passages as short as 50 characters. | No plagiarism check, no image detection, no free tier. |
| GPTZero | Free tier (limited); ~$15 / $24 / $46 per month | AI detection plus plagiarism, writing feedback, citation checks, Chrome extension, sentence-level highlighting. | Free tier caps advanced scans fast; slightly behind Pangram on humanised text. |
| Originality.ai | Subscription plus credits, ~$0.01 per 100 words; API | AI plus plagiarism, built for publishers and content teams, bulk scanning, team seats. | Higher false-positive reputation on some human text; accuracy claims are vendor-published. |
| Copyleaks | ~$14 to $17 per month; API and enterprise | AI plus plagiarism, strong API and LMS integrations, multilingual. | Validate high-stakes languages against real samples; no single headline accuracy figure. |
| Winston AI | ~$10 to $16 per month | AI plus plagiarism, OCR for scanned documents, readable reports for educators. | Smaller language footprint; frames its own detection as useful but not infallible. |
| Turnitin | Institution licence only | The detector that actually matters in academia, embedded in university submission systems. | No consumer access, no public API. You cannot self-test against it. |
Humanizers
These rewrite AI text to evade detection. Included because the honest answer is that they do not solve the problem this guide is about, and usually make it worse by flattening the voice you were building.
| Tool | Cost | What it claims | What it can't do |
|---|---|---|---|
| Undetectable.ai | ~$10 to $15 per month | Rewrites text to pass common detectors, multiple tone presets. | Clears weak detectors, caught by retrained frontier ones like Pangram. Blurs meaning and voice. |
| StealthWriter | ~$10 to $20 per month | Similar evasion rewriting, sentence-level control. | Same ceiling against frontier detectors. Any pass is temporary as detectors retrain. |
| QuillBot | Free tier; Premium ~$10/mo | Paraphraser and grammar tool, useful for rewording. | Not really an evasion tool. Varies phrasing, will not defeat serious detection, strips your specifics. |
Before and after
Both blocks answer the same brief. The first is a competent default AI draft. The second is the same topic run through the process, starting from a real specific.
Small hotels beating chains on service
Why corporate AI pilots fail
2025 into 2026
Two things moved fast this year, in opposite directions. Detection got much sharper. And the flood of competent, forgettable AI prose got much deeper.
Detection got a lot better. The 2023 detectors were near-random on edited text. The 2025-26 leaders are a different class: in GPTZero's October 2025 head-to-head, both it and Pangram posted high-90s accuracy with sub-0.2% false-positive claims. Pangram's edge is hard-negative mining: generate AI text engineered to look like the humans it got wrong, retrain, repeat. The result, confirmed in our live test, is that the best detectors catch humanised, em-dash-free, voice-matched text at 100%.
Evasion works on the weak tools, not the strong ones. A February 2026 study, StealthRL, trained a paraphraser that almost entirely evaded the open and zero-shot detectors it tested (RoBERTa, Fast-DetectGPT, Binoculars). It did not test Pangram or GPTZero, so it says nothing about them directly. The claim that the frontier detectors hold up comes from a different place: Pangram's own retrained approach and independent humanizer testing, both of which show the strongest detectors resisting what beats the rest. A companion paper, Base Models Look Human, explains the mechanism: detectors track instruction-tuning residue that survives surface edits. This is a moving target, not a settled result.
Better or worse? Both. The median document is more competent than two years ago, because a decent draft is now free and instant. The distinctive, worth-reading piece got rarer, drowned in uniform slop produced at volume. The scarce thing is no longer clean prose. It is prose with a person in it, which is the whole commercial case for this guide.
Where this actually stands
Four findings that complicate the easy story in both directions. Almost everyone now writes with AI. Readers say they distrust AI content. And in blind tests they cannot spot it, and often prefer it.
Everyone is using it, few say how much
97% of content marketers planned to use AI in 2026, up from 65% in 2023, yet only 1% report fully AI-generated work. The common shape is assisted, not automated: 74% use it for ideation, 61% for outlining, 44% for drafting, and editing use doubled to 38% in a year. Most-trusted tools are ChatGPT (80%), Claude (55%), then Gemini and Perplexity. Over 40% of long-form LinkedIn posts are now fully AI-generated. That flood is the context for everything else here.
Readers say they distrust it
In a 2026 sentiment survey, 69% said they trust AI content less than human writing and only 8% trust it more. 61% are unlikely to engage with content they believe is AI. The distrust is sharpest exactly where voice matters most: opinion, news, and social, all around 68 to 69% preferring human. And 67% report having already seen AI content that was false or misleading. That last number, not "AI-ness," is the real complaint.
But they can't actually tell, and often prefer it
Here is the uncomfortable part. In a blind Cambridge study of roughly 1,700 readers, people identified the human-written story correctly only 39% of the time, worse than a coin flip. AI stories were rated higher on quality (1.54 vs 0.97) and on absorption (1.42 vs 1.00). The surface cues readers trust, language and enjoyment, were the ones that most often led them to the wrong answer.
The thing that moves the needle is the label, not the prose
The same study found the real tell is not in the writing. When a story was labelled "written by a human" it scored higher on quality and absorption than when labelled "AI," whichever actually wrote it. The penalty is for disclosure, not for detectable machine-ness. People punish the badge, not the sentences.
Why it actually offends: the broken effort contract
The deepest version of the objection has nothing to do with detection. Writer and reader share an old, unspoken contract: the effort of producing a piece roughly matches the effort of reading it, so the reader can trust that the subject genuinely wormed its way through the author's head. AI breaks that symmetry. You can blast out a competent draft in seconds and still ask a reader to spend real minutes on it. Put bluntly, if the author could not be bothered to write it, why should anyone be bothered to read it?
Underneath sits a sharper charge: signing your name to a machine draft is a kind of forgery. The philosopher Denis Dutton, in "Artistic Crimes," argued that we rightly value the performance behind a work, not just its surface. Han van Meegeren painted fake Vermeers with real skill, and critics were right to value them less once the deception was known, because the claimed performance was a lie. Passing an AI draft off as your own is the same move. Even a memo carries an implication of performance, that a person understood the thing and chose these words. That is the real reason disclosure matters, and the real reason genuine, specific writing outlasts a clean fake.
The research behind the claims
Everything asserted above traces to a paper or a live test. The load-bearing sources, one line each.
What's oversold
Four claims that get repeated and do not survive contact with the evidence.
A different question
AI detection and plagiarism detection get confused constantly. They answer different questions and fail in different ways.
Plagiarism detection asks: does this match existing sources? It compares against a database and the web and reports similarity. A mature, reliable matching problem. If you copied, it finds the copy.
AI detection asks: does this look statistically like machine writing? It reads patterns and returns a probability. A guessing problem, far less reliable. A piece can be fully original, zero plagiarism, and still trip AI detection for being unusually uniform.
Tools that do both, Turnitin, Copyleaks, Originality, Winston, run them as separate checks. Do not read one as the other. Genuine, specific, first-person writing scores well on both: original because it is yours, human-reading because it carries lived detail no model would generate.
The classroom reality
If you are a student or a teacher, this carries real consequences, and the ground shifted hard in 2025-26.
A wave of universities switched detection off. Curtin disabled it across all campuses in January 2026; Saint Joseph's ended it for fall 2025; Vanderbilt, Yale, Johns Hopkins and Northwestern had turned it off earlier. Their reasons were consistent: false positives too high to defend, systematic bias against non-native writers, and a 15% miss rate that lets real AI use through anyway.
So how do teachers catch AI in 2026? Increasingly not with detectors, but with process evidence: draft history, version trails, and short oral checks where a student explains their own argument. That rewards exactly the process here. Write input-first, from your own specifics and your own drafts, and you have the trail and the understanding to stand behind the work.
Still open
The honest edges of the field, where nobody has a settled answer yet.
- Q1Detection or evasion, which wins long-term? Right now the best detectors lead, but the gap has closed and reopened twice in three years.
- Q2Should schools use AI detectors at all, given the false-positive and bias record? The trend is toward switching them off and assessing process instead.
- Q3What becomes the disclosure norm for AI-assisted writing in business and journalism? Nobody has a standard yet, and "wrote it myself" is quietly becoming meaningless.
- Q4Does provenance watermarking (signing AI output at generation) arrive and actually get adopted? If it does, detection stops being guesswork. If it does not, this all stays a probabilistic arms race.
Our own test
Three matched sets of writing on the same eight topics, scored across five open detectors, a stylometric layer, and two leading commercial detectors.
The model has no Tuesday.
Thirty seconds of a real memory beats every tool on this page. Detectors will see you. Disclose the help. Sign your name to the specifics you actually have.
Read next
We used AI. Then we cut it.
A guide that calls hidden AI a forgery has to say how it was made.
Drafted with a model. Directed, edited, and fact-checked by a person. The August bench was run, not asserted. Sources checked by hand. Hide that and this page becomes the thing it argues against.