humanizing-writing

AI writing tells in 2026, with a source for each

In 2026 the signs of AI writing are long sentences chained with “and”, nouns built from verbs, metaphor where a plain statement would do, and short closing lines written for effect. “Delve” has mostly gone. This page lists each tell with where it was measured or reported, and how far to trust it.

Two cautions before the list. A single tell proves little, since people write lists of three and closing lines too, and it’s several together that make a reader suspicious. And none of this is a way to accuse someone. Shan, Lee, and Hao (EMNLP 2026) found that text a model edits has a much weaker footprint than text a model generates, so evidence about generated text shouldn’t be turned on a person’s edited draft.

Last checked on 10 October 2026. The entries about specific Claude versions date fastest.

Measured against human writing

These come from published corpus comparisons. They are the strongest evidence on this page, and they still have limits: The Economist measured news journalism, and the Reinhart figures are from 2024 models.

Long, noun-heavy vocabulary

"the implementation of", "utilization", "significant", "increasingly"

Polysyllables, Latinate suffixes, and nouns made from verbs. All four models in the study do it; Gemini and Claude do it most. Reinhart et al. put nominalisations at 1.5 to 2 times the human rate.

Rule in the skill: Use the plain verb and the common word. Source: The Economist, "How to Spot AI Writing", 30 July 2026.

"Not X but Y", "not only but also", and lists of three

"It's not just a defensive measure — it's a foundational component."

ChatGPT and Claude use these most per 1,000 sentences.

Rule in the skill: Deny X only when this reader is likely to believe X. Source: The Economist, "How to Spot AI Writing", 30 July 2026.

The trailing "-ing" clause

"..., underscoring its importance"

Measured at 2 to 5 times the human rate. It attaches significance to a fact without giving evidence for it.

Rule in the skill: Make claims specific and checkable. Source: Reinhart et al., PNAS, 2025.

The em dash, in Claude only

Of the four models, only Claude uses more em dashes than human writers. ChatGPT uses fewer than any other writer in the study. The article gives no figure for the size of the gap.

Rule in the skill: Use an em dash only for a break that a comma would hide. Source: The Economist, "How to Spot AI Writing", 30 July 2026.

Described by the model's own vendor

Anthropic describing its own model. It is not a measurement, and it compares the model with its predecessor, not with people.

Mannered prose

"a dial worth turning" for "a parameter worth varying"

Metaphor and flourish in place of direct statement. Anthropic's guide names the habit and gives this example in its suggested prompt.

Rule in the skill: Replace a figure of speech with the fact it stands for. Source: Anthropic, "Prompting Claude Fable 5.1".

Dense prose

Sentences run longer and there are fewer paragraph breaks than in the previous model.

Rule in the skill: Break a paragraph where the subject changes. Source: Anthropic, "Prompting Claude Fable 5.1".

Reported by readers, unmeasured

These come from two blog posts summarising forum threads about Claude Opus 5. Nobody has counted them, and one widely shared post can make a habit look more common than it is.

Tells that have faded or did not hold up

A rule built on any of these is aimed at 2023.

"Delve", "tapestry", "testament", "intricate"

These sit on Wikipedia's list for 2023 to mid-2024. Its list for mid-2025 onward is four words: "emphasizing", "enhance", "highlighting", "showcasing". The page records that "delve" fell sharply in 2025.

Rule in the skill: No rule. The skill keeps a short table and bans no word. Source: Wikipedia, "Signs of AI writing", read October 2026.

Uniform sentence length

Version 1.x called this the strongest signal, on the strength of one 2024 study of Mistral, Falcon, and LLaMA. A 2026 test of 284 features across 27 models found most proposed indicators depended heavily on context.

Rule in the skill: Removed in 2.0.0. "Vary rhythm on purpose" produced closers and fragments. Source: El Attar et al., 2026.

Why fixing tells doesn’t fool a detector

Avoiding these tells does not get text past an AI detector, because detectors aren’t reading for them. Current commercial detectors are trained classifiers. Pangram’s technical report describes a model trained on paired human and synthetic text, and the product now has a separate head for text that has been through a humanizer. Pangram reports catching the output of nineteen humanizer tools more than 90% of the time, which is the vendor’s own figure.

Xu et al. (2026) found that GPTZero and Pangram often classify text from base models as human and text from the instruction-tuned versions of the same models as AI. Their conclusion is that detectors “are tracking artifacts of instruction tuning and local context”. Getting past them took a fine-tuned paraphraser applied repeatedly, and the published attacks of that kind use the detector’s own score to guide the rewrite.

A style guide read by an instruction-tuned model doesn’t change that model’s tuning. So humanizing-writing makes no claim about detector scores. The reader it’s written for is a person — and practised people are good at this. Russell, Karpinska, and Iyyer (2025) found that the majority vote of five frequent LLM users misclassified 1 of 300 articles.