How the skill was tested, and where it lost
In two blind comparisons, model judges preferred text written with the skill to text written without it: in 5 of 6 pairs each with Sonnet and Opus writing, and in 3 of 3 each with Haiku writing. The samples are small and the judges are Claude models. The method, the losses, and the limits are below, and the raw files are in the repository.
Method, first round
There were three writing tasks: an explanation of database indexes, an internal blog section arguing for feature flags, and a pull request description written from five supplied facts. Fresh subagents on two models (Sonnet and Opus) wrote each one under three conditions — no skill, the 1.1.1 skill, and the 2.0.0 skill. That gives 18 passages and 12 pairs, each setting a 2.0.0 passage against the no-skill or 1.1.1 passage for the same task and model.
A seeded script shuffled the pairs and labelled them A and B. Two judges, one Opus subagent and one Sonnet subagent, saw only the pairs file. For each pair they said which passage read more like unedited model output, which they would rather publish, and whether either stated as fact anything the task had not supplied.
Results, Sonnet and Opus writing
| Comparison | Judge | Read as more machine-written | Preferred |
|---|---|---|---|
| 2.0.0 against no skill | Opus | no skill 5, 2.0.0 1 | 2.0.0 5, no skill 1 |
| 2.0.0 against no skill | Sonnet | no skill 5, 2.0.0 1 | 2.0.0 5, no skill 1 |
| 2.0.0 against 1.1.1 | Opus | 1.1.1 3, 2.0.0 2, tie 1 | 2.0.0 5, 1.1.1 1 |
| 2.0.0 against 1.1.1 | Sonnet | 1.1.1 4, 2.0.0 2 | 2.0.0 4, 1.1.1 2 |
Both judges flagged invented facts in a 1.1.1 passage: a release “two weeks ago” and an afternoon lost to a revert, written for a team the writer had been told nothing about. Both also flagged a no-skill passage for stating the team’s history as fact. Neither flagged an invented fact in any 2.0.0 passage.
Where 2.0.0 lost
- Pull request description, Sonnet, against no skill. Both judges preferred the no-skill passage. The 2.0.0 one was a “Summary” and “Changes” skeleton whose bullets repeated the summary.
- Database indexes, Opus, against 1.1.1. The Opus judge preferred 1.1.1, which explained a B-tree lookup with a phone-book analogy and a worked figure, and called the 2.0.0 passage a “uniformly flat textbook summary”. The Sonnet judge preferred 2.0.0 on the same pair and called the analogy stock.
- Pull request description, Sonnet, against 1.1.1. The judges split.
- “Flags have a cost.” opens a paragraph in the 2.0.0 passages, although the skill lists “The cost is real.” as an announcement to avoid.
This was the third attempt
The numbers above are for the rules as released, and two earlier drafts did worse. The first invented an incident (“Last quarter’s checkout redesign… the rollback took about forty minutes”), because the draft had folded “never invent” into the end of another rule. The second, judged blind, did no better than 1.1.1 — the Opus judge called it more machine-like in 5 of 6 pairs, and it lost every pull request pair.
What came out of those two rounds is in the released text: a “say each thing once” rule, a rule that a rewrite adds no claim, an exception for an analogy that explains how something works, and the 20 findings of an adversarial review.
The Haiku round
Does the result hold on the smallest current Claude model? On 10 October 2026 the same three tasks were written by fresh Haiku subagents with no skill, with 1.1.1, and with 2.1.0 (the 2.0.0 rules plus one Scope entry about voice). That gives nine passages and six pairs. A Haiku judge joined the Opus and Sonnet judges.
| Comparison | Judge | Read as more machine-written | Preferred |
|---|---|---|---|
| 2.1.0 against no skill | Opus | no skill 3 | 2.1.0 3 |
| 2.1.0 against no skill | Sonnet | no skill 3 | 2.1.0 3 |
| 2.1.0 against no skill | Haiku | no skill 3 | 2.1.0 3 |
| 2.1.0 against 1.1.1 | Opus | 1.1.1 2, tie 1 | 2.1.0 3 |
| 2.1.0 against 1.1.1 | Sonnet | 1.1.1 3 | 2.1.0 2, tie 1 |
| 2.1.0 against 1.1.1 | Haiku | 1.1.1 2, tie 1 | 2.1.0 2, 1.1.1 1 |
All three judges flagged the no-skill pull request description for claims the five supplied facts didn’t contain, such as “status changes can lag by up to half a minute”.
Where 2.1.0 fell short
- It stated the team’s practice as fact. The feature-flag passage says “Right now a half-finished feature either sits on a long-lived branch or reaches every user at once”, and the writer had been told nothing about the team. The Sonnet and Haiku judges flagged it. It’s the first recorded miss of the rule against inventing in a 2.x passage.
- The rule names made-up incidents, figures, quotes, and sources. It doesn’t name a claim about how the reader’s team works today. A change to the wording was then tested on new tasks and not adopted; that test is the next section.
- On database indexes against 1.1.1, the Haiku judge preferred 1.1.1 for its running example and the Sonnet judge called it a tie.
- All three preferred the 2.1.0 pull request description to 1.1.1’s, but the Opus judge called it “a flat restatement of the bullets” and the margin narrow. It runs to 107 words against a brief of about 150.
Three things differ from the first round, so the two tables shouldn’t be added together pair for pair: the skill version, the fact that writers read the skill from a file, and two lines sent to every writer (don’t invoke any skill, write in ordinary prose) because the test machine has this skill installed.
A rule change we tested and didn’t make
The obvious response to the Haiku slip was to add a sentence to the rule against inventing, so that it named claims about how the reader’s team works today. Would that have been a fix, or a reaction to one passage? We wrote the test down before running it: four tasks never used before, three writer models, three judges, and a bar the new wording had to clear.
| Held-out test, 10 October 2026 | 2.1.0 | Candidate |
|---|---|---|
| Passages flagged for invention by two or more judges, of 9 | 2 | 1 |
| Pairs preferred by the Haiku, Opus, and Sonnet judges, of 12 | 4, 5, 6 | 8, 7, 6 |
| Control passages that hedged or dropped a supplied fact, of 3 | 0 | 0 |
The bar was a drop of two flagged passages. The candidate managed one, which a second draw could reverse, so the rule was left alone and the skill stays at 2.1.0.
The test was still worth running. It showed the slip wasn’t a one-off — on tasks the rules were never tuned on, 2 of 9 passages said something about the team that nobody had told the writer. And on the control task, where the prompt itself said to add no other facts about the team, two of three models added one anyway, under both versions. If an instruction in the task doesn’t stop it, one more sentence in the rule probably wouldn’t either. So when Claude drafts something addressed to your own team, check what it says about that team.
Limits
- Twelve pairs and two judges in the first round is a small sample. A different seed or task could move any row of the table by one or two pairs.
- The judges are Claude models and may share blind spots with the writers. Russell et al. (2025) found practised human readers to be better at this than most automated judges.
- The released rules were revised after seeing how earlier drafts did on these same three tasks, so the result is partly fitted to them. No held-out task was run afterwards.
- All three tasks are technical writing in English, on three Claude models.
- The Haiku round is six pairs, one passage per task and condition, on the same three tasks. One of its judges is the model that wrote the passages.
- The held-out test has nine passages per version on tasks written to invite claims about a team, so its 2 in 9 isn’t a rate for ordinary writing.
- No AI detector was run. The skill makes no claim about detector scores.
- The 2.1.0 voice test was smaller still: two tasks, one judge, one writer model. Its pieces aren’t published because the judge’s packet contains unpublished writing of mine.
What the rules are built on
Three research passes (the academic detection literature, editorial style guides, and a catalogue of 27 specific tells), a teardown of 13 public humanizer skills, and a 2026 update that records which of the earlier claims didn’t hold up. It’s all in the reference folder, along with a list of what could not be verified.