The desk  /  Editing

I rewrote the same articles sixteen times. Line editing made them worse.

I assumed word choice was what made writing read as machine-made. Sixteen scored rewrites later I do not believe it, and the ranking surprised me.

Ilse Brandt · Editor  /  August 16, 2026  /  4 min read  /  5 sources
What this piece concludes
  • Four rounds of careful line editing on one article produced no direction: 60.2, 66.8, 55.6, 60.1. Two of them made it worse.
  • Changing where the piece starts, from topic-first to a moment in first person, moved the score nine points on its own.
  • Rhythm was the biggest single lever: 60.3 written evenly against 44.7 with sharp length variation, on identical content.
  • Extending an article from 577 to 1,189 words raised the score from 43.0 to 58.1, because more text means more chance of an even stretch.
  • Five of the six pure line-editing passes scored worse than the version they replaced.

I have rewritten the same handful of articles sixteen times over the past week, scoring each version on a detector, and the results were not what I assumed going in.

The assumption was that word choice does it. Strip out the tells, the leverages and the robusts and the seamlesses, and a piece stops reading as machine-written. That is what most style guides imply and it is what I would have told you on Monday.

Sixteen scores later I do not believe it.

The first article went through four versions. I moved the structure around, rewrote the rhythm, and stripped punctuation habits. Scores: 60.2, then 66.8, then 55.6, then 60.1. Four rounds of careful editing produced no direction at all. Two of them made it worse.

Then I changed one thing that was not a word. Instead of opening with the topic, I opened with the moment somebody asked the question, and let the piece follow what actually happened that afternoon. Same facts, same figures, same conclusions. The score dropped nine points on that change alone.

The second lever was rhythm, and it was bigger than I expected. A version written in even, well-formed sentences scored 60.3. The same content, rewritten so that three-word sentences sit next to forty-word ones, scored 44.7. Fifteen and a half points, from nothing but variation in length. The detector even names it: it reports a stretch of roughly three hundred words with too little variation and calls that stretch generated.

Length worked against me too. I extended one article from 577 words to 1,189 by adding explanation, which felt like an improvement and read like one. It scored 58.1 against the short version’s 43.0. More text is simply more opportunity to produce an even stretch, and the machine finds it.

Sixteen versions, one detector

Here is the part that stung. Every time I edited at the sentence level, tightening phrasing and swapping words, the score went up. Not down. Six of my sixteen versions were pure line editing, and five of those six scored worse than what they replaced.

That is a strange thing to learn about your own craft. The moves that feel like editing, the ones you were trained to make, are invisible to this measurement or actively counterproductive. The moves that work are structural and slightly uncomfortable: begin somewhere odd, leave a digression in, put the practical material in the middle instead of the end, and stop without resolving the last thought.

So the workflow changed. I no longer polish and re-measure, because sixteen data points say polishing does nothing. If a piece scores badly I go back to its plan, not its paragraphs.

The honest limits, since I am asking you to trust a small experiment. Sixteen versions is not a study. It is one detector, and other detectors weight things differently. The scores wobble by a couple of points on identical text, so I only trust gaps larger than about five.

What I would say confidently is the ranking. Plan matters most, rhythm second, vocabulary a distant third. If your editing time is limited, and it always is, spend it where the first two live.

The thing I have not resolved is whether any of this makes the writing better for a person, or only less detectable by a machine. Those are different goals and I have been quietly treating them as the same one all week. The rhythm changes did make the pieces sharper to read, I think. The odd openings, I am much less sure about.

Questions people ask about this

What is the ranking, in one line?

Plan first, rhythm second, vocabulary a distant third. If editing time is limited, spend it on where the piece starts and how the sentence lengths move.

Why does length make the score worse?

A longer piece gives more opportunity to produce a stretch of evenly built sentences, and that is exactly what the detector reports as generated.

Is sixteen versions enough to conclude anything?

No. It is one detector and one writer, the scores wobble by a couple of points on identical text, and I only trust gaps larger than about five points.

Does any of this make the writing better for a reader?

The rhythm changes did, I think. The unusual openings I am much less sure about, and I have been treating two different goals as one all week.

Read next
Before the next draft

Build the brief before the draft. Type a topic, an audience and a goal, and the generator gives you the questions to answer, the sources to find and the number to measure.

Build a brief free Open the glossary