ATRIUMsearch → argument graph
MechanismArticle

LLMs are good at checking or producing individual components and making small edits, but they compound errors when asked to organize and connect those components into an entire technical chapter.

Models can handle sentences, equations, figures, typo detection, and localized editing, but whole chapters tend to contain confusing wording, muddled organization, and conceptual errors. The author describes this as an irreducible accumulation of errors across repeated additions. ✦ AI generated

the author · Interconnects · 2026-08-12 · original ↗

Today, the models seem genuinely horrible at long-form technical writing. They can get a sentence right, but if you try and get them to write an entire chapter it’ll be a mix of sprinkled with confusing wording, muddled in its organization, and generally a bit off. They try to be too cute where they don’t need to be and in the process make random conceptual errors. The models in the near future will get much better at the small errors, especially as models get bigger — which allows them to hold more world knowledge — but I do not expect their ability to utilize it to transform. For example, the GPT models have been incredible at finding typos and minor issues for a long time. I passed a near-final draft of my book as a PDF to GPT 5.5 Pro and it found deep, surprising minor typos across the manuscript that is 200-300 pages. On the other hand, the Claude models have been much more useful as an editor. They have a lot more taste, tend to understand the mental model of the task better, and have more interesting suggestions to unstick the different forms of writer’s block. The examples I’ve given above all have a sort of consistent theme. The models know how to check every unit of content, in this case usually a sentence or equation or figure, or make one, specific section where you are caught. With these skills, they don’t do a good job revisiting components and stringing them together as they make many additions on top of each other. It feels like a sort of irreducible compounding errors.

Read full article ↗excerpt · fair-use quotation

Around this claim