ATRIUMsearch → argument graph
Article · 2026-08-12 · 6 moments

I wrote an AI textbook — how long until AI can do it better?

Reflections on AI's writing ability and how AI models get more capable. ✦ AI generated

01
Example

AI can provide substantial value in nonfiction writing when an expert supervises it closely, but it should mainly handle repetitive, verifiable, or editorial tasks rather than the parts that convey a work's story and intellectual soul.

The author used models for formatting, copyediting, diagrams, manuscript review, synchronization, and other repetitive tasks, occasionally accepting suggestions under expert review. In scientific papers, models are useful for routine background or related-work sections, but the expert should retain the abstract, introduction, experiments, and conclusion.

transcript

the author: I am working through similar balances in my scientific work too. AI models are great for repetitive pieces of the paper, like drafting a related work or background section that you know by heart, but using them for the abstract, introduction, experiments, or conclusion is a shame. Those are where the story and soul of the work is communicated — it’s where you learn what your research is really about. I am confident I created a lot more net value by being able to have AI models create and check my non-fiction writing work. They make writing equations trivial, can help refactor the repository, port between languages, and many other things. At the beginning, it was very fun, until I was a bit worn down by the length of the publishing process, watching the field move on. For an example of why AI was crucial in this case, I had to maintain Markdown and LaTeX versions of my book simultaneously in two spots, as readers gave feedback on the web version and my Manning editorial team reviewed a forked copy. Without AI agents, syncing between the two of them would’ve easily taken me five times as long (and this task took tens of hours already).

supports · 3

02
Mechanism

LLMs are good at checking or producing individual components and making small edits, but they compound errors when asked to organize and connect those components into an entire technical chapter.

Models can handle sentences, equations, figures, typo detection, and localized editing, but whole chapters tend to contain confusing wording, muddled organization, and conceptual errors. The author describes this as an irreducible accumulation of errors across repeated additions.

transcript

the author: Today, the models seem genuinely horrible at long-form technical writing. They can get a sentence right, but if you try and get them to write an entire chapter it’ll be a mix of sprinkled with confusing wording, muddled in its organization, and generally a bit off. They try to be too cute where they don’t need to be and in the process make random conceptual errors. The models in the near future will get much better at the small errors, especially as models get bigger — which allows them to hold more world knowledge — but I do not expect their ability to utilize it to transform. For example, the GPT models have been incredible at finding typos and minor issues for a long time. I passed a near-final draft of my book as a PDF to GPT 5.5 Pro and it found deep, surprising minor typos across the manuscript that is 200-300 pages. On the other hand, the Claude models have been much more useful as an editor. They have a lot more taste, tend to understand the mental model of the task better, and have more interesting suggestions to unstick the different forms of writer’s block. The examples I’ve given above all have a sort of consistent theme. The models know how to check every unit of content, in this case usually a sentence or equation or figure, or make one, specific section where you are caught. With these skills, they don’t do a good job revisiting components and stringing them together as they make many additions on top of each other. It feels like a sort of irreducible compounding errors.

explains mechanism · 1rebuts · 1supports · 3

03
Claim

Models have improved rapidly in coding, mathematics, search, and research, but writing well has barely advanced because it is a difficult, relatively orthogonal skill with inadequate specialized training data.

The author expected much greater progress in nonfiction writing, given the models' advances in other tasks. Specialized harnesses and prompts may yield incremental gains, but the fundamental difficulty of writing and the lack of suitable training interventions limit improvement.

transcript

the author: I would’ve expected way more progress on non-fiction writing from the models. I almost thought I would look dumb publishing a non-fiction book in 2026, given how things looked in 2024. Today, some of the most famous models on writing ability are pretty old, examples include OpenAI’s big GPT 4.5 and Moonshot’s Kimi K2. In and around these releases, the models have gone from okay to superhuman at other tasks like coding and mathematics. Maybe a closer, but still imperfect, comparison is how the models went from incapable to decent at search and research tasks. The pace of progress on most other skills is steep, but writing well feels orthogonal to most of them. I do not think writing is just ignored, but rather it’s challenging and lacks good training data to specifically intervene on it. There is certainly some low-hanging fruit for making AI models better at writing — such as specialized harnesses like Claude Code, prompts, and training environments that make models spend a lot more inference tokens on the output, but I don’t think these will have a multiplicative impact on ability. Writing well is a very hard task! It’s a shame that we haven’t unlocked inference-time scaling for one of the great intellectual pursuits. Regardless, writing seems very different than what the models are good at.

supports · 2

04
Claim

Stagnation in long-form nonfiction writing is an alarming sign for expectations that AI will autonomously solve grand, open scientific problems soon.

Current models struggle to organize and compellingly present even well-established science. Until they overcome this limitation, scientific progress from LLMs is likely to remain concentrated on low-hanging fruit and connections between distant fields rather than revolutionary insight.

transcript

the author: Models being stagnant in long-form, non-fiction writing should be alarming to those reliant on models autonomously solving grand, open science problems in the near future. The models today struggle to organize and compellingly present some of the most established science in their area. This seems like a natural prerequisite that we should expect the models to master before they can solve broad, open-ended problems on their own. Until this is solved, the progress of LLMs for science will look closer to solving low-hanging fruit and merging distant connections across fields, rather than any sort of revolutionary insight.

05
Prediction

The best textbooks will remain heavily crafted by humans for at least two to five years because AI currently saves only 10–20% of the effort and excels at transforming existing knowledge rather than creating its core insight.

The author believes AI can help more experts share knowledge and can adapt content to students, but it cannot yet create the organizing insight and skeleton of a textbook. The expected near-term result is a frustrating local minimum in which average effort declines even as the best work remains human-crafted.

transcript

the author: The crux of the above paragraph and preceding section is that I would be happy if more of the world’s experts used AI models to write a tiny bit of their books in order to get more of their knowledge shared with the world. The problem is that you can only use AI models to save 10-20% of the effort today, and I don’t see that percentage becoming the majority anytime soon. There’s also the social pressure, where people expect LLMs to be the best, personalized educators out there, so they think working on a book or educational content is pointless. I think some of these opinions are aging out, as there’s a massive dearth in the highest quality educational work — and there always has been. AI is great at manipulating said content into the form that suits the student, not creating the content from scratch. In the meantime I feel that we are stuck in a frustrating local minimum, where AI models are going to on net reduce the average effort spent on non-fiction writing, but they could enable great expression. Fewer people will start and push through. So, in 2-5 years I still expect the best textbooks to be heavily crafted by the human hand. I’m not sure after then, but that’s longer than many would’ve predicted, given just how much knowledge these models have and their structural propensity to stream it.

supports · 3

06
Mechanism

Today's LLMs increase entropy rather than compressing and organizing knowledge in long-form nonfiction, so they will remain dependent on human guidance.

The author argues that organizing knowledge is a form of compression necessary for insight, but current models make long-form nonfiction more disorganized. This prevents their outputs from being endlessly layered into increasingly reliable knowledge without human guides.

transcript

the author: Organizing knowledge is a compression. This compression is needed to make insight. Today’s LLMs increase entropy in long-form non-fiction writing, and I don’t see how that can be stacked on top of itself endlessly. They’ll be reliant on humans acting as sort of guides.

explains mechanism · 2rebuts · 2supports · 1

Highlight slides
长篇非虚构写作停滞,是危险信号✦ from: Stagnation in long-form nonfiction writing is an alarming sign for expectations that AI will autonomously solve grand, open scientific problems soon.Best Textbooks Will Remain Human-Crafted✦ from: The best textbooks will remain heavily crafted by humans for at least two to five years because AI currently saves only 10–20% of the effort and excels at transforming existing knowledge rather than creating its core insight.Current LLMs Increase Entropy✦ from: Today's LLMs increase entropy rather than compressing and organizing knowledge in long-form nonfiction, so they will remain dependent on human guidance.Human Guidance Remains Necessary✦ from: Today's LLMs increase entropy rather than compressing and organizing knowledge in long-form nonfiction, so they will remain dependent on human guidance.科学进展或仍停留在低垂果实✦ from: Stagnation in long-form nonfiction writing is an alarming sign for expectations that AI will autonomously solve grand, open scientific problems soon.AI’s Near-Term Educational Role✦ from: The best textbooks will remain heavily crafted by humans for at least two to five years because AI currently saves only 10–20% of the effort and excels at transforming existing knowledge rather than creating its core insight.A Frustrating Local Minimum✦ from: The best textbooks will remain heavily crafted by humans for at least two to five years because AI currently saves only 10–20% of the effort and excels at transforming existing knowledge rather than creating its core insight.
Related episodes