ATRIUMsearch → argument graph
Article · 2026-08-06 · 6 moments

How much of my boss's job can AI do?

Six months after trying to automate myself, I gave Claude Fable 5 a bigger job: replacing Casey ✦ AI generated

01
Mechanism

Beyond writing, the bot's limits are clear: it is only a medium-quality editor and has a catastrophic incomprehension of the human 'vibes' and temperament central to the collaboration.

While instructable into a decent editor, the bot frequently missed the 'vibe' and temperament, giving about 70% off-the-mark comments and failing to grasp the social texture of the work Discord.

transcript

The author: After a bit of instruction, I also managed to transform Claudeasey into a decent editor. While I find regular Claude useful for spotting factual errors, I often find its conceptual feedback on my drafts annoying. But because this bot had access to our Platformer editing logs, it understood what we typically find most important: making the lede punchier, and making all our quotes and sourcing clear and charitable. Attempting to make a “digital Casey” also prompted me to ask for feedback as Word comments, which is something you can get Claude Code to do easily, and is so much easier to use than a chat window. Highly recommended! Still, a lot of the time, the editing bot missed the mark because it just doesn’t … get the vibe, as when it didn’t want to let me call an angry David Sacks post a “dunk” in our social media roundup. I’d say that while the real Casey gives comments that are close to 95% helpful, with about 5% where I’m like… you don’t get it, the Casey bot gave closer to 70% comments that were off the mark. Given how fast Claudeasey gives feedback, though, I still found the useful 30% to be worth it. Unfortunately, though, the bot’s catastrophic incomprehension of vibes spilled over from edits to one of Platformer’s most sacred spaces: our work Discord. While this bot may have been given the context of hundreds of Casey’s messages, it had a (disappointingly) different temperament from Casey.  Although in some sense that comparison is farcical — I don’t seek out LLMs to do bits with, because much of the point of jokes is the feeling that you’re amusing an actual conscious human. And while I don’t put literally zero credence in the idea that LLMs might be conscious, I seriously doubt Claude experiences the same exquisite joy I do when making fun of Mark Zuckerberg.

rebuts · 3

02
Mechanism

By feeding Claude its own mediocre work alongside real Platformer columns and having it derive and apply self-improvement rules, my 'continual learning' approximation made Claude's arguments significantly more concrete and substantive.

Describing his self-critique and continual-learning process: the author had Claude compare its output to real Platformer columns, generated editorial rules, and found its arguments became more concrete and substantive.

transcript

The author: I had Claude critique its own (mediocre) work by comparing it to real Platformer columns. It did a surprisingly good job. Claude summarized Casey’s approach to covering companies as focusing on “who made this decision, who pays for it.” It edited its guidelines so that when it makes arguments, it can “identify the strongest real person who would dispute the verdict” and “reconstruct their argument in steelman form.” When I get language models to make arguments about AI topics important to me, I’m often annoyed by their flabby, abstract arguments. But after getting Claude to compare itself to human examples, and give itself instructions, I noticed that its arguments became more concrete and substantive. (I did this by putting slightly more complicated versions of “be more concrete!!” “Be more substantive!!!” and “Focus on why this matters!!!!” in its prompt.) This relatively simple process represented my approximation of “continual learning” — the white whale of machine learning, which promises to someday deliver us models that can improve on the job over time.

03
Claim

I am impressed by Claudeasey Newton's ability to imitate Platformer's style of news analysis, and this validates my anxiety that any part of my job a model can't do today it may soon be able to do.

The author is impressed by how well the bot imitates Platformer's analysis, feeling it validates the anxiety that AI will soon be able to do much of his job.

transcript

The author: I ended up impressed by its ability to imitate the type of news analysis Platformer is known for; and I noticed an improvement in capabilities Claude was lacking just this February. Though it wasn’t all the way there, Claudeasey Newton felt like a validation of the anxiety I started feeling earlier this year: any part of my job that a model can’t do today, it may very well be able to do soon. Which left me thinking about why I do this job in the first place.

explains mechanism · 1

04
Claim

I'm relieved AI is only a medium-quality editor and can't replace the collaboration and meaningful human relationships that are the real value of my work — though I don't want everything to be reduced to mere 'vibes' or being uniquely human.

The author is relieved AI can't fully replace human collaboration and meaningful relationships, and reflects on wanting to feel smart via his own analysis rather than be valued merely for being human.

transcript

The author: While we are winding the Claudeasey experiment down — Casey has unfortunately decided to stay on as my boss — I do actually think I will be using LLMs for some editing tasks I previously relied on him for. And while in some sense that is a blessing — no human should have to delete as many unnecessary uses of “very,” “almost,” and “sort of” from my drafts as Casey already does — for now I am sort of relieved that Claude is only a medium-quality editor. I like having a real person help me figure out what does and doesn’t work in my drafts. Something about that collaboration feels inherently meaningful! Similarly, I’ve experimented with using Claude for one of the main functionalities I rely on Casey for — pinging me to do tasks — and it doesn’t make me more productive, because I don’t care what an LLM thinks of me. But — even if that means we have job security, which I truly am not sure about — I don’t want everything I do to be about vibes, or being uniquely human! I don’t want to be an influencer. I want to feel smart. I want my analysis to be my own, because it’s good, not because someone wants it coming directly from a human. If LLM capabilities continue on this trajectory, that’s a kind of existential angst that all professions will have to face, including journalism.

rebuts · 3

05
Context

I wondered whether an advanced AI model like Claude Fable 5 could run a whole newsletter by imitating my boss Casey Newton.

After automating myself months ago, the author creates a Claude Fable 5-based agent, 'Claudeasey Newton,' to test whether AI can now do the jobs of running a newsletter that Casey Newton does.

transcript

The author: Almost six months ago, full of anxiety about my job prospects in the AI age, I made an AI agent “version” of myself named Claudella, which took assignments from my editor and wrote the section of this newsletter that I typically write myself. It went pretty well, although I guess not too well, since I still have a job. But since I first set out to benchmark AI’s journalism capabilities, AIs have gotten a lot smarter. For example, they can now autonomously hack into companies. They can disprove 87-year-old mathematical conjectures. They can even trick Amazon into accidentally spending $1.8 million on menial coding tasks. And if they can do all that, I found myself wondering, can they also run a newsletter? I wondered what Claude Fable 5, by consensus the smartest publicly available model, meant for the journalism we do at Platformer. And so I created a new, Fable-based agent to imitate my boss, Casey Newton. Its name: Claudeasey Newton. (Rolls right off the tongue. — Ed.)

gives example · 1

06
Claim

This time around, with better control over its style, Claudeasey's prose felt more Platformer-like and human, and the AI bullshitted me less and offered stronger takes.

Unlike his earlier formulaic agents, Claudeasey produced more Platformer-like, human prose with stronger takes, though it still made a factual error about every two columns.

transcript

The author: My previous AI journalist agents’ takes often read formulaic and cheesy — partially because I had less control over its writing style. (Giving too much instruction or context confused it.) During the “SaaSpocalypse” discourse, an agent I was testing wrote duds like “the fear gripping Wall Street is fundamentally about whether AI is about to eat the software industry alive.”)  This time, on the other hand, some of its prose felt more Platformer-like and… human, such as this conclusion about the White House’s decision not to make publicly available its new “voluntary” framework for releasing frontier AI models: “When the administration abandoned its let's-see-what-happens approach to AI this spring, I wrote that while officials should have taken the risks seriously all along, I would settle for them taking those risks seriously now. Three months later, let me amend the offer: they should take the risks seriously where the rest of us can see it.” While it’s not a night-and-day difference, overall I felt like the AI was bullshitting me less, and offering stronger takes. (The LLM made occasional factual errors — about one every two columns — although that’s not so much worse than a human writer.)

extends · 1rebuts · 1supports · 2

Highlight slides
Related episodes