ATRIUMsearch → argument graph
Article · 2026-07-22 · 4 moments

Are AI labs pelicanmaxxing?

Are AI labs pelicanmaxxing? Excellent piece of work by Dylan Castillo, who took a deep-dive into the frequently pondered question of whether the AI labs have been deliberately training models to draw pelicans riding bicycles in response to my deeply unscientific benchmark. I've been randomly spot-checking this in the past by testing models against other animals riding other types of vehicle, but never with anything close to the diligence of Dylan's methodology here. Dylan took 8 animals × 6 v ✦ AI generated

01
Context

Dylan Castillo conducted a rigorous study of whether AI labs have been deliberately training models to draw pelicans riding bicycles.

The author highlights Dylan Castillo's methodical investigation into the 'pelicanmaxxing' question, contrasting it with the author's own ad-hoc spot-checking.

transcript

the author (via the article's title/description, the voice is an anonymous blogger or commentator): Excellent piece of work by Dylan Castillo, who took a deep-dive into the frequently pondered question of whether the AI labs have been deliberately training models to draw pelicans riding bicycles in response to my deeply unscientific benchmark.

provides context · 2

02
Data

Dylan's methodology tested 8 animals × 6 vehicles across 7 different models with repeated runs and automated evaluation.

The author describes the scale and rigor of Castillo's experimental design: 48 prompts, three runs each, seven models, plus AI-assisted evaluation.

transcript

the author (via the article's title/description, the voice is an anonymous blogger or commentator): Dylan took 8 animals × 6 vehicles = 48 prompts and ran them three times each through 7 different models ( GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro). He then used GPT-5.6 Luna and Gemini 3.1 Flash-Lite to help evaluate the results.

provides context · 1

03
Data

GLM-5.2 showed the largest individual boost on the exact pelican-bicycle combination, but the effect was small and not statistically significant.

The author notes GLM-5.2 showed the largest improvement on the specific pelican-bicycle test case, but cautions the effect is small and not significant.

transcript

the author (via the article's title/description, the voice is an anonymous blogger or commentator): GLM-5.2 comes closest: it has the largest boost on the exact pelican-bicycle cell, and its first pelican-on-bicycle sample caught my eye. But the effect is small and not significant, so I wouldn't put too much weight on it.

gives example · 1

04
Claim

The study found no evidence of 'pelicanmaxxing' — AI labs are not deliberately biasing models toward better pelican-on-bicycle drawings.

Across all models tested, the study found no evidence that labs are deliberately improving pelican-on-bicycle generation quality.

transcript

the author (via the article's title/description, the voice is an anonymous blogger or commentator): For the models he tested he could find no evidence of pelimaxxing: The pelicans on bicycles don't look any better. Labs are not better at drawing pelicans. Labs are not better at drawing bicycles. Labs are not better at drawing pelicans on bicycles, even adjusting for difficulty.

supports · 2

Highlight slides
Related episodes