DataArticle
Dylan's methodology tested 8 animals × 6 vehicles across 7 different models with repeated runs and automated evaluation.
The author describes the scale and rigor of Castillo's experimental design: 48 prompts, three runs each, seven models, plus AI-assisted evaluation. ✦ AI generated
the author (via the article's title/description, the voice is an anonymous blogger or commentator) · Simon Willison's Weblog · 2026-07-22 · original ↗
Dylan took 8 animals × 6 vehicles = 48 prompts and ran them three times each through 7 different models ( GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro). He then used GPT-5.6 Luna and Gemini 3.1 Flash-Lite to help evaluate the results.
Read full article ↗excerpt · fair-use quotation
Around this claim
This moment responds to
provides context → Dylan Castillo conducted a rigorous study of whether AI labs have been deliberately training models to draw pelicans riding bicycles.the author (via the article's title/description, the voice is an anonymous blogger or commentator) · Simon Willison's Weblogsupports → The study found no evidence of 'pelicanmaxxing' — AI labs are not deliberately biasing models toward better pelican-on-bicycle drawings.the author (via the article's title/description, the voice is an anonymous blogger or commentator) · Simon Willison's Weblog