ATRIUMsearch → argument graph
DataArticle

Dylan's methodology tested 8 animals × 6 vehicles across 7 different models with repeated runs and automated evaluation.

The author describes the scale and rigor of Castillo's experimental design: 48 prompts, three runs each, seven models, plus AI-assisted evaluation. ✦ AI generated

the author (via the article's title/description, the voice is an anonymous blogger or commentator) · Simon Willison's Weblog · 2026-07-22 · original ↗

Dylan took 8 animals × 6 vehicles = 48 prompts and ran them three times each through 7 different models ( GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro). He then used GPT-5.6 Luna and Gemini 3.1 Flash-Lite to help evaluate the results.

Read full article ↗excerpt · fair-use quotation

Around this claim