Data◆Article
GLM-5.2 showed the largest individual boost on the exact pelican-bicycle combination, but the effect was small and not statistically significant.
The author notes GLM-5.2 showed the largest improvement on the specific pelican-bicycle test case, but cautions the effect is small and not significant. ✦ AI generated
the author (via the article's title/description, the voice is an anonymous blogger or commentator) · Simon Willison's Weblog · 2026-07-22 · original ↗
GLM-5.2 comes closest: it has the largest boost on the exact pelican-bicycle cell, and its first pelican-on-bicycle sample caught my eye. But the effect is small and not significant, so I wouldn't put too much weight on it.
Read full article ↗excerpt · fair-use quotation
- ·GLM-5.2 has the largest boost on the exact pelican-bicycle cell
- ·Its first pelican-on-bicycle sample stood out to the author
- ·Effect is small and not statistically significant
- ·The improvement is not reliable by statistical standards
- ·Author advises not to put much weight on the result
- ·No other model matched this specific combination
Around this claim
This moment responds to
provides context → Dylan Castillo conducted a rigorous study of whether AI labs have been deliberately training models to draw pelicans riding bicycles.the author (via the article's title/description, the voice is an anonymous blogger or commentator) · Simon Willison's Weblogprovides context → Dylan's methodology tested 8 animals × 6 vehicles across 7 different models with repeated runs and automated evaluation.the author (via the article's title/description, the voice is an anonymous blogger or commentator) · Simon Willison's Weblog