The pelican benchmark's former correlation with actual model quality has broken down, since GLM-5.2's pelican outclasses those from GPT-5.6 and Claude Fable 5 despite GLM not being a frontier-class model.
Willison says the pelican test's early correlation with model quality is now mostly gone, since a weaker model (GLM-5.2) draws a better pelican than top models. ✦ AI generated
That connection has been mostly severed now. The GPT-5.6 and Claude Fable 5 pelicans are outclassed by GLM-5.2, and much as I love GLM I don't think that's a Fable-class model.