Despite being a flawed comparative benchmark, the pelican test remains valuable to Willison mainly as a forcing function that ensures he actually runs a prompt through each new model.
Willison argues the real value of the pelican test isn't as a rigorous benchmark but as a forcing function that gets him to actually try each new model.
transcript
Simon Willison: Firstly, it's a forcing function for actually trying the model. If I show you a pelican, that means I've managed to run a prompt through it. If the model has an official API I'll use that, if it's open weight... I'll try running it on my own machine.