ATRIUMsearch → argument graph
ClaimArticle

The pelican benchmark's biggest limitation is that it says nothing about agentic tool calling and reliable tool use over long conversations, which is what matters most for today's models.

Willison points out the pelican test entirely misses agentic tool-calling ability, now the most important quality for evaluating models. ✦ AI generated

Simon Willison · Simon Willison's Weblog · 2026-07-16 · original ↗

The biggest limitation of the pelican is that it doesn't touch at all on the thing that matters most for today's model: agentic tool calling and the ability to operate tools reliably as conversations grow in length.

Read full article ↗excerpt · fair-use quotation

Related moments