DataArticle
Laguna S 2.1 is the fastest 100B+ model I've tested with the best tool calling, but it invents facts under pressure.
A private agentic eval comparing Laguna-S-2.1 against Qwen3.5-122B on an RTX Pro 6000 found Laguna faster (109 tok/s vs 103) and better at tool mechanics, but weaker on grounding with 3 confirmed fabrications vs Qwen's 0. ✦ AI generated
Reddit (r/LocalLlama commenter) · Latent Space · 2026-07-23 · original ↗
The image is a technical benchmark chart from a private agentic eval comparing Laguna-S-2.1 118B-A8B vs Qwen3.5-122B on a single RTX Pro 6000 96GB under vLLM with NVFP4 weights and FP8 KV at 256k context. It visualizes the post's main finding: Laguna is faster and stronger at tool mechanics—109 tok/s vs Qwen's 103 tok/s, slightly better tool-call args, no JSON/streaming errors, deeper tool chains—but is weaker on grounding and breadth, especially sports/odds knowledge and 'grounding under pressure,' where the author reports 3 confirmed fabrications versus Qwen's 0.
Read full article ↗excerpt · fair-use quotation
Around this claim