ATRIUMsearch → argument graph
ClaimVideo · 46:09 — 47:39

AI benchmark scores have become meaningless white noise that don't reflect real quality; qualitative feedback and usage metrics like team-wide token consumption are far more reliable signals of whether a model is actually good.

Dax says he no longer looks at benchmarks since the numbers going up feels meaningless, and instead trusts qualitative feedback and rising token usage as the real indicator that a model is working for people. ✦ AI generated

Dax Raad · Syntax · 2026-07-15 · original ↗

starts at this moment · 46:09

Elicited by

do you think we'll ever get to like a a benchmark that makes a lot of sense?

at this point I don't think I look at benchmarks at all. I don't know if I ever really did. I don't think anyone ever really did. I think they got they just kind of became white noise at some point. We all know that the numbers go up. Like congratulations number went up, you know.

verbatim transcript · starts at 46:09

Transcript · around this moment

46:09we'll ever get to like a a benchmark that makes a lot of sense? >> Yeah, at this point I don't think I look at benchmarks at all. I don't know if I ever really did. I don't think anyone ever really did. I think they got they just kind of became white noise at some point. We all know that the numbers go up. Like congratulations number went up, you

46:26know. >> [laughter] >> And it's like it went up more than the other guy, but then when the other guy releases their number goes up more. So it's it's I don't know if that means anything to us. So I just look at I mean I love qualitative feedback. I love people being like sharing what they've been able to do or what they've built. We have it you know, you obviously don't

46:47get that at like a million data point scale, but these are like products at the end of the day and they're kind of fuzzy. And it comes down to are people happy or people frustrated? That's why I like looking at token usage on our team. Uh just this cuz cuz to me that's like our if we if I see it go up, that means something is working. Like they're

47:10they're liking something. Yeah, so yeah, I think I I still rely on a lot of qualitative feedback and people kind of demonstrating how they how they're using it. Uh Given our team is, you know, within our team we have cloud fans, we have GPT fans, we've got some open source model fans. So we have okay coverage and visibility and all that. >> Yeah. So Dax, this is your second time

47:34on the show, so you know the drill. We have sick picks and shameless plugs on this show. Sick pick being anything that you're just enjoying in life right now. Do you have something that you would like to share as a sick pick? >> Yeah, I mean it's got to be that thing I mentioned earlier, uh exa.dev. I think uh like yeah, if you want to experiment with this machine in the cloud thing,

47:55um it's a really clever product. It's put together really well. Like if you use Tailscale, you know how good Tailscale is. Like Tailscale just like freaking works. Um and this has like that same vibe to it. And And I I really love like businesses that are like I think a lot about positioning and this product is it's like positioned in like a weird vacuum that existed. Like

Around this claim