ATRIUMsearch → argument graph
ClaimVideo · 4:59 — 6:15

The percentage framing of what models can and cannot do is totally off; we need to think about continuous uncapped rewards for workflows like legal arguments and medical advice where you could always get better.

Oswald argues that binary percentage-based thinking about AI capability misses the point — for domains like legal and medical reasoning, there is no ceiling, and the right framework is continuous improvement. ✦ AI generated

Oswald Nitski · 20VC · 2026-07-25 · original ↗

starts at this moment · 4:59

Elicited by

When you look at that dispersion, what do you think is inaccurate? You said you didn't really believe the 90/10. What what do you believe a more accurate representation is?

Well, in our Apex benchmarks, we're getting closer to around 50% of long horizon workflows. Top models are scoring around that much. But I think that there's a class of workflows that are just sufficiency-based where you do it and it's done and you're good. This is something like updating a CRM. You couldn't really get much better at it. And then there's a class of workflows that we shouldn't even be thinking about in terms of binary, like, can the models do it or not. And these can be things like legal arguments or to an extent medical advice, where you could always get better. And in those cases, I think that the percentage framing is just totally off and we need to be thinking more about continuous uncapped rewards.

verbatim transcript · starts at 4:59

Transcript · around this moment

4:59models is that the uh the inference can happen in multiple places. So, um you could be you could make mistakes using them, but um you have more control. >> When you look at that dispersion, what do you think is inaccurate? You said you didn't really believe the 90/10. What what do you believe a more accurate representation is? >> Well, in our uh Apex uh benchmarks, we're getting closer to around uh 50% um

5:24of long horizon workflows. Um top models are scoring around around around that much. But, I think that um the percentage for there's [snorts] a there's a class of workflows that are just sufficiency-based where you do it and it's done and you're you're good. This is something like updating a CRM. Um you couldn't really get much better at it. And then there's a class of workflows that we shouldn't even be

5:50thinking about in terms of, you know, binary, like, can the models do it or not. Um and these can be things like legal arguments or uh to an extent medical advice, where you could always get better. Um and in those cases, I think that the percentage framing is is just totally off and we need to be thinking more about continuous uncapped rewards. >> When we think about it could be better,

6:15I had Lynn Qual, the founder of Filecoin, on the show the other day, and she was like, "Exactly that is why we'll have specialized models for every single company." Because it could be better, depends entirely on the company. One company wants to focus on growth, one company wants to focus on margin, another wants to focus on I don't know, if we're in Europe, uh work-life balance. Um

6:39and and so you need individual specialized models for every company. Do you buy that we will have specialized models for every company, or is that a little bit self-serving towards Filecoin? [laughter] >> Uh I buy it. I think it's also self-serving towards Mekor um in that we think that every specialized model will need uh enterprise uh specific eval training data to show the model how to perform in

Around this claim