The percentage framing of what models can and cannot do is totally off; we need to think about continuous uncapped rewards for workflows like legal arguments and medical advice where you could always get better.
Oswald argues that binary percentage-based thinking about AI capability misses the point — for domains like legal and medical reasoning, there is no ceiling, and the right framework is continuous improvement. ✦ AI generated
Oswald Nitski · 20VC · 2026-07-25 · original ↗
starts at this moment · 4:59
“When you look at that dispersion, what do you think is inaccurate? You said you didn't really believe the 90/10. What what do you believe a more accurate representation is?”
Well, in our Apex benchmarks, we're getting closer to around 50% of long horizon workflows. Top models are scoring around that much. But I think that there's a class of workflows that are just sufficiency-based where you do it and it's done and you're good. This is something like updating a CRM. You couldn't really get much better at it. And then there's a class of workflows that we shouldn't even be thinking about in terms of binary, like, can the models do it or not. And these can be things like legal arguments or to an extent medical advice, where you could always get better. And in those cases, I think that the percentage framing is just totally off and we need to be thinking more about continuous uncapped rewards.
verbatim transcript · starts at 4:59
4:59models is that the uh the inference can happen in multiple places. So, um you could be you could make mistakes using them, but um you have more control. >> When you look at that dispersion, what do you think is inaccurate? You said you didn't really believe the 90/10. What what do you believe a more accurate representation is? >> Well, in our uh Apex uh benchmarks, we're getting closer to around uh 50% um
5:24of long horizon workflows. Um top models are scoring around around around that much. But, I think that um the percentage for there's [snorts] a there's a class of workflows that are just sufficiency-based where you do it and it's done and you're you're good. This is something like updating a CRM. Um you couldn't really get much better at it. And then there's a class of workflows that we shouldn't even be
5:50thinking about in terms of, you know, binary, like, can the models do it or not. Um and these can be things like legal arguments or uh to an extent medical advice, where you could always get better. Um and in those cases, I think that the percentage framing is is just totally off and we need to be thinking more about continuous uncapped rewards. >> When we think about it could be better,
6:15I had Lynn Qual, the founder of Filecoin, on the show the other day, and she was like, "Exactly that is why we'll have specialized models for every single company." Because it could be better, depends entirely on the company. One company wants to focus on growth, one company wants to focus on margin, another wants to focus on I don't know, if we're in Europe, uh work-life balance. Um
6:39and and so you need individual specialized models for every company. Do you buy that we will have specialized models for every company, or is that a little bit self-serving towards Filecoin? [laughter] >> Uh I buy it. I think it's also self-serving towards Mekor um in that we think that every specialized model will need uh enterprise uh specific eval training data to show the model how to perform in
- ·Top models score ~50% on long-horizon workflows (Apex benchmarks)
- ·Sufficiency-based tasks like CRM updates have a ceiling
- ·Legal arguments and medical advice have no ceiling — always improvable
- ·Percentage framing is 'totally off' for high-stakes domains
- ·Need continuous uncapped rewards, not binary can/can't
- ·Infinite improvement curve for legal and medical reasoning
- ·Stop asking whether models can do a task
- ·Start asking how much better the output could be
- ·The ceiling doesn't exist for expert reasoning domains