A core, still-unsolved trade-off in building GPT-5 is between how long the model reasons (reasoning tokens/thinking time) and response latency versus intelligence — models get dramatically smarter given more inference time, which conflicts with product needs for fast responses.
Sherwin Wu identifies the reasoning-time-versus-latency trade-off as the hardest and still-unresolved design decision behind GPT-5, since more thinking time makes the model dramatically smarter but users won't always wait minutes for an answer. ✦ AI generated
Sherwin Wu · BG2 Pod · 2025-09-11 · original ↗
starts at this moment · 27:38
“Are there trade-offs that you made as you were going through it? Maybe what are the hardest trade-offs you made as you were building GPT-5?”
I actually think a very clear trade-off, which I honestly think we are still iterating on, is the trade-off between the reasoning tokens and how long it thinks versus performance. And honestly, this is something that I think we've been working on with our customers since the launch of the reasoning models, which is these models are so, so smart, especially if you give it all this thinking time.
verbatim transcript · starts at 27:38
27:38a model which is extremely intelligent, but an exquisite academic, essentially, way. Are there trade-offs that you made as you were going through it? Maybe what are the hardest trade-offs you made as you were building GPT-5? I actually think a very clear trade-off, which I honestly think we are still iterating on, is the trade-off between the reasoning tokens and how long it thinks versus performance. And honestly, this is something that I think we've been working on with our
28:04customers since the launch of the reasoning models, which is these models are so, so smart, especially if you give it all this thinking time. I think the feedback I've been seeing around GPT-5 Pro has been pretty crazy, too. It's just like these unsolved-- Andrej had a great tweet last night. Yeah, I saw that Sam retweeted it. But these unsolved problems that none of the other models could handle, you throw to GPT-5 Pro and it just one-shots it, it's pretty crazy. But the
28:30trade-off here is you're waiting for 10 minutes. It's quite a long time. And so these things just get so smart with more inference time. But on the product builder on the API side for some of these business use cases, I think it's pretty tough to manage that trade-off. And for us, it's been difficult to figure out where we want to fall on that spectrum. So we've had to make
28:50some trade-offs on how much of the model think versus how intelligent should it get. Because as a product builder, there's a real latency trade-off that you have to deal with where your user might not be happy waiting 10 minutes for the best answer in the world. It might be more okay with the substandard answer in no wait at all. Yeah, I mean even between GPT-5 and GPT-5 thinking,
29:09I have to toggle it now because sometimes I'm so impatient I just want it ASAP. I think there's an ability to skip, right? Yeah, that's right. And GPT where it's like I'm impatient, I just want a more simple answer. That's right, that's right. Well, four weeks in, GPT-5, how's the feedback? Yeah, I think feedback has been very positive, especially on the platform side, which has been really great to see. I think a lot of the things that Olivier mentioned have been,
29:33you know, coming up in feedback from customers. The model is extremely good at coding, extremely good at kind of like reasoning through different tasks. But especially for like coding use cases, especially at the, you know, when it thinks for a while, it'll usually solve problems that no other models can solve. So I think that's been a big positive point of feedback. The kind of robustness and the reduction in hallucinations has been a really big positive feedback. Yeah, yeah,