ATRIUMsearch → argument graph
MechanismVideo · 18:36 — 20:06

Domain-specific AI agents, like Codex for coding, already work well because coding is testable and RL-friendly, and this pattern will likely extend to other quantitative knowledge work before general-purpose agents that can handle any task become viable.

Turley says narrow, domain-specific agents (especially coding agents like Codex) have already hit escape velocity because their outputs are verifiable, while general-purpose agents for arbitrary tasks remain the harder, still-unsolved goal. ✦ AI generated

Nick Turley · BG2 Pod · 2026-03-15 · original ↗

starts at this moment · 18:36

Elicited by

is there a shape or ordinality of tasks and or agents that you think, hey, this is the kind of thing that's likely to come first whenever it does?

the thing that's already come first is the domain specific agents, right? If you look at what's happening in code, we're fully there... it's testable, you know, if it worked or not, it's very RL friendly... the domain specific ones already work. I think the thing everyone's working for is general purpose agents that just kind of work for anything.

verbatim transcript · starts at 18:36

Transcript · around this moment

18:16you said on actions and tasks >> got it on timing tough to Okay. But is there a shape or ordinality of tasks and or or agents that that you think, hey, this is the kind of thing that's likely to come first whenever it does? >> I mean, the thing that's already come first is the uh domain specific agents, right? If you look at what's happening in in code,

18:36>> we're we're we're fully there. >> You know, it's it's mindbending, but we've got so many engineers um who who don't open their IDE like ever. And right >> for me as someone who you know used to code and then unfortunately got very very busy it's brought me back in the game. >> So um codeex and you know products like it is clearly a product that has escape

18:56velocity where people are absolutely using it for all kinds of agentic work and if you just take what people are doing >> and make it work even better >> you kind of get all the way there. Um, you know, I won't be surprised if you see this happen for other forms of sort of quantitative knowledge work just because it happens to have the properties that code has. It's testable.

19:19You know, if it worked or not, >> um, it's, you know, very RL friendly. >> Um, >> but um, uh, the the domain specific ones already work. I think the thing everyone's working for is, you know, general purpose agents that just >> kind of work for anything. >> Yeah. And I you know that's why I think you need to to win a consumer because it's very hard to train people into like

19:41okay >> um it can work like deep research was a consumer product and it's really was our first agentic >> uh thing out there >> um but I think what consumers want is I can just ask it anything and we'll do what >> um what needs to be done without any sort of retraining and um >> we'll get there just a matter of time >> at least a psychological goal is uh

20:03flight bookings bookings restaurant bookings, shopping, >> all this stuff. Um there are so many consumer problems and those are just the type of things that you would kick off, right? >> Yeah. >> The minute you have productivity, there's there's things you don't even think of as agentic tasks. >> Um like you're trying to get in shape. You don't think of that as a task you would delegate.

20:22>> Um unless you have a trainer, in which case you do, but most people don't, right? But if the EI knew that, it could totally start working in the background for you over very long periods of time and and getting you, you know, um, here's your fitness plan. Okay, I actually signed you up for this thing. You could imagine it being quite helpful if it's aligned with your with your

Around this claim