ATRIUMsearch → argument graph
ClaimAudio · 3:33 — 3:58

When agent work doesn't succeed, it feels like a skill issue — you just haven't found the right instructions or tools, not that the capability is absent.

Karpathy explains that when agents fail, it feels like a skill issue in how you prompted or set things up, not a fundamental capability gap. ✦ AI generated

Andrej Karpathy · No Priors · 2026-03-20 · original ↗

plays this moment only · 3:33 — 3:58

Elicited by

What is it limited by?

Just I think everything, so many things, even if they don't work, I think to a large extent you feel like it's a skill issue. It's not that the capability is not there. It's that you just haven't found a way to string it together of what's available. Like I just don't, I didn't give good enough instructions in the agent's MD file or whatever it may be. I don't have a nice enough memory tool that I put in there or something like that. So it all kind of feels like a skill issue when it doesn't work to some extent.

verbatim transcript · starts at 3:33

Transcript · around this moment

(00:00:00) Code's not even the right verb anymore, right? (00:00:02) But I have to express my will to my agents for 16 hours a day. (00:00:07) Manifest. (00:00:08) How can I have not just a single session of clot code or codecs or some of these agent harnesses? (00:00:12) How can I have more of them? (00:00:13) How can I do that appropriately? (00:00:15) The agent part is now taken for granted. (00:00:17) Now the claw-like entities are taken for granted. (00:00:19) And now you can have multiple of them. (00:00:20) And now you can have instructions to them. (00:00:21) And now you can have optimization over the instructions. (00:00:24) But I mean, this is why it gets to the psychosis is that this is like infinite and everything is skill issue. (00:00:35) Hi, listeners. (00:00:36) Welcome back to No Priors. (00:00:37) Today I'm here with Andre Karpathy, and we have a wide-ranging conversation for you about code agents, the future of engineering and AI research, how more people can contribute to research, what's happening in robotics, his prediction for how agents can reach out into the real world, and education in this next age. (00:00:55) Welcome, Andre. (00:00:56) Andre, thanks for doing this. (00:00:57) Yeah, thank you for having me. (00:00:59) So it's been a very exciting couple of months in AI. (00:01:02) Oh yeah, you could say that. (00:01:04) I remember walking into the office at some point and you were like really locked in and I was asking what you were up to and you're like, I just, I have to code for 16 hours a day or code's not even the right verb anymore, right? (00:01:15) But I have to express my will to my agents for 16 hours a day. (00:01:20) Manifest. (00:01:22) Because like there's been a jump in capability. (00:01:26) What's happening? (00:01:27) Tell me about your experience. (00:01:28) Yeah, I kind of feel like I was just in this perpetual, I still am often in this state of AI psychosis, just like all the time, because there was a huge unlock in what you can achieve as a person, as an individual, right? (00:01:38) Because you were bottlenecked by, you know, your typing speed and so on. (00:01:41) But now with these agents, it really, I would say in December is when it really just, (00:01:45) something flipped where I kind of went from 80-20 of like, to 20-80 of writing code by myself versus just delegating to agents. (00:01:53) And I don't even think it's 20-80 by now. (00:01:54) I think it's a lot more than that. (00:01:55) I don't think I've typed like a line of code probably since December, basically, which is like an extremely large change. (00:02:05) I was talking to it, like for example, (00:02:07) I was talking about it too, for example, my parents and so on. (00:02:09) And I don't think like a normal person actually realizes that this happened or how dramatic it was. (00:02:13) Like literally, like if you just find a random software engineer or something like that at their desk and what they're doing, like their default workflow of building software is completely different as of basically December. (00:02:25) So (00:02:26) I'm just like in the state of psychosis of trying to figure out like what's possible, trying to push it to the limit. (00:02:31) How is it, how can I have not just a single session of, you know, clock code or codecs or some of these agent harnesses? (00:02:36) How can I have more of them? (00:02:37) How can I do that appropriately? (00:02:40) And then how can I use these claws? (00:02:41) What are these claws? (00:02:43) And so there's like a lot of new things. (00:02:46) I want to be at the forefront of it, you know, and I'm very, (00:02:48) and see that I'm not at the forefront of it. (00:02:50) And I see lots of people on Twitter doing all kinds of things and they all sound like really good ideas and I need to be at the forefront or I feel extremely nervous. (00:02:56) And so I guess I'm just in the psychosis of like, what's possible, like, because it's unexplored fundamentally. (00:03:01) Well, if you're nervous, the rest of us are nervous. (00:03:03) We have a team that we work with at Conviction that their setup is everybody is like, you know, none of the engineers write code by hand. (00:03:12) And they're all microphoned and they just like whisper to their agents all the time. (00:03:16) It's the strangest work setting ever. (00:03:19) And I thought they were crazy. (00:03:20) And now I like, I fully accept, I was like, this was the way. (00:03:23) Like you're just ahead of it. (00:03:25) What, how do you think about your own capacity now to like explore or to do projects? (00:03:31) Like what is it limited by? (00:03:33) what is it limited by? (00:03:34) Just I think everything, so many things, even if they don't work, I think to a large extent you feel like it's a skill issue. (00:03:40) It's not that the capability is not there. (00:03:41) It's that you just haven't found a way to string it together of what's available. (00:03:45) Like I just don't, I didn't give good enough instructions in the agent's MD file or whatever it may be. (00:03:51) I don't have a nice enough memory tool that I put in there or something like that. (00:03:55) So it all kind of feels like a skill issue when it doesn't work to some extent. (00:03:58) You want to see how you can paralyze them, et cetera. (00:04:00) And you want to be Peter Steinberg, basically. (00:04:02) So Peter is (00:04:03) He has a funny photo where he's in front of a monitor with lots of like he uses codecs, so lots of codecs agents styling the monitor, and they all take about 20 minutes if you prompt them correctly and use the high effort, and so they all take about 20 minutes to have multiple, you know, 10 repos checked out, and so he's just... (00:04:20) going between them and giving them more. (00:04:21) It's just like you can move in much larger macro actions. (00:04:25) It's not just like, here's a line of code, here's a new function. (00:04:27) It's like, here's a new functionality and delegate it to agent one. (00:04:30) Here's a new functionality that's not going to interfere with the other one. (00:04:33) Give it to agent two. (00:04:34) and then try to review their work as best as you can, depending on how much you care about that code. (00:04:39) Like what are these macro actions that I can like manipulate my software repository by? (00:04:43) And like another agent is doing some like research, another agent is writing code, another one is coming up with a plan for some new implementation. (00:04:50) And so everything just like happens in these like macro actions over your repository. (00:04:55) And you're just trying to become like really good at it and develop like a muscle memory for it is extremely... (00:05:01) Yeah, it's very rewarding, number one, because it actually works. (00:05:03) But it's also kind of like the new thing to learn. (00:05:05) So that's why, hence the psychosis. (00:05:08) Yeah, I do feel like my instinct is like, whenever I am waiting for an agent to complete something, the obvious thing to do is like, well, I can do more work, right? (00:05:16) Like if I have access to more tokens, then like I should just paralyze, add more tasks. (00:05:21) And so that's very stressful because if you don't feel very bounded by your ability to spend on tokens, then you are the bottleneck in the system that is max capability. (00:05:31) Yeah, if you're not maximizing your subscription, at least, and ideally for multiple agents, like if you run out of the code on Codex, you should switch to Claude or whatnot. (00:05:39) I don't know. (00:05:40) Like that's what I've been trying to do a little bit. (00:05:41) And I feel nervous when I have subscription left over. (00:05:44) That just means I haven't maximized my token throughput. (00:05:47) So I actually kind of experienced this when I was a PhD student. (00:05:49) You would feel nervous when your GPUs are not running. (00:05:51) Like you have GPU capability and you're not maximizing the available flops to you. (00:05:55) But now it's not about flops, it's about tokens. (00:05:57) So what is your token throughput and what token throughput do you command? (00:06:01) I would actually argue that it's very interesting that we had at least 10 years where... (00:06:08) In many engineering tasks, people just, they didn't feel compute bound, right? (00:06:13) And the entire industry feels that now. (00:06:15) They feel like they felt resource bound. (00:06:19) And now that you have this big capability jump, you're like, oh, actually it's not, you know, my ability to access the compute anymore. (00:06:26) Like I'm the binding constraint. (00:06:28) Yeah, it's a skill issue, which is very empowering because, yeah, because you could be getting better. (00:06:33) So that's why I think it's very addictive because there's unlocks when you get better. (00:06:37) Where do you think it goes? (00:06:38) if you just think about like, okay, Andre's iterating and everybody else's for 16 hours a day, getting better at using coding agents, like what does it look like in a year of like you've reached mastery? (00:06:49) Yeah, what does mastery look like, right? (00:06:50) At the end of the year or like two, three years, five years, 10 years, et cetera. (00:06:55) Well, I think everyone is basically interested in like going up the stack. (00:06:58) So I would say, yeah, it's not about a single session with your agent, multiple agents, how do they collaborate and teams and so on. (00:07:04) So everyone's trying to figure out what that looks like. (00:07:07) And then I would say claw is also kind of an interesting direction because it really, when I say a claw, I mean this like layer that kind of takes persistence to a whole new level. (00:07:14) Like it's something that like keeps looping. (00:07:16) It's like, it's not something that you are interactively in the middle of. (00:07:20) It kind of like has its own little sandbox, its own little, you know, it kind of like does stuff on your behalf, even if you're not looking kind of thing. (00:07:28) And then also has maybe more sophisticated memory systems, et cetera, that are not yet implemented in agents. (00:07:32) So OpenClaw has a lot more sophisticated memory, I would say, than what you would get by default, which is just a memory compaction when your context runs out, right? (00:07:40) You think that's the piece that resonated for more users versus like... (00:07:44) perhaps like broader tool access. (00:07:46) For OpenClaw? (00:07:46) Yeah. (00:07:47) There's like, I think there's at least five things. (00:07:49) There's a lot of really good ideas in here. (00:07:50) Yeah, good job, Peter. (00:07:51) I mean, Peter has done a really amazing job. (00:07:53) I saw him recently and I talked to him about it and he's very humble about it, but I think he innovated simultaneously in like 5 different ways and put it all together. (00:08:02) So for example, like the Soul MD document, like he actually really crafted a personality that is kind of compelling and interesting. (00:08:08) And I feel like a lot of the current agents, they don't get this correctly. (00:08:10) I actually think Claude has a pretty good personality. (00:08:12) It feels like a teammate. (00:08:14) and it's excited with you, et cetera. (00:08:16) I would say, for example, Codex is a lot more dry, which is kind of interesting because in ChatGPT, Codex is like a lot more upbeat and highly sycopantic. (00:08:25) But I would say Codex, the coding agent, is very dry. (00:08:27) It doesn't seem to care about what you're creating. (00:08:29) It's kind of like, oh, I implemented it. (00:08:31) It's like, okay, but do you understand what we're building? (00:08:34) It's true. (00:08:35) You know, it doesn't. (00:08:37) And the other thing I would say is, for example, with Claude, I think they dialed the psychophancy fairly well, where when Claude gives me praise, I do feel like I slightly deserve it. (00:08:44) Because sometimes I kind of give it like not very well-formed thoughts and I give it an idea that I don't think is fully baked and it doesn't actually react very strongly. (00:08:51) It's like, oh yeah, we can implement that. (00:08:53) But when it's a really good idea by my own account, it does seem to reward it a bit more. (00:08:58) And so I kind of feel like I'm trying to like earn its praise, which is really weird. (00:09:02) And so I do think the personality matters A lot. (00:09:04) And I think a lot of the other tools maybe don't appreciate it as much. (00:09:07) And I think in this aspect also, Peter really cares about this. (00:09:09) And so that was correct. (00:09:10) And then the memory system and then (00:09:12) just, he's just having fun with this. (00:09:15) And then the single WhatsApp portal to all of the automation. (00:09:18) Yeah. (00:09:18) Is there something that you have done personally with your claws beyond software engineering that you think is fun or interesting? (00:09:26) Yeah, so in January, I had a claw. (00:09:28) I went through a period of claw psychosis. (00:09:29) So I built, I have a claw basically that takes care of my home. (00:09:33) And I call him Dobby the Elf Claw. (00:09:37) And basically I used the agents to find all of the smart home subsystems of my home on the local area network, which I was kind of surprised that worked out-of-the-box. (00:09:46) Like I just told it that I think I have Sonos at home. (00:09:48) Like can you try to find it? (00:09:49) And it goes and like IP scan of all the basically computers on the local area network. (00:09:55) And it found the Sonos thing, the Sonos system. (00:09:58) And it turned out that there's no password protection or anything like that. (00:10:01) It just logged in and it's like, oh yeah, you have these Sonos systems installed. (00:10:04) I let me try to reverse engineer how it's working. (00:10:06) It does some web searches and it finds like, okay, these are the API endpoints. (00:10:10) And then it's like, do you want to try it? (00:10:11) And I'm like, whoa, like you just did that. (00:10:12) And I'm like, yeah, can you try to play something in the study? (00:10:15) And it does. (00:10:16) And music comes out. (00:10:17) And I'm like, I can't believe I just. (00:10:18) That's crazy. (00:10:19) That's like 3 prompts. (00:10:20) I can't believe I just typed in like, can you find my Sonos? (00:10:22) And that suddenly it's playing music. (00:10:23) And it did the same for lights. (00:10:25) And so basically like it kind of hacked in, figured out the whole thing, created APIs, created dashboard so I could see the command kind of center of like all of my lights in the home. (00:10:33) And then it was like. (00:10:34) like switching lights on and off. (00:10:35) And so I can ask it like, don't be at sleepy time. (00:10:38) And when it's sleepy time, that just means all the lights go off, et cetera, and so on. (00:10:42) So it controls all of my lights, my HVAC, my shades, the pool, and the spa, and also my security system. (00:10:49) So I have a camera pointed outside of the house. (00:10:51) And anytime someone rolls in, I have a Quinn model that looks at the videos. (00:10:56) So first of all, there's change detection. (00:10:58) And then based on change detection, it goes to Quinn. (00:11:00) And then it actually tells me, it sends me a text to my WhatsApp. (00:11:04) It shows an image from the outside, and it says, Hey, FedEx truck just pulled up, FedEx truck just pulled up, and you might want to check it, and you got an e-mail or something like that, and Dobby just texts me this is really incredible. (00:11:17) So Dobby is in charge of the house. (00:11:19) I text with it through WhatsApp. (00:11:22) And it's been really fun to have these macro actions that maintain my house. (00:11:25) I haven't really pushed it way more beyond that. (00:11:28) And I think people are doing a lot more crazy things with it. (00:11:30) But for me, even just the home automation setup, I used to use like 6 apps, completely different apps. (00:11:35) And I don't have to use these apps anymore. (00:11:36) Like Dobby controls everything in natural language. (00:11:38) It's amazing. (00:11:39) And so I think I haven't even pushed a paradigm fully, but already that is so helpful and so inspiring, I would say. (00:11:45) Do you think that's indicative of like what people want from a user experience perspective with software, right? (00:11:50) Because I don't think, you know, it's pretty ignored that it takes humans effort to like learn new software, like new UI. (00:11:57) Yeah, I think to some extent, that's right. (00:12:00) is like working backwards from how people think an AI should be. (00:12:03) Because what people have in their mind of like what an AI is, not actually what an LLM is by like in a raw sense. (00:12:09) Like LLM is a token generator, you know, like more tokens come out. (00:12:12) But what they think of is like this persona identity that they can tell stuff and it remembers it. (00:12:19) And it's just kind of an entity behind the WhatsApp. (00:12:21) It's like a lot more understandable. (00:12:23) So I think to some extent, it's like matching the expectations that humans already have for what an AI should behave. (00:12:28) But under the hood, there's like a lot of technical details go into that. (00:12:30) And LLMs are too raw of a primitive to actually type check as AI, I think, for most people, if that makes sense. (00:12:37) Yeah, I think that's like how we understand what the AI is and like the. (00:12:43) description of it as Dobby or some personality obviously resonates with people. (00:12:48) I also think that the unification that you did across your six different software systems for your home automation speaks to a different question of like, (00:12:57) Do people really want all the software that we have today? (00:13:00) Right? (00:13:00) Because I would argue like, well, you have the hardware, but you've now thrown away the software or the UX layer of it. (00:13:08) Do you think that's what people want? (00:13:10) Yeah, I think there's this like, there's this sense that these apps that are in the app store for using these smart home devices, et cetera, these shouldn't even exist kind of in a certain sense. (00:13:19) Like shouldn't it just be APIs and shouldn't agents be just using it directly? (00:13:23) And wouldn't it like, I can do all kinds of home automation stuff that any individual app will not be able to do, right? (00:13:30) And then LLM can actually drive the tools and call all the right tools and do pretty complicated things. (00:13:35) And so in a certain sense, it does point to this (00:13:39) maybe there's like an overproduction of lots of custom bespoke apps that shouldn't exist because agents kind of like crumble them up. (00:13:45) And everything should be a lot more just like exposed API endpoints. (00:13:49) And agents are the glue of the intelligence that actually like tool calls all the parts. (00:13:55) Another example is like my treadmill. (00:13:56) There's an app for my treadmill and I wanted to like keep track of how often I do my cardio. (00:14:01) But I don't want to log into a web UI and go through a flow and et cetera. (00:14:05) All this should just be like make APIs available. (00:14:07) And this is kind of going towards the agentic sort of web or like agent-first tools and all this kind of stuff. (00:14:14) So I think the industry just has to reconfigure in so many ways that it's like the customer is not the human anymore. (00:14:18) It's like agents who are acting on behalf of humans. (00:14:21) And this refactoring will be (00:14:22) will probably be substantial in a certain sense. (00:14:24) One way that people sometimes push back on this is like, do we expect people to vibe code some of these tools? (00:14:30) Do we expect normal people to do this kind of stuff that I described? (00:14:34) But I think to some extent, this is just technology as it exists today. (00:14:37) And right now, there is some vibe coding, and I'm actually watching it, and I'm working with the system. (00:14:42) But I kind of feel like this kind of stuff that I just talked about, this should be free. (00:14:46) in a year or two or three. (00:14:48) There's no web coding involved. (00:14:49) This is trivial. (00:14:49) This is table stakes. (00:14:50) This is like any AI, even the open source models, et cetera, can like do this. (00:14:54) You should be able to translate from a less technical human's intent very easily to this output. (00:14:59) Yeah, extremely easily. (00:15:00) Yeah. (00:15:01) Today it's web coding, it's involved and not many people are going to do it. (00:15:03) And you still have to make some design decisions, right? (00:15:05) We were talking about like, well, you take frames, for example. (00:15:08) Yeah. (00:15:08) But I kind of feel like this will just start to, the barrier will just come down and it's just ephemeral software on your behalf. (00:15:16) and some kind of like claw is handling all the details for you, but you're not involved. (00:15:20) Claw has a, claw has a machine and it will figure it out and it's just presenting your UIs and you're like saying stuff, you know. (00:15:27) Why haven't you, I guess, like pushed the boundaries of what you can do personally with claws? (00:15:32) Like is it, you know, you're focusing on more important projects, auto research, et cetera, or (00:15:38) you're climbing the hill to mastery or something else, right? (00:15:42) Yeah, I just feel like I'm so distracted by everything. (00:15:44) So I spend like a week on the class stuff and I have more to do almost. (00:15:49) But I will say that... (00:15:50) It's like Jensen tools, we're all just busier, unfortunately. (00:15:53) Yeah. (00:15:54) I didn't really take advantage of a lot of like e-mail and calendar and all this other stuff. (00:15:58) And I didn't leave it access because I'm still a little bit like suspicious and still very new and rough around the edges. (00:16:02) So I didn't want to give it like full access to my digital life yet. (00:16:05) And part of it is just less security, privacy and (00:16:08) just being very cautious in that realm. (00:16:11) And so some of it is like held back by that, I would say. (00:16:14) maybe that's like the dominant feature, but some of it is also just, I feel so distracted because I feel like I had a week of claw and then other stuff is happening. (00:16:22) What was the, I mean, you've talked about like being able to train or at least optimize a model as a task you want to see agents do for a long time. (00:16:32) Like what was the motivation behind auto research? (00:16:34) Auto research, yeah. (00:16:35) So I think like, (00:16:36) I had a tweet earlier where I kind of like said something along the lines of, to get the most out of the tools that have become available now, you have to remove yourself as the bottleneck. (00:16:45) You can't be there to prompt the next thing. (00:16:47) You need to take yourself outside. (00:16:50) You have to arrange things such that they're completely autonomous. (00:16:53) And the more, you know, how can you maximize your token throughput and not be in the loop? (00:16:56) This is the goal. (00:16:58) And so (00:16:59) I kind of mentioned that the name of the game now is to increase your leverage. (00:17:02) I put in just very few tokens just once in a while, and a huge amount of stuff happens on my behalf. (00:17:07) And so auto research, I tweeted that, and I think people liked it and whatnot, but they haven't maybe worked through the implications of that. (00:17:13) And for me, auto research is an example of an implication of that. (00:17:16) Where it's like, I don't want to be the researcher in the loop, looking at results, et cetera. (00:17:21) I'm holding the system back. (00:17:23) So the question is, how do I refactor all the abstractions so that I'm not, I have to arrange it once and hit go? (00:17:28) The name of the game is, how can you get more agents running for longer periods of time without your involvement doing stuff on your behalf? (00:17:34) And auto research is just, yeah, here's an objective, here's a metric, here's your boundaries of what you can and cannot do, and go. (00:17:41) And yeah. (00:17:42) You were surprised at its effectiveness. (00:17:45) Yeah, I didn't expect it to work because, so I have the Project Nana Chat. (00:17:50) And fundamentally, I think a lot of people are very confused with my obsession for training GPT-2 models and so on. (00:17:55) But for me, training GPT models and so on is just a little harness, a little playground for training LLMs. (00:17:59) And fundamentally, what I'm more interested in is like this idea of recursive self-improvement and to what extent you can actually have LLMs improving LLMs. (00:18:06) Because I think all the Frontier Labs, this is like the thing for obvious reasons. (00:18:10) And they're all trying to recursively self-improve, roughly speaking. (00:18:13) And so for me, this is kind of like a little playpen of that. (00:18:17) And I guess I'd tuned NamChat already quite a bit by hand in a good old-fashioned way that I'm used to. (00:18:21) Like I'm a researcher, I've done this for like 2 decades. (00:18:24) I have some amount of like, what is the opposite of humans? (00:18:28) Earned confidence. (00:18:30) Okay. (00:18:30) I have like 2 decades of like, oh, I've trained this model like thousands of times of like, so I've done a bunch of experiments, I've done hyper-parameter tuning, I've done all the things I'm very used to and I've done for two decades. (00:18:40) And I've gotten to a certain point and I thought it was like fairly well tuned. (00:18:44) And then I let our research go for overnight and it came back with tunings that I didn't see. (00:18:49) And I did forget the weight decay on the value embeddings and my atom betas were not sufficiently tuned. (00:18:54) And these things jointly interact. (00:18:56) So once you tune one thing, the other things have to potentially change too. (00:18:59) I shouldn't be a bottleneck. (00:19:00) I shouldn't be running these hyper-parameters such optimizations. (00:19:02) I shouldn't be looking at the results. (00:19:04) There's objective criteria in this case. (00:19:06) So you just have to arrange it so that it can just go forever. (00:19:09) So that's a single sort of version of auto research or like a single loop trying to improve. (00:19:13) And I was surprised that it found these things that I, you know, the repo was already fairly well tuned and still found something. (00:19:19) And that's just a single, it's a single loop. (00:19:21) Like these Frontier Labs, they have GPU clusters of 10s of thousands of them. (00:19:25) And so it's very easy to imagine how you would basically get a lot of this automation on smaller models. (00:19:32) And fundamentally everything around like frontier level intelligence is about extrapolation and scaling loss. (00:19:37) And so you basically do a ton of the exploration on the smaller models, and then you try to (00:19:42) extrapolate that. (00:19:43) So you're saying our research efforts are going to get more efficient, like we're going to have better direction for when we scale as well, if we can do this experimentation better. (00:19:50) Yeah, I would say that like the most interesting project and probably what the Frontier Labs are working on is, you know, you experiment on the smaller models, you try to make it as autonomous as possible, remove researchers from the loop. (00:20:01) They have way too much. (00:20:02) What is the opposite? (00:20:05) Yeah, they don't know. (00:20:06) They shouldn't be touching any of this, really. (00:20:08) And so you have to rewrite the whole thing, because right now, I mean, certainly they can contribute ideas. (00:20:13) But okay, they shouldn't actually be enacting those ideas. (00:20:16) There's a queue of ideas. (00:20:17) And there's maybe an automated scientist that comes up with ideas based on all the archive papers and GitHub repos, and it funnels ideas in. (00:20:24) Or researchers can contribute ideas, but it's a single queue, and there's workers that pull items, and they try them out. (00:20:30) And whatever works just gets sort of put on the feature branch, and maybe some people (00:20:35) monitor the feature branch and merge to the main branch sometimes. (00:20:39) So yeah, just removing humans from all the processes and automating as much as possible and getting high tokens per second throughputs. (00:20:46) And it does require rethinking of all the abstractions and everything has to be reshuffled. (00:20:52) So yeah, I think it's very exciting. (00:20:54) If we take one more recursive step here, when is the model going to write a better program MD than you? (00:21:01) Yeah. (00:21:03) We're not in the loop. (00:21:04) Yeah, exactly. (00:21:06) So program MD is my crappy attempt at describing like how the auto researcher should work. (00:21:11) Like, oh, do this, then do that, and that, and then try these kinds of ideas. (00:21:14) And then here's maybe some ideas like look at architecture, look at optimizer, et cetera. (00:21:18) But I just came up with this in markdown, right? (00:21:21) And so, yeah, exactly. (00:21:25) You want some kind of an auto research loop maybe that looks for, (00:21:29) You can imagine that different program.nds would give you different progress. (00:21:35) So basically every research organization is described by program MD. (00:21:38) A research organization is a set of markdown files that describe all the roles and how the whole thing connects. (00:21:44) And you can imagine having a better research organization. (00:21:46) So maybe they do fewer standups in the morning because they're useless. (00:21:50) And this is all just code, right? (00:21:52) And so you can, so one organization can have fewer standups, one organization can have more. (00:21:56) One organization can be very risk-taking. (00:21:58) One organization can be less. (00:22:00) As you can definitely imagine that you have multiple research orgs, and then they all have code. (00:22:04) And once you have code, then you can imagine tuning the code. (00:22:07) So 100%, there's like the meta layer of it. (00:22:09) Did you see my text about my contest idea? (00:22:12) My contest idea was like let people write different program MDs, right? (00:22:19) And so for same hardware, where do you get most improvement? (00:22:22) Oh, I see. (00:22:23) And then you can take all that data (00:22:25) and then give it to the model and say, write a better program MD. (00:22:27) Yes, yeah, exactly. (00:22:29) We're gonna get something better. (00:22:30) Like, there's no way we don't. (00:22:31) You could 100% look at where the improvements came from, and like, can I change the program MD such that more of these kinds of things would be done, or like things that didn't work? (00:22:40) It's meta optimization. (00:22:41) Yeah, you can 100% imagine doing that. (00:22:43) So I think this is a great idea. (00:22:45) It's like, I think you sort of go one step at a time, where you sort of have one process and then second process and then the next process, and these are all layers of an onion. (00:22:53) Like the LLM sort of part is now taken for granted, the agent part is now taken for granted, now the claw-like entities are taken for granted, and now you can have multiple of them, and now you can have instructions to them, and now you can have optimization over the instructions. (00:23:05) And it's just like it's a little too much, but I mean, this is why it gets to the psychosis is that this is like infinite and everything is skill issue. (00:23:11) And that's why I feel like, yeah, that's just coming back to, this is why it's so insane. (00:23:16) Okay, well, if we're just trying to like diagnose the current moment and what is a relevant skill right now, what do you like, what do you think is the implication that this is the loop we should be trying to achieve in different areas? (00:23:29) And then it works, right? (00:23:30) Like, you know, remove (00:23:33) create the metric or create the ability for agents to continue working on it without you. (00:23:38) Do we still have performance engineering? (00:23:41) Yeah, I mean, so there's a few caveats that I would put on top of the LM psycho system. (00:23:44) Number one, this is extremely well suited to anything that has objective metrics that are easy to evaluate. (00:23:49) So for example, like writing kernels for more efficient CUDA code for various parts of a model, etc. (00:23:55) are the perfect fit because you have inefficient code and

Around this claim