ATRIUMsearch → argument graph
MechanismAudio · 14:08 — 15:05

AI breakthroughs allow robots to correct their own mechanical inaccuracies, which lets us use cheap, compliant, imprecise actuators that are inherently safe, unlike traditional industrial robots.

Cheng Chi explains that industrial robots were designed to be fast, stiff, and precise because they were blind — they blindly followed programmed trajectories. But AI breakthroughs give robots eyes, so they can correct their own hardware inaccuracies, enabling cheap, safe, compliant actuators. ✦ AI generated

Cheng Chi · No Priors · 2025-11-19 · original ↗

plays this moment only · 14:08 — 15:05

Elicited by

Why isn't the obvious answer, it should just be like a full human eye?

Traditionally, most robots are designed for industrial use cases. And the robots are very fast, they are very stiff, and they're very precise. The reason is because all the industrial robots are blind. So they're blindly following a trajectory that's programmed by someone. It's not reaction to perception. Correct. But because of the breakthrough we had in AI, like now robot have eyes, so it can actually correct its own mechanical and hardware inaccuracies. So that kind of opened up a new different space of design. And intuitively, it should be like, I can't tell you exactly what the distance is here on a millimeter scale, but I'm going to get to the cup because I could stop. So that allows us to use these low cost actuators that's cheap, that's compliant, but they're imprecise. But because of the AI's algorithms and systems we build, it allows us to build robots that's mechanically inherently safe and compliant, while simultaneously be able to achieve the sufficient accuracy we need for the home tasks.

verbatim transcript · starts at 14:08

Transcript · around this moment

(00:00:06) Nobody wants to do their dishes. (00:00:07) Nobody wants to do their laundry. (00:00:09) People will love to spend more time with their family, with their loved ones. (00:00:14) So what we believe in is that if the robot is cheap, safe, and capable, everyone will want our robots. (00:00:20) And we see a future where we have more than 1 billion of these robots in people's homes within a decade. (00:00:30) Thanks, Memo. (00:00:36) Hi, listeners. (00:00:36) Welcome back to No Priors. (00:00:38) Today we're here with Tony Zhao and Cheng Chi, co-founders of Sundae, makers of Memo, the first general home robot. (00:00:45) We'll talk about AI and robotics, data collection, building a full-stack robotics company, and a world beyond toil. (00:00:51) Welcome. (00:00:52) Cheng, Tony, thanks for being here. (00:00:54) Thanks for having us. (00:00:55) Yeah. (00:00:56) Okay, first I want to ask, like, where are we here? (00:00:59) Because classical robotics has not been an area of great optimism over time or like massive velocity of work. (00:01:06) And now people are talking about a foundation model for robotics or a ChatGPT moment. (00:01:11) Can you just contextualize like the state of AI robotics and why we should be excited? (00:01:16) I would say, I think we're kind of in between the GPT moment and the ChatGPT moment, like in the context of LLMs. (00:01:23) What it means is that (00:01:25) It seems like we have a recipe that can be scaled, but we haven't scaled up yet. (00:01:31) And we haven't scaled up so much so that we can have a great consumer product out of it. (00:01:36) So this is what I mean like GPT, which is like a technology, and ChatGPT, which is a product. (00:01:42) Yeah, and so we're seeing across academia, there's consensus around what's the method for manipulation, but everybody's talking about scaling up. (00:01:51) It's like, we know there's sign of life for the algorithms. (00:01:54) people are picking, but people don't know if we have more data, like what happened to GPT-2, GP-3, what will happen. (00:02:02) But we see a clear trend that there's no reason to believe that robotic doesn't follow the trajectory of other AI fields, that scaling up is going to improve performance. (00:02:12) Maybe even if you took a step back, like what was the process for deploying a robot into the world like 10 years ago, like pre-set of generalizable AI algorithms? (00:02:22) Like why was it so slow as a field? (00:02:24) Yeah, so previously, classical robotics have this sense, plan, act modular approach where there's a human designing interface between each of the modules. (00:02:35) And those are need to be designed for each specific task and each specific environment. (00:02:39) In academia, that means for every task, (00:02:42) That means a paper. (00:02:43) So a paper is you design a task, design an environment, and you design interfaces, and then you produce engineering work for that specific task. (00:02:50) But once you move on to the next task, you throw away all your code, all your work, and you start over again. (00:02:55) And that's also kind of what happened to industry. (00:02:58) So for each application, people build a very specific software and hardware system around it, but it's not really generalizable. (00:03:04) And therefore, it's still like we're just running in loops. (00:03:07) We build A1 system and then we build the next one, but there's like no (00:03:10) synergy between them. (00:03:12) And as a result, the progress has been somewhat slow. (00:03:14) I feel like that's a good segue into some of the amazing research work that you guys have contributed over the last five years to the field. (00:03:21) Should we start with diffusion policy? (00:03:22) What was the impact of that? (00:03:24) Yeah, so a diffusion policy is like specific algorithm for a paradigm called imitation learning. (00:03:29) That's really like the most intuitive way of, you know, how to use machine learning for robotics. (00:03:34) So you collect paired data, action, observation, (00:03:37) of what the robot should do. (00:03:38) You use that to train a model with supervised learning, and then the robot do the same thing. (00:03:43) The problem is that in the field, it's known to be very finicky. (00:03:47) So when I talk to researchers, when I start into the field, people are like, the researcher themselves, the specific researcher, need to collect the data so that there's exactly one way to do everything. (00:03:58) Otherwise, the robot either, like, use the model training will diverge or the robot will behave some weird way. (00:04:03) And diffusion model really allows us to capture multiple modes of behavior for the same observation in a way that's still preserved training stability. (00:04:13) And that really kind of unlocked (00:04:15) more scalable training and more scalable data collection. (00:04:18) So it doesn't have to be you personally wearing a tele-ops set in order to make a robot learn. (00:04:23) Yep, So like we can have multiple people, sometimes even untrained people collecting data, and the result will still be great. (00:04:29) Where do Aloha and ACT play into this? (00:04:33) so these two papers are actually super close to each other. (00:04:36) They're like one month or two months away. (00:04:38) That's actually how me and Chen know each other. (00:04:40) It was by looking at each other's paper, like, and we met on Twitter, I think, when Chen is back in Colombia. (00:04:46) Before, Aloha, I think the typical way people collect data is with a like polyoperation setup with we are headset. (00:04:55) And it turns out to be very unintuitive to do. (00:04:58) and it's hard to collect data that is actually dexterous. (00:05:01) What Aloha brings is a very simple and reproducible setup. (00:05:04) So it's very intuitive. (00:05:05) Sorry, in terms of just for most people who haven't worn a tele-ops setup, is it the lag? (00:05:10) Is it like just, you know, how should I compare it to like playing a video game or something? (00:05:14) Yeah, I think Aloha make it feel more like playing a video game. (00:05:17) normally feels kind of disconnected, that you're just like moving in the free air and the robot is moving with some delays, but Aloha reduces that delay by a lot. (00:05:27) And that contribute to the kind of smoothness and how fast human can react. (00:05:32) Like once we get those really dexterous data, what it allows us to do is to investigate on algorithms that are actually solving things that are difficult. (00:05:42) In this case is sort of the introducing of using transformers. (00:05:47) in the case of robotics. (00:05:49) And there was a long period of time that I think robotics was stuck with failure MLPs and conv nets. (00:05:56) And as you make it deeper, it works worse. (00:05:58) But it turns out that once you have very strong and dexterous data sets, like just throw a transformer at it and it works (00:06:07) Actually, just in terms of progress of the industry over time, transformers didn't make sense without certain level of data collection capability. (00:06:15) Okay. (00:06:15) And also auto system around it, for example, action chunking, which is to predict a trajectory as opposed to predicting single samples of actions. (00:06:22) All these things kind of combined to make dexterous task by manual tasks more scalable. (00:06:28) Why is chunking important here if I think about like just the analogy to LLMs and like text sequence prediction? (00:06:36) I think it just kind of throws the model off if you're trying to force it to react every millisecond. (00:06:42) That's not how human act. (00:06:45) We perceive and we can actually move quite a bit without looking at things again. (00:06:51) And that turns out to make things, the motion a lot more consistent and oral performance to be a lot better. (00:06:57) And you discovered that actually transformers architecturally did apply to robotics. (00:07:02) Cheng, you felt then that data collection was still a problem, so enter Umi. (00:07:07) Yeah, so after Aloha and diffusion policy, I was super excited about imitation learning. (00:07:13) But at the time, both of us started still doing teleoperation, and that just feels super limiting. (00:07:18) I think the problem is that, you know, in your setup at a time, like a teleop setup, (00:07:24) It takes a PhD student a couple of hours to set up in a lab. (00:07:27) It pretty much restrict data collection to a lab. (00:07:30) But in order for the robot to actually work as a product, it needs to be work in the wild, in unseen environments that requires data to be also be collected in the wild. (00:07:39) And at the time I was thinking, okay, is there a way we can collect robotic data without actually using a robot? (00:07:45) That forced me to think, okay, what's the actual most essential part of the robotics data? (00:07:50) And after diffusion policy and act, actually, the paradigm is kind of simple. (00:07:55) You just need paired observation and action data. (00:07:57) In our case, observation is the video clip. (00:07:59) The action is the movement of your hand, plus, you know, how the finger moves. (00:08:05) I realized that all this information you can get from a GoPro. (00:08:08) You can track the movement of GoPro in space, and you can, you know, track the motion of the (00:08:14) of the gripper and also finger through images as well. (00:08:17) And that's why I built this Umi gripper. (00:08:20) It's 3D printed. (00:08:21) At the time, like the project had three PhD students. (00:08:24) Like we just took the grippers everywhere. (00:08:26) Like, you know, I think it was two weeks before the paper deadline. (00:08:29) Like every time it goes to a restaurant, before the wearer come in, we just collect some data. (00:08:33) And very quickly we got, you know, I think 1,500 video clips of this like espresso cup serving task. (00:08:41) And that turns out to be (00:08:43) one of the biggest data sets in robotics and just simply by three people. (00:08:47) And that's like how, that's where the power kind of shines. (00:08:50) And then with that amount of data, allows to train the first end-to-end model that can actually generalize to unseen environments. (00:08:57) So we can push the robot around in Stanford. (00:08:59) Actually, Tony was there as well. (00:09:01) You know, push the robot arm around the Stanford campus and then anywhere, you know, the robot can serve you a drink. (00:09:07) I think that is the moment I was like, hey, maybe we should start a company. (00:09:12) This is actually working so well. (00:09:14) I remember like just following Chung and... (00:09:16) It doesn't work well. (00:09:18) Yes. (00:09:18) I think the only exception I saw was when it's under direct sunlight. (00:09:22) Yeah. (00:09:23) And I think the reasoning was like over that whole like 2, three weeks of data collection. (00:09:27) That 2 weeks is all raining. (00:09:28) So there's no sunlight data. (00:09:30) So like it fails. (00:09:30) That also demonstrates the importance of distribution matching. (00:09:34) So in order for robot to work, (00:09:36) sunny environment, it must have seen sunny environments during the training data. (00:09:40) Yeah, it's really interesting because I remember when I first met you guys, it was like you spent like, I don't know, $200,000 across all of your academic research, and yet the scale of data collection as translated to model capability is leading, right? (00:09:55) So it's very interesting that, you know, we look at where we are, maybe going back to Tony's point of scaling at massive capital deployment, (00:10:04) But that entire paradigm actually wasn't relevant before people realized like you should train on all of the internet data. (00:10:11) And we just don't have that in robotics. (00:10:13) So the entire field is just blocked on having any scale of data that's relevant. (00:10:17) I think these days there are still like so many debates about like what is even the right way to scale. (00:10:22) There are like role models, there are simulations, there is teleoperation, there are like all these new ideas. (00:10:28) And I think this is the (00:10:30) the sort of area that we really want to innovate, that we want to differentiate. (00:10:34) We want to find out something that is both high quality and scalable. (00:10:38) And then you guys, you decide to start a company, pushing this cart around Stanford. (00:10:43) Tell me about that decision and congratulations on the launch and sort of the direction and team you've built. (00:10:49) Yeah, it's a very interesting journey. (00:10:52) I remember in the beginning, it's actually two of us in Chung's apartment. (00:10:56) On his desk, we were like climbed a robot there and tried to do some tasks. (00:11:00) And it soon becomes like, I think, an eight-person team towards the end of 2024. (00:11:06) And now we're at around like 30 to 40 people. (00:11:08) We're not the best at everything, right? (00:11:10) But starting a company allows us to find people who we really love working with and then bring all the expertise together from mechanical engineering, supply chain, like software engineering, like controls, and to build a system together that is not like a demo. (00:11:26) but a real product. (00:11:27) You've built this amazing team. (00:11:28) What are people actually signing up for? (00:11:30) What's the mission of Sundae? (00:11:31) Yes, it is to put a home robot in everyone's home. (00:11:34) I think there are a lot of AIs trying to make you more efficient during the work, but there is not enough AI that actually helps you with all these mundane things that are not creative, that really has nothing to do with what's making us more intrinsically human. (00:11:49) What's ideal for people to spend more time on is actually with their hobbies, with their passions, as opposed to spending more (00:11:56) I'm doing chores. (00:11:57) So if you guys are going from, (00:12:00) these amazing research breakthroughs to, we're actually going to ship a home robot, that's a product you have to talk about, cost and capability and robustness, like what's the design philosophy? (00:12:12) As these AI models becomes more capable and as hardware cost continues to go down, the home robots, all kinds of robots will be everywhere. (00:12:22) So if you start from the most surface level, which is the design of the robots, when we design it, we think about (00:12:30) what should the robot look like if it is ubiquitous? (00:12:33) You need to see it every single day. (00:12:35) What should it look like? (00:12:37) And what we end up with is that we really think the robot should have a face. (00:12:42) It should have a cute face. (00:12:43) And it should be very friendly. (00:12:45) So instead of like a Terminator doing your dishes, we want the robot to feel like it's out of a cartoon movie. (00:12:52) And then the huge decision is like, how many arms should the robot have? (00:12:57) Should it have like 4 arms? (00:12:57) Should it have like 1 arms? (00:12:59) Should it have legs? (00:13:00) Should it have like 5 fingers, 2 fingers, 3 fingers? (00:13:02) It's a huge space. (00:13:04) Why isn't the obvious answer, it should just be like a full human eye? (00:13:07) I think the core motivation for us is how can we build a useful robot as soon as possible. (00:13:14) So whenever we see something that we can accelerate it with simplification, we'll go simplify that. (00:13:21) So one example of that is the hand that we designed, which has three fingers. (00:13:26) We kind of combine the three of the fingers that we have together. (00:13:29) And the reasoning there is just that most of the time when we use those fingers, we use it together. (00:13:34) Let it be like rasping a handle, let it be opening the dishwasher. (00:13:38) So it really doesn't make sense to add the cost, like multiply by 3X to have separate it into three, when we can do one with most of the benefits. (00:13:48) So this is how we think about the whole robot as well. (00:13:51) It's kind of with the constraint that we're building a general purpose robot that can eventually do all your chores and will simplify everything we possibly can so that the robot can be as low cost and as easy to repair as possible. (00:14:04) Yeah, I just want to add a little bit more to the actuator and mechanical design. (00:14:08) Traditionally, most robots are designed for industrial use cases. (00:14:11) And the robots are very fast, they are very stiff, and they're very precise. (00:14:16) The reason is because all the industrial robots are blind. (00:14:19) So they're blindly following a trajectory that's programmed by someone. (00:14:22) It's not reaction to perception. (00:14:23) Correct. (00:14:24) But because of the breakthrough we had in AI, like now robot have eyes, so it can actually correct its own mechanical and hardware inaccuracies. (00:14:33) So that kind of opened up a new different space of design. (00:14:37) And intuitively, it should be like, I can't tell you exactly what the distance is here on a millimeter scale, but I'm going to get to the cup because I could stop. (00:14:45) Yeah, exactly. (00:14:46) So that allows us to use these low cost actuators that's cheap, that's compliant, but they're imprecise. (00:14:54) But because of the AI's algorithms and systems we build, it allows us to build robots that's mechanically inherently safe and compliant, while simultaneously be able to achieve the sufficient accuracy we need for the home tasks. (00:15:05) Where are we in that timeline? (00:15:07) You said we're between GPT and ChatGPT. (00:15:11) And so like when do consumers get ChatGPT and when will you guys ship something? (00:15:15) Yeah, it's actually a really exciting time because like we have so many prototypes internally. (00:15:19) What we will do next year, 2026, is actually start doing beta programs. (00:15:24) We'll have these robots, all kinds of different ones, into people's home and see how they react to it. (00:15:30) That will be when we learn the most about how people, like people want to talk to their robots. (00:15:36) Do people want to have their robots maybe teach their kids some new knowledge about the world? (00:15:43) And this will inform us what the eventual product should look like. (00:15:47) Internally, we just have an extremely high standard of what is the minimal consumer product we want to ship. (00:15:55) It needs to be extremely safe, it needs to be extremely capable, and low cost. (00:15:59) Do you feel like you know something now that you didn't when you started the company? (00:16:02) Absolutely. (00:16:04) I think at the beginning, I would describe it as like we see light at the end of the tunnel. (00:16:09) of there are two axes, there's dexterity, there's generalization. (00:16:12) When we add more data, things works better. (00:16:15) And what this company about is the cross product of these two. (00:16:18) How can we scale and have both dexterity and generalization? (00:16:23) And this is something we're able to show in our generalization demo, which is like we can pick up these like very precise like forks, like actual metallic forks, only ceramic plates with very high success rates. (00:16:35) And honestly, this is not something that (00:16:39) we thought that would work so easily just by having so much more data. (00:16:43) So actually, just want to expand a little bit, it's actually the process was long and painful. (00:16:50) So there's so many issues, like just scaling up a system, a robotic system is very, very hard. (00:16:56) There are mechanical issues, like reliability issues. (00:17:00) There's like data, you know, quality issues that come out of it. (00:17:03) In the beginning, I actually thought it's going to be much easier than this. (00:17:07) But really it just, it takes time and effort to grind out all the little details for this to work. (00:17:13) I also actually, compared to teleop, it's much harder to get the system scaled up. (00:17:16) But once it's scaled up, it's very powerful, very repeatable. (00:17:19) So it is both harder than you thought it would be to get to here, and you are further than you thought you would be. (00:17:24) Yes. (00:17:25) And I remember in the beginning, we were having this like funny conversation of, we're like, if we build this, someone can just like take our glove and they'll build the same thing. (00:17:35) Like, what more do we have? (00:17:36) Are we worried about that? (00:17:38) And in the beginning, actually, we were a little bit worried, because we thought, they can probably just replicate it. (00:17:44) But as we go along the path, it turns out things are so much harder than we thought it would. (00:17:48) There's so many, the small demands. (00:17:50) Yeah, yes. (00:17:51) And when you say scaling up the robotic system, you mean the data collection to training pipeline and the hardware itself. (00:17:57) Yeah, so actually, like, for this to work at all, we need the data collection system, we need the robotic and control system to be able to deliver the hand to where we want to go. (00:18:06) And you also need the data filtering pipeline and data cleaning pipeline and the training pipeline. (00:18:12) And all these things need to be iterated together. (00:18:14) So actually we've gone through several loop of these. (00:18:16) It's kind of hard to imagine without having a full stack team in-house how this can even be done. (00:18:22) Yeah. (00:18:22) The glove we're using right now is we call it like V5. (00:18:26) And for V0 to V5, each version has like around 20 iterations. (00:18:31) Okay. (00:18:32) And so 100. (00:18:33) Yes. (00:18:34) And also like just when you make these at scale, right now we have more than 500 people using these gloves in the wild. (00:18:41) Like all the things that could go wrong will go wrong. (00:18:45) For example, they did. (00:18:46) They did. (00:18:47) Yes. (00:18:48) For example, like how things are assembled. (00:18:51) If you don't (00:18:52) specify exactly how it should be done, people will assemble it in creative ways. (00:18:57) And the creativity doesn't help us here because we really want the data collection device to be extremely precise. (00:19:03) So you guys can't obviously know everything that's happening in every company in academia and industry, but from what you know, how would you compare the scale of training data that you have today relative to the industry? (00:19:16) At this point, we are almost 10 million. (00:19:21) trajectories being collected in the wild. (00:19:24) And those trajectories are not just like, oh, pick up a cup. (00:19:26) It's these long trajectories with like walking, with navigation, and then like doing these long horizon tasks. (00:19:34) Tony, as you mentioned, like, it's an open question, actually, what the right way to scale data up is. (00:19:40) And so there are strong theories around teleop, around like pure RL, around video and world models. (00:19:47) Like, how did you think about all of these? (00:19:51) so from our perspective, actually, it's kind of somewhat surprising. (00:19:54) So, in the beginning, we worried that, the data from Glob or Umi-like data has higher quantity but lower quality compared to TallyOp, because of TallyOp, you're using exactly the same hardware and software stack between training and testing is perfectly distribution matched. (00:20:07) But what we realized is actually, you know, this Glob form factor encouraged people to do more dexteration, more natural movement, and those actually result in a like a more intelligent behavior on the modeling side. (00:20:20) And in terms of (00:20:21) data quality, we don't really see a difference in terms of, how much, like, there's a gap between tidy up and glove data. (00:20:28) After we did the 20 engineering, like, serious message, yeah, like, 'cause, like, apparently there is a mismatch, right? (00:20:37) That's in the camera frame, there's a human inside of the robot, and there are just a lot of things that we need to do to kind of... (00:20:45) convert a human data one-to-one to, as if it is robot data, and have the model not being able to tell the difference. (00:20:53) Yeah, and that kind of relies on, again, the whole, the full cycle iteration between hardware and software. (00:20:58) What about RL? (00:21:00) We see a lot of great promise for RL in locomotion, and we think that will continue to be true for locomotion. (00:21:06) So what we see, really, like RL is a method, it's very powerful, but it is (00:21:12) much less sample efficient compared to imitation learning. (00:21:15) And we see that to work great in environments where it's easy to simulate. (00:21:18) For the case of locomotion, you don't need to worry about rigid body dynamics and rigid body contact between the robot and the ground. (00:21:26) And, you know, because you engineer a robot, you know everything. (00:21:30) But for manipulation, it's kind of hard for us to imagine, like, have this actually the same amount of diversity and the distribution of real object. (00:21:40) in terms of matching both appearance and physical properties. (00:21:45) And we think that it's going to be challenging compared to globe data collection and tidy up. (00:21:50) Yeah, I think it's really about which method can get us there faster. (00:21:55) There might be different methods that will eventually get there. (00:21:57) For example, like, you know, simulation world model, right? (00:22:01) And like it's almost a tautology to say that if I have a perfect world simulator, (00:22:07) Anything can be done there. (00:22:09) Like as long as you can do it in the real world, you can do it in a simulation and you can like, cure cancer in a simulator, right? (00:22:16) But what it turns out for robotics is that some things are harder than others and it really depends on the problem itself. (00:22:24) So in the case of locomotion, as I mentioned, all we need to model in a simulator are point contacts with a somewhat flat ground. (00:22:33) More like feet. (00:22:34) Yes. (00:22:35) But sort of the behavior we want out of it is actually very difficult to model. (00:22:40) Like it's all these reactive behaviors that when you feel like your leg is hitting something, you should retract and step again. (00:22:48) These are very, very hard to describe or try to learn from demonstrations directly. (00:22:53) But in the case of manipulation, I think the difficulty is flipped. (00:22:58) That it's a lot easier to capture the behavior itself, and it's a lot harder to simulate the world. (00:23:07) For example, if you were to grasp a transparent cup with some orange juice in it, it's ridiculously hard to simulate how your hand deforms around the cup and how all those ripplings, how those (00:23:23) like the color of the juice results in like the rendering and what the policy end up seeing. (00:23:30) Simulating that is very expensive and difficult. (00:23:33) But all we need to learn is just to like get your hand to be in front of the cup and then close with the appropriate amount of force. (00:23:41) And that's actually very easy to learn. (00:23:42) That's why like we see so many success of imitation learning in the case of robotics manipulation is because the behavior itself is actually (00:23:52) not as hard as simulating the world. (00:23:56) And that's why we see faster progress there.

Around this claim