ATRIUMsearch → argument graph
MechanismAudio · 18:30 — 19:27

Customer support voice agents are shifting from reactive (customers calling with a problem) to proactive, becoming the front-end of the entire shopping and discovery experience.

Mati describes how voice agents are evolving beyond handling refunds and tracking into proactive shopping assistants that help customers discover products, navigate sites, and complete purchases end-to-end. ✦ AI generated

Mati Staniszewski · No Priors · 2025-12-11 · original ↗

plays this moment only · 18:30 — 19:27

Elicited by

Can you talk about what's happening on the agent platform side?

I think that second exciting piece within that domain, which is happening is the shift from effectively a reactive customer support, I have a problem, I'm reaching out to customer support, into more of like a proactive part of the experience customer support. So to make it explicit, we work with the biggest e-commerce a shop in India, Michau, where they started working on the customer support side, where I want a refund, I want to see the tracking of the package, to actually having an agent be a front part of the experience. So if you go to the website, you have the widget, you can engage it through voice, and you can ask it, hey, can you help me navigate to item X, item Y, or can you explain to me what's the right thing for me to give up for a gift for this period of time? And then it will actually help you based on your questions, based on what is on the offer, show you those items, navigate to the right parts of the piece, maybe go all the way through the checkout.

verbatim transcript · starts at 18:30

Transcript · around this moment

(00:00:06) Hi, listeners. (00:00:07) Welcome back to No Priors. (00:00:08) Today I'm here with Madi Stenesus, the co-founder and CEO of Eleven Labs, which was founded to change the way we interact with each other and with computers with voice. (00:00:17) Over 3 short years, they've skyrocketed to more than 300 million in run rate. (00:00:22) Madi and I talk about the future of voice, (00:00:25) education, customer experience, and the other applications of this voice, as well as how to build a multi-segment from self-serve to enterprise and combined research and product company. (00:00:36) Welcome, Mari. (00:00:37) Sara, thanks for having me. (00:00:39) And thank you for doing this at 7:00 in the morning. (00:00:41) Our pleasure. (00:00:41) Thank you for doing that at 7:00 in the morning. (00:00:43) It's great we got to finally do this together. (00:00:46) I think a lot of our listeners will have used or played with Eleven at some point, but for everybody else, can you just reintroduce the company? (00:00:52) Definitely. (00:00:53) We at Eleven Labs, we are solving how humans and technology interact, how you can create seamlessly with that technology. (00:01:01) What this means in practice is we build foundational audio models, so models in a space to help you create speech that sounds human, understand speech in a much better way, or orchestrate all those components to make it interactive, and then build (00:01:15) products on top of that foundational models. (00:01:18) And we have our creative product, which is a platform for helping you with narrations, for audiobooks, with voiceovers, for ads or movies, or dabs of those movies to other languages. (00:01:26) And our agents. (00:01:28) a platform product, which is effectively an offering to help you elevate customer experience, build an agent for personal AI, education, new ways of immersive media. (00:01:38) But all is kind of underlied with that mission of solving how we can interact with technology on our terms in a better way. (00:01:45) You started the company in 2022. (00:01:48) That's right. (00:01:49) And you've had amazing like rocket ship growth since then. (00:01:51) I'm sure it's felt up and down different ways. (00:01:53) I want to ask you about that. (00:01:55) Can you give a sense of what the scale of the company is today? (00:01:58) So we've grown to 350 people globally. (00:02:01) We started from Europe. (00:02:03) We started as a remote company and are still remote first, but have hubs around the world with London being the biggest, New York being second biggest, Warsaw, San Francisco, and now Tokyo and one in Brazil. (00:02:13) We are at 300 million in ARR. (00:02:17) which is roughly 50-50 between self-serve, so a lot of subscription and creators using our creative platform, and then approaching 50% on the enterprise side using our agents platform work. (00:02:31) And that's on this sales-led, classic sales-led side. (00:02:34) And we serve more than 5 million monthly actives on that creative side of the work. (00:02:41) And then on the enterprise side, we have a few thousand customers from Fortune 500s to some of the fastest AI-growing startups. (00:02:46) I think this is such a, you're an amazing founder, but I also think it's such an interesting company because it is very unintuitive to, I think, many people and investors in particular. (00:02:57) I don't know if you faced this at the beginning, but I remember both there in 2022, there's a class of companies that allow creation in some way when we look at your first business beyond the research itself. (00:03:09) And I would put (00:03:11) Eleven and Midjourney and Suno and Asian in this category. (00:03:15) And I think there's like this overall sense of like, who really wants to do this? (00:03:20) What was your initial read of like how many people want to make voices or what made you believe that was going to be much broader than, you know, like if I look at dubbing, for example, it's not a huge market. (00:03:32) I think first piece was, which is, as you mentioned, there's like a very, (00:03:36) It's very tricky to do both the product and the research. (00:03:39) I'm in a lucky position that my co-founder and I known each other for 15 years. (00:03:44) I think he's the smartest person I know and has been able to create a lot of that research work to be able to create that foundation to then elevate that experience. (00:03:52) But both of us are from Poland originally, and the original belief came from Poland. (00:03:56) It's a very peculiar thing, but if you watch a movie in Polish language, a foreign movie in Polish language, (00:04:02) All the voices, whether it's a male voice or a female voice, are narrated with one single character. (00:04:07) So you have like a flat delivery for everything in a movie. (00:04:09) It's a terrible experience. (00:04:11) It is a terrible experience. (00:04:12) And it's still, like, if you grow up, as soon as you learn English, you're like, switch out and you don't want to watch content in this way. (00:04:20) And it's crazy that it still happens until today in this way for a majority of content. (00:04:25) Combining that, and I worked with Palantir, and my co-founder worked at Google, we knew that will change in the future and that all the information will be available globally. (00:04:35) And then as we started digging further, we realized. (00:04:37) So in every language in a high quality way, that was kind of the starting point. (00:04:42) Exactly. (00:04:42) And the big thing was like, instead of having it just translated, could you have the original voice, original emotions, original intonation carried across? (00:04:52) So like, (00:04:53) Imagine having this podcast, but say people could switch it over to Spanish and they still hear Sarah, they still hear Mati and the same voice, the same delivery, which is kind of exactly what we did with Lex back when he interviewed Narendra Modi and you could kind of immerse yourself in that story a lot better. (00:05:10) So that was the original kind of insight. (00:05:13) And we then started digging further, which is that (00:05:17) just so much of the technology we interact with will change. (00:05:21) Whether this is how you create, it's still relatively tricky to bring voice alive. (00:05:26) You would need to go through the expensive process of hiring a voice talent, having a studio space, having expensive tooling to then actually adjust it. (00:05:35) The tooling isn't intuitive to be able to do this. (00:05:38) So like all that creation process will and should change to make it easier for new people with keenness to bring that alive. (00:05:45) Then a lot of the (00:05:46) technology wasn't possible for you to be able to recreate a specific voice or be able to create that in that high quality way. (00:05:54) And then of course, as we dived into further and shifted away from the static piece, the whole interactive piece is still crazy in the way it functions where most of us... (00:06:03) seen this technological evolution over the last decades, but you still will spend most of your time on the keyboard, you will look at the screen, and that interface feels broken. (00:06:14) It should be where you can communicate with the devices through speech, for the most natural interface there is, one that kind of started when the humanity started, and we realized we want to solve that. (00:06:27) And (00:06:28) I think now fast forward from 2022, I feel like many people will carry that belief too that voice is the interface of the future. (00:06:34) As you think about the devices around us, whether it's smartphones, whether it's computers, whether it's robots, speech will be one of the key ones. (00:06:40) But I think 2022, it wasn't. (00:06:43) And as you think about the market for the creative side or for that interactive side, it was very clear it will be a huge, huge one. (00:06:53) So even when you think about just the (00:06:56) research part of your business, and then you have products for at least two different markets, and then you have this larger mission. (00:07:02) A lot has changed the last five or 10 years, but it used to be like a very strongly held traditional belief of like, one must do one thing well in a startup and there's no other path. (00:07:11) Like you're treating this like interaction company, platform company. (00:07:15) How did you think about sequencing like the research and the product effort? (00:07:20) Does that make sense? (00:07:20) Or like thinking about new markets? (00:07:22) And maybe wrapped up in that question too is just like, (00:07:25) where are we in quality on voice as well? (00:07:28) Because if I would sort of claim like the models are not good enough for certain use cases at all, like it kind of doesn't make sense. (00:07:36) Do product? (00:07:37) And I think that's right. (00:07:38) It's almost exactly like when we started. (00:07:40) Originally, what we did was try to actually use existing models that were in the market and kind of optimize them for our first use case was actually starting with combination of narration and dubbing and then on the creative side. (00:07:53) And (00:07:54) We realized pretty quickly that the models that existed just produce such a robotic and not good speech that people didn't want to listen to it. (00:08:02) And that's where my co-founder's genius came in, where he was able to assemble the team and do a lot of the research himself to actually create a new version of creating that work. (00:08:12) But to your question, I think that the way we are kind of organized internally and how we think about sequencing a lot of that was looking at the first problem, (00:08:20) and then creating effectively a lab around that problem, which is like a combination of mighty researchers, engineers, operators to go after that problem. (00:08:29) And the first problem was the problem of voice. (00:08:31) So how can we recreate that? (00:08:34) And the voice, and like you say, it needs to have that research expertise to be able to do that well. (00:08:38) So we started with effectively a voice lab, which was that mission of, can we narrate the work in a better way? (00:08:45) There was a combination of roughly 5 people that were doing that work. (00:08:49) And then sequence the research 1st and then build a simple layer on top of that work to allow people to use that work. (00:08:56) And then kind of expand it from there with a holistic suite for creating a full audiobook and then creating a full movie narration. (00:09:04) movie dub. (00:09:05) And then we moved to the next problem, which was the realization that, okay, we have solved the voice, great for making content sound human. (00:09:13) The first problem, for that to be useful for us to interact with the technology, you need to solve how you bring the knowledge on demand into that. (00:09:21) So we effectively started then the second team, which was a second lab, an agent lab effectively, which was a team that would combine researchers, engineers, and operators once more, which would try to fix, okay, we have text to speech, (00:09:34) how do they now combine this with LLMs and speech-to-text and orchestrate all those components together while integrating that with other systems to make it easier? (00:09:43) And then similarly, you know, you kind of expand from looking just at the voice layer into how those systems work together. (00:09:49) And here too, you need the research expertise to do that in a low-latency way, efficient way. (00:09:56) accurate way. (00:09:57) But at the same time, there's that product layer that starts forming that it's not only the orchestration that matters. (00:10:02) It's also the integrations of how you link up to the legacy systems, how you build functions around it, or how you deploy that in production and test, monitor, evaluate over time. (00:10:12) Do you feel like you were creating new use cases when you built the tools? (00:10:15) Do people know that they wanted to do this already? (00:10:19) Because one argument that I remember hearing was like, (00:10:23) enterprises don't know what to do with voice. (00:10:25) How many people really want to do it? (00:10:26) And then you're serving essentially like perhaps the like creator publisher side of your business. (00:10:31) Yeah. (00:10:31) It's definitely a combination of like initiatives that we believe will happen in the world and then like response to a lot of that. (00:10:37) So like as I think back, we, you know, of course voice, the internal voice lab or agents lab, then kind of that kickstarted so many of the other labs (00:10:46) in response to the problems. (00:10:47) We started a music lab because people wanted to create music with Eleven Labs. (00:10:52) So it's a fully licensed model where people wanted to use and create speech, but they wanted to add music in a simple way. (00:10:59) We wanted to deliver that. (00:11:00) And then, of course, that kind of came together through how do we combine music, audio, sounds. (00:11:06) We are now integrating partner models from image and video into that suite is how could you combine all of that in one? (00:11:13) And all of that was in response to the market saying us like, hey, we would love to (00:11:16) And then you will have completely different use cases, even in that space. (00:11:19) Let's say dubbing. (00:11:20) Dubbing is a use case that we didn't feel there was like a big push for that, but we knew that in the ideal world in the future, you will be able to have that content delivered naturally around the languages, still carrying that. (00:11:34) And I still think actually this market will be immense because it's not going to be only the static delivery in movies, but if you (00:11:41) travel around the world and want to communicate in real time, like the full babblefish idea from Hitchhiker's Guide of the Galaxy, this will happen. (00:11:48) It will be like the biggest, like the whole breaking down language barriers, the barriers to communication, to creation, like all of that will break. (00:11:56) And that will be like the foundational real-time dabbing concept. (00:11:59) So (00:12:00) I'm super excited about that part. (00:12:01) And similarly, on the agent side, you are like some obvious things that of course customers that we work with or partners will want to want to integrate, which is we want integrations with XYZ systems. (00:12:14) But then there are like other parts that might not be as easy to predict of as your interactive technology, of course want to understand what's happening, but you also want to understand how (00:12:24) the things are being said and bring that into the fold, which would be something we try to prioritize on our side. (00:12:28) So then the people, when they actually interact with the technology, they realize, oh, expressive thing is actually so much more enjoyable and beneficial and helpful. (00:12:36) So I want to ask you a question about this, which relates to quality. (00:12:40) You know, I work with a series of companies where we're (00:12:44) Selling a product to the buyers are generally not machine learning scientists, right? (00:12:50) And even the scientific community does not have the full suite of evals and benchmarks to understand every domain well. (00:12:57) There's a well-known problem, but I imagine for a lot of your customers, it's not like they know how to choose good voice. (00:13:03) So how do you deal with that problem? (00:13:05) Like, is it like, hey, I make a clone and that sounds like me and I believe it. (00:13:10) I'm going to try all of these different options or actually are you teaching (00:13:14) people who do eval? (00:13:16) It's a great question because I think there are two big problems. (00:13:19) One is how do you benchmark the general space in audio where, like you say, it's so dependent on the specific voice, let alone if you are training into interactive, then it's even more tricky. (00:13:32) And then the second piece, which is as you are working on a specific use case, how you select a voice. (00:13:36) So I'll take the second front first, which is we have like a voice sommelier effectively with us. (00:13:41) We work with enterprises. (00:13:43) We deploy that person to work with them and help them navigate. (00:13:48) That person is like a voice coach, has an incredible voice themselves. (00:13:51) And now we have like a team under that person that like will partner to help you find what's the right branding of branding voice. (00:13:58) And now you have like the celebrity marketplace. (00:13:59) And now you have a celebrity marketplace to like help. (00:14:02) you even get that iconic talent in there like Sir Michael Caine. (00:14:06) That piece was important because of course the voice will depend on the use case that you are trying to build, the language, all of that will have an impact of what's the right voice for your customer base. (00:14:15) So we have effectively a voice person helping those companies. (00:14:21) And some companies will be very opinionated on when they want. (00:14:23) So they will (00:14:25) sometimes selected themselves, sometimes give us a brief of, hey, we want a voice that sounds professional, neutral, is coming. (00:14:32) We recently had a company, one of the biggest European companies that wanted, that gave us a brief, which is very original, that they wanted as a robotic voice as possible. (00:14:43) Okay. (00:14:43) Which was counterintuitive. (00:14:46) Yeah. (00:14:46) So you're like, we can't do that anymore. (00:14:49) Almost, but then we are like trying to go backwards of like, how do we do that? (00:14:52) But I think we got a good result. (00:14:55) But recently we had a company in Japan where Japan and Korea, where they wanted to serve different voices depending on the customer that's calling in. (00:15:04) They have an older population and a very younger population. (00:15:08) The younger one, they wanted like one of the famous voices in the market that's very excitable and happy. (00:15:13) And for the older one, they wanted like a calm, slow speaking one. (00:15:15) We help a lot with that. (00:15:16) So that's on the voice piece. (00:15:18) And I do think it's going to be a big important one. (00:15:20) It's like a personalized choice and then it can even be dynamic in a customer. (00:15:24) Yes, exactly. (00:15:26) And then maybe in the future, it's like going to be like a fully, depending on your interaction, you will have a voice created. (00:15:32) It's as we understand the preferences of what people want. (00:15:35) So, you know, like let's say you're in the evening and you are tired and you want a slightly different, or maybe not. (00:15:40) Maybe that's like the best focus time that you have like a voice that's giving that energy and probably it's different when you wake up and gives you that morning news of what's happening or what's the weather. (00:15:50) So like all of those could be different. (00:15:53) Yesterday we had a dinner with some of our partners and one of them, the first thing they said is like, hey, I have a new request for you. (00:16:01) I want a New York voice with a Long Island voice accent, which I never knew is a thing. (00:16:06) And it's supposedly a thing. (00:16:07) So we have that. (00:16:09) And then on the first piece, I think it's unsolved problem still, where I think you have a good benchmarks, of course, in LMs, I think in image space, they are pretty good. (00:16:18) In voice space, you have, of course, the speech quality, but then so much of (00:16:23) whether you like or not, the speech depends on the voice. (00:16:26) That just if you compare model A to model B and you serve them different voices, even if the quality is very different, the voice itself can just make that sort of different. (00:16:36) We've seen this, I don't know if you know artificial analysis, benchmarks, I think they're pretty good. (00:16:41) Just switching the voice makes that, it's such a big thing. (00:16:44) That's so interesting. (00:16:45) Yeah. (00:16:45) And I wonder if, as you said, this is (00:16:48) the most dominant interaction mode we'd had for millennia of all of human history, right? (00:16:54) I'm biased and self-serving, but I think so. (00:16:57) We're just very sensitive to it. (00:16:58) And I think people are going to be very sensitive to their own personalization as well. (00:17:04) 100%. (00:17:04) I think there's also a third piece which (00:17:07) Maybe it's not directly to your note, but we've also realized that you have, so you have the benchmarks, you have like, how do I find the right voice for my audience? (00:17:15) But even the understanding of how you describe audio data is still lagging in the industry. (00:17:21) Like when we initially started, we of course went into the traditional players for them to help us label not only what was said, so like transcription, but also how it was said. (00:17:29) Like what are the emotions, use accent. (00:17:31) And most people just weren't able to do that work effectively because you kind of need to (00:17:37) here and have a little bit of a skill set of like how would I describe this specific delivery? (00:17:42) So we need to create that ourselves. (00:17:44) So I think that is that piece as well of like how do you effectively interpret the data of audio in a more qualitative basis? (00:17:52) That's yeah, trickier. (00:17:54) Can you talk about what's happening on the agent platform side? (00:17:59) Like what is challenging for (00:18:02) businesses or even creators that are trying to build agents and what maybe what the surprising or high traction use cases are. (00:18:08) I think everybody's kind of aware of the idea of agent-based customer support, but I imagine you're doing many things beyond that. (00:18:15) Yeah, so exactly. (00:18:16) Customer support is probably the one that's kicking off the quickest, and that's the one that we see overtaken so many use cases. (00:18:23) but it's where I work with Cisco or Twilio or Telus Digital, all of them are kind of elevating that to a high extent. (00:18:30) I think that second exciting piece within that domain, which is happening is the shift from effectively a reactive customer support, I have a problem, I'm reaching out to customer support, into more of like a proactive part of the experience customer support. (00:18:46) So to make it explicit, we work with the biggest e-commerce (00:18:50) a shop in India, Michau, where they started working on the customer support side, where I want a refund, I want to see the tracking of the package, to actually having an agent be a front part of the experience. (00:19:03) So if you go to the website, you have the widget, you can engage it through voice, and you can ask it, hey, (00:19:10) can you help me navigate to item X, item Y, or can you explain to me what's the right thing for me to give up for a gift for this period of time? (00:19:18) And then it will actually help you based on your questions, based on what is on the offer, show you those items, navigate to the right parts of the piece, maybe go all the way through the checkout. (00:19:27) And I think this will be a phenomenal thing of like elevating the full experience where that's more of an assistant across the whole thing. (00:19:33) We kicked off our work with Square that enables all the businesses to do that work, exactly the same pattern. (00:19:38) Started with voice ordering, (00:19:40) how can now this be part of the full discovery experience too, where you get items shown to you, can have a lot more explanation, which I think will be a phenomenal piece where it's effectively from the beginning to the end. (00:19:52) That's one category. (00:19:53) The second one is the wider shift from static to immersive media, where there's just so much incredible stories in IP that today exist in effectively one way of delivery, and now you'll be able to interact with that (00:20:08) content in a completely new way. (00:20:10) I think one of the incredible use cases was working with Epic Games. (00:20:14) We worked with them on bringing the voice of Darth Vader and Darth Vader into Fortnite, where millions of players could interact with Darth Vader life in the game, where you had like a full experience of Darth Vader in a new way. (00:20:29) And I think this will be a theme across, whether it's talking to a book, (00:20:33) talking to the character that you like, to the whole space shifting. (00:20:38) And then I think the one that I'm most excited about for the world and for the shift is going to be education, where you will just be able to have effectively a personal tutor on your headphone and could actually study something in an amazing way. (00:20:55) I'll give you like 2 quick examples. (00:20:57) One is (00:20:58) We recently worked with Chess.com. (00:21:01) I'm a huge fan of chess. (00:21:02) I'm a huge chess fan. (00:21:03) Okay, great. (00:21:04) So you can learn chess, but you can have Hikaru Nakamura or Magnus Carlsen be your teacher of how you deliver that, which is amazing. (00:21:13) Or even Botta's sisters, or it's like all the plethora of different players that engaged with that, which I think is great. (00:21:19) And then maybe a last one, which is a master class we worked with to shift from, you can of course have the content and go through step by step. (00:21:29) But you can also have an interactive experience. (00:21:31) And the best example of that was working with Chris Boss, the FBI negotiator, one of the top negotiators, who has a masterclass lesson. (00:21:38) But then you can actually call him and have a practice negotiation, which is crazy. (00:21:43) Yeah, got to get that hostage out. (00:21:44) We'll definitely try it. (00:21:46) Yeah. (00:21:47) Can I add one more? (00:21:47) I think the one last one, which combines all of them together, which I realized just recently is, which was crazy. (00:21:54) So recently I went to (00:21:57) to Ukraine, where we are working with Ministry of Transformation, where they are effectively creating a first agent in government. (00:22:03) And the crazy thing is they have all of those. (00:22:05) Agent in government. (00:22:07) Agent in government. (00:22:08) So they want to like rechange of how they run all the ministries. (00:22:13) And it sounds. (00:22:15) Like a big ambitious goal and lofty. (00:22:17) No, I think the baseline is like here. (00:22:19) So actually I'm by that immediately. (00:22:21) Yeah. (00:22:22) And the crazy thing is I think they are like so ahead and actually doing that. (00:22:26) And I think they are like 2 concrete things there. (00:22:29) One, they kind of combine all those use cases. (00:22:32) So they, we are looking into how they can have effectively customer support of government, whether it's asking about benefits or employment, about the process of how you leave the country. (00:22:42) All of that. (00:22:44) be run through effectively a digital app, then two, how you can have a proactive way of informing citizens of things that might be happening, but then having an education system that also run through this personal tutoring experience. (00:22:55) And all of that is happening. (00:22:57) So that was incredible to see. (00:22:58) And the second amazing thing was that the way they've done it, so they have the digital transformation piece, (00:23:03) But they have engineering leaders in each of the ministries that lead those efforts and then bring them back to that one central piece. (00:23:09) So that is like incredible to see and also proud to be able to be working with them on that shift. (00:23:16) But despite everything that's happening, they're like so amazing. (00:23:19) That's amazing. (00:23:20) That's really encouraging. (00:23:21) Can I ask you a business model question here? (00:23:23) Because looking at the strategic landscape, actually I have many questions here. (00:23:28) One of the observations I'd have is if I look at one of these rich voice and action agent experiences, it's a lot of, let's say, Fortune 500 Global 2000 leaders who listen to the pod. (00:23:42) I think a lot of them are going to buy the idea of like, I want this amazing, automatic, real-time available, 24-7, every language experience for my customer. (00:23:55) That's (00:23:56) consistent and high quality. (00:23:58) The ways I...

Around this claim