ATRIUMsearch → argument graph
MechanismAudio · 10:05 — 11:34

The enterprise harness defines the models, data, and tools in a loop — the hard lesson is that prepping the context layer is where the magic is.

Satya describes the enterprise AI harness as a multi-model system with tool access and rich context, where the hardest and most valuable work is preparing the context layer so plans execute efficiently. ✦ AI generated

Satya Nadella · No Priors · 2026-06-04 · original ↗

plays this moment only · 10:05 — 11:34

Elicited by

What is the harness for the enterprise? Is there an equivalent concept for broader productivity work?

In some sense, you kind of want the harness to define the models, the data, and the tools, and so that you have a loop across those three. And so what we are trying to, first of all, make sure is each of our products that we build, right, whether it's GitHub Copilot or the security, the stuff we showed with MDash, or even the Discovery for Science, it doesn't matter. All of them are multi-model harnesses with tools access so that you can do this progressive disclosure of tools even so that they're token efficient. And then you're feeding it with very rich context, because that's sort of the other hard lesson we have learned in the last two years is, oh my God, the amount of work you need to do to prep the context layer such that your plan can execute in the most efficient way is where the magic is.

verbatim transcript · starts at 10:05

Transcript · around this moment

(00:00:00) The world is going to be very skeptical of tech and tech companies that say, trust us, we've got it. The future is going to be glorious. You kind of have to deliver tangible benefits because it's too important this time around. It's too much of the economy for it not to be the case. True ambition is about making the impossible possible. I take great inspiration from sort of the people who were managing the Azure network. We built in the last 15 months more Azure capacity than we built in the first 15 years. I mean, it's crazy. (00:00:30) Our job is not to do Azure networking. Our job is to build the agentic system that does Azure networking. The way to get to information, way to educate yourself, way to continuously keep yourself updated has changed so much. Maybe the next big startup could be someone who builds a new university, a new pedagogy even of how to get someone to go through a curriculum and find economic opportunity that's highly valuable. (00:01:12) Please welcome SWX, Seregawa, Alaud Gill, and Chairman and Chief Executive Officer of Microsoft, Satya Nadella. Hello, (00:01:32) I'm so excited to be here. Welcome to a crossover episode of No Priors in Lane Space with Satya Nadella. Congratulations on an amazing build. No, thank you so much. And it's great to be with both of you. I listen to both of you or both the podcast all the time. It's great to be on it. Thank you so much. So you're just talking about these amazing announcements from across the Microsoft estate all morning for I think 3 hours. What is the what's the most important reflection or takeaway you have? I'd say there are. (00:02:02) Perhaps the biggest one for me is, let's sort of conceptualize this more as an ecosystem play as opposed to a single model or even a single platform, right? I mean, whatever, at least for me, having grown up at Microsoft, having seen whatever, four major platform shifts, I sort of fall into that (00:02:26) a camp where a platform is defined by fundamentally its ability to create more value about the platform versus what's captured in the platform. And so if you view what's happening right now, I think this morning's keynote was, how can any company, whether it's an AI native company or a traditional enterprise company, participate as a first class participant where they can point to AI they created? (00:02:55) right? It's not that they don't use other people's AI. Of course they will. But to me, what's the path? What's the recipe? How do I do it? What does a stack look like? What does the tooling look like? What is valuable? How do you do that? That's it. That's sort of our job to do. Yeah. Ecosystem strategy is very complicated, right? Because you end up building certain components, partnering for certain components, supporting them. You just announced this big suite of models. Like tell us a little bit about the (00:03:26) Yeah. So the thing that we wanted to do with the MAI models was to build, and as Mustafa talked about, first of all, a great lineage, right? Starting with pre-training, with very good data quality, doing all the ablations, making sure, because in some sense, it's becoming even harder to build a clean lineage model, because there's so much stuff out there. (00:03:54) that you truly need to ablate out to be able to have a fantastic pre-trained model. In fact, that's one of the challenges of a lot of the open weight models is they look great on one benchmark or two, but they're not great on practice. So that's why, in fact, even in our FDEs are pretty gone really excited about these MAI models, because how the heck can a small 5B model hill climb (00:04:19) And it goes back a little bit to what I think is ultimately the key thing to do, which is try to pursue finding that cognitive core. So to me, starting with a clean lineage, then creating that ability for companies to be able to use this, right, not just as a generalist, but to create their own specialist by building this hill climbing scaffold around it, right? So it's not just the model, (00:04:47) but you have a hill-kline scaffold around it, then you will start building your RLE. You will start collecting the traces. Most importantly, you'll have private evals because we know all the evals out there are good, interesting, but they're not really that critical at this point because they all can be maxed. And so the point is each company will have its own private eval. And so that end-to-end platform story around our models is sort of (00:05:15) what I think is interesting. And then the one other thing, Sarah, since you brought that up is I do feel there's a new frontier. Like people talk about the frontier, and are you operating at the frontier? Interestingly enough, if you add a little temporality to it, you can use, let's say, in fact, the Land O'Lakes demo we showed was pretty cool. We used whatever, GPT-55, right? Then you collected a bunch of traces, and then you took a 5B reasoning model and achieved higher. (00:05:43) So that is another aspect of what it means to appear, operate at the frontier. Yeah. I think, first of all, I have to congratulate you on basically building a Frontier NeoLab inside of Microsoft in two years. I'm wondering, you know, you have all this AI strategy that you're rolling out. I'm wondering, what do you know now that you wish you would tell yourself two years ago, or two, three years ago, three years for the Jensen partnership, two years for (00:06:07) I mean, I think the thing that I reflect quite a bit, right, which is sort of obviously I got into all this when I got excited by the scaling loss paper and when even the OpenAI partnership came about when those folks said, hey, we're going to really throw a lot of computer transformers. And they've helped. The thing that I always go back and say, wow, these things do have capability that they're climbing up. I mean, this (00:06:34) this crude way of saying it is intelligence is log of compute kind of works. Now, what I think we underestimated perhaps is the real world complexity of deploying these so that they actually deliver the value in the real world, right? So the outcomes as measured by any benchmark, (00:06:56) is interesting, important, but the true eval is when people out there are able to do unique things that they only can value. And it's very measurable, right? That I wish we had sort of even like had more in our consciousness, right? Which is as an industry. Because right now, I think when people say, wow, I don't want a token max, it's an artifact of us not having thought ourselves as an industry. (00:07:25) that we are using tokens to create value every step of the way. So I think that's kind of what I wish we had gotten there, but I'm glad we are here. What are some of the use cases that you've seen that have created the most value for your customers? Because I know that people talk a lot about code, and I think it's pretty clear that that's something that's having very large scale impact. Are there other areas that you find in common that your customers are really benefiting? Yeah, I think to your point, obviously coding, (00:07:49) is now got, but it's interesting, by the way, to even talk about the coding, right? Which is coding has worked so well that we now have to rebuild the IDE, right? I mean, it's kind of nuts to see what we launched is like, oh my God, I have these 100 agent sessions. The cognitive load, it transfers back to me as a human is so excessive that now I need a new UI. Oh, by the way, (00:08:14) the chat as the only artifact is also impossible. So that's why we need a canvas. So it's kind of interesting for all the things about where is software needed or where is UI needed. You kind of need that even for code, right, in a fully agentic world. But that said, one of the things that we are starting to see, we start seeing the co-work, but even some of the work we showed with autopilot, right, on what you see with clause, is a good one. Because if you sort of think about (00:08:43) A lot of human capital is doing the glue work, right? If you now can augment that with tokens slash agents that are long running, durable, right? Then your ability to scale, even what is still judgment and glue work gets amplified like coding does. So you can, I'm positive that six months from now, we'll all be saying, oh, wow, (00:09:11) all through the night, there was a bunch of stuff that all these autopilots that I have working on my behalf with my delegated authority, so to speak, right? I can sort of given even my identity, did a bunch of work. Then of course, I'll need my new ADE to say, what did you do? Did I do this work? And so on. So I think that that's where compressing of workflows, completing of tasks, (00:09:36) That's where I think a lot of the value gets created. I think you raised a really interesting point, which is there's the actual agent is doing the code and then there's a harness around it. And that's the environment, that's the context, that's everything you're setting up as a developer around actually a coding agent. What is the harness for the enterprise? Is there an equivalent concept for broader productivity work? Or how do you think about that concept sort of? That's right. So in some sense, you kind of want the harness to define the models, the data, (00:10:05) and the tools, and so that you have a loop across those three. And so what we are trying to, first of all, make sure is each of our products that we build, right, whether it's GitHub Copilot or the security, the stuff we showed with MDash, or even the Discovery for Science, it doesn't matter. All of them are multi-model harnesses with tools access so that you can do this progressive disclosure of tools even so that they're token efficient. (00:10:33) And then you're feeding it with very rich context, because that's sort of the other hard lesson we have learned in the last two years is, oh my God, the amount of work you need to do to prep the context layer such that your plan can execute in the most efficient way (00:10:54) is where the magic is. So we have, in our case, we have the GitHub harness, which essentially we're using across all our products. It's available in Foundry and we're open, like you can use your Llama harness, whatever, or you can use the, you know, any open harness or any harness of yours and train with your tools and multiple models and your context. And so that's the pitch, because right now a lot of dialogue is (00:11:17) hey, if I train the harness plus tools and the model together, you get evals. And what we are proving out is, and the best example of that is what we did with M-dash, right? Because when it launched, it found bugs or vulnerabilities that were not found by mythos. And so there is existence proof. I would claim that you can have a multimodal harness that can in fact be more (00:11:46) performant in the real world. So a premise behind the training at the independent Frontier Labs is really, you know, we're going to have these models and we'll have an API business and we'll support enterprise and (00:12:00) startups, but a first party product, be it productivity or code or search, drives the majority of revenue. That's a different value equation than you're describing, I think, with the Microsoft ecosystem. If that's the case, tell me if it's the case, because obviously you have first party products and you have enablement products. What is the role of the developer, like what's going to be hard and the set of skills and the value capture the developer has in that world? Yeah, so I think that there's always going to be the case that someone who is super successful (00:12:30) and as a platform builder, can also have first-party products. It was true with Windows. It was true with the SAS side and the cloud side as well with us and others and so on. But the thing that is, it should not be a limiter to other people achieving that same success, right? That I think is the core difference, which is the network effects this time around, intelligence as such, because they learn from data, (00:12:59) and not really lots of data. It's just a few samples that you have to see to understand what's novel about something. So that's why the game becomes how to protect. So that's why I would say every company, having private evals, maybe the biggest IP, right? I think about it, like what's that private eval that you can then use even a frontier model to hill climb on and not leak the traces, maybe one of the biggest drivers (00:13:28) of IP. So in other words, another acid test is you have an eval that's private. You're using a model A. Can you switch it to model B and climb up? If you can, then you're in control. If you can't, you're not in control. And that's where even the harness decision becomes super important. So therefore, having an open harness, letting all models come in, (00:13:54) having your evals, your context, your tools help you hill climb, I think is the skills that an AI native startup needs, a SaaS company needs, or every enterprise needs. Yeah, I think in a very real way, you are, Microsoft historically as an operating systems company, and then become a cloud company, maybe like the third act is that you're a harness or evals company, whatever the sort of conglomerate of concepts that you want to put together. (00:14:24) I think enabling every company to have frontier intelligence or the exact term that you use is the mission, right? That's it. That is the platform promise that you build with us. You will get your intelligence for your data. That's it. To me, that is the, like if there was one tagline for this entire developer conference is, can everybody operate (00:14:49) at the frontier with their frontier intelligence, right? To me, that is so important because otherwise, I don't know how you achieve stable equilibrium, right? Which is how do I then go and say, wow, my company is going to have a terminal value because I now know how to continuously compound on top of what's a platform that gets better, right? So when like Windows obviously came out, Adobe built, Autodesk built, (00:15:17) Or even like take what Jensen said, we built DX and he built, CUDA on top of it. Right? I mean, I always say to Jensen, God, I got the short handle of that, right? I wish we had recognized it. But nevertheless, but that idea that you can build a platform layer that someone else can then extend out (00:15:38) and build their own intelligence layer in this case, I think is everything, right? Without it, why have a developer conference? I can just come and have you all sort of just worship at the altar of 1 model. But that's not a developer conference. Backstage, we had a discussion about what is IP or what is the value in a company. It used to be the length of human experience at a company. And now it's this other thing, which is the evals, the (00:16:03) experience in sort of applying agents to the company. I just want you to like fresh it out a little bit, Marcus. Yeah, it's a great way to frame it, right? Because at the end of the day, every company is going to have both the human capital that is still going to be super valuable because humans and their ability to find (00:16:20) the gaps that exist at all times is going to be the way we all will create value, right? I mean, so I'm definitely in the camp that this is going to be about expressing new forms of human agency and ambition, even as token capital goes up, right? So let's say any corporation has lots of tokens and a lot of human capital. The question is, how do you compound the two? (00:16:44) So if you have a, if you take in teams, I have a bunch of agents doing work and a bunch of humans doing work, and the traces between those, that is really important context of how that enterprise is creating value. Then that goes back to train not a generalist model, but to train the company veteran agent. (00:17:06) right? That is super valuable again, right? Which is when a company says it should in fact go onto the balance sheet is how I think about it, right? That's what, in fact, there may be like human capital was never possible to go put on a balance sheet because you didn't know how to capture the tacit knowledge. Whereas now I think you can with the agents that have learned through time, through all the traces. So that's what at least we think will happen. I think the SEC is going to have to have accounting standards for token (00:17:36) expertise. You're talking about the equilibrium state and a stable equilibrium where companies have this compounding value and can see terminal value for themselves. Another challenge to the considered equilibrium of, okay, there are applications and workflows that are sort of common to a vertical or a horizontal. (00:17:58) And this was like the generation of SaaS companies. And Microsoft has lots of SaaS properties as well. And then there are things that are very specific to every enterprise that they're differentiated against. I'm sure you have heard much and participated in much of the debate about the end of software, because all these workflows are cheap to generate now. Do you think the equilibrium looks different between what agents (00:18:20) get built in enterprises versus in their vendors in the future? Yeah, so I think what's happening there is, see, we had a particular way we captured, I would say, workflow in apps, right? Because we built a data model. (00:18:37) right? We schematized some part of some business process. We then built a bunch of business logic, and then we put a bunch of UI on top of it, right? So that's kind of what every SaaS company. And a little configuration. For like 20 years, that was. And that was it. So interestingly enough, now you kind of get to re-litigate that vertical stacking, right? So I still think, for example, (00:19:00) That data model that you build underneath every SaaS application is super good, right? Like why reinvent it? Like my general ledger better be a general ledger. I don't need new schema creation. In fact, that entity relationship is actually pretty good, robust thing that I want to feed. And you want it to be stable. That's right. Then same thing with business logic, right? If you look at (00:19:24) We have this product called Power BI, right? It was like dashboards galore people created. The beauty underneath that dashboard is a very rich semantic model, right? Someone took the pain to create a dashboard and do all the measures. And you want that, that's business logic, right? I want that to be available to me. (00:19:46) So I think the challenge of the SAS business model is we packaged one way. We now have to learn how to unbundle these things and rebundle in new ways and discover new business models, right? I mean, if you look at it, what's happening today with Microsoft 365 is a great example, right? We have this thing called work IQ. (00:20:07) In fact, what we are realizing is, my God, if I look at it, in fact, there's a historical parallel to, right? (00:20:13) We sold first Exchange and SharePoint, and before Teams, we had a thing called Link Server and what have you. (00:20:21) And we thought, that's all going to move to the cloud. (00:20:22) But little did we realize that the number of people who will use servers in the cloud is 10x, 100x, right? (00:20:30) Because people were not buying servers, they were just buying a subscription. (00:20:33) The same thing is now happening with M365, because with work IQ, (00:20:38) we've exposed what is perhaps the most important database in a company that never got used as a database because it was only captive to our apps, right? (00:20:48) It was all e-mail operated on it, Teams operated on it, Word, Excel, PowerPoint, SharePoint. (00:20:53) But now, (00:20:54) This is one of the coolest things I get to do with work IQ. (00:20:57) I go to a GitHub repo and I say, hey, I attended a bunch of design meetings last week related to this repo. (00:21:04) Can you capture all that and tell me what changes I should make? (00:21:07) I mean, think about that. (00:21:09) It literally can go look at all those transcripts, come back with a plan to change a code base. (00:21:15) Previously, you could never have thought of using M365 for something like that. (00:21:20) So the value creation opportunity now in the agent world is in fact 10X more. (00:21:26) But it does require us to have, for example, there's going to be usage around M365, which is going to be perhaps more than even the end users. (00:21:35) And we have to even re-architect. (00:21:36) In fact, what I use to serve an inbox or a mailbox cannot be used to serve an agent. (00:21:45) And so that's sort of what we're doing. (00:21:47) I don't believe in permanent business models for any of these domains, but in the near term, do you have a prediction between outcomes-based pricing, token-based pricing, enterprise bundles? (00:22:01) Yeah, the way I think about this is always we've had, like, let's even take the per user pricing. (00:22:08) The per user pricing is (00:22:10) really an artifact of someone creating a budget needing certainty, right? (00:22:16) Because it's the most important thing. (00:22:18) Like somebody wants a budget, they need a per user. (00:22:21) And per user is just a set of entitlements to usage, right? (00:22:25) That's kind of what it is. (00:22:27) And so the way is, the first bundling will be, take some usage, bundle it into per user stacks, and then sell subscriptions. (00:22:35) So subscriptions, I think, are going to be there, per user is going to be there. (00:22:39) Then the next big thing will be consumption. (00:22:41) So people will say, I want consumption. (00:22:43) And it's also possible that people will say, I don't even want to pay for any of the subscriptions or the consumption's outcome. (00:22:49) But remember, most people love outcomes until they have an outcome. (00:22:53) Because once you have an outcome, it's like giving away royalty, right? (00:22:56) I mean, I've talked to customers who love, you know, outcome-based pricing. (00:23:00) I say, I'm all in. (00:23:02) Until they, my God, like, what are you talking about? (00:23:04) You're sharing in my outcome? (00:23:05) No, I want you to go back to per user pricing and I want you to consumption price, right? (00:23:10) So I think that debate will go on. (00:23:13) And all of these business models have a particular time and a place versus one to rule them all. (00:23:20) And if anything, if you're a SaaS vendor or you're a platform vendor, having that flexibility, and quite frankly, we face this with GitHub, right? (00:23:27) We just recently announced a per user pricing on GitHub. (00:23:31) Because little, GitHub Copilot was constructed at a per user level before we understood even the intensity of usage of agents, right? (00:23:43) It was an interactive way for a developer to use code complete, maybe task. (00:23:48) It was not like, oh, I launched 10,000 sort of agents that are going on all day, right? (00:23:54) So that is what the adjustment is about. (00:23:57) So now that we really want,

Around this claim
Mechanism · 4
This moment responds to