ATRIUMsearch → argument graph
Video · 2026-07-02 · 2h 51m · 48 moments

Fable's Back, AI Engineer Recap, & SambaNova

✦ AI generated

timeline · colored by role

01
Context

The full story behind Fable's (Claude Opus's) sudden shutdown — including whether Amazon panicked and called the government, and why — is still unclear and deserves more public disclosure.

Nathan says key facts about why Fable was pulled — including Amazon's alleged role in alerting the government — remain unexplained and should be discussed more openly.

transcript

Nathan: I mean, we're still kind of piecing it all together through reporting and hearsay and, you know, what exactly was it that caused this in the first place? Was it really Amazon that called the government in a panic? If so, like why? And you know, were they confused about what was going on? That would be weird.

02
Context

We still don't actually know what caused the government panic over Fable/Anthropic in the first place — whether Amazon really called the government, why, and whether there was some strange intrigue involved.

Nathan says the reporting on why Anthropic's frontier model got pulled and how the government got involved is still murky, with unresolved questions about Amazon's role.

transcript

Nathan: You know, a lot of open questions still about this, including like generally what the hell happened, right? I mean, we're still kind of piecing it all together through reporting and hearsay and, you know, what exactly was it that caused this in the first place? Was it really Amazon that called the government in a panic? If so, like why?

03
Context

Nobody outside Anthropic and the government actually knows what triggered the panic that got Fable pulled — whether it was really Amazon sounding an alarm, why, and how the situation was resolved.

Nathan says the reporting on why Fable got pulled (possibly triggered by an Amazon report to the government) is still full of gaps, and he wants a more public accounting of what actually happened.

transcript

Nathan: I have a hard time telling a story that doesn't leave me with weird open-ended questions around what was seen, why was it reported in this way, how did we end up in this sort of panic situation, and then how was it fixed?

04
Context

Anthropic mostly reassured the government about Fable's risky capabilities not by fixing anything new, but by pointing out that the observed capabilities had already existed in other models in the wild for some time.

Nathan describes the strange resolution of the Fable pause: Anthropic calmed government concerns largely by arguing the flagged capabilities weren't new, already existing in competitor models.

transcript

Nathan: the best information I've seen suggests that the way that they mostly reassured the government was to painstakingly point out that the things that were observed had already been capabilities that were in the wild with other models for some time.

05
Data

After the relaunch, Fable's fallback rate to Opus was actually lower than before because Anthropic loosened its safety classifiers, including for production-database access.

Pash reports that in his own testing, Fable's fallback rate dropped post-relaunch because Anthropic loosened classifiers, including allowing production-database work that previously triggered restrictions.

transcript

Pash: So in my testing the fallback rates were lower than before. And it it primarily because uh I was dealing with the production database etc. And the fable before uh the moment you mention the word production that's it's out uh and now the fable now is happy to uh address production databases etc.

06
Mechanism

Anthropic loosened its safety classifiers and added an appeal process for legitimate use cases like cybersecurity research after Fable's relaunch, rather than making the model itself less capable.

Pash reports that in his own testing, Fable's fallback rate to Opus actually dropped after the relaunch, suggesting Anthropic refined its safety classifiers and added an appeals process rather than nerfing the model.

transcript

Pash: I think actually what ended up happening is that they had a sit down and they refined they they actually got their classifiers uh looser for a lot of things and I've also seen online people say that they have uh when they get dropped out for cyber security research they can appeal and uh the company does a quick like why do you need this access and then they provided access

07
Mechanism

Anthropic quickly loosened Fable's safety classifiers and added an appeal process for restricted capabilities within days of the model's relaunch.

Pash says his testing showed lower fallback rates than before, suggesting Anthropic quickly loosened Fable's safety classifiers and set up an appeal process for restricted capabilities like cybersecurity research.

transcript

Pash: I think actually what ended up happening is that they had a sit down and they refined they they actually got their classifiers uh looser for a lot of things and I've also seen online people say that they have uh when they get dropped out for cyber security research they can appeal and uh the company does a quick like why do you need this access and then they provided access

08
Claim

AI labs are relying on iterative deployment and defense-in-depth safeguards that have worked so far, but nobody knows if or when these paradigms will break, and they could fail at especially critical moments.

Nathan warns that the iterative-deployment and defense-in-depth approach that has worked for AI labs so far is untested at higher capability levels, and could fail precisely when the stakes are highest.

transcript

Nathan: But you do wonder I mean all these models kind of start to come under some strain as you get to sufficiently powerful capabilities or or regimes where a small enough gap in the defenses is enough to create a huge problem. And yeah, it feels like in in multiple ways. We're kind of riding these paradigms that have worked well so far and we just don't know if and when they might break and if they do break it might they might be breaking at kind of critical times which is a strange juaposition on on multiple different levels.

09
Claim

AI models could hit sudden strain right at the moment capabilities become powerful enough that a small gap in defenses creates a huge problem, meaning the iterative-deployment paradigm might fail at the worst possible time.

Nathan warns that the iterative-deployment approach labs rely on has worked so far, but no one knows if or when it will break, and it could break precisely during high-stakes moments.

transcript

Nathan: all these models kind of start to come under some strain as you get to sufficiently powerful capabilities or or regimes where a small enough gap in the defenses is enough to create a huge problem. And yeah, it feels like in in multiple ways. We're kind of riding these paradigms that have worked well so far and we just don't know if and when they might break and if they do break it might they might be breaking at kind of critical times which is a strange juaposition

supports · 1

10
Claim

AI safety paradigms like iterative deployment and defense-in-depth safeguards have worked so far, but nobody actually knows when they'll break — and if they do break, it could happen at exactly the highest-stakes moment.

Nathan argues that Anthropic's safety paradigms (iterative deployment, defense-in-depth classifiers) have worked so far but are essentially untested at the frontier, and any failure could hit at a critical moment.

transcript

Nathan: But you do wonder I mean all these models kind of start to come under some strain as you get to sufficiently powerful capabilities or or regimes where a small enough gap in the defenses is enough to create a huge problem. And yeah, it feels like in in multiple ways we're kind of riding these paradigms that have worked well so far and we just don't know if and when they might break.

explains mechanism · 1

11
Claim

Frontier AI deployment paradigms (defense-in-depth safeguards, iterative deployment) have worked well so far, but nobody knows if or when they will break, and they could break at the worst possible moments.

Nathan argues that while iterative deployment and layered safeguards have worked so far for both OpenAI and Anthropic, there's no guarantee these paradigms won't fail right at a critical, high-stakes moment.

transcript

Nathan: We're kind of riding these paradigms that have worked well so far and we just don't know if and when they might break and if they do break it might they might be breaking at kind of critical times which is a strange juaposition on on multiple different levels.

12
Mechanism

Frontier labs can already route tasks internally, using a cheaper core model like Sonnet that pulls in a frontier model like Fable as an 'advisor' for harder tasks, and Anthropic's product team is actively promoting this pattern to API customers.

Pash explains that Anthropic already promotes an 'advisor' architecture where Sonnet/Haiku handle routine work and escalate to Fable only when needed, effectively enterprise-taught model routing.

transcript

Pash: So the Frontier Lab can always do this kind of advisor strategies where you have um you know sonnet 5 as the core model which then pulls in fable as an adviser for uh various tasks and in fact on the API that is actually what the product team at anthropic has been promoting.

13
Mechanism

Anthropic's own product team already tells enterprise customers to run cheap models like Sonnet and Haiku by default and only escalate to the flagship model when tasks get too hard, which is model routing driven by the customer.

Pash explains that Anthropic already promotes a routing pattern where Sonnet/Haiku handle classification and customer-service tasks while Fable is pulled in only for harder problems, meaning enterprise-driven model routing already exists.

transcript

Pash: They've been promoting to people if you're going to use Sonnet and Haiku use Sonnet and Haiku in your kind of classification systems for your customer service but when things get too difficult pull in a fable. And so this is basically model routing but model routing driven by the enterprise customer themselves.

14
Claim

Frontier AI labs already have the internal capability to do model routing themselves, so external multi-provider routing is mainly a short-term negotiating tactic on price rather than a durable long-term strategy.

Pash argues labs can route between their own models internally at will, so customer-side model routing is really just leverage in price negotiations, useful mainly because OpenAI and Anthropic remain close competitors.

transcript

Pash: So I don't I don't I don't really know that, you know, model routing is going to be a long-term thing, but it is a negotiating tactic against against the firms. Uh, negotiating against the price. So, it's it's very helpful right now that we have OpenAI and Anthropic kind of close to each other. It would be horrible if we just had anthropic.

supports · 1

15
Claim

Model routing across providers is not a durable long-term architecture but mainly a negotiating tactic enterprises and customers use to pressure frontier labs on price.

Pash argues that multi-model routing strategies are less a permanent architecture and more a leverage point customers use to negotiate better pricing from OpenAI and Anthropic.

transcript

Pash: I don't I don't I don't really know that, you know, model routing is going to be a long-term thing, but it is a negotiating tactic against against the firms. Uh, negotiating against the price.

supports · 1

16
Claim

Model routing between cheap and expensive models isn't really a long-term technical necessity right now — it functions mainly as a negotiating tactic against the frontier labs on price, and it only works this well because OpenAI and Anthropic are close competitors.

Pash argues that enterprise model-routing strategies are less a real technical requirement and more a negotiating lever against frontier labs, which only works because there's real competition between OpenAI and Anthropic.

transcript

Pash: So I don't I don't I don't really know that, you know, model routing is going to be a long-term thing, but it is a negotiating tactic against against the firms. Uh, negotiating against the price. So, it's it's very helpful right now that we have OpenAI and Anthropic kind of close to each other. It would be horrible if we just had anthropic.

supports · 1

17
Claim

AI disruption risk is concentrated almost entirely in 'paperwork' white-collar industries like banking, accounting, and compliance, since physical-product businesses like Nike have no exploitable IP for a model to absorb.

Pash argues that Alex Karp's warning about frontier AI 'stealing your business' only makes sense for paperwork-heavy industries (banking, compliance, accounting), not for physical businesses like Nike, since Anthropic has nothing to gain from entering shoemaking.

transcript

Pash: What does Nike have to fear from Anthropic? I mean, it's Anthropic is not going to go and make shoes. They're not going to like build a shoe brand. They don't have the, you know, expertise to go approach athletes to do sponsorships. like what what exactly is like anthropic threat threat to Nike or the threat to Nike's IP?

18
Claim

Frontier AI labs like Anthropic pose an existential threat mainly to 'paperwork'/IP-based businesses (software, banking, accounting, compliance) — not to physical, asset-based businesses like Nike or Walmart, since a model can't manufacture shoes or build a physical supply chain.

Pash pushes back on Alex Karp's warning that AI labs will 'steal your business,' arguing the real risk is concentrated in paperwork/IP-heavy industries, not physical-goods businesses like Nike.

transcript

Pash: What does Nike have to fear from Anthropic? I mean, it's Anthropic is not going to go and make shoes. They're not going to like build a shoe brand. They don't have the, you know, expertise to go approach athletes to do sponsorships. Like what exactly is anthropic threat to Nike or the threat to Nike's IP?

19
Claim

AI disruption risk from companies like Anthropic is concentrated in 'paperwork' IP-and-compliance businesses (software, banking, accounting, tax, regulatory), not physical-product businesses like Nike, Walmart, or mining companies.

Pash rebuts Alex Karp's warning that frontier AI labs will 'steal' enterprise IP, arguing the real threat is narrowly confined to white-collar paperwork industries — physical businesses like Nike have nothing to fear.

transcript

Pash: Like Walmart is not going to be disrupted by anthropic, right? You get all of the physical businesses out of the way and what you're left with is the pure IP businesses, right? Software, software production, um maybe pharma, maybe pharma, I'm not I'm not so sure. the paperwork businesses, banking, paperwork compliance businesses, um accounting, tax, uh compliance, regulatory, all of these things which are paperwork businesses, right?

20
Claim

Alex Karp's warning that frontier AI labs will 'steal' enterprises' IP mainly applies to paperwork/white-collar businesses (software, banking, compliance), not physical-product businesses like Nike or Walmart, which have little to fear from Anthropic.

Pash argues Karp's fear-based pitch about AI labs stealing enterprise IP only really threatens the paperwork-heavy, white-collar sector that has grown for decades, not physical businesses like Nike, Walmart, or mining companies.

transcript

Pash: he's really creating this aspect of fear and risk, but it's only going to affect a a a a portion of the economy. And unfortunately, that portion is the portion that has been growing dramatically in the last like 30 or 40 years. The paperwork, white collar, um, you know, type of professions have grown dramatically. The physical, the physical stuff has not grown a lot.

21
Claim

The AI disruption threat to businesses is concentrated in paperwork- and IP-based white-collar industries like software, banking, accounting, and compliance, not in physical-product businesses like Nike or Walmart.

Pash argues Alex Karp's warning that AI labs will 'steal your IP' only really applies to paperwork-heavy, white-collar businesses (banking, accounting, compliance, software) since physical businesses like Nike or Walmart have nothing an AI model can absorb.

transcript

Pash: Those are all of the businesses where, uh you have IP or relationships built up over years where if you have an anthropic go in and they read through your entire workflow and processing, they can basically absorb all of that into the model. So that is where I think the risk is.

22
Anecdote

Enterprise teams on the ground, outside the Silicon Valley bubble, are already seeing measurable AI ROI (falling exception rates, rising handling rates) even though it hasn't yet been recognized at the top of their organizations.

Pash recounts meeting non-Silicon-Valley implementers at the AI Engineer World's Fair, like a Midwest logistics CTO, who brought AI in-house after external vendors proved too slow, and are now seeing day-to-day drops in customer-service exception rates.

transcript

Pash: they have tried working with external vendors before and they have found it difficult because external vendors have not delivered as fast as they want them to deliver which you can imagine sitting out in the Midwest if you have a local external vendor and you outsource you know like a project to them and they have no idea what's going on they're like you know a year behind the frontier so his team internalize everything

23
Anecdote

Enterprise practitioners actually implementing AI on the front lines, such as in logistics and back-office accounting, are already seeing measurable ROI through falling exception rates and faster-handled customer service calls, even though this hasn't yet shown up in the aggregate numbers top executives look at.

Pash recounts conversations with non-Silicon-Valley implementers at the AI Engineer World's Fair — a logistics CTO and an accounting firm — who described day-to-day ROI from AI (falling exception rates, faster resolution) that hasn't yet filtered up to top-level enterprise metrics.

transcript

Pash: they see that customer service calls or exceptions that used to happen are now getting handled immediately and then all of the more difficult stuff that they used to have to like you know jump on they can now start to address and so they are seeing I think the return on investment on a day-to-day basis

rebuts · 2supports · 2

24
Anecdote

Frontline enterprise implementers deploying AI are already seeing clear day-to-day ROI — falling exception rates and faster handling of customer issues — even though this hasn't yet registered with executives who are far removed from the front line.

From conversations at the AI Engineer Worldfair, Pash says implementers outside Silicon Valley (logistics, accounting) are already seeing concrete ROI in falling exception rates, even though this hasn't filtered up to top executives.

transcript

Pash: they are very close to that edge they see that customer service calls or exceptions that used to happen are now getting handled immediately and then all of the more difficult stuff that they used to have to like you know jump on they can now start to address and so they are seeing I think the return on investment on a day-to-day basis

rebuts · 1

25
Anecdote

Frontline implementers of AI at non-tech companies are already seeing clear day-to-day ROI (falling exception rates, rising handling rates), even though this hasn't yet filtered up to the numbers that top enterprise executives see.

Pash recounts meeting non-Silicon-Valley operators at the AI Engineer World's Fair (a logistics CTO, an accounting firm) who are seeing immediate, measurable ROI from AI in daily operations, even though this hasn't yet shown up in top-level enterprise metrics.

transcript

Pash: they are actually in the day-to-day process of implementing and as they implement they automatically they see the results because they they are very close to that edge they see that customer service calls or exceptions that used to happen are now getting handled immediately

rebuts · 1supports · 2

26
Anecdote

Frontline implementers at non-Silicon-Valley companies are already seeing measurable day-to-day ROI from AI (falling exception rates, rising handling rates), but this hasn't filtered up to enterprise C-suites who only see rising token spend without visibility into the returns.

Pash recounts talking to non-tech implementers (logistics, accounting) at the AI Engineer World's Fair who see AI ROI daily on the ground, while big-company CTOs, removed from the front line, only see rising token spend.

transcript

Pash: they are seeing I think the return on investment on a day-to-day basis and this is very different from I think the story that you get from like the big enterprise c CTO's because they're so far away from like the the the front line that they don't actually know what's going on very closely. They're just looking at the numbers and by the numbers the token spend is going up but you know are you really seeing the return on investment.

27
Anecdote

Frontline teams actually implementing AI day-to-day are already seeing measurable ROI — falling exception rates and faster handling of customer calls — even though senior enterprise leadership hasn't recognized it yet because all they see is rising token spend.

From conversations at the AI Engineer World's Fair, Pash says implementers close to the ground already see AI paying off daily, but that signal hasn't reached top-level enterprise leadership, who only see token costs rising.

transcript

Pash: So they are seeing I think the return on investment on a day-to-day basis and this is very different from I think the story that you get from like the big enterprise CTO's because they're so far away from the front line that they don't actually know what's going on very closely. They're just looking at the numbers and by the numbers the token spend is going up but you know are you really seeing the return on investment.

rebuts · 1supports · 1

28
Anecdote

Frontline AI implementers at ordinary companies (logistics, accounting) are already seeing clear day-to-day ROI, even though this hasn't yet shown up in the aggregate numbers that top executives rely on.

Pash recounts meeting non-Silicon-Valley operators at the AI Engineer World's Fair — a Midwest logistics CTO and an accounting firm — who are seeing falling exception rates and rising handling rates from internally-built AI systems, even as C-suite leaders further from the front line remain skeptical about ROI.

transcript

Pash: the exception rate is falling the handling rate is going up and and they're seeing that on a day-to-day basis right so this is I think where we are uh the guys who are actually implementing and close to the implementations are actually seeing the results it hasn't really filtered up into the, you know, top layer of the enterprises yet, but the CEOs who are who are AI pilled kind of know what's going to happen.

gives example · 1

29
Claim

Enterprises are generally better off just paying for broad frontier-model (Fable) access across their whole employee base and letting people build ad hoc workflows, rather than investing time and committees into structuring cheaper open-source models to save money.

Nathan argues enterprises trying to economize by building workflows around cheaper open-source models are signing up to move slowly, versus just giving employees frontier-model tokens directly.

transcript

Nathan: It just feels to me like there's still a lot of advantage in just throwing some high value tokens out to your whole employee base and basically saying develop your own workflow solutions on a on a kind of as needed basis. And by the way, the frontier model which you have is really good at that.

supports · 1

30
Claim

Enterprises are better off paying for ad hoc frontier model access than rushing to build cheaper structured workflows, because cheaper models have only recently become good enough to deliver comparable results, and only after a lot of work.

Nathan argues that despite the appeal of economizing with cheaper models and structured workflows, frontier models like Fable still deliver more reliable value on an ad hoc basis right now.

transcript

Nathan: the frontier models still have the juice where if you're using them on an ad hoc basis, you're going to get value. You're going to know where that money went. And even if you're structuring things, which is a good goal to have, you probably don't want to rush into trying to structure them with a much cheaper model because it's only recently been possible with a lot of work to get comparable results.

rebuts · 1supports · 1

31
Claim

Enterprises should not rush to build cheaper structured workflows around lesser models to save money, because frontier models like Fable still deliver more value on an ad hoc basis and are themselves better at routing/delegating to cheaper models than most people are.

Nathan argues that despite cost pressure, enterprises are still better off just paying for frontier-model access rather than investing in workflows to route work to cheaper open-source models.

transcript

Nathan: I kind of still feel like all roads lead back to fable sort of uh if you're really going for performance. And if you're not going for performance or you're you're you're you know really prioritizing risk or you're really prioritizing IP then I think you do have another question on your hands.

rebuts · 1supports · 3

32
Prediction

Once a single AI model can hold the full context of an enterprise and allocate its own compute, the layers of standardized paperwork, permissions, and structured workflows that businesses built to manage risk and cost will become unnecessary.

Pash predicts that today's fragmented landscape of differently-sized models and rigid enterprise workflows will dissolve once a single model can hold an entire company's context and decide its own compute allocation, letting employees focus on judgment calls instead of paperwork.

transcript

Pash: So I think eventually all of this as you pointed out all of this stuff about like having different models addressed blah blah blah all of it kind of melts away because the model itself will allocate its own compute.

rebuts · 2

33
Prediction

Anthropic has already solved the hard 'zero-to-one' problem of a single AI maintaining shared multiplayer context across many people via Claude Tag; scaling that to more users is the comparatively easy 'one-to-n' problem.

Pash argues Anthropic has already cracked the hardest part of building a shared-context, enterprise-wide AI assistant (Claude Tag), and that scaling it further is the easy part — undercutting the case for elaborate model-routing setups.

transcript

Pash: They've solved they've solved the zero to one problem on multiplayer. Now the problem is the the the the one to n and the one to n as Peter the points out is actually easier than the zero to one. The zero to1 problem is the hard problem.

34
Prediction

Anthropic's Claude/Cloud Tag has already cracked the hard 'zero-to-one' problem of a shared-context multiplayer AI, which means SaaS software will progressively melt away as the model itself becomes the connective tissue of the enterprise.

Pash argues that Cloud Tag represents solving the genuinely hard 'zero-to-one' multiplayer AI problem (shared context, independent threads), and that scaling this from one-to-n is the easy part — implying SaaS businesses are about to be hollowed out.

transcript

Pash: Claude tag is the beginnings of that, right? Like I I've had this vision for for a while. Claude tag is the like they've solved multiplayer. They can they've solved like a able to handle independent threads, able to manage cont like large amounts of context. Now it's a question of just getting it better and better. They've solved they've solved the zero to one problem on multiplayer.

extends · 2gives example · 2rebuts · 1supports · 1

35
Prediction

Anthropic's 'Claude tag' has already solved the hard 'zero to one' problem of a multiplayer AI system with shared context and independent threads, making further scaling comparatively easy.

Pash argues Claude tag represents Anthropic having cracked the hard 'zero to one' problem of a shared-context, multiplayer AI system, meaning scaling it to more users ('one to n') is the easier remaining step.

transcript

Pash: They've solved they've solved the zero to one problem on multiplayer. Now the problem is the the the the one to n and the one to n as Peter the points out is actually easier than the zero to one. The zero to1 problem is the hard problem.

36
Prediction

Anthropic is coming for the entire SaaS industry because as the model itself becomes the connective tissue of the enterprise, the need for separate SaaS products melts away, a shift people like Alex Karp are struggling to grasp the speed of.

Pash argues that as a single model (via things like Claude's multiplayer 'Claude Tag') can hold an entire enterprise's context, the whole SaaS layer becomes unnecessary — a disruption he thinks is happening faster than people realize.

transcript

Pash: anthropic is coming for all of SAS right now because all of SAS melts away as the model as the model becomes the connective tissue of the enterprise all of SAS mel melts away and I think people are still having like Alex Karp are having their having trouble like wrapping their minds around what's going to and the pace at what's going to happen.

gives example · 2

37
Prediction

As Claude/Anthropic's multiplayer, shared-context product (Cloud Tag) matures, it will absorb the connective-tissue role currently played by SaaS software, causing most SaaS businesses to become obsolete.

Pash argues Anthropic has already solved the hard 'zero-to-one' problem of a shared-context, multiplayer AI (Cloud Tag), and that as this matures it will make most SaaS software redundant, which is why SaaS incumbents are alarmed.

transcript

Pash: anthropic is coming for all of SAS right now because all of SAS melts away as the model as the model becomes the connective tissue of the enterprise all of SAS mel melts away and I think people are still having like Alex Karp are having their having trouble like wrapping their minds around what's going to and the pace at what's going to happen.

gives example · 1

38
Claim

Alex Karp's anti-frontier-lab pitch is essentially the same enterprise lock-in sales strategy IBM used for decades with mainframes: sell an inferior, sheltered, owned-and-controlled solution to companies wary of dependency, and it can work for a long time even against superior alternatives.

Pash frames Palantir's anti-frontier-lab pitch as an IBM-style lock-in strategy that trades on enterprise fear of dependency rather than on cost or capability advantages, and notes it can succeed for decades regardless.

transcript

Pash: Alex Karp is like IBM now. He's basically trying to sell lock in to enterprise into these older inferior products and uh he and it and it works. I mean IBM has had a business for years, right? Like for 30 40 years they've been servicing mainframes while everyone has else has moved to cloud servers, right? So it it can work.

39
Claim

Alex Karp/Palantir is selling enterprise lock-in the same way IBM did with mainframes, pushing a 'trust and ownership' narrative because the cost argument alone can't compete with frontier AI labs.

Pash argues Alex Karp's warnings about frontier labs 'stealing your IP' are really an old-school enterprise lock-in sales pitch, comparing Palantir's positioning to IBM's decades-long mainframe servicing business.

transcript

Pash: Alex Karp is like IBM now. He's basically trying to sell lock in to enterprise into these older inferior products and uh he and it and it works. I mean IBM has had a business for years, right? Like for 30 40 years they've been servicing mainframes while everyone has else has moved to cloud servers, right?

40
Claim

Sam Altman's proposal to give roughly 5% of OpenAI equity to the public/households is a shrewd political move that pre-empts and defuses left-wing pressure to seize or heavily regulate the company's future value.

Pash describes Sam Altman's equity-sharing offer as savvy politics: by voluntarily giving away a stake, Altman removes the Democrats' rationale for forcibly taking more, while still capturing enormous future upside from ASI.

transcript

Pash: what Sam is doing is like look we can either have the Dems come into power and take it away from us forcibly or we can give it voluntarily and decide the terms of which you know the ownership is structured and that is really what he's aiming to do like the 5% is obviously a starting stake. If Bernie Sanders wanted to go to 20, I'm sure Sam would go to 20.

41
Mechanism

Inference is fundamentally a memory/data-movement problem rather than a compute problem, because once a model is trained you must move its weights and KV cache to the compute units, and GPUs were optimized for matrix-multiplication compute rather than this data-movement bottleneck.

Kunle Olukotun explains SambaNova's core architectural thesis: unlike training, inference is bottlenecked by moving weights and KV cache between memory and compute units, not by matrix-multiplication compute, which is why GPUs (optimized for training-era matmul) are inefficient at it.

transcript

Kunle Olukotun: the inference problem is not really a compute problem because as the models get bigger, you now need to move the weights and of course what we call the the KV cache, you know, into uh the compute units. And that is essentially a data movement problem, right? And it's a data movement problem, you know, from the memory to the uh compute units.

42
Mechanism

Running a trained AI model (inference) is fundamentally a data-movement/memory-bandwidth problem, not a compute/matrix-multiplication problem, because it requires continuously moving weights and the KV cache to compute units.

Kunle Olukotun explains that while training a model is compute-bound, running inference on a trained model is instead bottlenecked by moving weights and the KV cache between memory and compute units — the core problem SambaNova's dataflow chips are designed to solve.

transcript

Kunle Olukotun: once you've trained a model right and you train a model once you now need to use that model of course and that's the inference problem. And the inference problem is not really a compute problem because as the models get bigger, you now need to move the weights and of course what we call the the KV cache, you know, into uh the compute units. And that is essentially a data movement problem, right?

43
Mechanism

GPUs waste most of their memory bandwidth on inference because they synchronize data movement between kernels in software, whereas SambaNova's reconfigurable dataflow chips push memory-bandwidth utilization to 70-80% versus roughly 10-20% for GPUs.

Kunle Olukotun explains that LLM inference is fundamentally a memory-bandwidth problem rather than a compute problem, and that SambaNova's dataflow architecture keeps HBM bandwidth utilized 70-80% of the time by fusing decode into a hardware-orchestrated pipeline, versus GPUs typically running at only 10-20% utilization.

transcript

Kunle Olukotun: whereas uh GPU views are often running at you know maybe 10 to 20% of the uh capabilities of the resources right uh the bandwidth and and and and the uh the memory bandwidth and and and the communication resources our goal in a Samanova system is to push that to be 70 to 80% of of the peak.

44
Mechanism

Standard GPUs typically use only 10-20% of their available memory-bandwidth and communication capability during inference, whereas SambaNova's dataflow architecture is designed to push effective utilization to 70-80% of peak.

Kunle Olukotun explains that SambaNova's reconfigurable dataflow chip is built to keep memory bandwidth near-fully utilized during inference, versus the 10-20% utilization typical of GPUs.

transcript

Kunle Olukotun: whereas uh GPU views are often running at you know maybe 10 to 20% of the uh capabilities of the resources right uh the bandwidth and and and and the uh the memory bandwidth and and and the communication resources our goal in a Samanova system is to push that to be 70 to 80% of of the peak

extends · 1supports · 1

45
Mechanism

SambaNova's dataflow architecture pushes memory bandwidth utilization to 70-80% of peak by orchestrating data movement in hardware, compared to the roughly 10-20% GPUs typically achieve because they synchronize data movement in software.

Kunle Olukotun explains that SambaNova's core technical advantage is hardware-orchestrated data movement that keeps memory bandwidth utilization near 70-80%, versus GPUs' typical 10-20%, since inference is fundamentally a data-movement problem, not a compute problem.

transcript

Kunle Olukotun: whereas uh GPU views are often running at you know maybe 10 to 20% of the uh capabilities of the resources right uh the bandwidth and and and and the uh the memory bandwidth and and and the communication resources our goal in a Samanova system is to push that to be 70 to 80% of of the peak.

extends · 1supports · 1

46
Mechanism

The real bottleneck in AI inference isn't raw compute (FLOPs) but memory bandwidth — moving weights and the KV cache to the compute units — and winning that bottleneck is what lets you generate 'premium tokens': fast, accurate output from large models.

SambaNova co-founder Kunle Olukotun explains that inference is a memory-bandwidth problem, not a compute problem, and that solving it is what enables high-speed, high-value 'premium tokens' for agentic AI.

transcript

Kunle Olukotun: The speed of light of doing inference is really especially if you want high speed inference, right? Everybody knows that we're in the agentic era and so what one wants is what we call premium tokens. Premium tokens are tokens that you can charge the most money for because they are premium. And why are they premium? Because they come from very large models, so they're accurate, but they also are provided to you at high speed.

47
Mechanism

For high-speed agentic AI, the real bottleneck isn't compute (floating-point matrix multiplication) but memory bandwidth for moving model weights and KV cache, and dataflow chips are designed to maximize utilization of that bandwidth to deliver fast, expensive 'premium tokens.'

Kunle Olukotun explains that SambaNova's north star is delivering 'premium tokens' — fast, accurate outputs from large models — which requires maximizing memory bandwidth utilization rather than raw compute, since inference is fundamentally a data-movement problem.

transcript

Kunle Olukotun: everybody, you know, knows that we're in the aentic envir era and and so what one wants is what we call premium tokens, right? So, premium tokens are tokens that you can charge the most money for because they are premium. And why are they premium? because they are they come from from very large models.

48
Mechanism

SambaNova's RDU architecture fuses an entire model decode step into a single hardware-orchestrated kernel with a technique called kernel looping, keeping HBM bandwidth continuously utilized instead of the stop-start kernel-by-kernel data movement that limits GPUs.

Kunle Olukotun explains that SambaNova's RDU chips collapse the whole decoder into one continuously-looping kernel, avoiding the repeated HBM round-trips and kernel-launch overhead that waste bandwidth on GPUs.

transcript

Kunle Olukotun: the way that things work on a RDU in a data flow is essentially you take the decoder and you make that a single kernel, right? And then you go even further and you use a technique that we've developed called kernel looping.

Highlight slides
Current AI safety paradigms are working — for now✦ from: Frontier AI deployment paradigms (defense-in-depth safeguards, iterative deployment) have worked well so far, but nobody knows if or when they will break, and they could break at the worst possible moments.The risk: breaking at the worst moment✦ from: Frontier AI deployment paradigms (defense-in-depth safeguards, iterative deployment) have worked well so far, but nobody knows if or when they will break, and they could break at the worst possible moments.AI Threatens Paperwork, Not Physical Goods✦ from: Frontier AI labs like Anthropic pose an existential threat mainly to 'paperwork'/IP-based businesses (software, banking, accounting, compliance) — not to physical, asset-based businesses like Nike or Walmart, since a model can't manufacture shoes or build a physical supply chain.Why Nike Isn't Threatened by Anthropic✦ from: Frontier AI labs like Anthropic pose an existential threat mainly to 'paperwork'/IP-based businesses (software, banking, accounting, compliance) — not to physical, asset-based businesses like Nike or Walmart, since a model can't manufacture shoes or build a physical supply chain.Where AI Actually Threatens Businesses✦ from: The AI disruption threat to businesses is concentrated in paperwork- and IP-based white-collar industries like software, banking, accounting, and compliance, not in physical-product businesses like Nike or Walmart.Karp's AI Fear Pitch Has a Narrow Target✦ from: Alex Karp's warning that frontier AI labs will 'steal' enterprises' IP mainly applies to paperwork/white-collar businesses (software, banking, compliance), not physical-product businesses like Nike or Walmart, which have little to fear from Anthropic.Physical Businesses Are Different✦ from: The AI disruption threat to businesses is concentrated in paperwork- and IP-based white-collar industries like software, banking, accounting, and compliance, not in physical-product businesses like Nike or Walmart.Physical Businesses Largely Unaffected✦ from: Alex Karp's warning that frontier AI labs will 'steal' enterprises' IP mainly applies to paperwork/white-collar businesses (software, banking, compliance), not physical-product businesses like Nike or Walmart, which have little to fear from Anthropic.Frontline AI ROI Is Already Real✦ from: Frontline implementers of AI at non-tech companies are already seeing clear day-to-day ROI (falling exception rates, rising handling rates), even though this hasn't yet filtered up to the numbers that top enterprise executives see.Why the Gap Exists✦ from: Frontline implementers of AI at non-tech companies are already seeing clear day-to-day ROI (falling exception rates, rising handling rates), even though this hasn't yet filtered up to the numbers that top enterprise executives see.Gap Between Floor and C-Suite✦ from: Frontline implementers of AI at non-tech companies are already seeing clear day-to-day ROI (falling exception rates, rising handling rates), even though this hasn't yet filtered up to the numbers that top enterprise executives see.The Fragmentation Melts Away✦ from: Once a single AI model can hold the full context of an enterprise and allocate its own compute, the layers of standardized paperwork, permissions, and structured workflows that businesses built to manage risk and cost will become unnecessary.From Paperwork to Judgment✦ from: Once a single AI model can hold the full context of an enterprise and allocate its own compute, the layers of standardized paperwork, permissions, and structured workflows that businesses built to manage risk and cost will become unnecessary.What Melts Away✦ from: Once a single AI model can hold the full context of an enterprise and allocate its own compute, the layers of standardized paperwork, permissions, and structured workflows that businesses built to manage risk and cost will become unnecessary.Claude Tag Solved Multiplayer AI's Hard Part✦ from: Anthropic's Claude/Cloud Tag has already cracked the hard 'zero-to-one' problem of a shared-context multiplayer AI, which means SaaS software will progressively melt away as the model itself becomes the connective tissue of the enterprise.Claude Tag Solved the Hard Part✦ from: Anthropic has already solved the hard 'zero-to-one' problem of a single AI maintaining shared multiplayer context across many people via Claude Tag; scaling that to more users is the comparatively easy 'one-to-n' problem.Why It Threatens SaaS✦ from: Anthropic's Claude/Cloud Tag has already cracked the hard 'zero-to-one' problem of a shared-context multiplayer AI, which means SaaS software will progressively melt away as the model itself becomes the connective tissue of the enterprise.Scaling Is the Easy Part✦ from: Anthropic has already solved the hard 'zero-to-one' problem of a single AI maintaining shared multiplayer context across many people via Claude Tag; scaling that to more users is the comparatively easy 'one-to-n' problem.Anthropic Is Coming for All of SaaS✦ from: As Claude/Anthropic's multiplayer, shared-context product (Cloud Tag) matures, it will absorb the connective-tissue role currently played by SaaS software, causing most SaaS businesses to become obsolete.The Model Becomes the Enterprise's Connective Tissue✦ from: Anthropic is coming for the entire SaaS industry because as the model itself becomes the connective tissue of the enterprise, the need for separate SaaS products melts away, a shift people like Alex Karp are struggling to grasp the speed of.Incumbents Are Struggling to Grasp It✦ from: As Claude/Anthropic's multiplayer, shared-context product (Cloud Tag) matures, it will absorb the connective-tissue role currently played by SaaS software, causing most SaaS businesses to become obsolete.A Disruption Few Grasp the Speed Of✦ from: Anthropic is coming for the entire SaaS industry because as the model itself becomes the connective tissue of the enterprise, the need for separate SaaS products melts away, a shift people like Alex Karp are struggling to grasp the speed of.Karp's Playbook: The New IBM✦ from: Alex Karp/Palantir is selling enterprise lock-in the same way IBM did with mainframes, pushing a 'trust and ownership' narrative because the cost argument alone can't compete with frontier AI labs.The IBM Mainframe Parallel✦ from: Alex Karp/Palantir is selling enterprise lock-in the same way IBM did with mainframes, pushing a 'trust and ownership' narrative because the cost argument alone can't compete with frontier AI labs.Inference Is a Memory Problem, Not a Compute Problem✦ from: Inference is fundamentally a memory/data-movement problem rather than a compute problem, because once a model is trained you must move its weights and KV cache to the compute units, and GPUs were optimized for matrix-multiplication compute rather than this data-movement bottleneck.Why GPUs Struggle at Inference✦ from: Inference is fundamentally a memory/data-movement problem rather than a compute problem, because once a model is trained you must move its weights and KV cache to the compute units, and GPUs were optimized for matrix-multiplication compute rather than this data-movement bottleneck.Inference Is a Data-Movement Problem✦ from: Running a trained AI model (inference) is fundamentally a data-movement/memory-bandwidth problem, not a compute/matrix-multiplication problem, because it requires continuously moving weights and the KV cache to compute units.Why SambaNova Built Dataflow Chips✦ from: Running a trained AI model (inference) is fundamentally a data-movement/memory-bandwidth problem, not a compute/matrix-multiplication problem, because it requires continuously moving weights and the KV cache to compute units.Inference Is a Memory-Bandwidth Problem✦ from: GPUs waste most of their memory bandwidth on inference because they synchronize data movement between kernels in software, whereas SambaNova's reconfigurable dataflow chips push memory-bandwidth utilization to 70-80% versus roughly 10-20% for GPUs.GPUs Leave Most Bandwidth Unused✦ from: Standard GPUs typically use only 10-20% of their available memory-bandwidth and communication capability during inference, whereas SambaNova's dataflow architecture is designed to push effective utilization to 70-80% of peak.GPU vs. SambaNova: Bandwidth Utilization✦ from: GPUs waste most of their memory bandwidth on inference because they synchronize data movement between kernels in software, whereas SambaNova's reconfigurable dataflow chips push memory-bandwidth utilization to 70-80% versus roughly 10-20% for GPUs.SambaNova's Dataflow Fix✦ from: Standard GPUs typically use only 10-20% of their available memory-bandwidth and communication capability during inference, whereas SambaNova's dataflow architecture is designed to push effective utilization to 70-80% of peak.SambaNova's Memory Bandwidth Advantage✦ from: SambaNova's dataflow architecture pushes memory bandwidth utilization to 70-80% of peak by orchestrating data movement in hardware, compared to the roughly 10-20% GPUs typically achieve because they synchronize data movement in software.Why the Gap Exists✦ from: SambaNova's dataflow architecture pushes memory bandwidth utilization to 70-80% of peak by orchestrating data movement in hardware, compared to the roughly 10-20% GPUs typically achieve because they synchronize data movement in software.Inference's Real Bottleneck: Memory, Not Compute✦ from: The real bottleneck in AI inference isn't raw compute (FLOPs) but memory bandwidth — moving weights and the KV cache to the compute units — and winning that bottleneck is what lets you generate 'premium tokens': fast, accurate output from large models.What Makes a Token 'Premium'✦ from: The real bottleneck in AI inference isn't raw compute (FLOPs) but memory bandwidth — moving weights and the KV cache to the compute units — and winning that bottleneck is what lets you generate 'premium tokens': fast, accurate output from large models.The Real Bottleneck: Memory, Not Compute✦ from: For high-speed agentic AI, the real bottleneck isn't compute (floating-point matrix multiplication) but memory bandwidth for moving model weights and KV cache, and dataflow chips are designed to maximize utilization of that bandwidth to deliver fast, expensive 'premium tokens.'Chasing 'Premium Tokens'✦ from: For high-speed agentic AI, the real bottleneck isn't compute (floating-point matrix multiplication) but memory bandwidth for moving model weights and KV cache, and dataflow chips are designed to maximize utilization of that bandwidth to deliver fast, expensive 'premium tokens.'Dataflow Chips: Built for Bandwidth✦ from: For high-speed agentic AI, the real bottleneck isn't compute (floating-point matrix multiplication) but memory bandwidth for moving model weights and KV cache, and dataflow chips are designed to maximize utilization of that bandwidth to deliver fast, expensive 'premium tokens.'RDU: One Kernel for the Whole Decoder✦ from: SambaNova's RDU architecture fuses an entire model decode step into a single hardware-orchestrated kernel with a technique called kernel looping, keeping HBM bandwidth continuously utilized instead of the stop-start kernel-by-kernel data movement that limits GPUs.Why It Beats Stop-Start Execution✦ from: SambaNova's RDU architecture fuses an entire model decode step into a single hardware-orchestrated kernel with a technique called kernel looping, keeping HBM bandwidth continuously utilized instead of the stop-start kernel-by-kernel data movement that limits GPUs.Why It Matters✦ from: SambaNova's RDU architecture fuses an entire model decode step into a single hardware-orchestrated kernel with a technique called kernel looping, keeping HBM bandwidth continuously utilized instead of the stop-start kernel-by-kernel data movement that limits GPUs.
Related episodes