Data◆Audio · 14:37 — 16:07
Reinforcement-learning rollout workloads are so extremely bursty that they can require spinning up on the order of a hundred thousand sandboxes at once, far more bursty than typical agent workloads.
Akshat notes that while ordinary agent sandbox usage isn't especially bursty, reinforcement-learning rollouts are an extreme exception, sometimes demanding roughly 100,000 sandboxes simultaneously. ✦ AI generated
Akshat Bubna · Latent Space · 2026-07-08 · original ↗
plays this moment only · 14:37 — 16:07
Elicited by
“When you look at workloads like that, you're doing train runs where you need to scale up, scale down every hour thousands of GPUs, right? That's the example for we do need it, right?”
I agree that agents are not that bursty. sandboxes are, except when you're doing RL. RL is just insanely bursty. Yeah. Like when you're doing, rollouts, you sometimes need a hundred thousand sandboxes in your sandboxes.
verbatim transcript · starts at 14:37
Transcript · around this moment
15:24DeFlash, speculative decoding, and Auto Endpoints
- ·Typical agent sandbox usage isn't very bursty
- ·RL rollouts are the extreme exception
- ·Rollouts can demand ~100,000 sandboxes at once
- ·Far more bursty than standard agent workloads
- ·RL training triggers simultaneous rollout spikes
- ·Sandbox infra must scale instantly to ~100k
- ·Standard agent workloads never hit this scale
Around this claim