ByteByteGo Newsletter
System design explained, from Alex Xu.
34 episodes indexed · ~19.6/month · since 2026-07 · Visit original ↗
EP224: MCP vs RAG vs AI Agents
2026-09-05 · 0 moments
How Databases Keep Their Sanity with Concurrency Control
2026-09-03 · 0 moments
Why Your RAG System Is Only as Good as Its Translator Model
2026-09-02 · 0 moments
How to Shrink a Language Model Without Making it Too Dumb
2026-09-01 · 0 moments
What Happens Inside an AI Chatbot Between Enter and the First Word?
2026-08-31 · 0 moments
Background Work: From Cron Jobs to Distributed Systems
2026-08-27 · 0 moments
How to Make LLMs 3X Faster
2026-08-26 · 0 moments
How to Steal an AI Model’s Private Thoughts
2026-08-25 · 0 moments
Why Code Verification Matters More Than Ever in the Age of AI
2026-08-24 · 0 moments
EP223: Ollama vs vLLM vs SGLang
2026-08-22 · 0 moments
Schema Evolution: Changing the Contract Without Breaking What Runs
2026-08-20 · 0 moments
GraphRAG: How AI Answers Questions Hidden Across Many Documents
2026-08-19 · 6 moments
The New American AI Model Designed to be Customized
2026-08-18 · 6 moments
Waymo vs Tesla: Two Ways to Build Self-Driving Cars
2026-08-17 · 6 moments
EP222: What is Google’s TPU?
2026-08-15 · 3 moments
A Detailed Guide to API Composition Techniques
2026-08-13 · 3 moments
GitHub vs Vercel vs Replit: What Dev Platforms Do When AI Code Is Cheap
2026-08-12 · 6 moments
How Cloudflare Is Making AI Pay for Content
2026-08-11 · 6 moments
How to Fight Clickbait: Meta, LinkedIn & YouTube Case Studies
2026-08-10 · 6 moments
The Read Path versus the Write Path: Strategies and Techniques
2026-08-06 · 4 moments
How Big Models Teach Small Models to Be Smart
2026-08-05 · 6 moments
Why An LLM’s Memory Gets Expensive and How to Fix It
2026-08-04 · 6 moments
LLM Security Basics: The Full Threat Model
2026-08-03 · 6 moments
Hiring: Part Time Instructor, Write Production Grade Code with AI
2026-07-31 · 3 moments
A Detailed Guide to Idempotency, Delivery Semantics, and Deduplication
2026-07-30 · 2 moments
How ChatGPT Optimizes its Agent Loop: Harness, API, and Inference
2026-07-29 · 6 moments
Why DoorDash, Instacart, and Uber Eats Integrated LLMs Into Search Three Different Ways
2026-07-28 · 5 moments
How NVIDIA Builds Open Models for the Age of AI
2026-07-27 · 6 moments
A Beginner’s Guide to Clocks, Causality, and Ordering in Distributed Systems
2026-07-23 · 6 moments
Best Practices for Building AI Agents That Work in Production
2026-07-22 · 6 moments
Inside Roblox’s Bet on World Models
2026-07-21 · 6 moments
MCP vs A2A vs ACP: How AI Agents Actually Talk to Each Other
2026-07-18 · 6 moments
A Guide to Multi-Tenancy: Benefits and Challenges
2026-07-16 · 4 moments
AI Customer Support at Scale: The Travel Industry’s $Billion Bet
2026-07-15 · 18 moments