ATRIUMsearch → argument graph
Blog

ByteByteGo Newsletter

System design explained, from Alex Xu.

34 episodes indexed · ~19.6/month · since 2026-07 · Visit original ↗

EP224: MCP vs RAG vs AI Agents

2026-09-05 · 0 moments

How Databases Keep Their Sanity with Concurrency Control

2026-09-03 · 0 moments

Why Your RAG System Is Only as Good as Its Translator Model

2026-09-02 · 0 moments

How to Shrink a Language Model Without Making it Too Dumb

2026-09-01 · 0 moments

What Happens Inside an AI Chatbot Between Enter and the First Word?

2026-08-31 · 0 moments

Background Work: From Cron Jobs to Distributed Systems

2026-08-27 · 0 moments

How to Make LLMs 3X Faster

2026-08-26 · 0 moments

How to Steal an AI Model’s Private Thoughts

2026-08-25 · 0 moments

Why Code Verification Matters More Than Ever in the Age of AI

2026-08-24 · 0 moments

EP223: Ollama vs vLLM vs SGLang

2026-08-22 · 0 moments

Schema Evolution: Changing the Contract Without Breaking What Runs

2026-08-20 · 0 moments

GraphRAG: How AI Answers Questions Hidden Across Many Documents

2026-08-19 · 6 moments

The New American AI Model Designed to be Customized

2026-08-18 · 6 moments

Waymo vs Tesla: Two Ways to Build Self-Driving Cars

2026-08-17 · 6 moments

EP222: What is Google’s TPU?

2026-08-15 · 3 moments

A Detailed Guide to API Composition Techniques

2026-08-13 · 3 moments

GitHub vs Vercel vs Replit: What Dev Platforms Do When AI Code Is Cheap

2026-08-12 · 6 moments

How Cloudflare Is Making AI Pay for Content

2026-08-11 · 6 moments

How to Fight Clickbait: Meta, LinkedIn & YouTube Case Studies

2026-08-10 · 6 moments

The Read Path versus the Write Path: Strategies and Techniques

2026-08-06 · 4 moments

How Big Models Teach Small Models to Be Smart

2026-08-05 · 6 moments

Why An LLM’s Memory Gets Expensive and How to Fix It

2026-08-04 · 6 moments

LLM Security Basics: The Full Threat Model

2026-08-03 · 6 moments

Hiring: Part Time Instructor, Write Production Grade Code with AI

2026-07-31 · 3 moments

A Detailed Guide to Idempotency, Delivery Semantics, and Deduplication

2026-07-30 · 2 moments

How ChatGPT Optimizes its Agent Loop: Harness, API, and Inference

2026-07-29 · 6 moments

Why DoorDash, Instacart, and Uber Eats Integrated LLMs Into Search Three Different Ways

2026-07-28 · 5 moments

How NVIDIA Builds Open Models for the Age of AI

2026-07-27 · 6 moments

A Beginner’s Guide to Clocks, Causality, and Ordering in Distributed Systems

2026-07-23 · 6 moments

Best Practices for Building AI Agents That Work in Production

2026-07-22 · 6 moments

Inside Roblox’s Bet on World Models

2026-07-21 · 6 moments

MCP vs A2A vs ACP: How AI Agents Actually Talk to Each Other

2026-07-18 · 6 moments

A Guide to Multi-Tenancy: Benefits and Challenges

2026-07-16 · 4 moments

AI Customer Support at Scale: The Travel Industry’s $Billion Bet

2026-07-15 · 18 moments