ByteByteGo
Alex Xu explains system design and scalable architecture, visually.
- How Databases Keep Their Sanity with Concurrency ControlSeptember 3, 2026
- Why Your RAG System Is Only as Good as Its Translator ModelSeptember 2, 2026
- How to Shrink a Language Model Without Making it Too DumbSeptember 1, 2026
- What Happens Inside an AI Chatbot Between Enter and the First Word?August 31, 2026
- Background Work: From Cron Jobs to Distributed SystemsAugust 27, 2026
- How to Make LLMs 3X FasterAugust 26, 2026
- How to Steal an AI Model’s Private ThoughtsAugust 25, 2026
- Why Code Verification Matters More Than Ever in the Age of AIAugust 24, 2026
- EP223: Ollama vs vLLM vs SGLangAugust 22, 2026
- Schema Evolution: Changing the Contract Without Breaking What RunsAugust 20, 2026
- GraphRAG: How AI Answers Questions Hidden Across Many DocumentsAugust 19, 2026
- The New American AI Model Designed to be CustomizedAugust 18, 2026
- Waymo vs Tesla: Two Ways to Build Self-Driving CarsAugust 17, 2026
- EP222: What is Google’s TPU?August 15, 2026
- A Detailed Guide to API Composition TechniquesAugust 13, 2026
- GitHub vs Vercel vs Replit: What Dev Platforms Do When AI Code Is CheapAugust 12, 2026
- How Cloudflare Is Making AI Pay for ContentAugust 11, 2026
- How to Fight Clickbait: Meta, LinkedIn & YouTube Case StudiesAugust 10, 2026
- The Read Path versus the Write Path: Strategies and TechniquesAugust 6, 2026
- How Big Models Teach Small Models to Be SmartAugust 5, 2026