Chapter 17
Master reading list
Books — read in this order
- Designing Data-Intensive Applications — Martin Kleppmann. Still the single best book on these topics. Re-read every 2 years.
- Database Internals — Alex Petrov. Pairs with DDIA; goes deeper into storage engines and distributed protocols.
- Release It! 2nd ed — Michael Nygard. Failure-mode-first design. Bulkheads, circuit breakers, etc.
- Site Reliability Engineering + The SRE Workbook — Google. Free online. SLOs, error budgets, capacity. sre.google/books
- Streaming Systems — Tyler Akidau et al. For event-time, watermarks, exactly-once.
- Understanding Distributed Systems — Roberto Vitillo. Tighter than DDIA; modern.
Engineering blogs — subscribe
- AWS Builder's Library — highest signal-per-page on the internet.
- Stripe Engineering — idempotency, ledgers, online migrations.
- Discord Engineering — large-scale messaging.
- Cloudflare Blog — network and infra.
- Slack Engineering
- GitHub Engineering
- Figma Engineering
- Uber Engineering
- Honeycomb Blog — observability mindset.
- Notion Engineering
- Netflix Tech Blog
- High Scalability — older case studies; archive is gold.
Papers — read the originals
- Dynamo (2007) — Amazon's KV store. Consistent hashing in production.
- Bigtable (2006) — LSM-trees + Chubby.
- Spanner (2012) — global strong consistency via TrueTime.
- Raft (2014) — the consensus paper that's actually readable.
- Paxos Made Simple — Lamport's accessible Paxos.
- Cassandra (2009)
- A Critique of ANSI SQL Isolation Levels
- Time, Clocks, and the Ordering of Events (Lamport, 1978) — 8 pages, foundational.
Courses + lectures
- MIT 6.824 Distributed Systems — full lectures + labs. The course where you build Raft from scratch.
- CMU Database Group YouTube — Andy Pavlo's Intro and Advanced Database Systems are world-class.
- Hillel Wayne — formal methods, TLA+, clear-thinking essays.
People worth following
- Kyle Kingsbury (@aphyr) — Jepsen, distributed systems empiricist.
- Marc Brooker — AWS principal engineer; queues, retries, capacity.
- Cindy Sridharan (@copyconstruct) — observability, distributed tracing.
- Charity Majors (@mipsytipsy) — observability mindset; Honeycomb.
- Martin Kleppmann (@martinkl) — DDIA author; CRDTs, local-first.
- Hillel Wayne — formal methods, careful thinking.
- Colm MacCárthaigh — AWS principal; TLS, networking, fail-safe design.
- Andy Pavlo — database internals.
- Brendan Gregg — perf, observability, flamegraphs.
How to use this list
Don't read it all. Pick one topic from this guide where you feel weakest, read the corresponding chapter in DDIA, then read the 2-3 best blog posts on that topic (linked above), then write a 200-word summary in your own words. That sequence produces actual depth.
Reading without writing is shallow. Writing without reading is empty. Doing both, on one topic at a time, is how you go from "knows the words" to "thinks like a staff engineer."