Group mentoring · pilot cohort

Distributed systems that fail well.

For mid-level and senior backend developers who want to move beyond shipping features and start designing systems that stay useful through failures, spikes, and change.

In six weeks, you turn a functional but fragile architecture into a system with a failure model, resilience mechanisms, observability, and a recovery strategy.

This program is for you if you…

  • work in a monolith that is hard to test and evolve
  • need to decide between APIs, queues, databases, and cache
  • know the tools but want to handle failures and production better
  • want the practical judgment needed to operate as a senior engineer or tech lead

You will leave with

  • an architecture diagram and risk catalogue
  • a failure model and documented architectural decisions
  • implemented patterns in your context or study project
  • an observability and recovery plan

What we'll work on

  1. 01Failure models, timeouts, retries, backoff, and jitter
  2. 02Circuit breakers, bulkheads, and controlled degradation
  3. 03Events, idempotency, ordering, DLQs, and reprocessing
  4. 04Eventual consistency, outbox, contracts, and traceability
  5. 05Observability, tracing, and signals that matter during incidents
  6. 06Architecture Decision Records and a system-design presentation
Tell me about the next cohortView workshop →