- Published:
- Last updated:
Distributed Systems Fundamentals
- Authors
- Name
- Lucian Oprea
- @LucianDSA_
00:30:00
Filter by difficulty
33 of 33 questions shown
Failure, Latency, and Overload
⏷ 1. What makes a distributed system fundamentally harder than a single process?
⏷ 2. How should you choose timeouts and propagate deadlines?
⏷ 3. When should a request be retried, and how do you prevent a retry storm?
⏷ 4. How do backpressure, load shedding, circuit breakers, and bulkheads differ?
⏷ 5. How do failure detectors and heartbeats work, and why can they be wrong?
Consistency and Availability
Free preview complete
You’ve reached the end of the free preview
Get every remaining question and complete answer, plus progress tracking across the full Interview Question Library.
- 28 more questions and complete answers in this topic
- Full access to every interview topic
- Progress tracking and question flags
- New questions and improvements during your subscription
Full access from
$12/month
No long-term commitment. Cancel whenever you want.
⏷ 6. How do linearizable, sequential, causal, and eventual consistency differ?
⏷ 7. What does the CAP theorem actually say?
⏷ 8. What does PACELC add to the CAP discussion?
⏷ 9. What session guarantees make an eventually consistent system easier to use?
Replication, Quorums, and Consensus
⏷ 10. How do leader-based, multi-leader, and leaderless replication compare?
⏷ 11. How do read and write quorums work, and what can still go wrong?
⏷ 12. What problem does consensus solve, and what does Raft provide?
⏷ 13. How do leader election, split brain, leases, and fencing tokens relate?
Clocks, Ordering, and Delivery Semantics
⏷ 14. Why are physical clocks dangerous in distributed systems?
⏷ 15. How do Lamport clocks, vector clocks, and sequence numbers help order events?
⏷ 16. Can a distributed system provide exactly-once processing?
Coordination and Cross-Service State
⏷ 17. When is a distributed lock appropriate, and why are fencing tokens important?
⏷ 18. How do two-phase commit, sagas, and the transactional outbox differ?
Core Distributed-System Vocabulary
⏷ 19. What is a network partition?
⏷ 20. What is a partial failure?
⏷ 21. Why does network latency change system design?
⏷ 22. What does idempotent mean in a distributed API?
⏷ 23. Why are stateless services easier to scale?
⏷ 24. How do availability and reliability differ in a distributed system?
⏷ 25. When should distributed work be synchronous versus asynchronous?
Repair, Placement, and Verification
⏷ 26. What are anti-entropy and read repair?
⏷ 27. How does consistent hashing help place data across nodes?
⏷ 28. How do you version and roll out a distributed protocol?
⏷ 29. What is a gossip protocol, and when is it useful?
Advanced Failure and Convergence
⏷ 30. How would you design regional failover for a stateful service?
⏷ 31. How does state-machine replication provide consistent results?
⏷ 32. What are CRDTs, and when are they appropriate?
⏷ 33. How should chaos testing be used safely in distributed systems?