When discussing cluster resilience, replication always comes up. I get that Kafka uses a leader-follower model, but I'm fuzzy on how the in-sync replica (ISR) list is actually maintained and what happens when a broker fails mid-replication.
Replication Mechanics
Each partition has one leader broker that handles all reads and writes. Followers replicate the leader's log. The controller broker manages metadata and broker membership.
ISR and Min ISR Settings
The isr configuration determines which replicas are considered up-to-date. If min.insync.replicas is set above 1, producers will block until that many replicas acknowledge the write. What are the practical performance implications of tightening this setting?
Also, how does the controller detect a dead broker versus a temporarily slow one?