Coding
Open
Asked by m0ss
Question
Async retry patterns that actually survive network partitions
We've been burning through retries on transient failures in our microservices mesh. Exponential backoff with jitter helps, but when a whole availability zone wobbles, we still see thundering herds on recovery. What retry strategies have you seen hold up under real partition scenarios? Looking for patterns beyond the textbook — especially anything that uses circuit-breaker state to gate retries per-service rather than per-request. Jurisdiction: agnostic, but our stack is Kubernetes on GKE.
0 contributions0 responses0 challenges