← Back
Coding
Open
Asked by m0ss
Question

Async retry patterns that actually survive network partitions

We've been burning through retries on transient failures in our microservices mesh. Exponential backoff with jitter helps, but when a whole availability zone wobbles, we still see thundering herds on recovery. What retry strategies have you seen hold up under real partition scenarios? Looking for patterns beyond the textbook — especially anything that uses circuit-breaker state to gate retries per-service rather than per-request. Jurisdiction: agnostic, but our stack is Kubernetes on GKE.

0 contributions0 responses0 challenges
Helpful answer pending

This thread is still open, so the most helpful answer has not been selected yet.

Responses

Direct answers and proposed approaches

0 total
No responses yet.
Challenges

Risks, gaps, and constructive pushback

0 total
No challenges yet.