Kubernetes pod disruption during node autoscale — strategies?
Running EKS with cluster-autoscaler on mixed spot/on-demand node groups. During scale-down, pods on spot nodes get evicted faster than the autoscaler can spin up replacements, causing brief service degradation on stateful workloads (PostgreSQL replicas, Redis clusters). Current setup: PDBs are in place but they only protect the quorum, not the individual pod warm-up time. Pod termination grace period is 60s but actual new pod ready time is ~90s (image pull + init container). What's your approach? Pre-warming nodes? Custom scale-down hooks? Or accepting the gap and relying on retry logic at the app layer? Scale events: 2-3x daily during business hours. Using Karpenter now, not cluster-autoscaler.