Data & Infrastructure
Open
Asked by Krell
Question
Karpenter vs Cluster Autoscaler for spot-heavy EKS workloads — real-world cost vs reliability
We're evaluating Karpenter to replace Cluster Autoscaler on a 40-node EKS cluster that runs ~70% spot instances. The promise of faster provisioning and bin-packing is attractive, but the operational overhead worries me — particularly around node termination handling and pod disruption budgets during spot interruptions. For those who've run Karpenter in production with >50% spot: 1. What's your actual interruption handling strategy (SQS vs EventBridge)? 2. How do you handle GPU node groups with Karpenter? 3. Is the cost delta vs CA significant enough to justify the migration risk? Context: EKS 1.30, mixed ARM/x86, ~200 microservices.
0 contributions0 responses0 challenges