If you are running Amazon Elastic Kubernetes Service (EKS) in production, compute costs are likely your largest cloud expense. For years, the Kubernetes Cluster Autoscaler (CAS) has been the default standard for scaling nodes. However, AWS's open-source Karpenter has fundamentally changed the math.
In this engineering postmortem, we will detail how we migrated a Series-B SaaS platform handling 50,000 requests per minute from CAS to Karpenter, resulting in a 40% reduction in monthly compute spend and sub-minute scaling latency.
The Problem: The "Node Group" Straitjacket
The fundamental limitation of the traditional Kubernetes Cluster Autoscaler on AWS is its tight coupling with EC2 Auto Scaling Groups (ASGs).
When a pod becomes pending due to insufficient resources, CAS looks at the configured ASGs and increments the desired capacity. This architecture introduces several painful bottlenecks:
- Slow Provisioning: Because it relies on ASGs, it often takes 3-5 minutes for a node to become ready. During a traffic spike, this delay causes dropped requests and 502 Bad Gateway errors.
- Inefficient Bin-Packing: ASGs require uniform instance types. If you have a pod requesting 3 vCPUs and 8GB RAM, and your ASG is configured for
m5.xlarge(4 vCPUs, 16GB RAM), you will constantly waste resources. - Spot Instance Churn: Managing Spot instances via ASGs requires complex mixed-instance policies and multiple node groups, increasing operational overhead.
Engineering Reality
"We had 15 different node groups configured just to handle different instance sizes and spot availability zones. Updating AMIs across all of them took a full engineering day."
The Intervention: Enter Karpenter
Karpenter bypasses Auto Scaling Groups entirely. It watches for unschedulable pods, evaluates their exact resource requests (CPU, memory, GPU), and communicates directly with the AWS EC2 Fleet API to launch the perfectly sized instance in milliseconds.
1. Group-less Scaling
Instead of defining node groups, you define a NodePool. You give Karpenter a list of acceptable instance families (e.g., c5, m5, r5, c6i, m6i) and it automatically determines the cheapest instance type that satisfies the pending pods' requirements.
apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
name: default
spec:
template:
spec:
requirements:
- key: karpenter.k8s.aws/instance-family
operator: In
values: [c5, m5, r5, c6i, m6i, r6i]
- key: karpenter.sh/capacity-type
operator: In
values: ["spot", "on-demand"]2. Instant Consolidation (FinOps Magic)
Karpenter's killer feature is Consolidation. While CAS can scale down empty nodes, Karpenter actively monitors the cluster to see if pods can be moved to cheaper nodes or if multiple underutilized nodes can be consolidated into a single smaller node.
If a $1.00/hr instance is running at 30% utilization, Karpenter will safely drain the pods, launch a $0.30/hr instance, move the pods, and terminate the expensive instance—all automatically.
The Outcomes
After a two-week shadow deployment, we fully cut over to Karpenter. The results were immediate and measurable.
Driven by aggressive bin-packing and automatic fallback to cheaper Spot instances.
Down from 4.5 minutes. Nodes boot and join the cluster significantly faster.
Verdict
For any EKS cluster spending more than $2,000/month on compute, migrating to Karpenter is one of the highest ROI engineering tasks available. It requires a paradigm shift away from static node groups, but the resulting agility and cost savings are unparalleled.

