The Real Cost of Running Kubernetes on AWS: 12 Places EKS Teams Waste Money
Amazon Elastic Kubernetes Service (EKS) is incredibly powerful, but it is notoriously easy to misconfigure in ways that silently bleed money. When we conduct Architecture Audits for scaling startups, we routinely find companies wasting 30% to 50% of their AWS compute budget on easily fixable EKS anti-patterns.
Here are the 12 most common places we see engineering teams wasting money when running Kubernetes on AWS—and exactly how to fix them.
1. Grossly Over-Provisioned Resource Requests
The number one cause of EKS bloat is developers setting arbitrary CPU and memory `requests` in their deployment manifests. Because the Kubernetes scheduler reserves this capacity entirely, a cluster can be completely "full" from the scheduler's perspective while actual CPU utilization sits at 5%.
The Fix: Implement Vertical Pod Autoscaler (VPA) in recommendation mode to profile actual usage, or use tools like Robusta and Datadog to right-size requests to match the 95th percentile of historical usage.
2. Ignoring Karpenter for Node Provisioning
If you are still using Managed Node Groups and the legacy Kubernetes Cluster Autoscaler (CAS), you are likely overpaying. CAS scales by adding identical EC2 instances to a group, which often results in stranded capacity (e.g., you need 1 CPU, but CAS spins up a 16-CPU instance).
The Fix: Migrate to Karpenter. It bypasses Auto Scaling Groups to provision the exact instance size needed for pending pods, and continuously consolidates workloads to cheaper nodes.
3. Cross-AZ Data Transfer Leaks
AWS charges $0.01 per GB for traffic crossing Availability Zones. In a multi-AZ EKS cluster, microservices communicating randomly across AZs can easily generate thousands of dollars in hidden data transfer fees.
The Fix: Enable Kubernetes Topology Aware Routing (TAR) to keep pod-to-pod traffic within the same AZ whenever possible.
4. Unused EBS Volumes (The "Orphaned PVC" Problem)
When a developer deletes a StatefulSet or a Helm chart fails to uninstall cleanly, the underlying AWS EBS volumes often remain provisioned. You continue paying for these unattached disks forever.
The Fix: Use AWS Config rules or automated Lambda scripts (like AWS Compute Optimizer) to find and delete unattached EBS volumes older than 7 days.
5. NAT Gateway Extortion
If your EKS worker nodes are in private subnets, pulling container images from public registries (like Docker Hub or GitHub Container Registry) routes traffic through the AWS NAT Gateway. At $0.045 per GB, scaling up hundreds of large container images will cause your NAT Gateway bill to explode.
The Fix: Use VPC Endpoints for ECR, S3, and DynamoDB. For third-party images, mirror them to a private ECR repository within your VPC.
6. Wasting Money on Idle GPUs
AI/ML workloads require expensive GPU instances (e.g., P4 or G5). But if a model is only processing inferences for 2 hours a day, keeping a $10/hour GPU node running 24/7 is a massive waste.
The Fix: Use Karpenter to aggressively scale GPU nodes to zero when idle, and investigate Multi-Instance GPUs (MIG) or NVIDIA Time-Slicing to share a single GPU across multiple pods.
7. Lack of Spot Instance Strategy
Spot instances offer up to 90% discounts on EC2, but many teams are afraid to use them due to the risk of interruption.
The Fix: Identify stateless, fault-tolerant workloads (like background workers, CI/CD runners, and web APIs) and route them to Spot nodes using node selectors and tolerations. Keep stateful databases on On-Demand nodes.
8. The "Weekend Wasteland" (Non-Prod Environments)
Development, staging, and QA clusters are rarely used at 2 AM on a Sunday, yet most teams leave them running 24/7.
The Fix: Implement a cronjob using tools like Kube-Downscaler to scale deployments in non-production namespaces to 0 outside of business hours.
9. Hoarding CloudWatch Logs
By default, EKS control plane logging and Container Insights can generate massive volumes of CloudWatch data. At $0.50 per GB ingested, observability can quickly become the most expensive part of your infrastructure.
The Fix: Only enable the specific control plane logs you need for compliance (e.g., audit/authenticator). For application logs, bypass CloudWatch entirely and ship directly to an external, cheaper store using FluentBit or OpenTelemetry.
10. Too Many Load Balancers
Creating a `LoadBalancer` type service in Kubernetes provisions a dedicated AWS Classic or Network Load Balancer. If you have 50 microservices, you are paying the hourly baseline cost for 50 ALBs.
The Fix: Use an Ingress Controller (like NGINX or AWS ALB Ingress Controller) to route traffic through a single, shared Load Balancer based on hostnames or URL paths.
11. Suboptimal CPU Architectures (Ignoring Graviton)
AWS Graviton (ARM-based) instances offer up to 40% better price-performance than standard x86 instances, but teams delay migrating because it requires multi-architecture Docker builds.
The Fix: Update your CI/CD pipelines to build multi-arch images (`docker buildx`), then configure Karpenter to prioritize Graviton instance types for compatible workloads.
12. EKS Control Plane Sprawl
AWS charges $73/month per EKS cluster just for the control plane. We often see teams spinning up separate EKS clusters for every developer, every microservice, or every minor environment.
The Fix: Consolidate. Use Kubernetes namespaces, RBAC, and network policies to achieve multi-tenancy within a smaller number of larger clusters.
Stop Guessing, Start Optimizing
The difference between a heavily optimized EKS environment and a default setup is often thousands of dollars per month. Cost optimization isn't a one-time event; it's a continuous platform engineering discipline.
Want us to find your hidden AWS costs?
Velocity Stream helps growing engineering teams drastically reduce their AWS spend without sacrificing reliability.

