Velocity Stream LogoVelocity Stream Logo
Back to Insights
Kubernetes

AWS EKS Production Readiness Checklist

40 things to validate before routing live traffic to your Amazon EKS cluster. We don't write marketing fluff—this is the exact engineering checklist we use when building platforms for SaaS companies.

eks-architecture-diagram.tf
    VPC (10.0.0.0/16)
    │
    ├── Public Subnets (NAT Gateway, ALB, Bastion)
    │
    └── Private Subnets (EKS Nodes, RDS, Elasticache)
        │
        ├── EKS Control Plane (AWS Managed)
        │   ├── API Server Endpoint (Private/Public)
        │   └── etcd (AWS Managed)
        │
        ├── Worker Nodes (Karpenter Provisioned)
        │   ├── Core DNS & VPC CNI
        │   ├── Application Pods
        │   └── OTel Collectors (DaemonSet)
        │
        └── Security & IAM
            └── IRSA (IAM Roles for Service Accounts)

Building a Kubernetes cluster is easy. Building an EKS cluster that survives Black Friday traffic, passes SOC2 compliance, and doesn't bankrupt your startup is hard.

At Velocity Stream, we architect and scale cloud platforms for growing engineering teams. Before we sign off on an EKS environment, we validate it against this strict 8-part checklist.


1. Cluster Architecture

Decision: Multi-AZ deployment with Karpenter for compute orchestration.

Why: Hardcoding ASGs (Auto Scaling Groups) leads to slow scaling and wasted resources. Karpenter provisions right-sized nodes in milliseconds.

Trade-offs: Karpenter requires passing IAM roles to nodes and managing a separate controller, adding slight initial complexity over managed node groups.

  • Control plane spans at least 3 Availability Zones.
  • API Server endpoint access is restricted (Private only, or Public with strict CIDR allowlists).
  • Karpenter is installed and configured with appropriate NodePools (not just default NodeGroups).
  • CoreDNS replicas are scaled according to cluster size (using cluster-proportional-autoscaler).
  • Amazon VPC CNI is updated to the latest stable version and configured to avoid IP exhaustion.

2. Networking

Failure Mode: A spike in inter-AZ traffic or NAT Gateway usage bankrupts the startup because VPC Endpoints were not configured.

  • Worker nodes are placed strictly in private subnets.
  • VPC Endpoints (Gateway/Interface) are configured for S3, ECR, and DynamoDB to bypass NAT Gateway costs.
  • Security Groups strictly limit ingress to the ALB/NLB.
  • Network Policies (Calico or VPC CNI) restrict namespace-to-namespace communication.
  • Subnet CIDR blocks are large enough to handle peak pod counts (VPC CNI assigns IPs per pod).

3. Security & Access

Implementation: Never attach IAM policies directly to worker nodes. Always use IRSA.

  • IAM Roles for Service Accounts (IRSA) or EKS Pod Identities are configured for all AWS API access.
  • No hardcoded AWS credentials exist in any ConfigMap or Secret.
  • Secrets are encrypted at rest using a customer-managed AWS KMS key.
  • Container images are scanned for vulnerabilities before deployment.
  • Role-Based Access Control (RBAC) follows the principle of least privilege (no default cluster-admin).

4. Reliability & Availability

  • PodDisruptionBudgets (PDBs) are defined for all critical microservices to prevent eviction downtime.
  • Liveness and Readiness probes are correctly configured (not just returning 200 OK blindly).
  • Horizontal Pod Autoscaler (HPA) is configured based on CPU/Memory or custom metrics.
  • Stateful workloads use appropriate StorageClasses with cross-AZ backup strategies.
  • TopologySpreadConstraints are applied to force pods across multiple AZs.

5. Deployment & GitOps

Decision: Use ArgoCD for declarative deployments.

Why: Applying YAML via kubectl from a CI pipeline causes configuration drift. GitOps ensures the cluster always matches the repository.

  • All cluster state is declared in a Git repository.
  • ArgoCD or Flux is responsible for syncing state (pull-based, not push-based).
  • Helm charts or Kustomize are used to template environments cleanly.
  • Rollback procedures are documented and tested.
  • No human has permission to run kubectl apply in production.

6. Observability

Cost Alert: Sending all pod logs to CloudWatch or Datadog will result in massive bills. Filter spans and logs at the collector level.

  • OpenTelemetry collectors are deployed as DaemonSets to gather traces.
  • Prometheus/Grafana (or Datadog) is capturing kube-state-metrics.
  • Application logs are structured (JSON) and centrally aggregated.
  • Alerting rules exist for Node CPU > 80%, Pod Restarts > 5, and API Server Latency.
  • PagerDuty (or similar) is hooked up to high-priority alerts.

7. Cost & FinOps

  • Spot instances are utilized for stateless/fault-tolerant workloads.
  • Resource requests and limits are strictly set on every single namespace (using LimitRanges).
  • Karpenter Consolidation is enabled to aggressively pack pods and terminate idle nodes.
  • Cost allocation tags are applied at the EKS cluster and NodePool level.
  • Kubecost (or similar FinOps tooling) is installed to track cost-per-namespace.

8. Operational Readiness

  • Runbooks exist for common failure modes (e.g., "Node stuck in NotReady").
  • Incident response escalation paths are documented.
  • Disaster recovery plan exists to recreate the entire cluster from Terraform in under 30 minutes.
  • Clear ownership is established for who manages cluster upgrades versus application deployments.
  • Load testing has been executed against the production architecture.

The Velocity Stream Standard

Before you put an EKS cluster into production, every box above should have an owner, a documented decision, and a way to verify it.

Need help achieving production readiness?

We specialize in auditing, architecting, and rescuing Amazon EKS clusters. Let us handle the infrastructure so you can focus on shipping product.

Explore Kubernetes Consulting
Chat with an Engineer