Velocity Stream LogoVelocity Stream Logo
Back to Insights
Architecture

We Reviewed 10 AWS Architectures. Here Are the Problems We Keep Finding.

After conducting dozens of AWS infrastructure audits for Series A and B startups, distinct patterns of technical debt emerge. This is an engineering report on the 10 most common architectural anti-patterns, the risks they introduce, and how we remediate them.

The best AWS architecture isn't the one with the most services. It's the one your team can operate confidently at 3 AM.

Unfortunately, as engineering teams prioritize shipping features over infrastructure hygiene, architectural compromises accumulate. What begins as a "quick fix" hardens into a single point of failure that blocks scale, inflates the AWS bill, and destroys reliability.

Here are the 10 architectural problems we consistently find during our infrastructure assessments, and precisely how we fix them.


#1 The Single-Account Monolith

The Anti-Pattern
Development, Staging, and Production workloads all running in a single AWS account, separated only by VPCs (or worse, just tags).
The Risk
A developer running a destructive Terraform apply against "Staging" accidentally deletes Production RDS instances due to overlapping IAM permissions.
Our Remediation
Implement AWS Organizations with strong Service Control Policies (SCPs). Isolate Production, Staging, and Management (SSO, DNS) into strictly separated AWS accounts.

#2 EKS with Static, Over-Provisioned Nodes

The Anti-Pattern
Kubernetes clusters running on static AWS Auto Scaling Groups (ASGs) using uniform instance types (e.g., all m5.2xlarge), resulting in 20% utilization.
The Cost Impact
Burning thousands of dollars monthly on idle CPU cycles because ASGs cannot bin-pack efficiently.
Our Remediation
Deprecate ASGs and deploy Karpenter. Enable automatic consolidation to continuously repack pods onto cheaper, right-sized Spot instances.

#3 Private Workloads with Massive NAT Traffic

The Anti-Pattern
High-throughput services (like data processors or ML models) running in private subnets communicating with S3 or DynamoDB through a NAT Gateway.
The Cost Impact
Paying $0.045 per GB for internal AWS traffic. We frequently see NAT Gateway bills exceeding $5,000/month for no logical reason.
Our Remediation
Deploy VPC Gateway Endpoints for S3 and DynamoDB, and VPC Interface Endpoints (PrivateLink) for ECR and CloudWatch to keep traffic on the AWS backbone.

#4 "IaC Exists" but Production is ClickOps

The Anti-Pattern
Terraform was used to provision the baseline infrastructure, but engineers bypass it to manually tweak Security Groups or scaling policies in the AWS Console.
The Risk
Severe configuration drift. The next Terraform apply will either fail entirely or silently overwrite the manual production fixes, causing an outage.
Our Remediation
Revoke write access in the AWS Console. Implement Atlantis or GitHub Actions to require IaC for all changes. Run daily drift detection pipelines.

#5 Blind Kubernetes

The Anti-Pattern
Deploying microservices to EKS without centralized tracing or APM, relying solely on `kubectl logs` for debugging.
The Risk
Unexplainable 502 Bad Gateways. When a cross-service request fails, tracing the failure through the mesh is impossible, leading to multi-hour MTTRs.
Our Remediation
Deploy OpenTelemetry collectors as DaemonSets. Instrument applications to emit traces and centralize metrics in Prometheus/Grafana or Datadog.

#6 One-Way Deployments

The Anti-Pattern
CI/CD pipelines that execute `kubectl apply` or AWS ECS updates perfectly, but lack any automated mechanism to revert a bad deployment.
The Risk
Fear-driven development. Engineers stop deploying on Fridays because a bad release requires manually searching for the previous stable commit and pushing a hotfix.
Our Remediation
Implement ArgoCD for GitOps. Decouple CI (building the image) from CD (syncing the cluster state). Enable automated rollbacks based on Prometheus health checks.

#7 The Single-AZ Database

The Anti-Pattern
Running production PostgreSQL on a Single-AZ RDS instance "to save money," while the frontend application is distributed across 3 Availability Zones.
The Risk
When the underlying EC2 host for the database undergoes maintenance or fails, the entire application goes down hard for 15+ minutes.
Our Remediation
Enable RDS Multi-AZ. The standby instance ensures automatic failover in ~60 seconds. High availability is non-negotiable for stateful workloads.

#8 CloudWatch Eating the Observability Budget

The Anti-Pattern
Pumping 100% of application debug logs, VPC Flow Logs, and EKS audit logs into AWS CloudWatch Logs without lifecycle rules.
The Cost Impact
CloudWatch Logs ingestion costs $0.50 per GB. A chatty microservice can quickly generate $10,000+ monthly bills just in logging infrastructure.
Our Remediation
Filter logs at the collector level (FluentBit or OpenTelemetry). Route critical metrics to Prometheus, business logs to Elasticsearch/Loki, and archive audit logs directly to cold S3.

#9 Premature Microservices

The Anti-Pattern
A Series-A startup attempting to manage 40+ microservices communicating synchronously over gRPC, before finding product-market fit.
The Risk
Crippling operational complexity. Distributed monoliths result in cascading failures, impossible local development environments, and massive infrastructure overhead.
Our Remediation
Consolidate tightly coupled services back into a well-architected Modular Monolith, scaled via a simple ALB/ECS pattern until scaling pain dictates separation.

#10 "Highly Available" on Paper

The Anti-Pattern
The architecture diagrams show Multi-AZ RDS, autoscaling groups, and Route53 failover routing—but the team has never actually tested a failover event.
The Risk
When an AZ actually fails, you discover that your application caches were hardcoded to a single AZ IP, or your database takes 5 minutes to promote the replica, breaking SLAs.
Our Remediation
Game days. We deliberately kill nodes, trigger RDS failovers, and blackhole network traffic in staging (and eventually production) to validate resilience mechanisms.

The Conclusion

Production-ready doesn't mean the application is running. It means the engineering team knows exactly what happens when something goes wrong.

Are you scaling on technical debt?

We conduct deep-dive architecture reviews for companies scaling on AWS. We'll identify your single points of failure, cost anomalies, and security gaps—and deliver a prioritized remediation roadmap.

Explore Infrastructure Assessments
Chat with an Engineer