AWS New Services August 2026: 5 Performance Wins
Optimize your AWS infrastructure with five proven performance tuning strategies. Real bottlenecks, specific fixes, measurable results.

On this page
TL;DR — Key takeaways
- CPU throttling and memory pressure are the top two bottlenecks in most AWS workloads; identify them with CloudWatch metrics before scaling vertically
- Network latency between availability zones adds 1-3ms per call; colocate tightly coupled services in the same AZ and use VPC endpoints for AWS services
- Storage IOPS limits cause silent performance degradation; provision adequate baseline IOPS or switch to gp3 volumes with independent throughput tuning
- Cold start penalties in serverless functions can be reduced 60-80% by increasing memory allocation and using provisioned concurrency for critical paths
- Monitoring response time at the 99th percentile reveals real user impact that averages mask; set alarms on p99 latency, not mean
AWS infrastructure performance bottlenecks rarely announce themselves clearly. Your application slows down, users complain, and CloudWatch shows a wall of metrics that all look fine at first glance.
In the support tickets I handled over the past year, five patterns kept showing up: CPU throttling on burstable instances, cross-AZ latency eating into response times, storage IOPS limits that nobody provisioned for, Lambda cold starts killing API performance, and averages that hid the awful experience of the slowest 5% of requests. Every AWS account I've worked on hits at least two of these.
Identify CPU and Memory Pressure Before Scaling
The first place to look is compute resources. Not because it's always the problem, but because it's fast to check and cheap to fix.
Open CloudWatch and pull up the CPUUtilization and MemoryUtilization metrics for your EC2 instances or ECS tasks. Look at the maximum values over the past week, not the averages. If you see sustained periods above 80%, you have pressure. For T-series instances, check CPUCreditBalance too. When that metric hits zero, AWS throttles your CPU to the baseline rate, and your application turns into molasses.
Memory pressure is harder to spot because the OS will use swap before it fails. If your instances have swap configured, check the swap usage metric. Any consistent swap activity means you're out of physical RAM and performance is suffering. I've seen API response times drop from 200ms to 4 seconds just because the instance was swapping.
The fix: increase instance size or switch families. T3 instances are cheap but throttle. M6i or C6i instances cost more but deliver predictable performance. Before you resize, document your current CPU and memory usage so you can measure the improvement. After the change, expect response times to drop by 30-60% if CPU or memory was the bottleneck.
- Pull max CPU and memory metrics over 7 days, not averages
- Check CPUCreditBalance on T-series instances; zero means throttling
- Look for swap usage as a sign of memory exhaustion
- Switch to fixed-performance instance types (M6i, C6i) if workload needs predictability
- Measure response time before and after to confirm the bottleneck
Reduce Network Latency with AZ and VPC Tuning
Network latency shows up in two places: between your services and between your services and AWS APIs. Both are fixable.
Cross-AZ latency is 1-3ms per hop. That sounds small. Multiply it by 50 database calls per request and you've added 50-150ms to every page load. Use the AWS CLI or a simple ping test between instances in different zones to measure the actual latency in your environment. If your application server and database are in different AZs and latency matters, move them to the same zone. Yes, you lose a bit of availability. In practice, most availability problems come from software bugs, not entire AZ failures.
For AWS service calls—S3, DynamoDB, Secrets Manager—use VPC endpoints. Without them, traffic routes out to the internet gateway and back through AWS's public network. VPC endpoints keep everything on the backbone. I've seen S3 API latency drop from 30ms to under 5ms just by adding an endpoint. The cost is minimal and the setup takes five minutes.
- Measure cross-AZ latency with ping or application health checks
- Colocate application and database in the same AZ if latency > availability
- Deploy VPC endpoints for S3, DynamoDB, and other AWS services you call frequently
- Expect 5-10ms improvement on AWS service API calls with VPC endpoints
- Monitor network latency in CloudWatch with custom metrics if needed
Provision Storage IOPS to Match Your Workload
EBS volume performance is limited by two things: IOPS and throughput. Default gp2 volumes give you 3 IOPS per GB, with a minimum of 100 IOPS and a burst bucket for short spikes. That works fine until it doesn't.
Check the VolumeReadOps and VolumeWriteOps metrics in CloudWatch. If you're consistently hitting the IOPS limit—usually visible as maxed-out metrics combined with increased disk latency—your volume is the bottleneck. This shows up as slow database queries, long file writes, or application pauses that don't correlate with CPU or memory.
The fix depends on your volume type. For gp2, you can increase volume size to get more baseline IOPS, but that's wasteful if you don't need the storage. Switch to gp3 instead. gp3 lets you set IOPS and throughput independently. A 100 GB gp3 volume can deliver 16,000 IOPS if you pay for it. That's the same performance as a 5 TB gp2 volume.
Provision 20% above your current peak IOPS usage to leave headroom. Monitor the change for a week. If IOPS utilization stays below 80%, you're good. If it climbs back up, provision more.
- Check VolumeReadOps and VolumeWriteOps in CloudWatch for IOPS saturation
- Look for increased disk latency that coincides with IOPS limits
- Switch from gp2 to gp3 for independent IOPS and throughput tuning
- Provision 20% above peak observed IOPS to avoid future bottlenecks
- For databases, consider io2 volumes if you need sub-millisecond latency
Cut Lambda Cold Start Latency with Memory and Provisioned Concurrency
Cold starts kill Lambda performance. A function that runs in 50ms when warm can take 2 seconds when cold. That's a terrible user experience, especially for APIs.
First, understand what's causing the delay. Enable AWS X-Ray on your Lambda functions and look at the initialization trace. If most of the time is spent loading dependencies, you can't do much except optimize your imports. If the runtime is taking a long time to spin up, increasing memory helps. Lambda allocates CPU proportionally to memory. A 512 MB function gets twice the CPU of a 256 MB function and initializes faster.
I've measured this on production functions. Doubling memory from 512 MB to 1024 MB cut cold start time from 1.8 seconds to 700ms. Yes, it costs more per invocation, but the improvement in user-facing latency is worth it for anything customer-facing.
For critical paths, use provisioned concurrency. It keeps a pool of warm instances ready. You pay for the idle time, but cold starts disappear. Set provisioned concurrency to match your baseline traffic, not your peak. Let autoscaling handle the spikes.
- Enable X-Ray to measure cold start duration and identify bottlenecks
- Increase memory allocation to 1024 MB or higher for compute-bound initialization
- Expect 50-70% cold start reduction by doubling memory
- Use provisioned concurrency for user-facing APIs with strict latency requirements
- Set provisioned capacity to baseline traffic; let autoscaling handle peaks
Monitor P99 Latency, Not Averages
Averages lie. Your API can have a mean response time of 120ms and still deliver a miserable experience to 5% of users.
The problem is outliers. A few slow requests pull the average up a bit, but they destroy the experience for the people who hit them. If your p99 latency is 3 seconds, that means one in a hundred requests takes at least three seconds. At any meaningful scale, that's a lot of angry users.
CloudWatch supports percentile metrics. Add p95, p99, and p99.9 alarms for your critical endpoints. When p99 latency crosses your threshold—say, 500ms for an API—you know there's a performance problem even if the average still looks fine. In my experience, p99 issues come from resource contention, garbage collection pauses, or tail database queries that don't get indexed properly.
Once you have the alarms, dig into the slow requests. Enable detailed logging or tracing for requests above your p99 threshold. Find the common pattern—usually it's a specific query, a cold cache, or a downstream service that's having a bad day.
- Configure CloudWatch alarms for p95, p99, and p99.9 latency on all user-facing endpoints
- Set p99 thresholds based on user experience requirements, not system capacity
- Enable detailed logging or tracing for requests that exceed p99 latency
- Investigate common patterns in slow requests: queries, cache misses, downstream calls
- Track p99 improvements after each optimization to confirm impact
Testing Changes Safely and Measuring Results
Performance tuning fails when you change three things at once and can't tell which one helped. Make one change at a time. Measure it. Move on.
Before you touch production, test in staging with realistic load. Use tools like Apache Bench, wrk, or Locust to simulate your traffic patterns. If you don't have a staging environment that matches production, at least run the change on a canary instance or a single container and monitor it for an hour before rolling out.
Document your baseline metrics: CPU utilization, memory usage, network latency, storage IOPS, and response time at p50, p95, and p99. After you apply a change, wait at least 30 minutes for metrics to stabilize. Compare the new numbers to baseline. If performance improves by less than 10%, the change didn't address the real bottleneck. Revert and try something else.
Keep your previous configurations or launch templates. If something breaks, you need to roll back fast. CloudWatch alarms should catch problems automatically, but check the metrics yourself for the first few hours after a change.
- Change one variable at a time so you know what worked
- Test configuration changes in staging with production-like load before deploying
- Document baseline metrics for CPU, memory, network, storage, and response time
- Wait 30 minutes after a change for metrics to stabilize before evaluating impact
- Keep previous configurations available for fast rollback if needed
Quick troubleshooting checklist
- Review CloudWatch CPU and memory metrics for the past 7 days to identify throttling patterns
- Run network latency tests between services using ping or application-level health checks
- Check EBS volume IOPS utilization in CloudWatch; provision baseline IOPS if consistently above 80%
- Measure Lambda cold start duration and memory usage with X-Ray or custom CloudWatch logs
- Set up p95 and p99 latency alarms for user-facing endpoints
- Document baseline performance metrics before making changes
- Test configuration changes in a staging environment with production-like load
- Schedule a rollback window and keep previous launch configurations available
FAQ
How do I know if my EC2 instance is CPU throttled?
Check the CPUCreditBalance metric in CloudWatch for T-series instances. If the balance drops to zero and stays there, your instance is being throttled. For other instance types, sustained CPU utilization above 80% combined with increased application latency indicates you need more compute capacity. Switch to a larger instance type or move to a compute-optimized family if the workload is CPU-bound.
What causes high network latency between AWS services in the same region?
Cross-AZ traffic adds 1-3ms latency compared to same-AZ communication. If your application makes hundreds of calls between services, this adds up. Use VPC endpoints to keep traffic on the AWS backbone instead of routing through internet gateways. Place tightly coupled services like an application server and its database in the same availability zone when latency matters more than the marginal availability improvement from zone separation.
When should I upgrade from gp2 to gp3 EBS volumes?
Upgrade when you need consistent performance above the gp2 baseline or when storage costs matter. gp3 volumes let you provision IOPS and throughput independently from volume size, so you can tune performance without overprovisioning storage. For workloads with predictable I/O patterns, gp3 typically costs 20% less than gp2 for the same performance level. Migrate during a maintenance window using snapshots or live volume modification.
Related articles
- Hosting OperationsSelf-Hosted App Deployment Fails? Check DNS, SSL, Reverse Proxy, and Logs FirstTroubleshoot failed self-hosted app deployments by checking DNS, SSL, reverse proxy routing, container status, logs, and ports.
- Hosting OperationsSelf-Hosted PaaS on a VPS: What to Check Before Installing Coolify, Dokploy, or CapRoverA hosting support checklist for preparing a VPS before installing self-hosted PaaS tools like Coolify, Dokploy, or CapRover.
- Hosting OperationsLinux Server Security Lessons from the Arch Linux Malware Package IncidentPractical Linux server security checklist for VPS admins after package malware concerns, with safe checks, rollback steps, and support guidance.
- Hosting OperationsAWS Lightsail Hong Kong VPS Latency: Practical Hosting Guide for IndonesiaLearn how to test AWS Lightsail Hong Kong VPS latency, compare regions, migrate safely, and troubleshoot hosting performance.
- Hosting OperationsCloudflare Tomorrow Watchlist: A Practical Hosting Operations GuidePractical Cloudflare troubleshooting checklist for DNS, SSL, caching, WAF, origin health, safe testing, and rollback planning.
- Hosting OperationsNetwork Safety Checklist for AI Agent Skills in Hosting OperationsAudit AI agent skills safely with network checks, secret protection, sandbox testing, rollback steps, and hosting support troubleshooting guidance.