An EC2 instance can look comfortably sized in a daily CPU chart and still slow down during a busy period. I check the limits that apply to the actual workload before buying a larger instance or changing database settings.
The useful question is not simply “how much CPU have I got?” It is whether the server can sustain the CPU, storage and memory demand of its busiest normal tasks.
Use the time window where the problem happened
Record the slow request or job, its start and end times, and the instance and volume involved. Look at the smallest useful monitoring interval available. A brief burst can disappear inside a broad average.
Compare that period with a normal one. Was a backup, import, image rebuild or deployment running alongside visitor requests? A workload that is harmless by itself can compete with the database when the schedules overlap.
Keep a simple incident worksheet: observed symptom, affected resource, metric window, suspected constraint and the next check. This prevents a promising-looking graph becoming a conclusion before you have tested it.
Check CPU credits on burstable instances
For a burstable instance, inspect CPUCreditBalance and CPUCreditUsage alongside CPU utilisation. Credits describe the ability to operate above the instance’s baseline; a short test with a healthy balance is not proof of sustained capacity.
Also check whether the instance uses Standard or Unlimited mode. Exhausted earned credits do not imply identical behaviour in both modes. Unlimited can use surplus credits, with possible charges; its surplus balance and charged-credit metrics help explain that behaviour. AWS publishes the exact metric definitions.
I would not switch modes simply to make a warning disappear. First estimate the ongoing demand and cost, then compare a better-fitting instance with reducing avoidable work. A routine task that spends hours above baseline deserves a capacity review, not just a faster five-minute benchmark.
Separate volume limits from instance limits
EBS performance is constrained by both the attached volumes and the instance’s EBS capability. Increasing a volume’s provisioned performance cannot remove a lower limit on the instance connection.
Check IOPS and throughput separately, along with latency and queue behaviour. A workload reading many small records differs from one copying large files. Compare the workload with the relevant volume configuration and instance limits rather than assuming storage capacity in gigabytes describes its speed.
Where burst-balance metrics apply, inspect those too. Do not assume every volume type has the same credit mechanism. Use the AWS documentation for the volume and instance you actually run.
Do not forget the work inside the server
A cloud graph does not replace application evidence. Match it with Linux resource samples, slow requests and database activity. If memory pressure is the real problem, extra IOPS may only make the symptoms less painful while the underlying shortage remains.
Check the schedule before increasing capacity. Staggering a rebuild and a backup can be worth testing where they compete, but verify that both still complete within their required windows. Do not delay essential work indefinitely to make the graph quieter.
Likewise, a cache only helps the requests it can serve correctly. Confirm that the slow journey is cacheable and that the expected cache is actually being used.
Change the constraint you have demonstrated
Write down the expected result before resizing: for example, a particular batch should finish within its window without slowing the public site. Keep a rollback route and account for any interruption or temporary overlap cost.
After the change, repeat a representative workload and inspect the following busy period. Compare completion time, visitor response time and the relevant capacity metrics. Record the cost as well as the improvement.
The best outcome may be a different instance, a storage adjustment or less unnecessary work. The evidence should choose between them.
Further reading: EC2 CPU-credit metrics, instance and EBS performance limits, and EBS I/O monitoring.