Your server is slowing down. Which numbers explain why?

CPU is only part of the picture. Connect latency, errors, memory, disk and queues when looking for a bottleneck.

An aisle in a server room lined with metal equipment racks
Photography: İsmail Enes Ayhan / Unsplash

Pages load slowly and your first glance goes to CPU. The graph shows reasonable usage. The server could be waiting for storage, a database, a free connection or a remote API. Start with the symptom the user experiences, then look for the constraint. One number cannot explain an entire service.

Start with latency, errors and traffic

Compare the slow period with normal operation. Did request volume rise, did the mix of operations change or did new errors appear? Separate application work from time waiting on dependencies. Supplement average latency with internal percentiles; a small group of very slow requests can disappear inside an average.

External checks reveal worsening response times. Internal metrics and logs help explain where the delay originates. Use a consistent time reference so both can be compared with deployments and background jobs.

Read CPU and memory in context

High CPU during batch processing may be expected. Low CPU while the website is slow may mean waiting on another resource. With memory, distinguish useful caching from pressure that causes swapping or process termination. Inspect container limits as well as the machine's total capacity.

Disk and queues reveal hidden waiting

Check free space, available inodes and storage latency. A full disk can block both logs and database writes. Transfer volume alone does not explain how long operations wait. For the database, add slow queries, locks and the number of occupied connections.

A growing queue often signals that the service accepts work faster than it can process it. Track the age of the oldest task as well as queue length. A short queue containing a very old item may reveal a stuck process or a processing failure.

Add capacity after identifying the constraint

More application instances may not help if they all wait on the same database. They can even increase connection pressure and make the situation worse. Identify the constraint, make a small intervention and compare before and after. Even a temporary fix should have a measurable result.

  • Latency, error rate and request volume.
  • CPU, memory and process or container limits.
  • Free space, inodes and storage latency.
  • Database locks, connections and queue age.

What to take away

Collect a small set of measurements you can connect to user impact. The most useful view is not the one with the most graphs, but the one that points you towards the right intervention.

Documentation and further reading

Mgr. Martin Hlavaj, MBA

Software Engineer

All articles

Hear about an outage early.

Add your website or API to UpBot and choose who receives the alert.

Start monitoring for free