Performance monitoring
Load average, top and htop, free, vmstat, iostat, and finding CPU, memory, disk and network bottlenecks.
"The server is slow" is one of the most common tickets an engineer gets. Slowness almost always comes down to one of four resources: CPU, memory, disk I/O or network. In this lesson you'll learn a systematic, tool-by-tool way to find which one is the bottleneck, and which process is responsible.
The first 60 seconds#
When you log in to a slow machine, these commands give you the big picture quickly:
Let's look at what each one tells you.
CPU and load average#
The three numbers are the load average over the last 1, 5 and 15 minutes: the average number of tasks running or waiting to run (plus those stuck in uninterruptible I/O). Interpret it relative to the number of CPU cores (nproc):
Compare the three values to see the trend: 8.0, 4.0, 1.0 means load is rising right now, while 1.0, 4.0, 8.0 means it's calming down.
top
The %Cpu(s) line is the key:
A process can show more than 100% CPU: 182.4 means it's using almost two cores. Interactive keys: P sorts by CPU, M by memory, 1 shows per-core usage, c shows full command lines, k kills, q quits.
htop (sudo apt install htop / sudo dnf install htop) shows the same information with per-core bars, a tree view (F5) and easy filtering (F4). mpstat -P ALL 1 (sysstat) shows usage per core, which is useful for spotting a single-threaded program maxing out one core.
Memory#
Don't panic about a low free value. Linux deliberately uses spare RAM as page cache (buff/cache) to speed up disk access, and hands it back when programs need it. The number that matters is available: how much memory applications can still get without swapping.
Signs of real memory pressure:
availableclose to zero;- swap activity: non-zero
si/socolumns invmstat(swap used on its own is fine; constant swapping in and out is not); - the OOM killer: when memory runs out, the kernel kills a process (often your database). Check with
sudo dmesg -T | grep -i -E 'killed process|out of memory'orjournalctl -k | grep -i oom.
Find the memory hogs with ps aux --sort=-%mem | head or top (press M). RES (resident memory) is what a process actually uses; VIRT includes reserved-but-unused address space and is usually misleading.
vmstat: everything at a glance#
The first line is an average since boot; read the later ones.
r: runnable tasks. Consistently greater than the core count means CPU saturation.b: tasks blocked on I/O.si/so: swap in/out per second. Should be ~0.bi/bo: blocks read from / written to disk.us sy id wa st: as in top. In the last line,wajumped to 24 whilebispiked, which points to disk.
Disk I/O#
(Abridged: the real output has more columns.) Focus on:
%util: the share of time the device was busy. Near 100% on a single HDD means saturated. Fast NVMe/SSD drives handle parallel requests, so also check the latency.r_await/w_await: average milliseconds per read/write, including queueing. Single-digit ms is fine on SSD; tens to hundreds means trouble.aqu-sz: the average queue length.
To find which process is doing the I/O:
And don't forget the simplest disk problem of all: a full filesystem (df -h, df -i). See the disks lesson.
Network#
Watch for an interface close to its bandwidth limit, huge numbers of connections in TIME-WAIT or CLOSE-WAIT (an app not closing connections), and retransmissions (nstat -az | grep -i retrans).
Pressure stall information (PSI)#
Modern kernels report how much time tasks spend waiting for CPU, memory or I/O, which is a very direct "are we starved?" signal:
The avg10/avg60/avg300 values are percentages of time over 10 seconds, 1 minute and 5 minutes. some means at least one task was stalled; full means all non-idle tasks were. Sustained double-digit values are a real problem.
History: sar#
The tools above show now. For "what happened at 3 a.m.?", enable sysstat's collector, which records stats every 10 minutes:
For fleets of servers, teams use a monitoring stack such as Prometheus + node_exporter + Grafana, Netdata, or a hosted service, with alerts on CPU, memory, disk space and latency.
A troubleshooting walkthrough#
uptime: load is 9 on a 4-core box, so something's queueing.top: CPU is 30%us, 45%wa, so it's not compute; something is waiting on disk.iostat -xz 1:nvme0n1is at 99%%utilwith 80 msw_await.sudo iotop -o: a nightlymysqldumpand a log-compression job are both hammering the disk.- Fix: reschedule one of them, and run the backup with
nice -n 19 ionice -c3.
The pattern is always the same: load → which resource → which process → why.
Common mistakes#
- Panicking about "free" memory. Read "available".
- Reading load average without the core count. A load of 8 is idle for a 64-core server and terrible for a 2-core one.
- Only looking at CPU when the real problem is I/O wait, swapping, or a full disk.
- Trusting the first line of
vmstat/iostat: it's the average since boot. - Ignoring
st(steal) on cloud VMs. Burstable instances (t-series) get throttled when they run out of CPU credits. - No history. Install sysstat or monitoring before the incident.
What's next#
You can now watch a server's health. The final lesson brings everything together into a security and hardening checklist for any Linux server you run.
Check your understanding
Quick quiz
1.A 4-core server shows
load average: 3.80, 3.65, 3.70. What does that suggest?2.
free -hshows only 200 MB 'free' but 5 GB 'available'. Is the server out of memory?3.In
vmstat 1, a consistently highwacolumn points to what bottleneck?
Finished reading?
Mark this lesson complete to track your progress.