Monitor your server
Check CPU, memory, disk and logs with htop, df and journalctl, and set up a small uptime check that emails you when your site stops answering.
Most server trouble gives you some warning: a disk creeping towards full, memory running short, a service restarting over and over. A handful of commands show all of that, and a ten-line script can tell you when your site stops answering.
CPU and memory
htop is top with the rough edges filed off:
apt install -y htop
htopThe bars along the top are your CPU cores and memory. F6 sorts by a column (try PERCENT_MEM), and q quits.
For a quick look without the full screen:
uptime
free -huptime gives load averages for the last 1, 5 and 15 minutes. If they stay above your number of vCPUs, work is queueing. In free -h, the number that matters is available: what programs can still get.
If the kernel ran out of memory and killed something, it says so in the journal:
journalctl -k | grep -i 'out of memory'Disk
df -hWatch Use% on /. Past about 90%, things start breaking in odd ways. To find what's eating the space:
du -xh / --max-depth=2 2>/dev/null | sort -h | tail -n 20The usual suspects are logs, old Docker images (docker system df shows them), and backups someone forgot to delete. The system journal can be trimmed with journalctl --vacuum-size=200M.
Services and logs
List anything that has failed:
systemctl --failedRead a service's recent log, nginx in this case:
journalctl -u nginx -n 50 --no-pagerjournalctl -f follows everything live, which is handy while you test something. journalctl -b -p err shows only errors since the last boot.
A tiny uptime check
This has to run somewhere else. A monitor on the server that's down can't tell you it's down. A second server or an always-on machine at home is fine. Save this as /usr/local/bin/check-site:
#!/bin/sh
URL="https://example.com/"
if ! curl -fsS -o /dev/null --max-time 10 "$URL"; then
echo "$URL did not answer at $(date -u)" | mail -s "DOWN: $URL" [email protected]
fiMake it executable:
chmod 755 /usr/local/bin/check-siteAnd run it every five minutes with crontab -e:
*/5 * * * * /usr/local/bin/check-sitemail needs working mail on that machine, for example mailutils relaying through a mail service. If that's more setup than you want, swap the mail line for a curl to a chat webhook, or use a hosted uptime service. Any of them beats finding out from a customer.
What the numbers are telling you
Low available memory most of the day, or swap always full: find the process in htop, or move to a bigger plan. Disk over 80%: clean up now while it's not urgent. A service that keeps restarting: read its log from the first failure, not the latest one.
And if the load is high and you don't know why, look for processes you don't recognise. If you find some, What to do if your server is compromised is the next read.