Regular checks make server faults easier to catch while there is still time to choose a fix. I keep a short routine covering exposed services, disk space, logs, updates, access and scheduled jobs. Each check below has a specific question to answer.
1. Know what is listening
sudo ss -tlnpAnything bound to 0.0.0.0 or [::] is reachable from outside unless a firewall says otherwise. Caches, databases and admin backends should usually listen on 127.0.0.1 only. If something surprises you here, find out why before you do anything else.
2. Find what is filling the disk
# Biggest folders on this filesystem, smallest last
sudo du -xh --max-depth=1 / 2>/dev/null | sort -h | tail -15
# Or browse interactively
sudo apt install ncdu && sudo ncdu -x /A full disk takes databases down with it. I check this whenever usage creeps past 80%, not when it hits 100%.
3. Look for space that has not really been freed
sudo lsof +L1 2>/dev/null | sort -k7 -n | tailDelete a big log file while a process still has it open and the space stays used until that process restarts. This command lists those ghosts.
4. Keep the journal on a leash
sudo journalctl --disk-usage
sudo journalctl --vacuum-time=14d
# Make it permanent
sudo mkdir -p /etc/systemd/journald.conf.d
printf '[Journal]\nSystemMaxUse=500M\n' | sudo tee /etc/systemd/journald.conf.d/size.conf
sudo systemctl restart systemd-journald5. Let security updates install themselves
sudo apt install unattended-upgrades
sudo dpkg-reconfigure -plow unattended-upgrades
# Some updates need a reboot to take effect
[ -f /var/run/reboot-required ] && cat /var/run/reboot-required.pkgs6. Ban the people knocking
Every public SSH port is tried thousands of times a day. Fail2ban watches the logs and blocks addresses that keep failing:
sudo apt install fail2ban
printf '[sshd]\nenabled = true\nmaxretry = 5\nbantime = 1h\n' | sudo tee /etc/fail2ban/jail.d/sshd.local
sudo systemctl restart fail2ban
sudo fail2ban-client status sshd7. SSH with keys, not passwords
sudo tee /etc/ssh/sshd_config.d/10-hardening.conf > /dev/null <<'EOF'
PasswordAuthentication no
KbdInteractiveAuthentication no
PermitRootLogin prohibit-password
EOF
sudo sshd -t && sudo systemctl reload sshBefore you reload, open a second terminal and prove you can log in with your key. Keep your current session open until you have. Locking yourself out of a cloud server is an afternoon you will not get back.
8. Use systemd timers for jobs that matter
Cron is fine, but a systemd timer gives you logs, a clear status and catch-up runs after downtime:
# /etc/systemd/system/nightly-backup.service
[Unit]
Description=Nightly backup
[Service]
Type=oneshot
ExecStart=/usr/local/bin/nightly-backup
# /etc/systemd/system/nightly-backup.timer
[Unit]
Description=Run the nightly backup
[Timer]
OnCalendar=*-*-* 02:30
Persistent=true
RandomizedDelaySec=10m
[Install]
WantedBy=timers.targetsudo systemctl daemon-reload
sudo systemctl enable --now nightly-backup.timer
systemctl list-timers
journalctl -u nightly-backup.service --since today9. Trim SSDs on a schedule
Mounting with discard trims on every delete, which adds latency to busy disks. A weekly trim does the same job quietly:
sudo systemctl enable --now fstrim.timer10. Put /etc under version control
sudo apt install etckeeper
cd /etc && sudo git log --oneline | headEvery package install and config change gets a commit. When something stops working, git diff tells you what changed and when. It has saved me more than once.