Systems Troubleshooting, But Only the Useful Stuff
CONTENTS
A condensed reference for troubleshooting Linux hosts, services, processes, resources, containers, and common failures.
Not a Linux administration tutorial. Just the things that repeatedly matter when something is broken.
Mental Model
Start here:
symptom
→ scope
→ evidence
→ hypothesis
→ test
→ fix
→ verify
Useful questions:
What exactly is broken?
One user, one host, or everything?
What changed recently?
What does the evidence actually prove?
What do I still not know?
What is the cheapest test for the next hypothesis?
Do not change five things at once.
Evidence Has Limits
systemctl says active
→ process is alive
→ does not prove the app is healthy
CPU is idle
→ CPU is not the bottleneck right now
→ does not prove the machine is healthy
df shows free space
→ filesystem has free blocks
→ does not prove inodes, quotas, or another filesystem are fine
process exists
→ process has not exited
→ does not prove it is making progress
Keep facts and guesses separate.
First Look
uptime
free -h
df -h
df -i
systemctl --failed
ps aux --sort=-%cpu | head
ps aux --sort=-%mem | head
Live:
top
vmstat 1
Storage:
iostat -xz 1
The goal is not to diagnose everything here. It is to decide where to look next.
Processes
Find a process:
pgrep -af <name>
Inspect:
ps -fp <pid>
ps -o pid,ppid,user,stat,%cpu,%mem,rss,vsz,etime,cmd -p <pid>
Process tree:
pstree -ap
ps -ef --forest
Useful fields:
| Field | Meaning |
|---|---|
PID |
process ID |
PPID |
parent PID |
STAT |
process state |
RSS |
resident memory |
VSZ |
virtual address space |
ETIME |
elapsed runtime |
Do not confuse VSZ with actual RAM usage.
Process States
ps -eo pid,ppid,stat,wchan:32,comm
| State | Meaning |
|---|---|
R |
running / runnable |
S |
sleeping |
D |
uninterruptible sleep, often I/O |
T |
stopped |
Z |
zombie |
D state
Find blocked tasks:
ps -eo pid,stat,wchan:32,comm | awk '$2 ~ /^D/'
Think:
disk
NFS
network storage
filesystem
device I/O
High load + idle CPU + many D tasks usually points toward I/O.
Zombies
ps -eo pid,ppid,stat,comm | awk '$3 ~ /^Z/'
A zombie is already dead. Its parent has not reaped it yet.
Find the parent:
ps -fp <ppid>
CPU
top
ps aux --sort=-%cpu | head
pidstat 1
One process:
pidstat -p <pid> 1
top -H -p <pid>
Ask:
one process or many?
one core or all cores?
constant or intermittent?
user CPU or kernel CPU?
traffic spike?
retry loop?
bad deployment?
Useful top fields:
| Field | Meaning |
|---|---|
us |
user-space CPU |
sy |
kernel CPU |
id |
idle |
wa |
I/O wait |
st |
stolen CPU |
Load Average
uptime
Example:
load average: 12.4, 10.8, 8.1
Roughly 1, 5, and 15 minutes.
On Linux, load includes tasks that are runnable and tasks stuck in uninterruptible sleep.
So:
high load ≠ necessarily high CPU
High load with idle CPU:
vmstat 1
iostat -xz 1
ps -eo pid,stat,wchan:32,comm
In vmstat:
| Field | Meaning |
|---|---|
r |
runnable tasks |
b |
blocked tasks |
si |
swap in |
so |
swap out |
wa |
I/O wait |
Memory
free -h
The important number is usually:
available
not:
free
Linux uses spare RAM for cache. Low completely-unused memory is normal.
Find consumers:
ps aux --sort=-%mem | head
ps -eo pid,user,rss,%mem,etime,cmd --sort=-rss | head
Watch:
watch -n 1 free -h
vmstat 1
A memory leak is usually interesting because usage keeps growing.
Swap
swapon --show
free -h
vmstat 1
Swap being used is not automatically a problem.
More useful:
si
so
in vmstat.
Sustained swap-in/out plus poor performance can indicate memory pressure.
OOM
Process suddenly disappears with Killed?
Check:
dmesg -T | grep -Ei 'oom|out of memory|killed process'
journalctl -k | grep -Ei 'oom|out of memory|killed process'
Example:
Out of memory: Killed process 1234 (java)
That means the kernel killed it, not that the application exited normally.
Disk Space
Check the filesystem containing the failing path:
df -h /application/path
findmnt -T /application/path
Find large directories:
du -xhd1 /var
du -xhd1 /var/lib
Sort:
du -xhd1 /var | sort -h
-x avoids crossing into other filesystems.
Inodes
No space left on device does not always mean disk blocks are full.
Check:
df -i
df -i /application/path
A filesystem can have plenty of bytes free and zero inodes left.
Common cause:
millions of tiny files
df and du Disagree
Classic:
df says filesystem is full
du cannot find the space
Check deleted files still held open:
lsof +L1
Example:
java 1234 app 5w REG ... 20G /var/log/app.log (deleted)
Deleting the filename does not release the blocks until the process closes the file.
Disk I/O
iostat -xz 1
Useful fields:
| Field | Meaning |
|---|---|
await |
I/O latency |
aqu-sz |
queue depth |
%util |
device busy time |
High latency plus growing queues is more interesting than any single field alone.
Find the process:
pidstat -d 1
iotop -oPa
Do not just say “the disk is slow.” Find what is hammering it.
Filesystems
findmnt
findmnt -T /path
lsblk -f
Read-only filesystem:
findmnt -T /affected/path
dmesg -T | tail -n 100
If a filesystem remounted itself read-only after errors, do not blindly remount it read-write.
Find out why.
Permissions
ls -l /path/to/file
id
Check the entire path:
namei -l /path/to/file
Also consider:
parent directory permissions
service user
group membership
ACLs
SELinux
AppArmor
mount options
ACLs:
getfacl /path/to/file
Open File Descriptors
Too many open files:
cat /proc/<pid>/limits
ls /proc/<pid>/fd | wc -l
lsof -p <pid>
Look for:
Max open files
Current shell:
ulimit -n
ulimit -n for your shell does not necessarily match an already-running service.
systemd
Status:
systemctl status <service>
Failed services:
systemctl --failed
Logs:
journalctl -u <service> -n 100
journalctl -u <service> -f
journalctl -u <service> --since '-30 min'
Inspect the unit:
systemctl cat <service>
Useful properties:
systemctl show <service> \
-p User \
-p Group \
-p ExecStart \
-p MainPID \
-p Result \
-p NRestarts
active (running) means the process is alive. It does not prove the application works.
Works in Shell, Fails Under systemd
Compare:
user
group
PATH
environment
working directory
HOME
limits
Inspect:
systemctl cat <service>
systemctl show <service>
cat /proc/<pid>/limits
Interactive shells and systemd services often run with different environments.
Service Active, App Unavailable
systemctl status <service>
journalctl -u <service> -n 100
ss -ltnp
Think:
wrong port
wrong bind address
listener failed
wrapper alive, child dead
dependency unavailable
process alive but stuck
Process health is not application health.
Logs
File:
tail -n 100 application.log
tail -f application.log
grep -Ei 'error|fail|exception|timeout|killed' application.log
With context:
grep -C 5 -i error application.log
Prefer the incident time window over searching months of logs.
For systemd:
journalctl -u <service> \
--since '2026-09-04 13:00' \
--until '2026-09-04 13:30'
Look for the first meaningful error.
Later errors may just be consequences.
Kernel Logs
dmesg -T
journalctl -k
Useful for:
OOM kills
filesystem errors
disk errors
device failures
driver problems
kernel warnings
Recent:
journalctl -k --since '-30 min'
Process Alive but Hung
ps -o pid,stat,wchan:32,etime,cmd -p <pid>
top -H -p <pid>
lsof -p <pid>
When needed:
strace -p <pid>
Useful repeated calls may include:
connect()
read()
write()
futex()
poll()
openat()
strace can add overhead and expose sensitive data. Do not use it as the first command on every production process.
Signals
Graceful termination:
kill <pid>
This sends SIGTERM.
Force kill:
kill -9 <pid>
This sends SIGKILL.
Prefer graceful shutdown.
SIGKILL gives the process no chance to flush data or run cleanup handlers.
Recent Changes
System logs:
journalctl --since '-1 hour'
Debian / Ubuntu:
grep -iE 'install|upgrade|remove' /var/log/dpkg.log
RHEL / Fedora:
dnf history
Always ask:
what changed immediately before this started?
Clock Problems
Bad system time can break:
TLS
authentication
tokens
Kerberos
log ordering
distributed systems
Check:
date
timedatectl
Chrony:
chronyc tracking
chronyc sources
Docker
docker ps
docker ps -a
docker logs <container>
docker inspect <container>
docker stats
docker top <container>
Container keeps restarting:
application exits
bad command
missing env var
missing file
permission error
dependency unavailable
OOM
Do not assume Docker is broken because the application inside it keeps dying.
Kubernetes
First look:
kubectl get pods
kubectl describe pod <pod>
kubectl logs <pod>
kubectl get events --sort-by=.lastTimestamp
Previous crashed container:
kubectl logs <pod> --previous
Resources:
kubectl top pod <pod> --containers
kubectl top node
CrashLoopBackOff
Means roughly:
container starts
→ exits
→ restarts
→ exits again
→ backoff increases
It is a symptom, not the root cause.
Check:
kubectl describe pod <pod>
kubectl logs <pod> --previous
Think:
startup failure
bad command
missing secret
missing env var
permission error
dependency failure
probe failure
OOMKilled
OOMKilled
kubectl describe pod <pod>
kubectl logs <pod> --previous
kubectl top pod <pod> --containers
A pod can hit its own memory limit while the node still has plenty of free RAM.
Do not immediately increase the limit. Check whether usage is expected, spiky, or leaking.
Common Patterns
High load, low CPU
vmstat 1
iostat -xz 1
ps -eo pid,stat,wchan:32,comm
Think I/O, NFS, network storage, blocked tasks.
Memory looks full
free -h
Look at available, not just free.
Process suddenly disappeared
systemctl status <service>
journalctl -u <service>
journalctl -k | grep -Ei 'oom|out of memory|killed process'
No space left, but df looks fine
df -h /path
df -i /path
findmnt -T /path
lsof +L1
Think:
wrong filesystem
inodes
quota
/tmp
/dev/shm
deleted open file
Permission denied, but file looks correct
namei -l /full/path/to/file
id
getfacl /full/path/to/file
App keeps restarting
systemctl status <service>
journalctl -u <service>
systemctl show <service> -p NRestarts -p Result
Find why it exits before debugging the restart policy.
Do Not Restart Blindly
Restarting may restore service, but it can also destroy evidence.
When practical, capture:
logs
process state
resource usage
open files
kernel messages
current config
failure time
Then restart if needed.
Difference:
restart as mitigation
vs:
restart because I have no idea
Verify the Fix
Do not stop at:
service is active
Repeat the original failing action.
fix
→ repeat original test
→ check logs
→ check resources
→ watch for recurrence
Mitigation, root cause, and prevention are different things.
restart service
→ mitigation
memory leak
→ root cause
fix leak + alerting
→ prevention
Minimal Cheat Sheet
Host
uptime
free -h
df -h
df -i
Processes
pgrep -af name
ps aux
pstree -ap
top
pidstat 1
CPU / Load
top
uptime
vmstat 1
pidstat 1
Memory
free -h
vmstat 1
ps aux --sort=-%mem
Disk
df -h
df -i
du -xhd1 /path
findmnt -T /path
iostat -xz 1
pidstat -d 1
Services
systemctl status <service>
systemctl --failed
systemctl cat <service>
journalctl -u <service>
Kernel
dmesg -T
journalctl -k
Open files
lsof -p <pid>
lsof +L1
Permissions
id
namei -l /path
getfacl /path
Docker
docker ps -a
docker logs <container>
docker inspect <container>
docker stats
Kubernetes
kubectl get pods
kubectl describe pod <pod>
kubectl logs <pod>
kubectl logs <pod> --previous
kubectl get events --sort-by=.lastTimestamp
Most troubleshooting is not knowing a magical command.
It is narrowing the problem until the next test is obvious.