A condensed reference for troubleshooting Linux hosts, services, processes, resources, containers, and common failures.

Not a Linux administration tutorial. Just the things that repeatedly matter when something is broken.


Mental Model

Start here:

symptom
→ scope
→ evidence
→ hypothesis
→ test
→ fix
→ verify

Useful questions:

What exactly is broken?
One user, one host, or everything?
What changed recently?
What does the evidence actually prove?
What do I still not know?
What is the cheapest test for the next hypothesis?

Do not change five things at once.


Evidence Has Limits

systemctl says active
→ process is alive
→ does not prove the app is healthy
CPU is idle
→ CPU is not the bottleneck right now
→ does not prove the machine is healthy
df shows free space
→ filesystem has free blocks
→ does not prove inodes, quotas, or another filesystem are fine
process exists
→ process has not exited
→ does not prove it is making progress

Keep facts and guesses separate.


First Look

uptime
free -h
df -h
df -i
systemctl --failed
ps aux --sort=-%cpu | head
ps aux --sort=-%mem | head

Live:

top
vmstat 1

Storage:

iostat -xz 1

The goal is not to diagnose everything here. It is to decide where to look next.


Processes

Find a process:

pgrep -af <name>

Inspect:

ps -fp <pid>
ps -o pid,ppid,user,stat,%cpu,%mem,rss,vsz,etime,cmd -p <pid>

Process tree:

pstree -ap
ps -ef --forest

Useful fields:

Field Meaning
PID process ID
PPID parent PID
STAT process state
RSS resident memory
VSZ virtual address space
ETIME elapsed runtime

Do not confuse VSZ with actual RAM usage.


Process States

ps -eo pid,ppid,stat,wchan:32,comm
State Meaning
R running / runnable
S sleeping
D uninterruptible sleep, often I/O
T stopped
Z zombie

D state

Find blocked tasks:

ps -eo pid,stat,wchan:32,comm | awk '$2 ~ /^D/'

Think:

disk
NFS
network storage
filesystem
device I/O

High load + idle CPU + many D tasks usually points toward I/O.

Zombies

ps -eo pid,ppid,stat,comm | awk '$3 ~ /^Z/'

A zombie is already dead. Its parent has not reaped it yet.

Find the parent:

ps -fp <ppid>

CPU

top
ps aux --sort=-%cpu | head
pidstat 1

One process:

pidstat -p <pid> 1
top -H -p <pid>

Ask:

one process or many?
one core or all cores?
constant or intermittent?
user CPU or kernel CPU?
traffic spike?
retry loop?
bad deployment?

Useful top fields:

Field Meaning
us user-space CPU
sy kernel CPU
id idle
wa I/O wait
st stolen CPU

Load Average

uptime

Example:

load average: 12.4, 10.8, 8.1

Roughly 1, 5, and 15 minutes.

On Linux, load includes tasks that are runnable and tasks stuck in uninterruptible sleep.

So:

high load ≠ necessarily high CPU

High load with idle CPU:

vmstat 1
iostat -xz 1
ps -eo pid,stat,wchan:32,comm

In vmstat:

Field Meaning
r runnable tasks
b blocked tasks
si swap in
so swap out
wa I/O wait

Memory

free -h

The important number is usually:

available

not:

free

Linux uses spare RAM for cache. Low completely-unused memory is normal.

Find consumers:

ps aux --sort=-%mem | head
ps -eo pid,user,rss,%mem,etime,cmd --sort=-rss | head

Watch:

watch -n 1 free -h
vmstat 1

A memory leak is usually interesting because usage keeps growing.


Swap

swapon --show
free -h
vmstat 1

Swap being used is not automatically a problem.

More useful:

si
so

in vmstat.

Sustained swap-in/out plus poor performance can indicate memory pressure.


OOM

Process suddenly disappears with Killed?

Check:

dmesg -T | grep -Ei 'oom|out of memory|killed process'
journalctl -k | grep -Ei 'oom|out of memory|killed process'

Example:

Out of memory: Killed process 1234 (java)

That means the kernel killed it, not that the application exited normally.


Disk Space

Check the filesystem containing the failing path:

df -h /application/path
findmnt -T /application/path

Find large directories:

du -xhd1 /var
du -xhd1 /var/lib

Sort:

du -xhd1 /var | sort -h

-x avoids crossing into other filesystems.


Inodes

No space left on device does not always mean disk blocks are full.

Check:

df -i
df -i /application/path

A filesystem can have plenty of bytes free and zero inodes left.

Common cause:

millions of tiny files

df and du Disagree

Classic:

df says filesystem is full
du cannot find the space

Check deleted files still held open:

lsof +L1

Example:

java  1234  app  5w  REG  ...  20G  /var/log/app.log (deleted)

Deleting the filename does not release the blocks until the process closes the file.


Disk I/O

iostat -xz 1

Useful fields:

Field Meaning
await I/O latency
aqu-sz queue depth
%util device busy time

High latency plus growing queues is more interesting than any single field alone.

Find the process:

pidstat -d 1
iotop -oPa

Do not just say “the disk is slow.” Find what is hammering it.


Filesystems

findmnt
findmnt -T /path
lsblk -f

Read-only filesystem:

findmnt -T /affected/path
dmesg -T | tail -n 100

If a filesystem remounted itself read-only after errors, do not blindly remount it read-write.

Find out why.


Permissions

ls -l /path/to/file
id

Check the entire path:

namei -l /path/to/file

Also consider:

parent directory permissions
service user
group membership
ACLs
SELinux
AppArmor
mount options

ACLs:

getfacl /path/to/file

Open File Descriptors

Too many open files:

cat /proc/<pid>/limits
ls /proc/<pid>/fd | wc -l
lsof -p <pid>

Look for:

Max open files

Current shell:

ulimit -n

ulimit -n for your shell does not necessarily match an already-running service.


systemd

Status:

systemctl status <service>

Failed services:

systemctl --failed

Logs:

journalctl -u <service> -n 100
journalctl -u <service> -f
journalctl -u <service> --since '-30 min'

Inspect the unit:

systemctl cat <service>

Useful properties:

systemctl show <service> \
  -p User \
  -p Group \
  -p ExecStart \
  -p MainPID \
  -p Result \
  -p NRestarts

active (running) means the process is alive. It does not prove the application works.


Works in Shell, Fails Under systemd

Compare:

user
group
PATH
environment
working directory
HOME
limits

Inspect:

systemctl cat <service>
systemctl show <service>
cat /proc/<pid>/limits

Interactive shells and systemd services often run with different environments.


Service Active, App Unavailable

systemctl status <service>
journalctl -u <service> -n 100
ss -ltnp

Think:

wrong port
wrong bind address
listener failed
wrapper alive, child dead
dependency unavailable
process alive but stuck

Process health is not application health.


Logs

File:

tail -n 100 application.log
tail -f application.log
grep -Ei 'error|fail|exception|timeout|killed' application.log

With context:

grep -C 5 -i error application.log

Prefer the incident time window over searching months of logs.

For systemd:

journalctl -u <service> \
  --since '2026-09-04 13:00' \
  --until '2026-09-04 13:30'

Look for the first meaningful error.

Later errors may just be consequences.


Kernel Logs

dmesg -T
journalctl -k

Useful for:

OOM kills
filesystem errors
disk errors
device failures
driver problems
kernel warnings

Recent:

journalctl -k --since '-30 min'

Process Alive but Hung

ps -o pid,stat,wchan:32,etime,cmd -p <pid>
top -H -p <pid>
lsof -p <pid>

When needed:

strace -p <pid>

Useful repeated calls may include:

connect()
read()
write()
futex()
poll()
openat()

strace can add overhead and expose sensitive data. Do not use it as the first command on every production process.


Signals

Graceful termination:

kill <pid>

This sends SIGTERM.

Force kill:

kill -9 <pid>

This sends SIGKILL.

Prefer graceful shutdown.

SIGKILL gives the process no chance to flush data or run cleanup handlers.


Recent Changes

System logs:

journalctl --since '-1 hour'

Debian / Ubuntu:

grep -iE 'install|upgrade|remove' /var/log/dpkg.log

RHEL / Fedora:

dnf history

Always ask:

what changed immediately before this started?

Clock Problems

Bad system time can break:

TLS
authentication
tokens
Kerberos
log ordering
distributed systems

Check:

date
timedatectl

Chrony:

chronyc tracking
chronyc sources

Docker

docker ps
docker ps -a
docker logs <container>
docker inspect <container>
docker stats
docker top <container>

Container keeps restarting:

application exits
bad command
missing env var
missing file
permission error
dependency unavailable
OOM

Do not assume Docker is broken because the application inside it keeps dying.


Kubernetes

First look:

kubectl get pods
kubectl describe pod <pod>
kubectl logs <pod>
kubectl get events --sort-by=.lastTimestamp

Previous crashed container:

kubectl logs <pod> --previous

Resources:

kubectl top pod <pod> --containers
kubectl top node

CrashLoopBackOff

Means roughly:

container starts
→ exits
→ restarts
→ exits again
→ backoff increases

It is a symptom, not the root cause.

Check:

kubectl describe pod <pod>
kubectl logs <pod> --previous

Think:

startup failure
bad command
missing secret
missing env var
permission error
dependency failure
probe failure
OOMKilled

OOMKilled

kubectl describe pod <pod>
kubectl logs <pod> --previous
kubectl top pod <pod> --containers

A pod can hit its own memory limit while the node still has plenty of free RAM.

Do not immediately increase the limit. Check whether usage is expected, spiky, or leaking.


Common Patterns

High load, low CPU

vmstat 1
iostat -xz 1
ps -eo pid,stat,wchan:32,comm

Think I/O, NFS, network storage, blocked tasks.

Memory looks full

free -h

Look at available, not just free.

Process suddenly disappeared

systemctl status <service>
journalctl -u <service>
journalctl -k | grep -Ei 'oom|out of memory|killed process'

No space left, but df looks fine

df -h /path
df -i /path
findmnt -T /path
lsof +L1

Think:

wrong filesystem
inodes
quota
/tmp
/dev/shm
deleted open file

Permission denied, but file looks correct

namei -l /full/path/to/file
id
getfacl /full/path/to/file

App keeps restarting

systemctl status <service>
journalctl -u <service>
systemctl show <service> -p NRestarts -p Result

Find why it exits before debugging the restart policy.


Do Not Restart Blindly

Restarting may restore service, but it can also destroy evidence.

When practical, capture:

logs
process state
resource usage
open files
kernel messages
current config
failure time

Then restart if needed.

Difference:

restart as mitigation

vs:

restart because I have no idea

Verify the Fix

Do not stop at:

service is active

Repeat the original failing action.

fix
→ repeat original test
→ check logs
→ check resources
→ watch for recurrence

Mitigation, root cause, and prevention are different things.

restart service
→ mitigation

memory leak
→ root cause

fix leak + alerting
→ prevention

Minimal Cheat Sheet

Host

uptime
free -h
df -h
df -i

Processes

pgrep -af name
ps aux
pstree -ap
top
pidstat 1

CPU / Load

top
uptime
vmstat 1
pidstat 1

Memory

free -h
vmstat 1
ps aux --sort=-%mem

Disk

df -h
df -i
du -xhd1 /path
findmnt -T /path
iostat -xz 1
pidstat -d 1

Services

systemctl status <service>
systemctl --failed
systemctl cat <service>
journalctl -u <service>

Kernel

dmesg -T
journalctl -k

Open files

lsof -p <pid>
lsof +L1

Permissions

id
namei -l /path
getfacl /path

Docker

docker ps -a
docker logs <container>
docker inspect <container>
docker stats

Kubernetes

kubectl get pods
kubectl describe pod <pod>
kubectl logs <pod>
kubectl logs <pod> --previous
kubectl get events --sort-by=.lastTimestamp

Most troubleshooting is not knowing a magical command.

It is narrowing the problem until the next test is obvious.