Unix systems treat text as a universal interface.

Logs, configuration files, command outputs, and pipelines are often exposed as text streams.

Two classic tools dominate command-line text processing:

  • sed — stream editor for transforming text
  • awk / gawk — pattern-processing language for structured text

Understanding how these tools think about input makes everyday debugging much easier.


Stream Processing Philosophy

Unix tools operate on streams rather than requiring the entire input to be loaded into memory.

Each tool performs a small transformation and passes the result downstream.

Typical pipeline:

sed ... file | awk ... | sort | uniq

A useful mental shortcut:

sed  → transform text
awk  → analyze structured records

Examples:

sed 's/foo/bar/' file.txt
awk '{print $1}' file.txt

cat file | ... is sometimes useful when demonstrating a pipeline, but when a command can read the file directly, passing the filename is usually simpler.


sed Execution Model

sed processes input one cycle at a time.

For normal line-oriented input, the cycle is roughly:

read line → pattern space
apply commands
print pattern space
repeat

By default, sed automatically prints the pattern space at the end of each cycle.

The -n option disables this automatic printing:

sed -n '...' file.txt

When -n is used, output normally appears only when a command such as p explicitly prints it.

Two internal buffers control much of sed’s behaviour.


Pattern Space

The pattern space contains the text currently being processed.

Example:

sed 's/foo/bar/' file.txt

Processing:

pattern space = "foo hello"
apply substitution
pattern space = "bar hello"
output

Normally the pattern space contains one line, but commands such as N, G, and H can create multi-line buffers.


Hold Space

The hold space is persistent storage that survives between cycles.

Important commands:

h   pattern → hold, replacing hold space
H   pattern → hold, appending
g   hold → pattern, replacing pattern space
G   hold → pattern, appending
x   swap pattern and hold spaces

Mental model:

pattern space = working memory
hold space    = persistent memory

One subtle detail matters:

H and G append a newline before the copied content.

This newline behaviour explains many initially strange-looking multi-line sed scripts.


sed Primitives

Most sed scripts rely on a relatively small set of primitive operations.

s   substitute

p   print pattern space
P   print through the first newline

d   delete pattern space and start the next cycle
D   delete through the first newline and restart the current cycle

n   read the next line into pattern space
N   append the next line to pattern space

h   copy pattern space to hold space
H   append pattern space to hold space
g   copy hold space to pattern space
G   append hold space to pattern space
x   swap pattern and hold spaces

:   define label
b   unconditional branch
t   branch if substitution succeeded

q   quit early

The lowercase and uppercase commands form useful pairs:

n   next line, replace
N   next line, append

p   print everything
P   print through first newline

d   delete everything
D   delete through first newline

h   hold, replace
H   hold, append

g   get, replace
G   get, append

A useful way to group them mentally:

editing
    s

output and cycle control
    p P d D

multi-line input
    n N

memory
    h H g G x

control flow
    : b t q

Less frequently needed, but still useful:

a   append text
i   insert text
c   replace selected text
=   print current input line number
l   display pattern space unambiguously

Complex one-liners are usually combinations of a surprisingly small number of these primitives.


sed Patterns

Certain patterns appear repeatedly in real sed scripts.


Loop Pattern

sed supports simple loops using labels and conditional branching.

sed ':a; s/foo/bar/; ta' file.txt

Breakdown:

:a         → define label
s/foo/bar/ → replace first occurrence
ta         → jump back if substitution succeeded

Conceptually:

repeat substitution
until no more matches exist

Input:

foo foo foo

Output:

bar bar bar

A global substitution is simpler for this exact case:

sed 's/foo/bar/g' file.txt

The loop example is useful because the same :, t, and b primitives can drive more complicated transformations.


Reverse Stream Pattern

Classic example:

sed '1!G; h; $!d' file.txt

Input:

A
B
C

Output:

C
B
A

Conceptually:

1!G   → append the previously stored lines
h     → store the new accumulated state
$!d   → suppress output until the final input line

The hold space accumulates the stream in reverse order.

This is a good example of why understanding h, G, and d is more useful than memorizing the one-liner itself.


Sliding Window Pattern

sed can maintain a rolling multi-line buffer.

Example: print everything except the last five lines.

sed -n -e ':a; 1,5!{P; N; D}; N; ba' text.txt

The script gradually builds a multi-line pattern space.

Once enough lines are buffered:

P   → print the oldest buffered line
N   → append another input line
D   → remove the oldest line and restart the cycle

The final five lines remain buffered and are never printed.

This works, but it is also a good example of where sed starts becoming difficult to read.


Multi-Line Merge Pattern

Example: merge lines ending with a continuation character.

sed ':a; /\\$/N; s/\\\n//; ta' file.txt

Input:

hello world \
continued line

Output:

hello world continued line

Logic:

/\\$/N     → if the line ends in "\", append the next line
s/\\\n//   → remove the continuation marker and newline
ta         → repeat if the substitution succeeded

awk Execution Model

awk treats input as records containing fields.

By default:

record = one input line
field separator = whitespace

Processing loop:

read record
split into fields
evaluate pattern
execute action

Important variables:

$0   full record
$1   first field
$2   second field

NF   number of fields in current record
NR   total record number
FNR  record number within current file

FS   input field separator
OFS  output field separator

RS   input record separator
ORS  output record separator

Example:

awk '{print $1}' file.txt

awk vs gawk

awk is the language.

gawk is the GNU implementation of awk.

Most examples in this guide use portable awk syntax and therefore work with gawk, BSD awk, mawk, and other common implementations.

GNU gawk also provides useful extensions such as:

gensub()
FPAT
PROCINFO
BEGINFILE / ENDFILE

If portability matters, prefer standard awk features unless a GNU-specific feature clearly simplifies the task.


awk Primitives

Core awk operations include:

  • field extraction
  • conditional filtering
  • arithmetic
  • string substitution
  • associative arrays
  • record control
  • BEGIN and END blocks

Examples:

awk '{print $1}' file.txt

Filter:

awk '$3 == "ERROR"' file.txt

Run setup code before reading input:

awk 'BEGIN {FS="\t"} {print $2}' file.tsv

Run final aggregation after all input is processed:

awk '{sum += $1} END {print sum}' values.txt

Real Example: FASTQ to FASTA

A FASTQ record contains four logical lines:

@read_id
SEQUENCE
+
QUALITY

FASTA requires:

>read_id
SEQUENCE

Because FASTQ has a fixed four-line record structure, it is safer to process records by position rather than assuming that every line beginning with @ is a header.

A quality line may also begin with @.


sed — GNU Step Addressing

GNU sed supports first~step addressing:

sed -n '1~4s/^@/>/p; 2~4p' input.fastq > output.fasta

Explanation:

1~4 → lines 1, 5, 9, ...   → headers
2~4 → lines 2, 6, 10, ...  → sequences

The header lines are converted from:

@read_id

to:

>read_id

This syntax is concise, but first~step is a GNU sed extension and is not portable to every sed implementation, including the default BSD sed shipped with macOS.


sed — Portable Record Traversal

The same transformation can be expressed using n:

sed -n '
s/^@/>/p
n
p
n
n
' input.fastq > output.fasta

For every four-line FASTQ record:

line 1 → replace @ with > and print
line 2 → read and print
line 3 → read and discard
line 4 → read and discard

Then sed begins the next cycle at the following FASTQ header.

Conceptually:

header   → print
sequence → print
+        → skip
quality  → skip

This assumes the input is valid four-line FASTQ.

A tempting implementation is:

sed -e '/^@/!d; s//>/; N' input.fastq

but this identifies records only by a leading @.

That is not structurally safe for arbitrary FASTQ because quality strings may also begin with @.


awk Implementation

awk can express the four-line record structure directly:

awk '
NR % 4 == 1 {
    sub(/^@/, ">")
    print
    getline
    print
}
' input.fastq > output.fasta

Explanation:

NR % 4 == 1 → FASTQ header
sub(...)     → convert @ to >
getline      → read sequence
print        → output sequence

Here record position determines what each line means rather than its contents.

For production bioinformatics workflows, dedicated sequence parsers are preferable when records may contain wrapped sequences, malformed input, or other non-trivial cases.


awk Aggregation

awk associative arrays make quick aggregation easy.

Example input:

user1 200
user2 150
user1 300

Command:

awk '{sum[$1] += $2} END {for (u in sum) print u, sum[u]}' file.txt

Possible output:

user1 500
user2 150

Associative-array iteration order is not guaranteed.

If stable ordering matters, pipe the result to sort:

awk '{sum[$1] += $2} END {for (u in sum) print u, sum[u]}' file.txt | sort

Deduplicating While Preserving Order

awk can remove duplicates without sorting:

awk '!seen[$0]++' file.txt

How it works:

seen[$0]      → count occurrences of the current line
!seen[$0]++   → true only on the first occurrence

Unlike:

sort -u

this preserves the original input order.


Log Analysis Pipelines

These pipelines appear constantly during debugging.

Top IP addresses:

awk '{print $1}' access.log | sort | uniq -c | sort -nr | head

Count errors per service:

grep 'ERROR' application.log | awk '{print $3}' | sort | uniq -c

Slow requests when the last field contains latency in milliseconds:

awk '$NF > 1000' access.log

Equivalent counting can often be performed entirely inside awk:

awk '{count[$1]++} END {for (ip in count) print count[ip], ip}' access.log \
| sort -nr \
| head

Which form is clearer depends on the task.


sed and awk Together

In practice these tools are often chained so each stage performs one small transformation.

Example: count log entries per minute.

Input:

[2026-03-06 14:12:33] INFO request completed
[2026-03-06 14:12:40] ERROR timeout

Pipeline:

sed 's/^\[\(....-..-.. ..:..\):..]/\1/' application.log \
| awk '{count[$1" "$2]++} END {for (t in count) print t, count[t]}' \
| sort

Pipeline logic:

sed  → remove seconds from timestamp
awk  → count entries per minute
sort → order results chronologically

Example output:

2026-03-06 14:12 34
2026-03-06 14:13 27
2026-03-06 14:14 31

Here the roles are clear:

sed  → normalize text structure
awk  → perform aggregation

Each stage performs one transformation, which keeps the pipeline easier to reason about.


Common Pitfalls

Automatic Printing in sed

sed prints the pattern space automatically unless -n is used.

This command:

sed 's/foo/bar/p' file.txt

may print matching lines twice:

once from p
once from sed's normal end-of-cycle output

Use:

sed -n 's/foo/bar/p' file.txt

when you want only explicitly selected output.


Shell Quoting

Prefer single quotes around sed and awk programs:

sed 's/foo/bar/' file.txt
awk '{print $1}' file.txt

Double quotes allow shell expansion of characters such as $, backticks, and backslashes.

Sometimes expansion is intentional:

awk -v threshold="$LIMIT" '$NF > threshold' file.txt

Passing shell values through -v is generally safer than interpolating them directly into an awk program.


Field Separators

Simple comma-separated fields can be split with:

awk -F, '{print $2}' file.csv

But this is not a complete CSV parser.

Quoted fields such as:

alice,"hello, world",42

contain commas that are part of the field value.

For real CSV with quoting, escaping, or embedded newlines, use a CSV-aware parser.


Regex Expectations

sed traditionally uses Basic Regular Expressions by default.

For example:

.*

matches the longest possible sequence allowed by the surrounding pattern.

Some implementations support extended regular expressions with:

sed -E '...'

GNU sed also historically supported -r.

When portability matters, check which syntax is available on the target system.


GNU vs BSD sed

Not every sed behaves identically.

A common portability trap is in-place editing.

GNU sed commonly uses:

sed -i 's/foo/bar/' file.txt

BSD/macOS sed commonly requires an explicit backup suffix argument:

sed -i '' 's/foo/bar/' file.txt

Step addressing such as:

1~4

is another GNU extension.

For scripts that must run across Linux and macOS, test the exact sed implementation being used.


Overly Complex sed

sed excels at compact stream transformations.

Once scripts rely heavily on multi-line buffers, hold space, and control flow, readability drops quickly.

For example, this command prints everything except the last five lines:

sed -n -e ':a; 1,5!{P; N; D}; N; ba' text.txt

The logic is valid, but difficult to understand at first glance.

The same task can often be expressed more directly in awk:

awk '
{
    buffer[NR % 6] = $0

    if (NR > 5)
        print buffer[(NR - 5) % 6]
}
' text.txt

Explanation:

NR % 6      → store lines in a circular buffer
NR > 5      → wait until at least five lines have been read
(NR - 5)%6  → print the line that is five lines behind

Why % 6?

To print everything except the last five lines, the program needs five lines of look-ahead plus one slot for the newly read line.

skip last N lines → circular buffer size = N + 1

For five lines:

buffer size = 6

Both commands solve the same problem.

The sed version expresses it through pattern-space manipulation and cycle control.

The awk version expresses the state directly with an array and index arithmetic.

When transformations become stateful or algorithmic, awk is usually easier to read and maintain.


Choosing Between sed and awk

A practical rule of thumb:

Use sed when the task is mainly:

substitute text
delete or select lines
perform small local rewrites
join a few neighbouring lines

Use awk when the task involves:

fields
numbers
conditions
aggregation
state
associative arrays
record-oriented logic

And when either script starts becoming difficult to explain, consider moving the logic into a general-purpose language.

A Unix one-liner is valuable when it stays understandable.


Closing Thoughts

The real power of sed and awk comes from thinking in streams.

Instead of writing a large program for every small problem, engineers can compose transformations that progressively reshape data.

This makes it possible to:

  • inspect large log files quickly
  • reshape command output
  • prototype data-processing workflows
  • aggregate tabular data
  • validate assumptions directly from the shell
  • debug production systems with tools available almost everywhere

The important skill is not memorizing one-liners.

It is understanding the execution model well enough to build them, read them, and know when to stop using them.

sed gives you a tiny state machine for transforming streams.

awk gives you a small record-processing language.

Together they remain some of the most useful tools in the Unix toolbox.