Skip to content
elephantoo

sed & awk

Lesson 15 of 31 18 min read

Find-and-replace and in-place edits with sed; fields, patterns and reports with awk.


sed and awk have been around since the 1970s and are still on every Linux machine, in every Docker image and in countless deployment scripts. sed (stream editor) is the tool for search-and-replace and line edits. awk is a small programming language built for columns: filtering rows, summing values and printing reports. You don't need to master them; the 20% covered here handles 80% of real-world jobs.

Sample files: app.conf

Output
# App settings
port=8080
host=localhost
debug=true
# db
db_host=localhost

and sales.csv from the previous lesson:

Output
date,region,product,units,price
2026-09-01,north,laptop,3,55000
2026-09-01,south,phone,10,18000
2026-09-02,north,phone,7,18000
2026-09-02,east,tablet,4,25000
2026-09-03,south,laptop,2,55000
2026-09-03,north,laptop,5,55000
2026-09-04,east,phone,12,18000

sed: substitution#

The command you'll use 90% of the time is s/pattern/replacement/flags:

Terminal
echo "I love cats. cats are great" | sed 's/cats/dogs/'
echo "I love cats. cats are great" | sed 's/cats/dogs/g'
Output
I love dogs. cats are great
I love dogs. dogs are great

Without g, only the first match on each line is replaced. Other flags: I for case-insensitive (GNU sed), and a number like 2 for "only the 2nd match".

Terminal
sed 's/localhost/db.internal/' app.conf
Output
# App settings
port=8080
host=db.internal
debug=true
# db
db_host=db.internal

sed prints the result; the file is unchanged. When the pattern contains slashes, pick another delimiter so you don't drown in backslashes:

Terminal
echo "/usr/local/bin/tool" | sed 's|/usr/local|/opt|'
Output
/opt/bin/tool

Capture groups

With -E (extended regex, as in grep -E), parentheses capture parts of the match and \1, \2... reuse them:

Terminal
echo "2026-10-01" | sed -E 's/([0-9]{4})-([0-9]{2})-([0-9]{2})/\3\/\2\/\1/'
echo "John Smith" | sed -E 's/(\w+) (\w+)/\2, \1/'
Output
01/10/2026
Smith, John

& in the replacement stands for the whole match: sed 's/[0-9]\+/[&]/' wraps the first number in brackets.

sed: selecting and deleting lines#

Put an address before a command to choose which lines it applies to: a line number, a range, or a /regex/.

Terminal
sed -n '2,3p' app.conf        # -n: print nothing by default; p: print these lines
Output
port=8080
host=localhost
Terminal
sed '/^#/d' app.conf          # d: delete comment lines
Output
port=8080
host=localhost
debug=true
db_host=localhost

Several commands with -e:

Terminal
sed -e '/^#/d' -e '/^$/d' -e 's/=/ = /' app.conf
Output
port = 8080
host = localhost
debug = true
db_host = localhost

More one-liners:

Terminal
sed -n '/^db/p' app.conf               # like grep
sed '1i # generated file' app.conf     # insert a line before line 1
sed '/^port=/a timeout=30' app.conf    # append a line after the match
sed '5q' big.log                        # print first 5 lines, then quit (like head)
sed 's/\r$//' windows.txt               # strip Windows line endings

sed: editing files in place#

Terminal
sed -i.bak 's/^debug=true/debug=false/' app.conf
grep debug app.conf app.conf.bak
Output
app.conf:debug=false
app.conf.bak:debug=true

-i edits the file directly, and -i.bak keeps the original with a .bak suffix. Typical server automation:

Terminal
sudo sed -i.bak 's/^#\?PasswordAuthentication .*/PasswordAuthentication no/' /etc/ssh/sshd_config

Test without -i first and check the output. On macOS/BSD sed, -i requires an argument (sed -i '' ...), a common portability trap.

awk: thinking in columns#

An awk program is a list of pattern { action } rules. awk reads input line by line (each line is a record), splits it into fields $1, $2, ... ($0 is the whole line), and runs the action for every line where the pattern is true.

Terminal
awk -F, 'NR > 1 {print $3, $4}' sales.csv | head -3
Output
laptop 3
phone 10
phone 7

-F, sets the field separator (default: any run of spaces/tabs, which is perfect for command output). NR > 1 skips the header.

Built-inMeaning
$0The whole line
$1, $2...Fields
NFNumber of fields ($NF = last field)
NRCurrent line number
FS / OFSInput / output field separator
BEGIN {} / END {}Run before the first line / after the last

Filtering rows

A pattern on its own prints matching lines:

Terminal
awk -F, 'NR > 1 && $4 > 5' sales.csv
Output
2026-09-01,south,phone,10,18000
2026-09-02,north,phone,7,18000
2026-09-04,east,phone,12,18000

Regex matches on a field use ~:

Terminal
awk -F: '$7 ~ /bash$/ {print $1}' /etc/passwd     # users whose shell is bash
df -h / | awk 'NR == 2 {print "Root disk used:", $5}'
Output
root
ada
Root disk used: 17%

(Your users and disk figures will differ.)

Sums, counts and groups

Variables need no declaration and start at 0 (or empty):

Terminal
awk -F, 'NR > 1 {sum += $4} END {print "Total units:", sum}' sales.csv
Output
Total units: 43

Associative arrays group by any key. Here's the question cut | sort | uniq couldn't answer, units sold per product:

Terminal
awk -F, 'NR > 1 {units[$3] += $4} END {for (p in units) print p, units[p]}' sales.csv | sort
Output
laptop 10
phone 29
tablet 4

(for (p in ...) visits keys in no particular order, hence the sort.) Revenue per region, formatted with printf:

Terminal
awk -F, 'NR > 1 {r[$2] += $4 * $5}
         END {for (k in r) printf "%-6s %9.2f lakh\n", k, r[k] / 100000}' sales.csv | sort -k2 -nr
Output
north       5.66 lakh
east        3.16 lakh
south       2.90 lakh

printf works like in C: %s string, %d integer, %.2f two decimals, %-6s left-aligned in 6 columns.

Handy awk one-liners

Terminal
awk '{print NF, $NF}' file              # field count and last field of each line
awk 'END {print NR}' file               # count lines (like wc -l)
awk 'length > 80' file                  # lines longer than 80 characters
awk '!seen[$0]++' file                  # remove duplicates, keeping first-seen order (no sort needed)
awk -F, -v OFS='\t' '{print $2, $3}' sales.csv   # CSV to tab-separated
ps aux | awk '$3 > 50 {print $2, $11}'  # PID and command of processes using >50% CPU

-v name=value passes a shell value into awk safely: awk -v min="$LIMIT" '$4 > min' ....

sed or awk?#

  • sed: replace text, delete or insert lines, edit config files in place.
  • awk: anything about columns: pick fields, filter by value, sum, count, group, format reports.
  • Neither: for nested JSON use jq; for complex CSV with quoted commas, or logic longer than a line or two, write a Python script.

Common mistakes#

  • Forgetting g and replacing only the first match per line.
  • Redirecting sed's output into the input file (> same.txt), which empties it. Use -i.
  • Using double quotes around awk programs, so the shell expands $1 before awk sees it. Always single-quote awk programs.
  • Forgetting -F, for CSV (awk splits on whitespace by default).
  • Using basic regex in sed and wondering why + or () don't work: add -E.

What's next#

Next: bundling and shrinking files with archives and compression: tar, gzip, xz, zstd and zip.

Check your understanding

Quick quiz

0/3 answered
  1. 1.What does sed 's/cat/dog/' pets.txt do to a line containing cat cat?

  2. 2.In awk, what do $1, $NF and NR mean?

  3. 3.Which command safely edits app.conf in place while keeping a backup copy?

Finished reading?

Mark this lesson complete to track your progress.