sed & awk
Find-and-replace and in-place edits with sed; fields, patterns and reports with awk.
sed and awk have been around since the 1970s and are still on every Linux machine, in every Docker image and in countless deployment scripts. sed (stream editor) is the tool for search-and-replace and line edits. awk is a small programming language built for columns: filtering rows, summing values and printing reports. You don't need to master them; the 20% covered here handles 80% of real-world jobs.
Sample files: app.conf
and sales.csv from the previous lesson:
sed: substitution#
The command you'll use 90% of the time is s/pattern/replacement/flags:
Without g, only the first match on each line is replaced. Other flags: I for case-insensitive (GNU sed), and a number like 2 for "only the 2nd match".
sed prints the result; the file is unchanged. When the pattern contains slashes, pick another delimiter so you don't drown in backslashes:
Capture groups
With -E (extended regex, as in grep -E), parentheses capture parts of the match and \1, \2... reuse them:
& in the replacement stands for the whole match: sed 's/[0-9]\+/[&]/' wraps the first number in brackets.
sed: selecting and deleting lines#
Put an address before a command to choose which lines it applies to: a line number, a range, or a /regex/.
Several commands with -e:
More one-liners:
sed: editing files in place#
-i edits the file directly, and -i.bak keeps the original with a .bak suffix. Typical server automation:
Test without
-ifirst and check the output. On macOS/BSD sed,-irequires an argument (sed -i '' ...), a common portability trap.
awk: thinking in columns#
An awk program is a list of pattern { action } rules. awk reads input line by line (each line is a record), splits it into fields $1, $2, ... ($0 is the whole line), and runs the action for every line where the pattern is true.
-F, sets the field separator (default: any run of spaces/tabs, which is perfect for command output). NR > 1 skips the header.
Filtering rows
A pattern on its own prints matching lines:
Regex matches on a field use ~:
(Your users and disk figures will differ.)
Sums, counts and groups
Variables need no declaration and start at 0 (or empty):
Associative arrays group by any key. Here's the question cut | sort | uniq couldn't answer, units sold per product:
(for (p in ...) visits keys in no particular order, hence the sort.) Revenue per region, formatted with printf:
printf works like in C: %s string, %d integer, %.2f two decimals, %-6s left-aligned in 6 columns.
Handy awk one-liners
-v name=value passes a shell value into awk safely: awk -v min="$LIMIT" '$4 > min' ....
sed or awk?#
- sed: replace text, delete or insert lines, edit config files in place.
- awk: anything about columns: pick fields, filter by value, sum, count, group, format reports.
- Neither: for nested JSON use
jq; for complex CSV with quoted commas, or logic longer than a line or two, write a Python script.
Common mistakes#
- Forgetting
gand replacing only the first match per line. - Redirecting sed's output into the input file (
> same.txt), which empties it. Use-i. - Using double quotes around awk programs, so the shell expands
$1before awk sees it. Always single-quote awk programs. - Forgetting
-F,for CSV (awk splits on whitespace by default). - Using basic regex in sed and wondering why
+or()don't work: add-E.
What's next#
Next: bundling and shrinking files with archives and compression: tar, gzip, xz, zstd and zip.
Check your understanding
Quick quiz
1.What does
sed 's/cat/dog/' pets.txtdo to a line containingcat cat?2.In awk, what do
$1,$NFandNRmean?3.Which command safely edits
app.confin place while keeping a backup copy?
Finished reading?
Mark this lesson complete to track your progress.