Text processing: cut, sort, uniq, tr & more
Slice, sort, count, translate and join text, and build data pipelines from small tools.
Logs, CSV exports, command output, config files: so much of Linux is text arranged in lines and columns. A handful of small tools let you slice, sort, count and compare that text without opening a spreadsheet or writing a program. Combined with pipes, they answer surprisingly complex questions in one line.
Our sample file, sales.csv:
cut: pick columns#
-f2-4 selects a range and -f3- means "field 3 to the end". cut -c1-7 cuts by character position instead, useful for fixed-width output. cut only accepts a single-character delimiter and can't handle runs of spaces; for space-aligned output like ps or df, use awk (next lesson) or squeeze the spaces first with tr -s ' '.
Skip a header line with tail -n +2, which you'll see in every example below.
sort: order lines#
Sort the sales by units, biggest first:
Several keys: by region, then units descending within each region:
Always give the end field (-k4,4, not just -k4), otherwise the key runs to the end of the line. Note that the sort order of text depends on your locale (LANG); LC_ALL=C sort gives plain byte order and is faster on huge files.
uniq: collapse duplicates#
uniq removes adjacent duplicate lines, so you almost always sort first:
"What's most common?" is sort | uniq -c | sort -rn | head, one of the most useful pipelines you'll ever learn. It works on IP addresses in logs, error messages, HTTP status codes, words in a text, and anything else.
tr: translate or delete characters#
tr works on characters, and reads only from stdin:
-d deletes characters, -s squeezes runs, and -c complements the set ("everything except"). A common use is stripping Windows line endings: tr -d '\r' < win.txt > unix.txt.
wc, nl, paste#
paste joins files side by side, line by line:
paste -s joins all lines of one file into a single line, which makes a quick sum possible:
Comparing files: comm and diff#
comm needs sorted input and prints three columns: only in the first file, only in the second, in both. -23 gives "only in a.txt", which is great for "which users are in list A but not list B?".
diff shows how to turn one file into the other:
< lines are from the first file and > lines from the second. The unified format is easier to read and is what Git uses:
(The two header lines show the file names and modification times, so yours will differ.) diff exits with status 0 if the files are identical and 1 if they differ, so diff -q old.conf new.conf || echo "config changed" works in scripts. diff -r dir1 dir2 compares directories.
Putting it together#
Which product sold the most units overall? That needs a sum per group, which is where these tools reach their limit and awk shines. But plenty of questions fall to a pipeline:
Build pipelines one stage at a time: run the first command, check the output, add | next-command, check again.
Common mistakes#
uniqwithoutsort.- Text sort on numbers (
sortinstead ofsort -n). sort -k2without an end field, or forgetting-t,for CSV.- Expecting
cutto handle multiple spaces or quoted CSV fields with commas inside. For real-world CSV, use a proper tool (python3 -c "import csv...",csvkit, ormlr). - Using
common unsorted files.
What's next#
sed and awk are the power tools of text processing: search-and-replace across files, and a tiny programming language for columns, sums and reports.
Check your understanding
Quick quiz
1.Why does
cut -d, -f2 data.csv | uniq -coften give wrong counts?2.
printf '10\n9\n100\n' | sortprints10,100,9. How do you sort them as numbers?3.What does
tr -s ' 'do to the texttoo many spaces?
Finished reading?
Mark this lesson complete to track your progress.