Skip to content
elephantoo

Text processing: cut, sort, uniq, tr & more

Lesson 14 of 31 15 min read

Slice, sort, count, translate and join text, and build data pipelines from small tools.


Logs, CSV exports, command output, config files: so much of Linux is text arranged in lines and columns. A handful of small tools let you slice, sort, count and compare that text without opening a spreadsheet or writing a program. Combined with pipes, they answer surprisingly complex questions in one line.

Our sample file, sales.csv:

Output
date,region,product,units,price
2026-09-01,north,laptop,3,55000
2026-09-01,south,phone,10,18000
2026-09-02,north,phone,7,18000
2026-09-02,east,tablet,4,25000
2026-09-03,south,laptop,2,55000
2026-09-03,north,laptop,5,55000
2026-09-04,east,phone,12,18000

cut: pick columns#

Terminal
cut -d, -f2,3 sales.csv | head -3     # -d delimiter, -f fields
Output
region,product
north,laptop
south,phone

-f2-4 selects a range and -f3- means "field 3 to the end". cut -c1-7 cuts by character position instead, useful for fixed-width output. cut only accepts a single-character delimiter and can't handle runs of spaces; for space-aligned output like ps or df, use awk (next lesson) or squeeze the spaces first with tr -s ' '.

Skip a header line with tail -n +2, which you'll see in every example below.

sort: order lines#

Terminal
printf '10\n9\n100\n' | sort
printf '10\n9\n100\n' | sort -n
printf '2K\n1G\n500M\n' | sort -h
Output
10
100
9
9
10
100
2K
500M
1G
OptionMeaning
-nNumeric
-hHuman-readable sizes (du -h output)
-rReverse
-uUnique: drop duplicates
-t,Field separator (here a comma)
-k4,4Sort by field 4 only (start,end)
-fIgnore case
-VVersion order (v1.9 before v1.10)

Sort the sales by units, biggest first:

Terminal
tail -n +2 sales.csv | sort -t, -k4,4nr | head -2
Output
2026-09-04,east,phone,12,18000
2026-09-01,south,phone,10,18000

Several keys: by region, then units descending within each region:

Terminal
tail -n +2 sales.csv | sort -t, -k2,2 -k4,4nr | cut -d, -f2,4
Output
east,12
east,4
north,7
north,5
north,3
south,10
south,2

Always give the end field (-k4,4, not just -k4), otherwise the key runs to the end of the line. Note that the sort order of text depends on your locale (LANG); LC_ALL=C sort gives plain byte order and is faster on huge files.

uniq: collapse duplicates#

uniq removes adjacent duplicate lines, so you almost always sort first:

Terminal
tail -n +2 sales.csv | cut -d, -f2 | sort | uniq -c
Output
      2 east
      3 north
      2 south
OptionMeaning
-cPrefix each line with its count
-dOnly print lines that are duplicated
-uOnly print lines that appear once
-iIgnore case

"What's most common?" is sort | uniq -c | sort -rn | head, one of the most useful pipelines you'll ever learn. It works on IP addresses in logs, error messages, HTTP status codes, words in a text, and anything else.

tr: translate or delete characters#

tr works on characters, and reads only from stdin:

Terminal
echo "Hello World" | tr 'a-z' 'A-Z'           # upper-case
echo "too    many   spaces" | tr -s ' '       # squeeze repeats
echo "phone: +91 98765-43210" | tr -cd '0-9\n'  # delete everything except digits (and newline)
echo "a,b,,c" | tr ',' '\n'                   # one item per line
Output
HELLO WORLD
too many spaces
919876543210
a
b

c

-d deletes characters, -s squeezes runs, and -c complements the set ("everything except"). A common use is stripping Windows line endings: tr -d '\r' < win.txt > unix.txt.

wc, nl, paste#

Terminal
wc -l sales.csv               # lines (here: 1 header + 7 rows)
Output
8 sales.csv

paste joins files side by side, line by line:

Terminal
printf 'ada\nben\ncy\n' > names
printf '31\n27\n45\n' > ages
paste -d, names ages
Output
ada,31
ben,27
cy,45

paste -s joins all lines of one file into a single line, which makes a quick sum possible:

Terminal
paste -sd+ ages                # 31+27+45
echo $(( $(paste -sd+ ages) )) # let Bash do the arithmetic
Output
31+27+45
103

Comparing files: comm and diff#

Terminal
printf 'apple\nbanana\ncherry\n' > a.txt
printf 'banana\ncherry\ndate\n' > b.txt
comm -12 a.txt b.txt          # lines in BOTH (suppress columns 1 and 2)
Output
banana
cherry

comm needs sorted input and prints three columns: only in the first file, only in the second, in both. -23 gives "only in a.txt", which is great for "which users are in list A but not list B?".

diff shows how to turn one file into the other:

Terminal
diff a.txt b.txt
Output
1d0
< apple
3a3
> date

< lines are from the first file and > lines from the second. The unified format is easier to read and is what Git uses:

Terminal
diff -u a.txt b.txt
Output
--- a.txt	2026-10-01 10:30:00.000000000 +0530
+++ b.txt	2026-10-01 10:30:00.000000000 +0530
@@ -1,3 +1,3 @@
-apple
 banana
 cherry
+date

(The two header lines show the file names and modification times, so yours will differ.) diff exits with status 0 if the files are identical and 1 if they differ, so diff -q old.conf new.conf || echo "config changed" works in scripts. diff -r dir1 dir2 compares directories.

Putting it together#

Which product sold the most units overall? That needs a sum per group, which is where these tools reach their limit and awk shines. But plenty of questions fall to a pipeline:

Terminal
# How many distinct products?
tail -n +2 sales.csv | cut -d, -f3 | sort -u | wc -l

# Top 5 most frequent words in a text file
tr -cs '[:alpha:]' '\n' < book.txt | tr 'A-Z' 'a-z' | sort | uniq -c | sort -rn | head -5

# The 10 biggest items in the current directory
du -sh * | sort -rh | head -10

Build pipelines one stage at a time: run the first command, check the output, add | next-command, check again.

Common mistakes#

  • uniq without sort.
  • Text sort on numbers (sort instead of sort -n).
  • sort -k2 without an end field, or forgetting -t, for CSV.
  • Expecting cut to handle multiple spaces or quoted CSV fields with commas inside. For real-world CSV, use a proper tool (python3 -c "import csv...", csvkit, or mlr).
  • Using comm on unsorted files.

What's next#

sed and awk are the power tools of text processing: search-and-replace across files, and a tiny programming language for columns, sums and reports.

Check your understanding

Quick quiz

0/3 answered
  1. 1.Why does cut -d, -f2 data.csv | uniq -c often give wrong counts?

  2. 2.printf '10\n9\n100\n' | sort prints 10, 100, 9. How do you sort them as numbers?

  3. 3.What does tr -s ' ' do to the text too many spaces?

Finished reading?

Mark this lesson complete to track your progress.