Skip to content
elephantoo

Comprehensions & generator expressions

Lesson 14 of 38 15 min read

List, dict and set comprehensions, filtering, nesting, generator expressions and when not to use them.


A comprehension builds a new list, dict or set from an existing iterable in a single, readable expression. They're one of Python's most loved features: shorter than the equivalent loop, often faster, and — once you can read them — very clear. This lesson covers all four kinds, plus when not to use them.

From loop to list comprehension#

Here's a familiar pattern: start with an empty list, loop, append.

Python
squares = []
for n in range(1, 6):
    squares.append(n * n)
print(squares)
Output
[1, 4, 9, 16, 25]

The comprehension says the same thing in one line:

Python
squares = [n * n for n in range(1, 6)]
print(squares)
Output
[1, 4, 9, 16, 25]

Read it as: "a list of n * n, for each n in range(1, 6)". The general shape is:

Syntax
[expression for item in iterable]

The expression can be anything — a method call, an f-string, a function call:

Python
names = ["ada", "grace", "linus"]
print([name.title() for name in names])
print([len(name) for name in names])
print([f"{i}:{name}" for i, name in enumerate(names, 1)])
Output
['Ada', 'Grace', 'Linus']
[3, 5, 5]
['1:ada', '2:grace', '3:linus']

Filtering with if#

Add an if at the end to keep only some items:

Python
nums = [12, 7, 3, 18, 25, 4]
evens = [n for n in nums if n % 2 == 0]
big_squares = [n * n for n in nums if n > 10]
print(evens)
print(big_squares)

files = ["app.py", "README.md", "test_app.py", "setup.cfg"]
print([f for f in files if f.endswith(".py")])
Output
[12, 18, 4]
[144, 324, 625]
['app.py', 'test_app.py']

Transforming conditionally with if … else#

To change every item depending on a condition (rather than drop some), put a conditional expression at the front:

Python
nums = [12, 7, 3, 18]
labels = ["even" if n % 2 == 0 else "odd" for n in nums]
clamped = [min(n, 10) for n in nums]
print(labels)
print(clamped)
Output
['even', 'odd', 'odd', 'even']
[10, 7, 3, 10]

Remember the rule: if after for filters; if … else before for transforms.

Nested loops in a comprehension#

Multiple for clauses nest like loops — left to right is outer to inner:

Python
pairs = [(color, size) for color in ["red", "blue"] for size in "SM"]
print(pairs)

matrix = [[1, 2, 3], [4, 5, 6]]
flat = [x for row in matrix for x in row]
print(flat)

transposed = [[row[i] for row in matrix] for i in range(3)]
print(transposed)
Output
[('red', 'S'), ('red', 'M'), ('blue', 'S'), ('blue', 'M')]
[1, 2, 3, 4, 5, 6]
[[1, 4], [2, 5], [3, 6]]

flat is equivalent to:

Python
flat = []
for row in [[1, 2, 3], [4, 5, 6]]:
    for x in row:
        flat.append(x)
print(flat)
Output
[1, 2, 3, 4, 5, 6]

Beyond two levels, comprehensions get hard to read — switch to regular loops.

Dict comprehensions#

Use braces with a key: value expression:

Python
words = ["apple", "banana", "cherry"]
lengths = {w: len(w) for w in words}
print(lengths)

prices = {"apple": 40, "banana": 10, "cherry": 120}
discounted = {item: round(p * 0.9) for item, p in prices.items() if p > 20}
print(discounted)

codes = ["IN", "US", "JP"]
names = ["India", "United States", "Japan"]
print({c: n for c, n in zip(codes, names)})
Output
{'apple': 5, 'banana': 6, 'cherry': 6}
{'apple': 36, 'cherry': 108}
{'IN': 'India', 'US': 'United States', 'JP': 'Japan'}

(For the last one, dict(zip(codes, names)) is even shorter.)

Set comprehensions#

Braces without a colon build a set — duplicates disappear automatically:

Python
emails = ["Ada@Example.com", "grace@example.com", "ada@example.com"]
domains = {e.split("@")[1].lower() for e in emails}
unique_users = {e.lower() for e in emails}
print(domains)
print(sorted(unique_users))
Output
{'example.com'}
['ada@example.com', 'grace@example.com']

Generator expressions#

Swap the brackets for parentheses and you get a generator expression. It doesn't build a list; it produces values one at a time, on demand:

Python
nums = range(1, 1_000_001)
total = sum(n * n for n in nums)     # no million-item list in memory
print(total)

gen = (n * n for n in range(3))
print(gen)
print(list(gen))
print(list(gen))                     # already used up!
Output
333333833333500000
<generator object <genexpr> at 0x...>
[0, 1, 4]
[]

When a generator expression is the only argument to a function, you can drop the extra parentheses, as in sum(n * n for n in nums). Use generator expressions with sum, min, max, any, all, "".join and anything else that consumes an iterable once. You'll learn how generators work under the hood in a later lesson.

any and all with a generator short-circuit, stopping as soon as the answer is known:

Python
passwords = ["hunter2", "correct horse battery staple", "abc"]
print(any(len(p) < 6 for p in passwords))
print(all(p.islower() for p in passwords))
Output
True
True

The walrus operator in comprehensions#

The assignment expression := (Python 3.8+) lets you compute a value once and both test and keep it:

Python
import math

values = [4, 10, 25, 30]
roots = [r for v in values if (r := math.sqrt(v)).is_integer()]
print(roots)
Output
[2.0, 5.0]

Comprehensions have their own scope#

The loop variable doesn't leak out of a comprehension:

Python
x = "outer"
squares = [x * x for x in range(3)]
print(x)
Output
outer

When not to use a comprehension#

Comprehensions are for building a collection. Don't use them:

  • For side effects — [print(x) for x in items] builds a useless list of Nones. Use a for loop.
  • When logic is complex — several conditions, nested if … else, or more than two for clauses. A loop with good names is clearer.
  • When you only need one result — use any, sum, next(...) with a generator.
Python
users = [{"name": "Ada", "admin": False}, {"name": "Grace", "admin": True}]
first_admin = next((u for u in users if u["admin"]), None)
print(first_admin)
Output
{'name': 'Grace', 'admin': True}

Worked example: cleaning survey data#

Python
raw = ["  Yes", "no ", "YES", "", "maybe", "No", "yes  "]

cleaned = [r.strip().lower() for r in raw if r.strip()]
counts = {answer: cleaned.count(answer) for answer in sorted(set(cleaned))}
yes_share = sum(a == "yes" for a in cleaned) / len(cleaned)

print(cleaned)
print(counts)
print(f"{yes_share:.0%} said yes")
Output
['yes', 'no', 'yes', 'maybe', 'no', 'yes']
{'maybe': 1, 'no': 2, 'yes': 3}
50% said yes

sum(a == "yes" for a in cleaned) works because True counts as 1 and False as 0.

Common mistakes#

  • Putting the filter if in front ([x if x > 0 for x in xs]) — that's a syntax error; a front if needs an else.
  • Expecting (x for x in xs) to be a tuple — it's a generator. Use tuple(x for x in xs).
  • Reusing an exhausted generator.
  • Cramming too much into one line — readability beats cleverness.

What's next#

Your programs are growing. Next we'll split code across files with modules and packages, and learn exactly how import finds things.

Check your understanding

Quick quiz

0/3 answered
  1. 1.What does [n * 2 for n in range(5) if n % 2 == 0] produce?

  2. 2.What is the key difference between [x*x for x in data] and (x*x for x in data)?

  3. 3.Where does the else go in a comprehension that transforms every item conditionally?

Finished reading?

Mark this lesson complete to track your progress.