Skip to content
elephantoo

Iterators & generators

Lesson 24 of 38 15 min read

The iterator protocol, generator functions and yield, lazy pipelines, infinite sequences and itertools.


Every for loop you've written relies on a simple protocol: ask an object for an iterator, then repeatedly ask the iterator for the next value until it says "done". Understanding this protocol lets you write generators — functions that produce values lazily, one at a time. Generators let you process huge files, infinite sequences and data pipelines using almost no memory.

Iterables and iterators#

An iterable is anything you can loop over: lists, strings, dicts, files, ranges. An iterator is the object that actually does the walking, remembering where it is. The built-in iter() gets an iterator from an iterable, and next() pulls the next value:

Python
colors = ["red", "green", "blue"]
it = iter(colors)

print(next(it))
print(next(it))
print(next(it))
try:
    next(it)
except StopIteration:
    print("exhausted")
Output
red
green
blue
exhausted

When an iterator runs out, it raises StopIteration. A for loop is essentially:

Python
colors = ["red", "green", "blue"]
it = iter(colors)
while True:
    try:
        color = next(it)
    except StopIteration:
        break
    print(color.upper())
Output
RED
GREEN
BLUE

Key consequences:

  • A list can be looped over many times, because each for gets a fresh iterator.
  • An iterator is single-use. Once consumed, it's empty — this is why a zip, map, file object or generator can only be looped over once.
Python
squares = map(lambda n: n * n, [1, 2, 3])   # map returns an iterator
print(list(squares))
print(list(squares))                         # already consumed
Output
[1, 4, 9]
[]

Writing an iterator class#

To make your own iterator, implement __iter__ (return the iterator, usually self) and __next__ (return the next value or raise StopIteration):

Python
class Countdown:
    def __init__(self, start):
        self.current = start

    def __iter__(self):
        return self

    def __next__(self):
        if self.current <= 0:
            raise StopIteration
        value = self.current
        self.current -= 1
        return value


for n in Countdown(3):
    print(n)
Output
3
2
1

It works, but that's a lot of boilerplate for "count down". Generators do the same in a few lines.

Generator functions#

A function that contains yield is a generator function. Calling it doesn't run the body — it returns a generator (an iterator). Each next() runs the body until the next yield, hands back that value, and pauses, keeping all local variables intact until the next request:

Python
def countdown(start):
    print("starting")
    while start > 0:
        yield start
        start -= 1
    print("done")


gen = countdown(3)
print(gen)
print(next(gen))
print(next(gen))
for n in gen:          # continues where it left off
    print(n)
Output
<generator object countdown at 0x...>
starting
3
2
1
done

Note that "starting" was printed only when we first called next — generators are lazy.

Why laziness matters#

A list holds every value in memory at once. A generator holds only its current state:

Python
import sys

squares_list = [n * n for n in range(1_000_000)]
squares_gen = (n * n for n in range(1_000_000))

print(sys.getsizeof(squares_list) > 8_000_000)   # megabytes
print(sys.getsizeof(squares_gen) < 500)          # a couple of hundred bytes
print(sum(squares_gen))
Output
True
True
333332833333500000

Laziness also enables infinite sequences — you simply stop asking for values when you have enough:

Python
from itertools import islice


def fibonacci():
    a, b = 0, 1
    while True:            # infinite, but harmless: values are produced on demand
        yield a
        a, b = b, a + b


print(list(islice(fibonacci(), 10)))
print(next(n for n in fibonacci() if n > 1000))
Output
[0, 1, 1, 2, 3, 5, 8, 13, 21, 34]
1597

Generator pipelines#

Generators chain together beautifully. Each stage pulls items from the previous one, so data flows through one item at a time — like Unix pipes:

Python
from pathlib import Path

Path("access.log").write_text(
    "200 GET /home 12ms\n"
    "404 GET /missing 3ms\n"
    "200 POST /login 85ms\n"
    "500 GET /report 1200ms\n"
    "200 GET /home 9ms\n",
    encoding="utf-8",
)


def read_lines(path):
    with open(path, encoding="utf-8") as f:
        for line in f:
            yield line.rstrip("\n")


def parse(lines):
    for line in lines:
        status, method, path, ms = line.split()
        yield {"status": int(status), "path": path, "ms": int(ms.removesuffix("ms"))}


def slow(requests, threshold):
    return (r for r in requests if r["ms"] > threshold)


for r in slow(parse(read_lines("access.log")), threshold=50):
    print(r["status"], r["path"], f"{r['ms']}ms")
Output
200 /login 85ms
500 /report 1200ms

This pipeline would work identically on a 50 GB log file, because no stage ever holds more than one line.

yield from#

yield from delegates to another iterable — handy for flattening and recursion:

Python
def flatten(items):
    for item in items:
        if isinstance(item, list):
            yield from flatten(item)     # recurse into nested lists
        else:
            yield item


print(list(flatten([1, [2, [3, 4]], 5, [[6]]])))
Output
[1, 2, 3, 4, 5, 6]

The itertools toolbox#

The standard library's itertools module provides fast, lazy building blocks for working with iterators:

Python
from itertools import count, cycle, islice, chain, batched, pairwise, accumulate, groupby

print(list(islice(count(10, 5), 4)))           # 10, 15, 20, 25 ...
print(list(islice(cycle("AB"), 5)))            # repeat forever
print(list(chain([1, 2], (3, 4), "ab")))       # concatenate iterables
print(list(batched(range(7), 3)))              # chunks (Python 3.12+)
print(list(pairwise([1, 4, 9, 16])))           # overlapping pairs
print(list(accumulate([5, 10, 20])))           # running totals

words = ["apple", "avocado", "banana", "blueberry", "cherry"]
for letter, group in groupby(words, key=lambda w: w[0]):
    print(letter, list(group))
Output
[10, 15, 20, 25]
['A', 'B', 'A', 'B', 'A']
[1, 2, 3, 4, 'a', 'b']
[(0, 1, 2), (3, 4, 5), (6,)]
[(1, 4), (4, 9), (9, 16)]
[5, 15, 35]
a ['apple', 'avocado']
b ['banana', 'blueberry']
c ['cherry']

Note that groupby only groups consecutive items, so sort by the same key first if needed. You'll see more of itertools in the standard library tour.

Generators with send and return (briefly)#

Generators can also receive values via gen.send(value) and return a final value via return (which becomes the StopIteration.value). These features were the foundation of Python's coroutines, but in modern code you'll use async/await (covered in the concurrency lesson) instead. For everyday work, yield is all you need.

Common mistakes#

  • Iterating a generator twice — the second loop gets nothing. Recreate it, or store results in a list if you truly need them twice.
  • Calling len() on a generator — generators don't know their length. Use sum(1 for _ in gen) (which consumes it) or a list.
  • Forgetting to call the generator function: for x in countdown: fails; write countdown(3).
  • Expecting work to happen eagerly — a generator does nothing until someone iterates over it.
  • Using return value to emit items — use yield.

What's next#

Generators are functions with superpowers. Decorators — functions that wrap other functions — are the next superpower, and they rely on the closures you learned earlier.

Check your understanding

Quick quiz

0/3 answered
  1. 1.What happens when you call a generator function (one containing yield)?

  2. 2.What is the difference between an iterable and an iterator?

  3. 3.What does yield from other_generator() do?

Finished reading?

Mark this lesson complete to track your progress.