Iterators & generators
The iterator protocol, generator functions and yield, lazy pipelines, infinite sequences and itertools.
Every for loop you've written relies on a simple protocol: ask an object for an iterator, then repeatedly ask the iterator for the next value until it says "done". Understanding this protocol lets you write generators — functions that produce values lazily, one at a time. Generators let you process huge files, infinite sequences and data pipelines using almost no memory.
Iterables and iterators#
An iterable is anything you can loop over: lists, strings, dicts, files, ranges. An iterator is the object that actually does the walking, remembering where it is. The built-in iter() gets an iterator from an iterable, and next() pulls the next value:
When an iterator runs out, it raises StopIteration. A for loop is essentially:
Key consequences:
- A list can be looped over many times, because each
forgets a fresh iterator. - An iterator is single-use. Once consumed, it's empty — this is why a
zip,map, file object or generator can only be looped over once.
Writing an iterator class#
To make your own iterator, implement __iter__ (return the iterator, usually self) and __next__ (return the next value or raise StopIteration):
It works, but that's a lot of boilerplate for "count down". Generators do the same in a few lines.
Generator functions#
A function that contains yield is a generator function. Calling it doesn't run the body — it returns a generator (an iterator). Each next() runs the body until the next yield, hands back that value, and pauses, keeping all local variables intact until the next request:
Note that "starting" was printed only when we first called next — generators are lazy.
Why laziness matters#
A list holds every value in memory at once. A generator holds only its current state:
Laziness also enables infinite sequences — you simply stop asking for values when you have enough:
Generator pipelines#
Generators chain together beautifully. Each stage pulls items from the previous one, so data flows through one item at a time — like Unix pipes:
This pipeline would work identically on a 50 GB log file, because no stage ever holds more than one line.
yield from#
yield from delegates to another iterable — handy for flattening and recursion:
The itertools toolbox#
The standard library's itertools module provides fast, lazy building blocks for working with iterators:
Note that groupby only groups consecutive items, so sort by the same key first if needed. You'll see more of itertools in the standard library tour.
Generators with send and return (briefly)#
Generators can also receive values via gen.send(value) and return a final value via return (which becomes the StopIteration.value). These features were the foundation of Python's coroutines, but in modern code you'll use async/await (covered in the concurrency lesson) instead. For everyday work, yield is all you need.
Common mistakes#
- Iterating a generator twice — the second loop gets nothing. Recreate it, or store results in a list if you truly need them twice.
- Calling
len()on a generator — generators don't know their length. Usesum(1 for _ in gen)(which consumes it) or a list. - Forgetting to call the generator function:
for x in countdown:fails; writecountdown(3). - Expecting work to happen eagerly — a generator does nothing until someone iterates over it.
- Using
return valueto emit items — useyield.
What's next#
Generators are functions with superpowers. Decorators — functions that wrap other functions — are the next superpower, and they rely on the closures you learned earlier.
Check your understanding
Quick quiz
1.What happens when you call a generator function (one containing
yield)?2.What is the difference between an iterable and an iterator?
3.What does
yield from other_generator()do?
Finished reading?
Mark this lesson complete to track your progress.