Skip to content
elephantoo

Concurrency: threads, processes & asyncio

Lesson 33 of 38 20 min read

I/O-bound vs CPU-bound work, thread and process pools, the GIL, locks, and async/await with asyncio.


By default, a Python program does one thing at a time. That's wasteful when your code spends most of its time waiting — for web servers, databases or disks — or when you have a heavy computation and a CPU with eight idle cores. Concurrency lets a program make progress on several tasks at once. Python offers three main tools: threads, processes and asyncio. The trick is knowing which one fits your problem.

I/O-bound vs CPU-bound#

Kind of workExampleBottleneckBest tool
I/O-bounddownloading pages, calling APIs, querying databaseswaitingthreads or asyncio
CPU-boundresizing images, crunching numbers, parsing huge filesthe processorprocesses

Let's simulate an I/O-bound task — "download" that takes one second — and run it sequentially:

Python
import time


def download(n):
    time.sleep(1)               # stand-in for waiting on the network
    return f"page {n}"


start = time.perf_counter()
results = [download(n) for n in range(5)]
print(results)
print(f"sequential: {time.perf_counter() - start:.1f}s")
Output
['page 0', 'page 1', 'page 2', 'page 3', 'page 4']
sequential: 5.0s

Five seconds, almost all of it spent waiting. Let's fix that.

Threads with concurrent.futures#

A thread is a separate flow of execution inside the same process, sharing the same memory. The easiest, safest way to use threads is a thread pool from concurrent.futures:

Python
import time
from concurrent.futures import ThreadPoolExecutor


def download(n):
    time.sleep(1)
    return f"page {n}"


start = time.perf_counter()
with ThreadPoolExecutor(max_workers=5) as pool:
    results = list(pool.map(download, range(5)))
print(results)
print(f"threaded: {time.perf_counter() - start:.1f}s")
Output
['page 0', 'page 1', 'page 2', 'page 3', 'page 4']
threaded: 1.0s

Five times faster: while one thread sleeps (waits), the others run. pool.map returns results in input order. When you want to handle results as soon as each finishes — and deal with errors per task — use submit and as_completed:

Python
import time
from concurrent.futures import ThreadPoolExecutor, as_completed


def fetch(url):
    time.sleep(len(url) / 20)            # pretend longer URLs are slower
    if "bad" in url:
        raise ConnectionError(f"failed: {url}")
    return url.upper()


urls = ["a.com", "bad.example", "python.org", "elephantoo.com"]
with ThreadPoolExecutor(max_workers=4) as pool:
    futures = {pool.submit(fetch, u): u for u in urls}
    for future in as_completed(futures):
        url = futures[future]
        try:
            print("ok   ", future.result())
        except ConnectionError as e:
            print("error", e)
Output
ok    A.COM
ok    PYTHON.ORG
error failed: bad.example
ok    ELEPHANTOO.COM

In real code, fetch would use requests.get(url, timeout=10) — requests releases the GIL while waiting on the network, so threads work well with it.

The GIL and race conditions#

Standard CPython has a Global Interpreter Lock (GIL): only one thread executes Python bytecode at any instant. Threads still help with I/O, because waiting threads release the GIL. But pure-Python CPU work doesn't get faster with threads. (Python 3.13 introduced an experimental free-threaded build without the GIL; it's not yet the default.)

Because threads share memory, two threads updating the same data can interfere — a race condition. counter += 1 is really read → add → write, and a thread switch in between loses updates. Protect shared state with a Lock:

Python
import threading

counter = 0
lock = threading.Lock()


def work():
    global counter
    for _ in range(100_000):
        with lock:                    # only one thread at a time in here
            counter += 1


threads = [threading.Thread(target=work) for _ in range(4)]
for t in threads:
    t.start()
for t in threads:
    t.join()                          # wait for each thread to finish
print(counter)
Output
400000

Better still, avoid shared mutable state: have each task return its result (as the pool examples do) and combine results in the main thread. For producer/consumer pipelines, queue.Queue is a thread-safe way to pass work between threads.

Processes for CPU-bound work#

For heavy computation, use processes. Each process has its own Python interpreter and its own GIL, so they truly run in parallel on multiple cores. ProcessPoolExecutor has the same interface as the thread pool:

primes.py
import time
from concurrent.futures import ProcessPoolExecutor


def count_primes(limit):
    count = 0
    for n in range(2, limit):
        if all(n % d for d in range(2, int(n ** 0.5) + 1)):
            count += 1
    return count


if __name__ == "__main__":                 # required for multiprocessing!
    jobs = [150_000] * 4

    start = time.perf_counter()
    serial = [count_primes(j) for j in jobs]
    serial_time = time.perf_counter() - start

    start = time.perf_counter()
    with ProcessPoolExecutor() as pool:
        parallel = list(pool.map(count_primes, jobs))
    parallel_time = time.perf_counter() - start

    print(serial == parallel, parallel[0])
    print(f"speed-up: {serial_time / parallel_time:.1f}x")
Output
True 13848
speed-up: 3.7x

Your numbers will differ: with four equal jobs the best possible speed-up is 4x, and you need at least four CPU cores to get close to it. Important rules for processes:

  • Always guard the start-up code with if __name__ == "__main__":. On Windows and macOS, child processes re-import your module, and without the guard they'd spawn processes endlessly.
  • Arguments and results are pickled (serialised) and copied between processes, so send compact data, not giant objects. Functions must be defined at module top level.
  • Starting processes is far more expensive than starting threads. It only pays off for substantial work.

The lower-level multiprocessing module (Process, Pool, Queue, shared memory) offers more control, but ProcessPoolExecutor covers most needs. For numeric work, libraries like NumPy already run optimised C code that sidesteps the GIL.

asyncio: thousands of concurrent waits#

asyncio handles concurrency on a single thread using an event loop. You write coroutines with async def, and at every await a coroutine voluntarily pauses so others can run. It shines when you have very many simultaneous I/O operations — thousands of network connections — and it powers modern frameworks like FastAPI.

Python
import asyncio
import time


async def download(n):
    await asyncio.sleep(1)          # non-blocking wait: other coroutines run meanwhile
    return f"page {n}"


async def main():
    start = time.perf_counter()
    results = await asyncio.gather(*(download(n) for n in range(5)))
    print(results)
    print(f"asyncio: {time.perf_counter() - start:.1f}s")


asyncio.run(main())
Output
['page 0', 'page 1', 'page 2', 'page 3', 'page 4']
asyncio: 1.0s

Key concepts:

  • async def defines a coroutine function. Calling it returns a coroutine object; nothing runs until it's awaited or scheduled.
  • await pauses the current coroutine until the awaited operation completes, letting the event loop run others.
  • asyncio.run(main()) starts the event loop — call it once, at your program's entry point.
  • asyncio.gather(...) runs several coroutines concurrently and returns their results in order.

TaskGroup, timeouts and limiting concurrency

Python 3.11 added asyncio.TaskGroup, the modern structured way to run tasks: if one fails, the others are cancelled and the errors are raised together (as an ExceptionGroup). asyncio.timeout() bounds how long something may take, and a Semaphore limits how many operations run at once — essential so you don't open 10,000 connections to one server:

Python
import asyncio


async def fetch(i, sem):
    async with sem:                         # at most 3 at a time
        await asyncio.sleep(0.2)
        return i * i


async def slow():
    await asyncio.sleep(5)


async def main():
    sem = asyncio.Semaphore(3)
    async with asyncio.TaskGroup() as tg:
        tasks = [tg.create_task(fetch(i, sem)) for i in range(6)]
    print([t.result() for t in tasks])

    try:
        async with asyncio.timeout(0.5):
            await slow()
    except TimeoutError:
        print("gave up after 0.5s")


asyncio.run(main())
Output
[0, 1, 4, 9, 16, 25]
gave up after 0.5s

Real async HTTP with httpx

requests is synchronous, so for asyncio you use an async client such as httpx.AsyncClient (or aiohttp):

Python
import asyncio

import httpx


async def get_user(client, user_id):
    r = await client.get(f"https://jsonplaceholder.typicode.com/users/{user_id}")
    r.raise_for_status()
    return r.json()["name"]


async def main():
    async with httpx.AsyncClient(timeout=10) as client:
        names = await asyncio.gather(*(get_user(client, i) for i in range(1, 6)))
    print(names)


asyncio.run(main())
Output
['Leanne Graham', 'Ervin Howell', 'Clementine Bauch', 'Patricia Lebsack', 'Chelsey Dietrich']

The golden rule: never block the event loop

Inside a coroutine, a blocking call such as time.sleep(), requests.get() or a heavy calculation freezes every coroutine, because they all share one thread. Use async libraries (asyncio.sleep, httpx, asyncpg, aiomysql), and push unavoidable blocking work to a thread with await asyncio.to_thread(func, *args):

Python
import asyncio
import time


def legacy_blocking_call(n):
    time.sleep(0.5)                 # e.g. a library with no async version
    return n * 10


async def main():
    start = time.perf_counter()
    results = await asyncio.gather(*(asyncio.to_thread(legacy_blocking_call, n) for n in range(4)))
    print(results, f"{time.perf_counter() - start:.1f}s")


asyncio.run(main())
Output
[0, 10, 20, 30] 0.5s

Choosing the right tool#

  • A few dozen I/O tasks, existing synchronous code (requests, database drivers): ThreadPoolExecutor. Simple and effective.
  • Thousands of concurrent connections, or an async framework (FastAPI, websockets): asyncio.
  • CPU-heavy pure-Python work: ProcessPoolExecutor — or vectorise with NumPy/pandas first.
  • Background jobs in web apps (sending emails, generating reports): a task queue such as Celery, RQ or a cloud queue — not ad-hoc threads.

And remember: concurrency adds complexity. Measure first, and only reach for it when waiting or computation is a real bottleneck.

Common mistakes#

  • Using threads for CPU-bound Python code and expecting a speed-up.
  • Forgetting if __name__ == "__main__": with multiprocessing.
  • Calling blocking functions inside async def (time.sleep, requests.get).
  • Forgetting to await a coroutine — you get a "coroutine was never awaited" warning and nothing happens.
  • Sharing mutable state between threads without a lock.
  • Unbounded concurrency — limit with max_workers or a semaphore so you don't overwhelm servers or get rate-limited.

What's next#

Most applications store their data in a database. Next: databases with sqlite3 and MySQL.

Check your understanding

Quick quiz

0/3 answered
  1. 1.You need to download 200 web pages as fast as possible. Which approach fits best?

  2. 2.Why don't threads speed up CPU-bound pure-Python code in standard CPython?

  3. 3.What does await asyncio.gather(a(), b(), c()) do?

Finished reading?

Mark this lesson complete to track your progress.