Skip to content
elephantoo

Regular expressions

Lesson 29 of 38 16 min read

The re module: patterns, quantifiers, groups and named groups, greedy vs lazy, substitution and validation.


A regular expression (regex) is a tiny pattern language for describing text: "a word followed by an @ and a domain", "five digits, a dash, five digits", "any line starting with ERROR". With Python's built-in re module you can search, validate, extract and replace text in ways that would take dozens of lines of string methods. Regexes appear everywhere — log analysis, data cleaning, form validation, web scraping, editors' search-and-replace — so they're a genuinely job-ready skill.

A first example#

Python
import re

text = "Order #4521 shipped on 2026-09-30, order #4522 pending."

match = re.search(r"#(\d+)", text)
print(match.group(), match.group(1), match.start())

print(re.findall(r"#(\d+)", text))
print(re.findall(r"\d{4}-\d{2}-\d{2}", text))
Output
#4521 4521 6
['4521', '4522']
['2026-09-30']

Patterns are written as raw strings (r"...") so that backslashes go straight to the regex engine instead of being interpreted by Python.

The pattern language#

Characters and classes

PatternMatches
abcthe literal text "abc"
.any character except newline
\d / \Da digit / a non-digit
\w / \Wa word character (letter, digit, _) / anything else
\s / \Swhitespace / non-whitespace
[aeiou]any one of these characters
[a-z0-9]a range
[^0-9]anything except these
\. \$ \(a literal special character (escaped)

Quantifiers

PatternMeaning
*0 or more
+1 or more
?0 or 1 (optional)
{3}exactly 3
{2,5}between 2 and 5
{2,}2 or more

Anchors and alternation

PatternMeaning
^ / $start / end of the string (or line, with re.MULTILINE)
\bword boundary
a|ba or b
(...)group (capture)
(?:...)group without capturing
Python
import re

print(re.findall(r"\b\w+ing\b", "Singing and dancing, then sleeping in."))
print(re.findall(r"colou?r", "color colour colr"))
print(re.findall(r"\b[A-Z]{2,}\b", "The API uses JSON over HTTP, not xml."))
print(re.findall(r"^\w+", "first line\nsecond line", flags=re.MULTILINE))
print(re.findall(r"cat|dog", "hotdog, catalog, cow"))
Output
['Singing', 'dancing', 'sleeping']
['color', 'colour']
['API', 'JSON', 'HTTP']
['first', 'second']
['dog', 'cat']

The main functions#

Python
import re

s = "Python 3.12 and Python 3.13"

print(re.match(r"Python", s) is not None)     # only at the START
print(re.match(r"3\.12", s))                   # not at the start → None
print(re.search(r"3\.1\d", s).group())         # first match anywhere
print(re.fullmatch(r"\d+", "2026") is not None)  # whole string must match
print(re.findall(r"3\.1\d", s))                # all matches as strings
print([m.span() for m in re.finditer(r"Python", s)])   # match objects lazily
print(re.split(r"\s*[,;]\s*", "a , b;c ,d"))   # split on a pattern
print(re.sub(r"\d", "#", "PIN: 4821"))          # replace
Output
True
None
3.12
True
['3.12', '3.13']
[(0, 6), (16, 22)]
['a', 'b', 'c', 'd']
PIN: ####

search and match return a match object or None, so always check before calling .group():

Python
import re

if m := re.search(r"(\d+) items", "Cart: 3 items"):
    print(int(m.group(1)) * 2)
Output
6

Groups: extracting parts#

Parentheses capture parts of the match. Named groups (?P<name>...) make code far more readable:

Python
import re

log = "2026-09-30 22:14:05 ERROR [payments] Card declined for order 9912"

pattern = r"(?P<date>\S+) (?P<time>\S+) (?P<level>[A-Z]+) \[(?P<module>\w+)\] (?P<msg>.*)"
m = re.match(pattern, log)

print(m.group("level"), m.group("module"))
print(m.groupdict())
print(m.groups()[:2])
Output
ERROR payments
{'date': '2026-09-30', 'time': '22:14:05', 'level': 'ERROR', 'module': 'payments', 'msg': 'Card declined for order 9912'}
('2026-09-30', '22:14:05')

With multiple groups, findall returns tuples:

Python
import re

text = "ada@example.com, grace.hopper@navy.mil; bad@, linus@kernel.org"
for user, domain in re.findall(r"([\w.+-]+)@([\w-]+\.[\w.]+)", text):
    print(f"{user:<13} at {domain}")
Output
ada           at example.com
grace.hopper  at navy.mil
linus         at kernel.org

Greedy vs lazy#

Quantifiers are greedy by default — they match as much as possible. Add ? to make them lazy:

Python
import re

html = "<b>bold</b> and <i>italic</i>"
print(re.findall(r"<.+>", html))     # greedy: one giant match
print(re.findall(r"<.+?>", html))    # lazy: each tag
print(re.findall(r"<(\w+)>", html))  # better: be specific
Output
['<b>bold</b> and <i>italic</i>']
['<b>', '</b>', '<i>', '</i>']
['b', 'i']

Often the best fix is a more specific pattern ([^>]+ or \w+) rather than a lazy one. (And for real HTML, use a proper parser such as BeautifulSoup — regex isn't designed for nested structures.)

Substitution with groups and functions#

re.sub can reference captured groups with \1 or \g<name>, or call a function for each match:

Python
import re

# reformat dates from DD/MM/YYYY to ISO YYYY-MM-DD
print(re.sub(r"(\d{2})/(\d{2})/(\d{4})", r"\3-\2-\1", "Due 30/09/2026, paid 02/10/2026"))

# mask all but the last 4 digits of card numbers
print(re.sub(r"\b\d{12}(\d{4})\b", r"************\1", "Card 4111111111111111 charged"))

# use a function to compute the replacement
prices = "Mouse ₹799, Keyboard ₹2499"
print(re.sub(r"₹(\d+)", lambda m: f"₹{int(m.group(1)) * 1.18:.0f}", prices))

# collapse repeated whitespace
print(re.sub(r"\s+", " ", "  too   many\n\nspaces  ").strip())
Output
Due 2026-09-30, paid 2026-10-02
Card ************1111 charged
Mouse ₹943, Keyboard ₹2949
too many spaces

Flags and compiled patterns#

Flags change how a pattern behaves: re.IGNORECASE (re.I), re.MULTILINE (re.M, ^/$ per line), re.DOTALL (re.S, . matches newlines) and re.VERBOSE (re.X, allows whitespace and comments). For patterns you use repeatedly, re.compile creates a reusable pattern object — and verbose mode makes complex patterns maintainable:

Python
import re

PHONE = re.compile(r"""
    (?:\+91[\s-]?)?     # optional country code
    ([6-9]\d{4})        # first five digits (Indian mobiles start with 6-9)
    [\s-]?              # optional separator
    (\d{5})             # last five digits
""", re.VERBOSE)

for raw in ["+91 98765 43210", "9876543210", "98765-43210", "12345 67890"]:
    m = PHONE.fullmatch(raw)
    print(f"{raw:<16} ->", f"{m.group(1)}{m.group(2)}" if m else "invalid")

print(re.findall(r"python", "Python PYTHON python", flags=re.I))
Output
+91 98765 43210  -> 9876543210
9876543210       -> 9876543210
98765-43210      -> 9876543210
12345 67890      -> invalid
['Python', 'PYTHON', 'python']

Validation: use fullmatch#

When validating input, the whole string must match — otherwise "abc123xyz" would "contain" a valid number. fullmatch (or anchoring with ^...$) ensures that:

Python
import re

PINCODE = re.compile(r"[1-9]\d{5}")      # Indian PIN code: 6 digits, no leading 0
for code in ["560001", "060001", "5600012", "56000a"]:
    print(code, bool(PINCODE.fullmatch(code)))
Output
560001 True
060001 False
5600012 False
56000a False

Be pragmatic: perfect email validation by regex is famously impossible. A simple check like [^@\s]+@[^@\s]+\.[^@\s]+ plus sending a confirmation email is what real systems do.

Worked example: summarising a log#

Python
import re
from collections import Counter

log = """\
2026-09-30 10:00:01 INFO  GET /home 200 12ms
2026-09-30 10:00:03 INFO  GET /api/users 200 48ms
2026-09-30 10:00:04 WARN  GET /api/orders 200 950ms
2026-09-30 10:00:09 ERROR POST /api/pay 500 1203ms
2026-09-30 10:00:12 INFO  GET /home 304 3ms
"""

LINE = re.compile(r"(?P<level>[A-Z]+)\s+(?P<method>GET|POST) (?P<path>\S+) (?P<status>\d{3}) (?P<ms>\d+)ms")

entries = [m.groupdict() for m in LINE.finditer(log)]
slow = [e["path"] for e in entries if int(e["ms"]) > 500]
statuses = Counter(e["status"] for e in entries)

print(len(entries), "requests")
print("slow:", slow)
print("status codes:", dict(statuses))
Output
5 requests
slow: ['/api/orders', '/api/pay']
status codes: {'200': 3, '500': 1, '304': 1}

Common mistakes#

  • Forgetting the r prefix — "\b" is a backspace, not a word boundary.
  • Using match when you meant search (or search when validating — use fullmatch).
  • Not escaping special characters like ., ?, (, $. Use re.escape(user_text) when building patterns from input.
  • Greedy .* swallowing too much.
  • Writing unreadable one-liners — use re.VERBOSE and named groups.
  • Using regex when a string method would do: s.startswith("http") beats re.match(r"http", s).

What's next#

Text processing done. Next we'll talk to the outside world over the network: HTTP requests and web APIs with the requests library.

Check your understanding

Quick quiz

0/3 answered
  1. 1.What is the difference between re.match() and re.search()?

  2. 2.Why are regex patterns usually written as raw strings, like r"\d+"?

  3. 3.Given re.findall(r"<.+?>", "<b>hi</b>"), what is returned?

Finished reading?

Mark this lesson complete to track your progress.