Regular expressions
The re module: patterns, quantifiers, groups and named groups, greedy vs lazy, substitution and validation.
A regular expression (regex) is a tiny pattern language for describing text: "a word followed by an @ and a domain", "five digits, a dash, five digits", "any line starting with ERROR". With Python's built-in re module you can search, validate, extract and replace text in ways that would take dozens of lines of string methods. Regexes appear everywhere — log analysis, data cleaning, form validation, web scraping, editors' search-and-replace — so they're a genuinely job-ready skill.
A first example#
Patterns are written as raw strings (r"...") so that backslashes go straight to the regex engine instead of being interpreted by Python.
The pattern language#
Characters and classes
Quantifiers
Anchors and alternation
The main functions#
search and match return a match object or None, so always check before calling .group():
Groups: extracting parts#
Parentheses capture parts of the match. Named groups (?P<name>...) make code far more readable:
With multiple groups, findall returns tuples:
Greedy vs lazy#
Quantifiers are greedy by default — they match as much as possible. Add ? to make them lazy:
Often the best fix is a more specific pattern ([^>]+ or \w+) rather than a lazy one. (And for real HTML, use a proper parser such as BeautifulSoup — regex isn't designed for nested structures.)
Substitution with groups and functions#
re.sub can reference captured groups with \1 or \g<name>, or call a function for each match:
Flags and compiled patterns#
Flags change how a pattern behaves: re.IGNORECASE (re.I), re.MULTILINE (re.M, ^/$ per line), re.DOTALL (re.S, . matches newlines) and re.VERBOSE (re.X, allows whitespace and comments). For patterns you use repeatedly, re.compile creates a reusable pattern object — and verbose mode makes complex patterns maintainable:
Validation: use fullmatch#
When validating input, the whole string must match — otherwise "abc123xyz" would "contain" a valid number. fullmatch (or anchoring with ^...$) ensures that:
Be pragmatic: perfect email validation by regex is famously impossible. A simple check like [^@\s]+@[^@\s]+\.[^@\s]+ plus sending a confirmation email is what real systems do.
Worked example: summarising a log#
Common mistakes#
- Forgetting the
rprefix —"\b"is a backspace, not a word boundary. - Using
matchwhen you meantsearch(orsearchwhen validating — usefullmatch). - Not escaping special characters like
.,?,(,$. Usere.escape(user_text)when building patterns from input. - Greedy
.*swallowing too much. - Writing unreadable one-liners — use
re.VERBOSEand named groups. - Using regex when a string method would do:
s.startswith("http")beatsre.match(r"http", s).
What's next#
Text processing done. Next we'll talk to the outside world over the network: HTTP requests and web APIs with the requests library.
Check your understanding
Quick quiz
1.What is the difference between
re.match()andre.search()?2.Why are regex patterns usually written as raw strings, like
r"\d+"?3.Given
re.findall(r"<.+?>", "<b>hi</b>"), what is returned?
Finished reading?
Mark this lesson complete to track your progress.