You already know how to check if a word is in a string with
in, or split it up with .split(). Those tools are fast and simple. But when you need to find things that follow a pattern — like every 5-digit number, or all email addresses on a page — plain methods get clumsy.Regular expressions, called regex for short, let you describe patterns. The
Use regex when:
re module in the standard library does the matching work.Use regex when:
- You need to find a shape of text (digits followed by letters).
- You want every match at once (
findall). - You are replacing or validating based on structure.
in, .startswith(), or .split() already solves the problem.Your first pattern
import re
text = 'Order 4821 shipped'
match = re.search(r'\d+', text)
if match:
print(match.group()) # prints: 4821re.search(pattern, string) scans left to right and returns the first match. match.group() gives you that matched text.Note the
r'...' — a raw string so backslashes like \d are not eaten by Python's normal escape rules.Building blocks
# \d one digit (0-9)
# \w word char: letters, digits, underscore
# \s whitespace
#
# [A-Z] any single uppercase letter
# [aeiou] vowels only
#
# Quantifiers:
# + one or more
# * zero or more
# ? zero or one
# {2} exactly two
# {1,3} one to three
#
# Anchors:
# ^ the start of the string
# $ the end of the string# fullmatch checks that WHOLE string fits.
import re
time = '09:35'
ok = bool(re.fullmatch(r'\d{2}:\d{2}', time)) # True
bad = bool(re.fullmatch(r'\d{2}:\d{2}', '9x')) # FalseGroups and findall
import re
log = 'user=alice action=read user=bob action=write'
users = re.findall(r'user=(\w+)', log)
print(users) # ['alice', 'bob']
# findall returns a list of group(1) strings when there is one group.import re
m = re.search(r'(\d{2}):(\d{2})', 'Meet at 14:30 by the gate')
print(m.group(0)) # 14:30 the whole match
print(m.group(1)) # 14 the first group
print(m.group(2)) # 30 the second groupParentheses make a group.
match.group(0) is everything the pattern matched, and group(1), group(2) are the pieces inside each pair of parentheses, counted from the left.Replacing with re.sub
import re
msg = 'Call 555-1234 or 5550199 for help'
cleaned = re.sub(r'\d{3}-?\d{4}', '[REDACTED]', msg)
print(cleaned) # Call [REDACTED] or [REDACTED] for help