Legends ofPythos
Claim your name
The Standard Library

Regular expressions

Lesson 3 of 7

Watch the lesson1:33 · with Torsten
You already know how to check if a word is in a string with in, or split it up with .split(). Those tools are fast and simple. But when you need to find things that follow a pattern — like every 5-digit number, or all email addresses on a page — plain methods get clumsy.
Regular expressions, called regex for short, let you describe patterns. The re module in the standard library does the matching work.

Use regex when:
  • You need to find a shape of text (digits followed by letters).
  • You want every match at once (findall).
  • You are replacing or validating based on structure.
Skip it if in, .startswith(), or .split() already solves the problem.

Your first pattern

import re

text = 'Order 4821 shipped'
match = re.search(r'\d+', text)
if match:
    print(match.group())   # prints: 4821
re.search finds one match anywhere in a string.
re.search(pattern, string) scans left to right and returns the first match. match.group() gives you that matched text.

Note the r'...' — a raw string so backslashes like \d are not eaten by Python's normal escape rules.

Building blocks

# \d  one digit (0-9)
# \w  word char: letters, digits, underscore
# \s  whitespace
#
# [A-Z]   any single uppercase letter
# [aeiou] vowels only
#
# Quantifiers:
# +    one or more
# *    zero or more
# ?    zero or one
# {2}  exactly two
# {1,3} one to three
#
# Anchors:
# ^    the start of the string
# $    the end of the string
Common shorthand characters and character sets.
# fullmatch checks that WHOLE string fits.
import re
time = '09:35'
ok  = bool(re.fullmatch(r'\d{2}:\d{2}', time))   # True
bad = bool(re.fullmatch(r'\d{2}:\d{2}', '9x'))    # False
fullmatch only matches when the pattern covers the whole string, as if it were wrapped in ^ and $.

Groups and findall

import re
log = 'user=alice action=read user=bob action=write'
users = re.findall(r'user=(\w+)', log)
print(users)   # ['alice', 'bob']
# findall returns a list of group(1) strings when there is one group.
Parentheses create numbered groups. group(1) is the first one.
import re
m = re.search(r'(\d{2}):(\d{2})', 'Meet at 14:30 by the gate')
print(m.group(0))   # 14:30  the whole match
print(m.group(1))   # 14     the first group
print(m.group(2))   # 30     the second group
Pulling pieces out with groups
Parentheses make a group. match.group(0) is everything the pattern matched, and group(1), group(2) are the pieces inside each pair of parentheses, counted from the left.

Replacing with re.sub

import re
msg = 'Call 555-1234 or 5550199 for help'
cleaned = re.sub(r'\d{3}-?\d{4}', '[REDACTED]', msg)
print(cleaned)   # Call [REDACTED] or [REDACTED] for help
re.sub replaces every match in the string.

Your turn

0 of 3 solved

Exercise 1

+40 XP
Flight codes are two capital letters followed by digits, like AA482. Use re.search with two groups to find the code in text. Store the letters in airline and the digits, as an int, in number, then print f'{airline} {number}'.
import re


text = 'FLIGHT AA482 DEPARTED 09:40'


# search with one group for the letters and one for the digits

Run your code to check it against the tests.

Exercise 2

+40 XP
Write prices(text), which uses re.findall to return every price in text written like £3.50 (a pound sign, digits, a point and exactly two digits), as a list of floats. For example, prices('Bread £1.20, cheese £4.75') is [1.2, 4.75].
import re




def prices(text):
    # findall with a group around the number
    pass

Run your code to check it against the tests.

Exercise 3

+40 XP
Write is_time(s), which uses re.fullmatch to check that the whole string is a time like 09:40 (two digits, a colon, two digits) and returns True or False. Then write redact(text), which uses re.sub to replace every phone number like 555-0199 (three digits, a hyphen, four digits) with '***-****'.
import re




def is_time(s):
    pass




def redact(text):
    pass

Run your code to check it against the tests.