DS 208 · Week 10

Text Processing & Regular Expressions

Cleaning, searching, and extracting from messy text — with Python's string tools and the pattern language called regex.

Programming for Data Science · University of the Philippines Cebu

Session Map

From tidy strings to found patterns

Words for Today

Ten words you will hear this week

string method

A built-in action a piece of text can do, written after a dot.

e.g. " hi ".strip() gives back "hi"

regular expression (regex)

A short code that describes a shape of text, so Python can find every piece that fits.

e.g. \d{4} means "four digits in a row"

pattern

The shape you are looking for, written in regex. Most characters mean themselves; a few are special.

e.g. 09\d{9} = "09, then nine more digits"

match

A stretch of text that fits the pattern. re.search finds the first; re.findall finds them all.

e.g. re.findall(r"\d+", "3 cats 4 dogs") gives ['3', '4']

character class

A set of allowed characters that fills one position in the text.

e.g. [A-Z] = one capital letter; \d = one digit

quantifier

A symbol that says how many times the thing just before it may repeat.

e.g. {4} exactly four · + one or more · * zero or more

group

Round brackets around part of a pattern; the text inside is kept ("captured") so you can pull it out.

e.g. (\d{4})-(\d{2}) on 2024-08 keeps 2024 and 08

anchor

A symbol that matches a position, not a character: ^ start, $ end, \b edge of a word.

e.g. ^09 = "the text must begin with 09"

escape

A backslash \ in front of a character flips its meaning between special and ordinary.

e.g. \( = a real bracket; \d = any digit, not the letter d

raw string

A string written r"…": Python leaves every backslash alone, so the regex gets it untouched.

e.g. always write patterns as r"\d+", never "\d+"

Where We Left Off

You can load data — but it's rarely clean

Whatever the source, text arrives with stray spaces, mixed case, and values buried inside sentences. Before you can analyse it, you have to tidy and extract.

You already know a string is text in quotes (week 2) and a pandas column is a Series of values (week 6). Today adds the tools that clean and search the text inside them.

String methods

Enough for tidy, predictable text.

Regex

For when you need a pattern, not an exact string.

String methods are like Find & Replace in Word: you type the exact text. Regex is like telling a friend "find every 11-digit number that starts with 09" — you describe what it looks like, not what it is.

Part A

String Methods

The everyday cleaning toolkit.

You already store text in strings. This part adds string methods — built-in actions a string can do to itself — to trim, re-case, swap and split it.

Tidy First

Three methods clean most text

strip removes surrounding whitespace, lower normalises case so comparisons work, and replace swaps out unwanted pieces.

New words

method
a function that belongs to a value, called with a dot: text.lower() (week 4)
whitespace
the invisible characters: spaces, tabs and line breaks
normalise
make every value follow the same rules, so "Cebu" and "CEBU" compare equal
s = " Cebu City " s.strip() # "Cebu City" s.strip().lower() # "cebu city" s.replace("City", "") # " Cebu "
  • s = " Cebu City "store the text, with two spaces at each end, in a variable called s
  • s.strip()a new string with the spaces at both ends removed; s itself is unchanged
  • s.strip().lower()chaining: strip first, then make the result lowercase
  • s.replace("City", "")swap every "City" for an empty string; the spaces around it stay, so the result still has spaces at both ends
34,079 Real Descriptions

Free text, written by many hands

Every flood control project has a description. They are inconsistent, abbreviated, and full of exactly the structure regex is for.

Start with counting, not extracting

str.contains answers most questions and cannot fail silently the way extraction can. (Extracting = copying the matched piece out into its own column; later today you will see it quietly miss rows.)

d = df.description for kw in ["CONSTRUCTION", "REHABILITATION", "IMPROVEMENT", "REPAIR"]: print(kw, d.str.contains(kw, na=False).sum()) CONSTRUCTION 26,552 REHABILITATION 6,387 IMPROVEMENT 2,791 REPAIR 1,010
  • d = df.descriptiondf is the flood-control projects DataFrame from weeks 5–9; d is its description column (34,079 texts)
  • for kw in [...]:run the indented line once for each of the four keywords
  • d.str.contains(kw, na=False).str applies a string method to every row → one True/False per row; na=False counts a missing description as False
  • .sum()adds up the Trues (True counts as 1)
  • CONSTRUCTION 26,55226,552 descriptions contain the letters CONSTRUCTION somewhere
Text ↔ List

split breaks apart; join puts together

These two are inverses (each undoes the other). split turns one string into a list on a separator — the character that marks where to cut; join glues a list back into one string.

"Cebu,Davao,Iloilo".split(",") # ['Cebu', 'Davao', 'Iloilo'] ", ".join(["a", "b", "c"]) # "a, b, c"
  • "Cebu,Davao,Iloilo".split(",")cut the text at every comma; the commas are thrown away
  • ['Cebu', 'Davao', 'Iloilo']square brackets = a list (week 3) of three strings, each in quotes
  • ", ".join([...])the string before the dot is the glue: put ", " between the items
  • "a, b, c"one single string again
The Cheat Sheet

The string methods you'll reach for

MethodDoesExample → result
.strip()trim whitespace" hi ".strip() → "hi"
.lower()lowercase"HI".lower() → "hi"
.replace(a, b)swap text"a-b".replace("-","_") → "a_b"
.startswith(x)test the start"09xx".startswith("09") → True
x in scontains?"bu" in "Cebu" → True
Part B

Regular Expressions

Describe the shape; find every match.

You can now clean text whose exact value you know. This part adds regex: a way to describe text you only know the shape of — "four digits", "starts with 09" — and find every piece that fits.

Patterns, Not Literals

Search for a shape of text

A regex is a mini-language for describing patterns: "four digits," "a word," "starts with 09." search finds the first; findall returns them all.

The description itself is the pattern; each stretch of text that fits it is a match.

A regex is a wanted poster. It doesn't name anyone; it describes them — "all digits, one or more of them" — and Python checks every stretch of the text against the description.

import re re.search(r"\d+", "order 42") # matches '42' re.findall(r"\d+", "3 cats 4 dogs") # ['3', '4']
  • import reload re, Python's built-in regex module (a toolbox that ships with Python)
  • r"\d+"the pattern. \d = one digit, + = one or more of it → "a run of digits". The r makes it a raw string, so the backslash survives
  • re.search(pattern, text)find the first match. Python shows <re.Match object; span=(6, 8), match='42'>: a match object, found at positions 6–8; None if nothing fits
  • re.findall(pattern, text)find every match → a list of strings, ['3', '4']; an empty list [] if none
Word Boundaries

Your keyword count is 393 too high

Searching for CONSTRUCTION also matches it inside RECONSTRUCTION — which is different work, quietly added to your total.

\b is a zero-width assertion

An assertion is a check, not a character: it matches the position between a word character and a non-word one. It consumes nothing, so it does not change what you capture.

New words

word character
a letter, digit or underscore _ — what \w matches
word boundary \b
the invisible edge where a word starts or ends
zero-width / consume
a token "consumes" the characters it matches; \b matches a position, so it uses up none
# substring — catches RECONSTRUCTION too d.str.contains("CONSTRUCTION", na=False).sum() 26,552 # word boundary — the real count d.str.contains(r"\bCONSTRUCTION\b", na=False).sum() 26,159 26552 - 26159 393 # RECONSTRUCTION
  • "CONSTRUCTION"plain letters: found anywhere, including inside RECONSTRUCTION
  • r"\bCONSTRUCTION\b"read it: word edge, the letters CONSTRUCTION, word edge. In RECONSTRUCTION the "RE" sits before it, so there is no edge → no match
  • .str.contains(...)treats its text as a regex by default, so \b works here
  • 393descriptions counted only because they say RECONSTRUCTION
Anchors And Boundaries

Four things that match a position, not a character

TokenMatchesUse it when
^Start of the stringThe pattern must begin the field
$End of the stringValidating a whole value — e.g. the contract-ID pattern
\bA word boundaryCounting a keyword without catching it inside longer words
(?=...)A lookaheadRequiring something follows without consuming it
The Alphabet Of Regex

Seven tokens go a long way

TokenMatchesTokenMatches
\da digit 0–9+one or more
\wletter, digit, _*zero or more
.any character{4}exactly four
[A-Z]one in the set^ $start / end
A Pattern That Looks Right

And matches 9.7% of your data

Contract IDs. All 8 characters. This is the most common regex failure there is.

Build It From The Examples You Saw

Which is exactly the mistake

Look at the first few IDs — 22F00136 — and the structure seems obvious: two digits, a letter or two, then five digits.

It even matches your test cases

That is the trap. The pattern is validated against the handful of rows you happened to read.

pat = r"^(\d{2})([A-Z]{1,2})(\d{5})$" ex = df.contract_id.str.extract(pat) ex[0].notna().sum() 3,311 # of 34,079 ex[0].notna().mean() 0.097 # 9.7% # and NOT ONE error was raised
  • ^ … $anchors: the whole ID, start to end, must fit
  • (\d{2})group 1: exactly two digits (22)
  • ([A-Z]{1,2})group 2: one or two capital letters (F)
  • (\d{5})group 3: exactly five digits (00136)
  • .str.extract(pat)run the pattern on every ID; each group becomes a column 0, 1, 2; a row that does not match gets NaN (missing, week 6)
  • ex[0].notna().sum() / .mean()how many / what fraction of rows are not missing → 3,311 rows, 0.097 = 9.7%
Look At The Shapes Instead

Two formats, and the rigid pattern excludes the bigger one

ShapeRowsShareExample
DDLLDDDD30,76890.3%18BI0029 — 2 letters, 4 digits
DDLDDDDD3,3119.7%22F00136 — 1 letter, 5 digits
The Fix Is To Stop Counting

Quantify what varies, pin what does not

+ means "one or more". Using it where the length genuinely varies, and exact counts only where it does not, takes this from 9.7% to 100%.

Then verify the coverage

Always print what fraction matched. An extraction that silently returns NaN is the whole problem.

pat = r"^(\d{2})([A-Z]+)(\d+)$" ex = df.contract_id.str.extract(pat) ex[0].notna().mean() 1.0 # 100% # the year prefix decodes too sorted(ex[0].unique()) ['14', '15', ..., '25', '26']
  • (\d{2})still exact: every ID starts with exactly two year digits
  • ([A-Z]+)one or more capital letters: F and BI both fit
  • (\d+)one or more digits: four or five both fit
  • 1.0the fraction of rows that matched: all of them
  • sorted(ex[0].unique())the distinct year prefixes, in order: '14' … '26' = 2014–2026
And Now It Finds Something

1,266 IDs disagree with their own year column

With the pattern working on all rows, the extracted year prefix can be checked against the year column. They agree 96.3% of the time.

3.7% is a real inconsistency

Not a regex bug — a data one. Contract 20BH0131 carries a 2020 prefix on a row labelled 2019. Worth a footnote in anything that uses either field.

prefix_year = "20" + ex[0] bad = df[prefix_year != df.year.astype(str)] len(bad) 1,266 # 3.7% bad[["contract_id", "year"]].head(3) 20BH0131 2019 22MF0066 2021 20CH0073 2019
  • "20" + ex[0]+ on text glues it: "20" + "22" → "2022", for every row at once
  • df.year.astype(str)the year column holds numbers (int64); turn them into text so we compare like with like. Without it every row differs: 34,079
  • df[... != ...]keep only the rows where the two years disagree (a True/False mask, week 6)
  • len(bad)how many rows that is: 1,266
Reading A Pattern

A PH mobile number, decoded

09\d{9} literal "09" every PH mobile starts here \d — a digit any single 0–9 {9} — nine times so 2 + 9 = 11 digits total
Quick Check

Tap to reveal

What does re.findall(r"\d{4}", "1998 and 2020") return?

A · ['1', '9', '9', '8', ...]
B · ['1998', '2020']
C · ['19982020']
D · []

B — ['1998', '2020'].

\d{4} means "exactly four digits in a row," so it grabs each four-digit run as one match. \d alone would give answer A.

Break

  Five minutes

Back to pull fields out and apply patterns to a whole column.

Part C

Regex In Practice

Extract fields; clean whole columns.

You can now write a pattern and find its matches. This part adds groups, which pull out just the pieces you want, and runs patterns down a whole pandas column.

Capture The Pieces

Parentheses pull out parts

Wrap part of a pattern in ( ) to capture it: that bracketed part is a group. After a match, .group(1), .group(2) give you each captured piece separately.

It is like highlighting parts of a sentence you found: the whole match is found, but you copy out only the highlighted bits.

m = re.search(r"(\d{4})-(\d{2})", "2024-08") m.group(1) # '2024' (year) m.group(2) # '08' (month)
  • (\d{4})-(\d{2})group 1: four digits; then a literal hyphen; then group 2: two digits
  • m = re.search(...)m is the match object: the text found plus what each group caught
  • m.group(1)the text group 1 caught → '2024'. m.group(0) is the whole match, '2024-08'
  • m.groups()every group at once, as a tuple: ('2024', '08')
Greedy By Default

A real description, two bracket pairs

* and + are greedy: they take as much as they can while still allowing a match. With two sets of brackets, that is not what you want. Adding ? makes them lazy: as little as possible.

A Cebu project, with a ñ

Same character that broke the encoding in week 9. Real text keeps handing you the same problems.

# CONSTRUCTION OF OCAÑA RIVER # (GABIONS) (UPSTREAM), CARCAR CITY, CEBU # greedy: runs to the LAST ) re.search(r"\((.*)\)", t).group(1) 'GABIONS) (UPSTREAM' # lazy: stops at the FIRST ) re.search(r"\((.*?)\)", t).group(1) 'GABIONS' # 1,543 descriptions have 2+ brackets
  • tthe description text shown in the two comment lines
  • \( and \)escaped: a real "(" and ")" in the text, not a group
  • (.*)group: . any character, * zero or more, as many as possible → runs to the last ")"
  • (.*?)the extra ? = as few as possible → stops at the first ")"
findall Gets Them All

When there is more than one, ask for all of them

search returns the first match. findall returns every one — and with a lazy quantifier it gets the brackets individually.

In pandas: extractall

.str.extractall() is the vectorised version, returning a row per match with a MultiIndex.

New words

vectorised
runs on the whole column in one call, with no loop you write (week 5)
MultiIndex
row labels with two parts, here (original row, match number 0, 1, …)
re.findall(r"\((.*?)\)", t) ['GABIONS', 'UPSTREAM'] # across the whole column df.description.str.extractall(r"\((.*?)\)") one row per bracket pair, indexed by (row, match number) # vs .str.extract, which takes only # the FIRST match per row
Keep These Handy

Patterns you'll write again and again

You wantPatternRead it asMatches
A 4-digit year\d{4}exactly four digits2024
A PH mobile09\d{9}"09", then exactly nine digits09171234567
A date\d{4}-\d{2}-\d{2}4 digits, hyphen, 2 digits, hyphen, 2 digits2024-08-31
A word\w+one or more word characters (letters, digits, _)Cebu
On A Whole Column

.str applies a pattern to every row

pandas exposes string and regex tools through .str. contains filters; extract pulls a captured group into a new column — vectorized (the whole column in one call), no loop.

df[df["phone"].str.contains(r"^09\d{9}$")] # rows with a valid mobile df["code"].str.extract(r"(\d{4})") # the year, as a column
  • df["phone"], df["code"]columns of a small made-up example table (not the flood data)
  • r"^09\d{9}$"start, "09", nine digits, end: the whole value must be an 11-digit mobile
  • .str.contains(...)True/False for every row; df[...] around it keeps only the True rows
  • .str.extract(r"(\d{4})")the first four-digit run in each row, as a new column; NaN where there is none
Regex Responsibly

Powerful, and easy to overdo

Imperfect Extraction Is Normal

A place-name pattern that matches 36%

Free text written by thousands of people will not yield to one pattern. The goal is not 100% — it is knowing your coverage and saying so.

12,400 of 34,079

Good enough to describe where the bulk of work happens; not good enough to claim a complete geographic breakdown. Report which.

places = d.str.extract( r"\b(?:ALONG|IN)\s+([A-Z][A-Z\s\.]{3,40}?)(?:,|\s+\()" , expand=False) places.notna().sum() 12,400 # 36.4% places.dropna().str.strip().value_counts().head(3) AGNO RIVER 140 ANGAT RIVER 131 RIO CHICO RIVER 105
  • d.str.extract(pattern, expand=False)the place name from every description; expand=False gives back one column (a Series), not a table
  • places.notna().sum()12,400 rows matched; the other 21,679 are NaN
  • .value_counts().head(3)the three most common place names and how often each appears
Reading The Code

That place-name pattern, one piece at a time

PieceRead it as
\bat the edge of a word
(?:ALONG|IN)the word ALONG or IN. | means "or"; (?: ) groups without capturing (nothing is kept)
\s+one or more whitespace characters
( … )group 1: the place name we keep
[A-Z]it starts with a capital letter
[A-Z\s\.]{3,40}?then 3 to 40 more capitals, spaces or dots (\. = a real dot); the ? = lazy, stop as early as possible
(?:,|\s+\()stop at a comma, or at spaces followed by a real "("
Always Report Coverage

Three numbers with every extraction

When Not To Use Regex

Three cases where it is the wrong tool

Regex is for shapes of text. When the structure is already known, something else parses it better and more safely.

The classic

Do not parse HTML or JSON with regex. Both nest arbitrarily; a regular expression fundamentally cannot track nesting.

New words

parse / parser
read a known format into usable pieces; a parser is a tool built for one format
nesting
boxes inside boxes: a table inside a section inside a page
# NO — use a parser re.findall(r"<td>(.*?)</td>", html) -> pd.read_html(html) # NO — it is already structured re.search(r'"region": "(.*?)"', text) -> json.loads(text)["region"] # NO — a plain method is clearer re.search(r"^CONST", s) -> s.startswith("CONST")
  • ->not Python: read it as "use this instead"
  • pd.read_html, json.loadsthe parsers from week 9 for HTML tables and JSON text
Compiling

Do it for readability, not for the speed you were promised

The common advice is "compile for performance". Measured over 34,079 strings, the gain is 1.26× — because the re module already caches recent patterns.

The real reasons to compile

A named object you can reuse, pass around and test. And flags attach once rather than at every call site.

New words

compile
re.compile(pat) turns pattern text into a ready-to-use pattern object, once
cache
a short memory: re quietly keeps recently used patterns so it need not redo them
ms
milliseconds, thousandths of a second
# 34,079 searches for s in descriptions: re.search(pat, s) 14.3 ms rx = re.compile(pat) for s in descriptions: rx.search(s) 11.3 ms # 1.26x # worth doing — just not for the # reason usually given
  • for s in descriptions: re.search(pat, s)search each of the 34,079 texts, handing re the pattern text every time
  • rx = re.compile(pat)build the pattern object once and name it rx
  • rx.search(s)the same search, as a method of that object
The Tokens You Actually Need

Ten, and you can write most patterns

TokenMeansExample
\d \w \sdigit / word char / whitespace\d{2} → 22
[A-Z]any one of a set[A-Z]+ → BI
+ * ?one-or-more / zero-or-more / optional[A-Z]+
{n} {n,m}exactly n / between n and m\d{5}
.any character except newline(.*?)
? after a quantifiermake it lazy(.*?) vs (.*)
( )capture a group(?P<yr>\d{2})
^ $ \bpositions, not characters\bCONSTRUCTION\b
|alternation(CONST|REHAB)
\escape a special character\( for a literal bracket
Text Work, In Order

Clean first, then find, then check

Your Turn · 8 min

Break a pattern, then fix it

1 · Reproduce the failure

Run the rigid contract-ID pattern. Confirm you get 9.7% and that nothing raised.

2 · Diagnose with shapes

Map each ID to a D/L signature (D for a digit, L for a letter, as on the shapes slide) and value_counts() it. Find the two formats yourself.

3 · Fix and verify

Loosen the quantifiers. Do not stop until notna().mean() is 1.0.

4 · Extract something useful

Pull the work type (CONSTRUCTION / REHABILITATION / …) from the description. Report your coverage, and look at ten rows that did not match.

Quick Check

Tap to reveal

Your .str.extract() runs without error and the new column is 90% NaN. What happened?

A · The source column is 90% missing
B · The pattern did not match those rows — extract returns NaN rather than raising
C · You need expand=True
D · pandas silently dropped the rows
B. str.extract reports a non-match as NaN, so a pattern that fits only a minority of your data produces a mostly-empty column and no error anywhere. This is why you check notna().mean() after every extraction. Here the rigid contract-ID pattern matched 9.7% because it demanded exactly five trailing digits, and 90.3% of IDs have four.
Quick Check

Tap to reveal

On "OCAÑA RIVER (GABIONS) (UPSTREAM)", what does re.search(r"\((.*)\)", t) capture?

A · "GABIONS"
B · "GABIONS) (UPSTREAM" — greedy matching runs to the last closing bracket
C · ["GABIONS", "UPSTREAM"]
D · Nothing — the brackets need escaping
B. .* is greedy: it takes as much as it can while still permitting a match, so it runs past the first ) to the last one. Use the lazy .*? to stop at the first, and re.findall (or .str.extractall) when you want every pair. 1,543 descriptions in this dataset contain two or more brackets.
Putting It Together

Every mobile number in a message

One pattern, one call, and every Philippine mobile number falls out — regardless of the words around it. This is what regex is for.

text = "Call 09171234567 or 09228889999" re.findall(r"09\d{9}", text) # ['09171234567', '09228889999']
  • 09literal: the characters 0 then 9
  • \d{9}exactly nine more digits (11 in total)
  • re.findall(...)every match, left to right, as a list; the words "Call" and "or" are simply skipped
re.sub Cleans

The same pattern language, used to rewrite

Everything you can find, you can replace. Collapsing whitespace is the one you will use most, because real text is full of stray runs of it.

Chain with .strip()

\s+ squeezes internal runs but leaves the single leading and trailing space it produced.

New word

backreference
\1 in the replacement means "whatever group 1 caught"
s = "CONST. OF FLOOD CONTROL STRUCTURE, CEBU " re.sub(r"\s+", " ", s).strip() 'CONST. OF FLOOD CONTROL STRUCTURE, CEBU' # reference a group with \1 re.sub(r"\s+([,.])", r"\1", "STRUCTURE , CEBU .") 'STRUCTURE, CEBU.' # in pandas: .str.replace(..., regex=True)
  • re.sub(pattern, new, text)find every match and swap in new → a new string
  • r"\s+" → " "each run of spaces becomes a single space; .strip() then trims the ends
  • r"\s+([,.])"spaces, then group 1: one comma or full stop (inside [ ] a dot is just a dot)
  • r"\1"put back only the punctuation group 1 caught, so the spaces before it vanish
Named Groups

Give the pieces names instead of numbers

(?P<name>...) turns a positional capture into a named one — and in pandas, straight into column names.

Worth it past two groups

ex[0], ex[1], ex[2] is unreadable a week later. ex.yr is not.

pat = r"^(?P<yr>\d{2})(?P<office>[A-Z]+)(?P<seq>\d+)$" ex = df.contract_id.str.extract(pat) list(ex.columns) ['yr', 'office', 'seq'] ex.head(3) yr office seq 22 F 00136 18 BI 0029 19 CD 0029
  • (?P<yr>\d{2})a group named yr that catches two digits; ?P<name> is just how you attach the name
  • list(ex.columns)the new columns are called after the groups, not 0, 1, 2
  • ex.head(3)the first three IDs, split into their three parts
Flags

Change how the whole pattern behaves

A flag is an on/off option for the whole pattern. Three you will actually use. Pass them to re functions, or use the inline (?i) form inside a pandas pattern.

case=False is the pandas shortcut

.str.contains(..., case=False) is the same thing as re.IGNORECASE, and reads better in a chain.

re.search(pat, s, re.IGNORECASE) # or in pandas: d.str.contains(pat, case=False) re.VERBOSE # allows whitespace + comments re.MULTILINE # ^ and $ match each line # VERBOSE makes long patterns readable: rx = re.compile(r""" ^(\d{2}) # 2-digit year ([A-Z]+) # office code (\d+)$ # sequence """, re.VERBOSE)
  • re.IGNORECASEupper and lower case count as the same letter
  • re.VERBOSEspaces and # comments inside the pattern are ignored, so you can lay it out
  • re.MULTILINEfor text with line breaks: ^/$ mean start/end of each line
  • r"""…"""a raw string that may run over several lines
Quick Check

Tap to reveal

In a regex, what does \d match?

A · any letter
B · a single digit 0–9
C · a space
D · the literal letter d

B — a single digit 0–9.

\d is one digit; add a quantifier like + or {4} for more. \w is the one that also covers letters.

This Week's Lab

Clean text, then extract with regex

You'll tidy a messy column with string methods, write patterns to pull out years and phone numbers, and apply one across a whole DataFrame column. ~45 minutes.

You'll practise

strip/lower/split, re.findall, groups, and .str.extract.

Two new names: re.split cuts text wherever the pattern matches, and re.fullmatch succeeds only if the whole string fits (like wrapping the pattern in ^…$).

Stretch, if you want

Write one pattern that accepts both 0917... and +63917... formats.

Recap

Clean, then find the shape

Glossary

This week's words, one line each

Words for today

string method
built-in action on text, after a dot: s.strip()
regular expression (regex)
a short code describing a shape of text
pattern
the shape you search for, e.g. 09\d{9}
match
a stretch of text that fits the pattern
character class
allowed characters for one position: [A-Z], \d, \w, \s
quantifier
how many of the thing before: + * ? {4} {1,2}
group
brackets ( ) that capture part of a match
anchor
matches a position: ^ start, $ end, \b word edge
escape
backslash that flips special/ordinary: \( is a real bracket
raw string
r"…": Python leaves backslashes alone; use it for every pattern

Also new today

whitespace
spaces, tabs and line breaks
separator
the character split cuts at
literal
a character that just means itself, like 09
match object
what re.search returns; read it with .group()
word boundary
\b, the edge where a word starts or ends
greedy / lazy
.* takes as much as possible; .*? as little
alternation
| means "or"
backreference
\1: whatever group 1 caught
flag
an option for the whole pattern, e.g. re.IGNORECASE
compile
turn pattern text into a reusable pattern object
extract / coverage
copy matches into a column / the share of rows that matched
Before Next Week

Practice & reading

Next Week

Reproducible Workflows

Git for version history, virtual environments for pinned dependencies, and a calm, systematic way to debug — so your work runs the same for everyone.

DS 208 · Programming for Data Science