DS 227 · Week 3

Data Scraping I

Web fundamentals and HTML parsing — by the end of tonight you will have fetched and parsed a real, live website: this course's own portal.

Knowledge Discovery in Data · University of the Philippines Cebu

Session Map

Where does raw data come from?

Words for Today · 1 of 2

Nine words for how the web works

Each is explained again on the slide where it first matters, and all are in the glossary at the end.

URL

A web address: which server to ask, and which page on it.

e.g. https://portal.latarak.com/calendar

server

A computer that stores web pages and sends them out when asked.

e.g. the machine behind portal.latarak.com

browser

The program you read the web with: it asks servers for pages and draws them.

e.g. Chrome, Firefox, Safari

HTTP

The rules browsers and servers use to talk. https is the encrypted version.

e.g. GET /calendar = “please send the calendar page”

request & response

The request is your program asking for a page; the response is the server’s answer.

e.g. ask for /calendar → get back 24,337 characters of HTML

status code

A three-digit number on every response saying how it went.

e.g. 200 = OK, 404 = not found

HTML

The text format web pages are written in: content wrapped in tags.

e.g. <h1>Academic Calendar</h1>

tag & element

A tag is a label in angle brackets; an element is the opening tag, its content, and the closing tag.

e.g. <li>Aug 6</li> is one li element

attribute & class

An attribute is a name="value" setting inside an opening tag; class is the one that labels elements.

e.g. <table class="cal-table">

Words for Today · 2 of 2

Ten words for the code we write tonight

Tonight is the first week with real code on the slides. These are the Python words it uses (the same meanings DS 208 uses).

library & import

A library is ready-made code someone else wrote; import is the line that loads it.

e.g. import requests

variable

A name that holds a value, like a labelled box.

e.g. r = requests.get(URL) puts the response in r

function & call

A named piece of work. You call it with its name and brackets, inputs inside.

e.g. len(r.text) gives back 24337

string

Text, written inside quotes.

e.g. "html.parser", "cal-table"

list

An ordered collection in square brackets; positions count from 0.

e.g. tables[0] is the first table in the list

loop

Code that repeats a block once per item.

e.g. for tr in rows: runs once per table row

parse

Turn text into a structure a program can search.

e.g. HTML text → a tree of elements

BeautifulSoup

The Python library (bs4) that parses HTML and lets you search the tree.

e.g. soup = BeautifulSoup(r.text, "html.parser")

find / find_all

find gives the first matching element (or None); find_all gives a list of every match.

e.g. soup.find_all("tr") → every table row

get_text

A method that returns the text inside an element, without the tags.

e.g. td.get_text(strip=True) → 'Registration period'

Where We Left Off

The process assumed you had data

Weeks 1–2 gave you the map: definition and process. But CRISP-DM's very first phases need data in hand. Scraping is one way to get it when nobody hands you a CSV (a plain-text table file: commas between columns).

The web is the world's biggest dataset

Prices, listings, directories, records — most of it lives only inside pages.

But it's built for eyes, not analysis

Scraping is the craft of reversing that: page → structure → rows.

Definitions

Three words people mix up

Before Writing Any Code

Scraping is the last resort, not the first

A scraper is code you must maintain against a page you don't control. Check the cheaper doors first — in this order.

1 · A download exists?

Many portals (PSA, data.gov.ph, Kaggle) publish CSVs. Ten seconds beats ten hours.

2 · An API exists?

An API is a door a site opens for programs: ask, and it answers with structured JSON (a text format for data). Stabler than any scraper. Next week's topic.

3 · Only then: scrape

The data is visible on a page and nowhere else. Now it's worth building.

Part A

How the Web Works

One question, one answer — repeated billions of times a second.

Weeks 1–2 were about what to do with data once you have it. This part shows how a web page travels to your computer — the first step in collecting data yourself.

Version 1 · The Complete Picture

What actually happens in those 263 milliseconds

1 DNS lookup portal.latarak.com → 104.21.82.229 2 TCP + TLS handshake connection + encryption · ~44 ms 3 HTTP request GET /calendar HTTP/2 + headers 4 The server side (this portal's real chain) Cloudflare edge → Caddy → Gunicorn → Flask → PostgreSQL …builds the HTML, then back out 5 Response travels back 200 OK + headers + 24 kB of HTML 6 Browser renders — or your scraper parses same bytes, different reader
Version 1 · Decoded

The words on that diagram, in plain language

You do not need to remember these tonight. They are here so the diagram (and the error messages that mention them) are not a wall of jargon.

Fetching a page is like ordering food by phone: look up the number (DNS), dial and agree to talk privately (TCP + TLS), place the order (request), the kitchen cooks it (the server chain), the delivery arrives (response), and you unpack it (render or parse).

Version 2 · The Working Model

Your browser asks; a server answers

Your computer browser / Python Server holds the page GET /calendar → "please send this page" 200 OK + HTML ← "here it is"
Definitions · The Address

A URL has parts — and each part matters

https://portal.latarak.com/calendar?v=7
Scheme

The protocol (the rules for talking). https = encrypted HTTP. Always prefer it.

Host

Which server to ask. One scraper usually stays on one host.

Path

Which page on that server. The part you'll loop over.

Query string

Extra key=value options after ? — filters, pages, versions.

The Envelope

Every response carries a status code

Before you trust the body (the content itself, after the status and headers), check the code. A page that "looks empty" is often a 404 or 403 you didn't read.

A status code is the courier’s delivery slip: “delivered” (200), “moved, new address attached” (301), “recipient refused” (403), “no such address” (404). Read the slip before you open the box.

CodeMeansDo
200OKParse the body
301/302MovedFollow the new URL
403ForbiddenYou're not allowed — stop
404Not foundWrong URL; nothing to parse
429Too many requestsSlow down (next week)
500Server errorTheir bug, not yours — retry later
Definitions · The Small Print

Requests and responses carry headers

Headers are key–value metadata (labelled facts about the message, like Content-Type: text/html) riding along with every request and response. Two matter to scrapers from day one.

User-Agent — who's asking

Your identity string. Python announces itself as python-requests/2.34.2 unless you say otherwise. Some sites answer browsers only.

Content-Type — what came back

The portal answers text/html; charset=utf-8. An API would say application/json. Check it before parsing.

GET vs POST

Scrapers GET (read). POST sends data — forms, logins. Tonight is all GET.

The Payload

The body is just text — HTML text

"View Source" in any browser shows it. It looks like chaos, but it's a strict, nested structure a program can walk. That structure is Part B.

Here li marks one item of a list and span marks a small labelled piece of text; class is the label. Part B explains every tag.

<li> <span class="h-date">Aug 6</span> <span class="h-name">Cebu Provincial Charter Day</span> </li>

This is real

One holiday from portal.latarak.com/calendar — the exact markup we'll parse in Part C.

Quick Check

Tap to reveal

You request a page and get status 404. What should your scraper do with the response body?

A · Parse it anyway
B · Nothing — there's no page to parse
C · Retry immediately forever
D · Treat it as success

B — there's nothing to parse.

A 404 body is usually an error page, not your data. Parsing it produces garbage rows that look real. Check the status before the body.

Tonight's Specimen

We scrape the site you're looking at

portal.latarak.com — this course portal. Target: the academic calendar page, a real table of real dates you actually need.

Why it's a good first target

Public pages, no login, clean tables — and you can open the source next to the code.

Why it's a fair target

I run it, and I'm telling you to scrape it — permission is explicit. It even serves a robots.txt (a file where a site says which pages robots may visit; we'll peek at it later).

Ground rule

Public pages only — /calendar and the course hub. Nothing behind a login, ever.

Your Microscope

View Source and Inspect are different tools

Toolbox Setup

Two installs — on your own laptop

Scraping the open internet needs your machine — the in-portal sandbox can only reach this same site (the lab uses exactly that). One-time setup:

  • pip install …typed in a terminal (a window for typed commands), not in Python: download and install three libraries
  • import requests, bs4Python: load the two libraries so your code can use them
  • print(requests.__version__, …)show each library’s version number. 2.34.2 4.15.0 means both are installed; yours may be newer
# in your terminal (any one of these) pip install requests beautifulsoup4 pandas # or: uv pip install … / conda install … # check — in python import requests, bs4 print(requests.__version__, bs4.__version__) 2.34.2 4.15.0

Division of labour

requests fetches. bs4 parses. pandas holds the result. Three tools, one pipeline.

First Contact

Your first fetch — three lines

This output is real — captured against the live portal while building this deck. The calendar has been edited since, so the live numbers may differ; the frozen snapshot on the course page still gives exactly these.

  • import requestsload the requests library (ready-made code for fetching pages)
  • r = requests.get("https://…")call the function get with the URL as a string: it sends the request and gives back the response, stored in the variable r
  • r.status_code → 200the status code: 200 means OK
  • r.headers["Content-Type"]look up one header by its name, in square brackets: the server says it sent HTML
  • len(r.text) → 24337r.text is the whole page as one string; len counts its characters
import requests r = requests.get("https://portal.latarak.com/calendar") print(r.status_code) 200 print(r.headers["Content-Type"]) text/html; charset=utf-8 print(len(r.text)) 24337 # ~24k characters of HTML
Reading The Response

One response object, four things you'll use

AttributeTypeWhat it is
r.status_codeint200, 404… — check this first
r.textstrThe body decoded to text — what BeautifulSoup wants
r.contentbytesThe raw bytes — for images, PDFs, files
r.headersdict-likeThe response's metadata (Content-Type, …)

New words

attribute (Python)
a value stored on an object, read with a dot and no brackets: r.status_code
object
a value that carries its own attributes and methods, like the response r
int / str / bytes
a whole number / a string (text) / raw data not yet turned into text
dict-like
works like a dictionary: look a value up by its name, like a phone book
2xx
any status from 200 to 299 — all the “it worked” codes
Part B

HTML Structure

Tags, attributes, and the tree they build.

You can now fetch a page as one long string of HTML. This part shows the structure hiding in that string, so you know exactly what to ask for.

Anatomy

One element, four things to grab

<a href="/course/ds227">Knowledge Discovery in Data</a>
Tag name

What kind of element — here, a link (a).

Attribute name

The key you look up — here, href.

Attribute value

Read with tag["href"].

Text content

Read with tag.get_text().

Vocabulary

Eight tags cover most scraping

TagMeaningScraper's interest
<a>LinkText label + href target
<table> <tr> <td> <th>Table, row, cell, header cellReady-made rows and columns
<ul> / <ol> + <li>List + itemsRepeated records without a table
<div>Generic block containerThe box everything hides in
<span>Generic inline containerSmall labelled values (dates, prices)
<h1>…<h6>HeadingsSection titles, record names
<img>Imagesrc attribute — the file's URL
The DOM

Tags nest — so a page is a tree

<div> <ul class="holiday-list"> <li> <li> <li> span.h-date span.h-name

Like folders inside folders: the ul is a folder, each li a folder inside it, each span a file inside that.

How To Aim

Classes and ids are your handles

A page has hundreds of <div>s. What makes one findable is its class or id — the labels the site's own designers added for their CSS (the style rules that make a page look the way it does).

class — a shared label

Many elements can share one, like class="cal-table" on both portal calendar tables. Perfect for "give me all of these."

id — a unique label

Meant to appear once per page, like id="main-content". Perfect for "give me exactly this one."

data-* — machine-readable extras

Custom attributes (data-open="1988") often carry cleaner values than the visible text.

The Luckiest Case

An HTML table is already a dataset

thead holds the header row; tbody holds data rows; each tr is a row of td cells. Your DataFrame (pandas’ table of rows and named columns) is pre-drawn.

<table class="cal-table"> <thead><tr> <th>Milestone</th> <th>First Semester</th> … </tr></thead> <tbody> <tr><td>Registration period</td><td>Aug 3 – Aug 7</td>…</tr> … </tbody> </table>

Real source

Trimmed from View Source on /calendar — compare it yourself.

The Real Difficulty

Your 13 rows are buried in 24,000 characters

"A real page wraps the numbers you want in a hundred tags you don't. Scraping is mostly the patience to find the one class that marks your data."
The calendar page: ~24,300 characters; the data you want: ~1,200 of them
Break

  Ten minutes

Back to turn all this theory into a working scraper.

Part C

Parsing with BeautifulSoup

The library that turns markup back into a tree you can query.

You know how a page is built. This part hands its HTML to BeautifulSoup and pulls out exactly the pieces you want, one Python line at a time.

Step One

Text in, tree out

Hand BeautifulSoup the HTML string and a parser name. You get back a soup object you can search — no string slicing, ever.

  • from bs4 import BeautifulSoupload just the BeautifulSoup tool from the bs4 library
  • BeautifulSoup(r.text, "html.parser")parse the page text from the fetch, using Python’s built-in HTML reader; the searchable tree goes into the variable soup
  • soup.find("title")the first <title> element in the tree
  • .get_text()its text without the tags: the words on the browser tab
  • strip=Truea keyword argument (an input given by name): trim spaces and line breaks from both ends
from bs4 import BeautifulSoup soup = BeautifulSoup(r.text, "html.parser") print(soup.find("title").get_text()) Academic Calendar 2026–2027 - Learning Hub print(soup.find("h1").get_text(strip=True)) Academic Calendar 2026–2027

Parser choice

"html.parser" ships with Python. "lxml" is faster if installed. Same code either way.

The Two Workhorses

find gives one; find_all gives a list

New words

None
Python’s value for “nothing here”
type
the kind of value: one element, a list, a string, None…
keyword (Python)
a word Python reserves for itself, e.g. class, for, if
AttributeError
the error when you ask a value for something it does not have, e.g. None.find_all
The One Confusion To Kill

Text lives between tags; data often lives on them

You want…UseFrom <a href="/course/ds227">Knowledge Discovery…</a>
The visible labela.get_text()"Knowledge Discovery in Data"
The link targeta["href"]"/course/ds227"
A safe attribute reada.get("href")"/course/ds227" or None
A missing attributea["nope"]KeyError

Heads-up: “attribute” means three things in this course

in a dataset
a column, e.g. barangay (Week 1)
in HTML
a name="value" setting on a tag, e.g. href, read with a["href"]
in Python
a value stored on an object, read with a dot, e.g. r.status_code
Whitespace Reality

get_text(strip=True) — almost always

Pretty-printed HTML is full of newlines and indentation. They all count as text. strip=True trims the edges before you store the value.

  • td.get_text()all the text in the cell, including the line breaks and spaces from the HTML file
  • '\n …'the quotes mean the value is a string; \n is how Python shows a line break (a newline)
  • td.get_text(strip=True)the same text with spaces and newlines trimmed from both ends
td.get_text() '\n Registration period\n ' td.get_text(strip=True) 'Registration period'

Why it matters downstream

'Aug 17' and '\n Aug 17 ' are different strings — group-bys and joins silently split on the mess.

A Second Way To Aim

.select() speaks CSS

If you know the CSS a page uses, select() takes the same selector syntax — often shorter than chained find calls. A CSS selector is a short pattern describing which elements to pick: . means class, # means id, a space means “inside”.

# every holiday name on the calendar page soup.select("ul.holiday-list span.h-name") # the one main-content element, by id soup.select("#main-content") # . = class, # = id, space = descendant

find_all or select?

Same power for tonight's work. Pick one and be consistent; you'll read both in the wild.

Moving Around The Tree

Sometimes you land nearby and walk over

MoveCallPortal example
Down (first match)tag.find(...)row → its first td
Down (all matches)tag.find_all(...)table → every tr
Uptag.find_parent("a")course title h2 → the card's link
Sidewaystag.find_next_sibling()span.h-date → its span.h-name
Quick Check

Tap to reveal

You write soup.find("table", class_="caltable") (typo) then .find_all("tr"). What happens?

A · Returns an empty list
B · AttributeError on None
C · Returns all rows anyway
D · A syntax error

B — None has no .find_all().

find matched nothing, returned None, and calling a method on None raised AttributeError. A typo in a class name fails two lines later than you'd expect.

Worked Example · Live

Scraping the portal, end to end

Goal: the academic calendar's "Key Academic Dates" table, as a pandas DataFrame. Five steps, every output real.

You have met each tool on its own. Now we chain them into one working scraper, and read every line aloud. (The outputs match the frozen snapshot of the page; the live calendar has since grown and changed its columns, so expect different counts there.)

Walkthrough · Plan

Every scraper is these five steps

1

Fetch

requests.get + status check

2

Parse

BeautifulSoup(r.text)

3

Locate

find the tag + class marking your data

4

Extract

loop → text + attributes → dicts (dictionaries)

5

Store

DataFrame → CSV, then analyse

Walkthrough · Step 1 of 5

Fetch, then check before parsing

raise_for_status() converts a bad code into a loud error right here — not a mystery three functions later.

  • three import linesload requests (fetch), BeautifulSoup (parse) and pandas, nicknamed pd (tables)
  • URL = "https://…"a variable holding the address as a string; capitals are a habit for values that never change
  • r.raise_for_status()a method (a function that belongs to a value, after a dot): silent if the status is 2xx, otherwise stops with an error on this line
  • 200 24337status OK, and the page is 24,337 characters of HTML
import requests from bs4 import BeautifulSoup import pandas as pd URL = "https://portal.latarak.com/calendar" r = requests.get(URL) r.raise_for_status() # crashes here if not 2xx print(r.status_code, len(r.text)) 200 24337
Walkthrough · Steps 2–3 of 5

Parse, then locate the table

DevTools told us the class: cal-table. Always count what you found before trusting [0] — the page has two such tables.

  • soup.find_all("table", class_="cal-table")a list of every table whose class is cal-table; class_ has an underscore because plain class is reserved by Python
  • len(tables) → 2the list holds two tables
  • cal = tables[0]take the first item: an index is a position number, counted from 0
soup = BeautifulSoup(r.text, "html.parser") tables = soup.find_all("table", class_="cal-table") print(len(tables)) 2 # [0] key dates · [1] breaks & notable dates cal = tables[0]

Locating = Inspect work

The class name came from right-click → Inspect on the live page, nothing fancier.

Walkthrough · Step 4 of 5

Extract the header row first

Column names come from thead's th cells. A list comprehension (a one-line loop that collects its results into a new list) over find_all.

  • cal.find("thead")inside this table only, the header section
  • .find_all("th")every header cell in it, as a list
  • [th.get_text(strip=True) for th in …]for each cell (called th in turn), take its trimmed text; the square brackets collect the answers into a list
  • the outputa list of 4 strings: the table’s column names
headers = [th.get_text(strip=True) for th in cal.find("thead").find_all("th")] print(headers) ['Milestone', 'First Semester', 'Second Semester', 'Midyear']

Scoped search

cal.find(...) searches inside this table only — the other table's headers can't leak in.

Walkthrough · Step 4 of 5

Loop the body rows into dicts

One tr at a time: grab its td texts, zip them with the headers, keep the dict. The next slide reads it line by line.

rows = [] for tr in cal.find("tbody").find_all("tr"): cells = [td.get_text(strip=True) for td in tr.find_all("td")] rows.append(dict(zip(headers, cells))) print(len(rows), rows[0]) 13 {'Milestone': 'Registration period', 'First Semester': 'Aug 3 – Aug 7', 'Second Semester': 'Jan 11 – Jan 15', 'Midyear': 'Jun 9 – Jun 10'}
Walkthrough · Step 4, read aloud

The row loop, one line at a time

  • rows = []start with an empty list that will collect one result per row
  • for tr in cal.find("tbody").find_all("tr"):a loop: run the indented lines once for each body row; tr names the current row
  • cells = [td.get_text(strip=True) for td in tr.find_all("td")]the trimmed text of each cell in this row, as a list of strings
  • zip(headers, cells)pair them up in order: (‘Milestone’, ‘Registration period’), (‘First Semester’, ‘Aug 3 – Aug 7’)…
  • dict(…)turn the pairs into a dictionary: key → value pairs, like a phone book (name → number)
  • rows.append(…)append is a list method: add this row’s dictionary to the end of rows
  • the output13 rows were collected; rows[0] is the first one, printed in curly brackets as key: value pairs

The loop is a clerk copying a paper table into index cards: for each row, read the cells, write each one next to its column name on a fresh card, and add the card to the stack.

Walkthrough · Step 5 of 5

Store — and you're back in pandas-land

A list of dicts drops straight into a DataFrame. From here it's the pandas you met in the Week 1–2 labs: clean, filter, analyse.

  • pd.DataFrame(rows)each dictionary becomes a row; each key becomes a column name
  • df.shape → (13, 4)13 rows, 4 columns
  • df.head(3)the first 3 rows; ... means pandas hid some columns to fit the screen
  • df.to_csv(…, index=False)save the table as a CSV file; index=False leaves out the 0, 1, 2 row labels
df = pd.DataFrame(rows) print(df.shape) (13, 4) print(df.head(3)) Milestone First Semester ... 0 Registration period Aug 3 – Aug 7 ... 1 Start of classes Aug 10, Mon ... 2 Last day to withdraw enlistment Aug 17 ... df.to_csv("up_calendar.csv", index=False)
Walkthrough · All Together

The whole scraper is 14 lines

import requests from bs4 import BeautifulSoup import pandas as pd r = requests.get("https://portal.latarak.com/calendar") r.raise_for_status() soup = BeautifulSoup(r.text, "html.parser") cal = soup.find_all("table", class_="cal-table")[0] headers = [th.get_text(strip=True) for th in cal.find("thead").find_all("th")] rows = [dict(zip(headers, [td.get_text(strip=True) for td in tr.find_all("td")])) for tr in cal.find("tbody").find_all("tr")] df = pd.DataFrame(rows) df.to_csv("up_calendar.csv", index=False)
The Shortcut

For pure tables, pandas can do it alone

pd.read_html finds every <table> in the HTML and returns a list of DataFrames. When it works, it's four lines total.

from io import StringIO dfs = pd.read_html(StringIO(r.text)) print(len(dfs)) 2 dfs[0].head() # same 13×4 table, zero BS4

Real gotcha (pandas ≥ 2.1)

Passing r.text directly is rejected — wrap it in StringIO. A bare string is treated as a file path (a file’s address on disk). StringIO makes text behave like an open file.

Why Learn BS4 At All?

Most web data isn't in a <table>

The portal's holiday list is ul + li + spans. read_html returns nothing for it. Listings, cards, and directories all look like this.

read_html shines when…

The data is a genuine HTML table — Wikipedia, PSA tables, sports standings.

BS4 is required when…

Data lives in divs, lists, spans, attributes — i.e., most modern sites, including most of this portal.

Pro move

Try read_html first. If the list comes back empty or mangled, reach for BS4.

Worked Example 2

No table this time: the holiday list

Same page, different shape. Each li holds two labelled spans — and sometimes a <small> note (small print). Optional fields are where scrapers die.

<ul class="holiday-list"> <li><span class="h-date">Aug 6</span><span class="h-name">Cebu Provincial Charter Day <small>UP Cebu only</small></span></li> <li><span class="h-date">Aug 21</span><span class="h-name">Ninoy Aquino Day</span></li> ← no <small> </ul>

Spot the trap

In the first semester's list, exactly 1 of 12 items has the note. Code that assumes it crashes on the other 11 — or vice versa.

Worked Example 2

Handle the optional field explicitly

find("small") returns None when absent — so test it, and store None as the honest value.

  • soup.select("ul.holiday-list li")every li inside the list with class holiday-list
  • small = li.find("small")this item’s note, or None if it has none
  • hols.append({ … })add a dictionary with two keys, "date" and "note", to the list
  • … if small else Noneif a note was found, use its text; otherwise store None
  • the outputthe first holiday has a note; the second’s note is None: recorded as missing instead of crashing
hols = [] for li in soup.select("ul.holiday-list li"): small = li.find("small") hols.append({ "date": li.find("span", class_="h-date").get_text(strip=True), "note": small.get_text(strip=True) if small else None, }) print(hols[0]) {'date': 'Aug 6', 'note': 'UP Cebu only'} print(hols[1]) {'date': 'Aug 21', 'note': None}
The Hard Truth

Scrapers break — because pages change

Your loop assumes every row has the tags you saw today. The day the site redesigns — or one row is different — the whole loop stops.

Reading this error

Last line first
it names the problem: AttributeError = you asked a value for something it does not have
The message
'NoneType' is the type of None: li.find("small") found nothing, and None has no get_text
Line number
the full traceback (the error report) above it says line 2: the print line
Check first
which find can come back empty?
for li in holidays: print(li.find("small").get_text()) AttributeError: 'NoneType' object has no attribute 'get_text' # 11 of 12 holidays have no <small> — # the second one killed the loop
Write It To Survive

.get() over [], and check for None

Reproducibility

Freeze a snapshot; parse the file

Live pages change under you. For homework and debugging, save the HTML once, then parse the file — same code from Step 2 onward.

  • open("calendar.html", "w")open a file for writing ("w" erases it first); the name is its file path
  • .write(r.text)write the page’s text into the file
  • open("calendar.html").read()open it for reading (the default) and read all its text into html
# save once open("calendar.html", "w").write(r.text) # parse forever — no network, no drift html = open("calendar.html").read() soup = BeautifulSoup(html, "html.parser")

Frozen copy provided

A snapshot of the calendar page taken this week is linked on the course page — your results will match these slides even if the portal changes.

In-Class Activity · 25 min · Pairs

Your turn: two scrapes

Laptops out, same pipeline, new targets. Outputs to beat are on the right — they're real.

Task 1 · The second table

tables[1] is "Breaks & Notable Dates" (Event, Date). Build its DataFrame. Check: 5 rows; first event is the UP Cebu CU Anniversary.

Task 2 · The course hub (harder)

Scrape portal.latarak.com/: every h2.course-title plus the href of its parent link (find_parent("a")). Check: 7 courses; yours is /course/ds227.

Finished early?

Add the holiday lists — all three semesters, with the optional-note fix.

Before You Scrape Anything Else

Just because you can reach it…

Glossary · 1 of 2

This week’s words: the web and HTML

Glossary · 2 of 2

This week’s words: Python and BeautifulSoup

Recap + This Week's Lab

Five things to carry out

Next Week

Data Scraping II

APIs, dynamic pages, and the ethics of scraping — JSON instead of HTML, pages that build themselves, and knowing what you're allowed to take.

DS 227 · Knowledge Discovery in Data