Web fundamentals and HTML parsing — by the end of tonight you will have fetched and parsed a real, live website: this course's own portal.
Knowledge Discovery in Data · University of the Philippines Cebu
Requests, responses, URLs, status codes — and your first fetch in Python.
Tags, attributes and the tree — the structure hiding inside every page.
BeautifulSoup end to end: the course portal, scraped live into a DataFrame.
CRISP-DM starts with data you have. Often you don't have it yet — it's sitting in a web page someone built for a browser, not for you.
portal.latarak.com — the site you're reading these slides on. We scrape it
for real, together.
Each is explained again on the slide where it first matters, and all are in the glossary at the end.
A web address: which server to ask, and which page on it.
e.g. https://portal.latarak.com/calendar
A computer that stores web pages and sends them out when asked.
e.g. the machine behind portal.latarak.com
The program you read the web with: it asks servers for pages and draws them.
e.g. Chrome, Firefox, Safari
The rules browsers and servers use to talk. https is the encrypted version.
e.g. GET /calendar = “please send the calendar page”
The request is your program asking for a page; the response is the server’s answer.
e.g. ask for /calendar → get back 24,337 characters of HTML
A three-digit number on every response saying how it went.
e.g. 200 = OK, 404 = not found
The text format web pages are written in: content wrapped in tags.
e.g. <h1>Academic Calendar</h1>
A tag is a label in angle brackets; an element is the opening tag, its content, and the closing tag.
e.g. <li>Aug 6</li> is one li element
An attribute is a name="value" setting inside an opening tag; class is the one that labels elements.
e.g. <table class="cal-table">
Tonight is the first week with real code on the slides. These are the Python words it uses (the same meanings DS 208 uses).
A library is ready-made code someone else wrote; import is the line that loads it.
e.g. import requests
A name that holds a value, like a labelled box.
e.g. r = requests.get(URL) puts the response in r
A named piece of work. You call it with its name and brackets, inputs inside.
e.g. len(r.text) gives back 24337
Text, written inside quotes.
e.g. "html.parser", "cal-table"
An ordered collection in square brackets; positions count from 0.
e.g. tables[0] is the first table in the list
Code that repeats a block once per item.
e.g. for tr in rows: runs once per table row
Turn text into a structure a program can search.
e.g. HTML text → a tree of elements
The Python library (bs4) that parses HTML and lets you search the tree.
e.g. soup = BeautifulSoup(r.text, "html.parser")
find gives the first matching element (or None); find_all gives a list of every match.
e.g. soup.find_all("tr") → every table row
A method that returns the text inside an element, without the tags.
e.g. td.get_text(strip=True) → 'Registration period'
Weeks 1–2 gave you the map: definition and process. But CRISP-DM's very first phases need data in hand. Scraping is one way to get it when nobody hands you a CSV (a plain-text table file: commas between columns).
Prices, listings, directories, records — most of it lives only inside pages.
Scraping is the craft of reversing that: page → structure → rows.
Extracting structured data from pages meant for humans. One site, targeted fields, a table at the end. This course.
Discovering pages by following links at scale — what Google does. Crawlers find; scrapers extract.
Turning text into structure. The step inside scraping where markup (the tags in a page) becomes a tree you can search.
Tonight is scraping: one known site, a handful of pages, specific fields out.
Screen scraping (old term, includes non-web UIs) and data harvesting (scraping at scale, often said disapprovingly).
A scraper is code you must maintain against a page you don't control. Check the cheaper doors first — in this order.
Many portals (PSA, data.gov.ph, Kaggle) publish CSVs. Ten seconds beats ten hours.
An API is a door a site opens for programs: ask, and it answers with structured JSON (a text format for data). Stabler than any scraper. Next week's topic.
The data is visible on a page and nowhere else. Now it's worth building.
One question, one answer — repeated billions of times a second.
Weeks 1–2 were about what to do with data once you have it. This part shows how a web page travels to your computer — the first step in collecting data yourself.
All six steps, measured on the portal while building this deck: DNS 4 ms, handshake 44 ms, total 263 ms. Every https fetch you'll ever make does all of this.
The response headers admit it: server: cloudflare and
via: 1.1 Caddy. Run curl -sI portal.latarak.com/calendar and
look.
You do not need to remember these tonight. They are here so the diagram (and the error messages that mention them) are not a wall of jargon.
Fetching a page is like ordering food by phone: look up the number (DNS), dial and agree to talk privately (TCP + TLS), place the order (request), the kitchen cooks it (the server chain), the delivery arrives (response), and you unpack it (render or parse).
This is the complete picture with steps 1–2 and 4 folded away —
requests (the Python library we fetch pages with) handles the connection for you, and the server side is the site's
problem. You work at ask → answer.
An SSLError (an encryption problem) is step 2; a timeout (no answer in time) is step 4; a 429 is the edge
saying slow down. Error messages name the hidden layers.
The protocol (the rules
for talking). https = encrypted HTTP. Always prefer it.
Which server to ask. One scraper usually stays on one host.
Which page on that server. The part you'll loop over.
Extra
key=value options after ? — filters, pages, versions.
Relative links like /course/ds227 drop the
scheme and host — your code must add them back before requesting.
Loop over paths, filter with query strings — that's how one scraper covers a whole site.
Before you trust the body (the content itself, after the
status and headers), check the code. A page that "looks empty" is often a 404 or
403 you didn't read.
A status code is the courier’s delivery slip: “delivered” (200), “moved, new address attached” (301), “recipient refused” (403), “no such address” (404). Read the slip before you open the box.
| Code | Means | Do |
|---|---|---|
200 | OK | Parse the body |
301/302 | Moved | Follow the new URL |
403 | Forbidden | You're not allowed — stop |
404 | Not found | Wrong URL; nothing to parse |
429 | Too many requests | Slow down (next week) |
500 | Server error | Their bug, not yours — retry later |
Headers are key–value metadata
(labelled facts about the message, like Content-Type: text/html) riding along with
every request and response. Two matter to scrapers from day one.
User-Agent — who's askingYour identity string. Python announces itself as
python-requests/2.34.2 unless you say otherwise. Some sites answer
browsers only.
Content-Type — what came backThe portal answers text/html; charset=utf-8. An API would say
application/json. Check it before parsing.
Scrapers GET (read). POST sends data — forms, logins. Tonight
is all GET.
"View Source" in any browser shows it. It looks like chaos, but it's a strict, nested structure a program can walk. That structure is Part B.
Here li marks one item of a list and span marks
a small labelled piece of text; class is the label. Part B explains every tag.
One holiday from portal.latarak.com/calendar — the exact markup we'll parse
in Part C.
You request a page and get status
404. What should your scraper do with the response body?
B — there's nothing to parse.
A 404 body is usually an error
page, not your data. Parsing it produces garbage rows that look real. Check the
status before the body.
The most dangerous scraper bug is one that succeeds on the wrong page.
Status code first, body second. Always.
portal.latarak.com — this course portal. Target:
the academic calendar page, a real table of real dates you actually
need.
Public pages, no login, clean tables — and you can open the source next to the code.
I run it, and I'm telling you to scrape it — permission is explicit. It even serves a
robots.txt (a file where a site says which pages robots may visit; we'll peek
at it later).
Public pages only — /calendar and the course hub. Nothing behind a
login, ever.
The raw HTML the server sent — exactly what
requests will see. If your data is here, tonight's tools suffice.
The live tree after JavaScript ran (JavaScript: code that runs inside the page and can change it after it loads). If data shows here but not in View Source, it was loaded dynamically — next week's problem.
Workflow: Inspect to find the tag and class marking your data, then confirm it exists in View Source before writing any Python.
Open portal.latarak.com/calendar, right-click a date, Inspect. Find the
<table class="cal-table"> above it.
Scraping the open internet needs your machine — the in-portal sandbox can only reach this same site (the lab uses exactly that). One-time setup:
pip install …typed in a terminal (a window for typed commands), not in Python: download and install three librariesimport requests, bs4Python: load the two libraries so your code can use themprint(requests.__version__, …)show each library’s version number. 2.34.2 4.15.0 means both are installed; yours may be newerrequests fetches. bs4 parses. pandas holds the
result. Three tools, one pipeline.
This output is real — captured against the live portal while building this deck. The calendar has been edited since, so the live numbers may differ; the frozen snapshot on the course page still gives exactly these.
import requestsload the requests library (ready-made code for fetching pages)r = requests.get("https://…")call the function get with the URL as a string: it sends the request and gives back the response, stored in the variable rr.status_code → 200the status code: 200 means OKr.headers["Content-Type"]look up one header by its name, in square brackets: the server says it sent HTMLlen(r.text) → 24337r.text is the whole page as one string; len counts its characters| Attribute | Type | What it is |
|---|---|---|
r.status_code | int | 200, 404… — check this first |
r.text | str | The body decoded to text — what BeautifulSoup wants |
r.content | bytes | The raw bytes — for images, PDFs, files |
r.headers | dict-like | The response's metadata (Content-Type, …) |
r.ok is a handy shortcut: True for any 2xx
status. r.raise_for_status() turns bad codes into loud errors.
r = requests.get(url) → r.raise_for_status() → only then
parse r.text.
r.status_coderTags, attributes, and the tree they build.
You can now fetch a page as one long string of HTML. This part shows the structure hiding in that string, so you know exactly what to ask for.
What kind of element —
here, a link (a).
The key you look up —
here, href.
Read with
tag["href"].
Read with
tag.get_text().
This element is real — it's how the course hub links to this course. An
element is the whole thing: opening tag, text, closing
</a>. It has a name, attributes, and
text.
tag["href"] reads an attribute; tag.get_text() reads the text.
| Tag | Meaning | Scraper's interest |
|---|---|---|
<a> | Link | Text label + href target |
<table> <tr> <td> <th> | Table, row, cell, header cell | Ready-made rows and columns |
<ul> / <ol> + <li> | List + items | Repeated records without a table |
<div> | Generic block container | The box everything hides in |
<span> | Generic inline container | Small labelled values (dates, prices) |
<h1>…<h6> | Headings | Section titles, record names |
<img> | Image | src attribute — the file's URL |
div and span mean nothing by themselves (a block
starts on a new line; inline sits inside a line of text) — their class is what
tells you what's inside.
The calendar uses table.cal-table; holidays use
ul.holiday-list with span.h-date / span.h-name.
Parents contain children. Finding data means navigating
this tree: "each span.h-date inside each li inside the
ul." This one is the portal's holiday list.
BeautifulSoup turns the flat text back into this tree so you can walk it. DOM (Document Object Model) is the browser’s name for it.
Like folders inside folders: the ul is a folder, each
li a folder inside it, each span a file inside that.
A page has hundreds of <div>s. What makes one findable
is its class or id — the labels the site's own designers
added for their CSS (the style rules that make a page look the way it
does).
class — a shared labelMany elements can share one, like class="cal-table" on both portal
calendar tables. Perfect for "give me all of these."
id — a unique labelMeant to appear once per page, like id="main-content". Perfect for "give me
exactly this one."
data-* — machine-readable extrasCustom attributes (data-open="1988") often carry cleaner values than the
visible text.
thead holds the header row; tbody holds data
rows; each tr is a row of td cells. Your
DataFrame (pandas’ table of rows and named columns) is
pre-drawn.
Trimmed from View Source on /calendar — compare it yourself.
"A real page wraps the numbers you want in a hundred tags you don't. Scraping is mostly the patience to find the one class that marks your data."The calendar page: ~24,300 characters; the data you want: ~1,200 of them
Most of a page's markup is not the data — it's chrome (menus, headers, footers) around it.
Right-click → Inspect shows you the exact tag and class before you write code.
The tighter your selector (your description of which elements to pick), the less junk you have to clean later.
Back to turn all this theory into a working scraper.
The library that turns markup back into a tree you can query.
You know how a page is built. This part hands its HTML to BeautifulSoup and pulls out exactly the pieces you want, one Python line at a time.
Hand BeautifulSoup the HTML string and a parser name. You get back a
soup object you can search — no string slicing, ever.
from bs4 import BeautifulSoupload just the BeautifulSoup tool from the bs4 libraryBeautifulSoup(r.text, "html.parser")parse the page text from the fetch, using Python’s built-in HTML reader; the searchable tree goes into the variable soupsoup.find("title")the first <title> element in the tree.get_text()its text without the tags: the words on the browser tabstrip=Truea keyword argument (an input given by name): trim spaces and line breaks from both ends"html.parser" ships with Python. "lxml" is faster if
installed. Same code either way.
find gives one; find_all gives a listThe return type is the whole difference: one tag versus a
list of tags. That decides whether you can loop. Note the underscore in
class_ — plain class is a Python keyword.
find returns None when nothing matches — a common source of
AttributeError two lines later.
None…class, for, ifNone.find_all| You want… | Use | From <a href="/course/ds227">Knowledge Discovery…</a> |
|---|---|---|
| The visible label | a.get_text() | "Knowledge Discovery in Data" |
| The link target | a["href"] | "/course/ds227" |
| A safe attribute read | a.get("href") | "/course/ds227" or None |
| A missing attribute | a["nope"] | KeyError |
Mixing these up is the single most common BeautifulSoup mistake. The text is not an attribute; the attribute is not the text.
Prefer .get("href") when a page might omit the attribute.
KeyError is the error for “no such name here”.
name="value" setting on a tag, e.g. href, read with a["href"]r.status_codeget_text(strip=True) — almost alwaysPretty-printed HTML is full of newlines and indentation. They all count as
text. strip=True trims the edges before you store the value.
td.get_text()all the text in the cell, including the line breaks and spaces from the HTML file'\n …'the quotes mean the value is a string; \n is how Python shows a line break (a newline)td.get_text(strip=True)the same text with spaces and newlines trimmed from both ends'Aug 17' and '\n Aug 17 ' are different strings —
group-bys and joins silently split on the mess.
.select() speaks CSSIf you know the CSS a page uses, select() takes the same
selector syntax — often shorter than chained find calls. A
CSS selector is a short pattern describing which elements to pick:
. means class, # means id, a space means “inside”.
Same power for tonight's work. Pick one and be consistent; you'll read both in the wild.
| Move | Call | Portal example |
|---|---|---|
| Down (first match) | tag.find(...) | row → its first td |
| Down (all matches) | tag.find_all(...) | table → every tr |
| Up | tag.find_parent("a") | course title h2 → the card's link |
| Sideways | tag.find_next_sibling() | span.h-date → its span.h-name |
Every tag you find is itself searchable — find on a row
searches inside that row only. That scoping is what keeps loops correct. A
parent is the element directly around this one; a
sibling sits next to it under the same parent.
We'll need find_parent("a") for the course-hub cards in a few slides.
You write
soup.find("table", class_="caltable") (typo) then
.find_all("tr"). What happens?
AttributeError on NoneB — None has no
.find_all().
find matched nothing, returned
None, and calling a method on None raised
AttributeError. A typo in a class name fails two lines later than you'd
expect.
When find can miss, check for None before you use
the result.
assert table is not None right after the find, while prototyping.
Goal: the academic calendar's "Key Academic Dates" table, as a pandas DataFrame. Five steps, every output real.
You have met each tool on its own. Now we chain them into one working scraper, and read every line aloud. (The outputs match the frozen snapshot of the page; the live calendar has since grown and changed its columns, so expect different counts there.)
requests.get + status check
BeautifulSoup(r.text)
find the tag + class marking your data
loop → text + attributes → dicts (dictionaries)
DataFrame → CSV, then analyse
Steps 1–2 are boilerplate. Step 3 is detective work in DevTools. Step 4 is where scrapers differ. Step 5 hands over to pandas, the table library you met in the Week 1–2 labs.
Fetch → Parse → Locate → Extract → Store. Every scraper you'll ever write.
raise_for_status() converts a bad code into a loud error
right here — not a mystery three functions later.
import linesload requests (fetch), BeautifulSoup (parse) and pandas, nicknamed pd (tables)URL = "https://…"a variable holding the address as a string; capitals are a habit for values that never changer.raise_for_status()a method (a function that belongs to a value, after a dot): silent if the status is 2xx, otherwise stops with an error on this line200 24337status OK, and the page is 24,337 characters of HTMLDevTools told us the class: cal-table. Always count what you
found before trusting [0] — the page has two such tables.
soup.find_all("table", class_="cal-table")a list of every table whose class is cal-table; class_ has an underscore because plain class is reserved by Pythonlen(tables) → 2the list holds two tablescal = tables[0]take the first item: an index is a position number, counted from 0The class name came from right-click → Inspect on the live page, nothing fancier.
Column names come from thead's th cells. A
list comprehension (a one-line loop that collects its results into a
new list) over find_all.
cal.find("thead")inside this table only, the header section.find_all("th")every header cell in it, as a list[th.get_text(strip=True) for th in …]for each cell (called th in turn), take its trimmed text; the square brackets collect the answers into a listcal.find(...) searches inside this table only — the other table's
headers can't leak in.
One tr at a time: grab its td texts,
zip them with the headers, keep the dict. The next slide reads it line by
line.
rows = []start with an empty list that will collect one result per rowfor tr in cal.find("tbody").find_all("tr"):a loop: run the indented lines once for each body row; tr names the current rowcells = [td.get_text(strip=True) for td in tr.find_all("td")]the trimmed text of each cell in this row, as a list of stringszip(headers, cells)pair them up in order: (‘Milestone’, ‘Registration period’), (‘First Semester’, ‘Aug 3 – Aug 7’)…dict(…)turn the pairs into a dictionary: key → value pairs, like a phone book (name → number)rows.append(…)append is a list method: add this row’s dictionary to the end of rows13 rows were collected; rows[0] is the first one, printed in curly brackets as key: value pairsThe loop is a clerk copying a paper table into index cards: for each row, read the cells, write each one next to its column name on a fresh card, and add the card to the stack.
A list of dicts drops straight into a DataFrame. From here it's the pandas you met in the Week 1–2 labs: clean, filter, analyse.
pd.DataFrame(rows)each dictionary becomes a row; each key becomes a column namedf.shape → (13, 4)13 rows, 4 columnsdf.head(3)the first 3 rows; ... means pandas hid some columns to fit the screendf.to_csv(…, index=False)save the table as a CSV file; index=False leaves out the 0, 1, 2 row labelsFetch → parse → locate → extract → store, one screen. This exact script ran while building this deck; every output you saw came from it.
Linked on the course page as scrape_calendar.py — run it tonight.
pd.read_html finds every <table> in the
HTML and returns a list of DataFrames. When it works, it's four lines total.
Passing r.text directly is rejected — wrap it in
StringIO. A bare string is treated as a file path (a
file’s address on disk). StringIO makes text behave like an open file.
<table>The portal's holiday list is ul + li +
spans. read_html returns nothing for it. Listings, cards, and
directories all look like this.
read_html shines when…The data is a genuine HTML table — Wikipedia, PSA tables, sports standings.
Data lives in divs, lists, spans, attributes — i.e., most modern sites, including most of this portal.
Try read_html first. If the list comes back empty or mangled, reach for
BS4.
Same page, different shape. Each li holds two labelled
spans — and sometimes a <small> note (small print). Optional
fields are where scrapers die.
In the first semester's list, exactly 1 of 12 items has the note. Code that assumes it crashes on the other 11 — or vice versa.
find("small") returns None when absent — so
test it, and store None as the honest value.
soup.select("ul.holiday-list li")every li inside the list with class holiday-listsmall = li.find("small")this item’s note, or None if it has nonehols.append({ … })add a dictionary with two keys, "date" and "note", to the list… if small else Noneif a note was found, use its text; otherwise store NoneNone: recorded as missing instead of crashingYour loop assumes every row has the tags you saw today. The day the site redesigns — or one row is different — the whole loop stops.
AttributeError = you asked a value for something it does not have'NoneType' is the type of None: li.find("small") found nothing, and None has no get_textprint linefind can come back empty?.get() over [], and check for NoneA missing value you recorded is data. A crash halfway through a scrape is lost work — and a half-built table you might mistake for complete.
If a missing field means the row is meaningless, failing loudly is the honest choice.
Live pages change under you. For homework and debugging, save the HTML once, then parse the file — same code from Step 2 onward.
open("calendar.html", "w")open a file for writing ("w" erases it first); the name is its file path.write(r.text)write the page’s text into the fileopen("calendar.html").read()open it for reading (the default) and read all its text into htmlA snapshot of the calendar page taken this week is linked on the course page — your results will match these slides even if the portal changes.
Laptops out, same pipeline, new targets. Outputs to beat are on the right — they're real.
tables[1] is "Breaks & Notable Dates" (Event,
Date). Build its DataFrame. Check: 5 rows; first event is
the UP Cebu CU Anniversary.
Scrape portal.latarak.com/: every h2.course-title plus the
href of its parent link (find_parent("a")).
Check: 7 courses; yours is /course/ds227.
Add the holiday lists — all three semesters, with the optional-note fix.
Sites publish crawling rules at /robots.txt. The portal serves one —
read portal.latarak.com/robots.txt tonight; it's short.
Hammering a server is rude and gets you a 429 or a ban.
time.sleep(1) (pause for one second) between requests is the floor of politeness.
Terms of service and personal data both constrain what's fair to take. Public page ≠ public data.
Tonight you scraped with explicit permission, on public pages, a handful of requests. That's the model. The full ethics discussion is next week.
Week 4: APIs, dynamic pages, and scraping ethics in full.
<li> is a tag; tag + content + closing tag is an elementname="value" on a tag; class labels elementsul.holiday-list liimport requestsr = requests.get(URL)len(x)rows.append(x)r.status_code"cal-table"[ ] / a position counted from 0for tr in rows:[x for x in …]requests.get(url)None) / a list of every matchstrip=True trims the endsNone if it is missingStatus code first, body second. raise_for_status() is your friend.
Inspect finds the class; View Source confirms the server sent it.
Fetch → Parse → Locate → Extract → Store. Tables get the
read_html shortcut.
.get(), None checks, snapshots — pages change under
you.
Parse the health-station directory, then scrape the live calendar inside
the portal (~45 min) — then redo it with requests at
home.
BeautifulSoup "Quick Start"; requests "Quickstart"; MDN "What is HTTP?" — linked on the course page.
One sentence: scraping turns a tree built for eyes into rows built for analysis — carefully, because the tree can change under you.
Extraction is a Data Preparation decision, not free data.
APIs, dynamic pages, and the ethics of scraping — JSON instead of HTML, pages that build themselves, and knowing what you're allowed to take.
DS 227 · Knowledge Discovery in Data