APIs, dynamic pages, and scraping ethics — fetching data over the network, and knowing what you're allowed to take.
Knowledge Discovery in Data · University of the Philippines Cebu
URLs built for programs — ask politely, get clean JSON back.
When the HTML you fetch isn't the page you see — and what to do about it.
robots.txt, rate limits, keys and personal data — the rules of taking.
Last week you parsed a page you already held. This week you go get one — over the network, from a machine someone else pays for.
The best scraper is often the one you never write — because an API already exists.
Last week you fetched the portal's calendar with
requests and pulled 13 rows out of its tags. It worked — and it was
fragile. This week: the doors that don't make you read HTML at all.
Fetch HTML built for eyes, hunt for the right class, hope the page never changes.
Ask a web API — a URL built for programs — and get structured data with no tags at all.
Always check for an API before you scrape the page. It's faster, sturdier and fairer.
A web address built for programs, not people: ask in a fixed way, get data back in a fixed shape.
e.g. /api/demo on this portal
One specific address of an API. Each endpoint serves one kind of data.
e.g. /api/demo/sales/7 gives sale number 7
The message your program sends to a server (“GET me this”) and the reply that comes back.
e.g. r = requests.get(url) sends one; r holds the reply
The three-digit number on every response: 2xx worked, 4xx your mistake, 5xx the server’s.
e.g. 200 OK · 404 not found · 429 too many requests
Text written like Python lists and dictionaries; the usual format of API replies.
e.g. {"item": "kape", "qty": 2}
JSON’s two containers: an object { } holds named values (a dict); an array [ ] holds values in order (a list).
e.g. [{"id": 1}, {"id": 2}] is an array of two objects
Inside an object, the key is the name in quotes; the value is the data after the colon.
e.g. in "qty": 2 the key is qty, the value 2
JSON’s word for “nothing recorded here”. Python turns it into None.
e.g. "closed_on": null becomes None
A setting after ? in the address that narrows what you ask for.
e.g. /api/demo/sales?barangay=Lahug
A long answer split into pages: ask for page 1, then 2… until there is no next page.
e.g. "page": 1, "pages": 3, "next": …
A labelled note that travels with a request or response, outside the data itself.
e.g. Retry-After: 60
A password-like code that tells the API who is asking; often raises your limits.
e.g. X-API-Key: latarak-demo
The most requests an API accepts from you per minute. Go over and you get a 429.
e.g. 12 a minute on our demo API without a key
A page whose data is filled in after it loads, so the HTML your code fetches is nearly empty.
e.g. the browser shows 40 rows; find_all finds 0
The browser running the page’s own program (JavaScript) to fetch data and draw it.
e.g. visible in Inspect, missing from View Source
A plain-text file at a site’s root listing where automated programs should not go.
e.g. Disallow: /private/
The site’s written rules of use. They can forbid scraping even of public pages.
e.g. “no automated collection” clauses
Collecting slowly and openly: a delay between requests, and a User-Agent header saying who you are.
e.g. time.sleep(1)
The same request–response dance — but the answer is built for you.
You already know from week 3 how a program asks a server for a page and gets HTML back. This part keeps that exchange but swaps the HTML for data built for programs: JSON from an API.
HTTP (week 3: the rules browsers and programs use to talk to
web servers) is identical to last week. The GET arrow is your request
(“send me this”); the arrow back is the response. What changes is its
body, the data part: instead of HTML built for a browser, you get JSON built for code.
A web page is a restaurant’s dining room, arranged for people to look at. An API is the take-out window: hand over an order slip in the agreed format, get a packed box back.
API: Application Programming Interface — a contract: "ask this URL (web address) this way, get data back this shape." Each such URL is an endpoint.
The requests library (ready-made code, week 3) does the networking. Your job is the
habit from last week: status first, body second.
timeout=30Networks hang. Without a timeout your notebook waits forever; with one, a dead server becomes an error you can handle.
import requestsload the requests library, which sends web requests for your = requests.get("…/users", timeout=30)send a GET request to that address and keep the whole reply in r; give up after 30 secondsr.status_code → 200the status code: 200 means “OK, here is your data”data = r.json()read the reply’s body and turn its JSON text into Python lists and dictslen(data) → 10the body was a list of 10 records: ten usersCurly braces, square brackets, key–value pairs. One
r.json() call and the text becomes dicts and lists you already know how to
walk.
{ }"id": 1 the key is the name "id", the value is 1[ ]"address": {"city": …}Compare this to the <li> jungle from last week. The structure
is the data.
r.json() is the translator. After it runs, there is no JSON
left in your program — only ordinary Python objects (values). JSON’s
null means “no value recorded” and becomes Python’s None.
JSON is a shipping language every program can read. r.json() unpacks
the parcel into Python’s own containers: a dict (key → value, like a phone book) or a list
(items in order).
| JSON | Python | Example |
|---|---|---|
object { } | dict | u["name"] |
array [ ] | list | data[0] |
| string | str | "Cebu" |
| number | int / float | 27.5 |
true / false | True / False | "open": true |
null | None | None: the gap you must handle |
A sandbox is a safe place to practise, where mistakes cost nothing. Public APIs throttle you (slow you down), move, or die. Ours lives on the same server as this deck — hammer it all you like.
portal.latarak.com/api/demo
BASE = "https://portal.latarak.com"a variable holding the site’s address, typed once and reused on every slideBASE + "/api/demo"+ joins two pieces of text into one full address; /api/demo is the endpointSari-sari store sales that never existed — but the status codes, pagination, errors and rate limits are the genuine article.
Before reading any docs (the API’s instructions), ask the root endpoint:
the top address, /api/demo, with nothing after it. Well-built APIs describe themselves.
It names every endpoint, the auth header, and the rate limit — the three things you need before writing a loop.
d["name"]look up one key of the dict and get its valuefor name in d["endpoints"]:endpoints is a dict inside the dict; looping over it gives its keys, one endpoint per linethe outputthree endpoints, an optional key sent as a header, and the rate limit: 12 requests a minute without the keyA status code is the three-digit number on every response saying how it went.
Two are about you: your key (401, new) and your
pace (429, which last week's table promised to explain here).
| Code | Means | Do |
|---|---|---|
200 | OK | Parse the JSON |
400 | Bad request (a setting you sent is not allowed) | Read the error, fix the request |
401 | No / bad key | Fix your credentials |
404 | No such record | Nothing to parse — handle it |
429 | Too many requests | Slow down, wait, retry |
500 | Server broke | Not your fault; try later |
Read the first digit, like the stamp on a returned parcel: 2 = delivered, 4 = something about your request (wrong address, too many parcels today), 5 = the depot itself broke.
Most APIs give you a collection endpoint (many records) and a single-item endpoint (one record).
The id goes in the path (the part of the address after the site name, here
/api/demo/sales/7), not in a parameter.
A single-item endpoint returns one object. r.json()["total"] works directly — no indexing first.
The single quotes show it is already a Python dict, not JSON text.
There are only 60 sales. A good API does not crash or return an empty 200 — it says 404 (“not found”) and tells you why.
The hint ("Valid ids are 1..60") is the API teaching you its own shape. Most students never look.
Every one of these is our API, right now. The code tells you what to do next before you look at a single byte of body.
400 is new: “bad request”. The address exists,
but a setting you sent (here, limit=99) is not allowed.
200 parse it · 400 fix your request · 404 it does not exist · 429 slow down.
Each JSON object becomes a row; each key becomes a column. Then it is ordinary pandas — the same DataFrame you've used all along.
The column list [["id", "name", "email"]] is a selection — add a key there
to keep one more column.
URLthe users address from the earlier slide, kept in a variablepd.DataFrame(people)build a DataFrame (pandas’ table of rows and named columns): each dict becomes a row, each key a column[["id", "name", "email"]]keep only these three columns: the outer brackets select, the inner list names themdf.shape → (10, 3)the table’s size: 10 rows (users) and 3 columnsA user's address is itself a dict: a dict inside a dict is called
nested. To flatten it into a
column, reach in with two keys before building the frame.
You'll build exactly this name + city table. Chained keys —
nothing fancier needed.
personone record from the list, e.g. people[0]person["address"]its value is a whole dict, not a single piece of textperson["address"]["city"]two keys in a row: open address, then take city from inside itrows = [ {…} for p in people ]a list comprehension: a one-line loop that builds one small dict per person and collects them all in a listA query parameter is a setting added after ? in the
address, like ?postId=1; together they form the query string.
Pass a dict to params= and requests builds the
query string for you — correctly escaped (special characters safely coded), every time.
Spaces, symbols and Unicode all need escaping. params= gets it right;
string-glued URLs eventually don't.
params={"postId": 1}the filter as a dict: key = the parameter’s name, value = what you wantr.url → …/comments?postId=1the full address requests actually sent: it added ?postId=1 for youlen(r.json()) → 5the server sent back only the comments on post 1: five of themDownloading all 60 rows to keep 12 of them wastes their bandwidth (the data they pay to send) and your time. Ask for what you want.
d["total"] → 12d["pages"] → 3d["results"]Passing a dict lets requests encode spaces and symbols correctly. Building the URL by hand is how you get silent 400s.
Pagination means the API splits a long answer into pages
and sends one page per request. Each reply says which page you have, how many there are, and
where the next one is. On the last page next is None (JSON null).
Like shopping-site search results: 5 items per screen and a “Next” button. You keep pressing Next until there isn’t one.
for page in [1, 2, 3]:send the request three times, once per page number"page": pagethe query parameter that picks which page comes backlen(d["results"])records on this page: 5, 5, then the last 2 (12 = 5 + 5 + 2)d["next"] is NoneTrue only on page 3: no next page, so you have everythingAsk for 99 rows a page and our API refuses — but it tells you the cap and how to raise it: send a header (a labelled note that travels with the request, outside the address) carrying an API key (a password-like code that says who is asking).
A 400 body that only says "Bad Request" is a bad API. Ours names the limit and the header that lifts it.
Your loop of API calls suddenly starts
returning 429. What is the server telling you?
C — too many requests.
A 429 is a rate limit: the
server is protecting itself from you. Pause, add a delay between calls, and retry
later — hammering on regardless is how you get banned.
Codes starting with 4 mean your request is the problem; 5 means the server's.
The pace you request at is an ethical choice, not just a technical one.
Same verbs, same status codes. But this is 34,079 real DPWH (Department of Public Works and Highways) flood control projects, 2016–2026, and the numbers are public record.
d["source"]["licence"]source dict, then its licenceBetterGov.ph’s mirror of the DPWH transparency portal. CC0 — public domain, yours to use. The DPWH site itself returns 403 to scripts, which is why we do not demo against it live.
The discovery (root) endpoint carries a data_quality_warning. Most do not — but when one does, reading it saves you from publishing something false.
We are going to ignore this warning in a moment, on purpose, to see what happens to someone who does.
Every project has a contract id. Real money, a real contractor, a real place on the map.
One storm surge protection structure in Donsol, Sorsogon. That is the scale of a single row in this file.
…/projects/22F00136a single-item endpoint: the contract id goes in the pathfor k in ("status", …):repeat for each of the four key names in the bracketsprint(k, "=", p[k])print the key’s name, an equals sign, and that key’s value from the record34,079 rows is small enough to download — but the habit matters. Filter server-side and you move kilobytes instead of megabytes.
Filtering on the server is asking the librarian for the three books you need, instead of carrying the whole library home to search it.
The output: 4,424 completed Region III projects;
at limit 3 per page that is 1,475 pages.
"Region III" contains a space. Let requests encode it; glue the URL by hand and you get a silent 400.
Same pattern as the Talamban pages: keep asking until next is None. On real data it pulls five thousand rows without you ever computing a page number.
Without it, limit: 100 is refused with a 400: anonymous callers get 25 rows a page and 20 requests a minute, so this loop would need 217 requests and hit a 429. With the key: 100 a page, 200 a minute.
headers=KEYsend the API key as a header, so 100 rows a page are allowedwhile True: … breakrepeat for ever, until break jumps out of the looprows += d["results"]add this page’s records to the growing listif not d["next"]: breakno next page (None)? stop; otherwise page += 1 asks for the next55 541255 pages fetched, 5,412 projects collectedFive thousand rows into a DataFrame, and the first real question is already answerable.
pd.DataFrame(rows)the list of 5,412 dicts becomes a table, one row per projectdf["budget"].sum() / 1e9add up the budget column, divide by a billion (1e9 = 1 with 9 zeros)round(…, 1) → 267.7keep one decimal: ₱267.7 billion.value_counts()count how many rows have each status, biggest firstName: count, dtype: int64the result is named count and holds whole numbersdf[df["status"] == "Completed"]a boolean filter: keep only the rows where the test is Truecomp["amount_paid"] == 0True for each completed project whose amount_paid is 04,424 finished projects. Every one shows nothing paid out. In one region alone.
This is the moment your story writes itself. It is also the moment you are most likely to be wrong.
Before writing a word, ask the dullest possible question: does this column ever contain anything else?
.unique() lists each different value once.
[0.] is an array (a NumPy list of numbers) holding a single value, 0.0:
the column never says anything but zero.
amount_paid is not populated in this export. It says nothing about whether anyone was paid. The ₱201B story does not exist.
The code was correct. The filter was correct. The sum was correct. Every number on the previous slide is true.
“Zero in a column” and “zero pesos paid” are different claims. The data cannot tell them apart. You have to.
Publishing this would have accused named companies of something that never happened — on evidence that was an empty column.
A missing value (nothing was recorded) and a zero look identical in a spreadsheet. Ask what a column means before you let it mean something.
Open the API playground (/api-playground on the portal: a page that sends requests when you click) or a notebook. Everything here answers in one request — no key needed yet.
/api/demo/barangays. How many are there?id=42. What item, and what total?barangay=Lahug. How many rows?limit=99. What status, and what does the body tell you?5 barangays · id 42 is load, total 40 · Lahug has 12 rows · limit=99 gives 400 with the cap and the header to lift it.
Next: pages that hide their data, and the ethics of taking it.
When the HTML you fetched isn't the page you saw.
You can now read an API and parse a static page. This part adds the third case: pages that fill themselves in after loading, and how to tell in one minute.
You fetch the page, parse it, and find_all comes back empty —
yet the data is right there on screen. Nothing is broken. The page is
dynamic.
find_allThe server sent a nearly empty shell plus JavaScript. The browser ran the script, which fetched the data and built the page — after your snapshot.
Ctrl+U in most browsers. The raw response — exactly what requests receives. If your data
isn't in here, plain scraping cannot see it.
DevTools (the browser’s developer tools: F12, or right-click → Inspect)
show the live tree after JavaScript ran — what the browser built. This is
what your eyes see.
The one-minute test: open View Source and search for a value you can see on screen. Found? Static — last week's tools work. Missing? Dynamic — you need another door.
Sixty seconds of looking saves an afternoon of debugging an "empty" scraper.
That somewhere is almost always an API the site built for
itself. Open DevTools → Network → filter Fetch/XHR, reload, and watch
the JSON arrive.
Each row in the Network tab is a request the page made while building itself.
Click a Fetch/XHR row and preview the response — one of them holds your data,
clean.
Copy that URL into requests.get() — you're back in Part A, HTML skipped
entirely.
Browser automation means your program drives a real browser. Tools like Selenium and Playwright (Python libraries) run an actual browser from Python — the JavaScript executes, the page builds, and then you read the tree.
Like hiring an assistant to open each page, wait for it to finish loading and copy what is on screen. It always works, but it is slow and every assistant needs a desk (memory).
You see exactly what a user sees — any page, however dynamic.
A whole browser per script: slow, memory-hungry, fragile to site redesigns.
Know it exists; reach for it last. The hidden-API trick usually gets there first — and this week's lab needs neither.
Use it. Cleanest data, clearest rules — Part A is all you need.
Static page — requests + BeautifulSoup, last week's craft.
Call the site's own JSON endpoint directly.
Browser automation — powerful, heavy, last resort.
Each step down this ladder costs more code, more fragility and more server load. Stop at the first door that opens.
The Part C questions apply: are you allowed, and are you being gentle?
| Situation | Reach for | Why |
|---|---|---|
| There is a documented API | API | Clean JSON, paginated, stable. Always try this first. |
| Data is in the page source | requests + parse | Last week’s skill. Cheap, but breaks when markup changes. |
| Page is empty until JS runs | find the hidden API | Open DevTools → Network. The page is calling something. |
| Genuinely nothing else works | drive a browser | Slow, heavy, brittle. The last resort, not the first. |
Just because you can reach it doesn't mean you may take it.
You know three ways to get data. This part is about how fast, as whom, and whether you should: rate limits, keys, robots.txt, terms of service and personal data.
A rate limit is the most requests an API accepts from you in a time window. Our API allows 12 requests a minute anonymously. Run a loop of 15 and you will see the wall — on purpose.
A 429 means your request was fine and you simply sent too many. Retrying instantly makes it worse.
for i in range(15):repeat the indented lines 15 timesprint(r.status_code, end=" ")print each code followed by a space instead of a new linetwelve 200s, then 429the first 12 requests this minute are served; numbers 13, 14 and 15 are refusedDo not invent a sleep time. Retry-After is a response header: the number of seconds the server wants you to wait — honour it.
r.headers["Retry-After"]read one header of the response; it arrives as text, "60"int(…)turn that text into the number 60time.sleep(60)pause the program for 60 seconds, then retryif r.status_code == 429: time.sleep(int(r.headers["Retry-After"])) and then retry. That is a well-behaved scraper.
You get a 429 with Retry-After: 36. What should your code do?
B — wait exactly as long as the server asked.
Retrying instantly makes the limit worse and can get you blocked. The server named the number; honour it. Anything else is guessing.
Same request, one extra header, and the limit goes from 12/min to 200/min. Keys are how a site says "I know who you are."
headers={…}send one header with the request: its name is X-API-Key, its value is the key200the same second that gave 429 without a key now succeedsOurs is public and harmless. A real one belongs in an environment variable (a setting stored outside your code, e.g. in Colab Secrets) — a key pasted into a notebook ends up on GitHub.
robots.txt is a plain-text file where a site lists what automated
programs may not fetch. Nearly every site publishes one at /robots.txt. It's not a
lock — it's a posted request. Honest crawlers (programs that visit pages
automatically) read it first and comply.
"Everyone: stay out of /private/ and /search, and pause 10
seconds between requests."
User-agent: *the rules below are for every program (* = everyone); a user agent is a program’s name tagDisallow: /private/do not fetch any address that starts with /private/Crawl-delay: 10wait 10 seconds between requestsUser-agent: Googlebot / Allow: /Google’s crawler may fetch everythingA tight loop of requests is indistinguishable from an attack. A small pause (a delay) keeps you honest — and keeps you from being banned.
429 anywayStop, wait longer, and re-read the API's documented limits. Retrying instantly makes it worse.
for i in [1, 2, 3]:fetch users 1, 2 and 3, one request eachf"…/users/{i}"an f-string: {i} is replaced by 1, then 2, then 3names.append(u["name"])add each name to a list names made before the looptime.sleep(0.3)wait 0.3 seconds before the next requestPolite scraping means collecting the way the site owner would accept: a delay between requests, and a User-Agent header that names your program and how to reach you.
It is the difference between a stranger rattling every door at night and a visitor who knocks once, gives their name, and waits.
Without it, requests introduces itself only as python-requests/ plus a
version number. With a name and an email, a worried site owner can write to you instead of blocking you.
HEADERS = {"User-Agent": …}your name tag: the program’s name plus a way to contact youheaders=HEADERSsend that header with every requesttime.sleep(1)the delay: wait one second before the next requestr.request.headers[…]what was actually sent: the server saw your name tagMany APIs require one to track quota and identity. Everything done with it is done as you.
An environment variable or a secrets store — in Colab, the Secrets panel ( icon), read at runtime.
Not pasted in the notebook, not committed to git (saved into a project’s version history, often published on GitHub). A pushed key is public forever — history remembers.
You already use this pattern: your portal submit token lives in Colab Secrets, not in the notebook itself.
Revoke and reissue it immediately — don't just delete the file that exposed it.
A licence is a contractor’s government permit to build; REVOKED means it was cancelled. The source encodes licence status in the contractor name. It is public record — and it is still not a verdict on any single project.
When it was revoked, why, or whether it relates to these projects. It is a status at export time. Reporting it as guilt is the same error as the ₱201B story.
| Ask | Because |
|---|---|
| What does this column actually mean? | Empty, zero, and “not applicable” look identical once summed. |
| Could a missing value explain this? | The most dramatic findings are usually data problems. Check that first. |
| Who is named, and can they answer? | A company in your chart is a party in your story. They get to respond. |
| Would I bet my name on this number? | If not, it is not ready — no matter how good the chart looks. |
Terms of service (the site’s written rules of use, usually linked at the
foot of every page) and robots.txt both bind you. "The page loaded" is not
permission.
Personal data (anything that identifies a real person: names, emails, ID numbers) carries legal weight — in the Philippines, the Data Privacy Act of 2012. Public ≠ free to harvest.
Research and journalism weigh differently from reselling a copied database. Purpose shapes what's defensible.
A good test: could you explain your scrape — source, volume, purpose — to the site's owner and to the people in the data, without flinching?
Ethics & privacy get a full week. This is the working minimum until then.
No code. Argue the call — these are the judgements you will actually have to make.
/search. You only want /products.Each one trades your convenience against someone else’s cost or consent. Being able to defend the call matters more than the verdict.
In the Network tab you find a site's hidden endpoint that returns every registered user's name and email. Fair game for your project dataset?
time.sleep()C — reachable is not offered.
That endpoint exposes personal data of real people, almost certainly by mistake. Harvesting it fails every Part C test — being polite about the pace doesn't fix what's being taken. The right move: don't collect it; consider reporting the leak.
Technical access and permission are different questions. Ethics is what happens when only the first one says yes.
The hidden-API trick is for data the site already shows everyone — not for what it accidentally exposes.
/api/demo/sales/7/api/demo/sales/7{ } named values (dict) / [ ] values in order (list)Nonep["address"]["city"][… for p in people]df[test] keeps the rows where the test is True?: ?barangay=Lahugnext is NoneYou'll practise on captured API replies, then call this portal’s own live API: check the
status, turn JSON into a DataFrame, filter, and pace a polite loop. ~45 minutes, in the portal’s
lab page: fill each ____, press Run, then Check.
requests.get with timeout, r.status_code,
r.json(), params=, and pd.DataFrame(records).
Reach into nested JSON for each user's city, and handle a request for a record that doesn't exist.
Check for one before scraping. JSON beats tag-hunting every time.
Same habit as last week — 401 and 429 are about you.
View Source to diagnose; the Network tab usually reveals the site's own endpoint.
robots.txt, pace, keys in secrets — and personal data is never free just because it's reachable.
One sentence: fetch through the cheapest door that opens, and behave at every one of them.
With acquisition done, CRISP-DM's (week 2) Data Preparation phase begins — that's the next two weeks.
Requests docs — "Quickstart": exactly the functions from today, one page.
requests.readthedocs.io
MDN, "An overview of HTTP" — and open robots.txt on two sites you
visit daily; read what they ask of crawlers.
Both are linked on the course page beside this deck and the lab.
The Quickstart reads like a summary of code you've already run.
Cleaning, missing data and outliers — the raw material is in hand; now the real work of making it trustworthy begins.
DS 227 · Knowledge Discovery in Data