DS 227 · Week 4

Data Scraping II

APIs, dynamic pages, and scraping ethics — fetching data over the network, and knowing what you're allowed to take.

Knowledge Discovery in Data · University of the Philippines Cebu

Session Map

Three doors into someone else's data

Where We Left Off

You scraped a page — now stop scraping pages

Last week you fetched the portal's calendar with requests and pulled 13 rows out of its tags. It worked — and it was fragile. This week: the doors that don't make you read HTML at all.

The hard way (last week)

Fetch HTML built for eyes, hunt for the right class, hope the page never changes.

The easy way

Ask a web API — a URL built for programs — and get structured data with no tags at all.

Rule of thumb

Always check for an API before you scrape the page. It's faster, sturdier and fairer.

Words for Today · 1 of 2

Eight words for asking an API (Part A)

API

A web address built for programs, not people: ask in a fixed way, get data back in a fixed shape.

e.g. /api/demo on this portal

endpoint

One specific address of an API. Each endpoint serves one kind of data.

e.g. /api/demo/sales/7 gives sale number 7

request / response

The message your program sends to a server (“GET me this”) and the reply that comes back.

e.g. r = requests.get(url) sends one; r holds the reply

status code

The three-digit number on every response: 2xx worked, 4xx your mistake, 5xx the server’s.

e.g. 200 OK · 404 not found · 429 too many requests

JSON

Text written like Python lists and dictionaries; the usual format of API replies.

e.g. {"item": "kape", "qty": 2}

object / array

JSON’s two containers: an object { } holds named values (a dict); an array [ ] holds values in order (a list).

e.g. [{"id": 1}, {"id": 2}] is an array of two objects

key / value

Inside an object, the key is the name in quotes; the value is the data after the colon.

e.g. in "qty": 2 the key is qty, the value 2

null

JSON’s word for “nothing recorded here”. Python turns it into None.

e.g. "closed_on": null becomes None

Words for Today · 2 of 2

Ten words for asking well (Parts A–C)

query parameter

A setting after ? in the address that narrows what you ask for.

e.g. /api/demo/sales?barangay=Lahug

pagination

A long answer split into pages: ask for page 1, then 2… until there is no next page.

e.g. "page": 1, "pages": 3, "next": …

header

A labelled note that travels with a request or response, outside the data itself.

e.g. Retry-After: 60

API key

A password-like code that tells the API who is asking; often raises your limits.

e.g. X-API-Key: latarak-demo

rate limit

The most requests an API accepts from you per minute. Go over and you get a 429.

e.g. 12 a minute on our demo API without a key

dynamic page

A page whose data is filled in after it loads, so the HTML your code fetches is nearly empty.

e.g. the browser shows 40 rows; find_all finds 0

JavaScript rendering

The browser running the page’s own program (JavaScript) to fetch data and draw it.

e.g. visible in Inspect, missing from View Source

robots.txt

A plain-text file at a site’s root listing where automated programs should not go.

e.g. Disallow: /private/

terms of service

The site’s written rules of use. They can forbid scraping even of public pages.

e.g. “no automated collection” clauses

polite scraping

Collecting slowly and openly: a delay between requests, and a User-Agent header saying who you are.

e.g. time.sleep(1)

Part A

APIs

The same request–response dance — but the answer is built for you.

You already know from week 3 how a program asks a server for a page and gets HTML back. This part keeps that exchange but swaps the HTML for data built for programs: JSON from an API.

Same Model, New Payload

Your program asks; the server answers in JSON

Your program requests / pandas API server serves data, not pages GET /users?city=Cebu → "send me records" 200 OK + JSON ← "here they are"
The Whole Move

Ask, then check the status

The requests library (ready-made code, week 3) does the networking. Your job is the habit from last week: status first, body second.

Why timeout=30

Networks hang. Without a timeout your notebook waits forever; with one, a dead server becomes an error you can handle.

import requests r = requests.get( "https://jsonplaceholder.typicode.com/users", timeout=30) print(r.status_code) # 200 data = r.json() # body, parsed print(len(data)) # 10
  • import requestsload the requests library, which sends web requests for you
  • r = requests.get("…/users", timeout=30)send a GET request to that address and keep the whole reply in r; give up after 30 seconds
  • r.status_code → 200the status code: 200 means “OK, here is your data”
  • data = r.json()read the reply’s body and turn its JSON text into Python lists and dicts
  • len(data) → 10the body was a list of 10 records: ten users
The Payload

JSON is text that looks like Python

Curly braces, square brackets, key–value pairs. One r.json() call and the text becomes dicts and lists you already know how to walk.

New words

JSON
JavaScript Object Notation: a text format programs use to send each other data
object { }
named values in curly brackets; becomes a Python dict
key / value
in "id": 1 the key is the name "id", the value is 1
array [ ]
values in order in square brackets; becomes a Python list. Here: two objects
nested
a value that is itself an object: "address": {"city": …}
[ { "id": 1, "name": "Leanne Graham", "email": "[email protected]", "address": { "city": "Gwenborough" } }, { "id": 2, "name": "Ervin Howell", … } ]

No tags, no tree-walking

Compare this to the <li> jungle from last week. The structure is the data.

Translation Table

Every JSON piece has a Python twin

r.json() is the translator. After it runs, there is no JSON left in your program — only ordinary Python objects (values). JSON’s null means “no value recorded” and becomes Python’s None.

JSON is a shipping language every program can read. r.json() unpacks the parcel into Python’s own containers: a dict (key → value, like a phone book) or a list (items in order).

JSONPythonExample
object { }dictu["name"]
array [ ]listdata[0]
stringstr"Cebu"
numberint / float27.5
true / falseTrue / False"open": true
nullNoneNone: the gap you must handle
Our Sandbox

An API built for this class

A sandbox is a safe place to practise, where mistakes cost nothing. Public APIs throttle you (slow you down), move, or die. Ours lives on the same server as this deck — hammer it all you like.

portal.latarak.com/api/demo

import requests BASE = "https://portal.latarak.com" r = requests.get(BASE + "/api/demo", timeout=30) print(r.status_code) # 200
  • BASE = "https://portal.latarak.com"a variable holding the site’s address, typed once and reused on every slide
  • BASE + "/api/demo"+ joins two pieces of text into one full address; /api/demo is the endpoint

Fictional data, real behaviour

Sari-sari store sales that never existed — but the status codes, pagination, errors and rate limits are the genuine article.

Discovery

A good API tells you how to use it

Before reading any docs (the API’s instructions), ask the root endpoint: the top address, /api/demo, with nothing after it. Well-built APIs describe themselves.

Read this first, always

It names every endpoint, the auth header, and the rate limit — the three things you need before writing a loop.

d = r.json() # the reply, as a Python dict print(d["name"]) for name in d["endpoints"]: print(name) print(d["auth"]) print(d["rate_limit"]) Latarak Demo API GET /api/demo/barangays GET /api/demo/sales GET /api/demo/sales/{id} optional header X-API-Key: latarak-demo raises your rate limit 12/min anonymous, 200/min with the demo key
  • d["name"]look up one key of the dict and get its value
  • for name in d["endpoints"]:endpoints is a dict inside the dict; looping over it gives its keys, one endpoint per line
  • the outputthree endpoints, an optional key sent as a header, and the rate limit: 12 requests a minute without the key
The Envelope, Again

APIs speak in the same codes

A status code is the three-digit number on every response saying how it went.

Two are about you: your key (401, new) and your pace (429, which last week's table promised to explain here).

CodeMeansDo
200OKParse the JSON
400Bad request (a setting you sent is not allowed)Read the error, fix the request
401No / bad keyFix your credentials
404No such recordNothing to parse — handle it
429Too many requestsSlow down, wait, retry
500Server brokeNot your fault; try later

Read the first digit, like the stamp on a returned parcel: 2 = delivered, 4 = something about your request (wrong address, too many parcels today), 5 = the depot itself broke.

One Record

Fetch exactly one thing

Most APIs give you a collection endpoint (many records) and a single-item endpoint (one record). The id goes in the path (the part of the address after the site name, here /api/demo/sales/7), not in a parameter.

r = requests.get(BASE + "/api/demo/sales/7") print(r.json()) {'barangay': 'Mabolo', 'date': '2026-09-07', 'id': 7, 'item': 'noodles', 'price': 13, 'qty': 2, 'total': 26}

Dict, not list

A single-item endpoint returns one object. r.json()["total"] works directly — no indexing first. The single quotes show it is already a Python dict, not JSON text.

When It Is Not There

Ask for id 9999

There are only 60 sales. A good API does not crash or return an empty 200 — it says 404 (“not found”) and tells you why.

r = requests.get(BASE + "/api/demo/sales/9999") print(r.status_code) # 404 print(r.json()) {'error': 'No sale with id 9999.', 'hint': 'Valid ids are 1..60.'}

Read the body on failure

The hint ("Valid ids are 1..60") is the API teaching you its own shape. Most students never look.

Status Codes

Four requests, four verdicts

Every one of these is our API, right now. The code tells you what to do next before you look at a single byte of body.

400 is new: “bad request”. The address exists, but a setting you sent (here, limit=99) is not allowed.

# same API, four different requests /api/demo/sales/7 200 /api/demo/sales/9999 404 /api/demo/sales?limit=99 400 /api/demo/nope 404

Branch on the code

200 parse it · 400 fix your request · 404 it does not exist · 429 slow down.

Two Lines To Rows

A list of records drops straight into pandas

Each JSON object becomes a row; each key becomes a column. Then it is ordinary pandas — the same DataFrame you've used all along.

Keep only what you need

The column list [["id", "name", "email"]] is a selection — add a key there to keep one more column.

import requests, pandas as pd people = requests.get(URL, timeout=30).json() df = pd.DataFrame(people)[["id", "name", "email"]] print(df.shape) # (10, 3)
  • URLthe users address from the earlier slide, kept in a variable
  • pd.DataFrame(people)build a DataFrame (pandas’ table of rows and named columns): each dict becomes a row, each key a column
  • [["id", "name", "email"]]keep only these three columns: the outer brackets select, the inner list names them
  • df.shape → (10, 3)the table’s size: 10 rows (users) and 3 columns
One Wrinkle

Values can be whole objects

A user's address is itself a dict: a dict inside a dict is called nested. To flatten it into a column, reach in with two keys before building the frame.

In the lab's stretch

You'll build exactly this name + city table. Chained keys — nothing fancier needed.

person["address"] # {'city': 'Gwenborough', …} person["address"]["city"] # 'Gwenborough' rows = [{"name": p["name"], "city": p["address"]["city"]} for p in people]
  • personone record from the list, e.g. people[0]
  • person["address"]its value is a whole dict, not a single piece of text
  • person["address"]["city"]two keys in a row: open address, then take city from inside it
  • rows = [ {…} for p in people ]a list comprehension: a one-line loop that builds one small dict per person and collects them all in a list
Narrowing The Ask

Parameters filter on the server side

A query parameter is a setting added after ? in the address, like ?postId=1; together they form the query string. Pass a dict to params= and requests builds the query string for you — correctly escaped (special characters safely coded), every time.

Why not paste it by hand?

Spaces, symbols and Unicode all need escaping. params= gets it right; string-glued URLs eventually don't.

r = requests.get( "…/comments", params={"postId": 1}, timeout=30) print(r.url) # …/comments?postId=1 print(len(r.json())) # 5
  • params={"postId": 1}the filter as a dict: key = the parameter’s name, value = what you want
  • r.url → …/comments?postId=1the full address requests actually sent: it added ?postId=1 for you
  • len(r.json()) → 5the server sent back only the comments on post 1: five of them
Filtering

Let the server do the work

Downloading all 60 rows to keep 12 of them wastes their bandwidth (the data they pay to send) and your time. Ask for what you want.

Reading the output

d["total"] → 12
sales in Talamban altogether
d["pages"] → 3
pages needed at 5 sales per page, the default
d["results"]
this page’s records: a list of dicts
r = requests.get(BASE + "/api/demo/sales", params={"barangay": "Talamban"}) d = r.json() print(d["total"], d["pages"]) # 12 3 print(d["results"][0]) {'barangay': 'Talamban', 'date': '2026-09-03', 'id': 3, 'item': 'sardinas', 'price': 28, 'qty': 3, 'total': 84}

params=, not string-glue

Passing a dict lets requests encode spaces and symbols correctly. Building the URL by hand is how you get silent 400s.

Pagination

Long answers come one page at a time

Pagination means the API splits a long answer into pages and sends one page per request. Each reply says which page you have, how many there are, and where the next one is. On the last page next is None (JSON null).

Like shopping-site search results: 5 items per screen and a “Next” button. You keep pressing Next until there isn’t one.

for page in [1, 2, 3]: d = requests.get(BASE + "/api/demo/sales", params={"barangay": "Talamban", "page": page}).json() print(page, len(d["results"]), d["next"] is None) 1 5 False 2 5 False 3 2 True
  • for page in [1, 2, 3]:send the request three times, once per page number
  • "page": pagethe query parameter that picks which page comes back
  • len(d["results"])records on this page: 5, 5, then the last 2 (12 = 5 + 5 + 2)
  • d["next"] is NoneTrue only on page 3: no next page, so you have everything
When You Ask Wrong

A 400 is a conversation

Ask for 99 rows a page and our API refuses — but it tells you the cap and how to raise it: send a header (a labelled note that travels with the request, outside the address) carrying an API key (a password-like code that says who is asking).

requests.get(BASE + "/api/demo/sales", params={"limit": 99}) # 400 {"error": "limit is capped at 20 on your tier.", "hint": "Send X-API-Key: latarak-demo to raise the cap to 50."}

The fix is in the response

A 400 body that only says "Bad Request" is a bad API. Ours names the limit and the header that lifts it.

Quick Check

Tap to reveal

Your loop of API calls suddenly starts returning 429. What is the server telling you?

A · The record doesn't exist
B · Your key is wrong
C · You're asking too fast — back off
D · The server crashed

C — too many requests.

A 429 is a rate limit: the server is protecting itself from you. Pause, add a delay between calls, and retry later — hammering on regardless is how you get banned.

The Real Thing

Now the same moves, on public money

Same verbs, same status codes. But this is 34,079 real DPWH (Department of Public Works and Highways) flood control projects, 2016–2026, and the numbers are public record.

New words

d["source"]["licence"]
two keys in a row: the source dict, then its licence
CC0 / public domain
free for anyone to reuse; no permission needed
mirror
a copy of a dataset kept on another site
403
“Forbidden”: the server understood and refuses; here it refuses scripts
BASE = "https://portal.latarak.com" r = requests.get(BASE + "/api/flood", timeout=30) d = r.json() print(d["rows"], d["source"]["licence"]) 34079 CC0-1.0 (public domain)

Where it comes from

BetterGov.ph’s mirror of the DPWH transparency portal. CC0 — public domain, yours to use. The DPWH site itself returns 403 to scripts, which is why we do not demo against it live.

Read The Docs First

This API warns you about itself

The discovery (root) endpoint carries a data_quality_warning. Most do not — but when one does, reading it saves you from publishing something false.

New words

field
another word for a column (or a JSON key)
upstream export
the original file this API copied its data from
not populated
never filled in: the 0 is a placeholder, not a measurement
print(d["data_quality_warning"]) {'field': 'amount_paid', 'note': "This field is 0 on EVERY row, here and in the upstream export. It is not populated. It does NOT mean these projects were unpaid. Anyone reporting 'completed but never paid' from this column would be wrong."}

Hold that thought

We are going to ignore this warning in a moment, on purpose, to see what happens to someone who does.

One Project

A single contract, in full

Every project has a contract id. Real money, a real contractor, a real place on the map.

₱279.8 million

One storm surge protection structure in Donsol, Sorsogon. That is the scale of a single row in this file.

r = requests.get(BASE + "/api/flood/projects/22F00136") p = r.json() for k in ("status", "budget", "contractor", "region"): print(k, "=", p[k]) status = Completed budget = 279839557.31 contractor = CENTERWAYS CONSTRUCTION AND DEVELOPMENT INC. (34105) region = Region V
  • …/projects/22F00136a single-item endpoint: the contract id goes in the path
  • for k in ("status", …):repeat for each of the four key names in the brackets
  • print(k, "=", p[k])print the key’s name, an equals sign, and that key’s value from the record
Filtering

Ask the server, not your laptop

34,079 rows is small enough to download — but the habit matters. Filter server-side and you move kilobytes instead of megabytes.

Filtering on the server is asking the librarian for the three books you need, instead of carrying the whole library home to search it.

The output: 4,424 completed Region III projects; at limit 3 per page that is 1,475 pages.

r = requests.get(BASE + "/api/flood/projects", params={"region": "Region III", "status": "Completed", "limit": 3}) d = r.json() print(d["total"], d["pages"]) 4424 1475

params= handles the space

"Region III" contains a space. Let requests encode it; glue the URL by hand and you get a silent 400.

The Loop

Every Region III project, 55 requests

Same pattern as the Talamban pages: keep asking until next is None. On real data it pulls five thousand rows without you ever computing a page number.

Use the key

Without it, limit: 100 is refused with a 400: anonymous callers get 25 rows a page and 20 requests a minute, so this loop would need 217 requests and hit a 429. With the key: 100 a page, 200 a minute.

KEY = {"X-API-Key": "latarak-demo"} rows, page = [], 1 while True: d = requests.get(BASE + "/api/flood/projects", params={"region": "Region III", "limit": 100, "page": page}, headers=KEY).json() rows += d["results"] if not d["next"]: break page += 1 print(page, len(rows)) 55 5412
  • headers=KEYsend the API key as a header, so 100 rows a page are allowed
  • while True: … breakrepeat for ever, until break jumps out of the loop
  • rows += d["results"]add this page’s records to the growing list
  • if not d["next"]: breakno next page (None)? stop; otherwise page += 1 asks for the next
  • 55 541255 pages fetched, 5,412 projects collected
The Payoff

₱267.7 billion, in one region

Five thousand rows into a DataFrame, and the first real question is already answerable.

  • pd.DataFrame(rows)the list of 5,412 dicts becomes a table, one row per project
  • df["budget"].sum() / 1e9add up the budget column, divide by a billion (1e9 = 1 with 9 zeros)
  • round(…, 1) → 267.7keep one decimal: ₱267.7 billion
  • .value_counts()count how many rows have each status, biggest first
  • Name: count, dtype: int64the result is named count and holds whole numbers
import pandas as pd df = pd.DataFrame(rows) print(round(df["budget"].sum() / 1e9, 1)) 267.7 print(df["status"].value_counts()) status Completed 4424 On-Going 822 For Procurement 111 Not Yet Started 51 Terminated 4 Name: count, dtype: int64
The Find

You spot something

comp = df[df["status"] == "Completed"] print(len(comp)) 4424 unpaid = comp[comp["amount_paid"] == 0] print(len(unpaid)) 4424 # every single one print(round(unpaid["budget"].sum() / 1e9, 1)) 201.1 # billion pesos
  • df[df["status"] == "Completed"]a boolean filter: keep only the rows where the test is True
  • comp["amount_paid"] == 0True for each completed project whose amount_paid is 0

₱201 billion of completed work, never paid?

4,424 finished projects. Every one shows nothing paid out. In one region alone.

This is the moment your story writes itself. It is also the moment you are most likely to be wrong.

One Line

Check the column before you check the claim

Before writing a word, ask the dullest possible question: does this column ever contain anything else?

.unique() lists each different value once. [0.] is an array (a NumPy list of numbers) holding a single value, 0.0: the column never says anything but zero.

print(df["amount_paid"].unique()) [0.] # every row in the region. and nationally: # 0 non-zero values across all 248,220 rows # of the source export.

The field is empty, not damning

amount_paid is not populated in this export. It says nothing about whether anyone was paid. The ₱201B story does not exist.

The Rule

What just nearly happened

1 · The finding was real

The code was correct. The filter was correct. The sum was correct. Every number on the previous slide is true.

2 · The meaning was not

“Zero in a column” and “zero pesos paid” are different claims. The data cannot tell them apart. You have to.

3 · Real firms are named

Publishing this would have accused named companies of something that never happened — on evidence that was an empty column.

A missing value (nothing was recorded) and a zero look identical in a spreadsheet. Ask what a column means before you let it mean something.

Your Turn · 8 min

Pull it yourself

Open the API playground (/api-playground on the portal: a page that sends requests when you click) or a notebook. Everything here answers in one request — no key needed yet.

Try it now

  1. Fetch /api/demo/barangays. How many are there?
  2. Fetch sale id=42. What item, and what total?
  3. Filter sales to barangay=Lahug. How many rows?
  4. Now ask for limit=99. What status, and what does the body tell you?

Check yourself

5 barangays · id 42 is load, total 40 · Lahug has 12 rows · limit=99 gives 400 with the cap and the header to lift it.

Break — 10 minutes

Next: pages that hide their data, and the ethics of taking it.

Part B

Dynamic Pages

When the HTML you fetched isn't the page you saw.

You can now read an API and parse a static page. This part adds the third case: pages that fill themselves in after loading, and how to tell in one minute.

The Symptom

The browser shows data; your scraper finds none

You fetch the page, parse it, and find_all comes back empty — yet the data is right there on screen. Nothing is broken. The page is dynamic.

New words

dynamic page
a page whose data is filled in after it loads
JavaScript
the programming language that runs inside web pages, in your browser
JavaScript rendering
the browser running that code to fetch the data and draw it
find_all
week 3’s BeautifulSoup search: every tag that matches
soup = BeautifulSoup(r.text, "html.parser") rows = soup.find_all("li", class_="listing") print(len(rows)) # 0 ← but the browser shows 40!

What actually happened

The server sent a nearly empty shell plus JavaScript. The browser ran the script, which fetched the data and built the page — after your snapshot.

The Diagnosis

Two views of the same page

The Best Trick In This Deck

Dynamic pages get their data from somewhere

That somewhere is almost always an API the site built for itself. Open DevTools → Network → filter Fetch/XHR, reload, and watch the JSON arrive.

New words

Network tab
the DevTools panel listing every request the page makes
Fetch/XHR
the filter showing only background data requests made by the page’s JavaScript
hidden API
an API the site built for its own pages: not advertised, but reachable

1 · Watch the traffic

Each row in the Network tab is a request the page made while building itself.

2 · Find the JSON call

Click a Fetch/XHR row and preview the response — one of them holds your data, clean.

3 · Call it yourself

Copy that URL into requests.get() — you're back in Part A, HTML skipped entirely.

The Heavy Machinery

When all else fails: drive a real browser

Browser automation means your program drives a real browser. Tools like Selenium and Playwright (Python libraries) run an actual browser from Python — the JavaScript executes, the page builds, and then you read the tree.

Like hiring an assistant to open each page, wait for it to finish loading and copy what is on screen. It always works, but it is slow and every assistant needs a desk (memory).

What you gain

You see exactly what a user sees — any page, however dynamic.

What it costs

A whole browser per script: slow, memory-hungry, fragile to site redesigns.

In this course

Know it exists; reach for it last. The hidden-API trick usually gets there first — and this week's lab needs neither.

Decision Path

Cheapest door first

1

Documented API?

Use it. Cleanest data, clearest rules — Part A is all you need.

2

Data in View Source?

Static page — requests + BeautifulSoup, last week's craft.

3

Hidden API in Network tab?

Call the site's own JSON endpoint directly.

4

Still stuck?

Browser automation — powerful, heavy, last resort.

Decide

Which door, and when

SituationReach forWhy
There is a documented APIAPI Clean JSON, paginated, stable. Always try this first.
Data is in the page sourcerequests + parse Last week’s skill. Cheap, but breaks when markup changes.
Page is empty until JS runsfind the hidden API Open DevTools → Network. The page is calling something.
Genuinely nothing else worksdrive a browser Slow, heavy, brittle. The last resort, not the first.
Part C

Scraping Ethics

Just because you can reach it doesn't mean you may take it.

You know three ways to get data. This part is about how fast, as whom, and whether you should: rate limits, keys, robots.txt, terms of service and personal data.

Rate Limits

Watch it say no

A rate limit is the most requests an API accepts from you in a time window. Our API allows 12 requests a minute anonymously. Run a loop of 15 and you will see the wall — on purpose.

Not an error in your code

A 429 means your request was fine and you simply sent too many. Retrying instantly makes it worse.

for i in range(15): r = requests.get(BASE + "/api/demo/barangays") print(r.status_code, end=" ") 200 200 200 200 200 200 200 200 200 200 200 200 429 429 429
  • for i in range(15):repeat the indented lines 15 times
  • print(r.status_code, end=" ")print each code followed by a space instead of a new line
  • twelve 200s, then 429the first 12 requests this minute are served; numbers 13, 14 and 15 are refused
Back Off

The server tells you how long to wait

Do not invent a sleep time. Retry-After is a response header: the number of seconds the server wants you to wait — honour it.

  • r.headers["Retry-After"]read one header of the response; it arrives as text, "60"
  • int(…)turn that text into the number 60
  • time.sleep(60)pause the program for 60 seconds, then retry
print(r.status_code) # 429 print(r.headers["Retry-After"]) # 60 print(r.json()) {'error': 'Too many requests — slow down.', 'hint': "You get 12 requests/min on the 'anon' tier. Send header X-API-Key: latarak-demo for 200/min.", 'retry_after': 60}

The polite loop

if r.status_code == 429: time.sleep(int(r.headers["Retry-After"])) and then retry. That is a well-behaved scraper.

Quick Check

Tap to reveal

You get a 429 with Retry-After: 36. What should your code do?

A · Retry immediately
B · Sleep 36 seconds, then retry
C · Give up and raise an error
D · Switch to a different endpoint

B — wait exactly as long as the server asked.

Retrying instantly makes the limit worse and can get you blocked. The server named the number; honour it. Anything else is guessing.

Identify Yourself

A key raises the ceiling

Same request, one extra header, and the limit goes from 12/min to 200/min. Keys are how a site says "I know who you are."

r = requests.get( BASE + "/api/demo/barangays", headers={"X-API-Key": "latarak-demo"}) print(r.status_code) # 200 # same second, same endpoint — the key is # the only difference
  • headers={…}send one header with the request: its name is X-API-Key, its value is the key
  • 200the same second that gave 429 without a key now succeeds

Never hard-code a real key

Ours is public and harmless. A real one belongs in an environment variable (a setting stored outside your code, e.g. in Colab Secrets) — a key pasted into a notebook ends up on GitHub.

The Posted Sign

robots.txt: the site tells you where not to go

robots.txt is a plain-text file where a site lists what automated programs may not fetch. Nearly every site publishes one at /robots.txt. It's not a lock — it's a posted request. Honest crawlers (programs that visit pages automatically) read it first and comply.

Read it as

"Everyone: stay out of /private/ and /search, and pause 10 seconds between requests."

# https://example.com/robots.txt User-agent: * Disallow: /private/ Disallow: /search Crawl-delay: 10 User-agent: Googlebot Allow: /
  • User-agent: *the rules below are for every program (* = everyone); a user agent is a program’s name tag
  • Disallow: /private/do not fetch any address that starts with /private/
  • Crawl-delay: 10wait 10 seconds between requests
  • User-agent: Googlebot / Allow: /Google’s crawler may fetch everything
Pace Yourself

Your loop is load on someone's server

A tight loop of requests is indistinguishable from an attack. A small pause (a delay) keeps you honest — and keeps you from being banned.

If you get a 429 anyway

Stop, wait longer, and re-read the API's documented limits. Retrying instantly makes it worse.

import requests, time for i in [1, 2, 3]: u = requests.get(f"…/users/{i}", timeout=30).json() names.append(u["name"]) time.sleep(0.3) # be polite
  • for i in [1, 2, 3]:fetch users 1, 2 and 3, one request each
  • f"…/users/{i}"an f-string: {i} is replaced by 1, then 2, then 3
  • names.append(u["name"])add each name to a list names made before the loop
  • time.sleep(0.3)wait 0.3 seconds before the next request
Polite Scraping

Go slowly, and say who you are

Polite scraping means collecting the way the site owner would accept: a delay between requests, and a User-Agent header that names your program and how to reach you.

It is the difference between a stranger rattling every door at night and a visitor who knocks once, gives their name, and waits.

Why say who you are

Without it, requests introduces itself only as python-requests/ plus a version number. With a name and an email, a worried site owner can write to you instead of blocking you.

import requests, time HEADERS = {"User-Agent": "ds227-lab ([email protected])"} for page in [1, 2, 3]: r = requests.get(BASE + "/api/demo/sales", params={"page": page}, headers=HEADERS, timeout=30) print(page, r.status_code) time.sleep(1) # one request per second print(r.request.headers["User-Agent"]) 1 200 2 200 3 200 ds227-lab ([email protected])
  • HEADERS = {"User-Agent": …}your name tag: the program’s name plus a way to contact you
  • headers=HEADERSsend that header with every request
  • time.sleep(1)the delay: wait one second before the next request
  • r.request.headers[…]what was actually sent: the server saw your name tag
Credentials

A key identifies you — guard it like a password

Named Entities

Some contractors are marked REVOKED

A licence is a contractor’s government permit to build; REVOKED means it was cancelled. The source encodes licence status in the contractor name. It is public record — and it is still not a verdict on any single project.

r = requests.get(BASE + "/api/flood/contractors", params={"limit": 4}) for c in r.json()["results"]: print(c["projects"], c["licence_revoked"]) 876 False # "Unknown" — missing contractor 345 True # ST. TIMOTHY CONSTRUCTION 306 False # LEGACY CONSTRUCTION 283 True # ALPHA & OMEGA GEN. CONTRACTOR

What REVOKED does not tell you

When it was revoked, why, or whether it relates to these projects. It is a status at export time. Reporting it as guilt is the same error as the ₱201B story.

Before You Publish

Four questions, every time

AskBecause
What does this column actually mean? Empty, zero, and “not applicable” look identical once summed.
Could a missing value explain this? The most dramatic findings are usually data problems. Check that first.
Who is named, and can they answer? A company in your chart is a party in your story. They get to respond.
Would I bet my name on this number? If not, it is not ready — no matter how good the chart looks.
Beyond Manners

Three questions before any real scrape

Your Turn · 5 min

Would you scrape it?

No code. Argue the call — these are the judgements you will actually have to make.

Try it now

  1. A site has an API but it costs money. The HTML is free. Scrape it?
  2. robots.txt disallows /search. You only want /products.
  3. You need 50,000 pages by tomorrow morning.
  4. The data is public, but it is named individuals.

No clean answers

Each one trades your convenience against someone else’s cost or consent. Being able to defend the call matters more than the verdict.

Quick Check

Tap to reveal

In the Network tab you find a site's hidden endpoint that returns every registered user's name and email. Fair game for your project dataset?

A · Yes — it's publicly reachable
B · Yes, if I add time.sleep()
C · No — personal data the site never meant to expose
D · Only on weekends

C — reachable is not offered.

That endpoint exposes personal data of real people, almost certainly by mistake. Harvesting it fails every Part C test — being polite about the pace doesn't fix what's being taken. The right move: don't collect it; consider reporting the leak.

Glossary Recap

Every new word, one line each · Part A

Asking an API

API
a web address built for programs; returns data, not pages
endpoint
one address of an API, e.g. /api/demo/sales/7
root endpoint
the API’s top address; often describes the rest
request / response
the message you send / the reply that comes back
HTTP
the rules browsers and programs use to talk to servers
path
the address after the site name: /api/demo/sales/7
status code
three digits: 2xx worked, 4xx your mistake, 5xx theirs
body
the data part of a response (HTML or JSON)
timeout
how long to wait before giving up on a request
sandbox
a safe practice copy where mistakes cost nothing

Reading JSON

JSON
text shaped like Python lists and dicts
object / array
{ } named values (dict) / [ ] values in order (list)
key / value
the name before the colon / the data after it
null
JSON’s “nothing here”; becomes None
nested
a container inside a container: p["address"]["city"]
list comprehension
a one-line loop that builds a list: [… for p in people]
DataFrame
pandas’ table: each record a row, each key a column
boolean filter
df[test] keeps the rows where the test is True
missing value
nothing recorded; can hide behind a 0
Glossary Recap

Every new word, one line each · Parts A–C

Asking well

query parameter
a setting after ?: ?barangay=Lahug
pagination
results in pages; loop until next is None
header
a labelled note sent with a request or response
API key
a password-like code saying who is asking
rate limit
most requests allowed per minute; over it → 429
400 / 401 / 403
bad request / missing key / forbidden
404 / 429 / 500
not found / too many requests / server broke
environment variable
a setting kept outside the code, e.g. Colab Secrets

Pages and manners

dynamic page
data filled in after loading; fetched HTML is nearly empty
JavaScript rendering
the browser running page code to fetch and draw data
DevTools / Network tab
the browser’s inspection panel / its list of requests
browser automation
a program driving a real browser (Selenium, Playwright)
robots.txt
a site’s posted list of where programs may not go
terms of service
the site’s written rules of use
polite scraping
a delay between requests plus an honest User-Agent
User-Agent
the header naming your program: your name tag
personal data
anything identifying a real person; protected by law
This Week's Lab

Pull data from a real API

You'll practise on captured API replies, then call this portal’s own live API: check the status, turn JSON into a DataFrame, filter, and pace a polite loop. ~45 minutes, in the portal’s lab page: fill each ____, press Run, then Check.

You'll practise

requests.get with timeout, r.status_code, r.json(), params=, and pd.DataFrame(records).

Stretch, if you're quick

Reach into nested JSON for each user's city, and handle a request for a record that doesn't exist.

Recap

Four things to carry out

Readings

Before next week

Next Week

Data Preprocessing I

Cleaning, missing data and outliers — the raw material is in hand; now the real work of making it trustworthy begins.

DS 227 · Knowledge Discovery in Data