DS 227 · Week 12

Course Synthesis

The whole knowledge-discovery arc in one view — and your final presentations to bring it home.

Knowledge Discovery in Data · University of the Philippines Cebu

Session Map

Closing the loop

Words for Today · 1 of 2

The workflow, in eight plain words

pipeline

The fixed order of steps that turns raw data into an answer: load, prepare, explore, communicate.

e.g. one notebook that runs all four, top to bottom

load

Get the data into Python as a table you can work with.

e.g. df = pd.read_csv("flood.csv")

prepare

Fix what is wrong before analysing: types, missing values, outliers, joins, and privacy.

e.g. turn text dates into dates; drop the name column

explore

Summarise and chart the data to find patterns, then check them.

e.g. describe(), a histogram, a correlation

communicate

Turn one checked finding into a sentence, a chart and a talk a stranger can trust.

e.g. a finding-title with its number and its limit

reproducible

Someone else can rerun your work from the start and get the same numbers.

e.g. Restart & Run All finishes with no errors

provenance

Where the data came from: who collected it, when, why, and under what licence.

e.g. "DPWH export via the BetterGov.ph mirror", plus the date you downloaded it

limitation (caveat)

A plain statement of what your data or method cannot show.

e.g. "no outcome column, so I cannot say if flooding fell"

Words for Today · 2 of 2

Eight words for the final project

deliverable

Something you hand in and are graded on.

e.g. the notebook, the written claim, the talk

rubric

The list of criteria your work is marked against, with what earns full marks.

e.g. defensibility, correctness, method, communication, ethics

claim

Your finding written as four sentences: the finding, its number, the method, the limit.

e.g. the week-9 template

scope

How big the question is. A good scope is one sentence with the exact columns named.

e.g. not "analyse education" but "completion vs teachers per pupil, Region VII"

defensible

Able to survive a sceptical reader: every number traced to code, every limit stated.

e.g. "80.8% complete (n=34,079)", not "almost all complete"

null result

A checked finding that there is no clear difference or relationship. Still a result.

e.g. "completion rates do not track budget size"

peer review

A classmate checks your work before submission, looking for what you missed.

e.g. "which sentence goes further than the evidence?"

Restart & Run All

Wipe the notebook's memory, then run every cell from the top. The honest test that it works.

e.g. Colab: Runtime → Restart and run all

Part A

The Arc

What twelve weeks assembled.

You met every step one week at a time. This part lines them up into one workflow, load → prepare → explore → communicate, that your project follows.

One Pipeline

Every week was a station on the same line

1

Frame

Weeks 1–2 · KDD & CRISP-DM: the question comes first.

2

Acquire

Weeks 3–4 · Scraping, APIs, and the manners of taking.

3

Prepare

Weeks 5–6 · Gaps, outliers, merges, one honest table.

4

Explore

Weeks 7–8 · Describe, draw, correlate — carefully.

5

Tell & Guard

Weeks 9–11 · Findings, dashboards, ethics.

The Workflow, In Plain Words

Load → prepare → explore → communicate

StepIn plain wordsWhat you usedThe trap to watch
LoadGet the data into a table and look at itpd.read_csv, an API, a scraper; shape, head()A column that looks full but is empty (amount_paid)
PrepareFix types, gaps, outliers and joins; remove what identifies peopleto_datetime, isna(), merge, dropA cleaning step that quietly deletes a whole group
ExploreSummarise and chart; look for a pattern, then try to break itdescribe(), groupby, histograms, corr()A summary that silently skipped rows
CommunicateOne checked finding, as a sentence and a chart, with its limitfinding-titles, sorted bars, the four-sentence claimA verb stronger than the evidence ("proves", "causes")
One Dataset. Twelve Weeks.

34,079 flood control projects

Everything you learned, you learned on this.

The Whole Course, In One Table

Each week added a technique and exposed a trap

WeekThe techniqueThe trap it exposed
4APIs, JSON, paginationamount_paid is 0 on every row — unpopulated, not unpaid
5Missing data, outliersdropna() removed 25.6% of rows and every unfinished project
6Joins, types, derived columnsA project that finished 17 days before it started
7Descriptive statisticsdescribe() silently excluded 6,415 rows
8Correlation, visualisationr = −0.785 between latitude and longitude
9Data journalismA ₱75.2B headline the data could not support
10Storytelling, dashboardsA partial year drawn as a 47% collapse
11Ethics, privacyOne rounded column made 30.9% of records unique
The Whole Thing Is One Script

Every figure in the course, in forty lines

Not a summary of the course — the actual pipeline, shortened here. It runs in about a second and reproduces every number you were shown.

This is your project, in miniature

Load, convert, check, derive, summarise, relate, protect. The order matters and you now know why each step precedes the next.

  • pd.read_csv(path)load: read the CSV file into a DataFrame
  • pd.to_numeric(..., errors="coerce")turn text into numbers; anything unreadable becomes NaN (missing) instead of an error
  • (df.completion_date - df.start_date).dt.daysdays between two dates (the full script first converts both columns with pd.to_datetime, week 6)
  • df.isna().sum()how many missing cells in each column
  • ... and the # w7-8 linessteps left out to fit the slide; each is a week's lab
# w4: get it df = pd.read_csv(path) # w6: types BEFORE anything else df["budget"] = pd.to_numeric(df.budget, errors="coerce") df["duration"] = (df.completion_date - df.start_date).dt.days # w5: what is missing, and why df.isna().sum(); df.groupby("status")... # w7-8: describe, skew, correlate # w9-10: claim, limit, chart # w11: k-anonymity before release
And Its Output

Every number you were taught, reproduced in one run

$ python pipeline.py w4 loaded 34,079 rows w4 amount_paid unique [0] w5 naive dropna() 25,370 (25.6% lost) w5 completed, true 80.8% w5 completed, after drop 99.6% <- the lie w6 impossible durations 1 (16DC0070) w7 describe() count 27,664 (hid 6,415) w7 budget skew 5.13 w8 pearson / spearman 0.365 / 0.508 w8 median days by quartile [119,189,226,269] w9 contractors 4,841 w9 top 1% share 27.6% w11 unique at 1M 30.9% w11 unique at 50M bands 0.4%
Reading That Output

What each line of the run is telling you

LineWhat it means
amount_paid unique [0]Every row holds 0: the column was never filled in, so it says nothing about payment
naive dropna() 25,370 (25.6% lost)dropna() deletes any row with an empty cell; a quarter of the data went
completed, after drop 99.6%Unfinished projects have no end date, so dropping gaps dropped them: 80.8% became 99.6%
budget skew 5.13Skew above 0 means a long right tail: most budgets are small, a few are huge
pearson / spearman 0.365 / 0.508Two correlation scores (week 8): 0 = no relationship, 1 = perfect; bigger budgets tend to take longer
unique at 50M bands 0.4%With ₱50M budget bands and region instead of province, 0.4% of projects stand alone (week 11)
Writing It Found A Bug

In the synthesis, not in the course

The first run of that script reported 27,664 rows surviving dropna(). Week 5 had said 25,370. One of them was wrong.

  • csv.DictReader(f)Python's built-in CSV reader: each row becomes a dictionary of text
  • ''an empty string: text with nothing in it. It is a value, so dropna() keeps it
  • pd.read_csv(f)pandas' reader: an empty cell becomes NaN, the missing-value marker, which dropna() removes

It was the new script

csv.DictReader leaves an empty field as ''; pd.read_csv makes it NaN. Only the second is dropped — so the loader changed the answer.

# gives '' — dropna sees a value pd.DataFrame(csv.DictReader(f)) # then both date columns through # pd.to_datetime ('' becomes NaT, missing) # → dropna() keeps 27,664 # gives NaN — dropna removes it pd.read_csv(f) # → dropna() keeps 25,370 # same data. same method. different # answer, because of how it was read.
The Toolbox

Task to tool, twelve weeks compressed

This table is your project cheat-sheet. Every row is muscle memory by now — the labs made sure.

Reminders, in a few words

requests / r.json()
fetch a web page or API / read its JSON reply
BeautifulSoup
find things inside a page's HTML
IQR fences / clip
outlier cut-offs / cap values at a limit
merge / concat
join tables side by side on a key / stack them
groupby / pivot_table
summarise per group / reshape into rows × columns
Need to…Reach for
Fetch datarequests, BeautifulSoup, r.json()
Audit & cleanisna().sum(), IQR fences, clip
Combine & shapemerge, concat, groupby, pivot_table
Exploredescribe, histograms, corr, scatters
Communicatesmall multiples, sorted bars, finding-titles
Protectk-anonymity audit, banding, suppression
Deeper Than Tools

The habits were the real syllabus

The Seven Checks

One per week, and they are the deliverable

CheckFromCost
.unique() on the column your claim rests onw45 seconds
Row count before and after every filterw52 lines
Range check on every derived columnw61 assert
Compare count to len(df)w71 line
Recompute the relationship within groupsw81 groupby
Check the verb matches the evidencew91 reading
Compute k before releasingw111 groupby
What You Can Now Do

Stated plainly, because you should know it

Twelve weeks ago a 34,000-row government dataset was something you read about. You can now take one apart and say what it does and does not support.

That is a professional skill

Most people who work with data never learn the second half — the limits. It is what makes the first half trustworthy.

you can: get data from a file, an API or a database, and know what each door costs you find what is missing and decide what to do about it, with reasons summarise a distribution without being fooled by its shape tell the difference between a relationship and a coincidence write a claim that survives being checked by someone hostile release data without exposing the people in it
One Last Quick Check

Tap to reveal

A classmate's project: totals doubled after a merge, the "average" is quoted without spread, the headline says "proves", and one barangay's mean covers a single student. How many weeks of this course did that violate?

A · One — it's all the same mistake
B · Four — weeks 6, 7, 9 and 11, one each
C · None — data science is vibes
D · Impossible to tell without more charts

B — and you just diagnosed all four from one paragraph.

Unexplained row growth (week 6's merge audit), centre without spread (week 7), "proves" from observational data (week 9), and a group-of-one statistic (week 11). Being able to name the failures is what the course installed.

Part B

The Final Project

The whole pipeline, your dataset, your question.

You now have the whole workflow. This part applies it to data you choose: picking the dataset, sizing the question, and what exactly you hand in.

Your Skeleton

The week-12 lab is your project workbench

Not exercises — a scaffold. Four sections that mirror the arc, each with TODOs to replace with real work and answer cells for your written findings.

§1 · Understanding

Your question, your data's provenance, and an honest first look (shape, head, info).

§2 · Preparation

Gaps, outliers, merges, reductions — with every decision documented as you make it.

§3–4 · Explore & deliver

Univariate → multivariate → the comparison at your story's heart → one honest chart and one defensible sentence.

Your Project

One question. One dataset. One defensible claim.

Not the most impressive analysis — the most defensible one.

Choosing A Dataset

Four criteria, and "interesting" is not one

The most common project failure is choosing a dataset that cannot answer any question you care about, and discovering it in week three. Point 4 is the data's provenance: where it came from.

Check it can answer something FIRST

Before committing: load it, and write down one question it could answer. If you cannot, it is the wrong dataset, however interesting it looks.

1. it exists and you can load it (not "there's an API" — try it) 2. it has enough rows to group by something (hundreds, not dozens) 3. at least one numeric column and one categorical to split it by 4. you can state its provenance — who collected it, when, why # 4 is the one people skip and # then cannot write their method
Scoping

One question, finished, beats three, started

Your scope is how big the question is. You have weeks, not months. A complete small analysis demonstrates every skill in this course; an ambitious unfinished one demonstrates none of them.

The test of a good scope

You can state the question in one sentence and name the exact columns that answer it. If either is vague, narrow it further.

# too big "analyse Philippine education" # still too big "compare school performance by region" # scoped "Do schools in Region VII with more teachers per student report higher completion rates? (n=1,240 schools, DepEd 2024 data)" # one question. named columns. an n.
What You Hand In

Three artifacts, and they check each other

How It Is Assessed

Defensibility is weighted highest, on purpose

CriterionWhat earns full marks
DefensibilityEvery claim traceable to a computation; limits stated without being asked
CorrectnessThe seven checks were run; the notebook reproduces the numbers shown
MethodCleaning decisions justified, not just performed; denominators stated
CommunicationTitles state findings; charts labelled; three beats, not fifteen charts
EthicsProvenance given; privacy considered if personal; no claim beyond the evidence
The Rubric, In Plain Words

Each criterion, and how to check it yourself

CriterionIn plain wordsCheck it yourself by…
DefensibilityEvery number you say can be found in your notebook, and you say what the data cannot show before anyone asksPointing at each number on your slides and finding the cell that computed it; reading your limit sentence aloud
CorrectnessThe numbers are right, and rerunning gives the same onesRestart & Run All, then comparing every figure on the slides with the notebook; running the seven checks
MethodYou say why you cleaned each thing, and "out of how many" (the denominator) for every figureFinding a comment with a reason beside every drop or fill; asking "of what?" at every percentage
CommunicationA stranger gets your point from the titles and charts alonePasting your chart titles into a blank page: do they read as a summary? A five-second look by a friend
EthicsYou say where the data came from, protect any people in it, and claim no more than the evidencePublisher, date and licence under the chart; a k-audit (week 11) if people are in it; checking every verb (week 9)
First Decision

Choose a dataset you can finish with

Where To Get Data

Philippine sources that actually load

"There is an API" is not the same as "I downloaded it". Try the download before you commit to the question.

Check the licence too

You need to be able to say where it came from and whether you may republish it. That is part of your method section.

PSA (psa.gov.ph) census, labour, prices — the national statistics office data.gov.ph open data portal, mixed quality DepEd / DOH / DPWH portals sector data; often XLSX not CSV BetterGov.ph mirrors cleaned copies of government exports — this course's flood dataset came from here HDX, World Bank, HumData PH subsets, well documented
Judging A Source Before You Commit

Five minutes that saves three weeks

The Bar

Defensible beats impressive

A modest finding, cleanly derived, honestly caveated, and reproducible end-to-end outranks a spectacular claim with a shaky pipeline. Every time.

The reproducibility test

Runtime → Restart and run all. If it completes top-to-bottom on a fresh kernel (the Python program behind the notebook, with its memory wiped), a stranger can retrace you. That's §4's last checkbox for a reason.

The limits section is graded content

Sample size, confounders (a third factor that moves both things, week 8) you couldn't rule out, who the data misses — naming them is competence, not confession.

Commit early, commit often

The workbench says it; git (a tool that saves every version of your files; each saved version is a commit) makes your progress and your honesty visible.

A Plan That Finishes

Work backwards from the presentation

WhenDone by thenThe trap it avoids
NowQuestion written; dataset loaded; seven checks runDiscovering in week 3 that the data cannot answer it
+1 weekCleaning decisions made and written downRe-deciding the same thing three times
+2 weeksThe finding computed; the four-sentence claim draftedHaving analysis but no argument
+3 weeksThree charts, titles as sentences; rehearsed once aloudBuilding slides the night before
Before you hand inRestart & Run All on a fresh kernelA notebook that only runs on your machine
Peer Review Before Submission

Someone hostile, for twenty minutes

Peer review means a classmate checks your work before you submit. Swap projects with someone. Their job is not encouragement — it is to find the assumption you cannot see because you made it.

Give them the four questions

Unstructured "what do you think?" produces "looks good". Specific questions produce findings.

ask your reviewer to answer: 1. What population is this claim about? Can you tell from the slide alone? 2. What did they drop, and why? 3. Could the pattern be explained by something they did not check? 4. Which sentence goes further than the evidence? # question 4 finds the most. # it is almost always the verb.
"But My Finding Is Boring"

The most common week-11 panic

It is also usually wrong — and the fix is never to chase a better story.

A Null Result Is A Result

If it is checked, it is worth more than a shaky scoop

A null result is a checked finding that there is no clear difference or relationship. "Completion rates do not differ meaningfully across regions" is a finding. It is checkable, it is useful, and it will outscore a dramatic claim you could not verify.

The marking rewards defensibility

Look at the rubric again: novelty is not on it. A careful null result hits every criterion that is.

weak, exciting: "Region IX is being neglected!" (n=12, no comparison, causal verb) strong, quiet: "Completion rates range 65.6% to 92.0% across 18 regions (n=34,079). The spread is real but I could not identify a driver: rank correlations with median budget, project count and year are all below 0.2." # the second is publishable. # the first is a retraction.
Three Ways To Rescue A Flat Finding

Before you abandon it

Part C

Presenting

Eight minutes to make a stranger trust your discovery.

You learned to tell a data story in week 10. This part turns that into your talk: its outline, its timing, and the questions afterwards.

Structure

The week-10 arc is your presentation outline

1

Context — 1 min

The question and why it matters. No pipeline talk yet.

2

Journey — 2 min

Data source, the two hardest preparation calls, and why you made them.

3

Contrast — 3 min

The finding: your best chart, finding-title on it, number in your sentence.

4

Consequence — 1 min

Limits, what you'd check next, what someone should do with this.

The Talk

Eight minutes, five beats

Use the week-10 structure. The timing below is not a suggestion — it is what fits eight minutes. It is the same talk as the last slide with "the check" (how you tried to break your finding) given its own beat.

Rehearse once, out loud, timed

Silently in your head is always faster than reality. The first time you say it aloud should not be in front of the class.

1. the question ~1 min why it matters, what data 2. the data + one caveat ~1 min n, source, what it cannot say 3. the finding ~2 min ONE chart, title states it 4. the check ~2 min what you did to try to break it 5. the limit + question ~2 min what it does not show; what you would do next
Preparing For Questions

Predict three, and prepare the fourth answer

You can guess most of what you will be asked. The one you cannot guess is the one where the honest answer matters.

The three you will get

"How did you handle missing data?", "Could it be something else?", "How many rows?" — have all three ready with numbers.

Q: "Why did you drop those rows?" A: "1,240 of 8,900 had no score. They're concentrated in one region, so I kept them and analysed separately." Q: "Couldn't that be X instead?" A: "Yes — I checked by splitting on X and the effect held." ...or: "I couldn't rule it out. That's a limit." Q: anything you don't know A: "I don't know."
Five Ways Projects Go Wrong

All of them are avoidable this week

FailureHow it shows upPrevention
No questionA tour of the dataset with no argumentWrite the question before touching the data
Numbers that disagreeThe slide says 27,664; the notebook says 25,370Generate every figure from the notebook
Unstated denominator"the average" — of what population?Every figure gets its n
Claim exceeds evidenceA causal verb on correlational dataCheck the verb; week 9
Cannot be rerunFails on a fresh kernel or another machineRestart & Run All before submitting
The Hard Part

Questions are where trust is won

A panel probes exactly where the course taught you to probe yourself: merges, caveats, confounders, privacy. Arrive having already asked yourself their questions.

Know your soft spots

List your three weakest points before the talk. Two of them will be asked. Answer from the list, calmly.

The honest fallback

"I don't know — here's how I'd check" is a strong answer. Improvised certainty is the weak one.

Let the caveat lead

"With only ten weeks of data…" said by you sounds like rigor. Said by the panel first, it sounds like a catch.

On The Day

The logistics that go wrong

Every one of these has happened. None of them is about your analysis, and all of them cost you minutes you do not have.

Export to PDF as a backup

Live notebooks fail at the worst moment. A PDF of your three charts always opens.

- PDF backup on a USB and emailed to yourself - charts readable from the back (context="talk", not the default: week 10's one-line text scaling) - know your first sentence by heart — the nerves hit hardest there - eight minutes means eight. rehearse with a timer once. - do not open the notebook live unless the demo IS the point
Reading Feedback

Separate the two kinds

You will get comments on the analysis and comments on the communication. They need different responses, and conflating them wastes both.

"I did not understand" is always valid

You cannot argue someone into having understood. If a reader missed it, the chart or sentence failed — regardless of whether the analysis was right.

about the ANALYSIS: "did you check X?" -> go and check X. this is the useful kind; it is free review. about the COMMUNICATION: "I couldn't tell what the chart showed" -> not a defence opportunity. fix the title. about the SCOPE: "why didn't you do Y?" -> "out of scope, here's why" is a complete answer.
The Submission Checklist

Run it on a fresh clone, in a fresh environment

CheckHow to verify it
Notebook runs top to bottomRestart & Run All — on another machine if you can
Every figure on a slide comes from the notebookRegenerate them all; compare to what you pasted
Every number has its nRead your own slides and ask "of what?" at each figure
The limit is stated, unpromptedIt should appear before anyone asks
Sources cited: publisher, date, licenceUnder the chart, not in an appendix
No data or secrets committedgit ls-files — and remember .gitignore does not untrack
After DS 227

This course was the foundation layer

Where This Goes Next

What you are now prepared for

This course was deliberately the foundation layer. Everything above it assumes exactly what you just built.

The order that works

Statistics before machine learning. A model fit on data you have not interrogated is a faster way to be wrong.

The names, in a few words

SQL
the language for asking a database questions
inferential statistics
how far a sample's numbers hold for everyone
machine learning
programs that learn a prediction rule from data
causal inference
methods for claims about cause, not just association
immediately useful: SQL in depth inferential statistics — confidence intervals, tests experiment design builds on this directly: machine learning (CMSC 173) causal inference spatial analysis (you have coordinates now) the habit that carries: every dataset gets the seven checks, forever
If You Keep Only One Thing

Take the reflex, not the syntax

Your Turn · 8 min

Project workbench · the rest of the session

1 · State your question

One sentence, with the columns named. Show it to someone before you write any code.

2 · Load and run the seven checks

All of them, now. Better to find the trap today than the night before.

3 · Draft the four-sentence claim

Finding, number, method, limit. Even if the number is provisional.

4 · Sketch three charts

On paper. Titles as full sentences. If you cannot title it, you do not have the finding yet.

Quick Check

Tap to reveal

Your notebook computes 25,370 surviving rows. Your slide says 27,664. What is the correct response?

A · Use whichever supports the argument better
B · Find out why they differ — one of them is wrong, and the reason is usually a real bug
C · Average them
D · Note the discrepancy in a footnote and continue
B. This happened while building this session: the synthesis script used csv.DictReader (empty fields stay '') while the original used pd.read_csv (empty fields become NaN), so dropna() removed different rows. A discrepancy is information — it told us the loader changed the answer. Never reconcile by choosing.
The Course In Four Lines

If you keep only this

Glossary

This week's words, one line each

Words for today

pipeline
the fixed order of steps from raw data to answer
load
get the data into a table
prepare
fix types, gaps, outliers, joins; remove identifiers
explore
summarise and chart, then try to break the pattern
communicate
one checked finding as a sentence, a chart, a talk
reproducible
anyone can rerun it and get the same numbers
provenance
who collected the data, when, why, what licence
limitation (caveat)
what the data or method cannot show
deliverable
something you hand in and are graded on
rubric
the criteria your work is marked against
claim
finding, number, method, limit: four sentences
scope
how big the question is; one sentence, named columns
defensible
survives a sceptic: traced numbers, stated limits
null result
a checked "no clear difference": still a result
peer review
a classmate checks your work before you submit
Restart & Run All
wipe memory, run every cell from the top

Also used today (reminders)

KDD / CRISP-DM
the course's process / week 2's project cycle
kernel
the Python program behind a notebook
empty string ''
text with nothing in it; not NaN
assert
stop with an error if a condition is False
skew
lopsidedness; above 0 = long right tail
denominator
the "out of how many"
confounder
a third factor that moves both things
causal verb
"causes", "drives": stronger than correlation
statistical power
the chance your data could spot a real effect
git / commit
version saver / one saved version
fresh clone
a new copy of the project, as a stranger gets it
environment
the Python and packages where code runs
DS 227 · The End

Go discover something

Thank you for twelve weeks of honest work. The pipeline is yours now — point it at a question that matters.

DS 227 · Knowledge Discovery in Data · University of the Philippines Cebu