The whole knowledge-discovery arc in one view — and your final presentations to bring it home.
Knowledge Discovery in Data · University of the Philippines Cebu
Twelve weeks, one pipeline — what you can now do end to end.
The final workbench: expectations, dataset choice, and what "good" looks like.
Carrying your discovery through eight minutes and a Q&A.
Week 1 promised a full circuit: from a question, through raw and messy data, to something a stranger can trust. This week you run that circuit alone.
Nothing in the project needs a tool you haven't used. That's the point.
The fixed order of steps that turns raw data into an answer: load, prepare, explore, communicate.
e.g. one notebook that runs all four, top to bottom
Get the data into Python as a table you can work with.
e.g. df = pd.read_csv("flood.csv")
Fix what is wrong before analysing: types, missing values, outliers, joins, and privacy.
e.g. turn text dates into dates; drop the name column
Summarise and chart the data to find patterns, then check them.
e.g. describe(), a histogram, a correlation
Turn one checked finding into a sentence, a chart and a talk a stranger can trust.
e.g. a finding-title with its number and its limit
Someone else can rerun your work from the start and get the same numbers.
e.g. Restart & Run All finishes with no errors
Where the data came from: who collected it, when, why, and under what licence.
e.g. "DPWH export via the BetterGov.ph mirror", plus the date you downloaded it
A plain statement of what your data or method cannot show.
e.g. "no outcome column, so I cannot say if flooding fell"
Something you hand in and are graded on.
e.g. the notebook, the written claim, the talk
The list of criteria your work is marked against, with what earns full marks.
e.g. defensibility, correctness, method, communication, ethics
Your finding written as four sentences: the finding, its number, the method, the limit.
e.g. the week-9 template
How big the question is. A good scope is one sentence with the exact columns named.
e.g. not "analyse education" but "completion vs teachers per pupil, Region VII"
Able to survive a sceptical reader: every number traced to code, every limit stated.
e.g. "80.8% complete (n=34,079)", not "almost all complete"
A checked finding that there is no clear difference or relationship. Still a result.
e.g. "completion rates do not track budget size"
A classmate checks your work before submission, looking for what you missed.
e.g. "which sentence goes further than the evidence?"
Wipe the notebook's memory, then run every cell from the top. The honest test that it works.
e.g. Colab: Runtime → Restart and run all
What twelve weeks assembled.
You met every step one week at a time. This part lines them up into one workflow, load → prepare → explore → communicate, that your project follows.
Weeks 1–2 · KDD & CRISP-DM: the question comes first.
Weeks 3–4 · Scraping, APIs, and the manners of taking.
Weeks 5–6 · Gaps, outliers, merges, one honest table.
Weeks 7–8 · Describe, draw, correlate — carefully.
Weeks 9–11 · Findings, dashboards, ethics.
Read it left to right and it's your project plan. Read any single station and it's a week of this course. That correspondence is deliberate. (KDD, knowledge discovery in data, is the course's name for the whole line.)
CRISP-DM (week 2's standard project cycle) has arrows that go backward too: exploration exposes preparation flaws; findings reopen the question. Expect to revisit.
| Step | In plain words | What you used | The trap to watch |
|---|---|---|---|
| Load | Get the data into a table and look at it | pd.read_csv, an API, a scraper; shape, head() | A column that looks full but is empty (amount_paid) |
| Prepare | Fix types, gaps, outliers and joins; remove what identifies people | to_datetime, isna(), merge, drop | A cleaning step that quietly deletes a whole group |
| Explore | Summarise and chart; look for a pattern, then try to break it | describe(), groupby, histograms, corr() | A summary that silently skipped rows |
| Communicate | One checked finding, as a sentence and a chart, with its limit | finding-titles, sorted bars, the four-sentence claim | A verb stronger than the evidence ("proves", "causes") |
It is cooking: buy the ingredients (load), wash and chop (prepare), taste as you go (explore), then plate and serve (communicate). Skip the washing and the plating will not save the meal.
Write the question first (weeks 1–2). Every step above is chosen by what that question needs.
Everything you learned, you learned on this.
| Week | The technique | The trap it exposed |
|---|---|---|
| 4 | APIs, JSON, pagination | amount_paid is 0 on every row — unpopulated, not unpaid |
| 5 | Missing data, outliers | dropna() removed 25.6% of rows and every unfinished project |
| 6 | Joins, types, derived columns | A project that finished 17 days before it started |
| 7 | Descriptive statistics | describe() silently excluded 6,415 rows |
| 8 | Correlation, visualisation | r = −0.785 between latitude and longitude |
| 9 | Data journalism | A ₱75.2B headline the data could not support |
| 10 | Storytelling, dashboards | A partial year drawn as a 47% collapse |
| 11 | Ethics, privacy | One rounded column made 30.9% of records unique |
Eight techniques, eight ways to be confidently wrong. The traps were the syllabus. The techniques are in every textbook; knowing where each one lies to you is what you actually acquired.
Not a summary of the course — the actual pipeline, shortened here. It runs in about a second and reproduces every number you were shown.
Load, convert, check, derive, summarise, relate, protect. The order matters and you now know why each step precedes the next.
pd.read_csv(path)load: read the CSV file into a DataFramepd.to_numeric(..., errors="coerce")turn text into numbers; anything unreadable becomes NaN (missing) instead of an error(df.completion_date - df.start_date).dt.daysdays between two dates (the full script first converts both columns with pd.to_datetime, week 6)df.isna().sum()how many missing cells in each column... and the # w7-8 linessteps left out to fit the slide; each is a week's labIf any of those numbers ever disagrees with a slide, the slide is wrong. That is what reproducibility buys you — not elegance, but the ability to be corrected. Reproducible means anyone can rerun the script from the start and get these same numbers.
| Line | What it means |
|---|---|
amount_paid unique [0] | Every row holds 0: the column was never filled in, so it says nothing about payment |
naive dropna() 25,370 (25.6% lost) | dropna() deletes any row with an empty cell; a quarter of the data went |
completed, after drop 99.6% | Unfinished projects have no end date, so dropping gaps dropped them: 80.8% became 99.6% |
budget skew 5.13 | Skew above 0 means a long right tail: most budgets are small, a few are huge |
pearson / spearman 0.365 / 0.508 | Two correlation scores (week 8): 0 = no relationship, 1 = perfect; bigger budgets tend to take longer |
unique at 50M bands 0.4% | With ₱50M budget bands and region instead of province, 0.4% of projects stand alone (week 11) |
Each line is one week's lesson reduced to a number you can check. The $ in $ python pipeline.py is the terminal's prompt, not something you type.
The first run of that script reported 27,664 rows surviving dropna(). Week 5 had said 25,370. One of them was wrong.
csv.DictReader(f)Python's built-in CSV reader: each row becomes a dictionary of text''an empty string: text with nothing in it. It is a value, so dropna() keeps itpd.read_csv(f)pandas' reader: an empty cell becomes NaN, the missing-value marker, which dropna() removescsv.DictReader leaves an empty field as ''; pd.read_csv makes it NaN. Only the second is dropped — so the loader changed the answer.
This table is your project cheat-sheet. Every row is muscle memory by now — the labs made sure.
r.json()clipmerge / concatgroupby / pivot_table| Need to… | Reach for |
|---|---|
| Fetch data | requests, BeautifulSoup, r.json() |
| Audit & clean | isna().sum(), IQR fences, clip |
| Combine & shape | merge, concat, groupby, pivot_table |
| Explore | describe, histograms, corr, scatters |
| Communicate | small multiples, sorted bars, finding-titles |
| Protect | k-anonymity audit, banding, suppression |
Status code before body. Row counts around merges. describe() before
belief. Plots before conclusions.
Drop vs fill, keep vs cap, blur vs publish — every choice logged, every choice defensible.
Association not causation; number and caveat in the sentence; titles that report.
Polite requests, purpose-bound data, no group-of-one statistics, bias named in the report.
Libraries will change; these won't. A discoverer with these habits is trustworthy with any toolset.
Your final work is graded through exactly these four habits.
| Check | From | Cost |
|---|---|---|
.unique() on the column your claim rests on | w4 | 5 seconds |
| Row count before and after every filter | w5 | 2 lines |
| Range check on every derived column | w6 | 1 assert |
Compare count to len(df) | w7 | 1 line |
| Recompute the relationship within groups | w8 | 1 groupby |
| Check the verb matches the evidence | w9 | 1 reading |
| Compute k before releasing | w11 | 1 groupby |
None takes more than a minute. Together they are the difference between an analysis someone can rely on and one that happens to be right. Run all seven on your project before you present it.
assertassert (df.duration >= 0).all().unique()Twelve weeks ago a 34,000-row government dataset was something you read about. You can now take one apart and say what it does and does not support.
Most people who work with data never learn the second half — the limits. It is what makes the first half trustworthy.
A classmate's project: totals doubled after a merge, the "average" is quoted without spread, the headline says "proves", and one barangay's mean covers a single student. How many weeks of this course did that violate?
B — and you just diagnosed all four from one paragraph.
Unexplained row growth (week 6's merge audit), centre without spread (week 7), "proves" from observational data (week 9), and a group-of-one statistic (week 11). Being able to name the failures is what the course installed.
Run the same diagnostic on your own project before anyone else gets the chance.
Review your work as the course's toughest grader. Then submit.
The whole pipeline, your dataset, your question.
You now have the whole workflow. This part applies it to data you choose: picking the dataset, sizing the question, and what exactly you hand in.
Not exercises — a scaffold. Four sections that mirror the arc, each with TODOs to replace with real work and answer cells for your written findings.
Your question, your data's provenance, and an honest first look
(shape, head, info).
Gaps, outliers, merges, reductions — with every decision documented as you make it.
Univariate → multivariate → the comparison at your story's heart → one honest chart and one defensible sentence.
Not the most impressive analysis — the most defensible one.
The most common project failure is choosing a dataset that cannot answer any question you care about, and discovering it in week three. Point 4 is the data's provenance: where it came from.
Before committing: load it, and write down one question it could answer. If you cannot, it is the wrong dataset, however interesting it looks.
Your scope is how big the question is. You have weeks, not months. A complete small analysis demonstrates every skill in this course; an ambitious unfinished one demonstrates none of them.
You can state the question in one sentence and name the exact columns that answer it. If either is vague, narrow it further.
Runs top to bottom after Restart & Run All. Every decision has a comment saying why, not what.
Finding, number, method, limit — four sentences, the week-9 template.
Three beats, three charts, eight minutes. The week-10 structure.
Each is a deliverable: something you hand in and are graded on. If the three disagree — the notebook computes one number and the slide shows another — that is the single most damaging thing a reviewer can find. Regenerate the charts from the notebook, never by hand.
The due date, how to submit and your presentation slot are on the course page. If anything there differs from this deck, the course page wins.
| Criterion | What earns full marks |
|---|---|
| Defensibility | Every claim traceable to a computation; limits stated without being asked |
| Correctness | The seven checks were run; the notebook reproduces the numbers shown |
| Method | Cleaning decisions justified, not just performed; denominators stated |
| Communication | Titles state findings; charts labelled; three beats, not fifteen charts |
| Ethics | Provenance given; privacy considered if personal; no claim beyond the evidence |
A rubric is the list of criteria your work is marked against. Notice what is not on this list: novelty, technique count, dataset size. A careful analysis of 400 rows outscores a careless one of four million.
| Criterion | In plain words | Check it yourself by… |
|---|---|---|
| Defensibility | Every number you say can be found in your notebook, and you say what the data cannot show before anyone asks | Pointing at each number on your slides and finding the cell that computed it; reading your limit sentence aloud |
| Correctness | The numbers are right, and rerunning gives the same ones | Restart & Run All, then comparing every figure on the slides with the notebook; running the seven checks |
| Method | You say why you cleaned each thing, and "out of how many" (the denominator) for every figure | Finding a comment with a reason beside every drop or fill; asking "of what?" at every percentage |
| Communication | A stranger gets your point from the titles and charts alone | Pasting your chart titles into a blank page: do they read as a summary? A five-second look by a friend |
| Ethics | You say where the data came from, protect any people in it, and claim no more than the evidence | Publisher, date and licence under the chart; a k-audit (week 11) if people are in it; checking every verb (week 9) |
The weights of each criterion are not in this deck; they are on the course page. What the table promises is the direction: defensibility first.
The right-hand column is a to-do list for the night before. Every row you can tick is marks you already have.
You'll spend hours inside it. Local data — transport, weather, prices, sports — keeps you honest and interested.
Weeks 4 and 11 apply in full: permitted source, no personal data you can't protect, k-audit before sharing.
Enough rows to hold structure (hundreds+), small enough to understand every column. Complexity is not a rubric item.
Good hunting grounds: PSA and open-data portals, city open records, public APIs from week 4, or a page that passes the week-3 scraping checks.
If loading the data takes a week, the dataset is your project's first finding: pick another.
"There is an API" is not the same as "I downloaded it". Try the download before you commit to the question.
You need to be able to say where it came from and whether you may republish it. That is part of your method section.
Download it now. A portal that 403s scripted clients or serves a broken XLSX is a project-ending problem in week three.
If no column dictionary exists, you will be guessing what fields mean — and week 4 showed where that ends.
Publisher, date accessed, licence. If you cannot write those three, you cannot write your method.
The flood dataset passed all three, which is why the course used it. The upstream DPWH portal returns 403 to scripted clients — the mirror is what made it usable, and knowing that difference is itself part of the provenance.
A modest finding, cleanly derived, honestly caveated, and reproducible end-to-end outranks a spectacular claim with a shaky pipeline. Every time.
Runtime → Restart and run all. If it completes top-to-bottom on a fresh kernel (the Python program behind the notebook, with its memory wiped), a stranger can retrace you. That's §4's last checkbox for a reason.
Sample size, confounders (a third factor that moves both things, week 8) you couldn't rule out, who the data misses — naming them is competence, not confession.
The workbench says it; git (a tool that saves every version of your files; each saved version is a commit) makes your progress and your honesty visible.
| When | Done by then | The trap it avoids |
|---|---|---|
| Now | Question written; dataset loaded; seven checks run | Discovering in week 3 that the data cannot answer it |
| +1 week | Cleaning decisions made and written down | Re-deciding the same thing three times |
| +2 weeks | The finding computed; the four-sentence claim drafted | Having analysis but no argument |
| +3 weeks | Three charts, titles as sentences; rehearsed once aloud | Building slides the night before |
| Before you hand in | Restart & Run All on a fresh kernel | A notebook that only runs on your machine |
Notice the finding is due at the two-thirds mark. The last third is for communicating it and checking it — which is where most of the marks are, and what always gets squeezed.
Peer review means a classmate checks your work before you submit. Swap projects with someone. Their job is not encouragement — it is to find the assumption you cannot see because you made it.
Unstructured "what do you think?" produces "looks good". Specific questions produce findings.
It is also usually wrong — and the fix is never to chase a better story.
A null result is a checked finding that there is no clear difference or relationship. "Completion rates do not differ meaningfully across regions" is a finding. It is checkable, it is useful, and it will outscore a dramatic claim you could not verify.
Look at the rubric again: novelty is not on it. A careful null result hits every criterion that is.
A relationship that is absent overall may be strong within groups. Week 8: budget vs duration ranged 0.09 to 0.62 by region.
Counts hide rates. Region III looks dominant by total and ordinary per project.
"No relationship" plus what you checked and how much power you had (how likely your data was to spot an effect if one existed; more rows, more power) is a complete result.
What does not work is starting over in the final week. A finished, honest analysis of a modest question beats an unfinished one of an exciting question, every time, by design.
Eight minutes to make a stranger trust your discovery.
You learned to tell a data story in week 10. This part turns that into your talk: its outline, its timing, and the questions afterwards.
The question and why it matters. No pipeline talk yet.
Data source, the two hardest preparation calls, and why you made them.
The finding: your best chart, finding-title on it, number in your sentence.
Limits, what you'd check next, what someone should do with this.
One chart per point, every title a finding, every claim caveat-sized. Your slides should pass the five-second test a panel actually gives them. These four add up to 7 minutes, leaving one for slack; check your official slot on the course page.
If woken at 3 a.m., you should still produce your one-sentence finding, number and all.
Use the week-10 structure. The timing below is not a suggestion — it is what fits eight minutes. It is the same talk as the last slide with "the check" (how you tried to break your finding) given its own beat.
Silently in your head is always faster than reality. The first time you say it aloud should not be in front of the class.
You can guess most of what you will be asked. The one you cannot guess is the one where the honest answer matters.
"How did you handle missing data?", "Could it be something else?", "How many rows?" — have all three ready with numbers.
| Failure | How it shows up | Prevention |
|---|---|---|
| No question | A tour of the dataset with no argument | Write the question before touching the data |
| Numbers that disagree | The slide says 27,664; the notebook says 25,370 | Generate every figure from the notebook |
| Unstated denominator | "the average" — of what population? | Every figure gets its n |
| Claim exceeds evidence | A causal verb on correlational data | Check the verb; week 9 |
| Cannot be rerun | Fails on a fresh kernel or another machine | Restart & Run All before submitting |
A denominator is the "out of how many"; a causal verb ("causes", "drives", "leads to") claims more than a correlation can show. The second row, numbers that disagree, is the most damaging and the easiest to avoid. It is also exactly the bug that turned up while writing this session’s own synthesis script — nobody is above it.
A panel probes exactly where the course taught you to probe yourself: merges, caveats, confounders, privacy. Arrive having already asked yourself their questions.
List your three weakest points before the talk. Two of them will be asked. Answer from the list, calmly.
"I don't know — here's how I'd check" is a strong answer. Improvised certainty is the weak one.
"With only ten weeks of data…" said by you sounds like rigor. Said by the panel first, it sounds like a catch.
Every one of these has happened. None of them is about your analysis, and all of them cost you minutes you do not have.
Live notebooks fail at the worst moment. A PDF of your three charts always opens.
You will get comments on the analysis and comments on the communication. They need different responses, and conflating them wastes both.
You cannot argue someone into having understood. If a reader missed it, the chart or sentence failed — regardless of whether the analysis was right.
| Check | How to verify it |
|---|---|
| Notebook runs top to bottom | Restart & Run All — on another machine if you can |
| Every figure on a slide comes from the notebook | Regenerate them all; compare to what you pasted |
| Every number has its n | Read your own slides and ask "of what?" at each figure |
| The limit is stated, unprompted | It should appear before anyone asks |
| Sources cited: publisher, date, licence | Under the chart, not in an appendix |
| No data or secrets committed | git ls-files — and remember .gitignore does not untrack |
Six checks. They are the same discipline as the seven analytical ones, applied to the artifact instead of the data. Run them the day before, not the hour before.
git ls-files.gitignoreEverything here feeds models: clean tables, scaled features, honest evaluation. Your ML courses start where week 6 ended.
Weeks 9–10 scale into a discipline of their own — dashboards, visual essays, decision support.
When the pipeline must run every night at scale, the same stations become systems: ingestion, warehouses, orchestration.
Whichever branch you take, the discovery discipline travels with you — it's the part employers can't teach in onboarding.
Twelve working notebooks are a reference library you wrote yourself. Future-you will search them.
This course was deliberately the foundation layer. Everything above it assumes exactly what you just built.
Statistics before machine learning. A model fit on data you have not interrogated is a faster way to be wrong.
Even pandas rewrites its own rules between versions: pandas 3 changes how copies of a table behave. Libraries move; the questions do not.
"What is this column actually recording?" and "what did I just exclude?" work on any dataset, in any language, forever.
Most people producing charts have never asked either. That is the gap you now sit on the right side of.
Twelve weeks ago the ₱1.17 trillion headline (week 9: "completed but never paid", resting on an empty column) would have looked like a scoop. Knowing why it was not — and checking in five seconds — is the whole course.
One sentence, with the columns named. Show it to someone before you write any code.
All of them, now. Better to find the trap today than the night before.
Finding, number, method, limit. Even if the number is provisional.
On paper. Titles as full sentences. If you cannot title it, you do not have the finding yet.
Leave today with a question, a loaded dataset and three chart titles. That is most of the work; the rest is execution.
Your notebook computes 25,370 surviving rows. Your slide says 27,664. What is the correct response?
csv.DictReader (empty fields stay '') while the original used pd.read_csv (empty fields become NaN), so dropna() removed different rows. A discrepancy is information — it told us the loader changed the answer. Never reconcile by choosing.Discovery is aimed, not stumbled into.
In the fetching, cleaning and merging nobody applauds — where the errors are born.
Claim, number, caveat — everything else is supporting material.
Take gently, protect thoroughly, publish what you'd defend to their face.
One sentence for the whole course: knowledge discovery is a chain of small honest decisions that ends in something a stranger can trust.
The presentations close the loop. Make us trust it.
''assertThank you for twelve weeks of honest work. The pipeline is yours now — point it at a question that matters.
DS 227 · Knowledge Discovery in Data · University of the Philippines Cebu