From raw, messy data to valid, novel, and useful knowledge — the frameworks, craft, and ethics of finding what data has to say.
3 units · Blended learning (full online)
Two 10-minute breaks separate the hours — the break slides will tell you when.
The quizzes and the Hour-3 discussion are graded for engagement, not correctness. Wrong answers are welcome data.
KDD = Knowledge Discovery in Data. A precise definition of "knowledge discovery," and why it is more than running an algorithm.
The five-stage pipeline from raw data to interpreted knowledge — and why it loops.
How our 12 weeks map onto that pipeline, plus tools, logistics, and expectations.
By the end of today you should be able to explain what separates knowledge discovery from mere computation.
"When does a number become knowledge?" — we return to it at the end.
You will hear these all semester. Each one is explained again on the slide where it first matters, and they are all in the glossary at the end.
Recorded facts — numbers, text, dates, clicks — before anyone has interpreted them.
e.g. “7/1: 3 sachets kape” in a store’s notebook
A collection of data about one topic, usually arranged as a table.
e.g. two years of clinic appointments, saved as one file
One entry in a dataset: everything recorded about one thing or event. One row of the table.
e.g. one appointment: patient, date, attended or not
One property that every record has — one column of the table.
e.g. age, barangay, number of visits
The data-mining word for an attribute, especially one you made to help find a pattern.
e.g. “distance to the clinic”, worked out from the address
A cell where nothing was recorded. In code it shows up as NaN (“not a number”).
e.g. a patient with no contact number on file
No need to memorise them now. Read them once so none is a surprise later today.
A regularity that shows up across many records.
e.g. “no-shows rise on clinic days after heavy rain”
A pattern that has passed four tests — valid, novel, useful, understandable — so someone can act on it.
e.g. “remind far-away patients when rain is forecast”
Knowledge Discovery in Data: the whole multi-step process from raw data to knowledge.
e.g. choose the data, clean it, reshape it, search it, judge the result
The one step inside KDD where a method searches the data for patterns.
e.g. sorting shoppers into groups that buy alike
A fixed recipe of steps a computer follows, the same way every time.
e.g. “sort the list, then take the middle value”
A simplified summary of the data that describes a pattern or predicts an outcome.
e.g. the rule “lives over 4 km away + rain → likely to miss”
A chain of stages where each stage’s output is the next stage’s input.
e.g. raw file → cleaned table → summary → chart
Every jeepney GPS ping, GCash payment, weather station reading, and census form adds to a pile growing far faster than anyone can read it.
Storage grows exponentially (doubling again and again); human attention doesn't. Whatever is not processed into knowledge is effectively lost.
Picture a store notebook that fills a new page every second. Nobody can read it all, so most of it is never used — unless we have a method for finding what matters in it.
Sales, payments, enrollments, bookings — deliberately recorded, usually clean-ish, the classic raw material for data mining.
GPS traces, weather stations, server logs, CCTV — recorded as a side effect, enormous volume, rarely designed for analysis.
Posts, reviews, news articles, photos — rich in meaning, messy in structure. Weeks 3–4 teach you to harvest it.
Census, surveys, administrative registers — PSA OpenSTAT (the Philippine Statistics Authority’s free online tables) is our recurring example. Curated, documented, but aggregated: already added up into totals, so the individual records are gone.
Notice: only the last category was collected to be analyzed. Most data is exhaust — a by-product of some other activity — which is why preprocessing (cleaning and repairing data before analysis) gets four weeks of this course.
Which of the four quadrants does your research data live in? Keep the answer for the Hour-3 discussion.
| Shape | What it looks like | Example |
|---|---|---|
| Structured | Rows × labeled columns; a schema fixed in advance | Enrollment table, POS sales, OpenSTAT extracts |
| Semi-structured | Tagged and nested, but flexible — structure travels with the data | JSON from an API, HTML pages, XML feeds |
| Unstructured | No inherent fields; meaning must be extracted | News articles, complaint emails, photos, audio |
The KDD pipeline's early stages exist largely to convert the bottom two rows into the top one — analysis-ready tables.
Week 3 scrapes semi-structured HTML; Week 10's storytelling starts from structured summaries. You'll cross this table constantly.
<p> (Week 3){"name": "Ana", "visits": [3, 5]}A definition worth memorizing — one word at a time.
You have met the raw material: records and attributes in a table. This part asks when a pattern in that table deserves the word knowledge.
"The non-trivial process of identifying valid, novel, potentially useful, and ultimately understandable patterns in data."— Fayyad, Piatetsky-Shapiro & Smyth, "From Data Mining to Knowledge Discovery in Databases," AI Magazine, 1996
Nearly 30 years old, and still the reference definition. Every adjective in it rules something out. Non-trivial means it takes real work — not a single lookup or one sum.
KDD is not a single algorithm or query — it is a multi-stage, human-guided process.
Finding knowledge is like panning for gold. The river gives you plenty of sand (patterns); the four underlined words are the sieve that keeps only the gold.
The pattern holds on new data, not just the sample (the part of the data you happened to study). Guards against overfitting (learning quirks of this sample that will not repeat) and coincidence.
It tells us something we didn't already know. "Sales rise in December" is valid but not a discovery.
Someone can act on it — a decision, a policy, a product. Usefulness is judged by the domain, not the algorithm.
A human can grasp why it holds — immediately or after some extra explaining. A black-box score (a number from a method whose reasoning nobody can see) alone doesn't qualify.
A result that fails any one filter is not yet knowledge — it's a lead to investigate.
Be ready to judge a given "finding" against all four criteria.
The textbook classic: ice cream sales and drowning deaths rise and fall together, month after month. Strong pattern, real data — and utterly wrong as a causal claim.
Hot weather drives both. A pattern can be statistically real yet invalid as knowledge if it evaporates once you account for context.
Treat every exciting pattern as the start of an interrogation, never the end of one. The four filters are that interrogation.
DIKW stands for Data → Information → Knowledge → Wisdom. Each layer adds context and meaning. KDD is the machinery that moves us up the pyramid.
Between data and knowledge — turning recorded facts into patterns we can act on.
"7/1: 3 sachets kape · 7/1: 2 load ₱20 · 7/2: 5 sachets kape …" — raw entries in the tindera's notebook, one per sale.
Summarized and contextualized: "Coffee sachet sales double on weekdays before 7 AM."
A validated, actionable pattern: "Commuters buy coffee on the way to work — stock up and open earlier." That's discovery.
Scale that notebook up to millions of rows and the manual route breaks — which is exactly why KDD exists.
Knowing when to trust the pattern — e.g., not during a holiday week. Judgment stays human.
A mall's analytics team reports: "Customers who buy umbrellas are 340% more likely to also buy pancit canton during the same typhoon week. Recommendation: bundle them year-round."
Commit to an answer, then click and hold (or tap) to reveal
A — a validity failure.
The pattern is real within typhoon weeks (people stock up on storm supplies together), but the recommendation extrapolates it year-round. Tested on ordinary weeks, it would likely vanish. It's arguably not novel either — but the fatal flaw is validity of the generalization.
Stretch, hydrate, check nothing. When we return: the five-stage process that turns raw data into defensible knowledge.
Five stages, many loops.
You now know what counts as knowledge. This part adds the process that produces it — the pipeline from a raw file to a finding you can defend.
Practitioners consistently report that stages 1–3 consume the large majority of project time. Mining is the short, glamorous part.
Weeks 3–6 live in stages 1–3; weeks 7–10 in stages 4–5.
KDD is like cooking a meal: pick the ingredients (selection), wash them (preprocessing), chop them (transformation), cook (mining), then taste before serving (interpretation). If it tastes wrong, you go back a step.
Before any analysis: which data, and can we trust it?
"If this field is wrong, would my conclusion change?" If yes, it needs cleaning attention.
Reshape the cleaned data so patterns become findable — then find them.
Classification · regression · clustering · association rules · anomaly detection · summarization
Lots of new words here. Every one of them is decoded, in plain language, on the next slide.
Transformation is chopping the ingredients into a usable shape; mining is choosing the recipe that fits the dish you were asked for. A fancy recipe cannot rescue badly chopped ingredients.
Test the four filters: valid on held-out data? genuinely novel? useful to the stakeholder? explainable?
Translate into the stakeholder's language (the person or group who will use the result) — charts, narratives, dashboards (screens of live charts). (Weeks 9–10 are devoted to this.)
Most first findings fail a filter. Loop back: reselect, re-clean, re-transform, re-mine.
A pattern nobody understands or acts on is a computation, not a discovery.
Expect to traverse the loop several times per project — plan your time budget for it.
A barangay health center notices many patients miss their follow-up appointments. The captain asks: "Can the data tell us who is likely to miss, so we can send reminders where they matter?"
Small, local, humane — and it exercises every stage, every pitfall, and every ethical question this course covers. We'll revisit it all semester.
Not "predict no-shows" — it's "find valid, novel, useful, understandable patterns in who misses and why, so a human can design a better reminder policy."
Every bullet above is a judgment call that changes the final answer — which is why these stages can't be delegated to a script (a file of code) you didn't write.
Dropping rows with missing contacts silently removes the hardest-to-reach patients. Is that a statistics problem or an ethics problem? (Both.)
Engineer features: distance-to-center from address, days-since-last-visit, rainy-season flag, prior no-show count. Raw columns rarely carry the signal; derived ones do.
Start simple: cross-tabs (counts of one column against another) and a decision tree (a model made of yes/no questions, like a flowchart). Suppose it surfaces: "no-shows concentrate among patients >4 km away, on clinic days after heavy rain."
Valid? Check next quarter. Novel? Staff suspected distance, not the rain interaction. Useful? Reschedule far patients around forecasts. Understandable? Completely.
And the loop: the rain finding sends us back to Stage 1 — we now want PAGASA rainfall data joined in (matched to each appointment by date), which restarts selection and cleaning for a new source.
Real KDD projects are spirals, not lines. Budget for at least two full passes.
| Stage | Classic pitfall | Defense |
|---|---|---|
| Selection | Sampling bias — the data you have isn't the population (everyone you want to describe) you claim | Ask "who is missing from this table, and why?" |
| Preprocessing | Silent deletion of inconvenient rows | Cleaning journal; report row counts before/after |
| Transformation | Leakage — a feature that secretly contains the answer | Ask "would I know this before the outcome?" |
| Mining | Overfitting — memorizing noise as pattern | Hold-out data; prefer simpler models first |
| Interpretation | Storytelling past the evidence | The four filters, applied in writing |
Each of these gets a full treatment in its home week — today you only need to recognize their names.
It doubles as a review checklist for your final project's methodology section.
| Term | What it names | Relation to KDD |
|---|---|---|
| KDD | The whole discovery process, end to end | The umbrella — this course |
| Data mining | Pattern-extraction algorithms | Stage 4 of the KDD process |
| Machine learning | Algorithms that improve from data | Supplies many mining tools |
| Data science | The broad discipline & profession | KDD is its methodological core |
| Analytics / BI (business intelligence) | Reporting & decision support in orgs | Consumes KDD outputs |
In industry these words blur; in this course we keep them distinct so we can talk precisely about where in the process a problem lives.
"Data mining" used to mean the entire process — context usually disambiguates, but be precise in your writing.
You compute "days until next typhoon" as a feature for predicting appointment no-shows — using PAGASA's after-the-fact storm records. Which stage are you in, and what pitfall are you committing?
Commit to an answer, then click and hold (or tap) to reveal
B — transformation-stage leakage.
Feature engineering happens in Stage 3, and "days until next typhoon" uses information nobody could have known at prediction time. The model would score brilliantly in testing and fail in deployment — the signature symptom of leakage.
Final hour when we return: real applications, a 170-year-old case study, ethics — and what this course asks of you.
Who discovers what — and at what risk.
You know the definition and the five stages. This part watches them at work in the real world, where a wrong finding lands on real people.
Market-basket analysis (what gets bought together), churn prediction (which customers are about to leave), credit scoring, fraud detection in e-wallet transactions.
Disease surveillance (tracking cases as they are reported), genomics (DNA data), drug-interaction discovery, outbreak early warning.
Census & poverty mapping (PSA OpenSTAT), disaster response, and data journalism that holds power to account.
In this course, examples and projects lean on open Philippine data — PSA, PAGASA, DOH, and public web sources you'll scrape yourself.
Graduate students: bring data from your own field for the final project.
During London's 1854 cholera outbreak, physician John Snow mapped death locations by hand and saw them cluster around the Broad Street water pump — evidence against the reigning "bad air" theory. The pump handle was removed; the outbreak subsided.
Data collection, cleaning (deaths misattributed to other addresses), a visualization, an inference, an action — the whole KDD loop, 170 years early.
Every dashboard tracking dengue clusters or traffic fatalities is Snow's map with better plumbing. Weeks 9–10 teach you to build the map and the argument.
The same techniques that map poverty can profile individuals. The Philippines' Data Privacy Act of 2012 (RA 10173) sets legal duties for anyone processing personal data (any information that can identify a person) — including researchers.
Privacy law, anonymization (removing names and identifiers) and its limits, algorithmic bias (a method that treats some groups unfairly), and responsible scraping — but ethics threads through every week, starting with how we scrape in Week 3.
A telco's analytics team discovers that prepaid users who let their load expire twice in a month are 5× more likely to default on the company's new micro-loan product. They propose auto-rejecting those applicants.
Your job: prosecute or defend this as "knowledge" using the four filters and the three ethics questions from earlier.
We'll hear two groups aloud — one prosecuting, one defending — then move to the course roadmap.
That's deliberate. Graduate-level KDD is the discipline of defending judgment calls in public.
Twelve weeks along the pipeline.
You know the five stages. This part shows how the twelve weeks walk you through them, one stage at a time — and what the course asks of you.
| Weeks | Theme | Topics |
|---|---|---|
| 1–2 | KDD Foundations | This intro · CRISP-DM, SEMMA & process frameworks |
| 3–6 | Scraping & Preprocessing | HTML parsing · APIs & scraping ethics · cleaning · transformation |
| 7–8 | Data Exploration | Descriptive statistics · univariate & multivariate EDA · visualization |
| 9–10 | Journalism & Storytelling | Finding stories in data · narrative, dashboards & communication |
| 11–12 | Ethics & Synthesis | Privacy & responsible analytics · final presentations |
Notice the arc: it is exactly the KDD pipeline from Part 2, walked slowly with your hands on real data.
Weekly exercises build toward a final discovery project presented in Week 12.
Blended learning, fully online: recorded lectures + live (synchronous) discussions. 3 units, 3 contact hours/week.
Python (a programming language) and its libraries (bundles of ready-made code): pandas for tables, requests for fetching pages, BeautifulSoup for reading HTML, matplotlib for charts. Jupyter notebooks and this portal.
Graded (GRD): weekly hands-on exercises, a midterm synthesis, and the final discovery project.
No prerequisites are assumed beyond graduate standing — but comfort with basic Python will make Weeks 3+ far smoother. The slides teach every line of code they use, one line at a time.
DS 208 (Programming for Data Science) pairs naturally with this course.
| Component | Weight | What it rewards |
|---|---|---|
| Weekly exercises (best 8 of 10) | 40% | Consistent hands-on practice; two free passes for busy weeks |
| Midterm synthesis (Week 7) | 20% | Integrating scraping → cleaning → exploration on a fresh dataset |
| Final discovery project | 30% | Full KDD cycle on data from your field, defended in Week 12 |
| Participation | 10% | Quizzes, forum activity, discussion contributions |
Project proposals are due Week 6 — start auditioning datasets now; this week's assignment is the first step.
Standard graded course (GRD). Rubrics for each component are posted on the course page.
Small on purpose — the goal is to start seeing your world as data.
Browse PSA OpenSTAT or any open portal from your field. Pick one dataset that genuinely interests you.
What question might it answer? Which of the four filters (valid, novel, useful, understandable) would be hardest to satisfy, and why?
Install instructions are on the course page. We start writing code in Week 3; the weekly labs need no install at all — they run in your browser.
You know the words: dataset, record, attribute, missing value. The lab shows what they look like when Python opens a small clinic table — the first move of Stage 2 (preprocessing). The next two slides read that code aloud, line by line.
import io, pandas as pdload two libraries (ready-made code): pandas, for tables, and io; as pd gives pandas a short nicknameraw = """…"""a variable (a name that holds a value) called raw, holding the table as plain text: commas between columns, one record per linepd.read_csv(io.StringIO(raw))read that text as if it were a CSV file; the result is a DataFrame, pandas’ table, stored in dfprint(df)show the table underneath the codeNaN is the missing value: Talamban has no visits recorded. Visits show .0 because a column with a gap is stored as decimals.A DataFrame is a spreadsheet you control with code: each record is a row, each attribute is a named column.
How big is it? What is missing? What is a typical value? Same
df as the previous slide.
df.shapethe size, as (rows, columns): (5, 4) means 5 clinics, 4 attributesdf["visits"]pick one column by its name, in square brackets and quotes.isna().sum().isna() marks each missing cell True; .sum() counts the Trues. 1 = one clinic has no visits number.mean() / .median()the average / the middle value. pandas skips the missing cell, so both use 4 clinics.name() is a method: a function (a named piece of work) that belongs to a value, called with a dot and bracketsOne busy clinic (1,240 visits) drags the mean up to 623.75; the median, 433.5, stays close to a typical clinic. It is like a millionaire walking into a sari-sari store: the average customer gets richer, the middle customer does not.
Replace each ____ with your answer, press ▶ Run to see
the output, then Check. Two multiple-choice questions ask what the numbers
mean.
Valid, novel, useful, understandable — a finding must pass all four.
Five stages with feedback — and most of the work happens before mining.
Real people live inside our datasets; ethics is a stage-zero concern.
Back to the opening question: a number becomes knowledge when it survives the four filters and someone can act on it.
Re-watch Part 1, then bring your questions to the discussion forum.
NaNraw = "…".mean()Fayyad, U., Piatetsky-Shapiro, G., & Smyth, P. (1996). From Data Mining to Knowledge Discovery in Databases. AI Magazine, 17(3). — short, readable, and the source of this week's definition.
Han, Kamber & Pei, Data Mining: Concepts and Techniques — Ch. 1. · Republic Act No. 10173 (Data Privacy Act of 2012), full text at the National Privacy Commission site.
Readings are linked on the course page alongside this deck.
Read the Fayyad paper after this lecture — it will feel familiar, which is the point.
CRISP-DM, SEMMA, and how real teams structure the discovery loop — the project-management skeleton for everything that follows.
DS 227 · Knowledge Discovery in Data