DS 227 · Graduate · Week 1

Knowledge Discovery in Data

From raw, messy data to valid, novel, and useful knowledge — the frameworks, craft, and ethics of finding what data has to say.

3 units · Blended learning (full online)

Session Map

Three hours, three questions

Today

Where we're going in Week 1

Words for Today · 1 of 2

Six words for talking about data

You will hear these all semester. Each one is explained again on the slide where it first matters, and they are all in the glossary at the end.

data

Recorded facts — numbers, text, dates, clicks — before anyone has interpreted them.

e.g. “7/1: 3 sachets kape” in a store’s notebook

dataset

A collection of data about one topic, usually arranged as a table.

e.g. two years of clinic appointments, saved as one file

record (row)

One entry in a dataset: everything recorded about one thing or event. One row of the table.

e.g. one appointment: patient, date, attended or not

attribute (column)

One property that every record has — one column of the table.

e.g. age, barangay, number of visits

feature

The data-mining word for an attribute, especially one you made to help find a pattern.

e.g. “distance to the clinic”, worked out from the address

missing value

A cell where nothing was recorded. In code it shows up as NaN (“not a number”).

e.g. a patient with no contact number on file

Words for Today · 2 of 2

Seven words for turning data into knowledge

No need to memorise them now. Read them once so none is a surprise later today.

pattern

A regularity that shows up across many records.

e.g. “no-shows rise on clinic days after heavy rain”

knowledge

A pattern that has passed four tests — valid, novel, useful, understandable — so someone can act on it.

e.g. “remind far-away patients when rain is forecast”

KDD

Knowledge Discovery in Data: the whole multi-step process from raw data to knowledge.

e.g. choose the data, clean it, reshape it, search it, judge the result

data mining

The one step inside KDD where a method searches the data for patterns.

e.g. sorting shoppers into groups that buy alike

algorithm

A fixed recipe of steps a computer follows, the same way every time.

e.g. “sort the list, then take the middle value”

model

A simplified summary of the data that describes a pattern or predicts an outcome.

e.g. the rule “lives over 4 km away + rain → likely to miss”

pipeline

A chain of stages where each stage’s output is the next stage’s input.

e.g. raw file → cleaned table → summary → chart

The Motivation

We are rich in data, poor in knowledge

Every jeepney GPS ping, GCash payment, weather station reading, and census form adds to a pile growing far faster than anyone can read it.

The gap

Storage grows exponentially (doubling again and again); human attention doesn't. Whatever is not processed into knowledge is effectively lost.

Picture a store notebook that fills a new page every second. Nobody can read it all, so most of it is never used — unless we have a method for finding what matters in it.

Data collected Data we can study by hand Time → the gap
The Raw Material

Where does all this data come from?

Transactions

Sales, payments, enrollments, bookings — deliberately recorded, usually clean-ish, the classic raw material for data mining.

Sensors & logs

GPS traces, weather stations, server logs, CCTV — recorded as a side effect, enormous volume, rarely designed for analysis.

Human text & media

Posts, reviews, news articles, photos — rich in meaning, messy in structure. Weeks 3–4 teach you to harvest it.

Official statistics

Census, surveys, administrative registers — PSA OpenSTAT (the Philippine Statistics Authority’s free online tables) is our recurring example. Curated, documented, but aggregated: already added up into totals, so the individual records are gone.

Vocabulary

Structured, semi-structured, unstructured

Shape What it looks like Example
Structured Rows × labeled columns; a schema fixed in advance Enrollment table, POS sales, OpenSTAT extracts
Semi-structured Tagged and nested, but flexible — structure travels with the data JSON from an API, HTML pages, XML feeds
Unstructured No inherent fields; meaning must be extracted News articles, complaint emails, photos, audio

New words on this slide

schema
the fixed list of columns (and what kind of value each holds) a table must follow
HTML
the text format web pages are written in, full of tags like <p> (Week 3)
JSON / XML
text formats for nested data, e.g. {"name": "Ana", "visits": [3, 5]}
API
a door a website opens so programs can ask it for data directly (Week 4)
Part 1

What Is Knowledge Discovery?

A definition worth memorizing — one word at a time.

You have met the raw material: records and attributes in a table. This part asks when a pattern in that table deserves the word knowledge.

The Classic Definition

KDD, defined in one sentence (1996)

"The non-trivial process of identifying valid, novel, potentially useful, and ultimately understandable patterns in data."
— Fayyad, Piatetsky-Shapiro & Smyth, "From Data Mining to Knowledge Discovery in Databases," AI Magazine, 1996

Finding knowledge is like panning for gold. The river gives you plenty of sand (patterns); the four underlined words are the sieve that keeps only the gold.

Unpacking the Definition

Four adjectives, four filters

Valid

The pattern holds on new data, not just the sample (the part of the data you happened to study). Guards against overfitting (learning quirks of this sample that will not repeat) and coincidence.

Novel

It tells us something we didn't already know. "Sales rise in December" is valid but not a discovery.

Potentially useful

Someone can act on it — a decision, a policy, a product. Usefulness is judged by the domain, not the algorithm.

Understandable

A human can grasp why it holds — immediately or after some extra explaining. A black-box score (a number from a method whose reasoning nobody can see) alone doesn't qualify.

"Valid" Is the Hard One

Correlation loves to impersonate knowledge

The textbook classic: ice cream sales and drowning deaths rise and fall together, month after month. Strong pattern, real data — and utterly wrong as a causal claim.

The lurking variable

Hot weather drives both. A pattern can be statistically real yet invalid as knowledge if it evaporates once you account for context.

New words

correlation
two things that rise and fall together
causal claim
saying one thing makes the other happen
lurking variable
a hidden third thing driving both; also called a confounder
hold-out testing
checking a pattern on data kept aside and never searched

Validity checks you'll learn here

  • Hold-out testing: does it survive on unseen data? (Week 7+)
  • Confounder hunting: what else moves together? (Week 8)
  • Stability: does it hold across years, regions, subgroups?

Graduate-level habit

Treat every exciting pattern as the start of an interrogation, never the end of one. The four filters are that interrogation.

Data → Wisdom

The DIKW hierarchy

DIKW stands for Data → Information → Knowledge → Wisdom. Each layer adds context and meaning. KDD is the machinery that moves us up the pyramid.

The four layers

Data
raw recorded facts, one entry at a time
Information
data summarised and given context (who, when, how many)
Knowledge
a checked pattern someone can act on
Wisdom
judgment about when to trust and use that knowledge

Where KDD lives

Between data and knowledge — turning recorded facts into patterns we can act on.

Wisdom Knowledge Information Data
A Concrete Walk Up the Pyramid

From a sari-sari store's notebook to a decision

Quick Check · Hour 1

Which filter does this finding fail?

A mall's analytics team reports: "Customers who buy umbrellas are 340% more likely to also buy pancit canton during the same typhoon week. Recommendation: bundle them year-round."

A · Valid — the pattern may not hold outside typhoon weeks
B · Novel — everyone already knew this
C · Useful — no one can act on it
D · Understandable — the mechanism is opaque

Commit to an answer, then click and hold (or tap) to reveal

A — a validity failure.

The pattern is real within typhoon weeks (people stock up on storm supplies together), but the recommendation extrapolates it year-round. Tested on ordinary weeks, it would likely vanish. It's arguably not novel either — but the fatal flaw is validity of the generalization.

Break

  10 minutes

Stretch, hydrate, check nothing. When we return: the five-stage process that turns raw data into defensible knowledge.

Part 2

The KDD Process

Five stages, many loops.

You now know what counts as knowledge. This part adds the process that produces it — the pipeline from a raw file to a finding you can defend.

The Pipeline

Five stages from raw data to knowledge

1 · Selection choose target data 2 · Preprocessing clean & repair 3 · Transformation reshape & reduce 4 · Data Mining extract patterns 5 · Interpretation / Evaluation judge & explain feedback: any stage can send you back Raw data in Knowledge out

KDD is like cooking a meal: pick the ingredients (selection), wash them (preprocessing), chop them (transformation), cook (mining), then taste before serving (interpretation). If it tastes wrong, you go back a step.

Stages 1–2

Selection & preprocessing

Before any analysis: which data, and can we trust it?

Guiding question

"If this field is wrong, would my conclusion change?" If yes, it needs cleaning attention.

New words

domain expert
someone who knows the field the data describes, e.g. a nurse
field
another word for an attribute (column)
scraping
copying data out of web pages with code (Week 3)
duplicate
the same record entered twice
encoding
how text is stored in a file; a mismatch turns “Ñ” into junk

1 · Selection

  • Define the goal with the domain expert first
  • Choose tables, fields, time windows, populations
  • Scraping and acquisition live here (Weeks 3–4)

2 · Preprocessing

  • Handle missing values, duplicates, impossible entries
  • Reconcile units, encodings, spellings ("Q.C." vs "Quezon City")
  • Document every repair — cleaning is analysis (Weeks 5–6)
Stages 3–4

Transformation & mining

Reshape the cleaned data so patterns become findable — then find them.

Mining task menu

Classification · regression · clustering · association rules · anomaly detection · summarization

Lots of new words here. Every one of them is decoded, in plain language, on the next slide.

3 · Transformation

  • Feature engineering: derive age from birthdate, rates from counts
  • Normalize, encode categoricals, reduce dimensionality
  • The same data, reshaped, can reveal or hide a pattern

4 · Data mining

  • Apply the algorithm matched to the goal — not the trendiest one
  • Fit, tune, and check stability across samples
  • Output: candidate patterns, not yet knowledge
Stages 3–4 · Decoded

The words on that slide, in plain language

Transformation is chopping the ingredients into a usable shape; mining is choosing the recipe that fits the dish you were asked for. A fancy recipe cannot rescue badly chopped ingredients.

Stage 5

Interpretation: where patterns face judgment

Worked Example

One problem, all five stages

A barangay health center notices many patients miss their follow-up appointments. The captain asks: "Can the data tell us who is likely to miss, so we can send reminders where they matter?"

Why this example

Small, local, humane — and it exercises every stage, every pitfall, and every ethical question this course covers. We'll revisit it all semester.

What exists

  • Paper logbooks of appointments, partially digitized
  • A patient registry with age, address, contact number
  • SMS records of past reminder attempts

The KDD question hiding inside

Not "predict no-shows" — it's "find valid, novel, useful, understandable patterns in who misses and why, so a human can design a better reminder policy."

Worked Example · Stages 1–2

Selection and preprocessing decisions, made concrete

Worked Example · Stages 3–5

Transform, mine, interpret — and loop

Failure Modes

Where discoveries go wrong, stage by stage

Stage Classic pitfall Defense
Selection Sampling bias — the data you have isn't the population (everyone you want to describe) you claim Ask "who is missing from this table, and why?"
Preprocessing Silent deletion of inconvenient rows Cleaning journal; report row counts before/after
Transformation Leakage — a feature that secretly contains the answer Ask "would I know this before the outcome?"
Mining Overfitting — memorizing noise as pattern Hold-out data; prefer simpler models first
Interpretation Storytelling past the evidence The four filters, applied in writing
Disambiguation

KDD and its overlapping neighbors

Term What it names Relation to KDD
KDD The whole discovery process, end to end The umbrella — this course
Data mining Pattern-extraction algorithms Stage 4 of the KDD process
Machine learning Algorithms that improve from data Supplies many mining tools
Data science The broad discipline & profession KDD is its methodological core
Analytics / BI (business intelligence) Reporting & decision support in orgs Consumes KDD outputs
Quick Check · Hour 2

Name the stage

You compute "days until next typhoon" as a feature for predicting appointment no-shows — using PAGASA's after-the-fact storm records. Which stage are you in, and what pitfall are you committing?

A · Preprocessing — silent deletion
B · Transformation — leakage
C · Mining — overfitting
D · Selection — sampling bias

Commit to an answer, then click and hold (or tap) to reveal

B — transformation-stage leakage.

Feature engineering happens in Stage 3, and "days until next typhoon" uses information nobody could have known at prediction time. The model would score brilliantly in testing and fail in deployment — the signature symptom of leakage.

Break

  10 minutes

Final hour when we return: real applications, a 170-year-old case study, ethics — and what this course asks of you.

Part 3

KDD in the Real World

Who discovers what — and at what risk.

You know the definition and the five stages. This part watches them at work in the real world, where a wrong finding lands on real people.

Applications

The same pipeline, very different stakes

Case Study · 1854

Knowledge discovery before computers

During London's 1854 cholera outbreak, physician John Snow mapped death locations by hand and saw them cluster around the Broad Street water pump — evidence against the reigning "bad air" theory. The pump handle was removed; the outbreak subsided.

Why it's our founding story

Data collection, cleaning (deaths misattributed to other addresses), a visualization, an inference, an action — the whole KDD loop, 170 years early.

Run it through the four filters

  • Valid — clustering held up; later traced to a contaminated well
  • Novel — contradicted expert consensus (miasma theory)
  • Useful — one removed pump handle; lives saved
  • Understandable — a map anyone could read

The modern echo

Every dashboard tracking dengue clusters or traffic fatalities is Snow's map with better plumbing. Weeks 9–10 teach you to build the map and the argument.

A Standing Caution

Discovery is powerful — and not neutral

The same techniques that map poverty can profile individuals. The Philippines' Data Privacy Act of 2012 (RA 10173) sets legal duties for anyone processing personal data (any information that can identify a person) — including researchers.

Questions we'll keep asking

  • Was this data collected with meaningful consent?
  • Can individuals be re-identified (matched back to their names) from our outputs?
  • Who benefits, and who bears the risk, if we're wrong?

Week 11 goes deep

Privacy law, anonymization (removing names and identifiers) and its limits, algorithmic bias (a method that treats some groups unfairly), and responsible scraping — but ethics threads through every week, starting with how we scrape in Week 3.

Activity · 15 Minutes

Breakout: put a "finding" on trial

Part 4

This Course

Twelve weeks along the pipeline.

You know the five stages. This part shows how the twelve weeks walk you through them, one stage at a time — and what the course asks of you.

Roadmap

The 12 weeks at a glance

Weeks Theme Topics
1–2 KDD Foundations This intro · CRISP-DM, SEMMA & process frameworks
3–6 Scraping & Preprocessing HTML parsing · APIs & scraping ethics · cleaning · transformation
7–8 Data Exploration Descriptive statistics · univariate & multivariate EDA · visualization
9–10 Journalism & Storytelling Finding stories in data · narrative, dashboards & communication
11–12 Ethics & Synthesis Privacy & responsible analytics · final presentations

New words in the table

HTML parsing
having code read a web page’s text and pull out the parts you want
descriptive statistics
summary numbers such as the average and how spread out values are
EDA
exploratory data analysis: looking at data from many angles before concluding
univariate / multivariate
one column at a time / several columns together
Logistics

How the course runs

Assessment

How your grade is assembled

Component Weight What it rewards
Weekly exercises (best 8 of 10) 40% Consistent hands-on practice; two free passes for busy weeks
Midterm synthesis (Week 7) 20% Integrating scraping → cleaning → exploration on a fresh dataset
Final discovery project 30% Full KDD cycle on data from your field, defended in Week 12
Participation 10% Quizzes, forum activity, discussion contributions
Before Next Week

Your Week 1 assignment

Small on purpose — the goal is to start seeing your world as data.

1 · Find a dataset

Browse PSA OpenSTAT or any open portal from your field. Pick one dataset that genuinely interests you.

2 · Write half a page

What question might it answer? Which of the four filters (valid, novel, useful, understandable) would be hardest to satisfy, and why?

3 · Set up Python + Jupyter

Install instructions are on the course page. We start writing code in Week 3; the weekly labs need no install at all — they run in your browser.

This Week’s Lab

Your first look at a table, in code

You know the words: dataset, record, attribute, missing value. The lab shows what they look like when Python opens a small clinic table — the first move of Stage 2 (preprocessing). The next two slides read that code aloud, line by line.

Lab · Part 0

Five clinics, typed as text, opened as a table

  • import io, pandas as pdload two libraries (ready-made code): pandas, for tables, and io; as pd gives pandas a short nickname
  • raw = """…"""a variable (a name that holds a value) called raw, holding the table as plain text: commas between columns, one record per line
  • pd.read_csv(io.StringIO(raw))read that text as if it were a CSV file; the result is a DataFrame, pandas’ table, stored in df
  • print(df)show the table underneath the code
  • the output0–4 on the left are row labels, counted from 0. NaN is the missing value: Talamban has no visits recorded. Visits show .0 because a column with a gap is stored as decimals.
import io, pandas as pd raw = """clinic,barangay,visits,staff C-01,Lahug,412,6 C-02,Mabolo,388,5 C-03,Guadalupe,1240,6 C-04,Talamban,,4 C-05,Banilad,455,5""" df = pd.read_csv(io.StringIO(raw)) print(df) clinic barangay visits staff 0 C-01 Lahug 412.0 6 1 C-02 Mabolo 388.0 5 2 C-03 Guadalupe 1240.0 6 3 C-04 Talamban NaN 4 4 C-05 Banilad 455.0 5

A DataFrame is a spreadsheet you control with code: each record is a row, each attribute is a named column.

Lab · Parts 1–2

Three questions to ask any new table

How big is it? What is missing? What is a typical value? Same df as the previous slide.

  • df.shapethe size, as (rows, columns): (5, 4) means 5 clinics, 4 attributes
  • df["visits"]pick one column by its name, in square brackets and quotes
  • .isna().sum().isna() marks each missing cell True; .sum() counts the Trues. 1 = one clinic has no visits number
  • .mean() / .median()the average / the middle value. pandas skips the missing cell, so both use 4 clinics
  • the doteach .name() is a method: a function (a named piece of work) that belongs to a value, called with a dot and brackets
print(df.shape) (5, 4) print(df["visits"].isna().sum()) 1 print(df["visits"].mean(), df["visits"].median()) 623.75 433.5

One busy clinic (1,240 visits) drags the mean up to 623.75; the median, 433.5, stays close to a typical clinic. It is like a millionaire walking into a sari-sari store: the average customer gets richer, the middle customer does not.

In the lab

Replace each ____ with your answer, press ▶ Run to see the output, then Check. Two multiple-choice questions ask what the numbers mean.

Synthesis

Three things to leave with

Glossary · 1 of 2

This week’s words: data and the process

Glossary · 2 of 2

This week’s words: pitfalls and first code

Readings

Where these ideas come from

Required

Fayyad, U., Piatetsky-Shapiro, G., & Smyth, P. (1996). From Data Mining to Knowledge Discovery in Databases. AI Magazine, 17(3). — short, readable, and the source of this week's definition.

Reference (dip in as needed)

Han, Kamber & Pei, Data Mining: Concepts and Techniques — Ch. 1. · Republic Act No. 10173 (Data Privacy Act of 2012), full text at the National Privacy Commission site.

Next Week

KDD Frameworks

CRISP-DM, SEMMA, and how real teams structure the discovery loop — the project-management skeleton for everything that follows.

DS 227 · Knowledge Discovery in Data