DS 227 · Graduate · Week 2

KDD Frameworks

CRISP-DM, SEMMA, and the KDD Process — the project skeleton that turns "poke at the data" into work you can defend.

Knowledge Discovery in Data · University of the Philippines Cebu

Session Map

Three names for one idea

Words for Today · 1 of 2

Seven words for talking about process

Each is explained again where it first matters, and all are in the glossary at the end. Last week’s words (dataset, record, pattern, model, pipeline…) still apply.

process model

A named, step-by-step plan for running a data project. Also called a framework.

e.g. CRISP-DM, SEMMA, the KDD process

KDD process

The academic original (Fayyad, 1996): selection, preprocessing, transformation, data mining, interpretation.

e.g. the five-stage pipeline from Week 1

CRISP-DM

CRoss-Industry Standard Process for Data Mining: six phases in a loop, starting from the question.

e.g. the framework this course runs on

SEMMA

Sample, Explore, Modify, Model, Assess: five steps from the software company SAS.

e.g. starts with data already in hand

phase

One stage of a process model, with its own question and its own output.

e.g. Evaluation asks “could this be wrong?”

stakeholder

The person or group whose decision the analysis serves.

e.g. a town councillor deciding on water testing

iterate

Go round again: loop back to an earlier phase with what you just learned.

e.g. Evaluation finds a problem → back to Business Understanding

Words for Today · 2 of 2

CRISP-DM’s six phases, plus two words from the example

One line each for now; Part 2 gives every phase its own slide.

business understanding

Phase 1: whose decision is this for, and what exact question would settle it?

e.g. “did the monthly fail rate rise after May?”

data understanding

Phase 2: meet the data — what each column means, and what looks odd.

e.g. noticing the testing lab changed in May

data preparation

Phase 3: the decisions that turn raw rows into a table you can analyse.

e.g. what to do with a month that has no samples

modelling

Phase 4: choose and compute the summary or model that answers the question.

e.g. the average fail rate per lab

evaluation

Phase 5: does the result answer the question, and how could it be wrong?

e.g. “did the water change, or only the lab?”

deployment

Phase 6: put the result into real use, usually a sentence someone acts on.

e.g. a two-line note to the council, caveat included

rate

A count divided by what it is out of, often × 100 to make a percentage.

e.g. 9 failing out of 38 samples = 23.7%

confound

A third thing that changes together with what you are studying and can fake a pattern.

e.g. a new testing lab, the same month the rate jumped

Where We Left Off

Last week: what KDD is

Knowledge discovery is the non-trivial extraction of valid, novel, useful patterns from data. This week: the process that gets you there without fooling yourself.

Week 1 → Week 2

Definition → discipline. A definition tells you the destination; a framework tells you the order of the roads and where people crash.

Why order matters

Do the steps out of order and you get a confident answer to a question nobody asked.

The Motivation

Without a process, analysis is just vibes

"If you don't know where you're going, any road will get you there — and you won't be able to say why."
The failure mode a framework exists to prevent
Part 1

The KDD Process

Fayyad, Piatetsky-Shapiro & Smyth, 1996 — where the field's vocabulary comes from.

You met these five stages last week. This part revisits them as a process model and shows why data mining is only one of them.

A Common Confusion

Data mining is one step, not the whole thing

People say "data mining" to mean the entire project. In the KDD paper it is a single stage — the algorithm run — surrounded by far more work.

KDD = the whole process

Selecting, cleaning, transforming, mining, and interpreting — end to end.

Data mining = one stage inside it

The pattern-finding step. Powerful, but useless if the four stages before it were sloppy.

The Pipeline

Five stages, four transformations

Raw Data Selection Target data Pre- processing Transform Features Data Mining Interpret / Evaluate → Knowledge Each arrow discards or reshapes data — and each is a decision you own.
Stage by Stage

What actually happens at each step

StageYou start withYou end with
SelectionEverything recordedThe target subset worth studying
PreprocessingTarget data, warts and allCleaned data: missing values, noise handled
TransformationClean rowsFeatures — reduced, encoded, scaled
Data miningFeatures + a goalPatterns: rules, clusters, a model
InterpretationRaw patternsKnowledge you'd act on

New words in the table

noise
random errors and oddities that hide the real signal
encoded / scaled
categories turned into numbers / numbers rescaled to a common range
rules / clusters
“if this, then that” statements / groups of similar records
Untangling The Buzzwords

Data mining, statistics, machine learning — not synonyms

Where This Actually Lives

Three discoveries you've felt this week

Quick Check

Tap to reveal

"We used data mining to find the pattern." Which KDD stages does that sentence quietly skip?

A · Selection & preprocessing
B · Transformation
C · Interpretation
D · All of the above

D — all of them.

"We ran the algorithm" hides the four stages that decided what the algorithm even saw. That is where most errors live.

Part 2

CRISP-DM

The Cross-Industry Standard Process for Data Mining — the one you will actually run this week.

You know the academic five-stage pipeline. CRISP-DM adds what it leaves out: a first phase for the question and a last phase for using the answer.

Where It Came From

Built by practitioners, not a lab

CRISP-DM was written in the late 1990s by companies drowning in one-off analytics projects. They wanted a shared, tool-agnostic way to run them. Decades on, it is still the most-used framework in industry.

Cross-industry

Meant to work for a bank, a hospital, or a water utility alike.

Tool-agnostic

Says nothing about Python vs R (another programming language) vs a spreadsheet. It structures the thinking.

Still dominant

Surveys repeatedly rank it the most-followed data-science process.

The Cycle

Six phases, and it loops

Business Understanding Data Understanding Data Preparation Modelling Evaluation Deploy- ment

CRISP-DM is a recipe cycle: ask who you are cooking for (business), check the pantry (data), wash and chop (preparation), cook (modelling), taste (evaluation), serve (deployment). A bad taste sends you back to rethink the menu, not just add salt.

Phase 1 of 6

Business Understanding

Before a single number: whose decision does this serve, and what would change the answer? The most skipped phase, and the most expensive to skip. “Business” just means whoever the work is for: a town, a clinic, a school.

Ask the party’s host what they want to eat before you go shopping. Cooking a perfect dish nobody ordered is still a failure.

Turn the ask into a question the data can settle

"Is our water getting worse?" is a feeling. "Did the monthly fail rate rise after May?" is checkable.

Define success up front

Decide what result would count as an answer before you see it — or you will rationalise whatever you find.

Phase 2 of 6

Data Understanding

Meet the data before you trust it. Look at every column, ask what each one means, and hunt for the surprises that break naïve conclusions.

Check what is really in the fridge — and whether the labels are right — before you plan the meal around it.

Read the table, slowly

A jump from 6% to 25% failures has a boring explanation and an alarming one. Both start as hypotheses.

Find what quietly changed

If the measuring lab switched the same month the numbers jumped, the numbers may not be comparable at all.

Phase 3 of 6 · Where the time goes

Data Preparation is decisions, not tidying

"Preparation is not cleaning until it looks nice. It is a series of decisions, each of which changes the answer — and each of which you should be able to defend."
This week's lab, on the phase that eats 60–80% of a real project

Washing and chopping: slow, unglamorous, and it decides how the dish tastes more than the cooking does.

Phase 4 of 6

Modelling

"Modelling" need not mean machine learning. It means pick the summary that answers the question — sometimes a grouped average (one average per group, e.g. per lab) is the whole model.

Pick the cooking method the dish needs. Sometimes boiling beats an elaborate recipe — and you can explain exactly what you did.

Match the method to the question

Grouping fail rates by lab answers "do the labs differ?" — a different question from "did the water change?"

Simpler is a feature

A summary you can explain beats a model you can't. Reach for complexity only when the question demands it.

Phase 5 of 6

Evaluation

Does the result actually answer the Phase-1 question, and how could it be wrong? Evaluation is where you attack your own conclusion before anyone else does.

Taste before it leaves the kitchen — and ask whether the odd flavour comes from the soup or from the spoon you tasted it with.

What would make you confident?

Name one specific thing you'd need to see to believe the water genuinely worsened.

What would change your mind?

If a stricter threshold (the cut-off level that counts as a fail) at the new lab explains the jump, the water may not have changed at all.

Phase 6 of 6

Deployment

Usually a sentence someone acts on, not a model in production (running automatically inside a real system). The discovery only matters once it reaches the person making the decision.

Serve the dish. It only counts once someone eats it — and you tell them what is in it.

Say what the data shows

And, just as clearly, what you would not yet claim.

Earn your confidence

No hedging so heavy it says nothing; no certainty you haven't earned.

At a Glance

The six phases, one line each

1

Business

Whose decision? What question?

2

Data

Meet it; find surprises.

3

Preparation

Decisions that shape the answer.

4

Modelling

The summary that answers it.

5

Evaluation

Could this be wrong?

6

Deployment

A sentence someone acts on.

CRISP-DM in numbers

Let's walk the water table through the phases

MonthSamplesFailsFail rateLab
April4025.0%A
May38923.7%B
June411024.4%B
July39923.1%B
August00—B
Phase 1 · applied

The question, made checkable

"Is the water worse?" is a feeling. Business Understanding turns it into a test with a yes-or-no answer.

The precise question

"Did the mean monthly fail rate rise after April?" — a number we can compute and defend.

Success, decided up front

A rise from ~5% to ~24% would count as real only if nothing else about the measurement changed.

Phase 3 · applied

fail_rate is a decision, not a division

Phase 4 · applied

Group by lab — a summary is a model

No machine learning here. Just the right grouped average, which quietly answers a different question than the one we asked.

LabMonthsMean fail rate
AApril5.0%
BMay–July23.7%

Look what surfaced

The jump lines up perfectly with the lab switching from A to B. Suddenly "the water got worse" has a rival explanation.

Phase 5 · applied

Attack your own conclusion

"The failure rate quadrupled the same month a new lab took over the testing. Before blaming the water, rule out the ruler."
Evaluation is where the obvious story goes to be cross-examined
Phase 6 · applied

State it honestly, then ship it

Deployment isn't a server. Here it's the sentence you hand the town — claim, number, and the caveat that keeps it true.

The defensible finding

"Measured fail rates rose from ~5% to ~24% after April — but a lab change in May confounds it, so this is a flag to investigate, not proof the water worsened."

Why the caveat is the point

Drop it and you've raised a false alarm — or missed a real one. The honesty is the deliverable.

The Loop, Concretely

Evaluation just sent us back to the start

5

Evaluation found a confound

The lab switch could explain everything.

1

Back to Business Understanding

New question: "Does the water differ, holding the lab fixed?"

2

Back to Data

Go get comparable months — same lab, before and after.

Break

  Five minutes

Back to compare CRISP-DM with the tool that competes with it.

Part 3

SEMMA

SAS's five-step process — narrower on purpose, and worth knowing why.

You have walked CRISP-DM’s six phases. This part adds a third framework and lines all three up side by side, so you can see what each one leaves out.

The Acronym

Sample · Explore · Modify · Model · Assess

S

Sample

Take a workable slice of the data.

E

Explore

Visualise, spot patterns and oddities.

M

Modify

Clean, transform, engineer features.

M

Model

Fit the pattern-finding method.

A

Assess

Judge how well it holds up.

SEMMA is a cooking-show round: the ingredients are already on the counter and it ends when the judges taste. Nobody asks who the meal is for, and nobody serves it.

The Missing Bookends

SEMMA skips the why and the what-next

It is deliberately scoped to the technical core. That makes it tidy — and dangerous if you forget what it leaves out.

No business understanding

SEMMA assumes someone already framed the problem. CRISP-DM makes that Phase 1.

No deployment

It ends at "Assess." Turning the result into action lives outside the acronym.

Side by Side

Three frameworks, aligned

KDD ProcessCRISP-DMSEMMA
—Business Understanding— (missing)
SelectionData UnderstandingSample + Explore
Preprocessing + TransformationData PreparationModify
Data MiningModellingModel
Interpretation / EvaluationEvaluationAssess
Use of knowledgeDeployment— (missing)
Read The Gaps

SEMMA is missing both bookends

It opens at Sample — you already have data — and closes at Assess. No "why", no "now what". Powerful, and quietly dangerous.

No Business Understanding

Start at Sample and you can model beautifully — the wrong question, flawlessly.

No Deployment

End at Assess and a great result dies in a notebook nobody acts on.

Why it still earns its place

For the modelling middle, SEMMA's five steps are tight and repeatable — just don't mistake the middle for the whole.

Choosing In Practice

Which one do you actually reach for?

SituationReach forBecause
A stakeholder with a fuzzy askCRISP-DMIt forces the "why" first
Data already in hand, iterating fastSEMMATight modelling loop
Explaining the field to a newcomerKDD ProcessNames every transformation
Quick Check

Tap to reveal

You drop the August row because it has zero samples. Which CRISP-DM phase are you in?

A · Business Understanding
B · Data Preparation
C · Modelling
D · Deployment

B — Data Preparation.

Deciding what to do with a broken row — drop, fill, or flag — is preparation, and it's a judgment call you must be able to defend, not a tidy-up.

So Which One?

They're complementary, not rivals

You don't "pick a winner." Use CRISP-DM to structure the project; borrow SEMMA's crisp verbs for the technical middle; keep KDD's vocabulary for talking about stages precisely.

Default to CRISP-DM

It's the most complete and the most widely understood. This course runs on it.

The framework is a checklist, not a cage

Its value is catching the step you were about to skip — usually Phase 1 or Phase 5.

The Uncomfortable Truth

Most of the work isn't modelling

Understanding + Preparation Modelling Eval + Deploy ~60–80% of the effort the glamorous bit the payoff Effort across a typical KDD project
The Classic Mistake

Jumping straight to modelling

The seductive move: open the notebook (a document of text and runnable code), load the CSV (a plain-text table file), fit something (let a model learn from the data). It feels like progress and produces a confident answer to a question nobody asked.

Symptom

A polished result you can't tie back to any decision anyone needs to make.

Cure

Spend the first hour in Phases 1 and 2 with no code. The framework is nagging you to.

Quick Check

Tap to reveal

A team reports "fail rate rose from 6% to 25% after May." What CRISP-DM phase most needs revisiting before they publish?

A · Modelling — try a fancier method
B · Data Understanding — the lab changed in May
C · Deployment — write it up faster
D · Nothing — the numbers speak

B — Data Understanding.

If the measuring lab switched the same month, the before and after numbers may not be comparable. Loop back before claiming anything.

This Week's Lab

Walking a dataset through CRISP-DM

You'll take one small water-quality dataset through all six phases. Almost no code — the thinking is the work. ~45 minutes.

What you'll wrestle with

A month with zero samples, a lab that switched mid-year, and three different "averages" that each tell a different story.

The deliverable

Two honest sentences for a councillor: what the data shows, and what you won't yet claim.

The Lab’s Code, Read Aloud

Two decisions, in pandas

In the Week 1 lab you opened a table and counted its gaps. This week’s lab adds a new column (the fail rate) and makes Data Preparation choices about a month with no data. The next two slides read that code line by line.

Lab · Parts 0–1

A fail-rate column, and a month with nothing in it

  • pd.read_csv(io.StringIO(raw))read the ten-month table (typed as text in the variable raw) into a DataFrame, pandas’ table, called water
  • water["samples_failing"] / water["samples_taken"] * 100divide one column by another, row by row, then × 100: the fail rate as a percentage
  • water["fail_rate"] = (…).round(1)store the answer as a new column named fail_rate, rounded to 1 decimal
  • water.iloc[3:8]show rows by position: 3 up to, not including, 8 (April to August; positions count from 0)
  • the outputthe rate jumps from 6.7 to 25.4 exactly when the lab becomes Regional. August: 0 ÷ 0 has no answer, so pandas writes NaN (missing), not 0
import io, pandas as pd # raw holds the ten-month table as text (see the lab) water = pd.read_csv(io.StringIO(raw)) water["fail_rate"] = (water["samples_failing"] / water["samples_taken"] * 100).round(1) print(water.iloc[3:8]) month samples_taken samples_failing lab fail_rate 3 April 119 8 Municipal 6.7 4 May 122 31 Regional 25.4 5 June 120 29 Regional 24.2 6 July 117 33 Regional 28.2 7 August 0 0 Regional NaN

One line of pandas works on a whole column at once, like filling a formula down a spreadsheet column.

Lab · Part 2

The default is a choice

Two ways to average the same fail_rate column.

  • water["fail_rate"].mean().mean() is a method (a function that belongs to a value, after a dot). pandas skips NaN by default, so it averages the 9 months that have data
  • .fillna(0)first fill every missing cell with 0 — pretending August was a perfect month
  • round(…, 1)a function (a named piece of work) that rounds the number in its brackets to 1 decimal
  • the output16.8 vs 15.1: the invented perfect month makes the water look 1.7 points better than any real month showed
print(round(water["fail_rate"].mean(), 1)) 16.8 print(round(water["fail_rate"].fillna(0).mean(), 1)) 15.1

Leaving an unmarked exam out of the class average is honest; writing 0 on it changes the average and records something that never happened.

In the lab

You write the guard yourself (keep only months where samples were taken), so the choice is visible in your code instead of hidden in a default. Fill each ____, press ▶ Run, then Check.

Recap

Four things to carry out

Glossary

This week’s words, one line each

Readings

Before next week

Next Week

Data Scraping I

Web fundamentals and parsing HTML — how the raw material actually gets onto your machine before any of this process can start.

DS 227 · Knowledge Discovery in Data