CRISP-DM, SEMMA, and the KDD Process — the project skeleton that turns "poke at the data" into work you can defend.
Knowledge Discovery in Data · University of the Philippines Cebu
The original nine-step academic pipeline (Fayyad, 1996).
The six-phase industry standard you will use in the lab.
SAS's five-step tool-centred alternative, and how they compare.
All three describe the same journey: raw data on one end, a defensible decision on the other.
A checklist you can hold any messy analysis against — including your own.
Each is explained again where it first matters, and all are in the glossary at the end. Last week’s words (dataset, record, pattern, model, pipeline…) still apply.
A named, step-by-step plan for running a data project. Also called a framework.
e.g. CRISP-DM, SEMMA, the KDD process
The academic original (Fayyad, 1996): selection, preprocessing, transformation, data mining, interpretation.
e.g. the five-stage pipeline from Week 1
CRoss-Industry Standard Process for Data Mining: six phases in a loop, starting from the question.
e.g. the framework this course runs on
Sample, Explore, Modify, Model, Assess: five steps from the software company SAS.
e.g. starts with data already in hand
One stage of a process model, with its own question and its own output.
e.g. Evaluation asks “could this be wrong?”
The person or group whose decision the analysis serves.
e.g. a town councillor deciding on water testing
Go round again: loop back to an earlier phase with what you just learned.
e.g. Evaluation finds a problem → back to Business Understanding
One line each for now; Part 2 gives every phase its own slide.
Phase 1: whose decision is this for, and what exact question would settle it?
e.g. “did the monthly fail rate rise after May?”
Phase 2: meet the data — what each column means, and what looks odd.
e.g. noticing the testing lab changed in May
Phase 3: the decisions that turn raw rows into a table you can analyse.
e.g. what to do with a month that has no samples
Phase 4: choose and compute the summary or model that answers the question.
e.g. the average fail rate per lab
Phase 5: does the result answer the question, and how could it be wrong?
e.g. “did the water change, or only the lab?”
Phase 6: put the result into real use, usually a sentence someone acts on.
e.g. a two-line note to the council, caveat included
A count divided by what it is out of, often × 100 to make a percentage.
e.g. 9 failing out of 38 samples = 23.7%
A third thing that changes together with what you are studying and can fake a pattern.
e.g. a new testing lab, the same month the rate jumped
Knowledge discovery is the non-trivial extraction of valid, novel, useful patterns from data. This week: the process that gets you there without fooling yourself.
Definition → discipline. A definition tells you the destination; a framework tells you the order of the roads and where people crash.
Do the steps out of order and you get a confident answer to a question nobody asked.
"If you don't know where you're going, any road will get you there — and you won't be able to say why."The failure mode a framework exists to prevent
A named process lets a teammate — or you in six months — retrace the steps.
Each stage produces an artefact (something you can hand over: a table, a chart, a note) someone can check before the next stage builds on it.
When a stakeholder (whoever the analysis is for) pushes back, you can point to where a choice was made.
Fayyad, Piatetsky-Shapiro & Smyth, 1996 — where the field's vocabulary comes from.
You met these five stages last week. This part revisits them as a process model and shows why data mining is only one of them.
People say "data mining" to mean the entire project. In the KDD paper it is a single stage — the algorithm run — surrounded by far more work.
Selecting, cleaning, transforming, mining, and interpreting — end to end.
The pattern-finding step. Powerful, but useless if the four stages before it were sloppy.
The data gets smaller and cleaner left to right. What you throw away is as consequential as what you keep.
The loop back — you rarely go straight through — is the part beginners forget.
| Stage | You start with | You end with |
|---|---|---|
| Selection | Everything recorded | The target subset worth studying |
| Preprocessing | Target data, warts and all | Cleaned data: missing values, noise handled |
| Transformation | Clean rows | Features — reduced, encoded, scaled |
| Data mining | Features + a goal | Patterns: rules, clusters, a model |
| Interpretation | Raw patterns | Knowledge you'd act on |
Notice how little of this is "run the algorithm." Four of five stages are preparation and sense-making.
A pattern is not knowledge until a human decides it is valid and useful.
Starts with a hypothesis (a specific guess you can test) and tests it. "Is this difference real, or luck?"
Machine learning: methods that learn a rule from examples to predict the next case. One tool in the mining step.
The whole hunt for patterns worth knowing — using both of the above, plus everything around them.
KDD is the umbrella; statistics and ML are two tools under it. Confusing the tool for the process is how projects skip straight to modelling.
If someone says "just run a model," they've named a step — not a plan.
A charge two cities apart, minutes apart, gets held. A pattern mined from millions of normal ones.
Vitals drifting toward a known danger shape page a nurse before the crisis.
"Because you watched…" is a pattern discovered across everyone's history, aimed at you.
None of these is "a model." Each is a full KDD pipeline — a question, messy data, preparation, a method, and a decision at the end.
Raw records in, a defensible decision out. That shape is the whole course.
"We used data mining to find the pattern." Which KDD stages does that sentence quietly skip?
D — all of them.
"We ran the algorithm" hides the four stages that decided what the algorithm even saw. That is where most errors live.
Press and hold (or focus) the box to reveal the answer.
When someone says "data mining," ask what happened before the algorithm ran.
The Cross-Industry Standard Process for Data Mining — the one you will actually run this week.
You know the academic five-stage pipeline. CRISP-DM adds what it leaves out: a first phase for the question and a last phase for using the answer.
CRISP-DM was written in the late 1990s by companies drowning in one-off analytics projects. They wanted a shared, tool-agnostic way to run them. Decades on, it is still the most-used framework in industry.
Meant to work for a bank, a hospital, or a water utility alike.
Says nothing about Python vs R (another programming language) vs a spreadsheet. It structures the thinking.
Surveys repeatedly rank it the most-followed data-science process.
The outer ring is the project. The arrow from Evaluation back to Business Understanding is the point: one pass is rarely enough.
Because the whole loop exists to change what someone does — not to produce a notebook.
CRISP-DM is a recipe cycle: ask who you are cooking for (business), check the pantry (data), wash and chop (preparation), cook (modelling), taste (evaluation), serve (deployment). A bad taste sends you back to rethink the menu, not just add salt.
Before a single number: whose decision does this serve, and what would change the answer? The most skipped phase, and the most expensive to skip. “Business” just means whoever the work is for: a town, a clinic, a school.
Ask the party’s host what they want to eat before you go shopping. Cooking a perfect dish nobody ordered is still a failure.
"Is our water getting worse?" is a feeling. "Did the monthly fail rate rise after May?" is checkable.
Decide what result would count as an answer before you see it — or you will rationalise whatever you find.
Meet the data before you trust it. Look at every column, ask what each one means, and hunt for the surprises that break naïve conclusions.
Check what is really in the fridge — and whether the labels are right — before you plan the meal around it.
A jump from 6% to 25% failures has a boring explanation and an alarming one. Both start as hypotheses.
If the measuring lab switched the same month the numbers jumped, the numbers may not be comparable at all.
"Preparation is not cleaning until it looks nice. It is a series of decisions, each of which changes the answer — and each of which you should be able to defend."This week's lab, on the phase that eats 60–80% of a real project
Drop, fill, or flag? A month with zero samples is not a month with zero failures.
Letting a library (ready-made code, like pandas) decide how to
handle 0/0 (zero divided by zero, which has no answer) is still you
choosing — just silently.
Averaging ten months across two different labs may answer no real question at all.
Washing and chopping: slow, unglamorous, and it decides how the dish tastes more than the cooking does.
"Modelling" need not mean machine learning. It means pick the summary that answers the question — sometimes a grouped average (one average per group, e.g. per lab) is the whole model.
Pick the cooking method the dish needs. Sometimes boiling beats an elaborate recipe — and you can explain exactly what you did.
Grouping fail rates by lab answers "do the labs differ?" — a different question from "did the water change?"
A summary you can explain beats a model you can't. Reach for complexity only when the question demands it.
Does the result actually answer the Phase-1 question, and how could it be wrong? Evaluation is where you attack your own conclusion before anyone else does.
Taste before it leaves the kitchen — and ask whether the odd flavour comes from the soup or from the spoon you tasted it with.
Name one specific thing you'd need to see to believe the water genuinely worsened.
If a stricter threshold (the cut-off level that counts as a fail) at the new lab explains the jump, the water may not have changed at all.
Usually a sentence someone acts on, not a model in production (running automatically inside a real system). The discovery only matters once it reaches the person making the decision.
Serve the dish. It only counts once someone eats it — and you tell them what is in it.
And, just as clearly, what you would not yet claim.
No hedging so heavy it says nothing; no certainty you haven't earned.
Whose decision? What question?
Meet it; find surprises.
Decisions that shape the answer.
The summary that answers it.
Could this be wrong?
A sentence someone acts on.
Screenshot this slide. It is the structure of this week's lab.
Evaluation often sends you back to Phase 1 with a sharper question.
| Month | Samples | Fails | Fail rate | Lab |
|---|---|---|---|---|
| April | 40 | 2 | 5.0% | A |
| May | 38 | 9 | 23.7% | B |
| June | 41 | 10 | 24.4% | B |
| July | 39 | 9 | 23.1% | B |
| August | 0 | 0 | — | B |
Fail rate = fails ÷ samples × 100: the percentage of samples that failed. The story looks obvious — failures tripled after April. The phases are how we check whether that story is real. (The lab uses a longer, ten-month version of this table with the same traps.)
Zero samples. Not zero failures — no data. That distinction decides the whole analysis.
"Is the water worse?" is a feeling. Business Understanding turns it into a test with a yes-or-no answer.
"Did the mean monthly fail rate rise after April?" — a number we can compute and defend.
A rise from ~5% to ~24% would count as real only if nothing else about the measurement changed.
fail_rate = fails / samples. Fine for April–July. For August it is
0 / 0 — undefined: there is no number that answers “0 out of 0 is
what percent?”
Call August 0% and you invent a perfect month. Drop it and you admit you
simply don't know. Only one is honest.
This one line — keep only rows where samples > 0 — is the
whole difference between an honest table and a flattering one.
You'll write exactly this guard (a check that stops bad rows getting through), then explain why the default answer would mislead.
No machine learning here. Just the right grouped average, which quietly answers a different question than the one we asked.
| Lab | Months | Mean fail rate |
|---|---|---|
| A | April | 5.0% |
| B | May–July | 23.7% |
The jump lines up perfectly with the lab switching from A to B. Suddenly "the water got worse" has a rival explanation.
"The failure rate quadrupled the same month a new lab took over the testing. Before blaming the water, rule out the ruler."Evaluation is where the obvious story goes to be cross-examined
A confound is a third thing that moves with both. Here it's the lab change — change the measurer and the measurement can shift with nothing real behind it.
To trust the trend, you'd need months from Lab B before May, or Lab A after. The table gives you neither.
Deployment isn't a server. Here it's the sentence you hand the town — claim, number, and the caveat that keeps it true.
"Measured fail rates rose from ~5% to ~24% after April — but a lab change in May confounds it, so this is a flag to investigate, not proof the water worsened."
Drop it and you've raised a false alarm — or missed a real one. The honesty is the deliverable.
The lab switch could explain everything.
New question: "Does the water differ, holding the lab fixed?"
Go get comparable months — same lab, before and after.
This is why CRISP-DM is drawn as a circle. Finishing a lap doesn't mean you're done — it means you finally know the right question to ask.
A good evaluation usually creates work, not applause. That's the process working.
Back to compare CRISP-DM with the tool that competes with it.
SAS's five-step process — narrower on purpose, and worth knowing why.
You have walked CRISP-DM’s six phases. This part adds a third framework and lines all three up side by side, so you can see what each one leaves out.
Take a workable slice of the data.
Visualise, spot patterns and oddities.
Clean, transform, engineer features.
Fit the pattern-finding method.
Judge how well it holds up.
Born inside SAS Enterprise Miner (a commercial data-mining program from the company SAS) — it maps neatly onto the software's analysis nodes (the boxes you connect on its point-and-click screen).
It starts at Sample, not at a question. That omission is the whole debate.
SEMMA is a cooking-show round: the ingredients are already on the counter and it ends when the judges taste. Nobody asks who the meal is for, and nobody serves it.
It is deliberately scoped to the technical core. That makes it tidy — and dangerous if you forget what it leaves out.
SEMMA assumes someone already framed the problem. CRISP-DM makes that Phase 1.
It ends at "Assess." Turning the result into action lives outside the acronym.
| KDD Process | CRISP-DM | SEMMA |
|---|---|---|
| — | Business Understanding | — (missing) |
| Selection | Data Understanding | Sample + Explore |
| Preprocessing + Transformation | Data Preparation | Modify |
| Data Mining | Modelling | Model |
| Interpretation / Evaluation | Evaluation | Assess |
| Use of knowledge | Deployment | — (missing) |
The middle rows line up almost perfectly. The frameworks agree on the technical work.
CRISP-DM wraps the core in purpose (top) and action (bottom). That's why it endures.
It opens at Sample — you already have data — and closes at Assess. No "why", no "now what". Powerful, and quietly dangerous.
Start at Sample and you can model beautifully — the wrong question, flawlessly.
End at Assess and a great result dies in a notebook nobody acts on.
For the modelling middle, SEMMA's five steps are tight and repeatable — just don't mistake the middle for the whole.
| Situation | Reach for | Because |
|---|---|---|
| A stakeholder with a fuzzy ask | CRISP-DM | It forces the "why" first |
| Data already in hand, iterating fast | SEMMA | Tight modelling loop |
| Explaining the field to a newcomer | KDD Process | Names every transformation |
They overlap far more than they differ. The framework is scaffolding — pick the one that stops you skipping the step you're tempted to skip.
Most working teams use a personal blend and never say a framework's name out loud.
You drop the August row because it has zero samples. Which CRISP-DM phase are you in?
B — Data Preparation.
Deciding what to do with a broken row — drop, fill, or flag — is preparation, and it's a judgment call you must be able to defend, not a tidy-up.
Naming the phase you're in keeps you honest about what kind of decision you're making.
Every choice about a row or column is Preparation — and every one changes the answer.
You don't "pick a winner." Use CRISP-DM to structure the project; borrow SEMMA's crisp verbs for the technical middle; keep KDD's vocabulary for talking about stages precisely.
It's the most complete and the most widely understood. This course runs on it.
Its value is catching the step you were about to skip — usually Phase 1 or Phase 5.
The famous "80% of data science is cleaning data" is not a complaint — it's the phase where the answer is actually decided.
Judge a project by its preparation choices, not by which algorithm it used.
The seductive move: open the notebook (a document of text and runnable code), load the CSV (a plain-text table file), fit something (let a model learn from the data). It feels like progress and produces a confident answer to a question nobody asked.
A polished result you can't tie back to any decision anyone needs to make.
Spend the first hour in Phases 1 and 2 with no code. The framework is nagging you to.
A team reports "fail rate rose from 6% to 25% after May." What CRISP-DM phase most needs revisiting before they publish?
B — Data Understanding.
If the measuring lab switched the same month, the before and after numbers may not be comparable. Loop back before claiming anything.
This is exactly the trap the water dataset sets in this week's lab.
When a number jumps, ask what else changed at the same time.
You'll take one small water-quality dataset through all six phases. Almost no code — the thinking is the work. ~45 minutes.
A month with zero samples, a lab that switched mid-year, and three different "averages" that each tell a different story.
Two honest sentences for a councillor: what the data shows, and what you won't yet claim.
In the Week 1 lab you opened a table and counted its gaps. This week’s lab adds a new column (the fail rate) and makes Data Preparation choices about a month with no data. The next two slides read that code line by line.
pd.read_csv(io.StringIO(raw))read the ten-month table (typed as text in the variable raw) into a DataFrame, pandas’ table, called waterwater["samples_failing"] / water["samples_taken"] * 100divide one column by another, row by row, then × 100: the fail rate as a percentagewater["fail_rate"] = (…).round(1)store the answer as a new column named fail_rate, rounded to 1 decimalwater.iloc[3:8]show rows by position: 3 up to, not including, 8 (April to August; positions count from 0)NaN (missing), not 0One line of pandas works on a whole column at once, like filling a formula down a spreadsheet column.
Two ways to average the same fail_rate column.
water["fail_rate"].mean().mean() is a method (a function that belongs to a value, after a dot). pandas skips NaN by default, so it averages the 9 months that have data.fillna(0)first fill every missing cell with 0 — pretending August was a perfect monthround(…, 1)a function (a named piece of work) that rounds the number in its brackets to 1 decimalLeaving an unmarked exam out of the class average is honest; writing 0 on it changes the average and records something that never happened.
You write the guard yourself (keep only months where samples were taken), so the choice is
visible in your code instead of hidden in a default. Fill each ____, press
▶ Run, then Check.
Repeatable, reviewable, and traceable to the decision that was made where.
Six phases, and it loops. Business Understanding first; Deployment is the point.
60–80% of the work, and every choice changes the result.
Same technical core; they differ on whether they include the why and the what-next.
If you remember one sentence: the framework's job is to stop you skipping the step that decides the answer.
Phase 1 (framing) or Phase 5 (could I be wrong?).
.mean()Wirth & Hipp (2000), CRISP-DM: Towards a Standard Process Model for Data Mining — the paper that named the six phases.
Fayyad, Piatetsky-Shapiro & Smyth (1996), §2 — the KDD pipeline you saw in Part 1.
Both are linked on the course page beside this deck and the lab.
The reading lands better once you've felt the six phases on real (invented) data.
Web fundamentals and parsing HTML — how the raw material actually gets onto your machine before any of this process can start.
DS 227 · Knowledge Discovery in Data