Finding a pattern in data is the easy part. Most of this course is about deciding which patterns are real, and being able to say why.
The whole plan is on the next slide. Nothing on it is a surprise later.
Noel Jeffrey Pinton · Department of Computer Science
Each stretch feeds the next. You cannot clean data you have not collected, and you should not tell a story about data you have not looked at.
KDD (Knowledge Discovery in Data, this course’s name): the gap between a pattern and a finding, then the frameworks people argue about — CRISP-DM, SEMMA, and the classic KDD process.
Scraping web pages (copying data out of them with code), pulling from APIs, and the ethics of both. Then the unglamorous half: missing values, outliers (values far from all the others), and merging sources that disagree with each other.
Descriptive statistics and exploratory analysis — one variable at a time, then several at once, where the interesting things usually hide.
How data journalists find a story worth running, and how to build the narrative and dashboards that carry it to people who did not do the analysis.
Privacy, and the harm this work can do when nobody checks. Week 12 you present what you built.
You do not need the other names on this slide yet. Each one (CRISP-DM, API, dashboard…) is explained, in everyday words, in the week it arrives.
Anyone can find a pattern. The hard part is deciding which patterns deserve to be called knowledge — and being able to defend that decision.
Acquire and prepare data from the messy places it actually lives.
Put every finding through four filters: valid, novel, useful, understandable.
Communicate what you found — and what you are not entitled to claim.
Not “understand KDD”. These, specifically — they are what the weekly work and the final presentation actually ask of you.
Scrape a page, call an API, and know where the legal and ethical lines sit before you cross one.
Missing values and outliers force decisions. You should be able to say why you made each one, months later.
Run an analysis that holds up when somebody checks it. In this course somebody checks it.
Turn a finding into something a non-analyst will act on, and be just as clear about what it does not show.
We meet on Saturdays for three contact hours. Everything around that is yours to schedule.
Slides for every week stay here permanently. Nothing important is only said out loud, so missing a session is recoverable.
Work through material when it suits you. Arrive at each session having done the reading — that is what makes the live time worth spending.
Learn the layout once and you can study any week on your own, at your own speed.
Arrow keys or space bar (or the ‹ › buttons at the top). F is full screen. The bar along the bottom shows how far through the deck you are.
A new word is explained on the slide where it first appears. Each week opens with a Words for today slide and ends with a glossary of all the new words.
Pick an answer in your head first, then click and hold (or tap) the question to see the answer and why. Guessing wrong here is free — and it is how the idea sticks.
A dark slide starts a new part (and says how it connects to what you know) or marks a break. Stopping there is part of the plan.
Treat each deck like a textbook chapter you can click through: read, predict, reveal, and pause at the dark slides.
____, press Run, press CheckEach week has a short lab on the course page. The Python in it runs inside your web browser — nothing to install. Python is the programming language we use; code is written instructions the computer carries out exactly.
____ (four underscores) in the code with your answer.The slides show you an idea; the lab asks you to use it. A red error message is normal — read its last line first (it names the problem), change something, and run again.
There are no prerequisites here beyond graduate standing. Some of you write code daily; some have never opened a terminal (a window where you type commands instead of clicking). Both belong in this room.
Knowing what a number means in your field is exactly what stops a coincidence being mistaken for a discovery.
Week 1 has a short lab that mostly asks you to run code someone else wrote and fill in one word at a time. You start writing more of it yourself in Week 3. No prior programming is assumed at any point.
None of this is about being the best programmer in the room. It is mostly about habits.
Three hours goes fast. If everyone arrives having read, we spend the time arguing about hard cases instead of covering definitions.
The people who switch topic in Week 9 are the ones presenting something thin in Week 12.
Half this work is dead ends. Written down they become your methods section. Forgotten, they get repeated.
Being quietly stuck for a week is the most expensive thing you can do in a course that only meets once.
Two small things before we meet, and nothing to install — the Week 1 lab runs in your browser.
Open PSA OpenSTAT (the Philippine Statistics Authority’s free online tables), or any open data portal from your own field.
Choose a dataset (a collection of data on one topic, usually a table) that genuinely interests you. Curiosity beats convenience here.
Write down one question it might answer, and bring it to our first session.
No need to memorise them now — Week 1 explains each one again, with examples. This is just so none of them is a surprise.
Recorded facts — numbers, text, dates — before anyone has interpreted them.
e.g. “7/1: 3 sachets kape” in a store’s notebook
A collection of data about one topic, usually arranged as a table.
e.g. a year of clinic visits, saved as one file
One entry in a dataset: everything recorded about one thing or event.
e.g. one appointment; one sale
One property every record has — a column. Data miners also call it a feature.
e.g. age, barangay, number of visits
A regularity that shows up across many records.
e.g. “coffee sales double before 7 AM on weekdays”
A pattern that has passed four tests: valid, novel, useful, understandable.
e.g. “commuters buy coffee early: open at 6”
Knowledge Discovery in Data: the whole multi-step process from raw data to knowledge.
e.g. choose, clean, reshape, search, judge
The one step inside KDD where a method searches the data for patterns.
e.g. sorting shoppers into groups that buy alike
We open with the awkward question: when does a pattern earn the word knowledge? Then the five stages that take a raw file to something you would defend in front of someone who disagrees with you.
Week 1 slides are already on the course page. Read ahead if you want to.